Screaming in the Cloud: Recent Episodes

Corey Quinn

Screaming in the Cloud with Corey Quinn features conversations with domain experts in the world of Cloud Computing. Topics discussed include AWS, GCP, Azure, Oracle Cloud, and the "why" behind how businesses are coming to think about the Cloud.

View Details

Tony Baer, Principal at dbInsight, joins Corey on Screaming in the Cloud to discuss his definition of what is and isn’t a database, and the trends he’s seeing in the industry. Tony explains why it’s important to try and have an outsider’s perspective when evaluating new ideas, and the growing awareness of the impact data has on our daily lives. Corey and Tony discuss the importance of working towards true operational simplicity in the cloud, and Tony also shares why explainability in generative AI is so crucial as the technology advances.

About Tony

Tony Baer, the founder and CEO of dbInsight, is a recognized industry expert in extending data management practices, governance, and advanced analytics to address the desire of enterprises to generate meaningful value from data-driven transformation. His combined expertise in both legacy database technologies and emerging cloud and analytics technologies shapes how clients go to market in an industry undergoing significant transformation.

During his 10 years as a principal analyst at Ovum, he established successful research practices in the firm’s fastest growing categories, including big data, cloud data management, and product lifecycle management. He advised Ovum clients regarding product roadmap, positioning, and messaging and helped them understand how to evolve data management and analytic strategies as the cloud, big data, and AI moved the goal posts. Baer was one of Ovum’s most heavily-billed analysts and provided strategic counsel to enterprises spanning the Fortune 100 to fast-growing privately held companies.

With the cloud transforming the competitive landscape for database and analytics providers, Baer led deep dive research on the data platform portfolios of AWS, Microsoft Azure, and Google Cloud, and on how cloud transformation changed the roadmaps for incumbents such as Oracle, IBM, SAP, and Teradata. While at Ovum, he originated the term “Fast Data” which has since become synonymous with real-time streaming analytics.

Baer’s thought leadership and broad market influence in big data and analytics has been formally recognized on numerous occasions. Analytics Insight named him one of the 2019 Top 100 Artificial Intelligence and Big Data Influencers. Previous citations include Onalytica, which named Baer as one of the world’s Top 20 thought leaders and influencers on Data Science; Analytics Week, which named him as one of 200 top thought leaders in Big Data and Analytics; and by KDnuggets, which listed Baer as one of the Top 12 top data analytics thought leaders on Twitter. While at Ovum, Baer was Ovum’s IT’s most visible and publicly quoted analyst, and was cited by Ovum’s parent company Informa as Brand Ambassador in 2017. In raw numbers, Baer has 14,000 followers on Twitter, and his ZDnet “Big on Data” posts are read 20,000 – 30,000 times monthly. He is also a frequent speaker at industry conferences such as Strata Data and Spark Summit.

Links Referenced:

  • dbInsight: https://dbinsight.io/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is brought to us in part by our friends at RedHat.As your organization grows, so does the complexity of your IT resources. You need a flexible solution that lets you deploy, manage, and scale workloads throughout your entire ecosystem. The Red Hat Ansible Automation Platform simplifies the management of applications and services across your hybrid infrastructure with one platform. Look for it on the AWS Marketplace.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Back in my early formative years, I was an SRE sysadmin type, and one of the areas I always avoided was databases, or frankly, anything stateful because I am clumsy and unlucky and that’s a bad combination to bring within spitting distance of anything that, you know, can’t be spun back up intact, like databases. So, as a result, I tend not to spend a lot of time historically living in that world. It’s time to expand horizons and think about this a little bit differently. My guest today is Tony Baer, principal at dbInsight. Tony, thank you for joining me.

Tony: Oh, Corey, thanks for having me. And by the way, we’ll try and basically knock down your primal fear of databases today. That’s my mission.

Corey: We’re going to instill new fears in you. Because I was looking through a lot of your work over the years, and the criticism I have—and always the best place to deliver criticism is massively in public—is that you take a very conservative, stodgy approach to defining a database, whereas I’m on the opposite side of the world. I contain information. You can ask me about it, which we’ll call querying. That’s right. I’m a database.

But I’ve never yet found myself listed in any of your analyses around various database options. So, what is your definition of databases these days? Where do they start and stop?

Tony: Oh, gosh.

Corey: Because anything can be a database if you hold it wrong.

Tony: [laugh]. I think one of the last things I’ve ever been called as conservative and stodgy, so this is certainly a way to basically put the thumbtack on my share.

Corey: Exactly. I’m trying to normalize my own brand of lunacy, so we’ll see how it goes.

Tony: Exactly because that’s the role I normally play with my clients. So, now the shoe is on the other foot. What I view a database is, is basically a managed collection of data, and it’s managed to the point where essentially, a database should be transactional—in other words, when I basically put some data in, I should have some positive information, I should hopefully, depending on the type of database, have some sort of guidelines or schema or model for how I structure the data. So, I mean, database, you know, even though you keep hearing about unstructured data, the fact is—

Corey: Schemaless databases and data stores. Yeah, it was all the rage for a few years.

Tony: Yeah, except that they all have schemas, just that those schemaless databases just have very variable schema. They’re still schema.

Corey: A question that I have is you obviously think deeply about these things, which should not come as a surprise to anyone. It’s like, “Well, this is where I spend my entire career. Imagine that. I might think about the problem space a little bit.” But you have, to my understanding, never worked with databases in anger yourself. You don’t have a history as a DBA or as an engineer—

Tony: No.

Corey: —but what I find very odd is that unlike a whole bunch of other analysts that I’m not going to name, but people know who I’m talking about regardless, you bring actual insights into this that I find useful and compelling, instead of reverting to the mean of well, I don’t actually understand how any of these things work in reality, so I’m just going to believe whoever sounds the most confident when I ask a bunch of people about these things. Are you just asking the right people who also happen to sound confident? But how do you get away from that very common analyst trap?

Tony: Well, a couple of things. One is I purposely play the role of outside observer. In other words, like, the idea is that if basically an idea is supposed to stand on its own legs, it has to make sense. If I’ve been working inside the industry, I might take too many things for granted. And a good example of this goes back, actually, to my early days—actually this goes back to my freshman year in college where I was taking an organic chem course for non-majors, and it was taught as a logic course not as a memorization course.

And we were given the option at the end of the term to either, basically, take a final or do a paper. So, of course, me being a writer I thought, I can BS my way through this. But what I found—and this is what fascinated me—is that as long as certain technical terms were defined for me, I found a logic to the way things work. And so, that really informs how I approach databases, how I approach technology today is I look at the logic on how things work. That being said, in order for me to understand that, I need to know twice as much as the next guy in order to be able to speak that because I just don’t do this in my sleep.

Corey: That goes a big step toward, I guess, addressing a lot of these things, but it also feels like—and maybe this is just me paying closer attention—that the world of databases and data and analytics have really coalesced or emerged in a very different way over the past decade-ish. It used to be, at least from my perspective, that oh, that the actual, all the data we store, that’s a storage admin problem. And that was about managing NetApps and SANs and the rest. And then you had the database side of it, which functionally from the storage side of the world was just a big file or series of files that are the backing store for the database. And okay, there’s not a lot of cross-communication going on there.

Then with the rise of object store, it started being a little bit different. And even the way that everyone is talking about getting meaning from data has really seem to be evolving at an incredibly intense clip lately. Is that an accurate perception, or have I just been asleep at the wheel for a while and finally woke up?

Tony: No, I think you’re onto something there. And the reason is that, one, data is touching us all around ourselves, and the fact is, I mean, I’m you can see it in the same way that all of a sudden that people know how to spell AI. They may not know what it means, but the thing is, there is an awareness the data that we work with, the data that is about us, it follows us, and with the cloud, this data has—well, I should say not just with the cloud but with smart mobile devices—we’ll blame that—we are all each founts of data, and rich founts of data. And people in all walks of life, not just in the industry, are now becoming aware of it and there’s a lot of concern about can we have any control, any ownership over the data that should be ours? So, I think that phenomenon has also happened in the enterprise, where essentially where we used to think that the data was the DBAs’ issue, it’s become the app developers’ issue, it’s become the business analysts’ issue. Because the answers that we get, we’re ultimately accountable for. It all comes from the data.

Corey: It also feels like there’s this idea of databases themselves becoming more contextually aware of the data contained within them. Originally, this used to be in the realm of, “Oh, we know what’s been accessed recently and we can tier out where it lives for storage optimization purposes.” Okay, great, but what I’m seeing now almost seems to be a sense of, people like to talk about pouring ML into their database offerings. And I’m not able to tell whether that is something that adds actual value, or if it’s marketing-ware.

Tony: Okay. First off, let me kind of spill a couple of things. First of all, it’s not a question of the database becoming aware. A database is not sentient.

Corey: Niether are some engineers, but that’s neither here nor there.

Tony: That would be true, but then again, I don’t want anyone with shotguns lining up at my door after this—

Corey: [laugh].

Tony: —after this interview is published. But [laugh] more of the point, though, is that I can see a couple roles for machine learning in databases. One is a database itself, the logs, are an incredible font of data, of operational data. And you can look at trends in terms of when this—when the pattern of these logs goes this way, that is likely to happen. So, the thing is that I could very easily say we’re already seeing it: machine learning being used to help optimize the operation of databases, if you’re Oracle, and say, “Hey, we can have a database that runs itself.”

The other side of the coin is being able to run your own machine-learning models in database as opposed to having to go out into a separate cluster and move the data, and that’s becoming more and more of a checkbox feature. However, that’s going to be for essentially, probably, like, the low-hanging fruit, like the 80/20 rule. It’ll be like the 20% of an ana—of relatively rudimentary, you know, let’s say, predictive analyses that we can do inside the database. If you’re going to be doing something more ambitious, such as a, you know, a large language model, you probably do not want to run that in database itself. So, there’s a difference there.

Corey: One would hope. I mean, one of the inappropriate uses of technology that I go for all the time is finding ways to—as directed or otherwise—in off-label uses find ways of tricking different services into running containers for me. It’s kind of a problem; this is probably why everyone is very grateful I no longer write production code for anyone.

But it does seem that there’s been an awful lot of noise lately. I’m lazy. I take shortcuts very often, and one of those is that whenever AWS talks about something extensively through multiple marketing cycles, it becomes usually a pretty good indicator that they’re on their back foot on that area. And for a long time, they were doing that about data and how it’s very important to gather data, it unlocks the key to your business, but it always felt a little hollow-slash-hypocritical to me because you’re going to some of the same events that I have that AWS throws on. You notice how you have to fill out the exact same form with a whole bunch of mandatory fields every single time, but there never seems to be anything that gets spat back out to you that demonstrates that any human or system has ever read—

Tony: Right.

Corey: Any of that? It’s basically a, “Do what we say, not what we do,” style of story. And I always found that to be a little bit disingenuous.

Tony: I don’t want to just harp on AWS here. Of course, we can always talk about the two-pizza box rule and the fact that you have lots of small teams there, but I’d rather generalize this. And I think you really—what you’re just describing is been my trip through the healthcare system. I had some sports-related injuries this summer, so I’ve been through a couple of surgeries to repair sports injuries. And it’s amazing that every time you go to the doctor’s office, you’re filling the same HIPAA information over and over again, even with healthcare systems that use the same electronic health records software. So, it’s more a function of that it’s not just that the technologies are siloed, it’s that the organizations are siloed. That’s what you’re saying.

Corey: That is fair. And I think at some level—I don’t know if this is a weird extension of Conway’s Law or whatnot—but these things all have different backing stores as far as data goes. And there’s a—the hard part, it seems, in a lot of companies once they hit a certain point of maturity is not just getting the data in—because they’ve already done that to some extent—but it’s also then making it actionable and helping various data stores internal to the company reconcile with one another and start surfacing things that are useful. It increasingly feels like it’s less of a technology problem and more of a people problem.

Tony: It is. I mean, put it this way, I spent a lot of time last year, I burned a lot of brain cells working on data fabrics, which is an idea that’s in the idea of the beholder. But the ideal of a data fabric is that it’s not the tool that necessarily governs your data or secures your data or moves your data or transforms your data, but it’s supposed to be the master orchestrator that brings all that stuff together. And maybe sometime 50 years in the future, we might see that.

I think the problem here is both technical and organizational. [unintelligible 00:11:58] a promise, you have all these what we used call island silos. We still call them silos or islands of information. And actually, ironically, even though in the cloud we have technologies where we can integrate this, the cloud has actually exacerbated this issue because there’s so many islands of information, you know, coming up, and there’s so many different little parts of the organization that have their hands on that. That’s also a large part of why there’s such a big discussion about, for instance, data mesh last year: everybody is concerned about owning their own little piece of the pie, and there’s a lot of question in terms of how do we get some consistency there? How do we all read from the same sheet of music? That’s going to be an ongoing problem. You and I are going to get very old before that ever gets solved.

Corey: Yeah, there are certain things that I am content to die knowing that they will not get solved. If they ever get solved, I will not live to see it, and there’s a certain comfort in that, on some level.

Tony: Yeah.

Corey: But it feels like this stuff is also getting more and more complicated than it used to be, and terms aren’t being used in quite the same way as they once were. Something that a number of companies have been saying for a while now has been that customers overwhelmingly are preferring open-source. Open source is important to them when it comes to their database selection. And I feel like that’s a conflation of a couple of things. I’ve never yet found an ideological, purity-driven customer decision around that sort of thing.

What they care about is, are there multiple vendors who can provide this thing so I’m not going to be using a commercially licensed database that can arbitrarily start playing games with seat licenses and wind up distorting my cost structure massively with very little notice. Does that align with your—

Tony: Yeah.

Corey: Understanding of what people are talking about when they say that, or am I missing something fundamental? Which is again, always possible?

Tony: No, I think you’re onto something there. Open-source is a whole other can of worms, and I’ve burned many, many brain cells over this one as well. And today, you’re seeing a lot of pieces about the, you know, the—that are basically giving eulogies for open-source. It’s—you know, like HashiCorp just finally changed its license and a bunch of others have in the database world. What open-source has meant is been—and I think for practitioners, for DBAs and developers—here’s a platform that’s been implemented by many different vendors, which means my skills are portable.

And so, I think that’s really been the key to why, for instance, like, you know, MySQL and especially PostgreSQL have really exploded, you know, in popularity. Especially Postgres, you know, of late. And it’s like, you look at Postgres, it’s a very unglamorous database. If you’re talking about stodgy, it was born to be stodgy because they wanted to be an adult database from the start. They weren’t the LAMP stack like MySQL.

And the secret of success with Postgres was that it had a very permissive open-source license, which meant that as long as you don’t hold University of California at Berkeley, liable, have at it, kids. And so, you see, like, a lot of different flavors of Postgres out there, which means that a lot of customers are attracted to that because if I get up to speed on this Postgres—on one Postgres database, my skills should be transferable, should be portable to another. So, I think that’s a lot of what’s happening there.

Corey: Well, I do want to call that out in particular because when I was coming up in the naughts, the mid-2000s decade, the lingua franca on everything I used was MySQL, or as I insist on mispronouncing it, my-squeal. And lately, on same vein, Postgres-squeal seems to have taken over the entire universe, when it comes to the de facto database of choice. And I’m old and grumpy and learning new things as always challenging, so I don’t understand a lot of the ways that thing gets managed from the context coming from where I did before, but what has driven the massive growth of mindshare among the Postgres-squeal set?

Tony: Well, I think it’s a matter of it’s 30 years old and it’s—number one, Postgres always positioned itself as an Oracle alternative. And the early years, you know, this is a new database, how are you going to be able to match, at that point, Oracle had about a 15-year headstart on it. And so, it was a gradual climb to respectability. And I have huge respect for Oracle, don’t get me wrong on that, but you take a look at Postgres today and they have basically filled in a lot of the blanks.

And so, it now is a very cre—in many cases, it’s a credible alternative to Oracle. Can it do all the things Oracle can do? No. But for a lot of organizations, it’s the 80/20 rule. And so, I think it’s more just a matter of, like, Postgres coming of age. And the fact is, as a result of it coming of age, there’s a huge marketplace out there and so much choice, and so much opportunity for skills portability. So, it’s really one of those things where its time has come.

Corey: I think that a lot of my own biases are simply a product of the era in which I learned how a lot of these things work on. I am terrible at Node, for example, but I would be hard-pressed not to suggest JavaScript as the default language that people should pick up if they’re just entering tech today. It does front-end, it does back-end—

Tony: Sure.

Corey: —it even makes fries, apparently. There’s a—that is the lingua franca of the modern internet in a bunch of different ways. That doesn’t mean I’m any good at it, and it doesn’t mean at this stage, I’m likely to improve massively at it, but it is the right move, even if it is inconvenient for me personally.

Tony: Right. Right. Put it this way, we’ve seen—and as I said, I’m not an expert in programming languages, but we’ve seen a huge profusion of programming languages and frameworks. But the fact is that there’s always been a draw towards critical mass. At the turn of the millennium, we thought is between Java and .NET. Little did we know that basically JavaScript—which at that point was just a web scripting language—[laugh] we didn’t know that it could work on the server; we thought it was just a client. Who knew?

Corey: That’s like using something inappropriately as a database. I mean, good heavens.

Tony: [laugh]. That would be true. I mean, when I could have, you know, easily just use a spreadsheet or something like that. But so, I mean, who knew? I mean, just like for instance, Java itself was originally conceived for a set-top box. You never know how this stuff is going to turn out. It’s the same thing happen with Python. Python was also a web scripting language. Oh, by the way, it happens to be really powerful and flexible for data science. And whoa, you know, now Python is—in terms of data science languages—has become the new SaaS.

Corey: It really took over in a bunch of different ways. Before that, Perl was great, and I go, “Why would I use—why write in Python when Perl is available?” It’s like, “Okay, you know, how to write Perl, right?” “Yeah.” “Have you ever read anything a month later?” “Oh…” it’s very much a write-only language. It is inscrutable after the fact. And Python at least makes that a lot more approachable, which is never a bad thing.

Tony: Yeah.

Corey: Speaking of what you touched on toward the beginning of this episode, the idea of databases not being sentient, which I equate to being self-aware, you just came out very recently with a report on generative AI and a trip that you wound up taking on this. Which I’ve read; I love it. In fact, we’ve both been independently using the phrase [unintelligible 00:19:09] to, “English is the new most common programming language once a lot of this stuff takes off.” But what have you seen? What have you witnessed as far as both the ground truth reality as well as the grandiose statements that companies are making as they trip over themselves trying to position as the forefront leader and all of this thing that didn’t really exist five months ago?

Tony: Well, what’s funny is—and that’s a perfect question because if on January 1st you asked “what’s going to happen this year?” I don’t think any of us would have thought about generative AI or large language models. And I will not identify the vendors, but I did some that had— was on some advanced briefing calls back around the January, February timeframe. They were talking about things like server lists, they were talking about in database machine learning and so on and so forth. They weren’t saying anything about generative.

And all of a sudden, April, it changed. And it’s essentially just another case of the tail wagging the dog. Consumers were flocking to ChatGPT and enterprises had to take notice. And so, what I saw, in the spring was—and I was at a conference from SaaS, I’m [unintelligible 00:20:21] SAP, Oracle, IBM, Mongo, Snowflake, Databricks and others—that they all very quickly changed their tune to talk about generative AI. What we were seeing was for the most part, position statements, but we also saw, I think, the early emphasis was, as you say, it’s basically English as the new default programming language or API, so basically, coding assistance, what I’ll call conversational query.

I don’t want to call it natural language query because we had stuff like Tableau Ask Data, which was very robotic. So, we’re seeing a lot of that. And we’re also seeing a lot of attention towards foundation models because I mean, what organization is going to have the resources of a Google or an open AI to develop their own foundation model? Yes, some of the Wall Street houses might, but I think most of them are just going to say, “Look, let’s just use this as a starting point.”

I also saw a very big theme for your models with your data. And where I got a hint of that—it was a throwaway LinkedIn post. It was back in, I think like, February, Databricks had announced Dolly, which was kind of an experimental foundation model, just to use with your own data. And I just wrote three lines in a LinkedIn post, it was on Friday afternoon. By Monday, it had 65,000 hits.

I’ve never seen anything—I mean, yes, I had a lot—I used to say ‘data mesh’ last year, and it would—but didn’t get anywhere near that. So, I mean, that really hit a nerve. And other things that I saw, was the, you know, the starting to look with vector storage and how that was going to be supported was it was going be a new type of database, and hey, let’s have AWS come up with, like, an, you know, an [ADF 00:21:41] database here or is this going to be a feature? I think for the most part, it’s going to be a feature. And of course, under all this, everybody’s just falling in love, falling all over themselves to get in the good graces of Nvidia. In capsule, that’s kind of like what I saw.

Corey: That feels directionally accurate. And I think databases are a great area to point out one thing that’s always been more a little disconcerting for me. The way that I’ve always viewed databases has been, unless I’m calling a RAND function or something like it and I don’t change the underlying data structure, I should be able to run a query twice in a row and receive the same result deterministically both times.

Tony: Mm-hm.

Corey: Generative AI is effectively non-deterministic for all realistic measures of that term. Yes, I’m sure there’s a deterministic reason things are under the hood. I am not smart enough or learned enough to get there. But it just feels like sometimes we’re going to give you the answer you think you’re going to get, sometimes we’re going to give you a different answer. And sometimes, in generative AI space, we’re going to be supremely confident and also completely wrong. That feels dangerous to me.

Tony: [laugh]. Oh gosh, yes. I mean, I take a look at ChatGPT and to me, the responses are essentially, it’s a high school senior coming out with an essay response without any footnotes. It’s the exact opposite of an ACID database. The reason why we’re very—in the database world, we’re very strongly drawn towards ACID is because we want our data to be consistent and to get—if we ask the same query, we’re going to get the same answer.

And the problem is, is that with generative, you know, based on large language models, computers sounds sentient, but they’re not. Large language models are basically just a series of probabilities, and so hopefully those probabilities will line up and you’ll get something similar. That to me, kind of scares me quite a bit. And I think as we start to look at implementing this in an enterprise setting, we need to take a look at what kind of guardrails can we put on there. And the thing is, that what this led me to was that missing piece that I saw this spring with generative AI, at least in the data and analytics world, is nobody had a clue in terms of how to extend AI governance to this, how to make these models explainable. And I think that’s still—that’s a large problem. That’s a huge nut that it’s going to take the industry a while to crack.

Corey: Yeah, but it’s incredibly important that it does get cracked.

Tony: Oh, gosh, yes.

Corey: One last topic that I want to get into. I know you said you don’t want to over-index on AWS, which, fair enough. It is where I spend the bulk of my professional time and energy—

Tony: [laugh].

Corey: Focusing on, but I think this one’s fair because it is a microcosm of a broader industry question. And that is, I don’t know what the DBA job of the future is going to look like, but increasingly, it feels like it’s going to primarily be picking which purpose-built AWS database—or larger [story 00:24:56] purpose database is appropriate for a given workload. Even without my inappropriate misuse of things that are not databases as databases, they are legitimately 15 or 16 different AWS services that they position as database offerings. And it really feels like you’re spiraling down a well of analysis paralysis, trying to pick between all these things. Do you think the future looks more like general-purpose databases, or very purpose-built and each one is this beautiful, bespoke unicorn?

Tony: [laugh]. Well, this is basically a hit on a theme that I’ve been—you know, we’ve been all been thinking about for years. And the thing is, there are arguments to be made for multi-model databases, you know, versus a for-purpose database. That being said, okay, two things. One is that what I’ve been saying, in general, is that—and I wrote about this way, way back; I actually did a talk at the [unintelligible 00:25:50]; it was a throwaway talk, or [unintelligible 00:25:52] one of those conferences—I threw it together and it’s basically looking at the emergence of all these specialized databases.

But how I saw, also, there’s going to be kind of an overlapping. Not that we’re going to come back to Pangea per se, but that, for instance, like, a relational database will be able to support JSON. And Oracle, for instance, does has some fairly brilliant ideas up the sleeve, what they call a JSON duality, which sounds kind of scary, which basically says, “We can store data relationally, but superimpose GraphQL on top of all of this and this is going to look really JSON-y.” So, I think on one hand, you are going to be seeing databases that do overlap. Would I use Oracle for a MongoDB use case? No, but would I use Oracle for a case where I might have some document data? I could certainly see that.

The other point, though, and this is really one I want to hammer on here—it’s kind of a major concern I’ve had—is I think the cloud vendors, for all their talk that we give you operational simplicity and agility are making things very complex with its expanding cornucopia of services. And what they need to do—I’m not saying, you know, let’s close down the patent office—what I think we do is we need to provide some guided experiences that says, “Tell us the use case. We will now blend these particular services together and this is the package that we would suggest.” I think cloud vendors really need to go back to the drawing board from that standpoint and look at, how do we bring this all together? How would he really simplify the life of the customer?

Corey: That is, honestly, I think the biggest challenge that the cloud providers have across the board. There are hundreds of services available at this point from every hyperscaler out there. And some of them are brand new and effectively feel like they’re there for three or four different customers and that’s about it and others are universal services that most people are probably going to use. And most things fall in between those two extremes, but it becomes such an analysis paralysis moment of trying to figure out what do I do here? What is the golden path?

And what that means is that when you start talking to other people and asking their opinion and getting their guidance on how to do something when you get stuck, it’s, “Oh, you’re using that service? Don’t do it. Use this other thing instead.” And if you listen to that, you get midway through every problem for them to start over again because, “Oh, I’m going to pick a different selection of underlying components.” It becomes confusing and complicated, and I think it does customers largely a disservice. What I think we really need, on some level, is a simplified golden path with easy on-ramps and easy off-ramps where, in the absence of a compelling reason, this is what you should be using.

Tony: Believe it or not, I think this would be a golden case for machine learning.

Corey: [laugh].

Tony: No, but submit to us the characteristics of your workload, and here’s a recipe that we would propose. Obviously, we can’t trust AI to make our decisions for us, but it can provide some guardrails.

Corey: “Yeah. Use a graph database. Trust me, it’ll be fine.” That’s your general purpose—

Tony: [laugh].

Corey: —approach. Yeah, that’ll end well.

Tony: [laugh]. I would hope that the AI would basically be trained on a better set of training data to not come out with that conclusion.

Corey: One could sure hope.

Tony: Yeah, exactly.

Corey: I really want to thank you for taking the time to catch up with me around what you’re doing. If people want to learn more, where’s the best place for them to find you?

Tony: My website is dbinsight.io. And on my homepage, I list my latest research. So, you just have to go to the homepage where you can basically click on the links to the latest and greatest. And I will, as I said, after Labor Day, I’ll be publishing my take on my generative AI journey from the spring.

Corey: And we will, of course, put links to this in the [show notes 00:29:39]. Thank you so much for your time. I appreciate it.

Tony: Hey, it’s been a pleasure, Corey. Good seeing you again.

Corey: Tony Baer, principal at dbInsight. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an angry, insulting comment that we will eventually stitch together with all those different platforms to create—that’s right—a large-scale distributed database.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

View Details

Alex Gallego, CEO & Founder of Redpanda, joins Corey on Screaming in the Cloud to discuss his experience founding and scaling a successful data streaming company over the past 4 years. Alex explains how it’s been a fun and humbling journey to go from being an engineer to being a founder, and how he’s built a team he trusts to hand the production off to. Corey and Alex discuss the benefits and various applications of Redpanda’s data streaming services, and Alex reveals why it was so important to him to focus on doing one thing really well when it comes to his product strategy. Alex also shares details on the Hack the Planet scholarship program he founded for individuals in underrepresented communities.

About Alex

Alex Gallego is the founder and CEO of Redpanda, the streaming data platform for developers. Alex has spent his career immersed in deeply technical environments, and is passionate about finding and building solutions to the challenges of modern data streaming. Prior to Redpanda, Alex was a principal engineer at Akamai, as well as co-founder and CTO of Concord.io, a high-performance stream-processing engine acquired by Akamai in 2016. He has also engineered software at Factset Research Systems, Forex Capital Markets and Yieldmo; and holds a bachelor’s degree in computer science and cryptography from NYU.

Links Referenced:

  • Redpanda: https://redpanda.com/
  • Twitter: https://twitter.com/emaxerrno
  • Redpanda community Slack: https://redpandacommunity.slack.com/join/shared_invite/zt-1xq6m0ucj-nI41I7dXWB13aQ2iKBDvDw

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Tired of slow database performance and bottlenecks on MySQL or PostgresSQL when using Amazon RDS or Aurora? How’d you like to reduce query response times by ninety percent? Better yet, how would you like to get me to pronounce database names correctly? Join customers like Zscaler, Intel, Booking.com, and others that use OtterTune’s artificial intelligence to automatically optimize and keep their databases healthy. Go to ottertune dot com to learn more and start a free trial. That’s O-T-T-E-R-T-U-N-E dot com.

Corey: Welcome to Screaming in the Cloud, I’m Corey Quinn, and this promoted guest episode is brought to us by our friends at Redpanda, which I’m thrilled about because I have a personal affinity for companies that have cartoon mascots in the form of animals and are willing to at least be slightly creative with them. My guest is Alex Gallego, the founder and CEO over at Redpanda. Alex, thanks for joining me.

Alex: Corey, thanks for having me.

Corey: So, I’m not asking about the animal; I’m talking about the company, which I imagine is a frequent source of disambiguation when you meet people at parties and they don’t quite understand what it is that you do. And you folks are big in the data streaming space, but data streaming can mean an awful lot of things to an awful lot of people. What is it for you?

Alex: Largely it’s about enabling developers to build applications that can extract value of every single event, every click, every mouse movement, every transaction, every event that goes through your network. This is what Redpanda is about. It’s like how do we help you make more money with every single event? How do we help you be more successful? And you know, happy to give examples in finance, or IoT, or oil and gas, if it’s helpful for the audience, but really, to me, it’s like, okay, if we can give you the framework in which you can build a new application that allows you to extract value out of data, every single event that’s going through your network, to me, that’s what a streaming is about. It large, it’s you know, data contextualized with a timestamp and largely, a sort of a database of event streaming.

Corey: One of the things that I find curious about the space is that usually, companies wind up going one of two directions when you’re talking about data streaming. Either there, “Oh, just send it all to us and we’ll take care of it for you,” or otherwise, it’s a, great they more or less ship something that you’ve run in your own environment. In the olden days of data centers, that usually resembled a box of some sort. You’re one of those interesting split-the-difference companies where you offer both models. Do you find that one of those tends to be seeing more adoption these days or that there’s an increasing trend toward one direction or the other?

Alex: Yeah. So, right now, I think that to me, the future of all these data-intensive products—whether you’re a database or a streaming engine—will, because simply of cost of networks transferred between the hybrid clouds and your accounts, sending a gigabyte a second of data between, let’s say, you know, your data center and a vendor, it’s just so expensive that at some point, from just a cost perspective, like, running the infrastructure, it’s in the millions of dollars. And so, running the data inside your VPC, it’s sort of the next logical evolution of how we’ve used to consume services. And so, I actually think it’s just the evolution: people would self-host because of costs and then they would use services because of operational simplicity. “I don’t want to spend team skills and time building this. I want to pay a vendor.”

And so, BYOC, to be honest—which is what we call this offering—it was about [laugh] sidestepping the costs and of being stuck in the hybrid clouds, whether it’s Google or Amazon, where you’re paying egress and ingress costs and it’s just so expensive, in addition to this whole idea of data residency or data sovereignty and privacy. It’s like, yeah, why not both? Like, if I’m an engineer, I want low latency and I don’t want to pay you to transfer this thing to the next rack. I mean, my computer’s probably, like, you know, a hundred feet away from my customer's computer. Like, why [laugh] way is that so complicated? So, you know, my view is that the future of data-intensive products will be in this form of where it—like, data planes are actually owned by companies, and then you offer that as a Software as a Service.

Corey: One of the things that catches an awful lot of companies with telemetry use cases—or data streaming as another example of that—by surprise when they start building their own cloud-hosted offering is that they’re suddenly seeing a lot more cross-AZ data charges than they would have potentially expected. And that’s because unlike cross-region or the really expensive version of this with egress, it’s a penny in and a penny out per gigabyte in most of AWS regions. Which means that that isn’t also bound strictly to an AWS organization. So, you have customers co-located with you and you’re starting to pay ingress charges on customers throwing their data over to you. And, on some level, the most economical solution for you is well, we’re just going to put our listeners somewhere else far away so that we can just have them pay the steep egress fee but then we can just reflect it back to ourselves for free.

And that’s a terrible pattern, but it’s a byproduct of the absolutely byzantine cross-AZ data transfer pricing, in fact, all of the data transfer pricing that is at least AWS tends to present. And it shapes the architectural decisions you make as a result.

Alex: You know, as a user, it just didn’t make sense. When we launched this product, the number of people that says like, “Why wouldn’t your charge for, you know, effectively renting [unintelligible 00:05:14], and giving a markup to your customers?” That’s we don’t add any value on that, you know? I think people should really just pay us for the value that we create for them. And so, you know, for us competing with other companies is relatively easy.

Competing with MSK is it’s harder because MSK just has this, you know, muscle where they don’t charge you for some particular network traffic between you. And so, it forces companies like us that are trying to be innovative in the data space to, like, put our services in that so that we can actually compete in the market. And so, it’s a forcing function of the hybrid clouds having this strong muscle of being able to discount their services in a way that companies just simply don’t have access to. And then, you know, it becomes—for the others—latency and sovereignty.

Corey: This is the way that effectively all of AWS has first-party offerings of other things go. Replication traffic between AZs is not chargeable. And when I asked them about that, they say, “Oh, yeah. We just price that into the cost of the service.” I don’t know that I necessarily buy that because if I try and run this sort of thing on top of EC2, it would cost me more than using their crappy implementation of it, just in data transfer alone for an awful lot of use cases.

No third party can touch that level of cost-effectiveness and discounting. It really is probably the clearest example I can think of actual anti-competitive behavior in the market. But it’s also complex enough to explain, to, you know, regulators that it doesn’t make for exciting exposés and the basis for lawsuits. Yet. Hope springs eternal.

Alex: [laugh]. You know—okay, so here is how—if someone is listening to this podcast and is, like, “Okay, well, what can I do?” For us, S3 is the answer. S3 is basically you need to be able to lean in into S3 as a way of replication across [AZ 00:06:56], you need to be able to lean into S3 to read data. And so actually, when I wrote, originally, Redpanda, you know, it’s just like this C++ thing using [unintelligible 00:07:04], geared towards super low latency.

When we moved it into the cloud, what we realized is, this is cost prohibitive to run either on EBS volumes or local disk. I have to tier all the storage into S3, so that I can use S3’s cross-AZ network transfer, which is basically free, to be able to then bring a separate cluster on a different AZ, and then read from the bucket at zero cost. And so, you end up really—like, there are fundamental technical things that you have to do to just be able to compete in a way that’s cost-effective for you. And so, in addition to just, like, the muscle that they can enforce on the companies is—it—there are deep implications of what it translates to at the technical level. Like, at the code level.

Corey: In the cloud, more than almost anywhere else, it really does become apparent that cost and architecture are fundamentally the same thing. And I have a bit of an advantage here in that I’ve seen what you do deployed at least one customer of mine. It’s fun. When you have a bunch of logos on your site, it’s, “Hey, I recognize some of those.” And what I found interesting was the way that multiple people, when I spoke to them, described what it is that you do because some of them talked about it purely as a cost play, but other people were just as enthusiastic about it being a means of improving feature velocity and unlocking capabilities that they didn’t otherwise have or couldn’t have gotten to without a whole lot of custom work on their part. Which is it? How do you view what it is that you’re bringing to market? Is it a cost play or is it a capability story?

Alex: From our customer base, I would say 40% is—of our customer base—is about Redpanda enabling them to do things that they simply couldn’t do before. An example is, we have, you know, a Fortune 100 company that they basically run their hedge trading strategy on top of Redpanda. And the reason for that is because we give them a five-millisecond average latency with predictable flight latencies, right? And so, for them, that predictability of Redpanda, you know, and sort of like the architecture that came about from trying to invent a new storage engine, allows them to throw away a bunch of in-house, you know, custom-built pub/sub messaging that, you know, basically gave them the same or worse latency. And so, for them, there’s that.

For others, I think in the IoT space, or if you have flying vehicles around the world, we have some logos that, you know, I just can’t mention them. But they have this, like, flying computers around the world and they want to measure that. And so, like, the profile of the footprint, like, the mechanical footprint of being able to run on a single Pthread with a few megs of memory allows these new deployment models that, you know, simply, it’s just, it’s not possible with the alternatives where let’s say you have to have, you know, like, a zookeeper on the schema registry and an HTTP proxy and a broker and all of these things. That simply just, it cannot run on a single Pthread with a few megs of memory, if you put any sort of workload into that. And so, it’s like, the computational efficiencies simply enable new things that you couldn’t do before. And that’s probably 40%. And then the other, it’s just… money was really cheap last year [laugh] or the year before and I think now it’s less cheap [unintelligible 00:10:08] yeah.

Corey: Yeah, I couldn’t help but notice that in my own business, too. It turns out that not giving a shit about the AWS bill was a zero-interest-rate phenomenon. Who knew?

Alex: [laugh]. Yeah, exactly. And now people [unintelligible 00:10:17], you know, the CIOs in particular, it’s like, help. And so, that’s really 60%, and our business has boomed since.

Corey: Yeah, one thing that I find interesting is that you’ve been around for only four years. I know that’s weird to say ‘only,’ but time moves differently in tech. And you’ve started showing up in some very strange places that I would not have expected. You recently—somewhat recently; time is, of course, a flat circle—completed $100 million Series C, and I also saw you in places where I didn’t expect to see you in the form of, last week, one of your large competitor's earnings calls, where they were asked by an analyst about an unnamed company that had raised $100 million Series C, and the CEO [unintelligible 00:11:00], “Oh, you’re probably talking about Redpanda.” And then they gave an answer that was fine.

I mean, no one is going to be on an earnings call and not be prepared for questions like that and to not have an answer ready to go. No one’s going to say, “Well, we’re doomed if it works,” because I think that businesses are more sophisticated than that. But it was an interesting shout-out in a place where you normally don’t see competitors validate that you’re doing something interesting by name-checking you.

Alex: What was fundamentally interesting for me about that, is that I feel that as an investor, if you’re putting you know, 2, 3, 4, or $500 million check into a public position of a company, you want to know, is this money simply going to make returns? That’s basically what an investor cares about. And so, the reason for that question is, “Hey, there’s a Series C startup company that now has a bunch of these Fortune 2000 logos,” and you know, when we talked to them, like, their customer [unintelligible 00:11:51] phenomena, like, why is that the case? And then, you know, our competitor was forced to name, you know, [laugh] a single win. That’s as far as I remember it. We don’t know of any additional customers that have switched to that.

And so, I think when you have, like, you know, your win rate is above, whatever, 95%, 97% ratio, then I think, you know, they’re just sort of forced to answer that. And in a way, I just think that they focus on different things. And for me, it was like, “Okay, developer, hands on keyboard, behind the terminal, how do I make you successful?” And that seems to have worked out enough to be mentioned in the earnings call.

Corey: On some level, it’s a little bit of a dog-and-pony show. I think that as companies had a certain point of scale, they feel that they need to validate what they’re doing to investors at various points—which is always, on some level, of concern—and validate themselves to analysts, both financial—which, okay, whatever—and also, industry analysts, where they come with checklists that they believe is what customers want and is often a little bit off of the mark. But the validation that I think that matters, that actually determines whether or not something has legs is what your customers—you know, people paying you money for a thing—have to say and what they take away from what you’re doing. And having seen in a couple of cases now myself, that usage of Redpanda has increased after initial proofs of concept and putting things on to it, I already sort of know the answer to this, but it seems that you also have a vibrant community of boosters for people who are thrilled to use the thing you’re selling them.

Alex: You know, Jumptraders recently posted that there was a use case in the new stack where they, like, put for the most mission-critical. So, for those of you that listening, Jumptraders is financial company, and they’re super technical company. One of, like, the hardest things, they’ll probably put your [unintelligible 00:13:35] your product through some of the most rigorous testing [unintelligible 00:13:38]. So, when you start doing some of these logos, it gives confidence. And actually, the majority of our developers that we get to partner with, it was really a friend telling a friend, for [laugh] the longest time, my marketing department was super, super small.

And then what’s been fun, some, like, really different use case was the one I mentioned about on this, like, flying vehicles around the world. They fly both in outer space and in airplanes. That was really fun. And then the large one is when you have workloads at, like, 14-and-a-half gigabytes per second, where the alternative of using something like Kinesis in the case of Lacework—which, you know, they wrote a new stack article about—would be so exorbitantly expensive. And so, in a way, I think that, you know, just trying to make the developers successful, really focusing, honestly, on the person who just has to make things work. We don’t—by the time we get to the CIO, really the champion was the engineer who had to build an application. “I was just trying to figure it out the whack-a-mole of trying to debug alternative systems.”

Corey: One of the, I think, seductive problems with your entire space is that no one decides day one that they’re going to implement a data streaming solution for a very scaled-out, high-traffic site. The early adoption is always a small thing that you’re in the process of building. And at that scale at that speed, it just doesn’t feel like it’s that hard of a problem because scale introduces its own unique series of challenges, but it’s often one that people only really find out themselves when the simple thing that works in theory but not in production starts to cause problems internally. I used to work with someone who was a deeply passionate believer in Apache Kafka to a point where it almost became a problem, just because their answer to every problem—it almost didn’t matter if it was, “How do we get more coffee this morning?”—Kafka would be the answer for all of it.

And that’s great, but it turned out, they became one of these people that borderline took on a product or a technology as their identity. So, anything that would potentially take a workload away from that, I got a lot of internal resistance. I’m wondering if you find that you’re being brought in to replace existing systems or for completely greenfield stuff. And if the former, are you seeing a lot of internal resistance to people who have built a little niche for themselves?

Alex: It’s true, the people that have built a career, especially at large banks, were a pretty good fit for, you know, they actually get a team, they got a promotion cycle because they brought this technology and the technology sort of helped them make money. I personally tend to love to talk to these people. And there was a ca—to me, like, technically, let’s talk about, like, deeply technical. Let me help you. That obviously doesn’t scale because I can’t have the same conversation with ten people.

So, we do tend to see some of that. Actually, from our customers' standpoint, I would say that the large part of our customer base, you know, if I’m trying to put numbers, maybe 65%, I probably rip and replace of, you know, either upstream Apache Software or private companies or hosted services, et cetera. And so, I think you’re right in saying, “Hey, that resistance,” they probably handled the [unintelligible 00:16:38], but what changed in the last year is that the CIO now stepped in and says, “I am going to fire all of you or you have to come up with a $10 million savings. Help me.” [laugh]. And so, you know, then really, my job is to help them look like a hero.

It’s like, “Hey, look, try it tested, benchmark it in your with your own workload, and if it saves you money, then use it.” That’s been, you know, to sort of super helpful kind of on the macroeconomic environment. And then the last one is sometimes, you know, you do have to go with a greenfield, right? Like, someone has built a career, they want to gain confidence, they want to ask you questions, they want to trust you that you don’t lose data, they want to make sure that you do say the things that you want to say. And so, sometimes it’s about building trust and building that relationship.

And developers are right. Like, there’s a bunch of products out there. Like, why should I trust you? And so, a little easier time, probably now, that you know, with the CIOs wanting to cut costs, and now you have an excuse to go back to the executive team and say, “Look, I made you look smart. We get to [unintelligible 00:17:35], you know, our systems can scale to this.” That’s easy. Or the second one is we do, you know, we’ll start with some side use case or a greenfield. But both exists, and I would say 65% is probably rip-outs.

Corey: One question, I love to, I wouldn’t call it ambush, but definitely come up with, the catches some folks by surprise is one of the ways I like to sort out zealots from people who are focused on business problems. Do have an example of a data streaming workload for which Redpanda would not be a great fit?

Alex: Yeah. Database-style queries are not a fit. And so, think that there was a streaming engine before there was trying to build a database on top of it, and, like—and probably it does work in some low volume of traffic, like, say 5, 10 megabytes per second, but when you get to actual large scale, it just it doesn’t work. And it doesn’t work because but what Redpanda is, it gives you two properties as a developer. You can add data to the end or you can truncate the head, right?

And so, because those are your only two operations on the log, then you have to build this entire caching level to be able to give this database semantics. And so, do you know, I think for that the future isn’t for us to build a database, just as an example, it’s really to almost invert it. It’s like, hey, what if we make our format an open format like Apache Iceberg and then bring in your favorite database? Like, bring in, you know, Snowflake or Athena or Trina or Spark or [unintelligible 00:18:54] or [unintelligible 00:18:55] or whatever the other [unintelligible 00:18:56] of great databases that are better than we are, and doing, you know, just MPP, right, like a massively parallelizable database, do that, and then the job for us, for [unintelligible 00:19:05], let me just structure your log in a way that allows you to query, right? And so, for us, when we announced the $100 million dollar Series C funding, it’s like, I’m going to put the data in an iceberg format so you can go and query it with the other ten databases. And there are a better job than we are at that than we are.

Corey: It’s frankly, refreshing to see a vendor that knows where, okay, this is where we start and this is where we stop because it just seems that there’s been an industry-wide push for a while now to oh, you built a component in a larger system that works super well. Now, expand to do everything else in the architectural diagram. And you suddenly have databases trying to be network transport layers and queues trying to be data warehouses, and it just doesn’t work that way. It just it feels like oh, this is a terrible approach to solving this particular problem. And what’s worse, from my mind, is that people who hadn’t heard of you before look at you through this lens that does not put you in your best light, and, “Oh, this is a terrible database.” Well, it’s not supposed to be one.

Alex: [laugh].

Corey: But it also—it puts them off as a result. Have you faced pressure to expand beyond your core competency from either investors or customers or analysts or, I don’t know, the voices late at night that I hear and I assume everyone else does, too?

Alex: Exactly. The 3 a.m. voice that I have to take my phone and take a voice note because it’s like, I don’t want to lose this idea. Totally. For us. I think there’s pressures, like, hey, you built this great engine. Why don’t you add, like, the latest, you know, soup de jour in systems was like a vector database.

I was like, “This doesn’t even make any sense.” For me, it’s, I want to do one thing really well. And I generally call it internally, ‘the ring zero.’ It’s, if you think of the internet, right, like, as a computer, especially with this mode to what we talked about earlier in a BYOC, like, we could be the best ring zero, the best sort of like, you know, messaging platform for people to build real-time applications. And then that’s the case and there’s just so much low-hanging fruit for us.

Like, the developer experience wasn’t great for other systems, like, why don’t we focus on the last mile, like, making that developer, you know, successful at doing this one thing as opposed to be an average and a bunch of other a hundred products? And until we feel, honestly, that we’ve done a phenomenal job at that—I think we still have some roadmap to get there—I don’t want to expand. And, like, if there’s pressure, my answer is, like… look, the market is big enough. We don’t have to do it. We’re still, you know, growing.

I think it’s obviously not trivial and I’m kind of trivializing a bunch of problems from a business perspective. I’m not trying to degrade anyone else. But for us, it’s just being focused. This is what we do well. And bring every other technology that makes you successful. I don’t really care. I just want to make this part well.

Corey: I think that that is something that’s under-appreciated. I feel like I should get over at one point to something that’s been nagging at the back of my mind. Some would call it a personal attack and I suppose I’ll let them, but what I find interesting is your background. Historically, you were a distributed systems engineer at very large scale. And you apparently wrote the first version of Redpanda yourself in—was it C or C++?

Alex: C++.

Corey: Yeah. And now you are the CEO of a company that is clearly doing very well. Have you gotten the hell out of production yet? The reason I ask this is I have worked in a number of companies where the founder was also the initial engineer and then they invariably treated main as their feature branch and the rest of us all had to work around them to keep them from, you know, destroying everything we were trying to build around us, due to missing context. In other words, how annoyed with you are your engineers on any given afternoon?

Alex: [laugh]. Yeah. I would say that as a company builder now, if I may say that, is the team is probably the thing I’m the most proud of. They’re just so talented, such good [unintelligible 00:22:47] of humans. And so—group of humans—I stopped coding about two years ago, roughly.

So, the company is four-and-a-half years old, really the first two-and-a-half years old, the first one, two years, definitely, I was personally putting in, like, tons and tons of hours working on the code. It was a ton of fun. To me, one of the most rewarding technical projects I’ve ever had a chance to do. I still read pull requests, though, just so that when I have a conversation with a technical leader, I don’t be, like, I have no clue how the transactions work. So, I still have to read the code, but I don’t write any more code and my heart was a little broken when my dev prod team removed my write access to the GitHub repo.

We got SOC2 compliance, and they’re like, “You can’t have access to being an admin on Google domains, and you’re no longer able to write into main.” And so, I think as a—I don’t know, maybe my identity—myself identity is that of a builder, and I think as long as I personally feel like I’m building, today, it’s not code, but you know, is the company and [unintelligible 00:23:41] sort of culture, then I feel okay [laugh]. But yeah, I no longer write code. And the last story on that, is this—an engineer of ours, his name is [Stefan 00:23:51], he’s like, “Hey, so Alex wrote this semaphore”—this was actually two days ago—and so they posted a video, and I commented, I was like, “Hey, this was the context of semaphore. I’m sorry for this bug I caused.” But yeah, at least I still remember some context for them.

Corey: What’s fun is watching things continue to outpace and outgrow you. I mean, one of the hard parts of building a company is the realization that every person you hire for a thing that’s now getting off of your plate is better at that thing than you are. It’s a constant experience of being humbled. And at some point, things wind up outpacing you to the point where, at least in my case, I’ve been on calls with customers and I explained how we did some things and how it worked and had to be corrected by my team of, “Well. That used to be true, however…” like, “Oh, dear Lord. I’m falling behind.” And that’s always been a weird feeling for me.

Alex: Totally. You know, it’s the feeling of being—before I think I became a CEO, I was a highly comped engineer and did a competent, to the extent that it allowed me to build this product. And then you start doing all of these things and you’re incompetent, obviously, by definition because you haven’t done those things and so there’s like that discomfort [laugh]. But I have to get it done because no one else wants to do, whatever, like say, like, you know, rev ops or marketing or whatever.

And then you find somebody who’s great and you’re like, oh my God, I was like, I was so poor tactically at doing this thing. And it’s definitely humbling every day. And it’s almost it’s, like, gosh, you’re just—this year was kind of this role where you’re just, like, mediocre at, like, a whole lot of things as a company, but you’re the only person that has to do the job because you have the context and you just have to go and do it. And so, it’s definitely humbling. And in some ways, I’m learning, so for me today, it’s still a lot of fun to learn.

Corey: This is a little more in the weeds, I suppose, but I always love to ask people these questions. Because I used to be naive, which meant that I had hope and I saw a brighter future in technology. I now know that was all a lie. But I used to believe that out there was some company whose internal infrastructure for what they’d built was glorious and it would be amazing. And I knew I would never work there, nor what I want to, because when everything’s running perfectly, all I can really do is mess that up; there’s no way to win and a bunch of ways to lose.

But I found that place doesn’t exist. Every time I talk to someone about how they built the thing that they built and I ask them, “If you were starting over from scratch, what would you do differently?” The answer often distills down to, “Oh, everything.” Because it’s an organically evolving system that oh, yeah, everything’s easier the second time. At least you get to find new failure modes go in that way. When you look back at how you designed it originally, are there any missteps that you could have saved yourself a whole lot of grief by not making the first time?

Alex: Gosh, so many things. But if I were to give Hollywood highlights on these things, something that [unintelligible 00:26:35] is, does well is exposing these high-level data types of, like, streams, and lists and maps and et cetera. And I was like, “Well, why couldn’t streams offer this as a first-class citizen?” And we got some things well which I think would still do, like the whole [thread recorder 00:26:49] could—like, the fundamentals of the engine I will still do the same. But, you know, exposing new programming models earlier in the life of the product, I think would have allowed us to capture even more wildly different use cases.

But now we kind of have this production engine, we have to support Fortune 2000, so you know, it’s kind of like a very delicate evolution of the product. Definitely would have changed—I would have added, like, custom data types upfront, I would have pushed a little harder on I think WebAssembly than we did originally. Man, I could just go on for—like, [added detail 00:27:21], I would definitely have changed things. Like, I would have pressed on the first—on the version of the cloud that we talked about early on, that as the first deployment mode. If we go back through the stack of all of the products you had, it’s funny, like, 11 products that are surfaced to the customers to, like, business lines, I would change fundamental things about just [laugh], you know, everything else. I think that’s maybe the curse of the expert. Like, you know, you could always find improvements.

Corey: Oh, always. I still look back at my career before starting this place when I was working in a bunch of finance companies, and—I’ll never forget this; it was over a decade ago—we were building out our architecture in AWS, and doing a deal with a large finance company. And they said, “Cool, where’s your data center?” And I said, “Oh, it’s AWS.” And they said, “Ha ha ha ha. Where’s your data center?”

And that was oh, okay, great. Now, it feels like if that’s their reaction, they have not kept pace with the times. It feels it is easier to go to a lot of very serious enterprises with very serious businesses and serious workload concerns attendant to those and not get laughed out of the room because you didn’t wind up doing a multi-million dollar data center build out that, with an eye toward making it look as enterprise-y as possible.

Alex: Yeah. Okay, so here’s, I think, maybe something a little bit controversial. I think that’s true. People are moving to the cloud, and I don’t think that that idea, especially when we go when we talk to banks, is true. They’re like, “Hey, I have this contract with one of the hybrid clouds.”—you know, it’s usually with two of them, and then you’re like—“This is my workload. I want to spend $70 million or $100 million. Who could give me the biggest discount?” And then you kind of shop it around.

But what we are seeing is that effectively, the data transfer costs are so expensive and running this for so much this large volume of traffic is still so, so expensive, that there is an inverse [unintelligible 00:29:09] to host from some category of the workload where you don’t have dynamism. Actually hosted in your data center is, like, a huge boom in terms of cost efficiencies for the companies, especially where we are and especially in finances—you mentioned that—if you’re trying to trade and you have this, like, steady state line from nine to five, whatever, eight to four, whenever the markets open, it’s actually relatively cost-efficient because you can measure hey, look, you know, the New York Stock Exchange is 1.5 gigabytes per second at market close. Like, I could provision my hardware to beat this. And like, it’ll be that I don’t need this dynamism that the cloud gives me.

And so yeah, it’s kind of fascinating that for us because we offered the self-hosted Redpanda which can adapt to super low latencies with kernel parameter tuning, and the cloud due to the tiered storage, we talked about S3 being [unintelligible 00:29:52] to, so it’s been really fun to participate in deployments where we have both. And you couldn’t—they couldn’t look more different. I mean, it’s almost looks like two companies.

Corey: One last question before we wind up calling it an episode. I think I saw something fly by on Twitter a while back as I slowly returned to the platform—no, I’m not calling it X—something you’re doing involving a scholarship. Can you tell me a bit more about that?

Alex: Yeah. So, you know, I’m a Latino CEO, first generation in the States, and some of the things that I felt really frustrated with, growing up that, like, I feel fortunate because I got to [unintelligible 00:30:25] that is that, you know, people were just—that look like me are probably given some bullshit QA jobs, so like, you know, behemoth job, I think, for a bank. And so, I wanted to change that. And so, we give money and mentorship to people and we release all of the intellectual property. And so, we mentor someone—actually, anyone from underrepresented backgrounds—for three months.

We give then, like, 1200 bucks a month—or 1500, I can’t remember—mentorship from our top principal level engineers that have worked at Amazon and Google and Facebook and basically the world’s top companies. And so, they meet with them one hour a week, we give them money, they could sit in the couch if they want to. No one has to [unintelligible 00:31:06]. And all we’re trying to do is, like, “Hey, if you are part of this group, go and try to build something super hard.” [laugh].

And often their minds, which is great, and they’re like, “I want to build an OpenAI competitor in three months, and here’s the week-by-week progress.” Or, “I want to build a new storage engine, new database in three months.” And that’s the kind of people that we want to help, these like, super ambitious, that just hasn’t had a chance to be mentored by some of the world’s best engineers. And I just want to help them. Like, we—this is a non-scalable project. I meet with them once a week. I don’t want to have a team of, like, ten people.

Like, to me, I feel like their most valuable thing I could do is to give them my time and to help them mentor. I was like, “Hey, let’s think about this problem. Let’s decompose this. How do you think about this?” And then bring you the best engineers that I, you know, that work for—with me, and let me help you think about problems differently and give you some money.

And we just don’t care how you use the time or the money; we just want people to work on hard problems. So, it’s active. It runs once a year, and if anyone is listening to this, if you want to send it to your friends, we’d love to have that application. It’s for anyone in the world, too, as long as we can send the person a check [laugh]. You know, my head of finance is not going to walk to a Moneygram—which we have done in the past—but other than that, as long as you have a bank account that we can send the check to, you should be able to apply.

Corey: That is a compelling offer, particularly in the current macro environment that we find ourselves faced in. We’ll definitely put a link to that into the [show notes 00:32:32]. I really want to thank you for taking the time to, I guess, get me up to speed on what it is you’re doing. If people want to learn more where’s the best place for them to go?

Alex: On Twitter, my handle is @emaxerrno, which stands for the largest error in the kernel. I felt like that was apt for my handle. So, that’s one. Feel free to find me on the community Slack. There’s a Slack button on the website redpanda.com on the top right. I’m always there if you want to DM me. Feel free to stop by. And yeah, thanks for having me. This was a lot of fun.

Corey: Likewise. I look forward to the next time. Alex Gallego, CEO and founder at Redpanda. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an insulting comment that I will almost certainly never read because they have not figured out how to get data from one place to another.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

View Details

Kelsey Hightower joins Corey on Screaming in the Cloud to discuss his reflections on how the tech industry is progressing. Kelsey describes what he’s been getting out of retirement so far, and reflects on what he learned throughout his high-profile career - including why feature sprawl is such a driving force behind the complexity of the cloud environment and the tactics he used to create demos that are engaging for the audience. Corey and Kelsey also discuss the importance of remaining authentic throughout your career, and what it means to truly have an authentic voice in tech.

About Kelsey

Kelsey Hightower is a former Distinguished Engineer at Google Cloud, the co-chair of KubeCon, the world’s premier Kubernetes conference, and an open source enthusiast. He’s also the co-author of Kubernetes Up & Running: Dive into the Future of Infrastructure. Recently, Kelsey announced his retirement after a 25-year career in tech.

Links Referenced:

  • Twitter: https://twitter.com/kelseyhightower

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Do you wish there were cheat codes for database optimization? Well, there are – no seriously. If you’re using Postgres or MySQL on Amazon Aurora or RDS, OtterTune uses AI to automatically optimize your knobs and indexes and queries and other bits and bobs in databases. OtterTune applies optimal settings and recommendations in the background or surfaces them to you and allows you to do it. The best part is that there’s no cost to try it. Get a free, thirty-day trial to take it for a test drive. Go to ottertune dot com to learn more. That’s O-T-T-E-R-T-U-N-E dot com.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. You know, there’s a great story from the Bible or Torah—Old Testament, regardless—that I was always a big fan of where you wind up with the Israelites walking the desert for 40 years in order to figure out what comes next. And Moses led them but could never enter into what came next. Honestly, I feel like my entire life is sort of going to be that direction. Not the biblical aspects, but rather always wondering what’s on the other side of a door that I can never cross, and that door is retirement. Today I’m having returning guest Kelsey Hightower, who is no longer at Google. In fact, is no longer working and has joined the ranks of the gloriously retired. Welcome back, and what’s it like?

Kelsey: I’m happy to be here. I think retirement is just like work in some ways: you have to learn how to do it. A lot of people have no practice in their adult life what to do with all of their time. We have small dabs in it, like, you get the weekend off, depending on what your work, but you never have enough time to kind of unwind and get into something else. So, I’m being honest with myself. It’s going to be a learning curve, what to do with that much time.

You’re probably still going to do work, but it’s going to be a different type of work than you’re used to. And so, that’s where I am. 30 days into this, I’m in that learning mode, I’m on-the-job training.

Corey: What’s harder than you expected?

Kelsey: It’s not the hard part because I think mentally I’ve been preparing for, like, the last ten years, being a minimalist, learning how to kind of live within my means, learn to appreciate things that are just not work-related or status symbols. And so, to me, it felt like a smooth transition because I started to value my time more than anything else, right? Just waking up the next day became valuable to me. Spending time in the moment, right, you go to these conferences, there’s, like, 10,000 people, but you learn to value those one-on-one encounters, those one-off, kind of, let’s just go grab lunch situations. So, to me, retirement just makes more room for that, right? I no longer have this calendar that is super full, so I think for me, it was a nice transition in terms of getting more of that valuable time back.

Corey: It seems to me that you’re in a similar position to the one that I find myself in where the job that you were doing and I still am is tied, more or less, to a sense of identity as opposed to a particular task or particular role that you fill. You were Kelsey Hightower. That was a complete sentence. People didn’t necessarily need to hear the rest of what you were working on or what you were going to be talking about at a given conference or whatnot. So, it seemed, at least from the outside, that an awful lot of what you did was quite simply who you were. Do you feel that your sense of identity has changed?

Kelsey: So, I think when you have that much influence, when you have that much reputation, the words you say travel further, they tend to come with a little bit more respect, and so when you’re working with a team on new product, and you say, “Hey, I think we should change some things.” And when they hear those words coming from someone that they trust or has a name that is attached to reputation, you tend to be able to make a lot of impact with very few words. But what you also find is that no matter what you get involved in—configuration management, distributed systems, serverless, working with customers—it all is helped and aided by the reputation that you bring into that line of work. And so yes, who you are matters, but one thing that I think helped me, kind of greatly, people are paying attention maybe to the last eight years of my career: containers, Kubernetes, but my career stretches back to the converting COBOL into Python days; the dawn of DevOps, Puppet, Chef, and Ansible; the Golang appearance and every tool being rewritten from Ruby to Golang; the Docker era.

And so, my identity has stayed with me throughout those transitions. And so, it was very easy for me to walk away from that thing because I’ve done it three or four times before in the past, so I know who I am. I’ve never had, like, a Twitter bio that said, “Company X. X person from company X.” I’ve learned long ago to just decouple who I am from my current employer because that is always subject to change.

Corey: I was fortunate enough to not find myself in the public eye until I owned my own company. But I definitely remember times in my previous incarnations where I was, “Oh, today I’m working at this company,” and I believed—usually inaccurately—that this was it. This was where I really found my niche. And then surprise I’m not there anymore six months later for, either their decision, my decision, or mutual agreement. And I was always hesitant about hanging a shingle out that was tied too tightly to any one employer.

Even now, I was little worried about doing it when I went independent, just because well, what if it doesn’t work? Well, what if, on some level? I think that there’s an authenticity that you can bring with you—and you certainly have—where, for a long time now, whenever you say something, I take it seriously, and a lot of people do. It’s not that you’re unassailably correct, but I’ve never known you to say something you did not authentically believe in. And that is an opinion that is very broadly shared in this industry. So, if nothing else, you definitely were a terrific object lesson in speaking the truth, as you saw it.

Kelsey: I think what you describe is one way that, whether you’re an engineer doing QA, working in the sales department, when you can be honest with the team you’re working with, when you can be honest with the customers you’re selling into when you can be honest with the community you’re part of, that’s where the authenticity gets built, right? Companies, sometimes on the surface, you believe that they just want you to walk the party line, you know, they give you the lines and you just read them verbatim and you’re doing your part. To be honest, you can do that with the website. You can do that with a well-placed ad in the search queries.

What people are actually looking for are real people with real experiences, sharing not just fact, but I think when you mix kind of fact and opinion, you get this level of authenticity that you can’t get just by pure strategic marketing. And so, having that leverage, I remember back in the day, people used to say, “I’m going to do the right thing and if it gets me fired, then that’s just the way it’s going to be. I don’t want to go around doing the wrong thing because I’m scared I’m going to lose my job.” You want to find yourself in that situation where doing the right thing, is also the best thing for the company, and that’s very rare, so when I’ve either had that opportunity or I’ve tried to create that opportunity and move from there.

Corey: It resonates and it shows. I have never had a lot of respect for people who effectively are saying one thing today and another thing the next week based upon which way they think that the winds are blowing. But there’s also something to be said for being able and willing to publicly recant things you have said previously as technology evolves, as your perspective evolves and, in light of new information, I’m now going to change my perspective on something. I’ve done that already with multi-cloud, for example. I thought it was ridiculous when I heard about it. But there are also expressions of it that basically every company is using, including my own. And it’s a nuanced area. Where I find it challenging is when you see a lot of these perspectives that people are espousing that just so happen to deeply align with where their paycheck comes from any given week. That doesn’t ring quite as true to me.

Kelsey: Yeah, most companies actually don’t know how to deal with it either. And now there has been times at any number of companies where my authentic opinion that I put out there is against party line. And you get those emails from directors and VPs. Like, “Hey, I thought we all agree to think this way or to at least say this.” And that’s where you have to kind of have that moment of clarity and say, “Listen, that is undeniably wrong. It’s so wrong in fact that if you say this in public, whether a small setting or large setting, you are going to instantly lose credibility going forward for yourself. Forget the company for a moment. There’s going to be a situation where you will no longer be effective in your job because all of your authenticity is now gone. And so, what I’m trying to do and tell you is don’t do that. You’re better off saying nothing.”

But if you go out there, and you’re telling what is obviously misinformation or isn’t accurate, people are not dumb. They’re going to see through it and you will be classified as a person not to listen to. And so, I think a lot of people struggle with that because they believe that enterprise's consensus should also be theirs.

Corey: An argument that I made—we’ll call it a prediction—four-and-a-half years ago, was that in five years, nobody would really care about Kubernetes. And people misunderstood that initially, and I’ve clarified since repeatedly that I’m not suggesting it’s going away: “Oh, turns out that was just a ridiculous fever dream and we’re all going back to running bare metal with our hands again,” but rather that it would slip below the surface-level of awareness. And I don’t know that I got the timing quite right on that, I think it’s going to depend on the company and the culture that you find yourself in. But increasingly, when there’s an application to run, it’s easy to ask someone just, “Oh, great. Where’s the Kubernetes cluster live so we can throw this on there and just add it to the rest of the pile?”

That is sort of what I was seeing. My intention with that was not purely just to be controversial, as much fun as that might be, but also to act as a bit of a warning, where I’ve known too many people who let their identities become inextricably tangled with the technology. But technologies rise and fall, and at some point—like, you talk about configuration management days; I learned to speak publicly as a traveling trainer for Puppet. I wrote part of SaltStack once upon a time. But it was clear that that was not the direction the industry was going, so it was time to find something else to focus on. And I fear for people who don’t keep an awareness or their feet underneath them and pay attention to broader market trends.

Kelsey: Yeah, I think whenever I was personally caught up in linking my identity to technology, like, “I’m a Rubyist,” right?“, I’m a Puppeteer,” and you wear those names proudly. But I remember just thinking to myself, like, “You have to take a step back. What’s more important, you or the technology?” And at some point, I realized, like, it’s me, that is more important, right? Like, my independent thinking on this, my independent experience with this is far more important than the success of this thing.

But also, I think there’s a component there. Like when you talked about Kubernetes, you know, maybe being less relevant in five years, there’s two things there. One is the success of all infrastructure things equals irrelevancy. When flights don’t crash, when bridges just work, you do not think about them. You just use them because they’re so stable and they become very boring. That is the success criteria.

Corey: Utilities. No one’s wondering if the faucet’s going to work when they turn it on in the morning.

Kelsey: Yeah. So, you know, there’s a couple of ways to look at your statement. One is, you believe Kubernetes is on the trajectory that it’s going to stabilize itself and hit that success criteria, and then it will be irrelevant. Or there’s another part of the irrelevancy where something else comes along and replaces that thing, right? I think Cloud Foundry and Mesos are two good examples of Kubernetes coming along and stealing all of the attention from that because those particular products never gained that mass adoption. Maybe they got to the stable part, but they never got to the mass adoption part. So, I think when it comes to infrastructure, it’s going to be irrelevant. It’s just what side of that [laugh] coin do you land on?

Corey: It’s similar to folks who used to have to work at a variety of different companies on very specific Linux kernel subsystems because everyone had to care because there were significant performance impacts. Time went on and now there’s still a few of those people that very much need to care, but for the rest of us, it is below the level of things that we have to care about. For me, the signs of the unsustainability were, oh, you can run Kubernetes effectively in production? That’s a minimum of a quarter-million dollars a year in comp or up in some cases. Not every company is going to be able to field a team of those people and still remain a going concern in business. Nor frankly, should they have to.

Kelsey: I’m going to pull on that thread a little bit because it’s about—we’re hitting that ten-year mark of Kubernetes. So, when Kubernetes comes out, why were people drawn to it, right? Why did it even get the time of day to begin with? And I think Docker kind of opened Pandora’s box there. This idea of Chef, Puppet, Ansible, ten thousand package managers, and honestly, that trajectory was going to continue forever and it was helping no one. It was literally people doing duplicate work depending on the operating system you’re dealing with and we were wasting time copying bits to servers—literally—in a very glorified way.

So, Docker comes along and gives us this nicer, better abstraction, but it has gaps. It has no orchestration. It’s literally this thing where now we’ve unified the packaging situation, we’ve learned a lot from Red Hat, YUM, Debian, and the various package repo combinations out there and so we made this universal thing. Great. We also learned a little bit about orchestration through brute force, bash scripts, config management, you name it, and so we serialized that all into this thing we call Kubernetes.

It’s pretty simple on the surface, but it was probably never worthy of such fanfare, right? But I think a lot of people were relieved that now we finally commoditized this expertise that the Googles, the Facebooks of the world had, right, building these systems that can copy bits to other systems very fast. There you go. We’ve gotten that piece. But I think what the market actually wants is in the mobile space, if you want to ship software to 300 million people that you don’t even know, you can do it with the app store.

There’s this appetite that the boring stuff should be easy. Let’s Encrypt has made SSL certificates beyond easy. It’s just so easy to do the right thing. And I think for this problem we call deployments—you know, shipping apps around—at some point we have to get to a point where that is just crazy easy. And it still isn’t.

So, I think some of the frustration people express ten years later, they’re realizing that they’re trying to recreate a Rube Goldberg machine with Kubernetes is the base element and we still haven’t understood that this whole thing needs to simplify, not ten thousand new pieces so you can build your own adventure.

Corey: It’s the idea almost of what I’m seeing AWS go through, and to some extent, its large competitors. But building anything on top of AWS from scratch these days is still reminiscent of going to Home Depot—or any hardware store—and walking up and down the aisles and getting all the different components to piece together what you want. Sometimes just want to buy something from Target that’s already assembled and you have to do all of that work. I’m not saying there isn’t value to having a Home Depot down the street, but it’s also not the panacea that solves for all use cases. An awful lot of customers just want to get the job done and I feel that if we cling too tightly to how things used to be, we lose it.

Kelsey: I’m going to tell you, being in the cloud business for almost eight years, it’s the customers that create this. Now, I’m not blaming the customer, but when you start dealing with thousands of customers with tons of money, you end up in a very different situation. You can have one customer willing to pay you a billion dollars a year and they will dictate things that apply to no one else. “We want this particular set of features that only we will use.” And for a billion bucks a year times ten years, it’s probably worth from a business standpoint to add that feature.

Now, do this times 500 customers, each major provider. What you end up with is a cloud console that is unbearable, right? Because they also want these things to be first-class citizens. There’s always smaller companies trying to mimic larger peers in their segment that you just end up in that chaos machine of unbound features forever. I don’t know how to stop it. Unless you really come out maybe more Apple style and you tell people, “This is the one and only true way to do things and if you don’t like it, you have to go find an alternative.” The cloud business, I think, still deals with the, “If you have a large payment, we will build it.”

Corey: I think that that is a perspective that is not appreciated until you’ve been in the position of watching how large enterprises really interact with each other. Because it’s, “Well, what customer the world is asking for yet another way to run containers?” “Uh, this specific one and their constraints are valid.” Every time I think I’ve seen everything there is to see in the world of cloud, I just have to go talk to one more customer and I’m learning something new. It’s inevitable.

I just wish that there was a better way to explain some of this to newcomers, when they’re looking at, “Oh, I’m going to learn how this cloud thing works. Oh, my stars, look at how many services there are.” And then they wind up getting lost with analysis paralysis, and every time they get started and ask someone for help, they’re pushed in a completely different direction and you keep spinning your wheels getting told to start over time and time again when any of these things can be made to work. But getting there is often harder than it really should be.

Kelsey: Yeah. I mean, I think a lot of people don’t realize how far you can get with, like, three VMs, a load balancer, and Postgres. My guess is you can probably build pretty much any clone of any service we use today with at least 1 million customers. Most people never reached that level—I don’t even want to say the word scale—but that blueprint is there and most people will probably be better served by that level of simplicity than trying to mimic the behaviors of large customers—or large companies—with these elaborate use cases. I don’t think they understand the context there. A lot of that stuff is baggage. It’s not [laugh] even, like, best-of-breed or great design. It’s like happenstance from 20 years of trying to buy everything that’s been sold to you.

Corey: I agree with that idea wholeheartedly. I was surprising someone the other day when I said that if you were to give me a task of getting some random application up and running by tomorrow, I do a traditional three-tier architecture, some virtual machines, a load balancer, and a database service. And is that the way that all the cool kids are doing it today? Well, they’re not talking about it, but mostly. But the point is, is that it’s what I know, it’s where my background is, and the thing you already know when you’re trying to solve a new problem is incredibly helpful, rather than trying to learn everything along that new path that you’re forging down. Is that architecture the best approach? No, but it’s perfectly sufficient for an awful lot of stuff.

Kelsey: Yeah. And so, I mean, look, I’ve benefited my whole career from people fantasizing about [laugh] infrastructure—

Corey: [laugh].

Kelsey: And the truth is that in 2023, this stuff is so powerful that you can do almost anything you want to do with the simplest architecture that’s available to us. The three-tier architecture has actually gotten better over the years. I think people are forgotten: CPUs are faster, RAM is much bigger quantities, the networks are faster, right, these databases can store more data than ever. It’s so good to learn the fundamentals, start there, and worst case, you have a sound architecture people can reason about, and then you can go jump into the deep end, once you learn how to swim.

Corey: I think that people would be depressed to understand just how much the common case for the value that Kubernetes brings is, “Oh yeah, now we can lose a drive or a server and the application stays up.” It feels like it’s a bit overkill for that one somewhat paltry use case, but that problem has been hounding companies for decades.

Kelsey: Yeah, I think at some point, the whole ‘SSH is my only interface into these kinds of systems,’ that’s a little low level, that’s a little bare bones, and there will probably be a feature now where we start to have this not Infrastructure as Code, not cloud where we put infrastructure behind APIs and you pay per use, but I think what Kubernetes hints at is a future where you have APIs that do something. Right now the APIs give you pieces so you can assemble things. In the future, the APIs will just do something, “Run this app. I need it to be available and here’s my money budget, my security budget, and reliability budget.” And then that thing will say, “Okay, we know how to do that, and here’s roughly what is going to cost.”

And I think that’s what people actually want because that’s how requests actually come down from humans, right? We say, “We want this app or this game to be played by millions of people from Australia to New York.” And then for a person with experience, that means something. You kind of know what architecture you need for that, you know what pieces that need to go there. So, we’re just moving into a realm where we’re going to have APIs that do things all of a sudden.

And so, Kubernetes is the warm-up to that era. And that’s why I think that transition is a little rough because it leaks the pieces part, so where you can kind of build all the pieces that you want. But we know what’s coming. Serverless also hints at this. But that’s what people should be looking for: APIs that actually do something.

Corey: This episode is sponsored in part by Panoptica. Panoptica simplifies container deployment, monitoring, and security, protecting the entire application stack from build to runtime. Scalable across clusters and multi-cloud environments, Panoptica secures containers, serverless APIs, and Kubernetes with a unified view, reducing operational complexity and promoting collaboration by integrating with commonly used developer, SRE, and SecOps tools. Panoptica ensures compliance with regulatory mandates and CIS benchmarks for best practice conformity. Privacy teams can monitor API traffic and identify sensitive data, while identifying open-source components vulnerable to attacks that require patching. Proactively addressing security issues with Panoptica allows businesses to focus on mitigating critical risks and protecting their interests. Learn more about Panoptica today at panoptica.app.

Corey: You started the show by talking about how your career began with translating COBOL into Python. I firmly believe someone starting their career today listening to this could absolutely find that by the time their career starts drawing to their own close, that Kubernetes is right in there as far as sounding like the deprecated thing that no one really talks about or thinks about anymore. And I hope so. I want the future to be brighter than the past. I want getting a business or getting software together in a way that helps people to not require the amount of, “First, spend six weeks at a boot camp,” or, “Learn how to write just enough code that you can wind up getting funding and then have it torn apart.”

What’s the drag-and-drop story? What’s the describe the application to a robot and it builds it for you? I’m optimistic about the future of infrastructure, just because based upon its power to potentially make reliability and scale available to folks who have no idea of what’s involved with that. That’s kind of the point. That’s the end game of having won this space.

Kelsey: Well, you know what? Kubernetes is providing the metadata to make that possible, right? Like in the early days, people were writing one-off scripts or, you know, writing little for loops to get things in the right place. And then we get config management that kind of formalizes that, but it still had no metadata, right? You’d have things like Puppet report information.

But in the world of, like, Kubernetes, or any cloud provider, now you get semantic meaning. “This app needs this volume with this much space with this much memory, I need three of these behind this load balancer with these protocols enabled.” There is now so much metadata about applications, their life cycles, and how they work that if you were to design a new system, you can actually use that data to craft a much better API that made a lot of this boilerplate the defaults. Oh, that’s a web application. You do not need to specify all of this boilerplate. Now, we can give you much better nouns and verbs to describe what needs to happen.

So, I think this is that transition as all the new people coming up, they’re going to be dealing with semantic meaning to infrastructure, where we were dealing with, like, tribal knowledge and intuition, right? “Run this script, pipe it to this thing, and then this should happen. And if it doesn’t, run the script again with this flag.” Versus, “Oh, here’s the semantic meaning to a working system.” That’s a game-changer.

Corey: One other topic I wanted to ask you about—I’ve it’s been on my list of things to bring up the next time I ran into you and then you went ahead and retired, making it harder to run into you. But a little while back, I was at a tech conference and someone gave a demo, and it didn’t go as well as they had hoped. And a few of us were talking about it afterwards. We’ve all been speakers, we’ve all lived that life. Zero shade.

But someone brought you up in particular—unprompted; your legend does precede you—and the phrase that they used was that Kelsey’s demos were always picture-perfect. He was so lucky with how the demos worked out. And I just have to ask—because you don’t strike me as someone who is not careful, particularly when all eyes are upon you—and real experts make things look easy, did you have demos periodically go wrong that the audience just didn’t see going wrong along the way? Or did you just actually YOLO all of your demos and got super lucky every single time for the last eight years?

Kelsey: There was a musician who said, “Hey, your demos are like jazz. You improvise the whole thing.” There’s no script, there’s no video. The way I look at the demo is, like, you got this instrument, the command prompt, and the web browser. You can do whatever you want with them.

Now, I have working code. I wrote the code, I wrote the deployment scenarios, I delete it all and I put it all back. And so, I know how it’s supposed to work from the ground up. And so, what that means is if anything goes wrong, I can improvise. I could go into fixing the code. I can go into doing a redeploy.

And I’ll give you one good example. The first time Kubernetes came out, there was this small meetup in San Francisco with just the core contributors, right? So, there is no community yet, there’s no conference yet, just people hacking on Kubernetes. And so, we decided, we’re going to have the first Kubernetes meetup. And everyone got, like, six, seven minutes, max. That’s it. You got to move.

And so, I was like, “Hey, I noticed that in the lineup, there is no ‘What is Kubernetes?’ talk. We’re just getting into these nuts and bolts and I don’t think that’s fair to the people that will be watching this for the first time.” And I said, “All right, Kelsey, you should give maybe an intro to what it is.” I was like, “You know what I’ll do? I’m going to build a Kubernetes cluster from the ground up, starting with VMs on my laptop.”

And I’m in it and I’m feeling confident. So, confidence is the part that makes it look good, right? Where you’re confident in the commands you type. One thing I learned to do is just use your history, just hit the up arrow instead of trying to copy all these things out. So, you hit the up arrow, you find the right command and you talk through it and no one looks at what’s happening. You’re cycling through the history.

Or you have multiple tabs where you know the next up arrow is the right history. So, you give yourself shortcuts. And so, I’m halfway through this demo. We got three minutes left, and it doesn’t work. Like, VMware is doing something weird on my laptop and there’s a guy calling me off stage, like, “Hey, that’s it. Cut it now. You’re done.”

I’m like, “Oh, nope. Thou shalt not go out like this.” It’s time to improvise. And so, I said, “Hey, who wants to see me finish this?” And now everyone is locked in. It’s dead silent. And I blow the whole thing away. I bring up the VMs, I [pixie 00:28:20] boot, I installed the kubelet, I install Docker. And everyone’s clapping. And it’s up, it’s going, and I say, “Now, if all of this works, we run this command and it should start running the app.” And I do kubectl apply-f and it comes up and the place goes crazy.

And I had more to the demo. But you stop. You’ve gotten the point across, right? This is what Kubernetes is, here’s how it works, and look how you do it from scratch. And I remember saying, “And that’s the end of my presentation.” You need to know when to stop, you need to know when to pivot, and you need to have confidence that it’s supposed to work, and if you’ve seen it work a couple of times, your confidence is unshaken.

And when I walked off that stage, I remember someone from Red Hat was like—Clayton Coleman; that’s his name—Clayton Coleman walked up to me and said, “You planned that. You planned it to fail just like that, so you can show people how to go from scratch all the way up. That was brilliant.” And I was like, “Sure. That’s exactly what I did.”

Corey: “Yeah, I meant to do that.” I like that approach. I found there’s always things I have to plan for in demos. For example, I can never count on having solid WiFi from a conference hall. The show has to go on. It’s, okay, the WiFi doesn’t work. I’ve at one point had to give a talk where the projector just wasn’t working to a bunch of students. So okay, close the laptop. We’re turning this into a bunch of question-and-answer sessions, and it was one of the better talks I’ve ever given.

But the alternative is getting stuck in how you think a talk absolutely needs to go. Now, keynotes are a little harder where everything has been scripted and choreographed and at that point, I’ve had multiple fallbacks for demos that I’ve had to switch between. And people never noticed I was doing it for that exact reason. But it takes work to look polished.

Kelsey: I will tell you that the last Next keynote I gave was completely irresponsible. No dry runs, no rehearsals, no table reads, no speaker notes. And I think there were 30,000 people at that particular Next. And Diane Greene was still CEO, and I remember when marketing was like, “Yo, at least a backup recording.” I was like, “Nah, I don’t have anything.”

And that demo was extensive. I mean, I was building an app from scratch, starting with Postgres, adding the schema, building an app, deploying the app. And something went wrong halfway. And there’s this joke that I came up with just to pass over the time, they gave me a new Chromebook to do the demo. And so, it’s not mine, so none of the default settings were there, I was getting pop-ups all over the place.

And I came up with this joke on the way to the conference. I was like, “You know what’d be cool? When I show off the serverless stuff, I would just copy the code from Stack Overflow. That’d be like a really cool joke to say this is what senior engineers do.” And I go to Stack Overflow and it’s getting all of these pop-ups and my mouse couldn’t highlight the text.

So, I’m sitting there like a deer in headlights in front of all of these people and I’m looking down, and marketing is, like, “This is what… this is what we’re talking about.” And so, I’m like, “Man do I have to end this thing here?” And I remember I kept trying, I kept trying, and came to me. Once the mouse finally got in there and I cleared up all the popups, I just came up with this joke. I said, “Good developers copy.” And I switched over to my terminal and I took the text from Stack Overflow and I said, “Great developers paste,” and the whole room start laughing.

And I had them back. And we kept going and continued. And at the end, there was like this Google Assistant, and when it was finished, I said, “Thank you,” to the Google Assistant and it was talking back through the live system. And it said, “I got to admit, that was kind of dope.” So, I go to the back and Diane Greene walks back there—the CEO of Google Cloud—and she pats me on the shoulder. “Kelsey, that was dope.”

But it was the thrill because I had as much thrill as the people watching it. So, in real-time, I was going through all these emotions. But I think people forget, the demo is supposed to convey something. The demo is supposed to tell some story. And I’ve seen people overdo their demos with way too much code, way too many commands, almost if they’re trying to show off their expertise versus telling a story. And so, when I think about the demo, it has to complement the entire narrative. And so, sometimes you don’t need as many commands, you don’t need as much code. You can keep things simple and that gives you a lot more ins and outs in case something does go crazy.

Corey: And I think the key takeaway here that so many people lose sight of is you have to know the material well enough that whatever happens, well, things don’t always go the way I planned during the day, either, and talking through that is something that I think serves as a good example. It feels like a bit more of a challenge when you’re trying to demo something that a company is trying to sell someone, “Oh, yeah, it didn’t work. But that’s okay.” But I’m still reminded by probably one of the best conference demo fails I’ve ever seen on video. One day, someone was attempting to do a talk that hit Amazon S3 and it didn’t work.

And the audience started shouting at him that yeah, S3 is down right now. Because that was the big day that S3 took a nap for four hours. It was one of those foundational things you’d should never stop to consider. Like, well, what if the internet doesn’t work tomorrow when I’m doing my demo? That’s a tough one to work around. But rough timing.

Kelsey: [breathy sound]

Corey: He nailed the rest of the talk, though. You keep going. That’s the thing that people miss. They get stuck in the demo that isn’t working, they expect the audience knows as much as they do about what’s supposed to happen next. You’re the one up there telling a story. People forget it’s storytelling.

Kelsey: Now, I will be remiss to say, I know that the demo gods have been on my side for, like, ten, maybe fifteen years solid. So, I retired from doing live demos. This is why I just don’t do them anymore. I know I’m overdue as an understatement. But the thing I’ve learned though, is that what I found more impressive than the live demo is to be able to convey the same narratives through story alone. No slides. No demo. Nothing. But you can still make people feel where you would try to go with that live demo.

And it’s insanely hard, especially for technologies people have never seen before. But that’s that new challenge that I kind of set up for myself. So, if you see me at a keynote and you’ve noticed why I’ve been choosing these fireside chats, it’s mainly because I’m also trying to increase my ability to share narrative, technical concepts, but now in a new form. So, this new storytelling format through the fireside chat has been my substitute for the live demo, normally because I think sometimes, unless there’s something really to show that people haven’t seen before, the live demo isn’t as powerful to me. Once the thing is kind of known… the live demo is kind of more of the same. So, I think they really work well when people literally have never seen the thing before, but outside of that, I think you can kind of move on to, like, real-life scenarios and narratives that help people understand the fundamentals and the philosophy behind the tech.

Corey: An awful lot of tools and tech that we use on a day-to-day basis as well are thankfully optimized for the people using them and the ergonomics of going about your day. That is orthogonal, in my experience, to looking very impressive on stage. It’s the rare company that can have a product that not only works well but also presents well. And that is something I don’t tend to index on when I’m selecting a tool to do something with. So, it’s always a question of how can I make this more visually entertaining? For while I got out of doing demos entirely, just because talking about things that have more staying power than a screenshot that is going to wind up being irrelevant the next week when they decide to redo the console for some service yet again.

Kelsey: But you know what? That was my secret to doing software products and projects. When I was at CoreOS, we used to have these meetups we would used to do every two weeks or so. So, when we were building things like etcd, Fleet was a container management platform that came before Kubernetes, we would always run through them as a user, start install them, use them, and ask how does it feel? These command line flags, they don’t feel right. This isn’t a narrative you can present with the software alone.

But once we could, then the meetups were that much more engaging. Like hey, have you ever tried to distribute configuration to, like, a thousand servers? It’s insanely hard. Here’s how you do with Puppet. But now I’m going to show you how you do with etcd. And then the narrative will kind of take care of itself because the tool was positioned behind what people would actually do with it versus what the tool could do by itself.

Corey: I think that’s the missing piece that most marketing doesn’t seem to quite grasp is, they talk about the tool and how awesome it is, but that’s why I love customer demos so much. They’re showing us how they use a tool to solve a real-world problem. And honestly, from my snarky side of the world and the attendant perspective there, I can make an awful lot of fun about basically anything a company decides to show me, but put a customer on stage talking about how whatever they’ve built is solving a real-world problem for them, that’s the point where I generally shut up and listen because I’m going to learn something about a real-world story. Because you don’t generally get to tell customers to go on stage and just make up a story that makes us sound good, and have it come off with any sense of reality whatsoever. I haven’t seen that one happen yet, but I’m sure it’s out there somewhere.

Kelsey: I don’t know how many founders or people building companies listen in to your podcast, but this is right now, I think the number one problem that especially venture-backed startups have. They tend to have great technology—maybe it’s based off some open-source project—with tons of users who just know how that tool works, it’s just an ingredient into what they’re already trying to do. But that isn’t going to ever be your entire customer base. Soon, you’ll deal with customers who don’t understand the thing you have and they need more than technology, right? They need a product.

And most of these companies struggle painting that picture. Here’s what you can do with it. Or here’s what you can’t do now, but you will be able to do if you were to use this. And since they are missing that, a lot of these companies, they produce a lot of code, they ship a lot of open-source stuff, they raise a lot of capital, and then it just goes away, it fades out over time because they can bring on no newcomers. The people who need help the most, they don’t have a narrative for them, and so therefore, they’re just hoping that the people who have all the skills in the world, the early adopters, but unfortunately, those people are tend to be the ones that don’t actually pay. They just kind of do it themselves. It’s the people who need the most help.

Corey: How do we monetize the bleeding edge of adoption? In many cases you don’t. They become your community if you don’t hug them to death first.

Kelsey: Exactly.

Corey: Ugh. None of this is easy. I really want to thank you for taking the time to catch up and talk about how you seen the remains of a career well spent, and now you’re going off into that glorious sunset. But I have a sneaking suspicion you’ll still be around. Where should people go if they want to follow up on what you’re up to these days?

Kelsey: Right now I still use… I’m going to keep calling it Twitter.

Corey: I agree.

Kelsey: I kind of use that for my real-time interactions. And I’m still attending conferences, doing fireside chats, and just meeting people on those conference floors. But that’s what where I’ll be for now. So yeah, I’ll still be around, but maybe not as deep. And I’ll be spending more time just doing normal life stuff, maybe less building software.

Corey: And we will, of course, put a link to that in the show notes. Thank you so much for taking the time to catch up and share your reflections on how the industry is progressing.

Kelsey: Awesome. Thanks for having me, Corey.

Corey: Kelsey Hightower, now gloriously retired. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry comment that you’re going to type on stage as part of a conference talk, and then accidentally typo all over yourself while you’re doing it.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

View Details

Levi McCormick, Cloud Architect at Jamf, joins Corey on Screaming in the Cloud to discuss his work modernizing baseline cloud infrastructure and his experience being on the compliance side of cloud engineering. Levi explains how he works to ensure the different departments he collaborates with are all on the same page so that different definitions don’t end up in miscommunications, and why he feels a sandbox environment is an important tool that leads to a successful production environment. Levi and Corey also explore the ethics behind the latest generative AI craze.

About Levi

Levi is an automation engineer, with a focus on scalable infrastructure and rapid development. He leverages deep understanding of DevOps culture and cloud technologies to build platforms that scale to millions of users. His passion lies in helping others learn to cloud better.

Links Referenced:

  • Jamf: https://www.jamf.com/
  • Twitter: https://twitter.com/levi_mccormick
  • LinkedIn: https://www.linkedin.com/in/levimccormick/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. A longtime friend and person has been a while since he’s been on the show, Levi McCormick has been promoted or punished for his sins, depending upon how you want to slice that, and he is now the Director of Cloud Engineering at Jamf. Levi, welcome back.

Levi: Thanks for having me, Corey.

Corey: I have to imagine internally, you put that very pronounced F everywhere, and sometimes where it doesn’t belong, like your IAMf policies and whatnot.

Levi: It is fun to see how people like to interpret how to pronounce our name.

Corey: So, it’s been a while. What were you doing before? And how did you wind up stumbling your way into your current role?

Levi: [laugh]. When we last spoke, I was a cloud architect here, diving into just our general practices and trying to shore up some of them. In between, I did a short stint as director of FedRAMP. We are pursuing some certifications in that area and I led, kind of, the engineering side of the compliance journey.

Corey: That sounds fairly close to hell on earth from my particular point of view, just because I’ve dealt in the compliance side of cloud engineering before, and it sounds super interesting from a technical level until you realize just how much of it revolves around checking the boxes, and—at least in the era I did it—explaining things to auditors that I kind of didn’t feel I should have to explain to an auditor, but there you have it. Has the state of that world improved since roughly 2015?

Levi: I wouldn’t say it has improved. While doing this, I did feel like I drove a time machine to work, you know, we’re certifying VMs, rather than container-based architectures. There was a lot of education that had to happen from us to auditors, but once they understood what we were trying to do, I think they were kind of on board. But yeah, it was a [laugh] it was a journey.

Corey: So, one of the things you do—in fact, the first line in your bio talking about it—is you modernize baseline cloud infrastructure provisioning. That means an awful lot of things depending upon who it is that’s answering the question. What does that look like for you?

Levi: For what we’re doing right now, we’re trying to take what was a cobbled-together part-time project for one engineer, we’re trying to modernize that, turn it into as much self-service as we can. There’s a lot of steps that happen along the way, like a new workload needs to be spun up, they decide if they need a new AWS account or not, we pivot around, like, what does the access profile look like, who needs to have access to it, which things does it need to connect to, and then you look at the billing side, compliance side, and you just say, you know, “Who needs to be informed about these things?” We apply tags to the accounts, we start looking at lower-level tagging, depending on if it’s a shared workload account or if it’s a completely dedicated account, and we’re trying to wrap all of that in automation so that it can be as click-button as possible.

Corey: Historically, I found that when companies try to do this, the first few attempts at it don’t often go super well. We’ll be polite and say their first attempts resemble something artisanal and handcrafted, which might not be ideal for this. And then in many cases, the overreaction becomes something that is very top-down, dictatorial almost, is the way I would frame that. And the problem people learn then is that, “Oh, everyone is going to route around us because they don’t want to deal with us at all.” That doesn’t quite seem like your jam from what I know of you and your approach to things. How do you wind up keeping the guardrails up without driving people to shadow IT their way around you?

Levi: I always want to keep it in mind that even if it’s not an option, I want to at least pretend like a given team could not use our service, right? I try to bring a service mentality to it, so we’re talking Accounts as a Service. And then I just think about all of the things that they would have to solve if they didn’t go through us, right? Like, are they managing their finances w—imagine they had to go in and negotiate some kind of pricing deal on their own, right, all of these things that come with being part of our organization, being part of our service offering. And then just making sure, like, those things are always easier than doing it on their own.

Corey: How diverse would you say that the workloads are that are in your organization? I found that in many cases, you’ll have a SaaS-style company where there’s one primary workload that is usually bearing the name of the company, and that’s the thing that they provide to everyone. And then you have the enterprise side of the world where they have 1500 or 2000 distinct application teams working on different things, and the only thing they really have in common is, well, that all gets billed to the same company, eventually.

Levi: They are fairly diverse in how… they’re currently created. We’ve gone through a few acquisitions, we’ve pulled a bunch of those into our ecosystem, if you will. So, not everything has been completely modernized or brought over to, you know, standards, if you will, if such a thing even exists in companies. You know [laugh], you may pretend that they do, but you’re probably lying to yourself, right? But you know, there are varying platforms, we’ve got a whole laundry list of languages that are being used, we’ve got some containerized, some VM-based, some serverless workloads, so it’s all over the place. But you nailed it. Like, you know, the majority of our footprint lives in maybe a handful of, you know, SaaS offerings.

Corey: Right. It’s sort of a fun challenge when you start taking a looser approach to these things because someone gets back from re:Invent, like, “Well, I went to the keynote and now I have my new shopping list of things I’m going to wind up deploying,” and ehh, that never goes well, having been that person in a previous life.

Levi: Yeah. And you don’t want to apply too strict of governance over these things, right? You want people to be able to play, you want them to be inspired and start looking at, like, what would be—what’s something that’s going to move the needle in terms of our cloud architecture or product offerings or whatever we have. So, we have sandbox accounts that are pretty much wide open, we’ve got some light governance over those, [laugh] moreso for billing than anything. And all of our internal tooling is available, you know, like if you’re using containers or whatever, like, all of that stuff is in those sandbox accounts.

And that’s where our kind of service offering comes into play, right? Sandbox is still an account that we tried to vend, if you will, out of our service. So, people should be building in your sandbox environments just like they are in your production as much as possible. You know, it’s a place where tools can get the tires kicked and smooth out bugs before you actually get into, you know, roadmap-impacting problems.

Corey: One of the fun challenges you have is, as you said, the financial aspect of this. When you’ve got a couple of workloads that drive most things, you can reason about them fairly intelligently, but trying to predict the future—especially when you’re dealing with multi-year contract agreements with large cloud providers—becomes a little bit of a guessing game, like, “Okay. Well, how much are we going to spend on generative AI over the next three years?” The problem with that is that if you listen to an awful lot of talking heads or executive types, like, “Oh, yeah, if we’re spending $100 million a year, we’re going to add another 50 on top of that, just in terms of generative AI.” And it’s like, press X to doubt, just because it’s… I appreciate that you’re excited about these things and want to play with them, but let’s make sure that there’s some ‘there’ there before signing contracts that are painful to alter.

Levi: Yeah, it’s a real struggle. And we have all of these new initiatives, things people are excited for. Meanwhile, we’re bringing old architecture into a new platform, if you will, or a new footprint, so we have to constantly measure those against each other. We have a very active conversation with finance and with leadership every month, or even weekly, depending on the type of project and where that spend is coming from.

Corey: One of the hard parts has always been, I think, trying to get people on the finance side of the world, the engineering side of the world, and the folks who are trying to predict what the business was going to do next, all speaking the same language. It just feels like it’s too easy to wind up talking past each other if you’re not careful.

Levi: Yeah, it’s really hard. Recently taken over the FinOps practice. It’s been really important for me, for us to align on what our words mean, right? What are these definitions mean? How do we come to common consensus so that eventually the communication gets faster? But we can’t talk past each other. We have to know what our words mean, we have to know what each person cares about in this conversation, or what does their end goal look like? What do they want out of the conversation? So, that’s been—that’s taken a significant amount of time.

Corey: One of the problems I have is with the term FinOps as a whole, ignoring the fact entirely that it was an existing term of art within finance for decades; great, we’re just going to sidestep past that whole mess—the problem you’ll see is that it just seems like that it means something different to almost everyone who hears it. And it’s sort of become a marketing term more so that it has an actual description of what people are doing. Just because some companies will have a quote-unquote, “FinOps team,” that is primarily going to be run by financial analysts. And others, “Well, we have one of those lying around, but it’s mostly an engineering effort on our part.”

And I’ve seen three or four different expressions as far as team composition goes and I’m not convinced any of them are right. But again, it’s easy for me to sit here and say, “Oh, that’s wrong,” without having an environment of my own to run. I just tend to look at what my clients do. And, “Well, I’ve seen a lot of things, and they all work poorly in different ways,” is not uplifting and helpful.

Levi: Yeah. I try not to get too hung up on what it’s called. This is the name that a lot of people inside the company have rallied around and as long as people are interested in saving money, cool, we’ll call it FinOps, you know? I mean, DevOps is the same thing, right? In some companies, you’re just a sysadmin with a higher pay, and in some companies, you’re building extensive cloud architecture and pipelines.

Corey: Honestly, for the whole DevOps side of the world, I maintain we’re all systems administrators. The tools have changed, the methodologies have changed, the processes have changed, but the responsibility of ‘keep the site up’ generally has not. But if you call yourself a sysadmin, you’re just asking him to, “Please pay me less money in my next job.” No, thanks.

Levi: Yeah. “Where’s the Exchange Server for me to click on?” Right? That’s the [laugh]—if you call yourself a sysadmin [crosstalk 00:11:34]—

Corey: God. You’re sending me back into twitching catatonia from my early days.

Levi: Exactly [laugh].

Corey: So, you’ve been paying attention to this whole generative AI hype monster. And I want to be clear, I say this as someone who finds the technology super neat and I’m optimistic about it, but holy God, it feels like people have just lost all sense. If that’s you, my apologies in advance, but I’m still going to maintain the point.

Levi: I’ve played with all the various toys out there. I’m very curious, you know? I think it’s really fun to play with them, but to, like, make your entire business pivot on a dime and pursue it just seems ridiculous to me. I hate that the cryptocurrency space has pivoted so hard into it, you know? All the people that used to be shilling coins are now out there trying to cobble together a couple API calls and turn it into an AI, right?

Corey: It feels like it’s just a hype cycle that people are more okay with being a part of. Like, Andy Jassy, in the earnings call a couple of weeks ago saying that every Amazon team is working with generative AI. That’s not great. That’s terrifying. I’ve been playing with the toys as well and I’ve asked it things like, “Oh, spit out an IAM policy for me,” or, “Oh, great, what can I do to optimize my AWS bill?” And it winds up spitting out things that sound highly plausible, but they’re also just flat-out wrong. And that, it feels like a lot of these spaces, it’s not coming up with a plausible answer—that’s the hard part—is coming up with the one that is correct. And that’s what our jobs are built around.

Levi: I’ve been trying to explain to a lot of people how, if you only have surface knowledge of the thing that it’s telling you, it probably seems really accurate, but when you have deep knowledge on the topic that you’re interacting with this thing, you’re going to see all of the errors. I’ve been using GitHub’s Copilot since the launch. You know, I was in one of the previews. And I love it. Like, it speeds up my development significantly.

But there have been moments where I—you know, IAM policies are a great example. You know, I had it crank out a Lambda functions policy, and it was just frankly, wrong in a lot of places [laugh]. It didn’t quite imagine new AWS services, but it was really [laugh] close. The API actions were—didn’t exist. It just flat-out didn’t exist.

Corey: I love that. I’ve had some magic happen early on where it could intelligently query things against the AWS pricing API, but then I asked it the same thing a month later and it gave me something completely ridiculous. It’s not deterministic, which is part of the entire problem with it, too. But it’s also… it can help incredibly in some weird ways I didn’t see coming. But it can also cause you to spend more time chasing that thing than just doing it yourself the first time.

I found a great way to help it—you know, it helped me write blog posts with it. I tell it to write a blog post about a topic and give it some bullet points and say, “Write in my voice,” and everything it says I take issue with, so then I just copy that into a text editor and then mansplain-correct the robot for 20 minutes and, oh, now I’ve got a serviceable first draft.

Levi: And how much time did you save [laugh] right? It is fun, you know?

Corey: It does help because that’s better for me at least and staring at an empty page of what am I going to write? It gets me past the writer’s block problem.

Levi: Oh, that’s a great point, yeah. Just to get the ball rolling, right, once you—it’s easier to correct something that’s wrong, and you’re almost are spite-driven at that point, right? Like, “Let me show this AI how wrong it was and I’ll write the perfect blog post.” [laugh].

Corey: It feels like the companies jumping on this, if you really dig into what we’re talking about, it seems like they’re all very excited about the possibility of we don’t have to talk to customers anymore because the robots will all do that. And I don’t think that’s going to go the way you want to. We just have this minor hallucination problem. Yeah, that means that lies and tries to book customers to hotel destinations that don’t exist. Think about this a little more. The failure mode here is just massive.

Levi: It’s scary, yeah. Like, without some kind of review process, I wouldn’t ship that straight to my customers, right? I wouldn’t put that in front of my customer and say, like, “This is”—I’m going to take this generative output and put it right in front of them. That scares me. I think as we get deeper into it, you know, maybe we’ll see… I don’t know, maybe we’ll put some filters or review process, or maybe it’ll get better. I mean, who was it that said, you know, “This is the worst it’s ever going to be?” Right, it will only get better.

Corey: Well, the counterargument to that is, it will get far worse when we start putting this in charge [unintelligible 00:16:08] safety-critical systems, which I’m sure it’s just a matter of time because some of these boosters are just very, very convincing. It’s just thinking, how could this possibly go the worst? Ehhh. It’s not good.

Levi: Yeah, well, I mean, we’re talking impact versus quality, right? The quality will only ever get better. But you know, if we run before we walk, the impact can definitely get wider.

Corey: From where I sit, I want to see this really excel within bounded problem spaces. The one I keep waiting for is the AWS bill because it’s a vast space, yes, and it’s complicated as all hell, but it is bounded. There are a finite—though large—number of things you can see in an AWS bill, and there are recommendations you can make based on top of that. But everything I’ve seen that plays in this space gets way overconfident far too quickly, misses a bunch of very obvious lines of inquiry. Ah, I’m skeptical.

Then you pass that off to unbounded problem spaces like human creativity and that just turns into an absolute disaster. So, much of what I’ve been doing lately has been hamstrung by people rushing to put in safeguards to make sure it doesn’t accidentally say something horrible that it’s stripped out a lot of the fun and the whimsy and the sarcasm in the approach, of I—at one point, I could bully a number of these things into ranking US presidents by absorbency. That’s getting harder to do now because, “Nope, that’s not respectful and I’m not going to do it,” is basically where it draws the line.

Levi: The one thing that I always struggle with is, like, how much of the models are trained on intellectual property or, when you distill it down, pure like human suffering, right? Like, this is somebody’s art, they’ve worked hard, they’ve suffered for it, they put it out there in the world, and now it’s just been pulled in and adopted by this tool that—you know, how many of the examples of, “Give me art in the style of,” right, and you just see hundreds and hundreds of pieces that I mean, frankly, are eerily identical to the style.

Corey: Even down to the signature, in some cases. Yeah.

Levi: Yeah, exactly. You know, and I think that we can’t lose sight of that, right? Like, these tools are fun and you know, they’re fun to play with, it’s really interesting to explore what’s possible, but we can’t lose sight of the fact that there are ultimately people behind these things.

Corey: This episode is sponsored in part by Panoptica. Panoptica simplifies container deployment, monitoring, and security, protecting the entire application stack from build to runtime. Scalable across clusters and multi-cloud environments, Panoptica secures containers, serverless APIs, and Kubernetes with a unified view, reducing operational complexity and promoting collaboration by integrating with commonly used developer, SRE, and SecOps tools. Panoptica ensures compliance with regulatory mandates and CIS benchmarks for best practice conformity. Privacy teams can monitor API traffic and identify sensitive data, while identifying open-source components vulnerable to attacks that require patching. Proactively addressing security issues with Panoptica allows businesses to focus on mitigating critical risks and protecting their interests. Learn more about Panoptica today at panoptica.app.

Corey: I think it matters, on some level, what the medium is. When I’m writing, I will still use turns of phrase from time to time that I first encountered when I was reading things in the 1990s. And that phrase stuck with me and became part of my lexicon. And I don’t remember where I originally encountered some of these things; I just know I use those raises an awful lot. And that has become part and parcel of who and what I am.

Which is also, I have no problem telling it to write a blog post in the style of Corey Quinn and then ripping a part of that out, but anything that’s left in there, cool. I’m plagiarizing the thing that plagiarized from me and I find that to be one of those ethically just moments there. But written word is one thing depending on what exactly it’s taking from you, but visual style for art, that’s something else entirely.

Levi: There’s a real ethical issue here. These things can absorb far much more information than you ever could in your entire lifetime, right, so that you can only quote-unquote, you know, “Copy, borrow, steal,” from a handful of other people in your entire life, right? Whereas this thing could do hundreds or thousands of people per minute. I think that’s where the calculus needs to be, right? How many people can we impact with this thing?

Corey: This is also nothing new, where originally in the olden times, great, copyright wasn’t really a thing because writing a book was a massive, massive undertaking. That was something that you’d have to do by hand, and then oh, you want a copy of the book? You’d have to have a scribe go and copy the thing. Well then, suddenly the printing press came along, and okay, that changes things a bit.

And then we continue to evolve there to digital distribution where suddenly it’s just bits on a disk that I can wind up throwing halfway around the internet. And when the marginal cost of copying something becomes effectively zero, what does that change? And now we’re seeing, I think, another iteration in that ongoing question. It’s a weird world and I don’t know that we have the framework in place even now to think about that properly. Because every time we start to get a handle on it, off we go again. It feels like if they were doing be invented today, libraries would absolutely not be considered legal. And yet, here we are.

Levi: Yeah, it’s a great point. Humans just do not have the ethical framework in place for a lot of these things. You know, we saw it even with the days of Napster, right? It’s just—like you said, it’s another iteration on the same core problem. I [laugh] don’t know how to solve it. I’m not a philosopher, right?

Corey: Oh, yeah. Back in the Napster days, I was on that a fair bit in high school and college because I was broke, and oh, I wanted to listen to this song. Well, it came on an album with no other good songs on it because one-hit wonders were kind of my jam, and that album cost 15, 20 bucks, or I could grab the thing for free. There was no reasonable way to consume. Then they started selling individual tracks for 99 cents and I gorged myself for years on that stuff.

And now it feels like streaming has taken over the world to the point where the only people who really lose on this are the artists themselves, and I don’t love that outcome. How do we have a better tomorrow for all of this? I know we’re a bit off-topic from you know, cloud management, but still, this is the sort of thing I think about when everything’s running smoothly in a cloud environment.

Levi: It’s hard to get people to make good decisions when they’re so close to the edge. And I think about when I was, you know, college-age scraping by on minimum wage or barely above minimum wage, you know, it was hard to convince me that, oh yeah, you shouldn’t download an MP3 of that song; you should go buy the disc, or whatever. It was really hard to make that argument when my decision was buy an album or figure out where I’m going to, you know, get my lunch. So, I think, now that I’m in a much different place in my life, you know, these decisions are a lot easier to make in an ethical way because that doesn’t impact my livelihood nearly as much. And I think that is where solutions will probably come out of. The more people doing better, the easier it is for them to make good decisions.

Corey: I sure hope you’re right, but something I found is that okay we made it easy for people to make good decisions. Like, “Nope, you’ve just made it easier for me to scale a bunch of terrible ones. I can make 300,000 more terrible decisions before breakfast time now. Thanks.” And, “No, that’s not what I did that for.” Yet here we are. Have you been tracking lately what’s been going on with the HashiCorp license change?

Levi: Um, a little bit, we use—obviously use Terraform in the company and a couple other Hashi products, and it was kind of a wildfire of, you know, how does this impact us? We dove in and we realized that it doesn’t, but it is concerning.

Corey: You’re not effectively wrapping Terraform and then using that as the basis for how you do MDM across your customer fleets.

Levi: Yeah. You know, we’re not deploying customers' written Terraform into their environments or something kind of wild like that. Yeah, it doesn’t impact us. But it is… it is concerning to watch a company pivot from an open-source, community-based project to, “Oh, you can’t do that anymore.” It doesn’t impact a lot of people who use it day-to-day, but I’m really worried about just the goodwill that they’ve lit on fire.

Corey: One of the problems, too, is that their entire write-up on this was so vague that it was—there is no way to get an actual… piece of is it aimed at us or is it not without very deep analysis, and hope that when it comes to court, you’re going to have the same analysis as—that is sympathetic. It’s, what is considered to be a competitor? At least historically, it was pretty obvious. Some of these databases, “Okay great. Am I wrapping their database technology and then selling it as a service? No? I’m pretty good.”

But with HashiCorp, what they do is so vast in a few key areas that no one has the level of certainty. I was pretty freaking certain that I’m not shipping MongoDB with my own wrapper around it, but am I shipping something that looks like Terraform if I’m managing someone’s environment for them? I don’t know. Everything’s thrown into question. And you’re right. It’s the goodwill that currently is being set on fire.

Levi: Yeah, I think people had an impression of Hashi that they were one of the good guys. You know, the quote-unquote, “Good guys,” in the space, right? Mitchell Hashimoto is out there as a very prominent coder, he’s an engineer at heart, he’s in the community, pretty influential on Twitter, and I think people saw them as not one of the big, faceless corporations, so to see moves like this happen, it… I think it shook a lot of people’s opinions of them and scared them.

Corey: Oh, yeah. They’ve always been the good guys in this context. Mitch and Armon were fantastic folks. I’m sure they still are. I don’t know if this is necessarily even coming from them. It’s market forces, what are investors demanding? They see everyone is using Terraform. How does that compare to HashiCorp’s market value?

This is one of the inherent problems if I’m being direct, of the end-stages of capitalism, where it’s, “Okay, we’re delivering on a lot of value. How do we capture ever more of it and growing massively?” And I don’t know. I don’t know what the answer is, but I don’t think anyone’s thrilled with this outcome. Because, let’s be clear, it is not going to meaningfully juice their numbers at all. They’re going to be setting up a lot of ill will against them in the industry, but I don’t see the upside for them. I really don’t.

Levi: I haven’t really done any of the analysis or looked for it, I should say. Have you seen anything about what this might actually impact any providers or anything? Because you’re right, like, what kind of numbers are we actually talking about here?

Corey: Right. Well, there are a few folks that have done things around this that people have named for me: Spacelift being one example, Pulumi being another, and both of them are saying, “Nope, this doesn’t impact us because of X, Y, and Z.” Yeah, whether it does or doesn’t, they’re not going to sit there and say, “Well, I guess we don’t have a company anymore. Oh, well.” And shut the whole thing down and just give their customers over to HashiCorp.

Their own customers would be incensed if that happened and would not go to HashiCorp if that were to be the outcome. I think, on some level, they’re setting the stage for the next evolution in what it takes to manage large-scale cloud environments effectively. I think basically, every customer I’ve ever dealt with on my side has been a Terraform shop. I finally decided to start learning the ins and outs of it myself a few weeks ago, and well, it feels like I should have just waited a couple more weeks and then it would have become irrelevant. Awesome. Which is a bit histrionic, but still, this is going to plant seeds for people to start meaningfully competing. I hope.

Levi: Yeah, I hope so too. I have always awaited releases of Terraform Cloud with great anticipation. I generally don’t like managing my Terraform back-ends, you know, I don’t like managing the state files, so every time Terraform Cloud has some kind of release or something, I’m looking at it because I’m excited, oh finally, maybe this is the time I get to hand it off, right? Maybe I start to get to use their product. And it has never been a really compelling answer to the problems that I have.

And I’ve always said, like, the [laugh] cloud journey would be Google’s if they just released a managed Terraform [laugh] service. And this would be one way for them to prevent that from happening. Because Google doesn’t even have an Infrastructure as Code competitor. Not really. I mean, I know they have their, what, Plans or their Projects or whatever they… their Infrastructure as Code language was, but—

Corey: Isn’t that what Stackdriver was supposed to be? What happened with that? It’s been so long.

Levi: No, that’s a logging solution [laugh].

Corey: That’s the thing. It all runs together. Not it was their operations suite that was—

Levi: There we go.

Corey: —formerly Stackdriver. Yeah. Now, that does include some aspects—yeah. You’re right, it’s still hanging out in the observability space. This is the problem is all this stuff conflates and companies are terrible at naming and Google likes to deprecate things constantly. And yeah, but there is no real competitor. CloudFormation? Please. Get serious.

Levi: Hey, you’re talking to a member of the CloudFormation support group here. So, I’m still a huge fan [laugh].

Corey: Emotional support group, more like it, it seems these days.

Levi: It is.

Corey: Oh, good. It got for loops recently. We’ve been asking for basically that to make them a lot less wordy only for, what, ten years?

Levi: Yeah. I mean, my argument is that I’m operating at the account level, right? I need to deploy to 250, 300, 500 accounts. Show me how to do that with Terraform that isn’t, you know, stab your eyes out with a fork.

Corey: It can be done, but it requires an awful lot of setting things up first.

Levi: Exactly.

Corey: That’s sort of a problem. Like yeah, once you have the first 500 going, the rest are just like butter. But that’s a big step one is massive, and then step two becomes easy. Yeah… no, thank you.

Levi: [laugh]. I’m going to stick with my StacksSets, thank you.

Corey: [laugh]. I really want to thank you for taking the time to come back on and honestly kibitz about the state of the industry with me. If people want to learn more, where’s the best place for them to find you?

Levi: Well, I’m still active on the space normally known as—formerly known as Twitter. You can reach out to me there. DMs are open. I’m always willing to help people learn how to cloud better. Hopefully trying to make my presence known a little bit more on LinkedIn. If you happen to be over there, reach out.

Corey: And we will, of course, put links to that in the [show notes 00:30:16]. Thank you so much for taking the time to speak with me again. It’s always a pleasure.

Levi: Thanks, Corey. I always appreciate it.

Corey: Levi McCormick, Director of Cloud Engineering at Jamf. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, and along with an insulting comment that tells us that we completely missed the forest for the trees and that your programmfing is going to be far superior based upon generative AI.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

View Details

Alysha Love, Executive Editor and Co-Founder of Payette Media House, joins Corey on Screaming in the Cloud to discuss her career journey going from journalism to editing and how she works with Corey on his content. Alysha describes why she feels it’s so important to capture the voice of the person you’re editing, and why editing your content makes a difference to those reading it. Corey and Alysha also explore the differences in editing for something that will be read silently versus something that will be read out loud, as well as the different styles of editing.

About Alysha

Alysha Love is executive editor and co-founder of Payette Media House, an editorial agency serving startups and tech companies. Alysha is the treasurer of ACES: The Society for Editing, the nation's largest editing organization, and trains editors and writers in digital best practices.

She was an editor at CNN and POLITICO during the Obama and Trump administrations. Alysha has a bachelor's in journalism from the University of Missouri and a master's in leadership and organizational development from the University of Texas. She's a big fan of the humble ampersand.

Links Referenced:

  • Company website: https://payettemediahouse.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Human-scale teams use Tailscale to build trusted networks. Tailscale Funnel is a great way to share a local service with your team for collaboration, testing, and experimentation. Funnel securely exposes your dev environment at a stable URL, complete with auto-provisioned TLS certificates. Use it from the command line or the new VS Code extensions. In a few keystrokes, you can securely expose a local port to the internet, right from the IDE.

I did this in a talk I gave at Tailscale Up, their first inaugural developer conference. I used it to present my slides and only revealed that that’s what I was doing at the end of it. It’s awesome, it works! Check it out!

Their free plan now includes 3 users & 100 devices. Try it at snark.cloud/tailscalescream

Corey: Welcome to Screaming in the Cloud, I’m Corey Quinn. And one of the, I guess, illusions about what I do is that I sit down at a keyboard periodically, and I just start typing and then, you know, brilliance emerges, and then my work is done. It turns out that this is rarely true, not to deflate my own image overly much. And a big part of how that works comes down to my guest today. Alysha Love is the executive editor at Payette Media House and has been my editor for just about three years now. Alysha, thank you for tolerating me.

Alysha: Anytime.

Corey: So, I want to start by dispensing with a few illusions that I’m not saying other people have, I’m saying that I have, where I was fortunate enough—or unfortunate as the case may be—to grow up with an English teacher for a mother and understanding how to put together a grammatically correct sentence was not exactly optional in my house, so what possible value could an editor present to me? And one of the things I learned along the way is that there are multiple kinds of editors, as it turns out. What are they and which are you?

Alysha: So yeah, not only is editing a thing, we can look at your sentences, your story, and make it all better and clearer so that it really shines. But there are different types of editors who can do different specific functions. So, at the maybe most nitpicky level, you have proofreaders who are looking at what would be a final page, usually in something like a book, where it needs to look exactly right the way that it’s going up, it needs to be sure that every last little detail is in place. At the next level up, you have copy editors. They’re looking for things like spelling, grammar, punctuation, style, factual accuracy. That’s sort of what you think of usually when you think of somebody who might peer-review a piece for you, or who you might ask to edit something. And then at the next level, you have people who are able to do the copy editing, but in addition to that, they look at the overarching arc of the story or the blog piece, and they’re able to help look for some of those gaps and organize it into something that is clearer and easier to understand.

Corey: Something I’ve always been curious about is that you, previously in another life, were an editor at CNN and then Politico during the Obama and then Trump administrations. Is editing what I do significantly different than editing, you know, journalists?

Alysha: Yes, in a few key ways. One is that when we’re writing news, we always come out and say the most important thing first. It’s what we call the inverted pyramid style, so if you turn a pyramid on its nose and it’s standing on the tip, you have the biggest part of the triangle or the pyramid at the top, and that’s the most important thing that could never get cut, and you say it right out of the gate. I tease my husband a lot because he tends to bury the lede, and that’s what we’re talking about when it’s not the first thing you say.

Corey: Absolutely. And I do that, meanwhile, stylistically as a choice because, you know, don’t put the punch line in the title.

Alysha: Totally. So, that’s a big difference between editing for news and editing you. You also use significantly more voice than we would use in a CNN or Politico article. That’s also a choice. And it’s actually something I have a ton of fun with is emulating your voice as I make edits.

Corey: I found that as we’ve worked together, our comfort with one another has grown significantly over the past few years. At this point, just for folks who are wondering, anytime you have an edit that’s just a reordering or something that clarifies something slightly or is basically low-level stylistic, you don’t track those changes; you just go ahead and apply them because otherwise, I’ll wind up, “Oh, here are 600 changes to make.” It’s like, the article is 2000 words. Exactly how much was done? And so, much of it is white spaces and comma placement and the rest and just strange little things that frankly, are not that important to me past a certain point.

The exception, of course, was always great. First, if you’re making a change, tell me why. I have political opinions about the use of the Oxford comma, for example. I find it lends clarity to things. Fortunately, you and I aligned on that, so it’s a non-issue. But I am curious as far as what do you see that I tend to do the most that I guess either annoys you or you disagree with stylistically or, honestly, is flat-out wrong.

Alysha: So, there’s not much you do that’s flat-out wrong. I will say, like, the instances that I do see something… you’ve told me before that your mom was an English teacher and that these are things that you really pride yourself on being able to do well, so I don’t know how you feel about it, but I’ll usually leave a comment telling you about the change and linking you [laugh] to something that can explain it a little better. Maybe that annoys the shit out of you, or—sorry, can we cuss?

Corey: No, no—oh, you absolutely can, and—

Alysha: Great.

Corey: Because it’s—I was taught by a teacher, I want to say in third grade, that you leave two spaces at the end of every sentence before you begin the next sentence, and I only found out about six or seven years ago that’s not really a thing. It took me a year to break myself of that habit, but I would rather go through that effort than, “Well, I’ve been wrong this long. I may as well double down on it now.” Just seems like that’s not helping anyone.

Alysha: [laugh]. Right. And we all have those things that there was some English teacher somewhere along the way who taught us things that were just wrong. So, my favorite thing about my magazine editing class, when I was in school at the University of Missouri, was that she started from the very, very basics because she said everyone has learned things that are wrong about how we write and how we edit and so we’re just going to learn it all from scratch. And it was really brilliant. It was the best way to learn it all the right way.

Corey: I’m usually gratified when I am trying to figure out what is the proper tense of this particular verb in this particular phrase. And my wife and I will wind up in debates on this constantly because she’s an attorney and also writes a lot for a living. Who knew? And invariably whenever we finally get to an impasse and look it up—because, you know, we do have the sum total of human knowledge in the supercomputer that lives in our pockets—the answer is more often than not, it’s a matter of choice. And both are considered accepted because English is, of course, a language defined by its usage, or one way is British and one is American, or some other aspect where it’s not about wrong; it’s about which is preferred in certain contexts. So, I’ll take it.

Alysha: That’s—yeah, that’s totally accurate. And those are the kinds of choices that I feel like, if I were to change all of those things in your writing, you would not appreciate it because they’re preferences. So, those are the things that even if there’s a style that’s a little different unless what you’ve done is wrong in some sort of, like, widely accepted way, then I’m going to leave it the way you’ve got it.

Corey: One way that I have found that I am both strong and weak, I think, as a writer—and I’m thrilled to be criticized on any of this; please don’t spare my feelings any—is that I write like I speak. When—this is most noticeable on Twitter when I meet people who’ve never met me in person before, a very common refrain is, “Oh, you’re just like your Twitter feed.”

Alysha: [laugh].

Corey: And partially that’s because I’m sarcastic and irreverent and a class clown who never grew up, but another part of it is because I write like I would put together the sentence. In long-form writing, I feel like that can be something of a setback for me. When I’m making a sentence right now, for example, and talking to you, if I were to write this out as a literal transcript, it will be a long run-on sentence in a bunch of different ways. And it works conversationally, but it does not work that way in long-form writing. So, I feel like I have a bunch of clauses that continue to go on forever when I let myself. [transcriber note: yup]

Alysha: You do. And the thing that you also love to do stick a bunch of semicolons between all of them, which is technically correct, but I do have a whole thing about distracting punctuation, so I will take out many of your semicolons.

Corey: I would like credit though because before you were involved in this, Mike would periodically look at some of my blog posts before they went out and—because I wanted his perspective on, “Am I onto something here or am I a fool,” but then he’ll go back and edit some of the things he sees—which I get. If I see a misspelling in something, I itch until I can fix it, or a grammatical mistake. But at one point, he was constantly onto me about overusing commas. And in one case, he took a bunch out. And then I looked at the tracked changes on this and it’s forever one of my favorite things. You went through next and wound up returning all of the commas that he had removed. It’s, “That’s right.” But you got me on the semicolon thing. I’m trying to reduce usage and have shorter sentences.

Alysha: Yes. And that’s something that’s really good for digital best practices and having a wide and varied audience. You know, with a diverse audience, with audiences that don’t speak English as a first language, it’s helpful to have much shorter sentences. For folks who are consuming content on the internet, in general, it’s easier to skim and get the meaning out of a shorter sentence. However, when we think about your voice, it’s important to leave some of those really long sentences in because we want people to keep thinking and, like, “Oh, yeah. This is a Corey Quinn piece,” when they read your article, whether your name is at the top of it or not.

Corey: What I found is that varying the sentence structure and length also keeps reading from being fatiguing in some cases. And there are times I’ll do things that are, quote-unquote, “Incorrect” to make a point stylistically. Like, normally you wouldn’t put that word in italics and bold, but yeah, for this case, it is so egregious—probably Managed NAT Gateway or something—that I absolutely feel the need to wind up emphasizing the egregiousness of whatever it is I’m opining on that week.

Alysha: Yes. And I think that is also part of what makes editing for you really fun is that there’s a great balance of let’s keep to the rules as much as we can when it makes sense, but let’s be super strategic about how we break them to have better emphasis and to make it clear that this is a Corey Quinn piece.

Corey: One problem that I’ve had, too, is understanding the difference in medium. I mean, most of my engagement with writing, when I was growing up, was books I read enthusiastically. And then I started writing a lot of newsletters and mailing lists and various written fora. I spent entirely too many years on IRC over the course of my life. And there are different rules and all of those circumstances, but never having written a book myself, how differently do you approach the editing process when you’re writing something that is long-form or writing something that is essay length, or writing something that is a book?

Alysha: So, I’m actually working on my first book now as the editor. So, that’s a thing that I’m learning about, learning more about what that process looks like and how it’s different. I think there’s a lot more note-taking as you go along to track, you know, this is the story arc, these are the characters, what’s a first reference and a second reference?

Corey: When you overuse a phrase, it’s easy to figure out if it’s in a 2000-word essay. When you use it more than once, oh, great. Easy to spot. But okay, you write books—generally not in one sitting, I would assume—and you say, all right, that is the eighth time you have used that very particular turn of phrase. Stop it here’s a thesaurus.

Alysha: Totally. I don’t know, maybe this is just a me thing or an editor thing, but do you notice when you hear, like, a very unique word, that’s the thing, if—by the way [laugh], speaking of different, you know, if this weren’t a spoken word podcast, then I would never say very unique; I would edit out ‘very’ ahead of ‘unique’ because unique is unique.

Corey: Exactly. It’s a pet peeve.

Alysha: Totally. But I have very different rules for the way that we speak versus the way that we write. How fleeting are things? So, that gets back to your original question. Something that, you know, if I’m editing something quick, that’s a quick hit, it’s not going to live for very long. If you needed me to edit a tweet, I wouldn’t spend a lot of time on that. I’m going to spend more time on things that have longer legs or that are going to a bigger platform. Books, you spend way more time editing than you would a 2000-word essay.

Corey: I find that I don’t have people edit tweets very often because first, it’s moving too quickly for me to really take something out for opinion. The reason I’d have to do that is, “Is this too close to the edge?” Well, it turns out at this stage, if I have to ask that question, I already know the answer.

Alysha: Mm-hm.

Corey: Everything else is going to be more stylistic, like, “Is dogshit one word or two?” And you’re, “Ah, it’s a [unintelligible 00:13:24]. There we go. Excellent.” It’s not the typical kind of problem or question that you would run into.

Alysha: The BuzzFeed Style Guide has been a great resource for questions like dogshit [laugh].

Corey: I didn’t realize they had such a thing and that is absolutely amazing.

Alysha: It is fantastic. Most of the internet things you need to know are there. CNN is where—well, CNN and Politico both—that’s where we were always taking second eyes to look at a tweet before it goes out and you’re doing that in about ten seconds. But we’re looking at factual accuracy. Is there something that is about to be very wrong that we don’t want to embarrass the publication with?

Corey: I’m a prolific writer because I have to be. I have a content schedule that you could charitably call punishing. And that works super well with the way that I view the world, but the counterargument is that getting me to go back and review edits or go back and edit after I’ve written something is sort of like pulling teeth. So, something I found that works for me as a way around this is I record these essays as podcast episodes on the AWS Morning Brief. What that forces me to do is once the edits are in, I get one last read-through as I read it out loud in a normal speaking voice and don’t power my way through it, and I’m forced to pay attention to every word at that point.

And, “Oh, that doesn’t quite make the point that I thought it did.” And you’ve edited them by this point, so it’s not ever going to be something that is, “Oh, that’s a run-on sentence,” or, “Huh, punctuation is probably a good idea.” It’s something that is more abstract than that and often very tied to a domain-specific aspect of what I’ve written about. But I found that to be one of the best last filters for a lot of the stuff.

Alysha: Yeah. That’s a great tactic for catching errors, and… and not even errors, right? But it also comes back to, like, what’s a difference between the way that we’re going to write things and the way that you’re going to read it out loud? I try to edit, keeping in mind that you’re going to be reading these out loud, but then there are always going to be things that are going to sound better a particular way, and the way that we write them is going to be slightly different.

Corey: One thing that I find as well, given that I read a lot more than I write, is that when I’m looking at articles for inclusion in a whole bunch of different places because I’m looking for creative content from the community, it is very hard for me to go ahead and greenlight including something that is poorly written. If I can’t get through the first three sentences without seeing six mistakes in how the sentence is built, I judge the writing for it. It’s you’re talking about a technical topic, but if you can’t even get to a point where the sentence is coherent, then how do I know you don’t have typos littered throughout the code samples you’re about to put up, or whatnot? And I don’t think that that is an entirely fair assessment of mine, but it still feels like nails on a chalkboard, every time I encounter some of it.

Alysha: It’s actually something that’s backed by research. I’m on the board for ACES, the Society for Editing, and we commissioned research about 13 years ago now, so it’s getting updated. But what it showed was that readers can distinguish between edited and unedited content in significant ways. So, it may not be, like, “Oh, I know exactly how to fix this,” or, “I know exactly what’s wrong with this,” but they get the sense that that content is not as reliable if it hasn’t been edited. So, there’s true value in exactly what you just said, in having content that’s edited and the way that it makes people feel about the quality of the content that they’re reading.

Corey: It’s one of those important things—which I’m not trying to shame people, particularly those for whom English is not a first language; you speak more languages than I do. Good for you—but I also will judge corporate blog posts far more harshly for this because it’s no longer just one person. You should—in theory—have the ability to proofread and copy-edit the thing that is going out underneath your masthead. People are expensive. Writers are expensive unless you’re ripping people off, which I don’t advise. At least take the extra few steps to make sure that it doesn’t drive people away for reasons other than the content.

Alysha: Yep, I’m totally with you on that [laugh].

Corey: I find myself having that same negative reaction to typos on your landing page when you describe what you do. There have been security vendors that I won’t touch with a ten-foot pole because they talk about the standards that they follow, but they misspell the word ‘standards’ on the webpage. And in a lot of these areas, details very much matter. One area that I want to get into as well that I think you and I have always been aligned on. Because I’ve worked with a number of editors—all of them great in different ways, I want to be very clear; I’m not trying to shame anyone—but challenges I’ve had from time to time have been editors who come from the marketing world who like to embody what I can only refer to as the bullshit marketing voice.

And I don’t know what exactly the elements of it are, but I know it when I hear it or see it. You can see this on almost every billboard out there, every press release that goes out. If I were to talk to a human in a way that the press release talks about the product and company, it would not go well for me, just because I would come across as incredibly condescending, entirely too self-promotional, and there’s just something about the way that it’s written that feels off-putting so much of my online persona and approaches have come from simply calling out the subtext in an awful lot of unfortunate marketing communications. You’ve never had a problem with that. I have never once looked at something you’ve edited for me and put something in where it’s, “Ooh. That sounds a little bit too market-y.” And again, I consider myself something of a marketer. This is not me disliking marketing; it’s disliking bad marketing. I don’t believe that that’s the sort of thing that just emerges out of nowhere. What’s your history of marketing?

Alysha: So, I did start working in marketing at Intuit QuickBooks a few years ago, back in 2018, when I moved out of journalism. So, I think the way that I approach marketing, and content marketing in particular, is always very journalistic. My bullshit meter really goes off, too, when I read something that’s like, “You have a claim to back that,” or, “Oh, the evidence that you’re using to back that is really thin.” And it just… it’s just icky, right? Like, none of us like to feel like we’re being marketed to in that way.

And you’re right, you would like, you would turn tail and run if somebody started talking to you that way in real life. So yeah, so it’s just sort of a combination of journalistic instinct and like, you know, a lot of times, if you just say something straight, if you just say the truth, it comes out with even more impact than if you tried to fluff it up with marketing speak.

Corey: The thing that I wish companies would figure out is that when you go out and talk about your product and mention the things it’s bad at, it really engenders an awful lot of trust. Because it’s not like you’re going to hide that from the first people to use it, so call it out upfront that this is an area it’s weak in. And that is anathema to some folks where they believe that you can say something is good, something is great, or you can stop talking. But it is unhelpful to the people you’re trying to reach. I’m sure there are reasons for this. I don’t believe for a second that I know better than the entire field of marketing, let’s be clear here. But I know what I want to read what I’m trying to get when I’m presented with new information about a product or a service.

Alysha: I think it just scares the bejesus out of people to think that they are going to publicly admit to things that aren’t great. Yeah, and I don’t know what the idea is after that. Like, that we’re just going to sweep it under the rug and hope nobody notices or try to work on it in the background, and until then, we’ll just talk about, you know, our one huge talking point and tell you that it’s the best, most amazing in the world. It comes off as disingenuous to the rest of us. And that is something that you are not. You are… you’re definitely the antithesis of that. You’re very trusting because you call out all of the things that aren’t quite right in a very honest way.

Corey: And people love it until it’s their turn to the hot seat, I think. That tends to, “Well, hang on, my product is perfect.” I assure you it’s not but that’s okay.

Alysha: To being fair, you’re also very good at calling out what others do great, and maybe in a way that you don’t always get credit for, but—

Corey: Well, no, I’ve done experiments on this. When I am unflinchingly positive about some aspect of what a cloud provider or other vendor has done or a feature that I really like, it gets almost no notice. But when I say, “Oh, and this part is crap,” that’s the part that blows up and goes around the internet a bunch of times. And I think that’s human nature. I don’t know if, as an editor, you have a way to fix human nature, but if you do, I’m very interested to learn it.

Alysha: [laugh]. No, we just all love to bitch and to talk about our pain points, and when somebody says it and all you can say is, “Yes, plus one million,” then it’s going to get a lot of play.

Corey: One aspect of what you do did scare me initially when we first started working together, and that was, you do a few things: you write as well—which that’s not scary. I would expect someone who can edit would also know how to write. That does make sense. But you also do some SEO-facing work. And that in many ways feels like it is modern-day witch doctor-y because my approach to SEO has always been naive but also effective.

I write compelling, original content that people like and as a result, link against or refer to. And I find the rest of it really takes care of itself. I haven’t spent deep effort or large amounts of brain sweat figuring out how to appear at the top of Google search results for various terms.

Alysha: Yeah. And one of the things I was tasked with as soon as I came in, was, “Please write an SEO description for each of Corey’s pieces and make sure that we’re writing for digital best practices, including SEO.”

Corey: And I’ve read those descriptions and I’ve never had a problem with any of them because it’s not something that is aligned with… anything that I hate. So, good for you on that. It’s an active description using very direct, to-the-point phrasing about what it is I’m talking about. And yeah, that is, ideally what SEO should be. It’s about, this is what this is, but you shouldn’t have to read through a thousand meandering words while he circles the point to death like some sort of persistence hunter. I get it.

Alysha: Totally. We’re going to be direct and to the point, we’re going to use the key nouns, but we’re not going to be gross about it. And we’re definitely not going to jump on the latest trend because honestly, Google’s always looking to get ahead of what all of the SEO magicians are trying to magic up.

Corey: I get so many emails in the course of a week for people asking to contribute articles to my blog. Which again, we do have a guest author program, but that’s one of those, yeah, if that happens, we’re paying you and then throwing an editor—read as you—to whoever it is that’s contributing that so it comes out something that we’re thrilled to have up there. But money flows one direction in that and it’s from us to the guest author. Instead there, “Oh, we’re going to provide high-quality content,” or they’ll link to something on the site, usually a newsletter back issue, and say, “Hey, include our link to this because it’s relevant,” and it’s clearly for SEO juice. And first, I’m sure Google and the other search engines would just love if I suddenly have a bunch of crappy links to low-quality sites. But further, it doesn’t serve the audience in any meaningful way, and… it just irks me.

Alysha: And when you start playing that game, you get into the middle of all of that the link-swapping, and trying to up their SEO juice, and it is wild the amount of money that people will offer to pay for a link on a reputable site. It’s super valuable. So, the way that I approach linking in your pieces is exactly the same way that I did it in news, which is, where do we need to show our sources, where will people want to verify information? Let’s just go ahead and give them that link. And that’s about it. Like, what do people need to know?

Corey: I always worry, on some level, that I’m thinking about this all wrong. But if I’m being snarky and sarcastic with all of the SEO people emailing me who then try to offer me SEO services, it’s frankly, if that’s what I’m looking for, shouldn’t I just Google ‘SEO’ and pick whoever’s at the top of the list? Because they clearly get it in one.

Alysha: [laugh].

Corey: It turns out, for some reason, they don’t really have a good rejoinder to that when you ask them directly.

Alysha: [laugh]. I love that. And might I mention, when I do search for topics that I know you guys have articles on, I won’t necessarily include The Duckbill Group, but you do show up because you are a reputable and authoritative source who does not play the SEO game.

Corey: I do have one more question that lies down this path that I’m actually deeply curious about, and I’ve always found it to be something that is incredibly helpful for my purposes. But your background is in journalism and writing and editing. It is not—for some unknown reason—the world of cloud. Almost like you want to be happy or something.

Alysha: [laugh].

Corey: How approachable or unapproachable is my writing to someone who does not live in the space the way I do?

Alysha: Oh, that’s a really interesting question. So, I’m married to somebody who spent ten years, just about, working specifically around AWS, in the—

Corey: It took that long for them to stop the billing. I get it.

Alysha: [laugh]. And my brother-in-law is also a software engineer. So, I have witnessed enough conversations between the two of them that I had a decent idea of what I was getting into. And those two are my resources when I have stupid questions that I don’t want to ask you [laugh] in a Google Doc comment. So, I go to them, I get the lowdown, I do a little research sometimes, but by and large, we’re talking about bigger concepts, and I think sometimes it might even help that I’m not in the weeds on some of the details of things that you talk about because it helps me see patterns that I’m a little—I can make some connections that maybe you’re not making in the middle of the weeds.

Corey: It’s always tricky to figure out where to level-set what I’m talking about. I don’t want to turn every article to have the first 18 pages of it be a primer on what Cloud computing is. I have to assume, at least on some level, people have a baseline level of understanding. But there are times I go too far in the other direction where I assume that, “Oh, well, I used to be a software engineer, so I’m going to write as if everyone reading is.”

In fact, the audience is not overwhelmingly populated by purely software engineers. There’s a lot of systems folk, there’s a lot of managers at a variety of different levels, ranging from line management to executive, and it really takes all kinds. I’m always surprised when people reach out and mention they’ve been reading for a while, and then they describe what they do for a day job and it’s nothing I would have ever considered. It would not have occurred to me early on that people who spent their entire life in the finance department would find most of what I talk about that isn’t cost related to be interesting. But they assure me they do. Okay.

Alysha: That makes sense because it gives them insight into what the other half of the business is doing.

Corey: On some level, what I’ve found is you have to pick—and it can vary; it can honestly vary even within a piece, but at every given point, I feel like you have to have someone in mind that you’re writing for because otherwise you’re trying to write for everyone and in so doing, you write something that’s valuable to no one.

Alysha: Do you remember how much I hammered on you about who your audience was at the beginning of every single article when I first started editing?

Corey: Oh, yes. That’s what shaped the ideas. I mean, honestly, if you were telling me the same thing, now that you were two-and-a-half years ago, I’d wonder if—in your case—if I was even reading the notes that you put into these things. Editors make your writing better, but they also longer-term make you a better writer, is my firm belief.

Alysha: Oh, that’s lovely.

Corey: I’m assuming that the mistakes I make are at least more interesting now as opposed to some of the ones that we had long conversations about. I hope.

Alysha: Totally. It is interesting every time.

Corey: So, I have to ask, given as someone who is a big believer in writing, and because it’s a way of expressing myself and giving myself a platform that doesn’t require me to be in the same room as a bunch of other people or them to be willing to fire up a podcast and listen to me or watch a video, they can access it anywhere they are at any given point in time, I love the writing process, but the editing process is challenging for me. You have—seem to be on the other side of that where you are much happier editing than you are writing. At least that’s my perception of you and your background. If that is accurate, how do you think about this stuff because it’s foreign to me?

Alysha: That’s really funny. So, I actually started out thinking that I wanted—well, I wanted to work with words, and I thought that the way that you could do that was by writing, and specifically reporting. So, that was the track I went down. And it actually wasn’t until my first full-time job out of college—I was working as a copy editor at Politico—that I realized that I could wake up and edit every single day. That was what I had the energy for.

And when I say wake up and edit, I mean, it was 4 a.m. and we were editing the newsletters that had to go out by 6 a.m. so quite literally, it was the thing I could wake up and do. And I think what I really love about it is taking something that’s already good, that’s already great in a lot of instances, and making it better, so it’s just that little bit more clear, more understandable, that your message is getting across in a way that still feels authentic to you. Because I can tell you one of my least favorite things as a writer was having someone come through and edit and I could tell you every single spot that that editor had touched. And it sort of… it burned. It just didn’t feel quite right.

Corey: Suddenly the voice switches, like effectively, you have someone whose voice sounds like you, for example, and then for half a sentence, it suddenly sounds like James Earl Jones is delivering it, and then it goes back to your voice. It’s hmm.

Alysha: Totally. So, with my experience of editing as a writer, my goal is to make that as seamless as possible. So, I want to show you the changes that you’re going to be most interested in and that I think you might want to learn from. And the changes that I do make, I want them to sound just like you.

Corey: Honestly, because there’s usually a week or two in time that happens between me writing a draft or something, and then going back, when you’ve just automatically made some of the quick rewrites on the fly, unless I go looking, I never realize which parts you’ve touched or not. And I’m the one that wrote it. So, I guess, honestly, you’re in a terrific position to put words in my mouth if you want to. Have fun. But that is, to me, the mark of an editor who gets it.

I just find it scary, on some level, to the idea, from my perspective, of fading into the background. I always lived in fear of not having my name front and center and being in the spotlight, for good or bad, just because it’s that’s who I am. That’s what I bias for.

Alysha: That’s really interesting. Totally makes sense because you are very front and center. When I was working at Reuters in Brussels, one of the things I think is really cool that they do is they put their editors’ byline at the bottom of articles, so the editors do get a hat-tip of recognition. But I think as somebody who’s a little bit of a helper, I just get a lot of enjoyment out of making other people’s stuff better.

Corey: I can certainly say that you’ve been a smashing success from my perspective, although I’m sure now you’re going to be inundated with people who are urging you to, “Okay, now make what he says less bullshit or at least something that I can agree with.” Unfortunately, it doesn’t quite work that way most of the time.

Alysha: No.

Corey: Though you are getting very proficient at sanding off some of the more colorful metaphors. Thank you for that.

Alysha: [laugh]. Anytime. I got to keep my [unintelligible 00:33:28], too.

Corey: If people want to learn more, where’s the best place for them to find you, and—take this as a personal recommendation—hire you to edit their stuff, so I don’t have to claw my eyes out as much when I read their things?

Alysha: You can find me at payettemediahouse.com. P-A-Y-E-T-T-E Media House.

Corey: And we will, of course, put a link to that in the [show notes 00:33:50]. Thank you so much for taking the time to go through something different with me in a stranger way than we normally wind up communicating, which is via tracking changes.

Alysha: [laugh]. I love it. It’s nice to see your face.

Corey: It really is. I have a face for radio though, so it’s only for so long. Alysha Love, Executive Editor at Payette Media House. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an angry, insulting comment that I will absolutely not be reading because you’ve [BLEEP]-ed the subject-verb agreement in your first sentence.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

View Details

Forrest Brazeal, Head of Developer Media at Google Cloud, joins Corey on Screaming in the Cloud to discuss how AI, current job markets, and more are impacting software engineers. Forrest and Corey explore whether AI helps or hurts developers, and what impact it has on the role of a junior developer and the rest of the development team. Forrest also shares his viewpoints on how he feels AI affects people in creative roles. Corey and Forrest discuss the pitfalls of a long career as a software developer, and how people can break into a career in cloud as well as the necessary pivots you may need to make along the way. Forrest then describes why he feels workers are currently staying put where they work, and how he predicts a major shift will happen when the markets shift.

About Forrest

Forrest is a cloud educator, cartoonist, author, and Pwnie Award-winning songwriter. He currently leads the content marketing team at Google Cloud. You can buy his book, The Read Aloud Cloud, from Wiley Publishing or attend his talks at public and private events around the world.

Links Referenced:

  • Personal Website: https://goodtechthings.com
  • Newsletter signup: https://cloud.google.com/innovators

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn and I am thrilled to have a returning guest on, who has been some would almost say suspiciously quiet over the past year or so. Forrest Brazeal is the Head of Developer Media over at Google Cloud, and everyone sort of sits there and cocks their head, like, “What does that mean?” And then he says, “Oh, I’m the cloud bard.” And everyone’s, “Oh, right. Get it: the song guy.” Forrest, welcome back.

Forrest: Thanks, Corey. As always, it’s great to be here.

Corey: So, what have you been up to over the past, oh let’s call it, I don’t know, a year, since I think, is probably the last time you’re on the show.

Forrest: Well, gosh, I mean, for one thing, it seems like I can’t call myself the cloud bard anymore because Google rolled out this thing called Bard and I’ve started to get some DMs from people asking for, you know, tech support on Bard. So, I need to make that a little bit clearer that I do not work on Bard. I am a lowercase bard, but I was here first, so if anything, you know, Google has deprecated me.

Corey: Honestly, this feels on some level like it’s more cloudy if we define cloudy as what, you know, Amazon does because they launched a quantum computing service about six months after they launched some unrelated nonsense that they called [QuantumDB 00:01:44], which you’d think if you’re launching quantum stuff, you’d reserve the word quantum for that. But no, they’re going to launch things that stomp all over other service names as well internally, so customers just wind up remarkably confused. So, if you find a good name, just we’re going to slap it on everything, seems to be the way of cloud.

Forrest: Yeah, naming things has proven to be harder than either quantum computing or generative AI at this point, I think.

Corey: And in fairness, I will point out that naming things is super hard; making fun of names is not. So, that is—everyone’s like, “Wow, you’re so good at making fun of names. Can you name something well?” [laugh]. Absolutely not.

Forrest: Yeah, well, one of the things you know, that I have been up to over the past year or so is just, you know, getting to learn more about what it’s like to have an impact in a very, very large organizational context, right? I mean, I’ve worked in large companies before, but Google is a different size and scale of things and it takes some time honestly, to, you know, figure out how you can do the best for the community in an environment like that. And sometimes that comes down to the level of, like, what are things called? How do we express things in a way that makes sense to everyone and takes into account people’s different communication styles and different preferences, different geographies, regions? And that’s something that I’m still learning.

But you know, hopefully, we’re getting to a point where you’re going to start hearing some things come out of Google Cloud that answer your questions and makes sense to you. That’s supposed to be part of my job, anyway.

Corey: So, I want to talk a bit about the idea of generative AI because there has been an awful lot of hype in the space, but you have never given me a bum steer. You have always been a level-headed, reasonable voice. You are not—to my understanding—a VC trying desperately to prop up an industry that you may or may not believe in, but you are financially invested into. What is your take on the last, let’s call it, year of generative AI enhancements?

Forrest: So, to be clear, while I do have a master’s degree in interactive intelligence, which is kind of AI adjacent, this is not something that I build with day-to-day professionally. But I have spent a lot of time over the last year working with the people who do that and trying to understand what is the value that gen AI can bring to the domains that I do care about and have a lot of interest in, which of course, are cloud developers and folks trying to build meaningful enterprise applications, take established workloads and make them better, and as well work with folks who are new to their careers and trying to figure out, you know, what’s the most appropriate technology for me to bet on? What’s going to help me versus what’s going to hurt me?

And I think one of the things that I have been telling people most frequently—because I talk to a lot of, like, new cloud learners, and they’re saying, “Should I just drop what I’m doing? Should I stop building the projects I’m working on and should I instead just go and get really good at generating code through something like a Bard or a ChatGPT or what have you?” And I went down a rabbit hole with this, Corey, for a long time and spent time building with these tools. And I see the value there. I don’t think there’s any question.

But what has come very, very clearly to the forefront is, the better you already are at writing code, the more help a generative AI coding assistant is going to give you, like a Bard or a ChatGPT, what have you. So, that means the way to get better at using these tools is to get better at not using these tools, right? The more time you spend learning to code without AI input, the better you’ll be at coding with AI input.

Corey: I’m not sure I entirely agree because for me, the wake-up call that I had was a singular moment using I want to say it was either Chat-Gippity—yes, that’s how it’s pronounced—or else it was Gif-Ub Copilot—yes, also how it’s pronounced—and the problem that I was having was, I wanted to query probably the worst API in the known universe—which is, of course, the AWS pricing API: it returns JSON, that kind of isn’t, it returns really weird structures where you have to correlate between a bunch of different random strings to get actual data out of it, and it was nightmarish and of course, it’s not consistent. So, I asked it to write me a Python script that would contrast the hourly cost of a Managed NAT gateway in all AWS regions and return a table sorted by the most to least expensive. And it worked.

Now, this is something that I could have done myself in probably half a day because my two programming languages of choice remain brute force and enthusiasm, but it wound up taking away so much of the iterative stuff that doesn’t work of oh, that’s not quite how you’d handle that data structure. Oh, you think it’s a dict, but no, it just looks like one. It’s a string first; now you have to convert it, or all kinds of other weird stuff like that. Like, this is not senior engineering work, but it really wound up as a massive accelerator to get the answer I was after. It was almost an interface to a bad API. Or rather, an interface to a program—to a small script that became an interface itself to a bad API.

Forrest: Well, that’s right. But think for a minute, Corey, about what’s implicit in that statement though. Think about all the things you had to know to get that value out of ChatGPT, right? You had to know, A, what you were looking for: how these prices worked, what the right price [style 00:06:52] was to look for, right, why NAT gateway is something you needed to be caring about in the first place. There’s a pretty deep stack of things—actually, it’s what we call a context window, right, that you needed to know to make this query take a half-day of work away from you.

And all that stuff that you’ve built up through years and years of being very hands-on with this technology, you put that same sentence-level task in the hands of someone who doesn’t have that background and they’re not going to have the same results. So, I think there’s still tremendous value in expanding your personal mental context window. The more of that you have, the better and faster results you’re going to get.

Corey: Oh, absolutely. I do want to steer away from this idea that there needs to be this massive level of subject matter expertise because I don’t disagree with it, but you’re right, the question I asked was highly contextual to the area of expertise that I have. But everyone tends to have something like that. If you’re a marketer for example, and you wind up with an enormous pile of entrants on a feedback form, great. Can you just dump it all in and say, can you give me a sentiment analysis on this?

I don’t know how to run a sentiment analysis myself, but I’m betting that a lot of these generative AI models do, or being able to direct me in the right area on this. The question I have is—it can even be distilled down into simple language of, “Here’s a bunch of comments. Do people love the thing or hate the thing?” There are ways to get there that apply, even if you don’t have familiarity with the computer science aspects of it, you definitely have aspect to the problem in which you are trying to solve.

Forrest: Oh, yeah, I don’t think we’re disagreeing at all. Domain expertise seems to produce great results when you apply it to something that’s tangential to your domain expertise. But you know, I was at an event a month or two ago, and I was talking to a bunch of IT executives about ChatGPT and these other services, and it was interesting. I heard two responses when we were talking about this. The first thing that was very common was I did not hear any one of these extremely, let’s say, a little bit skeptical—I don’t want to say jaded—technical leaders—like, they’ve been around a long time; they’ve seen a lot of technologies come and go—I didn’t hear a single person say, “This is something that’s not useful to me.”

Every single one of them immediately was grasping the value of having a service that can connect some of those dots, can in-between a little bit, if you will. But the second thing that all of them said was, “I can’t use this inside my company right now because I don’t have legal approval.” Right? And then that’s the second round of challenges is, what does it look like to actually take these services and make them safe and effective to use in a business context where they’re load-bearing?

Corey: Depending upon what is being done with them, I am either sympathetic or dismissive of that concern. For example, yesterday, I wound up having fun with it, and—because I saw a query, a prompt that someone had put in of, “Create a table of the US presidents ranked by years that they were in office.” And it’s like, “Okay, that’s great.” Like, I understand the value here. But if you have a magic robot voice from the future in a box that you can ask it any question and as basically a person, why not have more fun with it?

So, I put to it the question of, “Rank the US presidents by absorbency.” And it’s like, “Well, that’s not a valid way of rating presidential performance.” I said, “It is if I have a spill and I’m attempting to select the US president with which to mop up the spill.” Like, “Oh, in that case, here you go.” And it spat out a bunch of stuff.

That was fun and exciting. But one example he gave was it ranked Theodore Roosevelt very highly. Teddy Roosevelt was famous for having a mustache. That might be useful to mop up a spill. Now, I never would have come up in isolation with the idea of using a president’s mustache to mop something up explicitly, but that’s a perfect writer’s room style Yes, And approach that I could then springboard off of to continue iterating on if I’m using that as part of something larger. That is a far cry from copying and pasting whatever it is to say into an email, whacking send before realizing it makes no sense.

Forrest: Yeah, that’s right. And of course, you can play with what we call the temperatures on these models, right, to get those very creative, off-the-wall kind of answers, or to make them very, kind of, dry and factual on the other end. And Google Cloud has been doing some interesting things there with Generative AI Studio and some of the new features that have come to Vertex AI. But it’s just—it’s going to be a delicate dance, honestly, to figure out how you tune those things to work in the enterprise.

Corey: Oh, absolutely. I feel like the temperature dial should instead be instead relabeled as ‘corporate voice.’ Like, do you want a lot of it or a little of it? And of course, they have to invert it. But yeah, the idea is that, for some things, yeah, you definitely just want a just-the-facts style of approach.

Another demo that I saw, for example, that I thought showed a lack of imagination was, “Here’s a transcript of a meeting. Extract all the to-do items.” Okay. Yeah, I suppose that works, but what about, here’s a transcript of the meeting. Identify who the most unpleasant, passive-aggressive person in this meeting is to work with.

And to its credit—because of course this came from something corporate, none of the systems that I wound up running that particular query through could identify anyone because of course the transcript was very bland and dry and not actually how human beings talk, other than in imagined corporate training videos.

Forrest: Yes, well again, I think that gets us into the realm of just because you can doesn’t mean you should use it for this.

Corey: Oh, I honestly, most of what I use this stuff for—or use anything for—should be considered a cautionary tale as opposed to guidance for the future. You write parody songs a fair bit. So do I, and I’ve had an attempt to write versions of, like, write parody lyrics for some random song about this theme. And it’s not bad, but for a lot of that stuff, it’s not great, either. It is a starting point.

Forrest: Now, hang on, Corey. You know, as well as I do that I don’t write parody songs. We’ve had this conversation before. A parody is using existing music and adding new lyrics to it. I write my own music and my own lyrics and I’ll have you know, that’s an important distinction. But—

Corey: True.

Forrest: I think you’re right on that, you know, having these services give you creative output. What you’re getting is an average of a lot of other creative output, right, which is—could give you a perfectly average result, but it’s difficult to get a first pass that gives you something that really stands out. I do also find, as a creative, that starting with something that’s very average oftentimes locks me into a place where I don’t really want to be. In other words, I’m not going to potentially come up with something as interesting if I’m starting with a baseline like that. It’s almost a little bit polluting to the creative process.

I know there’s a lot of other creatives that feel that way as well, but you’ve also got people that have found ways to use generative AI to stimulate some really interesting creative things. And I think maybe the example you gave of the president’s rank by absorbency is a great way to do that. Now, in that case, the initial creativity, a lot of it resided in the prompt, Corey. I mean, you’re giving it a fantastically creative, unusual, off-the-wall place to start from. And just about any average of five presidents that come out of that is going to be pretty funny and weird because of just how funny and weird the idea was to begin with. That’s where I think AI can give you that great writer’s room feel.

Corey: It really does. It’s a Yes, And approach where there’s a significant way that it can build on top of stuff. I’ve been looking for a, I guess, a writer’s room style of approach for a while, but it’s hard to find the right people who don’t already have their own platform and voice to do this. And again, it’s not a matter of payment. I’m thrilled to basically pay any reasonable out of money to build a writer’s room here of people who get the cloud industry to work with me and workshops on some of the bigger jokes.

The challenge is that those people are very hard to find and/or are conflicted out. Having just a robot who, with infinite patience for tomfoolery—because the writing process can look kind of dull and crappy until you find the right thing—has been awesome. There’s also a sense of psychological safety in not poisoning people. Like, “I thought you were supposed to be funny, but this stuff is all terrible. What’s the deal here?” I’ve already poisoned that well with my business partner, for example.

Forrest: Yeah, there’s only so many chances you get to make that first impression, so why not go with AI that never remembers you or any of your past mistakes?

Corey: Exactly. Although the weird thing is that I found out that when they first launched Chat-Gippity, it already knew who I was. So, it is in fact familiar, so at least my early work of my entire—I guess my entire life. So that’s—

Forrest: Yes.

Corey: —kind of worrisome.

Forrest: Well, I know it credited to me books I hadn’t written and universities I hadn’t attended and all kinds of good stuff, so it made me look better than I was.

Corey: So, what have you been up to lately in the context of, well I said generative AI is a good way to start, but I guess we can also call it at Google Cloud. Because I have it on good authority that, marketing to the contrary, all of the cloud providers do other things in addition to AI and ML work. It’s just that’s what’s getting in the headline these days. But I have noticed a disturbing number of virtual machines living in a bunch of customer environments relative to the amount of AI workloads that are actually running. So, there might be one or two other things afoot.

Forrest: That’s right. And when you go and talk to folks that are actively building on cloud services right now, and you ask them, “Hey, what is the business telling you right now? What is the thing that you have to fix? What’s the thing that you have to improve?” AI isn’t always in the conversation.

Sometimes it is, but very often, those modernization conversations are about, “Hey, we’ve got to port some of these services to a language that the people that work here now actually know how to write code in. We’ve got to find a way to make this thing a little faster. Or maybe more specifically, we’ve got to figure out how to make it run at the same speed while using less or less expensive resources.” Which is a big conversation right now. And those are things that they are conversations as old as time. They’re not going away, and so it’s up to the cloud providers to continue to provide services and features that help make that possible.

And so, you’re seeing that, like, with Cloud Run, where they’ve just announced this CPU Boost feature, right, that gives you kind of an additional—it’s like a boost going downhill or a push on the swing as you’re getting started to help you get over that cold-start penalty. Where you’re seeing the session affinity features for Cloud Run now where you have the sticky session ability that might allow you to use something like, you know, a container-backed service like that, instead of a more traditional load balancer service that you’d be using in the past. So, you know, just, you take your eye off the ball for a minute, as you know, and 10 or 20, more of these feature releases come out, but they’re all kind of in service of making that experience better, broadening the surface area of applications and workloads that are able to be moved to cloud and able to be run more effectively on cloud than anywhere else.

Corey: There’s been a lot of talk lately about how the idea of generative AI might wind up causing problems for people, taking jobs away, et cetera, et cetera. You almost certainly have a borderline unique perspective on this because of your work with, honestly, one of the most laudable things I’ve ever seen come out of anywhere which is The Cloud Resume Challenge, which is a build a portfolio site, then go ahead and roll that out into how you interview. And it teaches people how to use cloud, step-by-step, you have multi-cloud versions, you have them for specific clouds. It’s nothing short of astonishing. So, you find yourself talking to an awful lot of very early career folks, folks who are transitioning into tech from other places, and you’re seeing an awful lot of these different perspectives and AI plays come to the forefront. How do you wind up, I guess, making sense of all this? What guidance are you giving people who are worried about that?

Forrest: Yeah, I mean, I, you know—look, for years now, when I get questions from these, let’s call them career changers, non-traditional learners who tend to be a large percentage, if not a plurality, of the people that are working on The Cloud Resume Challenge, for years now, the questions that they’ve come to me with are always, like, you know, “What is the one thing I need to know that will be the magic technology, the magic thing that will unlock the doors and give me the inside track to a junior position?” And what I’ve always told them—and it continues to be true—is, there is no magic thing to know other than magically going and getting two years of experience, right? The way we hire juniors in this industry is broken, it’s been broken for a long time, it’s broken not because of any one person’s choice, but because of this sort of tragedy of the commons situation where everybody’s competing over a dwindling pool of senior staff level talent and hopes that the next person will, you know, train the next generation for them so they don’t have to expend their energy and interview cycles and everything else on it. And as long as that remains true, it’s just going to be a challenge to stand out.

Now, you’ll hear a lot of people saying that, “Well, I mean, if I have generative AI, I’m not going to need to hire a junior developer.” But if you’re saying that as a hiring manager, as a team member, then I think you always had the wrong expectation for what a junior developer should be doing. A junior developer is not your mini me who sits there and takes the little challenges, you know, the little scripts and things like that are beneath you to write. And if that’s how you treat your junior engineers, then you’re not creating an environment for them to thrive, right? A junior engineer is someone who comes in who, in a perfect world, is someone who should be able to come in almost in more of an apprentice context, and somebody should be able to sit alongside you learning what you know, right, and having education integrated into their actual job experience so that at the end of that time, they’re able to step back and actually be a full-fledged member of your team rather than just someone that you kind of throw tasks over the wall to, and they don’t have any career advancement potential out of that.

So, if anything, I think the advancement of generative AI, in a just world, ought to give people a wake-up call that, hey, training the next generation of engineers is something that we’re actually going to have to actively create programs around, now. It’s not something that we can just, you know, give them the scraps that fall off of our desks. Unfortunately, I do think that in some cases, the gen AI narrative more than the reality is being used to help people put off the idea of trying to do that. And I don’t believe that that’s going to be true long-term. I think that if anything, generative AI is going to open up more need for developers.

I mean, it’s generating a lot of code, right, and as we know, Jevons paradox says that when you make it easier to use something and there’s elastic demand for that thing, the amount of creation of that thing goes up. And that’s going to be true for code just like it was for electricity and for code and for GPUs and who knows what all else. So, you’re going to have all this code that has a much lower barrier of entry to creating it, right, and you’re going to need people to harden that stuff and operate it in production, be on call for it at three in the morning, debug it. Someone’s going to have to do all that, you know? And what I tell these junior developers is, “It could be you, and probably the best thing for you to do right now is to, like I said before, get good at coding on your own. Build as much of that personal strength around development as you can so that when you do have the opportunity to use generative AI tools on the job, that you have the maximum amount of mental context to put around them to be successful.”

Corey: I want to further point out that there are a number of folks whose initial reaction to a lot of this is defensiveness. I showed that script that wound up spitting out the Managed NAT gateway ranked-by-region table to one of our contract engineers, who’s very senior. And the initial response I got from them was almost defensive, were, “Okay, yeah. That’ll wind up taking over, like, a $20 an hour Upwork coder, but it’s not going to replace a senior engineer.” And I felt like that was an interesting response psychologically because it felt defensive for one, and two, not for nothing, but senior developers don’t generally spring fully formed from the forehead of some ancient God. They start off as—dare I say it—junior developers who learn and improve as they go.

So, I wonder what this means. If we want to get into a point where generative AI takes care of all the quote-unquote, “Easy programming problems,” and getting the easy scripts out, what does that mean for the evolution and development of future developers?

Forrest: Well, keep in mind—

Corey: And that might be a far future question.

Forrest: Right. That’s an argument as old as time, right, or a concern is old as time and we hear it anew with each new level of automation. So, people were saying this a few years ago about the cloud or about virtual machines, right? Well, how are people going to, you know, learn how to do the things that sit on top of that if they haven’t taken the time to configure what’s below the surface? And I’m sympathetic to that argument to some extent, but at the same time, I think it’s more important to deal with the reality we have now than try to create an artificial version of realities’ past.

So, here’s the reality right now: a lot of these simple programming tasks can be done by AI. Okay, that’s not likely to change anytime soon. That’s the new reality. So now, what does it look like to bring on juniors in that context? And again, I think that comes down to don’t look at them as someone who’s there just to, you know, be a pair of hands on a keyboard, spitting out tiny bits of low-level code.

You need to look at them as someone who needs to be, you know, an effective user of general AI services, but also someone who is being trained and given access to the things they’ll need to do on top of that, so the architectural decisions, the operational decisions that they’ll need to make in order to be effective as a senior. And again, that takes buy-in from a team, right, to make that happen. That is not going to happen automatically. So, we’ll see. That’s one of those things that’s very hard to automate the interactions between people and the growth of people. It takes people that are willing to be mentors.

Corey: I’m also curious as to how you see the guidance shifting as computers get better. Because right now, one of my biggest problems that I see is that if I have an idea for a company I want to start or a product I want to build that involves software, step one is, learn to write a bunch of code. And I feel like there’s a massive opportunity for skipping aspects of that, whereas effectively have the robot build me the MVP that I describe. Think drag-and-drop to build a web app style of approach.

And the obvious response to that is, well, that’s not going to go to hyperscale. That’s going to break in a bunch of different ways. Well, sure, but I can get an MVP out the door to show someone without having to spend a year building it myself by learning the programming languages first, just to throw away as soon as I hire someone who can actually write code. It cuts down that cycle time massively, and I can’t shake the feeling that needs to happen.

Forrest: I think it does. And I think, you know, you were talking about your senior engineer that had this kind of default defensive reaction to the idea that something like that could meaningfully intrude on their responsibilities. And I think if you’re listening to this and you are that senior engineer, you’re five or more years into the industry and you’ve built your employability on the fact that you’re the only person who can rough out these stacks, I would take a very, very hard look at yourself and the value that you’re providing. And you say, you know—let’s say that I joined a startup and the POC was built out by this technical—or possibly the not-that-technical co-founder, right—they made it work and that thing went from, you know, not existing to having users in the span of a week, which we’re seeing more now and we’re going to see more and more of. Okay, what does my job look like in that world? What am I actually coming on to help with?

Am I—I’m coming on probably to figure out how to scale that thing and make it maintainable, right, operate it in a way that is not going to cause significant legal and financial problems for the company down the road. So, your role becomes less about being the person that comes in and does this totally greenfield thing from scratch and becomes more about being the person who comes in as the adult in the room, technically speaking. And I think that role is not going away. Like I said, there’s going to be more of those opportunities rather than less. But it might change your conception of yourself a little bit, how you think about yourself, the value that you provide, now’s the time to get ahead of that.

Corey: I think that it is myopic and dangerous to view what you do as an engineer purely through the lens of writing code because it is a near certainty that if you are learning to write code today and build systems involving technology today, that you will have multiple careers between now and retirement. And in fact, if you’re entering the workforce now, the job that you have today will not exist in anything remotely approaching the same way by the time you leave the field. And the job you have then looks borderline unrecognizable, if it even exists at all today. That is the overwhelming theme that I’ve got on this ar—the tech industry moves quickly and is not solidified like a number of other industries have. Like, accountants: they existed a generation ago and will exist in largely the same form a generation from now.

But software engineering in particular—and cloud, of course, as well, tied to that—have been iterating so rapidly, with such sweepingly vast changes, that that is something that I think we’re going to have a lot of challenge with, just wrestling with. If you want a job that doesn’t involve change, this is the wrong field.

Forrest: Is it the wrong field. And honestly, software engineering is, has been, and will continue to be a difficult business to make a 40-year career in. And this came home to me really strongly. I was talking to somebody a couple of months ago who, if I were to say the name—which I won’t—you and I would both know it, and a lot of people listening to this would know as well. This is someone who’s very senior, very well respected is, by name, identified in large part with the creation of a significant movement in technology. So, someone who you would never think of would be having a problem getting a job.

Corey: Is it me? And is it Route 53 as a database, as the movement?

Forrest: No, but good guess.

Corey: Excellent.

Forrest: This is someone I was talking to because I had just given a talk where I was pleading with IT leaders to take more responsibility for building on-ramps for non-traditional learners, career changers, people that are doing something a little different with their career. And I was mainly thinking of it as people that had come from a completely non-technical background or maybe people that were you know, like, I don’t know, IT service managers with skills 20 years out of date, something like that. But this is a person who you and I would think of as someone at the forefront, the cutting edge, an incredibly employable person. And this person was a little bit farther on in their career and they came up to me and said, “Thank you so much for giving that talk because this is the problem I have. Every interview that I go into, I get told, ‘Oh, we probably can’t afford you,’ or, ‘Oh well, you say you want to do AI stuff now, but we see that all your experience is doing this other thing, and we’re just not interested in taking a chance on someone like that at the salary you need to be at.’” and this person’s, like, “What am I going to do? I don’t see the roadmap in front of me anymore like I did 10, 15, or 20 years ago.”

And I was so sobered to hear that coming from, again, someone who you and I would consider to be a luminary, a leading light at the top of the, let’s just broadly say IT field. And I had to go back and sit with that. And all I could come up with was, if you’re looking ahead and you say I want to be in this industry for 30 years, you may reach a point where you have to take a tremendous amount of personal control over where you end up. You may reach a point where there is not going to be a job out there for you, right, that has the salary and the options that you need. You may need to look at building your own path at some point. It’s just it gets really rough out there unless you want to continue to stagnate and stay in the same place. And I don’t have a good piece of advice for that other than just you’re going to have to find a path that’s unique to you. There is not a blueprint once you get beyond that stage.

Corey: I get asked questions around this periodically. The problem that I have with it is that I can’t take my own advice anymore. I wish I could. But what I used to love doing was, every quarter or so, I’d make it a point to go on at least one job interview somewhere else. This wound up having a few great features.

One, interviewing is a skill that atrophies if you don’t use it. Two, it gives me a finger on the pulse of what the market is doing, what the industry cares about. I dismissed Docker the first time I heard about it, but after the fourth interview where people were asking about Docker, okay, this is clearly a thing. And it forced me to keep my resume current because I’ve known too many people who spend seven years at a company and then wind up forgetting what they did years three, four, and five, where okay, then what was the value of being there? It also forces you to keep an eye on how you’re evolving and growing or whether you’re getting stagnant.

I don’t ever want to find myself in the position of the person who’s been at a company for 20 years and gets laid off and discovers to their chagrin that they don’t have 20 years of experience; they have one year of experience repeated 20 times. Because that is a horrifying and scary position to be in.

Forrest: It is horrifying and scary. And I think people broadly understand that that’s not a position they want to be in, hence why we do see people that are seeking out this continuing education, they’re trying to find—you know, trying to reinvent themselves. I see a lot of great initiative from people that are doing that. But it tends to be more on the company side where, you know, they get pigeonholed into a position and the company that they’re at says, “Yeah, no. We’re not going to give you this opportunity to do something else.”

So, we say, “Okay. Well, I’m going to go and interview other places.” And then other companies say, “No, I’m not going to take a chance on someone that’s mid-career to learn something brand new. I’m going to go get someone that’s fresh out of school.” And so again, that comes back to, you know, where are we as an industry on making space for non-traditional learners and career changers to take the maturity that they have, right, even if it’s not specific familiarity with this technology right now, and let them do their thing, let them get untracked.

You know, there’s tremendous potential being untapped there and wasted, I would say. So, if you’re listening to this and you have the opportunity to hire people, I would just strongly encourage you to think outside the box and consider people that are farther on in their careers, even if their technical skill set doesn’t exactly line up with the five pieces of technology that are on your job req, look for people that have demonstrated success and ability to learn at whatever [laugh] the things are that they’ve done in the past, people that are tremendously highly motivated to succeed, and let them go win on your behalf. There’s—you have no idea the amount of talent that you’re leaving on the table if you don’t do that.

Corey: I’d also encourage people to remember that job descriptions are inherently aspirational. If you take a job where you know how to do every single item on the list because you’ve done it before, how is that not going to be boring? I love being given problems. And maybe I’m weird like this, but I love being given a problem where people say, “Okay, so how are you going to solve this?” And the answer is, “I have no idea yet, but I can’t wait to find out.” Because at some level, being able to figure out what the right answer is, pick up the skill sets I don’t need, the best way to learn something that I’ve ever found, at least for me.

Forrest: Oh, I hear that. And what I found, you know, working with a lot of new learners that I’ve given that advice to is, typically the ones that advice works best for, unfortunately, are the ones who have a little bit of baked-in privilege, people that tend to skate by more on the benefit of the doubt. That is a tough piece of advice to fulfill if you’re, you know, someone who’s historically underrepresented or doesn’t often get the chance to prove that you can do things that you don’t already have a testament to doing successfully. So again, takes it back to the hiring side. Be willing to bet on people, right, and not just to kind of look at their resume and go from there.

Corey: So, I’m curious to see what you’ve noticed in the community because I have a certain perspective on these things, and a year ago, everyone was constantly grousing about dissatisfaction with their employers in a bunch of ways. And that seems to have largely vanished. I know, there have been a bunch of layoffs and those are tragic on both sides, let’s be very clear. No one is happy when a layoff hits. But I’m also seeing a lot more of people keeping their concerns to either private channels or to themselves, and I’m seeing what seems to be less mobility between companies than I saw previously. Is that because people are just now grateful to have a job and don’t want to rock the boat, or is it still happening and I’m just not seeing it in the same way?

Forrest: No, I think the vibe has shifted, for sure. You’ve got, you know, less opportunities that are available, you know that if you do lose your job that you’re potentially going to have fewer places to go to. I liken it to like if you bought a house with a sub-3% mortgage and 2021, let’s say, and now you want to move. Even though the housing market may have gone down a little bit, those interest rates are so high that you’re going to be paying more, so you kind of are stuck where you are until the market stabilizes a little bit. And I think there’s a lot of people in that situation with their jobs, too.

They locked in salaries at ’21, ’22 prices and now here we are in 2023 and those [laugh] those opportunities are just not open. So, I think you’re seeing a lot of people staying put—rationally, I would say—and waiting for the market to shift. But I think that at the point that you do see that shift, then yes, you’re going to see an exodus; you’re going to see a wave and there will be a whole bunch of new think pieces about the great resignation or something, but all it is just that pent up demand as people that are unhappy in their roles finally feel like they have the mobility to shift.

Corey: I really want to thank you for taking the time to speak with me. If people want to learn more, where’s the best place for them to find you?

Forrest: You can always find me at goodtechthings.com. I have a newsletter there, and I like to post cartoons and videos and other fun things there as well. If you want to hear my weekly take on Google Cloud, go to cloud.google.com/innovators and sign up there. You will get my weekly newsletter The Overwhelmed Person’s Guide to Google Cloud where I try to share just the Google Cloud news and community links that are most interesting and relevant in a given week. So, I would love to connect with you there.

Corey: I have known you for years, Forrest, and both of those links are new to me. So, this is the problem with being active in a bunch of different places. It’s always difficult to—“Where should I find you?” “Here’s a list of 15 places,” and some slipped through the cracks. I’ll be signing up for both of those, so thank you.

Forrest: Yeah. I used to say just follow my Twitter, but now there’s, like, five Twitters, so I don’t even know what to tell you.

Corey: Yes. The balkanization of this is becoming very interesting. Thanks [laugh] again for taking the time to chat with me and I look forward to the next time.

Forrest: All right. As always, Corey, thanks.

Corey: Forrest Brazeal, Head of Developer Media at Google Cloud, and of course the Cloud Bard. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an insulting comment that you undoubtedly had a generative AI model write for you and then failed to proofread it.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

View Details

AWS Morning Brief Extras edition for the week of August 16, 2023.

Want to give your ears a break and read this as an article? You’re looking for this link.

https://www.lastweekinaws.com/blog/prime-day-pricing

Never miss an episode

  • Join the Last Week in AWS newsletter
  • Subscribe wherever you get your podcasts

Help the show

  • Leave a review
  • Share your feedback
  • Subscribe wherever you get your podcasts

Buy our merch

  • https://store.lastweekinaws.com

What's Corey up to?

  • Follow Corey on Twitter (@quinnypig)
  • See our recent work at the Duckbill Group
  • Apply to work with Corey and the Duckbill Group to help lower your AWS bill

View Details

Josh Doody, Owner of Fearless Salary Negotiation, joins Corey on Screaming in the Cloud to discuss how important tonality and communication is, both in salary negotiations and everyday life. Josh describes how important it is to have a positive padding to your communications in order to make the person on the other end of the negotiation feel like a collaborator rather than a combatant. Corey and Josh also describe scenarios where tonality made a huge difference in the outcome, and Josh gives some examples of where and when to be mindful of how you’re coming across in modern communication methods. Josh also reveals how negotiating with companies multiple times allows him to understand their recruiters more than a person who is encountering their negotiation process for the first time.

About Josh

Josh is a salary negotiation coach who works with senior software engineers and engineering managers to negotiate job offers with big tech companies. He also wrote Fearless Salary Negotiation: A Step-by-Step Guide to Getting Paid What You're Worth, and recently launched Salary Negotiation Mastery to help folks who aren't able to work with Josh 1-on-1.

Links Referenced:

  • Fearless Salary Negotiation website: https://fearlesssalarynegotiation.com
  • Fearless Salary Negotiation: https://www.amazon.com/Fearless-Salary-Negotiation-step-step/dp/0692568689/
  • Twitter: https://twitter.com/joshdoody
  • LinkedIn: https://www.linkedin.com/in/joshdoody/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Human-scale teams use Tailscale to build trusted networks. Tailscale Funnel is a great way to share a local service with your team for collaboration, testing, and experimentation. Funnel securely exposes your dev environment at a stable URL, complete with auto-provisioned TLS certificates. Use it from the command line or the new VS Code extensions. In a few keystrokes, you can securely expose a local port to the internet, right from the IDE.

I did this in a talk I gave at Tailscale Up, their first inaugural developer conference. I used it to present my slides and only revealed that that’s what I was doing at the end of it. It’s awesome, it works! Check it out!

Their free plan now includes 3 users & 100 devices. Try it at snark.cloud/tailscalescream

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’m joined by recurring guest and friend Josh Doody, who among oh, so many things, is the owner of fearlesssalarynegotiation.com, and basically does exactly what it says on the tin. Josh, great to talk to you again.

Josh: Hey, Corey. Thanks for having me back. I appreciate it and I’m glad to be here.

Corey: So, you are, for those who have not heard me evangelize what you do—which is fine. No one listens to all of the backlog of episodes and whatnot—you are a salary negotiation coach, and you emphasize working with high earners who are negotiating new job offers, which is basically awesome. How did you stumble into this?

Josh: Yeah, a good question. Really, it started as what I would say is a series of interesting career choices that I made, where I started as an engineer. I was pretty quickly bored in engineering and I switched to—I wanted to be customer-facing and do stuff that had impact on the business, so I did that and ended up working for a software company that made HR software that happened to do among other things, compensation planning. And so, I kind of started learning how it worked behind the scenes.

And then over time, I started wising up and negotiating my own job offers. And noticed that wow that kind of worked pretty well, and I decided to write a book about it, a hundred percent just because I like to write stuff. I’ve been writing for 20 years on the internet, and I decided, why not just write a book about this? You know, five or six people will buy it, my mom will love it, I’ll get it out there and it’ll feel really good.

And then people started reading the book and asking me if they could hire me to do the methodology in the book for them. And I said, “Sure.”

Corey: When people try to give you money, say yes.

Josh: Yeah. Okay, you know, whatever, you know? My first person that ever hired me asked me what my rate was, and I didn’t have a rate because I had never considered doing that before. But she was a freelance writer and I said, “Well, whatever your rate is, that’s my rate.” [laugh]. So, that was my first rate that I charged someone.

And yeah, from there just, it took off as more people started hiring me. A number of friends were chirping in my ear that hey, you know, this seems like a really valuable thing that you’re doing and people are coming out of the woodwork to ask you to do it for them. Maybe you should do that thing instead of the other things you’re doing and trying to sell copies of the book and stuff like that. Like, why don’t you just be a salary negotiation coach? That was, I don’t know, like, seven years ago now, and here I am.

Corey: I don’t know if I ever told you this, but back when we met in the fall of 2016, I was trying to figure out what windmill I was going to tilt at before I stumbled upon the idea of AWS billing as being one of them. I thought that writing a book and being a sort of a coach of sorts on how to do job interviews with an emphasis, of course, on salary negotiation, would be a great topic for me because I’ve done it an awful lot. This is a byproduct of getting fired all the time because of my mouth. And then I started talking to you and my reaction was, “Oh, Josh is way better at this than I am. No, I’m going to go find something else instead.”

And now the world is what it is, and honestly, at this point, all the cloud providers really wish you hadn’t been there at that point in time because then they wouldn’t have to deal with the nonsense that I present to them now. But I always had a high opinion of what you do, just because it is in such a sweet spot where if I were to shut this place down and get a quote-unquote, “Real job” somewhere, I would hire you. And it’s not that I intellectually don’t know how to negotiate. Half my consulting now is negotiating large AWS contracts on behalf of AWS customers with AWS. A lot of these things tend to apply and go very hand-in-glove.

But there’s something to be said for having someone who sees this all the time in a consistent ongoing basis, who is able to be dispassionate. Because when you’re coaching someone, it’s not you in the same boat. For you, it’s okay, you want to have a happy customer, obviously, but for your client, it’s suddenly, wow, this is the next stage of my career. This matters. The stakes are infinitely higher for them than they are for you.

And that means you have the luxury of taking a step back and recognizing a bad deal when you see one. There is such value to that I can’t imagine not engaging you or someone like you the next time that I would go about changing jobs. Although these days, it’s probably an acquisition or I finally succumb to a cease and desist. I don’t really know that I’m employable anymore.

Josh: [laugh]. Yeah, I mean, you said a lot of really interesting things there. I think a common theme—you know, to work with me, there’s a short application that people fill out, and very frequently in the application, there are a couple of open-ended questions about you know, how can I help you? What’s your number one concern? That kind of stuff.

And frequently, they’ll say, “Yeah, I’ve negotiated before and I actually did okay, but I want to work with a professional this time,” is the gist of it, for I think reasons that you mentioned. And one of them is, there’s just a difference between negotiating for yourself and feeling all of that pressure and having somebody who can just objectively look at it and say, “No, I think you should ask for this instead.” Or, “No, I don’t think that you should give that information to the recruiter.” And the person instead of feeling, you know, personal subjective pressure can just say, “Well, the objective person that I hired and paid money to help me with this says, ‘don’t do that,’ or ‘do this instead,’ and it’s easier for me to just trust what they’re doing as a professional and let me be a professional at the other things that I’m a professional at.”

And so yeah, I think that’s a lot of—you know, for some people, it’s, “I have no idea how to negotiate. I don’t want to screw this up. Please help me, Josh.” And for some people, it’s, “Yeah, I’ve done this before. I did, okay, but I want you here to help me do this.”

And that includes people who come back and work with me two or three times. They know the methodology. They’ve been through it literally with me, and I’m very open about what we’re doing and why I’m collaborative with my clients. We’re talking about the decisions we make. I will bounce things off of them.

I’ll say, “Here’s what I think we should do. What does your intuition tell you about that? How do you feel about it?” Because it’s important to me because they’re in the game and I need to know what they think. And they’ll come back to me and we’ll do it again. They already know the playbook. And I think that’s because it’s easier to just have somebody who’s a professional there to objectively tell you, “You’re not asking for enough.” Or, “Did you think about asking for this instead?” Or, “Do you really care about that thing?” Stuff like that.

Corey: There is so much value to that, just because it’s a what’s normal in this? Because I’m sure you’ve seen before where—I’m probably—I should put this in more of a question, but I already know the answer because I’ve seen it just from people randomly sending me things out on the internet—of their times for companies say or ask for things that are just absolute clown shoes. It’s, I would barely consider it professional at that. It always feels like there’s value in being able to talk to someone who sees this all the time who can say, “Hold on. That is absolutely not normal. That is not a reasonable question. That is not an expectation that any sensible person is going to have.” Because the failure mode otherwise is you think it’s you.

Josh: Yeah, part of my value prop is, you know, I know how to negotiate with companies. I’m not afraid of them. I’ve negotiated with Fortune 5 companies, come out way ahead—just as you do frequently—and I know the playbook that they’re running. But part of it also is, you know, I have a compendium of recruiter responses. I know what they say, I know what their words mean, and so I can say things like, “Oh, here’s what they actually mean when they ask you for that.”

Or I can say, “That’s weird.” Which, you know, if I’ve done 20 negotiations with this company and all of a sudden a recruiter says something that’s weird, that makes my ears perk up and makes me wonder why. And so, I can dig in on my side and try and figure out what’s going on, see if we tripped some wire that I didn’t see or, you know, something like that. So, that’s part of the value too, is just all the reps that I’ve had, even like you said, I’m sure that you would do a wonderful job negotiating; I’ve talked to you about negotiating online and off, and I know that you know the game, you know how to do it, for your day job but also for compensation. But I probably have more reps negotiating with those companies than you do and therefore my compendium is a little bit deeper, so there might be things that I could recognize that you would not recognize that I could see, right, in the similar way that in your negotiation world that there are things that I certainly would not recognize that you would catch on to.

And I think that can be a very valuable thing. There could be something a recruiter says where I recognize, “Aha. That’s a technical term or that’s a key phrase that we can grab onto. And that is an opportunity to get more.”

Corey: Or, “What are you making now?” It’s like, yes, that’s the industry accepted one free pass that’s screwing the candidate. Yeah—

Josh: Right.

Corey: —let’s not do that.

Josh: Right. And we’re—here’s how to sidestep it and here’s what happens when they ask for it for the fourth time, and here’s what happens when they say the magic words and, you know, all that stuff. So yeah, a lot of it is just getting reps. It started with let me just run my playbook and then as I run the playbook, I get more data every time I do it, and I get to learn what the edge cases look like, and how to spot, you know, weird funky stuff coming from recruiters and that sort of thing.

Corey: One aspect of this that has been, I guess, capturing my imagination since you first talked to me about it, and I am certain I’m going to butcher this into something that sounds insulting and demeaning, which sort of cuts against the entire point. Specifically, the idea of a positive language, or, the term you used was ‘Positively Persuasive.’ What is that? Because it sounds like it’s just someone who’s setting me up, like, waving raw steak in front of a tiger, like, “Please maul me on this.” But there is more to it than that.

Josh: [laugh]. Yeah, so this is something that, to be honest with you, I have done almost intuitively throughout my career, but certainly as a salary negotiation coach. And what it is, is a tendency to use positive, meaning, you know, not negative words. So like, essentially, if you’re familiar at all with improv, which I would say probably half of the people listening probably have some idea what I’m talking about, you take improv classes, and they teach you an exercise called Yes, And. And the reason you do Yes, And is, you know, Corey says something wacky and I could shut it down.

I could say, “That’s not true.” You know, “My hair isn’t red.” And then we’re done improving. But if Corey says, “Josh your hair is red,” even if my hair is not red and I say, “Yes, and… it’s on fire right now,” then we have something going, right? And so, using those positive words—yes, and is a positive way of responding to that—opens up a further dialog and also makes it easier for you to engage with me in that improvisation. In a way, a negotiation is an improvisation; they’re all going to be different.

A business conversation is going to be an improvisation. It’s rare that you’re going to have a conversation where you could write the script completely before the conversation starts. Often there will be an opportunity to improv, to do something different. And so, positively persuasive is essentially my way of thinking about how to use those positive words to accomplish an objective while building rapport with the person that you’re talking to, and leaving the door open for that kind of positive collaboration and improvisation where you can work together with your co-party, with the person that you’re talking to in the negotiation. And so, that’s super abstract, and a concrete example of this would be for example, in a counteroffer email.

Frequently people will, kind of unsolicited, just send me their counteroffer emails. “I’m writing up this email. What do you think?” Somebody on my newsletter or my email list or something. And sometimes they’re okay, and sometimes it’s like, they’re giving an ultimatum and they’re saying, “You promised this when we first talked on the phone and you’re not giving me that. You offered me this and I want what you offered to start with.”

And they’re using all these negative words: “You promised this and didn’t give it to me.” “That’s not what I expected.” Whereas in the counteroffers that I’m writing, it says, “Hey, thanks for the offer.” Starts right away with something that looks like a throwaway line, a platitude, but really what it is is saying, “Hey, we’re on the same team here. We’re collaborating. Thanks for the offer. I appreciate it and I hope you’re having a good week so far.”

And then as it goes on, it says, “Here are the reasons that I’m super valuable to your team. I can’t wait to join this team and, you know, express that value.” And then, “You offered $100,000. I would be more comfortable if we could settle on $115,000.” And so, that’s a counteroffer. In some cases, the counter will be more than 15%. That’s kind of a middle-of-the-road one, but the way I say it is, “I would be more comfortable if,” and so there’s no sort of in-your-face, there’s no ultimatum, there’s no fist pounding on the desk—

Corey: There’s no, “No.” There’s no, “This is not acceptable.” There’s no, “I won’t accept this.” It’s a very soft approach that generally doesn’t put people on edge.

Josh: Puts it—it not only doesn’t put them on edge, but you’re sort of putting your arm around them saying, “Hey, you know, I’d be more comfortable if we could do this.” And they’re like, “Okay, you know, let me see what I can do for you.” So, you’re not making—you’re not turning them into, you know, an enemy combatant; you’re turning them into a collaborator. And now it’s you and then working together to try to make you comfortable so that you can join their team. So, that’s a subtle thing that happens in a counteroffer email and numerous other places.

But that’s the idea is that when you can, you’re choosing positive language so that your requests will be received better, so you build rapport with the person that you’re negotiating with, and so that they perceive you to be a collaborator and not an opponent.

Corey: It sounds hokey, but I’ve also watched it work. It’s weird in that we hear about things like this, we think, “Oh, that wouldn’t work on me at all,” except it the evidence very clearly shows that it does. There’s a reason that some people are considered charismatic and I think this is a large part of it. And I also wonder, I mean, you focus on salary negotiation for high earners, and that, historically at least, as included, you know, a fair few number of software developers and whatnot. And these days, let’s be very clear that communicating what you want, clearly, concisely, and in an understandable way that something or someone can action is such a lost foreign skill for some of these people that they call the entire field ‘prompt engineering’ because just communicate clearly is apparently a microaggression when you ask an engineer to do it without giving it a fancy name. Improved communication really feels like it has been part of a dawning awareness lately that, wait, this is actually important, not just one of those box-checking items that you say so that people don’t spit in your food.

Josh: I think you’re a hundred percent right about that. I mean, it’s interesting is you think about, you know, forms of communication that we have kind of experienced over the past, you know, however many years. But you know, at first, there was no writing, over, you know, thousands of years ago, or whatever, it was just all kind of oral tradition. And then we had writing and it was, like, long-form writing. And then, you know, fast forward to today and it’s like you’re sending a text with two letters and that means something right, or I’m about to head to my friend’s house, and I text him three letters: OMW, right?

It’s like, extremely terse, direct, and to the point. And there is a place for that, I think. I think that efficiency probably has some benefits. I mean, there’s not a lot of reason for me to spend six minutes, you know, writing a text to tell somebody that I’m heading to their house. But on the other hand, I think that sort of concision, that terse writing can also lose a lot in translation, and as we’re using more media that look like Slack, or Discord, or these other chat-based ways of communicating—including email, by the way; I mean, email can be a place where you can be as terse, or I guess, as pleonastic as you’d like—and you get more and more words in there.

And so, I think it’s important to be intentional with those words in contexts where tone and meaning and intent can matter. And a lot of that is in interpersonal communication. And again, it’s about how messages are received and what you’re conveying. I use a lot of—this is [laugh] not directly related—I use a lot of emoji and emoticons and stuff like that and I do that because I’m trying to convey tone in a medium that doesn’t really facilitate it, right?

If I’m talking to you, you and I can see each other’s faces right now, so you know if I’m being sarcastic, or telling a joke, or being very serious. And so, in emails, I’ll put a smiley face. And that’s me saying, “Hey, I’m not laying this on real thick. I’m just letting you know.” Right? So anyway, there are so many media that are available to us now that make it hard to convey tone that I think a lot of it is you’ve got to be intentional with your tone.

Corey: I have worked with more people over the course of my career that have what I’ve taken the call being the asshole-in-email problem, where I have—I think these people are just these absolute jerks. They are completely onerous to deal with and I despise dealing with these people, but then I’ll sit down with them and they are the nicest people and they are incredibly competent and effective. They just have a challenge where whatever they write an email, it sounds like there’s an implicit, “Listen up here dickhead…” that they’re starting the email with.

Josh: [laugh]. Yeah.

Corey: And, “You know what your problem is…” may as well be how they open these things. And it feels like effectively communicating and tone is becoming something of a lost art. I’ve talked to multiple people now who will wind up using Chat-Gippity to construct the bones of a work email and then they’ll just change a sentence or two in the center that actually is the substantive thing that they want to send so it winds up handling all the window-dressing there. Now, I’m wondering what the other side is going to look like when you have someone using Chat-Gippity to paste a work email into it. It’s like, “Okay, strip out the flattery. What are they actually asking from me here?” So, you effectively have, like, an API layer of padding provided by computers, where you could just like, say, the direct thing, but it comes with all the flowerly accouterments that has become expected in business correspondence.

Josh: Yeah. I mean, I love everything that you said there. It’s true. I mean, I’ve worked with people in the past where they would send me an email, or I would email with them frequently and then we were talking in person, I realized that oh, I totally misread what they were saying. Like, I misread what they meant to say, I misread what their outcome, their preferred outcome was, and it’s because the tone is just lost in email.

And I don’t think it was necessarily due to any sort of deficiency on their side. It was on—they have a way of communicating, I have a way of perceiving communications, and they were different, and so the message that I got was different. So, I think a lot of what I’m talking about with positively persuasive is how do I communicate in a way where it is not ambiguous, where it is very clear what I’m saying, what my intent is, what my tone is. And sometimes, like you said, [laugh] use ChatGPT to, like, strip out the flattery. I put the flattery in because I want them to know, like, “Look, I know that you’re a person. You and I are on the same team here. We’re working together.”

So, a lot of my emails will open with, “Hey, hope you’re having a good day.” And it’s like, do I care if they’re having a good day? Yeah, but I don’t need to say that out loud. The reason I’m saying it out loud is I want them—the opposite of everything you just described where I want them to read that email and think, “Okay, Josh isn’t coming at me. Even if he does have critiques of something that I’m doing, or he has a suggestion to improve something, he’s coming at it from the place of, ‘Hey, I hope you’re having a good day so far.’” Whatever I say at the beginning of the email.

And so, that’s filler, a hundred percent, but it’s filler with a purpose that is meant to convey the tone of the email, that is, I’m not coming down on you too hard. I’m trying to convey a message or ask a question and sincerely curious, and can we come together on this to figure out what the solution is or to move forward or to find the next steps or whatever the thing is that we’re trying to do?

Corey: It feels like this is an area that has massive application beyond the obvious negotiation piece of it, which is fundamentally where we sit down and try and convince people to do a thing that we want them to do that is in our interest. But it’s like, okay, well, that’s not just negotiation. That is, on some level, a disturbing number of human interactions that we tend to have. Where do you see this being applied? Is it something that just—that you’re looking at just through a lens of communicating effectively in a salary negotiation, or does it extend beyond that to your worldview?

Josh: I think it can get pretty broad. I mean, as you were describing, I was thinking kind of, as you were talking, like, when else do I use this? And the answer is a lot. But one place that I use this kind of thing a lot is when I’m emailing people who I don’t know, and trying to get them to either just give me something or to allow me more leeway than they otherwise necessarily have to allow. And so for—here, I’ll give you an example, which is, I recently switched homeowners insurance providers because I live in Florida and homeowner’s insurance in Florida is a nightmare.

And so, I changed providers. I thought I had crossed all my t’s and dotted all my I’s, but there was something that fell between the cracks, and that is that the mortgage holder—the bank that holds my mortgage—hadn’t sent the premium check to my new insurance provider. They didn’t get that memo. And it was essentially my responsibility, but I kind of goofed. So, the bank writes me an email and they say, “Hey, we see you changed providers but we don’t have an address for them. We can’t send them a check. Can you give it to me?”

And so, now I’m—there are two parties that I have to kind of keep on my side. One of them is obviously the bank, but also the insurance provider, who might be mad at me because I’m ten days late on this premium or whatever. So, my emails to them are places where I use this where it’s like, I’m basically going to make it so that the person who could get mad at me and cause me some kind of detriment is going to have to do it through a really thick cloud of, “Josh is a nice guy who isn’t trying to be a jerk to anybody here. He’s not trying to pull one over on anybody. There was an honest mistake that was made, he’s just trying to make everything right, and he’s hoping that I can help them.”

And they’re going to have to look at the way that I communicate with them and they’re going to have to push through it and say, “Nope. I’m going to be a jerk. I’m going to follow the letter of the law or I’m going to be as punitive as I can be.” That’s really hard to do when somebody like me is emailing, say, “Hey, listen, I know that we were supposed to get a check out to you last week. I’m working on it right now. I’ve already got everything to the bank. It’s going to be overnighted to you tonight. Is there anything else I could do to make this easy for you on your side?”

And then they’re going to be like, “No, just, you know, as soon as we get it, we’ll let you know.” Whereas if I’m, like, you know, mad at them or I’m mad at somebody or I’m being a jerk in email, then they don’t really have any reason to not be as punitive as they can be to me. And so, that’s just—it’s a little manipulative, I guess, but it’s also the way that I see life, right? Like, I’m like that with everyone, including people who are on the other side of that equation. I’m going to give them grace when I can.

And so, it’s a way of me saying, “Hey, can you extend some grace to me? I think you’re a human being who’s on the other side of this and you have a job to do and I understand that, and if you could be a little bit kind to me, that would be great.” And it works almost every time.

Corey: This episode is sponsored in part by Panoptica. Panoptica simplifies container deployment, monitoring, and security, protecting the entire application stack from build to runtime. Scalable across clusters and multi-cloud environments, Panoptica secures containers, serverless APIs, and Kubernetes with a unified view, reducing operational complexity and promoting collaboration by integrating with commonly used developer, SRE, and SecOps tools. Panoptica ensures compliance with regulatory mandates and CIS benchmarks for best practice conformity. Privacy teams can monitor API traffic and identify sensitive data, while identifying open-source components vulnerable to attacks that require patching. Proactively addressing security issues with Panoptica allows businesses to focus on mitigating critical risks and protecting their interests. Learn more about Panoptica today at panoptica.app.

Corey: There’s value as well, even everyday customer service interactions, if I have a bad customer experience buying something off of Amazon—I know, imagine that.j could that ever happen? Of course not. But in a magical world in which in hypothetically did, I can call up and they answer the phone, I’m probably going to be pretty steamed going into that conversation because this is effort I didn’t want to have to deal with. But stop and think about it for a second. Usually, when I call Amazon for a variety of things, it’s not Andy Jassy who’s answering the phone. Those are atypical moments for me.

Instead, it is generally some poor customer service schmo, who is basically given zero amount of autonomy to speak of in the course of their job, and surprisingly, does not set Amazon’s strategic priorities for them. And if I unload on this person, maybe I make myself feel better, I’ve made someone else’s day actively worse, but even if you want to set aside the story of being a good person—which I don’t suggest people do—but view it in a purely Machiavellian self-serving way, you’re still going to have a better outcome if you inspire people to like you by making yourself likable. Because when you’re a jerk—and I used to work helpdesk; I remember how this works—

Josh: Me, too.

Corey: Suddenly, I will fall back on every policy that I can have, “Oh, we’re not allowed to sit through a reboot. Bye.” As opposed to, “Eh, [unintelligible 00:22:31] say ever not to, but I’m enjoying this and I want to help you out and make sure you get there, so hang out. Why not?” There are ways people can bend the rules in your favor, but if you give them an excuse to fall back on that, they’re not going to go out of their way to help you at all. They’re going to make you go through every bit of procedural red tape they can possibly come up with. And again, you’ve made their day worse and that should not be lost on you. The outcomes are better for everyone when you’re a nice person.

Josh: As you were talking, it’s funny because I remembered, maybe the most frustrated I’ve ever been talking to customer service. This is several years—many years ago, but I had some student loan stuff going on. I don’t even remember specifically what it was, but it had to do with, you know, who was servicing the loan and I’m trying to pay off a loan and I can’t get the right person on the phone and they say, you know, “It’s this other place that owns that holds the loan.” Or, “You need to call this person,” and I’m getting the runaround and I’m not able to do the thing I want to do.

And after I think I’ve been hung up on, like, three times, and I was really steamed. Like you said, I’m legitimately, like, very frustrated. My voice had been cracking a little bit, which is how I know I’m, like, really getting heated is my [laugh] voice will start to crack a little bit. But I said to the person—and I became conscious in that moment of like, okay, I’m very frustrated. I could say something I regret I could really, like, hurt this person that I’m talking to.

As you said, they’re just somebody who’s a customer service representative for this bank or loan servicer, whoever they were. So, I said something like, “Listen, you can’t hear it in the tone of my voice right now, but I need you to know that I’m extremely frustrated and I’m going to [laugh] I’m going to get really upset, and so I’m asking you to help me before I do that before I escalate. I don’t want to talk to your manager, but I’m going to ask you to do that if you don’t help me right now. And you should know that I’m super frustrated. My voice is not betraying that right now, but understand that I am.”

And they snapped in and they were like, “Okay, I get it, I get it,” you know? And right there even as a place where I could have just started shouting at them or whatever it takes, you know, “I want to talk to your manager,” and, “I’m going to escalate,” and all this stuff. And instead, I was like, well, I’m going to give them one last chance, which is, let me just tell them how frustrated I am without using colorful language or mean words. And it worked. It was a subtle thing that actually, I think it got their attention more than anything else. They said, “Oh, this person is really angry. I should actually listen to them.”

Corey: Now, there is a dark side to this as well and that is human nature. I have done experiments on this over the years, most notably on Twitter, back when that was the central place people went to, and when I would say something nice about an AWS service, it got in most cases two likes and maybe a bot would retweet it. Whereas if I say, “This AWS service is a piece of garbage,” and I come up with some reason for it, it went around the internet three times and it was misconstrued, with me saying, “The entirety of AWS is terrible.” Not usually, no. There are some frustrating elements, but yeah, there’s context. It doesn’t fit into a single tweet.

The snarky negativity blows up and responds to—and resonates with something in human nature that the people love spreading that around and engaging with it, whereas the happy positivity does not work that way. On Twitter. I’ve noticed what seems to be the opposite effect on LinkedIn. Snark doesn’t do well over there, but almost saccharine-sweet sincerity does. And I don’t know what this says about various social media channels or human nature or what. All I know is that I’m confused.

Josh: I think you’re right. You know, I mean, as you were talking, I was thinking about clickbait, right? Like, there’s a reason that clickbait is called what it is, and it’s because you read it and you get annoyed or frustrated or angry, and I’m going to hate-read this article right now and I’m going to send it to six friends. There is something in human nature. I mean, you know, we talked—for decades, I’ve heard about how the local news is our news, “If it bleeds, it leads,” in news, right?

We’re not talking about how great the planet is or how things—like, this bad thing happened in New Orleans yesterday and you should be really upset about it, or wherever that place happens to be on that particular day. I do think there is something innate in us that allows us to gravitate towards those kinds of things and I have no idea what it is. But it is interesting, as you said, that there are places where either that’s frowned upon or there’s just a different mode of communication, which tells me that there’s something sort of in the cultural water there that causes people to perceive stuff differently in different kinds of social media environments, right? Twitter definitely is a place where things can go pretty negative. And there are other places that are significantly more negative, right, on the internet, if you want to go, they get really bad, and then there’s places that are really positive.

And it’s interesting how it’s like a maybe people self-select into those places, but also, I think, you know, I think there’s a big difference if you think about, like, who’s using Twitter and why and who’s using LinkedIn and why. I think that people correctly perceive on LinkedIn that for the most part, you’re probably not going to be somebody that’s at the top of a bunch of lists to be hired if your whole thing on LinkedIn is just being negative all the time and doom and gloom and snark and that kind of thing. It’ll be entertaining to some people, but you’re probably not going to get many job offers based on that because people are going to ask, “Do I want to work with this person 24 hours a day?” And they’ll read your posts and say, “No,” whereas at least a saccharine sweet person, everybody knows those people who are like that in real life, and they can be I don’t know, a little bit much, but also can generally be very good people to work with and it’s not difficult to sort of like manage that.

Corey: There’s a lot that can be done just by having people want to help you. It’s weird. Like, I take a look at some of the people that I identify publicly as the nicest in tech—Mark [unintelligible 00:27:48] is a good example. Kelsey Hightower is sort of the canonical example of all of this. These are just genuinely nice people. Ashley Willis, another good example.

There are so many different folks out there who are just beacons of positivity. And I look at that, and it’s like, first, that is admirable. Second, holy hell that is absolutely not me. No one is ever going to say, “That’s what I love about Corey. He’s so uplifting and positive all the time.” You know, I do strive to be a better person and inspire others to be better people, but I’m also willing to spare no quarter for corporate tomfoolery either. Which apparently means a lot of people think you’re a jerk as a result. I’ll take it.

Josh: Yeah, I think it’s, you know, everybody—that’s the nice thing about humans, right, is we’re all different. And there are lots of different types of person—if everybody had the same personality, what a boring place that we would live. And that’s true for, more or less, any human characteristic. If we were all the same and vanilla, I think it would be pretty boring. So, I think that having really positive people out there is great, and having some people who are snarky is great, and having people who have, you know, an ability to just point out absurdity is great. If everyone is pointing out absurdity all the time, then we’re not left with too much.

So, I do think it’s good that those people are out there and they’re very positive. And I think that, you know, even for myself, like, I try to be positive and helpful. Like, we were talking about customer service. I’m like, overly nice to customer service people. I tip more than I should most of the time. And a lot of that is just, you know, that’s a human; they have needs and feelings and this is a way for me to be kind to them.

And I know most people don’t think that way that I do. And I like that. And I think that some people don’t think that way and I think that’s totally fine, too. I think the variety is the spice of life and I think that makes it interesting and useful. I also think that being intentional with those different modes, having them all available to you, and exercising them in different environments can be, like, a level-up, right? It can be a superpower.

You can either be a person with a personality who exercises that same personality all the time, or you can choose to exercise, sort of, different personalities or different ways of communicating or different levels of positivity or negativity in different environments. And I think that makes it even more interesting where you’re able to essentially be a chameleon and find the right mode of communication for the environment or the situation that you’re in, which can enhance that situation for you or for other people that are around you.

Corey: I have to ask, do you find that this is something you do all the time or do you put on your negotiating phrasing the same way that I do when my children accuse me of putting on ‘podcast voice.’

Josh: All the time, definitely not. I am aware of it as a way of communicating that’s available to me and I do consciously use it a lot of the time. But you know, if I’m just sitting around with my buddies on, you know, Wednesday night watching the game, probably not. And a lot of that is because, you know, part of this is, it’s a default to positive because you don’t know sort of who’s on the other end of the line, whereas if you’re communicating with somebody that you’ve communicated with for hundreds of hours, you don’t need all that stuff, you don’t need all the tonal indicators and the padding and all that stuff because you know that person. So, a lot of what I’m describing, even like in a salary negotiation, I’m basically working from the default of I don’t know the counterparty, I don’t know the recruiter, and therefore we’re going to default to positive, and that’s going to essentially, you know, make things smoother.

It’s going to remove friction because there are things that I don’t know, whereas, you know, if I’m communicating with somebody I know really well for 20 years, we don’t need all that stuff. We can—that’s where the shorthand can come in handy. It can be really useful because we already know all of the background there. One place that I’m very conscious of this is, you know, every now and then somebody, with a personal friend or somebody that I know, well, I’ll have, like, a difficult conversation where they’ll say, “Hey, you know, this is something that happened to me recently. Can you help me out?” Or, “This is a difficult thing that I’m going through.”

And that’s a place where I am very conscious of this and it comes in different ways. One of them is using positive words, but one of them is also just, like, exercising extreme sympathy or empathy if it’s appropriate. Which is, again, it’s a conscious decision to say, this isn’t a time to point out, you know, for example, errors, or like, this person just needs someone that they want to talk to and I’m going to listen to them carefully, I’m going to try to give them reassurance that the situation will be resolved eventually, and that kind of thing, but it’s not a time for you know, critique or, you know, negative words or pointing out flaws and that kind of thing. And so, I think that’s also kind of a conscious place that I will exercise it. But to answer your question, no, I don’t do this all the time.

I would say without having ever thought about this before, the less familiar I am with the person or the situation, the more I will default to this, and the more familiar I am with the person or the situation, the less I will default to it. And I will just use more plain, kind of, direct language because that familiarity is there, and it assumes a lot that isn’t there when I don’t know the person well.

Corey: I really want to thank you for taking the time to speak with me about this. Where can people go to learn more?

Josh: Maybe follow me on Twitter [laugh], @joshdoody on Twitter.

Corey: It’s a harder problem these days than it once was.

Josh: Yeah. I really paused there. I am pretty active on LinkedIn these days. And fearlesssalarynegotiation.com isn’t explicitly about positive language or being positively persuasive, but you’ll see even just reading the articles that I write there that underlying most of what I write is this sort of implicit understanding that positivity is the way to make progress and to get closer to what your goals are. So, @joshdoody on Twitter; joshdoody on LinkedIn, of course, and then fearlesssalarynegotiation.com.

Corey: And we will, of course, put links to all of this in the show notes. Thank you so much for taking the time to speak with me. I appreciate it.

Josh: Thanks for having me on, Corey. This was a lot of fun. I always like talking to you.

Corey: I do, too. Josh [laugh] Doody, owner of Fearless Salary Negotiation. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry comment that rants itself sick, but also only uses positive language.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

View Details

Matthew Prince, Co-founder & CEO at Cloudflare, joins Corey on Screaming in the Cloud to discuss how and why Cloudflare is working to solve some of the Internet’s biggest problems. Matthew reveals some of his biggest issues with cloud providers, including the tendency to charge more for egress than ingress and the fact that the various clouds don’t compete on a feature vs. feature basis. Corey and Matthew also discuss how Cloudflare is working to change those issues so the Internet is a better and more secure place. Matthew also discusses how transparency has been key to winning trust in the community and among Cloudflare’s customers, and how he hopes the Internet and cloud providers will evolve over time.

About Matthew

Matthew Prince is co-founder and CEO of Cloudflare. Cloudflare’s mission is to help build a better Internet. Today the company runs one of the world's largest networks, which spans more than 200 cities in over 100 countries. Matthew is a World Economic Forum Technology Pioneer, a member of the Council on Foreign Relations, winner of the 2011 Tech Fellow Award, and serves on the Board of Advisors for the Center for Information Technology and Privacy Law. Matthew holds an MBA from Harvard Business School where he was a George F. Baker Scholar and awarded the Dubilier Prize for Entrepreneurship. He is a member of the Illinois Bar, and earned his J.D. from the University of Chicago and B.A. in English Literature and Computer Science from Trinity College. He’s also the co-creator of Project Honey Pot, the largest community of webmasters tracking online fraud and abuse.

Links Referenced:

  • Cloudflare: https://www.cloudflare.com/
  • Twitter: https://twitter.com/eastdakota

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. One of the things we talk about here, an awful lot is cloud providers. There sure are a lot of them, and there’s the usual suspects that you would tend to expect with to come up, and there are companies that work within their ecosystem. And then there are the enigmas.

Today, I’m talking to returning guest Matthew Prince, Cloudflare CEO and co-founder, who… well first, welcome back, Matthew. I appreciate your taking the time to come and suffer the slings and arrows a second time.

Matthew: Corey, thanks for having me.

Corey: What I’m trying to do at the moment is figure out where Cloudflare lives in the context of the broad ecosystem because you folks have released an awful lot. You had this vaporware-style announcement of R2, which was an S3 competitor, that then turned out to be real. And oh, it’s always interesting, when vapor congeals into something that actually exists. Cloudflare Workers have been around for a while and I find that they become more capable every time I turn around. You have Cloudflare Tunnel which, to my understanding, is effectively a VPN without the VPN overhead. And it feels that you are coming at building a cloud provider almost from the other side than the traditional cloud provider path. Is it accurate? Am I missing something obvious? How do you see yourselves?

Matthew: Hey, you know, I think that, you know, you can often tell a lot about a company by what they measure and what they measure themselves by. And so, if you’re at a traditional, you know, hyperscale public cloud, an AWS or a Microsoft Azure or a Google Cloud, the key KPI that they focus on is how much of a customer’s data are they hoarding, effectively? They’re all hoarding clouds, fundamentally. Whereas at Cloudflare, we focus on something of it’s very different, which is, how effectively are we moving a customer’s data from one place to another? And so, while the traditional hyperscale public clouds are all focused on keeping your data and making sure that they have as much of it, what we’re really focused on is how do we make sure your data is wherever you need it to be and how do we connect all of the various things together?

So, I think it’s exactly right, where we start with a network and are kind of building more functions on top of that network, whereas other companies start really with a database—the traditional hyperscale public clouds—and the network is sort of an afterthought on top of it, just you know, a cost center on what they’re delivering. And I think that describes a lot of the difference between us and everyone else. And so oftentimes, we work very much in conjunction with. A lot of our customers use hyperscale public clouds and Cloudflare, but increasingly, there are certain applications, there’s certain data that just makes sense to live inside the network itself, and in those cases, customers are using things like R2, they’re using our Workers platform in order to be able to build applications that will be available everywhere around the world and incredibly performant. And I think that is fundamentally the difference. We’re all about moving data between places, making sure it’s available everywhere, whereas the traditional hyperscale public clouds are all about hoarding that data in one place.

Corey: I want to clarify that when you say hoard, I think of this, from my position as a cloud economist, as effectively in an economic story where hoarding the data, they get to charge you for hosting it, they get to charge you serious prices for egress. I’ve had people mishear that before in a variety of ways, usually distilled down to, “Oh, and their data mining all of their customers’ data.” And I want to make sure that that’s not the direction that you intend the term to be used. If it is, then great, we can talk about that, too. I just want to make sure that I don’t get letters because God forbid we get letters for things that we say in the public.

Matthew: No, I mean, I had an aunt who was a hoarder and she collected every piece of everything and stored it somewhere in her tiny little apartment in the panhandle of Florida. I don’t think she looked at any of it and for the most part, I don’t think that AWS or Google or Microsoft are really using your data in any way that’s nefarious, but they’re definitely not going to make it easy for you to get it out of those places; they’re going to make it very, very expensive. And again, what they’re measuring is how much of a customer’s data are they holding onto whereas at Cloudflare we’re measuring how much can we enable you to move your data around and connected wherever you need it. And again, I think that that kind of gets to the fundamental difference between how we think of the world and how I think the hyperscale public clouds thing of the world. And it also gets to where are the places where it makes sense to use Cloudflare, and where are the places that it makes sense to use an AWS or Google Cloud or Microsoft Azure.

Corey: So, I have to ask, and this gets into the origin story trope a bit, but what radicalized you? For me, it was the realization one day that I could download two terabytes of data from S3 once, and it would cost significantly more than having Amazon.com ship me a two-terabyte hard drive from their store.

Matthew: I think that—so Cloudflare started with the basic idea that the internet’s not as good as it should be. If we all knew what the internet was going to be used for and what we’re all going to depend on it for, we would have made very different decisions in how it was designed. And we would have made sure that security was built in from day one, we would have—you know, the internet is very reliable and available, but there are now airplanes that can’t land if the internet goes offline, they are shopping transactions shut down if the internet goes offline. And so, I don’t think we understood—we made it available to some extent, but not nearly to the level that we all now depend on it. And it wasn’t as fast or as efficient as it possibly could be. It’s still very dependent on the geography of where data is located.

And so, Cloudflare started out by saying, “Can we fix that? Can we go back and effectively patch the internet and make it what it should have been when we set down the original protocols in the ’60s, ’70s, and ’80s?” But can we go back and say, can we build a new, sort of, overlay on the internet that solves those problems: make it more secure, make it more reliable, make it faster and more efficient? And so, I think that that’s where we started, and as a result of, again, starting from that place, it just made fundamental sense that our job was, how do you move data from one place to another and do it in all of those ways? And so, where I think that, again, the hyperscale public clouds measure themselves by how much of a customer’s data are they hoarding; we measure ourselves by how easy are we making it to securely, reliably, and efficiently move any piece of data from one place to another.

And so, I guess, that is radical compared to some of the business models of the traditional cloud providers, but it just seems like what the internet should be. And that’s our North Star and that’s what just continues to drive us and I think is a big reason why more and more customers continue to rely on Cloudflare.

Corey: The thing that irks me potentially the most in the entire broad strokes of cloud is how the actions of the existing hyperscalers have reflected mostly what’s going on in the larger world. Moore’s law has been going on for something like 100 years now. And compute continues to get faster all the time. Storage continues to cost less year over year in a variety of ways. But they have, on some level, tricked an entire generation of businesses into believing that network bandwidth is this precious, very finite thing, and of course, it’s going to be ridiculously expensive. You know, unless you’re taking it inbound, in which case, oh, by all means back the truck around. It’ll be great.

So, I’ve talked to founders—or prospective founders—who had ideas but were firmly convinced that there was no economical way to build it. Because oh, if I were to start doing real-time video stuff, well, great, let’s do the numbers on this. And hey, that’ll be $50,000 a minute, if I read the pricing page correctly, it’s like, well, you could get some discounts if you ask nicely, but it doesn’t occur to them that they could wind up asking for a 98% discount on these things. Everything is measured in a per gigabyte dimension and that just becomes one of those things where people are starting to think about and meter something that—from my days in data centers where you care about the size of the pipe and not what’s passing through it—to be the wrong way of thinking about things.

Matthew: A little of this is that everybody is colored by their experience of dealing with their ISP at home. And in the United States, in a lot of the world, ISPs are built on the old cable infrastructure. And if you think about the cable infrastructure, when it was originally laid down, it was all one-directional. So, you know, if you were turning on cable in your house in a pre-internet world, data fl—

Corey: Oh, you’d watch a show and your feedback was yelling at the TV, and that’s okay. They would drop those packets.

Matthew: And there was a tiny, tiny, tiny bit of data that would go back the other direction, but cable was one-directional. And so, it actually took an enormous amount of engineering to make cable bi-directional. And that’s the reason why if you’re using a traditional cable company as your ISP, typically you will have a large amount of download capacity, you’ll have, you know, a 100 megabits of down capacity, but you might only have a 10th of that—so maybe ten megabits—of upload capacity. That is an artifact of the cable system. That is not just the natural way that the internet works.

And the way that it is different, that wholesale bandwidth works, is that when you sign up for wholesale bandwidth—again, as you phrase it, you’re not buying this many bytes that flows over the line; you’re buying, effectively, a pipe. You know, the late Senator Ted Stevens said that the internet is just a series of tubes and got mocked mercilessly, but the internet is just a series of tubes. And when Cloudflare or AWS or Google or Microsoft buys one of those tubes, what they pay for is the diameter of the tube, the amount that can fit through it. And the nature of this is you don’t just get one tube, you get two. One that is down and one that is up. And they’re the same size.

And so, if you’ve got a terabit of traffic coming down and zero going up, that costs exactly the same as a terabit going up and zero going down, which costs exactly the same as a terabit going down and a terabit going up. It is different than your home, you know, cable internet connection. And that’s the thing that I think a lot of people don’t understand. And so, as you pointed out, but the great tragedy of the cloud is that for nothing other than business reasons, these hyperscale public cloud companies don’t charge you anything to accept data—even though that is actually the more expensive of the two operations for that because writes are more expensive than reads—but the inherent fact that they were able to suck the data in means that they have the capacity, at no additional cost, to be able to send that data back out. And so, I think that, you know, the good news is that you’re starting to see some providers—so Cloudflare, we’ve never charged for egress because, again, we think that over time, bandwidth prices go to zero because it just makes sense; it makes sense for ISPs, it makes sense for connectiv—to be connected to us.

And that’s something that we can do, but even in the cases of the cloud providers where maybe they’re all in one place and somebody has to pay to backhaul the traffic around the world, maybe there’s some cost, but you’re starting to see some pressure from some of the more forward-leaning providers. So Oracle, I think has done a good job of leaning in and showing how egress fees are just out of control. But it’s crazy that in some cases, you have a 4,000x markup on AWS bandwidth fees. And that’s assuming that they’re paying the same rates as what we would get at Cloudflare, you know, even though we are a much smaller company than they are, and they should be able to get even better prices.

Corey: Yes, if there’s one thing Amazon is known for, it as being bad at negotiating. Yeah, sure it is. I’m sure that they’re just a terrific joy to be a vendor to.

Matthew: Yeah, and I think that fundamentally what the price of bandwidth is, is tied very closely to what the cost of a port on a router costs. And what we’ve seen over the course of the last ten years is that cost has just gone enormously down where the capacity of that port has gone way up and the just physical cost, the depreciated cost that port has gone down. And yet, when you look at Amazon, you just haven’t seen a decrease in the cost of bandwidth that they’re passing on to customers. And so, again, I think that this is one of the places where you’re starting to see regulators pay attention, we’ve seen efforts in the EU to say whatever you charge to take data out is the same as what you should charge it to put data in. We’re seeing the FTC start to look at this, and we’re seeing customers that are saying that this is a purely anti-competitive action.

And, you know, I think what would be the best and healthiest thing for the cloud by far is if we made it easy to move between various cloud providers. Because right now the choice is, do I use AWS or Google or Microsoft, whereas what I think any company out there really wants to be able to do is they want to be able to say, “I want to use this feature at AWS because they’re really good at that and I want to use this other feature at Google because they’re really good at that, and I want to us this other feature at Microsoft, and I want to mix and match between those various things.” And I think that if you actually got cloud providers to start competing on features as opposed to competing on their overall platform, we’d actually have a much richer and more robust cloud environment, where you’d see a significantly improved amount of what’s going on, as opposed to what we have now, which is AWS being mediocre at everything.

Corey: I think that there’s also a story where for me, the egress is annoying, but so is the cross-region and so is the cross-AZ, which in many cases costs exactly the same. And that frustrates me from the perspective of, yes, if you have two data centers ten miles apart, there is some startup costs to you in running fiber between them, however you want to wind up with that working, but it’s a sunk cost. But at the end of that, though, when you wind up continuing to charge on a per gigabyte basis to customers on that, you’re making them decide on a very explicit trade-off of, do I care more about cost or do I care more about reliability? And it’s always going to be an investment decision between those two things, but when you make the reasonable approach of well, okay, an availability zone rarely goes down, and then it does, you get castigated by everyone for, “Oh it even says in their best practice documents to go ahead and build it this way.” It’s funny how a lot of the best practice documents wind up suggesting things that accrue primarily to a cloud provider’s benefit. But that’s the way of the world I suppose.

I just know, there’s a lot of customer frustration on it and in my client environments, it doesn’t seem to be very acute until we tear apart a bill and look at where they’re spending money, and on what, at which point, the dawning realization, you can watch it happen, where they suddenly realize exactly where their money is going—because it’s relatively impenetrable without that—and then they get angry. And I feel like if people don’t know what they’re being charged for, on some level, you’ve messed up.

Matthew: Yeah. So, there’s cost to running a network, but there’s no reason other than limiting competition why you would charge more to take data out than you would put data in. And that’s a puzzle. The cross-region thing, you know, I think where we’re seeing a lot of that is actually oftentimes, when you’ve got new technologies that come out and they need to take advantage of some scarce resource. And so, AI—and all the AI companies are a classic example of this—right now, if you’re trying to build a model, an AI model, you are hunting the world for available GPUs at a reasonable price because there’s an enormous scarcity of them.

And so, you need to move from AWS East to AWS West, to AWS, you know, Singapore, to AWS in Luxembourg and bounce around to find wherever there’s GPU availability. And then that is crossed against the fact that these training datasets are huge. You know, I mean, they’re just massive, massive, massive amounts of data. And so, what that is doing is you’re having these AI companies that are really seeing this get hit in the face, where they literally can’t get the capacity they need because of the fact that whatever cloud provider in whatever region they’ve selected to store their data isn’t able to have that capacity. And so, they’re getting hit not only by sort of a double whammy of, “I need to move my data to wherever there’s capacity. And if I don’t do that, then I have to pay some premium, an ever-escalating price for the underlying GPUs.” And God forbid, you have to move from AWS to Google to chase that.

And so, we’re seeing a lot of companies that are saying, “This doesn’t make any sense. We have this enormous training set. If we just put it with Cloudflare, this is data that makes sense to live in the network, fundamentally.” And not everything does. Like, we’re not the right place to store your long-term transaction logs that you’re only going to look at if you get sued. There are much better places, much more effective places do it.

But in those cases where you’ve got to read data frequently, you’ve got to read it from different places around the world, and you will need to decrease what those costs of each one of those reads are, what we’re seeing is just an enormous amount of demand for that. And I think these AI startups are really just a very clear example of what company after company after company needs, and why R2 has had—which is our zero egress cost S3 competitor—why that is just seeing such explosive growth from a broad set of customers.

Corey: Because I enjoy pushing the bounds of how ridiculous I can be on the internet, I wound up grabbing a copy of the model, the Llama 2 model that Meta just released earlier this week as we’re recording this. And it was great. It took a little while to download here. I have gigabit internet, so okay, it took some time. But then I wound up with something like 330 gigs of models. Great, awesome.

Except for the fact that I do the math on that and just for me as one person to download that, had they been paying the listed price on the AWS website, they would have spent a bit over $30, just for me as one random user to download the model, once. If you can express that into the idea of this is a model that is absolutely perfect for whatever use case, but we want to have it run with some great GPUs available at another cloud provider. Let’s move the model over there, ignoring the data it’s operating on as well, it becomes completely untenable. It really strikes me as an anti-competitiveness issue.

Matthew: Yeah. I think that’s it. That’s right. And that’s just the model. To build that model, you would have literally millions of times more data that was feeding it. And so, the training sets for that model would be many, many, many, many, many, many orders of magnitude larger in terms of what’s there. And so, I think the AI space is really illustrating where you have this scarce resource that you need to chase around the world, you have these enormous datasets, it’s illustrating how these egress fees are actually holding back the ability for innovation to happen.

And again, they are absolutely—there is no valid reason why you would charge more for egress than you do for ingress other than limiting competition. And I think the good news, again, is that’s something that’s gotten regulators’ attention, that’s something that’s gotten customers’ attention, and over time, I think we all benefit. And I think actually, AWS and Google and Microsoft actually become better if we start to have more competition on a feature-by-feature basis as opposed to on an overall platform. The choice shouldn’t be, “I use AWS.” And any big company, like, nobody is all-in only on one cloud provider. Everyone is multi-cloud, whether they want to be or not because people end up buying another company or some skunkworks team goes off and uses some other function.

So, you are across multiple different clouds, whether you want to be or not. But the ideal, and when I talk to customers, they want is, they want to say, “Well, you know that stuff that they’re doing over at Microsoft with AI, that sounds really interesting. I want to use that, but I really like the maturity and robustness of some of the EC2 API, so I want to use that at AWS. And Google is still, you know, the best in the world at doing search and indexing and everything, so I want to use that as well, in order to build my application.” And the applications of the future will inherently stitch together different features from different cloud providers, different startups.

And at Cloudflare, what we see is our, sort of, purpose for being is how do we make that stitching as easy as possible, as cost-effective as possible, and make it just make sense so that you have one consistent security layer? And again, we’re not about hording the data; we’re about connecting all of those things together. And again, you know, from the last time we talked to now, I’m actually much more optimistic that you’re going to see, kind of, this revolution where egress prices go down, you get competition on feature-by-features, and that’s just going to make every cloud provider better over the long-term.

Corey: This episode is sponsored in part by Panoptica. Panoptica simplifies container deployment, monitoring, and security, protecting the entire application stack from build to runtime. Scalable across clusters and multi-cloud environments, Panoptica secures containers, serverless APIs, and Kubernetes with a unified view, reducing operational complexity and promoting collaboration by integrating with commonly used developer, SRE, and SecOps tools. Panoptica ensures compliance with regulatory mandates and CIS benchmarks for best practice conformity. Privacy teams can monitor API traffic and identify sensitive data, while identifying open-source components vulnerable to attacks that require patching. Proactively addressing security issues with Panoptica allows businesses to focus on mitigating critical risks and protecting their interests. Learn more about Panoptica today at panoptica.app.

Corey: I don’t know that I would trust you folks to the long-term storage of critical data or the store of record on that. You don’t have the track record on that as a company the way that you do for being the network interchange that makes everything just work together. There are areas where I’m thrilled to explore and see how it works, but it takes time, at least from the sensible infrastructure perspective of trusting people with track records on these things. And you clearly have the network track record on these things to make this stick. It almost—it seems unfair to you folks, but I view you as Cloudflare is a CDN, that also dabbles in a few other things here in there, though, increasingly, it seems it’s CDN and security company are becoming synonymous.

Matthew: It’s interesting. I remember—and this really is going back to the origin story, but when we were starting Cloudflare, you know, what we saw was that, you know, we watched as software—starting with companies like Salesforce—transition from something that you bought in the box to something that you bought as a service [into 00:23:25] the cloud. We watched as, sort of, storage and compute transition from something that you bought from Dell or HP to something that you rented as a service. And so the fundamental problem that Cloudflare started out with was if the software and the storage and compute are going to move, inherently the security and the networking is going to move as well because it has to be as a service as well, there’s no way you can buy a you know, Cisco firewall and stick it in front of your cloud service. You have to be in the cloud as well.

So, we actually started very much as a security company. And the objection that everybody had to us as we would sort of go out and describe what we were planning on doing was, “You know, that sounds great, but you’re going to slow everything down.” And so, we became just obsessed with latency. And Michelle, my co-founder, and I were business students and we had an advisor, a guy named Tom [Eisenmann 00:24:26] in business school. And I remember going in and that was his objection as well and so we did all this work to figure it out.

And obviously, you know, I’d say computer science, and anytime that you have a problem around latency or speed caching is an obvious part of the solution to that. And so, we went in and we said, “Here’s how we’re going to do it: [unintelligible 00:24:47] all this protocol optimization stuff, and here’s how we’re going to distribute it around the world and get close to where users are. And we’re going to use caching in the places where we can do caching.” And Tom said, “Oh, you’re building a CDN.” And I remember looking at him and then I’m looking at Michelle. And Michelle is Canadian, and so I was like, “I don’t know that I’m building a Canadian, but I guess. I don’t know.”

And then, you know, we walked out in the hall and Michelle looked at me and she’s like, “We have to go figure out what the CDN thing is.” And we had no idea what a CDN was. And even when we learned about it, we were like, that business doesn’t make any sense. Like because again, the CDNs were the first ones to really charge for bandwidth. And so today, we have effectively built, you know, a giant CDN and are the fastest in the world and do all those things.

But we’ve always given it away basically for free because fundamentally, what we’re trying to do is all that other stuff. And so, we actually started with security. Security is—you know, my—I’ve been working in security now for over 25 years and that’s where my background comes from, and if you go back and look at what the original plan was, it was how do we provide that security as a service? And yeah, you need to have caching because caching makes sense. What I think is the difference is that in order to do that, in order to be able to build that, we had to build a set of developer tools for our own team to allow them to build things as quickly as possible.

And, you know, if you look at Cloudflare, I think one of the things we’re known for is just the rapid, rapid, rapid pace of innovation. And so, over time, customers would ask us, “How do you innovate so fast? How do you build things fast?” And part of the answer to that, there are lots of ways that we’ve been able to do that, but part of the answer to that is we built a developer platform for our own team, which was just incredibly flexible, allowed you to scale to almost any level, took care of a lot of that traditional SRE functions just behind the scenes without you having to think about it, and it allowed our team to be really fast. And our customers are like, “Wow, I want that too.”

And so, customer after customer after customer after customer was asking and saying, you know, “We have those same problems. You know, if we’re a big e-commerce player, we need to be able to build something that can scale up incredibly quickly, and we don’t have to think about spinning up VMs or containers or whatever, we don’t have to think about that. You know, our customers are around the world. We don’t want to have to pick a region for where we’re going to deploy code.” And so, where we built Cloudflare Workers for ourself first, customers really pushed us to make it available to them as well.

And that’s the way that almost any good developer platform starts out. That’s how AWS started. That’s how, you know, the Microsoft developer platform, and so the Apple developer platform, the Salesforce developer platform, they all start out as internal tools, and then someone says, “Can you expose this to us as well?” And that’s where, you know, I think that we have built this. And again, it’s very opinionated, it is right for certain applications, it’s never going to be the right place to run SAP HANA, but the company that builds the tool [crosstalk 00:27:58]—

Corey: I’m not convinced there is a right place to run SAP HANA, but that’s probably unfair of me.

Matthew: Yeah, but there is a startup out there, I guarantee you, that’s building whatever the replacement for SAP HANA is. And I think it’s a better than even bet that Cloudflare Workers is part of their stack because it solves a lot of those fundamental challenges. And that’s been great because it is now allowing customer after customer after customer, big and large startups and multinationals, to do things that you just can’t do with traditional legacy hyperscale public cloud. And so, I think we’re sort of the next generation of building that. And again, I don’t think we set out to build a developer platform for third parties, but we needed to build it for ourselves and that’s how we built such an effective tool that now so many companies are relying on.

Corey: As a Cloudflare customer myself, I think that one of the things that makes you folks standalone—it’s why I included security as well as CDN is one of the things I trust you folks with—has been—

Matthew: I still think CDN is Canadian. You will never see us use that term. It’s like, Gartner was like, “You have to submit something for the CDN-like ser—” and we ended up, like, being absolute top-right in it. But it’s a space that is inherently going to zero because again, if bandwidth is free, I’m not sure what—this is what the internet—how the internet should work. So yeah, anyway.

Corey: I agree wholeheartedly. But what I’ve always enjoyed, and this is probably going to make me sound meaner than I intend it to, it has been your outages. Because when computers inherently at some point break, which is what they do, you personally and you as a company have both taken a tone that I don’t want to say gleeful, but it’s sort of the next closest thing to it regarding the postmortem that winds up getting published, the explanation of what caused it, the transparency is unheard of at companies that are your scale, where usually they want to talk about these things as little as possible. Whereas you’ve turned these into things that are educational to those of us who don’t have the same scale to worry about but can take things from that are helpful. And that transparency just counts for so much when we’re talking about things as critical as security.

Matthew: I would definitely not describe it as gleeful. It is incredibly painful. And we, you know, we know we let customers down anytime we have an issue. But we tend not to make the same mistake twice. And the only way that we really can reliably do that is by being just as transparent as possible about exactly what happened.

And we hope that others can learn from the mistakes that we made. And so, we own the mistakes we made and we talk about them and we’re transparent, both internally but also externally when there’s a problem. And it’s really amazing to just see how much, you know, we’ve improved over time. So, it’s actually interesting that, you know, if you look across—and we measure, we test and measure all the big hyperscale public clouds, what their availability and reliability is and measure ourselves against it, and across the board, second half of 2021 and into the first half of 2022 was the worst for every cloud provider in terms of reliability. And the question is why?

And the answer is, Covid. I mean, the answer to most things over the last three years is in one way, directly or indirectly, Covid. But what happened over that period of time was that in April of 2020, internet traffic and traffic to our service and everyone who’s like us doubled over the course of a two-week period. And there are not many utilities that you can imagine that if their usage doubles, that you wouldn’t have a problem. Imagine the sewer system all of a sudden has twice as much sewage, or the electrical grid as twice as much demand, or the freeways have twice as many cars. Like, things break down.

And especially the European internet came incredibly close to just completely failing at that time. And we all saw where our bottlenecks were. And what’s interesting is actually the availability wasn’t so bad in 2020 because people were—they understood the absolute critical importance that while we’re in the middle of a pandemic, we had to make sure the internet worked. And so, we—there were a lot of sleepless nights, there’s a—and not just at with us, but with every provider that’s out there. We were all doing Herculean tasks in order to make sure that things came online.

By the time we got to the sort of the second half of 2021, what everybody did, Cloudflare included, was we looked at it, and we said, “Okay, here were where the bottlenecks were. Here were the problems. What can we do to rearchitect our systems to do that?” And one of the things that we saw was that we effectively treated large data centers as one big block, and if you had certain pieces of equipment that failed in a way, that you would take that entire data center down and then that could have cascading effects across traffic as it shifted around across our network. And so, we did the work to say, “Let’s take that one big data center and divide it effectively into multiple independent units, where you make sure that they’re all on different power suppliers, you make sure they’re all in different [crosstalk 00:32:52]”—

Corey: [crosstalk 00:32:51] harder than it sounds. When you have redundant things, very often, the thing that takes you down the most is the heartbeat that determines whether something next to it is up or not. It gets a false reading and suddenly, they’re basically trying to clobber each other to death. So, this is a lot harder than it sounds like.

Matthew: Yeah, and it was—but what’s interesting is, like, we took it all that into account, but the act of fixing things, you break things. And that was not just true at Cloudflare. If you look across Google and Microsoft and Amazon, everybody, their worst availability was second half of 2021 or into 2022. But it both internally and externally, we talked about the mistakes we made, we talked about the challenges we had, we talked about—and today, we’re significantly more resilient and more reliable because of that. And so, transparency is built into Cloudflare from the beginning.

The earliest story of this, I remember, there was a 15-year-old kid living in Long Beach, California who bought my social security number off of a Russian website that had hacked a bank that I’d once used to get a mortgage. He then use that to redirect my cell phone voicemail to a voicemail box he controlled. He then used that to get into my personal email. He then used that to find a zero-day vulnerability in Google’s corporate email where he could privilege-escalate from my personal email into Google’s corporate email, which is the provider that we use for our email service. And then he used that as an administrator on our email at the time—this is back in the early days of Cloudflare—to get into another administration account that he then used to redirect one of Cloud Source customers to a website that he controlled.

And thankfully, it wasn’t, you know, the FBI or the Central Bank of Brazil, which were all Cloudflare customers. Instead, it was 4chan because he was a 15-year-old hacker kid. And we fix it pretty quickly and nobody knew who Cloudflare was at the time. And so potential—

Corey: The potential damage that could have been caused at that point with that level of access to things, like, that is such a ridiculous way to use it.

Matthew: And—yeah [laugh]—my temptation—because it was embarrassing. He took a bunch of stuff from my personal email and he put it up on a website, which just to add insult to injury, was actually using Cloudflare as well. And I wanted to sweep it under the rug. And our team was like, “That’s not the right thing to do. We’re fundamentally a security company and we need to talk about when we make mistakes on security.” And so, we wrote a huge postmortem on, “Here’s all the stupid things that we did that caused this hack to happen.” And by the way, it wasn’t just us. It was AT&T, it was Google. I mean, there are a lot of people that ended up being involved.

Corey: It builds trust with that stuff. It’s painful in the short term, but I believe with the benefit of hindsight, it was clearly the right call.

Matthew: And it was—and I remember, you know, pushing ‘publish’ on the blog post and thinking, “This is going to be the end of the company.” And quite the opposite happened, which was all of a sudden, we saw just an incredible amount of people who signed up the next day saying, “If you’re going to be that transparent about something that was incredibly embarrassing when you didn’t have to be, then that’s the sort of thing that actually makes me trust that you’re going to be transparent the future.” And I think learning that lesson early on, has been just an incredibly valuable lesson for us and made us the company that we are today.

Corey: A question that I have for you about the idea of there being no reason to charge in one direction but not the other. There’s something that I’m not sure that I understand on this. If I run a website, to use your numbers of a terabit out—because it’s a web server—and effectively nothing in—because it’s a webserver; other than the request, nothing really is going to come in—that ingress bandwidth becomes effectively unused and also free. So, if I have another use case where I’m paying for it anyway, if I’m primarily caring about an outward direction, sure, you can send things in for free. Now, there’s a lot of nuance that goes into that. But I’m curious as to what the—is their fundamental misunderstanding in that analysis of the bandwidth market?

Matthew: No. And I think that’s exactly, exactly right. And it’s actually interesting. At Cloudflare, our infrastructure team—which is the one that manages our connections to the outside world, manages the hardware we have—meets on a quarterly basis with our product team. It’s called the Hot and Cold Meeting.

And what they do is they go over our infrastructure, and they say, “Okay, where are we hot? Where do we have not enough capacity?” If you think of any given server, an easy way to think of a server is that it has, sort of, four resources that are available to it. This is, kind of, vast simplification, but one is the connectivity to the outside world, both transit in and out. The second is the—

Corey: Otherwise it’s just a complicated space heater.

Matthew: Yeah [laugh]. The other is the CPU. The other is the longer-term storage. We use only SSDs, but sort of, you know, hard drives or SSD storage. And then the fourth is the short-term storage, or RAM that’s in that server.

And so, at any given moment, there are going to be places where we are running hot, where we have a sort of capacity level that we’re targeting and we’re over that capacity level, but we’re also going to be running cold in some of those areas. And so, the infrastructure team and the product team get together and the product team has requests on, you know, “Here’s some more places we would be great to have more infrastructure.” And we’re really good at deploying that when we need to, but the infrastructure team then also says, “Here are the places where we’re cold, where we have excess capacity.” And that turns into products at Cloudflare. So, for instance, you know, the reason that we got into the zero-trust space was very much because we had all this excess capacity.

We have 100 times the capacity of something like Zscaler across our network, and we can add that—that is primar—where most of our older products are all about outward traffic, the zero-trust products are all about inward traffic. And the reason that we can do everything that Zscaler does, but for, you know, a much, much, much more affordable prices, we going to basically just layer that on the network that already exists. The reason we don’t charge for the bandwidth behind DDoS attacks is DDoS attacks are always about inbound traffic and we have just a ton of excess capacity around that inbound traffic. And so, that unused capacity is a resource that we can then turn into products, and very much that conversation between our product team and our infrastructure team drives how we think about building new products. And we’re always trying to say how can we get as much utilization out of every single piece of equipment that we run everywhere in the world.

The way we build our network, we don’t have custom machines or different networks for every products. We build all of our machines—they come in generations. So, we’re on, I think, generation 14 of servers where we spec a server and it has, again, a certain amount of each of those four [bits 00:39:22] of capacity. But we can then deploy that server all around the world, and we’re buying many, many, many of them at any given time so we can get the best cost on that. But our product team is very much in constant communication with our infrastructure team and saying, “What more can we do with the capacity that we have?” And then we pass that on to our customers by adding additional features that work across our network and then doing it in a way that’s incredibly cost-effective.

Corey: I really want to thank you for taking the time to, basically once again, suffer slings and arrows about networking, security, cloud, economics, and so much more. If people want to learn more, where’s the best place for them to find you?

Matthew: You know, used to be an easy question to answer because it was just, you know, go on Twitter and find me but now we have all these new mediums. So, I’m @eastdakota on Twitter. I’m eastdakota.com on Bluesky. I’m @real_eastdakota on Threads. And so, you know, one way or another, if you search for eastdakota, you’ll come across me somewhere out there in the ether.

Corey: And we will, of course, put links to that in the show notes. Thank you so much for your time. I appreciate it.

Matthew: It’s great to talk to you, Corey.

Corey: Matthew Prince, CEO and co-founder of Cloudflare. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry, insulting comment that I will of course not charge you inbound data rates on.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

View Details

Richard Seroter, Director of Outbound Product Management at Google, joins Corey on Screaming in the Cloud to discuss what’s new at Google. Corey and Richard discuss how AI can move from a novelty to truly providing value, as well as the importance of people maintaining their skills and abilities rather than using AI as a black box solution. Richard also discusses how he views the DevRel function, and why he feels it’s so critical to communicate expectations for product launches with customers.

About Richard

Richard Seroter is Director of Outbound Product Management at Google Cloud. He’s also an instructor at Pluralsight, a frequent public speaker, and the author of multiple books on software design and development. Richard maintains a regularly updated blog (seroter.com) on topics of architecture and solution design and can be found on Twitter as @rseroter.

Links Referenced:

  • Google Cloud: https://cloud.google.com
  • Personal website: https://seroter.com
  • Twitter: https://twitter.com/rseroter
  • LinkedIn: https://www.linkedin.com/in/seroter/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Human-scale teams use Tailscale to build trusted networks. Tailscale Funnel is a great way to share a local service with your team for collaboration, testing, and experimentation. Funnel securely exposes your dev environment at a stable URL, complete with auto-provisioned TLS certificates. Use it from the command line or the new VS Code extensions. In a few keystrokes, you can securely expose a local port to the internet, right from the IDE.

I did this in a talk I gave at Tailscale Up, their first inaugural developer conference. I used it to present my slides and only revealed that that’s what I was doing at the end of it. It’s awesome, it works! Check it out!

Their free plan now includes 3 users & 100 devices. Try it at snark.cloud/tailscalescream

Corey: Welcome to Screaming in the Cloud, I’m Corey Quinn. We have returning guest Richard Seroter here who has apparently been collecting words to add to his job title over the years that we’ve been talking to him. Richard, you are now the Director of Product Management and Developer Relations at Google Cloud. Do I have all those words in the correct order and I haven’t forgotten any along the way?

Richard: I think that’s all right. I think my first job was at Anderson Consulting as an analyst, so my goal is to really just add more words to whatever these titles—

Corey: It’s an adjective collection, really. That’s what a career turns into. It’s really the length of a career and success is measured not by accomplishments but by word count on your resume.

Richard: If your business card requires a comma, success.

Corey: So, it’s been about a year or so since we last chatted here. What have you been up to?

Richard: Yeah, plenty of things here, still, at Google Cloud as we took on developer relations. And, but you know, Google Cloud proper, I think AI has—I don’t know if you’ve noticed, AI has kind of taken off with some folks who’s spending a lot the last year… juicing up services and getting things ready there. And you know, myself and the team kind of remaking DevRel for a 2023 sort of worldview. So, yeah we spent the last year just scaling and growing and in covering some new areas like AI, which has been fun.

Corey: You became profitable, which is awesome. I imagined at some point, someone wound up, like, basically realizing that you need to, like, patch the hole in the pipe and suddenly the water bill is no longer $8 billion a quarter. And hey, that works super well. Like, wow, that explains our utility bill and a few other things as well. I imagine the actual cause is slightly more complex than that, but I am a simple creature.

Richard: Yeah. I think we made more than YouTube last quarter, which was a good milestone when you think of—I don’t think anybody who says Google Cloud is a fun side project of Google is talking seriously anymore.

Corey: I misunderstood you at first. I thought you said that you’re pretty sure you made more than I did last year. It’s like, well, yes, if a multi-billion dollar company’s hyperscale cloud doesn’t make more than I personally do, then I have many questions. And if I make more than that, I have a bunch of different questions, all of which could be terrifying to someone.

Richard: You’re killing it. Yeah.

Corey: I’m working on it. So, over the last year, another trend that’s emerged has been a pivot away—thankfully—from all of the Web3 nonsense and instead embracing the sprinkle some AI on it. And I’m not—people are about to listen to this and think, wait a minute, is he subtweeting my company? No, I’m subtweeting everyone’s company because it seems to be a universal phenomenon. What’s your take on it?

Richard: I mean, it’s countercultural now to not start every conversation with let me tell you about our AI story. And hopefully, we’re going to get past this cycle. I think the AI stuff is here to stay. This does not feel like a hype trend to me overall. Like, this is legit tech with real user interest. I think that’s awesome.

I don’t think a year from now, we’re going to be competing over who has the biggest model anymore. Nobody cares. I don’t know if we’re going to hopefully lead with AI the same way as much as, what is it doing for me? What is my experience? Is it better? Can I do this job better? Did you eliminate this complex piece of toil from my day two stuff? That’s what we should be talking about. But right now it’s new and it’s interesting. So, we all have to rub some AI on it.

Corey: I think that there is also a bit of a passing of the buck going on when it comes to AI where I’ve talked to companies that are super excited about how they have this new AI story that’s going to be great. And, “Well, what does it do?” “It lets you query our interface to get an answer.” Okay, is this just cover for being bad UX?

Richard: [laugh]. That can be true in some cases. In other cases, this will fix UXes that will always be hard. Like, do we need to keep changing… I don’t know, I’m sure if you and I go to our favorite cloud providers and go through their documentation, it’s hard to have docs for 200 services and millions of pages. Maybe AI will fix some of that and make it easier to discover stuff.

So in some cases, UIs are just hard at scale. But yes, I think in some cases, this papers over other things not happening by just rubbing some AI on it. Hopefully, for most everybody else, it’s actually interesting, new value. But yeah, that’s a… every week it’s a new press release from somebody saying they’re about to launch some AI stuff. I don’t know how any normal human is keeping up with it.

Corey: I certainly don’t know. I’m curious to see what happens but it’s kind of wild, too, because there you’re right. There is something real there where you ask it to draw you a picture of a pony or something and it does, or give me a bunch of random analysis of this. I asked one recently to go ahead and rank the US presidents by absorbency and with a straight face, it did it, which is kind of amazing. I feel like there’s a lack of imagination in the way that people talk about these things and a certain lack of awareness that you can make this a lot of fun, and in some ways, make that a better showcase of the business value than trying to do the straight-laced thing of having it explain Microsoft Excel to you.

Richard: I think that’s fair. I don’t know how much sometimes whimsy and enterprise mix. Sometimes that can be a tricky part of the value prop. But I’m with you this some of this is hopefully returns to some more creativity of things. I mean, I personally use things like Bard or what have you that, “Hey, I’m trying to think of this idea. Can you give me some suggestions?” Or—I just did a couple weeks ago—“I need sample data for my app.”

I could spend the next ten minutes coming up with Seinfeld and Bob’s Burgers characters, or just give me the list in two seconds in JSON. Like that’s great. So, I’m hoping we get to use this for more fun stuff. I’ll be fascinated to see if when I write the keynote for—I’m working on the keynote for Next, if I can really inject something completely off the wall. I guess you’re challenging me and I respect that.

Corey: Oh, I absolutely am. And one of the things that I believe firmly is that we lose sight of the fact that people are inherently multifaceted. Just because you are a C-level executive at an enterprise does not mean that you’re not also a human being with a sense of creativity and a bit of whimsy as well. Everyone is going to compete to wind up boring you to death with PowerPoint. Find something that sparks the imagination and sparks joy.

Because yes, you’re going to find the boring business case on your own without too much in the way of prodding for that, but isn’t it great to imagine what if? What if we could have fun with some of these things? At least to me, that’s always been the goal is to get people’s attention. Humor has been my path, but there are others.

Richard: I’m with you. I think there’s a lot to that. And the question will be… yeah, I mean, again, to me, you and I talked about this before we started recording, this is the first trend for me in a while that feels purely organic where our customers, now—and I’ll tell our internal folks—our customers have much better ideas than we do. And it’s because they’re doing all kinds of wild things. They’re trying new scenarios, they’re building apps purely based on prompts, and they’re trying to, you know, do this.

And it’s better than what we just come up with, which is awesome. That’s how it should be, versus just some vendor-led hype initiative where it is just boring corporate stuff. So, I like the fact that this isn’t just us talking; it’s the whole industry talking. It’s people talking to my non-technical family members, giving me ideas for what they’re using this stuff for. I think that’s awesome. So yeah, but I’m with you, I think companies can also look for more creative angles than just what’s another way to left-align something in a cell.

Corey: I mean, some of the expressions on this are wild to me. The Photoshop beta with its generative AI play has just been phenomenal. Because it’s weird stuff, like, things that, yeah, I’m never going to be a great artist, let’s be clear, but being able to say remove this person from the background, and it does it, as best I can tell, seamlessly is stuff where yeah, that would have taken me ages to find someone who knows what the hell they’re doing on the internet somewhere and then pay them to do it. Or basically stumble my way through it for two hours and it somehow looks worse afterwards than before I started. It’s the baseline stuff of, I’m never going to be able to have it—to my understanding—go ahead just build me a whole banner ad that does this and hit these tones and the rest, but it is going to help me refine something in that direction, until I can then, you know, hand it to a professional who can take it from my chicken scratching into something real.

Richard: If it will. I think that’s my only concern personally with some of this is I don’t want this to erase expertise or us to think we can just get lazy. I think that I get nervous, like, can I just tell it to do stuff and I don’t even check the output, or I don’t do whatever. So, I think that’s when you go back to, again, enterprise use cases. If this is generating code or instructions or documentation or what have you, I need to trust that output in some way.

Or more importantly, I still need to retain the skills necessary to check it. So, I’m hoping people like you and me and all our —every—all the users out there of this stuff, don’t just offload responsibility to the machine. Like, just always treat it like a kind of slightly drunk friend sitting next to you with good advice and always check it out.

Corey: It’s critical. I think that there’s a lot of concern—and I’m not saying that people are wrong on this—but that people are now going to let it take over their jobs, it’s going to wind up destroying industries. No, I think it’s going to continue to automate things that previously required human intervention. But this has been true since the Industrial Revolution, where opportunities arise and old jobs that used to be critical are no longer centered in quite the same way. The one aspect that does concern me is not that kids are going to be used to cheat on essays like, okay, great, whatever. That seems to be floated mostly by academics who are concerned about the appropriate structure of academia.

For me, the problem is, is there’s a reason that we have people go through 12 years of English class in the United States and that is, it’s not to dissect of the work of long-dead authors. It’s to understand how to write and how to tell us a story and how to frame ideas cohesively. And, “The computer will do that for me,” I feel like that potentially might not serve people particularly well. But as a counterpoint, I was told when I was going to school my entire life that you’re never going to have a calculator in your pocket all the time that you need one. No, but I can also speak now to the open air, ask it any math problem I can imagine, and get a correct answer spoken back to me. That also wasn’t really in the bingo card that I had back then either, so I am a hesitant to try and predict the future.

Richard: Yeah, that’s fair. I think it’s still important for a kid that I know how to make change or do certain things. I don’t want to just offload to calculators or—I want to be able to understand, as you say, literature or things, not just ever print me out a book report. But that happens with us professionals, too, right? Like, I don’t want to just atrophy all of my programming skills because all I’m doing is accepting suggestions from the machine, or that it’s writing my emails for me. Like, that still weirds me out a little bit. I like to write an email or send a tweet or do a summary. To me, I enjoy those things still. I don’t want to—that’s not toil to me. So, I’m hoping that we just use this to make ourselves better and we don’t just use it to make ourselves lazier.

Corey: You mentioned a few minutes ago that you are currently working on writing your keynote for Next, so I’m going to pretend, through a vicious character attack here, that this is—you know, it’s 11 o’clock at night, the day before the Next keynote and you found new and exciting ways to procrastinate, like recording a podcast episode with me. My question for you is, how is this Next going to be different than previous Nexts?

Richard: Hmm. Yeah, I mean, for the first time in a while it’s in person, which is wonderful. So, we’ll have a bunch of folks at Moscone in San Francisco, which is tremendous. And I [unintelligible 00:11:56] it, too, I definitely have online events fatigue. So—because absolutely no one has ever just watched the screen entirely for a 15 or 30 or 60-minute keynote. We’re all tabbing over to something else and multitasking. And at least when I’m in the room, I can at least pretend I’ll be paying attention the whole time. The medium is different. So, first off, I’m just excited—

Corey: Right. It feels a lot ruder to get up and walk out of the front row in the middle of someone’s talk. Now, don’t get me wrong, I’ll still do it because I’m a jerk, but I’ll feel bad about it as I do. I kid, I kid. But yeah, a tab away is always a thing. And we seem to have taken the same structure that works in those events and tried to force it into more or less a non-interactive Zoom call, and I feel like that is just very hard to distinguish.

I will say that Google did a phenomenal job of online events, given the constraints it was operating under. Production value is great, the fact that you took advantage of being in different facilities was awesome. But yeah, it’ll be good to be back in person again. I will be there with bells on in Moscone myself, mostly yelling at people, but you know, that’s what I do.

Richard: It’s what you do. But we missed that hallway track. You missed this sort of bump into people. Do hands-on labs, purposely have nothing to do where you just walk around the show floor. Like we have been missing, I think, society-wise, a little bit of just that intentional boredom. And so, sometimes you need at conference events, too, where you’re like, “I’m going to skip that next talk and just see what’s going on around here.” That’s awesome. You should do that more often.

So, we’re going to have a lot of spaces for just, like, go—like, 6000 square feet of even just going and looking at demos or doing hands-on stuff or talking with other people. Like that’s just the fun, awesome part. And yeah, you’re going to hear a lot about AI, but plenty about other stuff, too. Tons of announcements. But the key is that to me, community stuff, learn from each other stuff, that energy in person, you can’t replicate that online.

Corey: So, an area that you have expanded into has been DevRel, where you’ve always been involved with it, let’s be clear, but it’s becoming a bit more pronounced. And as an outsider, I look at Google Cloud’s DevRel presence and I don’t see as much of it as your staffing levels would indicate, to the naive approach. And let’s be clear, that means from my perspective, all public-facing humorous, probably performative content in different ways, where you have zany music videos that, you know, maybe, I don’t know, parody popular songs do celebrate some exec’s birthday they didn’t know was coming—[fake coughing]. Or creative nonsense on social media. And the the lack of seeing a lot of that could in part be explained by the fact that social media is wildly fracturing into a bunch of different islands which, on balance, is probably a good thing for the internet, but I also suspect it comes down to a common misunderstanding of what DevRel actually is.

It turns out that, contrary to what many people wanted to believe in the before times, it is not getting paid as much as an engineer, spending three times that amount of money on travel expenses every year to travel to exotic places, get on stage, party with your friends, and then give a 45-minute talk that spends two minutes mentioning where you work and 45 minutes talking about, I don’t know, how to pick the right standing desk. That has, in many cases, been the perception of DevRel and I don’t think that’s particularly defensible in our current macroeconomic climate. So, what are all those DevRel people doing?

Richard: [laugh]. That’s such a good loaded question.

Corey: It’s always good to be given a question where the answers are very clear there are right answers and wrong answers, and oh, wow. It’s a fun minefield. Have fun. Go catch.

Richard: Yeah. No, that’s terrific. Yeah, and your first part, we do have a pretty well-distributed team globally, who does a lot of things. Our YouTube channel has, you know, we just crossed a million subscribers who are getting this stuff regularly. It’s more than Amazon and Azure combined on YouTube. So, in terms of like that, audience—

Corey: Counterpoint, you definitionally are YouTube. But that’s neither here nor there, either. I don’t believe you’re juicing the stats, but it’s also somehow… not as awesome if, say, I were to do it, which I’m working on it, but I have a face for radio and it shows.

Richard: [laugh]. Yeah, but a lot of this has been… the quality and quantity. Like, you look at the quantity of video, it overwhelms everyone else because we spend a lot of time, we have a specific media team within my DevRel team that does the studio work, that does the production, that does all that stuff. And it’s a concerted effort. That team’s amazing. They do really awesome work.

But, you know, a lot of DevRel as you say, [sigh] I don’t know about you, I don’t think I’ve ever truly believed in the sort of halo effect of if super smart person works at X company, even if they don’t even talk about that company, that somehow presents good vibes and business benefits to that company. I don’t think we’ve ever proven that’s really true. Maybe you’ve seen counterpoints, where [crosstalk 00:16:34]—

Corey: I can think of anecdata examples of it. Often though, on some level, for me at least, it’s been okay someone I tremendously respect to the industry has gone to work at a company that I’ve never heard of. I will be paying attention to what that company does as a direct result. Conversely, when someone who is super well known, and has been working at a company for a while leaves and then either trashes the company on the way out or doesn’t talk about it, it’s a question of, what’s going on? Did something horrible happen there? Should we no longer like that company? Are we not friends anymore? It’s—and I don’t know if that’s necessarily constructive, either, but it also, on some level, feels like it can shorthand to oh, to be working DevRel, you have to be an influencer, which frankly, I find terrifying.

Richard: Yeah. Yeah. I just—the modern DevRel, hopefully, is doing a little more of product-led growth style work. They’re focusing specifically on how are we helping developers discover, engage, scale, become advocates themselves in the platform, increasing that flywheel through usage, but that has very discreet metrics, it has very specific ownership. Again, personally, I don’t even think DevRel should do as much with sales teams because sales teams have hundreds and sometimes thousands of sales engineers and sales reps. It’s amazing. They have exactly what they need.

I don’t think DevRel is a drop in the bucket to that team. I’d rather talk directly to developers, focus on people who are self-service signups, people who are developers in those big accounts. So, I think the modern DevRel team is doing more in that respect. But when I look at—I just look, Corey, this morning at what my team did last week—so the average DevRel team, I look at what advocacy does, teams writing code labs, they’re building tutorials. Yes, they’re doing some in person events. They wrote some blog posts, published some videos, shipped a couple open-source projects that they contribute to in, like gaming sector, we ship—we have a couple projects there.

They’re actually usually customer zero in the product. They use the product before it ships, provides bugs and feedback to the team, we run DORA workshops—because again, we’re the DevOps Research and Assessment gang—we actually run the tutorial and Docs platform for Google Cloud. We have people who write code samples and reference apps. So, sometimes you see things publicly, but you don’t see the 20,000 code samples in the docs, many written by our team. So, a lot of the times, DevRel is doing work to just enable on some of these different properties, whether that’s blogs or docs, whether that’s guest articles or event series, but all of this should be in service of having that credible relationship to help devs use the platform easier. And I love watching this team do that.

But I think there’s more to it now than years ago, where maybe it was just, let’s do some amazing work and try to have some second, third-order effect. I think DevRel teams that can have very discrete metrics around leading indicators of long-term cloud consumption. And if you can’t measure that successfully, you’ve probably got to rethink the team.

[midroll 00:19:20]

Corey: That’s probably fair. I think that there’s a tremendous series of… I want to call it thankless work. Like having done some of those ridiculous parody videos myself, people look at it and they chuckle and they wind up, that was clever and funny, and they move on to the next one. And they don’t see the fact that, you know, behind the scenes for that three-minute video, there was a five-figure budget to pull all that together with a lot of people doing a bunch of disparate work. Done right, a lot of this stuff looks like it was easy or that there was no work at all.

I mean, at some level, I’m as guilty of that as anyone. We’re recording a podcast now that is going to be handed over to the folks at HumblePod. They are going to produce this into something that sounds coherent, they’re going to fix audio issues, all kinds of other stuff across the board, a full transcript, and the rest. And all of that is invisible to me. It’s like AI; it’s the magic box I drop a file into and get podcast out the other side.

And that does a disservice to those people who are actively working in that space to make things better. Because the good stuff that they do never gets attention, but then the company makes an interesting blunder in some way or another and suddenly, everyone’s out there screaming and wondering why these people aren’t responding on Twitter in 20 seconds when they’re finding out about this stuff for the first time.

Richard: Mm-hm. Yeah, that’s fair. You know, different internal, external expectations of even DevRel. We’ve recently launched—I don’t know if you caught it—something called Jump Start Solutions, which were executable reference architectures. You can come into the Google Cloud Console or hit one of our pages and go, “Hey, I want to do a multi-tier web app.” “Hey, I want to do a data processing pipeline.” Like, use cases.

One click, we blow out the entire thing in the platform, use it, mess around with it, turn it off with one click. Most of those are built by DevRel. Like, my engineers have gone and built that. Tons of work behind the scenes. Really, like, production-grade quality type architectures, really, really great work. There’s going to be—there’s a dozen of these. We’ll GA them at Next—but really, really cool work. That’s DevRel. Now, that’s behind-the-scenes work, but as engineering work.

That can be some of the thankless work of setting up projects, deployment architectures, Terraform, all of them also dropped into GitHub, ton of work documenting those. But yeah, that looks like behind-the-scenes work. But that’s what—I mean, most of DevRel is engineers. These are folks often just building the things that then devs can use to learn the platforms. Is it the flashy work? No. Is it the most important work? Probably.

Corey: I do have a question I’d be remiss not to ask. Since the last time we spoke, relatively recently from this recording, Google—well, I’d say ‘Google announced,’ but they kind of didn’t—Squarespace announced that they’d be taking over Google domains. And there was a lot of silence, which I interpret, to be clear, as people at Google being caught by surprise, by large companies, communication is challenging. And that’s fine, but I don’t think it was anything necessarily nefarious.

And then it came out further in time with an FAQ that Google published on their site, that Google Cloud domains was a part of this as well. And that took a lot of people aback, in the sense—not that it’s hard to migrate a domain from one provider to another, but it brought up the old question of, if you’re building something in cloud, how do you pick what to trust? And I want to be clear before you answer that, I know you work there. I know that there are constraints on what you can or cannot say.

And for people who are wondering why I’m not hitting you harder on this, I want to be very explicit, I can ask you a whole bunch of questions that I already know the answer to, and that answer is that you can’t comment. That’s not constructive or creative. So, I don’t want people to think that I’m not intentionally asking the hard questions, but I also know that I’m not going to get an answer and all I’ll do is make you uncomfortable. But I think it’s fair to ask, how do you evaluate what services or providers or other resources you’re using when you’re building in cloud that are going to be around, that you can trust building on top of?

Richard: It’s a fair question. Not everyone’s on… let’s update our software on a weekly basis and I can just swap things in left. You know, there’s a reason that even Red Hat is so popular with Linux because as a government employee, I can use that Linux and know it’s backwards compatible for 15 years. And they sell that. Like, that’s the value, that this thing works forever.

And Microsoft does the same with a lot of their server products. Like, you know, for better or for worse, [laugh] they will always kind of work with a component you wrote 15 years ago in SharePoint and somehow it runs today. I don’t even know how that’s possible. Love it. That’s impressive.

Now, there’s a cost to that. There’s a giant tax in the vendor space to make that work. But yeah, there’s certain times where even with us, look, we are trying to get better and better at things like comms. And last year we announced—I checked them recently—you know, we have 185 Cloud products in our enterprise APIs. Meaning they have a very, very tight way we would deprecate with very, very long notice, they’ve got certain expectations on guarantees of how long you can use them, quality of service, all the SLAs.

And so, for me, like, I would bank on, first off, for every cloud provider, whether they’re anchor services. Build on those right? You know, S3 is not going anywhere from Amazon. Rock solid service. BigQuery Goodness gracious, it’s the center of Google Cloud.

And you look at a lot of services: what can you bet on that are the anchors? And then you can take bets on things that sit around it. There’s times to be edgy and say, “Hey, I’ll use Service Weaver,” which we open-sourced earlier this year. It’s kind of a cool framework for building apps and we’ll deconstruct it into microservices at deploy time. That’s cool.

Would I literally build my whole business on it? No, I don’t think so. It’s early stuff. Now, would I maybe use it also with some really boring VMs and boring API Gateway and boring storage? Totally. Those are going to be around forever.

I think for me, personally, I try to think of how do I isolate things that have some variability to them. Now, to your point, sometimes you don’t know there’s variability. You would have just thought that service might be around forever. So, how are you supposed to know that that thing could go away at some point? And that’s totally fair. I get that.

Which is why we have to keep being better at comms, making sure more things are in our enterprise APIs, which is almost everything. So, you have some assurances, when I build this thing, I’ve got a multi-year runway if anything ever changes. Nothing’s going to stay the same forever, but nothing should change tomorrow on a dime. We need more trust than that.

Corey: Absolutely. And I agree. And the problem, too, is hidden dependencies. Let’s say what is something very simple. I want to log in to [unintelligible 00:25:34] brand new AWS account and spin of a single EC2 instance. The end. Well, I can trust that EC2 is going to be there. Great. That’s not one service you need to go through that critical path. It is a bare minimum six, possibly as many as twelve, depending upon what it is exactly you’re doing.

And it’s the, you find out after the fact that oh, there was that hidden dependency in there that I wasn’t fully aware of. That is a tricky and delicate balance to strike. And, again, no one is going to ever congratulate you—at all—on the decision to maintain a service that is internally painful and engineering-ly expensive to keep going, but as soon as you kill something, even it’s for this thing doesn’t have any customers, the narrative becomes, “They’re screwing over their customers.” It’s—they just said that it didn’t have any. What’s the concern here?

It’s a messaging problem; it is a reputation problem. Conversely, everyone knows that Amazon does not kill AWS services. Full stop. Yeah, that turns out everyone’s wrong. By my count, they’ve killed ten, full-on AWS services and counting at the moment. But that is not the reputation that they have.

Conversely, I think that the reputation that Google is going to kill everything that it touches is probably not accurate, though I don’t know that I’d want to have them over to babysit either. So, I don’t know. But it is something that it feels like you’re swimming uphill on in many respects, just due to not even deprecation decisions, historically, so much as poor communication around them.

Richard: Mm-hm. I mean, communication can always get better, you know. And that’s, it’s not our customers’ problem to make sure that they can track every weird thing we feel like doing. It’s not their challenge. If our business model changes or our strategy changes, that’s not technically the customer’s problem. So, it’s always our job to make this as easy as possible. Anytime we don’t, we have made a mistake.

So, you know, even DevRel, hey, look, it puts teams in a tough spot. We want our customers to trust us. We have to earn that; you will never just give it to us. At the same time, as you say, “Hey, we’re profitable. It’s great. We’re growing like weeds,” it’s amazing to see how many people are using this platform. I mean, even services, you don’t talk about having—I mean, doing really, really well. But I got to earn that. And you got to earn, more importantly, the scale. I don’t want you to just kick the tires on Google Cloud; I want you to bet on it. But we’re only going to earn that with really good support, really good price, stability, really good feeling like these services are rock solid. Have we totally earned that? We’re getting there, but not as mature as we’d like to get yet, but I like where we’re going.

Corey: I agree. And reputations are tricky. I mean, recently InfluxDB deprecated two regions and wound up turning them off and deleting data. And they wound up getting massive blowback for this, which, to their credit, their co-founder and CTO, Paul Dix—who has been on the show before—wound up talking about and saying, “Yeah, that was us. We’re taking ownership of this.”

But the public announcement said that they had—that data in AWS was not recoverable and they’re reaching out to see if the data in GCP was still available. At which point, I took the wrong impression from this. Like, whoa, whoa, whoa. Hang on. Hold the phone here. Does that mean that data that I delete from a Google Cloud account isn’t really deleted?

Because I have a whole bunch of regulators that would like a word if so. And Paul jumped onto that with, “No, no, no, no, no. I want to be clear, we have a backup system internally that we were using that has that set up. And we deleted the backups on the AWS side; we don’t believe we did on the Google Cloud side. It’s purely us, not a cloud provider problem.” It’s like, “Okay, first, sorry for causing a fire drill.” Secondly, “Okay, that’s great.” But the reason I jumped in that direction was just because it becomes so easy when a narrative gets out there to believe the worst about companies that you don’t even realize you’re doing it.

Richard: No, I understand. It’s reflexive. And I get it. And look, B2B is not B2C, you know? In B2B, it’s not, “Build it and they will come.” I think we have the best cloud infrastructure, the best security posture, and the most sophisticated managed services. I believe that I use all the clouds. I think that’s true. But it doesn’t matter unless you also do the things around it, around support, security, you know, usability, trust, you have to go sell these things and bring them to people. You can’t just sit back and say, “It’s amazing. Everyone’s going to use it.” You’ve got to earn that. And so, that’s something that we’re still on the journey of, but our foundation is terrific. We just got to do a better job on some of these intangibles around it.

Corey: I agree with you, when you s—I think there’s a spirited debate you could have on any of those things you said that you believe that Google Cloud is the best at, with the exception of security, where I think that is unquestionably. I think that is a lot less variable than the others. The others are more or less, “Who has the best cloud infrastructure?” Well, depends on who had what for breakfast today. But the simplicity and the approach you take to security is head and shoulders above the competition.

And I want to make sure I give credit where due: it is because of that simplicity and default posturing that customers wind up better for it as a result. Otherwise, you wind up in this hell of, “You must have at least this much security training to responsibly secure your environment.” And that is never going to happen. People read far less than we wish they would. I want to make very clear that Google deserves the credit for that security posture.

Richard: Yeah, and the other thing, look, I’ll say that, from my observation, where we do something that feels a little special and different is we do think in platforms, we think in both how we build and how we operate and how the console is built by a platform team, you—singularly. How—[is 00:30:51] we’re doing Duet AI that we’ve pre-announced at I/O and are shipping. That is a full platform experience covering a dozen services. That is really hard to do if you have a lot of isolation. So, we’ve done a really cool job thinking in platforms and giving that simplicity at that platform level. Hard to do, but again, we have to bring people to it. You’re not going to discover it by accident.

Corey: Richard, I will let you get back to your tear-filled late-night writing of tomorrow’s Next keynote, but if people want to learn more—once the dust settles—where’s the best place for them to find you?

Richard: Yeah, hopefully, they continue to hang out at cloud.google.com and using all the free stuff, which is great. You can always find me at seroter.com. I read a bunch every day and then I’ve read a blog post every day about what I read, so if you ever want to tune in on that, just see what wacky things I’m checking out in tech, that is good. And I still hang out on different social networks, Twitter at @rseroter and LinkedIn and things like that. But yeah, join in and yell at me about anything I said.

Corey: I did not realize you had a daily reading list of what you put up there. That is news to me and I will definitely track in, and then of course, yell at you from the cheap seats when I disagree with anything that you’ve chosen to include. Thank you so much for taking the time to speak with me and suffer the uncomfortable questions.

Richard: Hey, I love it. If people aren’t talking about us, then we don’t matter, so I would much rather we’d be yelling about us than the opposite there.

Corey: [laugh]. As always, it’s been a pleasure. Richard Seroter, Director of Product Management and Developer Relations at Google Cloud. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an angry comment that you had an AI system write for you because you never learned how to structure a sentence.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

View Details

Anna Belak, Director of The Office of Cybersecurity Strategy at Sysdig, joins Corey on Screaming in the Cloud to discuss the findings in this year’s newly-released Sysdig Global Cloud Threat Report. Anna explains the challenges that teams face in ensuring their cloud is truly secure, including quantity of data versus quality, automation, and more. Corey and Anna also discuss how much faster attacks are able to occur, and Anna gives practical insights into what can be done to make your cloud environment more secure.

About Anna

Anna has nearly ten years of experience researching and advising organizations on cloud adoption with a focus on security best practices. As a Gartner Analyst, Anna spent six years helping more than 500 enterprises with vulnerability management, security monitoring, and DevSecOps initiatives. Anna's research and talks have been used to transform organizations' IT strategies and her research agenda helped to shape markets. Anna is the Director of The Office of Cybersecurity Strategy at Sysdig, using her deep understanding of the security industry to help IT professionals succeed in their cloud-native journey.

Anna holds a PhD in Materials Engineering from the University of Michigan, where she developed computational methods to study solar cells and rechargeable batteries.

Links Referenced:

  • Sysdig: https://sysdig.com/
  • Sysdig Global Cloud Threat Report: https://www.sysdig.com/2023threatreport
  • duckbillgroup.com: https://duckbillgroup.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. This promoted guest episode is brought to us by our friends over at Sysdig. And once again, I am pleased to welcome Anna Belak, whose title has changed since last we spoke to Director of the Office of Cybersecurity Strategy at Sysdig. Anna, welcome back, and congratulations on all the adjectives.

Anna: [laugh]. Thank you so much. It’s always a pleasure to hang out with you.

Corey: So, we are here today to talk about a thing that has been written. And we’re in that weird time thing where while we’re discussing it at the moment, it’s not yet public but will be when this releases. The Sysdig Global Cloud Threat Report, which I am a fan of. I like quite a bit the things it talks about and the ways it gets me thinking. There are things that I wind up agreeing with, there are things I wind up disagreeing with, and honestly, that makes it an awful lot of fun.

But let’s start with the whole, I guess, executive summary version of this. What is a Global Cloud Threat Report? Because to me, it seems like there’s an argument to be made for just putting all three of the big hyperscale clouds on it and calling it a day because they’re all threats to somebody.

Anna: To be fair, we didn’t think of the cloud providers themselves as the threats, but that’s a hot take.

Corey: Well, an even hotter one is what I’ve seen out of Azure lately with their complete lack of security issues, and the attackers somehow got a Microsoft signing key and the rest. I mean, at this point, I feel like Charlie Bell was brought in from Amazon to head cybersecurity and spent the last two years trapped in the executive washroom or something. But I can’t prove it, of course. No, you target the idea of threats in a different direction, towards what people more commonly think of as threats.

Anna: Yeah, the bad guys [laugh]. I mean, I would say that this is the reason you need a third-party security solution, buy my thing, blah, blah, blah, but [laugh], you know? Yeah, so we are—we have a threat research team like I think most self-respecting security vendors these days do. Ours, of course, is the best of them all, and they do all kinds of proactive and reactive research of what the bad guys are up to so that we can help our customers detect the bad guys, should they become their victims.

Corey: So, there was a previous version of this report, and then you’ve, in long-standing tradition, decided to go ahead and update it. Unlike many of the terrible professors I’ve had in years past, it’s not just slap a new version number, change the answers to some things, and force all the students to buy a new copy of the book every year because that’s your retirement plan, you actually have updated data. What are the big changes you’ve seen since the previous incarnation of this?

Anna: That is true. In fact, we start from scratch, more or less, every year, so all the data in this report is brand new. Obviously, it builds on our prior research. I’ll say one clearly connected piece of data is, last year, we did a supply chain story that talked about the bad stuff you can find in Docker Hub. This time we upleveled that and we actually looked deeper into the nature of said bad stuff and how one might identify that an image is bad.

And we found that 10% of the malware scary things inside images actually can’t be detected by most of your static tools. So, if you’re thinking, like, static analysis of any kind, SCA, vulnerability scanning, just, like, looking at the artifact itself before it’s deployed, you actually wouldn’t know it was bad. So, that’s a pretty cool change, I would say [laugh].

Corey: It is. And I’ll also say what’s going to probably sound like a throwaway joke, but I assure you it’s not, where you’re right, there is a lot of bad stuff on Docker Hub and part of the challenge is disambiguating malicious-bad and shitty-bad. But there are serious security concerns to code that is not intended to be awful, but it is anyway, and as a result, it leads to something that this report gets into a fair bit, which is the ideas of, effectively, lateralling from one vulnerability to another vulnerability to another vulnerability to the actual story. I mean, Capital One was a great example of this. They didn’t do anything that was outright negligent like leaving an S3 bucket open; it was a determined sophisticated attacker who went from one mistake to one mistake to one mistake to, boom, keys to the kingdom. And that at least is a little bit more understandable even if it’s not great when it’s your bank.

Anna: Yeah. I will point out that in the 10% that these things are really bad department, it was 10% of all things that were actually really bad. So, there were many things that were just shitty, but we had pared it down to the things that were definitely malicious, and then 10% of those things you could only identify if you had some sort of runtime analysis. Now, runtime analysis can be a lot of different things. It’s just that if you’re relying on preventive controls, you might have a bad time, like, one times out of ten, at least.

But to your point about, kind of, chaining things together, I think that’s actually the key, right? Like, that’s the most interesting moment is, like, which things can they grab onto, and then where can they pivot? Because it’s not like you barge in, open the door, like, you’ve won. Like, there’s multiple steps to this process that are sometimes actually quite nuanced. And I’ll call out that, like, one of the other findings we got this year that was pretty cool is that the time it takes to get through those steps is very short. There’s a data point from Mandiant that says that the average dwell time for an attacker is 16 days. So like, two weeks, maybe. And in our data, the average dwell time for the attacks we saw was more like ten minutes.

Corey: And that is going to be notable for folks. Like, there are times where I have—in years past; not recently, mind you—I have—oh, I’m trying to set something up, but I’m just going to open this port to the internet so I can access it from where I am right now and I’ll go back and shut it in a couple hours. There was a time that that was generally okay. These days, everything happens so rapidly. I mean, I’ve sat there with a stopwatch after intentionally committing AWS credentials to Gif-ub—yes, that’s how it’s pronounced—and 22 seconds until the first probing attempt started hitting, which was basically impressively fast. Like, the last thing in the entire sequence was, and then I got an alert from Amazon that something might have been up, at which point it is too late. But it’s a hard problem and I get it. People don’t really appreciate just how quickly some of these things can evolve.

Anna: Yeah. And I think the main reason, from at least what we see, is that the bad guys are into the cloud saying, right, like, we good guys love the automation, we love the programmability, we love the immutable infrastructure, like, all this stuff is awesome and it’s enabling us to deliver cool products faster to our customers and make more money, but the bad guys are using all the same benefits to perpetrate their evil crimes. So, they’re building automation, they’re stringing cool things together. Like, they have scripts that they run that basically just scan whatever’s out there to see what new things have shown up, and they also have scripts for reconnaissance that will just send a message back to them through Telegram or WhatsApp, letting them know like, “Hey, I’ve been running, you know, for however long and I see a cool thing you may be able to use.” Then the human being shows up and they’re like, “All right. Let’s see what I can do with this credential,” or with this misconfiguration or what have you. So, a lot of their initial, kind of, discovery into what they can get at is heavily automated, which is why it’s so fast.

Corey: I feel like, on some level, this is an unpleasant sharp shock for an awful lot of executives because, “Wait, what do you mean attackers can move that quickly? Our crap-ass engineering teams can’t get anything released in less than three sprints. What gives?” And I don’t think people have a real conception of just how fast bad actors are capable of moving.

Anna: I think we said—actually [unintelligible 00:07:57] last year, but this is a business for them, right? They’re trying to make money. And it’s a little bleak to think about it, but these guys have a day job and this is it. Like, our guys have a day job, that’s shipping code, and then they’re supposed to also do security. The bad guys just have a day job of breaking your code and stealing your stuff.

Corey: And on some level, it feels like you have a choice to make in which side you go at. And it’s, like, which one of those do I spend more time in meetings with? And maybe that’s not the most legitimate way to pick a job; ethics do come into play. But yeah, there’s it takes a certain similar mindset, on some level, to be able to understand just how the security landscape looks from an attacker's point of view.

Anna: I’ll bet the bad guys have meetings too, actually.

Corey: You know, you’re probably right. Can you imagine the actual corporate life of a criminal syndicate? That’s a sitcom in there that just needs to happen. But again, I’m sorry, I shouldn’t talk about that. We’re on a writer’s strike this week, so there’s that.

One thing that came out of the report that makes perfect sense—and I’ve heard about it, but I haven’t seen it myself and I wanted to dive into on this—specifically that automation has been weaponized in the cloud. Now, it’s easy to misinterpret that the first time you read it—like I did—as, “Oh, you mean the bad guys have discovered the magic of shell scripts? No kidding.” It’s more than that. You have reports of people using things like CloudFormation to stand up resources that are then used to attack the rest of the infrastructure.

And it’s, yeah, it makes perfect sense. Like, back in the data center days, it was a very determined attacker that went through the process of getting an evil server stuffed into a rack somewhere. But it’s an API call away in cloud. I’m surprised we haven’t seen this before.

Anna: Yeah. We probably have; I don’t know if we’ve documented before. And sometimes it’s hard to know that that’s what’s happening, right? I will say that both of those things are true, right? Like the shell scripts are definitely there, and to your point about how long it takes, you know, to stopwatch, these things, on the short end of our dwell time data set, it’s zero seconds. It’s zero seconds from, like, A to B because it’s just a script.

And that’s not surprising. But the comment about CloudFormation specifically, right, is we’re talking about people, kind of, figuring out how to create policy in the cloud to prevent bad stuff from happening because they’re reading all the best practices ebooks and whatever, watching the YouTube videos. And so, you understand that you can, say, write policy to prevent users from doing certain things, but sometimes we forget that, like, if you don’t want a user to be able to attach user policy to something. If you didn’t write the rule that says you also can’t do that in CloudFormation, then suddenly, you can’t do it in command line, but you can do it in CloudFormation. So there’s, kind of, things like this, where for every kind of tool that allows this beautiful, programmable, immutable infrastructure, kind of, paradigm, you now have to make sure that you have security policies that prevent those same tools from being used against you and deploying evil things because you didn’t explicitly say that you can’t deploy evil things with this tool and that tool and that other tool in this other way. Because there’s so many ways to do things, right?

Corey: That’s part of the weird thing, too, is that back when I was doing the sysadmin dance, it was a matter of taking a bunch of tools that did one thing well—or, you know, aspirationally well—and then chaining them together to achieve things. Increasingly, it feels like that’s what cloud providers have become, where they have all these different services with different capabilities. One of the reasons that I now have a three-part article series, each one titled, “17 Ways to Run Containers on AWS,” adding up for a grand total of 51 different AWS services you can use to run containers with, it’s not just there to make fun of the duplication of efforts because they’re not all like that. But rather, each container can have bad acting behaviors inside of it. And are you monitoring what’s going on across that entire threatened landscape?

People were caught flat-footed to discover that, “Wait, Lambda functions can run malware? Wow.” Yes, effectively, anything that can bang two bits together and return a result is capable of running a lot of these malware packages. It’s something that I’m not sure a number of, shall we say, non-forward-looking security teams have really wrapped their heads around yet.

Anna: Yeah, I think that’s fair. And I mean, I always want to be a little sympathetic to the folks, like, in the trenches because it’s really hard to know all the 51 ways to run containers in the cloud and then to be like, oh, 51 ways to run malicious containers in the cloud. How do I prevent all of them, when you have a day job?

Corey: One point that it makes in the report here is that about who the attacks seem to be targeting. And this is my own level of confusion that I imagine we can probably wind up eviscerating neatly. Back when I was running, like, random servers for me for various projects I was working on—or working at small companies—there was a school of thought in some quarters that, well, security is not that important to us. We don’t have any interesting secrets. Nobody actually cares.

This was untrue because a lot of these things are running on autopilot. They don’t have enough insight to know that you’re boring and you have to defend just like everyone else does. But then you see what can only be described as dumb attacks. Like there was the attack on Twitter a few years ago where a bunch of influential accounts tweeted about some bitcoin scam. It’s like, you realize with the access you had, you had so many other opportunities to make orders of magnitude more money if you want to go down that path or to start geopolitical conflict or all kinds of other stuff. I have to wonder how much these days are attacks targeted versus well, we found an endpoint that doesn’t seem to be very well secured; we’re going to just exploit it.

Anna: Yeah. So, that’s correct intuition, I think. We see tons of opportunistic attacks, like, non-stop. But it’s just, like, hitting everything, honeypots, real accounts, our accounts, your accounts, like, everything. Many of them are pretty easy to prevent, honestly, because it’s like just mundane stuff, whatever, so if you have decent security hygiene, it’s not a big deal.

So, I wouldn’t say that you’re safe if you’re not special because none of us are safe and none of us are that special. But what we’ve done here is we actually deliberately wanted to see what would be attacked as a fraction, right? So, we deployed a honey net that was indicative of what a financial org would look like or what a healthcare org would look like to see who would bite, right? And what we expected to see is that we probably—we thought the finance would be higher because obviously, that’s always top tier. But for example, we thought that people would go for defense more or for health care.

And we didn’t see that. We only saw, like, 5% I think for health—very small numbers for healthcare and defense and very high numbers for financial services and telcos, like, around 30% apiece, right? And so, it’s a little curious, right, because you—I can theorize as to why this is. Like, telcos and finance, obviously, it’s where the money is, like, great [unintelligible 00:14:35] for fraud and all this other stuff, right?

Defense, again, maybe people don’t think defense and cloud. Healthcare arguably isn’t that much in cloud, right? Like a lot of health healthcare stuff is on-premise, so if you see healthcare in cloud, maybe, you, like, think it’s a honeypot or you don’t [laugh] think it’s worth your time? You know, whatever. Attacker logic is also weird. But yeah, we were deliberately trying to see which verticals were the most attractive for these folks. So, these attacks are infected targeted because the victim looked like the kind of thing they should be looking for if they were into that.

Corey: And how does it look in that context? I mean, part of me secretly suspects that an awful lot of terrible startup names where they’re so frugal they don’t buy vowels, is a defense mechanism. Because you wind up with something that looks like a cat falling on a keyboard as a company name, no attacker is going to know what the hell your company does, so therefore, they’re not going to target you specifically. Clearly, that’s not quite how it works. But what are those signals that someone gets into an environment and says, “Ah, this is clearly healthcare,” versus telco versus something else?

Anna: Right. I think you would be right. If you had, like… hhhijk as your company name, you probably wouldn’t see a lot of targeted attacks. But where we’re saying either the company and the name looks like a provider of that kind, and-slash-or they actually contain some sort of credential or data inside the honeypot that appears to be, like, a credential for a certain kind of thing. So, it really just creatively naming things so they look delicious.

Corey: For a long time, it felt like—at least from a cloud perspective because this is how it manifested—the primary purpose of exploiting a company’s cloud environment was to attempt to mine cryptocurrency within it. And I’m not sure if that was ever the actual primary approach, or rather, that was just the approach that people noticed because suddenly, their AWS bill looks a lot more like a telephone number than it did yesterday, so they can as a result, see that it’s happening. Are these attacks these days, effectively, just to mine Bitcoin, if you’ll pardon the oversimplification, or are they focused more on doing more damage in different ways?

Anna: The analyst answer: it depends. So, again, to your point about how no one’s safe, I think most attacks by volume are going to be opportunistic attacks, where people just want money. So, the easiest way right now to get money is to mine coins and then sell those coins, right? Obviously, if you have the infrastructure as a bad guy to get money in other ways, like, you could do extortion through ransomware, you might pursue that. But the overhead on ransomware is, like, really high, so most people would rather not if they can get money other ways.

Now, because by volume APTs, or Advanced Persistent Threats, are much smaller than all the opportunistic guys, they may seem like they’re not there or we don’t see them. They’re also usually better at attacking people than the opportunistic guys who will just spam everybody and see what they get, right? But even folks who are not necessarily nation states, right, like, we see a lot of attacks that probably aren’t nation states, but they’re quite sophisticated because we see them moving through the environment and pivoting and creating things and leveraging things that are quite interesting, right? So, one example is that they might go for a vulnerable EC2 instance—right, because maybe you have Log4J or whatever you have exposed—and then once they’re there, they’ll look around to see what else they can get. So, they’ll pivot to the Cloud Control Plane, if it’s possible, or they’ll try to.

And then in a real scenario we actually saw in an attack, they found a Terraform state file. So, somebody was using Terraform for provisioning whatever. And it requires an access key and this access key was just sitting in an S3 bucket somewhere. And I guess the victim didn’t know or didn’t think it was an issue. And so, this state file was extracted by the attacker and they found some [unintelligible 00:18:04], and they logged into whatever, and they were basically able to access a bunch of information they shouldn’t have been able to see, and this turned into a data [extraction 00:18:11] scenario and some of that data was intellectual property.

So, maybe that wasn’t useful and maybe that wasn’t their target. I don’t know. Maybe they sold it. It’s hard to say, but we increasingly see these patterns that are indicative of very sophisticated individuals who understand cloud deeply and who are trying to do intentionally malicious things other than just like, I popped [unintelligible 00:18:30]. I’m happy.

Corey: This episode is sponsored in part by our friends at Calisti.

Introducing Calisti. With Integrated Observability, Calisti provides a single pane of glass for accelerated root cause analysis and remediation. It can set, track, and ensure compliance with Service Level Objectives.

Calisti provides secure application connectivity and management from datacenter to cloud, making it the perfect solution for businesses adopting cloud native microservice-based architectures. If you’re running Apache Kafka, Calisti offers a turnkey solution with automated operations, seamless integrated security, high-availability, disaster recovery, and observability. So you can easily standardize and simplify microservice security, observability, and traffic management. Simplify your cloud-native operations with Calisti. Learn more about Calisti at calisti.app.

Corey: I keep thinking of ransomware as being a corporate IT side of problem. It’s a sort of thing you’ll have on your Windows computers in your office, et cetera, et cetera, despite the fact that intellectually I know better. There were a number of vendors talking about ransomware attacks and encrypting data within S3, and initially, I thought, “Okay, this sounds like exactly a story people would talk about some that isn’t really happening in order to sell their services to guard against it.” And then AWS did a blog post saying, “We have seen this, and here’s what we have learned.” It’s, “Oh, okay. So, it is in fact real.”

But it’s still taking me a bit of time to adapt to the new reality. I think part of this is also because back when I was hands-on-keyboard, I was unlucky, and as a result, I was kept from taking my aura near anything expensive or long-term like a database, and instead, it’s like, get the stateless web servers. I can destroy those and we’ll laugh and laugh about it. It’ll be fine. But it’s not going to destroy the company in the same way. But yeah, there are a lot of important assets in cloud that if you don’t have those assets, you will no longer have a company.

Anna: It’s funny you say that because I became a theoretical physicist instead of experimental physicist because when I walked into the room, all the equipment would stop functioning.

Corey: Oh, I like that quite a bit. It’s one of those ideas of, yeah, your aura just winds up causing problems. Like, “You are under no circumstances to be within 200 feet of the SAN. Is that clear?” Yeah, same type of approach.

One thing that I particularly like that showed up in the report that has honestly been near and dear to my heart is when you talk about mitigations around compromised credentials at one point when GitHub winds up having an AWS credential, AWS has scanners and a service that will catch that and apply a quarantine policy to those IAM credentials. The problem is, is that policy goes nowhere near far enough at all. I wound up having fun thought experiment a while back, not necessarily focusing on attacking the cloud so much as it was a denial of wallet attack. With a quarantined key, how much money can I cost? And I had to give up around the $26 billion dollar mark.

And okay, that project can’t ever see the light of day because it’ll just cause grief for people. The problem is that the mitigations around trying to list the bad things and enumerate them mean that you’re forever trying to enumerate something that is innumerable in and of itself. It feels like having a hard policy of once this is compromised, it’s not good for anything would be the right answer. But people argue with me on that.

Anna: I don’t think I would argue with you on that. I do think there are moments here—again, I have to have sympathy for the folks who are actually trying to be administrators in the cloud, and—

Corey: Oh God, it’s hard.

Anna: [sigh]. I mean, a lot of the things we choose to do as cloud users and cloud admins are things that are very hard to check for security goodness, if you will, right, like, the security quality of the naming convention of your user accounts or something like that, right? One of the things we actually saw in this report it—and it almost made me cry, like, how visceral my reaction was to this thing—is, there were basically admin accounts in this cloud environment, and they were named according to a specific convention, right? So, if you were, like, admincorey and adminanna, like, that, if you were an admin, you’ve got an adminanna account, right? And then there was a bunch of rules that were written, like, policies that would prevent you from doing things to those accounts so that they couldn’t be compromised.

Corey: Root is my user account. What are you talking about?

Anna: Yeah, totally. Yeah [laugh]. They didn’t. They did the thing. They did the good accounts. They didn’t just use root everybody. So, everyone had their own account, it was very neat. And all that happened is, like, one person barely screwed up the naming of their account, right? Instead of a lowercase admin, they use an uppercase Admin, and so all of the policy written for lowercase admin didn’t apply to them, and so the bad guy was able to attach all kinds of policies and basically create a key for themselves to then go have a field day with this admin account that they just found laying around.

Now, they did nothing wrong. It’s just, like, a very small mistake, but the attacker knew what to do, right? The attacker went and enumerated all these accounts or whatever, like, they see what’s in the environment, they see the different one, and they go, “Oh, these suckers created a convention, and like, this joker didn’t follow it. And I’ve won.” Right? So, they know to check with that stuff.

But our guys have so much going on that they might forget, or they might just you know, typo, like, whatever. Who cares. Is it case-sensitive? I don’t know. Is it not case-sensitive? Like, some policies are, some policies aren’t. Do you remember which ones are and which ones aren’t? And so, it’s a little hopeless and painful as, like, a cloud defender to be faced with that, but that’s sort of the reality.

And right now we’re in kind of like, ah, preventive security is the way to save yourself in cloud mode, and these things just, like, they don’t come up on, like, the benchmarks and, like the configuration checks and all this other stuff that’s just going, you know, canned, did you, you know, put MFA on your user account? Like, yeah, they did, but [laugh] like, they gave it a wrong name and now it’s a bad na—so it’s a little bleak.

Corey: There’s too much data. Filtering it becomes nightmarish. I mean, I have what I think of as the Dependabot problem, where every week, I get this giant list of Dependabot freaking out about every repository I have on Gif-ub and every dependency thereof. And some of the stuff hasn’t been deployed in years and I don’t care. Other stuff is, okay, I can see how that markdown parser could have malicious input passed to it, but it’s for an internal project that only ever has very defined things allowed to talk to it so it doesn’t actually matter to me.

And then at some point, it’s like, you expect to read, like, three-quarters of the way down the list of a thousand things, like, “Oh, and by the way, the basement’s on fire.” And then have it keep going on where it’s… filtering the signal from noise is such a problem that it feels like people only discover the warning signs after they’re doing forensics when something has already happened rather than when it’s early enough to be able to fix things. How do you get around that problem?

Anna: It’s brutal. I mean, I’m going to give you, like, my [unintelligible 00:24:28] vendor answer: “It’s just easy. Just do what we said.” But I think [laugh] in all honesty, you do need to have some sort of risk prioritization. I’m not going to say I know the answer to what your algorithm has to be, but our approach of, like, oh, let’s just look up the CVSS score on the vulnerabilities. Oh, look, 600,000 criticals. [laugh]. You know, you have to be able to filter past that, too. Like, is this being used by the application? Like, has this thing recently been accessed? Like, does this user have permissions? Have they used those permissions?

Like, these kinds of questions that we know to ask, but you really have to kind of like force the security team, if you will, or the DevOps team or whatever team you have to actually, instead of looking at the list and crying, being like, how can we pare this list down? Like anything at all, just anything at all. And do that iteratively, right? And then on the other side, I mean, it’s so… defense-in-depth, like, right? I know it’s—I’m not supposed to say that because it’s like, not cool anymore, but it’s so true in cloud, like, you have to assume that all these controls will fail and so you have to come up with some—

Corey: People will fail, processes will fail, controls will fail, and great—

Anna: Yeah.

Corey: How do you make sure that one of those things failing isn’t winner-take-all?

Anna: Yeah. And so, you need some detection mechanism to see when something’s failed, and then you, like, have a resilience plan because you know, if you can detect that it’s failed, but you can’t do anything about it, I mean, big deal, [laugh] right? So detection—

Corey: Good job. That’s helpful.

Anna: And response [laugh]. And response. Actually, mostly response yeah.

Corey: Otherwise, it’s, “Hey, guess what? You’re not going to believe this, but…” it goes downhill from there rapidly.

Anna: Just like, how shall we write the news headline for you?

Corey: I have to ask, given that you have just completed this report and are absolutely in a place now where you have a sort of bird’s eye view on the industry at just the right time, over the past year, we’ve seen significant macro changes affect an awful lot of different areas, the hiring markets, the VC funding markets, the stock markets. How has, I guess, the threat space evolved—if at all—during that same timeframe?

Anna: I’m guessing the bad guys are paying more than the good guys.

Corey: Well, there is part of that and I have to imagine also, crypto miners are less popular since sanity seems to have returned to an awful lot of people’s perspective on money.

Anna: I don’t know if they are because, like, even fractions of cents are still cents once you add up enough of them. So, I don’t think [they have stopped 00:26:49] mining.

Corey: It remains perfectly economical to mine Bitcoin in the cloud, as long as you use someone else’s account to do it.

Anna: Exactly. Someone else’s money is the best kind of money.

Corey: That’s the VC motto and then some.

Anna: [laugh]. Right? I think it’s tough, right? I don’t want to be cliche and say, “Look, oh automate more stuff.” I do think that if you’re in the security space on the blue team and you are, like, afraid of losing your job—you probably shouldn’t be afraid if you do your job at all because there’s a huge lack of talent, and that pool is not growing quick enough.

Corey: You might be out of work for dozens of minutes.

Anna: Yeah, maybe even an hour if you spend that hour, like, not emailing people, asking for work. So yeah, I mean, blah, blah, skill up in cloud, like, automate, et cetera. I think what I said earlier is actually the more important piece, right? We have all these really talented people sitting behind these dashboards, just trying to do the right thing, and we’re not giving them good data, right? We’re giving them too much data and it’s not good quality data.

So, whatever team you’re on or whatever your business is, like, you will have to try to pare down that list of impossible tasks for all of your cloud-adjacent IT teams to a list of things that are actually going to reduce risk to your business. And I know that’s really hard to do because you’re asking now, folks who are very technical to communicate with folks who are very non-technical, to figure out how to, like, save the business money and keep the business running, and we’ve never been good at this, but there’s no time like the present to actually get good at it.

Corey: Let’s see, what is it, the best time to plant a tree was 20 years ago. The second best time is now. Same sort of approach. I think that I’m seeing less of the obnoxious whining that I saw for years about how there’s a complete shortage of security professionals out there. It’s, “Okay, have you considered taking promising people and training them to do cybersecurity?” “No, that will take six months to get them productive.” Then they sit there for two years with the job rec open. It’s hmm. Now, I’m not a professor here, but I also sort of feel like there might be a solution that benefits everyone. At least that rhetoric seems to have tamped down.

Anna: I think you’re probably right. There’s a lot of awesome training out there too. So there’s, like, folks giving stuff away for free that’s super resources, so I think we are doing a good job of training up security folks. And everybody wants to be in security because it’s so cool. But yeah, I think the data problem is this decade’s struggle, more so than any other decades.

Corey: I really want to thank you for taking the time to speak with me. If people want to learn more, where can they go to get their own copy of the report?

Anna: It’s been an absolute pleasure, Corey, and thanks, as always for having us. If you would like to check out the report—which you absolutely should—you can find it ungated at www.sysdig.com/2023threatreport.

Corey: You had me at ungated. Thank you so much for taking the time today. It’s appreciated. Anna Belak, Director of the Office of Cybersecurity Strategy at Sysdig. This promoted guest episode has been brought to us by our friends at Sysdig and I’m Cloud Economist Corey Quinn.

If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an insulting comment that no doubt will compile into a malicious binary that I can grab off of Docker Hub.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

View Details

Colleen Coll, Account Executive at The Duckbill Group, joins Corey on Screaming in the Cloud to discuss her journey of breaking into tech and why it’s so important to make your presence known. Colleen explains how she wound up working for The Duckbill Group after taking the initiative to go and meet Corey at a networking event, and what motivates her to take risks and do things that might feel intimidating in order to advance her career. Colleen and Corey also discuss the power of influencer marketing, as well as the focus The Duckbill Group has on setting a high standard for employee onboarding and culture.

About Colleen

Colleen Coll is a native of Pittsburgh and wannabe tech geek working in tech media sales, events, writing and marketing. She’s an advocate for women and underrepresented communities in tech and is extremely proud of her efforts to learn coding languages and engage and connect diversity in the open source circle! When she’s not geeking out and traveling the globe (and virtually) producing/ hosting tech podcasts and livestreams, she enjoys trips to local wineries, binging sci-fi, and hosting bolognese dinner parties.

Links Referenced:

  • Twitter: https://twitter.com/colleencoll
  • LinkedIn: https://www.linkedin.com/in/colleen-coll-b971505/
  • Last Week in AWS sponsorship form: https://lastweekinaws.com/sponsorship

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Human-scale teams use Tailscale to build trusted networks. Tailscale Funnel is a great way to share a local service with your team for collaboration, testing, and experimentation. Funnel securely exposes your dev environment at a stable URL, complete with auto-provisioned TLS certificates. Use it from the command line or the new VS Code extensions. In a few keystrokes, you can securely expose a local port to the internet, right from the IDE.

I did this in a talk I gave at Tailscale Up, their first inaugural developer conference. I used it to present my slides and only revealed that that’s what I was doing at the end of it. It’s awesome, it works! Check it out!

Their free plan now includes 3 users & 100 devices. Try it at snark.cloud/tailscalescream

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Every once in a while, I like to do a bit of a behind-the-scenes episode where I talk to one of the people here at The Duckbill Group that makes the whole thing run. Because I’ve got to be honest, there’s a certain audience and public perception that everyone on our team's page more or less just sits around and claps as I do all the work. And that’s not true, at least, you know, 80% of the time. My guest today is a relatively new hire here at Duckbill. Colleen Coll is our account executive in media here at The Duckbill Group. Colleen, thank you for joining me, both in an employment sense and on the podcast.

Colleen: [laugh]. Thanks for having me, Corey. This is an honor and privilege [laugh].

Corey: You say that now we’ll see how you feel by the end of this conversation. There’s always that. So, I find when you’re trying to tell a story, one of the best places to begin is the beginning. And we look at people in the space who are doing things that are aspirational or admirable and we have this tendency to believe that they were always this way as if they were formed fully and sprung live from the forehead of some ancient god. And that isn’t how it works. Where do you come from? How do you find yourself now? And where did you start before getting here?

Colleen: I come from a long line—well, possibly a short line—of people who wanted to be a journalist. They wanted to be the All the President’s Men, and that’s what I grew up with. And I took the journalistic ethics classes for it and that’s what I wanted to do. I wanted to be hard-hitting. And then after I graduated, I found that it did not pay well at all, and [laugh] it was a hard jobs to get. It was just hard to get in. So, I had to figure out a plan. And somehow I went through nonprofits, hospitals, and all kinds of marketing gigs, and then I ended up in tech and I still don’t know how I got here. I think it was a dare basically. It was basically a dare [laugh].

Corey: Something that you didn’t do post-journalism that, as best I can tell, is the path that so many, I don’t want to use the term ‘fallen journalist,’ so I won’t, but so many folks who have gone through journalism decide to drift into is in many cases PR and corporate comms. And, on some level, we’ve now hit a point where I think there’s two or three PR professionals for every working journalist, at least. And that ratio is even more skewed in tech. You didn’t go in that direction. Was that not something that appealed? Did it not exist in the timeframe that you were making that transition in the same numbers that it does these days? What was it that wound up, I guess, making that not a path you went on?

Colleen: Actually Corey, I did.

Corey: Oh, wonderful.

Colleen: Yeah, for a short time. I went in as communications, doing a little bit of PR work. Not a lot because that’s not what my focus was. But because—this is what I found out—because you had that writing specialty that you could go into PR, somehow be dropped into it, but then it will lead you to other things like speech writing for whomever VP or C-level. Because I found that a lot of people with Business Economics degrees did not really know how to write or spell for that matter [laugh]. And so, I was sort of used from the bottom once they knew that I had these writing abilities to help in that manner. And that’s-I dropped into PR. It was okay, but I love being engaged in the community more, so they put me into events.

Corey: I did not realize that. And I apologize. It turns out that we don’t do the in-depth, totally invasive style of background check that it seems is becoming more and more common in some places. Imagine that. No, it’s one of those weird areas just because it’s—I deal with so many PR types in different ways and, on some level, it’s easy to fall into the trap of forming a dim view based upon the worst examples because those are the ones that stick in your minds.

But a lot of folks have serious challenges in the corporate comms space and communicating authentically and effectively to the audience, which in many cases is something that, as you said, you take it in a slightly different direction. You do know how to write, you know how to spell, and that sounds like I’m being incredibly sarcastic, like, “Yay, and you can tie your shoes. Triple threat, baby,” but no, it’s not. It is a vanishingly rare skill set in this space, and I see that and get frustrated with it every week when I try and put my newsletter together. And, “What am I going to link to?” I see promising titles that look like it’s going to be germane and if I can’t get through the first paragraph without facepalming, just based upon how poorly it’s written, I don’t want to inflict that on the audience.

People lose sight of the fact that we’re going to record a podcast or we’re going to write blog posts, we’re going to send out an email newsletter. How do we get the best production value for all of this without focusing on the most important part, which is, what are we going to talk about? How are we going to get people to care about it? And that’s the thing that I think gets overlooked the most is the audiences will forgive all kinds of weird production-style issues. It’s okay, this was filmed on your iPhone or it was just posted in plain text on a web dumped somewhere. People will read it if it’s compelling. If it’s not, it doesn’t matter how big the production budget was, it’s not interesting. And if it’s not interesting, no one’s going to care.

Colleen: And I think that’s where I’m sort of a professional in a sense where—and sort of, I’d like to say gifted in a way because I thought it was natural, I still think I’m a natural in storytelling. Regardless of where you are, it’s how you tell your story where people would listen is very, very important. Whether you’re doing a sales pitch or whether you’re trying to sell a bride—from my event management background—just tell the story of how this is going to be successful and to try to sell her the 12-top instead of something else. So yeah, you have to just make it compelling and try to—I like to call myself a kick-ass storyteller, and that’s what’s gotten me here so far, and will get me to where I need to be in the future.

Corey: I would agree with that. You’ve always had a fascinating curiosity, I think is the best approach to this. I still remember the first time we met. It was over a year ago at Monitorama where I threw the annual drink-up—or basically, the drink-up I throw whenever I’m in town somewhere, otherwise, people get annoyed that I didn’t remember to hang out with them personally. And then well, I’ll be there for six weeks. That’s how many lunches I need to book.

And you showed up and introduce yourself, and it was glorious. It was, wow, someone wants to have an actual conversation about things I said on the internet. And for once, I don’t think they’re about to punch me right in my snarky face. Like, this is amazing. How do we get more people like this showing up? And when you applied to work here it was, “Oh, wait is Colleen? Colleen, Colleen?” And that definitely raised eyebrows.

Colleen: Yeah. It was so weird. I love telling that story because it’s a great story to tell because I heard about you where I was before, in the tech industry. And I started following him, like, this guy is… he’s crazy [laugh]. I want to know more.

Corey: Undoubtedly.

Colleen: Exactly. And you were so tech-heavy, but you made it to a point where it was just, like, some kind of the humor and people got you, and I was just so—even though I didn’t understand—I will admit, I don’t understand half the stuff you’re talking about, but the engagement that you had, I was just it was just so compelling, I’m just so interested in that. Because in any kind of work, you want to see, when somebody says something and people respond, and a lot of people respond whether it’s bad or, “Good,” quote-unquote, but I was like, “I got to meet this guy.” And then one time, you were in Portland. Who knew? I thought you’d be somewhere else, and I was, “Like he’s at Momo’s.” Which I live in Portland; it’s right down the street.

I was like, I’m going to—I was on the couch, just laying down [laugh], doing nothing. I got dressed up, put on my eyelashes for you, Corey, and I went to the bar [laugh]. And I said, “You know what?” I was still nervous, I was [unintelligible 00:08:42], but you know what? I’m going to just say, “Hey, Corey. I’m Colleen. I work here. Nice to meet you finally.” So. It was so weird, you and Mike and everybody were just so nice and I just ended up having a—and I got the photo with your mouth open. So, that was awesome.

Corey: Well, that’s the important part. People walk up like I’m the mascot. Like, “Can I get a photo with you? Is that weird?” It’s like, “Yes, it’s extremely weird and you would not be the even fifth hundredth person to ask me for that this year. So sure, by all means.” My face has more or less become a cautionary tale to small children. “You know they say your face is going to freeze like this if you hold it?” “Yeah.”

Colleen: It got me so much street cred though. I was so happy. Like, I didn’t have that many followers in tech and then, soon as I put that picture and I tagged you in it, like, “Oh, my God. You met Corey.”

I was like, “Shut up.” So it’s, thank you. And then how we got here, I have no idea. Again, on Twitter. So [laugh].

Corey: One thing leads to another leads to another. And I have to ask, as someone who is explicitly bad at this, namely, approaching someone I don’t know and striking up a conversation, I have uniformly been terrible at this my entire life, I was bad at dating and honestly, it’s the reason I became a conference speaker because once you give a talk, everyone starts the conversation with you and you’re golden. Was it intimidating for you to come up to someone who you only knew is a loudmouth on Twitter who snarked about everything? And if so, what made you decide to do it anyway?

Colleen: Yes, it was intimidating, I’m not going to lie. Your presence and how you talk and your directness and I didn’t know who you were, I just knew your presence on social media. But you had a lot of followers. But you were here. And I was like, if I don’t take this opportunity, then screw it.

What made me do it is I had to do it for the women and the minorities who always wanted to be in tech and just let them know that you got to do it to make your presence, and you just get up and do it. It’s not that hard. And that’s how it started with me in tech in the first place. Somebody—it was a dare that I take a, I take [laugh] a Python class. And I didn’t—I was like, okay, we’ll do it. It should be easy.

And I did. And I got into it. And that’s how I learned the culture. So, when that opportunity arose and I came up to you, my heart was pounding so fast. I thought you would just like, “Oh, hey. You know, whatever.” And you were just, like, so engaging, and nice and you smiled, and I was like, “Wow. This is it. This is what I want to tell to people.” Even though it might seem hard, it really isn’t once you give your all and just do it and take that chance. I know it sounds very cliche, but that’s what happened.

Corey: It’s an interesting problem, just because the upside feels limited and the downside feels vast. And for me, I’ve gone through an awakening over the last few years as my Twitter audience got to a point where the baseline baked-in assumption I had no longer applied which was, when I started this company, I had less than 1500 followers. And no—all of them had seen me or knew who I was from conferences, so I always assumed that the people who are reading this know who I am, they know what I’m about. And I didn’t really think of the use case of this is someone’s only exposure to me. And I found out a few years back that I was inadvertently causing people pain, which is not what I set out to do, with the singular exception of Larry Ellison, who is not a person, nor does he feel pain.

But as for everyone else, it’s a, I’m not here to make your day actively worse. I’m here to advocate on behalf of, in most cases, customers of which I am one usually, and trying to make tomorrow’s experience better than it is today, and as a part of that, to send the elevator back down. And I realized I was abdicating some of that responsibility, so it started an intentional shift toward being more mindful about how things I say can resonate. And I still get it wrong a lot. And I spent more time apologizing, but that’s something that is going to be a lot more nuanced, tricky, and delicate than I think people give it credit for. And it turns out an apology is not just saying you’re sorry. You actually have to change the behavior.

Colleen: Hmm.

Corey: More people should be aware of that one.

Colleen: Yeah. Well, I think that’s part of the attraction is that you own up to it [laugh]. So.

Corey: I do try.

Colleen: Yeah.

Corey: During the interview process, you redistinguished yourself again and again. Like, one of my personal favorite memories of that was you asked about, I believe it was our event strategy of what do we do at events? And I said, “Yeah, that doesn’t work because the people that buy sponsorships are generally not the same people at the events physically.” And you had the politest framing I can remember of, “Oh, you sweet summer child.” I forget the exact phrasing that you used, but it was simultaneously clear that you did not agree with what I just said in the least, were willing to challenge me on it, but also weren’t going to come lunging across the table with, “Now, you listen here,” all of which I have seen before.

It was the perfect mix of, “Ah, yeah, you just said something that’s complete bullshit, but I’m just going to say something that lets you figure that out for yourself rather than leading you by hand down the journey of discovery of what a dumbass you’re being.” And lo and behold, you were absolutely right. That’s one of the more useful and also irritating aspects is when you have people who come in who are better in than you are at the job you hired them to do. It’s a constant humbling process of hey, I thought I knew what I was doing, and guess what? I absolutely did not. And now I just step out of the way and hope I don’t wind up causing problems accidentally.

Colleen: Well, you’re welcome [laugh]. It’s my pleasure.

Corey: It’s worked out rather well.

Colleen: [laugh]. Yes, yes. Yeah, I am glad that I had the opportunity to enlighten you and couple other people in the staff at the Duckbill Group. So yes, I definitely—I will definitely fall—what do you call it? A fall on the sword [laugh] or die on a hill that particular topic: event management, sales, community, it all works.

Corey: Before this, you spent a few years at The New Stack and that was honestly one of our biggest internal fears. We’re a big fan of Alex and what he’s been able to build over there. It’s like, is he going to hate us if we extend an offer to you? And of course, we could not ask him that question in advance because it’s, “So, one of your employees is interviewing and we’d like to make them an offer. Is this going to cause a problem for you?”

Yeah, that’s called how you potentially just absolutely gut someone who took a chance on talking to you. Confidentiality in those things is required because you never know the actual story. And there’s a power imbalance in the job interview process of, “Hey, do you mind if I talk to your current boss about bringing you aboard?” And what did they going to say? “Please don’t?” Because, oh, that’s going to potentially ring hollow. And there’s nothing nefarious in it, but we knew there was no way we could ask it. So it’s, well, we’ll send him a lovely fruit basket and an apology note. Which I still need to get out. If you’re listening to this, Alex, my apologies. It’s on my backlog.

Colleen: [laugh]. Yes, that’s awesome. Now, that was a difficult decision. I wasn’t even looking. I was on Twitter and I saw the opportunity. And I love being in the company of journalists for the first time. It’s funny, I started off as a journalist, like, post-college and just writing a few things as a freelancer, and I ended up in corporate world.

And then somehow I finally got back to working with journalists at The New Stack. I mean, these are writers talking about tech and writing about tech. And it was fantastic and I got to be a part of that. But when you get to phase where I am, my age [laugh], my experience, I wanted to do more. And when I saw that opportunity—and the opportunity to work with your brand because I met you last year—I was like, you know what? Well, we’ll see. It’s who’ll know—who—you know, we’ll see what happens. And I ended up having this won—

Corey: You meet me for 30 seconds, you’re like, “Well, I already had a perfect answer to ‘so why didn’t it work out?’” It’s like, “Have you met that jackhole? There we go.” No one is going to have a follow-up to that. You’re always going to be assured that yep, that one is not going to come back on you.

Colleen: Yeah. The New Stack is a wonderful place to start if you want to go into tech media, and just—and it’s not a niche market like we have here, but it’s just all around what’s going on and the trends. And the people there taught me so much. So, I think that’s kind of why I had the opportunity here. If I didn’t have that opportunity at The New Stack, I don’t think I would have had this opportunity to have these conversations with you, this team. So, yeah.

Corey: I have to wonder, on some level, though, there’s a this is niche upon niche because not only is it, we wind up focusing exclusively on the AWS market, but not only that, we also are incredibly sarcastic and generally make fun of you. Would you like to buy a sponsorship? I always thought that that was a ridiculous pitch that was never going to work, but sponsors have come in repeatedly and they still talk to us and still asked to give us money, which is frankly, somewhat surprising to me. And to my understanding that is no longer the exact pitch we give verbatim—because it’s not strictly true—but it does feel like it’s a harder conversation, then. “So, what do you do?” “We’re a news site.” “What do you do there?” “We report the news.” “How do I sponsor?” “You give us money, and we put banner ads on the website, and possibly sponsored content. The end.” This feels like it’s much more nuanced and as a result is probably a harder sales conversation.

Colleen: You think so? I’ve noticed just by being in the know of tech and going into and reading everything about media and everything, people are—especially when it comes to buying—people are more apt—and you know this—to buy from influencers and customers than the actual company. And when you have somebody that’s constantly in the know, tech-heavy like yourself or somebody with another product, whether it’s nail polish or something like that, and they used it, they gave this wonderful review or gave a bad review, they’re more apt to buy or not buy. I think that’s why, that’s the connection with the company, to somebody, the big company to purchase what we have to offer here at Last Week in AWS is because you’re buying sort of an influence, and you have the audience of customers who believe in you and believe most of the things that you say [laugh] when you’re not shitposting. And [laugh]—

Corey: Wait, when am I not shitposting?

Colleen: [laugh]. Exactly. And so, people a—I think that that’s what these sponsors are buying. They want to buy the influence in the customers’ view because they know that they’ll get more buy-in. So, I think that whole buy this product because I am Heinz ketchup, that whole generation is gone. It’s, buy the Heinz ketchup because what his name is using it on his hot dog all the time and he just absolutely loves it and Tik Toks about it all the time. I’m so I think that’s why, I think it’s a—I don’t want to say it’s an easy sell; everything’s not that easy, but it’s a fun and more compelling sell here.

Corey: This episode is sponsored in part by our friends at Calisti.

Introducing Calisti. With Integrated Observability, Calisti provides a single pane of glass for accelerated root cause analysis and remediation. It can set, track, and ensure compliance with Service Level Objectives.

Calisti provides secure application connectivity and management from datacenter to cloud, making it the perfect solution for businesses adopting cloud native microservice-based architectures. If you’re running Apache Kafka, Calisti offers a turnkey solution with automated operations, seamless integrated security, high-availability, disaster recovery, and observability. So you can easily standardize and simplify microservice security, observability, and traffic management. Simplify your cloud-native operations with Calisti. Learn more about Calisti at calisti.app.

Corey: I want to be very clear a nuance here that I’m not sure it’s fully understood, in that years ago, I made a very intentional choice of severing myself almost completely from the sponsor sales process here. And the reason is, is that I never wanted to find myself in a position of writing the weekly news and, I don’t know, let’s pick on a former sponsor, for example, Google Cloud does something that I’m about to dunk on. But oh, it turns out that Google is also sponsoring this issue so I probably shouldn’t do that. I built my own version of an editorial firewall so I did not have that conflict. So, I say what I want, I don’t find out who’s sponsoring something until afterwards and to be very clear, to this day, I have never had a single complaint or piece of pushback on anything I’ve written from a sponsor who is sponsoring that issue, which is, frankly, tells me that it’s sort of unnecessary from an external perspective, but it makes it work better for me.

So, I don’t know what a lot of the sales conversations look like. People reach out, “Hey, can I sponsor your stuff?” And it’s, “Have you met Colleen?” And I get the hell out of that critical path as fast as possible. Also because I’m bad at email. And that just means that I’m more or less have a mystery box that I throw all of those things into and then sponsors come out the other side. And, lovely, I’ll take it.

Colleen: Yeah, that’s—I mean… should you? You should get updates on who wants to buy and why, but most of the time, it’s the audience. They want to connect to the audience that you have created, well, the company has created. And basically they’ll know about this, I mean, eventually, they’ll know about the services we provide, whether it’s consulting, and then of course, the opportunities for ad placements at Last Week in WSat AWS. Oh, God, I need to really improve that [laugh].

But it’s the audience. So, why if you would shitpost something or say something that might make them uncomfortable, why they would buy it anyway, I mean, that’s a conversation that maybe we should keep having. But I think the answer is clear. I mean, it’s the audience who believe in you and believe in what you have to say and believe in our brand. So, they want to get close to it in order to sell their product. And they have the money and means to do so. So, I don’t question it too much.

Corey: No. It’s similar to the whole approach that I always take is I don’t think too hard about what keeps the airplane in the air when I’m mid-flight because if I do, it might stop working. Similar here. It’s like, I don’t know why these people keep showing up and listening to what I have to say and caring about it. I’m not going to look too closely at it because then the magic might break. But that is probably at this point not the most helpful instinct I could have.

Colleen: Yeah. That’s good point, yeah. I think what you’re doing is actually really great and it’s keeping tech media interesting.

Corey: I try anyway. To turn it around slightly, though, I have to ask you if you knew my public persona for probably entirely too long and then you got to actually work inside the sausage factory, and—which is a polite way of saying abattoir—and as you move aside as someone’s, like, shoving a cow into a food processor behind you or whatnot, I have to ask, what’s caught you by surprise once you got here that you did not know or expect before you joined?

Colleen: I think what caught me by surprise is two things. The onboarding was, for such a small company, was immaculate. I didn’t have to ask too many questions; things were there. And the questions that I asked were answered. So, the culture was so open to a point with feedback and how we do things that I didn’t have to do most of the work and info gathering when you’re going from 30, 60, to 90 days. So, I was shocked by that because usually in my past [laugh], I’m usually just, “Hey, got a job. Figure it out.” [laugh]. So, I went in that way, but it was just, it was fantastic.

Then I’m not used to this, especially when you’re the only… when people don’t look like you [laugh] and you’re usually the only one—especially when it comes to CEOs and founders—the amount of openness, friendliness, and direct feedback that wasn’t as condescending—because this is what we expect most of the time—was just fantastic. It was just friendly and I can do my job without having to be attacked or attack any other mindsets that they might have some stereotypes of how—who I am and how I got here. It was just, you were just so professionally, “Colleen, this is what we expect. How do you feel about this?” The, how do you feel about this? Do you have any other ways or feedback that this could be better?

And I know this sounds so corny, but when it comes to people like me, we don’t always get that opportunity before—like, so fast, before we have to prove ourselves. I know, that’s a long-winded way of saying that [laugh]. And so far, it’s been a delight to a point where I sort of have a little bit of PTSD because I’m, like, how do I operate in this non-toxic [laugh] environment?

Corey: And I don’t know if you recall this, but you made the observation that in many places, there is an undercurrent of bias, be it conscious or otherwise, that causes people to out of hand reject proposals or ideas that come from people who do not look like the traditional person you would expect in that role to be framing those ideas. And how much of that would you encounter here? And that I thought was a poignant question that deserves a great answer. And my answer then remains as it is now, which is, “I don’t know. I don’t believe that we have that type of culture here.”

But again, I wouldn’t believe that we have that type of culture here, even if it were rampant. So, I would consider it a personal favor if you see elements of that to please let someone you trust here know that because it is certainly not our intent, it is certainly not who we aspire to be, but societal and systemic patterns are incredibly hard to break. And I don’t know what a good answer to that would be. I know the bad answers are obvious of, “Oh, we don’t hire anyone like that.” Or, “Nope, that’s not a problem.” Or the worst, I suppose is, “How dare someone who doesn’t look like me ask me that question,” which I’m pretty sure gets the high score for terrible answers. But I don’t know what the good answer to that is, other than we’re always learning and trying.

Colleen: And that’s basically how you did respond. And it was eye-opening. And it’s not an easy question to ask. I’m like, “Will you have a problem with someone like me giving you feedback on something like this? Will you have a problem that someone that looks like me working this and doing this, and you know, just trying to do the job, or do I have to make you feel comfortable first?”

And I’m at a point in my life where I don’t have time to do that. I would rather just go on, let me do my work without, you know, making other people feel uncomfortable or feel comfortable. So, I will tell you, it’s almost been two months. This is the first time I haven’t had that feeling of trying to make people feel comfortable before I can actually do my job. And I am not kissing ass because you know, I’m really direct in that approach because—

Corey: I have not known you to ever kiss ass—

Colleen: [laugh].

Corey: Which is probably a good thing, and also, some of them may be disturbing, but I don’t know. It’s like, “Oh, you’re not authorized to fire me. It’s fine.” Which I’m in fact not, so… cool, by all means. But no, it is refreshing. It really is.

Colleen: It is. It’s very refreshing. I want to—if I can even tell people out there that there are places like that you don’t always have to use 50% of your time attacking stereotypes and you can actually do your job. They do exist out there. And this position so far has been living proof. And I do appreciate it.

But I also want them to make sure that they know that I worked hard at this to get here and it’s good to be appreciated, but it’s also good to be respected and valued. And I do believe that you, that’s why you hired me is because you saw what I was capable of and you valued my input, my feedback, and it’s still going on, and we keep having these conversations. And I did not expect this interview, which I was like, “Is he serious?” I mean really. Why [laugh]?

So, this is just another, like, example of how that—what I just talked about is being respected and valued, and regardless of if I don’t look like you. And one of the funniest parts of our interview is when you said that whole manel description of how if you were asked to be on a panel, but it was a bunch of white males, and you refer to it as a manel, I’ve never heard of that [laugh] before, so I u—

Corey: It’s not my term. I don’t want to claim credit for it at all. I heard from some wit on Twitter years ago that is lost to the mysteries of time. But it’s a perfect description.

Colleen: So, I use it. I steal it. It’s awesome.

Corey: It seems like one of those weird areas, too, where it’s a—like, we’re going to get stuff wrong. That is human nature. The question is, is when it’s pointed out, how do you react? Do you get hyper-defensive? Do you just, like, turn that into a cudgel to beat other people with? Or do you take the lesson? Do you pass it forward to folks in a way that is constructive and helpful?

And I believe one of the rejoinders I asked to you was, if you have an idea, we are absolutely going to hear it out, but there are going to be cases where… like, in the consulting side of our business, whenever I describe that we fix the AWS bill for our clients, I explain that to an engineer, and they think hard on that for two-and-a-half seconds and then they say the same thing in almost every case, which is, “You should charge a percentage of savings,” to a point where now my default reflexive response is, “Holy shit. I never thought of that. This is going to change everything.” Now, there are a variety of reasons that that doesn’t work, but it is an obvious line of inquiry. And the only concern I had was, understand that there are things like that scattered throughout the business that things are like they are for a reason.

And I’m thrilled to reevaluate and reexamine a bunch of those, but are you going to take the response of, “Well, this is why it is the way that it is,” as shutting down the line of inquiry? And your response was incredibly reassuring. You said, “No, that is strictly a business discussion. That’s fine. I just want to be heard.”

And I can’t commit to always agreeing with ideas you have. I can’t even give it to that to my business partner. In fact, correcting him is my favorite part of any given hour for me, but I can at least guarantee you’ll be heard or you’ll be—or I will hear you out on any of these concerns that you raise or ideas that you have. And I’d like to think that almost three months in, that we’ve lived up to that. And if not, please let us know. You don’t need to actually call me out on it now if you don’t want to. I realize, like, yeah, well, that’s the right time to ask someone for harsh feedback, at a point where they cannot possibly give it to you other than that a really flattering way. Go for it. You need not respond.

Colleen: [laugh]. No, I will keep that in mind. And so far, so good. We are good [laugh].

Corey: I really want to thank you for taking the time to sit down and basically have to, I suppose justify after the fact why you accepted the job that we offered to you, which feels very strange, and yet here we are. If people want to learn more, where’s the best place for them to find you?

Colleen: Oh, where the best place to find me? Shall I mention Twitter [laugh]? Or [laugh]—

Corey: That’s always a bit of a dicey thing these days.

Colleen: Well, Twitters, Threads, LinkedIn, wherever you, your heart’s desire, you can find me at Colleen Coll. It’s really easy.

Corey: We’ll put links to all of that in the [show notes 00:32:26].

Colleen: Yes. But you can also find me in Portland and sometimes in Europe, and always just being open. And I love to meet people. I know that sounds weird, but if I have the opportunity to network, it’s going to be—and please, if you ever see me at an event, just please walk up to me and say hello. And I—because I know that I would do the same with you. I did that with Corey [laugh].

Corey: Or there’s always the guaranteed way to make sure that you see something and that is to fill out the form at lastweekinaws.com/sponsorship. There’s a little self-interest behind that one I absolutely am aware of and I’m putting that in there.

Colleen: Nice.

Corey: Thank you again for your time. I appreciate it.

Colleen: No, thank you, Corey. And have fun with the squirrels and FedEx.

Corey: I’ll do my best. Colleen Coll, account executive here at The Duckbill Group. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an angry comment complaining about the episode disparaging the value of writing clearly and journalism is particular, and of course failing to have anything remotely resembling a coherent sentence structure while you do.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

View Details

Evelyn Grizzle, Senior Salesforce Developer, joins Corey on Screaming in the Cloud to discuss the often-misunderstood and always exciting world of Salesforce development. Evelyn explains why Salesforce Development is still seen as separate from traditional cloud development, and describes the work of breaking down barriers and silos between Salesforce developers and engineering departments. Corey and Evelyn discuss how a non-traditional background can benefit people who want to break into tech careers, and Evelyn reveals the best parts of joining the Salesforce community.

About Evelyn

Evelyn is a Salesforce Certified Developer and Application Architect and 2023 Salesforce MVP Nominee. They enjoy full stack Salesforce development, most recently having built a series of Lightning Web Components that utilize a REST callout to a governmental database to verify the licensure status of a cannabis dispensary. An aspiring Certified Technical Architect candidate, Evelyn prides themself on deploying secure and scalable architecture. With over ten years of customer service experience prior to becoming a Salesforce Developer, Evelyn is adept at communicating with both technical and non-technical internal and external stakeholders. When they are not writing code, Evelyn enjoys coaching for RADWomenCode, mentoring through the Trailblazer Mentorship Program, and rollerskating.

Links Referenced:

  • Another Salesforce Blog: https://anothersalesforceblog.com
  • RAD Women Code: https://radwomen.org/
  • Personal Website: https://evelyn.fyi
  • LinkedIn: https://www.linkedin.com/in/evelyngrizzle/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn, and this is Screaming in the Cloud. But what do we mean by cloud? Well, people have the snarky answer of, it’s always someone else’s computer. I tend to view it through a lens of being someone else’s call center, which is neither here nor there.

But it all seems to come back to Infrastructure as a Service, which is maddeningly incomplete. Today, we’re going in a slightly different direction in the world of cloud. My guest today is Evelyn Grizzle, who, among many other things, is also the author of anothersalesforceblog.com. I want to be clear, that is not me being dismissive. That is the actual name of the blog. Evelyn, thank you for joining me.

Evelyn: Hi, Corey. Thank you for having me.

Corey: So, I want to talk a little bit about one of the great unacknowledged secrets of the industry, which is that every company out there, sooner or later, uses Salesforce. They talk about their cloud infrastructure, but Salesforce is nowhere to be seen in it. But, for God’s sake, at The Duckbill Group, we are a Salesforce customer. Everyone uses Salesforce. How do you think that wound up not being included in the narrative of cloud in quite the same way as AWS or, heaven forbid, Azure?

Evelyn: So, Salesforce is kind of at the proverbial kid’s table in terms of the cloud infrastructure at most companies. And this is relatively because the end-users are, you know, sales reps. We’ve got people in call centers who are working on Salesforce, taking in information, taking in leads, opportunities, creating accounts for folks. And it’s kind of seen as a lesser service because the primary users of Salesforce are not necessarily the techiest people on the planet. So, I am really passionate about, like, making sure that end-users are respected.

Salesforce actually just added a new certification, the Sales Representative Certification that you can get. That kind of gives you insight to what it’s like to use Salesforce as an end-user. And given that Salesforce is for sales, a lot of times Salesforce is kind of grouped under the Financial Services portion of a company as opposed to, like, engineering. So again, kind of at the proverbial kid’s table; we’re over in finance, and the engineering team who’s working on the website, they have their engineering stuff.

And a lot of people don’t really know what Salesforce is. So, to give a rundown, basically, Salesforce development is, I lovingly referred to it as bastard Java full-stack development. Apex, the proprietary language, is based in Java, so you have your server-side Java interface with the Salesforce relational database. There’s the Salesforce Object Query Language and Salesforce Object Search Language that you can use to interact with the database. And then you build out front-end components using HTML and JavaScript, which a lot of people don’t know.

So, it’s not only an issue of the end-users are call center reps, their analysts, they’re working on stuff that isn’t necessarily considered techie, but there’s also kind of an institutional breakdown of, like, what is Salesforce? This person is just dragging and dropping when that isn’t true. It’s actually, you know, we’re writing code, we’re doing stuff, we’re basically writing full-stack Java. So, I like to call that out.

Corey: I mean, your undergraduate degree is in network engineering, let’s be very clear. This is—I’m not speaking to you as someone who’s non-technical trying to justify what they do as being technical. You have come from a very deep place that no one would argue is, “Well, that’s not real computering.” Oh, I assure you, networking is very much real computering, and so is Salesforce. I have zero patience for this gatekeeping nonsense we see in so many areas of tech, but I found this out firsthand when we started trying to get set up with Salesforce here. It took wailing and gnashing of teeth and contractor upon contractor. Some agencies did not do super well, some people had to come in and rescue the project. And now it mostly—I think—works.

Evelyn: Yeah, and that’s what we go for. And actually, so my degree is in network engineering, but an interesting story about me. I actually went to school for chemical engineering. I hated it. It was the worst. And I dropped out of school did, like, data analytics for a while. Worked my way up as a call center rep at a telephone company and made a play into database administration. And because I was working at the phone company, my degree is in network engineering because I was like, “I want to work at the phone company forever.” Of course that did not pan out. I got a job doing Salesforce development and really enjoy it. There’s always something to learn. I taught myself Salesforce while I was working at IBM, and with the Blue Wolf department that… they’re a big Salesforce consulting shop at IBM, and through their guidance and tutelage, I guess, I did a lot of training and worked up on Salesforce. And it’s been a lot of fun.

Corey: I do feel that I need to raise my hand here and say that I am in the group you described earlier of not really understanding what Salesforce is. My first real exposure to Salesforce in anything approaching a modern era was when I was at a small consulting company that has since been bought by IBM, which rather than opine on that, what I found interesting was the Salesforce use case where we wound up using that internally to track where all the consultants were deployed, how they wound up doing on their most recent refresher skills assessment, et cetera, so that when we had something strange, like a customer coming in with, “I need someone who knows the AS/400 really well,” we could query that database internally and say, “Ah. We happen to have someone coming off of a project who does in fact, know how that system works. Let’s throw them into the mix.” And that was incredibly powerful, but I never thought of it as being a problem that a tool that was aimed primarily at sales would be effective at solving. I was very clearly wrong.

Evelyn: Yeah. So, the thing about Salesforce is there’s a bunch of different clouds that you can access. So there’s, like, Health Cloud, Service Cloud, Sales Cloud is the most common, you know, Salesforce, Sales Cloud, obviously. But Service Cloud is going to be a service-based Salesforce organization that allows you to track folks, your HR components, you’re going to track your people. There’s also Field Service Lightning.

And an interesting use case I had for Field Service Lightning, which is a application that’s built on top of Salesforce that allows field technicians to access Salesforce, one of the coolest projects I’ve built in my career so far is, the use case is, there’s an HVAC company that wants to be able to charge customers when they go out into the field. And they want to have their technician pull out an iPad, swipe the credit card, and it charges the customer for however much duct tape they used, however much piping, whatever, duct work they do. Like I said, I’m a software engineer, I’m not a HVAC person, but—

Corey: It’s the AWS building equivalent for HVAC, as best I can tell. It’s like all right, “By the metric foot-pound—” “Isn’t that a torque measurement?” “Not anymore.” Yeah, that’s how we’re going to bill you for time and materials. It’ll be great.

Evelyn: Exactly. So, this project I built out, it connects with Square, which is awesome. And Field Service Lightning allows this technician to see where they’re supposed to go on the map, it pulls up all the information, a trigger in Salesforce, an automation, pulls all the information into Field Service Lightning, and then you run the card, it webhooks into Square, you send the information back. And it was a really fun project to work on. So, that was actually a use case I had not thought of for Salesforce is, you know, being able to do something like this in the field and making a technician’s job that much easier.

Corey: That’s really when I started to feel, as this Salesforce deployment we were doing here started rolling out, it wasn’t just—my opinion on it was like, “Wait, isn’t this basically just that Excel sheet somewhere that we can have?” And it starts off that way, sure, but then you have people—for example, we’ve made extensive use of aspects of this over on the media side of our business, where we have different people that we’ve reached out to who then matriculate on to other companies and become sponsors in that side of the world. And how do we track this? How do we wind up figuring out what’s currently in flight that doesn’t live in someone’s head, or God forbid, email inbox? How do we start reasoning about these things in a more holistic way?

We went in a slightly different direction before rolling it out to handle all of the production pieces and the various things we have in flight, but I could have easily seen a path whereas we instead went down that rabbit hole and used it as more or less the ERP, for lack of a better term, for running a services business.

Evelyn: Yeah. And that is one thing you can use Salesforce as an ERP. FinancialForce, now Certinia, exists, so it is possible to use Salesforce as an ERP, but there’s so much more to it than that. And Salesforce, at its heart, is a relational database with a fancy user interface. And when I say, “I’m a Salesforce developer,” they’re like, “Oh, you work at Salesforce?” And I’m like, “No, not quite. I customize Salesforce for companies that purchase Salesforce as a Salesforce customer.”

And the extensibility of the platform is really awesome. And you know, speaking of the external clients that want to use Salesforce, there’s, like, Community Cloud where you can come in and have guest users. You can have your—if you are, say at a phone company, you can have a troubleshooting help center. You can have chatbots in Salesforce. I have a lot of friends who are working on AI chatbots with the Einstein AI within Salesforce, which is actually really cool. So, there is a lot of functionality that is extensible within Salesforce beyond just a basic Excel spreadsheet. And it’s a lot of fun.

Corey: If I pull up your website, anothersalesforceblog.com, one of the first things that you mentioned on the About the Author page just below the fold, is that you are an eight-time Salesforce Certified Developer and application architect. Like, wow, “Eight different certifications? What is this, AWS, on some level?”

I think that there’s not a broad level of awareness in the ecosystem, just how vast the Salesforce-specific ecosystem really is. It seems like there’s an entire, I want to reprise the term that someone—I can’t recall who—used to describe Dark Matter developers, the people that you don’t generally see in most of the common developer watering holes like Stack Overflow, or historically shitposting on Twitter, but they’re out there. They rock in, they do their jobs. Why is it that we don’t see more Salesforce representation in, I guess, the usual tech watering holes?

Evelyn: So, we do have a Stack Overflow, a Stack Exchange as well. They are separate entities that are within the greater Stack websites. And I assure you, there’s lots of Salesforce shitposting on Twitter. I used to be very good at it, but no longer on Twitter due to personal reasons. We’ll leave it at that.

But yeah, Dreamforce is like a massive conference that happens in San Francisco every year. We are gearing up for that right now. And there’s not a lot of visibility into Salesforce outside of that it feels like. It’s kind of an insulated community. And that goes back to the Salesforce being at the kids’ table in the engineering departments.

And one of the things that I’ve been working on in my current role is really breaking down the barriers and the silos between the engineering department who’s working on JavaScript, they’re working on Node, they’re working on HTML, they’re, you know, building websites with React or whatever, and I’m coming in and saying, like, hey, we do the same thing. I can build a Heroku app in React, if I want to, I can do PHP, I can do this. And that’s one of the cool things about Salesforce is some days I get to write in, like, five or six different languages if I want to. So, that is something that, there’s not a lot of understanding. Because again, relational database with a fancy user interface.

To the outside, it may seem like we’re dragging and dropping stuff. Which yes, there is some stuff. I love Flows, which are… they’re drag-and-drop automations that you can do within Salesforce that are actually really powerful. In the most recent update, you can actually do an HTTP call-out in a Flow, which is something that’s, like, unheard of for a Salesforce admin with no coding background can come in, they can call an Apex class, they can do an HTTP call-out to an external resource and say, like, “Hey, I want to grab this information, pull it back into Salesforce, and get running off the ground with, like, zero development resources, if there are none available.”

Corey: I want to call out just for people who think this is more niche than it really is. I live in San Francisco. And I remember back in pre-Covid times, back when Dreamforce was in town. I started seeing a bunch of, you know, nerdy-looking people with badges. Oh, it’s a tech conference, what conference is it? It’s something called Dreamforce for Salesforce.

Oh, is that like the sad small equivalent of re:Invent in Las Vegas? And it’s no, no, it’s actually about three times the size. 170,000 people descend on San Francisco to attend this conference. It is massive. And it was a real eye-opener for me just to understand that. I mean, I have a background in sales before I got into tech and I did not realize that this entire ecosystem existed. It really does feel like it is more or less invisible and made me wonder what the hell else I’m missing, as I am too myopically focused on one particular giant cloud company to the point where it has now become a major facet of my personality.

Evelyn: And that’s the thing is there’s all kinds of community events as well. So, I’m actually speaking at Forcelandia which, it’s a Salesforce developer-focused event that is in Portland—Forcelandia, obviously—and I’m going to be speaking on a project that I built for my current company that is, like, REST APIs, we’ve got some encryption, we’ve got a front-end widget that you drop into a Salesforce object. Which, a Salesforce object is a table within the relational database, and being able to use polymorphic object relationships within Salesforce and really extending the functionality of Salesforce. So, if you’re in Portland, I will be at Forcelandia on July 13th and I’m really excited about it.

But it’s this really cool ecosystem that, you know, there’s events all over the world, every month, happening. And we’ve got Mile High Dreamin’ coming up in August, which I’ll be at as well, speaking there on how to break into the ecosystem from a non-tech role, which will be exciting. But yeah, it’s a really vibrant community like, and it’s a really close-knit community as well. Everyone is so super helpful. If I have a question on Stack Exchange, or, you know, back in my Twittering days, if I’d have something on Twitter, I could just post out and blast out, and the whole Salesforce community would come in with answers, which is awesome. I feel like the Stack Exchange is not the friendliest place on the planet, so to be able to have people who, like, I recognize that username and this person is going to come and help me out. And that’s really cool. I like that about the Salesforce community.

Corey: Yeah, a ding for a second on the whole Stack Exchange thing. That the Stack Overflow survey was fascinating, and last year, they showed that 92% of their respondents were male. So, this year, they fixed that problem and did not ask the question. So, I just refer to it nowadays as Stack Broverflow because that’s exactly how it seems.

Evelyn: [laugh].

Corey: And that is a giant problem. I just didn’t want that to pass uncommented-on in public. Thank you for giving me the opportunity to basically—

Evelyn: Fair enough.

Corey: —mouth off about that crappy misbehavior.

Evelyn: Oh, yeah. No. And that’s one of the things that I really like about the Salesforce community is there’s actually, like, a huge movement towards gender equity and parity. So, one of the organizations that I’m involved with is RAD Women Code, which is a nonprofit that Angela Mahoney and a couple of other women started that it seeks to upskill women and other marginalized genders from Salesforce admins, which are your declarative users within Salesforce that set up the security settings, they set up the database relationships, they make metadata changes within Salesforce, and take that relational database knowledge and then upskill them into Salesforce developers.

And right now, there is a two-part course that you can sign up for. If you have I believe it’s a year or two of Salesforce admin experience and you are a woman or other marginalized gender, you can sign up and take part one, which is a very intro to computer programming, you go over the basics of object-oriented programming, a little bit of Java, a little bit of SOQL, which is the Salesforce Object Query Language. And then you build projects, which is really awesome, which is, like, the most effective way to learn is actually building stuff. And then the second part of the course is, like, a more advanced, like, let’s get into our bash classes, which is like an automation that you can run every night. Let’s do advanced object-oriented programming topics like abstraction and polymorphism. And being able to teach that is really fun.

We’re also planning on adding a third course, which is going to be the front-end development in Salesforce, which is your HTML, your JavaScript. Salesforce uses vanilla JavaScript, which I love, personally. I know I’m alone in that. I know that’s the big meme on Facebook in the programming groups is ‘JavaScript bad,’ but I have fun with it. There’s a lot you can do with just native JavaScript in Salesforce. Like, you can grab the geolocation of a device and print it onto a Salesforce object record using just vanilla JavaScript. And it’s been really helpful. I’ve done that a few times on various projects.

But yeah, we’re planning on adding a third course. We are currently getting ready to launch the pilot program on that for RAD Women Code. So, if you are listening to this, and you are a Salesforce admin who is a marginalized gender, definitely hit me up on LinkedIn and I will send you some information because it’s a really good program and I love being able to help out with it.

Corey: We’ll definitely include links to that in the [show notes 00:18:59]. I mean, this does tie into the next question I have, which is, how do you go about giving a cohesive talk or even talking at all about Salesforce, given the tremendous variety in terms of technical skills people bring to bear with it, the backgrounds that they have going into it? It feels, on some level, like, it’s only a half-step removed from, “So, you’re into computers? Here’s a conference for that.” Which I understand, let’s be clear here, that I am speaking from the position of the AWS ecosystem, which is throwing stones in a very fragile glass house.

Evelyn: Yeah, so again, I said this already. When I say I’m a Salesforce developer, people say, “Oh, you work at Salesforce. That is so cool.” And I have to say, “No, no. No working at Salesforce. I work on Salesforce in the proprietary system.” But there’s always stuff to be learned. There’s obviously, like, two releases a year where they send updates to the Salesforce software that companies are running on and working on computers is kind of how I sum it up, but yeah, I don’t know [laugh].

Corey: No, I think that’s a fair place to come at from. It’s, I think that we all have a bit of a bias in that we tend to assume that other people, in the absence of data to the contrary, have similar backgrounds and experiences to our own. And that means in many cases, we paper over things that are not necessarily true. We find ourselves biasing for people whose paths resemble our own, which is not inherently a bad thing until it becomes exclusionary. But it does tend to occlude the fact that there are many paths to this broader industry.

Evelyn: Yeah. So, there is a term in the Salesforce ecosystem, we like to call people accidental admins, where they learn Salesforce on a job and like it so much that they become a Salesforce admin. And a lot of times these folks will then become developers and then architects, even, which is kind of how I got into it as well. I started at a phone company as a Salesforce end-user, worked my way up as a database admin, database coordinator doing e911 databases, and then transitioned into software engineering from there. So, there’s a lot of folks who find themselves within the Salesforce ecosystem, and yeah, there are people with, like, bonafide top-ten computer science school degrees, and you know, we’ve got a fair bit of that, but one thing that I really like about the Salesforce ecosystem is because everyone’s so friendly and helpful and because there’s so many resources to upskill folks, it’s really easy to get involved in the ecosystem.

Like Trailhead, the training platform for Salesforce is entirely free. You can sign up for an account, you can learn anything on Salesforce from end-user stuff to Salesforce architecture and anything in between. So, that’s how most people study for their certifications. And I love Trailhead. It’s a very fun little modules.

It gamifies learning and you get little, I call them Girl Scout badges because they resemble, you know, you have your Girl Scout vest and your Girl Scout sash, and you get the little badges. So, when you complete a project, you get a badge—or if you work on a big project, a super badge—that you can then put on your resume and say, “Hey, I built this 12-hour project in Salesforce Trailhead.” And some of them are required for certifications. So, you can say, “I did this. I got this certification, and I can actually showcase my skills and what I’ve been working on.”

So, it really makes a good entrance to the ecosystem. Because there’s a lot of people who want to break into tech that don’t necessarily have that background that are able to do so and really, really shine. And I tell people, like, let’s see, it’s 2023. Eight years ago, I was a barista. I was doing undergraduate research and working in a coffee shop. And that’s really helped me in my career.

And a lot of people don’t think about this, but the soft skills that you learn in, like, a food service job or a retail job are really helpful for communicating with those internal and external stakeholders, technical and non-technical stakeholders. And if you’ve ever been yelled at by a Karen on a Sunday morning, in a university town on graduation weekend, you can handle any project manager. So, that’s one thing that, like, because there’s so many resources in the ecosystem, there’s so many people with so many varied backgrounds in the ecosystem, it’s a really welcoming place. And there’s not, like… I don’t know, there’s not a lot of, like, degree shaming or school shaming or background shaming that I feel happens in some other tech spaces. You know, I see your face you’re making there. I know you know what I’m talking about. But—[laugh].

Corey: I have an eighth-grade education on paper. My 20s were very interesting. Now, it’s a fun story, but it was very tricky to get past a lot of that bias early on in my career. You’re not wrong.

Evelyn: Absolutely. And like I said, eight years ago, I was a barista. I went to school for chemical engineering. I have an engineering background, I have most of a chemical engineering degree. I just hated it so much.

But getting into Salesforce honestly changed my life because I worked my way up from a call center as an end-user on Salesforce. Being able to say I have worked as a consultant. I have worked as a staff software engineer, I have worked at an ISV partner, which if you don’t know what that is, Salesforce has an app store, kind of like the Google Play Store or the Apple App Store, but purely apps on Salesforce, and it’s called the Salesforce App Exchange. So, if you have Salesforce, you can extend your functionality by adding an app from the App Exchange to if you want to use Salesforce as an ERP, for example, you can add the Certinia app from the App Exchange. And I’ve worked on AppExchange apps before, and now I’m like, making a big kid salary and, like, it’s really, really kind of cool because ten years ago, I didn’t think my life was going to be like this, and I owe it to—I’m going to give my old boss Scott Bell a shout out on this because he hired me, and I’m happy about it, so thank you, Scott for taking a chance and letting me learn Salesforce. Because now I’m on Screaming in the Cloud, which is really cool, so—talking about Salesforce, which is dorky, but it’s really fun.

Corey: If it works, what’s wrong with it?

Evelyn: Exactly.

Corey: There’s a lot to be said for helping people find a path forward. One of the things that I’ve always been taken aback by has been just how much small gestures can mean to people. I mean, I’ve had people thanked me for things I’ve done for them in their career that I don’t even remember because it was, “You introduced me to someone once,” or, “You sat down with me at a conference and talked for 20 minutes about something that then changed the course of my career.” And honestly, I feel like a jerk when I don’t remember some of these things, but it’s a yeah, you asked me my opinion, I’m thrilled to give it to you, but the choices beyond that are yours. It still sticks out, though, that the things I do can have that level of impact for people.

Evelyn: Yeah, absolutely. And that’s one of the things about the Salesforce community is there are so many opportunities to make those potentially life-changing moments for people. You can give back by being a Trailblazer Mentor, you can sign up for Trailblazer Mentorship from any level of your career, from being a basic fresh, green admin to signing up for architecture lessons. And the highest level of certification in Salesforce is the Certified Technical Architect. There’s, like, 300 of them in the world and there are nonprofits that are entirely dedicated to helping marginalized genders and women and black and indigenous people of color to make these milestones and go for the Certified Technical Architect certification.

And there’s lots of opportunities to give back and create those moments for people. And I spoke at Forcelandia last year, and one of the things that I did—it was the Women in Tech breakfast, and we went over my LinkedIn—which is apparently very good, so if you don’t know what to do on LinkedIn, you can look at mine, it’s fine—we went through LinkedIn and your search engine optimization in LinkedIn and how you can do this, and you know, how to get recruiters to look at your LinkedIn profile. And I went through my salary history of, like, this is how much I was making ten years ago, this is how much I’m making now, and this is how much I made at every job on the way. And we went through and did that. And I had, like, ten women come up to me afterwards and say, “I have never heard someone say outright their salary numbers before. And I don’t know what to ask for when I’m in negotiations.”

Corey: It’s such a massive imbalance because all the companies know what other people are making because they get a holistic view. They know what they’re paying across the board. I think a lot of the pay transparency movement has been phenomenal. I’ve been in situations before myself, where my boss walks up to me out of nowhere, and gives me a unsolicited $10,000 raise. It’s, “Wow, thanks.” Followed immediately by, “Wait a minute.”

Evelyn: Mm-hm.

Corey: People generally don’t do that out of the goodness of their hearts. How underpaid, am I? And every time it was, yeah, here’s the $10,000 raise so you don’t go get 30 somewhere else.

Evelyn: Yeah. And that’s one of the things that, like, going into job negotiations, women and people of marginalized genders will apply for jobs that they’re a hundred percent qualified for, which means that they’re not growing in their positions. So, if you’re not kind of reaching when you’re applying for positions, you’re not going to get the salary you need, you’re not going to get that career growth you need, whereas, not to play this card, but like, white men will go in and be, like, “I’ve got 60% of the qualifications. I’m going to ask for this much money.” And then they get it.

And it’s like, why don’t I do that? It’s, you know, societal whatever is pressuring me not to. And being able to talk transparently about that stuff is, like, so important. And these women just, like, went into salary negotiations a couple weeks later, and I had one of them message me and say, like, “Yeah, I asked for the number you said at this conference and I got it.” And I was like, “Yes! congratulations.” Because that is life-changing, especially, like, because so many of us come from non-technical backgrounds in Salesforce, you don’t know how much money you can make in tech until you get it, and it’s absolutely life-changing.

Corey: Yeah, it’s wild to me, but that’s the way it works. I really want to thank you for taking the time to speak with me. If people want to learn more, where’s the best place for them to find you?

Evelyn: So, I am reachable at anothersalesforceblog.com, and evelyn.fyi, E-V-E-L-Y-N dot F-Y-I, which actually just links back to another Salesforce blog, which is fine. But I’m really [laugh] reachable on LinkedIn and really active there, so if you need any Salesforce mentorship, I do that. And I love doing it because so many people have helped me in my career that it’s really, like, anything I can do to give back. And that’s really kind of the attitude of the Salesforce ecosystem, so definitely feel free to reach out.

Corey: And we will, of course, put links to that in the [show notes 00:30:27]. Thank you so much for taking the time to, I guess, explain how an entire swath of the ecosystem views the world.

Evelyn: Yeah, absolutely. Thank you for having me, Corey.

Corey: Evelyn Grizzle, Senior Salesforce Developer. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry, insulting comment that I will one day aggregate somewhere, undoubtedly within Salesforce.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

View Details

Ricardo Gonzalez, Senior Principal Product Manager at Oracle, joins Corey on Screaming in the Cloud to discuss his approach to Product Management and cloud migration. Ricardo explains how a chance conversation landed him a role at Oracle, and why he feels it’s so important to always bring your A-game in any conversation. Corey and Ricardo discuss why being a good Product Manager involves empathy for your customers and being able to speak their language as well as the language of your product and development team. Ricardo also explains how he’s seen the Oracle product suite grow, and why he feels more and more companies are seeing the value of migrating their data to the cloud.

About Ricardo

Ricardo is a Product Manager at Oracle, in charge of Database Migration to the Cloud, and the ZDM and ACFS products.

Ricardo is a native Costa Rican and has lived in Mexico, Italy and currently resides in the United States.

He is passionate about technology, education, photography, music and cooking. He loves languages and connecting with people from all over the world. In a future life, Ricardo wants to own a taco truck, and share taco happiness with everybody.

Links Referenced:

  • Oracle: https://www.oracle.com/
  • LinkedIn: https://www.linkedin.com/in/ricardogonzaleza/
  • Twitter: https://twitter.com/productmanaged

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Human-scale teams use Tailscale to build trusted networks. Tailscale Funnel is a great way to share a local service with your team for collaboration, testing, and experimentation. Funnel securely exposes your dev environment at a stable URL, complete with auto-provisioned TLS certificates. Use it from the command line or the new VS Code extensions. In a few keystrokes, you can securely expose a local port to the internet, right from the IDE.

I did this in a talk I gave at Tailscale Up, their first inaugural developer conference. I used it to present my slides and only revealed that that’s what I was doing at the end of it. It’s awesome, it works! Check it out!

Their free plan now includes 3 users & 100 devices. Try it at snark.cloud/tailscalescream

Corey: Welcome to Screaming in the Cloud, I’m Corey Quinn. Some wit once said that 90% of life was just showing up. And I’m not going to suggest that today’s guest has only the fact that he shows up going for him, but I do want to say that when I first met him, it was at a drink-up that I threw here in San Francisco. And he kept turning up to a variety of community events, not just ones that I wound up hosting, but other people, too. By day, Ricardo Gonzalez is a Senior Principal Product Manager at Oracle. But in the community, he is also much more. Ricardo, thank you for joining me today.

Ricardo: Thank you so much for having me, Corey. It’s a great pleasure to be here with you today.

Corey: So, it is interesting watching you come to what I can only describe as the other side of the tracks. Because you work at Oracle. I make fun of AWS all the time, so yes, I suppose our companies do have that in common, but I digress. You also work in the database world, which is, I guess you could say I do and that I misuse things as databases, mostly for laughs and occasionally for production. And you’re over in the product manager side of the world, which for me, has always—may as well be a language that I do not understand, let alone speak. Yet you have consistently shown up and made great contributions to every conversation you’ve ever been a part in. Where did you come from? How did you start doing this?

Ricardo: Well, I’m originally from Costa Rica, right, which is I wouldn’t say uncommon, but then again, there’s just a few of us. And I was doing my master’s degree in Mexico when I showed up to a recruitment event dressed up like a business student and realize all of my peers were actually developers—although I’m a computer scientist by trade—looking for a job at Oracle’s Development Center in Mexico, right? And by showing up, something magical happened. I stayed at the session, they made a raffle with numbers. I didn’t win, but they asked us questions nobody answered, and as you can see, I talk a lot.

I raised my hand, and they said, “Okay, answer these questions.” And then it became, like, a competition, and I won. And back then I got, like, a tablet. I think it was an iPad; it was great. I thought, okay, no job for me because I wasn’t working—looking for a job in development.

And then this person, which is now an SVP in my company, which has been my mentor in many ways, approached me, and he said, “I really liked what you did. It seemed you do have some technical background. We need somebody that can talk like that with customers, but at the same time, understand the requirements for a technical product and work with engineers. Do you want to come to the office tomorrow?” And a week later, I got an offer my life change in ways, like, we’ve never foreseen.

Corey: This is a hard thing to talk about because it’s the way the world works, but when you say it, people love to come back and tear you down, like, “You just got lucky.” Or it—“Well, yeah, that works for you, but it doesn’t work for other people.” But I’ve found invariably that the seminal moments that happened in the course of my career have all come from conversations I had with people I didn’t need to be talking to at events I didn’t need to be attending, but one thing leads to another. Instead of sitting at home and brooding, I put myself in situations where I could, for lack of a better term, make my own luck. Sure, if only one conversation in a thousand winds up turning into a career opportunity, okay, but that means you need to have a thousand conversations to get there, so time to get started. And you are probably one of the best living embodiments of this that I’ve ever met.

Ricardo: Well, it’s interesting. You’re right. I mean, the luck part plays a factor, I guess, but you have to change your own luck. And it’s complicated to talk about that because there’s also privilege in both, and being part of—like, I was in college. I had the privilege to go to college, although, I mean, there’s a whole, like, list of things that made me get there and the sacrifices from family, et cetera.

And not everybody has the same level of field, right? But what I can say though, is that I heard somebody said something that really resonated with me, which is, “For some of us, right, we won’t be the main player in the game.” [reading 00:04:25], like, so imagine you have, like, a sports event where—whatever sport you want—and there’s a game of playing, right? The coach will not call you. But they might call you over the last five minutes, but when they do, you have to be there and score a goal, touchdown, whatever you want to call it, be the best player because that’s the opportunity you have and you have to make the most out of it. Some people were born and they have the opportunity to be in the starting lineup. Some of us will be just called at the last minute. But when you do, your A-Game has to be there on top and you have to be the best you can because that’s the only way you have to shine.

Corey: I think that you’re right. There’s a tremendous amount of privilege baked into all of this. And privilege is one of those things you can’t just set aside. It’s something that we wind up all manifesting in different ways to different degrees. But it’s a, “Oh, just be like me,” is fundamentally what a lot of advice comes down to, regardless of whoever it is the me in question that’s talking about it.

But there seems to be just certain things that lend themselves to better possibilities of success. One of the things that has always impressed me is that you just show up and start great conversations with people, left and right. That’s a skill that I honestly wish I had. I have to be noisy and public to get people to approach me, whereas you, ah, you don’t have the time for that. You just walk up and start talking to them. I’ve never been good at that.

Ricardo: I guess part of my upbringing—also, you know, my home country has a whole history of [horizontalness 00:05:47], but that’s a different discussion. And we are, I guess, not shy to just talk to people, right, which sometimes can bring into interesting conversations with management and, like—because if I disagree, I will let you know, right? I will be completely candid about things. But I think it’s important, right? Because like, we’re all human beings trying to do the same thing, right?

We all wake up in the morning with the same set of problems and then get to share moments in between each other. Why don’t make them as pleasant as possible and try to see how can we actually grow together? It’s important that you’re not only getting things and growing yourself but also see how can with that help others grow as well, right? So that’s, I think, part of what conversations can be—I mean, starting conversation with anybody just it’s really important to say, “Okay, nice to meet you. How can we, you know, make the most out of it for both of us?” And, you know, either even if it’s just, like, you’d have a great conversation or, you know, help each other or just me help you, et cetera.

Corey: So, I want to talk a little bit about your day job. Given that you work in product management, I have to assume that having people skills is kind of a prerequisite for the role. At least you would think. I’ve worked in places where that was apparently not the case, and not for nothing, it kind of showed.

Ricardo: Yeah, I mean, it’s really important. I think product management is one of these positions in which you are in the middle of things. When people ask me—and these are people that don’t work in technology—“What do you do?” I tell them, “I’m a translator,” right? And when they ask me, like, “Oh, so you do it between languages?” I said, “Well, yeah, I speak different languages with us.”

So, the point is, I am able to talk with people that have a less technical acumen or are actually just users of our product, right, and [unintelligible 00:07:17] highly skilled, and then go back to the engineers, which have a different point of view, right? So, I’m always back and forth. But that people skills, as you mentioned, is really important because otherwise you cannot do your job. The thing that is interesting for me is that product management itself is not really a thing that can be defined. I mean, yes, of course, there’s, like, books on it and people that have done their careers and, like, saying how it works, but it changes from company to company.

And even within the same company, there are different product managers doing different things. What I do—and I’ve been really fortunate to have really good managers that I’ve worked for the last seven years, I think—has a lot to do with the people skills that you mentioned, right? And it allows me to be as good as I can with my job and try to do me just, you know, grow every day.

Corey: It’s easier to sit here and reason about these things in the context of specifics, on some level. And it’s also easy for me at least to look at a company and think, “Oh, they do one thing,” but I have it on good authority that Oracle is a large-ish company that might have more than one product at any given point in time. What product do you work with? Where do you start and where do you stop?

Ricardo: Okay, well, I’ve been part of three different teams, if you could call it that way. Although, like, over the last seven years, I’ve been focusing mostly for—I mean, always within the database organization, so like database development. And then over the last, like, six, seven years, I’ve been on the high-availability team, which focus on a thing called maximum availability architecture, right, which is basically helping customers to achieve all their requirements. And we’re talking about, like, heavy usage of, you know, regular Oracle database with high availability, scalability, I mean, requirements for, like, 24/7, like, great uptime.

And I started working with them with the cluster file system, which I still do, but my main job over the last, like, let’s say, four, almost five years, have been working towards helping customers come to the cloud, right, to Oracle Cloud. And my product, I’m the product manager for protocols ZDM, Zero Downtime Migration, and it’s been in the market for the last three-and-a-half years, right? So, I was there since it’s all started as a whole interesting story about cross-work with different teams in Oracle getting together to get this product out. So, that’s my day-to-day job, just enabling customers on maximize the usage of the Oracle database in the high-availability realm, and also helping them move to the cloud, the Oracle Cloud, if that’s what they want and the mission they have right now in their organizations.

Corey: I know that people are going to have opinions about Oracle Cloud, and I’m just going to say something that I think is relatively uncontroversial, in that the technology is freaking solid. I have used it in a bunch of different ways, I’ve talked to folks who have, and there is remarkably little argument that when you use it as directed, that stuff works. And there’s a lot to be said for that. So, you focus a lot on the migration story, specifically, to my understanding, databases inward from a variety of other places. Do they tend to find themselves living in on-prem environments? Are they in other cloud providers? Are they, God forbid, well, we have this filing cabinet full of paperwork and we’re hoping you can help us digitize it all, which, yes, those projects exist. And no, I don’t want to be within 6000 miles of them.

Ricardo: Well, mostly, we’re talking about on-premises customers, right, that have large fleets of Oracle databases and we’re trying to help these customers, either as small businesses, it could be public or enterprises move to the Oracle Cloud when they deem that the strategy they’re doing, right? So, my product, what it does is it actually orchestrates, it automates that process for them so that when they’re actually doing the migration, it’s as seamless as possible for them. Because there’s a lot of, like, caveats and a lot of things to consider when we’re talking about database migration into the cloud.

Corey: When you take a look at what is going on in the larger ecosystem, it’s easy for me to sit here and say, “Well, I don’t see Oracle databases very often.” And yeah, in the context of companies that I work with, that are very often founded in the last few years and are born in a particular cloud provider—in my case, AWS—yeah, there doesn’t seem to be a lot of those things. But at the same time, Oracle rose to its current position by having database technology that was second to none. There’s a reason that all of these quote-unquote, “Legacy companies,” by which of course, we mean, companies that made money and had the temerity to be founded more than three years ago, have wound up standardizing across Oracle to a large extent. As a result then, we’re seeing a stupendous amount of those companies now looking and weighing, what does moving into the cloud actually look like because we have an increasingly dire raccoon problem in our data center?

Ricardo: Yeah, I mean, we have all the latest technology over the last 40 years. Like, Oracle, as you mentioned, right? It has impressive technology and it’s quite solid. Now, you’re asking me about companies that, you know, that might not be using Oracle or that you’re not aware of they’re using Oracle. The interesting thing, and when people asked me about this, right, is that it’s really easy, both me and you without knowing, use Oracle products today, right?

Because you check your bank account, you use certain financial services, you made phone calls, et cetera, right? And a lot of the underlying technology and infrastructure that runs the world today—either you took a plane, et cetera—is running on Oracle, right? There’s a lot of deployments there, right? It’s just that is not that maybe we’re not doing—you know, again, we’re talking about the whole ecosystem that runs a lot of infrastructure that normal people would do on a daily basis, but it’s right on the back end, so you might not hear about it, or it’s not as known, but it is there in the top companies all over the world. So, what we’re doing now is helping these companies, right, migrate to the cloud when their needs really adapt to exactly that goal.

And sometimes it’s actually more, “Okay, how can we actually modernize your data center, right?” So, Oracle actually has Cloud@Customer, and we also help them with that migration as well. So, we have a whole set of products and deployments that would work within the customer data center, but within a cloud managed by Oracle.

Corey: I think that that’s an interesting question in its own right, which is you have these companies that are doing incredibly important things. Like this, like Oracle databases, run hospitals, they run DMVs—

Ricardo: Yep.

Corey: At various states. They run basically everything big infrastructure that you can imagine a lot of places. They run banks, for example. And now these companies are looking at transforming into a cloud approach, on some level. How on earth you convince them to move something as critical as a workload on Oracle database, which in many cases, is a bedrock layer upon which aspects of society depend, to, “Oh, yeah, just go ahead and move it to this cloud thing. That’ll be fine.” It feels like an almost impossible goal, but it’s clearly not. What drives it.

Ricardo: Well, it’s happening all over the industry, right? People are realizing that cloud, it’s—I wouldn’t dare to say the future because it’s been happening for, you know, over the last years, but clearly for cost management, security, administration, resource scaling, you name it, it’s the way to go, right? So, it takes time, and depending on who you’re working with the projects could, you know, span, three years, et cetera. But people, that’s the way you like, you know, the whole ecosystem is going, right? So, what we’re doing is, and we didn’t reinvent the wheel here, at least with my product, right, was to take technology has been used for over 40 years as a standard for, you know, backup, export, data transfer, synchronization, security, database management, and integrate it into a single product that would be, like, automated and helping the customers.

And what we wanted to do, and it was really important for me is, like, we want you to be in our cloud, we want to help you, so let’s make this free. Even if we’re using other products that Oracle already has that have a cost, if you’re using the migration suite that we offer, it will not cost you money.

Corey: There’s a lot of value to being able to make assurances like that but, on some level, that feels like whatever someone migrates anything anywhere, but a few things are certain. One is that there’s going to be technical challenges with it. There always are. That is the nature of large systems, particularly systems built upon systems built upon systems. And too, as humans, as much as we love talking about the idea of blamelessness, everyone’s going to be looking for a scapegoat when something inherently goes wrong.

The database is always an easy thing to blame, and the cloud, aha, that’s stuff where it’s non-deterministic and we can’t go and put our meaty hands on it in the data center the way we used to when things start breaking. How do you avoid becoming the blame center in a scenario like that?

Ricardo: That’s a great question and it’s interesting because it could happen, right, that somebody says, “Well, because of the migration, things are not working as expected,” et cetera. So, we do help customer—there is a lot of implications when you’re talking about migration, right—to the proper planning, sizing, are there any architecture implications? Are you doing any cross-endianness? Then, you know, database-wise, Oracle has different architectures, so we have the previous model of non-containerized or no-container databases. Now, we’re going to tenant-based.

We’re working—are you doing an upgrade as well? Are you doing, you know, you’re coming from an older version to a newer version? Are there security implications? Because a lot of the database is on-premises might not have encryption, and we by default encrypt at the target level because it’s a requirement in the cloud, right?

So, what we work with the customers is two things. First of all, do all the planning and testing as possible before the migration so that you know what you’re doing is correct. Is the app certified with newer version and the environment you’re going into, right? And we can work with you to do all these tests. And then one thing that we realized was really important in the product is to have a way to have, like, knobs or control of what you’re doing, and you could actually do testing before the actual switch over into the cloud.

So, you will have, like, a standby database, like, a copy of your database, running in the cloud, [unintelligible 00:16:47] synchronization with your on-prem, your database, right, on your application, but you can use that to just do all the testing you want and then be sure. And only when you’re ready, then you will do a switchover, and then things would work as expected, right? But again, there’s a lot of process. And we’ve worked with customers that you know, they know what they’re doing, they were, like, super happy and they did it quite fast. There’s others that said, “You know what? I am going to do a nine-month testing process because my week that I’m going to be migrating and then the weekend that I’m going to do the switchover is crucial.” And then we work with them over those nine months. But then when it happened, it went, you know, perfectly, right? So, it really depends on the project. But we do ensure that everything is taken care of because as you mentioned, it’s a big change, it’s the big shift.

Corey: Tired of wrestling with Apache Kafka's complexity and cost? Feel like you're stuck in a Kafka novel, but with more latency spikes and less existential dread by at least 10%? You're not alone.

What if there was a way to 10x your streaming data performance without having to rob a bank? Enter Redpanda. It's not just another Kafka wannabe. Redpanda powers mission-critical workloads without making your AWS bill look like a phone number.

And with full Kafka API compatibility, migration is smoother than a fresh jar of peanut butter. Imagine cutting as much as 50% off your AWS bills. With Redpanda, it's not a pipedream, it's reality.

Visit go.redpanda.com/duckbill today. Redpanda: Because your data infrastructure shouldn’t give you Kafkaesque nightmares.

Corey: I think that there’s a very true story about how oh, we just try to close our eyes and cross our fingers and hope for the best and press the migrate button that everything will work out flawlessly. It doesn’t work that way. The way that we always wound up handling migrations in places that weren’t riddled with dysfunction up, down, and sideways—at least not in this particular way because everyone’s environment’s terrible—is that we would test these things out, we’d stage them, we would have rollbacks that were tested and known to work. In some cases, we’d begin with the rollback before we started the migration plan, just because we absolutely cannot have this system down outside of a maintenance window or outside of certain constraints. And it feels like a lot of that planning is wasted when things go well. But it’s not. It’s the reason that important things don’t crumble underneath us. Like, on some level it’s, do I feel like I wasted money on my airbags and seatbelts because I’ve never used them? Not really no.

Ricardo: Well, I mean, this is, like, the classic [unintelligible 00:18:23] ops thing and support thing, right? People always complain when things don’t work, but when they do work, they don’t realize it because of all the worries that all the people that infrastructure and planning and support and ops were doing, right? So it’s—yeah, there’s a lot of time that can be spent in planning and people would think that it’s actually wasted time, that actually is super important and crucial for these. The other thing I think it’s important is that you always should have a fallback plan. There’s, like, different configurations in which these might be more cumbersome or complex, but we do have the possibility to, like, keep replicating back to on-premises, so that if anything happens, people do have that option, right?

And we do have customers that like the idea of having a disaster recovery configuration in which they have, like, something in the cloud and then another thing on-premises, so there’s always an option for you. But planning is crucial, right? So, we even have a thing called, like, evaluation mode in which we could we dry run a migration without actually doing it, just to tell you what could happen. Of course, when you do things live, there’s always things, right, that could be related to many other factors, but we really, really try to dial in and be sure that when you’re doing the migration and you properly planned, things will be automated and work for you.

And so, we’ve grown over the last three-and-a-half years, and I was doing some research, right, and we’ve had, like, you know, thousands of databases migrated, great customer that have been using us, and surprises, so sometimes we don’t know, right, and we find out, oh, somebody’s doing a course in one of these learning platforms based on our product, which is really new, but it’s, oh, it’s cool. Like, when we’re not creating, like, even your [unintelligible 00:19:50] et cetera, right? And I’m really glad that what your doing has an impact and helps people. That’s all you want. You want to help people achieve their goals.

Corey: So, I have to ask. On some level, building something that migrates a database from one location to another naively would seem to folks to be a, “Okay, at some point, this gets declared feature-complete and then we go work on other interesting problems.” But yet the fact that you’ve remained employed in the role that you’re in where you continue to work on the problem would strongly suggest that this is not, in fact, true. How does the product continue to evolve once you are, let’s be clear, shipping this to paying customers?

Ricardo: Well, I mean, the product will evolve, as you mentioned, right—

Corey: And I want to be clear, that’s not just a rephrasing of, “Hey, quick, justify your job.” Obviously, this stuff has to evolve. This is not one of those, “So, what is it? You’d say these do here,” crappy questions that isn’t really a question so much as an accusation. Those come in a slightly different tone of voice.

Ricardo: That’s, you know, it’s a super valid question and I actually appreciate it a lot because it also makes me reflect on how much we’ve grown right? I mean, I think the magic of ZDM and the team behind it is that it’s kind of like a startup within Oracle, right? It all started because different teams [within 00:21:03] Oracle, right, you’re banded together, a team propose a prototype based on existing technology, right? So like, again, as I mentioned, like, Oracle technology for database has been over 40 years in the making. And, you know, a team said, “Okay, what are the standard tools to actually do a backup or an export of data transferred, you know, to a location”—in this case, the cloud—“Doing a whole synchronization, encryption, et cetera, and then the switchover?” Right?

So, the thing is that databases come in many flavors, there are different options, different ways for databases to work. There’s also different targets in the Oracle Cloud and those then change how you would be migrating into, you will have different workflows, physical, logical, you could use different backup locations, so of course, in Oracle Cloud, the standard is the object storage, right? You can do a direct data transfer; you have that technology as well. If you’re doing migrations to [cloud 00:21:51] customer, you’ll definitely will require, like, external storage, like NFS. If you’re doing a conversion from AIX or Solaris into, you know, the cloud target which is Linux, then again, there’s other implications, if you’re doing an inflight upgrade, if you’re changing architectures, from non-multi-tenancy to multi-tenancy, if you’re doing, you know, coming from other clouds, there are also certain considerations.

So, now that I mentioned all of these, you can see how a product from the get-go can have all those, right? So, we started with a subset of features and we’ve grown up to six releases now over the last three-and-a-half years that have incorporated everything that I’ve just mentioned. And we can do all those things, but it keeps getting better. And then there’s always, like, things that we realize that customers are using us in ways that maybe were not expected, which is great because oh, okay, cool, then this is something that we can actually, like, make better or enhance, right?

And there’s always requests from customers on what they want to do or see change in the product. We also integrate with our team. So, there is an advisor that does a pre-check for the database and checks, okay, what are, like, the recommendations on what you should do? So, those integrations and working with our teams across Oracle, again, take time, and hence why, you know, products keep growing and evolving. And you’re right, maybe at some point, we will be able to cover everything that there is to do, right?

What we’re doing now, and we’ve been working, again, in partnership with our teams at Oracle, right, is, like, be the engine of other Oracle migration strategies. So, there is a native service in Oracle Cloud infrastructure called DMS that has a subset of our features and it uses ZDM under the hood, right? So again, there’s always work to do and a lot of it sometimes is go to communication and working with customers, but there’s also a lot of, like, going back to the drawing board and see how can the product be improved.

Corey: I think that there’s a certain lack of attention also given to the fact that every time you think you’ve seen it all, all it takes is talking to one more customer, where they have a use case that you potentially hadn’t considered. And maybe it sounds ridiculous to you, but it’s ridiculous in load-bearing ways in an awful lot of these other places. Empathy becomes such a key aspect of this that I’m somewhat surprised that more folks don’t spend more time than they do thinking about these things.

Ricardo: Well, I think as a product manager, this is really important, right? You need to put yourself in the customer’s shoes. And you also need to use the product. Sometimes using the product, like, so I use it, like, to create my own workshops that we have. There’s a platform called Live Labs in Oracle that has, like, I don’t like 6000 labs that are free for you to use and learn about our technology, right?

So, in order to better the product understand, and then you know, when we’re doing a new release, et cetera, then see the key features, like, we create materials like that and we use it. But that doesn’t give us the whole scope of how a product customer would be using it. So, for all internal migration that we have within Oracle products into the Oracle Cloud, we use that and then that gives us a lot of insights. But then going to a customer and spending time with them, sometimes developing relationships that go more than a year because we were talking about, like, big [fleet 00:24:41] migrations, thousands of databases, you realize, oh, the scope is broader than we expect, but it’s actually a really—there’s a lot of satisfaction in learning from them and then getting back to the development team. Or even including. I think that’s really important as well.

I think a good PM would include development sometimes in the conversation with customers because they then—there is, like, a better understanding from both sides of the aisle. And even bring them to conferences, et cetera, so that the actual, you know, empathy of the customer requests and what they need, it’s created.

Corey: Yeah, I think that there’s also a presupposition that you can look at a company and say, “Oh, you’re using X technology? You must be crappy,” or whatnot. Something I’ve learned is that every company of a size that is remarkably small compared to what people often think he is using basically everything already. Like, I’m at this point at a company that has less than ten employees and we already have five different clouds that we have accounts with, doing different things in different ways. This explosion of different tools and different utilities is like it is for a reason. And it’s very tricky to really, I think, appreciate that until you’ve walked a mile in the shoes of someone who’s building things like that.

Ricardo: Yeah. It’s interesting. There’s a whole, like, view of product management, right, and having this idea of building and building and building products, but what you’re doing is actually helping people with their needs, and their needs can be really broad, so maybe the solution is not your product. And maybe the solution is not your technology. But I think good PM, and I think anyone in technology, a good person, would actually, like, help these users or customers to get where they need to, even if it’s not using your technology, right?

Corey: I would agree wholeheartedly. I really want to thank you for taking the time to go through what it is you’re up to and how you view the world. If people want to learn more, where’s the best place for them to find you other than, you know, local community meetups when you happen to be in town?

Ricardo: Well, I mean, of course, anybody can, like, go to LinkedIn and look me there. I have a Twitter account @productmanaged, so product manager, but without the R and a D instead because of course. Twitter handles are—or handles over on social media are hard to get, although I’m as active lately on Twitter. And I, you know, I opened an account on Bluesky, which is [@productmanager 00:26:54]. I did get that one. But um, I only starting now to use it, right, so, you know, I guess those three would be the places to.

Corey: Awesome. And we will, of course, put links to that in the [show notes 00:27:04]. Thank you so much for your time, I appreciate it.

Ricardo: Anytime. And one thing. If anyone is ever in San Francisco, you know, I’m more than happy to meet up. I love this city. It has changed my life tremendously and I’m happy to show you around. I consider myself now somebody that really, really, really cares for this place and happy to just, you know, have a good time, talk technology or not. I also love to cook. So anytime, I’m here.

Corey: I highly recommend that. He’s not just fun to hang out with, he is an excellent cook as well. But I don’t know if there's a good way to put that in show notes, so you’ll have to take my word [laugh] for it instead.

Ricardo: [laugh].

Corey: Ricardo Gonzalez, Senior Principal Product Manager at Oracle. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an angry, insulting comment that one day I will find a tool to migrate into a central database. I know not where.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

View Details

Salman Paracha, Founder & CEO at Katanemo Labs, joins Corey at Screaming in the Cloud to discuss his vision for the future of SaaS application development. Salman and Corey discuss what led him to take the leap into founding a start-up, and Salman shares how he believes the future of SaaS application development is at an inflection point. Salman also explains why it’s critical to focus on the outcome your customers experience over infrastructure, and shares his vision for future developers looking to build the next wave of SaaS applications.

About Salman

Building high-growth, high-tech software products that affect the lives of millions of customers. 15+ years of experience in building successful products and highly effective teams. I am deeply interested in bringing the power of the cloud to end customers, large scale data problems, and delivering scalable services on commodity hardware.

Links Referenced:

  • Katanemo: https://www.katanemo.com/
  • LinkedIn: https://www.linkedin.com/in/salmanparacha/
  • Email: mailto:salman@katanemo.com
  • Twitter: https://twitter.com/salman_paracha

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. And this promoted guest episode of Screaming in the Cloud is brought to us by our friends at Katanemo, who is—when you talk to small startups, like, “Who should we talk to?” They invariably look around the room, figure out who they should throw directly into the grist mill, and in this particular case, they have selected Salman Paracha, who is the founder and CEO. Salman, thank you for joining me.

Salman: Hey, thanks for having me. Second time.

Corey: It is. And every time we talk, it seems like there has been an interesting progression in your career. Originally, when we first started talking, you were the GM of the serverless application repository at AWS, and of AWS SAM, the Serverless Application Model that most people know because of the giant psychotic squirrel running around the expo hall at events. Then you went to be a group VP at Oracle Cloud, and now you look around the landscape and decide, you know, what I’ve done my entire career? Worked at big companies where everything is, you know, convenient in certain ways. And that sucks. I want everything to be three times harder, at least, so I’m going to go start a startup of my own. Presumably. I’m assuming that is the thought process that led you here. What’s the actual story behind why you decided to leave giant corporate entities and go to a small startup?

Salman: Thanks for that intro. The primary reason to sort of pursue this dream was something was pulling at me for the past four to five years. As a person who considers himself a builder, the most happiest I am when I’m actually trying to ship software out for customers. And so, I’ve been pulling on this thread for a very, very long time, that the world of the modern reference architecture, as it goes more microservices and explodes in the face of developers, has gotten to a point that we are now being inundated with all these micro-primitives, if you would, on infrastructure that actually slow the rate of innovation down. And why hasn’t there been a move and reversion to the other side?

And so, as I looked around, and at my time at Serverless particularly, where we were trying to champion this idea of serverless compute where you don’t manage your servers, I kind of was ruminating on this notion of how do you get to zero infrastructure? And the idea that how can we actually orchestrate out all the complexity behind the scenes and you can truly focus on what your application does. And in that part of the journey, I’ve been chatting with developers across swath of industries and varying degrees of sophistication if you would, and the thing that emerged is that the most complex, perhaps the most complex piece to build in the cloud is a SaaS application. And there’s inherent complexity in sort of thinking through the various concerns of the shapes and sizes of your customers that you’re serving, the security and safety controls that you have to give them, the operational burdens that you take to serve a very large customer versus a very small customer who is perhaps in your free tier.

And so, I was pulling on this thread for a very long time, even at my time, somewhat, at AWS, having—I chatted folks like Twilio and Slack at that time, and I said, “I think there has to be a much, much better way.” It cannot be more; it has to be less, and that less is actually getting us closer to what we believe is the future of cloud infrastructure, which is no infrastructure. So, that’s it. I mean, I think the core thesis was, “Hey, if I’ve operated at this intersection of hyperscale cloud infrastructure and SaaS applications for the past 20 years, what is the compression algorithm that I can apply and give to developers so that they can truly focus on building something phenomenal without having to worry about the complexity of the infrastructure, the security of the scaling of the operational, and the access logs, and all that stuff that they have to today focus on?” And then I’m very fortunate to have had a phenomenal team that have joined and humbled me in my journey here.

Since last year, we have folks across the spectrum who have built these things at scale and at Lyft, at Dropbox, at Meta, at AWS, at Cloudera, and et cetera. And so, we’ve been really fortunate that we have a very firm belief of where we want to take the future of infrastructure and who we want to serve in that market segment. And I said to myself, I don’t think I’m getting any younger. My parents, my South Asian parents, perhaps they’re going to be more happy to see me sort of fight it out and battle it out versus just naturally climb the corporate ladder. Nothing wrong with that, of course.

Corey: It’s not too late to go be a doctor. I say that as someone who grew up in a Jewish home where there were certain expectations and pressures placed upon me that I continue to disappoint four decades later.

Salman: Yeah, so anywho [laugh], on that front, so I think I’m kind of living to the expectations I had for myself 20 years ago when I joined the workforce, and I now have the great fortune to build alongside these amazing builders and see what we can unlock for the developer community.

Corey: One of the challenges with the approach that I found historically has been Heroku did something very similar and then everyone tried to build the next Heroku, except for the company that bought Heroku, they were content to let that thing sit and never think about it again, for whatever reason. But another example would be something like NPM, the Node Package Manager, where it abstracts away stupendous complexity. You tell it to npm install for some project and it just starts scrolling huge amounts of text past and doing all kinds of work and your computer fans start screaming, and you’re like, “Wow, it’s doing an awful lot of fascinating stuff underneath the hood, and I really hope this works. If it breaks halfway through, I haven’t first idea where to look under the hood to make sure that this actually works and doesn’t break my application.”

The problem that I have historically with the things in this space is it requires a certain element of trust. That said, looking at the things you’ve done before, the places you’ve been, I don’t have to explain that to you. You have clearly spent your entire career in environments where mistakes matter because they’re going to show very quickly to an awful lot of customers if they wind up getting out there. That feels to me like it’s a significant competitive advantage versus, not to be disparaging, but a couple of founders fresh out of a boot camp who have never worked in the industry before, but they have an idea, gosh darn it, this is what they’re going to build.

Salman: You know, you’ll find builders, and you’ll have builders surprise you. And I, you know, salute all those who come out and start something new. I have a whole bunch of respect for that, just the courage that takes it. But there’s an advantage that the team has, and we’re very fortunate on having that advantage, having seen things break. And I think we’re at this inflection point, perhaps now that there’s been an incredible amount of effort done in the open-source community relative to [dis-established 00:06:56] standards.

Like if you imagine, what, 25+ years ago, when HTTP and HTTP 1.1 came out, that created an explosion of people hosting these web services and HTTP-based applications. I think we’re at the point where we can preserve the developer experience, preserve the operator experience, but never have to sort of have you tinker in the bowels of the infrastructure … to build a SaaS application. And I think that the interesting part of this is knowing how successful these projects can be, but also how complex they can be to manage. But if you (the developer) can just focus on dev experience and operator experience and ask what’s the most pressing question to answer, which is…Can you know who (your customers) are and what they’re doing in your system, and have the ability to shape their experience versus shaping the infrastructure?”

I think we’ll be in a much better state as an industry, we’ll be much happier developers, we’ll just be in a much higher place than we are today. Where as I said earlier, which is the modern reference architecture of microservices perhaps gives you some powers, but it really explodes the amount of choices and results in this massive drag on innovation. And that’s that part of the lessons and learnings and insights that we have and we’re going to compress that, hopefully, on behalf of developers as we build out Katanemo, particularly, you know, going towards this future of no infrastructure, zero infrastructure. So yeah, all respect to everyone who’s building. You know, we’ve had the good fortune and we hope to pass that fortune back in terms of a product experience.

Corey: This feels like a problem that never really goes away, at any scale, for that matter. I want to build out something new. Maybe it’s just a ridiculous static site. Maybe it’s some serverless-powered shitposting app. I have several of those in existence.

And every time it’s like, “Oh, you have a great idea for an application. Cool. Step one: do a whole bunch of infrastructure provisioning nonsense along the way first because that’s going to be the important thing to get done.” And then, only then, do you get to start getting into the application logic and the rest. And it always feels like boilerplate, but it’s specific boilerplate, in that it has to be right for this environment with this constraint and this use case, and it just feels like it’s undifferentiated work that I don’t want to be doing.

Salman: I think that actually is magnified to a certain degree when you’re thinking about an enterprise-grade SaaS application. And my impression is it’s magnified of perhaps an order of magnitude more. Because in any modern SaaS experience, you would have to think through the list of concerns relative to your small customer base that’s trying your product, teams that are relying on your experience that their workflows don’t break, or perhaps large enterprises who you’re trying to serve and upsell to. And that inherent complexity then gets baked into the choices on “Hey, should I have more nodes or should I have more concurrency or should I have more isolation boundaries? How do I think about security for multiple customers within my system?”

And I think that’s the really hard nut to crack. And we’re focused there first because we believe we can serve that community really well, get off on the get-go, and then create the right level of experiences, perhaps for general business-to-consumer applications as well. But this problem, I think, it’s magnified even more for the [unintelligible 00:10:01] dot community that’s trying to start off with a developer-led motion but naturally wants to upsell to teams, organizations, and enterprises with their suite of services, perhaps a next-generation ChatGPT, if you would. So yeah, I’m with you. I hear you, and I think the problems amplified, in our view, to that other community that sort of struggles with this and has to hire specific talent to build that stack out.

Corey: I have to ask, because I alluded to, it seems like every company has been trying to build the next version of Heroku, which when you distill down what the value would actually deliver doesn’t sound that far removed from what it is that you’re proposing to build. Hasn’t this been done yet?

Salman: So, I think the way we think about this problem is it’s across multiple layers. And some components to this problem that’s worth talking about. Of course, when you say zero infrastructure and no infrastructure, what does that even mean? Like, I think people naturally get confused. So, three weeks ago, we actually launched what we call our first set of capabilities on behalf of this community as we break things out in components, which is zero-trust capability.

So, if you think about the space, there’s a whole bunch of these undifferentiated essentials that you need to build something meaningful and serve users, teams, organizations, and enterprises. And Heroku is this approach was an abstraction—and a fine one, if you want to build a general purpose app that is just serving the consumer, perhaps. And we’re sort of taking a very different position. We’re saying we’re here to solve you if you’re building something that’s going to serve developers, teams, or organizations. So, we are very different in terms of how we’re approaching the market relative to what we want to go solve for.

That’s just number one. And B, as the thing that we recently launched, is how do we break this problem down on behalf of the community and be targeted to solve a particular problem? So, when I connected with developers in my journey for the past six to nine months as we’ve been in business, is that they felt that the modern state and fragmented nature of identity and access management is really complex for their application. Why? Because now you have this very interesting usage patterns for your applications.

There’s no longer users using something you’re built. There are, of course, as I mentioned, teams, and of course, there’s an enterprise component to this. There are machine keys for your APIs. And all these vectors now of uses all naturally become a threat vector that you have to protect for and they have to be neatly thought through from a access management strategy. And so, what we’ve set out to do is, like, how do we unify this experience today, and solve a real problem, which is you can effortlessly onboard any customer of any size and upsell through zero-trust capabilities like role-based access control, attribute-based access control, and give your customers the ability to achieve least privileged access?

So, meaning how do you safeguard the most protected resources off your SaaS application and make sure they will be safeguarded, but if your users want to create for more sharing and collaboration experiences, you have the means for them to go achieve that without having to build custom logic, custom code, and perhaps spend, many months cycles and perfecting it? And that engineer that built it, and when it left, who’s going to take over and maintain that piece of code?

Corey: Not to mention you’re going to get it wrong, and as a result, mistakes there have security implications that can be dire.

Salman: I think that’s where developers tell us, this is why—you know, I was talking to one potential developer the other day and the thinking was, hey, you know, it was really hard for us to, perhaps, let go of these security controls because we want to build them ourselves. And I asked them this question: “Where do you store your username and passwords for your applications?” Like, “I don’t store them anymore.” Like, I think the reason why people have moved away from having these concerns is because it’s a compliance security risk, it’s a threat vector. And there are others who have hired teams and staff of experts to make sure that thing never breaks, on their behalf.

And similarly, I think as you think about this multimodal identity experiences, this permissions experience that we have built for developers, we are the experts in this domain. We have advisors, past advisors from AWS IAM, perhaps people know that’s a very popular. It serves billions of transactions a second, and securing cloud infrastructure at this rate of $100 billion worth of workloads. And so, we’ve got the expertise to help think through, like, what do developers need to create these safety guardrails, but with a phenomenal developer experience? And I’ll give you an example of that, Corey.

Like, in order for you to, sort of, interact with Katanemo, all you need to do is capture your API surface area in an OpenAPI specification or a GraphQL specification. And that submission of that specification means we know your resources, we know your resource model, your data paths, your access control mechanisms from the HTTP methods that you’re exposing, and then we create the entire identity, customer identity, and finding permissions experience that the developer can expose to their customers in a self-service way to construct their own roles, construct their own SSO, construct their own access log controls, if you would, and just move past this, like, can we get to an enterprise-grade experience instantly as we serve, users, teams as effortlessly with us, and through their business lifecycle. Like, no developer is going to serve necessarily an enterprise on day one; they’re going to get these teams really excited about their product and then they’re going to have an upsell motion. But having to build these by bespoke experiences on onboarding and safety for each different cohort of the customers that they want to serve, that’s just time away from stuff that they can build, cool things that are differentiating for them.

Corey: One of the things has always sucked for me about building applications, even from an infrastructure perspective, has been that I don’t know what I don’t know. And I always feel like I am making a bunch of decisions now that make perfect sense, but when I start scaling or having to take this into a more serious environment, I’m going to have to throw so much of it away and backtrack massively. And oh, I shouldn’t roll my own authentication subsystem and whatnot. But finding the right path forward that matches the current state of the art from an industry perspective really feels like a crapshoot, it’s you’re looking at all the horses, wondering which one you want to bet on and it carries a cost to get it wrong.

Salman: In my time at Serverless, at EC2, even my time at Oracle, the whole idea was to make sure that we reduce this crapshoot behavior on behalf of developers. Of course, at AWS, at Oracle, it was very wide and horizontal in its appeal to any type of developer, but we have felt that if you sort of flipped on its head and go with a verticalized approach, and particularly target one persona and their use cases and their needs, that actually helps us, sort of, look at the problem very holistically and solve that thing just for them. And as I mentioned, we sort of focused on that SaaS use case, particularly, because we believe there’s inherent and unbounded complexity there. So, this is just for playing from the experiences and learnings I’ve had in the past, which is, yeah, this stuff is hard. It’s incredibly hard to get right, and just as, you know, the industry moved to hey, I can trust somebody else who’s an expert here, we’re saying we complete that story. And we look to the modern ways people access your applications through APIs and API keys, or users, or teams, or SSL, whatever, and we compress it, saying single API call to us and you get those capabilities out of the box so you can focus on what matters: moving fast, closing customers even faster.

Corey: I think that is the grail that people are chasing. The problem I found, especially in the enterprise space, has been that it sounds great in theory, but in practice, it’s a oh great, the old Model T story, you can get in any color you want, as long as it’s black. And it’s well, okay, that’s a path, but it doesn’t comport with our security requirements and our guardrails and our compliance objectives, et cetera, et cetera, et cetera. Rightly or—more often—wrongly, people tend to believe that they are bespoke unicorns whose problems could never possibly be realized by anything that wasn’t brewed in-house at their own company. I don’t find that to be true, but I imagine you’re getting a lot of pushback from that direction.

Salman: I think there are two pieces of feedback that we normally hear. “Oh, hey. We built some of this stuff. How do we sort of untangle the mess that we have?” That’s fine. We can help them we have some components that easily wrap around their experience and give them the ability to sort of move to a better state.

But if we build this stuff as a meaningful framework using open standards, like OpenAPI and GraphQL, as the only way you interact with us today, that means that your customers can now build, have a framework in which they set their own security standards against your service, against your application. And I think that makes you getting out of the business of defining the security posture to giving them the ability to construct their security posture is using open standards so their teams can plug down their own SEIM tools if they have to. But you have that framework powering your security and safety experience, your identity and access management experience, without having to build it.

Going back to the earlier thing that we talked about, we believe we’re in an inflection point where standards do establish a lot of innovation, specifically in infrastructure, and we’re going to leverage as much as we can on behalf of developers to bet on those standards. Like I said, OpenAPI, GraphQL, AsyncAPI, so that their customers can say, “Yeah, I get it. I understand your surface area. I can construct these things at least privileged or coarse grained. That’s my choice. You’re going to give me access logs so I know what I did, or who did what, when, and how, so, you know, I can confirm for my compliance requirements.”

And they’re off the hook. They’re actually truly off the hook without having to think about, I think I can do it better because my customers are pi—or second, their customers put these requirements that take them and create [sort of 00:19:29] Rube Goldberg type of scenario in terms of their own stack. So, we think we have something to really serve the market and make it such that it’s not necessarily bespoke.

Corey: I think that you’re probably right that there’s a lot of opportunity to develop those things. I mean, you spent enough time at Amazon, for example, to have benefited from the realization of some of this. One of the nice things I have to imagine, about building a product or a service at AWS is so much of the infrastructure work has already been done. You’re not going to convince me that individual service teams have to sit there and come up with, well, we need to implement a global, highly available block store. S3 already exists. It’s right there. You can use it.

Same with authentication in the form of IAM, et cetera, et cetera, et cetera, a bunch of internal infrastructure stuff that’s there and ready to go. Now, the counterargument, of course, is, as you’re building this out, you don’t have that, I guess, luxury anymore of big company, massive, awesome infrastructure there and ready to go, other than what is available to the rest of us mere mortals. So, I have to ask, is that the big part of what sucks about building SaaS these days or are you finding the friction and challenging parts somewhere else?

Salman: So, it’s a good question because Katanemo is built on Katanemo. It’s a very [mind-tingling 00:20:46] type of discussion, but the one principle that we took is if we’re going to build something on behalf of the community, then our product and service has to consume it as well, and specifically in talking about identity and access management for our SaaS service. Because there’s nothing in the market that neatly solves this problem today. And should we rely on the cloud infrastructure and build on top of AWS and perhaps others in the market like Azure or GCP trying to do? Yeah, absolutely.

We’re not here to reinvent the primitives that are there for low-level infrastructure. We have a very strong non-religious belief that hey, we should leverage what we have, so we can move faster into market. So, we have a whole bunch of usage on, you know, openly speaking, we, when customers ask us, “How do you [unintelligible 00:21:27]? I’m like, “We use KMS for securing some of the things that we do on your behalf.” We have architected around the complexity on [unintelligible 00:21:34] groups and pools and multiples and trying keys and all that stuff. And so, we are trying to use as much as we can, but as I go back to this earlier notion, we’re trying to develop a purpose-built experience that dramatically simplifies for that developer community.

And tomorrow, as we go in towards our [unintelligible 00:21:51] infrastructure future, we will then design something very particular for that next community. And perhaps it’s going to be a gaming community if we want to solve their problems. And that’s going to be the ethos of the company. It has to be purpose-built, it has to be developer experiences phenomenal, not just digging any large cloud provider, but that is a missing component tree and how to think about it, and make sure that we can compress our infrastructure and systems knowledge so that they don’t have to build it. And so, that’s the mission that we’re on. And we’re, of course, very excited about what we’re doing and very fortunate to have both the team and the backing that we have so far to pursue this a little bit further.

Corey: You’re putting your finger right on a very painful spot that has been resonating with me for a long time, which is that it feels like building something on top of AWS natively is a lot like going to the Home Depot and building a cabinet. Well, you go walk up and down the aisles and you pick the exact wood you want, the exact stain for it, the fasteners, et cetera, et cetera, et cetera, whereas sometimes you just want something to store some bowls, so going to Target is going to be the better solution. But now you’re so forced to go and build these things yourself from parts. And that just feels like it has been such a heavy lift for folks because there’s so much you need to understand. And it’s more or less a shipping of AWS’s internal product culture.

But containers, databases, networks, compute, et cetera, are all things that any customer building even a Hello World app has to think about. But that falls across five different talk tracks at re:Invent, for example. It’s too much burden that has been put on the customer and as a result, I think that there’s a lot of value being left on the table. I spend roughly equivalent amounts of money every month on AWS and on Retool. For AWS, I spent about 450 bucks to get about 450 bucks worth of infrastructure services.

Retool, which is basically a WYSIWYG app that designs in-house applications charges me about 400 bucks for which I receive probably about 20 cents worth of infrastructure services, but the value it presents by stringing those things together for me means I am happy to pay it. I really feel like there’s a massive untapped value in being able to deliver not building blocks, but conceived solutions that get out of the way and let people build the differentiated thing that they’re in business to build.

Salman: We feel the same way. I think part of this realization is developers who are building these things continue to stumble upon the explosion of courses and certification material and all that stuff to train themselves to do something. As of course, naturally, AI comes into play and the way that you know the future of applications continues to press upon, you have to build something quickly, you will see that this notion of just [hugging 00:24:32] your primitives or hugging these low-level infrastructure primitives is going to go away because the world is moving at an incredibly breakneck pace. And that will be true, but there is truly now an inflection point where everyone wants to move even faster.

And our talk track with, I guess, our customers is, focus on what really matters to grow your business. And if you are a SaaS developer, or perhaps you’re a gaming developer, or perhaps you’re thinking very specifically in terms of vertical industry that you want to unlock, like, a healthcare company, for example, you should focus on great patient care, you should focus on great gaming experience, you should focus on great X, Y, Z. Don’t focus on infrastructure. Infrastructure is not the outcome. The outcome is your customers are happy and you’re going to serve them.

And your customers are not all equal size, equal shape, and never will be, but you need to give equal shape, equal size, type of price performance or great experience to them. Because you’re not necessarily going to spend the effort to make sure that your free tier is the most highly performant place for you serve your customers and leave your perhaps platinum or enterprise customers hanging dry, as an example. But yeah, I mean, I think that’s the ethos of our company and the spirit of what we are trying to go build. As I said, we’re humbled to be—I am humbled to be surrounded by folks who are much smarter than me and been better builders, and customers who are so excited about our journey. So, this is a good time for us at the moment.

Corey: I understand the grass is always greener when it comes to looking at the road not taken. For me, I see one of the advantages of running a services business as I do, in that, well, I can start a services business on Monday and by you know, Friday or so, I have my first client lined up and I’m ready to start performing work and get paid immediately. SaaS on the other hand feels a lot more like a real estate adjacent, where you have to go ahead and buy the land and get everyone lined up and sink the massive investment into it to get it built up, and you won’t know for years in some cases whether this is something that is going to catch on, much less even justify the cost of building it in the first place. Where are you on that journey as far as validating that you’re building something that’s resonating?

Salman: So, we have design partners, we call them because they’re shaping our product experience. And we don’t call them customers yet, just because we’re in sort of the early stages. But we have designed partners across four critical industries. One of them which is AI, as the booming next-generation AI company is going to be API-first, we have that use case that we can target really well. They’re really early in their days and they need support across their business lifecycle. Hey, I’m just going to support three users tinkering of my product to 3000 customers in an enterprise.

But that’s one. We are very much engaged in the healthcare space because the healthcare is actually going through a very massive legal transformation through—well, what’s happening there’s this HL7 FHIR standard which is actually making healthcare records more interoperable. So, you actually can get patient records if you go from one doctor to the other and not be blocked by the healthcare Gods to say, “No, you cannot do that.” And that is actually creating a very net-new experience in the healthcare space, so we have very customers excited about how we can self-solve their problems in terms of identity and authorization. We have customers in the Web3 off-chain space.

So, on-chain is all permissionless and it’s a whole bunch of different type of development experience, but off-chain has very much of the same characteristics that you will find on a traditional SaaS application. They [need 00:27:56] about safety, you think about privacy, you think about users and teams and API keys and a whole bunch of stuff that sort of baked into it. And the general developer tools who are going from an open-source experience to perhaps a cloud service experience, they’ve got a really great project in the GitHub, they got a bunch of stars and they now have to think about how to provide a better value to customers? And they have to go through a journey.

So, in those four general sort of in buckets is where we are operating right now. We’re very excited about that. And, you know, this opportunity to talk to you is to connect with more folks, especially as we, as I travel in the to AWS New York Summit, or perhaps just meeting up through one-on-ones through Calendly, or whatever have you, and figuring out how we can unlock more value for customers in these use case verticals, or perhaps something that we haven’t necessarily thought through yet.

Corey: I think that one of the clear signs of someone who used to work at Amazon is that—I don’t even have to ask; I already know the answer—of are you talking to prospective customers before you start building things? Whereas start to finish everyone I’ve ever met at AWS is highly focused on the customer experience, whereas when you talk to people building things who have not been through that, a depressing amount of the time, your question is, okay, so what do your prospective customers think about this? Like, “Oh, we haven’t talked to any of those people, yet. Talking to people is scary and we’re here to write code.” It’s, “You might be surprised by what you learn.”

And there’s no immunity to it. When I started this place, I thought I knew pretty well what people thought about their AWS bill, and it turns out, I was way off. There were nuances of the way customers talked about it that I didn’t fully understand. So, to that end, in fact, we can prove it relatively easily. What is something you have learned about your space since you started the company from customer conversations?

Salman: Oh, we actually made a pivot into this space that we are in at the moment because customers told us that’s something that they do not want to focus their efforts on. Repeatedly. We did not write a single line of code all up until November of last year, but once we got the signal from our, as I said, as I mentioned, design partners, they’re like, “This is a problem worth solving.” They’re like, “We’re going to get to work for you. You have these use cases, you have these scenarios that are coming up in your conversations with your customers. Let us be that accelerant for you and be an extension of your team in some ways, so that you can focus on what’s really, really, really important.”

So, you know, I think that’s just survival, Corey. Part, of course—naturally, of course, you work backwards from customers and that was the framework I used when I joined Amazon back in 2012. And even in my time at Oracle, that’s been the ethos of my, I guess, my personal self. But in our case, particularly, we actually talked about a very different idea, we wanted to start, but then customers told us, “You know what? Don’t start there. Start here.”

And I think that’s obviously, just the nature of surviving in through the first few years of your company existence is… getting people to say yes and getting people to say no, and then no, is actually really valuable in many cases because it tells you what to adjust to. And so, we adjusted here as a result of those conversations.

Corey: That may be the best answer to that question I think I’ve ever gotten. That is a phenomenal way to approach things. We started building a SaaS product here and two months later, we sunset the SaaS product because it turned out that what we were building and what customers wanted were not necessarily aligned. I like you said didn’t even write a line of code until last November, just because of the conversations were still shaping what was actually needed in the marketplace. You would be astonished how rare that is.

Salman: I guess. The startup founders that I have the privilege to call peers, they actually taught me some of the stuff. So, we’ve got the startup founders we want to just connect on the founder journey, we’re happy to connect. Just, but yeah, I think that the strength of the team is sort of making sure that we have our ears to the ground. Get out of the building. You got to get out of the building. And we’ve been trying to get out of the building as much as we can with Katanemo. And I think that journey just continues. The learning journey, the evolution of what we’re doing on behalf of SaaS developers continues, and we hope to delight them.

Corey: I want to thank you for taking the time to speak with me. If people want to learn more, where should they go?

Salman: So, they can go to katanemo.com, which is where our website is, and they can learn a little bit about what we do today and also where we’re headed with the venture. They can reach out to me directly on LinkedIn. Salman Paracha. I’m not super hard to find on LinkedIn. You search for me and say Katanemo or AWS and Oracle, I think you’ll be able to get to me. I’m also going to the AWS New York Summit, which happens on July 26, I believe. I might run into you there.

Corey: Oh, yes. The night before I’ll be hosting a drink up at Vol de Nuit at eight o’clock. You’re welcome there, as anyone who’s listening. And oh, it’s always a pleasure to go and talk to people doing interesting things and just talk shop. But that’s the reason I throw the drink up.

Salman: Ah, okay. I’ll take you up on that. And good, we’ll get to see each other face-to-face after some time. You can reach out to me, as I said, even basic email, and I’ll say that to you, and LinkedIn if you’re just a chat. And there’s just so many ways to get to me. On Twitter, I’m @salmanparacha, and it should be a bit easier to find me.

Don’t hesitate to reach out or search or connect with us. We are eager to talk to folks who are trying to solve or crack this Gordian Knot on terms of the what they’re building. And especially if you’re building towards the next-generation AI application and think through safety, we believe we are years ahead in terms of thinking about safety in that space. It’s early days for us there, but we’re obviously interacting with customers and developers who are trying to think through, how do I now take what was understood to be a table stakes, okay, API-first experiences, [user seems 00:33:31], keys, all that good jazz, and provide safety for that? But I think the new world that we’re going to live in is not only going to just be deterministic responses from APIs; it’s going to be probabilistic responses from large language models. And we got something going on in that space, particularly. We feel fairly bullish on it. But more, customer conversations before we write a piece of code is important. So, just connect with us. I’m salman@katanemo.com, on LinkedIn, Twitter, and I will be quick to reach back out to you.

Corey: And I will, of course, put links to that in the [show notes 00:34:02]. And I’ve also filled out the contact us form on katanemo.com because I have a couple of problems it sounds like this might absolutely be a way to solve. Because otherwise, God help us all. I’m writing another login page.

Salman: Right. So, just see Corey Quinn just signed up for our access. So, I will give you access. So.

Corey: You think I’m kidding. I assure you I’m not. That’s the scariest part is that I’m often being completely serious and people think I’m making a joke. Thank you so much for taking the time to speak with me. I really appreciate it.

Salman: Hey, thanks for the time. I appreciate the opportunity.

Corey: Salman Paracha, founder and CEO at Katanemo. I’m Cloud Economist Corey Quinn and this has been a promoted guest episode of Screaming in the Cloud brought to us by our friends at Katanemo. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with it insulting comment talking about how difficult it was to build that platform yourself from scratch because of all the infrastructure moving parts before it would take that insulting comment.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

View Details

Martin Mao, CEO & Cofounder at Chronosphere, joins Corey on Screaming in the Cloud to discuss the trends he sees in the observability industry. Martin explains why he feels measuring observability costs isn’t nearly as important as understanding the velocity of observability costs increasing, and why he feels efficiency is something that has to be built into processes as companies scale new functionality. Corey and Martin also explore how observability can now be used by business executives to provide top line visibility and value, as opposed to just seeing observability as a necessary cost.

About Martin

Martin is a technologist with a history of solving problems at the largest scale in the world and is passionate about helping enterprises use cloud native observability and open source technologies to succeed on their cloud native journey. He's now the Co-Founder & CEO of Chronosphere, a Series C startup with $255M in funding, backed by Greylock, Lux Capital, General Atlantic, Addition, and Founders Fund. He was previously at Uber, where he led the development and SRE teams that created and operated M3. Previously, he worked at AWS, Microsoft, and Google. He and his family are based in the Seattle area, and he enjoys playing soccer and eating meat pies in his spare time.

Links Referenced:

  • Chronosphere: https://chronosphere.io/
  • LinkedIn: https://www.linkedin.com/in/martinmao/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Human-scale teams use Tailscale to build trusted networks. Tailscale Funnel is a great way to share a local service with your team for collaboration, testing, and experimentation. Funnel securely exposes your dev environment at a stable URL, complete with auto-provisioned TLS certificates. Use it from the command line or the new VS Code extensions. In a few keystrokes, you can securely expose a local port to the internet, right from the IDE.

I did this in a talk I gave at Tailscale Up, their first inaugural developer conference. I used it to present my slides and only revealed that that’s what I was doing at the end of it. It’s awesome, it works! Check it out!

Their free plan now includes 3 users & 100 devices. Try it at snark.cloud/tailscalescream

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. This promoted guest episode is brought to us by our friends at Chronosphere. It’s been a couple of years since I got to talk to their CEO and co-founder, Martin Mao, who is kind enough to subject himself to my slings and arrows today. Martin, great to talk to you.

Martin: Great to talk to you again, Corey, and looking forward to it.

Corey: I should probably disclose that I did run into you at Monitorama a week before this recording. So, that was an awful lot of fun to just catch up and see people in person again. But one thing that they started off the conference with, in the welcome-to-the-show style of talk, was the question about benchmarking: what observability spend should be as a percentage of your infrastructure spent. And from my perspective, that really feels a lot like a question that looks like, “Well, how long should a piece of string be?” It’s always highly contextual.

Martin: Mm-hm.

Corey: Agree, disagree, or are you hopelessly compromised because you are, in fact, an observability vendor, and it should always be more than it is today?

Martin: [laugh]. I would say, definitely agree with you from a exact number perspective. I don’t think there is a magic number like 13.82% that this should be. It definitely depends on the context of how observability is used within a company, and really, ultimately, just like anything else you pay for, it really gets derived from the value you get out of it. So, I feel like if you feel like you’re getting the value out of it, it’s sort of worth the dollars that you put in.

I do see why a lot of companies out there and people are interested because they’re trying to benchmark, to trying to see, am I doing best practice? So, I do think that there are probably some best practice ranges that I’d say most typical organizations out there that we see. This is one thing I’d say. The other thing I’d say when it comes to observability costs is one of the concerns we’ve seen talking with companies is that the relative proportion of that cost to the infrastructure is rising over time. And that’s probably a bad sign for companies because if you extrapolate, you know, if the relative cost of observability is growing faster than infrastructure, and you extrapolate that out a few years, then the direction in which this is going is bad. So, it’s probably more the velocity of growth than the absolute number that folks should be worried about.

Corey: I think that that is probably a fair assessment. I get it all the time, at least in years past, where companies will say, “For every 1000 daily active users, what should it cost to service them?” And I finally snapped in one of my talks that I gave at DevOps Enterprise Summit, and said, I think it was something like $7.34.

Martin: [laugh]. Right, right.

Corey: It’s an arbitrary number that has no context on your business, regardless of whether those users are, you know, Twitter users or large banks you have partnerships with. But now you have something to cite. Does it help you? Not really. But we’ll it get people to leave you alone and stop asking you awkward questions?

Martin: Right, right.

Corey: Also not really, but at least now you have a number.

Martin: Yeah, a hundred percent. And again, like I said, there’s no—and glad magic numbers weren’t too far away from each other. But yeah, I mean, there’s no exact number there, for sure. One pattern I’ve been seeing more recently is, like, rather than asking for the number, there’s been a lot more clarity in companies on figuring out, “Well, okay, before even pick what the target should be, how much am I spending on this per whatever unit of efficiency is?” Right?

And generally, that unit of efficiency, I’ve actually seen it being mapped more to the business side of things, so perhaps to the number of customers or to customer transactions and whatnot. And those things are generally perhaps easier to model out and easier to justify as opposed to purely, you know, the number of seats or the number of end-users. But I’ve seen a lot more companies at least focus on the measurement of things. And again, it’s been more about this sort of, rather than the absolute number, the relative change in number because I think a lot of these companies are trying to figure out, is my business scaling in a linear fashion or sub-linear fashion or perhaps an exponential fashion, if it’s—the costs are, you know, you can imagine growing exponentially, that’s a really bad thing that you want to get ahead of.

Corey: That I think is probably the real question people are getting at is, it seems like this number only really goes up and to the right, it’s not something that we have any real visibility into, and in many cases, it’s just the pieces of it that rise to the occasion. A common story is someone who winds up configuring a monitoring system, and they’ll be concerned about how much they’re paying that vendor, but ignore the fact that, well, it’s beating up your CloudWatch API charges all the time on this other side as well, and data egress is not free—surprise, surprise. So, it’s the direct costs, it’s the indirect costs. And the thing people never talk about, of course, is the cost of people to feed and maintain these systems.

Martin: Yeah, a hundred percent, you’re spot on. There’s the direct costs, there’s the indirect costs. Like you mentioned, in observability, network egress is a huge indirect cost. There’s the people that you mentioned that need to maintain these systems. And I think those are things that companies definitely should take into account when they think about the total cost of ownership there.

I think what’s more in observability actually is, and this is perhaps a hard thing to measure, as well, is often we ask companies, “Well, what is the cost of downtime?” Right? Like if you’re, if your business is impacted and your customers are impacted and you’re down, what is the cost of each additional minute of downtime, perhaps, right? And then the effectiveness of the tool can be evaluated against that because you know, observability is one of these, it’s not just any other developer tool; it’s the thing that’s giving you insight into, is my business or my product or my service operating in the way that I intend. And, you know, is my infrastructure up, for example, as well, right? So, I think there’s also the piece of, like, what is the tool really doing in terms of, like, a lost revenue or brand impact? Those are often things that are sort of quite easily overlooked as well.

Corey: I am curious to see whether you have noticed a shifting in the narrative lately, where, as someone who sells AWS cost optimization consulting as a service, something that I’ve noticed is that until about a year ago, no one really seemed to care overly much about what the AWS bill was. And suddenly, my phone’s been ringing off the hook. Have you found that the same is true in the observability space, where no one really cared what the observability cost, until suddenly, recently, everyone does or has this been simmering for a while?

Martin: We have found that exact same phenomenon. And what I tell most companies out there is, we provide an observability platform that’s targeted at cloud-native platforms. So, if your—a cloud-native architecture, so if you’re running microservices-oriented architecture on containers, that’s the type of architecture that we’ve optimized our solution for. And historically, we’ve always done two things to try to differentiate: one is, provide a better tool to solve that particular problem in that particular architecture, and the second one is to be a more cost-efficient solution in doing so. And not just cost-efficient, but a tool that shows you the cost and the value of the data that you’re storing.

So, we’ve always had both sides of that equation. And to your point, in conversations in the past years, they’ve generally been led with, “Look, I’m looking for a better solution. If you just happen to be cheaper, great. That’s a nice cherry on top.” Whereas this year, the conversations have flipped 180, in which case, most companies are looking for a more cost-efficient solution. If you just happen to be a better tool at the same time, that’s more of a nice-to-have than anything else. So, that conversation has definitely flipped 180 for us. And we found a pretty similar experience to what you’ve been seeing out in the market right now.

Corey: Which makes a tremendous amount of sense. I think that there’s an awful lot of—oh, we’ll just call it strangeness, I think. That’s probably the best way to think about it—in terms of people waking up to the grim reality that not caring about your bills was functionally a zero-interest-rate phenomenon in the corporate sense. Now, suddenly, everyone really has to think about this in some unfortunate and some would say displeasing ways.

Martin: Yeah, a hundred percent. And, you know, it was a great environment for tech for over a decade, right? So, it was an environment that I think a lot of companies and a lot of individuals got used to, and perhaps a lot of folks that have entered the market in the last decade don’t know of another situation or another set of conditions where, you know, efficiency and cost really do matter. So, it’s definitely top of mind, and I do think it’s top of mind for good reason. I do think a lot of companies got fairly inefficient over the last few years chasing that top-line growth.

Corey: Yeah, that has been—and I think it makes sense in the context with which people were operating. Because before a lot of that wound up hitting, it was, well grow, grow, grow at all costs. “What do you mean you’re not doing that right now? You should be doing that right now. Are you being irresponsible? Do we need to come down there and talk to you?”

Martin: A hundred percent.

Corey: Yeah, it’s like eating your vegetables. Now, it’s time to start paying attention to this.

Martin: Yeah, a hundred percent. It’s always a trade-off, right? It’s like in an individual company and individual team, you only have so many resources and prioritization. I do think, to your point, in a zero interest environment, trying to grow that top line was the main thing to do, and hence, everything was pushed on how quickly can we deliver new functionality, new features, or grow that top line. Whereas, the efficiency is always something I think a lot of companies looked at as something I can go deal with later on and go fix. And you know, I feel like that that time has now just come.

Corey: I will say that I somewhat recently had the distinct privilege of working with a company whose observability story was effectively, “We wait for customers to call and tell us there’s a problem and then we go looking in into it.” And on the one hand, my immediate former SRE reflexes kicked in, and I recoiled, but this company has been in this industry longer than I have. They clearly have a model that is working for them and for their customers. It’s not the way I would build something, but it does seem that for some use cases, you absolutely are going to be okay with something like that. And I should probably point out, they were not, for example, a bank where yeah, you kind of want to get some early warning on things that could destabilize the economy.

Martin: Right, right. I mean, to your point, depending on the context, and the company, it could definitely make sense, and depending on how they execute it as well, right? So, you know, you called out an example already, where if they were a bank or if any correctness or timeliness of a response was important to that business, perhaps not the best thing to do to have your customers find out, especially if you have a ton of customers at the same time. But however, you know, if it’s a different type of business where, you know, the responses are perhaps more asynchronous or you don’t have a lot of users encountering at the same time or perhaps you have a great A/B experimentation platform and testing platform, you know, there are definitely conditions in which that could be potentially a viable option.

Especially when you weigh up the cost and the benefit, right? If the cost to having a few bad customers have a bad experience is not that much to the business and the benefit is that you don’t have to spend a ton on observability, perhaps that’s a trade-off that the company is willing to make. In most of the businesses that we’ve been working with, I would say that probably not been the case, but I do think that there’s probably some bias and some skew there in the sense that you can imagine a company that cares about these things, perhaps it’s more likely to talk to an observability vendor like us to try to fix these problems.

Corey: When we spoke a few years back, you definitely were focused on the large, one would say, almost hyperscale style of cloud-native build-out. Is that still accurate or has the bar to entry changed since we last spoke? I know you’ve raised an awful lot of money, which good for you. It’s a sign of a healthy, robust VC ecosystem. What the counterpoint to that is, they’re probably not investing in a company whose total addressable market is, like, 15 companies that must be at least this big.

Martin: [laugh]. Yeah, a hundred percent. A hundred percent. So, I’ll tell you that the bar to entry definitely has changed, but it’s not due to a business decision on our end. If you think about how we started and, you know, the focus area, we’re really targeting accounts that are adopting cloud-native technology.

And it just so happens that the large tech, [decacorns 00:12:35], and the hyperscalers were the earliest adopters of cloud-native. So containerization, microservices, they were the earliest adopters of that, so hence, there was a high correlation in the companies that had that problem and the companies that we could serve. Luckily, for us, the trend has been that more of the rest of the industry has gone down this route as well. And it’s not just new startups; you can imagine any new startup these days probably starts off cloud-native from day one, but what we’re finding is the more established, larger enterprises are doing this shift as well. And I think the folks out there like Gartner have studied this and predicted that, you know, by about 2028, I believe was the date, about 95% of applications are going to be containerized in large enterprises. So, it’s definitely a trend that the rest of the industry will go on. And as they continue down that trend, that’s when, sort of, our addressable market will grow because the amount of use cases where our technology shines will grow along with that as well.

Corey: I’m also curious about your description of being aimed at cloud-native companies. You gave one example of microservices powered by containers, if I understood correctly. What are the prerequisites for this? When you say that it almost sounds like you’re trying to avoid defining a specific architecture that you don’t want to deal well with or don’t want to support for a variety of reasons? Is that what it is or is there certain you must be built in these ways or the product does not work super well for you? What is it you’re trying to say with that, is what I’m trying to get at here.

Martin: Yeah, a hundred percent. If you look at the founding story here, it’s really myself and my co-founder, found Uber going through this transition of both a new architecture, in the sense that, you know, they were going containers, they were building microservices-oriented architecture there, were also adopting a DevOps mentality as well. So, it was just a new way of building software, almost. And what we found is that when you develop software in this particular way—so you can imagine when you’re developing a tiny piece of functionality as a microservice and you’re a individual developer, and you’re—you know, you can imagine rolling that out into production multiple times a day, in that way of developing software, what we found was that the traditional tools, the application performance monitoring tools, the IT monitoring tools that used to exist pre this way of both architecture and way of developing software just weren’t a good fit.

So, the whole reason we exist is that we had to figure out a better way of solving this particular problem for the way that Uber built software, which was more of a cloud-native approach. And again, it just so happens that the rest of the industry is moving down this path as well and hence, you know, that problem is larger for a larger portion of the companies out there. You know, I would say some of the things when you look into why the existing solutions can’t solve these problems well, you know, if you look at a application performance monitoring tool, an APM tool, it’s really focused on introspecting into that application and its interaction with the operating system or the underlying hardware. And yet, these days, that is less important when you’re running inside the container. Perhaps you don’t even have access to the underlying hardware, or the operating system and what you care about—you can imagine—is how that piece of functionality interacts with all the other pieces of functionality out there, over a network core.

So, just the architecture and the conditions ask for a different type of observability, a different type of monitoring, and hence, you just need a different type of solution to go solve for this new world. Along with this, which is sort of related to the cost as well, is that, you know, as we go from virtual machines onto containers, you can imagine the sheer volume of data that gets produced now because everything is much smaller than it was before and a lot more ephemeral than it was before, and hence, every small piece of infrastructure, every small piece of code, you can imagine still needs as much monitoring and observability as it did before, as well. So, just the sheer volume of data is so much larger for the same amount of infrastructure, for the same amount of hardware that that you used to have, and that’s really driving a huge problem in terms of being able to scale for it and also being able to pay for these systems as well.

Corey: Tired of Apache Kafka's complexity making your AWS bill look like a phone number? Enter Redpanda. You get 10x your streaming data performance without having to rob a bank. And migration? Smoother than a fresh jar of peanut butter. Imagine cutting as much as 50% off your AWS bills. With Redpanda, it's not a dream, it's reality. Visit go.redpanda.com/duckbill. Redpanda: Because Kafka shouldn't cause you nightmares.

Corey: I think that there’s a common misconception in the industry that people are going to either have ancient servers rotting away in racks, or they’re going to build something greenfield, the way that we see done on keynote stage is all the time of companies that have been around with this architecture for less than 18 months. In practice, I find it’s awfully frequent that this is much more of a spectrum, and a case-by-case per-workload basis. I haven’t met too many data center companies where everything’s the disaster that the cloud companies like to paint it as, and vice versa, I also have never yet seen a architecture that really existed as described in a keynote presentation.

Martin: A hundred percent agree with you there. And you know, it’s not clean-cut from that perspective. And also, you’re also forgetting the messy middle as well, right? Like, often what happens is, there’s a transition. If you don’t start off cloud-native from day one, you do need to transition there from your monolithic applications, from your VM-based architectures, and often the use case can’t transform over perfectly.

What ends up happening is you start moving some functionality and containerizing some functionality and that still has dependencies between the old architecture and the new architecture. And companies have to live in this middle state, perhaps for a very long time. So, it’s definitely true. It’s not a clean-cut transition. But you can think about that middle state is actually one that a lot of companies struggle with because all of a sudden, you only have a partial view of the world, or what’s happening with your old tools, they’re not well suited for the new environments. Perhaps you got to start bringing new tools and new ways of doing things in your new environments, and they’re not perhaps the best suited for the old environment as well.

So, you do actually end up in this middle state where you need a good solution that can really handle both because there are a lot of interdependencies between the two. And it’s actually one of the things that we strive to do here at Chronosphere is to help companies through that transition. So, it’s not just all of the new use cases and it’s not just all of your new environments. It’s actually helping companies through this transition is actually pretty critical as well.

Corey: My question for you is, given that you have, I don’t want to say a preordained architecture that your customers have to use, but there are certain assumptions you’ve made based upon both their scale and the environment in which they’re operating. How heavy of a lift is it for them to wind up getting Chronosphere into their environments? Just because seems to me that it’s not that hard to design an architecture on a whiteboard that can meet almost any requirement. The messy part is figuring out how to get something that resembles that into place on a pre-existing, extant architecture.

Martin: Yeah. I’d say it’s something we spent a lot of time on. The good thing for the industry overall, for the observability industry, is that open-source standards are now created and now exist when they didn’t before. So, if you look at the APM-based view, it was all proprietary agents producing the data themselves that would only really work with one vendor product, whereas if you’ve look at a modern environment, the production of the data has actually been shifted from the vendor down to the companies themselves, and there’ll be producing these pieces of data in open-source standard formats like OpenTelemetry for distributed traces, or perhaps Prometheus for metrics.

So, the good thing is that for all of the new environments, there’s a standard way to produce all of this data and you can send all that data to whichever vendor you want on the back end. So, it just makes the implementation for the new environments so much easier. Now, for the legacy environments, or if a company is shifting over from an existing tool, there is actually a messy migration there because often you’re trying to replace proprietary formats and proprietary ways of producing data with open-source standard ones. So, just something that us as Chronosphere just come in and we view that as a particular problem that we need to solve and we take the responsibility of solving for a company because what we’re trying to sell companies is not just a tool, what we’re really trying to solve them is the solution to the problem, and the problem is they need an observability solution end to end. So, this often involves us coming in and helping them, you can imagine, not just convert the data types over but also move over existing dashboards, existing alerts.

There’s a huge piece of lift that the end—that perhaps every developer in a company would have to do if we didn’t come in and do it on behalf of those companies. So, it’s just an additional responsibility. It’s not an easy thing to do. We’ve built some tooling that helps with it, and we just spend a lot of manual hours going through this, but it’s a necessary one in order to help a company transition. Now, the good thing is, once they have transitioned into the new way of doing things and they are dependent on open-source standard formats, they are no longer locked in. So, you know, you can imagine future transitions will be much easier, however the current one does have to go through a little bit of effort.

Corey: I think that’s probably fair. And then there’s no such thing, in my experience, as a easy deployment for something that is large enough to matter. And let’s be clear, people are not going to be deploying something as large scale as Chronosphere on a lark. This is going to be when they have a serious application with serious observability challenges. So, it feels like, on some level, that even doing a POC is a tricky proposition, just due to the instrumentation part of it. Something I’ve seen is that very often, enterprise sales teams will decide that by the time that they can get someone to successfully pull off a POC, at that point, the deal win rate is something like 95% just because no one wants to try that in a bake-off with something else.

Martin: Yeah, I’d say that we do see high pilot conversion rates, to your point. For us, it’s perhaps a little bit easier than other solutions out there, in the sense that I think with our type of observability tooling, the good thing is, an individual team could pick this up for their one use case and they could get value out of it. It’s not that every team across an environment or every team in an organization needs to adopt. So, while generally, we do see that, you know, a company would want to pilot and it’s not something you can play around online with by yourself because it does need a particular deployment, it does need a little bit of setup, generally one single team can come and perform that and see value out of the tool. And that sort of value can be extrapolated and applied to all the other teams as well. So, you’re correct, but it hasn’t been a huge lift. And you know, these processes end to end, we’ve seen be as short as perhaps 30-something days end to end, which is generally a pretty fast-moving process there.

Corey: Now, I guess, on some level, I’m still trying to wrap my head around the idea of the scale that you operate at, just because as you mentioned, this came out of Uber—which is beyond imagining for most people—and you take a look at a wide variety of different use cases. And in my experience it’s never been, “Holy crap, we have no observability and we need to fix that.” It’s, “There are a variety of systems in place that just are not living up to the hopes, dreams, and potential that they had when they were originally deployed.” Either due to growth or due to lack of product fit, or the fact that it turns out in a post zero-interest-rate world, most people don’t want to have a pipeline of 20 discrete observability tools.

Martin: Yep, yep. No, a hundred percent. And, to your point there, ultimately, it’s our goal and, you know, in many companies were replacing up to six to eight tools in a single platform. And so, it’s great to do. That definitely doesn’t happen overnight. It takes time.

You know, you can imagine in a pilot or when you’re looking at it, we’re picking a few of the use cases to demonstrate what our tool could do across many other use cases, and then generally on the onboarding, during the onboarding time or perhaps over a period of months or perhaps even a year plus, we then go on board these use cases a piece by piece. So, it’s definitely not a quick overnight process there, but, you know, you can imagine something that can help each end developer in that particular company be more effective and it’s something that can really help move the bottom line in terms of far better price efficiency. These things are generally not things that are quick fixes; these are generally things that do take some time and a little bit of investment to achieve the results.

Corey: So, a question I do have for you, given that I just watched an awful lot of people talking about observability for three days at Monitorama, what are people not talking about? What did you not see discussed that you think should be?

Martin: Yeah, one thing I think often gets overlooked, and especially in today’s climate is, I think observability gets relegated to a cost center. It’s something that every company must have, every company has today, and it’s often looked at a tool that gives you insights about your infrastructure and your applications and it’s a backend tool, something you have to have, something you have to pay for and it doesn’t really move the direct needle for the business top line. And I think that’s often something that companies don’t talk about enough. And you know, from our experience at Uber and through most of the companies that we work with here at Chronosphere, yes, there are infrastructure problems and application level problems that we help companies solve, but ultimately, the more mature organizations, or when it comes to observability, are often starting to get real-time insights into the business more than the application layer and the infrastructure layer.

And if you think about it, companies that are cloud-native architected, there’s not one single endpoint or one single application that fulfills a single customer request. So, even if you could look at all the individual pieces, the actual what we have to do for customers in our products and services span across so many of them that often you need to introduce a new view, a view that’s just focused on your customers, just focused on the business, and sort of apply the same type of techniques on your backend infrastructure as you do for your business. Now, this isn’t a replacement for your BI tools, you still need those, but what we find is that BI tools are more used for longer-term strategic decisions, whereas you may need to do a lot of sort of tactical, more tactical, business operational functions based on having a live view of the business. So, what we find is often observability is only ever thought about for infrastructure, it’s only ever thought about as a cost center, but ultimately observability tooling can actually add a lot directly to your top line by giving you visibility into the products and services that make up that top line. And I would say the more mature organizations that we work with here at Chronosphere all had their executives looking at, you know, monitoring dashboards to really get a good sense of what’s happening in their business in real-time. So, I think that’s something that hopefully a lot more companies evolve into over time and they really see the full benefit of observability and what it can do to a business’s top line.

Corey: I think that’s probably a fair way of approaching it. It seems similar, in some respects, to what I tend to see over in the cloud cost optimization space. People often want to have something prescriptive of, do this, do that do the other thing, but it depends entirely what the needs of the business are internally, it depends upon the stories that they wind up working with, it depends really on what their constraints are, what their architectures are doing. Very often it’s a let’s look and figure out what’s going on and accidentally, they discover they can blow 40% off their spend by just deleting things that aren’t in use anymore. That becomes increasingly uncommon with scale, but it’s still one of those questions of, “What do we do here and how?”

Martin: Yep, a hundred percent.

Corey: I really want to thank you for taking the time to speak with me today about what you’re seeing. If people want to learn more, where’s the best place for them to find you?

Martin: Yeah, the best place is probably going to our website Chronosphere.io to find out more about the company, or if you want to chat with me directly, LinkedIn is probably the best place to come find me, via my name.

Corey: And we will, of course, put links to both of those things in the [show notes 00:28:49]. Thank you so much for suffering the slings and arrows I was able to throw at you today.

Martin: Thank you for having me Corey. Always a pleasure to speak with you, and looking forward to our next conversation.

Corey: Likewise. Martin Mao, CEO and co-founder of Chronosphere. This promoted guest episode has been brought to us by Chronosphere, here on Screaming in the Cloud. And I’m Cloud Economist Corey Quinn. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an insulting comment that I will never notice because I have an observability gap.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

View Details

Andreas Wittig, Co-Author of Amazon Web Services in Action and Co-Founder of marbot, joins Corey on Screaming in the Cloud to discuss ways to keep a book up to date in an ever-changing world, the advantages of working with a publisher, and how he began the journey of writing a book in the first place. Andreas also recalls how much he learned working on the third edition of Amazon Web Services in Action and how teaching can be an excellent tool for learning. Since writing the first edition, Adreas’s business has shifted from a consulting business to a B2B product business, so he and Corey also discuss how that change came about and the pros and cons of each business model.

About Andreas

Andreas is the Co-Author of Amazon Web Services in Action and Co-Founder of marbot - AWS Monitoring made simple! He is also known on the internet as cloudonaut through the popular blog, podcast, and youtube channel he created with his brother Michael.

Links Referenced:

  • Amazon Web Services in Action: https://www.amazon.com/Amazon-Services-Action-Andreas-Wittig/dp/1617295116
  • Rapid Docker on AWS: https://cloudonaut.io/rapid-docker-on-aws/
  • bucket/av: https://bucketav.com/
  • marbot: https://marbot.io/
  • cloudonaut.io: https://cloudonaut.io

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. It’s been a few years since I caught up with Andreas Wittig, who is also known in the internet as cloudonaut, and much has happened since then. Andreas, how are you?

Andreas: Hey, absolutely. Thank you very much. I’m happy to be here in the show. I’m doing fine.

Corey: So, one thing that I have always held you in some high regard for is that you have done what I have never had the attention span to do: you wrote a book. And you published it a while back through Manning, it was called Amazon Web Services in Action. That is ‘in action’ two words, not Amazon Web Services Inaction of doing absolutely nothing about it, which is what a lot of companies in the space seem to do instead.

Andreas: [laugh]. Yeah, absolutely. So. And it was not only me. I’ve written the book together with my brother because back in 2015, Manning, for some reason, wrote in and asked us if we would be interested in writing the book.

And we had just founded our own consulting company back then and we had—we didn’t have too many clients at the very beginning, so we had a little extra time free. And then we decided, okay, let’s do the book. And let’s write a book about Amazon Web Services, basically, a deep introduction into all things AWS. So, this was 2015, and it was indeed a lot of work, much more [laugh] than we expected. So, first of all, the hard part is, what do you want to have in the book? So, what’s the TOC? What is important and must be in?

And then you start writing and have examples and everything. So, it was really an interesting journey. And doing it together with a publisher like Manning was also really interesting because we learned a lot about writing. You have kind of a coach, an editor that helps you through that process. So, this was really a hard and fun experience.

Corey: There’s a lot of people that have said very good things about writing the book through a traditional publisher. And they also say that one of the challenges is it’s a blessing and a curse, where you basically have someone standing over your shoulder saying, “Is it done yet? Is it done yet? Is it done yet?” The consensus that seems to have emerged from people who have written books is, “That was great, please don’t ever ask me to do it again.”

And my operating theory is that no one wants to write a book. They want to have written a book. Which feels like two very different things most of the time. But the reason you’re back on now is that you have gone the way of the terrible college professor, where you’re going to update the book, and therefore you get to do a whole new run of textbooks and make everyone buy it and kill the used market, et cetera. And you’ve done that twice now because you have just recently released the third edition. So, I have to ask, how different is version one from version two and from version three? Although my apologies; we call them ‘editions’ in the publishing world.

Andreas: [laugh]. Yeah, yeah. So, of course, as you can imagine, things change a lot in AWS world. So, of course, you have to constantly update things. So, I remember from first to second edition, we switched from CloudFormation in JSON to YAML. And now to the third edition, we added two new chapters. This was also important to us, so to keep also the scope of the book in shape.

So, we have in the third edition, two new chapters. One is about automating deployments, recovering code deploy, [unintelligible 00:03:59], CloudFormation rolling updates in there. And then there was one important topic missing at all in the book, which was containers. And we finally decided to add that in, and we have now container chapter, starting with App Runner, which I find quite an interesting service to observe right now, and then our bread and butter service: ECS and Fargate. So, that’s basically the two new chapters. And of course, then reworking all the other chapters is also a lot of work. And so, many things change over time. Cannot imagine [laugh].

Corey: When was the first edition released? Because I believe the second one was released in 2018, which means you’ve been at this for a while.

Andreas: Yeah. So, the first was 2015, the second 2018, three years later, and then we had five years, so now this third edition was released at the beginning of this year, 2023.

Corey: Eh, I think you’re right on schedule. Just March of 2020 lasted three years. That’s fine.

Andreas: Yeah [laugh].

Corey: So, I have to ask, one thing that I’ve always appreciated about AWS is, it feels like with remarkably few exceptions, I can take a blog post written on how to do something with AWS from 2008 and now in 2023, I can go through every step along with that blog post. And yeah, I might have trouble getting some of the versions and services and APIs up and running, but the same steps will absolutely work. There are very few times where a previously working API gets deprecated and stops working. Is this the best way to proceed? Absolutely not.

But you can still spin up the m1.medium instance sizes, or whatever it was, or [unintelligible 00:05:39] on small or whatever the original only size that you could get was. It’s just there are orders of magnitude and efficiency gains you can do by—you can go through by using more modern approaches. So, I have to ask, was there anything in the book as you revised it—two times now—that needed to come out because it was now no longer working?

Andreas: So, related to the APIs that’s—they are really very stable, you’re right about that. So, the problem is, our first few chapters where we have screenshots of how you go through the management—

Corey: Oh no.

Andreas: —console [laugh]. And you can probably, you can redo them every three months, probably, because the button moves or a step is included or something like that. So, the later chapters in the book, where we focus a lot on the CLI or CloudFormation and stuff like—or SDKs, they are pretty stable. But the first few [ones 00:06:29] are a nightmare to update all those screenshots. And then sometimes, so I was going through the book, and then I noticed, oh, there’s a part of this chapter that I can completely remove nowadays.

So, I can give you an example. So, I was going through the chapter about Simple Storage Service S3, and I—there was a whole section in the chapter about read-after-write consistency. Because back then, it was important that you knew that after updating an object or reading an object before it was created the first time, you could get outdated versions for a little while, so this was eventually consistent. But nowadays, AWS has changed that and basically now, S3 has this strong read-after-write consistency. So, I basically could remove that whole part in the chapter which was quite complicated to explain to the reader, right, so I [laugh] put a lot of effort into that.

Corey: You think that was confusing? I look at the sea of systems I had to oversee at one company, specifically to get around that problem. It’s like, well, we can now take this entire application and yeet it into the ocean because it was effectively a borderline service to that just want to ens—making consistency guarantees. It’s not a common use case, but it is one that occurs often enough to be a problem. And of course, when you need it, you really need it. That was a nice under-the-hood change that was just one day, surprise, it works that way. But I’m sure it was years of people are working behind the scenes, solving for impossible problems to get there, and cetera, et cetera.

Andreas: Yeah, yeah. But that’s really cool is to remove parts of the book that are now less complicated. This is really cool. So, a few other examples. So, things change a lot. So, for example, EFS, so we have EFS, Elastic File System, in the book as well. So, now we have new throughput modes, different limits. So, there’s really a lot going on and you have to carefully go through all the—

Corey: Oh, when EFS launched, it was terrible. Now, it’s great just because it’s gotten so much more effective and efficient as a service. It’s… AWS releases things before they’re kind of ready, it feels like sometimes, and then they improve with time. I know there have been feature deprecations. For example, for some reason, they are no longer allowing us to share out a bucket via BitTorrent, which, you know, in 2006 when it came out, seemed like a decent idea to save on bandwidth. But here in 2023, no one cares about it.

But I’m also keeping a running list of full-on AWS services that have been deprecated or have the deprecations announced. Are any of those in the book in any of its editions? And if and when there’s a fourth edition, will some of those services have to come out?

Andreas: [laugh]. Let’s see. So, right after the book was published—because the problem with books is they get printed, right; that’s the issue—but the target of the book, AWS, changes. So, a few weeks after the printed book was out, we found out that we have an issue in our one of our examples because now S3 buckets, when you create them, they have locked public access enabled by default. And this was not the case before. And one of our example relies on that it can create object access control lists, and this is not working now anymore. [laugh].

So yeah, there are things changing. And we have, the cool thing about Manning is they have that what they call a live book, so you can read it online and you can have notes from other readers and us as the authors along the text, and there we can basically point you in the right direction and explain what happened here. So, this is how we try to keep the book updated. Of course, the printed one stays the same, but the ebook can change over time a little bit.

Corey: Yes, ebooks are… at least keeping them updated is a lot easier, I would imagine. It feels like that—speaking of continuous builds and automatic CI/CD approaches—yeah, well, we could build a book just by updating some text in a Git repo or its equivalent, and pressing go, but it turns out that doing a whole new print run takes a little bit more work.

Andreas: Yeah. Because you mentioned the experience of writing a book with a publisher and doing it on your own with self-publishing, so we did both in the past. We have Amazon Web Services in Action with Manning and we did another book, Rapid Docker on AWS in self-publishing. And what we found out is, there’s really a lot of effort that goes into typesetting and layouting a book, making sure it looks consistent.

And of course, you can just transform some markdown into a epub and PDF versions, but if a publisher is doing that, the results are definitely different. So, that was, besides the other help that we got from the publisher, very helpful. So, we enjoyed that as well.

Corey: What is the current state of the art—since I don’t know the answer to this one—around updating ebook versions? If I wind up buying an ebook on Kindle, for example, will they automatically push errata down automatically through their system, or do they reserve that for just, you know, unpublishing books that they realized shouldn’t be on the Marketplace after people have purchased them?

Andreas: [laugh]. So—

Corey: To be fair, that only happened once, but I’m still giving them grief for it a decade and change later. But it was 1984. Of all the books to do that, too. I digress.

Andreas: So, I’m not a hundred percent sure how it works with the Kindle. I know that Manning pushes out new versions by just emailing all the customers who bought the book and sending them a new version. Yeah.

Corey: Yeah. It does feel, on some level, like there needs to be at least a certain degree of substantive change before they’re going to start doing that. It’s like well, good news. There was a typo on page 47 that we’re going to go ahead and fix now. Two letters were transposed in a word. Now, that might theoretically be incredibly important if it’s part of a code example, which yes, send that out, but generally, A, their editing is on point, so I didn’t imagine that would sneak through, and 2, no one cares about a typo release and wants to get spammed over it?

Andreas: Definitely, yeah. Every time there’s a reprint of the book, you have the chance to make small modifications, to add something or remove something. That’s also a way to keep it in shape a little bit.

Corey: I have to ask, since most people talk about AWS services to a certain point of view, what is your take on databases? Are you sticking to the actual database services or are you engaged in my personal hobby of misusing everything as a database by holding it wrong?

Andreas: [laugh]. So, my favorite database for starting out is DynamoDB. So, I really like working with DynamoDB and I like the limitations and the thing that you have to put some thoughts into how to structure your data set in before. But we also use a lot of Aurora, which really find an interesting technology. Unfortunately, Aurora Serverless, it’s not becoming a product that I want to use. So, version one is now outdated, version two is much too expensive and restricted. So—

Corey: I don’t even know that it’s outdated because I’m seeing version one still get feature updates to it. It feels like a divergent service. That is not what I would expect a version one versus version two to be. I’m with you on Dynamo, by the way. I started off using that and it is cheap is free for most workloads I throw at it. It’s just a great service start to finish. The only downside is that if I need to move it somewhere else, then I have a problem.

Andreas: That’s true. Yeah, absolutely.

Corey: I am curious, as far as you look across the sea of change—because you’ve been doing this for a while and when you write a book, there’s nothing that I can imagine that would be better at teaching you the intricacies of something like AWS than writing a book on it. I got a small taste of this years ago when I shot my mouth off and committed to give a talk about Git. Well, time to learn Git. And teaching it to other people really solidifies a lot of the concepts yourself. Do you think that going through the process of writing this book has shaped how you perform as an engineer?

Andreas: Absolutely. So, it’s really interesting. So,I added the third edition and I worked on it mostly last year. And I didn’t expect to learn a lot during that process actually, because I just—okay, I have to update all the examples, make sure everything work, go through the text, make sure everything is up to date. But I learned things, not only new things, but I relearned a lot of things that I wasn’t aware of anymore. Or maybe I’ve never been; I don’t know exactly [laugh].

But it’s always, if you go into the details and try to explain something to others, you learn a lot about that. So, teaching is a very good way to, first of all gather structure and a deep understanding of a topic and also dive into the details. Because when you write a book, every time you write a sentence, ask the question, is that really correct? Do I really know that or do I just assume that? So, I check the documentation, try to find out, is that really the case or is that something that came up myself?

So, you’ll learn a lot by doing that. And always come to the limits of the AWS documentation because sometimes stuff is just not documented and you need to figure out, what is really happening here? What’s the real deal? And then this is basically the research part. So, I always find that interesting. And I learned a lot in during the third edition, while was only adding two new chapters and rewriting a lot of them. So, I didn’t expect that.

Corey: Do you find that there has been an interesting downstream effect from having written the book, that for better or worse, I’ve always no—I always notice myself responding to people who have written a book with more deference, more acknowledgment for the time and effort that it takes. And some books, let’s be clear, are terrible, but I still find myself having that instinctive reaction because they saw something through to be published. Have you noticed it changing other aspects of your career over the past, oh, dear Lord, it would have been almost ten years now.

Andreas: So, I think it helped us a lot with our consulting business, definitely. Because at the very beginning, so back in 2015, at least here in Europe and Germany, AWS was really new in the game. And being the one that has written a book about AWS was really helping… stuff. So, it really helped us a lot for our consulting work. I think now we are into that game of having to update the book [laugh] every few years, to make sure it stays up to date, but I think it really helped us for starting our consulting business.

Corey: And you’ve had a consulting business for a while. And now you have effectively progressed to the next stage of consulting business lifecycle development, which is, it feels like you’re becoming much more of a product company than you were in years past. Is that an accurate perception from the outside or am I misunderstanding something fundamental?

Andreas: You know, absolutely, that’s the case. So, from the very beginning, basically, when we founded our company, so eight years ago now, so we always had to go to do consulting work, but also do product work. And we had a rule of thumb that 20% of our time goes into product development. And we tried a lot of different things. So, we had just a few examples that failed completely.

So, we had a Time [Series 00:17:49] as a Service offering at the very beginning of our journey, which failed completely. And now we have Amazon Timestream, which makes that totally—so now the market is maybe there for that. We tried a lot of things, tried content products, but also as we are coming from the software development world, we always try to build products. And over the years, we took what we learned from consulting, so we learned a lot about, of course, AWS, but also about the market, about the ecosystem. And we always try to bring that into the market and build products out of that.

So nowadays, we really transitioned completely from consulting to a product company, as you said. So, we do not do any consulting anymore with one few exception with one of our [laugh] best or most important clients. But we are now a product company. And we only a two-person company. So, the idea was always how to scale a company without growing the team or hiring a lot of people, and a consulting business is definitely not a good way to do that, so yeah, this was why always invested into products.

And now we have two products in the AWS Marketplace which works very well for us because it allows us to sell worldwide and really easily get a relationship up and running with our customers, and that pay through their AWS bill. So, that’s really helping us a lot. Yeah.

Corey: A few questions on that. At first it always seems to me that writing software or building a product is a lot like real estate in that you’re doing a real estate development—to my understanding since I live in San Francisco and this is a [two exit 00:19:28] town; I still rent here—I found though, that you have to spend a lot of money and effort upfront and you don’t get to start seeing revenue on that for years, which is why the VC model is so popular where you’ll take $20 million, but then in return they want to see massive, outsized returns on that, which—it feels—push an awful lot of perfectly sustainable products into things that are just monstrous.

Andreas: Hmm, yeah. Definitely.

Corey: And to my understanding, you bootstrapped. You didn’t take a bunch of outside money to do this, right?

Andreas: No, no, we have completely bootstrapping and basically paying the bills with our consulting work. So yeah, I can give you one example. So, bucketAV is our solution to scan S3 buckets for malware, and basically, this started as an open-source project. So, this was just a side project we are working on. And we saw that there is some demand for that.

So, people need ways to scan their objects—for example, user uploads—for malware, and we just tried to publish that in the AWS Marketplace to sell it through the Marketplace. And we don’t really expect that this is a huge deal, and so we just did, I don’t know, Michael spent a few days to make sure it’s possible to publish that and get in shape. And over time, this really grew into an important, really substantial part of our business. And this doesn’t happen overnight. So, this adds up, month by month. And you get feedback from customers, you improve the product based on that. And now this is one of the two main products that we sell in the Marketplace.

Corey: I wanted to ask you about the Marketplace as well. Are you finding that that has been useful for you—obviously, as a procurement vehicle, it means no matter what country a customer is in, they can purchase it, it shows up on the AWS bill, and life goes on—but are you finding that it has been an effective way to find new customers?

Andreas: Yes. So, I definitely would think so. It’s always funny. So, we have completely inbound sales funnel. So, all customers find us through was searching the Marketplace or Google, probably. And so, what I didn’t expect that it’s possible to sell a B2B product that way. So, we don’t know most of our customers. So, we know their name, we know the company name, but we don’t know anyone there. We don’t know the person who buys the product.

This is, on the one side, a very interesting thing as a two-person company. You cannot build a huge sales process and I cannot invest too much time into the sales process or procurement process, so this really helps us a lot. The downside of it is a little bit that we don’t have a close relationship with our customers and sometimes it’s a little tricky for us to find important person to talk to, to get feedback and stuff. But on the other hand, yeah, it really helps us to sell to businesses all over the world. And we sell to very small business of course, but also to large enterprise customers. And they are fine with that process as well. And I think, even the large enterprises, they enjoy that it’s so easy [laugh] to get a solution up and running and don’t have to talk to any salespersons. So, enjoy it and I think our customers do as well.

Corey: This is honestly the first time I’ve ever heard a verifiable account a vendor saying, “Yeah, we put this thing on the Marketplace, and people we’ve never talked to find us on the Marketplace and go ahead and buy.” That is not the common experience, let’s put it that way. Now true, an awful lot of folks are selling enterprise software on this and someone—I forget who—many years ago had a great blog post on why no enterprise software costs $5,000. It either is going to cost $500 or it’s going to cost 100 grand and up because the difference is, is at some point, you’d have a full-court press enterprise sales motion to go and sell the thing. And below a certain point, great, people are just going to be able to put it on their credit card and that’s fine. But that’s why you have this giant valley of there is very little stuff priced in that sweet spot.

Andreas: Yeah. So, I think maybe it’s important to mention that our products are relatively simple. So, they are just for a very small niche, a solution for a small problem. So, I think that helps a lot. So, we’re not selling a full-blown cloud security solution; we only focus on that very small part: scanning S3 objects for malware.

For example, on marbot,f the other product that we sell, which is monitoring of AWS accounts. Again, we focus on a very simple way to monitor AWS workloads. And so, I think that is probably why this is a successful way for us to find new customers because it’s not a very complicated product where you have to explain a lot. So, that’s probably the differentiator here.

Corey: Having spent a fair bit of time doing battle with compliance goblins—which is, to be clear, I’m not describing people; I’m describing processes—in many cases, we had to do bucket scanning for antivirus, just to check a compliance box. From our position, there was remarkably little risk of a user-generated picture of a receipt that is input sanitized to make sure it is in fact a picture, landing in an S3 bucket and then somehow infecting one of the Linux servers through which it passed. So, we needed something that just checked the compliance box or we would not be getting the gold seal on our website today. And it was, more or less, a box-check as opposed to something that solved a legitimate problem. This was also a decade and change ago. Has that changed to a point now where there are legitimate threats and concerns around this, or is it still primarily just around make the auditor stop yelling at me, please?

Andreas: Mmm. I think it’s definitely to tick the checkbox, to be compliant with this, some regulation. On the other side, I think there are definitely use cases where it makes a lot of sense, especially when it comes to user-generated content of all kinds, especially if you’re not only consuming it internally, but maybe also others can immediately start downloading that. So, that is where we see many of our customers are coming with that scenario that they want to make sure that the files that people upload and others can download are not infected. So, that is probably the most important use case.

Corey: There’s also, on some level, an increasing threat of ransomware. And for a long time, I was very down on the ideas of all these products that hit the market to defend S3 buckets against ransomware. Until one day, there was an AWS security blog post talking about how they found it. And yeah, we’ve we have seen this in the wild; it is causing problems for companies; here’s what to do about it. Because it’s one of those areas where I can’t trust a vendor who’s trying to sell me something to tell me that this problem exists.

I mean, not to cast aspersions, but they’re very interested, they’re very incentivized to tell that story, whereas AWS is not necessarily incentivized to tell a story like that. So, that really brought it home for me that no, this is a real thing. So, I just want to be clear that my opinion on these things does in fact, evolve. It’s not, “Well, I thought it was dumb back in 2012, so clearly it’s still dumb now.” That is not my position, I want to be very clear on that.

I do want to revisit for a moment, the idea of going from a consultancy that is a services business over to a product business because we’ve toyed with aspects of that here at The Duckbill Group a fair bit. We’ve not really found long-term retainer services engagements that add value that we are comfortable selling. And that means as a result that when you sell fixed duration engagements, it’s always a sell, sell, sell, where’s the next project coming from? Whereas with product businesses, it’s oh, the grass is always greener on the other side. It’s recurring revenue. Someone clicks, the revenue sticks around and never really goes away. That’s the dream from where I sit on the services side of the fence, wistfully looking across and wondering what if. Now that you’ve made that transition, what sucks about product businesses that you might not have seen going into it?

Andreas: [laugh]. Yeah, that a good question. So, on the one side, it was really also our dream to have a product business because it really changes the way we work. We can block large parts of our calendar to do deep-focus work, focus on things, find new solutions, and really try to make a solution that really fits to problem and uses all the AWS capabilities to do so. And on the other side, a product business involves, of course, selling the product, which is hard.

And we are two software engineers, [laugh] and really making sure that we optimize our sales and there’s search engine optimization, all that stuff, this is really hard for us because we don’t know anything about that and we always have to find an expert, or we need to build a knowledge ourself, try things out, and so on. So, that whole part of selling the product, this is really a challenge for us. And then of course, product business evolves a lot of support work. So, we get support emails multiple times per hour, and we have to answer them and be as fast as possible with that. So, that is, of course, something that you do not have to do with consulting work.

And not always that, the questions are many times really simple questions that pointed people in the right direction, find part of the documentation that answers the question. So, that is a constant stream of questions coming in that you have to answer. So, the inbox is always full [laugh]. So, that is maybe a small downside of a product business. But other than that, yeah, compared to a consulting business, it really gives us many flexibilities with planning our work day around the rest of our lives. That’s really what we enjoy about a product company.

Corey: I was very careful to pick an expensive problem that was only a business-hours problem. So, I don’t wind up with a surprise, middle-of-the-night panic phone call. It’s yeah, it turns out that AWS billing operate during business hours in the US Pacific Time. The end. And there are no emergencies here; there are simply curiosities that will, in the fullness of time take weeks to get resolved.

Andreas: Mmm. Yeah.

Corey: I spent too many years on call, in that sense. Everyone who’s built a product company the first time always says the second time, the engineering? Meh, there are ways to solve that. Solving the distribution problem. That’s the thing I want to focus on next.

And I feel like I sort of went into this backwards in that I don’t really have a product to sell people but I somehow built an audience. And to be honest, it’s partly why. It’s because I didn’t know what I was going to be doing after 18 months and I knew that whatever it was going to be, I needed an audience to tell about it, so may as well start the work of building the audience now. So, I have to imagine if nothing else, your book has been a tremendous source of building a community. When I mentioned the word cloudonaut to people who have been learning AWS, more often than not, they know who you are.

Andreas: Yeah.

Corey: Although I admit they sometimes get you confused with your brother.

Andreas: [laugh]. Yes, that’s not too hard. Yeah, yeah, cloudonaut is definitely—this was always our, also a side project of we was just writing about things that we learned about AWS. Whenever we, I don’t know, for example, looked into a new series, we wrote a blog post about that. Later, we did start a podcast and YouTube videos during the pandemic, of course, as everyone did. And so, I think this was always fun stuff to do. And we like sharing what we learn and getting into discussion with the community, so this is what we still do and enjoy as well, of course. Yeah.

Corey: I really want to thank you for taking the time to catch up and see what you’ve been up to these last few years with a labor of love and the pivot to a product company. If people want to learn more, where’s the best place for them to find you?

Andreas: So definitely, the best place to find me is cloudonaut.io. So, this basically points you to all [laugh] what I do. Yeah, that’s basically the one domain and URL that you need to know.

Corey: Excellent. And we will put that in the show notes, of course. Thank you so much for taking the time to speak with me today. I really appreciate it.

Andreas: Yeah, it was a pleasure to be back here. I’m big fan of podcasts and also of Screaming in the Cloud, of course, so it was a pleasure to be here again.

Corey: [laugh]. You are always welcome. Andreas Wittig, co-author of Amazon Web Services in Action, now up to its third edition. And of course, the voice behind cloudonaut. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry, insulting comment that I will at one point be able to get to just as soon as I find something to deal with your sarcasm on the AWS Marketplace.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

View Details

Brandon Sherman, Cloud Security Engineer at Temporal Technologies Inc., joins Corey on Screaming in the Cloud to discuss his experiences at recent cloud conferences and the ongoing changes in cloud computing. Brandon shares why he enjoyed fwd:cloudsec more than this year’s re:Inforce, and how he’s seen AWS events evolve over the years. Brandon and Corey also discuss how the cloud has matured and why Brandon feels ongoing change can be expected to be the continuing state of cloud. Brandon also shares insights on how his perspective on Google Cloud has changed, and why he’s excited about the future of Temporal.io.

About Brandon

Brandon is currently a Cloud Security Engineer at Temporal Technologies Inc. One of Temporal’s goals is to make our software as reliable as running water, but to stretch the metaphor it must also be clean water. He has stared into the abyss and it stared back, then bought it a beer before things got too awkward. When not at work, he can be found playing with his kids, working on his truck, or teaching his kids to work on his truck.

Links Referenced:

  • Temporal: https://temporal.io/
  • Personal website: https://brandonsherman.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: In the cloud, ideas turn into innovation at virtually limitless speed and scale. To secure innovation in the cloud, you need Runtime Insights to prioritize critical risks and stay ahead of unknown threats. What's Runtime Insights, you ask? Visit sysdig.com/screaming to learn more. That's S-Y-S-D-I-G.com/screaming.

My thanks as well to Sysdig for sponsoring this ridiculous podcast.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’m joined today by my friend who I am disappointed to say I have not dragged on to this show before. Brandon Sherman is a cloud security engineer over at Temporal. Brandon, thank you for finally giving in.

Brandon: Thanks, Corey, for finally pestering me enough to convince me to join. Happy to be here.

Corey: So, a few weeks ago as of this recording—I know that time is a flexible construct when it comes to the podcast production process—you gave a talk at fwd:cloudsec, the best cloud security conference named after an email subject line. Yes, I know re:Inforce also qualifies; this one’s better. Tell me about what you talked about.

Brandon: Yeah, definitely agree on this being the better the two conferences. I gave a talk about how the ground shifts underneath us, kind of touching on how these cloud services that we operate—and I’m mostly experienced in AWS and that’s kind of the references that I can give—but these services work as a contract basis, right? We use their APIs and we don’t care how they’re implemented behind the scenes. At this point, S3 has been rewritten I don’t know how many times. I’m sure that other AWS services, especially the longer-lived ones have gone through that same sort of rejuvenation cycle.

But as a security practitioner, these implementation details that get created are sort of byproducts of, you know, releasing an API or releasing a managed service can have big implications to how you can either secure that service or respond to actions or activities that happen in that service. And when I say actions and activity, I’m kind of focused on, like, security incidents, breaches, your ability to do incident response from that.

Corey: One of the reasons I’ve always felt that cloud providers have been cagey around how the services work under the hood is not because they don’t want to talk about it so much as they don’t want to find themselves committed to certain patterns that are not guaranteed as a part of the definition of the service. So if, “Yeah, this is how it works under the hood,” and you start making plans and architecting in accordance with that and they rebuild the service out from under you like they do with S3, then very often, those things that you depend upon being true could very easily no longer be true. And there’s no announcement around those things.

Brandon: No. It’s very much Amazon is… you know, they’re building a service to meet the needs of their customers. And they’re trying to grow these services as the customers grow along with them. And it’s absolutely within their right to act that way, to not have to tell us when they make a change because in some contexts, right, Amazon’s feature update might be me as a customer a breaking change. And Amazon wants to try and keep that, what they need to tell me, as small as possible, probably not out of malice, but just because there’s a lot of people out there using their services and trying to figure out what they’ve promised to each individual entity through either literal contracts or their API contracts is hard work. And that’s not the job I would want.

Corey: No. It seems like it’s one of those thankless jobs where you don’t get praise for basically anything. Instead, all you get to do is deal with the grim reality that people either view as invisible or a problem.

Brandon: Yeah. It sort of feels like documentation. Everyone wants more and better documentation, but it’s always an auxiliary part of the service creation process. The best documentation always starts out when you write the documentation first and then kind of build backwards from that, but that’s rarely how I’ve seen software get made.

Corey: No. I feel like I left them off the hook, on some level, when we say this, but I also believe in being fair. I think there’s a lot of things that cloud providers get right and by and large, with any of the large cloud providers, they are going to do a better job of securing the fundamentals than you are yourself. I know that that is a controversial statement to some folks who spent way too much time in the data centers, but I stand by it.

Brandon: Yeah, I agree. I’ve had to work in both environments and some of the easiest, best wins in security is just what do I have, so that way I know what I have to protect, what that is there. But even just that asset inventory, that’s the sort of thing that back in the days of data centers—and still today; it was data centers all over the place—to do an inventory you might need to go and send an actual human with an actual clipboard or iPad or whatever, to the actual physical location and hope that they read the labels on hundreds of thousands of servers correctly and get their serial numbers and know what you have. And that doesn’t even tell you what’s running on them, what ports are open, what stuff you have to care about. In AWS, I can run a couple of describe calls or list calls and that forms the backbone of my inventory.

There’s no server that, you know, got built into a wall or lost behind and some long-forgotten migration. A lot of those basic stuff that really, really helps. Not to mention then the user-managed service like S3, you never have to care about patch notes or what an update might do. Plenty of times I’ve, like, hesitated upgrading a software package because I didn’t know what was going to happen. Control Tower, I guess, is kind of an exception to that where you do have to care about the version of your cloud service, but stuff like, yeah, these other services is absolutely right. The undifferentiated heavy lifting it’s taken care of. And hopefully, we always kind of hope that the undifferentiated heavy lifting doesn’t become differentiated and heavy and lands on us.

Corey: So, now that we’ve done the obligatory be nice to cloud providers thing, let’s potentially be a little bit harsher. While you were speaking at fwd:cloudsec, did you take advantage of the fact that you were in town to also attend re:Inforce?

Brandon: I did because I was given a ticket, and I wanted to go see some people who didn’t have tickets to fwd:cloudsec. Yeah, we’ve been nice to cloud providers, but as—I haven’t found I’ve learned a lot from the re:Inforce sessions. They’re all recorded anyway. There’s not even an open call for papers, right, for talking about at a re:Inforce session, “Hey, like, this would be important and fresh or things that I would be wanting to share.” And that’s not the sort of thing that Amazon does with their conferences.

And that’s something that I think would be really interesting to change if there was a more community-minded track that let people submit, not just handpicked—although I suppose any kind of Amazon selection committee is going to be involved, but to pick out, from the community, stories or projects that are interesting that can be, not just have to get filtered through your TAM but something you can actually talk to and say, “Hey, this is something I’d like to talk about. Maybe other people would find it useful.”

Corey: One of the things that I found super weird about re:Inforce this year has been that, in a normal year, it would have been a lot more notable, I think. I know for a fact that if I had missed re:Invent, for example, I would have had to be living in a cave not to see all of the various things coming out of that conference on social media, in my email, in all the filters I put out there. But unless you’re looking for it, you’ve would not know that they had a conference that costs almost as much.

Brandon: Yeah. The re:Invent-driven development cycle is absolutely a real thing. You can always tell in the lead up to re:Invent when there’s releases that get pushed out beforehand and you think, “Oh, that’s cool. I wonder why this doesn’t get a spot at re:Invent, right, some kind of announcement or whatever.” And I was looking for that this year for re:Inforce and didn’t see any kind of announcement or that kind of pre-release trickle of things that are like, oh, there’s a bunch of really cool stuff. And that’s not to say that cool stuff didn’t happen; it just there was a very different marketing feel to it. Hard to say, it’s just the vibes around felt different [laugh].

Corey: Would you recommend that people attend next year—well let me back up. I’ve heard that they had not even announced a date for next year. Do you think there will be a re:Inforce next year?

Brandon: Making me guess, predict the future, something that I’m—

Corey: Yeah, do a prediction. Why not?

Brandon: [laugh]. Let’s engage in some idle speculation, right? I think that not announcing it was kind of a clue that there’s a decent chance it won’t happen because in prior years, it had been pre-announced at the—I think it was either at closing or opening ceremonies. Or at some point. There’s always the, “Here’s what you can look forward to next year.”

And that didn’t happen, so I think that’s there’s a decent chance this may have been the last re:Inforce, especially once all the data is crunched and people look at the numbers. It might just be… I don’t know, I’m not a marketing-savvy kind of person, but it might just be that a day at re:Invent next year is dedicated to security. But then again, security is always job zero at Amazon so maybe re:Invent just becomes re:Inforce all the time, right? Do security, everybody.

Corey: It just feels like a different type of conference. Whenever re:Invent there’s something for everyone. At re:Inforce, there’s something for everyone as long as they work in InfoSec. Because other than that, you wind up just having these really unfortunate spiels of them speaking to people that are not actually present, and it winds up missing the entire forest for the trees, really.

Brandon: I don’t know if I’d characterize it as that. I feel like some of the re:Inforce content was people who were maybe curious about the cloud or making progress in their companies and moving to the cloud—and in Amazon’s case when they say the cloud, they mean themselves. They don’t mean any other cloud. And re:Inforce tries to dispel the notion there are any other clouds.

But at the same time, it feels like an attempt to try and make people feel better. There’s a change underway in the industry and it still is going to continue for a while. There’s still all kinds of non-cloud environments people are going to operate for probably until the end of time. But at the same time, a lot of these are moving to the cloud and they want the people who are thinking about this or engaged in it, to be comforted by that Amazon that either has these services, or there’s a pattern you can follow to do something in a secure manner. I think that’s that was kind of the primary audience of re:Inforce was people who were charged with doing cloud security or were exploring moving their corporate systems to AWS and they wanted some assurance that they’re going to actually be doing things the right way, or someone else hadn’t made those mistakes first. And if that audience has been sort of saturated, then maybe there isn’t a need for that style of conference anymore.

Corey: It feels like it’s not intended to be the same thing at re:Invent, which is probably I guess, a bigger problem. Re:Invent for a long time has attempted to be all things to all people, and it has grown to a scale where that is no longer possible. So, they’ve also done a poor job of signaling that, so you wind up attending Adam Selipsky’s keynote, and in many cases, find yourself bored absolutely to tears. Or you go in expecting it to be an Andy Jassy style of, “Here are 200 releases, four of them good,” and instead, you wind up just having what feels like a relatively paltry number doled out over a period of days. And I don’t know that their wrong to do it; I just think it doesn’t align with pre-existing expectations. I also think people expecting to go to re:Inforce to see a whole bunch of feature releases are bound to be disappointed.

Brandon: Like, both of those are absolutely correct. The number of releases on the slide must always increase up and the right; away we go; we’re pushing more code and making more changes to services. I mean, if you look at the history, there’s always new instance types. Do they count each instance type as a new release, or they not do that?

Corey: Yeah, it honestly feels like that sometimes. They also love to do price cuts where they—you wind up digging into them and something like 90% of them are services you’ve never heard of in regions you couldn’t find on a map if your life depended on it. It’s not quite the, “Yeah, the bill gets lower all the time,” that they’d love to present it as being.

Brandon: Yeah. And you may even find that there’s services that had updates that you didn’t know about until you go and check the final bill, the Cost and Usage Report, and you look and go, “Oh, hey. Look at all the services that we were using, that our engineers started using after they heard announcements at re:Invent.” And then you find out how much you’re actually paying for them. [pause]. Or that they were in use in the first place. There’s no better way to find what is actually happening in your environment than, look at the bill.

Corey: It’s depressing that that’s true. At least they finally stopped doing the slides where they talk about year-over-year, they have a histogram of number of feature and service releases. It’s, no one feels good about that, even the people building the services and features because they look at that and think, “Oh, whatever I do is going to get lost in the noise.” And they’re not wrong. Customers see it and freak out because how am I ever going to keep current with all this stuff? I take a week off and I spend a month getting caught back up again.

Brandon: Yeah. And are you going to—you know, what’s your strategy for dealing with all these new releases and features? Do you want to have a strategy of saying, “No, you can’t touch any of those until we’ve vetted and understand them?” I mean, you don’t even have to talk about security in that context; just the cost alone, understanding it’s someone, someone going to run an experiment that bankrupts your company by forgetting about it or by growing into some monster in the bill. Which I suspect helps [laugh] helps you out when those sorts of things happen, right, for companies don’t have that strategy.

But at the same time, all these things are getting released. There’s not really a good way of understanding which of these do I need to care about. Which of these is going to really impact my operational flow, my security impacts? What does this mean to me as a user of the service when there’s, I don’t know, an uncountable number really, or at least a number that’s so big, it stops mattering that it got any bigger?

Corey: One thing that I will say was great about re:Invent, I want to say 2021, was how small it felt. It felt like really a harkening back to the old re:Invents. And then you know, 2022 hit, and we go there and half of us wound up getting Covid because of course we did. But it was also this just this massive rush of, we’re talking with basically the population of a midsize city just showing up inside of this entire enormous conference. And you couldn’t see the people you wanted to see, it was difficult to pay attention to all there was to pay attention to, and it really feels like we’ve lost something somewhere.

Brandon: Yeah, but at the same time is that just because there are more people in this ecosystem now? You know, 2021 may have been a callback to that a decade ago. And these things were smaller when it was still niche, but growing in kind of the whole ecosystem. And parts of—let’s say, the ecosystem there, I’m talking about like, how—when I say that ecosystem there, I’m kind of talking about how in general, I want to run something in technology, right? I need a server, I need an object store, I need compute, whatever it is that you need, there is more attractive services that Amazon offers to all kinds of customers now.

So, is that just because, right, we’ve been in this for a while and we’ve seen the cloud grow up and like, oh, wow, you’re now in your awkward teenage phase of cloud computing [laugh]? Have we not yet—you know, we’re watching the maturity to adulthood, as these things go? I really don’t know. But it definitely feels a little, uh… feels a little like we’ve watched this cloud thing grow from a half dozen services to now, a dozen-thousand services all operating different ways.

Corey: Part of me really thinks that we could have done things differently, had we known, once upon a time, what the future was going to hold. So, much of the pain I see in Cloud is functionally people trying to shove things into the cloud that weren’t designed with Cloud principles in mind. Yeah, if I was going to build a lot of this stuff from scratch myself, then yeah, I would have absolutely made a whole universe of different choices. But I can’t predict the future. And yet, here we are.

Brandon: Yep. If I could predict the future, I would have definitely won the lottery a lot more times, avoided doing that one thing I regretted that once back in my history [laugh]. Like, knowing the future change a lot of things. But at least unless you’re not letting on with something, then that’s something that no one’s got the ability to, do not even at Amazon.

Corey: So, one of the problems I’ve always had when I come back from a conference, especially re:Invent, it takes me a few… well, I’ll be charitable and say days, but it’s more like weeks, to get back into the flow of my day-to-day work life. Was there any of that with you and re:Inforce? I mean, what is your day job these days anyway? What are you up to?

Brandon: What is my day job? There’s a lot. So, Temporal is a small, but quickly growing company. A lot of really cool customers that are doing really cool things with our technology and we need to build a lot of basics, essentially, making sure that when we grow, that we’re going to kind of grow into our security posture. There’s not anything talking about predicting the future. My prediction is that the company I work for is going to do well. You can hold your analysis on that [laugh].

So, while I’m predicting what the company that I’m working at is going to do well, part of it is also what are the things that I’m going to regret not having in two or three years’ time. So, some baseline cloud monitoring, right? I want that asset inventory across all of our accounts; I want to know what’s going on there. There’s other things that are sort of security adjacent. So, things like DNS records, domain names, a lot of those things where if we can capture this and centralize it early and build it in a way—especially that users are less unhappy about, like, not everyone, for example, is hosting their own—buying their own domains on personal cards and filing for reimbursement, that DNS records aren’t scattered across a dozen different software projects and manipulated in different ways, then that sets us up.

It may not be perfect today, but in a year, year-and-a-half, two years, we have the ability to then say, “Okay, we know what we’re pointing at. What are the dangling subdomains? What are the things that are potential avenues of being taken over? What do we have? What are people doing?” And trying to understand how we can better help users with their needs day-to-day.

Also as a side part of my day job is advising a startup Common Fate. Does just-in-time access management. And that’s been a lot of fun to do as well because fundamentally—this is maybe a hot take—that, in a lot of cases, you really only need admin access and read-only access when you’re doing really intensive work. In Temporal day job, we’ve got infrastructure teams that are building stuff, they need lots of permissions and it’d be very silly to say you can’t do your job just because you could potentially use IAM and privilege escalate yourself to administrator. Let’s cut that out. Let’s pretend that you are a responsible adult. We can monitor you in other ways, we’re not going to put restrictions between you and doing your job. Have admin access, just only have it for a short period of time, when you say you’re going to need it and not all the time, every account, every service, all the time, all day.

Corey: I do want to throw a shout-in for that startup you advise, Common Fate. I’ve been a big fan of their Granted offering for a while now. granted.dev for those who are unfamiliar. I use that to automatically generate console logins, do all kinds of other things. When you’re moving between a bunch of different AWS accounts, which it kind of feels like people building the services don’t have to do somehow because of their Isengard system handling it for them. Well, as a customer, can I just say that experience absolutely sucks and Granted goes a long way toward making it tolerable, if not great.

Brandon: Mm-hm. Yeah, I remember years ago, the way that I would have to handle this is I would have probably a half-dozen different browsers at the same time, Safari, Chrome, the Safari web developer preview, just so I could have enough browsers to log into with, to see all the accounts I needed to access. And that was an extremely painful experience. And it still feels so odd that the AWS console today still acts like you have one account. You can switch roles, you can type in a [role 00:21:23] on a different account, but it’s very clunky to use, and having software out there that makes this easier is definitely, definitely fills a major pain point I have with using these services.

Corey: Tired of Apache Kafka's complexity making your AWS bill look like a phone number? Enter Redpanda. You get 10x your streaming data performance without having to rob a bank. And migration? Smoother than a fresh jar of peanut butter. Imagine cutting as much as 50% off your AWS bills. With Redpanda, it's not a dream, it's reality. Visit go.redpanda.com/duckbill. Redpanda: Because Kafka shouldn't cause you nightmares.

Corey: Do you believe that there’s hope? Because we have seen some changes where originally AWS just had the AWS account you’d log into, it’s the root user. Great. Then they had IAM. Now, they’re using what used to be known as AWS SSO, which they wound up calling IAM Access Identity Center, or—I forget the exact words they put in order, but it’s confusing and annoying. But it does feel like the trend is overall towards something that’s a little bit more coherent.

Brandon: Mm-hm.

Corey: Is the future five years from now better than it looks like today?

Brandon: That’s certainly the hope. I mean, we’ve talked about how we both can’t predict the future, but I would like to hope that the future gets better. I really like GCP’s project model. There’s complaints I have with how Google Cloud works, and it’s going to be here next year, and if the permission model is exactly how I’d like to use it, but I do like the mental organization that feels like Google was able to come in and solve a lot of those problems with running projects and having a lot of these different things. And part of that is, there’s still services in AWS that don’t really respect resource-based permissions or tag-based permissions, or I think the new one is attribute-based access control.

Corey: One of the challenges I see, too, is that I don’t think that there’s been a lot of thought put into how a lot of these things are going to work between different AWS accounts. One of my bits of guidance whenever I’m talking to someone who’s building anything, be it at AWS or external is, imagine an architecture diagram and now imagine that between any two resources in that diagram is now an account boundary. Because someone somewhere is going to have one there, so it sounds ridiculous, but you can imagine a microservices scenario where every component is in its own isolated account. What are you going to do now as a result? Because if you’re going to build something that scales, you’ve got to respect those boundaries. And usually, that just means the person starts drinking.

Brandon: Not a bad place to start, the organizational structure—lowercase organizations, not the Amazon service, Organizations—it’s still a little tricky to get it in a way that sort of… I guess, I always kind of feel that these things are going to change and that the—right, the only constant is change. That’s true. The services we use are going to change. The way that we’re going to want to organize them is going to change. Our researcher is going to come out with something and say, “Hey, I found a really cool way to do something really terrible to the stuff in your cloud environment.”

And that’s going to happen eventually, in the fullness of time. So, how do we be able to react quickly to those kinds of changes? And how can we make sure that if you know, suddenly, we do need to separate out these services to go, you know, to decompose the monolith even more, or whatever the cool, current catchphrase is, and we have those account boundaries, which are phenomenal boundaries, they make it so much easier to do—if you can do multi-account then you’ve solved multi-regional on the way, you’ve sold failover, you’ve solve security issues. You have not solved the fact that your life is considerably more challenging at the moment, but I would really hope that in you know, even next year, but by the time five years comes around, that that’s really been taken to heart within Amazon and it’s a lot easier to be working creating services in different accounts that can talk to each other, especially in the current environment where it’s kind of a mess to wire these things all together. ClickOps has its place, but some console applications just don’t want to believe that you have a KMS key in another account because well, why would you put that over there? It’s not like if your current account has a problem, you want to lose all your data that’s encrypted.

Corey: It’s one of those weird things, too, where the clouds almost seem to be arguing against each other. Like, I would be hard-pressed to advise someone not to put a ‘rehydrate the entire business’ level of backups into a different cloud provider entirely, but there’s so steeped in the orthodoxy of no other clouds ever, that that message is not something that they can effectively communicate. And I think they’re doing their customers a giant disservice by that, just because it is so much easier to explain to your auditor that you’ve done it than to explain why it’s not necessary. And it’s never true; you always have the single point of failure of the payment instrument, or the contract with that provider that could put things at risk.

Is it a likely issue? No. But if you’re running a publicly traded company on top of it, you’d be negligent not to think about it that way. So, why pretend otherwise?

Brandon: Is that a question for me because [laugh]—

Corey: Oh, that was—no, absolutely. That was a rant ending in a rhetorical question. So, don’t feel you have to answer it. But getting the statement out there because hopefully, someone at Amazon is listening to this.

Brandon: That’s, uh, hopefully, if you find out who’s the one that listens to this and can affect it, then yeah, I’d like to send them a couple of emails because absolutely. There’s room out there, there will always be room for at least two providers.

Corey: Yeah, I’d say a third, but I don’t know that Google is going to have the attention span to still have a cloud offering by lunchtime today.

Brandon: Yeah. I really wish that I had more faith in the services and that they weren’t going—you know, speaking of services changing underneath you, that’s definitely a—speaking of services changing underneath, you definitely a major disservice if you don’t know—if you’re going to put into work into architecting and really using cloud providers as they’re meant to be used. Not in a, sort of, least common denominator sense, in which case, you’re not in good shape.

Corey: Right. You should not be building something with an idea toward what if this gets deprecated. You shouldn’t have to think about that on a consistent basis.

Brandon: Mm-hm. Absolutely. You should expect those things to change because they will, right, the performance impact. I mean, the performance of these services is going to change, the underlying technology that the providers use is going to change, but you should still be able to mostly expect that at least the API calls you make are going to still be there and still be consistent come this time next year.

Corey: The thing that really broke me was the recent selling off of Google domains to Squarespace. Nothing against Squarespace, but they have a different target market in many respects. And oh, I’m a Google customer, you’re now going to give all of my information to a third party I never asked to deal with. Great. And more to the point, if I recommend Google to folks because as has happened in years past, then they canceled the thing that I recommended, then I looked like a buffoon. So, we’ve gotten to a point now where it has become so steady and so consistent, that I fear I cannot, in good conscience, recommend a Google product without massive caveats. Otherwise, I look like a clown or worse, a paid shill.

Brandon: Yeah. And when you want to start incorporating these things into the core of your business, to take that point about, you know, total failover scenarios, you should, you know, from you want it to have a domain registered in a Google service that was provisioned to Google Cloud services, that whole sort of ecosystem involved there, that’s now gone, right? If I want to use Google Cloud with a Google Cloud native domain name hosting services, I can’t. How am—I just—now I can’t [laugh]. There’s, like, not workarounds available.

I’ve got to go to some other third-party and it just feels odd that an organization would sort of take those core building blocks and outsource them. [I know 00:29:05] that Google’s core offering isn’t Google Cloud; it’s not their primary focus, and it kind of reflects that, which was a shame. There’s things that I’d love to see grow out of Google Cloud and get better. And, you know, competition is good for the whole cloud computing industry.

Corey: I think that it’s a sad thing, but it’s real, that there are people who were passionate defenders of Google over the years. I used to be one. We saw a bunch of them with Stadia fans coming out of the woodwork, and then all those people who have defended Google and said, “No, no, you can trust Google on this service because it’s different,” for some reason or other, then wind up looking ridiculous. And some of the staunchest Google defenders that I’ve seen are starting to come around to my point of view. Eventually, you’ve run out of people who are willing to get burned if you burn them all.

Brandon: Yeah. I’ve always been a little, uh… maybe this is the security Privacy part of me; I’ve always been a little leery of the services that really want to capture and gather your data. But I always respected the Google engineering that went into building these things at massive scale. It’s something beyond my ability to understand as I haven’t worked in something that big before. And Google made it look… maybe not effortless, but they made it look like they knew what they were doing, they could build something really solid.

And I don’t know if that’s still true because it feels like they might know how to build something, and then they’ll just dismantle it and turn it over to somebody else, or just dismantle it completely. And I think humans, we do a lot of things because we don’t want to look foolish and… now recommending Google Cloud starts to make you wonder, “Am I going to look foolish?” Is this going to be a reflection on me in a year or two years, when you got to come in to say, “Hey, I guess that whole thing we architected around, it’s being sold to someone else. It’s being closed down. We got to transfer and rearchitect our whole whatever we built because of factors out of our control.” I want to be rearchitecting things because I screwed it up. I want to be rearchitecting things because I made an interesting novel mistake, not something that’s kind of mundane, like, oh, I guess the thing we were going to use got shut down. Like, that makes it look like not only can I not predict the future, but I can’t even pretend to read the tea leaves.

Corey: And that’s what’s hard is because, on some level, our job, when we work in operations and cloud and try and make these decisions, is to convince the business we know what we’re talking about. And when we look foolish, we don’t make that same mistake again.

Brandon: Mm-hm. Billing and security are oftentimes frequently aligned with each other. We’re trying to convince the business that we need to build things a certain way to get a certain outcome, right? Either lower costs or more performance for the dollar, so that way, we don’t wind up in the front page of newspapers, any kinds of [laugh] any kind of those things.

Corey: Oh, yes. I really want to thank you for taking the time to speak with me. If people want to learn more, where’s the best place for them to find you?

Brandon: The best place to find me, I have a website about me, [brandonsherman.com 00:32:13]. That’s where I post stuff. There’s some links to—I have a [Mastodon 00:32:18] profile. I’m not much of a social, sort of post your information out there kind of person, but if you want to get a hold of me, then that’s probably the best way to find me and contact me. Either that or head out to the desert somewhere, look for a silver truck out in the dunes and without technology around. It’s another good spot if you can find me there.

Corey: And I will include a link to that, of course, in the [show notes 00:32:45]. Thank you so much for taking the time to speak with me today. As always, I appreciate it.

Brandon: Thank you very much for having me, Corey. Good to chat with you.

Corey: Brandon Sherman, cloud security engineer at Temporal. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry comment that will somehow devolve into you inviting me to your new uninspiring cloud security conference that your vendor is putting on, and is of course named after an email subject line.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

View Details

Jonathan (Koz) Kozolchyk, General Manager for Certificate Services at AWS, joins Corey on Screaming in the Cloud to discuss the best practices he recommends around certificates. Jonathan walks through when and why he recommends private certs, and the use cases where he’d recommend longer or unusual expirations. Jonathan also highlights the importance of knowing who’s using what cert and why he believes in separating expiration from rotation. Corey and Jonathan also discuss their love of smart home devices as well as their security concerns around them and how they hope these concerns are addressed moving forward.

About Jonathan

Jonathan is General Manager of Certificate Services for AWS, leading the engineering, operations, and product management of AWS certificate offerings including AWS Certificate Manager (ACM) AWS Private CA, Code Signing, and Encryption in transit. Jonathan is an experienced leader of software organizations, with a focus on high availability distributed systems and PKI. Starting as an intern, he has built his career at Amazon, and has led development teams within our Consumer and AWS businesses, spanning from Fulfillment Center Software, Identity Services, Customer Protection Systems and Cryptography. Jonathan is passionate about building high performing teams, and working together to create solutions for our customers. He holds a BS in Computer Science from University of Illinois, and multiple patents for his work inventing for customers. When not at work you’ll find him with his wife and two kids or playing with hobbies that are hard to do well with limited upside, like roasting coffee.

Links Referenced:

  • AWS website: https://www.aws.com
  • Email: mailto:koz@amazon.com
  • Twitter: https://twitter.com/seakoz

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: In the cloud, ideas turn into innovation at virtually limitless speed and scale. To secure innovation in the cloud, you need Runtime Insights to prioritize critical risks and stay ahead of unknown threats. What's Runtime Insights, you ask? Visit sysdig.com/screaming to learn more. That's S-Y-S-D-I-G.com/screaming.

My thanks as well to Sysdig for sponsoring this ridiculous podcast.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. As I record this, we are about a week and a half from re:Inforce in Anaheim, California. I am not attending, not out of any moral reason not to because I don’t believe in cloud security or conferences that Amazon has that are named after subject lines, but rather because I am going to be officiating a wedding on the other side of the world because I am an ordained minister of the Church of There Is A Problem With This Website’s Security Certificate. So today, my guest is going to be someone who’s a contributor, in many ways, to that religion, Jonathan Kozolchyk—but, you know, we all call him Koz—is the general manager for Certificate Services at AWS. Koz, thank you for joining me.

Koz: Happy to be here, Corey.

Corey: So, one of the nice things about ACM historically—the managed service that handles certificates from AWS—is that for anything public-facing, it’s free—which is always nice, you should not be doing upcharges for security—but you also don’t let people have the private portion of the cert. You control all of the endpoints that terminate SSL. Whereas when I terminate SSL myself, it terminates on the floor because I’ve dropped things here and there, which means that suddenly the world of people exposing things they shouldn’t or expiry concerns just largely seemed to melt away. What was the reason that Amazon looked around at the landscape and said, “Ah, we’re going to launch our own certificate service, but bear with me here, we’re not going to charge people money for it.” It seems a little bit out of character.

Koz: Well, Amazon itself has been battling with certificates for years, long before even AWS was a thing, and we learned that you have to automate. And even that’s not enough; you have to inspect and you have to audit, you need a controlled loop. And we learned that you need a closed loop to truly manage it and make sure that you don’t have outages. And so, when we built ACM, we built it saying, we need to provide that same functionality to our customers, that certificates should not be the thing that makes them go out. Is that we need to keep them available and we need to minimize the sharp edges customers have to deal with.

Corey: I somewhat recently caught some flack on one of the Twitter replacement social media sites for complaining about the user experience of expired SSL certs. Because on the one hand, if I go to my bank’s website, and the response is that instead, the server is sneakyhackerman.com, it has the exact same alert and failure mode as, holy crap, this certificate reached its expiry period 20 minutes ago. And from my perspective, one of those is a lot more serious than the other. What also I wind up encountering is not just when I’m doing banking, but when I’m trying to read some random blog on how to solve a technical problem. I’m not exactly putting personal information into the thing. It feels like that was a missed opportunity, agree or disagree?

Koz: Well, I wouldn’t categorize it as a missed opportunity. I think one of the things you have to think about with security is you have to keep it simple so that everyone, whether they’re a technologist or not, can abide by the rules and be safe. And so, it’s much easier to say to somebody, “There’s something wrong. Period. Stop.” versus saying there are degrees of wrongness. Now, that said, boy, do I wish we had originally built PKI and TLS such that you could submit multiple certificates to somebody, in a connection for example, so that you could always say, you know, my certificates can expire, but I’ve got two, and they’re off by six months, for example. Or do something so that you don’t have to close failed because the certificate expired.

Corey: It feels like people don’t tend to think about what failure modes are going to look like. Because, pfhh, as an expired certificate? What kind of irresponsible buffoon would do such a thing? But I’ve worked in enough companies where you have historically, the wildcard cert because individual certs cost money, once upon a time. So, you wound up getting the one certificate that could work on all of the stuff that ends in the same domain.

And that was great, but then whenever it expired, you had to go through and find all the places that you put it and you always miss some, so things would break for a while and the corporate response was, “Ugh, that was awful. Instead of a one-year certificate, let’s get a five-year or a ten-year certificate this time.” And that doesn’t make the problem better; it makes it absolutely worse because now it proliferates forever. Everyone who knows where that thing lives is now long gone by the time it hits again. Counterintuitively, it seems the industry has largely been moving toward short-lived certs. Let’s Encrypt, for example, winds up rotating every 90 days, by my estimation. ACM is a year, if memory serves.

Koz: So, ACM certs are 13 months, and we start rotating them around the 11th month. And Let’s Encrypt offers you 90-day certs, but they don’t necessarily require you to rotate every 90 days; they expire in 90 days. My tip for everybody is divorce expiration from rotation. So, if your cert is a 90-day cert, rotate it at 45 days. If your cert is a year cert, give yourself a couple of months before expiration to start the rotation. And then you can alarm on it on your own timeline when something fails, and you still have time to fix it.

Corey: This makes a lot of sense in—you know, the second time because then you start remembering, okay, everywhere I use this cert, I need to start having alarms and alerts. And people are bad at these things. What ACM has done super well is that it removes that entire human from the loop because you control all of the endpoints. You folks have the ability to rotate it however often you’d like. You could have picked arbitrary timelines of huge amounts of time or small amounts of time and it would have been just fine.

I mean, you log into an EC2 instance role and I believe the credentials get passed out of either a 6 or a 12-hour validity window, and they’re consistently rotating on the back end and it’s completely invisible to the customer. Was there ever thought given to what that timeline should be,j what that experience should be? Or did you just, like, throw a dart at a wall? Like, “Yeah, 13 months feels about right. We’re going to go with that.” And never revisited it. I have a guess which—

Koz: [laugh].

Corey: Side of that it was. Did you think at all about what you were doing at the time, or—yeah.

Koz: So, I will admit, this happened just before I got there. I got to ACM after—

Corey: Ah, blame the predecessor. Always a good call.

Koz: —the launch. It’s a God-given right to blame your predecessor.

Corey: Oh, absolutely. It’s their entire job.

Koz: I think they did a smart job here. What they did was they took the longest lifetime cert that was then allowed, at 13 months, knowing that we were going to automate the rotation and basically giving us as much time as possible to do it, right, without having to worry about scaling issues or having to rotate overly frequently. You know, there are customers who while I don’t—I strongly disagree with [pinning 00:07:35], for example, but there are customers out there who don’t like certs to change very often. I don’t recommend pinning at all, but I understand these cases are out there, and changing it once every year can be easier on customers than changing it every 20 minutes, for example. If I were to pick an ideal rotation time, it’d probably be under ten days because an OCSP response is good for ten days and if you rotate before, then I never have to update an OCSP response, for example. But changing that often would play havoc with many systems because of just the sheer frequency you’re rotating what is otherwise a perfectly valid certificate.

Corey: It is computationally expensive to generate certificates at scale, I would imagine.

Koz: It starts to be a problem. You’re definitely putting a lot of load on the HSMs at that point, [laugh] when you’re generating. You know, when you have millions of certs out in deployment, you’re generating quite a few at a time.

Corey: There is an aspect of your service that used to be part of ACM and now it’s its own service—which I think is probably the right move because it was confusing for a lot of customers—Amazon looks around and sees who can we compete with next, it feels like sometimes. And it seemed like you were squarely focused on competing against your most desperate of all enemies, my crappy USB key where I used to keep the private CA I used at any given job—at the time; I did not keep it after I left, to be very clear—for whatever I’m signing things for certificates for internal use. You’re, like, “Ah, we can have your crappy USB key as a service.” And sure enough, you wound up rolling that out. It seems like adoption has been relatively brisk on that, just because I see it in almost every client account I work with.

Koz: Yeah. So, you’re talking about the private CA offering which is—

Corey: I—that’s right. Private CA was the new service name. Yes, it used to be a private certificate authority was an aspect of ACM, and now you’re—mmm, we’re just going to move that off.

Koz: And we split it out because like you said customers got confused. They thought they had to only use it with ACM. They didn’t understand it was a full standalone service. And it was built as a standalone service; it was not built as part of ACM. You know, before we built it, we talked to customers, and I remember meeting with people running fairly large startups, saying, “Yes, please run this for me. I don’t know why, but I’ve got this piece of paper in my sock drawer that one of my security engineers gave me and said, ‘if something goes wrong with our CA, you and two other people have to give me this piece of paper.’” And others were like, “Oh, you have a piece of paper? I have a USB stick in my sock drawer.” And like, this is what, you know, the startup world was running their CAs from sock drawers as far as I can tell.

Corey: Yeah. A piece of paper? Someone wrote out the key by hand? That sounds like hell on earth.

Koz: [sigh]. It was a sharding technique where you needed, you know, three of five or something like that to—

Corey: Oh, they, uh, Shamir’s Secret Sharing Service.

Koz: Yes.

Corey: The SSSS. Yeah.

Koz: Yes. You know, and we looked at it. And the other alternative was people would use open-source or free certificate authorities, but without any of the security, you’d want, like, HSM backing, for example, because that gets really expensive. And so yeah, we did what our customers wanted: we built this service. We’ve been very happy with the growth it’s taken and, like you said, we love the places we’ve seen it. It’s gone into all kinds of different things, from the traditional enterprise use cases to IoT use cases. At one point, there’s a company that tracks sheep and every collar has one of our certs in it. And so, I am active in the sheep-tracking industry.

Corey: I am certain that some wit is going to comment on this. “Oh, there’s a company out there that tracks sheep. Yeah, it’s called Apple,” or Facebook, or whatever crappy… whatever axe someone has to grind against any particular big company. But you’re talking actual sheep as in baa, smell bad, count them when going to sleep?

Koz: Yes. Actual sheep.

Corey: Excellent, excellent.

Koz: The certs are in drones, they’re in smart homes, so they’re everywhere now.

Corey: That is something I want to ask you about because I found that as a competition going on between your service, ACM because you won’t give me the private keys for reasons that we already talked about, and Let’s Encrypt. It feels like you two are both competing to not take my money, which is, you know, an odd sort of competition. You’re not actually competing, you’re both working for a secure internet in different ways, but I wind up getting certificates made automatically for me for all of my internal stuff using Let’s Encrypt, and with publicly resolvable domain names. Why would someone want a private CA instead of an option that, okay, yeah, we’re only using it internally, but there is public validity to the certificate?

Koz: Sure. And just because I have to nitpick, I wouldn’t say we’re competing with them. I personally love Let’s Encrypt; I use them at home, too. Amazon supports them financially; we give them resources. I think they’re great. I think—you know, as long as you’re getting certs I’m happy. The world is encrypted and I—people use private CA because fundamentally, before you get to the encryption, you need secure identity. And a certificate provides identity. And so, Let’s Encrypt is great if you have a publicly accessible DNS endpoint that you can prove you own and get a certificate for and you’re willing to update it within their 90-day windows. Let’s use the sheep example. The sheep don’t have publicly valid DNS endpoints and so—

Corey: Or to be very direct with you, they also tend to not have terrific operational practices around updating their own certificates.

Koz: Right. Same with drones, same with internal corporate. You may not want your DNS exposed to the internet, your internal sites. And so, you use a private certificate where you own both sides of the connection, right, where you can say—because you can put the CA in the trust store and then that gets you out of having to be compliant with the CA browser form and the web trust rules. A lot of the CA browser form dictates what a public certificate can and can’t do and the rules around that, and those are built very much around the idea of a browser connecting to a client and protecting that user.

Corey: And most people are not banking on a sheep.

Koz: Most people are not banking on a sheep, yes. But if you have, for example, a database that requires a restart to pick up a new cert, you’re not going to want to redo that every 90 days. You’re probably going to be fine with a five-year certificate on that because you want to minimize your downtime. Same goes with a lot of these IoT devices, right? You may want a thousand-year cert or a hundred-year cert or cert that doesn’t expire because this is a cert that happens at—that is generated at creation for the device. And it’s at birth, the machine is manufactured and it gets a certificate and you want it to live for the life of that device.

Or you have super-secret-project.internal.mycompany.com and you don’t want a publicly visible cert for that because you’re not ready to launch it, and so you’ll start with a private cert. Really, my advice to customers is, if you own both pieces of the connection, you know, if you have an API that gets called by a client you own, you’re almost always better off with a private certificate and managing that trust store yourself because then you are subject not to other people’s rules, but the rules that fit the security model and the threat assessment you’ve done.

Corey: For the publication system for my newsletter, when I was building it out, I wanted to use client certificates as a way of authenticating that it was me. Because I only have a small number of devices that need to talk to this thing; other people don’t, so how do I submit things into my queue and manage it? And back in those ancient days, the API Gateways didn’t support TLS authentication. Now, they do. I would redo it a bunch of different ways. They did support API key as an authentication mechanism, but the documentation back then was so terrible, or I was so new to this stuff, I didn’t realize what it was and introduced it myself from first principles where there’s a hard-coded UUID, and as long as there’s the right header with that UUID, I accept it, otherwise drop it on the floor. Which… there are probably better ways to do that.

Koz: Sure. Certificates are, you know, a very popular way to handle that situation because they provide that secure identity, right? You can be assured that the thing connecting to you can prove it is who they say they are. And that’s a great use of a private CA.

Corey: Changing gears slightly. As we record this, we are about two weeks before re:Inforce, but I will be off doing my own thing on that day. Anything interesting and exciting coming out of your group that’s going to be announced, with the proviso, of course, that this will not air until after re:Inforce.

Koz: Yes. So, we are going to be pre-announcing the launch of a connector for Active Directory. So, you will be able to tie your private CA instance to your Active Directory tree and use private CA to issue certificates for use by Active Directory for all of your Windows hosts for the users in that Active Directory tree.

Corey: It has been many years since I touched Windows in anger, but in 2003 or so, I was a mediocre Small Business Windows Server Admin. Doesn’t Active Directory have a private CA built into it by default for whenever you’re creating a new directory?

Koz: It does.

Corey: Is that one of the FSMO roles? I’m trying to remember offhand.

Koz: What’s a Fimal?

Corey: FSMO. F-S-M-O. There are—I forget, it’s some trivia question that people love to haze each other with in Microsoft interviews. “What are the seven FSMO roles?” At least back then. And have to be moved before you decommission a domain controller or you’re going to have tears before bedtime.

Koz: Ah. Yeah, so Microsoft provides a certificate authority for use with Active Directory. They’ve had it for years and they had to provide it because back then nobody had a certificate authority, but AD needed one. The difference here is we manage it for you. And it’s backed by HSMs. We ensure that the keys are kept secure. It’s a serverless connection to your Active Directory tree, you don’t have to run any software of ours on your hosts. We take care of all of it.

And it’s been the top requests from customers for years now. It’s been quite [laugh] a bit of effort to build it, but we think customers are going to love it because they’re going to get all the security and best practices from private CA that they’re used to and they can decommission their on-prem certificate authority and not have to go through the hassle of running it.

Corey: A big area where I see a lot of private CA work has been in the realm of desktops for corporate environments because when you can pass out your custom trusted root or trusted CA to all of the various nodes you have and can control them, it becomes a lot easier. I always tended to shy away from it, just because in small businesses like the one that I own, I don’t want to play corporate IT guy more than I absolutely have to.

Koz: Yeah. Trust or management is always a painful part of PKI. As if there weren’t enough painful things in PKI. Trust store management is yet another one. Thankfully, in the large enterprises, there are good tooling out there to help you manage it for the corporate desktops and things like that.

And with private CA, you can also, if you already have an offline root that is in all of your trust stores in your enterprise, you can cross-sign the route that we give you from private CA into that hierarchy. And so, then you don’t have to distribute a new trust store out if you don’t want to.

Corey: This is a tricky release and I’m very glad I’m taking the week off it’s getting announced because there are two reactions that are going to happen to any snarking I can do about this. The first is no one knows what the hell this is and doesn’t have any context for the rest, and the other folks are going to be, “Yes, shut up clown. This is going to change my workflow in amazing ways. I’ll deal with your nonsense later. I want to do this.” And I feel like one of those constituencies is very much your target market and the other isn’t. Which is fine. No service that AWS offers—except the bill—is for every customer, but every service is for someone.

Koz: That’s right. We’ve heard from a lot of our customers, especially as they—you know, the large international ones, right, they find themselves running separate Active Directory CAs in different countries because they have different regulatory requirements and separations that they want to do. They are chomping at the bit to get this functionality because we make it so easy to run a private CA in these different regions. There’s certainly going to be that segment at re:Inforce, that’s just happy certificates happen in the background and they don’t think anything about where they come from and this won’t resonate with them, but I assure you, for every one of them, they have a colleague somewhere else in the building that is going to do a happy dance when this launches because there’s a great deal of customer heavy-lifting and just sharp edges that we’re taking away from them. And we’ll manage it for them, and they’re going to love it.

[midroll 0:21:08]

Corey: One thing that I have seen the industry shift to that I love is the Let’s Encrypt model, where the certificate expires after 90 days. And I love that window because it is a quarter, which means yes, you can do the crappy thing and have a calendar reminder to renew the thing. It’s not something you have to do every week, so you will still do it, but you’re also not going to love it. It’s just enough friction to inspire people to automate these things. And that I think is the real win.

There’s a bunch of things like Certbot, I believe the protocol is called ACME A-C-M-E, always in caps, which usually means an acronym or someone has their caps lock key pressed—which is of course cruise control for cool. But that entire idea of being able to have a back-and-forth authentication pass and renew certificates on a schedule, it’s transformative.

Koz: I agree. ACM, even Amazon before ACM, we’ve always believed that automation is the way out of a lot of this pain. As you said earlier, moving from a one-year cert to a five-year cert doesn’t buy you anything other than you lose even more institutional knowledge when your cert expires. You know, I think that the move to further automation is great. I think ACME is a great first step.

One of the things we’ve learned is that we really do need a closed loop of monitoring to go with certificate issuance. So, at Amazon, for example, every cert that we issue, we also track and the endpoints emit metrics that tell us what cert they’re using. And it’s not what’s on disk, it’s what’s actually in the endpoint and what they’re serving from memory. And we know because we control every cert issued within the company, every cert that’s in use, and if we see a cert in use that, for example, isn’t the latest one we issued, we can send an alert to the team that’s running it. Or if we’ve issued a cert and we don’t see it in use, we see the old ones still in use, we can send them an alert, they can alarm and they can see that, oh, we need to do something because our automation failed in this case.

And so, I think ACME is great. I think the push Let’s Encrypt did to say, “We’re going to give you a free certificate, but it’s going to be short-lived so you have to automate,” that’s a powerful carrot and stick combination they have going, and I think for many customers Certbot’s enough. But you’ll see even with ACM where we manage it for our customers, we have that closed loop internally as well to make sure that the cert when we issue a new cert to our client, you know, to the partner team, that it does get picked up and it does get loaded. Because issuing you a cert isn’t enough; we have to make sure that you’re actually using the new certificate.

Corey: I also have learned as a result of this, for example, that AWS certificate manager—Amazon Certificate Manager, the ACM, the certificate thingy that you run, that so many names, so many acronyms. It’s great—but it has a limit—by default—of 2500 certificates. And I know this because I smacked into it. Why? I wasn’t sitting there clicking and adding that many certificates, but I had a delightful step function pattern called ‘The Lambda invokes itself.’ And you can exhaust an awful lot of resources that way because I am bad at programming. That is why for safety, I always recommend that you iterate development-wise in an account that is not production, and preferably one that belongs to someone else.

Koz: [laugh]. We do have limits on cert issuance.

Corey: You have limits on everything in AWS. As it should because it turns out that whatever there’s not a limit, A, free database just dropped, and B, things get hammered to death. You have to harden these things. And it’s one of those things that’s obvious once you’ve operated at a certain point of scale, but until you do, it just feels arbitrary and capricious. It’s one of those things where I think Amazon is still—and all the cloud companies who do this—are misunderstood.

Koz: Yeah. So, in the case of the ACM limits, we look at them fairly regularly. Right now, they’re high enough that most of our customers, vast majority, never come close to hitting it. And the ones that do tend to go way over.

Corey: And it’s been a mistake, as in my case as well. This was not a complaint, incidentally. It was like, well, I want to wind up having more waste and more ridiculous nonsense. It was not my concern.

Koz: No no no, but we do, for those customers who have not mistake use cases but actual use cases where they need more, we’re happy to work with their account teams and with the customer and we can up those limits.

Corey: I’ve always found that limit increases, with remarkably few exceptions, the process is, “Explain to you what your use case is here.” And I feel like that is a screen for, first, are you doing something horrifying for which there’s a better solution? And two, it almost feels like it’s a bit of a customer research approach where this is fine for most customers. What are you folks doing over there and is there a use case we haven’t accounted for in how we use the service?

Koz: I always find we learned something when we look at the [P100 00:26:05] accounts that they use the most certificates, and how they’re operating.

Corey: Every time I think I’ve seen it all on AWS, I just talk to one more customer, and it’s back to school I go.

Koz: Yep. And I thank them for that education.

Corey: Oh, yeah. That is the best part of working with customers and honestly being privileged enough to work with some of these things and talk to the people who are building really neat stuff. I’m just kibitzing from the sideline most of the time.

Koz: Yeah.

Corey: So, one last topic I want to get into before we call it a show. You and I have been talking a fair bit, out of school, for lack of a better term, around a couple of shared interests. The one more germane to this is home automation, which is always great because especially in a married situation, at least as I am and I know you are as well, there’s one partner who is really into home automation and the other partner finds himself living in a haunted house.

Koz: [laugh]. I knew I had won that battle when my wife was on a work trip and she was in a hotel and she was talking to me on the phone and she realized she had to get out of bed to turn the lights off because she didn’t have our Alexa Good Night routine available to her to turn all the lights off and let her go to bed. And so, she is my core customer when I do the home automation stuff. And definitely make sure my use cases and my automations work for her. But yeah, I’m… I love that space.

Coincidentally, it overlaps with my work life quite a bit because identity in smart home is a challenge. We’re really excited about the Matter standard. For those listening who aren’t sure what that is, it’s a new end-all be-all smart home standard for defining devices in a protocol-independent way that lets your hubs talk to devices without needing drivers from each company to interact with them. And one of the things I love about it is every device needs a certificate to identify it. And so, private CA has been a great partner with Matter, you know, it goes well with it.

In fact, we’re one of the leading certificate authorities for Matter devices. Customers love the pricing and the way they can get started without talking to anybody. So yeah, I’m excited to see, you know, as a smart home junkie and as a PKI guy, I’m excited to see Matter take off. Right now I have a huge amalgamation of smart home devices at home and seeing them all go to Matter will be wonderful.

Corey: Oh, it’s fantastic. I am a little worried about aspects of this, though, where you have things that get access to the internet and then act as a bridge. So suddenly, like, I have a IoT subnet with some controls on it for obvious reasons and honestly, one of the things I despise the most in this world has been the rise of smart TVs because I just want you to be a big dumb screen. “Well, how are you going to watch your movies?” “With the Apple TV I’ve plugged into the thing. I just want you to be a screen. That’s it.” So, I live a bit in fear of the day where these things find alternate ways to talk to the internet and, you know, report on what I’m watching.

Koz: Yeah, I think Matter is going to help a lot with this because it’s focused on local control. And so, you’ll have to trust your hub, whether that’s your TV or your Echo device or what have you, but they all communicate securely amongst themselves. They use certificates for identification, and they’re building into Matter a robust revocation mechanism. You know, in my case at home, my TV’s not connected to the internet because I use my Fire TV to talk to it, similar to your Apple TV situation. I want a device I control not my TV, doing it. I’m happy with the big dumb screen.

And I think, you know, what you’re going to end up doing is saying there’s a device out there you’ll trust maybe more than others and say, “That’s what I’m going to use as my hub for my Matter devices and that’s what will speak to the internet,” and otherwise my Matter devices will talk directly to my hub.

Corey: Yeah, there’s very much a spectrum of trust. There’s the, this is a Linux distribution on a computer that I installed myself and vetted and wound up contributing to at one point on the one end of the spectrum, and the other end of the spectrum of things you trust the absolute least in this world, which are, of course, printers. And most things fall somewhere in between.

Koz: Yes, right, now, it is a Wild West of rebranded white-label applications, right? You have all kinds of companies spitting out reference designs as products and white labeling the control app for it. And so, your phone starts collecting these smart home applications to control each one of these things because you buy different switches from different people. I’m looking forward to Matter collapsing that all down to having one application and one control model for all of the smart home devices.

Corey: Wemo explicitly stated that they’re not going to be pursuing this because it doesn’t let them differentiate the experience. Read as, cash grab. I also found out that Wemo—which is, of course, a Belkin subsidiary—had a critical vulnerability in some of the light switches it offered, including the one built into the wall in this room—until a week ago—where they’re not going to be releasing a patch for it because those are end-of-life. Really? Because I log into the Wemo app and the only way I would have known this has been the fact that it’s been a suspiciously long time since there was a firmware update available for it. But that’s it. Like, the only way I found this out was via a security advisory, at which point that got ripped out of the wall and replaced with something that isn’t, you know, horrifying. But man did that bother me.

Koz: Yeah. I think this is still an open issue for the smart home world.

Corey: Every company wants a moat of some sort, but I don’t want 15 different apps to manage this stuff. You turned me on to Home Assistant, which is an open-source, home control automation system and, on some level, the interface is very clearly built by a bunch of open-source people—good for them; they could benefit from a graphic designer or three to—or user experience person to tie it all together, but once you wrap your head around it, it works really well, where I have automations let me do different things. They even have an Apple Watch app [without its 00:32:14] complications on it. So, I can tap the thing and turn on the lights in my office to different levels if I don’t want to talk to the robot that runs my house. And because my daughter has started getting very deeply absorbed into some YouTube videos from time to time, after the third time I asked her what—I call her name, I tap a different one and the internet dies to her iPad specifically, and I wait about 30 to 45 seconds, and she’ll find me immediately.

Koz: That’s an amazing automation. I love Home Assistant. It’s certainly more technical than I could give to my parents, for example, right now. I think things like Matter are going to bring a lot of that functionality to the easier-to-use hubs. And I think Home Assistant will get better over time as well.

I think the only way to deal with these devices that are going to end-of-life and stop getting support is have them be local control only and so then it’s your hub that keeps getting support and that’s what talks to the internet. And so, you don’t—you know, if there’s a vulnerability in the TCP stack, for example, in your light switch, but your light switch only talks to the hub and isn’t allowed to talk to anything else, how severe is that? I don’t think it’s so bad. Certainly, I wall off all of my IoT devices so that they don’t talk to the rest of my network, but now you’re getting a fairly complicated networking… mojo that listeners to your podcast I’m sure capable of, but many people aren’t.

Corey: I had something that did something very similar and then I had to remove a lot of those restrictions, try to diagnose a phantom issue that it appears was an unreported bug in the wireless AP when you use its second ethernet port as a bridge, where things would intermittently not be able to cross VLANs when passing through that. As in, the initial host key exchange for SSH would work and then it would stall and resets on both sides and it was a disaster. It was, what is going on here? And the answer was it was haunted. So, a small architecture change later, and the problem has not recurred. I need to reapply those restrictions.

Koz: I mean, these are the kinds of things that just make me want to live in a shack in the woods, right? Like, I don’t know how you manage something like that. Like, these are just pain points all over. I think over time, they’ll get better, but until then, that shack in the woods with not even running water sounds pretty appealing.

Corey: Yeah, at some level, having smart lights, for example, one of the best approaches that all the manufacturers I’ve seen have taken, it still works exactly as you would expect when you hit the light switch on the wall because that’s something that you really need to make work or it turns out for those of us who don’t live alone, we will not be allowed to smart home things anymore.

Koz: Exactly. I don’t have any smart bulbs in my house. They’re all smart switches because I don’t want to have to put tape over something and say, “Don’t hit that switch.” And then watch one of my family members pull the tape off and hit the switch anyways.

Corey: I have floor lamps with smart bulbs in them, but I wind up treating them all as one device. And I mean, I’ve taken the switch out from the root because it’s, like, too many things to wind up slicing and dicing. But yeah, there’s a scaling problem because right now a lot of this stuff—because Matter is not quite there all winds up using either Zigbee—which is fine; I have no problem with that it feels like it’s becoming Matter quickly—or WiFi. And there is an upper bound to how many devices you want or can have on some fairly limited frequency.

Koz: Yeah. I think this is still something that needs to be resolved. You know, I’ve got hundreds of devices in my house. Thankfully, most of them are not WiFi or Zigbee. But I think we’re going to see this evolve over time and I’m excited for it.

Corey: I was talking to someone where I was explaining that, well, how this stuff works. Like, “Well, how many devices could you possibly have on your home network?” And at the time it was about 70 or 80. And they just stared at me for the longest time. I mean, it used to be that I could name all the computers in my house. I can no longer do that.

Koz: Sure. Well, I mean, every light switch ends up being a computer.

Corey: And that’s the weirdest thing is that it’s, I’m used to computers, being a thing that requires maintenance and care and feeding and security patches and—yes, relevant to your work—an SSL certificate. It’s like, so what does all of that fancy wizardry do? Well, when it receives a signal, it completes a circuit. The end. And it’s, are really better off for some of these things? There are days we wonder.

Koz: Well, my light bill, my electric bill, is definitely better off having these smart switches because nobody in my house seems to know how to turn a light switch off. And so, having the house do it itself helps quite a bit.

Corey: To be very clear, I would skewer you if you worked on an AWS service that actually charged money for anything for what you just said about the complaining about light bills and optimizing light bills and the rest—

Koz: [laugh].

Corey: —but I’ve never had to optimize your service’s certificate bill beca—after you’ve spun off the one thing that charges—because you can’t cost optimize free, as it turns out, and I’ve yet to find a way to the one optimization possible where now you start paying customers money. I’m sure there’s a way to do that somewhere but damned if I can find it.

Koz: Well, if you find a way to optimize free, please let me know and I’ll share it with all of our customers.

Corey: [laugh]. Isn’t that the truth? I really want to thank you for taking the time to speak with me today. If people want to learn more, where’s the best place for them to find you?

Koz: I can give you the standard AWS answer.

Corey: Yeah, www.aws.com. Yeah.

Koz: Well, I would have said koz@amazon.com. I’m always happy to talk about certs and PKI. I find myself less active on social media lately. You can find me, I guess, on Twitter as @seakoz and on Bluesky as [kozolchyk.com 00:38:03].

Corey: And we will put links to all of that in the [show notes 00:38:06]. Thank you so much for being so generous with your time. I appreciate it.

Koz: Always happy, Corey.

Corey: Jonathan Kozolchyk, or Koz as we all call him, general manager for Certificate Services at AWS. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry, insulting comment that then will fail to post because your podcast platform of choice has an expired security certificate.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

View Details

Jake Gold, Infrastructure Engineer at Bluesky, joins Corey on Screaming in the Cloud to discuss his experience helping to build Bluesky and why he’s so excited about it. Jake and Corey discuss the major differences when building a truly open-source social media platform, and Jake highlights his focus on reliability. Jake explains why he feels downtime can actually be a huge benefit to reliability engineers, and why how he views abstractions based on the size of the team he’s working on. Corey and Jake also discuss whether cloud is truly living up to its original promise of lowered costs.

About Jake

Jake Gold leads infrastructure at Bluesky, where the team is developing and deploying the decentralized social media protocol, ATP. Jake has previously managed infrastructure at companies such as Docker and Flipboard, and most recently, he was the founding leader of the Robot Reliability Team at Nuro, an autonomous delivery vehicle company.

Links Referenced:

  • Bluesky: https://blueskyweb.xyz/
  • Bluesky waitlist signup: https://bsky.app

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. In case folks have missed this, I spent an inordinate amount of time on Twitter over the last decade or so, to the point where my wife, my business partner, and a couple of friends all went in over the holidays and got me a leather-bound set of books titled The Collected Works of Corey Quinn. It turns out that I have over a million words of shitpost on Twitter. If you’ve also been living in a cave for the last year, you’ll notice that Twitter has basically been bought and driven into the ground by the world’s saddest manchild, so there’s been a bit of a diaspora as far as people trying to figure out where community lives.

Jake Gold is an infrastructure engineer at Bluesky—which I will continue to be mispronouncing as Blue-ski because that’s the kind of person I am—which is, as best I can tell, one of the leading contenders, if not the leading contender to replace what Twitter was for me. Jake, welcome to the show.

Jake: Thanks a lot, Corey. Glad to be here.

Corey: So, there’s a lot of different angles we can take on this. We can talk about the policy side of it, we can talk about social networks and things we learn watching people in large groups with quasi-anonymity, we can talk about all kinds of different nonsense. But I don’t want to do that because I am an old-school Linux systems administrator. And I believe you came from the exact same path, given that as we were making sure that I had, you know, the right person on the show, you came into work at a company after I’d left previously. So, not only are you good at the whole Linux server thing; you also have seen exactly how good I am not at the Linux server thing.

Jake: Well, I don’t remember there being any problems at TrueCar, where you worked before me. But yeah, my background is doing Linux systems administration, which turned into, sort of, Linux programming. And these days, we call it, you know, site reliability engineering. But yeah, I discovered Linux in the late-90s, as a teenager and, you know, installing Slackware on 50 floppy disks and things like that. And I just fell in love with the magic of, like, being able to run a web server, you know? I got a hosting account at, you know, my local ISP, and I was like, how do they do that, right?

And then I figured out how to do it. I ran Apache, and it was like, still one of my core memories of getting, you know, httpd running and being able to access it over the internet and telling my friends on IRC. And so, I’ve done a whole bunch of things since then, but that’s still, like, the part that I love the most.

Corey: The thing that continually surprises me is just what I think I’m out and we’ve moved into a fully modern world where oh, all I do is I write code anymore, which I didn’t realize I was doing until I realized if you call YAML code, you can get away with anything. And I get dragged—myself getting dragged back in. It’s the falling back to fundamentals in these weird moments of yes, yes, immutable everything, Infrastructure is code, but when the server is misbehaving and you want to log in and get your hands dirty, the skill set rears its head yet again. At least that’s what I’ve been noticing, at least as far as I’ve gone down a number of interesting IoT-based projects lately. Is that something you experience or have you evolved fully and not looked back?

Jake: Yeah. No, what I try to do is on my personal projects, I’ll use all the latest cool, flashy things, any abstraction you want, I’ll try out everything, and then what I do it at work, I kind of have, like, a one or two year, sort of, lagging adoption of technologies, like, when I’ve actually shaken them out in my own stuff, then I use them at work. But yeah, I think one of my favorite quotes is, like, “Programmers first learn the power of abstraction, then they learn the cost of abstraction, and then they’re ready to program.” And that’s how I view infrastructure, very similar thing where, you know, certain abstractions like container orchestration, or you know, things like that can be super powerful if you need them, but like, you know, that’s generally very large companies with lots of teams and things like that. And if you’re not that, it pays dividends to not use overly complicated, overly abstracted things. And so, that tends to be [where 00:04:22] I follow up most of the time.

Corey: I’m sure someone’s going to consider this to be heresy, but if I’m tasked with getting a web application up and running in short order, I’m putting it on an old-school traditional three-tier architecture where you have a database server, a web server or two, and maybe a job server that lives between them. Because is it the hotness? No. Is it going to be resume bait? Not really.

But you know, it’s deterministic as far as where things live. When something breaks, I know where to find it. And you can miss me with the, “Well, that’s not webscale,” response because yeah, by the time I’m getting something up overnight, to this has to serve the entire internet, there’s probably a number of architectural iterations I’m going to be able to go through. The question is, what am I most comfortable with and what can I get things up and running with that’s tried and tested?

I’m also remarkably conservative on things like databases and file systems because mistakes at that level are absolutely going to show. Now, I don’t know how much you’re able to talk about the Blue-ski infrastructure without getting yelled at by various folks, but how modern versus… reliable—I guess that’s probably a fair axis to put it on: modernity versus reliability—where on that spectrum, does the official Blue-ski infrastructure land these days?

Jake: Yeah. So, I mean, we’re in a fortunate position of being an open-source company working on an open protocol, and so we feel very comfortable talking about basically everything. Yeah, and I’ve talked about this a bit on the app, but the basic idea we have right now is we’re using AWS, we have auto-scaling groups, and those auto-scaling groups are just EC2 instances running Docker CE—the Community Edition—for the runtime and for containers. And then we have a load balancer in front and a Postgres multi-AZ instance in the back on RDS, and it is really, really simple.

And, like, when I talk about the difference between, like, a reliability engineer and a normal software engineer is, software engineers tend to be very feature-focused, you know, they’re adding capabilities to a system. And the goal and the mission of a reliability team is to focus on reliability, right? Like, that’s the primary thing that we’re worried about. So, what I find to be the best resume builder is that I can say with a lot of certainty that if you talk to any teams that I’ve worked on, they will say that the infrastructure I ran was very reliable, it was very secure, and it ended up being very scalable because you know, the way we solve the, sort of, integration thing is you just version your infrastructure, right? And I think this works really well.

You just say, “Hey, this was the way we did it now and we’re going to call that V1. And now we’re going to work on V2. And what should V2 be?” And maybe that does need something more complicated. Maybe you need to bring in Kubernetes, you maybe need to bring in a super-cool reverse proxy that has all sorts of capabilities that your current one doesn’t.

Yeah, but by versioning it, you just—it takes away a lot of the, sort of, interpersonal issues that can happen where, like, “Hey, we’re replacing Jake’s infrastructure with Bob’s infrastructure or whatever.” I just say it’s V1, it’s V2, it’s V3, and then I find that solves a huge number of the problems with that sort of dynamic. But yeah, at Bluesky, like, you know, the big thing that we are focused on is federation is scaling for us because the idea is not for us to run the entire global infrastructure for AT Proto, which is the protocol that Bluesky is based on. The idea is that it’s this big open thing like the web, right? Like, you know, Netscape popularized the web, but they didn’t run every web server, they didn’t run every search engine, right, they didn’t run all the payment stuff. They just did all of the core stuff, you know, they created SSL, right, which became TLS, and they did all the things that were necessary to make the whole system large, federated, and scalable. But they didn’t run it all. And that’s exactly the same goal we have.

Corey: The obvious counterexample is, no, but then you take basically their spiritual successor, which is Google, and they build the security, they build—they run a lot of the servers, they have the search engine, they have the payments infrastructure, and then they turn a lot of it off for fun and… I would say profit, except it’s the exact opposite of that. But I digress. I do have a question for you that I love to throw at people whenever they start talking about how their infrastructure involves auto-scaling. And I found this during the pandemic in that a lot of people believed in their heart-of-hearts that they were auto-scaling, but people lie, mostly to themselves. And you would look at their daily or hourly spend of their infrastructure and their user traffic dropped off a cliff and their spend was so flat you could basically eat off of it and set a table on top of it. If you pull up Cost Explorer and look through your environment, how large are the peaks and valleys over the course of a given day or week cycle?

Jake: Yeah, no, that’s a really good point. I think my basic approach right now is that we’re so small, we don’t really need to optimize very much for cost, you know? We have this sort of base level of traffic and it’s not worth a huge amount of engineering time to do a lot of dynamic scaling and things like that. The main benefit we get from auto-scaling groups is really just doing the refresh to replace all of them, right? So, we’re also doing the immutable server concept, right, which was popularized by Netflix.

And so, that’s what we’re really getting from auto-scaling groups. We’re not even doing dynamic scaling, right? So, it’s not keyed to some metric, you know, the number of instances that we have at the app server layer. But the cool thing is, you can do that when you’re ready for it, right? The big issue is, you know, okay, you’re scaling up your app instances, but is your database scaling up, right, because there’s not a lot of use in having a whole bunch of app servers if the database is overloaded? And that tends to be the bottleneck for, kind of, any complicated kind of application like ours. So, right now, the bill is very flat; you could eat off, and—if it wasn’t for the CDN traffic and the load balancer traffic and things like that, which are relatively minor.

Corey: I just want to stop for a second and marvel at just how educated that answer was. It’s, I talk to a lot of folks who are early-stage who come and ask me about their AWS bills and what sort of things should they concern themselves with, and my answer tends to surprise them, which is, “You almost certainly should not unless things are bizarre and ridiculous. You are not going to build your way to your next milestone by cutting costs or optimizing your infrastructure.” The one thing that I would make sure to do is plan for a future of success, which means having account segregation where it makes sense, having tags in place so that when, “Huh, this thing’s gotten really expensive. What’s driving all of that?” Can be answered without a six-week research project attached to it.

But those are baseline AWS Hygiene 101. How do I optimize my bill further, usually the right answer is go build. Don’t worry about the small stuff. What’s always disturbing is people have that perspective and they’re spending $300 million a year. But it turns out that not caring about your AWS bill was, in fact, a zero interest rate phenomenon.

Jake: Yeah. So, we do all of those basic things. I think I went a little further than many people would where every single one of our—so we have different projects, right? So, we have the big graph server, which is sort of like the indexer for the whole network, and we have the PDS, which is the Personal Data Server, which is, kind of, where all of people’s actual social data goes, your likes and your posts and things like that. And then we have a dev staging, sandbox, prod environment for each one of those, right? And there’s more services besides. But the way we have it is those are all in completely separated VPCs with no peering whatsoever between them. They are all on distinct IP addresses, IP ranges, so that we could do VPC peering very easily across all of them.

Corey: Ah, that’s someone who’s done data center work before with overlapping IP address ranges and swore, never again.

Jake: Exactly. That is when I had been burned. I have cleaned up my mess and other people’s messes. And there’s nothing less fun than renumbering a large complicated network. But yeah, so once we have all these separate VPCs and so it’s very easy for us to say, hey, we’re going to take this whole stack from here and move it over to a different region, a different provider, you know?

And the other thing is that we’re doing is, we’re completely cloud agnostic, right? I really like AWS, I think they are the… the market leader for a reason: they’re very reliable. But we’re building this large federated network, so we’re going to need to place infrastructure in places where AWS doesn’t exist, for example, right? So, we need the ability to take an environment and replicate it in wherever. And of course, they have very good coverage, but there are places they don’t exist. And that’s all made much easier by the fact that we’ve had a very strong separation of concerns.

Corey: I always found it fun that when you had these decentralized projects that were invariably NFT or cryptocurrency-driven over the past, eh, five or six years or so, and then AWS would take a us-east-1 outage in a variety of different and exciting ways,j and all these projects would go down hard. It’s, okay, you talk a lot about decentralization for having hard dependencies on one company in one data center, effectively, doing something right. And it becomes a harder problem in the fullness of time. There is the counterargument, in that when us-east-1 is having problems, most of the internet isn’t working, so does your offering need to be up and running at all costs? There are some people for whom that answer is very much, yes. People will die if what we’re running is not up and running. Usually, a social network is not on that list.

Jake: Yeah. One of the things that is surprising, I think, often when I talk about this as a reliability engineer, is that I think people sometimes over-index on downtime, you know? They just, they think it’s much bigger deal than it is. You know, I’ve worked on systems where there was credit card processing where you’re losing a million dollars a minute or something. And like, in that case, okay, it matters a lot because you can put a real dollar figure on it, but it’s amazing how a few of the bumps in the road we’ve already had with Bluesky have turned into, sort of, fun events, right?

Like, we had a bug in our invite code system where people were getting too many invite codes and it was sort of caused a problem, but it was a super fun event. We all think back on it fondly, right? And so, outages are not fun, but they’re not life and death, generally. And if you look at the traffic, usually what happens is after an outage traffic tends to go up. And a lot of the people that joined, they’re just, they’re talking about the fun outage that they missed because they weren’t even on the network, right?

So, it’s like, I also like to remind people that eBay for many years used to have, like, an outage Wednesday, right? Whereas they could put a huge dollar figure on how much money they lost every Wednesday and yet eBay did quite well, right? Like, it’s amazing what you can do if you relax the constraints of downtime a little bit. You can do maintenance things that would be impossible otherwise, which makes the whole thing work better the rest of the time, for example.

Corey: I mean, it’s 2023 and the Social Security Administration’s website still has business hours. They take a nightly four to six-hour maintenance window. It’s like, the last person out of the office turns off the server or something. I imagine some horrifying mainframe job that needs to wind up sweeping after itself are running some compute jobs. But yeah, for a lot of these use cases, that downtime is absolutely acceptable.

I am curious as to… as you just said, you’re building this out with an idea that it runs everywhere. So, you’re on AWS right now because yeah, they are the market leader for a reason. If I’m building something from scratch, I’d be hard-pressed not to pick AWS for a variety of reasons. If I didn’t have cloud expertise, I think I’d be more strongly inclined toward Google, but that’s neither here nor there. But the problem is these large cloud providers have certain economic factors that they all treat similarly since they’re competing with each other, and that causes me to believe things that aren’t necessarily true.

One of those is that egress bandwidth to the internet is very expensive. I’ve worked in data centers. I know how 95th percentile commit bandwidth billing works. It is not overwhelmingly expensive, but you can be forgiven for believing that it is looking at cloud environments. Today, Blue-ski does not support animated GIFs—however you want to mispronounce that word—they don’t support embedded videos, and my immediate thought is, “Oh yeah, those things would be super expensive to wind up sharing.”

I don’t know that that’s true. I don’t get the sense that those are major cost drivers. I think it’s more a matter of complexity than the rest. But how are you making sure that the large cloud provider economic models don’t inherently shape your view of what to build versus what not to build?

Jake: Yeah, no, I kind of knew where you’re going as soon as you mentioned that because anyone who’s worked in data centers knows that the bandwidth pricing is out of control. And I think one of the cool things that Cloudflare did is they stopped charging for egress bandwidth in certain scenarios, which is kind of amazing. And I think it’s—the other thing that a lot of people don’t realize is that, you know, these network connections tend to be fully symmetric, right? So, if it’s a gigabit down, it’s also a gigabit up at the same time, right? There’s two gigabits that can be transferred per second.

And then the other thing that I find a little bit frustrating on the public cloud is that they don’t really pass on the compute performance improvements that have happened over the last few years, right? Like computers are really fast, right? So, if you look at a provider like Hetzner, they’re giving you these monster machines for $128 a month or something, right? And then you go and try to buy that same thing on the public, the big cloud providers, and the equivalent is ten times that, right? And then if you add in the bandwidth, it’s another multiple, depending on how much you’re transferring.

Corey: You can get Mac Minis on EC2 now, and you do the math out and the Mac Mini hardware is paid for in the first two or three months of spinning that thing up. And yes, there’s value in AWS’s engineering and being able to map IAM and EBS to it. In some use cases, yeah, it’s well worth having, but not in every case. And the economics get very hard to justify for an awful lot of work cases.

Jake: Yeah, I mean, to your point, though, about, like, limiting product features and things like that, like, one of the goals I have with doing infrastructure at Bluesky is to not let the infrastructure be a limiter on our product decisions. And a lot of that means that we’ll put servers on Hetzner, we’ll colo servers for things like that. I find that there’s a really good hybrid cloud thing where you use AWS or GCP or Azure, and you use them for your most critical things, you’re relatively low bandwidth things and the things that need to be the most flexible in terms of region and things like that—and security—and then for these, sort of, bulk services, pushing a lot of video content, right, or pushing a lot of images, those things, you put in a colo somewhere and you have these sort of CDN-like servers. And that kind of gives you the best of both worlds. And so, you know, that’s the approach that we’ll most likely take at Bluesky.

Corey: I want to emphasize something you said a minute ago about CloudFlare, where when they first announced R2, their object store alternative, when it first came out, I did an analysis on this to explain to people just why this was as big as it was. Let’s say you have a one-gigabyte file and it blows up and a million people download it over the course of a month. AWS will come to you with a completely straight face, give you a bill for $65,000 and expect you to pay it. The exact same pattern with R2 in front of it, at the end of the month, you will be faced with a bill for 13 cents rounded up, and you will be expected to pay it, and something like 9 to 12 cents of that initially would have just been the storage cost on S3 and the single egress fee for it. The rest is there is no egress cost tied to it.

Now, is Cloudflare going to let you send petabytes to the internet and not charge you on a bandwidth basis? Probably not. But they’re also going to reach out with an upsell and they’re going to have a conversation with you. “Would you like to transition to our enterprise plan?” Which is a hell of a lot better than, “I got Slashdotted”—or whatever the modern version of that is—“And here’s a surprise bill that’s going to cost as much as a Tesla.”

Jake: Yeah, I mean, I think one of the things that the cloud providers should hopefully eventually do—I hope Cloudflare pushes them in this direction—is to start—the original vision of AWS when I first started using it in 2006 or whenever launched, was—and they said this—they said they’re going to lower your bill every so often, you know, as Moore’s law makes their bill lower. And that kind of happened a little bit here and there, but it hasn’t happened to the same degree that you know, I think all of us hoped it would. And I would love to see a cloud provider—and you know, Hetzner does this to some degree, but I’d love to see these really big cloud providers that are so great in so many ways, just pass on the savings of technology to the customer so we’ll use more stuff there. I think it’s a very enlightened viewpoint is to just say, “Hey, we’re going to lower the costs, increase the efficiency, and then pass it on to customers, and then they will use more of our services as a result.” And I think Cloudflare is kind of leading the way in there, which I love.

Corey: I do need to add something there—because otherwise we’re going to get letters and I don’t think we want that—where AWS reps will, of course, reach out and say that they have cut prices over a hundred times. And they’re going to ignore the fact that a lot of these were a service you don’t use in a region you couldn’t find a map if your life depended on it now is going to be 10% less. Great. But let’s look at the general case, where from C3 to C4—if you get the same size instance—it cut the price by a lot. C4 to C5, somewhat. C5 to C6 effectively is no change. And now, from C6 to C7, it is 6% more expensive like for like.

And they’re making noises about price performance is still better, but there are an awful lot of us who say things like, “I need ten of these servers to live over there.” That workload gets more expensive when you start treating it that way. And maybe the price performance is there, maybe it’s not, but it is clear that the bill always goes down is not true.

Jake: Yeah, and I think for certain kinds of organizations, it’s totally fine the way that they do it. They do a pretty good job on price and performance. But for sort of more technical companies—especially—it’s just you can see the gaps there, where that Hetzner is filling and that colocation is still filling. And I personally, you know, if I didn’t need to do those things, I wouldn’t do them, right? But the fact that you need to do them, I think, says kind of everything.

Corey: Tired of wrestling with Apache Kafka's complexity and cost? Feel like you're stuck in a Kafka novel, but with more latency spikes and less existential dread by at least 10%? You're not alone.

What if there was a way to 10x your streaming data performance without having to rob a bank? Enter Redpanda. It's not just another Kafka wannabe. Redpanda powers mission-critical workloads without making your AWS bill look like a phone number.

And with full Kafka API compatibility, migration is smoother than a fresh jar of peanut butter. Imagine cutting as much as 50% off your AWS bills. With Redpanda, it's not a pipedream, it's reality.

Visit go.redpanda.com/duckbill today. Redpanda: Because your data infrastructure shouldn’t give you Kafkaesque nightmares.

Corey: There are so many weird AWS billing stories that all distill down to you not knowing this one piece of trivia about how AWS works, either as a system, as a billing construct, or as something else. And there’s a reason this has become my career of tracing these things down. And sometimes I’ll talk to prospective clients, and they’ll say, “Well, what if you don’t discover any misconfigurations like that in our account?” It’s, “Well, you would be the first company I’ve ever seen where that [laugh] was not true.” So honestly, I want to do a case study if we do.

And I’ve never had to write that case study, just because it’s the tax on not having the forcing function of building in data centers. There’s always this idea that in a data center, you’re going to run out of power, space, capacity, at some point and it’s going to force a reckoning. The cloud has what distills down to infinite capacity; they can add it faster than you can fill it. So, at some point it’s always just keep adding more things to it. There’s never a let’s clean out all of the cruft story. And it just accumulates and the bill continues to go up and to the right.

Jake: Yeah, I mean, one of the things that they’ve done so well is handle the provisioning part, right, which is kind of what you’re getting out there. One of the hardest things in the old days, before we all used AWS and GCP, is you’d have to sort of requisition hardware and there’d be this whole process with legal and financing and there’d be this big lag between the time you need a bunch more servers in your data center and when you actually have them, right, and that’s not even counting the time takes to rack them and get them, you know, on network. The fact that basically, every developer now just gets an unlimited credit card, they can just, you know, use that’s hugely empowering, and it’s for the benefit of the companies they work for almost all the time. But it is an uncapped credit card. I know, they actually support controls and things like that, but in general, the way we treated it—

Corey: Not as much as you would think, as it turns out. But yeah, it’s—yeah, and that’s a problem. Because again, if I want to spin up $65,000 an hour worth of compute right now, the fact that I can do that is massive. The fact that I could do that accidentally when I don’t intend to is also massive.

Jake: Yeah, it’s very easy to think you’re going to spend a certain amount and then oh, traffic’s a lot higher, or, oh, I didn’t realize when you enable that thing, it charges you an extra fee or something like that. So, it’s very opaque. It’s very complicated. All of these things are, you know, the result of just building more and more stuff on top of more and more stuff to support more and more use cases. Which is great, but then it does create this very sort of opaque billing problem, which I think, you know, you’re helping companies solve. And I totally get why they need your help.

Corey: What’s interesting to me about distributed social networks is that I’ve been using Mastodon for a little bit and I’ve started to see some of the challenges around a lot of these things, just from an infrastructure and architecture perspective. Tim Bray, former Distinguished Engineer at AWS posted a blog post yesterday, and okay, well, if Tim wants to put something up there that he thinks people should read, I advise people generally read it. I have yet to find him wasting my time. And I clicked it and got a, “Server over resource limits.” It’s like wow, you’re very popular. You wound up getting—got effectively Slashdotted.

And he said, “No, no. Whatever I post a link to Mastodon, two thousand instances all hidden at the same time.” And it’s, “Oh, yeah. The hug of death. That becomes a challenge.” Not to mention the fact that, depending upon architecture and preferences that you make, running a Mastodon instance can be extraordinarily expensive in terms of storage, just because it’ll, by default, attempt to cache everything that it encounters for a period of time. And that gets very heavy very quickly. Does the AT Protocol—AT Protocol? I don’t know how you pronounce it officially these days—take into account the challenges of running infrastructures designed for folks who have corporate budgets behind them? Or is that really a future problem for us to worry about when the time comes?

Jake: No, yeah, that’s a core thing that we talked about a lot in the recent, sort of, architecture discussions. I’m going to go back quite a ways, but there were some changes made about six months ago in our thinking, and one of the big things that we wanted to get right was the ability for people to host their own PDS, which is equivalent to, like, posting a WordPress or something. It’s where you post your content, it’s where you post your likes, and all that kind of thing. We call it your repository or your repo. But that we wanted to make it so that people could self-host that on a, you know, four or five $6-a-month droplet on DigitalOcean or wherever and that not be a problem, not go down when they got a lot of traffic.

And so, the architecture of AT Proto in general, but the Bluesky app on AT Proto is such that you really don’t need a lot of resources. The data is all signed with your cryptographic keys—like, not something you have to worry about as a non-technical user—but all the data is authenticated. That’s what—it’s Authenticated Transfer Protocol. And because of that, it doesn’t matter where you get the data, right? So, we have this idea of this big indexer that’s looking at the entire network called the BGS, the Big Graph Server and you can go to the BGS and get the data that came from somebody’s PDS and it’s just as good as if you got it directly from the PDS. And that makes it highly cacheable, highly conducive to CDNs and things like that. So no, we intend to solve that problem entirely.

Corey: I’m looking forward to seeing how that plays out because the idea of self-hosting always kind of appealed to me when I was younger, which is why when I met my wife, I had a two-bedroom apartment—because I lived in Los Angeles, not San Francisco, and could afford such a thing—and the guest bedroom was always, you know, 10 to 15 degrees warmer than the rest of the apartment because I had a bunch of quote-unquote, “Servers” there, meaning deprecated desktops that my employer had no use for and said, “It’s either going to e-waste or your place if you want some.” And, okay, why not? I’ll build my own cluster at home. And increasingly over time, I found that it got harder and harder to do things that I liked and that made sense. I used to have a partial rack in downtown LA where I ran my own mail server, among other things.

And when I switched to Google for email solutions, I suddenly found that I was spending five bucks a month at the time, instead of the rack rental, and I was spending two hours less a week just fighting spam in a variety of different ways because that is where my technical background lives. Being able to not have to think about problems like that, and just do the fun part was great. But I worry about the centralization that that implies. I was opposed to it at the idea because I didn’t want to give Google access to all of my mail. And then I checked and something like 43% of the people I was emailing were at Gmail-hosted addresses, so they already had my email anyway. What was I really doing by not engaging with them? I worry that self-hosting is going to become passe, so I love projects that do it in sane and simple ways that don’t require massive amounts of startup capital to get started with.

Jake: Yeah, the account portability feature of AT Proto is super, super core. You can backup all of your data to your phone—the [AT 00:28:36] doesn’t do this yet, but it most likely will in the future—you can backup all of your data to your phone and then you can synchronize it all to another server. So, if for whatever reason, you’re on a PDS instance and it disappears—which is a common problem in the Mastodon world—it’s not really a problem. You just sync all that data to a new PDS and you’re back where you were. You didn’t lose any followers, you didn’t lose any posts, you didn’t lose any likes.

And we’re also making sure that this works for non-technical people. So, you know, you don’t have to host your own PDS, right? That’s something that technical people can self-host if they want to, non-technical people can just get a host from anywhere and it doesn’t really matter where your host is. But we are absolutely trying to avoid the fate of SMTP and, you know, other protocols. The web itself, right, is sort of… it’s hard to launch a search engine because the—first of all, the bar is billions of dollars a year in investment, and a lot of websites will only let us crawl them at a higher rate if you’re actually coming from a Google IP, right? They’re doing reverse DNS lookups, and things like that to verify that you are Google.

And the problem with that is now there’s sort of this centralization with a search engine that can’t be fixed. With AT Proto, it’s much easier to scrape all of the PDSes, right? So, if you want to crawl all the PDSes out on the AT Proto network, they’re designed to be crawled from day one. It’s all structured data, we’re working on, sort of, how you handle rate limits and things like that still, but the idea is it’s very easy to create an index of the entire network, which makes it very easy to create feed generators, search engines, or any other kind of sort of big world networking thing out there. And then without making the PDSes have to be very high power, right? So, they can do low power and still scrapeable, still crawlable.

Corey: Yeah, the idea of having portability is super important. Question I’ve got—you know, while I’m talking to you, it’s, we’ll turn this into technical support hour as well because why not—I tend to always historically put my Twitter handle on conference slides. When I had the first template made, I used it as soon as it came in and there was an extra n in the @quinnypig username at the bottom. And of course, someone asked about that during Q&A.

So, the answer I gave was, of course, n+1 redundancy. But great. If I were to have one domain there today and change it tomorrow, is there a redirect option in place where someone could go and find that on Blue-ski, and oh, they’ll get redirected to where I am now. Or is it just one of those 404, sucks to be you moments? Because I can see validity to both.

Jake: Yeah, so the way we handle it right now is if you have a, something.bsky.social name and you switch it to your own domain or something like that, we don’t yet forward it from the old.bsky.social name. But that is totally feasible. It’s totally possible. Like, the way that those are stored in your what’s called your [DID record 00:31:16] or [DID document 00:31:17] is that there’s, like, a list that currently only has one item in general, but it’s a list of all of your different names, right? So, you could have different domain names, different subdomain names, and they would all point back to the same user. And so yeah, so basically, the idea is that you have these aliases and they will forward to the new one, whatever the current canonical one is.

Corey: Excellent. That is something that concerns me because it feels like it’s one of those one-way doors, in the same way that picking an email address was a one-way door. I know people who still pay money to their ancient crappy ISP because they have a few mails that come in once in a while that are super-important. I was fortunate enough to have jumped on the bandwagon early enough that my vanity domain is 22 years old this year. And my email address still works,which, great, every once in a while, I still get stuff to, like, variants of my name I no longer use anymore since 2005. And it’s usually spam, but every once in a blue moon, it’s something important, like, “Hey, I don’t know if you remember me. We went to college together many years ago.” It’s ho-ly crap, the world is smaller than we think.

Jake: Yeah.j I mean, I love that we’re using domains, I think that’s one of the greatest decisions we made is… is that you own your own domain. You’re not really stuck in our namespace, right? Like, one of the things with traditional social networks is you’re sort of, their domain.com/yourname, right?

And with the way AT Proto and Bluesky work is, you can go and get a domain name from any registrar, there’s hundreds of them—you know, we’d like Namecheap, you can go there and you can grab a domain and you can point it to your account. And if you ever don’t like anything, you can change your domain, you can change, you know which PDS you’re on, it’s all completely controlled by you. And there’s nearly no way we as a company can do anything to change that. Like, that’s all sort of locked into the way that the protocol works, which creates this really great incentive where, you know, if we want to provide you services or somebody else wants to provide you services, they just have to compete on doing a really good job; you’re not locked in. And that’s, like, one of my favorite features of the network.

Corey: I just want to point something out because you mentioned oh, we’re big fans of Namecheap. I am too, for weird half-drunk domain registrations on a lark. Like, “Why am I poor?” It’s like, $3,000 a month of my budget goes to domain purchases, great. But I did a quick whois on the official Bluesky domain and it’s hosted at Route 53, which is Amazon’s, of course, premier database offering.

But I’m a big fan of using a enterprise registrar for enterprise-y things. Wasabi, if I recall correctly, wound up having their primary domain registered through GoDaddy, and the public domain that their bucket equivalent would serve data out of got shut down for 12 hours because some bad actor put something there that shouldn’t have been. And GoDaddy is not an enterprise registrar, despite what they might think—for God’s sake, the word ‘daddy’ is in their name. Do you really think that’s enterprise? Good luck.

So, the fact that you have a responsible company handling these central singular points of failure speaks very well to just your own implementation of these things. Because that’s the sort of thing that everyone figures out the second time.

Jake: Yeah, yeah. I think there’s a big difference between corporate domain registration, and corporate DNS and, like, your personal handle on social networking. I think a lot of the consumer, sort of, domain registries are—registrars—are great for consumers. And I think if you—yeah, you’re running a big corporate domain, you want to make sure it’s, you know, it’s transfer locked and, you know, there’s two-factor authentication and doing all those kinds of things right because that is a single point of failure; you can lose a lot by having your domain taken. So, I completely agree with you on there.

Corey: Oh, absolutely. I am curious about this to see if it’s still the case or not because I haven’t checked this in over a year—and they did fix it. Okay. As of at least when we’re recording this, which is the end of May 2023, Amazon’s Authoritative Name Servers are no longer half at Oracle. Good for them. They now have a bunch of Amazon-specific name servers on them instead of, you know, their competitor that they clearly despise. Good work, good work.

I really want to thank you for taking the time to speak with me about how you’re viewing these things and honestly giving me a chance to go ambling down memory lane. If people want to learn more about what you’re up to, where’s the best place for them to find you?

Jake: Yeah, so I’m on Bluesky. It’s invite only. I apologize for that right now. But if you check out bsky.app, you can see how to sign up for the waitlist, and we are trying to get people on as quickly as possible.

Corey: And I will, of course, be talking to you there and will put links to that in the show notes. Thank you so much for taking the time to speak with me. I really appreciate it.

Jake: Thanks a lot, Corey. It was great.

Corey: Jake Gold, infrastructure engineer at Bluesky, slash Blue-ski. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an angry comment that will no doubt result in a surprise $60,000 bill after you posted.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

View Details

Avi Freedman, CEO at Kentik, joins Corey on Screaming in the Cloud to discuss the fun of solving for observability. Corey and Avi discuss how great simplicity can be deceiving, and Avi points out that with great simplicity comes great complexity. Avi discusses examples of this that he sees in Kentik customer environments, as well as the differences he sees in cloud environments from traditional data center environments. Avi also reveals his predictions for the future and how enterprise M&A will affect the way companies view data centers and VPCs.

About Avi

Avi Freedman is the co-founder and CEO of network observability company Kentik. He has decades of experience as a networking technologist and executive. As a network pioneer in 1992, Freedman started Philadelphia’s first ISP, known as netaxs. He went on to run network operations at Akamai for over a decade as VP of network infrastructure and then as chief network scientist. He also ran the network at AboveNet and was the CTO of ServerCentral.

Links Referenced:

  • Kentik: https://kentik.com
  • Email: avi@kentik.com
  • Twitter: https://twitter.com/avifreedman
  • LinkedIn: https://www.linkedin.com/in/avifreedman

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Most Companies find out way too late that they’ve been breached. Thinkst Canary changes this. Deploy Canaries and Canarytokens in minutes and then forget about them. Attackers tip their hand by touching ’em giving you the one alert, when it matters. With 0 admin overhead and almost no false-positives, Canaries are deployed (and loved) on all 7 continents. Check out what people are saying at canary.love today!

Corey: Welcome to Screaming in the Cloud, I’m Corey Quinn. This promoted guest episode is brought to us by our friends at Kentik. And into my social grist mill, they have thrown Avi Freedman, their CEO. Avi, thank you for joining me.

Avi: Thank you for having me, Corey. I’ve been a big fan for some time, I have never actually fallen off my seat laughing, but I’ve come close a couple times on some of your threads.

Corey: You must have a great chair.

Avi: I should probably upgrade it [laugh].

Corey: [laugh]. I have been looking forward to this conversation for a while because you are one of those rare creatures who comes from a similar world to what I did where we were grumpy and old before our time because we worked on physical infrastructure in data centers, we basically wrangled servers into doing the things that we wanted them to do when hardware reliability was an aspiration rather than a reality. And we also moved on from that, in many ways. We are not blind to the modern order of how computers work. But you still run a lot of what you do in data centers, but many of your customers are in cloud. You speak both languages very fluently because of the unifying thread between all of this, which is, of course, the network. How did you wind up in, I guess we’ll call it network hell.

Avi: [laugh]. I mean, network hell was truly… in the ’90s, when the internet was—I mean, the internet is sort of like the human body: the more you study it, the more amazing it is that it ever worked in the first place, not that it breaks sometimes—was the bugs, and trying to put together the technology back then, you know, that we had the life is pretty good nowadays, other than the [laugh] immense complexity that has been unleashed on us by everyone taking the same technology and then writing it in their own software and giving it their own marketing names. And thus, you have multi-cloud networking. So, got into it because it’s a problem that needs to be solved, right? There’s no ESP that connects the applications together; the network still needs to make it work. And now people own some of it, and then more of it, they don’t own, but they’re still responsible for it. So, it’s a fun problem to solve.

Corey: The timing of this episode is apt because I’ve used Kentik myself for a few things over the years. And to be fair, using it for any of my personal networking problems is a bit like noticing, “Oh, I have a loose thread here on my shirt. Pass me the chainsaw.” It’s, my environment is tiny and it’s over-scoped. But I just earlier this week wound up having to analyze a day’s worth of Flow Logs from one of my clients, and to do this, I had to spin up an EC2 instance with 128 gigs of RAM and then load the Flow Logs for that day into RAM, and then—not kidding—I ran into OOM Killer because I ran out of RAM on this thing.

Avi: [laugh].

Corey: It is, like, yeah, that’s right. The network is chatty, the logs are immense, and it’s easy to forget. Because the reason I was doing this was just to figure out what are the things that are talking to each other in this environment to drive up some aspects of data transfer costs. But that is an esoteric use case for this; it’s not why most people tend to think about network observability. So, I’m going to ask you the blunt question up front here because it might be a really short episode. Do we have to care about networking in the least now that cloud is the default in most locations? It is just an API call away, isn’t it?

Avi: With great simplicity comes great complexity. So, to the people running infrastructure, to developers or architects, turning it all on, it looks like just API calls. But did you set the policies right? Can the things talk to each other? Are they talking in patterns that are causing you wild data transfer costs?

All these things ultimately come back to some team that actually has to make it go. And can be pretty hard to figure that out, right, because it’s not just the VPC Flow Logs. It’s, what’s the policy? It’s, what are they talking to that maybe isn’t in that cloud, that’s maybe in another cloud? So, how do you bring it all together? Like, you could have—and maybe you should have—used Athena, right? You can put VPC Flow Logs in S3 buckets and use Athena and run SQL queries if all you want is your top talker.

Corey: Oh, I did. That’s how I started, but Athena is, uh… it has some challenges. Let’s just put it that way and leave it there. DuckDB is what I was using and I’m much happier with it for a variety of excellent reasons.

Avi: Okay. Well, I’ll tease you another time about, you know—I lost this battle at Kentik. We actually don’t use swap, but I’m a big fan of having swap and monitoring it so the OOM Killer only does what you want or doesn’t fire at all. But that’s a separate religious debate.

Corey: There’s a counterargument of running an in-memory data store. And then oh, we’re going to use it as swap though, so it’s like, hang on, this just feels like running a normal database with extra steps.

Avi: Computers allow you to do amazing things and only occasionally slap you nowadays with it. It’s pretty amazing. But back to the question. APIs make it easy to turn on, but not so easy to run. The observability that you get within a given cloud is typically very limited.

Google actually has the best. They show some topology and other things. I mean, a lot of what we do involves scraping API calls in the cloud to figure out what does this all mean, then convolving it with the VPC Flow Logs and making it look like a network, and what are the gateways, and what are the rules being applied and what can’t talk to itself? If you just look at VPC Flow Logs like it’s Syslog, good luck trying to figure out what VPCs are talking to each other. It’s exactly the problem that you were describing.

So, the ease of turning it on is exactly inversely proportional to the ease of running it. And, you know, as a vendor, we think it’s an awesome [laugh] problem, but we feel for our customers. And you know, occasionally it’s a pain to get the IAM roles set up to scrape things and help them, but that’s you know, that’s just part of the job.

Corey: It’s fascinating to me, just looking from an AWS perspective, just how much work clearly has to be done to translate their Byzantine and very strange networking environment and concepts into things that customers see. Because in many cases, the things that the virtual machines that we’ve run on top of EC2, let alone anything higher level, is being lied to the entire time about what the actual topology of the environment is. It was most notable, for me at least, at re:Invent 2022, the most recent one, where they announced they have a TCP replacement, scalable, reliable data grammar SRD. It’s a new protocol entirely. It’s, “Oh, wow, can we use it?” “No.” “Okay.” Like, I get that it’s a lot of work, I get you’re excited about it. Are you going to talk to us about how it actually works? “Oh, absolutely not.” So… okay, good for you, I guess.

Avi: Doesn’t Amazon have to write a press release before they build anything, and doesn’t the press release have to say, like, why people give a shit, why people care?

Corey: Yep. And their story on this was oh, it enables us to be a lot faster at letting EBS volumes talk to some of our beefier instances.

Avi: [laugh].

Corey: And that’s all well and good, don’t get me wrong, but it’s also, “Yay, it’s more reliable,” is a difficult message to send. I mean, it’s hard enough when—and it’s necessary because you’ve got to tacitly admit that reliability and performance haven’t been all they could be. But when it’s no longer an issue for most folks, now you’re making them wonder, like, wait, how bad was it? It’s just a strange message.

Avi: Yeah. One of my projects for this weekend is, I actually got a gaming PC and I’m going to try compression offload to the CUDA cores because right now, we do compress and decompress with Intel cores. And like, if I’m successful there and we can get 30% faster subqueries—which doesn’t really matter, you know, on the kind of massive queries we run—and 20% more use out of the computers that we actually run, I’m probably not going to do a press release about it. But good to see the pattern.

But you know, what you said is pretty interesting. As people like Kentik, we have to put together, well, on Azure, you can have VPCs that cross regions, right? And in other places, you can’t. And in Google, you have performance metrics that come out and you can get it very frequently, and in Amazon and Azure, you can’t. Like, how do you take these kinds of telemetry that are all the same stuff underneath, but packaged up differently in different quantos and different things and make it all look the same is actually pretty fun and interesting.

And it’s pretty—you know, if you give some cloud engineers who focus on the infrastructure layer enough beers or alcohol or just room to talk, you can hear some funny stories. And it all made sense to somebody in the first place, but unpacking it and actually running it as a common infrastructure can be quite fun.

Corey: One of the things that I have found notable about your perspective, as particularly, you’re running all of the network ingest, to my understanding, in your data center environment. Because we talked about this when you were kind enough to invite me to your company all-hands offsite, presumably I assume when people do that, it’s so they can beat me up in the alley, but that only happened twice. I was very pleasantly surprised.

Avi: [And you 00:09:23] made fun of us only three times, so you know, you beat us—

Corey: Exactly.

Avi: —but it was all enjoyed.

Corey: But always with love. Now, what I found fascinating was you and I sat down for a while and you talked about your data center architecture. And you asked me—since I don’t have anything to sell you—is there an economical way that I could see running your environment on top of AWS? And the answer was sure, if by economical you mean an absolute minimum of six times what you’re currently paying a year, sure you can get there. But it just does not make sense for any realistic approach to doing this.

And the reason I bring this up is that you’re in a data center not because of religious beliefs, “Of, well, this is good enough for my grandpappy, so it’s good enough for me.” It’s because it solves the problem you have in a way that the cloud providers clearly cannot. But you also are not anti-cloud. So, many folks who are all-in on data centers seem to be doing it out of pure self-interest where, well, if everyone goes all-in on cloud, then we have nothing left to sell them. I’ve used AWS VPC Flow Logs. They have nothing that could even remotely be termed network observability. Your future is assured as long as people understand what it is that you’re providing them and what are you that adds. So yeah, people keep going in a cloud direction, you’re happy as houses.

Avi: We’ll use the best tools for building our infrastructure that we can, right? We use cloud. In fact, we’re just buying some reserved instances, which always, you know, I give it the hairy eyeball, but you know, we’re probably always going to have our CI/CD bursty stuff in the cloud. We have performance testing regions on all the major clouds so that we can tell people what performance is to and from cloud. Like, that we have to use cloud for.

And if there’s an always-on model, which starts making sense in the cloud, then I try not to be the first to use anything, but [laugh] we’ll be one of the first to use it. But every year, we talk to, you know, the major clouds because we’re customers of all them, for as I said, our testing infrastructure if nothing else, and you know, some of them for some other parts, you know, for example, proxying VPC Flow Logs, we run infrastructure on Kubernetes in all—in the three biggest to proxy VPC Flow Logs, you know, and so that’s part of our bill. But if something’s always on, you know, one of our storage servers, it’s a $15,000 machine that, you know, realistically runs five years, but even if you assume it runs three years, we get financing for it, cost a couple $100 a month to host, and that’s inclusive of our ops team that runs, sort of, everything, you just do the math. That same machine would be, you know, even not including data transfer would be maybe 3500 a month on cloud. The economics just don’t quite make sense.

For burst, for things like CI/CD, test, seasonality, I think it’s great. And if we have patterns like that, you know, we’re the first to use it. So, it’s just a question of using what’s best. And a lot of our customers are in that realm, too. I would say some of them are a little over-rotated, you know, they’ve had big mandates to go one way or the other and don’t have the right, you know, sort of nuanced view, but I think over time, that’s going to fix itself. And yeah, as you were saying, like, the more people use cloud, the better we do, so it’s just really a question of what’s the best for us at our infrastructure and at any given time.

Corey: I think that that is something that is not fully appreciated or well understood is that I work with cloud technologies because for what I do, it makes an awful lot of sense. But I’ve been lately doing a significant build-out in my home network on the perspective of yeah, this makes sense for what I do. And I now have increased number of workloads that I’m running here and I got to say, it feels a little strange, on some level, not to be paying AWS on something metered by the second whenever I’m running a job here. That always feels a little on the weird side. But I’m not suggesting I filled my house with servers either.

Avi: [unintelligible 00:13:18] going to report you to the House on Cloudian Activities Committee [laugh] for—

Corey: [laugh].

Avi: To straighten you out about your infrastructure use and beliefs. I do have to ask you, and I do have some foreknowledge of this, where is the controller for your network running? Is it running in your house or—

Corey: Oh, the WiFi controller lives in Ohio with all the other unpleasant things. I mean, even data transfer between Ohio and Virginia—if you’re on AWS—is half-price because data wants to get out of Ohio just as much as the people do. And that’s fine, but it can also fail out of band. I can chill that thing for a while and I’m not able to provision new equipment, I can’t spin up new SSIDs, but—

Avi: Right. It’s the same as [kale scale 00:14:00], which is, like, sufficiently indistinguishable from magic, but it’s nice there’s [head scale 00:14:05] in case something happened to them. But yeah, you know, you just can’t set up new stuff without your SSHing old way while it’s down. So.

Corey: And worst case, it goes away irretrievably, I can spin a new one up, I can pair everything locally, do it by repointing DNS locally, and life will go on. It’s one of those areas where, like, I would not have this in Ohio if latency was a concern if it was routing every packet out halfway across the country before it hit the general internet. That would be a challenge for me. But that’s not what I’m doing.

Avi: Yeah, yeah. No, that makes sense. And I think also—

Corey: And I certainly pay AWS by the second for that thing. That’s—I have a three-year savings plan for that thing, and if nothing else, it was useful for me just to figure out what the living hell was going on with the savings plan purchase project one year. That was just, it was challenged to get that straightened out in some ways. Turns out that the high watermark of the console is a hundred-and-some-odd-thirty-million dollars you can add to cart and click the buy button. Have fun.

Avi: My goodness. Okay, well.

Corey: The API goes up to $26.2 billion. Try that in a free tier account, preferably someone else’s.

Avi: I would love to have such problems. Right now, that is not one of them. We don’t spend that much on infrastructure.

Corey: Oh, that is more than Amazon’s—AWS’s at least—quarterly revenue. So, if you wind up doing a $26.2 billion, it’s like—it’s that old saw. You owe Amazon a million dollars, you have a problem. If you owe Amazon $26 billion, Amazon has a problem. Yeah, that’s when Andy Jassy calls you 20 minutes after you make that purchase, and at least to me, he yells at me with a, “Listen here, asshole,” and it sort of devolves from there.

Avi: Well, I do live in Seattle, so you know, they send the posse out, I’m pretty sure.

Corey: [laugh] I will be keynoting DevOpsDays Seattle on August 1st with a talk that might very well resonate with your perspective, “The Modern Devops: A Million Ways to Die in Production.”

Avi: That is very cool. I mean, ultimately, I think that’s what cloud comes back to. When cloud was being formed, it’s just other people’s computers, storage, and network. I don’t know if you’d argue that there’s a politics, control plane, or a—

Corey: Oh, I would say, “Cloud? There’s no cloud; just someone else’s cost center.”

Avi: Exactly. And so, how do you configure it? And back to the question of, should everything be on-prem or does cloud abstract at all, it’s all the same stuff that we’ve been doing for decades and decades, just with other people’s software and names, which you help decode. And then it’s the question we’ve always had: what’s the best thing to do? Do you like [Wellfleet 00:16:33] or [Protion 00:16:35]? Now, do you like Azure [laugh] or Google or Amazon or somebody else or running your own?

Corey: It’s almost this generation's equivalent of Vi versus Emacs.

Avi: Yes. I guess there could be a crowd equivalent. I use VI, but only because I’m a lisp addict and I don’t want to get stuck refining Eliza macros and connecting to the ChatGPT in Emacs. So, you know. Someone just did a Emacs as PID 0. So basically, no init, just, you know, the kernel boots into Emacs, and then someone of course had to do a VI as PID 0. And I have to admit, Emacs would be a lot more useful as a PID 0, even though I use VI.

Corey: I would say that—I mean, you wind up in writing in Emacs and writing lisp in it, then I’ve got to say every third thing you say becomes a parenthetical.

Avi: Exactly. Ha.

Corey: But I want to say that there’s also a definite moving of data going on there that I think is a scale that, for those of us working mostly in home labs and whatnot, can be hard to imagine. And I see that just in terms of the volume of Flow Logs, which to be clear, are smaller than the data transfer they are representing in almost every case.

Avi: Almost every.

Corey: You see so much of the telemetry that comes out of there and what customers are seeing and what their problems are, in different ways. It’s not just Flow Logs, you ingest a whole bunch of different telemetry through a variety of modern and ancient and everything in between variety of protocols to support, you know, the horror that is network equipment interoperability. And just, I can’t—I feel like I can’t do a terrific job of doing justice to describing just how comprehensive Kentik is, once you get it set up as a product. What is on the wire has always been for me the arbiter of truth because computers will lie to you, but it’s very tricky to get them to lie and get the network story to cover for it.

Avi: Right. I mean, ultimately, that’s one of the sources of truth. There’s routing, there’s performance testing, there’s a whole lot of different things, and as you were saying, in any one of these slices of your, let’s just pick the network. There’s many different things that all mean the same, but look different that you need to put together. You could—the nerd term would be, you know, normalizing. You need to take all this stuff and normalize it.

But traffic, we agree, that’s where we started with. We call it the what if what is. What’s actually happening on the infrastructure and that’s the ancient stuff like IPFIX and NetFlow and sFlow. Some people that would argue that, you know, the [IATF 00:19:04] would say, “Oh, we’re still innovating and it’s still current,” but you know, it’s certainly on-prem only. The major cloud vendors would say, “Oh, well, you can run the router—cloud routers—or you could run cloud versions of the big routers,” but we don’t really see that as a super common pattern today.

But what’s really the difference between NetFlow and the VPC Flow Log? Well, some VPC Flow Logs have permit deny because they’re really firewall logs, but ultimately, it’s something went from here to there. There might not be a TCP flag, but there might be something else in cloud. And, you know, maybe there’s rum data, which is also another kind of traffic. And ultimately, all together, we try to take that and then the business metadata to say, whether it’s NetBox in the old world or Kubernetes in the new world, or some other [unintelligible 00:19:49], what application is this? What user is this?

So, you can ask questions about why am I blowing up between these cloud regions? What applications are doing it, right? VPC Flow Logs by themselves don’t know that, so you need to add that kind of metadata in. And then there’s performance testing, which is sort of the what is. Something we do, Thousand Eyes does, some other people do.

It’s not the actual source of truth, but for example, if you’re having a performance problem getting between, you know, us-east and Azure in the east, well, there’s three other ways you can get there. If your actual traffic isn’t getting there that way, then how do you know which one to use? Well, let’s fire up some tests. There’s all the metrics on what all of the devices are reporting, just like you get metrics from your machines and from your applications, and then there’s stuff even up at the routing layer, which God help you, hopefully you don’t need to actually get in and debug, but sometimes you do. And sometimes, you know, your neighbor tells the mailman that that mail is for me and not for you and they believe them and then you have a big problem when your bills don’t get paid.

The same thing happens in the cloud, the same thing happens on the internet [unintelligible 00:20:52] at the routing. So, the goal is, take all the different sources of it, make it the same within each type, and then pull it all together so you can look at a single place, you can look at a map, you can look at everything, whether it’s the cloud, whether it’s your own data centers, your own WAN, into the internet and in between in a coherent way that understands your application. So, it’s a small task that we’ve bit off, but you know, we have fun solving it.

Corey: Do you find that when you look at customer environments, that they are, and I don’t mean to be disparaging here, truly I don’t, but if you were to ask me to design something today, I would probably not even be using VPCs if I’m doing this completely greenfield. I would be a lot more cloud-first, et cetera, et cetera. Whereas in many cases, that is not the right path, especially if you know, customers have the temerity to not be founded within the last 18 months before AWS existed in some ways. Do you find that the majority of what they’re doing looks like they’re treating the cloud like data centers or do you find that they are leveraging cloud in ways that surprise you and would not be possible in traditional data centers? Because I can’t shake the feeling that the network has a source of truth for figuring out what’s really going on [is very hard to beat 00:22:05].

Avi: Yes, for the most part, to both your assertion at the end and sort of the question. So, in terms of the question, for the most part, people think of VPCs as… you know, they could just equivalent be VLANs and [unintelligible 00:22:21], right? I’ve got policies, and I have these things that are talking to each other, and everything else is not local. And I’ve got—you know, it’s not a perfect mapping to physical interfaces in VLANs but it’s the equivalent of that.

And that is sort of how people think about it. In the data center, you’d call it micro-segmentation, in the cloud, you call it clouding, but you know, just applying all the same policies and saying this stuff can talk to each other and not. Which is always sort of interesting, if you don’t actually know what is talking [laugh] to each other to apply those policies. Which is a lot of what you know, Kentik gets brought in for first. I think where we see the cloud-native thinking, which is overlaid on top of that—you could call it overlay, I guess—which is service mesh.

Now, putting aside the question of what’s going to be a service mesh, what’s going to be a network mesh, where there’s something like [unintelligible 00:23:13] sit, the idea that there’s a way that you look at traffic above the packets at, you know, layers three to more layer seven, that can do things like load balancing, do things like telemetry, do things like policy enforcement, that is a layer that we see very commonly that a lot of the old school folks have—you know, they want their lsu F5s and they want their F5 script. And they’re like, “Why can’t I have this in the cloud?”—which I guess you could buy it from F5 if you really want—but that’s pretty common. Now, not everything’s a sidecar anymore and there’s still debates about what’s going on there, but that’s pretty common, even where the underlying cloud just looks like it could just be a data center.

And that seems to be state of the art, I would say, our traditional enterprise customers, for sure. Our web company customers, and you know, service providers use cloud more for their OTT and some other things. As we work with them, they’re a little bit more likely to be on-prem, you know, historic. But remember, in the enterprise, there’s still a lot of M&A going on, I think that’s even going to pick up in the next couple of years and a lot of what they’re doing is lift-and-shift of [laugh] actual data centers. And my theory is, it’s got to be easier to just make it look like VPCs than completely redo it.

Corey: I’d say that there’s reasons that things are the way that they are. Like, ignoring that this is the better approach from a technical perspective entirely because that’s often not the only answer, it’s we have assurances we made as part of audit compliance regimes, of our SOC 2, of how we handle certain things and what those controls are. And yeah, it’s not hard for even a junior employee, most of the time, to design a reasonable architecture on a whiteboard. The problem is, how do you take something pre-existing and get it to a state that closely resembles that while not turning it off for a long time?

Avi: Right. And I think we’re starting to see some things that probably shouldn’t exist, like, people trying to do VXLAN as overlays into and between VPCs because that was how their data s—you know, they’re more modern on the data center side and trying to do that. But generally, I think people have an understanding they need to be designing architecture for greenfield things that aren’t too far bleeding edge, unless it’s like a pure developer shop, and also can map to the least common denominator kinds of infrastructure that people have. Now, sometimes that may be serverless, which means, you know, more CDN use and abstracted layers in front, but for, you know, running your own components, we see a lot of differences but also a lot of commonality. It’s differences at the micro but commonality the macro. And I don’t know what you see in your practice. So.

Corey: I will say that what I see in practice is that there’s a dichotomy where you have born-in-the-cloud companies where 80% of their spend is on a single workload and you can do a whole bunch of deep optimizations. And then you see the conglomerate approach where it’s giant spend, but it’s all very diffuse across 1500 different applications. And different philosophies, different processes, different cultures give rise to a lot of these things. I will say that if I had a magic wand, I would—and again, the fact that you sponsor and promote this episode is deeply appreciated. Thank you—

Avi: You’re welcome.

Corey: —but it does not mean that you get to compromise my authenticity and objectivity. You can’t buy my opinion, just my attention. But I will say this, that I would love it if my customers used Kentik because it is one of the best things I’ve ever seen to describe what is talking to what that scale and in volume without going super deep into the weeds. Now, obviously, I’m not allowed to start rolling out random things into customer environments. That’s how I get sued to death. But, ugh, I wish it was there.

Avi: You probably shouldn’t set up IAM rules without asking them, yes. That wouldn’t be bad.

Corey: There’s a reason that the only writable stuff that I have access to is generating reports in Cost Explorer.

Avi: [laugh]. Okay.

Corey: Everything else is read-only. All we do is to have conversations with folks. It sets context for those conversations. I used to think that we’d be doing this as a software offering. I no longer believe that actually solves the in-depth problems that people have.

Avi: Well, I appreciate the praise. I even take some of the backhanded praise slash critique at the beginning because we think a lot about, you know, we did design for these complex and often hybrid infrastructures and it’s true, we didn’t design it for the two or four router, you know, infrastructure. If we had bootstrapped longer, if we’d done some other things, we might have done it that way. We don’t want to be exclusionary. It’s just sort of how we focus.

But in the kind of customers that you have, these are things that we’re thinking about what can we do to make it easier to onboard because people have these massive challenges seeing the traffic and understanding it and the cost and security and the performance, but to do more with the VPC Flow Logs, we need to get some of those metrics. We think about should we make an open-source thing. I don’t know how much you’ve seen the concern that people have universally across cloud providers that they turn on something like Kentik, and they’re going to hit their API rate limiter. Which is like, really, you can’t build a cache for that at the scale that these guys run at, the large cloud providers. I don’t really understand that. But it is what it is.

We spent a lot of time thinking about that because of security policy, and getting the kind of metrics that we need. You know, if we open-source some of that, would it make it easier, plug it into people’s observability infrastructure, we’d like to get that onboarding time down, even for those more complex infrastructures. But you know, the payoff is there, you know? It only takes a day of elapsed time and one hour or so. It’s just you got to get a lot of approvals to get the kind of telemetry that you need to make sense of this in some environments.

Corey: Oh, yes. And that’s part of the problem, too, is like, you could talk about one of those big environments where you have 1500 apps all talking to each other. You can’t make sense of any of it without talking to people and having contacts and occasionally get a little bit of [unintelligible 00:29:07] just what these things are named. But at that point, you’re just speculating wildly. And, you know, it’s an engineering trap, where I’m just going to guess rather than asking someone who knows the answer because I don’t want to look foolish. It’s… you just three weeks chasing your own tail. Who’s the foolish one?

Avi: We’re not in a competitive business to yours—

Corey: [laugh].

Avi: But I do often ask when we’re starting off, “So, can you point us at the source of truth that describes what all your applications are?” And usually, they’re, like, “[laugh]. No.” But you know, at the same time to make sense of this stuff, you also need that metadata and that’s something that we designed to be able to take.

Now, Kubernetes has some of that. You may have some of it in ServiceNow, a lot of people use. You may have it in your own text file, CSV somewhere. It may be in NetBox, which we’ve seen people actually use for the cloud, more on the web company and service provider side, but even some traditional enterprise is starting to use it. So, a lot of what we have to do as a vendor is put all that together because yeah, when you’re running across multiple environments and thousands of applications, ultimately scrying at IP addresses and VPC IDs is not going to be sufficient.

So, the good news is, almost everybody has those sources and we just tried to drag it out of them and pull it back together. And for now, we refuse to actually try to get into that business because it’s not a—seems sort of like, you know, SAP where you’re going to be sending consultants forever, and not as interesting as the problems we’re trying to solve.

Corey: I really want to thank you, not just for supporting the show of course, but also for coming here to suffer my slings and arrows. If people want to learn more, where’s the best place for them to find you? And please don’t respond with an IP address.

Avi: 127.0.0.1. You’re welcome at my home at any time.

Corey: There’s no place like localhost.

Avi: There’s no place like localhost. Indeed. So, the company is kentik.com, K-E-N-T-I-K. I am avi@kentik.com. I am@avifriedman on Twitter and LinkedIn and some other things. And happy to chat with nerds, infrastructure nerds, cloud nerds, network nerds, software nerds, debate, maybe not VI versus Emacs, but should you swap space or not, and what should your cloud architecture look like?

Corey: And we will, of course, put links to that in the [show notes 00:31:20].

Avi: Thank you.

Corey: Thank you so much for being so generous with your time. I really appreciate it.

Avi: Thank you for having this forum. And I will let you know when I am down in San Francisco with some time.

Corey: I would be offended if you didn’t take the time to at least say hello. Avi Friedman, CEO at Kentik. I’m Cloud Economist Corey Quinn, and this has been a promoted guest episode of Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a all five-star review on your podcast platform of choice, along with an angry comment saying how everything, start to finish, is somehow because of the network.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

View Details

Jeremy Snyder, Founder of FireTail, joins Corey on Screaming in the Cloud to discuss his career journey and what led him to start FireTail. Jeremy reveals what’s changed in cloud since he was an AE and AWS, and walks through how the need for customization in cloud security has led to a boom in the number of security companies out there. Corey and Jeremy also discuss the costs of cloud security, and Jeremy points out some of his observations in the world of cloud security pricing and packaging.

About Jeremy

Jeremy is the founder and CEO of FireTail.io, an end-to-end API security startup. Prior to FireTail, Jeremy worked in M&A at Rapid7, a global cyber leader, where he worked on the acquisitions of 3 companies during the pandemic. Jeremy previously led sales at DivvyCloud, one of the earliest cloud security posture management companies, and also led AWS sales in southeast Asia. Jeremy started his career with 13 years in cyber and IT operations. Jeremy has an MBA from Mason, a BA in computational linguistics from UNC, and has completed additional studies in Finland at Aalto University. Jeremy speaks 5 languages and has lived in 5 countries. Once, Jeremy went 5 days without seeing another human, but saw plenty of reindeer.

Links Referenced:

  • Firetail: https://firetail.io
  • Email: jeremy@firetail.io

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. My guest today is Jeremy Snyder, who’s the founder at Firetail. Jeremy, thank you for joining me today. I appreciate you taking the time from your day to suffer my slings and arrows.

Jeremy: My pleasure, Corey. I’m really happy to be here.

Corey: So, we’ll get to a point where we talk about what you’re up to these days, but first, I want to dive into the jobs of yesteryear because over a decade ago, you did a stint at AWS doing sales. And not to besmirch your hard work, but it feels like at the time, that must have been a very easy job. Because back then it really felt across the board like the sales motion was basically responding to, “Well, why should we do business with you?” And the response is, “Oh, you misunderstand. You have 87 different accounts scattered throughout your organization. I’m just here to give you visibility, governance, and possibly some discounting over that.” It feels like times have changed in a lot of ways since then. Is that accurate?

Jeremy: Well, yeah, but I will correct a couple of things in there. In my days—

Corey: Oh, please.

Jeremy: —almost nobody had more than one account. I was in the one account, no VPCs, you know, you only separate your workloads by tagging days of AWS. So, our job was a lot, actually, harder at the time because people couldn’t wrap their heads around the lack of subnetting, the lack of workload segregation. All of that was really, like, brand new to people, and so you were trying to tell them like, “Hey, you’re going to be launching something on an EC2 instance that’s in the same subnet as everybody else’s EC2 instance.” And people were really worried about lateral traffic and sniffing and what could their neighbors or other customers on AWS see. And by the way, I mean, this was the customers who even believed it was real. You know, a lot of the conversations we went into with people was, “Oh, so Amazon bought too many servers and you’re trying to sell us excess capacity.”

Corey: That legend refuses to die.

Jeremy: And, you know, it is a legend. That is not at all the genesis of AWS. And you know, the genesis is pretty well publicized at this point; you can go just google, “how did AWS started?” You can find accurate stuff around that.

Corey: I did it a few years ago with multiple Amazon execs and published it, and they said definitively that that story was not true. And you can say a lot about AWS folks, and I assure you, I do, but I also do not catch them lying to my face, ever. And as soon as that changes, well, now we’re going to have a different series of [laugh] conversations that are a lot more pointed. But they’ve earned some trust there.

Jeremy: Yeah, I would agree. And I mean, look, I saw it internally, the way that Amazon built stuff was at such a breakneck pace, that challenge that they had that was, you know, the published version of events for why AWS got created, developers needed a place to test code. And that was something that they could not get until they got EC2, or could not get in a reasonably enough timeframe for it to be, you know, real-time valid or relevant for what was going on with the company. So, you know, that really is the genesis of things, and you know, the early services, SQS, S3, EC2, they all really came out of that journey. But yeah, in our days at AWS, there was a lot of ease, in the sense that lots of customers had pent-up frustrations with their data center providers or their colo providers and lots of customers would experience bursts and they would have capacity constraints and they would need a lot of the features that AWS offered, but we had to overcome a lot of technical misunderstandings and trust issues and, you know, oh, hey, Amazon just wants to sniff our data and they want to see what we’re up to, and explain to them how encryption works and why they have their own keys and all these things. You know, we had to go through a lot of that. So, it wasn’t super easy, but there was some element of it where, you know, just demand actually did make some aspects easy.

Corey: What have you seen change since, well I guess ten years ago and change now? And let’s be clear, you don’t work in AWS sales, but you also are not oblivious to what the market is doing.

Jeremy: For sure. For sure. I left AWS in 2011 and I’ve stayed in the cloud ecosystem pretty much ever since. I did spend some time working for a system integrator where all we did was migrate customers to AWS. And then I spent about five, six years working on cloud security primarily focused on AWS, a lot of GCP, a little bit of Azure.

So yeah, I mean, I certainly stay up to date with what’s going on in the state of cloud. I mean, look, Cloud has evolved from this kind of, you know, developer-centric, very easy-to-launch type of platform into a fully-fledged enterprise IT platform and all of the management structures and all of the kind of bells and whistles that you would want that you probably wanted from your old VMware networks but never really got, they’re all there now. It is a very different ballgame in terms of what the platform actually enables you to do, but fundamentally, a lot of the core building block constructs and the primitives are still kind of driving the heart of it. It’s just a lot of nicer packaging.

What I think is really interesting is actually how customers’ usage of cloud platforms has changed over time. And I always think of it and kind of like the, going back to my days, what did I see from my customers? And it was kind of like the month zero, “I just don’t believe you.” Like, “This thing can’t be real, I don’t trust it, et cetera.” Month one is, I’m going to assign some developer to work on some very low-priority, low-risk workload. In my days, that was SharePoint, by the way. Like, nine times out of ten, the first workload that customers stood up was a SharePoint instance that they had to share across multiple locations.

Corey: That thing falls over all the time anyway. May as well put it in the cloud where it can do so without taking too much else down with it. Was that the thinking or?

Jeremy: Well, and the other thing about it at the time, Corey, was that, like, so many customers worked in this, like, remote-first world, right? And so, SharePoint was inevitably hosted at somebody’s office. And so, the workers at that office were so privileged over the workers everywhere else. The performance gap between consuming SharePoint in one location versus another was like, night and day. So, you know, employees in headquarters were like, “Yeah, SharePoint’s great.” Employees in branch offices were like, “This thing is terrible,” you know? “It’s so slow. I hate it, I hate it, I hate it.”

And so, Cloud actually became, like, this neutral location to move SharePoint to that kind of had an equal performance for every office. And so, that was, I think, one of the reasons and it was also, you know, it had capacity problems, and customers were right at that point, uploading tons of static documents to it, like Word documents, Office attachments, et cetera, and so they were starting to have some of these, like, real disk sprawl problems with SharePoint. So, that was kind of the month one problem. And only after they get through kind of month two, three, and four, and they go through, “I don’t understand my bill,” and, “Help me understand security implications,” then they think about, like, “Hey, should we go back and look at how we’re running that SharePoint stuff and maybe do it more efficiently and, like, move those static Office documents onto S3?” And so on, and so on.

And that’s kind of one of the big things that I’ve changed that I would say is very different from, like, 2011 to now, is there’s enough sophistication around understanding that, like, you don’t just translate what you’re doing in your office or in your data center to what you’re doing on cloud. Or if you do, you’re not getting the most out of your investment.

Corey: I’m curious to get your take on how you have seen cloud adoption patterns differ, specifically tied to geo. I mean, I tend to see it from a world where there’s a bifurcation of between born-in-the-cloud SaaS-type companies where one workload is 80% of their bill or whatnot, and of the big enterprises where the largest single component is 3%. So, it’s a very different slice there. But I’m curious what you would see from a sales perspective, looking across a lot of different geographic boundaries because we’re all, on some level, biased based upon where we tend to spend our time doing business. I’m in San Francisco, which is its very own strange universe that has a certain perspective about itself that is occasionally accurate, but not usually. But it’s a big world out there.

Jeremy: It is. One thing that I would say it’s interesting. I spent my AWS days based in Singapore, living in Singapore at the time, and I was working with customers across Southeast Asia. And to your point, Corey, one of the most interesting things was this little bit of a leapfrog effect. Data centers in Asia-Pac, especially in places like the Philippines, were just terrible.

You know, the Philippines had, like, the second highest electricity rates in Asia at the time, only behind Japan, even though the GDP per capita gap between those two countries is really large. And yet you’re paying, like, these super-high electricity rates. Secondarily, data centers in the Philippines were prone to flooding. And so, a lot of companies in the Philippines never went the data center route. You know, they just hosted servers in their offices, you know, they had a bunch of desktop machines in a cubicle, that kind of situation because, like, data centers themselves were cost prohibitive.

So, you saw this effect a little bit like cell phones in a lot of the developing world. Landline infrastructure was too expensive or never got done for whatever reason, and people went straight to cell phones. So actually, what I saw in a lot of emerging markets in Asia was, screw the data center; we’re going to go straight to cloud. So, I saw a lot of Asia-Pac get a little bit ahead of places like Europe where you had, for instance, a lot of long-term data center contracts and you had customers really locked in. And we saw this over the next, let’s say between, like, say, 2014 and 2018 when I was working with a systems integrator, and then started working on cloud security.

We saw that US customers and Asia-Pac customers didn’t have these obligations; European customers, a lot of them were still working off their lease, and still, you know, I’m locked into let’s just say Equinix Frankfurt for another five years before I can think about cloud migration. So, that’s definitely one aspect that I observed. Second thing I think is, like, the earlier you started, the earlier you reached the point where you realize that actually there is value in a lot of managed services and there actually is value in getting away from the kind of server mindset around EC2.

Corey: It feels like there’s a lot of, I want to call it legacy thinking, in some ways, except that’s unfair because legacy remains a condescending engineering term for something that makes money. The problem that you have is that you get bound by choices you didn’t necessarily realize you were making, and then something becomes revenue-bearing. And now there’s a different way to do it, or you learn more about the platform, or the platform itself evolves, and, “Oh, I’m going to rewrite everything to take advantage of this,” isn’t happening. So, it winds up feeling like, yeah, we’re treating the cloud like a data center. And sometimes that’s right; sometimes that’s a problem, but ultimately, it still becomes a significant challenge. I mean, there’s no way around it. And I don’t know what the right answer is, I don’t know what the fix is going to be, but it always feels like I’m doing something wrong somewhere.

Jeremy: I think a lot of customers go through that same set of feelings and they realize that they have the active runway problem, where you know, how do you do maintenance on an active runway? You kind of can’t because you’ve got flights going in and out. And I think you’re seeing this in your part of the world at SFO with a lot of the work that got done in, like, 2018, 2019 where they kind of had to close down a runway and had, like, near misses because they consolidated all flights onto the one active runway, right? It is a challenge. And I actually think that some of the evolution that I’ve seen our customers go through over the last, like, two, three years, is starting to get away from that challenge.

So, to your point, when you have revenue-bearing workloads that you can’t really modify and things are pretty tightly coupled, it is very hard to make change. But when you start to have it where things are broken down into more microservices, it makes it a lot easier to cycle out Service A for Service B, or let’s say more accurately, Service A1 with Service A2 where you can kind of just, like, plug and play different APIs, and maybe, you know, repoint services at the new stuff as they come online. But getting to that point is definitely a painful process. It does require architectural changes and often those architectural changes aren’t at the infrastructure level; they’re actually inside the application or they’re between things like applications and third-party dependencies where the customers may not have full control over the dependencies, and that does become a real challenge for people to break down and start to attack. You’ve heard of the Strangler Methodology?

Corey: Oh, yes. Both in terms of the Boston Strangler, as well—

Jeremy: [laugh]. Right.

Corey: As the Strangler design pattern.

Jeremy: Yeah, yeah. But I think, like, getting to that is challenging until, like, once you understand that you want to do that, it makes a lot of sense. But getting to the starting point for that journey can be really challenging for a lot of customers because it involves stakeholders that are often not involved on infrastructure conversations, and organizational dysfunction can really creep in there, where you have teams that don’t necessarily play nice together, not for any particular reason, but just because historically they haven’t had to. So, that’s something that I’ve seen and definitely takes a little bit of cultural work to overcome.

Corey: When you take a look across the board of cloud adoption, it’s interesting to have seen the patterns that wound up unfolding. Your career path, though, seem to have gotten away from the selling cloud and into some strange directions leading up to what you’re doing now, where you founded Firetail. What do you folks do?

Jeremy: We do API security. And it really is kind of the culmination of, like, the last several years and what we saw. I mean, to your point, we saw customers going through kind of Phase One, Two, Three of cloud adoption. Phase One, the, you know, for lack of a better phrase, lift-and-shift and Phase Two, the kind of first step on the path towards quote-unquote, “Enlightenment,” where they start to see that, like, actually, we can get better operational efficiency if we, you know, move our databases off of EC2 and on to RDS and we move our static content onto S3.

And then Phase Three, where they realize actually EC2 kind of sucks, and it’s a lot of management overhead, it’s a lot of attack surface, I hate having to bake AMIs. What I really want to do is just drop some code on a platform and run my application. And that might be serverless. That might be containerized, et cetera. But one path or the other, where we pretty much always see customers ending up is with an API sitting on a network.

And that API is doing two things. It front-ends a data set and at front-ends a set of functionality, and most cases. And so, what that really means is that the thing that sits on the network that does represent the attack surface, both in terms of accessing data or in terms of let’s say, like, abusing an application is an API. And that’s what led us to where I am today, what led me and my co-founder Riley to, you know, start the company and try to make it easier for customers to build more secure APIs. So yeah, that’s kind of the change that I’ve observed over the last few years that really, as you said, lead to what I’m doing now.

Corey: There is a lot of, I guess, challenge in the entire space when we bound that to—even API security, though as soon as you going down the security path it starts seeming like there’s a massive problem, just in terms of proliferation of companies that each do different things, that each focus on different parts of the story. It feels like everything winds up spitting out huge amounts of security-focused, or at least security-adjacent telemetry. Everything has findings on top of that, and at least in the AWS universe, “Oh, we have a service that spits out a lot of that stuff. We’re going to launch another service on top of it that, of course, cost more money that then winds up organizing it for you. And then another service on top of that that does the same thing yet again.” And it feels like we’re building a tower of these things that are just… shouldn’t just be a feature in the original underlying thing that turns down the noise? “Well, yes, but then we couldn’t sell you three more things around it.”

Jeremy: Yeah, I mean—

Corey: Agree? Disagree?

Jeremy: I don’t entirely disagree. I think there is a lot of validity on what you just said there. I mean, if you look at like the proliferation of even the security services, and you see GuardDuty and Config and Security Hub, or things like log analysis with Athena or log analysis with an ELK stack, or OpenSearch, et cetera, I mean, you see all these proliferation of services around that. I do think the thing to bear in mind is that for most customers, like, security is not a one size fits all. Security is fundamentally kind of a risk management exercise, right? If it wasn’t a risk management exercise, then all security would really be about is, like, keeping your data off of networks and making sure that, like, none of your data could ever leave.

But that’s not how companies work. They do interact with the outside world and so then you kind of always have this decision and this trade-off to make about how much data you expose. And so, when you have that decision, then it leads you down a path of determining what data is important to your organization and what would be most critical if it were breached. And so, the point of all of that is honestly that, like, security is not the same for you as it is for me, right? And so, to that end, you might be all about Security Hub, and Config instead of basic checks across all your accounts and all your active regions, and I might be much more about, let’s say I’m quote-unquote, “Digital-native, cloud-native,” blah, blah, blah, I really care about detection and response on top of events.

And so, I only care about log aggregation and, let’s say, GuardDuty or Athena analysis on top of that because I feel like I’ve got all of my security configurations in Infrastructure as Code. So, there’s not a right and wrong answer and I do think that’s part of why there are a gazillion security services out there.

Corey: On some level, I’ve been of the opinion for a while now that the cloud providers themselves should not necessarily be selling security services directly because, on some level, that becomes an inherent conflict of interest. Why make the underlying platform more secure or easier to use from a security standpoint when you can now turn that into a revenue source? I used to make comments that Microsoft Defender was a classic example of getting this right because they didn’t charge for it and a bunch of antivirus companies screamed and whined about it. And then of course, Microsoft’s like, “Oh, Corey saying nice things about us. We can’t have that.” And they started charging for it. So okay, that more or less completely subverts my entire point. But it still feels squicky.

Jeremy: I mean, I kind of doubt that’s why they started charging for it. But—

Corey: Oh, I refuse to accept that I’m not that influential. There we are.

Jeremy: [laugh]. Fair enough.

Corey: Yeah, I just can’t get away from the idea that it feels squicky when the company providing the infrastructure now makes doing the secure thing on top of it into an investment decision.

Jeremy: Yeah.

Corey: “Do you want the crappy, insecure version of what we build or do you want the top-of-the-line secure version?” That shouldn’t be a choice people have to make. Because people don’t care about security until right after they really should have cared about security.

Jeremy: Yeah. Look, and I think the changes to S3 configuration, for instance, kind of bear out your point. Like, it shouldn’t be the case that you have to go through a lot of extra steps to not make your S3 data public, it should always be the case that, like, you have to go through a lot of steps if you want to expose your data. And then you have explicitly made a set of choices on your own to make some data public, right? So, I kind of agree with the underlying logic. I think the counterargument, if there is one to be made, is that it’s not up to them to define what is and is not right for your organization.

Because again, going back to my example, what is secure for you may not be secure for me because we might have very different modes of operation, we might have very different modes of building our infrastructure, deploying our infrastructure, et cetera. And I think every cloud provider would tell you, “Hey, we’re just here to enable customers.” Now, do I think that they could be doing more? Do I think that they could have more secure defaults? You know, in general, yes, of course, they could. And really, like, the fundamentals of what I worry about are people building insecure applications, not so much people deploying infrastructure with bad configurations.

Corey: It’s funny, we talk about this now. Earlier today, I was lamenting some of the detritus from some of my earlier builds, where I’ve been running some of these things in my old legacy single account for a while now. And the build service is dramatically overscoped, just because trying to get the security permissions right, was an exercise in frustration at the time. It was, “Nope, that’s not it. Nope, blocked again.”

So, I finally said to hell with it, overscope it massively, and then with a, “Todo: fix this later,” which of course, never happened. And if there’s ever a breach on something like that, I know that I’ll have AWS wagging its finger at me and talking about the shared responsibility model, but it’s really kind of a disaster plan of their own making because there’s not a great way to say easily and explicitly—or honestly, by default the way Google Cloud does—of okay, by default, everything in this project can talk to everything in this project, but the outside world can’t talk to any of it, which I think is where a lot of people start off. And the security purists love to say, “That’s terrible. That won’t work at a bank.” You’re right, it won’t, but a bank has a dedicated security apparatus, internally. They can address those things, whereas your individual student learner does not. And that’s how you wind up with open S3 bucket monstrosities left and right.

Jeremy: I think a lot of security fundamentalists would say that what you just described about that Google project structure, defeats zero trust, and you know, that on its own is actually a bad thing. I might counterargue and say that, like, hey, you can have a GCP project as a zero trust, like, first principle, you know? That can be the building block of zero trust for your organization and then it’s up to you to explicitly create these trust relationships to other projects, and so on. But the thing that I think in what you said that really kind of does resonate with me in particular as an area that AWS—and really this case, just AWS—should have done better or should do better, is IAM permissions. Because every developer in the world that I know has had that exact experience that you described, which is, they get to a point where they’re like, “Okay, this thing isn’t working. It’s probably something with IAM.”

And then they try one thing, two things, and usually on the third or fourth try, they end up with a star permission, and maybe a comment in that IAM policy or maybe a Jira ticket that, you know, gets filed into backlog of, “Review those permissions at some point in the future,” which pretty much never happens. So, IAM in particular, I think, is one where, like, Amazon should do better, or should at least make it, like, easy for us to kind of graphically build an IAM policy that is scoped to least permissions required, et cetera. That one, I’ll a hundred percent agree with your comments and your statement.

Corey: As you take a look across the largest, I guess, environments you see, and as well as some of the folks who are just getting started in this space, it feels like, on some level, it’s two different universes. Do you see points of commonality? Do you see that there is an opportunity to get the individual learner who’s just starting on their cloud journey to do things that make sense without breaking the bank that they then can basically have instilled in them as they start scaling up as they enter corporate environments where security budgets are different orders of magnitude? Because it seems to me that my options for everything that I’ve looked at start at tens of thousands of dollars a year, or are a bunch of crappy things I find on GitHub somewhere. And it feels like there should be something between those two.

Jeremy: In terms of training, or in terms of, like, tooling to build—

Corey: In terms of security software across the board, which I know—

Jeremy: Yeah.

Corey: —is sort of a vague term. Like, I first discovered this when trying to find something to make sense of CloudTrail logs. It was a bunch of sketchy things off GitHub or a bunch of very expensive products. Same thing with VPC flow logs, same thing with trying to parse other security alerting and aggregate things in a sensible way. Like, very often it’s, oh, there’s a few very damning log lines surrounded by a million lines of nonsense that no one’s going to look through. It’s the needle in a haystack problem.

Jeremy: Yeah, well, I’m really sorry if you spent much time trying to analyze VPC flow logs because that is just an exercise in futility. First of all, the level of information that’s in them is pretty useless, and the SLA on actually, like, log delivery, A, whether it’ll actually happen, and B, whether it will happen in a timely fashion is just pretty much non-existent. So—

Corey: Oh, from a security perspective I agree wholeheartedly, but remember, I’m coming from a billing perspective, where it’s—

Jeremy: Ah, fair enough.

Corey: —huh, we’re taking a petabyte in and moving 300 petabytes between availability zones. It’s great. It’s a fun game called find whatever is chatty because, on some level, it’s like, run two of whatever that is—or three—rather than having it replicate. What is the deal here? And just try to identify, especially in the godforsaken hellscape that is Kubernetes, what is that thing that’s talking? And sometimes flow logs are the only real tool you’ve got, other than oral freaking tradition.

Jeremy: But God forbid you forgot to tag your [ENI 00:24:53] so that the flow log can actually be attributed to, you know, what workload is responsible for it behind the scenes. And so yeah, I mean, I think that’s a—boy that’s a case study and, like, a miserable job that I don’t think anybody would really want to have in this day and age.

Corey: The timing of this is apt. I sent out my newsletter for the week a couple hours before this recording, and in the bottom section, I asked anyone who’s got an interesting solution for solving what’s talking to what with VPC flow logs, please let me know because I found this original thing that AWS put up as part of their workshops and a lab to figure this out, but other than that, it’s more or less guess-and-check. What is the hotness? It’s been a while since I explored the landscape. And now we see if the audience is helpful or disappoints me. It’s all on you folks.

Jeremy: Isn’t the hotness to segregate every microservice into an account and run it through a load balancer so that it’s like much more properly tagged and it’s also consumable on an account-by-account basis for better attribution?

Corey: And then everything you see winds up incurring a direct fee when passing through that load balancer, instead of the same thing within the same subnet being able to talk to one another for free.

Jeremy: Yeah, yeah.

Corey: So, at scale—so yes, for visibility, you’re absolutely right. From a, I would like to spend less money giving it directly to Amazon, not so much.

Jeremy: [unintelligible 00:26:08] spend more money for the joy of attribution of workload?

Corey: Not to mention as well that coming into an environment that exists and is scaled out—which is sort of a prerequisite for me going in on a consulting project—and saying, “Oh, you should rebuild everything using serverless and microservice principles,” is a great way to get thrown out of the engagement in the first 20 minutes. Because yes, in theory, anyone can design something great, that works, that solves a problem on a whiteboard, but most of us don’t get to throw the old thing away and build fresh. And when we do great, I’m greenfielding something; there’s always constraints and challenges down the road that you don’t see coming. So, you finally wind up building the most extensible thing in the universe that can handle all these things, and your business dies before you get to MVP because that takes time, energy and effort. There are many more companies that have died due to failure to find product-market fit than have died because, “Oh hey, your software architecture was terrible.” If you hit the market correctly, there is budget to fix these things down the road, whereas your code could be pristine and your company’s still dead.

Jeremy: Yeah. I don’t really have a solution for you on that one, Corey [laugh].

Corey: [laugh].

Jeremy: I will come back to your one question—

Corey: I was hoping you did.

Jeremy: Yeah, sorry. I will come back to the question about, you know, how should people kind of get started in thinking about assessing security. And you know, to your point, look, I mean, I think Config is a low-ish cost, but should it cost anything? Probably not, at least for, like, basic CIS foundation benchmark checks. I mean, like, if the best practice that Amazon tells everybody is, “Turn on these 40-ish checks at last count,” you know, maybe those 40-ish checks should just be free and included and on in everybody’s account for any account that you tag as production, right?

Like, I will wholeheartedly agree with that sentiment, and it would be a trivial thing for Amazon to do, with one kind of caveat—and this is something that I think a lot of people don’t necessarily understand—collecting all the required data for security is actually really expensive. Security is an extremely data-intensive thing at this day and age. And I have a former coworker who used to hate the expression that security is data science, but there is some truth in it at this point, other than the kind of the magic around it is not actually that big because there’s not a lot of, let’s say, heuristic analysis or magic that goes into what queries, et cetera. A lot of security is very rule-based. It’s a lot of, you know, just binary checks: is this bit set to zero or one?

And some of those things are like relatively simple, but what ends up inevitably happening is that customers want more out of it. They don’t just want to know, is my security good or bad? They want to know things like is it good or bad now relative to last week? Has it gotten better or worse over time? And so, then you start accumulating lots of data and time series data, and that becomes really expensive.

And secondarily, the thing that’s really starting to happen more and more in the security world is correlation of multiple layers of data, infrastructure with applications, infrastructure with operating system, infrastructure with OS and app vulnerabilities, infrastructure plus vulnerabilities plus Kubernetes configurations plus API sitting at the edge of that. Because realistically, like, so many organizations that are built out at scale, the truth of the matter is, is just like on their operating system vulnerabilities, they’re going to have tens of thousands, if not millions of individual items to deal with and no human can realistically prioritize those without some context around it. And that is where the data, kind of, management becomes really expensive.

Corey: I hear you. Particularly the complaints about AWS Config, which many things like Control Tower setup for you. And on some level, it is a tax on using the cloud as the cloud should be used because it charges for evaluation of changes to your environment. So, if you’re spinning things up all the time and then turning them down when they’re not in use, that incurs a bunch of Config charges, whereas if you’ve treat it like a big dumb version of your data center where you just spin [unintelligible 00:30:13] things forever, your Config charge is nice and low. When you start seeing it entering the top ten of your spend on services, something is very wrong somewhere.

Jeremy: Yeah. I would actually say, like, a good compromise in my mind would be that we should be included with something like business support. If you pay for support with AWS, why not include Config, or some level of Config, for all the accounts that are in scope for your production support? That would seem like a very reasonable compromise.

Corey: For a lot of folks that have it enabled but they don’t see any direct value from it either, so it’s one of those things where not knowing how to turn it off becomes a tax on what you’re doing, in some cases. In SCPs, but often with Control Tower don’t allow you to do that. So, it’s your training people who are learning this in their test environments to avoid it, but you want them to be using it at scale in an enterprise environment. So, I agree with you, there has to be a better way to deliver that value to customers. Because, yeah, this thing is now, you know, 3 or 4% of your cloud bill, it’s not adding that much value, folks.

Jeremy: Yeah, one thing I will say just on that point, and, like, it’s a super small semantic nitpick that I have, I hate when people talk about security as a tax because I think it tends to kind of engender the wrong types of relationships to security. Because if you think about taxes, two things about them, I mean, one is that they’re kind of prescribed for you, and so in some sense, this kind of Control Tower implementation is similar because, like you know, it’s hard for you to turn off, et cetera, but on the other hand, like, you don’t get to choose how that tax money is spent. And really, like, you get to set your security budget as an organization. Maybe this Control Tower Config scenario is a slight outlier on that side, but you know, there are ways to turn it off, et cetera.

The other thing, though, is that, like, people tend to relate to tax, like, this thing that they really, really hate. It comes once a year, you should really do everything you can to minimize it and to, like, not spend any time on it or on getting it right. And in fact, like, there’s a lot of people who kind of like to cheat on taxes, right? And so, like, you don’t really want people to have that kind of mindset of, like, pay as little as possible, spend as little time as possible, and yes, let’s cheat on it. Like, that’s not how I hope people are addressing security in their cloud environments.

Corey: I agree wholeheartedly, but if you have a service like Config, for example—that’s what we’re talking about—and it isn’t adding value to you, and you just you don’t know what it does, how it works, than it [unintelligible 00:32:37]—or more or less how to turn it off, then it does effectively become directly in line of a tax, regardless of how people want to view the principle of taxation. It’s a—yeah, security should not be a tax. I agree with you wholeheartedly. The problem is, is it is—

Jeremy: It should be an enabler.

Corey: —unclea—yeah, the relationship between Config and security in many cases is fairly attenuated in a lot of people’s minds.

Jeremy: Yeah. I mean, I think if you don’t have, kind of, ideas in mind for how you want to use it or consume it, or how you want to use it, let’s say as an assessment against your own environment, then it’s particularly vexing. So, if you don’t know, like, “Hey, I’m going to use Config. I’m going to use Config for this set of rules. This is how I’m going to consume that data and how I’m going to then, like, pass the results on to people to make change in the organization,” then it’s particularly useless.

Corey: Yeah. I really want to thank you for taking the time to speak with me. If people want to learn more, where’s the best place for them to find you?

Jeremy: Easy, breezy. We are just firetail.io. That’s ‘fire’ like the, you know, flaming substance, and ‘tail’ like the tail of an animal, not like a story. But yeah, just firetail.io.

And if you come now, we’ve actually got, like, a white paper that we just put out around API security and kind of analyzing ten years of API-based data breaches and trying to understand what actually went wrong in most of those cases. And you’re more than welcome to grab that off of our website. And if you have any questions, just reach out to me. I’m just jeremy@firetail.io.

Corey: And we’ll put links to all of that in the [show notes 00:34:03]. Thank you so much for your time. I appreciate it.

Jeremy: My pleasure, Corey. Thanks so much for having me.

Corey: Jeremy Snyder, founder and CEO at Firetail. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an angry comment pointing out that listening to my nonsense is a tax on you going about your day.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

View Details

Kenneth Rose, CTO at OpsLevel, joins Corey on Screaming in the Cloud to discuss how OpsLevel is helping developer teams to scale effectively. Kenneth reveals what a developer portal is, how he thinks about the functionality of a developer portal, and the problems a developer portal solves for large developer teams. Corey and Kenneth discuss how to drive adoption of a developer portal, and Kenneth explains why it’s so necessary to have executive buy-in throughout that process. Kenneth also discusses how using their own portal internally along with seeking out customer feedback has allowed OpsLevel to make impactful innovations.

About Ken

Kenneth (Ken) Rose is the CTO and Co-Founder of OpsLevel. Ken has spent over 15 years scaling engineering teams as an early engineer at PagerDuty and Shopify. Having in-the-trenches experience has allowed Ken a unique perspective on how some of the best teams are built and scaled and lends this viewpoint to building products for OpsLevel, a service ownership platform built to turn chaos into consistency for engineering leaders.

Links Referenced:

  • OpsLevel: https://www.opslevel.com/
  • LinkedIn: https://www.linkedin.com/company/opslevel/
  • Twitter: https://twitter.com/OpsLevelHQ

Transcript
Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud, I’m Corey Quinn, about, oh I don’t know, two years ago and change, I wound up writing a blog post titled, “Developer Portals are An Anti Pattern,” and I haven’t really spent a lot of time thinking about them since. This promoted guest episode is brought to us by our friends at OpsLevel, and they have sent their CTO and co-founder Ken Rose, presumably in an attempt to change my perspective on these things. Let’s find out. Ken, thank you for agreeing to, well, run the gauntlet, for lack of a better term.

Ken: Hey, Corey. Thanks again for having me. And I’ve heard, you know, heard and listened to your show a bunch, and really excited to be here today.

Corey: Let’s begin with defining our terms. I’m curious to know what a developer portal is. ‘What would you say a developer portal means to you?’ Like it’s a college entrance essay.

Ken: Right? Definitely. You know, so really, a developer portal is this consolidated place for developers to come to, especially in large organizations to be able to get their jobs done more easily, right? A large challenge that developers have in large organizations, there’s just a lot to do and a lot to take care of. So, a developer portal is a place for developers to be able to better own, manage, and run the services, they’re responsible for that run in production, and they can do that through access, easy access to self-service tooling.

Corey: I guess, on some level, this turns into one of those alignment charts of, like, what is a database and, like, how prescriptive you want to be. It’s like, well is a senior engineer a database because you can query them and they have information? Would you consider, for example, Kubernetes be a developer platform, and/or would the AWS console?

Ken: Yeah, that’s actually an interesting question, right? So, I think there’s actually two—we’re going to get really niggly here—there’s developer platform and developer portal, right? And the word portal for me is something that sits above a developer platform. I don’t know if you remember, like, the late-90s, early-2000s, like, portals were all the rage.

Like, Yahoo and AltaVistas were like search portals, they were trying to, at the time, consolidate all this information on a much smaller internet to make it easy to access. A developer portal is sort of the same thing, but custom-built for developers and trying to consolidate a lot of the tooling that exists. Now, in terms of the AWS console? Yeah, maybe. Like, it has a suite of tools and suite of offerings. It doesn’t do a lot on the well, how do I quickly find out what’s running in production and who is responsible for it? I don’t know, unless AWS shipped, like, their, you know, three-hundredth new offering in the last week that I haven’t, you know, kept on top of.

But you know, there’s definitely some spectrum in terms of what goes into a developer portal. For me, there’s kind of three main things you need. You do need some kind of a catalog, like, what’s out there who owns it; you need some kind of a way to measure, like, how good are those services, like, how well built are they; and then you need some access to self-service tooling. And that last part is where, like, the Kubernetes or AWS could be, you know, sort of a dev portal as well.

Corey: My experience with developer portals—there was a time when I loved it. RightScale was what I used—at some depth—back in I want to say 2010, 2011 because the EC2 console was clearly not built or designed by anyone who had not built EC2 themselves with their bare hands and sweat of their brow. And in time, the EC2 console got better where it wasn’t written in hieroglyphics, as best we could tell, and it became ‘click button to launch instance.’ And RightScale really didn’t have a second act and they wound up getting acquired by our friends over at Flexera years later. And I haven’t seen their developer portal in at least eight years as a direct result of this.

So, the problem, at least when I was viewing it purely in the context of AWS services, it feels like you are competing against AWS iterating forward on developer experience, which they iterate slowly, sometimes, and unevenly across their breadth of services, but it does feel like at some level by building an internal portal, you are, first, trying to out-innovate AWS, in some ways, and two, you are inherently making the trade-off of not using recent features and enhancements that have not themselves been incorporated into the portal. That’s where the, I guess the start, the genesis of my opposition to the developer portal approach comes from. Is that philosophy valid these days? Not as much. Because I can see an argument for it shifting.

Ken: Yeah, I think it’s slightly different. I think of a developer portal as again, it’s something that sort of sits on top of AWS or Google Cloud or whatever cloud provider use, right? You give an example for example with RightScale and EC2. So, provisioning instances is one part of the activity you have to do as a developer. Now, in most modern organizations, you have, like, your product developers that ship features. They don’t actually care about provisioning instance themselves. There are another group called the platform engineers or platform group that are responsible for building automation and tooling to help spin up instances and create CI/CD pipelines and get everything you need set up.

And they might use AWS under the covers to do that, but the automation built on top and making that accessible to developers, that’s really what a developer portal can provide. In addition, it also provides links to operational tooling that you need, technical documentation, it’s everything you need as a developer to do your job, in one place. And though AWs bills itself is that, I think of them as more, they have a lot of platform offerings, right, they have a lot of infra-offerings, but they still haven’t been able to, I think, customize that, unless you’re an organization that builds—that has kind of gone in-all on AWS and doesn’t build any of your own tooling, that’s where a developer portal helps. It really helps by consolidating all that information in one place, by making that information discoverable for every developer so they have less… less cognitive load, right? We’ve asked developers to kind of do too much that we don’t… we’ve asked to shift left and well, how do we make that information more accessible?

Regarding the point of, you know, AWS adds new features or new capabilities all the time and, like, well you have this dev portal, that’s sort of your interface for how to get things done. Like, how will you use those? Dev portal doesn’t stop you from doing that, right? So, my mental model is, if I’m a developer, and I want to spin up a new service, I can just press a button inside of my dev portal in my company and do that. And I have a service that is built according to the latest standards, it has a CI/CD pipeline, it already has a—you know, it’s registered in PagerDuty, it’s registered in Datadog, it has all the various bits.

And then there’s something else that I want to do that isn’t really on the golden path because maybe this is some new service or some experiment, nothing stops us from doing that. Like, you still can use all those tools from AWS, you know, kind of raw. And if those prove to be valuable for the rest of the organization, great. They can make their way into the dev portal; they can actually become a source of leverage. But if they’re not, then they can also just sit there on the vine. Like, not everything that eight of us ever produces will be used by every company.

Corey: Many years ago, I got a Cisco pair of certifications because recession was hitting and I needed to be better at networking. And taking those certifications, in those days before Cisco became the sad corporate dragon with no friends we all know today, they were highly germane and relevant. But I distinctly remember, even now, 15 years later, that there was this entire philosophy of pretend that the entire world is Cisco only, which in networking is absolutely never true. It feels like a lot of the AWS designs and patterns tend to assume, oh yeah, you’re going to use AWS services for everything. I have never yet found that to be true, other than when I’m just trying to be obstinate.

And hell is interoperability between a bunch of different things. Yes, I may want to spin up an EC2 instance and an AWS load balancer and some S3 storage or whatnot, but I’m also going to want to monitor it with PagerDuty, I’m going to want to have a CDN that isn’t CloudFront because most CDN these days don’t hate you in quite the same economic ways and are simpler to work with, et cetera, et cetera, et cetera. So, there’s definitely a story wherein I’ve found that there’s an—the interoperability of tying these things together is helpful. How do you avoid falling down the trap of oh, everyone should be multi-cloud, single pane of glass at cetera, et cetera? In practice that always seems to turn to custard.

Ken: Yeah, I think multi-cloud and single pane of glass are actually two different things. So multi-cloud, like, I agree with you to some sense. Like, pick a cloud and go with it, like, unless you have really good business reasons to go for multi-cloud. And sometimes you do, like, years ago, I worked at PagerDuty, they were multi-cloud for a reliability reason, that hey, if one cloud provider goes down, you don’t want [crosstalk 00:08:40]—

Corey: They were an example I used all the time for that story—

Ken: Right.

Corey: —specifically the thing woke you up was homed in a bunch of different places, whereas the marketing site, the onboarding flow, the periphery stuff around it was not because it didn’t need to be.

Ken: Exactly.

Corey: Like, the core business need of wake you up was very much multi-cloud because once upon a time, it wasn’t and it went down with the rest of us-east-1 and people weren’t woken up to be told their site was on fire.

Ken: A hundred percent. And on the kind of like application side where, even then, pick a cloud and go with it, unless there’s a really compelling business reason for your business to go multi-cloud. Maybe there’s something credits or compliance or availability, right? There might be reasons, but you have to be articulate about whether they’re right for you.

Now, single pane of glass, I think that’s different, right? I do think that’s something that, ultimately, is a net boon for developers. In any large organization, there is a myriad of internal tools that have been built. And it’s like, well, how do I provision a new topic in the Kafka cluster? How do I actually get access to the AWS console? How do I spin up a new service, right? How do I kind of do these things?

And if I’m a developer, I just want to ship features. Like, that’s what I’m incented to do, that’s what I’m optimizing for. And all this other stuff I have to do as part of my job, but I don’t want to have to become, like, a Kubernetes guru to be able to do it, right? So, what a developer portal is trying to do is be that single pane of glass, bringing all these common set of tools and responsibilities that you have as a developer in one place. They’re easy to search for, they’re easy to find, they’re easy to query, they’re easy to use.

Corey: I should probably have asked this earlier on, but let’s disambiguate for a little bit here. Because when I’m setting up to use a new service or product and kick the tires on it, no two explorations really look the same. Whereas at most responsible mature companies that are building products that are—services that are going to production use, they’ve standardized around a number of different approaches. What does your target customer look like? Is there a certain point of scale, a certain level of complexity, a certain maturity of process?

Ken: Absolutely. So, a tool like OpsLevel or a developer portal really only makes sense when you hit some critical mass in terms of the number of services you have running in production, or the number of developers that you have. So, when you hit 20, 30, 50 developers or 20, 30, 50 services, an important part of a developer portal is this catalog of what’s out there. Once you kind of hit the Dunbar number of services, like, when you have more than you keep in your head, that’s when you start to need tooling like this. If you look at our customer base, they’re all you know, kind of medium to large-sized companies. If you’re a startup with, like, ten people, OpsLevel is probably not right for you. We use all playable internally at OpsLevel, and you know, like, we’re still a small company. It’s like, we make it work for us because we know how to get the most out of it, but like, it’s not the perfect fit because it’s not really meant for, you know, smaller companies.

Corey: Oh, I hear you. I think I’m probably… I have a better AWS bill analytic system running internally here at The Duckbill Group than some banks do. So, I hear you on that front.

Ken: I believe it.

Corey: But also implies to me that there’s no OpsLevel prospect or customer deployment that has ever been greenfield. It’s always you’re building existing things, there’s already infrastructure in place, vendors have been selected across the board. You aren’t—don’t to want to starting a company day one, they’re going to all right, time to spin up our AWS account and we’re also going to wind up signing up for OpsLevel, from the sound of it.

Ken: Correct—

Corey: Accurate? Inaccurate?

Ken: I think that’s actually accurate. Like, a lot of the problems, we solve other problems that come as you start to scale both your product and your engineering team. And it’s the problems of complexity.

Corey: What do those painful problems look like? In other words, what is someone sitting at home right now listening to this, or driving to work debating whether want to ram a bridge abutment or go into the office depending on their mental state today, what painful problem did they have that OpsLevel is designed to fix?

Ken: Yeah, for sure. So, let’s help people self-select. So, here’s my mental model for any [unintelligible 00:12:25]. There are product developers, platform developers, and engineering leaders. Product developers, if you’re asking questions like, “I just got paged for the service. I don’t know what this does.” Or, “It’s upstream from here. Where do I find the technical documentation?” Or, “I think I have to do something with the payment service. Where do I find the API for that?”

You know, when you get to that scale, a developer portal can help you. If you’re a platform engineer and you have questions like, “Okay, we got to migrate. We’re migrating, I don’t know, from Datadog to Honeycomb, right? We got to get these fifty or a hundred or thousands of services and all these different owners to, like, switch to some new tool.” Or, “Hey, we’ve done all this work to ship the golden path. Like, how to actually measure the adoption of all this work that we’re doing and if it’s actually valuable?” Right?

Like, we want everybody to be on a certain set of CI tooling or a certain minimum version of some library or framework. How do we do that? How do we measure that? OpsLevel is for you, right? We have a whole bunch of stuff around maturity.

And if you’re engineering leader, ultimately, questions you care about, like, “How fast are my developers working? I have this massive team, we’ve made this massive investment in hiring all these humans to write software and bring value for our customers. How can we be more efficient as a business in terms of that value delivery?” And that’s where OpsLevel can help as well.

Corey: Guardrails, whether they be economic, regulatory, or otherwise, have to make it easier than doing things incorrectly because one of the miracle aspects of cloud also turns into a bit of a problem, which is shadow IT is only ever a corporate credit card away. Make it too difficult to comply with corporate policies and people won’t. And they’re good actors; they’re trying to get work done. They’re not trying to make people’s lives harder, but they don’t want to spend six weeks provisioning an EC2 cluster. So, there’s always that weird trade-off.

Now, it feels—and please correct me if I’m wrong—once someone has rolled out OpsLevel at their organization, where it really shines is spinning up a new service where okay, great, you’re going to spin up the automatic observability portion of it, you’re going to spin up the underlying infrastructure in certain ways that comply with our policies, it’s going to build the CI/CD pipelines around it, you’re going to wind up having the various cost instrumentation rolled out to it. But for services that are already excellent within the environment, is there an OpsLevel story for them?

Ken: Oh, absolutely. So, I look at it as, like, the first problem OpsLevel helps solve is the catalog and what’s out there and who owns it. So, not even getting developers to spin up new services that are kind of on the golden path, but just understanding the taxonomy of what are the services we have? How do those services compose into higher-level things like systems or domains? What’s the whole set of infrastructure we have?

Like, I have 50 AWS accounts, maybe a handful of GCP ones, also, some Azure. I have all this infrastructure that, like, how do I start to get a handle on, like, what’s out there in prod and who’s responsible for it. And that helps you get in front of compliance risks, security risks. That’s really the starting point for OpsLevel building that catalog. And we have a bunch of integrations that kind of slurp all this data to automatically assemble that catalog, or YAML as well if that’s your thing. But that’s the starting point is building that catalog and figuring out this assignment of, like, okay, this service and this human, or this—sorry—team, like, they’re paired together.

Corey: A number of offerings in this space, which honestly, my exposure to it is bounded simultaneously to things that are ten years old and no one uses anymore, or a bunch of things I found on GitHub. And the challenge that both of those products tend to have is that they assume certain things to be true about a given environment: that they’re using Terraform to manage everything, or they’re always going to be using CloudFormation, or everyone there knows Python or something else like that. What are the prerequisites to get started with OpsLevel?

Ken: Yeah, so we worked pretty hard to build just a ton of integrations. I would say integrations is are just continuing thing we have going on in the background. Like, when we started, like, we only supported a GitHub. Now, we support all the gits, you know, like GitHub, GitLab, Bitbucket, Azure DevOps, like, we’re building [unintelligible 00:16:19]. There’s just a whole, like, long tail of integrations.

The same with APM tooling. The same with vulnerability management tooling, right? And the reason we do that is because there’s just this huge vendor footprint, and people, you know, want OpsLevel to work for them. Now, the other thing we try to do is we also build APIs. So, anything we have as, like, a core integration, we also have kind of like an underlying API for, so that there’s, no matter what you have an escape hatch. If like, you’re using some tool that we don’t support or you have some homegrown thing, there’s always a way to try to be able to integrate that into OpsLevel.

Corey: When people think about developer portals, the most common one that pops to mind is Backstage, which Spotify wound up building, internally, championing, open-sourcing, and I believe, on some level, turned into a product because if there’s one thing people want, it’s to have their podcast music company become a SaaS vendor, which is weird to me. But the criticisms that I’ve seen about and across the board have rung relatively true, including from people internal at Spotify who have used the thing, which is, well first is underestimating the amount of effort that is necessary to maintain Backstage itself, that the build versus buy discussion is always harder to bu—engineers love to build, but they shouldn’t be building things outside of their core competency half the time, and the other is driving adoption within the org where you can have the most amazing developer portal in the known universe, but if people don’t use it, it may as well not exist and doing the carrot and stick approach often doesn’t work. I think you have a pretty good answer that I need not even ask you to elaborate on, “Well, how do we avoid having to maintain this ourselves,” since you have a company that does this, but how do you find companies are driving adoption successfully once they have deployed OpsLevel?

Ken: Yeah, that’s a great question. So, absolutely. Like, I think the biggest thing you need first, is kind of cultural buy-in and that this is a tool that we want to invest in, right? I think one of the reasons Spotify was successful with Backstage and I think it was System Z before that was that they had this kind of flywheel of, like, they saw that their developers were getting, you know better faster, working happier, by using this type of tooling, by reducing the cognitive load. The way that we approach it is sort of similar, right?

We want to make sure that there is executive buy-in that, like, everybody agrees this is, like, a problem that’s worth solving. The first step we do is trying to build out that catalog again and helping assign ownership. And that helps people understand, like, hey, these are the services I’m responsible for. Oh, look, and now here’s this other context that I didn’t have before. And then helping organizations, you know, what—it depends on the problem we’re trying to solve, but whether it’s rolling out self-serve automation to help developers, like, reduce what was before a ton of cognitive load or if it’s helping platform teams define what good looks like so they can start to level up the overall health of what’s running in production, we kind of work on different problems, but it’s picking one problem and then you know, kind of working with the customers and driving it forward.

Corey: On some level, I think that this is going to be looked down upon inherently just by automatic reflex of folks with infrastructure engineering backgrounds. It’s taken me some time to learn to overcome my own negative reaction to it. Because it’s, I’m here to build things and I want to build things out in such a way that it’s portable and reusable without having to be tied to a particular vendor and move on. And it took me a long time to realize that what that instinct was whispering in my ear was in fact, no, you should be your own cloud provider. If that’s really what I want to do, I probably should just brush up on you know, computer science trivia from 20 years ago and then go see if I can pass Google’s SRE interview.

I’m not here to build the things that just provision infrastructure from scratch every company I wind up landing at. It feels like there’s more important, impactful work that I can do. And let’s be clear, people are never going to follow guardrails themselves when they have to do a bunch of manual steps. It has to be something that is done for them. And I don’t know how you necessarily get there without having some form of blueprint or something like that, provided for them with something that is self-service because otherwise, it’s not going to work.

Ken: I a hundred percent agree, by the way, Corey. Like, the take that, like, automation is the only way to drive a lot of this forward is true, right? If for every single thing you’re trying—like, we have a concept called a rubric and it’s basically how you measure the service health. And you can—it’s very customizable, you have different dimensions. But if, for any check that’s on your rubric, it requires manual effort from all your developers, that is going to be harder than something you can just automate away.

So, vulnerability management is a great example. If you tell developers, “Hey, you have to go upgrade this library,” okay, some percentage [unintelligible 00:20:47], if you give developers, “Here’s a pull request that’s already been done and has a test passing and now you just need to merge it,” you’re going to have a much better adoption rate with that. Similarly with, like, applying templates being able to [up-level 00:20:57], you know, kind of apply the latest version of a template to an existing service, those types of capabilities, anything where you can automate what the fixes are, absolutely you’re going to get better adoption.

Corey: As you take a look at your existing reference customers—which is something I always look for on vendor websites because, like, oh, we have many customers who will absolutely not admit to being customers, it’s like, that sounds like something that’s easy to say—you have actual names tied to these things. Not just companies, but also individuals. If you were to sit down and ask your existing customer base, “So, why did you wind up implementing OpsLevel and what has the value that’s delivered to you been since that implementation?” What do they say?

Ken: Definitely. I actually had to check our website because we, you know, land new customers and put new logos on it. I was like, “Oh, I wonder what the current set is out right now?”

Corey: I have the exact same challenge. Like oh, we have some mutual customers. And it’s okay. I don’t know if I can mention them by name because I haven’t checked our own list of testimonials [unintelligible 00:21:51] lately because say the wrong thing and that’s how you wind up being sued and not having a company anymore.

Ken: Yeah. So, I don’t—I definitely, you know, want to stay [on side 00:22:00] on that part, but in terms of, like, kind of sample reference customer, a lot of the folks that we initially worked with are the platform teams, right? They’re the teams that care about what’s out there, and they need to know who’s responsible for it because they’re trying to drive some kind of cross-cutting change across the entire, you know, production footprint. And so, the first thing that generally people will say is—and I love this quote. This came—I won’t name them, but like, it’s in one of our case studies.

It was like, “I had, like, 50 different attempts at making a spreadsheet and they’re all, like, in the graveyard, like, to be able to capture what’s out there and who’s responsible for it.” And just OpsLevel helping automate that has been one of the biggest values that they’ve gotten. The second point, then is now be able to drive maturity and be able to measure how well those services are being built. And again, it’s sort of this interesting thing where we start with the platform teams. And then sometime later security teams find out about OpsLevel, and they’re like, “Oh, this is a tool I can use to, like, get developers to do stuff? Like, I’ve been trying to get developers to do stuff for the longest time.”

And they—I file Jira tickets and they just sit there and nothing gets done. But when it becomes part of this, like, overall health score that you’re trying to increase a part of the across the board, yeah, it’s just a way to kind of drive action.

Corey: I think that there’s a dichotomy of companies that emerge. And I tend to see the world through a lens of AWS bills, so let’s go down that path. I feel like there are some companies presumably like OpsLevel, whereas if I—assuming you’re running on top of AWS—if I were to pull your AWS bill, I would see upwards of 80% of your spend is going to be on this application called OpsLevel, the service that you provide to people. As opposed to the other side of the world, which is large enterprises, where they’re spending hundreds of millions of dollars a year, but the largest application they have is a million-and-a-half a year in spend because just, they have thousands of these things scattered everywhere. That latter case is where I tend to see more platform teams, where I start to see a lot of managing a whole bunch of relatively small workloads. And developer platforms really seem to be where a lot of solutions lead, whereas 80% of our workload is one application, we don’t feel the need for that as much. Is that accurate? Am I misunderstanding some aspect of it?

Ken: No, a hundred percent you’d hit the nail on the head. Like, okay, think about the typical, like, microservices adoption journey. Like, you started with, you know, some small company—like us—you started with a monolith. Ah, maybe you built out a second app—

Corey: Then you read on Hacker News and realize, “Oh, if we want to hire people, we’ve got to be doing what all the cool kids are up to.”

Ken: Right. We got a microservice all the thing—but that’s actually you know, microservices should come later, right, as a response to you needing scale your org and scale your—

Corey: As someone who started building some application with microservices, I could not agree more.

Ken: A hundred percent. So, it’s as you’re starting to take steps to having just more moving parts in your production infrastructure, right? If you have one moving part, unless it’s like a really large moving part that you can internally break down, like, kind of this majestic monolith where you do have kind of like individual domains that are owned by different teams, but really the problem we’re trying to solve, it’s more about, like, who owns what. Now, if that’s a single atomic unit, great, but can you decompose that? But if you just have, like, one small application, kind of like the whole team is owning everything, again, a developer portal is probably not the right tool for you. It really is a tool that you need as you start to scale your engineer work and as you start to scale the number of moving parts in your production infrastructure.

Corey: I tended to use to think of that in terms of boring companies versus innovative ones and I don’t think that’s accurate. I think it is the question of maturity and where companies lead to. On some level, of OpsLevel starts growing and becomes larger and larger in different ways and starts doing acquisitions and launching into other areas, at some point, you don’t have just one product offering, you have a multitude of them. At which point having something like that is going to be critical. But I have to ask, given that you are sort of not exactly your target customer profile, what are the sharp edges been on using it for your use case?

Ken: Yeah. So, we actually have an internal Slack channel, we call OpsLevel on OpsLevel. And finding those sharp edges actually has been really useful for us. You know, all the good stuff, dogfooding and it makes your own product better. Okay, so we have our main app, we also do have a bunch of smaller things and it’s like, oh yeah, you know, we have, like, I don’t know, various Hackaday things that go on, it’s important we kind of wind those down for, you know, compliance, we have our marketing site, we have, like, our Terraform.

Like, so there’s, like, stuff. It’s not, like, hundreds or thousands of things, but there’s more than just the main app. The second though, is it’s really on the maturity piece that we really try to get a lot of value out of our own product, right? Helping—we have our own platform team. They’re also trying to drive certain initiatives with our product developers.

There is that usual tension of our, like, our own product developers are like, “I want to ship features.” What’s this security thing I have to go take care of right now? But OpsLevel itself, like, helps reflect that. We had an operational review today and it was like, “Oh, this one service is actually now”—we have platinum as a level. It’s in gold instead of platinum. It’s like, “Why?” “Oh, there’s this thing that came up. We got to go fix that.” “Great. Let’s go actually go fix that so we’re back into platinum.”

Corey: Do you find that there’s often a choice you have to make internally, where you could make the product more effective for your specific use case, but that also diverges from where your typical customer needs or wants the product to go?

Ken: No, I think a lot of the things we find for our use case are, like, they’re more small paper cuts, right? They’re just as we’re using it, it’s like, “Hey, like, as I’m using this, I want to see the report for this particular check. Why do I have to click six times to get?” You know, like, “Wouldn’t it be great if we had a button?” Right?

And so, it’s those type of, like, small innovations that kind of come up. And those ultimately lead to, you know, a better product for our customers. We also work really closely with our customers and developers are not shy about telling you what they don’t like about your product. And I say this with love, like, a lot of our customers give us phenomenal feedback just on how our product can be better and we try to internalize that and you know, roll that feedback into the product.

Corey: You have a number of integrations of different SaaS providers, infrastructure providers, et cetera, that you wind up working with. I imagine that given your scale and scope and whatnot, those offerings are dictated by what customers say, “Hey, we’re using this thing. Are you going to support that or are you not going to maintain our business?” Which is a great way to wind up financing a lot of product development and figuring out what matters to people. My question for you is, if you look across the totality of your user base, what are the most popularly used integrations, if you can say?

Ken: Yeah, for sure. I think right now—I could actually dive in to pull the numbers—GitHub and GitLab—or… I think GitHub, like, has slightly more adoption across our customer base. At least with our customers, almost nobody uses Bitbucket. I mean, we have, like, a small number, but, like, it’s… I think, single-digit percentage. A lot of people use PagerDuty, which you know, hey, I’m an ex-PagerDuty person [crosstalk 00:28:24] and I’m glad to see that.

Corey: I have a free tier PagerDuty account that will automatically page me for my home automation stuff. Specifically, if you know, the fire alarm goes off. Like, yeah, okay, there are certain things I want to be woken up for, but it’s a very short list.

Ken: Yeah, it’s funny, the running default message when we use a test PagerDuty was, “The server is on fire.” [unintelligible 00:28:44] be like, “The house is on fire.” Like you know, go get that taken care of. There’s one other tool so that’s used a lot. Datadog actually is used a ton by just across our entire customer base, despite its… we’re also Data—we’re a Datadog partner, we’re a Datadog customer, you know? It’s not cheap, but it’s a good product for, you know, monitoring and logs and there are [crosstalk 00:29:01]—

Corey: No other than cloud infrastructure providers, I get the number one most common source of inquiries is Datadog optimization. It has now risen to a board-level concern in many cases because observability is expensive. That’s a sign of success, on some level. Meanwhile, I’m sitting here, like, Date-a-dog? Oh, my God, that’s disgusting. It’s like Tinder for Pets. Which it turns out is not at all what they do.

Ken: Nice.

Corey: Yeah.

[audio break 00:29:23]—optimizing their Slack integrations, their GitHub integration, et cetera. Or are they starting with the spinning up the servers piece of it?

Ken: A lot of the time—and again, that first problem they’re trying to solve is just get me a handle on everything we have running in production. You know, if you have multiple AWS accounts, multiple Kubernetes clusters, dozens or even hundreds of teams, God help you if you’re going to try to, like, build a list manually to consolidate all that information. That’s really the first part is, like, integrate Kubernetes, integrate your CI/CD pipelines, integrate Git, integrate your Cloud account, like, will integrate with everything and will try to build that map of, like, here’s everything that’s out there, and start to try to assign it to, like, and here’s people that we think might be responsible in terms of owning the software. That’s generally the starting point.

Corey: Which makes an awesome amount of sense. I think going at it from the infrastructure first perspective is where I’ve seen most developer platforms founder. And to be fair, the job is easier now than it was years ago because it used to be that you were being out-innovated by AWS constantly. Innovation has slow down there. And you know that because of how much they say the pace of innovation has only sped up.

And whenever AWS says something in a marketing context, they’re insecure about it. I’ve learned this through the fullness of time observing that company. And these days, most customers do not use the majority of features available for any given service. They have solidified to a point where you can responsibly build on top of these things. Now, it seems that the problem is all the ‘yes, and’ stuff that gets built on top of it.

Ken: Yeah. Do you have an example, actually, like, one of the kinds of, like, ‘yes, and’ tools that you’re thinking about?

Corey: Oh, absolutely. We have a bunch of AWS environment stuff so we should configure CloudWatch to look at all these things from an observability perspective. No, you should not. You should set up Datadog. And the first time someone does that by hand, they enable all have the observability and the rest and suddenly get charged approximately the GDP of Guam.

And okay, maybe we shouldn’t do that because then you have the downstream impact of that on your CloudWatch bill. So okay, how do we optimize this for the observability piece directly tied to that? How do we make sure that we get woken up when the site is down or preferably before that, but not every time basically, a EBS volume starts to get a little bit toasty? You have to start dialing this stuff in. And once you’ve found a lot of those aspects, being able to templatize that and roll that out on an ongoing basis and having the integrations all work together feels like it’s the right problem to be solving.

Ken: Yeah, absolutely. And the group that I think is responsible for that kind of—because it’s a set of problems you described—is really, like, platform teams. Sometimes service owners for like, how should we get paged, but really, what you’re describing are these kind of cross-cutting engineering concerns that platform teams are uniquely poised to help solve in an [unintelligible 00:32:03] organization, right? I was thinking what you said earlier. Like, nobody just wants to rebuild the same info over and over, but it’s sort of like, it’s not just building an [unintelligible 00:32:09]; it’s kind of like solving this, like, how do we ship? Can we actually run stuff in prod? And not just run it but get observability and ensure that we’re woken up for it and, like, what’s that total end-to-end look like from, like, developers writing code to running software in production that’s serving traffic? And solving all the problems [unintelligible 00:32:24], that’s what I think of was platform engineering.

Corey: So, my last question before we wind up wrapping this episode comes down to, I am very adept at two different programming languages, and those are brute force and enthusiasm. What implementation language is most of what you find yourself working with? And why is it in invariably going to be YAML?

Ken: Yeah, that’s a great question. So, I think there’s, in terms of implementing OpsLevel and implementing a service catalog, we support YAML. Like, you know, there’s this very common workflow, you just drop a YAML spec, basically, in your repo, if you’re a service owner. And that, we can support that. I don’t think that’s a great take, though.

Like, we have other integrations. Again, if the problem you’re trying to solve is I want to build a catalog of everything that’s out there, asking each of your developers hey, can you please all write YAML files that, like, describe the services you own and drop them into this repo? You’ve inverted this, like, database that essentially you’re trying to build, like, what’s out there and stored it in Git, potentially across several hundreds or thousands of repos. You put a lot of toil now on individual product developers to go write and maintain these files. And if you ever had to, like, make a blanket update to these files, there’s no atomic way to kind of do that, right?

So, I look at YAML as, like, I get it, you know? Like, we use the YAML for all the things in DevOps, so why not their service catalog as well, but I think it’s toil. Like, there are easier ways to build a catalog. By, kind of, just integrate. Like, hook up AWS, hook up GitHub, hook up Kubernetes, hook up your CI/CD pipeline, hook up all these different sources that have information about what’s running in prod, and let the software, let the tool, automatically infer what’s actually running as opposed to requiring humans to manually enter data.

Corey: I find that there are remarkably few technical holy wars that I cannot unify both sides on by nominating something far worse. Like, the VI versus Emacs stuff, the tabs versus spaces, and of course, the JSON versus YAML folks. My JSON versus YAML answer is XML: God’s language. I find that as soon as you suggest that, people care a hell of a lot less about the differences between JSON and YAML because their job is to now kill the apostate, which is me.

Ken: Right. Yeah. I remember XML, like, oh, man, 2002. SOAP. I remember SOAP as a protocol. That was a thing.

Corey: Some of the earliest S3 API calls were done in SOAP, and I think they finally just used it to wash their mouths out when all was said and done.

Ken: Nice. Yeah.

Corey: I really want to thank you for taking the time to do your level best to attempt to convert me, and I would argue in many respects, you have succeeded. I’m thinking about this differently than I did half an hour ago. If people want to learn more, where’s the best place for them to find you?

Ken: Absolutely. So, you can always check out our website, opslevel.com. We’re also fairly active on LinkedIn. If Twitter hasn’t imploded by the time this episode becomes launched, then they can also check us out at twitter.com/OpsLevelHQ. We’re always posting, just different content on, like, how to be successful with service maturity, DevOps, developer productivity, so that you know, ultimately, that you can ship out to customers faster.

Corey: And we will, of course, put links to that in the [show notes 00:35:23]. Thank you so much for taking the time, not just to speak with me, but also for sponsoring this episode. It is appreciated.

Ken: Cheers.

Corey: Ken Rose, CTO and co-founder at OpsLevel. I’m Cloud Economist Corey Quinn and this has been a promoted guest episode of Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an angry comment which, upon further reflection, you could have posted to all of the podcast platforms if only you had the right developer platform to pull it off.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

View Details

Nickolas Means, VP Engineering at Sym, joins Corey on Screaming in the Cloud to discuss how Sym is looking to solve the most common and most frustrating elements of compliance. Nick reveals why he finds it valuable to focus on making it easy for people to do the right thing over preventing them from doing the wrong thing, and why he feels the true spirit of compliance involves helping teams collaboratively come up with mutually beneficial solutions. Corey and Nick also dive into the common problems that engineers experience as a result of traditional compliance methods, and why historically the compliance industry has gotten a bad rap.

About Nickolas

Nickolas Means loves nothing more than a story of engineering triumph (except maybe a story of engineering disaster). When he’s not stuck in a Wikipedia loop reading about plane crashes, he leads the engineering team at Sym, helping create the building blocks engineering teams need to build delightful developer access and approval workflows.

Nick has been leading software engineering teams for more than a decade in the healthtech and devtools spaces. His focus is on building distributed organizations defined by their cultures of high trust and autonomy. He’s also an international keynote speaker, having shared his unique brand of storytelling with audiences around the world. He works remotely from Austin, TX, and spends his spare time going on adventures with his wife and kids, running very slowly, and trying to brew the perfect cup of coffee.

Links Referenced:

  • symops.com: https://symops.com
  • Twitter: https://twitter.com/nmeans

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Developers are responsible for more than ever these days. Not just the code they write, but also the containers and cloud infrastructure their apps run on. And probably the billing on top of that - which is neither here nor there. And a big part of that responsibility is app security — from code to cloud.

That’s where Snyk comes in. Snyk is a frictionless security platform that meets teams where they are, automating application security controls across their existing tools, workflows, and the AWS application stack — including seamless integrations with AWS CodePipeline, Amazon EKS, Amazon Inspector and several others.

Deploy on AWS. Secure with Snyk. Learn more at snyk.co/scream. That’s S-N-Y-K-dot-C-O/scream.

And my thanks to them for sponsoring this ridiculous nonsense!

Corey: LANs of the late 90’s and early 2000’s were a magical place to learn about computers, hang out with your friends, and do cool stuff like share files, run websites & game servers, and occasionally bring the whole thing down with some ill-conceived software or network configuration. That’s not how things are done anymore, but what if we could have a 90’s style LAN experience along with the best parts of the 21st century internet? (Most of which are very hard to find these days.) Tailscale thinks we can, and I’m inclined to agree. With Tailscale I can use trusted identity providers like Google, or Okta, or GitHub to authenticate users, and automatically generate & rotate keys to authenticate devices I've added to my network. I can also share access to those devices with friends and teammates, or tag devices to give my team broader access. And that’s the magic of it, your data is protected by the simple yet powerful social dynamics of small groups that you trust.Try now - it's free forever for personal use. I’ve been using it for almost two years personally, and am moderately annoyed that they haven’t attempted to charge me for what’s become an essential-to-my-workflow service.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. This promoted guest episode is brought to us by our friends over at Sym, and into my verbal grist mill, they have thrown their VP of Engineering, Nickolas Means. Nickolas, thank you for joining me.

Nickolas: Thank you so much for having me, Corey. And feel free to call me Nick.

Corey: I certainly shall. So, let’s begin at a high level. When you’re starting a company and trying to, sort of, bootstrap and raise initial rounds of funding and the rest, you’re trying to save money in a bunch of places. And one of the most expensive things you can buy when starting a company is, of course, a vowel. You wound up not naming the company—or the vowel, really—the y is sometimes a vowel, sometimes not. It’s S-Y-M. What is it you folks do exactly? What do you folks start? Where do you stop?

Nickolas: So, the name of the company comes from the idea of helping humans and machines work together more effectively. And that’s really nice and high level; it doesn’t tell you any information about what we do.

Corey: It feels like we’re—we’d assume that most startups pivot at some point; we’re just going to set—

Nickolas: [laugh].

Corey: —[crosstalk 00:01:33] seeds for that nice and early on, and dive on in.

Nickolas: So, what we actually do, the two co-founders and myself all have a background in highly compliant industries. I’ve done VPN stints at a couple of health tech startups; they’ve done similarly. And all three of us ended up building sort of a certain set of things every time we were at one of these companies. Because you have to be compliant with things, and in order to be compliant with things, you have to have a set of controls, you have to restrict certain things: how people get to production, how people access customer data. And those controls, by and large, all suck. They’re all painful and every company ends up building something from scratch at some point to make them not suck quite so bad. And it seemed like there was a product opportunity there.

Corey: I would argue there absolutely is. One of the big problems that I’ve found throughout the time that I’ve been fixing AWS bills on a consultancy basis has been, we’re really talking about cloud governance. But even now, by using the phrase cloud governance, three-quarters of the audience immediately wound up skipping to the next podcast over on their playlist because it sounds like it is one of those incredibly boring things. And to be fair, usually, when it comes to compliance, you want some of the most boring, least creative people in the world overseeing that. Like, when you wind up talking to someone at a company and they have a great sense of humor and they are constantly cracking jokes constantly, it’s like, “What do you do?” Like, “Oh, I’m the CFO.” All you hear from that is, “Oh, I’m about to go to prison. Awesome.”

Like, you want the wild, cutting-loose CEO to have three drinks and then confide, “I really like typing the number six.” You want them [laugh] to be predictable in a whole bunch of ways. And it always feels like compliance takes that entire mindset of, it’s always about risk management, it’s about wanting to make sure that people don’t go off script in a bunch of weird ways, but as an engineer, what I always heard from that is slow down, don’t be creative, go ahead and do things in very predictable ways. Only release things once a quarter, et cetera, et cetera. And yes, that’s one way to meet compliance goals, but it’s a crappy way, in my experience. I’m going to guess, though, that you have a lot more experience with the compliance world than I do because having worked a few times now, for big regulated finance companies, I wanted to get the hell out of the compliance universe.

Nickolas: Yeah, I mean, you used an interesting turn of phrase there. You used the phrase, “Avoid going off script,” and I think there’s a subtle turn there that actually makes all of this work a lot better. Instead of focusing on keeping people from going off script, you focus on keeping them on script. You focus on making it easier to do the right thing than to do the wrong thing. And that takes away a significant amount of the pain involved in compliance stuff.

You look at implementing controls—and everybody has the exact same reaction you just brought up about governance—because there’s so much FUD around this stuff. Everybody has been slowed down by one of these silly rules that makes no sense, that’s checking a box and not actually meeting the spirit of any kind of meaningful improvement.

Corey: Oh, cloud has absolutely doubled our speed of iteration because it used to take six weeks to get a server racked in the data center and we moved our processes to cloud and now to spin up an EC2 instance, it only takes three weeks of approvals. And at that point, it’s what are you really doing? You wind up with people building on shadow IT. It’s part of what contributed to the rise of cloud in the first place. Well, I can go through the annoying thing that this company wants me to do, or I have a corporate credit card and by the time it raises the level of spend to a point where it gets scrutiny, it’s in production serving customers and what are they going to do?

Some of the very early AWS sales conversations with customers started off as, “Well, why should we build on top of your cloud?” asks the exec, and they say, “Oh, sorry, you have 87 different accounts throughout your organization currently with us. We’re just trying to give you some unified view into it and possibly some discounting if you want.” Yeah, these days, that’s a fast track to getting yourself fired in some companies, if you wind up deviating from that story. But also, people are not doing this out of malfeasance; they’re trying to get their job done.

And as soon as guardrails start increasing friction, making it harder to do things the right way than to go around it, people will not comply. I strongly believe that, whether it’s cost—which is my universe, and frankly, only a business hours problem—or actual governance issues with some compliance regimes, which get those wrong and hope you enjoy some time in prison.

Nickolas: Yeah, exactly. I mean, you know, if you look at SOC 2, for example, there’s a lot of companies out there that are willing to sell you a program that will help you become SOC 2 compliant. They show you all the steps you need to take, all the programs you need to put in place. The thing they don’t do is help you establish the controls that are required. They’ll tell you that you have to have somebody formally approving before software goes out to production. They won’t give you any guidance whatsoever on how to put that control in place. And so, it’s really easy for a compliance person that’s not looking to collaborate with engineering just to go, “Okay, I need you to put a button in the deploy process and I need the CTO to click that button.”

Corey: Yes. We’ve always seen that as reactions to different things. I was at a company once where there were some outages caused by bad deploys, so they decided that a VP had to sign off on every deploy. Now, I come from the sysadmin ops world, which explains so much about my cynical perspective on life, so the way we got that overturned within two days is we did the malicious compliance thing, where oh, we need to deploy this. Great, we are walking into the middle of a senior leadership team meeting to get them to—with a tablet or comput—laptop—“I need you to click the button right now.”

And doing that out of hours and all kinds of other things, it’s oh. Yeah. How about we wind up only doing that for significant large changes? How about that? Maybe you don’t need to wake someone up at home in the middle of the night when there’s a deploy going out that fixes a typo on the marketing page; little things like that.

And at some point, you’re always felt like the goal of governance was either ossified scar tissue around all the ways that things have failed before, or through a, frankly, misguided belief that if we wind up distilling everything down to processes and procedures, eventually, someday, we can have a bunch of trained monkeys doing this job instead of people who are expensive and, you know, cynical, and difficult to please. I feel like that is not the right way to think about these things.

Nickolas: Well, I mean, the thing about those controls, you know, it’s exactly what you just said. Nowhere in SOC 2 does it say that your VPN [unintelligible 00:07:56] or CTO has to approve all code deploys, that’s not in there. But that’s the reality of life at a bunch of companies. In reality, if you just follow a software development life cycle that has multiple people looking at code before it gets deployed, multiple people signing off on that code being okay to deploy and you have a staging environment before you hit production, you’ve met the control. And SOC 2 gives you so much flexibility in how you write the control.

So, I think the thing that I’ve seen that makes compliance so much less painful, is when you have somebody that is 95% the boring persona like you’re talking about, but 5% creative. 5% willing to kind of get their hands dirty, empathize with the engineering team, collaborate with the engineering team, and find a way to put some of these controls in place that doesn’t just bring things to a grinding halt.

Corey: I have to assume that, given that you’ve built an entire product slash company around this idea, that you have some opinions other than doing what I do, which is sitting in my lofty ivory tower and oh, you should, in this idealized case, do things a little bit differently. But it’s going to be bespoke and the answer to any complex question, the more senior you get is, “It depends.” You, of course, have built something that scales out in a bunch of different ways. How do you view that in a way that makes it not either completely useless or overly prescriptive?

Nickolas: We focus on giving the power to engineering teams and giving the security complexity [unintelligible 00:09:23] the power to oversee those things. You know, it would be easy to give somebody, like, a clickbox UI, let them design controls for SOC 2 or whatever, end-user interface, but that’s not how engineers think; engineers think and express ideas and code. So, we’ve made the rather controversial decision in the face of a bunch of no-code tools to go low-code instead. So, to build a compliance workflow in Sym, you’re going to write some Terraform, you’re going to write a little bit of Python—a lot less than if you were building it from scratch—but you’re going to end up with something that perfectly fits the way that you already work versus having to shift your work practices around to fit the tool.

Corey: If you have inadvertently stumbled upon one of my hot buttons. There’s a lot of people that take a perspective around low code. And I just want to say that that perspective is often garbage. Like, oh, that’s not a real program—great. Hypothetically, if you have an idea for a business or a product or something, and involve software as most things seem to these days, maybe having to go to a boot camp for six months first as a prerequisite is not the best path forward.

“Well, you’re never going to build something hyperscale in a low-code environment.” Great, how many things that we built that actually need to be hyperscale that don’t go through 16 different architectural iterations between ridiculous idea one day and thing that is actually hyperscale? It’s an early optimization. I have an entire production pipeline in Retool that I built using low code. I think that that is a very powerful thing. And this idea that, “Oh, that’s not real code.” Cool. What’s your point?

Nickolas: Well, and for us, one of the things that we’re trying to enable is for software engineering teams, ops teams, whoever is building these controls, to interact with a security person or a compliance person, for them to be able to read the code, understand what it does, understand the way that the control has been implemented. And so, we provide a bunch of frameworks around that and a bunch of things. Like, you don’t have to go and build a Slack workflow from scratch and nobody has to understand that code because it’s buried in the platform. The only thing that the security or compliance person has to understand is the business logic that’s been put into place. Who can approve it? Who can’t approve it? How does that change after hours? How does that change if there’s an incident? All of that is in very simple Python that you don’t have to be an experienced programmer to be able to read.

Corey: One of the big powerful things behind that is it really reduces the interrupt volume of someone coming by to an engineer who is deep in the middle of something else, and, “Hey, guess what I have? A surprise context switch for something that’s going to take you probably 30 seconds, but then you’re going to be distracted by all of this.” If you give people the ability to self-serve, everything tends to work a lot more smoothly.

Nickolas: Yeah, absolutely. And, you know, that’s one of the ways we use Sym at Sym: we’ve got it in front of our AWS production environment, so if you need to go and do anything in production, you just have to get approval from any other engineer that happens to be in the approval channel, sort of a two-keys-to-launch-a-missile model. And that works fine for our compliance needs and it avoids there being a single point of failure that every time you need to go and get into production, you have to go and say, “Mother, may I?”

Corey: Exactly. It’s one of those things where every time you wind up with something that injects friction, people are going to find ways around it. And in some cases, this leads to positive outcomes where, when you’re subject to PCI, which is a lot more prescriptive than a number of other compliance regimes, it’s, great; this is a lot of things that don’t necessarily reflect how we work, how we want to work, et cetera. We can ignore it, which is not a great plan, we can wind up having to slow everything down, which is the common case, or the right answer is, we’re going to build the PCI environment that is very self-contained, just the critical stuff that needs to be in there is going to be in there, and then we can build everything that touches it around it in ways that are a lot more aligned with how we believe software should be built.

Nickolas: Yeah, absolutely. I mean, you silo off those high-control places, but there are controls that have to extend into the rest of the business. And one of the things that I’m a very firm believer in is, if you’re going to impose a control upon somebody, they need to have the agency to shape and to change that control so that it lets them work the way that they want to work.

Corey: I just want to call out how wonderful that is because I had a belief that looked borderline heretical, 12 years ago, when I said that, “Okay, simple rule. If you want me on call, I am empowered to change the thing that wakes me up.” Whether that is the code itself, the system itself, the paging threshold and frequency, or ultimately, I’m turning the physical pager off. It’s one of those things where I decide what’s an emergency outside of hours on that point. If it’s going to wake me up, I need the power to make sure it never does. Otherwise, you have no agency. It just feels like you’re being victimized by the stuff.

Nickolas: Yeah, absolutely. I mean, there was a wave of on-call regimes that ran through large companies for a while where there would be a centralized on-call team that would be responsible for responding to hundreds of services. And thankfully, we are maturing past that; we’re distributing on-call rotations so that teams that actually build services are responsible for them. And it’s the same mindset, right? If you’re going to be participating, if you’re going to be working with a system or working with a control, then you need to be able to change it, you need to be able to make it work the way that you think that it ought to work.

And in the context of compliance, you need to bring somebody along with you. You need to bring the person that’s responsible for the controls that actually has to sign on the dotted line at the end of the audit period, saying that we do all of these things. So, you have to be able to explain what you’re doing to them. But you have to be able to iterate.

Corey: I have to ask, given that what you are building is going to have heavy involvement from engineering, how do you respond to the probably most common engineering objection I imagine you get, which is, “Well, this doesn’t look hard. I could build this in a weekend.”

Nickolas: You know, it’s funny. We joke that our biggest competitor is build in-house, right? It’s pretty easy to start looking at what it takes to build a from scratch workflow in Slack to build a Slack app, to understand the cost of building it in-house. Because nothing about building an elegant user interface in Slack is easy or cheap. That API is difficult to work with and hard to get good user experience out of.

And we’ve spent a lot of time polishing a lot of places in the platform: we’ve got good documentation, we’ve got a good SDK, we’ve got good integration with third-party services that make all of this stuff easy to do. And it does look easy on the surface, it does look like ‘I can build it,’ but we’ve had customers that have had that objection gone and tried to build it and come back. Because it’s not as easy as anybody thinks.

Corey: My biggest competitor for fixing AWS bills has always been Microsoft Excel. It’s the, we’re going to do it ourselves—badly—internally. Okay, great. If that works for you, terrific—

Nickolas: Yeah.

Corey: —but very often it doesn’t. I mean, I think a classic case study of this is, in the terms of something that is well designed but is almost mind-bogglingly complex—and we’re getting a case study in it this year—is Twitter because it looks from the outside, very simple. I wind up writing a thing and I hit the post button and it shows up in a timeline. And then other people can subscribe to it or not, and they see it themselves. That sounds like something you can build on a weekend. And we look at all the ways it’s now exploding and collapsing and having weird bugs that no one anticipated, to realize, oh, this is a very challenging, very sophisticated application. But because it was well designed at one point, it looks easy.

Nickolas: Yeah. Yeah, it continues to run despite the fact that it’s having less than a quarter of the staff that originally maintained it, maintaining it because the services were well designed in the first place. They’re resilient on their own and they’re self-healing in a lot of cases. It’s the same thing with Sym. You can build these tools in-house, you can build them yourself, but then you’ve got more software to maintain. Because once you build something, you own it, forever. And the cheapest code is no code; the cheapest code is code that you don’t have to write.

It’s easy to look at a simple use case and understand a little bit of the cost of this. If you want a Slack workflow that gives you access to production in AWS, you can wire that up fairly quickly. Those APIs are not all that difficult. Now, let’s say you want to add an integration where if you’re on-call in PagerDuty, you can get to production without having to get an approval. Okay, well, now you’ve got a new API that you need to wire in.

And let’s say that every time that happens, you want to open a Jira ticket so that you can record that that’s happened. Well, there’s another API that you’ve got to wire in.j, whereas with Sym, it’s just, it’s right there. It’s a few lines of code to wire it all together. And it deploys in Terraform alongside the rest of your infrastructure, so you manage it the same way you’re used to managing things.

Corey: It reminds me of my earlier career when I was deep in the configuration management weeds with Puppet and SaltStack, where the biggest competitor we had any of those projects was always someone writing a bash script to do it themselves. And yes, you can do that, but then the requirements change, or you’re going to hit a point of scale that was surprising. And one of the valuable parts of it is that when the future is uncertain, as it always is—

Nickolas: Always.

Corey: Having folks who work in environments that aren’t just yours who encounter a lot of those edge cases you’re going to stumble into and can build things in is incredibly valuable. I don’t think I’ve ever met anyone who ran an infrastructure that said, “I would build it the same way if I had to start over again.” They always want to, “I would fix these annoying things.” Well, by having a product focused on a space like this, it’s yeah, today, you can have that VP click the approve button inside the GitHub Actions workflow. Good for you.

But when you get just a little bit further down the path, you aren’t going to want to do that anymore. There needs to be some decision-making it builds into it, and for certain high-risk changes, maybe a second person and so on. How do you build that logic engine? How do you build that workflow approach? How do you have a break glass thing for middle of the night when the site is down? Et cetera, et cetera, et cetera.

And that’s exactly the sort of thing that I would expect something like Sym to get very right, just because there’s always a bigger fish. You’ve seen this [unintelligible 00:19:17] before in other shops. And more to the point, if there’s something I want to do as a part of this that Sym doesn’t support and you are looking at me strangely if I asked how to do it, that’s usually a good early warning sign that maybe there’s something I’m not thinking about here. Because whatever the problem space is, I’m probably not the only person that has to do this. How are other companies solving for this? And it turns out that all my copy of our SOC 2 report has a typo on it. That would explain a lot. That’s a ‘can’ instead of ‘can’t.’ Nevermind. Or something like that.

Nickolas: Well, and the flip side of that is also true. I mean, the interesting thing about working on something that is sort of wide open with what you can wire up and build with it is we’re always learning from our customers. We’re always learning from the things that they’re doing. And so, you know, when somebody approaches us of, “Hey, we need to solve this particular problem,” if we don’t have a ready answer, we brainstorm and help figure that out. And to your point, that always extrapolates to other customers finding the same sort of thing useful.

The other bit of this that’s really interesting beyond the durability and the ability to kind of rapidly evolve these workflows is the audibility. It’s helpful in a lot of these compliance regimes to have a third-party tracking this data for you. So, when somebody accesses AWS production, who approved that access? When somebody deploys code, who approved those deploys? Well, we sit there as kind of a third party on the side, observing all of this, taking all these notes for you, and piping them into whatever audit tool that you want.

So, you’ve got that data long-term and when it comes time to audit, you’ve got all the evidence you need; it’s already there, already collected. You don’t have to go through and write a regex to parse a bunch of logs to get the information you need.

Corey: And invariably, that regex is always going to be different, depending upon the log stuff. It’s great having a unified central approach that is the trusted repository for this stuff. As you’ve been going to market and talking to your earlier customers and seeing the problems that you folks solve, what have you learned about the market space since you’ve gone into this direction? Because I feel like this is one of those products where you start designing and thinking you know a lot about the space, and you learn so much more just from the customer conversations and seeing that you can build the most finely crafted torque wrench in the world and the customer complains because it turns out, you built a crappy hammer.

Nickolas: So, I think what’s been really interesting to me is how much use our Lambda integration gets. We have a lot of first-party integrations with things like IAM and IAM identity center and Aptible and a bunch of tools that you can interact with, but a lot of our customers have wanted to do very specific things inside their infrastructure and put those things behind an approval. And the Lambda integration turns out to be a great Swiss army knife to do that because you can wire it up—it runs inside your firewall—to take essentially whatever action that you need it to. And that gets a ton of use. Probably more than half of our customers have at least one Lambda workflow in production, and I would not have expected that going in.

Corey: It’s wild to me just how pervasive Lambda has become. And even from a compliance perspective, it’s great because unlike, “Well, it’s a script that runs on a server somewhere,” yeah, it’s immutable. It’s versioned. There’s a way to conclusively prove that at invocation, this is the code that ran, the end, with the following parameters. Done.

There’s no, “Well, looking at the timestamp on the file”—like, no. None of that nonsense. It’s arguable that something that I have seen has been that Lambda is one of those rare technologies where you’re seeing faster adoption in the enterprise and you are in startup land.

Nickolas: Yeah, I would say that’s true. I mean, it’s so great for running undifferentiated workloads. I just need this one thing to happen really quickly and I don’t want to mess with standing up a server to run this thing that runs once a week. Okay, well, here’s a computer that will run just long enough for you to run this thing and then go away. It tracks exactly what ran, exactly when it ran, exactly how it got kicked off.

And in our case, it has access to all of the internal AWS APIs that we wall off in our platform because we obviously don’t want you using those things in the Sym runtime. But you can do anything that you want to your AWS environment from your own Lambda and we will gladly provide the approval step ahead of kicking that job off.

Corey: Are you seeing people use Lambda-based workflows to manage on-premises things or is it more heavily in environments that are already within the AWS boundary?

Nickolas: The Lambda stuff that we see is almost entirely—I think it is entirely for things that are within the AWS boundary. I can’t think of an instance when somebody is managing something on-prem with it.

Corey: I am increasingly discovering, through the magic of Tailscale—among a few other things—that I can use that for things on-premises that talk directly and interface with my Raspberry Pi in the spare room, et cetera. Which is—I think some people call it hybrid, which is the business enterprise term for ‘horrifying—

Nickolas: Yep.

Corey: —because it’s a terrible pattern in some ways. But it’s so convenient and it’s so nice not to have to worry about some of these things, just an infrastructure point of view. One thing that I think that AWS has done very well at, as they’ve evolved, has been with AWS Artifact, which ties directly to their own compliance reports, where in the early days when I was responsible for SOC 2 controls at a company, I found myself answering security questionnaires from vendors as if I was running in a data center. And sure enough, they wanted to tour us-east-1. And it turns out, you can’t really do that.

So now, just pointing them to the stuff that comes out of Artifact, it’s written by auditors for auditors and they go away and leave you alone without having to explain your bespoke artisanal nonsense to them. There’s something very pleasant about being able to throw the lion’s share of the work over to someone who already knows how to do it.

Nickolas: Our audit period is ending here shortly and I have recently been and spending time in Artifact. So yes, a hundred percent.

Corey: It used to be that you would only be able to get those things under explicit NDAs, you’d have to talk to your account manager for every one, it was a back-and-forth process, and you didn’t really know if what you were going to get was going to answer the questions that they had. Now it’s, you show up, you click things three times, and you’re done. The hardest part is sorting out which ones you need from the hundreds of things available within Artifact.

It’s like, okay, that’s great, but this one is in Spanish for some reason. And that’s awesome, but on some level, it feels like that should be an easy filter option. But yeah, no one ever accused AWS of building a good user interface. But once you get the thing you need and can pass it off, great. Job over. It’s one of my favorite services that most people who are what we know as ‘happy’ don’t know exist.

Nickolas: Yeah well, and that, it points to a larger industry trend, right, that companies are getting SOC 2 specifically earlier and earlier because it is becoming table stakes to be able to sell into other companies. They want to see your SOC 2 report before they’re willing to work with you before they’re willing to let your software touch their infrastructure. And there is a lot of value in these compliance programs as essentially a stamp of approval that you’re taking these things seriously, even with as much flexibility as SOC 2 has, just the stamp that we’ve thought about these things and we have serious answers to them is a pretty important signal to be able to send to somebody that’s wanting to buy your software.

Corey: We’ve toyed with the idea of going through the process ourselves because we get asked about it all the time, but it feels like the procurement processes that ask us for it expect us to come in with a whole software suite and the rest. And yeah, if that’s the world we’re operating in, it makes a lot of sense. We’re a services-based consultancy; we come in as individuals, we have conversations with people, and we talk about this and we have no write access to anything in your environment and give you scoped-down permissions for what we talk to because we don’t want the responsibility of that stuff.

And a lot of companies get that intrinsically, but there’s occasionally a few you have to go round and round and round with. It just it feels like it’s one of those, okay, you’re not quite there yet. You’re trying to view everything through this very specific worldview. Maybe it works for your constraints and requirements, but I’ve never understood it. And I’ve learned the older I get, the more time I spend around this, I used to have such a negative perspective on compliance.

And now it’s, you know, everything’s nuanced. There’s a reason that these things are there. It’s not just a make-work project for an industry that wants to slow everyone else down. It’s, there are risks here; these things exist for a reason. There’s a reason that you can go start Twitter for Pets tonight and not be regulated, but the same is not true of First Bank of Twitter Pets.

It’s okay, yeah, one of those things is going to require a fair bit of regulatory scrutiny, and as a society, we want that. Now, the counterargument that I don’t necessarily want to get too far into is, should Twitter for Pets be regulated?

Nickolas: [laugh].

Corey: And that’s a can of worms that I think we’ll leave for another episode.

Nickolas: Yeah, I mean, that’s—you know, the people that hate compliance the most are the people that are on the sharp end of compliance, people that are having to actually deal with the controls that are imposed upon them by these compliance regimes and by somebody who’s taking a very literal view in interpreting the things that some of these compliance programs say that you’d have to put in place. And I think, you know, that’s—kind of bring the conversation full circle—that’s the thing that we want to change more than anything. If we can wave a magic wand and change the compliance universe, the thing that I most want is to help compliance and security people collaborate with their engineering teams and come up with mutually beneficial solutions. Things that actually—the spirit of compliance.

Corey: Oh, yeah. My first PCI audit was a little bit of a challenge, just because the auditor wasn’t really conversant with anything that wasn’t a large company. So, they show up at our twelve-person start off, and, “Okay, where’s the Active Directory?” It’s like, “We don’t have one of those.” “Okay, well how do you authenticate to the WiFi?” It’s like, “Oh, the password’s on the wall.”

It’s, “Well, what happens if I get on that WiFi?” It’s, “What can I do that I couldn’t do from anywhere else?” Like, “Use that printer over there. That’s it.” Because everything else was the idea of the security boundary was built on identity, not on what blessed network you happened to be on; there was no special permissioning that didn’t apply to the Starbucks WiFi next town over.

But that was one of those things where at first they thought this was a horrifying problem and they were not going to be able to certify us, and it turned into no, we had significantly advanced culture of security compliance, oversight, separation of duties, all the things you really care about. We just didn’t have the trappings that usually came across with when you’re thinking about this or starting—or having the temerity to start a company, you know, longer than 18 months ago at a place that wasn’t San Francisco on the latest version of a MacBook Pro running the bleeding edge version of Chrome. It turns out that there’s a big universe out there. And not that there’s anything wrong with either side of it, until they start forgetting that not everyone operates the way that they do.

Nickolas: Yeah. I mean, you know, we talked about checkbox compliance a lot and I think that’s probably the biggest problem is there is a lot of checkbox compliance out there. And people have seen it not actually solve anything and just make everything harder. And so, compliance gets a bad rap.

Corey: Oh, for me, the one that I’ve been picking fights on social media about for a few years now is encryption-at-rest in the cloud. Like, yes, you want full-disk encryption turned on your laptops, your phones, your tablets, et cetera. Someone steals it from the coffee shop, you want to be out the cost the hardware. The end. But if you can get a hard drive intact out of an AWS facility and then reassemble it with the right number of drives in the right places, without… and hasn’t been encrypted. Congratulations, you earned it. As far as I’m concerned, that’s yours. You can keep it.

Because AWS employees aren’t able to do that, let alone third parties. But it is easier by far to click the box to enable encryption-at-rest and not spend half an hour arguing with the auditor… and just get on with your day. And recently in S3, for example, they wound up making that a default. Good for them. It’s just, can we please focus on the part of the story that’s relevant and germane to our business? Because that is not the threat model of modern attacks.

Nickolas: Yeah, I mean, for a long time, how much of the internet ran on unencrypted HTTP, but it was being served off of an encrypted disk? Great. What have we solved?

Corey: Oh, absolutely. It’s wild to me. Even now, I still we feel like there should be a reasonable way to handle—to [unintelligible 00:31:17] basically encryption between two points that doesn’t depend on the third-party CA’s with expiring certs and the rest. Drives me up a wall every time because it’s always the worst possible time. It causes the strangest issues and there is something deeply and profoundly wrong with the fact that the failure mode from the user perspective between, “Your connection is being intercepted by a third party,” and, “Holy shit. This certificate expired two hours ago.” Like, those are very different use cases, but the scary warnings have trained people to treat them the same way.

Nickolas: Yep. Yep, exactly the same. Ugh.

Corey: I really want to thank you for being so generous with your time. If people want to learn more, where’s the best place for them to find you?

Nickolas: Yeah, so the best place to find out more about Sym is our website, symops.com, SYMOPS dot com. And I should mention that Sym is completely free for teams of up to ten people. If any of you out there listening check it out, please reach out. We’d love to hear about your experiences, help any way we can. And if you want to get in touch with me directly, the best place to do that for now, while it lasts is still Twitter. I’m on there as @nmeans.

Corey: And we will, of course, include a link to that in the [show notes 00:32:27]. Thank you so much for agreeing to talk to me about all this stuff. I really appreciate it.

Nickolas: Yeah. Thanks so much for having me on, Corey. It’s been a lot of fun.

Corey: Nick Means, VP of Engineering at Sym. I’m Cloud Economist Corey Quinn, and this has been a promoted guest episode, brought to us by our friends at Sym. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an angry bitter comment that will get posted in six weeks, after you track down your elusive VP to click the approve button.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

View Details

Chris Farris, Cloud Security Nerd at PrimeHarbor Technologies, LLC, joins Corey on Screaming in the Cloud to discuss his new project, breaches.cloud, and why he feels having a centralized location for cloud security breach information is so important. Corey and Chris also discuss what it means to dive into entrepreneurship, including both the benefits of not having to work within a corporate structure and the challenges that come with running your own business. Chris also reveals what led him to start breaches.cloud, and what he’s learned about some of the biggest cloud security breaches so far.

About Chris

Chris Farris is a highly experienced IT professional with a career spanning over 25 years. During this time, he has focused on various areas, including Linux, networking, and security. For the past eight years, he has been deeply involved in public-cloud and public-cloud security in media and entertainment, leveraging his expertise to build and evolve multiple cloud security programs.

Chris is passionate about enabling the broader security team’s objectives of secure design, incident response, and vulnerability management. He has developed cloud security standards and baselines to provide risk-based guidance to development and operations teams. As a practitioner, he has architected and implemented numerous serverless and traditional cloud applications, focusing on deployment, security, operations, and financial modeling.

He is one of the organizers of the fwd:cloudsec conference and presented at various AWS conferences and BSides events. Chris shares his insights on security and technology on social media platforms like Twitter, Mastodon and his website https://www.chrisfarris.com.

Links Referenced:

  • fwd:cloudsec: https://fwdcloudsec.org/
  • breaches.cloud: https://breaches.cloud
  • Twitter: https://twitter.com/jcfarris
  • Company Site: https://www.primeharbor.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud, I’m Corey Quinn. My returning guest today is Chris Farris, now at PrimeHarbor, which is his own consultancy. Chris, welcome back. Last time we spoke, you were a Turbot, and now you’ve decided to go independent because you don’t like sleep anymore.

Chris: Yeah, I don’t like sleep.

Corey: [laugh]. It’s one of those things where when I went independent, at least in my case, everyone thought that it was, oh, I have this grand vision of what the world could be and how I could look at these things, and that’s going to just be great and awesome and everyone’s going to just be a better world for it. In my case, it was, no, just there was quite literally nothing else for me to do that didn’t feel like an exact reframing of what I’d already been doing for years. I’m a terrible employee and setting out on my own was important. It was the only way I found that I could wind up getting to a place of not worrying about getting fired all the time because that was my particular skill set. And I look back at it now, almost seven years in, and it’s one of those things where if I had known then what I know now, I never would have started.

Chris: Well, that was encouraging. Thank you [laugh].

Corey: Oh, of course. And in sincerity, it’s not one of those things where there’s any one thing that stops you, but it’s the, a lot of people get into the independent consulting dance because they want to do a thing and they’re very good at that thing and they love that thing. The problem is, when you’re independent, and at least starting out, I was spending over 70% of my time on things that were not billable, which included things like go and find new clients, go and talk to existing clients, the freaking accounting. One of the first hires I made was a fractional CFO, which changed my life. Up until that, my business partner and I were more or less dead reckoning of looking at the bank account and how much money is in there to determine if we could afford things. That’s a very unsophisticated way of navigating. It’s like driving by braille.

Chris: Yeah, I think I went into it mostly as a way to define my professional identity outside of my W-2 employer. I had built cloud security programs for two major media companies and felt like that was my identity: I was the cloud security person for these companies. And so, I was like, ehh, why don’t I just define myself as myself, rather than define myself as being part of a company that, in the media space, they are getting overwhelmed by change, and job security, job satisfaction, wasn’t really something that I could count on.

Corey: One of the weird things that I found—it’s counterintuitive—is that when you’re independent, you have gotten to a point where you have hit a point of sustainability, where you’re not doing the oh, I’m just going to go work for 40 billable hours a week for a client. It’s just like being an employee without a bunch of protections and extra steps. That doesn’t work super well. But now, at the point where I’m at where the largest client we have is a single-digit percentage of revenue, I can’t get fired anymore, without having a whole bunch of people suddenly turn on me because I’ve done something monstrous, in which case, I probably deserve not to have business anymore, or there’s something systemic in the macro environment, which given that I do the media side and I do the cost-cutting side, I work on the way up, I work on the way down, I’m questioning what that looks like in a scenario that doesn’t involve me hunting for food. But it’s counterintuitive to people who have been employees their whole life, like I was, where, oh, it’s risky and dangerous to go out on your own.

Chris: It’s risky and dangerous to be, you know, tied to a single, yeah, W-2 paycheck. So.

Corey: Yeah. The question I’d like to ask is, how many people need to be really pissed off before you have one of those conversations with HR that doesn’t involve giving you a cup of coffee? That’s the tell: when you don’t get coffee, it’s a bad conversation.

Chris: Actually, that you haven’t seen [unintelligible 00:04:25] coffee these days. You don’t want the cup of coffee, you know. That’s—

Corey: Even when they don’t give you the crappy percolator navy coffee, like, midnight hobo diner style, it’s still going to be a bad meeting because [unintelligible 00:04:37] pretend the coffee’s palatable.

Chris: Perhaps, yes. I like not having to deal with my own HR department. And I do agree that yeah, getting out of the W-2 space allows me to work on side projects that interests me or, you know, volunteer to do things like continuing the fwd:cloudsec, developing breaches.cloud, et cetera.

Corey: I’ll never forget, one of my last jobs I had a boss who walked past and saw me looking at Reddit and asked me if that was really the best use of my time. At first—it was in, I think, the sysadmin forum at the time, so yes, it was very much the best use of my time for the problem I was focusing on, but also, even if it wasn’t, I spent an inordinate amount of time on social media, just telling stories and building audiences, on some level. That’s the weird thing is that what counts as work versus what doesn’t count as work gets very squishy when you’re doing your own marketing.

Chris: True. And even when I was a W-2 employee, I spent a lot of time on Twitter because Twitter was an intel source for us. It was like, “Hey, who’s talking about the latest cloud security misconfigurations? Who’s talking about the latest data breach? What is Mandiant tweeting about?” It was, you know—I consider it part of my job to be on Twitter and watching things.

Corey: Oh, people ask me that. “So, you’re on Twitter an awful lot. Don’t you have a newsletter to write?” Like, yeah, where do you think that content comes from, buddy?

Chris: Exactly. Twitter and Mastodon. And Reddit now.

Corey: There’s a whole argument to be had about where to find various things. For me at least, because I’m only security adjacent, I was always trying to report the news that other people had, not make the news myself.

Chris: You don’t want to be the one making the news in security.

Corey: Speaking of, I’d like to talk a bit about what you just alluded to breaches.cloud. I don’t think I’ve seen that come across my desk yet, which tells me that it has not been making a big splash just yet.

Chris: I haven’t been really announcing it; it got published the other night and so basically, yeah, is this is sort of a inaugural marketing push for breaches.cloud. So, what we’re looking to do is document all the public cloud security breaches, what happened, why, and more importantly, what the companies did or didn’t do that led to the security incident or the security breach.

Corey: How are you slicing the difference between broad versus deep? And what I mean by that is, there are some companies where there are indictments and massive deep dives into everything that happens with timelines and blows-by-blows, and other times you wind up with the email that shows up one day of, “Security is very important to us. Now, listen to how we completely dropped the ball on it.” And it just makes the biggest description that they can get away with of what happened. Occasionally, you find out oh, it was an open S3 buckets, or they’ll allude to something that sounds like it. Does that count for inclusion? Does it not? How do you make those editorial decisions?

Chris: So, we haven’t yet built a page around just all of the recipients of the Bucket Negligence Award. We’re looking at the specific ones where there’s been something that’s happened that’s usually involving IAM credentials—oftentimes involving IAM credentials found in GitHub—and what led to that. So, in a lot of cases, if there’s a detailed company postmortem that they send their customers that said, “Hey, we goofed up, but complete transparency—” and then they hit all the bullet points of how they goofed up. Or in the case of certain others, like Uber, “Hey, we have court transcripts that we can go to,” or, “We have federal indictments,” or, “We have court transcripts, and federal indictments and FTC civil actions.” And so, we go through those trying to suss out what the company did or did not do that led to the breach. And really, the goal here is to be able to articulate as security practitioners, hey, don’t attach S3 full access to this role on EC2. That’s what got Capital One in trouble.

Corey: I have a lot of sympathy for the Capital One breach and I wish they would talk about it more than they do, for obvious reasons, just because it was not, someone showed up and made a very obvious dumb decision, like, “Oh, that was what that giant red screaming thing in the S3 console means.” It was a series of small misconfigurations that led to another one, to another one, to another one, and eventually gets to a point where a sophisticated attacker was able to chain them all together. And yes, it’s bad, yes, they’re a bank and the rest, but I look at that and it’s—that’s the sort of exploit that you look at and it’s okay, I see it. I absolutely see it. Someone was very clever, and a bunch of small things that didn’t rise to the obvious. But they got dragged and castigated as if they basically had a four-character password that they’d left on the back of the laptop on a Post-It note in an airport lounge when their CEO was traveling. Which is not the case.

Chris: Or all of the highlighting the fact that Paige Thompson was a former Amazon employee, making it seem like it was her insider abilities that lead to the incident, rather than she just knew that, hey, there’s a metadata service and it gives me creds if I ask it.

Corey: Right. That drove me nuts. There was no maleficence as an employee. And to be very direct, from what I understand of internal AWS controls, had there been, it would have been audited, flagged, caught, interdicted. I have talked to enough Amazonians that either a lot of them are lying to me very consistently despite not knowing each other, or they’re being honest when they say that you can’t get access to customer data using secret inside hacks.

Chris: Yeah. I have reasonably good faith in AWS and their ability to not touch customer data in most scenarios. And I’ve had cases that I’m not allowed to talk about where Amazon has gone and accessed customer data, and the amount of rigmarole and questions and drilling that I got as a customer to have them do that was pretty intense and somewhat, actually, annoying.

Corey: Oh, absolutely. And, on some level, it gets frustrating when it’s a, look, this is a test account. I have nothing of sensitive value in here. I want the thing that isn’t working to start working. Can I just give you a whole, like, admin-powered user account and we can move on past all of this? And their answer is always absolutely not.

Chris: Yes. Or, “Hey, can you put this in our bucket?” “No, we can’t even write to a public bucket or a bucket that, you know, they can share too.” So.

Corey: An Amazonian had to mail me a hard drive because they could not send anything out of S3 to me.

Chris: There you go.

Corey: So, then I wound up uploading it back to S3 with, you know, a Snowball Edge because there’s no overkill like massive overkill.

Chris: No, the [snowmobile 00:11:29] would have been the massive overkill. But depending on where you live, you know, you might not have been able to get a permit to park the snowmobile there.

Corey: They apparently require a loading dock. Same as with the outposts. I can’t fake having one of those on my front porch yet.

Chris: Ah. Well, there you go. I mean, you know it’s the right height though, and you don’t mind them ruining your lawn.

Corey: So, help me understand. It makes sense to me at least, on some level, why having a central repository of all the various cloud security breaches in one place that’s easy to reference is valuable. But what caused you to decide, you know, rather than saying it’d be nice to have, I’m going to go build that thing?

Chris: Yeah, so it was actually right before the last time we spoke, Nicholas Sharp was indicted. And there was like, hey, this person was indicted for, you know, this cloud security case. And I’m like, that name rings a bell, but I don’t remember who this person was. And so, I kind of realized that there’s so many of these things happening now that I forget who is who. And so, when a new piece of news comes along, I’m like, where did this come from and how does this fit into what my knowledge of cloud security is and cloud security cases?

So, I kind of realized that these are all running together in my mind. The Department of Justice only referenced ‘Company One,’ so it wasn’t clear to me if this even was a new cloud incident or one I already knew about. And so basically, I decided, okay, let’s build this. Breaches.cloud was available; I think I kind of got the idea from hackingthe.cloud.

And I had been working with some college students through the Collegiate Cyber Defense Competition, and I was like, “Hey, anybody want a spring research project that I will pay you for?” And so yeah, PrimeHarbor funded two college students to do quite a bit of the background research for me, I mentored them through, “Hey, so here’s what this means,” and, “Hey, have we noticed that all of these seem to relate to credentials found in GitHub? You know, maybe there’s a pattern here.” So, if you’re not yet scanning for secrets in GitHub, I recommend you start scanning for secrets in your GitHub, private and public repos.

Corey: Also, it makes sense to look at the history. Because, oh, I committed a secret. I’m going to go ahead and revert that commit and push that. That solves the problem, right?

Chris: No, no, it doesn’t. Yes, apparently, you can force push and delete an entire commit, but you really want to use a tool that’s going to go back through the commit history and dig through it because as we saw in the Uber incident, when—the second Uber incident, the one that led to the CSOs conviction—yeah, the two attackers, [unintelligible 00:14:09] stuffed a Uber employee’s personal GitHub account that they were also using for Uber work, and yeah, then they dug through all the source code and dug through the commit histories until they found a set of keys, and that’s what they used for the second Uber breach.

Corey: Awful when that hits. It’s one of those things where it’s just… [sigh], one thing leads to another leads to another. And on some level, I’m kind of amazed by the forensics that happen around all of these things. With the counterpoint, it is so… freakishly difficult, I think, for lack of a better term, just to be able to say what happened with any degree of certainty, so I can’t help but wonder in those dark nights when the creeping dread starts sinking in, how many things like this happen that we just never hear about because they don’t know?

Chris: Because they don’t turn on CloudTrail. Probably a number of them. Once the data gets out and shows up on the dark web, then people start knocking on doors. You know, Troy Hunt’s got a large collection of data breach stuff, and you know, when there’s a data breach, people will send him, “Hey, I found these passwords on the dark web,” and he loads them into Have I Been Pwned, and you know, [laugh] then the CSO finds out. So yeah, there’s probably a lot of this that happens in the quiet of night, but once it hits the dark web, I think that data starts becoming available and the victimized company finds out.

Corey: I am profoundly cynical, in case that was unclear. So, I’m wondering, on some level, what is the likelihood or commonality, I suppose, of people who are fundamentally just viewing security breach response from a perspective of step one, make sure my resume is always up to date. Because we talk about these business continuity plans and these DR approaches, but very often it feels like step one, secure your own mask before assisting others, as they always say on the flight. Where does personal preservation come in? And how does that compare with company preservation?

Chris: I think down at the [IaC 00:16:17] level, I don’t know of anybody who has not gotten a job because they had Equifax on their resume back in, what, 2017, 2018, right? Yes, the CSO, the CEO, the CIO probably all lost their jobs. And you know, now they’re scraping by book deals and speaking engagements.

Corey: And these things are always, to be clear, nuanced. It’s rare that this is always one person’s fault. If you’re a one-person company, okay, yeah, it’s kind of your fault, let’s be clear here, but there are controls and cost controls and audit trails—presumably—for all of these things, so it feels like that’s a relatively easy thing to talk around, that it was a process failure, not that one person sucked. “Well, didn’t you design and implement the process?” “Yes. But it turned out there were some holes in it and my team reported that those weren’t there and it turned out that they were and, well, live and learn.” It feels like that’s something that could be talked around.

Chris: It’s an investment failure. And again, you know, if we go back to Harry Truman, “The buck stops here,” you know, it’s the CEO who decides that, hey, we’re going to buy a corporate jet rather than buy a [SIIM 00:17:22]. And those are the choices that happen at the top level that define, do you have a capable security team, and more importantly, do you have a capable security culture such that your security team isn’t the only ones who are actually thinking about security?

Corey: That’s, I guess, a fair question. I saw a take on Twitter—which is always a weird thing—or maybe was Blue-ski or somewhere else recently, that if you don’t have a C-level executive responsible for security with security in their title, your company does not take security seriously. And I can see that past a certain point of scale, but as a one-person company, do you have a designated CSO?

Chris: As a one-person company and as a security company, I sort of do have a designated CSO. I also have, you know, the person who’s like, oh, I’m going to not put MFA on the root of this one thing because, while it’s an experiment and it’s a sandbox and whatever else, but I also know that that’s not where I’m going to be putting any customer data, so I can measure and evaluate the risk from both a security perspective and a business existential investment perspective. When you get to the larger the organization, the more detached the CEO gets from the risk and what the company is building and what the company is doing, is where you get into trouble. And lots of companies have C-level somebody who’s responsible for security. It’s called the CSO, but oftentimes, they report four levels down, or even more, from the chief executive who is actually the one making the investment decisions.

Corey: On some level, the oh yeah, that’s my responsibility, too, but it feels like it’s a trap that falls into. Like, well, the CTO is responsible for security at a publicly traded company. Like, well… that tends to not work anymore, past certain points of scale. Like when I started out independently, yes, I was the CSO. I was also the accountant. I was also the head of marketing. I was also the janitor. There’s a bunch of different roles; we all wear different hats at different times.

I’m also not a big fan of shaming that oh, yeah. This is a universal truth that applies to every company in existence. That’s also where I think Twitter started to go wrong where you would get called out whenever making an observation or witticism or whatnot because there was some vertex case to which it did not necessarily apply and then people would ‘well, actually,’ you to death.

Chris: Yeah. Well, and I think there’s a lot of us in the security community who are in the security one-percenters. We’re, “Hey, yes, I’m a cloud security person on a 15-person cloud security team, and here’s this awesome thing we’re doing.” And then you’ve got most of the other companies in this country that are probably below the security poverty line. They may or may not have a dedicated security person, they certainly don’t have a SIIM, they certainly don’t have anybody who’s monitoring their endpoints for malware attacks or anything else, and those are the companies that are getting hit all the time with, you know, a lot of this ransomware stuff. Healthcare is particularly vulnerable to that.

Corey: When you take a look across the industry, what is it that you’re doing now at PrimeHarbor that you feel has been an unmet need in the space? And let me be clear, as of this recording earlier today, we signed a contract with you for a project. There’s more to come on that in the future. So, this is me asking you to tell a story, not challenging, like, what do you actually do? This is not a refund request, let’s be very clear here. But what’s the unmet need that you saw?

Chris: I think the unmet need that I see is we don’t talk to our builder community. And when I say builder, I mean, developers, DevOps, sysadmins, whatever. AWS likes the term builder and I think it works. We don’t talk to our builder community about risk in a way that makes sense to them. So, we can say, “Hey, well, you know, we have this security policy and section 24601 says that all data’s classifications must be signed off by the data custodian,” and a developer is going to look at you with their head tilted, and be like, “Huh? What? I just need to get the sprint done.”

Whereas if we can articulate the risk—and one of the reasons I wanted to do breaches.cloud was to have that corpus of articulated risk around specific things—I can articulate the risk and say, “Hey, look, you know how easy it is for somebody to go in and enumerate an S3 bucket? And then once they’ve enumerated and guessed that S3 bucket exists, they list it, and oh, hey, look, now that they’ve listed it, they know all of the objects and all of the juicy PII that you just made public.” If you demonstrate that to them, then they’re going to be like, “Oh, I’m going to add the extra story point to this story to go figure out how to do CloudFront origin access identity.” And now you’ve solved, you know, one more security thing. And you’ve done in a way that not just giving a man a fish or closing the bucket for them, but now they know, hey, I should always use origin access identity. This is why I need to do this particular thing.

Corey: One of the challenges that I’ve seen in a variety of different sites that have tried to start cataloging different breaches and other collections of things happening in public is the discoverability or the library management problem. The most obvious example of this is, of course, the AWS console itself, where when it paginates things like, oh, there are 3000 things here, ten at a time, through various pages for it. Like, the marketplace is just a joke of discoverability. How do you wind up separating the stuff that is interesting and notable, rather than, well, this has about three sentences to it because that’s all the company would say?

Chris: So, I think even the ones where there’s three sentences, we may actually go ahead and add it to the repo, or we may just hold it as a draft, so that we know later on when, “Hey, look, here’s a federal indictment for Company Three. Oh, hey, look. Company Three was actually this breach announcement that we heard about three months ago,” or even three years ago. So like, you know, Chegg is a great example of, you know, one of those where, hey, you know, there was an incident, and they disclosed something, and then, years later, FTC comes along and starts banging them over the head. And in the FTC documentation, or in the FTC civil complaint, we got all sorts of useful data.

Like, not only were they using root API keys, every contractor and employee there was sharing the root API keys, so when they had a contractor who left, it was too hard to change the keys and share it with everybody, so they just didn’t do that. The contractor still had the keys, and that was one of the findings from the FTC against Chegg. Similar to that, Cisco didn’t turn off contractors’ access, and I think—this is pure speculation—I think the poor contractor one day logged into his Google Cloud Shell, cd’ed into a Terraform directory, ran ‘terraform destroy’, and rather than destroying what he thought he was destroying, it had the access keys back to Cisco WebEx and took down 400 EC2 instances that made up all of WebEx. These are the kinds of things that I think it’s worth capturing because the stories are going to come out over time.

Corey: What have you seen in your, I guess, so far, a limited history of curating this that—I guess, first what is it you’ve learned that you’ve started seeing as far as patterns go, as far as what warrants inclusion, what doesn’t, and of course, once you started launching and going a bit more public with it, I’m curious to hear what the response from companies is going to be.

Chris: So, I want to be very careful and clear that if I’m going to name somebody, that we’re sourcing something from the criminal justice system, that we’re not going to say, “Hey, everybody knows that it was Paige Thompson who was behind it.” No, no, here’s the indictment that said it was Paige Thompson that was, you know, indicted for this Capital One sort of thing. All the data that I’m using, it all comes from public sources, it’s all sited, so it’s not like, hey, some insider said, “Hey, this is what actually happened.” You know? I very much learned from the Ubiquiti case that I don’t want to be in the position of Brian Krebs, where it’s the attacker themselves who’s updating the site and telling us everything that went wrong, when in fact, it’s not because they’re in fact the perpetrator.

Corey: Yeah, there’s a lot of lessons to be learned. And fortunately, for what it’s s—at least it seems… mostly, that we’ve moved past the battle days of security researchers getting sued on a whim from large companies for saying embarrassing things about them. Of course, watch me be tempting fate and by the time this publishes, I’ll get sued by some company, probably Azure or whatnot, telling me that, “Okay, we’ve had enough of you saying bad things about our security.” It’s like, well, cool, but I also read the complaint before you file because your security is bad. Buh-dum-tss. I’m kidding. I’m kidding. Please don’t sue me.

Chris: So, you know, whether it’s slander or libel, depending on whether you’re reading this or hearing it, you know, truth is an actual defense, so I think Microsoft doesn’t have a case against you. I think for what we’re doing in breaches, you know—and one of the reasons that I’m going to be very clear on anybody who contributes—and just for the record, anybody is welcome to contribute. The GitHub repo that runs breaches.cloud is public and anybody can submit me a pull request and I will take their write-ups of incidents. But whatever it is, it has to be sourced.

One of the things that I’m looking to do shortly, is start soliciting sponsorships for breaches so that we can afford to go pull down the PACER documents. Because apparently in this country, while we have a right to a speedy trial, we don’t have a right to actually get the court transcripts for less than ten cents a page. And so, part of what we need to do next is download those—and once we’ve purchased them, we can make them public—download those, make them public, and let everybody see exactly what the transcript was from the Capital One incident, or the Joey Sullivan trial.

Corey: You’re absolutely right. It drives me nuts that I have to wind up budgeting money for PACER to pull up court records. And at ten cents a page, it hasn’t changed in decades, where it’s oh, this is the cost of providing that data. It’s, I’m not asking someone to walk to the back room and fax it to me. I want to be very clear here. It just feels like it’s one of those areas where the technology and government is not caught up and it’s—part of the problem is, of course, having no competition.

Chris: There is that. And I think I read somewhere that the ent—if you wanted to download the entire PACER, it would be, like, $100 million. Not that you would do that, but you know, it is the moneymaker for the judicial system, and you know, they do need to keep the lights on. Although I guess that’s what my taxes are for. But again, yes, they’re a monopoly; they can do that.

Corey: Wildly frustrating, isn’t it?

Chris: Yeah [sigh]… yeah, yeah, yeah. Yeah, I think there’s a lot of value in the court transcripts. I’ve held off on publishing the Capital One case because one, well, already there’s been a lot of ink spilled on it, and two, I think all the good detail is going to be in the trial transcripts from Paige Thompson’s trial.

Corey: So, I am curious what your take is on… well, let’s called the ‘FTX thing.’ I don’t even know how to describe it at this point. Is it a breach? Is it just maleficence? Is it 15,000 other things? But I noticed that it’s something that breaches.cloud does talk about a bit.

Chris: Yeah. So, that one was a fascinating one that came out because as I was starting this project, I heard you know, somebody who was tweeting was like, “Hey, they were storing all of the crypto private keys in AWS Secrets Manager.” And I was like, “Errr?” And so, I went back and I read John J. Ray III’s interim report to the creditors.

Now, John Ray is the man who was behind the cleaning up of Enron, and his comment was “FTX is the”—“Never in my career have I seen such a complete failure of corporate controls and such a complete absence of trustworthy information as occurred here.” And as part of his general, broad write-up, they went into, in-depth, a lot of the FTX AWS practices. Like, we talk about, hey, you know, your company should be multi-account. FTX was worse. They had three or four different companies all operating in the same AWS account.

They had their main company, FTX US, Alameda, all of them had crypto keys in Secrets Manager and there was no access control between any of those. And what ended up happening on the day that SBF left and Ray came in as CEO, the $400 million worth of crypto somehow disappeared out of FTX’s wallets.

Corey: I want to call this out because otherwise, I will get letters from the AWS PR spin doctors. Because on the surface of it, I don’t know that there’s necessarily a lot wrong with using Secrets Manager as the backing store for private keys. I do that with other things myself. The question is, what other controls are there? You can’t just slap it into Secrets Manager and, “Well, my job is done. Let’s go to lunch early today.”

There are challenges [laugh] around the access levels, there are—around who has access, who can audit these things, and what happens. Because most of the secrets I have in Secrets Manager are not the sort of thing that is, it is now a viable strategy to take that thing and abscond to a country with a non-extradition treaty for the rest of my life, but with private keys and crypto, there kind of is.

Chris: That’s it. It’s like, you know, hey, okay, the RDS database password is one thing, but $400 million in crypto is potentially another thing. Putting it in and Secrets Manager might have been the right answer, too. You get KMS customer-managed keys, you get full auditability with CloudTrail, everything else, but we didn’t hear any of that coming out of Ray’s report to the creditors. So again, the question is, did they even have CloudTrail turned on? He did explicitly say that FTX had not enabled GuardDuty.

Corey: On some level, even if GuardDuty doesn’t do anything for you, which in my case, it doesn’t, but I want to be clear, you should still enable it anyway because you’re going to get dragged when there’s inevitable breach because there’s always a breach somewhere, and then you get yelled at for not having turned on something that was called GuardDuty. You already sound negligent, just with that sentence alone. Same with Security Hub. Good name on AWS’s part if you’re trying to drive service adoption. Just by calling it the thing that responsible people would use, you will see adoption, even if people never configure or understand it.

Chris: Yeah, and then of course, hey, you had Security Hub turned on, but you ignore the 80,000 findings in it. Why did you ignore those 80,000 findings? I find Security Hub to probably be a little bit too much noise. And it’s not Security Hub, it’s ‘Compliance Hub.’ Everything—and I’m going to have a blog post coming out shortly—on this, everything that Security Hub looks at, it looks at it from a compliance perspective.

If you look at all of its scoring, it’s not how many things are wrong; it’s how many rules you are a hundred percent compliant to. It is not useful for anybody below that AWS security poverty line to really master or to really operationalize.

Corey: I really want to thank you for taking the time to catch up with me once again. Although now that I’m the client, I expect I can do this on demand, which is just going to be delightful. If people want to learn more, where can they find you?

Chris: So, they can find breaches.cloud at, well https://breaches.cloud. If you’re looking for me, I am either on Twitter, still, at @jcfarris, or you can find me and my consulting company, which is www.primeharbor.com.

Corey: And we will, of course, put links to all of that in the [show notes 00:33:57]. Thank you so much for taking the time to speak with me. As always, I appreciate it.

Chris: Oh, thank you for having me again.

Corey: Chris Farris, cloud security nerd at PrimeHarbor. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an angry, insulting comment that you’re also going to use as the storage back-end for your private keys.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

View Details

About Rick

Rick is the Product Leader of the AWS Optimization team. He previously led the cloud optimization product organization at Turbonomic, and previously was the Microsoft Azure Resource Optimization program owner.

Links Referenced:

  • AWS: https://console.aws.amazon.com
  • LinkedIn: https://www.linkedin.com/in/rick-ochs-06469833/
  • Twitter: https://twitter.com/rickyo1138

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Chronosphere. Tired of observability costs going up every year without getting additional value? Or being locked in to a vendor due to proprietary data collection, querying and visualization? Modern day, containerized environments require a new kind of observability technology that accounts for the massive increase in scale and attendant cost of data. With Chronosphere, choose where and how your data is routed and stored, query it easily, and get better context and control. 100% open source compatibility means that no matter what your setup is, they can help. Learn how Chronosphere provides complete and real-time insight into ECS, EKS, and your microservices, whereever they may be at snark.cloud/chronosphere That’s snark.cloud/chronosphere

Corey: This episode is bought to you in part by our friends at Veeam. Do you care about backups? Of course you don’t. Nobody cares about backups. Stop lying to yourselves! You care about restores, usually right after you didn’t care enough about backups. If you’re tired of the vulnerabilities, costs and slow recoveries when using snapshots to restore your data, assuming you even have them at all living in AWS-land, there is an alternative for you. Check out Veeam, thats V-E-E-A-M for secure, zero-fuss AWS backup that won't leave you high and dry when it’s time to restore. Stop taking chances with your data. Talk to Veeam. My thanks to them for sponsoring this ridiculous podcast.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. For those of you who’ve been listening to this show for a while, the theme has probably emerged, and that is that one of the key values of this show is to give the guest a chance to tell their story. It doesn’t beat the guests up about how they approach things, it doesn’t call them out for being completely wrong on things because honestly, I’m pretty good at choosing guests, and I don’t bring people on that are, you know, walking trash fires. And that is certainly not a concern for this episode.

But this might devolve into a screaming loud argument, despite my best effort. Today, I’m joined by Rick Ochs, Principal Product Manager at AWS. Rick, thank you for coming back on the show. The last time we spoke, you were not here you were at, I believe it was Turbonomic.

Rick: Yeah, that’s right. Thanks for having me on the show, Corey. I’m really excited to talk to you about optimization and my current role and what we’re doing.

Corey: Well, let’s start at the beginning. Principal product manager. It sounds like one of those corporate titles that can mean a different thing in every company or every team that you’re talking to. What is your area of responsibility? Where do you start and where do you stop?

Rick: Awesome. So, I am the product manager lead for all of AWS Optimizations Team. So, I lead the product team. That includes several other product managers that focus in on Compute Optimizer, Cost Explorer, right-sizing recommendations, as well as Reservation and Savings Plan purchase recommendations.

Corey: In other words, you are the person who effectively oversees all of the AWS cost optimization tooling and approaches to same?

Rick: Yeah.

Corey: Give or take. I mean, you could argue that oh, every team winds up focusing on helping customers save money. I could fight that argument just as effectively. But you effectively start and stop with respect to helping customers save money or understand where the money is going on their AWS bill.

Rick: I think that’s a fair statement. And I also agree with your comment that I think a lot of service teams do think through those use cases and provide capabilities, you know? There’s, like, S3 storage lines. You know, there’s all sorts of other products that do offer optimization capabilities as well, but as far as, like, the unified purpose of my team, it is, unilaterally focused on how do we help customers safely reduce their spend and not hurt their business at the same time.

Corey: Safely being the key word. For those who are unaware of my day job, I am a partial owner of The Duckbill Group, a consultancy where we fix exactly one problem: the horrifying AWS bill. This is all that I’ve been doing for the last six years, so I have some opinions on AWS bill reduction as well. So, this is going to be a fun episode for the two of us to wind up, mmm, more or less smacking each other around, but politely because we are both professionals. So, let’s start with a very high level. How does AWS think about AWS bills from a customer perspective? You talk about optimizing it, but what does that mean to you?

Rick: Yeah. So, I mean, there’s a lot of ways to think about it, especially depending on who I’m talking to, you know, where they sit in our organization. I would say I think about optimization in four major themes. The first is how do you scale correctly, whether that’s right-sizing or architecting things to scale in and out? The second thing I would say is, how do you do pricing and discounting, whether that’s Reservation management, Savings Plan Management, coverage, how do you handle the expenditures of prepayments and things like that?

Then I would say suspension. What that means is turn the lights off when you leave the room. We have a lot of customers that do this and I think there’s a lot of opportunity for more. Turning EC2 instances off when they’re not needed if they’re non-production workloads or other, sort of, stateful services that charge by the hour, I think there’s a lot of opportunity there.

And then the last of the four methods is clean up. And I think it’s maybe one of the lowest-hanging fruit, but essentially, are you done using this thing? Delete it. And there’s a whole opportunity of cleaning up, you know, IP addresses unattached EBS volumes, sort of, these resources that hang around in AWS accounts that sort of getting lost and forgotten as well. So, those are the four kind of major thematic strategies for how to optimize a cloud environment that we think about and spend a lot of time working on.

Corey: I feel like there’s—or at least the way that I approach these things—that there are a number of different levels you can look at AWS billing constructs on. The way that I tend to structure most of my engagements when I’m working with clients is we come in and, step one: cool. Why do you care about the AWS bill? It’s a weird question to ask because most of the engineering folks look at me like I’ve just grown a second head. Like, “So, why do you care about your AWS bill?” Like, “What? Why do you? You run a company doing this?”

It’s no, no, no, it’s not that I’m being rhetorical and I don’t—I’m trying to be clever somehow and pretend that I don’t understand all the nuances around this, but why does your business care about lowering the AWS bill? Because very often, the answer is they kind of don’t. What they care about from a business perspective is being able to accurately attribute costs for the service or good that they provide, being able to predict what that spend is going to be, and also yes, a sense of being good stewards of the money that has been entrusted to them by via investors, public markets, or the budget allocation process of their companies and make sure that they’re not doing foolish things with it. And that makes an awful lot of sense. It is rare at the corporate level that the stated number one concern is make the bills lower.

Because at that point, well, easy enough. Let’s just turn off everything you’re running in production. You’ll save a lot of money in your AWS bill. You won’t be in business anymore, but you’ll be saving a lot of money on the AWS bill. The answer is always deceptively nuanced and complicated.

At least, that’s how I see it. Let’s also be clear that I talk with a relatively narrow subset of the AWS customer totality. The things that I do are very much intentionally things that do not scale. Definitionally, everything that you do has to scale. How do you wind up approaching this in ways that will work for customers spending billions versus independent learners who are paying for this out of their own personal pocket?

Rick: It’s not easy [laugh], let me just preface that. The team we have is incredible and we spent so much time thinking about scale and the different personas that engage with our products and how they’re—what their experience is when they interact with a bill or AWS platform at large. There’s also a couple of different personas here, right? We have a persona that focuses in on that cloud cost, the cloud bill, the finance, whether that’s—if an organization is created a FinOps organization, if they have a Cloud Center of Excellence, versus an engineering team that maybe has started to go towards decentralized IT and has some accountability for the spend that they attribute to their AWS bill. And so, these different personas interact with us in really different ways, where Cost Explorer downloading the CUR and taking a look at the bill.

And one thing that I always kind of imagine is somebody putting a headlamp on and going into the caves in the depths of their AWS bill and kind of like spelunking through their bill sometimes, right? And so, you have these FinOps folks and billing people that are deeply interested in making sure that the spend they do have meets their business goals, meaning this is providing high value to our company, it’s providing high value to our customers, and we’re spending on the right things, we’re spending the right amount on the right things. Versus the engineering organization that’s like, “Hey, how do we configure these resources? What types of instances should we be focused on using? What services should we be building on top of that maybe are more flexible for our business needs?”

And so, there’s really, like, two major personas that I spend a lot of time—our organization spends a lot of time wrapping our heads around. Because they’re really different, very different approaches to how we think about cost. Because you’re right, if you just wanted to lower your AWS bill, it’s really easy. Just size everything to a t2.nano and you’re done and move on [laugh], right? But you’re [crosstalk 00:08:53]—

Corey: Aw, t3 or t4.nano, depending upon whether regional availability is going to save you less. I’m still better at this. Let’s not kid ourselves I kid. Mostly.

Rick: For sure. So t4.nano, absolutely.

Corey: T4g. Remember, now the way forward is everything has an explicit letter designator to define which processor company made the CPU that underpins the instance itself because that’s a level of abstraction we certainly wouldn’t want the cloud provider to take away from us any.

Rick: Absolutely. And actually, the performance differences of those different processor models can be pretty incredible [laugh]. So, there’s huge decisions behind all of that as well.

Corey: Oh, yeah. There’s so many factors that factor in all these things. It’s gotten to a point of you see this usually with lawyers and very senior engineers, but the answer to almost everything is, “It depends.” There are always going to be edge cases. Easy example of, if you check a box and enable an S3 Gateway endpoint inside of a private subnet, suddenly, you’re not passing traffic through a 4.5 cent per gigabyte managed NAT Gateway; it’s being sent over that endpoint for no additional cost whatsoever.

Check the box, save a bunch of money. But there are scenarios where you don’t want to do it, so always double-checking and talking to customers about this is critically important. Just because, the first time you make a recommendation that does not work for their constraints, you lose trust. And make a few of those and it looks like you’re more or less just making naive recommendations that don’t add any value, and they learn to ignore you. So, down the road, when you make a really high-value, great recommendation for them, they stop paying attention.

Rick: Absolutely. And we have that really high bar for recommendation accuracy, especially with right sizing, that’s such a key one. Although I guess Savings Plan purchase recommendations can be critical as well. If a customer over commits on the amount of Savings Plan purchase they need to make, right, that’s a really big problem for them.

So, recommendation accuracy must be above reproach. Essentially, if a customer takes a recommendation and it breaks an application, they’re probably never going to take another right-sizing recommendation again [laugh]. And so, this bar of trust must be exceptionally high. That’s also why out of the box, the compute optimizer recommendations can be a little bit mild, they’re a little time because the first order of business is do no harm; focus on the performance requirements of the application first because we have to make sure that the reason you build these workloads in AWS is served.

Now ideally, we do that without overspending and without overprovisioning the capacity of these workloads, right? And so, for example, like if we make these right-sizing recommendations from Compute Optimizer, we’re taking a look at the utilization of CPU, memory, disk, network, throughput, iops, and we’re vending these recommendations to customers. And when you take that recommendation, you must still have great application performance for your business to be served, right? It’s such a crucial part of how we optimize and run long-term. Because optimization is not a one-time Band-Aid; it’s an ongoing behavior, so it’s really critical that for that accuracy to be exceptionally high so we can build business process on top of it as well.

Corey: Let me ask you this. How do you contextualize what the right approach to optimization is? What is your entire—there are certain tools that you have… by ‘you,’ I mean, of course, as an organization—have repeatedly gone back to and different approaches that don’t seem to deviate all that much from year to year, and customer to customer. How do you think about the general things that apply universally?

Rick: So, we know that EC2 is a very popular service for us. We know that sizing EC2 is difficult. We think about that optimization pillar of scaling. It’s an obvious area for us to help customers. We run into this sort of industry-wide experience where whenever somebody picks the size of a resource, they’re going to pick one generally larger than they need.

It’s almost like asking a new employee at your company, “Hey, pick your laptop. We have a 16 gig model or a 32 gig model. Which one do you want?” That person [laugh] making the decision on capacity, hardware capacity, they’re always going to pick the 32 gig model laptop, right? And so, we have this sort of human nature in IT of, we don’t want to get called at two in the morning for performance issues, we don’t want our apps to fall over, we want them to run really well, so we’re going to size things very conservatively and we’re going to oversize things.

So, we can help customers by providing those recommendations to say, you can size things up in a different way using math and analytics based on the utilization patterns, and we can provide and pick different instance types. There’s hundreds and hundreds of instance types in all of these regions across the globe. How do you know which is the right one for every single resource you have? It’s a very, very hard problem to solve and it’s not something that is lucrative to solve one by one if you have 100 EC2 instances. Trying to pick the correct size for each and every one can take hours and hours of IT engineering resources to look at utilization graphs, look at all of these types available, look at what is the performance difference between processor models and providers of those processors, is there application compatibility constraints that I have to consider? The complexity is astronomical.

And then not only that, as soon as you make that sizing decision, one week later, it’s out of date and you need a different size. So, [laugh] you didn’t really solve the problem. So, we have to programmatically use data science and math to say, “Based on these utilization values, these are the sizes that would make sense for your business, that would have the lowest cost and the highest performance together at the same time.” And it’s super important that we provide this capability from a technology standpoint because it would cost so much money to try to solve that problem that the savings you would achieve might not be meaningful. Then at the same time… you know, that’s really from an engineering perspective, but when we talk to the FinOps and the finance folks, the conversations are more about Reservations and Savings Plans.

How do we correctly apply Savings Plans and Reservations across a high percentage of our portfolio to reduce the costs on those workloads, but not so much that dynamic capacity levels in our organization mean we all of a sudden have a bunch of unused Reservations or Savings Plans? And so, a lot of organizations that engage with us and we have conversations with, we start with the Reservation and Savings Plan conversation because it’s much easier to click a few buttons and buy a Savings Plan than to go institute an entire right-sizing campaign across multiple engineering teams. That can be very difficult, a much higher bar. So, some companies are ready to dive into the engineering task of sizing; some are not there yet. And they’re a little maybe a little earlier in their FinOps journey, or the building optimization technology stacks, or achieving higher value out of their cloud environments, so starting with kind of the low hanging fruit, it can vary depending on the company, size of company, technical aptitudes, skill sets, all sorts of things like that.

And so, those finance-focused teams are definitely spending more time looking at and studying what are the best practices for purchasing Savings Plans, covering my environment, getting the most out of my dollar that way. Then they don’t have to engage the engineering teams; they can kind of take a nice chunk off the top of their bill and sort of have something to show for that amount of effort. So, there’s a lot of different approaches to start in optimization.

Corey: My philosophy runs somewhat counter to this because everything you’re saying does work globally, it’s safe, it’s non-threatening, and then also really, on some level, feels like it is an approach that can be driven forward by finance or business. Whereas my worldview is that cost and architecture in cloud are one and the same. And there are architectural consequences of cost decisions and vice versa that can be adjusted and addressed. Like, one of my favorite party tricks—although I admit, it’s a weird party—is I can look at the exploded PDF view of a customer’s AWS bill and describe their architecture to them. And people have questioned that a few times, and now I have a testimonial on my client website that mentions, “It was weird how he was able to do this.”

Yeah, it’s real, I can do it. And it’s not a skill, I would recommend cultivating for most people. But it does also mean that I think I’m onto something here, where there’s always context that needs to be applied. It feels like there’s an entire ecosystem of product companies out there trying to build what amount to a better Cost Explorer that also is not free the way that Cost Explorer is. So, the challenge I see there’s they all tend to look more or less the same; there is very little differentiation in that space. And in the fullness of time, Cost Explorer does—ideally—get better. How do you think about it?

Rick: Absolutely. If you’re looking at ways to understand your bill, there’s obviously Cost Explorer, the CUR, that’s a very common approach is to take the CUR and put a BI front-end on top of it. That’s a common experience. A lot of companies that have chops in that space will do that themselves instead of purchasing a third-party product that does do bill breakdown and dissemination. There’s also the cross-charge show-back organizational breakdown and boundaries because you have these super large organizations that have fiefdoms.

You know, if HR IT and sales IT, and [laugh] you know, product IT, you have all these different IT departments that are fiefdoms within your AWS bill and construct, whether they have different ABS accounts or say different AWS organizations sometimes, right, it can get extremely complicated. And some organizations require the ability to break down their bill based on those organizational boundaries. Maybe tagging works, maybe it doesn’t. Maybe they do that by using a third-party product that lets them set custom scopes on their resources based on organizational boundaries. That’s a common approach as well.

We do also have our first-party solutions, they can do that, like the CUDOS dashboard as well. That’s something that’s really popular and highly used across our customer base. It allows you to have kind of a dashboard and customizable view of your AWS costs and, kind of, split it up based on tag organizational value, account name, and things like that as well. So, you mentioned that you feel like the architectural and cost problem is the same problem. I really don’t disagree with that at all.

I think what it comes down to is some organizations are prepared to tackle the architectural elements of cost and some are not. And it really comes down to how does the customer view their bill? Is it somebody in the finance organization looking at the bill? Is it somebody in the engineering organization looking at the bill? Ideally, it would be both.

Ideally, you would have some of those skill sets that overlap, or you would have an organization that does focus in on FinOps or cloud operations as it relates to cost. But then at the same time, there are organizations that are like, “Hey, we need to go to cloud. Our CIO told us go to cloud. We don’t want to pay the lease renewal on this building.” There’s a lot of reasons why customers move to cloud, a lot of great reasons, right? Three major reasons you move to cloud: agility, [crosstalk 00:20:11]—

Corey: And several terrible ones.

Rick: Yeah, [laugh] and some not-so-great ones, too. So, there’s so many different dynamics that get exposed when customers engage with us that they might or might not be ready to engage on the architectural element of how to build hyperscale systems. So, many of these customers are bringing legacy workloads and applications to the cloud, and something like a re-architecture to use stateless resources or something like Spot, that’s just not possible for them. So, how can they take 20% off the top of their bill? Savings Plans or Reservations are kind of that easy, low-hanging fruit answer to just say, “We know these are fairly static environments that don’t change a whole lot, that are going to exist for some amount of time.”

They’re legacy, you know, we can’t turn them off. It doesn’t make sense to rewrite these applications because they just don’t change, they don’t have high business value, or something like that. And so, the architecture part of that conversation doesn’t always come into play. Should it? Yes.

The long-term maturity and approach for cloud optimization does absolutely account for architecture, thinking strategically about how you do scaling, what services you’re using, are you going down the Kubernetes path, which I know you’re going to laugh about, but you know, how do you take these applications and componentize them? What services are you using to do that? How do you get that long-term scale and manageability out of those environments? Like you said at the beginning, the complexity is staggering and there’s no one unified answer. That’s why there’s so many different entrance paths into, “How do I optimize my AWS bill?”

There’s no one answer, and every customer I talk to has a different comfort level and appetite. And some of them have tried suspension, some of them have gone heavy down Savings Plans, some of them want to dabble in right-sizing. So, every customer is different and we want to provide those capabilities for all of those different customers that have different appetites or comfort levels with each of these approaches.

Corey: This episode is sponsored in part by our friends at Redis, the company behind the incredibly popular open source database. If you’re tired of managing open source Redis on your own, or if you are looking to go beyond just caching and unlocking your data’s full potential, these folks have you covered. Redis Enterprise is the go-to managed Redis service that allows you to reimagine how your geo-distributed applications process, deliver, and store data. To learn more from the experts in Redis how to be real-time, right now, from anywhere, visit redis.com/duckbill. That’s R - E - D - I - S dot com slash duckbill.

Corey: And I think that’s very fair. I think that it is not necessarily a bad thing that you wind up presenting a lot of these options to customers. But there are some rough edges. An example of this is something I encountered myself somewhat recently and put on Twitter—because I have those kinds of problems—where originally, I remember this, that you were able to buy hourly Savings Plans, which again, Savings Plans are great; no knock there. I would wish that they applied to more services rather than, “Oh, SageMaker is going to do its own Savings Pla”—no, stop keeping me from going from something where I have to manage myself on EC2 to something you manage for me and making that cost money. You nailed it with Fargate. You nailed it with Lambda. Please just have one unified Savings Plan thing. But I digress.

But you had a limit, once upon a time, of $1,000 per hour. Now, it’s $5,000 per hour, which I believe in a three-year all-up-front means you will cheerfully add $130 million purchase to your shopping cart. And I kept adding a bunch of them and then had a little over a billion dollars a single button click away from being charged to my account. Let me begin with what’s up with that?

Rick: [laugh]. Thank you for the tweet, by the way, Corey.

Corey: Always thrilled to ruin your month, Rick. You know that.

Rick: Yeah. Fantastic. We took that tweet—you know, it was tongue in cheek, but also it was a serious opportunity for us to ask a question of what does happen? And it’s something we did ask internally and have some fun conversations about. I can tell you that if you clicked purchase, it would have been declined [laugh]. So, you would have not been—

Corey: Yeah, American Express would have had a problem with that. But the question is, would you have attempted to charge American Express, or would something internally have gone, “This has a few too many commas for us to wind up presenting it to the card issuer with a straight face?”

Rick: [laugh]. Right. So, it wouldn’t have gone through and I can tell you that, you know, if your account was on a PO-based configuration, you know, it would have gone to the account team. And it would have gone through our standard process for having a conversation with our customer there. That being said, we are—it’s an awesome opportunity for us to examine what is that shopping cart experience.

We did increase the limit, you’re right. And we increased the limit for a lot of reasons that we sat down and worked through, but at the same time, there’s always an opportunity for improvement of our product and experience, we want to make sure that it’s really easy and lightweight to use our products, especially purchasing Savings Plans. Savings Plans are already kind of wrought with mental concern and risk of purchasing something so expensive and large that has a big impact on your AWS bill, so we don’t really want to add any more friction necessarily the process but we do want to build an awareness and make sure customers understand, “Hey, you’re purchasing this. This has a pretty big impact.” And so, we’re also looking at other ways we can kind of improve the ability for the Savings Plan shopping cart experience to ensure customers don’t put themselves in a position where you have to unwind or make phone calls and say, “Oops.” Right? We [laugh] want to avoid those sorts of situations for our customers. So, we are looking at quite a few additional improvements to that experience as well that I’m really excited about that I really can’t share here, but stay tuned.

Corey: I am looking forward to it. I will say the counterpoint to that is having worked with customers who do make large eight-figure purchases at once, there’s a psychology element that plays into it. Everyone is very scared to click the button on the ‘Buy It Now’ thing or the ‘Approve It.’ So, what I’ve often found is at that scale, one, you can reduce what you’re buying by half of it, and then see how that treats you and then continue to iterate forward rather than doing it all at once, or reach out to your account team and have them orchestrate the buy. In previous engagements, I had a customer do this religiously and at one point, the concierge team bought the wrong thing in the wrong region, and from my perspective, I would much rather have AWS apologize for that and fix it on their end, than from us having to go with a customer side of, “Oh crap, oh, crap. Please be nice to us.”

Not that I doubt you would do it, but that’s not the nervous conversation I want to have in quite the same way. It just seems odd to me that someone would want to make that scale of purchase without ever talking to a human. I mean, I get it. I’m as antisocial as they come some days, but for that kind of money, I kind of just want another human being to validate that I’m not making a giant mistake.

Rick: We love that. That’s such a tremendous opportunity for us to engage and discuss with an organization that’s going to make a large commitment, that here’s the impact, here’s how we can help. How does it align to our strategy? We also do recommend, from a strategic perspective, those more incremental purchases. I think it creates a better experience long-term when you don’t have a single Savings Plan that’s going to expire on a specific day that all of a sudden increases your entire bill by a significant percentage.

So, making staggered monthly purchases makes a lot of sense. And it also works better for incremental growth, right? If your organization is growing 5% month-over-month or year-over-year or something like that, you can purchase those incremental Savings Plans that sort of stack up on top of each other and then you don’t have that risk of a cliff one day where one super-large SP expires and boom, you have to scramble and repurchase within minutes because every minute that goes by is an additional expense, right? That’s not a great experience. And so that’s, really, a large part of why those staggered purchase experiences make a lot of sense.

That being said, a lot of companies do their math and their finance in different ways. And single large purchases makes sense to go through their process and their rigor as well. So, we try to support both types of purchasing patterns.

Corey: I think that is an underappreciated aspect of cloud cost savings and cloud cost optimization, where it is much more about humans than it is about math. I see this most notably when I’m helping customers negotiate their AWS contracts with AWS, where they are often perspectives such as, “Well, we feel like we really got screwed over last time, so we want to stick it to them and make them give us a bigger percentage discount on something.” And it’s like, look, you can do that, but I would much rather, if it were me, go for something that moves the needle on your actual business and empowers you to move faster, more effectively, and lead to an outcome that is a positive for everyone versus the well, we’re just going to be difficult in this one point because they were difficult on something last time. But ego is a thing. Human psychology is never going to have an API for it. And again, customers get to decide their own destiny in some cases.

Rick: I completely agree. I’ve actually experienced that. So, this is the third company I’ve been working at on Cloud optimization. I spent several years at Microsoft running an optimization program. I went to Turbonomic for several years, building out the right-sizing and savings plan reservation purchase capabilities there, and now here at AWS.

And through all of these journeys and experiences working with companies to help optimize their cloud spend, I can tell you that the psychological needle—moving the needle is significantly harder than the technology stack of sizing something correctly or deleting something that’s unused. We can solve the technology part. We can build great products that identify opportunities to save money. There’s still this psychological component of IT, for the last several decades has gone through this maturity curve of if it’s not broken, don’t touch it. Five-nines, six sigma, all of these methods of IT sort of rationalizing do no harm, don’t touch anything, everything must be up.

And it even kind of goes back several decades. Back when if you rebooted a physical server, the motherboard capacitors would pop, right? So, there’s even this anti—or this stigma against even rebooting servers sometimes. In the cloud really does away with a lot of that stuff because we have live migration and we have all of these, sort of, stateless designs and capabilities, but we still carry along with us this mentality of don’t touch it; it might fall over. And we have to really get past that.

And that means that the trust, we went back to the trust conversation where we talk about the recommendations must be incredibly accurate. You’re risking your job, in some cases; if you are a DevOps engineer, and your commitments on your yearly goals are uptime, latency, response time, load time, these sorts of things, these operational metrics, KPIs that you use, you don’t want to take a downsized recommendation. It has a severe risk of harming your job and your bonus.

Corey: “These instances are idle. Turn them off.” It’s like, yeah, these instances are the backup site, or the DR environment, or—

Rick: Exactly.

Corey: —something that takes very bursty but occasional traffic. And yeah, I know it costs us some money, but here’s the revenue figures for having that thing available. Like, “Oh, yeah. Maybe we should shut up and not make dumb recommendations around things,” is the human response, but computers don’t have that context.

Rick: Absolutely. And so, the accuracy and trust component has to be the highest bar we meet for any optimization activity or behavior. We have to circumvent or supersede the human aversion, the risk aversion, that IT is built on, right?

Corey: Oh, absolutely. And let’s be clear, we see this all the time where I’m talking to customers and they have been burned before because we tried to save money and then we took a production outage as a side effect of a change that we made, and now we’re not allowed to try to save money anymore. And there’s a hidden truth in there, which is auto-scaling is something that a lot of customers talk about, but very few have instrumented true auto-scaling because they interpret is we can scale up to meet demand. Because yeah, if you don’t do that you’re dropping customers on the floor.

Well, what about scaling back down again? And the answer there is like, yeah, that’s not really a priority because it’s just money. We’re not disappointing customers, causing brand reputation, and we’re still able to take people’s money when that happens. It’s only money; we can fix it later. Covid shined a real light on a lot of the stuff just because there are customers that we’ve spoken to who’s—their user traffic dropped off a cliff, infrastructure spend remained constant day over day.

And yeah, they believe, genuinely, they were auto-scaling. The most interesting lies are the ones that customers tell themselves, but the bill speaks. So, getting a lot of modernization traction from things like that was really neat to watch. But customers I don’t think necessarily intuitively understand most aspects of their bill because it is a multidisciplinary problem. It’s engineering, its finance, its accounting—which is not the same thing as finance—and you need all three of those constituencies to be able to communicate effectively using a shared and common language. It feels like we’re marriage counseling between engineering and finance, most weeks.

Rick: Absolutely, we are. And it’s important we get it right, that the data is accurate, that the recommendations we provide are trustworthy. If the finance team gets their hands on the savings potential they see out of right-sizing, takes it to engineering, and then engineering comes back and says, “No, no, no, we can’t actually do that. We can’t actually size those,” right, we have problems. And they’re cultural, they’re transformational. Organizations’ appetite for these things varies greatly and so it’s important that we address that problem from all of those angles. And it’s not easy to do.

Corey: How big do you find the optimization problem is when you talk to customers? How focused are they on it? I have my answers, but that’s the scale of anec-data. I want to hear your actual answer.

Rick: Yeah. So, we talk with a lot of customers that are very interested in optimization. And we’re very interested in helping them on the journey towards having an optimal estate. There are so many nuances and barriers, most of them psychological like we already talked about.

I think there’s this opportunity for us to go do better exposing the potential of what an optimal AWS estate would look like from a dollar and savings perspective. And so, I think it’s kind of not well understood. I think it’s one of the biggest areas or barriers of companies really attacking the optimization problem with more vigor is if they knew that the potential savings they could achieve out of their AWS environment would really align their spend much more closely with the business value they get, I think everybody would go bonkers. And so, I’m really excited about us making progress on exposing that capability or the total savings potential and amount. It’s something we’re looking into doing in a much more obvious way.

And we’re really excited about customers doing that on AWS where they know they can trust AWS to get the best value for their cloud spend, that it’s a long-term good bet because their resources that they’re using on AWS are all focused on giving business value. And that’s the whole key. How can we align the dollars to the business value, right? And I think optimization is that connection between those two concepts.

Corey: Companies are generally not going to greenlight a project whose sole job is to save money unless there’s something very urgent going on. What will happen is as they iterate forward on the next generation of services or a migration of a service from one thing to another, they will make design decisions that benefit those optimizations. There’s low-hanging fruit we can find, usually of the form, “Turn that thing off,” or, “Configure this thing slightly differently,” that doesn’t take a lot of engineering effort in place. But, on some level, it is not worth the engineering effort it takes to do an optimization project. We’ve all met those engineers—speaking is one of them myself—who, left to our own devices, will spend two months just knocking a few hundred bucks a month off of our AWS developer environment.

We steal more than office supplies. I’m not entirely sure what the business value of doing that is, in most cases. For me, yes, okay, things that work in small environments work very well in large environments, generally speaking, so I learned how to save 80 cents here and that’s a few million bucks a month somewhere else. Most folks don’t have that benefit happening, so it’s a question of meeting them where they are.

Rick: Absolutely. And I think the scale component is huge, which you just touched on. When you’re talking about a hundred EC2 instances versus a thousand, optimization becomes kind of a different component of how you manage that AWS environment. And while single-decision recommendations to scale an individual server, the dollar amount might be different, the percentages are just about the same when you look at what is it to be sized correctly, what is it to be configured correctly? And so, it really does come down to priority.

And so, it’s really important to really support all of those companies of all different sizes and industries because they will have different experiences on AWS. And some will have more sensitivity to cost than others, but all of them want to get great business value out of their AWS spend. And so, as long as we’re meeting that need and we’re supporting our customers to make sure they understand the commitment we have to ensuring that their AWS spend is valuable, it is meaningful, right, they’re not spending money on things that are not adding value, that’s really important to us.

Corey: I do want to have as the last topic of discussion here, how AWS views optimization, where there have been a number of repeated statements where helping customers optimize their cloud spend is extremely important to us. And I’m trying to figure out where that falls on the spectrum from, “It’s the thing we say because they make us say it, but no, we’re here to milk them like cows,” all the way on over to, “No, no, we passionately believe in this at every level, top to bottom, in every company. We are just bad at it.” So, I’m trying to understand how that winds up being expressed from your lived experience having solved this problem first outside, and then inside.

Rick: Yeah. So, it’s kind of like part of my personal story. It’s the main reason I joined AWS. And, you know, when you go through the interview loops and you talk to the leaders of an organization you’re thinking about joining, they always stop at the end of the interview and ask, “Do you have any questions for us?” And I asked that question to pretty much every single person I interviewed with. Like, “What is AWS’s appetite for helping customers save money?”

Because, like, from a business perspective, it kind of is a little bit wonky, right? But the answers were varied, and all of them were customer-obsessed and passionate. And I got this sense that my personal passion for helping companies have better efficiency of their IT resources was an absolute primary goal of AWS and a big element of Amazon’s leadership principle, be customer obsessed.

Now, I’m not a spokesperson, so [laugh] we’ll see, but we are deeply interested in making sure our customers have a great long-term experience and a high-trust relationship. And so, when I asked these questions in these interviews, the answers were all about, “We have to do the right thing for the customer. It’s imperative. It’s also in our DNA. It’s one of the most important leadership principles we have to be customer-obsessed.”

And it is the primary reason why I joined: because of that answer to that question. Because it’s so important that we achieve a better efficiency for our IT resources, not just for, like, AWS, but for our planet. If we can reduce consumption patterns and usage across the planet for how we use data centers and all the power that goes into them, we can talk about meaningful reductions of greenhouse gas emissions, the cost and energy needed to run IT business applications, and not only that, but most all new technology that’s developed in the world seems to come out of a data center these days, we have a real opportunity to make a material impact to how much resource we use to build and use these things. And I think we owe it to the planet, to humanity, and I think Amazon takes that really seriously. And I’m really excited to be here because of that.

Corey: As I recall—and feel free to make sure that this comment never sees the light of day—you asked me before interviewing for the role and then deciding to accept it, what I thought about you working there and whether I would recommend it, whether I wouldn’t. And I think my answer was fairly nuanced. And you’re working there now and we still are on speaking terms, so people can probably guess what my comments took the shape of, generally speaking. So, I’m going to have to ask now; it’s been, what, a year since you joined?

Rick: Almost. I think it’s been about eight months.

Corey: Time during a pandemic is always strange. But I have to ask, did I steer you wrong?

Rick: No. Definitely not. I’m very happy to be here. The opportunity to help such a broad range of companies get more value out of technology—and it’s not just cost, right, like we talked about. It’s actually not about the dollar number going down on a bill. It’s about getting more value and moving the needle on how do we efficiently use technology to solve business needs.

And that’s been my career goal for a really long time, I’ve been working on optimization for, like, seven or eight, I don’t know, maybe even nine years now. And it’s like this strange passion for me, this combination of my dad taught me how to be a really good steward of money and a great budget manager, and then my passion for technology. So, it’s this really cool combination of, like, childhood life skills that really came together for me to create a career that I’m really passionate about. And this move to AWS has been such a tremendous way to supercharge my ability to scale my personal mission, and really align it to AWS’s broader mission of helping companies achieve more with cloud platforms, right?

And so, it’s been a really nice eight months. It’s been wild. Learning AWS culture has been wild. It’s a sharp diverging culture from where I’ve been in the past, but it’s also really cool to experience the leadership principles in action. They’re not just things we put on a website; they’re actually things people talk about every day [laugh]. And so, that journey has been humbling and a great learning opportunity as well.

Corey: If people want to learn more, where’s the best place to find you?

Rick: Oh, yeah. Contact me on LinkedIn or Twitter. My Twitter account is @rickyo1138. Let me know if you get the 1138 reference. That’s a fun one.

Corey: THX 1138. Who doesn’t?

Rick: Yeah, there you go. And it’s hidden in almost every single George Lucas movie as well. You can contact me on any of those social media platforms and I’d be happy to engage with anybody that’s interested in optimization, cloud technology, bill, anything like that. Or even not [laugh]. Even anything else, either.

Corey: Thank you so much for being so generous with your time. I really appreciate it.

Rick: My pleasure, Corey. It was wonderful talking to you.

Corey: Rick Ochs, Principal Product Manager at AWS. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an angry comment, rightly pointing out that while AWS is great and all, Azure is far more cost-effective for your workloads because, given their lack security, it is trivially easy to just run your workloads in someone else’s account.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Sam

Sam Nicholls: Veeam’s Director of Public Cloud Product Marketing, with 10+ years of sales, alliance management and product marketing experience in IT. Sam has evolved from his on-premises storage days and is now laser-focused on spreading the word about cloud-native backup and recovery, packing in thousands of viewers on his webinars, blogs and webpages.

Links Referenced:

  • Veeam AWS Backup: https://www.veeam.com/aws-backup.html
  • Veeam: https://veeam.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Chronosphere. Tired of observability costs going up every year without getting additional value? Or being locked in to a vendor due to proprietary data collection, querying and visualization? Modern day, containerized environments require a new kind of observability technology that accounts for the massive increase in scale and attendant cost of data. With Chronosphere, choose where and how your data is routed and stored, query it easily, and get better context and control. 100% open source compatibility means that no matter what your setup is, they can help. Learn how Chronosphere provides complete and real-time insight into ECS, EKS, and your microservices, whereever they may be at snark.cloud/chronosphere That’s snark.cloud/chronosphere

Corey: This episode is brought to us by our friends at Pinecone. They believe that all anyone really wants is to be understood, and that includes your users. AI models combined with the Pinecone vector database let your applications understand and act on what your users want… without making them spell it out. Make your search application find results by meaning instead of just keywords, your personalization system make picks based on relevance instead of just tags, and your security applications match threats by resemblance instead of just regular expressions. Pinecone provides the cloud infrastructure that makes this easy, fast, and scalable. Thanks to my friends at Pinecone for sponsoring this episode. Visit Pinecone.io to understand more.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. This promoted guest episode is brought to us by and sponsored by our friends over at Veeam. And as a part of that, they have thrown one of their own to the proverbial lion. My guest today is Sam Nicholls, Director of Public Cloud over at Veeam. Sam, thank you for joining me.

Sam: Hey. Thanks for having me, Corey, and thanks for everyone joining and listening in. I do know that I’ve been thrown into the lion’s den, and I am [laugh] hopefully well-prepared to answer anything and everything that Corey throws my way. Fingers crossed. [laugh].

Corey: I don’t think there’s too much room for criticizing here, to be direct. I mean, Veeam is a company that is solidly and thoroughly built around a problem that absolutely no one cares about. I mean, what could possibly be wrong with that? You do backups; which no one ever cares about. Restores, on the other hand, people care very much about restores. And that’s when they learn, “Oh, I really should have cared about backups at any point prior to 20 minutes ago.”

Sam: Yeah, it’s a great point. It’s kind of like taxes and insurance. It’s almost like, you know, something that you have to do that you don’t necessarily want to do, but when push comes to shove, and something’s burning down, a file has been deleted, someone’s made their way into your account and, you know, running a right mess within there, that’s when you really, kind of, care about what you mentioned, which is the recovery piece, the speed of recovery, the reliability of recovery.

Corey: It’s been over a decade, and I’m still sore about losing my email archives from 2006 to 2009. There’s no way to get it back. I ran my own mail server; it was an iPhone setting that said, “Oh, yeah, automatically delete everything in your trash folder—or archive folder—after 30 days.” It was just a weird default setting back in that era. I didn’t realize it was doing that. Yeah, painful stuff.

And we learned the hard way in some of these cases. Not that I really have much need for email from that era of my life, but every once in a while it still bugs me. Which gets speaks to the point that the people who are the most fanatical about backing things up are the people who have been burned by not having a backup. And I’m fortunate in that it wasn’t someone else’s data with which I had been entrusted that really cemented that lesson for me.

Sam: Yeah, yeah. It’s a good point. I could remember a few years ago, my wife migrated a very aging, polycarbonate white Mac to one of the shiny new aluminum ones and thought everything was good—

Corey: As the white polycarbonate Mac becomes yellow, then yeah, all right, you know, it’s time to replace it. Yeah. So yeah, so she wiped the drive, and what happened?

Sam: That was her moment where she learned the value and importance of backup unless she backs everything up now. I fortunately have never gone through it. But I’m employed by a backup vendor and that’s why I care about it. But it’s incredibly important to have, of course.

Corey: Oh, yes. My spouse has many wonderful qualities, but one that drives me slightly nuts is she’s something of a digital packrat where her hard drives on her laptop will periodically fill up. And I used to take the approach of oh, you can be more efficient and do the rest. And I realized no, telling other people they’re doing it wrong is generally poor practice, whereas just buying bigger drives is way easier. Let’s go ahead and do that. It’s small price to pay for domestic tranquility.

And there’s a lesson in that. We can map that almost perfectly to the corporate world where you folks tend to operate in. You’re not doing home backup, last time I checked; you are doing public cloud backup. Actually, I should ask that. Where do you folks start and where do you stop?

Sam: Yeah, no, it’s a great question. You know, we started over 15 years ago when virtualization, specifically VMware vSphere, was really the up-and-coming thing, and, you know, a lot of folks were there trying to utilize agents to protect their vSphere instances, just like they were doing with physical Windows and Linux boxes. And, you know, it kind of got the job done, but was it the best way of doing it? No. And that’s kind of why Veeam was pioneered; it was this agentless backup, image-based backup for vSphere.

And, of course, you know, in the last 15 years, we’ve seen lots of transitions, of course, we’re here at Screaming in the Cloud, with you, Corey, so AWS, as well as a number of other public cloud vendors we can help protect as well, as a number of SaaS applications like Microsoft 365, metadata and data within Salesforce. So, Veeam’s really kind of come a long way from just virtual machines to really taking a global look at the entirety of modern environments, and how can we best protect each and every single one of those without trying to take a square peg and fit it in a round hole?

Corey: It’s a good question and a common one. We wind up with an awful lot of folks who are confused by the proliferation of data. And I’m one of them, let’s be very clear here. It comes down to a problem where backups are a multifaceted, deep problem, and I don’t think that people necessarily think of it that way. But I take a look at all of the different, even AWS services that I use for my various nonsense, and which ones can be used to store data?

Well, all of them. Some of them, you have to hold it in a particularly wrong sort of way, but they all store data. And in various contexts, a lot of that data becomes very important. So, what service am I using, in which account am I using, and in what region am I using it, and you wind up with data sprawl, where it’s a tremendous amount of data that you can generally only track down by looking at your bills at the end of the month. Okay, so what am I being charged, and for what service?

That seems like a good place to start, but where is it getting backed up? How do you think about that? So, some people, I think, tend to ignore the problem, which we’re seeing less and less, but other folks tend to go to the opposite extreme and we’re just going to backup absolutely everything, and we’re going to keep that data for the rest of our natural lives. It feels to me that there’s probably an answer that is more appropriate somewhere nestled between those two extremes.

Sam: Yeah, snapshot sprawl is a real thing, and it gets very, very expensive very, very quickly. You know, your snapshots of EC2 instances are stored on those attached EBS volumes. Five cents per gig per month doesn’t sound like a lot, but when you’re dealing with thousands of snapshots for thousands machines, it gets out of hand very, very quickly. And you don’t know when to delete them. Like you say, folks are just retaining them forever and dealing with this unfortunate bill shock.

So, you know, where to start is automating the lifecycle of a snapshot, right, from its creation—how often do we want to be creating them—from the retention—how long do we want to keep these for—and where do we want to keep them because there are other storage services outside of just EBS volumes. And then, of course, the ultimate: deletion. And that’s important even from a compliance perspective as well, right? You’ve got to retain data for a specific number of years, I think healthcare is like seven years, but then you’ve—

Corey: And then not a day more.

Sam: Yeah, and then not a day more because that puts you out of compliance, too. So, policy-based automation is your friend and we see a number of folks building these policies out: gold, silver, bronze tiers based on criticality of data compliance and really just kind of letting the machine do the rest. And you can focus on not babysitting backup.

Corey: What was it that led to the rise of snapshots? Because back in my very early days, there was no such thing. We wound up using a bunch of servers stuffed in a rack somewhere and virtualization was not really in play, so we had file systems on physical disks. And how do you back that up? Well, you have an agent of some sort that basically looks at all the files and according to some ruleset that it has, it copies them off somewhere else.

It was slow, it was fraught, it had a whole bunch of logic that was pushed out to the very edge, and forget about restoring that data in a timely fashion or even validating a lot of those backups worked other than via checksum. And God help you if you had data that was constantly in the state of flux, where anything changing during the backup run would leave your backups in an inconsistent state. That on some level seems to have largely been solved by snapshots. But what’s your take on it? You’re a lot closer to this part of the world than I am.

Sam: Yeah, snapshots, I think folks have turned to snapshots for the speed, the lack of impact that they have on production performance, and again, just the ease of accessibility. We have access to all different kinds of snapshots for EC2, RDS, EFS throughout the entirety of our AWS environment. So, I think the snapshots are kind of like the default go-to for folks. They can help deliver those very, very quick RPOs, especially in, for example, databases, like you were saying, that change very, very quickly and we all of a sudden are stranded with a crash-consistent backup or snapshot versus an application-consistent snapshot. And then they’re also very, very quick to recover from.

So, snapshots are very, very appealing, but they absolutely do have their limitations. And I think, you know, it’s not a one or the other; it’s that they’ve got to go hand-in-hand with something else. And typically, that is an image-based backup that is stored in a separate location to the snapshot because that snapshot is not independent of the disk that it is protecting.

Corey: One of the challenges with snapshots is most of them are created in a copy-on-write sense. It takes basically an instant frozen point in time back—once upon a time when we ran MySQL databases on top of the NetApp Filer—which works surprisingly well—we would have a script that would automatically quiesce the database so that it would be in a consistent state, snapshot the file and then un-quiesce it, which took less than a second, start to finish. And that was awesome, but then you had this snapshot type of thing. It wasn’t super portable, it needed to reference a previous snapshot in some cases, and AWS takes the same approach where the first snapshot it captures every block, then subsequent snapshots wind up only taking up as much size as there have been changes since the first snapshots. So, large quantities of data that generally don’t get access to a whole lot have remarkably small, subsequent snapshot sizes.

But that’s not at all obvious from the outside, and looking at these things. They’re not the most portable thing in the world. But it’s definitely the direction that the industry has trended in. So, rather than having a cron job fire off an AWS API call to take snapshots of my volumes as a sort of the baseline approach that we all started with, what is the value proposition that you folks bring? And please don’t say it’s, “Well, cron jobs are hard and we have a friendlier interface for that.”

Sam: [laugh]. I think it’s really starting to look at the proliferation of those snapshots, understanding what they’re good at, and what they are good for within your environment—as previously mentioned, low RPOs, low RTOs, how quickly can I take a backup, how frequently can I take a backup, and more importantly, how quickly can I restore—but then looking at their limitations. So, I mentioned that they were not independent of that disk, so that certainly does introduce a single point of failure as well as being not so secure. We’ve kind of touched on the cost component of that as well. So, what Veeam can come in and do is then take an image-based backup of those snapshots, right—so you’ve got your initial snapshot and then your incremental ones—we’ll take the backup from that snapshot, and then we’ll start to store that elsewhere.

And that is likely going to be in a different account. We can look at the Well-Architected Framework, AWS deeming accounts as a security boundary, so having that cross-account function is critically important so you don’t have that single point of failure. Locking down with IAM roles is also incredibly important so we haven’t just got a big wide open door between the two. But that data is then stored in a separate account—potentially in a separate region, maybe in the same region—Amazon S3 storage. And S3 has the wonderful benefit of being still relatively performant, so we can have quick recoveries, but it is much, much cheaper. You’re dealing with 2.3 cents per gig per month, instead of—

Corey: To start, and it goes down from there with sizeable volumes.

Sam: Absolutely, yeah. You can go down to S3 Glacier, where you’re looking at, I forget how many points and zeros and nines it is, but it’s fractions of a cent per gig per month, but it’s going to take you a couple of days to recover that da—

Corey: Even infrequent access cuts that in half.

Sam: Oh yeah.

Corey: And let’s be clear, these are snapshot backups; you probably should not be accessing them on a consistent, sustained basis.

Sam: Well, exactly. And this is where it’s kind of almost like having your cake and eating it as well. Compliance or regulatory mandates or corporate mandates are saying you must keep this data for this length of time. Keeping that—you know, let’s just say it’s three years' worth of snapshots in an EBS volume is going to be incredibly expensive. What’s the likelihood of you needing to recover something from two years—actually, even two months ago? It’s very, very small.

So, the performance part of S3 is, you don’t need to take it as much into consideration. Can you recover? Yes. Is it going to take a little bit longer? Absolutely. But it’s going to help you meet those retention requirements while keeping your backup bill low, avoiding that bill shock, right, spending tens and tens of thousands every single month on snapshots. This is what I mean by kind of having your cake and eating it.

Corey: I somewhat recently have had a client where EBS snapshots are one of the driving costs behind their bill. It is one of their largest single line items. And I want to be very clear here because if one of those people who listen to this and thinking, “Well, hang on. Wait, they’re telling stories about us, even though they’re not naming us by name?” Yeah, there were three of you in the last quarter.

So, at that point, it becomes clear it is not about something that one individual company has done and more about an overall driving trend. I am personalizing it a little bit by referring to as one company when there were three of you. This is a narrative device, not me breaking confidentiality. Disclaimer over. Now, when you talk to people about, “So, tell me why you’ve got 80 times more snapshots than you do EBS volumes?” The answer is as, “Well, we wanted to back things up and we needed to get hourly backups to a point, then daily backups, then monthly, and so on and so forth. And when this was set up, there wasn’t a great way to do this natively and we don’t always necessarily know what we need versus what we don’t. And the cost of us backing this up, well, you can see it on the bill. The cost of us deleting too much and needing it as soon as we do? Well, that cost is almost incalculable. So, this is the safe way to go.” And they’re not wrong in anything that they’re saying. But the world has definitely evolved since then.

Sam: Yeah, yeah. It’s a really great point. Again, it just folds back into my whole having your cake and eating it conversation. Yes, you need to retain data; it gives you that kind of nice, warm, cozy feeling, it’s a nice blanket on a winter’s day that that data, irrespective of what happens, you’re going to have something to recover from. But the question is does that need to be living on an EBS volume as a snapshot? Why can’t it be living on much, much more cost-effective storage that’s going to give you the warm and fuzzies, but is going to make your finance team much, much happier [laugh].

Corey: One of the inherent challenges I think people have is that snapshots by themselves are almost worthless, in that I have an EBS snapshot, it is sitting there now, it’s costing me an undetermined amount of money because it’s not exactly clear on a per snapshot basis exactly how large it is, and okay, great. Well, I’m looking for a file that was not modified since X date, as it was on this time. Well, great, you’re going to have to take that snapshot, restore it to a volume and then go exploring by hand. Oh, it was the wrong one. Great. Try it again, with a different one.

And after, like, the fifth or six in a row, you start doing a binary search approach on this thing. But it’s expensive, it’s time-consuming, it takes forever, and it’s not a fun user experience at all. Part of the problem is it seems that historically, backup systems have no context or no contextual awareness whatsoever around what is actually contained within that backup.

Sam: Yeah, yeah. I mean, you kind of highlighted two of the steps. It’s more like a ten-step process to do, you know, granular file or folder-level recovery from a snapshot, right? You’ve got to, like you say, you’ve got to determine the point in time when that, you know, you knew the last time that it was around, then you’re going to have to determine the volume size, the region, the OS, you’re going to have to create an EBS volume of the same size, region, from that snapshot, create the EC2 instance with the same OS, connect the two together, boot the EC2 instance, mount the volume search for the files to restore, download them manually, at which point you have your file back. It’s not back in the machine where it was, it’s now been downloaded locally to whatever machine you’re accessing that from. And then you got to tear it all down.

And that is again, like you say, predicated on the fact that you knew exactly that that was the right time. It might not be and then you have to start from scratch from a different point in time. So, backup tooling from backup vendors that have been doing this for many, many years, knew about this problem long, long ago, and really seek to not only automate the entirety of that process but make the whole e-discovery, the search, the location of those files, much, much easier. I don’t necessarily want to do a vendor pitch, but I will say with Veeam, we have explorer-like functionality, whereby it’s just a simple web browser. Once that machine is all spun up again, automatic process, you can just search for your individual file, folder, locate it, you can download it locally, you can inject it back into the instance where it was through Amazon Kinesis or AWS Kinesis—I forget the right terminology for it; some of its AWS, some of its Amazon.

But by-the-by, the whole recovery process, especially from a file or folder level, is much more pain-free, but also much faster. And that’s ultimately what people care about how reliable is my backup? How quickly can I get stuff online? Because the time that I’m down is costing me an indescribable amount of time or money.

Corey: This episode is sponsored in part by our friends at Redis, the company behind the incredibly popular open source database. If you’re tired of managing open source Redis on your own, or if you are looking to go beyond just caching and unlocking your data’s full potential, these folks have you covered. Redis Enterprise is the go-to managed Redis service that allows you to reimagine how your geo-distributed applications process, deliver, and store data. To learn more from the experts in Redis how to be real-time, right now, from anywhere, visit redis.com/duckbill. That’s R - E - D - I - S dot com slash duckbill.

Corey: Right, the idea of RPO versus RTO: recovery point objective and recovery time objective. With an RPO, it’s great, disaster strikes right now, how long is acceptable to it have been since the last time we backed up data to a restorable point? Sometimes it’s measured in minutes, sometimes it’s measured in fractions of a second. It really depends on what we’re talking about. Payments databases, that needs to be—the RPO is basically an asymptotically approaches zero.

The RTO is okay, how long is acceptable before we have that data restored and are back up and running? And that is almost always a longer time, but not always. And there’s a different series of trade-offs that go into that. But both of those also presuppose that you’ve already dealt with the existential question of is it possible for us to recover this data. And that’s where I know that you are obviously—you have a position on this that is informed by where you work, but I don’t, and I will call this out as what I see in the industry: AWS backup is compelling to me except for one fatal flaw that it has, and that is it starts and stops with AWS.

I am not a proponent of multi-cloud. Lord knows I’ve gotten flack for that position a bunch of times, but the one area where it makes absolute sense to me is backups. Have your data in a rehydrate-the-business level state backed up somewhere that is not your primary cloud provider because you’re otherwise single point of failure-ing through a company, through the payment instrument you have on file with that company, in the blast radius of someone who can successfully impersonate you to that vendor. There has to be a gap of some sort for the truly business-critical data. Yes, egress to other providers is expensive, but you know what also is expensive? Irrevocably losing the data that powers your business. Is it likely? No, but I would much rather do it than have to justify why I’m not doing it.

Sam: Yeah. Wasn’t likely that I was going to win that 2 billion or 2.1 billion on the Powerball, but [laugh] I still play [laugh]. But I understand your standpoint on multi-cloud and I read your newsletters and understand where you’re coming from, but I think the reality is that we do live in at least a hybrid cloud world, if not multi-cloud. The number of organizations that are sole-sourced on a single cloud and nothing else is relatively small, single-digit percentage. It’s around 80-some percent that are hybrid, and the remainder of them are your favorite: multi-cloud.

But again, having something that is one hundred percent sole-source on a single platform or a single vendor does expose you to a certain degree of risk. So, having the ability to do cross-platform backups, recoveries, migrations, for whatever reason, right, because it might not just be a disaster like you’d mentioned, it might also just be… I don’t know, the company has been taken over and all of a sudden, the preference is now towards another cloud provider and I want you to refactor and re-architect everything for this other cloud provider. If all that data is locked into one platform, that’s going to make your job very, very difficult. So, we mentioned at the beginning of the call, Veeam is capable of protecting a vast number of heterogeneous workloads on different platforms, in different environments, on-premises, in multiple different clouds, but the other key piece is that we always use the same backup file format. And why that’s key is because it enables portability.

If I have backups of EC2 instances that are stored in S3, I could copy those onto on-premises disk, I could copy those into Azure, I could do the same with my Azure VMs and store those on S3, or again, on-premises disk, and any other endless combination that goes with that. And it’s really kind of centered around, like control and ownership of your data. We are not prescriptive by any means. Like, you do what is best for your organization. We just want to provide you with the toolset that enables you to do that without steering you one direction or the other with fee structures, disparate feature sets, whatever it might be.

Corey: One of the big challenges that I keep seeing across the board is just a lack of awareness of what the data that matters is, where you see people backing up endless fleets of web server instances that are auto-scaled into existence and then removed, but you can create those things at will; why do you care about the actual data that’s on these things? It winds up almost at the library management problem, on some level. And in that scenario, snapshots are almost certainly the wrong answer. One thing that I saw previously that really changed my way of thinking about this was back many years ago when I was working at a startup that had just started using GitHub and they were paying for a third-party service that wound up backing up Git repos. Today, that makes a lot more sense because you have a bunch of other stuff on GitHub that goes well beyond the stuff contained within Git, but at the time, it was silly. It was, why do that? Every Git clone is a full copy of the entire repository history. Just grab it off some developer's laptop somewhere.

It’s like, “Really? You want to bet the company, slash your job, slash everyone else’s job on that being feasible and doable or do you want to spend the 39 bucks a month or whatever it was to wind up getting that out the door now so we don’t have to think about it, and they validate that it works?” And that was really a shift in my way of thinking because, yeah, backing up things can get expensive when you have multiple copies of the data living in different places, but what’s really expensive is not having a company anymore.

Sam: Yeah, yeah, absolutely. We can tie it back to my insurance dynamic earlier where, you know, it’s something that you know that you have to have, but you don’t necessarily want to pay for it. Well, you know, just like with insurances, there’s multiple different ways to go about recovering your data and it’s only in crunch time, do you really care about what it is that you’ve been paying for, right, when it comes to backup?

Could you get your backup through a git clone? Absolutely. Could you get your data back—how long is that going to take you? How painful is that going to be? What’s going to be the impact to the business where you’re trying to figure that out versus, like you say, the 39 bucks a month, a year, or whatever it might be to have something purpose-built for that, that is going to make the recovery process as quick and painless as possible and just get things back up online.

Corey: I am not a big fan of the fear, uncertainty, and doubt approach, but I do practice what I preach here in that yeah, there is a real fear against data loss. It’s not, “People are coming to get you, so you absolutely have to buy whatever it is I’m selling,” but it is something you absolutely have to think about. My core consulting proposition is that I optimize the AWS bill. And sometimes that means spending more. Okay, that one S3 bucket is extremely important to you and you say you can’t sustain the loss of it ever so one zone is not an option. Where is it being backed up? Oh, it’s not? Yeah, I suggest you spend more money and back that thing up if it’s as irreplaceable as you say. It’s about doing the right thing.

Sam: Yeah, yeah, it’s interesting, and it’s going to be hard for you to prove the value of doing that when you are driving their bill up when you’re trying to bring it down. But again, you have to look at something that’s not itemized on that bill, which is going to be the impact of downtime. I’m not going to pretend to try and recall the exact figures because it also varies depending on your business, your industry, the size, but the impact of downtime is massive financially. Tens of thousands of dollars for small organizations per hour, millions and millions of dollars per hour for much larger organizations. The backup component of that is relatively small in comparison, so having something that is purpose-built, and is going to protect your data and help mitigate that impact of downtime.

Because that’s ultimately what you’re trying to protect against. It is the recovery piece that you’re buying is the most important piece. And like you, I would say, at least be cognizant of it and evaluate your options and what can you live with and what can you live without.

Corey: That’s the big burning question that I think a lot of people do not have a good answer to. And when you don’t have an answer, you either backup everything or nothing. And I’m not a big fan of doing either of those things blindly.

Sam: Yeah, absolutely. And I think this is why we see varying different backup options as well, you know? You’re not going to try and apply the same data protection policies each and every single workload within your environment because they’ve all got different types of workload criticality. And like you say, some of them might not even need to be backed up at all, just because they don’t have data that needs to be protected. So, you need something that is going to be able to be flexible enough to apply across the entirety of your environment, protect it with the right policy, in terms of how frequently do you protect it, where do you store it, how often, or when are you eventually going to delete that and apply that on a workload by workload basis. And this is where the joy of things like tags come into play as well.

Corey: One last thing I want to bring up is that I’m a big fan of watching for companies saying the quiet part out loud. And one area in which they do this—because they’re forced to by brevity—is in the title tag of their website. I pull up veeam.com and I hover over the tab in my browser, and it says, “Veeam Software: Modern Data Protection.”

And I want to call that out because you’re not framing it as explicitly backup. So, the last topic I want to get into is the idea of security. Because I think it is not fully appreciated on a lived-experience basis—although people will of course agree to this when they’re having ivory tower whiteboard discussions—that every place your data lives is a potential for a security breach to happen. So, you want to have your data living in a bunch of places ideally, for backup and resiliency purposes. But you also want it to be completely unworkable or illegible to anyone who is not authorized to have access to it.

How do you balance those trade-offs yourself given that what you’re fundamentally saying is, “Trust us with your Holy of Holies when it comes to things that power your entire business?” I mean, I can barely get some companies to agree to show me their AWS bill, let alone this is the data that contains all of this stuff to destroy our company.

Sam: Yeah. Yeah, it’s a great question. Before I explicitly answer that piece, I will just go to say that modern data protection does absolutely have a security component to it, and I think that backup absolutely needs to be a—I’m going to say this an air quotes—a “first class citizen” of any security strategy. I think when people think about security, their mind goes to the preventative, like how do we keep these bad people out?

This is going to be a bit of the FUD that you love, but ultimately, the bad guys on the outside have an infinite number of attempts to get into your environment and only have to be right once to get in and start wreaking havoc. You on the other hand, as the good guy with your cape and whatnot, you have got to be right each and every single one of those times. And we as humans are fallible, right? None of us are perfect, and it’s incredibly difficult to defend against these ever-evolving, more complex attacks. So backup, if someone does get in, having a clean, verifiable, recoverable backup, is really going to be the only thing that is going to save your organization, should that actually happen.

And what’s key to a secure backup? I would say separation, isolation of backup data from the production data, I would say utilizing things like immutability, so in AWS, we’ve got Amazon S3 object lock, so it’s that write once, read many state for whatever retention period that you put on it. So, the data that they’re seeking to encrypt, whether it’s in production or in their backup, they cannot encrypt it. And then the other piece that I think is becoming more and more into play, and it’s almost table stakes is encryption, right? And we can utilize things like AWS KMS for that encryption.

But that’s there to help defend against the exfiltration attempts. Because these bad guys are realizing, “Hey, people aren’t paying me my ransom because they’re just recovering from a clean backup, so now I’m going to take that backup data, I’m going to leak the personally identifiable information, trade secrets, or whatever on the internet, and that’s going to put them in breach compliance and give them a hefty fine that way unless they pay me my ransom.” So encryption, so they can’t read that data. So, not only can they not change it, but they can’t read it is equally important. So, I would say those are the three big things for me on what’s needed for backup to make sure it is clean and recoverable.

Corey: I think that is one of those areas where people need to put additional levels of thought in. I think that if you have access to the production environment and have full administrative rights throughout it, you should definitionally not—at least with that account and ideally not you at all personally—have access to alter the backups. Full stop. I would say, on some level, there should not be the ability to alter backups for some particular workloads, the idea being that if you get hit with a ransomware infection, it’s pretty bad, let’s be clear, but if you can get all of your data back, it’s more of an annoyance than it is, again, the existential business crisis that becomes something that redefines you as a company if you still are a company.

Sam: Yeah. Yeah, I mean, we can turn to a number of organizations. Code Spaces always springs to mind for me, I love Code Spaces. It was kind of one of those precursors to—

Corey: It’s amazing.

Sam: Yeah, but they were running on AWS and they had everything, production and backups, all stored in one account. Got into the account. “We’re going to delete your data if you don’t pay us this ransom.” They were like, “Well, we’re not paying you the ransoms. We got backups.” Well, they deleted those, too. And, you know, unfortunately, Code Spaces isn’t around anymore. But it really kind of goes to show just the importance of at least logically separating your data across different accounts and not having that god-like access to absolutely everything.

Corey: Yeah, when you talked about Code Spaces, I was in [unintelligible 00:32:29] talking about GitHub Codespaces specifically, where they have their developer workstations in the cloud. They’re still very much around, at least last time I saw unless you know something I don’t.

Sam: Precursor to that. I can send you the link—

Corey: Oh oh—

Sam: You can share it with the listeners.

Corey: Oh, yes, please do. I’d love to see that.

Sam: Yeah. Yeah, absolutely.

Corey: And it’s been a long and strange time in this industry. Speaking of links for the show notes, I appreciate you’re spending so much time with me. Where can people go to learn more?

Sam: Yeah, absolutely. I think veeam.com is kind of the first place that people gravitate towards. Me personally, I’m kind of like a hands-on learning kind of guy, so we always make free product available.

And then you can find that on the AWS Marketplace. Simply search ‘Veeam’ through there. A number of free products; we don’t put time limits on it, we don’t put feature limitations. You can backup ten instances, including your VPCs, which we actually didn’t talk about today, but I do think is important. But I won’t waste any more time on that.

Corey: Oh, configuration of these things is critically important. If you don’t know how everything was structured and built out, you’re basically trying to re-architect from first principles based upon archaeology.

Sam: Yeah [laugh], that’s a real pain. So, we can help protect those VPCs and we actually don’t put any limitations on the number of VPCs that you can protect; it’s always free. So, if you’re going to use it for anything, use it for that. But hands-on, marketplace, if you want more documentation, want to learn more, want to speak to someone veeam.com is the place to go.

Corey: And we will, of course, include that in the show notes. Thank you so much for taking so much time to speak with me today. It’s appreciated.

Sam: Thank you, Corey, and thanks for all the listeners tuning in today.

Corey: Sam Nicholls, Director of Public Cloud at Veeam. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an angry insulting comment that takes you two hours to type out but then you lose it because you forgot to back it up.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Brian

Brian is an accomplished dealmaker with experience ranging from developer platforms to mobile services. Before InfluxData, Brian led business development at Twilio. Joining at just thirty-five employees, he built over 150 partnerships globally from the company’s infancy through its IPO in 2016. He led the company’s international expansion, hiring its first teams in Europe, Asia, and Latin America. Prior to Twilio Brian was VP of Business Development at Clearwire and held management roles at Amp’d Mobile, Kivera, and PlaceWare.

Links Referenced:

  • InfluxData: https://www.influxdata.com/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is bought to you in part by our friends at Veeam. Do you care about backups? Of course you don’t. Nobody cares about backups. Stop lying to yourselves! You care about restores, usually right after you didn’t care enough about backups. If you’re tired of the vulnerabilities, costs and slow recoveries when using snapshots to restore your data, assuming you even have them at all living in AWS-land, there is an alternative for you. Check out Veeam, thats V-E-E-A-M for secure, zero-fuss AWS backup that won't leave you high and dry when it’s time to restore. Stop taking chances with your data. Talk to Veeam. My thanks to them for sponsoring this ridiculous podcast.

Corey: This episode is brought to us by our friends at Pinecone. They believe that all anyone really wants is to be understood, and that includes your users. AI models combined with the Pinecone vector database let your applications understand and act on what your users want… without making them spell it out.
Make your search application find results by meaning instead of just keywords, your personalization system make picks based on relevance instead of just tags, and your security applications match threats by resemblance instead of just regular expressions. Pinecone provides the cloud infrastructure that makes this easy, fast, and scalable. Thanks to my friends at Pinecone for sponsoring this episode. Visit Pinecone.io to understand more.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. It’s been a year, which means it’s once again time to have a promoted guest episode brought to us by our friends at InfluxData. Joining me for a second time is Brian Mullen, CMO over at InfluxData. Brian, thank you for agreeing to do this a second time. You’re braver than most.

Brian: Thanks, Corey. I’m happy to be here. Second time is the charm.

Corey: So, it’s been an interesting year to put it mildly and I tend to have the attention span of a goldfish of most days, so for those who are similarly flighty, let’s start at the very top. What is an InfluxDB slash InfluxData slash Influx—when you’re not sure which one to use, just shorten it and call it good—and why might someone need it?

Brian: Sure. So, InfluxDB is what most people understand our product as, a pretty popular open-source product, been out for quite a while. And then our company, InfluxData is the company behind InfluxDB. And InfluxDB is where developers build IoT real-time analytics and cloud applications, typically all based on time series. It’s a time-series data platform specifically built to handle time-series data, which we think about is any type of data that is stamped in time in some way.

It could be metrics, like, taken every one second, every two seconds, every three seconds, or some kind of event that occurs and is stamped in time in some way. So, our product and platform is really specialized to handle that technical problem.

Corey: When last we spoke, I contextualized that in the realm of an IoT sensor that winds up reporting its device ID and its temperature at a given timestamp. That is sort of baseline stuff that I think aligns with what we’re talking about. But over the past year, I started to see it in a bit of a different light, specifically viewing logs as time-series data, which hadn’t occurred to me until relatively recently. And it makes perfect sense, on some level. It’s weird to contextualize what Influx does as being a logging database, but there’s absolutely no reason it couldn’t be.

Brian: Yeah, it certainly could. So typically, we see the world of time-series data in kind of two big realms. One is, as you mentioned the, you know, think of it as the hardware or, you know, physical realm: devices and sensors, these are things that are going to show up in a connected car, in a factory deployment, in renewable energy, you know, wind farm. And those are real devices and pieces of hardware that are out in the physical world, collecting data and emitting, you know, time-series every one second, or five seconds, or ten minutes, or whatever it might be.

But it also, as you mentioned, applies to, call it the virtual world, which is really all of the software and infrastructure that is being stood up to run applications and services. And so, in that world, it could be the same—it’s just a different type of source, but is really kind of the same technical problem. It’s still time-series data being stamped, you know, data being stamped every, you know, one second, every five seconds, in some cases, every millisecond, but it is coming from a source that is actually in the infrastructure. Could be, you know, virtual machines, it could be containers, it could be microservices running within those containers. And so, all of those things together, both in the physical world and this infrastructure world are all emitting time-series data.

Corey: When you take a look at the broader ecosystem, what is it that you see that has been the most misunderstood about time-series data as a whole? For example, when I saw AWS talking about a lot of things that they did in the realm of for your data lake, I talked to clients of mine about this and their response is, “Well, that’d be great genius, if we had a data lake.” It’s, “What do you think those petabytes of nonsense in S3 are?” “Oh, those are the logs and the assets and a bunch of other nonsense.” “Yeah, that’s what other people are calling a data lake.” “Oh.” Do you see similar lights-go-on moment when you talk to clients and prospective clients about what it is that they’re doing that they just hadn’t considered to be time-series data previously?

Brian: Yeah. In fact, that’s exactly what we see with many of our customers is they didn’t realize that all of a sudden, they are now handling a pretty sizable time-series workload. And if you kind of take a step back and look at a couple of pretty obvious but sometimes unrecognized trends in technology, the first is cloud applications in general are expanding, they’re both—horizontally and vertically. So, that means, like, the workloads that are being run in the Netflix’s of the world, or all the different infrastructure that’s being spun up in the cloud to run these various, you know, applications and services, those workloads are getting bigger and bigger, those companies and their subscriber bases, and the amount of data they’re generating is getting bigger and bigger. They’re also expanding horizontally by region and geography.

So Netflix, for example, running not just in the US, but in every continent and probably every cloud region around the world. So, that’s happening in the cloud world, and then also, in the IoT world, there’s this massive growth of connected devices, both net-new devices that are being developed kind of, you know, the next Peloton or the next climate control unit that goes in an apartment or house, and also these longtime legacy devices that are been on the factory floor for a couple of decades, but now are being kind of modernized and coming online. So, if you look at all of that growth of the data sources now being built up in the cloud and you look at all that growth of these connected devices, both new and existing, that are kind of coming online, there’s a huge now exponential growth in the sources of data. And all of these sources are emitting time-series data. You can just think about a connected car—not even a self-driving car, just a connected car, your everyday, kind of, 2022 model, and nearly every element of the car is emitting time-series data: its engine components, you know, your tires, like, what the climate inside of the car is, statuses of the engine itself, and it’s all doing that in real-time, so every one second, every five seconds, whatever.

So, I think in general, people just don’t realize they’re already dealing with a substantial workload of time series. And in most cases, unless they’re using something like Influx, they’re probably not, you know, especially tuned to handle it from a technology perspective.

Corey: So, it’s been a year. What has changed over on your side of the world since the last time we spoke? It seems that well, things continue and they’re up and to the right. Well, sure, generally speaking, you’re clearly still in business. Good job, always appreciative of your custom, as well as the fact that oh, good, even in a world where it seems like there’s a macro recession in progress, that there are still companies out there that continue to persist and in some cases, dare I say, even thrive? What have you folks been up to?

Brian: Yeah, it’s been a big year. So first, we’ve seen quite a bit of expansion across the use cases. So, we’ve seen even further expansion in IoT, kind of expanding into consumer, industrial, and now sustainability and clean energy, and that pairs with what we’ve seen on FinTech and cryptocurrency, gaming and entertainment applications, network telemetry, including some of the biggest names in telecom, and then a little bit more on the cloud side with cloud services, infrastructure, and dev tools and APIs. So, quite a bit more broad set of use cases we’re now seeing across the platform. And the second thing is—you might have seen it in the last month or so—is a pretty big announcement we had of our new storage engine.

So, this was just announced earlier this month in November and was previously introduced to our community as what we call an IOx, which is how it was known in the open-source. And think of this really as a rebuilt and reimagined storage engine which is built on that open-source project, InfluxDB IOx that allows us to deliver faster queries, and now—pretty exciting for the first time—unlimited time-series, or cardinality as it’s known in the space. And then also we introduced SQL for writing queries and BI tool support. And this is, for the first time we’re introducing SQL, which is world’s most popular data programming language to our platform, enabling developers to query via the API our language Flux, and InfluxQL in addition.

Corey: A long time ago, it really seems that the cloud took a vote, for lack of a better term, and decided that when it comes to storage, object store is the way forward. It was a bit of a reimagining from how we all considered using storage previously, but the economics are at minimum of ten to one in favor of objects store, the latency is far better, the durability is off the charts better, you don’t have to deal—at least in AWS-land—with the concept of availability zones and the rest, just from an economic and performance perspective, provided the use case embraces it, there’s really no substitute.

Brian: Yeah, I mean, the way we think about storage is, you know, obviously, it varies quite a bit from customer to customer with our use cases. Especially in IoT, we see some use cases where customers want to have data around for months and in some cases, years. So, it’s a pretty substantial data set you’re often looking at. And sometimes those customers want to downsample those, they don’t necessarily need every single piece of minutia that they may need in real-time, but not in summary, looking backward. So, you really—we’re in this kind of world where we’re dealing with both hive fidelity—usually in the moment—data and lower fidelity, when people can downsample and have a little bit more of a summarized view of what happened.

So, pretty unique for us and we have to kind of design the product in a way that is able to balance both of those because that’s what, you know, the customer use cases demand. It’s a super hard problem to solve. One of the reasons that you have a product like InfluxDB, which is specialized to handle this kind of thing, is so that you can actually manage that balance in your application service and setting your retention policy, et cetera.

Corey: That’s always been something that seemed a little on the odd side to me when I’m looking at a variety of different observability tools, where it seems that one of the key dimensions that they all tend to, I guess, operate on and price on is retention period. And I get it; you might not necessarily want to have your load balancer logs from 2012 readily available and paying for the privilege, but it does seem that given the dramatic fall of archival storage pricing, on some level, people do want to be able to retain that data just on the off chance that will be useful. Maybe that’s my internal digital packrat chiming in at this point, but I do believe strongly that there is a correlation between how recent the data is and how useful it is, for a variety of different use cases. But that’s also not a global truth. How do you view the divide? And what do you actually see people saying they want versus what they’re actually using?

Brian: It’s a really good question and not a simple problem to solve. So, first of all, I would say it probably really depends on the use case and the extent to which that use case is touching real world applications and services. So, in a pure observability setting where you’re looking at, perhaps more of a, kind of, operational view of infrastructure monitoring, you want to understand kind of what happened and when those tend to be a little bit more focused on real-time and recent. So, for example, you of course, want to know exactly what’s happening in the moment, zero in on whatever anomaly and kind of surrounding data there is, perhaps that means you’re digging into something that happened in you know, fairly recent time. So, those do tend to be, not all of them, but they do tend to be a little bit more real-time and recent-oriented.

I think it’s a little bit different when we look at IoT. Those generally tend to be longer timeframes that people are dealing with. Their physical out-in-the-field devices, you know, many times those devices are kind of coming online and offline, depending on the connectivity, depending on the environment, you can imagine a connected smart agriculture setup, I mean, those are a pretty wide array of devices out and in, you know, who knows what kind of climate and environment, so they tend to be a little bit longer in retention policy, kind of, being able to dig into the data, what’s happening. The time frame that people are dealing with is just, in general, much longer in some of those situations.

Corey: One story that I’ve heard a fair bit about observability data and event data is that they inevitably compose down into metrics rather than events or traces or logs, and I have a hard time getting there because I can definitely see a bunch of log entries showing the web servers return codes, okay, here’s the number of 500 errors and number of different types of successes that we wind up seeing in the app. Yeah, all right, how many per minute, per second, per hour, whatever it is that makes sense that you can look at aberrations there. But in the development process at least, I find that having detailed log messages tell me about things I didn’t see and need to understand or to continue building the dumb thing that I’m in the process of putting out. It feels like once something is productionalized and running, that its behavior is a lot more well understood, and at that point, metrics really seem to take over. How do you see it, given that you fundamentally live at that intersection where one can become the other?

Brian: Yeah, we are right at that intersection and our answer probably would be both. Metrics are super important to understand and have that regular cadence and be kind of measuring that state over time, but you can miss things depending on how frequent those metrics are coming in. And increasingly, when you have the amount of data that you’re dealing with coming from these various sources, the measurement is getting smaller and smaller. So, unless you have, you know, perfect metrics coming in every half-second, or you know, in some sub-partition of that, in milliseconds, you’re likely to miss something. And so, events are really key to understand those things that pop up and then maybe come back down and in a pure metric setting, in your regular interval, you would have just completely missed. So, we see most of our use cases that are showing a balance of the two is kind of the most effective. And from a product perspective, that’s how we think about solving the problem, addressing both.

Corey: One of the things that I struggled with is it seems that—again, my approach to this is relatively outmoded. I was a systems administrator back when that title was not considered disparaging by a good portion of the technical community the way that it is today. Even though the job is the same, we call them something different now. Great. Okay, whatever smile, nod, and accept the larger paycheck.

But my way of thinking about things are okay, you have the logs, they live on the server itself. And maybe if you want to be fancy, you wind up putting them to a centralized rsyslog cluster or whatnot. Yes, you might send them as well to some other processing system for visibility or a third-party monitoring system, but the canonical truth slash source of logs tends to live locally. That said, I got out of running production infrastructure before this idea of ephemeral containers or serverless functions really became a thing. Do you find that these days you are the source of truth slash custodian of record for these log entries, or do you find that you are more of a secondary source for better visibility and analysis, but not what they’re going to bust out when the auditor comes calling in three years?

Brian: I think, again, it—[laugh] I feel like I’m answering the same way [crosstalk 00:15:53]

Corey: Yeah, oh, and of course, let’s be clear, use cases are going to vary wildly. This is not advice on anyone’s approach to compliance and the rest [laugh]. I don’t want to get myself in trouble here.

Brian: Exactly. Well, you know, we kind of think about it in terms of profiles. And we see a couple of different profiles of customers using InfluxDB. So, the first is, and this was kind of what we saw most often early on, still see quite a bit of them is kind of more of that operator profile. And these are folks who are going to—they’re building some sort of monitor, kind of, source of truth for—that’s internally facing to monitor applications or services, perhaps that other teams within their company built.

And so that’s, kind of like, a little bit more of your kind of pure operator. Yes, they’re building up in the stack themselves, but it’s to pay attention to essentially something that another team built. And then what we’ve seen more recently, especially as we’ve moved more prominently into the cloud and offered a usage-based service with a, you know, APIs and endpoint people can hit, as we see more people come into it from a builder’s perspective. And similar in some ways, except that they’re still building kind of a, you know, a source of truth for handling this kind of data. But they’re also building the applications and services themselves are taken out to market that are in the hands of customers.

And so, it’s a little bit different mindset. Typically, there’s, you know, a little bit more comfort with using one of many services to kind of, you know, be part of the thing that they’re building. And so, we’ve seen a little bit more comfort from that type of profile, using our service running in the cloud, using the API, and not worrying too much about the kind of, you know, underlying setup of the implementation.

Corey: Love how serverless helps you scale big and ship fast, but hate debugging your serverless apps? With Lumigo’s serverless observability, it’s fast and easy (and maybe a little fun, too). End-to-end distributed tracing gives developers full clarity into their most complex serverless and containerized applications, connecting every service from AWS Lambda and Amazon ECS to DynamoDB, API Gateways, Step Functions and more. Try Lumigo free and debug 3x faster, reduce error rate and speed up development. Visit snark.cloud/lumigo That’s snark.cloud/L-U-M-I-G-O

Corey: So, I’ve been on record a lot saying that the best database is TXT records stuffed into Route 53, which works super well as a gag, let's be clear, don’t actually build something on top of this, that’s a disaster waiting to happen. I don’t want to destroy anyone’s career as I do this. But you do have a much more viable competitive threat on the landscape. And that is quite simply using the open-source version of InfluxDB. What is the tipping point where, “Huh, I can run this myself,” turns into, “But I shouldn’t. I should instead give money to other people to run it for me.”

Because having been an engineer, where I believe I’m the world’s greatest everything, when it comes to my environment—a fact provably untrue, but that hubris never quite goes away entirely—at what point am I basically being negligent not to start dealing with you in a more formalized business context?

Brian: First of all, let me say that we have many customers, many developers out there who are running open-source and it works perfectly for them. The workload is just right, the deployment makes sense. And so, there are many production workloads we’re using open-source. But typically, the kind of big turning point for people is on scale, scale, and overall performance related to that. And so, that’s typically when they come and look at one of the two commercial offers.

So, to start, open-source is a great place to, you know, kind of begin the journey, check it out, do that level of experimentation and kind of proof of concept. We also have 60,000-plus developers using our introductory cloud service, which is a free service. You simply sign up and can begin immediately putting data into the platform and building queries, and you don’t have to worry about any of the setup and running servers to deploy software. So, both of those, the open-source and our cloud product are excellent ways to get started. And then when it comes time to really think about building in production and moving up in scale, we have our two commercial offers.

And the first of those is InfluxDB Cloud, which is our cloud-native fully managed by InfluxData offering. We run this not only in AWS but also in Google Cloud and Microsoft Azure. It’s a usage-based service, which means you pay exactly for what you use, and the three components that people pay for our data in, number of queries, and the amount of data you store in storage. We also for those who are interested in actually managing it themselves, we have InfluxDB Enterprise, which is a software subscription-base model, and it is self-managed by the customer in their environment. Now, that environment could be their own private cloud, it also could be on-premises in their own data center.

And so, lots of fun people who are a little bit more oriented to kind of manage software themselves rather than using a service gear toward that. But both those commercial offers InfluxDB Cloud and InfluxDB Enterprise are really designed for, you know, massive scale. In the case of Cloud, I mentioned earlier with the new storage engine, you can hit unlimited cardinality, which means you have no limit on the number of time series you can put into the platform, which is a pretty big game-changing concept. And so, that means however many time-series sources you have and however many series they’re emitting, you can run that without a problem without any sort of upper limit in our cloud product. Over on the enterprise side with our self-managed product, that means you can deploy a cluster of whatever size you want. It could be a two-by-four, it could be a four-by-eight, or something even larger. And so, it gives people that are managing in their own private cloud or in a data center environment, really their own options to kind of construct exactly what they need for their particular use case.

Corey: Does your object storage layer make it easier to dynamically change clusters on the fly? I mean, historically, running things in a pre-provisioned cluster with EBS volumes or local disk was, “Oh, great. You want to resize something? Well, we’re going to be either taking an outage or we’re going to be building up something, migrating data live, and there’s going to be a knife-switch cutover at some point that makes things relatively unfortunate.” It seems that once you abstract the storage layer away from anything resembling an instance that you would be able to get away from some of those architectural constraints.

Brian: Yeah, that’s really the promise, and what is delivered in our cloud product is that you no longer, as a developer, have to think about that if you’re using that product. You don’t have to think about how big the cluster is going to be, you don’t have to think about these kind of disaster scenarios. It is all kind of pre-architected in the service. And so, the things that we really want to deliver to people, in addition to the elimination of that concern for what the underlying infrastructure looks like and how its operating. And so, with infrastructure concerns kind of out of the way, what we want to deliver on are kind of the things that people care most about: real-time query speed.

So, now with this new storage engine, you can query data across any time series within milliseconds, 100 times faster queries against high cardinality data that was previously impossible. And we also have unlimited time-series volume. Again, any total number of time series you have, which is known as cardinality, is now able to run without a problem in the platform. And then we also have kind of opening up, we’re opening up the aperture a bit for developers with SQL language support. And so, this is just a whole new world of flexibility for developers to begin building on the platform. And again, this is all in the way that people are using the product without having to worry about the underlying infrastructure.

Corey: For most companies—and this does not apply to you—their core competency is not running time-series databases and the infrastructure attendant thereof, so it seems like it is absolutely a great candidate for, “You know, we really could make this someone else’s problem and let us instead focus on the differentiated thing that we are doing or building or complaining about.”

Brian: Yeah, that’s a true statement. Typically what happens with time-series data is that people first of all, don’t realize they have it, and then when they realize they have time-series data, you know, the first thing they’ll do is look around and say, “Well, what do I have here?” You know, I have this relational database over here or this document database over here, maybe even this, kind of, search database over here, maybe that thing can handle time series. And in a light manner, it probably does the job. But like I said, the sources of data and just the volume of time series is expanding, really across all these different use cases, exponentially.

And so, pretty quickly, people realize that thing that may be able to handle time series in some minor manner, is quickly no longer able to do it. They’re just not purpose-built for it. And so, that’s where really they come to a product like Influx to really handle this specific problem. We’re built specifically for this purpose and so as the time-series workload expands when it kind of hits that tipping point, you really need a specialized tool.

Corey: Last question, before I turn you loose to prepare for re:Invent, of course—well, I guess we’ll ask you a little bit about that afterwards, but first, we can talk a lot theoretically about what your product could or might theoretically do. What are you actually seeing? What are the use cases that other than the stereotypical ones we’ve talked about, what have you seen people using it for that surprised you?

Brian: Yeah, some of it is—it’s just really interesting how it connects to, you know, things you see every day and/or use every day. I mean, chances are, many people listening have probably use InfluxDB and, you know, perhaps didn’t know it. You know, if anyone has been to a home that has Tesla Powerwalls—Tesla is a customer of ours—then they’ve seen InfluxDB in action. Tesla’s pulling time-series data from these connected Powerwalls that are in solar-powered homes, and they monitor things like health and availability and performance of those solar panels and the battery setup, et cetera. And they’re collecting this at the edge and then sending that back into the hub where InfluxDB is running on their back end.

So, if you’ve ever seen this deployed like that’s InfluxDB running behind the scenes. Same goes, I’m sure many people have a Nest thermostat in their house. Nest monitors the infrastructure, actually the powers that collection of IoT data collection. So, you think of this as InfluxDB running behind the scenes to monitor what infrastructure is standing up that back-end Nest service. And this includes their use of Kubernetes and other software infrastructure that’s run in their platform for collection, managing, transforming, and analyzing all of this aggregate device data that’s out there.

Another one, especially for those of us that streamed our minds out during the pandemic, Disney+ entertainment, streaming, and delivery of that to applications and to devices in the home. And so, you know, this hugely popular Disney+ streaming service is essentially a global content delivery network for distributing all these, you know, movies and video series to all the users worldwide. And they monitor the movement and performance of that video content through this global CDN using InfluxDB. So, those are a few where you probably walk by something like this multiple times a week, or in our case of Disney+ probably watching it once a day. And it’s great to see InfluxDB kind of working behind the scenes there.

Corey: It’s one of those things where it’s, I guess we’ll call it plumbing, for lack of a better term. It’s not the sort of thing that people are going to put front-and-center into any product or service that they wind up providing, you know, except for you folks. Instead, it’s the thing that empowers a capability behind that product or service that is often taken for granted, just because until you understand the dizzying complexity, particularly at scale, of what these things have to do under the hood, it just—well yeah, of course, it works that way. Why shouldn’t it? That’s an expectation I have of the product because it’s always had that. Yeah, but this is how it gets there.

Brian: Our thesis really is that data is best understood through the lens of time. And as this data is expanding exponentially, time becomes increasingly the, kind of, common element, the common component that you’re using to kind of view what happened. That could be what’s running through a telecom network, what’s happening with the devices that are connected that network, the movement of data through that network, and when, what’s happening with subscribers and content pushing through a CDN on a streaming service, what’s happening with climate and home data in hundreds of thousands, if not millions of homes through common device like a Nest thermostat. All of these things they attach to some real-world collection of data, and as long as that’s happening, there’s going to be a place for time-series data and tools that are optimized to handle it.

Corey: So, my last question—for real this time—we are recording this the week before re:Invent 2022. What do you hope to see, what do you expect to see, what do you fear to see?

Brian: No fears. Even though it’s Vegas, no fears.

Corey: I do have the super-spreader event fear, but that’s a separate—

Brian: [laugh].

Corey: That’s a separate issue. Neither one of us are deep into the epidemiology weeds, to my understanding. But yeah, let’s just bound this to tech, let’s be clear.

Brian: Yeah, so first of all, we’re really excited to go there. We’ll have a pretty big presence. We have a few different locations where you can meet us. We’ll have a booth on the main show floor, we’ll be in the marketplace pavilion, as I mentioned, InfluxDB Cloud is offered across the marketplaces of each of the clouds, AWS, obviously in this case, but also in Azure and Google. But we’ll be there in the AWS Marketplace pavilion, showcasing the new engine and a lot of the pretty exciting new use cases that we’ve been seeing.

And we’ll have our full team there, so if you’re looking to kind of learn more about InfluxDB, or you’ve checked it out recently and want to understand kind of what the new capability is, we’ll have many folks from our technical teams there, from our development team, some our field folks like the SEs and some of the product managers will be there as well. So, we’ll have a pretty great collection of experts on InfluxDB to answer any questions and walk people through, you know, demonstrations and use cases.

Corey: I look forward to it. I will be doing my traditional Wednesday afternoon tour through the expo halls and nature walk, so if you’re listening to this and it’s before Wednesday afternoon, come and find me. I am kicking off and ending at the [unintelligible 00:29:15] booth, but I will make it a point to come by the Influx booth and give you folks a hard time because that’s what I do.

Brian: We love it. Please. You know, being on the tour is—on the walking tour is excellent. We’ll be mentally prepared. We’ll have some comebacks ready for you.

Corey: Therapists are standing by on both sides.

Brian: Yes, exactly. Anyway, we’re really looking forward to it. This will be my third year on your walking tour. So, the nature walk is one of my favorite parts of AWS re:Invent.

Corey: Well, I appreciate that. Thank you. And thank you for your time today. I will let you get back to your no doubt frenzied preparations. At least they are on my side.

Brian: We will. Thanks so much for having me and really excited to do it.

Corey: Brian Mullen, CMO at InfluxData, I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an insulting comment that you naively believe will be stored as a TXT record in a DNS server somewhere rather than what is almost certainly a time-series database.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Perry

Perry Krug currently leads the Shared Services team which is focused on building tools and managing infrastructure and data to increase the productivity of Couchbase’s Sales and Field organisations. Perry has been with Couchbase for over 12 years and has served in many customer-facing technical roles, helping hundreds of customers understand, deploy, and maintain Couchbase's NoSQL database technology. He has been working with high performance caching and database systems for over 15 years.

Links Referenced:

  • Couchbase: https://www.couchbase.com/
  • Perry’s LinkedIn: https://www.linkedin.com/in/perrykrug/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is brought to us by our friends at Pinecone. They believe that all anyone really wants is to be understood, and that includes your users. AI models combined with the Pinecone vector database let your applications understand and act on what your users want… without making them spell it out. Make your search application find results by meaning instead of just keywords, your personalization system make picks based on relevance instead of just tags, and your security applications match threats by resemblance instead of just regular expressions. Pinecone provides the cloud infrastructure that makes this easy, fast, and scalable. Thanks to my friends at Pinecone for sponsoring this episode. Visit Pinecone.io to understand more.

Corey: InfluxDB is the smart data platform for time series. It’s built from the ground-up to handle the massive volumes and countless sources of time-stamped data produced by sensors, applications, and systems. You probably think of these as logs.
InfluxDB is programmable and performant, has a common API across the platform, and handles high granularity data–at scale and with high availability. Use InfluxDB to build real-time applications for analytics, IoT, and cloud-native services, all in less time and with less code. So go ahead–turn your apps up to 11 and start your journey to Awesome for free at InfluxData.com/screaminginthecloud

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Today’s episode is a promoted guest episode brought to us by our friends at Couchbase. Now, I want to start off by saying that this week is AWS re:Invent. And there is Last Week in AWS swag available at their booth. More on that to come throughout the next half hour or so of conversation. But let’s get right into it. My guest today is Perry Krug, Director of Shared Services over at Couchbase. Perry, thanks for joining me.

Perry: Hey, Corey, thank you so much for having me. It’s a pleasure.

Corey: So, we’re recording this before re:Invent, so the fact that we both have, you know, personality and haven’t lost our voices yet should probably be a bit of a giveaway on this. But I want to start at the very beginning because unlike people who are academically successful, I tend to suck at doing the homework, across the board. Couchbase has been around for a long time. We’ve seen the company do a bunch of different things, most importantly and notably, sponsoring my ridiculous nonsense for which I thank you. But let’s start at the beginning. What is Couchbase?

Perry: Yeah, you’re very welcome, Corey. And it’s again, it’s a pleasure to be here. So, Couchbase is an enterprise database company at the very top level. We make database software and we distribute that to our customers. We have two flavors, two ways of getting your hands on it.

One is the kind of legacy, what we call self-managed, where you the user, the customer, downloads the software, installs it themselves, sets it up, manages the cluster monitoring, scaling all of that. And that’s, you know, a big part of our business. Over the last few years we’ve identified, and certainly others in the industry have, as well the desire for users to access database and other technology in a hosted Software-as-a-Service pay-as-you-go, cloud-native, buzzword, et cetera, et cetera, vehicle. And so, we’ve released the Couchbase Capella, which is our fully managed, fully hosted database-as-a-service, running in—currently—Amazon and Google, soon to be Azure as well. And it wraps and extends our core Couchbase Server product into a, as I mentioned, hosted and managed platform that our users can now come to and consume as developers and build their applications while leaving all of the operational and administration—monitoring, managing failover expansion, all of that—to us as the experts.

Corey: So, you folks are non-relational database, NoSQL in the common parlance, which is odd because they call it NoSQL, yet. They keep making more of them, so I feel like that’s sort of the Hollywood model where okay, that was so good. We’re going to do it again. Where did NoSQL come from? Because back when I was learning databases, when dinosaurs roamed the earth, it was all about relational models, like we’re going to use a relational database because when the only tool you have is an axe, every problem looks like hours of fun. What gave rise to this, I guess, Cambrian explosion that we’ve seen of NoSQL options that proliferate o’er the land?

Perry: Yeah, a really, really good question, and I like the axe-throwing metaphor. So sure, 20, 30, 40 now years ago, as digital applications needed a place to store their data, the world invented relational databases. And those were used and continue to be used very well for what they were designed for, for data that follows a very strict structure that doesn’t need to be served at significant scale, does not need to be replicated geographically, does not need to handle data coming in from different sources and those sources changing their formats of things all the time. And so, I’m probably as old as you are and been around when the dinosaurs were there. We remember this term called ‘Web 2.0.’ Kids, you’re going to have to go look that up in the dictionary or TikTok it or something.

But Web 2.0 really was the turning point when websites became web applications. And suddenly, there was the introduction of MySpace and Facebook and Amazon and Google and LinkedIn, and a number of others, and they realized that relational databases we’re not going to meet their needs, whether it be performance, whether it be flexibility, whether it be changing of data models, whether it be introducing new features at a rapid pace. They tried; they stretched them, they added a bunch of different databases together, and really was not going to be a viable solution. So, 10 now, maybe 15 years ago, you started to see the rise of these tech giants—although we didn’t call them tech giants back then but they were the precursors to today’s—invent their own new databases.

So, Amazon had theirs, Google has theirs, LinkedIn, and a number of others. These companies had reached a level of scale and reached a level of user base, had reached a level of data requirement, had reached a level of expectation with their customers. These customers, us, the users, us consumers, we expect things to be fast, we expect them to be always available. We expect Facebook to give us our news feed in milliseconds. We expect Google to give us our website or our search results in immediate, with more and more information coming along with them.

And so, it was these companies that hit those requirements first. The only solution for them was to start from scratch and rewrite their own databases. Fast forward five, six, seven years, and we as consumers turned around and said, “Look, I really liked the way Facebook does things. I really like the way Google does things. I really like the way Amazon does things.

“Bank of America, can you do the same? IRS, can you do the same? Health care vendor number one, two, three, and four, government body, can you all give me the same experience? I want my taxi to tell me exactly where it’s going to take me from one place to another, I want it to give me a receipt immediately after I finish my ride. Actually, I want to be able to change my payment method after I paid for that ride because I used the wrong one.”

All of these are expectations that we as consumers have taken from the tech giants—Apple, LinkedIn, Facebook—and turned around to nearly every other service that we interact with on a daily basis. And all of a sudden, the requirements that Facebook had, that Google had, that no other company had, you know, outside of the top five, suddenly were needed by every single industry, nearly every single company, in order to be competitive in their markets.

Corey: And there’s no way to scale relational to get to a point where it can wind up handling those type workloads efficiently?

Perry: Correct, correct. And it’s not just that the technology cannot do it—everything is technically feasible—but the cost both financially and time-to-market-wise in order to do that in a relational database was untenable. It either cost too much money, or it costs too much developers time, or cost too much of everybody’s time to try to shoehorn something into it. And then you have the rise of cloud and containers, which relational databases, you know, never even had the inkling of a thought that they might need to be able to handle someday. And so, these requirements that consumers have been placed on everything else that they interact with really led to the rise of NoSQL as a commodity or as a database for the masses.

LinkedIn is not in the business of developing a database and then selling it to everybody else to use as a database, right? They built it for themselves, they made their service better. And so, what you see is some of those founding fathers created databases, but then had no desire to sell them to others. And then after that followed the rise of companies like Couchbase and a number of others who said, “Look, we think we can provide those capabilities, we think we can meet those requirements for everybody.” And thereby rose the plethora of NoSQL databases because everybody had a little bit different of an approach to it.

If you ask ten people what NoSQL is about, you’re going to get eleven or twelve different answers. But you can kind of distill that into two categories. One is performance and operations. So, I need it to be faster, I need it to be scalable, I need it to be replicated geographically. And that’s what NoSQL is to me. And that’s the right answer.

And so, you have things like Cassandra and Redis that are meant to be fast and scalable and replicated. You ask another group and they’re going to tell you, “No, no, no. NoSQL needs to be flexible. I need to get rid of the rigid database schemas, I need to bring JSON or other data formats in and munge all this data together and create something cool and new out of it.” And thereby you have the rise of things like MongoDB, who focused nearly exclusively on the developer experience of working with data.

And for a long time, those two were in opposite camps, where you have the databases that did performance and the databases that did flexibility. I’m not here to say that Couchbase is the ultimate kitchen sink for everything, but we’ve certainly tried to approach both of those challenges together so that you can have something that scales and performs and can be flexible enough in data model. And everybody else is trying to do the same thing, right? But all these databases are competing for that same nirvana of the best of both worlds.

Corey: And it almost feels like there’s a convergence play in place where everything now is trying to go away from the idea of, “Oh, yeah, we started off as a purpose-built database, but you can use this for everything.” And I don’t necessarily know that is going to be the path that a lot of companies want to go down. What do you view Couchbase as I guess, falling down? In other words, what workloads is Couchbase inappropriate for?

Perry: Yeah, that’s a good question. And my [crosstalk 00:10:35]—

Corey: Anyone who can’t answer that one is a zealot and that’s one of those okay, let’s be very careful and not take our eyes off you for one second, while smiling and backing away slowly.

Perry: Let’s cut to commercial. No, I mean, there certainly are workloads that you know, in the past, we’ve not been good for that we’ve made improvements to address. There are workloads that we had not address well today that we will try to address in the future, and there are workloads that we may never see as fitting in our wheelhouse. The biggest category group that comes to mind is Couchbase is not an archival database. We are not meant to have data put in us that you don’t care about, that you don’t want to—that you just need to keep it around, but you don’t ever need to access.

And there are systems that do that well, they do that at a solid total cost of ownership. And Couchbase is meant for operational data. It’s meant for data that needs to be interacted with, read and/or written, at scale and at a reasonable performance to serve a user-facing or system-facing application. And we call ourselves a general-purpose database. Bongo and others call themselves as well. Oracle calls itself a general-purpose database, and yet, not everybody uses Oracle for everything.

So, there are reasons that you—

Corey: Who could afford that?

Perry: Who could? Exactly. It comes down to cost, ultimately. So, I’m not here to say that Couchbase does everything. We like to think, and we’re trying to target and strive towards an 80%, right? If we can do 80% of an application or an organization’s workloads, there is certainly room for 20% of other workloads, other applications, other requirements that can be met or need to be met by purpose-built databases.

But if you rewind four or five years, there was this big push towards polyglot persistence. It’s a buzzword that came and kind of has gone out of fashion, but it presented the idea that everybody is going to use 15 different databases and everybody is going to pick the right one for exactly the workload and they’re going to somehow stitch them all together. And that really hasn’t come to fruition either. So, I think there’s some balance, where it’s not one to rule them all, but it’s also not 15 for every company. Some organizations just have a set of requirements that they want to be met and our database can do that.

Corey: Let’s continue our tour of the competitive landscape here now that we’ve handled the relational side of the world. The best database, as anyone who’s listened to this show knows, is of course, Amazon’s Route 53 TXT records stuffed into DNS, especially in the NoSQL land. Clearly, you’re all fighting for second place after that. How do you stack up against the idea of legitimately using that approach? And for those who are not in on the joke, please don’t do this. It is not the right answer. But I’m curious to get your take as to why DNS TXT records are an inappropriate NoSQL option.

Perry: Well, it’s a joke, right? And let’s be clear about that. But—

Corey: I have to say that because otherwise, someone tries it in production. I’ve gotten that wrong a few times, historically, so now I put a disclaimer in because yeah, it’s only funny, so long as people are in on the joke. If not, and I lead someone down the primrose path to disaster, I feel bad. So, let’s be very clear. We’re kidding.

Perry: And I’m laughing. I’m laughing here behind the camera. I am. I am.

Corey: Yeah.

Perry: So, the element of truth that I think Couchbase is in a position, or I’m in a position to kind of talk about is, 12 years ago, when Couchbase started, we were a key-value database and that’s where we saw the best part of the market in those days, and where we were able to achieve the best scale and replication and performance, and fairly quickly realized that simple key-value, though extremely valuable and easy to manage, was not broad enough in requirements-meeting. And that’s where we set our sights on and identified the larger, kind of, document database group, which is really just a derivative of key-value, where still everything is a key and a value; it’s just now a document that you can reason about, that you can create an index on, that you can query, that you can run full-text search on, you can do much more with the data. So, at our core, we are still a key-value database. When that value is JSON, we become a document database. And so, if Route 53 decided that they wanted to enter into the document database market, they would need to be adding things that allowed you to introspect and ask questions of the data within that text which you can’t, right?

Corey: Well, not with that attitude. But yeah, I agree with you.

Perry: [laugh].

Corey: Moving up the stack, let’s talk about a much more fearsome competitor here that I’m certain you see an awful lot of deals that you wind up closing, specifically your own open-source product. You historically have wound up selling software into environments, I believe, you referred to as your legacy offering where it’s the hosted version of your commercial software. And now of course, you also have Capella, your cloud-hosted version. But open-source looks surprisingly compelling for an awful lot of use cases and an awful lot of folks. What’s the distinction?

Perry: Sure. Just to correct a little bit the distinction, we have Couchbase Server, which we provide as a what we call self-managed, where you can download it and install it yourself. Now, you could do that with the open-source version or you could do that with our Enterprise Edition. What we’ve then done is wrapped that Enterprise Edition in a hosted bottle, and that’s Capella. So, the open-source version is something we’ve long been supporters of; it’s been a core part of our go-to-market for the last 12 or 13 years or so and we still see it as a strong offering for organizations that don’t need the added features, the added capabilities, don’t need the support of the experts that wrote the software behind them.

Certainly, we contribute and support our community through our forums and Discord and other channels, but that’s a very big difference than two o’clock in the morning, something’s not working and I need a ticket to track. We don’t do that for our community edition. So, we see lots of users downloading that, picking it up building it into their applications, especially applications that are in their infancy or are with organizations that they simply can’t afford the added cost and therefore they don’t get the added benefit. We’re not here to gouge and carve out every dollar that we can, but if you need the benefit that we can provide, we think there’s value in that and that’s what we’re trying to run a business as.

Corey: Oh, absolutely. It doesn’t work when you’re trying to wind up charging a license fee for something that someone is doing in their spare time project for funsies just to learn the technology. It’s like, and then you show up. It’s like, “That’ll be $700. Surprise.”

Yeah, that’s sort of the AWS billing model approach, where—it’s not a viable onramp for most folks. So, the open-source direction down there make sense. Counterpoint. If you’re running a bank on top of it, “Well, we’re running it ourselves and really hoping for the best. I mean, we have access to the code and all.” Great, but there are times you absolutely want some of the best minds in the world, with respect to that particular product, able to help troubleshoot so the ATM start working again before people riot in the streets.

Perry: Yeah, yeah. And ultimately, it’s a question of core competencies. Are you an organization that wants to be in the database development market? Great, by all means, we’d love to support you in that. If you want to focus on doing what you do best be at a bank or an e-commerce website, you worry about your application, you let us worry about the database and everybody gets along very well.

Corey: There’s definitely something to be said for outsourcing some of the pain, some of the challenge around an awful lot of it.

Perry: There’s a natural progression to the cloud for that and Software-as-a-Service, database-as-a-service where you’re now outsourcing even more by running on our hosting platform. No longer do you have to download the binary and install yourself, no longer do you have to setup the cluster and watch it in case it has a blip or the statistic goes up too far. We’re taking care of that for you. So yes, you’re paying for that service, but you’re getting the value of not having to be a database manager, let alone database developer for them.

Corey: Love how serverless helps you scale big and ship fast, but hate debugging your serverless apps? With Lumigo’s serverless observability, it’s fast and easy (and maybe a little fun, too). End-to-end distributed tracing gives developers full clarity into their most complex serverless and containerized applications, connecting every service from AWS Lambda and Amazon ECS to DynamoDB, API Gateways, Step Functions and more. Try Lumigo free and debug 3x faster, reduce error rate and speed up development. Visit snark.cloud/lumigo That’s snark.cloud/L-U-M-I-G-O

Corey: What is the point of distinction between Couchbase Server and Couchbase Capella? To be clear, your self-hosted versus managed cloud offerings. When is one appropriate versus the other?

Perry: Well, I’m supposed to say that Capella is always the appropriate choice, but there are currently a number of situations where Capella is not available in particular regions or cloud providers and so downloading running the software yourself certainly in your own—yes, there are people who still run their own data centers. I know it’s taboo and we don’t like to talk about that, but there are people who have on-premise. And so, Couchbase Capella is not available for them. But Couchbase Server is the original Couchbase database and it is the core of Couchbase Capella. So, wrapping is not giving it enough credit; we use Couchbase Server to power Couchbase Capella.

And so, there’s an enormous amount of value added around the core database, but ultimately, it’s the behind the scenes of Couchbase Capella. Which I think is a nice benefit in that when an application is connecting to either one, it gets the same experience. You can point an application at one versus the other and because it’s the same database running behind the scenes, the behavior, the data model, the query language, the APIs are all the same, so it adds a nice level of flexibility four customers that are either moving from one to another or have to have some sort of hybrid approach, which we see in the market today.

Corey: Let’s talk economics for a second. I can see scenarios where especially you have a high volume environment where you’re sending tremendous amounts of data back and forth and as soon as it crosses an availability zone boundary or a region boundary, or God forbid, goes out to the internet via standard egress fees over in AWS-land, there’s a radically different economic modeling that comes into play as opposed to having something in the same availability zone, in the same subnet just where that—or all traffic back and forth is free. Do you see that in your customer base, that that is a model that is driving people towards self-hosting?

Perry: No. And I’d say no because Capella allows you to peer and run your application in the same availability zone as the as a database. And so, as long as that’s an option for you that we have, you know, our offering in the right region, in the right AZ, and you can put your application there, then that’s not a not an issue. We did have a customer not too long ago that didn’t set that up correctly, they thought they did, and we noticed some high data transfer charges. Again, the benefit of running a hosted service, we detected that for them and were able to turn around and say, “Hmm, you might want to change this to over there so that we all save some money in doing so.”

If we were not there watching it, they might not have noticed that themselves if they were running it self-managed; they might not have known what to do about it. And so, there’s a benefit to working with us and using that hosted platform that we can keep an eye out. And we can apply all of our learning and best practices and bug fixes, we give that to everybody, rather than each person having to stumble across those hurdles themselves.

Corey: That’s one of those fun, weird corner-case trivia things about AWS data transfer. When you’re transferring data within the same region between availability zones, it costs a penny on the sending side and a penny on the receiving side. Everything else is one side or the other that winds up getting the charge. And what makes this especially fun is that when it shows up on your bill, if you transfer a petabyte, it shows as cross-AZ data transfer: two petabytes.

Perry: Two. Yeah.

Corey: So, it double-counts so they can bill for it appropriately, but it leads to some really weird hunting it down, like, “Okay, well, we found half of it, but where’s the other half hiding?” It’s always obnoxious to trace this stuff down. The fact that you see it on your bill, well, that’s testament to the fact that yeah, they’re using the service. Good for them and good for you. Being able to track it down on a per-customer basis that does speak to your level of insight into what exactly is going on in your environment and where. As someone who does this for a living, let me confirm that is absolutely non-trivial.

Perry: No, definitely not trivial. And you know, we’ve learned over the last four or five years, we’ve learned an enormous amount about how cloud providers work, how AWS works, but guess what, Google does it completely differently. And Azure does it—

Corey: Yep.

Perry: —completely differently. And so, on the surface level, they’re all just cloud providers and they give you a VM, and you put some stuff on it, but integrating with the APIs, integrating with the different systems and naming of things, and then understanding the intricacies of the ins and outs, and, yeah, these cloud providers have their own bugs as well. And so, sometimes you stumble across that for them. And it’s been a significant learning exercise that I think we’re all better off for, having Couchbase gone through it for you.

Corey: Let’s get this a little bit more germane for this week for those of you who are listening to this during re:Invent. You folks are clearly here at the show—it’s funny to talk about ‘here,’ even though when we’re recording this, it is not near here; we’re actually home and enjoying ourselves, but welcome to temporal dislocation; here we are—here at the show, you folks are—among other things—being kind enough to pass out the Last Week in AWS swag from your booth, which, thank you. So, that is obviously the primary reason that you were at the show. What are the other reasons? What are the secondary reasons that you decided to come here?

Perry: Yeah [laugh]. Well, I guess I have to think about this now since you already called out the primary reason.

Corey: Exactly. Wait, we can have more than one reason for things? My God.

Perry: Can we? Can we? AWS has long been a huge partner of ours, even before Capella itself was released. I remember sometime in, you know, five years or so ago, some 30% of our customers were running Couchbase inside of AWS, and some of our largest were some of your largest at times, like Viber, the messaging platform. And so, we’ve always had a very strong relationship with AWS, and the better that we can be presenting ourselves to your customers, and our customers can feel that we are jointly supporting them, the better. And so, you know, coming to re:Invent is a testament to that long-standing and very solid partnership, and also it’s meant to get more exposure for us to let it be clear that Couchbase runs very well on AWS.

Corey: It’s one of those areas where when someone says, “Oh yeah, this is a great service offering, but it doesn’t run super well on AWS.” It’s like, “Okay, so are you bad computers or is what you have built so broken and Byzantine that it has to live somewhere else?” Or occasionally, the use case is absolutely not supported by AWS. Not to beat them up some more on their egress fees, but I’m absolutely about to if you’re building a video streaming site, you don’t want it living in AWS. It won’t run super well there. Well, it’ll run well, it’ll just run extortionately expensively and that means that it’s a non-starter.

Perry: Yeah, why do you think Netflix raises their fees?

Corey: Netflix, to their credit, has been really rather public about this, where they do all of their egress via their Open Connect, custom-built CDN appliances that they drop all over the place. They don’t stream a single byte from AWS, and we know this from the outside because they are clearly still solvent.

Perry: [laugh].

Corey: I do the math on that. So, if I had been streaming at on-demand prices one month with my Netflix usage, I would have wound up spending four times my subscription fee just in their raw costs for data transfer. And I have it on good authority that is not just data transfer that is their only bill in the entire company; they also have to pay people and content and the analytics engine and whatnot. And it’s kind of a weird, strange world.

Perry: Real estate.

Corey: Yeah. Because it’s one of those strange stories because they are absolutely a showcase customer for AWS. They’ve been a marquee customer trotted out year after year to talk about what they’re doing. But if you attempt to replicate their business purely on top of AWS, it will not work. Full stop. The economics preclude that happening.

What is your philosophy these days on what historically has felt like an existential threat to most vendors that I’ve spoken to in a variety of ways: what if Amazon decides to enter your market? I’d ask you the same thing. Do you have fears that they’re going to wind up effectively taking your open-source offering and turning it into Amazon Basics Couchbase, for lack of a better term? Is that something that is on your threat radar, or is that not really something you concern yourselves about?

Perry: So, I mean, there’s no arguing, there’s no illusion that Amazon and Google and Microsoft are significant competitors in the database space, along with Oracle and IBM and Mongo and a handful of others.

Corey: Anything’s a database if you hold it wrong.

Perry: This is true. This specific point of open-source is something that we have addressed in the same ways that others have addressed. And that’s by choosing and changing our license model so that it precludes cloud providers from using the open-source database to produce their own service on the back of it. Let me be clear, it does not impact our existing open-source users and anybody that wants to use the Community Edition or download the software, the source code, and build it themselves. It’s only targeted at Amazon because they have a track record of doing that to things like Elastic and Redis and Mongo, all of whom who have made similar to Couchbase moves to prevent that by the licensing of the open-source code.

Corey: So, one of the things I do see at re:Invent every year is—and I believe wholeheartedly this comes historically from a lot of AWS’s requirements for vendors on the show floor that have become public through a variety of different ways—where you for a long time, you are not allowed to mention multi-cloud or reference the fact that you work on any other cloud provider there. So, there’s been a theme of this is why, for whatever it is we sell or claim to sell or hope one day to sell, AWS is the absolute best place for you to run it, full stop. And in some cases, that’s absolutely true because people build primarily for a certain cloud provider and then when they find customers and other places, they learn to run it over there, too. If I’m approaching this from the perspective of I have a database problem—because looking at my philosophy on databases is hard to imagine I don’t have database problems—then is my experience going to be better or even materially different between any of the cloud providers if I become a Couchbase Capella customer?

Perry: I’d like to say no. We’ve done our best to abstract and to leverage the best of all of the cloud providers underneath to provide Couchbase in the best form that they will allow us to. And as far as I can see, there’s no difference amongst those. Your application and what you do with the data, that may be better suited to one provider or another, but it’s always been Couchbase is philosophy—sort of say, strategy—to make our software available to wherever our customers and users want to, to consume it. And that goes everything from physical hardware running in a data center, virtual machines on top of that, containers, cloud, and different cloud providers, different regions, different availability zones, all the way through to edge and other infrastructures. We’re not in a position to say, “If you want Couchbase, you should use AWS.” We’re in a position to say, “If you are using AWS, you can have Couchbase.”

Corey: I really want to thank you for being so generous with your time, and of course, your sponsorship dollars, which are deeply appreciated. Once again, swag is available at the Couchbase booth this week at re:Invent. If people want to learn more and if for some unfathomable reason, they’re not at re:Invent, probably because they make good life choices, where can they go to find you?

Perry: couchbase.com. That’ll to be the best place to land on. That takes you to our documentation, our resources, our getting help, our contact pages, directly into Capella if you want to sign in or login. I would go there.

Corey: And we will, of course, put links to that in the show notes. Thank you so much for your time. I really appreciate it.

Perry: Corey, it’s been a pleasure. Thank you for your questions and banter, and I really appreciate the opportunity to come and share some time with you.

Corey: We’ll have to have you back in the near future. Perry Krug, Director of Shared Services at Couchbase. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an angry and insulting comment berating me for being nowhere near musical enough when referencing [singing] Couchbase Capella.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Andi

Andi Gutmans is the General Manager and Vice President for Databases at Google. Andi’s focus is on building, managing and scaling the most innovative database services to deliver the industry’s leading data platform for businesses.

Before joining Google, Andi was VP Analytics at AWS running services such as Amazon Redshift. Before his tenure at AWS, Andi served as CEO and co-founder of Zend Technologies, the commercial backer of open-source PHP.

Andi has over 20 years of experience as an open source contributor and leader. He co-authored open source PHP. He is an emeritus member of the Apache Software Foundation and served on the Eclipse Foundation’s board of directors. He holds a bachelor’s degree in Computer Science from the Technion, Israel Institute of Technology.

Links Referenced:

  • LinkedIn: https://www.linkedin.com/in/andigutmans/
  • Twitter: https://twitter.com/andigutmans

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Sysdig. Sysdig secures your cloud from source to run. They believe, as do I, that DevOps and security are inextricably linked. If you wanna learn more about how they view this, check out their blog, it's definitely worth the read. To learn more about how they are absolutely getting it right from where I sit, visit Sysdig.com and tell them that I sent you. That's S Y S D I G.com. And my thanks to them for their continued support of this ridiculous nonsense.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. This promoted episode is brought to us by our friends at Google Cloud, and in so doing, they have gotten a guest to appear on this show that I have been low-key trying to get here for a number of years. Andi Gutmans is VP and GM of Databases at Google Cloud. Andi, thank you for joining me.

Andi: Corey, thanks so much for having me.

Corey: I have to begin with the obvious. Given that one of my personal passion projects is misusing every cloud service I possibly can as a database, where do you start and where do you stop as far as saying, “Yes, that’s a database,” so it rolls up to me and, “No, that’s not a database, so someone else can deal with the nonsense?”

Andi: I’m in charge of the operational databases, so that includes both the managed third-party databases such as MySQL, Postgres, SQL Server, and then also the cloud-first databases, such as Spanner, Big Table, Firestore, and AlloyDB. So, I suggest that’s where you start because those are all awesome services. And then what doesn’t fall underneath, kind of, that purview are things like BigQuery, which is an analytics, you know, data warehouse, and other analytics engines. And of course, there’s always folks who bring in their favorite, maybe, lesser-known or less popular database and self-manage it on GCE, on Compute.

Corey: Before you wound up at Google Cloud, you spent roughly four years at AWS as VP of Analytics, which is, again, one of those very hazy type of things. Where does it start? Where does it stop? It’s not at all clear from the outside. But even before that, you were, I guess, something of a legendary figure, which I know is always a weird thing for people to hear.

But you were partially at least responsible for the Zend Framework in the PHP world, which I didn’t realize what the heck that was, despite supporting it in production at a couple of jobs, until after I, for better or worse, was no longer trusted to support production environments anymore. Which, honestly, if you can get out, I’m a big proponent of doing that. You sleep so much better without a pager. How did you go from programming languages all the way on over to databases? It just seems like a very odd mix.

Andi: Yeah. No, that’s a great question. So, I was one of the core developers of PHP, and you know, I had been in the PHP community for quite some time. I also helped ideate. The Zend Framework, which was the company that, you know, I co-founded Zend Technologies was kind of the company behind PHP.

So, like Red Hat supports Linux commercially, we supported PHP. And I was very much focused on developers, programming languages, frameworks, IDEs, and that was, you know, really exciting. I had also done quite a bit of work on interoperability with databases, right, because behind every application, there’s a database, and so a lot of what we focused on is a great connectivity to MySQL, to Postgres, to other databases, and I got to kind of learn the database world from the outside from the application builders. We sold our company in I think it was 2015 and so I had to kind of figure out what’s next. And so, one option would have been, hey, stay in programming languages, but what I learned over the many years that I worked with application developers is that there’s a huge amount of value in data.

And frankly, I’m a very curious person; I always like to learn, so there was this opportunity to join Amazon, to join the non-relational database side, and take myself completely out of my comfort zone. And actually, I joined AWS to help build the graph database Amazon Neptune, which was even more out of my comfort zone than even probably a relational database. So, I kind of like to do different things and so I joined and I had to learn, you know how to build a database pretty much from the ground up. I mean, of course, I didn’t do the coding, but I had to learn enough to be dangerous, and so I worked on a bunch of non-relational databases there such as, you know, Neptune, Redis, Elasticsearch, DynamoDB Accelerator. And then there was the opportunity for me to actually move over from non-relational databases to analytics, which was another way to get myself out of my comfort zone.

And so, I moved to run the analytic space, which included services like Redshift, like EMR, Athena, you name it. So, that was just a great experience for me where I got to work with a lot of awesome people and learn a lot. And then the opportunity arose to join Google and actually run the Google transactional databases including their older relational databases. And by the way, my job actually have two jobs. One job is running Spanner and Big Table for Google itself—meaning, you know, search ads and YouTube and everything runs on these databases—and then the second job is actually running external-facing databases for external customers.

Corey: How alike are those two? Is it effectively the exact same thing, just with different API endpoints? Are they two completely separate universes? It’s always unclear from the outside when looking at large companies that effectively eat versions of their own dog food, where their internal usage of these things starts and stops.

Andi: So, great question. So, Cloud Spanner and Cloud Big Table do actually use the internal Spanner and Big Table. So, at the core, it’s exactly the same engine, the same runtime, same storage, and everything. However, you know, kind of, internally, the way we built the database APIs was kind of good for scrappy, you know, Google engineers, and you know, folks are kind of are okay, learning how to fit into the Google ecosystem, but when we needed to make this work for enterprise customers, we needed a cleaner APIs, we needed authentication that was an external, right, and so on, so forth. So, think about we had to add an additional set of APIs on top of it, and management, right, to really make these engines accessible to the external world.

So, it’s running the same engine under the hood, but it is a different set of APIs, and a big part of our focus is continuing to expose to enterprise customers all the goodness that we have on the internal system. So, it’s really about taking these very, very unique differentiated databases and democratizing access to them to anyone who wants to.

Corey: I’m curious to get your position on the idea that seems to be playing it’s—I guess, a battle that’s been playing itself out in a number of different customer conversations. And that is, I guess, the theoretical decision between, do we go towards general-purpose databases and more or less treat every problem as a nail in search of a hammer or do you decide that every workload gets its own custom database that aligns the best with that particular workload? There are trade-offs in either direction, but I’m curious where you land on that given that you tend to see a lot more of it than I do.

Andi: No, that’s a great question. And you know, just for the viewers who maybe aren’t aware, there’s kind of two extreme points of view, right? There’s one point of view that says, purpose-built for everything, like, every specific pattern, like, build bespoke databases, it’s kind of a best-of-breed approach. The problem with that approach is it becomes extremely complex for customers, right? Extremely complex to decide what to use, they might need to use multiple for the same application, and so that can be a bit daunting as a customer. And frankly, there’s kind of a law of diminishing returns at some point.

Corey: Absolutely. I don’t know what the DBA role of the future is, but I don’t think anyone really wants it to be, “Oh, yeah. We’re deciding which one of these three dozen manage database services is the exact right fit for each and every individual workload.” I mean, at some point it feels like certain cloud providers believe that not only every workload should have its own database, but almost every workload should have its own database service. It’s at some point, you’re allowed to say no and stop building these completely, what feel like to me, Byzantine, esoteric database engines that don’t seem to have broad applicability to a whole lot of problems.

Andi: Exactly, exactly. And maybe the other extreme is what folks often talk about as multi-model where you say, like, “Hey, I’m going to have a single storage engine and then map onto that the relational model, the document model, the graph model, and so on.” I think what we tend to see is if you go too generic, you also start having performance issues, you may not be getting the right level of abilities and trade-offs around consistency, and replication, and so on. So, I would say Google, like, we’re taking a very pragmatic approach where we’re saying, “You know what? We’re not going to solve all of customer problems with a single database, but we’re also not going to have two dozen.” Right?

So, we’re basically saying, “Hey, let’s understand that the main characteristics of the workloads that our customers need to address, build the best services around those.” You know, obviously, over time, we continue to enhance what we have to fit additional models. And then frankly, we have a really awesome partner ecosystem on Google Cloud where if someone really wants a very specialized database, you know, we also have great partners that they can use on Google Cloud and get great support and, you know, get the rest of the benefits of the platform.

Corey: I’m very curious to get your take on a pattern that I’ve seen alluded to by basically every vendor out there except the couple of very obvious ones for whom it does not serve their particular vested interests, which is that there’s a recurring narrative that customers are demanding open-source databases for their workloads. And when you hear that, at least, people who came up the way that I did, spending entirely too much time on Freenode, back when that was not a deeply problematic statement in and of itself, where, yes, we’re open-source, I guess, zealots is probably the best terminology, and yeah, businesses are demanding to participate in the open-source ecosystem. Here in reality, what I see is not ideological purity or anything like that and much more to do with, “Yeah, we don’t like having a single commercial vendor for our databases that basically plays the insert quarter to continue dance whenever we’re trying to wind up doing something new. We want the ability to not have licensing constraints around when, where, how, and how quickly we can run databases.” That’s what I hear when customers are actually talking about open-source versus proprietary databases. Is that what you see or do you think that plays out differently? Because let’s be clear, you do have a number of database services that you offer that are not open-source, but are also absolutely not tied to weird licensing restrictions either?

Andi: That’s a great question, and I think for years now, customers have been in a difficult spot because the legacy proprietary database vendors, you know, knew how sticky the database is, and so as a result, you know, the prices often went up and was not easy for customers to kind of manage costs and agility and so on. But I would say that’s always been somewhat of a concern. I think what I’m seeing changing and happening differently now is as customers are moving into the cloud and they want to run hybrid cloud, they want to run multi-cloud, they need to prove to their regulator that it can do a stressed exit, right, open-source is not just about reducing cost, it’s really about flexibility and kind of being in control of when and where you can run the workloads. So, I think what we’re really seeing now is a significant surge of customers who are trying to get off legacy proprietary database and really kind of move to open APIs, right, because they need that freedom. And that freedom is far more important to them than even the cost element.

And what’s really interesting is, you know, a lot of these are the decision-makers in these enterprises, not just the technical folks. Like, to your point, it’s not just open-source advocates, right? It’s really the business people who understand they need the flexibility. And by the way, even the regulators are asking them to show that they can flexibly move their workloads as they need to. So, we’re seeing a huge interest there and, as you said, like, some of our services, you know, are open-source-based services, some of them are not.

Like, take Spanner, as an example, it is heavily tied to how we build our infrastructure and how we build our systems. Like, I would say, it’s almost impossible to open-source Spanner, but what we’ve done is we’ve basically embraced open APIs and made sure if a customer uses these systems, we’re giving them control of when and where they want to run their workloads. So, for example, Big Table has an HBase API; Spanner now has a Postgres interface. So, our goal is really to give customers as much flexibility and also not lock them into Google Cloud. Like, we want them to be able to move out of Google Cloud so they have control of their destiny.

Corey: I’m curious to know what you see happening in the real world because I can sit here and come up with a bunch of very well-thought-out logical reasons to go towards or away from certain patterns, but I spent years building things myself. I know how it works, you grab the closest thing handy and throw it in and we all know that there is nothing so permanent as a temporary fix. Like, that thing is load-bearing and you’ll retire with that thing still in place. In the idealized world, I don’t think that I would want to take a dependency on something like—easy example—Spanner or AlloyDB because despite the fact that they have Postgres-squeal—yes, that’s how I pronounce it—compatibility, the capabilities of what they’re able to do under the hood far exceed and outstrip whatever you’re going to be able to build yourself or get anywhere else. So, there’s a dataflow architectural dependency lock-in, despite the fact that it is at least on its face, Postgres compatible. Counterpoint, does that actually matter to customers in what you are seeing?

Andi: I think it’s a great question. I’ll give you a couple of data points. I mean, first of all, even if you take a complete open-source product, right, running them in different clouds, different on-premises environments, and so on, fundamentally, you will have some differences in performance characteristics, availability characteristics, and so on. So, the truth is, even if you use open-source, right, you’re not going to get a hundred percent of the same characteristics where you run that. But that said, you still have the freedom of movement, and with I would say and not a huge amount of engineering investment, right, you’re going to make sure you can run that workload elsewhere.

I kind of think of Spanner in the similar way where yes, I mean, you’re going to get all those benefits of Spanner that you can’t get anywhere else, like unlimited scale, global consistency, right, no maintenance downtime, five-nines availability, like, you can’t really get that anywhere else. That said, not every application necessarily needs it. And you still have that option, right, that if you need to, or want to, or we’re not giving you a reasonable price or reasonable price performance, but we’re starting to neglect you as a customer—which of course we wouldn’t, but let’s just say hypothetically, that you know, that could happen—that you still had a way to basically go and run this elsewhere. Now, I’d also want to talk about some of the upsides something like Spanner gives you. Because you talked about, you want to be able to just grab a few things, build something quickly, and then, you know, you don’t want to be stuck.

The counterpoint to that is with Spanner, you can start really, really small, and then let’s say you’re a gaming studio, you know, you’re building ten titles hoping that one of them is going to take off. So, you can build ten of those, you know, with very minimal spend on Spanner and if one takes off overnight, it’s really only the database where you don’t have to go and re-architect the application; it’s going to scale as big as you need it to. And so, it does enable a lot of this innovation and a lot of cost management as you try to get to that overnight success.

Corey: Yeah, overnight success. I always love that approach. It’s one of those, “Yeah, I became an overnight success after only ten short years.” It becomes this idea people believe it’s in fits and starts, but then you see, I guess, on some level, the other side of it where it’s a lot of showing up and doing the work. I have to confess, I didn’t do a whole lot of admin work in my production years that touched databases because I have an aura and I’m unlucky, and it turns out that when you blow away some web servers, everyone can laugh and we’ll reprovision stateless things.

Get too close to the data warehouse, for example, and you don’t really have a company left anymore. And of course, in the world of finance that I came out of, transactional integrity is also very much a thing. A question that I had [centers 00:17:51] really around one of the predictions you gave recently at Google Cloud Next, which is your prediction for the future is that transactional and analytical workloads from a database perspective will converge. What’s that based on?

Andi: You know, I think we’re really moving from a world where customers are trying to make real-time decisions, right? If there’s model drift from an AI and ML perspective, want to be able to retrain their models as quickly as possible. So, everything is fast moving into streaming. And I think what you’re starting to see is, you know, customers don’t have that time to wait for analyzing their transactional data. Like in the past, you do a batch job, you know, once a day or once an hour, you know, move the data from your transactional system to analytical system, but that’s just not how it is always-on businesses run anymore, and they want to have those real-time insights.

So, I do think that what you’re going to see is transactional systems more and more building analytical capabilities, analytical systems building, and more transactional, and then ultimately, cloud platform providers like us helping fill that gap and really making data movement seamless across transactional analytical, and even AI and ML workloads. And so, that’s an area that I think is a big opportunity. I also think that Google is best positioned to solve that problem.

Corey: Forget everything you know about SSH and try Tailscale. Imagine if you didn't need to manage PKI or rotate SSH keys every time someone leaves. That'd be pretty sweet, wouldn't it? With Tailscale SSH, you can do exactly that. Tailscale gives each server and user device a node key to connect to its VPN, and it uses the same node key to authorize and authenticate SSH.

Basically you're SSHing the same way you manage access to your app. What's the benefit here? Built-in key rotation, permissions as code, connectivity between any two devices, reduce latency, and there's a lot more, but there's a time limit here. You can also ask users to reauthenticate for that extra bit of security. Sounds expensive?

Nope, I wish it were. Tailscale is completely free for personal use on up to 20 devices. To learn more, visit snark.cloud/tailscale. Again, that's snark.cloud/tailscale

Corey: On some level, I’ve found that, at least in my own work, that once I wind up using a database for something, I’m inclined to try and stuff as many other things into that database as I possibly can just because getting a whole second data store, taking a dependency on it for any given workload tends to be a little bit on the, I guess, challenging side. Easy example of this. I’ve talked about it previously in various places, but I was talking to one of your colleagues, [Sarah Ellis 00:19:48], who wound up at one point making a joke that I, of course, took way too far. Long story short, I built a Twitter bot on top of Google Cloud Functions that every time the Azure brand account tweets, it simply quote-tweets that translates their tweet into all caps, and then puts a boomer-style statement in front of it if there’s room. This account is @cloudboomer.

Now, the hard part that I had while doing this is everything stateless works super well. Where do I wind up storing the ID of the last tweet that it saw on his previous run? And I was fourth and inches from just saying, “Well, I’m already using Twitter so why don’t we use Twitter as a database?” Because everything’s a database if you’re either good enough or bad enough at programming. And instead, I decided, okay, we’ll try this Firebase thing first.

And I don’t know if it’s Firestore, or Datastore or whatever it’s called these days, but once I wrap my head around it incredibly effective, very fast to get up and running, and I feel like I made at least a good decision, for once in my life, involving something touching databases. But it’s hard. I feel like I’m consistently drawn toward the thing I’m already using as a default database. I can’t shake the feeling that that’s the wrong direction.

Andi: I don’t think it’s necessarily wrong. I mean, I think, you know, with Firebase and Firestore, that combination is just extremely easy and quick to build awesome mobile applications. And actually, you can build mobile applications without a middle tier which is probably what attracted you to that. So, we just see, you know, huge amount of developers and applications. We have over 4 million databases in Firestore with just developers building these applications, especially mobile-first applications. So, I think, you know, if you can get your job done and get it done effectively, absolutely stick to them.

And by the way, one thing a lot of people don’t know about Firestore is it’s actually running on Spanner infrastructure, so Firestore has the same five-nines availability, no maintenance downtime, and so on, that has Spanner, and the same kind of ability to scale. So, it’s not just that it’s quick, it will actually scale as much as you need it to and be as available as you need it to. So, that’s on that piece. I think, though, to the same point, you know, there’s other databases that we’re then trying to make sure kind of also extend their usage beyond what they’ve traditionally done. So, you know, for example, we announced AlloyDB, which I kind of call it Postgres on steroids, we added analytical capabilities to this transactional database so that as customers do have more data in their transactional database, as opposed to having to go somewhere else to analyze it, they can actually do real-time analytics within that same database and it can actually do up to 100 times faster analytics than open-source Postgres.

So, I would say both Firestore and AlloyDB, are kind of good examples of if it works for you, right, we’ll also continue to make investments so the amount of use cases you can use these databases for continues to expand over time.

Corey: One of the weird things that I noticed just looking around this entire ecosystem of databases—and you’ve been in this space long enough to, presumably, have seen the same type of evolution—back when I was transiting between different companies a fair bit, sometimes because I was consulting and other times because I’m one of the greatest in the world at getting myself fired from jobs based upon my personality, I found that the default standard was always, “Oh, whatever the database is going to be, it started off as MySQL and then eventually pivots into something else when that starts falling down.” These days, I can’t shake the feeling that almost everywhere I look, Postgres is the answer instead. What changed? What did I miss in the ecosystem that’s driving that renaissance, for lack of a better term?

Andi: That’s a great question. And, you know, I have been involved in—I’m going to date myself a bit—but in PHP since 1997, pretty much, and one of the things we kind of did is we build a really good connector to MySQL—and you know, I don’t know if you remember, before MySQL, there was MS SQL. So, the MySQL API actually came from MS SQL—and we bundled the MySQL driver with PHP. And so, kind of that LAMP stack really took off. And kind of to your point, you know, the default in the web, right, was like, you’re going to start with MySQL because it was super easy to use, just fun to use.

By the way, I actually wrote—co-authored—the tab completion in the MySQL client. So like, a lot of these kinds of, you know, fun, simple ways of using MySQL were there, and frankly, was super fast, right? And so, kind of those fast reads and everything, it just was great for web and for content. And at the time, Postgres kind of came across more like a science project. Like the folks who were using Postgres were kind of the outliers, right, you know, the less pragmatic folks.

I think, what’s changed over the past, how many years has it been now, 25 years—I’m definitely dating myself—is a few things: one, MySQL is still awesome, but it didn’t kind of go in the direction of really, kind of, trying to catch up with the legacy proprietary databases on features and functions. Part of that may just be that from a roadmap perspective, that’s not where the owner wanted it to go. So, MySQL today is still great, but it didn’t go into that direction. In parallel, right, customers wanting to move more to open-source. And so, what they found this, the thing that actually looks and smells more like legacy proprietary databases is actually Postgres, plus you saw an increase of investment in the Postgres ecosystem, also very liberal license.

So, you have lots of other databases including commercial ones that have been built off the Postgres core. And so, I think you are today in a place where, for mainstream enterprise, Postgres is it because that is the thing that has all the features that the enterprise customer is used to. MySQL is still very popular, especially in, like, content and web, and mobile applications, but I would say that Postgres has really become kind of that de facto standard API that’s replacing the legacy proprietary databases.

Corey: I’ve been on the record way too much as saying, with some justification, that the best database in the world that should be used for everything is Route 53, specifically, TXT records. It’s a key-value store and then anyone who’s deep enough into DNS or databases generally gets a slightly greenish tinge and feels ill. That is my simultaneous best and worst database. I’m curious as to what your most controversial opinion is about the worst database in the world that you’ve ever seen.

Andi: This is the worst database? Or—

Corey: Yeah. What is the worst database that you’ve ever seen? I know, at some level, since you manage all things database, I’m asking you to pick your least favorite child, but here we are.

Andi: Oh, that’s a really good question. No, I would say probably the, “Worst database,” double-quotes is just the file system, right? When folks are basically using the file system as regular database. And that can work for, you know, really simple apps, but as apps get more complicated, that’s not going to work. So, I’ve definitely seen some of that.

I would say the most awesome database that is also file system-based kind of embedded, I think was actually SQLite, you know? And SQLite is actually still very, very popular. I think it sits on every mobile device pretty much on the planet. So, I actually think it’s awesome, but it’s, you know, it’s on a database server. It’s kind of an embedded database, but it’s something that I, you know, I’ve always been pretty excited about. And, you know, their stuff [unintelligible 00:27:43] kind of new, interesting databases emerging that are also embedded, like DuckDB is quite interesting. You know, it’s kind of the SQLite for analytics.

Corey: We’ve been using it for a few things around a bill analysis ourselves. It’s impressive. I’ve also got to say, people think that we had something to do with it because we’re The Duckbill Group, and it’s DuckDB. “Have you done anything with this?” And the answer is always, “Would you trust me with a database? I didn’t think so.” So no, it’s just a weird coincidence. But I liked that a lot.

It’s also counterintuitive from where I sit because I’m old enough to remember when Microsoft was teasing the idea of WinFS where they teased a future file system that fundamentally was a database—I believe it’s an index or journal for all of that—and I don’t believe anything ever came of it. But ugh, that felt like a really weird alternate world we could have lived in.

Andi: Yeah. Well, that’s a good point. And by the way, you know, if I actually take a step back, right, and I kind of half-jokingly said, you know, file system and obviously, you know, all the popular databases persist on the file system. But if you look at what’s different in cloud-first databases, right, like, if you look at legacy proprietary databases, the typical setup is wright to the local disk and then do asynchronous replication with some kind of bounded replication lag to somewhere else, to a different region, or so on. If you actually start to look at what the cloud-first databases look like, they actually write the data in multiple data centers at the same time.

And so, kind of joke aside, as you start to think about, “Hey, how do I build the next generation of applications and how do I really make sure I get the resiliency and the durability that the cloud can offer,” it really does take a new architecture. And so, that’s where things like, you know, Spanner and Big Table, and kind of, AlloyDB databases are truly architected for the cloud. That’s where they actually think very differently about durability and replication, and what it really takes to provide the highest level of availability and durability.

Corey: On some level, I think one of the key things for me to realize was that in my own experiments, whenever I wind up doing something that is either for fun or I just want see how it works in what’s possible, the scale of what I’m building is always inherently a toy problem. It’s like the old line that if it fits in RAM, you don’t have a big data problem. And then I’m looking at things these days that are having most of a petabyte’s worth of RAM sometimes it’s okay, that definition continues to extend and get ridiculous. But I still find that most of what I do in a database context can be done with almost any database. There’s no reason for me not to, for example, uses a SQLite file or to use an object store—just there’s a little latency, but whatever—or even a text file on disk.

The challenge I find is that as you start scaling and growing these things, you start to run into limitations left and right, and only then it’s one of those, oh, I should have made different choices or I should have built-in abstractions. But so many of these things comes to nothing; it just feels like extra work. What guidance do you have for people who are trying to figure out how much effort to put in upfront when they’re just more or less puttering around to see what comes out of it?

Andi: You know, we like to think about ourselves at Google Cloud as really having a unique value proposition that really helps you future-proof your development. You know, if I look at both Spanner and I look at BigQuery, you can actually start with a very, very low cost. And frankly, not every application has to scale. So, you can start at low cost, you can have a small application, but everyone wants two things: one is availability because you don’t want your application to be down, and number two is if you have to scale you want to be able to without having to rewrite your application. And so, I think this is where we have a very unique value proposition, both in how we built Spanner and then also how we build BigQuery is that you can actually start small, and for example, on Spanner, you can go from one-tenth of what we call an instance, like, a small instance, that is, you know, under $65 a month, you can go to a petabyte scale OLTP environment with thousands of instances in Spanner, with zero downtime.

And so, I think that is really the unique value proposition. We’re basically saying you can hold the stick at both ends: you can basically start small and then if that application doesn’t need to scale, does need to grow, you’re not reengineering your application and you’re not taking any downtime for reprovisioning. So, I think that’s—if I had to give folks, kind of, advice, I say, “Look, what’s done is done. You have workloads on MySQL, Postgres, and so on. That’s great.”

Like, they’re awesome databases, keep on using them. But if you’re truly building a new app, and you’re hoping that app is going to be successful at some point, whether it’s, like you said, all overnight successes take at least ten years, at least you built in on something like Spanner, you don’t actually have to think about that anymore or worry about it, right? It will scale when you need it to scale and you’re not going to have to take any downtime for it to scale. So, that’s how we see a lot of these industries that have these potential spikes, like gaming, retail, also some use cases in financial services, they basically gravitate towards these databases.

Corey: I really want to thank you for taking so much time out of your day to talk with me about databases and your perspective on them, especially given my profound level of ignorance around so many of them. If people want to learn more about how you view these things, where’s the best place to find you?

Andi: Follow me on LinkedIn. I tend to post quite a bit on LinkedIn, I still post a bit on Twitter, but frankly, I’ve moved more of my activity to LinkedIn now. I find it’s—

Corey: That is such a good decision. I envy you.

Andi: It’s a more curated [laugh], you know, audience and so on. And then also, you know, we just had Google Cloud Next. I recorded a session there that kind of talks about database and just some of the things that are new in database-land at Google Cloud. So, that’s another thing that if folks more interested to get more information, that may be something that could be appealing to you.

Corey: We will, of course, put links to all of this in the [show notes 00:34:03]. Thank you so much for your time. I really appreciate it.

Andi: Great. Corey, thanks so much for having me.

Corey: Andi Gutmans, VP and GM of Databases at Google Cloud. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry, insulting comment, then I’m going to collect all of those angry, insulting comments and use them as a database.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Ashish

Ashish has over 13+yrs experience in the Cybersecurity industry with the last 7 focusing primarily helping Enterprise with managing security risk at scale in cloud first world and was the CISO of a global Cloud First Tech company in his last role. Ashish is also a keynote speaker and host of the widely poplar Cloud Security Podcast, a SANS trainer for Cloud Security & DevSecOps. Ashish currently works at Snyk as a Principal Cloud Security Advocate. He is a frequent contributor on topics related to public cloud transformation, Cloud Security, DevSecOps, Security Leadership, future Tech and the associated security challenges for practitioners and CISOs.

Links Referenced:

  • Cloud Security Podcast: https://cloudsecuritypodcast.tv/
  • Personal website: https://www.ashishrajan.com/
  • LinkedIn: https://www.linkedin.com/in/ashishrajan/
  • Twitter: https://twitter.com/hashishrajan
  • Cloud Security Podcast YouTube: https://www.youtube.com/c/CloudSecurityPodcast
  • Cloud Security Podcast LinkedIn: https://www.linkedin.com/company/cloud-security-podcast/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Thinkst Canary. Most folks find out way too late that they’ve been breached. Thinkst Canary changes this. Deploy canaries and canary tokens in minutes, and then forget about them. Attackers tip their hand by touching them, giving you one alert, when it matters. With zero administrative overhead to this and almost no false positives, Canaries are deployed and loved on all seven continents. Check out what people are saying at canary.love today.

Corey: This episode is bought to you in part by our friends at Veeam. Do you care about backups? Of course you don’t. Nobody cares about backups. Stop lying to yourselves! You care about restores, usually right after you didn’t care enough about backups. If you’re tired of the vulnerabilities, costs and slow recoveries when using snapshots to restore your data, assuming you even have them at all living in AWS-land, there is an alternative for you. Check out Veeam, thats V-E-E-A-M for secure, zero-fuss AWS backup that won't leave you high and dry when it’s time to restore. Stop taking chances with your data. Talk to Veeam. My thanks to them for sponsoring this ridiculous podcast.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. This promoted episode is brought to us once again by our friends at Snyk. Snyk does amazing things in the world of cloud security and terrible things with the English language because, despite raising a whole boatload of money, they still stubbornly refuse to buy a vowel in their name. I’m joined today by Principal Cloud Security Advocate from Snyk, Ashish Rajan. Ashish, thank you for joining me.

Corey: Your history is fascinating to me because you’ve been around for a while on a podcast of your own, the Cloud Security Podcast. But until relatively recently, you were a CISO. As has become relatively accepted in the industry, the primary job of the CISO is to get themselves fired, and then, “Well, great. What’s next?” Well, failing upward is really the way to go wherever possible, so now you are at Snyk, helping the rest of us fix our security. That’s my headcanon on all of that anyway, which I’m sure bears scant, if any, resemblance to reality, what’s your version?

Ashish: [laugh]. Oh, well, fortunately, I wasn’t fired. And I think I definitely find that it’s a great way to look at the CISO job to walk towards the path where you’re no longer required because then I think you’ve definitely done your job. I moved into the media space because we got an opportunity to go full-time. I spoke about this offline, but an incident inspired us to go full-time into the space, so that’s what made me leave my CISO job and go full-time into democratizing cloud security as much as possible for anyone and everyone. So far, every day, almost now, so it’s almost like I dream about cloud security as well now.

Corey: Yeah, I dream of cloud security too, but my dreams are of a better world in which people didn’t tell me how much they really care about security in emails that demonstrate how much they failed to care about security until it was too little too late. I was in security myself for a while and got out of it because I was tired of being miserable all the time. But I feel that there’s a deep spiritual alignment between people who care about cost and people who care about security when it comes to cloud—or business in general—because you can spend infinite money on those things, but it doesn’t really get your business further. It’s like paying for fire insurance. It isn’t going to get you to your next milestone, whereas shipping faster, being more effective at launching a feature into markets, that can multiply revenue. That’s what companies are optimized around. It’s, “Oh, right. We have to do the security stuff,” or, “We have to fix the AWS billing piece.” It feels, on some level, like it’s a backburner project most of the time and it’s certainly invested in that way. What’s your take on that?

Ashish: I tend to disagree with that, for a couple reasons.

Corey: Excellent. I love arguments.

Ashish: I feel this in a healthy way as well. A, I love the analogy of spiritual animals where they are cost optimization as well as the risk aversion as well. I think where I normally stand—and this is what I had to unlearn after doing years of cybersecurity—was that initially, we always used to be—when I say ‘we,’ I mean cybersecurity folks—we always used to be like police officers. Is that every time there’s an incident, it turns into a crime scene, and suddenly we’re all like, “Pew, pew, pew,” with trying to get all the evidence together, let’s make this isolated as much—as isolated as possible from the rest of the environment, and let’s try and resolve this.

I feel like in Cloud has asked people to become more collaborative, which is a good problem to have. It also encourages that, I don’t know how many people know this, but the reason we have brakes in our cars is not because we can slow down the car; it’s so that we can go faster. And I feel security is the same thing. The guardrails we talk about, the risks that you’re trying to avert, the reason you’re trying to have security is not to slow down but to go faster. Say for example in an ideal world, to quote what you were saying earlier if we were to do the right kind of encryption—I’m just going to use the most basic example—if we just do encryption, right, and just ensure that as a guardrail, the entire company needs to have encryption at rest, encryption in transit, period, nothing else, no one cares about anything else.

But if you just lay that out as a framework and this is our guardrail, no one brakes this, and whoever does, hey we—you know, slap on the wrist and come back on to the actual track, but keep going forward. That just means any project that comes in that meets [unintelligible 00:04:58] criteria. Keeps going forward, as many times we want to go into production. Doesn’t matter. So, that is the new world of security that we are being asked to move towards where Amazon re:Invent is coming in, there will be another, I don’t know, three, four hundred services that will be released. How many people, irrespective of security, would actually know all of those services? They would not. So, [crosstalk 00:05:20]—

Corey: Oh, we’ve long since passed the point where I can convincingly talk about AWS services that don’t really exist and not get called out on it by Amazon employees. No one keeps them on their head. Except me because I’m sad.

Ashish: Oh, no, but I think you’re right, though. I can’t remember who was it—maybe Andrew Vogel or someone—they didn’t release a service which didn’t exist, and became, like, a thing on Twitter. Everyone—

Corey: Ah, AWS’s Infinidash. I want to say that was Joe Nash out of Twilio at the time. I don’t recall offhand if I’m right on that, but that’s how it feels. Yeah, it was certainly not me. People said that was my idea. Nope, nope, I just basically amplified it to a huge audience.

But yeah, it was a brilliant idea, just because it’s a fake service so everyone could tell stories about it. And amazing product feedback, if you look at it through the right lens of how people view your company and your releases when they get this perfect, platonic ideal of what it is you might put out there, what do people say about it?

Ashish: Yeah. I think that’s to your point, I will use that as an example as well to talk about things that there will always be a service which we will be told about for the first time, which we will not know. So, going back to the unlearning part, as a security team, we just have to understand that we can’t use the old ways of, hey, I want to have all the controls possible, cover all there is possible. I need to have a better understanding of all the cloud services because I’ve done, I don’t know, 15 years of cloud, there is no one that has 10, 15 years of cloud unless you’re I don’t know someone from Amazon employee yourself. Most people these days still have five to six years experience and they’re still learning.

Even the cloud engineering folks or the DevOps folks, they’re all still learning and the tooling is continuing to evolve. So yeah, I think I definitely find that the security in this cloud world a lot more collaborative and it’s being looked at as the same function as a brake would have in car: to help you go faster, not to just slam the brake every time it’s like, oh, my God, is the situation isolated and to police people.

Corey: One of the points I find that is so aligned between security and cost—and you alluded to it a minute ago—is the idea of helping companies go faster safely. To that end, guardrails have to be at least as easy as just going off and doing it cow-person style. Because if it’s not, it’s more work in any way, shape, or form, people won’t do it. People will not tag their resources by hand, people will not go through and use the dedicated account structure you’ve got that gets in their way and screams at them every time they try to use one of the native features built into the platform. It has to get out of their way and make things easier, not worse, or people fight it, they go around it, and you’re never going to get buy-in.

Ashish: Do you feel like cost is something that a lot more people pay a lot more attention to because, you know, that creeps into your budget? Like, as people who’ve been leaders before, and this was the conversation, they would just go, “Well, I only have, I don’t know, 100,000 to spend this quarter,” or, “This year,” and they are the ones who—are some of them, I remember—I used to have this manager, once, a CTO would always be conscious about the spend. It’s almost like if you overspend, where do you get the money from? There’s no money to bring in extra. Like, no. There’s a set money that people plan for any year for a budget. And to your point about if you’re not keeping an eye on how are we spending this in the AWS context because very easy to spend the entire money in one day, or in the cloud context. So, I wonder if that is also a big driver for people to feel costs above security? Where do you stand on that?

Corey: When it comes to cost, one of the nice things about it—and this is going to sound sarcastic, but I swear to you it’s not—it’s only money.

Ashish: Mmm.

Corey: Think about that for a second because it’s true. Okay, we wound up screwing up and misconfiguring something and overspending. Well, there are ways around that. You can call AWS, you can get credits, you can get concessions made for mistakes, you can sign larger contracts and get a big pile of proof of concept credit et cetera, et cetera. There are ways to make that up, whereas with security, it’s there are no do-overs on security breaches.

Ashish: No, that’s a good point. I mean, you can always get more money, use a credit card, worst case scenario, but you can’t do the same for—there’s a security breach and suddenly now—hopefully, you don’t have to call New York Times and say, “Can you undo that article that you just have posted that told you it was a mistake. We rewinded what we did.”

Corey: I’m curious to know what your take is these days on the state of the cloud security community. And the reason I bring that up is, well, I started about a year-and-a-half ago now doing a podcast every Thursday. Which is Last Week in AWS: Security Edition because everything else I found in the industry that when I went looking was aimed explicitly at either—driven by the InfoSec community, which is toxic and a whole bunch of assumed knowledge already built in that looks an awful lot like gatekeeping, which is the reason I got out of InfoSec in the first place, or alternately was completely vendor-captured, where, okay, great, we’re going to go ahead and do a whole bunch of interesting content and it’s all brought to you by this company and strangely, all of the content is directly align with doing some pretty weird things that you wouldn’t do unless you’re trying to build a business case for that company’s product. And it just feels hopelessly compromised. I wanted to find something that was aimed at people who had to care about security but didn’t have security as part of their job title. Think DevOps types and you’re getting warmer.

That’s what I wound up setting out to build. And when all was said and done, I wasn’t super thrilled with, honestly, how alone it still felt. You’ve been doing this for a while, and you’re doing a great job at it, don’t get me wrong, but there is the question that—and I understand they’re sponsoring this episode, but the nice thing about promoted guest episodes is that they can buy my attention, not my opinion. How do you retain creative control of your podcast while working for a security vendor?

Ashish: So, that’s a good question. So, Snyk by themselves have not ever asked us to change any piece of content; we have been working with them for the past few months now. The reason we kind of came along with Snyk was the alignment. And we were talking about this earlier for I totally believe that DevSecOps and cloud security are ultimately going to come together one day. That may not be today, that may not be tomorrow, that may not be in 2022, or maybe 2023, but there will be a future where these two will sit together.

And the developer-first security mentality that they had, in this context from cloud prospective—developers being the cloud engineers, the DevOps people as you called out, the reason you went in that direction, I definitely want to work with them. And ultimately, there would never be enough people in security to solve the problem. That is the harsh reality. There would never be enough people. So, whether it’s cloud security or not, like, for people who were at AWS re:Inforce, the first 15 minutes by Steve Schmidt, CSO of Amazon, was get a security guardian program.

So, I’ve been talking about it, everyone else is talking about right now, Amazon has become the first CSP to even talk about this publicly as well that we should have security guardians. Which by the way, I don’t know why, but you can still call it—it is technically DevSecOps what you’re trying to do—they spoke about a security champion program as part of the keynote that they were running. Nothing to do with cloud security, but the idea being how much of this workload can we share? We can raise, as a security team—for people who may be from a security background listening to this—how much elevation can we provide the risk in front of the right people who are a decision-maker? That is our role.

We help them with the governance, we help with managing it, but we don’t know how to solve the risk or close off a risk, or close off a vulnerability because you might be the best person because you work in that application every day, every—you know the bandages that are put in, you know all the holes that are there, so the best threat model can be performed by the person who works on a day-to-day, not a security person who spent, like, an hour with you once a week because that’s the only time they could manage. So, going back to the Snyk part, that’s the mission that we’ve had with the podcast; we want to democratize cloud security and build a community around neutral information. There is no biased information. And I agree with what you said as well, where a lot of the podcasts outside of what we were finding was more focused on, “Hey, this is how you use AWS. This is how you use Azure. This is how you use GCP.”

But none of them were unbiased in the opinion. Because real life, let’s just say even if I use the AWS example—because we are coming close to the AWS re:Invent—they don’t have all the answers from a security perspective. They don’t have all the answers from an infrastructure perspective or cloud-native perspective. So, there are some times—or even most times—people are making a call where they’re going outside of it. So, unbiased information is definitely required and it is not there enough.

So, I’m glad that at least people like yourself are joining, and you know, creating the world where more people are trying to be relatable to DevOps people as well as the security folks. Because it’s hard for a security person to be a developer, but it’s easy for a developer or an engineer to understand security. And the simplest example I use is when people walk out of their house, they lock the door. They’re already doing security. This is the same thing we’re asking when we talk about security in the cloud or in the [unintelligible 00:14:49] as well. Everyone is, it just it hasn’t been pointed out in the right way.

Corey: I’m curious as to what it is that gets you up in the morning. Now, I know you work in security, but you’re also not a CISO anymore, so I’m not asking what gets you up at 2 a.m. because we know what happens in the security space, then. There’s a reason that my area of business focus is strictly a business hours problem. But I’d love to know what it is about cloud security as a whole that gets you excited.

Ashish: I think it’s an opportunity for people to get into the space without the—you know, you said gatekeeper earlier, those gatekeepers who used to have that 25 years experience in cybersecurity, 15 years experience in cybersecurity, Cloud has challenged that norm. Now, none of that experience helps you do AWS services better. It definitely helps you with the foundational pieces, definitely helps you do identity, networking, all of that, but you still have to learn something completely new, a new way of working, which allows for a lot of people who earlier was struggling to get into cybersecurity, now they have an opening. That’s what excites me about cloud security, that it has opened up a door which is beyond your CCNA, CISSP, and whatever else certification that people want to get. By the way, I don’t have a CISSP, so I can totally throw CISSP under the bus.

But I definitely find that cloud security excites me every morning because it has shown me light where, to what you said, it was always a gated community. Although that’s a very huge generalization. There’s a lot of nice people in cybersecurity who want to mentor and help people get in. But Cloud security has pushed through that door, made it even wider than it was before.

Corey: I think there’s a lot to be said for the concept of sending the elevator back down. I really have remarkably little patience for people who take the perspective of, “Well, I got mine so screw everyone else.” The next generation should have it easier than we did, figuring out where we land in the ecosystem, where we live in the space. And there are folks who do a tremendous job of this, but there are also areas where I think there is significant need for improvement. I’m curious to know what you see as lacking in the community ecosystem for folks who are just dipping their toes into the water of cloud security.

Ashish: I think that one, there’s misinformation as well. The first one being, if you have never done IT before you can get into cloud security, and you know, you will do a great job. I think that is definitely a mistake to just accept the fact if Amazon re:Invent tells you do all these certifications, or Azure does the same, or GCP does the same. If I’ll be really honest—and I feel like I can be honest, this is a safe space—that for people who are listening in, if you’re coming to the space for the first time, whether it’s cloud or cloud security, if you haven’t had much exposure to the foundational pieces of it, it would be a really hard call. You would know all the AWS services, you will know all the Azure services because you have your certification, but if I was to ask you, “Hey, help me build an application. What would be the architecture look like so it can scale?”

“So, right now we are a small pizza-size ten-people team”—I’m going to use the Amazon term there—“But we want to grow into a Facebook tomorrow, so please build me an architecture that can scale.” And if you regurgitate what Amazon has told you, or Azure has told you, or GCP has told you, I can definitely see that you would struggle in the industry because that’s not how, say every application is built. Because the cloud service provider would ask you to drink the Kool-Aid and say they can solve all your problems, even though they don’t have all the servers in the world. So, that’s the first misinformation.

The other one too, for people who are transitioning, who used to be in IT or in cybersecurity and trying to get into the cloud security space, the challenge over there is that outside of Amazon, Google, and Microsoft, there is not a lot of formal education which is unbiased. It is a great way to learn AWS security on how amazing AWS is from AWS people, the same way Microsoft will be [unintelligible 00:19:10], however, when it comes down to actual formal education, like the kind that you and I are trying to provide through a podcast, me with the Cloud Security Podcast, you with Last Week in AWS in the Security Edition, that kind of unbiased formal education, like free education, like what you and I are doing does definitely exist and I guess I’m glad we have company, that you and I both exist in this space, but formal education is very limited. It’s always behind, say an expensive paid wall sometimes, and rightly so because it’s information that would be helpful. So yeah, those two things.

Corey: This episode is sponsored in part by our friends at Uptycs. Attackers don’t think in silos, so why would you have siloed solutions protecting cloud, containers, and laptops distinctly? Meet Uptycs - the first unified solution prioritizes risk across your modern attack surface—all from a single platform, UI, and data model. Stop by booth 3352 at AWS re:Invent in Las Vegas to see for yourself and visit uptycs.com. That’s U-P-T-Y-C-S.com.

Corey: One of the problems that I have with the way a lot of cloud security stuff is situated is that you need to have something running to care about the security of. Yeah, I can spin up a VM in the free tier of most of these environments, and okay, “How do I secure a single Linux box?” Okay, yes, there are a lot of things you can learn there, but it’s very far from a holistic point of view. You need to have the infrastructure running at reasonable scale first, in order to really get an effective lab that isn’t contrived.

Now, Snyk is a security company. I absolutely understand and have no problem with the fact that you charge your customers money in order to get security outcomes that are better than they would have otherwise. I do not get why AWS and GCP charge extra for security. And I really don’t get why Azure charges extra for security and then doesn’t deliver security by dropping the ball on it, which is neither here nor there.

Ashish: [laugh].

Corey: It feels like there’s an economic form of gatekeeping, where you must spend at least this much money—or work for someone who does—in order to get exposure to security the way that grownups think about it. Because otherwise, all right, I hit my own web server, I have ten lines in the logs. Now, how do I wind up doing an analysis run to figure out what happened? I pull it up on my screen and I look at it. You need a point of scale before anything that the modern world revolves around doesn’t seem ludicrous.

Ashish: That’s a good point. Also because we don’t talk about the responsibility that the cloud service provider has themselves for security, like the encryption example that I used earlier, as a guardrail, it doesn’t take much for them to enable by default. But how many do that by default? I feel foolish sometimes to tell people that, “Hey, you should have encryption enabled on your storage which is addressed, or in transit.”

It should be—like, we have services like Let’s Encrypt and other services, which are trying to make this easily available to everyone so everyone can do SSL or HTTPS. And also, same goes for encryption. It’s free and given the choice that you can go customer-based keys or your own key or whatever, but it should be something that should be default. We don’t have to remind people, especially if you’re the providers of the service. I agree with you on the, you know, very basic principle of why do I pay extra for security, when you should have already covered this for me as part of the service.

Because hey, technically, aren’t you also responsible in this conversation? But the way I see shared responsibility is that—someone on the podcast mentioned it and I think it’s true—shared responsibility means no one’s responsible. And this is the kind of world we’re living in because of that.

Corey: Shared responsibility has always been an odd concept to me because AWS is where I first encountered it and they, from my perspective, turn what fits into a tweet into a 45-minute dog-and-pony show around, “Ah, this is how it works. This is the part we’re responsible for. This is the part where the customer responsibility is. Now, let’s have a mind-numbingly boring conversation around it.” Whereas, yeah, there’s a compression algorithm here. Basically, if the cloud gets breached, it is overwhelmingly likely that you misconfigured something on your end, not the provider doing it, unless it’s Azure, which is neither here nor there, once again.

The problem with that modeling, once you get a little bit more business sophistication than I had the first time I made the observation, is that you can’t sit down with a CISO at a company that just suffered a data breach and have your conversation be, “Doesn’t it suck to be you—[singing] duh, duh—because you messed up. That’s it.” You need that dog-and-pony show of being able to go in-depth and nuance because otherwise, you’re basically calling out your customer, which you can’t really do. Which I feel occludes a lot of clarity for folks who are not in that position who want to understand these things a bit better.

Ashish: You’re right, Corey. I think definitely I don’t want to be in a place where we’re definitely just educating people on this, but I also want to call out that we are in a world where it is true that Amazon, Azure, Google Cloud, they all have vulnerabilities as well. Thanks to research by all these amazing people on the internet from different companies out there, they’ve identified that, hey, these are not pristine environments that you can go into. Azure, AWS, Google Cloud, they themselves have vulnerabilities, and sometimes some of those vulnerabilities cannot be fixed until the customer intervenes and upgrades their services. We do live in a world where there is not enough education about this as well, so I’m glad you brought this up because for people who are listening in, I mean, I was one of those people who would always say, “When was the last time you heard Amazon had a breach?” Or, “Microsoft had a breach?” Or, “Google Cloud had a breach?”

That was the idea when people were just buying into the concept of cloud and did not trust cloud. Every cybersecurity person that I would talk to they’re like, “Why would you trust cloud? Doesn’t make sense.” But this is, like, seven, eight years ago. Fast-forward to today, it’s almost default, “Why would you not go into cloud?”

So, for people who tend to forget that part, I guess, there is definitely a journey that people came through. With the same example of multi-factor authentication, it was never a, “Hey, let’s enable password and multi-factor authentication.” It took a few stages to get there. Same with this as well. We’re at that stage where now cloud service providers are showing the kinks in the armor, and now people are questioning, “I should update my risk matrix for what if there’s actually a breach in AWS?”

Now, Capital One is a great example where the Amazon employee who was sentenced, she did something which has—never even [unintelligible 00:25:32] on before, opened up the door for that [unintelligible 00:25:36] CISO being potentially sentenced. There was another one. Because it became more primetime news, now people are starting to understand, oh, wait. This is not the same as it used to be. Cloud security breaches have evolved as well.

And just sticking to the Uber point, when Uber has that recent breach where they were talking about, “Hey, so many data records were gone,” what a lot of people did not talk about in that same message, it also mentioned the fact that, hey, they also got access to the AWS console of Uber. Now, that to me, is my risk metrics has already gone higher than where it was before because it just not your data, but potentially your production, your pre-prod, any development work that you were doing for, I don’t know, self-driving cars or whatever that Uber [unintelligible 00:26:18] is doing, all that is out on the internet. But who was talking about all of that? That’s a much worse a breach than what was portrayed on the internet. I don’t know, what do you think?

Corey: When it comes to trusting providers, where I sit is that I think, given their scale, they need to be a lot more transparent than they have been historically. However, I also believe that if you do not trust that these companies are telling you the truth about what they’re doing, how they’re doing it, what their controls are, then you should not be using them as a customer, full stop. This idea of confidential computing drives me nuts because so much of it is, “Well, what if we assume our cloud provider is lying to us about all of these things?” Like, hypothetically there’s nothing stopping them from building an exact clone of their entire control plane that they redirect your request to that do something completely different under the hood. “Oh, yeah, of course, we’re encrypting it with that special KMS key.” No, they’re not. For, “Yeah, sure we’re going to put that into this region.” Nope, it goes right back to Virginia. If you believe that’s what’s going on and that they’re willing to do that, you can’t be in cloud.

Ashish: Yeah, a hundred percent. I think foundational trust need to exist and I don’t think the cloud service providers themselves do a great job of building that trust. And maybe that’s where the drift comes in because the business has decided they’re going to cloud. The cyber security people are trying to be more aware and asking the question, “Hey, why do we trust it so blindly? I don’t have a pen test report from Amazon saying they have tested service.”

Yes, I do have a certificate saying it’s PCI compliant, but how do I know—to what you said—they haven’t cloned our services? Fortunately, businesses are getting smarter. Like, Walmart would never have their resources in AWS because they don’t trust them. It’s a business risk if suddenly they decide to go into that space. But the other way around, Microsoft may decides tomorrow that they want to start their own Walmart. Then what do you do?

So, I don’t know how many people actually consider that as a real business risk, especially because there’s a word that was floating around the internet called supercloud. And the idea behind this was—oh, I can already see your reaction [laugh].

Corey: Yeah, don’t get me started on that whole mess.

Ashish: [laugh]. Oh no, I’m the same. I’m like, “What? What now?” Like, “What are you—” So, one thing I took away which I thought was still valuable was the fact that if you look at the cloud service providers, they’re all like octopus, they all have tentacles everywhere.

Like, if you look at the Amazon of the world, they not only a bookstore, they have a grocery store, they have delivery service. So, they are into a lot of industries, the same way Google Cloud, Microsoft, they’re all in multiple industries. And they can still have enough money to choose to go into an industry that they had never been into before because of the access that they would get with all this information that they have, potentially—assuming that they [unintelligible 00:29:14] information. Now, “Shared responsibility,” quote-unquote, they should not do it, but there is nothing stopping them from actually starting a Walmart tomorrow if they wanted to.

Corey: So, because a podcast and a day job aren’t enough, what are you going to be doing in the near future given that, as we record this, re:Invent is nigh?

Ashish: Yeah. So, podcasting and being in the YouTube space has definitely opened up the creative mindset for me. And I think for my producer as well. We’re doing all these exciting projects. We have something called Cloud Security Villains that is coming up for AWS re:Invent, and it’s going to be released on our YouTube channel as well as my social media.

And we’ll have merchandise for it across the re:Invent as well. And I’m just super excited about the possibility that media as a space provides for everyone. So, for people who are listening in and thinking that, I don’t know, I don’t want to write for a blog or email newsletter or whatever the thing may be, I just want to put it out there that I used to be excited about AWS re:Invent just to understand, hey, hopefully, they will release a new security service. Now, I get excited about these events because I get to meet community, help them, share what they have learned on the internet, and sound smarter [laugh] as a result of that as well, and get interviewed where people like yourself. But I definitely find that at the moment with AWS re:Invent coming in, a couple of things that are exciting for me is the release of the Cloud Security Villains, which I think would be an exciting project, especially—hint, hint—for people who are into comic books, you will definitely enjoy it, and I think your kids will as well. So, just in time for Christmas.

Corey: We will definitely keep an eye out for that and put a link to that in the show notes. I really want to thank you for being so generous with your time. If people want to learn more about what you’re up to, where’s the best place for them to find you?

Ashish: I think I’m fortunate enough to be at that stage where normally if people Google me—and it’s simply Ashish Rajan—they will definitely find me [laugh]. I’ll be really hard for them not find me on the internet. But if you are looking for a source of unbiased cloud security knowledge, you can definitely hit up cloudsecuritypodcast.tv or our YouTube and LinkedIn channel.

We go live stream every week with a new guest talking about cloud security, which could be companies like LinkedIn, Twilio, to name a few that have come on the show already, and a lot more than have come in and been generous with their time and shared how they do what they do. And we’re fortunate that we get ranked top 100 in America, US, UK, as well as Australia. I’m really fortunate for that. So, we’re doing something right, so hopefully, you get some value out of it as well when you come and find me.

Corey: And we will, of course, put links to all of that in the show notes. Thank you so much for being so generous with your time. I really appreciate it.

Ashish: Thank you, Corey, for having me. I really appreciate this a lot. I enjoyed the conversation.

Corey: As did I. Ashish Rajan, Principal Cloud Security Advocate at Snyk who is sponsoring this promoted guest episode. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an insulting comment pointing out that not every CISO gets fired; some of them successfully manage to blame the intern.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Clinton

Clinton Herget is Field CTO at Snyk, the leader is Developer Security. He focuses on helping Snyk's strategic customers on their journey to DevSecOps maturity. A seasoned technnologist, Cliton spent his 20-year career prior to Snyk as a web software developer, DevOps consultant, cloud solutions architect, and engineering director. Cluinton is passionate about empowering software engineering to do their best work in the chaotic cloud-native world, and is a frequent conference speaker, developer advocate, and technical thought leader.

Links Referenced:

  • Snyk: https://snyk.io/
  • duckbillgroup.com: https://duckbillgroup.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is brought to us in part by our friends at Pinecone. They believe that all anyone really wants is to be understood, and that includes your users. AI models combined with the Pinecone vector database let your applications understand and act on what your users want… without making them spell it out.

Make your search application find results by meaning instead of just keywords, your personalization system make picks based on relevance instead of just tags, and your security applications match threats by resemblance instead of just regular expressions. Pinecone provides the cloud infrastructure that makes this easy, fast, and scalable. Thanks to my friends at Pinecone for sponsoring this episode. Visit Pinecone.io to understand more.

Corey: This episode is bought to you in part by our friends at Veeam. Do you care about backups? Of course you don’t. Nobody cares about backups. Stop lying to yourselves! You care about restores, usually right after you didn’t care enough about backups. If you’re tired of the vulnerabilities, costs and slow recoveries when using snapshots to restore your data, assuming you even have them at all living in AWS-land, there is an alternative for you. Check out Veeam, thats V-E-E-A-M for secure, zero-fuss AWS backup that won't leave you high and dry when it’s time to restore. Stop taking chances with your data. Talk to Veeam. My thanks to them for sponsoring this ridiculous podcast.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. One of the fun things about establishing traditions is that the first time you do it, you don’t really know that that’s what’s happening. Almost exactly a year ago, I sat down for a previous promoted guest episode much like this one, With Clinton Herget at Snyk—or Synic; however you want to pronounce that. He is apparently a scarecrow of some sorts because when last we spoke, he was a principal solutions engineer, but like any good scarecrow, he was outstanding in his field, and now, as a result, is a Field CTO. Clinton, Thanks for coming back, and let me start by congratulating you on the promotion. Or consoling you depending upon how good or bad it is.

Clinton: You know, Corey, a little bit of column A, a little bit of column B. But very glad to be here again, and frankly, I think it’s because you insist on mispronouncing Snyk as Synic, and so you get me again.

Corey: Yeah, you could add a couple of new letters to it and just call the company [Synack 00:01:27]. Now, it’s a hard pivot to a networking company. So, there’s always options.

Clinton: I acknowledge what you did there, Corey.

Corey: I like that quite a bit. I wasn’t sure you’d get it.

Clinton: I’m a nerd going way, way back, so we’ll have to go pretty deep in the stack for you to stump me on some of this stuff.

Corey: As we did with the, “I wasn’t sure you’d get it.” See that one sailed right past you. And I win. Chalk another one up for me and the networking pun wars. Great, we’ll loop back for that later.

Clinton: I don’t even know where I am right now.

Corey: [laugh]. So, let’s go back to a question that one would think that I’d already established a year ago, but I have the attention span of basically a goldfish, let’s not kid ourselves. So, as I’m visiting the Snyk website, I find that it says different words than it did a year ago, which is generally a sign that is positive; when nothing’s been updated including the copyright date, things are going really well or really badly. One wonders. But no, now you’re talking about Snyk Cloud, you’re talking about several other offerings as well, and my understanding of what it is you folks do no longer appears to be completely accurate. So, let me be direct. What the hell do you folks do over there?

Clinton: It’s a really great question. Glad you asked me on a year later to answer it. I would say at a very high level, what we do hasn’t changed. However, I think the industry has certainly come a long way in the past couple years and our job is to adapt to that Snyk—again, pronounced like a pair of sneakers are sneaking around—it’s a developer security platform. So, we focus on enabling the people who build applications—which as of today, means modern applications built in the cloud—to have better visibility, and ultimately a better chance of mitigating the risk that goes into those applications when it matters most, which is actually in their workflow.

Now, you’re exactly right. Things have certainly expanded in that remit because the job of a software engineer is very different, I think this year than it even was last year, and that’s continually evolving over time. As a developer now, I’m doing a lot more than I was doing a few years ago. And one of the things I’m doing is building infrastructure in the cloud, I’m writing YAML files, I’m writing CloudFormation templates to deploy things out to AWS. And what happens in the cloud has a lot to do with the risk to my organization associated with those applications that I’m building.

So, I’d love to talk a little bit more about why we decided to make that move, but I don’t think that represents a watering down of what we’re trying to do at Snyk. I think it recognizes that developer security vision fundamentally can’t exist without some understanding of what’s happening in the cloud.

Corey: One of the things that always scares me is—and sets the spidey sense tingling—is when I see a company who has a product, and I’m familiar—ish—with what they do. And then they take their product name and slap the word cloud at the end, which is almost always codes to, “Okay, so we took the thing that we sold in boxes in data centers, and now we’re making a shitty hosted version available because it turns out you rubes will absolutely pay a subscription for it.” Yeah, I don’t get the sense that at all is what you’re doing. In fact, I don’t believe that you’re offering a hosted managed service at the moment, are you?

Clinton: No, the cloud part, that fundamentally refers to a new product, an offering that looks at the security or potentially the risks being introduced into cloud infrastructure, by now the engineers who were doing it who are writing infrastructure as code. We previously had an infrastructure-as-code security product, and that served alongside our static analysis tool which is Snyk Code, our open-source tool, our container scanner, recognizing that the kinds of vulnerabilities you can potentially introduce in writing cloud infrastructure are not only bad to the organization on their own—I mean, nobody wants to create an S3 bucket that’s wide open to the world—but also, those misconfigurations can increase the blast radius of other kinds of vulnerabilities in the stack. So, I think what it does is it recognizes that, as you and I think your listeners well know, Corey, there’s no such thing as the cloud, right? The cloud is just a bunch of fancy software designed to abstract away from the fact that you’re running stuff on somebody else’s computer, right?

Corey: Unfortunately, in this case, the fact that you’re calling it Snyk Cloud does not mean that you’re doing what so many other companies in that same space do it would have led to a really short interview because I have no faith that it’s the right path forward, especially for you folks, where it’s, “Oh, you want to be secure? You’ve got to host your stuff on our stuff instead. That’s why we called it cloud.” That’s the direction that I’ve seen a lot of folks try and pivot in, and I always find it disastrous. It’s, “Yeah, well, at Snyk if we run your code or your shitty applications here in our environment, it’s going to be safer than if you run it yourself on something untested like AWS.” And yeah, those stories hold absolutely no water. And may I just say, I’m gratified that’s not what you’re doing?

Clinton: Absolutely not. No, I would say we have no interest in running anyone’s applications. We do want to scan them though, right? We do want to give the developers insight into the potential misconfigurations, the risks, the vulnerabilities that you’re introducing. What sets Snyk apart, I think, from others in that application security testing space is we focus on the experience of the developer, rather than just being another tool that runs and generates a bunch of PDFs and then throws them back to say, “Here’s everything you did wrong.”

We want to say to developers, “Here’s what you could do better. Here’s how that default in a CloudFormation template that leads to your bucket being, you know, wide open on the internet could be changed. Here’s the remediation that you could introduce.” And if we do that at the right moment, which is inside that developer workflow, inside the IDE, on their local machine, before that gets deployed, there’s a much greater chance that remediation is going to be implemented and it’s going to happen much more cheaply, right? Because you no longer have to do the round trip all the way out to the cloud and back.

So, the cloud part of it fundamentally means completing that story, recognizing that once things do get deployed, there’s a lot of valuable context that’s happening out there that a developer can really take advantage of. They can say, “Wait a minute. Not only do I have a Log4Shell vulnerability, right, in one of my open-source dependencies, but that artifact, that application is actually getting deployed to a VPC that has ingress from the internet,” right? So, not only do I have remote code execution in my application, but it’s being put in an enclave that actually allows it to be exploited. You can only know that if you’re actually looking at what’s really happening in the cloud, right?

So, not only does Snyk cloud allows us to provide an additional layer of security by looking at what’s misconfigured in that cloud environment and help your developers make remediations by saying, “Here’s the actual IAC file that caused that infrastructure to come into existence,” but we can also say, here’s how that affects the risk of other kinds of vulnerabilities at different layers in the stack, right? Because it’s all software; it’s all connected. Very rarely does a vulnerability translate one-to-one into risk, right? They’re compound because modern software is compound. And I think what developers lack is the tooling that fits into their workflow that understands what it means to be a software engineer and actually helps them make better choices rather than punishing them after the fact for guessing and making bad ones.

Corey: That sounds awesome at a very high level. It is very aligned with how executives and decision-makers think about a lot of these things. Let’s get down to brass tacks for a second. Assume that I am the type of developer that I am in real life, by which I mean shitty. What am I going to wind up attempting to do that Snyk will flag and, in other words, protect me from myself and warn me that I’m about to commit a dumb?

Clinton: First of all, I would say, look, there’s no such thing as a non-shitty developer, right? And I built software for 20 years and I decided that’s really hard. What’s a lot easier is talking about building software for a living. So, that’s what I do now. But fundamentally, the reason I’m at Snyk, is I want to help people who are in the kinds of jobs that I had for a very long time, which is to say, you have a tremendous amount of anxiety because you recognize that the success of the organization rests on your shoulders, and you’re making hundreds, if not thousands of decisions every day without the right context to understand fully how the results of that decision is going to affect the organization that you work for.

So, I think every developer in the world has to deal with this constant cognitive dissonance of saying, “I don’t know that this is right, but I have to do it anyway because I need to clear that ticket because that release needs to get into production.” And it becomes really easy to short-sightedly do things like pull an open-source dependency without checking whether it has any CVEs associated with it because that’s the version that’s easiest to implement with your code that already exists. So, that’s one piece. Snyk Open Source, designed to traverse that entire tree of dependencies in open-source all the way down, all the hundreds and thousands of packages that you’re pulling in to say, not only, here’s a vulnerability that you should really know is going to end up in your application when it’s built, but also here’s what you can do about it, right? Here’s the upgrade you can make, here’s the minimum viable change that actually gets you out of this problem, and to do so when it’s in the right context, which is in you know, as you’re making that decision for the first time, right, inside your developer environment.

That also applies to things like container vulnerabilities, right? I have even less visibility into what’s happening inside a container than I do inside my application. Because I know, say, I’m using an Ubuntu or a Red Hat base image. I have no idea, what are all the Linux packages that are on it, let alone what are the vulnerabilities associated with them, right? So, being able to detect, I’ve got a version of OpenSSL 3.0 that has a potentially serious vulnerability associated with it before I’ve actually deployed that container out into the cloud very much helps me as a developer.

Because I’m limiting the rework or the refactoring I would have to do by otherwise assuming I’m making a safe choice or guessing at it, and then only finding out after I’ve written a bunch more code that relies on that decision, that I have to go back and change it, and then rewrite all of the things that I wrote on top of it, right? So, it’s the identifying the layer in the stack where that risk could be introduced, and then also seeing how it’s affected by all of those other layers because modern software is inherently complex. And that complexity is what drives both the risk associated with it, and also things like efficiency, which I know your audience is, for good reason, very concerned about.

Corey: I’m going to challenge you on aspect of this because on the tin, the way you describe it, it sounds like, “Oh, I already have something that does that. It’s the GitHub Dependabot story where it winds up sending me a litany of complaints every week.” And we are talking, if I did nothing other than read this email in that day, that would be a tremendously efficient processing of that entire thing because so much of it is stuff that is ancient and archived, and specific aspects of the vulnerabilities are just not relevant. And you talk about the OpenSSL 3.0 issues that just recently came out.

I have no doubt that somewhere in the most recent email I’ve gotten from that thing, it’s buried two-thirds of the way down, like all the complaints like the dishwasher isn’t loaded, you forgot to take the trash out, that baby needs a change, the kitchen is on fire, and the vacuuming, and the r—wait, wait. What was that thing about the kitchen? Seems like one of those things is not like the others. And it just gets lost in the noise. Now, I will admit to putting my thumb a little bit on the scale here because I’ve used Snyk before myself and I know that you don’t do that. How do you avoid that trap?

Clinton: Great question. And I think really, the key to the story here is, developers need to be able to prioritize, and in order to prioritize effectively, you need to understand the context of what happens to that application after it gets deployed. And so, this is a key part of why getting the data out of the cloud and bringing it back into the code is so important. So, for example, take an OpenSSL vulnerability. Do you have it on a container image you’re using, right? So, that’s question number one.

Question two is, is there actually a way that code can be accessed from the outside? Is it included or is it called? Is the method activated by some other package that you have running on that container? Is that container image actually used in a production deployment? Or does it just go sit in a registry and no one ever touches it?

What are the conditions required to make that vulnerability exploitable? You look at something like Spring Shell, for example, yes, you need a certain version of spring-beans in a JAR file somewhere, but you also need to be running a certain version of Tomcat, and you need to be packaging those JARs inside a WAR in a certain way.

Corey: Exactly. I have a whole bunch of Lambda functions that provide the pipeline system that I use to build my newsletter every week, and I get screaming concerns about issues in, for example, a version of the markdown parser that I’ve subverted. Yeah, sure. I get that, on some level, if I were just giving it random untrusted input from the internet and random ad hoc users, but I’m not. It’s just me when I write things for that particular Lambda function.

And I’m not going to be actively attempting to subvert the thing that I built myself and no one else should have access to. And looking through the details of some of these things, it doesn’t even apply to the way that I’m calling the libraries, so it’s just noise, for lack of a better term. It is not something that basically ever needs to be adjusted or fixed.

Clinton: Exactly. And I think cutting through that noise is so key to creating developer trust in any kind of tool that scanning an asset and providing you what, in theory, are a list of actionable steps, right? I need to be able to understand what is the thing, first of all. There’s a lot of tools that do that, right, and we tend to mock them by saying things like, “Oh, it’s just another PDF generator. It’s just another thousand pages that you’re never going to read.”

So, getting the information in the right place is a big part of it, but filtering out all of the noise by saying, we looked at not just one layer of the stack, but multiple layers, right? We know that you’re using this open-source dependency and we also know that the method that contains the vulnerability is actively called by your application in your first-party code because we ran our static analysis tool against that. Furthermore, we know because we looked at your cloud context, we connected to your AWS API—we’re big partners with AWS and very proud of that relationship—but we can tell that there’s inbound internet access available to that service, right? So, you start to build a compound case that maybe this is something that should be prioritized, right? Because there’s a way into the asset from the outside world, there’s a way into the vulnerable functions through the labyrinthine, you know, spaghetti of my code to get there, and the conditions required to exploit it actually exist in the wild.

But you can’t just run a single tool; you can’t just run Dependabot to get that prioritization. You actually have to look at the entire holistic application context, which includes not just your dependencies, but what’s happening in the container, what’s happening in your first-party, your proprietary code, what’s happening in your IAC, and I think most importantly for modern applications, what’s actually happening in the cloud once it gets deployed, right? And that’s sort of the holy grail of completing that loop to bring the right context back from the cloud into code to understand what change needs to be made, and where, and most importantly why. Because it’s a priority that actually translates into organizational risk to get a developer to pay attention, right? I mean, that is the key to I think any security concern is how do you get engineering mindshare and trust that this is actually what you should be paying attention to and not a bunch of rework that doesn’t actually make your software more secure?

Corey: One of the challenges that I see across the board is that—well, let’s back up a bit here. I have in previous episodes talked in some depth about my position that when it comes to the security of various cloud providers, Google is number one, and AWS is number two. Azure is a distant third because it figures out what Crayons tastes the best; I don’t know. But the reason is not because of any inherent attribute of their security models, but rather that Google massively simplifies an awful lot of what happens. It automatically assumes that resources in the same project should be able to talk to one another, so I don’t have to painstakingly configure that.

In AWS-land, all of this must be done explicitly; no one has time for that, so we over-scope permissions massively and never go back and rein them in. It’s a configuration vulnerability more than an underlying inherent weakness of the platform. Because complexity is the enemy of security in many respects. If you can’t fit it all in your head to reason about it, how can you understand the security ramifications of it? AWS offers a tremendous number of security services. Many of them, when taken in some totality of their pricing, cost more than any breach, they could be expected to prevent. Adding more stuff that adds more complexity in the form of Snyk sounds like it’s the exact opposite of what I would want to do. Change my mind.

Clinton: I would love to. I would say, fundamentally, I think you and I—and by ‘I,’ I mean Snyk and you know, Corey Quinn Enterprises Limited—I think we fundamentally have the same enemy here, right, which is the cyclomatic complexity of software, right, which is how many different pathways do the bits have to travel down to reach the same endpoint, right, the same goal. The more pathways there are, the more risk is introduced into your software, and the more inefficiency is introduced, right? And then I know you’d love to talk about how many different ways is there to run a container on AWS, right? It’s either 30 or 400 or eleventy-million.

I think you’re exactly right that that complexity, it is great for, first of all, selling cloud resources, but also, I think, for innovating, right, for building new kinds of technology on top of that platform. The cost that comes along with that is a lack of visibility. And I think we are just now, as we approach the end of 2022 here, coming to recognize that fundamentally, the complexity of modern software is beyond the ability of a single engineer to understand. And that is really important from a security perspective, from a cost control perspective, especially because software now creates its own infrastructure, right? You can’t just now secure the artifact and secure the perimeter that it gets deployed into and say, “I’ve done my job. Nobody can breach the perimeter and there’s no vulnerabilities in the thing because we scanned it and that thing is immutable forever because it’s pets, not cattle.”

Where I think the complexity story comes in is to recognize like, “Hey, I’m deploying this based on a quickstart or CloudFormation template that is making certain assumptions that make my job easier,” right, in a very similar way that choosing an open-source dependency makes my job easier as a developer because I don’t have to write all of that code myself. But what it does mean is I lack the visibility into, well hold on. How many different pathways are there for getting things done inside this dependency? How many other dependencies are brought on board? In the same way that when I create an EKS cluster, for example, from a CloudFormation template, what is it creating in the background? How many VPCs are involved? What are the subnets, right? How are they connected to each other? Where are the potential ingress points?

So, I think fundamentally, getting visibility into that complexity is step number one, but understanding those pathways and how they could potentially translate into risk is critically important. But that prioritization has to involve looking at the software holistically and not just individual layers, right? I think we lose when we say, “We ran a static analysis tool and an open-source dependency scanner and a container scanner and a cloud config checker, and they all came up green, therefore the software doesn’t have any risks,” right? That ignores the fundamental complexity in that all of these layers are connected together. And from an adversaries perspective, if my job is to go in and exploit software that’s hosted in the cloud, I absolutely do not see the application model that way.

I see it as it is inherently complex and that’s a good thing for me because it means I can rely on the fact that those engineers had tremendous anxiety, we’re making a lot of guesses, and crossing their fingers and hoping something would work and not be exploitable by me, right? So, the only way I think we get around that is to recognize that our engineers are critical stakeholders in that security process and you fundamentally lack that visibility if you don’t do your scanning until after the fact. If you take that traditional audit-based approach that assumes a very waterfall, legacy approach to building software, and recognize that, hey, we’re all on this infinite loop race track now. We’re deploying every three-and-a-half seconds, everything’s automated, it’s all built at scale, but the ability to do that inherently implies all of this additional complexity that ultimately will, you know, end up haunting me, right? If I don’t do anything about it, to make my engineer stakeholders in, you know, what actually gets deployed and what risks it brings on board.

Corey: This episode is sponsored in part by our friends at Uptycs. Attackers don’t think in silos, so why would you have siloed solutions protecting cloud, containers, and laptops distinctly? Meet Uptycs - the first unified solution that prioritizes risk across your modern attack surface—all from a single platform, UI, and data model. Stop by booth 3352 at AWS re:Invent in Las Vegas to see for yourself and visit uptycs.com. That’s U-P-T-Y-C-S.com. My thanks to them for sponsoring my ridiculous nonsense.

Corey: When I wind up hearing you talk about this—I’m going to divert us a little bit because you’re dancing around something that it took me a long time to learn. When I first started fixing AWS bills for a living, I thought that it would be mostly math, by which I mean arithmetic. That’s the great secret of cloud economics. It’s addition, subtraction, and occasionally multiplication and division. No, turns out it’s much more psychology than it is math. You’re talking in many aspects about, I guess, what I’d call the psychology of a modern cloud engineer and how they think about these things. It’s not a technology problem. It’s a people problem, isn’t it?

Clinton: Oh, absolutely. I think it’s the people that create the technology. And I think the longer you persist in what we would call the legacy viewpoint, right, not recognizing what the cloud is—which is fundamentally just software all the way down, right? It is abstraction layers that allow you to ignore the fact that you’re running stuff on somebody else’s computer—once you recognize that, you realize, oh, if it’s all software, then the problems that it introduces are software problems that need software solutions, which means that it must involve activity by the people who write software, right? So, now that you’re in that developer world, it unlocks, I think, a lot of potential to say, well, why don’t developers tend to trust the security tools they’ve been provided with, right?

I think a lot of it comes down to the question you asked earlier in terms of the noise, the lack of understanding of how those pieces are connected together, or the lack of context, or not even frankly, caring about looking beyond the single-point solution of the problem that solution was designed to solve. But more importantly than that, not recognizing what it’s like to build modern software, right, all of the decisions that have to be made on a daily basis with very limited information, right? I might not even understand where that container image I’m building is going in the universe, let alone what’s being built on top of it and how much critical customer data is being touched by the database, that that container now has the credentials to access, right? So, I think in order to change anything, we have to back way up and say, problems in the cloud or software problems and we have to treat them that way.

Because if we don’t if we continue to represent the cloud as some evolution of the old environment where you just have this perimeter that’s pre-existing infrastructure that you’re deploying things onto, and there’s a guy with a neckbeard in the basement who is unplugging cables from a switch and plugging them back in and that’s how networking problems are solved, I think you missed the idea that all of these abstraction layers introduced the very complexity that needs to be solved back in the build space. But that requires visibility into what actually happens when it gets deployed. The way I tend to think of it is, there’s this firewall in place. Everybody wants to say, you know, we’re doing DevOps or we’re doing DevSecOps, right? And that’s a lie a hundred percent of the time, right? No one is actually, I think, adhering completely to those principles.

Corey: That’s why one of the core tenets of ClickOps is lying about doing anything in the console.

Clinton: Absolutely, right? And that’s why shadow IT becomes more and more prevalent the deeper you get into modern development, not less and less prevalent because it’s fundamentally hard to recognize the entirety of the potential implications, right, of a decision that you’re making. So, it’s a lot easier to just go in the console and say, “Okay, I’m going to deploy one EC2 to do this. I’m going to get it right at some point.” And that’s why every application that’s ever been produced by human hands has a comment in it that says something like, “I don’t know why this works but it does. Please don’t change it.”

And then three years later because that developer has moved on to another job, someone else comes along and looks at that comment and says, “That should really work. I’m going to change it.” And they do and everything fails, and they have to go back and fix it the original way and then add another comment saying, “Hey, this person above me, they were right. Please don’t change this line.” I think every engineer listening right now knows exactly where that weak spot is in the applications that they’ve written and they’re terrified of that.

And I think any tool that’s designed to help developers fundamentally has to get into the mindset, get into the psychology of what that is, like, of not fundamentally being able to understand what those applications are doing all of the time, but having to write code against them anyway, right? And that’s what leads to, I think, the fear that you’re going to get woken up because your pager is going to go off at 3 a.m. because the building is literally on fire and it’s because of code that you wrote. We have to solve that problem and it has to be those people who’s psychology we get into to understand, how are you working and how can we make your life better, right? And I really do think it comes with that the noise reduction, the understanding of complexity, and really just being humble and saying, like, “We get that this job is really hard and that the only way it gets better is to begin admitting that to each other.”

Corey: I really wish that there were a better way to articulate a lot of these things. This the reason that I started doing a security newsletter; it’s because cost and security are deeply aligned in a few ways. One of them is that you care about them a lot right after you failed to care about them sufficiently, but the other is that you’ve got to build guardrails in such a way that doing the right thing is easier than doing it the wrong way, or you’re never going to gain any traction.

Clinton: I think that’s absolutely right. And you use the key term there, which is guardrails. And I think that’s where in their heart of hearts, that’s where every security professional wants to be, right? They want to be defining policy, they want to be understanding the risk posture of the organization and nudging it in a better direction, right? They want to be talking up to the board, to the executive team, and creating confidence in that risk posture, rather than talking down or off to the side—depending on how that org chart looks—to the engineers and saying, “Fix this, fix that, and then fix this other thing.” A, B, and C, right?

I think the problem is that everyone in a security role or an organization of any size at this point, is doing 90% of the latter and only about 10% of the former, right? They’re acting as gatekeepers, not as guardrails. They’re not defining policy, they’re spending all of their time creating Jira tickets and all of their time tracking down who owns the piece of code that got deployed to this pod on EKS that’s throwing all these errors on my console, and how can I get the person to make a decision to actually take an action that stops these notifications from happening, right? So, all they’re doing is throwing footballs down the field without knowing if there’s a receiver there, right, and I think that takes away from the job that our security analysts really shouldn’t be doing, which is creating those guardrails, which is having confidence that the policy they set is readily understood by the developers making decisions, and that’s happening in an automated way without them having to create friction by bothering people all the time. I don’t think security people want to be [laugh] hated by the development teams that they work with, but they are. And the reason they are is I think, fundamentally, we lack the tooling, we lack—

Corey: They are the barrier method.

Clinton: Exactly. And we lacked the processes to get the right intelligence in a way that’s consumable by the engineers when they’re doing their job, and not after the fact, which is typically when the security people have done their jobs.

Corey: It’s sad but true. I wish that there were a better way to address these things, and yet here we are.

Clinton: If only there were better way to address these things.

Corey: [laugh].

Clinton: Look, I wouldn’t be here at Snyk if I didn’t think there were a better way, and I wouldn’t be coming on shows like yours to talk to the engineering communities, right, people who have walked the walk, right, who have built those Terraform files that contain these misconfigurations, not because they’re bad people or because they’re lazy, or because they don’t do their jobs well, but because they lacked the visibility, they didn’t have the understanding that that default is actually insecure. Because how would I know that otherwise, right? I’m building software; I don’t see myself as an expert on infrastructure, right, or on Linux packages or on cyclomatic complexity or on any of these other things. I’m just trying to stay in my lane and do my job. It’s not my fault that the software has become too complex for me to understand, right?

But my management doesn’t understand that and so I constantly have white knuckles worrying that, you know, the next breach is going to be my fault. So, I think the way forward really has to be, how do we make our developers stakeholders in the risk being introduced by the software they write to the organization? And that means everything we’ve been talking about: it means prioritization; it means understanding how the different layers of the stack affect each other, especially the cloud pieces; it means an extensible platform that lets me write code against it to inject my own reasoning, right? The piece that we haven’t talked about here is that risk calculation doesn’t just involve technical aspects, there’s also business intelligence that’s involved, right? What are my critical applications, right, what actually causes me to lose significant amounts of money if those services go offline?

We at Snyk can’t tell that. We can’t run a scanner to say these are your crown jewel services that can’t ever go down, but you can know that as an organization. So, where we’re going with the platform is opening up the extensible process, creating APIs for you to be able to affect that risk triage, right, so that as the creators have guardrails as the security team, you are saying, “Here’s how we want our developers to prioritize. Here are all of the factors that go into that decision-making.” And then you can be confident that in their environment, back over in developer-land, when I’m looking at IntelliJ, or, you know, or on my local command line, I am seeing the guardrails that my security team has set for me and I am confident that I’m fixing the right thing, and frankly, I’m grateful because I’m fixing it at the right time and I’m doing it in such a way and with a toolset that actually is helping me fix it rather than just telling me I’ve done something wrong, right, because everything we do at Snyk focuses on identifying the solution, not necessarily identifying the problem.

It’s great to know that I’ve got an unencrypted S3 bucket, but it’s a whole lot better if you give me the line of code and tell me exactly where I have to copy and paste it so I can go on to the next thing, rather than spending an hour trying to figure out, you know, where I put that line and what I actually have to change it to, right? I often say that the most valuable currency for a developer, for a software engineer, it’s not money, it’s not time, it’s not compute power or anything like that, it’s the right context, right? I actually have to understand what are the implications of the decision that I’m making, and I need that to be in my own environment, not after the fact because that’s what creates friction within an organization is when I could have known earlier and I could have known better, but instead, I had to guess I had to write a bunch of code that relies on the thing that was wrong, and now I have to redo it all for no good reason other than the tooling just hadn’t adapted to the way modern software is built.

Corey: So, one last question before we wind up calling it a day here. We are now heavily into what I will term pre:Invent where we’re starting to see a whole bunch of announcements come out of the AWS universe in preparation for what I’m calling Crappy Cloud Hanukkah this year because I’m spending eight nights in Las Vegas. What are you doing these days with AWS specifically? I know I keep seeing your name in conjunction with their announcements, so there’s something going on over there.

Clinton: Absolutely. No, we’re extremely excited about the partnership between Snyk and AWS. Our vulnerability intelligence is utilized as one of the data sources for AWS Inspector, particularly around open-source packages. We’re doing a lot of work around things like the code suite, building Snyk into code pipeline, for example, to give developers using that code suite earlier visibility into those vulnerabilities. And really, I think the story kind of expands from there, right?

So, we’re moving forward with Amazon, recognizing that it is, you know, sort of the de facto. When we say cloud, very often we mean AWS. So, we’re going to have a tremendous presence at re:Invent this year, I’m going to be there as well. I think we’re actually going to have a bunch of handouts with your face on them is my understanding. So, please stop by the booth; would love to talk to folks, especially because we’ve now released the Snyk Cloud product and really completed that story. So, anything we can do to talk about how that additional context of the cloud helps engineers because it’s all software all the way down, those are absolutely conversations we want to be having.

Corey: Excellent. And we will, of course, put links to all of these things in the [show notes 00:35:00] so people can simply click, and there they are. Thank you so much for taking all this time to speak with me. I appreciate it.

Clinton: All right. Thank you so much, Corey. Hope to do it again next year.

Corey: Clinton Herget, Field CTO at Snyk. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an angry comment telling me that I’m being completely unfair to Azure, along with your favorite tasting color of Crayon.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Chen

Chen Goldberg is GM and Vice President of Engineering at Google Cloud, where she leads the Cloud Runtimes (CR) product area, helping customers deliver greater value, effortlessly. The CR portfolio includes both Serverless and Kubernetes based platforms on Google Cloud, private cloud and other public clouds. Chen is a strong advocate for customer empathy, building products and solutions that matter. Chen has been core to Google Cloud’s open core vision since she joined the company six years ago. During that time, she has led her team to focus on helping development teams increase their agility and modernize workloads. Prior to joining Google, Chen wore different hats in the tech industry including leadership positions in IT organizations, SI teams and SW product development, contributing to Chen’s broad enterprise perspective. She enjoys mentoring IT talent both in and outside of Google. Chen lives in Mountain View, California, with her husband and three kids. Outside of work she enjoys hiking and baking.

Links Referenced:

  • Twitter: https://twitter.com/GoldbergChen
  • LinkedIn: https://www.linkedin.com/in/goldbergchen/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Forget everything you know about SSH and try Tailscale. Imagine if you didn't need to manage PKI or rotate SSH keys every time someone leaves. That'd be pretty sweet, wouldn't it? With Tailscale SSH, you can do exactly that. Tailscale gives each server and user device a node key to connect to its VPN, and it uses the same node key to authorize and authenticate SSH.

Basically you're SSHing the same way you manage access to your app. What's the benefit here? Built-in key rotation, permissions as code, connectivity between any two devices, reduce latency, and there's a lot more, but there's a time limit here. You can also ask users to reauthenticate for that extra bit of security. Sounds expensive?

Nope, I wish it were. Tailscale is completely free for personal use on up to 20 devices. To learn more, visit snark.cloud/tailscale. Again, that's snark.cloud/tailscale

Corey: Welcome to Screaming in the Cloud, I’m Corey Quinn. When I get bored and the power goes out, I find myself staring at the ceiling, figuring out how best to pick fights with people on the internet about Kubernetes. Because, well, I’m basically sad and have a growing collection of personality issues. My guest today is probably one of the best people to have those arguments with. Chen Goldberg is the General Manager of Cloud Runtimes and VP of Engineering at Google Cloud. Chen, Thank you for joining me today.

Chen: Thank you so much, Corey, for having me.

Corey: So, Google has been doing a lot of very interesting things in the cloud, and the more astute listener will realize that interesting is not always necessarily a compliment. But from where I sit, I am deeply vested in the idea of a future where we do not have a cloud monoculture. As I’ve often said, I want, “What cloud should I build something on in five to ten years?” To be a hard question to answer, and not just because everything is terrible. I think that Google Cloud is absolutely a bright light in the cloud ecosystem and has been for a while, particularly with this emphasis around developer experience. All of that said, Google Cloud is sort of a big, unknowable place, at least from the outside. What is your area of responsibility? Where do you start? Where do you stop? In other words, what can I blame you for?

Chen: Oh, you can blame me for a lot of things if you want to. I [laugh] might not agree with that, but that’s—

Corey: We strive for accuracy in these things, though.

Chen: But that’s fine. Well, first of all, I’ve joined Google about seven years ago to lead the Kubernetes and GKE team, and ever since, continued at the same area. So evolved, of course, Kubernetes, and Google Kubernetes Engine, and leading our hybrid and multi-cloud strategy as well with technologies like Anthos. And now I’m responsible for the entire container runtime, which includes Kubernetes and the serverless solutions.

Corey: A while back, I, in fairly typical sarcastic form, wound up doing a whole inadvertent start of a meme where I joked about there being 17 ways to run containers on AWS. And then as that caught on, I wound up listing out 17 services you could use to do that. A few months went past and then I published a sequel of 17 more services you can use to run Kubernetes. And while that was admittedly tongue-in-cheek, it does lead to an interesting question that’s ecosystem-wide. If I look at Google Cloud, I have Cloud Run, I have GKE, I have GCE if I want to do some work myself.

It feels like more and more services are supporting Docker in a variety of different ways. How should customers and/or people like me—though, I am sort of a customer as well since I do pay you folks every month—how should we think about containers and services in which to run them?

Chen: First of all, I think there’s a lot of credit that needs to go to Docker that made containers approachable. And so, Google has been running containers forever. Everything within Google is running on containers, even our VMs, even our cloud is running on containers, but what Docker did was creating a packaging mechanism to improve developer velocity. So, that’s on its own, it’s great. And one of the things, by the way, that I love about Google Cloud approach to containers and Docker that yes, you can take your Docker container and run it anywhere.

And it’s actually really important to ensure what we call interoperability, or low barrier to entry to a new technology. So, I can take my Docker container, I can move it from one platform to another, and so on. So, that’s just to start with on a containers. Between the different solutions, so first of all, I’m all about managed services. You are right, there are many ways to run a Kubernetes. I’m taking a lot of pride—

Corey: The best way is always to have someone else run it for you. Problem solved. Great, the best kind of problems are always someone else’s.

Chen: Yes. And I’m taking a lot of pride of what our team is doing with Kubernetes. I mean, we’ve been working on that for so long. And it’s something that you know, we’ve coined that term, I think back in 2016, so there is a success disaster, but there’s also what we call sustainable success. So, thinking about how to set ourselves up for success and scale. Very proud of that service.

Saying that, not everybody and not all your workloads you need the flexibility that Kubernetes gives you in all the ecosystem. So, if you start with containers your first time, you should start with Cloud Run. It’s the easiest way to run your containers. That’s one. If you are already in love with Kubernetes, we won’t take it away from you. Start with GKE. Okay [laugh]? Go all-in. Okay, we are all in loving Kubernetes as well. But what my team and I are working on is to make sure that those will work really well together. And we actually see a lot of customers do that.

Corey: I’d like to go back a little bit in history to the rise of Docker. I agree with you it was transformative, but containers had been around in various forms—depending upon how you want to define it—dating back to the ’70s with logical partitions on mainframes. Well, is that a container? Is it not? Well, sort of. We’ll assume yes for the sake of argument.

The revelation that I found from Docker was the developer experience, start to finish. Suddenly, it was a couple commands and you were just working, where previously it had taken tremendous amounts of time and energy to get containers working in that same context. And I don’t even know today whether or not the right way to contextualize containers is as sort of a lite version of a VM, as a packaging format, as a number of other things that you could reasonably call it. How do you think about containers?

Chen: So, I’m going to do, first of all, a small [unintelligible 00:06:31]. I actually started my career as a system mainframe engineer—

Corey: Hmm.

Chen: And I will share that when you know, I’ve learned Kubernetes, I’m like, “Huh, we already have done all of that, in orchestration, in workload management on mainframe,” just to the side. The way I think about containers is as a—two things: one, it is a packaging of an application, but the other thing which is also critical is the decoupling between your application and the OS. So, having that kind of abstraction and allowing you to portable and move it between environments. So, those are the two things that are when I think about containers. And what technologies like Kubernetes and serverless gives on top of that is that manageability and making sure that we take care of everything else that is needed for you to run your application.

Corey: I’ve been, how do I put this, getting some grief over the past few years, in the best ways possible, around a almost off-the-cuff prediction that I made, which was that in five years, which is now a lot closer to two, basically, nobody is going to care about Kubernetes. And I could have phrased that slightly more directly because people think I was trying to say, “Oh, Kubernetes is just hype. It’s going to go away. Nobody’s going to worry about it anymore.” And I think that is a wildly inaccurate prediction.

My argument is that people are not going to have to think about it in the same way that they are today. Today, if I go out and want to go back to my days of running production services in anger—and by ‘anger,’ I of course mean in production—then it would be difficult for me to find a role that did not at least touch upon Kubernetes. But people who can work with that technology effectively are in high demand and they tend to be expensive, not to mention then thinking about all of the intricacies and complexities that Kubernetes brings to the foreground, that is what doesn’t feel sustainable to me. The idea that it’s going to have to collapse down into something else is, by necessity, going to have to emerge. How are you seeing that play out? And also, feel free to disagree with the prediction. I am thrilled to wind up being told that I’m wrong it’s how I learn the most.

Chen: I don’t know if I agree with the time horizon of when that will happen, but I will actually think it’s a failure on us if that won’t be the truth, that the majority of people will not need to know about Kubernetes and its internals. And you know, we keep saying that, like, hey, we need to make it more, like, boring, and easy, and I’ve just said like, “Hey, you should use managed.” And we have lots of customers that says that they’re just using GKE and it scales on their behalf and they don’t need to do anything for that and it’s just like magic. But from a technology perspective, there is still a way to go until we can make that disappear.

And there will be two things that will push us into that direction. One is—you mentioned that is as well—the talent shortage is real. All the customers that I speak with, even if they can find those great people that are experts, they’re actually more interesting things for them to work on, okay? You don’t need to take, like, all the people in your organization and put them on building the infrastructure. You don’t care about that. You want to build innovation and promote your business.

So, that’s one. The second thing is that I do expect that the technology will continue to evolve and are managed solutions will be better and better. So hopefully, with these two things happening together, people will not care that what’s under the hood is Kubernetes. Or maybe not even, right? I don’t know exactly how things will evolve.

Corey: From where I sit, what are the early criticisms I had about Docker, which I guess translates pretty well to Kubernetes, are that they solve a few extraordinarily painful problems. In the case of Docker, it was, “Well, it works on my machine,” as a grumpy sysadmin, the way I used to be, the only real response we had to that was, “Well. Time to backup your email, Skippy, because your laptop is going into production, then.” Now, you can effectively have a high-fidelity copy of production, basically anywhere, and we’ve solved the problem of making your Mac laptop look like a Linux server. Great, okay, awesome.

With Kubernetes, it also feels, on some level, like it solves for very large-scale Google-type of problems where you want to run things across at least a certain point of scale. It feels like even today, it suffers from having an easy Hello World-style application to deploy on top of it. Using it for WordPress, or some other form of blogging software, for example, is stupendous overkill as far as the Hello World story tends to go. Increasingly as a result, it feels like it’s great for the large-scale enterprise-y applications, but the getting started story of how do I have a service I could reasonably run in production? How do I contextualize that, in the world of Kubernetes? How do you respond to that type of perspective?

Chen: We’ll start with maybe a short story. I started my career in the Israeli army. I was head of the department and one of the lead technology units and I was responsible for building a PAS. In essence, it was 20-plus years ago, so we didn’t really call it a PAS but that’s what it was. And then at some point, it was amazing, developers were very productive, we got innovation again, again. And then there was some new innovation just at the beginning of web [laugh] at some point.

And it was actually—so two things I’ve noticed back then. One, it was really hard to evolve the platform to allow new technologies and innovation, and second thing, from a developer perspective, it was like a black box. So, the developers team that people were—the other development teams couldn’t really troubleshoot environment; they were not empowered to make decisions or [unintelligible 00:12:29] in the platform. And you know, when it was just started with Kubernetes—by the way, beginning, it only supported 100 nodes, and then 1000 nodes. Okay, it was actually not for scale; it actually solved those two problems, which I’m—this is where I spend most of my time.

So, the first one, we don’t want magic, okay? To be clear on, like, what’s happening, I want to make sure that things are consistent and I can get the right observability. So, that’s one. The second thing is that we invested so much in the extensibility an environment that it’s, I wouldn’t say it’s easy, but it’s doable to evolve Kubernetes. You can change the models, you can extend it you can—there is an ecosystem.

And you know, when we were building it, I remember I used to tell my team, there won’t be a Kubernetes 2.0. Which is for a developer, it’s [laugh] frightening. But if you think about it and you prepare for that, you’re like, “Huh. Okay, what does that mean with how I build my APIs? What does that mean of how we build a system?” So, that was one. The second thing I keep telling my team, “Please don’t get too attached to your code because if it will still be there in 5, 10 years, we did something wrong.”

And you can see areas within Kubernetes, again, all the extensions. I'm very proud of all the interfaces that we’ve built, but let’s take networking. This keeps to evolve all the time on the API and the surface area that allows us to introduce new technologies. I love it. So, those are the two things that have nothing to do with scale, are unique to Kubernetes, and I think are very empowering, and are critical for the success.

Corey: One thing that you said that resonates most deeply with me is the idea that you don’t want there to be magic, where I just hand it to this thing and it runs it as if by magic. Because, again, we’ve all run things in anger in production, and what happens when the magic breaks? When you’re sitting around scratching your head with no idea how it starts or how it stops, that is scary. I mean, I recently wound up re-implementing Google Cloud Distinguished Engineer Kelsey Hightower’s “Kubernetes the Hard Way” because he gave a terrific tutorial that I ran through in about 45 minutes on top of Google Cloud. It’s like, “All right, how do I make this harder?”

And the answer is to do it on AWS, re-implement it there. And my experiment there can be found at kubernetesthemuchharderway.com because I have a vanity domain problem. And it taught me he an awful lot, but one of the challenges I had as I went through that process was, at one point, the nodes were not registering with the controller.

And I ran out of time that day and turned everything off—because surprise bills are kind of what I spend my time worrying about—turn it on the next morning to continue and then it just worked. And that was sort of the spidey sense tingling moment of, “Okay, something wasn’t working and now it is, and I don’t understand why. But I just rebooted it and it started working.” Which is terrifying in the context of a production service. It was understandable—kind of—and I think that’s the sort of thing that you understand a lot better, the more you work with it in production, but a counterargument to that is—and I’ve talked about it on this show before—for this podcast, I wind up having sponsors from time to time, who want to give me fairly complicated links to go check them out, so I have the snark.cloud URL redirector.

That’s running as a production service on top of Google Cloud Run. It took me half an hour to get that thing up and running; I haven’t had to think about it since, aside from a three-second latency that was driving me nuts and turned out to be a sleep hidden in the code, which I can’t really fault Google Cloud Run for so much as my crappy nonsense. But it just works. It’s clearly running atop Kubernetes, but I don’t have to think about it. That feels like the future. It feels like it’s a glimpse of a world to come, we’re just starting to dip our toes into. That, at least to me, feels like a lot more of the abstractions being collapsed into something easily understandable.

Chen: [unintelligible 00:16:30], I’m happy you say that. When talking with customers and we’re showing, like, you know, yes, they’re all in Kubernetes and talking about Cloud Run and serverless, I feel there is that confidence level that they need to overcome. And that’s why it’s really important for us in Google Cloud is to make sure that you can mix and match. Because sometimes, you know, a big retail customer of ours, some of their teams, it’s really important for them to use a Kubernetes-based platform because they have their workloads also running on-prem and they want to serve the same playbooks, for example, right? How do I address issues, how do I troubleshoot, and so on?

So, that’s one set of things. But some cloud only as simple as possible. So, can I use both of them and still have a similar developer experience, and so on? So, I do think that we’ll see more of that in the coming years. And as the technology evolves, then we’ll have more and more, of course, serverless solutions.

By the way, it doesn’t end there. Like, we see also, you know, databases and machine learning, and like, there are so many more managed services that are making things easy. And that’s what excites me. I mean, that’s what’s awesome about what we’re doing in cloud. We are building platforms that enable innovation.

Corey: I think that there’s an awful lot of power behind unlocking innovation from a customer perspective. The idea that I can use a cloud provider to wind up doing an experiment to build something in the course of an evening, and if it works, great, I can continue to scale up without having to replace, you know, the crappy Raspberry Pi-level hardware in my spare room with serious enterprise servers in a data center somewhere. The on-ramp and the capability and the lack of long-term commitments is absolutely magical. What I’m also seeing that is contributing to that is the de facto standard that’s emerged of most things these days support Docker, for better or worse. There are many open-source tools that I see where, “Oh, how do I get this up and running?”

“Well, you can go over the river and through the woods and way past grandmother’s house to build this from source or run this Docker file.” I feel like that is the direction the rest of the world is going. And as much fun as it is to sit on the sidelines and snark, I’m finding a lot more capability stories emerging across the board. Does that resonate with what you’re seeing, given that you are inherently working at very large scale, given the [laugh] nature of where you work?

Chen: I do see that. And I actually want to double down on the open standards, which I think this is also something that is happening. At the beginning, we talked about I want it to be very hard when I choose the cloud provider. But innovation doesn’t only come from cloud providers; there’s a lot of companies and a lot of innovation happening that are building new technologies on top of those cloud providers, and I don’t think this is going to stop. Innovation is going to come from many places, and it’s going to be very exciting.

And by the way, things are moving super fast in our space. So, the investment in open standard is critical for our industry. So, Docker is one example. Google is in [unintelligible 00:19:46] speaking, it’s investing a lot in building those open standards. So, we have Docker, we have things like of course Kubernetes, but we are also investing in open standards of security, so we are working with other partners around [unintelligible 00:19:58], defining how you can secure the software supply chain, which is also critical for innovation. So, all of those things that reduce the barrier to entry is something that I’m personally passionate about.

Corey: Scaling containers and scaling Kubernetes is hard, but a whole ‘nother level of difficulty is scaling humans. You’ve been at Google for, as you said, seven years and you did not start as a VP there. Getting promoted from Senior Director to VP at Google is a, shall we say, heavy lift. You also mentioned that you previously started with, I believe, it was a seven-person team at one point. How have you been able to do that? Because I can see a world in which, “Oh, we just write some code and we can scale the computers pretty easily,” I’ve never found a way to do that for people.

Chen: So yes, I started actually—well not 7, but the team was 30 people [laugh]. And you can imagine how surprised I was when I joining Google Cloud with Kubernetes and GKE and it was a pretty small team, to the beginning of those days. But the team was already actually on the edge of burning out. You know, pings on Slack, the GitHub issues, there was so many things happening 24/7.

And the thing was just doing everything. Everybody were doing everything. And one of the things I’ve done on my second month on the team—I did an off-site, right, all managers; that’s what we do; we do off-sites—and I brought the team in to talk about—the leadership team—to talk about our team values. And in the beginning, they were a little bit pissed, I would say, “Okay, Chen. What’s going on? You’re wasting two days of our lives to talk about those things. Why we are not doing other things?”

And I was like, “You know guys, this is really important. Let’s talk about what’s important for us.” It was an amazing it worked. By the way, that work is still the foundation of the culture in the team. We talked about the three values that we care about and how that will look like.

And the reason it’s important is that when you scale teams, the key thing is actually to scale decision-making. So, how do you scale decision-making? I think there are two things there. One is what you’re trying to achieve. So, people should know and understand the vision and know where we want to get to.

But the second thing is, how do we work? What’s important for us? How do we prioritize? How do we make trade-offs? And when you have both the what we’re trying to do and the how, you build that team culture. And when you have that, I find that you’re set up more for success for scaling the team.

Because then the storyteller is not just the leader or the manager. The entire team is a storyteller of how things are working in this team, how do we work, what you’re trying to achieve, and so on. So, that’s something that had been a critical. So, that’s just, you know, from methodology of how I think it’s the right thing to scale teams. Specifically, with a Kubernetes, there were more issues that we needed to work on.

For example, building or [recoding 00:23:05] different functions. It cannot be just engineering doing everything. So, hiring the first product managers and information engineers and marketing people, oh my God. Yes, you have to have marketing people because there are so many events. And so, that was one thing, just you know, from people and skills.

And the second thing is that it was an open-source project and a product, but what I was personally doing, I was—with the team—is bringing some product engineering practices into the open-source. So, can we say, for example, that we are going to focus on user experience this next release? And we’re not going to do all the rest. And I remember, my team was like worried about, like, “Hey, what about that, and what about this, and we have—” you know, they were juggling everything together. And I remember telling them, “Imagine that everything is on the floor. All the balls are on the floor. I know they’re on the floor, you know they’re on the floor. It’s okay. Let’s just make sure that every time we pick something up, it never falls again.” And that idea is a principle that then evolved to ‘No Heroics,’ and it evolved to ‘Sustainable Success.’ But building things towards sustainable success is a principle which has been very helpful for us.

Corey: This episode is sponsored in part by our friend at Uptycs. Attackers don’t think in silos, so why would you have siloed solutions protecting cloud, containers, and laptops distinctly? Meet Uptycs - the first unified solution that prioritizes risk across your modern attack surface—all from a single platform, UI, and data model. Stop by booth 3352 at AWS re:Invent in Las Vegas to see for yourself and visit uptycs.com. That’s U-P-T-Y-C-S.com. My thanks to them for sponsoring my ridiculous nonsense.

Corey: When I take a look back, it’s very odd to me to see the current reality that is Google, where you’re talking about empathy, and the No Heroics, and the rest of that is not the reputation that Google enjoyed back when a lot of this stuff got started. It was always oh, engineers should be extraordinarily bright and gifted, and therefore it felt at the time like our customers should be as well. There was almost an arrogance built into, well, if you wrote your code more like Google will, then maybe your code wouldn’t be so terrible in the cloud. And somewhat cynically I thought for a while that oh Kubernetes is Google’s attempt to wind up making the rest of the world write software in a way that’s more Google-y. I don’t think that observation has aged very well. I think it’s solved a tremendous number of problems for folks.

But the complexity has absolutely been high throughout most of Kubernetes life. I would argue, on some level, that it feels like it’s become successful almost in spite of that, rather than because of it. But I’m curious to get your take. Why do you believe that Kubernetes has been as successful as it clearly has?

Chen: [unintelligible 00:25:34] two things. One about empathy. So yes, Google engineers are brilliant and are amazing and all great. And our customers are amazing, and brilliant, as well. And going back to the point before is, everyone has their job and where they need to be successful and we, as you say, we need to make things simpler and enable innovation. And our customers are driving innovation on top of our platform.

So, that’s the way I think about it. And yes, it’s not as simple as it can be—probably—yet, but in studying the early days of Kubernetes, we have been investing a lot in what we call empathy, and the customer empathy workshop, for example. So, I partnered with Kelsey Hightower—and you mentioned yourself trying to start a cluster. The first time we did a workshop with my entire team, so then it was like 50 people [laugh], their task was to spin off a cluster without using any scripts that we had internally.

And unfortunately, not many folks succeeded in this task. And out of that came the—what you you call it—a OKR, which was our goal for that quarter, is that you are able to spin off a cluster in three commands and troubleshoot if something goes wrong. Okay, that came out of that workshop. So, I do think that there is a lot of foundation on that empathetic engineering and the open-source of the community helped our Google teams to be more empathetic and understand what are the different use cases that they are trying to solve.

And that actually bring me to why I think Kubernetes is so successful. People might be surprised, but the amount of investment we’re making on orchestration or placement of containers within Kubernetes is actually pretty small. And it’s been very small for the last seven years. Where do we invest time? One is, as I mentioned before, is on the what we call the API machinery.

So, Kubernetes has introduced a way that is really suitable for a cloud-native technologies, the idea of reconciliation loop, meaning that the way Kubernetes is—Kubernetes is, like, a powerful automation machine, which can automate, of course, workload placement, but can automate other things. Think about it as a way of the Kubernetes API machinery is observing what is the current state, comparing it to the desired state, and working towards it. Think about, like, a thermostat, which is a different automation versus the ‘if this, then that,’ where you need to anticipate different events. So, this idea about the API machinery and the way that you can extend it made it possible for different teams to use that mechanism to automate other things in that space.

So, that has been one very powerful mechanism of Kubernetes. And that enabled all of innovation, even if you think about things like Istio, as an example, that’s how it started, by leveraging that kind of mechanism to separate storage and so on. So, there are a lot of operators, the way people are managing their databases, or stateful workloads on top of Kubernetes, they’re extending this mechanism. So, that’s one thing that I think is key and built that ecosystem. The second thing, I am very proud of the community of Kubernetes.

Corey: Oh, it’s a phenomenal community success story.

Chen: It’s not easy to build a community, definitely not in open-source. I feel that the idea of values, you know, that I was talking about within my team was actually a big deal for us as we were building the community: how we treat each other, how do we help people start? You know, and we were talking before, like, am I going to talk about DEI and inclusivity, and so on. One of the things that I love about Kubernetes is that it’s a new technology. There is actually—[unintelligible 00:29:39] no, even today, there is no one with ten years experience in Kubernetes. And if anyone says they have that, then they are lying.

Corey: Time machine. Yes.

Chen: That creates an opportunity for a lot of people to become experts in this technology. And by having it in open-source and making everything available, you can actually do it from your living room sofa. That excites me, you know, the idea that you can become an expert in this new technology and you can get involved, and you’ll get people that will mentor you and help you through your first PR. And there are some roles within the community that you can start, you know, dipping your toes in the water. It’s exciting. So, that makes me really happy, and I know that this community has changed the trajectory of many people’s careers, which I love.

Corey: I think that’s probably one of the most impressive things that it’s done. One last question I have for you is that we’ve talked a fair bit about the history and how we see it progressing through the view toward the somewhat recent past. What do you see coming in the future? What does the future of Kubernetes look like to you?

Chen: Continue to be more and more boring. There is the promise of hybrid and multi-cloud, for example, is only possible by technologies like Kubernetes. So, I do think that, as a technology, it will continue to be important by ensuring portability and interoperability of workloads. I see a lot of edge use cases. If you think about it, it’s like just lagging a bit around, like, innovation that we’ve seen in the cloud, can we bring that innovation to the edge, this will require more development within Kubernetes community as well.

And that’s really actually excites me. I think there’s a lot of things that we’re going to see there. And by the way, you’ve seen it also in KubeCon. I mean, there were some announcements in that space. In Google Cloud, we just announced before, like, with customers like Wendy’s and Rite Aid as well. So, taking advantage of this technology to allow innovation everywhere.

But beyond that, my hope is that we’ll continue and hide the complexity. And our challenge will be to not make it a black box. Because that will be, in my opinion, a failure pattern, doesn’t help those kinds of platforms. So, that will be the challenge. Can we scope the project, ensure that we have the right observability, and from a use case perspective, I do think edge is super interesting.

Corey: I would agree. There are a lot of workloads out there that are simply never going to be hosted in the cloud provider region, for a variety of reasons of varying validity, but it is the truth. I think that the focus on addressing customers where they are has been an emerging best practice for cloud providers and I’m thrilled to see Google leading the charge on that.

Chen: Yeah. And you just reminded me, the other thing that we see also more and more is definitely AI and ML workloads running on Kubernetes, which is part of that, right? So, Google Cloud is investing a lot in making an AI/ML easy. And I don’t know if many people know, but, like, even Vertex AI, our own platform, is running on GKE. So, that’s part of seeing how do we make sure that platform is suitable for these kinds of workloads and really help customers do the heavy lifting.

So, that’s another set of workloads that are very relevant at the edge. And one of our customers—MLB, for example—two things are interesting there. The first one, I think a lot of people sometimes say, “Okay, I’m going to move to the cloud and I want to know everything right now, how that will evolve.” And one of the things that’s been really exciting with working with MLB for the last four years is the journey and the iterations. So, they started somewhat, like, at one phase and then they saw what’s possible, and then moved to the next one, and so on. So, that’s one. The other thing is that, really, they have so much ML running at the stadium with Google Cloud technology, which is very exciting.

Corey: I’m looking forward to seeing how this continues to evolve and progress, particularly in light of the recent correction we’re seeing in the market where a lot of hype-driven ideas are being stress test, maybe not in the way we might have hoped that they would, but it’ll be really interesting to see what shakes out as far as things that deliver business value and are clear wins for customers versus a lot of the speculative stories that we’ve been hearing for a while now. Maybe I’m totally wrong on this. And this is going to be a temporary bump in the road, and we’ll see no abatement in the ongoing excitement around so many of these emerging technologies, but I’m curious to see how it plays out. But that’s the beautiful part about getting to be a pundit—or whatever it is people call me these days that’s at least polite enough to say on a podcast—is that when I’m right, people think I’m a visionary, and when I’m wrong, people don’t generally hold that against you. It seems like futurist is the easiest job in the world because if you predict and get it wrong, no one remembers. Predict and get it right, you look like a genius.

Chen: So, first of all, I’m optimistic. So usually, my predictions are positive. I will say that, you know, what we are seeing, also what I’m hearing from our customers, technology is not for the sake of technology. Actually, nobody cares [laugh]. Even today.

Okay, so nothing needs to change for, like, nobody would c—even today, nobody cares about Kubernetes. They need to care, unfortunately, but what I’m hearing from our customers is, “How do we create new experiences? How we make things easy?” Talent shortage is not just with tech people. It’s also with people working in the warehouse or working in the store.

Can we use technology to help inventory management? There’s so many amazing things. So, when there is a real business opportunity, things are so much simpler. People have the right incentives to make it work. Because one thing we didn’t talk about—right, we talked about all these new technologies and we talked about scaling team and so on—a lot of time, the challenge is not the technology.

A lot of time, the challenge is the process. A lot of time, the challenge is the skills, is the culture, there’s so many things. But when you have something—going back to what I said before—how you unite teams, when there’s something a clear goal, a clear vision that everybody’s excited about, they will make it work. So, I think this is where having a purpose for the innovation is critical for any successful project.

Corey: I think and I hope that you’re right. I really want to thank you for spending as much time with me as you have. If people want to learn more, where’s the best place for them to find you?

Chen: So, first of all, on Twitter. I’m there or on LinkedIn. I will say that I’m happy to connect with folks. Generally speaking, at some point in my career, I recognized that I have a voice that can help people, and I’ve experienced that can also help people build their careers. I’m happy to share that and [unintelligible 00:36:54] folks both in the company and outside of it.

Corey: I think that’s one of the obligations on a lot of us, once we wanted to get into a certain position or careers to send the ladder back down, for lack of a better term. It’s I’ve never appreciated the perspective, “Well, screw everyone else. I got mine.” The whole point the next generation should have it easier than we did.

Chen: Yeah, definitely.

Corey: Chen Goldberg, General Manager of Cloud Runtimes and VP of Engineering at Google. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an angry rant of a comment talking about how LPARs on mainframes are absolutely not containers, making sure it’s at least far too big to fit in a reasonably-sized Docker container.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Andy

Andy is on a lifelong journey to understand, invent, apply, and leverage technology in our world. Both personally and professionally technology is at the root of his interests and passions.

Andy has always had an interest in understanding how things work at their fundamental level. In addition to figuring out how something works, the recursive journey of learning about enabling technologies and underlying principles is a fascinating experience which he greatly enjoys.

The early Internet afforded tremendous opportunities for learning and discovery. Andy’s early work focused on network engineering and architecture for regional Internet service providers in the late 1990s – a time of fantastic expansion on the Internet.

Since joining Akamai in 2000, Akamai has afforded countless opportunities for learning and curiosity through its practically limitless globally distributed compute platform. Throughout his time at Akamai, Andy has held a variety of engineering and product leadership roles, resulting in the creation of many external and internal products, features, and intellectual property.

Andy’s role today at Akamai – Senior Vice President within the CTO Team - offers broad access and input to the full spectrum of Akamai’s applied operations – from detailed patent filings to strategic company direction. Working to grow and scale Akamai’s technology and business from a few hundred people to roughly 10,000 with a world-class team is an amazing environment for learning and creating connections.

Personally Andy is an avid adventurer, observer, and photographer of nature, marine, and astronomical subjects. Hiking, typically in the varied terrain of New England, with his family is a common endeavor. He enjoys compact/embedded systems development and networking with a view towards their applications in drone technology.

Links Referenced:

  • Macrometa: https://www.macrometa.com/
  • Akamai: https://www.akamai.com/
  • LinkedIn: https://www.linkedin.com/in/andychampagne/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Forget everything you know about SSH and try Tailscale. Imagine if you didn't need to manage PKI or rotate SSH keys every time someone leaves. That'd be pretty sweet, wouldn't it? With Tailscale SSH, you can do exactly that. Tailscale gives each server and user device a node key to connect to its VPN, and it uses the same node key to authorize and authenticate SSH.

Basically you're SSHing the same way you manage access to your app. What's the benefit here? Built-in key rotation, permissions as code, connectivity between any two devices, reduce latency, and there's a lot more, but there's a time limit here. You can also ask users to reauthenticate for that extra bit of security. Sounds expensive?

Nope, I wish it were. Tailscale is completely free for personal use on up to 20 devices. To learn more, visit snark.cloud/tailscale. Again, that's snark.cloud/tailscale

Corey: Managing shards. Maintenance windows. Overprovisioning. ElastiCache bills. I know, I know. It's a spooky season and you're already shaking. It's time for caching to be simpler. Momento Serverless Cache lets you forget the backend to focus on good code and great user experiences. With true autoscaling and a pay-per-use pricing model, it makes caching easy. No matter your cloud provider, get going for free at gomomento.co/screaming That's GO M-O-M-E-N-T-O dot co slash screaming

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I like doing promoted guest episodes like this one. Not that I don’t enjoy all of my promoted guest episodes. But every once in a while, I generally have the ability to wind up winning an argument with one of my customers. Namely, it’s great to talk to you folks, but why don’t you send me someone who doesn’t work at your company? Maybe a partner, maybe an investor, maybe a customer. At Macrometa who’s sponsoring this episode said, okay, my guest today is Andy Champagne, SVP at the CTO office at Akamai. Andy, thanks for joining me.

Andy: Thanks, Corey. Appreciate you having me. And appreciate Macrometa letting me come.

Corey: Let’s start with talking about you, and then we’ll get around to the Macrometa discussion in the fullness of time. You’ve been at an Akamai for 22 years, which in tech company terms, it’s like staying at a normal job for 75 years. What’s it been like being in the same place for over two decades?

Andy: Yeah, I’ve got several gold watches. I’ve been retired twice. Nobody—you know, Akamai—so in the late-90s, I was in the ISP universe, right? So, I was in network engineering at regional ISPs, you know, kind of cutting teeth on, you know, trying to scale networks and deal with the flux of user traffic coming in from the growth of the web. And, you know, frankly, it wasn’t working, right?

Companies were trying to scale up at the time by adding bigger and bigger servers, and buying literally, you know, servers, the size of refrigerators. And all of a sudden, there was this company that was coming together out in Cambridge, I’m from Massachusetts, and Akamai started in Cambridge, Massachusetts, still headquartered there. And Akamai was forming up and they had a totally different solution to how to solve this, which was amazing. And it was compelling and it drew me there, and I am still there, 22-odd years in, trying to solve challenging problems.

Corey: Akamai is one of those companies that I often will describe to people who aren’t quite as inclined in the network direction as I’ve been previously, as one of the biggest companies of the internet that you’ve never heard of. You are—the way that I think of you historically, I know this is not how you folks frame yourself these days, but I always thought of you as the CDN that you use when it really mattered, especially in the earlier days of the internet where there were not a whole lot of good options to choose from, and the failure mode that Akamai had when I was looking at it many years ago, is that, well, it feels enterprise-y. Well, what does that mean exactly because that’s usually used as a disparaging term by any developer in San Francisco. What does that actually unpack to? And to my mind, it was, well, it was one of the more expensive options, which yes, that’s generally not a terrible thing, and also that it felt relatively stodgy, for lack of a better term, where it felt like updating things through an API was more of a JSON API—namely a guy named Jason—who would take a ticket, possibly from Jira if they were that modern or not, and then implement it by hand. I don’t believe that it is quite that bad these days because, again, this was circa 2012 that we’re talking here. But how do you view what Akamai is and does in 2022?

Andy: Yeah. Awesome question. There’s a lot to unpack in there, including a few clever jabs you threw in. But all good.

Corey: [laugh].

Andy: [laugh]. I think Akamai has been through a tremendous, tremendous series of evolutions on the internet. And really the one that, you know, we’re most excited about today is, you know, earlier this year, we kind of concluded our acquisition of Linode. And if we think about Linode, which brings compute into our platform, you know, ultimately Akamai today is a compute company that has a security offering and has a delivery offering as well. We do more security than delivery, so you know, delivery is kind of something that was really important during our first ten or twelve years, and security during the last ten, and we think compute during the next ten.

The great news there is that if you look at Linode, you can’t really find a more developer-focused company than Linode. You essentially fall into a virtual machine, you may accidentally set up a virtual machine inadvertently it’s so easy. And that is how we see the interface evolving. We see a compute-centric interface becoming standard for people as time moves on.

Corey: I’m reminded of one of those ancient advertisements, I forget, I think would have been Sun that put it out where the network is the computer or the computer is the network. The idea of that a computer sitting by itself unplugged was basically just this side of useless, whereas a bunch of interconnected computers was incredibly powerful. That today and 2022 sounds like an extraordinarily obvious statement, but it feels like this is sort of a natural outgrowth of that, where, okay, you’ve wound up solving the CDN piece of it pretty effectively. Now, you’re expanding out into, as you say, compute through the Linode acquisition and others, and the question I have is, is that because there’s a larger picture that’s currently unfolding, or is this a scenario where well, we nailed the CDN side of the world, well, on that side of the universe, there’s no new worlds left to conquer. Let’s see what else we can do. Next, maybe we’ll start making toasters.

Andy: Bunch of bored guys in Cambridge, and we’re just like, “Hey, let’s go after compute. We don’t know what we’re doing.” No. There’s a little bit more—

Corey: Exactly. “We have money and time. Let’s combine the two and see what we can come up with.”

Andy: [laugh]. Hey, folks, compute: it’s the new thing. No, it’s more than that. And you know, Akamai has a very long history with the edge, right? And Akamai started—and again, arrogantly saying, we invented the concept of the edge, right, out there in ’99, 2000, deploying hundreds and then to thousands of different locations, which is what our CDN ran on top of.

And that was a really new, novel concept at the time. We extended that. We’ve always been flirting with what is called edge computing, which is how do we take pieces of application logic and move them from a centralized point and move them out to the edge. And I mean, cripes, if you go back and Google, like, ‘Akamai edge computing,’ we were working on that in 2003, which is a bit like ancient history, right? And we are still on a quest.

And literally, we think about it in the company this way: we are on a quest to make edge computing a reality, which is how do you take applications that have centralized chokepoints? And how do you move as much of those applications as possible out to the edge of the network to unblock user performance and experience, and then see what folks developers can enable with that kind of platform?

Corey: For me, it seems that the rise of AWS—which is, by extension, the rise of cloud—has been, okay, you wind up building whatever you want for the internet and you stuff it into an AWS region, and oh, that’s far away from your customers and/or your entire architecture is terrible so it has to make 20 different calls to the data center in series rather than in parallel. Great, how do we reduce the latency as much as possible? And their answer has largely seemed to be, ah, we’ll build more regions, ever closer to you. One of these days, I expect to wake up and find that there’s an announcement that they’re launching a new region in my spare room here. It just seems to get closer and closer and closer. You look around, and there’s a cloud construction crew stalking you to the mall and whatnot. I don’t believe that is the direction that the future necessarily wants to be going in.

Andy: Yeah, I think there’s a lot there. And I would say it this way, which is, you know, having two-ish dozen uber-large data centers is probably not the peak technology of the internet, right? There’s more we need to do to be able to get applications truly distributed. And, you know, just to be clear, I mean, Amazon AWS’s done amazing stuff, they’ve projected phenomenal scale and they continue to do so. You know, but at Akamai, the problem we’re trying to solve is really different than how do we put a bunch of stuff in a small number of data centers?

It’s, you know, obviously, there’s going to be a centralized aspect, but there also needs to be incredibly integrated and seamless, moves through a gradient of compute, where hey, maybe you’re in a very large data center for your AI/ML, kind of, you know, offline data lake type stuff. And then maybe you’re in hundreds of locations for mid-tier application processing, and, you know, reconciliation of databases, et cetera. And then all the way out at the edge, you know, in thousands of locations, you should be there for user interactivity. And when I say user interactivity, I don’t just mean, you know, read-only, but you’ve got to be able to do a read-write operation in synchronous fashion with the edge. And that’s what we’re after is building ultimately a platform for that and looking at tools, technology, and people along the way to help us with it.

Corey: I’ve built something out, my lasttweetinaws.com threading Twitter client, and that’s… it’s fine. It’s stateless, but it’s a little too intricate to effectively run in the Lambda@Edge approach, so using their CloudFront offering is simply a non-starter. So, in order to get low latency for people using it around the world, I now have to deploy it simultaneously to 20 different AWS regions.

And that is, to be direct, a colossal pain in the ass. No one is really doing stuff like that, that I can see. I had to build a whole lot of customs tooling just to get a CI/CD system up and working. Their strong regional isolation is great for containing blast radii, but obnoxious when you’re trying to get something deployed globally. It’s not the only way.

Combine that with the reality that ingress data transfer to any of their regions is free—generally—but sending data to the internet is a jewel beyond price because all my stars, that is egress bandwidth; there is nothing more valuable on this planet or any other. And that doesn’t quite seem right. Because if that were actively true, a whole swath of industries and apps would not be able to exist.

Andy: Yeah, you know, Akamai, a huge part of our business is effectively distributing egress bandwidth to the world, right? And that is a big focus of ours. So, when we look at customers that are well positioned to do compute with Akamai, candidly, the filtering question that I typically ask with customers is, “Hey, do you have a highly distributed audience that you want to engage with, you know, a lot of interactivity or you’re pushing a lot of content, video, updates, whatever it is, to them?” And that notion of highly distributed applications that have high egress requirements is exactly the sweet spot that we think Akamai has, you know, just a great advantage with, between our edge platform that we’ve been working on for the last 20-odd years and obviously, the platform that Linode brings into the conversation.

Corey: Let’s talk a little bit about Macrometa.

Andy: Sure.

Corey: What is the nature of your involvement with those folks? Because it seems like you sort of crossed into a whole bunch of different areas simultaneously, which is fascinating and great to see, but to my understanding, you do not own them.

Andy: No, we don’t. No, they’re an independent company doing their thing. So, one of the fun hats that I get to wear at Akamai is, I’m responsible for our Akamai Ventures Program. So, we do our corporate investing and all this kind of thing. And we work with a wide array of companies that we think are contributing to the progression of the internet.

So, there’s a bunch of other folks out there that we work with as well. And Macrometa is on that list, which is we’ve done an investment in Macrometa, we’re board observers there, so we get to sit in and give them input on, kind of, how they’re doing things, but they don’t have to listen to us since we’re only observers. And we’ve also struck a preferred partnership with them. And what that means is that as our customers are building solutions, or as we’re building solutions for our customers, utilizing the edge, you know, we’re really excited and we’ve got Macrometa at the table to help with that. And Macrometa is—you know, just kind of as a refresher—is trying to solve the problem of distributed data access at the edge in a high-performance and almost non-blocking, developer-friendly way. And that is very, very exciting to us, so that’s the context in which they’re interesting to our continuing evolution of how the edge works.

Corey: One of the questions I always like to ask, and it’s usually not considered a personal attack when I asked the question—

Andy: Oh, good.

Corey: But it’s, “Describe what the company does.” Now, at some places like the latter days of Yahoo, for example, it’s very much a personal attack. But what is it that Macrometa does?

Andy: So, Macrometa provides a worldwide, high-speed distributed database that is resident on what today, you could call the edge of the network. And the advantage here is, instead of having one SQL server sitting somewhere, or what you would call a distributed SQL Server, which is two SQL Servers sitting next to one another, Macrometa has a high-speed data store that allows you to, instead of having that centralized SQL Server, have it run natively at the edge of the network. And when you’re building applications that run on the edge or anywhere, you need to try to think about how do you have the data as close to the user or to the access point as possible. And that’s the problem Macrometa is after and that’s what their products today solve. It’s an incredibly bright team over there, a fantastic founder-CEO team, and we’re really excited to be working with him.

Corey: It wasn’t intentionally designed this way as a setup when I mentioned a few minutes ago, but yeah, my Twitter client works across the 20-some-odd AWS regions, specifically because it’s stateless. All of the state, other than a couple of API keys at provision time, wind up living in the user’s browser. If this was something that needed to retain state in any way, like, you know, basically every real application under the sun, this strategy would absolutely not work unless I wound up with some heinous form of circular replication, and then you wind up with a single region going down and everything explodes. Having a cohesive, coherent data layer that spans all of that is key.

Andy: Yeah, and you’re on to the classical, you know, CompSci issue here around edge, which is if you have 100 edge regions, how do you have consistent state storage between applications running on N of those? And that is the problem Macrometa is after, and, you know, Akamai has been working on this and other variants of the edge problem for some time. We’re very excited to be working with the folks at Macrometa. It’s a cool group of folks. And it’s an interesting approach to the technology. And from what we’ve seen so far, it’s been working great.

Corey: The idea of how do I wind up having persistent, scalable state across a bunch of different edge locations is not just a hard computer science problem; it’s also a hard cloud economics problem, given the cost of data transit in a bunch of different directions between different providers. It turns, “How much does it cost?” In most cases to a question that can only be answered by well let’s run it for a few days and find out. Which is not usually the best way to answer some questions. Like, “Is that power socket live?” “Let’s touch it and find out.” Yeah, there are ways you learn that are extraordinarily painful.

Andy: Yeah no, nobody should be doing that with power sockets. I think this is one of these interesting areas, which is this is really right in Akamai’s backyard but it’s not realized by a lot of folks. So, you know, Akamai has, for the last 20-odd-years, been all about how do we egress as much as possible to the entire internet. The weird areas, the big areas, the small areas, the up-and-coming areas, we serve them all. And in doing that, we’ve built a very large global fabric network, which allows us to get between those locations at a very low cost because we have to move our own content around.

And hooking those together, having a essentially private network fabric that hooks the vast majority of our big locations together and then having very high-speed egress out of all of the locations to the internet, you know, that’s been how we operate our business at scale effectively and economically for years, and utilizing that for compute data replication, data synchronization tasks is what we’re doing.

Corey: There are a lot of different solutions that could be used to solve a lot of the persistent data layer question. For example, when you had to solve a similar problem with compute, you had a few options in front of you. Well, we could buy a whole bunch of computers and stuff them in a rack somewhere because, eh, cloud; how hard could it be? Saner heads prevailed, and no, no, no, we’re going to buy Linode, which was honestly a genius approach on about three different levels, and I’m still unconvinced the industry sees that for the savvy move that it was. I’m confident that’ll change in time.

Why not build it yourself? Or alternately, acquire another company that was working on something similar? Instead, you’re an investor in a company that’s doing this effectively, but not buying them outright?

Andy: Yeah, you know, and I think that’s—Akamai is beyond at this point in thinking that it’s just about ownership, right? I think that this—we don’t have to own everything in order to have a successful ecosystem. You know, certainly, we’re going to want to own key parts of it and that’s where you saw the Linode acquisition, where we felt that was kind of core. But ultimately, we believe in promoting customer choice here. And there’s a pretty big role that we have that we think we can help with companies, such as folks like Macrometa where they have, you know, really interesting technology, but they can use leverage, they can use some of our go-to-market, they can use, you know, some of our, you know, kind of guidance and expertise on running a startup—which, by the way, it’s not an easy job for these folks—and that’s what we’re there to do.

So, with things like Linode, you know, we want to bring it in, and we want to own it because we think it’s just so compelling, and it fits so well with where we want to go. With folks like Macrometa, you know, that’s still a really young area. I mean, you know, Linode was in business for many, many, many years and was a good-sized business, you know, before we bought them.

Corey: Yeah, there’s something to be said, for letting the market shake something out rather than having to do it all yourself as trailblazers. I’m a big believer in letting other companies do things. I mean, one of the more annoying things, from my position, is this idea where AWS takes a product strategy of, “Yes.” That becomes a bit of a challenge when they’re trying to wind up building compete decks, and how do we defeat the competition? And it’s like, “Wh—oh, you’re talking about the other hyperscalers?” “No, we’re talking with the service team one floor away.”

That just seems a little on the strange side to—some companies get too big and too expensive on some level. I think that there’s a very real risk of Akamai trying to do everything on the internet if you continue to expand and start listing out things that are not currently in your portfolio. And, oh, we should do that, too, and we should do that, too, and we should do that, too. And suddenly, it feels pretty closely aligned with you’re trying to do everything.

Andy: Yeah. I think we’ve been a company who has been really disciplined and not doing everything. You know, we started with CDN. And you know, we’re talking ’98 to 2010, you know, CDN was really our thing, and we feel we executed really well on that. We probably executed quite quietly and well, but feel we executed pretty well on that.

Really from 2010, 2012 to 2020, it was all about security, right? And, you know, we built, you know, pretty amazing security business, hundred percent of SaaS business, on top of our CDN platform with security. And now we’re thinking about—we did that route relatively quietly, as well, and now we’re thinking about the next ten years and how do we have that same kind of impact on cloud. And that is exciting because it’s not just centralized cloud; it’s about a distributed cloud vision. And that is really compelling and that’s why you know, we’ve got great folks that are still here and working on it.

Corey: I’m a big believer in the idea that you can start getting distilled truth out of folks, particularly companies, the more you compress the space they have to wind up saying. Something that’s why Twitter very often lets people tip their hands. But a commonplace that I look for is the title field on a company’s website. So, when I go over to akamai.com, you position yourself as something that fits in a small portion of a tweet, which is good. Whenever have a Tolstoy-length paragraph in the tooltip title for the browser tab, that’s a problem.

But you say simply, “Security, cloud delivery, performance. Akamai.” Which is beautifully well done, but security comes first. I have a mental model of Akamai as being a CDN and some other stuff that I don’t fully understand. But again, I first encountered you folks in the early-2000s.

It turns out that it’s hard to change existing opinions. Are you a CDN Company or are you a security company?

Andy: Oh, super—

Corey: In other words, if someone wind up mis-alphabetizing that and they’re about to get censured after this show because, “No, we’re a CDN, first; why did you put security first?”

Andy: You know, so all those things feed off each other, right? And this has been a question where it’s like, you know, our security layer and our distributed WAF and other security offerings run on top of the CDN layer. So, it’s all about building a common compute edge and then leveraging that for new applications. CDN was the first application. The next and second application was security.

And we think the third application, but probably not the final one, is compute. So, I think I don’t think anyone in marketing will be fired by the ordering that they did on that. I think that ultimately now, you know, for—just if we look at it from a monetary perspective, right, we do more security than we do CDN. So, there’s a lot that we have in the security business. And you know, compute’s got a long way to go, especially because it’s not just one big data center of compute; it is a different flavor than I think folks have seen before.

Corey: When I was at RSA, you folks were one of the exhibitors there. And I like to make the common observation that there are basically six companies that exhibit at RSA. Yeah, there are hundreds of booths, but it’s the same six products, all marketed are different logos with different words. And they all seem to approach it from a few relatively expectable personas and positions. I’ve always found myself agreeing with the things that you folks say, and maybe it’s because of my own network-centric background, but it doesn’t seem like you take the same approach that a number of other companies do or it’s, “Oh, it has to start with the way that developers write their first line of code.” Instead, it seems to take a holistic view that comes from the starting position of everything talks to each other on a network basis, and from here, let’s move forward. Is that accurate to how you view the security space?

Andy: Yeah, you know, our view of the security space is—again, it’s a network-centric one, right? And our work in the security space initially came from really big DDoS attacks, right? And how do we stop Distributed Denial of Service attacks from impacting folks? And that was the initial benefit that we brought. And from there, we evolved our story around, you know, how do we have a more sophisticated WAF? How do we have predictive capabilities at the edge?

So ultimately, we’re not about ingraining into your process of how your thing was written or telling you how to write it. We’re about, you know, essentially being that perimeter edge that is watching and monitoring everything that comes into you to make sure that, you know, hey, we’re not seeing Log4j-type exploits coming at you, and we’ll let you know if we do, or to block malicious activity. So, we fit on anything, which is why our security business has been so successful. If you have an application on the edge, you can put Akamai Security in front of it and it’s going to make your application better. That’s been super compelling for the last, you know, again, last decade or so that we’ve really been focused on security.

Corey: I think that it is a mistake to take a security model that starts with a view of what people have in front of them day-to-day—like, I look at my laptop and say, “Oh, this is what I spend my time on. This is where all security must start and stop.” Because yeah, okay, great. If you get physical access to my laptop, it’s pretty much game over on some level. But yeah, if you’re at a point where you’re going to bust into my house and threaten me in order to get access to my laptop, here you go.

There are no secrets that I am in possession of that are worth dying for. It’s just money and that’s okay. But looking at it through a lens of the internet has gone from science experiment to thing that the nerds love to use to a cornerstone of the fabric of modern society. And that’s not because of the magic supercomputer that we all have in our pockets, but rather because those magic supercomputers can talk to the sum total of human knowledge and any other human anywhere on the planet, basically, ever. And I don’t know that that evolution has been really appreciated by society at large as far as just how empowering that can be. But it completely changes the entire security paradigm from back in the ’80s when I got started, don’t put untrusted floppy disks into your computer or it might literally explode on your desk.

Andy: [laugh]. So, we’re talking about floppy disks now? Yes. So, first of all, the scope of impact of the internet has increased, meaning what you can do with it has increased. And directly proportional to that increase the threat vectors have increased, right? And the more systems are connected, the more vulnerabilities there are.

So listen, it’s easy to scare anybody about security on the internet. It is a topic that is an infinite well of scariness. At the same time, you know, and not just Akamai, but there’s a lot of companies out there that can, whether it’s making your development more secure, making your pipeline, your digital supply chain a more secure, or then you know where Akamai is, we’re at the end, which is you know, helping to wrap around your entire web presence to make it more secure, there’s a variety of companies that are out there really making the internet work from a security perspective. And honestly, there’s also been tremendous progress on the operating system front in the last several years, which previously was not as good—probably is way to characterize it—as it is today. So, and you know, at the end of the day, the nerds are still out there working, right?

We are out here still working on making the internet, you know, scale better, making it more secure, making it more robust because we’re probably not done, right? You know, phones are awesome, and tablet devices, et cetera, are awesome, but we’ve probably got more coming. We don’t quite know what that is yet, but we want to have the capacity, safety, and compute to power it.

Corey: How does Macrometa as a persistent data layer tie into your future vision of security first as what Akamai does? I can see a few directions, but I’m going to go out on a limb and guess that before you folks decided to make an investment in such a thing, you probably gave it more than the 30 seconds or whatnot or so a thought that I’ve had to wind up putting these pieces together.

Andy: So, a few things there. First of all, Macrometa, ultimately, we see them coming in the front door with our compute solution, right? Because as folks are building capabilities on the edge, “Hey, I want to run compute on the edge. How do I interoperate with data?” The worst answer possible is, “Well, call back to the centralized data store.”

So, we want to ensure that customers have choice and performance options for distributed data access. Macrometa fits great there. However, now pause that; let’s transition back to the security point you raised, which is, you know, coordinating an edge data security platform is a really complicated thing. Because you want to make sure that threats that are coming in on one side of the network, or you know, in one given country, you know, are also understood throughout the network. And there’s a definite role for a data platform in doing that.

We obviously, you know, for the last ten years have built several that help accomplish that at scale for our network, but we also recognize that, you know, innovation in data platforms is probably not done. And you know, Macrometa’s got some pretty interesting approaches. So, we’re very interested in working with them and talking jointly with customers, which we’ve done a bunch of, to see how that progresses. But there’s tie-ins, I would say, mostly on compute, but secondarily, there’s a lot of interesting areas with real-time security intel, they can be very useful as well.

Corey: Since I have you here, I would love to ask you something that’s a little orthogonal to the rest of this conversation, but I don’t even care about that because that’s why it’s my show; I can ask what I want.

Andy: Oh, no.

Corey: Talk to me a little bit about the Linode acquisition. Because when it first came out, I thought, “Oh, Linode must not be doing well, so it’s an acqui-hire scenario.” Followed by, “Wait a minute, that doesn’t seem quite right.” And I dug deeper, and suddenly, I started to see a bunch of things that made sense. But that’s just my outside perspective. I prefer to see you justify what it is that you’ve done.

Andy: Justify what we’ve done. Well, with that positive framing—

Corey: Exactly. “Explain yourself. How dare you, sir?”

Andy: [laugh]. “What are you doing?” So, to take that, which is first of all, Linode was doing great when we bought them and they’re continuing to do great now. You know, backstory here is actually a fun one. So, I personally have been a customer of Linode for about 13 years, and you know, super familiar with their offerings, as we’re a bunch of other folks at Akamai.

And what ultimately attracted us to Linode was, first of all, from a strategic perspective, is we talked about how Akamai thinks about Compute being a gradient of compute: you’ve got the edge, you’ve got kind of a middle tier, and you’ve got more centralized locations. Akamai has the edge, we’ve got the middle, we didn’t have the central. Linode has got the central. And obviously, you know, we’re going to see some significant expansion of capacity and scale there, but they’ve got the central location. And, you know, ultimately, we feel that there’s a lot of passion in Linode.

You know, they’re a Linux open-source-centric company, and believe it or not Akamai is, too. I mean, you know, that’s kind of how it works. And there was a great connection between the sorts of folks that they had and how they think about customers. Linode was a really customer-driven company. I mean, they were fanatical.

I mean, I as a, you know, customer of $30 a month personally, could open a ticket and I’d get an answer in five minutes. And that’s very similar to kind of how Akamai is driven, which is we’re very customer-centric, and when a customer has a problem or need something different, you know, we’re on it. So, there’s literally nothing bad there and it’s a super exciting beginning of a new chapter for Akamai, which is really how do we tackle compute? We’re super excited to have the Linode team. You know, they’re still mostly down in Philadelphia doing their thing.

And, you know, we’ve hired substantially and we’re continuing to do so, so if you want to work there, drop a note over. And it’s been fantastic. And it’s one of our, you know, really large acquisitions that we’ve done, and I think we were really lucky to find a great company in such a good position and be able to make it work.

Corey: From my perspective, one of the areas that has me excited about the acquisition stems from what I would consider to be something of a customer-base culture misalignment between the two companies. One of the things that I have always enjoyed about Linode—and in the interest of full transparency, they have been a periodic sponsor over the last five or six years of my ridiculous nonsense. I believe that they are not at the moment which I expect you to immediately rectify after this conversation, of course.

Andy: I’ll give you my credit card. Yeah.

Corey: Excellent. Excellent. We do not get in the way of people trying to give you money. But it was great because that’s exactly it. I could take a credit card in the middle of the night and spin up things on Linode.

And it was one of those companies that aligned very closely to how I tended to view cloud infrastructure from the perspective of, I need a Linux box, or I need a bunch of Linux boxes right there, right now, and I don’t have 12 weeks to go to cloud school to learn the intricacies of a given provider. It more or less just worked in a whole bunch of easy ways. Whereas if I wanted to roll out at Akamai, it was always I would pull up the website, and it’s, “Click here to talk to our enterprise sales team.” And that tells me two things. One, it is probably going to be outside of my signing authority because no one trusts me with money for obvious reasons, when I was an employee, and two, you will not be going to space today because those conversations always take time.

And it’s going to be—if I’m in a hurry and trying to get something out the door, that is going to act as a significant drag on capability. Now, most of your customers do not launch things by the seat of their pants, three hours after the idea first occurs to them, but on Linode, that often seems to be the case. The idea of addressing developers early on in the ‘it’s just an idea’ phase. I can’t shake the feeling that there’s a definite future in which Linode winds up being able to speak much more effectively to enterprise, while Akamai also learns to speak to, honestly, half-awake shitposters at 2 a.m. when we’re building something heinous.

Andy: I feel like you’ve been sitting in on our strategy presentations. Maybe not the shitposters, but the rest of it. And I think the way that I would couch it, my corporate-speak of that, would be that there’s a distinct yin and yang, there a complementary nature between the customer bases of Akamai, which has, you know, an incredible list of enterprise customers—I mean, the who’s-who of enterprise customers, Akamai works with them—but then, you know, Linode, who has really tremendous representation of developers—that’s what we’ll use for the name posts—like, folks like myself included, right, who want to throw something together, want to spin up a VM, and then maybe tear it down and never do it again, or maybe set up 100 of them. And, to your point, the crossover opportunities there, which is, you know, Linode has done a really good job of having small customers that grow over time. And by having Akamai, you know, you can now grow, and never have to leave because we’re going to be able to bring enough scale and throughput and, you know, professional help services as you need it to help you stay in the ecosystem.

And similarly, Akamai has a tremendous—you know, the benefit of a tremendous set of enterprise customers who are out there, you know, frankly, looking to solve their compute challenges, saying, “Hey, I have a highly distributed application. Akamai, how can you help me with this?” Or, “Hey, I need presence in x or y.” And now we have, you know, with Linode, the right tools to support that. And yes, we can make all kinds of jokes about, you know, Akamai and Linode and different, you know, people and archetypes we appeal to, but ultimately, there’s an alignment between Akamai and Linode on how we approach things, which is about Linux, open-source, it’s about technical honesty and simplicity. So, great group of folks. And secondly, like, I think the customer crossover, you’re right on it. And we’re very excited for how that goes.

Corey: I also want to call out that Macrometa seems to have split this difference perfectly. One of the first things I visit on any given company’s page when I’m trying to understand them is the pricing page. It’s one of those areas where people spend the least time, early on, but it’s also where they tend to be the most honest. Maybe that’s why. And I look for two things, and Macrometa has both of them.

The first is a ‘try it for free, right now, get started.’ It’s a free-tier approach. Because even if you charge $10 or whatnot, there are many developers working on things in odd hours where they don’t necessarily either have the ability to make that purchase decision, know that they have the ability to make that purchase decision, or are willing to do that by the seat of their pants. So, ‘get started for free’ is important; it means you can develop right now. Conversely, there are a bunch of enterprise procurement departments out there who will want a whole bunch of custom things.

Custom SLAs, custom support responses, custom everything, and they also don’t know how to sign a check that doesn’t have two commas in it. So, you don’t probably want to avoid those customers, but what they’re looking for is an enterprise offering that is no price. There should not be a price tag on that because you will never get it right for everyone, but what they want to see is ‘click here to contact sales.’ That is coded language for, “We are serious professionals and know who you are and how you like to operate.” They’ve got both and I think that is absolutely the right decision.

Andy: It do—

Corey: And whatever you have in between those two is almost irrelevant.

Andy: No, I think you’re on it. And Macrometa, their pricing philosophy allows you to get in and try it with zero friction, which is super important. Like, I don’t even have to use a credit card. I can experiment for free, I can try it for free, but then as I grow their pricing tier kind of scales along with that. And it’s a—you know, that is the way that folks try applications.

I always try to think about, hey, you know, if I’m on a team and we’re tasked with putting together a proof of concept for something in two days, and I’ve got, you know, a couple folks working with me, how do I do that? And you don’t have time for procurement, you might need to use the free thing to experiment. So, there is a lot that they can do. And you know, their pricing—this transparency of pricing that they have is fantastic. Now, Linode, also very transparent, we don’t have a free tier, but you know, you can get in for very low friction and try that as well.

Corey: Yeah, companies tend to go through a maturity curve evolution on these things. I’ve talked to companies that purely view it is how much money a given customer is spending determines how much attention they get. And it’s like, “Yeah, maybe take a look through some of your smaller users or new signups there.” Yeah, they’re spending $10 a month or whatnot, but their email address is@cocacola.com. Just spitballing here; maybe you might want a white-glove a few of those folks, just because not everyone comes in the door via an RFP.

Andy: Yep. We look at customers for what your potential is, right? Like, you know, how much could you end up spending with us, right? You know, so if you’re building your application on Linode, and you’re going to spend $20, for the first couple months, that’s totally fine. Get in there, experiment, and then you know, in the next several years, let’s see where it goes. So, you’re exactly right, which is, you know, that username@enterprisedomain.com is often much more indicative than what the actual bill is on a monthly basis.

Corey: I always find it a little strange when I have a vendor that I’m doing business with, and then suddenly, an account person reaches out, like, hey, let’s just have a call for half an hour to talk about what you’re doing and how you’re doing it. It’s my immediate response to that these days, just of too many years doing that, as, “I really need to look at that bill. How much are we spending, again?” And I honestly, usually not that much because believe it or not, when you focus on cloud economics for a living, you pay attention to your credit card bills, but it is always interesting to see who reaches out and who doesn’t. That’s been a strange approach, and there is no one right answer for all of this.

If every free tier account user of any given cloud provider wound up getting constant emails from their account managers, it’s how desperate are you to grow revenue, and what are you about to do to pricing? At some level of becomes… unhelpful.

Andy: I can see that. I’ve had, personally, situations where I’m a trial user of something, and all of a sudden I get emails—you know, using personal email addresses, no Akamai involvement—all of a sudden, I’m getting emails. And I’m like, “Really? Did I make the priority list for you to call me and leave me a voicemail, and then email me?” I don’t know how that’s possible.

So, from a personal perspective, totally see that. You know, from an account development perspective, you know, kind of with the Akamai hat on, it’s challenging, right? You know, folks are out there trying to figure out where business is going to come from. And I think if you’re able to get an indicator that somebody, you know, maybe you’re going to call that person at enterprisedomain.com to try to figure out, you know, hey, is this real and is this you with a side project or is this you with a proof of concept for something that could be more fruitful? And, you know, Corey, they’re probably just calling you because you’re you.

Corey: One of the things that I was surprised by where I saw the exact same thing. I started getting a series of emails from my account manager for Google Workspaces. Okay, and then I really did a spit-take when I realized this was on my personal address. Okay… so I read this carefully because what the hell is happening? Oh, they’re raising prices and it’s a campaign. Great.

Now, my one-user vanity domain is going to go from $6 a month to $8 a month or whatever. Cool, I don’t care. This is not someone actively trying to reach out as a human being. It’s an outreach campaign. Cool, fair. But that’s the problem, on some level, for super-tiny customers. It’s a, what is it, is it a shakedown? What are they about to yell at me for?

Andy: No, I got the same thing. My Google Workspace personal account, which is, like, two people, right? Like, and I got an email and then I think, like, a voicemail. And I’m like, I read the email and I’m like—you know, it’s going—again, it’s like, it was like six something and now it’s, like, eight something a month. So, it’s like, “Okay. You’re all right.”

Corey: Just go—that’s what you have a credit card for. Go ahead and charge it. It’s fine. Now, yeah, counterpoint if you’re a large company, and yeah, we’re just going to be raising prices by 20% across the board for everyone, and you look at this and like, that’s a phone number. Yeah, I kind of want some special outreach and conversations there. But it’s odd.

Andy: It’s interesting. Yeah. They’re great.

Corey: Last question before we call this an episode. In 22 years, how have you seen the market change from your perspective? Most people do not work in the industry from one company’s perspective for as long as you have. That gives you a somewhat privileged position to see, from a point of relative stability, what the industry has done.

Andy: So—

Corey: What have you noticed?

Andy: —and I’m going to give you an answer, which is about, like, the sales cycle, which is it used to be about meetings and about everybody coming together and used to have to occasionally wear a suit. And there would be, you know, meetings where you would need to get a CEO or CFO to personally see a presentation and decide something and say, “Okay, we’re going with X or Y. We’re going to make a decision.” And today, those decisions are, pretty far and wide, made much, much further down in the organization. They’re made by developers, team leads, project managers, program managers.

So, the way people engage with customers today is so different. First of all, like, most meetings are still virtual. I mean, like, yeah, we have physical meetings and we get together for things, but like, so much more is done virtually, which is cool because we built the internet so we wouldn’t have to go anywhere, so it’s nice that we got that landed. It’s unfortunate that we had to do with Covid to get there, but ultimately, I think that purchasing decisions and technology decisions are distributed so much more deeply into the organization than they were. It used to be a, like, C-level thing. We’re now seeing that stuff happened much further down in the organization.

We see that inside Akamai and we see it with our customers as well. It’s been, honestly, refreshing because you tend to be able to engage with technical folks when you’re talking about technical products. And you know, the business folks are still there and they’re helping to guide the discussions and all that, but it’s a much better time, I think, to be a technical person now than it probably was 20 years ago.

Corey: I would say that being a technical person has gotten easier in a bunch of ways; it’s gotten harder in a bunch of ways. I would say that it has transformed. I was very opposed to the idea that oh, as a sysadmin, why should I learn to write code? And in retrospect, it was because I wasn’t sure I could do it and it felt like the rising tide was going to drown me. And in hindsight, yeah, it was the right direction for the industry to go in.

But I’m also sensitive to folks who don’t want to, midway through their career, pick up an entirely new skill set in order to remain relevant. I think that it is a lot easier to do some things. Back when Akamai started, it took an intimate knowledge of GCC compiler flags, in most cases, to host a website. Now, it is checking a box on a web page and you’re done. Things have gotten easier.

The abstractions continue to slip below the waterline, so the things we have to care about getting more and more meaningful to the business. We’re nowhere near our final form yet, but I’m very excited about how accessible this industry is to folks that previously would not have been, while also disheartened by just how much there is to know. Otherwise, “Oh yeah, that entire aspect of the way that this core thing that runs my business, yeah, that’s basically magic and we just hope the magic doesn’t stop working, or we make a sacrifice to the proper God, which is usually a giant trillion-dollar company.” And the sacrifice is, of course, engineering time combined with money.

Andy: You know, technology is all about abstraction layers, right? And I think—that’s my view, right—and we’ve been spending the last several decades, not, ‘we’ Akamai; ‘we’ the technology industry—on, you know, coming up with some pretty solid abstraction layers. And you’re right, like, the, you know, GCC j6—you know, -j6—you know, kind of compiler tags not that important anymore, we could go back in time and talk about inetd, the first serverless. But other than that, you know, as we get to the present day, I think what’s really interesting is you can contribute technically without being a super coding nerd. There’s all kinds of different technical approaches today and technical disciplines that aren’t just about development.

Development is super important, but you know, frankly, the sysadmin skill set is more valuable today if you look at what SREs have become and how important they are to the industry. I mean, you know, those are some of the most critical folks in the entire piping here. So, don’t feel bad for starting out as a sysadmin. I think that’s my closing comment back to you.

Corey: I think that’s probably a good place to leave it. I really want to thank you for being so generous with your time.

Andy: Anytime.

Corey: If people want to learn more about how you see the world, where can they find you?

Andy: Yeah, I mean, I guess you could check me out on LinkedIn. Happy to shoot me something there and happy to catch up. I’m pretty much read-only on social, so I don’t pontificate a lot on Twitter, but—

Corey: Such a good decision.

Andy: Feel free to shoot me something on LinkedIn if you want to get in touch or chat about Akamai.

Corey: Excellent. And of course, our thanks goes well, to the fine folks at Macrometa who have promoted this episode. It is always appreciated when people wind up supporting this ridiculous nonsense that I do. My guest has been Andy Champagne SVP at the CTO office over at Akamai. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an insulting comment that will not post successfully because your podcast provider of choice wound up skimping out on a provider who did not care enough about a persistent global data layer.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Ben

Ben Whaley is a staff software engineer at Chime. Ben is co-author of the UNIX and Linux System Administration Handbook, the de facto standard text on Linux administration, and is the author of two educational videos: Linux Web Operations and Linux System Administration. He is an AWS Community Hero since 2014. Ben has held Red Hat Certified Engineer (RHCE) and Certified Information Systems Security Professional (CISSP) certifications. He earned a B.S. in Computer Science from Univ. of Colorado, Boulder.

Links Referenced:

  • Chime Financial: https://www.chime.com/
  • alternat.cloud: https://alternat.cloud
  • Twitter: https://twitter.com/iamthewhaley
  • LinkedIn: https://www.linkedin.com/in/benwhaley/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Forget everything you know about SSH and try Tailscale. Imagine if you didn't need to manage PKI or rotate SSH keys every time someone leaves. That'd be pretty sweet, wouldn't it? With Tailscale SSH, you can do exactly that. Tailscale gives each server and user device a node key to connect to its VPN, and it uses the same node key to authorize and authenticate SSH.

Basically you're SSHing the same way you manage access to your app. What's the benefit here? Built-in key rotation, permissions as code, connectivity between any two devices, reduce latency, and there's a lot more, but there's a time limit here. You can also ask users to reauthenticate for that extra bit of security. Sounds expensive?

Nope, I wish it were. Tailscale is completely free for personal use on up to 20 devices. To learn more, visit snark.cloud/tailscale. Again, that's snark.cloud/tailscale

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn and this is an episode unlike any other that has yet been released on this august podcast. Let’s begin by introducing my first-time guest somehow because apparently an invitation got lost in the mail somewhere. Ben Whaley is a staff software engineer at Chime Financial and has been an AWS Community Hero since Andy Jassy was basically in diapers, to my level of understanding. Ben, welcome to the show.

Ben: Corey, so good to be here. Thanks for having me on.

Corey: I’m embarrassed that you haven’t been on the show before. You’re one of those people that slipped through the cracks and somehow I was very bad at following up slash hounding you into finally agreeing to be here. But you certainly waited until you had something auspicious to talk about.

Ben: Well, you know, I’m the one that really should be embarrassed here. You did extend the invitation and I guess I just didn’t feel like I had something to drop. But I think today we have something that will interest most of the listeners without a doubt.

Corey: So, folks who have listened to this podcast before, or read my newsletter, or follow me on Twitter, or have shared an elevator with me, or at any point have passed me on the street, have heard me complain about the Managed NAT Gateway and it’s egregious data processing fee of four-and-a-half cents per gigabyte. And I have complained about this for small customers because they’re in the free tier; why is this thing charging them 32 bucks a month? And I have complained about this on behalf of large customers who are paying the GDP of the nation of Belize in data processing fees as they wind up shoving very large workloads to and fro, which is I think part of the prerequisite requirements for having a data warehouse. And you are no different than the rest of these people who have those challenges, with the singular exception that you have done something about it, and what you have done is so, in retrospect, blindingly obvious that I am embarrassed the rest of us never thought of it.

Ben: It’s interesting because when you are doing engineering, it’s often the simplest solution that is the best. I’ve seen this repeatedly. And it’s a little surprising that it didn’t come up before, but I think it’s in some way, just a matter of timing. But what we came up with—and is this the right time to get into it, do you want to just kind of name the solution, here?

Corey: Oh, by all means. I’m not going to steal your thunder. Please, tell us what you have wrought.

Ben: We’re calling it AlterNAT and it’s an alternative solution to a high-availability NAT solution. As everybody knows, NAT Gateway is sort of the default choice; it certainly is what AWS pushes everybody towards. But there is, in fact, a legacy solution: NAT instances. These were around long before NAT Gateway made an appearance. And like I said they’re considered legacy, but with the help of lots of modern AWS innovations and technologies like Lambdas and auto-scaling groups with max instance lifetimes and the latest generation of networking improved or enhanced instances, it turns out that we can maybe not quite get as effective as a NAT Gateway, but we can save a lot of money and skip those data processing charges entirely by having a NAT instance solution with a failover NAT Gateway, which I think is kind of the key point behind the solution. So, are you interested in diving into the technical details?

Corey: That is very much the missing piece right there. You’re right. What we used to use was NAT instances. That was the thing that we used because we didn’t really have another option. And they had an interface in the public subnet where they lived and an interface hanging out in the private subnet, and they had to be configured to wind up passing traffic to and fro.

Well, okay, that’s great and all but isn’t that kind of brittle and dangerous? I basically have a single instance as a single point of failure and these are the days early on when individual instances did not have the level of availability and durability they do now. Yeah, it’s kind of awful, but here you go. I mean, the most galling part of the Managed NAT Gateway service is not that it’s expensive; it’s that it’s expensive, but also incredibly good at what it does. You don’t have to think about this whole problem anymore, and as of recently, it also supports ipv4 to ipv6 translation as well.

It’s not that the service is bad. It’s that the service is stonkingly expensive, particularly at scale. And everything that we’ve seen before is either oh, run your own NAT instances or bend your knee and pays your money. And a number of folks have come up with different options where this is ridiculous. Just go ahead and run your own NAT instances.

Yeah, but what happens when I have to take it down for maintenance or replace it? It’s like, well, I guess you’re not going to the internet today. This has the, in hindsight, obvious solution, well, we just—we run the Managed NAT Gateway because the 32 bucks a year in instance-hour charges don’t actually matter at any point of scale when you’re doing this, but you wind up using that for day in, day out traffic, and the failover mode is simply you’ll use the expensive Managed NAT Gateway until the instance is healthy again and then automatically change the route table back and forth.

Ben: Yep. That’s exactly it. So, the auto-scaling NAT instance solution has been around for a long time well, before even NAT Gateway was released. You could have NAT instances in an auto-scaling group where the size of the group was one, and if the NAT instance failed, it would just replace itself. But this left a period in which you’d have no internet connectivity during that, you know, when the NAT instance was swapped out.

So, the solution here is that when auto-scaling terminates an instance, it fails over the route table to a standby NAT Gateway, rerouting the traffic. So, there’s never a point at which there’s no internet connectivity, right? The NAT instance is running, processing traffic, gets terminated after a certain period of time, configurable, 14 days, 30 days, whatever makes sense for your security strategy could be never, right? You could choose that you want to have your own maintenance window in which to do it.

Corey: And let’s face it, this thing is more or less sitting there as a network traffic router, for lack of a better term. There is no need to ever log into the thing and make changes to it until and unless there’s a vulnerability that you can exploit via somehow just talking to the TCP stack when nothing’s actually listening on the host.

Ben: You know, you can run your own AMI that has been pared down to almost nothing, and that instance doesn’t do much. It’s using just a Linux kernel to sit on two networks and pass traffic back and forth. It has a translation table that kind of keeps track of the state of connections and so you don’t need to have any service running. To manage the system, we have SSM so you can use Session Manager to log in, but frankly, you can just disable that. You almost never even need to get a shell. And that is, in fact, an option we have in the solution is to disable SSM entirely.

Corey: One of the things I love about this approach is that it is turnkey. You throw this thing in there and it’s good to go. And in the event that the instance becomes unhealthy, great, it fails traffic over to the Managed NAT Gateway while it terminates the old node and replaces it with a healthy one and then fails traffic back. Now, I do need to ask, what is the story of network connections during that failover and failback scenario?

Ben: Right, that’s the primary drawback, I would say, of the solution is that any established TCP connections that are on the NAT instance at the time of a route change will be lost. So, say you have—

Corey: TCP now terminates on the floor.

Ben: Pretty much. The connections are dropped. If you have an open SSH connection from a host in the private network to a host on the internet and the instance fails over to the NAT Gateway, the NAT Gateway doesn’t have the translation table that the NAT instance had. And not to mention, the public IP address also changes because you have an Elastic IP assigned to the NAT instance, a different Elastic IP assigned to the NAT Gateway, and so because that upstream IP is different, the remote host is, like, tracking the wrong IP. So, those connections, they’re going to be lost.

So, there are some use cases where this may not be suitable. We do have some ideas on how you might mitigate that, for example, with the use of a maintenance window to schedule the replacement, replaced less often so it doesn’t have to affect your workflow as much, but frankly, for many use cases, my belief is that it’s actually fine. In our use case at Chime, we found that it’s completely fine and we didn’t actually experience any errors or failures. But there might be some use cases that are more sensitive or less resilient to failure in the first place.

Corey: I would also point out that a lot of how software is going to behave is going to be a reflection of the era in which it was moved to cloud. Back in the early days of EC2, you had no real sense of reliability around any individual instance, so everything was written in a very defensive manner. These days, with instances automatically being able to flow among different hardware so we don’t get instance interrupt notifications the way we once did on a semi-constant basis, it more or less has become what presents is bulletproof, so a lot of people are writing software that’s a bit more brittle. But it’s always been a best practice that when a connection fails okay, what happens at failure? Do you just give up and throw your hands in the air and shriek for help or do you attempt to retry a few times, ideally backing off exponentially?

In this scenario, those retries will work. So, it’s a question of how well have you built your software. Okay, let’s say that you made the worst decisions imaginable, and okay, if that connection dies, the entire workload dies. Okay, you have the option to refactor it to be a little bit better behaved, or alternately, you can keep paying the Manage NAT Gateway tax of four-and-a-half cents per gigabyte in perpetuity forever. I’m not going to tell you what decision to make, but I know which one I’m making.

Ben: Yeah, exactly. The cost savings potential of it far outweighs the potential maintenance troubles, I guess, that you could encounter. But the fact is, if you’re relying on Managed NAT Gateway and paying the price for doing so, it’s not as if there’s no chance for connection failure. NAT Gateway could also fail. I will admit that I think it’s an extremely robust and resilient solution. I've been really impressed with it, especially so after having worked on this project, but it doesn’t mean it can’t fail.

And beyond that, upstream of the NAT Gateway, something could in fact go wrong. Like, internet connections are unreliable, kind of by design. So, if your system is not resilient to connection failures, like, there’s a problem to solve there anyway; you’re kind of relying on hope. So, it’s a kind of a forcing function in some ways to build architectural best practices, in my view.

Corey: I can’t stress enough that I have zero problem with the capabilities and the stability of the Managed NAT Gateway solution. My complaints about it start and stop entirely with the price. Back when you first showed me the blog post that is releasing at the same time as this podcast—and you can visit that at alternat.cloud—you sent me an early draft of this and what I loved the most was that your math was off because of a not complete understanding of the gloriousness that is just how egregious the NAT Gateway charges are.

Your initial analysis said, “All right, if you’re throwing half a terabyte out to the internet, this has the potential of cutting the bill by”—I think it was $10,000 or something like that. It’s, “Oh no, no. It has the potential to cut the bill by an entire twenty-two-and-a-half thousand dollars.” Because this processing fee does not replace any egress fees whatsoever. It’s purely additive. If you forget to have a free S3 Gateway endpoint in a private subnet, every time you put something into or take something out of S3, you’re paying four-and-a-half cents per gigabyte on that, despite the fact there’s no internet transitory work, it’s not crossing availability zones. It is simply a four-and-a-half cent fee to retrieve something that has only cost you—at most—2.3 cents per month to store in the first place. Flip that switch, that becomes completely free.

Ben: Yeah. I’m not embarrassed at all to talk about the lack of education I had around this topic. The fact is I’m an engineer primarily and I came across the cost stuff because it kind of seemed like a problem that needed to be solved within my organization. And if you don’t mind, I might just linger on this point and kind of think back a few months. I looked at the AWS bill and I saw this egregious ‘EC2 Other’ category. It was taking up the majority of our bill. Like, the single biggest line item was EC2 Other. And I was like, “What could this be?”

Corey: I want to wind up flagging that just because that bears repeating because I often get people pushing back of, “Well, how bad—it’s one Managed NAT Gateway. How much could it possibly cost? $10?” No, it is the majority of your monthly bill. I cannot stress that enough.

And that’s not because the people who work there are doing anything that they should not be doing or didn’t understand all the nuances of this. It’s because for the security posture that is required for what you do—you are at Chime Financial, let’s be clear here—putting everything in public subnets was not really a possibility for you folks.

Ben: Yeah. And not only that but there are plenty of services that have to be on private subnets. For example, AWS Glue services must run in private VPC subnets if you want them to be able to talk to other systems in your VPC; like, they cannot live in public subnet. So essentially, if you want to talk to the internet from those jobs, you’re forced into some kind of NAT solution. So, I dug into the EC2 Other category and I started trying to figure out what was going on there.

There’s no way—natively—to look at what traffic is transiting the NAT Gateway. There’s not an interface that shows you what’s going on, what’s the biggest talkers over that network. Instead, you have to have flow logs enabled and have to parse those flow logs. So, I dug into that.

Corey: Well, you’re missing a step first because in a lot of environments, people have more than one of these things, so you get to first do the scavenger hunt of, okay, I have a whole bunch of Managed NAT Gateways and first I need to go diving into CloudWatch metrics and figure out which are the heavy talkers. Is usually one or two followed by a whole bunch of small stuff, but not always, so figuring out which VPC you’re even talking about is a necessary prerequisite.

Ben: Yeah, exactly. The data around it is almost missing entirely. Once you come to the conclusion that it is a particular NAT Gateway—like, that’s a set of problems to solve on its own—but first, you have to go to the flow logs, you have to figure out what are the biggest upstream IPs that it’s talking to. Once you have the IP, it still isn’t apparent what that host is. In our case, we had all sorts of outside parties that we were talking to a lot and it’s a matter of sorting by volume and figuring out well, this IP, what is the reverse IP? Who is potentially the host there?

I actually had some wrong answers at first. I set up VPC endpoints to S3 and DynamoDB and SQS because those were some top talkers and that was a nice way to gain some security and some resilience and save some money. And then I found, well, Datadog; that’s another top talker for us, so I ended up creating a nice private link to Datadog, which they offer for free, by the way, which is more than I can say for some other vendors. But then I found some outside parties, there wasn’t a nice private link solution available to us, and yet, it was by far the largest volume. So, that’s what kind of started me down this track is analyzing the NAT Gateway myself by looking at VPC flow logs. Like, it’s shocking that there isn’t a better way to find that traffic.

Corey: It’s worse than that because VPC flow logs tell you where the traffic is going and in what volumes, sure, on an IP address and port basis, but okay, now you have a Kubernetes cluster that spans two availability zones. Okay, great. What is actually passing through that? So, you have one big application that just seems awfully chatty, you have multiple workloads running on the thing. What’s the expensive thing talking back and forth? The only way that you can reliably get the answer to that I found is to talk to people about what those workloads are actually doing, and failing that you’re going code spelunking.

Ben: Yep. You’re exactly right about that. In our case, it ended up being apparent because we have a set of subnets where only one particular project runs. And when I saw the source IP, I could immediately figure that part out. But if it’s a K8s cluster in the private subnets, yeah, how are you going to find it out? You’re going to have to ask everybody that has workloads running there.

Corey: And we’re talking about in some cases, millions of dollars a month. Yeah, it starts to feel a little bit predatory as far as how it’s priced and the amount of work you have to put in to track this stuff down. I’ve done this a handful of times myself, and it’s always painful unless you discover something pretty early on, like, oh, it’s talking to S3 because that’s pretty obvious when you see that. It’s, yeah, flip switch and this entire engagement just paid for itself a hundred times over. Now, let’s see what else we can discover.

That is always one of those fun moments because, first, customers are super grateful to learn that, oh, my God, I flipped that switch. And I’m saving a whole bunch of money. Because it starts with gratitude. “Thank you so much. This is great.” And it doesn’t take a whole lot of time for that to alchemize into anger of, “Wait. You mean, I’ve been being ridden like a pony for this long and no one bothered to mention that if I click a button, this whole thing just goes away?”

And when you mention this to your AWS account team, like, they’re solicitous, but they either have to present as, “I didn’t know that existed either,” which is not a good look, or, “Yeah, you caught us,” which is worse. There’s no positive story on this. It just feels like a tax on not knowing trivia about AWS. I think that’s what really winds me up about it so much.

Ben: Yeah, I think you’re right on about that as well. My misunderstanding about the NAT pricing was data processing is additive to data transfer. I expected when I replaced NAT Gateway with NAT instance, that I would be substituting data transfer costs for NAT Gateway costs, NAT Gateway data processing costs. But in fact, NAT Gateway incurs both data processing and data transfer. NAT instances only incur data transfer costs. And so, this is a big difference between the two solutions.

Not only that, but if you’re in the same region, if you’re egressing out of your say us-east-1 region and talking to another hosted service also within us-east-1—never leaving the AWS network—you don’t actually even incur data transfer costs. So, if you’re using a NAT Gateway, you’re paying data processing.

Corey: To be clear you do, but it is cross-AZ in most cases billed at one penny egressing, and on the other side, that hosted service generally pays one penny ingressing as well. Don’t feel bad about that one. That was extraordinarily unclear and the only reason I know the answer to that is that I got tired of getting stonewalled by people that later turned out didn’t know the answer, so I ran a series of experiments designed explicitly to find this out.

Ben: Right. As opposed to the five cents to nine cents that is data transfer to the internet. Which, add that to data processing on a NAT Gateway and you’re paying between thirteen-and-a-half cents to nine-and-a-half cents for every gigabyte egressed. And this is a phenomenal cost. And at any kind of volume, if you’re doing terabytes to petabytes, this becomes a significant portion of your bill. And this is why people hate the NAT Gateway so much.

Corey: I am going to short-circuit an angry comment I can already see coming on this where people are going to say, “Well, yes. But it’s a multi-petabyte scale. Nobody’s paying on-demand retail price.” And they’re right. Most people who are transmitting that kind of data, have a specific discount rate applied to what they’re doing that varies depending upon usage and use case.

Sure, great. But I’m more concerned with the people who are sitting around dreaming up ideas for a company where I want to wind up doing some sort of streaming service. I talked to one of those companies very early on in my tenure as a consultant around the billing piece and they wanted me to check their napkin math because they thought that at their numbers when they wound up scaling up, if their projections were right, that they were going to be spending $65,000 a minute, and what did they not understand? And the answer was, well, you didn’t understand this other thing, so it’s going to be more than that, but no, you’re directionally correct. So, that idea that started off on a napkin, of course, they didn’t build it on top of AWS; they went elsewhere.

And last time I checked, they’d raised well over a quarter-billion dollars in funding. So, that’s a business that AWS would love to have on a variety of different levels, but they’re never going to even be considered because by the time someone is at scale, they either have built this somewhere else or they went broke trying.

Ben: Yep, absolutely. And we might just make the point there that while you can get discounts on data transfer, you really can’t—or it’s very rare—to get discounts on data processing for the NAT Gateway. So, any kind of savings you can get on data transfer would apply to a NAT instance solution, you know, saving you four-and-a-half cents per gigabyte inbound and outbound over the NAT Gateway equivalent solution. So, you’re paying a lot for the benefit of a fully-managed service there. Very robust, nicely engineered fully-managed service as we’ve already acknowledged, but an extremely expensive solution for what it is, which is really just a proxy in the end. It doesn’t add any value to you.

Corey: The only way to make that more expensive would be to route it through something like Splunk or whatnot. And Splunk does an awful lot for what they charge per gigabyte, but it just feels like it’s rent-seeking in some of the worst ways possible. And what I love about this is that you’ve solved the problem in a way that is open-source, you have already released it in Terraform code. I think one of the first to-dos on this for someone is going to be, okay now also make it CloudFormation and also make it CDK so you can drop it in however you want.

And anyone can use this. I think the biggest mistake people might make in glancing at this is well, I’m looking at the hourly charge for the NAT Gateways and that’s 32-and-a-half bucks a month and the instances that you recommend are hundreds of dollars a month for the big network-optimized stuff. Yeah, if you care about the hourly rate of either of those two things, this is not for you. That is not the problem that it solves. If you’re an independent learner annoyed about the $30 charge you got for a Managed NAT Gateway, don’t do this. This will only add to your billing concerns.

Where it really shines is once you’re at, I would say probably about ten terabytes a month, give or take, in Managed NAT Gateway data processing is where it starts to consider this. The breakeven is around six or so but there is value to not having to think about things. Once you get to that level of spend, though it’s worth devoting a little bit of infrastructure time to something like this.

Ben: Yeah, that’s effectively correct. The total cost of running the solution, like, all-in, there’s eight Elastic IPs, four NAT Gateways, if you’re—say you’re four zones; could be less if you’re in fewer zones—like, n NAT Gateways, n NAT instances, depending on how many zones you’re in, and I think that’s about it. And I said right in the documentation, if any of those baseline fees are a material number for your use case, then this is probably not the right solution. Because we’re talking about saving thousands of dollars. Any of these small numbers for NAT Gateway hourly costs, NAT instance hourly costs, that shouldn’t be a factor, basically.

Corey: Yeah, it’s like when I used to worry about costing my customers a few tens of dollars in Cost Explorer or CloudWatch or request fees against S3 for their Cost and Usage Reports. It’s yeah, that does actually have a cost, there’s no real way around it, but look at the savings they’re realizing by going through that. Yeah, they’re not going to come back and complaining about their five-figure consulting engagement costing an additional $25 in AWS charges and then lowering it by a third. So, there’s definitely a difference as far as how those things tend to be perceived. But it’s easy to miss the big stuff when chasing after the little stuff like that.

This is part of the problem I have with an awful lot of cost tooling out there. They completely ignore cost components like this and focus only on the things that are easy to query via API, of, oh, we’re going to cost-optimize your Kubernetes cluster when they think about compute and RAM. And, okay, that’s great, but you’re completely ignoring all the data transfer because there’s still no great way to get at that programmatically. And it really is missing the forest for the trees.

Ben: I think this is key to any cost reduction project or program that you’re undertaking. When you look at a bill, look for the biggest spend items first and work your way down from there, just because of the impact you can have. And that’s exactly what I did in this project. I saw that ‘EC2 Other’ slash NAT Gateway was the big item and I started brainstorming ways that we could go about addressing that. And now I have my next targets in mind now that we’ve reduced this cost to effectively… nothing, extremely low compared to what it was, we have other new line items on our bill that we can start optimizing. But in any cost project, start with the big things.

Corey: You have come a long way around to answer a question I get asked a lot, which is, “How do I become a cloud economist?” And my answer is, you don’t. It’s something that happens to you. And it appears to be happening to you, too. My favorite part about the solution that you built, incidentally, is that it is being released under the auspices of your employer, Chime Financial, which is immune to being acquired by Amazon just to kill this thing and shut it up.

Because Amazon already has something shitty called Chime. They don’t need to wind up launching something else or acquiring something else and ruining it because they have a Slack competitor of sorts called Amazon Chime. There’s no way they could acquire you [unintelligible 00:27:45] going to get lost in the hallways.

Ben: Well, I have confidence that Chime will be a good steward of the project. Chime’s goal and mission as a company is to help everyone achieve financial peace of mind and we take that really seriously. We even apply it to ourselves and that was kind of the impetus behind developing this in the first place. You mentioned earlier we have Terraform support already and you’re exactly right. I’d love to have CDK, CloudFormation, Pulumi supports, and other kinds of contributions are more than welcome from the community.

So, if anybody feels like participating, if they see a feature that’s missing, let’s make this project the best that it can be. I suspect we can save many companies, hundreds of thousands or millions of dollars. And this really feels like the right direction to go in.

Corey: This is easily a multi-billion dollar savings opportunity, globally.

Ben: That’s huge. I would be flabbergasted if that was the outcome of this.

Corey: The hardest part is reaching these people and getting them on board with the idea of handling this. And again, I think there’s a lot of opportunity for the project to evolve in the sense of different settings depending upon risk tolerance. I can easily see a scenario where in the event of a disruption to the NAT instance, it fails over to the Managed NAT Gateway, but fail back becomes manual so you don’t have a flapping route table back and forth or a [hold 00:29:05] downtime or something like that. Because again, in that scenario, the failure mode is just well, you’re paying four-and-a-half cents per gigabyte for a while until you wind up figuring out what’s going on as opposed to the failure mode of you wind up disrupting connections on an ongoing basis, and for some workloads, that’s not tenable. This is absolutely, for the common case, the right path forward.

Ben: Absolutely. I think it’s an enterprise-grade solution and the more knobs and dials that we add to tweak to make it more robust or adaptable to different kinds of use cases, the best outcome here would actually be that the entire solution becomes irrelevant because AWS fixes the NAT Gateway pricing. If that happens, I will consider the project a great success.

Corey: I will be doing backflips like you wouldn’t believe. I would sing their praises day in, day out. I’m not saying reduce it to nothing, even. I’m not saying it adds no value. I would change the way that it’s priced because honestly, the fact that I can run an EC2 instance and be charged $0 on a per-gigabyte basis, yeah, I would pay a premium on an hourly charge based upon traffic volumes, but don’t meter per gigabyte. That’s where it breaks down.

Ben: Absolutely. And why is it additive to data transfer, also? Like, I remember first starting to use VPC when it was launched and reading about the NAT instance requirement and thinking, “Wait a minute. I have to pay this extra management and hourly fee just so my private hosts could reach the internet? That seems kind of janky.”

And Amazon established a norm here because Azure and GCP both have their own equivalent of this now. This is a business choice. This is not a technical choice. They could just run this under the hood and not charge anybody for it or build in the cost and it wouldn’t be this thing we have to think about.

Corey: I almost hate to say it, but Oracle Cloud does, for free.

Ben: Do they?

Corey: It can be done. This is a business decision. It is not a technical capability issue where well, it does incur cost to run these things. I understand that and I’m not asking for things for free. I very rarely say that this is overpriced when I’m talking about AWS billing issues. I’m talking about it being unpredictable, I’m talking about it being impossible to see in advance, but the fact that it costs too much money is rarely my complaint. In this case, it costs too much money. Make it cost less.

Ben: If I’m not mistaken. GCPs equivalent solution is the exact same price. It’s also four-and-a-half cents per gigabyte. So, that shows you that there’s business games being played here. Like, Amazon could get ahead and do right by the customer by dropping this to a much more reasonable price.

Corey: I really want to thank you both for taking the time to speak with me and building this glorious, glorious thing. Where can we find it? And where can we find you?

Ben: alternat.cloud is going to be the place to visit. It’s on Chime’s GitHub, which will be released by the time this podcast comes out. As for me, if you want to connect, I’m on Twitter. @iamthewhaley is my handle. And of course, I’m on LinkedIn.

Corey: Links to all of that will be in the podcast notes. Ben, thank you so much for your time and your hard work.

Ben: This was fun. Thanks, Corey.

Corey: Ben Whaley, staff software engineer at Chime Financial, and AWS Community Hero. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry rant of a comment that I will charge you not only four-and-a-half cents per word to read, but four-and-a-half cents to reply because I am experimenting myself with being a rent-seeking schmuck.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Mike

Besides his duties as The Duckbill Group’s CEO, Mike is the author of O’Reilly’s Practical Monitoring, and previously wrote the Monitoring Weekly newsletter and hosted the Real World DevOps podcast. He was previously a DevOps Engineer for companies such as Taos Consulting, Peak Hosting, Oak Ridge National Laboratory, and many more. Mike is originally from Knoxville, TN (Go Vols!) and currently resides in Portland, OR.

Links Referenced:

  • Twitter: https://twitter.com/Mike_Julian
  • mikejulian.com: https://mikejulian.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is brought to us in part by our friends at Datadog. Datadog is a SaaS monitoring and security platform that enables full-stack observability for modern infrastructure and applications at every scale. Datadog enables teams to see everything: dashboarding, alerting, application performance monitoring, infrastructure monitoring, UX monitoring, security monitoring, dog logos, and log management, in one tightly integrated platform. With 600-plus out-of-the-box integrations with technologies including all major cloud providers, databases, and web servers, Datadog allows you to aggregate all your data into one platform for seamless correlation, allowing teams to troubleshoot and collaborate together in one place, preventing downtime and enhancing performance and reliability. Get started with a free 14-day trial by visiting datadoghq.com/screaminginthecloud, and get a free t-shirt after installing the agent.

Corey: Forget everything you know about SSH and try Tailscale. Imagine if you didn't need to manage PKI or rotate SSH keys every time someone leaves. That'd be pretty sweet, wouldn't it? With Tailscale SSH, you can do exactly that. Tailscale gives each server and user device a node key to connect to its VPN, and it uses the same node key to authorize and authenticate SSH.

Basically you're SSHing the same way you manage access to your app. What's the benefit here? Built in key rotation permissions is code connectivity between any two devices, reduce latency and there's a lot more, but there's a time limit here. You can also ask users to reauthenticate for that extra bit of security. Sounds expensive?

Nope, I wish it were. tail scales. Completely free for personal use on up to 20 devices. To learn more, visit snark.cloud/tailscale. Again, that's snark.cloud/tailscale

Corey: Welcome to Screaming in the Cloud, I’m Corey Quinn and I’m having something of a crisis of faith based upon a recent conversation I’ve had with my returning yet again guest, Mike Julian, my business partner and CEO of The Duckbill Group. Welcome back, Mike.

Mike: Hi, everyone.

Corey: So, the revelation that had surfaced unexpectedly was, based upon a repeated talking point where I am a terrible employee slash expensive to manage, et cetera, et cetera, and you pointed out that you’ve been managing me for four years or so now, at which point I did a spit take, made all the more impressive by the fact that I wasn’t drinking anything at the time, and realized, “Oh, my God, you’re right, but I haven’t had any of the usual problems slash friction with you that I have with basically every boss I’ve ever had in my entire career.” So, I’m spiraling. Let’s talk about that.

Mike: My recollection of that conversation is slightly different than yours. Mine is that you called me and said, “Mike, I just realized that you’re my boss.” And I’m like, “How do you feel about that?” He’s like, “I’m not really sure.”

Corey: And I’m still not entirely sure how I feel if I’m being fully honest with you. Just because it’s such a weird thing to have to deal with. Because historically, I always view a managerial relationship as starting from a place of a power imbalance. And that is the one element that is missing from our relationship. We each own half the company, we can fire each other, but it takes the form of tearing the company apart, and that isn’t something that we’re really set up to entertain.

Mike: And you know, I actually think it’s deeper than that because you owning the other half of the company is not really… it’s not really power in itself. Like, yeah, it is, but you could easily own half the company and have no power. Because, like, really when we talk about power, we’re talking about political power, influence, and I think the reason that there is no power imbalance is because each of us does something in the company that is just as important as the other. And they’re both equally valuable to the company and we both recognize the other’s contributions, as that, as being equally valuable to the company. It’s less to do about how much we own and more about the work that we do.

Corey: Oh, of course. The ownership starts and stops entirely with the fact that neither one of us can force the other out. So it’s, as opposed to well, I own 51% of the company, so when I’m tired of your bullshit, you’re leaving. And that is a dynamic that’s never entered into it. I’m also going to add one more thing onto what you just said, which is, both of us would sooner tear off our own skin than do the other’s job.

Mike: Yeah. God, I would hate to do your job, but I know you’d hate to do mine.

Corey: You look at my calendar on a busy meeting day and you have a minor panic attack just looking at it where, “Oh, my God, talking to that many people.” And you are going away for a while and you come back with a whole analytical model where your first love language feels like it’s spreadsheets on some days, and I look at this and it’s like, “Yeah, I know what some of those numbers mean.” And it just drives me up a wall, the idea of building out a plan and an execution thing and then delegating a lot of it to other people, it does not work for my worldview in so many different ways. It’s the reason I think that you and I get along. That and our shared values.

Mike: I remember the first time that you and I did a consulting engagement together. We went on a multi-day trip. And at the end of, like, three days of nonstop conversations, you made a comment, it was like, “Cool. So, what are we going to do that again?” Like, you were excited by it. I can tell you’re energized. And I was just thinking, “Please for love of God, I want to die right now.”

Corey: One of the weirdest parts about all of it, though, is neither one of us is in a scenario where what we do for a living and how we go about it works without the other.

Mike: Right. Yeah, like, this is one of the interesting things about the company we have built is that it would not work with just you or just me; it’s us being co-founders is what makes it successful.

Corey: The thing that I do not understand and I don’t think I ever will is the idea of co-founder speed dating, where you basically go to some big networking mixer event, pick some rando off the street, and congratulations, that’s your business partner. Have fun. It is not that much of an exaggeration to say that co-founding a company with someone else is like a marriage. You are creating a legal entity that without very specific controls and guidelines, you are opening yourself up to massive liability issues if the other person decides to screw you over. That is part of the reason that the values match was so important for us.

Mike: Yeah, it is surprising to me how similar being co-founders and business partners is to being married. I did not expect how close those two things were. You and I spend an incredible amount of time just on the relationship for each of us, which I never expected, but makes sense in hindsight.

Corey: That’s I think part of it makes the whole you managing me type of relationship work is because not only can you not, “Fire me,” quote-unquote, but I can’t quit without—

Mike: [laugh].

Corey: Leaving behind a giant pile of effort with nothing to show for it over the last four years. So, it’s one of those conversation styles where we go into the conversation knowing, regardless of how heated it gets or how annoyed we are with each other, that we are not going to blow the company up because one of us is salty that week.

Mike: Right. Yeah, I remember from the legal perspective, when we put together a partnership agreement, our attorneys were telling us that we really needed to have someone at the 51% owner, and we were both adamant that no, that doesn’t work for us. And finally, the way that we handled it is if you and I could not handle a dispute, then the only remedy left was to shut the entire thing down. And that would be an automatic trigger. We’ve never ever, ever even got close to that point.

But like, I like that’s the structure because it really means that if you and I can’t agree on something and it’s a substantial thing, then there’s no business, which really kind of sets the stage for how important the conversations that we have are. And of course, you and I, we’re close, we have a great relationship, so that’s never come up. But I do like that it’s still there.

Corey: I like the fact that there’s always going to be an option to get out. It’s not a suicide pact, for lack of a better term. But it’s also something that neither one of us would ever entertain lightly. And credit where due, there have been countless conversations where you and I were diametrically opposed; we each talk through it, and one or the other of us will just do a complete one-eighty our position where, “Okay, you convinced me,” and that’s it. What’s so odd about that is because we don’t have too many examples of that in public society, it just seems like there’s now this entire focus on, “Oh, if you make an observation or a point, that’s wrong, you’ve got to double down on it.” Why would you do that? That makes zero sense. When you’ve considered something of a different angle and change your mind, why waste more time on it?

Mike: I think there’s other interesting ones, too, where you and I have come at something from a different angle and one of us will realize that we just actually don’t care as much as we thought we did. And we’ll just back down because it’s not the hill we want to die on.

Corey: Which brings us to a good point. What hill do we want to die on?

Mike: Hmm. I think we’ve only got a handful. I mean, as it should; like, there should not be there should not be many of them.

Corey: No, no because most things can change, in the fullness of time. Just because it’s not something we believe is right for the business right now does not mean it never will be.

Mike: Yeah. I think all of them really come down to questions of values, which is why you and I worked so well together, in that we don’t have a lot of common interests, we’re at completely different stages in our lives, but we have very tightly aligned values. Which means that when we go into a discussion about something, we know where the other stands right away, like, we could generally make a pretty good guess about it. And there’s often very little question about how some values discussion is going to go. Like, do we take on a certain client that is, I don’t know, they build landmines? Is that a thing that we’re going to do? Of course not. Like—

Corey: I should clarify, we’re talking here about physical landmines; not whatever disastrous failure mode your SaaS application has.

Mike: [laugh]. Yeah.

Corey: We know what those are.

Mike: Yeah, and like, that sort of thing, you and I would never even pose the question to each other. We would just make the decision. And maybe we tell each other later because and, like, “Hey, haha, look what happened,” but there will never be a discussion around it because it just—our values are so tightly aligned that it wouldn’t be necessary.

Corey: Whenever we’re talking to someone that’s in a new sector or a company that has a different expression, we always like to throw it past each other just to double-check, you don’t have a problem with—insert any random thing here; the breadth of our customer base just astounds me—and very rarely as either one of us thrown a flag on something just because we do have this affinity for saying[ yes and making money.

Mike: Yeah. But you actually wanted to talk about the terribleness of managing you.

Corey: Yeah. I am very curious as to what your experience has been.

Mike: [laugh].

Corey: And before we dive into it, I want to call out a couple of things that make me a little atypical for your typical problem employee. I am ADHD personified. My particular expression of that means that my energy level is very different at different times of day, there are times where I will get nothing done for a day or two, and then in four hours, get three weeks of work done. It is hard to predict and it’s hard to schedule around and it’s never clear exactly what that energy level is going to be at any given point in time. That’s the starting point of this nonsense. Now, take it away.

Mike: Yeah. What most people know about Corey is what everyone sees on Twitter, which is what I would call the high highs. Everyone sees you as your most energetic, or at least perceived as the most energetic. If they see you in person at a conference, it’s the same sort of thing. What people don’t see are your lows, which are really, really low lows.

And it’s not a matter of, like, you don’t get anything done. Like, you know, we can handle that; it’s that you disappear. And it may be for a couple hours, it may be for a couple of days, and we just don’t really know what’s going on. That’s really hard. But then, in your high highs, they’re really high, but they’re also really unpredictable.

So, what that means is that because you have ADHD, like, the way that your brain thinks, the way your brain works, is that you don’t control what you’re going to focus on, and you never know what you’re going to focus on. It may be exactly what you should be focusing on, which is a huge win for everyone involved, but sometimes you focus on stuff that doesn’t matter to anyone except you. Sometimes really interesting stuff comes out of that, but oftentimes it doesn’t. So, helping build a structure to work around those sorts of things and to also support those sorts of things, has been one of the biggest challenges that I’ve had. And most of my job is really about building a support structure for you and enabling you to do your best work.

So, that’s been really interesting and really challenging because I do not think that way. Like, if I need to focus on something, I just say, “Great. I’m just going to focus on this thing,” and I’ll focus on it until I’m done. But you don’t work that way, and you couldn’t conceivably work that way, ever. So, it’s always been hard because I say things like, “Hey, Corey, I need you to go write this series of emails.” And you’ll write them when your brain decides that wants to write them, which might be never.

Corey: That’s part of the problem. I’ve also found that if I have an idea floating around too long, it’ll linger for years and I’ll never write anything about it, whereas there are times when I have—the inspiration strikes, I write a one- to 2000-word blog post every week that goes out, and there are times it takes me hours and there are times I bust out the entire thing in first draft form in 20 minutes or less. Like, if it’s Domino’s, like, there’s not going to be a refund on it. So, it’s kind of wild and I wish I could harness that somehow I don’t know how, but… that’s one of the biggest challenges.

Mike: I wish I could too, but it’s one of the things that you learn to get used to. And with that, because we’ve worked together for so long, I’ve gotten to be able to tell in what state of mind you are. Like, are you in a state where if I put something in front of you, you’re going to go after it hard, and like, great things are going to happen, or are you more likely to ignore that I said anything? And I can generally tell within the first sentence or so of bringing something up. But that also means that I have other—I have to be careful with how I structure requests that I have for you.

In some cases, I come with a punch list of, like, here’s six things I need to get through and I’m going to sit on this call while we go through them. In other cases, I have to drip them out one at a time over the span of a week just because that’s how your mind is those days. That makes it really difficult because that’s not how most people are managed and it’s not how most people expect to manage. So, coming up with different ways to do that has been one of the trickiest things I’ve done.

Corey: Let’s move on a little bit other than managing my energy levels because that does not sound like a particularly difficult employee to manage. “Okay, great. We’ve got to build some buffer room into the schedule in case he winds up not delivering for a few days. Okay, we can live with that.” But oh, working with me gets so much worse.

Mike: [laugh]. It absolutely does.

Corey: This is my performance review. Please hit me with it.

Mike: Yeah. The other major concern that has been challenging to work through that makes you really frustrating to work with, is you hate conflict. Actually, I don’t actually—let me clarify that further. You avoid conflict, except your definition of conflict is more broad than most. Because when most people think of conflicts, like, “Oh, I have to go have this really hard conversation, it’s going to be uncomfortable, and, like—”

Corey: “Time to go fire Steven.”

Mike: Right, or things like, “I have to have our performance conversation with someone.” Like, everyone hates those, but, like, there’s good ways and bad ways to them, like, it’s uncomfortable even at the best of times. But with you, it’s more than that, it’s much more broad. You avoid giving direction because you perceive giving direction as potential for conflict, and because you’re so conflict-avoidant, you don’t give direction to people.

Which means that if someone does something you don’t like, you don’t say anything and then it leaves everyone on the team to say, like, “I really wish Corey would be more explicit about what he wants. I wish he was more vocal about the direction he wanted to go.” Like, “Please tell us something more.” But you’re so conflict-avoidant that you don’t, and no amount of begging or we’re asking for it has really changed that, so we end up with these two things where you’re doing most of the work yourself because you don’t want to direct other people to do it.

Corey: I will push back slightly on one element of that, which is when I have a strong opinion about something, I am not at all hesitant about articulating that. I mean, this is not—like, my Twitter is not performance art; it’s very much what I believe. The challenge is that for so much of what we talk about internally on a day-to-day basis, I don’t really have a strong opinion. And what I’ve always shied away from is the idea of telling people how to do their jobs. So, I want to be very clear that I’m not doing that, except when it’s important.

Because we’ve all been in environments in the corporate world where the president of the company wanders past or your grand-boss walks into the room and asks an idle question, or, “Maybe we should do this,” and it never feels like it’s really just idle pondering. It’s, “Welp, new strategic priority just dropped from on high.”

Mike: Right.

Corey: And every senior manager has a story about screwing that one up. And I have led us down that path once or twice previously. So—

Mike: That’s true.

Corey: When I don’t have a strong opinion, I think what I need to get better at is saying, “I don’t give a shit,” but when I frame it like that it causes different problems.

Mike: Yeah. Yeah, that’s very true. I still don’t completely agree with your disagreement there, but I understand your perspective. [laugh].

Corey: Oh, he’s not like you can fire me, so it doesn’t really matter. I kid. I kid.

Mike: Right. Yeah. So, I think those are the two major areas that make you a real challenge to manage and a challenge to direct. But one of the reasons why I think we’ve been successful at it, or at least I’ll say I’ve been successful at managing you, is I do so with such a gentle touch that you don’t realize that I’m doing anything, and I have all these different—

Corey: Well, it did take me four years to realize what was going on.

Mike: Yeah, like, I have all these different ways of getting you to do things, and you don’t realize I’m doing them. And, like, I’ve shared many of them here for you for the first time. And that’s really is what has worked out well. Like, a lot of the ways that I manage you, you don’t realize are management.

Corey: Managing shards. Maintenance windows. Overprovisioning. ElastiCache bills. I know, I know. It's a spooky season and you're already shaking. It's time for caching to be simpler. Momento Serverless Cache lets you forget the backend to focus on good code and great user experiences. With true autoscaling and a pay-per-use pricing model, it makes caching easy. No matter your cloud provider, get going for free at gomomento.co/screaming That's GO M-O-M-E-N-T-O dot co slash screaming

Corey: What advice would you have for someone for whom a lot of these stories are resonating? Because, “Hey, I have a direct report is driving me to distraction and a lot sounds like what you’re describing.” What do you wish you’d known sooner about how to coax performance out of me, for lack of a better phrasing?

Mike: When we first started really working together, I knew what ADHD was, but I knew it from a high school paper that I did on ADHD, and it’s um—oh, what was it—“The Overdiagnosis of ADHD,” which was a thing when you and I were at high school. That’s all I knew is just that ADHD was suspected to be grossly overdiagnosed and that most people didn’t have it. What I have learned is that yeah, that might have been true—maybe; I don’t know—but for people that do have any ADHD, it’s a real thing. Like, it does have some pretty substantial impact.

And I wish I had known more about how that manifests, and particularly how it manifests in different people. And I wish I’d known more earlier on about the coping mechanisms that different people create for themselves and how they manage and how they—[sigh], I’m struggling to come up with the right word here, but many people who are neurodivergent in some way create coping mechanisms and ways to shift themselves to appear more neurotypical. And I wish I had understood that better. Particularly, I wish I had understood that better for you when we first started because I’ve kind of learned about it over time. And I spent so much time trying to get you to work the way that I work rather than understand that you work different. Had I spent more time to understand how you work and what your coping mechanisms were, the earlier years of Duckbill would have been so much smoother.

Corey: And again, having this conversation has been extraordinarily helpful. On my side of it, one of the things that was absolutely transformative and caused a massive reduction in our interpersonal conflict was the very simple tool of, it’s not necessarily a problem when I drop something on the floor and don’t get to it, as long as I throw a hand up and say, “I’m dropping this thing,” and so someone else can catch it as we go. I don’t know how much of this is ADHD speaking versus how much of it is just my own brokenness in some ways, but I feel like everyone has this neverending list of backlog tasks that they’ll get to someday that generally doesn’t ever seem to happen. More often than not, I wind up every few months, just looking at my ever-growing list, reset to zero and we’ll start over. And every once in a while, I’ll be really industrious and knock a thing or two off the list. But so many that don’t necessarily matter or need to be me doing them, but it drives people to distraction when something hits my email inbox, it just dies there, for example.

Mike: Yeah. One of the systems that we set up here is that if there’s something that Corey does not immediately want to do, I have you send it to someone else. And generally it’s to me and then I become a router for you. But making that more explicit and making that easier for you—I’m just like, “If this is not something that you’re going to immediately take care of yourself, forward it to me.” And that was huge. But then other things, like when you take time off, no one knows you’re taking time off. And it’s an—the easiest thing is no one cares that you’re taking time off; just, you know, tell us you’re doing it.

Corey: Yeah, there’s a difference between, “I’m taking three days off,” and your case, the answer is generally, “Oh, thank God. He’s finally using some of that vacation.”

Mike: [laugh].

Corey: The problem is there’s a world of difference between, “Oh, I’m going to take these three days off,” and just not showing up that day. That tends to cause problems for people.

Mike: Yeah. They’re just waving a hand in the air and saying, “Hey, this is happening,” that’s great. But not waving it, not saying anything at all, that’s where the pain really comes from.

Corey: When you take a look across your experience managing people, which to my understanding your first outing with it was at this company—

Mike: Yeah.

Corey: What about managing me is the least surprising and the most surprising that you’ve picked up during that pattern? Because again, the story has always been, “Oh, yeah, you’re a terrible manager because you’ve never done it before,” but I look back and you’re clearly the best manager I’ve ever had, if for no other reason than neither one of us can rage-quit. But there’s a lot of artistry to how you’ve handled a lot of challenges that I present to you.

Mike: I’m the best manager you’ve had because I haven’t fired you. [laugh].

Corey: And also, some of the best ones I have had fired me. That doesn’t necessarily disqualify someone.

Mike: Yeah. I want to say, I am by no means experienced as a manager. As you mentioned, this is my first outing into doing management. As my coach tells me, I’m getting better every day. I am not terrible [laugh].

The—let’s see—most surprising; least surprising. I don’t think I have anything for least surprising. I think most surprising is how easy it is for you to accept feedback and how quickly you do something about it, how quickly you take action on that feedback. I did not expect that, given all your other proclivities for not liking managers, not liking to be managed, for someone to give feedback to you and you say, “Yep, that sounds good,” and then do it, like, that was incredibly surprising.

Corey: It’s one of those areas where if you’re not embracing or at least paying significant attention to how you are being perceived, maybe that’s a problem, maybe it’s not, let’s be very clear. However, there’s also a lot of propensity there to just assume, “Oh, I’m right and screw everyone else.” You can do an awful lot of harm that way. And that is something I’ve had to become incredibly aware of, especially during the pandemic, as the size of my audience at this point more than quadrupled from the start of the pandemic. These are a bunch of people now who have never met me in person, they have no context on what I do.

And I tend to view the world the way you might expect a dog to behave, who caught a car that he has absolutely no idea how to drive, and he’s sort of winging it as he goes. Like, step one, let’s not kill people. Step two, eh, we’ll figure that out later. Like, step one is the most important.

Mike: Mm-hm. Yeah.

Corey: And feedback is hard to get, past a certain point. I often lament from time to time that it’s become more challenging for me to weed out who the jerks are because when you’re perceived to have a large platform and more or less have no problem calling large companies and powerful folk to account, everyone’s nice to you. And well, “Really? He’s terrible and shitty to women. That’s odd. He’s always been super nice to me.” Is not the glowing defense that so many people seem to think that it is. It’s I have learned to listen a lot more clearly the more I speak.

Mike: That’s a challenge for me as well because, as we’ve mentioned, my first foray into management. As we’ve had more people in the company, that has gotten more of a challenge of I have to watch what I say because my word carries weight on its own, by virtue of my position. And you have the same problem, except yours is much more about your weight in public, rather than your weight internally.

Corey: I see it as different sides of the same coin. I take it as a personal bit of a badge of honor that almost every person I meet, including the people who’ve worked here, have come away, very surprised by just how true to life my personality on Twitter is to how actually am when I interact with humans. You’re right, they don’t see the low sides, but I also try not to take that out on the staff either.

Mike: [laugh]. Right.

Corey: We do the best of what we have, I think, and it’s gratifying to know that I can still learn new tricks.

Mike: Yeah. And I’m not firing anytime soon.

Corey: That’s right. Thank you again for giving me the shotgun performance review. It’s always appreciated. If people want to learn more, where can they find you, to get their own performance preview, perhaps?

Mike: Yeah, you can find me on Twitter at @Mike_Julian. Or you can sign up for our newsletter, where I’m talking about my upcoming book on consulting at mikejulian.com.

Corey: And we will put links to that into the show notes. Thanks again, sir.

Mike: Thank you.

Corey: Mike Julian, CEO of The Duckbill Group, my business partner, and apparently my boss. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry comment that demonstrates the absolute worst way to respond to a negative performance evaluation.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Chetan

Chetan Venkatesh is a technology startup veteran focused on distributed data, edge computing, and software products for enterprises and developers. He has 20 years of experience in building primary data storage, databases, and data replication products. Chetan holds a dozen patents in the area of distributed computing and data storage.

Chetan is the CEO and Co-Founder of Macrometa – a Global Data Network featuring a Global Data Mesh, Edge Compute, and In-Region Data Protection. Macrometa helps enterprise developers build real-time apps and APIs in minutes – not months.

Links Referenced:

  • Macrometa: https://www.macrometa.com
  • Macrometa Developer Week: https://www.macrometa.com/developer-week

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Forget everything you know about SSH and try Tailscale. Imagine if you didn't need to manage PKI or rotate SSH keys every time someone leaves. That'd be pretty sweet, wouldn't it? With Tailscale SSH, you can do exactly that. Tailscale gives each server and user device a node key to connect to its VPN, and it uses the same node key to authorize and authenticate SSH.

Basically you're SSHing the same way you manage access to your app. What's the benefit here? Built in key rotation permissions is code connectivity between any two devices, reduce latency and there's a lot more, but there's a time limit here. You can also ask users to reauthenticate for that extra bit of security. Sounds expensive?

Nope, I wish it were. tail scales. Completely free for personal use on up to 20 devices. To learn more, visit snark.cloud/tailscale. Again, that's snark.cloud/tailscale

Corey: Managing shards. Maintenance windows. Overprovisioning. ElastiCache bills. I know, I know. It's a spooky season and you're already shaking. It's time for caching to be simpler. Momento Serverless Cache lets you forget the backend to focus on good code and great user experiences. With true autoscaling and a pay-per-use pricing model, it makes caching easy. No matter your cloud provider, get going for free at gomomento.co/screaming That's GO M-O-M-E-N-T-O dot co slash screaming

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Today, this promoted guest episode is brought to us basically so I can ask a question that has been eating at me for a little while. That question is, what is the edge? Because I have a lot of cynical sarcastic answers to it, but that doesn’t really help understanding. My guest today is Chetan Venkatesh, CEO and co-founder at Macrometa. Chetan, thank you for joining me.

Chetan: It’s my pleasure, Corey. You’re one of my heroes. I think I’ve told you this before, so I am absolutely delighted to be here.

Corey: Well, thank you. We all need people to sit on the curb and clap as we go by and feel like giant frauds in the process. So let’s start with the easy question that sets up the rest of it. Namely, what is Macrometa, and what puts you in a position to be able to speak at all, let alone authoritatively, on what the edge might be?

Chetan: I’ll answer the second part of your question first, which is, you know, what gives me the authority to even talk about this? Well, for one, I’ve been trying to solve the same problem for 20 years now, which is build distributed systems that work really fast and can answer questions about data in milliseconds. And my journey’s sort of been like the spiral staircase journey, you know, I keep going around in circles, but the view just keeps getting better every time I do one of these things. So I’m on my fourth startup doing distributed data infrastructure, and this time really focused on trying to provide a platform that’s the antithesis of the cloud. It’s kind of like taking the cloud and flipping it on its head because instead of having a single region application where all your stuff runs in one place, on us-west-1 or us-east-1, what if your apps could run everywhere, like, they could run in hundreds and hundreds of cities around the world, much closer to where your users and devices and most importantly, where interesting things in the real world are happening?

And so we started Macrometa about five years back to build a new kind of distributed cloud—let’s call the edge—that kind of looks like a CDN, a Content Delivery Network, but really brings very sophisticated platform-level primitives for developers to build applications in a distributed way around primitives for compute, primitives for data, but also some very interesting things that you just can’t do in the cloud anymore. So that’s Macrometa. And we’re doing something with edge computing, which is a big buzzword these days, but I’m sure you’ll ask me about that.

Corey: It seems to be. Generally speaking, when I look around and companies are talking about edge, it feels almost like it is a redefining of what they already do to use a term that is currently trending and deep in the hype world.

Chetan: Yeah. You know, I think humans just being biologically social beings just tend to be herd-like, and so when we see a new trend, we like to slap it on everything we have. We did that 15 years back with cloud, if you remember, you know? Everybody was very busy trying to stick the cloud label on everything that was on-prem. Edge is sort of having that edge-washing moment right now.

But I define edge very specifically is very different from the cloud. You know, where the cloud is defined by centralization, i.e., you’ve got a giant hyperscale data center somewhere far, far away, where typically electricity, real estate, and those things are reasonably cheap, i.e., not in urban centers, where those things tend to be expensive.

You know, you have platforms where you run things at scale, it’s sort of a your mess for less business in the cloud and somebody else manages that for you. The edge is actually defined by location. And there are three types of edges. The first edge is the CDN edge, which is historically where we’ve been trying to make things faster with the internet and make the internet scale. So Akamai came about, about 20 years back and created this thing called the CDN that allowed the web to scale. And that was the first killer app for edge, actually. So that’s the first location that defines the edge where a lot of the peering happens between different network providers and the on-ramp around the cloud happens.

The second edge is the telecom edge. That’s actually right next to you in terms of, you know, the logical network topology because every time you do something on your computer, it goes through that telecom layer. And now we have the ability to actually run web services, applications, data, directly from that telecom layer.

And then the third edge is—sort of, people have been familiar with this for 30 years. The third edge is your device, just your mobile phone. It’s your internet gateway and, you know, things that you carry around in your pocket or sit on your desk, where you have some compute power, but it’s very restricted and it only deals with things that are interesting or important to you as a person, not in a broad range. So those are sort of the three things. And it’s not the cloud. And these three things are now becoming important as a place for you to build and run enterprise apps.

Corey: Something that I think is often overlooked here—and this is sort of a natural consequence of the cloud’s own success and the joy that we live in a system that we do where companies are required to always grow and expand and find new markets—historically, for example, when I went to AWS re:Invent, which is a cloud service carnival in the desert that no one in the right mind should ever want to attend but somehow we keep doing, it used to be that, oh, these announcements are generally all aligned with people like me, where I have specific problems and they look a lot like what they’re talking about on stage. And now they’re talking about things that, from that perspective, seem like Looney Tunes. Like, I’m trying to build Twitter for Pets or something close to it, and I don’t understand why there’s so much talk about things like industrial IoT and, “Machine learning,” quote-unquote, and other things that just do not seem to align with. I’m trying to build a web service, like it says on the name of a company; what gives?

And part of that, I think, is that it’s difficult to remember, for most of us—especially me—that what they’re coming out with is not your shopping list. Every service is for someone, not every service is for everyone, so figuring out what it is that they’re talking about and what those workloads look like, is something that I think is getting lost in translation. And in our defense—collective defense—Amazon is not the best at telling stories to realize that, oh, this is not me they’re talking to; I’m going to opt out of this particular thing. You figure it out by getting it wrong first. Does that align with how you see the market going?

Chetan: I think so. You know, I think of Amazon Web Services, or even Google, or Azure as sort of Costco and, you know, Sam’s Wholesale Club or whatever, right? They cater to a very broad audience and they sell a lot of stuff in bulk and cheap. And you know, so it’s sort of a lowest common denominator type of a model. And so emerging applications, and especially emerging needs that enterprises have, don’t necessarily get solved in the cloud. You’ve got to go and build up yourself on sort of the crude primitives that they provide.

So okay, go use your bare basic EC2, your S3, and build your own edgy, or whatever, you know, cutting edge thing you want to build over there. And if enough people are doing it, I’m sure Amazon and Google start to pay interest and you know, develop something that makes it easier. So you know, I agree with you, they’re not the best at this sort of a thing. The edge is phenomenon also that’s orthogonally, and diametrically opposite to the architecture of the cloud and the economics of the cloud.

And we do centralization in the cloud in a big way. Everything is in one place; we make giant piles of data in one database or data warehouse slice and dice it, and almost all our computer science is great at doing things in a centralized way. But when you take data and chop it into 50 copies and keep it in 50 different places on Earth, and you have this thing called the internet or the wide area network in the middle, trying to keep all those copies in sync is a nightmare. So you start to deal with some very basic computer science problems like distributed state and how do you build applications that have a consistent view of that distributed state? So you know, there have been attempts to solve these problems for 15, 18 years, but none of those attempts have really cracked the intersection of three things: a way for programmers to do this in a way that doesn’t blow their heads with complexity, a way to do this cheaply and effectively enough where you can build real-world applications that serve billions of users concurrently at a cost point that actually is economical and make sense, and third, a way to do this with adequate levels of performance where you don’t die waiting for the spinning wheel on your screen to go away.

So these are the three problems with edge. And as I said, you know, me and my team, we’ve been focused on this for a very long while. And me and my co-founder have come from this world and we created a platform very uniquely designed to solve these three problems, the problems of complexity for programmers to build in a distributed environment like this where data sits in hundreds of places around the world and you need a consistent view of that data, being able to operate and modify and replicate that data with consistency guarantees, and then a third one, being able to do that, at high levels of performance, which translates to what we call ultra-low latency, which is human perception. The threshold of human perception, visually, is about 70 milliseconds. Our finest athletes, the best Esports players are about 70 to 80 milliseconds in their twitch, in their ability to twitch when something happens on the screen. The average human is about 100 to 110 milliseconds.

So in a second, we can maybe do seven things at rapid rates. You know, that’s how fast our brain can process it. Anything that falls below 100 milliseconds—especially if it falls into 50 to 70 milliseconds—appears instantaneous to the human mind and we experience it as magic. And so where edge computing and where my platform comes in is that it literally puts data and applications within 50 milliseconds of 90% of humans and devices on Earth and allows now a whole new set of applications where latency and location and the ability to control those things with really fine-grained capability matters. And we can talk a little more about what those apps are in a bit.

Corey: And I think that’s probably an interesting place to dive into at the moment because whenever we talk about the idea of new ways of building things that are aimed at decentralization, first, people at this point automatically have a bit of an aversion to, “Wait, are you talking about some of the Web3 nonsense?” It’s one of those look around the poker table and see if you can spot the sucker, and if you can’t, it’s you. Because there are interesting aspects to that entire market, let’s be clear, but it also seems to be occluded by so much of the grift and nonsense and spam and the rest that, again, sort of characterize the early internet as well. The idea though, of decentralizing out of the cloud is deeply compelling just to anyone who’s really ever had to deal with the egress charges, or even the data transfer charges inside of one of the cloud providers. The counterpoint is it feels that historically, you either get to pay the tax and go all-in on a cloud provider and get all the higher-level niceties, or otherwise, you wind up deciding you’re going to have to more or less go back to physical data centers, give or take, and other than the very baseline primitives that you get to work with of VMs and block storage and maybe a load balancer, you’re building it all yourself from scratch. It seems like you’re positioning this as setting up for a third option. I’d be very interested to hear it.

Chetan: Yeah. And a quick comment on decentralization: good; not so sure about the Web3 pieces around it. We tend to talk about computer science and not the ideology of distributing data. There are political reasons, there are ideological reasons around data and sovereignty and individual human rights, and things like that. There are people far smarter than me who should explain that.

I fall personally into the Nicholas Weaver school of skepticism about Web3 and blockchain and those types of things. And for readers who are not familiar with Nicholas Weaver, please go online. He teaches at UC Berkeley is just one of the finest minds of our time. And I think he’s broken down some very good reasons why we should be skeptical about, sort of, Web3 and, you know, things like that. Anyway, that’s a digression.

Coming back to what we’re talking about, yes, it is a new paradigm, but that’s the challenge, which is I don’t want to introduce a new paradigm. I want to provide a continuum. So what we’ve built is a platform that looks and feels very much like Lambdas, and a poly-model database. I hate the word multi. It’s a pretty dumb word, so I’ve started to substitute ‘multi’ with ‘poly’ everywhere, wherever I can find it.

So it’s not multi-cloud; it’s poly-cloud. And it’s not multi-model; it’s poly-model. Because what we want is a world where developers have the ability to use the best paradigm for solving problems. And it turns out when we build applications that deal with data, data doesn’t just come in one form, it comes in many different forms, it’s polymorphic, and so you need a data platform, that’s also, you know, polyglot and poly-model to be able to handle that. So that’s one part of the problem, which is, you know, we’re trying to provide a platform that provides continuity by looking like a key-value store like Redis. It looks like a document database—

Corey: Or the best database in the world Route 53 TXT records. But please, keep going.

Chetan: Well, we’ve got that too, so [laugh] you know? And then we’ve got a streaming graph engine built into it that kind of looks and behaves like a graph database, like Neo4j, for example. And, you know, it’s got columnar capabilities as well. So it’s sort of a really interesting data platform that is not open-source; it’s proprietary because it’s designed to solve these problems of being able to distribute data, put it in hundreds of locations, keep it all in sync, but it looks like a conventional NoSQL database. And it speaks PostgreSQL, so if you know PostgreSQL, you can program it, you know, pretty easily.

What it’s also doing is taking away the responsibility for engineers and developers to understand how to deal with very arcane problems like conflict resolution in data. I made a change in Mumbai; you made a change in Tokyo; who wins? Our systems in the cloud—you know, DynamoDB, and things like that—they have very crude answers for this something called last writer wins. We’ve done a lot of work to build a protocol that brings you ACID-like consistency in these types of problems and makes it easy to reason with state change when you’ve got an application that’s potentially running in 100 locations and each of those places is modifying the same record, for example.

And then the second part of it is it’s a converged platform. So it doesn’t just provide data; it provides a compute layer that’s deeply integrated directly with the data layer itself. So think of it as Lambdas running, like, stored procedures inside the database. That’s really what it is. We’ve built a very, very specialized compute engine that exposes containers in functions as stored procedures directly on the database.

And so they run inside the context of the database and so you can build apps in Python, Go, your favorite language; it compiles down into a [unintelligible 00:15:02] kernel that actually runs inside the database among all these different polyglot interfaces that we have. And the third thing that we do is we provide an ability for you to have very fine-grained control on your data. Because today, data’s become a political tool; it’s become something that nation-states care a lot about.

Corey: Oh, do they ever.

Chetan: Exactly. And [unintelligible 00:15:24] regulated. So here’s the problem. You’re an enterprise architect and your application is going to be consumed in 15 countries, there are 13 different frameworks to deal with. What do you do? Well, you spin up 13 different versions, one for each country, and you know, build 13 different teams, and have 13 zero-day attacks and all that kind of craziness, right?

Well, data protection is actually one of the most important parts of the edge because, with something like Macrometa, you can build an app once, and we’ll provide all the necessary localization for any region processing, data protection with things like tokenization of data so you can exfiltrate data securely without violating potentially PII sensitive data exfiltration laws within countries, things like that, i.e. It’s solving some really hard problems by providing an opinionated platform that does these three things. And I’ll summarize it as thus, Corey, we can kind of dig into each piece. Our platform is called the Global Data Network. It’s not a global database; it’s a global data network. It looks like a frickin database, but it’s actually a global network available in 175 cities around the world.

Corey: The challenge, of course, is where does the data actually live at rest, and—this is why people care about—well, they’re two reasons people care about that; one is the data residency locality stuff, which has always, honestly for me, felt a little bit like a bit of a cloud provider shakedown. Yeah, build a data center here or you don’t get any of the business of anything that falls under our regulation. The other is, what is the egress cost of that look like? Because yeah, I can build a whole multicenter data store on top of AWS, for example, but minimum, we’re talking two cents, a gigabyte of transfer, even with inside of a region in some cases, and many times that externally.

Chetan: Yeah, that’s the real shakedown: the egress costs [laugh] more than the other example that you talked about over there. But it’s a reality of how cloud pricing works and things like that. What we have built is a network that is completely independent of the cloud providers. We’re built on top of five different service providers. Some of them are cloud providers, some of them are telecom providers, some of them are CDNs.

And so we’re building our global data network on top of routes and capacity provided by transfer providers who have different economics than the cloud providers do. So our cost for egress falls somewhere between two and five cents, for example, depending on which edge locations, which countries, and things that you’re going to use over there. We've got a pretty generous egress fee where, you know, for certain thresholds, there’s no egress charge at all, but over certain thresholds, we start to charge between two to five cents. But even if you were to take it at the higher end of that spectrum, five cents per gigabyte for transfer, the amount of value our platform brings in architecture and reduction in complexity and the ability to build apps that are frankly, mind-boggling—one of my customers is a SaaS company in marketing that uses us to inject offers while people are on their website, you know, browsing. Literally, you hit their website, you do a few things, and then boom, there’s a customized offer for them.

In banking that’s used, for example, you know, you’re making your minimum payments on your credit card, but you have a good payment history and you’ve got a decent credit score, well, let’s give you an offer to give you a short-term loan, for example. So those types of new applications, you know, are really at this intersection where you need low latency, you need in-region processing, and you also need to comply with data regulation. So when you building a high-value revenue-generating app like that egress cost, even at five cents, right, tends to be very, very cheap, and the smallest part of you know, the complexity of building them.

Corey: One of the things that I think we see a lot of is that the tone of this industry is set by the big players, and they have done a reasonable job, by and large, of making anything that isn’t running in their blessed environments, let me be direct, sound kind of shitty, where it’s like, “Oh, do you want to be smart and run things in AWS?”—or GCP? Or Azure, I guess—“Or do you want to be foolish and try and build it yourself out of popsicle sticks and twine?” And, yeah, on some level, if I’m trying to treat everything like it’s AWS and run a crappy analog version of DynamoDB, for example, I’m not going to have a great experience, but if I also start from a perspective of not using things that are higher up the stack offerings, that experience starts to look a lot more reasonable as we start expanding out. But it still does present to a lot of us as well, we’re just going to run things in VM somewhere and treat them just like we did back in 2005. What’s changed in that perspective?

Chetan: Yeah, you know, I can’t talk for others but for us, we provide a high-level Platform-as-a-Service, and that platform, the global data network, has three pieces to it. First piece is—and none of this will translate into anything that AWS or GCP has because this is the edge, Corey, is completely different, right? So the global data network that we have is composed of three technology components. The first one is something that we call the global data mesh. And this is Pub/Sub and event processing on steroids. We have the ability to connect data sources across all kinds of boundaries; you’ve got some data in Germany and you’ve got some data in New York. How do you put these things together and get them streaming so that you can start to do interesting things with correlating this data, for example?

And you might have to get across not just physical boundaries, like, they’re sitting in different systems in different data centers; they might be logical boundaries, like, hey, I need to collaborate with data from my supply chain partner and we need to be able to do something that’s dynamic in real-time, you know, to solve a business problem. So the global data mesh is a way to very quickly connect data wherever it might be in legacy systems, in flat files, in streaming databases, in data warehouses, what have you—you know, we have 500-plus types of connectors—but most importantly, it’s not just getting the data streaming, it’s then turning it into an API and making that data fungible. Because the minute you put an API on it and it’s become fungible now that data is actually got a lot of value. And so the data mesh is a way to very quickly connect things up and put an API on it. And that API can now be consumed by front-ends, it can be consumed by other microservices, things like that.

Which brings me to the second piece, which is edge compute. So we’ve built a compute runtime that is Docker compatible, so it runs containers, it’s also Lambda compatible, so it runs functions. Let me rephrase that; it’s not Lambda-compatible, it’s Lambda-like. So no, you can’t take your Lambda and dump it on us and it won’t just work. You have to do some things to make it work on us.

Corey: But so many of those things are so deeply integrated to the ecosystem that they’re operating within, and—

Chetan: Yeah.

Corey: That, on the one hand, is presented by cloud providers as, “Oh, yes. This shows how wonderful these things are.” In practice, talk to customers. “Yeah, we’re using it as spackle between the different cloud services that don’t talk to one another despite being made by the same company.”

Chetan: [laugh] right.

Corey: It’s fun.

Chetan: Yeah. So the second edge compute piece, which allows you now to build microservices that are stateful, i.e., they have data that they interact with locally, and schedule them along with the data on our network of 175 regions around the world. So you can build distributed applications now.

Now, your microservice back-end for your banking application or for your HR SaaS application or e-commerce application is not running in us-east-1 and Virginia; it’s running literally in 15, 18, 25 cities where your end-users are, potentially. And to take an industrial IoT case, for example, you might be ingesting data from the electricity grid in 15, 18 different cities around the world; you can do all of that locally now. So that’s what the edge functions does, it flips the cloud model around because instead of sending data to where the compute is in the cloud, you’re actually bringing compute to where the data is originating, or the data is being consumed, such as through a mobile app. So that’s the second piece.

And the third piece is global data protection, which is hey, now I’ve got a distributed infrastructure; how do I comply with all the different privacy and regulatory frameworks that are out there? How do I keep data secure in each region? How do I potentially share data between regions in such a way that, you know, I don’t break the model of compliance globally and create a billion-dollar headache for my CIO and CEO and CFO, you know? So that’s the third piece of capabilities that this provides.

All of this is presented as a set of serverless APIs. So you simply plug these APIs into your existing applications. Some of your applications work great in the cloud. Maybe there are just parts of that app that should be on our edge. And that’s usually where most customers start; they take a single web service or two that’s not doing so great in the cloud because it’s too far away; it has data sensitivity, location sensitivity, time sensitivity, and so they use us as a way to just deal with that on the edge.

And there are other applications where it’s completely what I call edge native, i.e., no dependancy on the cloud comes and runs completely distributed across our network and consumes primarily the edges infrastructure, and just maybe send some data back on the cloud for long-term storage or long-term analytics.

Corey: And ingest does remain free. The long-term analytics, of course, means that once that data is there, good luck convincing a customer to move it because that gets really expensive.

Chetan: Exactly, exactly. It’s a speciation—as I like to say—of the cloud, into a fast tier where interactions happen, i.e., the edge. So systems of record are still in the cloud; we still have our transactional systems over there, our databases, data warehouses.

And those are great for historical types of data, as you just mentioned, but for things that are operational in nature, that are interactive in nature, where you really need to deal with them because they’re time-sensitive, they’re depleting value in seconds or milliseconds, they’re location sensitive, there’s a lot of noise in the data and you need to get to just those bits of data that actually matter, throw the rest away, for example—which is what you do with a lot of telemetry in cybersecurity, for example, right—those are all the things that require a new kind of a platform, not a system of record, a system of interaction, and that’s what the global data network is, the GDN. And these three primitives, the data mesh, Edge compute, and data protection, are the way that our APIs are shaped to help our enterprise customers solve these problems. So put it another way, imagine ten years from now what DynamoDB and global tables with a really fast Lambda and Kinesis with actually Event Processing built directly into Kinesis might be like. That’s Macrometa today, available in 175 cities.

Corey: This episode is brought to us in part by our friends at Datadog. Datadog is a SaaS monitoring and security platform that enables full-stack observability for modern infrastructure and applications at every scale. Datadog enables teams to see everything: dashboarding, alerting, application performance monitoring, infrastructure monitoring, UX monitoring, security monitoring, dog logos, and log management, in one tightly integrated platform. With 600-plus out-of-the-box integrations with technologies including all major cloud providers, databases, and web servers, Datadog allows you to aggregate all your data into one platform for seamless correlation, allowing teams to troubleshoot and collaborate together in one place, preventing downtime and enhancing performance and reliability. Get started with a free 14-day trial by visiting datadoghq.com/screaminginthecloud, and get a free t-shirt after installing the agent.

Corey: I think it’s also worth pointing out that it’s easy for me to fall into a trap that I wonder if some of our listeners do as well, which is, I live in, basically, downtown San Francisco. I have gigabit internet connectivity here, to the point where when it goes out, it is suspicious and more a little bit frightening because my ISP—Sonic.net—is amazing and deserves every bit of praise that you never hear any ISP ever get. But when I travel, it’s a very different experience. When I go to oh, I don’t know, the conference center at re:Invent last year and find that the internet is patchy at best, or downtown San Francisco on Verizon today, I discover that the internet is almost non-existent, and suddenly applications that I had grown accustomed to just working suddenly didn’t.

And there’s a lot more people who live far away from these data center regions and tier one backbones directly to same than don’t. So I think that there’s a lot of mistaken ideas around exactly what the lower bandwidth experience of the internet is today. And that is something that feels inadvertently classist if that make sense. Are these geographically bigoted?

Chetan: Yeah. No, I think those two points are very well articulated. I wish I could articulate it that well. But yes, if you can afford 5G, some of those things get better. But again, 5G is not everywhere yet. It will be, but 5G can in many ways democratize at least one part of it, which is provide an overlap network at the edge, where if you left home and you switched networks, on to a wireless, you can still get the same quality of service that you used to getting from Sonic, for example. So I think it can solve some of those things in the future. But the second part of it—what did you call it? What bigoted?

Corey: Geographically bigoted. And again, that’s maybe a bit of a strong term, but it’s easy to forget that you can’t get around the speed of light. I would say that the most poignant example of that I had was when I was—in the before times—giving a keynote in Australia. So ah, I know what I’ll do, I’ll spin up an EC2 instance for development purposes—because that’s how I do my development—in Australia. And then I would just pay my provider for cellular access for my iPad and that was great.

And I found the internet was slow as molasses for everything I did. Like, how do people even live here? Well, turns out that my provider would backhaul traffic to the United States. So to log into my session, I would wind up having to connect with a local provider, backhaul to the US, then connect back out from there to Australia across the entire Pacific Ocean, talk to the server, get the response, would follow that return path. It’s yeah, turns out that doing laps around the world is not the most efficient way of transferring any data whatsoever, let alone in sizable amounts.

Chetan: And that’s why we decided to call our platform the global data network, Corey. In fact, it’s really built inside of sort of a very simple reason is that we have our own network underneath all of this and we stop this whole ping-pong effect of data going around and help create deterministic guarantees around latency, around location, around performance. We’re trying to democratize latency and these types of problems in a way that programmers shouldn’t have to worry about all this stuff. You write your code, you push publish, it runs on a network, and it all gets there with a guarantee that 95% of all your requests will happen within 50 milliseconds round-trip time, from any device, you know, in these population centers around the world.

So yeah, it’s a big deal. It’s sort of one of our je ne sais quoi pieces in our mission and charter, which is to just democratize latency and access, and sort of get away from this geographical nonsense of, you know, how networks work and it will dynamically switch topology and just make everything slow, you know, very non-deterministic way.

Corey: One last topic that I want to ask you about—because I near certain given your position, you will have an opinion on this—what’s your take on, I guess, the carbon footprint of clouds these days? Because a lot of people been talking about it; there has been a lot of noise made about, justifiably so. I’m curious to get your take.

Chetan: Yeah, you know, it feels like we’re in the ’30s and the ’40s of the carbon movement when it comes to clouds today, right? Maybe there’s some early awareness of the problem, but you know, frankly, there’s very little we can do than just sort of put a wet finger in the air, compute some carbon offset and plant some trees. I think these are good building blocks; they’re not necessarily the best ways to solve this problem, ultimately. But one of the things I care deeply about and you know, my company cares a lot about is helping make developers more aware off what kind of carbon footprint their code tangibly has on the environment. And so we’ve started two things inside the company. We’ve started a foundation that we call the Carbon Conscious Computing Consortium—the four C’s. We’re going to announce that publicly next year, we’re going to invite folks to come and join us and be a part of it.

The second thing that we’re doing is we’re building a completely open-source, carbon-conscious computing platform that is built on real data that we’re collecting about, to start with, how Macrometa’s platform emits carbon in response to different types of things you build on it. So for example, you wrote a query that hits our database and queries, you know, I don’t know, 20 billion objects inside of our database. It’ll tell you exactly how many micrograms or how many milligrams of carbon—it’s an estimate; not exactly. I got to learn to throttle myself down. It’s an estimate, you know, you can’t really measure these things exactly because the cost of carbon is different in different places, you know, there are different technologies, et cetera.

Gives you a good decent estimate, something that reliably tells you, “Hey, you know that query that you have over there, that piece of SQL? That’s probably going to do this much of micrograms of carbon at this scale.” You know, if this query was called a million times every hour, this is how much it costs. A million times a day, this is how much it costs and things like that. But the most important thing that I feel passionate about is that when we give developers visibility, they do good things.

I mean, when we give them good debugging tools, the code gets better, the code gets faster, the code gets more efficient. And Corey, you’re in the business of helping people save money, when we give them good visibility into how much their code costs to run, they make the code more efficient. So we’re doing the same thing with carbon, we know there’s a cost to run your code, whether it’s a function, a container, a query, what have you, every operation has a carbon cost. And we’re on a mission to measure that and provide accurate tooling directly in our platform so that along with your debug lines, right, where you’ve got all these print statements that are spitting up stuff about what’s happening there, we can also print out, you know, what did it cost in carbon.

And you can set budgets. You can basically say, “Hey, I want my application to consume this much of carbon.” And down the road, we’ll have AI and ML models that will help us optimize your code to be able to fit within those carbon budgets. For example. I’m not a big fan of planting—you know, I love planting trees, but don’t get me wrong, we live in California and those trees get burned down.

And I was reading this heartbreaking story about how we returned back into the atmosphere a giant amount of carbon because the forest reserve that had been planted, you know, that was capturing carbon, you know, essentially got burned down in a forest fire. So, you know, we’re trying to just basically say, let’s try and reduce the amount of carbon, you know, that we can potentially create by having better tooling.

Corey: That would be amazing, and I think it also requires something that I guess acts almost as an exchange where there’s a centralized voice that can make sure that, well, one, the provider is being honest, and two, being able to ensure you’re doing an apples-to-apples comparison and not just discounting a whole lot of negative externalities. Because, yes, we’re talking about carbon released into the environment. Okay, great. What about water effects from what’s happening with your data centers are located? That can have significant climate impact as well. It’s about trying to avoid the picking and choosing. It’s hard, hard problem, but I’m unconvinced that there’s anything more critical in the entire ecosystem right now to worry about.

Chetan: So as a startup, we care very deeply about starting with the carbon part. And I agree, Corey, it’s a multi-dimensional problem; there’s lots of tentacles. The hydrocarbon industry goes very deeply into all parts of our lives. I’m a startup, what do I know? I can’t solve all of those things, but I wanted to start with the philosophy that if we provide developers with the right tooling, they’ll have the right incentives then to write better code. And as we open-source more of what we learn and, you know, our tooling, others will do the same. And I think in ten years, we might have better answers. But someone’s got to start somewhere, and this is where we’d like to start.

Corey: I really want to thank you for taking as much time as you have for going through what you’re up to and how you view the world. If people want to learn more, where’s the best place to find you?

Chetan: Yes, so two things on that front. Go to www.macrometa.com—M-A-C-R-O-M-E-T-A dot com—and that’s our website. And you can come and experience the full power of the platform. We’ve got a playground where you can come, open an account and build anything you want for free, and you can try and learn. You just can’t run it in production because we’ve got a giant network, as I said, of 175 cities around the world. But there are tiers available for you to purchase and build and run apps. Like I think about 80 different customers, some of the biggest ones in the world, some of the biggest telecom customers, retail, E-Tail customers, [unintelligible 00:34:28] tiny startups are building some interesting things on.

And the second thing I want to talk about is November 7th through 11th of 2022, just a couple of weeks—or maybe by the time this recording comes out, a week from now—is developer week at Macrometa. And we’re going to be announcing some really interesting new capabilities, some new features like real-time complex event processing with low, ultra-low latency, data connectors, a search feature that allows you to build search directly on top of your applications without needing to spin up a giant Elastic Cloud Search cluster, or providing search locally and regionally so that, you know, you can have search running in 25 cities that are instant to search rather than sending all your search requests back in one location. There’s all kinds of very cool things happening over there.

And we’re also announcing a partnership with the original, the OG of the edge, one of the largest, most impressive, interesting CDN players that has become a partner for us as well. And then we’re also announcing some very interesting experimental work where you as a developer can build apps directly on the 5G telecom cloud as well. And then you’ll hear from some interesting companies that are building apps that are edge-native, that are impossible to build in the cloud because they take advantage of these three things that we talked about: geography, latency, and data protection in some very, very powerful ways. So you’ll hear actual customer case studies from real customers in the flesh, not anonymous BS, no marchitecture. It’s a week-long of technical talk by developers, for developers. And so, you know, come and join the fun and let’s learn all about the edge together, and let’s go build something together that’s impossible to do today.

Corey: And we will, of course, put links to that in the [show notes 00:36:06]. Thank you so much for being so generous with your time. I appreciate it.

Chetan: My pleasure, Corey. Like I said, you’re one of my heroes. I’ve always loved your work. The Snark-as-a-Service is a trillion-dollar market cap company. If you’re ever interested in taking that public, I know some investors that I’d happily put you in touch with. But—

Corey: Sadly, so many of those investors lack senses of humor.

Chetan: [laugh]. That is true. That is true [laugh].

Corey: [laugh]. [sigh].

Chetan: Well, thank you. Thanks again for having me.

Corey: Thank you. Chetan Venkatesh, CEO and co-founder at Macrometa. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an angry and insulting comment about why we should build everything on the cloud provider that you work for and then the attempt to challenge Chetan for the title of Edgelord.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Kevin

Kevin Miller is currently the global General Manager for Amazon Simple Storage Service (S3), an object storage service that offers industry-leading scalability, data availability, security, and performance. Prior to this role, Kevin has had multiple leadership roles within AWS, including as the General Manager for Amazon S3 Glacier, Director of Engineering for AWS Virtual Private Cloud, and engineering leader for AWS Virtual Private Network and AWS Direct Connect. Kevin was also Technical Advisor to the Senior Vice President for AWS Utility Computing. Kevin is a graduate of Carnegie Mellon University with a Bachelor of Science in Computer Science.

Links Referenced:

  • snark.cloud/shirt: https://snark.cloud/shirt
  • aws.amazon.com/s3: https://aws.amazon.com/s3

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is brought to us in part by our friends at Datadog. Datadog is a SaaS monitoring and security platform that enables full-stack observability for modern infrastructure and applications at every scale. Datadog enables teams to see everything: dashboarding, alerting, application performance monitoring, infrastructure monitoring, UX monitoring, security monitoring, dog logos, and log management, in one tightly integrated platform. With 600-plus out-of-the-box integrations with technologies including all major cloud providers, databases, and web servers, Datadog allows you to aggregate all your data into one platform for seamless correlation, allowing teams to troubleshoot and collaborate together in one place, preventing downtime and enhancing performance and reliability. Get started with a free 14-day trial by visiting datadoghq.com/screaminginthecloud, and get a free t-shirt after installing the agent.

Corey: Managing shards. Maintenance windows. Overprovisioning. ElastiCache bills. I know, I know. It’s a spooky season and you’re already shaking. It’s time for caching to be simpler. Momento Serverless Cache lets you forget the backend to focus on good code and great user experiences. With true autoscaling and a pay-per-use pricing model, it makes caching easy. No matter your cloud provider, get going for free at gomomento.co/screaming. That’s GO M-O-M-E-N-T-O dot co slash screaming.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Right now, as I record this, we have just kicked off our annual charity t-shirt fundraiser. This year’s shirt showcases S3 as the eighth wonder of the world. And here to either defend or argue the point—we’re not quite sure yet—is Kevin Miller, AWS’s vice president and general manager for Amazon S3. Kevin, thank you for agreeing to suffer the slings and arrows that are no doubt going to be interpreted, misinterpreted, et cetera, for the next half hour or so.

Kevin: Oh, Corey, thanks for having me. And happy to do that, and really flattered for you to be thinking about S3 in this way. So more than happy to chat with you.

Corey: It’s absolutely one of those services that is foundational to the cloud. It was the first AWS service that was put into general availability, although the beta folks are going to argue back and forth about no, no, that was SQS instead. I feel like now that Mai-Lan handles both SQS and S3 as part of her portfolio, she is now the final arbiter of that. I’m sure that’s an argument for a future day. But it’s impossible to imagine cloud without S3.

Kevin: I definitely think that’s true. It’s hard to imagine cloud, actually, with many of our foundational services, including SQS, of course, but we are—yes, we were the first generally available service with S3. And pretty happy with our anniversary being Pi Day, 3/14.

Corey: I’m also curious, your own personal trajectory has been not necessarily what folks would expect. You were the general manager of Amazon Glacier, and now you’re the general manager and vice president of S3. So, I’ve got to ask, because there are conflicting reports on this depending upon what angle you look at, are Glacier and S3 the same thing?

Kevin: Yes, I was the general manager for S3 Glacier prior to coming over to S3 proper, and the answer is no, they are not the same thing. We certainly have a number of technologies where we’re able to use those technologies both on S3 and Glacier, but there are certainly a number of things that are very distinct about Glacier and give us that ability to hit the ultra-low price points that we do for Glacier Deep Archive being as low as $1 per terabyte-month. And so, that definitely—there’s a lot of actual ingenuity up and down the stack, from hardware to software, everywhere in between, to really achieve that with Glacier. But then there’s other spots where S3 and Glacier have very similar needs, and then, of course, today many customers use Glacier through S3 as a storage class in S3, and so that’s a great way to do that. So, there’s definitely a lot of shared code, but certainly, when you get into it, there’s [unintelligible 00:04:59] to both of them.

Corey: I ran a number of obnoxiously detailed financial analyses, and they all came away with, unless you have a very specific very nuanced understanding of your data lifecycle and/or it is less than 30 or 60 days depending upon a variety of different things, the default S3 storage class you should be using for virtually anything is Intelligent Tiering. That is my purely economic analysis of it. Do you agree with that? Disagree with that? And again, I understand that all of these storage classes are like your children, and I am inviting you to tell me which one of them is your favorite, but I’m absolutely prepared to do that.

Kevin: Well, we love Intelligent Tiering because it is very simple; customers are able to automatically save money using Intelligent Tiering for data that’s not being frequently accessed. And actually, since we launched it a few years ago, we’ve already saved customers more than $250 million using Intelligent Tiering. So, I would say today, it is our default recommendation in almost every case. I think that the cases where we would recommend another storage class as the primary storage class tend to be specific to the use case where—and particularly for use cases where customers really have a good understanding of the access patterns. And we saw some customers do for their certain dataset, they know that it’s going to be heavily accessed for a fixed period of time, or this data is actually for archival, it’ll never be accessed, or very rarely if ever access, just maybe in an emergency.

And those kinds of use cases, I think actually, customers are probably best to choose one of the specific storage classes where they’re, sort of, paying that the lower cost from day one. But again, I would say for the vast majority of cases that we see, the data access patterns are unpredictable and customers like the flexibility of being able to very quickly retrieve the data if they decide they need to use it. But in many cases, they’ll save a lot of money as the data is not being accessed, and so, Intelligent Tiering is a great choice for those cases.

Corey: I would take it a step further and say that even when customers believe that they are going to be doing a deeper analysis and they have a better understanding of their data flow patterns than Intelligent Tiering would, in practice, I see that they rarely do anything about it. It’s one of those things where they’re like, “Oh, yeah, we’re going to set up our own lifecycle policies real soon now,” whereas, just switch it over to Intelligent Tiering and never think about it again. People’s time is worth so much more than the infrastructure they’re working on in almost every case. It doesn’t seem to make a whole lot of sense unless you have a very intentioned, very urgent reason to go and do that stuff by hand in most cases.

Kevin: Yeah, that’s right. I think I agree with you, Corey. And certainly, that is the recommendation we lead with customers.

Corey: In previous years, our charity t-shirt has focused on other areas of AWS, and one of them was based upon a joke that I’ve been telling for a while now, which is that the best database in the world is Route 53 and storing TXT records inside of it. I don’t know if I ever mentioned this to you or not, but the first iteration of that joke was featuring around S3. The challenge that I had with it is that S3 Select is absolutely a thing where you can query S3 with SQL which I don’t see people doing anymore because Athena is the easier, more, shall we say, well-articulated version of all of that. And no, no, that joke doesn’t work because it’s actually true. You can use S3 as a database. Does that statement fill you with dread? Regret? Am I misunderstanding something? Or are you effectively running a giant subversive database?

Kevin: Well, I think that certainly when most customers think about a database, they think about a collection of technology that’s applied for given problems, and so I wouldn’t count S3 as providing the whole range of functionality that would really make up a database. But I think that certainly a lot of the primitives and S3 Select as a great example of a primitive are available in S3. And we’re looking at adding, you know, additional primitives going forward to make it possible to, you know, to build a database around S3. And as you see, other AWS services have done that in many ways. For example, obviously with Amazon Redshift having a lot of capability now to just directly access and use data in S3 and make that a super seamless so that you can then run data warehousing type queries on top of S3 and on top of your other datasets.

So, I certainly think it’s a great building block. And one other thing I would actually just say that you may not know, Corey, is that one of the things over the last couple of years we’ve been doing a lot more with S3 is actually working to directly contribute improvements to open-source connector software that uses S3, to make available automatically some of the performance improvements that can be achieved either using both the AWS SDK, and also using things like S3 Select. So, we started with a few of those things with Select; you’re going to see more of that coming, most likely. And some of that, again, the idea there as you may not even necessarily know you’re using Select, but when we can identify that it will improve performance, we’re looking to be able to contribute those kinds of improvements directly—or we are contributing those directly to those open-source packages. So, one thing I would definitely recommend customers and developers do is have a capability of sort of keeping that software up-to-date because although it might seem like those are sort of one-and-done kind of software integrations, there’s actually almost continuous improvement now going on, and around things like that capability, and then others we come out with.

Corey: What surprised me is just how broadly S3 has been adopted by a wide variety of different clients’ software packages out there. Back when I was running production environments in anger, I distinctly remember in one Ubuntu environment, we wound up installing a specific package that was designed to teach apt how to retrieve packages and its updates from S3, which was awesome. I don’t see that anymore, just because it seems that it is so easy to do it now, just with the native features that S3 offers, as well as an awful lot of software under the hood has learned to directly recognize S3 as its own thing, and can react accordingly.

Kevin: And just do the right thing. Exactly. No, we certainly see a lot of that. So that’s, you know—I mean, obviously making that simple for end customers to use and achieve what they’re trying to do, that’s the whole goal.

Corey: It’s always odd to me when I’m talking to one of my clients who is looking to understand and optimize their AWS bill to see outliers in either direction when it comes to S3 itself. When they’re driving large S3 bills as in a majority of their spend, it’s, okay, that is very interesting. Let’s dive into that. But almost more interesting to me is when it is effectively not being used at all. When, oh, we’re doing everything with EBS volumes or EFS.

And again, those are fine services. I don’t have any particular problem with them anymore, but the problem I have is that the cloud long ago took what amounts to an economic vote. There’s a tax savings for storing data in an object store the way that you—and by extension, most of your competitors—wind up pricing this, versus the idea of on a volume basis where you have to pre-provision things, you don’t get any form of durability that extends beyond the availability zone boundary. It just becomes an awful lot of, “Well, you could do it this way. But it gets really expensive really quickly.”

It just feels wild to me that there is that level of variance between S3 just sort of raw storage basis, economically, as well as then just the, frankly, ridiculous levels of durability and availability that you offer on top of that. How did you get there? Was the service just mispriced at the beginning? Like oh, we dropped to zero and probably should have put that in there somewhere.

Kevin: Well, no, I wouldn’t call it mispriced. I think that the S3 came about when we took a—we spent a lot of time looking at the architecture for storage systems, and knowing that we wanted a system that would provide the durability that comes with having three completely independent data centers and the elasticity and capability where, you know, customers don’t have to provision the amount of storage they want, they can simply put data and the system keeps growing. And they can also delete data and stop paying for that storage when they’re not using it. And so, just all of that investment and sort of looking at that architecture holistically led us down the path to where we are with S3.

And we’ve definitely talked about this. In fact, in Peter’s keynote at re:Invent last year, we talked a little bit about how the system is designed under the hood, and one of the thing you realize is that S3 gets a lot of the benefits that we do by just the overall scale. The fact that it is—I think the stat is that at this point more than 10,000 customers have data that’s stored on more than a million hard drives in S3. And that’s how you get the scale and the capability to do is through massive parallelization. Where customers that are, you know, I would say building more traditional architectures, those are inherently typically much more siloed architectures with a relatively small-scale overall, and it ends up with a lot of resource that’s provisioned at small-scale in sort of small chunks with each resource, that you never get to that scale where you can start to take advantage of the some is more than the greater of the parts.

And so, I think that’s what the recognition was when we started out building S3. And then, of course, we offer that as an API on top of that, where customers can consume whatever they want. That is, I think, where S3, at the scale it operates, is able to do certain things, including on the economics, that are very difficult or even impossible to do at a much smaller scale.

Corey: One of the more egregious clown-shoe statements that I hear from time to time has been when people will come to me and say, “We’ve built a competitor to S3.” And my response is always one of those, “Oh, this should be good.” Because when people say that, they generally tend to be focusing on one or maybe two dimensions that doesn’t work for a particular use case as well as it could. “Okay, what was your story around why this should be compared to S3?” “Well, it’s an object store. It has full S3 API compatibility.” “Does it really because I have to say, there are times where I’m not entirely convinced that S3 itself has full compatibility with the way that its API has been documented.”

And there’s an awful lot of magic that goes into this too. “Okay, great. You’re running an S3 competitor. Great. How many buildings does it live in?” Like, “Well, we have a problem with the s at the end of that word.” It’s, “Okay, great. If it fits on my desk, it is not a viable S3 competitor. If it fits in a single zip code, it is probably not a viable S3 competitor.” Now, can it be an object store? Absolutely. Does it provide a new interface to some existing data someone might have? Sure why not. But I think that, oh, it’s S3 compatible, is something that gets tossed around far too lightly by folks who don’t really understand what it is that drives S3 and makes it special.

Kevin: Yeah, I mean, I would say certainly, there’s a number of other implementations of the S3 API, and frankly we’re flattered that customers recognize and our competitors and others recognize the simplicity of the API and go about implementing it. But to your point, I think that there’s a lot more; it’s not just about the API, it’s really around everything surrounding S3 from, as you mentioned, the fact that the data in S3 is stored in three independent availability zones, all of which that are separated by kilometers from each other, and the resilience, the automatic failover, and the ability to withstand an unlikely impact to one of those facilities, as well as the scalability, and you know, the fact that we put a lot of time and effort into making sure that the service continues scaling with our customers need. And so, I think there’s a lot more that goes into what is S3. And oftentimes just in a straight-up comparison, it’s sort of purely based on just the APIs and generally a small set of APIs, in addition to those intangibles around—or not intangibles, but all of the ‘-ilities,’ right, the elasticity and the durability, and so forth that I just talked about. In addition to all that also, you know, certainly what we’re seeing for customers is as they get into the petabyte and tens of petabytes, hundreds of petabytes scale, their need for the services that we provide to manage that storage, whether it’s lifecycle and replication, or things like our batch operations to help update and to maintain all the storage, those become really essential to customers wrapping their arms around it, as well as visibility, things like Storage Lens to understand, what storage do I have? Who’s using it? How is it being used?

And those are all things that we provide to help customers manage at scale. And certainly, you know, oftentimes when I see claims around S3 compatibility, a lot of those advanced features are nowhere to be seen.

Corey: I also want to call out that a few years ago, Mai-Lan got on stage and talked about how, to my recollection, you folks have effectively rebuilt S3 under the hood into I think it was 235 distinct microservices at the time. There will not be a quiz on numbers later, I’m assuming. But what was wild to me about that is having done that for services that are orders of magnitude less complex, it absolutely is like changing the engine on a car without ever slowing down on the highway. Customers didn’t know that any of this was happening until she got on stage and announced it. That is wild to me. I would have said before this happened that there was no way that would have been possible except it clearly was. I have to ask, how did you do that in the broad sense?

Kevin: Well, it’s true. A lot of the underlying infrastructure that’s been part of S3, both hardware and software is, you know, you wouldn’t—if someone from S3 in 2006 came and looked at the system today, they would probably be very disoriented in terms of understanding what was there because so much of it has changed. To answer your question, the long and short of it is a lot of testing. In fact, a lot of novel testing most recently, particularly with the use of formal logic and what we call automated reasoning. It’s also something we’ve talked a fair bit about in re:Invent.

And that is essentially where you prove the correctness of certain algorithms. And we’ve used that to spot some very interesting, the one-in-a-trillion type cases that S3 scale happens regularly, that you have to be ready for and you have to know how the system reacts, even in all those cases. I mean, I think one of our engineers did some calculations that, you know, the number of potential states for S3, sort of, exceeds the number of atoms in the universe or something so crazy. But yet, using methods like automated reasoning, we can test that state space, we can understand what the system will do, and have a lot of confidence as we begin to swap, you know, pieces of the system.

And of course, nothing in S3 scale happens instantly. It’s all, you know, I would say that for a typical engineering effort within S3, there’s a certain amount of effort, obviously, in making the change or in preparing the new software, writing the new software and testing it, but there’s almost an equal amount of time that goes into, okay, and what is the process for migrating from System A to System B, and that happens over a timescale of months, if not years, in some cases. And so, there’s just a lot of diligence that goes into not just the new systems, but also the process of, you know, literally, how do I swap that engine on the system. So, you know, it’s a lot of really hard working engineers that spent a lot of time working through these details every day.

Corey: I still view S3 through the lens of it is one of the easiest ways in the world to wind up building a static web server because you basically stuff the website files into a bucket and then you check a box. So, it feels on some level though, that it is about as accurate as saying that S3 is a database. It can be used or misused or pressed into service in a whole bunch of different use cases. What have you seen from customers that has, I guess, taught you something you didn’t expect to learn about your own service?

Kevin: Oh, I’d say we have those [laugh] meetings pretty regularly when customers build their workloads and have unique patterns to it, whether it’s the type of data they’re retrieving and the access pattern on the data. You know, for example, some customers will make heavy use of our ability to do [ranged gets 00:22:47] on files and [unintelligible 00:22:48] objects. And that’s pretty good capability, but that can be one where that’s very much dependent on the type of file, right, certain files have structure, as far as you know, a header or footer, and that data is being accessed in a certain order. Oftentimes, those may also be multi-part objects, and so making use of the multi-part features to upload different chunks of a file in parallel. And you know, also certainly when customers get into things like our batch operations capability where they can literally write a Lambda function and do what they want, you know, we’ve seen some pretty interesting use cases where customers are running large-scale operations across, you know, billions, sometimes tens of billions of objects, and this can be pretty interesting as far as what they’re able to do with them.

So, for something is sort of what you might—you know, as simple and basics, in some sense, of GET and PUT API, just all the capability around it ends up being pretty interesting as far as how customers apply it and the different workloads they run on it.

Corey: So, if you squint hard enough, what I’m hearing you tell me is that I can view all of this as, “Oh, yeah. S3 is also compute.” And it feels like that as a fast-track to getting a question wrong on one of the certification exams. But I have to ask, from your point of view, is S3 storage? And whether it’s yes or no, what gets you excited about the space that it’s in?

Kevin: Yeah well, I would say S3 is not compute, but we have some great compute services that are very well integrated with S3, which excites me as well as we have things like S3 Object Lambda, where we actually handle that integration with Lambda. So, you’re writing Lambda functions, we’re executing them on the GET path. And so, that’s a pretty exciting feature for me. But you know, to sort of take a step back, what excites me is I think that customers around the world, in every industry, are really starting to recognize the value of data and data at large scale. You know, I think that actually many customers in the world have terabytes or more of data that sort of flows through their fingers every day that they don’t even realize.

And so, as customers realize what data they have, and they can capture and then start to analyze and make ultimately make better business decisions that really help drive their top line or help them reduce costs, improve costs on whether it’s manufacturing or, you know, other things that they’re doing. That’s what really excites me is seeing those customers take the raw capability and then apply it to really just to transform how they not just how their business works, but even how they think about the business. Because in many cases, transformation is not just a technical transformation, it’s people and cultural transformation inside these organizations. And that’s pretty cool to see as it unfolds.

Corey: One of the more interesting things that I’ve seen customers misunderstand, on some level, has been a number of S3 releases that focus around, “Oh, this is for your data lake.” And I’ve asked customers about that. “So, what’s your data lake strategy?” “Well, we don’t have one of those.” “You have, like, eight petabytes and climbing in S3? What do you call that?” It’s like, “Oh, yeah, that’s just a bunch of buckets we dump things into. Some are logs of our assets and the rest.” It’s—

Kevin: Right.

Corey: Yeah, it feels like no one thinks of themselves as having anything remotely resembling a structured place for all of the data that accumulates at a company.

Kevin: Mm-hm.

Corey: There is an evolution of people learning that oh, yeah, this is in fact, what it is that we’re doing, and this thing that they’re talking about does apply to us. But it almost feels like a customer communication challenge, just because, I don’t know about you, but with my legacy AWS account, I have dozens of buckets in there that I don’t remember what the heck they’re for. Fortunately, you folks don’t charge by the bucket, so I can smile, nod, remain blissfully ignorant, but it does make me wonder from time to time.

Kevin: Yeah, no, I think that what you hear there is actually pretty consistent with what the reality is for a lot of customers, which is in distributed organizations, I think that’s bound to happen, you have different teams that are working to solve problems, and they are collecting data to analyze, they’re creating result datasets and they’re storing those datasets. And then, of course, priorities can shift, and you know, and there’s not necessarily the day-to-day management around data that we might think would be expected. I feel [we 00:26:56] sort of drew an architecture on a whiteboard. And so, I think that’s the reality we are in. And we will be in, largely forever.

I mean, I think that at a smaller-scale, that’s been happening for years. So, I think that, one, I think that there’s a lot of capability just being in the cloud. At the very least, you can now start to wrap your arms around it, right, where used to be that it wasn’t even possible to understand what all that data was because there’s no way to centrally inventory it well. In AWS with S3, with inventory reports, you can get a list of all your storage and we are going to continue to add capability to help customers get their arms around what they have, first off; understand how it’s being used—that’s where things like Storage Lens really play a big role in understanding exactly what data is being accessed and not. We’re definitely listening to customers carefully around this, and I think when you think about broader data management story, I think that’s a place that we’re spending a lot of time thinking right now about how do we help customers get their arms around it, make sure that they know what’s the categorization of certain data, do I have some PII lurking here that I need to be very mindful of?

And then how do I get to a world where I’m—you know, I won’t say that it’s ever going to look like the perfect whiteboard picture you might draw on the wall. I don’t think that’s really ever achievable, but I think certainly getting to a point where customers have a real solid understanding of what data they have and that the right controls are in place around all that data, yeah, I think that’s directionally where I see us heading.

Corey: As you look around how far the service has come, it feels like, on some level, that there were some, I guess, I don’t want to say missteps, but things that you learned as you went along. Like, back when the service was in beta, for example, there was no per-request charge. To my understanding that was changed, in part because people were trying to use it as a file system, and wow, that suddenly caused a tremendous amount of load on some of the underlying systems. You originally launched with a BitTorrent endpoint as an option so that people could download through peer-to-peer approaches for large datasets and turned out that wasn’t really the way the internet evolved, either. And I’m curious, if you were to have to somehow build this off from scratch, are there any other significant changes you would make in how the service was presented to customers in how people talked about it in the early days? Effectively given a mulligan, what would you do differently?

Kevin: Well, I don’t know, Corey, I mean, just given where it’s grown to in macro terms, you know, I definitely would be worried taking a mulligan, you know, that I [laugh] would change the sort of the overarching trajectory. Certainly, I think there’s a few features here and there where, for whatever reason, it was exciting at the time and really spoke to what customers at the time were thinking, but over time, you know, sort of quickly those needs move to something a little bit different. And, you know, like you said things like the BitTorrent support is one where, at some level, it seems like a great technical architecture for the internet, but certainly not something that we’ve seen dominate in the way things are done. Instead, you know, we’ve largely kind of have a world where there’s a lot of caching layers, but it still ends up being largely client-server kind of connections. So, I don’t think I would do a—I certainly wouldn’t do a mulligan on any of the major functionality, and I think, you know, there’s a few things in the details where obviously, we’ve learned what really works in the end. I think we learned that we wanted bucket names to really strictly conform to rules for DNS encoding. So, that was the change that was made at some point. And we would tweak that, but no major changes, certainly.

Corey: One subject of some debate while we were designing this year’s charity t-shirt—which, incidentally, if you’re listening to this, you can pick up for yourself at snark.cloud/shirt—was the is S3 itself dependent upon S3? Because we know that every other service out there is as well, but it is interesting to come up with an idea of, “Oh, yeah. We’re going to launch a whole new isolated region of S3 without S3 to lean on.” That feels like it’s an almost impossible bootstrapping problem.

Kevin: Well, S3 is not dependent on S3 to come up, and it’s certainly a critical dependency tree that we look at and we track and make sure that we’d like to have an acyclic graph as we look at dependencies.

Corey: That is such a sophisticated way to say what I learned the hard way when I was significantly younger and working in production environments: don’t put the DNS servers needed to boot the hypervisor into VMs that require a working hypervisor. It’s one of those oh, yeah, in hindsight, that makes perfect sense, but you learn it right after that knowledge really would have been useful.

Kevin: Yeah, absolutely. And one of the terms we use for that, as well as is the idea of static stability, or that’s one of the techniques that can really help with isolating a dependency is what we call static stability. We actually have an article about that in the Amazon Builder Library, which there’s actually a bunch of really good articles in there from very experienced operations-focused engineers in AWS. So, static stability is one of those key techniques, but other techniques—I mean, just pure minimization of dependencies is one. And so, we were very, very thoughtful about that, particularly for that core layer.

I mean, you know, when you talk about S3 with 200-plus microservices, or 235-plus microservices, I would say not all of those services are critical for every single request. Certainly, a small subset of those are required for every request, and then other services actually help manage and scale the kind of that inner core of services. And so, we look at dependencies on a service by service basis to really make sure that inner core is as minimized as possible. And then the outer layers can start to take some dependencies once you have that basic functionality up.

Corey: I really want to thank you for being as generous with your time as you have been. If people want to learn more about you and about S3 itself, where should they go—after buying a t-shirt, of course.

Kevin: Well, certainly buy the t-shirt. First, I love the t-shirts and the charity that you work with to do that. Obviously, for S3, it’s aws.amazon.com/s3. And you can actually learn more about me. I have some YouTube videos, so you can search for me on YouTube and kind of get a sense of myself.

Corey: We will put links to that into the show notes, of course. Thank you so much for being so generous with your time. I appreciate it.

Kevin: Absolutely. Yeah. Glad to spend some time. Thanks for the questions, Corey.

Corey: Kevin Miller, vice president and general manager for Amazon S3. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry, ignorant comment talking about how your S3 compatible service is going to blow everyone’s socks off when it fails.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Victor

Victor is an Independent Senior Cloud Infrastructure Architect working mainly on Amazon Web Services (AWS), designing: secure, scalable, reliable, and cost-effective cloud architectures, dealing with large-scale and mission-critical distributed systems. He also has a long experience in Cloud Operations, Security Advisory, Security Hardening (DevSecOps), Modern Applications Design, Micro-services and Serverless, Infrastructure Refactoring, Cost Saving (FinOps).

Links Referenced:

  • Zoph: https://zoph.io/
  • unusd.cloud: https://unusd.cloud
  • Twitter: https://twitter.com/zoph
  • LinkedIn: https://www.linkedin.com/in/grenuv/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is brought to us in part by our friends at Datadog. Datadog's SaaS monitoring and security platform that enables full stack observability for developers, IT operations, security, and business teams in the cloud age. Datadog's platform, along with 500 plus vendor integrations, allows you to correlate metrics, traces, logs, and security signals across your applications, infrastructure, and third party services in a single pane of glass.

Combine these with drag and drop dashboards and machine learning based alerts to help teams troubleshoot and collaborate more effectively, prevent downtime, and enhance performance and reliability. Try Datadog in your environment today with a free 14 day trial and get a complimentary T-shirt when you install the agent.

To learn more, visit datadoghq.com/screaminginthecloud to get. That's www.datadoghq.com/screaminginthecloud

Corey: Managing shards. Maintenance windows. Overprovisioning. ElastiCache bills. I know, I know. It's a spooky season and you're already shaking. It's time for caching to be simpler. Momento Serverless Cache lets you forget the backend to focus on good code and great user experiences. With true autoscaling and a pay-per-use pricing model, it makes caching easy. No matter your cloud provider, get going for free at gomomento.co/screaming That's GO M-O-M-E-N-T-O dot co slash screaming

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. One of the best parts about running a podcast like this and trolling the internet of AWS things is every once in a while, I get to learn something radically different than what I expected. For a long time, there’s been this sort of persona or brand in the AWS space, specifically the security side of it, going by Zoph—that’s Z-O-P-H—and I just assumed it was a collective or a whole bunch of people working on things, and it turns out that nope, it is just one person. And that one person is my guest today. Victor Grenu is an independent AWS architect. Victor, thank you for joining me.

Victor: Hey, Corey, thank you for having me. It’s a pleasure to be here.

Corey: So, I want to start by diving into the thing that first really put you on my radar, though I didn’t realize it was you at the time. You have what can only be described as an army of Twitter bots around the AWS ecosystem. And I don’t even know that I’m necessarily following all of them, but what are these bots and what do they do?

Victor: Yeah. I have a few bots on Twitter that I push some notification, some tweets, when things happen on AWS security space, especially when the AWS managed policies are updated from AWS. And it comes from an initial project from Scott Piper. He was running a Git command on his own laptop to push the history of AWS managed policy. And it told me that I can automate this thing using a deployment pipeline and so on, and to tweet every time a new change is detected from AWS. So, the idea is to monitor every change on these policies.

Corey: It’s kind of wild because I built a number of somewhat similar Twitter bots, only instead of trying to make them into something useful, I’d make them into something more than a little bit horrifying and extraordinarily obnoxious. Like there’s a Cloud Boomer Twitter account that winds up tweeting every time Azure tweets something only it quote-tweets them in all caps and says something insulting. I have an AWS releases bot called AWS Cwoud—so that’s C-W-O-U-D—and that winds up converting it to OwO speak. It’s like, “Yay a new auto-scawowing growp.” That sort of thing is obnoxious and offensive, but it makes me laugh.

Yours, on the other hand, are things that I have notifications turned on for just because when they announce something, it’s generally fairly important. The first one that I discovered was your IAM changes bot. And I found some terrifying things coming out of that from time to time. What’s the data source for that? Because I’m just grabbing other people’s Twitter feeds or RSS feeds; you’re clearly going deeper than that.

Victor: Yeah, the data source is the official AWS managed policy. In fact, I run AWS CLI in the background and I’m doing just a list policy, the list policy command, and with this list I’m doing git of each policy that is returned, so I can enter it in a git repository to get the full history of the time. And I also craft a list of deprecated policy, and I also run, like, a dog-food initiative, the policy analysis, validation analysis from AWS tools to validate the consistency and the accuracy of the own policies. So, there is a policy validation with their own tool. [laugh].

Corey: You would think that wouldn’t turn up anything because their policy validator effectively acts as a linter, so if it throws an error, of course, you wouldn’t wind up pushing that. And yet, somehow the fact that you have bothered to hook that up and have findings from it indicates that that’s not how the real world works.

Victor: Yeah, there is some, let’s say, some false positive because we are running the policy validation with their own linter then own policies, but this is something that is documented from AWS. So, there is an official page where you can find why the linter is not working on each policy and why. There is a an explanation for each findings. I thinking of [unintelligible 00:05:05] managed policy, which is too long, and policy analyzer is crashing because the policy is too long.

Corey: Excellent. It’s odd to me that you have gone down this path because it’s easy enough to look at this and assume that, oh, this must just be something you do for fun or as an aspect of your day job. So, I did a little digging into what your day job is, and this rings very familiar to me: you are an independent AWS consultant, only you’re based out of Paris, whereas I was doing this from San Francisco, due to an escalatingly poor series of life choices on my part. What do you focus on in the AWS consulting world?

Victor: Yeah. I’m running an AWS consulting boutique in Paris and I’m working for a large customer in France. And I’m doing mostly infrastructure stuff, infrastructure design for cloud-native application, and I’m also doing some security audits and [unintelligible 00:06:07] mediation for my customer.

Corey: It seems to me that there’s a definite divide as far as how people find the AWS consulting experience to be. And I’m not trying to cast judgment here, but the stories that I hear tend to fall into one of two categories. One of them is the story that you have, where you’re doing this independently, you’ve been on your own for a while working specifically on this, and then there’s the stories of, “Oh, yeah, I work for a 500 person consultancy and we do everything as long as they’ll pay us money. If they’ve got money, we’ll do it. Why not?”

And it always seems to me—not to be overly judgy—but the independent consultants just seem happier about it because for better or worse, we get to choose what we focus on in a way that I don’t think you do at a larger company.

Victor: Yeah. It’s the same in France or in Europe; there is a lot of consulting firms. But with the pandemic and with the market where we are working, in the cloud, in the cloud-native solution and so on, that there is a lot of demands. And the natural path is to start by working for a consulting firm and then when you are ready, when you have many AWS certification, when you have the experience of the customer, when you have a network of well-known customer, and you gain trust from your customer, I think it’s natural to go by yourself, to be independent and to choose your own project and your own customer.

Corey: I’m curious to get your take on what your perception of being an AWS consultant is when you’re based in Paris versus, in my case, being based in the West Coast of the United States. And I know that’s a bit of a strange question, but even when I travel, for example, over to the East Coast, suddenly, my own newsletter sends out three hours later in the day than I expect it to and that throws me for a loop. The AWS announcements don’t come out at two or three in the afternoon; they come out at dinnertime. And for you, it must be in the middle of the night when a lot of those things wind up dropping. The AWS stuff, not my newsletter. I imagine you’re not excitedly waiting on tenterhooks to see what this week’s issue of Last Week in AWS talks about like I am.

But I’m curious is that even beyond that, how do you experience the market? From what you’re perceiving people in the United States talking about as AWS consultants versus what you see in Paris?

Victor: It’s difficult, but in fact, I don’t have so much information about the independent in the US. I know that there is a lot, but I think it’s more common in Europe. And yeah, it’s an advantage to whoever ten-hour time [unintelligible 00:08:56] from the US because a lot of stuff happen on the Pacific time, on the Seattle timezone, on San Francisco timezone. So, for example, for this podcast, my Monday is over right now, so, so yeah, I have some advantage in time, but yeah.

Corey: This is potentially an odd question for you. But I find an awful lot of the AWS documentation to be challenging, we’ll call it. I don’t always understand exactly what it’s trying to tell me, and it’s not at all clear that the person writing the documentation about a service in some cases has ever used the service. And in everything I just said, there is no language barrier. This documentation was written—theoretically—in English and I, most days, can stumble through a sentence in English and almost no other language. You obviously speak French as a first language. Given that you live in Paris, it seems to be a relatively common affliction. How do you find interacting with AWS in French goes? Or is it just a complete nonstarter, and it all has to happen in English for you?

Victor: No, in fact, the consultants in Europe, I think—in fact, in my part, I’m using my laptop in English, I’m using my phone in English, I’m using the AWS console in English, and so on. So, the documentation for me is a switch on English first because for the other language, there is sometimes some automated translation that is very dangerous sometimes, so we all keep the documentation and the materials in English.

Corey: It’s wild to me just looking at how challenging so much of the stuff is. Having to then work in a second language on top of that, it just seems almost insurmountable to me. It’s good they have automated translation for a lot of this stuff, but that falls down in often hilariously disastrous ways, sometimes. It’s wild to me that even taking most programming languages that folks have ever heard of, even if you program and speak no English, which happens in a large part of the world, you’re still using if statements even if the term ‘if’ doesn’t mean anything to you localized in your language. It really is, in many respects, an English-centric industry.

Victor: Yeah. Completely. Even in French for our large French customer, I’m writing the PowerPoint presentation in English, some emails are in English, even if all the folks in the thread are French. So yeah.

Corey: One other area that I wanted to explore with you a bit is that you are very clearly focused on security as a primary area of interest. Does that manifest in the work that you do as well? Do you find that your consulting engagements tend to have a high degree of focus on security?

Victor: Yeah. In my design, when I’m doing some AWS architecture, my main objective is to design some security architecture and security patterns that apply best practices and least privilege. But often, I’m working for engagement on security audits, for startups, for internal customer, for diverse company, and then doing some accommodation after all. And to run my audit, I’m using some open-source tooling, some custom scripts, and so on. I have a methodology that I’m running for each customer. And the goal is to sometime to prepare some certification, PCI DSS or so on, or maybe to ensure that the best practice are correctly applied on a workload or before go-live or, yeah.

Corey: One of the weird things about this to me is that I’ve said for a long time that cost and security tend to be inextricably linked, as far as being a sort of trailing reactive afterthought for an awful lot of companies. They care about both of those things right after they failed to adequately care about those things. At least in the cloud economic space, it’s only money as opposed to, “Oops, we accidentally lost our customers’ data.” So, I always found that I find myself drifting in a security direction if I don’t stop myself, just based upon a lot of the cost work I do. Conversely, it seems that you have come from the security side and you find yourself drifting in a costing direction.

Your side project is a SaaS offering called unusd.cloud, that’s U-N-U-S-D dot cloud. And when you first mentioned this to me, my immediate reaction was, “Oh, great. Another SaaS platform for costing. Let’s tear this one apart, too.” Except I actually like what you’re building. Tell me about it.

Victor: Yeah, and unusd.cloud is a side project for me and I was working since, let’s say one year. It was a project that I’ve deployed for some of my customer on their local account, and it was very useful. And so, I was thinking that it could be a SaaS project. So, I’ve worked at [unintelligible 00:14:21] so yeah, a few months on shifting the product to assess [unintelligible 00:14:27].

The product aim to detect the worst on AWS account on all AWS region, and it scan all your AWS accounts and all your region, and you try to detect and use the EC2, LDS, Glue [unintelligible 00:14:45], SageMaker, and so on, and attach a EBS and so on. I don’t craft a new dashboard, a new Cost Explorer, and so on. It’s it just cost awareness, it’s just a notification on email or Slack or Microsoft Teams. And you just add your AWS account on the project and you schedule, let’s say, once a day, and it scan, and it send you a cost of wellness, a [unintelligible 00:15:17] detection, and you can act by turning off what is not used.

Corey: What I like about this is it cuts at the number one rule of cloud economics, which is turn that shit off if you’re not using it. You wouldn’t think that I would need to say that except that everyone seems to be missing that, on some level. And it’s easy to do. When you need to spin something up and it’s not there, you’re very highly incentivized to spin that thing up. When you’re not using it, you have to remember that thing exists, otherwise it just sort of sits there forever and doesn’t do anything.

It just costs money and doesn’t generate any value in return for that. What you got right is you’ve also eviscerated my most common complaint about tools that claim to do this, which is you build in either a explicit rule of ignore this resource or ignore resources with the following tags. The benefit there is that you’re not constantly giving me useless advice, like, “Oh, yeah, turn off this idle thing.” It’s, yeah, that’s there for a reason, maybe it’s my dev box, maybe it’s my backup site, maybe it’s the entire DR environment that I’m going to need at little notice. It solves for that problem beautifully. And though a lot of tools out there claim to do stuff like this, most of them really failed to deliver on that promise.

Victor: Yeah, I just want to keep it simple. I don’t want to add an additional console and so on. And you are correct. You can apply a simple tag on your asset, let’s say an EC2 instances, you apply the tag in use and the value of, and then the alerting is disabled for this asset. And the detection is based on the CPU [unintelligible 00:17:01] and the network health metrics, so when the instances is not used in the last seven days, with a low CPU every [unintelligible 00:17:10] and low network out, it comes as a suspect. [laugh].

[midroll 00:17:17]

Corey: One thing that I like about what you’ve done, but also have some reservations about it is that you have not done with so many of these tools do which is, “Oh, just give us all the access in your account. It’ll be fine. You can trust us. Don’t you want to save money?” And yeah, but I also still want to have a company left when all sudden done.

You are very specific on what it is that you’re allowed to access, and it’s great. I would argue, on some level, it’s almost too restrictive. For example, you have the ability to look at EC2, Glue, IAM—just to look at account aliases, great—RDS, Redshift, and SageMaker. And all of these are simply list and describe. There’s no gets in there other than in Cost Explorer, which makes sense. You’re not able to go rummaging through my data and see what’s there. But that also bounds you, on some level, to being able to look only at particular types of resources. Is that accurate or are you using a lot of the CloudWatch stuff and Cost Explorer stuff to see other areas?

Victor: In fact, it’s the least privilege and read-only permission because I don’t want too much question for the security team. So, it’s full read-only permission. And I’ve only added the detection that I’m currently supports. Then if in some weeks, in some months, I’m adding a new detection, let’s say for Snapshot, for example, I will need to update, so I will ask my customer to update their template. There is a mechanisms inside the project to tell them that the template is obsolete, but it’s not a breaking change.

So, the detection will continue, but without the new detection, the new snapshot detection, let’s say. So yeah, it’s least privilege, and all I need is the get-metric-statistics from CloudWatch to detect unused assets. And also checking [unintelligible 00:19:16] Elastic IP or [unintelligible 00:19:19] EBS volume. So, there is no CloudWatching in this detection.

Corey: Also, to be clear, I am not suggesting that what you have done is at all a mistake, even if you bound it to those resources right now. But just because everyone loves to talk about these exciting, amazing, high-level services that AWS has put up there, for example, oh, what about DocumentDB or all these other—you know, Amazon Basics MongoDB; same thing—or all of these other things that they wind up offering, but you take a look at where customers are spending money and where they’re surprised to be spending money, it’s EC2, it’s a bit of RDS, occasionally it’s S3, but that’s a lot harder to detect automatically whether that data is unused. It’s, “You haven’t been using this data very much.” It’s, “Well, you see how the bucket is labeled ‘Archive Backups’ or ‘Regulatory Logs?’” imagine that. What a ridiculous concept.

Yeah. Whereas an idle EC2 instance sort of can wind up being useful on this. I am curious whether you encounter in the wild in your customer base, folks who are having idle-looking EC2 instances, but are in fact, for example, using a whole bunch of RAM, which you can’t tell from the outside without custom CloudWatch agents.

Victor: Yeah, I’m not detecting this behavior for larger usage of RAM, for example, or for maybe there is some custom application that is low in CPU and don’t talk to any other services using the network, but with this detection, with the current state of the detection, I’m covering large majority of waste because what I see from my customer is that there is some teams, some data scientists or data teams who are experimenting a lot with SageMaker with Glue, with Endpoint and so on. And this is very expensive at the end of the day because they don’t turn off the light at the end of the day, on Friday evening. So, what I’m trying to solve here is to notify the team—so on Slack—when they forgot to turn off the most common waste on AWS, so EC2, LTS, Redshift.

Corey: I just now wound up installing it while we’ve been talking on my dedicated shitposting account, and sure enough, it already spat out a single instance it found, which yeah was running an EC2 instance on the East Coast when I was just there, so that I had a DNS server that was a little bit more local. Okay, great. And it’s a T4g.micro, so it’s not exactly a whole lot of money, but it does exactly what it says on the tin. It didn’t wind up nailing the other instances I have in that account that I’m using for a variety of different things, which is good.

And it further didn’t wind up falling into the trap that so many things do, which is the, “Oh, it’s costing you zero and your spend this month is zero because this account is where I dump all of my AWS credit codes.” So, many things say, “Oh, well, it’s not costing you anything, so what’s the problem?” And then that’s how you accidentally lose $100,000 in activate credits because someone left something running way too long. It does a lot of the right things that I would hope and expect it to do, and the fact that you don’t do that is kind of amazing.

Victor: Yeah. It was a need from my customer and an opportunity. It’s a small bet for me because I’m trying to do some small bets, you know, the small bets approach, so the idea is to try a new thing. It’s also an excuse for me to learn something new because building a SaaS is a challenging.

Corey: One thing that I am curious about, in this account, I’m also running the controller for my home WiFi environment. And that’s not huge. It’s T3.small, but it is still something out there that it sits there because I need it to exist. But it’s relatively bored.

If I go back and look over the last week of CloudWatch metrics, for example, it doesn’t look like it’s usually busy. I’m sure there’s some network traffic in and out as it updates itself and whatnot, but the CPU peeks out at a little under 2% used. It didn’t warn on this and it got it right. I’m just curious as to how you did that. What is it looking for to determine whether this instance is unused or not?

Victor: It’s the magic [laugh]. There is some intelligence artif—no, I’m just kidding. It just statistics. And I’m getting two metrics, the superior average from the last seven days and the network out. And I’m getting the average on those metrics and I’m doing some assumption that this EC2, this specific EC2 is not used because of these metrics, this server average.

Corey: Yeah, it is wild to me just that this is working as well as it is. It’s just… like, it does exactly what I would expect it to do. It’s clear that—and this is going to sound weird, but I’m going to say it anyway—that this was built from someone who was looking to answer the question themselves and not from the perspective of, “Well, we need to build a product and we have access to all of this data from the API. How can we slice and dice it and add some value as we go?” I really liked the approach that you’ve taken on this. I don’t say that often or lightly, particularly when it comes to cloud costing stuff, but this is something I’ll be using in some of my own nonsense.

Victor: Thanks. I appreciate it.

Corey: So, I really want to thank you for taking as much time as you have to talk about who you are and what you’re up to. If people want to learn more, where can they find you?

Victor: Mainly on Twitter, my handle is @zoph [laugh]. And, you know, on LinkedIn or on my company website, as zoph.io.

Corey: And we will, of course, put links to that in the [show notes 00:25:23]. Thank you so much for your time today. I really appreciate it.

Victor: Thank you, Corey, for having me. It was a pleasure to chat with you.

Corey: Victor Grenu, independent AWS architect. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an insulting comment that is going to cost you an absolute arm and a leg because invariably, you’re going to forget to turn it off when you’re done.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Mike

Beside his duties as The Duckbill Group’s CEO, Mike is the author of O’Reilly’s Practical Monitoring, and previously wrote the Monitoring Weekly newsletter and hosted the Real World DevOps podcast. He was previously a DevOps Engineer for companies such as Taos Consulting, Peak Hosting, Oak Ridge National Laboratory, and many more. Mike is originally from Knoxville, TN (Go Vols!) and currently resides in Portland, OR.

Links Referenced:

  • @Mike_Julian: https://twitter.com/Mike_Julian
  • mikejulian.com: https://mikejulian.com
  • duckbillgroup.com: https://duckbillgroup.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at AWS AppConfig. Engineers love to solve, and occasionally create, problems. But not when it’s an on-call fire-drill at 4 in the morning. Software problems should drive innovation and collaboration, NOT stress, and sleeplessness, and threats of violence. That’s why so many developers are realizing the value of AWS AppConfig Feature Flags. Feature Flags let developers push code to production, but hide that that feature from customers so that the developers can release their feature when it’s ready. This practice allows for safe, fast, and convenient software development. You can seamlessly incorporate AppConfig Feature Flags into your AWS or cloud environment and ship your Features with excitement, not trepidation and fear. To get started, go to snark.cloud/appconfig. That’s snark.cloud/appconfig.

Corey: Forget everything you know about SSH and try Tailscale. Imagine if you didn't need to manage PKI or rotate SSH keys every time someone leaves. That'd be pretty sweet, wouldn't it? With Tailscale SSH, you can do exactly that. Tailscale gives each server and user device a node key to connect to its VPN, and it uses the same node key to authorize and authenticate SSH.

Basically you're SSHing the same way you manage access to your app. What's the benefit here? Built in key rotation, permissions is code, connectivity between any two devices, reduce latency and there's a lot more, but there's a time limit here. You can also ask users to reauthenticate for that extra bit of security. Sounds expensive?

Nope, I wish it were. Tailscale is completely free for personal use on up to 20 devices. To learn more, visit snark.cloud/tailscale. Again, that's snark.cloud/tailscale

Corey: Welcome to Screaming in the Cloud. I’m Cloud Economist Corey Quinn, and my guest is a returning guest on this show, my business partner and CEO of The Duckbill Group, Mike Julian. Mike, thanks for making the time.

Mike: Lucky number three, I believe?

Corey: Something like that, but numbers are hard. I have databases for that of varying quality and appropriateness for the task, but it works out. Anything’s a database. If you’re brave enough.

Mike: With you inviting me this many times, I’m starting to think you’d like me or something.

Corey: I know, I know. So, let’s talk about something that is going to put that rumor to rest.

Mike: [laugh].

Corey: Clearly, you have made some poor choices in the course of your career, like being my business partner being the obvious one. But what’s really in a dead heat for which is the worst decision is you’ve written a book previously. And now you are starting the process of writing another book because, I don’t know, we don’t keep you busy enough or something. What are you doing?

Mike: Making very bad decisions. When I finished writing Practical Monitoring—O’Reilly, and by the way, you should go buy a copy if interested in monitoring—I finished the book and said, “Wow, that was awful. I’m never doing it again.” And about a month later, I started thinking of new books to write. So, that was 2017, and Corey and I started Duckbill and kind of stopped thinking about writing books because small companies are basically small children. But now I’m going to write a book about consulting.

Corey: Oh, thank God. I thought you’re going to go down the observability path a second time.

Mike: You know, I’m actually dreading the day that O’Reilly asks me to do a second edition because I don’t really want to.

Corey: Yeah. Effectively turn it into an entire story where the only monitoring tool you really need is the AWS bill. That’ll go well.

Mike: [laugh]. Yeah. So yeah, like, basically, I’ve been doing consulting for such a long time, and most of my career is consulting in some form or fashion, and I head up all the consulting at Duckbill. I’ve learned a lot about consulting. And I’ve found that people have a lot of questions about consulting, particularly at the higher-end levels. Once you start getting into advisory sort of stuff, there’s not a lot of great information out there aimed at engineering.

Corey: There’s a bunch of different views on what consulting is. You have independent contractors billing by the hour as staff replacement who call what they do consulting; you have the big consultancies, like Bain or BCG; you’ve got what we do in an advisory sense, and of course, you have a bunch of MBA new grads going to a lot of the big consultancies who are going to see a book on consulting and think that it’s potentially for them. I don’t know that you necessarily have a lot of advice for the new grad type, so who is this for? What is your target customer for this book?

Mike: If you’re interested in joining McKinsey out of college, I don’t have a lot to add; I don’t have a lot to tell you. The reason for that is kind of twofold. One is that shops like McKinsey and Deloitte and Accenture and BCG and Bain, all those, are playing very different games than what most of us think about when we think consulting. Their entire model revolves around running a process. And it’s the same process for every client they work with. But, like, you’re buying them because of their process.

And that process is nothing new or novel. You don’t go to those firms because you want the best advice possible. You go to those firms because it’s the most defensible advice. It’s sort of those things like, “No one gets fired for buying Cisco,” no one got fired for buying IBM, like, that sort of thing, it’s a very defensible choice. But you’re not going to get great results from it.

But because of that, their entire model revolves around throwing dozens, in some cases, hundreds of new grads at a problem and saying, “Run this process. Have fun. Let us know if you need help.” That’s not consulting I have any experience with. It’s honestly not consulting that most of us want to do.

Most of that is staffed by MBAs and accountants. When I think consulting, I think about specialized advice and providing that specialized advice to people. And I wager that most of us think about that in the same way, too. In some cases, it might just be, “I’m going to write code for you as a freelancer,” or I’m just going to tell you like, “Hey, put the nail in here instead of over here because it’s going to be better for you.” Like, paying for advice is good.

But with that, I also have a… one of the first things I say in the beginning of the book, which [laugh] I’ve already started writing because I’m a glutton for punishment, is I don’t think junior people should be consultants. I actually think it’s really bad idea because to be a consultant, you have to have expertise in some area, and junior staff don’t. They haven’t been in their careers long enough to develop that yet. So, they’re just going to flounder. So, my advice is generally aimed at people that have been in their careers for quite some time, generally, people that are 10, 15, 20 years into their career, looking to do something.

Corey: One of the problems that we see when whenever we talk about these things on Twitter is that we get an awful lot of people telling us that we’re wrong, that it can’t be made to work, et cetera, et cetera. But following this model, I’ve been independent for—well, I was independent and then we became The Duckbill Group; add them together because figuring out exactly where that divide happened is always a mental leap for me, but it’s been six years at this point. We’ve definitely proven our ability to not go out of business every month. It’s kind of amazing. Without even an exception case of, “That one time.”

Mike: [laugh]. Yeah, we are living proof that it does work, but you don’t really have to take just our word for it because there are a lot of other firms that exist entirely on an advisory-only, high-expertise model. And it works out really well. We’ve worked with several of them, so it does work; it just isn’t very common inside of tech and particularly inside of engineering.

Corey: So, one of the things that I find is what differentiates an expert from an enthusiastic amateur is, among other things, the number of mistakes that they’ve made. So, I guess a different way of asking this is what qualifies you to write this book, but instead, I’m going to frame it in a very negative way. What have you screwed up on that puts you in a position of, “Ah, I’m going to write a book so that someone else can make better choices.”

Mike: One of my favorite stories to tell—and Corey, I actually think you might not have heard this story before—

Corey: That seems unlikely, but give it a shot.

Mike: Yeah. So, early in my career, I was working for a consulting firm that did ERP implementations. We worked with mainly large, old-school manufacturing firms. So, my job there was to do the engineering side of the implementation. So, a lot of rack-and-stack, a lot of Windows Server configuration, a lot of pulling cables, that sort of thing. So, I thought I was pretty good at this. I quickly learned that I was actually not nearly as good as I thought I was.

Corey: A common affliction among many different people.

Mike: A common affliction. But I did not realize that until this one particular incident. So, me and my boss are both on site at this large manufacturing facility, and the CFO pulls my boss aside and I can hear them talking and, like, she’s pretty upset. She points at me and says, “I never want this asshole in my office ever again.” So, he and I have a long drive back to our office, like an hour and a half.

And we had a long chat about what that meant for me. I was not there for very long after that, as you might imagine, but the thing is, I still have no idea to this day what I did to upset her. I know that she was pissed and he knows that she was pissed. And he never told me exactly what it was, only that’s you take care of your client. And the client believes that I screwed up so massively that she wanted me fired.

Him not wanting to argue—he didn’t; he just kind of went with it—and put me on other clients. But as a result of that, it really got me thinking that I screwed something up so badly to make this person hate me so much and I still have no idea what it was that I did. Which tells me that even at the time, I did not understand what was going on around me. I did not understand how to manage clients well, and to really take care of them. That was probably the first really massive mistake that I’ve made my career—or, like, the first time I came to the realization that there’s a whole lot I don’t know and it’s really costing me.

Corey: From where I sit, there have been a number of things that we have done as we’ve built our consultancy, and I’m curious—you know, let’s get this even more personal—in the past, well, we’ll call it four years that we have been The Duckbill Group—which I think is right—what have we gotten right and what have we gotten wrong? You are the expert; you’re writing a book on this for God’s sake.

Mike: So, what I think we’ve gotten right is one of my core beliefs is never bill hourly. Shout out to Jonathan Stark. He wrote I really good book that is a much better explanation of that than I’ve ever been able to come up with. But I’ve always had the belief that billing hourly is just a bad idea, so we’ve never done that and that’s worked out really well for us. We’ve turned down work because that’s the model they wanted and it’s like, “Sorry, that’s not what we do. You’re going to have to go work for someone else—or hire someone else.”

Other things that I think we’ve gotten right is a focus on staying on the advisory side and not doing any implementation. That’s allowed us to get really good at what we do very quickly because we don’t get mired in long-term implementation detail-level projects. So, that’s been great. Where we went a little wrong, I think—or what we have gotten wrong, lessons that we’ve learned. I had this idea that we could build out a junior and mid-level staff and have them overseen by very senior people.

And, as it turns out, that didn’t work for us, entirely because it didn’t work for me. That was really my failure. I went from being an IC to being the leader of a company in one single step. I’ve never been a manager before Duckbill. So, that particular mistake was really about my lack of abilities in being a good manager and being a good leader.

So, building that out, that did not work for us because it didn’t work for me and I didn’t know how to do it. So, I made way too many mistakes that were kind of amateur-level stuff in terms of management. So, that didn’t work. And the other major mistake that I think we’ve made is not putting enough effort into marketing. So, we get most of our leads by inbound or referral, as is common with boutique consulting firms, but a lot of the income that we get comes through Last Week in AWS, which is really awesome.

But we don’t put a whole lot of effort into content or any marketing stuff related to the thing that we do, like cost management. I think a lot of that is just that we don’t really know how, aside from just creating content and publishing it. We don’t really understand how to market ourselves very well on that side of things. I think that’s a mistake we’ve made.

Corey: It’s an effective strategy against what’s a very complicated problem because unlike most things, if—let’s go back to your old life—if we have an observability problem, we will talk about that very publicly on Twitter and people will come over and get—“Hey, hey, have you tried to buy my company’s product?” Or they’ll offer consulting services, or they’ll point us in the right direction, all of which is sometimes appreciated. Whereas when you have a big AWS bill, you generally don’t talk about it in public, especially if you’re a serious company because that’s going to, uh, I think the phrase is, “Shake investor confidence,” when you’re actually live tweeting slash shitposting about your own AWS bill. And our initial thesis was therefore, since we can’t wind up reaching out to these people when they’re having the pain because there’s no external indication of it, instead what we have to do is be loud enough and notable in this space, where they find us where it shouldn’t take more than them asking one or two of their friends before they get pointed to us. What’s always fun as the stories we hear is, “Okay, so I asked some other people because I wanted a second opinion, and they told us to go to you, too.” Word of mouth is where our customers come from. But how do you bootstrap that? I don’t know. I’m lucky that I got it right the first time.

Mike: Yeah, and as I mentioned a minute ago, that a lot of that really comes through your content, which is not really cost management-related. It’s much more AWS broad. We don’t put out a lot of cost management specific content. And honestly, I think that’s to our detriment. We should and we absolutely can. We just haven’t. I think that’s one of the really big things that we’ve missed on doing.

Corey: There’s an argument that the people who come to us do not spend their entire day thinking about AWS bills. I mean, I can’t imagine what that would be like, but they don’t for whatever reason; they’re trying to do something ridiculous, like you know, run a profitable company. So, getting in front of them when they’re not thinking about the bills means, on some level, that they’re going to reach out to us when the bill strikes. At least that’s been my operating theory.

Mike: Yeah, I mean, this really just comes down to content strategy and broader marketing strategy. Because one of the things you have to think about with marketing is how do you meet a customer at the time that they have the problem that you solve? And what most marketing people talk about here is what’s called the triggering event. Something causes someone to take an action. What is that something? Who is that someone, and what is that action?

And for us, one of the things that we thought early on is that well, the bill comes out the first week of the month, every month, so people are going to opened the bill freak out, and a big influx of leads are going to come our way and that’s going to happen every single month. The reality is that never happened. That turns out was not a triggering event for anyone.

Corey: And early on, when we didn’t have that many leads coming in, it was a statistical aberration that I thought I saw, like, “Oh, out of the three leads this month, two of them showed up in the same day. Clearly, it’s an AWS billing day thing.” No. It turns out that every company’s internal cadence is radically different.

Mike: Right. And I wish I could say that we have found what our triggering events are, but I actually don’t think we have. We know who the people are and we know what they reach out for, but we haven’t really uncovered that triggering event. And it could also be there, there isn’t a one. Or at least, if there is one, it’s not one that we could see externally, which is kind of fine.

Corey: Well, for the half of our consulting that does contract negotiation for large-scale commitments with AWS, it comes up for renewal or the initial discount contract gets offered, those are very clear triggering events but the challenge is that we don’t—

Mike: You can’t see them externally.

Corey: —really see that from the outside. Yeah.

Mike: Right. And this is one of those things where there are triggering events for basically everything and it’s probably going to be pretty consistent once you get down to specific services. Like we provide cost optimization services and contract negotiation services. I’m willing to bet that I can predict exactly what the trigger events for both of those will be pretty well. The problem is, you can never see those externally, which is kind of fine.

Ideally, you would be able to see it externally, but you can’t, so we roll with it, which means our entire strategy has revolved around always being top-of-mind because at the time where it happens, we’re already there. And that’s a much more difficult strategy to employ, but it does work.

Corey: All it takes is time and being really lucky and being really prolific, and, and, and. It’s one of those things where if I were to set out to replicate it, I don’t even know how I’d go about doing it.

Mike: People have been asking me. They say, “I want to create The Duckbill Group for X. What do I do?” And I say, “First step, get yourself a Corey Quinn.” And they’re like, “Well, I can’t do that. There’s only one.” I’m like, “Yep. Sucks to be you.” [laugh].

Corey: Yeah, we called the Jerk Store. They’re running out of him. Yeah, it’s a problem. And I don’t think the world needs a whole lot more of my type of humor, to be honest, because the failure mode that I have experienced brutally and firsthand is not that people don’t find me funny; it’s that it really hurts people’s feelings. I have put significant effort into correcting those mistakes and not repeating them, but it sucks every time I get it wrong.

Mike: Yeah.

Corey: Another question I have for you around the book targeting, are you aiming this at individual independent consultants or are you looking to advise people who are building agencies?

Mike: Explicitly not the latter. My framing around this is that there are a number of people who are doing consulting right now and they’ve kind of fell into it. Often, they’ll leave one job and do a little consulting while they’re waiting on their next thing. And in some cases, that might be a month or two. In some cases, it might go on years, but that whole time, they’re just like, “Oh, yeah, I’m doing consulting in between things.”

But at some point, some of those think, “You know what? I want this to be my thing. I don’t want there to be a next thing. This is my thing. So therefore, how do I get serious about doing consulting? How do I get serious about being a consultant?”

And that’s where I think I can add a lot of value because casually consulting of, like, taking whatever work just kind of falls your way is interesting for a while, but once you get serious about it, and you have to start thinking, well, how do I actually deliver engagements? How do I do that consistently? How do I do it repeatedly? How to do it profitably? How do I price my stuff? How do I package it? How do I attract the leads that I want? How do I work with the customers I want?

And turning that whole thing from a casual, “Yeah, whatever,” into, “This is my business,” is a very different way of thinking. And most people don’t think that way because they didn’t really set out to build a business. They set out to just pass time and earn a little bit of money before they went off to the next job. So, the framing that I have here is that I’m aiming to help people that are wanting to get serious about doing consulting. But they generally have experience doing it already.

Corey: Managing shards. Maintenance windows. Overprovisioning. ElastiCache bills. I know, I know. It's a spooky season and you're already shaking. It's time for caching to be simpler. Momento Serverless Cache lets you forget the backend to focus on good code and great user experiences. With true autoscaling and a pay-per-use pricing model, it makes caching easy. No matter your cloud provider, get going for free at gomomento.co/screaming That's GO M-O-M-E-N-T-O dot co slash screaming

Corey: We went from effectively being the two of us on the consulting delivery side, two scaling up to, I believe, at one point we were six of us, and now we have scaled back down to largely the two of us, aided by very specific external folk, when it makes sense.

Mike: And don’t forget April.

Corey: And of course. I’m talking delivery.

Mike: [laugh].

Corey: There’s a reason I—

Mike: Delivery. Yes.

Corey: —prefaced it that way. There’s a lot of support structure here, let’s not get ourselves, and they make this entire place work. But why did we scale up? And then why did we scale down? Because I don’t believe we’ve ever really talked about that publicly.

Mike: No, not publicly. In fact, most people probably don’t even notice that it happened. We got pretty big for—I mean, not big. So, we hit, I think, six full-time people at one point. And that was quite a bit.

Corey: On the delivery side. Let’s be clear.

Mike: Yeah. No, I think actually with support structure, too. Like, if you add in everyone that we had with the sales and marketing as well, we were like 11 people. And that was a pretty sizable company. But then in July this year, it kind of hit a point where I found that I just wasn’t enjoying my job anymore.

And I looked around and noticed that a lot of other people was kind of feeling the same way, is just things had gotten harder. And the business wasn’t suffering at all, it was just everything felt more difficult. And I finally realized that, for me personally at least, I started Duckbill because I love working with clients, I love doing consulting. And what I have found is that as the company grew larger and larger, I spent most of my time keeping the trains running and taking care of the staff. Which is exactly what I should be doing when we’re that size, like, that is my job at that size, but I didn’t actually enjoy it.

I went into management as, like, this job going from having never done it before. So, I didn’t have anything to compare it to. I didn’t know if I would like it or not. And once I got here, I realized I actually don’t. And I spent a lot of efforts to get better at it and I think I did. I’ve been working with a leadership coach for years now.

But it finally came to a point where I just realized that I wasn’t actually enjoying it anymore. I wasn’t enjoying the job that I had created. And I think that really panned out to you as well. So, we decided, we had kind of an opportune time where one of our team decided that they were also wanting to go back to do independent consulting. I’m like, “Well, this is actually pretty good time. Why don’t we just start scaling things back?” And like, maybe we’ll scale it up again in the future; maybe we won’t. But like, let’s just buy ourselves some breathing room.

Corey: One of the things that I think we didn’t spend quite enough time really asking ourselves was what kind of place do we want to work at. Because we’ve explicitly stated that you and I both view this as the last job either of us is ever going to have, which means that we’re not trying to do the get big quickly to get acquired, or we want to raise a whole bunch of other people’s money to scale massively. Those aren’t things either of us enjoy. And it turns out that handling the challenges of a business with as many people working here as we had wasn’t what either one of us really wanted to do.

Mike: Yeah. You know what—[laugh] it’s funny because a lot of our advisors kept asking the same thing. Like, “So, what kind of company do you want?” And like, we had some pretty good answers for that, in that we didn’t want to build a VC-backed company, we didn’t ever want to be hyperscale. But there’s a wide gulf of things between two-person company and hyperscale and we didn’t really think too much about that.

In fact, being a ten-person company is very different than being a three-person company, and we didn’t really think about that either. We should have really put a lot more thought into that of what does it mean to be a ten-person company, and is that what we want? Or is three, four, or five-person more our style? But then again, I don’t know that we could have predicted that as a concern had we not tried it first.

Corey: Yeah, that was very much something that, for better or worse, we pay advisors for their advice—that’s kind of definitionally how it works—and then we ignored it, on some level, though we thought we were doing something different at the time because there’s some lessons you’ve just got to learn by making the mistake yourself.

Mike: Yeah, we definitely made a few of those. [laugh].

Corey: And it’s been an interesting ride and I’ve got zero problem with how things have shaken out. I like what we do quite a bit. And honestly, the biggest fear I’ve got going forward is that my jackass business partner is about to distract the hell out of himself by writing a book, which is never as easy as even the most pessimistic estimates would be. So, that’s going to be awesome and fun.

Mike: Yeah, just wait until you see the dedication page.

Corey: Yeah, I wasn’t mentioned at all in the last book that you wrote, which I found personally offensive. So, if I’m not mentioned this time, you’re fired.

Mike: Oh, no, you are. It’s just I’m also adding an anti-dedication page, which just has a photo of you.

Corey: Oh, wonderful, wonderful. This is going to be one of those stories of the good consultant and the bad consultant, and I’m going to be the Goofus to your Gallant, aren’t I?

Mike: [laugh]. Yes, yes. You are.

Corey: “Goofus wants to bill by the hour.”

Mike: It’s going to have a page of, like, “Here’s this [unintelligible 00:25:05] book is dedicated to. Here’s my acknowledgments. And [BLEEP] this guy.”

Corey: I love it. I absolutely love it. I think that there is definitely a bright future for telling other people how to consult properly. May just suggest as a subtitle for the book is Consulting—subtitle—You Have Problems and Money. We’ll Take Both.

Mike: [laugh]. Yeah. My working title for this is Practical Consulting, but only because my previous book was Practical Monitoring. Pretty sure O’Reilly would have a fit if I did that. I actually have no idea what I’m going to call the book, still.

Corey: Naming things is super hard. I would suggest asking people at AWS who name services and then doing the exact opposite of whatever they suggest. Like, take their list of recommendations and sort by reverse order and that’ll get you started.

Mike: Yeah. [laugh].

Corey: I want to thank you for giving us an update on what you’re working on and why you have less hair every time I see you because you’re mostly ripping it out due to self-inflicted pain. If people want to follow your adventures, where’s the best place to keep updated on this ridiculous, ridiculous nonsense that I cannot talk you out of?

Mike: Two places. You can follow me on Twitter, @Mike_Julian, or you can sign up for the newsletter on my site at mikejulian.com where I’ll be posting all the updates.

Corey: Excellent. And I look forward to skewering the living hell out of them.

Mike: I look forward to ignoring them.

Corey: Thank you, Mike. It is always a pleasure.

Mike: Thank you, Corey.

Corey: Mike Julian, CEO at The Duckbill Group, and my unwilling best friend. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry, annoying comment in which you tell us exactly what our problem is, and then charge us a fixed fee to fix that problem.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Richard

Richard "RichiH" Hartmann is the Director of Community at Grafana Labs, Prometheus team member, OpenMetrics founder, OpenTelemetry member, CNCF Technical Advisory Group Observability chair, CNCF Technical Oversight Committee member, CNCF Governing Board member, and more. He also leads, organizes, or helps run various conferences from hundreds to 18,000 attendess, including KubeCon, PromCon, FOSDEM, DENOG, DebConf, and Chaos Communication Congress. In the past, he made mainframe databases work, ISP backbones run, kept the largest IRC network on Earth running, and designed and built a datacenter from scratch. Go through his talks, podcasts, interviews, and articles at https://github.com/RichiH/talks or follow him on Twitter at https://twitter.com/TwitchiH for musings on the intersection of technology and society.

Links Referenced:

  • Grafana Labs: https://grafana.com/
  • Twitter: https://twitter.com/TwitchiH
  • Richard Hartmann list of talks: https://github.com/richih/talks

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at AWS AppConfig. Engineers love to solve, and occasionally create, problems. But not when it’s an on-call fire-drill at 4 in the morning. Software problems should drive innovation and collaboration, NOT stress, and sleeplessness, and threats of violence. That’s why so many developers are realizing the value of AWS AppConfig Feature Flags. Feature Flags let developers push code to production, but hide that that feature from customers so that the developers can release their feature when it’s ready. This practice allows for safe, fast, and convenient software development. You can seamlessly incorporate AppConfig Feature Flags into your AWS or cloud environment and ship your Features with excitement, not trepidation and fear. To get started, go to snark.cloud/appconfig. That’s snark.cloud/appconfig.

Corey: This episode is brought to us in part by our friends at Datadog. Datadog's SaaS monitoring and security platform that enables full stack observability for developers, IT operations, security, and business teams in the cloud age. Datadog's platform, along with 500 plus vendor integrations, allows you to correlate metrics, traces, logs, and security signals across your applications, infrastructure, and third party services in a single pane of glass.

Combine these with drag and drop dashboards and machine learning based alerts to help teams troubleshoot and collaborate more effectively, prevent downtime, and enhance performance and reliability. Try Datadog in your environment today with a free 14 day trial and get a complimentary T-shirt when you install the agent.

To learn more, visit datadoghq/screaminginthecloud to get. That's www.datadoghq/screaminginthecloud

Corey: Welcome to Screaming in the Cloud, I’m Corey Quinn. There are an awful lot of people who are incredibly good at understanding the ins and outs and the intricacies of the observability world. But they didn’t have time to come on the show today. Instead, I am talking to my dear friend of two decades now, Richard Hartmann, better known on the internet as RichiH, who is the Director of Community at Grafana Labs, here to suffer—in a somewhat atypical departure for the theme of this show—personal attacks for once. Richie, thank you for joining me.

Richard: And thank you for agreeing on personal attacks.

Corey: Exactly. It was one of your riders. Like, there have to be the personal attacks back and forth or you refuse to appear on the show. You’ve been on before. In fact, the last time we did a recording, I believe you were here in person, which was a long time ago. What have you been up to?

You’re still at Grafana Labs. And in many cases, I would point out that, wow, you’ve been there for many years; that seems to be an atypical thing, which is an American tech industry perspective because every time you and I talk about this, you look at folks who—wow, you were only at that company for five years. What’s wrong with you—you tend to take the longer view and I tend to have the fast twitch, time to go ahead and leave jobs because it’s been more than 20 minutes approach. I see that you’re continuing to live what you preach, though. How’s it been?

Richard: Yeah, so there’s a little bit of Covid brains, I think. When we talked in 2018, I was still working at SpaceNet, building a data center. But the last two-and-a-half years didn’t really happen for many people, myself included. So, I guess [laugh] that includes you.

Corey: No, no you’re right. You’ve only been at Grafana Labs a couple of years. One would think I would check the notes for shooting my mouth off. But then, one wouldn’t know me.

Richard: What notes? Anyway, I’ve been around Prometheus and Grafana Since 2015. But it’s like, real, full-time everything is 2020. There was something in between. Since 2018, I contracted to do vulnerability handling and everything for Grafana Labs because they had something and they didn’t know how to deal with it.

But no, full time is 2020. But as to the space in the [unintelligible 00:02:45] of itself, it’s maybe a little bit German of me, but trying to understand the real world and trying to get an overview of systems and how they actually work, and if they are working correctly and as intended, and if not, how they’re not working as intended, and how to fix this is something which has always been super important to me, in part because I just want to understand the world. And this is a really, really good way to automate understanding of the world. So, it’s basically a work-saving mechanism. And that’s why I’ve been sticking to it for so long, I guess.

Corey: Back in the early days of monitoring systems—so we called it monitoring back then because, you know, are using simple words that lack nuance was sort of de rigueur back then—we wound up effectively having tools. Nagios is the one that springs to mind, and it was terrible in all the ways you would expect a tool written in janky Perl in the early-2000s to be. But it told you what was going on. It tried to do a thing, generally reach a server or query it about things, and when things fell out of certain specs, it screamed its head off, which meant that when you had things like the core switch melting down—thinking of one very particular incident—you didn’t get a Nagios alert; you got 4000 Nagios alerts. But start to finish, you could wrap your head rather fully around what Nagios did and why it did the sometimes strange things that it did.

These days, when you take a look at Prometheus, which we hear a lot about, particularly in the Kubernetes space and Grafana, which is often mentioned in the same breath, it’s never been quite clear to me exactly where those start and stop. It always feels like it’s a component in a larger system to tell you what’s going on rather than a one-stop shop that’s going to, you know, shriek its head off when something breaks in the middle of the night. Is that the right way to think about it? The wrong way to think about it?

Richard: It’s a way to think about it. So personally, I use the terms monitoring and observability pretty much interchangeably. Observability is a relatively well-defined term, even though most people won’t agree. But if you look back into the ’70s into control theory where the term is coming from, it is the measure of how much you’re able to determine the internal state of a system by looking at its inputs and its outputs. Depending on the definition, some people don’t include the inputs, but that is the OG definition as far as I’m aware.

And from this, there flow a lot of things. This question of—or this interpretation of the difference between telling that, yes, something’s broken versus why something’s broken. Or if you can’t ask new questions on the fly, it’s not observability. Like all of those things are fundamentally mapped to this definition of, I need enough data to determine the internal state of whatever system I have just by looking at what is coming in, what is going out. And that is at the core the thing. Now, obviously, it’s become a buzzword, which is oftentimes the fate of successful things. So, it’s become a buzzword, and you end up with cargo culting.

Corey: I would argue periodically, that observability is hipster monitoring. If you call it monitoring, you get yelled at by Charity Majors. Which is tongue and cheek, but she has opinions, made, nonetheless shall I say, frustrating by the fact that she is invariably correct in those opinions, which just somehow makes it so much worse. It would be easy to dismiss things she says if she weren’t always right. And the world is changing, especially as we get into the world of distributed systems.

Is the server that runs the app working or not working loses meaning when we’re talking about distributed systems, when we’re talking about containers running on top of Kubernetes, which turns every outage into a murder mystery. We start having distributed applications composed of microservices, so you have no idea necessarily where an issue is. Okay, is this one microservice having an issue related to the request coming into a completely separate microservice? And it seems that for those types of applications, the answer has been tracing for a long time now, where originally that was something that felt like it was sprung, fully-formed from the forehead of some God known as one of the hyperscalers, but now is available to basically everyone, in theory.

In practice, it seems that instrumenting applications still one of the hardest parts of all of this. I tried hooking up one of my own applications to be observed via OTEL, the open telemetry project, and it turns out that right now, OTEL and AWS Lambda have an intersection point that makes everything extremely difficult to work with. It’s not there yet; it’s not baked yet. And someday, I hope that changes because I would love to interchangeably just throw metrics and traces and logs to all the different observability tools and see which ones work, which ones don’t, but that still feels very far away from current state of the art.

Richard: Before we go there, maybe one thing which I don’t fully agree with. You said that previously, you were told if a service up or down, that’s the thing which you cared about, and I don’t think that’s what people actually cared about. At that time, also, what they fundamentally cared about: is the user-facing service up, or down, or impacted? Is it slow? Does it return errors every X percent for requests, something like this?

Corey: Is the site up? And—you’re right, I was hand-waving over a whole bunch of things. It was, “Okay. First, the web server is returning a page, yes or no? Great. Can I ping the server?” Okay, well, there are ways of server can crash and still leave enough of the TCP/IP stack up or it can respond to pings and do little else.

And then you start adding things to it. But the Nagios thing that I always wanted to add—and had to—was, is the disk full? And that was annoying. And, on some level, like, why should I care in the modern era how much stuff is on the disk because storage is cheap and free and plentiful? The problem is, after the third outage in a month because the disk filled up, you start to not have a good answer for well, why aren’t you monitoring whether the disk is full?

And that was the contributors to taking down the server. When the website broke, there were what felt like a relatively small number of reasonably well-understood contributors to that at small to midsize applications, which is what I’m talking about, the only things that people would let me touch. I wasn’t running hyperscale stuff where you have a fleet of 10,000 web servers and, “Is the server up?” Yeah, in that scenario, no one cares. But when we’re talking about the database server and the two application servers and the four web servers talking to them, you think about it more in terms of pets than you do cattle.

Richard: Yes, absolutely. Yet, I think that was a mistake back then, and I tried to do it differently, as a specific example with the disk. And I’m absolutely agreeing that previous generation tools limit you in how you can actually work with your data. In particular, once you’re with metrics where you can do actual math on the data, it doesn’t matter if the disk is almost full. It matters if that disk is going to be full within X amount of time.

If that disk is 98% full and it sits there at 98% for ten years and provides the service, no one cares. The thing is, will it actually run out in the next two hours, in the next five hours, what have you. Depending on this, is this currently or imminently a customer-impacting or user-impacting then yes, alert on it, raise hell, wake people, make them fix it, as opposed to this thing can be dealt with during business hours on the next workday. And you don’t have to wake anyone up.

Corey: Yeah. The big filer with massive amounts of storage has crossed the 70% line. Okay, now it’s time to start thinking about that, what do you want to do? Maybe it’s time to order another shelf of discs for it, which is going to take some time. That’s a radically different scenario than the 20 gigabyte root volume on your server just started filling up dramatically; the rate of change is such that’ll be full in 20 minutes.

Yeah, one of those is something you want to wake people up for. Generally speaking, you don’t want to wake people up for what is fundamentally a longer-term strategic business problem. That can be sorted out in the light of day versus, “[laugh] we’re not going to be making money in two hours, so if I don’t wake up and fix this now.” That’s the kind of thing you generally want to be woken up for. Well, let’s be honest, you don’t want that to happen at all, but if it does happen, you kind of want to know in advance rather than after the fact.

Richard: You’re literally describing linear predict from Prometheus, which is precisely for this, where I can look back over X amount of time and make a linear prediction because everything else breaks down at scale, blah, blah, blah, to detail. But the thing is, I can draw a line with my pencil by hand on my data and I can predict when is this thing going to it. Which is obviously precisely correct if I have a TLS certificate. It’s a little bit more hand-wavy when it’s a disk. But still, you can look into the future and you say, “What will be happening if current trends for the last X amount of time continue in Y amount of time.” And that’s precisely a thing where you get this more powerful ability of doing math with your data.

Corey: See, when you say it like that, it sounds like it actually is a whole term of art, where you’re focusing on an in-depth field, where salaries are astronomical. Whereas the tools that I had to talk about this stuff back in the day made me sound like, effectively, the sysadmin that I was grunting and pointing: “This is gonna fill up.” And that is how I thought about it. And this is the challenge where it’s easy to think about these things in narrow, defined contexts like that, but at scale, things break.

Like the idea of anomaly detection. Well, okay, great if normally, the CPU and these things are super bored and suddenly it gets really busy, that’s atypical. Maybe we should look into it, assuming that it has a challenge. The problem is, that is a lot harder than it sounds because there are so many factors that factor into it. And as soon as you have something, quote-unquote, “Intelligent,” making decisions on this, it doesn’t take too many false positives before you start ignoring everything it has to say, and missing legitimate things. It’s this weird and obnoxious conflation of both hard technical problems and human psychology.

Richard: And the breaking up of old service boundaries. Of course, when you say microservices, and such, fundamentally, functionally a microservice or nanoservice, picoservice—but the pendulum is already swinging back to larger units of complexity—but it fundamentally does not make any difference if I have a monolith on some mainframe or if I have a bunch of microservices. Yes, I can scale differently, I can scale horizontally a lot more easily, vertically, it’s a little bit harder, blah, blah, blah, but fundamentally, the logic and the complexity, which is being packaged is fundamentally the same. More users, everything, but it is fundamentally the same. What’s happening again, and again, is I’m breaking up those old boundaries, which means the old tools which have assumptions built in about certain aspects of how I can actually get an overview of a system just start breaking down, when my complexity unit or my service or what have I, is usually congruent with a physical piece, of hardware or several services are congruent with that piece of hardware, it absolutely makes sense to think about things in terms of this one physical server. The fact that you have different considerations in cloud, and microservices, and blah, blah, blah, is not inherently that it is more complex.

On the contrary, it is fundamentally the same thing. It scales with users' everything, but it is fundamentally the same thing, but I have different boundaries of where I put interfaces onto my complexity, which basically allow me to hide all of this complexity from the downstream users.

Corey: That’s part of the challenge that I think we’re grappling with across this entire industry from start to finish. Where we originally looked at these things and could reason about it because it’s the computer and I know how those things work. Well, kind of, but okay, sure. But then we start layering levels of complexity on top of layers of complexity on top of layers of complexity, and suddenly, when things stop working the way that we expect, it can be very challenging to unpack and understand why. One of the ways I got into this whole space was understanding, to some degree, of how system calls work, of how the kernel wound up interacting with userspace, about how Linux systems worked from start to finish. And these days, that isn’t particularly necessary most of the time for the care and feeding of applications.

The challenge is when things start breaking, suddenly having that in my back pocket to pull out could be extremely handy. But I don’t think it’s nearly as central as it once was and I don’t know that I would necessarily advise someone new to this space to spend a few years as a systems person, digging into a lot of those aspects. And this is why you need to know what inodes are and how they work. Not really, not anymore. It’s not front and center the way that it once was, in most environments, at least in the world that I live in. Agree? Disagree?

Richard: Agreed. But it’s very much unsurprising. You probably can’t tell me how to precisely grow sugar cane or corn, you can’t tell me how to refine the sugar out of it, but you can absolutely bake a cake. But you will not be able to tell me even a third of—and I’m—for the record, I’m also not able to tell you even a third about the supply chain which just goes from I have a field and some seeds and I need to have a package of refined sugar—you’re absolutely enabled to do any of this. The thing is, you’ve been part of the previous generation of infrastructure where you know how this underlying infrastructure works, so you have more ability to reason about this, but it’s not needed for cloud services nearly as much.

You need different types of skill sets, but that doesn’t mean the old skill set is completely useless, at least not as of right now. It’s much more a case of you need fewer of those people and you need them in different places because those things have become infrastructure. Which is basically the cloud play, where a lot of this is just becoming infrastructure more and more.

Corey: Oh, yeah. Back then I distinctly remember my elders looking down their noses at me because I didn’t know assembly, and how could I possibly consider myself a competent systems admin if I didn’t at least have a working knowledge of assembly? Or at least C, which I, over time, learned enough about to know that I didn’t want to be a C programmer. And you’re right, this is the value of cloud and going back to those days getting a web server up and running just to compile Apache’s httpd took a week and an in-depth knowledge of GCC flags.

And then in time, oh, great. We’re going to have rpm or debs. Great, okay, then in time, you have apt, if you’re in the dev land because I know you are a Debian developer, but over in Red Hat land, we had yum and other tools. And then in time, it became oh, we can just use something like Puppet or Chef to wind up ensuring that thing is installed. And then oh, just docker run. And now it’s a checkbox in a web console for S3.

These things get easier with time and step by step by step we’re standing on the shoulders of giants. Even in the last ten years of my career, I used to have a great challenge question that I would interview people with of, “Do you know what TinyURL is? It takes a short URL and then expands it to a longer one. Great, on the whiteboard, tell me how you would implement that.” And you could go up one side and down the other, and then you could add constraints, multiple data centers, now one goes offline, how do you not lose data? Et cetera, et cetera.

But these days, there are so many ways to do that using cloud services that it almost becomes trivial. It’s okay, multiple data centers, API Gateway, a Lambda, and a global DynamoDB table. Now, what? “Well, now it gets slow. Why is it getting slow?”

“Well, in that scenario, probably because of something underlying the cloud provider.” “And so now, you lose an entire AWS region. How do you handle that?” “Seems to me when that happens, the entire internet’s kind of broken. Do people really need longer URLs?”

And that is a valid answer, in many cases. The question doesn’t really work without a whole bunch of additional constraints that make it sound fake. And that’s not a weakness. That is the fact that computers and cloud services have never been as accessible as they are now. And that’s a win for everyone.

Richard: There’s one aspect of accessibility which is actually decreasing—or two. A, you need to pay for them on an ongoing basis. And B, you need an internet connection which is suitably fast, low latency, what have you. And those are things which actually do make things harder for a variety of reasons. If I look at our back-end systems—as in Grafana—all of them have single binary modes where you literally compile everything into a single binary and you can run it on your laptop because if you’re stuck on a plane, you can’t do any work on it. That kind of is not the best of situations.

And if you have a huge CI/CD pipeline, everything in this cloud and fine and dandy, but your internet breaks. Yeah, so I do agree that it is becoming generally more accessible. I disagree that it is becoming more accessible along all possible axes.

Corey: I would agree. There is a silver lining to that as well, where yes, they are fraught and dangerous and I would preface this with a whole bunch of warnings, but from a cost perspective, all of the cloud providers do have a free tier offering where you can kick the tires on a lot of these things in return for no money. Surprisingly, the best one of those is Oracle Cloud where they have an unlimited free tier, use whatever you want in this subset of services, and you will never be charged a dime. As opposed to the AWS model of free tier where well, okay, it suddenly got very popular or you misconfigured something, and surprise, you now owe us enough money to buy Belize. That doesn’t usually lead to a great customer experience.

But you’re right, you can’t get away from needing an internet connection of at least some level of stability and throughput in order for a lot of these things to work. The stuff you would do locally on a Raspberry Pi, for example, if your budget constrained and want to get something out here, or your laptop. Great, that’s not going to work in the same way as a full-on cloud service will.

Richard: It’s not free unless you have hard guarantees that you’re not going to ever pay anything. It’s fine to send warning, it’s fine to switch the thing off, it’s fine to have you hit random hard and soft quotas. It is not a free service if you can’t guarantee that it is free.

Corey: I agree with you. I think that there needs to be a free offering where, “Well, okay, you want us to suddenly stop serving traffic to the world?” “Yes. When the alternative is you have to start charging me through the nose, yes I want you to stop serving traffic.” That is definitionally what it says on the tin.

And as an independent learner, that is what I want. Conversely, if I’m an enterprise, yeah, I don’t care about money; we’re running our Superbowl ad right now, so whatever you do, don’t stop serving traffic. Charge us all the money. And there’s been a lot of hand wringing about, well, how do we figure out which direction to go in? And it’s, have you considered asking the customer?

So, on a scale of one to bank, how serious is this account going to be [laugh]? Like, what are your big concerns: never charge me or never go down? Because we can build for either of those. Just let’s make sure that all of those expectations are aligned. Because if you guess you’re going to get it wrong and then no one’s going to like you.

Richard: I would argue this. All those services from all cloud providers actually build to address both of those. It’s a deliberate choice not to offer certain aspects.

Corey: Absolutely. When I talk to AWS, like, “Yeah, but there is an eventual consistency challenge in the billing system where it takes”—as anyone who’s looked at the billing system can see—“Multiple days, sometimes for usage data to show up. So, how would we be able to stop things if the usage starts climbing?” To which my relatively direct responses, that sounds like a huge problem. I don’t know how you’d fix that, but I do know that if suddenly you decide, as a matter of policy, to okay, if you’re in the free tier, we will not charge you, or even we will not charge you more than $20 a month.

So, you build yourself some headroom, great. And anything that people are able to spin up, well, you’re just going to have to eat the cost as a provider. I somehow suspect that would get fixed super quickly if that were the constraint. The fact that it isn’t is a conscious choice.

Richard: Absolutely.

Corey: And the reason I’m so passionate about this, about the free space, is not because I want to get a bunch of things for free. I assure you I do not. I mean, I spend my life fixing AWS bills and looking at AWS pricing, and my argument is very rarely, “It’s too expensive.” It’s that the billing dimension is hard to predict or doesn’t align with a customer’s experience or prices a service out of a bunch of use cases where it’ll be great. But very rarely do I just sit here shaking my fist and saying, “It costs too much.”

The problem is when you scare the living crap out of a student with a surprise bill that’s more than their entire college tuition, even if you waive it a week or so later, do you think they’re ever going to be as excited as they once were to go and use cloud services and build things for themselves and see what’s possible? I mean, you and I met on IRC 20 years ago because back in those days, the failure mode and the risk financially was extremely low. It’s yeah, the biggest concern that I had back then when I was doing some of my Linux experimentation is if I typed the wrong thing, I’m going to break my laptop. And yeah, that happened once or twice, and I’ve learned not to make those same kinds of mistakes, or put guardrails in so the blast radius was smaller, or use a remote system instead. Yeah, someone else’s computer that I can destroy. Wonderful. But that was on we live and we learn as we were coming up. There was never an opportunity for us, to my understanding, to wind up accidentally running up an $8 million charge.

Richard: Absolutely. And psychological safety is one of the most important things in what most people do. We are social animals. Without this psychological safety, you’re not going to have long-term, self-sustaining groups. You will not make someone really excited about it. There’s two basic ways to sell: trust or force. Those are the two ones. There’s none else.

Corey: Managing shards. Maintenance windows. Overprovisioning. ElastiCache bills. I know, I know. It's a spooky season and you're already shaking. It's time for caching to be simpler. Momento Serverless Cache lets you forget the backend to focus on good code and great user experiences. With true autoscaling and a pay-per-use pricing model, it makes caching easy. No matter your cloud provider, get going for free at gomomento.co/screaming That's GO M-O-M-E-N-T-O dot co slash screaming

Corey: Yeah. And it also looks ridiculous. I was talking to someone somewhat recently who’s used to spending four bucks a month on their AWS bill for some S3 stuff. Great. Good for them. That’s awesome. Their credentials got compromised. Yes, that is on them to some extent. Okay, great.

But now after six days, they were told that they owed $360,000 to AWS. And I don’t know how, as a cloud company, you can sit there and ask a student to do that. That is not a realistic thing. They are what is known, in the United States at least, in the world of civil litigation as quote-unquote, “Judgment proof,” which means, great, you could wind up finding that someone owes you $20 billion. Most of the time, they don’t have that, so you’re not able to recoup it. Yeah, the judgment feels good, but you’re never going to see it.

That’s the problem with something like that. It’s yeah, I would declare bankruptcy long before, as a student, I wound up paying that kind of money. And I don’t hear any stories about them releasing the collection agency hounds against people in that scenario. But I couldn’t guarantee that. I would never urge someone to ignore that bill and see what happens.

And it’s such an off-putting thing that, from my perspective, is beneath of the company. And let’s be clear, I see this behavior at times on Google Cloud, and I see it on Azure as well. This is not something that is unique to AWS, but they are the 800-pound gorilla in the space, and that’s important. Or as I just to mention right now, like, as I—because I was about to give you crap for this, too, but if I go to grafana.com, it says, and I quote, “Play around with the Grafana Stack. Experience Grafana for yourself, no registration or installation needed.”

Good. I was about to yell at you if it’s, “Oh, just give us your credit card and go ahead and start spinning things up and we won’t charge you. Honest.” Even your free account does not require a credit card; you’re doing it right. That tells me that I’m not going to get a giant surprise bill.

Richard: You have no idea how much thought and work went into our free offering. There was a lot of math involved.

Corey: None of this is easy, I want to be very clear on that. Pricing is one of the hardest things to get right, especially in cloud. And it also, when you get it right, it doesn’t look like it was that hard for you to do. But I fix [sigh] I people’s AWS bills for a living and still, five or six years in, one of the hardest things I still wrestle with is pricing engagements. It’s incredibly nuanced, incredibly challenging, and at least for services in the cloud space where you’re doing usage-based billing, that becomes a problem.

But glancing at your pricing page, you do hit the two things that are incredibly important to me. The first one is use something for free. As an added bonus, you can use it forever. And I can get started with it right now. Great, when I go and look at your pricing page or I want to use your product and it tells me to ‘click here to contact us.’ That tells me it’s an enterprise sales cycle, it’s got to be really expensive, and I’m not solving my problem tonight.

Whereas the other side of it, the enterprise offering needs to be ‘contact us’ and you do that, that speaks to the enterprise procurement people who don’t know how to sign a check that doesn’t have to commas in it, and they want to have custom terms and all the rest, and they’re prepared to pay for that. If you don’t have that, you look to small-time. When it doesn’t matter what price you put on it, you wind up offering your enterprise tier at some large number, it’s yeah, for some companies, that’s a small number. You don’t necessarily want to back yourself in, depending upon what the specific needs are. You’ve gotten that right.

Every common criticism that I have about pricing, you folks have gotten right. And I definitely can pick up on your fingerprints on a lot of this. Because it sounds like a weird thing to say of, “Well, he’s the Director of Community, why would he weigh in on pricing?” It’s, “I don’t think you understand what community is when you ask that question.”

Richard: Yes, I fully agree. It’s super important to get pricing right, or to get many things right. And usually the things which just feel naturally correct are the ones which took the most effort and the most time and everything. And yes, at least from the—like, I was in those conversations or part of them, and the one thing which was always clear is when we say it’s free, it must be free. When we say it is forever free, it must be forever free. No games, no lies, do what you say and say what you do. Basically.

We have things where initially you get certain pro features and you can keep paying and you can keep using them, or after X amount of time they go away. Things like these are built in because that’s what people want. They want to play around with the whole thing and see, hey, is this actually providing me value? Do I want to pay for this feature which is nice or this and that plugin or what have you? And yeah, you’re also absolutely right that once you leave these constraints of basically self-serve cloud, you are talking about bespoke deals, but you’re also talking about okay, let’s sit down, let’s actually understand what your business is: what are your business problems? What are you going to solve today? What are you trying to solve tomorrow?

Let us find a way of actually supporting you and invest into a mutual partnership and not just grab the money and run. We have extremely low churn for, I would say, pretty good reasons. Because this thing about our users, our customers being successful, we do take it extremely seriously.

Corey: It’s one of those areas that I just can’t shake the feeling is underappreciated industry-wide. And the reason I say that this is your fingerprints on it is because if this had been wrong, you have a lot of… we’ll call them idiosyncrasies, where there are certain things you absolutely will not stand for, and misleading people and tricking them into paying money is high on that list. One of the reasons we’re friends. So yeah, but I say I see your fingerprints on this, it’s yeah, if this hadn’t been worked out the way that it is, you would not still be there. One other thing that I wanted to call out about, well, I guess it’s a confluence of pricing and logging in the rest, I look at your free tier, and it offers up to 50 gigabytes of ingest a month.

And it’s easy for me to sit here and compare that to other services, other tools, and other logging stories, and then I have to stop and think for a minute that yeah, discs have gotten way bigger, and internet connections have gotten way faster, and even the logs have gotten way wordier. I still am not sure that most people can really contextualize just how much logging fits into 50 gigs of data. Do you have any, I guess, ballpark examples of what that looks like? Because it’s been long enough since I’ve been playing in these waters that I can’t really contextualize it anymore.

Richard: Lord of the Rings is roughly five megabytes. It’s actually less. So, we’re talking literally 10,000 Lord of the Rings, which you can just shove in us and we’re just storing this for you. Which also tells you that you’re not going to be reading any of this. Or some of it, yes, but not all of it. You need better tooling and you need proper tooling.

And some of this is more modern. Some of this is where we actually pushed the state of the art. But I’m also biased. But I, for myself, do claim that we did push the state of the art here. But at the same time you come back to those absolute fundamentals of how humans deal with data.

If you look back basically as far as we have writing—literally 6000 years ago, is the oldest writing—humans have always dealt with information with the state of the world in very specific ways. A, is it important enough to even write it down, to even persist it in whatever persistence mechanisms I have at my disposal? If yes, write a detailed account or record a detailed account of whatever the thing is. But it turns out, this is expensive and it’s not what you need. So, over time, you optimize towards only taking down key events and only noting key events. Maybe with their interconnections, but fundamentally, the key events.

As your data grows, as you have more stuff, as this still is important to your business and keeps being more important to—or doesn’t even need to be a business; can be social, can be whatever—whatever thing it is, it becomes expensive, again, to retain all of those key events. So, you turn them into numbers and you can do actual math on them. And that’s this path which you’ve seen again, and again, and again, and again, throughout humanity’s history. Literally, as long as we have written records, this has played out again, and again, and again, and again, for every single field which humans actually cared about. At different times, like, power networks are way ahead of this, but fundamentally power networks work on metrics, but for transient load spike, and everything, they have logs built into their power measurement devices, but those are only far in between. Of course, the main thing is just metrics, time-series. And you see this again, and again.

You also were sysadmin in internet-related all switches have been metrics-based or metrics-first for basically forever, for 20, 30 years. But that stands to reason. Of course the internet is running at by roughly 20 years scale-wise in front of the cloud because obviously you need the internet because as you wouldn’t be having a cloud. So, all of those growing pains why metrics are all of a sudden the thing, “Or have been for a few years now,” is basically, of course, people who were writing software, providing their own software services, hit the scaling limitations which you hit for Internet service providers two decades, three decades ago. But fundamentally, you have this complete system. Basically profiles or distributed tracing depending on how you view distributed tracing.

You can also argue that distributed tracing is key events which are linked to each other. Logs sit firmly in the key event thing and then you turn this into numbers and that is metrics. And that’s basically it. You have extremes at the and where you can have valid, depending on your circumstances, engineering trade-offs of where you invest the most, but fundamentally, that is why those always appear again in humanity’s dealing with data, and observability is no different.

Corey: I take a look at last month’s AWS bill. Mine is pretty well optimized. It’s a bit over 500 bucks. And right around 150 of that is various forms of logging and detecting change in the environment. And on the one hand, I sit here, and I think, “Oh, I should optimize that,” because the value of those logs to me is zero.

Except that whenever I have to go in and diagnose something or respond to an incident or have some forensic exploration, they then are worth an awful lot. And I am prepared to pay 150 bucks a month for that because the potential value of having that when the time comes is going to be extraordinarily useful. And it basically just feels like a tax on top of what it is that I’m doing. The same thing happens with application observability where, yeah, when you just want the big substantial stuff, yeah, until you’re trying to diagnose something. But in some cases, yeah, okay, then crank up the verbosity and then look for it.

But if you’re trying to figure it out after an event that isn’t likely or hopefully won’t recur, you’re going to wish that you spent a little bit more on collecting data out of it. You’re always going to be wrong, you’re always going to be unhappy, on some level.

Richard: Ish. You could absolutely be optimizing this. I mean, for $500, it’s probably not worth your time unless you take it as an exercise, but outside of due diligence where you need specific logs tied to—or specific events tied to specific times, I would argue that a lot of the problems with logs is just dealing with it wrong. You have this one extreme of full-text indexing everything, and you have this other extreme of a data lake—which is just a euphemism of never looking at the data again—to keep storage vendors happy. There is an in between.

Again, I’m biased, but like for example, with Loki, you have those same label sets as you have on your metrics with Prometheus, and you have literally the same, which means you only index that part and you only extract on ingestion time. If you don’t have structured logs yet, only put the metadata about whatever you care about extracted and put it into your label set and store this, and that’s the only thing you index. But it goes further than just this. You can also turn those logs into metrics.

And to me this is a path of optimization. Where previously I logged this and that error. Okay, fine, but it’s just a log line telling me it’s HTTP 500. No one cares that this is at this precise time. Log levels are also basically an anti-pattern because they’re just trying to deal with the amount of data which I have, and try and get a handle on this on that level whereas it would be much easier if I just counted every time I have an HTTP 500, I just up my counter by one. And again, and again, and again.

And all of a sudden, I have literally—and I did the math on this—over 99.8% of the data which I have to store just goes away. It’s just magic the way—and we’re only talking about the first time I’m hitting this logline. The second time I’m hitting this logline is functionally free if I turn this into metrics. It becomes cheap enough that one of the mantras which I have, if you need to onboard your developers on modern observability, blah, blah, blah, blah, blah, the whole bells and whistles, usually people have logs, like that’s what they have, unless they were from ISPs or power companies, or so; there they usually start with metrics.

But most users, which I see both with my Grafana and with my Prometheus [unintelligible 00:38:46] tend to start with logs. They have issues with those logs because they’re basically unstructured and useless and you need to first make them useful to some extent. But then you can leverage on this and instead of having a debug statement, just put a counter. Every single time you think, “Hey, maybe I should put a debug statement,” just put a counter instead. In two months time, see if it was worth it or if you delete that line and just remove that counter.

It’s so much cheaper, you can just throw this on and just have it run for a week or a month or whatever timeframe and done. But it goes beyond this because all of a sudden, if I can turn my logs into metrics properly, I can start rewriting my alerts on those metrics. I can actually persist those metrics and can more aggressively throw my logs away. But also, I have this transition made a lot easier where I don’t have this huge lift, where this day in three months is to be cut over and we’re going to release the new version of this and that software and it’s not going to have that, it’s going to have 80% less logs and everything will be great and then you missed the first maintenance window or someone is ill or what have you, and then the next Big Friday is coming so you can’t actually deploy there. I mean Black Friday. But we can also talk about deploying on Fridays.

But the thing is, you have this huge thing, whereas if you have this as a continuous improvement process, I can just look at, this is the log which is coming out. I turn this into a number, I start emitting metrics directly, and I see that those numbers match. And so, I can just start—I build new stuff, I put it into a new data format, I actually emit the new data format directly from my code instrumentation, and only then do I start removing the instrumentation for the logs. And that allows me to, with full confidence, with psychological safety, just move a lot more quickly, deliver much more quickly, and also cut down on my costs more quickly because I’m just using more efficient data types.

Corey: I really want to thank you for spending as much time as you have. If people want to learn more about how you view the world and figure out what other personal attacks they can throw your way, where’s the best place for them to find you?

Richard: Personal attacks, probably Twitter. It’s, like, the go-to place for this kind of thing. For actually tracking, I stopped maintaining my own website. Maybe I’ll do again, but if you go on github.com/richih/talks, you’ll find a reasonably up-to-date list of all the talks, interviews, presentations, panels, what have you, which I did over the last whatever amount of time. [laugh].

Corey: And we will, of course, put links to that in the [show notes 00:41:23]. Thanks again for your time. It’s always appreciated.

Richard: And thank you.

Corey: Richard Hartmann, Director of Community at Grafana Labs. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an insulting comment. And then when someone else comes along with an insulting comment they want to add, we’ll just increment the counter by one.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Michael

Michael is the Director of Threat Research at Sysdig, managing a team of experts tasked with discovering and defending against novel security threats. Michael has more than 20 years of industry experience in many different roles, including incident response, threat intelligence, offensive security research, and software development at companies like Rapid7, ThreatQuotient, and Mantech. Prior to joining Sysdig, Michael worked as a Gartner analyst, advising enterprise clients on security operations topics.

Links Referenced:

  • Sysdig: https://sysdig.com/
  • “2022 Sysdig Cloud-Native Threat Report”: https://sysdig.com/threatreport

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Something interesting about this particular promoted guest episode that is brought to us by our friends at Sysdig is that when they reached out to set this up, one of the first things out of their mouth was, “We don’t want to sell anything,” which is novel. And I said, “Tell me more,” because I was also slightly skeptical. But based upon the conversations that I’ve had, and what I’ve seen, they were being honest. So, my guest today—surprising as though it may be—is Mike Clark, Director of Threat Research at Sysdig. Mike, how are you doing?

Michael: I’m doing great. Thanks for having me. How are you doing?

Corey: Not dead yet. So, we take what we can get sometimes. You folks have just come out with the “2022 Sysdig Cloud-Native Threat Report”, which on one hand, it feels like it’s kind of a wordy title, on the other it actually encompasses everything that it is, and you need every single word of that report. At a very high level, what is that thing?

Michael: Sure. So, this is our first threat report we’ve ever done, and it’s kind of a rite of passage, I think for any security company in the space; you have to have a threat report. And the cloud-native part, Sysdig specializes in cloud and containers, so we really wanted to focus in on those areas when we were making this threat report, which talks about, you know, some of the common threats and attacks we were seeing over the past year, and we just wanted to let people know what they are and how they protect themselves.

Corey: One thing that I’ve found about a variety of threat reports is that they tend to excel at living in the fear, uncertainty, and doubt space. And invariably, they paint a very dire picture of the internet about become cascading down. And then at the end, there’s always a, “But there is hope. Click here to set up a meeting with us.” It’s basically a very thinly- veiled cover around what is fundamentally a fear, uncertainty, and doubt-driven marketing strategy, and then it tries to turn into a sales pitch.

This does absolutely none of that. So, I have to ask, did you set out to intentionally make something that added value in that way and have contributed to the body of knowledge, or is it because it’s your inaugural report; you didn’t realize you were supposed to turn it into a terrible sales pitch.

Michael: We definitely went into that on purpose. There’s a lot of ways to fix things, especially these days with all the different technologies, so we can easily talk about the solutions without going into specific products. And that’s kind of way we went about it. There’s a lot of ways to fix each of the things we mentioned in the report. And hopefully, the person reading it finds a good way to do it.

Corey: I’d like to unpack a fair bit of what’s in the report. And let’s be clear, I don’t intend to read this report into a microphone; that is generally not a great way of conveying information that I have found. But I want to highlight a few things that leapt out to me that I find interesting. Before I do that, I’m curious to know, most people who write reports, especially ones of this quality, are not sitting there cogitating in their office by themselves, and they set pen to paper and emerge four days later with the finished treatise. There’s a team involved, there’s more than one person that weighs in. Who was behind this?

Michael: Yeah, it was a pretty big team effort across several departments. But mostly, it came to the Sysdig threat research team. It’s about ten people right now. It’s grown quite a bit through the past year. And, you know, it’s made up of all sorts of backgrounds and expertise.

So, we have machine learning people, data scientists, data engineers, former pen-testers and red team, a lot of blue team people, people from the NSA, people from other government agencies as well. And we’re also a global research team, so we have people in Europe and North America working on all of this. So, we try to get perspectives on how these threats are viewed by multiple areas, not just Silicon Valley, and express fixes that appeal to them, too.

Corey: Your executive summary on this report starts off with a cloud adversary analysis of TeamTNT. And my initial throwaway joke on that, it was going to be, “Oh, when you start off talking about any entity that isn’t you folks, they must have gotten the platinum sponsorship package.” But then I read the rest of that paragraph and I realized that wait a minute, this is actually interesting and germane to something that I see an awful lot. Specifically, they are—and please correct me if I’m wrong on any of this; you are definitionally the expert whereas I am, obviously the peanut gallery—but you talk about TeamTNT as being a threat actor that focuses on targeting the cloud via cryptojacking, which is a fanciful word for, “Okay, I’ve gotten access to your cloud environment; what am I going to do with it? Mine Bitcoin and other various cryptocurrencies.” Is that generally accurate or have I missed the boat somewhere fierce on that? Which is entirely possible.

Michael: That’s pretty accurate. We also think it just one person, actually, and they are very prolific. So, they were pretty hard to get that platinum support package because they are everywhere. And even though it’s one person, they can do a lot of damage, especially with all the automation people can make now, one person can appear like a dozen.

Corey: There was an old t-shirt that basically encompassed everything that was wrong with the culture of the sysadmin world back in the naughts, that said, “Go away, or I will replace you with a very small shell script.” But, on some level, you can get a surprising amount of work done on computers, just with things like for loops and whatnot. What I found interesting was that you have put numbers and data behind something that I’ve always taken for granted and just implicitly assumed that everyone knew. This is a common failure mode that we all have. We all have blind spots where we assume the things that we spend our time on is easy and the stuff that other people are good at and you’re not good at, those are the hard things.

It has always been intuitively obvious to me as a cloud economist, that when you wind up spending $10,000 in cloud resources to mine cryptocurrency, it does not generate $10,000 of cryptocurrency on the other end. In fact, the line I’ve been using for years is that it’s totally economical to mine Bitcoin in the cloud; the only trick is you have to do it in someone else’s account. And you’ve taken that joke and turned it into data. Something that you found was that in one case, that you were able to attribute $8,100 of cryptocurrency that were generated by stealing $430,000 of cloud resources to do it. And oh, my God, we now have a number and a ratio, and I can talk intelligently and sound four times smarter. So, ignoring anything else in this entire report, congratulations, you have successfully turned this into what is beginning to become a talking point of mine. Value unlocked. Good work. Tell me more.

Michael: Oh, thank you. Cryptomining is kind of like viruses in the old on-prem environment. Normally it just cleaned up and never thought of again; the antivirus software does its thing, life goes on. And I think cryptominers are kind of treated like that. Oh, there’s a miner; let’s rebuild the instance or bring a new container online or something like that.

So, it’s often considered a nuisance rather than a serious threat. It also doesn’t have the, you know, the dangerous ransomware connotation to it. So, a lot of people generally just think of as a nuisance, as I said. So, what we wanted to show was, it’s not really a nuisance and it can cost you a lot of money if you don’t take it seriously. And what we found was for every dollar that they make, it costs you $53. And, you know, as you mentioned, it really puts it into view of what it could cost you by not taking it seriously. And that number can scale very quickly, just like your cloud environment can scale very quickly.

Corey: They say this cloud scales infinitely and that is not true. First, tried it; didn’t work. Secondly, it scales, but there is an inherent limit, which is your budget, on some level. I promise they can add hard drives to S3 faster than you can stuff data into it. I’ve checked.

One thing that I’ve seen recently was—speaking of S3—I had someone reach out in what I will charitably refer to as a blind panic because they were using AWS to do something. Their bill was largely $4 a month in S3 charges. Very reasonable. That carries us surprisingly far. And then they had a credential leak and they had a threat actor spin up all the Lambda functions in all of the regions, and it went from $4 a month to $60,000 a day and it wasn’t caught for six days.

And then AWS as they tend to do, very straight-faced, says, “Yeah, we would like our $360,000, please.” At which point, people start panicking because a lot of the people who experience this are not themselves sophisticated customers; they’re students, they’re learning how this stuff works. And when I’m paying $4 a month for something, it is logical and intuitive for me to think that, well, if I wind up being sloppy with their credentials, they could run that bill up to possibly $25 a month and that wouldn’t be great, so I should keep an eye on it. Yeah, you dropped a whole bunch of zeros off the end of that. Here you go. And as AWS spins up more and more regions and as they spin up more and more services, the ability to exploit this becomes greater and greater. This problem is not getting better, it is only getting worse, by a lot.

Michael: Oh, yeah, absolutely. And I feel really bad for those students who do have that happen to them. I’ve heard on occasion that the cloud providers will forgive some debts, but there’s no guarantee of that happening, from breaches. And you know, the more that breaches happen, the less likely they are going to forgive it because they still to pay for it; someone’s paying for it in the end. And if you don’t improve and fix your environment and it keeps happening, one day, they’re just going to stick you with the bill.

Corey: To my understanding, they’ve always done the right thing when I’ve highlighted something to them. I don’t have intimate visibility into it and of course, they have a threat model themselves of, okay, I’m going to spin up a bunch of stuff, mine cryptocurrency for a month—cry and scream and pretend I got hacked because fraud is very much a thing, there is a financial incentive attached to this—and they mostly seem to get it right. But the danger that I see for the cloud provider is not that they’re going to stop being nice and giving money away, but assume you’re a student who just winds up getting more than your entire college tuition as a surprise bill for this month from a cloud provider. Even assuming at the end of that everything gets wiped and you don’t owe anything. I don’t know about you, but I’ve never used that cloud provider again because I’ve just gotten a firsthand lesson in exactly what those risks are, it’s bad for the brand.

Michael: Yeah, it really does scare people off of that. Now, some cloud providers try to offer more proactive protections against this, try to shut down instances really quick. And you know, you can take advantage of limits and other things, but they don’t make that really easy to do. And setting those up is critical for everybody.

Corey: The one cloud provider that I’ve seen get this right, of all things, has been Oracle Cloud, where they have an always free tier. Until you affirmatively upgrade your account to chargeable, they will not charge you a penny. And I have experimented with this extensively, and they’re right, they will not charge you a penny. They do have warnings plastered on the site, as they should, that until you upgrade your account, do understand that if you exceed a threshold, we will stop serving traffic, we will stop servicing your workload. And yeah, for a student learner, that’s absolutely what I want. For a big enterprise gearing up for a giant Superbowl commercial or whatnot, it’s, “Yeah, don’t care what it costs, just make sure you continue serving traffic. We don’t get a redo on this.” And without understanding exactly which profile of given customer falls into, whenever the cloud provider tries to make an assumption and a default in either direction, they’re wrong.

Michael: Yeah, I’m surprised that Oracle Cloud of all clouds. It’s good to hear that they actually have a free tier. Now, we’ve seen attackers have used free tiers quite a bit. It all depends on how people set it up. And it’s actually a little outside the threat report, but the CI/CD pipelines in DevOps, anywhere there’s free compute, attackers will try to get their miners in because it’s all about scale and not quality.

Corey: Well, that is something I’d be curious to know. Because you talk about focusing specifically on cloud and containers as a company, which puts you in a position to be authoritative on this. That Lambda story that I mentioned about, surprise $60,000 a day in cryptomining, what struck me about that and caught me by surprise was not what I think would catch most people who didn’t swim in this world by surprise of, “You can spend that much?” In my case, what I’m wondering about is, well hang on a minute. I did an article a year or two ago, “17 Ways to Run Containers On AWS” and listed 17 AWS services that you could use to run containers.

And a few months later, I wrote another article called “17 More Ways to Run Containers On AWS.” And people thought I was belaboring the point and making a silly joke, and on some level, of course I was. But I was also highlighting very clearly that every one of those containers running in a service could be mining cryptocurrency. So, if you get access to someone else’s AWS account, when you see those breaches happen, are people using just the one or two services they have things ready to go for, or are they proliferating as many containers as they can through every service that borderline supports it?

Michael: From what we’ve seen, they usually just go after a compute, like EC2 for example, as it's most well understood, it gets the job done, it’s very easy to use, and then get your miner set up. So, if they happen to compromise your credentials versus the other method that cryptominers or cryptojackers do is exploitation, then they’ll try to spread throughout their all their EC2 they can and spin up as much as they can. But the other interesting thing is if they get into your system, maybe via an exploit or some other misconfiguration, they’ll look for the IAM metadata service as soon as they get in, to try to get your IAM credentials and see if they can leverage them to also spin up things through the API. So, they’ll spin up on the thing they compromised and then actively look for other ways to get even more.

Corey: Restricting the permissions that anything has in your cloud environment is important. I mean, from my perspective, if I were to have my account breached, yes, they’re going to cost me a giant pile of money, but I know the magic incantations to say to AWS and worst case, everyone has a pet or something they don’t want to see unfortunate things happen to, so they’ll waive my fee; that’s fine. The bigger concern I’ve got—in seriousness—I think most companies do is the data. It is the access to things in the account. In my case, I have a number of my clients’ AWS bills, given that that is what they pay me to work on.

And I’m not trying to undersell the value of security here, but on the plus side that helps me sleep at night, that’s only money. There are datasets that are far more damaging and valuable about that. The worst sleep I ever had in my career came during a very brief stint I had about 12 years ago when I was the director of TechOps at Grindr, the gay dating site. At that scenario, if that data had been breached, people could very well have died. They live in countries where that winds up not being something that is allowed, or their family now winds up shunning them and whatnot. And that’s the stuff that keeps me up at night. Compared to that, it’s, “Well, you cost us some money and embarrassed a company.” It doesn’t really rank on the same scale to me.

Michael: Yeah. I guess the interesting part is, data requires a lot of work to do something with for a lot of attackers. Like, it may be opportunistic and come across interesting data, but they need to do something with it, there’s a lot more risk once they start trying to sell the data, or like you said, if it turns into something very unfortunate, then there’s a lot more risk from law enforcement coming after them. Whereas with cryptomining, there’s very little risk from being chased down by the authorities. Like you said, people, they rebuild things and ask AWS for credit, or whoever, and move on with their lives. So, that’s one reason I think cryptomining is so popular among threat actors right now. It’s just the low risk compared to other ways of doing things.

Corey: It feels like it’s a nuisance. One thing that I was dreading when I got this copy of the report was that there was going to be what I see so often, which is let’s talk about ransomware in the cloud, where people talk about encrypting data in S3 buckets and sneakily polluting the backups that go into different accounts and how your air -gapping and the rest. And I don’t see that in the wild. I see that in the fear-driven marketing from companies that have a thing that they say will fix that, but in practice, when you hear about ransomware attacks, it’s much more frequently that it is their corporate network, it is on-premises environments, it is servers, perhaps running in AWS, but they’re being treated like servers would be on-prem, and that is what winds up getting encrypted. I just don’t see the attacks that everyone is warning about. But again, I am not primarily in the security space. What do you see in that area?

Michael: You’re absolutely right. Like we don’t see that at all, either. It’s certainly theoretically possible and it may have happened, but there just doesn’t seem to be that appetite to do that. Now, the reasoning? I’m not a hundred percent sure why, but I think it’s easier to make money with cryptomining, even with the crypto markets the way they are. It’s essentially free money, no expenses on your part.

So, maybe they’re not looking because again, that requires more effort to understand especially if it’s not targeted—what data is important. And then it’s not exactly the same method to do the attack. There’s versioning, there’s all this other hoops you have to jump through to do an extortion attack with buckets and things like that.

Corey: Oh, it’s high risk and feels dirty, too. Whereas if you’re just, I guess, on some level, psychologically, if you’re just going to spin up a bunch of coin mining somewhere and then some company finds it and turns it off, whatever. You’re not, as in some cases, shaking down a children’s hospital. Like that’s one of those great, I can’t imagine how you deal with that as a human being, but I guess it takes all types. This doesn’t get us to sort of the second tentpole of the report that you’ve put together, specifically around the idea of supply chain attacks against containers. There have been such a tremendous number of think pieces—thought pieces, whatever they’re called these days—talking about a software bill of materials and supply chain threats. Break it down for me. What are you seeing?

Michael: Sure. So, containers are very fun because, you know, you can define things as code about what gets put on it, and they become so popular that sharing sites have popped up, like Docker Hub and other public registries, where you can easily share your container, it has everything built, set up, so other people can use it. But you know, attackers have kind of taken notice of this, too. Where anything’s easy, an attacker will be. So, we’ve seen a lot of malicious containers be uploaded to these systems.

A lot of times, they’re just hoping for a developer or user to come along and use them because your Docker Hub does have the official designation, so while they can try to pretend to be like Ubuntu, they won’t be the official. But instead, they may try to see theirs and links and things like that to entice people to use theirs instead. And then when they do, it’s already pre-loaded with a miner or, you know, other malware. So, we see quite a bit of these containers in Docker Hub. And they’re disguised as many different popular packages.

They don’t stand up to too much scrutiny, but enough that, you know, a casual looker, even Docker file may not see it. So yeah, we see a lot of—and embedded credentials and other big part that we see in these containers. That could be an organizational issue, like just a leaked credential, but you can put malicious credentials into Docker files, to0, like, say an SSH private key that, you know, if they start this up, the attacker can now just log—SSH in. Or other API keys or other AWS changing commands you can put in there. You can put really anything in there, and wherever you load it, it’s going to run. So, you have to be really careful.

[midroll 00:22:15]

Corey: Years ago, I gave a talk at the conference circuit called, “Terrible Ideas in Git” that purported to teach people how to get worked through hilarious examples of misadventure. And the demos that I did on that were, well, this was fun and great, but it was really annoying resetting them every time I gave the talk, so I stuffed them all into a Docker image and then pushed that up to Docker Hub. Great. It was awesome. I didn’t publicize it and talk about it, but I also just left it as an open repository there because what are you going to do? It’s just a few directories in the route that have very specific contrived scenarios with Git, set up and ready to go.

There’s nothing sensitive there. And the thing is called, “Terrible Ideas.” And I just kept watching the download numbers continue to increment week over week, and I took it down because it’s, I don’t know what people are going to do with that. Like, you see something on there and it says, “Terrible Ideas.” For all I know, some bank is like, “And that’s what we’re running in production now.” So, who knows?

But the idea o—not that there was necessarily anything wrong with that, but the fact that there’s this theoretical possibility someone could use that or put the wrong string in if I give an example, and then wind up running something that is fairly compromisable in a serious environment was just something I didn’t want to be a part of. And you see that again, and again, and again. This idea of what Docker unlocks is amazing, but there’s such a tremendous risk to it. I mean, I’ve never understood 15 years ago, how you’re going to go and spin up a Linux server on top of EC2 and just grab a community AMI and use that. It’s yeah, I used to take provisioning hardware very seriously to make sure that I wasn’t inadvertently using something compromised. Here, it’s like, “Oh, just grab whatever seems plausible from the catalog and go ahead and run that.” But it feels like there’s so much of that, turtles all the way down.

Michael: Yeah. And I mean, even if you’ve looked at the Docker file, with all the dependencies of the things you download, it really gets to be difficult. So, I mean, to protect yourself, it really becomes about, like, you know, you can do the static scanning of it, looking for bad strings in it or bad version numbers for vulnerabilities, but it really comes down to runtime analysis. So, when you start to Docker container, you really need the tools to have visibility to what’s going on in the container. That’s the only real way to know if it’s safe or not in the end because you can’t eyeball it and really see all that, and there could be a binary assortment of layers, too, that’ll get run and things like that.

Corey: Hell is other people’s workflows, as I’m sure everyone’s experienced themselves, but one of mine has always been that if I’m doing something as a proof of concept to build it up on a developer box—and I do keep my developer environments for these sorts of things isolated—I will absolutely go and grab something that is plausible- looking from Docker Hub as I go down that process. But when it comes time to wind up putting it into a production environment, okay, now we’re going to build our own resources. Yeah, I’m sure the Postgres container or whatever it is that you’re using is probably fine, but just so I can sleep at night, I’m going to take the public Docker file they have, and I’m going to go ahead and build that myself. And I feel better about doing that rather than trusting some rando user out there and whatever it is that they’ve put up there. Which on the one hand feels like a somewhat responsible thing to do, but on the other, it feels like I’m only fooling myself because some rando putting things up there is kind of what the entire open-source world is, to a point.

Michael: Yeah, that’s very true. At some point, you have to trust some product or some foundation to have done the right thing. But what’s also true about containers is they’re attacked and use for attacks, but they’re also used to conduct attacks quite a bit. And we saw a lot of that with the Russian-Ukrainian conflict this year. Containers were released that were preloaded with denial-of-service software that automatically collected target lists from, I think, GitHub they were hosted on.

So, all a user to get involved had to do was really just get the container and run it. That’s it. And now they’re participating in this cyberwar kind of activity. And they could also use this to put on a botnet or if they compromise an organization, they could spin up at all these instances with that Docker container on it. And now that company is implicated in that cyber war. So, they can also be used for evil.

Corey: This gets to the third point of your report: “Geopolitical conflict influences attacker behaviors.” Something that happened in the early days of the Russian invasion was that a bunch of open-source maintainers would wind up either disabling what their software did or subverting it into something actively harmful if it detected it was running in the Russian language and/or in a Russian timezone. And I understand the desire to do that, truly I do. I am no Russian apologist. Let’s be clear.

But the counterpoint to that as well is that, well, to make a reference I made earlier, Russia has children’s hospitals, too, and you don’t necessarily know the impact of fallout like that, not to mention that you have completely made it untenable to use anything you’re doing for a regulated industry or anyone else who gets caught in that and discovers that is now in their production environment. It really sets a lot of stuff back. I’ve never been a believer in that particular form of vigilantism, for lack of a better term. I’m not sure that I have a better answer, let’s be clear. I just, I always knew that, on some level, the risk of opening that Pandora’s box were significant.

Michael: Yeah. Even if you’re doing it for the right reasons. It still erodes trust.

Corey: Yeah.

Michael: Especially it erodes trust throughout open-source. Like, not just the one project because you’ll start thinking, “Oh, how many other projects might do this?” And—

Corey: Wait, maybe those dirty hippies did something in our—like, I don’t know, they’ve let those people anywhere near this operating system Linux thing that we use? I don’t think they would have done that. Red Hat seems trustworthy and reliable. And it’s yo, [laugh] someone needs to crack open a history book, on some level. It’s a sticky situation.

I do want to call out something here that it might be easy to get the wrong idea from the summary that we just gave. Very few things wind up raising my hackles quite like companies using tragedy to wind up shilling whatever it is they’re trying to sell. And I’ll admit when I first got this report, and I saw, “Oh, you’re talking about geopolitical conflict, great.” I’m not super proud of this, but I was prepared to read you the riot act, more or less when I inevitably got to that. And I never did. Nothing in this entire report even hints in that direction.

Michael: Was it you never got to it, or, uh—

Corey: Oh, no. I’ve read the whole thing, let’s be clear. You’re not using that to sell things in the way that I was afraid you were. And simultaneously I want to say—I want to just point that out because that is laudable. At the same time, I am deeply and bitterly resentful that that even is laudable. That should be the common state.

Capitalizing on tragedy is just not something that ever leaves any customer feeling good about one of their vendors, and you’ve stayed away from that. I just want to call that out is doing the right thing.

Michael: Thank you. Yeah, it was actually a big topic about how we should broach this. But we have a good data point on right after it started, there was a huge spike in denial-of-service installs. And that we have a bunch of data collection technology, honeypots and other things, and we saw the day after cryptomining started going down and denial-of-service installs started going up. So, it was just interesting how that community changed their behaviors, at least for a time, to participate in whatever you want to call it, the hacktivism.

Over time, though, it kind of has gone back to the norm where maybe they’ve gotten bored or something or, you know, run out of funds, but they’re starting cryptomining again. But these events can cause big changes in the hacktivism community. And like I mentioned, it’s very easy to get involved. We saw over 150,000 downloads of those pre-canned denial-of-service containers, so it’s definitely something that a lot of people participated in.

Corey: It’s a truism that war drives innovation and different ways of thinking about things. It’s a driver of progress, which says something deeply troubling about us. But it’s also clear that it serves as a driver for change, even in this space, where we start to see different applications of things, we see different threat patterns start to emerge. And one thing I do want to call out here that I think often gets overlooked in the larger ecosystem and industry as a whole is, “Well, no one’s going to bother to hack my nonsense. I don’t have anything interesting for them to look at.”

And it’s, on some level, an awful lot of people running tools like this aren’t sophisticated enough themselves to determine that. And combined with your first point in the report as well that, well, you have an AWS account, don’t you? Congratulations. You suddenly have enormous piles of money—from their perspective—sitting there relatively unguarded. Yay. Security has now become everyone’s problem, once again.

Michael: Right. And it’s just easier now. It means, it was always everyone’s problem, but now it’s even easier for attackers to leverage almost everybody. Like before, you had to get something on your PC. You had to download something. Now, your search of GitHub can find API keys, and then that’s it, you know? Things like that will make it game over or your account gets compromised and big bills get run up. And yeah, it’s very easy for all that to happen.

Corey: Ugh. I do want to ask at some point, and I know you asked me not to do it, but I’m going to do it anyway because I have this sneaking suspicion that given that you’ve spent this much time on studying this problem space, that you probably, as a company, have some answers around how to address the pain that lives in these problems. What exactly, at a high level, is it that Sysdig does? Like, how would you describe that in an elevator without sabotaging the elevator for 45 minutes to explain it in depth to someone?

Michael: So, I would describe it as threat detection and response for cloud containers and workloads in general. And all the other kind of acronyms for cloud, like CSPM, CIEM.

Corey: They’re inventing new and exciting acronyms all the time. And I honestly at this point, I want to have almost an acronym challenge of, “Is this a cybersecurity acronym or is it an audio cable? Which is it?” Because it winds up going down that path, super easily. I was at RSA walking the expo floor and I had I think 15 different companies I counted pitching XDR, without a single one bothering to explain what that meant. Okay, I guess it’s just the thing we’ve all decided we need. It feels like security people selling to security people, on some level.

Michael: I was a Gartner analyst.

Corey: Yeah. Oh… that would do it then. Terrific. So, it’s partially your fault, then?

Michael: No. I was going to say, don’t know what it means either.

Corey: Yeah.

Michael: So, I have no idea [laugh]. I couldn’t tell you.

Corey: I’m only half kidding when I say in many cases, from the vendor perspective, it seems like what it means is whatever it is they’re trying to shoehorn the thing that they built into filling. It’s kind of like observability. Observability means what we’ve been doing for ten years already, just repurposed to catch the next hype wave.

Michael: Yeah. The only thing I really understand is: detection and response is a very clear detect things and respond to things. So, that’s a lot of what we do.

Corey: It’s got to beat the default detection mechanism for an awful lot of companies who in years past have found out that they have gotten breached in the headline of The New York Times. Like it’s always fun when that, “Wait, what? What? That’s u—what? How did we not know this was coming?”

It’s when a third party tells you that you’ve been breached, it’s never as positive—not that it’s a positive experience anyway—than discovering yourself internally. And this stuff is complicated, the entire space is fraught, and it always feels like no matter how far you go, you could always go further, but left to its inevitable conclusion, you’ll burn through the entire company budget purely on security without advancing the other things that company does.

Michael: Yeah.

Corey: It’s a balance.

Michael: It’s tough because it’s a lot to know in the security discipline, so you have to balance how much you’re spending and how much your people actually know and can use the things you’ve spent money on.

Corey: I really want to thank you for taking the time to go through the findings of the report for me. I had skimmed it before we spoke, but talking to you about this in significantly more depth, every time I start going to cite something from it, I find myself coming away more impressed. This is now actively going on my calendar to see what the 2023 version looks like. Congratulations, you’ve gotten me hooked. If people want to download a copy of the report for themselves, where should they go to do that?

Michael: They could just go to sysdig.com/threatreport. There’s no email blocking or gating, so you just download it.

Corey: I’m sure someone in your marketing team is twitching at that. Like, why can’t we wind up using this as a lead magnet? But ugh. I look at this and my default is, oh, wow, you definitely understand your target market. Because we all hate that stuff. Every mandatory field you put on those things makes it less likely I’m going to download something here. Click it and have a copy that’s awesome.

Michael: Yep. And thank you for having me. It’s a lot of fun.

Corey: No, thank you for coming. Thanks for taking so much time to go through this, and thanks for keeping it to the high road, which I did not expect to discover because no one ever seems to. Thanks again for your time. I really appreciate it.

Michael: Thanks. Have a great day.

Corey: Mike Clark, Director of Threat Research at Sysdig. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry comment pointing out that I didn’t disclose the biggest security risk at all to your AWS bill, an AWS Solutions Architect who is working on commission.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Steve:

Steve Rice is Principal Product Manager for AWS AppConfig. He is surprisingly passionate about feature flags and continuous configuration. He lives in the Washington DC area with his wife, 3 kids, and 2 incontinent dogs.

Links Referenced:

  • AWS AppConfig: https://go.aws/awsappconfig

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at AWS AppConfig. Engineers love to solve, and occasionally create, problems. But not when it’s an on-call fire-drill at 4 in the morning. Software problems should drive innovation and collaboration, NOT stress, and sleeplessness, and threats of violence. That’s why so many developers are realizing the value of AWS AppConfig Feature Flags. Feature Flags let developers push code to production, but hide that that feature from customers so that the developers can release their feature when it’s ready. This practice allows for safe, fast, and convenient software development. You can seamlessly incorporate AppConfig Feature Flags into your AWS or cloud environment and ship your Features with excitement, not trepidation and fear. To get started, go to snark.cloud/appconfig. That’s snark.cloud/appconfig.

Corey: Forget everything you know about SSH and try Tailscale. Imagine if you didn't need to manage PKI or rotate SSH keys every time someone leaves. That'd be pretty sweet, wouldn't it? With tail scale, ssh, you can do exactly that. Tail scale gives each server and user device a node key to connect to its VPN, and it uses the same node key to authorize and authenticate.

S. Basically you're SSHing the same way you manage access to your app. What's the benefit here? Built in key rotation permissions is code connectivity between any two devices, reduce latency and there's a lot more, but there's a time limit here. You can also ask users to reauthenticate for that extra bit of security. Sounds expensive?

Nope, I wish it were. tail scales. Completely free for personal use on up to 20 devices. To learn more, visit snark.cloud/tailscale. Again, that's snark.cloud/tailscale

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. This is a promoted guest episode. What does that mean? Well, it means that some people don’t just want me to sit here and throw slings and arrows their way, they would prefer to send me a guest specifically, and they do pay for that privilege, which I appreciate. Paying me is absolutely a behavior I wish to endorse.

Today’s victim who has decided to contribute to slash sponsor my ongoing ridiculous nonsense is, of all companies, AWS. And today I’m talking to Steve Rice, who’s the principal product manager on AWS AppConfig. Steve, thank you for joining me.

Steve: Hey, Corey, great to see you. Thanks for having me. Looking forward to a conversation.

Corey: As am I. Now, AppConfig does something super interesting, which I’m not aware of any other service or sub-service doing. You are under the umbrella of AWS Systems Manager, but you’re not going to market with Systems Manager AppConfig. You’re just AWS AppConfig. Why?

Steve: So, AppConfig is part of AWS Systems Manager. Systems Manager has, I think, 17 different features associated with it. Some of them have an individual name that is associated with Systems Manager, some of them don’t. We just happen to be one that doesn’t. AppConfig is a service that’s been around for a while internally before it was launched externally a couple years ago, so I’d say that’s probably the origin of the name and the service. I can tell you more about the origin of the service if you’re curious.

Corey: Oh, I absolutely am. But I just want to take a bit of a detour here and point out that I make fun of the sub-service names in Systems Manager an awful lot, like Systems Manager Session Manager and Systems Manager Change Manager. And part of the reason I do that is not just because it’s funny, but because almost everything I found so far within the Systems Manager umbrella is pretty awesome. It aligns with how I tend to think about the world in a bunch of different ways. I have yet to see anything lurking within the Systems Manager umbrella that has led to a tee-hee-hee bill surprise level that rivals, you know, the GDP of Guam. So, I’m a big fan of the entire suite of services. But yes, how did AppConfig get its name?

Steve: [laugh]. So, AppConfig started about six years ago, now, internally. So, we actually were part of the region services department inside of Amazon, which is in charge of launching new services around the world. We found that a centralized tool for configuration associated with each service launching was really helpful. So, a service might be launching in a new region and have to enable and disable things as it moved along.

And so, the tool was sort of built for that, turning on and off things as the region developed and was ready to launch publicly; then the regions launch publicly. It turned out that our internal customers, which are a lot of AWS services and then some Amazon services as well, started to use us beyond launching new regions, and started to use us for feature flagging. Again, turning on and off capabilities, launching things safely. And so, it became massively popular; we were actually a top 30 service internally in terms of usage. And two years ago, we thought we really should launch this externally and let our customers benefit from some of the goodness that we put in there, and some of—those all come from the mistakes we’ve made internally. And so, it became AppConfig. In terms of the name itself, we specialize in application configuration, so that’s kind of a mouthful, so we just changed it to AppConfig.

Corey: Earlier this year, there was a vulnerability reported around I believe it was AWS Glue, but please don’t quote me on that. And as part of its excellent response that AWS put out, they said that from the time that it was disclosed to them, they had patched the service and rolled it out to every AWS region in which Glue existed in a little under 29 hours, which at scale is absolutely magic fast. That is superhero speed and then some because you generally don’t just throw something over the wall, regardless of how small it is when we’re talking about something at the scale of AWS. I mean, look at who your customers are; mistakes will show. This also got me thinking that when you have Adam, or previously Andy, on stage giving a keynote announcement and then they mention something on stage, like, “Congratulations. It’s now a very complicated service with 14 adjectives in his name because someone’s paid by the syllable. Great.”

Suddenly, the marketing pages are up, the APIs are working, it’s showing up in the console, and it occurs to me only somewhat recently to think about all of the moving parts that go on behind this. That is far faster than even the improved speed of CloudFront distribution updates. There’s very clearly something going on there. So, I’ve got to ask, is that you?

Steve: Yes, a lot of that is us. I can’t take credit for a hundred percent of what you’re talking about, but that’s how we are used. We’re essentially used as a feature-flagging service. And I can talk generically about feature flagging. Feature flagging allows you to push code out to production, but it’s hidden behind a configuration switch: a feature toggle or a feature flag. And that code can be sitting out there, nobody can access it until somebody flips that toggle. Now, the smart way to do it is to flip that toggle on for a small set of users. Maybe it’s just internal users, maybe it’s 1% of your users. And so, the features available, you can—

Corey: It’s your best slash worst customers [laugh] in that 1%, in some cases.

Steve: Yeah, you want to stress test the system with them and you want to be able to look and see what’s going to break before it breaks for everybody. So, you release us to a small cohort, you measure your operations, you measure your application health, you measure your reputational concerns, and then if everything goes well, then you maybe bump it up to 2%, and then 10%, and then 20%. So, feature flags allow you to slowly release features, and you know what you’re releasing by the time it’s at a hundred percent. It’s tempting for teams to want to, like, have everybody access it at the same time; you’ve been working hard on this feature for a long time. But again, that’s kind of an anti-pattern. You want to make sure that on production, it behaves the way you expect it to behave.

Corey: I have to ask what is the fundamental difference between feature flags and/or dynamic configuration. Because to my mind, one of them is a means of achieving the other, but I could also see very easily using the terms interchangeably. Given that in some of our conversations, you have corrected me which, first, how dare you? Secondly, okay, there’s probably a reason here. What is that point of distinction?

Steve: Yeah. Typically for those that are not eat, sleep, and breathing dynamic configuration—which I do—and most people are not obsessed with this kind of thing, feature flags is kind of a shorthand for dynamic configuration. It allows you to turn on and off things without pushing out any new code. So, your application code’s running, it’s pulling its configuration data, say every five seconds, every ten seconds, something like that, and when that configuration data changes, then that app changes its behavior, again, without a code push or without restarting the app.

So, dynamic configuration is maybe a superset of feature flags. Typically, when people think feature flags, they’re thinking of, “Oh, I’m going to release a new feature, so it’s almost like an on-off switch.” But we see customers using feature flags—and we use this internally—for things like throttling limits. Let’s say you want to be able to throttle TPS transactions per second. Or let’s say you want to throttle the number of simultaneous background tasks, and say, you know, I just really don’t want this creeping above 50; bad things can start to happen.

But in a period of stress, you might want to actually bring that number down. Well, you can push out these changes with dynamic configuration—which is, again, any type of configuration, not just an on-off switch—you can push this out and adjust the behavior and see what happens. Again, I’d recommend pushing it out to 1% of your users, and then 10%. But it allows you to have these dials and switches to do that. And, again, generically, that’s dynamic configuration. It’s not as fun to term as feature flags; feature flags is sort of a good mental picture, so I do use them interchangeably, but if you’re really into the whole world of this dynamic configuration, then you probably will care about the difference.

Corey: Which makes a fair bit of sense. It’s the question of what are you talking about high level versus what are you talking about implementation detail-wise.

Steve: Yep. Yep.

Corey: And on some level, I used to get… well, we’ll call it angsty—because I can’t think of a better adjective right now—about how AWS was reluctant to disclose implementation details behind what it did. And in the fullness of time, it’s made a lot more sense to me, specifically through a lens of, you want to be able to have the freedom to change how something works under the hood. And if you’ve made no particular guarantee about the implementation detail, you can do that without potentially worrying about breaking a whole bunch of customer expectations that you’ve inadvertently set. And that makes an awful lot of sense.

The idea of rolling out changes to your infrastructure has evolved over the last decade. Once upon a time you’d have EC2 instances, and great, you want to go ahead and make a change there—or this actually predates EC2 instances. Virtual machines in a data center or heaven forbid, bare metal servers, you’re not going to deploy a whole new server because there’s a new version of the code out, so you separate out your infrastructure from the code that it runs. And that worked out well. And increasingly, we started to see ways of okay, if we want to change the behavior of the application, we’ll just push out new environment variables to that thing and restart the service so it winds up consuming those.

And that’s great. You’ve rolled it out throughout your fleet. With containers, which is sort of the next logical step, well, okay, this stuff gets baked in, we’ll just restart containers with a new version of code because that takes less than a second each and you’re fine. And then Lambda functions, it’s okay, we’ll just change the deployment option and the next invocation will wind up taking the brand new environment variables passed out to it. How do feature flags feature into those, I guess, three evolving methods of running applications in anger, by which I mean, of course, production?

Steve: [laugh]. Good question. And I think you really articulated that well.

Corey: Well, thank you. I should hope so. I’m a storyteller. At least I fancy myself one.

Steve: [laugh]. Yes, you are. Really what you talked about is the evolution of you know, at the beginning, people were—well, first of all, people probably were embedding their variables deep in their code and then they realized, “Oh, I want to change this,” and now you have to find where in my code that is. And so, it became a pattern. Why don’t we separate everything that’s a configuration data into its own file? But it’ll get compiled at build time and sent out all at once.

There was kind of this breakthrough that was, why don’t we actually separate out the deployment of this? We can separate the deployment from code from the deployment of configuration data, and have the code be reading that configuration data on a regular interval, as I already said. So now, as the environments have changed—like you said, containers and Lambda—that ability to make tweaks at microsecond intervals is more important and more powerful. So, there certainly is still value in having things like environment variables that get read at startup. We call that static configuration as opposed to dynamic configuration.

And that’s a very important element in the world of containers that you talked about. Containers are a bit ephemeral, and so they kind of come and go, and you can restart things, or you might spin up new containers that are slightly different config and have them operate in a certain way. And again, Lambda takes that to the next level. I’m really excited where people are going to take feature flags to the next level because already today we have people just fine-tuning to very targeted small subsets, different configuration data, different feature flag data, and allows them to do this like at we’ve never seen before scale of turning this on, seeing how it reacts, seeing how the application behaves, and then being able to roll that out to all of your audience.

Now, you got to be careful, you really don’t want to have completely different configurations out there and have 10 different, or you know, 100 different configurations out there. That makes it really tough to debug. So, you want to think of this as I want to roll this out gradually over time, but eventually, you want to have this sort of state where everything is somewhat consistent.

Corey: That, on some level, speaks to a level of operational maturity that my current deployment adventures generally don’t have. A common reference I make is to my lasttweetinaws.com Twitter threading app. And anyone can visit it, use it however they want.

And it uses a Route 53 latency record to figure out, ah, which is the closest region to you because I’ve deployed it to 20 different regions. Now, if this were a paid service, or I had people using this in large volume and I had to worry about that sort of thing, I would probably approach something that is very close to what you describe. In practice, I pick a devoted region that I deploy something to, and cool, that’s sort of my canary where I get things working the way I would expect. And when that works the way I want it to I then just push it to everything else automatically. Given that I’ve put significant effort into getting deployments down to approximately two minutes to deploy to everything, it feels like that’s a reasonable amount of time to push something out.

Whereas if I were, I don’t know, running a bank, for example, I would probably have an incredibly heavy process around things that make changes to things like payment or whatnot. Because despite the lies, we all like to tell both to ourselves and in public, anything that touches payments does go through waterfall, not agile iterative development because that mistake tends to show up on your customer's credit card bills, and then they’re also angry. I think that there’s a certain point of maturity you need to be at as either an organization or possibly as a software technology stack before something like feature flags even becomes available to you. Would you agree with that, or is this something everyone should use?

Steve: I would agree with that. Definitely, a small team that has communication flowing between the two probably won’t get as much value out of a gradual release process because everybody kind of knows what’s going on inside of the team. Once your team scales, or maybe your audience scales, that’s when it matters more. You really don’t want to have something blow up with your users. You really don’t want to have people getting paged in the middle of the night because of a change that was made. And so, feature flags do help with that.

So typically, the journey we see is people start off in a maybe very small startup. They’re releasing features at a very fast pace. They grow and they start to build their own feature flagging solution—again, at companies I’ve been at previously have done that—and you start using feature flags and you see the power of it. Oh, my gosh, this is great. I can release something when I want without doing a big code push. I can just do a small little change, and if something goes wrong, I can roll it back instantly. That’s really handy.

And so, the basics of feature flagging might be a homegrown solution that you all have built. If you really lean into that and start to use it more, then you probably want to look at a third-party solution because there’s so many features out there that you might want. A lot of them are around safeguards that makes sure that releasing a new feature is safe. You know, again, pushing out a new feature to everybody could be similar to pushing out untested code to production. You don’t want to do that, so you need to have, you know, some checks and balances in your release process of your feature flags, and that’s what a lot of third parties do.

It really depends—to get back to your question about who needs feature flags—it depends on your audience size. You know, if you have enough audience out there to want to do a small rollout to a small set first and then have everybody hit it, that’s great. Also, if you just have, you know, one or two developers, then feature flags are probably something that you’re just kind of, you’re doing yourself, you’re pushing out this thing anyway on your own, but you don’t need it coordinated across your team.

Corey: I think that there’s also a bit of—how to frame this—misunderstanding on someone’s part about where AppConfig starts and where it stops. When it was first announced, feature flags were one of the things that it did. And that was talked about on stage, I believe in re:Invent, but please don’t quote me on that, when it wound up getting announced. And then in the fullness of time, there was another announcement of AppConfig now supports feature flags, which I’m sitting there and I had to go back to my old notes. Like, did I hallucinate this? Which again, would not be the first time I’d imagine such a thing. But no, it was originally how the service was described, but now it’s extra feature flags, almost like someone would, I don’t know, flip on a feature-flag toggle for the service and now it does a different thing. What changed? What was it that was misunderstood about the service initially versus what it became?

Steve: Yeah, I wouldn’t say it was a misunderstanding. I think what happened was we launched it, guessing what our customers were going to use it as. We had done plenty of research on that, and as I mentioned before we had—

Corey: Please tell me someone used it as a database. Or am I the only nutter that does stuff like that?

Steve: We have seen that before. We have seen something like that before.

Corey: Excellent. Excellent, excellent. I approve.

Steve: And so, we had done our due diligence ahead of time about how we thought people were going to use it. We were right about a lot of it. I mentioned before that we have a lot of usage internally, so you know, that was kind of maybe cheating even for us to be able to sort of see how this is going to evolve. What we did announce, I guess it was last November, was an opinionated version of feature flags. So, we had people using us for feature flags, but they were building their own structure, their own JSON, and there was not a dedicated console experience for feature flags.

What we announced last November was an opinionated version that structured the JSON in a way that we think is the right way, and that afforded us the ability to have a smooth console experience. So, if we know what the structure of the JSON is, we can have things like toggles and validations in there that really specifically look at some of the data points. So, that’s really what happened. We’re just making it easier for our customers to use us for feature flags. We still have some customers that are kind of building their own solution, but we’re seeing a lot of them move over to our opinionated version.

Corey: This episode is brought to us in part by our friends at Datadog. Datadog's SaaS monitoring and security platform that enables full stack observability for developers, IT operations, security, and business teams in the cloud age. Datadog's platform, along with 500 plus vendor integrations, allows you to correlate metrics, traces, logs, and security signals across your applications, infrastructure, and third party services in a single pane of glass.

Combine these with drag and drop dashboards and machine learning based alerts to help teams troubleshoot and collaborate more effectively, prevent downtime, and enhance performance and reliability. Try Datadog in your environment today with a free 14 day trial and get a complimentary T-shirt when you install the agent.

To learn more, visit datadoghq/screaminginthecloud to get. That's www.datadoghq/screaminginthecloud

Corey: Part of the problem I have when I look at what it is you folks do, and your use cases, and how you structure it is, it’s similar in some respects to how folks perceive things like FIS, the fault injection service, or chaos engineering, as is commonly known, which is, “We can’t even get the service to stay up on its own for any [unintelligible 00:18:35] period of time. What do you mean, now let’s intentionally degrade it and make it work?” There needs to be a certain level of operational stability or operational maturity. When you’re still building a service before it’s up and running, feature flags seem awfully premature because there’s no one depending on it. You can change configuration however your little heart desires. In most cases. I’m sure at certain points of scale of development teams, you have a communications problem internally, but it’s not aimed at me trying to get something working at 2 a.m. in the middle of the night.

Whereas by the time folks are ready for what you’re doing, they clearly have that level of operational maturity established. So, I have to guess on some level, that your typical adopter of AppConfig feature flags isn’t in fact, someone who is, “Well, we’re ready for feature flags; let’s go,” but rather someone who’s come up with something else as a stopgap as they’ve been iterating forward. Usually something homebuilt. And it might very well be you have the exact same biggest competitor that I do in my consulting work, which is of course, Microsoft Excel as people try to build their own thing that works in their own way.

Steve: Yeah, so definitely a very common customer of ours is somebody that is using a homegrown solution for turning on and off things. And they really feel like I’m using the heck out of these feature flags. I’m using them on a daily or weekly basis. I would like to have some enhancements to how my feature flags work, but I have limited resources and I’m not sure that my resources should be building enhancements to a feature-flagging service, but instead, I’d rather have them focusing on something, you know, directly for our customers, some of the core features of whatever your company does. And so, that’s when people sort of look around externally and say, “Oh, let me see if there’s some other third-party service or something built into AWS like AWS AppConfig that can meet those needs.”

And so absolutely, the workflows get more sophisticated, the ability to move forward faster becomes more important, and do so in a safe way. I used to work at a cybersecurity company and we would kind of joke that the security budget of the company is relatively low until something bad happens, and then it’s, you know, whatever you need to spend on it. It’s not quite the same with feature flags, but you do see when somebody has a problem on production, and they want to be able to turn something off right away or make an adjustment right away, then the ability to do that in a measured way becomes incredibly important. And so, that’s when, again, you’ll see customers starting to feel like they’re outgrowing their homegrown solution and moving to something that’s a third-party solution.

Corey: Honestly, I feel like so many tools exist in this space, where, “Oh, yeah, you should definitely use this tool.” And most people will use that tool. The second time. Because the first time, it’s one of those, “How hard could that be out? I can build something like that in a weekend.” Which is sort of the rallying cry of doomed engineers who are bad at scoping.

And by the time that they figure out why, they have to backtrack significantly. There’s a whole bunch of stuff that I have built that people look at and say, “Wow, that’s a really great design. What inspired you to do that?” And the absolute honest answer to all of it is simply, “Yeah, I worked in roles for the first time I did it the way you would think I would do it and it didn’t go well.” Experience is what you get when you didn’t get what you wanted, and this is one of those areas where it tends to manifest in reasonable ways.

Steve: Absolutely, absolutely.

Corey: So, give me an example here, if you don’t mind, about how feature flags can improve the day-to-day experience of an engineering team or an engineer themselves. Because we’ve been down this path enough, in some cases, to know the failure modes, but for folks who haven’t been there that’s trying to shave a little bit off of their journey of, “I’m going to learn from my own mistakes.” Eh, learn from someone else’s. What are the benefits that accrue and are felt immediately?

Steve: Yeah. So, we kind of have a policy that the very first commit of any new feature ought to be the feature flag. That’s that sort of on-off switch that you want to put there so that you can start to deploy your code and not have a long-lived branch in your source code. But you can have your code there, it reads whether that configuration is on or off. You start with it off.

And so, it really helps just while developing these things about keeping your branches short. And you can push the mainline, as long as the feature flag is off and the feature is hidden to production, which is great. So, that helps with the mess of doing big code merges. The other part is around the launch of a feature.

So, you talked about Andy Jassy being on stage to launch a new feature. Sort of the old way of doing this, Corey, was that you would need to look at your pipelines and see how long it might take for you to push out your code with any sort of code change in it. And let’s say that was an hour-and-a-half process and let’s say your CEO is on stage at eight o’clock on a Friday. And as much as you like to say it, “Oh, I’m never pushing out code on a Friday,” sometimes you have to. The old way—

Corey: Yeah, that week, yes you are, whether you want to or not.

Steve: [laugh]. Exactly, exactly. The old way was this idea that I’m going to time my release, and it takes an hour-and-a-half; I’m going to push it out, and I’ll do my best, but hopefully, when the CEO raises her arm or his arm up and points to a screen that everything’s lit up. Well, let’s say you’re doing that and something goes wrong and you have to start over again. Well, oh, my goodness, we’re 15 minutes behind, can you accelerate things? And then you start to pull away some of these blockers to accelerate your pipeline or you start editing it right in the console of your application, which is generally not a good idea right before a really big launch.

So, the new way is, I’m going to have that code already out there on a Wednesday [laugh] before this big thing on a Friday, but it’s hidden behind this feature flag, I’ve already turned it on and off for internals, and it’s just waiting there. And so, then when the CEO points to the big screen, you can just flip that one small little configuration change—and that can be almost instantaneous—and people can access it. So, that just reduces the amount of stress, reduces the amount of risk in pushing out your code.

Another thing is—we’ve heard this from customers—customers are increasing the number of deploys that they can do per week by a very large percentage because they’re deploying with confidence. They know that I can push out this code and it’s off by default, then I can turn it on whenever I feel like it, and then I can turn it off if something goes wrong. So, if you’re into CI/CD, you can actually just move a lot faster with a number of pushes to production each week, which again, I think really helps engineers on their day-to-day lives. The final thing I’m going to talk about is that let’s say you did push out something, and for whatever reason, that following weekend, something’s going wrong. The old way was oop, you’re going to get a page, I’m going to have to get on my computer and go and debug things and fix things, and then push out a new code change.

And this could be late on a Saturday evening when you’re out with friends. If there’s a feature flag there that can turn it off and if this feature is not critical to the operation of your product, you can actually just go in and flip that feature flag off until the next morning or maybe even Monday morning. So, in theory, you kind of get your free time back when you are implementing feature flags. So, I think those are the big benefits for engineers in using feature flags.

Corey: And the best way to figure out whether someone is speaking from a position of experience or is simply a raving zealot when they’re in a position where they are incentivized to advocate for a particular way of doing things or a particular product, as—let’s be clear—you are in that position, is to ask a form of the following question. Let’s turn it around for a second. In what scenarios would you absolutely not want to use feature flags? What problems arise? When do you take a look at a situation and say, “Oh, yeah, feature flags will make things worse, instead of better. Don’t do it.”

Steve: I’m not sure I wouldn’t necessarily don’t do it—maybe I am that zealot—but you got to do it carefully.

Corey: [laugh].

Steve: You really got to do things carefully because as I said before, flipping on a feature flag for everybody is similar to pushing out untested code to production. So, you want to do that in a measured way. So, you need to make sure that you do a couple of things. One, there should be some way to measure what the system behavior is for a small set of users with that feature flag flipped to on first. And it could be some canaries that you’re using for that.

You can also—there’s other mechanisms you can do that to: set up cohorts and beta testers and those kinds of things. But I would say the gradual rollout and the targeted rollout of a feature flag is critical. You know, again, it sounds easy, “I’ll just turn it on later,” but you ideally don’t want to do that. The second thing you want to do is, if you can, is there some sort of validation that the feature flag is what you expect? So, I was talking about on-off feature flags; there are things, as when I was talking about dynamic configuration, that are things like throttling limits, that you actually want to make sure that you put in some other safeguards that say, “I never want my TPS to go above 1200 and never want to set it below 800,” for whatever reason, for example. Well, you want to have some sort of validation of that data before the feature flag gets pushed out. Inside Amazon, we actually have the policy that every single flag needs to have some sort of validation around it so that we don’t accidentally fat-finger something out before it goes out there. And we have fat-fingered things.

Corey: Typing the wrong thing into a command structure into a tool? “Who would ever do something like that?” He says, remembering times he’s taken production down himself, exactly that way.

Steve: Exactly, exactly, yeah. And we’ve done it at Amazon and AWS, for sure. And so yeah, if you have some sort of structure or process to validate that—because oftentimes, what you’re doing is you’re trying to remediate something in production. Stress levels are high, it is especially easy to fat-finger there. So, that check-and-balance of a validation is important.

And then ideally, you have something to automatically roll back whatever change that you made, very quickly. So AppConfig, for example, hooks up to CloudWatch alarms. If an alarm goes off, we’re actually going to roll back instantly whatever that feature flag was to its previous state so that you don’t even need to really worry about validating against your CloudWatch. It’ll just automatically do that against whatever alarms you have.

Corey: One of the interesting parts about working at Amazon and seeing things in Amazonian scale is that one in a million events happen thousands of times every second for you folks. What lessons have you learned by deploying feature flags at that kind of scale? Because one of my problems and challenges with deploying feature flags myself is that in some cases, we’re talking about three to five users a day for some of these things. That’s not really enough usage to get insights into various cohort analyses or A/B tests.

Steve: Yeah. As I mentioned before, we build these things as features into our product. So, I just talked about the CloudWatch alarms. That wasn’t there originally. Originally, you know, if something went wrong, you would observe a CloudWatch alarm and then you decide what to do, and one of those things might be that I’m going to roll back my configuration.

So, a lot of the mistakes that we made that caused alarms to go off necessitated us building some automatic mechanisms. And you know, a human being can only react so fast, but an automated system there is going to be able to roll things back very, very quickly. So, that came from some specific mistakes that we had made inside of AWS. The validation that I was talking about as well. We have a couple of ways of validating things.

You might want to do a syntactic validation, which really you’re validating—as I was saying—the range between 100 and 1000, but you also might want to have sort of a functional validation, or we call it a semantic validation so that you can make sure that, for example, if you’re switching to a new database, that you’re going to flip over to your new database, you can have a validation there that says, “This database is ready, I can write to this table, it’s truly ready for me to switch.” Instead of just updating some config data, you’re actually going to be validating that the new target is ready for you. So, those are a couple of things that we’ve learned from some of the mistakes we made. And again, not saying we aren’t making mistakes still, but we always look at these things inside of AWS and figure out how we can benefit from them and how our customers, more importantly, can benefit from these mistakes.

Corey: I would say that I agree. I think that you have threaded the needle of not talking smack about your own product, while also presenting it as not the global panacea that everyone should roll out, willy-nilly. That’s a good balance to strike. And frankly, I’d also say it’s probably a good point to park the episode. If people want to learn more about AppConfig, how you view these challenges, or even potentially want to get started using it themselves, what should they do?

Steve: We have an informational page at go.aws/awsappconfig. That will tell you the high-level overview. You can search for our documentation and we have a lot of blog posts to help you get started there.

Corey: And links to that will, of course, go into the [show notes 00:31:21]. Thank you so much for suffering my slings, arrows, and other assorted nonsense on this. I really appreciate your taking the time.

Steve: Corey thank you for the time. It’s always a pleasure to talk to you. Really appreciate your insights.

Corey: You’re too kind. Steve Rice, principal product manager for AWS AppConfig. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry comment. But before you do, just try clearing your cookies and downloading the episode again. You might be in the 3% cohort for an A/B test, and you [want to 00:32:01] listen to the good one instead.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Nipun

Nipun Agarwal is a Senior Vice President, MySQL HeatWave Development, Oracle. His interests include distributed data processing, machine learning, cloud technologies and security. Nipun was part of the Oracle Database team where he introduced a number of new features. He has been awarded over 170 patents.

Links Referenced:

  • Oracle: https://oracle.com
  • MySQL HeatWave info: https://www.oracle.com/mysql/
  • MySQL Service on AWS and OCI login (Oracle account required): https://cloud.mysql.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is brought to us in part by our friends at Datadog. Datadog's SaaS monitoring and security platform that enables full stack observability for developers, IT operations, security, and business teams in the cloud age. Datadog's platform, along with 500 plus vendor integrations, allows you to correlate metrics, traces, logs, and security signals across your applications, infrastructure, and third party services in a single pane of glass.

Combine these with drag and drop dashboards and machine learning based alerts to help teams troubleshoot and collaborate more effectively, prevent downtime, and enhance performance and reliability. Try Datadog in your environment today with a free 14 day trial and get a complimentary T-shirt when you install the agent.

To learn more, visit datadoghq.com/screaminginthecloud to get. That's www.datadoghq.com/screaminginthecloud

Corey: This episode is sponsored in part by our friends at Sysdig. Sysdig secures your cloud from source to run. They believe, as do I, that DevOps and security are inextricably linked. If you wanna learn more about how they view this, check out their blog, it's definitely worth the read. To learn more about how they are absolutely getting it right from where I sit, visit Sysdig.com and tell them that I sent you. That's S Y S D I G.com. And my thanks to them for their continued support of this ridiculous nonsense.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. This promoted episode is sponsored by our friends at Oracle, and back for a borderline historic third round going out and telling stories about these things, we have Nipun Agarwal, who is, as opposed to his first appearance on the show, has been promoted to senior vice president of MySQL HeatWave. Nipun, thank you for coming back. Most people are not enamored enough with me to subject themselves to my slings and arrows a second time, let alone a third. So first, thanks. And are you okay, over there?

Nipun: Thank you, Corey. Yeah, very happy to be back.

Corey: [laugh]. So, since the last time we’ve spoken, there have been some interesting developments that have happened. It was pre-announced by Larry Ellison on a keynote stage or an earnings call, I don’t recall the exact format, that HeatWave was going to be coming to AWS. Now, you’ve conducted a formal announcement, this usual media press blitz, et cetera, talking about it with an eye toward general availability later this year, if I’m not mistaken, and things seem to be—if you’ll forgive the term—heating up a bit.

Nipun: That is correct. So, as you know, we have had MySQL HeatWave on OCI for just about two years now. Very good reception, a lot of people who are using MySQL HeatWave, are migrating from other clouds, specifically from AWS, and now we have announced availability of MySQL HeatWave on AWS.

Corey: So, for those who have not done the requisite homework of listening to the entire back catalog of nearly 400 episodes of this show, what exactly is MySQL HeatWave, just so we make sure that we set the stage for what we’re going to be talking about? Because I sort of get the sense that without a baseline working knowledge of what that is, none of the rest of this is going to make a whole lot of sense.

Nipun: MySQL HeatWave is a managed MySQL service provided by Oracle. But it is different from other MySQL-based services in the sense that we have significantly enhanced the service such that it can very efficiently process transactions, analytics, and in-database machine learning. So, what customers get with the service, with MySQL HeatWave, is a single MySQL database which can process OLTP, transaction processing, real-time analytics, and machine learning. And they can do this without having to move the data out of MySQL into some other specialized database services who are running analytics or machine learning. And all existing tools and applications which work with MySQL work as is because this is something that enhances the server. In addition to that, it provides very good performance and very good price performance compared to other similar services out there.

Corey: The idea historically that some folks were pushing around the idea of multi-cloud was that you would have workloads that—oh, they live in one cloud, but the database was going to be all the way across the other side of the internet, living in a different provider. And in practice, what we generally tend to see is that where the data lives is where the compute winds up living. By and large, it’s easier to bring the compute resources to the data than it is to move the data to the compute, just because data egress in most of the cloud providers—notably exempting yours—is astronomically expensive. You are, if I recall correctly, less than 10% of AWS’s data egress charge on just retail pricing alone, which is wild to me. So first, thank you for keeping that up and not raising prices because I would have felt rather annoyed if I’d been saying such good things. And it was, haha, it was a bait and switch. It was not. I’m still a big fan. So, thank you for that, first and foremost.

Nipun: Certainly. And what you described is absolutely correct that while we have a lot of customers migrating from AWS to use MySQL HeatWave and OCI, a class of customers are unable to, and the number one reason they’re unable to is that AWS charges these customers all very high egress fees to move the data out of AWS into OCI for them to benefit from MySQL HeatWave. And this has definitely been one of the key incentives for us, the key motivation for us, to offer MySQL HeatWave on AWS so that customers don’t need to pay this exorbitant data egress fees.

Corey: I think it’s fair to disclose that I periodically advise a variety of different cloud companies from a perspective of voice-of-the-customer feedback, which essentially distills down to me asking really annoying slash obnoxious questions that I, as a customer, legitimately want to know, but people always frown at me when I asked that in vendor pitches. For some reason, when I’m doing this on an advisory basis, people instead nod thoughtfully and take notes, so that at least feels better from my perspective. Oracle Cloud has been one of those, and I’ve been kicking the tires on the AWS offering that you folks have built out for a bit of time now. I have to say, it is legitimate. I was able to run a significant series of tests on this, and what I found going through that process was interesting on a bunch of different levels.

I’m wondering if it’s okay with you, if we go through a few of them, just things that jumped out to me as we went through a series of conversations around, “So, we’re going to run a service on AWS.” And my initial answer was, “Is this Oracle? Are you sure?” And here we are today; we are talking about it and press releases.

Nipun: Yes, certainly fine with me. Please go ahead.

Corey: So, I think one of the first questions I had when you said, “We’re going to run a database service on AWS itself,” was, if I’m true to type, is going to be fairly sarcastic, which is, “Oh, thank God. Finally, a way to run a MySQL database on AWS. There’s never been one of those before.” Unless you count EC2 or Aurora or Redshift depending upon how you squint at it, or a variety of other increasingly strange things. It feels like that has been a largely saturated market in many respects.

I generally don’t tend to advise on things that I find patently ridiculous, and your answer was great, but I don’t want to put words in your mouth. What was it that you saw that made you say, “Ah, we’re going to put a different database offering on AWS, and no, it’s not a terrible decision.”

Nipun: Got it. Okay, so if you look at it, the value proposition which MySQL HeatWave offers is that customers of MySQL or customers have MySQL compatible databases, whether Aurora, or whether it’s RDS MySQL, right, or even, like, you know, customers of Redshift, they have been migrating to MySQL HeatWave on OCI. Like, for the reasons I said: it’s a single database, customers don’t need to have multiple databases for managing different kinds of workloads, it’s much faster, it’s a lot less expensive, right? So, there are a lot of value propositions. So, what we found is that if you were to offer MySQL HeatWave on AWS, it will significantly ease the migration of other customers who might be otherwise thinking that it will be difficult for them to migrate, perhaps because of the high egress cost of AWS, or because of the high latency some of the applications in the AWS incur when the database is running somewhere else.

Or, if they really have an ecosystem of applications already running on AWS and they just want to replace the database, it’ll be much easier for them if MySQL HeatWave was offered on AWS. Those are the reasons why we feel it’s a compelling proposition, that if existing customers of AWS are willing to migrate the cloud from AWS to OCI and use MySQL HeatWave, there is clearly a value proposition we are offering. And if we can now offer the same service in AWS, it will hopefully increase the number of customers who can benefit from MySQL HeatWave.

Corey: One of the next questions I had going in was, “Okay, so what actually is this under the hood?” Is this you effectively just stuffing some software into a machine image or an AMI—or however they want to mispronounce that word over an AWS-land—and then just making it available to your account and running that, where’s the magic or mystery behind this? Like, it feels like the next more modern cloud approach is to stuff the whole thing into a Docker container. But that’s not what you wound up doing.

Nipun: Correct. So, HeatWave has been designed and architected for scale-out processing, and it’s been optimized for the cloud. So, when we decided to offer MySQL HeatWave on AWS, we have actually gone ahead and optimize our server for the AWS architecture. So, the processor we are running on, right, we have optimized our software for that instance types in AWS, right? So, the data plane has been optimized for AWS architecture.

The second thing is we have a brand new control plane layer, right? So, it’s not the case that we’re just taking what we had in OCI and running it on AWS. We have optimized the data plane for AWS, we have a native control plane, which is running on AWS, which is using the respective services on AWS. And third, we have a brand new console which we are offering, which is a very interactive console where customers can run queries from the console. They can do data management from the console, they’re able to use Autopilot from the console, and we have performance monitoring from the console, right? So, data plane, control plane, console. They’re all running natively in AWS. And this provides for a very seamless integration or seamless experience for the AWS customers.

Corey: I think it’s also a reality, however much we may want to pretend otherwise, that if there is an opportunity to run something in a different cloud provider that is better than where you’re currently running it now, by and large, customers aren’t going to do it because it needs to not just be better, but so astronomically better in ways that are foundational to a company’s business model in order to justify the tremendous expense of a cloud migration, not just in real, out of pocket, cost in dollars and cents that are easy to measure, but also in terms of engineering effort, in terms of opportunity cost—because while you’re doing that you’re not doing other things instead—and, on some level, people tend to only do that when there’s an overwhelming strategic reason to do it. When folks already have existing workloads on AWS, as many of them do, it stands to reason that they are not going to want to completely deviate from that strategy just because something else offers a better database experience any number of axes. So, meeting customers where they are is one of the, I guess, foundational shifts that we’ve really seen from the entire IT industry over the last 40 years, rather than you will buy it from us and you will tolerate it. It’s, now customers have choice, and meeting them where they are and being much more, I guess, able to impedance-match with them has been critical. And I’m really optimistic about what the launch of this service portends for Oracle.

Nipun: Indeed, but let me give you another data point. We find a very large number of Aurora customers migrating to MySQL HeatWave on OCI, right? And this is the same workload they were running on Aurora, but now they want to run the same workload on MySQL HeatWave on OCI. They are willing to undertake this journey of migration because their applications, they get much faster, and for a lot less price, but they get much faster. Then the second aspect is, there’s another class of customers who are for instance running, on Aurora or other transactions or workloads, but then they have to keep moving the data, they’ll keep performing the ETL process into some other service, whether it’s Snowflake, or whether it’s Redshift for analytics.

Now, with this migration, when they move to MySQL HeatWave, customers don’t need to, like, have multiple databases, and they get real-time analytics, meaning that if any data changes inside the server inside the OLTP as a database service, right? If they were to run a query, that query is giving them the latest results, right? It’s not stale. Whereas with an ETL process, it gets to be stale. So, given that we already found that there were so many customers migrating to OCI to use MySQL HeatWave, I think there’s a clear value proposition of MySQL HeatWave, and there’s a lot of demand.

But like, as I was mentioning earlier, by having MySQL HeatWave be offered on AWS, it makes the proposition even more compelling because, as you said, yes, there is some engineering work that customers will need to do to migrate between clouds, and if they don’t want to, then absolutely now they have MySQL HeatWave which they can now use in AWS itself.

Corey: I think that one of the things I continually find myself careening into, perhaps unexpectedly, is a failure to really internalize just how vast this entire industry really is. Every time I think I’ve seen it all, all I have to do is talk to one more cloud customer and I learn something completely new and different. Sometimes it’s an innovative, exciting use of a thing. Other times, it’s people holding something fearfully wrong and trying to use it as a hammer instead. And you know, if it’s dumb and it works, is it really dumb? There are questions around that.

And this in turn gave rise to one of my next obnoxious questions as I was looking at what you were building at the time because a lot of your pricing and discussions and framing of this was targeting very large enterprise-style customers, and the price points reflected that. And then I asked the question that Big E enterprise never quite expects, for whatever reason, it’s like, “That looks awesome if I have a budget with many commas in it. What can I get for $4?” And as of this recording, pricing has not been finalized slash published for the service, but everything that you have shown me so far absolutely makes developing on this for a proof of concept or an evening puttering around, completely tenable: it is not bound to a fixed period of licensing; it’s, use it when you want to use it, turn it off when you’re done; and the hourly pricing is not egregious. I think that is something that historically, Oracle Database offerings have not really aligned with.

OCI very much has, particularly with an eye toward its extraordinarily awesome free tier that’s always free. But this feels like it’s a weird blending of the OCI model versus historical Oracle Database pricing models in a way that, honestly I’m pretty excited about.

Nipun: So, we react to what the customer requirements and needs are. So, for this class of customers who are using, say, RDS, MySQL, Aurora, we understand that they are very cost sensitive, right? So, one of the things which we have done in addition to offering MySQL HeatWave on AWS is based on the customer feedback and such. We are now offering a small shape of HeatWave instance in addition to the regular large shape. So, if customers want to just, you know, kick the tires, if developers just want to get started, they can get a MySQL node with HeatWave for less than ten cents an hour. So, for less than ten cents an hour, they get the ability to run transaction processing, analytics, and machine learning.

And if you were to compare the corresponding cost of Aurora for the same, like, you know, core count, it’s, like, you know, 12-and-a-half cents. And that’s just Aurora, without Redshift or without SageMaker. So yes, you’re right that based on the feedback and we have found that it would be much more attractive to have this low-end shape for the AWS developers. We are offering this smaller shape. And yeah, it’s very, very affordable. It’s about just shy of ten cents an hour.

Corey: This brings up another question that I raised pretty early on in the process because you folks kept talking about shapes, and it turns out that is the Oracle Cloud term that applies to instance size over an AWS-land. And as we dug into this a bit further, it does make sense for how you think about these things and how you build them to customers. Specifically, if I want to run this, I log into cloud.oracle.com and sign up for it there, and pay you over on that side of the world, this does not show up on my AWS bill. What drove that decision?

Nipun: Okay, so a couple of things. One clarification is that the site people log in to is cloud.mysql.com. So, that’s where they come to: cloud.mysql.com.

Corey: Oh, my apologies. I keep forgetting that you folks have multiple cloud offerings and domains. They’re kind of a thing. How do they work? Given I have a bad domain by habit myself, I have no room to judge.

Nipun: So, they come to cloud.mysql.com. From there, they can provision an instance. And we, as, like, you know, Oracle or MySQL, go ahead and create an instance in AWS, in the Oracle tenancy. From there, customers can then, you know, access their data on AWS and such. Now, what we want to provide the customers is a very seamless experience, that they just come to cloud.mysql.com, and from there, they can do everything: provisioning an instance, running the queries, payment and such. So, this is one of the reasons that we want customers just to be able to come to the site, cloud.mysql.com, and take care of the billing and such.

Now, the other thing is that, okay, why not allow customers to pay from AWS, right? Now, one of the things over there is that if you were to do that and there’s a customer, they’ll be like, “Hey, I got to pay something to AWS, something to Oracle, so we’d prefer, it’d be better to have a one-stop shop.” And since many of these are already Oracle customers, it’s helpful to do it this way.

Corey: Another approach you could have taken—and I want to be very clear here that I am not suggesting that this would have been a good idea—but an approach that you could have taken would have been to go down the weird AWS partner rabbit hole, and we’re going to provide this to customers on the AWS Marketplace. Because according to AWS, that’s where all of their customers go to discover new softwares. Yeah, first, that’s a lie. They do not. But aside from that, what was it about that Marketplace model that drove you to a decision point where okay, at launch, we are not going to be offering this on the AWS Marketplace? And to be clear, I’m not suggesting that was the wrong decision.

Nipun: Right. The main reason is we want to offer the MySQL HeatWave service at the least expensive cost to the user, right, or like, the least cost. If you were to, like, have MySQL HeatWave in the Marketplace, AWS charges a premium. This the customers would need to pay. So, we just didn’t want the customers to have to pay this additional premium just because they can now source this thing from the Marketplace. So, it’s really to, like, save costs for the customer.

Corey: The value of the Marketplace, from my perspective, has been effectively not having to deal as much with customer procurement departments because well, AWS is already on the procurement approved list, so we’re just going to go ahead and take the hit to wind up making it accessible from that perspective and calling it good. The downside to this is that increasingly, as customers are making larger and longer-term commitments that are tied to certain levels of spend on AWS, they’re increasingly trying to drag every vendor with whom they do business into the your AWS bill so they can check those boxes off. And the problem that I keep seeing with that is vendors who historically have been doing just fine, have great working relationships with a customer are reporting that suddenly customers are coming back with, “Yeah, so for our contract renewal, we want to go through the AWS Marketplace.” In return, effectively, these companies are then just getting a haircut off whatever it is they’re able to charge their customers but receiving no actual value for any of this. It attenuates the relationship by introducing a third party into the process, and it doesn’t make anything better from the vendor’s point of view because they already had something functional and working; now they just have to pay a commission on it to AWS, who, it seems, is pathologically averse to any transaction happening where they don’t get a cut, on some level. But I digress. I just don’t like that model very much at all. It feels coercive.

Nipun: That’s absolutely right. That’s absolutely right. And we thought that, yes, there is some value to be going to Marketplace, but it’s not worth the additional premium customers would need to pay. Totally agree.

Corey: This episode is sponsored in part by our friends at AWS AppConfig. Engineers love to solve, and occasionally create, problems. But not when it’s an on-call fire-drill at 4 in the morning. Software problems should drive innovation and collaboration, NOT stress, and sleeplessness, and threats of violence. That’s why so many developers are realizing the value of AWS AppConfig Feature Flags. Feature Flags let developers push code to production, but hide that that feature from customers so that the developers can release their feature when it’s ready. This practice allows for safe, fast, and convenient software development. You can seamlessly incorporate AppConfig Feature Flags into your AWS or cloud environment and ship your Features with excitement, not trepidation and fear. To get started, go to snark.cloud/appconfig. That’s snark.cloud/appconfig.

Corey: It’s also worth pointing out that in Oracle’s historical customer base, by which I mean the last 40 years that you folks have been in business, you do have significant customers with very sizable estates. A lot of your cloud efforts have focused around, I guess, we’ll call it an Oracle-specific currency: Oracle Credits. Which is similar to the AWS style of currency just for a different company in different ways. One of the benefits that you articulated to me relatively early on was that by going through cloud.mysql.com, customers with those credits—which can be in sizable amounts based upon various differentiating variables that change from case to case—and apply that to their use of MySQL HeatWave on AWS.

Nipun: Right. So, in fact, just for starters, right, what we give to customers is we offer some free credits for customers to try a service on OCI of, you know, $300. And that’s the same thing, the same experience you would like customers who are trying HeatWave on AWS to get. Yes, so you’re right, this is the kind of consistency we want to have, and yet another reason why cloud.mysql.com makes sense is the entry point for customers to try the service.

Corey: There was a time where I would have struggled not to laugh in your face at the idea that we’re talking about something in the context of an Oracle database, and well, there’s $300 in credit. That’s, “What can I get for that? Hung up on?” No. A surprising amount, when it comes to these things.

I feel like that opens up an entirely new universe of experimentation. And, “Let’s see how this thing actually works with his workload,” and lets people kick the tires on it for themselves in a way that, “Oh, we have this great database. Well, can I try it? Sure, for $8 million, you absolutely can.” “Well, it can stay great and awesome over there because who wants to take that kind of a bet?” It feels like it’s a new world and in a bunch of different respects, and I just can’t make enough noise about how glad I am to see this transformation happening.

Nipun: Yeah. Absolutely, right? So, just think about it. So, you’re getting MySQL and HeatWave together for just shy of ten cents an hour, right? So, what you could get for $300 is 3000 hours for MySQL HeatWave instance, which is very good for people to try for free. And then, you know, decide if they want to go ahead with it.

Corey: One other, I guess, obnoxious question that I love to ask—it’s not really a question so much as a statement; that that’s part of the first thing that makes it really obnoxious—but it always distills down to the following that upsets product people left and right, which is, “I don’t get it.” And one of the things that I didn’t fully understand at the outset of how you were structuring things was the idea of separating out HeatWave from its constituent components. I believe it was Autopilot if I’m not mistaken, and it was effectively different SKUs that you could wind up opting to go for. And okay, if I’m trying to kick the tires on this and contextualize it as someone for whom the world’s best database is Route 53, then it really felt like an additional decision point that I wasn’t clear on the value of. And I’m still not entirely sure on the differentiation point and the value there, but now you offer it bundled as a default, which I think is so much better, from the user experience perspective.

Nipun: Okay, so let me clarify a couple of things.

Corey: Please. Databases are not my forte, so expect me to wind up getting most of the details hilariously wrong.

Nipun: Sure. So, MySQL Autopilot provides machine-learning-based automation for various aspects of the MySQL service; very popular. There is no charge for it. It is built into MySQL HeatWave; there is no additional charge for it, right, so there never was any SKU for it. What you’re referring to is, we have had a SKU for the MySQL node or the MySQL instance, and there’s a separate SKU for HeatWave.

The reason there is a need to have a different SKU for these two is because you always only have one node of MySQL. It could be, like, you know, running on one core, or like, you know, multiple cores, but it’s always, like, you know, one node. But with HeatWave, it’s a scale-out architecture, so you can have multiple nodes. So, the users need to be able to express how many nodes of HeatWave are they provisioning, right? So, that’s why there is a need to have two SKUs, and we continue to have those two SKUs.

What we are doing now differently is that when users instantiate a MySQL instance, by default, they always get the HeatWave node associated with it, right? So, they don’t need to, like, you know, make the decision to—okay when to add HeatWave; they always get HeatWave along with the MySQL instance, and that’s what I was saying a combination of both of these is, you know, like, just about ten cents an hour. If for whatever reason, they decide that they do not want HeatWave, they can turn it off, and then the price drops to half. But what we’re providing is the AWS service that HeatWave is turned on by default.

Corey: Which makes an awful lot of sense. It’s something that lets people opt out if they decide they don’t need this as they continue to scale out, but for the newcomer who does not, in many cases—in my particular case—have a nuanced understanding of where this offering starts and stops, it’s clearly the right decision of—rather than, “Oh, yeah. The thing you were trying and it didn’t work super well? Well, yeah. If you enable this other thing, it would have been awesome.” “Well, great. Please enable it for me by default and let me opt out later in time as my level of understanding deepens.”

Nipun: That’s right. And that’s exactly what we are doing. Now, this was a feedback we got because many, if not most, of our customers would want to have HeatWave, and we just kind of, you know, mitigating them from going through one more step, it’s always enabled by default.

Corey: As far as I’m aware, you folks are running this effectively as any other AWS customer might, where you establish a private link connection to your customers, in some cases, or give them a public or private endpoint where they can wind up communicating with this service. It doesn’t require any favoritism or special permissions from AWS themselves that they wouldn’t give to any other random customer out there, correct?

Nipun: Yes, that is correct. So, for now, we are exposing this thing as a public endpoint. In the future, we have plans to support the private endpoint as well, but for now, it’s public.

Corey: Which means that foundationally what you’re building out is something that fits into a model that could work extraordinarily well across a variety of different environments. How purpose-tuned is the HeatWave installation you have running on AWS for the AWS environment, versus something that is relatively agnostic, could be dropped into any random cloud provider, up to and including the terrifyingly obsolete rack I have in the spare room?

Nipun: So, as I mentioned, when we decided to offer MySQL HeatWave on AWS, the idea was that okay, for the AWS customers, we now want to have an offering which is completely optimized for AWS, provides the best price-performance on AWS. So, we have determined which instance types underneath will provide the best price performance, and that’s what we have optimized for, right? So, I can tell you, like, in terms of many of—for instance, take the case of the cache size of the underlying processor that we’re using on AWS is different than what we’re using for OCI. So, we have gone ahead, made these optimizations in our code, and we believe that our code is really optimized now for the AWS infrastructure.

Corey: I think that makes a fair deal of sense because, again, one of the big problems AWS has had is the proliferation of EC2 instance types to the point now where the answer is super easy, too, “Are you using the correct instance type for your workload?” Because that answer now is, “Of course not. Who could possibly say that they were with any degree of confidence?” But when you take the time to look at a very specific workload that’s going to be scaled out, it’s worth the time investment to figure out exactly how to optimize things for price and performance, given the constraints. Let’s be very clear here, I would argue that the better price performance for HeatWave is almost certainly not going to be on AWS themselves, if for no other reason than the joy that is their data transfer pricing, even for internal things moving around from time to time.

Personally, I love getting charged data transfer for taking data from S3, running it through AWS Glue, putting it into a different S3 bucket, accessing it with Athena, then hooking that up to Tableau as we go down and down and down the spiraling rabbit hole that never ends. It’s not exactly what I would call well-optimized economically. Their entire system feels almost like it’s a rigged game, on some level. But given those constraints, yeah, dialing in it and making it cost-effective is absolutely something that I’ve watched you folks put significant time and effort into.

Nipun: So, I’ll make two points, right, to the questions. First is yes, I just want to, like, be clear about it, that when a user provisions MySQL HeatWave via cloud.mysql.com and we create an instance in AWS, we don’t give customers a multitude of things to, like, you know, choose from.

We have determined which instance type is going to provide the customer the best price performance, and that’s what we provision. So, the customer doesn’t even need to know or care, is it going to be, like, you know, AMD? Is it going to be Intel? Is it going to be, like, you know, ARM, right? So, it’s something which we have predetermined and we have optimized for it. That’s first.

The second point is in terms of the price performance. So, you’re absolutely right, that for the class of customers who cannot migrate away from AWS because of the egress costs or because of the high latency because of AWS, right, sure, MySQL HeatWave on AWS will provide the best price-performance compared to other services out in AWS like Redshift, or Aurora, or Snowflake. But if customers have the flexibility to choose a cloud of their choice, it is indeed the case that customers are going to find that running MySQL HeatWave on OCI is going to provide them, by far, the best price performance, right? So, the price performance of running MySQL HeatWave on OCI is indeed better than MySQL HeatWave on AWS. And just because of the fact that when we are running the service in AWS, we are paying the list price, right, on AWS; that’s how we get the gear. Whereas with OCI, like, you know, things are a lot less expensive for us.

But even when you’re running on AWS, we are very, very price competitive with other services. And you know, as you’ve probably seen from the performance benchmarks and such, what I’m very intrigued about is that we’re able to run a standard workload, like some, like, you know, TPC-H and offer seven times better price-performance while running in AWS compared to Redshift. So, what this goes to show is that we are really passing on the savings to the customers. And clearly, Redshift is not doing a good job of performance or, like, you know, they’re charging too much. But the fact that we can offer seven times better price performance than Redshift in AWS speaks volumes, both about architecture and how much of savings we are passing to our customers.

Corey: What I love about this story is that it makes testing the waters of what it’s like to run MySQL HeatWave a lot easier for customers because the barrier to entry is so much lower. Where everything you just said I agree with it is more cost-effective to run on Oracle Cloud. I think there are a number of workloads that are best placed on Oracle Cloud. But unless you let people kick the tires on those things, where they happen to be already, it’s difficult to get them to a point where they’re going to be able to experience that themselves. This is a massive step on that path.

Nipun: Yep. Right.

Corey: I really want to thank you for taking time out of your day to walk us through exactly how this came to be and what the future is going to look like around this. If people want to learn more, where should they go?

Nipun: Oh, they can go to oracle.com/mysql, and there they can get a lot more information about the capabilities of MySQL HeatWave, what we are offering in AWS, price-performance. By the way, all the price performance numbers I was talking about, all the scripts are available publicly on GitHub. So, we welcome, we encourage customers to download the scripts from GitHub, try for themselves, and all of this information is available from oracle.com/mysql where they can get this detailed information.

Corey: And we will, of course, put links to that in the show notes. Thank you so much for your time. I appreciate it.

Nipun: Sure thing, Corey. Thank you for the opportunity.

Corey: Nipun Agarwal, Senior Vice President of MySQL HeatWave. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry insulting comment. You will then be overcharged for the data transfer to submit that insulting comment, and then AWS will take a percentage of that just because they’re obnoxious and can.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Thomas

Thomas Hazel is Founder, CTO, and Chief Scientist of ChaosSearch. He is a serial entrepreneur at the forefront of communication, virtualization, and database technology and the inventor of ChaosSearch's patented IP. Thomas has also patented several other technologies in the areas of distributed algorithms, virtualization and database science. He holds a Bachelor of Science in Computer Science from University of New Hampshire, Hall of Fame Alumni Inductee, and founded both student & professional chapters of the Association for Computing Machinery (ACM).

Links Referenced:

  • ChaosSearch: https://www.chaossearch.io/
  • Twitter: https://twitter.com/ChaosSearch
  • Facebook: https://www.facebook.com/CHAOSSEARCH/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at AWS AppConfig. Engineers love to solve, and occasionally create, problems. But not when it’s an on-call fire-drill at 4 in the morning. Software problems should drive innovation and collaboration, NOT stress, and sleeplessness, and threats of violence. That’s why so many developers are realizing the value of AWS AppConfig Feature Flags. Feature Flags let developers push code to production, but hide that that feature from customers so that the developers can release their feature when it’s ready. This practice allows for safe, fast, and convenient software development. You can seamlessly incorporate AppConfig Feature Flags into your AWS or cloud environment and ship your Features with excitement, not trepidation and fear. To get started, go to snark.cloud/appconfig. That’s snark.cloud/appconfig.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. This promoted episode is brought to us by our returning sponsor and friend, ChaosSearch. And once again, the fine folks at ChaosSearch has seen fit to basically subject their CTO and Founder, Thomas Hazel, to my slings and arrows. Thomas, thank you for joining me. It feels like it’s been a hot minute since we last caught up.

Thomas: Yeah, Corey. Great to be on the program again, then. I think it’s been almost a year. So, I look forward to these. They’re fun, they’re interesting, and you know, always a good time.

Corey: It’s always fun to just take a look at companies’ web pages in the Wayback Machine, archive.org, where you can see snapshots of them at various points in time. Usually, it feels like this is either used for long-gone things and people want to remember the internet of yesteryear, or alternately to deliver sick burns with retorting a “This you,” when someone winds up making an unpopular statement. One of the approaches I like to use it for, which is significantly less nefarious—usually—is looking back in time at companies’ websites, just to see how the positioning of the product evolves over time.

And ChaosSearch has had an interesting evolution in that direction. But before we get into that, assuming that there might actually be people listening who do not know the intimate details of exactly what it is you folks do, what is ChaosSearch, and what might you folks do?

Thomas: Yeah, well said, and I look forward to [laugh] doing the Wayback Time because some of our ideas, way back when, seemed crazy, but now they make a lot of sense. So, what ChaosSearch is all about is transforming customers’ cloud object stores like Amazon S3 into an analytical database that supports search and SQL-type use cases. Now, where’s that apply? In log analytics, observability, security, security data lakes, operational data, particularly at scale, where you just stream your data into your data lake, connect our service, our SaaS service, to that lake and automagically we index it and provide well-known APIs like Elasticsearch and integrate with Kibana or Grafana, and SQL APIs, something like, say, a Superset or Tableau or Looker into your data. So, you stream it in and you get analytics out. And the key thing is the time-cost complexity that we all know that operational data, particularly at scale, like terabytes and a day and up causes challenges, and we all know how much it costs.

Corey: They certainly do. One of the things that I found interesting is that, as I’ve mentioned before, when I do consulting work at The Duckbill Group, we have absolutely no partners in the entire space. That includes AWS, incidentally. But it was easy in the beginning because I was well aware of what you folks were up to, and it was great when there was a use case that matched of you’re spending an awful lot of money on Elasticsearch; consider perhaps migrating some of that—if it makes sense—to ChaosSearch. Ironically, when you started sponsoring some of my nonsense, that conversation got slightly trickier where I had to disclose, yeah our media arm is does have sponsorships going on with them, but that has no bearing on what I’m saying.

And if they take their sponsorships away—please don’t—then we would still be recommending them because it’s the right answer, and it’s what we would use if we were in your position. We receive no kickbacks or partner deal or any sort of reseller arrangement because it just clouds the whole conflict of interest perception. But you folks have been fantastic for a long time in a bunch of different ways.

Thomas: Well, you know, I would say that what you thought made a lot of sense made a lot of sense to us as well. So, the ChaosSearch idea just makes sense. Now, you had to crack some code, solve some problems, invent some technology, and create some new architecture, but the idea that Elasticsearch is a useful solution with all the tooling, the visualization, the wonderful community around that, was a good place to start, but here’s the problem: setting it up, scaling it out, keep it up, when things are happening, things go bump in the night. All those are real challenges, and one of them was just the storaging of the data. Well, what if you could make S3 the back-end store? One hundred percent; no SSDs or HDDs. Makes a lot of sense.

And then support the APIs that your tooling uses. So, it just made a lot of sense on what we were trying to do, just no one thought of it. Now, if you think about the Northstar you were talking about, you know, five, six years ago, when I said, transforming cloud storage into an analytical database for search and SQL, people thought that was crazy and mad. Well, now everyone’s using Cloud Storage, everyone’s using S3 as a data lake. That’s not in question anymore.

But it was a question five, six, you know, years ago. So, when we met up, you’re like, “Well, that makes sense.” It always made sense, but people either didn’t think was possible, or were worried, you know, I’ll just try to set up an Elastic cluster and deal with it. Because that’s what happens when you particularly deal with large-scale implementations. So, you know, to us, we would love the Elastic API, the tooling around it, but what we all know is the cost, the time the complexity, to manage it, to scale it out, just almost want to pull your hair out. And so, that’s where we come in is, don’t change what you do, just change how you do it.

Corey: Every once in a while, I’ll talk to a client who’s running an Amazon Elasticsearch cluster, and they have nothing but good things to say about it. Which, awesome. On the one hand, part of me wishes that I had some of their secrets, but often what’s happened is that they have this down to a science, they have a data lifecycle that’s clearly defined and implemented, the cluster is relatively static, so resizes aren’t really a thing, and it just works for their use cases. And in those scenarios, like, “Do you care about the bill?” “Not overly. We don’t have to think about it.”

Great. Then why change? If there’s no pain, you’re not going to sell someone something, especially when we’re talking, this tends to be relatively smaller-scale as well. It’s okay, great, they’re spending $5,000 a month on it. It doesn’t necessarily justify the engineering effort to move off.

Now, when you start looking at this, and, “Huh, that’s a quarter million bucks a month we’re spending on this nonsense, and it goes down all the time,” yeah, that’s when it starts to be one of those logical areas to start picking apart and diving into. What’s also muddied the waters since the last time we really went in-depth on any of this was it used to be we would be talking about it exactly like we are right now, about how it's Elasticsearch-compatible. Technically, these days, we probably shouldn’t be saying it is OpenSearch compatible because of the trademark issues between Elastic and AWS and the Schism of the OpenSearch fork of the Elasticsearch project. And now it feels like when you start putting random words in front of the word search, ChaosSearch fits right in. It feels like your star is rising.

Thomas: Yeah, no, well said. I appreciate that. You know, it’s funny when Elastic changed our license, we all didn’t know what was going to happen. We knew something was going to happen, but we didn’t know what was going to happen. And Amazon, I say ironically, or, more importantly, decided they’ll take up the open mantle of keeping an open, free solution.

Now, obviously, they recommend running that in their cloud. Fair enough. But I would say we don’t hear as much Elastic replacement, as much as OpenSearch replacement with our solution because of all the benefits that we talked about. Because the trigger points for when folks have an issue with the OpenSearch or Elastic stack is got too expensive, or it was changing so much and it was falling over, or the complexity of the schema changing, or all the above. The pipelines were complex, particularly at scale.

That’s both for Elasticsearch, as well as OpenSearch. And so, to us, we want either to win, but we want to be the replacement because, you know, at scale is where we shine. But we have seen a real trend where we see less Elasticsearch and more OpenSearch because the community is worried about the rules that were changed, right? You see it day in, day out, where you have a community that was built around open and fair and free, and because of business models not working or the big bad so-and-so is taking advantage of it better, there’s a license change. And that’s a trust change.

And to us, we’re following the OpenSearch path because it’s still open. The 600-pound gorilla or 900-pound gorilla of Amazon. But they really held the mantle, saying, “We’re going to stay open, we assume for as long as we know, and we’ll follow that path. But again, at that scale, the time, the costs, we’re here to help solve those problems.” Again, whether it’s on Amazon or, you know, Google et cetera.

Corey: I want to go back to what I mentioned at the start of this with the Wayback Machine and looking at how things wound up unfolding in the fullness of time. The first time that it snapshotted your site was way back in the year 2018, which—

Thomas: Nice. [laugh].

Corey: Some of us may remember, and at that point, like, I wasn’t doing any work with you, and later in time I would make fun of you folks for this, but back then your brand name was in all caps, so I would periodically say things like this episode is sponsored by our friends at [loudly] CHAOSSEARCH.

Thomas: [laugh].

Corey: And once you stopped capitalizing it and that had faded from the common awareness, it just started to look like I had the inability to control the volume of my own voice. Which, fair, but generally not mid-sentence. So, I remember those early days, but the positioning of it was, “The future of log management and analytics,” back in 2018. Skipping forward a year later, you changed this because apparently in 2019, the future was already here. And you were talking about, “Log search analytics, purpose-built for Amazon S3. Store everything, ask anything all on your Amazon S3.”

Which is awesome. You were still—unfortunately—going by the all caps thing, but by 2020, that wound up changing somewhat significantly. You were at that point, talking for it as, “The data platform for scalable log analytics.” Okay, it’s clearly heading in a log direction, and that made a whole bunch of sense. And now today, you are, “The data lake platform for analytics at scale.” So, good for you, first off. You found a voice?

Thomas: [laugh]. Well, you know, it’s funny, as a product mining person—I’ll take my marketing hat off—we’ve been building the same solution with the same value points and benefits as we mentioned earlier, but the market resonates with different terminology. When we said something like, “Transforming your Cloud Object Storage like S3 into an analytical database,” people were just were like, blown away. Is that even possible? Right? And so, that got some eyes.

Corey: Oh, anything is a database if you hold that wrong. Absolutely.

Thomas: [laugh]. Yeah, yeah. And then you’re saying log analytics really resonated for a few years. Data platform, you know, is more broader because we do more broader things. And now we see over the last few years, observability, right? How do you fit in the observability viewpoint, the stack where log analytics is one aspect to it?

Some of our customers use Grafana on us for that lens, and then for the analysis, alerting, dashboarding. You can say that Kibana in the hunting aspect, the log aspects. So, you know, to us, we’re going to put a message out there that resonates with what we’re hearing from our customers. For instance, we hear things like, “I need a security data lake. I need that. I need to stream all my data. I need to have all the data because what happens today that now, I need to know a week, two weeks, 90 days.”

We constantly hear, “I need at least 90 days forensics on that data.” And it happens time and time again. We hear in the observability stack where, “Hey, I love Datadog, but I can’t afford it more than a week or two.” Well, that’s where we come in. And we either replace Datadog for the use cases that we support, or we’re auxiliary to it.

Sometimes we have an existing Grafana implementation, and then they store data in us for the long tail. That could be the scenario. So, to us, the message is around what resonates with our customers, but in the end, it’s operational data, whether you want to call it observability, log analytics, security analytics, like the data lake, to us, it’s just access to your data, all your data, all the time, and supporting the APIs and the tooling that you’re using. And so, to me, it’s the same product, but the market changes with messaging and requirements. And this is why we always felt that having a search and SQL platform is so key because what you’ll see in Elastic or OpenSearch is, “Well, I only support the Elastic API. I can’t do correlations. I can’t do this. I can’t do that. I’m going to move it over to say, maybe Athena but not so much. Maybe a Snowflake or something else.”

Corey: “Well, Thomas, it’s very simple. Once you learn our own purpose-built, domain-specific language, specifically for our product, well, why are you still sitting here, go learn that thing.” People aren’t going to do that.

Thomas: And that’s what we hear. It was funny, I won’t say what the company was, a big banking company that we’re talking to, and we hear time and time again, “I only want to do it via the Elastic tooling,” or, “I only want to do it via the BI tooling.” I hear it time and time again. Both of these people are in the same company.

Corey: And that’s legitimate as well because there’s a bunch of pre-existing processes pointing at things and we’re not going to change 200 different applications in their data model just because you want to replace a back-end system. I also want to correct myself. I was one tab behind. This year’s branding is slightly different: “Search and analyze unlimited log data in your cloud object storage.” Which is, I really like the evolution on this.

Thomas: Yeah, yeah. And I love it. And what was interesting is the moving, the setting up, the doubling of your costs, let’s say you have—I mean, we deal with some big customers that have petabytes of data; doubling your petabytes, that means, if your Elastic environment is costing you tens of millions and then you put into Snowflake, that’s also going to be tens of millions. And with a solution like ours, you have really cost-effective storage, right? Your cloud storage, it’s secure, it’s reliable, it’s Elastic, and you attach Chaos to get the well-known APIs that your well-known tooling can analyze.

So, to us, our evolution has been really being the end viewpoint where we started early, where the search and SQL isn’t here today—and you know, in the future, we’ll be coming out with more ML type tooling—but we have two sides: we have the operational, security, observability. And a lot of the business side wants access to that data as well. Maybe it’s app data that they need to do analysis on their shopping cart website, for instance.

Corey: The thing that I find curious is, the entire space has been iterating forward on trying to define observability, generally, as whatever people are already trying to sell in many cases. And that has seemed to be a bit of a stumbling block for a lot of folks. I figured this out somewhat recently because I’ve built the—free for everyone to use—the lasttweetinaws.com, Twitter threading client.

That’s deployed to 20 different AWS regions because it’s go—the idea is that should be snappy for people, no matter where they happen to be on the planet, and I use it for conferences when I travel, so great, let’s get ahead of it. But that also means I’ve got 20 different sources of logs. And given that it’s an omnibus Lambda function, it’s very hard to correlate that to users, or user sessions, or even figure out where it’s going. The problem I’ve had is, “Oh, well, this seems like something I could instrument to spray logs somewhere pretty easily, but I don’t want to instrument it for 15 different observability vendors. Why don’t I just use otel—or Open Telemetry—and then tell that to throw whatever I care about to various vendors and do a bit of a bake-off?” The problem, of course, is that open telemetry and Lambda seem to be in just the absolute wrong directions. A lot.

Thomas: So, we see the same trend of otel coming out, and you know, this is another API that I’m sure we’re going to go all-in on because it’s getting more and more talked about. I won’t say it’s the standard that I think is trending to all your points about I need to normalize a process. But as you mentioned, we also need to correlate across the data. And this is where, you know, there are times where search and hunting and alerting is awesome and wonderful and solves all your needs, and sometimes correlation. Imagine trying to denormalize all those logs, set up a pipeline, put it into some database, or just do a SELECT *, you know, join this to that to that, and get your answers.

And so, I think both OpenTelemetry and SQL and search all need to be played into one solution, or at least one capability because if you’re not doing that, you’re creating some hodgepodge pipeline to move it around and ultimately get your questions answered. And if it takes weeks—maybe even months, depending on the scale—you may sometimes not choose to do it.

Corey: One other aspect that has always annoyed me about more or less every analytics company out there—and you folks are no exception to this—is the idea of charging per gigabyte ingested because that inherently sets up a weird dichotomy of, well, this is costing a lot, so I should strive to log less. And that is sort of the exact opposite, not just of the direction you folks want customers to go in, but also where customers themselves should be going in. Where you diverge from an awful lot of those other companies because of the nature of how you work, is that you don’t charge them again for retention. And the idea that, yeah, the fact that anything stored in ChaosSearch lives in your own S3 buckets, you can set your own lifecycle policies and do whatever you want to do with that is a phenomenal benefit, just because I’ve always had a dim view of short-lived retention periods around logs, especially around things like audit logs. And these days, I would consider getting rid of audit logging data and application logging data—especially if there’s a correlation story—any sooner than three years feels like borderline malpractice.

Thomas: [laugh]. We—how many times—I mean, we’ve heard it time and time again is, “I don’t have access to that data because it was too costly.” No one says they don’t want the data. They just can’t afford the data. And one of the key premises that if you don’t have all the data, you’re at risk, particularly in security—I mean, even audits. I mean, so many times our customers ask us, you know, “Hey, what was this going on? What was that go on?” And because we can so cost-effectively monitor our own service, we can provide that information for them. And we hear this time and time again.

And retention is not a very sexy aspect, but it’s so crucial. Anytime you look in problems with X solution or Y solution, it’s the cost of the data. And this is something that we wanted to address, officially. And why do we make it so cost-effective and free after you ingest it was because we were using cloud storage. And it was just a great place to land the data cost-effective, securely.

Now, with that said, there are two types of companies I’ve seen. Everybody needs at least 90 days. I see time and time again. Sure, maybe daily, in a weeks, they do a lot of their operation, but 90 days is where it lands. But there’s also a bunch of companies that need it for years, for compliance, for audit reasons.

And imagine trying to rehydrate, trying to rebuild—we have one customer—again I won’t say who—has two petabytes of data that they rehydrate when they need it. And they say it’s a nightmare. And it’s growing. What if you just had it always alive, always accessible? Now, as we move from search to SQL, there are use cases where in the log world, they just want to pay upfront, fixed fee, this many dollars per terabyte, but as we get into the more ad hoc side of it, more and more folks are asking for, “Can I pay per query?”

And so, you’ll see coming out soon, about scenarios where we have a different pricing model. For logs, typically, you want to pay very consistent, you know, predetermined cost structure, but in the case of more security data lakes, where you want to go in the past and not really pay for something until you use it, that’s going to be an option as well coming out soon. So, I would say you need both in the pricing models, but you need the data to have either side, right?

Corey: This episode is sponsored in part by our friends at ChaosSearch. You could run Elasticsearch or Elastic Cloud—or OpenSearch as they’re calling it now—or a self-hosted ELK stack. But why? ChaosSearch gives you the same API you’ve come to know and tolerate, along with unlimited data retention and no data movement. Just throw your data into S3 and proceed from there as you would expect. This is great for IT operations folks, for app performance monitoring, cybersecurity. If you’re using Elasticsearch, consider not running Elasticsearch. They’re also available now in the AWS marketplace if you’d prefer not to go direct and have half of whatever you pay them count towards your EDB commitment. Discover what companies like Equifax, Armor Security, and Blackboard already have. To learn more, visit chaossearch.io and tell them I sent you just so you can see them facepalm, yet again.

Corey: You’d like to hope. I mean, you could always theoretically wind up just pulling what Ubiquiti apparently did—where this came out in an indictment that was unsealed against an insider—but apparently one of their employees wound up attempting to extort them—which again, that’s not their fault, to be clear—but what came out was that this person then wound up setting the CloudTrail audit log retention to one day, so there were no logs available. And then as a customer, I got an email from them saying there was no evidence that any customer data had been accessed. I mean, yeah, if you want, like, the world’s most horrifyingly devilish best practice, go ahead and set your log retention to nothing, and then you too can confidently state that you have no evidence of anything untoward happening.

Contrast this with what AWS did when there was a vulnerability reported in AWS Glue. Their analysis of it stated explicitly, “We have looked at our audit logs going back to the launch of the service and have conclusively proven that the only time this has ever happened was in the security researcher who reported the vulnerability to us, in their own account.” Yeah, one of those statements breeds an awful lot of confidence. The other one makes me think that you’re basically being run by clowns.

Thomas: You know what? CloudTrail is such a crucial—particularly Amazon, right—crucial service because of that, we see time and time again. And the challenge of CloudTrail is that storing a long period of time is costly and the messiness the JSON complexity, every company struggles with it. And this is how uniquely—how we represent information, we can model it in all its permutations—but the key thing is we can store it forever, or you can store forever. And time and time again, CloudTrail is a key aspect to correlate—to your question—correlate this happened to that. Or do an audit on two years ago, this happened.

And I got to tell you, to all our listeners out there, please store your CloudTrail data—ideally in ChaosSearch—because you’re going to need it. Everyone always needs that. And I know it’s hard. CloudTrail data is messy, nested JSON data that can explode; I get it. You know, there’s tricks to do it manually, although quite painful. But CloudTrail, every one of our customers is indexing with us in CloudTrail because of stories like that, as well as the correlation across what maybe their application log data is saying.

Corey: I really have never regretted having extra logs lying around, especially with, to be very direct, the almost ridiculously inexpensive storage classes that S3 offers, especially since you can wind up having some of the offline retrieval stuff as part of a lifecycle policy now with intelligent tiering. I’m a big believer in just—again—the Glacier Deep Archive I’ve at the cost of $1,000 a month per petabyte, with admittedly up to 12 hours of calling that as a latency. But that’s still, for audit logs and stuff like that, why would I ever want to delete things ever again?

Thomas: You’re exactly right. And we have a bunch of customers that do exactly that. And we automate the entire process with you. Obviously, it’s your S3 account, but we can manage across those tiers. And it’s just to a point where, why wouldn’t you? It’s so cost-effective.

And the moments where you don’t have that information, you’re at risk, whether it’s internal audits, or you’re providing a service for somebody, it’s critical data. With CloudTrail, it’s critical data. And if you’re not storing it and if you’re not making it accessible through some tool like an Elastic API or Chaos, it’s not worth it. I think, to your point about your story, it’s epically not worth it.

Corey: It’s really not. It’s one of those areas where that is not a place to overly cost optimize. This is—I mean we talked earlier about my business and perceptions of conflict of interest. There’s a reason that I only ever charge fixed-fee and not percentage of savings or whatnot because, at some point, I’ll be placed in a position of having to say nonsense, like, “Do you really need all of these backups?” That doesn’t make sense at that point.

I do point out things like you have hourly disk snapshots of your entire web fleet, which has no irreplaceable data on them dating back five years. Maybe cleaning some of that up might be the right answer. The happy answer is somewhere in between those two, and it’s a business decision around exactly where that line lies. But I’m a believer in never regretting having kept logs almost into perpetuity. Until and unless I start getting more or less pillaged by some particularly rapacious vendor that’s oh, yeah, we’re going to charge you not just for ingest, but also for retention. And for how long you want to keep it, we’re going to treat it like we’re carving it into platinum tablets. No. Stop that.

Thomas: [laugh]. Well, you know, it’s funny, when we first came out, we were hearing stories that vendors were telling customers why they didn’t need their data, to your point, like, “Oh, you don’t need that,” or, “Don’t worry about that.” And time and time again, they said, “Well, turns out we didn’t need that.” You know, “Oh, don’t index all your data because you just know what you know.” And the problem is that life doesn’t work out that way business doesn’t work out that way.

And now what I see in the market is everyone’s got tiering scenarios, but the accessibility of that data takes some time to get access to. And these are all workarounds and bandaids to what fundamentally is if you design an architecture and a solution is such a way, maybe it’s just always hot; maybe it’s just always available. Now, we talked about tiering off to something very, very cheap, then it’s like virtually free. But you know, our solution was, whether it’s ultra warm, or this tiering that takes hours to rehydrate—hours—no one wants to live in that world, right? They just want to say, “Hey, on this date on this year, what was happening? And let me go look, and I want to do it now.”

And it has to be part of the exact same system that I was using already. I didn’t have to call up IT to say, “Hey, can you rehydrate this?” Or, “Can I go back to the archive and look at it?” Although I guess we’re talking about archiving with your website, viewing from days of old, I think that’s kind of funny. I should do that more often myself.

Corey: I really wish that more companies would put themselves in the customers’ shoes. And for what it’s worth, periodically, I’ve spoken to a number of very happy ChaosSearch customers. I haven’t spoken to any angry ones yet, which tells me you’re either terrific at crisis comms, or the product itself functions as intended. So, either way, excellent job. Now, which team of yours is doing that excellent job, of course, is going to depend on which one of those outcomes it is. But I’m pretty good at ferreting out stories on those things.

Thomas: Well, you know, it’s funny, being a company that’s driven by customer ask, it’s so easy build what the customer wants. And so, we really take every input of what the customer needs and wants—now, there are cases where we relace Splunk. They’re the Cadillac, they have all the bells and whistles, and there’s times where we’ll say, “Listen, that’s not what we’re going to do. We’re going to solve these problems in this vector.” But they always keep on asking, right? You know, “I want this, I want that.”

But most of the feedback we get is exactly what we should be building. People need their answers and how they get it. It’s really helped us grow as a company, grow as a product. And I will say ever since we went live now many, many years ago, all our roadmap—other than our Northstar of transforming cloud storage into a search SQL big data analytics database has been customer-driven, market customer-driven, like what our customer is asking for, whether it’s observability and integrating with Grafana and Kibana or, you know, security data lakes. It’s just a huge theme that we’re going to make sure that we provide a solution that meets those needs.

So, I love when customers ask for stuff because the product just gets better. I mean, yeah, sometimes you have to have a thick skin, like, “Why don’t you have this?” Or, “Why don’t you have that?” Or we have customers—and not to complain about customers; I love our customers—but they sometimes do crazy things that we have to help them on crazy-ify. [laugh]. I’ll leave it at that. But customers do silly things and you have to help them out. I hope they remember that, so when they ask for a feature that maybe takes a month to make available, they’re patient with us.

Corey: We sure can hope. I really want to thank you for taking so much time to once again suffer all of my criticisms, slings and arrows, blithe market observations, et cetera, et cetera. If people want to learn more, where’s the best place to find you?

Thomas: Well, of course, chaossearch.io. There’s tons of material about what we do, use cases, case studies; we just published a big case study with Equifax recently. We’re in Gartner and a whole bunch of Hype Cycles that you can pull down to see how we fit in the market.

Reach out to us. You can set up a trial, kick the tires, again, on your cloud storage like S3. And ChaosSearch on Twitter, we have a Facebook, we have all this classic social medias. But our website is really where all the good content and whether you want to learn about the architecture and how we’ve done it, and use cases; people who want to say, “Hey, I have a problem. How do you solve it? How do I learn more?”

Corey: And we will, of course, put links to that in the show notes. For my own purposes, you could also just search for the term ChaosSearch in your email inbox and find one of their sponsored ads in my newsletter and click that link, but that’s a little self-serving as we do it. I’m kidding. I’m kidding. There’s no need to do that. That is not how we ever evaluate these things. But it is funny to tell that story. Thomas, thank you so much for your time. As always, it’s appreciated.

Thomas: Corey Quinn, I truly enjoyed this time. And I look forward to upcoming re:Invent. I’m assuming it’s going to be live like last year, and this is where we have a lot of fun with the community.

Corey: Oh, I have no doubt that we’re about to go through that particular path very soon. Thank you. It’s been an absolute pleasure.

Thomas: Thank you.

Corey: Thomas Hazel, CTO and Founder of ChaosSearch. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an angry, insulting comment that I will then set to have a retention period of one day, and then go on to claim that I have received no negative feedback.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Donovan

Donovan Brady is the Director of Solutions Architecture at Logicworks. He began his career at Logicworks six years ago as a Solutions Architect, fast forward to today, Donovan now manages a team of highly skilled and certified AWS and Azure Solutions Architects. During his time at Logicworks, Donovan has had the opportunity to work with companies in a variety of verticals to solve their most complex IT and business challenges. Donovan is originally from New York and has been a professional musician since the age of six. He is also a self-proclaimed 90’s video game nerd.

Links Referenced:

  • LogicWorks: https://www.logicworks.com/
  • Donovan’s LinkedIn: https://www.linkedin.com/in/donovan-brady-9403a583/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. My guest on this promoted episode of Screaming in the Cloud is Donovan Brady, director of solutions architecture at Logicworks and something of a kindred spirit in that he tends to also focus on something that I more or less spend my time being obnoxious about on Twitter, which is in many cases going towards cloud for the right reasons with an outcome in mind, which is rarely having the most interesting and clever technical stack imaginable. Donovan, thank you for joining me today.

Donovan: Yeah, thanks for having me. Corey, really excited for the conversation and looking forward to getting into it.

Corey: Let’s start by establishing the bona fides, for lack of a better term here. What does Logicworks do that they require a director of solutions architecture?

Donovan: Logicworks is a managed services and professional services cloud provider specializing in hyper-compliant workloads migrating, optimizing, and operating in the cloud, right? That means that we work with primarily Software-as-a-Service companies or HealthTech or FinTech companies with compliances, like PCI, HIPAA, HITRUST, SOC, ISO, GDPR, pretty much you name it and we help customers in their cloud journey operate and optimize in the cloud.

Corey: It’s weird, when you talk about optimizing in cloud, people always hear that as, “Oh, we’re going to fix the bill, we’re going to—” because again, that is the context in which I operate where, “Oh, great. We’re going to optimize your cloud bill.” Which makes sense, but I think that people have lost sight of the forest for the trees in many respects, where when they hear, “I’m going to optimize your bill,” that often comes across, “Oh, we’re going to make it smaller,” because generally, it’s not very well optimized, and one of those optimal things you can do is turn something that’s unneeded or unnecessary off, and surprise, there’s a side effect of saving money. But it often means that in some cases, it’s time to start spending more money on things like, oh I don’t know, backups, and resiliency, and figure out what it is that the business is aligned around. Because you can always cut the bill to zero by turning everything off. It seems like that’s not really the alignment that—or the reason that companies go to cloud in the first place. So, what’s your take on that? Why do most companies say, “Ah. We have a problem and we’re going to go with cloud,” in the hopes that it fixes that problem?

Donovan: Yeah, it’s an interesting point. So, a lot of times we hear customers say exactly what you just mentioned: “We want to move to cloud so that we can save costs,” or, “Oh, we’re in the cloud and we’re primarily concerned about costs.” Unfortunately, that’s the wrong motivation. Saving cost is definitely a byproduct of moving to the cloud if you do it the right way, but the primary business reasons why somebody or an organization would want to move to the cloud are slightly different, right? The business objectives that most customers or companies want to increase are their agility, they want to increase their profitability, they want to decrease their time-to-revenue.

Let’s say—as I mentioned, I work with a lot of Software-as-a-Service companies—they have a monolithic application right now that takes a long time to update, it takes a long time to patch. If you can modernize that application, if you can use some more cutting edge or bleeding edge tools like what’s provided in the public cloud, you’d be able to significantly decrease that timetable for deploying new updates, acquiring new companies, decreasing your time to revenue. Those are the primary business drivers that customers should be focused on. And then when you’re optimizing in the cloud, you’re really looking at the five categories of the Well-Architected Framework, which as you mentioned, isn’t just cost. There’s security, there’s reliability, there’s performance, and there’s operations, and all of these are kind of intertwined, and can eventually lead to a decrease in [laugh] costs, right, but if you were to just lift-and-shift into the cloud, you’re probably not going to save that much money.

Corey: Let’s not forget the sixth pillar of the Well-Architected Framework, which was recently added which is sustainability. Unfortunately, it does seem a little true to life where it seems like it’s been bolted on after the fact. I would have expected that to also be the security pillar, but that’s a little sensitive for some folks. It’s one of those areas where sustainability and cost optimization tend to go hand-in-glove because turning something off benefits everyone except the cloud providers hoping that you don’t turn things off in some cases. I don’t necessarily believe that’s where the major hyperscalers are sitting today, but there’s no denying that they do benefit from things sitting there going unused, as do most companies that charge money for a thing that they provide you.

Donovan: Yeah, that’s exactly true. There are some other tools that make it easier to save costs and still achieve your expected goals, and that’s the more of those more cutting-edge technologies like a serverless deployment, right? A Lambda function is a point-in-time deployment of your code without needing to rely on an ongoing virtual machine, or database, or whatever it is that is running that application, and it runs for just a couple minutes, couple seconds, however long you need it. And I was actually just speaking with a customer recently, we did this CXO dinner, and we were talking about the benefits of the cloud. It was almost as if we planted him there.

We didn’t; he randomly showed up. But he was also promoting the cloud because he said he runs his entire organization almost in the free tier of AWS because the entire thing is serverless-based, you know? So, there are definitely ways to optimize your costs and therefore sustainability, as you mentioned: turning things off or maybe not have them permanently running in the first place. But you can only achieve that once you’ve actually done the shift after the lift.

Corey: The website that hosts this podcast is lastweekinaws.com. The first version of that that I put up was entirely Lambda and S3-driven. It was traditionally serverless; it cost pennies a month to run.

And I’ve migrated it a few years ago to where it currently is, living on WordPress at WP Engine. Now, a lot of the technology purists will look at that and say that I went in the exact wrong direction. Why would I ever do that? Well, because things like integrating a podcast feed into the website are, grab a plugin; call it good. I can find a universe of people who are better at working with WordPress from a development perspective than I am as opposed to building something myself out of basically popsicle sticks and string and spending all of my time maintaining that.

Like, this website does not directly bring in any revenue to the business. It is ancillary. I need to have it in order to empower what the business does, but it is not a core competency of what we do, so outsourcing that to someone who does specialize in this makes an awful lot of business sense. And very often I’ll see people who are missing the point in cloud when it comes to losing sight of what the actual business objective is.

Donovan: Absolutely. Oftentimes, we hear the primary business goal is, “We want to decrease costs.” They hear for a number of reasons, “We can decrease costs.” “Oh, we can cut some of our operations teams because the cloud just magically runs.” You know, it doesn’t exactly work that way. But they’re missing sight of the actual business drivers that can help them grow the business, right?

What I like to say is that the cloud is not just a data center to host your infrastructure or host your applications; it’s actually a platform to grow your business. And because of these more streamlined and automated tools that AWS, Azure, Google, and these other hyperscaler clouds provide you, you can actually hit an immense amount of profitability by just leveraging these tools, but it requires you to transform your business. And that is what cloud adoption really is. Many people think that cloud adoption is a project, “Oh, I migrated to the cloud,” and then you leave it alone. Right?

But if you do that you’re subject to a lot of issues. Back to that Well-Architected Framework that we said, if you don’t have any guardrails in place to make sure that you have the proper security posture or the proper high availability or reliability concerns, right, it leads to cloud sprawl. Now, cloud sprawl is when an environment lacks the necessary guardrails and governance to limit the deployment of resources, right? This has a number of impacts, including a larger attack surface—that’s primarily security—there are tons of resources that can get deployed—there’s a lot of cost there—and then a lot of management overhead, you know? So, this means that cloud adoption is not a point-in-time project; it’s really a process; it’s a methodology; it’s an ongoing, continuous process, to make sure that you are cost-optimized, to make sure that you’re reliable, to make sure that you’re secure.

And if you can maintain an optimized environment that’s well-architected, you will inevitably grow your business because as we mentioned, it’s going to lead to better innovation, more agile development teams that can go to market quicker, increasing your overall revenue.

Corey: I think that people like to lose sight of the fact that in almost every case, unless you’re doing something absolutely bizarre, payroll is going to cost more than your cloud bill. Now, please don’t take that as a personal challenge if you’re listening to this. The goal is not to run up the score and see how high you can get just for funsies. Although I do admit, I play that game occasionally with, you know, accounts that are not my own. But in practice, for an easy example, on this, people like to turn up their nose at RDS sometimes because well, that charges a lot of money to run MySQL and I can run MySQL on top of EC2 instances.

Yes, you can, but what is your time worth comparatively? And at certain points of scale, when you’re making extraordinary demands on it, the economics do come back around where yeah, you probably should be running it on EC2 instead of as a managed service, but so much of that is going to be contextual. It’s basically impossible to look at an AWS estate—or any cloud estate for that matter—and unequivocally say, “Oh, this is a bad design. This is something that should not be because of X, Y, and Z,” because invariably, there’s context that you’re missing. I look at cloud accounts all the time where I could, in a vacuum, go ahead and optimize the living hell out of it, except there are reasons that things are the way they are, and the best way to look like a junior consultant is to show up and start throwing shade without understanding why things are the way they are.

Donovan: Exactly. And I think the number one issue that people run into, that organizations run into post-migration, is being able to track or tie back their migration and their optimization efforts to the business drivers. Logicworks actually ran an anonymous campaign, it was a survey. We hired some consultants, and some of the information that we found was absolutely shocking. They said that about 63% of organizations that have migrated to the cloud don’t actually understand or see the value of the cloud.

Now, again, if you’re listening to this, that doesn’t mean that there isn’t the value; it means that these organizations haven’t been able to track that, they haven’t been able to identify, okay, what are the agility metrics? What is the intended time to market? How long do we want to go with patches? Is this going to be a 24-hour timetable or is it going to be a week-long timetable? How many pushes are we making a day?

And then you can actually track that to the revenue that’s been generated? And then you can understand. But that’s a lot of work, and people are mostly concerned about the hard details of the technical. You know, “Okay, just put me in there. Let’s see how it goes. We got to get out of our data center for one reason or another,” but they missed the overall idea of this adoption that we’re referencing here.

Corey: On some level, it feels like the real business value of a cloud migration has little to do with the cloud itself and everything to do with let’s get out of the environment we’re currently in and into a new environment because that turns over a whole bunch of rocks and lets you finally sunset some things that really should have been turned off decades ago. I don’t know that’s too cynical of a take, but I can’t shake the feeling there’s some validity to it somewhere.

Donovan: [laugh]. Yeah, there definitely is some validity to that we’ve worked with a number of customers that we’ve migrated, and one perfect example. Last year, we were working with this disaster cleanup organization. They are one of the largest in the country. They have 1700 franchises and they needed to get out of their primary data center because there were a ton of issues.

There was actually a case where a squirrel had chewed through one of their power cables to the data center and shut off their air conditioning, so their data center was overheating. We were talking to them about the entire scope that they understood the migration to include, it had about 100 distinct applications and about maybe three to 400 virtual machines. Once we actually got into the assessment, there were over 700 virtual machines and 320 applications; distinct applications. There’s a mixture between custom off-the-shelf builds or homegrown applications that they’ve been making themselves, but they didn’t understand that they had over 60% of the IT estate that was just residing in that data center. Because sometimes it’s difficult to keep track of all of that.

That’s one of the things that cloud helps with. Cloud provides a ton of management and visibility and observability and traceability tools, but again, they need to be enabled, you know? And I think when companies are concerned about migrating and they’re just concerned about the re-hosting or the re-platforming of their applications, the management of this as an afterthought, and the actual, “How does this change our business,” is kind of a, “Oh man,” question mark in their head because that wasn’t something that was considered. And then that’s usually when we get the call.

Corey: It’s always fun when the bat phone goes off because people are generally not calling because, “Hey, things are great here. We just really wanted to boast about it some.” It turns into a, “Oh crap, we have a problem.” That honestly is one of my favorite parts of consulting is when you wind up being able to solve what feels like a monumental problem for someone, and sure, it’s relatively easy for you—presumably because when you do the same thing again and again and become a subject matter expert in it, it’s just a question of style more than how do I solve this intractable problem. But it’s nice to be able to walk away with a client saying, “That was awesome. I wish we could do it again.”

Donovan: Yeah.

Corey: Helping people is really what the game is about. I think that some folks tend to lose sight of that. And again, consulting doesn’t have the best reputation in the world, due to some of the larger shops, in some cases, doing what presents as, “Oh, I don’t know how to fix this, but I’m going to show up and prolong the problem forever.” I don’t see that that is true consulting in the traditional sense, but maybe I’m just playing games with words.

Donovan: No, I agree with that. I think one of my favorite parts of consulting is taking a step back and evaluating the entire landscape of what the problem is that we’re trying to solve, right? Oftentimes, as you are aware, somebody comes to you and they say, “This is the problem and this is what I need you to do to fix it.” Eight out of ten times, the solution that they’ve come up with probably isn’t the right solution. That will fix a symptom, right, but it’s not going to fix the actual cause or the biggest issue that’s plaguing them, right?

They have this continuous issue and they know that that’s going to cause them heartache and we need somebody to just do this, we don’t have time to do this. But if you peel back the onion and understand why you’re having that issue, that’s where I think a better consulting company would come in and help them discern.

Corey: There’s value in perspective. Very often, when you are the client organization, you’re too close to the problem, and/or you are extraordinarily familiar with your own context, but you’re missing some of the larger context in the greater ecosystem. I mean, one of the ways that I tend to look like a wizard from the future when I see something odd on their AWS bill and can just call it out, like, “Oh yeah, that’s this weird side effect of when you wind up having additional CloudTrail Management Events, yadda, yadda, yadda, yadda,” and they look at me like I’m a wizard from the future because I can just bust that out off the top of my head. Yeah, but the reason I can do that is because these things repeat themselves, these patterns continue to emerge, and the first time I saw that, it took me two weeks to get to the bottom of it. Now, when I see the symptoms of it, it’s oh, yeah, it’s that thing again.

And I feel like consulting is just a collection of stories like that again, and again, and again. And again, part of the trick is you don’t let the client ever see you sweat or do the research. You just, “All right, my next availability is in two weeks. I have a thing coming up.” And that thing, as it turns out, is spending two weeks of deep-dive research. But at least in my case I’d never charge by the hour, so it’s all about being mysterious and perceived as being good at these things, at least back in the early days. Then for my sins, I actually became good at it. Oops.

Donovan: [laugh]. Doesn’t that always seem to happen? It’s just like, “Oh, I didn’t expect to be an expert in this,” and then I don’t know where one day you’re like, “Oh, I guess I am the expert.”

Corey: Yeah. There’s also value in being an outside voice where you’re not beholden to individual stakeholders of, “Oh, yeah, we’re never allowed to wind up talking about that one system because someone went empire-building and they’re powerful here and you can’t ever talk about that thing in any way that doesn’t lead to more headcount and the rest.” Awesome. I don’t have the energy or time for those things, and to be direct, I’m not very good at it, as an employee. I can mind my manners for the duration of a relatively short duration consulting engagement just fine, but I’m also not necessarily there to look at things that are clearly suboptimal and say, “Oh, yeah, this is the way that it should be.”

But I’m also not the type to come in and say, “Well, why didn’t you build this in Lambda?” “Well, genius because then when this thing was built, Lambda didn’t exist, for one,” is a perfectly valid answer to that. And why would you go back and refactor something that’s already doing its job, unless your primary business objective is to bolster your own resume? Which I would suggest, it probably shouldn’t be?

Donovan: Yeah, exactly. And that’s where that outside perspective comes in because there are seven Rs of migration, right? Used to be six Rs; now there’s seven hours of migration, and you need to continuously reevaluate all of your application portfolio and see which of the seven Rs is most applicable to you, right? This is why that cloud adoption idea is an ongoing process. Maybe you’re beholden to some legacy applications that they really did serve their purpose when you were first migrating and there was nothing wrong with them, but now a few years later, it’s time to take a deeper look at all of your applications. Maybe we don’t need that application anymore. Maybe we can repurchase that. Maybe there’s a SaaS version of that, that we can go leverage now.

Or maybe it’s completely unrelated. You know, I was working with a prospect recently, who came to us—they are currently hosted in AWS—they came to us by way of Microsoft because we’re both an AWS and an Azure Partner, and Microsoft said, “Hey, they have some concerns. They wanted us to do a little assessment and see what’s going on.” We talked to them, and they said that they were looking to migrate away from AWS and into Azure because they wanted better pricing—

Corey: Oh, jeez.

Donovan: —and that is—exactly [laugh]. And that is the exact reason why you don’t migrate from one to the other because you’re setting yourself up for failure, right? It’s not about the cost; it’s not about the sticker price, the retail price. At the end of the day, all the cloud providers are going to be plus or minus a 2% difference. It’s, how did you architect this environment?

And when we started peeling back that onion, we started realizing they didn’t really have any guardrails or governance, and they started experiencing some of this cloud sprawl. And they expected that by transitioning cloud providers, they would have solved this magic wand solution, and now they’re just going to have better costs without putting any processes or frameworks in place to manage the environment to ensure that doesn’t happen again.

Corey: Let’s take it a step further. If you’re migrating to cloud to save money, you are probably not going to achieve any cost savings within five years at the soonest. That ignores as well the opportunity cost of all that energy that could be spent on other projects. If you’re moving to the cloud, it has to be based on a capability story, not because, “Oh, we’re going to save money by moving to the cloud,” in almost every case. I mean, the one time that actually did result in saving money was my own story, where instead of renting a rack in downtown Los Angeles—or part of a rack—for what was it, I think $300 a month or whatnot, suddenly, I wound up just spinning up a couple of Gmail accounts, and oh, this is costing me $10 a month instead, that actually did save some money. On a hobby project where my time was effectively free.

That is not usually the case for any functioning business. Don’t lose sight of the fact that as technologists, we tend to view our time is free, but to our employer, it’s more expensive than the AWS bill. It’s spend the money where it makes sense to spend the money.

Donovan: Exactly. And that’s a great tie-back to those business drivers, right? I think the cloud is a no-brainer for an organization who is looking to expand. “We’re right now only in these couple states and we want to be national,” or, “Now, we’re going to have a global presence, and it’s way easier to leverage the global footprint of the cloud than it is to build your own data centers or whatever your own solution would be.” Right?

But those are the primary business drivers that you’re looking to achieve. And I think, as a solutions architect, you know, solutions architects are really those consultants that take that higher-level approach and dig deeper to understand okay, but what are we trying to solve here? And this is how we can solve that, right? Not just, oh, I want to save some costs.

Corey: I’m taking a look at one of the projects that I’m working on right now and from—this is objectively the wrong direction along almost every axis—I am building something new that I’m not quite ready to talk about yet, but it is going to be revenue-bearing and thus production, and tied to a web app that I’m in the process of constructing. I am intentionally setting out from the beginning to break from my usual serverless pattern and build this to run on top of Kubernetes. I have been, therefore, learning Kubernetes for the last few weeks, and I have many thoughts on it, few of them flattering. And I’m looking at this going, “This seems over-engineered,” and for my use case, it certainly is. However, the reason I’m doing this is because every client I’m dealing with these days runs some Kubernetes stuff themselves, and I should understand it better. It gives me a production workload that I can use for demos for a variety of different things. The staging environment will contain no personally identifiable information, so I can deploy that anywhere that there’s a Kubernetes-esque environment or control plane as a dummy workload that I can use to kick the tires on this.

And for those reasons, it makes an awful lot of sense. In practice, doing something serverlessly would be the better option. Or the best answer would be to find someone who provides this sort of thing as a white label-able service that I can just pay a few hundred bucks a month to and not think about this thing again. But those are the constraints that make what otherwise looks like a ridiculous idea make sense in this context. But oh God, looking at this for people who run Kubernetes just so they can host their blog, it’s, what are you doing over there? Is that just sort of the Hello World-style application? Because not for nothing, the value of a blog is in the content, not in the magic that winds up making that content visible to readers in almost every case.

Donovan: Yeah, [laugh] this is one of my favorite topics to talk about, actually, because so many times, we engage these customers, and they’re like, “We want to containerize.” And I’m like, “Oh, that’s awesome. I love the idea. There’s so many benefits to containerization that pretty much checks every box within the Well-Architected Framework.” And then their next sentence is, “And so, we’re going to move to Kubernetes.”

And I’m like, “Well, why Kubernetes?” Right? Because as you said, this is just a blog, or this is just a static website, or whatever the application they have is. Containers are great, but do you need a full-fledged Kubernetes deployment? Do you need every open-source software possible so that you can integrate?

Or would you be just fine with ECS that is pretty streamlined, pretty much out of the box just works from AWS? Sometimes it’s necessary. Sometimes EKS or a Kubernetes deployment or maybe your own self-managed Kubernetes deployment is necessary, but it’s another one of these misconceptions, I’d say, when people hear the new buzzwords, they’re like, “Oh, we need to do that because that’s the best thing to do.” And it’s often best as you—I loved your word ‘perspective,’ the outside perspective—it’s usually best to get that outside perspective and understand really how—what specific solution might help you.

Corey: From where I sit, I think that people often tend to skip over that. It gets to the idea of resume-driven development. And honestly, it’s hard to tell people they’re necessarily going in the wrong direction given how fever-pitched the hype around Kubernetes has gotten. Every company is using it for something, and on some level, wanting to get that on your resume is a logical next step. Looking at cloud bills, I would never suggest someone [laugh] wind up doing this on their own dime, so yeah, it does make sense in that context.

The goal, I think of being a technology executive is to be able to understand that and, one, provide pathways for your team to develop, but also to make sure that their objectives align with the business’s objectives. And often I do not see that leading in the Kubernetes direction, for better or worse. There’s a reason that I own the domain kubernetestheeasyway.com and I repoint it to Amazon’s ECS.

Although to those listening, I am thrilled to repoint that to the highest bidder; my email is open. Jokes and witticisms aside, I am curious based upon your perspective in the market—which is broader than mine because imagine that, you don’t focus on one very specific problem—what are most organizations looking at for business drivers that they generally tend to either not realize or not quite achieve, or, “Well, that didn’t go quite the way that we wanted to?” Because we see the outcome of people moving to cloud; we don’t necessarily see the reasons behind it.

Donovan: Yeah, that’s a great point. So, the first and foremost concern that I’d say most companies experience is how do we make sure that our website or application is always up and always able to generate revenue, right? Based on the types of customers that we get, they’re hyper-compliant, they’re mission critical, they’re 24x7, they need to always be on, they need to make sure that whatever the platform that they’re running their infrastructure on is able to be up 24x7. Much harder to do that in your own world than it is to do that in cloud, right, so that’s one large business driver.

Another large business driver, again, with hyper-compliance, is security. They’ve experienced some security incidents and they want to make sure that they have a more scalable, secure platform as they grow because, again, maybe they’re becoming national, or they’re becoming global, or they need to meet some compliance regulation, right? And then additionally, as I’ve mentioned a couple times here, they want to increase their overall agility, which will in turn decrease their time to market their time to revenue. That is an often miscalculated correlation, right? People say, “Okay, we’ll increase our agility,” but without realizing why they’re going to try to increase their agility.

They want to move from pushing code 20 times a month to 20 times a day, but why, right? That’s the agility piece, but why do you want to do that? Well, because it allows us to more quickly debug our code. It allows us to more quickly identify issues or rollbacks and improve our overall efficiency, increase customer retention, increase customer satisfaction, things like that.

Corey: I really wish that people would, I guess, tell those stories more in conference talks, rather than, “We moved to the cloud because of reasons, and it was awesome.” And it’s always depressing to me because you hear them tell this beautiful story, and you turn next to the person in the audience often and be like, “Oh, I wish I could work in a place that ran a project like that.” And they say, “Yeah, me too.” And you check their badge and they work at the same company the speaker does. It’s the idea of telling this fanciful, imagined versioning rather than addressing the reality that the real world is very messy.

Donovan: Exactly. And it’s really hard, you know? I’m not going to sit here and say, “Well, you know, everything that we’re talking about here and completely transforming your business is simple. And wow, you guys aren’t smart for doing it or for not doing it.” You know? It is difficult.

But back to that idea of the outside perspective, there are tried and true methods, right? Like we talked about with the Well-Architected Framework, with the Cloud Adoption Methodology, with the Seven Rs of Migration, right? There’s a lot of content out there are a lot of people out there that can help you, but it’s definitely a journey.

Corey: It really is. And I want to thank you for sharing your view of it with us on the show. If people want to learn more, where’s the best place to find you?

Donovan: Visit our website, logicworks.com. You could visit us across all of our social media platforms. You could reach out to me directly; happy to talk to anybody, even if you just wanted to say, “Hey, how’s the weather?” Right now, it’s raining. But yeah, definitely reach out to us via our website.

Corey: And we will, of course, put a link to that in the show notes. Thank you so much for being so generous with your time. I appreciate it.

Donovan: Thank you so much, Corey. I had a great time, and looking forward to the next one.

Corey: Donovan Brady, director of solutions architecture at Logicworks on this promoted guest episode. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry rant as a negative comment on this, presumably because you are Donovan’s antithesis, the director of problems architecture.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Amy

Amy Tobey has worked in tech for more than 20 years at companies of every size, working with everything from kernel code to user interfaces. These days she spends her time building an innovative Site Reliability Engineering program at Equinix, where she is a principal engineer. When she's not working, she can be found with her nose in a book, watching anime with her son, making noise with electronics, or doing yoga poses in the sun.

Links Referenced:

  • Equinix: https://metal.equinix.com
  • Twitter: https://twitter.com/MissAmyTobey

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud, I’m Corey Quinn, and this episode is another one of those real profiles in shitposting type of episodes. I am joined again from a few months ago by Amy Tobey, who is a Senior Principal Engineer at Equinix, back for more. Amy, thank you so much for joining me.

Amy: Welcome. To your show. [laugh].

Corey: Exactly. So, one thing that we have been seeing a lot over the past year, and you struck me as one of the best people to talk about what you’re seeing in the wilderness perspective, has been the idea of cloud repatriation. It started off with something that came out of Andreessen Horowitz toward the start of the year about the trillion-dollar paradox, how, at a certain point of scale, repatriating to a data center is the smart and right move. And oh, my stars that ruffle some feathers for people?

Amy: Well, I spent all this money moving to the cloud. That was just mean.

Corey: I know. Why would I want to leave the cloud? I mean, for God’s sake, my account manager named his kid after me. Wait a minute, how much am I spending on that? Yeah—

Amy: Good question.

Corey: —there is that ever-growing problem. And there have been the examples that people have given of Dropbox classically did a cloud repatriation exercise, and a second example that no one can ever name. And it seems like okay, this might not necessarily be the direction that the industry is going. But I also tend to not be completely naive when it comes to these things. And I can see repatriation making sense on a workload-by-workload basis.

What that implies is that yeah, but a lot of other workloads are not going to be going to a data center. They’re going to stay in a cloud provider, who would like very much if you never read a word of this to anyone in public.

Amy: Absolutely, yeah.

Corey: So, if there are workloads repatriating, it would occur to me that there’s a vested interest on the part of every major cloud provider to do their best to, I don’t know if saying suppress the story is too strongly worded, but it is directionally what I mean.

Amy: They aren’t helping get the story out. [laugh].

Corey: Yeah, it’s like, “That’s a great observation. Could you maybe shut the hell up and never make it ever again in public, or we will end you?” Yeah. Your Amazon. What are you going to do, launch a shitty Amazon Basics version of what my company does? Good luck. Have fun. You’re probably doing it already.

But the reason I want to talk to you on this is a confluence of a few things. One, as I mentioned back in May when you were on the show, I am incensed and annoyed that we’ve been talking for as long as we have, and somehow I never had you on the show. So, great. Come back, please. You’re always welcome here. Secondly, you work at Equinix, which is, effectively—let’s be relatively direct—it is functionally a data center as far as how people wind up contextualizing this. Yes, you have higher level—

Amy: Yeah I guess people contextualize it that way. But we’ll get into that.

Corey: Yeah, from the outside. I don’t work there, to be clear. My talking points don’t exist for this. But I think of oh, Equinix. Oh, that means you basically have a colo or colo equivalent. The pricing dynamics have radically different; it looks a lot closer to a data center in my imagination than it does a traditional public cloud. I would also argue that if someone migrates from AWS to Equinix, that would be viewed—arguably correctly—as something of a repatriation. Is that directionally correct?

Amy: I would argue incorrectly. For Metal, right?

Corey: Ah.

Amy: So, Equinix is a data center company, right? Like that’s why everybody knows us as. Equinix Metal is a bare metal primitive service, right? So, it’s a lot more of a cloud workflow, right, except that you’re not getting the rich services that you get in a technically full cloud, right? Like, there’s no RDS; there’s no S3, even. What you get is bare metal primitives, right? With a really fast network that isn’t going to—

Corey: Are you really a cloud provider without some ridiculous machine-learning-powered service that’s going to wind up taking pictures, perform incredibly expensive operations on it, and then return something that’s more than a little racist? I mean, come on. That’s not—you’re not a cloud until you can do that, right?

Amy: We can do that. We have customers that do that. Well, not specifically that, but um—

Corey: Yeah, but they have to build it themselves. You don’t have the high-level managed service that basically serves as, functionally, bias laundering.

Amy: Yeah, you don’t get it in a box, right? So, a lot of our customers are doing things that are unique, right, that are maybe not exactly fit into the cloud well. And it comes back down to a lot of Equinix’s roots, which is—we talk but going into the cloud, and it’s this kind of abstract environment we’re reaching for, you know, up in the sky. And it’s like, we don’t know where it is, except we have regions that—okay, so it’s in Virginia. But the rule of real estate applies to technology as often as not, which is location, location, location, right?

When we’re talking about a lot of applications, a challenge that we face, say in gaming, is that the latency from the customer, so that last mile to your data center, can often be extremely important, right, so a few milliseconds even. And a lot of, like, SaaS applications, the typical stuff that really the cloud was built on, 10 milliseconds, 50 milliseconds, nobody’s really going to notice that, right? But in a gaming environment or some very low latency application that needs to run extremely close to the customer, it’s hard to do that in the cloud. They’re building this stuff out, right? Like, I see, you know, different ones [unintelligible 00:05:53] opening new regions but, you know, there’s this other side of the cloud, which is, like, the edge computing thing that’s coming alive, and that’s more where I think about it.

And again, location, location, location. The speed of light is really fast, but as most of us in tech know, if you want to go across from the East Coast to the West Coast, you’re talking about 80 milliseconds, on average, right? I think that’s what it is. I haven’t checked in a while. Yeah, that’s just basic fundamental speed of light. And so, if everything’s in us-east-1—and this is why we do multi-region, sometimes—the latency from the West Coast isn’t going to be great. And so, we run the application on both sides.

Corey: It has improved though. If you want to talk old school things that are seared into my brain from over 20 years ago, every person who’s worked in data centers—or in technology, as a general rule—has a few IP addresses seared. And the one that I’ve always had on my mind was 130.111.32.11. Kind of arbitrary and ridiculous, but it was one of the two recursive resolvers provided at the University of Maine where I had my first help desk job.

And it lives on-prem, in Maine. And generally speaking, I tended to always accept that no matter where I was—unless I was in a data center somewhere—it was about 120 milliseconds. And I just checked now; it is 85 and change from where I am in San Francisco. So, the internet or the speed of light have improved. So, good for whichever one of those it was. But yeah, you’ve just updated my understanding of these things. All of this is, which is to say, yes, latency is very important.

Amy: Right. Let’s forget repatriation to really be really honest. Even the Dropbox case or any of them, right? Like, there’s an economic story here that I think all of us that have been doing cloud work for a while see pretty clearly that maybe not everybody’s seeing that—that’s thinking from an on-prem kind of situation, which is that—you know, and I know you do this all the time, right, is, you don’t just look at the cost of the data center and the servers and the network, the technical components, the bill of materials—

Corey: Oh, lies, damned lies, and TCO analyses. Yeah.

Amy: —but there’s all these people on top of it, and the organizational complexity, and the contracts that you got to manage. And it’s this big, huge operation that is incredibly complex to do well that is almost nobody’s business.

So the way I look at this, right, and the way I even talk to customers about it is, like, “What is your produ—” And I talk to people internally about this way? It’s like, “What are you trying to build?” “Well, I want to build a SaaS.” “Okay. Do you need data center expertise to build a SaaS?” “No.” “Then why the hell are you putting it in a data center?” Like we—you know, and speaking for my employer, right, like, we have Equinix Metal right here. You can build on that and you don’t have to do all the most complex part of this, at least in terms of, like, the physical plant, right? Like, right, getting a bare metal server available, we take care of all of that. Even at the primitive level, where we sit, it’s higher level than, say, colo.

Corey: There’s also the question of economics as it ties into it. It’s never just a raw cost-of-materials type of approach. Like, my original job in a data center was basically to walk around and replace hard drives, and apparently, to insult people. Now, the cloud has taken one of those two aspects away, and you can follow my Twitter account and figure out which one of those two it is, but what I keep seeing now is there is value to having that task done, but in a cloud environment—and Equinix Metal, let’s be clear—that has slipped below the surface level of awareness. And well, what are the economic implications of that?

Well, okay, you have a whole team of people at large companies whose job it is to do precisely that. Okay, we’re going to upskill them and train them to use cloud. Okay. First, not everyone is going to be capable or willing to make that leap from hard drive replacement to, “Congratulations and welcome to JavaScript. You’re about to hate everything that comes next.”

And if they do make that leap, their baseline market value—by which I mean what the market is willing to pay for them—approximately will double. And whether they wind up being paid more by their current employer or they take a job somewhere else with those skills and get paid what they are worth, the company still has that economic problem. Like it or not, you will generally get what you pay for whether you want to or not; that is the reality of it. And as companies are thinking about this, well, what gets into the TCO analysis and what doesn’t, I have yet to see one where the outcome was not predetermined. They’re less, let’s figure out in good faith whether it’s going to be more expensive to move to the cloud, or move out of the cloud, or just burn the building down for insurance money. The outcome is generally the one that the person who commissioned the TCO analysis wants. So, when a vendor is trying to get you to switch to them, and they do one for you, yeah. And I’m not saying they’re lying, but there’s so much judgment that goes into this. And what do you include and what do you not include? That’s hard.

Amy: And there’s so many hidden costs. And that’s one of the things that I love about working at a cloud provider is that I still get to play with all that stuff, and like, I get to see those hidden costs, right? Like you were talking about the person who goes around and swaps out the hard drives. Or early in my career, right, I worked with someone whose job it was this every day, she would go into data center, she’d swap out the tapes, you know, and do a few things other around and, like, take care of the billing system. And that was a job where it was kind of going around and stewarding a whole bunch of things that kind of kept the whole machine running, but most people outside of being right next to the data center didn’t have any idea that stuff even happen, right, that went into it.

And so, like you were saying, like, when you go to do the TCO analysis, I mean, I’ve been through this a couple of times prior in my career, where people will look at it and go like, “Well, of course we’re not going to list—we’ll put, like, two headcount on there.” And it’s always a lie because it’s never just to headcount. It’s never just the network person, or the SRE, or the person who’s racking the servers. It’s also, like, finance has to do all this extra work, and there’s all the logistic work, and there is just so much stuff that just is really hard to include. Not only do people leave it out, but it’s also just really hard for people to grapple with the complexity of all the things it takes to run a data center, which is, like, one of the most complex machines on the planet, any single data center.

Corey: I’ve worked in small-scale environments, maybe a couple of mid-sized ones, but never the type of hyperscale facility that you folks have, which I would say is if it’s not hyperscale, it’s at least directionally close to it. We’re talking thousands of servers, and hundreds of racks.

Amy: Right.

Corey: I’ve started getting into that, on some level. Now, I guess when we say ‘hyperscale,’ we’re talking about AWS-size things where, oh, that’s a region and it’s going to have three dozen data center facilities in it. Yeah, I don’t work in places like that because honestly, have you met me? Would you trust me around something that’s that critical infrastructure? No, you would not, unless you have terrible judgment, which means you should not be working in those environments to begin with.

Amy: I mean, you’re like a walking chaos exercise. Maybe I would let you in.

Corey: Oh, I bring my hardware destruction aura near anything expensive and things are terrible. It’s awful. But as I looked at the cloud, regardless of cloud, there is another economic element that I think is underappreciated, and to be fair, this does, I believe, apply as much to Equinix Metal as it does to the public hyperscale cloud providers that have problems with naming things well. And that is, when you are provisioning something as a customer of one of these places, you have an unbounded growth problem. When you’re in a data center, you are not going to just absentmindedly sign an $8 million purchase order for new servers—you know, a second time—and then that means you’re eventually run out of power, space, places to put things, and you have to go find it somewhere.

Whereas in cloud, the only limit is basically your budget where there is no forcing function that reminds you to go and clean up that experiment from five years ago. You have people with three petabytes of data they were using for a project, but they haven’t worked there in five years and nothing’s touched it since. Because the failure mode of deleting things that are important, or disasters—

Amy: That’s why Glacier exists.

Corey: Oh, exactly. But that failure mode of deleting things that should not be deleted are disastrous for a company, whereas if you’ve leave them there, well, it’s only money. And there’s no forcing function to do that, which means you have this infinite growth problem with no natural limit slash predator around it. And that is the economic analysis that I do not see playing out basically anywhere. Because oh, by the time that becomes a problem, we’ll have good governance in place. Yeah, pull the other one. It has bells on it.

Amy: That’s the funny thing, right, is a lot of the early drive in the cloud was those of us who wanted to go faster and we were up against the limitations of our data centers. And then we go out and go, like, “Hey, we got this cloud thing. I’ll just, you know, put the credit card in there and I’ll spin up a few instances, and ‘hey, I delivered your product.’” And everybody goes, “Yeah, hey, happy.” And then like you mentioned, right, and then we get down the road here, and it’s like, “Oh, my God, how much are we spending on this?”

And then you’re in that funny boat where you have both. But yeah, I mean, like, that’s just typical engineering problem, where, you know, we have to deal with our constraints. And the cloud has constraints, right? Like when I was at Netflix, one of the things we would do frequently is bump up against instance limits. And then we go talk to our TAM and be like, “Hey, buddy. Can we have some more instance limit?” And then take care of that, right?

But there are some bounds on that. Of course, in the cloud providers—you know, if I have my cloud provider shoes on, I don’t necessarily want to put those limits to law because it’s a business, the business wants to hoover up all the money. That’s what businesses do. So, I guess it’s just a different constraint that is maybe much too easy to knock down, right? Because as you mentioned, in a data center or in a colo space, I outgrow my cage and I filled up all that space I have, I have to either order more space from my colo provider, I expand to the cloud, right?

Corey: The scale I was always at, the limit was not the space because I assure you with enough shoving all things are possible. Don’t believe me? Look at what people are putting in the overhead bin on any airline. Enough shoving, you’ll get a Volkswagen in there. But it was always power constrained is what I dealt with it. And it’s like, “Eh, they’re just being conservative.” And the whole building room dies.

Amy: You want blade servers because that’s how you get blade servers, right? That movement was about bringing the density up and putting more servers in a rack. You know, there were some management stuff and [unintelligible 00:16:08], but a lot of it was just about, like, you know, I remember I’m picturing it, right—

Corey: Even without that, I was still power constrained because you have to remember, a lot of my experiences were not in, shall we say, data center facilities that you would call, you know, good.

Amy: Well, that brings up a fun thing that’s happening, which is that the power envelope of servers is still growing. The newest Intel chips, especially the ones they’re shipping for hyperscale and stuff like that, with the really high core counts, and the faster clock speeds, you know, these things are pulling, like, 300 watts. And they also have to egress all that heat. And so, that’s one of the places where we’re doing some innovations—I think there’s a couple of blog posts out about it around—like, liquid cooling or multimode cooling. And what’s interesting about this from a cloud or data center perspective, is that the tools and skills and everything has to come together to run a, you know, this year’s or next year’s servers, where we’re pushing thousands of kilowatts into a rack. Thousands; one rack right?

The bar to actually bootstrap and run this stuff successfully is rising again, compared to I take my pizza box servers, right—and I worked at a gaming company a long time ago, right, and they would just, like, stack them on the floor. It was just a stack of servers. Like, they were in between the rails, but they weren’t screwed down or anything, right? And they would network them all up. Because basically, like, the game would spin up on the servers and if they died, they would just unplug that one and leave it there and spin up another one.

It was like you could just stack stuff up and, like, be slinging cables across the data center and stuff back then. I wouldn’t do it that way now, but when you add, say liquid cooling and some of these, like, extremely high power situations into the mix, now you need to have, for example, if you’re using liquid cooling, you don’t want that stuff leaking, right? And so, it’s good as the pressure fittings and blind mating and all this stuff that’s coming around gets, you still have that element of additional training, and skill, and possibility for mistakes.

Corey: The thing that I see as I look at this across the space is that, on some level, it’s gotten harder to run a data center than it ever did before. Because again, another reason I wanted to have you on this show is that you do not carry a quota. Although you do often carry the conversation, when you have boring people around you, but quotas, no. You are not here selling things to people. You’re not actively incentivized to get people to see things a certain way.

You are very clearly an engineer in the right ways. I will further point out though, that you do not sound like an engineer, by which I mean, you’re going to basically belittle people, in many cases, in the name of being technically correct. You’re a human being with a frickin soul. And believe me, it is noticed.

Amy: I really appreciate that. If somebody’s just listening to hearing my voice and in my name, right, like, I have a low voice. And in most of my career, I was extremely technical, like, to the point where you know, if something was wrong technically, I would fight to the death to get the right technical solution and maybe not see the complexity around the decisions, and why things were the way they were in the way I can today. And that’s changed how I sound. It’s changed how I talk. It’s changed how I look at and talk about technology as well, right? I’m just not that interested in Kubernetes. Because I’ve kind of started looking up the stack in this kind of pursuit.

Corey: Yeah, when I say you don’t sound like an engineer, I am in no way shape or form—

Amy: I know.

Corey: —alluding in any respect to your technical acumen. I feel the need to clarify that statement for people who might be listening, and say, “Hey, wait a minute. Is he being a shithead?” No.

Amy: No, no, no.

Corey: Well, not the kind you’re worried I’m being anyway; I’m a different breed of shithead and that’s fine.

Amy: Yeah, I should remember that other people don’t know we’ve had conversations that are deeply technical, that aren't on air, that aren’t context anybody else has. And so, like, I bring that deep technical knowledge, you know, the ability to talk about PCI Express, and kilovolts [unintelligible 00:19:58] rack, and top-of-rack switches, and network topologies, all of that together now, but what’s really fascinating is where the really big impact is, for reliability, for security, for quality, the things that me as a person, that I’m driven by—products are cool, but, like, I like them to be reliable; that’s the part that I like—really come down to more leadership, and business acumen, and understanding the business constraints, and then being able to get heard by an audience that isn’t necessarily technical, that doesn’t necessarily understand the difference between PCI, PCI-X, and PCI Express. There’s a difference between those. It doesn’t mean anything to the business, right, so when we want to go and talk about why are we doing, for example, multi-region deployment of our application? If I come in and say, “Well, because we want to use Raft.” That’s going to fall flat, right?

The business is going to go, “I don’t care about Raft. What does that have to do with my customers?” Which is the right question to always ask. Instead, when I show up and say, “Okay, what’s going on here is we have this application sits in a single region—or in a single data center or whatever, right? I’m using region because that’s probably what most of the people listening understand—you know, so I put my application in a single region and it goes down, our customers are going to be unhappy. We have the alternative to spend, okay, not a little bit more money, probably a lot more money to build a second region, and the benefit we will get is that our customers will be able to access the service 24x7, and it will always work and they’ll have a wonderful experience. And maybe they’ll keep coming back and buy more stuff from us.”

And so, when I talk about it in those terms, right—and it’s usually more nuanced than that—then I start to get the movement at the macro level, right, in the systemic level of the business in the direction I want it to go, which is for the product group to understand why reliability matters to the customer, you know? For the individual engineers to understand why it matters that we use secure coding practices.

[midroll 00:21:56]

Corey: Getting back to the reason I said that you are not quota-carrying and you are not incentivized to push things in a particular way is that often we’ll meet zealots, and I’ve never known you to be one, you have always been a strong advocate for doing the right thing, even if it doesn’t directly benefit any given random employer that you might have. And as a result, one of the things that you’ve said to me repeatedly is if you’re building something from scratch, for God’s sake, put it in cloud. What is wrong with you? Do that. The idea of building it yourself on low-lying, underlying primitives for almost every modern SaaS style workload, there’s no reason to consider doing something else in almost any case. Is that a fair representation of your position on this?

Amy: It is. I mean, the simpler version right, “Is why the hell are you doing undifferentiated lifting?” Right? Things that don’t differentiate your product, why would you do it?

Corey: The thing that this has empowered then is I can build an experiment tonight—I don’t have to wait for provisioning and signed contracts and do all the rest. I can spend 25 cents and get the experiment up and running. If it takes off, though, it has changed how I move going forward as well because there’s no difference in the way that there was back when we were in data centers. I’m going to try and experiment I’m going to run it in this, I don’t know, crappy Raspberry Pi or my desktop or something under my desk somewhere. And if it takes off and I have to scale up, I got to do a giant migration to real enterprise-grade hardware. With cloud, you are getting all of that out of the box, even if all you’re doing with it is something ridiculous and nonsensical.

Amy: And you’re often getting, like, ridiculously better service. So, 20 years ago, if you and I sat down to build a SaaS app, we would have spun up a Linux box somewhere in a colo, and we would have spun up Apache, MySQL, maybe some Perl or PHP if we were feeling frisky. And the availability of that would be one machine could do, what we could handle in terms of one MySQL instance. But today if I’m spinning up a new stack for some the same kind of SaaS, I’m going to probably deploy it into an ASG, I’m probably going to have some kind of high availability database be on it—and I’m going to use Aurora as an example—because, like, the availability of an Aurora instance, in terms of, like, if I’m building myself up with even the very best kit available in databases, it’s going to be really hard to hit the same availability that Aurora does because Aurora is not just a software solution, it’s also got a team around it that stewards that 24/7. And it continues to evolve on its own.

And so, like, the base, when we start that little tiny startup, instead of being that one machine, we’re actually starting at a much higher level of quality, and availability, and even security sometimes because of these primitives that were available. And I probably should go on to extend on the thought of undifferentiated lifting, right, and coming back to the colo or the edge story, which is that there are still some little edge cases, right? Like I think for SaaS, duh right? Like, go straight to. But there are still some really interesting things where there’s, like, hardware innovations where they’re doing things with GPUs and stuff like that.

Where the colo experience may be better because you’re trying to do, like, custom hardware, in which case you are in a colo. There are businesses doing some really interesting stuff with custom hardware that’s behind an application stack. What’s really cool about some of that, from my perspective, is that some of that might be sitting on, say, bare metal with us, and maybe the front-end is sitting somewhere else. Because the other thing Equinix does really well is this product we call a Fabric which lets us basically do peering with any of the cloud providers.

Corey: Yeah, the reason, I guess I don’t consider you as a quote-unquote, “Cloud,” is first and foremost, rooted in the fact that you don’t have a bandwidth model that is free and grass and criminally expensive to send it anywhere that isn’t to you folks. Like, are you really a cloud if you’re not just gouging the living piss out of your customers every time they want to send data somewhere else?

Amy: Well, I mean, we like to say we’re part of the cloud. And really, that’s actually my favorite feature of Metal is that you get, I think—

Corey: Yeah, this was a compliment, to be very clear. I’m a big fan of not paying 1998 bandwidth pricing anymore.

Amy: Yeah, but this is the part where I get to do a little bit of, like, showing off for Metal a little bit, in that, like, when you buy a Metal server, there’s different configurations, right, but, like, I think the lowest one, you have dual 10 Gig ports to the server that you can get either in a bonded mode so that you have a single 20 Gig interface in your operating system, or you can actually do L3 and you can do BGP to your server. And so, this is a capability that you really can’t get at all on the other clouds, right? This lets you do things with the network, not only the bandwidth, right, that you have available. Like, you want to stream out 25 gigs of bandwidth out of us, I think that’s pretty doable. And the rates—I’ve only seen a couple of comparisons—are pretty good.

So, this is like where some of the business opportunities, right—and I can’t get too much into it, but, like, this is all public stuff I’ve talked about so far—which is, that’s part of the opportunity there is sitting at the crossroads of the internet, we can give you a server that has really great networking, and you can do all the cool custom stuff with it, like, BGP, right? Like, so that you can do Anycast, right? You can build Anycast applications.

Corey: I miss the days when that was a thing that made sense.

Amy: [laugh].

Corey: I mean that in the context of, you know, with the internet and networks. These days, it always feels like the network engineering as slipped away within the cloud because you have overlays on top of overlays and it’s all abstractions that are living out there right until suddenly you really need to know what’s going on. But it has abstracted so much of this away. And that, on some level, is the surprise people are often in for when they wind up outgrowing the cloud for a workload and wanting to move it someplace that doesn’t, you know, ride them like naughty ponies for bandwidth. And they have to rediscover things that we’ve mostly forgotten about.

I remember having to architect significantly around the context of hard drive failures. I know we’ve talked about that a fair bit as a thing, but yeah, it’s spinning metal, it throws off heat and if you lose the wrong one, your data is gone and you now have serious business problems. In cloud, at least AWS-land, that’s not really a thing anymore. The way EBS is provisioned, there’s a slight tick in latency if you’re looking at just the right time for what I think is a hard drive failure, but it’s there. You don’t have to think about this anymore.

Migrate that workload to a pile of servers in a colo somewhere, guess what? Suddenly your reliability is going to decrease. Amazon, and the other cloud providers as well, have gotten to a point where they are better at operations than you are at your relatively small company with your nascent sysadmin team. I promise. There is an economy of scale here.

Amy: And it doesn’t have to be good or better, right? It’s just simply better resourced—

Corey: Yeah.

Amy: Than most anybody else can hope. Amazon can throw a billion dollars at it and never miss it. In most organizations out there, you know, and most of the especially enterprise, people are scratching and trying to get resources wherever they can, right? They’re all competing for people, for time, for engineering resources, and that’s one of the things that gets freed up when you just basically bang an API and you get the thing you want. You don’t have to go through that kind of old world internal process that is usually slow and often painful.

Just because they’re not resourced as well; they’re not automated as well. Maybe they could be. I’m sure most of them could, in theory be, but we come back to undifferentiated lifting. None of this helps, say—let me think of another random business—Claire’s, whatever, like, any of the shops in the mall, they all have some kind of enterprise behind them for cash processing and all that stuff, point of sale, none of this stuff is differentiating for them because it doesn’t impact anything to do with where the money comes in. So again, we’re back at why are you doing this?

Corey: I think that’s also the big challenge as well, when people start talking about repatriation and talking about this idea that they are going to, oh, that cloud is too expensive; we’re going to move out. And they make the economics work. Again, I do firmly believe that, by and large, businesses do not intentionally go out and make poor decisions. I think when we see a company doing something inscrutable, there’s always context that we’re missing, and I think as a general rule of thumb, that at these companies do not hire people who are fools. And there are always constraints that they cannot talk about in public.

My general position as a consultant, and ideally as someone who aspires to be a decent human being, is that when I see something I don’t understand, I assume that there’s simply a lack of context, not that everyone involved in this has been foolish enough to make giant blunders that I can pick out in the first five seconds of looking at it. I’m not quite that self-confident yet.

Amy: I mean, that’s a big part of, like, the career progression into above senior engineer, right, is, you don’t get to sit in your chair and go, like, “Oh, those dummies,” right? You actually have—I don’t know about ‘have to,’ but, like, the way I operate now, right, is I remember in my youth, I used to be like, “Oh, those business people. They don’t know, nothing. Like, what are they doing?” You know, it’s goofy what they’re doing.

And then now I have a different mode, which is, “Oh, that’s interesting. Can you tell me more?” The feeling is still there, right? Like, “Oh, my God, what is going on here?” But then I get curious, and I go, “So, how did we get here?” [laugh]. And you get that story, and the stories are always fascinating, and they always involve, like, constraints, immovable objects, people doing the best they can with what they have available.

Corey: Always. And I want to be clear that very rarely is it the right answer to walk into a room and say, look at the architecture and, “All right, what moron built this?” Because always you’re going to be asking that question to said moron. And it doesn’t matter how right you are, they’re never going to listen to another thing out of your mouth again. And have some respect for what came before even if it’s potentially wrong answer, well, great. “Why didn’t you just use this service to do this instead?” “Yeah, because this thing predates that by five years, jackass.”

There are reasons things are the way they are, if you take any architecture in the world and tell people to rebuild it greenfield, almost none of them would look the same as they do today because we learn things by getting it wrong. That’s a great teacher, and it hurts. But it’s also true.

Amy: And we got to build, right? Like, that’s what we’re here to do. If we just kind of cycle waiting for the perfect technology, the right choices—and again, to come back to the people who built it at the time used—you know, often we can fault people for this—used the things they know or the things that are nearby, and they make it work. And that’s kind of amazing sometimes, right?

Like, I’m sure you see architectures frequently, and I see them too, probably less frequently, where you just go, how does this even work in the first place? Like how did you get this to work? Because I’m looking at this diagram or whatever, and I don’t understand how this works. Maybe that’s a thing that’s more a me thing, like, because usually, I can look at a—skim over an architecture document and be, like, be able to build the model up into, like, “Okay, I can see how that kind of works and how the data flows through it.” I get that pretty quickly.

And comes back to that, like, just, again, asking, “How did we get here?” And then the cool part about asking how did we get here is it sets everybody up in the room, not just you as the person trying to drive change, but the people you’re trying to bring along, the original architects, original engineers, when you ask, how did we get here, you’ve started them on the path to coming along with you in the future, which is kind of cool. But until—that storytelling mode, again, is so powerful at almost every level of the stack, right? And that’s why I just, like, when we were talking about how technical I bring things in, again, like, I’m just not that interested in, like, are you Little Endian or Big Endian? How did we get here is kind of cool. You built a Big Endian architecture in 2022? Like, “Ohh. [laugh]. How do we do that?”

Corey: Hey, leave me to my own devices, and I need to build something super quickly to get it up and running, well, what I’m going to do, for a lot of answers is going to look an awful lot like the traditional three-tier architecture that I was running back in 2008. Because I know it, it works well, and I can iterate rapidly on it. Is it a best practice? Absolutely not, but given the constraints, sometimes it’s the fastest thing to grab? “Well, if you built this in serverless technologies, it would run at a fraction of the cost.” It’s, “Yes, but if I run this thing, the way that I’m running it now, it’ll be $20 a month, it’ll take me two hours instead of 20. And what exactly is your time worth, again?” It comes down to the better economic model of all these things.

Amy: Any time you’re trying to make a case to the business, the economic model is going to always go further. Just general tip for tech people, right? Like if you can make the better economic case and you go to the business with an economic case that is clear. Businesses listen to that. They’re not going to listen to us go on and on about distributed systems.

Somebody in finance trying to make a decision about, like, do we go and spend a million bucks on this, that’s not really the material thing. It’s like, well, how is this going to move the business forward? And how much is it going to cost us to do it? And what other opportunities are we giving up to do that?

Corey: I think that’s probably a good place to leave it because there’s no good answer. We can all think about that until the next episode. I really want to thank you for spending so much time talking to me again. If people want to learn more, where’s the best place for them to find you?

Amy: Always Twitter for me, MissAmyTobey, and I’ll see you there. Say hi.

Corey: Thank you again for being as generous with your time as you are. It’s deeply appreciated.

Amy: It’s always fun.

Corey: Amy Tobey, Senior Principal Engineer at Equinix Metal. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry comment that tells me exactly what we got wrong in this episode in the best dialect you have of condescending engineer with zero people skills. I look forward to reading it.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Shinji

Shinji Kim is the Founder & CEO of Select Star, an automated data discovery platform that helps you to understand & manage your data. Previously, she was the Founder & CEO of Concord Systems, a NYC-based data infrastructure startup acquired by Akamai Technologies in 2016. She led the strategy and execution of Akamai IoT Edge Connect, an IoT data platform for real-time communication and data processing of connected devices. Shinji studied Software Engineering at University of Waterloo and General Management at Stanford GSB.

Links Referenced:

  • Select Star: https://www.selectstar.com/
  • LinkedIn: https://www.linkedin.com/company/selectstarhq/
  • Twitter: https://twitter.com/selectstarhq

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at AWS AppConfig. Engineers love to solve, and occasionally create, problems. But not when it’s an on-call fire-drill at 4 in the morning. Software problems should drive innovation and collaboration, NOT stress, and sleeplessness, and threats of violence. That’s why so many developers are realizing the value of AWS AppConfig Feature Flags. Feature Flags let developers push code to production, but hide that that feature from customers so that the developers can release their feature when it’s ready. This practice allows for safe, fast, and convenient software development. You can seamlessly incorporate AppConfig Feature Flags into your AWS or cloud environment and ship your Features with excitement, not trepidation and fear. To get started, go to snark.cloud/appconfig. That’s snark.cloud/appconfig.

Corey: I come bearing ill tidings. Developers are responsible for more than ever these days. Not just the code that they write, but also the containers and the cloud infrastructure that their apps run on. Because serverless means it’s still somebody’s problem. And a big part of that responsibility is app security from code to cloud. And that’s where our friend Snyk comes in. Snyk is a frictionless security platform that meets developers where they are - Finding and fixing vulnerabilities right from the CLI, IDEs, Repos, and Pipelines. Snyk integrates seamlessly with AWS offerings like code pipeline, EKS, ECR, and more! As well as things you’re actually likely to be using. Deploy on AWS, secure with Snyk. Learn more at Snyk.co/scream That’s S-N-Y-K.co/scream

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Every once in a while, I encounter a company that resonates with something that I’ve been doing on some level. In this particular case, that is what’s happened here, but the story is slightly different. My guest today is Shinji Kim, who’s the CEO and founder at Select Star.

And the joke that I was making a few months ago was that Select Stars should have been the name of the Oracle ACE program instead. Shinji, thank you for joining me and suffering my ridiculous, basically amateurish and sophomore database-level jokes because I am bad at databases. Thanks for taking the time to chat with me.

Shinji: Thanks for having me here, Corey. Good to meet you.

Corey: So, Select Star despite being the only query pattern that I’ve ever effectively been able to execute from memory, what you do as a company is described as an automated data discovery platform. So, I’m going to start at the beginning with that baseline definition. I think most folks can wrap their heads around what the idea of automated means, but the rest of the words feel like it might mean different things to different people. What is data discovery from your point of view?

Shinji: Sure. The way that we define data discovery is finding and understanding data. In other words, think about how discoverable your data is in your company today. How easy is it for you to find datasets, fields, KPIs of your organization data? And when you are looking at a table, column, dashboard, report, how easy is it for you to understand that data underneath? Encompassing on that is how we define data discovery.

Corey: When you talk about data lurking around the company in various places, that can mean a lot of different things to different folks. For the more structured data folks—which I tend to think of as the organized folks who are nothing like me—that tends to mean things that live inside of, for example, traditional relational databases or things that closely resemble that. I come from a grumpy old sysadmin perspective, so I’m thinking, oh, yeah, we have a Jira server in the closet and that thing’s logging to its own disk, so that’s going to be some information somewhere. Confluence is another source of data in an organization; it’s usually where insight and a knowledge of what’s going on goes to die. It’s one of those write once, read never type of things.

And when I start thinking about what data means, it feels like even that is something of a squishy term. From the perspective of where Select Start starts and stops, is it bounded to data that lives within relational databases? Does it go beyond that? Where does it start? Where does it stop?

Shinji: So, we started the company with an intention of increasing the discoverability of data and hence providing automated data discovery capability to organizations. And the part where we see this as the most effective is where the data is currently being consumed today. So, this is, like, where the data consumption happens. So, this can be a data warehouse or data lake, but this is where your data analysts, data scientists are querying data, they are building dashboards, reports on top of, and this is where your main data mart lives.

So, for us, that is primarily a cloud data warehouse today, usually has a relational data structure. On top of that, we also do a lot of deep integrations with BI tools. So, that includes tools like Tableau, Power BI, Looker, Mode. Wherever these queries from the business stakeholders, BI engineers, data analysts, data scientists run, this is a point of reference where we use to auto-generate documentation, data models, lineage, and usage information, to give it back to the data team and everyone else so that they can learn more about the dataset they’re about to use.

Corey: So, given that I am seeing an increased number of companies out there talking about data discovery, what is it the Select Star does that differentiates you folks from other folks using similar verbiage in how they describe what they do?

Shinji: Yeah, great question. There are many players that popping up, and also, traditional data catalog’s definitely starting to offer more features in this area. The main differentiator that we have in the market today, we call it fast time-to-value. Any customer that is starting with Select Star, they get to set up their instance within 24 hours, and they’ll be able to get all the analytics and data models, including column-level lineage, popularity, ER diagrams, and how other people are—top users and how other people are utilizing that data, like, literally in few hours, max to, like, 24 hours. And I would say that is the main differentiator.

And most of the customers I have pointed out that setup and getting started has been super easy, which is primarily backed by a lot of automation that we’ve created underneath the platform. On top of that, just making it super easy and simple to use. It becomes very clear to the users that it’s not just for the technical data engineers and DBAs to use; this is also designed for business stakeholders, product managers, and ops folks to start using as they are learning more about how to use data.

Corey: Mapping this a little bit toward the use cases that I’m the most familiar with, this big source of data that I tend to stumble over is customer AWS bills. And that’s not exactly a big data problem, given that it can fit in memory if you have a sufficiently exciting computer, but using Tableau don’t wind up slicing and dicing that because at some point, Excel falls down. From my perspective, problem with Excel is that it doesn’t tend to work on huge datasets very well, and from the position of Salesforce, the problem with Excel is that it doesn’t cost a giant pile of money every month. So, those two things combined, Tableau is the answer for what we do. But that’s sort of the end-all for us of, that’s where it stops.

At that point, we have dashboards that we build and queries that we run that spit out the thing we’re looking at, and then that goes back to inform our analysis. We don’t inherently feed that back into anything else that would then inform the rest of what we do. Now, for our use case, that probably makes an awful lot of sense because we’re here to help our customers with their billing challenges, not take advantage of their data to wind up informing some giant model and mispurposing that data for other things. But if we were generating that data ourselves as a part of our operation, I can absolutely see the value of tying that back into something else. You wind up almost forming a reinforcing cycle that improves the quality of data over time and lets you understand what’s going on there. What are some of the outcomes that you find that customers get to by going down this particular path?

Shinji: Yeah, so just to double-click on what you just talked about, the way that we see this is how we analyze the metadata and the activity logs—system logs, user logs—of how that data has been used. So, part of our auto-generated documentation for each table, each column, each dashboard, you’re going to be able to see the full data lineage: where it came from, how it was transformed in the past, and where it’s going to. You will also see what we call popularity score: how many unique users are utilizing this data inside the organization today, how often. And utilizing these two core models and analysis that we create, you can start looking at first mapping out the data flow, and then determining whether or not this dataset is something that you would want to continue keeping or running the data pipelines for. Because once you start mapping these usage models of tables versus dashboards, you may find that there are recurring jobs that creates all these materialized views and tables that are feeding dashboards that are not being looked at anymore.

So, with this mechanism by looking initially data lineage as a concept, a lot of companies use data lineage in order to find dependencies: what is going to break if I make this change in the column or table, as well as just debugging any of issues that is currently happening in their pipeline. So, especially when you will have to debug a SQL query or pipeline that you didn’t build yourself but you need to find out how to fix it, this is a really easy way to instantly find out, like, where the data is coming from. But on top of that, if you start adding this usage information, you can trace through where the main compute is happening, which largest route table is still being queried, instead of the more summarized tables that should be used, versus which are the tables and datasets that is continuing to get created, feeding the dashboards and is those dashboards actually being used on the business side. So, with that, we have customers that have saved thousands of dollars every month just by being able to deprecate dashboards and pipelines that they were afraid of deprecating in the past because they weren’t sure if anyone’s actually using this or not. But adopting Select Star was a great way to kind of do a full spring clean of their data warehouse as well as their BI tool. And this is an additional benefit to just having to declutter so many old, duplicated, and outdated dashboards and datasets in their data warehouse.

Corey: That is, I guess, a recurring problem that I see in many different pockets of the industry as a whole. You see it in the user visibility space, you see it in the cost control space—I even made a joke about Confluence that alludes to it—this idea that you build a whole bunch of dashboards and use it to inform all kinds of charts and other systems, but then people are busy. It feels like there’s no ‘and then.’ Like, one of the most depressing things in the universe that you can see after having spent a fair bit of effort to build up those dashboards is the analytics for who internally has looked at any of those dashboards since the demo you gave showing it off to everyone else. It feels like in many cases, we put all these projects and amount of effort into building these things out that then don’t get used.

People don’t want to be informed by data they want to shoot from their gut. Now, sometimes that’s helpful when we’re talking about observability tools that you use to trace down outages, and, “Well, our site’s really stable. We don’t have to look at that.” Very awesome, great, awesome use case. The business insight level of dashboard just feels like that’s something you should really be checking a lot more than you are. How do you see that?

Shinji: Yeah, for sure. I mean, this is why we also update these usage metrics and lineage every 24 hours for all of our customers automatically, so it’s just up-to-date. And the part that more customers are asking for where we are heading to—earlier, I mentioned that our main focus has been on analyzing data consumption and understanding the consumption behavior to drive better usage of your data, or making data usage much easier. The part that we are starting to now see is more customers wanting to extend those feature capabilities to their staff of where the data is being generated. So, connecting the similar amount of analysis and metadata collection for production databases, Kafka Queues, and where the data is first being generated is one of our longer-term goals. And then, then you’ll really have more of that, up to the source level, of whether the data should be even collected or whether it should even enter the data warehouse phase or not.

Corey: One of the challenges I see across the board in the data space is that so many products tend to have a very specific point of the customer lifecycle, where bringing them in makes sense. Too early and it’s, “Data? What do you mean data? All I have are these logs, and their purpose is basically to inflate my AWS bill because I’m bad at removing them.” And on the other side, it’s, “Great. We pioneered some of these things and have built our own internal enormous system that does exactly what we need to do.” It’s like, “Yes, Google, you’re very smart. Good job.” And most people are somewhere between those two extremes. Where are customers on that lifecycle or timeline when using Select Star makes sense for them?

Shinji: Yeah, I think that’s a great question. Also the time, the best place where customers would use Select Star for is that after they have their cloud data warehouse set up. Either they have finished their migration, they’re starting to utilize it with their BI tools, and they’re starting to notice that it’s not just, like, you know, ten to fifty tables that they’re starting with; most of them have more than hundreds of tables. And they’re feeling that this is starting to go out of control because we have all these data, but we are not a hundred percent sure what exactly is in our database. And this usually just happens more in larger companies, companies at thousand-plus employees, and they usually find a lot of value out of Select Star right away because, like, we will start pointing out many different things.

But we also see a lot of, like, forward-thinking, fast-growing startups that are at the size of a few hundred employees, you know, they now have between five to ten-person data team, and they are really creating the right single source of truth of their data knowledge through a Select Star. So, I think you can start anywhere from when your data team size is, like, beyond five and you’re continuing to grow because every time you’re trying to onboard a data analyst, data scientist, you will have to go through, like, basically the same type of training of your data model, and it might actually look different because the data models and the new features, new apps that you’re integrating this changes so quickly. So, I would say it’s important to have that base early on and then continue to grow. But we do also see a lot of companies coming to us after having thousands of datasets or tens of thousands of datasets that it’s really, like, very hard to operate and onboard anyone. And this is a place where we really shine to help their needs, as well.

Corey: Sort of the, “I need a database,” to the, “Help, I have too many databases,” pipeline, where [laugh] at some point people start to—wanting to bring organization to the chaos. One thing I like about your model is that you don’t seem to be making the play that every other vendor in the data space tends to, which is, “Oh, we want you to move your data onto our systems. The end.” You operate on data that is in place, which makes an awful lot of sense for the kinds of things that we’re talking about. Customers are flat out not going to move their data warehouse over to your environment, just because the data gravity is ludicrous. Just the sheer amount of money it would take to egress that data from a cloud provider, for example, is monstrous.

Shinji: Exactly. [laugh]. And security concerns. We don’t want to be liable for any of the data—and this is, like, a very specific decision we’ve made very early on the company—to not access data, to not egress any of the real data, and to provide as much value as possible just utilizing the metadata and logs. And depending on the types of data warehouses, it also can be really efficient because the query history or the metadata systems tables are indexed separately. Usually, it’s much lighter load on the compute side. And that definitely has, like, worked well for our advantage, especially being a SaaS tool.

Corey: This episode is sponsored in part by our friends at Sysdig. Sysdig secures your cloud from source to run. They believe, as do I, that DevOps and security are inextricably linked. If you wanna learn more about how they view this, check out their blog, it's definitely worth the read. To learn more about how they are absolutely getting it right from where I sit, visit Sysdig.com and tell them that I sent you. That's S Y S D I G.com. And my thanks to them for their continued support of this ridiculous nonsense.

Corey: What I like is just how straightforward the integrations are. It’s clear you’re extraordinarily agnostic as far as where the data itself lives. You integrate with Google’s BigQuery, with Amazon Redshift, with Snowflake, and then on the other side of the world with Looker, and Tableau, and other things as well. And one of the example use cases you give is find the upstream table in BigQuery that a Looker dashboard depends on. That’s one of those areas where I see something like that, and, oh, I can absolutely see the value of that.

I have two or three DynamoDB tables that drive my newsletter publication system that I built—because I have deep-seated emotional problems and I take it out and everyone else via code—but as a small, contained system that I can still fit in my head. Mostly. And I still forget which table is which in some cases. Down the road, especially at scale, “Okay, where is the actual data source that’s informing this because it doesn’t necessarily match what I’m expecting,” is one of those incredibly valuable bits of insight. It seems like that is something that often gets lost; the provenance of data doesn’t seem to work.

And ideally, you know, you’re staffing a company with reasonably intelligent people who are going to look at the results of something and say, “That does not align with my expectations. I’m going to dig.” As opposed to the, “Oh, yeah, that seems plausible. I’ll just go with whatever the computer says.” There’s an ocean of nuance between those two, but it’s nice to be able to establish the validity of the path that you’ve gone down in order to set some of these things up.

Shinji: Yeah, and this is also super helpful if you’re tasked to debug a dashboard or pipeline that you did not build yourself. Maybe the person has left the company, or maybe they’re out-of-office, but this dashboard has been broken and you’re quote-unquote, “On call,” for data. What are you going to do? You’re going to—without a tool that can show you a full lineage, you will have to start digging through somebody else’s SQL code and try to map out, like, where the data is coming from, if this is calculating correctly. Usually takes, you know, few hours to just get to the bottom of the issue. And this is one of the main use cases that our customers bring up every single time, as more of, like, this is now the go-to place every time there is any data questions or data issues.

Corey: The first and golden rule of cloud economics is step one, turn that shit off.

Shinji: [laugh].

Corey: When people are using something, you can optimize the hell out of it however you want, but nothing’s going to beat turning it off. One challenge is when we’re looking at various accounts and we see a Redshift cluster, and it’s, “Okay. That thing’s costing a few million bucks a year and no one seems to know anything about it.” They keep pointing to other teams, and it turns into this giant, like, finger-pointing exercise where no one seems to have responsibility for it. And very often, our clients will choose not to turn that thing off because on the one hand, if you don’t turn it off, you’re going to spend a few million bucks a year that you otherwise would not have had to.

On the other, if you delete the data warehouse, and it turns out, oh, yeah, that was actually kind of important, now we don’t have a company anymore. It’s a question of which is the side you want to be wrong on. And in some levels, leaving something as it is and doing something else is always a more defensible answer, just because the first time your cost-saving exercises take out production, you’re generally not allowed to save money anymore. This feels like it helps get to that source of truth a heck of a lot more effectively than tracing individual calls and turning into basically data center archaeologists.

Shinji: [laugh]. Yeah, for sure. I mean, this is why from the get go, we try to give you all your tables, all of your database, just ordered by popularity. So, you can also see overall, like, from all the tables, whether that’s thousands or tens of thousands, you’re seeing the most used, has the most number of dependencies on the top, and you can also filter it by all the database tables that hasn’t been touched in the last 90 days. And just having this, like, high-level view gives a lot of ideas to the data platform team about how they can optimize usage of their data warehouse.

Corey: From where I tend to sit, an awful lot of customers are still relatively early in their data journey. An awful lot of the marketing that I receive from various AWS mailing lists that I found myself on because I’ve had the temerity to open accounts has been along the lines of oh, data discovery is super important, but first, they presuppose that I’ve already bought into this idea that oh, every company must be a completely data-driven company. The end. Full stop.

And yeah, we’re a small bespoke services consultancy. I don’t necessarily know that that’s the right answer here. But then it takes it one step further and starts to define the idea of data discovery as, ah, you will use it to find a PII or otherwise sensitive or restricted data inside of your datasets so you know exactly where it lives. And sure, okay, that’s valuable, but it also feels like a very narrow definition compared to how you view these things.

Shinji: Yeah. Basically, the way that we see data discovery is it’s starting to become more of an essential capability in order for you to monitor and understand how your data is actually being used internally. It basically gives you the insights around sure, like, what are the duplicated datasets, what are the datasets that have that descriptions or not, what are something that may contain sensitive data, so on and so forth, but that’s still around the characteristics of the physical datasets. Whereas I think the part that’s really important around data discovery that is not being talked about as much is how the data can actually be used better. So, have it as more of a forward-thinking mechanism and in order for you to actually encourage more people to utilize data or use the data correctly, instead of trying to contain this within just one team is really where I feel like data discovery can help.

And in regards to this, the other big part around data discovery is really opening up and having that transparency just within the data team. So, just within the data team, they always feel like they do have that access to the SQL queries and you can just go to GitHub and just look at the database itself, but it’s so easy to get lost in the sea of metadata that is just laid out as just the list; there isn’t much context around the data itself. And that context and with along with the analytics of the metadata is what we’re really trying to provide automatically. So eventually, like, this can be also seen as almost like a way to, like, monitor the datasets, like, how you’re currently monitoring your applications through Datadog or your website with your Google Analytics, this is something that can be also used as more of a go-to source of truth around what your state of the data is, how that’s defined, and how that’s being mapped to different business processes, so that there isn’t much confusion around data. Everything can be called the same, but underneath it actually can mean very different things. Does that make sense?

Corey: No, it absolutely does. I think that this is part of the challenge in trying to articulate value that is, I guess, specific to this niche across an entire industry. The context that drives data is going to be incredibly important, and it feels like so much of the marketing in the space is aimed at one or two pre-imagined customer profiles. And that has the side effect of making customers for whom that model doesn’t align, look and feel like either doing something wrong, or makes it look like the vendor who’s pitching this is somewhat out of touch. I know that I work in a relatively bounded problem space, but I still learn new things about AWS billing on virtually every engagement that I go on, just because you always get to learn more about how customers view things and how they view not just their industry, but also the specificities of their own business and their own niche.

I think that is one of the challenges historically, with the idea of letting software do everything. Do you find the problems that you’re solving tend to be global in nature or are you discovering strange depths of nuance on a customer-by-customer basis at this point?

Shinji: Overall, a lot of the problems that we solve and the customers that we work with is very industry agnostic. As long as you are having many different datasets that you need to manage, there are common problems that arises, regardless of the industry that you’re in. We do observe some industry-specific issues because your data is either, it’s an unstructured data, or your data is primarily events, or you know, depending on how the data looks like, but primarily because of most of the BI solutions and data warehouses are operating as a relational databases, this is a part where we really try to build a lot of best practices, and the common analytics that we can apply to every customer that’s using Select Star.

Corey: I really want to thank you for taking so much time to go through the ins and outs of what it is you’re doing these days. If people want to learn more, where’s the best place to find you?

Shinji: Yeah, I mean, it’s been fun [laugh] talking here. So, we are at selectstar.com. That’s our website. You can sign up for a free trial. It’s completely self-service, so you don’t need to get on a demo but, like, we’ll also help you onboard and happy to give a free demo to whoever that is interested.

We are also on LinkedIn and Twitter under selectstarhq. Yeah, I mean, we’re happy to help for any companies that have these issues around wanting to increase their discoverability of data, and want to help their data team and the rest of the company to be able to utilize data better.

Corey: And we will, of course, put links to all of that in the [show notes 00:28:58]. Thank you so much for your time today. I really appreciate it.

Shinji: Great. Thanks for having me, Corey.

Corey: Shinji Kim, CEO and founder at Select Star. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry comment that I won’t be able to discover because there are far too many podcast platforms out there, and I have no means of discovering where you’ve said that thing unless you send it to me.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Scott

With more than 28 years of successful leadership in building high technology companies and delivering advanced products to market, Scott provides the overall strategic leadership and visionary direction for Azul Systems.

Scott has a consistent proven track record of vision, leadership, and success in enterprise, consumer and scientific markets. Prior to co-founding Azul Systems, Scott founded 3dfx Interactive, a graphics processor company that pioneered the 3D graphics market for personal computers and game consoles. Scott served at 3dfx as Vice President of Engineering, CTO and as a member of the board of directors and delivered 7 award-winning products and developed 14 different graphics processors. After a successful initial public offering, 3dfx was later acquired by NVIDIA Corporation.

Prior to 3dfx, Scott was a CPU systems architect at Pellucid, later acquired by MediaVision. Before Pellucid, Scott was a member of the technical staff at Silicon Graphics where he designed high-performance workstations.

Scott graduated from Princeton University with a bachelor of science, earning magna cum laude and Phi Beta Kappa honors. Scott has been granted 8 patents in high performance graphics and computing and is a regularly invited keynote speaker at industry conferences.

Links Referenced:

  • Azul: https://www.azul.com/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: I come bearing ill tidings. Developers are responsible for more than ever these days. Not just the code that they write, but also the containers and the cloud infrastructure that their apps run on. Because serverless means it’s still somebody’s problem. And a big part of that responsibility is app security from code to cloud. And that’s where our friend Snyk comes in. Snyk is a frictionless security platform that meets developers where they are - Finding and fixing vulnerabilities right from the CLI, IDEs, Repos, and Pipelines. Snyk integrates seamlessly with AWS offerings like code pipeline, EKS, ECR, and more! As well as things you’re actually likely to be using. Deploy on AWS, secure with Snyk. Learn more at Snyk.co/scream That’s S-N-Y-K.co/scream

Corey: This episode is sponsored in part by our friends at AWS AppConfig. Engineers love to solve, and occasionally create, problems. But not when it’s an on-call fire-drill at 4 in the morning. Software problems should drive innovation and collaboration, NOT stress, and sleeplessness, and threats of violence. That’s why so many developers are realizing the value of AWS AppConfig Feature Flags. Feature Flags let developers push code to production, but hide that that feature from customers so that the developers can release their feature when it’s ready. This practice allows for safe, fast, and convenient software development. You can seamlessly incorporate AppConfig Feature Flags into your AWS or cloud environment and ship your Features with excitement, not trepidation and fear. To get started, go to snark.cloud/appconfig. That’s snark.cloud/appconfig.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. My guest on this promoted episode today is Scott Sellers, CEO and co-founder of Azul. Scott, thank you for joining me.

Scott: Thank you, Corey. I appreciate the opportunity in talking to you today.

Corey: So, let’s start with what you’re doing these days. What is Azul? What do you folks do over there?

Scott: Azul is an enterprise software and SaaS company that is focused on delivering more efficient Java solutions for our customers around the globe. We’ve been around for 20-plus years, and as an entrepreneur, we’ve really gone through various stages of different growth and different dynamics in the market. But at the end of the day, Azul is all about adding value for Java-based enterprises, Java-based applications, and really endearing ourselves to the Java community.

Corey: This feels like the sort of space where there are an awful lot of great business cases to explore. When you look at what’s needed in that market, there are a lot of things that pop up. The surprising part to me is that this is the direction that you personally went in. You started your career as a CPU architect, to my understanding. You were then one of the co-founders of 3dfx before it got acquired by Nvidia.

You feel like you’ve spent your career more as a hardware guy than working on the SaaS side of the world. Is that a misunderstanding of your path, or have things changed, or is this just a new direction? Help me understand how you got here from where you were.

Scott: I’m not exactly sure what the math would say because I continue to—can’t figure out a way to stop time. But you’re correct that my academic background, I was an electrical engineer at Princeton and started my career at Silicon Graphics. And that was when I did a lot of fantastic and fascinating work building workstations and high-end graphics systems, you know, back in the day when Silicon Graphics really was the who’s who here in Silicon Valley. And so, a lot of my career began in the context of hardware. As you mentioned, I was one of the founders of graphics company called 3dfx that was one of, I think, arguably the pioneer in terms of bringing 3d graphics to the masses, if you will.

And we had a great run of that. That was a really fun business to be a part of just because of what was going on in the 3d world. And we took that public and eventually sold that to Nvidia. And at that point, my itch, if you will, was really learning more about the enterprise segment. I’d been involved with professional graphics with SGI, I had been involved with consumer graphics with 3dfx.

And I was fascinated just to learn about the enterprise segment. And met a couple people through a mutual friend around the 2001 timeframe, and they started talking about this thing called Java. And you know, I had of course heard about Java, but as a consumer graphics guy, didn’t have a lot of knowledge about it or experience with it. And the more I learned about it, recognized that what was going on in the Java world—and credit to Sun for really creating, obviously, not only language, but building a community around Java—and recognized that new evolutions of developer paradigms really only come around once a decade if then, and was convinced and really got excited about the opportunity to ride the wave of Java and build a company around that.

Corey: One of the blind spots that I have throughout the entire world of technology—and to be fair, I have many of them, but the one most relevant to this conversation, I suppose, is the Java ecosystem as a whole. I come from a background of being a grumpy Unix sysadmin—because I’ve never met a happy one of those in my entire career—and as a result, scripting languages is where everything that I worked with started off. And on the rare occasions, I worked in Java shops, it was, “Great. We’re going to go—here’s a WAR file. Go ahead and deploy this with Tomcat,” or whatever else people are going to use. But basically, “Don’t worry your pretty little head about that.”

At most, I have to worry about how to configure a heap or whatnot. But it’s from the outside looking in, not having to deal with that entire ecosystem as a whole. And what I’ve seen from that particular perspective is that every time I start as a technologist, or even as a consumer trying to install some random software package in the depths of the internet, and I have to start thinking about Java, it always feels like I’m about to wind up in a confusing world. There are a number of software packages that I installed back in, I want to say the early-2010s or whatnot. “Oh, you need to have a Java runtime installed on your Mac,” for example.

And okay, going through Oracle site, do I need the JRE? Do I need the JDK? Oh, there’s OpenJDK, which kind of works, kind of doesn’t. Amazon got into the space with Corretto, which because that sounds nothing whatsoever, like Java, but strange names coming from Amazon is basically par for the course for those folks. What is the current state of the Java ecosystem, for those of us who have—basically the closest we’ve ever gotten is JavaScript, which is nothing alike except for the name.

Scott: And you know, frankly, given the protection around the name Java—and you know, that is a trademark that’s owned by Oracle—it’s amazing to me that JavaScript has been allowed to continue to be called JavaScript because as you point out, JavaScript has nothing to do with Java per se.

Corey: Well, one thing they do have in common I found out somewhat recently is that Oracle also owns the trademark for JavaScript.

Scott: Ah, there you go. Maybe that’s why it continues.

Corey: They’re basically a law firm—three law firms in a trench coat, masquerading as a tech company some days.

Scott: Right. But anyway, it is a confusing thing because you know, I think, arguably, JavaScript, by the numbers, probably has more programmers than any other language in the world, just given its popularity as a web language. But to your question about Java specifically, it’s had an evolving life, and I think the state where it is today, I think it’s in the most exciting place it’s ever been. And I’ll walk you through kind of why I believe that to be the case.

But Java has evolved over time from its inception back in the days when it was called, I think it was Oak when it was originally conceived, and Sun had eventually branded it as Java. And at the time, it truly was owned by Sun, meaning it was proprietary code; it had to be licensed. And even though Sun gave it away, in most cases, it still at the end of the day, it was a commercially licensed product, if you will, and platform. And if you think about today’s world, it would not be conceivable to create something that became so popular with programmers that was a commercially licensed product today. It almost would be mandated that it would be open-source to be able to really gain the type of traction that Java has gained.

And so, even though Java was really garnering interest, you know, not only within the developer community, but also amongst commercial entities, right, everyone—and the era now I’m talking about is around the 2000 era—all of the major software vendors, whether it was obviously Sun, but then you had Oracle, you had IBM, companies like BEA, were really starting to blossom at that point. It was a—you know, you could almost not find a commercial software entity that was not backing Java. But it was still all controlled by Sun. And all that success ultimately led to a strong outcry from the community saying this has to be open-source; this is too important to be beholden to a single vendor. And that decision was made by Sun prior to the Oracle acquisition, they actually open-sourced the Java runtime code and they created an open-source project called OpenJDK.

And to Oracle’s credit, when they bought Sun—which I think at the time when you really look back, Oracle really did not have a lot of track record, if you will, of being involved with an open-source community—and I think when Oracle acquired Sun, there was a lot of skepticism as to what’s going to happen to Java. Is Oracle going to make this thing, you know, back to the old days, proprietary Oracle, et cetera? And really—

Corey: I was too busy being heartbroken over Solaris at that point to pay much attention to the Java stuff, but it felt like it was this—sort of the same pattern, repeated across multiple ecosystems.

Scott: Absolutely. And even though Sun had also open-sourced Solaris, with the OpenSolaris project, that was one of the kinds of things that it was still developed very much in a closed environment, and then they would kind of throw some code out into the open world. And no one really ran OpenSolaris because it wasn’t fully compatible with Solaris. And so, that was a faint attempt, if you will.

But Java was quite different. It was truly all open-sourced, and the big difference that—and again, I give Oracle a lot of credit for this because this was a very important time in the evolution of Java—that Oracle, maintained Sun’s commitment to not only continue to open-source Java but most importantly, develop it in the open community. And so, you know, again, back and this is the 2008, ‘09, ‘10 timeframe, the evolution of Java, the decisions, the standards, you know, what goes in the platform, what doesn’t, decisions about updates and those types of things, that truly became a community-led world and all done in the open-source. And credit to Oracle for continuing to do that. And that really began the transition away from proprietary implementations of Java to one that, very similar to Linux, has really thrived because of the true open-source nature of what Java is today.

And that’s enabled more and more companies to get involved with the evolution of Java. If you go to the OpenJDK page, you’ll see all of the not only, you know, incredibly talented individuals that are involved with the evolution of Java, but again, a who’s who in pretty much every major commercial entities in the enterprise software world is also somehow involved in the OpenJDK community. And so, it really is a very vibrant, evolving standard. And some of the tactical things that have happened along the way in terms of changing how versions of Java are released still also very much in the context of maintaining compatibility and finding that careful balance of evolving the platform, but at the same time, recognizing that there is a lot of Java applications out there, so you can’t just take a right-hand turn and forget about the compatibility side of things. But we as a community overall, I think, have addressed that very effectively, and the result has been now I think Java is more popular than ever and continues to—we liken it kind of to the mortar and the brick walls of the enterprise. It’s a given that it’s going to be used, certainly by most of the enterprises worldwide today.

Corey: There’s a certain subset of folk who are convinced the Java, “Oh, it’s this a legacy programming language, and nothing modern or forward-looking is going to be built in it.” Yeah, those people generally don’t know what the internal language stack looks like at places like oh, I don’t know, AWS, Google, and a few others, it is very much everywhere. But it also feels, on some level, like, it’s a bit below the surface-level of awareness for the modern full-stack developer in some respects, right up until suddenly it’s very much not. How is Java evolving in a cloud these days?

Scott: Well, what we see happening—you know, this is true for—you know, I’m a techie, so I can talk about other techies. I mean as techies, we all like the new thing, right? I mean, it’s not that exciting to talk about a language that’s been around for 20-plus years. But that doesn’t take away from the fact that we still all use keyboards. I mean, no one really talks about what keyboard they use anymore—unless you’re really into keyboards—but at the end of the day, it’s still a fundamental tool that you use every single day.

And Java is kind of in the same situation. The reason that Java continues to be so fundamental is that it really comes back to kind of reinventing the wheel problem. Are there are other languages that are more efficient to code in? Absolutely. Are there other languages that, you know, have some capabilities that the Java doesn’t have? Absolutely.

But if you have the ability to reinvent everything from scratch, sure, go for it. And you also don’t have to worry about well, can I find enough programmers in this, you know, new hot language, okay, good luck with that. You might be able to find dozens, but when you need to really scale a company into thousands or tens of thousands of developers, good luck finding, you know, everyone that knows, whatever your favorite hot language of the day is.

Corey: It requires six years experience in a four-year-old language. Yeah, it’s hard to find that, sometimes.

Scott: Right. And you know, the reality is, is that really no application ever is developed from scratch, right? Even when an application is, quote, new, immediately, what you’re using is frameworks and other things that have written long ago and proven to be very successful.

Corey: And disturbing amounts of code copied and pasted from Stack Overflow.

Scott: Absolutely.

Corey: But that’s one of those impolite things we don’t say out loud very often.

Scott: That’s exactly right. So, nothing really is created from scratch anymore. And so, it’s all about building blocks. And this is really where this snowball of Java is difficult to stop because there is so much third-party code out there—and by that, I mean, you know, open-source, commercial code, et cetera—that is just so leveraged and so useful to very quickly be able to take advantage of and, you know, allow developers to focus on truly new things, not reinventing the wheel for the hundredth time. And that’s what’s kind of hard about all these other languages is catching up to Java with all of the things that are immediately available for developers to use freely, right, because most of its open-source. That’s a pretty fundamental Catch-22 about when you start talking about the evolution of new languages.

Corey: I’m with you so far. The counterpoint though is that so much of what we’re talking about in the world of Java is open-source; it is freely available. The OpenJDK, for example, says that right on the tin. You have built a company and you’ve been in business for 20 years. I have to imagine that this is not one of those stories where, “Oh, all the things we do, we give away for free. But that’s okay. We make it up in volume.” Even the venture capitalist mindset tends to run out of patience on those kinds of timescales. What is it you actually do as a business that clearly, obviously delivers value for customers but also results in, you know, being able to meet payroll every week?

Scott: Right? Absolutely. And I think what time has shown is that, with one very notable exception and very successful example being Red Hat, there are very, very few pure open-source companies whose business is only selling support services for free software. Most successful businesses that are based on open-source are in one-way shape or form adding value-added elements. And that’s our strategy as well.

The heart of everything we do is based on free code from OpenJDK, and we have a tremendous amount of business that we are following the Red Hat business model where we are selling support and long-term access and a huge variety of different operating system configurations, older Java versions. Still all free software, though, right, but we’re selling support services for that. And that is, in essence, the classic Red Hat business model. And that business for us is incredibly high growth, very fast-moving, a lot of that business is because enterprises are tired of paying the very high price to Oracle for Java support and they’re looking for an open-source alternative that is exactly the same thing, but comes in pure open-source form and with a vendor that is as reputable as Oracle. So, a lot of our businesses based on that.

However, on top of that, we also have value-added elements. And so, our product that is called Azul Platform Prime is rooted in OpenJDK—it is OpenJDK—but then we’ve added value-added elements to that. And what those value-added elements create is, in essence, a better Java platform. And better in this context means faster, quicker to warm up, elimination of some of the inconsistencies of the Java runtime in terms of this nasty problem called garbage collection which causes applications to kind of bounce around in terms of performance limitations. And so, creating a better Java is another way that we have monetized our company is value-added elements that are built on top of OpenJDK. And I’d say that part of the business is very typical for the majority of enterprise software companies that are rooted in open-source. They’re typically adding value-added components on top of the open-source technology, and that’s our similar strategy as well.

And then the third evolution for us, which again is very tried-and-true, is evolving the business also to add SaaS offerings. So today, the majority of our customers, even though they deploy in the cloud, they’re stuck customer-managed and so they’re responsible for where do I want to put my Java runtime on building out my stack and cetera, et cetera. And of course, that could be on-prem, but like I mentioned, the majority are in the cloud. We’re evolving our product offerings also to have truly SaaS-based solutions so that customers don’t even need to manage those types of stacks on their own anymore.

Corey: On some level, it feels like we’re talking about two different things when we talk about cloud and when we talk about programming languages, but increasingly, I’m starting to see across almost the entire ecosystem that different languages and different cloud providers are in many ways converging. How do you see Java changing as cloud-native becomes the default rather than the new thing?

Scott: Great question. And I think the thing to recognize about, really, most popular programming languages today—I can think of very few exceptions—these languages were created, envisioned, implemented if you will, in a day when cloud was not top-of-mind, and in many cases, certainly in the case of Java, cloud didn’t even exist when Java was originally conceived, nor was that the case when you know, other languages, such as Python, or JavaScript, or on and on. So, rethinking how these languages should evolve in very much the context of a cloud-native mentality is a really important initiative that we certainly are doing and I think the Java community is doing overall. And how you architect not only the application, but even the Java runtime itself can be fundamentally different if you know that the application is going to be deployed in the cloud.

And I’ll give you an example. Specifically, in the world of any type of runtime-based language—and JavaScript is an example of that; Python is an example of that; Java is an example of that—in all of those runtime-based environments, what that basically means is that when the application is run, there’s a piece of software that’s called the runtime that actually is running that application code. And so, you can think about it as a middleware piece of software that sits between the operating system and the application itself. And so, that runtime layer is common across those languages and those platforms that I mentioned. That runtime layer is evolving, and it’s evolving in a way that is becoming more and more cloud-native in it’s thinking.

The process itself of actually taking the application, compiling it into whatever underlying architecture it may be running on—it could be an x86 instance running on Amazon; it could be, you know, for example, an ARM64, which Amazon has compute instances now that are based on an ARM64 processor that they call Graviton, which is really also kind of altering the price-performance of the compute instances on the AWS platform—that runtime layer magically takes an application that doesn’t have to be aware of the underlying hardware and transforms that into a way that can be run. And that’s a very expensive process; it’s called just-in-time compiling, and that just-in-time compilation, in today’s world—which wasn’t really based on cloud thinking—every instance, every compute instance that you deploy, that same JIT compilation process is happening over and over again. And even if you deploy 100 instances for scalability, every one of those 100 instances is doing that same work. And so, it’s very inefficient and very redundant. Contrast that to a cloud-native thinking: that compilation process should be a service; that service should be done once.

The application—you know, one instance of the application is actually run and there are the other ninety-nine should just reuse that compilation process. And that shared compiler service should be scalable and should be able to scale up when applications are launched and you need more compilation resources, and then scaled right back down when you’re through the compilation process and the application is more moving into the—you know, to the runtime phase of the application lifecycle. And so, these types of things are areas that we and others are working on in terms of evolving the Java runtime specifically to be more cloud-native.

Corey: This episode is sponsored in part by our friends at Sysdig. Sysdig secures your cloud from source to run. They believe, as do I, that DevOps and security are inextricably linked. If you wanna learn more about how they view this, check out their blog, it's definitely worth the read. To learn more about how they are absolutely getting it right from where I sit, visit Sysdig.com and tell them that I sent you. That's S Y S D I G.com. And my thanks to them for their continued support of this ridiculous nonsense.

Corey: This feels like it gets even more critical when we’re talking about things like serverless functions across basically all the cloud providers these days, where there’s the whole setup, everything in the stack, get it running, get it listening, ready to go, to receive a single request and then shut itself down. It feels like there are a lot of operational efficiencies possible once you start optimizing from a starting point of yeah, this is what that environment looks like, rather than us big metal servers sitting in a rack 15 years ago.

Scott: Yeah. I think the evolution of serverless appears to be headed more towards serverless containers as opposed to serverless functions. Serverless functions have a bunch of limitations in terms of when you think about it in the context of a complex, you know, microservices-based deployment framework. It’s just not very efficient, to spin up and spin down instances of a function if that actually is being—it is any sort of performance or latency-sensitive type of applications. If you’re doing something very rarely, sure, it’s fine; it’s efficient, it’s elegant, et cetera.

But any sort of thing that has real girth to it—and girth probably means that’s what’s driving your application infrastructure costs, that’s what’s driving your Amazon bill every month—those types of things typically are not going to be great for starting and stopping functional instances. And so, serverless is evolving more towards thinking about the container itself not having to worry about the underlying operating system or the instance on Amazon that it’s running on. And that’s where, you know, we see more and more of the evolution of serverless is thinking about it at a container-level as opposed to a functional level. And that appears to be a really healthy steady state, so it gets the benefits of not having to worry about all the underlying stuff, but at the same time, doesn’t have the downside of trying to start and stop functional influences at a given point in time.

Corey: It seems to me that there are really two ways of thinking about cloud. The first is what I think a lot of companies do their first outing when they’re going into something like AWS. “Okay, we’re going to get a bunch of virtual machines that they call instances in AWS, we’re going to run things just like it’s our data center except now data transfer to the internet is terrifyingly expensive.” The more quote-unquote, “Cloud-native” way of thinking about this is what you’re alluding to where there’s, “Here’s some code that I wrote. I want to throw it to my cloud provider and just don’t tell me about any of the infrastructure parts. Execute this code when these conditions are met and leave me alone.”

Containers these days seem to be one of our best ways of getting there with a minimum of fuss and friction. What are you seeing in the enterprise space as far as adoption of those patterns go? Or are we seeing cloud repatriation showing up as a real thing and I’m just not in the right place to see it?

Scott: Well, I think as a cloud journey evolves, there’s no question that—and in fact it’s even silly to say that cloud is here to stay because I think that became a reality many, many years ago. So really, the question is, what are the challenges now with cloud deployments? Cloud is absolutely a given. And I think you stated earlier, it’s rare that, whether it’s a new company or a new application, at least in most businesses that don’t have specific regulatory requirements, that application is highly, highly likely to be envisioned to be initially and only deployed in the cloud. That’s a great thing because you have so many advantages of not having to purchase infrastructure in advance, being able to tap into all of the various services that are available through the cloud providers. No one builds databases anymore; you’re just tapping into the service that’s provided by Azure or AWS, or what have you.

And, you know, just that specific example is a huge amount of savings in terms of just overhead, and license costs, and those types of stuff, and there’s countless examples of that. And so, the services that are available in the cloud are unquestioned. So, there’s countless advantages of why you want to be in the cloud. The downside, however, the cloud that is, if at the end of the day, AWS, Microsoft with Azure, Google with GCP, they are making 30% margin on that cloud infrastructure. And in the days of hardware, when companies would actually buy their servers from Dell, or HP, et cetera, those businesses are 5% margin.

And so, where’s that 25% going? Well, the 25% is being paid for by the users of cloud, and as a result of that, when you look at it purely from an operational cost perspective, it is more expensive to run in the cloud than it is back in the legacy days, right? And that’s not to say that the industry has made the wrong choice because there’s so many advantages of being in cloud, there’s no doubt about it. And there should be—you know, and the cloud providers deserve to take some amount of margin to provide the services that they provide; there’s no doubt about that. The question is, how do you do the best of all worlds?

And you know, there is a great blog by a couple of the partners in Andreessen Horowitz, they called this the Cloud Paradox. And the Cloud Paradox really talks about the challenges. It’s really a Catch-22; how do you get all the benefits of cloud but do that in a way that is not overly taxing from a cost perspective? And a lot of it comes down to good practices and making sure that you have the right monitoring and culture within an enterprise to make sure that cloud cost is a primary thing that is discussed and metric, but then there’s also technologies that can help so that you don’t have to even think about what you really don’t ever want to do: repatriating, which is about the concept of actually moving off the cloud back to the old way of doing things. So certainly, I don’t believe repatriation is a practical solution for ongoing and increasing cloud costs. I believe technology is a solution to that.

And there are technologies such as our product, Azul Platform Prime, that in essence, allows you to do more with less, right, get all the benefits of cloud, deploy in your Amazon environment, deploy in your Azure environment, et cetera, but imagine if instead of needing a hundred instances to handle your given workload, you could do that with 50 or 60. Tomorrow, that means that you can start savings and being able to do that simply by changing your JVM from a standard OpenJDK or Oracle JVM to something like Platform Prime, you can immediately start to start seeing the benefits from that. And so, a lot of our business now and our growth is coming from companies that are screaming under the ongoing cloud costs and trying to keep them in line, and using technology like Azul Platform Prime to help mitigate those costs.

Corey: I think that there is a somewhat foolish approach that I’m seeing taken by a lot of folks where there are some companies that are existentially anti-cloud, if for no other reason than because if the cloud wins, then they don’t really have a business anymore. The problem I see with that is that it seems that their solution across the board is to turn back the clock where if I’m going to build a startup, it’s time for me to go buy some servers and a rack somewhere and start negotiating with bandwidth providers. I don’t see that that is necessarily viable for almost anyone. We aren’t living in 1995 anymore, despite how much some people like to pretend we are. It seems like if there are workloads—for which I agree, cloud is not necessarily an economic fit, first, I feel like the market will fix that in the fullness of time, but secondly, on an individual workload belonging in a certain place is radically different than, “Oh, none of our stuff should live on cloud. Everything belongs in a data center.” And I just think that companies lose all credibility when they start pretending that it’s any other way.

Scott: Right. I’d love to see the reaction of the venture capitalists’ face when an entrepreneur walks in and talks about how their strategy for deploying their SaaS service is going to be buying hardware and renting some space in the local data center.

Corey: Well, there is a good cost control method, if you think about it. I mean very few engineers are going to accidentally spin up an $8 million cluster in a data center a second time, just because there’s no space left for it.

Scott: And you’re right; it does happen in the cloud as well. It’s just, I agree with you completely that as part of the evolution of cloud, in general, is an ever-improving aspect of cost and awareness of cost and building in technologies that help mitigate that cost. So, I think that will continue to evolve. I think, you know, if you really think about the cloud journey, cost, I would say, is still in early phases of really technologies and practices and processes of allowing enterprises to really get their head around cost. I’d still say it’s a fairly immature industry that is evolving quickly, just given the importance of it.

And so, I think in the coming years, you’re going to see a radical improvement in terms of cost awareness and technologies to help with costs, that again allows you to the best of all worlds. Because, you know, if you go back to the Dark Ages and you start thinking about buying servers and infrastructure, then you are really getting back to a mentality of, “I’ve got to deploy everything. I’ve got to buy software for my database. I’ve got to deploy it. What am I going to do about my authentication service? So, I got to buy this vendor’s, you know, solution, et cetera.” And so, all that stuff just goes away in the world of cloud, so it’s just not practical, in this day and age I think, to think about really building a business that’s not cloud-native from the beginning.

Corey: I really want to thank you for spending so much time talking to me about how you view the industry, the evolution we’ve seen in the Java ecosystem, and what you’ve been up to. If people want to learn more, where’s the best place for them to find you?

Scott: Well, there’s a thing called a website that you may not have heard of, it’s really cool.

Corey: Can I build it in Java?

Scott: W-W-dot—[laugh]. Yeah. Azul website obviously has an awful lot of information about that, Azul is spelled A-Z-U-L, and we sometimes get the question, “How in the world did you name a company—why did you name it Azul?”

And it’s kind of a funny story because back in the days of Azul when we thought about, hey, we want to be big and successful, and at the time, IBM was the gold standard in terms of success in the enterprise world. And you know, they were Big Blue, so we said, “Hey, we’re going to be a little blue. Let’s be Azul.” So, that’s where we began. So obviously, go check out our site.

We’re very present, also, in the Java community. We’re, you know, many developer conferences and talks. We sponsor and run many of what’s called the Java User Groups, which are very popular 10-, 20-person meetups that happen around the globe on a regular basis. And so, you know, come check us out. And I appreciate everyone’s time in listening to the podcast today.

Corey: No, thank you very much for spending as much time with me as you have. It’s appreciated.

Scott: Thanks, Corey.

Corey: Scott Sellers, CEO and co-founder of Azul. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an entire copy of the terms and conditions from Oracle’s version of the JDK.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Allen

Allen is a cloud architect at Tyler Technologies. He helps modernize government software by creating secure, highly scalable, and fault-tolerant serverless applications.

Allen publishes content regularly about serverless concepts and design on his blog - Ready, Set Cloud!

Links Referenced:

  • Ready, Set, Cloud blog: https://readysetcloud.io
  • Tyler Technologies: https://www.tylertech.com/
  • Twitter: https://twitter.com/allenheltondev
  • Linked: https://www.linkedin.com/in/allenheltondev/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at AWS AppConfig. Engineers love to solve, and occasionally create, problems. But not when it’s an on-call fire-drill at 4 in the morning. Software problems should drive innovation and collaboration, NOT stress, and sleeplessness, and threats of violence. That’s why so many developers are realizing the value of AWS AppConfig Feature Flags. Feature Flags let developers push code to production, but hide that that feature from customers so that the developers can release their feature when it’s ready. This practice allows for safe, fast, and convenient software development. You can seamlessly incorporate AppConfig Feature Flags into your AWS or cloud environment and ship your Features with excitement, not trepidation and fear. To get started, go to snark.cloud/appconfig. That’s snark.cloud/appconfig.

Corey: I come bearing ill tidings. Developers are responsible for more than ever these days. Not just the code that they write, but also the containers and the cloud infrastructure that their apps run on. Because serverless means it’s still somebody’s problem. And a big part of that responsibility is app security from code to cloud. And that’s where our friend Snyk comes in. Snyk is a frictionless security platform that meets developers where they are - Finding and fixing vulnerabilities right from the CLI, IDEs, Repos, and Pipelines. Snyk integrates seamlessly with AWS offerings like code pipeline, EKS, ECR, and more! As well as things you’re actually likely to be using. Deploy on AWS, secure with Snyk. Learn more at Snyk.co/scream That’s S-N-Y-K.co/scream

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Every once in a while I wind up stumbling into corners of the internet that I previously had not traveled. Somewhat recently, I wound up having that delightful experience again by discovering readysetcloud.io, which has a whole series of, I guess some people might call it thought leadership, I’m going to call it instead how I view it, which is just amazing opinion pieces on the context of serverless, mixed with APIs, mixed with some prognostications about the future.

Allen Helton by day is a cloud architect at Tyler Technologies, but that’s not how I encountered you. First off, Allen, thank you for joining me.

Allen: Thank you, Corey. Happy to be here.

Corey: I was originally pointed towards your work by folks in the AWS Community Builder program, of which we both participate from time to time, and it’s one of those, “Oh, wow, this is amazing. I really wish I’d discovered some of this sooner.” And every time I look through your back catalog, and I click on a new post, I see things that are either I’ve really agree with this or I can’t stand this opinion, I want to fight about it, but more often than not, it’s one of those recurring moments that I love: “Damn, I wish I had written something like this.” So first, you’re absolutely killing it on the content front.

Allen: Thank you, Corey, I appreciate that. The content that I make is really about the stuff that I’m doing at work. It’s stuff that I’m passionate about, stuff that I’d spend a decent amount of time on, and really the most important thing about it for me, is it’s stuff that I’m learning and forming opinions on and wants to share with others.

Corey: I have to say, when I saw that you were—oh, your Tyler Technologies, which sounds for all the world like, oh, it’s a relatively small consultancy run by some guy presumably named Tyler, and you know, it’s a petite team of maybe 20, 30 people on the outside. Yeah, then I realized, wait a minute, that’s not entirely true. For example, for starters, you’re publicly traded. And okay, that does change things a little bit. First off, who are you people? Secondly, what do you do? And third, why have I never heard of you folks, until now?

Allen: Tyler is the largest company that focuses completely on the public sector. We have divisions and products for pretty much everything that you can imagine that’s in the public sector. We have software for schools, software for tax and appraisal, we have software for police officers, for courts, everything you can think of that runs the government can and a lot of times is run on Tyler software. We’ve been around for decades building our expertise in the domain, and the reason you probably haven’t heard about us is because you might not have ever been in trouble with the law before. If you [laugh] if you have been—

Corey: No, no, I learned very early on in the course of my life—which will come as a surprise to absolutely no one who spent more than 30 seconds with me—that I have remarkably little filter and if ten kids were the ones doing something wrong, I’m the one that gets caught. So, I spent a lot of time in the principal’s office, so this taught me to keep my nose clean. I’m one of those squeaky-clean types, just because I was always terrified of getting punished because I knew I would get caught. I’m not saying this is the right way to go through life necessarily, but it did have the side benefit of, no, I don’t really engage with law enforcement going throughout the course of my life.

Allen: That’s good. That’s good. But one exposure that a lot of people get to Tyler is if you look at the bottom of your next traffic ticket, it’ll probably say Tyler Technologies on the bottom there.

Corey: Oh, so you’re really popular in certain circles, I’d imagine?

Allen: Super popular. Yes, yes. And of course, you get all the benefits of writing that code that says ‘if defendant equals Allen Helton then return.’

Corey: I like that. You get to have the exception cases built in that no one’s ever going to wind up looking into.

Allen: That’s right. Yes.

Corey: The idea of what you’re doing makes an awful lot of sense. There’s a tremendous need for a wide variety of technical assistance in the public sector. What surprises me, although I guess it probably shouldn’t, is how much of your content is aimed at serverless technologies and API design, which to my way of thinking, isn’t really something that public sector has done a lot with. Clearly I’m wrong.

Allen: Historically, you’re not wrong. There’s an old saying that government tends to run about ten years behind on technology. Not just technology, but all over the board and runs about ten years behind. And until recently, that’s really been true. There was a case last year, a situation last year where one of the state governments—I don’t remember which one it was—but they were having a crisis because they couldn’t find any COBOL developers to come in and maintain their software that runs the state.

And it’s COBOL; you’re not going to find a whole lot of people that have that skill. A lot of those people are retiring out. And what’s happening is that we’re getting new people sitting in positions of power and government that want innovation. They know about the cloud and they want to be able to integrate with systems quickly and easily, have little to no onboarding time. You know, there are people in power that have grown up with technology and understand that, well, with everything else, I can be up and running in five or ten minutes. I cannot do this with the software I’m consuming now.

Corey: My opinion on it is admittedly conflicted because on the one hand, yeah, I don’t think that governments should be running on COBOL software that runs on mainframes that haven’t been supported in 25 years. Conversely, I also don’t necessarily want them being run like a seed series startup, where, “Well, I wrote this code last night, and it’s awesome, so off I go to production with it.” Because I can decide not to do business anymore with Twitter for Pets, and I could go on to something else, like PetFlicks, or whatever it is I choose to use. I can’t easily opt out of my government. The decisions that they make stick and that is going to have a meaningful impact on my life and everyone else’s life who is subject to their jurisdiction. So, I guess I don’t really know where I believe the proper, I guess, pace of technological adoption should be for governments. Curious to get your thoughts on this.

Allen: Well, you certainly don’t want anything that’s bleeding edge. That’s one of the things that we kind of draw fine lines around. Because when we’re dealing with government software, we’re dealing with, usually, critically sensitive information. It’s not medical records, but it’s your criminal record, and it’s things like your social security number, it’s things that you can’t have leaking out under any circumstances. So, the things that we’re building on are things that have proven out to be secure and have best practices around security, uptime, reliability, and in a lot of cases as well, and maintainability. You know, if there are issues, then let’s try to get those turned around as quickly as we can because we don’t want to have any sort of downtime from the software side versus the software vendor side.

Corey: I want to pivot a little bit to some of the content you’ve put out because an awful lot of it seems to be, I think I’ll call it variations on a theme. For example, I just read some recent titles, and to illustrate my point, “Going API First: Your First 30 Days,” “Solutions Architect Tips how to Design Applications for Growth,” “3 Things to Know Before Building A Multi-Tenant Serverless App.” And the common thread that I see running through all of these things are these are things that you tend to have extraordinarily strong and vocal opinions about only after dismissing all of them the first time and slapping something together, and then sort of being forced to live with the consequences of the choices that you’ve made, in some cases you didn’t realize you were making at the time. Are you one of those folks that has the wisdom to see what’s coming down the road, or did you do what the rest of us do and basically learn all this stuff by getting it hilariously wrong and having to careen into rebound situations as a result?

Allen: [laugh]. I love that question. I would like to say now, I feel like I have the vision to see something like that coming. Historically, no, not at all. Let me talk a little bit about how I got to where I am because that will shed a lot of context on that question.

A few years ago, I was put into a position at Tyler that said, “Hey, go figure out this cloud thing.” Let’s figure out what we need to do to move into the cloud safely, securely, quickly, all that rigmarole. And so, I did. I got to hand-select team of engineers from people that I worked with at Tyler over the past few years, and we were basically given free rein to learn. We were an R&D team, a hundred percent R&D, for about a year’s worth of time, where we were learning about cloud concepts and theory and building little proof of concepts.

CI/CD, serverless, APIs, multi-tenancy, a whole bunch of different stuff. NoSQL was another one of the things that we had to learn. And after that year of R&D, we were told, “Okay, now go do something with that. Go build this application.” And we did, building on our theory our cursory theory knowledge. And we get pretty close to go live, and then the business says, “What do you do in this scenario? What do you do in that scenario? What do you do here?”

Corey: “I update my resume and go work somewhere else. Where’s the hard part here?”

Allen: [laugh].

Corey: Turns out, that’s not a convincing answer.

Allen: Right. So, we moved quickly. And then I wouldn’t say we backpedaled, but we hardened for a long time before the—prior to the go-live, with the lessons that we’ve learned with the eyes of Tyler, the mature enterprise company, saying, “These are the things that you have to make sure that you take into consideration in an actual production application.” One of the things that I always pushed—I was a manager for a few years of all these cloud teams—I always push do it; do it right; do it better. Right?

It’s kind of like crawl, walk, run. And if you follow my writing from the beginning, just looking at the titles and reading them, kind of like what you were doing, Corey, you’ll see that very much. You’ll see how I talk about CI/CD, you’ll see me how I talk about authorization, you’ll see me how I talk about multi-tenancy. And I kind of go in waves where maybe a year passes and you see my content revisit some of the topics that I’ve done in the past. And they’re like, “No, no, no, don’t do what I said before. It’s not right.”

Corey: The problem when I’m writing all of these things that I do, for example, my entire newsletter publication pipeline is built on a giant morass of Lambda functions and API Gateways. It’s microservices-driven—kind of—and each microservice is built, almost always, with a different framework. Lately, all the new stuff is CDK. I started off with the serverless framework. There are a few other things here and there.

And it’s like going architecting, back in time as I have to make updates to these things from time to time. And it’s the problem with having done all that myself is that I already know the answer to, “What fool designed this?” It’s, well, you’re basically watching me learn what I was, doing bit by bit. I’m starting to believe that the right answer on some level, is to build an inherent shelf-life into some of these things. Great, in five years, you’re going to come back and re-architect it now that you know how this stuff actually works rather than patching together 15 blog posts by different authors, not all of whom are talking about the same thing and hoping for the best.

Allen: Yep. That’s one of the things that I really like about serverless, I view that as a giant pro of doing Serverless is that when we revisit with the lessons learned, we don’t have to refactor everything at once like if it was just a big, you know, MVC controller out there in the sky. We can refactor one Lambda function at a time if now we’re using a new version of the AWS SDK, or we’ve learned about a new best practice that needs to go in place. It’s a, “While you’re in there, tidy up, please,” kind of deal.

Corey: I know that the DynamoDB fanatics will absolutely murder me over this one, but one of the reasons that I have multiple Dynamo tables that contain, effectively, variations on the exact same data, is because I want to have the dependency between the two different microservices be the API, not, “Oh, and under the hood, it’s expecting this exact same data structure all the time.” But it just felt like that was the wrong direction to go in. That is the justification I use for myself why I run multiple DynamoDB tables that [laugh] have the same content. Where do you fall on the idea of data store separation?

Allen: I’m a big single table design person myself, I really like the idea of being able to store everything in the same table and being able to create queries that can return me multiple different types of entity with one lookup. Now, that being said, one of the issues that we ran into, or one of the ambiguous areas when we were getting started with serverless was, what does single table design mean when you’re talking about microservices? We were wondering does single table mean one DynamoDB table for an entire application that’s composed of 15 microservices? Or is it one table per microservice? And that was ultimately what we ended up going with is a table per microservice. Even if multiple microservices are pushed into the same AWS account, we’re still building that logical construct of a microservice and one table that houses similar entities in the same domain.

Corey: So, something I wish that every service team at AWS would do as a part of their design is draw the architecture of an application that you’re planning to build. Great, now assume that every single resource on that architecture diagram lives in its own distinct AWS account because somewhere in some customer, there’s going to be an account boundary at every interconnection point along the way. And so, many services don’t do that where it’s, “Oh, that thing and the other thing has to be in the same account.” So, people have to write their own integration shims, and it makes doing the right thing of putting different services into distinct bounded AWS accounts for security or compliance reasons way harder than I feel like it needs to be.

Allen: [laugh]. Totally agree with you on that one. That’s one of the things that I feel like I’m still learning about is the account-level isolation. I’m still kind of early on, personally, with my opinions in how we’re structuring things right now, but I’m very much of a like opinion that deploying multiple things into the same account is going to make it too easy to do something that you shouldn’t. And I just try not to inherently trust people, in the sense that, “Oh, this is easy. I’m just going to cross that boundary real quick.”

Corey: For me, it’s also come down to security risk exposure. Like my lasttweetinaws.com Twitter shitposting thread client lives in a distinct AWS account that is separate from the AWS account that has all of our client billing data that lives within it. The idea being that if you find a way to compromise my public-facing Twitter client, great, the blast radius should be constrained to, “Yay, now you can, I don’t know, spin up some cryptocurrency mining in my AWS account and I get to look like a fool when I beg AWS for forgiveness.”

But that should be the end of it. It shouldn’t be a security incident because I should not have the credit card numbers living right next to the funny internet web thing. That sort of flies in the face of the original guidance that AWS gave at launch. And right around 2008-era, best practices were one customer, one AWS account. And then by 2012, they had changed their perspective, but once you’ve made a decision to build multiple services in a single account, unwinding and unpacking that becomes an incredibly burdensome thing. It’s about the equivalent of doing a cloud migration, in some ways.

Allen: We went through that. We started off building one application with the intent that it was going to be a siloed application, a one-off, essentially. And about a year into it, it’s one of those moments of, “Oh, no. What we’re building is not actually a one-off. It’s a piece to a much larger puzzle.”

And we had a whole bunch of—unfortunately—tightly coupled things that were in there that we’re assuming that resources were going to be in the same AWS account. So, we ended up—how long—I think we took probably two months, which in the grand scheme of things isn’t that long, but two months, kind of unwinding the pieces and decoupling what was possible at the time into multiple AWS accounts, kind of, segmented by domain, essentially. But that’s hard. AWS puts it, you know, it’s those one-way door decisions. I think this one was a two-way door, but it locked and you could kind of jimmy the lock on the way back out.

Corey: And you could buzz someone from the lobby to let you back in. Yeah, the biggest problem is not necessarily the one-way door decisions. It’s the one-way door decisions that you don’t realize you’re passing through at the time that you do them. Which, of course, brings us to a topic near and dear to your heart—and I only recently started have opinions on this myself—and that is the proper design of APIs, which I’m sure will incense absolutely no one who’s listening to this. Like, my opinions on APIs start with well, probably REST is the right answer in this day and age. I had people, like, “Well, I don’t know, GraphQL is pretty awesome.” Like, “Oh, I’m thinking SOAP,” and people look at me like I’m a monster from the Black Lagoon of centuries past in XML-land. So, my particular brand of strangeness side, what do you see that people are doing in the world of API design that is the, I guess, most common or easy to make mistakes that you really wish they would stop doing?

Allen: If I could boil it down to one word, fundamentalism. Let me unpack that for you.

Corey: Oh, please, absolutely want to get a definition on that one.

Allen: [laugh]. I approach API design from a developer experience point of view: how easy is it for both internal and external integrators to consume and satisfy the business processes that they want to accomplish? And a lot of times, REST guidelines, you know, it’s all about entity basis, you know, drill into the appropriate entities and name your endpoints with nouns, not verbs. I’m actually very much onto that one.

But something that you could easily do, let’s say you have a business process that given a fundamentally correct RESTful API design takes ten API calls to satisfy. You could, in theory, boil that down to maybe three well-designed endpoints that aren’t, quote-unquote, “RESTful,” that make that developer experience significantly easier. And if you were a fundamentalist, that option is not even on the table, but thinking about it pragmatically from a developer experience point of view, that might be the better call. So, that’s one of the things that, I know feels like a hot take. Every time I say it, I get a little bit of flack for it, but don’t be a fundamentalist when it comes to your API designs. Do something that makes it easier while staying in the guidelines to do what you want.

Corey: For me the problem that I’ve kept smacking into with API design, and it honestly—let me be very clear on this—my first real exposure to API design rather than API consumer—which of course, I complain about constantly, especially in the context of the AWS inconsistent APIs between services—was when I’m building something out, and I’m reading the documentation for API Gateway, and oh, this is how you wind up having this stage linked to this thing, and here’s the endpoint. And okay, great, so I would just populate—build out a structure or a schema that has the positional parameters I want to use as variables in my function. And that’s awesome. And then I realized, “Oh, I might want to call this a different way. Aw, crap.” And sometimes it’s easy; you just add a different endpoint. Other times, I have to significantly rethink things. And I can’t shake the feeling that this is an entire discipline that exists that I just haven’t had a whole lot of exposure to previously.

Allen: Yeah, I believe that. One of the things that you could tie a metaphor to for what I’m saying and kind of what you’re saying, is AWS SAM, the Serverless Application Model, all it does is basically macros CloudFormation resources. It’s just a transform from a template into CloudFormation. CDK does same thing. But what the developers of SAM have done is they’ve recognized these business processes that people do regularly, and they’ve made these incredibly easy ways to satisfy those business processes and tie them all together, right?

If I want to have a Lambda function that is backed behind a endpoint, an API endpoint, I just have to add four or five lines of YAML or JSON that says, “This is the event trigger, here’s the route, here’s the API.” And then it goes and does four, five, six different things. Now, there’s some engineers that don’t like that because sometimes that feels like magic. Sometimes a little bit magic is okay.

Corey: This episode is sponsored in part by our friends at Sysdig. Sysdig secures your cloud from source to run. They believe, as do I, that DevOps and security are inextricably linked. If you wanna learn more about how they view this, check out their blog, it's definitely worth the read. To learn more about how they are absolutely getting it right from where I sit, visit Sysdig.com and tell them that I sent you. That's S Y S D I G.com. And my thanks to them for their continued support of this ridiculous nonsense.

Corey: I feel like one of the benefits I’ve had with the vast majority of APIs that I’ve built is that because this is all relatively small-scale stuff for what amounts to basically shitposting for the sake of entertainment, I’m really the only consumer of an awful lot of these things. So, I get frustrated when I have to backtrack and make changes and teach other microservices to talk to this thing that has now changed. And it’s frustrating, but I have the capacity to do that. It’s just work for a period of time. I feel like that equation completely shifts when you have published this and it is now out in the world, and it’s not just users, but in many cases paying customers where you can’t really make those changes without significant notice, and every time you do you’re creating work for those customers, so you have to be a lot more judicious about it.

Allen: Oh, yeah. There is a whole lot of governance and practice that goes into production-level APIs that people integrate with. You know, they say once you push something out the door into production that you’re going to support it forever. I don’t disagree with that. That seems like something that a lot of people don’t understand.

And that’s one of the reasons why I push API-first development so hard in all the content that I write is because you need to be intentional about what you’re letting out the door. You need to go in and work, not just with the developers, but your product people and your analysts to say, what does this absolutely need to do, and what does it need to do in the future? And you take those things, and you work with analysts who want specifics, you work with the engineers to actually build it out. And you’re very intentional about what goes out the door that first time because once it goes out with a mistake, you’re either going to version it immediately or you’re going to make some people very unhappy when you make a breaking change to something that they immediately started consuming.

Corey: It absolutely feels like that’s one of those things that AWS gets astonishingly right. I mean, I had the privilege of interviewing, at the time, Jeff Barr and then Ariel Kelman, who was their head of marketing, to basically debunk a bunch of old myths. And one thing that they started talking about extensively was the idea that an API is fundamentally a promise to your customers. And when you make a promise, you’d better damn well intend on keeping it. It’s why API deprecations from AWS are effectively unique whenever something happens.

It’s the, this is a singular moment in time when they turn off a service or degrade old functionality in favor of new. They can add to it, they can launch a V2 of something and then start to wean people off by calling the old one classic or whatnot, but if I built something on AWS in 2008 and I wound up sleeping until today, and go and try and do the exact same thing and deploy it now, it will almost certainly work exactly as it did back then. Sure, reliability is going to be a lot better and there’s a crap ton of features and whatnot that I’m not taking advantage of, but that fundamental ability to do that is awesome. Conversely, it feels like Google Cloud likes to change around a lot of their API stories almost constantly. And it’s unplanned work that frustrates the heck out of me when I’m trying to build something stable and lasting on top of it.

Allen: I think it goes to show the maturity of these companies as API companies versus just vendors. It’s one of the things that I think AWS does [laugh]—

Corey: You see the similar dichotomy with Microsoft and Apple. Microsoft’s new versions of Windows generally still have functionalities in them to support stuff that was written in the ’90s for a few use cases, whereas Apple’s like, “Oh, your computer’s more than 18-months old? Have you tried throwing it away and buying a new one? And oh, it’s a new version of Mac OS, so yeah, maybe the last one would get security updates for a year and then get with the times.” And I can’t shake the feeling that the correct answer is in some way, both of those, depending upon who your customer is and what it is you’re trying to achieve.

If Microsoft adopted the Apple approach, their customers would mutiny, and rightfully so; the expectation has been set for decades that isn’t what happens. Conversely, if Apple decided now we’re going to support this version of Mac OS in perpetuity, I don’t think a lot of their application developers wouldn’t quite know what to make of that.

Allen: Yeah. I think it also comes from a standpoint of you better make it worth their while if you’re going to move their cheese. I’m not a Mac user myself, but from what I hear for Mac users—and this could be rose-colored glasses—but is that their stuff works phenomenally well. You know, when a new thing comes out—

Corey: Until it doesn’t, absolutely. It’s—whenever I say things like that on this show, I get letters. And it’s, “Oh, yeah, really? They’ll come up with something that is a colossal pain in the ass on Mac.” Like, yeah, “Try building a system-wide mute key.”

It’s yeah, that’s just a hotkey away on windows and here in Mac land. It’s, “But it makes such beautiful sounds. Why would you want them to be quiet?” And it’s, yeah, it becomes this back-and-forth dichotomy there. And you can even explain it to iPhones as well and the Android ecosystem where it’s, oh, you’re going to support the last couple of versions of iOS.

Well, as a developer, I don’t want to do that. And Apple’s position is, “Okay, great.” Almost half of the mobile users on the planet will be upgrading because they’re in the ecosystem. Do you want us to be able to sell things those people are not? And they’re at a point of scale where they get to dictate those terms.

On some level, there are benefits to it and others, it is intensely frustrating. I don’t know what the right answer is on the level of permanence on that level of platform. I only have slightly better ideas around the position of APIs. I will say that when AWS deprecates something, they reach out individually to affected customers, on some level, and invariably, when they say, “This is going to be deprecated as of August 31,” or whenever it is, yeah, it is going to slip at least twice in almost every case, just because they’re not going to turn off a service that is revenue-bearing or critical-load-bearing for customers without massive amounts of notice and outreach, and in some cases according to rumor, having engineers reach out to help restructure things so it’s not as big of a burden on customers. That’s a level of customer focus that I don’t think most other companies are capable of matching.

Allen: I think that comes with the size and the history of Amazon. And one of the things that they’re doing right now, we’ve used Amazon Cloud Cams for years, in my house. We use them as baby monitors. And they—

Corey: Yea, I saw this I did something very similar with Nest. They didn’t have the Cloud Cam at the right time that I was looking at it. And they just announced that they’re going to be deprecating. They’re withdrawing them for sale. They’re not going to support them anymore. Which, oh at Amazon—we’re not offering this anymore. But you tell the story; what are they offering existing customers?

Allen: Yeah, so slightly upset about it because I like my Cloud Cams and I don’t want to have to take them off the wall or wherever they are to replace them with something else. But what they’re doing is, you know, they gave me—or they gave all the customers about eight months head start. I think they’re going to be taking them offline around Thanksgiving this year, just mid-November. And what they said is as compensation for you, we’re going to send you a Blink Cam—a Blink Mini—for every Cloud Cam that you have in use, and then we are going to gift you a year subscription to the Pro for Blink.

Corey: That’s very reasonable for things that were bought years ago. Meanwhile, I feel like not to be unkind or uncharitable here, but I use Nest Cams. And that’s a Google product. I half expected if they ever get deprecated, I’ll find out because Google just turns it off in the middle of the night—

Allen: [laugh].

Corey: —and I wake up and have to read a blog post somewhere that they put an update on Nest Cams, the same way they killed Google Reader once upon a time. That’s slightly unfair, but the fact that joke even lands does say a lot about Google’s reputation in this space.

Allen: For sure.

Corey: One last topic I want to talk with you about before we call it a show is that at the time of this recording, you recently had a blog post titled, “What does the Future Hold for Serverless?” Summarize that for me. Where do you see this serverless movement—if you’ll forgive the term—going?

Allen: So, I’m going to start at the end. I’m going to work back a little bit on what needs to happen for us to get there. I have a feeling that in the future—I’m going to be vague about how far in the future this is—that we’ll finally have a satisfied promise of all you’re going to write in the future is business logic. And what does that mean? I think what can end up happening, given the right focus, the right companies, the right feedback, at the right time, is we can write code as developers and have that get pushed up into the cloud.

And a phrase that I know Jeremy Daly likes to say ‘infrastructure from code,’ where it provisions resources in the cloud for you based on your use case. I’ve developed an application and it gets pushed up in the cloud at the time of deploying it, optimized resource allocation. Over time, what will happen—with my future vision—is when you get production traffic going through, maybe it’s spiky, maybe it’s consistently at a scale that outperforms the resources that it originally provisioned. We can have monitoring tools that analyze that and pick that out, find the anomalies, find the standard patterns, and adjust that infrastructure that it deployed for you automatically, where it’s based on your production traffic for what it created, optimizes it for you. Which is something that you can’t do on an initial deployment right now. You can put what looks best on paper, but once you actually get traffic through your application, you realize that, you know, what was on paper might not be correct.

Corey: You ever noticed that whiteboard diagrams never show the reality, and they’re always aspirational, and they miss certain parts? And I used to think that this was the symptom I had from working at small, scrappy companies because you know what, those big tech companies, everything they build is amazing and awesome. I know it because I’ve seen their conference talks. But I’ve been a consultant long enough now, and for a number of those companies, to realize that nope, everyone’s infrastructure is basically a trash fire at any given point in time. And it works almost in spite of itself, rather than because of it.

There is no golden path where everything is shiny, new and beautiful. And that, honestly, I got to say, it was really [laugh] depressing when I first discovered it. Like, oh, God, even these really smart people who are so intelligent they have to have extra brain packs bolted to their chests don’t have the magic answer to all of this. The rest of us are just screwed, then. But we find ways to make it work.

Allen: Yep. There’s a quote, I wish I remembered who said it, but it was a military quote where, “No battle plan survives impact with the enemy—first contact with the enemy.” It’s kind of that way with infrastructure diagrams. We can draw it out however we want and then you turn it on in production. It’s like, “Oh, no. That’s not right.”

Corey: I want to mix the metaphors there and say, yeah, no architecture survives your first fight with a customer. Like, “Great, I don’t think that’s quite what they’re trying to say.” It’s like, “What, you don’t attack your customers? Pfft, what’s your customer service line look like?” Yeah, it’s… I think you’re onto something.

I think that inherently everything beyond the V1 design of almost anything is an emergent property where this is what we learned about it by running it and putting traffic through it and finding these problems, and here’s how it wound up evolving to account for that.

Allen: I agree. I don’t have anything to add on that.

Corey: [laugh]. Fair enough. I really want to thank you for taking so much time out of your day to talk about how you view these things. If people want to learn more, where is the best place to find you?

Allen: Twitter is probably the best place to find me: @AllenHeltonDev. I have that username on all the major social platforms, so if you want to find me on LinkedIn, same thing: AllenHeltonDev. My blog is always open as well, if you have any feedback you’d like to give there: readysetcloud.io.

Corey: And we will, of course, put links to that in the show notes. Thanks again for spending so much time talking to me. I really appreciate it.

Allen: Yeah, this was fun. This was a lot of fun. I love talking shop.

Corey: It shows. And it’s nice to talk about things I don’t spend enough time thinking about. Allen Helton, cloud architect at Tyler Technologies. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry comment that I will reject because it was not written in valid XML.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Ian

Ian Smith is Field CTO at Chronosphere where he works across sales, marketing, engineering and product to deliver better insights and outcomes to observability teams supporting high-scale cloud-native environments. Previously, he worked with observability teams across the software industry in pre-sales roles at New Relic, Wavefront, PagerDuty and Lightstep.

Links Referenced:

  • Chronosphere: https://chronosphere.io
  • Last Tweet in AWS: lasttweetinaws.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Every once in a while, I find that something I’m working on aligns perfectly with a person that I wind up basically convincing to appear on this show. Today’s promoted guest is Ian Smith, who’s Field CTO at Chronosphere. Ian, thank you for joining me.

Ian: Thanks, Corey. Great to be here.

Corey: So, the coincidental aspect of what I’m referring to is that Chronosphere is, despite the name, not something that works on bending time, but rather an observability company. Is that directionally accurate?

Ian: That’s true. Although you could argue it probably bend a little bit of engineering time. But we can talk about that later.

Corey: [laugh]. So, observability is one of those areas that I think is suffering from too many definitions, if that makes sense. And at first, I couldn’t make sense of what it was that people actually meant when they said observability, this sort of clarified to me at least when I realized that there were an awful lot of, well, let’s be direct and call them ‘legacy monitoring companies’ that just chose to take what they were already doing and define that as, “Oh, this is observability.” I don’t know that I necessarily agree with that. I know a lot of folks in the industry vehemently disagree.

You’ve been in a lot of places that have positioned you reasonably well to have opinions on this sort of question. To my understanding, you were at interesting places, such as LightStep, New Relic, Wavefront, and PagerDuty, which I guess technically might count as observability in a very strange way. How do you view observability and what it is?

Ian: Yeah. Well, a lot of definitions, as you said, common ones, they talk about the three pillars, they talk really about data types. For me, it’s about outcomes. I think observability is really this transition from the yesteryear of monitoring where things were much simpler and you, sort of, knew all of the questions, you were able to define your dashboards, you were able to define your alerts and that was really the gist of it. And going into this brave new world where there’s a lot of unknown things, you’re having to ask a lot of sort of unique questions, particularly during a particular instance, and so being able to ask those questions in an ad hoc fashion layers on top of what we’ve traditionally done with monitoring. So, observability is sort of that more flexible, more dynamic kind of environment that you have to deal with.

Corey: This has always been something that, for me, has been relatively academic. Back when I was running production environments, things tended to be a lot more static, where, “Oh, there’s a problem with the database. I will SSH into the database server.” Or, “Hmm, we’re having a weird problem with the web tier. Well, there are ten or 20 or 200 web servers. Great, I can aggregate all of their logs to Syslog, and worst case, I can log in and poke around.”

Now, with a more ephemeral style of environment where you have Kubernetes or whatnot scheduling containers into place that have problems you can’t attach to a running container very easily, and by the time you see an error, that container hasn’t existed for three hours. And that becomes a problem. Then you’ve got the Lambda universe, which is a whole ‘nother world pain, where it becomes very challenging, at least for me, in order to reason using the old style approaches about what’s actually going on in your environment.

Ian: Yeah, I think there’s that and there’s also the added complexity of oftentimes you’ll see performance or behavioral changes based on even more narrow pathways, right? One particular user is having a problem and the traffic is spread across many containers. Is it making all of these containers perform badly? Not necessarily, but their user experience is being affected. It’s very common in say, like, B2B scenarios for you to want to understand the experience of one particular user or the aggregate experience of users at a particular company, particular customer, for example.

There’s just more complexity. There’s more complexity of the infrastructure and just the technical layer that you’re talking about, but there’s also more complexity in just the way that we’re handling use cases and trying to provide value with all of this software to the myriad of customers in different industries that software now serves.

Corey: For where I sit, I tend to have a little bit of trouble disambiguating, I guess, the three baseline data types that I see talked about again and again in observability. You have logs, which I think I’ve mostly I can wrap my head around. That seems to be the baseline story of, “Oh, great. Your application puts out logs. Of course, it’s in its own unique, beautiful format. Why wouldn’t it be?” In an ideal scenario, they’re structured. Things are never ideal, so great. You’re basically tailing log files in some cases. Great. I can reason about those.

Metrics always seem to be a little bit of a step beyond that. It’s okay, I have a whole bunch of log lines that are spitting out every 500 error that my app is throwing—and given my terrible code, it throws a lot—but I can then ideally count the number of times that appears and then that winds up incrementing counter, similar to the way that we used to see with StatsD, for example, and Collectd. Is that directionally correct? As far as the way I reason about, well so far, logs and metrics?

Ian: I think at a really basic level, yes. I think that, as we’ve been talking about, sort of greater complexity starts coming in when you have—particularly metrics in today’s world of containers—Prometheus—you mentioned StatsD—Prometheus has become sort of like the standard for expressing those things, so you get situations where you have incredibly high cardinality, so cardinality being the interplay between all the different dimensions. So, you might have, my container is a label, but also the type of endpoint is running on that container as a label, then maybe I want to track my customer organizations and maybe I have 5000 of those. I have 3000 containers, and so on and so forth. And you get this massive explosion, almost multiplicatively.

For those in the audience who really live and read cardinality, there’s probably someone screaming about well, it’s not truly multiplicative in every sense of the word, but, you know, it’s close enough from an approximation standpoint. As you get this massive explosion of data, which obviously has a cost implication but also has, I think, a really big implication on the core reason why you have metrics in the first place you alluded to, which is, so a human being can reason about it, right? You don’t want to go and look at 5000 log lines; you want to know, out of those 5000 log lines of 4000 errors and I have 1000, OKs. It’s very easy for human beings to reason about that from a numbers perspective. When your metrics start to re-explode out into thousands, millions of data points, and unique sort of time series more numbers for you to track, then you’re sort of losing that original goal of metrics.

Corey: I think I mostly have wrapped my head around the concept. But then that brings us to traces, and that tends to be I think one of the hardest things for me to grasp, just because most of the apps I build, for obvious reasons—namely, I’m bad at programming and most of these are proof of concept type of things rather than anything that’s large scale running in production—the difference between a trace and logs tends to get very muddled for me. But the idea being that as you have a customer session or a request that talks to different microservices, how do you collate across different systems all of the outputs of that request into a single place so you can see timing information, understand the flow that user took through your application? Is that again, directionally correct? Have I completely missed the plot here? Which is again, eminently possible. You are the expert.

Ian: No, I think that’s sort of the fundamental premise or expected value of tracing, for sure. We have something that’s akin to a set of logs; they have a common identifier, a trace ID, that tells us that all of these logs essentially belong to the same request. But importantly, there’s relationship information. And this is the difference between just having traces—sorry, logs—with just a trace ID attached to them. So, for example, if you have Service A calling Service B and Service C, the relatively simple thing, you could use time to try to figure this out.

But what if there are things happening in Service B at the same time there are things happening in Service C and D, and so on and so forth? So, one of the things that tracing brings to the table is it tells you what is currently happening, what called that. So oh, I know that I’m Service D. I was actually called by Service B and I’m not just relying on timestamps to try and figure out that connection. So, you have that information and ultimately, the data model allows you to fully sort of reflect what’s happening with the request, particularly in complex environments.

And I think this is where, you know, tracing needs to be sort of looked at as not a tool for—just because I’m operating in a modern environment, I’m using some Kubernetes, or I’m using Lambda, is it needs to be used in a scenario where you really have troubles grasping, from a conceptual standpoint, what is happening with the request because you need to actually fully document it. As opposed to, I have a few—let’s say three Lambda functions. I maybe have some key metrics about them; I have a little bit of logging. You probably do not need to use tracing to solve, sort of, basic performance problems with those. So, you can get yourself into a place where you’re over-engineering, you’re spending a lot of time with tracing instrumentation and tracing tooling, and I think that’s the core of observability is, like, using the right tool, the right data for the job.

But that’s also what makes it really difficult because you essentially need to have this, you know, huge set of experience or knowledge about the different data, the different tooling, and what influential architecture and the data you have available to be able to reason about that and make confident decisions, particularly when you’re under a time crunch which everyone is familiar with a, sort of like, you know, PagerDuty-style experience of my phone is going off and I have a customer-facing incident. Where is my problem? What do I need to do? Which dashboard do I need to look at? Which tool do I need to investigate? And that’s where I think the observability industry has become not serving the outcomes of the customers.

Corey: I had a, well, I wouldn’t say it’s a genius plan, but it was a passing fancy that I’ve built this online, freely available Twitter client for authoring Twitter threads—because that’s what I do is that of having a social life—and it’s available at lasttweetinaws.com. I’ve used that as a testbed for a few things. It’s now deployed to roughly 20 AWS regions simultaneously, and this means that I have a bit of a problem as far as how to figure out not even what’s wrong or what’s broken with this, but who’s even using it?

Because I know people are. I see invocations all over the planet that are not me. And sometimes it appears to just be random things crawling the internet—fine, whatever—but then I see people logging in and doing stuff with it. I’d kind of like to log and see who’s using it just so I can get information like, is there anyone I should talk to about what it could be doing differently? I love getting user experience reports on this stuff.

And I figured, ah, this is a perfect little toy application. It runs in a single Lambda function so it’s not that complicated. I could instrument this with OpenTelemetry, which then, at least according to the instructions on the tin, I could then send different types of data to different observability tools without having to re-instrument this thing every time I want to kick the tires on something else. That was the promise.

And this led to three weeks of pain because it appears that for all of the promise that it has, OpenTelemetry, particularly in a Lambda environment, is nowhere near ready for being able to carry a workload like this. Am I just foolish on this? Am I stating an unfortunate reality that you’ve noticed in the OpenTelemetry space? Or, let’s be clear here, you do work for a company with opinions on these things. Is OpenTelemetry the wrong approach?

Ian: I think OpenTelemetry is absolutely the right approach. To me, the promise of OpenTelemetry for the individual is, “Hey, I can go and instrument this thing, as you said and I can go and send the data, wherever I want.” The sort of larger view of that is, “Well, I’m no longer beholden to a vendor,”—including the ones that I’ve worked for, including the one that I work for now—“For the definition of the data. I am able to control that, I’m able to choose that, I’m able to enhance that, and any effort I put into it, it’s mine. I own that.”

Whereas previously, if you picked, say, for example, an APM vendor, you said, “Oh, I want to have some additional aspects of my information provider, I want to track my customer, or I want to track a particular new metric of how much dollars am I transacting,” that effort really going to support the value of that individual solution, it’s not going to support your outcomes. Which is I want to be able to use this data wherever I want, wherever it’s most valuable. So, the core premise of OpenTelemetry, I think, is great. I think it’s a massive undertaking to be able to do this for at least three different data types, right? Defining an API across a whole bunch of different languages, across three different data types, and then creating implementations for those.

Because the implementations are the thing that people want, right? You are hoping for the ability to, say, drop in something. Maybe one line of code or preferably just, like, attach a dependency, let’s say in Java-land at runtime, and be able to have the information flow through and have it complete. And this is the premise of, you know, vendors I’ve worked with in the past, like New Relic. That was what New Relic built on: the ability to drop in an agent and get visibility immediately.

So, having that out-of-the-box visibility is obviously a goal of OpenTelemetry where it makes sense—Go, it’s very difficult to attach things at runtime, for example—but then saying, well, whatever is provided—let’s say your gRPC connections, database, all these things—well, now I want to go and instrument; I want to add some additional value. As you said, maybe you want to track something like I want to have in my traces the email address of whoever it is or the Twitter handle of whoever is so I can then go and analyze that stuff later. You want to be able to inject that piece of information or that instrumentation and then decide, well, where is the best utilized? Is it best utilized in some tooling from AWS? Is it best utilized in something that you’ve built yourself? Is it best of utilized an open-source project? Is it best utilized in one of the many observability vendors, or is even becoming more common, I want to shove everything in a data lake and run, sort of, analysis asynchronously, overlay observability data for essentially business purposes.

All of those things are served by having a very robust, open-source standard, and simple-to-implement way of collecting a really good baseline of data and then make it easy for you to then enhance that while still owning—essentially, it’s your IP right? It’s like, the instrumentation is your IP, whereas in the old world of proprietary agents, proprietary APIs, that IP was basically building it, but it was tied to that other vendor that you were investing in.

Corey: One thing that I was consistently annoyed by in my days of running production infrastructures at places, like, you know, large banks, for example, one of the problems I kept running into is that this, there’s this idea that, “Oh, you want to use our tool. Just instrument your applications with our libraries or our instrumentation standards.” And it felt like I was constantly doing and redoing a lot of instrumentation for different aspects. It’s not that we were replacing one vendor with another; it’s that in an observability, toolchain, there are remarkably few, one-size-fits-all stories. It feels increasingly like everyone’s trying to sell me a multifunction printer, which does one thing well, and a few other things just well enough to technically say they do them, but badly enough that I get irritated every single time.

And having 15 different instrumentation packages in an application, that’s either got security ramifications, for one, see large bank, and for another it became this increasingly irritating and obnoxious process where it felt like I was spending more time seeing the care and feeding of the instrumentation then I was the application itself. That’s the gold—that’s I guess the ideal light at the end of the tunnel for me in what OpenTelemetry is promising. Instrument once, and then you’re just adjusting configuration as far as where to send it.

Ian: That’s correct. The organization’s, and you know, I keep in touch with a lot of companies that I’ve worked with, companies that have in the last two years really invested heavily in OpenTelemetry, they’re definitely getting to the point now where they’re generating the data once, they’re using, say, pieces of the OpenTelemetry pipeline, they’re extending it themselves, and then they’re able to shove that data in a bunch of different places. Maybe they’re putting in a data lake for, as I said, business analysis purposes or forecasting. They may be putting the data into two different systems, even for incident and analysis purposes, but you’re not having that duplication effort. Also, potentially that performance impact, right, of having two different instrumentation packages lined up with each other.

Corey: There is a recurring theme that I’ve noticed in the observability space that annoys me to no end. And that is—I don’t know if it’s coming from investor pressure, from folks never being satisfied with what they have, or what it is, but there are so many startups that I have seen and worked with in varying aspects of the observability space that I think, “This is awesome. I love the thing that they do.” And invariably, every time they start getting more and more features bolted onto them, where, hey, you love this whole thing that winds up just basically doing a tail-F on a log file, so it just streams your logs in the application and you can look for certain patterns. I love this thing. It’s great.

Oh, what’s this? Now, it’s trying to also be the thing that alerts me and wakes me up in the middle of the night. No. That’s what PagerDuty does. I want PagerDuty to do that thing, and I want other things—I want you just to be the log analysis thing and the way that I contextualize logs. And it feels like they keep bolting things on and bolting things on, where everything is more or less trying to evolve into becoming its own version of Datadog. What’s up with that?

Ian: Yeah, the sort of, dreaded platform play. I—[laugh] I was at New Relic when there were essentially two products that they sold. And then by the time I left, I think there was seven different products that were being sold, which is kind of a crazy, crazy thing when you think about it. And I think Datadog has definitely exceeded that now. And I definitely see many, many vendors in the market—and even open-source solutions—sort of presenting themselves as, like, this integrated experience.

But to your point, even before about your experience of these banks it oftentimes become sort of a tick-a-box feature approach of, “Hey, I can do this thing, so buy more. And here’s a shared navigation panel.” But are they really integrated? Like, are you getting real value out of it? One of the things that I do in my role is I get to work with our internal product teams very closely, particularly around new initiatives like tracing functionality, and the constant sort of conversation is like, “What is the outcome? What is the value?”

It’s not about the feature; it’s not about having a list of 19 different features. It’s like, “What is the user able to do with this?” And so, for example, there are lots of platforms that have metrics, logs, and tracing. The new one-upmanship is saying, “Well, we have events as well. And we have incident response. And we have security. And all these things sort of tie together, so it’s one invoice.”

And constantly I talk to customers, and I ask them, like, “Hey, what are the outcomes that you’re getting when you’ve invested so heavily in one vendor?” And oftentimes, the response is, “Well, I only need to deal with one vendor.” Okay, but that’s not an outcome. [laugh]. And it’s like the business having a single invoice.

Corey: Yeah, that is something that’s already attainable today. If you want to just have one vendor with a whole bunch of crappy offerings, that’s what AWS is for. They have AmazonBasics versions of everything you might want to use in production. Oh, you want to go ahead and use MongoDB? Well, use AmazonBasics MongoDB, but they call it DocumentDB because of course they do. And so, on and so forth.

There are a bunch of examples of this, but those companies are still in business and doing very well because people often want the genuine article. If everyone was trying to do just everything to check a box for procurement, great. AWS has already beaten you at that game, it seems.

Ian: I do think that, you know, people are hoping for that greater value and those greater outcomes, so being able to actually provide differentiation in that market I don’t think is terribly difficult, right? There are still huge gaps in let’s say, root cause analysis during an investigation time. There are huge issues with vendors who don’t think beyond sort of just the one individual who’s looking at a particular dashboard or looking at whatever analysis tool there is. So, getting those things actually tied together, it’s not just, “Oh, we have metrics, and logs, and traces together,” but even if you say we have metrics and tracing, how do you move between metrics and tracing? One of the goals in the way that we’re developing product at Chronosphere is that if you are alerted to an incident—you as an engineer; doesn’t matter whether you are massively sophisticated, you’re a lead architect who has been with the company forever and you know everything or you’re someone who’s just come out of onboarding and is your first time on call—you should not have to think, “Is this a tracing problem, or a metrics problem, or a logging problem?”

And this is one of those things that I mentioned before of requiring that really heavy level of knowledge and understanding about the observability space and your data and your architecture to be effective. And so, with the, you know, particularly observability teams and all of the engineers that I speak with on a regular basis, you get this sort of circumstance where well, I guess, let’s talk about a real outcome and a real pain point because people are like, okay, yeah, this is all fine; it’s all coming from a vendor who has a particular agenda, but the thing that constantly resonates is for large organizations that are moving fast, you know, big startups, unicorns, or even more traditional enterprises that are trying to undergo, like, a rapid transformation and go really cloud-native and make sure their engineers are moving quickly, a common question I will talk about with them is, who are the three people in your organization who always get escalated to? And it’s usually, you know, between two and five people—

Corey: And you can almost pick those perso—you say that and you can—at least anyone who’s worked in environments or through incidents like this more than a few times, already have thought of specific people in specific companies. And they almost always fall into some very predictable archetypes. But please, continue.

Ian: Yeah. And people think about these people, they always jump to mind. And one of the things I asked about is, “Okay, so when you did your last innovation around observably”—it’s not necessarily buying a new thing, but it maybe it was like introducing a new data type or it was you’re doing some big investment in improving instrumentation—“What changed about their experience?” And oftentimes, the most that can come out is, “Oh, they have access to more data.” Okay, that’s not great.

It’s like, “What changed about their experience? Are they still getting woken up at 3 am? Are they constantly getting pinged all the time?” One of the vendors that I worked at, when they would go down, there were three engineers in the company who were capable of generating list of customers who are actually impacted by damage. And so, every single incident, one of those three engineers got paged into the incident.

And it became borderline intolerable for them because nothing changed. And it got worse, you know? The platform got bigger and more complicated, and so there were more incidents and they were the ones having to generate that. But from a business level, from an observability outcomes perspective, if you zoom all the way up, it’s like, “Oh, were we able to generate the list of customers?” “Yes.”

And this is where I think the observability industry has sort of gotten stuck—you know, at least one of the ways—is that, “Oh, can you do it?” “Yes.” “But is it effective?” “No.” And by effective, I mean those three engineers become the focal point for an organization.

And when I say three—you know, two to five—it doesn’t matter whether you’re talking about a team of a hundred or you’re talking about a team of a thousand. It’s always the same number of people. And as you get bigger and bigger, it becomes more and more of a problem. So, does the tooling actually make a difference to them? And you might ask, “Well, what do you expect from the tooling? What do you expect to do for them?” Is it you give them deeper analysis tools? Is it, you know, you do AI Ops? No.

The answer is, how do you take the capabilities that those people have and how do you spread it across a larger population of engineers? And that, I think, is one of those key outcomes of observability that no one, whether it be in open-source or the vendor side is really paying a lot of attention to. It’s always about, like, “Oh, we can just shove more data in. By the way, we’ve got petabyte scale and we can deal with, you know, 2 billion active time series, and all these other sorts of vanity measures.” But we’ve gotten really far away from the outcomes. It’s like, “Am I getting return on investment of my observability tooling?”

And I think tracing is this—as you’ve said, it can be difficult to reason about right? And people are not sure. They’re feeling, “Well, I’m in a microservices environment; I’m in cloud-native; I need tracing because my older APM tools appear to be failing me. I’m just going to go and wriggle my way through implementing OpenTelemetry.” Which has significant engineering costs. I’m not saying it’s not worth it, but there is a significant engineering cost—and then I don’t know what to expect, so I’m going to go on through my data somewhere and see whether we can achieve those outcomes.

And I do a pilot and my most sophisticated engineers are in the pilot. And they’re able to solve the problems. Okay, I’m going to go buy that thing. But I’ve just transferred my problems. My engineers have gone from solving problems in maybe logs and grepping through petabytes worth of logs to using some sort of complex proprietary query language to go through your tens of petabytes of trace data but actually haven’t solved any problem. I’ve just moved it around and probably just cost myself a lot, both in terms of engineering time and real dollars spent as well.

Corey: One of the challenges that I’m seeing across the board is that observability, for certain use cases, once you start to see what it is and its potential for certain applications—certainly not all; I want to hedge that a little bit—but it’s clear that there is definite and distinct value versus other ways of doing things. The problem is, is that value often becomes apparent only after you’ve already done it and can see what that other side looks like. But let’s be honest here. Instrumenting an application is going to take some significant level of investment, in many cases. How do you wind up viewing any return on investment that it takes for the very real cost, if only in people’s time, to go ahead instrumenting for observability in complex environments?

Ian: So, I think that you have to look at the fundamentals, right? You have to look at—pretend we knew nothing about tracing. Pretend that we had just invented logging, and you needed to start small. It’s like, I’m not going to go and log everything about every application that I’ve had forever. What I need to do is I need to find the points where that logging is going to be the most useful, most impactful, across the broadest audience possible.

And one of the useful things about tracing is because it’s built in distributed environments, primarily for distributed environments, you can look at, for example, the biggest intersection of requests. A lot of people have things like API Gateways, or they have parts of a monolith which is still handling a lot of requests routing; those tend to be areas to start digging into. And I would say that, just like for anyone who’s used Prometheus or decided to move away from Prometheus, no one’s ever gone and evaluated Prometheus solution without having some sort of Prometheus data, right? You don’t go, “Hey, I’m going to evaluate a replacement for Prometheus or my StatsD without having any data, and I’m simultaneously going to generate my data and evaluate the solution at the same time.” It doesn’t make any sense.

With tracing, you have decent open-source projects out there that allow you to visualize individual traces and understand sort of the basic value you should be getting out of this data. So, it’s a good starting point to go, “Okay, can I reason about a single request? Can I go and look at my request end-to-end, even in a relatively small slice of my environment, and can I see the potential for this? And can I think about the things that I need to be able to solve with many traces?” Once you start developing these ideas, then you can have a better idea of, “Well, where do I go and invest more in instrumentation? Look, databases never appear to be a problem, so I’m not going to focus on database instrumentation. What’s the real problem is my external dependencies. Facebook API is the one that everyone loves to use. I need to go instrument that.”

And then you start to get more clarity. Tracing has this interesting network effect. You can basically just follow the breadcrumbs. Where is my biggest problem here? Where are my errors coming from? Is there anything else further down the call chain? And you can sort of take that exploratory approach rather than doing everything up front.

But it is important to do something before you start trying to evaluate what is my end state. End state obviously being sort of nebulous term in today’s world, but where do I want to be in two years’ time? I would like to have a solution. Maybe it’s open-source solution, maybe it’s a vendor solution, maybe it’s one of those platform solutions we talked about, but how do I get there? It’s really going to be I need to take an iterative approach and I need to be very clear about the value and outcomes.

There’s no point in doing a whole bunch of instrumentation effort in things that are just working fine, right? You want to go and focus your time and attention on that. And also you don’t want to go and burn just singular engineers. The observability team’s purpose in life is probably not to just write instrumentation or just deploy OpenTelemetry. Because then we get back into the land where engineers themselves know nothing about the monitoring or observability they’re doing and it just becomes a checkbox of, “I dropped in an agent. Oh, when it comes time for me to actually deal with an incident, I don’t know anything about the data and the data is insufficient.”

So, a level of ownership supported by the observability team is really important. On that return on investment, sort of, though it’s not just the instrumentation effort. There’s product training and there are some very hard costs. People think oftentimes, “Well, I have the ability to pay a vendor; that’s really the only cost that I have.” There’s things like egress costs, particularly volumes of data. There’s the infrastructure costs. A lot of the times there will be elements you need to run in your own environment; those can be very costly as well, and ultimately, they’re sort of icebergs in this overall ROI conversation.

The other side of it—you know, return and investment—return, there’s a lot of difficulty in reasoning about, as you said, what is the value of this going to be if I go through all this effort? Everyone knows a sort of, you know, meme or archetype of, “Hey, here are three options; pick two because there’s always going to be a trade off.” Particularly for observability, it’s become an element of, I need to pick between performance, data fidelity, or cost. Pick two. And when data fidelity—particularly in tracing—I’m talking about the ability to not sample, right?

If you have edge cases, if you have narrow use cases and ways you need to look at your data, if you heavily sample, you lose data fidelity. But oftentimes, cost is a reason why you do that. And then obviously, performance as you start to get bigger and bigger datasets. So, there’s a lot of different things you need to balance on that return. As you said, oftentimes you don’t get to understand the magnitude of those until you’ve got the full data set in and you’re trying to do this, sort of, for real. But being prepared and iterative as you go through this effort and not saying, “Okay, well, I’m just going to buy everything from one vendor because I’m going to assume that’s going to solve my problem,” is probably that undercurrent there.

Corey: As I take a look across the entire ecosystem, I can’t shake the feeling—and my apologies in advance if this is an observation, I guess, that winds up throwing a stone directly at you folks—

Ian: Oh, please.

Corey: But I see that there’s a strong observability community out there that is absolutely aligned with the things I care about and things I want to do, and then there’s a bunch of SaaS vendors, where it seems that they are, in many cases, yes, advancing the state of the art, I am not suggesting for a second that money is making observability worse. But I do think that when the tool you sell is a hammer, then every problem starts to look like a nail—or in my case, like my thumb. Do you think that there’s a chance that SaaS vendors are in some ways making this entire space worse?

Ian: As we’ve sort of gone into more cloud-native scenarios and people are building things specifically to take advantage of cloud from a complexity standpoint, from a scaling standpoint, you start to get, like, vertical issues happening. So, you have things like we’re going to charge on a per-container basis; we’re going to charge on a per-host basis; we’re going to charge based off the amount of gigabytes that you send us. These are sort of like more horizontal pricing models, and the way the SaaS vendors have delivered this is they’ve made it pretty opaque, right? Everyone has experiences, or has jerks about overages from observability vendors’ massive spikes. I’ve worked with customers who have used—accidentally used some features and they’ve been billed a quarter million dollars on a monthly basis for accidental overages from a SaaS vendor.

And these are all terrible things. Like, but we’ve gotten used to this. Like, we’ve just accepted it, right, because everyone is operating this way. And I really do believe that the move to SaaS was one of those things. Like, “Oh, well, you’re throwing us more data, and we’re charging you more for it.” As a vendor—

Corey: Which sort of erodes your own value proposition that you’re bringing to the table. I mean, I don’t mean to be sitting over here shaking my fist yelling, “Oh, I could build a better version in a weekend,” except that I absolutely know how to build a highly available Rsyslog cluster. I’ve done it a handful of times already and the technology is still there. Compare and contrast that with, at scale, the fact that I’m paying 50 cents per gigabyte ingested to CloudWatch logs, or a multiple of that for a lot of other vendors, it’s not that much harder for me to scale that fleet out and pay a much smaller marginal cost.

Ian: And so, I think the reaction that we’re seeing in the market and we’re starting to see—we’re starting to see the rise of, sort of, a secondary class of vendor. And by secondary, I don’t mean that they’re lesser; I mean that they’re, sort of like, specifically trying to address problems of the primary vendors, right? Everyone’s aware of vendors who are attempting to reduce—well, let’s take the example you gave on logs, right? There are vendors out there whose express purpose is to reduce the cost of your logging observability. They just sit in the middle; they are a middleman, right?

Essentially, hey, use our tool and even though you’re going to pay us a whole bunch of money, it’s going to generate an overall return that is greater than if you had just continued pumping all of your logs over to your existing vendor. So, that’s great. What we think really needs to happen, and one of the things we’re doing at Chronosphere—unfortunate plug—is we’re actually building those capabilities into the solution so it’s actually end-to-end. And by end-to-end, I mean, a solution where I can ingest my data, I can preprocess my data, I can store it, query it, visualize it, all those things, aligned with open-source standards, but I have control over that data, and I understand what’s going on with particularly my cost and my usage. I don’t just get a bill at the end of the month going, “Hey, guess what? You’ve spent an additional $200,000.”

Instead, I can know in real time, well, what is happening with my usage. And I can attribute it. It’s this team over here. And it’s because they added this particular label. And here’s a way for you, right now, to address that and cap it so it doesn’t cost you anything and it doesn’t have a blast radius of, you know, maybe degraded performance or degraded fidelity of the data.

That though is diametrically opposed to the way that most vendors are set up. And unfortunately, the open-source projects tend to take a lot of their cues, at least recently, from what’s happening in the vendor space. One of the ways that you can think about it is a sort of like a speed of light problem. Everyone knows that, you know, there’s basic fundamental latency; everyone knows how fast disk is; everyone knows the, sort of like, you can’t just make your computations happen magically, there’s a cost of running things horizontally. But a lot of the way that the vendors have presented efficiency to the market is, “Oh, we’re just going to incrementally get faster as AWS gets faster. We’re going to incrementally get better as compression gets better.”

And of course, you can’t go and fit a petabyte worth of data into a kilobyte, unless you’re really just doing some sort of weird dictionary stuff, so you feel—you’re dealing with some fundamental constraints. And the vendors just go, “I’m sorry, you know, we can’t violate the speed of light.” But what you can do is you can start taking a look at, well, how is the data valuable, and start giving the people controls on how to make it more valuable. So, one of the things that we do with Chronosphere is we allow you to reshape Prometheus metrics, right? You go and express Prometheus metrics—let’s say it’s a business metric about how many transactions you’re doing as a business—you don’t need that on a per-container basis, particularly if you’re running 100,000 containers globally.

When you go and take a look at that number on a dashboard, or you alert on it, what is it? It's one number, one time series. Maybe you break it out per region. You have five regions, you don’t need 100,000 data points every minute behind that. It’s very expensive, it’s not very performant, and as we talked about earlier, it’s very hard to reason about as a human being.

So, giving the tools to be able to go and condense that data down and make it more actionable and more valuable, you get performance, you get cost reduction, and you get the value that you ultimately need out of the data. And it’s one of the reasons why, I guess, I work at Chronosphere. Which I’m hoping is the last observability [laugh] venture I ever work for.

Corey: Yeah, for me a lot of the data that I see in my logs, which is where a lot of this stuff starts and how I still contextualize these things, is nonsense that I don’t care about and will never care about. I don’t care about load balance or health checks. I don’t particularly care about 200 results for the favicon when people visit the site. I care about other things, but just weed out the crap, especially when I’m paying by the pound—or at least by the gigabyte—in order to get that data into something. Yeah. It becomes obnoxious and difficult to filter out.

Ian: Yeah. And the vendors just haven’t done any of that because why would they, right? If you went and reduced the amount of log—

Corey: Put engineering effort into something that reduces how much I can charge you? That sounds like lunacy. Yeah.

Ian: Exactly. They’re business models entirely based off it. So, if you went and reduced every one’s logging bill by 30%, or everyone’s logging volume by 30% and reduced the bills by 30%, it’s not going to be a great time if you’re a publicly traded company who has built your entire business model on essentially a very SaaS volume-driven—and in my eyes—relatively exploitative pricing and billing model.

Corey: Ian, I want to thank you for taking so much time out of your day to talk to me about this. If people want to learn more, where can they find you? I mean, you are a Field CTO, so clearly you’re outstanding in your field. But if, assuming that people don’t want to go to farm country, where’s the best place to find you?

Ian: Yeah. Well, it’ll be a bunch of different conferences. I’ll be at KubeCon this year. But chronosphere.io is the company website. I’ve had the opportunity to talk to a lot of different customers, not from a hard sell perspective, but you know, conversations like this about what are the real problems you’re having and what are the things that you sort of wish that you could do?

One of the favorite things that I get to ask people is, “If you could wave a magic wand, what would you love to be able to do with your observability solution?” That’s, A, a really great part, but oftentimes be being able to say, “Well, actually, that thing you want to do, I think I have a way to accomplish that,” is a really rewarding part of this particular role.

Corey: And we will, of course, put links to that in the show notes. Thank you so much for being so generous with your time. I appreciate it.

Ian: Thanks, Corey. It’s great to be here.

Corey: Ian Smith, Field CTO at Chronosphere on this promoted guest episode. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry comment, which going to be super easy in your case, because it’s just one of the things that the omnibus observability platform that your company sells offers as part of its full suite of things you’ve never used.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Dan

Dan Moore is head of developer relations for FusionAuth, where he helps share information about authentication, authorization and security with developers building all kinds of applications.

A former CTO, AWS certification instructor, engineering manager and a longtime developer, he's been writing software for (checks watch) over 20 years.

Links Referenced:

  • FusionAuth: https://fusionauth.io
  • Twitter: https://twitter.com/mooreds

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at AWS AppConfig. Engineers love to solve, and occasionally create, problems. But not when it’s an on-call fire-drill at 4 in the morning. Software problems should drive innovation and collaboration, NOT stress, and sleeplessness, and threats of violence. That’s why so many developers are realizing the value of AWS AppConfig Feature Flags. Feature Flags let developers push code to production, but hide that that feature from customers so that the developers can release their feature when it’s ready. This practice allows for safe, fast, and convenient software development. You can seamlessly incorporate AppConfig Feature Flags into your AWS or cloud environment and ship your Features with excitement, not trepidation and fear. To get started, go to snark.cloud/appconfig. That’s snark.cloud/appconfig.

Corey: This episode is sponsored in part by our friends at Sysdig. Sysdig secures your cloud from source to run. They believe, as do I, that DevOps and security are inextricably linked. If you wanna learn more about how they view this, check out their blog, it's definitely worth the read. To learn more about how they are absolutely getting it right from where I sit, visit Sysdig.com and tell them that I sent you. That's S Y S D I G.com. And my thanks to them for their continued support of this ridiculous nonsense.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I am joined today on this promoted episode, which is brought to us by our friends at FusionAuth by Dan Moore, who is their head of DevRel at same. Dan, thank you for joining me.

Dan: Corey, thank you so much for having me.

Corey: So, you and I have been talking for a while. I believe it predates not just you working over at FusionAuth but me even writing the newsletter and the rest. We met on a leadership Slack many years ago. We’ve kept in touch ever since, and I think, I haven’t run the actual numbers on this, but I believe that you are at the top of the leaderboard right now for the number of responses I have gotten to various newsletter issues that I’ve sent out over the years.

And it’s always something great. It’s “Here’s a link I found that I thought that you might appreciate.” And we finally sat down and met each other in person, had a cup of coffee somewhat recently, and the first thing you asked was, “Is it okay that I keep doing this?” And at the bottom of the newsletter is “Hey, if you’ve seen something interesting, hit reply and let me know.” And you’d be surprised how few people actually take me up on it. So, let me start by thanking you for being as enthusiastic a contributor of the content as you have been.

Dan: Well, I appreciate that. And I remember the first time I ran across your newsletter and was super impressed by kind of the breadth of it. And I guess my way of thanking you is to just send you interesting tidbits that I run across. And it’s always fun when I see one of the links that I sent go into the newsletter because what you provide is just such a service to the community. So, thank you.

Corey: The fun part, too, is that about half the time that you send a link in, I already have it in my queue, or I’ve seen it before, but not always. I talked to Jeff Barr about this a while back, and apparently, a big Amazonian theme that he lives by is two is better than zero. He’d rather two people tell him about a thing than no one tells him about the thing. And I’ve tried to embody that. It’s the right answer, but it’s also super tricky to figure out what people have heard or haven’t heard. It leads to interesting places. But enough about my nonsense. Let’s talk about your nonsense instead. So, FusionAuth; what do you folks do over there?

Dan: So, FusionAuth is an auth provider, and we offer a Community Edition, which is downloadable for free; we also offer premium editions, but the space we play in is really CIAM, which is Customer Identity Access Management. Very similar to Auth0 or Cognito that some of your listeners might have heard of.

Corey: If people have heard about Cognito, it’s usually bracketed by profanity, in one direction or another, but I’m sure we’ll get there in a minute. I will say that I never considered authentication to be a differentiator between services that I use. And then one day I was looking for a tool—I’m not going to name what it was just because I don’t really want to deal with the angry letters and whatnot—but I signed up for this thing to test it out, and “Oh, great. So, what’s my password?” “Oh, we don’t use passwords. We just every time you want to log in, we’re going to email you a link and then you go ahead and click the link.”

And I hadn’t seen something like that before. And my immediate response to that was, “Okay, this feels like an area they’ve decided to innovate in.” Their core business is basically information retention and returning it to you—basically any CRUD app. Yay. I don’t think this is where I want them to be innovating.

I want them to use the tried and true solutions, not build their own or be creative on this stuff, so it was a contributor to me wanting to go in a different direction. When you start doing things like that, there’s no multi-factor authentication available and you start to wonder, how have they implemented this? What corners have they cut? Who’s reviewed this? It just gave me a weird feeling.

And that was sort of the day I realized that authentication for me is kind of like crypto, by which I mean cryptography, not cryptocurrency, I want to be very clear on, here. You should not roll your own cryptography, you should not roll your own encryption, you should buy off-the-shelf unless you’re one of maybe five companies on the planet. Spoiler, if you’re listening to this, you are almost certainly not one of them.

Dan: [laugh]. Yeah. So, first of all, I’ve been at FusionAuth for a couple of years. Before I came to FusionAuth, I had rolled my own authentication a couple of times. And what I’ve realized working there is that it really is—there a couple of things worth unpacking here.

One is you can now buy or leverage open-source libraries or other providers a lot more than you could 15 or 20 years ago. So, it’s become this thing that can be snapped into your architecture. The second is, auth is the front door to application. And while it isn’t really that differentiated—I don’t think most applications, as you kind of alluded to, should innovate there—it is kind of critical that it runs all the time that it’s safe and secure, that it’s accessible, that it looks like your application.

So, at the same time, it’s undifferentiated, right? Like, at the end of the day, people just want to get through authentication and authorization schemes into your application. That is really the critical thing. So, it’s undifferentiated, it’s critical, it needs to be highly available. Those are all things that make it a good candidate for outsourcing.

Corey: There are a few things to unpack there. First is that everything becomes commoditized in the fullness of time. And this is a good thing. Back in the original dotcom bubble, there were entire teams of engineers at all kinds of different e-commerce companies that were basically destroying themselves trying to build an online shopping cart. And today you wind up implementing Shopify or something like it—which is usually Shopify—and that solves the problem for you. This is no longer a point of differentiation.

If I want to start selling physical goods on the internet, it feels like it’ll take me half an hour or so to wind up with a bare-bones shopping cart thing ready to go, and then I just have to add inventory. Authentication feels like it was kind of the same thing. I mean, back in that song from early on in internet history “Code Monkey” talks about building a login page as part of it, and yeah, that was a colossal pain. These days, there are a bunch of different ways to do that with folks who spend their entire careers working on this exact problem so you can go and work on something that is a lot more core and central to the value that your business ostensibly provides. And that seems like the right path to go down.

But this does lead to the obvious counter-question of how is it that you differentiate other than, you know, via marketing, which again, not the worst answer in the world, but it also turns into skeezy marketing. “Yes, you should use this other company’s option, or you could use ours and we don’t have any intentional backdoors in our version.” “Hmm. That sounds more suspicious and more than a little bit frightening. Tell me more.” “No, legal won’t let me.” And it’s “Okay.” Aside from the terrible things, how do you differentiate?

Dan: I liked that. That was an oddly specific disclaimer, right? Like, whenever a company says, “Oh, yeah, no.” [laugh].

Corey: “My breakfast cereal has less arsenic than leading brands.”

Dan: Perfect. So yeah, so FusionAuth realizes that, kind of, there are a lot of options out there, and so we’ve chosen to niche down. And one of the things that we really focus on is the CIAM market. And that stands for Customer Identity Access Management. And we can dive into that a little bit later if you want to know more about that.

We have a variety of deployment options, which I think differentiates us from a lot of the SaaS providers out there. You can run us as a self-hosted option with, by the way, professional-grade support, you can use us as a SaaS provider if you don’t want to run it yourself. We are experts in operating this piece of software. And then thirdly, you can move between them, right? It’s your data, so if you start out and you’re bare bones and you want to save money, you can start with self-hosted, when you grow, move to the SaaS version.

Or we actually have some bigger companies that kickstart on the SaaS version because they want to get going with this integration problem and then later, as they build out their capabilities, they want the option to move it in-house. So, that is a really key differentiator for us. The last one I’d say is we’re really dev-focused. Who isn’t, right? Everyone says they’re dev-focused, but we live that in terms of our APIs, in terms of our documentation, in terms of our open development process. Like, there’s actually a GitHub issues list you can go look on the FusionAuth GitHub profile and it shows exactly what we have planned for the next couple of releases.

Corey: If you go to one of my test reference applications, lasttweetinaws.com, as of the time of this recording at least, it asks you to authenticate with your Twitter account. And you can do that, and it’s free; I don’t charge for any of these things. And once you’re authenticated, you can use it to author Twitter threads because I needed it to exist, first off, and secondly, it makes a super handy test app to try out a whole bunch of different things.

And one of the reasons you can just go and use it without registering an account for this thing or anything else was because I tried to set that up in an early version with Cognito and immediately gave the hell up and figured, all right, if you can find the URL, you can use this thing because the experience was that terrible. If instead, I had gone down the path of using FusionAuth, what would have made that experience different, other than the fact that Cognito was pretty clearly a tech demo at best rather than something that had any care, finish, spit and polish went into it.

Dan: So, I’ve used Cognito. I’m not going to bag on Cognito, I’m going to leave that to—[laugh].

Corey: Oh, I will, don’t worry. I’ll do all the bagging on Cognito you’d like because the problem is, and I want to be clear on this point, is that I didn’t understand what it was doing because the interface was arcane, and the failure mode of everything in this entire sector, when the interface is bad, the immediate takeaway is not “This thing’s a piece of crap.” It’s, “Oh, I’m bad at this. I’m just not smart enough.” And it’s insulting, and it sets me off every time I see it. So, if I feel like I’m coming across as relatively annoyed by the product, it’s because it made me feel dumb. That is one of those cardinal sins, from my perspective. So, if you work on that team, please reach out. I would love to give you a laundry list of feedback. I’m not here to make you feel bad about your product; I’m here to make you feel bad about making your customers feel bad. Now please, Dan, continue.

Dan: Sure. So, I would just say that one of the things that we’ve strived to do for years and years is translate some of the arcane IAM Identity Access Management jargon into what normal developers expect. And so, we don’t have clients in our OAuth implementation—although they really are clients if you’re an RFC junkie—we have applications, right? We have users, we have groups, we have all these things that are what users would expect, even though underlying them they’re based on the same standards that, frankly, Cognito and Auth0 and a lot of other people use as well.

But to get back to your question, I would say that, if you had chosen to use FusionAuth, you would have had a couple of advantages. The first is, as I mentioned, kind of the developer friendliness and the extensive documentation, example applications. The second would be a themeability. And this is something that we hear from our clients over and over again, is Cognito is okay if you stay within the lines in terms of your user interface, right? If you just want to login form, if you want to stay between lines and you don’t want to customize your application’s login page at all.

We actually provide you with HTML templates. It’s actually using a language called FreeMarker, but they let you do whatever the heck you want. Now, of course, with great power comes great responsibility. Now, you own that piece, right, and we do have some more simple customization you can do if all you want to do is change the color. But most of our clients are the kind of folks who really want their application login screen to look exactly like their application, and so they’re willing to take on that slightly heavier burden. Unfortunately, Cognito doesn’t give you that option at all, as far as I can tell when I’ve kicked the tires on it. The theming is—how I put this politely—some of our clients have found the theming to be lacking.

Corey: That’s part of the issue where when I was looking at all the reference implementations, I could find for Cognito, it went from “Oh, you have your own app, and its branding, and the rest,” and bam, suddenly, you’re looking right, like, you’re logging into an AWS console sub-console property because of course they have those. And it felt like “Oh, great. If I’m going to rip off some company’s design aesthetic wholesale, I’m sorry, Amazon is nowhere near anywhere except the bottom 10% of that list, I’ve got to say. I’m sorry, but it is not an aesthetically pleasing site, full stop. So, why impose that on customers?”

It feels like it’s one of those things where—like, so many Amazon service teams say, “We’re going to start by building a minimum lovable product.” And it’s yeah, it’s a product that only a parent could love. And the problem is, so many of them don’t seem to iterate beyond that do a full-featured story. And this is again, this is not every AWS service. A lot of them are phenomenal and grow into themselves over time.

One of the best rags-to-riches stories that I can recall is EFS, their Elastic File System, for an example. But others, like Cognito just sort of seem to sit and languish for so long that I’ve basically given up hope. Even if they wind up eventually fixing all of these problems, the reputation has been cemented at this point. They’ve got to give it a different terrible name.

Dan: I mean, here’s the thing. Like, EFS, if it looks horrible, right, or if it has, like, a toughest user experience, guess what? Your users are devs. And if they’re forced to use it, they will. They can sometimes see the glimmers of the beauty that is kind of embedded, right, the diamond in the rough. If your users come to a login page and see something ugly, you immediately have this really negative association. And so again, the login and authentication process is really the front door of your application, and you just need to make sure that it shines.

Corey: For me at least, so much of what’s what a user experience or user takeaway is going to be about a company’s product starts with their process of logging into it, which is one of the reasons that I have challenges with the way that multi-factor auth can be presented, like, “Step one, login to the thing.” Oh, great. Now, you have to fish out your YubiKey, or you have to go check your email for a link or find a code somewhere and punch it in. It adds friction to a process. So, when you have these services or tools that oh, your session will expire every 15 minutes and you have to do that whole thing again to log back in, it’s ugh, I’m already annoyed by the time I even look at anything beyond just the login stuff.

And heaven forbid, like, there are worse things, let’s be very clear here. For example, if I log in to a site, and I’m suddenly looking at someone else’s account, yeah, that’s known as a disaster and I don’t care how beautiful the design aesthetic is or how easy to use it is, we’re done here. But that is job zero: the security aspect of these things. Then there’s all the polish that makes it go from something that people tolerate because they have to into something that, in the context of a login page I guess, just sort of fades into the background.

Dan: That’s exactly what you want, right? It’s just like the old story about the sysadmin. People only notice when things are going wrong. People only care about authentication when it stops them from getting into what they actually want to do, right? No one ever says, “Oh, my gosh, that login experience was so amazing for that application. I’m going to come back to that application,” right? They notice when it’s friction, they noticed when it’s sand in the gears.

And our goal at FusionAuth, obviously, security is job zero because as you said, last thing you want is for a user to have access to some other user’s data or to be able to escalate their privileges, but after that, you want to fade in the background, right? No one comes to FusionAuth and builds a whole application on top of it, right? We are one component that plugs into your application and lets you get on to the fundamentals of building the features that your users really care about, and then wraps your whole application in a blanket of security, essentially.

Corey: I’ll take even one more example before we just drive this point home in a way that I hope resonates with folks. Everyone has an opinion on logging into AWS properties because “Oh, what about your Amazon account?” At which point it’s “Oh, sit down. We’re going for a ride here. Are you talking about amazon.com account? Are you talking about the root account for my AWS account? Are you talking about an IAM user? Are you talking about the service formerly known as AWS SSO that’s now IAM Identity Center users? Are you talking about their Chime user account? Are you talking about your repost forum account?” And so, on and so on and so on. I’m sure I’m missing half a dozen right now off the top of my head.

Yeah, that’s awful. I’ve been also developing lately on top of Google Cloud, and it is so far to the opposite end of that spectrum that it’s suspicious and more than a little bit frightening. When I go to console.cloud.google.com, I am boom, there. There is no login approach, which on the one hand, I definitely appreciate, just from a pure perspective of you’re Google, you track everything I do on the internet. Thank you for not insulting my intelligence by pretending you don’t know who I am when I log into your Cloud Console.

Counterpoint, when I log into the admin portal for my Google Workspaces account, admin.google.com, it always re-prompts for a password, which is reasonable. You’d think that stuff running production might want to do something like that, in some cases. I would not be annoyed if it asked me to just type in a password again when I get to the expensive things that have lasting repercussions.

Although, given my personality, logging into Gmail can have massive career repercussions as soon as I hit send on anything. I digress. It is such a difference from user experience and ease-of-use that it’s one of those areas where I feel like you’re fighting something of a losing battle, just because when it works well, it’s glorious to the point where you don’t notice it. When authentication doesn’t work well, it’s annoying. And there’s really no in between.

Dan: I don’t have anything to say to that. I mean, I a hundred percent agree that it’s something that you could have to get right and no one cares, except for when you get it wrong. And if your listeners can take one thing away from this call, right, I know it’s we’re sponsored by FusionAuth, I want to rep Fusion, I want people to be aware of FusionAuth, but don’t roll your own, right? There are a lot of solutions out there. I hope you evaluate FusionAuth, I hope you evaluate some other solutions, but this is such a critical thing and Corey has laid out [laugh] in multiple different ways, the ways it can ruin your user experience and your reputation. So, look at something that you can build or a library that you can build on top of. Don’t roll your own. Please, please don’t.

Corey: This episode is sponsored in part by Honeycomb. When production is running slow, it’s hard to know where problems originate. Is it your application code, users, or the underlying systems? I’ve got five bucks on DNS, personally. Why scroll through endless dashboards while dealing with alert floods, going from tool to tool to tool that you employ, guessing at which puzzle pieces matter? Context switching and tool sprawl are slowly killing both your team and your business. You should care more about one of those than the other; which one is up to you. Drop the separate pillars and enter a world of getting one unified understanding of the one thing driving your business: production. With Honeycomb, you guess less and know more. Try it for free at honeycomb.io/screaminginthecloud. Observability: it’s more than just hipster monitoring.

Corey: So, tell me a little bit more about how it is that you folks think about yourselves in just in terms of the market space, for example. The idea of CIAM, customer IAM, it does feel viscerally different than traditional IAM in the context of, you know, AWS, which I use all the time, but I don’t think I have the vocabulary to describe it without sounding like a buffoon. What is the definition between the two, please? Or the divergence, at least?

Dan: Yeah, so I mean, not to go back to AWS services, but I’m sure a lot of your listeners are familiar with them. AWS SSO or the artist formerly known as AWS SSO is IAM, right? So, it’s Workforce, right, and Workforce—

Corey: And it was glorious, to the point where I felt like it was basically NDA’ed from other service teams because they couldn’t talk about it. But this was so much nicer than having to juggle IAM keys and sessions that timeout after an hour in the console. “What do you doing in the console?” “I’m doing ClickOps, Jeremy. Leave me alone.”

It’s just I want to make sure that I’m talking about this the right way. It feels like AWS SSO—creature formerly known as—and traditional IAM feels like they’re directionally the same thing as far as what they target, as far as customer bases, and what they empower you to do.

Dan: Absolutely, absolutely. There are other players in that same market, right? And that’s the market that grew up originally: it’s for employees. So, employees have this very fixed lifecycle. They have complicated relationships with other employees and departments in organizations, you can tell them what to do, right, you can say you have to enroll your MFA key or you are no longer employed with us.

Customers have a different set of requirements, and yet they’re crucial to businesses because customers are, [laugh] who pay you money, right? And so, things that customers do that employees don’t: they choose to register; they pick you, you don’t pick them; they have a wide variety of devices and expectations; they also have a higher expectation of UX polish. Again, with an IAM solution, you can kind of dictate to your employees because you’re paying them money. With a customer identity access management solution, it is part of your product, in the same way, you can’t really dictate features unless you have something that the customer absolutely has to have and there are no substitutes for it, you have to adjust to the customer demands. CIAM is more responsive to those demands and is a smoother experience.

The other thing I would say is CIAM, also, frankly, has a simpler model. Most customers have access to applications, maybe they have a couple of roles that you know, an admin role, an editor role, a viewer role if you’re kind of a media conglomerate, for an example, but they don’t have necessarily the thicket of complexity that you might have to have an eye on, so it’s just simpler to model.

Corey: Here’s an area that feels like it’s on the boundary between them. I distinctly remember being actively annoyed a while back that I had to roll my marketing person her own entire AWS IAM account solely so that she could upload assets into an S3 bucket that was driving some other stuff. It feels very much like that is a better use case for something that is a customer IAM solution. Because if I screw up those permissions even slightly, well, congratulations, now I’ve inadvertently given someone access to wind up, you know, taking production down. It feels like it is way too close to things that are going to leave a mark, whereas the idea of a customer authentication story for something like that is awesome.

And no please if you’re listening to this, don’t email me with this thing you built and put on the Marketplace that “Oh, it uses signed URLs and whatnot to wind up automatically federating an identity just for this one per—” Yes. I don’t want to build something ridiculous and overwrought so a single person can update assets within S3. I promise I don’t want to do that. It just ends badly.

Dan: Well, that was the promise of Cognito, right? And that is actually one of the reasons you should stick with Cognito if you have super-detailed requirements that are all about AWS and permissions to things inside AWS. Cognito has that tight integration. And I assume—I haven’t looked at some of the other big cloud providers, but I assume that some of the other ones have that similar level of integration. So yeah, so that my answer there would be Cognito is the CIAM solution that AWS has, so that is what I would expect it to be able to handle, relatively smoothly.

Corey: A question I have for you about the product itself is based on a frustration I originally had with Cognito, which is that once you’re in there and you are using that for authentication and you have users, there’s no way for me to get access to the credentials of my users. I can’t really do an export in any traditional sense. Is that possible with FusionAuth?

Dan: Absolutely. So, your data is your data. And because we’re a self-hosted or SaaS solution, if you’re running it self-hosted, obviously you have access to the password hashes in your database. If you are—

Corey: The hashes, not the plaintext passwords to be explicitly clear on this. [laugh].

Dan: Absolutely the hashes. And we have a number of guides that help you get hashes from other providers into ours. We have a written export guide ourselves, but it’s in the database and the schema is public. You can go download our schema right now. And if—

Corey: And I assume you’ve used an industry standard hashing algorithm for this?

Dan: Yeah, we have a number of different options. You can bring your own actually, if you want, and we’ve had people bring their own options because they have either special needs or they have an older thing that’s not as secure. And so, they still want their users to be able to log in, so they write a plugin and then they import the users’ hashes, and then we transparently re-encrypt with a more modern one. The default for us is PDK.

Corey: I assume you do the re-encryption at login time because there’s no other way for you to get that.

Dan: Exactly. Yeah yeah yeah—

Corey: Yeah.

Dan: —because that’s the only time we see the password, right? Like we don’t see it any other time. But we support Bcrypt and other modern algorithms. And it’s entirely configurable; if you want to set a factor, which basically is how—

Corey: I want to use MD5 because I’m still living in 2003.

Dan: [laugh]. Please don’t use MD5. Second takeaway: don’t roll your own and don’t use MD5. Yeah, so it’s very tweakable, but we shipped with a secured default, basically.

Corey: I just want to clarify as well why this is actively important. I don’t think people quite understand that in many cases, picking an authentication provider is one of those lasting decisions where migrations take an awful lot of work. And they probably should. There should be no mechanism by which I can export the clear text passwords. If any authentication provider advertises or offers such a thing, don’t use that one. I’m going to be very direct on that point.

The downside to this is that if you are going to migrate from any other provider to any other provider, it has to happen either slowly as in, every time people log in, it’ll check with the old system and then migrate that user to the new one, or you have to force password resets for your entire customer base. And the problem with that is I don’t care what story you tell me. If I get an email from one of my vendors saying “You now have to reset your password because we’re migrating to their auth thing,” or whatnot, there’s no way around it, there’s no messaging that solves this, people will think that you suffered a data breach that you are not disclosing. And that is a heavy, heavy lift. Another pattern I’ve seen is it for a period of three months or whatnot, depending on user base, you will wind up having the plug in there, and anyone who logs in after that point will, “Ohh you need to reset your password. And your password is expired. Click here to reset.” That tends to be a little bit better when it’s not the proactive outreach announcement, but it’s still a difficult lift and it adds—again—friction to the customer experience.

Dan: Yep. And the third one—which you imply it—is you have access to your password hashes. They’re hashed in a secure manner. And trust me, even though they’re hashed securely, like, if you contact FusionAuth and say, “Hey, I want to move off FusionAuth,” we will arrange a way to get you your database in a secure manner, right? It’s going to be encrypted, we’re going to have a separate password that we communicate with you out-of-band because this is—even if it is hashed and salted and handled correctly, it’s still very, very sensitive data because credentials are the keys to the kingdom.

So, but those are the three options, right? The slow migration, which is operationally expensive, the requiring the user to reset their password, which is horribly expensive from a user interface perspective, right, and the customer service perspective, or export your password hashes. And we think that the third option is the least of the evils because guess what? It’s your data, right? It’s your user data. We will help you be careful with it, but you own it.

Corey: I think that there’s a lot of seriously important nuance to the whole world of authentication. And the fact that this is such a difficult area to even talk about with folks who are not deeply steeped in that ecosystem should be an indication alone that this is the sort of thing that you definitely want to outsource to a company that knows what the hell they’re doing. And it’s not like other areas of tech where you can basically stumble your way through something. It’s like “Well, I’m going to write a Lambda to go ahead and post some nonsense on Twitter.” “Okay, are you good at programming?” “Not even slightly, but I am persistent and brute force is a viable strategy, so we’re going to go with that one.” “Great. Okay, that’s awesome.”

But authentication is one of those areas where mistakes will show. The reputational impact of losing data goes from merely embarrassing to potentially life-ruining for folks. The most stressful job I’ve ever had from a data security position wasn’t when I was dealing with money—because that’s only money, which sounds like a weird thing to say—it was when I did a brief stint at Grindr where people weren’t out. In some countries, users could have wound up in jail or have been killed if their sexuality became known. And that was the stuff that kept me up at night.

Compared to that, “Okay, you got some credit card numbers with that. What the hell do I care about that, relatively speaking?” It’s like, “Yeah, it’s well, my credit card number was stolen.” “Yeah, but did you die, though?” “Oh, you had to make a phone call and reset some stuff.” And I’m not trivializing the importance of data security. Especially, like, if you’re a bank, and you’re listening to this, and you’re terrified, yeah, that’s not what I’m saying at all. I’m just saying there are worse things.

Dan: Sure. Yeah. I mean, I think that, unfortunately, the pandemic showed us that we’re living more and more of our lives online. And the identity online and making sure that safe and secure is just critical. And again, not just for your employees, although that’s really important, too, but more of your customer interactions are going to be taking place online because it’s scalable, because it makes people money, because it allows for capabilities that weren’t previously there, and you have to take that seriously. So, take care of your users’ data. Please, please do that.

Corey: And one of the best ways you can do that is by not touching the things that are commoditized in your effort to apply differentiation. That’s why I will never again write my own auth system, with a couple of asterisks next to it because some of what I do is objectively horrifying, intentionally so. But if I care about the authentication piece, I have the good sense to pay someone else to do it for me.

Dan: From personal experience, you mentioned at the beginning that we go back aways. I remember when I first discovered RDS, and I thought, “Oh, my God. I can outsource all this scut work, all of the database backups, all of the upgrades, all of the availability checking, right? Like, I can outsource this to somebody else who will take this off my plate.” And I was so thankful.

And I don’t—outside of, again, with some asterisks, right, there are places where I could consider running a database, but they’re very few and far between—I feel like auth has entered that category. There are great providers like FusionAuth out there that are happy to take this off your plate and let you move forward. And in some ways, I’m not really sure which is more dangerous; like, not running a database properly or not running an auth system properly. They both give me shivers and I would hate to [laugh] hate to be forced to choose. But they’re comparable levels of risk, so I a hundred percent agree, Corey.

Corey: Dan, I really want to thank you for taking so much time to talk to me about your view of the world. If people want to learn more because you’re not in their inboxes responding to newsletters every week, where’s the best place to find you?

Dan: Sure, you can find more about me at Twitter. I’m @mooreds, M-O-O-R-E-D-S. And you can learn more about FusionAuth and download it for free at fusionauth.io.

Corey: And we will put links to all of that in the show notes. I really want to thank you again for just being so generous with your time. It’s deeply appreciated.

Dan: Corey, thank you so much for having me.

Corey: Dan Moore, Head of DevRel at FusionAuth. I’m Cloud Economist Corey Quinn. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an angry, insulting comment that will be attributed to someone else because they screwed up by rolling their own authentication.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Anaïs

Anaïs is a Developer Advocate at Aqua Security, where she contributes to Aqua’s cloud native open source projects. When she is not advocating DevOps best practices, she runs her own YouTube Channel centered around cloud native technologies. Before joining Aqua, Anais worked as SRE at Civo, a cloud native service provider, where she helped enhance the infrastructure for hundreds of tenant clusters. As CNCF ambassador of the year 2021, her passion lies in making tools and platforms more accessible to developers and community members.

Links Referenced:

  • Aqua Security: https://www.aquasec.com/
  • Aqua Open Source YouTube channel: https://www.youtube.com/c/AquaSecurityOpenSource
  • Personal blog: https://anaisurl.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at AWS AppConfig. Engineers love to solve, and occasionally create, problems. But not when it’s an on-call fire-drill at 4 in the morning. Software problems should drive innovation and collaboration, NOT stress, and sleeplessness, and threats of violence. That’s why so many developers are realizing the value of AWS AppConfig Feature Flags. Feature Flags let developers push code to production, but hide that that feature from customers so that the developers can release their feature when it’s ready. This practice allows for safe, fast, and convenient software development. You can seamlessly incorporate AppConfig Feature Flags into your AWS or cloud environment and ship your Features with excitement, not trepidation and fear. To get started, go to snark.cloud/appconfig That’s snark.cloud/appconfig.

Corey: This episode is sponsored in part by Honeycomb. When production is running slow, it’s hard to know where problems originate. Is it your application code, users, or the underlying systems? I’ve got five bucks on DNS, personally. Why scroll through endless dashboards while dealing with alert floods, going from tool to tool to tool that you employ, guessing at which puzzle pieces matter? Context switching and tool sprawl are slowly killing both your team and your business. You should care more about one of those than the other; which one is up to you. Drop the separate pillars and enter a world of getting one unified understanding of the one thing driving your business: production. With Honeycomb, you guess less and know more. Try it for free at honeycomb.io/screaminginthecloud. Observability: it’s more than just hipster monitoring.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Every once in a while, when I start trying to find guests to chat with me and basically suffer my various slings and arrows on this show, I encounter something that I’ve never really had the opportunity to explore further. And today’s guest leads me in just such a direction. Anaïs is an open-source developer advocate at Aqua Security, and when I was asking her whether or not she wanted to talk about various topics, one of the first thing she said was, “Don’t ask me much about AWS because I’ve never used it,” which, oh my God. Anaïs, thank you for joining me. You must be so very happy never to have dealt with the morass of AWS.

Anaïs: [laugh]. Yes, I’m trying my best to stay away from it. [laugh].

Corey: Back when I got into the cloud space, for lack of a better term, AWS was sort of really the only game in town unless you wanted to start really squinting hard at what you define cloud as. I mean yes, I could have gone into Salesforce or something, but I was already sad and angry all the time. These days, you can very much go all in-on cloud. In fact, you were a CNCF ambassador, if I’m not mistaken. So, you absolutely are in the infrastructure cloud space, but you haven’t dealt with AWS. That is just an interesting path. Have you found others who have gone down that same road, or are you sort of the first of a new breed?

Anaïs: I think to find others who are in a similar position or have a similar experience, as you do, you first have to talk about your experience, and this is the first time, or maybe the second, that I’m openly [laugh] saying it on something that will be posted live, like, to the internet. Before I, like, I tried to stay away from mentioning it at all, do the best that I can because I’m at this point where I’m so far into my cloud-native Kubernetes journey that I feel like I should have had to deal with AWS by now, and I just didn’t. And I’m doing my best and I’m very successful in avoiding it. [laugh]. So, that’s where I am. Yeah.

Corey: We’re sort of on opposite sides of a particular fence because I spend entirely too much time being angry at AWS, but I’ve never really touched Kubernetes and anger. I mean, I see it in a lot of my customer accounts and I get annoyed at its data transfer bills and other things that it causes in an economic sense, but as far as the care and feeding of a production cluster, back in my SRE days, I had very old-school architectures. It’s, “Oh, this is an ancient system, just like grandma used to make,” where we had the entire web tier, then a job applic—or application server tier, and then a database at the end, and everyone knew where everything was. And then containers came out of nowhere, and it seemed like okay, this solves a bunch of problems and introduces a whole bunch more. How do I orchestrate them? How do I ensure that they’re healthy?

And then ah, Kubernetes was the answer. And for a while, it seemed like no matter what the problem was, Kubernetes was going to be the answer because people were evangelizing it pretty hard. And now I see it almost everywhere that I turn. What’s your journey been like? How did you get into the weeds of, “You know what I want to do when I grow up? That’s right. I want to work on container orchestration systems.” I have a five-year-old. She has never once said that because I don’t abuse my children by making them learn how clouds work. How did you wind up doing what you do?

Anaïs: It’s funny that you mention that. So, I’m actually of the generation of engineers who doesn’t know anything else but Kubernetes. So, when you mentioned that you used to use something before, I don’t really know what that looks like. I know that you can still deploy systems without Kubernetes, but I have no idea how. My journey into the cloud-native space started out of frustration from the previous industry that I was working at.

So, I was working for several years as developer advocate in the open-source blockchain cryptocurrency space and it’s highly similar to all of the cliches that you hear online and across the news. And out of this frustration, [laugh] I was looking at alternatives. One of them was either going into game development, into the gaming industry, or the cloud-native space and infrastructure development and deployment. And yeah, that’s where I ended up. So, at the end of 2020, I joined a startup in the cloud-native space and started my social media journey.

Corey: One of the things that I found that Kubernetes solved for—and to be clear, Kubernetes really came into its own after I was doing a lot more advisory work and a lot more consulting style activity rather than running my own environments, but there’s an entire universe of problems that the modern day engineer never has to think about due to, partially cloud and also Kubernetes as well, which is the idea of hardware or node failure. I’ve had middle of the night driving across Los Angeles in a panic getting to the data center because the disk array on the primary database had degraded because the drive failed. That doesn’t happen anymore. And clouds have mostly solved that. It’s okay, drives fail, but yeah, that’s the problem for some people who live in Virginia or Oregon. I don’t have to think about it myself.

But you do have to worry about instances failing; what if the primary database instance dies? Well, when everything lives in a container then that container gets moved around in the stateless way between things, well great, you really only have to care instead about okay, what if all of my instances die? Or, what if my code is really crappy? To which my question is generally, what do you mean, ‘if?’ All of us write crappy code.

That’s the nature of the universe. We open-source only the small subset that we are not actively humiliated by, which is, in a lot of ways, what you’re focusing on now, over at Aqua Sec, you are an advocate for open-source. One of the most notable projects that come out of that is Trivy, if I’m pronouncing that correctly.

Anaïs: Yeah, that’s correct. Yeah. So, Trivy is our main open-source project. It’s an all-in-one cloud-native security scanner. And it’s actually—it’s focused on misconfiguration issues, so it can help you to build more robust infrastructure definitions and configurations.

So ideally, a lot of the things that you just mentioned won’t happen, but it obviously, highly depends on so many different factors in the cloud-native space. But definitely misconfigurations of one of those areas that can easily go wrong. And also, not just that you have data might cease to exist, but the worst thing or, like, as bad might be that it’s completely exposed online. And they are databases of different exposures where you can see all the kinds of data of information from just health data to dating apps, just being online available because the IP address is not protected, right? Things like that. [laugh].

Corey: We all get those emails that start with, “Your security is very important to us,” and I know just based on that opening to an email, that the rest of that email is going to explain how security was not very important to you folks. And it’s the apology, “Oops, we have messed up,” email. Now, the whole world of automated security scanners is… well, it’s crowded. There are a number of different services out there that the cloud providers themselves offer a bunch of these, a whole bunch of scareware vendors at the security conferences do as well. Taking a quick glance at Trivy, one of the problems I see with it, from a cloud provider perspective, is that I see nothing that it does that winds up costing extra money on your cloud bill that you then have to pay to the cloud provider, so maybe they’ll put a pull request in for that one of these days. But my sarcasm aside, what is it that differentiates Trivy from a bunch of other offerings in various spaces?

Anaïs: So, there are multiple factors. If we’re looking from an enterprise perspective, you could be using one of the in-house scanners from any of the cloud providers available, depending which you’re using. The thing is, they are not generally going to be the ones who have a dedicated research team that provides the updates based on the vulnerabilities they find across the space. So, with an open-source security scanner or from a dedicated company, you will likely have more up-to-date information in your scans. Also, lots of different companies, they’re using Trivy under the hood ultimately, or for their own scans.

I can link a few where you can also find them in a Trivy repository. But ultimately, a lot of companies rely on Trivy and other open-source security scanners under the hood because they are from dedicated companies. Now, the other part to Trivy and why you might want to consider using Trivy is that in larger teams, you will have different people dealing with different components of your infrastructure, of your deployments, and you could end up having to use multiple different security scanners for all your different components from your container images that you’re using, whether or not they are secure, whether or not they’re following best practices that you defined to your infrastructure-as-code configurations, to you’re running deployments inside of your cluster, for instance. So, each of those different stages across your lifecycle, from development to runtime, will maybe either need different security scanners, or you could use one security scanner that does it all. So, you could have in a team more knowledge sharing, you could have dedicated people who know how to use the tool and who can help out across a team across the lifecycle, and similar. So, that’s one of the components that you might want to consider.

Another thing is how mature is a tool, right? A lot of cloud providers, what they end up doing is they provide you with a solution, but it’s nice to decoupled from anything else that you’re using. And especially in the cloud-native space, you’re heavily reliant on open-source tools, such as for your observability stack, right? Coming from Site Reliability Engineering also myself, I love using metrics and Grafana. And for me, if anything open-source from Loki to accessing my logs, to Grafana to dashboards, and all their integrations.

I love that and I want to use the same tools that I’m using for everything else, also for my security tools. I don’t want to have the metrics for my security tools visualized in a different solution to my reliability metrics for my application, right? Because that ultimately makes it more difficult to correlate metrics. So, those are, like, some of the factors that you might want to consider when you’re choosing a security scanner.

Corey: When you talk about thinking about this, from the perspective of an SRE is—I mean, this is definitely an artifact of where you come from and how you approach this space. Because in my world, when you have ten web servers, five application servers, and two database servers and you wind up with a problem in production, how do you fix this? Oh, it’s easy. You log into one of those nodes and poke around and start doing diagnostics in production. In a containerized world, you generally can’t do that, or there’s a problem on a container, and by the time you’re aware of that, that container hasn’t existed for 20 minutes.

So, how do you wind up figuring out what happens? And instrumenting for telemetry and metrics and observability, particularly at scale becomes way more important than it ever was, for me. I mean, my version of monitoring was always Nagios, which was the original Call of Duty that wakes you up at two in the morning when the hard drive fails. The world has thankfully moved beyond that and a bunch of ways. But it’s not first nature for me. It’s always, “Oh, yeah, that’s right. We have a whole telemetry solution where I can go digging into.” My first attempt is always, oh, how do I get into this thing and poke it with a stick? Sometimes that’s helpful, but for modern applications, it really feels like it’s not.

Anaïs: Totally. When we’re moving to an infrastructure to an environment where we can deploy multiple times a day, right, and update our application multiple times a day, multiple times a day, we can introduce new security issues or other things can go wrong, right? So, I want to see—as much as I want to see all of the other failures, I want to see any security-related issues that might be deployed alongside those updates at the same frequency, right?

Corey: The problem that I see across all this stuff, though, is there are a bunch of tools out there that people install, but then don’t configure because, “Oh, well, I bought the tool. The end.” I mean, I think it was reported almost ten years ago or so on the big Target breach that they wound up installing some tool—I want to say FireEye, but please don’t quote me on that—and it wound up firing off a whole bunch of alerts, and they figured was just noise, so they turned it all off. And it turned out no, no, this was an actual breach in progress. But people are so used to all the alarms screaming at them, that they don’t dig into this.

I mean, one of the original security scanners was Nessus. And I seen a lot of Nessus reports because for a long time, what a lot of crappy consultancies would do is they would white-label the output of whatever it was that Nessus said and deliver that in as the report. So, you’d wind up with 700 pages of quote-unquote, “Security issues.” And you’d have to flip through to figure out that, ah, this supports a somewhat old SSL negotiation protocol, and you’re focusing on that instead of the oh, and by the way, the primary database doesn’t have a password set. Like, it winds up just obscuring it because there is so much. How does Trivy approach avoiding the information overload problem?

Anaïs: That’s a great question because everybody’s complaining about vulnerability fatigue, of them, for the first time, scanning their container images and workloads and seeing maybe even hundreds of vulnerabilities. And one of the things that can be done to counteract that right from the beginning is investing your time into looking at the different flags and configurations that you can do before actually deploying Trivy to, for example, your cluster. That’s one part of it. The other part is I mentioned earlier, you would use a security scan at different parts of your deployment. So, it’s really about integrating scanning not just once you—like, in your production environment, once you’ve deployed everything, but using it already before and empowering engineers to actually use it on their machines.

Now, they can either decide to do it or not; it’s not part of most people’s job to do security scanning, but as you move along, the more you do, the more you can reduce the noise and then ultimately, when you deploy Trivy, for example, inside of your cluster, you can do a lot of configuration such as scanning just for critical vulnerabilities, only scanning for vulnerabilities that already have a fix available, and everything else should be ignored. Those are all factors and flags that you can place into Trivy, for instance, and make it easier. Now, with Trivy, you won’t have automated PRs and everything out of the box; you would have to set up the actions or, like, the ways to mitigate those vulnerabilities manually by yourself with tools, as well as integrating Trivy with your existing stack, and similar. But then obviously, if you want to have something more automated, if you want to have something that does more for you in the background, that’s when you want to use to an enterprise solution and shift to something like Aqua Security Enterprise Platform that actually provides you with the automated way of mitigating vulnerabilities where you don’t have to know much about it and it just gives you the solution and provides you with a PR with the updates that you need in your infrastructure-as-code configurations to mitigate the vulnerability [unintelligible 00:15:52]?

Corey: I think that’s probably a very fair answer because let’s be serious when you’re running a bank or someone for whom security matters—and yes, yes, I know, security should matter for everyone, but let’s be serious, I care a little bit less about the security impact of, for example, I don’t know, my Twitter for Pets nonsense, than I do a dating site where people are not out about their orientation or whatnot. Like, there is a world of difference between the security concerns there. “Oh, no, you might be able to shitpost as me if you compromise my lasttweetinaws.com Twitter client that I put out there for folks to use.” Okay, great. That is not the end of the world compared to other stuff.

By the time you’re talking about things that are critically important, yeah, you want to spend money on this, and you want to have an actual full-on security team. But open-source tools like this are terrific for folks who are just getting started or they’re building something for fun themselves and as it turns out, don’t have a full security budget for their weird late-night project. I think that there’s a beautiful, I guess, spectrum, as far as what level of investment you can make into security. And it’s nice to see the innovation continued happening in the space.

Anaïs: And you just mentioned that dedicated security companies, they likely have a research team that’s deploying honeypots and seeing what happens to them, right? Like, how are attackers using different vulnerabilities and misconfigurations and what can be done to mitigate them. And that ultimately translates into the configurations of the open-source tool as well. So, if you’re using, for instance, a security scanner that doesn’t have an enterprise company with a research team behind it, then you might have different input into the data of that security scanner than if you do, right? So, these are, like, additional considerations that you might want to take when choosing a scanner. And also that obviously depends on what scanning you want to do, on the size of your company, and similar, right?

Corey: This episode is sponsored in part by our friend EnterpriseDB. EnterpriseDB has been powering enterprise applications with PostgreSQL for 15 years. And now EnterpriseDB has you covered wherever you deploy PostgreSQL on-premises, private cloud, and they just announced a fully-managed service on AWS and Azure called BigAnimal, all one word. Don’t leave managing your database to your cloud vendor because they’re too busy launching another half-dozen managed databases to focus on any one of them that they didn’t build themselves. Instead, work with the experts over at EnterpriseDB. They can save you time and money, they can even help you migrate legacy applications—including Oracle—to the cloud. To learn more, try BigAnimal for free. Go to biganimal.com/snark, and tell them Corey sent you.

Corey: Something that I do find fairly interesting is that you started off, as you say, doing DevRel in the open-source blockchain world, then you went to work as an SRE, and then went back to doing DevRel-style work. What got you into SRE and what got you out of SRE, other than the obvious having worked in SRE myself and being unhappy all the time? I kid, but what was it that got you into that space and then out of it?

Anaïs: Yeah. Yeah, but no, it’s a great question. And it’s, I guess, also was shaped my perspective on different tools and, like, the user experience of different tools. But ultimately, I first worked in the cloud-native space for an enterprise tool as developer advocate. And I did not like the experience of working for a paid solution. Doing developer advocacy for it, it felt wrong in a lot of ways. A lot of times you were required to do marketing work in those situations.

And that kind of got me out of developer advocacy into SRE work. And now I was working partially or mainly as SRE, and then on the side, I was doing some presentations in developer advocacy. However, that split didn’t quite work, either. And I realized that the value that I add to a project is really the way I convey information, which I can’t do if I’m busy fixing the infrastructure, right? I can’t convey the information of as much of how the infrastructure has been fixed as I can if I’m working with an engineering team and then doing developer advocacy, solely developer advocacy within the engineering team.

So, how I ultimately got back into developer advocacy was just simply by being reached out to by my manager at Aqua Security, and Itay telling me, him telling me that he has a role available and if I want to join his team. And it was open-source-focused. Given that I started my career for several years working in the open-source space and working with engineers, contributing to open-source tools, it was kind of what I wanted to go back to, what I really enjoy doing. And yeah, that’s how that came about [laugh].

Corey: For me, I found that I enjoy aspects of the technology part, but I find I enjoy talking to people way more. And for me, the gratifying moment that keeps me going, believe it or not, is not necessarily helping giant companies spend slightly less money on another giant company. It’s watching people suddenly understand something they didn’t before, it’s watching the light go on in their eyes. And that’s been addictive to me for a long time. I’ve also found that the best way for me to learn something is to teach someone else.

I mean, the way I learned Git was that I foolishly wound up proposing a talk, “Terrible Ideas in Git”—we’ll teach it by counterexample—four months before the talk. And they accepted it, and crap, I’d better learn enough get to give this talk effectively. I don’t recommend this because if you miss the deadline, I checked, they will not move the conference for you. But there really is something to be said for watching someone learn something by way of teaching it to them.

Anaïs: It’s actually a common strategy for a lot of developer advocates of making up a talk and then waiting whether or not it will get accepted. [laugh] and once it gets accepted, that’s when you start learning the tool and trying to figure it out. Now, it’s not a good strategy, obviously, to do that because people can easily tell that you just did that for a conference. And—

Corey: Sounds to me, like, you need to get better at bluffing. I kid.

Anaïs: [laugh].

Corey: I kid. Don’t bluff your way through conference talks as a general rule. It tends not to go well. [laugh].

Anaïs: No. It’s a bad idea. It’s a really bad idea. And so, I ultimately started learning the technologies or, like, the different tools and projects in the cloud-native space. And there are lots, if you look at the CNCF landscape, right? But just trying to talk myself through them on my YouTube channel. So, my early videos on my channel, it’s just very much on the go of me looking for the first time at somebody’s documentation and not making any sense out of them.

Corey: It’s surprising to me how far that gets you. I mean, I guess I’m always reminded of that Tom Hanks movie from my childhood Big where he wakes up—the kid wakes up as an adult one day, goes to work, and bluffs his way into working at a toy company. He’s in a management meeting and just they’re showing their new toy they’re going to put out there and he’s, “I don’t get it.” Everyone looks at him like how dare you say it? And, “I don’t get it. What’s fun about this?” Because he’s a kid.

And he wants to getting promoted to vice president because wow, someone pointed out the obvious thing. And so often, it feels like using a tool or a product, be it open-source or enterprise, it is clearly something different in my experience of it when I try to use this thing than the person who developed it. And very often it’s that I don’t see the same things or think of the problem space the same way that the developers did, but also very often—and I don’t mean to call anyone in particular out here—it’s a symptom of a terrible user interface or user experience.

Anaïs: What you’ve just said, a lot of times, it’s just about saying the thing that nobody that dares to say or nobody has thought of before, and that gets you obviously, easier, further [laugh] then repeating what other people have already mentioned, right? And a lot of what you see a lot of times in these—also an open-source projects, but I think more even in closed-source enterprise organizations is that people just repeat whatever everybody else is saying in the room, right? You don’t have that as much in the open-source world because you have more input or easier input in public than you do otherwise, but it still happens that I mean, people are highly similar to each other. If you’re contributing to the same project, you probably have a similar background, similar expertise, similar interests, and that will get you to think in a similar way. So, if there’s somebody like, like a high school student maybe, somebody just graduated, somebody from a completely different industry who’s looking at those tools for the first time, it’s like, “Okay, I know what I’m supposed to do, but I don’t understand why I should use this tool for that.” And just pointing that out, gets you a response, most of the time. [laugh].

Corey: I use Twitter and use YouTube. And obviously, I bias more for short, pithy comments that are dripping in sarcasm, whereas in a long-form video, you can talk a lot more about what you’re seeing. But the problem I have with bad user experience, particularly bad developer experience, is that when it happens to me—and I know at a baseline level, that I am reasonably competent in technical spaces, but when I encounter a bad interface, my immediate instinctive reaction is, “Oh, I’m dumb. And this thing is for smart people.” And that is never, ever true, except maybe with quantum computing. Great, awesome. The Hello World tutorial for that stuff is a PhD from Berkeley. Good luck if you can get into that. But here in the real world where the rest of us play, it’s just a bad developer experience, but my instinctive reaction is that there’s stuff I don’t know, and I’m not good enough to use this thing. And I get very upset about that.

Anaïs: That’s one of the things that you want to do with any technical documentation is that the first experience that anybody has, no matter the background, with your tool should be a success experience, right? Like people should look at it, use maybe one command, do one thing, one simple thing, and be like, “Yeah, this makes sense,” or, like, this was fun to do, right? Like, this first positive interaction. And it doesn’t have to be complex. And that’s what many people I think get wrong, that they try to show off how powerful a tool is, of like, oh, “My God, you can do all those things. It’s so exciting, right?” But [laugh] ultimately, if nobody can use it or if most of the people, 99% of the people who try it for the first time have a bad experience, it makes them feel uncomfortable or any negative emotion, then it’s really you’re approaching it from the wrong perspective, right?

Corey: That’s very apt. I think it’s so much of whether people stick with something long enough to learn it and find the sharp edges has to do with what their experience looks like. I mean, back when I was more or less useless when it comes to anything that looked like programming—because I was a sysadmin type—I started contributing to SaltStack. And what was amazing about that was Tom Hatch, the creator of the project had this pattern that he kept up for way too long, where whenever anyone submitted an issue, he said, “Great, well, how about you fix it?” And because we had a patch, like, “Well, I’m not good at programming.” He’s like, “That’s okay. No one is. Try it and we’ll see.”

And he accepted every patch and then immediately, you’d see another patch come in ten minutes later that fixed the problems in your patch. But it was the most welcoming and encouraging experience, and I’m not saying that’s a good workflow for an open-source maintainer, but he still remains one of the best humans I know, just from that perspective alone.

Anaïs: That’s amazing. I think it’s really about pointing out that there are different ways of doing open-source [laugh] and there is no one way to go about it. So, it’s really about—I mean, it’s about the community, ultimately. That’s what it boils down to, of you are dependent, as an open-source project, on the community, so what is the best experience that you can give them? If that’s something that you want to and can invest in, then yeah [laugh] that’s probably the best outcome for everybody.

Corey: I do have one more question, specifically around things that are more timely. Now, taking a quick look at Trivy and recent features, it seems like you’ve just now—now-ish—started supporting cloud scanning as well. Previously, it was effectively, “Oh, this scans configuration and containers. Okay, great.” Now, you’re targeting actually scanning cloud providers themselves. What does this change and what brought you to this place, as someone who very happily does not deal with AWS?

Anaïs: Yeah, totally. So, I just started using AWS, specifically to showcase this feature. So, if you look at the Aqua Open Source YouTube channel, you will find several tutorials that show you how to use that feature, among others.

Now, what I mentioned earlier in the podcast already is that Trivy is really versatile, it allows you to scan different aspects of your stack at different stages of your development lifecycle. And that’s made possible because Trivy is ultimately using different open-source projects under the hood. For example, if you want to scan your infrastructure-as-code misconfigurations, it’s using a tool called tfsec, specifically for Terraform. And then other tools for other scanning, for other security scanning. Now, we have—or had; it’s going to be probably deprecated—a tool called CloudSploit in the Aqua open-source project suite.

Now, that’s going to, kind of like, the functionality that CloudSploit was providing is going to get converted to become part of Trivy, so everything scanning-related is going to become part of Trivy that really, like, once you understand how Trivy works and all of the CLI commands in Trivy have exactly the same structure, it’s really easy to scan from container images to infrastructure-as-code, to generating s-bombs to scanning also now, your cloud infrastructure and Trivy can scan any of your AWS services for misconfigurations, and it’s using basically the AWS client under the hood to connect with the services of everything you have set up there, and then give you the list of misconfigurations. And once it has done the scan, you can then drill down further into the different aspects of your misconfigurations without performing the entire scan again, since you likely have lots and lots of resources, so you wouldn’t want to scan them every time again, right, when you perform the scan. So, once something has been scanned, Trivy will know whether the resource changed or not, it won’t scan it again. That’s the same way that in-classes scanning works right now. Once a container image has been scanned for vulnerabilities, it won’t scan the same container image again because that would just waste time. [laugh]. So yeah, do check it out. It’s our most recent feature, and it’s going to come out also to the other cloud providers out there. But we’re starting with AWS and this kind of forced me to finally [laugh] look at it for the sake of it. But I’m not going to be happy. [laugh].

Corey: No, I don’t think anyone is. It’s every time I see on a resume that someone says, “Oh, I’m an expert in AWS,” it’s, “No you’re not.” They have 400-some-odd services now. We have crossed the point long ago, where I can very convincingly talk about AWS services that do not exist to Amazonians and not get called out for it because who in the world knows what they run? And half of their services sound like something I made up to be funny, but they’re very real. It’s wild to me that it is a sprawling as it is and apparently continues to work as a viable business.

But no one knows all of it and everyone feels confused, lost, and overwhelmed every time they look at the AWS console. This has been my entire career in life for the last six years, and I still feel that way. So, I’m sure everyone else does, too.

Anaïs: And this is how misconfigurations happen, right? You’re confused about what you’re actually supposed to do and how you’re supposed to do it. And that’s, for example, with all the access rights in Google Cloud, something that I’m very familiar with, that completely overwhelms you and you get super frustrated by, and you don’t even know what you give access to. It’s like, if you’ve ever had to configure Discord user roles, it’s a similar disaster. You will not know which user has access to which. They kind of changed it and try to improve it over the past year, but it’s a similar issue that you face in cloud providers, just on a much larger-scale, not just on one chat channel. [laugh]. So.

Corey: I think that is probably a fair place to leave it. I really want to thank you for spending as much time with me as you have talking about the trials and travails of, well, this industry, for lack of a better term. If people want to learn more, where’s the best place to find you?

Anaïs: So, I have a weekly DevOps newsletter on my blog, which is anaisurl—like, how you spell U-R-L—and then dot com. anaisurl.com. That’s where I have all the links to my different channels, to all of the resources that are published where you can find out more as well. So, that’s probably the best place. Yeah.

Corey: And we will, of course, put a link to that in the show notes. I really want to thank you for being as generous with your time as you have been. Thank you.

Anaïs: Thank you for having me. It was great.

Corey: Anaïs, open-source developer advocate at Aqua Security. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry, insulting comment that I will never see because it’s buried under a whole bunch of minor or false-positive vulnerability reports.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Alex

Alex is the Chief Product Officer of Twingate, which he cofounded in 2019. Alex has held a range of product leadership roles in the enterprise software market over the last 16 years, including at Dropbox, where he was the first enterprise hire in the company's transformation from consumer to enterprise business. A focus of his product career has been using the power of design thinking to make technically complex products intuitive and easy to use. Alex graduated from Stanford University with a degree in Electrical Engineering.

Links Referenced:

  • twingate.com: https://twingate.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Sysdig. Sysdig secures your cloud from source to run. They believe, as do I, that DevOps and security are inextricably linked. If you wanna learn more about how they view this, check out their blog, it's definitely worth the read. To learn more about how they are absolutely getting it right from where I sit, visit Sysdig.com and tell them that I sent you. That's S Y S D I G.com. And my thanks to them for their continued support of this ridiculous nonsense.

Corey: This episode is sponsored in part by Honeycomb. When production is running slow, it’s hard to know where problems originate. Is it your application code, users, or the underlying systems? I’ve got five bucks on DNS, personally. Why scroll through endless dashboards while dealing with alert floods, going from tool to tool to tool that you employ, guessing at which puzzle pieces matter? Context switching and tool sprawl are slowly killing both your team and your business. You should care more about one of those than the other; which one is up to you. Drop the separate pillars and enter a world of getting one unified understanding of the one thing driving your business: production. With Honeycomb, you guess less and know more. Try it for free at honeycomb.io/screaminginthecloud. Observability: it’s more than just hipster monitoring.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. This promoted episode is brought to us by our friends at Twingate, and in addition to bringing you this episode, they also brought me a guest. Alex Marshall is the Chief Product Officer at Twingate. Alex, thank you for joining me, and what is a Twingate?

Alex: Yeah, well, thanks. Well, it’s great to be here. What is Twingate? Well, the way to think about Twingate is we’re really a network overlay layer. And so, the experience you have when you’re running Twingate as a user is that network resources or network destinations that wouldn’t otherwise be accessible to you or magically accessible to you and you’re properly authenticated and authorized to access them.

Corey: When you say it’s a network overlay, what I tend to hear and the context I usually see that in, in the real world is, “Well, we’re running some things in AWS and some things in Google Cloud, and I don’t know because of a sudden sharp blow to the head, maybe Azure as well, and how do you get all of the various security network models of security groups on one side to talk to their equivalent on the other side?” And the correct answer is generally that you don’t and you use something else that more or less makes the rest of that irrelevant. Is that the direction you’re coming at this from, or do you view it differently?

Alex: Yeah, so I think the way that we view this in terms of, like, why we decide to build a product in the first place is that if you look at, sort of like, the internet in 2022, like, there’s one thing that’s missing from the network routing table, which is authentication and authorization on each row [laugh]. And so, the way that we designed the product is we said, “Okay, we’re not going to worry about everything, basically, above the network layer and we’re going to focus on making sure that what we’re controlling with the client is looking at outbound network connections and making sure that when someone accesses something and only when they access it, that we check to make sure that they’re allowed access.” We’re basically holding those network connections until someone’s proven that they’re allowed to access to, then we let it go. And so, from the standpoint of, like, figuring out, like, security groups and all that kind of stuff, we’re basically saying, like, “Yeah, if you’re allowed to access the database in AWS, or your home assistant on your home network, fine, we’ll let you do that, but we’ll only let you go there once you’ve proven you’re allowed to. And then once you’re there, then you know, we’ll let you figure out how you want to authenticate into the destination system.” So, our view is, like, let’s start at the network layer, and then that solves a lot of problems.

Corey: When I call this a VPN, I know a couple of things are going to be true. One, you’re almost certainly going to correct me on that because this is all about Zero Trust. This is the Year of our Lord 2022, after all. But also what I round to what basically becomes a VPN to my mind, there are usually two implementations or implementation patterns that I think about. One of them is the idea of client access, where I have a laptop; I’m in a Starbucks; I want to connect to a thing. And the other has historically been considered, site to site, or I have a data center that I want to have constantly connected to my cloud environment. Which side of that mental model do you tend to fall in? Or is that the wrong way to frame it?

Alex: Mm-hm. The way we look at it and sort of the vision that we have for what the product should be, the problem that we should be solving for customers is what we want to solve for customers is that Twingate is a product that lets you be certain that your employees can work securely from anywhere. And so, you need a little bit of a different model to do that. And the two examples you gave are actually both entirely valid, especially given the fact that people just work from everywhere now. Like, resources everywhere, they use a lot of different devices, people work from lots of different networks, and so it’s a really hard problem to solve.

And so, the way that we look at it is that you really want to be running something or have a system in place that’s always taking into account the context that user is in. So, in your example of someone’s at a Starbucks, you know, in the public WiFi, last time I checked, Starbucks WiFi was unencrypted, so it’s pretty bad for security. So, what we should do is you should take that context into account and then make sure that all that traffic is encrypted. But at the same time, like, you might be in the corporate office, network is perfectly safe, but you still want to make sure that you’re authorizing people at the point in time they try to access something to make sure that they actually are entitled to access that database in the AWS network. And so, we’re trying to get people away from thinking about this, like, point-to-point connection with a VPN, where you know, the usual experience we’ve all had as employees is, “Great. Now, I need to fire up the VPN. My internet traffic is going to be horrible. My battery’s probably going to die. My—”

Corey: Pull out the manual token that rotates with an RSA—

Alex: Exactly.

Corey: —token that spits out a different digital code every 30 seconds if the battery hasn’t died or they haven’t gotten their seeds leaked again, and then log in and the rest; in some horrible implementations type that code after your password for some Godforsaken reason. Yeah, we’ve all been down that path and it’s like, “Yeah, just sign into the corporate VPN.” It’s like, “Did you just tell me to go screw myself because that’s what I heard.”

Alex: [laugh]. Exactly. And that is exactly the situation that we’re in. And the fact is, like, VPNs were invented a long time ago and they were designed to connect to networks, right? They were designed to connect a branch office to a corporate office, and they’re just to join all the devices on the network.

So, we’re really, like—everybody has had this experience of VPN is suffering from the fact that it’s the wrong tool for the job. Going back to, sort of like, this idea of, like, us being the network overlay, we don’t want to touch any traffic that isn’t intended to go to something that the company or the organization or the team wants to protect. And so, we’re only going to gate traffic that goes to those network destinations that you actually want to protect. And we’re going to make sure that when that happens, it’s painless. So, for example, like, you know, I don’t know, again, like, use your example again; you’ve been at Starbucks, you’ve been working your email, you don’t really need to access anything that’s private, and all of a sudden, like, you need to as part of your work that you’re doing on the Starbucks WiFi is access something that’s in AWS.

Well, then the moment you do that, then maybe you’re actually fine to access it because you’ve been authenticated, you know, and you’re within the window, it’s just going to work, right, so you don’t have to go through this painful process of firing up the VPN like you’re just talking about.

Corey: There are a number of companies out there that, first, self-described as being, “Oh, we do Zero Trust.” And when I hear that, what I immediately hear in my own mind is, “I have something to sell you,” which, fair enough, we live in an industry. We’re trying to have a society here. I get it. The next part that I wind up getting confused by then is, it seems like one of those deeply overloaded terms that exists to, more or less—in some cases to be very direct—well, we’ve been selling this thing for 15 years and that’s the buzzword, so now we’re going to describe it as the thing we do with a fresh coat of paint on it.

Other times it seems to be something radically different. And, on some level, I feel like I could wind up building an entire security suite out of nothing other than things self-billing themselves as Zero Trust. What is it that makes Twingate different compared to a wide variety of other offerings, ranging from Seam to whatever the hell an XDR might be to, apparently according to RSA, a breakfast cereal?

Alex: So, you’re right. Like, Zero Trust is completely, like, overused word. And so, what’s different about Twingate is that really, I think goes back to, like, why we started the company in the first place, which is that we started looking at the remote workspace. And this is, of course, before the pandemic, before everybody was actually working remotely and it became a really urgent problem.

Corey: During the pandemic, of course, a lot of the traditional VPN companies are, “Huh. Why is the VPN concentrator glowing white in the rack and melting? And it sounds like screaming. What’s going on?” Yeah, it turns out capacity provisioning and bottlenecking of an entire company tends to be a thing at scale.

Alex: And so, you’re right, like, that is exactly the conversation. We’ve had a bunch of customers over the last couple years, it’s like their VPN gateway is, like, blowing up because it used to be that 10% of the workforce used it on average, and all of a sudden everybody had to use it. What’s different about our approach in terms of what we observed when we started the company, is that what we noticed is that this term Zero Trust is kind of floating out there, but the only company that actually implemented Zero Trust was Google. So, if you think about the situations that you look at, Zero Trust is like, obvious. It’s like, it’s what you would want to do if you redesigned the internet, which is you’d want to say every network connection has to be authorized every single time it’s made.

But the internet isn’t actually designed that way. It’s designed default open instead of default closed. And so, we looked at the industry are, like, “Great. Like, Google’s done it. Google has, like, tons and tons of resources. Why hasn’t anyone else done it?”

And the example that I like to talk about when we talk about inception of the business is we went to some products that are out there that were implementing the right technological approach, and one of these products is still in use today, believe it or not, but I went to the documentation page, and I hit print, and it was almost 50 pages of documentation to implement it. And so, when you look at that, you’re, like, okay, like, maybe there’s a usability problem here [laugh]. And so, what we really, really focus on is, how do we make this product as easy as possible to deploy? And that gets into, like, this area of change management. And so, if you’re in IT or DevOps or engineering or security and you’re listening to this, I’m sure you’ve been through this process where it’s taken months to deploy something because it was just really technically difficult and because you had to change user behavior. So, the thing that we focus on is making sure that you didn’t have to change user behavior.

Corey: Every time you expect people to start doing things completely differently, congratulations, you’ve already lost before you’ve started.

Alex: Yes, exactly. And so, the difference with our product is that you can switch off the VPN one day, have people install a Twingate client, and then tomorrow, they still access things with exactly the same addresses they used before. And this seems like such a minor point, but the fact that I don’t have to rewrite scripts, I don’t have to change my SSH proxy configuration, I don’t have to do anything, all of those private DNS addresses or those private IP address, they’ll still work because of the way that our client works on the device.

Corey: So, what you’re saying is fundamental; you could even do a slow rollout. It doesn’t need to be a knife-switch cutover at two in the morning where you’re scrambling around and, “Oh, my God, we forgot the entire accounting department.”

Alex: Yep, that’s exactly right. And that is, like, an attraction of deploying this is that you can actually deploy it department by department and not have to change all your infrastructure at the same time. So again, it’s like pretty fundamental point here. It’s like, if you’re going to get adoption technology, it’s not just about how cool the technology is under the hood and how advanced it is; it’s actually thinking about from a customer and a business standpoint, like, how much is actually going to cost time-wise and effort-wise to move over to the new solution. So, we’ve really, really focused on that.

Corey: Yeah. That is generally one of those things, that seems to be the hardest approach. I mean, let’s back up a little bit here because I will challenge—likely—something that you said a few minutes ago, which is Google was the first and only company for a little while doing Zero Trust. Back in 2012, it turned out that we weren’t calling it that then, but that is fundamentally what I built out of the ten-person startup that I was at, where I was the first ops hire, which generally comes in right around Series B when developers realize, okay, we can no longer lie to ourselves that we know what we’re doing on an ops side. Everything’s on fire and no one can sleep through the night. Help, help, help. Which is fine.

I’ve never had tolerance or patience for ops people who insult people in those situations. It’s, “Well, they got far enough along to hire you, didn’t they? So, maybe show some respect.” But one of the things that I did was, being on the corporate network got you access to the printer in the corner and that was it. There was no special treatment of that network.

And I didn’t think much of it at the time, but I got some very strange looks and had some—uh, will call it interesting a decade later; most of the pain has faded—discussions with our auditor when we were going through some PCI work, and they showed up and said, “Great. Okay, where are the credentials for your directory?” And my response was, “Our what now?” And that’s when I realized there’s a certain point of scale. Back when I started as an independent consultant, everything I did for single-sign-on, for example, was my 1Password vault. Easy enough.

Now, that we’ve scaled up beyond that, I’m starting to see the value of things like single-sign-on in a way that I never did before, and in hindsight, I’d like to go back and do things very differently as a result. Scale matters. What is the point of scale that you find is your sweet spot? Is it one person trying to connect to a whole bunch of nonsense? Is it small to midsize companies—and we should probably bound that because to me, a big company is still one that has 200 people there?

Alex: To your original interesting point, which is that yeah, kudos to you for, like, implementing that, like, back then because we’ve had probably—

Corey: I was just being lazy and it was what was there. It’s like, “Why do I want to maintain a server in the closet? Honestly, I’m not sure that the office is that secure. And all it’s going to do—what I’m I going to put on that? A SharePoint server? Please. We’re using Macs.”

Alex: Yeah, exactly. Yeah. So it’s, we’ve had, like, I don’t know at this point, thousands of customer conversations. The number of people have actually gone down that route implementing things themselves as a very small number. And I think that just shows how hard it is. So again, like, kudos.

And I think the scale point is, I think, really critical. So, I think it’s changed over time, but actually, the point at which a customer gets to a scale where I think a solution has, like, leveraged high value is when you get to maybe only 50, 75 people, which is a pretty small business. And the reason is that that’s the point at which a bunch of tools start getting implemented a company, right? When you’re five people, you’re not going to install, like, an MDM or something on people’s devices, right? When you get to 50, 75, 100, you start hiring your first IT team members. That’s the point where them being able to, like, centralize management of things at the company becomes really critical.

And so, one of the other aspects that makes this a little bit different terms of approach is that what we see is that there’s a huge number of tools that have to be managed, and they have different configuration settings. You can’t even get consistency on MDM is across different platforms, necessarily, right? Like, Linux, Windows, and Mac are all going to have slight differences, and so what we’ve been working with the platform towards is actually being the centralization point where we integrate with these different systems and then pull together, like, a consistent way to create those authentication authorization policies I was talking about before. And the last thing on SSO, just to sort of reiterate that, I think that you’re talking about you’re seeing the value of that, the other thing that we’ve, like, made a deliberate decision on is that we’re not going to try to, like, re-solve, like, a bunch of these problems. Like, some of the things that we do on the user authentication point is that we rely on there being an SSO, like, user directory, that handles authentication, that handles, like, creating user groups. And we want to reuse that when people are using Twingate to control access to network destinations.

So, for us, like, it’s actually, you know, that point of scale comes fairly early. It only gets harder from there, and it’s especially when that IT team is, like, a relatively small number of people compared to number of employees where it becomes really critical to be able to leverage all the technology they have to deploy.

Corey: I guess this might be one of those areas where I’m not deep enough in your space to really see it the same way that you do, which is the whole reason I have people like you on the show: so I can ask these questions directly. What is the painful position that I find myself in that I should say, “Ah, I should bring Twingate in to solve this obnoxious, painful problem so I never have to think about it again.” What is it that you solve?

Alex: Yeah, I mean, I think for what our customers tell us, it’s providing a, like, consistent way to get access into, like, a wide variety of internal resources, and generally in multi-cloud environments. That’s where it gets, like, really tricky. And the consistency is, like, really important because you’re trying to provide access to your team—often like it’s DevOps teams, but all kinds of people can access these things—trying to write access is a multiple different environments, again, there’s a consistency problem where there are multiple different ways to provide that, and there isn’t a single place to manage all that. And so, it gets really challenging to understand who has access to what, makes sure that credentials expire when they’re supposed to expire, make sure that all the routing inside those remote destinations is set up correctly. And it just becomes, like, a real hassle to manage those things.

So, that’s the big one. And usually where people are coming from is that they’ve been using VPN to do that because they didn’t know anything better exists, or they haven’t found anything that’s easy enough to deploy, right? So, that’s really the problem that they’re running into.

Corey: There’s also a lot of tribal knowledge that gets passed down. The oral tradition of, “I have this problem. What should I do? I know, I will consult the wise old sage.” “Well, where can you find the wise old sage?” “Under the rack of servers, swearing at them.” “Great, cool. Well, use a VPN. That’s what we’ve used since time immemorial.” And then the sins are visited onto yet another generation.

There’s a sense that I have that companies that are started now are going to have a radically different security posture and a different way of thinking about these things than the quote-unquote, “Legacy companies.”—legacy, of course, being that condescending engineering term for ‘it makes money—who are migrating their way into a brave new world because they had the temerity to found themselves as companies before 2012.

Alex: Absolutely. When we’re working with customers, there is a sort of a sweet spot, both in terms of, like, the size and role that we were talking about before, but also just in terms of, like, where they are, in, sort of like, the sort of lifecycle of their company. And I think one of the most exciting things for us is that we get to work with companies that are kind of figuring this stuff out for the first time and they’re taking a fresh look at, like, what the capabilities are out there in the landscape. And that’s, I think, what makes this whole space, like, super, super interesting.

There’s some really, really fantastic things you can do. Just give you an example, again, that I think might resonate with your audience quite a bit is this whole topic of automation, right? Your time at the tribal knowledge of, like, “Oh, of course. You know, we set up a VPN and so on.” One of the things that I don’t think is necessarily obvious in this space is that for the teams that—at companies that are deploying, configuring, managing internal network infrastructure, is that in the past, you’ve had to make compromises on infrastructure in order to accommodate access, right?

Because it’s kind of a pain to deploy a bunch of, like, VPN gateways, mostly for the end-user because they got to, like, choose which one they’re connecting to. You potentially had to open up traffic routes to accommodate a VPN gateway that you wouldn’t otherwise want to open up. And so, one of the things that’s, like, really sort of fascinating about, like, a new way of looking at things is that what we allow with Twingate—and part of this is because we’ve really made sure that the product is, like, API-first in the very beginning, which allows us to very easily integrate in with things, like, Terraform and Pulumi for deployment automation, is that now you have a new way of looking at things, which is that you can build a network infrastructure that you want with the data flow rules that you want, and very easily provide access into, like, points of that infrastructure, whether that’s an entire subnet or just a single host somewhere. I think these are the ways, like, the capabilities have been realized are possible until they, sort of like, understand some of these new technologies.

Corey: This episode is sponsored in part by our friend EnterpriseDB. EnterpriseDB has been powering enterprise applications with PostgreSQL for 15 years. And now EnterpriseDB has you covered wherever you deploy PostgreSQL on-premises, private cloud, and they just announced a fully-managed service on AWS and Azure called BigAnimal, all one word. Don’t leave managing your database to your cloud vendor because they’re too busy launching another half-dozen managed databases to focus on any one of them that they didn’t build themselves. Instead, work with the experts over at EnterpriseDB. They can save you time and money, they can even help you migrate legacy applications—including Oracle—to the cloud. To learn more, try BigAnimal for free. Go to biganimal.com/snark, and tell them Corey sent you.

Corey: This feels like one of those technologies where the place that a customer starts from and where they wind up going are very far apart. Because I can see the metaphorical camel’s nose under the tent flap being, “Ah, this is a VPN except it doesn’t suck. Great.” But once you wind up with effectively an overlay network connecting all the things that you care about within an organization, it feels like that unlocks a whole universe of possibility.

Alex: Mm-hm. Yeah, definitely. I mean, I think you hit the nail on the head there. Like, a lot of people approach us because they’re having a lot of pain with VPN and all the operational difficulties they were talking about earlier, but I think what sort of starts to open up is there’s some, sort of like, not obvious things that happen. And one of them is that all of a sudden, when you can limit access at a network connection level, you start to think about, like, credentials and access management a little differently, right?

So, one of the problems that well-known is people set a bastion host. And they set bastion host so that there’s, like, a limited way into the network and all the, you know, keys are stored in that bastion host and so on. So, you basically have a system where fine, we had bastion host set up because, A, we want limited ingress, and B, we want to make sure that we know exactly who has access to our internal resources. You could do away with that and with a simple, like, configuration change, you can basically say, “Even if this employee for whatever reason, we’ve forgotten to remove—revoke their SSH keys, even if they still have those keys, they can’t access the destination because we’re blocking network access at their actual device,” then you have a very different way to restrict access. So, it’s still important to manage credentials, but you now have a way to actually block things out at a network level. And I think it’s like when people start to realize that these capabilities are possible that they definitely start thinking about things a little bit differently. VPNs just don’t allow this, like, level of granularity.

Corey: I am a firm believer in the idea that any product with any kind of longevity gets an awful lot of its use case and product-market fit not from the people building it, but from the things that those folks learn from their customers. What did you learn from customers rolling out Twingate that reshaped how you thought about the space, or surprised you as far as use cases go?

Alex: Yeah, so I think it’s a really interesting question because one of the benefits of having a small business and being early on is that you have very close relationships with all your customers and they’re really passionate about your product. And what that leads to is just a lot of, sort of like, knowledge sharing around, like, how they’re using your product, which then helps inform the types of things that we build. So, one of the things that we’ve done internally to help us learn, but then also help us respond more quickly to customers, is we have this group called Twingate Labs. And it’s really just a group of folks that are outside the engineering org that are just allowed to build whatever they want to try to prove out, like, interesting concepts. And a lot of those—I say a lot; honestly, probably all of those concepts have come from our customers, and so we’ve been able to, like, push the boundaries on that.

And so, it just gave you an example, I mean, AWS can be sometimes a challenging product to manage and interact with, and so that team has, for example, built capabilities, again, using that just the regular Twingate API to show that it’s possible to automatically configure resources in AWS based on tags. Now, that’s not something that’s in our product, but it’s us showing our customers that, you know, we can respond quickly to them and then they actually, like, try to accommodate some, like, these special use cases they have. And if that works out, then great, we’ll pull it into the product, right? So, I think that’s, like, the nice thing about serving a smaller businesses is that you get a lot of that back and forth to your customers and they help us generate ideas, too.

Corey: One thing that stands out to me from the testimonials from customers you have on your website has been a recurring theme that crops up that speaks to I guess, once I spend more than ten seconds thinking about it, one of the most obvious reasons that I would say, “Oh, Twingate? That sounds great for somebody else. We’re never rolling it out here.” And that is the ease of adoption into environments that are not greenfield because I don’t believe that something like this product will ever get deployed to something greenfield because this is exactly the kind of problem that you don’t realize exists and don’t have to solve for until it’s too late because you already have that painful problem. It’s an early optimization until suddenly, it’s something you should have done six months ago. What is the rolling it out process for a company that presumably already is built out, has hired a bunch of people, and they already have something that, quote-unquote, “Works,” for granting access to things?

Alex: Mm-hm. Yeah, so the beauty is that you can really deploy this side-by-side with an existing solution, so—whatever it happens to be; I mean, whether it’s a VPN or something else—is you can put the side-by-side and the deployment process, just to talk a little bit about the architecture; we’ve talked a lot about this client that runs on the user’s device, but on the remote network side, just to be really clear on this, there’s a component called a connector that gets deployed inside the remote network, and it does not have to be installed on every single destination host. You’re sort of thinking about it, sort of like this routing point inside that network, and that connector controls what traffic is allowed to go to internal locations based on the rules. So, from a deployment standpoint, it’s really just put a connector in place and put it in place in whatever subnet you want to provide access to.

And so you’re—unlikely, but if your entire company has one subnet, great. You’re done with one connector. But it does mean you can sort of gradually roll it out as it goes. And the connector can be deployed in a bunch of different environments, so we’re just talking with AWS. Maybe it’s inside a VPC, but we have a lot of people that actually just want to control access to specific services inside a Kubernetes cluster, and so you can deploy it as a container, right inside Kubernetes. And so, you can be, like, really specific about how you do that and then gradually roll it out to teams as they need it and without having to necessarily on that day actually shut off the old solution.

So, just to your comment, by the way, on the greenfield versus, sort of like, brownfield, I think the greenfield story, I think, is changing a little bit, I think, especially to your comment earlier around younger companies. I think younger companies are realizing that this type of capability is an option and that they want to get in earlier. But the reality is that, you know, 98% of people are really in the established network situation, and so that’s where that rollout process is really important.

Corey: As you take a look throughout what you’re seeing customers doing, what you see the industry doing as a result of that—because customers are, in fact, the industry, let’s be clear here—what do you think is, I guess, the next wave of security offerings? I guess what I’m trying to do here is read the tea leaves and predict what the buzzwords will be all over the place that next RSA. But on a slightly more serious note, what do you see this is building towards? What are the trends that you’re identifying in the space?

Alex: There’s a couple of things that we see. So one, sort of, way to look at this is that we’re sort of in this, like, Third Wave. And I think these things change more slowly than—with all due respect to marketers—than marketers would [laugh] have you believe. And so, thinking about where we are, there’s, like, Wave One is, like, good old happy days, we’re all in the office, like, your computer can’t move, like, all the data is in the office, like, everything is in one place, right?

Corey: What if someone steals your desktop? Well, they’re probably going to give themselves a hernia because that thing’s heavy. Yeah.

Alex: Exactly. And is it really worth stealing, right? But the Wave One was really, like, network security was actually just physical security, to that point; that’s all it was, just, like, physically secure the premises.

Wave Two—and arguably you could say we’re kind of still in this—is actually the transition to cloud. So, let’s convert all CapEx to OpEx, but that also introduces a different problem, which is that everything is off-network. So, you have to, like, figure out, you know, what you do about that.

But Wave Three is really I think—and again, just to be clear, I think Wave Two, there are, like, multi-decade things that happen—and I’d say we’re in the middle of, like, Wave Three. And I think that everyone is still, like, gradually adapting to this, which is what we describe it as sort of people everywhere, applications are everywhere, people are using a whole bunch of different devices, right? There is no such thing as BYOD in the early-2000s, late-90s, and people are accessing things from all kinds of different networks. And this presents a really, really challenging problem. So, I would argue, to your question, I think we’re still in the middle of that Wave Three and it’s going to take a long time to see that play through the industry. Just, things change slowly. That tribal knowledge takes time to change.

The other thing that I think we very strongly believe in is that—and again, this is, sort of like, coming from our customers, too—is that people basically with security industry have had a tough time trying things out and adopting them because a lot of vendors have put a lot of blockers in place of doing that. There’s no public documentation; you can’t just go use the product. You got to talk to a salesperson who then filters you through—

Corey: We have our fifth call with the sales team. We’re hoping this is the one where they’ll tell us how much it costs.

Alex: Exactly. Or like, you know, now you get to the sales engineer, so you gradually adopt this knowledge. But ultimately, people just want to try the darn thing [laugh], right? So, I think we’re big believers that I think hopefully, what we’ll see in the security industry is that—we’re trying to set an example here—is really that there’s an old way of doing things, but a new way of doing things is make the product available for people to use, document the heck out of it, explain all the different use cases that exist for how to be successful your product, and then have these users actually then reach out to you when they want to have more in-depth conversation about things. So, those are the two big things, I’d say. I don’t know if those are translated buzzwords at RSA, but those are two big trends we see.

Corey: I look forward to having you back in a year or two and seeing how close we get to the reality. “Well, I guess we didn’t see that acronym coming, but don’t worry. They’ve been doing it for the last 15 years under different names, so it works out.” I really want to thank you for being as generous with your time as you have been. If people want to learn more, where should they go?

Alex: Well, as we’re just talking about, you try the product at twingate.com. So, that should be your first stop.

Corey: And we will of course put links to that in the show notes. Thank you so much for being as forthcoming as you have been about all this stuff. I really appreciate your time.

Alex: Yeah, thank you, Corey. I really appreciate it. Thanks.

Corey: Alex Marshall, Chief Product Officer at Twingate. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with a long angry ranty comment about what you hated about the episode, which will inevitably get lost when it fails to submit because your crappy VPN concentrator just dropped it on the floor.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Martin

Martin Casado is a general partner at the venture capital firm Andreessen Horowitz where he focuses on enterprise investing. He was previously the cofounder and chief technology officer at Nicira, which was acquired by VMware for $1.26 billion in 2012. While at VMware, Martin was a fellow, and served as senior vice president and general manager of the Networking and Security Business Unit, which he scaled to a $600 million run-rate business by the time he left VMware in 2016.

Martin started his career at Lawrence Livermore National Laboratory where he worked on large-scale simulations for the Department of Defense before moving over to work with the intelligence community on networking and cybersecurity. These experiences inspired his work at Stanford where he created the software-defined networking (SDN) movement, leading to a new paradigm of network virtualization. While at Stanford he also cofounded Illuminics Systems, an IP analytics company, which was acquired by Quova Inc. in 2006.

For his work, Martin was awarded both the ACM Grace Murray Hopper award and the NEC C&C award, and he’s an inductee of the Lawrence Livermore Lab’s Entrepreneur’s Hall of Fame. He holds both a PhD and Masters degree in Computer Science from Stanford University.

Martin serves on the board of ActionIQ, Ambient.ai, Astranis, dbt Labs, Fivetran, Imply, Isovalent, Kong, Material Security, Netlify, Orbit, Pindrop Security, Preset, RapidAPI, Rasa, Tackle, Tecton, and Yubico.

Links:

  • Yet Another Infra Group Discord Server: https://discord.gg/f3xnJzwbeQ
  • “The Cost of Cloud, a Trillion Dollar Paradox” - https://a16z.com/2021/05/27/cost-of-cloud-paradox-market-cap-cloud-lifecycle-scale-growth-repatriation-optimization/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Honeycomb. When production is running slow, it’s hard to know where problems originate. Is it your application code, users, or the underlying systems? I’ve got five bucks on DNS, personally. Why scroll through endless dashboards while dealing with alert floods, going from tool to tool to tool that you employ, guessing at which puzzle pieces matter? Context switching and tool sprawl are slowly killing both your team and your business. You should care more about one of those than the other; which one is up to you. Drop the separate pillars and enter a world of getting one unified understanding of the one thing driving your business: production. With Honeycomb, you guess less and know more. Try it for free at honeycomb.io/screaminginthecloud. Observability: it’s more than just hipster monitoring.

Corey: This episode is sponsored in part by our friends at Sysdig. Sysdig secures your cloud from source to run. They believe, as do I, that DevOps and security are inextricably linked. If you wanna learn more about how they view this, check out their blog, it's definitely worth the read. To learn more about how they are absolutely getting it right from where I sit, visit Sysdig.com and tell them that I sent you. That's S Y S D I G.com. And my thanks to them for their continued support of this ridiculous nonsense.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’m joined today by someone who has taken a slightly different approach to being—well, we’ll call it cloud skepticism here. Martin Casado is a general partner at Andreessen Horowitz and has been on my radar starting a while back, based upon a piece that he wrote focusing on the costs of cloud and how repatriation is going to grow. You wrote that in conjunction with your colleague, Sarah Wang. Martin, thank you so much for joining me. What got you onto that path?

Martin: So, I want to be very clear, just to start with is, I think cloud is the biggest innovation that we’ve seen in infrastructure, probably ever. It’s a core part of the industry. I think it’s very important, I think every company’s going to be using cloud, so I’m very pro-cloud. I just think the nature of how you use clouds is shifting. And that was the focus.

Corey: When you first put out your article in conjunction with your colleague as well, like, I saw it and I have to say that this was the first time I’d really come across any of your work previously. And I have my own biases that I started from, so my opening position on reading it was this is just some jerk who’s trying to say something controversial and edgy to get attention. That’s my frickin job. Excuse me, sir. And who is this clown?

So, I started digging, and what I found really changed my perspective because as mentioned at the start of the show, you are a general partner at Andreessen Horowitz, which means you are a VC. You are definitionally almost the archetype of a VC in that sense. And to me, being a venture capitalist means the most interesting thing about you is that you write a large check consisting of someone else’s money. And that’s never been particularly interesting.

Martin: [laugh].

Corey: You kind of cut against that grain and that narrative. You have a master’s and a PhD in computer science from Stanford; you started your career at one of the national labs—Laurence Livermore, if memory serves—you wound up starting a business, Nicira, if I’m pronouncing that correctly—

Martin: Yeah, yeah, yeah.

Corey: That you then sold to VMware in 2012, back at a time when that was a noble outcome, rather than a state of failure because VMware is not exactly what it once was. You ran a $600 million a year business while you were there. Basically, the list of boards that you’re on is lengthy enough and notable enough that it sounds almost like you’re professionally bored, so I don’t—

Martin: [laugh].

Corey: So, looking at this, it’s okay, this is someone who actually knows what he is talking about, not just, “Well, I talked to three people in pitch meetings and I now think I know what is going on in this broader industry.” You pay attention, and you’re connected, disturbingly well, to what’s going on, to the point where if you see something, it is almost certainly rooted in something that is happening. And it’s a big enough market that I don’t think any one person can keep their finger on the pulse of everything. So, that’s when I started really digging into it, paying attention, and more or less took a lot of what you wrote as there are some theses in here that I want to prove or disprove. And I spent a fair bit of time basically threatening, swindling, and bribing people with infinite cups of coffee in order to start figuring out what is going on.

And I am begrudgingly left with no better conclusion than you have a series of points in here that are very challenging to disprove. So, where do you stand today, now that, I guess, the whole rise and fall of the hype around your article on cloud repatriation—which yes, yes, we’ll put a link to it in the show notes if people want to go there—but you’ve talked about this in a lot of different contexts. Having had the conversations that you’ve had, and I’m sure some very salty arguments with people who have a certain vested interest in you being wrong, do you wind up continuing to stand by the baseline positions that you’ve laid out, or have they evolved into something more nuanced?

Martin: So yeah, I definitely want to point out, so this was work done with Sarah Wang was also at Andreessen Horowitz; she’s also a GP. She actually did the majority of the analysis and she’s way smarter than I am. [laugh]. And so, I’m just very—feel very lucky to work with her on this. And I want to make sure she gets due credit on this.

So, let’s talk about the furor. So like, I actually thought that this was kind of interesting and it started a good discussion, but instead, like, [laugh] the amount of, like, response pieces and, like, angry emails I got, and [laugh] like, I mean it just—and I kind of thought to myself, like, “Why are people so upset?” I think there’s three reasons. I’m going to go through them very quickly because they’re interesting.

So, the first one is, like, you’re right, like, I’m a VC. I think people see a VC and they’re like, oh, lack of credibility, lack of accountability, [laugh], you know, doesn’t know what they’re doing, broad pattern matcher. And, like, I will say, like, I did not necessarily write this as a VC; I wrote this as somebody that’s, like, listen, my PhD is an infrastructure; my company was an infrastructure. It’s all data center stuff. I had a $600 million a year data center business that sold infrastructure into data centers. I’ve worked with all of the above. Like, I’ve worked with Amazon, I’ve—

Corey: So, you sold three Cisco switches?

Martin: [laugh]. That’s right.

Corey: I remember those days. Those were awesome, but not inexpensive.

Martin: [laugh]. That’s right. Yeah, so like, you know, I had 15 years. It’s kind of a culmination of that experience. So, that was one; I just think that people see VC and they have a reaction.

The second one is, I think people still have the first cloud wars fresh in their memories and so they just don’t know how to think outside of that. So, a lot of the rebuttals were first cloud war rebuttals. Like, “Well, but internal IT is slow and you can’t have the expertise.” But like, they just don’t apply to the new world, right? Like, listen, if you’re Cloudflare, to say that you can’t run, like, a large operation is just silly. If you went to Cloudflare and you’re like, “Listen, you can’t run your own infrastructure,” like, they’d take out your sucker and pat you on the head. [laugh].

Corey: And not for nothing, if you try to run what they’re doing on other cloud providers from a pure bandwidth perspective, you don’t have a company anymore, regardless of how well funded you are. It’s a never-full money pit that just sucks all of the money. And I’ve talked to a number of very early idea stage companies that aren’t really founded yet about trying to do things like CDN-style work or streaming video, and a lot of those questions start off with well, we did some back-of-the-envelope math around AWS data transfer pricing, and if our numbers are right, when we scale, we’ll be spending $65,000 on data transfer every minute. What did we get wrong?

And it’s like, “Oh, yeah, you realize that one thing is per hour not per minute, so slight difference there. But no, you’re basically correct. Don’t do it.” And yeah, no one pays retail price at that volume, but they’re not going to give you a 99.999% discount on these things, so come up with a better plan. Cloudflare’s business will not work on AWS, full stop.

Martin: Yep, yep. So, I legitimately know, basically, household name public companies that are software companies that anybody listening to this knows the name of these companies, who have product lines who have 0% margins because they’re [laugh] basically, like, for every dollar they make, they pay a dollar to Amazon. Like, this is a very real thing, right? And if you go to these companies, these are software infrastructure companies; they’ve got very talented teams, they know how to build, like, infrastructure. To tell them that like, “Well, you know, you can’t build your own infrastructure,” or something is, I mean, it’s like telling, like, an expert in the business, they can’t do what they do; this is what they do. So, I just think that part of the furor, part of the uproar, was like, I just think people were stuck in this cloud war 1.0 mindset.

I think the third thing is, listen, we’ve got an oligopoly, and they employ a bunch of people, and they’ve convinced a bunch of people they’re right, and it’s always hard to change that. And I also think there’s just a knee-jerk reaction to these big macro shifts. And it was the same thing we did to software-defined networking. You know, like, my grad school work was trying to change networking to go from hardware to software. I remember giving a talk at Cisco, and I was, like, this kind of like a naive grad student, and they literally yelled at me out of the room. They’re like, it’ll never work.

Corey: They tried to burn you as a witch, as I recall.

Martin: [laugh]. And so, your specific question is, like, have our views evolved? But the first one is, I think that this macro downturn really kind of makes the problem more acute. And so, I think the problem is very, very real. And so, I think the question is, “Okay, so what happens?”

So, let’s say if you’re building a new software company, and you have a choice of using, like, one of the Big Three public clouds, but it impacts your margins so much that it depresses your share price, what do you do? And I think that we thought a lot more about what the answers there are. And the ones that I think that we’re seeing is, some actually are; companies are building their own infrastructure. Like, very famously MosaicML is building their own infrastructure. Fly.io, -building their own infrastructure.

Mighty—you know, Suhail’s company—building his own infrastructure. Cloudflare has their own infrastructure. So, I think if you’re an infrastructure provider, a very reasonable thing to do is to build your own infrastructure. If you’re not a core infrastructure provider, you’re not; you can still use somebody’s infrastructure that’s built at a better cost point.

So, for example, if I’m looking at a CDN tier, I’m going to use Fly.io, right? I mean, it’s like, it’s way cheaper, the multi-region is way better, and so, like, I do think that we’re seeing, like, almost verticalized clouds getting built out that address this price point and, like, these new use cases. And I think this is going to start happening more and more now. And we’re going to see basically almost the delamination of the cloud into these verticalized clouds.

Corey: I think there’s also a question of scale, where if you’re starting out in the evening tonight, to—I want to build, I don’t know Excel as a service or something. Great. You’re pretty silly if you’re not going to start off with a cloud provider, just because you can get instant access to resources, and if your product catches on, you scale out without having to ever go back and build it as quote-unquote “Enterprise grade,” as opposed to having building it on cheap servers or Raspberry Pis or something floating around. By the time that costs hit a certain point—and what that point is going to depend on your stage of company and lifecycle—you’re remiss if you don’t at least do an analysis on is this the path we want to continue on for the service that we’re offering?

And to be clear, the answer to this is almost entirely going to be bounded by the context of your business. I don’t believe that companies as a general rule, make ill-reasoned decisions. I think that when we see a decision a company makes, by and large, there’s context or constraints that we don’t see that inform that. I know, it’s fun to dunk on some of the large companies’ seemingly inscrutable decisions, but I will say, having had the privilege to talk to an awful lot of execs in an awful lot of places—particularly on this show—I don’t find myself encountering a whole lot of people in those roles who I come away with thinking that they’re a few fries short of a Happy Meal. They generally are very well reasoned in why they do what they do. It’s just a question of where we think the future is going on some level.

Martin: Yep. So, I think that’s absolutely right. So, to be a little bit more clear on what I think is happening with the cloud, which is I think every company that gets created in tech is going to use the cloud for something, right? They’ll use it for development, the website, test, et cetera. And many will have everything in the cloud, right?

So, the cloud is here to stay, it’s going to continue to grow, it’s a very important piece of the ecosystem, it’s very important piece of IT. I’m very, very pro cloud; there’s a lot of value. But the one area that’s under pressure is if your product is SaaS if your product is selling Software as a Service, so then your product is basically infrastructure, now you’ve got a product cost model that includes the infrastructure itself, right? And if you reduce that, that’s going to increase your margin. And so, every company that’s doing that should ask the question, like, A, is the Big Three the right one for me?

Maybe a verticalized cloud—like for example, whatever Fly or Mosaic or whatever is better because the cost is better. And I know how to, you know, write software and run these things, so I’ll use that. They’ll make that decision or maybe they’ll build their own infrastructure. And I think we’re going to see that decision happening more and more, exactly because now software is being offered as a service and they can do that. And I just want to make the point, just because I think it’s so important, that the clouds did exactly this to the hardware providers. So, I just want to tell a quick story, just because for me, it’s just so interesting. So—

Corey: No, please, I was only really paying attention to this market from 2016 or so. There was a lot of the early days that I was using as a customer, but I wasn’t paying attention to the overall industry trends. Please, storytime. This is how I learned things. I hang out with smart people and I come away a little bit smarter than when I started.

Martin: [laugh]. This is, like, literally my fa—this is why this is one of my favorite topics is what I’m about to tell you, which is, so the clouds have always had this argument, right? The big clouds, three clouds, they’re like, “Listen, why would you build your own cloud? Because, like, you don’t have the expertise, and it’s hard and you don’t have economies of scale.” Right?

And the answer is you wouldn’t unless it impacts your share price, right? If it impacts your share price, then of course you would because it makes economic sense. So, the clouds had that exact same dilemma in 2005, right? So, in 2005, Google and Amazon and Microsoft, they looked at their COGS, they looked like, “Okay, I’m offering a cloud. If I look at the COGS, who am I paying?”

And it turns out, there was a bunch of hardware providers that had 30% margins or 70% margins. They’re like, “Why am I paying Cisco these big margins? Why am I paying Dell these big margins?” Right? So, they had the exact same dilemma.

And all of the arguments that they use now applied then, right? So, the exact same arguments, for example, “AWS, you know nothing about hardware. Why would you build hardware? You don’t have the expertise. These guys sell to everybody in the world, you don’t have the economies of scale.”

So, all of the same arguments applied to them. And yet… and yes because it was part of COGS] that it impacted the share price, they can make the economic argument to actually build hardware teams and build chips. And so, they verticalized, right? And so, it just turns out if the infrastructure becomes parts of COGS, it makes sense to optimize that infrastructure. And I would say, the Big Three’s foray into OEMs and hardware is a much, much, much bigger leap than an infrastructure company foraying into building their own infrastructure.

Corey: There’s a certain startup cost inherent to all these things. And the small version of that we had in every company that we started in a pre-cloud era: renting virtual computers from vendors was a thing, but it was still fraught and challenging and things that we use, then, like, GoGrid no longer exist, for good reason. But the alternative was, “Great, I’m going to start building and seeing if this thing has any traction.” Well, you need to go lease a rack somewhere and buy servers from Dell, and they’re going to do the fast expedited option, which means only six short weeks until they show up in the data center and then gets sent away because they weren’t expecting to receive them. And you wind up with this entire universe of hell between cross-connects and all the rest.

And that’s before you can ever get anything in front of customers or users to see what happens. Now, it’s a swipe of a credit card away and your evening’s experiments round up to 25 cents. That was significant. Having to make these significant tens of thousands of dollars of investment just to launch is no longer true. And I feel like that was a great equalizer in some respects.

Martin: Yeah, I think that—

Corey: And that cost has been borne by the astonishing level of investment that the cloud providers themselves have made. And that basically means that we don’t have to. But it does come at a cost.

Martin: I think it’s also worth pointing out that it’s much easier to stand up your own infrastructure now than it has been in the past, too. And so, I think that there’s a gradient here, right? So, if you’re building a SaaS app, [laugh] you would be crazy not to use the cloud, you just be absolutely insane, right? Like, what do you know about core infrastructure? You know, what do you know about building a back-end? Like, what do you know about operating these things? Go focus on your SaaS app.

Corey: The calluses I used to have from crimping my own Ethernet patch cables in data centers have faded by now. I don’t want them to come back. Yeah, we used to know how to do these things. Now, most people in most companies do not have that baseline of experience, for excellent reasons. And I wouldn’t wish that on the current generation of engineers, except for the ones I dislike.

Martin: However, that is if you’re building an application. Almost all of my investments are people that are building infrastructure. [laugh]. They’re already doing these hardcore backend things; that’s what they do: they sell infrastructure. Would you think, like, someone, like, at Databricks doesn’t understand how to run infr—of course it does. I mean, like, or Snowflake or whatever, right?

And so, this is a gradient. On the extreme app end, you shouldn’t be thinking about infrastructure; just use the cloud. Somewhere in the middle, maybe you start on the cloud, maybe you don’t. As you get closer to being a cloud service, of course you’re going to build your own infrastructure.

Like, for example—listen, I mean, I’ve been mentioning Fly; I just think it’s a great example. I mean, Fly is a next-generation CDN, that you can run compute on, where they build their own infrastructure—it’s a great developer experience—and they would just be silly. Like, they couldn’t even make the cost model work if they did it on the cloud. So clearly, there’s a gradient here, and I just think that you would be remiss and probably negligent if you’re selling software not to have this conversation, or at least do the analysis.

Corey: This episode is sponsored in part by our friend EnterpriseDB. EnterpriseDB has been powering enterprise applications with PostgreSQL for 15 years. And now EnterpriseDB has you covered wherever you deploy PostgreSQL on-premises, private cloud, and they just announced a fully-managed service on AWS and Azure called BigAnimal, all one word. Don’t leave managing your database to your cloud vendor because they’re too busy launching another half-dozen managed databases to focus on any one of them that they didn’t build themselves. Instead, work with the experts over at EnterpriseDB. They can save you time and money, they can even help you migrate legacy applications—including Oracle—to the cloud. To learn more, try BigAnimal for free. Go to biganimal.com/snark, and tell them Corey sent you.

Corey: I think there’s also a philosophical shift, where a lot of the customers that I talk to about their AWS bills want to believe something that is often not true. And what they want to believe is that their AWS bill is a function of how many customers they have.

Martin: Oh yeah.

Corey: In practice, it is much more closely correlated with how many engineers they’ve hired. And it sounds like a joke, except that it’s not. The challenge that you have when you choose to build in a data center is that you have bounds around your growth because there are capacity concerns. You are going to run out of power, cooling, and space to wind up having additional servers installed. In cloud, you have an unbounded growth problem.

S3 is infinite storage, and the reason I’m comfortable saying that is that they can add hard drives faster than you can fill them. For all effective purposes, it is infinite amounts of storage. There is no forcing function that forces you to get rid of things. You spin up an instance, the natural state of it in a data center as a virtual machine or a virtual instance, is that it’s going to stop working two to three years left on maintain when a raccoon hauls it off into the woods to make a nest or whatever the hell raccoons do. In cloud, you will retire before that instance does is it gets migrated to different underlying hosts, continuing to cost you however many cents per hour every hour until the earth crashes into the sun, or Amazon goes bankrupt.

That is the trade-off you’re making. There is no forcing function. And it’s only money, which is a weird thing to say, but the failure mode of turning something off mistakenly that takes things down, well that’s disastrous to your brand and your company. Just leaving it up, well, it’s only money. It’s never a top-of-mind priority, so it continues to build and continues to build and continues to build until you’re really forced to reckon with a much larger problem.

It is a form of technical debt, where you’ve kicked the can down the road until you can no longer kick that can. Then your options are either go ahead and fix it or go back and talk to you folks, and it’s time for more money.

Martin: Yeah. Or talk to you. [laugh].

Corey: There is that.

Martin: No seriously, I think everybody should, honestly. I think this is a board-level concern for every compa—I sit on a lot of boards; I see this. And this has organically become a board-level concern. I think it should become a conscious board-level concern of, you know, cloud costs, impact COGS. Any software company has it; it always becomes an issue, and so it should be treated as a first-class problem.

And if you’re not thinking through your options—and I think by the way, your company is a great option—but if you’re not thinking to the options, then you’re almost fiduciarily negligent. I think the vast, vast majority of people and vast majority of companies are going to stay on the cloud and just do some basic cost controls and some just basic hygiene and they’re fine and, like, this doesn’t touch them. But there are a set of companies, particularly those that sell infrastructure, where they may have to get more aggressive. And that ecosystem is now very vibrant, and there’s a lot of shifts in it, and I think it’s the most exciting place [laugh] in all of IT, like, personally in the industry.

Corey: One question I have for you is where do you draw the line around infrastructure companies. I tend to have an evolving view of it myself, where things that are hard and difficult do not become harder with time. It used to require a deep-level engineer with a week to kill to wind up compiling and building a web server. Now, it is evolved and evolved and evolved; it is check a box on a webpage somewhere and you’re serving a static website. Managed databases, I used to think, were something that were higher up the stack and not infrastructure. Today, I’d call them pretty clearly infrastructure.

Things seem to be continually, I guess, a slipping beneath the waves to borrow an iceberg analogy. And it’s only the stuff that you can see that is interesting and differentiated, on some level. I don’t know where the industry is going at all, but I continue to think of infrastructure companies as being increasingly broad.

Martin: Yeah, yeah, yeah. This is my favorite question. [laugh]. I’m so glad you asked. [laugh].

Corey: This was not planned to be clear.

Martin: No, no, no. Listen, I am such an infrastructure maximalist. And I’ve changed my opinion on this so much in the last three years. So, it used to be the case—and infrastructure has a long history of, like, calling the end of infrastructure. Like, every decade has been the end of infrastructure. It’s like, you build the primitives and then everything else becomes an app problem, you know?

Like, you build a cloud, and then we’re done, you know? You build the PC and then we’re done. And so, they are even very famous talks where people talk about the end of systems when we’ve be built everything right then. And I’ve totally changed my view. So, here’s my current view.

My current view is, infrastructure is the only, really, differentiation in systems, in all IT, in all software. It’s just infrastructure. And the app layer is very important for the business, but the app layer always sits on infrastructure. And the differentiations in app is provided by the infrastructure. And so, the start of value is basically infrastructure.

And the design space is so huge, so huge, right? I mean, we’ve moved from, like, PCs to cloud to data. Now, the cloud is decoupling and moving to the CDN tier. I mean, like, the front-end developers are building stuff in the browser. Like, there’s just so much stuff to do that I think the value is always going to accrue to infrastructure.

So, in my view, anybody that’s improving the app accuracy or performance or correctness with technology is an infrastructure company, right? And the more of that you do, [laugh] the more infrastructure you are. And I think, you know, in 30 years, you and I are going to be old, and we’re going to go back on this podcast. We’re going to talk and there’s going to be a whole bunch of infrastructure companies that are being created that have accrued a lot of value. I’m going to say one more thing, which is so—okay, this is a sneak preview for the people listening to this that nobody else has heard before.

So Sarah, and I are back at it again, and—the brilliant Sarah, who did the first piece—and we’re doing another study. And the study is if you look at public companies and you look at ones that are app companies versus infrastructure companies, where does the value accrue? And there’s way, way more app companies; there’s a ton of app companies, but it turns out that infrastructure companies have higher multiples and accrue more value. And that’s actually a counter-narrative because people think that the business is the apps, but it just turns out that’s where the differentiation is. So, I’m just an infra maximalist. I think you could be an infra person your entire career and it’s the place to be. [laugh].

Corey: And this is the real value that I see of looking at AWS bills. And our narrative is oh, we come in and we fix the horrifying AWS bill. And the naive pass is, “Oh, you cut the bill and make it lower?” Not always. Our primary focus has been on understanding it because you get a phone-number-looking bill from AWS. Great, you look at it, what’s driving the cost? Storage.

Okay, great. That doesn’t mean anything to the company. They want to know what teams are doing this. What’s it going to cost for them to add another thousand monthly active users? What is the increase in cost? How do they wind up identifying their bottlenecks? How do they track and assign portions of their COGS to different aspects of their service? How do they trace the flow of capital for their organization as they’re serving their customers?

And understanding the bill and knowing what to optimize and what not to becomes increasingly strategic business concern.

Martin: Yeah.

Corey: That’s the fun part. That’s the stuff I don’t see that software has a good way of answering, just because there’s no way to use an API to gain that kind of business context. When I started this place, I thought I was going to be building software. It turns out, there’s so many conversations that have to happen as a part of this that cannot be replicated by software. I mean, honestly, my biggest competitor for all this stuff is Microsoft Excel because people want to try and do it themselves internally. And sometimes they do a great job, sometimes they don’t, but it’s understanding their drivers behind their cost. And I think that is what was often getting lost because the cloud obscures an awful lot of that.

Martin: Yeah. I think even just summarize this whole thing pretty quickly, which is, like, I do think that organically, like, cloud cost has become a board-level issue. And I think that the shift that founders and execs should make is to just, like, treat it like a first-class problem upfront. So, what does that mean? Minimally, it means understanding how these things break down—A, to your point—B, there’s a number of tools that actually help with onboarding of this stuff. Like, Vantage is one that I’m a fan of; it just provides some visibility.

And then the third one is if you’re selling Software as a Service, that’s your core product or software, and particularly it’s a infrastructure, if you don’t actually do the analysis on, like, how this impacts your share price for different cloud costs, if you don’t do that analysis, I would say your fiduciarily negligent, just because the impact would be so high, especially in this market. And so, I think, listen, these three things are pretty straightforward and I think anybody listening to this should consider them if you’re running a company, or you’re an executive company.

Corey: Let’s be clear, this is also the kind of problem that when you’re sitting there trying to come up with an idea for a business that you can put on slide decks and then present to people like you, these sounds like the paradise of problems to have. Like, “Wow, we’re successful and our business is so complex and scaled out that we don’t know where exactly a lot of these cost drivers are coming from.” It’s, “Yeah, that sounds amazing.” Like, I remember those early days, back when all I was able to do and spend time on and energy on was just down to the idea of, ohh, I’m getting business cards. That’s awesome. That means I’ve made it as a business person.

Spoiler: it did not. Having an aggressive Twitter presence, that’s what made me as a business person. But then there’s this next step and this next step and this next step and this next step, and eventually, you look around and realize just how overwrought everything you’ve built is and how untangling it just becomes a bit of a challenge and a hell of a mess. Now, the good part is at that point of success, you can bring people in, like, a CFO and a finance team who can do some deep-level analysis to help identify what COGS is—or in some cases, have some founders, explain what COGS is to you—and understand those structures and how you think about that. But it always feels like it’s a trailing problem, not an early problem that people focus on.

Martin: I’ll tell you the reason. The reason is because this is a very new phenomenon that it’s part of COGS. It’s literally five years new. And so, we’re just catching up. Even now, this discussion isn’t what it was when we first wrote the post.

Like, now people are pretty educated on, like, “Oh yeah, like, this is really an issue. Oh, yeah. It contributes to COGS. Oh, yeah. Like, our stock price gets hit.” Like, it’s so funny to watch, like, the industry mature in real-time. And I think, like, going forward, it’s just going to be obvious that this is a board-level issue; it’s going to be obvious this is, like, a first-class consideration. But I agree with you. It’s like, listen, like, the industry wasn’t ready for it because we didn’t have public companies. A lot of public companies, like, this is a real issue. I mean really we’re talking about the last five, seven years.

Corey: It really is neat, just in real time watching how you come up with something that sounds borderline heretical, and in a relatively short period of time, becomes accepted as a large-scale problem, and now it’s now it is fallen off of the hype train into, “Yeah, this is something to be aware of.” And people’s attention spans have already jumped to the next level and next generation of problem. It feels like this used to take way longer for these cycles, and now everything is so rapid that I almost worry that between the time we’re recording this and the time that it publishes in a few weeks, what is going to have happened that makes this conversation irrelevant? I didn’t used to have to think like that. Now, I do.

Martin: Yeah, yeah, yeah, for sure. Well, just a couple of things. I want to talk about, like, one of the reasons that accelerated this, and then when I think is going forward. So, one of the reasons this was accelerated was just the macro downturn. Like, when we wrote the post, you could make the argument that nobody cares about margins because it’s all about growth, right?

And so, like—and even then, it still saved a bunch of money, but like, a lot of people were like, “Listen, the only thing that matters is growth.” Now, that’s absolutely not the case if you look at public market valuations. I mean, people really care about free cash flow, they really care about profitability, and they really care about margins. And so, it’s just really forced the issue. And it also, like, you know, made kind of what we were saying very, very clear.

I would say, you know, as far as shifts that are going, I think one of the biggest shifts is for every back-end developer, there’s, like, a hundred front-end developers. It’s just crazy. And those front-end developers—

Corey: A third of a DevOps engineer.

Martin: [laugh]. True. I think those front-end developers are getting, like, better tools to build complete apps, right? Like, totally complete apps, right? Like they’ve got great JavaScript frameworks that coming out all the time.

And so, you could argue that actually a secular technology change—which is that developers are now rebuilding apps as kind of front-end applications—is going to pull compute away from the clouds anyways, right? Like if instead of, like, the app being some back-end thing running in AWS, but instead is a front-end thing, you know, running in a browser at the CDN tier, while you’re still using the Big Three clouds, it’s being used in a very different way. And we may have to think about it again differently. Now, this, again, is a five-year going forward problem, but I do feel like there are big shifts that are even changing the way that we currently think about cloud now. And we’ll see.

Corey: And if those providers don’t keep up and start matching those paradigms, there’s going to be an intermediary shim layer of companies that wind up converting their resources and infrastructure into things that suit this new dynamic, and effectively, they’re going to become the next version of, I don’t know, Level 3, one of those big underlying infrastructure companies that most people have never heard of or have to think about because they’re not doing anything that’s perceived as interesting.

Martin: Yeah, I agree. And I honestly think this is why Cloudflare and Cloudflare work is very interesting. This is why Fly is very interesting. It’s a set of companies that are, like, “Hey, listen, like, workloads are moving to the front-end and, you know, you need compute closer to the user and multi-region is really important, et cetera.” So, even as we speak, we’re seeing kind of shifts to the way the cloud is moving, which is just exciting. This is why it’s, like, listen, infrastructure is everything. And, like, you and I like if we live to be 200, we can do [laugh] a great infrastructure work every year.

Corey: I’m terrified, on some level, that I’ll still be doing the exact same type of thing in 20 years.

Martin: [laugh].

Corey: I like solving different problems as we go. I really want to thank you for spending so much time talking to me today. If people want to learn more about what you’re up to, slash beg you for other people’s money or whatnot, where’s the best place for them to find you?

Martin: You know, we’ve got this amazing infrastructure Discord channel. [laugh].

Corey: Really? I did not know that.

Martin: I love it. It’s, like, the best. Yeah, my favorite thing to do is drink coffee and talk about infrastructure. And like, I posted this on Twitter and we’ve got, like, 600 people. And it’s just the best thing. So, that’s honestly the best way to have these discussions. Maybe can you put, like, the link in, like, the show notes?

Corey: Oh, absolutely. It is already there in the show notes. Check the show notes. Feel free to join the infrastructure Discord. I will be there waiting for you.

Martin: Yeah, yeah, yeah. That’ll be fantastic.

Corey: Thank you so much for being so generous with your time. I appreciate it.

Martin: This was great. Likewise, Corey. You’re always a class act and I really appreciate that about you.

Corey: I do my best. Martin Casado, general partner at Andreessen Horowitz. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry comment telling me that I got it completely wrong and what check you wrote makes you the most interesting.

Announcer: The content here is for informational purposes only and should not be taken as legal, business, tax, or investment advice, or be used to evaluate any investment or security and is not directed at any investors or potential investors in any a16z fund. For more details, please see a16z.com/disclosures.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Matt

Matt is a Sr. Architect in Belfast, an AWS DevTools Hero, Serverless Architect, Author and conference speaker.

He is focused on creating the right environment for empowered teams to rapidly deliver business value in a well-architected, sustainable and serverless-first way.

You can usually find him sharing reusable, well architected, serverless patterns over at cdkpatterns.com or behind the scenes bringing CDK Day to life.

Links Referenced:

  • Previous guest appearance: https://www.lastweekinaws.com/podcast/screaming-in-the-cloud/slinging-cdk-knowledge-with-matt-coulter/
  • The CDK Book: https://thecdkbook.com/
  • Twitter: https://twitter.com/NIDeveloper

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. One of the best parts about, well I guess being me, is that I can hold opinions that are… well, I’m going to be polite and call them incendiary, and that’s great because I usually like to back them in data. But what happens when things change? What happens when I learn new things?

Well, do I hold on to that original opinion with two hands at a death grip or do I admit that I was wrong in my initial opinion about something? Let’s find out. My guest today returns from earlier this year. Matt Coulter is a senior architect since he has been promoted at Liberty Mutual. Welcome back, and thanks for joining me.

Matt: Yeah, thanks for inviting me back, especially to talk about this topic.

Corey: Well, we spoke about it a fair bit at the beginning of the year. And if you’re listening to this, and you haven’t heard that show, it’s not that necessary to go into; mostly it was me spouting uninformed opinions about the CDK—the Cloud Development Kit, for those who are unfamiliar—I think of it more or less as what if you could just structure your cloud resources using a programming language you claim to already know, but in practice, copy and paste from Stack Overflow like the rest of us? Matt, you probably have a better description of what the CDK is in practice.

Matt: Yeah, so we like to say it’s imperative code written in a declarative way, or declarative code written in an imperative way. Either way, it lets you write code that produces CloudFormation. So, it doesn’t really matter what you write in your script; the point is, at the end of the day, you still have the CloudFormation template that comes out of it. So, the whole piece of it is that it’s a developer experience, developer speed play, that if you’re from a background that you’re more used to writing a programming language than a YAML, you might actually enjoy using the CDK over writing straight CloudFormation or SAM.

Corey: When I first kicked the tires on the CDK, my first initial obstacle—which I’ve struggled with in this industry for a bit—is that I’m just good enough of a programmer to get myself in trouble. Whenever I wind up having a problem that StackOverflow doesn’t immediately shine a light on, my default solution is to resort to my weapon of choice, which is brute force. That sometimes works out, sometimes doesn’t. And as I went through the CDK, a couple of times in service to a project that I’ll explain shortly, I made a bunch of missteps with it. The first and most obvious one is that AWS claims publicly that it has support in a bunch of languages: .NET, Python, there’s obviously TypeScript, there’s Go support for it—I believe that went generally available—and I’m sure I’m missing one or two, I think? Aren’t I?

Matt: Yeah, it’s: TypeScript, JavaScript, Python Java.Net, and Go. I think those are the currently supported languages.

Corey: Java. That’s the one that I keep forgetting. It’s the block printing to the script that is basically Java cursive. The problem I run into, and this is true of most things in my experience, when a company says that we have deployed an SDK for all of the following languages, there is very clearly a first-class citizen language and then the rest that more or less drift along behind with varying degrees of fidelity. In my experience, when I tried it for the first time in Python, it was not a great experience for me.

When I learned just enough JavaScript, and by extension TypeScript, to be dangerous, it worked a lot better. Or at least I could blame all the problems I ran into on my complete novice status when it comes to JavaScript and TypeScript at the time. Is that directionally aligned with what you’ve experienced, given that you work in a large company that uses this, and presumably, once you have more than, I don’t know, two developers, you start to take on aspects of a polyglot shop no matter where you are, on some level?

Matt: Yeah. So personally, I jump between Java, Python, and TypeScript whenever I’m writing projects. So, when it comes to the CDK, you’d assume I’d be using all three. I typically stick to TypeScript and that’s just because personally, I’ve had the best experience using it. For anybody who doesn’t know the way CDK works for all the languages, it’s not that they have written a custom, like, SDK for each of these languages; it’s a case of it uses a Node process underneath them and the language actually interacts with—it’s like the compiled JavaScript version is basically what they all interact with.

So, it means there are some limitations on what you can do in that language. I can’t remember the full list, but it just means that it is native in all those languages, but there are certain features that you might be like, “Ah,” whereas, in TypeScript, you can just use all of TypeScript. And my first inclination was actually, I was using the Python one and I was having issues with some compiler errors and things that are just caused by that process. And it’s something that talking in the cdk.dev Slack community—there is actually a very active—

Corey: Which is wonderful, I will point out.

Matt: [laugh]. Thank you. There is actually, like, an awesome Python community in there, but if you ask them, they would all ask for improvements to the language. So, personally if someone’s new, I always recommend they start with TypeScript and then branch out as they learn the CDK so they can understand is this a me problem, or is this a problem caused by the implementation?

Corey: From my perspective, I didn’t do anything approaching that level of deep dive. I took a shortcut that I find has served me reasonably well in the course of my career, when I’m trying to do something in Python, and you pull up a tutorial—which I’m a big fan of reading experience reports, and blog posts, and here’s how to get started—and they all have the same problem, which is step one, “Run npm install.” And that’s “Hmm, you know, I don’t recall that being a standard part of the Python tooling.” It’s clearly designed and interpreted and contextualized through a lens of JavaScript. Let’s remove that translation layer, let’s remove any weird issues I’m going to have in that transpilation process, and just talk in the language it written in. Will this solve my problems? Oh, absolutely not, but it will remove a subset of them that I am certain to go blundering into like a small lost child trying to cross an eight-lane freeway.

Matt: Yeah. I’ve heard a lot of people say the same thing. Because the CDK CLI is a Node process, you need it no matter what language you use. So, if they were distributing some kind of universal binary that just integrated with the languages, it would definitely solve a lot of people’s issues with trying to combine languages at deploy time.

Corey: One of the challenges that I’ve had as I go through the process of iterating on the project—but I guess I should probably describe it for those who have not been following along with my misadventures; I write blog posts about it from time to time because I need a toy problem to kick around sometimes because my consulting work is all advisory and I don’t want to be a talking head-I have a Twitter client called lasttweetinaws.com. It’s free; go and use it. It does all kinds of interesting things for authoring Twitter threads.

And I wanted to deploy that to a bunch of different AWS regions, as it turns out, 20 or so at the moment. And that led to a lot of interesting projects and having to learn how to think about these things differently because no one sensible deploys an application simultaneously to what amounts to every AWS region, without canary testing, and having a phased rollout in the rest. But I’m reckless, and honestly, as said earlier, a bad programmer. So, that works out. And trying to find ways to make this all work and fit together led iteratively towards me discovering that the CDK was really kind of awesome for a lot of this.

That said, there were definitely some fairly gnarly things I learned as I went through it, due in no small part to help I received from generous randos in the cdk.dev Slack team. And it’s gotten to a point where it’s working, and as an added bonus, I even mostly understand what he’s doing, which is just kind of wild to me.

Matt: It’s one of those interesting things where because it’s a programming language, you can use it out of the box the way it’s designed to be used where you can just write your simple logic which generates your CloudFormation, or you can do whatever crazy logic you want to do on top of that to make your app work the way you want it to work. And providing you’re not in a company like Liberty, where I’m going to do a code review, if no one’s stopping you, you can do your crazy experiments. And if you understand that, it’s good. But I do think something like the multi-region deploy, I mean, with CDK, if you’d have a construct, it takes in a variable that you can just say what the region is, so you can actually just write a for loop and pass it in, which does make things a lot easier than, I don’t know, try to do it with a YAML, which you can pass in parameters, but you’re going to get a lot more complicated a lot quicker.

Corey: The approach that I took philosophically was I wrote everything in a region-agnostic way. And it would be instantiated and be told what region to run it in as an environment variable that CDK deploy was called. And then I just deploy 20 simultaneous stacks through GitHub Actions, which invoke custom runners that runs inside of a Lambda function. And that’s just a relatively basic YAML file, thanks to the magic of GitHub Actions matrix jobs. So, it fires off 20 simultaneous processes and on every commit to the main branch, and then after about two-and-a-half minutes, it has been deployed globally everywhere and I get notified on anything that fails, which is always fun and exciting to learn those things.

That has been, overall, just a really useful experiment and an experience because you’re right, you could theoretically run this as a single CDK deploy and then wind up having an iterate through a list of regions. The challenge I have there is that unless I start getting into really convoluted asynchronous concurrency stuff, it feels like it’ll just take forever. At two-and-a-half minutes a region times 20 regions, that’s the better part of an hour on every deploy and no one’s got that kind of patience. So, I wound up just parallelizing it a bit further up the stack. That said, I bet they are relatively straightforward ways, given the async is a big part of JavaScript, to do this simultaneously.

Matt: One of the pieces of feedback I’ve seen about CDK is if you have multiple stacks in the same project, it’ll deploy them one at a time. And that’s just because it tries to understand the dependencies between the stacks and then it works out which one should go first. But a lot of people have said, “Well, I don’t want that. If I have 20 stacks, I want all 20 to go at once the way you’re saying.” And I have seen that people have been writing plugins to enable concurrent deploys with CDK out of the box. So, it may be something that it’s not an out-of-the-box feature, but it might be something that you can pull in a community plug-in to actually make work.

Corey: Most of my problems with it at this point are really problems with CloudFormation. CloudFormation does not support well, if at all, secure string parameters from the AWS Systems Manager parameter store, which is my default go-to for secret storage, and Secrets Manager is supported, but that also cost 40 cents a month per secret. And not for nothing, I don’t really want to have all five secrets deployed to Secrets Manager in every region this thing is in. I don’t really want to pay $20 a month for this basically free application, just to hold some secrets. So, I wound up talking to some folks in the Slack channel and what we came up with was, I have a centralized S3 bucket that has a JSON object that lives in there.

It’s only accessible from the deployment role, and it grabs that at deploy time and stuffs it into environment variables when it pushes these things out. That’s the only stateful part of all of this. And it felt like that is, on some level, a pattern that a lot of people would benefit from if it had better native support. But the counterargument that if you’re only deploying to one or two regions, then Secrets Manager is the right answer for a lot of this and it’s not that big of a deal.

Matt: Yeah. And it’s another one of those things, if you’re deploying in Liberty, we’ll say, “Well, your secret is unencrypted at runtime, so you probably need a KMS key involved in that,” which as you know, the costs of KMS, it depends on if it’s a personal solution or if it’s something for, like, a Fortune 100 company. And if it’s personal solution, I mean, what you’re saying sounds great that it’s IAM restricted in S3, and then that way only at deploy time can be read; it actually could be a custom construct that someone can build and publish out there to the construct library—or the construct hub, I should say.

Corey: To be clear, the reason I’m okay with this, from a security perspective is one, this is in a dedicated AWS account. This is the only thing that lives in that account. And two, the only API credentials we’re talking about are the application-specific credentials for this Twitter client when it winds up talking to the Twitter API. Basically, if you get access to these and are able to steal them and deploy somewhere else, you get no access to customer data, you get—or user data because this is not charge for anything—you get no access to things that have been sent out; all you get to do is submit tweets to Twitter and it’ll have the string ‘Last Tweet in AWS’ as your client, rather than whatever normal client you would use. It’s not exactly what we’d call a high-value target because all the sensitive to a user data lives in local storage in their browser. It is fully stateless.

Matt: Yeah, so this is what I mean. Like, it’s the difference in what you’re using your app for. Perfect case of, you can just go into the Twitter app and just withdraw those credentials and do it again if something happens, whereas as I say, if you’re building it for Liberty, that it will not pass a lot of our Well-Architected reviews, just for that reason.

Corey: If I were going to go and deploy this at a more, I guess, locked down environment, I would be tempted to find alternative approaches such as having it stored encrypted at rest via KMS in S3 is one option. So, is having global DynamoDB tables that wind up grabbing those things, even grabbing it at runtime if necessary. There are ways to make that credential more secure at rest. It’s just, I look at this from a real-world perspective of what is the actual attack surface on this, and I have a really hard time just identifying anything that is going to be meaningful with regard to an exploit. If you’re listening to this and have a lot of thoughts on that matter, please reach out I’m willing to learn and change my opinion on things.

Matt: One thing I will say about the Dynamo approach you mentioned, I’m not sure everybody knows this, but inside the same Dynamo table, you can scope down a row. You can be, like, “This row and this field in this row can only be accessed from this one Lambda function.” So, there’s a lot of really awesome security features inside DynamoDB that I don’t think most people take advantage of, but they open up a lot of options for simplicity.

Corey: Is that tied to the very recent announcement about Lambda getting SourceArn as a condition key? In other words, you can say, “This specific Lambda function,” as opposed to, “A Lambda in this account?” Like that was a relatively recent Advent that I haven’t fully explored the nuances of.

Matt: Yeah, like, that has opened a lot of doors. I mean, the Dynamo being able to be locked out in your row has been around for a while, but the new Lambda from SourceArn is awesome because, yeah, as you say, you can literally say this thing, as opposed to, you have to start going into tags, or you have to start going into something else to find it.

Corey: So, I want to talk about something you just alluded to, which is the Well-Architected Framework. And initially, when it launched, it was a whole framework, and AWS made a lot of noise about it on keynote stages, as they are want to do. And then later, they created a quote-unquote, “Well-Architected Tool,” which let’s be very direct, it’s the checkbox survey form, at least the last time I looked at it. And they now have the six pillars of the Well-Architected Framework where they talk about things like security, cost, sustainability is the new pillar, I don’t know, absorbency, or whatever the remainders are. I can’t think of them off the top of my head. How does that map to your experience with the CDK?

Matt: Yeah, so out of the box, the CDK from day one was designed to have sensible defaults. And that’s why a lot of the things you deploy have opinions. I talked to a couple of the Heroes and they were like, “I wish it had less opinions.” But that’s why whenever you deploy something, it’s got a bunch of configuration already in there. For me, in the CDK, whenever I use constructs, or stacks, or deploying anything in the CDK, I always build it in a well-architected way.

And that’s such a loaded sentence whenever you say the word ‘well-architected,’ that people go, “What do you mean?” And that’s where I go through the six pillars. And in Liberty, we have a process, it used to be called SCORP because it was five pillars, but not SCORPS [laugh] because they added sustainability. But that’s where for every stack, we’ll go through it and we’ll be like, “Okay, let’s have the discussion.” And we will use the tool that you mentioned, I mean, the tool, as you say, it’s a bunch of tick boxes with a text box, but the idea is we’ll get in a room and as we build the starter patterns or these pieces of infrastructure that people are going to reuse, we’ll run the well-architected review against the framework before anybody gets to generate it.

And then we can say, out of the box, if you generate this thing, these are the pros and cons against the Well-Architected Framework of what you’re getting. Because we can’t make it a hundred percent bulletproof for your use case because we don’t know it, but we can tell you out of the box, what it does. And then that way, you can keep building so they start off with something that is well documented how well architected it is, and then you can start having—it makes it a lot easier to have those conversations as they go forward. Because you just have to talk about the delta as they start adding their own code. Then you can and you go, “Okay, you’ve added these 20 lines. Let’s talk about what they do.” And that’s why I always think you can do a strong connection between infrastructure-as-code and well architected.

Corey: As I look through the actual six pillars of the Well-Architected Framework: sustainability, cost optimization, performance, efficiency, reliability, security, and operational excellence, as I think through the nature of what this shitpost thread Twitter client is, I am reasonably confident across all of those pillars. I mean, first off, when it comes to the cost optimization pillar, please, don’t come to my house and tell me how that works. Yeah, obnoxiously the security pillar is sort of the thing that winds up causing a problem for this because this is an account deployed by Control Tower. And when I was getting this all set up, my monthly cost for this thing was something like a dollar in charges and then another sixteen dollars for the AWS config rule evaluations on all of the deploys, which is… it just feels like a tax on going about your business, but fine, whatever. Cost and sustainability, from my perspective, also tend to be hand-in-glove when it comes to this stuff.

When no one is using the client, it is not taking up any compute resources, it has no carbon footprint of which to speak, by my understanding, it’s very hard to optimize this down further from a sustainability perspective without barging my way into the middle of an AWS negotiation with one of its power companies.

Matt: So, for everyone listening, watch as we do a live well-architected review because—

Corey: Oh yeah, I expect—

Matt: —this is what they are. [laugh].

Corey: You joke; we should do this on Twitter one of these days. I think would be a fantastic conversation. Or Twitch, or whatever the kids are using these days. Yeah.

Matt: Yeah.

Corey: And again, if so much of it, too, is thinking about the context. Security, you work for one of the world’s largest insurance companies. I shitpost for a living. The relative access and consequences of screwing up the security on this are nowhere near equivalent. And I think that’s something that often gets lost, per the perfect be the enemy of the good.

Matt: Yeah that’s why, unfortunately, the Well-Architected Tool is quite loose. So, that’s why they have the Well-Architected Framework, which is, there’s a white paper that just covers anything which is quite big, and then they wrote specific lenses for, like, serverless or other use cases that are shorter. And then when you do a well-architected review, it’s like loose on, sort of like, how are you applying the principles of well-architected. And the conversation that we just had about security, so you would write that down in the box and be, like, “Okay, so I understand if anybody gets this credential, it means they can post this Last Tweet in AWS, and that’s okay.”

Corey: The client, not the Twitter account, to be clear.

Matt: Yeah. So, that’s okay. That’s what you just mark down in the well-architected review. And then if we go to day one on the future, you can compare it and we can go, “Oh. Okay, so last time, you said this,” and you can go, “Well, actually, I decided to—” or you just keep it as a note.

Corey: “We pivoted. We’re a bank now.” Yeah.

Matt: [laugh]. So, that’s where—we do more than tweets now. We decided to do microtransactions through cryptocurrency over Twitter. I don’t know but if you—

Corey: And that ends this conversation. No no. [laugh].

Matt: [laugh]. But yeah, so if something changes, that’s what the well-architected reviews for. It’s about facilitating the conversation between the architect and the engineer. That’s all it is.

Corey: This episode is sponsored in part by our friend EnterpriseDB. EnterpriseDB has been powering enterprise applications with PostgreSQL for 15 years. And now EnterpriseDB has you covered wherever you deploy PostgreSQL on-premises, private cloud, and they just announced a fully-managed service on AWS and Azure called BigAnimal, all one word. Don’t leave managing your database to your cloud vendor because they’re too busy launching another half-dozen managed databases to focus on any one of them that they didn’t build themselves. Instead, work with the experts over at EnterpriseDB. They can save you time and money, they can even help you migrate legacy applications—including Oracle—to the cloud. To learn more, try BigAnimal for free. Go to biganimal.com/snark, and tell them Corey sent you.

Corey: And the lens is also helpful in that this is a serverless application. So, we’re going to view it through that lens, which is great because the original version of the Well-Architected Tool is, “Oh, you built this thing entirely in Lambda? Have you bought some reserved instances for it?” And it’s, yeah, why do I feel like I have to explain to AWS how their own systems work? This makes it a lot more streamlined and talks about this, though, it still does struggle with the concept of—in my case—a stateless app. That is still something that I think is not the common path. Imagine that: my code is also non-traditional. Who knew?

Matt: Who knew? The one thing that’s good about it, if anybody doesn’t know, they just updated the serverless lens about, I don’t know, a week or two ago. So, they added in a bunch of more use cases. So, if you’ve read it six months ago, or even three months ago, go back and reread it because they spent a good year updating it.

Corey: Thank you for telling me that. That will of course wind up in next week’s issue of Last Week in AWS. You can go back and look at the archives and figure out what week record of this then. Good work. One thing that I have learned as well as of yesterday, as it turns out, before we wound up having this recording—obviously because yesterday generally tends to come before today, that is a universal truism—is it I had to do a bit of refactoring.

Because what I learned when I was in New York live-tweeting the AWS Summit, is that the Route 53 latency record works based upon where your DNS server is. Yeah, that makes sense. I use Tailscale and wind up using my Pi-hole, which lives back in my house in San Francisco. Yeah, I was always getting us-west-1 from across the country. Cool.

For those weird edge cases like me—because this is not the common case—how do I force a local region? Ah, I’ll give it its own individual region prepend as a subdomain. Getting that to work with both the global lasttweetinaws.com domain as well as the subdomain on API Gateway through the CDK was not obvious on how to do it.

Randall Hunt over at Caylent was awfully generous and came up with a proof-of-concept in about three minutes because he’s Randall, and that was extraordinarily helpful. But a challenge I ran into was that the CDK deploy would fail because the way that CloudFormation was rendered in the way it was trying to do stuff, “Oh, that already has that domain affiliated in a different way.” I had to do a CDK destroy then a CDK deploy for each one. Now, not the end of the world, but it got me thinking, everything that I see around the CDK more or less distills down to either greenfield or a day one experience. That’s great, but throw it all away and start over is often not what you get to do.

And even though Amazon says it’s always day one, those of us in, you know, real companies don’t get to just treat everything as brand new and throw away everything older than 18 months. What is the day two experience looking like for you? Because you clearly have a legacy business. By legacy, I of course, use it in the condescending engineering term that means it makes actual money, rather than just telling really good stories to venture capitalists for 20 years.

Matt: Yeah. We still have mainframes running that make a lot of money. So, I don’t mock legacy at all.

Corey: “What’s that piece of crap do?” “Well, about $4 billion a year in revenue. Perhaps show some respect.” It’s a common refrain.

Matt: Yeah, exactly. So yeah, anyone listening, don’t mock legacy because as Corey says, it is running the business. But for us when it comes to day two, it’s something that I’m actually really passionate about this in general because it is really easy. Like I did it with CDK patterns, it’s really easy to come out and be like, “Okay, we’re going to create a bunch of starter patterns, or quickstarts”—or whatever flavor that you came up with—“And then you’re going to deploy this thing, and we’re going to have you in production and 30 seconds.” But even day one later that day—not even necessarily day two—it depends on who it was that deployed it and how long they’ve been using AWS.

So, you hear these stories of people who deployed something to experiment, and they either forget to delete, it cost them a lot of money or they tried to change it and it breaks because they didn’t understand what was in it. And this is where the community starts to diverge in their opinions on what AWS CDK should be. There’s a lot of people who think that at the minute CDK, even if you create an abstraction in a construct, even if I create a construct and put it in the construct library that you get to use, it still unravels and deploys as part of your deploy. So, everything that’s associated with it, you don’t own and you technically need to understand that at some point because it might, in theory, break. Whereas there’s a lot of people who think, “Okay, the CDK needs to go server side and an abstraction needs to stay an abstraction in the cloud. And then that way, if somebody is looking at a 20-line CDK construct or stack, then it stays 20 lines. It never unravels to something crazy underneath.”

I mean, that’s one pro tip thing. It’d be awesome if that could work. I’m not sure how the support for that would work from a—if you’ve got something running on the cloud, I’m pretty sure AWS [laugh] aren’t going to jump on a call to support some construct that I deployed, so I’m not sure how that will work in the open-source sense. But what we’re doing at Liberty is the other way. So, I mean, we famously have things like the software accelerator that lets you pick a pattern or create your pipelines and you’re deployed, but now what we’re doing is we’re building a lot of telemetry and automated information around what you deployed so that way—and it’s all based on Well-Architected, common theme. So, that way, what you can do is you can go into [crosstalk 00:26:07]—

Corey: It’s partially [unintelligible 00:26:07], and partially at a glance, figure out okay, are there some things that can be easily remediated as we basically shift that whole thing left?

Matt: Yeah, so if you deploy something, and it should be good the second you deploy it, but then you start making changes. Because you’re Corey, you just start adding some stuff and you deploy it. And if it’s really bad, it won’t deploy. Like, that’s the Liberty setup. There’s a bunch of rules that all go, “Okay, that’s really bad. That’ll cause damage to customers.”

But there’s a large gap between bad and good that people don’t really understand the difference that can cost a lot of money or can cause a lot of grief for developers because they go down the wrong path. So, that’s why what we’re now building is, after you deploy, there’s a dashboard that’ll just come up and be like, “Hey, we’ve noticed that your Lambda function has too little memory. It’s going to be slow. You’re going to have bad cold starts.” Or you know, things like that.

The knowledge that I have had the gain through hard fighting over the past couple of years putting it into automation, and that way, combined with the well-architected reviews, you actually get me sitting in a call going, “Okay, let’s talk about what you’re building,” that hopefully guides people the right way. But I still think there’s so much more we can do for day two because even if you deploy the best solution today, six months from now, AWS are releasing ten new services that make it easier to do what you just did. So, someone also needs to build something that shows you the delta to get to the best. And that would involve AWS or somebody thinking cohesively, like, these are how we use our products. And I don’t think there’s a market for it as a third-party company, unfortunately, but I do think that’s where we need to get to, that at day two somebody can give—the way we’re trying to do for Liberty—advice, automated that says, “I see what you’re doing, but it would be better if you did this instead.”

Corey: Yeah, I definitely want to spend more time thinking about these things and analyzing how we wind up addressing them and how we think about them going forward. I learned a lot of these lessons over a decade ago. I was fairly deep into using Puppet, and came to the fair and balanced conclusion that Puppet was a steaming piece of crap. So, the solution was that I was one of the very early developers behind SaltStack, which was going to do everything right. And it was and it was awesome and it was glorious, right up until I saw an environment deployed by someone else who was not as familiar with the tool as I was, at which point I realized hell is other people’s use cases.

And the way that they contextualize these things, you craft a finely balanced torque wrench, it’s a thing of beauty, and people complain about the crappy hammer. “You’re holding it wrong. No, don’t do it that way.” So, I have an awful lot of sympathy for people building platform-level tooling like this, where it works super well for the use case that they’re in, but not necessarily… they’re not necessarily aligned in other ways. It’s a very hard nut to crack.

Matt: Yeah. And like, even as you mentioned earlier, if you take one piece of AWS, for example, API Gateway—and I love the API Gateway team; if you’re listening, don’t hate on me—but there’s, like, 47,000 different ways you can deploy an API Gateway. And the CDK has to cover all of those, it would be a lot easier if there was less ways that you could deploy the thing and then you can start crafting user experiences on a platform. But whenever you start thinking that every AWS component is kind of the same, like think of the amount of ways you’re can deploy a Lambda function now, or think of the, like, containers. I’ll not even go into [laugh] the different ways to run containers.

If you’re building a platform, either you support it all and then it sort of gets quite generic-y, or you’re going to do, like, what serverless cloud are doing though, like Jeremy Daly is building this unique experience that’s like, “Okay, the code is going to build the infrastructure, so just build a website, and we’ll do it all behind it.” And I think they’re really interesting because they’re sort of opposites, in that one doesn’t want to support everything, but should theoretically, for their slice of customers, be awesome, and then the other ones, like, “Well, let’s see what you’re going to do. Let’s have a go at it and I should hopefully support it.”

Corey: I think that there’s so much that can be done on this. But before we wind up calling it an episode, I had one further question that I wanted to explore around the recent results of the community CDK survey that I believe is a quarterly event. And I read the analysis on this, and I talked about it briefly in the newsletter, but it talks about adoption and a few other aspects of it. And one of the big things it looks at is the number of people who are contributing to the CDK in an open-source context. Am I just thinking about this the wrong way when I think that, well, this is a tool that helps me build out cloud infrastructure; me having to contribute code to this thing at all is something of a bug, whereas yeah, I want this thing to work out super well—Docker is open-source, but you’ll never see me contributing things to Docker ever, as a pull request, because it does, as it says on the tin; I don’t have any problems that I’m aware of that, ooh, it should do this instead. I mean, I have opinions on that, but those aren’t pull requests; those are complete, you know, shifts in product strategy, which it turns out is not quite done on GitHub.

Matt: So, it’s funny I, a while ago, was talking to a lad who was the person who came up with the idea for the CDK. And CDK is pretty much the open-source project for AWS if you look at what they have. And the thought behind it, it’s meant to evolve into what people want and need. So yes, there is a product manager in AWS, and there’s a team fully dedicated to building it, but the ultimate aspiration was always it should be bigger than AWS and it should be community-driven. Now personally, I’m not sure—like you just said it—what the incentive is, given that right now CDK only works with CloudFormation, which means that you are directly helping with an AWS tool, but it does give me hope for, like, their CDK for Terraform, and their CDK for Kubernetes, and there’s other flavors based on the same technology as AWS CDK that potentially could have a thriving open-source community because they work across all the clouds. So, it might make more sense for people to jump in there.

Corey: Yeah, I don’t necessarily think that there’s a strong value proposition as it stands today for the idea of the CDK becoming something that works across other cloud providers. I know it technically has the capability, but if I think that Python isn’t quite a first-class experience, I don’t even want to imagine what other providers are going to look like from that particular context.

Matt: Yeah, and that’s from what I understand, I haven’t personally jumped into the CDK for Terraform and we didn’t talk about it here, but in CDK, you get your different levels of construct. And is, like, a CloudFormation-level construct, so everything that’s in there directly maps to a property in CloudFormation, and then L2 is AWS’s opinion on safe defaults, and then L3 is when someone like me comes along and turns it into something that you may find useful. So, it’s a pattern. As far as I know, CDK for Terraform is still on L1. They haven’t got the rich collection—

Corey: And L4 is just hiring you as a consultant—

Matt: [laugh].

Corey: —to come in fix my nonsense for me?

Matt: [laugh]. That’s it. L4 could be Pulumi recently announced that you can use AWS CDK constructs inside it. But I think it’s one of those things where the constructs, if they can move across these different tools the way AWS CDK constructs now work inside Pulumi, and there’s a beta version that works inside CDK for Terraform, then it may or may not make sense for people to contribute to this stuff because we’re not building at a higher level. It’s just the vision is hard for most people to get clear in their head because it needs articulated and told as a clear strategy.

And then, you know, as you said, it is an AWS product strategy, so I’m not sure what you get back by contributing to the project, other than, like, Thorsten—I should say, so Thorsten who wrote the book with me, he is the number three contributor, I think, to the CDK. And that’s just because he is such a big user of it that if he sees something that annoys him, he just comes in and tries to fix it. So, the benefit is, he gets to use the tool. But he is a super user, so I’m not sure, outside of super users, what the use case is.

Corey: I really want to thank you for, I want to say spending as much time talking to me about this stuff as you have, but that doesn’t really go far enough. Because so much of how I think about this invariably winds up linking back to things that you have done and have been advocating for in that community for such a long time. If it’s not you personally, just, like, your fingerprints are all over this thing. So, it’s one of those areas where the entire software developer ecosystem is really built on the shoulders of others who have done a lot of work that came before. Often you don’t get any visibility of who those people are, so it’s interesting whenever I get to talk to someone whose work I have directly built upon that I get to say thank you. So, thank you for this. I really do appreciate how much more straightforward a lot of this is than my previous approach of clicking in the console and then lying about it to provision infrastructure.

Matt: Oh, no worries. Thank you for the thank you. I mean, at the end of the day, all of this stuff is just—it helps me as much as it helps everybody else, and we’re all trying to do make everything quicker for ourselves, at the end of the day.

Corey: If people want to learn more about what you’re up to, where’s the best place to find you these days? They can always take a job at Liberty; I hear good things about it.

Matt: Yeah, we’re always looking for people at Liberty, so come look up our careers. But Twitter is always the best place. So, I’m @NIDeveloper on Twitter. You should find me pretty quickly, or just type Matt Coulter into Google, you’ll get me.

Corey: I like it. It’s always good when it’s like, “Oh, I’m the top Google result for my own name.” On some level, that becomes an interesting thing. Some folks into it super well, John Smith has some challenges, but you know, most people are somewhere in the middle of that.

Matt: I didn’t used to be number one, but there’s a guy called the Kangaroo Kid in Australia, who is, like, a stunt driver, who was number one, and [laugh] I always thought it was funny if people googled and got him and thought it was me. So, it’s not anymore.

Corey: Thank you again for, I guess, all that you do. And of course, taking the time to suffer my slings and arrows as I continue to revise my opinion of the CDK upward.

Matt: No worries. Thank you for having me.

Corey: Matt Coulter, senior architect at Liberty Mutual. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice and leave an angry comment as well that will not actually work because it has to be transpiled through a JavaScript engine first.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Adam

Adam is an independent cloud consultant that helps startups build products on AWS. He’s also the host of AWS FM, a podcast with guests from around the AWS community, and an AWS DevTools Hero.

Adam is passionate about open source and has made a handful of contributions to the AWS CDK over the years. In 2020 he created Ness, an open source CLI tool for deploying web sites and apps to AWS.

Previously, Adam co-founded StatMuse—a Disney backed startup building technology that answers sports questions—and served as CTO for five years. He lives in Nixa, Missouri, with his wife and two children.

Links Referenced:

  • 17 Ways to Run Containers On AWS: https://www.lastweekinaws.com/blog/the-17-ways-to-run-containers-on-aws/
  • Twitter: https://twitter.com/aeduhm
  • Twitch: https://www.twitch.tv/adamelmore

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Every once in a while, I encounter someone in the wild that… well, I’ll just be direct, makes me feel a little bit uneasy, almost like someone’s walking over my grave. And I think I’ve finally figured out elements of what that is. It feels sometimes like I run into people—ideally not while driving—who are trying to occupy sort of the same space in the universe, and I never quite know how to react to that.

Today’s guest is just one such person. Adam Elmore is an independent AWS consultant, has been all over the Twitters for a while, recently started live streaming basically his every waking moment because he is just that interesting. Adam, thank you for suffering my slings and arrows—

Adam: [laugh].

Corey: —and agreeing to chat with me today.

Adam: I would say first of all, you don’t need to be worried about anyone walking over your grave. [laugh]. That was very flattering.

Corey: No, honestly, I have big enterprise companies looking to put me in my grave, but that’s a separate threat model. We’re good on that, for now.

Adam: [laugh]. I got to set myself up here to—I’m just going to laugh a lot, and your editor or somebody’s going to have to deal with that. And maybe the audience will see—[laugh].

Corey: Hey, I prefer that as opposed to talking to people who have absolutely no sense of humor of which they are aware. Awesome, I have a list of companies that they should apply for immediately. So, when I say that we’re trying to occupy elements of the same space in the universe, let me talk a little bit about what I mean by that. You are independent as a consultant, which is how I started this whole nonsense, and then I started gathering a company around me almost accidentally. You are an AWS Dev Tools Hero, whereas I am an AWS community villain, which is kind of a polar opposite slash anti-hero approach, and it’s self-granted in my case. How did you stumble into the universe of AWS? You just realized one day you were too happy and what can you do to make yourself miserable, and this was the answer, or what?

Adam: Yeah, I guess. So. I mean, I’ve been a software developer for 15 years, like, my whole career, that’s kind of what I’ve done. And at some point, I started a startup called StatMuse. And I was able, as sort of a co-founder there, with venture backing, like, I was able to just kind of play with the cloud.

And we deployed everything on AWS, so that was—like, I was there five years; it was sort of five years of running this, I would call it like a Digital Media Studio. Like, we built technology, but we did lots of experiments, so it felt like playing on AWS. Because we built kind of weird one-offs, these digital experiences for various organizations. The Hall of Fame was one of them. We did, like, a, like, a 3-D Talking bust of John Madden, so it was like all kinds of weird technology involved.

But that was sort of five years of, I guess, spending venture money [laugh] to play on AWS. And some of that was Google money; I guess I never thought about that, but Google was an investor in StatMuse. [laugh]. Yeah, so we sort of like—I ran that for five years and was able to learn just a lot of AWS stuff that really excited me. I guess, coming from normal web development stuff, it was exciting just how much leverage you have with AWS, so I sort of dove in pretty hard. And then yeah, when I left StatMuse in 2019 I’ve just been, I guess, going even harder into that direction. I just really enjoy it.

Corey: My first real exposure to AWS was at a company where the CTO was a, I guess we’ll call him an extraordinarily early cloud evangelist. I was there as a contractor, and he was super excited and would tweet nonsensical things like, “I’m never going to rack a server ever again.” And I was a grumpy sysadmin type; I came from the ops world where anything that is new shouldn’t be treated with disdain and suspicion because once you’ve been a sysadmin for 20 minutes, you’ve been there long enough to see today’s shiny new shit become tomorrow’s legacy garbage that you’re stuck supporting. So, “Oh, great. What now?”

I was very down on Cloud in those days and I encountered it with increasing frequency as I stumbled my way through my career. And at the end of 2016, I wound up deciding to go out independent and fix… well, what problems am I good at fixing that I can articulate in a sentence, and well, I’d gotten surprised by AWS bills from time to time—fortunately with someone else’s money; the best kind of mistake to make—and well I know a few things. Let’s get really into it. In time, I came to learn that cost and architecture the same thing in cloud, and now I don’t know how the hell to describe myself. Other people love to describe me, usually with varying forms of profanity, but here we are. It really turns into the idea of forging something of your own path. And you’ve absolutely been doing that for at least the last three years as you become someone who’s increasingly well known and simultaneously harder to describe.

Adam: Yeah, I would say if you figure it out, if you know how to describe me, I would love to know because just coming up with the title—for this episode you needed, like, my title, I don’t know what my title is. I’m also—like, we talked about independent, so nobody sort of gives me a title. I would love to just receive one if you think of one, [laugh] if anyone listening thinks of one… it’s increasingly hard to, sort of like, even decide what I care most about. I know I need to, like, probably niche down, I feel like you’ve kind of niched into the billing stuff. I can’t just be like, “I’m an AWS guy,” because AWS is so big. But yeah, I have no idea.

Corey: Anyone who claims, “Oh, I’m an expert in AWS,” is lying or trying to sell something.

Adam: [laugh]. Exactly.

Corey: I love that. It’s, “Really? I have some questions to establish that for you.” As far as naming what it is, you do, first piece of advice, never ever, ever, ever listen to someone who works at AWS; those people are awful at naming things, as evidenced by basically every service they’ve ever launched. But you are actually fairly close to being an AWS expert. You did a six-week speed-run through every certification that they offer and that is nothing short of astonishing. How’d it come about?

Adam: It’s a unique intersection of skills that I think I have. And I’m not very self-aware, I don’t know all my strengths and weaknesses and I struggle to sort of nail those down, but I think one of my strengths is just ability to, like, consume information, I guess at a high volume. So, I’m like an auditory learner; I can listen to content really fast and sort of retain enough. And then I think the other skill I have is just I’m good at tests. I’ve always said that, like, going back to school, like, high school, I always felt like I was really good at multiple-choice tests. I don’t know if that’s a skill or some kind of innate talent.

But I think those two things combined, and then, like, eight years of building on AWS, and that sort of frames how I was able to take all that on. And I don’t know that I really set out thinking I will do it in six weeks. I took the first few and then did them pretty fast and thought, “I wonder how quickly I could do all of them.” And I just kind of at that point, it became this sort of goal. I have to take on certain challenges occasionally that just sound fun for no reason other than they sound fun and that was kind of the thing for those six weeks. [laugh].

Corey: I have two certifications: Cloud Practitioner and the SysOps Administrator Associate. Those were interesting.

Adam: You took the new one, right? The new SysOps with the labs and stuff I’d love to hear about that.

Corey: I did, back when it was in beta. That was a really interesting experience and I’ll definitely get to that, but I wound up, for example, getting a question wrong in the Cloud Practitioner exam four years ago or so, when it was, “How long does it take to restore an RDS instance from backup?” And I gave the honest answer instead of the by-the-book, correct answer. That’s part of the problem is that I’ve been doing this stuff too long and I know how these things break and what the real world looks like. Certifications are also very much a snapshot at a point in time.

Because I write the Last Week in AWS newsletter, I’m generally up-to-the-minute on what has changed, and things that were not possible yesterday, suddenly are possible today, so I need to know when was this certification launched. Oh, it was in early 2021. Yeah, I needed to be a lot more specific; which week? And then people look at me very strangely and here we are.

The Systems Administrator Certification was interesting because this is the first one, to my knowledge, where they started doing a live lab as a—

Adam: Yeah.

Corey: Component of this. And I don’t think it’s a breach of the NDA to point out that one of the exams was, “Great. Configure CloudWatch out of the box to do this thing that it’s supposed to do out of the box.” And I’ve got to say that making the service do what it’s supposed to do with no caveats is probably the sickest shade I’ve ever seen anyone throw at AWS, like, configuring the service is so bad that it is going to be our test to prove you know what you’re doing. That is amazing.

Adam: [laugh]. Yeah, I don’t have any shade through I’m not as good with the, like, ability to come off, like, witty and kind while still criticizing things. So, I generally just try not to because I’m bad at it. [laugh].

Corey: It’s why I generally advise people don’t try, in seriousness. It’s not that people can’t be clever; it’s that the failure mode of clever is ‘asshole’ and I’m not a big fan of making people feel worse based upon the things that I say and do. It’s occasionally I wind up getting yelled at by Amazonians saying that the people who built a service didn’t feel great about something I said, and my instinctive immediate reaction is, “Oh, shit, that wasn’t my intention. How did I screw this up?” Given a bit of time, I realized that well hang on a minute because I’m not—they’re not my target audience. I’m trying to explain this to other customers.

And, on some level, if you’re going to charge tens of millions of dollars a month for a service or more, maybe make a better one, not for nothing. So, I see both sides of it. I’m not intentionally trying to cause pain, but I’m also not out here insulting people individually. Like, sometimes people make bad decisions, sometimes individually, sometimes in a group. And then we have a service name we have to live with, and all right, I guess I’m going to make fun of that forever. It’s fun that keeps it engaging for me because otherwise, it’s boring.

Adam: No, I hear you. No, and somebody’s got to do it. I’m glad you do it and do it so well because, I mean, you got to keep them honest. Like, that’s the thing. Keep AWS in check.

Corey: Something that I went through somewhat recently was a bit of an awakening. I have no problem revisiting old opinions and discovering that huh, I no longer agree with it; it’s time to evolve that opinion. The CDK specifically was one of those where I looked at it and thought this thing looks a little hokey. So, I started using it in Python and sure enough, the experience was garbage. So cool, the CDK is a piece of crap. There we go. My job is easy.

I was convinced to take a second look at it via TypeScript, a language I do not know and did not have any previous real experience with. So, I spent a few days just powering through it, and now I’m a convert. I think it’s amazing. It is my default go-to for building AWS infrastructure. And all it took was a little bit of poking and prodding to get me to change my mind on that. You’ve taken it to another level and you started actively contributing to the AWS CDK. What was your journey with that, honestly, remarkable piece of software?

Adam: Yeah, so I started contributing to CDK when I was actually doing a lot of Python development. So, I worked with a company that was doing—there was a Python shop. So actually, the first thing I contributed was a Python function construct, which is sort of the equivalent of the Node.js function construct, which like, you can just basically point at a TypeScript file and it transpiles it, bundles it, and does all that, right? So, it makes it easy to deploy TypeScript as a Lambda function.

Well, I mean, it ends up being a JavaScript Lambda function, but anyway, that was the Python function construct. And then I sort of got really into it. So, I got pretty hooked on using the CDK in every place that I could. I’m a huge fan, and I do primarily write in TypeScript these days. I love being able to write TypeScript front-end and back, so built a lot of, like, Next.JS front-ends, and then I’m building back-ends with CDK TypeScript.

Yeah, I’ve had, like, a lot of conversations about CDK. I think there’s definitely a group that’s sort of, against the CDK, if you’re thinking in terms of, like, beginners. And I do see where, for people who aren’t as familiar with AWS, or maybe this is their entry point into cloud development, it does a lot of things that maybe you’re not aware of that, you know, you’re now kind of responsible for. So, it’s deploying—like, it makes it really easy to write, like, three lines of TypeScript that stand up an entire VPC with all this configuration and Managed NAT Gateways and [laugh] everything else. And you may not be aware of all the things you just stood up.

So, CloudFormation maybe is a little more—sort of gives you that better visibility into what you’re creating. So, I’ve definitely seen that pushback. But I think for people who really, like, have built a lot of applications on AWS, I think the CDK is just such a time-saver. I mean, I spend so much less time building the same things in the CDK versus CloudFormation. I’m a big fan.

Corey: For me, I’ve learned enough about JavaScript to be dangerous and it seems like TypeScript is more or less trying to automate a bunch of people’s jobs away, which is basically, from I can tell, their job is to go on the internet and complain about someone’s JavaScript. So great, that that’s really all it does is it complains, “Oh, this ambiguous. You should be more specific about it.” And great. Awesome. I still haven’t gotten into scenarios where I’ve been caught out by typing issues, and very often I find that it just feels like sheer bloodymindedness, but I smile, nod, bend the knee and life goes on.

Adam: [laugh]. When you’ve got a project that’s, like, I don’t know, a few months old—or better, a few years old—and you need to do, like, major refactoring, that’s when TypeScript really saves you just a ton of time. Like, when you can make a change in a type or in actual implementation stuff and then see the ripple effects and then sort of go around the codebase and fix those things, it’s just a lot easier than doing it in JavaScript and discovering stuff at runtime. So, I’m a big TypeScript fan. I don’t know where it’s all headed. I know there’s people that are not fans of, like, transpiling your Lambda functions, for instance. Like, why not just ship good JavaScript? And I get that case, too. Yeah, but I’ve definitely—I felt the productivity boost, I guess—if that’s the thing—from TypeScript.

Corey: For me, I’m still at a point where I’m learning the edges of where things start and where they stop. But one of the big changes I made was that I finally, after 15 years, gave up my beloved Vim as my editor for this and started using VS Code. Because the reasons that I originally went with Vi were understandable when you realize what I was. I’m always going to be remoting into network gear or random—on maintained Unix boxes. Vi is going to be everywhere on everything and that’s fine.

Yeah, I don’t do that anymore, and increasingly, I find that everything I’m writing is local. It is not something that is tied to a remote thing that I need to login and edit by hand. At that point, we are in disaster area. And suddenly it’s nice. I mean things like tab completion, where it just winds up completing the rest of the variable name or, once you enable Copilot and absolutely not CodeWhisperer yet, it winds up you tab complete your entire application. Why not? It’s just outsourcing it to Stack Overflow without that pesky copy and paste step.

Adam: Yeah, I don’t know how in the weeds you want to get on your p—I don’t know, in terms of technical stuff, but Copilot both blows me away—there are days where it autocompletes something that I just, I can’t fathom how—it pulled in not just, like, the patterns that it found, obviously, in training, but, like, the context in the file I’m working and sort of figured out what I was trying to do. Sometimes it blows me away. A lot of times, though, it frustrates me because of TypeScript. Like, I’m used to Typescript and types saving me from typing a lot. Like, I can tab-complete stuff because I have good types defined, right, or it’s just inferred from the libraries I’m using.

It’s tough though when GitHub is fighting with TypeScript and VS Code. But it’s funny that you came from Vim and you now live in VS Code. I really am trying to move from VS Code to, like, the Vim world, mostly because of Twitch streamers that blow my mind with what they can do in Vim [laugh] and how fast they can move. I do—every time I move my hand, like, over to the arrow keys, I feel a little sad and I wish I just did Vim.

Corey:This episode is sponsored in part by our friends at Lambda Cloud. They offer GPU instances with pricing that’s not only scads better than other cloud providers, but is also accessible and transparent. Also, check this out, they get a lot more granular in terms of what’s available. AWS offers NVIDIA A100 GPUs on instances that only come in one size and cost $32/hour. Lambda offers instances that offer those GPUs as single card instances for $1.10/hour. That’s 73% less per GPU. That doesn’t require any long term commitments or predicting what your usage is gonna look like years down the road. So if you need GPUs, check out Lambda. In beta, they’re offering 10TB of free storage and, this is key, data ingress and egress are both free. Check them out at lambdalabs.com/cloud. That's l-a-m-b-d-a-l-a-b-s.com/cloud.

Corey: There are people who have just made it into an entire lifestyle, on some level. And I’m fair to middling; I’ve known people who are dark wizards at it. In practice, I found that my productivity was never constrained by how quickly I can type. It’s one of those things where it’s, I actually want to stop and have my brain catch up sometimes, believe it or not, for those who follow me on Twitter. It’s the idea of wanting to make sure that I am able to intelligently and rationally wrap my head around what it is I’m doing.

And okay, just type out a whole bunch of boilerplate is, like, the least valuable use of anything and that is where I find things like Copilot working super well, where I, if I’m doing CloudFormation, for example, the fact that it tab-completes all the necessary attributes and can go back and change them or whatnot, that’s an enormous time saver. Same story with the CDK, although with some constructs, it doesn’t quite understand which ones get certain values to it. And I really liked the idea behind it. I think this is in some ways, the future of IDEs, to a point.

Adam: Oh, for sure. I think, like, the case, you call that with CloudFormation, you don’t have really typeahead in VS Code, at least I’m not using anything. Maybe there are extensions that give you that in VS Code. But to have Copilot fill in required prompts on a CloudFormation template, that’s a lifesaver. Because I just, every time I write CloudFormation, I’ve just got the docs up and I’m copying stuff I’ve done before or whatever; like, to save that time it’s huge. But CodeWhisperer, not so much? Is it not, I guess, up to snuff? I haven’t seen it or played with it at all.

Corey: It’s still very early days and it hasn’t had exposure outside of Amazonian codebases to my understanding, so it’s, like, “Learn to code like an Amazonian.” And you can fill in your own joke here on that one. I imagine it’s like—isn’t that—aren’t they primarily a Java shop, for one? And all right. It turns out most of my code doesn’t need to operate the way that there’s does.

Adam: I didn’t know that they were training it just internally. Like, I’m assuming Copilot is trained on, like, Stack Overflow or something, right? Or just all of GitHub, I guess.

Corey: And GitHub and a bunch of other things, and people are yelling at them for it, and I haven’t been tracking that. But honestly, the CodeWhisperer announcement taught me things about Copilot, which is weird, which tells me that none of these companies are great at explaining this. Like I can just write a comment in this of, “Add an S3 bucket,” and then Copilot will tab-complete the entirety of adding an S3 bucket, usually even secure, which is awesome. They also fix the early Copilot teething problems of tab-completing people’s AWS API credentials. You know, the—yeah, they’ve fixed a lot of that, thankfully.

Adam: Yeah.

Corey: But it’s still one of those neat things that you can just basically start—it gets a little bit closer to describe what you want the application to do and then it’ll automatically write it for you on the back-end. Sure, sometimes it makes naive decisions that do not bear out, but again, it’s still early days. I’m optimistic.

Adam: Yeah, that reminds me of, like, the, I mean, the serverless cloud, so serverless framework folks, like, what they’re doing where they’re sort of inferring your infrastructure based on you just write an app and it sort of creates the infrastructure as code for you, or just sort of infers it all from your code. So, if you start using a bucket, it’ll create a bucket for that. That definitely seems to be a movement as well, where just do less as a developer [laugh] seems to be the theme.

Corey: Yeah, just move up the stack. We see this time and time again. I mean, look at the—I use this analogy from time to time from the sysadmin world, but in the late-90s, if you wanted to build a web server, you needed a spare week and an intimate knowledge of GCC compiler flags. In time, it became oh, great, now it’s rpm install, then yum install, then ensure present with something like Puppet, and then Docker has it, and now it’s just a checkbox on the S3 page, and you’re running a static site. Things don’t get harder with time, and I don’t think that as a developer, your time is best spent writing by hand the proper syntax for a for loop or whatnot.

It’s not the differentiated value. Talk to me instead about what you want that thing to do. That was my big problem with Lambda when it first came out and I spent two weeks writing my first Lambda function—because I’m bad at programming—where I had to learn the exact format of expected for input and output, and now any Lambda function I write takes me a couple of minutes to write because I’m also bad at programming and don’t know what tests are.

Adam: [laugh]. Tests are overrated, I don’t spend a lot of time writing t—I mean, I do a lot of stuff alone and I do a lot of stuff for myself, so in those contexts, I’m not writing tests if I’m being honest. I stream now and everyone on the stream is constantly asking, “Where are the tests?” Like, there are no tests. I’m sorry. [laugh]. Was someone else’s stream.

Corey: Oh yeah, it used to be though, that you had to be a little sneakier to have other people do work for you. Copilot makes it easier and presumably CodeWhisperer will, too. Used to be that if AWS launched new service and I didn’t know how to configure it, all I would do is restrict a role down to only being able to work with that service, attach that to a user and then just drop the credentials on Twitter or GitHub. And I waited 20 minutes and I came back and sure enough, someone configured it and was already up and mining Bitcoin. So, turn that off, take what they built, and off the production with it. Problem solved. Oh, and rotate those credentials, unless you enjoy pain. Problem solved. The end. And I don’t know if it’s a best practice, but it sure was effective.

Adam: Yeah, that would do it. Well, they’re just like scanners now, right, like they’re just scanning GitHub public repos for any credentials that are leaked like that, and they’re available within seconds. You can literally, like, push a public repo with credentials and it is being [laugh] used within minutes. It’s nuts.

Corey: GitHub has some automatic back channel thing—I believe; I haven’t done an experiment lately, but I believe that AWS will intentionally shoot down the credential as soon as it gets reported, which is kind of amazing. I really should do some more experiments with it just to see how disastrous this can get.

Adam: Yeah. No, I’d be curious. Please let me know. I guess you’ll tweet about it so I’ll see it.

Corey: Can I borrow your account for a few minutes?

Adam: Yeah. [laugh].

Corey: Yeah, it’s fun. Now, the secret to my 17 Ways to Run Containers On AWS is in almost every case, those containers can be crypto miners, so it’s not just about having too many services do the same thing; it’s the attack surface continues to grow and expand in the fullness of time. I’m not saying this is right or wrong; it is what it is, but it’s also something that I think people have an understated appreciation for.

Let’s change topic a little bit. Something you’ve been doing lately and talking about is the idea of building a course on AWS. You’re clearly capable of doing the engineering work. That’s not in question. You’ve been a successful consultant for years, which tells me you also know how to deliver software that meets customer requirements, as opposed to, “Well, the spec was shitty, but I wrote it anyway,” because you don’t last long as a consultant if you enjoy being able to afford to eat if that’s the direction you go in. Now, you’re drifting toward becoming a teacher. Tell me about that. First, what makes you think that’s something you’re good at?

Adam: So, I don’t know. I don’t know that I’m good at it and I guess I’ll find out. I’ve been streaming, like, on Twitch just my work days, and that’s been early signs that I think I’m okay at it, at least. I think it’s very different, obviously, like, a self-paced course are going to be very different from streaming for hours, so there’s a lot more editing and thoughtfulness involved, but I do think, like, I’ve always wanted to teach. So, even before I got into technology—I was pretty late into technology; it was after high school. Like back in high school, I always thought I wanted to be a professor.

I just enjoyed, I guess the idea of presenting ideas in ways that people understood. And I live in an area—so I live in the Ozarks, it’s not a very tech literate area. It became, like, this thing where I felt like I could really explain technology to people who are non-technical. And that’s not necessarily what my course—what I’m aiming to do. I’m trying to teach web developers how to leverage AWS, and then sort of get out of the maybe front-end only or maybe traditional web frameworks—like, they’ve only worked with stuff that they deploy to Heroku or whatever—trying to teach that crowd, how to leverage AWS and all these wonderful primitives that we have.

So, that’s not exactly the same thing, but that’s sort of like, I feel like I do have the ability to translate technology to non-technical folks. And then I guess, like, for me, at this stage of my career, you know, I’ve done a lot of work for a company, for startups, for individual clients, and it feels very, like—I just always feel like I’m going in a hole. Like, I feel like, I’m doing this little thing and I’m serving this one customer, but the idea of being able to, I guess, serve more people and sort of spread my reach, the idea of creating something that I can share with a lot of developers who would maybe benefit from it, it just feels better, I guess. [laugh]. I don’t know exactly all the reasons why that feels better, but like, at the end of the day, my consulting kind of feels like this thing I do because I just need money.

And now that I need money less and less, I just feel like I’d rather do stuff that I actually am excited about. I’m actually really excited about the outcomes for creating a course where, you know, I think I can maybe—my style of teaching or something could resonate with some group of people. Yeah, so that’s it. It’s AWS for web devs. The thought is that I’m going to create courses after this. Like, I hope to move into more education, less consulting. That’s where I’m at.

Corey: I would say you’re probably selling yourself fairly short. I’ve seen a lot of the content you’ve put out over the years and I learned a lot from it every time. I think that there are some folks who put courses out where, one, they don’t have the baseline knowledge around what it is that they’re teaching, it just feels like a grift, and another failure mode is that people know how to do the thing, but they have no idea how to teach it to someone who isn’t them. And there’s nothing inherently wrong with not knowing how to teach; it is its own distinct skill. The problem is when you don’t recognize that about yourself and in turn, wind up having some somewhat significant challenges.

Adam: Yeah. No, I know that one of the struggles is, I work with pretty obscure technologies on AWS. Not obscure, but like, I have a very specific way I build APIs on AWS and I don’t know that’s generally, if you’re taking a bunch of web developers and trying to move them into AWS is probably not the stack that I use. So, that is part of it, but that’s also kind of to my benefit, I guess. It works for me a little bit in that I’m less familiar with maybe the more beginner-friendly way to enter into AWS.

It’s been years, so I think I can kind of come at it a little fresh and that’ll help me produce a course that maybe meets them where they’re at better. Yeah, the grifting thing, I’m definitely sensitive to just this idea of putting out a course. It was hard for me to really go out there and say I was making a course, even on Twitter, because I just feel like there’s, like, some stereotype—I don’t know, there’s an association with that, for me at least, for my perception of course creation. But I know that there are people who’ve done it right and do it for the right reasons. And I think to the extent that I could hit that, you know, both those things, do it right and do it for the right reasons, then it’s exciting to me. And if I can’t, and it turned out not good at teaching, then I’ll move on and do more consulting, I guess, [laugh] or streaming on Twitch.

Corey: You are very clearly self-aware enough that if you put something out and it isn’t effective, I have zero doubt that you won’t just stop selling it, you’ll take it down and reach out to people. Because you, more so than most, seem very cognizant of the fact that a poor experience learning something does not in most people’s cases, translate to, “Oh, my teacher is shitty.” Instead, it’s, “Oh, I’m bad at this and I’m not smart enough to figure it out.” That’s still the problem I run into with bad developer experience on a bunch of things that get launched. If I have a bad time, I assume it’s, “Oh, I’m stupid. I wish someone had told me.”

And first, they did, secondly, it’s the sense that no, it’s just not being very clearly explained and the folks who wrote the documentation or talking about it are too close to what they’ve built to understand what it’s like to look at this thing from fresh eyes. They’re doing a poor job of setting the stage to explain the value it brings and in what scenario, you should be using this.

Adam: It’s a long process. I want to launch the course in the fall, but in the process of building out the course, I’m really going to be doing workshops and individual—like, I just have a lot of friends that are web developers and I’m going to be kind of getting on with them and teaching them this material and just trying to see what resonates. I’m going to a lot of trouble, I guess, to make sure I’m not just putting out a thing just to say I made a course. Like, I don’t actually want to say I made a course, so if I’m going to do it, it’s like most things I do I really kind of throw myself into. And I know if I spend enough energy and effort, I think I can make something that at least helps some people. I guess we’ll see.

Corey: I look forward to it. Any idea as far as rough timeline goes?

Adam: Yeah, I hope to launch in the fall. But if it takes longer, I don’t know. I’ve heard people say, to do a course right, you should spend a year on it. And maybe that’s what I do.

Corey: No, I love that answer. It’s great. You’re just saying I want to launch in the fall, which is sufficiently vague, and if that winds up not being vague enough, you could always qualify with, “Well, I didn’t say what year.”

Adam: [laugh].

Corey: So, great you know, it’s always going to be the fall somewhere.

Adam: [laugh]. I just know, like, when someone says you should spend a year I just do things very hard. Like I really, like, throw a lot of time and obsess, like, I’m very obsessive. And when I do something, it’s hard for me imagine doing any one thing for a year because I burn myself out. Like, I obsess very hard for usually, like, three months, it’s usually, like, a quarter, and then I fall off the face of the earth for three months and I basically mope around the house and I’m just too tired to do anything else. So, I think right now I’m streaming and that’s kind of been my obsession. I’m three weeks in so we got a few more months and then we’ll see, [laugh] we’ll see how I maintain it.

Corey: Well, I look forward to seeing how it comes out. You’ll have to come back and let us know when it’s ready for launch.

Adam: Yeah, that sounds great.

Corey: I really want to thank you for being so generous with your time and taking me through what you’re up to. If people want to learn more, what’s the best place for them to find you?

Adam: Yeah, I think Twitter. I mean, I mostly hang out on Twitter, and these days Twitch. So, Twitter my handle—I guess you’ll put it, like, in the thing description or something. It’s like the phonetic—

Corey: Oh, we will absolutely toss it into the show notes, where useful content goes to linger.

Adam: [laugh]. It’s like A-E-D-U-H-M. It’s like a—it’s the phonetic way of saying Adam, I guess. And then on Twitch, I’m adamelmore. So, those are the two places I spend most my time.

Corey: And off to the show notes it goes. Thank you so much for being so generous with your time. I really appreciate it, Adam.

Adam: Thank you so much for having me, Corey. I really appreciate it.

Corey: Adam Elmore, independent AWS consultant. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an insulting comment that attempts to teach us exactly what we got wrong, but fails utterly because you’re terrible at teaching things.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Alex

Alex holds a Ph.D. in Computer Science and Engineering from UC San Diego, and has spent over a decade building high-performance, robust data management and processing systems. As an early member of a couple fast-growing startups, he’s had the opportunity to wear a lot of different hats, serving at various times as an individual contributor, tech lead, manager, and executive. He also had a brief stint as a Cloud Economist with the Duckbill Group, helping AWS customers save money on their AWS bills. He's currently a freelance data engineering consultant, helping his clients build, manage, and maintain their data infrastructure. He lives in Los Angeles, CA.

Links Referenced:

  • Company website: https://bitsondisk.com
  • Twitter: https://twitter.com/alexras
  • LinkedIn: https://www.linkedin.com/in/alexras/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: I come bearing ill tidings. Developers are responsible for more than ever these days. Not just the code that they write, but also the containers and the cloud infrastructure that their apps run on. Because serverless means it’s still somebody’s problem. And a big part of that responsibility is app security from code to cloud. And that’s where our friend Snyk comes in. Snyk is a frictionless security platform that meets developers where they are - Finding and fixing vulnerabilities right from the CLI, IDEs, Repos, and Pipelines. Snyk integrates seamlessly with AWS offerings like code pipeline, EKS, ECR, and more! As well as things you’re actually likely to be using. Deploy on AWS, secure with Snyk. Learn more at Snyk.co/scream That’s S-N-Y-K.co/scream

Corey: DoorDash had a problem. As their cloud-native environment scaled and developers delivered new features, their monitoring system kept breaking down. In an organization where data is used to make better decisions about technology and about the business, losing observability means the entire company loses their competitive edge. With Chronosphere, DoorDash is no longer losing visibility into their applications suite. The key? Chronosphere is an open-source compatible, scalable, and reliable observability solution that gives the observability lead at DoorDash business, confidence, and peace of mind. Read the full success story at snark.cloud/chronosphere. That's snark.cloud slash C-H-R-O-N-O-S-P-H-E-R-E.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I am joined this week by a returning guest, who… well, it’s a little bit complicated and more than a little bittersweet. Alex Rasmussen was a principal cloud economist here at The Duckbill Group until he committed an unforgivable sin. That’s right. He gave his notice. Alex, thank you for joining me here, and what have you been up to, traitor?

Alex: [laugh]. Thank you for having me back, Corey.

Corey: Of course.

Alex: At time of recording, I am restarting my freelance data engineering business, which was dormant for the sadly brief time that I worked with you all at The Duckbill Group. And yeah, so that’s really what I’ve been up to for the last few days. [laugh].

Corey: I want to be very clear that I am being completely facetious when I say this. When someone is considering, “Well, am I doing what I really want to be doing?” And if the answer is no, too many days in a row, yeah, you should find something that aligns more with what you want to do. And anyone who’s like, “Oh, you’re leaving? Traitor, how could you do that?” Yeah, those people are trash. You don’t want to work with trash.

I feel I should clarify that this is entirely in jest and I could not be happier that you are finding things that are more aligned with aspects of what you want to be doing. I am serious when I say that, as a company, we are poorer for your loss. You have been transformative here across a number of different axes that we will be going into over the course of this episode.

Alex: Well, thank you very much, I really appreciate that. And I came to a point where I realized, you know, the old saying, “You don’t know what you got till it’s gone?” I realized, after about six months of working with Duckbill Group that I missed building stuff, I missed building data systems, I missed being a full-time data person. And I’m really excited to get back to that work, even though I’ll definitely miss working with everybody on the team. So yeah.

Corey: There are a couple of things that I found really notable about your time working with us. One of them was that even when you wound up applying to work here, you were radically different than—well, let’s be direct here—than me. We are almost polar opposites in a whole bunch of ways. I have an eighth-grade education; you have a PhD in computer science and engineering from UCSD. And you are super-deep into the world of data, start to finish, whereas I have spent my entire career on things that are stateless because I am accident prone, and when you accidentally have a problem with the database, you might not have a company anymore, but we can all laugh as we reprovision the web server fleet.

We just went in very different directions as far as what we found interesting throughout our career, more or less. And we were not quite sure how it was going to manifest in the context of cloud economics. And I can say now that we have concluded the experiment, that from my perspective, it went phenomenally well. Because the exact areas that I am weak at are where you excel. And, on some level, I would say that you’re not necessarily as weak in your weak areas as I am in mine, but we want to reinforce it and complementing each other rather than, “Well, we now have a roomful of four people who are all going to yell at you about the exact same thing.” We all went in different directions, which I thought was really neat.

Alex: I did too. And honestly, I learned a tremendous, tremendous amount in my time at Duckbill Group. I think the window into just how complex and just how vast the ecosystem of services within AWS is, and kind of how they all ping off of each other in these very complicated ways was really fascinating, fascinating stuff. But also just an insight into just what it takes to get stuff done when you’re talking with—you know, so most of my clientele to date have been small to medium-sized businesses, you know, small as two people; as big as a few hundred people. But I wasn’t working with Fortune 1000 companies like Duckbill Group regularly does, and an insight into just, number one, what it takes to get things done inside of those organizations, but also what it takes to get things done with AWS when you’re talking about, you know, for instance, contracts that are tens, or hundreds of millions of dollars in total contract value. And just what that involves was just completely eye-opening for me.

Corey: From my perspective, what I found—I guess, in hindsight, it should have been more predictable than it was—but you talk about having a background and an abiding passion for the world of data, and I’m sitting here thinking, that’s great. We have all this data in the form of the Cost and Usage Reports and the bills, and I forgot the old saw that yeah, if it fits in RAM, it’s not a big data problem. And yeah, in most cases, what we have tends to fit in RAM. I guess you don’t tend to find things interesting until Microsoft Excel gives up and calls uncle.

Alex: I don’t necessarily know that that’s true. I think that there are plenty of problems to be had in the it fits in RAM space, precisely because so much of it fits in RAM. And I think that, you know, particularly now that, you know—I think there’s it’s a very different world that we live in from the world that we lived in ten years ago, where ten years ago—

Corey: And right now I’m talking to you on a computer with 128 gigs of RAM, and it—

Alex: Well, yeah.

Corey: —that starts to look kind of big data-y.

Alex: Well, not only that, but I think on the kind of big data side, right? When you had to provision your own Hadoop cluster, and after six months of weeping tears of blood, you managed to get it going, right, at the end of that process, you went, “Okay, I’ve got this big, expensive thing and I need this group of specialists to maintain it all. Now, what the hell do I do?” Right? In the intervening decade, largely due to the just crushing dominance of the public clouds, that problem—I wouldn’t call that problem solved, but for all practical purposes, at all reasonable scales, there’s a solution that you can just plug in a credit card and buy.

And so, now the problem, I think, becomes much more high level, right, than it used to be. Used to be talking about how well you know, how do I make this MapReduce job as efficient as it possibly can be made? Nobody really cares about that anymore. You’ve got a query planner; it executes a query; it’ll probably do better than you can. Now, I think the big challenges are starting to be more in the area of, again, “How do I know what I have? How do I know who’s touched it recently? How do I fix it when it breaks? How do I even organize an organization that can work effectively with data at petabyte scale and say anything meaningful about it?”

And so, you know, I think that the landscape is shifting. One of the reasons why I love this field so much is that the landscape is shifting very rapidly and as soon as we think, “Ah yes. We have solved all of the problems.” Then immediately, there are a hundred new problems to solve.

Corey: For me, what I found, I guess, one of the most eye-opening things about having you here is your actual computer science background. Historically, we have biased for folks who have come up from the ops side of the world. And that lends itself to a certain understanding. And, yes, I’ve worked with developers before; believe it or not, I do understand how folks tend to think in that space. I have not a complete naive fool when it comes to these things.

But what I wasn’t prepared for was the nature of our internal, relatively casual conversations about a bunch of different things, where we’ll be on a Zoom chat or something, and you will just very casually start sharing your screen, fire up a Jupyter Notebook and start writing code as you’re talking to explain what it is you’re talking about and watching it render in real time. And I’m sitting here going, “Huh, I can’t figure out whether we should, like, wind up giving him a raise or try to burn him as a witch.” I could really see it going either way. Because it was magic and transformative from my perspective.

Alex: Well, thank you. I mean, I think that part of what I am very grateful for is that I’ve had an opportunity to spend a considerable period of time in kind of both the academic and industrial spaces. I got a PhD, basically kept going to school until somebody told me that I had to stop, and then spent a lot of time at startups and had to do a lot of different kinds of work just to keep the wheels attached to the bus. And so, you know, when I arrived at Duckbill Group, I kind of looked around and said, “Okay, cool. There’s all the stuff that’s already here. That’s awesome. What can I do to make that better?” And taking my lens so to speak, and applying it to those problems, and trying to figure out, like, “Okay, well as a cloud economist, what do I need to do right now that sucks? And how do I make it not suck?”

Corey: It probably involves a Managed NAT Gateway.

Alex: Whoa, God. And honestly, like, I spent a lot of time developing a bunch of different tools that were really just there in the service of that. Like, take my job, make it easier. And I’m really glad that you liked what you saw there.

Corey: It was interesting watching how we wound up working together on things. Like, there’s a blog post that I believe is out by the time this winds up getting published—but if not, congratulations on listening to this, you get a sneak preview—where I was looking at the intelligent tiering changes in pricing, where any object below 128 kilobytes does not have a monitoring charge attached to it, and above it, it does. And it occurred to me on a baseline gut level that, well wait a minute, it feels like there is some object sizes, where regardless of how long it lives in storage and transition to something cheaper, it will never quite offset that fee. So, instead of having intelligent tiering for everything, that there’s some cut-off point below which you should not enable intelligent tiering because it will always cost you more than it can possibly save you.

And I mentioned that to you and I had to do a lot of articulating with my hands because it’s all gut feelings stuff and this stuff is complicated at the best of times. And your response was, “Huh.” Then it felt like ten minutes later you came back with a multi-page blog post written—again—in a Python notebook that has a dynamic interactive graph that shows the breakeven and cut-off points, a deep dive math showing exactly where in certain scenarios it is. And I believe the final takeaway was somewhere between 148 to 161 kilobytes, somewhere in that range is where you want to draw the cut-off. And I’m just looking at this and marveling, on some level.

Alex: Oh, thanks. To be fair, it took a little bit more than ten minutes. I think it was something where it kind of went through a couple of stages where at first I was like, “Well, I bet I could model that.” And then I’m like, “Well, wait a minute. There’s actually, like—if you can kind of put the compute side of this all the way to the side and just remove all API calls, it’s a closed form thing. Like, you can just—this is math. I can just describe this with math.”

And cue the, like, Beautiful Mind montage where I’m, like, going onto the whiteboard and writing a bunch of stuff down trying to remember the point intercept form of a line from my high school algebra days. And at the end, we had that blog post. And the reason why I kind of dove into that headfirst was just this, I have this fascination for understanding how all this stuff fits together, right? I think so often, what you see is a bunch of little point things, and somebody says, “You should use this at this point, for this reason.” And there’s not a lot in the way of synthesis, relatively speaking, right?

Like, nobody’s telling you what the kind of underlying thing is that makes it so that this thing is better in these circumstances than this other thing is. And without that, it’s a bunch of, kind of, anecdotes and a bunch of kind of finger-in-the-air guesses. And there’s a part of that, that just makes me sad, fundamentally, I guess, that humans built all of this stuff; we should know how all of it fits together. And—

Corey: You would think, wouldn’t you?

Alex: Well, but the thing is, it’s so enormously complicated and it’s been developed over such an enormously long period of time, that—or at least, you know, relatively speaking—it’s really, really hard to kind of get that and extract it out. But I think when you do, it’s very satisfying when you can actually say like, “Oh no, no, we’ve actually done—we’ve done the analysis here. Like, this is exactly what you ought to be doing.” And being able to give that clear answer and backing it up with something substantial is, I think, really valuable from the customer’s point of view, right, because they don’t have to rely on us kind of just doing the finger-in-the-air guess. But also, like, it’s valuable overall. It extends the kind of domain where you don’t have to think about whether or not you’ve got the right answer there. Or at least you don’t have to think about it as much.

Corey: My philosophy has always been that when I have those hunches, they’re useful, and it’s an indication that there’s something to look into here. Where I think it goes completely off the rails is when people, like, “Well, I have a hunch and I have this belief, and I’m not going to evaluate whether or not that belief is still one that is reasonable to hold, or there has been perhaps some new information that it would behoove me to figure out. Nope, I’ve just decided that I know—I have a hunch now and that’s enough and I’ve done learning.” That is where people get into trouble.

And I see aspects of it all the time when talking to clients, for example. People who believe things about their bill that at one point were absolutely true, but now no longer are. And that’s one of those things that, to be clear, I see myself doing this. This is not something—

Alex: Oh, everybody does, yeah.

Corey: —I’m blaming other people for it all. Every once in a while I have to go on a deep dive into our own AWS bill just to reacquaint myself with an understanding of what’s going on over there.

Alex: Right.

Corey: And I will say that one thing that I was firmly convinced was going to happen during your tenure here was that you’re a data person; hiring someone like you is the absolute most expensive thing you can ever do with respect to your AWS bill because hey, you’re into the data space. During your tenure here, you cut the bill in half. And that surprises me significantly. I want to further be clear that did not get replaced by, “Oh, yeah. How do you cut your AWS bill by so much?” “We moved everything to Snowflake.” No, we did not wind up—

Alex: [laugh].

Corey: Just moving the data somewhere else. It’s like, at some level, “Great. How do I cut the AWS bill by a hundred percent? We migrate it to GCP.” Technically correct; not what the customer is asking for.

Alex: Right? Exactly, exactly. I think part of that, too—and this is something that happens in the data part of the space more than anywhere else—it’s easy to succumb to shiny object syndrome, right? “Oh, we need a cloud data warehouse because cloud data warehouse, you know? Snowflake, most expensive IPO in the history of time. We got to get on that train.”

And, you know, I think one of the things that I know you and I talked about was, you know, where should all this data that we’re amassing go? And what should we be optimizing for? And I think one of the things that, you know, the kind of conclusions that we came to there was, well, we’re doing some stuff here, that’s kind of designed to accelerate queries that don’t really need to be accelerated all that much, right? The difference between a query taking 500 milliseconds and 15 seconds, from our point of view, doesn’t really matter all that much, right? And that realization alone, kind of collapsed a lot of technical complexity, and that, I will say we at Duckbill Group still espouse, right, is that cloud cost is an architectural problem, it’s not a right-sizing your instances problem. And once we kind of got past that architectural problem, then the cost just sort of cratered. And honestly, that was a great feeling, to see the estimate in the billing console go down 47% from last month, and it’s like, “Ah, still got it.” [laugh].

Corey: It’s neat to watch that happen, first off—

Alex: For sure.

Corey: But it also happened as well, with increasing amounts of utility. There was a new AWS billing page that came out, and I’m sure it meets someone’s needs somewhere, somehow, but the things that I always wanted to look at when I want someone to pull up their last month’s bill is great, hit the print button—on the old page—and it spits out an exploded pdf of every type of usage across their entire AWS estate. And I can skim through that thing and figure out what the hell’s going on at a high level. And this new thing did not let me do that. And that’s a concern, not just for the consulting story because with our clients, we have better access than printing a PDF and reading it by hand, but even talking to randos on the internet who were freaking out about an AWS bill, they shouldn’t have to trust me enough to give me access into their account. They should be able to get a PDF and send it to me.

Well, I was talking with you about this, and again, in what felt like ten minutes, you wound up with a command line tool, run it on an exported CSV of a monthly bill and it spits it out as an HTML page that automatically collapses in and allocates things based upon different groups and service type and usage. And congratulations, you spent ten minutes to create a better billing experience than AWS did. Which feels like it was probably, in fairness to AWS, about seven-and-a-half minutes more time than they spent on it.

Alex: Well, I mean, I think that comes back to what we were saying about, you know, not all the interesting problems in data are in data that doesn’t fit in RAM, right? I think, in this case, that came from two places. I looked at those PDFs for a number of clients, and there were a few things that just made my brain hurt. And you and Mike and the rest of the folks at Duckbill could stare at the PDF, like, reading the matrix because you’ve seen so many of them before and go, ah, yes, “Bill spikes here, here, here.” I’m looking at this and it’s just a giant grid of numbers.

And what I wanted was I wanted to be able to say, like, don’t show me the services in alphabetical order; show me the service is organized in descending order by spend. And within that, don’t show me the operations in alphabetical order; show me the operations in decreasing order by spend. And while you’re at it, group them into a usage type group so that I know what usage type group is the biggest hitter, right? The second reason, frankly, was I had just learned that DuckDB was a thing that existed, and—

Corey: Based on the name alone, I was interested.

Alex: Oh, it was an incredible stroke of luck that it was named that. And I went, “This thing lets me run SQL queries against CSV files. I bet I can write something really fast that does this without having to bash my head against the syntactic wall that is Pandas.” And at the end of the day, we had something that I was pretty pleased with. But it’s one of those examples of, like, again, just orienting the problem toward, “Well, this is awful.”

Because I remember when we first heard about the new billing experience, you kind of had pinged me and went, “We might need something to fix this because this is a problem.” And I went, “Oh, yeah, I can build that.” Which is kind of how a lot of what I’ve done over the last 15 years has been. It’s like, “Oh. Yeah, I bet I could build that.” So, that’s kind of how that went.

Corey: This episode is sponsored in part by our friend EnterpriseDB. EnterpriseDB has been powering enterprise applications with PostgreSQL for 15 years. And now EnterpriseDB has you covered wherever you deploy PostgreSQL on-premises, private cloud, and they just announced a fully-managed service on AWS and Azure called BigAnimal, all one word. Don’t leave managing your database to your cloud vendor because they’re too busy launching another half-dozen managed databases to focus on any one of them that they didn’t build themselves. Instead, work with the experts over at EnterpriseDB. They can save you time and money, they can even help you migrate legacy applications—including Oracle—to the cloud. To learn more, try BigAnimal for free. Go to biganimal.com/snark, and tell them Corey sent you.

Corey: The problem that I keep seeing with all this stuff is I think of it in terms of having to work with the tools I’m given. And yeah, I can spin up infrastructure super easily, but the idea of, I’m going to build something that manipulates data and recombines it in a bunch of different ways, that’s not something that I have a lot of experience with, so it’s not my instinctive, “Oh, I bet there’s an easier way to spit this thing out.” And you think in that mode. You effectively wind up automatically just doing those things, almost casually. Which does make a fair bit of sense, when you understand the context behind it, but for those of us who don’t live in that space, it’s magic.

Alex: I’ve worked in infrastructure in one form or another my entire career, data infrastructure mostly. And one of the things—I heard this from someone and I can’t remember who it was, but they said, “When infrastructure works, it’s invisible.” When you walk in the room and flip the light switch, the lights come on. And the fact that the lights come on is a minor miracle. I mean, the electrical grid is one of the most sophisticated, globally-distributed engineering systems ever devised, but we don’t think about it that way, right?

And the flip side of that, unfortunately, is that people really pay attention to infrastructure most when it breaks. But they are two edges of the same proverbial sword. It’s like, I know, when I’ve done a good job, if the thing got built and it stayed built and it silently runs in the background and people forget it exists. That’s how I know that I’ve done a good job. And that’s what I aim to do really, everywhere, including with Duckbill Group, and I’m hoping that the stuff that I built hasn’t caught on fire quite yet.

Corey: The smoke is just the arising of the piles of money it wound up spinning up.

Alex: [laugh].

Corey: It’s like, “Oh yeah, turns out that maybe we shouldn’t have built a database out of pure Managed NAT Gateways. Yeah, who knew?”

Alex: Right, right. Maybe I shouldn’t have filled my S3 bucket with pure unobtainium. That was a bad idea.

Corey: One other thing that we do here that I admit I don’t talk about very often because people get the wrong idea, but we do analyst projects for vendors from time to time. And the reason I don’t say that is, when people hear about analysts, they think about something radically different, and I do not self-identify as an analyst. It’s, “Oh, I’m not an analyst.” “Really? Because we have analyst budget.” “Oh, you said analyst. I thought you said something completely different. Yes, insert coin to continue.”

And that was fine, but unlike the vast majority of analysts out there, we don’t form our opinions based upon talking to clients and doing deeper dive explorations as our primary focus. We’re a team of engineers. All right, you have a product. Let’s instrument something with it, or use your product for something and we’ll see how it goes along the way. And that is something that’s hard for folks to contextualize.

What was really fun was bringing you into a few of those engagements just because it was interesting; at the start of those calls. “It was all great, Corey is here and—oh, someone else’s here. Is this a security problem?” “It’s no, no, Alex is with me.” And you start off those calls doing what everyone should do on those calls is, “How can we help?” And then we shut up and listen. Step one, be a good consultant.

And then you ask some probing questions and it goes a little bit deeper and a little bit deeper, and by the end of that call, it’s like, “Wow, Alex is amazing. I don’t know what that Corey clown is doing here, but yeah, having Alex was amazing.” And every single time, it was phenomenal to watch as you, more or less, got right to the heart of their generally data-oriented problems. It was really fun to be able to think about what customers are trying to achieve through the lens that you see the world through.

Alex: Well, that’s very flattering, first of all. Thank you. I had a lot of fun on those engagements, honestly because it’s really interesting to talk to folks who are building these systems that are targeting mass audiences of very deep-pocketed organizations, right? Because a lot of those organizations, the companies doing the building are themselves massive. And they can talk to their customers, but it’s not quite the same as it would be if you or I were talking to the customers because, you know, you don’t want to tell someone that their baby is ugly.

And note, now, to be fair, we under no circumstances were telling people that their baby was ugly, but I think that the thing that is really fun for me is to kind of be able to wear the academic database nerd hat and the practitioner hat simultaneously, and say, like, “I see why you think this thing is really impressive because of this whiz-bang, technical thing that it does, but I don’t know that your customers actually care about that. But what they do care about is this other thing that you’ve done as an ancillary side effect that actually turns out is a much more compelling thing for someone who has to deal with this stuff every day. So like, you should probably be focusing attention on that.” And the thing that I think was really gratifying was when you know that you’re meeting someone on their level and you’re giving them honest feedback and you’re not just telling them, you know, “The Gartner Magic Quadrant says that in order to move up and to the right, you must do the following five features.” But instead saying, like, “I’ve built these things before, I’ve deployed them before, I’ve managed them before. Here’s what sucks that you’re solving.” And seeing the kind of gears turn in their head is a very gratifying thing for me.

Corey: My favorite part of consulting—and I consider analyst style engagements to be a form of consulting as well—is watching someone get it, watching that light go on, and they suddenly see the answer to a problem that’s been vexing them I love that.

Alex: Absolutely. I mean, especially when you can tell that this is a thing that has been keeping them up at night and you can say, “Okay. I see your problem. I think I understand it. I think I might know how to help you solve it. Let’s go solve it together. I think I have a way out.”

And you know, that relief, the sense of like, “Oh, thank God somebody knows what they’re doing and can help me with this, and I don’t have to think about this anymore.” That’s the most gratifying part of the job, in my opinion.

Corey: For me, it has always been twofold. One, you’ve got people figuring out how to solve their problem and you’ve made their situation better for it. But selfishly, the thing I like the most personally has been the thrill you get from solving a puzzle that you’ve been toying with and finally it clicks. That is the endorphin hit that keeps me going.

Alex: Absolutely.

Corey: And I didn’t expect when I started this place is that every client engagement is different enough that it isn’t boring. It’s not the same thing 15 times. Which it would be if it were, “Hi, thanks for having us. You haven’t bought some RIs. You should buy some RIs. And I’m off.” It… yeah, software can do that. That’s not interesting.

Alex: Right. Right. But I think that’s the other thing about both cloud economics and data engineering, they kind of both fit into that same mold. You know, what is it? “All happy families are alike, but each unhappy family is unhappy in its own way.” I’m butchering Chekhov, I’m sure. But like—if it’s even Chekhov.

But the general kind of shape of it is this: everybody’s infrastructure is different. Everybody’s organization is different. Everybody’s optimizing for a different point in the space. And being able to come in and say, “I know that you could just buy a thing that tells you to buy some RIs, but it’s not going to know who you are; it’s not going to know what your business is; it’s not going to know what your challenges are; it’s not going to know what your roadmap is. Tell me all those things and then I’ll tell you what you shouldn’t pay attention to and what you should.”

And that’s incredibly, incredibly valuable. It’s why, you know, it’s why they pay us. And that’s something that you can never really automate away. I mean, you hear this in data all the time, right? “Oh, well, once all the infrastructure is managed, then we won’t need data infrastructure people anymore.”

Well, it turns out all the infrastructure is managed now, and we need them more than we ever did. And it’s not because this managed stuff is harder to run; it’s that the capabilities have increased to the point that they’re getting used more. And the more that they’re getting used, the more complicated that use becomes, and the more you need somebody who can think at the level of what does the business need, but also, what the heck is this thing doing when I hit the run key? You know? And that I think, is something, particularly in AWS where I mean, my God, the amount and variety and complexity of stuff that can be deployed in service of an organization’s use case is—it can’t be contained in a single brain.

And being able to make sense of that, being able to untangle that and figure out, as you say, the kind of the aha moment, the, “Oh, we can take all of this and just reduce it down to nothing,” is hugely, hugely gratifying and valuable to the customer, I’d like to think.

Corey: I think you’re right. And again, having been doing this in varying capacities for over five years—almost six now; my God—the one thing has been constant throughout all of that is, our number one source for new business has always been word of mouth. And there have been things that obviously contribute to that, and there are other vectors we have as well, but by and large, when someone winds up asking a colleague or a friend or an acquaintance about the problem of their AWS bill, and the response almost universally, is, “Yeah, you should go talk to The Duckbill Group,” that says something that validates that we aren’t going too far wrong with what we’re approaching. Now that you’re back on the freelance data side, I’m looking forward to continuing to work with you, if through no other means and being your customer, just because you solve very interesting and occasionally very specific problems that we periodically see. There’s no reason that we can’t bring specialists in—and we do from time to time—to look at very specific aspects of a customer problem or a customer constraint, or, in your case for example, a customer data set, which, “Hmm, I have some thoughts on here, but just optimizing what storage class that three petabytes of data lives within seems like it’s maybe step two, after figuring what the heck is in it.” Baseline stuff. You know, the place that you live in that I hand-wave over because I’m scared of the complexity.

Alex: I am very much looking forward to continuing to work with you on this. There’s a whole bunch of really, really exciting opportunities there. And in terms of word of mouth, right, same here. Most of my inbound clientele came to me through word of mouth, especially in the first couple years. And I feel like that’s how you know that you’re doing it right.

If someone hires you, that’s one thing, and if someone refers you, to their friends, that’s validation that they feel comfortable enough with you and with the work that you can do that they’re not going to—you know, they’re not going to pass their friends off to someone who’s a chump, right? And that makes me feel good. Every time I go, “Oh, I heard from such and such that you’re good at this. You want to help me with this?” Like, “Yes, absolutely.”

Corey: I’ve really appreciated the opportunity to work with you and I’m super glad I got the chance to get to know you, including as a person, not just as the person who knows the data, but there’s a human being there, too, believe it or not.

Alex: Weird. [laugh].

Corey: And that’s the important part. If people want to learn more about what you’re up to, how you think about these things, potentially have you looked at a gnarly data problem they’ve got, where’s the best place to find you now?

Alex: So, my business is called Bits on Disk. The website is bitsondisk.com. I do write occasionally there. I’m also on Twitter at @alexras. That’s Alex-R-A-S, and I’m on LinkedIn as well. So, if your lovely listeners would like to reach me through any of those means, please don’t hesitate to reach out. I would love to talk to them more about the challenges that they’re facing in data and how I might be able to help them solve them.

Corey: Wonderful. And we will of course, put links to that in the show notes. Thank you again for taking the time to speak with me, spending as much time working here as you did, and honestly, for a lot of the things that you’ve taught me along the way.

Alex: My absolute pleasure. Thank you very much for having me.

Corey: Alex Rasmussen, data engineering consultant at Bits on Disk. I’m Cloud Economist Corey Quinn. This is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry comment that is so large it no longer fits in RAM.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Steren

Steren is a Group Product Manager at Google Cloud. He is part of the serverless team, leading Cloud Run. He is also working on sustainability, leading the Google Cloud Carbon Footprint product.

Steren is an engineer from École Centrale (France). Before joining Google, he was CTO of a startup building connected objects and multi device solutions.

Links Referenced:

  • previous episode: https://www.lastweekinaws.com/podcast/screaming-in-the-cloud/google-cloud-run-satisfaction-and-scalability-with-steren-giannini/
  • Google Cloud Region Picker: https://cloud.withgoogle.com/region-picker/
  • Google Cloud regions: https://cloud.google.com/sustainability/region-carbon

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: DoorDash had a problem. As their cloud-native environment scaled and developers delivered new features, their monitoring system kept breaking down. In an organization where data is used to make better decisions about technology and about the business, losing observability means the entire company loses their competitive edge. With Chronosphere, DoorDash is no longer losing visibility into their applications suite. The key? Chronosphere is an open-source compatible, scalable, and reliable observability solution that gives the observability lead at DoorDash business, confidence, and peace of mind. Read the full success story at snark.cloud/chronosphere. That's snark.cloud slash C-H-R-O-N-O-S-P-H-E-R-E.

Corey: This episode is sponsored in part by our friend EnterpriseDB. EnterpriseDB has been powering enterprise applications with PostgreSQL for 15 years. And now EnterpriseDB has you covered wherever you deploy PostgreSQL on-premises, private cloud, and they just announced a fully-managed service on AWS and Azure called BigAnimal, all one word. Don’t leave managing your database to your cloud vendor because they’re too busy launching another half-dozen managed databases to focus on any one of them that they didn’t build themselves. Instead, work with the experts over at EnterpriseDB. They can save you time and money, they can even help you migrate legacy applications—including Oracle—to the cloud. To learn more, try BigAnimal for free. Go to biganimal.com/snark, and tell them Corey sent you.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. My guest today was recently on the show. Steren Giannini is the product lead for Google Cloud Run, and we talked about that in a previous episode. If you haven’t listened to it, you might wish to go back and listen to it, but it’s not a prerequisite for what we’re about to talk about today. Because apparently Google still does it’s 20% time, and one of the things that Steren decided to do—because, you know, everyone needs a hobby—you decided to go ahead and start the Google Cloud Carbon Footprint, which is—well, Steren, thanks for coming back. What the hell is that?

Steren: Thanks for having me back on the show. So yes, we started with Cloud Carbon Footprint, and this is a product that now has launched publicly, available to every Google Cloud customer right out of the box of the Google Cloud Console.

Corey: I should also point out, because people always wonder and it’s the first thing I always check, yes, this one is free. I’m trying to imagine a scenario which you charge for this and I wasn’t incensed by it, and I can’t. So, good work, you aren’t charging anything for it. Good job. Please continue.

Steren: So, Google Cloud Carbon Footprint helps a Google Cloud customer understand and reduce their gross carbon emissions linked to their Google cloud usage. So yeah, what do we mean by carbon emission? Just so that we are all on the same page, these are the greenhouse gases that are emitted due to the activity of using Google Cloud that are notably responsible for climate change. And we report them in equivalent of carbon dioxide—CO2—and you know, the shortcut is just to say ‘carbon.’

Corey: Now, I’m going to start with something relatively controversial. It’s an opinion I have around this sort of thing. And I should also disclaim, I am not in any way, shape, or form, disputing the climate change as caused by humans is real. It is. If you don’t believe that, please go listen to something else, maybe Infowars. I don’t know and I don’t care. I just don’t want you around.

Now, the problem that I have with this is, on some level, it feels like a cloud provider talking to its customers about their carbon footprint is, on some level, shifting the onus of responsibilities in some way away from the cloud provider and onto the customer. Now, I freely admit that this is a nuanced topic, but how do you view this?

Steren: What I mentioned is that we are exposing to customer their gross carbon emissions, but what about their net carbon emissions? Well, Google Cloud customers, net operational carbon emissions are simply zero. Why? Because if you open Google’s environmental report, you will see that Google is purchasing as much renewable energy globally for the year as it is using. So, that means that on a yearly basis worldwide, every kilowatt hour of electricity has been matched with renewable energy.

And you know, this Google has been doing since 2017. Since 2007, Google was already matching its carbon footprint with carbon offsets. But 2017, Google went beyond and is matching the purchase of the electricity with renewable energy. So, in a sense, your net operational emissions are zero.

Now, that’s not sufficient for our customers. They have some reporting obligations; they need to know before this renewable matching, what were their gross emissions? And they also need to know what are their emissions coming from, not only the electricity usage, but maybe the data center or manufacturing. And this is all of what we expose in Google Cloud Carbon Footprint. They are before offset, before renewable energy matching.

And you’re right also to say that this is not only the customer’s problem, and indeed, Google itself has set a goal to get to a hundred percent carbon-free electricity for every hour in every location. The big goal for 2030 is that at every hour, every location, the electricity comes from carbon-free sources. This is very ambitious and never done before, of course, at the scale of Google, but this is the next goal for Google.

Corey: The challenge that I have—in the abstract—with cloud providers, more or less, shaming customers—not to say that’s what you’re doing here—about their carbon usage and their carbon footprint is, okay, I appreciate that this is everyone’s problem, and yes, it is something that we should be focusing on extensively. The counterargument is that I don’t recall ever getting a meeting invite to a Google or Amazon or Microsoft or Oracle negotiation with any of your power bills or power companies or power sourcing. I have no input whatsoever as a customer on those things. And, on some level, it’s “Ooh, you’re causing a particular amount of carbon to be used by your usage of these services.” Like, well, at some level, it feels like that is more of a you thing than a me thing.

And I want to be clear, I’m speaking more in the abstract to the industry rather than the specifics of Google Cloud, not to unfairly put you in the position of having to speak for everyone.

Steren: No, but you’re right. If you were to do nothing, Google is constantly working hard to sign more power purchase agreements with some renewable energy sources or optimizing its data centers. Google Cloud data centers are one of the most optimized data centers in the industry with a power usage effectiveness of 1.1, which is basically saying that the energy that is used to power the facility over the energy used to actually power the server is 1.1. So, not that much loss in between.

So, all of that to say, Google Cloud and Google are working very hard anyway to reduce Google Cloud’s carbon footprints and the carbon footprint of Google Cloud customers. So, if you were to do nothing, the charts that you’re seeing on Google Cloud Carbon Footprint should trend to zero. But in the meantime, you know, that’s not the case, so that’s why we show the data. And, like, many customers want to know or have the obligation to report on this data.

Corey: One of the challenges that I see—and I believe this might even be related to the carbon footprint tool you have built out on top of Google Cloud—is when I am looking at… at where to place something—first, let me just say the region experience is wildly different than I’m used to with AWS. In the AWS universe, every region is basically its own island; it is very difficult to get a holistic view between regions. Google Cloud does not have that approach. There are advantages and disadvantages to both. I’m not passing any particular value judgment—for once—on this topic in this context. But where do I want to spin something up? And I have a dropdown of the regions that I can put it in. And some of these now have a green leaf next to them and others do not. I’m going to go out on a limb and assume you had a hand in this somewhere.

Steren: Exactly. That’s something I worked on with the team. So, you will see on the Google Cloud Console location selectors on the Google Cloud location page, on the Google Cloud documentation, you will see a small low CO2 indicator next to some regions. And this indicator is basically saying that this region meets some criteria of high carbon-free energy percentage or low grid carbon intensity. So, you don’t need to go into the details; you just need to know that if you see this small leaf, that means that for a given workload, the emissions in that particular region will be way lower than on another region which doesn’t have the leaf.

Often at Google, when we do a change we A/B test it. We A/B tested those small low CO2 indicators because, you know, that’s a console-wide change so we want to make sure that it’s worth is. And well, it turns out that for people who were in the experiment—so people will be seeing the leaf—among new Google Cloud users, they were 50% more likely to pick a low-carbon region when the leaf was displayed. And among all users, it was 19%. So, you see how just by surfacing the information, we were able to significantly influence customers’ behavior towards reducing their carbon emissions.

And, you know, if you ask me, I think picking the cleanest region is probably one of the simplest action you can take—if possible, of course—to reduce your gross carbon emissions because, you know, they don’t require to change your architecture or your infrastructure; it just requires you to make the right choice in the first place. And just by letting people know that some regions are emitting much less carbon than others we basically allow them to reduce their footprint.

Corey: A question I have is that as you continue to move up the stack, one of the things that Google has done extraordinarily well is the global network. And we talked previously about how I run the snark.cloud URL shortener in Google Cloud. That is homed out of us-central1 as far as regions go. But given that thing is effectively stateless, it just talks to Google Sheets for its source of truth, but then just runs a Docker invocation on every request, cool, I can see a scenario in which that becomes much more of a global service.

In other words, if you can run that in pops in every region around the world on some level, there is no downside, from my perspective, on doing that. What I’m wondering then, as a result of that, is as you start seeing the evolution of services becoming more and more global, instead of highly region-specific, does that change the way that we should be thinking potentially about carbon footprint and regional selection? Or is that too much of a niche edge case to really be top of radar right now?

Steren: Oh, there are many things to talk about here. The first one is that you might be hinting at something that Google is already doing, which is location shifting of workloads in order to optimize power usage, and, you know of course, carbon emissions. So, Google itself is already doing that. For example, I guess, to process YouTube videos, that can be done, not necessarily right away and that can be done in the location in which, for example, the sun is shining. So, there are some very interesting things that can be done if you allow the workloads to be run in not necessarily a specific region.

Now, that being said, I think there are many other things that people consider when they pick a region. First, well, maybe they have some data locality constraints, right? This is very much the case in European countries where the data must stay in a given region, by law. Second, well, maybe they care about the price. And as you probably know, [laugh] the price of cloud providers is not the same in every region.

Corey: I’ve noticed that and in fact, I was going to get into that as our next transition, but you’ve just opened Pandora’s Box early. It’s great to have the carbon-friendly indicator next to the region, but I also want number of dollar signs next to it as well. Like in AWS-land, do you have the tier one regions where everything is the lowest price: us-east-1, us-west-2, and a few others escaped me from time to time, where Managed NAT Gateways are really expensive. And then you go under some others and they get even more expensive, somehow. Like, talk about pushing the bounds of cloud economics. It’s astonishing to me.

Steren: Yes. And so—

Corey: Because I want that display, on some level—

Steren: Exactly.

Corey: —as a customer, in many cases.

Steren: So, there is price, there is carbon, but of course, you know, if you are serving web requests, there is probably also latency that you care about, right? Even if—for example, Finland is very low carbon. You might not host your workloads in Finland if you want to serve US customers. So, in a sense, there are many dimensions to optimize when you pick a region. And I just sent you a link to something that I built, which is called Google Cloud Region Picker.

It’s basically a tool with three sliders. First one is carbon footprint; you tell us how much you care about that. Hopefully, you put it to the right. The second one is lower price. So, how much do you want the tool to optimize to lower your bill? And third one is latency, and then you tell us where your users are coming from and if you care about latency.

Because some workloads are not subject to latency requirements. Like, if you do batch jobs, well, that doesn’t serve a user request, so that can be done asynchronously at a later time or in a different place. And what this tool does is that it takes your inputs and it basically tells you which Google Cloud region is the best fit for you. And if you use it, you will see it has very small symbols like three dollars for the most expensive regions, one dollar for the least expensive ones, three leaves for the greenest regions, and zero leaves for the non-green one.

Corey: This is awesome. I’m a little bit disappointed that I hadn’t seen this before. This is a thing of beauty.

Steren: Yeah. Again, done by me as a 20%. [laugh]. And, you know, the goal is to educate, like, of course, it’s way more complex. Like, you know that price optimization is way more complex than a slider, but the goal of this tool is to educate and to give a first result. Like, okay, if you are in France and care about carbon, then go here. If you are in Oregon, go here. And so, many parameters that this tool help you optimize in a very simple way.

Corey: One of the challenges I think I get into when I look at this across the board, is that you have a couple of very different ends on a confusing spectrum, by which I mean that one of the things I would care about from a region picker, for example, is there sufficient capacity in that region for the things I want to run. At my scale of things where right now on Google Cloud I run a persistent VM that hangs out all the time, and I run some Google Cloud Run stuff. Great. If you have capacity problems with either one of those, are you really a cloud?

But then we have other folks who are spinning up tens or hundreds of thousands of a very particular instance type in a very specific region. That’s the sort of thing that requires a bit more in the way of capacity planning and the rest. So, I have to imagine for those types of use cases, this tool is insufficient. The obvious reason, of course, if you’re spinning up that much of anything, for God’s sake, reach out and talk to your account manager before trying to do it willy-nilly but yes.

Steren: That’s exactly right. So, as I said, this tool is simplified tool to give, like, the vast majority of users a sense of where to put their workloads. But of course, if you’re a very big enterprise customer who is going to sign a very big deal with Google Cloud, talk to your account manager because if you do need a lot of capacity, Google Cloud might need to plan for it. And not every regions have the same capacity and we are always working with our customers to make sure we direct them in the right place and have enough capacity. A real-life example of a very high profile Google Cloud customer was that they were selecting a region without knowing its carbon impact, and when we started to disclose the carbon characteristics of Google Cloud regions—which is another link we can send to the audience—this customer realized that the region they selected—you know, maybe because it was close to their user base—was really not the most carbon friendly.

So, they decided to switch to another one. And if we take an example, if you take Las Vegas, it has a carbon-free energy percentage of 20%. So, that basically means that on average, 20% of the time, the electricity comes from carbon-free sources. If you are to move to Oregon, this same workload, Oregon has a carbon-free energy percentage of 90%. So, you can see how just by staying on the West Coast, moving from Las Vegas to Oregon, you have drastically reduced your carbon emissions. And your bill, by the way because it turns out Oregon is one of the cheapest Google Cloud Data Center. So, you see how just being aware of those numbers led some very important customers who care about sustainability to make some fundamental choices when it comes to the regions they select.

Corey: I guess that leads to my big obvious question, where I wind up pulling up my own footprint in Google Cloud—again, I don’t run much there—and apparently over the last year, I’ve had something on the order of two kilograms of carbon. Great. It feels like for this scale, I almost certainly burn more carbon than that just searching Google for carbon-related questions about where to place things. So, for my use case, this is entirely academic. You can almost run my workloads based upon, I don’t know, burning baby seals or something, and the ecological footprint does not materially change.

Then we go to the other extreme end of the spectrum with the hundreds of thousands of instances, where this stuff absolutely makes a significant and massive difference. My question is, when should people begin thinking about the carbon footprint of their cloud workload at what point of scale?

Steren: So, as you said, a good order of magnitude is one transatlantic flight is a thousand kilogram of equivalent CO2. So, you see how just by flying once, you’re already, like, completely overshadowing your Google Cloud carbon footprint. But that’s because you are not using a lot of Google Cloud resources. Overall, you know, I think your question is basically the same as when should individuals try to optimize reducing their carbon footprint? And here I always recommend there are tons of things you can optimize.

Start by the most impactful ones. And impactful means an action will have a lot of impact in reducing the footprint, but also the footprint reduction will be significant by itself. And two kilograms of CO2, yes indeed, it is very low, but if you start reaching out into the thousands of kilograms of CO2 that starts to represent, like, one flight, for example. So, you should definitely care about it. And as I said, some actions might be rather easy, like picking the right region might be something you can do pretty easily for your business and then you will see your carbon emissions being divided by, you know, sometimes five.

This episode is sponsored in part by our friends at Lambda Cloud. They offer GPU instances with pricing that’s not only scads better than other cloud providers, but is also accessible and transparent. Also, check this out, they get a lot more granular in terms of what’s available. AWS offers NVIDIA A100 GPUs on instances that only come in one size and cost $32/hour. Lambda offers instances that offer those GPUs as single card instances for $1.10/hour. That’s 73% less per GPU. That doesn’t require any long term commitments or predicting what your usage is gonna look like years down the road. So if you need GPUs, check out Lambda. In beta, they’re offering 10TB of free storage and, this is key, data ingress and egress are both free. Check them out at lambdalabs.com/cloud. That's l-a-m-b-d-a-l-a-b-s.com/cloud.

Corey: I want to challenge your assertion, incidentally. You say that I’m not using a whole lot of Google Cloud resources. I disagree. I use roughly a dozen different Google Cloud resources tied together for some of these things, but they’re built on serverless design patterns, which means that they scale to nothing. I’m not sitting there with an idle VM—except that one—that is existing on a persistent basis.

For example, I look at the things that show up on the top five list. Compute Engine is number one, Cloud Run, Cloud Logging, Cloud Storage, and App Engine are the rest that are currently being used. I think there’s a significant untold story around the idea of building in a serverless way for climate purposes.

Steren: Yes. So, maybe for those who are not aware of what you are seeing on the dashboard, so when you open this Google Cloud Carbon Footprint tool on the Cloud Console, you saw a breakdown of your yearly carbon footprint and monthly carbon footprint across a few dimensions. The first one is the regions because as we said, this matters a lot; like, the regions have a lot of impact. The second one are the month; of course, you can see over time, how you’re trending. The third one is a concept called Google Cloud Project, which is, for those who are not aware, it’s a way to group Google Cloud resources into buckets.

And the third one is Google Cloud services. So, what you described here is, which of your services emits the most and therefore which ones should you optimize first? Like, again, to go back to impactful actions. And to your point, yes, it is very interesting that if you use products which auto-scale, basically, the carbon attributed to you, the customer, will really follow this auto-scaling behavior. Compare that to a virtual machine that is always on, burning some CPU for almost nothing because you have a server that doesn’t process requests. That is wasting, in a sense, resources.

So, what you describe here is very interesting, which is basically the most optimized products you’re going to pick, the less waste you’re going to have. Now, I also want to be careful because comparing one CPU hour of Cloud Run and one CPU hour of Compute Engine is not comparing apples to apples. Why? Because when you use Cloud Run, I’m not sure if you know, but you are using a regional product. So, a product which has built-in redundancy, which is safe in case of one zone going down in a region.

But that means the Cloud Run infrastructure has to provision a little bit more machines than if it was a zonal product. While Compute Engine, your virtual machine lives in one zone and there is only one machine for you. So, you see how we should also be careful comparing products with other products because fundamentally, they are not offering the same value and they are not running on the same infrastructure. But overall, I think you are correct to say that, you know, avoiding waste, using auto-scaling products, is a good way to reduce your footprint.

Corey: I do want to ask—and this is always a delicate topic because you’re talking about cultural things—how much headwind did you have internally at Google when you had the idea to start exposing this? How difficult was it to bring this to fruition?

Steren: I think we are lucky that our leadership cares about reducing carbon emissions and understood that our customers needed our help to understand their cloud emissions. Like, many customers before we had this tool, we’re trying to some kind of estimate their cloud emissions. And it was—you know, Google Cloud was a black box for them. They did not have access to what you said, to some data that only Google has access to.

And you know, to build that tool, we are using energy measurement of every machine in every data center. We are using, you know, customer-wide resource usage. And that is something that we use to divide the footprint among customers. So, there is some data used to compute those numbers that only Google Cloud has access to. And indeed, you’re correct; it required some executive approval which we received because many of our leaders believe that, you know, this is the right thing to do, and this is helping customers towards the same goal as Google, which is being net-zero and carbon-free.

Many of our customers have made some sustainability commitments, and they need our help to meet those goals. So yeah, we did receive approval, first to share the per-region characteristics. This was already, you know, a first in the industry where a cloud provider disclosed that not every region is equal and some are emitting more carbon than others. And second, another approval which was to disclose a per customer carbon footprint, which is broken down by service, project, region, using some, you know, if you touch a little bit on the methodology, you know, it uses energy consumption, resource usage, and carbon intensity coming from a partner of ours to compute, basically, a per customer footprint.

Corey: My question for you is, on some level, given that Google is already committed to being net-zero across the board for all of its usage, why do customers care? Why should they care? Effectively, haven’t you made that entirely something that is outside of their purview? You’ve solved the problem, either way.

Steren: This is where we should explore it a bit more the kinds of carbon emissions that exist. For a customer, their emissions linked to the cloud usage is all considered the indirect emissions. This, in the Greenhouse Gas Protocol Standard, this is called Scope 3. So, our Google Cloud emissions are the customers’ Scope 3 emissions; they are all indirect for them. But those indirect emissions, what I mentioned as being net-zero are the emissions coming from electricity usage.

So, to power those data centers, those data centers are located in certain electricity grids. Those electricity grids might be using energy sources that emit more or less carbon, right? Simply put, if in a given place, the electricity comes from coal, it will be emitting a lot of carbon compared to when electricity comes from solar, for example. So, you see how the location itself determines the carbon intensity. And these are the emissions coming from electricity usage, right?

So, these are neutralized by Google purchasing as much renewable energy. But there are also types of emissions. For example, when a data center loses connection to the grid, we startup diesel generators. Those diesel generators will directly emit carbon. This is called Scope 1 emissions.

And finally, there is the carbon emissions that are coming from the manufacturing of those servers, of those data centers. And this is called Scope 3 emissions. And the goal of Google is for the emissions coming from electricity to be always coming from carbon-free sources.

So, this is a change that we’ve recently released to Google Cloud Carbon Footprint, which is now we also break down your emissions by scope. So, they are all Scope 3 for you, the customer, they are all indirect emissions for you, the customer, but now those indirect emissions, you can see how much is coming from diesel generators, how much is coming from electricity consumption, and how much is coming from manufacturing of the data center, and other, like, upstream, downstream activities. And yeah, overall, this is something that customers do need to report on.

Corey: I think that’s very fair. I do want to thank you for taking so much time to speak with me. And instead of the usual question I’d like to ask here of where can people go to find out more because we have a bunch of links for that, instead, I want to ask something a little bit different here, which is, what are the takeaways that customers or prospective customers should really have around their carbon footprint when it comes to cloud?

Steren: So, I would recommend our audience to consider carbon emissions in your cloud infrastructure decisions. And my advice is, first, move to the cloud. Like, we’ve talked that Google Cloud has very well-optimized data centers. Like, your cloud gross carbon emissions are anyway going to be much lower than any on-premise carbon emissions. And by the way, if you use Google Cloud, your net operational emissions are zero.

Second action is pick the region with the lowest carbon impact. Like we discussed that this is probably a low-effort action, if possible, that will have a lot of impact on your gross carbon emissions. And you know, if you want to go further, try to schedule those workloads when the electricity is the greenest, you know, when the sun is shining, the wind is blowing, for example, or try to schedule those workloads in regions which have the lowest impact. And yeah, Google Cloud gives you all the tools to do that, the tools to optimize your region selection, and the tools to report and reduce your gross carbon emissions. We haven’t talked about it, but Google Cloud Carbon Footprint will even send you some proactive recommendations of things to do to reduce your emissions.

For example, if you have a project, a machine that you forgotten, Google Cloud Carbon Footprint, will recommend you to delete it and we’ll tell you how much carbon you would save by deleting it, as well as dollar, of course.

Corey: It’s funny because I feel like there’s a definite alignment between my view of cloud economics and the carbon perspective on this, which is step one, everyone wins if you turn things off when you’re not using them. What a concept. I sometimes try and take it too far of, ‘turn off all of production because your company’s terrible.’ Yeah, it turns out, that doesn’t work super well. But the idea of step one, turn it off, especially when you’re not using it. And if you’re never using it, why would you want to pay for it? That becomes a very clear win for everyone involved. I think that in the fullness of time, economics are what are going to move the needle on driving further adoption about this. I have to guess that you see the same thing from where you are?

Steren: Yes, very often working to reduce your carbon footprint is also working to reduce your bill. And we’ve also observed—not always—but some correlation between regions that have the lowest carbon impact and regions that are the cheapest. So, in a sense, this region selection, optimizing for price and carbon is often optimizing for the same thing. It’s not always true, but it is often true.

Corey: I really want to thank you for spending so much time to talk with me about this. This has definitely giving me a lot of food for thought, and I have to imagine that this will not be our last conversation around the topic.

Steren: Well, thanks for having me. And I’m very happy to talk to you in the podcast, of course.

Corey: Steren Giannini, product lead for Google Cloud Carbon Footprint and Google Cloud Run. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an angry screed about how climate change isn’t real as you sit there wondering why it’s 120 degrees in March.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Alex

Alex Su is a lawyer who's currently the Head of Community Development at Ironclad, the #1 contract lifecycle management technology company that's backed by Accel, Sequoia, Y Combinator, and other leading investors. Prior to joining Ironclad, Alex sold cloud software to legal departments and law firms on behalf of early stage startups. Alex maintains an active presence on social media, with over 180,000 followers across Twitter, LinkedIn, Instagram, and TikTok.

Links Referenced:

  • Ironclad: https://ironcladapp.com/
  • LinkedIn: https://www.linkedin.com/in/alexander-su/
  • Twitter: https://twitter.com/heyitsalexsu
  • Instagram: https://www.instagram.com/heyitsalexsu/
  • TikTok: https://www.tiktok.com/@legaltechbro

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Honeycomb. When production is running slow, it’s hard to know where problems originate. Is it your application code, users, or the underlying systems? I’ve got five bucks on DNS, personally. Why scroll through endless dashboards while dealing with alert floods, going from tool to tool to tool that you employ, guessing at which puzzle pieces matter? Context switching and tool sprawl are slowly killing both your team and your business. You should care more about one of those than the other; which one is up to you. Drop the separate pillars and enter a world of getting one unified understanding of the one thing driving your business: production. With Honeycomb, you guess less and know more. Try it for free at honeycomb.io/screaminginthecloud. Observability: it’s more than just hipster monitoring.

Corey: I come bearing ill tidings. Developers are responsible for more than ever these days. Not just the code that they write, but also the containers and the cloud infrastructure that their apps run on. Because serverless means it’s still somebody’s problem. And a big part of that responsibility is app security from code to cloud. And that’s where our friend Snyk comes in. Snyk is a frictionless security platform that meets developers where they are - Finding and fixing vulnerabilities right from the CLI, IDEs, Repos, and Pipelines. Snyk integrates seamlessly with AWS offerings like code pipeline, EKS, ECR, and more! As well as things you’re actually likely to be using. Deploy on AWS, secure with Snyk. Learn more at Snyk.co/scream That’s S-N-Y-K.co/scream

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’ve been off the beaten path from the traditional people building things in cloud by the sweat of their brow and the snark on their Twitters. I’m joined today by Alex Su, who’s the Head of Community Development at Ironclad, and also relatively well-renowned on the TikToks, as the kids say. Alex, thank you for joining me.

Alex: Thank you so much for having me on the show.

Corey: It’s always been an interesting experience because I joined TikTok about six months or so ago, due to an escalatingly poor series of life choices that continue to fail me, and I have never felt older in my life. But your videos consistently tend to show up there. You are @legaltechbro, which sounds like wow, I hate all of those things, and yet your content is on fire.

How long have you been doing the public dance thing, for lack of a better term? I don’t even know what they call it. I know how to talk about Twitter. I know how to talk about LinkedIn—sad. LinkedIn is sad—but TikTok is still something I’m trying to wrap my ancient brain around.

Alex: Yeah, I felt out of place when I first made my first TikTok. And by the way, I’m known for making funny skits. I have actually never danced. I’ve always wanted to, but I don’t think I have that… that talent. I started posting TikToks in, I will call it—let’s call it the fall of 2020. So, after the pandemic.

Before that, I had been posting consistently on LinkedIn for, gosh, ever since 2016, when I got into legal tech. And during the pandemic, I tried a bunch of different things including making funny skits. I’d seen something somewhere online if somebody’s making fun of the doctor life. And so, I thought, hey, I could do that for legal too. And so, I made one with iMovie. You know, I recorded it on Zoom.

And then people started telling me, “Hey, you should get on this thing called TikTok.” And so, I resisted it for a while because I was like, “This is not for me.” But at some point, I said, “I’ll try this out. The editing seems pretty easy.” So, I made a couple of videos poking fun at the life of a law firm lawyer or a lawyer working for a corporate legal department.

And on my fourth video, I went massively viral. Like, unexpected went viral, like, millions of—I think two million or so views. And I found myself with a following. So, I thought, “Hey, I guess this is what I’m doing now.” And so, it’s been, I don’t know, a year-and-a-half since then, and I’ve been continuously posting these skits.

Corey: It’s like they say the worst thing can happen when you go into a casino and play for the first time is you win.

Alex: [laugh].

Corey: You get that dopamine hit, and suddenly, well now, guess what you’re doing for the rest of your life? There you go. It sounds like it worked out for you in a lot of fun ways. Your skits about big law of life definitely track. My wife used to work in that space, and we didn’t meet till she was leaving that job because who has time to date in those environments?

But I distinctly remember one of our early dates, we went out to meet a bunch of her soon-to-be-former coworkers at something like eight or nine o’clock in Los Angeles on a Friday night. And at the end of it, we went back to one of our places, and they went back to work. Because that is the lifestyle, apparently, of being in big law. I don’t have the baseline prerequisites to get into law school, to let alone get the JD and then go to work in big law, and looking at that lifestyle, it’s, “Yeah, you know, I don’t think that’s for me.” Of course, I say that, and then three days later, I was doing a middle of the night wake up because the pager went off.

Like, “Oh, are you a doctor?” And the pager is like, “Holy shit. This SSL certificate expires in 30 days.” It’s, yeah. Again, life has been fun, but it’s always been one of those things that was sort of, I guess, held in awe. And you’re putting a very human face on it.

Alex: Yeah. You know, I never expected to be in big law either, Corey. Like, I was never good at school, but as I got older, I found a way to talk my way into, like, a good school. I hustled my way into a job at a firm that I never imagined I could get a job at. But once I got in, that’s when I was like, “Okay, I don’t feel like I fit in.”

And so, I struggled but I still you know grinded it out. I stayed at the job for a couple of years. And I left because I was like, “This is not right for me.” But I never imagined that all of those experiences in big law ended up being the source material for my content, like, eight years after I’d left. So, I’m very thankful that I had that experience even if it wasn’t a good fit for me. [laugh].

Corey: And on some level, it feels like, “Where do you get your material from?” It’s, “Oh, the terrible things that happened to me. Why do you ask?”

Alex: That’s basically it. And people ask me, they say, you know, “You haven’t worked in that environment for eight years. It’s probably different now, right?” Well, no. You know, the legal industry is not like the tech industry. Like, things move very slowly there.

The jokes that made people laugh back then, you know, 10 years ago, even 20 years ago, people still laugh at today because it’s the same way things have always worked. So, again, I’m very thankful that that’s been the case. And, you know, I feel like, the reason why my content is popular is because a lot of people can resonate with it. Things that a lot of people don’t really talk about publicly, about the lifestyle, the culture, how things work in a large firm, but I make jokes about it, so people feel comfortable laughing about it, or commenting and sharing.

Corey: I want to get into that a little bit because when you start seeing someone pop up again and again and again on TikTok, you’re one of those, “Okay, I should stalk this person and figure out what the hell their story is.” And I didn’t have to look very far in your case because you’re very transparent about it. You’re the head of community development at a company called Ironclad, and that one threw me for a little bit of a loop. So, let’s start with the easy question, I suppose. What is Ironclad?

Alex: We’re a digital contracting technology that helps accelerate business contracts. Companies deal with contracts of all types; a lot of times it gets bogged down in legal review. We just help with that process to make that process move faster. And I never expected I’d be in this space. You know, I always thought I was going to be a trial lawyer.

But I left that world, you know, maybe six years ago to go into the legal technology space, and I quickly saw that contracts was kind of a growing challenge, contracting, whether it’s for sales or for procurement. So, I found myself as a salesperson in legal tech selling, first e-discovery software, and then contracting software. And then I found my way to Ironclad as part of the community team, really to talk about how we can help, but also speaking up about the challenges of the legal profession, of working at a law firm or at a legal department. So, I feel like it’s all been the culmination of all my experiences, both in law and technology.

Corey: In the world in which I’ve worked, half of my consulting work has been helping our clients negotiate their large-scale AWS contracts and the other half is architectural nonsense of, “Hey, if you make these small changes, that cuts your bill in half. Maybe consider doing them.” But something that I’ve learned that is almost an industry-wide and universal truism, is that you want to keep the salespeople and the lawyers relatively separate just due to the absolute polar opposites of incentives. Salespeople are incentivized to sell anything that holds still long enough or they can outrun, whereas lawyers are incentivized to protect the company from risk. No, is the easy answer and everything else is risk that has to be managed. You are one of those very rare folks who has operated successfully and well by blending the two. How the hell did that happen?

Alex: I’m not sure to this day how it happened. But I think part of the reason why I left law in the first place was because I don’t think I fit in. I think there’s a lot of good about having a law degree and being part of the legal profession, but I just wanted to be around people, I wanted to work with people, I didn’t want to always worry about things. And so, that led me to technology sales, which took me to the other extreme. And so, you know, I carried a sales quota for five years and that was such an interesting experience to see where—to both sell technology, but also to see where legal fit into that process.

And so, I think by having the legal training, but also having been part of a sales team, that’s given me appreciation for what both teams do. And I think they’re often at tension with one another, but they’re both there to serve the greater goals of the company, whether it’s to generate revenue or protect against risk.

Corey: I think that there’s also a certain affinity that you may have—I’m just spitballing wildly—one of the things that sales folks and attorneys tend to have in common is that in the public imagination, as those roles are not, shall we call it, universally beloved. There tend to be a fair number of well, jokes, in which case, both sides of that tend to be on the receiving end. I mean, at some level, all you have to do is become an IRS auditor and you’ve got the holy trifecta working for you.

Alex: [laugh]. I don’t know why I gravitated to these professions, but I do think that it’s partly because both of these roles hold a significant amount of power. And if you look at just contracting in general, a salesperson at a company, they’re really the driver of the sales process. Like, if there’s no sale to be made, there’s no contract. On the flip side, the law person, the lawyer, knows everything about what’s inside of the contract.

They understand the legal terms, the jargon, and so they hold an immense amount of power over advising people on what’s going to happen. And so, I think sometimes, salespeople and legal people take it too far and either spend too much time reviewing a contract and lording it over the business folks, or maybe the salesperson is too blase about getting a deal done and maybe bypasses legal and doesn’t go through the right processes. By the way, Corey, these are jokes that I make in my TikToks all the time and they always go viral because it’s so relatable to people. But yeah, that’s probably why people always make jokes about lawyers and salespeople. There’s probably some element of ridiculing people with a significant amount of power within a company to determine these transactions.

Corey: Do you find that you have a better affinity for the folks doing contract work on the seller side or the buyer side? Something they don’t tell you when you run companies is, yeah, you’re going to spend a lot of time working on contracts, not just when selling things, but also when buying things and going back and forth. Aspects of what you’re talking about so far in this conversation have resonated, I guess, with both sides of that for me. What do you have the affinity for?

Alex: I think on the sales side, just because of my experience, you know, I think when you go through a transaction and you’re trying to convince someone to doing something, and this is probably why I wanted to go to law school in the first place. Like I watched those movies, right? I watched A Few Good Men and I thought I’d be standing up in court convincing a jury of something. Little did I know that that sort of interest [crosstalk 00:10:55]—

Corey: Like, Perry Mason breakthrough moment.

Alex: That moment where—the gotcha moment, right? I found that in sales. And so, it was really a thrill to be able to, like, talk to someone, listen to them, and then kind of convince them that, based on what challenges they’re facing, for them to buy some technology. I love that. And I think that was again, tied to why I went to law school in the first place.

I didn’t even know sales was a possible profession because I grew up in an immigrant community that was like, you just go to school, and that’ll lead to your career. But there’s a lot of different careers that are super interesting that don’t require formal schooling, or at least the seven years of schooling you need for law. So, I always identify with the sales side. And maybe that’s just how I am, but obviously, the folks who deal with the buy side, it’s a pretty important job, too.

Corey: There’s a lot of surprise when I start talking to folks in the engineering world. First, they’re in for a rough awakening at times when they learn exactly how much qualified enterprise salespeople can make. But also because being a lawyer without, you know, the appropriate credentials to tie into that, you’re going to have a bad time. There are regulatory requirements imposed on lawyers, whereas to be a salesperson, forget the law degree, forget the bachelor’s, forget the high school diploma, all you really need to be able to do from an academic credential standpoint is show up.

The rest of it is, can you actually sell? Can you have the conversations that convince people to see the outcome that benefits everyone? And I don’t know what that it’s possible, or advised necessarily, to be able to find a way to teach that in some formalized way. It almost feels like folks either have that spark or they don’t. Do you think it’s one of those things that can be taught? Do you think it’s something that people have to have a pre-existing affinity for?

Alex: It’s both, right, because part of it is some people will just—they don’t have the personality to really sell. It’s also like their interest; they don’t want to do that. But what I found that’s interesting is that what I thought would make a good salesperson didn’t end up being true when I looked at the most effective sellers. Like, in my head, I thought, “Oh, this is somebody who’s very boisterous, very extroverted,” but I found that in my experience in B2B SaaS that the most effective sellers are very, very much active listeners. They’re not the people showing up and talking at you. They are asking you about your day-to-day asking about processes, understanding the context of your situation, before making a small suggestion about what you might want to do.

I was very impressed the first time I saw one of these enterprise sellers who was just so good at that. Like, I saw him, and he looked nothing like what I imagined an effective sales guy to look like. And he was really kind and he just, kind of, just talked to me, like, I was a human being, and listened to my answers. So, I do think that there is some element of nature, your talent when it comes to that, but it can also be trained because I think a lot of folks who have sales talent, they don’t realize that they could be good at it. They think that they’ve got to be this extroverted, happy hour, partying, storyteller, where —

Corey: The Type A personality that interrupts people as they’re having the conversation.

Alex: Yeah, yeah.

Corey: Yeah.

Alex: So anyways, I think that’s why it’s a mix of both.

Corey: The conversations that I’ve learned the most from when I’m talking to prospects and clients have been when I asked the quote-unquote, dumb question that I already know the answer to, and then I shut up and I listen. And wow, I did not expect that answer. And when you dig a little further, you realize there’s nuance that—at least in my case—that I’ve completely missed to the entire problem space. I think that is really one of the key differentiators to my mind, that separate people who are good at this role from folks who just misunderstand what the role is based upon mass media, or in other cases—same problem with lawyers—the worst examples, in some cases, of the profession. The pushy used car salesperson or the lawyer they see advertising on the back of a bus for personal injury cases. The world is far more nuanced than that.

Alex: Absolutely. And I think you hit the nail on the head when you said, you know, you ask those questions and let them talk. Because that’s an entire process within the sales process. It’s called discovery, and you’re really asking questions to understand the person’s situation. More broadly, though, I think pitching at people doesn’t seem to work as well as understanding the situation.

And you know, I’ve kind of done that with my content, my TikToks because, you know, if you look at LinkedIn, a lot of people in our space, they’re always prescribing solutions, giving advice, posting content about teaching people things. I don’t do that. As a marketer, what I do is I talk about the problems and create discussions. So, I’ll create a funny video—

Corey: I think you’re teaching a whole generation that maybe law school isn’t what they want to be doing, after all there is that.

Alex: There is that. There is that. It’s a mix of things. But one of the things I think I focus on is talking about the challenges of working with a sales team if you’re an in-house lawyer. And I don’t prescribe technology, I don’t prescribe Ironclad, I don’t say this is what you need to do, but by having people talk about it, they realize, right—and I think this is why the videos are popular—as opposed to me coming out and saying, “I think you need technology because of XYZ.” I think, like, facilitating the conversation of the problem space, that leads people to naturally say, “Hey, I might need something. What do you guys do, by the way?”

Corey: This episode is sponsored in part by our friend EnterpriseDB. EnterpriseDB has been powering enterprise applications with PostgreSQL for 15 years. And now EnterpriseDB has you covered wherever you deploy PostgreSQL on-premises, private cloud, and they just announced a fully-managed service on AWS and Azure called BigAnimal, all one word. Don’t leave managing your database to your cloud vendor because they’re too busy launching another half-dozen managed databases to focus on any one of them that they didn’t build themselves. Instead, work with the experts over at EnterpriseDB. They can save you time and money, they can even help you migrate legacy applications—including Oracle—to the cloud. To learn more, try BigAnimal for free. Go to biganimal.com/snark, and tell them Corey sent you.

Corey: It sounds ridiculous for me to say that, “Oh, here’s my entire business strategy: step one, I shitpost on the internet about cloud computing; step two, magic happens here; and step three people reach out to talk about their AWS bills.” But it’s also true. Is that the pattern that you go through: step one, shitpost on TikTok; step two, magic happens here; and step three people reach out asking to learn more about what your company does? Or is there more nuance to do it?

Alex: I’m still figuring out this whole thing myself, but I will say shitposting is incredibly effective. Because I’m active on Twitter. Twitter is where I start my shitposts. TikTok, I also shitpost, but in video format, I think the number one thing to do is figure out what resonates with people, whether it’s the whole contracting thing or if it’s frustrations about law school. Once you create something that’s compelling, the conversation gets going and you start learning about what people are thinking.

And I think that what I’m trying to figure out is how that can lead to a deeper conversation that can lead to a business transaction or lead to a sale. I haven’t figured it out, right, but I didn’t know that when I started creating content that spoke to people when I was a quota-carrying salesperson, people reached out to me for demo requests, for sales conversations. There is something that is happening in this quote-unquote, “Dark funnel,” that I’m sure you’re very familiar with. There’s something that’s happening that I’m trying to understand, and I’m starting to see.

Corey: This is probably a good thing to the zero in on a bit because to most people’s understanding of the sales process, it would seem that you going out and making something of a sensation out of yourself on the internet, well what are you doing that for? That’s not sales work? How is that sales? That’s just basically getting distracted and going to do something fun. Shouldn’t you be picking up the phone and cold calling people or mass-emailing folks who don’t want to hear from you because you trick them into having a badge scanned somewhere? I don’t necessarily think that is accurate. How do you see the interplay of what you do and sales?

Alex: When you’re selling something like makeup or clothing, it’s a pretty transactional process. You create a video; people will buy, right? That’s B2C. In B2B, it’s a much more complex processes. There’s so many touchpoints. The start of a sales conversation and when they actually buy may take six months, 12 months, years. And so, there’s got to be a lot of touch points in between.

I remember when I was starting out in my content journey, I had this veteran enterprise sales leader, like, your classic, like, CRO. He said to me, “Hey, Alex, your content’s very funny, but shouldn’t you be making cold calls and emails? Like, why are you spending your time doing this?” And I said, “Hey, listen, do you notice that I’m actually sourcing more outbound sales calls than any other sales rep? Like, have you noticed that?”

And he’s like, “Actually, yeah, I did notice that. You know, how are you doing it?” And I was like, “Do you not see that these two are tied? These are not people I just started calling. They are people who have seen my content over time. And this is how it works.”

And so, I think that the B2B world is starting to wise up to this. I think, for example, Ironclad is leading the way on creating a community team to create those conversations, but plenty of B2B companies are doing the same thing. And so, I think by inserting themselves in a conversation—a two-way conversation—during that process, that’s become incredibly effective, far more so than, like, cold-calling a lawyer or a developer who doesn’t want to be bothered by some pushy salesperson.

Corey: Busy, expensive professionals generally don’t want to spend all their time doing that. The cold outreach emails that drive me nuts are, “Hey, can we talk for half an hour?” Yeah, I don’t tend to think in terms of billable hours because that’s not how I do anything that I do, but there is an internal rate that I used to benchmark and it’s what you want me just reach into my pocket and give you how much money for a random opportunity to pitch me on something that you haven’t even qualified whether I need or not? It’s like, asking people for time is worse, in some ways, than asking for money because they can always make more money, but no one can make more time.

Alex: Right, right. That’s absolutely right.

Corey: It’s the lack of awareness of understanding the needs and motivations of your target market. One thing that I found that really aided me back when I was working for other folks was trying to find a company or a management structure that understood and appreciated this. Easy example, when I was setting out as an independent consultant after a few months I’d been doing this and people started to hear about me. But you know, it turns out that there are challenges to running a business that are not recommended for most people. And I debated, do I take a job somewhere else?

So, I interviewed at a few places, and I was talking to one company that’s active in the cloud costing space at the time and they wanted me to come aboard. But discussions broke down because they thought I was, quote, “More interested in thought leadership than I was and actually fixing the bills themselves.” And looking at this now, four years later or so, yeah, they were right. And amazing how that whole thing played out, but that the lack of vision around, there’s an opportunity here, if we can chase it, at least in the places I was at, was relatively hard to come by. Did you luck out in finding a role that works for you in this way or did you basically have to forge it for yourself from the sweat of your brow and the strength of your TikTok account?

Alex: It was uphill at first, but eventually, I got lucky. And you know, part of it was engineered luck. And I’ll explain what I mean. When I first started out doing this, I didn’t expect this to lead to any jobs. I just thought it would support my sales career.

Over time, as the content got more popular, I never wanted to do anything else because I was like, I don’t want to be a marketer. I’m not a—I don’t know anything about demand gen. All I know is how to make funny videos that get people talking. The interesting that happened was that these videos created this awareness, this energy in our space, in the legal space. And it wasn’t long before Ironclad found me.

And you know, Ironclad has always been big on community, has always done things like—like, our CEO, our founder, he said that he used to host these dinners, never talking about Ironclad, but just kind of talking about law school and law with potential clients. And it would lead to business. Like, it’s almost the same concept of, like, not pushing sales on people. And so, Ironclad has always had that in its DNA. And one of our investors, our board members, Jessica Lee from Sequoia, she is a huge believer in community.

I mean, she was the CEO of another company that leveraged community, and so there’s this community element all throughout the DNA of Ironclad. Now, had I not put myself out there with this content, I may not have been discovered by Ironclad. But they saw me, they found me, and they said, “We don’t think about these things like many other companies. We really want to invest in this function.” And so, it’s almost like when you put yourself out there, yes, sometimes some people will say, “What are you doing? Like, this makes no sense. Like, stop doing that.” But there’s going to be some true believers who come out and seek you out and find you.

And that’s been my experience here, like, at Ironclad. Like, people were like, “When you go there, are they going to censor you? Is your content going to be less edgy?” No. Like, they pulled me aside multiple times and said, “Keep being yourself. This is what we want.” And I think that is so special and unique. And part of it is very much lucky, but it’s also when you put yourself out there kind of in a big way, like-minded people will seek you out as well.

Corey: I take the position that part of marketing, part of the core of marketing, is you’ve got to have an opinion. But as soon as you have an opinion, people are going to disagree with you. They’re going to, effectively, forget the human on the other side of it and start taking you for a drag on social media and whatnot. So, the default reaction a lot of people have is oh, I shouldn’t venture opinions forward.

No. People are always going to dislike you for something and you may as well have it be for who you are and what you want to be doing rather than who you’re pretending to be. That’s always been my approach. For me, the failure mode was not someone on Twitter is going to get mad about what I wrote. No one’s going to read it. That’s the failure mode. And the way to avoid that is make it interesting.

Alex: That is a hundred percent relatable to me because I think when I was younger, I was scared. I did worry that I would get in trouble for what I posted. But I realized these people I was worried about, they weren’t going to help me anyways. These are not people who are going to seek me out and help me but then say, “Oh, I saw your content, so now I can’t help you.” They were not going to help me anyways.

But by being authentic to myself and putting things out there, I attracted my own tribe of people who have helped me, right? A lot of my early results from content came not because I reached my target customers; it was because somebody resonated with what I put out there and they carried my message and said, “Hey, you should talk to Alex.” Something special happens when you kind of put yourself out there and say an opinion or share a perspective that not everyone agrees with because that tribe you build ends up helping you a lot. And meanwhile, these other people that might not like it, they probably weren’t going to help you either.

Corey: I maintain that one of the most valuable commodities in the universe is attention. And so, often there’s so much information overload that’s competing for our attention every minute of every day that trying to blend in with the rest of it feels like the exact wrong approach. I’m not a large company here. I don’t have a full marketing department to wind up doing ad buys, and complicated campaigns, and train a team of attacking interns to wind up tackling people to scan their badges at conferences. I’ve got to work with what I’ve got.

So, the goal I’ve always had is trigger the Rolodex moment where someone hears about a problem in the AWS billing space—ideally—and, “Oh, my God, you need to talk to Corey about that.” And it worked, for better or worse. And a lot of it was getting lucky, let’s be very clear here, and people doing me favors that they had no reason to do and I’ll never be able to repay. But being able to be in that space really is what made the difference. Now, the downside, of course, when you start doing that is, how do you go back to what happened before?

If you decide okay, well, it’s been a fun run for you and Ironclad. And yeah, TikTok. Turns out that is, in fact, for kids; time to go somewhere else. Like, I don’t know that you would fit into your old type of job.

Alex: Yeah. No, I wouldn’t. But very early on, I realized, I said, “If I’m going to find meaningful work, it’s okay to be wrong.” And when I went to big law, I realized this is not right for me. That’s okay. I’m just not going to get another big law job.

And so, when people ask me, “Hey, now that you’ve put yourself out there, you probably can’t get a job at a big firm anymore.” And that’s okay to me because I wasn’t going to go back anyways. But what I have found, Corey, is that there’s this other universe of people, whether it’s a entrepreneur, smaller businesses, technology companies, they would be interested in working with me. And so, by being myself, I may have blocked out a certain level of opportunities or a safety net, but now I’m kind of in this other world where I feel very confident that I won’t have trouble finding a job. So, I feel very lucky to have that, but that’s why I also don’t worry about the possibility of not going back.

Corey: Yeah, I’ve never had to think about the idea of, well, what if I go have to get a job again? Because at that point, it means well, it’s time to let every one at the company who is depending on the go, and that’s the bigger obstacle because, let’s be honest, I’m a white guy in tech, and I look like it. My failure mode is basically a board seat and a book deal because of inherent bias in the system.

Alex: [laugh]. Oh, my god.

Corey: That’s the outcome that, for me personally, I will be just fine. It’s the other people took a chance on me. I’m terrified of letting them down. So far, knock on wood, I haven’t said anything too offensive in public is going to wind up there. That’s also not generally my style.

But it is the… it is something that has weighed on me that has kept me from I guess, thinking about what would my next job be? I’m convinced this is the last job I’ll ever have, if for no other reason that I’ve made myself utterly unemployable.

Alex: [laugh]. Well, I think many of us aspire to find that perfect intersection of what you love doing and what pays the bills. Sounds like you’ve found it, I really do feel like I found it, too. I never imagined I’d be doing what I do now. Which is also sometimes hard to describe.

I’m not making TikToks for a living; I’m just on the community team, doing events—I’m getting to work with people. I’m basically doing the things that I wanted to do that led me to quit that job many years ago, that big law job many years ago. So, I feel very blessed and for anybody who’s, like, looking for that type of path, I do think that at some point, you do need to kind of shed the safety nets because if you always hang on to the safety nets, whether it’s a big tech job or a big law job, there’s going to be elements of that that don’t fit in with your personality, and you’re never going to be able to find that if you kind of stay there. But if you venture out—and, you know, I admire you for what you’ve done; it sounds like you’re very successful at what you do and get to do what you love every day—I think great things can happen.

Corey: Yeah, I get to insult Amazon for a living. It’s what I love. It’s what I would do if I weren’t being paid. So, here we are. Yeah—

Alex: [laugh].

Corey: I have no sense of self-preservation. It’s kind of awesome.

Alex: I love it.

Corey: But you’re right. It’s… there’s something to be said for finding the thing that winds up resonating with you and what you want to be doing.

Alex: It really does. And you know, I think when I first made the move to technology, to sales, there was no career path. I thought I would—maybe I thought I might be a VP of Sales. But the thing is, when you put yourself out there, the opportunities that show up might not be the ones that you had always seen from the beginning. Like if you ask a lawyer, like, “What can I do if I don’t practice law?” They’re going to give you these generic answers. “Work here. Work there. Work for that company. I’ve seen a lot of people do this.”

But once you put yourself out there in the wilderness, these opportunities arise. And I’ve been very lucky. I mean, I never imagined I’d be a TikTokker. And by the way, I also make memes on Twitter. Couldn’t imagine I’d be doing that either. I learned, like, Mematic, these tools. Like, you know, like, I’m immersed in this internet culture now.

Corey: It is bizarre to me and I never saw it coming either. For better or worse, though, here we are, stuck at it.

Alex: [laugh].

Corey: I really want to thank you for taking so much time to speak with me today. If people want to learn more about what you’re up to and follow along for the laughs, if nothing else, where’s the best place for them to find you?

Alex: The best way to find me is on LinkedIn; just look up Alex Su. But I’m around and on lots of social media platforms. You can find me on Twitter, on Instagram, and on TikTok, although I might be a little bit embarrassed of what I put on TikTok. I put some crazy gnarly stuff out there. But yeah, LinkedIn is probably the best place to find me.

Corey: And we will put links to all of it in the show notes, and let people wind up making their own decisions. Thanks so much for your time, Alex. I really appreciate it.

Alex: Corey, thank you so much for having me. This was so much fun.

Corey: Alex Su, Head of Community Development at Ironclad. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an angry insipid comment talking about how unprofessional everything we talked about is that you will not be able to post for the next six months because it’ll be hung up in legal review.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Jon

A husband, father of 3 wonderful kids who turned Podcaster during the pandemic. If you told me in early 2020 I would be making content or doing a podcast, I probably would have said "Nah, I couldn't see myself making YouTube videos". In fact, I told my kids, no way am I going to make videos for YouTube. Well, a year later I'm over 100 uploads and my subscriber count is growing.

Links Referenced:

  • LinkedIn: https://www.linkedin.com/in/jon-myer/
  • Twitter: https://twitter.com/_JonMyer
  • jonmyer.com: https://jonmyer.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Honeycomb. When production is running slow, it’s hard to know where problems originate. Is it your application code, users, or the underlying systems? I’ve got five bucks on DNS, personally. Why scroll through endless dashboards while dealing with alert floods, going from tool to tool to tool that you employ, guessing at which puzzle pieces matter? Context switching and tool sprawl are slowly killing both your team and your business. You should care more about one of those than the other; which one is up to you. Drop the separate pillars and enter a world of getting one unified understanding of the one thing driving your business: production. With Honeycomb, you guess less and know more. Try it for free at honeycomb.io/screaminginthecloud. Observability: it’s more than just hipster monitoring.

Corey: DoorDash had a problem. As their cloud-native environment scaled and developers delivered new features, their monitoring system kept breaking down. In an organization where data is used to make better decisions about technology and about the business, losing observability means the entire company loses their competitive edge. With Chronosphere, DoorDash is no longer losing visibility into their applications suite. The key? Chronosphere is an open-source compatible, scalable, and reliable observability solution that gives the observability lead at DoorDash business, confidence, and peace of mind. Read the full success story at snark.cloud/chronosphere. That's snark.cloud slash C-H-R-O-N-O-S-P-H-E-R-E.

Corey: Welcome to Screaming in the Cloud, I’m Corey Quinn. Every once in a while I get to talk to a guest who has the same problem that I do. Now, not that they’re a loud, obnoxious jerk, but rather that describing what they do succinctly is something of a challenge. It’s not really an elevator pitch anymore if you have to sabotage the elevator before you start giving it. I’m joined by Jon Myer. Jon, thank you for joining me. What the hell do you do?

Jon: Corey, thanks for that awesome introduction. What do I do? I get to talk into a microphone. And sometimes I get to stare at myself on camera, whether it makes a recording or not. And either I talk to myself or I talk to awesome people like you. And I get to interview and tell other people’s stories on my show; I pull out the interesting parts and we have a lot of freaking fun doing it.

Corey: I suddenly feel like I’ve tumbled down the rabbit hole and I’m in the wrong side of the conversation. Are we both trying to stand in the same part of the universe? My goodness.

Jon: Is this your podcast or mine? Maybe I should do an introduction right now to introduce you onto it and we’ll see how this works.

Corey: The dueling podcast banjo. I liked the approach quite a bit. So, you have done a lot of very interesting things. For example, once upon a time, you worked at AWS. But you have to go digging to figure that out because everything I’m seeing about you in your professional bio and the rest is forward-looking, as opposed to Former Company A, Former Company B, and this one time I was an early investor in Company C, which means, that’s right, one of the most interesting things about me is that I wrote a check once upon a time, which is never something I ever want to say about myself, ever. You’re very forward-looking, and I strive to do the same. How do you wind up coming at it from that position?

Jon: When I first left AWS—it’s been a year ago, so I served my time—and I actually used to have ex-Amazonian on it and listed on it. But as I continuously look at it, I used to have a podcast called The AWS Blogger. And it was all about AWS and everything, and there’s nothing wrong with them. And what I would hear—

Corey: Oh, there’s plenty wrong with them, but please continue.

Jon: [laugh]. We won’t go there. But anyway, you know, kind of talking about it and thinking about it ex-Amazonian, yeah, that’s great, you put it on your resume, put it on your stuff, and it, you know, allows you that foot in the door. But I want to look at and separate myself from AWS, in that I am my own independent voice. Yes, I worked for them; great company, I’ve learned so much from them, worked with some awesome people there, but my voice in the community has become very engaging and trustworthy. I don’t want to say I’m no longer an Amazonian; I still have some of the guidelines, some of the stuff that’s instilled in me, but I’m independent. And I want that to speak for itself when I come into a room.

Corey: It’s easy as hell, by the way, for me to sit here and cast stones at folks who, “Oh, you’re going to talk about this big company you worked for, even though you don’t work there anymore.” Yeah, I really haven’t worked anywhere that most people would recognize unless they’re, you know, professionally sad all the time. So, I don’t have that luxury; I had to wind up telling a story that was forward-looking just because I didn’t really have much of a better option. You have that option and decided to go in a direction where it presents, honestly as your viewpoint is that your best days are yet to come. And I want to be clear that for folks who are constantly challenged in our space to justify their existence there, usually because they don’t look like our wildly over-represented selves, Jon, they need that credibility.

And when they say that it’s necessary for them, I am not besmirching that. I’m speaking from my own incredibly privileged position that you share. That is where I’m coming from on this, so I don’t want people to hear this as shaming folks who are not themselves wildly over-represented. I’m not talking about you fine folks, I assure you.

Jon: You can have ex-Amazonian on your resume and be very proud of it. You can remove it and still be very proud of the company. There’s nothing wrong with either approach. There are some conversations that I’ll be in, and I’ll be on with AWS folks and I’ll say, “I completely understand where you’re coming from. I’m an ex-Amazonian.” And they’re like, “Oh, you get us. You get the process. You get the everything.”

I just want to look forward that I will be that voice in the community and that I have an understanding of what AWS is and will continuously be. And I have so much that I’m working towards that I’m very proud of where I’ve come from, but I do want to look forward.

Corey: One of these days, I really feel like I should hang out with some Amazonians or ex-Amazonians who don’t know who I am—which is easier to find than you think—and pretend that I used to work there and wonder how long I can keep the ruse going. Just because I’ve been told a few times that I am suspiciously Amazonian for someone who’s never worked there.

Jon: You have a lot of insights on the AWS processes and understanding. I think you could probably keep it going for quite a while. You will have to get that orange lanyard though, when you go to, like—

Corey: I got one once when I was at a New York Summit a couple years ago. My affiliation then, before I started The Duckbill Group, was Last Week in AWS, and apparently, someone saw that and thought that I was the director of Take-this-Job-and-Shove-it, but I’ll serve out my notice until Friday. So, cool; employee lanyard, it was. And I thought this is going to be awesome because I’ll be able to walk around and I’ll get the inside track if people think I work there. And they treated me like crap until I put the customer lanyard back on. It’s, “Oh, it’s better to be a customer at an AWS event than it is to be an employee.” I learned that when the fun way.

Jon: There is one day that I hope to get the press or analyst lanyard. I think it would be an accomplishment for me. But you get to experience that firsthand, and I hate to switch the tables because I know it’s your podcast recording, not mine, but—

Corey: Having the press analyst lanyard is interesting because a lot of people are not allowed to speak to you unless they’ve gone through training. Which, okay, great. I will say that it is a lot nicer walking the expo floor because most of the people working the booths know that means that person is press, generally—they’re not quite as familiar with analysts—but they know that regardless that they’re not going to sell you a damn thing, so they basically give you a little bit of breathing room, which is awesome, especially in these pandemic times. But the challenge I have with it is that very often I want to talk to folks who are AWS employees who may not have gone through press training. And I’ve never gotten anyone in trouble or taken advantage of things that I hear in those conversations and write about them.

Everything I write about is what I’ve experienced in public or as a customer, not based upon privileged inside information. I have so many NDAs at this point, I can’t keep track, so I just make sure everything I talked about publicly cited I have that already.

Jon: Corey, I got to flip the script real quick. I got to give you a shout-out because everybody sees you on Twitter and sees, like, “Oh, my God, he’s saying this negative, that negative towards AWS.” You and I had, I don’t know, it was a 30, 45 minute at the San Francisco Summit, and I think every Summit, we try to connect for a little bit. But that was really the premise I kicked off a lot of our conversations when you joined my podcast. No, this is not my podcast, this is Corey’s, but anyway—

Corey: And just you remember that. Please continue.

Jon: [laugh]. But you know, kind of going off it you have so much insight, so much value, and you kind of really understand the entire processes and all the behind the scenes and everything that’s going on that I was like, “Corey, I got to get your voice out there and show the other side of you, that you’re not there trying to get people in trouble, you never poke fun of an AWS employee. I heard there was some guy named Larry that you do, but we won’t jump into that.”

Corey: One of the things that I think happened is, first and foremost, there is an algorithmic bias towards outrage. When I say nice things about AWS or other providers, which I do periodically, they get basically no engagement. When I say something ridiculous, inflammatory, and insulting about a company, oh, goes around the internet three times. One of the things that I’m slowly waking up to is that when I went into my Covid hibernation, my audience was a quarter of the size it is now. People don’t have the context of knowing what I’ve been up to for the last five or six years. All they see is a handful of tweets.

And yeah, of course, you wind up taking some of my more aggravated moment tweets and put a few of those on a board, and yeah, I start to look a fair bit like a jerk if you’re not aware of what’s going on inside-track-wise. That’s not anyone else’s fault, except my own, and I guess understanding and managing that perception does become something of a challenge. I mean, it’s weird; Amazon is a company that famously prides itself on being misunderstood for long periods of time. I guess I never thought that would apply to me.

Jon: Well, it does. Maybe that’s why most people think you’re an Amazonian.

Corey: You know, honestly, I’ve got to say, there are a lot of worse things people can and do call me. Amazon has a lot to recommend it in different ways. What I find interesting now is that you’ve gone from large companies to sort of large companies. You were at Spot for a hot minute, then you were doing the nOps thing. But one thing that you’ve been focusing on a fair bit has been getting your own voice and brand out there—and we talked about this a bit at the Summit when we encountered each other which is part of what sparked this conversation—you’re approaching what you’re doing next in a way that I don’t ever do myself. I will not do it justice, but what are you working on?

Jon: All right. So Corey, when we talked at the New York Summit, things are actually moving pretty good. And some of the things that I am doing, and I’ve actually had a couple of really nice engagements kind of kick off is, that I’m creating highly engageable, trustworthy content for the community. Now, folks, you’re asking, like, what is that? What is that really about? You do podcasts?

Well, just think about some of the videos that you’re seeing on customer sites right now. How are they doing? How’s the views? How’s the engagement? Can you actually track those back to, like, even a sales engagement in utilizing those videos?

Well, as Jon Myer—and yes, this is highly scalable because guess what I am in talks with other folks to join the crew and to create these from a brand awareness portion, right? So, think about it. You have customers that you want to get engaged with: you have products, you have demos, you have reviews that you want to do, but you can’t get them turned around in a quick amount of time. We take the time to actually dive into your product and pull out the value prop of the exact product, a demo, maybe a review, all right? We do sponsors as well; I have a number of them that I can talk about, soVeeam on AWS, Diabolical Coffee, there’s a couple of other I cannot release just yet, but don’t worry, they will be hitting out there on social pretty soon.

But we take that and we make it an engaging kind of two to three-minute videos. And we say, “Listen, here’s the value of it. We’re going to turn this around, we’re going to make this pop.” And putting this stuff, right, so we’ll take the podcast and I’ll put it on to my YouTube channel, you will get all my syndication, you’ll get all my viewers, you’ll get all my views, you’ll get my outreach. Now, the kicker with that is I don’t just pick any brand; I pick a trusted brand to work with because obviously, I don’t want to tarnish mine or your brand. And we create these podcasts and we create these videos and we turn them around in days, not weeks, not months. And we focus on those who really need to actually present the value of their product in the environment.

Corey: It sounds like you’re sort of the complement to the way that I tend to approach these things. I’ll periodically do analyst engagements where I’ll kick the tires on a product in the space—that’s usually tied to a sponsorship scenario, but not always—where, “Oh, great. You want me to explain your product to people. Great, could I actually kick the tires on it so I understand at first? Otherwise, I’m just parroting what may as well be nonsense. Maybe it’s true, maybe it’s not.”

Very often small companies, especially early stage, do a relatively poor job of explaining the value of their product because everyone who works there knows the product intimately and they’re too close to the problem. If you’re going to explain what this does in a context where you have to work there and with that level of intensity on the problem space, you’re really only pitching to the already converted as opposed to folks who have the expensive problem that gets in the way of them doing their actual job. And having those endless style engagements is great; they periodically then ask me, “Hey, do you want to build a bunch of custom content for us?” And the answer is, “No, because I’m bad at deadlines in that context.”

And finding intelligent and fun and creative ways to tell stories takes up a tremendous amount of time and is something that I find just gets repetitive in a bunch of ways. So, I like doing the typical sponsorships that most people who listen to this are used to: “This episode is sponsored by our friends at Chex Mix.” And that’s fine because I know how to handle that and I have that down to a set of study workflows. Every time I’ve done custom content, I find it’s way more work than I anticipated, and honestly, I get myself in trouble with it.

Jon: Well, when you come across it, you send them our way because guess what, we are actually taking those and we’re diving deep with them. And yes, I used an Amazon term. But if you take their product—yeah [laugh]. I love the reaction I got from you. But we dive into the product. And you said it exactly: those people who are there at the facility, they understand it, they can say, “Yeah, it does this.”

Well, that’s not going to have somebody engaged. That’s not going to get somebody excited. Let me give you an example. Yesterday, I had a call with an awesome company that I want to use their product. And I was like, “Listen, I want to know about your product a little bit more.”

We demoed it for my current company, and I was like, “But how do you work for people like me: podcasters who do a lot of the work themselves? Or a social media expert?” You know, how do I get my content out there? How does that work? What’s your pricing?

And they’re, like, “You know, we thought about getting it and see if there was a need in that space, and you’re validating that there’s a need.” I actually turned it around and I pitched them. I was like, “Listen, I’d love for you guys to be a sponsor on my show. I’d love for you to—let me do this. Let me do some demos. Let’s get together.”

And I pitched them this idea that I can be a spokesperson for their product because I actually believed in it that much just from two calls, 30 minutes. And I said, “This is going to be great for people like me out there and getting the voice, getting the volume out there, how to use it.” I said, “I can show some quick integration setups. You don’t have to have the full-blown product that you sell out the businesses, us as individuals or small groupings, we’re only going to use certain features because, one, is going to be overwhelming, and two, it’s going to be costly. So, give us these features in a nice package and let's do this.” And they’re like, “Let’s set something up. I think we got to do this.”

Corey: How do you avoid the problem where if you do a few pieces of content around a particular brand, you start to become indelibly linked to that brand? And I found that in my early days when I was doing a lot of advisory work and almost DevRel-for-hire as part of the sponsorship story thing that I was doing, and I found that that did not really benefit the larger thing I was trying to build, which is part of the reason that I got out of it. Because it makes sense for the first one; yeah, it’s a slam dunk. And the second one, sure, but sooner or later, it feels like wow, I have five different sponsors in various ways that want me to be building stories and talking about their stuff as I travel the world. And now I feel like I’m not able to do any of them a decent service, while also confusing the living hell out of the audience of, “Who is it you work for again anyway?” It was the brand confusion, for lack of a better term.

Jon: Okay, so you have two questions there. One of them is, how do you do this without being associated with the brand? I don’t actually see a problem with that. Think of a race car; NASCAR drivers are walking around with all their stuff on their jackets, you know, sponsored by this person, this group, that group. Yeah, it’s kind of overwhelming at times, but what’s wrong with being tied to a couple of brands as long as the brands are trustworthy, like yourself? Or you believing those, right? So, there’s nothing wrong with that.

Second is the scalability that you’re talking about where you’re traveling all over the world and doing this and that. And that’s where I’m looking for other leaders and trustworthy community members that are doing this type of thing to join a highly visible team, right? So, now you have a multitude and a diverse group of individuals who can get the same message out that’s ultimately tied to—and I’m actually going to call it out here, I have it already as Myer Media, right? So, it’s going to be under the Jon Myer Podcast; everything’s going to be grouped in together under Myer Media, and then we’re going to have a group of highly engaging individuals that enjoy doing this for a living, but also trust what they’re talking about.

Corey: If you can find a realistic way to scale that, that sounds like it’s going to have some potential significant downstream consequences just as far as building almost a, I guess, a DevRel workshop, for lack of a better term. And I mean, that in the sense of an Andy Warhol workshop style approach, not just a training course. But you wind up with people in your orbit who become associated, affiliated with a variety of different brands. I mean, last time I did the numbers, I had something like 110 sponsors over the last five years. If I become deeply linked to those brands, no one knows what the hell I do because every company in the space, more or less, has at some level done a sponsorship with me at some point.

Jon: I guess I’ll cross that when it happens, or keep that in the top-of-mind as it moves forward. I mean, it’s a good point of view, but I think if we keep our individualism, that’s what’s going to separate us as associated. So, think of advertising, you have a, you know, actor, actress that actually gets on there, and they’re associated with a certain brand. Did they do it forever? I am looking at long-term relationships because that will help me understand the product in-depth and I’ll be able to jump in there and provide them value in a expedited version.

So, think about it. Like, they are launching a new version of their product or they’re talking about something different. And they’re, like, “Jon, we need to get this out ASAP.” I’ve had this long-term relationship with them that I’m able to actually turn it around rather quickly, but create highly engaging out of it. I guess, to really kind of signify that the question that you’re asking is, I’m not worried about it yet.

Corey: What stage or scale of company do you find is, I guess, the sweet spot for what you’re trying to build out?

Jon: I like the small to medium. And looking at it, the small to medium—

Corey: Define your terms because to my mind, I’m still stuck in this ancient paradigm that I was in as an employee, where a big company is anything that has more than 200 people, which is basically everyone these days.

Jon: So, think about startups. Startups, they are usually relatively 100 or less; medium, 200 or less. The reason I like that type of—is because we’re able to move fast. As you get bigger, you’re stuck in processes and you have to go through so many steps. If you want speed and you want scalability, you got to pay attention to some of the stuff that you’re doing and the processes that are slowing it down.

Granted, I will evaluate, you know, the enterprise companies, but the individuals who know the value of doing this will ultimately seek me and say, “Hey, listen, we need this because we’re just kicking this off and we need highly visible content, and we want to engage with our current community, and we don’t know how.”

Corey: This episode is sponsored in part by our friend EnterpriseDB. EnterpriseDB has been powering enterprise applications with PostgreSQL for 15 years. And now EnterpriseDB has you covered wherever you deploy PostgreSQL on-premises, private cloud, and they just announced a fully-managed service on AWS and Azure called BigAnimal, all one word. Don’t leave managing your database to your cloud vendor because they’re too busy launching another half-dozen managed databases to focus on any one of them that they didn’t build themselves. Instead, work with the experts over at EnterpriseDB. They can save you time and money, they can even help you migrate legacy applications—including Oracle—to the cloud. To learn more, try BigAnimal for free. Go to biganimal.com/snark, and tell them Corey sent you.

Corey: I think that there’s a fair bit of challenge somewhere in there. I’m not quite sure how to find it, that you’re going to, I think, find folks that are both too small and too big, that are going to think that they’re ready for this. I feel like this doesn’t, for example, have a whole lot of value until a company has found product-market fit unless what you’re proposing to do helps get them to that point. Conversely, at some point, you have some of the behemoth companies out there, it’s, “Yeah, we can’t hire DevRel people fast enough. We’ve hired 500 of them. Cool, can you come do some independent work for us?” At which point, it’s… great, good luck standing out from the crowd in any meaningful way at that point.

Jon: Well, even a high enterprise as hired X number of DevRels, the way you stand out is your personality and everything that you built behind your personal brand, and your value brand, and what you’re trying to do, and the voice that you’re trying to achieve out there. So, think about it—and this is very difficult for me to, kind of, boost and say, “Hey, listen, if I were to go to a DevRel of, like, say, 50 people, I will stand out. I might be one of the top five, or I might be two at the top five.” It doesn’t matter. But for me why and what I do, the value that I am actually driving across is what will stand out, the engaging conversations.

Every interview, every podcast that I do, at the end, everybody’s like, “Oh, my God, you’re, like, really good at it; you kind of keep us engaging, you know when to ask a question; you jump in there and you dive even deeper.” I literally have five bullet points on any conversation, and these are just, like, two or three sentences, maybe. And they’re not exact questions. They’re just topics that we need to talk about, just like we did going into this conversation. There is nothing that scripted. Everything that’s coming across the questions that you’re pulling out from me giving an answer to one of your questions and then you’re diving deep on it.

Corey: I think that that’s probably a fair approach. And it’s certainly going to lead to a better narrative than the organic storytelling that tends to arise internally. I mean, there’s no better view to see a lot of these things than working on bills. One of my favorite aspects of what I do is I get to see the lies that clients tell to themselves, where it’s—like, they believe these things, but it no longer matches the reality. Like developer environments being far too expensive as a proportion of the rest of their environment. It’s miniscule just because production has scaled since you last really thought about it.

Or the idea that a certain service is incredibly expensive. Well, sure. The way that it was originally configured and priced, it was and that has changed. Once people learn something, they tend to stop keeping current on that thing because now they know it. And that’s a bit of a tricky thing.

Jon: That’s why we keep doing podcasts, you keep doing interviews, you keep talking with folks is because if you look at when you and I actually started doing these podcasts—and aka, like, webinars, and I hate to say webinars because it’s always negative and—you know because they’re not as highly engaging, but taking that story and that narrative and creating a conversation out of it and clicking record. There are so many times that when I go to a summit or an event, I will tell people, they’re like, “So, what am I supposed to do for your podcast?” And we were talking for, like, ten minutes, I said, “You know, I would have clicked record and we would have ten minutes of conversation.” And they’re like, “What?” I was like, “That’s exactly what it is.”

My podcast is all about the person that I’m interviewing, what they’re doing, what they’re trying to achieve, what’s their message that they’re trying to get across? Same thing, Corey. When you kick this off, you asked me a bunch of questions and then that’s why we took it. And that’s where this conversation went because it’s—I mean, yeah, I’m spinning it around and making it about you, sometimes because obviously, it’s fun to do that, and that’s normally—I’m on the other side.

Corey: No, it’s always fun to wind up talking to people who have their own shows just because it’s fun watching the narrative flow back and forth. It’s kind of a blast.

Jon: It’s almost like commentators, though. You think about it at a sporting event. There’s two in the booth.

Corey: Do a team-up at some point, yeah.

Jon: Yeah.

Corey: In fact, doing the—what is it like the two old gentlemen in the Sesame Street box up in the corner? I forget their names… someone’s going to yell at me for that one. But yeah, the idea of basically kibitzing back and forth. I feel like at some level, we should do a team up and start doing a play-by-play of the re:Invent keynotes.

Jon: Oh… you know what, Corey, maybe we should talk about this offline. Having a huge event there, VIP receptions, a podcasting booth is set up at a villa that we have ready to go. We’re going to be hosting social media influencers, live-tweeting happening for keynotes. Now, you don’t have to go to the keynotes personally. You can come to this room, you can click record, we’ll record a live session right there, totally unscripted, like everything else we do, right? We’ll have a VIP reception, come in chat, do introductions. So, Corey, love to have you come into that and we can do a live one right there.

Corey: Unfortunately, I’m going to be spending most of re:Invent this year dressed in my platypus costume, but you know how it works.

Jon: [laugh]. Oh man, you definitely got to go for that because oh, I have a love to put that on the show. I’m actually doing something not similar, but in true style that I’ve been going to the last couple of re:Invents I will be doing something unique and standing out.

Corey: I’m looking forward to it. It’s always fun seeing how people continue to successfully exceed what they were able to do previously. That’s the best part, on some level, is just watching it continually iterate until you’re at a point where it just becomes, well frankly, either ridiculous or you flame out or it hits critical mass and suddenly you launch an entire TV network or something.

Jon: Stay tuned. Maybe I will.

Corey: You know, it’s always interesting to see how that entire thing plays out. Last question before we call it a show. Talk to me about your process for building content, if you don’t mind. What is your process when you sit down and stare at—at least from my perspective—that most accursed of all enemies, a blank screen? “All right time to create some content, Jackwagon, better be funny. And by the way, you’re on a deadline.” That is the worst part of my job.

Jon: All right, so the worst part of your job is the best part of my job. I have to tell you, I actually don’t—and I’m going to have to knock on wood because I don’t get content block. I don’t sit at a screen when I’m doing it. I actually will go for a walk or, you know, I’ll have my weirdest ideas at the weirdest time, like at the gym, I might have a quick idea of something like that and I’ll have a backlog of these ideas that I write down. The thing that I do is I come down, I open up a document and I’ll just drop this idea.

And I’ll write it out as almost as it seems like a script. And I’ll never read it verbatim because I look at it and be like, “I know what I’m going to say right now.” An example, if you take a look at my intros that I do for my podcast, they are done after the recording because I recap what we do on a recording.

So, let’s take this back. Corey will talk about the one you and I just did. And you and I we hopped on, we did a recording. Afterwards, I put together the intro. And what I’m going to say the intro, I have no freaking clue until I actually get to it, and then all of a sudden, I think of something—not at my desk, but away from my desk—what I’m going to say about you or the guest.

An example, there was a gentleman I did his name’s called Mat Batterbee, and he’s from the UK. And he’s a Social Media Finalist. And he has this beard and he always wears, like, this hat or something. And I saw somebody on Twitter make a comment about, you know, following in his footsteps or looking like him. So, they spoofed him with a hat and everything—glasses.

I actually bought a beard off of Amazon, put it on, glasses, hat, and I spoofed him for the intro. I had this idea, like, the day before. So, thank goodness for Prime delivery, that I was able to get this beard ASAP, put it on. One take; I only tried to do one take. I don’t think I’ve ever recorded any more.

Corey: I have a couple of times sometimes because the audio didn’t capture—

Jon: Yeah.

Corey: —but that’s neither here nor there. But yeah, I agree with you, I find that the back-and-forth with someone else is way easier from a content perspective for me. Because when you and I started talking, on this episode, for example, I had, like, three or four bullet points I wanted to cover and that’s about it. The rest of it becomes this organic freewheeling conversation and that just tends to work when it’s just me free-associating in front of the camera, it doesn’t work super well. I need something that’s a bit more structured in that sense. So apparently, my answer is just never be alone, ever.

Jon: [laugh]. The content that I create, like how-to tutorials, demos, reviews, I’ll take a lot more time on them and I’ll put them together in the flow. And I record those in certain sections. I’ll actually record the demo of walking through and clicking on everything and going through the process, and then I will actually put that in my recording software, and then I will record against it like a voiceover.

But I don’t record a script. I actually follow the flow that I did and in order to do that, I understand the product, so I’ll dive deep on it, I’ll figure out some of the things using keywords along the way to highlight the value of utilizing it. And I like to create these in, like, two to three minutes. So, my entire process of creating content—podcast—you know what we hop on, I give everybody the spiel, I click record and I say, “Welcome.” And I do the introduction. I cut that out later. We talk. I’ll tell you what, I never edited anything throughout the entire length of it because whatever happens happens in his natural and comes across.

And then I slap on an ending. And I try to make it as quick and as efficiently as possible because if I start doing cuts, people are going to be, like, “Oh, there’s a cut there. What did he cut out?” Oh, there’s this. It’s a full-on free flow. And so, if I mess up and flub or whatever it is, I poke fun of myself and we move on.

Corey: Oh, I have my own favorite punching bag. And I honestly think about that for a second. If I didn’t mock myself the way that I do, I would be insufferable. The entire idea of being that kind of a blowhard just doesn’t work. From my perspective, I am always willing to ask the quote-unquote dumb question.

It just happens to turn out but I’m never the only person wondering about that thing and by asking it out loud, suddenly I’m giving a whole bunch of other folks air cover to say, “Yeah, I don’t know the answer to that either.” I have no problem whatsoever doing that. I don’t have any technical credibility to worry about burning.

Jon: When you start off asking and say, “Hey, dumb question or dumb question,” you start being unsure of yourself. Start off and just ask the question. Never say it’s a dumb question because I’ll tell you what, like you said, there’s probably 20 other people in that room that have the same question and they’re afraid to ask it. You can be the one that just jumps up there and says it and then you’re well-respected for it. I have no problem asking questions.

Corey: Honestly, the problem I’ve got is I wish people would ask more questions. I think that it leads to such a better outcome. But people are always afraid to either admit ignorance. Or worse, when they do ask questions just for the joy they get from hearing themselves talk. We’ve all been conference talks where you there’s someone who’s just asking the question because they love the sound of their own voice. I say, they, but let’s be serious; it’s always a dude.

Jon: That is very true.

Corey: So, if people want to learn more about what you’re up to, where’s the best place to go?

Jon: All right, so the best place to go is to follow me on LinkedIn. LinkedIn is my primary one, right? Jon Myer; can’t miss me. At all. Twitter, I am active on Twitter. Not as well as Corey; I would love to get there one day, but my audience right now is LinkedIn.

Else you can go to jonmyer.com. Yes, that’s right, jonmyer.com. Because why not? I found I have to talk about this just a little bit. And the reason that I changed it—I actually do own the domain awsblogger, by the way and I still have it—is that when I was awsblogger, I had to chan—I didn’t have to change anything’ nobody required me to, but I changed it to, like, thedailytechshow. And that was pretty cool but then I just wanted to associated with me, and I felt that going with jonmyer, it allowed me not having to change the name ever again because, let’s face it, I’m not changing my name. And I want to stick with it so I don’t have to do a whole transition and when this thing takes off really huge, like it is doing right now, I don’t have to change the name.

Corey: Yeah. I would have named it slightly differently had I known was coming. But again, this far in—400 some-odd episodes in last I checked recorded—though I don’t know what episode this will be when it airs—I really get the distinct impression that I am going to learn as I go and, you know, you can’t change that this far in anymore.

Jon: I am actually rounding so I’m not as far as you are with the episodes, but I’m happy to say that I did cross number 76—actually 77; I recorded yesterday, so it’s pretty good. And 78 tomorrow, so I am very busy with all the episodes and I love it. I love everybody reaching out and enjoying the conversations that I have. And just the naturalness and the organicness of the podcast. It really puts people at ease and comfortable to start sharing more and more of their stories and what they want to talk about.

Corey: I really want to thank you for being so generous with your time and speak with me today. Thanks. It’s always a pleasure to talk with you and I look forward to seeing what you wind up building next.

Jon: Thanks, Corey. I really appreciate you having me on. This is very entertaining, informative. I had a lot of fun just having a conversation with you. Thanks for having me on, man.

Corey: Always a pleasure. Jon Myer, podcaster extraordinaire and content producer slash creator. The best folks really have no idea what to refer to themselves and I am no exception, so I made up my own job title. I am Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry, insulting comment telling me that I’m completely wrong and that you are a very interesting person. And then tell me what company you wrote a check to once upon a time.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Tim

Tim’s tech career spans over 20 years through various sectors. Tim’s initial journey into tech started as a US Marine. Later, he left government contracting for the private sector, working both in large corporate environments and in small startups. While working in the private sector, he honed his skills in systems administration and operations for large Unix-based datastores.

Today, Tim leverages his years in operations, DevOps, and Site Reliability Engineering to advise and consult with clients in his current role. Tim is also a father of five children, as well as a competitive Brazilian Jiu-Jitsu practitioner. Currently, he is the reigning American National and 3-time Pan American Brazilian Jiu-Jitsu champion in his division.

Links Referenced:

  • Twitter: https://twitter.com/elchefe

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Honeycomb. When production is running slow, it’s hard to know where problems originate. Is it your application code, users, or the underlying systems? I’ve got five bucks on DNS, personally. Why scroll through endless dashboards while dealing with alert floods, going from tool to tool to tool that you employ, guessing at which puzzle pieces matter? Context switching and tool sprawl are slowly killing both your team and your business. You should care more about one of those than the other; which one is up to you. Drop the separate pillars and enter a world of getting one unified understanding of the one thing driving your business: production. With Honeycomb, you guess less and know more. Try it for free at honeycomb.io/screaminginthecloud. Observability: it’s more than just hipster monitoring.

Corey: I come bearing ill tidings. Developers are responsible for more than ever these days. Not just the code that they write, but also the containers and the cloud infrastructure that their apps run on. Because serverless means it’s still somebody’s problem. And a big part of that responsibility is app security from code to cloud. And that’s where our friend Snyk comes in. Snyk is a frictionless security platform that meets developers where they are - Finding and fixing vulnerabilities right from the CLI, IDEs, Repos, and Pipelines. Snyk integrates seamlessly with AWS offerings like code pipeline, EKS, ECR, and more! As well as things you’re actually likely to be using. Deploy on AWS, secure with Snyk. Learn more at Snyk.co/scream That’s S-N-Y-K.co/scream

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. A bit of a sad episode Today. I am joined by Duckbill Group principal cloud economist, Tim Banks, but by the time this publishes, he will have left the Duckbill nest, as it were. Tim, thank you for joining me, and can I just start by saying, this is sad?

Tim: It is. I have really enjoyed being with Duckbill and I will never forget that message you sent me. It’s like, “Hey, would you like to do this?” And I was like, “Boy would I.” It’s been a fantastic ride and I have enjoyed working with a friend. And I’m glad that we remain friends to this day and always will be, so far as I can tell.

Corey: Yes, yes. What you can’t see while recording this, I’m actually sitting in the same room as Tim with a weapon pointed at him to make sure that he stays exactly on message. Yeah, I kid. There’s been a lot that’s happened over the last year. We only got to spend time together in person once at re:Invent. I think because re:Invent is such a blur for me, I don’t remember who the hell I talk to.

Someone can walk up and say, “Oh yeah, we met at re:Invent,” and I’ll nod and say, “Oh yeah,” and I will have no recollection of that whatsoever. But you don’t argue with people. But I do distinctly remember hanging out with you there. But since then, it’s been a purely distributed company, purely distributed work.

Tim: Yeah, that’s the only time I’ve seen you since I’ve worked here. It’s the only time I met Mike. But it’s weird because it’s like, someone you work with you see every day virtually and talk to, and then you actually get to, like, IRL them and like, “Oh, wow. I had all these, kind of, conceptions of, you know, what you are or who you are as a person, and then you get to, like, check yourself. Was I right? Was I wrong?” I was like, “Oh, you’re taller than I thought; you’re shorter than I thought,” you know, whatever it was.

But I think the fun part about it was we all end up being so close by the nature of how we work that it was just like going back and seeing family after a while; you already know who they are and how they are and about them. So, it felt good, but it felt familiar. That’s a great feeling to have. To me, that’s a sign of a very successful distributed culture.

Corey: Yeah, it’s weird the kinds of friendships we’ve built during the pandemic. When I was in New York for the summit, I got to meet Linda Haviv at AWS for the first time, despite spending the past year or so talking to her repeatedly. As I referred to her the entire time I was in New York, this is Linda, my new old friend because that is exactly how it felt. It’s the idea of meeting someone in person that you’ve had a long-term ongoing friendship with. It’s just a really—it’s a strange way Everything’s new but it’s not, all at the same time.

It reminds me of the early days of the internet culture where I had more friends online than off, which in my case was not hard. And finally meeting them, some people were exactly like they were described and others were nothing at all like they presented. Now that we have Zoom and this constant level of Slack chatter and whatnot, it’s become a lot easier to get a read on what someone is like, I think.

Tim: I think so too, you know, we’ve gotten away—and I think largely because of the pandemic—of just talking about work at work, right? The idea of embracing, you know, almost a cliche of the whole person. But it’s become a very necessary thing as people have dealt with pandemic, social upheaval, political climates, and whatever, while they’re working from home. You can’t compartmentalize that safely in perpetuity, right? So, you do end up getting to know people very well, especially in what their concerns are, what their anxieties are, what makes them happy, what makes them sad, things that go on in their lives.

You bring all that to your distributed culture because it’s not like you leave it at the door, when you walk out. You’re not walking out anymore; you’re walking to another room, and it’s hard to walk away from those things in this day and age. And we shouldn’t have to, right? I feel like for a successful and nurturing culture—whatever it is, whether it’s tech culture, whether it’s whatever kind of work culture—you can’t say, “I only want your productivity and nothing else about you,” and expect people to sustain that. So, you see these companies are, like, you know, “We don’t have political discussions. We don’t have personal discussions. We’re just about the work.” I’m like, “All right, well, that’s not going to last.” A person cannot just be an automaton in perpetuity and expect them to grow and thrive.

Corey: And this is why you’re leaving. And I want to give that a little context because without, sounds absolutely freaking horrifying. You’ve been a strong advocate for an awful lot of bringing the human to work, on your philosophy around leadership, around management. And you’ve often been acting in that capacity throughout, I would say, the majority of your career. But here at The Duckbill Group, we don’t have a scale of team where you being the director of the team or leader of the team is going to happen in anything approaching the near or mid-term.

And so, much of your philosophy is great and all because it’s easy to sit here at a small company and start talking about, “Oh, this is how you should be doing it.” You have the opportunity to wind up making a much deeper impact on a lot more people from a management perspective, but you do in fact, need a team to manage as opposed to sitting around there, “Oh, yeah. Who do you manage?” “This one person and I’m doing all of these things to make their life and job awesome.” It’s like, “Yeah, how many hours a week are you spending in one-on-ones?” “20 to 25.”

Okay, maybe you need a slightly larger team so you can diffuse that out a little bit. And we are definitely sad to be losing you; super excited to see where you wind up going next. This has been a long time coming where there are things that you have absolutely knocked out of the park here at The Duckbill Group, but you also have that growing—from what I picked up on anyway—need to set a good management example. And lord knows this industry needs more of those. So first, sad to lose you. Secondly, very excited for where you wind up next and what they’re in for, even though it has a strong likelihood that they don’t know the half of it yet.

Tim: One of the things that I like about The Duckbill Group and how my time here has been is the first thing that I was asked in the interview was very sincere, like, “Well, what’s your next job?” And I was very clear. It’s like, “After this, I want to be a director or VP of engineering because I would like to be a force multiplier, right?” I would like to make engineering orgs better. I would like to make engineering practices better. I want to make the engineers better, right?

And not by driving KPIs and not by management, right, not administrative functions. I want to do it via leadership. I want to do it by setting examples, making safe places for people, making people feel like they’re important and invested in, nurturing them, right? I’ve said this before I—this analogy was getting me somewhere else and I love, it’s like, if I plant a tree and I want it to grow apples, right, I’m not going to sit there and put a number down of apples it’s expected to produce, and then put it on a performance plan if it doesn’t get that number of apples, right? I need to nurture the tree, I need to fertilize it, I need to protect it, I need to keep it safe, I need to keep it safe from the elements, I need to make sure that it doesn’t have parasites, I need to take care of that tree.

And if that tree grows and it’s healthy and it’s thriving, it will produce, right? But I’m not—I can’t just expect apples if I’m not taking care of the tree. Now, people are not trees, but you still have to take care of the people if you want them to do things. And if you can’t take care of the people, if you can’t manage the environment that they’re in to make it safe, if you can’t give them the things they need to be successful, then you’re just going to be holding numbers over someone and expecting to hit them.

And that doesn’t work. That’s not something that’s sustainable. And it doesn’t really—it’s not even about how much you pay them. You must pay them well, right, but it has to be more than just that if you want people to succeed. And that doesn’t necessarily mean—like, one thing is at the Duckbill Group I love, succeeding doesn’t necessarily mean that I’m going to stay at—or your engineer is going to stay at one place in perpetuity. If you mentor and train and coach and give an engineer opportunity to grow and thrive and what they do is they go to another job for a title increase and a pay increase or something like that, you did your job.

Corey: A lot of companies love to tell that lie and they almost convince themselves of it where I look at your resume, and great you have not generally crossed the two-year mark at companies for the last decade. I never did until I started at this place. But we magically always liked to pretend in job interviews that, “Oh yeah, this is my forever job—” like you’re a rescue dog getting adopted or something, “—and I’m going to work here for 25 years and get a gold watch and a pension at the end of it.” It’s lunacy. I have never seen the value in lying to ourselves like that, which is why we start our interviews with, “What’s the job after this and how do we help you get there?”

It’s important that we ask those questions and acknowledge that reality. And the downside to it—if you can call it a downside—is you’ve got to live by it. It’s not just words, you can slop onto an interview questionnaire; you actually have to mean it. People can see through insincerity.

Tim: And it’s one of the things, like, if you run an org and you grow your people and you don’t have a place for them to grow into, you should expect and encourage them to find those opportunities elsewhere. It is not reasonable, I feel like, as a leader for you expect people to stay in a place where they have grown past or grown out of. You need to either need to give them a new pot to grow into or you need to let them move elsewhere and thrive and grow. And moving elsewhere—like, if you have a retention problem where you can’t retain anybody, that’s a problem, but if you have your junior engineers who become senior engineers at other places, right, and everyone leaves on good terms, and they got the role and you gave them a great recommendation and they give glowing recommendations to you, there’s nothing wrong with that. That’s not a failure; that’s success.

Corey: One bit of I would say pushback that I suspect you might get when talking to people about what’s next is that, “Well, you are just a consultant, on some level, for a year.” You always know that someone is really arguing in good faith when they describe what you did with the word ‘just,’ but we’ll skip past that part. And it’s, “You’re just a consultant. What would you possibly know about team management and team dynamics?” And there is a little bit of truth to that insofar as the worst place in the world to get management advice is very clearly on Twitter.

It turns out that most interpersonal scenarios are, one, far too personal to wind up tweeting about, and two, do not lend themselves to easy solutions that succinctly fit within 280 characters. Imagine that. The counter-argument though, is that you have—correctly from where I sit—identified a number of recurring dynamics on teams that you have encountered and worked with deeply as a large number of engagements. And these are recurring things, I want to be clear. So, I’m not talking about one particular client. If you’re one of our clients and listen to this thinking that we’re somehow subtweeting you with our voices—I don’t know what that is; subwoofing, maybe?

Tim: [crosstalk 00:12:05]—

Corey: Is that what a subwoofer is? I’m not an audio person.

Tim: Throwing shade, we’ll just say—call it throwing shade.

Corey: Yeah, we’re not throwing shade at any one person, team, or group in particular; these are recurring things. Tim, what have you seen?

Tim: And so, I think the biggest thing I see is folks that are on the precipice of a big technological change, right, and there is an extraordinary amount of anxiety, right? I’ve seen a number of customers through our engagements that, “We are moving away from this legacy platform,” or from this thing that we have been doing for X amount of time. And everyone has staked the other domains, staked out their areas of expertise and control and we’re going to change that. And the solution to that is not a technical solution. You don’t fix that by Helm charts, or Terraform, or CloudFormation. You fix that by conversations, and you fix that by listening. You fix that by finding ways to reassure folks and giving them confidence in their ability to adjust and thrive in a new environment.

If you take somebody who’s been, you know, an Oracle admin for 20 years, and you going to say, “Great. Now, you’re going to learn, you’re going to do this an RDS,” that’s a whole new animal, and folks feel like, well, you know, I can’t learn something new like that? Well, yeah you can. If you can learn Oracle, you can learn anything. I firmly believe that.

But that’s one of the conversations we have, it’s never, almost never a technical problem folks have. We need to reassure people, right? And so, folks who reach out to us, it’s typically folks who are trying to get their organizations in that direction. Another thing we see sometimes is that we find that there’s a disconnect between leadership and the engineers. They have either different priorities or different understandings of what’s going on. And we come in to solve a problem, which may be cost but that’s not the problem we actually solve. The problem we actually solve is fixing this communication bridges between management and leadership.

And that’s almost an every time occurred. At some point or another, there’s some disconnect there. And that’s the best part of the job. Like, the reason I do this consulting gig is not because I want to bang away at code. If I’ve had to do that, that’s an anomaly for sure because I want to have these conversations.

And people want to have these conversations; they want to get these problems solved and sometimes they don’t know how to. And that is the common thing, I think, through all of our customers. Like, we need some amount of expertise to help us find solutions to these things that aren’t necessarily technical problems. And I think that’s where we run into problems as an industry, right, where we think a lot of things are technical problems or have technical solutions, and they don’t. There are people problems. They’re—

Corey: Here at The Duckbill Group, we’re basically marriage counseling for engineering and finance in many cases.

Tim: We really are.

Corey: This is why were people not software.

Tim: Yeah. And I will say this very firmly and you can quote me on this: like, you cannot replace us. You cannot replace the kind of engagements we do with software. You can’t. Can’t be done, right? Software is not empathetic.

Corey: There are a whole series of questions we ask our clients at the start of an engagement and the answers to those questions change what we ask them going forward. In fact, even the level-setting in the conversation that we have at the start of that changes the nature of those. We’re not reading from a list; we’re trying to build an understanding. There is a process around what we do, but it’s not process that can ever be scoped down to the point where it’s just a list of questions or a questionnaire that isn’t maddening for people to fill out because it’s so deeply and clearly misses the mark around context of what they’re actually doing.

Tim: Mm-hm. Our engaged with their conversations. That’s all they are. They’re really in-depth conversations where we’re going to start asking questions and we’re going to ask questions about those answers. We’re start pulling out strings and kicking over rocks and seeing what we find.

And that’s the kind of thing that, you know, you would expect anyone to do who’s coming in and saying, “Okay, we have a problem. Now, let’s figure it out.” Right? Well, you can’t just look at something on the surface, and say, “Oh, I know what this is.” Right? You know, for someone to say, “Oh, I know how to fix this,” when they walk in is the surest way to know that someone doesn’t know what they’re talking about, right?

Corey: Oh, easiest thing in the world is to walk in and say, “This is broken and wrong.” That can translate directly to, “Hi, I am very junior. Please feed my own ass to me.” Because no one shows up at work thinking they’re going to do a crap job today on purpose. There’s a reason things are the way that they are.

Tim: Mm-hm. And that’s the biggest piece of context we get from our customers is we can understand what the best practices are. You can go Google them right now and say, “This is the ten things you’re supposed to do all the time,” right? And we would be really, really crappy consultants if we just read off that list, right? We need to have context: does this thing make sense? Is this the best practice? Maybe, but we want to know why you did it this way.

And after you tell us that way, I’m like, “You know what? I would do it the exact same way for this use case.” And that’s great. We can say like, “This is the best way to do that. Good job.” It’s atypical; it’s unusual, but it solves the problems that you need solving.

And that’s where I think a lot of people miss. Like, you know, you can go—and not to throw shade at AWS’s Trusted Advisor, but we’re going to throw shade at AWS’s Trusted Advisor—and the fact that it will give you—

Corey: It is Plausible Advisor at absolute best.

Tim: [laugh]. It will give you suggestions that have no context. And a lot of the automated AI things that will recommend that you do this and this and this and this are pretty much all the same. And they have no context because they don’t understand what you’re trying to do. And that’s what makes the difference between people. There’s these people problems.

And so, one of the things that I think is really interesting is that we have moved into doing a shorter engagement style that is very short. It’s very quick, it’s very kind of almost tactical, but we go in, we look at your bill, we ask you some questions, and we’re going to give you a list of suggestions that are going to save you a significant amount of money right away, right? So, a lot of times, folks when they need quick wins, or they don’t really need us to deep-dive into all their DynamoDB access patterns, right? They just want like, “Hey, what are the five things we can do to save us some money?” And we’re like, “Well, here they are. And here’s what we think they’re going to save you.” And folks who really enjoyed that type of engagement. And it’s one of my favorite ones to do.

Corey: This episode is sponsored in part by LaunchDarkly. Take a look at what it takes to get your code into production. I’m going to just guess that it’s awful because it’s always awful. No one loves their deployment process. What if launching new features didn’t require you to do a full-on code and possibly infrastructure deploy? What if you could test on a small subset of users and then roll it back immediately if results aren’t what you expect? LaunchDarkly does exactly this. To learn more, visit launchdarkly.com and tell them Corey sent you, and watch for the wince.

Corey: I can also predict that people are going to have questions for you—probably inane—of, well, you were a consultant, how are your actual technical chops? And I love answering these questions with data. So, I have here pulled up the last six months of The Duckbill Group’s AWS bills. And for those who are unaware, every cloud economist has their own dedicated test account for testing out strange things that we come across. And again, can the correct answer in many consulting engagements is, “I don’t know, but I’ll find out.”

Well, this is how we find out. We run tests and learn these things ourselves. I suppose we could extend this benefit, if you want to call it that, to people who aren’t cloud economists but I’m not entirely sure what, I don’t know, an audio engineer is going to do with an AWS account that isn’t, you know, kind of horrifying. To the audio engineer that is editing this podcast, my condolences if you take that as a slight, and if there is something you would use an AWS account for, please let me know. We’ll come talk about it here.

But back to topic, looking at the last six months of your bill for your account—that’s right, a ritualistic shaming of the AWS bill—in January you spent $16.06. In February, you spent 44 cents. And you realized that was too high, so back in March, you then spent 19 cents. And then $3.01 back in April. May wound up $10.02, and now you’re $9.84 as of June. July has not yet finalized as of this recording.

And what I want to highlight—and what that tells me when I look at these types of bills—and I assure you as the world’s leading self-described expert in AWS billing, I’m right; listening to me is a best practice on these things—that shows the exact opposite of a steady-state workload. There’s a lot of dynamism to those giant swings because we don’t have cloud economists who are going to just run these things steady-state for the rest of our lives. Those are experiments of building and testing out new and exciting things in a whole bunch of very weird, very strange ways. Whenever I wind up talking to someone in one of the overarching AWS services at AWS and I pull up my account, a common refrain is, “Wow, you use an awful lot of services.” Right. I’m not just sitting here run and EC2 instances forever. Imagine that. And your account is a perfect microcosm of that entire philosophy.

Tim: Well, I don’t know all the answers, right? And I will never profess all the answers. And before I say, “You should do this—” or maybe I will say, “You might be able to do this. Let me go save as possible.” [laugh]. Right? And so, just let me just see, can you do this? Does this work? No, I guess it doesn’t. Or AWS docu—especially, “The AWS documentation says this. Let me see if that’s actually the case.”

Corey: I don’t believe that they intend to lie, but—

Tim: No.

Corey: —they also certainly don’t get it correct all the time.

Tim: And to be fair, they have, what, 728 services by this point, and that’s a lot of documentation you’re not going to get—

Corey: Three more have launched since the start of this recording.

Tim: I—yeah, actually—well, by the time this hits, they’re probably going to have 22. But we’ll [laugh] see. But yeah, no. And that’s fine. And they’re not going to have every use case, and every edge, kind of like, concern handled, and so that’s why we need to kick the tires a little bit.

And what I think more than anything else is, you know, sometimes we just do things out of convenience. Like, “Well, I don’t want to run this on this; let me just fire it up because it’s not my money.” [laugh]. But we also want to be fairly concerned about you know, how we do things. You don’t want to run a fleet of z1ds, obviously.

But there is a certain amount of tire-kicking and infrastructure spinning up that you have to do in order to maintain freshness, right? And it’s not a thing where I’m going to say, “Oh, I know YAML off the top of my head, and I need to do—you know, I’m up to speed on every single possible API call that you can make.” No. My technical prowess has always been in architecture and operations. So, I think when we have these conversations, folks mostly tend to be impressed by not only business acumen and strategy, but also being able to get down to the weeds and talking with the developers and the engineers about the minutia. And you will have seen you know, the feedback that I’ve gotten about my technical prowess has always been good. You know, I can hang with anybody, I feel like.

Corey: I would agree wholeheartedly. It’s been really interesting watching you in conversations, internally and with our clients, where you will just idly bust out something fricking brilliant out of left field. And most of the time, I don’t think you even realize it. It’s just one of those things that makes intuitive and instinctive sense to you. And you basically just leave people stunned and their scribbling notes and trying to wrap their heads around what you just said.

And it’s adorable because sometimes you wind up almost, like, looking embarrassed, like, “Did I say something rude and not realize it? Like, I wasn’t trying to be insulting.” It’s like, “Nope, nope. You’re just doing your thing, Tim. Just keep on doing it. That’s fine.”

Tim: Yeah, it’s funny because, like you, one of the things that I’ve really enjoyed about it is, like, we’ll just start bouncing ideas off of each other and come up with something brilliant. “Yeah, let’s do that.” And then, “Okay, this is now a thing.” And it’s like, you know, there’s something to be said about being around smart people. So, it’s not just me coming up with something brilliant; these are almost always fruits of a conversation and discussion being had, and then you formulate something great in your head.

But again, this is why I love the aspect of talking and having conversations with people, so that way you can come up with something kind of brilliant. None of this is done in a silo. Like we’re not really, really good at what we do because we don’t rely or talk to or have conversations with other people.

Corey: One thing that you did that I think is one of the most transformative things that has happened in company history in some respects has been when you started, and for the first half of your tenure here, we had two engagement types that we would wind up giving our consulting clients. There’s contract negotiation, where we help companies negotiate their long-term commitment contracts with AWS—and we’re effective at it and that’s fun; that’s basically what you would more or less expected to be—and the other is our cost optimization project engagements. And those tend to look six to eight weeks where we wind up going in deep-dives into the intricacies of an organization’s AWS accounts, bills, strategy, growth plan, et cetera, et cetera, et cetera, to an exhaustive level of detail. And in an interest of being probably overly transparent here, I didn’t like working on those engagements myself. I like coming in, finding the big things that will be transformative to reduce the bills—it’s like solving a puzzle—and then the relatively in-depth analysis for things that are a relatively paltry portion of the AWS bill does not really lead me to enjoying the work very much.

And I beat my head against that one for years. And you busted out one day with an idea that became our third type of engagement, which is the first pass, where we charge significantly less for the engagement and it essentially distills down into you get us to talk to your engineering teams for a day. Bring us any questions, give us access in advance to these things, and we will basically go on a whirlwind guided tour and lay waste to your AWS bill and highlight different opportunities that we see to optimize these things. And it has been an absolute smash success. People love the engagements.

Very often, it leads to that second full-bore engagement that I was describing earlier, but it also aligns very well with the way that I like to think about these things. I’m a great consultant, specifically because once I’ve delivered the value, I like to leave. Whereas as an employee, I just sort of linger around, and then I go cause problems and other people’s departments—ideally, not on purpose, but you know, I am me—and this really emphasizes that and keeps me moving quickly. I really, really like that engagement style and I have you to thank for coming up with the idea and finding a way to do it that didn’t either not resonate with the market—in which case, we’re not selling a damn thing—or wound up completely eviscerating the value of the longer-term deep-dive engagements, and you threaded that needle perfectly.

Tim: I thank you; I appreciate that. There was this kind of vacuum that I saw where, both from a cost and from a resource point where six to eight weeks is a long time for an engineering org to dedicate to any one thing, especially if that one thing isn’t directly making money. But engineering orgs are also very interested in saving money. But it’s especially in smaller orgs where that velocity is very important, they don’t have six to eight weeks for that. They can’t dedicate the resources to those deep-dives all the time, and all the conversations we—and when we do a COP, it is exhaustive. We are exploring every avenue to almost an absurd level, right?

And that’s not the right engagement for a lot of orgs, right? So, coming in and saying, “Hey, you know, this is a quick one; these are the things that you can do. This is 90% of the savings you’re going to realize. These things: bam, bam, bam, bam, bam.” Right?

And then we give it to the folks and we let them work on it, and then they’re like, “Hey, we need this because we want to negotiate EDP,” or, “We need this because, you know, we’re just trying to make sure that our costs are in line so we can be more agile, so we can do this project, or whatever.” Right? And then there are a lot of other orgs that do need that exhaustive kind of thing, larger orgs especially, right? Larger, more complex orgs, orgs that are trying to maybe—like, if you’re trying to make a play to get acquired, you want to get this very, very in-depth study so you know all your liabilities and all your assets, so that way you can fix those problems and make it very attractive for someone to buy you, right? Or orgs that just have, like, we are not having an impending EDP; we have a lot of time to be able to focus on these things, and we can build this into the roadmap, right?

Then we can do a very exhaustive study of those things. But for a lot of times, people are just like, “Look, I just need to save X amount of money on my AWS bill and can you do that?” Well, sure. We can go in there and have those conversations and give you a lot of savings. And I’m very much in the camp of, you know, ‘perfect is the enemy of good.’ I don’t have to save down to the nth penny on your DynamoDB bill. But if I can, shave—cut it in half, that’s great. Most people are very happy about those kinds of things. And that’s a very routine finding for us.

Corey: One other aspect that I really liked about it, too, is that it let us move down market a bit, away from companies that are spending millions of dollars a month. Because yeah, the ROI for those customers is a slam dunk on virtually any engagement that we could put together, but what about the smaller companies, the ones that are not spending that much money, yet? They’ve never felt great talk to them and say, “Oh, just go screw up your AWS bill some more. Then, then you will absolutely be able to generate some value. Maybe turn off MFA and post your credentials to GitHub or something. That’ll speed up the process nicely.”

That’s terrible advice and we can’t do it. But this enables us to move down to smaller companies that are earlier in their cloud estate build-out or are growing organically rather than trying to do a giant migration as sort of greenfield growth approach. I really, really like our ability to help companies that are a bit earlier in their cloud journey, as well as in smaller environments, just because I guess, on some level, for me, at least, when you see enormous multimillion-dollar levels of spend, the misconfigurations are generally less fun to find; they’re less exciting. Because, yeah at a small scale, you can screw up and your Managed NAT Gateway bill is a third of your spend. When you’re spending $80 million a year, you’re not wasting that kind of money on Managed NAT Gateways because that misconfiguration becomes visible from frickin’ orbit.

So, someone has already found that stuff. And it’s always then it’s almost certainly EC2, RDS, and storage. Great. Then there’s some weird data transfer stuff and it starts to look a lot more identical. Smaller accounts, at least from my perspective, tend to have a lot more of interesting things to learn hiding in the shadows.

Tim: Oh, absolutely. And I think the impact that you make for the future for small companies much higher, right? You go in there and you have an engagement, you can say, “Okay, I understand the business reason why you did this here, but if you make these changes—bam, bam, bam—12 to 18 months and on, right, this is going to make a huge difference in your business. You’re going to save a tremendous amount of money and you’re going to be much more agile.”

You did this thing because it worked for the POC, it worked for the MVP, right? That’s great, but before it gets too big and becomes load-bearing technical debt, let’s make some changes to put you in a better position, both for cost optimization and an architectural future that you don’t have to then break a bone that’s already set to try and fix it. So, getting in there before there’s a tremendous load on their architecture—or rather on their infrastructure, it’s super, super fun because you know that when you’ve done this, you have given that company more runway, or you’ve given them the things they need to actually be more successful, and so they can focus their time and efforts on growth and not on trying to stop the bleeding with their AWS bill.

Corey: Tim, it’s been an absolute pleasure to work with you. I’m going to miss working with you, but we are definitely going to remain in touch. Where can people find you to follow along with your continuing adventures?

Tim: The best way to find me is on Twitter, I am @elchefe—E-L-C-H-E-F-E. And yeah, I will definitely keep in touch with you, Corey. Again, you have been a tremendous friend and I really appreciate you, your insights, and your honesty. Our partners are friends with each other and I do not think that they will let us ever drift too far apart. So.

Corey: No, I think it is pretty clear that we are basically going to be both of their plus-ones forever.

Tim: [laugh]. I think so.

Corey: I’m just waiting for them when they pulled the prank of dressing us the exact same way because our styles are somewhat different, and I’m pretty sure that there’s not a whole lot of convergence where we both wind up looking great. So, it’s going to be hilarious regardless of what direction it goes in.

Tim: Well, you do have velour tracksuits too, right?

Corey: Not yet, but please don’t tell that to Bethany.

Tim: [laugh].

Corey: Tim, it has been an absolute pleasure.

Tim: The pleasure has been all mine, Corey. I really appreciate it.

Corey: Tim Banks, for one last time, principal cloud economist at The Duckbill Group. I am Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice and an insulting comment that says that we are completely wrong in our approach to management and the real answer is as follows, making sure to keep that answer less than 280 characters.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Anton

Dr. Anton Chuvakin is now involved with security solution strategy at Google Cloud, where he arrived via Chronicle Security (an Alphabet company) acquisition in July 2019.

Anton was, until recently, a Research Vice President and Distinguished Analyst at Gartner for Technical Professionals (GTP) Security and Risk Management Strategies team. (see chuvakin.org for more)

Links Referenced:

  • Google Cloud: https://cloud.google.com/
  • Cloud Security Podcast: https://cloud.withgoogle.com/cloudsecurity/podcast/
  • Twitter: https://twitter.com/anton_chuvakin
  • Medium blog: https://medium.com/anton.chuvakin

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friend EnterpriseDB. EnterpriseDB has been powering enterprise applications with PostgreSQL for 15 years. And now EnterpriseDB has you covered wherever you deploy PostgreSQL on-premises, private cloud, and they just announced a fully-managed service on AWS and Azure called BigAnimal, all one word. Don’t leave managing your database to your cloud vendor because they’re too busy launching another half-dozen managed databases to focus on any one of them that they didn’t build themselves. Instead, work with the experts over at EnterpriseDB. They can save you time and money, they can even help you migrate legacy applications—including Oracle—to the cloud. To learn more, try BigAnimal for free. Go to biganimal.com/snark, and tell them Corey sent you.

Corey: Let’s face it, on-call firefighting at 2am is stressful! So there’s good news and there’s bad news. The bad news is that you probably can’t prevent incidents from happening, but the good news is that incident.io makes incidents less stressful and a lot more valuable. incident.io is a Slack-native incident management platform that allows you to automate incident processes, focus on fixing the issues and learn from incident insights to improve site reliability and fix your vulnerabilities. Try incident.io, recover faster and sleep more.

Corey: Welcome to Screaming in the Cloud, I’m Corey Quinn. My guest today is Anton Chuvakin, who is a Security Strategy Something at Google Cloud. And I absolutely love the title, given, honestly, how anti-corporate it is in so many different ways. Anton, first, thank you for joining me.

Anton: Sure. Thanks for inviting me.

Corey: So, you wound up working somewhere else—according to LinkedIn—for two months, which in LinkedIn time is about 20 minutes because their date math is always weird. And then you wound up going—according to LinkedIn, of course—leaving and going to Google. Now, that was an acquisition if I’m not mistaken, correct?

Anton: That’s correct, yes. And it kind of explains that timing in a little bit of a title story because my original title was Head of Security Solution Strategy, and it was for a startup called Chronicle. And within actually three weeks, if I recall correctly, I was acquired into Google. So, title really made little sense of Google, so I kind of go with, like, random titles that include the word security, and occasionally strategy if I feel generous.

Corey: It’s pretty clear, the fastest way to get hired at Google, given their famous interview process is to just get acquired. Like, “I’m going to start a company and raise it to, like, a little bit of providence, and then do an acquihire because that will be faster than going through the loop, and ideally, there will be less algorithm solving on whiteboards.” But I have to ask, did you have to solve algorithms on whiteboards for your role?

Anton: Actually, no, but it did come close to that for some other people who were seen as non-technical and had to join technical roles. I think they were forced to solve coding questions and stuff, but I was somehow grandfathered into a technical role. I don’t know exactly how it happened.

Corey: Yeah, how you wound up in a technical role. Let’s be clear, you are Doctor Anton Chuvakin, and you have written multiple books, you were a research VP at Gartner for many years, and once upon a time, that was sort of a punchline in the circles I hung out with, and then I figured out what Gartner actually does. And okay, that actually is something fairly impressive, let’s be clear here. Even as someone who categorically defines himself as not an analyst, I find myself increasingly having a lot of respect for the folks who are actually analysts and the laborious amount of work that they do that remarkably few people understand.

Anton: That’s correct. And I don’t want to boost my ego too much. It’s kind of big enough already, obviously, but I actually made it all the way to Distinguished Analyst, which is the next rank after VP.

Corey: Ah, my apologies. I did not realize it. This [challenges 00:02:53] the internal structure.

Anton: [laugh]. Yeah.

Corey: It’s like, “Oh, I went from Senior to Staff,” or Staff to Senior because I’m external; I don’t know the direction these things go in. It almost feels like a half-step away from oh, I went from [SDE3 to SDE4 00:03:02]. It’s like, what do those things mean? Nobody knows. Great.

Anton: And what’s the top? Is it 17 or is it 113? [laugh].

Corey: Exactly. It’s like, oh okay, so you’re Research VP—or various kinds of VPs—the real question is, how many people have to die before you’re the president? And it turns out that that’s not how companies think. Who knew?

Anton: That’s correct. And I think Gartner was a lot of hard work. And it’s the type of work that a lot of people actually don’t understand. Some people understand it wrong, and some people understand it wrong, kind of, for corrupt reasons. So, for example, a lot of Gartner machinery involves soaking insight from the outside world, organizing it, packaging it, writing it, and then giving it as advice to other people.

So, there’s nothing offensive about that because there is a lot of insight in the outside world, and somebody needs to be a sponge slash filter slash enrichment facility for that insight. And that, to me, is a good analyst firm, like Gartner.

Corey: Yeah. It’s a very interesting world. But you historically have been doing a lot of, well, let’s I don’t even know how to properly describe it because Gardner’s clientele historically has not been startups because let’s face it, Gartner is relatively expensive. And let’s be clear, you’re at Google Cloud now, which is a different kind of expensive, but in a way that works for startups, so good for you; gold star. But what was interesting there is that the majority of the Gartner clientele that I’ve spoken to tend to be big-E Enterprise, which runs legacy businesses, which is a condescending engineering term for ‘it makes money.’

And they had the temerity to start their company before 15 years ago, so they built data centers and did things in a data center environment, and now they’re moving in a cloudy direction. Your emphasis has always been on security, so my question for you to start with all this is where do you see security vendors fitting in? Because when I walk the RSA expo hall and find myself growing increasingly depressed, it seems like an awful lot of what vendors are selling looks very little removed from, “We took a box, now we shoved in a virtual machine and here you go; it’s in your cloud environment. Please pay us money.” The end. And it feels, if I’m looking at this from a pure cloud-native, how I would build things in the cloud from scratch perspective, to be the wrong design. Where do you stand on it?

Anton: So, this has been one of the agonizing questions. So, I’m going to kind of ignore some of the context. Of course, I’ll come back to it later, but want to kind of frame it—

Corey: I love ignoring context. My favorite thing; it’s what makes me a decent engineer some days.

Anton: So, the frame was this. One of the more agonizing questions for me as an analyst was, a client calls me and says, “We want to do X.” Deep in my heart, I know that X is absolutely wrong, however given their circumstances and how they got to decided to do X, X is perhaps the only thing they can logically do. So, do you tell them, “Don’t do X; X is bad,” or you tell them, “Here’s how you do X in a manner that aligns with your goals, that’s possible, that’s whatever.”

So, cloud comes up a lot in this case. Somebody comes and says, I want to put my on-premise security information management tool or SIM in the cloud. And I say, deep in my heart, I say, “No, get cloud-native tool.” But I tell them, “Okay, actually, here’s how you do it in a less painful manner.” So, this is always hard. Do you tell them they’re on their own path, but you help them tread their own path with least pain? So, as an analyst, I agonized over that. This was almost like a moral decision. What do I tell them?

Corey: It makes sense. It’s a microcosm of the architect’s dilemma, on some level, because if you ask a typical Google-style interview whiteboard question, one of my favorites in years past was ‘build a URL shortener.’ Great. And you can scale it out and turn it into different things and design things on the whiteboard, and that’s great. Most mid-level people can wind up building a passable designed for most things in a cloud sense, when you’re starting from scratch.

That’s not hard. The problem is that the real world is messy and doesn’t fit on a whiteboard. And when you’re talking about taking a thing that exists in a certain state—for whatever reason, that’s the state that it’s in—and migrating it to a new environment or a new way of operating, there are so many assumptions that have to break, and in most cases, you don’t get the luxury of just taking the thing down for 18 months so you can rework it. And even that it’s never as easy as people think it is, so it’s going to be 36. Great.

You have to wind up meeting people where they are as they’re contextualizing these things. And I always feel like the first step of the cloud migration has been to improve your data center environment at the cost of worsening your cloud environment. And that’s okay. We don’t all need to be the absolute vanguard of how everything should be built and pushing the bleeding edge. You’re an insurance company, for God’s sake. Maybe that’s not where you want to spend your innovation energies.

Anton: Yeah. And that’s why I tend to lean towards helping them get out of this situation, or maybe build a five-step roadmap of how to become a little bit more cloud-native, rather than tell them, “You’re wrong. You should just rewrite the app in a cloud-native way.” That advice almost never actually works in real world. So, I see a lot of the security people move their security stacks to the cloud.

And if I see this, I deepen my heart and say, “Holy cow. What do you mean, you want to IDS every packet between Cloud instances? You want to capture every packet in cloud instances? Why? It’s all encrypted anyway.” But I don’t say that. I say, “Okay, I see how this is the first step for you. Let’s describe the next seven steps.”

Corey: The problem I keep smacking into is that very often folks who are pushing a lot of these solutions are, yes, they’re meeting customers where they are, and that makes an awful lot of sense; I’m not saying that there’s anything inherently wrong about that. The challenge is it also feels on the high end, when those customers start to evolve and transform, that those vendors act as a drag. Because if you wind up going in a full-on cloud-native approach, in the fullness of time, there’s an entire swath of security vendors that do not have anything left to sell you.

Anton: Yes, that is correct. And I think that—I had a fight with an EDR vendor, Endpoint Detection Response, vendor one day when they said, “Oh, we’re going to be XDR and we’ll do cloud.” And I told them, “You do realize that in a true cloud-native environment, there’s no E? There is no endpoint the way you understand it? There is no OS. There is no server. And 99% of your IP isn’t working on the clients and servers. How are you going to secure a cloud again?”

And I get some kind of rambling answer from them, but the point is that you’re right, I do see a lot of vendors that meet clients where they are during their first step in the cloud, and then they may become a drag, or the customer has to show switch to a cloud-native vendor, or to both sometimes, and pay into two mouths. Well, shove money into two pockets.

Corey: Well, first, I just want to interject for a second here because when I was walking the RSA expo floor, there were something like 15 different vendors that were trying to sell me XDR. Not a single one of them bothered to expand the acronym—

Anton: Just 15? You missed half of them.

Corey: Well, yeah—

Anton: Holy cow.

Corey: As far as I know XDR cable. It’s an audio thing right? I already have a bunch of those for my microphone. What’s the deal here? Like, “I believe that’s XLR.” It’s like, “I believe you should expand your acronyms.” What is XDR?

Anton: So, this is where I’m going to be very self-serving and point to a blog that I’ve written that says we don’t know what’s XDR. And I’m going to—

Corey: Well, but rather than a spiritual meaning, I’m going to ask, what does the acronym stands for? I don’t actually know the answer to that.

Anton: Extended Detection and Response.

Corey: Ah.

Anton: Extended Detection and Response. But the word ‘extended’ is extended by everybody in different directions. There are multiple camps of opinion. Gartner argues with Forrester. If they ever had a pillow fight, it would look really ugly because they just don’t agree on what XDR is.

Many vendors don’t agree with many other vendors, so at this point, if you corner me and say, “Anton, commit to a definition of XDR,” I would not. I will just say, “TBD. Wait two years.” We don’t have a consensus definition of XDR at this point. And RSA notwithstanding, 30 booths with XDRs on their big signs… still, sorry, I don’t have it.

Corey: The problem that I keep running into again and again and again, has been pretty consistently that there are vendors willing to help customers in a very certain position, and for those customers, those vendors are spot on the right thing to do.

Anton: Mmm, yep.

Corey: But then they tried to expand and instead of realizing that the market has moved on and the market that they’re serving is inherently limited and long-term is going to be in decline, they instead start trying to fight the tide and saying, “Oh, no, no, no, no. Those new cloud things, can’t trust them.” And they start out with the FU, the Fear, Uncertainty, and Doubt marketing model where, “You can’t trust those newfangled cloud things. You should have everything on-prem,” ignoring entirely the fact that in their existing data centers, half the time the security team forgets to lock the door.

Anton: Yeah, yeah.

Corey: It just feels like there is so much conflict of interest about in the space. I mean, that’s the reason I started my Thursday Last Week in AWS newsletter that does security round-ups, just because everything else I found was largely either community-driven where it understood that it was an InfoSec community thing—and InfoSec community is generally toxic—or it was vendor-captured. And I wanted a round-up of things that I had to care about running an infrastructure, but security is not in my job title, even if the word something is or is not there. It’s—I have a job to do that isn’t security full time; what do I need to know? And that felt like an underserved market, and I feel like there’s no equivalent of that in the world of the emerging cloud security space.

Anton: Yes, I think so. But it has a high chance of being also kind of captured by legacy vendors. So, when I was at Gartner, there was a lot of acronyms being made with that started with a C: Cloud. There was CSPM, there was CWBP, and after I left the coined SNAPP with double p at the end. Cloud-Native Application Protection Platform. And you know, in my time at Gartner, five-letter acronyms are definitely not very popular. Like, you shouldn’t have done a five-letter acronym if you can help yourself.

So, my point is that a lot of these vendors are more from legacy vendors. They are not born in the cloud. They are born in the 1990s. Some are born in the cloud, but it’s a mix. So, the same acronym may apply to a vendor that’s 2019, or—wait for it—1989.

Corey: That is… well, I’d say on the one hand, it’s terrifying, but on the other, it’s not that far removed from the founding of Google.

Anton: True, true. Well, ’89, kind of, it’s another ten years. I think that if you’re from the ’90s, maybe you’re okay, but if you’re from the ’80s… you really need to have superpowers of adaptation. Again, it’s possible. Funny aside: at Gartner, I met somebody who was an analyst for 32 years.

So, he was I think, at Gartner for 32 years. And how do you keep your knowledge current if you are always in an ivory tower? The point is that this person did do that because he had a unique ability to absorb knowledge from the outside world. You can adapt; it’s just hard.

Corey: It always is. I’m going to pivot a bit and put you in a little bit of a hot seat here. Not intentionally so. But it is something that I’ve been really kicking around for a while. And I’m going to basically focus on Google because that’s where you work.

I yeah, I want you to go and mouth off about other cloud companies. Yeah, that’s—

Anton: [laugh]. No.

Corey: Going to go super well and no one will have a problem with that. No, it’s… we’ll pick on Google for a minute because Google Cloud offers a whole bunch of services. I think it’s directionally the right number of services because there are areas that you folks do not view as a core competency, and you actually—imagine that—partner with third parties to wind up delivering something great rather than building this shitty knockoff version that no one actually wants. Ehem, I might be some subtweeting someone here with this, only out loud.

Anton: [laugh].

Corey: The thing that resonates with me though, is that you do charge for a variety of security services. My perspective, by and large, is that the cloud vendors should not be viewing security as a profit center but rather is something that comes baked into the platform that winds up being amortized into the cost of everything else, just because otherwise you wind up with such a perverse set of incentives.

Anton: Mm-hm.

Corey: Does that sound ridiculous or is that something that aligns with your way of thinking. I’m willing to take criticism that I’m wrong on this, too.

Anton: Yeah. It’s not that. It’s I almost start to see some kind of a magic quadrant in my mind that kind of categorizes some things—

Corey: Careful, that’s trademarked.

Anton: Uhh, okay. So, some kind of vis—

Corey: It’s a mystical quadrilateral.

Anton: Some kind of visual depiction, perhaps including four parts—not quadrants, mind you—that is focused on things that should be paid and aren’t, things that should be paid and are paid, and whatever else. So, the point is that if you’re charging for encryption, like basic encryption, you’re probably making a mistake. And we don’t, and other people, I think, don’t as well. If you’re charging for logging, then it’s probably also wrong—because charging for log retention, keeping logs perhaps is okay because ultimately you’re spending resources on this—charging for logging to me is kind of in the vile territory. But how about charging for a tool that helps you secure your on-premise environment? That’s fair game, right?

Corey: Right. If it’s something you’re taking to another provider, I think that’s absolutely fair. But the idea—and again, I’m okay with the reality of, “Okay, here’s our object storage costs for things, and by the way, when you wind up logging things, yeah, we’ll charge you directionally what it costs to store that an object store,” that’s great, but I don’t have the Google Cloud price list shoved into my head, but I know over an AWS land that CloudWatch logs charge 50 cents per gigabyte, for ingress. And the defense is, “Well, that’s a lot less expensive than most other logging vendors out there.” It’s, yeah, but it’s still horrifying, and at scale, it makes me want to do some terrifying things like I used to, which is build out a cluster of Rsyslog boxes and wind up having everything logged to those because I don’t have an unbounded growth problem.

This gets worse with audit logs because there’s no alternative available for this. And when companies start charging for that, either on a data plane or a management plane level, that starts to get really, really murky because you can get visibility into what happened and reconstruct things after the fact, but only if you pay. And that bugs me.

Anton: That would bug me as well. And I think these are things that I would very clearly push into the box of this is security that you should not charge for. But authentication is free. But, like, deeper analysis of authentication patterns, perhaps costs money. This to me is in the fair game territory because you may have logs, you may have reports, but what if you want some kind of fancy ML that analyzes the logs and gives you some insights? I don’t think that’s offensive to charge for that.

Corey: I come bearing ill tidings. Developers are responsible for more than ever these days. Not just the code that they write, but also the containers and the cloud infrastructure that their apps run on. Because serverless means it’s still somebody’s problem. And a big part of that responsibility is app security from code to cloud. And that’s where our friend Snyk comes in. Snyk is a frictionless security platform that meets developers where they are - Finding and fixing vulnerabilities right from the CLI, IDEs, Repos, and Pipelines. Snyk integrates seamlessly with AWS offerings like code pipeline, EKS, ECR, and more! As well as things you’re actually likely to be using. Deploy on AWS, secure with Snyk. Learn more at Snyk.co/scream That’s S-N-Y-K.co/scream

Corey: I think it comes down to what you’re doing with it. Like, the baseline primitives, the things that no one else is going to be in a position to do because honestly, if I can get logging and audit data out of your control plane, you have a different kind of security problem, and—

Anton: [laugh].

Corey: That is a giant screaming fire in the building, as it should be. The other side of it, though, is that if we take a look at how much all of this stuff can cost, and if you start charging for things that are competitive to other log analytics tools, great because at that point, we’re talking about options. I mean, I’d like to see, in an ideal world, that you don’t charge massive amounts of money for egress but ingress is free. I’d like to see that normalized a bit.

But yeah, okay, great. Here’s the data; now I can run whatever analytics tools I want on it and then you’re effectively competing on a level playing field, as opposed to, like, okay, this other analytics tool is better, but it’ll cost me over ten times as much to migrate to it, so is it ten times better? Probably not; few things are, so I guess I’m sticking with the stuff that you’re offering. It feels like the cloud provider security tools never quite hit the same sweet spot that third-party vendors tend to as far as usability, being able to display things in a way that aligns with various stakeholders at those companies. But it still feels like a cash grab and I have to imagine without having insight into internal costing structures, that the security services themselves are not a significant revenue driver for any of the cloud companies. And the rare times where they are is almost certainly some horrifying misconfiguration that should be fixed.

Anton: That’s fair, but so to me, it still fits into the bucket of some things you shouldn’t charge for and most people don’t. There is a bucket of things that you should not charge for, but some people do. And there’s a bucket of things where it’s absolutely fair to charge for I don’t know the amount I’m not a pricing person, but I also seen things that are very clearly have cost to a provider, have value to a client, have margins, so it’s very clear it’s a product; it’s not just a feature of the cloud to be more secure. But you’re right if somebody positions as, “I got cloud. Hey, give me secure cloud. It costs double.” I’d be really offended because, like, what is your first cloud is, like, broken and insecure? Yeah. Replace insecure with broken. Why are you selling broken to me?

Corey: Right. You tried to spin up a service in Google Cloud, it’s like, “Great. Do you want the secure version or the shitty one?”

Anton: Yeah, exactly.

Corey: Guess which one of those costs more. It’s… yeah, in the fullness of time, of course, the shitty one cost more because you find out about security breaches on the front page of The New York Times, and no one’s happy, except maybe The Times. But the problem that you hit is that I don’t know how to fix that. I think there’s an opportunity there for some provider—any provider, please—to be a trendsetter, and, “Yeah, we don’t charge for security services on our own stuff just because it’d be believed that should be something that is baked in.” Like, that becomes the narrative of the secure cloud.

Anton: What about tiers? What about some kind of a good, better, best, or bronze, gold, platinum, where you have reasonable security, but if you want superior security, you pay money? How do you feel, what’s your gut feel on this approach? Like, I can’t think of example—log analysis. You’re going to get some analytics and you’re going to get fancy ML. Fancy ML costs money; yay, nay?

Corey: You’re bringing up an actually really interesting point because I think I’m conflating too many personas at once. Right now, just pulling up last months bill on Google Cloud, it fits in the free tier, but my Cloud Run bill was 13 cents for the month because that’s what runs my snark.cloud URL shortener. And it’s great. And I wound up with—I think my virtual machine costs dozen times that much. I don’t care.

Over in AWS-land, I was building out a serverless nonsense thing, my Last Tweet In AWS client, and that cost a few pennies a month all told, plus a whopping 50 cents for a DNS zone. Whatever. But because I was deploying it to all regions and the way that configural evaluations work, my config bill for that was 16 bucks. Now, I don’t actually care about the dollar figures on this. I assure you, you could put zeros on the end of that for days and it doesn’t really move the needle on my business until you get to a very certain number there, and then suddenly, I care a lot.

Anton: [laugh]. Yeah.

Corey: And large enterprises, this is expected because even the sheer cost of people’s time to go through these things is valuable. What I’m thinking of is almost a hobby-level side project instead, where I’m a student, and I’m learning this in a dorm room or in a bootcamp or in my off hours, or I’m a career switcher and I’m doing this on my own dime out of hours. And I wind up getting smacked with the bill for security services that, for a company, don’t even slightly matter. But for me, they matter, so I’m not going to enable them. And when I transition into the workforce and go somewhere, I’m going to continue to work the same way that I did when I was an independent learner, like, having a wildly generous free tier for small-scale accounts, like, even taking a perspective until you wind up costing, I don’t know, five, ten—whatever it is—thousand dollars a month, none of the security stuff is going to be billable for you because it’s it is not aimed at you and we want you comfortable with and using these things.

This is a whole deep dive into the weeds of economics and price-driven behavior and all kinds of other nonsense, but every time I wind up seeing that, like, in my actual production account over at AWS land for The Duckbill Group, all things wrapped up, it’s something like 1100 bucks a month. And over a third of it is monitoring, audit, and observability services, and a few security things as well. And on the one hand, I’m sitting here going, “I don’t see that kind of value coming from it.” Now, the day there’s an incident and I have to look into this, yeah, it’s absolutely going to be worth having, but it’s insurance. But it feels like a disproportionate percentage of it. And maybe I’m just sitting here whining and grousing and I sound like a freeloader who doesn’t want to pay for things, but it’s one of those areas where I would gladly pay more for a just having this be part of the cost and not complain at all about it.

Anton: Well, if somebody sells me a thing that costs $1, and then they say, “Want to make it secure?” I say yes, but I’m already suspicious, and they say, “Then it’s going to be 16 bucks.” I’d really freak out because, like, there are certain percentages, certain ratios of the actual thing plus security or a secure version of it; 16x is not the answer expect. 30%, probably still not the answer I expect, frankly. I don’t know. This is, like, an ROI question [crosstalk 00:23:46]—

Corey: Let’s also be clear; my usage pattern is really weird. You take a look at most large companies at significant scale, their cloud environments from a billing perspective look an awful lot like a crap ton of instances—or possibly containers running—and smattering of other things. Yeah, you also database and storage being the other two tiers and because of… reasons data transfer loves to show up too, but by and large, everything else was more or less a rounding error. I have remarkably few of those things, just given the weird way that I use services inappropriately, but that is the nature of me, so don’t necessarily take that as being gospel. Like, “Oh, you’ll spend a third of your bill.”

Like, I’ve talked to analyst types previously—not you, of course—who will hear a story like this and that suddenly winds up as a headline in some report somewhere. And it’s, “Yeah, if your entire compute is based on Lambda functions and you get no traffic, yeah, you’re going to see some weird distortions in your bill. Welcome to the conversation.” But it’s a problem that I think is going to have to be addressed at some point, especially we talked about earlier, those vendors who are catering to customers who are not born in the cloud, and they start to see their business erode as the cloud-native way of doing things continues to accelerate, I feel like we’re in for a time where they’re going to be coming at the cloud providers and smacking them for this way harder than I am with my, “As a customer, wouldn’t it be nice to have this?” They’re going to turn this into something monstrous. And that’s what it takes, that’s what it takes. But… yeah.

Anton: It will take more time than than we think, I think because again, back in the Gartner days, I loved to make predictions. And sometimes—I’ve learned that predictions end up coming true if you’re good, but much later.

Corey: I’m learning that myself. I’m about two years away from the end of it because three years ago, I said five years from now, nobody will care about Kubernetes. And I didn’t mean it was going to go away, but I meant that it would slip below the surface level of awareness to point where most people didn’t have to think about it in the same way. And I know it’s going to happen because it’s too complex now and it’s going to be something that just gets handled in the same way that Linux kernels do today, but I think I was aggressive on the timeline. And to be clear, I’ve been misquoted as, “Oh, I don’t think Kubernetes is going to be relevant.”

It is, it’s just going to not be something that you need to spend the quarter million bucks an engineer on to run in production safely.

Anton: Yeah.

Corey: So, we’ll see. I’m curious. One other question I had for you while I’ve got you here is you run a podcast of your own: the Cloud Security Podcast if I’m not mistaken, which is—

Anton: Sadly, you are not. [laugh].

Corey: —the Cloud Se—yeah. Interesting name on that one, yeah. It’s like what the Cloud Podcast was taken?

Anton: Essentially, we had a really cool name [Weather Insecurity 00:26:14]. But the naming team here said, you must be descriptive as everybody else at Google, and we ended up with the name, Cloud Security Podcast. Very, very original.

Corey: Naming is challenging. I still maintain that the company is renamed Alphabet, just so it could appear before Amazon in the yellow pages, but I don’t know how accurate that one actually is. Yeah, to be clear, I’m not dunking on your personal fun podcast, for those without context. This is a corporate Google Cloud podcast and if you want to make the argument that I’m punching down by making fun of Google, please, I welcome that debate.

Anton: [laugh]. Yes.

Corey: I can’t acquire companies as a shortcut to hire people. Yet. I’m sure it’ll happen someday, but I can aspire to that level of budgetary control. So, what are you up to these days? You spent seven years at Gartner and now you’re doing a lot of cloud security… I’ll call it storytelling, and I want to be clear that I mean that as a compliment, not the, “Oh, you just tell stories rather than build things?”

Anton: [laugh].

Corey: Yeah, it turns out that you have to give people a reason to care about what you’ve built or you don’t have your job for very long. What are you talking about these days? What narratives are you looking at going forward?

Anton: So, one of the things that I’ve been obsessed with lately is a lot of people from more traditional companies come in in the cloud with their traditional on-premise knowledge, and they’re trying to do cloud the on-premise way. On our podcast, we do dedicate quite some airtime to people who do cloud as if it were a rented data center, and sometimes we say, the opposite is called—we don’t say cloud-native, I think; we say you’re doing the cloud the cloudy way. So, if you do cloud, the cloudy way, you’re probably doing it right. But if you’re doing the cloud is rented data center, when you copy a security stack, you lift and shift your IDS, and your network capture devices, and your firewalls, and your SIM, you maybe are okay, as a first step. People here used to be a little bit more enraged about it, but to me, we meet customers where they are, but we need to journey with them.

Because if all you do is copy your stack—security stack—from a data center to the cloud, you are losing effectiveness, you’re spending money, and you’re making other mistakes. I sometimes joke that you copy mistakes, not just practices. Why copy on-prem mistakes to the cloud? So, that’s been bugging me quite a bit and I’m trying to tell stories to guide people out of a situation. Not away, but out.

Corey: A lot of people don’t go for the idea of the lift and shift migration and they say that it’s a terrible pattern and it causes all kinds of problems. And they’re right. The counterpoint is that it’s basically the second-worst approach and everything else seems to tie itself for first place. I don’t mean to sound like I’m trying to pick a fight on these things, but we’re going to rebuild an application while we move it. Great.

Then it doesn’t work or worse works intermittently and you have no idea whether it’s the rewrite, the cloud provider, or something else you haven’t considered. It just sounds like a recipe for disaster.

Anton: For sure. And so, imagine that you’re moving the app, you’re doing cut-and-paste to the cloud of the application, and then you cut-and-paste security, and then you end up with sizeable storage costs, possibly egress costs, possibly mistakes you used to make beyond five firewalls, now you make this mistake straight on the edge. Well, not on the edge edge, but on the edge of the public internet. So, some of the mistakes do become worse when you copy them from the data center to the cloud. So, we do need to, kind of, help people to get out of the situation but not by telling them don’t do it because they will do it. We need to tell them what step B; what’s step 1.5 out of this?

Corey: And cost doesn’t drive it and security doesn’t drive it. Those are trailing functions. It has to be a capability story. It has to be about improving feature velocity or it does not get done. I have learned this the painful way.

Anton: Whatever 10x cost if you do something in the data center-ish way in the cloud, and you’re ten times more expensive, cost will drive it.

Corey: To an extent, yes. However, the problem is that companies are looking at this from the perspective of okay, we can cut our costs by 90% if we make these changes. Okay, great. It cuts the cloud infrastructure cost that way. What is the engineering time, what is the opportunity cost that they gets baked into that, and what are the other strategic priorities that team has been tasked with this year? It has to go along for the ride with a redesign that unlocks additional capability because a pure cost savings play is something I have almost never found to be an argument that carries the day.

There are always exceptions, to be clear, but the general case I found is that when companies get really focused on cost-cutting, rather than expanding into new markets, on some level, it feels like they are not in the best of health, corporately speaking. I mean, there’s a reason I’m talking about cost optimization for what I do and not cost-cutting.

It’s not about lowering the bill to zero at all cost. “Cool. Turn everything off. Your bill drops to zero.” “Oh, you don’t have a company anymore? Okay, so there’s a constraint. Let’s talk more about that.” Companies are optimized to increase revenue as opposed to reduce costs. And engineers are always more expensive than the cloud provider resources they’re using, unless you’ve done something horrifying.

Anton: And some people did, by replicating their mistakes for their inefficient data centers straight into the cloud, occasionally, yeah. But you’re right, yeah. It costs the—we had the same pattern of Gartner. It’s like, it’s not about doing cheaper in the cloud.

Corey: I really want to thank you for spending so much time talking to me. If people want to learn more about what you’re up to, how you view the world, and what you’re up to next, where’s the best place for them to find you?

Anton: At this point, it’s probably easiest to find me on Twitter. I was about to say Podcast, I was about to say my Medium blog, but frankly, all of it kind of goes into Twitter at some point. And so, I think I am twitter.com/anton_chuvakin, if I recall correctly. Sorry, I haven’t really—

Corey: You are indeed. It’s always great; it’s one of those that you have a sizable audience, and you’re like, “What is my Twitter handle, again? That’s a good question. I don’t know.” And it’s your name. Great. Cool. “So, you’re going to spell that for you, too, while you’re at it?” We will, of course, put a link to that in the [show notes 00:32:09]. I really want to thank you for being so generous with your time. I appreciate it.

Anton: Perfect. Thank you. It was fun.

Corey: Anton Chuvakin, Security Strategy Something at Google Cloud. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry comment because people are doing it wrong, but also tell me which legacy vendor you work for.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Anadelia

Anadelia is a B2B marketing leader passionate about building tech brands and growing revenue. She is currently the Sr. Director of Demand Generation at Teleport. In her spare time she enjoys live music and craft beer.

Links Referenced:

  • Teleport: https://goteleport.com/
  • @anadeliafadeev: https://twitter.com/anadeliafadeev
  • LinkedIn: https://www.linkedin.com/in/anadeliafadeev/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: DoorDash had a problem. As their cloud-native environment scaled and developers delivered new features, their monitoring system kept breaking down. In an organization where data is used to make better decisions about technology and about the business, losing observability means the entire company loses their competitive edge. With Chronosphere, DoorDash is no longer losing visibility into their applications suite. The key? Chronosphere is an open-source compatible, scalable, and reliable observability solution that gives the observability lead at DoorDash business, confidence, and peace of mind. Read the full success story at snark.cloud/chronosphere. That's snark.cloud slash C-H-R-O-N-O-S-P-H-E-R-E.

Corey: Let’s face it, on-call firefighting at 2am is stressful! So there’s good news and there’s bad news. The bad news is that you probably can’t prevent incidents from happening, but the good news is that incident.io makes incidents less stressful and a lot more valuable. incident.io is a Slack-native incident management platform that allows you to automate incident processes, focus on fixing the issues and learn from incident insights to improve site reliability and fix your vulnerabilities. Try incident.io, recover faster and sleep more.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. This may surprise some of you to realize, but every once in a while, I mention how these episodes are sponsored by different companies. Well, to peel back a little bit of the mystery behind that curtain, I should probably inform some of you that when I say that, that means that companies have paid me to talk about them. I know, shocking.

This is a revelation that will topple the podcast industry if it gets out. That’s why it’s just between us. My guest today knows this better than most. Anadelia Fadeev is the Senior Director of Demand Generation at Teleport, who does in fact sponsor a number of different things that I do, but this is not a sponsored episode in that context. Anadelia, thank you for joining me today.

Anadelia: Thank you for having me.

Corey: It’s interesting. I always have to double-check where it is that you happen to be working because when we first met you were a Senior Marketing Manager, also in Demand Gen, at InfluxData, then you were a Director of Demand Generation at LightStep, and then you became a Director of Demand Gen and Growth and then a Senior Director of Demand Gen, where you are now at Teleport. And the couple of things that I’ve noticed are, one, you seem to more or less be not only doing the same role, but advancing within it, and also—selfishly—it turns out that every time you wind up working somewhere, that company winds up sponsoring some of my nonsense. So first, thank you for your business. It’s always appreciated. Now, what is demand gen exactly? Because I have to say, when I started podcasting and newslettering and shooting my mouth off on the internet, I had no clue.

Anadelia: [laugh]. Well, to put it very simply, demand generation, our goal is to drive awareness and interest in your products or services. It’s as simple as that. Now, how we do that, we could definitely dive into the specifics, but it’s all about generating awareness and interest. Especially when you work for an early-stage startup, it’s all about awareness, right? Just getting your name out there.

Corey: Marketing is one of those things that I suspect in some ways is kind of like engineering, where you take a look at, “Oh, what do you do? I’m a software engineer.” Okay, great. For someone who is in that space, does that mean front-end? Does that mean back-end? Does that mean security? Oh, wait, you’re crying and awake at weird hours and you’re angry all the time. You’re a DevOps, aren’t you?

And you start to realize that there are these breakdowns within engineering. And we realize this and we get offended when people in some cases miscategorize us as, “I am not that kind of engineer. How dare you?” Which I think is unwarranted and ridiculous, but it also sort of slips under our notice in the engineering space that marketing is every bit as divided into different functions, different roles, and the rest. For those of us who think of marketing in the naive approach, like I did when I started this place—“Oh, marketing. So basically, you do Super Bowl ads, right?” And it turns out, there might be more than one or two facets to marketing. What’s your journey been like in the wide world of marketing? Where did you start? Where does it stop?

Anadelia: Yeah. I have not gotten to the Super Bowl ads phase yet but on my way there. No, but when you think about the different core areas within marketing, right, you have your product marketing team, and this is the team that sets the positioning, the messaging, and the information about who your ideal audience is, what pain points are they having, and how is your product solving those pain points? Right, so they sort of set the direction for the rest of the team, you have another core function, which is the content team, right? So, with the direction from Product Marketing, now that we know what the pain points are and what our value prop for our product is, how do we tell that to the world in a compelling way, right? So, this is where content marketing really comes into play.

And then you have your demand generation teams. And some companies might call it growth or revenue or… I guess those two are the ones that come to mind. But this team is taking the direction from Product Marketing, taking the content produced by the content team, and then just making sure that people actually see it, right? And across all those teams, you have a lot of support from operations making sure that there’s processes and systems in place to support all of those marketing efforts, you have teams that help support web development and design, and brand.

Corey: One of the challenges that I think people have when they don’t really understand what marketing is they think back on what they know—maybe they’ve seen Mad Men, which to my understanding does not much resemble modern all workplaces, but then again, I’ve been on my own for five years, so one wonders—and they also see things in the context of companies that are targeting more mass-market, in some respects. If you’re trying to advertise Coca-Cola, every person on the planet—give or take—knows what Coca-Cola is. And the job is just to resurface it, on some level, in people’s awareness, so the correct marketing answer there apparently, is to slap the logo on a bunch of things, be it a stadium, be it a billboard, be it almost anything, whereas when we’re talking about earlier stage companies—oh, I don’t know Teleport, for example—if you were to slap the Teleport logo on a stadium somewhere for some sports game, I have the impression that most people looking at that, if they notice it at all, would instead respond to some level of confusion of, “Teleport, what is that exactly? Have scientists cracked the way of getting me to Miami from San Francisco in less than ten seconds? Because I feel like I would have heard about that.”

There’s a matter of targeting beyond just the general public or human beings walking around and starting to target people who might have a problem that you know how to solve. And then, of course, figuring out where those people are gathering and how to get in front of them in a way that resonates instead of being annoying. At least that has been my lived experience of watching the challenges that marketing people have talked to me about over the years. Is that directionally correct or are they all just shining me on and, like, “Oh, Corey, you’re adorable, you almost understand how this stuff works. Now, go insult some more things on Twitter. It’ll be fine.”

Anadelia: [laugh]. The reality is that advertising is a big part of a demand generation program, but it’s not all, right? So, good demand generation is meeting people where they are. So, the right channels, the right mediums, the right physical places. So, when you look at it from an inbound and outbound approach, inbound, you have a sign outside of your door inviting people to your house, right, and this is in the form of your website. And outbound is you go out to where people are and you knock on their door to introduce yourself.

So, when we look at it from that approach, so on the inbound side, right, the goal is to get people to come to your website because that is where you are telling them what you do and giving them the option to start using your product. So, what reason are you giving people to come to you, right? How are you helping them become better at something or achieve certain results, right? So, understanding the motivations behind it is extremely important.

And how are you driving people to you? Well, that’s where SEO comes in, right? Search engine optimization.So, what content are you producing that is driving the right search results to get your website to show up and get people to come to you, right? There’s also SEM or Search Engine Marketing. So, when people are searching for certain keywords that are relevant to you, are you showing up in those search results?

And on the outbound side of things is, what do you do to contribute to existing communities, right? So, this is where things like advertising comes into play. So, I know you have a huge following and I want to be where you are. So, of course, I’m going to sponsor your podcast and your newsletters. And similarly, I’m looking for what events are out there where I know that our potential customers are spending their time and what can we do to join that conversation in a way that adds value?

So, that can be in the form of supporting community events and meetups, giving community members a platform to share their experiences, and even supporting local businesses, right, it’s all about adding value, and by doing so, you are building trust that will allow you to then talk about how your product can help these communities solve their problems.

Corey: It’s interesting because when we look at the places that you have been, you were at InfluxData, they are a time-series database company; you were at LightStep, which was effectively an observability company, and now you’re at Teleport where you are an authentication and access company. And forgive me, none of these are your terms. These are my understandings of having talked to these folks. And on the one hand, from a product perspective, it sounds like you’re hopping between this and that and doing all those other things, and yet, we had conversations about all three of those products and how the companies around them are structured and built, and you’ve advertised all three of those on this show and others and all three of those companies and products speak specifically to problems that I have dealt with personally in the way I go through my engineering existence as well. So, instead of specializing on a particular product or on a particular niche, it almost feels like you’re specializing on a particular audience. Is that how you think about it, or is that just one of those happy accident, or in retrospect, we’re just going to retcon everything, and, “Yeah, that’s exactly why I did it.” And you’re like, “Let we jot that down. That belongs on my resume somewhere.”

Anadelia: [laugh]. No, so prior to me joining InfluxData, I was at other companies that were marketing to sales, HR, finance, different audiences, right? And the moment I joined Influx, it was really eye-opening for me to be part of a product that has an open-source community, and between that and marketing to a highly technical audience that probably very likely doesn’t want to hear from marketers, I found that to be a really good challenge for myself because it challenged me to elevate my own technical knowledge. And also personally, I just want to be surrounded by people that are smarter than me, and so I know that by being part of a community that markets to a developer audience, I am putting myself in a position where I’m having to constantly continue to learn. So, it’s a good challenge for a marketer in our industry. Just like in any others, there’s always the latest buzzword or the latest trend, and so it’s really easy to get caught up in those things. And I think that being a marketer whose audience is developers really forces you to kind of look at what you’re doing and sort of remove the fluff. This happens everywhere.

Corey: Well, I have to be careful about selling yourself too short on this because I’ve talked to a lot of different people who want to wind up promoting what it is that their companies do, and people come from all kinds of different places, and some of the less likely to be successful—in many cases, I turn the business down—are, “Well, this is our first real experience with marketing.” And the reason for that is people expect unrealistic things. I describe what I do as top-of-funnel where we get people’s attention and we give them a glimpse and a hook of what it is the product does. And I do that by talking about the painful problem that the product solves. So, when people hear their pain reflected in what we talk about, then that gives them the little bit of a push to go and take a look and see if this solves it.

And that’s great, but there has to be a process on the other side, where oh, a prospect comes in and starts looking at what it is we do. Do we have a sales funnel that moves them from someone just idly browsing to someone who might sign up for a trial, or try this in their own time, or start to understand how the community views it and the rest because just dropping a bunch of traffic on someone’s website doesn’t, in isolation, achieve anything without a means to convert that traffic into something that’s a bit more meaningful and material to the business? I’ve talked to other folks who are big on oh, well, we want to wind up just instrument in the living crap out of everything we put out there, so I want to know, when someone clicks on the ad, who they are, what they do for a living, what their signing authority is, et cetera, et cetera, et cetera. And my answer, that’s super easy, “Cool. We don’t do any of that.”

Part of the reason that people like hearing from me, is because I generally tend to respect their time, I’m not supporting invasive tracking of what they do, they don’t see my dumb face smiling with a open mouth grin as they travel across the internet on every property. Although one of these days I will see myself on the side of a bus; I’m just waiting for it. And it’s really nice to be able to talk to people who get the nuances and the peculiarities of the audience that I tend to speak to the most. You’ve always had that unlocked, even since our first conversation.

Anadelia: Yeah, well, first of all, thank you. And yeah, the reality is that, especially within my world, right—and demand generation, we are very metrics-driven because our goal [tends 00:13:00] to be pipeline, right? Pipeline for the sales team, so we want to generate sales opportunities, and in order to do that, we need to be able to measure what’s working and what is not working. But the reality is that good marketing is all about building trust, right? So, that’s why I stress the importance of providing something of value to your prospect so that you’re not wasting their time, right? The message that you have for them is something that can help them in the future.

And if building trust sometimes means I’m not able to measure the direct results of the activity that you’re doing, then that is okay, right? Because when you’re driving people to your website, there are things that you can measure, like, you have some web visits, and you know that percentage of those visitors might be interested in continue further, right? So, when you look at the journey across the buyer stages, you have to have a compelling offer for a person on each of the possible stages, right? So, if they are just learning about you today because this is the first time that heard your ad, it’s probably not expected that they would immediately go to your website and fill out your form, right? They’ve just heard about you, and now you start building that recognition.

Now, if all the stars align, and I actually have a need for a solution that’s like yours today, then, of course, you can expect a conversion to happen in that time point. But the reality is that having offers that are aimed at every stage of the buyer's journey is important.

Corey: I’m glad to hear you say this. And the reason is that I often feel like when I say it, it sounds incredibly self-serving. But if you imagine the ideal buyer and their journey, they have the exact problem that your product does and there’s an ad on my podcast that mentions it. Well, I imagine—and maybe this isn’t accurate, but it’s how I engage with podcasts myself—I’m probably not sitting in front of a computer ready to type in whatever it is that gets talked about.

I’m probably doing dishes or outside harassing a dog or something. And if it resonates is, “Oh, I should look into that.” In an ideal world. I’ll remember the short URL that I can go to, but in practice, I might just Google the company name. And oh, this does solve the problem.

If it’s not just me and there’s a team I have to have a buy-in on, I might very well mention it in our next group meeting. And, “Okay, we’re going to go ahead and try it out with an open-source version or whatnot.” And, “Oh, this seems to be working. We’ll have procurement reach out and see what it takes to wind up generating a longer-term deal.” And the original attribution of the engineer who heard it on a podcast, or the DevOps director who read it in my newsletter, or whatever it is, is long since lost. I’ve commiserated with marketing people over this, and the adage that I picked up that I love quoting is half your marketing budget is wasted, but you can spend an entire career trying to figure out which half and get nowhere by the end of it.

Anadelia: And this sort of touches on the buyer's journey is not linear. On the other side of that ad, or that marketing offer is a human, right? So, of course, as marketers, we’re going to try to build this path of once you landed on our website, we want to guide you through all the steps until you do the thing that we want you to do, but the reality is, that does not happen in your example, right? You see something, you come back to it later through another channel, there’s no way for us to measure those. And that’s okay because that’s just the reality of how humans behave.

And also, I think it’s worth noting that it takes multiple touch points until a person is ready to even hear what you have to say, right? And it sort of goes back to that point of building trust, right? It takes many times until you’ve gained that person’s trust enough for them to listen to what you have to say.

Corey: Building trust is important.

Anadelia: SIt is very important. And that’s why I think that running brand awareness programs are an extremely important part of a marketing mix. And sometimes there’s not going to be any direct attribution, and we just have to be okay with it.

Corey: I come bearing ill tidings. Developers are responsible for more than ever these days. Not just the code that they write, but also the containers and the cloud infrastructure that their apps run on. Because serverless means it’s still somebody’s problem. And a big part of that responsibility is app security from code to cloud. And that’s where our friend Snyk comes in. Snyk is a frictionless security platform that meets developers where they are - Finding and fixing vulnerabilities right from the CLI, IDEs, Repos, and Pipelines. Snyk integrates seamlessly with AWS offerings like code pipeline, EKS, ECR, and more! As well as things you’re actually likely to be using. Deploy on AWS, secure with Snyk. Learn more at Snyk.co/scream That’s S-N-Y-K.co/scream

Corey: I tend to take a perspective that trust is paramount, on some level, where we have our standard rules of, you know, don’t break the law, et cetera, et cetera, that we do require our sponsors to conform to, but there are really two rules that I have that I care about. The first is you’re not allowed to lie to the audience. Because if I wind up saying something is true in an ad or whatnot, and it’s not, that damages my credibility. And I take this old world approach of, well, I believe trust is built over time, and you continually demonstrate a pattern of doing the right thing, and people eventually are willing to extend a little bit of credulousness when you say something that sounds that might be a little bit beyond their experience.

The other is, and this is very nebulous, and difficult to define so I don’t think we even have this in writing, but you have to be able to convince me if you’re going to advertise something in one of my shows, that it will not, when used as directed, leave the user worse off than they were when they started. And that is a very strange thing. Like, a security product that has a bunch of typos on its page and is rolling its own crypto, for example—if you want an easy example—is one of those things that I will very gracefully decline not to wind up engaging with, just because I have the sneaking suspicion that if you trust that thing, you might very well live to regret it. In other cases, though—and this is almost never a problem because most companies that you have heard of and have established themselves as brands in this space already instinctively get that you’re not able to build a lasting business by lying to people and then ripping them off.

So, it’s a relatively straightforward approach, but every once in a while, I see something that makes me raise an eyebrow. And it’s not always bad. Sometimes I think that’s a little odd. Teleport is a good example of this because, “Oh, really? You wound up doing access and authentication? That sounds exactly like the kind of thing I want something old and boring, not new and exciting, around, so let’s dig into this and figure out whether this might be the one company you work at that doesn’t get to sponsor stuff that I do.”

But of course you do. You’re absolutely focusing on an area that is relevant, useful, and having talked to people on your side of the world, you’re doing the right thing. And okay, I would absolutely not be opposed to deploying this in the right production environment. But having that credulousness, having that exploratory conversation, makes it clear that I’m talking to people who know what they’re doing and not effectively shilling for the highest bidder, which is not really a position I ever want to find myself in.

Anadelia: And look, you have only one opportunity to make a first impression, right? So, being clear about what it is that you can do, and also being clear about what it is that you cannot do is extremely important, right? It kind of goes back to the point of just be a good human, don’t waste people’s time. You want to provide something of value to your audience. And so, setting those expectations early on is extremely important.

And I don’t know anyone that does this, but if your goal is only to drive people to your website, you can do that, probably very easily, but nothing will come out of it unless you have the right message.

Corey: Oh, all you do is write something incendiary and offensive, and you’ll have a lot of traffic. They won’t buy anything and they’ll hate you, but you’ll get traffic, so maybe you want to be a little bit more intentional. It’s the same reason that the companies that advertise on what I do pick me to advertise with as opposed to other things. It is more expensive than the mass-market podcasts and whatnot that speak to everyone. But you take a look at those podcasts and the things that they’re advertising are things that actually apply to an awful lot more people, things like mattresses, and click-and-design website services, and the baseline stuff that a lot of people would be interested in, whereas the things that advertise on what I do tend to look a lot more like B2B SaaS companies where they’re talking to folks who spend a lot of time working in cloud computing.

And one of the weird things to think about from that perspective, at least for me, is if one person is listening to a show that I’m putting out and they go through the journey and become a customer, well, at the size of some of these B2B contracts between large companies, that one customer has basically paid for everything I can sell for advertising for the next decade and change, just because the long-term value of some of these customers is enormous. But it’s why, for example—and I kept expecting it to happen, but it didn’t—I’ve never been subjected to outreach from the mattress companies of, “Hey, you want to go talk about that to your guests?” No, because for those folks, it is pure raw numbers: how many millions of subscribers do you have? Here, it’s—the newsletter is the easy one to get numbers on because lies, damned lies, and podcast statistics. I have 31,000 people that receive emails. Great, that’s not the biggest newsletter in the world by a longshot, but the people who are the type of person to sign up for cloud computing-style newsletters, that alone says something very specific about them and it doesn’t require anyone do anything creepy to wind up reaching out from that perspective.

It doesn’t require spying on customers to intuit that, hmm, maybe people who care about what AWS is up to and have big AWS-sized problems might sign up to a newsletter called Last Week in AWS. That’s the sort of easy thinking about advertising that I tend to go for, which yeah, admittedly sounds a lot like something out of that Mad Men era. But I think that we got a lot right back then, and everything’s new all the time.

Anadelia: [laugh]. And actually, that’s exactly what demand generation is, right? We want to find the right channels to reach our audience. And so, for a consumer company that sells mattresses, right, anyone might be on the market for a mattress, right? You want to go as broad as possible. But for something that’s more specific, you want to find what are the right channels to reach that audience where you know that there’s—it might be a smaller audience size, but it’s the right people.

And we’ve talked about the other core areas of marketing. So, with demand generation, it’s all about finding people where they are, right, and providing them their message to you and attracting them to come to you, right? It kind of goes back to that inbound and outbound motion that I mentioned earlier. But at the end of the day also, if you don’t have the right messaging to keep them engaged, once you got them to your website, then that’s a different problem, right? So, demand gen alone cannot be successful without really strong product marketing and without really strong content, and everything else that’s needed to support that, right? I mentioned the—if your website is not loading fast enough, then you’re losing people if your form is not working. So, there’s so many, so many different factors that come into play.

Corey: Oh, God, the forms. Don’t get me started on the forms. Hey, we have a great report that’s super useful. Okay, cool. I’ll click the link and I’ll follow that. I talk to sponsors about this all the time. And it’s, you have 30 mandatory fields on that website that I need to fill out. I am never going to do that.

What is the absolute bare minimum that you need in an ideal world? Don’t put any sort of gateway in front of it and just make it that good that I will reach out to thank you for it or something, but just make it an email address or something and that’s it. You don’t need to know the size of my company, the industry we’re in, the level of my signing authority, et cetera, et cetera, et cetera. Because if this is good, I might very well be in touch. And if it’s not, all you’re going to do is harass me forever with pointless calls and emails and whatnot, and I don’t want to deal with that. There’s something to be said for adding value early in the conversation and letting other people sometimes make the first move. But this is also, to be clear, a very inbound type of approach.

Anadelia: It’s a never-ending debate, to gate or not to gate. And I don’t know if there is a right answer. My approach is that if your content is good, people will come back to you. They’ll keep coming back, and they’ll want to take the next step with you. And so, I have some gated assets, and I have some that are not, and—but—

Corey: But your gates have also never been annoying of the type that I’m talking about where it’s the, “Oh, great. You need to, like, put in, like, how big is your company? What’s the budget?” It feels like I’m answering a survey at some point. AWS is notorious for this.

I counted once; there are 19 mandatory fields I had to fill out in order to watch a webinar that AWS was putting on.

Anadelia: [laugh].

Corey: And the worst part is they asked me the same questions every time I want to watch a different webinar. It’s like, for a company that says the data is so valuable, you’d really think they’d be better at managing it.

Anadelia: You know, like, some of the questions keep getting stranger. Like, I would not be surprised if people start asking what’s your favorite color, or what’s the answer to your—

Corey: The one they always ask now for, like, big data seminars and whatnot, is where this really gets me, is this in relation to your professional interests or your personal interests? It’s… “What do you think my hobbies are over there? Oh, yeah, I like big enterprise software. That’s my hobby.” “Okay, I guess.” But I really do wonder what happens if someone checks the personal interest [vibe 00:25:33]. Do they wind up just with various AWS employees showing up want to hang out on the weekends and go surfing or something? I don’t know.

Anadelia: As somebody who has been on the receiving end of lists like this—for example, we sponsor a conference and we get people stop by to talk to us, and now we get the list of those people. And there’s 25 columns. Like, honestly, that data does not come in helpful because at the end of the day, whatever you’ve marked on the required question is not going to change how I am going to communicate to you after, right, because we just had a conversation in person at this event.

Corey: My budget is not material to the reason I let you scan my badge. The reason I let you scan my badge because I really wanted one of those fun plastic toy things, so I waited in line for 45 minutes to get it. But that doesn’t mean that I’m going to be a buyer; it just means that now I’m in your funnel, although I could not possibly care less about what you do. One thing I do at re:Invent and a couple other conferences, for example, is I will have swag at a booth—because I don’t tend to get booths myself, I don’t have the staff to man it and I’m bad at that type of thing. But when people come up to get a sticker for Last Week in AWS or when of our data transfer diagram things or whatnot, the rule that we’ve always put in place is, you’re not going to mandate a badge scan for that.

And the kind of company I like doing that with gets it because the people who walk by and are interested will say, “Hey, can you scan my badge as well?” But they don’t want to pollute their own lead lists with a bunch of people who are only there to get a sticker featuring a sarcastic platypus, as opposed to getting them confused with people actually care about what it is that they’re solving for. And that’s a delicate balance to strike sometimes, but the nice thing about being me is I have customers who come back again and again and again. Although I will argue that I probably got better at being a service provider when I started also being a customer at the same time, where I hired out a marketing department here because it turns out that fixing the AWS bill is something that does a fair bit of marketing work. It’s not something people talk about at large scale in public, so you have to be noisy enough so that inbound finds its path to you a bunch of times. That’s always tricky.

And learning about how no matter what it is you do, in the case of my consulting work, we are quite honestly selling money, bring us in for an engagement, you will turn a profit on that engagement and we don’t come back with a whole bunch of extra add-ons after the fact to basically claw back more things. It’s one of the easiest sales in the world. And it’s still nuanced, and challenging, and finding the right way to talk about it to the right people at the right time explains why marketing is the industry that it is. It’s hard. None of this is easy.

Anadelia: It is. And you know, in your example, you’re not scanning that badge, but giving the person the sticker, right? Like, it’s all about making a good first impression, and if the person’s not ready to talk to you, that is okay. But there are ways that you can stay top-of-mind so that the moment that they have a need, they’ll come to you. It kind of goes back again to my earlier points of adding value in supporting existing communities, right? So, what are you doing to stay top-of-mind with that person that wasn’t quite ready back then, but the moment they have a need, they’ll think of you first because you made a good first impression.

Corey: And that’s really what it comes down to. It’s nice to talk to people who actually work in marketing because a lot of what I do in the marketing space, I’ve got to be honest, is terrible. Because I’ve done the old engineering thing of, well, I’m no marketer, but I know how to write code, so how hard could marketing really be and I invent this theory of marketing from first principles, which not only is mostly wrong, but also has a way of being incredibly insulting to people who have actually made this their profession and excel at it. But it’s an evolutionary process and trying to figure out the right way to do things and how to think about things from particular point of view has been transformative. Really easy example of this: when I first started selling sponsorships, I was constantly worried that a sponsor was going to reach out and say, “Well, hang on a second. We didn’t get the number of clicks that we expected to on this campaign. What do you have to say about that?”

Because I’m a consultant. I am used to clients not getting results that they expected having some harsh words for me. In practice, I don’t believe I’ve ever had a deep conversation about that with a marketing person. I’ve talked to them and they’ve said, “Well, some of these things worked. Some of these things didn’t. Here’s what works; here’s what didn’t, and for our next round, here’s what we want to try instead.” Those are the great constructive conversations.

The ones that I was fearing somehow would assume that I held this iron grip of control over exactly how many people would be clicking on a thing in a newsletter, and I’m not. We barely provide click-tracking at this point in the aggregate, let alone anything more specific, just because it’s so hard to actually tell and get value out of it. You talk as well, about there being brand awareness. Even if someone doesn’t click an ad, they’re potentially reading it, they’re starting to associate your company with the problem space. That’s one of those things that are effectively impossible to track, but it does pay dividends.

When you suddenly have a problem in a particular area. And there’s one or two companies off the top of your mind that you know work in that space. Well, what do you think marketing is? There has been huge money put into making that association in your mind. It’s not just about click the link; it’s not just about buy the thing; it’s about shaping the way that we think about different things.

Anadelia: And I spend a lot of time thinking about how people think we talk about what are the things that motivate you. When you have a problem, where do you go to look for a solution, or who do you go to, right? So, just understanding what the thought process is when someone is trying to solve a problem or making a purchasing decision, I think that a lot of demand generation is what are the different ways by which someone is trying to solve a problem that they’re having? And I had an interest in psychology growing up; both my parents are psychologists, and I think that marketing tends to bring some aspects of that in business and creativity, which is what led me to a career in marketing.

And you ended up being sort of a connector, right? Like your job was to connect to people who would benefit from meeting each other. Just one of them happens to be a product, or you know, it depends on your company, right, but you’re just introducing people and making sure they know about each other because there’s going to be a mutually beneficial relationship between them.

Corey: That seems to be what so many jobs ultimately distilled down to in the final analysis of things. I really want to thank you for being so generous with your time and talking about how you view the world slash industry in which we live. If people want to learn more about what you’re up to and how you think about these things, where’s the best place to find you?

Anadelia: You can follow me on Twitter at @anadeliafadeev, or connect with me on LinkedIn.

Corey: Oh, you’re one of the LinkedIn peoples. I used to do that a bit, and then I just started getting deluged with all kinds of nonsense, and let me adjust my notification settings, and there are 600 of them. And no, no, no, no, no. And I basically have quit the field, by and large, on LinkedIn. But power to you for not having done that. Links to that will of course be in the [show notes 00:32:38]. Thank you so much for being so generous with your time.

Anadelia: Thank you for having me. I appreciate it.

Corey: Anadelia Fadeev, Senior Director of Demand Generation at Teleport. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry ranting comment about how we got it completely wrong and that marketing does not work on you in the least. And by the way, when you close out that ranting comment, tell me what kind of brand of shoes you’re wearing today.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Jeff

Jeff Smith has been in the technology industry for over 20 years, oscillating between management and individual contributor. Jeff currently serves as the Director of Production Operations for Basis Technologies (formerly Centro), an advertising software company headquartered in Chicago, Illinois. Before that he served as the Manager of Site Reliability Engineering at Grubhub.

Jeff is passionate about DevOps transformations in organizations large and small, with a particular interest in the psychological aspects of problems in companies. He lives in Chicago with his wife Stephanie and their two kids Ella and Xander.

Jeff is also the author of Operations Anti-Patterns, DevOps Solutions with Manning publishing. (https://www.manning.com/books/operations-anti-patterns-devops-solutions)

Links Referenced:

  • Basis Technologies: https://basis.net/
  • Operations Anti-Patterns: https://attainabledevops.com/book
  • Personal Site: https://attainabledevops.com
  • LinkedIn: https://www.linkedin.com/in/jeffery-smith-devops/
  • Twitter: https://twitter.com/DarkAndNerdy
  • Medium: https://medium.com/@jefferysmith
  • duckbillgroup.com: https://duckbillgroup.com

Transcript
Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored by our friends at Fortinet. Fortinet’s partnership with AWS is a better-together combination that ensures your workloads on AWS are protected by best-in-class security solutions powered by comprehensive threat intelligence and more than 20 years of cybersecurity experience. Integrations with key AWS services simplify security management, ensure full visibility across environments, and provide broad protection across your workloads and applications. Visit them at AWS re:Inforce to see the latest trends in cybersecurity on July 25-26 at the Boston Convention Center. Just go over to the Fortinet booth and tell them Corey Quinn sent you and watch for the flinch. My thanks again to my friends at Fortinet.

Corey: Let’s face it, on-call firefighting at 2am is stressful! So there’s good news and there’s bad news. The bad news is that you probably can’t prevent incidents from happening, but the good news is that incident.io makes incidents less stressful and a lot more valuable. incident.io is a Slack-native incident management platform that allows you to automate incident processes, focus on fixing the issues and learn from incident insights to improve site reliability and fix your vulnerabilities. Try incident.io, recover faster and sleep more.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. One of the fun things about doing this show for long enough is that you eventually get to catch up with people and follow up on previous conversations that you’ve had. Many years ago—which sounds like I’m being sarcastic, but is increasingly actually true—Jeff Smith was on the show talking about a book that was about to release. Well, time has passed and things have changed. And Jeff Smith is back once again. He’s the Director of Product Operations at Basis Technologies, and the author of DevOps Anti-Patterns? Or what was the actual title of the book it was—

Jeff: Operations Anti-Patterns.

Corey: I got hung up in the anti-patterns part because it’s amazing. I love the title.

Jeff: Yeah, Operations Anti-Patterns, DevOps Solutions.

Corey: Got you. Usually in my experience, alway been operations anti-patterns, and here I am to make them worse, probably by doing something like using DNS as a database or some godforsaken thing. But you were talking about the book aspirationally a few years ago, and now it’s published and it has been sent out to the world. And it went well enough that they translated it to Japanese, I believe, and it has seen significant uptick. What was your experience of it? How did it go?

Jeff: You know, it was a great experience. This is definitely the first book that I’ve written. And the Manning process was extremely smooth. You know, they sort of hold your hand through the entire process. But even after launch, just getting feedback from readers and hearing how it resonated with folks was extremely powerful.

I was surprised to find out that they turned it into an audiobook as well. So, everyone reaches out and says, “Did you read the audiobook? I was going to buy it, but I wasn’t sure.” I was like, “No, unfortunately, I don’t read it.” But you know, still cool to have it out there.

Corey: My theory has been for a while now that no one wants to actually write a book; they want to have written a book. Now that you’re on the other side, how accurate is that? Are you in a position of, “Wow, sure glad that’s done?” Or are you, “That was fun. Let’s do it again because I like being sad all the time.” I mean, you do work Kubernetes for God’s sake. I mean, there’s a bit of masochism inherent to all of us in this space.

Jeff: Yeah. Kubernetes makes me cry a little bit more than the writing process. But it’s one of the things when you look back on it, you’re like, “Wow, that was fun,” but not in the heat of the moment, right? So, I totally agree with the sentiment that people want to have written a book but not actually gone through the process. And that’s evident by the fact that how many people try to start a book on their own without a publisher behind them, and they end up writing it for 15 years. The process is pretty grueling. The feedback is intense at first, but you start to get into a groove and you—I could see, you know, in a little while wanting to write another book. So, I can see the appeal.

Corey: And the last time you were on the show, I didn’t really bother to go in a particular topical direction because, what’s the point? It didn’t really seem like it was a top-of-mind issue to really bring up because what’s it matter; it’s a small percentage of the workforce. Now I feel like talking about remote work is suddenly taking on a bit of a different sheen than it was before the dark times arrived. Where do you land on the broad spectrum of opinions around the idea of remote work, given that you have specialized in anti-patterns, and well, as sarcastic as I am, I tend to look at almost every place I’ve ever worked is expressing different anti-patterns from time to time. So, where do you land on the topic?

Jeff: So, it’s funny, I started as a staunch office supporter, right? I like being in the office. I like collaborating in person; I thought we were way more productive. Since the pandemic, all of us are forced into remote work, I’ve hired almost half of my team now as remote. And I am somewhat of a convert, but I’m not on the bandwagon of remote work is just as good or is better as in person work.

I’ve firmly landed in the camp of remote work is good. It’s got its shortcomings, but it’s worth the trade off. And I think acknowledging what those trade-offs are important to keeping the team afloat. We just recently had a conversation with the team where we were discussing, like, you know, there’s definitely been a drop in productivity over the past six months to a year. And in that conversation, a lot of the things that came up were things that are different remote that were better in person, right, Slack etiquette—which is something, you know, I could talk a little bit about as well—but, you know, Slack etiquette in terms of getting feedback quickly, just the sort of camaraderie and the lack of building that camaraderie with new team members as they come on board and not having those rituals to replace the in-person rituals. But through all that, oddly enough, no one suggested going back into the office. [laugh].

Corey: For some strange reason, yeah. I need to be careful what I say here, I want to disclaim the position that I’m in. There is a power imbalance and nothing I say is going to be able to necessarily address that because I own the company and if my team members are listening to this, they’re going to read a lot into what I say that I might not necessarily intend. But The Duckbill Group, since its founding, has been a fully distributed company. My business partner lives in a different state than I do so there’s never been the crappy version of remote, which is, well, we’re all going to be in the same city, except for Theodore. Theodore is going to be timezones away and then wonder why he doesn’t get to participate in some of the conversations where the real decisions get made.

Like that’s crappy. I don’t like that striated approach to things. We don’t have many people who are co-located in any real sense, nor have we for the majority of the company’s life. But there are times when I am able to work on a project in a room with one of my colleagues, and things go a lot more smoothly. As much as we want to pretend that video is the same, it quite simply isn’t.

It is a somewhat poor substitute for the very high bandwidth of a face-to-face interaction. And yes, I understand this is also a somewhat neurotypical perspective, let’s be clear with that as well, and it’s not for everyone. But I think that for the base case, a lot of the remote work advocates are not being fully, I guess, honest with themselves about some of the shortcomings remote has. That is where I’ve mostly landed on this. Does that generally land with where you are?

Jeff: Yeah, that’s exactly where I’m at. I completely agree. And when we take work out of the equation, I think the shortcomings lay themselves bare, right? Like I was having a conversation with a friend and we were like, well, if you had a major breakup, right, I would never be like, “Oh, man. Grab a beer and hop on Zoom,” right? [laugh]. “Let’s talk it out.”

No, you’re like, hey, let’s get in person and let’s talk, right? We can do all of that conversation over Zoom, but the magic of being in person and having that personal connection, you know, can’t be replaced. So, you know, if it’s not going to work, commiserating over beers, right? I can’t imagine it’s going to work, diagramming some complex workflows and trying to come to an answer or a solution on that. So again, not to say that, you know, remote work is not valuable, it’s just different.

And I think organizations are really going to have to figure out, like, okay, if I want to entice people back into the office, what are the things that I need to do to make this realistic? We’ve opened the floodgates on remote hiring, right, so now it’s like, okay, everyone’s janky office setup needs to get fixed, right? So, I can’t have a scenario where it’s like, “Oh, just point your laptop at the whiteboard, right?” [laugh]. Like that can’t exist, we have to have office spaces that are first-class citizens for our remote counterparts as well.

Corey: Right because otherwise, the alternative is, “Great, I expect you to take the home that you pay for and turn it into an area fit for office use. Of course, we’re not going to compensate you for that, despite the fact that, let’s be realistic, rent is often larger than the AWS bill.” Which I know, gasp, I’m as shocked as anyone affected by that, but it’s true. “But oh, you want to work from home? Great. That just means you can work more hours.”

I am not of the school of thought where I consider time in the office to be an indicator of anything meaningful. I care if the work gets done and at small-scale, this works. Let me also be clear, we’re an 11-person company. A lot of what I’m talking about simply will not scale to companies that are orders of magnitude larger than this. And from where I sit, that’s okay. It doesn’t need to.

Jeff: Right. And I think a lot of the things that you talk about will scale, right? Because in most scenarios, you’re not scaling it organizationally so much as you are with a handful of teams, right? Because when I think about all the different teams I interact with, I never really interact with the organization as a whole, I interact with my little neighborhood in the organization. So, it is definitely something that scales.

But again, when it comes to companies, like, enticing people back into the office, now that I’m talking about working from home five days a week, I’ve invested in my home setup. I’ve got the monitor I want, I’ve got the chair that I want, I’ve got the mouse and keyboard that I want. So, you’re going to bring me back to the office so I can have some standard Dell keyboard and mouse with some janky, you know—maybe—21-inch monitor or something like that, right? Like, you really have to decide, like, okay, we’re going to make the office a destination, we’re going to make it where people want to go there where it’s not just even about the collaboration aspect, but people can still work and be effective.

And on top of that, I think how we look at what the office delivers is going to change, right? Because now when I go to the office now, I do very little work. It’s connections, right? It’s like, you know, “Oh, I haven’t seen you in forever. Let’s catch up.” And a lot of that stuff is valuable. You know, there’s these hallway conversations that exist that just weren’t happening previously because how do I accidentally bump into you on Slack? [laugh]. Right, it has to be much more it of a—

Corey: Right. It takes some contrivance to wind up making that happen. I remember back in the days of working in offices, I remember here in San Francisco where we had unlimited sick time and unlimited PTO, I would often fake a sick day, but just stay home and get work done. Because I knew if I was in the office, I’d be constantly subjected to drive-bys the entire time of just drive-by requests, people stopping by to ask, “Oh, can you just help me with this one thing,” that completely derails my train of thought. Then at the end of the day, they’d tell me, “You seem distractible and you didn’t get a lot of work done.”

It’s, “Well, no kidding. Of course not. Are you surprised?” And one of the nice things about starting your own company—because there are a lot of downsides, let me be very clear—one of the nice things is you get to decide how you want to work. And that was a study in, first, amazement, and then frustration.

It was, “All right, I just landed a big customer. I’m off to the races and going to take this seriously for a good six to twelve months. Great sky’s the limit, I’m going to do up my home office.” And then you see how little money it takes to have a nice chair, a good standing desk, a monitor that makes sense and you remember fighting tooth-and-nail for nothing that even approached this quality at companies and they acted like it was going to cost them 20-grand. And here, it’s two grand at most, when I decorated this place the first time.

And it was… “What the hell?” Like, it feels like the scales fall away from your eyes, and you start seeing things that you didn’t realize were a thing. Now I worry that five years in, there’s no way in the world I’m ever fit to be an employee again, so this is probably the last job I’ll ever have. Just because I’ve basically made myself completely unemployable across six different axes.

Jeff: [laugh]. And I think one of the things when it comes to, like, furniture, keyboard, stuff like that, I feel like part of it was just, like, this sort of enforced conformity, right, that the office provided us the ability to do. We can make sure everyone’s got the same monitor, the same keyboard that way, when it breaks, we can replace it easily. In a lot of organizations that I’ve been in, you know, that sort of like, you know, even if it was the same amount or ordering a custom keyboard was a big exception process, right? Like, “Oh, we’ve got to do a whole thing.” And it’s just like, “Well, it doesn’t have to be that complicated.”

And like you said, it doesn’t cost much to allow someone to get the tools that they want and prefer and they’re going to be more productive with. But to your point really quickly about work in the office, until the pandemic, I personally didn’t recognize how difficult it actually was to get work done in the office. I don’t think I appreciated it. And now that I’m remote, I’m like, wow, it is so much easier for me to close this door, put my headphones on, mute Slack and go heads down. You know, the only drive-by I’ve got is my wife wondering if I want to go for a walk, and that’s usually a text message that I can ignore and come back to later.

Corey: The thing that just continues to be strange for me and breaks in some of the weirdest ways has just been the growing awareness of how much of office life is unnecessary and ridiculous. When you’re in the office every day, you have to find a way to make it work and be productive and you have this passive-aggressive story of this open office, it’s for collaboration purposes. Yeah, I can definitively say that is not true. I had a boss who once told me that there was such benefits to working in an open plan office that if magically it were less expensive to give people individual offices, he would spare the extra expense for open plan. That was the day I learned he would lie to me while looking me in the eye. Because of course you wouldn’t.

And it’s for collaboration. Yeah, it means two loud people—often me—are collaborating and everyone else wears noise-canceling headphones trying desperately to get work done, coming in early, hours before everyone else to get things done before people show up and distracted me. What the hell kind of day-to-day work environment is that?

Jeff: What’s interesting about that, though, is those same distractions are the things that get cited as being missed from the perspective of the person doing the distracting. So, everyone universally hates that sort of drive-by distractions, but everyone sort of universally misses the ability to say like, “Hey, can I just pull on your ear for a second and get your feedback on this?” Or, “Can we just walk through this really quickly?” That’s the thing that people miss, and I don’t think that they ever connect it to the idea that if you’re not the interruptee, you’re the interruptor, [laugh] and what that might do to someone else’s productivity. So, you would think something like Slack would help with that, but in reality, what ends up happening is if you don’t have proper Slack etiquette, there’s a lot of signals that go out that get misconstrued, misinterpreted, internalized, and then it ends up impacting morale.

Corey: And that’s the most painful part of a lot of that too. Is that yeah, I want to go ahead and spend some time doing some nonsense—as one does; imagine that—and I know that if I’m going to go into an office or meet up with my colleagues, okay, that afternoon or that day, yeah, I’m planning that I’m probably not going to get a whole lot of deep coding done. Okay, great. But when that becomes 40 hours a week, well, that’s a challenge. I feel like being full remote doesn’t work out, but also being in the office 40 hours a week also feels a little sadistic, more than almost anything else.

I don’t know what the future looks like and I am privileged enough that I don’t have to because we have been full remote the entire time. But what we don’t spend on office space we spend on plane tickets back and forth so people can have meetings. In the before times, we were very good about that. Now it’s, we’re hesitant to do it just because it’s we don’t want people traveling before the feel that it’s safe to do so. We’ve also learned, for example, when dealing with our clients, that we can get an awful lot done without being on site with them and be extraordinarily effective.

It was always weird have traveled to some faraway city to meet with the client, and then you’re on a Zoom call from their office with the rest of the team. It’s… I could have done this from my living room.

Jeff: Yeah. I find those sorts of hybrid meetings are often worse than if we were all just remote, right? It’s just so much easier because now it’s like, all right, three of us are going to crowd around one person’s laptop, and then all of the things that we want to do to take advantage of being in person are excluding the people that are remote, so you got to do this careful dance. The way we’ve been sort of tackling it so far—and we’re still experimenting—is we’re not requiring anyone to come back into the office, but some people find it useful to go to the office as a change of scenery, to sort of, like break things up from their typical routine, and they like the break and the change. But it’s something that they do sort of ad hoc.

So, we’ve got a small group that meets, like, every Thursday, just as a day to sort of go into the office and switch things up. I think the idea of saying everyone has to come into the office two or three days a week is probably broken when there’s no purpose behind it. So, my wife technically should go into the office twice a week, but her entire team is in Europe. [laugh]. So, what point does that make other than I am a body in a chair? So, I think companies are going to have to get flexible with this sort of hybrid environment.

But then it makes you wonder, like, is it worth the office space and how many people are actually taking advantage of it when it’s not mandated? We find that our office time centers around some event, right? And that event might be someone in town that’s typically remote. That might be a particular project that we’re working on where we want to get ideas and collaborate and have a workshop. But the idea of just, like, you know, we’re going to systematically require people to be in the office x many days, I don’t see that in our future.

Corey: No, and I hope you’re right. But it also feels like a lot of folks are also doing some weird things around the idea of remote such as, “Oh, we’re full remote but we’re going to pay you based upon where you happen to be sitting geographically.” And we find that the way that we’ve done this—and again, I’m not saying there’s a right answer for everyone—but we wind up paying what the value of the work is for us. In many cases, that means that we would be hard-pressed to hire someone in the Bay Area, for example. On the other hand, it means that when we hire people who are in places with relatively low cost of living, they feel like they’ve just hit the lottery, on some level.

And yeah, some of them, I guess it does sort of cause a weird imbalance if you’re a large Amazon-scale company where you want to start not disrupting local economies. We’re not hiring that many people, I promise. So, there’s this idea of figuring out how that works out. And then where does the headquarters live? And well, what state laws do we wind up following on what we’re doing? Just seems odd.

Jeff: Yeah. So, you know, one thing I wanted to comment on that you’d mentioned earlier, too, was the weird things that people are doing, and organizations are doing with this, sort of, remote work thing, especially the geographic base pay. And you know, a lot of it is, how can we manipulate the situation to better us in a way that sounds good on paper, right? So, it sounds perfectly reasonable. Like, oh, you live in New York, I’m going to pay you in New York rates, right?

But, like, you live in Des Moines, so I’m going to pay you Des Moines rates. And on the surface, when you just go you’re like, oh, yeah, that makes sense, but then you think about it, you’re like, “Wait, why does that matter?” Right? And then, like, how do I, as a manager, you know, level that across my employees, right? It’s like, “Oh, so and so is getting paid 30 grand less. Oh, but they live in a cheaper area, right?” I don’t know what your personal situation is, and how much that actually resonates or matters.

Corey: Does the value that they provide to your company materially change based upon where they happen to be sitting that week?

Jeff: Right, exactly. But it’s a good story that you can tell, it sounds fair at first examination. But then when you start to scratch the surface, you’re like, “Wait a second, this is BS.” So, that’s one thing.

Corey: It’s like tipping on some level. If you can’t afford the tip, you can’t afford to eat out. Same story here. If you can’t afford to compensate people the value that they’re worth, you can’t afford to employ people. And figure that out before you wind up disappointing people and possibly becoming today’s Twitter main character.

Jeff: Right. And then the state law thing is interesting. You know, when you see states like California adopting laws similar to, like, GDPR. And it’s like, do you have to start planning for the most stringent possibility across every hire just to be safe and to avoid having to have this sort of patchwork of rules and policies based on where someone lives? You might say like, “Okay, Delaware has the most stringent employer law, so we’re going to apply Delaware’s laws across the board.” So, it’ll be interesting to see how that sort of plays out in the long run. Luckily, that’s not a problem I have to solve, but it’ll be interesting to see how it shakes out.

Corey: It is something we had to solve. We have an HR consultancy that helps out with a lot of these things, but the short answer is that we make sure that we obey with local laws, but the way that we operate is as if everyone were a San Francisco employee because that is—so far—the locale that, one, I live here, but also of every jurisdiction we’ve looked at in the United States, it tends to have the most advantageous to the employee restrictions and requirements. Like one thing we do is kind of ridiculous—and we have to do for me and one other person, but almost no one else, but we do it for everyone—is we have to provide stipends every month for electricity, for cellphone usage, for internet. They have to be broken out for each one of those categories, so we do 20 bucks a month for each of those. It adds up to 100 bucks, as I recall, and we call it good. And employees say, “Okay. Do we just send you receipts? Please don’t.”

I don’t want to look at your cell phone bill. It’s not my business. I don’t want to know. We’re doing this to comply with the law. I mean, if it were up to me, it would be this is ridiculous. Can we just give everyone $100 a month raise and call it good? Nope. The forms must be obeyed. So, all right.

We do the same thing with PTO accrual. If you’ve acquired time off and you leave the company, we pay it out. Not every state requires that. But paying for cell phone access and internet access as well, is something Amazon is currently facing a class action about because they didn’t do that for a number of their California employees. And even talking to Amazonians, like, “Well, they did, but you had to jump through a bunch of hoops.”

We have the apparatus administratively to handle that in a way that employees don’t. Why on earth would we make them do it unless we didn’t want to pay them? Oh, I think I figured out this sneaky, sneaky plan. I’m not here to build a business by exploiting people. If that’s the only way to succeed, and the business doesn’t deserve to exist. That’s my hot take of the day on that topic.

Jeff: No, I totally agree. And what’s interesting is these insidious costs that sneak up that employees tend to discount, like, one thing I always talk about with my team is all that time you’re thinking about a problem at work, right, like when you’re in the shower, when you’re at dinner, when you’re talking it over with your spouse, right? That’s work. That’s work. And it’s work that you’re doing on your time.

But we don’t account for it that way because we’re not typing; we’re not writing code. But, like, think about how much more effective as people, as employees, we would be if we had time dedicated to just sit and think, right? If I could just sit and think about a problem without needing to type but just critically think about it. But then it’s like, well, what does that look like in the office, right? If I’m just sitting there in my chair like this, it doesn’t look like I’m doing anything.

But that’s so important to be able to, like, break down and digest some of the complex problems that we’re dealing with. And we just sort of write it off, right? So, I’m like, you know, you got to think about how that bleeds into your personal time and take that into account. So yeah, maybe you leave three hours early today, but I guarantee you, you’re going to spend three hours throughout the week thinking about work. It’s the same thing with these cellphone costs that you’re talking about, right? “Oh, I’ve got a cell phone anyways; I’ve got internet anyways.” But still, that’s something that you’re contributing to the business that they’re not on the hook for, so it seems fair that you get compensated for that.

Corey: I just think about that stuff all the time from that perspective, and now that I you know, own the place, it’s one of those which pocket of mine does it come out of? But I hold myself to a far higher standard about that stuff than I do the staff, where it’s, for example, I could theoretically justify paying my internet bill here because we have business-class internet and an insane WiFi system because of all of the ridiculous video production I do. Now. It’s like, like, if anyone else on the team was doing this, yes, I will insist we pay it, but for me, it just it feels a little close to the edge. So, it’s one of those areas where I’m very conservative around things like that.

The thing that also continues to just vex me, on some level, is this idea that time in a seat is somehow considered work. I’ll never forget one of the last jobs I had before I started this place. My boss walked past me and saw that I was on Reddit. And, “Is that really the best use of your time right now?” May I use the bathroom when I’m done with this, sir?

Yeah, of course it is. It sounds ridiculous, but one of the most valuable things I can do for The Duckbill Group now is go on the internet and start shit posting on Twitter, which sounds ridiculous, but it’s also true. There’s a brand awareness story there, on some level. And that’s just wild to me. It’s weird, we start treating people like adults, they start behaving that way. And if you start micromanaging them, they live up or down to the expectations you tend to hold. I’m a big believer in if I have to micromanage someone, I should just do the job myself.

Jeff: Yeah. The Reddit story makes me think of, like, how few organizations have systematic ways of getting vital information. So, the first thing I think about is, like, security and security vulnerabilities, right? So, how does Basis Technologies, as an organization, know about these things? Right now, it’s like, well, my team knows because we’re plugged into Reddit and Twitter, right, but if we were gone Basis, right, may not necessarily get that information.

So, that’s something we’re trying to correct, but it just sort of highlights the importance of freedom for these employees, right? Because yeah, I’m on Reddit, but I’m on /r/sysadmin. I’m on /r/AWS, right, I’m on /r/Atlassian. Now I’m finding out about this zero-day vulnerability and it’s like, “Oh, guys, we got to act. I just heard about this thing.” And people are like, “Oh, where did this come from?” And it’s like it came from my network, right? And my network—

Corey: Mm-hm.

Jeff: Is on Twitter, LinkedIn, Reddit. So, the idea that someone browsing the internet on any site, really, is somehow not a productive use of their time, you better be ready to itemize exactly what that means and what that looks like. “Oh, you can do this on Reddit but you can’t do that on Reddit.”

Corey: I have no boss now, I have no oversight, but somehow I still show up with a work ethic and get things done.

Jeff: Right. [laugh].

Corey: Wow, I guess I didn’t need someone over my shoulder the whole time. Who knew?

Jeff: Right. That’s all that matters, right? And if you do it in 30 hours or 40 hours, that doesn’t really matter to me, you know? You want to do it at night because you’re more productive there, right, like, let’s figure out a way to make that happen. And remote work is actually empowering us ways to really retain people that wasn’t possible before I had an employee that was like, you know, I really want to travel. I’m like, “Dude, go to Europe. Work from Europe. Just do it. Work from Europe,” right? We’ve got senior leaders on the C-suite that are doing it. One of the chief—

Corey: I’m told they have the internet, even there. Imagine that?

Jeff: Yeah. [laugh]. So, our chief program officer, she was in Greece for four weeks. And it worked. It worked great. They had a process. You know, she would spent one week on and then one week off on vacation. But you know, she was able to have this incredible, long experience, and still deliver. And it’s like, you know, we can use that as a model to say, like—

Corey: And somehow the work got done. Wow, she must be amazing. No, that’s the baseline expectation that people can be self-managing in that respect.

Jeff: Right.

Corey: They aren’t toddlers.

Jeff: So, if she can do that, I’m sure you can figure out how to code in China or wherever you want to visit. So, it’s a great way to stay ahead of some of these companies that have a bit more lethargic policies around that stuff, where it’s like, you know, all right, I’m not getting that insane salary, but guess what, I’m going to spend three weeks in New Zealand hanging out and not using any time off or anything like that, and you know, being able to enjoy life. I wish this pandemic had happened pre-kids because—

Corey: Yeah. [laugh].

Jeff: —you know, we would really take advantage of this.

Corey: You and me both. It would have very different experience.

Jeff: Yeah. [laugh]. Absolutely, right? But with kids in school, and all that stuff, we’ve been tethered down. But man, I you know, I want to encourage the young people or the single people on my team to just, like, hey, really, really embrace this time and take advantage of it.

Corey: I come bearing ill tidings. Developers are responsible for more than ever these days. Not just the code that they write, but also the containers and the cloud infrastructure that their apps run on. Because serverless means it’s still somebody’s problem. And a big part of that responsibility is app security from code to cloud. And that’s where our friend Snyk comes in. Snyk is a frictionless security platform that meets developers where they are - Finding and fixing vulnerabilities right from the CLI, IDEs, Repos, and Pipelines. Snyk integrates seamlessly with AWS offerings like code pipeline, EKS, ECR, and more! As well as things you’re actually likely to be using. Deploy on AWS, secure with Snyk. Learn more at Snyk.co/scream That’s S-N-Y-K.co/scream

Corey: One last topic I want to get into before we call it an episode is, I admit, I read an awful lot of books, it’s a guilty pleasure. And it’s easy to fall into the trap, especially when you know the author, of assuming that snapshot of their state of mind at a very fixed point in time is somehow who they are, like a fly frozen in amber, and it’s never true. So, my question for you is, quite simply, what have you learned since your book came out?

Jeff: Oh, man, great question. So, when I was writing the book, I was really nervous about if my audience was as big as I thought it was, the people that I was targeting with the book.

Corey: Okay, that keeps me up at night, too. I have no argument there.

Jeff: Yeah. You know what I mean?

Corey: Please, continue.

Jeff: I’m surrounded, you know, by—

Corey: Is anyone actually listening to this? Yeah.

Jeff: Right. [laugh]. So, after the book got finished and it got published, I would get tons of feedback from people that so thoroughly enjoyed the book, they would say things like, you know, “It feels like you were in our office like a fly on the wall.” And that was exciting, one, because I felt like these were experiences that sort of resonated, but, two, it sort of proved this thesis that sometimes you don’t have to do something revolutionary to be a positive contribution to other people, right? So, like, when I lay out the tips and things that I do in the book, it’s nothing earth-shattering that I expect Google to adopt. Like, oh, my God, this is the most unique view ever.

But being able to talk to an audience in a way that resonates with them, that connects with them, that shows that I understand their problem and have been there, it was really humbling and enlightening to just see that there are people out there that they’re not on the bleeding edge, but they just need someone to talk to them in a language that they understand and resonate with. So, I think the biggest thing that I learned was this idea that your voice is important, your voice matters, and how you tell your story may be the difference between someone understanding a concept and someone not understanding a concept. So, there’s always an audience for you out there as you’re writing, whether it be your blog post, the videos that you produce, the podcasts that you make, somewhere there’s someone that needs to hear what you have to say, and the unique way that you can say it. So, that was extremely powerful.

Corey: Part of the challenge that I found is when I start talking to other people, back in the before times, trying to push them into conference talks and these days, write blog posts, the biggest objection I get sometimes is, “Well, I don’t have anything worth saying.” That is provably not true. One of my favorite parts about writing Last Week in AWS is as I troll the internet looking for topics about AWS that I find interesting, I keep coming across people who are very involved in one area or another of this ecosystem and have stories they want to tell. And I love, “Hey, would you like to write a guest post for Last Week in AWS?” It’s always invite only and every single one of them has been paid because people die of exposure and I’m not about that exploitation lifestyle.

A couple have said, “Oh, I can’t accept payment for a variety of reasons.” Great. Pick a charity that you would like it to go to instead because we do not accept volunteer work, we are a for-profit entity. That is the way it works here. And that has been just one of the absolute favorite parts about what I do just because you get to sort of discover new voices.

And what I find really neat is that for a lot of these folks, this is their start to writing and telling the story, but they don’t stop there, they start telling their story in other areas, too. It leads to interesting career opportunities for them, it leads to interesting exposure that they wouldn’t have necessarily had—again, not that they’re getting paid in exposure, but the fact that they are able to be exposed to different methodologies, different ways of thinking—I love that. It’s one of my favorite parts about doing what I do. And it seems to scale a hell of a lot better than me sitting down with someone for two hours to help them build a CFP that they wind up not getting accepted or whatnot.

Jeff: Right. It’s a great opportunity that you provide folks, too, because of, like, an instant audience, I think that’s one of the things that has made Medium so successful as, like, a blogging platform is, you know, everyone wants to go out and build their own WordPress site and launch it, but then it like, you write your blog post and it’s crickets. So, the ability for you to, you know, use your platform to also expose those voices is great and extremely powerful. But you’re right, once they do it, it lights a fire in a way that is admirable to watch. I have a person that I’m mentoring and that was my biggest piece of advice I can give. It was like, you know, write. Just write.

It’s the one thing that you can do without anyone else. And you can reinforce your own knowledge of a thing. If you just say, you know, I’m going to teach this thing that I just learned, just the writing process helps you solidify, like, okay, I know this stuff. I’m demonstrating that I know it and then four years from now, when you’re applying for a job, someone’s like, “Oh, I found your blog post and I see that you actually do know how to set up a Kubernetes cluster,” or whatever. It’s just extremely great and it—

Corey: It’s always fun. You’re googling for how to do something and you find something you wrote five years ago.

Jeff: Right, yeah. [laugh]. And it’s like code where you’re like, “Oh, man, I would do that so much differently now.”

Corey: Since we last spoke, one of the things I’ve been doing is I have been on the hook to write between a one to two-thousand-word blog post every week, and I’ve done that like clockwork, for about a year-and-a-half now. And I was no slouch at storytelling before I started doing that. I’ve given a few hundred conference talks in the before times. And I do obviously long Twitter threads in the past and I write reports a lot. But forcing me to go through that process every week and then sit with an editor and go ahead and get it improved, has made me a far better writer, it’s made me a better storyteller, I am far better at articulating my point of view.

It is absolutely just unlocking a host of benefits that I would have thought I was, oh, I passed all this. I’m already good at these things. And I was, but I’m better now. I think that writing is one of those things that people need to do a lot more of.

Jeff: Absolutely. And it’s funny that you mentioned that because I just recently, back in April, started to do the same thing I said, I’m going to write a blog post every week, right? I’m going to get three or four in the can, so that if life comes up and I miss a beat, right, I’m not actually missing the production schedule, so I have a steady—and you’re right. Even after writing a book, I’m still learning stuff through the writing process, articulating my point of view.

It’s just something that carries over, and it carries over into the workforce, too. Like, if you’ve ever read a bad piece of documentation, right, that comes from—

Corey: No.

Jeff: Right? [laugh]. That comes from an inability to write. Like, you know, you end up asking these questions like who’s the audience for this? What is ‘it’ in this sentence? [laugh].

Corey: Part of it too, is that people writing these things are so close to the problem themselves that the fact that, “Well, I’m not an expert in this.” That’s why you should write about it. Talk about your experience. You’re afraid everyone’s going to say, “Oh, you’re a fool. You didn’t understand how this works.”

Yeah, my lived experiences instead—and admittedly, I have the winds of privilege of my back on this—but it’s also yeah, I didn’t understand that either. It turns out that you’re never the only person who has trouble with a concept. And by calling it out, you’re normalizing it and doing a tremendous service for others in your shoes.

Jeff: Especially when you’re not an expert because I wrote some documentation about the SSL process and it didn’t occur to me that these people don’t use the AWS command line, right? Like, you know, in our organization, we sort of mask that from them through a bunch of in-house automation. Now we’re starting to expose it to them and simple things like oh, you need to preface the AWS command with a profile name. So, then when we’re going through the setup, we’re like, “Oh. What if they already have an existing profile, right?” Like, we don’t want to clobber that.

SSo, it just changed the way you write the documentation. But like, that’s not something that initially came to mind for me. It wasn’t until someone went through the docs, and they’re like, “Uh, this is blowing up in a weird way.” And I was like, “Oh, right. You know, like, I need to also teach you about profile management.”

Corey: Also, everyone has a slightly different workflow for the way they interact with AWS accounts, and their shell prompts, and the way they set up local dev environments.

Jeff: Yeah, absolutely. So, not being an expert on a thing is key because you’re coming to it with virgin eyes, right, and you’re able to look at it from a fresh perspective.

Corey: So, much documentation out there is always coming from the perspective of someone who is intimately familiar with the problem space. Some of the more interesting episodes that I have, from a challenge perspective, are people who are deep technologists in a particular area and they love they fallen in love with the thing that they are building. Great. Can you explain it to the rest of us mere mortals so that we can actually we can share your excitement on this? And it’s very hard to get them to come down to a level where it’s coherent to folks who haven’t spent years thinking deeply about that particular problem space.

Jeff: Man, the number one culprit for that is, like, the AWS blogs where they have, like, a how-to article. You follow that thing and you’re like, “None of this is working.” [laugh]. Right? And then you realize, oh, they made an assumption that I knew this, but I didn’t right?

So, it’s like, you know, I didn’t realize this was supposed to be, like, a handwritten JSON document just jammed into the value field. Because I didn’t know that, I’m not pulling those values out as JSON. I’m expecting that just to be, like, a straight string value. And that has happened more and more times on the AWS blog than I can count. [laugh].

Corey: Oh, yeah, very often. And then there’s other problems, too. “Oh, yeah. Set up your IAM permissions properly.” That’s left as an exercise for the reader. And then you wonder why everything’s full of stars. Okay.

Jeff: Right. Yep, exactly, exactly.

Corey: Ugh. It’s so great to catch up with you and see what you’ve been working on. If people want to learn more, where’s the best place to find you?

Jeff: So, the best place is probably my website, attainabledevops.com. That’s a place where you can find me on all the other places. I don’t really update that site much, but you can find me on LinkedIn, Twitter, from that jumping off point, links to the book are there if anyone’s interested in that. Perfect stocking stuffers. Mom would love it, grandma would love it, so definitely, definitely buy multiple copies of that.

Corey: Yeah, it’s going to be one of my two-year-old’s learning to read books, it’d be great.

Jeff: Yeah, it’s perfect. You know, you just throw it in the crib and walk away, right? They’re asleep at no time. Like I said, I’ve also been taking to, you know, blogging on Medium, so you can catch me there, the links will be there on Attainable DevOps as well.

Corey: Excellent. And that link will of course, be in the show notes. Thank you so much for being so generous with your time. I really do appreciate it. And it’s great to talk to you again.

Jeff: It was great to catch up.

Corey: Really was. Jeff Smith, Director of Product Operations at Basis Technologies. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice or smash the like and subscribe buttons on the YouTubes, whereas if you’ve hated this podcast, do the exact same thing—five-star review, smash the buttons—but also leave an angry, incoherent comment that you’re then going to have edited and every week you’re going to come back and write another incoherent comment that you get edited. And in the fullness of time, you’ll get much better at writing angry, incoherent comments.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Benjamin

Benjamin Anderson is CTO, Cloud at EDB, where he is responsible for developing and driving strategy for the company’s Postgres-based cloud offerings. Ben brings over ten years’ experience building and running distributed database systems in the cloud for multiple startups and large enterprises. Prior to EDB, he served as chief architect of IBM’s Cloud Databases organization, built an SRE practice at database startup Cloudant, and founded a Y Combinator-funded hardware startup.

Links Referenced:

  • EDB: https://www.enterprisedb.com/
  • BigAnimal: biganimal.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

I come bearing ill tidings. Developers are responsible for more than ever these days. Not just the code that they write, but also the containers and the cloud infrastructure that their apps run on. Because serverless means it’s still somebody’s problem. And a big part of that responsibility is app security from code to cloud. And that’s where our friend Snyk comes in. Snyk is a frictionless security platform that meets developers where they are - Finding and fixing vulnerabilities right from the CLI, IDEs, Repos, and Pipelines. Snyk integrates seamlessly with AWS offerings like code pipeline, EKS, ECR, and more! As well as things you’re actually likely to be using. Deploy on AWS, secure with Snyk. Learn more at Snyk.co/scream That’s S-N-Y-K.co/scream

Corey: This episode is sponsored by our friends at Fortinet. Fortinet’s partnership with AWS is a better-together combination that ensures your workloads on AWS are protected by best-in-class security solutions powered by comprehensive threat intelligence and more than 20 years of cybersecurity experience. Integrations with key AWS services simplify security management, ensure full visibility across environments, and provide broad protection across your workloads and applications. Visit them at AWS re:Inforce to see the latest trends in cybersecurity on July 25-26 at the Boston Convention Center. Just go over to the Fortinet booth and tell them Corey Quinn sent you and watch for the flinch. My thanks again to my friends at Fortinet.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. This promoted guest episode is brought to us by our friends at EDB. And not only do they bring us this promoted episode, they bring me their CTO for Cloud, Benjamin Anderson. Benjamin, thank you so much for agreeing to suffer the slings and arrows that I will no doubt throw at you in a professional context, because EDB is a database company, and I suck at those things.

Benjamin: [laugh]. Thanks, Corey. Nice to be here.

Corey: Of course. So, databases are an interesting and varied space. I think we can all agree—or agree to disagree—that the best database is, of course, Route 53, when you misuse TXT records as a database. Everything else is generally vying for number two. EDB was—back in the days that I was your customer—was EnterpriseDB, now rebranded as EDB, which is way faster to say, and I approve of that.

But you were always the escalation point of last resort. When you’re stuck with a really weird and interesting Postgres problem, EDB was where you went because if you folks couldn’t solve the problem, it was likely not going to get solved. I always contextualized you folks as a consulting shop. That’s not really what you do. You are the CTO for Cloud.

And, ah, interesting. Do databases behave differently in cloud environments? Well, they do when you host them as a managed service, which is an area you folks have somewhat recently branched into. How’d you get there?

Benjamin: Ah, that’s interesting. So, there’s a bunch of stuff to unpack there. I think EDB has been around for a long time. It’s something like 13, 14, 15 years, something like that, and really it's just been kind of slowly growing, right? We did start very much as a product company. We built some technology to help customers get from Oracle database on to Postgres, way back in 2007, 2008.

That business has just slowly been growing. It’s been going quite well. Frankly, I only joined about 18 months ago, and it’s really cool tech, right? We natively understand some things that Oracle is doing. Customers don’t have to change their schemas to migrate from Oracle to Postgres. There’s some cool technology in there.

But as you point out, I think a lot of our position in the market has not been that product focused. There’s been a lot of people seeing us as the Postgres experts, and as people who can solve Postgres problems, in general. We have, for a long time, employed a lot of really sharp Postgres people. We still employ a lot of really sharp Postgres people. That’s very much, in a lot of ways, our bread and butter. That we’re going to fix Postgres problems as they come up.

Now, over the past few years, we’ve definitely tried to shift quite a bit into being more of a product company. We’ve brought on a bunch of people who’ve been doing more enterprise software product type development over the past few years, and really focusing ourselves more and more on building products and investing in ourselves as a product company. We’re not a services company. We’re not a consulting company. We do, I think, provide the best Postgres support in the market. But it’s been a journey. The cloud has been a significant part of that as well, right? You can’t get away.

Corey: Oh, yeah. These days, when someone’s spinning up a new workload, it’s unlikely—in most cases—they’re going to wind up spinning up a new data center, if they don’t already have one. Yes, there’s still a whole bunch of on-prem workloads. But increasingly, the default has become cloud. Instead of, “Why cloud?” The question’s become, “Why not?”

Benjamin: Right, exactly. Then, as people are more and more accepting of managed services, you have to be a product company. You have to be building products in order to support your database customers because what they want his managed services. I was working in managed databases and service, something like, ten years ago, and it was like pulling teeth. This is after RDS launched. This was still pulling teeth trying to get people to think about, oh, I’m going to let you run my database. Whereas, now obviously, it’s just completely different. We have to build great products in order to succeed in the database business, in general.

Corey: One thing that jumped out at me when you first announced this was the URL is enterprisedb.com. That doesn’t exactly speak to, you know, non-large companies, and EDB is what you do. You have a very corporate logo, but your managed service is called BigAnimal, which I absolutely love. It actually expresses a sense of whimsy and personality that I can no doubt guess that a whole bunch of people argued against, but BigAnimal, it is. It won through. I love that. Was that as contentious as I’m painting it to be, or people actually have a sense of humor sometimes?

Benjamin: [laugh]. Both, it was extremely contentious. I, frankly, was one of the people who was not in favor of it at first. I was in favor of something that was whimsical, but maybe not quite that whimsical.

Corey: Well, I call it Postgres-squeal, so let’s be very clear here that we’re probably not going to see eye-to-eye on most anything in pronunciation things. But we can set those differences aside and have a conversation.

Benjamin: Absolutely, no consider that. It was deliberate, though, to try to step away a little bit from the blue-suit-and-tie, enterprise, DB-type branding. Obviously, a lot of our customers are big enterprises. We’re good at that. We’re not trying to be the hip, young startup targeting business in a lot of ways. We have a wide range of customers, but we want to branch out a little bit.

Corey: One of the challenges right now is if I spin up an environment inside of AWS, as one does, and I decide I certainly don’t want to take the traditional approach of running a database on top of an EC2 instance—the way that we did in the olden days—because RDS was crappy. Now that it’s slightly less crappy, that becomes a not ideal path. I start looking at their managed database offerings, and there are something like 15 distinct managed databases that they offer, and they never turn anything off. And they continue to launch things into the far future. And it really feels, on some level, like 20 years from now—what we call a DBA today—their primary role is going to look a lot more like helping a company figure out which of Amazon’s 40 managed databases is the appropriate fit for this given workload. Yet, when I look around at what the industry has done, it seems that when we’re talking about relational databases. Postgres has emerged back when I was, more or less, abusing servers in person in my data center days, it was always MySQL. These days, Postgres is the de facto standard, full stop. I admit that I was mostly keeping my aura away from any data that was irreplaceable at that time. What happened? What did I miss?

Benjamin: It’s a really good question. And I certainly am not a hundred percent on all the trends that went on there. I know there’s a lot of folks that are not happy about the MySQL acquisition by Oracle. I think there’s a lot of energy that was adopted by the NoSQL movement, as well. You have people who didn’t really care about transactional semantics that were using MySQL because they needed a place to store their data. And then, things like MongoDB and that type of system comes along where it’s significantly easier than MySQL, and that subset of the population just sort of drifts away from MySQL.

Corey: And in turn, those NoSQL projects eventually turn into something where, okay, now we’re trying to build a banking system on top of it, and it’s, you know, I guess you can use a torque wrench as a hammer if you’re really creative about it, but it seems like there’s a better approach.

Benjamin: Yeah, exactly. And those folks are coming back around to the relational databases, exactly. At the same time, the advancements in Postgres from the early eight series to today are significant, right? We shouldn’t underestimate how much Postgres has really moved forward. It wasn’t that long ago that replication was hardly a thing and Postgres, right? It’s been a journey.

Corey: One thing that your website talks about is that you accelerate your open-sourced database transformation. And this is a bit of a hobby horse I get on from time to time. I think that there are a lot of misunderstandings when people talk about this. You have the open-source purists—of which I shamefully admit I used to be one—saying that, “Oh, it’s about the idea of purity and open and free as in software.” Great. Okay, awesome. But when I find that corporate customers are talking about when they say open-source database, they don’t particularly care if they have access to the source code because they’re not going to go in and patch a database engine, we hope. But what they do care about is regardless of where they are today—even if they’re perfectly happy there—they don’t want to wind up beholden to a commercial database provider, and/or they don’t want to wind up beholden to the environment that is running within. There’s a strategic Exodus that’s available in theory, which on some level serves to make people feel better about not actually Exodus-ing, but it also means if they’re doing a migration at some point, they don’t also have to completely redo their entire data plan.

Benjamin: Yeah, I think that’s a really good point. I mean, I like to talk—there’s a big rat’s nest of questions and problems in here—but I generally like talk to about open APIs, talk about standards, talk about how much is going to have to change if you eliminate this vendor. We’re definitely not open-source purists. Well, we employ a lot of open-source purists. I also used to be an open—

Corey: Don’t let them hear you say that, then. Fair enough. Fair enough.

Benjamin: [laugh] we have proprietary software at EDB, as well. There’s a kind of wide range of businesses that we participate in. Glad to hear you also mention this where-it’s-hosted angle, as well. I think there’s some degree to which people are—they figured out that having at least open APIs or an open-source-ish database is a good idea rather than being beholden to proprietary database. But then, immediately forget that when they’re picking a cloud vendor, right? And realizing that putting their data in Cloud Vendor A versus Cloud Vendor B is also putting them in a similar difficult situation. They need to be really wary of when they’re doing that. Now, obviously, I work at an independent software company, and I have some incentive to say this, but I do think it’s true. And you know, there’s meaningful data gravity risk.

Corey: I assure you, I have no incentive. I don’t care what cloud provider you’re on. My guidance has been, for years, to—as a general rule—pick a provider, I care about which one, and go all in until there’s a significant reason to switch. Trying to build an optionality, “Oh, everything we do should be fully portable at an instance notice.” Great. Unless you’re actually doing it, you’re more or less, giving up a whole bunch of shortcuts and feature velocity you could otherwise have, in the hopes of one day you’ll do a thing, but all the assumptions you’re surrounded by baked themselves in regardless. So, you’re more or less just creating extra work for yourself for no defined benefit. This is not popular in some circles, where people try to sell something that requires someone to go multi-cloud, but here we are.

Benjamin: No, I think you’re right. I think people underestimate the degree to which the abstractions are just not very good, right, and the degree to which those cloud-specific details are going to leak in if you’re going to try to get anything done, you end up in kind of a difficult place. What I see more frequently is situations where we have a big enterprise—not even big, even medium-sized companies where maybe they’ve done an acquisition or two, they’ve got business units that are trying to do things on their own. And they end up in two or three clouds, sort of by happenstance. It’s not like they’re trying to do replication live between two clouds, but they’ve got one business unit in AWS and one business unit and Azure, and somebody in the corporate—say enterprise architect or something like that—really would like to make things consistent between the two so they get a consistent security posture and things like that. So, there are situations where the multi-cloud is a reality at a certain level, but maybe not at a concrete technical level. But I think it’s still really useful for a lot of customers.

Corey: You position your cloud offering in two different ways. One of them is the idea of BigAnimal, and the other—well, it sort of harkens back to when I was in sixth grade going through the American public school system. They had a cop come in and talk to us and paint to this imaginary story of people trying to push drugs. “Hey, kid. You want to try some of this?” And I’m reading this and it says EDB, Postgres for Kubernetes. And I’m sent back there, where it’s like, “Hey, kid. You want to run your stateful databases on top of Kubernetes?” And my default answer to that is good lord, no. What am I missing?

Benjamin: That’s a good question. Kubernetes has come a long way—I think is part of that.

Corey: Oh, truly. I used to think of containers as a pure story for stateless things. And then, of course, I put state into them, and then, everything exploded everywhere because it turns out, I’m bad at computers. Great. And it has come a long way. I have been tracking a lot of that. But it still feels like the idea being that you’d want to have your database endpoints somewhere a lot less, I guess I’ll call it fickle, if that makes sense.

Benjamin: It’s an interesting problem because we are seeing a lot of people who are interested in our Kubernetes-based products. It’s actually based on—we recently open-sourced the core of it under a project called cloud-native PG. It’s a cool piece of technology. If you think about sort of two by two. In one corner, you’ve got self-managed on-premise databases. So, you’re very, very slow-moving, big-iron type, old-school database deployments. And on the opposite corner, you’ve got fully-managed, in the cloud, BigAnimal, Amazon RDS, that type of thing. There’s a place on that map where you’ve got customers that want a self-service type experience. Whether that’s for production, or maybe it’s even for dev tests, something like that. But you don’t want to be giving the management capability off to a third party.

For folks that want that type of experience, trying to build that themselves by, like, wiring up EC2 instances, or doing something in their own data center with VMware, or something like that, can be extremely difficult. Whereas if you’ve go to a Kubernetes-based product, you can get that type of self-service experience really easily, right? And customers can get a lot more flexibility out of how they run their databases and operate their databases. And what sort of control they give to, say application developers who want to spin up a new database for a test or for some sort of small microservice, that type of thing. Those types of workloads tend to work really well with this first-party Kubernetes-based offering. I’ve been doing databases on Kubernetes in managed services for a long time as well. And I don’t, frankly, have any concerns about doing it. There are definitely some sharp edges. And if you wanted to do to-scale, you need to really know what you’re doing with Kubernetes because the naive thing will shoot you in the foot.

Corey: Oh, yes. So, some it feels almost like people want to cosplay working for Google, but they don’t want to pass the technical interview along the way. It’s a bit of a weird moment for it.

Benjamin: Yeah, I would agree.

Corey: I have to go back to my own experiences with using RDS back at my last real job before I went down this path. We were migrating from EC2-Classic to VPC. So, you could imagine what dates me reasonably effectively. And the big problem was the database. And the joy that we had was, “Okay, we have to quiesce the application.” So, the database is now quiet, stop writes, take a snapshot, restore that snapshot into the environment. And whenever we talk to AWS folks, it’s like, “So, how long is this going to take?” And the answer was, “Guess.” And that was not exactly reassuring. It went off without a hitch because every migration has one problem. We were sideswiped in an Uber on the way home. But that’s neither here nor there. This was two o’clock in the morning, and we finished in half the maintenance time we had allotted. But it was the fact that, well, guess we’re going to have to take the database down for many hours with no real visibility, and we hope it’ll be up by morning. That wasn’t great. But that was the big one going on, on an ongoing basis, there were maintenance windows with a database. We just stopped databasing for a period of time during a fairly broad maintenance window. And that led to a whole lot of unfortunate associations in my mind with using relational databases for an awful lot of stuff. How do you handle maintenance windows and upgrading and not tearing down someone’s application? Because I have to assume, “Oh, we just never patch anything. It turns out that’s way easier,” is in fact, the wrong answer.

Benjamin: Yeah, definitely. As you point out, there’s a bunch of fundamental limitations here, if we start to talk about how Postgres actually fits together, right? Pretty much everybody in RDS is a little bit weird. The older RDS offerings are a little bit weird in terms of how they do replication. But most folks are using Postgres streaming replication, to do high availability, Postgres in managed services. And honestly, of course—

Corey: That winds up failing over, or the application’s aware of both endpoints and switches to the other one?

Benjamin: Yeah—

Corey: Sort of a database pooling connection or some sort of proxy?

Benjamin: Right. There’s a bunch of subtleties that get into their way. You say, well, did the [vit 00:16:16] failover too early, did the application try to connect and start making requests before the secondaries available? That sort of thing.

Corey: Or you misconfigure it and point to the secondary, suddenly, when there’s a switchover of some database, suddenly, nothing can write, it can only read, then you cause a massive outage on the weekend?

Benjamin: Yeah. Yeah.

Corey: That may have been of an actual story I made up.

Benjamin: [laugh] yeah, you should use a managed service.

Corey: Yeah.

Benjamin: So, it’s complicated, but even with managed services, you end up in situations where you have downtime, you have maintenance windows. And with Postgres, especially—and other databases as well—especially with Postgres, one of the biggest concerns you have is major version upgrades, right? So, if I want to go from Postgres 12 to 13, 13 to 14, I can’t do that live. I can’t have a single cluster that is streaming one Postgres version to another Postgres version, right?

So, every year, people want to put things off for two years, three years sometimes—which is obviously not to their benefit—you have this maintenance, you have some sort of downtime, where you perform a Postgres upgrade. At EDB, we’ve got—so this is a big problem, this is a problem for us. We’re involved in the Postgres community. We know this is challenging. That’s just a well-known thing. Some of the folks that are working EDB are folks who worked on the Postgres logical replication tech, which arrived in Postgres 10. Logical replication is really a nice tool for doing things like change data capture, you can do Walter JSON, all these types of things are based on logical replication tech.

It’s not really a thing, at least, the code that’s in Postgres itself doesn’t really support high availability, though. It’s not really something that you can use to build a leader-follower type cluster on top of. We have some techs, some proprietary tech within EDB that used to be called bi-directional replication. There used to be an open-source project called bi-directional replication. This is a kind of a descendant of that. It’s now called Postgres Distributed, or EDB Postgres Distributed is the product name. And that tech actually allows us—because it’s based on logical replication—allows us to do multiple major versions at the same time, right? So, we can upgrade one node in a cluster to Postgres 14, while the other nodes in the clusters are at Postgres 13. We can then upgrade the next node. We can support these types of operations in a kind of wide range of maintenance operations without taking a cluster down from maintenance.

So, there’s a lot of interesting opportunities here when we start to say, well, let’s step back from what your typical assumptions are for Postgres streaming replication. Give ourselves a little bit more freedom by using logical replication instead of physical streaming replication. And then, what type of services, and what type of patterns can we build on top of that, that ultimately help customers build, whether it’s faster databases, more highly available databases, so on and so forth.

Corey: Let’s face it, on-call firefighting at 2am is stressful! So there’s good news and there’s bad news. The bad news is that you probably can’t prevent incidents from happening, but the good news is that incident.io makes incidents less stressful and a lot more valuable. incident.io is a Slack-native incident management platform that allows you to automate incident processes, focus on fixing the issues and learn from incident insights to improve site reliability and fix your vulnerabilities. Try incident.io, recover faster and sleep more.

Corey: One approach that I took for, I guess you could call it backup sort of, was intentionally staggering replication between the primary and the replica about 15 minutes or so. So, if I drop a production table or something like that, I have 15 short minutes to realize what has happened and sever the replication before it is now committed to the replica and now I’m living in hell. It felt like this was not, like, option A, B, or C, or the right way to do things. But given that meeting customers where they are as important, is that the sort of thing that you support with BigAnimal, or do you try to talk customers into not being ridiculous?

Benjamin: That’s not something we support now. It’s not actually something that I hear that many asks for these days. It’s kind of interesting, that’s a pattern that I’ve run into a lot in the past.

Corey: I was an ancient, grumpy sysadmin. Again, I’m dating myself here. These days, I just store everything at DNS text records, and it’s way easier. But I digress.

Benjamin: [laugh] yeah, it’s something that we see a lot for and we had support for a point-in-time restore, like pretty much anybody else in the business at this point. And that’s usually the, “I fat-fingered something,” type response. Honestly, I think there’s room to be a bit more flexible and room to do some more interesting things. I think RDS is setting a bar and a lot of database services out there and kind of just meeting that bar. And we all kind of need to be pushing a little bit more into more interesting spaces and figuring out how to get customers more value, get customers to get more out of their money for the database, honestly.

Corey: One of the problems we tend to see, in the database ecosystem at large, without naming names or companies or anything like that, is that it’s a pretty thin and blurry line between database advocate, database evangelist, and database zealot. Where it feels like instead, we’re arguing about religion more than actual technical constraints and concerns. So, here’s a fun question that hopefully isn’t too much of a gotcha. But what sort of workloads would you actively advise someone not to use BigAnimal for in the database world? But yes, again, if you try to run a DNS server, it’s probably not fit for purpose without at least a shim in the way there. But what sort of workloads are you not targeting that a customer is likely to have a relatively unfortunate time with?

Benjamin: Large-scale analytical workloads is the easy answer to that, right? If you’ve got a problem where you’re choosing between Postgres and Snowflake, you’re seriously considering—you actually have as much data that you seriously be considering Snowflake? You probably don’t want to be using Postgres, right? You want to be using something that’s column, or you want to be using a query planner that really understands a columnar layout that’s going to get you the sorts of performance that you need for those analytical workloads. We don’t try to touch that space.

Corey: Yeah, we’re doing some of that right now with just the sheer volume of client AWS bills we have. We don’t really need a relational model for a lot of it. And Athena is basically fallen down on the job in some cases, and, “Oh, do you want to use Redshift, that’s basically Postgres.” It’s like, “Yeah, it’s Postgres, if it decided to run on bars of gold.” No, thank you. It just becomes this ridiculously overwrought solution for what feels like it should be a lot similar. So, it’s weird, six months ago or so I wouldn’t have had much of an idea what you’re talking about. I see it a lot better now. Generally, by virtue of trying to do something the precise wrong way that someone should.

Benjamin: Right. Yeah, exactly. I think there’s interesting room for Postgres to expand here. It’s not something that we’re actively working on. I’m not aware of a lot happening in the community that Postgres is, for better or worse, extremely extensible, right? And if you see the JSON-supported Postgres, it didn’t exist, I don’t know, five, six years ago. And now it’s incredibly powerful. It’s incredibly flexible. And you can do a lot of quote-unquote, schemaless stuff straight in Postgres. Or you look at PostGIS, right, for doing GIS geographical data, right? That’s really a fantastic integration directly in the database.

Corey: Yeah, before that people start doing ridiculous things almost looks similar to a graph database or a columnar store somehow, and yeah.

Benjamin: Yeah, exactly. I think sometimes somebody will do a good column store that’s an open-source deeply integrated into Postgres, rather than—

Corey: I’ve seen someone build one on top of S3 bucket with that head, a quarter of a trillion objects in it. Professional advice, don’t do that.

Benjamin: [laugh]. Unless you’re Snowflake. So, I mean, it’s something that I’d like to see Postgres expand into. I think that’s an interesting space, but not something that, at least especially for BigAnimal, and frankly, for a lot of EDB customers. It’s not something we’re trying to push people toward.

Corey: One thing that I think we are seeing a schism around is the idea that some vendors are one side of it, some are on the other, where on the one side, you have, oh, every workload should have a bespoke, purpose-built database that is exactly for this type of workload. And the other school of thought is you should generally buy us for a general-purpose database until you have a workload that is scaled and significant to a point where running that on its own purpose-built database begins to make sense. I don’t necessarily think that is a binary choice, where do you tend to fall on that spectrum?

Benjamin: I think everybody should use Postgres. And I say not just because I work in a Postgres company.

Corey: Well, let’s be clear. Before this, you were at IBM for five years working on a whole bunch of database stuff over there, not just Postgres. And you, so far, have not struck me as the kind of person who’s like, “Oh, so what’s your favorite database?” “The one that pays me.” We’ve met people like that, let’s be very clear. But you seem very even-handed in those conversations.

Benjamin: Yeah, I got my start in databases, actually, with Apache CouchDB. I am a committer on CouchDB. I worked on a managed at CouchDB service ten years ago. At IBM, I worked on something in nine different open-source databases and managed services. But I love having conversations about, like, well, I’ve got this workload, should I use Postgres, rr should I use Mongo, should I use Cassandra, all of those types of discussions. Frankly, though, I think in a lot of cases people are—they don’t understand how much power they’re missing out on if they don’t choose a relational database. If they don’t understand the relational model well enough to understand that they really actually want that. In a lot of cases, people are also just over-optimizing too early, right? It’s just going to be much faster for them to get off the ground, get product in customers hands, if they start with something that they don’t have to think twice about. And they don’t end up with this architecture with 45 different databases, and there’s only one guy in the company that knows how to manage the whole thing.

Corey: Oh, the same story of picking a cloud provider. It’s, “Okay, you hire a team, you’re going to build a thing. Which cloud provider do you pick?” Every cloud provider has a whole matrix and sales deck, and the rest. The right answer, of course, is the one your team’s already familiar with because learning a new cloud provider while trying not to run out of money at your startup, can’t really doesn’t work super well.

Benjamin: Exactly. Yeah.

Corey: One thing that I think has been sort of interesting, and when I saw it, it was one of those, “Oh, I sort of like them.” Because I had that instinctive reaction and I don’t think I’m alone in this. As of this recording a couple of weeks ago, you folks received a sizable investment from private equity. And default reaction to that is, “Oh, well, I guess I put a fork in the company, they’re done.” Because the narrative is that once private equity takes an investment, well, that company’s best days are probably not in front of it. Now, the counterpoint is that this is not the first time private equity has invested in EDB, and you folks from what I can tell are significantly better than you were when I was your customer a decade ago. So clearly, there is something wrong with that mental model. What am I missing?

Benjamin: Yeah. Frankly, I don’t know. I’m no expert in funding models and all of those sorts of things. I will say that my experience has been what I’ve seen at EDB, has definitely been that maybe there’s private equity, and then there’s private equity. We’re in this to build better products and become a better product company. We were previously owned by a private equity firm for the past four years or so. And during the course of those four years, we brought on a bunch of folks who were very product-focused, new leadership. We made a significant acquisition of a company called 2ndQuadrant, which they employed a lot of the European best Postgres company. Now, they’re part of EDB and most of them have stayed with us. And we built the managed cloud service, right? So, this is a pretty significant—private equity company buying us to invest in the company. I’m optimistic that that’s what we’re looking at going forward.

Corey: I want to be clear as well, I’m not worried about what I normally would be in a private equity story about this, where they’re there to save money and cut costs, and, “Do we really need all these database replicas floating around,” and, “These backups, seems like that’s something we don’t need.” You have, at last count, 32 Postgres contributors, 7 Postgres committers, and 3 core members. All of whom would run away screaming loudly and publicly, in the event that such a thing were taking place. Of all the challenges and concerns I might have about someone running a cloud service in the modern day. I do not have any fear that you folks are not doing what will very clearly be shown to be the right thing by your customers for the technology that you’re building on top of. That is not a concern. There are companies I do not have that confidence in, to be clear.

Benjamin: Yeah, I’m glad to hear that. I’m a hundred percent on board as well. I work here, but I think we’re doing the right thing, and we’re going to be doing great stuff going forward.

Corey: One last topic I do want to get into a little bit is, on some level, launching in this decade, a cloud-hosted database offering at a time when Amazon—whose product strategy of yes is in full display—it seems like something ridiculous, that is not necessarily well thought out that why would you ever try to do this? Now, I will temper that by the fact that you are clearly succeeding in this direction. You have customers who say nice things about you, and the reviews have been almost universally positive anywhere I can see things. The negative ones are largely complaining about databases, which I admit might be coming from me.

Benjamin: Right, it is a crowded space. There’s a lot of things happening. Obviously, Amazon, Microsoft, Google are doing great things, both—

Corey: Terrible things, but great, yes. Yes.

Benjamin: [laugh] right, there’s good products coming in. I think AlloyDB is not necessarily a great product. I haven’t used it myself yet, but it’s an interesting step in the direction. I’m excited to see development happening. But at the end of the day, we’re a database company. Our focus is on building great databases and supporting great databases. We’re not entering this business to try to take on Amazon from an infrastructure point of view. In fact, the way that we’re structuring the product is really to try to get the strengths of both worlds. We want to give customers the ability to get the most out of the AWS or Azure infrastructure that they can, but come to us for their database.

Frankly, we know Postgres better than anybody else. We have a greater ability to get bugs fixed in Postgres than anybody else. We’ve got folks working on the database in the open. We got folks working on the database proprietary for us. So, we give customers things like break/fix support on that database. If there is a bug in Postgres, there’s a bug in the tech that sits around Postgres. Because obviously, Postgres is not a batteries-included system, really. We’re going to fix that for you. That’s part of the contract that we’re giving to our customers. And I know a lot of smaller companies maybe haven’t been burned by this sort of thing very much. We start to talk about enterprise customers and medium, larger-scale customers, this starts to get really valuable. The ability to have assurance on top of your open-source product. So, I think there’s a lot of interesting things there, a lot of value that we can provide there.

I think also that I talked a little bit about this earlier, but like the box, this sort of RDS-shaped box, I think is a bit too small. There’s an opportunity for smaller players to come in and try to push the boundaries of that. For example, giving customers more support by default to do a good job using their database. We have folks on board that can help consult with customers to say, “No, you shouldn’t be designing your schemas that way. You should be designing your schemas this way. You should be using indexes here,” that sort of stuff. That’s been part of our business for a long time. Now, with a managed service, we can bake that right into the managed service. And that gives us the ability to kind of make that—you talk about shared responsibility between the service writer and the customer—we can change the boundaries of that shared responsibility a little bit, so that customers can get more value out of the managed database service than they might expect otherwise.

Corey: There aren’t these harsh separations and clearly defined lines across which nothing shall pass, when it makes sense to do that in a controlled responsible way.

Benjamin: Right, exactly. Some of that is because we’re a database company, and some of that is because, frankly, we’re much smaller.

Corey: I’ll take it a step further beyond that, as well, that I have seen this pattern evolve a number of times where you have a customer running databases on EC2, and their AWS account managers suggests move to RDS. So, they do. Then, move to Aurora. So, they do. Then, I move this to DynamoDB. At which point, it’s like, what do you think your job is here, exactly? Because it seems like every time we move databases, you show up in a nicer car. So, what exactly is the story here, and what are the incentives? Where it just feels like there is a, “Whatever you’re doing is not the way that it should be done. So, it’s time to do, yet, another migration.”

There’s something to be said for companies who are focused around a specific aspect of things. Then once that is up and working and running, great. Keep on going. This is fine. As opposed to trying to chase the latest shiny, on some level. I have a big sense of, I guess, affinity for companies that wind up knowing where they start, and most notably, where they stop.

Benjamin: Yeah, I think that’s a really good point. I don’t think that we will be building an application platform anytime soon.

Corey: “We’re going to run Lambda functions on top of a database.” It’s like, “Congratulations. That is the weirdest stored procedure I can imagine this week, but I’m sure we can come up with a worse one soon.”

Benjamin: Exactly.

Corey: I really want to thank you for taking the time to speak with me so much about how you’re thinking about this, and what you’ve been building over there. If people want to learn more, where’s the best place to go to find you?

Benjamin: biganimal.com.

Corey: Excellent. We will throw a link to that in the show notes and it only just occurred to me that the Postgres mascot is an elephant, and now I understand why it’s called BigAnimal. Yeah, that’s right. He who laughs last, thinks slowest, and today, that’s me. I really want to thank you for being so generous with your time. I appreciate it.

Benjamin: Thank you. I really appreciate it.

Corey: Benjamin Anderson, CTO for Cloud at EDB. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an angry comment that you then wind up stuffing into a SQLite database, converting to Base64, and somehow stuffing into the comment field.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Brandon

Brandon West was raised in part by video games and BBSes and has been working on web applications since 1999. He entered the world of Developer Relations in 2011 as an evangelist for a small startup called SendGrid and has since held leadership roles at companies like AWS. At Datadog, Brandon is focused on helping developers improve the performance and developer experience of the things they build. He lives in Seattle where enjoys paddle-boarding, fishing, and playing music.

Links Referenced:

  • Datadog: https://www.datadoghq.com/
  • Twitter: https://twitter.com/bwest

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Honeycomb. When production is running slow, it’s hard to know where problems originate. Is it your application code, users, or the underlying systems? I’ve got five bucks on DNS, personally. Why scroll through endless dashboards while dealing with alert floods, going from tool to tool to tool that you employ, guessing at which puzzle pieces matter? Context switching and tool sprawl are slowly killing both your team and your business. You should care more about one of those than the other; which one is up to you. Drop the separate pillars and enter a world of getting one unified understanding of the one thing driving your business: production. With Honeycomb, you guess less and know more. Try it for free at honeycomb.io/screaminginthecloud. Observability: it’s more than just hipster monitoring.

Corey: This episode is sponsored in part by our friends at Fortinet. Fortinet’s partnership with AWS is a better-together combination that ensures your workloads on AWS are protected by best-in-class security solutions powered by comprehensive threat intelligence and more than 20 years of cybersecurity experience. Integrations with key AWS services simplify security management, ensure full visibility across environments, and provide broad protection across your workloads and applications. Visit them at AWS re:Inforce to see the latest trends in cybersecurity on July 25-26 at the Boston Convention Center. Just go over to the Fortinet booth and tell them Corey Quinn sent you and watch for the flinch. My thanks again to my friends at Fortinet.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. My guest today is someone I’ve been trying to get on the show for years, but I’m very bad at, you know, following up and sending the messages and all the rest because we all struggle with our internal demons. My guest instead struggles with external demons. He is the team lead for developer experience and tools advocacy at what I can only assume is a Tinder for Pets style company, Date-A-Dog. Brendon West, thank you for joining me today.

Brandon: Hey, Corey, thanks for having me. I’m excited to be here. Finally, like you said, it’s been a couple of years. But glad that it’s happening. And yeah, I’m on the DevRel team at Datadog.

Corey: Yes, I’m getting a note here in the headset of breaking news coming in. Yes, you’re not apparently a dog dating company, you are a monitoring slash observability slash whatever the cool kids are calling it today telemetry outputer dingus nonsense. Anyone who has ever been to a community or corporate event has no doubt been tackled by one of the badge scanners that you folks have orbiting your booth, but what is it that you folks do?

Brandon: Well, the observability, the monitoring, the distributed tracing, all that stuff that you mentioned. And then a lot of other interesting things that are happening. Security is a big focus—InfoSec—so we’re adding some products around that, automated security monitoring, very cool. And then the sort of stuff that I’m representing is stuff that helps developers provide a better experience to their end-users. So, things like front-end monitoring, real-time user monitoring, synthetic testing of your APIs, whatever it might be.

Corey: Your path has been somewhat interesting because you—well, everyone’s path has been somewhat interesting; yours has been really interesting because back in 2011, you entered the world of developer relations, or being a DevReloper as I insist on calling it. And you were in a—you call it a small startup called SendGrid. Which is, on some level, hilarious from my point of view. I’ve been working with you folks—you folks being SendGrid—for many years now. I cared a lot about email once upon a time.

And now I send an email newsletter every week, that deep under the hood, through a couple of vendor abstraction layers is still SendGrid, and I don’t care about email because that’s something that I can pay someone else to worry about. You went on as well to build out DevRel teams at AWS. You decided okay, you’re going to take some time off after that. You went to a small scrappy startup and ah, nice. You could really do things right and you have a glorious half of the year and then surprise, you got acquired by Datadog. Congratu-dolances on that because now you’re right back in the thick of things at big company-style approaches. Have I generally nailed the trajectory of the past decade for you?

Brandon: Yeah, I think the broad strokes are all correct there. SendGrid was a small company when I joined, you know? There were 30 of us or so. So, got to see that grow into what it is today, which was super, super awesome. But other than that, yeah, I think that’s the correct path.

Corey: It’s interesting to me, in that you were more or less doing developer relations before that was really a thing in the ecosystem. And I understand the challenge that you would have in a place like SendGrid because that is large-scale email sending, transactional or otherwise, and that is something that by and large, has slipped below the surface level of awareness for an awful lot of folks in your target market. It’s, “Oh, okay, and then we’ll just have the thing send an email,” they say, hand-waving over what is an incredibly deep and murky pool. And understanding that is a hard thing requires a certain level of technical sophistication. So, you started doing developer relations for something that very clearly needed some storytelling chops. How did you fall into it originally?

Brandon: Well, I wanted to do something that let me use those storytelling chops, honestly. I had been writing code at an agency for coal mines and gold mines and really actively inserting evil into the world, power plants, and that sort of thing. And, you know, I went to school for English literature. I loved writing. I played in thrash metal bands when I was a kid, so I’ve been up on stage being cussed at and told that I suck. So I—

Corey: Oh, I get that conference talks all the time.

Brandon: Yeah, right? So, that’s why when people ask me to speak, I’m like, “Absolutely.” There’s no way I can bomb harder than I’ve bombed before. No fear, right? So yeah, I wanted to use those skills. I wanted to do something different.

And one of my buddies had a company that he had co-founded that was going through TechStars in Boulder. SendGrid was the first accelerator-backed company to IPO which is pretty cool. But they had gone through TechStars in 2009. They were looking for a developer evangelist. So, SendGrid was looking for developer evangelist and my friend introduced me said, “I think you’d be good at this. You should have a conversation.” My immediate thought was what the hell is a developer evangelist?

Corey: And what might a SendGrid be? And all the rest. Yes, it’s that whole, “Oh, how do I learn to swim?” Someone throws you off the end of the dock and then retrospect, it’s, “I don’t think they were trying to teach me how to swim.” Yeah. Hindsight.

Brandon: Yeah. It worked out great. I will say, though, that I think DevRel has been around for a long time, you know? The title has been around since the original Macintosh at Apple in 1980-ish. There’s a whole large part of the tech world that would like you to think that it’s new because of all the terrible things that their DevRel team did at Microsoft in the late-90s.

And you can go read all about this. There were trials about it. These documents were released to the public, James Plamondon is the lead architect of all of this nastiness. But I think there was then a concerted effort to memory-hole that and say, “No, DevRel is new and shiny.” And then Google came along and said, “Well, it’s not evangelism anymore. It’s advocacy.”

Corey: It’s not sysadmin work anymore. It’s SRE. It’s not on-prem, it’s Sparkling Kubernetes, et cetera, et cetera.

Brandon: Yeah, so there’s this sense in a lot of places that DevRel is new, but it’s actually been around a long time. And you can learn a lot from reading about the history and understanding it, something I’ve given a talk on and written a bit about. So.

Corey: My philosophy around developer relations for a while has been that in many cases, its biggest obstacle is the way that it is great at telling stories about fantastically complex, deeply technical things; it can tell stories about almost anything except itself. And I keep seeing similar expressions of the same problem again, and again, and again. I mean, AWS, where you worked, as an example: they love to talk about their developer advocates, and you read the job descriptions and these are high-level roles with sweeping responsibilities, broad basis of experience being able to handle things at a borderline executive level. And then they almost neuter the entire thing by slapping a developer advocate title on top of those people, which means that some of the people that would be most effectively served by talking to them will dismiss them as, “Well, I’m a director”—or a VP—“What am I going to do talking to a developer advocate?” It feels like there’s a swing and a miss as far as encapsulating the value that the function provides.

I want to be clear, I am not sitting here shitting on DevRel or its practitioners, I see a problem with how it [laugh] is being expressed. Now, feel free to argue with me and just scream at me for the next 20 minutes, and this becomes a real short show. But—

Brandon: [laugh].

Corey: —It’ll be great. Hit me.

Brandon: No, you’re correct in many ways, which makes me sad because these are the same conversations that I’ve been having for the 11, 12 years that I’ve been in DevRel now. And I thought we would have moved past this at some point, but the problem is that we are bad at advocating for advocacy. We do a bad job of relating to people about DevRel because we spend so much time worried about stuff that doesn’t really matter. And we get very loud voices in the echo chamber screaming about titles and evangelism versus advocate versus community manager, and which department you should report up to, and all of these things that ultimately don’t matter. And it just seems like bickering from the outside. I think that the core of what we do is super awesome. And I don’t think it’s very hard to articulate. It’s just that we don’t spend the time to do that.

Corey: It’s always odd to me when I talk to someone like, “Oh, you’re in DevRel. What does that mean?” And their immediate response is, “Well, it’s not marketing, I’ll tell you that.” It’s feels like there might be some trauma that is being expressed in some strange ways. I do view it as marketing, personally, and people who take umbrage at that don’t generally tend to understand what marketing is.

Yeah, you can look at any area of business or any function and judge it by some of the worst examples that we’ve all seen, but when someone tells me they work in sales, I don’t automatically assume that they are sending me horrifyingly passive-aggressive drip campaigns, or trying to hassle me in a car lot. It’s no, there’s a broad spectrum of people. Just like I don’t assume that you’re an engineer. And I immediately think, oh, you can’t solve FizzBuzz on a whiteboard. No, there’s always going to be a broad spectrum of experience.

Marketing is one of those awesome areas of business that’s dramatically misunderstood a lot. Similarly to the fact that, you know, DevRel can’t tell stories, you think marketing could tell stories about itself, but it’s still struggles, too, in a bunch of ways. But I do believe that even if they’re not one of the same, developer relations and marketing are aligned around an awful lot of things like being able to articulate value that is hard to quantify.

Brandon: I completely agree with that. And if I meet someone in DevRel that starts off the conversation by saying that they’re not in marketing, then I know they’re probably not that great at their job. I mean, I think there’s a place of tech hubris, where we want to disrespect anything that’s not a hard skill where it’s not putting zeros and ones into a chip—

Corey: And spoiler, they’re all very hard skills.

Brandon: [laugh]. Yeah. And so, first off, like, stop disrespecting marketing. It’s important; your business probably wouldn’t survive if you didn’t have it. And second of all, you’re not immune to it, right?

Like, Heartbleed had a logo and a name for vulnerability because tech people are so susceptible to it, right? People don’t just wake up and wait in line for three days for a new iPhone because tech marketing doesn’t work, right?

Corey: “Oh, tech marketing doesn’t work on me,” says someone who’s devoted last five years of their life to working on Kubernetes. Yeah, sure it doesn’t.

Brandon: Yeah exactly. So, that whole perspective is silly. I think part of the problem is that they don’t want to invest in learning how to communicate what they do to a marketing org. They don’t want to spend the time to say, “Here’s how the marketing world thinks, and here’s how we can fit into that perspective.” They want to come in and say, “Well, you don’t understand DevRel. Let me define DevRel for you and tell you what we do.” And all those sorts of things. It’s too prescriptive and less collaborative.

Corey: Anytime you start getting into the idea of metrics around how do you measure someone in a developer advocacy role, the answer is, “Well, your metrics that you’re using are wrong, and any metrics you use are wrong, and there’s no good way to do it.” And I am sympathetic to that. When I started this place, I knew that if I went to a bunch of events and did my thing, good things would happen for the business. And how did I articulate that? Gut feel, but when you own the place, you can do that.

Whereas when you are a function inside of another org, inside of another org, and you start looking at from the executive leadership position at these things, it’s, “Okay, so let me get this straight. You cost as much as an engineer, you cost as much as that again, in your expenses because you’re traveling all the time, you write zero production code, whenever people ask you what it is you do here, you have a very strange answer, and from what we can tell, it looks like you hang out with your friends in exotic locations, give a 15-minute talk from time to time that mentions our name at the beginning, and nothing else relevant to our business, and then you go around and the entire story is ‘just trust me, I’m adding value.’” Yeah, when it’s time to tighten belts and start cutting back, is it any wonder that the developer advocacy is often one of the first departments hit from that perspective?

Brandon: It doesn’t surprise me. I mean, I’ve been a part of DevRel teams where we had some large number of events that we had attended for the year—I think 450-something—and the director of the team was very excited to show that off, right, you should have seen the CFOs face when he heard that, right, because all he sees is outgoing dollar signs. Like, how much expense? What’s the ROI on 450 events?

Corey: Yeah, “450 events? That’s more than one a day. Okay, great. That’s a big number and I already know what we’re spending. Great. How much business came out of that?”

And that’s when the hemming and hawing starts. Like, well, sort of, and yadda—and yeah, it doesn’t present well in the language that they are prepared to speak. But marketing can tell those stories because they have for ages. Like, “Okay, how much business came from our Superbowl ad?” “I dunno. The point is, is that there’s a brand awareness play, there’s the chance to remain top of the mental stack when people think about this space. And over the next few months, we can definitely see there’s been a dramatic uptick in our business. Now, how do we attribute that back? Well, I don’t know.”

There’s a saying in marketing, that half of your marketing budget is wasted. Now, figuring out which half will spend the rest of your career, you’ll never get even close. Because people don’t know the journey that customers go through, not really. Even customers don’t often see it.

Take this podcast, for example. I have sponsors that I do love and appreciate who say things from time to time on this show. And people will hear it and occasionally will become customers of those sponsors. But very often, it’s, “Oh, I heard about that on the podcast. I’ll Google it when I get to work and then I’ll have a conversation with my team and we’ll agree to investigate that.”

And any UTM tracking has long since fallen by the wayside. You might get to that from discussions with users in their interview process, but very often, they won’t remember where it came up. And it’s one of those impossible to quantify things. Now, I sound like one of those folks where I’m trying to say, “Oh, buy sponsorships that you can never prove add value.” But that is functionally how advertising tends to work, back in the days before it spied on you.

Brandon: Yeah, absolutely. And we’ve added a bunch of instrumentation to allow us to try and put that multi-touch attribution model together after the fact, but I’m still not sure that that’s worth the squeeze, right? You don’t get much juice out. One of the problems with metrics in DevRel is that the things that you can measure are very production-focused. It’s how many talks did you give? How many audience members did you reach?

Some developer relations folks do actually write production code, so it might be how many of the official SDK that you support got downloaded? That can be more directly attributed to business impact, those sorts of things are fantastic. But a lot of it is kind of fuzzy and because it’s production-focused, it can lead to burnout because it’s disconnected from business impact. “It’s how many widgets did your line produce today?” “Well, we gave all these talks and we had 150,000 engaged developer hours.” “Well, cool, what was the business outcome?” And if you can’t answer that for your own team and for your own self in your role, that leads pretty quickly to burnout.

Corey: Anytime you start measuring something and grading people based on it, they’re going to optimize for what you measure. For example, I send an email newsletter out, at time of this recording, to 31,000 people every week and that’s awesome. I also periodically do webinars about the joys of AWS bill optimization, and you know, 50 people might show up to one of those things. Okay, well, from a broad numbers perspective, yeah, I’d much rather go and send something out to those 31,000, folks until you realize that the kind of person that’s going to devote half an hour, forty-five minutes to having a discussion with you about AWS bill optimization is far likelier to care about this to the point where they become a customer than someone who just happens to be in an audience for something that is orthogonally-related. And that is the trick because otherwise, we would just all be optimizing for the single biggest platforms out there if oh, I’m going to go talk at this conference and that conference, not because they’re not germane to what we do, but because they have more people showing up.

And that doesn’t work. When you see that even on the podcast world, you have Joe Rogan, as the largest podcast in the world—let’s not make too many comparisons in different ways because I don’t want to be associated with that kind of tomfoolery—but there’s a reason that his advertisers, by and large, are targeting a mass-market audience, whereas mine are targeting B2B SaaS, by and large. I’m not here shilling for various mattress companies. I’m instead talking much more about things that solve the kind of problem that listeners to this show are likely to have. It’s the old-school of thought of advertising, where this is a problem that is germane to a certain type of audience, and that certain type of audience listens to shows like this. That was my whole school of thought.

Brandon: Absolutely. I mean, the core value that you need to do DevRel, in my opinion is empathy. It’s all about what Maya Angelou said, right? “People may not remember what you said, but they’ll definitely remember how you made them feel.” And I found that to be incredibly true.

Like, the moments that I regret the most in DevRel are the times when someone that I’ve met and spent time with before comes up to have a conversation and I don’t remember them because I met 200 people that night. And then I feel terrible, right? So, those are the metrics that I use internally. It’s hearts and minds. It’s how do people feel? Am I making them feel empowered and better at their craft through the work that I do?

That’s why I love DevRel. If I didn’t get that fulfillment, I’d go write code again. But I don’t get that sense of satisfaction, and wow, I made an impact on this person’s trajectory through their career that I do from DevRel. So.

Corey: I come bearing ill tidings. Developers are responsible for more than ever these days. Not just the code that they write, but also the containers and the cloud infrastructure that their apps run on. Because serverless means it’s still somebody’s problem. And a big part of that responsibility is app security from code to cloud. And that’s where our friend Snyk comes in. Snyk is a frictionless security platform that meets developers where they are - Finding and fixing vulnerabilities right from the CLI, IDEs, Repos, and Pipelines. Snyk integrates seamlessly with AWS offerings like code pipeline, EKS, ECR, and more! As well as things you’re actually likely to be using. Deploy on AWS, secure with Snyk. Learn more at Snyk.co/scream That’s S-N-Y-K.co/scream

Corey: The way that I tend to see it, too, is that there’s almost a bit of a broadening of DevRel. And let’s be clear, it’s a varied field with a lot of different ways to handle that approach. I’m have a terrible public speaker, so I’m not going to ever succeed in DevRel. Well, that’s certainly not true. People need to write blog posts; people need to wind up writing some of the sample code, in some cases; people need to talk to customers in a small group environment, as opposed to in front of 3000 people and talk about the things that they’re seeing, and the rest.

There’s a broad field and different ways that it applies. But I also see that there are different breeds of developer advocate as well. There are folks, like you for example. You and I have roughly the same amount of time in the industry working on different things, whereas there’s also folks who it seems like they graduate from a boot camp, and a year later, they’re working in a developer advocacy role. Does that mean that they’re bad developer advocates?

I don’t think so, but I think that if they try and present things the same way that you were I do from years spent in the trenches working on these things, they don’t have that basis of experience to fall back on, so they need to take a different narrative path. And the successful ones absolutely do.

Brandon: Yeah.

Corey: I think it’s a nuanced and broad field. I wish that there was more acceptance and awareness of that.

Brandon: That’s absolutely true. And part of the reason people criticize DevRel and don’t take it seriously, as they say, “Well, it’s inconsistent. This org, it reports to product; or, this org, it reports up to marketing; this other place, it’s part of engineering.” You know, it’s poorly defined. But I think that’s true of a lot of roles in tech.

Like, engineering is usually done a different way, very differently at some orgs compared to others. Product teams can have completely different methodologies for how they track and manage and estimate their time and all of those things. So, I would like to see people stop using that as a cudgel against the whole profession. It just doesn’t make any sense. At the same time, two of the best evangelist I ever hired were right out of university, so you’re completely correct.

The key thing to keep in mind there is, like, who’s the audience, right, because ultimately, it’s about building trust with the audience. There’s a lot of rooms where if you and I walk into the room; if it’s like a college hackathon, we’re going to have a—[laugh], we’re going to struggle.

Corey: Yeah, we have some real, “Hello fellow kids,” energy going on when we do that.

Brandon: Yeah. Which is also why I think it’s incredibly important for developer relations teams to be aware of the makeup of their team. Like, how diverse is your team, and how diverse are the audiences you’re speaking to? And if you don’t have someone who can connect, whether it’s because of age or lived experience or background, then you’re going to fail because like I said that the number one thing you need to be successful in this role is empathy, in my opinion.

Corey: I think that a lot of the efforts around a lot of this—trying to clarify what it is—some cases gone in well, I guess I’m going to call it the wrong direction. And I know that sounds judgy and I’m going to have to live with that, I suppose, but talk to me a bit about the, I guess, rebranding that we’ve seen in some recent years around developer advocates. Specifically, like, I like calling folks DevRelopers because it’s cutesy, it’s a bit of a portmanteau. Great. But it’s also not something I seriously suggest most people put on business cards.

But there are people who are starting to, I think, take a similar joke and actually identify with it where they call themselves developer avocados, which I don’t fully understand. I have opinions on it, but again, having opinions that are not based in data is something I try not to start shouting from the rooftops wherever I can. You live in that world a lot more posted than I do, where do you stand?

Brandon: So, I think it was well-intentioned and it was an attempt to do some of the awareness and brand building for DevRel, broadly, that we had lacked. But I see lots of problems with it. One, we already struggle to be taken seriously in many instances, as we’ve been discussing, and I don’t think we do ourselves any favors by giving ourselves cutesy nicknames that sort of infantilize the role like I can’t think of any other job that has a pet name for the work that they do.

Corey: Yeah. The “ooh-woo accounting”. Yeah, I sort of don’t see that happening very often in most business orgs.

Brandon: Yeah. It’s strange to me at the same time, a lot of the people who came up with it and popularized it are people that I consider friends and good colleagues. So hopefully, they won’t be too offended, but I really think that it kind of set us back in many ways. I don’t want to represent the work that I do with an emoji.

Corey: Funny, you bring that up. As we record this through the first recording, I have on my new ridiculous desktop computer thing from Apple, which I have named after a—you know, the same naming convention that you would expect from an AWS region—it’s us-shitpost-one. Instead of the word shit, it has the poop emoji. And you’d be amazed at the number of things that just melt when you start trying to incorporate that. GitHub has a problem with that being the name of an SSH key, for example.

I don’t know if I’ll keep it or I’ll just fall back to just spelling words out, but right now, at least, it really is causing all kinds of strange computer problems. Similarly, it causes strange cultural problems when you start having that dissonance and seeing something new and different like that in a business context. Because in some cases, yeah, it helps you interact with your audience and build rapport; in many others, it erodes trust and confidence that you know what you’re talking about because people expect things to be cast a certain way. I’m not saying they’re right. There’s a shitload of bias that bakes into that, but at the same time, I’d like to at least bias for choosing when and where I’m going to break those expectations.

There’s a reason that increasingly, my Duckbillgroup.com website speaks in business terms, rather than in platypus metaphors, whereas lastweekinaws.com, very much leans into the platypus. And that is the way that the branding is breaking down, just because people expect different things in different places.

Brandon: Yeah and, you know, this framing matters. And I’ve gone through two exercises now where I’ve helped rename an evangelism team to an advocacy team, not because I think it’s important to me—it’s a bunch of bikeshedding—but it has external implications, right? Especially evangelism, in certain parts of the world, has connotations. It’s just easier to avoid those. And how we present ourselves, the titles that we choose are important.

I wish we would spend way less time arguing about them, you know, advocacy has won evangelism, don’t use it. DevRel, if you don’t want to pick one, great. DevRel is broader umbrella. If you’ve got community managers, people who can’t write code that do things involving your events or whatever, program managers, if they’re on your team, DevRel, great description. I wish we could just settle that. Lots of wasted air discussing that one.

Corey: Constantly. It feels like this is a giant distraction that detracts from the value of DevRel. Because I don’t know about you, but when I pick what I want to do next in my career, the things I want to explain to people and spend that energy on are never, I want to explain what it is that I do. Like I’ve never liked those approaches where you have to first educate someone before they’re going to be in a position where they want to become your customer.

I think, honestly, that’s one of the things that Datadog has gotten very right. One of the early criticisms lobbed against Datadog when it first came out was, “Oh, this is basically monitoring by Fisher-Price.” Like, “This isn’t the deep-dive stuff.” Well yeah, but it turns out a lot of your buying audience are fundamentally toddlers with no visibility into what’s going on. For an awful lot of what I do, I want it to be click, click, done.

I am a Datadog customer for a reason. It’s not because I don’t have loud and angry opinions about observability; it’s because I just want there to be a dashboard that I can look at and see what’s working, what’s not, and do I need to care about things today? And it solves that job admirably because if I have those kinds of opinions about every aspect, I’m never going to be your customer anyway, or anyone’s customer. I’m going to go build my own and either launch a competitor or realize this is my what I truly love doing and go work at a company in this space, possibly yours. There’s something to be said for understanding the customer journey that those customers do not look like you.

And I think that’s what’s going on with a lot of the articulation around what developer relations is or isn’t. The people on stage who go to watch someone in DevRel give a talk, do not care, by and large, what DevRel is. They care about the content that they’re about to hear about, and when the first half of it is explaining what the person’s job is or isn’t, people lose interest. I don’t even like intros at the beginning of a talk. Give me a hook. Talk for 45 seconds. Give me a story about why I should care before you tell me who you are, what your credentials are, what your job title is, who you work for. Hit me with something big upfront and then we’ll figure it out from there.

Brandon: Yeah, I agree with you. I give this speaking advice to people constantly. Do not get up on stage and introduce yourself. You’re not a carnival hawker. You’re not trying to get people to roll up and see the show.

They’re already sitting in the seat. You’ve established your credibility. If they had questions about it, they read your abstract, and then they went and checked you out on LinkedIn, right? So, get to the point; make it engaging and entertaining.

Corey: I have a pet theory about what’s going on in some cases where, I think, on some level, it’s an outgrowth of an impostor-syndrome-like behavior, where people don’t believe that they deserve to be onstage talking about things, so they start backing up their bona fides to almost reassure themselves because they don’t believe that they should be up there and if they don’t believe it, why would anyone else. It’s the wrong approach. By holding the microphone, you inherently deserve to hold the microphone. And go ahead and tell your story. If people care enough to dig into you and who you are and well, “What is this person’s background, really?” Rest assured the internet is pretty easy to use these days, people will find out. So, let them do that research if they care. If they don’t, then there’s an entire line of people in this world who are going to dislike you or say you’re not qualified for what it is you’re doing or you don’t deserve it. Don’t be in that line, let alone at the front of it.

Brandon: So, you mentioned imposter syndrome and it got me thinking a little bit. And hopefully this doesn’t offend anyone, but I kind of starting to think that imposter syndrome is in many ways invented by people to put the blame on you for something that’s their fault. It’s like a carbon footprint to the oil and gas industry, right? These companies can’t provide you psychological safety and now they’ve gone and convinced you that it’s your fault and that you’re suffering from this syndrome, rather than the fact that they’re not actually making you feel prepared and confident and ready to get up on that stage, even if it’s your first time giving a talk, right?

Corey: I hadn’t considered it like that before. And again, I do tend to avoid straying into mental health territory on this show because I’m not an—

Brandon: Yes.

Corey: Expert. I’m a loud, confident white guy in tech. My failure mode is a board seat and a book deal, but I am not board-certified, let’s be clear. But I think you’re onto something here because early on in my career, I was very often faced with a whole lot of nebulous job description-style stuff and I was never sure if I was working on the right thing. Now that I’m at this stage of my career, and as you become more senior, you inherently find yourselves in roles, most of the time, that are themselves mired in uncertainty. That is, on some level, what seniority leads to.

And that’s fine, but early on in your career, not knowing if you’re succeeding or failing, I got surprise-fired a number of times when I thought I was doing great. There are also times that I thought I was about to be fired on the spot and, “Come on in; shut the door.” And yeah, “Here’s a raise because you’re just killing it.” And it took me a few years after that point to realize, wait a minute. They were underpaying me. That’s what that was, and they hope they didn’t know.

But it’s that whole approach of just trying to understand your place in the world. Do I rock? Do I suck? And it’s that constant uncertainty and unknowing. And I think companies do a terrible job, by and large, of letting people know that they’re okay, they’re safe, and they belong.

Brandon: I completely agree. And this is why I would strongly encourage people—if you have the privilege—please do not work at a company that does not want you to bring your whole self to work, or that bans politics, or however they want to describe it. Because that’s just a code word for we won’t provide you psychological safety. Or if they’re going to, it ends at a very hard border somewhere between work and life. And I just don’t think anyone can be successful in those environments.

Corey: I’m sure it’s possible, but it does bias for folks who, frankly, have a tremendous amount of privilege in many respects where I mentioned about, like, I’m a white dude in tech—you are too—and when we say things, we are presumed competent and people don’t argue with us by default. And that is a very easy to forget thing. Not everyone who looks like us is going to have very similar experiences. I have gotten it hilariously wrong before when I gave talks on how to wind up negotiating for salaries, for example, because well, it worked for me, what’s the problem? Yeah, I basically burned that talk with fire, redid the entire thing and wound up giving it with a friend of mine who was basically everything that I am not.

She was an attorney, she was a woman of color, et cetera, et cetera. And suddenly, it was a much stronger talk because it wasn’t just, “How to Succeed for White Guys.” There’s value in that, but you also have to be open to hearing that and acknowledging that you were born on third; you didn’t hit a triple. There’s a difference. And please forgive the sports metaphor. They do not sound natural coming from me.

Brandon: [laugh]. I don’t think I have anything more interesting to add on that topic.

Corey: [laugh]. So, I really want to thank you for taking the time to speak with me today. If people want to learn more about what you’re up to and how you view the world, what’s the best place to find you.

Brandon: So, I’m most active on Twitter at @bwest, but you know, it’s a mix of things so you may or may not just get tech. Most recently, I’ve been posting about a—

Corey: Oh, heaven forbid you bring your whole self to school.

Brandon: Right? I think most recently, I’ve been posting about a drill press that I’m restoring. So, all kinds of fun stuff on there.

Corey: I don’t know it sounds kind of—wait for it—boring to me. Bud-dum-tiss.

Brandon: [laugh]. [sigh]. I can’t believe I missed that one.

Corey: You’re welcome.

Brandon: Well, done. Well, done. And then I also will be hiring for a couple of developer relations folks at Datadogs soon, so if that’s interesting and you like the words I say about how to do DevRel, then reach out.

Corey: And you can find all of that in the show notes, of course. I want to thank you for being so generous with your time. I really appreciate it.

Brandon: Hey, thank you, Corey. I’m glad that we got to catch up after all this time. And hopefully get to chat with you again sometime soon.

Corey: Brandon West, team lead for developer experience and tools advocacy at Datadog. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry and insulting comment that is talking about how I completely misunderstand the role of developer advocacy. And somehow that rebuttal features no fewer than 400 emoji shoved into it.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Chris

Chris Short has been a proponent of open source solutions throughout his over two decades in various IT disciplines, including systems, security, networks, DevOps management, and cloud native advocacy across the public and private sectors. He currently works on the Kubernetes team at Amazon Web Services and is an active Kubernetes contributor and Co-chair of OpenGitOps. Chris is a disabled US Air Force veteran living with his wife and son in Greater Metro Detroit. Chris writes about Cloud Native, DevOps, and other topics at ChrisShort.net. He also runs the Cloud Native, DevOps, GitOps, Open Source, industry news, and culture focused newsletter DevOps’ish.

Links Referenced:

  • DevOps’ish: https://devopsish.com/
  • EKS News: https://eks.news/
  • Containers from the Couch: https://containersfromthecouch.com
  • opengitops.dev: https://opengitops.dev
  • ChrisShort.net: https://chrisshort.net
  • Twitter: https://twitter.com/ChrisShort

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Coming back to us since episode two—it’s always nice to go back and see the where are they now type of approach—I am joined by Senior Developer Advocate at AWS Chris Short. Chris, been a few years. How has it been?

Chris: Ha. Corey, we have talked outside of the podcast. But it’s been good. For those that have been listening, I think when we recorded I wasn’t even—like, when was season two, what year was that? [laugh].

Corey: Episode two was first pre-pandemic and the rest. I believe—

Chris: Oh. So, yeah. I was at Red Hat, maybe, when I—yeah.

Corey: Yeah. You were doing Red Hat stuff, back when you got to work on open-source stuff, as opposed to now, where you’re not within 1000 miles of that stuff, right?

Chris: Actually well, no. So, to be clear, I’m on the EKS team, the Kubernetes team here at AWS. So, when I joined AWS in October, they were like, “Hey, you do open-source stuff. We like that. Do more.” And I was like, “Oh, wait, do more?” And they were like, “Yes, do more.” “Okay.”

So, since joining AWS, I’ve probably done more open-source work than the three years at Red Hat that I did. So, that’s kind of—you know, like, it’s an interesting point when I talk to people about it because the first couple months are, like—you know, my friends are like, “So, are you liking it? Are you enjoying it? What’s going on?” And—

Corey: Do they beat you with reeds? Like, all the questions people have about companies? Because—

Chris: Right. Like, I get a lot of random questions about Amazon and AWS that I don’t know the answer to.

Corey: Oh, when I started telling people, I fixed Amazon bills, I had to quickly pivot that to AWS bills because people started asking me, “Well, can you save me money on underpants?” It’s I—

Chris: Yeah.

Corey: How do you—fine. Get the prime credit card. It docks 5% off the bill, so there you go. But other than that, no, I can’t.

Chris: No.

Corey: It’s—

Chris: Like, I had to call my bank this morning about a transaction that I didn’t recognize, and it was from Amazon. And I was like, that’s weird. Why would that—

Corey: Money just flows one direction, and that’s the wrong direction from my employer.

Chris: Yeah. Like, what is going on here? It shouldn’t have been on that card kind of thing. And I had to explain to the person on the phone that I do work at Amazon but under the Web Services team. And he was like, “Oh, so you’re in IT?”

And I’m like, “No.” [laugh]. “It’s actually this big company. That—it’s a cloud company.” And they’re like, “Oh, okay, okay. Yeah. The cloud. Got it.” [laugh]. So, it’s interesting talking to people about, “I work at Amazon.” “Oh, my son works at Amazon distribution center,” blah, blah, blah. It’s like, cool. “I know about that, but very little. I do this.”

Corey: Your son works in Amazon distribution center. Is he a robot? Is normally my next question on that? Yeah. That’s neither here nor there.

So, you and I started talking a while back. We both write newsletters that go to a somewhat similar audience. You write DevOps’ish. I write Last Week in AWS. And recently, you also have started EKS News because, yeah, the one thing I look at when I’m doing these newsletters every week is, you know what I want to do? That’s right. Write more newsletters.

Chris: [laugh].

Corey: So, you are just a glutton for punishment? And, yeah, welcome to the addiction, I suppose. How’s it been going for you?

Chris: It’s actually been pretty interesting, right? Like, we haven’t pushed it very hard. We’re now starting to include it in things. Like we did Container Day; we made sure that EKS news was on the landing page for Container Day at KubeCon EU. And you know, it’s kind of just grown organically since then.

But it was one of those things where it’s like, internally—this happened at Red Hat, right—when I started live streaming at Red Hat, the ultimate goal was to do our product management—like, here’s what’s new in the next version thing—do those live so anybody can see that at any point in time anywhere on Earth, the second it’s available. Similar situation to here. This newsletter actually is generated as part of a report my boss puts together to brief our other DAs—or developer advocates—you know, our solutions architects, the whole nine yards about new EKS features. So, I was like, why can’t we just flip that into a weekly newsletter, you know? Like, I can pull from the same sources you can.

And what’s interesting is, he only does the meeting bi-weekly. So, there’s some weeks where it’s just all me doing it and he ends up just kind of copying and pasting the newsletter into his document, [laugh] and then adds on for the week. But that report meeting for that team is now getting disseminated to essentially anyone that subscribes to eks.news. Just go to the site, there’s a subscribe thing right there. And we’ve gotten 20 issues in and it’s gotten rave reviews, right?

Corey: I have been a subscriber for a while. I will say that it has less Chris Short personality—

Chris: Mm-hm.

Corey: —to it than DevOps’ish does, which I have to assume is by design. A lot of The Duckbill Group’s marketing these days is no longer in my voice, rather intentionally, because it turns out that being a sarcastic jackass and doing half-billion dollar AWS contracts can not to be the most congruent thing in the world. So okay, we’re slowly ameliorating that. It’s professional voice versus snarky voice.

Chris: Well, and here’s the thing, right? Like, I realized this year with DevOps’ish that, like, if I want to take a week off, I have to do, like, what you did when your child was born. You hired folks to like, do the newsletter for you, or I actually don’t do the newsletter, right? It’s binary: hire someone else to do it, or don’t do it. So, the way I structured this newsletter was that any developer advocate on my team could jump in and take over the newsletter so that, you know, if I’m off that week, or whatever may be happening, I, Chris Short, am not the voice. It is now the entire developer advocate team.

Corey: I will challenge you on that a bit. Because it’s not Chris Short voice, that’s for sure, but it’s also not official AWS brand voice either.

Chris: No.

Corey: It is clearly written by a human being who is used to communicating with the audience for whom it is written. And that is no small thing. Normally, when oh, there’s a corporate newsletter; that’s just a lot of words to say it’s bad. This one is good. I want to be very clear on that.

Chris: Yeah, I mean, we have just, like, DevOps’ish, we have sections, just like your newsletter, there’s certain sections, so any new, what’s new announcements, those go in automatically. So, like, that can get delivered to your inbox every Friday. Same thing with new blog posts about anything containers related to EKS, those will be in there, then Containers from the Couch, our streaming platform, essentially, for all things Kubernetes. Those videos go in.

And then there’s some ecosystem news as well that I collect and put in the newsletter to give people a broader sense of what’s going on out there in Kubernetes-land because let’s face it, there’s upstream and then there’s downstream, and sometimes those aren’t in sync, and that’s normal. That’s how Kubernetes kind of works sometimes. If you’re running upstream Kubernetes, you are awesome. I appreciate you, but I feel like that would cause more problems and it’s worse sometimes.

Corey: Thank you for being the trailblazers. The rest of us can learn from your misfortune.

Chris: [laugh]. Yeah, exactly. Right? Like, please file your bugs accordingly. [laugh].

Corey: EKS is interesting to me because I don’t see a lot of it, which is, probably, going to get a whole lot of, “Wait, what?” Moments because wait, don’t you deal with very large AWS bills? And I do. But what I mean by that is that EKS, until you’re using its Fargate expression, charges for the control plane, which rounds to no money, and the rest is running on EC2 instances running in a company’s account. From the billing perspective, there is no difference between, “We’re running massive fleets of EKS nodes.” And, “We’re managing a whole bunch of EC2 instances by hand.”

And that feels like an interesting allegory for how Kubernetes winds up expressing itself to cloud providers. Because from a billing perspective, it just looks like one big single-tenant application that has some really strange behaviors internally. It gets very chatty across AZs when there’s no reason to, and whatnot. And it becomes a very interesting study in how to expose aspects of what’s going on inside of those containers and inside of the Kubernetes environment to the cloud provider in a way that becomes actionable. There are no good answers for this yet, but it’s something I’ve been seeing a lot of. Like, “Oh, I thought you’d be running Kubernetes. Oh, wait, you are and I just keep forgetting what I’m looking at sometimes.”

Chris: So, that’s an interesting point. The billing is kind of like, yeah, it’s just compute, right? So—

Corey: And my insight into AWS and the way I start thinking about it is always from a billing perspective. That’s great. It’s because that means the more expensive the services, the more I know about it. It’s like, “IAM. What is that?” Like, “Oh, I have no idea. It’s free. How important could it be?” Professional advice: do not take that philosophy, ever.

Chris: [laugh]. No. Ever. No.

Corey: Security: it matters. Oh, my God. It’s like you’re all stars. Your IAM policy should not be. I digress.

Chris: Right. Yeah. Anyways, so two points I want to make real quick on that is, one, we’ve recently released an open-source project called Carpenter, which is really cool in my purview because it looks at your Kubernetes file and says, “Oh, you want this to run on ARM instance.” And you can even go so far as to say, right, here’s my limits, and it’ll find an instance that fits those limits and add that to your cluster automatically. Run your pod on that compute as long as it needs to run and then if it’s done, it’ll downsize—eventually, kind of thing—your cluster.

So, you can basically just throw a bunch of workloads at it, and it’ll auto-detect what kind of compute you will need and then provision it for you, run it, and then be done. So, that is one-way folks are probably starting to save money running EKS is to adopt Carpenter as your autoscaler as opposed to the inbuilt Kubernetes autoscaler. Because this is instance-aware, essentially, so it can say, like, “Oh, your massive ARM application can run here,” because you know, thank you, Graviton. We have those processors in-house. And you know, you can run your ARM64 instances, you can run all the Intel workloads you want, and it’ll right size the compute for your workloads.

And I’ll look at one container or all your containers, however you want to configure it. Secondly, the good folks over at Kubecost have opencost, which is the open-source version of Kubecost, basically. So, they have a service that you can run in your clusters that will help you say, “Hey, maybe this one notes too heavy; maybe this one notes too light,” and you know, give you some insights into Kubernetes spend that are a little bit more granular as far as usage and things like that go. So, those two projects right there, I feel like, will give folks an optimal savings experience when it comes to Kubernetes. But to your point, it’s just compute, right? And that’s really how we treat it, kind of, here internally is that it’s a way to run… compute, Kubernetes, or ECS, or any of those tools.

Corey: A fairly expensive one because ignoring entirely for a second the actual raw cost of compute, you also have the other side of it, which is in every environment, unless you are doing something very strange or pre-funding as a one-person startup in your spare time, your payroll costs will it—should—exceed your AWS bill by a fairly healthy amount. And engineering time is always more expensive than services time. So, for example, looking at EKS, I would absolutely recommend people use that rather than rolling their own because—

Chris: Rolling their own? Yeah.

Corey: —get out of that engineering space where your time is free. I assure you from a business context, it is not. So, there’s always that question of what you can do to make things easier for people and do more of the heavy lifting.

Chris: Yeah, and to your rather cheeky point that there’s 17 ways to run a container on AWS, it is answering that question, right? Like those 17 ways, like, how much of this do you want to run yourself, you could run EKS distro on EC2 instances if you want full control over your environment.

Corey: And then run IoT Greengrass core on top within that cluster—

Chris: Right.

Corey: So, I can run my own Lambda function runtime, so I’m not locked in. Also, DynamoDB local so I’m not locked into AWS. At which point I have gone so far around the bend, no one can help me.

Chris: Well—

Corey: Pro tip, don’t do that. Just don’t do that.

Chris: But to your point, we have all these options for compute, and specifically containers because there’s a lot of people that want to granularly say, “This is where my engineering team gets involved. Everything else you handle.” If I want EKS on Spot Instances only, you can do that. If you want EKS to use Carpenter and say only run ARM workloads, you can do that. If you want to say Fargate and not have anything to manage other than the container file, you can do that.

It’s how much does your team want to manage? That’s the customer obsession part of AWS coming through when it comes to containers is because there’s so many different ways to run those workloads, but there’s so many different ways to make sure that your team is right-sized, based off the services you’re using.

Corey: I do want to change gears a bit here because you are mostly known for a couple of things: the DevOps’ish newsletter because that is the oldest and longest thing you’ve been doing the time that I’ve known you; EKS, obviously. But when prepping for this show, I discovered you are now co-chair of the OpenGitOps project.

Chris: Yes.

Corey: So, I have heard of GitOps in the context of, “Oh, it’s just basically your CI/CD stuff is triggered by Git events and whatnot.” And I’m sitting here going, “Okay, so from where you’re sitting, the two best user interfaces in the world that you have discovered are YAML and Git.” And I just have to start with the question, “Who hurt you?”

Chris: [laugh]. Yeah, I share your sentiment when it comes to Git. Not so much with YAML, but I think it’s because I’m so used to it. Maybe it’s Stockholm Syndrome, maybe the whole YAML thing. I don’t know.

Corey: Well, it’s no XML. We’ll put it that way.

Chris: Thankfully, yes because if it was, I would have way more, like, just template files laying around to build things. But the—

Corey: And rage. Don’t forget rage.

Chris: And rage, yeah. So, GitOps is a little bit more than just Git in IaC—infrastructure as Code. It’s more like Justin Garrison, who’s also on my team, he calls it infrastructure software because there’s four main principles to GitOps, and if you go to opengitops.dev, you can see them. It’s version one.

So, we put them on the website, right there on the page. You have to have a declared state and that state has to live somewhere. Now, it’s called GitOps because Git is probably the most full-featured thing to put your state in, but you could use an S3 bucket and just version it, for example. And make it private so no one else can get to it.

Corey: Or you could use local files: copy-of-copy-of-this-thing-restored-parentheses-use-this-one-dot-final-dot-doc-dot-zip. You know, my preferred naming convention.

Chris: Ah, yeah. Wow. Okay. [laugh]. Yeah.

Corey: Everything I touch is terrifying.

Chris: Yes. Geez, I’m sorry. So first, it’s declarative. You declare your state. You store it somewhere. It’s versioned and immutable, like I said. And then pulled automatically—don’t focus so much on pull—but basically, software agents are applying the desired state from source. So, what does that mean? When it’s—you know, the fourth principle is implemented, continuously reconciled. That means those software agents that are checking your desired state are actually putting it back into the desired state if it’s out of whack, right? So—

Corey: You’re talking about agents running it persistently on instances, validating—

Chris: Yes.

Corey: —a checkpoint on a cron. How is this meaningfully different than a Puppet agent running in years past? Having spent I learned to speak publicly by being a traveling trainer for Puppet; same type of model, and in fact, when I was at Pinterest, we wound up having a fair bit—like, that was their entire model, where they would have—the Puppet’s code would live in an S3 bucket that was then copied down, I believe, via Git, and then applied to the instance on a schedule. Like, that sounds like this was sort of a early days GitOps.

Chris: Yeah, exactly. Right? Like so it’s, I like to think of that as a component of GitOps, right? DevOps, when you talk about DevOps in general, there’s a lot of stuff out there. There’s a lot of things labeled DevOps that maybe are, or maybe aren’t sticking to some of those DevOps core things that make you great.

Like the stuff that Nicole Forsgren writes about in books, you know? Accelerate is on my desk for a reason because there’s things that good, well-managed DevOps practices do. I see GitOps as an actual implementation of DevOps in an open-source manner because all the tooling for GitOps these days is open-source and it all started as open-source. Now, you can get, like, Flux or Argo—Argo, specifically—there’s managed services out there for it, you can have Flux and not maintain it, through an add-on, on EKS for example, and it will reconcile that state for you automatically. And the other thing I like to say about GitOps, specifically, is that it moves at the speed of the Kubernetes Audit Log.

If you’ve ever looked at a Kubernetes audit log, you know it’s rather noisy with all these groups and versions and kinds getting thrown out there. So, GitOps will say, “Oh, there’s an event for said thing that I’m supposed to be watching. Do I need to change anything? Yes or no? Yes? Okay, go.”

And the change gets applied, or, “Hey, there’s a new Git thing. Pull it in. A change has happened inGit I need to update it.” You can set it to reconcile on events on time. It’s like a cron or it’s like an event-driven architecture, but it’s combined.

Corey: How does it survive the stake through the heart of configuration management? Because before I was doing all this, I wasn’t even a T-shaped engineer: you’re broad across a bunch of things, but deep in one or two areas, and one of mine was configuration management. I wrote part of SaltStack, once upon a time—

Chris: Oh.

Corey: —due to a bunch of very strange coincidences all hitting it once, like, I taught people how to use Puppet. But containers ultimately arose and the idea of immutable infrastructure became a thing. And these days when we were doing full-on serverless, well, great, I just wind up deploying a new code bundle to the Lambdas function that I wind up caring about, and that is a immutable version replacement. There is no drift because there is no way to log in and change those things other than through a clear deployment of this as the new version that goes out there. Where does GitOps fit into that imagined pattern?

Chris: So, configuration management becomes part of your approval process, right? So, you now are generating an audit log, essentially, of all changes to your system through the approval process that you set up as part of your, how you get things into source and then promote that out to production. That’s kind of the beauty of it, right? Like, that’s why we suggest using Git because it has functions, like, requests and issues and things like that you can say, “Hey, yes, I approve this,” or, “Hey, no, I don’t approve that. We need changes.” So, that’s kind of natively happening with Git and, you know, GitLab, GitHub, whatever implementation of Git. There’s always, kind of—

Corey: Uh, JIF-ub is, I believe, the pronunciation.

Chris: JIF-ub? Oh.

Corey: Yeah. That’s what I’m—

Chris: Today, I learned. Okay.

Corey: Exactly. And that’s one of the things that I do for my lasttweetinaws.com Twitter client that I build—because I needed it, and if other people want to use it, that’s great—that is now deployed to 20 different AWS commercial regions, simultaneously. And that is done via—because it turns out that that’s a very long to execute for loop if you start down that path—

Chris: Well, yeah.

Corey: I wound up building out a GitHub Actions matrix—sorry a JIF-ub—actions matrix job that winds up instantiating 20 parallel builds of the CDK deploy that goes out to each region as expected. And because that gets really expensive with native GitHub Actions runners for, like, 36 cents per deploy, and I don’t know how to test my own code, so every time I have a typo, that’s another quarter in the jar. Cool, but that was annoying for me so I built my own custom runner system that uses Lambda functions as runners running containers pulled from ECR that, oh, it just runs in parallel, less than three minutes. Every time I commit something between I press the push button and it is out and running in the wild across all regions. Which is awesome and also terrifying because, as previously mentioned, I don’t know how to test my code.

Chris: Yeah. So, you don’t know what you’re deploying to 20 regions sometime, right?

Corey: But it also means I have a pristine, re-composable build environment because I can—

Chris: Right.

Corey: Just automatically have that go out and the fact that I am making a—either merging a pull request or doing a direct push because I consider main to be my feature branch as whenever something hits that, all the automation kicks off. That was something that I found to be transformative as far as a way of thinking about this because I was very tired of having to tweak my local laptop environment to, “Oh, you didn’t assume the proper role and everything failed again and you broke it. Good job.” It wound up being something where I could start developing on more and more disparate platforms. And it finally is what got me away from my old development model of everything I build is on an EC2 instance, and that means that my editor of choice was Vim. I use the VS Code now for these things, and I’m pretty happy with it.

Chris: Yeah. So, you know, I’m glad you brought up CDK. CDK gives you a lot of the capabilities to implement GitOps in a way that you could say, like, “Hey, use CDK to declare I need four Amazon EKS clusters with this size, shape, and configuration. Go.” Or even further, connect to these EKS clusters to RDS instances and load balancers and everything else.

But you put that state into Git and then you have something that deploys that automatically upon changes. That is infrastructure as code. Now, when you say, “Okay, main is your feature branch,” you know, things happen on main, if this were running in Kubernetes across a fleet of clusters or the globe-wide in 20 regions, something like Flux or Argo would kick in and say, “There’s been a change to source, main, and we need to roll this out.” And it’ll start applying those changes. Now, what do you get with GitOps that you don’t get with your configuration?

I mean, can you rollback if you ever have, like, a bad commit that’s just awful? I mean, that’s really part of the process with GitOps is to make sure that you can, A, roll back to the previous good state, B, roll forward to a known good state, or C, promote that state up through various environments. And then having that all done declaratively, automatically, and immutably, and versioned with an audit log, that I think is the real power of GitOps in the sense that, like, oh, so-and-so approve this change to security policy XYZ on this date at this time. And that to an auditor, you just hand them a log file on, like, “Here’s everything we’ve ever done to our system. Done.” Right?

Like, you could get to that state, if you want to, which I think is kind of the idea of DevOps, which says, “Take all these disparate tools and processes and procedures and culture changes”—culture being the hardest part to adopt in DevOps; GitOps kind of forces a culture change where, like, you can’t do a CAB with GitOps. Like, those two things don’t fly. You don’t have a configuration management database unless you absolutely—

Corey: Oh, you CAB now but they’re all the comments of the pull request.

Chris: Right. Exactly. Like, don’t push this change out until Thursday after this other thing has happened, kind of thing. Yeah, like, that all happens in GitHub. But it’s very democratizing in the sense that people don’t have to waste time in an hour-long meeting to get their five minutes in, right?

Corey: DoorDash had a problem. As their cloud-native environment scaled and developers delivered new features, their monitoring system kept breaking down. In an organization where data is used to make better decisions about technology and about the business, losing observability means the entire company loses their competitive edge. With Chronosphere, DoorDash is no longer losing visibility into their applications suite. The key? Chronosphere is an open-source compatible, scalable, and reliable observability solution that gives the observability lead at DoorDash business, confidence, and peace of mind. Read the full success story at snark.cloud/chronosphere. That's snark.cloud slash C-H-R-O-N-O-S-P-H-E-R-E.

Corey: So, would it be overwhelmingly cynical to suggest that GitOps is the means to implement what we’ve all been pretending to have implemented for the last decade when giving talks at conferences?

Chris: Ehh, I wouldn’t go that far. I would say that GitOps is an excellent way to implement the things you’ve been talking about at all these conferences for all these years. But keep in mind, the technology has changed a lot in the, what 11, 12 years of the existence of DevOps, now. I mean, we’ve gone from, let’s try to manage whole servers immutably to, “Oh, now we just need to maintain an orchestration platform and run containers.” That whole compute interface, you go from SSH to a Docker file, that’s a big leap, right?

Like, you don’t have bespoke sysadmins; you have, like, a platform team. You don’t have DevOps engineers; they’re part of that platform team, or DevOps teams, right? Like, which was kind of antithetical to the whole idea of DevOps to have a DevOps team. You know, everybody’s kind of in the same boat now, where we see skill sets kind of changing. And GitOps and Kubernetes-land is, like, a platform team that manages the cluster, and its state, and health and, you know, production essentially.

And then you have your developers deploying what they want to deploy in when whatever namespace they’ve been given access to and whatever rights they have. So, now you have the potential for one set of people—the platform team—to use one set of GitOps tooling, and your applications teams might not like that, and that’s fine. They can have their own namespaces with their own tooling in it. Like, Argo, for example, is preferred by a lot of developers because it has a nice UI with green and red dots and they can show people and it looks nice, Flux, it’s command line based. And there are some projects out there that kind of take the UI of Argo and try to run Flux underneath that, and those are cool kind of projects, I think, in my mind, but in general, right, I think GitOps gives you the choice that we missed somewhat in DevOps implementations of the past because it was, “Oh, we need to go get cloud.” “Well, you can only use this cloud.” “Oh, we need to go get this thing.” “Well, you can only use this thing in-house.”

And you know, there’s a lot of restrictions sometimes placed on what you can use in your environment. Well, if your environment is Kubernetes, how do you restrict what you can run, right? Like you can’t have an easily configured say, no open-source policy if you’re running Kubernetes. [laugh] so it becomes, you know—

Corey: Well, that doesn’t stop some companies from trying.

Chris: Yeah, that’s true. But the idea of, like, enabling your developers to deploy at will and then promote their changes as they see fit is really the dream of DevOps, right? Like, same with production and platform teams, right? I want to push my changes out to a larger system that is across the globe. How do I do that? How do I manage that? How do I make sure everything’s consistent?

GitOps gives you those ways, with Kubernetes native things like customizations, to make consistent environments that are robust and actually going to be reconciled automatically if someone breaks the glass and says, “Oh, I need to run this container immediately.” Well, that’s going to create problems because it’s deviated from state and it’s just that one region, so we’ll put it back into state.

Corey: It’ll be dueling banjos, at some point. You’ll try and doing something manually, it gets reverted automatically. I love that pattern. You’ll get bored before the computer does, always.

Chris: Yeah. And GitOps is very new, right? When you think about the lifetime of GitOps, I think it was coined in, like, 2018. So, it’s only four years old, right? When—

Corey: I prefer it to ChatOps, at least, as far as—

Chris: Well, I mean—

Corey: —implementation and expression of the thing.

Chris: —ChatOps was a way to do DevOps. I think GitOps—

Corey: Well, ChatOps is also a way to wind up giving whoever gets access to your Slack workspace root in production.

Chris: Mmm.

Corey: But that’s neither here nor there.

Chris: Mm-hm.

Corey: It’s yeah, we all like to pretend that’s not a giant security issue in our industry, but that’s a topic for another time.

Chris: Yeah. And that’s why, like, GitOps also depends upon you having good security, you know, and good authorization and approval processes. It enforces that upon—

Corey: Yeah, who doesn’t have one of those?

Chris: Yeah. If it’s a sole operation kind of deal, like in your setup, your case, I think you kind of got it doing right, right? Like, as far as GitOps goes—

Corey: Oh, to be clear, we are 11 people and we do have dueling pull requests and all the rest.

Chris: Right, right, right.

Corey: But most of the stuff I talk about publicly is not our production stuff, so it really is just me. Just as a point of clarity there. I’ve n—the 11 people here do not all—the rest of you don’t just sit there and clap as I do all the work.

Chris: Right.

Corey: Most days.

Chris: No, I’m sure they don’t. I’m almost certain they don’t clap… for you. I mean, they would—

Corey: No. No, they try and talk me out of it in almost every case.

Chris: Yeah, exactly. So, the setup that you, Corey Quinn, have implemented to deploy these 20 regions is kind of very GitOps-y, in the sense that when main changes, it gets updated. Where it’s not GitOps-y is what if the endpoint changes? Does it get reconciled? That’s the piece you’re probably missing is that continuous reconciliation component, where it’s constantly checking and saying, “This thing out there is deployed in the way I want it. You know, the way I declared it to be in my source of truth.”

Corey: Yeah, when you start having other people getting involved, there can—yeah, that’s where regressions enter. And it’s like, “Well, I know where things are so why would I change the endpoint?” Yeah, it turns out, not everyone has the state of the entire application in their head. Ideally it should live in—

Chris: Yeah. Right. And, you know—

Corey: —you know, Git or S3.

Chris: —when I—yeah, exactly. When I think about interactions of the past coming out as a new DevOps engineer to work with developers, it’s always been, will developers have access to prod or they don’t? And if you’re in that environment with—you’re trying to run a multi-billion dollar operation, and your devs have direct—or one Dev has direct access to prod because prod is in his brain, that’s where it’s like, well, now wait a minute. Prod doesn’t have to be only in your brain. You can put that in the codebase and now we know what is in your brain, right?

Like, you can almost do—if you document your code, well, you can have your full lifecycle right there in one place, including documentation, which I think is the best part, too. So, you know, it encourages approval processes and automation over this one person has an entire state of the system in their head; they have to go in and fix it. And what if they’re not on call, or in Jamaica, or on a cruise ship somewhere kind of thing? Things get difficult. Like, for example, I just got back from vacation. We were so far off the grid, we had satellite internet. And let me tell you, it was hard to write an email newsletter where I usually open 50 to 100 tabs.

Corey: There’s a little bit of internet out Californ-ie way.

Chris: [laugh].

Corey: Yeah it’s… it’s always weird going from, like, especially after pandemic; I have gigabit symmetric here and going even to re:Invent where I’m trying to upload a bunch of video and whatnot.

Chris: Yeah. Oh wow.

Corey: And the conference WiFi was doing its thing, and well, Verizon 5G was there but spotty. And well, yeah. Usual stuff.

Chris: Yeah. It’s amazing to me how connectivity has become so ubiquitous.

Corey: To the point where when it’s not there anymore, it’s what do I do with myself? Same story about people pushing back against remote development of, “Oh, I’m just going to do it all on my laptop because what happens if I’m on a plane?” It’s, yeah, the year before the pandemic, I flew 140,000 miles domestically and I was almost never hamstrung by my ability to do work. And my only local computer is an iPad for those things. So, it turns out that is less of a real world concern for most folks.

Chris: Yeah I actually ordered the components to upgrade an old Nook that I have here and turn it into my, like, this is my remote code server, that’s going to be all attached to GitHub and everything else. That’s where I want to be: have Tailscale and just VPN into this box.

Corey: Tailscale is transformative.

Chris: Yes. Tailscale will change your life. That’s just my personal opinion.

Corey: Yep.

Chris: That’s not an AWS opinion or anything. But yeah, when you start thinking about your network as it could be anywhere, that’s where Tailscale, like, really shines. So—

Corey: Tailscale makes the internet work like we all wanted to believe that it worked.

Chris: Yeah. And Wireguard is an excellent open-source project. And Tailscale consumes that and puts an amazingly easy-to-use UI, and troubleshooting tools, and routing, and all kinds of forwarding capabilities, and makes it kind of easy, which is really, really, really kind of awesome. And Tailscale and Kubernetes—

Corey: Yeah, ‘network’ and ‘easy’ don’t belong in the same sentence, but in this case, they do.

Chris: Yeah. And trust me, the Kubernetes story in Tailscale, there is a lot of there. I understand you might want to not open ports in your VPC, maybe, but if you use Tailscale, that node is just another thing on your network. You can connect to that and see what’s going on. Your management cluster is just another thing on the network where you can watch the state.

But it’s all—you’re connected to it continuously through Tailscale. Or, you know, it’s a much lighter weight, kind of meshy VPN, I would say, if I had to sum it up in one sentence. That was not on our agenda to talk about at all. Anyways. [laugh]

Corey: No, no. I love how many different topics we talk about on these things. We’ll have to have you back soon to talk again. I really want to thank you for being so generous with your time. If people want to learn more about what you’re up to and how you view these things, where can they find you?

Chris: Go to ChrisShort.net. So, Chris Short—I’m six-four so remember, it’s Short—dot net, and you will find all the places that I write, you can go to devopsish.com to subscribe to my newsletter, which goes out every week. This year. Next year, there’ll be breaks. And then finally, if you want to follow me on Twitter, Chris Short: at @ChrisShort on Twitter. All one word so you see two s’s. Like, it’s okay, there’s two s’s there.

Corey: Links to all of that will of course be in the show notes. It’s easier for people to do the clicky-clicky thing as a general rule.

Chris: Clicky things are easier than the wordy things, yes.

Corey: Says the Kubernetes guy.

Chris: Yeah. Says the Kubernetes guy. Yeah, you like that, huh? Like I said, Argo gives you a UI. [laugh].

Corey: Thank you [laugh] so much for your time. I really do appreciate it.

Chris: Thank you. This has been fun. If folks have questions, feel free to reach out. Like, I am not one of those people that hides behind a screen all day and doesn’t respond. I will respond to you eventually.

Corey: I’m right here, Chris. Come on, come on. You’re calling me out in front of myself. My God.

Chris: Egh. It might take a day or two, but I will respond. I promise.

Corey: Thanks again for your time. This has been Chris Short, senior developer advocate at AWS. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice and if it’s YouTube, click the thumbs-up button. Whereas if you’ve hated this podcast, same thing, smash the buttons five-star review and leave an insulting comment that is written in syntactically correct YAML because it’s just so easy to do.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Sheeri

After almost 2 decades as a database administrator and award-winning thought leader, Sheeri Cabral pivoted to technical product management. Her super power of “new customer” empathy informs her presentations and explanations. Sheeri has developed unique insights into working together and planning, having survived numerous reorganizations, “best practices”, and efficiency models. Her experience is the result of having worked at everything from scrappy startups such as Guardium – later bought by IBM – to influential tech companies like Mozilla and MongoDB, to large established organizations like Salesforce.

Links Referenced:

  • Collibra: https://www.collibra.com
  • WildAid GitHub: https://github.com/wildaid
  • Twitter: https://twitter.com/sheeri
  • Personal Blog: https://sheeri.org

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored by our friends at Fortinet. Fortinet’s partnership with AWS is a better-together combination that ensures your workloads on AWS are protected by best-in-class security solutions powered by comprehensive threat intelligence and more than 20 years of cybersecurity experience. Integrations with key AWS services simplify security management, ensure full visibility across environments, and provide broad protection across your workloads and applications. Visit them at AWS re:Inforce to see the latest trends in cybersecurity on July 25-26 at the Boston Convention Center. Just go over to the Fortinet booth and tell them Corey Quinn sent you and watch for the flinch. My thanks again to my friends at Fortinet.

Corey: Let’s face it, on-call firefighting at 2am is stressful! So there’s good news and there’s bad news. The bad news is that you probably can’t prevent incidents from happening, but the good news is that incident.io makes incidents less stressful and a lot more valuable. incident.io is a Slack-native incident management platform that allows you to automate incident processes, focus on fixing the issues and learn from incident insights to improve site reliability and fix your vulnerabilities. Try incident.io, recover faster and sleep more.

Corey: Welcome to Screaming in the Cloud, I’m Corey Quinn. My guest today is Sheeri Cabral, who’s a Senior Product Manager of ETL lineage at Collibra. And that is an awful lot of words that I understand approximately none of, except maybe manager. But we’ll get there. The origin story has very little to do with that.

I was following Sheeri on Twitter for a long time and really enjoyed the conversations that we had back and forth. And over time, I started to realize that there were a lot of things that didn’t necessarily line up. And one of the more interesting and burning questions I had is, what is it you do, exactly? Because you’re all over the map. First, thank you for taking the time to speak with me today. And what is it you’d say it is you do here? To quote a somewhat bizarre and aged movie now.

Sheeri: Well, since your listeners are technical, I do like to match what I say with the audience. First of all, hi. Thanks for having me. I’m Sheeri Cabral. I am a product manager for technical and ETL tools and I can break that down for this technical audience. If it’s not a technical audience, I might say something—like if I’m at a party, and people ask what I do—I’ll say, “I’m a product manager for technical data tool.” And if they ask what a product manager does, I’ll say I helped make sure that, you know, we deliver a product the customer wants. So, you know, ETL tools are tools that transform, extract, and load your data from one place to another.

Corey: Like AWS Glue, but for some of them, reportedly, you don’t have to pay AWS by the gigabyte-second.

Sheeri: Correct. Correct. We actually have an AWS Glue technical lineage tool in beta right now. So, the technical lineage is how data flows from one place to another. So, when you’re extracting, possibly transforming, and loading your data from one place to another, you’re moving it around; you want to see where it goes. Why do you want to see where it goes? Glad you asked. You didn’t really ask. Do you care? Do you want to know why it’s important?

Corey: Oh, I absolutely do. Because it’s—again, people who are, like, “What do you do?” “Oh, it’s boring, and you won’t care.” It’s like when people aren’t even excited themselves about what they work on, it’s always a strange dynamic. There’s a sense that people aren’t really invested in what they do.

I’m not saying you have to have this overwhelming passion and do this in your spare time, necessarily, but you should, at least in an ideal world, like what you do enough to light up a bit when you talk about it. You very clearly do. I’m not wanting to stop you. Please continue.

Sheeri: I do. I love data and I love helping people. So, technical lineage does a few things. For example, a DBA—which I used to be a DBA—can use technical lineage to predict the impact of a schema update or migration, right? So, if I’m going to change the name of this column, what uses it downstream? What’s going to be affected? What scripts do I need to change? Because if the name changes other thing—you know, then I need to not get errors everywhere.

And from a data governance perspective, which Collibra is data governance tool, it helps organizations see if, you know, you have private data in a source, does it remain private throughout its journey, right? So, you can take a column like email address or government ID number and see where it’s used down the line, right? GDPR compliance, CCPA compliance. The CCPA is a little newer; people might not know that acronym. It’s California Consumer Privacy Act.

I forget what GDPR is, but it’s another privacy act. It also can help the business see where data comes from so if you have technical lineage all the way down to your reports, then you know whether or not you can trust the data, right? So, you have a report and it shows salary ranges for job titles. So, where did the data come from? Did it come from a survey? Did it come from job sites? Or did it come from a government source like the IRS, right? So, now you know, like, what you get to trust the most.

Corey: Wait, you can do that without a blockchain? I kid, I kid, I kid. Please don’t make me talk about blockchains. No, it’s important. The provenance of data, being able to establish a almost a chain-of-custody style approach for a lot of these things is extraordinarily important.

Sheeri: Yep.

Corey: I was always a little hazy on the whole idea of ETL until I started, you know, working with large-volume AWS bills. And it turns out that, “Well, why do you have to wind up moving and transforming all of these things?” “Oh, because in its raw form, it’s complete nonsense. That’s why. Thank you for asking.” It becomes a problem—

Sheeri: [laugh]. Oh, I thought you’re going to say because AWS has 14 different products for things, so you have to move it from one product to the other to use the features.

Corey: And two of them are good. It’s a wild experience.

Sheeri: [laugh].

Corey: But this is also something of a new career for you. You were a DBA for a long time. You’re also incredibly engaging, you have a personality, you’re extraordinarily creative, and that—if I can slander an entire profession for a second—does not feel like it is a common DBA trait. It’s right up there with an overly creative accountant. When your accountant has done a stand-up comedy, you’re watching and you’re laughing and thinking, “I am going to federal prison.” It’s one of those weird things that doesn’t quite gel, if we’re speaking purely in terms of stereotypes. What has your career been like?

Sheeri: I was a nerd growing up. So, to kind of say, like, I have a personality, like, my personality is very nerdish. And I get along with other nerdy people and we have a lot of fun, but when I was younger, like, when I was, I don’t know, seven or eight, one of the things I really love to do is I had a penny collection—you know, like you do—and I love to sort it by date. So, in the states anyway, we have these pennies that have the date that they were minted on it. And so, I would organize—and I probably had, like, five bucks worth a pennies.

So, you’re talking about 500 pennies and I would sort them and I’d be like, “Oh, this is 1969. This was 1971.” And then when I was done, I wanted to sort things more, so I would start to, like, sort them in order how shiny the pennies were. So, I think that from an early age, it was clear that I wanted to be a DBA from that sorting of my data and ordering it, but I never really had a, like, “Oh, I want to be this when I grew up.” I kind of had a stint when I was in, like, middle school where I was like, maybe I’ll be a creative writer and I wasn’t as creative a writer as I wanted to be, so I was like, “Ah, whatever.”

And I ended up actually coming to computer science just completely through random circumstance. I wanted to do neuroscience because I thought it was completely fascinating at how the brain works and how, like, you and I are, like, 99.999—we’re, like, five-nines the same except for, like, a couple of genetic, whatever. But, like, how our brain wiring right how the neuron, how the electricity flows through it—

Corey: Yeah, it feels like I want to store a whole bunch of data, that’s okay. I’ll remember it. I’ll keep it in my head. And you’re, like, rolling up the sleeves and grabbing, like, the combination software package off the shelf and a scalpel. Like, “Not yet, but you’re about to.” You’re right, there is an interesting point of commonality on this. It comes down to almost data organization and the—

Sheeri: Yeah.

Corey: —relationship between data nodes if that’s a fair assessment.

Sheeri: Yeah. Well, so what happened was, so I went to university and in order to take introductory neuroscience, I had to take, like, chemistry, organic chemistry, biology, I was basically doing a pre-med track. And so, in the beginning of my junior year, I went to go take introductory neuroscience and I got a D-minus. And a D-minus level doesn’t even count for the major. And I’m like, “Well, I want to graduate in three semesters.”

And I had this—I got all my requirements done, except for the pesky little major thing. So, I was already starting to take, like, a computer science, you know, basic courses and so I kind of went whole-hog, all-in did four or five computer science courses a semester and got my degree in computer science. Because it was like math, so it kind of came a little easy to me. So taking, you know, logic courses, and you know, linear algebra courses was like, “Yeah, that’s great.” And then it was the year 2000, when I got my bachelor’s, the turn of the century.

And my university offered a fifth-year master’s degree program. And I said, I don’t know who’s going to look at me and say, conscious bias, unconscious bias, “She’s a woman, she can’t do computer science, so, like, let me just get this master’s degree.” I, like, fill out a one page form, I didn’t have to take a GRE. And it was the year 2000. You were around back then.

You know what it was like. The jobs were like—they were handing jobs out like candy. I literally had a friend who was like, “My company”—that he founded. He’s like, just come, you know, it’s Monday in May—“Just start, you will just bring your resume the first day and we’ll put it on file.” And I was like, no, no, I have this great opportunity to get a master’s degree in one year at 25% off the cost because I got a tuition reduction or whatever for being in the program. I was like, “What could possibly go wrong in one year?”

And what happened was his company didn’t exist the next year, and, like, everyone was in a hiring freeze in 2001. So, it was the best decision I ever made without really knowing because I would have had a job for six months had been laid off with everyone else at the end of 2000 and… and that’s it. So, that’s how I became a DBA is I, you know, got a master’s degree in computer science, really wanted to use databases. There weren’t any database jobs in 2001, but I did get a job as a sysadmin, which we now call SREs.

Corey: Well, for some of the younger folks in the audience, I do want to call out the fact that regardless of how they think we all rode dinosaurs to school, databases did absolutely exist back in that era. There’s a reason that Oracle is as large as it is of a company. And it’s not because people just love doing business with them, but technology was head and shoulders above everything else for a long time, to the point where people worked with them in spite of their reputation, not because of it. These days, it seems like in the database universe, you have an explosion of different options and different ways that are great at different things. The best, of course, is Route 53 or other DNS TXT records. Everything else is competing for second place on that. But no matter what it is, you’re after, there are options available. This was not the case back then. It was like, you had a few options, all of them with serious drawbacks, but you had to pick your poison.

Sheeri: Yeah. In fact, I learned on Postgres in university because you know, that was freely available. And you know, you’d like, “Well, why not MySQL? Isn’t that kind of easier to learn?” It’s like, yeah, but I went to college from ’96 to 2001. MySQL 1.0 or whatever was released in ’95. By the time I graduated, it was six years old.

Corey: And academia is not usually the early adopter of a lot of emerging technologies like that. That’s not a dig on them any because otherwise, you wind up with a major that doesn’t exist by the time that the first crop of students graduates.

Sheeri: Right. And they didn’t have, you know, transactions. They didn’t have—they barely had replication, you know? So, it wasn’t a full-fledged database at the time. And then I became a MySQL DBA. But yeah, as a systems administrator, you know, we did websites, right? We did what web—are they called web administrators now? What are they called? Web admins? Webmaster?

Corey: Web admins, I think that they became subsumed into sysadmins, by and large and now we call them DevOps, or SRE, which means the exact same thing except you get paid 60% more and your primary job is arguing about which one of those you’re not.

Sheeri: Right. Right. Like we were still separated from network operations, but database stuff that stuff and, you know, website stuff, it’s stuff we all did, back when your [laugh] webmail was your Horde based on PHP and you had a database behind it. And yeah, it was fun times.

Corey: I worked at a whole bunch of companies in that era. And that’s where basically where I formed my early opinion of a bunch of DBA-leaning sysadmins. Like the DBA in and a lot of these companies, it was, I don’t want to say toxic, but there’s a reason that if I were to say, “I’m writing a memoir about a career track in tech called The Legend of Surly McBastard,” people are going to say, “Oh, is it about the DBA?”

There’s a reason behind this. It always felt like there was a sense of elitism and a sense of, “Well, that’s not my job, so you do your job, but if anything goes even slightly wrong, it’s certainly not my fault.” And to be fair, all of these fields have evolved significantly since then, but a lot of those biases that started early in our career are difficult to shake, particularly when they’re unconscious.

Sheeri: They are. I’d never ran into that person. Like, I never ran into anyone who—like a developer who treated me poorly because the last DBA was a jerk and whatever, but I heard a lot of stories, especially with things like granting access. In fact, I remember, my first job as an actual DBA and not as a sysadmin that also the DBA stuff was at an online gay dating site, and the CTO rage-quit. Literally yelled, stormed out of the office, slammed the door, and never came back.

And a couple of weeks later, you know, we found out that the customer service guys who were in-house—and they were all guys, so I say guys although we also referred to them as ladies because it was an online gay dating site.

Corey: Gals works well too, in those scenarios. “Oh, guys is unisex.” “Cool. So’s ‘gals’ by that theory. So gals, how we doing?” And people get very offended by that and suddenly, yeah, maybe ‘folks’ is not a terrible direction to go in. I digress. Please continue.

Sheeri: When they hired me, they were like, are you sure you’re okay with this? I’m like, “I get it. There’s, like, half-naked men posters on the wall. That’s fine.” But they would call they’d be, like, “Ladies, let’s go to our meeting.” And I’m like, “Do you want me also?” Because I had to ask because that was when ladies actually might not have included me because they meant, you know.

Corey: I did a brief stint myself as the director of TechOps at Grindr. That was a wild experience in a variety of different ways.

Sheeri: Yeah.

Corey: It’s over a decade ago, but it was still this… it was a very interesting experience in a bunch of ways. And still, to this day, it remains the single biggest source of InfoSec nightmares that kept me awake at night. Just because when I’m working at a bank—which I’ve also done—it’s only money, which sounds ridiculous to say, especially if you’re in a regulated profession, but here in reality where I’m talking about it, it’s I’m dealing instead, with cool, this data leaks, people will die. Most of what I do is not life or death, but that was and that weighed very heavily on me.

Sheeri: Yeah, there’s a reason I don’t work for a bank or a hospital. You know, I make mistakes. I’m human, right?

Corey: There’s a reason I work on databases for that exact same reason. Please, continue.

Sheeri: Yeah. So, the CTO rage-quit. A couple of weeks later, the head of customer service comes to me and be like, “Can we have his spot as an admin for customer service?” And I’m like, “What do you mean?” He’s like, “Well, he told us, we had, like, ten slots of permission and he was one of them so we could have have, like, nine people.”

And, like, I went and looked, and they put permission in the htaccess file. So, this former CTO had just wielded his power to be like, “Nope, can’t do that. Sorry, limitations.” When there weren’t any. I’m like, “You could have a hundred. You want every customer service person to be an admin? Whatever. Here you go.” So, I did hear stories about that. And yeah, that’s not the kind of DBA I was.

Corey: No, it’s the more senior you get, the less you want to have admin rights on things. But when I leave a job, like, the number one thing I want you to do is revoke my credentials. Not—

Sheeri: Please.

Corey: Because I’m going to do anything nefarious; because I don’t want to get blamed for it. Because we have a long standing tradition in tech at a lot of places of, “Okay, something just broke. Whose fault is it? Well, who’s the most recent person to leave the company? Let’s blame them because they’re not here to refute the character assassination and they’re not going to be angling for a raise here; the rest of us are so let’s see who we can throw under the bus that can’t defend themselves.” Never a great plan.

Sheeri: Yeah. So yeah, I mean, you know, my theory in life is I like helping. So, I liked helping developers as a DBA. I would often run workshops to be like, here’s how to do an explain and find your explain plan and see if you have indexes and why isn’t the database doing what you think it’s supposed to do? And so, I like helping customers as a product manager, right? So…

Corey: I am very interested in watching how people start drifting in a variety of different directions. It’s a, you’re doing product management now and it’s an ETL lineage product, it is not something that is directly aligned with your previous positioning in the market. And those career transitions are always very interesting to me because there’s often a mistaken belief by people in their career realizing they’re doing something they don’t want to do. They want to go work in a different field and there’s this pervasive belief that, “Oh, time for me to go back to square one and take an entry level job.” No, you have a career. You have experience. Find the orthogonal move.

Often, if that’s challenging because it’s too far apart, you find the half-step job that blends the thing you do now with something a lot closer, and then a year or two later, you complete the transition into that thing. But starting over from scratch, it’s why would you do that? I can’t quite wrap my head around jumping off the corporate ladder to go climb another one. You very clearly have done a lateral move in that direction into a career field that is surprisingly distant, at least in my view. How’d that happen?

Sheeri: Yeah, so after being on call for 18 years or so, [laugh] I decided—no, I had a baby, actually. I had a baby. He was great. And then I another one. But after the first baby, I went back to work, and I was on call again. And you know, I had a good maternity leave or whatever, but you know, I had a newborn who was six, eight months old and I was getting paged.

And I was like, you know, this is more exhausting than having a newborn. Like, having a baby who sleeps three hours at a time, like, in three hour chunks was less exhausting than being on call. Because when you have a baby, first of all, it’s very rare that they wake up and crying in the midnight it’s an emergency, right? Like they have to go to the hospital, right? Very rare. Thankfully, I never had to do it.

But basically, like, as much as I had no brain cells, and sometimes I couldn’t even go through this list, right: they need to be fed; they need to be comforted; they’re tired, and they’re crying because they’re tired, right, you can’t make them go to sleep, but you’re like, just go to sleep—what is it—or their diaper needs changing, right? There’s, like, four things. When you get that beep of that pager in the middle of the night it could be anything. It could be logs filling up disk space, you’re like, “Alright, I’ll rotate the logs and be done with it.” You know? It could be something you need snoozed.

Corey: “Issue closed. Status, I no longer give a shit what it is.” At some point, it’s one of those things where—

Sheeri: Replication lag.

Corey: Right.

Sheeri: Not actionable.

Corey: Don’t get me started down that particular path. Yeah. This is the area where DBAs and my sysadmin roots started to overlap a bit. Like, as the DBA was great at data analysis, the table structure and the rest, but the backups of the thing, of course that fell to the sysadmin group. And replication lag, it’s, “Okay.”

“It’s doing some work in the middle of the night; that’s normal, and the network is fine. And why are you waking me up with things that are not actionable? Stop it.” I’m yelling at the computer at that point, not the person—

Sheeri: Right,right.

Corey: —to be very clear. But at some point, it’s don’t wake me up with trivial nonsense. If I’m getting woken up in the middle of the night, it better be a disaster. My entire business now is built around a problem that’s business-hours only for that explicit reason. It’s the not wanting to deal with that. And I don’t envy that, but product management. That’s a strange one.

Sheeri: Yeah, so what happened was, I was unhappy at my job at the time, and I was like, “I need a new job.” So, I went to, like, the MySQL Slack instance because that was 2018, 2019. Very end of 2018, beginning of 2019. And I said, “I need something new.” Like, maybe a data architect, or maybe, like, a data analyst, or data scientist, which was pretty cool.

And I was looking at data scientist jobs, and I was an expert MySQL DBA and it took a long time for me to be able to say, “I’m an expert,” without feeling like oh, you’re just ballooning yourself up. And I was like, “No, I’m literally a world-renowned expert DBA.” Like, I just have to say it and get comfortable with it. And so, you know, I wasn’t making a junior data scientist’s salary. [laugh].

I am the sole breadwinner for my household, so at that point, I had one kid and a husband and I was like, how do I support this family on a junior data scientist’s salary when I live in the city of Boston? So, I needed something that could pay a little bit more. And a former I won’t even say coworker, but colleague in the MySQL world—because is was the MySQL Slack after all—said, “I think you should come at MongoDB, be a product manager like me.”

Corey: This episode is sponsored in part by Honeycomb. When production is running slow, it’s hard to know where problems originate. Is it your application code, users, or the underlying systems? I’ve got five bucks on DNS, personally. Why scroll through endless dashboards while dealing with alert floods, going from tool to tool to tool that you employ, guessing at which puzzle pieces matter? Context switching and tool sprawl are slowly killing both your team and your business. You should care more about one of those than the other; which one is up to you. Drop the separate pillars and enter a world of getting one unified understanding of the one thing driving your business: production. With Honeycomb, you guess less and know more. Try it for free at honeycomb.io/screaminginthecloud. Observability: it’s more than just hipster monitoring.

Corey: If I’ve ever said, “Hey, you should come work with me and do anything like me,” people will have the blood drain from their face. And like, “What did you just say to me? That’s terrible.” Yeah, it turns out that I have very hard to explain slash predict, in some ways. It’s always fun. It’s always wild to go down that particular path, but, you know, here we are.

Sheeri: Yeah. But I had the same question everybody else does, which was, what’s a product manager? What does the product manager do? And he gave me a list of things a product manager does, which there was some stuff that I had the skills for, like, you have to talk to customers and listen to them.

Well, I’ve done consulting. I could get yelled at; that’s fine. You can tell me things are terrible and I have to fix it. I’ve done that. No problem with that. Then there are things like you have to give presentations about how features were okay, I can do that. I’ve done presentations. You know, I started the Boston MySQL Meetup group and ran it for ten years until I had a kid and foisted it off on somebody else.

And then the things that I didn’t have the skills in, like, running a beta program were like, “Ooh, that sounds fascinating. Tell me more.” So, I was like, “Yeah, let’s do it.” And I talked to some folks, they were looking for a technical product manager for MongoDB’s sharding product. And they had been looking for someone, like, insanely technical for a while, and they found me; I’m insanely technical.

And so, that was great. And so, for a year, I did that at MongoDB. One of the nice things about them is that they invest in people, right? So, my manager left, the team was like, we really can’t support someone who doesn’t have the product management skills that we need yet because you know, I wasn’t a master in a year, believe it or not. And so, they were like, “Why don’t you find another department?” I was like, “Okay.”

And I ended up finding a place in engineering communications, doing, like, you know, some keynote demos, doing some other projects and stuff. And then after—that was a kind of a year-long project, and after that ended, I ended up doing product management for developer relations at MongoDB. Also, this was during the pandemic, right, so this is 2019, until ’21; beginning of 2019, to end of 2020, so it was, you know, three full years. You know, I kind of like woke up from the pandemic fog and I was like, “What am I doing? Do I want to really want to be a content product manager?” And I was like, “I want to get back to databases.”

One of the interesting things I learned actually in looking for a job because I did it a couple of times at MongoDB because I changed departments and I was also looking externally when I did that. I had the idea when I became a product manager, I was like, “This is great because now I’m product manager for databases and so, I’m kind of leveraging that database skill and then I’ll learn the product manager stuff. And then I can be a product manager for any technical product, right?”

Corey: I like the idea. Of some level, it feels like the product managers likeliest to succeed at least have a grounding or baseline in the area that they’re in. This gets into the age-old debate of how important is industry-specific experience? Very often you’ll see a bunch of job ads just put that in as a matter of course. And for some roles, yeah, it’s extremely important.

For other roles it’s—for example, I don’t know, hypothetically, you’re looking for someone to fix the AWS bill, it doesn’t necessarily matter whether you’re a services company, a product company, or a VC-backed company whose primary output is losing money, it doesn’t matter because it’s a bounded problem space and that does not transform much from company to company. Same story with sysadmin types to be very direct. But the product stuff does seem to get into that industry specific stuff.

Sheeri: Yeah, and especially with tech stuff, you have to understand what your customer is saying when they’re saying, “I have a problem doing X and Y,” right? The interesting part of my folly in that was that part of the time that I was looking was during the pandemic, when you know, everyone was like, “Oh, my God, it’s a seller’s market. If you’re looking for a job, employers are chomping at the bit for you.” And I had trouble finding something because so many people were also looking for jobs, that if I went to look for something, for example, as a storage product manager, right—now, databases and storage solutions have a lot in common; databases are storage solutions, in fact; but file systems and databases have much in common—but all that they needed was one person with file system experience that had more experience than I did in storage solutions, right? And they were going to choose them over me. So, it was an interesting kind of wake-up call for me that, like, yeah, probably data and databases are going to be my niche. And that’s okay because that is literally why they pay me the literal big bucks. If I’m going to go niche that I don’t have 20 years of experience and they shouldn’t pay me as big a bucks right?

Corey: Yeah, depending on what you’re doing, sure. I don’t necessarily believe in the idea that well you’re new to this particular type of role so we’re going to basically pay you a lot less. From my perspective it’s always been, like, there’s a value in having a person in a role. The value to the company is X and, “Well, I have an excuse now to pay you less for that,” has never resonated with me. It’s if you’re not, I guess, worth—the value-added not worth being paid what the stated rate for a position is, you are probably not going to find success in that role and the role has to change. That has always been my baseline operating philosophy. Not to yell at people on this, but it’s, uh, I am very tired of watching companies more or less dunk on people from a position of power.

Sheeri: Yeah. And I mean, you can even take the power out of that and take, like, location-based. And yes, I understand the cost of living is different in different places, but why do people get paid differently if the value is the same? Like if I want to get a promotion, right, my company is going to be like, “Well, show me how you’ve added value. And we only pay your value. We don’t pay because—you know, you don’t just automatically get promoted after seven years, right? You have to show the value and whatever.” Which is, I believe, correct, right?

And yet, there are seniority things, there are this many years experience. And you know, there’s the old caveat of do you have ten years experience or do you have two years of experience five times?

Corey: That is the big problem is that there has to be a sense of movement that pushes people forward. You’re not the first person that I’ve had on the show and talked to about a 20 year career. But often, I do wind up talking to folks as I move through the world where they basically have one year of experience repeated 20 times. And as the industry continues to evolve and move on and skill sets don’t keep current, in some cases, it feels like they have lost touch, on some level. And they’re talking about the world that was and still is in some circles, but it’s a market in long-term decline as opposed to keeping abreast of what is functionally a booming industry.

Sheeri: Their skills have depreciated because they haven’t learned more skills.

Corey: Yeah. Tech across the board is a field where I feel like you have to constantly be learning. And there’s a bit of an evolve-or-die dinosaur approach. And I have some, I do have some fallbacks on this. If I ever decide I am tired of learning and keeping up with AWS, all I have to do is go and work in an environment that uses GovCloud because that’s, like, AWS five years ago.

And that buys me the five years to find something else to be doing until a GovCloud catches up with the modern day of when I decided to make that decision. That’s a little insulting and also very accurate for those who have found themselves in that environment. But I digress.

Sheeri: No, and I find it to with myself. Like, I got to the point with MySQL where I was like, okay, great. I know MySQL back and forth. Do I want to learn all this other stuff? Literally just today, I was looking at my DMs on Twitter and somebody DMed me in May, saying, “Hi, ma’am. I am a DBA and how can I use below service: Lambda, Step Functions, DynamoDB, AWS Session Manager, and CloudWatch?”

And I was like, “You know, I don’t know. I have not ever used any of those technologies. And I haven’t evolved my DBA skills because it’s been, you know, six years since I was a DBA.” No, six years, four or five? I can’t do math.

Corey: Yeah. Which you think would be a limiting factor to a DBA but apparently not. One last question that [laugh] I want to ask you, before we wind up calling this a show. You’ve done an awful lot across the board. As you look at all of it, what is it you would say that you’re the most proud of?

Sheeri: Oh, great question. What I’m most proud of is my work with WildAid. So, when I was at MongoDB—I referenced a job with engineering communications, and they hired me to be a product manager because they wanted to do a collaboration with a not-for-profit and make a reference application. So, make an application using MongoDB technology and make it something that was going to be used, but people can also see it. So, we made this open-source project called o-fish.

And you know, we can give GitHub links: it’s github.com/wildaid, and it has—that’s the organization’s GitHub which we created, so it only has the o-fish projects in it. But it is a mobile and web app where governments who patrol waters, patrol, like, marine protected areas—which are like national parks but in the water, right, so they are these, you know, wildlife preserves in the water—and they make sure that people aren’t doing things they shouldn’t do: they’re not throwing trash in the ocean, they’re not taking turtles out of the Galapagos Island area, you know, things like that. And they need software to track that and do that because at the time, they were literally writing, you know, with pencil on paper, and, you know, had stacks and stacks of this paper to do data entry.

And MongoDB had just bought the Realm database and had just integrated it, and so there was, you know, some great features about offline syncing that you didn’t have to do; it did all the foundational plumbing for you. And then the reason though, that I’m proud of that project is not just because it’s pretty freaking cool that, you know, doing something that actually makes a difference in the world and helps fight climate change and all that kind of stuff, the reason I was proud of it is I was the sole product manager. It was the first time that I’d really had sole ownership of a product and so all the mistakes were my own and the credit was my own, too. And so, it was really just a great learning experience and it turned out really well.

Corey: There’s a lot to be said for pitching in and helping out with good causes in a way that your skill set winds up benefitting. I found that I was a lot happier with a lot of the volunteer stuff that I did when it was instead of licking envelopes, it started being things that I had a bit of proficiency in. “Hey, can I fix your AWS bill?” It turns out as some value to certain nonprofits. You have to be at a certain scale before it makes sense, otherwise it’s just easier to maybe not do it that way, but there’s a lot of value to doing something that puts good back into the world. I wish more people did that.

Sheeri: Yeah. And it’s something to do in your off-time that you know is helping. It might feel like work, it might not feel like work, but it gives you a sense of accomplishment at the end of the day. I remember my first job, one of the interview questions was—no, it wasn’t. [laugh]. It wasn’t an interview question until after I was hired and they asked me the question, and then they made it an interview question.

And the question was, what video games do you play? And I said, “I don’t play video games. I spend all day at work staring at a computer screen. Why would I go home and spend another 12 hours till three in the morning, right—five in the morning—playing video games?” And they were like, we clearly need to change our interview questions. This was again, back when the dinosaurs roamed the earth. So, people are are culturally sensitive now.

Corey: These days, people ask me, “What’s your favorite video game?” My answer is, “Twitter.”

Sheeri: Right. [laugh]. Exactly. It’s like whack-a-mole—

Corey: Yeah.

Sheeri: —you know? So, for me having a tangible hobby, like, I do a lot of art, I knit, I paint, I carve stamps, I spin wool into yarn. I know that’s not a metaphor for storytelling. That is I literally spin wool into yarn. And having something tangible, you work on something and you’re like, “Look. It was nothing and now it’s this,” is so satisfying.

Corey: I really want to thank you for taking the time to speak with me today about where you’ve been, where you are, and where you’re going, and as well as helping me put a little bit more of a human angle on Twitter, which is intensely dehumanizing at times. It turns out that 280 characters is not the best way to express the entirety of what makes someone a person. You need to use a multi-tweet thread for that. If people want to learn more about you, where can they find you?

Sheeri: Oh, they can find me on Twitter. I’m @sheeri—S-H-E-E-R-I—on Twitter. And I’ve started to write a little bit more on my blog at sheeri.org. So hopefully, I’ll continue that since I’ve now told people to go there.

Corey: I really want to thank you again for being so generous with your time. I appreciate it.

Sheeri: Thanks to you, Corey, too. You take the time to interview people, too, so I appreciate it.

Corey: I do my best. Sheeri Cabral, Senior Product Manager of ETL lineage at Collibra. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice or smash the like and subscribe buttons on the YouTubes, whereas if you’ve hated it, do exactly the same thing—like and subscribe, hit those buttons, five-star review—but also leave a ridiculous comment where we will then use an ETL pipeline to transform it into something that isn’t complete bullshit.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Chris

Chris is the Co-founder and Chief Product Officer at incident.io, where they're building incident management products that people actually want to use. A software engineer by trade, Chris is no stranger to gnarly incidents, having participated (and caused!) them at everything from early stage startups through to enormous IT organizations.

Links Referenced:

  • incident.io: https://incident.io
  • Practical Guide to Incident Management: https://incident.io/guide/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: DoorDash had a problem. As their cloud-native environment scaled and developers delivered new features, their monitoring system kept breaking down. In an organization where data is used to make better decisions about technology and about the business, losing observability means the entire company loses their competitive edge. With Chronosphere, DoorDash is no longer losing visibility into their applications suite. The key? Chronosphere is an open-source compatible, scalable, and reliable observability solution that gives the observability lead at DoorDash business, confidence, and peace of mind. Read the full success story at snark.cloud/chronosphere. That's snark.cloud slash C-H-R-O-N-O-S-P-H-E-R-E.

Corey: Let’s face it, on-call firefighting at 2am is stressful! So there’s good news and there’s bad news. The bad news is that you probably can’t prevent incidents from happening, but the good news is that incident.io makes incidents less stressful and a lot more valuable. incident.io is a Slack-native incident management platform that allows you to automate incident processes, focus on fixing the issues and learn from incident insights to improve site reliability and fix your vulnerabilities. Try incident.io, recover faster and sleep more.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Today’s promoted guest is Chris Evans, who’s the CPO and co-founder of incident.io. Chris, first, thank you very much for joining me. And I’m going to start with an easy question—well, easy question, hard answer, I think—what is an incident.io exactly?

Chris: Incident.io is a software platform that helps entire organizations to respond to recover from and learn from incidents.

Corey: When you say incident, that means an awful lot of things. And depending on where you are in the ecosystem in the world, that means different things to different people. For example, oh, incident. Like, “Are you talking about the noodle incident because we had an agreement that we would never speak about that thing again,” style, versus folks who are steeped in DevOps or SRE culture, which is, of course, a fancy way to say those who are sad all the time, usually about computers. What is an incident in the context of what you folks do?

Chris: That, I think, is the killer question. I think if you look at organizations in the past, I think incidents were those things that happened once a quarter, maybe once a year, and they were the thing that brought the entirety of your site down because your big central database that was in a data center sort of disappeared. The way that modern companies run means that the definition has to be very, very different. So, most places now rely on distributed systems and there is no, sort of, binary sense of up or down these days. And essentially, in the general case, like, most companies are continually in a sort of state of things being broken all of the time.

And so, for us, when we look at what an incident is, it is essentially anything that takes you away from your planned work with a sense of urgency. And that’s the sort of the pithy definition that we use there. Generally, that can mean anything—it means different things to different folks, and, like, when we talk to folks, we encourage them to think carefully about what that threshold is, but generally, for us at incident.io, that means basically a single error that is worthwhile investigating that you would stop doing your backlog work for is an incident. And also an entire app being down, that is an incident.

So, there’s quite a wide range there. But essentially, by sort of having more incidents and lowering that threshold, you suddenly have a heap of benefits, which I can go very deep into and talk for hours about.

Corey: It’s a deceptively complex question. When I talk to folks about backups, one of the biggest problems in the world of backup and building a DR plan, it’s not building the DR plan—though that’s no picnic either—it’s okay. In the time of cloud, all your planning figures out, okay. Suddenly the site is down, how do we fix it? There are different levels of down and that means different things to different people where, especially the way we build apps today, it’s not is the service or site up or down, but with distributed systems, it’s how down is it?

And oh, we’re seeing elevated error rates in us-tire-fire-1 region of AWS. At what point do we begin executing on our disaster plan? Because the worst answer, in some respects is, every time you think you see a problem, you start failing over to other regions and other providers and the rest, and three minutes in, you’ve irrevocably made the cutover and it’s going to take 15 minutes to come back up. And oh, yeah, then your primary site comes back up because whoever unplugged something, plugged it back in and now you’ve made the wrong choice. Figuring out all the things around the incident, it’s not what it once was.

When you were running your own blog on a single web server and it’s broken, it’s pretty easy to say, “Is it up or is it down?” As you scale out, it seems like that gets more and more diffuse. But it feels to me that it’s also less of a question of how the technology has scaled, but also how the culture and the people have scaled. When you’re the only engineer somewhere, you pretty much have no choice but to have the entire state of your stack shoved into your head. When that becomes 15 or 20 different teams of people, in some cases, it feels like it’s almost less than a technology problem than it is a problem of how you communicate and how you get people involved. And the issues in front of the people who are empowered and insightful in a certain area that needs fixing.

Chris: A hundred percent. This is, like, a really, really key point, which is that organizations themselves are very complex. And so, you’ve got this combination of systems getting more and more complicated, more and more sort of things going wrong and perpetually breaking but you’ve got very, very complicated information structures and communication throughout the whole organization to keep things up and running. The very best orgs are the ones where they can engage the entire, sort of, every corner of the organization when things do go wrong. And lived and breathed this firsthand when various different previous companies, but most recently at Monzo—which is a bank here in the UK—when an incident happened there, like, one of our two physical data center locations went down, the bank wasn’t offline. Everything was resilient to that, but that required an immediate response.

And that meant that engineers were deployed to go and fix things. But it also meant the customer support folks might be required to get involved because we might be slightly slower processing payments. And it means that risk and compliance folks might need to get involved because they need to be reporting things to regulators. And the list goes on. There’s, like, this need for a bunch of different people who almost certainly have never worked together or rarely worked together to come together, land in this sort of like empty space of this incident room or virtual incident room, and figure out how they’re going to coordinate their response and get things back on track in the sort of most streamlined way and as quick as possible.

Corey: Yeah, when your bank is suddenly offline, that seems like a really inopportune time to be introduced to the database team. It’s, “Oh, we have one of those. Wonderful. I feel like you folks are going to come in handy later today.” You want to have those pathways of communication open well in advance of these issues.

Chris: A hundred percent. And I think the thing that makes incidents unique is that fact. And I think the solution to that is this sort of consistent, level playing field that you can put everybody on. So, if everybody understands that the way that incidents are dealt with is consistent, we declare it like this, and under these conditions, these things happen. And, you know, if I flag this kind of level of impact, we have to pull in someone else to come and help make a decision.

At the core of it, there’s this weird kind of duality to incidents where they are both kind of semi-formulaic and that you can basically encode a lot of the processes that happen, but equally, they are incredibly chaotic and require a lot of human impact to be resilient and figure these things out because stuff that you have never seen happen before is happening and failing in ways that you never predicted. And so, this is where incident.io plays into this is that we try to take the first half of that off of your hands, which is, we will help you run your process so that all of the brain capacity you have, it goes on to the bit that humans are uniquely placed to be able to do, which is responding to these very, very chaotic, sort of, surprise events that have happened.

Corey: I feel as well—because I played around in this space a bit before I used to run ops teams—and, more or less I really should have had a t-shirt then that said, “I am the root cause,” because yeah, I basically did a lot of self-inflicted outages in various environments because it turns out, I’m not always the best with computers. Imagine that. There are a number of different companies that play in the space that look at some part of the incident lifecycle. And from the outside, first, they all look alike because it’s, “Oh, so you’re incident.io. I assume you’re PagerDuty. You’re the thing that calls me at two in the morning to make sure I wake up.”

Conversely, for folks who haven’t worked deeply in that space, as well, of setting things on fire, what you do sounds like it’s highly susceptible to the Hacker News problem. Where, “Wait, so what you do is effectively just getting people to coordinate and talk during an incident? Well, that doesn’t sound hard. I could do that in a weekend.” And no, no, you can’t.

If this were easy, you would not have been in business as long as you have, have the team the size that you do, the customers that you do. But it’s one of those things that until you’ve been in a very specific set of a problem, it doesn’t sound like it’s a real problem that needs solving.

Chris: Yeah, I think that’s true. And I think that the Hacker News point is a particularly pertinent one and that someone else, sort of, in an adjacent area launched on Hacker News recently, and the amount of feedback they got around, you know, “You’re a Slack bot. How is this a company?” Was kind of staggering. And I think generally where that comes from is—well, first of all that bias that engineers have, which is just everything you look at as an engineer is like, “Yeah, I can build that in a weekend.” I think there’s often infinite complexity under the hood that just gets kind of brushed over. But yeah, I think at the core of it, you probably could build a Slack bot in a weekend that creates a channel for you in Slack and allows you to post somewhere that some—

Corey: Oh, good. More channels in Slack. Just when everyone wants.

Chris: Well, there you go. I mean, that’s a particular pertinent one because, like, our tool does do that. And one of the things—so I built at Monzo, a version of incident.io that we used at the company there, and that was something that I built evenings and weekends. And among the many, many things I never got around to building, archiving and cleaning up channels was one of the ones that was always on that list.

And so, Monzo did have this problem of littered channels everywhere, I think that sort of like, part of the problem here is, like, it is easy to look at a product like ours and sort of assume it is this sort of friendly Slack bot that helps you orchestrate some very basic commands. And I think when you actually dig into the problems that organizations above a certain size have, they’re not solved by Slack bots. They’re solved by platforms that help you to encode your processes that otherwise have to live on a Google Doc somewhere which is five pages long and when it’s 2 a.m. and everything’s on fire, I guarantee you not a single person reads that Google Doc, so your process is as good as not in place at all. That’s the beauty of a tool like ours. We have a powerful engine that helps you basically to encode that and take some load off of you.

Corey: To be clear, I’m also not coming at this from a position of judging other people. I just look right now at the Slack workspace that we have The Duckbill Group, and we have something like a ten-to-one channel-to-human ratio. And the proliferation of channels is a very real thing. And the problem that I’ve seen across the board with other things that try to address incident management has always been fanciful at best about what really happens when something breaks. Like, you talk about, oh, here’s what happens. Step one: you will pull up the Google Doc, or you will pull up the wiki or the rest, or in some aspirational places, ah, something seems weird, I will go open a ticket in Jira.

Meanwhile, here in reality, anyone who’s ever worked in these environments knows that step one, “Oh shit, oh shit, oh shit, oh shit, oh shit. What are we going to do?” And all the practices and procedures that often exist, especially in orgs that aren’t very practiced at these sorts of things, tend to fly out the window and people are going to do what they’re going to do. So, any tool or any platform that winds up addressing that has to accept the reality of meeting people where they are not trying to educate people into different patterns of behavior as such. One of the things I like about your approach is, yeah, it’s going to be a lot of conversation in Slack that is a given we can pretend otherwise, but here in reality, that is how work gets communicated, particularly in extremis. And I really appreciate the fact that you are not trying to, like, fight what feels almost like a law of nature at this point.

Chris: Yeah, I think there’s a few things in that. The first point around the document approach or the clearly defined steps of how an incident works. In my experience, those things have always gone wrong because—

Corey: The data center is down, so we’re going to the wiki to follow our incident management procedure, which is in the data center just lost power.

Chris: Yeah.

Corey: There’s a dependency problem there, too. [laugh].

Chris: Yeah, a hundred percent. [laugh]. A hundred percent. And I think part of the problem that I see there is that very, very often, you’ve got this situation where the people designing the process are not the people following the process. And so, there’s this classic, I’ve heard it through John Allspaw, but it’s a bunch of other folks who talk about the difference between people, you know, at the sharp end or the blunt end of the work.

And I think the problem that people are facing the past is you have these people who sit in the, sort of, metaphorical upstairs of the office and think that they make a company safe by defining a process on paper. And they ship the piece of paper and go, “That is a good job for me done. I’m going to leave and know that I’ve made the bank—the other whatever your organization does—much, much safer.” And I think this is where things fall down because—

Corey: I want to ambush some of those people in their performance reviews with, “Cool. Just for fun, all the documentation here, we’re going to pull up the analytics to see how often that stuff gets viewed. Oh, nobody ever sees it. Hmm.”

Chris: It’s frustrating. It’s frustrating because that never ever happens, clearly. But the point you made around, like, meeting people where you are, I think that is a huge one, which is incidents are founded on great communication. Like, as I said earlier, this is, like, a form of team with someone you’ve never ever worked with before and the last thing you want to do is be, like, “Hey, Corey, I’ve never met you before, but let’s jump out onto this other platform somewhere that I’ve never been or haven’t been for weeks and we’ll try and figure stuff out over there.” It’s like, no, you’re going to be communicating—

Corey: We use Slack internally, but we have a WhatsApp chat that we wind up using for incident stuff, so go ahead and log into WhatsApp, which you haven’t done in 18 months, and join the chat. Yeah, in the dawn of time, in the mists of antiquity, you vaguely remember hearing something about that your first week and then never again. This stuff has to be practiced and it’s important to get it right. How do you approach the inherent and often unfortunate reality that incident response and management inherently becomes very different depending upon the specifics of your company or your culture or something like that? In other words, how cookie-cutter is what you have built versus adaptable to different environments it finds itself operating in?

Chris: Man, the amount of time we spent as a founding team in the early days deliberating over how opinionated we should be versus how flexible we should be was staggering. The way we like to describe it as we are quite opinionated about how we think incidents should be run, however we let you imprint your own process into that, so putting some color onto that. We expect incidents to have a lead. That is something you cannot get away from. However, you can call the lead whatever makes sense for you at your organization. So, some folks call them an incident commander or a manager or whatever else.

Corey: There’s overwhelming militarization of these things. Like, oh, yes, we’re going to wind up taking a bunch of terms from the military here. It’s like, you realize that your entire giant screaming fire is that the lights on the screen are in the wrong pattern. You’re trying to make them in the right pattern. No one dies here in most cases, so it feels a little grandiose for some of those terms being tossed around in some cases, but I get it. You’ve got to make something that is unpleasant and tedious in many respects, a little bit more gripping. I don’t envy people. Messaging is hard.

Chris: Yeah, it is. And I think if you’re overly virtuoustic and inflexible, you’re sort of fighting an uphill battle here, right? So, folks are going to want to call things what they want to call things. And you’ve got people who want to import [ITIL 00:15:04] definitions for severity ease into the platform because that’s what they’re familiar with. That’s fine.

What we are opinionated about is that you have some severity levels because absent academic criticism of severity levels, they are a useful mechanism to very coarsely and very quickly assess how bad something is and to take some actions off of it. So yeah, we basically have various points in the product where you can customize and put your own sort of flavor on it, but generally, we have a relatively opinionated end-to-end expectation of how you will run that process.

Corey: The thing that I find that annoys me—in some cases—the most is how heavyweight the process is, and it’s clearly built by people in an ivory tower somewhere where there’s effectively a two-day long postmortem analysis of the incident, and so on and so forth. And okay, great. Your entire site has been blown off the internet, yeah, that probably makes sense. But as soon as you start broadening that to things like okay, an increase in 500 errors on this service for 30 minutes, “Great. Well, we’re going to have a two-day postmortem on that.” It’s, “Yeah, sure would be nice if we could go two full days without having another incident of that caliber.” So, in other words, whose foot—are we going to hire a new team whose full-time job it is, is to just go ahead and triage and learn from all these incidents? Seems to me like that’s sort of throwing wood behind the wrong arrows.

Chris: Yeah, I think it’s very reductive to suggest that learning only happens in a postmortem process. So, I wrote a blog, actually, not so long ago that is about running postmortems and when it makes sense to do it. And as part of that, I had a sort of a statement that was [laugh] that we haven’t run a single postmortem when I wrote this blog at incident.io. Which is probably shocking to many people because we’re an incident company, and we talk about this stuff, but we were also a company of five people and when something went wrong, the learning was happening and these things were sort of—we were carving out the time, whether it was called a postmortem, or not to learn and figure out these things. Extrapolating that to bigger companies, there is little value in following processes for the sake of following processes. And so, you could have—

Corey: Someone in compliance just wound up spitting their coffee over their desktop as soon as you said that. But I hear you.

Chris: Yeah. And it's those same folks who are the ones who care about the document being written, not the process and the learning happening. And I think that’s deeply frustrating to me as—

Corey: All the plans, of course, assume that people will prioritize the company over their own family for certain kinds of disasters. I love that, too. It’s divorced from reality; that’s ridiculous, on some level. Speaking of ridiculous things, as you continue to grow and scale, I imagine you integrate with things beyond just Slack. You grab other data sources and over in the fullness of time.

For example, I imagine one of your most popular requests from some of your larger customers is to integrate with their HR system in order to figure out who’s the last engineer who left, therefore everything immediately their fault because lord knows the best practice is to pillory whoever was the last left because then they’re not there to defend themselves anymore and no one’s going to get dinged for that irresponsible jackass’s decisions, even if they never touched the system at all. I’m being slightly hyperbolic, but only slightly.

Chris: Yeah. I think [laugh] that's an interesting point. I am definitely going to raise that feature request for a prefilled root cause category, which is, you know, the value is just that last person who left the organization. That it’s a wonderful scapegoat situation there. I like it.

To the point around what we do integrate with, I think the thing is actually with incidents that’s quite interesting is there is a lot of tooling that exists in this space that does little pockets of useful, valuable things in the shape of incidents. So, you have PagerDuty is this system that does a great job of making people’s phone making noise, but that happens, and then you’re dropped into this sort of empty void of nothingness and you’ve got to go and figure out what to do. And then you’ve got things like Jira where clearly you want to be able to track actions that are coming out of things going wrong in some cases, and that’s a great tool for that. And various other things in the middle there. And yeah, our value proposition, if you want to call it that, is to bring those things together in a way that is massively ergonomic during an incident.

So, when you’re in the middle of an incident, it is really handy to be able to go, “Oh, I have shipped this horrible fix to this thing. It works, but I must remember to undo that.” And we put that at your fingertips in an incident channel from Slack, that you can just log that action, lose that cognitive load that would otherwise be there, move on with fixing the thing. And you have this sort of—I think it’s, like, that multiplied by 1000 in incidents that is just what makes it feel delightful. And I cringe a little bit saying that because it’s an incident at the end of the day, but genuinely, it feels magical when some things happen that are just like, “Oh, my gosh, you’ve automatically hooked into my GitHub thing and someone else merged that PR and you’ve posted that back into the channel for me so I know that that happens. That would otherwise have been a thing where I jump out of the incident to go and figure out what was happening.”

Corey: This episode is sponsored in part by our friend EnterpriseDB. EnterpriseDB has been powering enterprise applications with PostgreSQL for 15 years. And now EnterpriseDB has you covered wherever you deploy PostgreSQL on-premises, private cloud, and they just announced a fully-managed service on AWS and Azure called BigAnimal, all one word. Don’t leave managing your database to your cloud vendor because they’re too busy launching another half-dozen managed databases to focus on any one of them that they didn’t build themselves. Instead, work with the experts over at EnterpriseDB. They can save you time and money, they can even help you migrate legacy applications—including Oracle—to the cloud. To learn more, try BigAnimal for free. Go to biganimal.com/snark, and tell them Corey sent you.

Corey: The problem with the cloud, too, is the first thing that, when there starts to be an incident happening is the number one decision—almost the number one decision point is this my shitty code, something we have just pushed in our stuff, or is it the underlying provider itself? Which is why the AWS status page being slow to update is so maddening. Because those are two completely different paths to go down and you are having to pursue both of them equally at the same time until one can be ruled out. And that is why time to identify at least what side of the universe it’s on is so important. That has always been a bit of a tricky challenge.

I want to talk a bit about circular dependencies. You target a certain persona of customer, but I’m going to go out on a limb and assume that one explicit company that you are not going to want to do business with in your current iteration is Slack itself because a tool to manage—okay, so our service is down, so we’re going to go to Slack to fix it doesn’t work when the service is Slack itself. So, that becomes a significant challenge. As you look at this across the board, are you seeing customers having problems where you have circular dependency issues with this? Easy example: Slack is built on top of AWS.

When there’s an underlying degradation of, huh, suddenly us-east-1 is not doing what it’s supposed to be doing, now, Slack is degraded as well, as well as the customer site, it seems like at that point, you’re sort of in a bit of tricky positioning as a customer. Counterpoint, when neither Slack nor your site are working, figuring out what caused that issue doesn’t seem like it’s the biggest stretch of the imagination at that point.

Chris: I’ve spent a lot of my career working in infrastructure, platform-type teams, and I think you can end up tying yourself in knots if you try and over-optimize for, like, avoiding these dependencies. I think it’s one of those, sort of, turtles all the way down situations. So yes, Slack are unlikely to become a customer because they are clearly going to want to use our product when they are down.

Corey: They reach out, “We’d like to be your customer.” Your response is, “Please don’t be.” None of us are going to be happy with this outcome.

Chris: Yeah, I mean, the interesting thing that is that we’re friends with some folks at Slack, and they believe it or not, they do use Slack to navigate their incidents. They have an internal tool that they have written. And I think this sort of speaks to the point we made earlier, which is that incidents and things failing or not these sort of big binary events. And so—

Corey: All of Slack is down is not the only kind of incident that a company like Slack can experience.

Chris: I’d go as far as that it’s most commonly not that. It’s most commonly that you’re navigating incidents where it is a degradation, or some edge case, or something else that’s happened. And so, like, the pragmatic solution here is not to avoid the circular dependencies, in my view; it’s to accept that they exist and make sure you have sensible escape hatches so that when something does go wrong—so a good example, we use incident.io at incident.io to manage incidents that we’re having with incident.io. And 99% of the time, that is absolutely fine because we are having some error in some corner of the product or a particular customer is doing something that is a bit curious.

And I could count literally on one hand the number of times that we have not been able to use our products to fix our product. And in those cases, we have a fallback which is jump into—

Corey: I assume you put a little thought into what happened. “Well, what if our product is down?” “Oh well, I guess we’ll never be able to fix it or communicate about it.” It seems like that’s the sort of thing that, given what you do, you might have put more than ten seconds of thought into.

Chris: We’ve put a fair amount of thought into it. But at the end of the day, [laugh] it’s like if stuff is down, like, what do you need to do? You need to communicate with people. So, jump on a Google Chat, jump on a Slack huddle, whatever else it is we have various different, like, fallbacks in different order. And at the core of it, I think this is the thing is, like, you cannot be prepared for every single thing going wrong, and so what you can be prepared for is to be unprepared and just accept that humans are incredibly good at being resilient, and therefore, all manner of things are going to happen that you’ve never seen before and I guarantee you will figure them out and fix them, basically.

But yeah, I say this; if my SOC 2 auditor is listening, we also do have a very well-defined, like, backup plan in our SOC 2 [laugh] in our policies and processes that is the thing that we will follow that. But yeah.

Corey: The fact that you’re saying the magic words of SOC 2, yes, exactly. Being in a responsible adult and living up to some baseline compliance obligations is really the sign of a company that’s put a little thought into these things. So, as I pull up incident.io—the website, not the company to be clear—and look through what you’ve written and how you talk about what you’re doing, you’ve avoided what I would almost certainly have not because your tagline front and center on your landing page is, “Manage incidents at scale without leaving Slack.” If someone were to reach out and say, well, we’re down all the time, but we’re using Microsoft Teams, so I don’t know that we can use you, like, the immediate instinctive response that I would have for that to the point where I would put it in the copy is, “Okay, this piece of advice is free. I would posit that you’re down all the time because you’re the kind of company to use Microsoft Teams.” But that doesn’t tend to win a whole lot of friends in various places. In a slightly less sarcastic bent, do you see people reaching out with, “Well, we want to use you because we love what you’re doing, but we don’t use Slack.”

Chris: Yeah. We do. A lot of folks actually. And we will support Teams one day, I think. There is nothing especially unique about the product that means that we are tied to Slack.

It is a great way to distribute our product and it sort of aligns with the companies that think in the way that we do in the general case but, like, at the core of what we’re building, it’s a platform that augments a communication platform to make it much easier to deal with a high-stress, high-pressure situation. And so, in the future, we will support ways for you to connect Microsoft Teams or if Zoom sought out getting rich app experiences, talk on a Zoom and be able to do various things like logging actions and communicating with other systems and things like that. But yeah, for the time being very, very deliberate focus mechanism for us. We’re a small company with, like, 30 people now, and so yeah, focusing on that sort of very slim vertical is working well for us.

Corey: And it certainly seems to be working to your benefit. Every person I’ve talked to who is encountered you folks has nothing but good things to say. We have a bunch of folks in common listed on the wall of logos, the social proof eye chart thing of here’s people who are using us. And these are serious companies. I mean, your last job before starting incident.io was at Monzo, as you mentioned.

You know what you’re doing in a regulated, serious sense. I would be, quite honestly, extraordinarily skeptical if your background were significantly different from this because, “Well, yeah, we worked at Twitter for Pets in our three-person SRE team, we can tell you exactly how to go ahead and handle your incidents.” Yeah, there’s a certain level of operational maturity that I kind of just based upon the name of the company there; don’t think that Twitter for Pets is going to nail. Monzo is a bank. Guess you know what you’re talking about, given that you have not, basically, been shut down by an army of regulators. It really does breed an awful lot of confidence.

But what’s interesting to me is the number of people that we talk to in common are not themselves banks. Some are and they do very serious things, but others are not these highly regulated, command-and-control, top-down companies. You are nimble enough that you can get embedded at those startup-y of startup companies once they hit a certain point of scale and wind up helping them arrive at a better outcome. It’s interesting in that you don’t normally see a whole lot of tools that wind up being able to speak to both sides of that very broad spectrum—and most things in between—very effectively. But you’ve somehow managed to thread that needle. Good work.

Chris: Thank you. Yeah. What else can I say other than thank you? I think, like, it’s a deliberate product positioning that we’ve gone down to try and be able to support those different use cases. So, I think, at the core of it, we have always tried to maintain the incident.io should be installable and usable in your very first incident without you having to have a very steep learning curve, but there is depth behind it that allows you to support a much more sophisticated incident setup.

So, like, I mean, you mentioned Monzo. Like, I just feel incredibly fortunate to have worked at that company. I joined back in 2017 when they were, I don’t know, like, 150,000 customers and it was just getting its banking license. And I was there for four years and was able to then see it scale up to 6 million customers and all of the challenges and pain that goes along with that both from building infrastructure on the technical side of things, but from an organizational side of things. And was, like, front-row seat to being able to work with some incredibly smart people and sort of see all these various different pain points.

And honestly, it feels a little bit like being in sort of a cheat mode where we get to this import a lot of that knowledge and pain that we felt at Monzo into the product. And that happens to resonate with a bunch of folks. So yeah, I feel like things are sort of coming out quite well at the moment for folks.

Corey: The one thing I will say before we wind up calling this an episode is just how grateful I am that I don’t have to think about things like this anymore. There’s a reason that the problem that I chose to work on of expensive AWS bills being very much a business-hours only style of problem. We’re a services company. We don’t have production infrastructure that is externally facing. “Oh, no, one of our data analysis tools isn’t working internally.”

That’s an interesting curiosity, but it’s not an emergency in the same way that, “Oh, we’re an ad network and people are looking at ads right now because we’re broken,” is. So, I am grateful that I don’t have to think about these things anymore. And also a little wistful because there’s so much that you do it would have made dealing with expensive and dangerous outages back in my production years a lot nicer.

Chris: Yep. I think that’s what a lot of folks are telling us essentially. There’s this curious thing with, like, this product didn’t exist however many years ago and I think it’s sort of been quite emergent in a lot of companies that, you know, as sort of things have moved on, that something needs to exist in this little pocket of space, dealing with incidents in modern companies. So, I’m very pleased that what we’re able to build here is sort of working and filling that for folks.

Corey: Yeah. I really want to thank you for taking so much time to go through the ethos of what you do, why you do it, and how you do it. If people want to learn more, where’s the best place for them to go? Ideally, not during an incident.

Chris: Not during an incident, obviously. Handily, the website is the company name. So, incident.io is a great place to go and find out more. We’ve literally—literally just today, actually—launched our Practical Guide to Incident Management, which is, like, a really full piece of content which, hopefully, will be useful to a bunch of different folks.

Corey: Excellent. We will, of course, put a link to that in the [show notes 00:29:52]. I really want to thank you for being so generous with your time. Really appreciate it.

Chris: Thanks so much. It’s been an absolute pleasure.

Corey: Chris Evans, Chief Product Officer and co-founder of incident.io. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this episode, please leave a five-star review on your podcast platform of choice along with an angry comment telling me why your latest incident is all the intern’s fault.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Maish

Maish Saidel-Keesing is a Senior Enterprise Developer Advocate @AWS working on containers and has been working in IT for the past 20 years and with a stronger focus on cloud and automation for the past 7.

He has extensive experience with AWS Cloud technologies, DevOps and Agile practices and implementations, containers, Kubernetes, virtualization and, and a number of fun things he has done along the way

He is constantly trying to bridge the gap between Developers and Operators to allow all of us provide a better service for our customers (and not wake up from pages in the middle of the night). He is an avid practitioner of dissolving silos - educating Ops how to code and explaining to Devs what the hell is Operations

Links Referenced:

  • @maishsk: https://twitter.com/maishsk
  • duckbillgroup.com: https://duckbillgroup.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored by our friends at Fortinet. Fortinet’s partnership with AWS is a better-together combination that ensures your workloads on AWS are protected by best-in-class security solutions powered by comprehensive threat intelligence and more than 20 years of cybersecurity experience. Integrations with key AWS services simplify security management, ensure full visibility across environments, and provide broad protection across your workloads and applications. Visit them at AWS re:Inforce to see the latest trends in cybersecurity on July 25-26 at the Boston Convention Center. Just go over to the Fortinet booth and tell them Corey Quinn sent you and watch for the flinch. My thanks again to my friends at Fortinet.

Corey: Let’s face it, on-call firefighting at 2am is stressful! So there’s good news and there’s bad news. The bad news is that you probably can’t prevent incidents from happening, but the good news is that incident.io makes incidents less stressful and a lot more valuable. incident.io is a Slack-native incident management platform that allows you to automate incident processes, focus on fixing the issues and learn from incident insights to improve site reliability and fix your vulnerabilities. Try incident.io, recover faster and sleep more.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’m a cloud economist at The Duckbill Group, and that was a fun thing for me to become because when you’re starting to set out to solve a problem, well, what do you call yourself? I find that if you create a job title for yourself, well, no one knows quite how to categorize you and it leads to really interesting outcomes as a result. My guest today did something very similar. Maish Saidel-Keesing is an EntReloper, or Enterprise Developer Advocate, specifically for container services at AWS. Maish, thank you for joining me.

Maish: Thank you for having me on the show, Corey. It’s great to be here.

Corey: So, how did you wind up taking a whole bunch of words such as enterprise, developer, advocate because I feel like the way you really express seniority at big companies, almost as a display of dominance, is to have additional words in your job title, which all those words are very enterprise-y, very business-y, and very serious. And in container services to boot, which is a somewhat interesting culture, just looking at the enterprise adoption of the pattern. And then at AWS, whose entire sense of humor can be distilled down into, “That’s not funny.” You have the flexibility to refer to yourself as an EntReloper in public. I love it. Is it just something you started doing? Was there, like, 18 forms of approval you had to go through to do it? How did this happen? I love it.

Maish: So, no. I didn’t have to go through approval, of course. Same way, you didn’t call yourself a cloud economist with anybody else’s approval. But I got the idea mostly from you because I love your term of coining everybody who’s in developer advocacy or developer relations as a DevReloper. And specifically, the reason that I coined the term of an EntReloper—and actually looked it up on Google to see if anybody had actually used that term before, and no they haven’t—it’s the fact of I came into the position on the premise of trying to bring the enterprise voice of the customer into developer advocacy.

When we speak about developer advocates today, most of them are the people who are the small startups, developers who write the code, and we kind of forget that there is a whole big world outside of, besides small startups, which are these big, massive, behemoth sort of enterprise companies who kind of do things differently because they’ve been around for many, many years; they have many, many silos inside their organizations. And it’s not the most simple thing to open up your laptop, and install whatever software you want on, because some of these people don’t even give you admin rights on your laptop, or you’re allowed to ssh out to a computer in the cloud because also the same thing: everything is blocked by corporate firewalls where you have to put in a ticket in order to get access to the outside world. I worked in companies like that when I was—before I moved to Amazon. So, I want to bring that perspective to the table on behalf of our customers.

Corey: Bias is a very funny thing. I spent the overwhelming majority of my career in small environments like you describe. To me a big company is one that has 200 people there, and it turns out that there’s a whole ‘nother sense of scale that goes beyond that. And there’s, like, 18 different tiers beyond. But I still bias based upon my own experiences when I talk about how I do things and how I think about things to a certain persona that closely resembles my own experiences where, “Just install this thing as a tool and it’ll be great,” ignoring entirely, the very realistic fact that you’ve got an entire universe out there of people who are not empowered to install things on their own laptop, for example.

How is developer advocacy different within enterprises than in the common case of, “We’re a startup. We’re going to change the world with our amazing SaaS.” Great, maybe you will. Statistically, you won’t. But enterprises have different concerns, different challenges, and absolutely a different sense of scale. How is the practice of advocacy different in those environments?

Maish: So, I think the fact is, mostly working on standardization from the get-go that these big enterprises want things to work in a standard way where they can control it, they can monitor it, they can log everything, they can secure it mostly, of course, the most important thing. But it’s also the fact that as a developer advocate, you don’t always talk to developers within the enterprise. You also have to talk to the security team and to the network team and to the business itself or the C-level to understand. And as you also probably have found out as well in your job, you connect the people with inside the business one to another, these different groups, and get them talking to each other to make these decisions together. So, we act as kind of a bridge in between the people with inside their own company where they don’t really talk to each other, or don’t have the right connections, or the right conduit in order for them to start that conversation and make things better for themselves.

Corey: On some level, my line about developer relations, developer advocacy, has generally taken the track of, “What does that mean? Well, it means you work in marketing, but they’re scared to tell you that.” Do you view what you do as being within marketing, aligned with marketing, subtly different and I’m completely wrong, et cetera, et cetera? All positions are legitimate, by the way.

Maish: So, I think, at the position that I’m currently in, which is a developer advocate but for the service team, is slightly different than a marketing developer advocate. The marketing developer advocate—and we have many of them which are amazing people and doing amazing work within AWS—their job is to teach everybody about the services and the capabilities available within AWS. That is also part of my job, but I would think that is the 40% of my job. I also go on stage, I go on podcasts like this, I present at conferences, I write blog posts. I also do the kind of marketing work as well.

But the other 60% of my job as a service developer advocate is to seek out the feedback, or the signals, or the sentiment from our customers, and bring that back into the service teams, into the product management, into the engineering teams. And, as I said, sit as the enterprise customer in the chair in those meetings, to voice their concerns… their opinions, how they would like the products to go, how we can make the products better. So, the 60% is mostly what we call inbound, which is taking feedback from our customers back into the service teams directly in order to have some influence on the roadmap. And 40% is the outbound work, which we do, as I said, conferences, blog posts, and things like this.

Corey: I have a perception. And I am thrilled to be corrected on this because it’s not backed by data; it’s backed by my own biases—and some people tend to conflate the two; I strive not to—that there’s a—I think the term that I heard bandied around at one point was ‘the dark matter developers.’ These are folks that primarily work in .NET or Java. They work for companies that are not themselves tech companies, but rather tech is a supporting function, usually in a central IT-style organization, that supports what the business actually does, and they generally are not visible to a lot of traditional developer advocacy approaches.

They, by and large, don’t go to conferences, they don’t go on Twitter to yell at people about things, they commit the terrible sin—according to many startup folks—of daring to view the craft of writing software as this artistic thing, and they just view it as a job and a thing to make money for—filthy casuals—as opposed to this higher calling that’s changing the world. Which I think is wild take. But there are a tremendous number of people out there who do fit the profile of they show up, they do their jobs working on this stuff, they don’t go to conferences, they don’t go out into the community, and they just do their job and go home. The end. Is that an accurate perception? Are there large swaths of folks like that in the industry, and if so, do they centralize or congregate more around enterprises than they do around smaller companies?

Maish: I think that your perception is correct. Specifically, for my experience, when I worked, for example, my first two years before I was a developer advocate, I was an enterprise solutions architect which I worked with financial institutions, which are banks, which usually have software which are older than me, which are written in languages, which are older than I am. So, there are people which, as they say, they come there to—they do their job. They’re not interested in looking at Twitter, or writing blog posts, or participating in any kind of thing which is outgoing. And they just, they’re there to write the code. They go home at the end of the day.

They also usually don’t have pagers that page them in middle of the night because that’s what you have operations teams for, not developers because they’re completely different entities. So, I do think your perception might be correct, yes. There are people like that when you say, these dark matter people, dark matter developers.

Corey: And I don’t have any particular problem. I’m not here to cast shade on anything that they’re doing, to be very clear.

Maish: Not at all.

Corey: Everyone makes different choices and that’s great. I don’t think necessarily everyone should have a job that is all-consuming, that eats them alive. I wish I didn’t, some days. [laugh]. The challenge I have for you then is, as an EntReloper, how do you reach folks in positions like that? Or don’t you?

Maish: I think the way to reach those people is to firstly, expose them to technology, expose them to the capabilities that they can use in AWS in the cloud, specifically with my position in container services, and gain their trust because that’s one of the LPs in Amazon itself: customer obsession. And we work consistently in order to—with our customers to gain their trust and help them along their journey, whatever it may be. If it might be the fact, okay, I only want to write software for nine to five and go home and do everything afterwards, which most normal people do without having to worry about work, or they still want to continue working and adopt the full model of you build it, you own it; manage everything in production on their own and go into the new world of modern software, which many enterprises, unfortunately, are not all the way there yet, but hopefully, they will get there sooner than later.

Corey: There’s a misguided perception in many corners that you have to be able to reach everyone at all times; wherever they are, you have to be able to go there. I don’t think that’s true. I think that showing up and badgering people who are just trying to get a job done into, “Hey, have you heard the good word of cloud?” It’s like, evangelists knocking on your door at seven o’clock in the morning on a weekend and you’re trying to sleep in because the kids are somewhere else for the week. Yeah, I might be projecting a little bit on that.

I think that is the wrong direction to go. And I find that being able and willing to meet people where they are is key to success on this. I’m also a big believer in the idea that in any kind of developer advocacy role, regardless whether their targets are large, small, or in my case, patently ridiculous because my company is in fact ridiculous in some ways, you have to meet them where they are. There’s no choice around that. Do you find that there are very different concerns that you have to wind up addressing with your audience versus a more, “Mainstream,” quote-unquote, developer advocacy role?

Maish: For the enterprise audience, they need to, I would say, relate to what we’re talking to. For an example, I gave a talk a couple of weeks ago on the AWS Summit here in Tel Aviv, of how to use App Runner. So, instead of explaining to the audiences how you use the console, this is what it does, you can deploy here, this is how the deployments work, blue, green, et cetera, et cetera, I made up an imaginary company and told the story of how the three people in the startup of this company would start working using App Runner in order to make the thing more relatable, something which people can hopefully remember and understand, okay, this is something which I would do as a startup, or this is what my project, which I’m doing or starting to work on, something I can use. So, to answer your question, in two words, tell stories instead of demo products.

Corey: It feels like that’s a… heavy lift, in many cases, because I guess it’s also partially a perception issue on my part, where I’m looking at this across the board, where I see a company that has 5000 developers working there and, like, how do you wind up getting them to adopt cloud, or adopt new practices, or change anything? It feels like it’s a Herculean, impossible task. But in practice, I feel like you don’t try and do all of that at once. You start with small teams, you start with specific projects, and move on. Is that directionally accurate?

Maish: Completely accurate. There’s no way to move a huge mothership in one direction at one time. You have to do, as you say, start small, find the projects, which are going to bring value to the company or the business, and start small with those projects and those small teams, and continue that education within the organization and help the people with your teaching or introducing them to the cloud, to help others within inside their own organization. Make them, or enable them, or empower them to become leaders within their own organization. That’s what I tried to do, at least.

Corey: You and I have a somewhat similar background, which is weird given that we’ve just spent a fair bit of time talking about how different our upbringings were in tech at scales of companies and whatnot, but we’re alike in that we are both fairly crusty, old operations-side folks, sysadmins—

Maish: [laugh]. Yep.

Corey: —grumpy people.

Maish: Grumpy old sysadmins. Yeah, exactly.

Corey: Exactly. Because do you ever notice there’s never a happy one? Imagine that. And DevOps was always a meeting of the development and operations, meaning everyone’s unhappy. And there’s a school of thought that—like, I used to think that, “Oh, this is just what we call sysadmins once they want a better title and more money, but it’s still the same job.”

But then I started meeting a bunch of DevOps types who had come from the exact opposite of our background, where they were software developers and then they wound up having to learn not so much how the code stuff works the way that we did, but rather how systems work, how infrastructure works. Compare and contrast those for me. Who makes, I guess, the more successful DevOps engineer when you look at it through that lens?

Maish: So, I might be crucified for this on the social media from a number of people from the other side of the fence, but I have the firm belief that the people who make the best DevOps engineers—and I hate that term—but people who move DevOps initiatives or changes or transformations with organization is actually the operations people because they usually have a broader perspective of what is going on around them besides writing code. Too many times in my career, I’ve been burned by DNS, by a network cable, by a power outage, by somebody making a misconfiguration in the Puppet module, or whatever it might have been, somebody wrote it to deploy to 15,000 machines, whatever it may be. These are things where developers, at least my perception of what developers have been doing up until now, don’t really do that. In a previous organization I used to work for, the fact was, there was a very, very clear delineation about between the operations people, and the developers who wrote the software. We had very hard times getting them into rotations for on-call, we had very hard times educating them about the fact that not every single log line has to be written to the log because it doesn’t interest anybody.

But from developer perspective, of course, we need that log because we need to know what’s happening in the end. But there are 15, different thousand… turtles all the way down, which have implications about the number of log lines which are written into a piece of software. So, I am very much of the belief that the people that make the best DevOps engineers—if we can use that term still today—are actually people which come from an operations background because it’s easier to teach them how to write code or become a programmer than the other way round of teaching a developer how to become an operations person. So, the change or the move from one direction from operations to adding the additional toil of writing software is much easier to accomplish than the other way around, from a developer learning how to run infrastructure at scale.

Corey: This episode is sponsored in part by our friend EnterpriseDB. EnterpriseDB has been powering enterprise applications with PostgreSQL for 15 years. And now EnterpriseDB has you covered wherever you deploy PostgreSQL on-premises, private cloud, and they just announced a fully-managed service on AWS and Azure called BigAnimal, all one word. Don’t leave managing your database to your cloud vendor because they’re too busy launching another half-dozen managed databases to focus on any one of them that they didn’t build themselves. Instead, work with the experts over at EnterpriseDB. They can save you time and money, they can even help you migrate legacy applications—including Oracle—to the cloud. To learn more, try BigAnimal for free. Go to biganimal.com/snark, and tell them Corey sent you.

Corey: I once believed much the same because—and it made sense coming from the background that I was in. Everyone intellectually knows that if you’re having trouble with a piece of equipment, have you made sure that it’s plugged in? Yes, everyone knows that intellectually. But there’s something about having worked on a thing for three hours that wasn’t working and only discover it wasn’t plugged in, that really sears that lesson into your bones. The most confidence-inspiring thing you can ever hear from someone an operations role is, “Oh, I’ve seen this problem before. Here’s how we fixed it.”

It feels like there are no junior DevOps engineers, for lack of a better term. And for a long time, I believed that the upcoming and operational side of the world were in fact, the better DevOps types. And in the fullness of time, I think a lot of that—at least my position on it—was rooted in some level of insecurity because I didn’t know how to write code and the thing that I saw happening was my job that I had done historically was eroding. Today, I don’t know that it’s possible to be in the operation space and not be at least basically conversant with how code works. There’s a reason most of these job interviews turn into algorithm hazing.

And my articulation of it was rooted, for me at least—at least in a small way—in a sense of defensiveness and wanting to validate the thing that I had done with my career that I defined myself by, I was under threat. And obviously, the thing that I do is the best thing because otherwise it’s almost a tacit admission that I made poor career choices at some point. And I don’t think that’s true, either. But for me, at least psychologically, it was very much centered in that. And honestly, I found that the right answer for me was, in fact, neither of those two things because I have met a couple of people in my life that I would consider to be full-stack engineers.

And there’s a colloquialism these days, that means oh, you do front-end and back-end. Yeah. The people I’m thinking of did front-end, they did back-end, they did mobile software, they did C-level programming, they wrote their own freaking device drivers at one point. Like, they have done basically everything. And they were the sort of person you could throw any technical issue whatsoever at and get out of their way because it was going to get solved. Those people are, as it turns out, the best. Like, who does a better job developer or operations, folks? Yes. Specifically, both of those things together.

Maish: Exactly.

Corey: And I think that is a hard thing to talk about. I think that it’s a hard—it’s certainly a hard thing to find. It turns out that there’s a reason that I only know two or three of those folks in the course of my entire career. They’re out there, but they’re really, really hard to track down.

Maish: I completely agree.

Corey: A challenge that I hear articulated in some cases—and while we’re saying things that are going to get us yelled out on social media, let’s go for the fences on this one—a concern that comes up when talking about enterprises moving to cloud is that they have a bunch of existing sysadmin types—while we’re on the topic—and well, those people need to learn to work within cloud. And the reality is, in many cases that first, that’s a whole new skill set that not everyone is going to be willing or able to pick up. For those who can they have just found that their market rate has effectively doubled. And that seems, on some level, to pose a significant challenge to companies undergoing this, and the larger the company, the more significant the challenge.

Because it’s my belief that you pay market rate for the talent you have whether you want to or not. And if companies don’t increase compensation, these people will leave for things that double their income. And if they raise compensation internally, good for them, but that does have a massive drag on their budget that may not have been accounted for in a lot of the TCO analyses. How do you find that the companies you talk to wind up squaring that circle?

Maish: I don’t think I have a correct answer for that. I do completely agree—

Corey: Oh, I’m not convinced there’s a correct answer at all. I’m just trying [laugh] to figure out how to even think about it.

Maish: I… have seen this as well in companies which I used to work for and companies that were customers that I have also worked with as part of my tenure in AWS. It’s the fact of, when companies are trying to move to the cloud and they start upskilling their people, there’s always the concern in the back of their mind of the fact, “Okay, I’m now training this person with new technology. I’m investing time, I’m investing money. And why would I do this if I know that, for example, as soon as I finish this, I’m going to have to just say, I have to pay them more because they can go somewhere else and get the same job with a better pay? So, why would we invest amount of time and resources into upscaling the people?”

And these are questions which I have received and conversations which I’ve had with customers many times over the last two, three years. And the answer, from my perspective always, is the fact is because, number one, you’re making the world a better place. Number two, you’re making your employees feel more appreciated, giving them better knowledge. And if you’re afraid of the fact of teaching somebody to become better is going to have negative effects on your organization then, unfortunately, you deserve to have that person leave and let them find a better job because you’re not taking good care of your people. And it’s sometimes hard for companies to hear that.

Sometimes we get, “You know what? You’re completely right.” Sometimes I don’t agree with you because I need to compete there, get to the bottom line, and make sure that I stay within my budget or my TCO. But the most important thing is to have the conversation, let people hear different ideas, see how it can benefit them, not only by giving people more options to maybe leave the company, but it can actually make their whole organization a lot better in the long run.

Corey: I think that you have to do right by people because reputations last a long time. Even at big companies it becomes a very slow thing to change and almost impossible to do in the short term. So, people tell stories when they feel wronged. That becomes a problem. I do want to pivot a little bit because you’re not merely an EntReloper; you are an EntReloper specifically focusing on container services.

Maish: Correct.

Corey: Increasingly, I am viewing containers as what amounts to effectively a packaging format. That is the framework through which I am increasingly seeing. How are you seeing customers use containers? Is that directionally correct? Is it completely moonbat stuff compared to what you’re seeing in the wild, or something else?

Maish: I don’t think it’s a packaging format; I think it’s more as an accelerator to enable the customers to develop in a more modern way with using twelve-factor apps with modern technology and not necessarily have their own huge, sticky, big monolith of whatever it might be, written in C# or whatever, or C++ whatever it may be, as they’ve been using up until now, but they now have the option and the technology and the background in order to split it up into smaller services and develop in the way that most of the modern world—or at least, the what we perceive as the modern world—is developing and creating applications today.

Corey: I feel like on some level, containers were a radical change to how companies envisioned software. They definitely provide a path of modernizing things that were very tied to hardware previously. It let some companies even just leapfrog the virtualization migration that they’d been considering doing. But, on some level, I also feel like it runs counter to the ideas of DevOps, where you have development and operations working in partnership, where now it’s like, welp, inside the container is a development thing and outside the container, ops problem now. It feels almost, on some level, like, it reinforces a wall. But in a lot of cultures and a lot of companies, that wall is there and there’s no getting rid of it anytime soon. So, I confess that I’m conflicted on that.

Maish: I think you might be right, and it depends, of course, on the company and the company culture, but what I think that companies need to do is understand that there will never be one hundred percent of people writing software that want to know one hundred percent of how the underlying infrastructure works. And the opposite direction as well: that there will never be people which maintain infrastructure and understand how computers and CPUs and memory buses and NUMA works on motherboards, that they don’t need to know how to write the most beautiful enticing and wonderful software for programs, for the world. There’s always going to have to be a compromise of who’s going to be doing this or who’s going to be doing that, and how comfortable they are with taking at least part of the responsibility of the other side into their own realm of what they should be doing. So, there’s going to be a compromise on both sides, but there is some kind of divide today of separating, okay, you just write the Helm chart for your Kubernetes Pod spec, or your ECS task, or whatever task definition, whatever you would like. And don’t worry about the things in the background because they’re just going to magically happen in the end. But they do have to understand exactly what is happening at the background in the end because if something goes wrong, and of course, something will go wrong, eventually, one day somewhere, somehow, they’re going to have to know how to take care of it.

Corey: I really want to thank you for taking the time to speak with me today about, well, I guess a wide ranging variety of topics, some of which will absolutely inspire people to take to their feet—or at least their Twitter accounts—and tell us, “You know what your problem is?” And I honestly live for that. If you don’t evoke that kind of reaction on some level, have you ever really had an opinion in the first place? So, I’m looking forward to that. If people want to learn more about you, your beliefs, call set beliefs misguided, et cetera, et cetera, where’s the best place to find you?

Maish: So, I’m on Twitter under @maishsk. I assume that will be in the [show notes 00:26:31]. I pontificate some time on technology, on cooking every now and again, on Friday before the end of the weekend, a little bit of politics, but you can find me @maishsk on Twitter. Or maishsk everywhere else social that’s possible.

Corey: Excellent. We will toss links to that, of course, in the [show notes 00:26:50]. Thank you so much for being so generous with your time. I appreciate it.

Maish: Thank you very much, Corey. It was fun.

Corey: Maish Saidel-Keesing, EntReloper of container services at AWS. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated it, please leave a five-star review on your podcast platform of choice along with an angry comment that your 5000 enterprise developer colleagues can all pile on.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Chris

Chris is a robotics engineer turned cloud security practitioner. From building origami robots for NASA, to neuroscience wearables, to enterprise software consulting, he is a passionate builder at heart. Chris is a cofounder of Common Fate, a company with a mission to make cloud access simple and secure.

Links:

  • Common Fate: https://commonfate.io/
  • Granted: https://granted.dev
  • Twitter: https://twitter.com/chr_norm

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Let’s face it, on-call firefighting at 2am is stressful! So there’s good news and there’s bad news. The bad news is that you probably can’t prevent incidents from happening, but the good news is that incident.io makes incidents less stressful and a lot more valuable. incident.io is a Slack-native incident management platform that allows you to automate incident processes, focus on fixing the issues and learn from incident insights to improve site reliability and fix your vulnerabilities. Try incident.io, recover faster and sleep more.

Corey: This episode is sponsored in part by Honeycomb. When production is running slow, it’s hard to know where problems originate. Is it your application code, users, or the underlying systems? I’ve got five bucks on DNS, personally. Why scroll through endless dashboards while dealing with alert floods, going from tool to tool to tool that you employ, guessing at which puzzle pieces matter? Context switching and tool sprawl are slowly killing both your team and your business. You should care more about one of those than the other; which one is up to you. Drop the separate pillars and enter a world of getting one unified understanding of the one thing driving your business: production. With Honeycomb, you guess less and know more. Try it for free at honeycomb.io/screaminginthecloud. Observability: it’s more than just hipster monitoring.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. It doesn’t matter where you are on your journey in cloud—you could never have heard of Amazon the bookstore—and you encounter AWS and you spin up an account. And within 20 minutes, you will come to the realization that everyone in this space does. “Wow, logging in to AWS absolutely blows goats.”

Today, my guest, obviously had that reaction, but unlike most people I talked to, decided to get up and do something about it. Chris Norman is the co-founder of Common Fate and most notably to how I know him is one of the original authors of the tool, Granted. Chris, thank you so much for joining me.

Chris: Hey, Corey, thank you for having me.

Corey: I have done podcasts before; I have done a blog post on it; I evangelize it on Twitter constantly, and even now, it is challenging in a few ways to explain holistically what Granted is. Rather than trying to tell your story for you, when someone says, “Oh, Granted, that seems interesting and impossible to Google for in isolation, so therefore, we know it’s going to be good because all the open-source projects with hard to find names are,” what is Granted and what does it do?

Chris: Granted is a command-line tool which makes it really easy for you to get access and assume roles when you’re working with AWS. For me, when I’m using Granted day-to-day, I wake up, go to my computer—I’m working from home right now—crack open the MacBook and I log in and do some development work. I’m going to go and start working in the cloud.

Corey: Oh, when I start first thing in the morning doing development work and logging into the cloud, I know. All right, I’m going to log in to AWS and now I know that my day is going downhill from here.

Chris: [laugh]. Exactly, exactly. I think maybe the best days are when you don’t need to log in at all. But when you do, I go and I open my terminal and I run this command. Using Granted, I ran this assume command and it authenticates me with single-sign-on into AWS, and then it opens up a console window in a particular account.

Now, you might ask, “Well, that’s a fairly standard thing.” And in fact, that’s probably the way that the console and all of the tools work by default with AWS. Why do you need a third-party tool for this?

Corey: Right. I’ve used a bunch of things that do varying forms of this and unlike Granted, you don’t see me gushing about them. I want to be very clear, we have no business relationship. You’re not sponsoring anything that I do. I’m not entirely clear on what your day job entails, but I have absolutely fallen in love with the Granted tool, which is why I’m dragging you on to this show, kicking and screaming, mostly to give me an excuse to rave about it some more.

Chris: [laugh]. Exactly. And thank you for the kind words. And I’d say really what makes it special or why I’ve been so excited to be working on it is that it makes this access, particularly when you’re working with multiple accounts, really, really easy. So, when I run assume and I open up that console window, you know, that’s all fine and that’s very similar to how a lot of the other tools and projects that are out there work, but when I want to open that second account and that second console window, maybe because I’m looking at like a development and a staging account at the same time, then Granted allows me to view both of those simultaneously in my browser. And we do that using some platform sort of tricks and building into the way that the browser works.

Corey: Honestly, one of the biggest differences in how you describe what Granted is and how I view it is when you describe it as a CLI application because yes, it is that, but one of the distinguishing characteristics is you also have a Firefox extension that winds up leveraging the multi-container functionality extension that Firefox has. So, whenever I wind up running a single command—assume with a-c’ flag, then I give it the name of my AWS profile, it opens the web console so I can ClickOps my heart’s content inside of a tab that is locked to a container, which means I can have one or two or twenty different AWS accounts and/or regions up running simultaneously side-by-side, which is basically impossible any other way that I’ve ever looked at it.

Chris: Absolutely, yeah. And that’s, like, the big differentiating factor right now between Granted and between this sort of default, the native experience, if you’re just using the AWS command line by itself. With Granted, you can—with these Firefox containers, all of your cookies, your profile, everything is all localized into that one container. It’s actually it’s a privacy features that are built into Firefox, which keeps everything really separate between your different profiles. And what we’re doing with Granted is that we make it really easy to open a specific profiles that correspond with different AWS profiles that you’re using.

So, you’d have one which could be your development account, one which could be production or staging. And you can jump between these and navigate between them just as separate tabs in your browser, which is a massive improvement over, you know, what I’ve previously had to use in the past.

Corey: The thing that really just strikes me about this is first, of course, the functionality and the rest, so I saw this—I forget how I even came across it—and immediately I started using it. On my Mac, it was great. I started using it when I was on the road, and it was less great because you built this thing in Go. It can compile and install on almost anything, but there were some assumptions that you had built into this in its early days that did not necessarily encompass all of the use cases that I use. For example, it hadn’t really occurred to you that some lunatic would try and only use an iPad when they’re on the road, so they have to be able to run this to get federated login links via SSHing into an EC2 instance running somewhere and not have it open locally.

You seemed almost taken aback when I brought it up. Like, “What lunatic would do that?” Like, “Hi, I’m such a lunatic. Let’s talk about this.” And it does that now, and it’s awesome. It does seem to me though, and please correct me if I’m wrong on this assumption slash assessment that this is first and foremost aimed at desktop users, specifically people running Mac on the desktop, is that the genesis of it?

Chris: It is indeed. And I think part of the cause behind that is that we originally built a tool for ourselves. And as we were building things and as we were working using the cloud, we were running things—you know, we like to think that we’re following best practices when we’re using AWS, and so we’d set up multiple accounts, we’d have a special account for development, a separate one for staging, a separate one for production, even internal tools that we would build, we would go and spin up an individual account for those. And then you know, we had lots of accounts. and to go and access those really easily was quite difficult.

So, we definitely, we built it for ourselves first and I think that that’s part of when we released it, it actually a little bit of cause for some of the initial problems. And some of the feedback that we had was that it’s great to build tools for yourself, but when you’re working in open-source, there’s a lot of different diversity with how people are using things.

Corey: We take different approaches. You want to try to align with existing best practices, whereas I am a loudmouth white guy who works in tech. So, what I do definitionally becomes a best practice in the ecosystem. It’s easier to just comport with the ones that are already existing that smart people put together rather than just trying to competence your way through it, so you took a better path than I did.

But there’s been a lot of evolution to Granted as I’ve been using it for a while. I did a whole write-up on it and that got a whole bunch of eyes onto the project, which I can now admit was a nefarious plan on my part because popping into your community Slack and yelling at you for features I want was all well and good, but let’s try and get some people with eyes on this who are smarter than me—which is not that high of a bar when it comes to SSO, and IAM, and federated login, and the rest—and they can start finding other enhancements that I’ll probably benefit from. And sure enough, that’s exactly what happened. My sneaky plan has come to fruition. Thanks for being a sucker, I guess. I mean—[laugh] it worked. I’m super thrilled by the product.

Chris: [laugh]. I guess it’s a great thing I think that the feedback and particularly something that’s always been really exciting is just seeing new issues come through on GitHub because it really shows the kinds of interesting use cases and the kinds of interesting teams and companies that are using Granted to make their lives a little bit easier.

Corey: When I go to the website—which again is impossible to Google—the website for those wondering is granted.dev. It’s short, it’s concise, I can say it on a podcast and people automatically know how to spell it. But at the top of the website—which is very well done by the way—it mentions that oh, you can, “Govern access to breakglass roles with Common Fate Cloud,” and it also says in the drop shadow nonsense thing in the upper corner, “Brought to you by Common Fate,” which is apparently the name of your company.

So, the question I’ll get to in a second is what does your company do, but first and foremost, is this going to be one of those rug-pull open-source projects where one day it’s, “Oh, you want to log into your AWS accounts? Insert quarter to continue.” I’m mostly being a little over the top with that description, but we’ve all seen things that we love turn into molten garbage. What is the plan around this? Are you about to ruin this for the rest of us once you wind up raising a round or something? What’s the deal?

Chris: Yeah, it’s a great question, Corey. And I think that to a degree, releasing anything like this that sits in the access workflow and helps you assume roles and helps you day-to-day, you know, we have a responsibility to uphold stability and reliability here and to not change things. And I think part of, like, not changing things includes not [laugh] rug-pulling, as you’ve alluded to. And I think that for some companies, it ends up that open-source becomes, like, a kind of a lead-generation tool, or you end up with, you know, now finally, let’s go on add another login so that you have to log into Common Fate to use Granted. And I think that, to be honest, a tool like this where it’s all about improving the speed of access, the incentives for us, like, it doesn’t even make sense to try and add another login for to try to get people to, like, to say, login to Common Fate because that would make your signing process for AWS take even longer than it already does.

Corey: Yeah, you decided that you know, what’s the biggest problem? Oh, you can sleep at night, so let’s go ahead and make it even worse, by now I want you to be this custodian of all my credentials to log into all of my accounts. And now you’re going to be critical path, so if you’re down, I’m not able to log into anything. And oh, by the way, I have to trust you with full access to my bank stuff. I just can’t imagine that is a direction that you would be super excited about diving head-first into.

Chris: No, no. Yeah, certainly not. And I think that the, you know, building anything in this space, and with what we’re doing with Common Fate, you know, we’re building a cloud platform to try to make IAM a little bit easier to work with, but it’s really sensitive around granting any kind of permission and I think that you really do need that trust. So, trying to build trust, I guess, with our open-source projects is really important for us with Granted and with this project, that it’s going to continue to be reliable and continue to work as it currently does.

Corey: The way I see it, one of the dangers of doing anything that is particularly open-source—or that leans in the direction of building in Amazon’s ecosystem—it leads to the natural question of, well, isn’t this just going to be some people say stolen—and I don’t think those people understand how open-source works—by AWS themselves? Or aren’t they going to build something themselves at AWS that’s going to wind up stomping this thing that you’ve built? And my honest and remarkably cynical answer is that, “You have built a tool that is a joy to use, that makes logging into AWS accounts streamlined and efficient in a variety of different patterns. Does that really sound like something AWS would do?” And followed by, “I wish they would because everyone would benefit from that rising tide.”

I have to be very direct and very clear. Your product should not exist. This should be something the provider themselves handles. But nope. Instead, it has to exist. And while I’m glad it does, I also can’t shake the feeling that I am incredibly annoyed by the fact that it has to.

Chris: Yeah. Certainly, certainly. And it’s something that I think about a little bit. I like to wonder whether there’s maybe like a single feature flag or some single sort of configuration setting in AWS where they’re not allowing different tabs to access different accounts, they’re not allowing this kind of concurrent access. And maybe if we make enough noise about Granted, maybe one of the engineers will go and flick that switch and they’ll just enable it by default.

And then Granted itself will be a lot less relevant, but for everybody who’s using AWS, that’ll be a massive win because the big draw of using Granted is mainly just around being able to access different accounts at the same time. If AWS let you do that out of the box, hey, that would be great and, you know, I’d have a lot less stuff to maintain.

Corey: Originally, I had you here to talk about Granted, but I took a glance at what you’re actually building over at Common Fate and I’m about to basically hijack slash derail what probably is going to amount the rest of this conversation because you have a quick example on your site for by developers, for developers. You show a quick Python script that tries to access a S3 bucket object and it’s denied. You copy the error message, you paste it into what you’re building over a Common Fate, and in return, it’s like, “Oh. Yeah, this is the policy that fixes it. Do you want us to apply it for you?”

And I just about fell out of my chair because I have been asking for this explicit thing for a very long time. And AWS doesn’t do it. Their IAM access analyzer claims to. Like, “Oh, just go look at CloudTrail and see what permissions it uses and we’ll build a policy to scope it down.” “Okay. So, it’s S3 access. Fair enough. To what object or what bucket?” “Guess,” is what it tells you there.

And it’s, this is crap. Who thinks this is a good user experience? You have built the thing that I wish AWS had built in natively. Because let’s be honest here, I do what an awful lot of people do and overscope permissions massively just because messing around with the bare minimum set of permissions in many cases takes more time than building the damn thing in the first place.

Chris: Oh, absolutely. Absolutely. And in fact, this—was a few years ago when I was consulting—I had a really similar sort of story where one of the clients that we were working with, the CTO of this company, he was needing to grant us access to AWS and we were needing to build a particular service. And he said, “Okay, can you just let me know the permissions that you will need and I’ll go and deploy the role for this.” And I came back and I said, “Wait. I don’t even know the permissions that I’m going to need because the damn thing isn’t even built yet.”

So, we went sort of back and forth around this. And the compromise ended up just being you know, way too much access. And that was sort of part of the inspiration for, you know, really this whole project and what we’re building with Common Fate, just trying to make that feedback loop around getting to the right level of permissions a lot faster.

Corey: Yeah, I am just so overwhelmingly impressed by the fact that you have built—and please don’t take this as a criticism—but a set of very simple tools. Not simple in the terms of, “Oh, that’s, like, three lines of bash, and a fool could write that on a weekend.” No. Simple in the sense of it solves a problem elegantly and well and it’s straightforward—well, straightforward as anything in the world of access control goes—to wrap your head around exactly what it does. You don’t tend to build these things by sitting around a table brainstorming with someone you met at co-founder dating pool or something and wind up figuring out, “Oh, we should go and solve that. That sounds like a billion-dollar problem.”

This feels very much like the outcome of when you’re sitting around talking to someone and let’s start by drinking six beers so we become extraordinarily honest, followed immediately by let’s talk about what sucks. What pisses you off the most? It feels like this is sort of the low-hanging fruit of things that upset people when it comes to AWS. I mean, if things had gone slightly differently, instead of focusing on AWS bills, IAM was next on my list of things to tackle just because I was tired of smacking my head into it.

This is very clearly a problem space that you folks have analyzed deeply, worked within, and have put a lot of thought into. I want to be clear, I’ve thrown a lot of feature suggestions that you for Granted from start to finish. But all of them have been around interface stuff and usability and expanding use cases. None of them have been, “Well, that seems screamingly insecure.” Because it hasn’t been.

Chris: [laugh].

Corey: It has been effective, start to finish, I think that from a security posture, you make terrific choices, in many cases better than ones I would have made a starting from scratch myself. Everything that I’m looking at in what you have built is from a position of this is absolutely amazing and it is transformative to my own workflows. Now, how can we improve it?

Chris: Mmm. Thank you, Corey. And I’ll say as well, maybe around the security angle, that one of the goals with Granted was to try and do things a little bit better than the default way that AWS does them when it comes to security. And it’s actually been a bit of a source for challenges with some of the users that we’ve been working with with Granted because one of the things we wanted to do was encrypt the SSO token. And this is the token that when you sign in to AWS, kind of like, it allows you to then get access to all of the rest of the accounts.

So, it’s like a pretty—it’s a short-lived token, but it’s a really sensitive one. And you know, by default, it’s just stored in plain text on your disk. So, we dump to a file and, you know, anything that can go and read that, they can go and get it. It’s also a little bit hard to revoke and to lock people out. There’s not really great workflows around that on AWS’s side.

So, we thought, “Okay, great. One of the goals for Granted can be that we will go and store this in your keychain in your system and we’ll work natively with that.” And that’s actually been a cause for a little bit of a hassle for some users, though, because by doing that and by storing all of this information in the keychain, it’s actually broken some of the integrations with the rest of the tooling, which kind of expects tokens and things to be in certain places. So, we’ve actually had to, as part of dealing with that with Granted, we’ve had to give users the ability to opt out for that.

Corey: DoorDash had a problem. As their cloud-native environment scaled and developers delivered new features, their monitoring system kept breaking down. In an organization where data is used to make better decisions about technology and about the business, losing observability means the entire company loses their competitive edge. With Chronosphere, DoorDash is no longer losing visibility into their applications suite. The key? Chronosphere is an open-source compatible, scalable, and reliable observability solution that gives the observability lead at DoorDash business, confidence, and peace of mind. Read the full success story at snark.cloud/chronosphere. That's snark.cloud slash C-H-R-O-N-O-S-P-H-E-R-E.

Corey: That’s why I find this so, I think, just across the board, fantastic. It’s you are very clearly engaged with your community. There’s a community Slack that you have set up for this. And I know, I know, too many Slacks; everyone has this problem. This is one of those that is worth hanging in, at least from my perspective, just because one of the problems that you have, I suspect, is on my Mac it’s great because I wind up automatically updating it to whatever the most recent one is every time I do a brew upgrade.

But on the Linux side of the world, you’ve discovered what many of us have discovered, and that is that packaging things for Linux is a freaking disaster. The current installation is, “Great. Here’s basically a curl bash.” Or, “Here, grab this tarball and install it.” And that’s fine, but there’s no real way of keeping that updated and synced.

So, I was checking the other day, oh wow, I’m something like eight versions behind on this box. But it still just works. I upgraded. Oh, wow. There’s new functionality here. This is stuff that’s actually really handy. I like this quite a bit. Let’s see what else we can do.

I’m just so impressed, start to finish, by just how receptive you’ve been to various community feedbacks. And as well—I want to be very clear on this point, too—I’ve had folks who actually know what they’re doing in an InfoSec sense look at what you’re up to, and none of them had any issues of note. I’m sure that they have a pile of things like, with that curl bash, they should really be doing a GPG check. Yes, yes, fine. Whatever. If that’s your target threat model, okay, great. Here in reality-land for what I do, this is awesome.

And they don’t seem to have any problems with, “Oh, yeah. By the way, sending analytics back up”—which, okay, fine, whatever. “And it’s not disclosing them.” Okay, that’s bad. “And it’s including the contents of your AWS credentials.”

Ahhhh. I did encounter something that was doing that on the back-end once. [cough]—Serverless Framework—sorry, something caught in my throat for a second.

Chris: [laugh].

Corey: No faster way I can think of to erode trust in that. But everything you’re doing just makes sense.

Chris: Oh, I do remember that. And that was a little bit of a fiasco, really, around all of that, right? And it’s great to hear actually around that InfoSec folks and security people being, you know, not unhappy, I guess, with a tool like this. It’s been interesting for me personally. We’ve really come from a practitioner’s background.

You know, I wouldn’t call myself a security engineer at all. I would call myself as a sometimes a software developer, I guess. I have been hacking my way around Go and definitely learning a lot about how the cloud has worked over the past seven, eight years or so, but I wouldn’t call myself a security engineer, so being very cautious around how all of these things work. And we’ve really tried to defer to things like the system keychain and defer to things that we know are pretty safe and work.

Corey: The thing that I also want to call out as well is that your licensing is under the MIT license. This is not one of those, “Oh, you’re required to wind up doing a bunch of branding stuff around it.” And, like some people say, “Oh, you have to own the trademark for all of these things.” I mean, I’m not an expert in international trademark law, let’s be very clear, but I also feel that trademarking a term that is already used heavily in the space such as the word ‘Granted,’ feels like kind of an uphill battle. And let’s further be clear that it doesn’t matter what you call this thing.

In fact, I will call attention to an oddity that I’ve encountered a fair bit. After installing it, the first thing you do is you run the command ‘granted.’ That sets it up, it lets you configure your browser, what browser you want to use, and it now supports standard out for that headless, EC2 use case. Great. Awesome. Love it. But then the other binary that ships with it is Assume. And that’s what I use day-to-day. It actually takes me a minute sometimes when it’s been long enough to remember that the tool is called Granted and not Assume what’s up with that?

Chris: So, part of the challenge that we ran into when we were building the Granted project is that we needed to export some environment variables. And these are really important when you’re logging into AWS because you have your access key, your secret key, your session token. All of those, when you run the assume command, need to go into the terminal session that you called it. This doesn’t matter so much when you’re using the console mode, which is what we mentioned earlier where you can open 100 different accounts if you want to view all of those at the same time in your browser. But if you want to use it in your terminal, we wanted to make it look as really smooth and seamless as possible here.

And we were really inspired by this approach from—and I have to shout them out and kind of give credit to them—a tool called AWSume—they’re spelled A-W-S-U-M-E—Python-based tool that they don’t do as much with single-sign-on, but we thought they had a really nice, like, general approach to the way that they did the scripting and aliasing. And we were inspired by that and part of that means that we needed to have a shell script that called this executable, which then will export things back out into the shell script. And we’re doing all this wizardry under the hood to make the user experience really smooth and seamless. Part of that meant that we separated the commands into granted and assume and the other part of the naming for everything is that I felt Granted had a far better ring to it than calling the whole project Assume.

Corey: True. And when you say assume, is it AWS or not? I’ve used the AWSume project before; I’ve used AWS Vault out of 99 Designs for a while. I’ve used—for three minutes—the native AWS SSO config, and that is just trash. Again, they’re so good at the plumbing, so bad at the porcelain, I think is the criticism that I would levy toward a lot of this stuff.

Chris: Mmm.

Corey: And it’s odd to think there’s an entire company built around just smoothing over these sharp, obnoxious edges, but I’m saying this as someone who runs a consultancy and have five years that just fixes the bill for this one company. So, there’s definitely a series of cottage industries that spring up around these things. I would be thrilled, on some level, if you wound up being completely subsumed by their product advancements, but it’s been 15 years for a lot of this stuff and we’re still waiting. My big failure mode that I’m worried about is that you never are.

Chris: Yeah, exactly, exactly. And it’s really interesting when you think about all of these user experience gaps in AWS being opportunities for, I guess, for companies like us, I think, trying to simplify a lot of the complexity for things. I’m interested in sort of waiting for a startup to try and, like, rebuild the actual AWS console itself to make it a little bit faster and easier to use.

Corey: It’s been done and attempted a bunch of different times. The problem is that the console is a lot of different things to a lot of different people, and as you step through that, you can solve for your use case super easily. “Yeah, what do I care? I use RDS, I use some VPC nonsense, and I use EC2. The end.” “Great. What about IAM?”

Because I promise you’re using that whether you know it or not. And okay, well, I’m talking to someone else who’s DynamoDB, and someone else is full-on serverless, and someone else has more money than sense, so they mostly use SageMaker, and so on and so forth. And it turns out that you’re effectively trying to rebuild everything. I don’t know if that necessarily works.

Chris: Yeah, and I think that’s a good point around maybe while we haven’t seen anything around that sort of space so far. You go to the console, and you click down, you see that list of 200 different services and all of those have had teams go and actually, like, build the UI and work with those individual APIs. Yeah.

Corey: Any ideas as far as what’s next for features on Granted?

Chris: I think that, for us, it’s continuing to work with everybody who’s using it, and with a focus of stability and performance. We actually had somebody in the community raise an issue because they have an AWS config file that’s over 7000 lines long. And I kind of pity that person, potentially, for their day-to-day. They must deal with so much complexity. Granted is currently quite slow when the config files get very big. And for us, I think, you know, we built it for ourselves; we don’t have that many accounts just yet, so working to try to, like, make it really performant and really reliable is something that’s really important.

Corey: If you don’t mind a feature request while we’re at it—and I understand that this is more challenging than it looks like—I’m willing to fund this as a feature bounty that makes sense. And this also feels like it might be a good first project for a very particular type of person, I would love to get tab completion working in Zsh. You have it—

Chris: Oh.

Corey: For Fish because there’s a great library that automatically populates that out, but for the Zsh side of it, it’s, “Oh, I should just wind up getting Zsh completion working,” and I fell down a rabbit hole, let me tell you. And I come away from this with the perception of yeah, I’m not going to do it. I have not smart enough to check those boxes. But a lot of people are so that is the next thing I would love to see. Because I will change my browser to log into the AWS console for you, but be damned if I’m changing my shell.

Chris: [laugh]. I think autocomplete probably should be higher on our roadmap for the tool, to be honest because it’s really, like, a key metric and what we’re focusing on is how easy is it to log in. And you know, if you’re not too sure what commands to use or if we can save you a few keystrokes, I think that would be the, kind of like, reaching our goals.

Corey: From where I’m sitting, you definitely have. I really want to thank you for taking the time to not only build this in the first place, but also speak with me about it. If people want to learn more, where’s the best place to find you?

Chris: So, you can find me on Twitter, I’m @chr_norm, or you can go and visit granted.dev and you’ll have a link to join the Slack community. And I’m very active on the Slack.

Corey: You certainly are, although I will admit that I fall into the challenge of being in just the perfectly opposed timezone from you and your co-founder, who are in different time zones to my understanding; one of you is on Australia and one of you was in London; you’re the London guy as best I’m aware. And as a result, invariably, I wind up putting in feature requests right when no one’s around. And, for better or worse, in the middle of the night is not when I’m usually awake trying to log into AWS. That is Azure time.

Chris: [laugh]. Yeah, no, we don’t have the US time zone properly covered yet for our community support and help. But we do have a fair bit of the world timezone covered. The rest of the team for Common Fate is all based in Australia and I’m out here over in London.

Corey: Yeah. I just want to thank you again, for just being so accessible and, like, honestly receptive to feedback. I want to be clear, there’s a way to give feedback and I do strive to do it constructively. I didn’t come crashing into your Slack one day with a, “You know what your problem is?” I prefer to take the, “This is awesome. Here’s what I think would be even better. Does that make sense?” As opposed to the imperious demands and GitHub issues and whatnot? It’s, “I’d love it if it did this thing. Doesn’t do this thing. Can you please make it do this thing?” Turns out that’s the better way to drive change. Who knew?

Chris: Yeah. [laugh]. Yeah, definitely. And I think that one of the things that’s been the best around our journey with Granted so far has been listening to feedback and hearing from people how they would like to use the tool. And a big thank you to you, Corey, for actually suggesting changes that make it not only better for you, but better for everybody else who’s using Granted.

Corey: Well, at least as long as we’re using my particular byzantine workload patterns in some way, or shape, or form, I’ll hear that. But no, it’s been an absolute pleasure and I really want to thank you for your time as well.

Chris: Yeah, thank you for having me.

Corey: Chris Norman, co-founder of Common Fate, as well as one of the two primary developers originally behind the Granted project that logs you into AWS without you having to lose your mind. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry, incensed, raging comment that talks about just how terrible all of this is once you spend four hours logging into your AWS account by hand first.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

Full Description / Show Notes

  • Corey and Linda talk about Tiktok and the online developer community (1:18)
  • Linda talks about what prompted her to want to work at AWS (5:29)
  • Linda discusses navigating the change from just being part of the developer community to being an employee of AWS (10:37)
  • Linda talks about moving AWS more in the direction of short form content, and Corey and Linda talk about the Tiktok algorithm (15:56)
  • Linda talks about the potential struggle of going from short form to long form content (25:21)

About Linda

Linda Vivah is a Site Reliability Engineer for a major media organization in NYC, a tech content creator, an AWS community builder member, a part-time wedding singer, and the founder of a STEM jewelry shop called Coding Crystals. At the time of this recording she was about to join AWS in her current position as a Developer Advocate.

Linda had an untraditional journey into tech. She was a Philosophy major in college and began her career in journalism. In 2015, she quit her tv job to attend The Flatiron School, a full stack web development immersive program in NYC. She worked as a full-stack developer building web applications for 5 years before shifting into SRE to work on the cloud end internally.

Throughout the years, she’s created tech content on platforms like TikTok & Instagram and believes that sometimes the best way to learn is to teach.

Links Referenced:

  • lindavivah.com: https://lindavivah.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Honeycomb. When production is running slow, it’s hard to know where problems originate. Is it your application code, users, or the underlying systems? I’ve got five bucks on DNS, personally. Why scroll through endless dashboards while dealing with alert floods, going from tool to tool to tool that you employ, guessing at which puzzle pieces matter? Context switching and tool sprawl are slowly killing both your team and your business. You should care more about one of those than the other; which one is up to you. Drop the separate pillars and enter a world of getting one unified understanding of the one thing driving your business: production. With Honeycomb, you guess less and know more. Try it for free at honeycomb.io/screaminginthecloud. Observability: it’s more than just hipster monitoring.

Corey: Let’s face it, on-call firefighting at 2am is stressful! So there’s good news and there’s bad news. The bad news is that you probably can’t prevent incidents from happening, but the good news is that incident.io makes incidents less stressful and a lot more valuable. incident.io is a Slack-native incident management platform that allows you to automate incident processes, focus on fixing the issues and learn from incident insights to improve site reliability and fix your vulnerabilities. Try incident.io, recover faster and sleep more.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. We talk a lot about how people go about getting into this ridiculous industry of ours, and I’ve talked a little bit about how I go about finding interesting and varied guests to show up and help me indulge my ongoing love affair on this show with the sound of my own voice. Today, we’re going to be able to address both of those because today I’m speaking to Linda Haviv, who, as of this recording, has accepted a job as a Developer Advocate at AWS, but has not started. Linda, welcome to the show.

Linda: Thank you so much for having me, Corey. Happy to be here.

Corey: So, you and I have been talking for a while and there’s been a lot of interesting things I learned along the way. You were one of the first people I encountered when I joined the TikToks, as all the kids do these days, and was trying to figure out is there a community of folks who use AWS. Which really boils down to, “So, where are these people that are sad all the time?” Well, it turns out, they’re on TikTok, so there we go. We found my people.

And that was great. And we started talking, and it turns out that we were both in the AWS community builder program. And we’ve developed a bit of a rapport. We talk about different things. And then, I guess, weird stuff started happening, in the context of you were—you’re doing very well at building an audience for yourself on TikTok.

I tried it, and it was—my sense of humor sometimes works, sometimes doesn’t. I’ve had challenges in finding any reasonable way to monetize it because a 30-second video doesn’t really give nuance for a full ad read, for example. And you’ve been looking at it from the perspective of a content creator looking to build the audience slash platform is step one, and then, eh, step two, you’ll sort of figure out aspects of monetization later. Which, honestly, is a way easier way to do it in hindsight, but, yeah, the things that we learn. Now, that you’re going to AWS, first, you planning to still be on the TikToks and whatnot?

Linda: Absolutely. So, I really look at TikTok as a funnel. I don’t think it’s the main place, you’re going to get that deep-dive content but I think it’s a great way, especially for things that excite you or get you into understanding it, especially beginner-type audience, I think there’s a lot of untapped market of people looking to into tech, or technologists that aren’t in the cloud. I mean, even when I worked—I worked as a web developer and then kind of learned more about the cloud, and I started out as a front-end developer and shifted into, like, SRE and infrastructure, so even for people within tech, you can have a huge tech community which there is on TikTok, with a younger community—but not all of them really understand the cloud necessarily, depending on their job function. So, I think it’s a great way to kind of expose people to that.

For me, my exposure came from community. I met somebody at a meetup who was working in cloud, and it wasn’t even on the job that I really started getting into cloud because many times in corporations, you might be working on a specific team and you’re not really encountering other ends, and it seems kind of like a mystery. Although it shouldn’t seem like magic, many times when you’re doing certain job functions—especially the DevOps—could end up feeling like magic. So, [laugh] for the good and the bad. So sometimes, if you’re not working on that end, you really sometimes take it for granted.

And so, for me, I actually—meetups were the way I got exposed to that end. And then I brought it back into my work and shifted internally and did certifications and started, even, lunch-and-learns where I work to get more people in their learning journey together within the company, and you know, help us as we’re migrating to the cloud, as we’re building on the cloud. Which, of course, we have many more roles down the road. I did it for a few years and saw the shift. But I worked at a media company for many years and now shifting to AWS, and so I’ve seen that happen on different ends.

Not—oh, I wasn’t the one doing the migration because I was on the other end of that time, but now for the last two years, I was working on [laugh] the infrastructure end, and so it’s really fascinating. And many people actually—until now I feel like—that will work on maybe the web and mobile and don’t always know as much about the cloud. I think it’s a great way to funnel things in a quick manner. I think also society is getting used to short videos, and our attention span is very low, and I think for—

Corey: No argument here.

Linda: —[crosstalk 00:04:39] spending so mu—yeah, and we’re spending so much time on these platforms, we might as well, you know, learn something. And I think it depends what content. Some things work well, some things doesn’t. As with anything content creation, you kind of have to do trial and error, but I do find the audience to be a bit different on TikTok versus Twitter versus Instagram versus YouTube. Which is interesting how it’s going to play out on YouTube, too, which is a whole ‘nother topic conversation.

Corey: Well, it’s odd to me watching your path. It’s almost the exact opposite of mine where I started off on the back-end, grumpy sysadmin world and, “Oh, why would I ever need to learn JavaScript?” “Well, genius, because as the world progresses, guess what? That’s right. The entire world becomes JavaScript. Welcome.”

And it took me a long time to come around to that. You started with the front-end world and then basically approached from the exact opposite end. Let’s be clear, back in my day, mine was the common path. These days, yours is very much the common path.

Linda: Yeah.

Corey: I also want to highlight that all of those transitions and careers that you spoke about, you were at the same company for nine years, which in tech is closer to 30. So, I have to ask, what was it that inspired you, after nine years, to decide, “I’m going to go work somewhere else. But not just anywhere; I’m going to AWS.” Because normally people don’t almost institutionalized lifers past a certain point.

Linda: [laugh].

Corey: Like, “Oh, you’ll be there till you retire or die.” Whereas seeing significant career change after that long in one place, even if you’ve moved around internally and experienced a lot of different roles, is not common at all what sparked that?

Linda: Yeah. Yeah, no, it’s such a good question. I always think about that, too, especially as I was reflecting because I’m, you know, in the midst of this transition, and I’ve gotten a lot of reflecting over the last two weeks [laugh], or more. But I think the main thing for me is, I always, wherever I was—and this kind of something that—I’m very proactive when it comes to trying to transition. I think, even when I was—right, I held many roles in the same company; I used to work in TV production and actually left for three months to go to a coding boot camp and then came back on the other end, but I understood the product in a different way.

So, for that time period, it was really interesting to work on the other end. But, you know, as I kind of—every time I wanted to progress further, I always made a move that was actually new and put me in an uncomfortable place, even within the same company. And I’m at the point now that I’m in my career, I felt like this next step really needs to be, you know, at AWS. It's not, like, the natural progression for me. I worked alongside—on the client end—with AWS and have seen so many projects come through and how much our own workloads have changed.

And it’s just been an incredible journey, also dealing with accounts team. On that end, I’ve worked alongside them, so for me, it was kind of a natural progression. I was very passionate about cloud computing at AWS and I kind of wanted to take it to that next place, and I felt like—also, dealing with the community as part of my job is a dream part to me because I was always doing that on the side on social media. So, it wasn’t part of my day-to-day job. I was working as an SRE and an infrastructure engineer, so I didn’t get to do that as part of my day-to-day.

I was making videos at 2 a.m. and, you know, kind of trying to, like, do—you know, interact with the community like that. And I think—I come from a performing background, the people background, I was singing since I was four years old. I always go to—I was a wedding singer, so I go into a room and I love making people happy or giving value. And I think, like, education has a huge part of that. And in a way, like making that content and—

Corey: You got to get people’s attention—

Linda: Yeah.

Corey: —you can’t teach them a damn thing.

Linda: Right. Exactly. So, it’s kind of a mix of everything. It’s like that performance, the love of learning. You know, between you and I, like, I wanted to be a lawyer before I thought I was going to—before I went to tech.

I thought I was going to be a lawyer purely because I loved the concept of going to law school. I never took time to think about the law part, like, being the lawyer part. I always thought, “Oh, school.” I’m a student at heart. I always call myself a professional student. I really think that’s part of what you need to be in this world, in this tech industry, and I think for me, that’s what keeps my fire going.

I love to experiment, to learn, to build. And there’s something very fulfilling about building products. If you take a step back, like, you’re kind of—you know, for me that part, every time I look back at that, that always is what kind of keeps me going. When I was doing front-end, it felt a lot more like I was doing smaller things than when I was doing infrastructure, so I felt like that was another reason why I shifted. I love doing the front-end, but I felt like I was spending two days on an Internet Explorer bug and it just drove me—[laugh] it just made it feel unfulfilling versus spending two days on, you know, trying to understand why, you know, something doesn’t run the infrastructure or, like, there’s—you know, it’s failing blindly, you know? Stuff like that. Like, I don’t know, for me that felt more fulfilling because the problem was more macro. But I think I needed both. I have a love for both, but I definitely prefer being back-end. So. [laugh]. Well, I’m saying that now but—[laugh].

Corey: This might be a weakness on my part where I’m basically projecting onto others, and this is—I might be completely wrong on this, but I tend to take a bit of a bifurcated view of community. I mean, community is part of the reason that I know the things I know and how I got to this place that I am, so use that as a cautionary tale if you want. But when I talk to someone like you at this moment, where you’re in the community, I’m in the community, and I’m talking to you about a problem I’m having and we’re working on ways to potentially solve that or how to think about that. I view us as basically commiserating on these things, whereas as soon as you start on day one—and yes, it’s always day one—at AWS and this becomes your day job and you work there, on some level, for me, there’s a bit shift that happens and a switch gets flipped in my head where, oh, you actually work at this company. That means you’re the problem.

And I’m not saying that in a way of being antagonistic. Please, if you’re watching or listening to this, do not antagonize the developer advocates. They have a very hard job understanding all this so they can explain that to the rest of us. But how do you wind up planning to navigate, or I guess your views on, I guess, handling the shift between, “One of the customers like the rest of us,” to, as I say, “Part of the problem,” for lack of a better term.

Linda: Or, like, work because you kind of get the—you know. I love this question and it’s something I’ve been pondering a lot on because I think the messaging will need to be a little different [coming from me 00:10:44] in the sense of, there needs to be—just in anything, you have to kind of create trust. And to create trust, you have to be vulnerable and authentic. And I think I, for example, utilize a lot of things outside of just the AWS cloud topic to do that now, even, when I—you know, kind of building it without saying where I work or anything like that, going into this role and it being my job, it’s going to be different kind of challenge as far as the messaging, but I think it still holds true that part, that just developing trust and authenticity, I might have to do more of that, you know? I might have to really share more of that part, share other things to really—because it’s more like people come, it doesn’t matter how much somet—how many times you explain it, many times, they will see your title and they will judge you for it, and they don’t know what happened before. Every TikTok, for example, you have to act like it’s a new person watching. There is no series, you know? Like, yes, there’s a series but, like, sometimes you can make that but it’s not really the way TikTok functions or a short-form video functions. So, you kind of have to think this is my first time—

Corey: It works really terribly when you’re trying to break it out that way on TikTok.

Linda: [laugh]. Yeah.

Corey: Right. Here’s part 17 of my 80-TikTok-video saga. And it’s, “Could you just turn this into a blog post or put this on YouTube or something? I don’t have four hours to spend learning how all this stuff works in your world.”

Linda: Yeah. And you know, I think repeating certain things, too, is really important. So, they say you have to repeat something eight times for people to see it or [laugh] something like that. I learned that in media [crosstalk 00:12:13]—

Corey: In a row, or—yeah. [laugh].

Linda: I mean, the truth is that when you, kind of like, do a TikTok maybe, like, there’s something you could also say or clarify because I think there’s going to be—and I’m going to have to—there’s going to be a lot of trial and error for me; I don’t know if I have answers—but my plan is going into it very much testing that kind of introduction, or, like, clarifying what that role is. Because the truth is, the role is advocating on behalf of the community and really helping that community, so making sure that—you don’t have to say it as far as a definition maybe, but, like, making sure that comes across when you create a video. And I think that’s going to be really important for me, and more important than the prior even creating content going forward. So, I think that’s one thing that I definitely feel like is key.

As well as creating more raw interaction. So, it depends on the platform, too. Instagram, for example, is much more community—how do I put this? Instagram is much more easy to navigate as far as reaching the same community because you have something, like, called Instagram Stories, right? So, on Instagram Stories, you’re bringing those stories, mostly the same people that follow you. You’re able to build that trust through those stories.

On TikTok, they just released Stories. I haven’t really tried them much and I don’t play with it a lot, but I think that’s something I will utilize because those are the people that are already follow you, meaning they have seen a piece of content. So, I think addressing it differently and knowing who’s watching what and trying to kind of put yourself in their shoes when you’re trying to, you know, teach something, it’s important for you to have that trust with them. And I think—key to everything—being raw and authentic. I think people see through that. I would hope they do.

And I think, uh, [laugh] that’s what I’m going to be trying to do. I’m just going to be really myself and real, and try to help people and I hope that comes through because that’s—I’m passionate about getting more people into the cloud and getting them educated. And I feel like it’s something that could also allow you to build anything, just from anywhere on your computer, brings people together, the world is getting smaller, really. And just being able to meet people through that and there’s just a way to also change your life. And people really could change their life.

I changed my life, I think, going into tech and I’m in the United States and I, you know—I’m in New York, you know, but I feel like so many people in the States and outside of the States, you know, all over the world, you know, have access to this, and it’s powerful to be able to build something and contribute and be a part of the future of technology, which AWS is.

Corey: I feel like, in three years or whatever it is that you leave AWS in the far future, we’re going to basically pull this video up and MST3k came together. It’s like, “Remember how naive you were talking about these things?” And I’m mostly kidding, but let’s be serious. You are presumably going to be focusing on the idea of short-form content. That is—

Linda: Yeah.

Corey: What your bread-and-butter of audience-building has been around, and that is something that is new for AWS.

Linda: Yeah.

Corey: And I’m always curious as to how companies and their cultures continue to evolve. I can only imagine there’s a lot of support structure in place for that. I personally remember giving a talk at an AWS event and I had my slides reviewed by their legal team, as they always do, and I had a slide that they were looking at very closely where I was listing out the top five AWS services that are bullshit. And they don’t really have a framework for that, so instead, they did their typical thing of, “Okay, we need to make sure that each of those services starts with the appropriate AWS or Amazon naming convention and are they capitalized properly?” Because they have a framework for working on those things.

I’m really curious as to how the AWS culture and way of bringing messaging to where people are is going to be forced to evolve now that they, like it or not, are going to be having significantly increased presence on TikTok and other short-form platforms.

Linda: I mean, it’s really going to be interesting to see how this plays out. There’s so much content that’s put out, but sometimes it’s just not reaching the right audience, so making sure that funnel exists to the right people is important and reaching those audiences. So, I think even YouTube Shorts, for example. Many people in tech use YouTube to search a question.

They do not care about the intro, sometimes. It depends what kind of following, it depends if [in gaming 00:16:30], but if you’re coming and you’re building something, it’s like a Stack Overflow sometimes. You want to know the answer to your question. Now, YouTube Shorts is a great solution to that because many times people want the shortest possible answer. Now, of course, if it’s a tutorial on how to build something, and it warrants ten minutes, that’s great.

Even ten minutes is considered, now, Shorts because TikTok now has ten-minute videos, but I think TikTok is now searchable in the way YouTube is, and I think let’s say YouTube Shorts is short-form, but very different type of short-form than TikTok is. TikTok, hooks matter. YouTube answers to your questions, especially in chat. I wouldn’t say everything in YouTube is like that; depends on the niche. But I think even within short-form, there’s going to be a different strategy regarding that.

So, kind of like having that mix. I guess, depending on platform and audience, that’s there. Again, trial and error, but we’ll see how this plays out and how this will evolve.

Corey: This episode is sponsored in part by our friends at Vultr. Optimized cloud compute plans have landed at Vultr to deliver lightning-fast processing power, courtesy of third-gen AMD EPYC processors without the IO or hardware limitations of a traditional multi-tenant cloud server. Starting at just 28 bucks a month, users can deploy general-purpose, CPU, memory, or storage optimized cloud instances in more than 20 locations across five continents. Without looking, I know that once again, Antarctica has gotten the short end of the stick. Launch your Vultr optimized compute instance in 60 seconds or less on your choice of included operating systems, or bring your own. It’s time to ditch convoluted and unpredictable giant tech company billing practices and say goodbye to noisy neighbors and egregious egress forever. Vultr delivers the power of the cloud with none of the bloat. Screaming in the Cloud listeners can try Vultr for free today with a $150 in credit when they visit getvultr.com/screaming. That’s G-E-T-V-U-L-T-R dot com slash screaming. My thanks to them for sponsoring this ridiculous podcast.

Corey: I feel like there are two possible outcomes here. One is that AWS—

Linda: Yeah.

Corey: Nails this pivot into short-form content, and the other is that all your TikTok videos start becoming ten minutes long, which they now support, welcome to my TED Talk. It’s awful, and then you wind up basically being video equivalent for all of your content, of recipes when you search them on the internet where first they circle the point to death 18 times with, “Back when I was a small child growing up in the hinterlands, we wound—my grandmother would always make the following stew after she killed the bison with here bare hands. Why did grandma kill a bison? We don’t know.” And it just leads down this path so they can get, like, long enough content or they can have longer and longer articles to display more ads.

And then finally at the end, it’s like ingredient one: butter. Ingredient two, there is no ingredient two. Okay. That explains why it’s delicious. Awesome. But I don’t like having people prolong it. It’s just, give me the answer I’m looking for.

Linda: Yeah.

Corey: Get to the point. Tell me the story. And—

Linda: And this is—

Corey: —I’m really hoping that is not the direction your content goes in. Which I don’t think it would, but that is the horrifying thing and if for some chance I’m right, I will look like Nostradamus when we do that MST3k episode.

Linda: No, no. I mean, I really am—I always personally—even when I was creating content these last few years and testing different things, I’m really a fan of the shortest way possible because I don’t have the patience to watch long videos. And maybe it’s because I’m a New Yorker that can’t sit down from the life of me—apart from when I code of course—but, you know, I don’t like wasting time, I’m always on the go, I’m with my coffee, I’m like—that’s the kind of style I prefer to bring in videos in the sense of, like, people have no time. [laugh]. You know?

The amount of content we’re consuming is just, uh, bonkers. So, I don’t think our mind is really a built for consuming [laugh] this much content every time you open your phone, or every time you look, you know, online. It’s definitely something that is challenging in a whole different way. But I think where my content—if it’s ten minutes, it better be because I can’t shorten it. That’s my thing. So, you can hold me accountable to that because—

Corey: Yeah, I want ten minutes of—

Linda: I’m not a—

Corey: Content, not three minutes of content in a ten-minute bag.

Linda: Exactly. Exactly. So, if it’s a ten-minute video, it would have been in one hour that I cut down, like, meaning a tutorial, a very much technical types of content. I think things that are that long, especially in tech, would be something like, on that end—unless, of course, you know, I’m not talking about, like, longer videos on YouTube which are panels or that kind of thing. I’m talking more like if I’m doing something on TikTok specifically.

TikTok also cares about your watch time, so if people aren’t interested in it, it’s not going to do well, it doesn’t matter how many followers you have. Which is what I do like about the way TikTok functions as opposed to, let’s say, Instagram. Instagram is more like it gives it to your following—and this is the current state, I don’t know if it always evolves—but the current state is, Instagram Reels kind of functions in a way where it goes first to the people that follow you, but, like, in a way that’s more amplified than TikTok. TikTox tests people that follows you, but if it’s not a good video, it won’t do well. And honestly, they’re many good videos videos that don’t go viral. I’m not talking about that.

Sometimes it’s also the topic and the niche and the sound and the title. I mean, there’s so many people who take a topic and do it in three different ways and one of them goes viral. I mean, there’s so many factors that play into it and it’s hard to really, like, always, you know, kind of reverse engineer but I do think that with TikTok, things won’t do well, more likely if it’s not a good piece of content as opposed to—or, like, too long, right? Not—I shouldn’t say not good a good piece of content—it’s too long.

Corey: The TikTok algorithm is inscrutable to me. TikTok is firmly convinced, based upon what it shows me, that I am apparently a lesbian. Which okay, fine. Awesome. Whatever. I’m also—it keeps showing me ads for ADHD stuff, and it was like, “Wow, like, how did it know that?” Followed by, “Oh, right. I’m on TikTok. Nevermind.”

And I will say at one point, it recommended someone to me who, looking at the profile picture, she’s my nanny. And it’s, I have a strong policy of not, you know, stalking my household employees on social media. We are not Facebook friends, we are not—in a bunch of different areas. Like, how on earth would they have figured this out? I’m filling the corkboard with conspiracy and twine followed by, “Wait a minute. We probably both connect from the same WiFi network, which looks like the same IP address and it probably doesn’t require a giant data science team to put two and two together on those things.” So, it was great. I was all set to do the tinfoil hat conspiracy, but no, no, that’s just very basic correlation 101.

Linda: And also, this is why I don’t enable contacts on TikTok. You know, how it says, “Oh, connect your contacts?”

Corey: Oh, I never do that. Like, “Can we look at your contacts?”

Linda: Never.

Corey: “No.” “Can we look at all of your photos?” “Absolutely not.” “Can we track you across apps?” “Why would anyone say yes to this? You’re going to do it anyway, but I’ll say no.” Yeah.

Linda: Got to give the least privilege. [laugh]. Definitely not—

Corey: Oh absolutely.

Linda: Yeah. I think they also help [crosstalk 00:22:40]—

Corey: But when I’m looking at—the monetization problem is always a challenge on things like this, too, because when I’m—my guilty TikTok scrolling pleasures hit, it’s basically late at night, I just want to see—I want something to want to wind down and decompress. And I’m not about ready to watch, “Hey, would you like to migrate your enterprise database to this other thing?” It’s, I… no. There’s a reason that the ads that seem to be everywhere and doing well are aimed at the mass market, they’re generally impulse buys, like, “Hey, do you want to set that thing over there on fire, but you’re not close enough to get the job done? But this flame thrower today. Done.”

And great, like, that is something everyone can enjoy, but these nuanced database products and anything else is B2B SaaS style stuff, it feels like it’s a very tough sell and no one has quite cracked that nut, yet.

Linda: Yeah, and I think the key there—this is, I’m guessing based on, like, what I want to try out a lot—is the hook and the way you’re presenting it has to be very product-focused in the sense that it needs to be very relatable. Even if you don’t know anything about tech, you need to be—like, for example, in the architecture page on AWS, there’s a video about the Emirates going to Mars mission. Space is a very interesting topic, right? I think, a hook, like, “Do want to see how, like, how this is bu—” like, it’s all, like, freely available to see exactly [laugh] how this was built. Like, it might—in the right wording, of course—it might be interesting to someone who’s looking for fun-fact-style content.

Now, is it really addressing the people that are building everyday? Not really always, depends who’s on there and the mass market there. But I feel like going on the product and the things that are mass-market, and then working backwards to the tech part of it, even if they learn something and then want to learn more, that’s really where I see TikTok. I don’t think every platform would be, maybe, like this, but that’s where I see getting people: kind of inviting them in to learn more, but making it cool and fun. It’s very important, but it feels cool and fun. [laugh]. So.

Because you’re right, you’re scrolling at 2 a.m. who wants to start seeing that. Like, it’s all about how you teach. The content is there, the content has—you know, that’s my thing. It’s like, the content is there. You don’t need to—it’s yes, there’s the part where things are always evolving and you need to keep track of that; that’s whole ‘nother type thing which you do very well, right?

And then there’s a part where, like, the content that already exists, which part is evergreen? Meaning, which part is, like, something that could be re—also is not timely as far as update, for example, well-architected framework. Yes, it evolves all the time, you always have new pillars, but the guide, the story, that is an evergreen in some sense because that guide doesn’t, you know, that whole concept isn’t going anywhere. So, you know, why should someone care about that?

Corey: Right. How to turn on two-factor authentication for your AWS account.

Linda: Right.

Corey: That’s evergreen. That’s the sort of thing that—and this is the problem, I think, AWS has had for a long time where they’re talking about new features, new enhancements, new releases. But you look what people are actually doing and so much of it is just the same stuff again and again because yeah, that is how most of the cloud works. It turns out that three-quarters of company’s production infrastructures tends to run on EC2 more frequently than it tends to run on IoT Greengrass. Imagine that.

So, there’s this idea of continuing to focus on these things. Now, one of my predictions is that you’re going to have a lot of fun with this and on some level, it’s going to really work for you. In others, it’s going to be hilariously—well, its shortcomings might be predictable. I can just picture now you’re at re:Invent; you have a breakout talk and terrific. And you’ve successfully gotten your talk down to one minute and then you’re sitting there with—

Linda: [laugh].

Corey: —the remainder of maybe 59. Like, oh, right. Yeah. Turns out not everything is short-form. Are you predicting any—

Linda: Yep.

Corey: Problems going from short-form to long-form in those instances?

Linda: I think it needs to go hand-in-hand, to be honest. I think when you’re creating any short-form content, you have—you know, maybe something short is actually sometimes in some ways, right, harder because you really have to make sure, especially in a technical standpoint, leaving things out is sometimes—leaves, like, a blind spot. And so, making sure you’re kind of—whatever you’re educating, you kind of, to be clear, “Here’s where you learn more. Here’s how I’m going to answer this next question for you: go here.” Now, in a longer-form content, you would cover all that.

So, there’s always that longevity. I think even when I write a script, and there’s many scripts I’m still [laugh] I’ve had many ideas until now I’ve been doing this still at 2 a.m. so of course, there’s many that didn’t, you know, get released, but those are the things that are more time consuming to create because you’re taking something that’s an hour-long, and trying to make sure you’re pulling out the things that are most—that are hook-style, that invite people in, that are accurate, okay, that really give you—explain to you clearly where are the blind spots that I’m not explaining on this video are. So, “XYZ here is, like, the high level, but by the way, there’s, like, this and this.” And in a long-form, you kind of have to know the long-form version of it to make the short-form, in some ways, depending on what—you’re doing because you’re funneling them to somewhere. That’s my thing. Because I don’t think there should be [crosstalk 00:27:36]—

Corey: This is the curse of Twitter, on some level. It’s, “Well, you forgot about this corner case.” “Yeah, I had 280 characters to get into.” Like, the whole point of short-form content—which I do consider Twitter to be—is a glimpse and a hook, and get people interested enough to go somewhere and learn more.

For something like AWS, this makes a lot of sense. When you highlight a capability or something interesting, it’s something relevant, whereas on the other side of it, where it’s this, “Oh, great. Now, here’s an 8000-word blog post on how I did this thing.” Yeah, I’m going to get relatively fewer amounts of traffic through that giant thing, but the people who are they’re going to be frickin’ invested because that’s going to be a slog.

Linda: Exactly.

Corey: “And now my eight-hour video on how exactly I built this thing with TypeScript.” Badly—

Linda: Exactly.

Corey: —as it turns out because I’m a bad programmer.

Linda: [laugh]. No, you’re not. I love your shit-posting. It’s great.

Corey: Challenge accepted.

Linda: [laugh]. I love what you just mentioned because I think you’re hitting the nail on the head when it comes to the quality content that’s niche focus, like, there needs to be a good healthy mix. I think always doing that, like, mass-market type video, it doesn’t give you, also, the credibility you need. So, doing those more niche things that might not be relevant to everybody, but here and there, are part of that is really key for your own knowledge and for, like, the com—you know, as far as, like, helping someone specific. Because it’s almost like—right, when you’re selling a service and you’re using social media, right, not everybody’s going to buy your service. It doesn’t matter what business you’re in right? The deep-divers are going to be the people that pay up. It’s just a numbers game, right? The more people you, kind of, address from there, you’ll find—

Corey: It’s called a funnel for a reason.

Linda: Right. Exactly.

Corey: Free content, paid content. Almost anyone will follow me on Twitter; fewer than will sign up for a newsletter; fewer will listen to a podcast; fewer will watch a video, and almost none of them will buy a consulting engagement. But ‘almost’ and ‘actually none of them,’ it turns out is a very different world.

Linda: Exactly. [laugh]. So FYI, I think there’s—

Corey: And that’s fine. That’s the way it works.

Linda: That’s the way it works. And I think there needs to be that niche content that might not be, like, the most viral thing, but viral doesn’t mean quality, you know? It doesn’t. There’s many things that play into what viral is, but it’s important to have the quality content for the people that need that content, and finding those people, you know, it’s easier when you have that kind of mass engagement. Like, who knows? I’m a student. I told you; I’m a professional student. I’m still [laugh] learning every day.

Corey: Working with AWS almost makes it a requirement. I wish you luck—

Linda: Yeah.

Corey: —in the new gig and I also want to thank you for taking time out of your day to speak with me about how you got to this point. And we’re all very eager to see where you go from here.

Linda: Thank you so much, Corey, for having me. I’m a huge fan, I love your content, I’m an avid reader of your newsletter and I am looking forward to very much being in touch and on the Twitterverse and beyond. So. [laugh].

Corey: If people want to learn more about what you’re up to, and other assorted nonsense, where’s the best place they can go to find you?

Linda: So, the best place they could go is lindavivah.com. I have all my different social handles listed on there as well a little bit about me, and I hope to connect with you. So, definitely go to lindavivah.com.

Corey: And that link will, of course, be in the [show notes 00:30:39]. Thank you so much for taking the time to speak with me. I really appreciate it.

Linda: Thank you, Corey. Have a wonderful rest of the day.

Corey: Linda Haviv, AWS Developer Advocate, very soon now anyway. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, smash the like and subscribe buttons, and of course, leave an angry comment that you have broken down into 40 serialized TikTok videos.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

Full Description / Show Notes

  • Steren and Corey talk about how Google Cloud Run got its name (00:49)
  • Corey talks about his experiences using Google Cloud (2:42)
  • Corey and Steven discuss Google Cloud’s cloud run custom domains (10:01)
  • Steren talks about Cloud Run’s high developer satisfaction and scalability (15:54)
  • Corey and Steven talk about Cloud Run releases at Google I/O (23:21)
  • Steren discusses the majority of developer and customer interest in Google’s cloud product (25:33)
  • Steren talks about his 20% projects around sustainability (29:00)

About Steren

Steren is a Senior Product Manager at Google Cloud. He is part of the serverless team, leading Cloud Run. He is also working on sustainability, leading the Google Cloud Carbon Footprint product.

Steren is an engineer from École Centrale (France). Prior to joining Google, he was CTO of a startup building connected objects and multi device solutions.

Links Referenced:

  • Google Cloud Run: https://cloud.run
  • sheets-url-shortener: https://github.com/ahmetb/sheets-url-shortener
  • snark.cloud/run: https://snark.cloud/run
  • Twitter: https://twitter.com/steren

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’m joined today by Steren Giannini, who is a senior product manager at Google Cloud, specifically on something called Google Cloud Run. Steren, thank you for joining me today.

Steren: Thanks for inviting me, Corey.

Corey: So, I want to start at the very beginning of, “Oh, a cloud service. What are we going to call it?” “Well, let’s put the word cloud in it.” “Okay, great. Now, it is cloud, so we have to give it a vague and unassuming name. What does it do?” “It runs things.” “Genius. Let’s break and go for work.” Now, it’s easy to imagine that you spent all of 30 seconds on a name, but it never works that way. How easy was it to get to Cloud Run as a name for the service?

Steren: [laugh]. Such a good question because originally it was not named Cloud Run at all. The original name was Google Serverless Engine. But a few people know that because they’ve been helping us since the beginning, but originally it was Google Serverless Engine. Nobody liked the name internally, and I think at one point, we wondered, “Hey, can we drop the engine structure and let’s just think about the name. And what does this thing do?” “It runs things.”

We already have Cloud Build. Well, wouldn’t it be great to have Cloud Run to pair with Cloud Build so that after you’ve built your containers, you can run them? And that’s how we ended up with this very simple Cloud Run, which today seems so obvious, but it took us a long time to get to that name, and we actually had a lot of renaming to do because we were about to ship with Google Serverless Engine.

Corey: That seems like a very interesting last-minute change because it’s not just a find and replace at that point, it’s—

Steren: No.

Corey: —“Well, okay, if we call it Cloud Run, which can also be a verb or a noun, depending, is that going to change the meaning of some sentences?” And just doing a find and replace without a proofread pass as well, well, that’s how you wind up with funny things on Twitter.

Steren: API endpoints needed to be changed, adding weeks of delays to the launch. That is why we—you know, [laugh] announced in 2018 and publicly launched in 2019.

Corey: I’ve been doing a fair bit of work in cloud for a while, and I wound up going down a very interesting path. So, the first native Google Cloud service—not things like WP Engine that ride on top of GCP—but my first native Google Cloud Service was done in service of this podcast, and it is built on Google Cloud Run. I don’t think I’ve told you part of this story yet, but it’s one of the reasons I reached out to invite you onto the show. Let me set the stage here with a little bit of backstory that might explain what the hell I’m talking about.

As listeners of this show are probably aware, we have sponsors whom we love and adore. In the early days of this show, they would say, “Great, we want to tell people about our product”—which is the point of a sponsorship—“And then send them to a URL.” “Great. What’s the URL?” And they would give me something that was three layers deep, then with a bunch of UTM tracking parameters at the end.

And it’s, “You do realize that no one is going to be sitting there typing all of that into a web browser?” At best, you’re going to get three words or so. So, I built myself a URL redirector, snark.cloud. I can wind up redirecting things in there anywhere it needs to go.

And for a long time, I did this on top of S3 and then put CloudFront in front of it. And this was all well and good until, you know, things happened in the fullness of time. And now holy crap, I have an operations team involved in things, and maybe I shouldn’t be the only person that knows how to work on all of these bits and bobs. So, it was time to come up with something that had a business user-friendly interface that had some level of security, so I don’t wind up automatically building out a spam redirect service for anything that wants to, and it needs to be something that’s easy to work with. So, I went on an exploration.

So, at first it showed that there were—like, I have an article out that I’ve spoken about before that there are, “17 Ways to Run Containers on AWS,” and then I wrote the sequel, “17 More Ways to Run Containers on AWS.” And I’m keeping a list, I’m almost to the third installation of that series, which is awful. So, great. There’s got to be some ways to build some URL redirect stuff with an interface that has an admin panel. And I spent three days on this trying a bunch of different things, and some were running on deprecated versions of Node that wouldn’t build properly and others were just such complex nonsense things that had got really bad. I was starting to consider something like just paying for Bitly or whatnot and making it someone else’s problem.

And then I stumbled upon something on GitHub that really was probably one of the formative things that changed my opinion of Google Cloud for the better. And within half an hour of discovering this thing, it was up and running. I did the entire thing, start to finish, from my iPad in a web browser, and it just worked. It was written by—let me make sure I get his name correct; you know, messing up someone’s name is a great way to say that we don’t care about them—Ahmet Balkan used to work at Google Cloud; now he’s over at Twitter. And he has something up on GitHub that is just absolutely phenomenal about this, called sheets-url-shortener.

And this is going to sound wild, but stick with me. The interface is simply a Google Sheet, where you have one column that has the shorthand slug—for example, run; if you go to snark.cloud/run, it will redirect to Google Cloud Run’s website. And the second column is where you want it to go. The end.

And whenever that gets updated, there’s of course some caching issues, which means it can take up to five seconds from finishing that before it will actually work across the entire internet. And as best I can tell, that is fundamentally magic. But what made it particularly useful and magic, from my perspective, was how easy it was to get up and running. There was none of this oh, but then you have to integrate it with Google Sheets and that’s a whole ‘nother team so there’s no way you’re going to be able to figure that out from our Docs. Go talk to them and then come back in the day.

They were the get started, click here to proceed. It just worked. And it really brought back some of the magic of cloud for me in a way that I hadn’t seen in quite a while. So, all which is to say, amazing service, I continue to use it for all of these sponsored links, and I am still waiting for you folks to bill me, but it fits comfortably in the free tier because it turns out that I don’t have hundreds of thousands of people typing it in every week.

Steren: I’m glad it went well. And you know, we measure tasks success for Cloud Run. And we do know that most new users are able to deploy their apps very quickly. And that was the case for you. Just so you know, we’ve put a lot of effort to make sure it was true, and I’ll be glad to tell you more about all that.

But for that particular service, yes, I suppose Ahmet—who I really enjoyed working with on Cloud Run, he was really helpful designing Cloud Run with us—has open-sourced this side project. And basically, you might even have clicked on a deploy to Cloud Run button on GitHub, right, to deploy it?

Corey: That is exactly what I did and it somehow just worked and—

Steren: Exactly.

Corey: And it knew, even logging into the Google Cloud Console because it understands who I am because I use Google Docs and things, I’m already logged in. None of this, “Oh, which one of these 85 credential sets is it going to be?” Like certain other clouds. It was, “Oh, wow. Wait, cloud can be easy and fun? When did that happen?”

Steren: So, what has happened when you click that deploy to Google Cloud button, basically, the GitHub repository was built into a container with Cloud Build and then was deployed to Cloud Run. And once on Cloud Run, well, hopefully, you have forgotten about it because that’s what we do, right? We—give us your code, in a container if you know containers if you don’t just—we support, you know, many popular languages, and we know how to build them, so don’t worry about that. And then we run it. And as you said, when there is low traffic or no traffic, it scales to zero.

When there is low traffic, you’re likely going to stay under the generous free tier. And if you have more traffic for, you know, Screaming in the Cloud suddenly becoming a high destination URL redirects, well, Cloud Run will scale the number of instances of this container to be able to handle the load. Cloud Run scales automatically and very well, but only—as always—charging you when you are processing some requests.

Corey: I had to fork and make a couple of changes myself after I wound up doing some testing. The first was to make the entire thing case insensitive, which is—you know, makes obvious sense. And the other was to change the permanent redirect to a temporary redirect because believe it or not, in the fullness of time, sometimes sponsors want to change the landing page in different ways for different campaigns and that’s fine by me. I just wanted to make sure people’s browser cache didn’t remember it into perpetuity. But it was easy enough to run—that was back in the early days of my exploring Go, which I’ve been doing this quarter—and in the couple of months this thing has been running it has been effectively flawless.

It’s set it; it’s forget it. The only challenges I had with it are it was a little opaque getting a custom domain set up that—which is still in beta, to be clear—and I’ve heard some horror stories of people saying it got wedged. In my case, no, I deployed it and I started refreshing it and suddenly, it start throwing an SSL error. And it’s like, “Oh, that’s not good, but I’m going to break my own lifestyle here and be patient for ten minutes.” And sure enough, it cleared itself and everything started working. And that was the last time I had to think about any of this. And it just worked.

Steren: So first, Cloud Run is HTTPS only. Why? Because it’s 2020, right? It’s 2022, but—

Corey: [laugh].

Steren: —it’s launched in 2020. And so basically, we have made a decision that let’s just not accept HTTP traffic; it’s only HTTPS. As a consequence, we need to provision a cert for your custom domain. That is something that can take some time. And as you said, we keep it in beta or in preview because we are not yet satisfied with the experience or even the performance of Cloud Run custom domains, so we are actively working on fixing that with a different approach. So, expect some changes, hopefully, this year.

Corey: I will say it does take a few seconds when people go to a snark.cloud URL for it to finish resolving, and it feels on some level like it’s almost like a cold start problem. But subsequent visits, the same thing also feel a little on the slow and pokey side. And I don’t know if that’s just me being wildly impatient, if there’s an optimization opportunity, or if that’s just inherent to the platform that is not under current significant load.

Steren: So, it depends. If the Cloud Run service has scaled down to zero, well of course, your service will need to be started. But what we do know, if it’s a small Go binary, like something that you mentioned, it should really take less than, let’s say, 500 milliseconds to go from zero to one of your container instance. Latency can also be due to the way the code is running. If it occurred is fetching things from Google Sheets at every startup, that is something that could add to the startup latency.

So, I would need to take a look, but in general, we are not spinning up a virtual machine anytime we need to scale horizontally. Like, our infrastructure is a multi-tenant, rapidly scalable infrastructure that can materialize a container in literally 300 milliseconds. The rest of the latency comes from what does the container do at startup time?

Corey: Yeah, I just ran a quick test of putting time in front of a curl command. It looks like it took 4.83 seconds. So, enough to be perceptive. But again, for just a quick redirect, it’s generally not the end of the world and there’s probably something I’m doing that is interesting and odd. Again, I did not invite you on the show to file a—

Steren: [laugh].

Corey: Bug report. Let’s be very clear here.

Steren: Seems on the very high end of startup latencies. I mean, I would definitely expect under the second. We should deep-dive into the code to take a look. And by the way, building stuff on top of spreadsheets. I’ve done that a ton in my previous lives as a CTO of a startup because well, that’s the best administration interface, right? You just have a CRUD UI—

Corey: [unintelligible 00:12:29] world and all business users understand it. If people in Microsoft decided they were going to change Microsoft Excel interface, even a bit, they would revert the change before noon of the same day after an army of business users grabbed pitchforks and torches and marched on their headquarters. It’s one of those things that is how the world runs; it is the world’s most common IDE. And it’s great, but I still think of databases through the lens of thinking about it as a spreadsheet as my default approach to things. I also think of databases as DNS, but that’s neither here nor there.

Steren: You know, if you have maybe 100 redirects, that’s totally fine. And by the way, the beauty of Cloud Run in a spreadsheet, as you mentioned is that Cloud Run services run with a certain identity. And this identity, you can grant it permissions. And in that case, what I would recommend if you haven’t done so yet, is to give an identity to your Cloud Run service that has the permission to read that particular spreadsheet. And how you do that you invite the email of the service account as a reader of your spreadsheet, and that’s probably what you did.

Corey: The click button to the workflow on Google Cloud automatically did that—

Steren: Oh, wow.

Corey: —and taught me how to do it. “Here’s the thing that look at. The end.” It was a flawless user-onboarding experience.

Steren: Very nicely done. But indeed, you know, there is this built-in security which is the principle of minimal permission, like each of your Cloud Run service should basically only be able to read and write to the backing resources that they should. And by default, we give you a service account which has a lot of permissions, but our recommendation is to narrow those permissions to basically only look at the cloud storage buckets that the service is supposed to look at. And the same for a spreadsheet.

Corey: Yes, on some level, I feel like I’m going to write an analysis of my own security approach. It would be titled, “My God, It's Full Of Stars” as I look at the IAM policies of everything that I’ve configured. The idea of least privilege is great. What I like about this approach is that it made it easy to do it so I don’t have to worry about it. At one point, I want to go back and wind up instrumenting it a bit further, just so I can wind up getting aggregate numbers of all right, how many times if someone visited this particular link? It’ll be good to know.

And I don’t know… if I have to change permissions to do that yet, but that’s okay. It’s the best kind of problem: future Corey. So, we’ll deal with that when the time comes. But across the board, this has just been a phenomenal experience and it’s clear that when you were building Google Cloud Run, you understood the assignment. Because I was looking for people saying negative things about it and by and large, all of its seem to come from a perspective of, “Well, this isn’t going to be the most cost-effective or best way to run something that is hyperscale, globe-spanning.”

It’s yes, that’s the thing that Kubernetes was originally built to run and for some godforsaken reason people run their blog on it instead now. Okay. For something that is small, scales to zero, and has long periods where no one is visiting it, great, this is a terrific answer and there’s absolutely nothing wrong with that. It’s clear that you understood who you were aiming at, and the migration strategy to something that is a bit more, I want to say robust, but let’s be clear what I mean when I’m saying that if you want something that’s a little bit more impressive on your SRE resume as you’re trying a multi-year project to get hired by Google or pretend you got hired by Google, yeah, you can migrate to something else in a relatively straightforward way. But that this is up, running, and works without having to think about it, and that is no small thing.

Steren: So, there are two things to say here. The first is yes, indeed, we know we have high developer satisfaction. You know, we measure this—in Google Cloud, you might have seen those small satisfaction surveys popping up sometimes on the user interface, and you know, we are above 90% satisfaction score. We hire third parties to help us understand how usable and what satisfaction score would users get out of Cloud Run, and we are constantly getting very, very good results, in absolute but also compared to the competition.

Now, the other thing that you said is that, you know, Cloud Run is for small things, and here while it is definitely something that allows you to be productive, something that strives for simplicity, but it also scales a lot. And contrary to other systems, you do not have any pre-provisioning to make. So, we have done demos where we go from zero to 10,000 container instances in ten seconds because of the infrastructure on which Cloud Run runs, which is fully managed and multi-tenant, we can offer you this scale on demand. And many of our biggest customers have actually not switched to something like Kubernetes after starting with Cloud Run because they value the low maintenance, the no infrastructure management that Cloud Run brings them.

So, we have like Ikea, ecobee… for example ecobee, you know, the smart thermostats are using Cloud Run to ingest events from the thermostat. I think Ikea is using Cloud Run more and more for more of their websites. You know, those companies scale, right? This is not, like, scale to zero hobby project. This is actually production e-commerce and connected smart objects production systems that have made the choice of being on a fully-managed platform in order to reduce their operational overhead.

[midroll 00:17:54]

Corey: Let me be clear. When I say scale—I think we might be talking past each other on a small point here. When I say scale, I’m talking less about oh tens or hundreds of thousands of containers running concurrently. I’m talking in a more complicated way of, okay, now we have a whole bunch of different microservices talking to one another and affinity as far as location to each other for data transfer reasons. And as you start beginning to service discovery style areas of things, where we build a really complicated applications because we hired engineers and failed to properly supervise them, and that type of convoluted complex architecture.

That’s where it feels like Cloud Run increasingly, as you move in that direction, starts to look a little bit less like the tool of choice. Which is fine, I want to be clear on that point. The sense that I’ve gotten of it is a great way to get started, it’s a great way to continue running a thing you don’t have to think about because you have a day job that isn’t infrastructure management. And it is clear to—as your needs change—to either remain with the service or pivot to a very close service without a whole lot of retooling, which is key. There’s not much of a lock-in story to this, which I love.

Steren: That was one of the key principles when we started to design Cloud Run was, you know, we realized the industry had agreed that the container image was the standard for the deployment artifact of software. And so, we just made the early choice of focusing on deploying containers. Of course, we are helping users build those containers, you know, we have things called build packs, we can continuously deploy from GitHub, but at the end of the day, the thing that gets auto-scaled on Cloud Run is a container. And that enables portability.

As you said. You can literally run the same container, nothing proprietary in it, I want to be clear. Like, you’re just listening on a port for some incoming requests. Those requests can be HTTP requests, events, you know, we have products that can push events to Cloud Run like Eventarc or Pub/Sub. And this same container, you can run it on your local machine, you can run it on Kubernetes, you can run it on another cloud. You’re not locked in, in terms of API of the compute.

We even went even above and beyond by having the Cloud Run API looks like a Kubernetes API. I think that was an extra effort that we made. I’m not sure people care that much, but if you look at the Cloud Run API, it is actually exactly looking like Kubernetes, Even if there is no Kubernetes at all under the hood; we just made it for portability. Because we wanted to address this concern of serverless which was lock-in. Like, when you use a Function as a Service product, you are worried that the architecture that you are going to develop around this product is going to be only working in this particular cloud provider, and you’re not in control of the language, the version that this provider has decided to offer you, you’re not in control of more of the complexity that can come as you want to scan this code, as you want to move this code between staging and production or test this code.

So, containers are really helping with that. So, I think we made the right choice of this new artifact that to build Cloud Run around the container artifact. And you know, at the time when we launched, it was a little bit controversial because back in the day, you know, 2018, 2019, serverless really meant Functions as a Service. So, when we launched, we little bit redefined serverless. And we basically said serverless containers. Which at the time were two worlds that in the same sentence were incompatible. Like, many people, including internally, had concerns around—

Corey: Oh, the serverless versus container war was a big thing for a while. Everyone was on a different side of that divide. It’s… containers are effectively increasingly—and I know, I’ll get email for this, and I don’t even slightly care, they’re a packaging format—

Steren: Exactly.

Corey: —where it solves the problem of how do I build this thing to deploy on Debian instances? And Ubuntu instances, and other instances, God forbid, Windows somewhere, you throw a container over the wall. The end. Its DevOps is about breaking down the walls between Dev and Ops. That’s why containers are here to make them silos that don’t have to talk to each other.

Steren: A container image is a glorified zip file. Literally. You have a set of layers with files in them, and basically, we decided to adopt that artifact standard, but not the perceived complexity that existed at the time around containers. And so, we basically merged containers with serverless to make something as easy to use as a Function as a Service product but with the power of bringing your own container. And today, we are seeing—you mentioned, what kind of architecture would you use Cloud Run for?

So, I would say now there are three big buckets. The obvious one is anything that is a website or an API, serving public internet traffic, like your URL redirect service, right? This is, you have an API, takes a request and returns a response. It can be a REST API, GraphQL API. We recently added support for WebSockets, which is pretty unique for a service offering to support natively WebSockets.

So, what I mean natively is, my client can open a socket connection—a bi-directional socket connection—with a given instance, for up to one hour. This is pretty unique for something that is as fully managed as Cloud Run.

Corey: Right. As we’re recording this, we are just coming off of Google I/O, and there were a number of announcements around Cloud Run that were touching it because of, you know, strange marketing issues. I only found out that Google I/O was a thing and featured cloud stuff via Twitter at the time it was happening. What did you folks release around Cloud Run?

Steren: Good question, actually. Part of the Google I/O Developer keynote, I pitched a story around how Cloud Run helps developers, and the I/O team liked the story, so we decided to include that story as part of the live developer keynote. So, on stage, we announced Cloud Run jobs. So now, I talked to you about Cloud Run services, which can be used to expose an API, but also to do, like, private microservice-to-microservice communication—because cloud services don’t have to be public—and in that case, we support GRPC and, you know, a very strong security mechanism where only Service A can invoke Service B, for example, but Cloud Run jobs are about non-request-driven containers. So, today—I mean, before Google I/O a few days ago, the only requirement that we imposed on your container image was that it started to listen for requests, or events, or GRPC—

Corey: Web requests—

Steren: Exactly—

Corey: It speaks [unintelligible 00:24:35] you want as long as it’s HTTP. Yes.

Steren: That was the only requirement we asked you to have on your container image. And now we’ve changed that. Now, if you have a container that basically starts and executes to completion, you can deploy it on a Cloud Run job. So, you will use Cloud Run jobs for, like, daily batch jobs. And you have the same infrastructure, so on-demand, you can go from zero to, I think for now, the maximum is a hundred tasks in parallel, for—of course, you can run many tasks in sequence, but in parallel, you can go from zero to a hundred, right away to run your daily batch job, daily admin job, data processing.

But this is more in the batch mode than in streaming mode. If you would like to use a more, like, streaming data processing, than a Cloud Run service would still be the best fit because you can literally push events to it, and it will auto-scale to handle any number of events that it receives.

Corey: Do you find that the majority of customers are using Cloud Run for one-off jobs that barely will get more than a single container, like my thing, or do you find that they’re doing massively parallel jobs? Where’s the lion’s share of developer and customer interest?

Steren: It’s both actually. We have both individual developers, small startups—which really value the scale to zero and pay per use model of Cloud Run. Your URL redirect service probably is staying below the free tier, and there are many, many, many users in your case. But at the same time, we have big, big, big customers who value the on-demand scalability of Cloud Run. And for these customers, of course, they will probably very likely not scale to zero, but they value the fact that—you know, we have a media company who uses Cloud Run for TV streaming, and when there is a soccer game somewhere in the world, they have a big spike of usage of requests coming in to their Cloud Run service, and here they can trust the rapid scaling of Cloud Run so they don’t have to pre-provision things in advance to be able to serve that sudden traffic spike.

But for those customers, Cloud Run is priced in a way so that if you know that you’re going to consume a lot of Cloud Run CPU and memory, you can purchase Committed Use Discounts, which will lower your bill overall because you know you are going to spend one dollar per hour on Cloud Run, well purchase a Committed Use Discount because you will only spend 83 cents instead of one dollar. And also, Cloud Run and comes with two pricing model, one which is the default, which is the request-based pricing model, which is basically you only have CPU allocated to your container instances if you are processing at least one request. But as a consequence of that, you are not paying outside of the processing of those requests. Those containers might stay up for you, one, ready to receive new requests, but you’re not paying for them. And so, that is—you know, your URL redirect service is probably in that mode where yes when you haven’t used it for a while, it will scale down to zero, but if you send one request to it, it will serve that request and then it will stay up for a while until it decides to scale down. But you the user only pays when you are processing these specific requests, a little bit like a Function as a Service product.

Corey: Scales to zero is one of the fundamental tenets of serverless that I think that companies calling something serverless, but it always charges you per hour anyway. Yeah, that doesn’t work. Storage, let’s be clear, is a separate matter entirely. I’m talking about compute. Even if your workflow doesn’t scale down to zero ever as a workload, that’s fine, but if the workload does, you don’t get to keep charging me for it.

Steren: Exactly. And so, in that other mode where you decide to always have CPU allocated to your Cloud Run container instances, then you pay for the entire lifecycle of this container instances. You still benefit from the auto-scaling of Cloud Run, but you will pay for the lifecycle and in that case, the price points are lower because you pay for a longer period of time. But that’s more the price model that those bigger customers will take because at their scale, they basically always receive requests, so they already to pay always, basically.

Corey: I really want to thank you for taking the time to chat with me. Before you go, one last question that we’ll be using as a teaser for the next episode that we record together. It seems like this is a full-time job being the product manager on Cloud Run, but no Google, contrary to popular opinion, does in fact, still support 20% projects. What’s yours?

Steren: So, I’ve been looking to work on Cloud Run since it was a prototype, and you know, for a long time, we’ve been iterating privately on Cloud Run, launching it, seeing it grow, seeing it adopted, it’s great. It’s my full-time job. But on Fridays, I still find the time to have a 20% project, which also had quite a bit of impact. And I work on some sustainability efforts for Google Cloud. And notably, we’ve released two things last year.

The first one is that we are sharing some carbon characteristics of Google Cloud regions. So, if you have seen those small leaves in the Cloud Console next to the regions that are emitting the less carbon, that’s something that I helped bring to life. And the second one, which is something quite big, is we are helping customers report and reduce their gross carbon emissions of their Google Cloud usage by providing an out of the box reporting tool called Google Cloud Carbon Footprint. So, that’s something that I was able to bootstrap with a team a little bit on the side of my Cloud Run project, but I was very glad to see it launched by our CEO at the last Cloud Next Conference. And now it is a fully-funded project, so we are very glad that we are able to help our customers better meet their sustainability goals themselves.

Corey: And we will be talking about it significantly on the next episode. We’re giving a teaser, not telling the whole story.

Steren: [laugh].

Corey: I really want to thank you for being as generous with your time as you are. If people want to learn more, where can they find you?

Steren: Well, if they want to learn more about Cloud Run, we talked about how simple was that name. It was obviously not simple to find this simple name, but the domain is https://cloud.run.

Corey: We will also accept snark.cloud/run, I will take credit for that service, too.

Steren: [laugh]. Exactly.

Corey: There we are.

Steren: And then, people can find me on Twitter at @steren, S-T-E-R-E-N. I’ll be happy—I’m always happy to help developers get started or answer questions about Cloud Run. And, yeah, thank you for having me. As I said, you successfully deployed something in just a few minutes to Cloud Run. I would encourage the audience to—

Corey: In spite of myself. I know, I’m as surprised as anyone.

Steren: [laugh].

Corey: The only snag I really hit was the fact that I was riding shotgun when we picked up my daughter from school and went through a dead zone. It’s like, why is this thing not loading in the Google Cloud Console? Yeah, fix the cell network in my area, please.

Steren: I’m impressed that you did all of that from an iPad. But yeah, to the audience give Cloud Run the try. You can really get started connecting your GitHub repository or deploy your favorite container image. And we’ve worked very hard to ensure that usability was here, and we know we have pretty strong usability scores. Because that was a lot of work to simplicity, and product excellence and developer experience is a lot of work to get right, and we are very proud of what we’ve achieved with Cloud Run and proud to see that the developer community has been very supportive and likes this product.

Corey: I’m a big fan of what you’ve built. And well, of course, it links to all of that in the show notes. I just want to thank you again for being so generous with your time. And thanks again for building something that I think in many ways showcases the best of what Google Cloud has to offer.

Steren: Thanks for the invite.

Corey: We’ll talk again soon. Steren Giannini is a senior product manager at Google Cloud, on Cloud Run. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice. If it’s on YouTube, put the thumbs up and the subscribe buttons as well, but in the event that you hated it also include an angry comment explaining why your 20% project is being a shithead on the internet.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

Full Description / Show Notes

  • Gafnit explains how she found a vulnerability in RDS, an Amazon database service (1:40)
  • Gafnit and Corey discuss the concept of not being able to win in cloud security (7:20)
  • Gafnit talks about transparency around security breaches (11:02)
  • Corey and Gafnit discuss effectively communicating with customers about security (13:00)
  • Gafnit answers the question “Did you come at the RDS vulnerability exploration from a perspective of being deeper on the Postgres side or deeper on the AWS side? (18:10)
  • Corey and Gafnit talk about the risk of taking a pre-existing open source solution and offering it as a managed service (19:07)
  • Security measures in cloud-native approaches versus cloud-hosted (22:41)
  • Gafnit and Corey discuss the security community (25:04)

About Gafnit

Gafnit Amiga is the Director of Security Research at Lightspin. Gafnit has 7 years of experience in Application Security and Cloud Security Research. Gafnit leads the Security Research Group at Lightspin, focused on developing new methods to conduct research for new cloud native services and Kubernetes. Previously, Gafnit was a lead product security engineer at Salesforce focused on their core platform and a security researcher at GE Digital. Gafnit holds a Bs.c in Computer Science from IDC Herzliya and a student for Ms.c in Data Science.

Links Referenced:

  • Lightspin: https://www.lightspin.io/
  • Twitter: https://twitter.com/gafnitav
  • LinkedIn: https://www.linkedin.com/in/gafnit-amiga-b1357b125/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. We’ve taken a bit of a security bent to the conversations that we’ve been having on this show and over the past year or so and, well, today’s episode is no different. In fact, we’re going a little bit deeper than we normally tend to. My guest today is Gafnit Amiga, who’s the Director of Security Research at Lightspin. Gafnit, thank you for joining me.

Gafnit: Hey, Corey. Thank you for inviting me to the show.

Corey: You sort of burst onto the scene—and by ‘scene,’ I of course mean the cloud space, at least to the level of community awareness—back, I want to say in April of 2022 when you posted a very in-depth blog post about exploiting RDS and some misconfigurations on AWS’s side to effectively display internal service credentials for the RDS service itself. Now, that sounds like it’s one of those incredibly deep, incredibly murky things because it is, let’s be clear. At a high level, can you explain to me exactly what it is that you found and how you did it?

Gafnit: Yes, so, RDS is database service of Amazon. It’s a managed service where you can choose the engine that you prefer. One of them is Postgres. There, I found the vulnerability. The vulnerability was in the extension in the log_fdw—so it’s for—like, stands for Foreign Data Wrapper—where this extension is, therefore reading the logs directly of the engine, and then you can query it using SQL queries, which should be simpler and easy to use.

And this extension enables you to provide a path. And there was a path traversal, but the traversal happened only when you dropped a validation of the wrapper. And this is how I managed to read local files from the database EC2 machine, which shouldn’t happen because this is a managed service and you shouldn’t have any access to the underlying host.

Corey: It’s always odd when the abstraction starts leaking, from an AWS perspective. I know that a friend of mine was on Aurora during the beta and was doing some high-performance work and suddenly started seeing SQL errors about /var/temp filling up, which is, for those who are not well versed in SQL, and even for those who are, that’s not the sort of thing you tend to expect to show up on there. It feels like the underlying system tends to leak in—particularly in RDS sense—into what is otherwise at least imagined to be a fully-managed service.

Gafnit: Yes because sometimes they want to give you an informative error so you will be able to realize what happened and what caused to the error, and sometimes they prefer not to give you too many information because they don’t want you to get to the underlying machine. This is why, for example, you don’t get a regular superuser; you have an RDS superuser in the database.

Corey: It seems to me that this is sort of a problem of layering different security models on top of each other. If you take a cloud-native database that they designed, start to finish, themselves, like DynamoDB, the entire security model for Dynamo, as best I can determine, is wrapped up within IAM. So, if you know IAM—spoiler, nobody knows IAM completely, it seems—but if you have that on lock you’ve got it; there’s nothing else you need to think about. Whereas with RDS, you have to layer on IAM to get access to the database and what you’re allowed to do with it.

But then there’s an entirely separate user management system, in many respects, of local users for other Postgres or MySQL or any other systems that were using, to a point where even when they started supporting IRM for authentication to RDS at the database user level. It was flagged in the documentation with a bunch of warnings of, “Don’t do this for high-volume stuff; only do this in development style environments.” So, it’s clear that it has been a difficult marriage, for lack of a better term. And then you have to layer on all the other stuff that if God forbid, you’re in a multi-cloud style environment or working with Kubernetes on top of all of this, and it seems like you’re having to pick and choose between four or five different levels of security modeling, as well as understand how all of those things interplay together. How come we don’t see things like this happening four times a day as a result?

Gafnit: Well, I guess that there are more issues being found, but not always published but I think that this is what makes it more complex for both sides. Creating managed services with resources and third parties that everybody knows. To make it easy for them to use requires a deep understanding of the existing permission models of the service where you want to integrate it with your permission model and how the combination works. So, you actually need to understand how every change is going to affect the restrictions that you want to have. So, for example, if you don’t want the database users to be able to read-write or do a network activity, so you really need to understand the permission model of the Postgres itself. So, it makes it more complicated for development, but it’s also good for researchers because they already know Postgres and they have a good starting point.

Corey: My philosophy has always been when you’re trying to secure something, you need to have at least a topical level of understanding of the entire system, start to finish. One of the problems I’ve had with the idea of microservices as is frequently envisioned is that there’s separation, but not real separation, so you have to hand-wave over a whole bunch of the security model. If you don’t understand something, I believe it’s very difficult to secure it. And let’s be honest, even if you do understand [laugh] something, it can be very difficult to secure it. And the cloud vendors with IAM and similar systems don’t seem to be doing themselves any favors, given the sheer complexity and the capabilities that they’re demanding of themselves, even for having one AWS service talk to another one, but in the right way.

And it’s finicky, and it’s nuanced, and debugging it becomes a colossal pain. And finally, at least those of us who are bad at these things, finally say, “The hell with it,” and they just grant full access from Service A to Service B—in the confines of a test environment. I’m not quite that nuts myself, most days. And then it’s the biggest lie we always tell ourselves is once we have something overscoped like that, usually for CI/CD, it’s, “Oh, todo: I’ll go back and fix that later.” Yeah, I’m looking back five years ago and that’s still on my todo list.

For some reason, it’s never been the number one priority. And in all likelihood, it won’t be until right after it really should have been my number one priority. It feels like in cloud security particularly, you can’t win, you can only not lose. I always found that to be something of a depressing perspective and I didn’t accept it for the longest time. But increasingly, these days, it started to feel like that is the state of the world. Am I wrong on that? Am I just being too dour?

Gafnit: What do you mean by you cannot lose?

Corey: There’s no winning in security from my perspective because no one is going to say, “All right. We won the security. Problem solved. The end.” Companies don’t view security as a value-add. It is only about a downside risk mitigation play.

It’s, “Yay, another day of not getting breached.” And the failure mode from there is, “Okay, well, we got breached, but we found out about it ourselves immediately internally, rather than reading about it in The New York Times in two weeks.” The winning is just the steady-state, the status quo. It’s just all different flavors of losing beyond that.

Gafnit: So, I don’t think it’s quite the case because I can tell that they do do always an active work on securing the services and their structure because I went over other extensions before reaching to the log foreign data wrapper, and they actually excluded high-risk functionalities that could help me to achieve privileged access to the underlying host. And they do it with other services as well because they do always do the security review before having it integrated externally. But you know, it’s an endless zone. You can always have something. Security vulnerabilities are always [arrays 00:09:06]. So everyone, whenever they can help and to search and to give their value, it’s appreciated.

Corey: I feel like I need to clarify a bit of nuance. When your blog post first came out talking about this, I was, well let’s say a little irritated toward AWS on Twitter and other places. And Twitter is not a place for nuance, it is easy to look at that and think, “Oh, I was upset at AWS for having a vulnerability.” I am not, I want to be very clear on that. Now, it’s certainly not good, but these are computers; that is the nature of how they work.

If you want to completely secure computer, cut the power to it, sink it in concrete and then drop it in the ocean. And even then, there are exceptions to all of that. So, it’s always a question of not blocking all risk; it’s about trade-offs and what risk is acceptable. And to AWS is credit, they do say that they practice defense-in-depth. Being able to access the credentials for the running RDS service on top of the instance that it was running on, while that’s certainly not good, isn’t as if you’d suddenly had keys to everything inside of AWS and all their security model crumbles away before you.

They do the right thing and the people working on these things are incredibly good. And they work very hard at these things. My concern and my complaint is, as much as I enjoy the work that you do and reading these blog posts talking about how you did it, it bothers me that I have to learn about a vulnerability in a service for which I pay not small amounts of money—RDS is the number one largest charge in my AWS bill every month—and I have to hear about it from a third-party rather than the vendor themselves. In this case, it was a full day later, where after your blog post went up, and they finally had a small security disclosure on AWS’s site talking about it. And that pattern feels to me like it leads nowhere good.

Gafnit: So, transparency is a key word here. And when I wrote the post, I asked if they want to add anything from their side, and they told that they already reached out to the vulnerable customers and they helped them to migrate to their fixed version. So, from their side, it didn’t felt it’s necessary to add it over there. But I did mention the fact that I did the investigation and no customer data was hurt. Yeah, but I think that if there will be maybe a more organized process for any submission of any vulnerability that where all the steps are aligned, it will help everyone and anyone can be informed with everything that happens.

Corey: I have always been extraordinarily impressed by people who work at AWS and handle a lot of the triaging of vulnerability reports. Zack Glick, before he left, was doing an awful lot of that Dan [Erson 00:12:05] continues to be a one of the bright lights of AWS, from my perspective, just as far as customer communication and understanding exactly what the customer perspective is. And as individuals, I see nothing but stars over at AWS. To be clear, ‘Nothing but Stars’ is also the name of most of my IAM policies, but that’s neither here nor there.

It seems like, on some level, there’s a communications and policy misalignment, on some level, because I look at this and every conversation I ever have with AWS’s security folks, they are eminently reasonable, they’re incredibly intelligent, and they care. There’s no mistaking that they legitimately care. But somewhere at the scale of company they’re at, incentives get crossed, and everyone has a different position they’re looking at these things from, and it feels like that disjointedness leads to almost a misalignment as far as how to effectively communicate things like this to customers.

Gafnit: Yes, it looks like this is the case, but if more things will be discovered and published, I think that they will have eventually an organized process for that. Because I guess the researchers do find things over there, but they’re not always being published for several reasons. But yes, they should work on that. [laugh].

Corey: And that is part of the challenge as well, where AWS does not have a public vulnerability disclosure program. [unintelligible 00:13:30] hacker one, they don’t have a public bug bounty program. They have a vulnerability disclosure email address, and the people working behind that are some of the hardest working folks in tech, but there is no unified way of building a community of researchers around the idea of exploring this. And that is a challenge because you have reported vulnerabilities, I have reported significantly fewer vulnerabilities, but it always feels like it’s a hurry up and wait scenario where the communication is not always immediate and clear. And at best, it feels like we often get a begrudging, “Thank you.”

Versus all right, if we just throw ethics completely out the window and decide instead that now we’re going to wind up focusing on just effectively selling it to the highest bidder, the value of, for example, a hypervisor escape on EC2 for example, is incalculable. There is no amount of money that a bug bounty program could offer for something like that compared to what it is worth to the right bad actor at the right time. So, the vulnerabilities that we hear about are already we’re starting from a basis of people who have a functioning sense of ethics, people who are not deeply compromised trying to do something truly nefarious. What worries me is the story of—what are the stories that we aren’t seeing? What are the things that are being found where instead of fighting against the bureaucracy around disclosure and the rest, people just use them for their own ends? And I’m gratified by the level of response I see from AWS on the things that they do find out about, but I always have to wonder, what aren’t we seeing?

Gafnit: That’s a good question. And it really depends on their side if they choose to expose it or not.

Corey: Part of the challenge too, is the messaging and the communication around it and who gets credit and the rest. And it’s weird, whenever they release some additional feature to one of their big headline services, there are blog posts, there are keynote speeches, there are customer references, they go on speaking tours, and the emails, oh, God, they never stopped the emails talking about how amazing all of these things are. But whenever there’s a security vulnerability or a disclosure like this—and to be fair, AWS’s response to this speaks very well of them—it’s like you have to go sneak down into the dark sub-basement, with the filing cabinet behind the leopard sign and the rest, to even find out that these things exist. And I feel like they’re not doing themselves any favors by developing that reputation for lack of transparency around these things. “Well, while there was no customer impact, so why would we talk about it?”

Because otherwise, you’re setting up a myth that there never is a vulnerability on the side of—what is it that you’re building as a cloud provider. And when there is a problem down the road—because there always is going to be; nothing is perfect—people are going to say, “Hey, wait a minute. You didn’t talk about this. What else haven’t you talked about?”

And it rebounds on them with sometimes really unfortunate side effects. With Azure as a counterexample here, we see a number of Azure exploits where, “Yeah, turned out that we had access to other customers’ data and Azure had no idea until we told them.” And Azure does it statements about, “Oh, we have no evidence of any of this stuff being used improperly.” Okay, that can mean that you’ve either check your logs and things are great or you don’t have logging. I don’t know that necessarily is something I trust.

Conversely, AWS has said in the past, “We have looked at the audit logs for this service dating back to its launch years ago, and have validated that none of that has never been used like this.” One of those responses breeds an awful lot of customer trust. The other one doesn’t. And I just wish AWS knew a little bit more how good crisis communication around vulnerabilities can improve customer trust rather than erode it.

Gafnit: Yes, and I think that, as you said, there will always be vulnerabilities. And I think that we are expecting to find more, so being able to communicate as clearly as you can and to expose things about maybe the fakes and how the investigation is being done, even in a high level, for all the vulnerabilities can gain more trust from the customer side.

Corey: DoorDash had a problem. As their cloud-native environment scaled and developers delivered new features, their monitoring system kept breaking down. In an organization where data is used to make better decisions about technology and about the business, losing observability means the entire company loses their competitive edge. With Chronosphere, DoorDash is no longer losing visibility into their applications suite. The key? Chronosphere is an open-source compatible, scalable, and reliable observability solution that gives the observability lead at DoorDash business, confidence, and peace of mind. Read the full success story at snark.cloud/chronosphere. That's snark.cloud slash C-H-R-O-N-O-S-P-H-E-R-E.

Corey: You have experience in your background specifically around application security and cloud security research. You’ve been doing this for seven years at this point. When you started looking into this, did you come at the RDS vulnerability exploration from a perspective of being deeper on the Postgres side or deeper on the AWS side of things?

Gafnit: So, it was both. I actually came to the RDS lead from another service where there was something [about 00:18:21] in the application level. But then I reached to an RDS and thought, well, it will be really nice to find thing over here and to reach the underlying machine. And when I entered to the RDS zone, I started to look at it from the application security eyes, but you have to know the cloud as well because there are integrations with S3, you need to understand the IAM model. So, you need a mix of both to exploit specifically this kind of issue. But you can also be database experts because the payload is a pure SQL.

Corey: It always seems to me that this is an inherent risk in trying to take something that is pre-existing is an open-source solution—Postgres is one example but there are many more—and offer it as a managed service. Because I think one of the big misunderstandings is that when—well, AWS is just going to take something like Redis and offer that as a managed service, it’s okay, I accept that they will offer a thing that respects the endpoints and then acts as if it were Redis, but under the hood, there is so much in all of these open-source projects that is built for optionality of wherever you want to run this thing, it will run there; whatever type of workload you want to throw at it, it can work. Whereas when you have a cloud provider converting these things into a managed service, they are going to strip out an awful lot of those things. An easy example might be okay, there’s this thing that winds up having to calculate for the way the hard drives on a computer work and from a storage perspective.

Well, all the big cloud providers already have interesting ways that they have solved storage. Every team does not reimplement that particular wheel; they use in-house services. Chubby’s file locking, for example, over on Google side is a classic example of this that they’ve talked about an awful lot so every team building something doesn’t have to rediscover all of that. So, the idea that, oh, we’re just going to take up this open-source thing, clone it off a GitHub, fork it, and then just throw it into production as a managed service seems more than a little naive. What’s your experience around seeing, as you get more [laugh] into the weeds of these things than most customers are allowed to get, what’s your take on this?

Do you find that this looks an awful lot like the open-source version that we all use? Or is it something that looks like it has been heavily customized to take advantage of what AWS is offering internally as underlying bedrock services?

Gafnit: So, from what I saw until now, they do want to save the functionality so you will have the same experience as you’re working with the same service that not on AWS because you’re you are used to that. So, they are not doing dramatic changes, but they do want to reduce the risk in the security space. So, there will be some functionalities that they will not let you to do. And this is because of the managed party in areas where the full workload is deployed in your account and you can access it anyway, so they will not have the same security restrictions because you can access the workload anyway. But when it’s managed, they need to prevent you from accessing the underlying host, for example. And they do the changes, but they’re really picked to the specific actions that can lead you to that.

Corey: It also feels like RDS is something of a, I don’t want to call it a legacy service because it is clearly still very much actively developed, but it’s what we’ll call it a ‘classic service.’ When I look at a new AWS launch, I tend to mentally bucket them into two things. There’s the cloud-native approach, and we’ve already talked about DynamoDB. That would be one example of this. And there’s the cloud-hosted model where you have to worry about things like instances and security groups and the networking stuff, and so on and so forth, where it’s basically feels like they’re running their thing on top of a pile of EC2 instances, and that abstraction starts leaking.

Part of me wonders if looking at some of these older services like RDS, they made decisions in the design and build out of these things that they might not if they were to go ahead and build it out today. I mean, Aurora is an example of what that might look like. Have you found as you start looking around the various security foibles of different cloud services, that the security posture of some of the more cloud-native approaches is better or worse or the same as the cloud-hosted world?

Gafnit: Well, so for example, in the several issues that were found, and also here in the RDS where you can see credentials in a file, this is not a best practice in security space. And so, definitely there are things to improve, even if it’s developed on the provider side. But it’s really hard to answer this question because in a managed area where you don’t have any access, it’s hard to tell how it’s configured and if it’s configured properly. So, you need to have some certification from their side.

Corey: This is, on some level, part of the great security challenge, especially for something that is not itself open-source, where they obviously have terrific security teams, don’t get me wrong. At no point do I want to ever come across a saying, “Oh, those AWS people don’t know how security works.” That is provably untrue. But there is something to be said for the value of having a strong community in the security space focusing on this from the outside of looking at these things, of even helping other people contextualize these things. And I’m a little disheartened that none of the major cloud providers seem to have really embraced the idea of a cloud security community, to the point where the one that I’m most familiar with, the cloud security forum Slack team seems to be my default place where I go for context on things.

Because I dabble. I keep my hand in when it comes to security, but I’m certainly no expert. That’s what people like you are for. I make fun of clouds and I work on the billing parts of it and that’s about as far as it goes for me. But being able to get context around is this a big deal? Is this description that a company is giving, is it accurate?

For example, when your post came out, I had not heard of Lightspin in this context. So, reaching out to a few people I trusted, is this legitimate? The answer was, “Yes. It’s legitimate and it’s brilliant. That’s a company that keep your eye on.” Great. That’s useful context and there’s no way to buy that. It has to come from having those conversations with people in the [broader 00:24:57] sense of the community. What’s your experience been looking at the community side of the world of security?

Gafnit: Well, so I think that the cloud security has a great community, and this is one of the things that we at Lightspin really want to increase and push forward. And we see ourselves as a security-driven company. We always do the best to publish a post, even detailed posts, not about vulnerabilities, about how things works in the cloud and how things are being evaluated, to release open-source tools where you can use them to check your environment even if you’re not a customer. And I think that the community is always willing to explain and to investigate together. And it’s a welcome effort, but I think that the messaging should be also for all layers, you know, also for the DevOps and the developers because it can really help if it will start from this point from their side, as well.

Corey: It needs to be baked in, from start to finish.

Gafnit: Yeah, exactly.

Corey: I really want to thank you for taking the time out of your day to speak with me today. If people want to learn more about what you’re up to, where’s the best place for them to find you?

Gafnit: So, you can find me on Twitter and on LinkedIn, and feel free to reach out.

Corey: We will, of course, put links to that in the [show notes 00:26:25]. Thank you so much for being so generous with your time today. I appreciate it.

Gafnit: Thank you, Corey.

Corey: Gafnit Amiga, Director of Security Research at Lightspin. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, and if it’s on the YouTubes, smash the like and subscribe buttons, which I’m told are there. Whereas if you’ve hated this podcast, same story, like and subscribe and the buttons, leave a five-star review on a various platform, but also leave an insulting, angry comment about how my observation that our IAM policies are all full of stars is inaccurate. And then I will go ahead and delete that comment later because you didn’t set a strong password.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

Full Description / Show Notes

  • Marie talks about Oki Doki’s primary product, Notion Mastery (2:38)
  • Corey and Marie talk ADHD diagnosis and how it has impacted their lives and work (4:26)
  • Marie and Corey discuss techniques they’ve developed for coping with ADHD (11:22)
  • Corey and Marie talk about workarounds for people with ADHD who want to adopt something like Notion (16:13)
  • Marie discusses the importance of being excited about the tools you’re employing (18:54)
  • Corey and Marie talk about finding tools that work for you (26:43)
  • Marie and Corey discuss the unique challenge of teaching skills versus dumping knowledge (30:35)

About Marie Poulin

Marie teaches business owners to level up their digital systems, workflow, and knowledge management processes using Notion.

She’s the co-founder of Oki Doki and creator of Notion Mastery, an online program and community that helps creators, entrepreneurs and small teams tame their work + life chaos by building life and business management systems with Notion.

Diagnosed with ADHD at age 37, Marie is especially passionate about helping folks customize their workflows and workspaces to meet their unique needs and preferences.

She believes that Notion is especially powerful for neurodivergent folks who have long struggled to adhere to traditional or rigid project management processes, and may need a little extra customization and flexibility.

When she's not tinkering in Notion or doing live trainings, you can find her in the garden, playing video games, or cooking up some delicious vegetarian tacos.

Links Referenced:

  • Oki Doki: https://weareokidoki.com/
  • Personal website: https://mariepoulin.com
  • Notion Mastery: https://notionmastery.com
  • Twitter: https://twitter.com/mariepoulin

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Honeycomb. When production is running slow, it’s hard to know where problems originate. Is it your application code, users, or the underlying systems? I’ve got five bucks on DNS, personally. Why scroll through endless dashboards while dealing with alert floods, going from tool to tool to tool that you employ, guessing at which puzzle pieces matter? Context switching and tool sprawl are slowly killing both your team and your business. You should care more about one of those than the other; which one is up to you. Drop the separate pillars and enter a world of getting one unified understanding of the one thing driving your business: production. With Honeycomb, you guess less and know more. Try it for free at honeycomb.io/screaminginthecloud. Observability: it’s more than just hipster monitoring.

Corey: DoorDash had a problem. As their cloud-native environment scaled and developers delivered new features, their monitoring system kept breaking down. In an organization where data is used to make better decisions about technology and about the business, losing observability means the entire company loses their competitive edge. With Chronosphere, DoorDash is no longer losing visibility into their applications suite. The key? Chronosphere is an open-source compatible, scalable, and reliable observability solution that gives the observability lead at DoorDash business, confidence, and peace of mind. Read the full success story at snark.cloud/chronosphere. that’s snark.cloud slash C-H-R-O-N-O-S-P-H-E-R-E.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Today I’m joined by Marie Poulin, the CEO of Oki Doki. Marie, thank you for joining me.

Marie: Thank you for having me. I’m excited.

Corey: So, let’s start at the very beginning. What does Oki Doki do? And for folks listening that is O-K-I D-O-K-I, so you might want to have to think about that if you’re doing the Google approach of, “What is this thing?”

Marie: Well, at the moment, the majority of our products and services are surrounded by helping people learn how to use Notion to manage their life and business. So, it’s only a pivot that we took in the last couple of years, and so our signature program is a course called Notion Mastery. So, there’s four full-time employees now and that’s what we do. We design live trainings, we have a forum, we have a curriculum. It’s all products and services related to Notion.

Corey: That is an interesting pivot that you can wind up going through. Please tell me I’m not the first person to make the observation that you called it Oki Doki and you’ve turned yourself around.

Marie: [laugh]. You are the first, Corey? [laugh].

Corey: Oh, good. I am broken like that, so that’s kind of awesome. So, you’ve been more or less doing—I don’t know the best way to frame this, so my apologies if I’m getting it wrong—but the idea of well, what are you selling? Knowledge. You’re selling an understanding of how to improve things, you’re selling a better outcome.

And it’s easy to look at that and say, “Oh, you’re selling education.” No, you’re selling understanding. Education is the way that you get there because at least at the moment, you can’t just jack gigabytes of data directly into people’s head without going to prison for it. Or raising a whole boatload of VC money.

Marie: [laugh]. I mean, you can also say you’re kind of selling an outcome, right? You’re selling this future version of who someone wants to be. And so, we talk a lot about—you know, on our sales page, we get a lot of compliments on our sales page, but just speaking to the scattered mind, you know, feeling like a shitshow, feeling like you don’t really have all your data in one place. You know, it’s learning how to improve your workflow at work but also in life as well.

And so, a lot of our language speaks to the sort of future version of yourself. Like, stop feeling scattered, stop feeling stretched thin. Let’s actually get it so that you turn things into a well-oiled machine. So, you could say we’re selling a dream. [laugh].

Corey: This is an interesting direction to take this conversation in because I don’t normally talk about this. But why not; we’ll give it a shot. It’s been sufficiently long since the last time. Last year—you’ve been very public about this—you were diagnosed with ADHD. I periodically talk about the fact that I was diagnosed with it myself—back when it was called ADD—when I was five years old.

So, growing up I always knew that there was something neurodivergent about me. And the lesson I took away from this, as someone growing up with a lot of the limitations—yes, there are advantages but at the time, all I saw were limitations—about, “Well, what is ADHD?” It’s like, oh, okay. They sat down and explained it to me. And it’s not what they said, but it was, “See, this is the medical reason why you suck.”

And that was not the most constructive way of framing it. In adulthood, talking to other people who have been diagnosed with this, especially later in life. There’s a—it’s a spectrum disorder. It winds up impacting an awful lot of people differently, but the universal experience that I hear is, wait, you mean there’s a reason that I am the way that I am? It’s not that I’m lazy. It’s not that I’m shitty at things. It’s not that I’m—

Marie: Yeah.

Corey: —careless. And that is one of those things that just is transformative. I didn’t realize at the time how fortunate I was to be diagnosed that early on because trying to try to figure out why am I getting fired all the time? Why do I get bored doing the same thing too many days in a row, so I start causing problems for other people? What is going on with this? Why do I have this incredible opposition to anything that remotely resembles authority, et cetera, et cetera?

Not all of this might be ADHD traits, but here I am. And my only solution after, you know, deciding that I didn’t really want to set a world record for number of times getting fired was, well, I guess I’ll start my own company because that at least to get fired, it’s going to take some work. You figured this out while you were already self-employed.

Marie: Yes.

Corey: What was that like?

Marie: What was it like to find out that I finally had an answer or reason for, maybe, past behaviors? [laugh].

Corey: Right. Because it’s the simultaneous, “Oh, my God, there’s a reason that I am like I am,” and then followed immediately by, “I still am the way that I am. Huh. Okay.” It feels like it helps things, but it also doesn’t help things. But it does, and it comes back around. What was your experience with it?

Marie: Yeah, it started because I was doing research to understand my sister better because she had been diagnosed with ADHD for a couple years. It made so much sense once I kind of understood and started researching a little bit more about it. And then, of course, doing my deep-dive research. I’m hearing all these traits that I’m like, “Oh. Wait, that does really sound like me.” The not being able to wake—

Corey: What do you [mean 00:07:01]—

Marie: —up in the morning—

Corey: ADHD trait? Everyone does that. Wait.

Marie: [laugh]. Yeah. When you said that enough times, you’re like, “Oh, wait. Maybe this is not normal.” Or you don’t really know what is—what is normal anyway, right? So, in doing that research, trying to connect with her, trying to understand her experience better, I just started learning about more and more of these traits.

I also knew a shit ton of people in our course, had mentioned that they had ADHD in their intake form, and I was like, what is it about people that ADHD that are actually drawn to my YouTube videos or my way of explaining things? And I started to learn a little bit more; it’s quite common for folks with ADHD to be drawn to one another, probably because of our communication styles, even the sort of mild interrupting, or kind of the way we banter together. There’s different styles of communicating that I think often folks with ADHD are maybe drawn to one another or have an easier time understanding one another. So, listening to some of these symptoms, I was like, “Wait a second.” Because my sister and I are so different in the way our symptoms present.

I thought, “Well, that’s what ADHD looks like.” It’s pure unbridled chaos and unfiltered. And I just had this idea of what it looked like because she was one of the few examples that I had. Meanwhile, I’m skipping grades, I’m in the gifted program, I’m off, you know, doing my own thing. It looked very different.

I thought, “Oh, people with ADHD don’t thrive in university,” or whatnot. So, I had a lot of assumptions that I had to unpack. And I think the one, sort of, I don’t know, symptom that kind of twinged something in my brain was extreme difficulty getting up in the morning and even sort of waking up your brain in the morning. This has been a problem with jobs, it’s been a problem was school, getting to school on time, getting to work on time. Similar to you, it has caused job loss, it has caused tension with partners. They don’t understand, like, why can’t you get out of bed and seize the day?

And I just thought, “There’s something weird going on there with my body.” But I can be, you know, wide awake at 7 p.m. and I’m, like, ready to go. And I can hyperfocus for days on end. So, just noticing some of these symptoms and kind of unpacking it a bit, I thought, “Okay, there’s something to go a little deeper in here.”

Corey: I have trouble getting up, but I’m almost never late. That one does not hit me in quite the same way. In fact—

Marie: Well—

Corey: —my first consulting clients, and I’d been building—I was independent for two weeks at that point, and I was in an in-person meeting in San Francisco and one day, I showed up 20 minutes late, and he just stared at me. “You’re never late. What’s the deal here?” And it’s like, “Yeah, I had trouble getting up this morning.” That was a lie.

I was able to tell him about three or four months later, that morning, I found out I was going to be a father. And that was an—you know, it turns out that I was going to be okay being late, but it was so early, you didn’t want to tell anyone, yet. But it was—yeah, it’s one of those things where that was more important than—

Marie: Absolutely.

Corey: —doing the work thing. But I still remember, yeah, I feel like I’m always about to be late but apparently my reputation is, I never am, so okay. I’ll take it. That is a—again, it is a spectrum disorder. I also—

Marie: Absolutely.

Corey: —further there want to call out for viewers, listeners, et cetera, a couple of things. One, this is not mental health advice. If any of the stories we’re telling resonate, talk to a qualified mental health professional. Secondly, I want to be clear as well here, Marie, that you and I both have significant advantages when it comes to dealing with these things. We both run our own companies, we can effectively restructure the way that we work in ways that are more accommodating for what we do.

It turns out that in my employment days, that was never really a solution where, “Yeah, I decided I’m not going to wind up doing the on-call checklist every day. It doesn’t resonate with me.”

Marie: “Just not feeling like it.”

Corey: “It’s doing the same thing too many days in a row. And yeah, I’m not going to check the backups, either. What do you mean ‘I’m fired?’” yeah, it turns out, you’re not able to—you’re empowered to make those kinds of sweeping changes in the same way.

Marie: Exactly.

Corey: So, this is not advice for people. This is simply a pair of experience reports, the way I view it.

Marie: Absolutely. I sort of feel like self-employment wasn’t necessarily a choice, in a way. It just felt like that’s the only way I'm going to be able to operate in this world. I need some more sense of control and say in how I structure my days, how I structure my work, being able to switch things up, being able to pivot quickly. I knew that I was going to need more control over that. So yeah, pretty unemployable over here. [laugh].

Corey: So, once you wound up with the diagnosis, what happened next? What changes did you make that wound up resonating for you, things that were actionable? And, yeah, you’ve been very public about it as well. I want to highlight that. I’m not, for the most part.

And part of that is because I internalized growing up that it was somehow a shameful thing that we don’t talk about. And the other part of it, too, on some level, was I didn’t want to turn it into a part of my brand identity, where, “Oh, yeah, Corey is very hard to describe.” So, people thrash around and look for labels to slap on me. ‘Shitposter’ seems to have stuck rather well. Because as soon as people feel that they have a label for something, it becomes easier to classify and then dismiss it.

It’s aspects of my personality. It’s who I am. I don’t think of it as a disorder so much as it is part and parcel of who and what I am. And it turns out that being me is not—yet—a medically recognized diagnosis. So, I’m cautious to avoid the labeling aspect of it.

But you have very publicly not, if not going for the label, you at least embraced it as an aspect of who you are, and you’ve been very vocal about your experiences and telling people how you have overcome aspects of this. It’s admirable. I wish I did more of it, honestly.

Marie: I think it’s kind of essential, I think, in the nature of what we’re teaching. Like, when we’re teaching people to become more organized and we know that executive dysfunction is one of the signs or, you know, issues with ADHD, to me it sort of recontextualized why I became so freakin’ obsessed with systems and organization: because I never felt organized. I always felt the sense of what is the stuff come so easy to other people? Why is it taking me so much longer? Why am I spending nights, evenings, taking courses about systems like I’m trying to understand how to give my life structure?

And so, in a way, the way I have become organized was trial by fire, just teaching myself, learning, you know, getting coaches. Like, I literally had a systems coach to teach myself how to get my business organized. So, I had kind of obsessed over it, like a hyperfocus. And so, realizing that other people are struggling with this and there’s a reason that people with ADHD are coming to the course seeking that sense of control. And so, learning that I had it, I was like, oh, this actually [laugh] does explain, in a way, my obsession with this or my curiosity about this, of, like, why does this come easy to some other people? Why do some people need to study this and learn this? Like, what is it about that?

And so, I sort of felt like it would be doing a disservice if I didn’t kind of name it and talk about it and say, well, this actually colors a lot of my opinions. This actually influences the way I approach organization or even productivity, not from a timing perspective, but from an energy management perspective. I didn’t realize that was something that I’m doing. I’m not managing time, we’re managing Marie’s energy. And even my team is learning how to do that, too.

So, I was like, “Oh, that actually makes a ton of sense.” And it also makes sense why some people won’t resonate with this energy management thing or might think I’m going way too far down a rabbit hole on something and they’re like, “Why can’t people just do what they say?” Like, you don’t understand, some of us need to trick ourselves into being productive. And this is how I’ve learned to do that. So, it was just kind of a funny recontextualizing or uncovering, oh, our brains operate very differently. And even within ADHD, people’s brains operate differently, so how do we get people moving toward progress, but knowing that we kind of need different ways of doing that. So, it’s just been kind of an interesting process.

Corey: There’s a fairly common experience report from folks who have ADHD that when they’re kids, their memory is generally very good with a number of expressions of it, so we form our self-image in a lot of those times. And then for the rest of our lives, we tell ourselves the same lie, regardless of how many times it has proven to be a lie. And that lie is, “I don’t need to write this down. I’ll remember it.”

Marie: Oh yes.

Corey: “No, Corey, you will not remember it. You need to write it down. I promise.” And, for example, right now—I finally gave in and technology leapt ahead to the point where my entire life is run by Google Calendar—specifically three or four of them—that all route through Fantastical—which is the app I use—but it winds up grabbing my attention at the right time. It tells me what I need to do, when, and how, and it’s wonderful.

Because if it’s not on my calendar, it does not happen.

Marie: Yes.

Corey: Like, I will forget our anniversary, my kids’ birthdays, to pick my children up from school. We are talking about, if it is not on my calendar, it does not happen. That is the one system that has been forced on me that worked. Then we—let’s talk about Notion for a minute because I looked at it briefly a few years ago, and it is one in the long, long, long list of tools or approaches or systems that I have played with and then discarded to act as basically an auxiliary brain pack. I used Evernote for a while and that sort of worked because I just would do different notes all the time and I’d wind up with 3000 of those things, and then the app gets bloaty and I move on to something else.

For the last five years or so I’ve been using Drafts, a Mac slash iOS app, that only does text, which makes image management and attaching things kind of hard, but okay. And that’s great, and now I have 5000 of those in my [back 00:16:25] folder, not categorized or organized anyway, so I focus instead on well, search for terms and hope I use the term I thought I did at the time. And so, every time I’ve tried to use something like Notion, it’s yeah, this requires a way of thinking that I know I will get excited about if I look at it, and in a month, I’ll be right back to where I am now. So, there’s only so many times you go on the same ride before you know how it ends. How do y—like, that feels like a very common experience. How did you fix it?

Marie: I think at the core though, you kind of have to be excited about the tool that you’re using. And so, I don’t think—Notion is not going to be an exciting fun tool for everyone. Some people are going to be like, “I don’t want to frickin’ build my productivity system. Are you kidding me? Like, just give me something that works out of the box.” Absolutely.

But I think there’s something about the visual components of Notion. Like, I am a designer; I went to design school. I think I’m—it’s almost like something doesn’t click until I see it in the way that I need to see it. And that’s something I’ve learned about my brain is just, sometimes the same information can be presented to me, but if it’s not in a visual way, or whether it’s not spaced in the right way, my brain just kind of ignores it or it gets overwhelmed by it. And so, for me that visual aspect actually helps me learn.

I’m priming my brain, I’m making my goals front and center. The fact that I can design it the way I need my brain to see it is part of its appeal to me. But I also recognize that’s not something everyone gets excited about. They’re not drawn to it. I’m all for using the tool that works the way that your brain is going to work.

I get excited about making databases. I get excited about building glossaries of information to help me learn things. Like, for me, that’s part of my learning and part of my process and it’s just kind of what I’m used to, but I fully acknowledge, like, that stuff does not get everybody excited.

[midroll 00:18:03]

Corey: This episode is sponsored in part by our friend EnterpriseDB. EnterpriseDB has been powering enterprise applications with PostgreSQL for 15 years. And now EnterpriseDB has you covered wherever you deploy PostgreSQL on-premises, private cloud, and they just announced a fully-managed service on AWS and Azure called BigAnimal, all one word. Don’t leave managing your database to your cloud vendor because they’re too busy launching another half-dozen managed databases to focus on any one of them that they didn’t build themselves. Instead, work with the experts over at EnterpriseDB. They can save you time and money, they can even help you migrate legacy applications—including Oracle—to the cloud. To learn more, try BigAnimal for free. Go to biganimal.com/snark, and tell them Corey sent you.

Corey: There’s something very key you’re talking about here, which is the idea of having to be excited about what it is that you do. I look at the things that I do professionally, and if I didn’t deeply enjoy them, they would not get done, and I would have pivoted long ago to something else. People wonder why—

Marie: Absolutely.

Corey: —I make fun of so many things in the tech ecosystem. The honest answer is because if I just tell the dry, boring version of it, I will get bored because it’s a fairly boring field. Whereas instead, okay, someone releases a new thing. Great. How do I keep it interesting for me? How do I find a way to tell that story?

How do I find a way to, in turn, build that into something that, in turn, I can start dragging in different directions and opening up to new ways of talking without going too far? It’s always a razor’s edge, it’s always a bit of a mind puzzle, and it’s always different. I love that. That’s why I do it. It’s not for the audience so much as it is for myself. Because if I’m not engaged, no one else is going to care what I have to say.

Marie: Absolutely. And I think that’s a huge part of ADHD as well which is that interest-based nervous system, right? It’s like we have to [laugh] trick ourselves into finding the excitement in it or whatever that looks like for each of us. But just if I’m not motivated, if I’m not excited about it—writing email newsletters doesn’t get me excited; I’m like, “Okay, do I need to hire someone to do this?” Or how can I find a way to do it, whether it’s—if making a video is more fun or easy, great.

How can I, you know, make content do double-duty in that way? So yeah, I’m always trying to find ways to incentivize myself to do the things that need to get done, even though they may not be the most exciting. But step one is actually run a business that is based on something that you love doing. Which not everyone, maybe, has the privilege to do, but I think everything about the way I’ve designed my business model and the services that we offer is, don’t offer services you don’t really want to offer. Don’t make products that you don’t want to maintain, you’re not excited about. So, it’s definitely a core part of kind of how we design our whole business model.

Corey: For me, a big part of it has always been just trying to make sure that I’m doing the things that engage me. And this is where that whole idea of being in a very privileged position enters into it. Take this podcast slash video right now, as a terrific example. I’m having this conversation, I have an entire system when I wind up sending a link to someone, it fires off Calendly, that hides webhooks and gets a whole bunch of other things set up. I show up, we have a conversation before the show to figure out just this is the general ebb and flow of the show. Here’s the generalized topics we want to talk about. Let’s dive in.

And we finish the recording session. Great, I wind up closing the window and that’s the last time I generally think about it. Because everything else has been automated. If anything other than me having this conversation with you does not need to be me, I there is no differentiated value in me being the person that does the audio engineering. It turns out, I can pay people who are world’s better than I am at that, who actually enjoy it as opposed to viewing it as unnecessary chore, and I can do things that I find more appealing, like shitposting about a $1.108 trillion—

Marie: Exactly.

Corey: Company. It comes down to find the thing, the differentiation point, and find ways to make sure you don’t have to do the other parts of it. But that is not a path that’s available to everyone in every context. And again, I’m talking about this in a professional sense. I still have to do a whole bunch of stuff as I go through the course of my life that is not differentiated, but I can’t very well hire someone to get me dressed in the morning. Well, I can but I feel like that becomes a little bit out of the scope of the lived human experience most of the [crosstalk 00:22:29].

Marie: [laugh]. Absolutely. I feel like that’s one thing I sort of regret not doing earlier is hiring someone to work with. So, the very first hire that I made was my chief of operations, and oh my gosh, the things that she took on that I used to do that I’m like, how on earth did I do that before? Because now that you do that, and you do it way faster, I just got to wonder, like, how the heck did I ever convince myself to do those activities?

I don’t want to do touch spreadsheets, I don’t want to [laugh] deal with that stuff. I don’t want to, you know, email reminders, or whatever it is. There’s so many activities that she handles that I just… I would be happy to never touch again. And so, I sort of wish I had explored that earlier, but I was in that lone wolf, like, I got this. I’m going to run my own business solo forever.

And, you know, I just sort of thought it’s difficult to work with me or because of the way that I work, I don’t know how to delegate. Like, it’s all in your head. I just didn’t really know how to do that. So, that process, I think, takes a while. That first hire when you’re going from solo person to okay, now we’re two; how do we work together? Okay, who else can we hire? What other activities can I get other people to do? So, that’s been a process, for sure.

Corey: Mike Julian, my business partner who you know, is a very process-driven person. He is very organized. His love language is Microsoft Excel, as I frequently tease him with. And one of the—not the only factor by a landslide, but one of the big early factors of what would—okay, I know what I’d do. What would Mike do here?

Part of it is the never-ending litany of mail I get from the state around things like taxes, business registration, the rest. And normally my response when I get those, is I look at it, and it’s like, “Welp, I’m going to fucking prison. That’s the end of it. The end.” Because it’s not that I don’t have the money to pay my taxes, I assure you. What, I don’t have it—because I—financial planning is kind of part and parcel of how we think about cloud economics.

But no, it’s the fact that I’m not going to sit there, fill out the form, put a stamp on it—or God forbid, fax it somewhere—and the rest. It’s not the paying of the taxes that bothers me it is the paperwork and the process and the heavy lift associated with getting the executive function necessary to do it. So, it never gets done and deadlines slide by. And Mike was good at that for a time, and then he took the more reasonable approach about this of, “Huh. Seems to me like a lot of this stuff is not differentiated value that I need to be doing either.”

So, we have a CFO who handles a lot of that stuff now and other operational folks. And it turns out that yeah, wow, there’s a lot—I can—the quality of what I put out is a lot better because I get to focus on things instead of having to deal with the ebb and flow minutia of running payroll myself every week.

Marie: Oh, yeah. All of that is very relatable. And this is why I can’t do paper in the office. I think this is why I just moved my entire brain online. It’s like if there’s paper, stamps, anything related to having to go [laugh] to a post office to mail something. I think I still have the stack of thank you cards from our wedding from, you know, five years ago. So, yeah. [laugh].

Corey: That you haven’t sent out yet. Of course.

Marie: Yes, exactly.

Corey: Exact same—sorry, people 13—11 years ago, whenever it was.

Marie: I’m so sorry.

Corey: Yeah, one of these years. Yeah, and see, that’s exactly how I treat things like Drafts or Notion, if I were to use it, or something else is great, it’s still going to be the digital equivalent of a giant pile of paper. The thing is that computers can search through the contents of that paper a hell of a lot faster than I can, even with my own, at times, uncanny reading speed. There’s some value to that. So, understanding how the systems work and having them bend to accommodate you, rather than trying to fool yourself in half to work within the confines of an existing system, that seems to be the direction that you’re taking Notion in, specifically in the context of it is not prescriptive.

And, on some level, that’s kind of the problem I have with it. Whenever I try the getting started for us, it’s, “Great, you can build your own system.” It’s like, “Isn’t that your job? What am I missing here?” Because the scariest thing I ever see when it’s time for you to write a blog post or whatnot is an empty editor. It’s, where do I get started? Where’s the rest?

I even built a template that I wind up sometimes using text expander to autofill, that gets me started. And it’s just get—once I get started, it’s great. It’s hard to get me started; it’s hard to get me to stop, in case no one has been aware of that. But it’s been understanding how I work and how that integrates with it. I’m curious, given that you do talk to people who are trying to build these systems for a living for themselves? How common is my perspective on this? Am I out there completely, this unique, beautiful Snowflake? Is it yeah, that’s basically everyone? Or somewhere in between?

Marie: Oh, I definitely don’t think you’re alone with that. And again, I often will dissuade people from taking on Notion. I’m like, “Oh, if you’re just looking for a note-taker, or you’re just looking for something else,” or, “Your tools are already working for you, great. Keep using them.” So, I think it’s quite common. I don’t think Notion is the right tool for everyone.

I think it’s great for very visual people like myself, people that it matters how you are seeing your information, and how much information you’re seeing, and you want more control over that, that’s great. For me, I like the integration. I know that as soon as I’m bouncing around to different tools, like, I just already feel kind of scattered, so I was like, how can I pull everything that I need into these, sort of, singular dashboards. So, my approach is very dashboard-focused. Okay, Marie is going into content mode, it’s time to write a blog. Go to the content hub. On the content hub is your list of most recent ideas, your templates for how to write a blog post. There’s resources for creating video. It’s already there for me; I’m not having to start from scratch like you said.

But again, it took time to build that up for myself. So, I think you’re not alone, and I think some people get excited about that building process; other people get irritated by it, and I don’t think there’s a right or wrong answer. It’s just how do our brains work? Know thyself. And, yeah, I’ve sort of—I think also in a way, something that’s a little different, maybe, about the way that I use Notion is I think of it as a personal development tool.

It is a tool for making me better in different ways. It’s for exploring my interests, it’s for feeding my curiosity, it’s for looking at change over time. I track my feelings every day. I’ve been journaling for 1300 days in a row, which is probably the only thing I’ve done consistently in my life [laugh] in the last couple of years. But now I can look and I can see trends over time in a really beautiful and visual way. And I just, to me, it’s like a curiosity tool, to see, like, where am I going? Where have I been? What do I want more of?

Corey: I need to look into this a bit more because my idea of a well-designed user interface is—I’m very opinionated on this—but it comes down to the idea of where do you use nouns versus verbs in command-line arguments to things you’re running in the terminal. Because I was a grumpy Unix sysadmin for the first part of my career—because there’s no other kind of Unix sysadmin—and going down that path was great. Okay, everything I’m interacting with is basically a text file piped together to do different things. And it took a while for me to realize, you know, maybe—just spitballing here—there’s a better way to convey information than a wall of text, sometimes. Blasphemy.

And no, no, it turns out that just because it’s hard using the tools I’m used to doesn’t mean that’s the best way to convey information. And even now, these days, I’m spending more time getting the color theme and the font choices and typeface choices of what I’m doing in the terminal to represent something that’s a bit more aesthetically pleasing. Does it actually account for anything? I don’t know, but it feels better and there’s almost a Feng Shui element of it. Similar to work in a—

Marie: Yes.

Corey: Clean office versus a messy one.

Marie: A hundred percent. I think that’s kind of how I think of an approach. I am much more likely to get the things done. If, when I come in and I open Notion, it’s like, “Here’s what’s on today, Marie.” And it’s like speaking nicely to me, there’s little positive messages, there’s beautiful imagery.

It just makes me feel good when I’m starting my day. And knowing that how I feel is going to very much influence what I’m likely to accomplish in the day, again, I’m constantly tricking myself into getting [laugh] more excited and amped up about what’s on the schedule for the day. So, I really liked that about it. It feels beautiful to me.

Corey: I’m going to have to take another look at it at some point. I think that there’s a lot of interesting directions to go into on this. I also have the privilege of having known you for a little while, back when you were more or less just getting started. One of the things that you said at the time that absolutely resonated with me was the idea of, wait, you mean build a business around teaching people how to use Notion? Like an info product or a training approach?

And a lot of your concerns are the ones that I’ve harbored for a while, too, which is the idea of there’s a proliferation of info products in technical and other spaces, and an awful lot of them—without naming any names or talking in any particular direction—are not the highest quality. People are building these courses while learning the thing themselves. And when they tell stories about it, it’s all about, “And this is how I’m making money quickly.” I don’t find that admirable; I don’t necessarily want to learn how to do a thing from someone who does not have themselves at least a decent understanding themselves of what they’re working on so they can address questions that go a bit off into the weeds. And so mu—again, knowing how to do a thing and knowing how to teach a thing are orthogonal concepts. And very often a lot of these info products are being created by people who don’t really know how to do either, as best I can tell.

Marie: Yes. So, I think you’ve nailed a point to that, knowing a thing deeply and then knowing how to teach that thing really well are two totally different skills. And I definitely bumped up against that myself. I’m like, I know, Notion inside and out. Like, you know, name something, I can make it, I can optimize it, I can, you know, build a system out of thin air really fast, no problem. I’m a problem solver that way.

But to teach someone else how to do that requires very different skills. And I knew [laugh] as I was starting to teach people stuff, I’m like, “You could do this. You could do that.” And I’m like kind of bouncing around and I’m all over the place because I’m so excited about the possibilities. But wait a second.

Beginners that are just learning how to use Notion don’t need to know every frickin’ possible way that you could use it. So, knowing that instructional design, curriculum design is a whole other skill, and I care about student results, it’s like, this is a gap that I have, and I want to be an excellent teacher. It matters to me. I actually do want to become a better teacher. I want to have higher quality YouTube videos, I want to make sure that I’m not losing people along the way.

I don’t just care about making a shit ton of money with an info product; I care about peoples’ experience and kind of having that, I don’t know, that prestige element. Like, that’s something that does matter in terms of producing quality products. So, I hired experts to help me do that because again, it’s a not necessarily a strength of mine. So, I think I hired three different people in the course of six months to various consultants and people who understand learning design and that sort of thing. And I think that’s something a lot of info product creators. They think of it as just packaging a blog and selling it, right?

It’s different. When you’re teaching a course, for example, your formatting matters, how you display information matters, how you design activities matters. What separates a course from a passive income product or blog, right? We need to think about those things, and I think a lot of people are just like, what’s the quickest, you know, buck that I can make on these products and just kind of turn them out. And I don’t think every course creator has maybe done the extra legwork to really understand what makes students actually follow through and complete a course. It’s hard. It’s really hard.

Corey: And these are also very different products. There’s what you are teaching, which is here’s how to contextualize these things and how to build a system around it. There’s another offering out there that would be something that would also be very compelling from my perspective where, cool, I appreciate the understanding and the deep systems design approach that goes into this. Can I just give you a brain dump of all the problems that I have with this? You go away and build a system that accounts for all of that.

And again, it’s the outcome that I care about. There’s this belief that oh we want consultants to build by the hour and work hard. No. I don’t care. If you listen to this, nod and do the great customer service thing, the Zoom call, and just like, “Okay, that’s template number three with three one-line changes. Done. Now, we’re going to sit on it for a week so it looks hard.”

Which we’ve all got that as consultants in the early days. And then you turn that around because it’s the outcome that I really care about. But that’s a different business, that is a different revenue model, that is different—

Marie: Yes.

Corey: That is not nearly so much a one-to-many, like an info product. That is a one-to-one or one-to-few.

Marie: And I did that for the whole first year that the course was being developed and was out there. I was simultaneously consulting with people one-on-one all the time, with teams, with individuals. So, I’m learning about what are all those common challenges that keep popping up over and over again? What are the unique challenges? What are the common ones?

And in my experience, what I bumped up against is people think they want to just pay someone to solve that, but then when you give someone a very fleshed out, organized system that they didn’t participate in the building, it’s a lot harder to get somebody to use it, to plug into a ready-made system. So, in our experience, there’s a sort of back and forth. It has to happen in tandem; we do it over time. And you know, in my partner’s case, Ben does consulting with companies as well, so he’ll meet with them on a weekly basis and working with the different members of the team. So, there is some element of we built you a thing. Let’s have you use it, notice where there’s gaps, friction, whatever, because it’s not a one-and-done process.

It’s not like, “You gave me all the info. We’re good to go.” It’s not until people are using it that you’re like, “Oh, okay, that’s close, but I’m finding myself doing this, or avoiding this, or clicking around too much.” And so, to me, it’s a really organic process. But that’s not something that I’m as keen to do. And maybe it’s because I did it for, like, two years and kind of burnt out on it. I’m like, “I’m done. Like, I’d rather teach folks to do it themselves.” But so a partner does the consulting; I’m doing more of the teaching.

Corey: That’s what happened to an awful lot of our consulting work here at The Duckbill Group where it was exciting and fun for me for years, and at some point it turned into, I am interested in teaching how to do this a little bit more and systematizing it because I’m starting to get bored with aspects of it. And I was thinking, “Well, do I build a course?” It’s, “Well, no. As it turns out that if you have the right starting point, I can hire people who I can teach how to do AWS bill analysis if they have the right starting point.” And it turns out that a lot of those people—read as all of them—are going to be way better at doing the systemic deep-dive across the board, rather than just finding the things that they find personally interesting and significant, and then, “Well, there you go. I did a consulting engagement.” And the output is basically three bullet points scrawled on the back of an envelope.

Yeah, turns out that that’s not quite the level of professionalism clients expect. Great, so our product is better, we’re getting better insight into it, and I get to scratch my itch of teaching people how to do things internally without becoming a critical path blocker.

Marie: Yeah, absolutely.

Corey: I mean, I have shitposting to get back to. Come on.

Marie: Yeah exactly. [laugh]. The important things. Love it.

Corey: I really want to thank you for taking so much time to speak with me about all of these things. If people want to learn more—

Marie: Absolutely.

Corey: —where’s the best place to find you?

Marie: Yeah, you can find me at mariepoulin.com is where my personal blog, or weareokidoki.com, or notionmastery.com. You can also catch me on Twitter.

Corey: And we will put links to—

Marie: That’s where I am most active. Yeah.

Corey: Oh, of course. And all the links wind up going into the [show notes 00:37:42], as always. Thank you so much for your time. I appreciate it.

Marie: Thanks for having me, Corey. It was awesome.

Corey: Marie Poulin, CEO of Oki Doki. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice—and if it’s on the YouTubes smash the like and subscribe buttons—whereas if you’ve hated this podcast episode, great, same thing, five-star review on whatever platform, smash the two buttons, but also leave an insulting comment and then turn that comment into an info product that you wind up selling to a whole bunch of people primarily to boost your own Twitter threads about how successful you are as a creator.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

Full Description / Show Notes

  • Guillermo talks about how he came to work at OCI and what it was like helping to pioneer Oracle’s cloud product (1:40)
  • Corey and Guillermo discuss the challenges and realities of multi-cloud (6:00)
  • Corey asks about OCI’s dedicated region approach (8:27)
  • Guillermo discusses the problem of awareness (12:40)
  • Corey and Guillermo talk cloud providers and cloud migration (14:40)
  • Guillermo shares about how OCI’s cost and customer service is unique among cloud providers (16:56)
  • Corey and Guillermo talk about IoT services and 5G (23:58)

About Guillermo Ruiz

Guillermo Ruiz gets into trouble more often than he would like. During his career Guillermo has seen many horror stories while building data centers worldwide. In 2007 he dreamed with space-based internet and direct routing between satellites, but he could only reach “the Cloud”. And there he is, helping customer build their business in someone else servers since 2011.

Beware of his sense of humor...If you ever see him in a tech event, run, he will get you in problems.

Links:

  • Twitter: https://twitter.com/IaaSgeek, https://twitter.com/OracleStartup
  • LinkedIn: https://www.linkedin.com/in/gruizesteban/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’ve been meaning to get a number of folks on this show for a while and today is absolutely one of those episodes. I’m joined by Guillermo Ruiz who is the Director of OCI Developer Evangelism, slash the Director of Oracle for Startups. Guillermo, thank you for joining me, and is Oracle for Startups an oxymoron because it kind of feels like it in some weird way, in the fullness of time.

Guillermo: [laugh]. Thanks, Corey. It’s a pleasure being in your show.

Corey: Well, thank you. I enjoy having you here. I’ve been trying to get you on for a while. I’m glad I finally wore you down.

Guillermo: [laugh]. Thanks. As I said, well, startup, I think, is the future of the industry, so it’s a fundamental piece of our building blocks for the next generation of services.

Corey: I have to say that I know that you folks at Oracle Cloud have been a recurring sponsor of the show. Thank you for that, incidentally. This is not a promoted guest episode. I invited you on because I wanted to talk to you about these things, which means that I can say more or less whatever I damn well want. And my experience with Oracle Cloud has been one of constantly being surprised since I started using it a few years ago, long before I was even taking sponsorships for this show. It was, “Oh, Oracle has a cloud. This ought to be rich.”

And I started kicking the tires on it and I came away consistently and repeatedly impressed by the technical qualities the platform has. The always-free tier has a model of cloud economics that great. I have a sizable VM running there and have for years and it’s never charged me a dime. Your data egress fees aren’t, you know, a 10th of what a lot of the other cloud providers are charging, also known as, you know, you’re charging in the bounds of reality; good for that. And the platform continues to—although it is different from other cloud providers, in some respects, it continues to impress.

Honestly, I keep saying one of the worst problems that has is the word Oracle at the front of it because Oracle has a 40-some-odd-year history of big enterprise systems, being stodgy, being difficult to work with, all the things you don’t generally tend to think of in terms of cloud. It really is a head turn. How did that happen? And how did you get dragged into the mess?

Guillermo: Well, this came, like, back in five, six years ago, when they started building this whole thing, they picked people that were used to build cloud services from different hyperscalers. They dropped them into a single box in Seattle. And it’s like, “Guys, knowing what you know, how you would build the next generation cloud platform?” And the guys came up with OCI, which was a second generation. And when I got hired by Oracle, they showed me the first one, that classic.

It was totally bullshit. It was like, “Guys, there’s no key differentiator with what’s there in the market.” I didn’t even know Oracle had a cloud, and I’ve been in this space since late-2010. And I had to sign, like, a bunch of NDAs a lot of papers, and they show me what they were cooking in the oven, and oh my gosh, when I saw that SDN out of the box directly in the physical network, CPUs assign, it was [BLEEP] [unintelligible 00:03:45]. It was, like, bare metal. I saw that the future was there. And I think that they built the right solution, so I joined the company to help them leverage the cloud platform.

Corey: The thing that continually surprises me is that, “Oh, we have a cloud.” It has a real, “Hello fellow kids,” energy. Yes, yeah, so does IBM; we’ve seen how that played out. But the more I use it, the more impressed I am. Early on in the serverless function days, you folks more or less acquired Iron.io, and you were streets ahead as far as a lot of the event-driven serverless function style of thing tended to go.

And one of the challenges that I see in the story that’s being told about Oracle Cloud is, the big enterprise customer wins. These are the typical global Fortune 2000s, who have been around for, you know—which is weird for those of us in San Francisco, but apparently, these companies have been around longer than 18 months and they’ve built for platforms that are not the latest model MacBook Pro running the current version of Chrome. What is that? What is that legacy piece of garbage? What does it do? It’s like, “Oh, it does about $4 billion a quarter so maybe show some respect.”

It’s the idea of companies that are doing real-world things, and they absolutely have cloud power. Problems and needs that are being met by a variety of different companies. It’s easy to look at that narrative and overlook the fact that you could come up with some ridiculous Twitter for Pets-style business idea and build it on top of Oracle Cloud and I would not, at this point, call that a poor decision. I’m not even sure how it got there, and I wish that story was being told a little bit better. Given that you are a developer evangelist focusing specifically on startups and run that org, how do you see it?

Guillermo: Well, the thing here is, you mentioned, you know, about Oracle, many startup doesn’t even know we have a cloud provider. So, many of the question comes is like, how we can help on your business. It’s more on the experience, you know, what are the challenges, the gaps, and we go in and identify and try to use our cloud. And even though if I’m not able to fill that gap, that’s why we have this partnership with Microsoft. It’s the first time to cloud providers connect both clouds directly without no third party in between, router to router.

It’s like, let’s leverage the best of these clouds together. I’m a truly believer of multi-cloud. Non-single cloud is perfect. We are evolving, we’re getting better, we are adding services. I don’t want to get to 500 services like other guys do. It’s like, just have a set of things that really works and works really, really well.

Corey: Until you have 40 distinct managed database services and 80 ways to run containers, are you’re really a full cloud provider? I mean, there’s always that question that, at some point, the database Java, the future is going to have to be disambiguating between all the different managed database services on a per workload basis, and that job sounds terrible. I can’t let the multi-cloud advocacy pass unchallenged here because I’m often misunderstood on this, and if I don’t say something, I will get emails, and nobody wants that. I think that the idea of building a workload with the idea that it can flow seamlessly between cloud providers is a ridiculous fantasy that basically no one achieves. The number of workloads that can do that are very small.

That said, the idea of independent workloads living on different cloud providers as is the best fit for placement for those is not just a good idea, it is the—whether it’s a good idea or not as irrelevant because that’s the reality in which we all live now. That is the world we have to deal with.

Guillermo: If you want distributed system, obviously you need to have multiple cloud providers in your strategy. How you federate things—if you go down to the Kubernetes side, how you federate multi-clusters and stuff, that’s a challenge out there where people have. But you mentioned that having multiple apps and things, we have customers that they’ve been running Google Cloud, for example, and we build [unintelligible 00:07:40] that cloud service out there. And the thing is that when they run the network throughput and the performance test, they were like, “Damn, this is even better than what I have in my data center.” It’s like, “Guys, because we are room by room.” It’s here is Google, here it’s Oracle; we land in the same data center, we can provide better connectivity that what you even have.

So, that kind of perception is not well seen in some customers because they realize that they’re two separate clouds, but the reality is that most of us have our infrastructure in the same providers.

Corey: It’s kind of interesting, just to look at the way that the industry is misunderstanding a lot of these things. When you folks came out with your cloud at customer initiatives—the one that jumps out to my mind is the dedicated region approach—a lot of people started making fun of that because, “What is this nonsense? You’re saying that you can deploy a region of your cloud on site at the customer with all of the cloud services? That’s ridiculous. You folks don’t understand cloud.”

My rejoinder to that is people saying that don’t understand customers. You take a look at for example… AWS has their Outpost which is a rack or racks with a subset of services in them. And that, from their perspective, as best I can tell, solves the real problem that customers have, which is running virtual machines on-premises that do not somehow charge an hourly cost back to AWS—I digress—but it does bring a lot of those services closer to customers. You bring all of your services closer to customers and the fact that is a feasible thing is intensely appealing to a wide variety of customer types. Rather than waiting for you to build a region in a certain geographic area that conforms with some regulatory data requirement, “Well, cool, we can ship some racks. Does that work for you?” It really is a game-changer in a whole bunch of respects and I don’t think that the industry is paying close enough attention to just how valuable that is.

Guillermo: Indeed. I’ve been at least hearing since 2010 that next year is the boom; now everybody will move into the cloud. It has been 12 years and still 75% of customers doesn’t have their critical workloads in the cloud. They have developer environments, some little production stuff, but the core business is still relying in the data center. If I come and say, “Hey, what if I build this behind your firewall?”

And it’s not just that you have the whole thing. I’m removing all your operational expenses. Now, you don’t need to think about hardware refresh, upgrade staff, just focus on your business. I think when we came up with a dedicated region, it was awesome. It was one of the best thing I’ve seen their Outpost is a great solution, to be honest, but if you lose the one connectivity, the control plane is still in the cloud.

In our site, you have the control plane inside your data center so you can still operate and manage your services, even if there is an outage on your one site. One of the common questions we find on that area is, like, “Damn, this is great, but we would like to have a smaller size of this dedicated region.” Well, stay tuned because maybe we come with smaller versions of our dedicated regions so you guys can go and deploy whatever you need there.

Corey: It turns out that, in the fullness of time, I like this computer but I want it to be smaller is generally a need that gets met super well. One thing that I’ve looked into recently has been the evolution of companies, in the fullness of time—which this is what completely renders me a terrible analyst in any traditional sense; I think more than one or two quarters ahead, and I look at these things—the average tenure of a company in the S&P 500 index is 21 years or so. Which means that if we take a look at what’s going on 20 years or so from now in the 2040s, roughly half—give or take—of the constituency of the S&P 500 may very well not have been founded yet. So, when someone goes out and founds a company tomorrow as an idea that they’re kicking around, let’s be clear, with a couple of very distinct exceptions, they’re going to build it on Cloud. There’s a lot of reasons to do that until you hit certain inflection points.

So, this idea that, oh, we’re going to rent a rack, and we’re going to go build some nonsense, and yadda, yadda, yadda. It’s just, it’s a fantasy. So, the question that I see for a lot of companies is the longtail legacy where if I take that startup and found it tomorrow and drive it all the way toward being a multinational, at what point did they become a customer for whatever these companies are selling? A lot of the big E enterprise vendors don’t have a story for that, which tells me long-term, they have problems. Looking increasingly at what Oracle Cloud is doing, I have to level with you, I viewed Oracle as being very much in that slow-eroding dinosaur perspective until I started using the platform in some depth. I am increasingly of the mind that there’s a bright future. I’m just not sure that has sunk into the industry’s level of awareness these days.

Guillermo: Yeah, I can agree with you in that sense. Mainly, I think we need to work on that awareness side. Because for example, if I go back to the other products we have in the company, you know, like the database, what the database team has done—and I’m not a database guy—and it’s like, “Guys, even being an infrastructure guy, customers doesn’t care about infrastructure. They just want to run their service, that it doesn’t fail, you don’t have a disruption; let me evolve my business.” But even though they came with this converged database, I was really impressed that you can do everything in a single-engine rather than having multiple database implemented. Now, you can use the MongoDB APIs.

It’s like, this is the key of success. When you remove the learning curve and the frictions for people to use your services. I’m a [unintelligible 00:13:23] guy and I always say, “Guys, click, click, click. In three clicks, I should have my service up and running.” I think that the world is moving so fast and we have so much information today, that’s just 24 hours a day that I have to grab the right information. I don’t have time to go and start learning something from scratch and taking a course of six months because results needs to be done in the next few weeks.

Corey: One thing that I think that really reinforces this is—so as I mentioned before, I have a free tier account with you folks, have for years, whenever I log into the thing, I’m presented with the default dashboard view, which recommends a bunch of quickstarts. And none of the quickstarts that you folks are recommending to me involve step one, migrate your legacy data center or mainframe into the cloud. It’s all stuff like using analytics to predict things with AI services, it’s about observability, it’s about governance of deploy a landing zone as you build these things out. Here’s how to do a low-code app using Apex—which is awesome, let’s be clear here—and even then launching resources is all about things that you would tend to expect of launch database, create a stack, spin up some VMs, et cetera. And that’s about as far as it goes toward a legacy way of thinking.

It is very clear that there is a story here, but it seems that all the cloud providers these days are chasing the migration story. But I have to say that with a few notable exceptions, the way that those companies move to cloud, it always starts off by looking like an extension of their data center. Which is fine. In that phase, they are improving their data center environment at the expense of being particularly cloudy, but I don’t think that is necessarily an adoption model that puts any of these platforms—Oracle Cloud included—in their best light.

Guillermo: Yeah, well, people was laughing to us, when we released Layer 2 in the network in the cloud. They were like, “Guys, you’re taking the legacy to the cloud. It’s like, you’re lifting the shit and putting the shit up there.” Is like, “Guys, there are customers that cannot refactor and do anything there. They need to still run Layer 2 there. Why not giving people options?”

That’s my question is, like, there’s no right answers to the cloud. You just need to ensure that you have the right options for people that they can choose and build their strategy around that.

Corey: This has been a global problem where so many of these services get built and launched from all of the vendors that it becomes very unclear as a customer, is this thing for me or not? And honestly, sometimes one of the best ways to figure that out is to all right, what does it cost because that, it turns out, is going to tell me an awful lot. When it comes to the price tag of millions of dollars a year, this is probably not for my tiny startup. Whereas when it comes to a, oh, it’s in the always free tier or it winds up costing pennies per hour, okay, this is absolutely something I want to wind up exploring and seeing what happens. And it becomes a really polished experience across the board.

I also will say this is your generation two cloud—Gen 2, not to be confused with Gentoo, the Linux distribution for people with way more time on their hands than they have sense—and what I find interesting about it is, unlike a lot of the—please don’t take this the wrong way—late-comers to cloud compared to the last 15 years of experience of Amazon being out in front of everyone, you didn’t just look at what other providers have done and implement the exact same models, the exact same approaches to things. You’ve clearly gone in your own direction and that’s leading to some really interesting places.

Guillermo: Yeah, I think that doing what others are doing, you just follow the chain, no? That will never position you as a top number one out there. Being number one so many years in the cloud space as other cloud providers, sometimes you lose the perception of how to treat and speak to customers you know? It’s like, “I’m the number one. Who cares if this guy is coming with me or not?” I think that there’s more on the empathy side on how we treat customers and how we try to work and solve.

For example, in the startup team, we find a lot of people that hasn’t have infrastructure teams. We put for free our architects that will give you your GitHub or your GitLab account and we’ll build the Terraform modules and give that for you. It’s like now you can reuse it, spin up, modify whatever you want. Trying to make life easier for people so they can adopt and leverage their business in the cloud side, you know?

[midroll 00:14:45]

Corey: There’s so much that we folks get right. Honestly, one of the best things that recommends this is the always free tier does exactly what it says on the tin. Yeah, sure. I don’t get to use every edge case service that you’ve built across the board, but I’ve also had this thing since 2019, and never had to pay a penny for any of it, whereas recently—as we’re recording this, it was a week or two ago—that I saw someone wondering what happened to their AWS account because over the past week, suddenly they went from not using SageMaker to being charged $270,000 on SageMaker. And it’s… yeah, that’s not the kind of thing that is going to endear the platform to frickin’ anyone.

And I can’t believe I’m saying this, but the thing says Oracle on the front of it and I’m recommending it because it doesn’t wind up surprising you with a bill later. It feels like I’ve woken up in bizarro world. But it’s great.

Guillermo: Yep. I think that’s one of the clever things we’ve done on that side. We’ve built a very robust platform, really cool services. But it’s key on how people can start learning and testing the flavors of your cloud. But not only what you have in the fleet here, you have also the Ampere instances.

We’re moving into a more sustainable world, and I think that having, like, the ARM architectures in the cloud and providing that on the free space of people can just go and develop on top, I think that was one of the great things we’ve done in the last year-and-a-half, something like that. Definitely a full fan of a free tier.

Corey: You also, working over in the Developer Evangelist slash advocacy side of the world—devrelopers, as I tend to call it much to the irritation of basically everyone who works in developer relations—one of the things that I think is a challenge for you is that when I wind up trying to do something ridiculous—I don’t know maybe it’s a URL shortener; maybe it is build a small app that does something that’s fairly generic—with a lot of the other platforms. There’s a universe of blog posts out there, “Here’s how I did it on this platform,” and then it’s more or less you go to GitHub—or gif-UB, and I have mispronounced that too—and click the button and I wind up getting a deploy, whereas in things that are rapidly emerging with the Oracle Cloud space, it feels like, on some level, I wind up getting to be a bit of a trailblazer and figure some of these things out myself. That is diminishing. I’m starting to see more and more content around this stuff. I have to assume that is at least partially due to your organization’s work.

Guillermo: Oh, yeah, but things have changed. For example, we used to have our GitHub repository just as a software release, and we push to have that as a content management, you know, it’s like, I always say that give—let people steal the code. You just put the example that will come with other ideas, other extensions, plug-in connectors, but you need to have something where you can start. So, we created this DevRel Quickstart that now is managed by the new DevRel organization where we try to put those examples. So, you just can go and put it.

I’ve been working with the community on building, like, a content aggregator of how people is using our technology. We used to have ocigeek.com, that was a website with more than 1000 blog and, like, 500 visits a day looking after what other people were doing, but unfortunately, we had to, because of… the amount of X reasons we have to pull it off.

But we want to come with something like that. I think that information should be available. I don’t want people to think when it comes to my cloud is like, “Oh, how you use this product?” It’s like no, guys how I can build with Angular, React the content management system? You will do it in my cloud because that example I’m doing, but I want you to learn the basics and the context of running Python and doing other things there rather than go into oh, no, this is something specific to me. No, no, that will never work.

Corey: That was the big problem I found with doing a lot of the serverless stuff in years past where my first Lambda application took me two weeks to build because I’m terrible at programming. And now it takes me ten minutes to build because I’m terrible at programming and don’t know what tests are. But the problem I ran into for that first one was, what is the integration format? What is the event structure? How do I wind up accessing that?

What is the thing that I’m integrating with expecting because, “Mmm, that’s not it; try again,” is a terrible error message. And so, much of it felt like it was the undifferentiated gluing things together. The only way to make that stuff work is good documentation and numerous examples that come at the problem from a bunch of different ways. And increasingly, Oracle’s documentation is great.

Guillermo: Yeah, well, in my view, for example, you have the Three-Tier Oracle. We should have a catalog of 100 things that you can do in the free tier, even though when I propose some of the articles, I was even talking about VMware, and people was like, “[unintelligible 00:22:34], you cannot deploy VMware.” It’s like, “Yeah, but I can connect my [crosstalk 00:22:39]—”

Corey: Well, not with that attitude.

Guillermo: Yeah. And I was like, “Yeah, but I can connect to the cloud and just use it as a backup place where I can put my image and my stuff. Now, you’re connecting to things: VMware with free tier.” Stuff like that. There are multiple things that you can do.

And just having three blocks is things that you can do in the free tier, then having developer architectures. Show me how you can deploy an architecture directly from the command line, how I can run my DevOps service without going to the console, just purely using SDKs and stuff like that. And give me the option of how people is working and expanding that content and things there. If you put those three blocks together, I think you’re done on how people can adopt and leverage your cloud. It’s like, I want to learn; I don’t want to know the basics of I don’t know, it’s—I’m not a database guy, so I don’t understand those things and I don’t want to go into details.

I just they just need a database to store my profiles and my stuff so I can pick that and do computer vision. How I can pick and say, “Hey, I’m speaking with Corey Quinn and I have a drone flying here, he recommends your face and give me your background from all the different profiles.” That’s the kind of solutions I want to build. But I don’t want to be an expert on those areas.

Corey: Because with all the pictures of me with my mouth open, you wouldn’t be able to under—it would make no sense of me until I make that pose. There’s method to—

Guillermo: [laugh].

Corey: —my insane madness over here.

Guillermo: [laugh] [unintelligible 00:23:58].

Corey: Yeah. But yeah, there’s a lot of value as you move up the stack on these things. There’s also something to be said, as well, for a direction that you folks have been moving in recently, that I—let me be fair here—I think it’s clown shoes because I tend to think in terms of software because I have more or less the hardware destruction bunny level of aura when it comes to being near expensive things. And I look around the world and I don’t have a whole lot of problems that I can legally solve with an army of robots.

But there are customers who very much do. And that’s why we see sort of the twin linking of things like IoT services and 5G, which when I first started seeing cloud providers talking about this, I thought was Looney Tunes. And you folks are getting into it too, so, “Oh, great. The hype wound up affecting you too.” And the thing that changed my mind was not anything cloud providers have to say—because let’s be clear, everyone has an agenda they’re trying to push for—but who doesn’t have an agenda is the customers talking about these things and the neat things that they’re able to achieve with it, at which point I stopped making fun, I shut up and listen in the hopes that I might learn something. How have you seen that whole 5G slash IoT slash internet of Nonsense space evolving?

Guillermo: That’s the future. That’s what we’re going to see in the next five years. I run some innovation sessions with a lot of customers and one of the main components I speak about is this area. With 5G, the number of IoT devices will exponentially grow. That means that you’re going to have more data points, more data volume out there.

How can you provide the real value, how you can classify, index, and provide the right information in just 24 hours, that’s what people is looking. Things needs to be instant. If you say to the kids today, they cannot watch a football match, 90 minutes. If you don’t get the answer in ten, they move to the next thing. That’s how this society is moving [unintelligible 00:25:50].

Having all these solutions from a data perspective, and I think that Oracle has a great advantage in that space because we’ve been doing that for 43 years, right? It’s like, how we do the abstraction? How I can pick all that information and provide added value? We build the robot as a service. I can configure it from my browser, any robot anywhere in the world.

And I can do it in Python, Java. I can [unintelligible 00:26:14] applications. Two weeks ago, we were testing on connecting IoT devices and flashing the firmware. And it was working. And this is something that we didn’t do it alone. We did it with a startup.

The guys came and had a sandbox already there, is like, “let’s enable this on [unintelligible 00:26:28]. Let’s start working together.” Now, I can go to my customers and provide them a solution that is like, hey, let’s connect Boston Dynamics, or [unintelligible 00:26:37] Robotics. Let’s start doing those things and take the benefits of using Oracle’s AI and ML services. Pick that, let’s do computer vision, natural language processing.

Now, you’re connecting what I say, an end-to-end solution that provides real value for customers. Connected cars, we turn our car into a wallet. I can go and pay on the petrol station without leaving my car. If I’m taking the kids to takeaway, I can just pay these kind of things is like, “Whoa, this is really cool.” But what if I [laugh] get that information for your insurance company.

Next year, Corey, you will pay double because you’re a crazy driver. And we know how you drive in the car because we have all that information in place. That’s how the things will roll out in the next five to ten years. And [unintelligible 00:27:24] healthcare. We build something for emergencies that if you have a car crash, they have the guys that go and attend can have your blood type and some information about your car, where to cut the chassis and stuff when you get prisoner inside.

And I got people saying, “Oh gee, GDPR because we are in Europe.” It’s like, “Guys, if I’m going to die, I don’t care if they have my information.” That’s the point where people really need to balance the whole thing, right? Obviously, we protect the information and the whole thing, but in those situations is like hey, there’s so many things we can do. There are countless opportunities out there.

Corey: The way that I square that circle personally has always been it’s about informed consent, when if people are given a choice, then an awful lot of those objections that people have seemed to melt away. Provided, of course, that is an actual choice and it’s not one of those, “Well, you can either choose to”—quote-unquote—“Choose to do this, or you can pay $9,000 a month extra.” Which is, that’s not really a choice. But as long as there’s a reasonable way to get informed consent, I think that people don’t particularly mind, I think it’s when they wind up feeling that they have been spied upon without their knowledge, that’s when everything tends to blow up. It turns out, if you tell people in advance what you’re going to do with their information, they’re a lot less upset. And I don’t mean burying it deep and the terms and conditions.

Guillermo: And that’s a good example. We run a demo with one of our customers showing them how dangerous the public information you have out there. You usually sign and click and give rights to everybody. We found in Stack Overflow, there was a user that you just have the username there, nothing else. And we build a platform with six terabytes of information grabbing from Stack Overflow, LinkedIn, Twitter, and many other social media channels, and we show how we identify that this guy was living in Bangalore in India and was working for a specific company out there.

So, people was like, “Damn, just having that name, you end up knowing that?” It’s like there’s so much information out there of value. And we’ve seen other companies doing that illegally in other places, you know, Cambridge Analytics and things like that. But that’s the risk of giving your information for free out there.

Corey: It’s always a matter of trade-offs. There is no one-size-fits-all solution and honestly, if there were it feels like we wouldn’t have cloud providers; we would just have the turnkey solution that gives the same thing that everyone needs and calls it good. I dream of such a day, but it turns out that customers are different, people are different, and there’s no escaping that.

Guillermo: [laugh]. Well, you mentioned dreamer; I dream direct routing between satellites, and look where I am; I’m just in the cloud, one step lower. [laugh].

Corey: You know, bit by bit, we’re going to get there one way or another, for an altitude perspective. I really want to thank you for taking so much time to speak with me today. If people want to learn more, where’s the right place to find you?

Guillermo: Well, I have the @IaaSgeek Twitter account, and you can find me on LinkedIn gruizesteban there. Just people wants to talk about anything there, I’m open to any kind of conversation. Just feel free to reach out. And it was a pleasure finally meeting you, in person. Not—well in person; through a camera, at least being in the show with you.

Corey: Other than on the other side of a Twitter feed. No, I hear you.

Guillermo: [laugh].

Corey: We will, of course, put links to all of that in the [show notes 00:30:43]. Thank you so much for your time. I really do appreciate it.

Guillermo: Thanks very much. So, you soon.

Corey: Guillermo Ruiz, Director of OCI Developer Evangelism. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an insulting comment, to which I will respond with a surprise $270,000 bill.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About swyx

swyx has worked on React and serverless JavaScript at Two Sigma, Netlify and AWS, and now serves as Head of Developer Experience at Temporal.io. He has started and run communities for hundreds of thousands of developers, like Svelte Society, /r/reactjs, and the React TypeScript Cheatsheet. His nontechnical writing was recently published in the Coding Career Handbook for Junior to Senior developers.

Links Referenced:

  • “Learning Gears” blog post: https://www.swyx.io/learning-gears
  • The Coding Career Handbook: https://learninpublic.org
  • Personal Website: https://swyx.io
  • Twitter: https://twitter.com/swyx

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friend EnterpriseDB. EnterpriseDB has been powering enterprise applications with PostgreSQL for 15 years. And now EnterpriseDB has you covered wherever you deploy PostgreSQL on-premises, private cloud, and they just announced a fully-managed service on AWS and Azure called BigAnimal, all one word. Don’t leave managing your database to your cloud vendor because they’re too busy launching another half-dozen managed databases to focus on any one of them that they didn’t build themselves. Instead, work with the experts over at EnterpriseDB. They can save you time and money, they can even help you migrate legacy applications—including Oracle—to the cloud. To learn more, try BigAnimal for free. Go to biganimal.com/snark, and tell them Corey sent you.

Corey: Let’s face it, on-call firefighting at 2am is stressful! So there’s good news and there’s bad news. The bad news is that you probably can’t prevent incidents from happening, but the good news is that incident.io makes incidents less stressful and a lot more valuable. incident.io is a Slack-native incident management platform that allows you to automate incident processes, focus on fixing the issues and learn from incident insights to improve site reliability and fix your vulnerabilities. Try incident.io, recover faster and sleep more.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Some folks are really easy to introduce when I have them on the show because, “My name is, insert name here. I built thing X, and my job is Y at company Z.” Then we have people like today’s guest.

swyx is currently—and recently—the head of developer experience at Airbyte, but he’s also been so much more than that in so many different capacities that you’re very difficult to describe. First off, thank you for joining me. And secondly, what’s the deal with you?

swyx: [laugh]. I have professional ADD, just like you. Thanks for having me, Corey. I’m a—

Corey: It works out.

swyx: a big fan. Longtime listener, first time caller. Love saying that. [laugh].

Corey: You have done a lot of stuff. You have a business and finance background, which… okay, guilty; it’s probably why I feel some sense of affinity for a lot of your work. And then you went into some interesting directions. You were working on React and serverless YahvehScript—which is, of course, how I insist on pronouncing it—at Two Sigma, Netlify, AWS—a subject near and dear to my heart—and most recently temporal.io.

And now you’re at Airbyte. So, you’ve been focusing on a lot of, I won’t say the same things, but your area of emphasis has definitely consistently rhymed with itself. What is it that drives you?

swyx: So, I have been recently asking myself a lot of this question because I had to interview to get my new role. And when you have multiple offers—because the job market is very hot for DevRel managers—you have to really think about it. And so, what I like to say is: number one, working with great people; number two, working on great products; number three, making a lot of money.

Corey: There’s entire school of thought that, “Oh, that’s gauche. You shouldn’t mention trying to make money.” Like, “Why do you want to work here because I want to make money.” It’s always true—

swyx: [crosstalk 00:03:46]—

Corey: —and for some reason, we’re supposed to pretend otherwise. I have a lot of respect for people who can cut to the chase on that. It’s always been something that has driven me nuts about the advice that we give a new folks to the industry and peop—and even students figuring out their career path of, “Oh, do something you love and the money will follow.” Well, that’s not necessarily true. There are ways to pivot something you’d love into something lucrative and there are ways to wind up more or less borderline starving to death. And again, I’m not saying money is everything, but for a number of us, it’s hard to get to where we want to be without it.

swyx: Yeah, yeah. I think I’ve been cast with the kind of judgmental label of being very financially motivated—that’s what people have called me—for simply talking about it. And I’m like, “No. You know, it’s number three on my priority list.” Like, I will leave positions where I have a lot of money on the table because I don’t enjoy the people or the products, but having it up there and talking openly about it somehow makes you [laugh] makes you sort of greedy or something. And I don’t think that’s right. I tried to set an example for the people that I talk to or people who follow me.

Corey: One of the things I’ve always appreciated about, I guess, your online presence, which has remained remarkably consistent as you’ve been working through a bunch of different, I guess, stages of life and your career, is you have always talked in significant depth about an area of tech that I am relatively… well, relatively crap at, let’s be perfectly honest. And that is the wide world of most things front-end. Every time I see a take about someone saying, “Oh, front-end is junior or front-end is somehow less than,” I’d like to know what the hell it is they know because every time I try and work with it, I wind up more confused than I was when I started. And what I really appreciate is that you have always normalized the fact that this stuff is hard. As of the time that we’re recording this a day or so ago, you had a fantastic tweet thread about a friend of yours spun up a Create React App and imported the library to fetch from an endpoint and immediately got stuck. And then you pasted this ridiculous error message.

He’s a senior staff engineer, ex-Google, ex-Twitter; he can solve complex distributed systems problems and unable to fetch from a REST endpoint without JavaScript specialist help. And I talk about this a lot in other contexts, where the reason I care so much about developer experience is that a bad developer experience does not lead people to the conclusion of, “Oh, this is a bad interface.” It leads people to the conclusion, “Oh, I’m bad at this and I didn’t realize it.” No. I still fall into that trap myself.

I was under the impression that there was just this magic stuff that JS people know. And your tweet did so much to help normalize from my perspective, the fact that no, no, this is very challenging. I recently went on a Go exploration. Now, I’m starting to get into JavaScript slash TypeScript, which I think are the same thing but I’m not entirely certain of that. Like, oh, well, one of them is statically typed, or strongly typed. It’s like, “Well, I have a loud mechanical keyboard. Everything I do is typing strongly, so what’s your point?”

And even then we’re talking past each other in these things. I don’t understand a lot of the ecosystem that you live your career in, but I have always had a tremendous and abiding respect for your ability to make it accessible, understandable, and I guess for lack of a better term, to send the elevator back down.

swyx: Oh, I definitely think about that strongly, especially that last bit. I think it’s a form of personal growth. So, I think a lot of people, when they talk about this sending the elevator back down, they do it as a form of charity, like I’m giving back to the community. But honestly, you actually learn a lot by trying to explain it to others because that’s the only way that you truly know if you’ve learned something. And if you ever get anything wrong, you’ll—people will never let you forget it because it is the internet and people will crawl over broken glass to remind you that you’re wrong.

And once you’ve got it wrong, you will—you know, you’ve been so embarrassed that you’ll never forget it. So, I think it’s just a really good way to learn in public. And that’s kind of the motto that I’m kind of known for. Yeah, we can take the direction anywhere you want to go in JavaScript land. Happy to talk about it all day. [laugh].

Corey: Well, I want to start by something you just said where you’re doing the learning in public thing. And something I’ve noticed is that there are really two positions you can take—in the general sense—when you set out to make a bit of a reputation for yourself in a particular technical space. You can either do the, “I’m a beginner here, same as the rest of you, and I’m learning in public,” or you can position yourself as something of an expert. And there are drawbacks and advantages to both. I think that if you don’t look as wildly over-represented as I do, both of them are more fraught in different ways, where it’s, “Oh, you’re learning in public. Ah, look at the new person, she’s dumb.”

Or if you’re presenting yourself as an expert, you get nibbled to death by ducks on a lot of the deep technical nuances and well, actually’ed to death. And my position has always been and this is going to be a radical concept for some folks, is that I’m genuinely honest. I tend to learn in public about the things that I don’t know, but the things that I am something of a subject matter expert in—like, I don’t know, cloud billing—I don’t think that false modesty necessarily serves me particularly well. It’s yeah, I know exactly what I’m talking about here. Pretending otherwise it’s just being disingenuous.

swyx: I try to think of it as having different gears of learning in public. So, I’ve called this “Learning Gears” in a previous blog post of mine, where you try to fit your mode of learning to the terrain that you’re on, your domain expertise, and you should never over-represent the amount that you know because I think people are very rightly upset when there are a lot of people—let’s say on Twitter, or YouTube, or Udemy even—who present themselves as experts who are actually—they just read the docs the previous night. So, you should try not to over-represent your expertise.

But at the same time, don’t let your imposter syndrome stop you from sharing what you are currently learning and taking corrections when you’re wrong. And I think that’s the tricky balance to get which is constantly trying to put yourself out there while accepting that you might be wrong and not getting offended when or personally attacked when someone corrects you, inevitably. And sometimes people will—especially if you have a lot of followers, people will try to say—you know, someone of your following—you know, it’s—I kind of call this follower shaming, like, you should act, uh—invulnerable, or run every tweet through committee before you tweet after a certain sort of following size. So, I try to not do that and try to balance responsibility with authenticity.

Corey: I think that there’s something incredibly important about that, where there’s this idea that you either become invulnerable and get defensive and you yell at people, and down that path lies disaster because, believe it or not, we all get it wrong from time to time, and doubling down and doubling down and doubling down again, suddenly, you’re on an island all by yourself and no one respectable is going to be able to get there to help you. And the other side of it is going too far in the other direction, where you implicitly take any form of criticism whatsoever as being de facto correct. And I think that both paths don’t lead to super great places. I think it’s a matter of finding our own voices and doing a little bit of work as far as the validity of accepting a given piece of feedback goes. But other than that, I’m a big fan of being able to just more or less be as authentic as possible.

And I get that I live in a very privileged position where I have paths open to me that are not open to most folks. But in many respects so to you are one of the—easily—first five people I would think of if someone said, “Hey if I need to learn JavaScript for someone, who should I talk to first?” You’re on that list. And you’ve done a lot of things in this area, but you’ve never—you alluded to it a few minutes ago, but I’m going to call it out a little more pointedly—without naming names, let’s be clear—and that you’re never presented as a grifter, which is sort of the best way I can think of it of, “Well, I just learned this new technology stack yesterday and now I’m writing a book that I’m going to sell to people on how to be an expert at this thing.” And I want to be clear, this is very distinct from gatekeeping because I think that, “Oh, well, you have to be at least this much of an expert—” No, but I think that holding yourself out as I’m going to write a book on how to be proud of how to become a software engineer.

Okay, you were a software engineer for six months, and more to the point, knowing how to do a thing and knowing how to teach a thing are orthogonal skill sets, and I think that is not well understood. If I ever write a book or put something—or some sort of info product out there, I’m going to have to be very careful not to fall into that trap because I don’t want to pretend to be an expert in things that I’m not. I barely think I’m an expert in things that I provable am.

swyx: there are many ways to answer that. So, I have been accused a couple of times of that. And it’s never fun, but also, if you defend yourself well, you can actually turn a critic into a fan, which I love doing.

Corey: Mm-hm.

swyx: [laugh].

Corey: Oh yes.

swyx: what I fall back to, so I have a side interest in philosophy, based on one of my high school teachers giving us, like, a lecture in philosophy. I love him, he changed my life. [Lino Barnard 00:13:20], in case—in the off chance that he’s listening. So, there’s a theory of knowledge of, like, how do you know what you know, right? And if you can base your knowledge on truth—facts and not opinions, then people are arguing with the facts and not the opinions.

And so, getting as close to ground truth as possible and having certainty in your collection of facts, I think is the basis of not arguing based on identity of, like, “Okay, I have ten years experience; you have two years experience. I am more correct than you in every single opinion.” That’s also not, like, the best way to engage in the battlefield of ideas. It’s more about, do you have the right amount of evidence to support the conclusions that you’re trying to make? And oftentimes, I think, you know, that is the basis, if you don’t have that ability.

Another thing that I’ve also done is to collect the opinions of others who have more expertise and present them and curate them in a way that I think adds value without taking away from the individual original sources. So, I think there’s a very academic way [laugh] you can kind of approach this, but that defends your intellectual integrity while helping you learn faster than the typical learning rate. Which is kind of something I do think about a lot, which is, you know, why do we judge people by the number of years experience? It’s because that’s usually the only metric that we have available that is quantifiable. Everything else is kind of fuzzy.

But I definitely think that, you know, better algorithms for learning let you progress much faster than the median rate, and I think people who apply themselves can really get up there in terms of the speed of learning with that. So, I spend a lot of time thinking about this stuff. [laugh].

Corey: It's a hard thing to solve for. There’s no way around it. It’s, what is it that people should be focusing on? How should they be internalizing these things? I think a lot of it starts to with an awareness, even if not in public, just to yourself of, “I would like advice on some random topic.” Do you really? Are you actually looking for advice or are you looking—

swyx: right.

Corey: For validation? Because those are not the same thing, and you are likely to respond very differently when you receive advice, depending on which side of that you’re coming from.

swyx: Yeah. And so, one way to do that is to lay out both sides, to actually demonstrate what you’re split on, and ask for feedback on specific tiebreakers that would help your decision swing one way or another. Yeah, I mean, there are definitely people who ask questions that are just engagement bait or just looking for validation. And while you can’t really fix that, I think it’s futile to try to change others’ behavior online. You just have to be the best version of yourself you can be. [laugh].

Corey: DoorDash had a problem. As their cloud-native environment scaled and developers delivered new features, their monitoring system kept breaking down. In an organization where data is used to make better decisions about technology and about the business, losing observability means the entire company loses their competitive edge. With Chronosphere, DoorDash is no longer losing visibility into their applications suite. The key? Chronosphere is an open source compatible, scalable, and reliable observability solution that gives the observability lead at DoorDash business, confidence, and peace of mind. Read the full success story at snark.cloud/chronosphere. That's snark.cloud slash C-H-R-O-N-O-S-P-H-E-R-E.

Corey: So, you wrote a book that is available at learninpublic.org, called The Coding Career Handbook. And to be clear, I have not read this myself because at this point, if I start reading a book like that, and you know, the employees that I have see me reading a book like that, they’re going to have some serious questions about where this company is going to be going soon. But scrolling through the site and the social proof, the testimonials from various people who have read it, more or less read like a who’s-who of people that I respect, who have been on this show themselves.

Emma Bostian is fantastic at explaining a lot of these things. Forrest Brazeal is consistently a source to me of professional envy. I wish I had half his musical talent; my God. And your going down—it explains, more or less, the things that a lot of folks people are all expected to know but no one teaches them about every career stage, ranging from newcomer to the industry to senior. And there’s a lot that—there’s a lot of gatekeeping around this and I don’t even know that it’s intentional, but it has to do with the idea that people assume that folks, quote-unquote, “Just know” the answer to some things.

Oh, people should just know how to handle a technical interview, despite the fact that the skill set is completely orthogonal to the day-to-day work you’ll be doing. People should just know how to handle a performance review, or should just know how to negotiate for a raise, or should just know how to figure out is this technology that I’m working on no longer the direction the industry is going in, and eventually I’m going to wind up, more or less, waiting for the phone to ring because there’s only three companies in the world left who use it. Like, how do you keep—how do you pay attention to what’s going on around you? And it’s the missing manual that I really wish that people would have pointed out to me back when I was getting started. Would have made life a lot easier.

swyx: Oh, wow. That’s high praise. I actually didn’t know we’re going to be talking about the book that much. What I will say is—

Corey: That’s the problem with doing too much. You never know what people have found out about you and what they’re going to say when they drag you on to a podcast.

swyx: got you, got you. Okay. I know, I know, I know where this is going. Okay. So, one thing that I really definitely believe is that—and this happened to me in my first job as well, which is most people get the mentors that they’re assigned at work, and sometimes you have a bad roll the dice. [laugh].

And you’re supposed to pick up all the stuff they don’t teach you in school at work or among your friend group, and sometimes you just don’t have the right network at work or among your friend group to tell you the right things to help you progress your career. And I think a lot of this advice is written down in maybe some Hacker News posts, some Reddit posts, some Twitter posts, and there’s not really a place you to send people to point to, that consolidates that advice, particularly focused at the junior to senior stage, which is the stage that I went through before writing the book. And so, I think that basically what I was going for is targeting the biggest gap that I saw, which is, there a lot of interview prep type books like Crack the Coding Career, which is kind of—Crack the Coding Interview, which is kind of the book title that I was going after. But once you got the job, no one really tells you what to do after you got that first job. And how do you level up to the senior that everyone wants to hire, right? There’s—

Corey: “Well, I’ve mastered cracking the coding interview. Now, I’m really trying to wrap my head around the problem of cracking the showing up at work on time in the morning.” Like, the baseline stuff. And I had so many challenges with that early in my career. Not specifically punctuality, but just the baseline expectation that it’s just assumed that by the time you’re in the workplace earning a certain amount of money, it’s just assumed that you have—because in any other field, you would—you have several years of experience in the workplace and know how these things should play out.

No, the reason that I’m sometimes considered useful as far as giving great advice on career advancement and the rest is not because I’m some wizard from the future, it’s because I screwed it all up myself and got censured and fired and rejected for all of it. And it’s, yeah, I’m not smart enough to learn from other people’s mistakes; I got to make them myself. So, there’s something to be said for turning your own missteps into guidance so that the next person coming up has an easier time than you did. And that is a theme that, from what I have seen, runs through basically everything that you do.

swyx: I tried to do a lot of research, for sure. And so, one way to—you know, I—hopefully, I try not to make mistakes that others have learned, have made, so I tried to pick from, I think I include 1500 quotes and sources and blog posts and tweets to build up that level of expertise all in one place. So hopefully, it gives people something to bootstrap your experience off of. So, you’re obviously going to make some mistakes on your own, but at least you have the ability to learn from others, and I think this is my—you know, I’m very proud of the work that I did. And I think people have really appreciated it.

Because it’s a very long book, and nobody reads books these days, so what am I doing [laugh] writing a book? I think it’s only the people that really need this kind of advice, that they find themselves not having the right mentorship that reach out to me. And, you know, it’s good enough to support a steady stream of sales. But more importantly, like, you know, I am able to mentor them at various levels from read my book, to read my free tweets, to read the free chapters, or join the pay community where we have weekly sessions going through every chapter and I give feedback on what people are doing. Sometimes I’ve helped people negotiate their jobs and get that bump up to senior staff—senior engineer, and I think more than doubled their salary, which was very personal proud moment for me.

But yeah, anyway, I think basically, it’s kind of like a third place between the family and work that you could go to the talk about career stuff. And I feel like, you know, maybe people are not that open on Twitter, but maybe they can be open in a small community like ours.

Corey: There’s a lot to be said for a sense of professional safety and personal safety around being—having those communities. I mean, mine, when I was coming up was the freenode IRC network. And that was great; it’s pseudo-anonymous, but again, I was Corey and network staff at the time, which was odd, but it was great to be able to reach out and figure out am I thinking about this the wrong way, just getting guidance. And sure, there are some channels that basically thrived on insulting people. I admittedly was really into that back in the early-two-thousand-nothings.

And, like, it was always fun to go to the Debian channel. It’s like, “Yeah, can you explain to me how to do this or should I just go screw myself in advance?” Yeah, it’s always the second one. Like, community is a hard thing to get right and it took me a while to realize this isn’t the energy I want in the world. I like being able to help people come up and learn different things.

I’m curious, given your focus on learning in public and effectively teaching folks as well as becoming a better engineer yourself along the way, you’ve been focusing for a while now on management. Tell me more about that.

swyx: I wouldn’t say it’s been, actually, a while. Started dabbling in it with the Temporal job, and then now fully in it with Airbyte.

Corey: You have to know, it has been pandemic time; it has stood still. Anything is—

swyx: exactly.

Corey: —a while it given that these are the interminable—this is the decade of Zoom meetings.

swyx: [laugh]. I’ll say I have about a year-and-a-half of it. And I’m interested in it partially because I’ve really been enjoying the mentoring side with the coding career community. And also, I think, some of the more effective parts of what I do have to be achieved in the planning stages with getting the right resources rather than doing the individual contributor work. And so, I’m interested in that.

I’m very wary of the fact that I don’t love meetings myself. Meetings are a means to an end for me and meetings are most of the job in management time. So, I think for what’s important to me there, it is that we get stuff done. And we do whatever it takes to own the outcomes that we want to achieve and try to manage people’s—try to not screw up people’s careers along the way. [laugh]. Better put, I want people to be proud of what they get done with me by the time they’re done with me. [laugh].

Corey: So, I know you’ve talked to me about this very briefly, but I don’t know that as of the time of this recording, you’ve made any significant public statements about it. You are now over at Airbytes, which I confess is a company I had not heard of before. What do y’all do over there?

swyx: [laugh]. “What is it we do here?” So Airbyte—

Corey: Exactly. Consultants want to know.

swyx: Airbyte’s a data integration company, which means different things based on your background. So, a lot of the data engineering patterns in, sort of, the modern data stack is extracting from multiple sources and loading everything into a data warehouse like a Snowflake or a Redshift, and then performing analysis with tools like dbt or business intelligence tools out there. We like to use MetaBase, but there’s a whole there’s a whole bunch of these stacks and they’re all sort of advancing at different rates of progress. And what Airbyte would really like to own is the data integration part, the part where you load a bunch of sources, every data source in the world.

What really drew me to this was two things. One, I really liked the vision of data freedom, which is, you have—you know, as—when you run a company, like, a typical company, I think at Temporal, we had, like, 100, different, like, you know, small little SaaS vendors, all of them vying to be the sources of truth for their thing, or a system of record for the thing. Like, you know, Salesforce wants to be a source of truth for customers, and Google Analytics want to be source of truth for website traffic, and so on and so forth. Like, and it’s really hard to do analysis across all of them unless you dump all of them in one place.

So one, is the mission of data freedom really resonates with me. Like, your data should be put in put somewhere where you can actually make something out of it, and step one is getting it into a format in a place that is amenable for analysis. And data warehouse pattern has really taken hold of the data engineering discipline. And I find, I think that’s a multi-decade trend that I can really get behind. That’s the first thing.

Corey: I will say that historically, I’m bad at data. All jokes about using DNS as a database aside, one of the reasons behind that is when you work on stateless things like web servers and you blow trunks and one of them, oops. We all laugh, we take an outage, so maybe we’re not laughing that hard, but we can reprovision web servers and things are mostly fine. With data and that going away, there are serious problems that could theoretically pose existential risk to the business. Now, I was a sysadmin and a, at least mediocre one, which means that after the first time I lost data, I was diligent about doing backups.

Even now, the data work that we do have deep analysis on our customers’ AWS bills, which doesn’t sound like a big data problem, but I assure you it is, becomes something where, “Okay, step one. We don’t operate on it in place.” We copy it into our own secured environment and then we begin the manipulations. We also have backups installed on these things so that in the event that I accidentally the data, it doesn’t wind up causing horrifying problems for our customers. And lastly, I wind up also—this is going to surprise people—I might have securing the access to that data by not permitting writes.

Turns out it’s really hard—though apparently not impossible—to delete data with read-only calls.

swyx: [crosstalk 00:28:12].

Corey: It tends to be something of just building guardrails against myself. But the data structures, the understanding the analysis of certain things, I would have gotten into Go way sooner than I did if the introduction to Go tutorial on how to use it wasn’t just a bunch of math problems talking about this is how you do it. And great, but here in the year of our lord 2022, I mostly want a programming language to smack a couple of JSON objects together and ideally come out with something resembling an answer. I’m not doing a whole lot of, you know, calculating prime numbers in the course of my week. And that is something that took a while for me to realize that, no, no, it’s just another example of not being a great way of explaining something that otherwise could be incredibly accessible to folks who have real problems like this.

I think the entire field right now of machine learning and the big data side of the universe struggles with this. It’s, “Oh, yeah. If you have all your data, that’s going to absolutely change the world for you.” “Cool. Can you explain how?” “No. Not effectively anyway.” Like, “Well, thanks for wasting everyone’s time. It’s appreciated.”

swyx: Yeah, startup is sitting on a mountain of data that they don’t use and I think everyone kind of feels guilty about it because everyone who is, like, a speaker, they’re always talking about, like, “Oh, we used our data to inform this presidential campaign and look at how amazing we are.” And then you listen to the podcasts where the data scientists, you know, talk amongst themselves and they’re like, “Yeah, it’s bullshit.” Like, [laugh], “We’re making it up as we go along, just like everyone else.” But, you know, I definitely think, like, some of the better engineering practices are arising under this. And it’s professionalizing just like front-end professionalized maybe ten years ago, DevOps professionalized also, roughly in that timeframe, I think data is emerging as a field that is just a standalone discipline with its own tooling and potentially a lot of money running through it, especially if you look at the Snowflake ecosystem.

So, that’s why I’m interested in it. You know, I will say there’s also—I talked to you about the sort of API replication use case, but also there’s database replication, which is kind of like the big use case, which, for example, if you have a transactional sort of SQL database and you want to replicate that to an analytical database for queries, that’s a very common one. So, I think basically data mobility from place to place, reshaping it and transferring it in as flexible manner as possible, I think, is the mission, and I think there’s a lot of tooling that starts from there and builds up with it. So, Airbyte integrates pretty well with Airflow, Dexter, and all the other orchestration tools, and then, you know, you can use dbt, and everything else in that data stack to run with it. So, I just really liked that composition of tools because basically when I was a hedge fund analyst, we were doing the ETL job without knowing the name for it or having any tooling for it.

I just ran a Python script manually on a cron job and whenever it failed, I would have to get up in the middle of night to go kick it again. It’s, [laugh] it was that bad in 2014, ’15. So, I really feel the pain. And, you know, the more data that we have to play around with, the more analysis we can do.

Corey: I’m looking forward to seeing what becomes of this field as folks like you get further and further into it. And by, “Well, what do you mean, folks like me?” Well, I’m glad you asked, or we’re about to as I put words in your mouth. I will tell you. People who have a demonstrated ability not just to understand the technology—which is hard—but then have this almost unicorn gift of being able to articulate and explain it to folks who do not have that level of technical depth in a way that is both accessible and inviting. And that is no small thing.

If you were to ask me to draw a big circle around all the stuff that you’ve done in your career and define it, that’s how I would do it. You are a storyteller who is conversant with the relevant elements of the story in a first-person perspective. Which is probably a really wordy way to put it. We should get a storyteller to workshop that, but you see the point.

swyx: I try to call it, like, accessibly smart. So, it’s a balance that you want to make, where you don’t want to talk down to your audience because I think there are a lot of educators out there who very much stay at the basics and never leave that. You want to be slightly aspirational and slightly—like, push people to the bounds of their knowledge, but then not to go too far and be inaccessible. And that’s my sort of polite way of saying that I dumb things down as service. [laugh].

Corey: But I like that approach. The term dumbing it down is never a phrase to use, as it turns out, when you’re explaining it to someone. It’s like, “Let me dumb that down for you.” It’s like, yeah, I always find the best way to teach someone is to first reach them and get their attention. I use humor, but instead we’re going to just insult them. That’ll get their attention all right.

swyx: No. Yeah. It does offend some people who insist on precision and jargon. And I’m quite against that, but it’s a constant fight because obviously there is a place at time for jargon.

Corey: “Can you explain it to me using completely different words?” If the answer is, “No,” the question then is, “Do you actually understand it or are you just repeating it by rote?”

swyx: right.

Corey: There’s—people learn in different ways and reaching them is important. [sigh].

swyx: Exactly.

Corey: Yeah. I really want to thank you for being so generous with your time. If people want to learn more about all the various things you’re up to, where’s the best place to find you?

swyx: Sure, they can find me at my website swyx.io, or I’m mostly on Twitter at @swyx.

Corey: And we will include links to both of those in the [show notes 00:33:37]. Thank you so much for your time. I really appreciate it.

swyx: Thanks so much for having me, Corey. It was a blast.

Corey: swyx, head of developer experience at Airbyte, and oh, so much more. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice or if it’s on the YouTubes thumbs up and subscribe, whereas if you’ve hated this podcast, same thing, five-star review wherever you want, hit the buttons on the YouTubes, but also leaving insulting comment that is hawking your book: Why this Episode was Terrible that you’re now selling as a legitimate subject matter expert in this space.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Alyssa

Alyssa Miller, Business Information Security Officer (BISO) for S&P Global, is the global executive leader for cyber security across the Ratings division, connecting corporate security objectives to business initiatives. She blends a unique mix of technical expertise and executive presence to bridge the gap that can often form between security practitioners and business leaders. Her goal is to change how security professionals of all levels work with our non-security partners throughout the business.

A life-long hacker, Alyssa has a passion for technology and security. She bought her first computer herself at age 12 and quickly learned techniques for hacking modem communications and software. Her serendipitous career journey began as a software developer which enabled her to pivot into security roles. Beginning as a penetration tester, her last 16 years have seen her grow as a security leader with experience across a variety of organizations. She regularly advocates for improved security practices and shares her research with business leaders and industry audiences through her international public speaking engagements, online content, and other media appearances.

Links Referenced:

  • Cybersecurity Career Guide: https://alyssa.link/book
  • A-L-Y-S-S-A dot link—L-I-N-K slash book: https://alyssa.link/book
  • Twitter: https://twitter.com/AlyssaM_InfoSec
  • alyssasec.com: https://alyssasec.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Vultr. Optimized cloud compute plans have landed at Vultr to deliver lightning-fast processing power, courtesy of third-gen AMD EPYC processors without the IO or hardware limitations of a traditional multi-tenant cloud server. Starting at just 28 bucks a month, users can deploy general-purpose, CPU, memory, or storage optimized cloud instances in more than 20 locations across five continents. Without looking, I know that once again, Antarctica has gotten the short end of the stick. Launch your Vultr optimized compute instance in 60 seconds or less on your choice of included operating systems, or bring your own. It’s time to ditch convoluted and unpredictable giant tech company billing practices and say goodbye to noisy neighbors and egregious egress forever. Vultr delivers the power of the cloud with none of the bloat. Screaming in the Cloud listeners can try Vultr for free today with a $150 in credit when they visit getvultr.com/screaming. That’s G-E-T-V-U-L-T-R dot com slash screaming. My thanks to them for sponsoring this ridiculous podcast.

Corey: This episode is sponsored in part by Honeycomb. When production is running slow, it’s hard to know where problems originate. Is it your application code, users, or the underlying systems? I’ve got five bucks on DNS, personally. Why scroll through endless dashboards while dealing with alert floods, going from tool to tool to tool that you employ, guessing at which puzzle pieces matter? Context switching and tool sprawl are slowly killing both your team and your business. You should care more about one of those than the other; which one is up to you. Drop the separate pillars and enter a world of getting one unified understanding of the one thing driving your business: production. With Honeycomb, you guess less and know more. Try it for free at honeycomb.io/screaminginthecloud. Observability: it’s more than just hipster monitoring.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. One of the problems that many folks experience in the course of their career, regardless of what direction they’re in, is the curse of high expectations. And there’s no escaping for that. Think about CISOs for example, the C-I-S-O, the Chief Information Security Officer.

It’s generally a C-level role. Well, what’s better than a C in the academic world? That’s right, a B. My guest today is breaking that mold. Alyssa Miller is the BISO—B-I-S-O—at S&P Global. Alyssa, thank you for joining me to suffer my slings and arrows—

Alyssa: [laugh].

Corey: —as we go through a conversation that is certain to be no less ridiculous than it has begun to be already.

Alyssa: I mean, I’m good with ridiculous, but thanks for having me on. This is awesome. I’m really excited to be here.

Corey: Great. What the heck’s BISO?

Alyssa: [laugh]. I never get that question. So, this is—

Corey: “No one’s ever asked me that before.” [crosstalk 00:03:38]—

Alyssa: Right?

Corey: —the same thing as, “Do you know you’re really tall?” “No, you’re kidding.” Same type of story. But I wasn’t clear. That means I’m really the only person left wondering.

Alyssa: Exactly. I mean, I wrote a whole blog on it the day I got the job, right? So, Business Information Security Officer, Basically what it means is I am like the CISO but for my division, the Ratings Division at S&P Global. So, I lead our cyber security efforts within that division, work closely with our information security teams, our corporate IT teams, whatever, but I don’t report to them; I report into the business line.

I’m in the divisional CTO’s org structure. And so, I’m the one bridging that gap between that business side where hey, we make all the money and that corporate InfoSec side where hey, we’re trying to protect all the things, and there’s usually that little bit of a gap where they don’t always connect. That’s me building the bridge across that.

Corey: Someone who speaks both security and business is honestly in a bit of rare supply these days. I mean, when I started my Thursday newsletter podcast nonsense Last Week in AWS: Security, the problem I kept smacking into was everything I saw was on one side of that divide or the other. There was the folks who have the word security in their job title, and there tends to be this hidden language of corporate speak. It’s a dialect I don’t fully understand. And then you have the community side of actual security practitioners who are doing amazing work, but also have a cultural problem that more or less distills down to being an awful lot of shitheads in them there waters.

And I wanted something that was neither of those and also wasn’t vendor captured, which is why I decided to start storytelling in that space. But increasingly, I’m seeing that there’s a significant problem with people who are able to contextualize security in the context of business. Because if you’re secure enough, you can stop all work from ever happening, whereas if you’re pure business side and only care about feature velocity and the rest, like, “Well, what happens if we get breached?” It’s, “Oh, don’t worry, I have my resume up to date.” Not the most reassuring answer to give people. You have to be able to figure out where that line lies. And it seems like that figuring out where that line is, is more or less your entire stock-in-trade.

Alyssa: Oh absolutely, yeah. I mean, I can remember my earliest days as a developer, my cynical attitude towards security myself was, you know, their Utopia would be an impenetrable room full of servers that have no connections to anything, right? Like that would be wildly secure, yet completely useless. And so yeah, then I got into security and now I was one of them. And, you know, it’s one of those things, you sit in, say a board meeting sometime and you listen to a CISO, a typical CISO talk to the board, and they just don’t get it.

Like, there’s so much, “Hey, we’re implementing this technology and we’re doing this thing, and here’s our vulnerability counts, and here’s how many are overdue.” And none of that means anything. I mean, I actually had a board member ask me once, “What is a CISO?” I kid you not. Like, that’s where they’re at.

Like, so don’t tell them what you’re doing, but tell them why connected back to, like, “Hey, the business needs this and this, and in order to do it, we’ve got to make sure it’s secure, so we’re going to implement these couple of things. And here’s the roadmap of how we get from where we are right now to where we need to be so they can launch that new service or product,” or whatever the hell it is that they’re going to do.

Corey: It feels like security is right up there with accounting, in the sense of fields of endeavor where you don’t want someone with too much personality involved. Because if the CISO’s sitting there talking to the board, it’s like, “So, what do you do here, exactly?” And the answer is the honest, “Hey, remember last month how we were in The New York Times for that giant data breach?” And they do a split take, “No, no, I don’t.” “Exactly. You’re welcome.” On some level, it is kind of honest, but it also does not instill confidence when you’re that cavalier with the description of what it is you do here.

Alyssa: Oh there’s—

Corey: At least there’s some corners. I prefer—

Alyssa: —there’s so much—

Corey: —places where that goes over well, but that’s me.

Alyssa: Yeah. But there’s so much of that too, right? Like, here’s the one I love. “Well, you know, it’s not if you get breached, it’s when. Oh, by the way, give me millions and millions of dollars, so I can make sure we don’t get breached.”

But wait, you just told me we’re going to get breached no matter what we do. [laugh]. We do that in security. Like, and then you wonder why they don’t give you funding for the initiative. Like, “Hello?” You know?

And that’s the thing that gets me it’s like, can we just sit back and understand, like, how do you message to these people? Yeah I mean, you bring up the accounting thing; the funny thing is, at least all of them understand some level of accounting because most of them have MBAs and business degrees where they had to do some accounting. They didn’t go through cyber security in their MBA program.

So, one of my favorite questions on Twitter once was somebody asked me, you know, if I want to get into cyber security leadership, what is the one thing that I should focus on or what skills should I study? I said, “Go study MBA concepts.” Like, forget all the cyber security stuff. You probably have plenty of that technolog—go understand what they learn in MBA programs. And if you can start to speak that language, that’s going to pay dividends for bridging that gap.

Corey: So, you don’t look like the traditional slovenly computer geek showing up at those meetings who does not know how to sound as if they belong in the room. Like, it’s unfair, on some level, and I used to have bitter angst about that. Like, “Why should how I dress matter how people perceive me?” Yeah, in an absolute sense you’re absolutely right, however, I can talk about the way the world is or the way I wish it were and there has to be a bit of a divide there.

Alyssa: Oh, for sure. Yeah. I mean, you can’t deny that you have to be prepared for the audience you’re walking into. Now, I work in big conservative financial services on Wall Street. You know, and I had this conversation with a prominent member of our community when I started the job.

I’m like, “Boy, I guess I can’t really put stickers on my laptop. I’m going to have to get, you know, a protector or something to put stickers on.” Because the last thing I want to do is go into a boardroom with my laptop and whip out a bunch of hacker stickers on the backside of my laptop. Like, in a lot of spaces that will work, but you can’t really do that when you’re, you know, at, you know, the executive level and you’re in a conservative, financial [unintelligible 00:10:16]. It just, I would love to say they should deal with that, I should be able to have pink hair, and you know, face tattoos and everything else, but the reality is, yeah, I can do all that, but these are still human beings who are going to react to that.

And it’s the same when talking about cyber security, then. Like, I have to understand as a security practitioner that all they know about cyber security is it’s big and scary. It’s the thing that keeps them up at night. I’ve had board members tell me exactly that. And so, how do I make it a little less scary, or at least get them to have some confidence in me that I’ll, like, carry the shield in front of them and protect them. Like, that’s my job. That’s why I’m there.

Corey: When I was starting my consultancy five years ago, I was trying to make a choice between something in the security cloud direction or the cost cloud direction. And one of the things that absolutely tipped the balance for me was the fact that the AWS bill is very much a business-hours-only problem. No one calls me at two in the morning screaming their head off. Usually. But there’s a lot of alignment between those two directions in that you can spend all your time and energy fixing security issues and/or reducing the bill, but past a certain point, knock it off and go do the thing that your company is actually there to do.

And you want to be responsible to a point on those things, but you don’t want it to be the end-all-be-all because the logical outcome of all of that, if you keep going, is your company runs out of money and dies because you’re not going to either cost optimize or security optimize your business to its next milestone. And weighing those things is challenging. Now, too many people hear that and think, “See, I don’t have to worry about those things at all.” It’s, “Oh, you will sooner or later. I promise.”

Alyssa: So, here’s the fallacy in that. There is this assumption that everything we do in security is going to hamper the business in some way and so we have to temper that, right? Like, you’re not wrong. And we talked about before, right? You know, security in a traditional sense, like, we could do all of the puristic things and end up just, like, screeching the world to a halt.

But the reality is, we can do security in a way that actually grows the business, that actually creates revenue, or I should say enables the creation of revenue in that, you know, we can empower the business to do more things and to be more innovative by how we approach security in the organization. And that’s the big thing that we miss in security is, like, look, yes, we will always be a quote-unquote, “Cost center,” right? I mean, we in security don’t—unless you work for a security organization—we’re not getting revenue attributed to us, we’re not creating revenue. But we are enabling those people who can if we approach it right.

Corey: Well, the Red Team might if they go a little off-script, but that’s neither here nor there.

Alyssa: I—yeah, I mean, I’ve had that question. “Like, couldn’t we just sell resell our Red Team services?” No. No. That’s not our core [crosstalk 00:13:14]

Corey: Oh, I was going the other direction. Like, oh, we’re just going to start extorting other businesses because we got bored this week. I’m kidding. I’m kidding. Please don’t do an investigation, any law enforcement—

Alyssa: I was going to say, I think my [crosstalk 00:13:22]—

Corey: —folks that happen to be listening to this.

Alyssa: [crosstalk 00:13:24] is calling me right now. They’re want to know what I’m [laugh] talking about. But no—

Corey: They have some inquiries they would like you to assist them with and they’re not really asking.

Alyssa: Yeah, yeah, they’re good at that. No, I love them, though. They’re great. [laugh]. But no, seriously, like, I mean, we always think about it that way because—and then we wonder why do we have the reputation of, you know, the Department of No.

Well, because we kind of look at it that way ourselves; we don’t really look at, like how can we be a part of the answer? Like, when we look at, like, DevSecOps, for instance. Okay, I want to bring security into my pipeline. So, what do we say? “Oh, shared responsibility. That’s a DevOps thing.” So, that means security is everybody’s responsibility. Full stop.

Corey: Right. It’s a—

Alyssa: Well—

Corey: And there, I agree with you wholeheartedly. Cost is—

Alyssa: But—

Corey: —aligned with this. It has to be easier to do it the right way than to just go off half-baked and do it yourself off the blessed path. And that—

Alyssa: So there—

Corey: —means there’s that you cannot make it harder to do the right thing; you have to make it easier because you will not win against human psychology. Depending on someone when they’re done with an experiment to manually go in and turn things off. It will not happen. And my argument has been that security and cost are aligned constantly because the best way to secure something and save money on at the same time is to turn that shit off. You wouldn’t think it would be that simple, but yet here we are.

Alyssa: But see, here’s the thing. This is what kills me. It’s so arrogant of security people to look at it and say that right? Because shared responsibility means shared. Okay, that means we have responsibilities we’re going to share. Everybody is responsible for security, yes.

Our developers have responsibilities now that we have to take a share in as well, which is get that shit to production fast. Period. That is their goal. How fast can I pop user stories off the backlog and get them to deployment? My SRE is on the ops side. They’re, like, “We just got to keep that stuff running. That’s all we that’s our primary focus.”

So, the whole point of DevOps and DevSecOps was everybody’s responsible for every part of that, so if I’m bringing security into that message, I, as security, have to be responsible for site’s stability; I, in security, have to be responsible for efficient deployment and the speed of that pipeline. And that’s the part that we miss.

Corey: This episode is sponsored in parts by our friend EnterpriseDB. EnterpriseDB has been powering enterprise applications with PostgreSQL for 15 years. And now EnterpriseDB has you covered wherever you deploy PostgreSQL on-premises, private cloud, and they just announced a fully-managed service on AWS and Azure called BigAnimal, all one word. Don’t leave managing your database to your cloud vendor because they’re too busy launching another half-dozen managed databases to focus on any one of them that they didn’t build themselves. Instead, work with the experts over at EnterpriseDB. They can save you time and money, they can even help you migrate legacy applications—including Oracle—to the cloud. To learn more, try BigAnimal for free. Go to biganimal.com/snark, and tell them Corey sent you.

Corey: I think you might be the first person I’ve ever spoken to that has that particular take on the shared responsibility model. Normally, when I hear it, it’s on stage from an AWS employee doing a 45-minute song-and-dance about what the secured responsibility model is, and generally, that is interpreted as, “If you get breached, it’s your fault, not ours.”

Alyssa: [laugh].

Corey: Now, you can’t necessarily say it that directly to someone who has just suffered a security incident, which is why it takes 45 minutes and slides and diagrams and excel sheets and the rest. But that is what it fundamentally distills down to, and then you wind up pointing out security things that they’ve had that [unintelligible 00:17:11] security researchers have pointed out and they are very tight-lipped about those things. And it’s, “Oh, it’s not that you’re otherworldly good at security; it’s that you’re great at getting people to shut up.” You know, not me, for whatever reason because I’m noisy and obnoxious, but most people who actually care about not getting fired from their jobs, generally don’t want to go out there making big cloud companies look bad. Meanwhile, that’s kind of my entire brand.

Alyssa: I mean, it’s all about lines of liability, right?

Corey: Oh yeah.

Alyssa: I mean, where am I liable, where am I not? And yeah, well, if I tell you you’re responsible for security on all these things, and I can point to any part of that was part of the breach, well, hey, then it’s out of my hands. I’m not liable. I did what I said I would; you didn’t secure your stuff. Yeah, it’s—and I mean, and some of that is to be fair.

Like, I mean, okay, I’m going to host my stuff on your computer—the whole cloud is just somebody else’s computer model is still ultimately true—but, yeah, I mean, I’m expecting you to provide me a stable and secure environment and then I’m going to deploy stuff on it, and you are expecting me to deploy things that are stable and secure as well. And so, when they say shared model or shared responsibility model, but it—really if you listen to that message, it’s the exact opposite. They’re telling you why it’s a separate responsibility model. Here’s our responsibilities; here’s yours. Boom. It’s not about shared; it’s about separated.

Corey: One of the most formative, I guess, contributors to my worldview was 13 years ago, I went on a date and met someone lovely. We got married. We’ve been together ever since, and she’s an attorney. And it is been life-changing to understand a lot of that perspective, where it turns out when you’re dealing with legal, they are not—and everyone says, “Oh, and the lawyers insisted on these things.”

No, they didn’t. A lawyer’s entire role in a company is to identify risk, and then it is up to the business to make a decision around what is acceptable and what is not. If your lawyers ever insist on something, what that actually means in my experience is, you have said something profoundly ignorant that is one of those, like—that is—they’re doing the legal equivalent of slapping the gun out of the toddler’s hand of, “No, you cannot go and tweet that because you’ll go to prison,” level of ridiculous nonsense where it is, “That will violate the law.” Everything else is different shades of the same answer: it depends. Here’s what to consider.

Alyssa: Yes.

Corey: And then you choose—and the business chooses its own direction. So, when you have companies doing what appeared to be ridiculous things, like Oracle, for example, loves to begin every keynote with a disclaimer about how nothing they’re about to say is true, the lawyers didn’t insist on that—though they are the world’s largest law firm, Kirkland Ellison. But instead, it’s this entire story of given the risk and everything that we know about how we say things onstage and people gunning for us, yeah, we are going to [unintelligible 00:20:16] this disclaimer first. Most other tech companies do not do that exact thing, which I’ve got to say when you’re sitting in the audience ready to see the new hotness that’s about to get rolled out and it starts with a disclaimer, that is more or less corporate-speak for, “You are about to hear some bullshit,” in my experience.

Alyssa: [laugh]. Yes. I mean and that’s the thing, like, [clear throat], you know, we do deride legal teams a lot. And you know, I can find you plenty of security people who hate the fact that when you’re breached, who’s the first call you make? Well, it’s your legal team.

Why? Because they’re the ones who are going to do everything in their power to limit the amount that you can get sued on the back-end for anything that got exposed, that you know, didn’t meet service levels, whatever the heck else. And that all starts with legal privilege.

Corey: They’re reporting responsibilities. Guess who keeps up on what those regulatory requirements are? Spoiler, it’s probably not you, whoever’s listening to this, unless you’re an attorney because that is their entire job.

Alyssa: Yes, exactly. And, you know, work in a highly regulated environment—like mine—and you realize just how critical that is. Like, how do I know—I mean, there are times there’s this whole discussion of how do you determine if something is a material impact or not? I don’t want to be the one making that, and I’m glad I don’t have to make that decision. Like, I’ll tell you all the information, but yes, you lawyers, you compliance people, I want you to make the decision of if it’s a material impact or not because as much as I understand about the business, y’all know way more about that stuff than I do.

I can’t say. I can only say, “Look, this is what it impacted. This is the data that was impacted. These are the potential exposures that occurred here. Please take that information now and figure out what that means, and is there any materiality to that that now we have to report that to the street.”

Corey: Right, right. You can take my guesses on this or you can get it take an attorney’s. I am a loud, confident-sounding white guy. Attorneys are regulated professionals who carry malpractice insurance. If they give wrong advice that is wrong enough in these scenarios, they can be sanctioned for it; they can lose their license to practice law.

And there are challenges with the legal profession and how much of a gatekeeper the Bar Association is and the rest, but this is what it is [done 00:22:49] for itself. That is a regulated industry where they have continuing education requirements they need to certify in a test that certain things are true when they say it, whereas it turns out that I don’t usually get people even following up on a tweet that didn’t come true very often. There’s a different level of scrutiny, there’s a different level of professional bar it raises to, and it turns out that if you’re going to be legally held to account for things you say, yeah, turns out a lot of your answers to are going to be flavors of, “It depends.”

Alyssa: [laugh].

Corey: Imagine that.

Alyssa: Don’t we do that all the time? I mean, “How critical is this?” “Well, you know, it depends on what kind of data, it depends on who the attacker is. It depends.” Yeah, I mean, that’s our favorite word because no one wants to commit to an absolute, and nor should we, I mean, if we’re speaking in hyperbole and absolutes, boy, we’re doing all the things wrong in cyber.

We got to understand, like, hey, there is nuance here. That’s how you run—no business runs on absolutes and hyperbole. Well, maybe marketing sometimes, but that’s a whole other story.

Corey: Depends on if it’s done well or terribly.

Alyssa: [laugh]. Right. Exactly. “Hey, you can be unhackable. You can be breached-proof.” Oh, God.

Corey: Like, what’s your market strategy? We’re going to paint a big freaking target in the front of the building. Like, I still don’t know how Target the company was ever surprised by a data breach that they had when they have a frickin’ bullseye as their logo.

Alyssa: “Come get us.”

Corey: It’s, like, talk about poking the bear. But there we are.

Alyssa: [unintelligible 00:24:21] no. I mean, hey, [unintelligible 00:24:23] like that was so long ago.

Corey: It still casts a shadow.

Alyssa: I know.

Corey: People point to that as a great example of, like, “Well, what’s going to happen if we get breached?” It’s like, well look at Target because they wound up—like, their stock price a year later was above where it had been before and it seemed to have no lasting impact. Yeah, but they effectively replaced all of the execs, so you know, let’s have some self-interest going on here by named officers of the company. It’s, “Yeah, the company will be fine. Would you like to still be here what it is?”

Alyssa: And how many lawsuits do you think happened that you never heard about because they got settled before they were filed?

Corey: Oh, yes. There’s a whole world of that.

Alyssa: That’s what’s really interesting when people talk about, like, the cost of breach and stuff, it’s like, we don’t even know. We can’t know because there is so much of that. I mean, think about it, any organization that gets breached, the first thing they’re trying to do is keep as much of it out of the news as they can, and that includes the lawsuits. And so, you know, it’s like, all right, well, “Hey, let’s settle this before you ever file.”

Okay, good. No one will ever know about that. That will never show up anywhere. It is going to show up on a balance sheet anywhere, right? I mean, it’s there, but it’s buried in big categories of lots of other things, and how are you ever going to track that back without, you know, like, a full-on audit of all of their accounting for that year? Yeah, it’s—so I always kind of laugh when people start talking about that and they want to know, what’s the average cost of a breach. I’m like, “There’s no way to measure that. There is none.”

Corey: It’s not cheap, and the reputational damage gets annoying. I still give companies grief for these things all the time because it’s—again, the breach is often about information of mine that I did not consciously choose to give to you and the, “Oh, I’m going to blame a third-party process.” No, no, you can outsource work, but not responsibility. You can’t share that one.

Alyssa: Ah, third-party diligence, uh, that seems to be a thing. You know, I think we’re supposed to make sure our third parties are trustworthy and doing the right things too, right? I mean, it’s—

Corey: Best example I ever saw that was an article in the Wall Street Journal about the Pokemon company where they didn’t name the vendor, but they said they declined to do business with them in part based upon their lax security policy around S3 buckets. That is the first and so far only time I have had an S3 Bucket Responsibility Award engraved and sent to their security director. Usually, it’s the ignoble prize of the S3 Bucket Negligence Award, and there are oh so many of those.

Alyssa: Oh, and it’s hard, right? Because you’re standing—I mean, I’m in that position a lot, right? You know, you’re looking at a vendor and you’ve got the business saying, “God, we want to use this vendor. All their product is great.” And I’m sitting there saying, but, “Oh, my God, look at what they’re doing. It’s a mess. It’s horrible. How do I how do we get around this?”

And that’s where, you know, you just have to kind of—I wish I could say no more, but at the end of the day, I know what that does. That just—okay, well, we’ll go file an exception and we’ll use it anyway. So, maybe instead, we sit and work on how to do this, or maybe there is an alternative vendor, but let’s sort it out together. So yeah, I mean, I do applaud them. Like that’s great to, like, be able to look at a vendor and say, “No, we ain’t touching you because what you’re doing over there is nuts.” And I think we’re learning more and more how important that is, with a lot of the supply chain attacks.

Corey: Actually, I’m worried about having emailed you, you’re going to leak my email address when your inbox inevitably gets popped. Come on. It’s awful stuff.

Alyssa: Yeah, exactly. So, I mean, it’s we there’s—but like everything, it's a balance again, right? Like, how can we keep that business going and also make sure that their vendors—so that’s where it just comes down to, like, okay, let’s talk contracts now. So, now we’re back to legal.

Corey: We are. And if you talk to a lawyer and say, “I’m thinking about going to law school,” the answer is always the same. “No… don’t do it.” Making it clear that is apparently a terrible life and professional decision, which of course, brings us to your most recent terrible life and professional decision. As we record this, we are reportedly weeks away from you having a physical copy in your hands of a book.

And the segue there is because no one wants to write a book. Everyone wants to have written a book, but apparently—unless you start doing dodgy things and ghost-writing and exploiting people in the rest—one is a necessary prerequisite for the other. So, you’ve written a book. Tell me about it.

Alyssa: Oof, well, first of all, spot on. I mean, I think there are people who really do, like, enjoy the act of writing a book—

Corey: Oh, I don’t have the attention span to write a tweet. People say, “Oh, you should write a book, Corey,” which I think is code for them saying, “You should shut up and go away for 18 months.” Like, yeah, I wish.

Alyssa: Writing a book has been the most eye-opening experience of my life. And yeah, I’m not a hundred percent sure it’s one I’ll ever—I’ve joked with people already, like, I’ll probably—if I ever want another book, I’ll probably hire a ghostwriter. But no, I do have a book coming out: Cybersecurity Career Guide. You know, I looked at this cyber skills gap, blah, blah, blah, blah, blah, we hear about it, 4 million jobs are going to be left open.

Whatever, great. Well, then how come none of these college grads can get hired? Why is there this glut of people who are trying to start careers in cyber security and we can’t get them in?

Corey: We don’t have six months to train you, so we’re going to spend nine months trying to fill the role with someone experienced?

Alyssa: Exactly. So, 2020 I did a bunch of research into that because I’m like, I got to figure this out. Like, this is bizarre. How is this disconnect happening? I did some surveys. I did some interviews. I did some open-source research. Ended up doing a TED Talk based off of that—or TEDx Talk based off of that—and ultimately that led into this book. And so yeah, I mean, I just heard from the publisher yesterday, in fact that we’re, like, in that last stage before they kick it out to the printers, and then it’s like three weeks and I should have physical copies in my hands.

Corey: I will be getting one when it finally comes out. I have an almost, I believe, perfect track record of having bought every book that a guest on this show has written.

Alyssa: Well, I appreciate that.

Corey: Although, God help me if I ever have someone, like, “So, what have you done?” “I’ve written 80 books.” Like, “Well, thank you, Stephen King. I’m about to go to have a big—you’re going to see this number of the company revenue from orbit at this point with that many.” But yeah, it’s impressive having written a book. It’s—

Alyssa: I mean, for me, it’s the reward is already because there are a lot of people have—so my publisher does really cool thing they call it early acc—or electronic access program, and where there are people who bought the book almost a year ago now—which is kind of, I feel bad about that, but that’s as much my publisher as it is me—but where they bought it a year ago and they’ve been able to read the draft copy of the book as I’ve been finishing the book. And I’m already hearing from them, like, you know, I’m hearing from people who really found some value from it and who, you know, have been recommending it other people who are trying to start careers and whatever. And it’s like, that’s where the reward is, right?

Like, it was, it’s hell writing a book. It was ten times worse during Covid. You know, my publisher even confirmed that for me that, like, look, yeah, you know, authors around the globe are having problems right now because this is not a good environment conducive to writing. But, yeah, I mean, it’s rewarding to know that, like, all right, there’s going to be this thing out there, that, you know, these pages that I wrote that are helping people get started in their careers, that are helping bring to light some of the real challenges of how we hire in cyber security and in tech in general. And so, that’s the thing that’s going to make it worthwhile. And so yeah, I’m super excited that it’s looking like we’re mere weeks now from this thing being shipped to people who have bought it.

Corey: So, now it’s racing, whether this gets published before the book does. So, we’ll see. There is a bit of a production lag here because, you know, we have to make me look pretty and that takes a tremendous amount of effort.

Alyssa: Oh, stop. Come on now. But it will be interesting to see. Like, that would actually be really cool if they came out at about the same time. Like, you know, I’m just saying.

Corey: Yeah. We’ll see how it goes. Where’s the best place for people to find you if they want to learn more?

Alyssa: About the book or in general?

Corey: Both.

Alyssa: So—

Corey: Links will of course be in the [show notes 00:32:49]. Let’s not kid ourselves here.

Alyssa: The book is real easy. Go to Alyssa—A-L-Y-S-S-A, back here behind me for those of you seeing the video. Um—I can’t point the right direction. There we go. That one. A-L-Y-S-S-A dot link—L-I-N-K slash book. It’s that simple. It’ll take you right to Manning’s site, you can get in.

Still in that early access program, so if you bought it today, you would still be able to start reading the draft versions of it. If you want to know more about me, honestly, the easiest way is to find me on Twitter. You can hear all the ridiculousness of flight school and barbecue and some security topics, too, once in a while. But at @alyssam_infosec. Or if you want to check out the website where I blog, every rare occasion, it’s alyssasec.com.

Corey: And all of that will be in the [show notes 00:33:41]. Thank you—

Alyssa: There’s a lot. [laugh].

Corey: I’m looking forward to seeing it, too. Thank you so much for taking the time to deal with my nonsense today. I really appreciate it.

Alyssa: Oh, that was nonsense? Are you kidding me? This was a great discussion. I really appreciate it.

Corey: As have I. Thanks again for your time. It is always great to talk to people smarter than I am—which is, let’s be clear, most people—Alyssa Miller, BISO at S&P Global. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice—or smash the like and subscribe button if this is on the YouTubes—whereas if you’ve hated the podcast, same thing, five-star review, platform of choice, smash both of the buttons, but also leave an angry comment, either on the YouTube video or on the podcast platform, saying that this was a waste of your time and what you didn’t like about it because you don’t need to read Alyssa’s book; you’re going to get a job the tried and true way, by printing out a copy of your resume and leaving it on the hiring manager’s pillow in their home.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Sharone

I'm Sharone Zitzman, a marketing technologist and open source community builder, who likes to work with engineering teams that are building products that developers love. Having built both the DevOps Israel and Cloud Native Israel communities from the ground up, today I spend my time finding the places where technology and people intersect and ensuring that this is an excellent experience. You can find my talks, articles, and employment experience at rtfmplease.dev. Find me on Twitter or Github as @shar1z.

Links Referenced:

  • Personal Twitter: https://twitter.com/shar1z
  • Website: https://rtfmplease.dev
  • LinkedIn: https://www.linkedin.com/in/sharonez/
  • @TLVCommunity: https://twitter.com/TLVcommunity
  • @DevOpsDaysTLV: https://twitter.com/devopsdaystlv

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: DoorDash had a problem as their cloud native environment scaled and developers delivered new features, their monitoring system kept breaking down. In an organization where data is used to make better decisions about technology and about the business, losing observability means the entire company loses their competitive edge. With Chronosphere, DoorDash is no longer losing visibility into their applications suite. The key? Chronosphere is an open source compatible, scalable, and reliable observability solution that gives the observability lead at DoorDash business, competence, and peace of mind. Read the full success story at snark.cloud/chronosphere. That's snark.cloud/C-H-R-O-N-O-S-P-H-E-R-E.

Corey: The company 0x4447 builds products to increase standardization and security in AWS organizations. They do this with automated pipelines that use well-structured projects to create secure, easy-to-maintain and fail-tolerant solutions, one of which is their VPN product built on top of the popular OpenVPN project which has no license restrictions; you are only limited by the network card in the instance. To learn more visit: snark.cloud/deployandgo

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn and I have been remiss by not having today’s guest on years ago because back before I started this ridiculous nonsense that, well, whatever it is you’d call what I do for a living, I did other things instead. I did the DevOps, which means I was sad all the time. And the thing that I enjoyed was the chance to go and speak on conference stages. One of those stages, early on in my speaking career, was at DevOpsDays Tel Aviv.

My guest today is Sharone Zitzman, who was an organizer of DevOpsDays Tel Aviv, who started convincing me to come back. And today is in fact, in the strong tradition here of making up your own job titles in ways that make people smile, she is the Chief Manual Reader at RTFM Please Ltd. Sharone, thank you for joining me.

Sharone: Thank you for having me, Corey. Israelis love the name of my company, but Americans think it has a lot of moxie and chutzpah. [laugh].

Corey: It seems a little direct and aggressive. It’s like, oh, good, you are familiar with how this is going to go. There’s something to be said for telling people what you do on the tin upfront. I’ve never been a big fan of trying to hide that. I mean, the first iteration of my company was the Quinn Advisory Group because I thought, you know, let’s make it look boring and sedate and like I can talk to finance people. And yeah, that didn’t last more than ten seconds of people talking to me.

Also, in hindsight, the logo of a big stylized Q. Yeah, I would have had to change that anyway, for the whole QAnon nonsense because I don’t want to be mistaken for that particular brand of nuts.

Sharone: Yeah, I decided to do away with the whole formalities and upfront, just go straight [laugh]. For the core of who we are, Corey; you are very similar in that. So, yes. Being a dev first company, I thought the developers would appreciate such a title and name for my company. And I have to give a shout out here to Avishai Ish-Shalom, who’s my friend from the community who you also know from the DevOpsDays community.

Corey: Oh, yeah @nukemberg on Twitter—

Sharone: Yes exactly.

Corey: For those who are not familiar.

Sharone: [laugh]. Yep. He coined the name.

Corey: The problem that I found is that people when they start companies or they manage their careers, they don’t bias for the things that they’re really good at. And it took me a long time to realize this, I finally discovered, “Ah, what am I the best at? That’s right, getting myself fired for my personality, so why don’t I build a business where that stops being a liability?” So, I started my own company. And I can tell this heroic retcon of what happened, but no, it’s because I had nowhere else to go at that point.

And would you hire me? Think about this for a minute. You, on the other hand, had options. You are someone with a storied history in community building, in marketing to developers without that either coming across as insincere or that marked condescending accent that so many companies love to have of, “Oh, you’re a developer. Let me look at you and get down on my hands and knees like we’re going camping and tell a story in ways that actively and passively insult you.”

No, you have always gotten that pitch-perfect. The world was your oyster. And for some godforsaken reason, you looked around and decided, “Ah, I’m going to go out independently because you know what I love? Worrying.” Because let’s face it, running your own company is an exercise in finding new and exciting things to worry about that 20 minutes ago, you didn’t know existed. I say this from my own personal experience. Why would you ever do such a thing?

Sharone: [laugh]. That’s a great question. It was a long one, but a good one. And I do a thing where I hit the mic a lot because I also have. I can’t control my hand motions.

Corey: I too speak with my hands. It’s fine.

Sharone: [laugh]. Yeah, so it’s interesting because I wanted to be independent for a really long time. And I wasn’t sure, you know, if it was
something that I could do if I was a responsible enough adult to even run my own company, if I could make it work, if I could find the business, et cetera. And I left the job in December 2020, and it was the first time that I hadn’t figured out what I was doing next yet. And I wanted to take some time off.

And then immediately, like, maybe a week after I started to get a lot of, like, kind of people reaching out. And I started to interview places and I started to look into possibly being a co-founder at places and I started to look at all these different options. And then just, I was like, “Well…. This is an opportunity, right? Maybe I should finally—that thing that’s gnawing at the back of my head to see if, like, you know if I should go for this dream that I’ve always wanted, maybe now I can just POC it and see if, you know, it’ll work.”

And it just, like, kind of exploded on me. It was like there was so much demand, like, I just put a little, like, signal out to the world that this is something that I’m interested in doing, and everyone was like, “Ahh, I need that.” [laugh]. I wanted to take a quarter off and I signed my first clients already on February 1st, which was, like, a month after. I left in December and that—it was crazy. And since then, I’ve been in business. So, yeah. So, and since then, it’s also been a really crazy ride; I got to discover some really exciting companies. So.

Corey: How did you get into this? I found myself doing marketing-adjacent work almost entirely by accident. I started the newsletter and this podcast, and I was talking to sponsors periodically and they’d come back with, “Here’s the thing we want you to talk about in the sponsor read.” And it’s, “Okay, you want to give people a URL to go to that has four sub-directories and entire UTM code… okay, have you considered, I don’t know, not?” And because so much of what they were talking about did not resonate.

Because I have the engineering background, and it was, I don’t understand what your company does and you’re spending all your time talking about you instead of my painful problem. Because as your target market, I don’t give the slightest of shits about you, I care about my problem, so tell me how you’re going to solve my problem and suddenly I’m all ears. Spend the whole time talking about you, and I could not possibly care less and I’ll fast-forward through the nonsense. That was my path to it. How did you get into it?

Sharone: How did I get into it? It’s interesting. So, I started my journey in typical marketing, enterprise B2B marketing. And then at GigaSpaces, we kickstarted the open-source project Cloudify, and that’s when I found myself leading this project as the open-source community team leader, building, kind of, the community from the ground floor. And I discovered a whole new world of, like, how to build experience into your marketing, kind of making it really experiential and making sure that everyone has a really, really easy and frictionless way of using your product, and that the product—putting the product at the center and letting it speak for itself. And then you discover this whole new world of marketing where it’s—and today, you know, it has more of a name and a title, PLG, and people—it has a whole methodology and practice, but then it was like we were—

Corey: PLG? I’m unfamiliar with the acronym. I thought tech was bad for acronyms.

Sharone: Right? [laugh]. So, product-led growth. But then, you know, like, kind of wasn’t solidified yet. And so, a lot of what we were doing was making sure that developers had a really great experience with the product then it kind of sold itself and marketed itself.

And then you understood what they wanted to hear and how they wanted to consume the product and how they wanted it to be and to learn about it and to kind of educate themselves and get into it. And so, a lot of the things that I learned in the context of marketing was very guerilla, right, from the ground up and kind of getting in front of people and in the way they wanted to consume it. And that taught me a lot about how developers consume technology, the different channels that they’re involved in, and the different tools that they need in order to succeed, and the different, you know, all the peripheral experience, that makes marketing really, really great. And it’s not about what you’re selling to somebody; it’s making your product shine and making the experience shine, making them ensure that it’s a really, really easy and frictionless experience. You know, I like how [Donald Bacon 00:08:00] says it; he calls it, like, mean time to hello world, and that to me is the best kind of marketing, right? When you enable people to succeed very, very quickly.

Corey: Yeah, there’s something to be said for the ring of authenticity and the rest. Periodically I’ll promote guest episodes on this, where it’s a sponsored episode where people get up and they talk about what they’re working on. And they’re like, “Great. So, here’s the sales pitch I want to give,” and it’s no you won’t because first, it won’t work. And secondly, I’m sorry, whether it’s a promoted episode or not, I will not publish something that isn’t good because I have a reputation to uphold here.

And people run into challenges an awful lot when they’re trying to effectively tell their story. If you have a startup that was founded by an engineer, for example, as so many of these technical startups were, the engineer is often so deeply and profoundly in love with this problem space and the solution and the rest, but if they talk about that, no one cares about the how. I mean, I fix AWS bills, and people don’t care—as a general rule—how I do that at all if they’re in my target market. They don’t care if it’s through clever optimization, amazing tooling, doing it on-site, or taking hostages in Seattle. They care about their outcome much more than they ever do about the how.

The only people who care about the how are engineers who very often are going to want to build it themselves, or work for you, or start a competitor. And it doesn’t resonate in quite the same way. It’s weird because all these companies are in slightly different spaces; all of them tend to do slightly different things—or very different things—but so many of the challenges that I see in the way that they’re articulating what they do to customers rhymes with one another.

Sharone: Yeah. So, I agree completely that developers will talk often about how it works. How it works. How does it work under the hood? What are the bits and bytes, you know?

Like, nobody cares about how it works. People care about how will this make my life better, right? How will this improve my life? How will this change my life? [laugh]. As an operations engineer, if I’m, you know, crunching through logs, how will this tool change that? What my days look like? What will my on-call rotation look like? What will—you know, how are you changing my life for the better?

So, I think that that’s the question. When you learn how to crystallize the answer to that question and you hit it right on the mark—you know, and it takes a long time to understand the market, and to understand the buying persona, and t—and there’s so much that you have to do in the background, and so much research you have to do to understand who is that person that needs to have that question answered? But once you do and you crystallize that answer, it lands. And that’s the fun part about marketing, really trying to understand the person who’s going to consume your product and how you can help them understand that you will make their life better.

Corey: Back when I was starting out as a consultant myself, I would tell stories that I had seen in the AWS billing environment, and I occasionally had clients reach out to me, “Hey, why don’t you tell our story in public?” It’s, “Because that wasn’t your story. That was something I saw on six different accounts in the same month. It is something that everyone is feeling.” It’s, people think that you’re talking about them.

So, with that particular mindset on this, without naming specific companies, what themes are you seeing emerging? What are companies getting wrong when they are attempting and failing to market effectively to developers?

Sharone: So, exactly what we’re talking about in terms of the product pitch, in that they’re talking at developers from this kind of marketing speak and this business language that, you know, developers often—you know, unless a company does a really, really good job of translating, kind of, the business value—which they should do, by the way—to engineers, but oftentimes, it’s a little bit far from them in the chain, and so it’s very hard for them to understand the business fluff. If you talk to them in bits and bytes of this is what my day-to-day developer workflow looks like and if we do these things, it’ll cut down the time that I’m working on these things, it’ll make these things easier, it’ll help streamline whatever processes that are difficult, remove these bottlenecks, and help them understand, like I said, how it improves their life.

But the things that I’ve seen breakdown is also in the authenticity, right? So obviously, the world is built on a lot of the same gimmicks and it’s just a matter of whether you’re doing it right or not, right? So, there’s so much content out there and webcasts and webinars, and I don’t know what and podcasts and whatever it is, but a lot of the time, people, their most valuable asset is their time. And if you end up wasting their time, without it being, like, really deeply valuable—if you’re going to write content, make sure that there is a valuable takeaway; if you’re going to create a webinar, make sure that somebody learned something. That if they’re investing their time to join your marketing activities, make sure that they come away with something meaningful and then they’ll really appreciate you.

And it’s the same idea behind the whole DevOpsDays movement with the law of mobility and open spaces that people if they find value, they’ll join this open space and they’ll participate meaningfully and they’ll be a part of your event, and they’ll come back to your event from year to year. But if you’re not going to provide that tangible value that somebody takes away, and it’s like, okay, well, I can practically apply this in my specific tech stack without using your tool, without having to have this very deterministic or specific kind of tech stack that they’re talking about. You want to give people something—or even if it is, but even how to do it with or without, or giving them, like, kind of practical tools to try it. Or if there’s an open-source project that they can check out first, or some kind of lean utility that gives them a good indication of the value that this will give them, that’s a lot more valuable, I think. And practically understandable to somebody who wants to eventually
consume your product or use your products.

Corey: The way that I see things, at least in the past couple of years, the pandemic has sharpened an awful lot of the messaging that needs to happen. Because in most environments, you’re sitting at a DevOpsDays in the front row or whatnot, and it’s time for the sponsor talks and someone gets up and starts babbling and wasting your time, most people are not going to get up and leave. Okay, they will in Israel, but in most places, they’re not going to get up and leave, whereas in pandemic land, it’s you are one tab away from something I actually want—

Sharone: Exactly.

Corey: To be doing, so if you become even slightly boring, it’s not going to go well. So, you have to be on message, you have to be on point or no one cares. People are like, “Oh, well what if we say the wrong thing and people wind up yelling about us on Twitter?” It’s like unless it is for something horrifying, you should be so lucky because people are then talking about you. The failure mode isn’t that people don’t like your product, it’s no one talks about it.

Sharone: Yeah. No such thing as bad publicity [crosstalk 00:14:32] [laugh]—

Corey: Oh, there very much is such a thing is bad publicity. Like, “I could be tweeting about your product most days,” is apparently a version of that, according to some folks. But it’s a hard problem to solve for. And one of the things that continually surprises me is the things I’m still learning about this entire industry. The reason that people sponsor this show—and the rates they pay, to be direct—have little bearing to the actual size of the audience—as best we can tell; lies, damn lies, and podcast statistics; if you’re listening to this, let me know. I’d love to know if anyone listens to this nonsense—but when you see all of that coming out, why are we able to charge the rates that we do?

It’s because the long-term value of someone who is going to buy a long-term subscription or wind up rolling out something like ChaosSearch or whatnot that is going to be a fundamental tenet of their product, one prospect becoming a customer pays for anything, I can sell a company, it will sponsor—they can pay me to sponsor for the next ten years, as opposed to the typical mass-market audience where well, I’m here to sling Casper mattresses today or something. It’s a different audience and there’s a different perception there. People are starting to figure out the value of—in an age where tracking is getting harder and harder to do and attribution will drive you nuts, instead of go where your audience is. Go where the people who care about the problem that you have and will experience that problem are going to hang out. And it always is wild to me to see companies missing out on that.

It’s, “Okay, so you’re going to do a $25 million billboard ad in spotted in airports around the world talking about your company… but looking at your billboard, it makes no sense. I don’t understand what it’s there for.” Even as a brand awareness play, it fails because your logo is tiny in the corner or something. It’s you spent that much money on ads, and maybe a buck on messaging because it seems like with all that attention you just bought, you had nothing worthwhile to say. That’s the cardinal sin to me at least.

Sharone: Yeah. One thing that I found—and back to our community circuit and things that we’ve done historically—but that’s one thing that, you know, as a person comes from community, I’ve seen so much value, even from the smaller events. I mean, today, like with Covid and the pandemic and everything has changed all the equilibrium and the way things are happening. But some meetups are getting smaller, face-to-face events are getting smaller, but I’ve had people telling me that even from small, 30 to 40 people events, they’ll go up and they’ll do a talk and great, okay, a talk; everybody does talks, but it’s like, kind of, the hallway track or the networking that you do after the talk and you actually talk to real users and hear their real problems and you tap into the real community. And some people will tell me like, I had four concrete leads from a 30-person meet up just because they didn’t even know that this was a real challenge, or they didn’t know that there was a tool that solves this problem, or they didn’t understand that this can actually be achieved today.

Or there’s so many interesting technologies and emerging technologies. I’m privileged to be able to be at the forefront of that and discover it all, and I if I could, I would drop names of all of the awesome companies that work for me, that I work with, and just give them a shout out. But really, there’s so many amazing companies doing, like, developer metrics, and all kinds of troubleshooting and failure analysis that’s, like, deeply intelligent—and you’re going to love this one: I have a Git replacement client apropos to your closing keynote of DevOpsDays 2015—and tapping into the communities and tapping into the real users.

And sometimes, you know, it’s just a matter of really understanding how developers are working, what processes look like, what workflows look like, what teams look like, and being able to architect your products and things around real use cases. And that you can only discover by really getting in front of actual users, or potential users, and learning from them and feedback loops, and that’s the little core behind DevRel and developer advocacy is really understanding your actual users and your consumers, and encouraging them to you know, give you feedback and try things, and beta programs and a million things that are a lot more experiential today that help you understand what your users need, eventually, and how to actually architect that into your products. And that’s the important part in terms of marketing. And it’s a whole different marketing set. It’s a whole different skill set. It’s not talking at people, it’s actually… ingesting and understanding and hearing and implementing and bringing it into your products.

Corey: And it takes time. And you have to make yourself synonymous with a painful problem. And those problems are invariably very point-in-time specific. I don’t give a crap about log aggregation today, but in two weeks from now, when I’m trying to chase down 18 different Lambdas function trying to figure out what the hell’s broken this week, I suddenly will care very much about log aggregation. Who was that company that’s in that space that’s doing interesting things? And maybe it’s Cribl, for example; they do a lot of stuff in that space and they’ve been a good sponsor. Great.

I start thinking about those things in that light because it is—when I started having these problems, it sticks in your head and it resonates. And there’s value and validity to that, but you’re never going to be able to attribute that either, which is where people often lose their minds. Because for anything even slightly complicated—you’re going to be selling things to big bank—great, good on you. Most of those customers are not going to go and spin up a trial in the dead of night. They’re going to hear about you somewhere and think, “Ohh, this is interesting.”

They’re going to talk about a meeting, they’re going to get approval, and at that point, you have long since lost any tracking opportunity there.
So, the problem is that by saying it like this, as someone who is a publisher, let’s be very clear here, it sounds like you’re trying to justify your entire business model. I feel like that half the time, but I’ve been reassured by people who are experts in doing these things, like, oh, yeah, we have data on this; it’s working. So, the alternative is either I accept that they’re right or I sit here and arrogantly presume I know more about marketing than people who’ve devoted their entire careers to it. I’m not that bold. I am a white guy in tech, but not that much.

Sharone: Yeah, I mean, the DevRel measurement problem is a known problem. We have people like [unintelligible 00:20:21] who have written about it. We have [Sarah Drasner 00:20:23], we have a million people that have written really, really great content about how do you really measure DevRel and the quality. And one of the things that I liked, Philipp Krenn, the dev advocate at Elastic once said in one of his talks that, you know, “If you’re measuring your developer advocates on leads, you’re a marketing organization. If you’re measuring them on revenue, you’re a sales organization. It’s about reach, engagement, and awareness, and a lot of things that it’s much, much harder to measure.”

And I can say that, like, once upon a time, I used to try and attribute it at Cloudify. Like, I remember thinking, like, “Okay, maybe I could really track this back to, you know, the first touch that I actually had with this user.” It’s really, really difficult, but I do remember, like, when we used to go out into the events and we were really active in the OpenStack community, in the DevOps community, and many other things, and I remember, like, even after events, like, you get all those lead gen emails. All I would say now is, like, “Hey, if you missed us at the booth, you know, and you want still want a t-shirt, you know, reach out and I’ll ship it to you.” And some of those eventually, after we continued the relationship, and we, you know, when we were friends and community friends, six months later, when they moved to their next role at their next job, they were like, “Oh, now I have an opportunity to use Cloudify and I’m going to check it out.”

And it’s very long relationship that you have to cultivate. It has to be, you know, mutual. You have to be, you have to give be giving something and eventually is going to come back to you. Good deeds come back to you. So, I—that’s my credo, by the way, good deeds come back to you. I believe in that and I try to live by that.

Corey: This episode is sponsored in parts by our friend EnterpriseDB. EnterpriseDB has been powering enterprise applications with PostgreSQL for 15 years. And now EnterpriseDB has you covered wherever you deploy PostgreSQL on-premises, private cloud, and they just announced a fully-managed service on AWS and Azure called BigAnimal, all one word. Don’t leave managing your database to your cloud vendor because they’re too busy launching another half-dozen managed databases to focus on any one of them that they didn’t build themselves. Instead, work with the experts over at EnterpriseDB. They can save you time and money, they can even help you migrate legacy applications—including Oracle—to the cloud. To learn more, try BigAnimal for free. Go to biganimal.com/snark, and tell them Corey sent you.

Corey: So, I have one last question for you and it is pointed and the reason I buried it this deep in the episode is so that if I open with it, I will get letters and I’m hoping to get fewer of them. But I met you, again, at DevOpsDays Tel Aviv, and it was glorious. And then you said, “This is fun. Come help me organize it next year.”

And I, like an idiot said, “Sure, that sounds awesome because I love going to conferences and it’s great. So, what’s involved?” “Oh, a whole bunch of meetings.” “Okay, great.” “And planning”—things I’m terrible at—“Okay.” And then the big day finally arrives where, “Great, when do we get to get on stage and tell a story?” Like, “That’s the neat part. We don’t.” So, I have to ask, given that it is all behind-the-scenes work that is fairly thankless unless you really screw it up because then it’s very visible, what is the point of being so involved in the community?

Sharone: Wow, that’s a big question, Corey.

Corey: It really is.

Sharone: [laugh].

Corey: Because you’ve been involved in community for a long time and you’re very good at it.

Sharone: It’s true. It’s true. Appreciate it, thank you. So, for me, first of all, I enjoy, kind of, the people aspect of it, absolutely. And that people aspect of it actually has played out in so many different ways.

Corey: Oh, you mean great people, and also me.

Sharone: [laugh]. Particularly you, Corey, and we will bring you back. [laugh]. And we will make sure you chop wood and carry water because eventually it’ll fill your soul, you’ll see. [laugh] one of the things that really I have had the privilege and honor, and having come out of, like, kind of all my community work is really the network I’ve built and the people that I’ve met.

And I’ve learned so much and I’ve grown so much, but I’ve also had the opportunity to connect people, connect things that you wouldn’t imagine, un—seemingly-related things. So, there are so many friends of mine that have grown up with me in this community, it’s been already ten years now, and a lot of folks have now been going on to new adventures and are looking to kickstart their new startup and I can connect them to this investor, I can connect them to this other person who is maybe a good, you know, partner for their startup, and hiring opportunities, and something—I’ve had this, like, privilege of kind of being able to connect Israel to the outer world and other things and the global kind of community, and also bring really intelligent folks into the community. And this has just created this amazing flywheel of opportunity that I’m really happy to be at the center of. And I think I’ve grown as a person, I think our community has grown, has learned, and there’s a lot of value in that, I think, yeah. We got to meet wonderful folks like you, Corey. [laugh].

Corey: It has its moments. Again, you’re one of those rarities in that it’s almost become a trope in VC land where VCs always like, “How may I be useful?” And it’s this self-serving transparent thing. Every single time you have deigned to introduce me to someone, it’s been a productive conversation and I’m always glad I took the meeting. That is no small thing.

A lot of people say, “I’m good at community,” which is sort of cover for, “I’m not good at anything,” but in your case, it—

Sharone: [laugh]. [I’m an entrepreneur 00:24:48].—

Corey: Is very much not true. Oh, yeah. I’m a big believer that ‘entrepreneur’ and ‘hero’ and other terms like that are things people call you; you don’t call yourself that. It always feels weird for, “Oh, he’s an entrepreneur.” It’s like, that’s a pretty lofty word for shitposting, but okay, we’ll roll with it.

It doesn’t work that way. You’ve clearly invested long-term in a building reputation for yourself by building a name for yourself in the space, and I know that whenever you reach out to me as a result, you are not there to waste my time or shill some bullshit. It is always something that is going to, even if I don’t love every aspect of it or agree with the core of the message you’re sending, great, it is never not going to be worth my time, which is why I’m so glad I got the chance to talk to you this show.

Sharone: I appreciate that. It’s something that I really believe in, I don’t want to waste people’s time and I really only will connect folks or only really will reach out to someone if I do think that there’s something meaningful for both sides. It’s never only what’s in it for me, also. I also want to make sure that there’s something in it for the other person and it’s something that makes sense and it’s meaningful for both sides. I’ve had the opportunity of meeting such interesting folks, and sometimes it’s just like, “You must meet. [laugh]. You will love each other.” You will have so much to do together or it’s so much collaboration opportunity.

And so yeah, I really am that type of person. And I’ll even say from a personal perspective, you know, I know a lot of people, and I’ve even been asked from the flip side, “Okay, is this a toxic manager? Or is this a, you know, a good hire? Is this”—and I tried to provide really authentic input so people make the right decisions, or make, you know, the right contacts, or make—and that’s something I really value. And I managed to build trust with a lot of really great folks—

Corey: And also me—

Sharone: —and it’s come back to me, also. And—[laugh] and particularly you, again. [laugh].

Corey: If people want to learn more about how you see the world and the space and otherwise bask in your wisdom, where’s the best place to find you?

Sharone: So, I’m on Twitter as @shar1z, which is SharoneZ. Basically, everyone thinks it’s such a smart, or I don’t know what, like, or an esoteric screen name. And I’m like, no, it’s just my name, I just—the O-N-E is… the one. [laugh].

So yes, shar1z on Twitter, but also my website, rtfmplease.dev, you can reach out, there’s a contact form there. You can find me on the web anywhere—LinkedIn. Reach out, I answer almost all my DMs when I can. It’s very rare that I don’t answer DMs. Maybe there’ll be a slight lag, but I do. And I really do like when folks reach out to me. I do like it when people try and make contact.

Corey: And you can also be found, of course, wherever find DevOps products are sold, on stage apparently.

Sharone: [laugh]. The DevOps community, that’s right. @TLVCommunity, @DevOpsDaysTLV—don’t out me. All those are—yes, those are also
handles that I run on Twitter, it’s true.

Corey: Excellent.

Sharone: So, when you see them all retweeting the same tweet, yes, it’s happening within same five minutes, it’s me.

Corey: Oh, that would have made it way easier to go viral. My God, I should have just thought of that earlier.

Sharone: [laugh].

Corey: Thank you so much for your time. I appreciate it.

Sharone: Thank you, Corey, for having me. It’s been a privilege and honor being on your show and I really do think that you are doing wonderful things in the cloud space. You’re teaching us, and we’re all learning, and you—keep up the good work.

Corey: Well, thank you. I appreciate that.

Sharone: I also want to add that on proposed marketing and whatever, I do actually listen to all of your openings of all of your shows because they’re not fluffy and I like that you do, like, kind of a deep explanation, a deep technical explanation of what your sponsoring product does, and it gives a lot more insight into why is this important. So, I think you’re doing that right. So, anybody who’s sponsoring this show, listen. Corey knows what he’s doing.

Corey: Well, thank you. I appreciate that. Yay, “I know what I’m doing.” That one’s going in the testimonial kit. My God.

Sharone: [laugh]. That’s the name of this episode, “Corey knows what he’s doing.”

Corey: We’re going to roll with it, you know. No take-backsies. Sharone Zitzman, Chief Manual Reader at RTFM Please. I’m Cloud Economist
Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review of your podcast platform of choice, or if it’s on the YouTubes smash the like and subscribe buttons, whereas if you’ve hated this show, exact same thing—five-star review wherever you happen to find it, smash both the buttons—but also leave an insulting comment telling me that I’m completely wrong which then devolves into an 18-page diatribe about exactly how your nonsense, bullshit product is built and works.

Sharone: [laugh].

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Rafal

Rafal is Serverless Engineer at Stedi by day, and Dynobase founder by night - a modern DynamoDB UI client. When he is not coding or answering support tickets, he loves climbing and tasting whiskey (not simultaneously).

Links Referenced:

  • Company Website: https://dynobase.dev

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored by our friends at Revelo. Revelo is the Spanish word of the day, and its spelled R-E-V-E-L-O. It means “I reveal.” Now, have you tried to hire an engineer lately? I assure you it is significantly harder than it sounds. One of the things that Revelo has recognized is something I’ve been talking about for a while, specifically that while talent is evenly distributed, opportunity is absolutely not. They’re exposing a new talent pool to, basically, those of us without a presence in Latin America via their platform. It’s the largest tech talent marketplace in Latin America with over a million engineers in their network, which includes—but isn’t limited to—talent in Mexico, Costa Rica, Brazil, and Argentina. Now, not only do they wind up spreading all of their talent on English ability, as well as you know, their engineering skills, but they go significantly beyond that. Some of the folks on their platform are hands down the most talented engineers that I’ve ever spoken to. Let’s also not forget that Latin America has high time zone overlap with what we have here in the United States, so you can hire full-time remote engineers who share most of the workday as your team. It’s an end-to-end talent service, so you can find and hire engineers in Central and South America without having to worry about, frankly, the colossal pain of cross-border payroll and benefits and compliance because Revelo handles all of it. If you’re hiring engineers, check out revelo.io/screaming to get 20% off your first three months. That’s R-E-V-E-L-O dot I-O slash screaming.

Corey: The company 0x4447 builds products to increase standardization and security in AWS organizations. They do this with automated pipelines that use well-structured projects to create secure, easy-to-maintain and fail-tolerant solutions, one of which is their VPN product built on top of the popular OpenVPN project which has no license restrictions; you are only limited by the network card in the instance. To learn more visit: snark.cloud/deployandgo

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. It’s not too often that I wind up building an episode here out of a desktop application. I’ve done it once or twice, and I’m sure that the folks at Microsoft Excel are continually hoping for an invite to talk about things. But we’re going in a bit of a different direction today. Rafal Wilinski is a serverless engineer at Stedi and, in apparently what is the job requirement at Stedi, he also has a side project that manifests itself as a desktop app. Rafal, thank you for joining me today. I appreciate it.

Rafal: Yeah. Hi, everyone. Thanks for having me, Corey.

Corey: I first heard about you when you launched Dynobase, which is awesome. It sounds evocative of dinosaurs unless you read it, then it’s D-Y-N-O, and it’s, “Ah, this sounds a lot like DynamoDB. Let me see what it is.” And sure enough, it was. As much as I love misusing things as databases, DynamoDB is actually a database that is decent and good at what it does.

And please correct me if I get any of this wrong, but Dynobase is effectively an Electron app that you install, at least on a Mac, in my case; I don’t generally use other desktops, that’s other people’s problems. And it provides a user-friendly interface to DynamoDB that is not actively hostile to the customer.

Rafal: Yeah, exactly. That was the goal. That’s how I envisioned it, and I hope I executed correctly.

Corey: It was almost prescient in some ways because they recently redid the DynamoDB console in AWS to actively make it worse, to wind up working with individual items, to modify things. It feels like they are validating your market for you by, “Oh, we really like Dynobase. How do we drive more traffic to it? We’re going to make this thing worse.” But back then when you first created this, the console was his previous version. What was it that inspired you to say, “You know what I’m going to build? A desktop application for a cloud service.” Because on the surface, it seems relatively close to psychotic, but it’s brilliant.

Rafal: [laugh]. Yeah, sure. So, a few years ago, I was freelancing on AWS. I was jumping between clients and my side projects. That also involved jumping between regions, and AWS doesn’t have a good out-of-the-box solution for switching your accounts and switching your regions, so when you want it to work on your client table in Australia and simultaneously on my side project in Europe, there was no other solution than to have two browser windows open or to, even, browsers open.

And it was super frustrating. So, I was like, hey, “DynamoDB has SDK. Electron is this thing that allows you to make a desktop application using HTML and JS and some CSS, so maybe I can do something with it.” And I was so naive to think that it’s going to be a trivial task because it’s going to be—come on, it’s like, a couple of SDK calls, displaying some lists and tables, and that’s pretty much it, right?

Corey: Right. I use Retool as my system to build my newsletter every week, and that is the front-end I use to interact with DynamoDB. And it’s great. It has a table component that just—I run a query that, believe it or not, is a query, not a scan—I know, imagine that, I did something slightly right this one time—and it populates things for the current issue into it, and then I basically built a CRUD API around it and have components that let me update, delete, remove, the usual stuff. And it’s great, it works for my purposes, and it’s fine.

And that’s what I use most of the time until I, you know, hit an edge case or a corner case—because it turns out, surprise everyone, I’m bad at programming—and I need to go in and tweak the table myself manually. And that’s where Dynobase, at least for my use case, really comes into its own.

Rafal: Good to hear. Good to hear. Yeah, that was exactly same case why I built it because yeah, I was also, a few years ago, I started working on some project which was really crazy. It was before AppSync times. We wanted to have GraphQL serverless API using single table design and testing principles [unintelligible 00:04:38] there.

So, we’ve been verifying many things by just looking at the contents of the table, and sometimes fixing them manually. So, that was also the thing that motivated me to make the editing experience a little bit better.

Corey: One thing I appreciate about the application is that it does things right. I mean, there’s no real other way to frame that. When I fire up the application myself and I go to the account that I’ve been using it with—because in this case, there’s really only one account that I have that contains the data that I spent that my time working with—and I get access to it on my machine via Granted, which because it’s a federated SSO login. And it says, “Ah, this is an SSL account. Click here to open the browser tab and do the thing.”

I didn’t have to configure Dynobase. It is automatically reading my AWS config file in my user directory. It does a lot of things right. There’s no duplication of work. From my perspective. It doesn’t freak out because it doesn’t know how SSO works. It doesn’t have run into these obnoxious edge case problems that so many early generation desktop interfaces for AWS things seem to.

Rafal: Wow, it seems like it works for you even better than for me. [laugh].

Corey: Oh, well again, how I get into accounts has always been a little weird. I’ve ranted before about Granted, which is something that Common Fate puts out. It is a binary utility that winds up logging into different federated SSO accounts, opens them in Firefox containers so you could have you know, two accounts open, side-by-side. It’s some nice affordances like that. But it still uses the standard AWS profile syntax which Dynobase does as well.

There are a bunch of different ways I’ve logged into things, and I’ve never experienced friction [unintelligible 00:06:23] using Dynobase for this. To be clear, you haven’t paid me a dime. In fact, just the opposite. I wind up paying my monthly Dynobase subscription with a smile on my face. It is worth every penny, just because on those rare moments when I have to work with something odd in DynamoDB, it’s great having the tool.

I want to be very clear here. I don’t recall what the current cost on this is, but I know for a fact it is more than I spend every month on DynamoDB itself, which is fine. You pay for utility, not for the actual raw cost of the underlying resources on it. Some people tend to have issues with that and I think it’s the wrong direction to go in.

Rafal: Yeah, exactly. So, my logic was that it’s a productivity improvement. And a lot of programmers are simply obsessed with productivity, right? We tend to write those obnoxious nasty Bash and Python scripts to automate boring tasks in our day jobs. So, if you can eliminate this chore of logging to different AWS accounts and trying to find them, and even if it takes, like, five or ten seconds, if I can shave that five or ten seconds every time you try to do something, that over time accumulates into a big number and it’s a huge time investment. So, even if you save, like, I don’t know, maybe one hour a month or one hour a quarter, I think it’s still a fair price.

Corey: Your pricing is very interesting, and the reason I say that is you do not have a free tier as such, you have a free seven-day trial, which is great. That is the way to do it. You can sign up with no credit card, grab the thing, and it’s awesome. Dynobase.dev for folks who are wondering.

And you have a solo yearly plan, which is what I’m on, which is $9 a month. Which means that you end up, I think, charging me $108 a year billed annually. You have a solo lifetime option for 200 bucks—and I’m going to fight with you about that one in a second; we’re going to come back to it—then you have a team plan that is for I think for ten licenses at 79 bucks a month, and for 20 licenses it’s 150 bucks a month. Great. And then you have an enterprise option for 250 a month, the end. Billed annually. And I have problems with that, too.

So, I like arguing with pricing, I [unintelligible 00:08:43] about pricing with people just because I find that is one of those underappreciated aspects of things. Let’s start with my own decisions on this, if I may. The reason that I go for the solo yearly plan instead of a lifetime subscription of I buy this and I get to use it forever in perpetuity. I like the tool but, like, the AWS service that underlies it, it’s going to have to evolve in the fullness of time. It is going to have to continue to support new DynamoDB functionality, like the fact that they have infrequent access storage classes now, for tables, as an example. I’m sure they’re coming up with other things as well, like, I don’t know, maybe a sane query syntax someday. That might be nice if they ever built one of those.

Some people don’t like the idea of a subscription software. I do just because I like the fact that it is a continual source of revenue. It’s not the, “Well, five years ago, you paid me that one-off thing and now you expect feature enhancements for the rest of time.” How do you think about that?

Rafal: So, there are a couple of things here. First thing is that the lifetime support, it doesn’t mean that I will be always implementing to my death all the features that are going to appear in DynamoDB. Maybe there is going to be a some feature and I’m not going to implement it. For instance, it’s not possible to create the global tables via Dynobase right now, and it won’t be possible because we think that majority of people dealing with cloud are using infrastructure as a code, and creating tables via Dynobase is not a super useful feature. And we also believe that it’s not going to break even without support. [laugh]. I know it sounds bad; it sounds like I’m not going to support it at some point, but don’t worry, there are no plans to discontinue support [crosstalk 00:10:28]—

Corey: We all get hit by buses from time to time, let’s be clear.

Rafal: [laugh].

Corey: And I want to also point out as well that this is a graphical tool that is a front-end for an underlying AWS service. It is extremely convenient, there is tremendous value in it, but it is not critical path as if suddenly I cannot use Dynobase, my production app is down. It doesn’t work that way, in the sense—

Rafal: Yes.

Corey: Of a SaaS product. It is a desktop application. And huge fan of that as well. So, please continue.

Rafal: Yeah, exactly—

Corey: I just want to make sure that I’m not misleading people into thinking it’s something it’s not here. It’s, “Oh, that sounds dangerous if that’s critical pa”—yeah, it’s not designed to be. I imagine, at least. If so it seems like a very strange use case.

Rafal: Yeah. Also, you have to keep in mind that AWS isn’t basically introducing breaking changes, especially in a service that is so popular as DynamoDB. I cannot imagine them, like, announcing, like, “Hey, in a month, we are going to deprecate this API, so you’d better start, you know,
using this new API because this one is going to be removed.” I think that’s not going to happen because of the millions of clients using DynamoDB actively. So, I think that makes Dynobase safe. It’s built on a rock-solid foundation that is going to change only additively. No features are going to be just being removed.

Corey: I think that there’s a direction in a number of at least consumer offerings where people are upset at the idea of software subscriptions, the idea of why should I pay in perpetuity for a thing? And I want to call out my own bias here. For something like this, where you’re charging $9 a month, I do not care about the price, truly I don’t. I am a price inflexible customer. It could go and probably as high as 50 bucks a month and I would neither notice nor care.

That is probably not the common case customer, and it’s certainly not over in consumer-land. I understand that I am significantly in a privileged position when it comes to being able to acquire the tools that I need. It turns out compared to the AWS bill I have to deal with, I don’t have to worry about the small stuff, comparatively. Not everyone is in that position, so I am very sympathetic to that. Which is why I want to deviate here a little bit because somewhat recently, Dynobase showed up on the AWS Marketplace.

And I can go into the Marketplace now and get a yearly subscription for a single seat for $129. It is slightly more than buying it directly through your website, but there are some advantages for many folks in getting it on the Marketplace. AWS is an approved vendor, for example, so there’s no procurement dance. It counts toward your committed spend on contracts if someone is trying to wind up hitting certain levels of spend on their EDP. It provides a centralized place to manage things, as far as those licenses go when people are purchasing it. What was it that made you decide to put this on the Marketplace?

Rafal: So, this decision was pretty straightforward. It’s just, you know, yet another distribution channel for us. So, imagine you’re a software engineer that works for a really, really big company and it’s super hard to approve some kind of expense using traditional credit card. You basically cannot go to my site and check out with a company credit card because of the processes, or maybe it takes two years. But maybe it’s super easy to click this subscribe on your AWS account. So yeah, we thought that, hey, maybe it’s going to unlock some engineers working at those big corporations, and maybe this is the way that they are going to start using Dynobase.

Corey: Are you seeing significant adoption yet? Or is it more or less a—it’s something that’s still too early to say? And beyond that, are you finding that people are discovering the product via the AWS Marketplace, or is it strictly just a means of purchasing it?

Rafal: So, when it comes to discovering, I think we don’t have any data about it yet, which is supported by the fact that we also have zero subscriptions from the Marketplace yet. But it’s also our fault because we haven’t actually actively promoted the fact, apart from me sending just a tweet on Twitter, which is in [crosstalk 00:14:51]—

Corey: Which did not include a link to it as well, which means that Google was our friend for this because let’s face it, AWS Marketplace search is bad.

Rafal: Well, maybe. I didn’t know. [laugh]. I was just, you know, super relieved to see—

Corey: No, I—you don’t need to agree with that statement. I’m stating it as a fact. I am not a fan of Marketplace search. It irks me because for whatever reason whenever I’m in there looking for something, it does not show me the things I’m looking for, it shows me the biggest partners first that AWS has and it seems like the incentives are misaligned. I’m sure someone is going to come on the show to yell about me. I’m waiting
for your call.

Rafal: [laugh].

Corey: Do you find that if someone is going to purchase it, do you have a preference that they go directly, that they go through the Marketplace? Is there any direction for you that makes more sense than another?

Rafal: So ideally, would like to continue all the customers to purchase the software using the classical way, using the subscriptions for our website because it’s just one flow, one system, it’s simpler, it’s cleaner, but we want it to give that option and to have more adoption. We’ll see if that’s going to work.

Corey: I was going to say there were two issues I had with the pricing. That was one of them. The other is at the high end, the enterprise pricing being $250 a month for unlimited licenses, that doesn’t feel like it is the right direction, and the reason I say that is a 50-person company would wind up being able to spend 250 bucks a month to get this for their entire team, and that’s great and they’re happy. So, could AWS or Coca-Cola, and at that very high level, it becomes something that you are signing up for significant amount of support work, in theory, or a bunch of other directions.

I’ve always found that from where I stand, especially dealing with those very large companies with very specific SLA requirements and the rest, the pricing for enterprise that I always look for as the right answer for my mind is ‘click here to contact us.’ Because procurement departments, for example, we want this, this, this, this, and this around data guarantees and indemnities and all the rest. And well, yeah, that’s going to be expensive. And well, yeah. We’re a procurement company at a Fortune 50. We don’t sign contracts that don’t have two commas in them.

So, it feels like there’s a dialing it in with some custom optionality that feels like it is signaling to the quote-unquote, ‘sophisticated buyer,’ as patio11 likes to say on Twitter from time to time, that might be the right direction.

Rafal: That’s really good feedback. I haven’t thought about it this way, but you really opened my eyes on this issue.

Corey: I’m glad it was helpful. The reason I think about it this way is that more and more I’m realizing that pricing is one of the most key parts of marketing and messaging around something, and that is not really well understood, even by larger companies with significant staff and full marketing teams. I still see the pricing often feels like an afterthought, but personally, when I’m trying to figure out is this tool for me, the first thing I do is—I don’t even read the marketing copy of the landing page; I look for the pricing tab and click because if the only prices ‘call for details,’ I know, A, it’s going to be expensive, be it’s going to be a pain in the neck to get to use it because it’s two in the morning; I’m trying to get something done. I want to use it right now. If I had to have a conversation with your sales team first, that’s not going to be cheap and it’s not going to be something I’m going to be able to solve my problem this week. And that is the other end of it. I yell at people on both sides on that one.

Rafal: Okay.

Corey: Again, none of this stuff is intuitive; all of this stuff is complicated, and the way that I tend to see the world is, granted, a little bit different than the way that most folks who are kicking around databases and whatnots tend to view the world. Do you have plans in the future to extend Dynobase beyond strictly DynamoDB, looking to explore other fine database options like Redis, or MongoDB, or my personal favorite Route 53 TXT records?

Rafal: [laugh]. Yeah. So, we had plans. Oh, we had really big plans. We felt that we are going to create a second JetBrains company. We started analyzing the market when it comes to MongoDB, when it comes to Cassandra, when it comes to Redis. And our first pick was Cassandra because it seemed, like, to have really, really similar structure of the table.

I mean, it’s also no secret it also has a primary index, secondary global indexes, and things like that. But as always, reality surprises us over the amount of detail that we cannot see from the very top. And it isn’t as simple as just an install AWS SDK and install Cassandra Connector on—or Cassandra SDK and just roll with that. It requires a really big and significant investment. And we decided to focus just on one thing and nail this one thing and do this properly.

It’s like, if you go into the cloud, you can try to build a service that is agnostic, it’s not using the best features of the cloud. And you can move your containers, for instance, across the clouds and say, “Hey, I’m cloud-agnostic,” but at the same time, you’re missing out all the best features. And this is the same way we thought about Dynabase. Hey, we can provide an agnostic core, but then the agnostic application isn’t going to be as good and as sophisticated as something tailored specifically for the needs of this database and user using this exact database.

Corey: This episode is sponsored in parts by our friend EnterpriseDB. EnterpriseDB has been powering enterprise applications with PostgreSQL for 15 years. And now EnterpriseDB has you covered wherever you deploy PostgreSQL on premises, private cloud, and they just announced a fully managed service on AWS and Azure called BigAnimal, all one word.

Don't leave managing your database to your cloud vendor because they're too busy launching another half dozen manage databases to focus on any one of them that they didn't build themselves. Instead, work with the experts over at EnterpriseDB. They can save you time and money, they can even help you migrate legacy applications, including Oracle, to the cloud.

To learn more, try BigAnimal for free. Go to biganimal.com/snark, and tell them Corey sent you.

Corey: Some of the things that you do just make so much sense that I get actively annoyed that there aren’t better ways to do it and other places for other things. For example, when I fire up a table in a particular region within Dynobase, first it does a scan, which, okay, that’s not terrible. But on some big tables, that can get really expensive. But you cap it automatically to a thousand items. And okay, great.

Then it tells me, how long did it take? In this case because, you know, I am using on-demand and the rest and it’s a little bit of a pokey table, that scan took about a second-and-a-half. Okay. You scanned a thousand items. Well, there’s a lot more than a thousand items in this table. Ah, you limited it, so you didn’t wind up taking all that time.

It also says that it took 51-and-a-half RCUs—or Read Credit Units—because you know, why use normal numbers when you’re AWS and doing pricing dimensions on this stuff.

Rafal: [laugh].

Corey: And to be clear, I forget the exact numbers for reads, but it’s something like a million read RCUs cost me a dollar or something like that.
It is trivial; it does not matter, but because it is consumption-based pricing, I always live in a little bit of a concern that, okay, if I screw up and just, like, scan the entire 10-megabyte table every time I want to make an operation here, and I make a lot of operations in the course of a week, that’s going to start showing up in the bill in some really unfortunate ways. This sort of tells me as an ongoing basis of what it is that I’m going to wind up encountering.

And these things are all configurable, too. The initial stream limit that you have configured as a thousand. I can set that to any number I want if I think that’s too many or too few. You have a bunch of pagination options around it. And you also help people build out intelligent queries, [unintelligible 00:22:11] can export that to code. It’s not just about the graphical interface clickety and done—because I do love my ClickOps but there are limits to it—it helps formulate what kind of queries I want to build and then wind up implementing in code. And that is no small thing.

Rafal: Yeah, exactly. This is how we also envision that. The language syntax in DynamoDB is really… hard.

Corey: Awful. The term is awful.

Rafal: [laugh]. Yeah, especially for people—

Corey: I know, people are going to be mad at me, but they’re wrong. It is not intuitive, it took a fair bit of wrapping my head around. And more than once, what I found myself doing is basically just writing a thin CRUD API in Lambda in front of it just so I can query it in a way that I think about it as opposed to—now I’m not even talking changing the query modeling; I just want better syntax. That’s all it is.

Rafal: Yeah. You also touch on modeling; that’s also very important thing, especially—or maybe even scan or query. Suppose I’m an engineer with tens years of experience. I come to the DynamoDB, I jump straight into the action without reading any of the documentation—at least that’s my way of working—and I have no idea what’s the difference between a scan and query. So, in Dynobase, when I’m going to enter all those filtering parameters into the UI, I’m going to hit scan, Dynobase is automatically going to figure out for you what’s the best way to query—or to scan if query is not possible—and also give you the code that actually was behind that operation so you can just, like, copy and paste that straight to your code or service or API and have exactly the same result.

So yeah, we want to abstract away some of the weird things about DynamoDB. Like, you know, scan versus query, expression attribute names, expression attribute values, filter, filtering conditions, all sorts of that stuff. Also the DynamoDB JSON, that’s also, like, a bizarre thing. This JSON-type thing we should get out of the box, we also take care of that. So, yeah. Yeah, that’s also our mission to make the DynamoDB as approachable as possible. Because it’s a great database, but to truly embrace it and to truly use it, it’s hard.

Corey: I want to be clear, just for folks who are not seeing some of the benefits of it the way that I’ve described it thus far. Yes, on some level, it basically just provides a attractive, usable interface to wind up looking at items in a DynamoDB table. You can also use it to wind up refining queries to look at very specific things. You can export either a selection or an entire table either to a local file—or to S3, which is convenient—but it goes beyond on that because once you have the query dialed in and you’re seeing the things you want to see, there’s a generate code button that spits it out in—for Python, for JavaScript, for Golang.

And there are a few things that the AWS CLI is coming soon, according to the drop-down itself. Java; ooh, you do like pain. And Golang for example, it effectively exports the thing you have done by clicking around as code, which is, for some godforsaken reason, anathema to most AWS services. “Oh, you clicked around to the console to do a thing. Good job. Now, throw it all away and figure out how to do it in code.” As opposed to, “Here’s how to do what you just did programmatically.” My God, the console could be the best IDE in the world, except that they don’t do it for some reason.

Rafal: Yeah, yeah.

Corey: And I love the fact that Dynobase does.

Rafal: Thank you.

Corey: I’m a big fan of this. You can also import data from a variety of formats, export data, as well. And one of the more obnoxious—you talk about weird problems I have with DynamoDB that I wish to fix: I would love to move this table to a table in a different AWS account. Great, to do that, I effectively have to pause the service that is in front of this because I need to stop all writes—great—export the table, take the table to the new account, import the table, repoint the code to talk to that thing, and then get started again. Now, there are ways to do it without that, and they all suck because you have to either write a shim for it or you have to wind up doing a stream that winds up feeding from one to the other.

And in many cases, well okay, I want to take the table here, I do a knife-edge cutover so that new rights go to the new thing, and then I just want to backfill this old table data into it. How do I do that? The official answer is not what you would expect it to be, the DynamoDB console of ‘import this data.’ Instead, it’s, “Oh, use AWS Glue to wind up writing an ETL function to do all of this.” And it’s… what? How is that the way to do these things?

There are import and export buttons in Dynobase that solve this problem beautifully without having to do all of that. It really is such a different approach to thinking about this, and I am stunned that this had to be done as a third party. It feels like you were using the native tooling and the native console the same way the rest of us do, grousing about it the same way the rest of us do, and then set out to fix it like none of us do. What was it that finally made you say, “You know, I think there’s a better way and I’m going to prove it.” What pushed you over the edge?

Rafal: Oh, I think I was spending, just, hours in the console, and I didn’t have a really sophisticated suite of tests, which forced me [unintelligible 00:27:43] time to look at the data a lot and import data a lot and edit it a lot. And it was just too much. I don’t know, at some point I realized, like, hey, there’s got to be a better way. I browsed for the solutions on the internet; I realized that there is nothing on the market, so I asked a couple of my friends saying like, “Hey, do you also have this problem? Is this also a problem for you? Do you see the same challenges?”

And basically, every engineer I talked to said, “Yeah. I mean, this really sucks. You should do something about it.” And that was the moment I realized that I’m really onto something and this is a pain that I’m not alone. And so… yeah, that gave me a lot of motivation. So, there was a lot of frustration, but there was also a lot of motivation to push me to create a first product in my life.

Corey: It’s your first product, but it does follow an interesting pattern that seems to be emerging, Cloudash—Tomasz and Maciej—wound up doing that as well. They’re also working at Stedi and they have their side project which is an Electron-based desktop application that winds up, we’re interfacing with AWS services. And it’s. What are your job requirements over at Stedi, exactly?

People could be forgiven for seeing these things and not knowing what the hell EDI is—which guilty—and figure, “Ah, it’s just a very fancy term for a DevRels company because they’re doing serverless DevRel as a company.” It increasingly feels an awful lot like that.j, what’s going on over there where that culture just seems to be an emergent property?

Rafal: So, I feel like Stedi just attracts a lot of people that like challenges and the people that have a really strong sense of ownership and like to just create things. And this is also how it feels inside. There is plenty of individuals that basically have tons of energy and motivation to solve so many problems not only in Stedi, but as you can see also outside of Stedi, which is a result—Cloudash is a result, the mapping tool from Zack Charles is also a result, and Michael Barr created a scheduling service. So, yeah, I think the principles that we have at Stedi basically attract top-notch builders.

Corey: It certainly seems so. I’m going to have to do a little more digging and see what some of those projects are because they’re new to me. I really want to thank you for taking so much time to speak with me about what you’re building. If people want to learn more or try to kick the tires on Dynobase which I heartily recommend, where should they go?

Rafal: Go to dynobase.dev, and there’s a big download button that you cannot miss. You download the software, you start it. No email, no credit card required. You just run it. It scans your credentials, profiles, SSOs, whatever, and you can play with it. And that’s pretty much it.

Corey: Excellent. And we will put a link to that in the [show notes 00:30:48]. Thank you so much for your time. I really appreciate it.

Rafal: Yeah. Thanks for having me.

Corey: Rafal Wilinski, serverless engineer at Stedi and creator of Dynobase. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice—or a thumbs up and like and subscribe buttons on the YouTubes if that’s where you’re watching it—whereas if you’ve hated this podcast, same thing—five-star review, hit the buttons and such—but also leave an angry, bitter comment that you’re not going to be able to find once you write it because no one knows how to put it into DynamoDB by hand.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Carla

Carla Stickler is a professional multi-hyphenate advocating for the inclusion of artists in STEM. Currently, she works as a software engineer at G2 in Chicago. She loves chatting with folks interested in shifting gears from the arts to programming and especially hopes to get more women into the field. Carla spent over 10 years performing in Broadway musicals, most notably, “Wicked,” “Mamma Mia!” and “The Sound of Music.” She recently made headlines for stepping back into the role of Elphaba on Broadway for a limited time to help out during the covid surge after not having performed the role for 7 years. Carla is passionate about reframing the narrative of the “starving artist” and states, “When we choose to walk away from a full-time pursuit of the arts, it does not make us failed artists. The possibilities for what we can do and who we can be are unlimited.”

Links Referenced:

  • G2: https://www.g2.com/
  • Personal website: https://carlastickler.com
  • Instagram: https://www.instagram.com/sticklercarla/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Honeycomb. When production is running slow, it’s hard to know where problems originate. Is it your application code, users, or the underlying systems? I’ve got five bucks on DNS, personally. Why scroll through endless dashboards while dealing with alert floods, going from tool to tool to tool that you employ, guessing at which puzzle pieces matter? Context switching and tool sprawl are slowly killing both your team and your business. You should care more about one of those than the other; which one is up to you. Drop the separate pillars and enter a world of getting one unified understanding of the one thing driving your business: production. With Honeycomb, you guess less and know more. Try it for free at honeycomb.io/screaminginthecloud. Observability: it’s more than just hipster monitoring.

Corey: What if there were a single place to get an inventory of what you're running in the cloud that wasn't "the monthly bill?" Further, what if there were a way to compare that inventory to what you were already managing via Terraform, Pulumi, or CloudFormation, but then automatically add the missing unmanaged or drifted parts to it? And what if there were a policy engine to immediately flag and remediate a wide variety of misconfigurations? Well, stop dreaming and start doing; visit snark.cloud/firefly to learn more.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn, there seems to be a trope in our industry that the real engineers all follow what more or less looks like the exact same pattern, where it’s you wind up playing around with computers as a small child and then you wind up going to any college you want—as long as it’s Stanford—and getting a degree in anything under the sun—as long as it’s computer science—and then all of your next jobs are based upon how well you can re-implement algorithms on the whiteboard. A lot of us didn’t go through that path. We wound up finding our own ways to tech. My guest today has one of the more remarkable stories that I’ve come across. Carla Stickler is a software engineer at G2. Carla, thank you for agreeing to suffer my slings and arrows today. It’s appreciated.

Carla: Thanks so much for having me, Corey.

Corey: So, before you entered tech—I believe this is your first job as an engineer and as of the time we’re recording this, it’s been just shy of a year that you’ve done in the role. What were you doing before now?

Carla: Oh, boy, Corey. What was I doing? I definitely was not doing software engineering. I was a Broadway actress. So, I spent about 15 years in New York doing musical theater, touring around the country and Asia in big Broadway shows. And that was pretty much all I did.

I guess, I also was a teacher. I was a voice teacher and I taught voice lessons, and I had a studio and I taught it a couple of faculties in New York. But I was one hundred percent ride-or-die, like, all the way to the end musical theater or bust, from a very, very early age. So, it’s been kind of a crazy time changing careers. [laugh].

Corey: What inspired that? I mean, it doesn’t seem like it’s a common pattern of someone who had an established career as a Broadway actress to wake up one day and say, “You know what I don’t like anymore. That’s right being on stage, doing the thing that I spent 15 years doing. You know what I want to do instead? That’s right, be mad at computers all the time and angry because some of the stuff is freaking maddening.” What was the catalyst that—

Carla: Yeah, sounds crazy. [laugh].

Corey: —inspired you to move?

Carla: It sounds crazy. It was kind of a long time coming. I love performing; I do, and it’s like, my heart and soul is with performing. Nothing else in my life really can kind of replace that feeling I get when I’m on stage. But the one thing they don’t really talk about when you are growing up and dreaming of being a performer is how physically and emotionally taxing it is.

I think there’s, like, this narrative around, like, “Being an actor is really hard, and you should only do it if you can’t see yourself doing anything else,” but they don’t actually ever explain to you what hard means. You know, you expect that, oh, there’s going to be a lot of other people doing it in, I’m going to be auditioning all the time, and I’m going to have a lot of competition, but you never quite grasp the physical and emotional toll that it takes on your body and your—you know, just ongoing in auditions and getting rejections all the time. And then when you’re working in a show eight times a week and you’re wearing four-inch heels on a stage that is on a giant angle, and you’re wearing wigs that are, like, really, really massive, you don’t really—no one ever tells you how hard that is on your body. So, for me, I just hit a point where I was performing nonstop and I was so tired. I was, like, living at my physical therapist’s office, I was living at, like, my head therapist’s office.

I was just trying to, like, figure out why I was so miserable. And so, I actually left in 2015, performing full time. So, I went to get my Master’s in Education at NYU thinking that teaching was my way out of performing full-time.

Corey: It does seem that there’s some congruities—there’s some congruities there between your—instead of performing in front of a giant audience, you’re performing in front of a bunch of students. And whether it’s performing slash educating, well that comes down to almost stylistic differences. But I have a hard time imagining you just reading from your slides.

Carla: Yeah, no, I loved it because it allowed me to create connections with my students, and I found I like to help inspire them on their journeys, and I really like to help influence them in a positive way. And so yeah, it came really natural to me. And my family—or I have a bunch of teachers in my family so, you know, teaching was kind of a thing I just assumed I would be good at, and I think I fell naturally into. But the thing that was really hard for me was while I was teaching, I was still… kind of—I had, like, one foot in performing. I was still, like, going in and out of the show that I’ve been working on, which I didn’t mention.

So, I was in Wicked for, like, ten years, that’s kind of like my claim to fame. And I had been with that show for a really long time, and that was why—when I left to go teach, that was kind of my way out of that big show because it was hard for me to explain to people why it was leaving such a giant show. And teaching was just, like, a natural thing to go into. I felt like it was like a justifiable action, [laugh] you know, that I could explain to, like, my parents for why I was quitting Broadway.

So, you know, I love teaching and—but I—and so I kept that one foot kind of in Broadway, and I was still going in and out of the show. It’s like a vacation cover, filling in whenever they needed me, and I was still auditioning. But I was like, I was still so burned out, you know? Like, I still had those feelings of, like—and I wasn’t booking work; I think my heart just wasn’t really in it. Like, every time I’d go into audition, I would just feel awful about myself every time I left.

And I was starting to really reject that feeling in my life because I was also starting to find there were other things in my life that made me really happy. Like, just having a life. Like, I had—for the first time in a very long time, I had friends that I could hang out with on the weekends because I wasn’t working on the weekend. And I was able to, like, go to, you know, birthdays and weddings and I was having, like, this social life. And then every time I would go on an audition—

Corey: And they did other things with their lives, and it wasn’t—

Carla: Yeah.

Corey: All shop talk all the time—

Carla: Right.

Corey: Which speaking as someone who lives in San Francisco and worked in normal companies before starting this ridiculous one, it seems that your entire social circle can come out of your workplace. And congratulations, it’s now all shop talk, all the time. And anyone you know or might be married to who’s not deeply in tech just gets this long-suffering attitude on all of it. It’s nice to be able to have varied conversations about different things.

Carla: Yes. And so, I was like having all these, like—I was, like, having these life moments that felt really good, and then I would go to an audition and I would leave being, like, “Why do I do that to myself? Why do I need to feel like that?” Because I just feel awful every time I go. And so, then I was having trouble teaching my students because I was feeling really negative about it, and I was like, “I don’t know how to encourage you to go into a business that’s just going to, like, tear you down and make you feel awful about yourself all the time.”

Corey: And then you got into tech?

Carla: [laugh]. And then I was just, like, “Tech. That’s great.” No, I—do you know what—

Corey: Like, “I’m sad all the time and I feel like less than constantly. You know what I’m going to use to fix that? I’m going to learn JavaScript.” Oh, my God.

Carla: Yeah. I’m going to just challenge myself and do the hardest thing I can think of because that’s fun. But ki—I mean, sort of I [laugh] I, I was not ever—like, being an engineer was never, like, on my radar. My dad was an engineer for a long time, and he kind of always would be, like, “You’re good at math. You should do engineering.”

And I was like, “No, I’m an actor. [laugh]. I don’t want to do that.” And so, I kind of always just, like, shooed it away. And when a friend of mine came to my birthday party in the summer of 2018, who had been a songwriter and I had done some readings of a musical of his, and he was like, “I’m an engineer now at Forbes. Isn’t that great?”

And I was like, “What? How does that happen? I need you to back up, explain to me what’s going on.” And I just, like—but I went home and I could not stop thinking about it. I don’t know if it was like my dad’s voice in the back of my head, or there was like the stars aligned.

My misery that I was feeling in my life, and, like, this new thing that just got thrown in my face was just such an exciting, interesting idea. I was like, “That sounds—I don’t know what—I don’t even know what that looks like or I don’t even know what’s involved in that, but I need to figure out how to do it.” And I went home when I first started teaching myself how to do it. And I would just sit on my couch and I would do, like, little coding challenges, and before I knew it, like, hours would have passed by, I forgot to eat, I forget to go to the bathroom. Like, I would just be, like, groove on the couch from where I was sitting for too long.

And I was like, oh, I guess I really liked this. [laugh]. It’s interesting, it’s creative. Maybe I should do something with it.

Corey: And then from there, did you decide at some point to pursue—like, a lot of paths into tech these days. There’s a whole sea of boot camps, for example, that depending on how you look at them are either inspirational stories of how people can transform their lives, slash money-grabbing scams. And it really depends on the boot camp in particular, is that the path you took? Did you—

Carla: Yes.

Corey: Remain self-taught? How did you proceed from—there’s a whole Couch-to-5k running program; what is about—I guess we’ll call getting to tech—but what was your Couch-to-100k path?

Carla: Yeah, I was just going to say, Couch-to-100k tech gig.

Corey: Yeah.

Carla: So, my friend to had gone to Flatiron School, which is a boot camp. I think they have a few locations around the country, and so I initially started looking at their program just because he had gone there, and it sounded great. And I was like, “Cool, great.” And they had a lot of free resources online. They have, like, this whole free, like, boot camp prep program that you can do that teaches Rails and JavaScript.

And so, I started doing that online. And then I—at the time, they had, like, a part-time class. I like learning in person, which is funny because now I just work remote and I do everything on Google… it’s like, Google and Stack Overflow. So—but I knew at the time—

Corey: I have bad news about the people who are senior. It doesn’t exactly change that much.

Carla: Yeah, that’s what I’ve heard, so I don’t feel bad about telling people that I do it. [laugh].

Corey: We’re all Full Stack Overflow developers. It happens.

Carla: Exactly. So yeah, I just. They had, like, a part-time front-end class that was, like, in person two nights a week for a couple months. And I was like, “Okay, that’ll be a really good way to kind of get my feet wet with, like, a different kind of learning environment.”

And I loved it. I fell in love with it. I loved being in a room of people trying to figure out how to do something hard. I liked talking about it with other people. I liked talking about it with my teachers.

So, I was like, “Okay, I guess I’m going to invest in a boot camp.” And I did their, like, immersive, in-person boot camps. This was 2019 before everything shut down, so I was able to actually do it in person. And it was great. It was like, nine to six, five days a week, and it was really intense.

Did I remember everything I learned when it was over? No. And did I have to, like, spend a lot of time relearning a lot of things just so I could have, like, a deeper understanding of it. Yes. But, like, I also knew that was part of it, you know? It’s like, you throw a lot of information out you, hope some of it sticks, and then it’s your job to make sure that you actually remember it and then know how to use it when you have to.

Corey: One of the challenges that I’ve always found is that when I have a hobby that I’m into, similar to the way that you were doing this just for fun on your couch, and then it becomes your full-time focus, first as a boot camp and later as a job, that it has a tendency in some cases to turn a thing that you love into a thing that you view is this obligation or burden. Do you still love it? Is it still something that you find that’s fun and challenging and exciting? Or is it more a means to an end for you? And there is no wrong answer there.

Carla: Yeah, I think it’s a little bit of both, right? Like, I found it was a creative thing I could do that I enjoy doing. Am I the most passionate software engineer that ever lived? No. Do I have aspirations to be, like, an architect one day? Absolutely not. I really, like, the small tickets that I do that are just, like, refactoring a button or, you know, like, I find that stuff creative and I think it’s fun. Do I necessarily want to—

Corey: You can see—

Carla: —no.

Corey: The results immediately as [crosstalk 00:15:15]—

Carla: Yeah.

Corey: More abstract stuff. It’s like, “Well, when this 18 months migration finishes, and everything is 10% faster, oh, then I’ll be vindicated.”

Carla: Yeah. No.

Corey: It’s a little more attenuated from the immediate feedback.

Carla: Yeah. I’m not that kind of developer, I’m learning. But I’m totally fine with that. I have no issue. Like, I am a very humble person about it. I don’t have aspirations to be amazing.

Don’t ask me to do algorithm challenges. I’m terrible at them. I know that I’m terrible at them. But I also know that you can be a good developer and be terrible algorithm, like, challenges. So, I don’t feel bad about it.

Corey: The algorithm challenge is inherently biased for people who not only have a formal computer science education but have one relatively recently. I look back at some of the technical challenges I used to give candidates and take myself for jobs ten years ago, and I don’t remember half of it because it’s not my day-to-day anymore. It turns out that most of us don’t have a job implementing quicksort. We just use the one built into the library and we move on with our lives to do something interesting and much more valuable, like, moving that button three pixels left, but because of CSS, that’s now a two-week project.

Carla: Yeah. Add a little border-radius, changes the su—you know. There are some database things I like. You know, I’m trying to get better at SQL. Rails is really nice because we use Active Record, and I don’t really have to know SQL.

But I find there are some things that you can do in Rails that are really cool, and I enjoyed working in their console. And that’s exciting. You know when you write, like, a whole controller and then you make something but you can only see it in the console? That’s cool. I think to me, that’s fun. Being able to, like, generate things is fun. I don’t have to always see them, like, on the page in a visual, pretty way, even though I tend to be more visual.

Corey: This episode is sponsored in parts by our friend EnterpriseDB. EnterpriseDB has been powering enterprise applications with PostgreSQL for 15 years. And now EnterpriseDB has you covered wherever you deploy PostgreSQL on premises, private cloud, and they just announced a fully managed service on AWS and Azure called BigAnimal, all one word.

Don't leave managing your database to your cloud vendor because they're too busy launching another half dozen manage databases to focus on any one of them that they didn't build themselves. Instead, work with the experts over at EnterpriseDB. They can save you time and money, they can even help you migrate legacy applications, including Oracle, to the cloud.

To learn more, try BigAnimal for free. Go to biganimal.com/snark, and tell them Corey sent you.

Corey: One of the big fictions that we tend to have as an industry is when people sit down and say, “Oh, so why did you get into tech?” And everyone expects it to be this aspirational story of the challenge, and I’ve been interested in this stuff since I was a kid. And we’re all supposed to just completely ignore the very present reality of well, looking at all of my different opportunities, this is the one that pays three times what the others do. Like, we’re supposed to pretend that money doesn’t matter and we’re all following our passion. That is actively ridiculous from where I sit.

Carla: Mm-hm.

Corey: Do you find that effectively going from the Broadway actress side of the world to—where, let’s be clear, in the world of entertaining and arts—to my understanding—90% of people in that space are not able to do that as their only gig without side projects to basically afford to eat, whereas in tech, the median developer makes an extremely comfortable living that significantly outpaces the average median income for a family of four in the United States. Do you find that it has changed your philosophy on life in any meaningful way?

Carla: Oh, my God, yeah. I love talking about on all of my social platforms the idea that you can learn tech skills and you can—like, there are so many different jobs that exist for an engineer, right? There are full-time jobs. There are full-time job that are flexible and they’re remote, and nobody cares what time you’re working as long as you get the work done. And because of that and because of the nature of how performing and being an artist works, where you also have a lot of downtime in between jobs or even when you are working, that I feel like the two go very, very well together, and that it allows—if an artist can spend a little bit of time learning the skill, they now have the ability to feel stable in their lives, also be creative how they want to, and decide what the art looks like for them without struggling and freaking out all the time about where’s my next meal going to come from, or can I pay my rent?

And, like, I sometimes think back to when I was on tour—I was on tour for three years with Wicked—and I had so much free time, Corey. Like, if I had known that I could have spent some time when I was just like hanging out in my hotel room watching TV all day, like, learning how to code. I would have been—I would have done this years ago. If I had known it was even, I don’t even know actually if it was an option back then in, like, the early-2010s. I feel like boot camps kind of started around then, but they were mostly in person.

But if I was—today, if I was right now starting my career as an artist, I would absolutely learn how to code as a side hustle. Because why wait tables? [laugh]. Why make, like, minimum wage in a terrible job that you hate when you can I have a skillset that you can do from home now because everything is remote for the most part? Why not?

It doesn’t make sense to me that anybody would go back to those kind of awful side gigs, side hustle jobs. Because at the end of the day, side hustle jobs end up actually being the things that you spend more time doing, just because theater jobs and art jobs and music jobs are so, you know, far apart when you have them. That might as well pick something that’s lucrative and makes you feel less stressed out, you know, in the interim, between gigs. I see it as kind of a way to give artists a little more freedom in what they can choose to do with their art. Which I think is… it’s kind of magical, right?

Like, it takes away that narrative of if you can’t see yourself—if you can see yourself doing anything else, you should do it, right? That’s what we tell kids when they go into the arts. If you can see yourself doing any other thing, you know, you have to struggle to be an artist; that is part of the gig. That’s what you sign up for. And I just call bullshit on it, Corey. I don’t know if I can swear on this, but I call bullshit on [crosstalk 00:21:06]—[laugh].

Corey: Oh, you absolutely can.

Carla: I just think it’s so unfair to young people, to how they get to view themselves and their creativity, right? Like, you literally stunt them when you tell them that. You say, “You can only do this one thing.” That’s like the opposite of creative, right? That’s like telling somebody that they can only do one thing without imagining that they can do all these other things. The most interesting artists that I know do, like, 400 things, they are creative people and they can’t stop, right? They’re like multi-hyphenates [crosstalk 00:21:39].

Corey: It feels like it’s setting people up for failure, on some level, in a big way where when you’re building your entire life toward this make-or-break thing and then you don’t get it, it’s, well, what happens then?

Carla: Yeah.

Corey: I’ve always liked the idea of failure as a step forward. And well, that thing didn’t work out; let’s see if we can roll into it and see what comes out next. It’s similar to the idea of a lot of folks who are career-changing, where they were working somewhere else in a white-collar environment, well time to go back to square one for an entry-level world. Hell with that. Pivot; take a half step toward what you want to be doing in your next role, and then a year or so later, take the other half step, and now you’re doing it full time without having to start back at square one.

I think that there are very few things in this world that are that binary as far as you either succeed or you’re done and your whole life was a waste. It is easy get stuck in this idea that if your childhood dream doesn’t come true, well give up and prepare for a life of misery. I just don’t accept that.

Carla: Yeah, I—

Corey: But maybe it’s because I have no choice because getting fired is my stock-in-trade. So, it wasn’t until I built a company where I can’t get fired from it that I really started to feel a little bit secure in that. But it does definitely leave its marks and its damages. I spent 12 years waiting for the surprise meeting with my boss and someone I didn’t recognize from HR where they don’t offer you coffee—that’s always the tell when they don’t offer coffee—and to realize it while I’m back on the job market again; time to find something new. It left me feeling more mercenary that I probably should have, which wasn’t great for the career.

What about you? Do you think that—did it take, on some level, a sense of letting go of old dreams? Was it—and did it feel like a creeping awareness that this was, like—that you felt almost cornered into it? Or how did you approach it?

Carla: Yeah, I think I was the same way. I think I especially when you were younger because of that narrative, right, we tell people that if they decide to go into the arts, they have to be one hundred percent committed to it, and if they aren’t one hundred percent and then they don’t succeed, it is their fault, right? Like, if you give it everything that you have, and then it doesn’t work out, you have clearly done something wrong, therefore you are a failure. You failed at your dream because you gave it everything that you have, so you kind of set yourself up for failure because you don’t allow yourself to, you know, be more of who you are in other ways.

For me, I just spent so many I had so many moments in my life where I thought that the world was over, right? Like, when I was—right out of college, I went to school to study opera. And I was studying at Cincinnati Conservatory of Music, it was, like, the great, great conservatory, and halfway through my freshman year, I got diagnosed with a cyst on my vocal cords. So, basically what this meant was that I had to have surgery to have it removed, and the doctor told me that I probably would never sing opera. And I was devastated.

Like, I was—this was the thing I wanted to do with my life; I had committed myself one hundred percent, and now all of a sudden this thing happened, and I panicked. I thought it was my fault—because there was nobody to help me understand that it wasn’t—and I was like, “I have failed this thing. I have failed my dream. What am I going to do with my life?” And I said, “Okay I’ll be an actor because acting is a noble thing.” And that’s sort of like act—that’s sort of like performing; it’s performing in a different way, it’s just not singing.

And I was terrified to sing again because I had this narrative in my head that I was a failed singer if I co,uldn’t be an opera singer. And so, it took me, like, years, three years before I finally started singing again I got a voice teacher, and he—I would cry through all of my lessons. He was like, “Carly, you really have a—should be singing. Like, this is something that you’re good at.” And I was like, no because if I can’t sing, like, the way I want to sing, why would I sing?

And he really kind of pushed me and helped me, like, figure out what my voice could do in a new way. And it was really magical for me. It made me realize that this narrative that I’ve been telling myself of what I thought that I was supposed to be didn’t have to be true. It didn’t have to be the only one that existed; there could be other possibilities for what I could do and they could look different. But I closed myself off to that idea because I had basically been told no, you can’t do this thing that you want to do.

So, I didn’t even consider the possibilities of the other things that I could do. And when I relearned how to sing, it just blew my mind because I was like, “Oh, my God, I didn’t know this was possible. I didn’t know in my body it was possible of this. I didn’t know if I could do this.” And, like, overcoming that and making me realize that I could do other things, that there were other versions of what I wanted, kind of blew my mind a little bit.

And so, when I would hit road bumps and I’d hit these walls, I was like, “Okay, well, maybe I just need to pivot. Maybe the direction I’m going in isn’t quite the right one, but maybe if I just, like, open my eyes a little bit, there’s another—there’s something else over here that is interesting and will be creative and will take me in a different way, an unexpected way that I wasn’t expecting.” And so, I’ve kind of from that point on sort of living my life like that, in this way that, well, this might be a roadblock, and many people might view this thing as a failure, but for me, it allowed me to open up all these other new things that I didn’t even know I could do, right? Like, what I’m doing now is something I never would have imagined I’d be doing five years ago. And now I’m also in a place where not only am I doing something completely different as a software engineer, but I have this incredible opportunity to also start incorporating art back into my life in a way that I can own and I can do for myself instead of having to do for other people.

Which is also something I never thought because I thought it was all or nothing. I thought if I was an artist, I was an artist; I’m a software engineer, I’m a software engineer. And so, now I have the ability to kind of live in this weird gray area of getting to make those decisions for myself, and recognize that those little failures were, you know—like, I like to call them, like, the lowercase failures instead of the uppercase failure, right? Like, I am not a failure because I experienced failure. Those little failures are kind of what led me to grow my strength and my resilience and my ability to recognize it more free—like, more quickly when I see it so that I can bounce back faster, right?

Like, when I hit a wall, instead of living in that feeling of, like, “Ugh, God, this is the worst thing that ever happened,” I allow myself to move faster through it and recognize that there will be light on the other side. I will get there. And I know that it’s going to be okay, and I can trust that because it’s always been okay. I always figure it out. And so, that’s something—taken me a long time to, like, realize, you know? To, like, really learn, you have to fail a lot to learn that you’re going to be okay every time it happens. [laugh].

Corey: Yeah, what’s the phrase? “Sucking at something is the first step to being kind of good at it?”

Carla: Yeah. You got to let yourself suck at it. When I used to teach voice, I would make my students make just, like, the ugliest sounds because I was like, if we can just get past the fact that no matter what, when you sing you’re going to sound awful at some point. We’re going to try something, you’re going to crack, it’s not going to come out right, and if we can’t own that it’s going to suck a little bit on the journey to being good, like, you’re going to have a really hard time getting there because you’re just going to beat yourself up every time it sucks. Like, it’s going to suck a lot [laugh] before you get good. And that’s just part of it. That’s, like, it is just a part of the process, and you have to kind of own it.

Corey: I think that as people we are rarely as one-dimensional as we imagine we are when. And for example, I like working with cloud services, let’s not kid ourselves on this. But I have a deep and abiding love affair with the sound of my own voice, so I’m always going to find ways to work that into it. I have a hard time seeing a future career for you that does not in some way, shape or form, tie back to your performing background because even now, talking about singing, you lit up when talking about that in a way that no one does—or at least should—light up when they’re talking about React. So, do you think that there’s a place between the performing side of the world and the technical side of the world, or those phases of your life, that’s going to provide interesting paths for you down the road?

Carla: That is a good question, Corey. And I don’t know if I have the answer. You know, I think one thing—if there’s anything I learned from all the crazy things that happened to me, is that I just kind of have to be open. You know, I like to say yes to things. And also learning to say no, which has been really a big deal for me.

Corey: Oh, yes.“, no,” is a complete sentence and people know that sometimes at their own peril.

Carla: Yes, I have said no to some things lately, and it’s felt very good. But I like to be open, you know? I like to feel like if I’m putting out good things into the world, good things will come back to me, and so I’m just trying to keep that open. You know, I’m trying to be the best engineer that I can be. And I’m trying to also, you know—if I can use my voice and my platform to help inspire other people to see that there are other ways of being an artist, there are, you know, there are other paths in this world to take.

I hope that, you know, I can, other things will come up to me, there’ll be opportunities. And I don’t know what those look like, but I’m open. So, if anybody out there hears this and you want to collaborate, hit me up. [laugh].

Corey: Careful what you offer. People don’t know—people have a disturbing tendency of saying, “Well, all right, I have an idea.” That’s where a lot of my ridiculous parody music videos came from. It’s like, “So, what’s the business case for doing?” It’s like, “Mmm, I think it’ll be funny.”

It’s like, “Well, how are you going to justify the expense?” “Oh, there’s a line item and the company budget labeled ‘Spite.’ That’s how.” And it’s this weird combination of things that lead to a path that on some level makes perfect sense, but at the time you’re building this stuff out, it feels like you’re directionless and doing all these weird things. Like, one of the, I guess, strange parts of looking back at a path you take in the course of your career is, in retrospect, it feels like every step for the next was obvious and made intuitive sense, but going through it it’s, “I have no idea what I’m doing. I’m like the dog that caught the car, and they need to desperately figure out how to drive the thing before it hits the wall.”

It’s just a—I don’t pretend to understand how the tapestry of careers tie together, but I do know that I’m very glad to see people in this space, who do not all have the same ridiculous story for how they got in here. That’s the thing that I find continually obnoxious, this belief that there’s only one way to do it, or you’re somehow less than because you didn’t grow up programming in the ’90s. Great. There’s a lot of people like that. And yes, it is okay to just view computers as a job that pays the bills; there is nothing inherently wrong with that.

Carla: Yeah. And I mean, and I—

Corey: I just wish people were told that early on.

Carla: Yeah, why not? Right? Why didn’t anybody tell us that? Like, you don’t—the thing that I did not—it took me a long time to realize is that you do not have to be passionate about your job. And that’s like, that’s okay, right? All you have to do is enjoy it enough to do it, but it does not have to be, like—

Corey: You have to like it, on some level [crosstalk 00:33:10]—

Carla: Yeah, you just do have to like it. [laugh].

Corey: —dreading the 40 hours a week, that’s a miserable life on some level.

Carla: Like, I sit in front of a computer now all day, and I enjoy it. Like, I enjoy what I’m doing. But again, like, I don’t need to be the greatest software engineer that ever lived; I have other things that I like to do, and it allows me to also do those things. And that is what I love about it. It allows me that ability to just enjoy my weekends and have a stable career and have a stable life and have health insurance. And then when I want—

Corey: Oh, the luxuries of modern life.

Carla: [laugh]. Yeah, the luxuries of modern life. Health insurance, who knew? Yeah, you know, so it’s great. And then when creative projects come up, I can choose to say yes or no to them, and that’s really exciting for me.

Corey: I have a sneaking suspicion—I’ll just place my bet now—that the world of performing is not quite done with you yet.

Carla: Probably not. I would be lying if I said it was. I—so before all this stuff, I don’t know if your listeners know this, but in January, the thing that kind of happened to me that went a little viral where I went back to Broadway after not being on Broadway for a little while, and the news media and everybody picked up on it, and there were like these headlines of, “Software engineer plays Elphaba on Broadway after seven years.” It surprised me, but it also didn’t surprise me, you know? Like, when I left, I left thinking I was done.

And I think it was easy to leave when I left because of the pandemic, right? There was nothing going on when I—like, I started my journey before the pandemic, but I fully shifted into software engineering during the pandemic. So, I never had feelings of, like, “I’m missing out on performing,” because performing didn’t exist. There was no Broadway for a while. And so, once it kind of started to come back last year in the fall, I was like, “Oh, maybe I miss it a little bit.”

And maybe I accidentally manifested it, but, you know, when Wicked called and I flew back to New York for those shows, and I was like, “Oh, this is really wonderful.” Also, I’m really glad I don’t have to do this eight times a week. I’m so excited to go home. And I was like, having a little taste of it made me realize, “Oh, I can do this if I want to do this. I also don’t have to do this if I don’t want to do this.” And that was pretty—it was very empowering. I was like, “That feels nice.”

Corey: I really appreciate your taking so much time to talk about how you’ve gone through what at the time has got to have felt like a very strange set of career steps, but it’s starting to form into something that appears to have an arc to it. If people want to learn more and follow along as you continue to figure out what you’re going to do next, where’s the best place to find you?

Carla: Oh, good question, Corey. I do a website, carlastickler.com. Because I’ve had a lot of people—artists, in particular—reaching out and asking how I did this, I’m starting to build some resources, and so you can sign up for my mailing list.

I also am pretty big on Instagram if we’re going to choose social media. So, my Instagram is stiglercarla. And there’s links to all that stuff on my website. But—

Corey: And they will soon be in the [show notes 00:36:26] as well.

Carla: Ah yes, add them to the show notes. [laugh]. Yeah, and I want to make sure that I… I want—a lot of people who’ve seen my story and
felt very inspired by it. A lot of artists who have felt that they, too, were failures because they chose not to go into art and get a regular nine to five. And so, I’m trying to, like, kind of put a little bit more of that out there so that people see that they’re not alone.

And so, on my social media, I do post a lot of stories that people send to me, just telling me their story about how they made the transition and how they keep art in their life in different ways. And so, that’s something that also really inspires me. So, I tried to put their voices up, too. So, if anybody is interested in feeling not alone, feeling like there are other people out there, all of us, quote-unquote, “Failed artists,” and there’s a lot of us. And so, I’m just trying to create a little space for all of us.

Corey: I look forward to seeing it continue to evolve.

Carla: Thank you.

Corey: Thank you so much for your time. I appreciate it.

Carla: Thanks, Corey.

Corey: Carla Stickler, software engineer at G2 and also very much more. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, and if it’s on the YouTubes, smash the like and subscribe buttons, as the kids of today are saying, whereas if you’ve hated this podcast, same thing: Five-star review, smash the buttons,
but also leave an angry comment telling me exactly what you didn’t like about this, and I will reply with the time and date for your audition.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Simon

Founder and CEO of SnapShooter a backup company

Links Referenced:

  • SnapShooter.com: https://SnapShooter.com
  • MrSimonBennett: https://twitter.com/MrSimonBennett

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Finding skilled DevOps engineers is a pain in the neck! And if you need to deploy a secure and compliant application to AWS, forgettaboutit! But that’s where DuploCloud can help. Their comprehensive no-code/low-code software platform guarantees a secure and compliant infrastructure in as little as two weeks, while automating the full DevSecOps lifestyle. Get started with DevOps-as-a-Service from DuploCloud so that your cloud configurations are done right the first time. Tell them I sent you and your first two months are free. To learn more visit: snark.cloud/duplo. Thats’s snark.cloud/D-U-P-L-O-C-L-O-U-D.

Corey: What if there were a single place to get an inventory of what you're running in the cloud that wasn't "the monthly bill?" Further, what if there were a way to compare that inventory to what you were already managing via Terraform, Pulumi, or CloudFormation, but then automatically add the missing unmanaged or drifted parts to it? And what if there were a policy engine to immediately flag and remediate a wide variety of misconfigurations? Well, stop dreaming and start doing; visit snark.cloud/firefly to learn more.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. One of the things that I learned early on in my career as a grumpy Unix systems administrator is that there are two kinds of people out there: those who care about backups an awful lot, and people who haven’t lost data yet. I lost a bunch of data once upon a time and then I too fell on the side of backups are super important. Here to talk with me about them a bit today is Simon Bennett, founder and CEO of SnapShooter.com. Simon, thanks for joining me.

Simon: Thanks for having me. Thank you very much.

Corey: It’s fun to be able to talk to people who are doing business in the cloud space—in this sense too—that is not venture-backed, that is not, “Well, we have 600 people here that are building this thing out.” And similar to the way that I handle things at The Duckbill Group, you are effectively one of those legacy things known as a profitable business that self-funds. What made you decide to pursue that model as opposed to, well, whatever the polite version of bilking venture capitalists out of enormous piles of money for [unintelligible 00:01:32]?

Simon: I think I always liked the idea of being self-sufficient and running a business, so I always wanted to start a physical business when I was younger, but when I got into software, I realized that that’s a really easy way, no capital needed, to get started. And I tried for years and years to build products, all of which failed until finally SnapShooter actually gained a customer. [laugh].

Corey: “Oh, wait, someone finally is paying money for this, I guess I’m onto something.”

Simon: Yeah.

Corey: And it’s sort of progressed from there. How long have you been in business?

Simon: We started in 2017, as… it was an internal project for a company I was working at who had problems with DigitalOcean backups, or they had problems with their servers getting compromised. So, I looked at DigitalOcean API and realized I could build something. And it took less than a week to build a product [with billing 00:02:20]. And I put that online and people started using it. So, that was how it worked.

Every other product I tried before, I’d spent months and months developing it and never getting a customer. And the one time I spent less than [laugh] less than a week’s worth of evenings, someone started paying. I mean, admittedly, the first person was only paying a couple of dollars a month, but it was something.

Corey: There’s a huge turning point where you just validate the ability and willingness for someone to transfer one dollar from their bank account to yours. It speaks to validation in a way that social media nonsense generally doesn’t. It’s the oh, someone is actually willing to pay because I’m adding value to what they do. That’s no small thing.

Simon: Yeah. There’s definitely a big difference between people saying they’re going to and they’d love it, and actually doing it. So.

Corey: I first heard about you when Patrick McKenzie—or @patio11, as he goes by on Twitter—wound up doing a mini-thread on you about, “I’ve now used SnapShooter.com for real, and it was such a joy, including making a server migration easier than it would otherwise have been. Now, I have automatically monitored backups to my own S3 account for a bunch of things, which already had a fairly remote risk of failure.” And he keeps talking about the awesome aspects of it. And okay, when Patrick says, “This is neat,” that usually means it’s time for me to at least click the link and see what’s going on.

And the thing that jumped out at me was a few things about what it is that you offer. You talk about making sure that people can sleep well at night, that it’s about why backups are important, about—you obviously check the boxes and talk about how you do things and why you do them the way that you do, but it resonates around the idea of helping people sleep well at night. Because no one wants to think about backups. Because no one cares about backups; they just care an awful lot about restores, usually right after they should have cared about the backups.

Simon: Yeah. This is actually a big problem with getting customers because I don’t think it’s on a lot of people’s minds, getting backups set up until, as you said in the intro, something’s gone wrong. [laugh]. And then they’re happy to be a customer for life.

Corey: I started clicking around and looking at your testimonials, for example, on your website. And the first one I saw was from the CEO of Transistor.fm. For those who aren’t familiar with what they do, they are the company that hosts this podcast. I pay them as a vendor for all the
back issues and whatnot.

Whenever you download the show. It’s routing through their stuff. So yeah, I kind of want them to have backups of these things because I really don’t want to have all these conversations [laugh] again with everyone. That’s an important thing. But Transistor’s business is not making sure that the data is safe and secure; it’s making podcasts available, making it easy to publish to them.

And in your case, you’re handling the backup portion of it so they can pay their money and they set it up effectively once—set it and forget it—and then they can go back to doing the thing that they do, and not having to fuss with it constantly. I think a lot of companies get it wrong, where they seem to think that people are going to make sustained, engaged efforts in whatever platform or tool or service they build. People have bigger fish to fry; they just want the thing to work and not take up brain sweat.

Simon: Yeah. Customers hardly ever log in. I think it’s probably a good sign when they don’t have to log in. So, they get their report emails, and that’s that. And they obviously come back when they got new stuff to set up, but from a support point of view is pretty, pretty easy, really, people don’t—[laugh] constantly on there.

Corey: From where I sit, the large cloud providers—and some of the small ones, too—they all have backup functionality built into the offering that they’ve got. And some are great, some are terrible. I assume—perhaps naively—that all of them do what it says on the tin and actually back up the data. If that were sufficient, you wouldn’t have any customers. You clearly have customers. What is it that makes those things not work super well?

Simon: Some of them are inflexible. So, some of the providers have built-in server backups that only happen weekly, and six days of no backups can be a big problem when you’ve made a mistake. So, we offer a lot of flexibility around how often you backup your data. And then another key part is that we let you store your data where you want. A lot of the providers have either vendor lock-in, or they only store it in themselves. So… we let you take your data from one side of the globe to the other if you want.

Corey: As anyone who has listened to the show is aware, I’m not a huge advocate for multi-cloud for a variety of excellent reasons. And I mean that on a per-workload basis, not, “Oh, we’re going to go with one company called Amazon,” and you use everything that they do, including their WorkMail product. Yeah, even Amazon doesn’t use WorkMail; they use Exchange like a real company would. And great, pick the thing that works best for you, but backups have always been one of those areas.

I know that AWS has great region separation—most of the time. I know that it is unheard of for there to be a catastrophic data loss story that transcends multiple regions, so the story from their side is very often, oh, just back it up to a different region. Problem solved. Ignoring the data transfer aspect of that from a pricing perspective, okay. But there’s also a risk element here where everyone talks about the single point of failure with the AWS account that it’s there, people don’t talk about as much: it’s your payment instrument; if they suspend your account, you’re not getting into any region.

There’s also the story of if someone gets access to your account, how do you back that up? If you’re going to be doing backups, from my perspective, that is the perfect use case, to put it on a different provider. Because if I’m backing up from, I don’t know, Amazon to Google Cloud or vice versa, I have a hard time envisioning a scenario in which both of those companies simultaneously have lost my data and I still care about computers. It is very hard for me to imagine that kind of failure mode, it’s way out of scope for any disaster recovery or business continuity plan that I’m coming up with.

Simon: Yeah, that’s right. Yeah, I haven’t—[laugh] I don’t have that in my disaster recovery plan, to be honest about going to a different cloud, as in, we’ll solve that problem when it happens. But the data is, as you say, in two different places, or more. But yeah, the security one is a key one because, you know, there’s quite a lot of surface area on your AWS account for compromising, but if you’re using either—even a separate AWS account or a different provider purely for storage, that can be very tightly controlled.

Corey: I also appreciate the idea that when you’re backing stuff up between different providers, the idea of owning both sides of it—I know you offer a solution where you wind up hosting the data as well, and that has its value, don’t get me wrong, but there are also times, particularly for regulated industries, where yeah, I kind of don’t want my backup data just hanging out with someone else’s account with whatever they choose to do with it. There’s also the verification question, which again, I’m not accusing you of in any way, shape, or form of being nefarious, but it’s also one of those when I have to report to a board of directors of like, “Are you sure that they’re doing what they say they’re doing?” It’s a, “Well, he seemed trustworthy,” is not the greatest answer. And the boards ask questions like that all the time. Netflix has talked about this where they backup a rehydrate-the-business level of data to Google Cloud from AWS, not because they think Amazon is going to disappear off the face of the earth, but because it’s easier to do that and explain it than having to say, “Well, it’s extremely unlikely and here’s why,” and not get torn to pieces by auditors, shareholders, et cetera. It’s the path of least resistance, and there is some validity to it.

Simon: Yeah, when you see those big companies who’ve been with ransomware attacks and they’ve had to either pay the ransom or they’ve literally got to build the business from scratch, like, the cost associated with that is almost business-ending. So, just one backup for their data, off-site [laugh] they could have saved themselves millions and millions of pounds. So.

Corey: It’s one of those things where an ounce of prevention is worth a pound of cure. And we’re still seeing that stuff continue to evolve and continue to exist out in the ecosystem. There’s a whole host of things that I think about like, “Ooh, if I lost, that would be annoying but not disastrous.” When I was going through some contractual stuff when we were first setting up The Duckbill Group and talking to clients about this, they would periodically ask questions about, “Well, what’s your DR policy for these things?” It’s, “Well, we have a number of employees; no more than two are located in the same city anywhere, and we all work from laptops because it is the 21st century, so if someone’s internet goes out, they’ll go to a coffee shop. If everyone’s internet goes out, do you really care about the AWS bill that month?”

It’s a very different use case and [unintelligible 00:11:02] with these things. Now, let’s be clear, we are a consultancy that fixes AWS bills; we’re not a hospital. There’s a big difference in the use case and what is acceptable in different ways. But what I like is that you have really build something out that lets people choose their own adventure in how managed they want it to be, what the source is, what the target should be. And it gives people enough control but without having to worry about the finicky parts of aligning a bunch of scripts that wind up firing off in cron jobs.

Simon: Yeah. I’d say a fair few people run into issues running scripts or, you know, they silently fail and then you realize you haven’t actually been running backups for the last six months until you’re trying to pull them, even if you were trying to—

Corey: Bold of you to think that I would notice it that quickly.

Simon: [laugh]. Yeah, right. True. Yeah, that’s presuming you have a disaster recovery plan that you actually test. Lots of small businesses have never even heard of that as a thing. So, having as us, kind of, manage backups sort of enables us to very easily tell people that backups of, like—we couldn’t take the backup. Like, you need to address this.

Also, to your previous point about the control, you can decide completely where data flows between. So, when people ask us about what’s GDPR policies around data and stuff, we can say, “Well, we don’t actually handle your data in that sense. It goes directly from your source through almost a proxy that you control to your storage.” So.

Corey: The best answer: GDPR is out of scope. Please come again. And [laugh] yeah, just pass that off to someone else.

Simon: In a way, you’ve already approved those two: you’ve approved the person that you’re managing servers with and you’ve already approved the people that are doing storage with. You kind of… you do need to approve us, but we’re not handling the data. So, we’re handling your data, like your actual customer; we’re not handling your customer’s customer’s data.

Corey: Oh, yeah. Now, it’s a valuable thing. One of my famous personal backup issues was okay, “I’m going to back this up onto the shared drive,” and I sort of might have screwed up the backup script—in the better way, given the two possible directions this can go—but it was backing up all of its data and all the existing backup data, so you know, exponential growth of your backups. Now, my storage vendor was about to buy a boat and name it after me when I caught that. “Oh, yeah, let’s go ahead and fix that.”

But this stuff is finicky, it’s annoying, and in most cases, it fails in silent ways that only show up as a giant bill in one form or another. And not having to think about that is valuable. I’m willing to spend a few hours setting up a backup strategy and the rest; I’m not willing to tend it on an ongoing basis, just because I have other things I care about and things I need to get done.

Simon: Yeah. It’s such a kind of simple and trivial thing that can quickly become a nightmare [laugh] when you’ve made a mistake. So, not doing it yourself is a good [laugh] solution.

Corey: So, it wouldn’t have been a @patio11 recommendation to look at what you do without having some insight into the rest of the nuts and bolts of the business and the rest. Your plans are interesting. You have a free tier of course, which is a single daily backup job and half a gig of storage—or bring your own to that it’s unlimited storage—

Simon: Yep. Yeah.

Corey: Unlimited: the only limits are your budget. Yeah. Zombo.com got it slightly wrong. It’s not your mind, it’s your budget. And then it goes from Light to Startup to Business to Agency at the high end.

A question I have for you is at the high end, what I’ve found has been sort of the SaaS approach. The top end is always been a ‘Contact Us’ form where it’s the enterprise scope of folks where they tend to have procurement departments looking at this, and they’re going to have a whole bunch of custom contract stuff, but they’re also not used to signing checks with fewer than two commas in them. So, it’s the signaling and the messaging of, “Reach out and talk to us.” Have you experimented with that at all, yet? Is it something you haven’t gotten to yet or do you not have interest in serving that particular market segment?

Simon: I’d say we’ve been gearing the business from starting off very small with one solution to, you know, last—and two years ago, we added the ability to store data from one provider to a different provider. So, we’re sort of stair-stepping our way up to enterprise. For example, at the end of last year, we went and got certificates for ISO 27001 and… one other one, I can’t remember the name of them, and we’re probably going to get SOC 2 at some point this year. And then yes, we will be pushing more towards enterprises. We add, like, APIs as well so people can set up backups on the fly, or so they can put it as part of their provisioning.

That’s hopefully where I’m seeing the business go, as in we’ll become under-the-hood backup provider for, like, a managed hosting solution or something where their customers won’t even realize it’s us, but we’re taking the backups away from—responsibility away from businesses.

Corey: For those listeners who are fortunate enough to not have to have spent as long as I have in the woods of corporate governance, the correct answer to, “Well, how do we know that vendor is doing what they say that they’re doing,” because the, “Well, he seemed like a nice guy,” is not going to carry water, well, here are the certifications that they have attested to. Here’s copies under NDA, if their audit reports that call out what controls they claim to have and it validates that they are in fact doing what they say that they’re doing. That is corporate-speak that attests that you’re doing the right things. Now, you’re going to, in most cases, find yourself spending all your time doing work for no real money if you start making those things available to every customer spending 50 cents a year with you. So generally, the, “Oh, we’re going to go through the compliance, get you the reports,” is one of the higher, more expensive tiers where you must spend at least this much for us to start engaging down this rabbit hole of various nonsense.

And I don’t blame you in the least for not going down that path. One of these years, I’m going to wind up going through at least one of those certification approaches myself, but historically, we don’t handle anything except your billing data, and here’s how we do it has so far been sufficient for our contractual needs. But the world’s evolving; sophistication of enterprise buyers is at varying places and at some point, it’ll just be easier to go down that path.

Simon: Yeah, to be honest, we haven’t had many, many of those customers. Sometimes we have people who come in well over the plan limits, and that’s where we do a custom plan for them, but we’ve not had too many requests for certification. But obviously, we have the certification now, so if anyone ever [laugh] did want to see it under NDA, we could add some commas to any price. [laugh].

Corey: This episode is sponsored in parts by our friend EnterpriseDB. EnterpriseDB has been powering enterprise applications with PostgreSQL for 15 years. And now EnterpriseDB has you covered wherever you deploy PostgreSQL on premises, private cloud, and they just announced a fully managed service on AWS and Azure called BigAnimal, all one word.

Don't leave managing your database to your cloud vendor because they're too busy launching another half dozen manage databases to focus on any one of them that they didn't build themselves. Instead, work with the experts over at EnterpriseDB. They can save you time and money, they can even help you migrate legacy applications, including Oracle, to the cloud.

To learn more, try BigAnimal for free. Go to biganimal.com/snark, and tell them Corey sent you.

Corey: What I like as well is that you offer backups for a bunch of different things. You can do snapshots from, effectively, every provider. I’m sorry, I’m just going to call out because I love this: AWS and Amazon LightSail are called out as two distinct things. And Amazonians will say, “Oh, well, under the hood, they’re really the same thing, et cetera.” Yeah, the user experience is wildly different, so yeah, calling those things out as separate things make sense.

But it goes beyond that because it’s not just, “Well, I took a disk image. There we go. Come again.” You also offer backup recipes for specific things where you could, for example, back things up to a local file and external storage where someone is. Great, you also backup WordPress and MongoDB and MySQL and a whole bunch of other things.

A unified cloud controller, which is something I have in my house, and I keep thinking I should find a way to back that up. Yeah, this is great. It’s not just about the big server thing; it’s about having data living in managed services. It’s about making sure that the application data is backed up in a reasonable, responsible way. I really liked that approach. Was that an evolution or is that something you wound up focusing on almost from the beginning?

Simon: It was an evolution. So, we started with the snapshots, which got the business quite far to be honest and it was very simple. It was just DigitalOcean to start with, actually, for the first two years. Pretty easy to market in a way because it’s just focused on one thing. Then the other solutions came in, like the other providers and, you know, once you add one, it was easy to add many.

And then came database backups and file backups. And I just had those two solutions because that was what people were asking for. Like, they wanted to make sure their whole server snapshot, if you have a whole server snapshot, the point in time data for MySQL could be corrupt. Like, there could be stuff in RAM that a MySQL dump would have pulled out, for example. Like… there’s a possibility that the database could be corrupt from a snapshot, so people were asking for a bit of, more, peace of mind with doing proper backups of MySQL.

So, that’s what we added. And it soon became apparent when more customers were asking for more solutions that we really needed to, like, step back and think about what we’re actually offering. So, we rebuilt this whole, kind of like, database engine, then that allowed us to consume data from anywhere. So, we can easily add more backup types. So, the reason you can see all the ones you’ve listed there is because that’s kind of what people have been asking for. And every time someone comes up with a new, [laugh], like, a new open-source project or database or whatever, we’ll add support, even ones I’ve never heard of before. When people ask for some weird file—

Corey: All it takes is just waiting for someone to reach out and say, hey, can you back this thing up, please?

Simon: Yeah, exactly, some weird file-based database system that I’ve never ever heard of. Yeah, sure. Just give us [laugh] a test server to mess around with and we’ll build, essentially, like, we use bash in the background for doing the backups; if you can stream the data from a command, we can then deal with the whole management process. So, that’s the reason why. And then, I was seeing in, like, the Laravel space, for example, people were doing MySQL backups and they’d have a script, and then for whatever reason, someone rotated the passwords on the database and the backup script… was forgotten about.

So, there it is, not working for months. So, we thought we could build a backup where you could just point it at where the Laravel project is. We can get all the config we need at the runtime because it’s all there with the project anyway, and then thus, you never need to tell us the password for your database and that problem goes away. And it’s the same with WordPress.

Corey: I’m looking at this now just as you go through this, and I’m a big believer in disclaiming my biases, conflicts of interest, et cetera. And until this point, neither of us have traded a penny in either direction between us that I’m ever aware of—maybe you bought a t-shirt or something once upon a time—but great, I’m about to become a customer of this because I already have backup solutions for a lot of the things that you currently support, but again, when you’re a grumpy admin who’s lost data in the past, it’s, “Huh, you know what I would really like? That’s right, another backup.” And if that costs me a few hundred bucks a year for the peace of mind is money well spent because the failure mode is I get to rewrite a whole lot of blog posts and re-record all podcasts and pay for a whole bunch of custom development again. And it’s just not something that I particularly want to have to deal with. There’s something to be said for a holistic backup solution. I wish that more people thought about these things.

Simon: Can you imagine having to pull all the blog posts off [unintelligible 00:22:19]? [laugh]—

Corey: Oh, my got—

Simon: —to try and rebuild it.

Corey: That is called the crappiest summer internship someone has ever had.

Simon: Yeah.

Corey: And that is just painful. I can’t quite fathom having to do that as a strategy. Every once in a while some big site will have a data loss incident or go out of business or something, and there’s a frantic archiving endeavor that happens where people are trying to copy the content out of the Google Search Engine’s cache before it expires at whatever timeline that is. And that looks like the worst possible situation for any sort of giant backup.

Simon: At least that’s one you can fix. I mean, if you were to lose all the payment information, then you’ve got to restitch all that together, or anything else. Like, that’s a fixable solution, but a lot of these other ones, if you lose the data, yeah, there’s no two ways around it, you’re screwed. So.

Corey: Yeah, it’s a challenging thing. And it’s also—the question also becomes one of, “Well, hang on. I know about backups on this because I have this data, but it’s used to working in an AWS environment. What possible good would it do me sitting somewhere else?” It’s, yeah, the point is, it’s sitting somewhere else, at least in my experience. You can copy it back to that sort of environment.

I’m not suggesting this is a way that you can run your AWS serverless environment on DigitalOcean, but it’s a matter of if everything turns against you, you can rebuild from those backups. That’s the approach that I’ve usually taken. Do you find that your customers understand that going in or is there an education process?

Simon: I’d say people come for all sorts of reasons for why they want backup. So, having your data in two places for that is one of the reasons but, you know, I think there’s a lot of reasons why people want peace of mind: for either developer mistakes or migration mistakes or hacking, all these things. So, I guess the big one we come up with a lot is people talking about databases and they don’t need backups because they’ve got replication. And trying to explain that replication between two databases isn’t the same as a backup. Like, you make a mistake you drop—[laugh] you run your delete query wrong on the first database, it’s gone, replicated or not.

Corey: Right, the odds of me fat-fingering an S3 bucket command are incredibly likelier than the odds of AWS losing an entire region’s S3 data irretrievably. I make mistakes a lot more than they tend to architecturally, but let’s also be clear, they’re one of the best. My impression has always been the big three mostly do a decent job of this. The jury’s still out, in my opinion, on other third-party clouds that are not, I guess, tier one. What’s your take?

Simon: I have to be careful. I’ve got quite good relationships with some of these. [laugh].

Corey: Oh, of course. Of course. Of course.

Simon: But yes, I would say most customers do end up using S3 as their storage option, and I think that is because it is, I think, the best. Like, is in terms of reliability and performance, some storage can be a little slow at times for pulling data in, which could or could not be a problem depending on what your use case is. But there are some trade-offs. Obviously, S3, if you’re trying to get your data back out, is expensive. If you were to look at Backblaze, for example, as well, that’s considerably cheaper than S3, especially, like, when you’re talking in the petabyte-scale, there can be huge savings there. So… they all sort of bring their own thing to the table. Personally, I store the backups in S3 and in Backblaze, and in one other provider. [laugh].

Corey: Oh, yeah. Like—

Simon: I like to have them spread.

Corey: Like, every once in a while in the industry, there’s something that happens that’s sort of a watershed moment where it reminds everyone, “Oh, right. That’s why we do backups.” I think the most recent one—and again, love to them; this stuff is never fun—was when that OVH data center burned down. And OVH is a somewhat more traditional hosting provider, in some respects. Like, their pricing is great, but they wind up giving you what amounts to here as a server in a rack. You get to build all this stuff yourself.

And that backup story is one of those. Oh, okay. Well, I just got two of them and I’ll copy backups to each other. Yeah, but they’re in the same building and that building just burned down. Now, what? And a lot of people learned a very painful lesson. And oh, right, that’s why we have to do that.

Simon: Yeah. The other big lesson from that was that even if the people with data in a different region—like, they’d had cross-regional backups—because of the demand at the time for accessing backups, if you wanted to get your data quickly, you’re in a queue because so many other people were in the same boat as you’re trying to restore stored backups. So, being off-site with a different provider would have made that a little easier. [laugh].

Corey: It’s a herd of elephants problem. You test your DR strategy on a scheduled basis; great, you’re the only person doing it—give or take—at that time, as opposed to a large provider has lost a region and everyone is hitting their backup service simultaneously. It generally isn’t built for that type of scale and provisioning. One other question I have for you is when I make mistakes, for better or worse, they’re usually relatively small-scale. I want to restore a certain file or I will want to, “Ooh, that one item I just dropped out of that database really should not have been dropped.” Do you currently offer things that go beyond the entire restore everything or nothing? Or right now are you still approaching this from the perspective of this is for the catastrophic case where you’re in some pain already?

Simon: Mostly the catastrophic stage. So, we have MySQL [bin logs 00:27:57] as an option. So, if you wanted to do, like, a point-in-time of store, which… may be more applicable to what you’re saying, but generally, its whole, whole website recovery. For example, like, we have a WordPress backup that’ll go through all the WordPress websites on the server and we’ll back them up individually so you can restore just one. There are ways that we have helped customers in the past just pull one table, for example, from a backup.

But yeah, we geared towards, kind of, the set and the forget. And people don’t often restore backups, to be honest. They don’t. But when they do, it’s obviously [laugh] very crucial that they work, so I prefer to back up the whole thing and then help people, like, if you need to extract ten megabytes out of an entire gig backup, that’s a bit wasteful, but at least, you know, you’ve got the data there. So.

Corey: Yeah. I’m a big believer in having backups in a variety of different levels. Because I don’t really want to do a whole server restore when I remove a file. And let’s be clear, I still have that grumpy old Unix admin of before I start making changes to a file, yeah, my editor can undo things and remembers that persistently and all. But I have a disturbing number of files and directories whose names end in ‘.bac’ with then, like, a date or something on it, just because it’s—you know, like, “Oh, I have to fix something in Git. How do I do this?”

Step one, I’m going to copy the entire directory so when I make a pig’s breakfast out of this and I lose things that I care about, rather than having to play Git surgeon for two more days, I can just copy it back over and try again. Disk space is cheap for those things. But that’s also not a holistic backup strategy because I have to remember to do it every time and the whole point of what you’re building and the value you’re adding, from my perspective, is people don’t have to think about it.

Simon: Yes. Yeah yeah yeah. Once it’s there, it’s there. It’s running. It’s as you say, it’s not the most efficient thing if you wanted to restore one file—not to say you couldn’t—but at least you didn’t have to think about doing the backup first.

Corey: I really want to thank you for taking the time out of your day to talk to me about all this. If people want to learn more for themselves, where can they find you?

Simon: So, SnapShooter.com is a great place, or on Twitter, if you want to follow me. I am @MrSimonBennett.

Corey: And we will, of course, put links to that in the [show notes 00:30:11]. Thank you once again. I really appreciate it.

Simon: Thank you. Thank you very much for having me.

Corey: Simon Bennett, founder and CEO of SnapShooter.com. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve
enjoyed this episode, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this episode, please leave a five-star review on your podcast platform of choice, along with an angry insulting comment that, just like your backup strategy, you haven’t
put enough thought into.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Wesley

Wesley Faulkner is a first-generation American, public speaker, and podcaster. He is a founding member of the government transparency group Open Austin and a staunch supporter of racial justice, workplace equity, and neurodiversity. His professional experience spans technology from AMD, Atlassian, Dell, IBM, and MongoDB. Wesley currently works as a Developer Advocate, and in addition, co-hosts the developer relations focused podcast Community Pulse and serves on the board for SXSW.

Links Referenced:

  • Twitter: https://twitter.com/wesley83
  • Polywork: https://polywork.com/wesley83
  • Personal Website: https://www.wesleyfaulkner.com/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Finding skilled DevOps engineers is a pain in the neck! And if you need to deploy a secure and compliant application to AWS, forgettaboutit! But that’s where DuploCloud can help. Their comprehensive no-code/low-code software platform guarantees a secure and compliant infrastructure in as little as two weeks, while automating the full DevSecOps lifestyle. Get started with DevOps-as-a-Service from DuploCloud so that your cloud configurations are done right the first time. Tell them I sent you and your first two months are free. To learn more visit: snark.cloud/duplo. Thats’s snark.cloud/D-U-P-L-O-C-L-O-U-D.

Corey: What if there were a single place to get an inventory of what you're running in the cloud that wasn't "the monthly bill?" Further, what if there were a way to compare that inventory to what you were already managing via Terraform, Pulumi, or CloudFormation, but then automatically add the missing unmanaged or drifted parts to it? And what if there were a policy engine to immediately flag and remediate a wide variety of misconfigurations? Well, stop dreaming and start doing; visit snark.cloud/firefly to learn more.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I am joined again for a second time this year by Wesley Faulkner. Last time we spoke, he was a developer advocate. And since then, as so many have, he’s changed companies. Wesley, thank you for joining me again. You’re the Head of Community at SingleStore, now. Congrats on the promotion.

Wesley: Thank you. It’s been a very welcome change. I love developer advocates and developer advocacy. But I love people, too, so it’s almost, I think, very analogous to the ebbs and flow that we all have gone through, through the pandemic, and leaning into my strong suits.

Corey: It’s a big deal having a ‘head of’ in a role title, as opposed to Developer Advocate, Senior Developer Advocate. And it is a different role. It’s easy to default into the world of thinking that it’s a promotion. Management is in many ways orthogonal to what it takes to succeed in an actual role. And further, you’re not the head of DevRel, or DevRelopers or whatever you want to call the term. You are instead the Head of Community. How tied is that to developer relations, developer advocacy, or other things that we are used to using as terms of art in this space?

Wesley: If we’re talking about other companies, I would say the Head of Community is something that’s under the umbrella of developer relations, where it’s just a peer to some of the other different elements or columns of developer relations. But in SingleStore specifically, I have to say that developer relations in terms of what you think about whole umbrella is very new to the company. And so, I consider myself the first person in the role of developer relations by being the Head of Community. So, a lot of the other parts are being bolted in, but under the focus of developer as a community. So, I’m liaisoning right now as helping with spearheading some of the design of the activities that the advocates do, as well as architecting the platform and the experiences of people coming in and experiencing SingleStore through the community’s perspective.

So, all that to say is, what I’m doing is extremely structured, and a lot of stuff that we’re doing with the efficacy, I’m using some of my expertise to help guide that, but it’s still something that’s kind of like an offshoot and not well integrated at the moment.

Corey: How has it changed the way that you view the function of someone who’s advocating to developers, which is from my cynical perspective, “Oh, it’s marketing, but we don’t tell people it’s marketing because they won’t like it.” And yes, I know, I’ll get emails about that. But how does it differ from doing that yourself versus being the head of the function of a company? Because leadership is a heck of a switch? I thought earlier in my career that oh, yeah, it’s a natural evolution of being a mediocre engineer. Time to be a mediocre manager. And oh, no, no, I aspired to be a mediocre manager. It’s a completely different skill set and I got things hilariously wrong. What’s it like for you going through that shift?

Wesley: First of all, it is kind of like advertising, and people may not think of it that way. Just to give an example, movie trailers is advertising. The free samples at the grocery store is advertising. But people love those because it gives an experience that they like in a package that they are accustomed to. And so, it’s the same with developer relations; it’s finding the thing that makes the experience worthwhile.

On the community side, this is not new to me. I’ve done several different roles, maybe not in this combination. But when I was at MongoDB, I was a technical community manager, which is like a cog in the whole giant machine. But before that, in my other life, I managed social and community interactions for Walmart, and I had, at the slow period, around 65, but during the holidays, it would ramp up to 95 direct reports that I managed.

It’s almost—if you’re a fan of The Princess Bride, it’s different than fighting one person. Sometimes it’s easier to fight, like, a squad or a gang of people. So, being Head of Community with such a young company is definitely a lot different than. In some ways, harder to deal with this type of community where we’re just growing and emerging, rather than something more well-established.

Corey: It probably gives you an interesting opportunity. Because back when I was doing engineering work as an SRE or whatever we call them in that era, it was, “Yeah, wow, my boss is terrible and has no idea what the hell they’re doing.” So, then I found myself in the role, and it’s, “Cool. Now, do all the things that you said you would do. Put up or shut up.”

And it turns out that there’s a lot you don’t see that our strategic considerations. I completely avoided things like managing up or managing laterally or balancing trade-offs in different ways. Yeah, you’re right. If you view the role of management as strictly being something that is between you and your direct reports, you can be an amazing manager from their perspective, but completely ineffective organizationally at accomplishing the goals that have been laid out for you.

Wesley: Yeah. The good thing about being head of and the first head of is that you help establish those goals. And so, when you take a role with another company saying, “Hey, we have headcount for this,” and it’s an established role, then you’re kind of like streamlining into a process that’s already underway. What’s good about this role specifically, a ‘head of,’ is that I help with not only designing what are the goals and the OKRs but deciding what the teams and what the team structure should look like. And so, I’m hiring for a specific position based on how it interacts with everything else.

So, when I’m coming in, I don’t say, “Well, what do you do?” Or, “How do you do it?” I said, “This is what needs to be done.” And that makes it so much easier just to say that if everything is working the way it should and to give marching orders based on the grand vision, instead of hitting the numbers this quarter or next quarter. Because what is core to my belief, and what’s core, too, of how I approach things is at the heart of what I’m trying to do, which is really great, in terms of making something that didn’t exist before.

Corey: The challenge, too, is that everyone loves to say—and I love to see this at different ways—is the evolution and understanding of the DevRel folks who I work with and I have great relationships with realizing that you have to demonstrate business value. Because I struggle with this my entire career where I know intrinsically, that if I get on stage and tell a story about a thing that is germane to what my company does, that good things are going to happen. But it’s very hard to do any form of attribution to it. In a different light, this podcast is a great example of this.

We have sponsors. And people are listening. Ideally, they aren’t fast-forwarding through sponsor messages; I do have interesting thoughts about the sponsors that I put into these ads. And that’s great, but I also appreciate that people are driving while they’re listening to this, and they are doing the dishes, they are mowing the lawn, and hopefully not turning up the volume too loudly so it damages their hearing. And the idea that they’re going to suddenly stop any of those things and go punch in the link that I give is a little out to lunch there.

Instead, it’s partially brand awareness and it is occasionally the, “Wait. That resonates exactly with the problem that I have.” So, they get to work or they get back in front of a computer and the odds are terrific they’re not going to punch in that URL of whatever I wound up giving; they’re going to type in whatever phrases they remember and the company name into Google. Now—and doing attribution on something like that is very hard.

It gets even more hard when we’re talking about something that is higher up the stack that requires a bit more buy-in than individual developers. There’s often a meeting or two about it. And then someone finally approaches the company to have a conversation. Now, does it work? Yes. There are companies that are sponsoring this stuff that spend a lot of time, effort, and money on that.

I don’t know how you do that sort of attribution; I don’t pretend to know, but I know that it works. Because these people whose entire job is making sure that it does tell me it does. So, I smile, I nod, and that’s great. But it’s very hard to wind up building out a direct, “If you spend X dollars sponsoring this, you will see Y dollars in response.” But in the DevOps world, when your internal doing these things, well, okay because to the company, I look an awful lot like an expensive developer except I don’t ever write production code.

And then—at least in the before times—“So, what does your job do? Because looking at the achievements and accomplishments last quarter, it looks an awful lot like you traveled to exotic places on the company dime, give talks that are of only vague relevance to what we do, and then hang out at parties with your friends? Nice job, how can I get that?” But it’s also first on the chopping block when okay, how do we trim expenses go? And I think it’s a mistake to do that. I just don’t think that story of the value of developer relations is articulated super-well. And I say that, but I don’t know how to do a much better job of it myself.

Wesley: Well, that’s why corporate or executive buy-in is important because if they know from the get-go while you’re there, it makes it a little bit easier to sell. But you do have to show that you are executing. So, there are always two parts to presenting a story, and that’s one, the actual quantitative, like, I’ve done this many talks—so that output part—I’ve written this many blog posts, or I’ve stood up this many events that people can attend to. And then there’s the results saying, people did read this post, people did show up to my event, people did listen to my talk that I gave. But you also need to give the subjective ones where people respond back and say, “I loved your talk,” or, “I heard you on Corey’s podcast,” or, “I read your blog posts,” because even though you might not understand that it goes all the way down in a conversion funnel to a purchase, you can least use that stand-in to say there’s probably, like, 20, 30 people behind this person to have that same sentiment, so you can see that your impact is reaching people and that it’s having some sort of lasting effect.

That said, you have to keep it up. You have to try to increase your output and increase your sphere of influence. Because when people go to solve their problem, they’re going to look into their history and their own Rolodex of saying what was the last thing that I heard? What was the last thing that’s relevant?

There is a reason that Pepsi and Coke still do advertising. It’s not because people don’t know those brands, but being easily recalled, or a center of relevance based on how many touchpoints or how many times that you’ve seen them, either from being on American Idol and the logo facing the camera, or seeing a whole display when you go into the grocery store. Same with display advertising. All of this stuff works hand in hand so that you can be front-of-mind with the people and the decision-makers who will make that decision. And we went through this through the pandemic where… that same sentiment, it was like, “You just travel and now you can’t travel, so we’re just going to get rid of the whole department.”

And then those same companies are hunting for those people to come back or to rebuild these departments that are now gone because maybe you don’t see what we do, but when it’s gone, you definitely notice a dip. And that trust is from the top-up. You have to do not just external advocacy, but you have to do internal advocacy about what impacts you’re having so that at least the people who are making that decision can hopefully understand that you are working hard and the work is paying off.

Corey: Since the last time that we spoke, you’ve given your first keynote, which—

Wesley: Yes.

Corey: Is always an interesting experience to go through. It was at a conference called THAT Conference. And I feel the need to specify that because otherwise, we’re going to wind up with a ‘who’s on first’ situation. But THAT Conference is the name.

Wesley: Specify THAT. Yes.

Corey: Exactly. Better specify THAT. Yes. So, what was your keynote about? And for a bit of a behind-the-scenes look, what was that like for you?

Wesley: Let me do the behind-the-scenes because it’s going to lead up to actual the execution.

Corey: Excellent.

Wesley: So, I’ve been on several different podcasts. And one of the ones that I loved for years is one called This Week in Tech with Leo Laporte. Was a big fan of Leo Laporte back in the Screen Saver days back in TechTV days. Loved his opinion, follow his work. And I went to a South by Southwest… three, four years ago where I actually met him.

And then from that conversation, he asked me to be on his show. And I’ve been on the show a handful of times, just talking about tech because I love tech. Tech is my passion, not just doing it, but just experiencing and just being on either side of creating or consuming. When I moved—I moved recently also since, I think, from the last time I was on your show—when I moved here to Wisconsin, the organizer of THAT Conference said that he’s been following me for a while, since my first appearance on This Week in Tech, and loved my outlook and my take on things. And he approached me to do a keynote.

Since I am now Wisconsin—THAT Conference is been in Wisconsin since inception and it’s been going on for ten years—and he wanted me to just basically share my knowledge. Clean slate, have enough time to just say whatever I wanted. I said, “Yes, I can do that.” So, my experience on my end was like sheer excitement and then quickly sheer terror of not having a framework of what I was going to speak on or how I was going to deliver it. And knowing as a keynote, that it would be setting the tone for the whole conference.

So, I decided to talk on the thing that I knew the most about, which was myself. Talked about my journey growing up and learning what my strengths, what my weaknesses are, how to navigate life, as well as the corporate jungle, and deciding where I wanted to go. Do I want to be the person that I feel like I need to be in order to be successful, which when we look at structures and examples and the things that we hold on a pedestal, we feel that we have to be perfect, or we have to be knowledgeable, and we have to do everything, well rounded in order to be accepted. Especially being a minority, there’s a lot more caveats in terms of being socially acceptable to other people. And then the other path that I could have taken, that I chose to take, was to accept my things that are seen as false, but my own quirkiness, my own uniqueness and putting that front and center about, this is me, this is my person that over the years has formed into this version of myself.

I’m going to make sure that is really transparent and so if I go anywhere, they know what they’re getting, and they know what they’re signing up for by bringing me on board. I have an opinion, I will share my opinion, I will bring my whole self, I won’t just be the person that is technical or whimsical, or whatever you’re looking for. You have to take the good with the bad, you have to take the I really understand technology, but I have ADHD and I might miss some deadlines. [laugh].

Corey: This episode is sponsored in parts by our friend EnterpriseDB. EnterpriseDB has been powering enterprise applications with PostgreSQL for 15 years. And now EnterpriseDB has you covered wherever you deploy PostgreSQL on premises, private cloud, and they just announced a fully managed service on AWS and Azure called BigAnimal, all one word.

Don't leave managing your database to your cloud vendor because they're too busy launching another half dozen manage databases to focus on any one of them that they didn't build themselves. Instead, work with the experts over at EnterpriseDB. They can save you time and money, they can even help you migrate legacy applications, including Oracle, to the cloud.

To learn more, try BigAnimal for free. Go to biganimal.com/snark, and tell them Corey sent you.

Corey: I have a very similar philosophy, and how I approach these things where it’s there is no single speaking engagement that I can fathom even being presented to me, let alone me accepting that is going to be worth me losing the reputation I have developed for authenticity. It’s you will not get me to turn into a shill for whatever it is that I am speaking in front of this week. Conversely, whether it’s a paid speaking engagement or not, I have a standing policy of not using a platform that is being given to me by a company or organization to make them look foolish. In other words, I will not make someone regret inviting me to speak at their events. Full stop.

And I have spoken at events for AWS; I have spoken at events for Oracle, et cetera, et cetera, and there’s no company out there that I’m not
going to be able to get on stage and tell an entertaining and engaging story, but it requires me to dunk on them. And that’s fine. Frankly, if there is a company like that where I could not say nice things about them—such as Facebook—I would simply decline to pursue the speaking opportunity. And that is the way that I view it. And very few companies are on that list, to be very honest with you.

Now, there are exceptions to this, if you’re having a big public keynote, I will do my traditional live-tweet the keynote and make fun of people because that is, A, expected and, B, it’s live-streamed anywhere on the planet I want to be sitting at that point in time, and yeah, if you’re saying things in public, you can basically expect that to be the way that I approach these things. But it’s a nuanced take, and that is something that is not fully understood by an awful lot of folks who run events. I’ll be the first to admit that aspects of who and what I am mean that some speaking engagements are not open to me. And I’m okay with that, on some level, I truly am. It’s a different philosophy.

But I do know that I am done apologizing for who I am and what I’m about. And at some point that required a tremendous amount of privilege and a not insignificant willingness to take a risk that it was going to work out all right. I can’t imagine going back anymore. Now, that road is certainly not what I would recommend to everyone, particularly folks earlier in their career, particularly for folks who don’t look just like I do and have a failure mode of a board seat and a book deal somewhere, but figuring out where you will and will not compromise is always an important thing to get straight for yourself before you’re presented with a situation where you have to make those decisions, but now there’s a whole bunch of incentive to decide in one way or another.

Wesley: And that’s a journey. You can’t just skip sections, right? You didn’t get to where you are unless you went through the previous experience that you went through. And it’s true for everyone. If you see those success books or how-to books written by people who are extremely rich, and, like, how to become successful and, like, okay, well, that journey is your own. It doesn’t make it totally, like, inaccessible to everyone else, but you got to realize that not everyone can walk that path. And—

Corey: You were in the right place at the right time, an early employee at a company that did phenomenally well and that catapulted you into reach beyond the wildest dreams of avarice territory. Good for you, but fundamentally, when you give talks like that as a result, what it often presents as is, “I won the lottery, and here’s how you can too.” It doesn’t work that way. The road you walked was unique to you and that opportunity is closed, not open anyone else, so people have to find their own paths.

Wesley: Yeah, and lightning doesn’t strike in the same place twice. But there are some things where you can understand some fundamentals. And depending on where you go, I think you do need to know yourself, you do need to know—like, be able to access yourself, but being able to share that, of course, you have to be at a point where you feel comfortable. And so, even if you’re in a space where you don’t feel that you can be your authentic self or be able to share all parts of you, you yourself should at least know yourself and then make that decision. I agree that it’s a point of privilege to be able to say, “Take me how I am.”

I’m lucky that I’ve gotten here, not everyone does, and just because you don’t doesn’t mean that you’re a failure. It just means that the world hasn’t caught up yet. People who are part of marginalized society, like, if you are, let’s say trans, or if you are even gay, you take the same person, the same stance, the same yearning to be accepted, and then transport it to 50 years ago, you’re not safe. You will not necessarily be accepted, or you may not even be successful. And if you have a lane where you can do that, all the power to you, but not everyone could be themselves, and you just need to make sure that at least you can know yourself, even if you don’t share that with the world.

Corey: It takes time to get there, and I think you’re right that it’s impossible to get there without walking through the various steps. It’s one of the reasons I’m somewhat reluctant to talk overly publicly about my side project gig of paid speaking engagements, for instance, is that the way to get those is you start off by building a reputation as a speaker, and that takes an awful lot of time. And speaking at events where there’s no budget even to pay you a speaking fee out of anyway. And part of what gets the keynote invitations to, “Hey, we want you to come and give a talk,” is the fact that people have seen you speak elsewhere and know what you’re about and what to expect. Here’s a keynote presented by someone who’s never presented on stage before is a recipe for a terrifying experience, if not for the speaker or the audience, definitely [laugh] for the event organizers because what if they choke.?

Easy example of this, even now hundreds of speaking engagements in, the adrenaline hit right before I go on stage means that sometimes my knees shake a bit before I walk out on stage. I make it a point to warn the people who are standing with me backstage, “Oh, this is a normal thing. Don’t worry, it is absolutely expected. It happens every time. Don’t sweat it.”

And, like, “Thank you for letting us know. That is the sort of thing that’s useful.” And then they see me shake, and they get a little skeptical. Like, I thought this guy was a professional. What’s the story and I walk on stage and do my thing and I come back. Like, “That was incredible. I was worried at the beginning.” “I told you, we all have our rituals before going on stage. Mine is to shake like a leaf.”

But the value there is that people know what to generally expect when I get on stage. It’s going to have humor, there’s going to be a point interwoven throughout what I tend to say, and in the case of paid speaking engagements, I always make sure I know where the boundaries are of things I can make fun of a big company for. Like, I can get on stage and make fun of service naming or I can make fun of their deprecation policy or something like that, but yeah, making fun of the way that they wind up handling worker relations is probably not going to be great and it could get the person who championed me fired or centered internally. So, that is off the table.

Like, even on this podcast, for example, I sometimes get feedback from listeners of, “Well, you have someone from company X on and you didn’t beat the crap out of them on this particular point.” It’s yeah, you do understand that by having people on the show I’m making a tacit agreement not to attack them. I’m not a journalist. I don’t pretend to be. But if I beat someone up with questions about their corporate policy, yeah, very rarely do I have someone who is in a position in those companies to change that policy, and they’re certainly not authorized to speak on the record about those things.

So, I can beat them up on it, they can say, “I can’t answer that,” and we’re not going to go anywhere. What is the value of that? It looks like it’s not just gotcha journalism, but ineffective gotcha journalism. It doesn’t work that way. And that’s never been what this show is about.

But there’s that consistent effort behind the scenes of making sure that people will be entertained, will enjoy what they’re seeing, but also are not going to deeply regret giving me a microphone, has always been the balancing act, at least for me. And I want to be clear, my style is
humor. It is not for everyone. And my style of humor has a failure mode of being a jerk and making people feel bad, so don’t think that my path is the only or even a recommended way for folks who want to get more into speaking to proceed.

Wesley: You also mention, though, about, like, punching up versus punching down. And if you really tear down a company after you’ve been invited to speak, what you’re doing is you’re punching down at the person who booked you. They’re not the CEO; they’re not the owner of the company; they’re the person who’s in charge of running an event or booking speakers. And so, putting that person and throwing them under the bus is punching down because now you’re threatening their livelihood, and it doesn’t make any market difference in terms of changing the corporate’s values or how they execute. So yeah, I totally agree with you in that one.

And, like you were saying before, if there’s a company you really thought was abhorrent, why speak there? Why give them or lend your reputation to this company if you absolutely feel that it’s something you don’t want to be associated with? You can just choose not to do that. For me, when I look at speaking, it is important for me to really think about why I’m speaking as well. So, not just the company who’s hiring me, but the audience that I’ll be serving.

So, if I’m going to help with inspiring the next generation of developers, or helping along the thought of how to make the world a better place, or how people themselves can be better people so that we can just change the landscape and make it a lot friendlier, that is also its own… form of compensation and not just speaking for a speaker’s fee. So, I do agree that you need to not just be super Negative Nancy, and try to fight all fights. You need to embrace some of the good things and try to make more of those experiences good for everyone, not just the people who are inviting you there, but the people who are attending. And when I started speaking, I was not a good speaker as well. I made a lot of mistakes, and still do, but I think speaking is easier than some people think and if someone truly wants to do it, they should go ahead and get started.

What is the saying? If there’s something is truly important, you’ll be bad at it [laugh] and you’ll be okay with it. I started speaking because of my role as a developer advocate. And if you just do a Google search for ‘CFPs,’ you can start speaking, too. So, those who are not public speakers and want to get into it, just Google ‘CFP’ and then start applying.

And then you’ll get better at your submissions, you’ll get better at your slides, and then once you get accepted, then you’ll get better at preparing, then you’ll get better at actually speaking. There’s a lot of steps between starting and stopping and it’s okay to get started doing that route. The other thing I wanted to point out is I feel public speaking is the equivalent of lifting your own bodyweight. If you can do it, you’re one of the small few of the population that is willing to do so or that can do it. If you start public speaking, that in itself is an accomplishment and an experience that is something that is somewhat enriching. And being bad at it doesn’t take the passion away from you. If you just really want to do it, just keep doing it, even if you’re a bad speaker.

Corey: Yeah. The way to give a great talk because you have a bunch of terrible talks first.

Wesley: Yeah. And it’s okay to do that.

Corey: And it’s not the in entirety of community. It’s not even a requirement to be involved with the community. If you’re one of those people that absolutely dreads the prospect of speaking publicly, fine. I’m not suggesting that, oh, you need to get over that and get on stage. That doesn’t help anyone. Don’t do the things you dread doing because you know that it’s not going to go well for you.

That’s the reason I don’t touch actual databases. I mean, come on, let’s be realistic. I will accidentally the data, and then we won’t have a company anymore. So, I know what things I’m good at and things I’m not. I also don’t do hostage negotiations, for obvious reasons.

Wesley: And also, here’s a little, like, secret tip. If you really want to do public speaking and you start doing public speaking and you’re not so good at it from other peoples’ perspective, but you still love doing it and you think you’re getting better, doing public speaking is one of those things where you can say that you do it and no one will really question how good you are at it. [laugh]. If you’re just in casual conversation, it’s like, “Hey, I wrote a book.” People like, “Oh, wow. This person wrote the book on blah, blah, blah.”

Corey: It’s a self-published book that says the best way to run Kubernetes. It’s a single page; it says, “Don’t.” In 150-point type. “The end.” But I wrote a book.

Wesley: Yeah.

Corey: Yeah.

Wesley: People won’t probe too much and it’ll help you with your development. So, go ahead and get started. Don’t worry about doing that thing where, like, I have to be the best before I can present it. Call yourself a public speaker. Check, done.

Corey: Always. We are the stories we tell, and nowhere is it more true than in the world of public speaking. I really want to thank you for taking the time out of your day to speak with me about this for a second time in a single year. Oh, my goodness. If people want to learn more about what you’re up to, where can they find you?

Wesley: I’m on Twitter, @wesley83 on Twitter. And you can find me also on PolyWork. So, polywork.com/wesley83. Or just go to wesleyfaulkner.com which redirects you there. I list pretty much everything that I am working on and any upcoming speaking opportunities, hopefully when they release that feature, will also be on that Polywork page.

Corey: Excellent. And of course, I started Polywork recently, and I’m at thoughtleader.cloud because of course I am, which is neither here nor there. Thank you so much for taking the time to speak about this side of the industry that we never really get to talk about much, at least not publicly and not very often.

Wesley: Well, thank you for having me on the show. And I wanted to take some time to say thank you for the work that you’re doing. Not just elevating voices like myself, but talking truth to power, like we mentioned before, but being yourself and being a great representation of how people should be treating others: being honest without being mean, being snarky without being rude. And other companies and other people
who’ve given me a chance, and given me a platform, I wanted to say thank you to you too, and I wouldn’t be here unless it was people like you acknowledging the work that I’ve been doing.

Corey: All it takes is just recognizing what you’re doing and acknowledging it. People often want to thank me for this stuff, but it’s just, what, for keeping my eyes open? I don’t know, I feel like it’s just the job; it’s not something that is above and beyond any expected normal behavior. The only challenge is I look around the industry and I realize just how wrong that impression is, apparently. But here we are. It’s about finding people doing interesting work and letting them tell their story. That’s all this podcast has ever tried to be.

Wesley: Yeah. And you do it. And doing the work is part of the reward, and I really appreciate you just going through the effort. Even having your ears open is something that I’m glad that you’re able to at least know who the people are and who are making noises—or making noise to raise their profile up and then in turn, sharing that with the world. And so, that’s a great service that you’re providing, not just for me, but for everyone.

Corey: Well, thank you. And as always, thank you for your time. Wesley Faulkner, Head of Community at SingleStore. I’m Cloud Economist
Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with a rambling comment telling me exactly why DevRel does not need success metrics of any kind.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About James

James has been part of AWS for over 15 years. During that time he's led software engineering for Amazon EC2 and more recently leads the AWS Commerce Platform group that runs some of the largest systems in the world, handling volumes of data and request rates that would make your eyes water. And AWS customers trust us to be right all the time so there's no room for error.

Links Referenced:

  • Email: jamesg@amazon.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Vultr. Optimized cloud compute plans have landed at Vultr to deliver lightning-fast processing power, courtesy of third-gen AMD EPYC processors without the IO or hardware limitations of a traditional multi-tenant cloud server. Starting at just 28 bucks a month, users can deploy general-purpose, CPU, memory, or storage optimized cloud instances in more than 20 locations across five continents. Without looking, I know that once again, Antarctica has gotten the short end of the stick. Launch your Vultr optimized compute instance in 60 seconds or less on your choice of included operating systems, or bring your own. It’s time to ditch convoluted and unpredictable giant tech company billing practices and say goodbye to noisy neighbors and egregious egress forever. Vultr delivers the power of the cloud with none of the bloat. “Screaming in the Cloud” listeners can try Vultr for free today with a $150 in credit when they visit getvultr.com/screaming. That’s G-E-T-V-U-L-T-R dot com slash screaming. My thanks to them for sponsoring this ridiculous podcast.

Corey: Finding skilled DevOps engineers is a pain in the neck! And if you need to deploy a secure and compliant application to AWS, forgettaboutit! But that’s where DuploCloud can help. Their comprehensive no-code/low-code software platform guarantees a secure and compliant infrastructure in as little as two weeks, while automating the full DevSecOps lifestyle. Get started with DevOps-as-a-Service from DuploCloud so that your cloud configurations are done right the first time. Tell them I sent you and your first two months are free. To learn more visit: snark.cloud/duplo. Thats’s snark.cloud/D-U-P-L-O-C-L-O-U-D.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. And I’ve been angling to get someone from a particular department at AWS on this show for nearly its entire run. If you were to find yourself in an Amazon building and wander through the various dungeons and boiler rooms and subterranean basements—I presume; I haven’t seen nearly as many of you inside of those buildings as people might think—you pass interesting departments labeled things like ‘Spline Reticulation,’ or whatnot. And then you come to a very particular group called Commerce Platform.

Now, I’m not generally one to tell other people’s stories for them. My guest today is James Greenfield, the VP of Commerce Platform at AWS. James, thank you for joining me and suffering the slings and arrows I will no doubt be hurling at you.

James: Thanks for having me. I’m looking forward to it.

Corey: So, let’s start at the very beginning—because I guarantee you, you’re going to do a better job of giving the chapter and verse answer than I would from a background mired deeply in snark—what is Commerce Platform? It sounds almost like it’s the retail website that sells socks, books, and underpants.

James: So, Commerce Platform actually spans a bunch of different things. And so, I’m going to try not to bore you with a laundry list of all of the things that we do—it’s a much longer list than most people assume even internal to AWS—at its core, Commerce Platform owns all of the infrastructure and processes and software that takes the fact that you’ve been running an EC2 instance, or you’re storing an object in S3 for some period of time, and turns it into a number at the end of the month. That is what you asked for that service and then proceeds to try to give you as many ways to pay us as easily as possible. There are a few other bits in there that are maybe less obvious. One is we’re also responsible for protecting the platform and our customers from fraudulent activity. And then we’re also responsible for helping collect all of the data that we need for internal reporting to support some of the back-ends services that a business needs to do things like revenue recognition and general financial reporting.

Corey: One of the interesting aspects about the billing system is just how deeply it permeates everything that happens within AWS. I frequently say that when it comes to cloud, cost and architecture are foundationally and fundamentally the same exact thing. If your entire service goes down, a few interesting things happen. One, I don’t believe a single customer is going to complain other than maybe a few accountants here and there because the books aren’t reconciling, but also you’ve removed a whole bunch of constraints around why things are the way that they are. Like, what is the most efficient way to run this workload?

Well, if all the computers suddenly become free, I don’t really care about efficiency, so much is, “Oh, hey. There’s a fly, what do I have as a flyswatter? That’s right, I’m going to drop a building on it.” And those constraints breed almost everything. I’ve said, for example, that S3 has infinite storage because it does.

They can add drives faster than we’re able to fill them—at least historically; they added some more replication services—but they’re going to be able to buy hard drives faster than the rest of us are going to be able to stretch our budgets. If that constraint of the budget falls away, all bets are really off, and more or less, we’re talking about the destruction of the cloud as a viable business entity. No pressure or anything.

James: [laugh].

Corey: You’re also a recent transplant into AWS billing as a whole, Commerce Platform in general. You spent 15 years at the company, the vast majority of that over an EC2. So, either it was you’ve been exiled to a basically digital Siberia or it was one of those, “Okay, keeping all the EC2 servers up, this is easy. I don’t see what people stress about.” And they say, “Oh, ho ho, try this instead.” How did you find yourself migrating over to the Commerce Platform?

James: That’s actually one I’ve had a lot from folks that I’ve worked with. You’re right, I spent the first 15 or so years of my career at AWS in EC2, responsible for various things over there. And when the leadership role in Commerce Platform opened up, the timing was fortuitous, and part of it, I was in the process of relocating my family. We moved to Vancouver in the middle of last year. And we had an opening in the role and started talking about, potentially, me stepping into that role.

The reason that I took it—there’s a few reasons, but the primary reason is that if I look back over my career, I’ve kind of naturally gravitated towards owning things where people only really remember that they exist when they’re not working. And for some reason, you know, I enjoy the opportunity to try to keep those kinds of services ticking over to the point where people don’t notice them. And so, Commerce Platform lands squarely in that space. I’ve always been attracted to opportunities to have an impact, and it’s hard to imagine having much more of an impact than in the Commerce Platform space. It underpins everything, as you said earlier.

Every single one of our customers depends on the service, whether they think about it or realize it. Every single service that we offer to customers depends on us. And so, that really is the sort of nexus within AWS. And I’m a platform guy, I’ve always been a platform guy. I like the force multiplier nature of platforms, and so Commerce Platform, you know, as I kind of thought through all of those elements, really was a great opportunity to step in.

And I think there’s something to be said for, I’ve been a customer of Commerce Platform internally for a long time. And so, a chance to cross over and be on the other side of that was something that I didn’t want to pass up. And so, you know, I’m digging in, and learning quickly, ramping up. By no means an expert, very dependent on a very smart, talented, committed group of people within the team. That’s kind of the long and short of how and why.

Corey: Let’s say that I am taking on the role of an AWS product team, for the sake of argument. I know, keep the cringe down for a second, as far as oh, God, the wince is just inevitable when the idea of me working there ever comes up to anyone. But I have an idea for a service—obviously, it runs containers, and maybe it does some other things as well—going from idea to six-pager to MVP to barely better than MVP day-one launch, and at some point, various things happen to that service. It gets staff with a team, objectives and a roadmap get built, a P&L and budget, and a pricing model and the rest. One the last thing that happens, apparently, is someone picks the worst name off of a list of candidates, slaps it on the product, and ships it off there.

At what point does the billing system and figuring out the pricing dimensions for a given service tend to factor in? Is that a last-minute story? Is that almost from the beginning? Where along that journey does, “Oh, by the way, we’re building this thing. Maybe we should figure out, I don’t know, how to make money from it.” Factor into the conversation?

James: There are two parts to that answer. Pretty early on as we’re trying to define what that service is going to look like, we’re already typically thinking about what are the dimensions that we might charge along. The actual pricing discussions typically happen fairly late, but identifying those dimensions and, sort of, the right way to present it to customers happens pretty early on. The thing that doesn’t happen early enough is actually pulling the Commerce Platform team in. but it is something that we’re going to work this year to try to get a little bit more in front of.

Corey: Have you found historically that you have a pretty good idea of how a service is going to be priced, everything is mostly thought through, a service goes to either private preview or you’re discussing about a launch, and then more or less, I don’t know, someone like me crops up with a, “Hey, yeah, let’s disregard 90% of what the service does because I see a way to misuse the remaining 10% of it as a database.” And you run some mental math and realize, “Huh. We’re suddenly giving, like, eight petabytes of storage per customer away for free. Maybe we should guard against that because otherwise, it’s rife with misuse.” It used to be that I could find interesting ways to sneak through the cracks of various services—usually in pursuit of a laugh—those are getting relatively hard to come by and invariably a lot more trouble than they’re worth. Is that just better comprehensive diligence internally, is that learning from customers, or am I just bad at this?

James: No, I mean, what you’re describing is almost a variant of the Defender’s Dilemma. They are way more ways to abuse something than you can imagine, and so defending against that is pretty challenging. And it’s important because, you know, if you turn the economics of something upside down, then it just becomes harder for us to offer it to customers who want to use it legitimately. I would say 90% of that improvement is us learning. We make plenty of mistakes, but I think, you know, one of the things that I’ve always been impressed by over my time here is how intentional we are trying to learn from those mistakes.

And so, I think that’s what you’re seeing there. And then we try very hard to listen to customers, talk to folks like you, because one of the best ways to tackle anything it smells of the Defender’s Dilemma is to harness that collective creativity of a large number of smart people because you really are trying to cover as much ground as possible.

Corey: There was a fun joke going around a while back of what is the most expensive environment you can get running on a free tier account before someone from AWS steps in, and I think I got it to something like half a billion dollars in the first month. Now, I haven’t actually tested this for reasons that mostly have to do with being relatively poor compared to, you know, being able to buy Guam. And understanding as well the fraud protections built into something like AWS are largely built around defending against getting service usage for free that in some way, shape or form, benefits the attacker. The easy example of that would be mining cryptocurrency, which is just super-economic as long as you use someone else’s AWS account to do it. Whereas a lot of my vectors are, “Yeah, ignore all of that. How do I just make the bill artificially high? What can I do to misuse data transfer? And passing a single gigabyte through, how much can I make that per gigabyte cost be?” And, “Oh, circular replication and the Lambda invokes itself pattern,” and basically every bad architectural decision you can possibly make only this time, it’s intentional.

And that shines some really interesting light on it. And I have to give credit where due, a lot of that didn’t come from just me sitting here being sick and twisted nearly so much as it did having seen examples of that type of misconfiguration—by mistake—in a variety of customer accounts, most confidently my own because it turns out that the way I learn things is by screwing them up first.

James: Yeah, you’ve touched on a couple of different things in there. So, you know, maybe the first one is, I typically try to draw a line between fraud and abuse. And fraud is essentially trying to spend somebody else’s money to get something for free. And we spent a lot of time trying to shut that down, and we’re getting really good at catching it. And then abuse is either intentional or unintentional. There’s intentional abuse: You find a chink in our armor and you try to take advantage of it.

But much more commonly is unintentional abuse. It’s not really abuse, you know. Abuse has very negative connotations, but it’s unintentionally setting something up so that you run up a much larger bill than you intended. And we have a number of different internal efforts, and we’re working on a bunch more this year, to try to catch those early on because one of my personal goals is to minimize the frequency with which we surprise customers. And the least favorite kind of surprise for customers is a [laugh] large bill. And so, what you’re talking about there is, in a sufficiently complex system, there’s always going to be weaknesses and ways to get yourself tied up in knots.

We’re trying both at the service team level, but also within my teams to try to find ways to make it as hard as possible to accidentally do that to yourself and then catch when you do so that we can stop it. And even more on the intentional abuse side of things, if somebody’s found a way to do something that’s problematic for our services, then you know, that’s pretty much on us. But we will often reach out and engage with whoever’s doing and try to understand what they’re trying to do and why. Because often, somebody’s trying to do something legitimate, they’ve got a problem to solve, they found a creative way to solve it, and it may put strain on the service because it’s just not something we designed for, and so we’ll try to work with them to use that to feed into either new services, or find a better place for that workload, or just bolster what they’re using. And maybe that’s something that eventually becomes a fully-fledged feature that we offer the customers. We’re always open to learning from our customers. They have found far more creative ways to get really cool things done with our services than we’ve ever imagined. And that’s true today.

Corey: I mean, most of my service criticisms come down to the fact that you have more-or-less built a very late model, high performing iPad, and I’m out there complaining about, “What a shitty hammer this thing is, it barely works at all, and then it breaks in my hand. What gives?” I would also challenge something you said a minute ago that the worst day for some customers is to get a giant surprise bill, but [unintelligible 00:13:53] to that is, yeah, but, on some level, that kind of only money; you do have levers on your side to fix those issues. A worse scenario is you have a customer that exhibits fraud-like behavior, they’re suddenly using far more resources than they ever did before, so let’s go ahead and turn them off or throttle them significantly, and you call them up to tell them you saved them some money, and, “Our Superbowl ad ran. What exactly do you think you’re doing?” Because they don’t get a second bite at that kind of Apple.

So, there’s a parallel on both sides of this. And those are just two examples. The world is full of nuances, and at the scale that you folks operate at. The one-in-a-million events happen multiple times a second, the corner cases become common cases, and I’m surprised—to be direct—how little I see you folks dropping the ball.

James: Credit to all of the teams. I think our secret sauce, if anything, really does come down to our people. Like, a huge amount of what you see as hopefully relatively consistent, good execution comes down to people behind the scenes making sure. You know, like, some of it is software that we built and made sure it’s robust and tested to scale, but there’s always an element of people behind the scenes, when you hit those edge cases or something doesn’t quite go the way that you planned, making sure that things run smoothly. And that, if anything, is something that I’m immensely proud of and is kind of amazing to watch from the inside.

Corey: And, on some level, it’s the small errors that are the bigger concern than the big ones. Back a couple years ago, when they announced GP3 volumes at re:Invent, well, great, well spin up a test volume and kick the tires on it for an hour. And I think it was 80 or 100 gigs or whatnot, and the next day in the bill, it showed up as about $5,000. And it was, “Okay, that’s not great. Not great at all.” And it turned out that it was a mispricing error by I think a factor of a million.

And okay, at least it stood out. But there are scenarios where we were prepared to pay it because, oops, you got one over on us. Good job. That’s never been the mindset I’ve gotten about AWS’s philosophy for pricing. The better example that I love because no one took it seriously, was a few years before that when there was a LightSail bug in the billing system, and it made the papers because people suddenly found that for their LightSail instance, they were getting predicted bills of $4 billion.

And the way I see it, you really only had to make that work once and then you’ve made your numbers for the year, so why not? Someone’s going to pay for it, probably. But that was such out-of-the-world numbers that no one saw that and ever thought it was anything other than a bug. It’s the small pernicious things that creep in. Because the billing system is vast; I had no idea when I started working with AWS bills just how complicated it really was.

James: Yeah, I remember both of those, and there’s something in there that you touched on that I think is really important. That’s something that I realized pretty early on at Amazon, and it’s why customer obsession is our flagship leadership principle. It’s not because it’s love and butterflies and unicorns; customer obsession is key to us because that’s how you build a long-term sustainable business is your customers depend on you. And it drives how we think about everything that we do. And in the billing space, small errors, even if there are small errors in the customer’s favor, slowly erode that trust.

So, we take any kind of error really seriously and we try to figure out how we can make sure that it doesn’t happen again. We don’t always get that right. As you said, we’ve built an enormous, super-complex business to growing really quickly, and really quick growth like that always acts as kind of a multiplier on top of complexity. And on the pricing points, we’re managing millions of pricing points at the moment.

And our tools that we use internally, there’s always room for improvement. It’s a huge area of focus for us. We’re in the beginning of looking at applying things like formal methods to make sure that we can make very hard guarantees about the correctness of some of those. But at the end of the day, people are plugging numbers in and you need as many belts and braces as possible to make sure that you don’t make mistakes there.

Corey: One of the things that struck me by surprise when I first started getting deep into this space was the fact that the finalized bill was—what does it mean to have this be ‘finalized?’ It can hit the Cost and Usage Report in an S3 bucket and it can change retroactively after the month closed periodically. And that’s when I started to have an inkling of a few things: Not just the sheer scale and complexity inherent to something like the billing system that touches everything, but the sheer data retention stories where you clearly have to be able to go back and reconstruct a bill from the raw data years ago. And I know what the output of all of those things are in the form of Cost and Usage Reports and the billing data from our client accounts—which is the single largest expense in all of our AWS accounts; we spent thousands and thousands and thousands of dollars a year just on storing all of that data, let alone the processing piece of it—the sheer scale is staggering. I used to wonder why does it take you a day to record me using something to it’s showing up in the bill? And the more I learned the more it became a how can you do that in only a day?

James: Yes, the scale is actually mind-boggling. I’m pretty sure that the core of our billing system is—I’m reasonably confident it’s the largest or one of the largest data processing systems on the planet. I remember pretty early on when I joined Commerce Platform and was still starting to wrap my head around some of these things, Googling the definition of quadrillion because we measured the number of metering events, which is how we record usage in services, on a daily basis in the quadrillions, which is a billion billions. So, it’s just an absolutely staggering number. And so, the scale here is just out of this world.

That’s saying something because it’s not like other services across AWS are small in their own right. But I’m still reasonably sure that being one of a handful of services that is kind of at the nexus of AWS and kind of deals with the aggregate of AWS’s scale, this is probably one of the biggest systems on the planet. And that shows up in all sorts of places. You start with that input, just the sheer volume of metering events, but that has to produce as an output pretty fine-grained line item detailed information, which ultimately rolls up into the total that a customer will see in their bill. But we have a number of different systems further down the pipeline that try to do things like analyze your usage, make sensible recommendations, look for opportunities to improve your efficiency, give you the ability to slice and dice your data and allocate it out to different parts of your business in whatever way it makes sense for your business. And so, those systems have to deal with anywhere from millions to billions to recently, we were talking about trillions of data points themselves. And so, I was tangentially aware of some of the scale of this, but being in the thick of it having joined the team really just does underscore just how vast the systems are.

Corey: I think it’s, on some level, more than a little unfortunate that that story isn’t being more widely told, more frequently. Because when Commerce Platform has job postings that are available on the website, you read it and it’s very vague. It doesn’t tend to give hard numbers about a lot of these things, and people who don’t play in these waters can easily be forgiven for thinking the way that you folks do your job is you fire up one of those 24 terabyte of RAM instances that—you know, those monstrous things that you folks offer—and what do you do next? Well, Microsoft Excel. We have a special high memory version that we’ve done some horse-trading with our friends over at Microsoft for.

It’s, yeah, you’re several steps beyond that, at this point. It’s a challenging problem that every one of your customers has to deal with, on some
level, as well. But we’re only dealing with the output of a lot of the processing that you folks are doing first.

James: You’re exactly right. And a big focus for some of my teams is figuring out how to help customers deal with that output. Because even if you’re talking about couple of orders of magnitude reduction, you’re still talking about very large numbers there. So, to help customers make sense of that, we have a range of tools that exist, we’re investing in.

There’s another dimension of complexity in the space that I think is one that’s also very easy to miss. And I think of it as arbitrary complexity. And it’s arbitrary because some of the rules that we have to box within here are driven by legislative changes. As you operate more and more countries around the world, you want to make sure that we’re tax compliant, that we help our customers be tax compliant. Those rules evolve pretty rapidly, and Country A may sit next to Country B, but that doesn’t mean that they’re talking to one another. They’ve all got their own ideas. They’re trying to accomplish r—00:22:47

Corey: A company is picking up and relocating from India to Germany. How do we—

James: Exactly.

Corey: —change that on the AWS side and the rest? And it’s, “Hoo boy, have you considered burning it all down and filing an insurance claim
to start over?” And, like, there’s a lot of complexity buried underneath that that just doesn’t rise to the notice of 99% of your customers.

James: And the fact that it doesn’t rise to the notice is something that we strive for. Like, these shouldn’t be things that customers have to worry about. Because it really is about clearing away the things that, as far as possible, you don’t want to have to spend time thinking about so that you can focus on the thing that your business does that differentiates you. It’s getting rid of that undifferentiated heavy lifting. And there’s a ton of that in this space, and if you’re blissfully unaware of it, then hopefully that means that we’re doing our job.

Corey: What I’m, I think, the most surprised about, and I have been for a long time. And please don’t take this as an insult to various other folks—engineers, the rest, not just in other parts of AWS but throughout the other industry—but talking to the people who work within Commerce Platform has always been just a fantastic experience. The caliber of people that you have managed to attract and largely retain—we don’t own people, they do matriculate out eventually—but the caliber of people that you’ve retained on your teams has just been out of this world. And at first, I wondered, why are these awesome people working on something as boring and prosaic as billing? And then I started learning a little bit more as I went, and, “Oh, wow. How did they learn all the stuff that they have to hold in their head in tension at once to be able to build things like this?” It’s incredibly inspiring just watching the caliber of the people that you’ve been able to bring in.

James: I’ve been really, really excited joining this team, as I’ve gotten other folks on the team because there’s some super-smart people here. But what’s really jumped out to me is how committed the team is. This is, for the most part, a team that has been in the space for many years. Many of them have—we talk about boomerangs, folks who live AWS, go spend some time somewhere else and come back and there’s a surprisingly high proportion of folks in Commerce Platform who have spent time somewhere else and then come back because they enjoy the space, they find that challenging, folks are attracted to the ability to have an impact because it is so foundational. But yeah, there’s a super-committed core to this team. And I really enjoy working with teams where you’ve got that because then you really can take the long view and build something great. And I think we have tons of opportunities to do that here.

Corey: It sounds ridiculous, but I’ve reached out to team members before to explain two-cent variances in my bill, and never once have I been confronted with a, “It’s two cents. What do you care?” They understand the requirement that these things be accurate, not just, “Eh, take our word for it.” And also, frankly, they understand that two cents on a $20 bill looks a little different on a $20 million bill. So yeah, let us figure out if this is systemic or something I have managed to break.

It turns out the Cost and Usage Report processing systems don’t love it when there’s a cost allocation tag whose name contains an emoji. Who knew? It’s the little things in life that just have this fun way of breaking when you least expect it.

James: They’re also a surprisingly interesting problem. So like, it turns out something as simple as rounding numbers consistently across a distributed system at this scale, is a non-trivial problem. And if you don’t, then you do get small seventh or eighth decimal place differences that add up to something that then shows up as a two-cent difference somewhere. And so, there’s some really, really interesting problems in the space. And I think the team often takes these kinds of things as a personal challenge. It should be correct, and it’s not, so we should go make sure it is correct. The interesting problems abound here, but at the end of the day, it’s the kind of thing that any engineering team wants to go and make sure it’s correct because they know that it can be.

Corey: This episode is sponsored in parts by our friend EnterpriseDB. EnterpriseDB has been powering enterprise applications with PostgreSQL for 15 years. And now EnterpriseDB has you covered wherever you deploy PostgreSQL on premises, private cloud, and they just announced a fully managed service on AWS and Azure called BigAnimal, all one word.

Don't leave managing your database to your cloud vendor because they're too busy launching another half dozen manage databases to focus on any one of them that they didn't build themselves. Instead, work with the experts over at EnterpriseDB. They can save you time and money, they can even help you migrate legacy applications, including Oracle, to the cloud.

To learn more, try BigAnimal for free. Go to biganimal.com/snark, and tell them Corey sent you.

Corey: On the one hand, I love people who just round and estimate—we all do that, let’s be clear; I sit there and I back-of-the-envelope everything first. But then I look at some of your pricing pages and I count the digits after the zeros. Like, you’re talking about trillionths of a dollar on some of your pricing points. And you add it up in the course of a given hour and it’s like, oh, it’s $250 a month, most months. And it’s you work backwards to way more decimal places of precision than is required, sometimes.

I’m also a personal fan of the bill that counts, for example, number of Route 53 zones. Great. And it counts them to four decimal places of precision. Like, I don’t even know what half of it Route 53 zone is at this point, let alone something to, like, ah the 1,000th of the zone is going to cause this. It’s all an artifact of what the underlying systems are.

Can you by any chance shed a little light on what the evolution of those systems has been over a period of time? I have to imagine that anything you built in the early days, 16 years ago or so from the time of this recording when S3 launched to general availability, you probably didn’t have to worry about this scope and scale of what you do, now. In fact, I suspect if you tried to funnel this volume through S3 back then, the whole thing would have collapsed under its own weight. What’s evolved over the time that you had the billing system there? Because changes come slowly to your environment. And frankly, I appreciate that as a customer. I don’t like surprising people in finance.

James: Yeah, you’re totally right. So, I joined the EC2 team as an engineer myself, some 16 years ago, and the very first thing that I did was our billing integration. And so, my relationship with the Commerce Platform organization—what was the billing team way back when—it goes back over my entire career at AWS. And at the time, the billing team was similar, you know, [unintelligible 00:28:34] eight people. And that was everything. There was none of the scale and complexity; it was all one system.

And much like many of our biggest, oldest services—EC2 is very similar, S3 is as well—there’s been significant growth over the last decade-and-a-half. A lot of that growth has been rapid, and rapid growth presents its own challenges. And you live with decisions that you make early on that you didn’t realize were significant decisions that have pretty deep implications 15 years later. We’re still working through some of those; they present their own challenges. Evolving an existing system to keep up with the growth of business and a customer base that’s as varied and complex as ours is always challenging.

And also harder but I also think more fun than a clean sheet redo at this point. Like, that’s a great thought exercise for, well, if we got to do this again today, what would we do now that we’ve learned so much over the last 15 years? But there’s this—I find it personally fascinating challenge with evolving a live system where it’s like, “No, no, like, things exist, so how do we go from there to where we want to be next?”

Corey: Turn the billing system off for 18 months, rebuild—

James: Yeah. [laugh].

Corey: The whole thing from first principles. Light it up. I’m sure you’d have a much better billing system, and also not a company left anymore.

James: [laugh]. Exactly, exactly. I’ve always enjoyed that challenge. You know, even prior to AWS, my previous careers have involved similar kinds of constraints where you’ve got a live system, or you’ve got an existing—in the one case, it was an existing SDK that was deployed to tens of thousands of customers around the world, and so backwards compatibility was something that I spent the first five years of my career thinking about it way more detail than I think most people do. And it’s a very similar mindset. And I enjoy that challenge. I enjoy that: How do I evolve from here to there without breaking customers along the way?

And that’s something that we take pretty seriously across AWS. I think SimpleDB is the poster child for we never turn things off. But that applies equally to the services that are maybe less visible to customers, and billing is definitely one of them. Like, we don’t get to switch stuff off. We don’t get to throw things away and start again. It’s this constant state of evolution.

Corey: So, let’s say that I were to find a way to route data through a series of two Managed NAT Gateways and then egress to internet, and the sheer density of the expense of that traffic tears a hole in the fabric of space-time, it goes back 15 years ago, and you can make a single change to how the billing system was built. What would it be? What pisses you off the most about the current constraints that you have to work within or around?

James: I think one of the biggest challenges we’ve got, actually, is the concept of an account. Because an account means half-a-dozen different things. And way back, when it seemed like a great idea, you just needed an account; an account was your customer, and it was the same thing as the boundary that you put all your resources inside. And of course, it’s the same thing that you’re going to roll all of your usage up and issue a bill against. And that has been one of the areas that’s seen the most evolution and probably still has a pretty long way to go.

And what’s interesting about that is, that’s probably something we could have seen coming because we watched the retail business go through, kind of, the same evolution because they started with, well, a customer is a customer is a customer and had to evolve to support the concept of sellers and partners. And then users are different than customers, and you want to log in and that’s a different thing. So, we saw that kind of bifurcation of a single entity into a wide range of different related but separate entities, and I think if we’d looked at that, you know, thought out 15 years, then yeah, we could probably have learned something from that. But at the same time, when AWS first kicked off, we had wild ambitions for it, but there was no guarantee that it was going to be the monster that it is today. So, I’m always a little bit reluctant to—like, it’s a great thought exercise, but it’s easy to end up second-guessing a pretty successful 15 years, so I’m always a little bit careful to walk that line. But I think account is one of the things that we would probably go back and think about a little bit more.

Corey: I want to be very clear with this next question that it is intentionally setting up a question I suspect you get a lot. It does not mirror my own thinking on the matter even slightly, but I get a version of it myself all the time. “AWS bills, that sounds boring as hell. Why would you choose to work on such a thing?” Now, I have a laundry list of answers to that aren’t nearly as interesting as I suspect yours are going to be. What makes working on this problem space interesting to you?

James: There’s a bunch of different things. So, first and foremost, the scale that we’re talking about here is absolutely mind-blowing. And for any engineer who wants to get stuck into problems that deal with mind-blowingly large volumes of data, incredibly rich dimensions, problems where, honestly, applying techniques like statistical reasoning or machine learning is really the only way to chip away at it, that exists in spades in the space. It’s not always immediately obvious, and I think from the outside, it’s easy to assume this is actually pretty simple. So, the scale is a huge part of that.

Corey: “Oh, petabytes. How quaint.”

James: [laugh]. Exactly. Exactly I mean, it’s mind-blowing every time I see some of the numbers in various parts of the Commerce Platform space. I talked about quadrillions earlier. Trillions is a pretty common unit of measure.

The complexity that I talked about earlier, that’s a result of external environments is another one. So, imposed by external entities, whether it’s a government or a tax authority somewhere, or a business requirement from customers, or ourselves. I enjoy those as well. Those are different kinds of challenge. They really keep you on your toes.

I enjoy thinking of them as an engineering problem, like, how do I get in front of them? And that’s something we spend a lot of time doing in Commerce Platform. And when we get it right, customers are just unaware of it. And then the third one is, I personally am always attracted to the opportunity to have an impact. And this is a space where we get to hopefully positively impact every single customer every day. And that, to me is pretty fulfilling.

Those are kind of the three standout reasons why I think this is actually a super-exciting space. And I think it’s often an underestimated space. I think once folks join the team and sort of start to dig in, I’ve never heard anybody after they’ve joined, telling me that what they’re doing is boring. Challenging, yes. Is frustrating, sometimes. Hard, absolutely, but boring never comes up.

Corey: There’s almost no service, other than IAM, that I can think of that impacts every customer simultaneously. And it’s easy for me to sit in the cheap seats and say, “Oh, you should change this,” or, “You should change that.” But every change you have is so massive in scale that it’s going to break a whole bunch of companies’ automations around the bill processing in different ways. You have an entire category of user persona who is used to clicking a certain button in this certain place in the console to generate the report every month, and if that button moves or changes color, or has a different font, suddenly that renders their documentation invalid, and they’re scrambling because it’s not their core competency—nor should it be—and every change you make is so constricted, just based upon all the different concerns that you’ve got to be juggling with. How do you get anything done at all? I find that to be one of the most impressive aspects about your organization, bar none.

James: Yeah, I’m not going to lie and say that it isn’t a challenge, but a lot of it comes down to the talent that we have on the team. We have a super-motivated, super-smart, super-engaged team, and we spend a lot of time figuring out how to make sure that we can keep moving, keep up with the business, keep up with a world that’s getting more complicated [laugh] with every passing day. So, you’ve kind of hit on one of the core challenges there, which is, how do we keep up with all of those different dimensions that are demanding an increasing amount of engineering and new support and new investment from us, while we keep those customers happy?

And I think you touched on something else a little bit indirectly there, which is, a lot of our customers are actually pretty technical across AWS. The customers that Commerce Platform supports, are often the least technical of our customers, and so often need the most help understanding why things are the way they are, where the constraints are.

Corey: “A big bill from Amazon. How many books did you people buy last month?”—

James: [laugh]. Exactly.

Corey: —is still very much level of understanding in some cases. And it’s not because they’re dumb; far from it. It’s just, imagine that some
people view there as being more to life than understanding the nuances and intricacies of cloud computing. How dare they?

James: Exactly. Who would have thought?

Corey: So, as you look now over all of your domain, such as it is, what sucks the most? What are you looking to fix as far as impactful changes that the rest of the world might experience? Because I’m not going to accept one of those questions like, “Oh, yeah, on the back-end, we have
this storage subsystem for a tertiary thing that just annoys me because it wakes us up once in a whi”—no, no, I want something customer-facing. What’s the painful thing you’re looking at fixing next?

James: I don’t like surprising customers. And free tier is, sort of, one of those buckets of surprises, but there are others. Another one that’s pretty squarely in my sights is, whether we like it or not, customer accounts get compromised. Usually, it’s a password got reused somewhere or was accidentally committed into a GitHub repository somewhere.

And we have pretty established, pretty effective mechanisms for finding all of those, we’ll scan for passwords and credentials, and alert customers to those, and help them correct that pretty quickly. We’re also actually pretty good at detecting when an account does start to do something that suggests that it’s been compromised. Usually, the first thing that a compromised account starts to do is cryptocurrency mining. We’re pretty quick to catch those; we catch those within a matter of hours, much faster most days.

What we haven’t really cracked and where I’m focused at the moment is getting back to the customer in a way that’s effective. And by that I mean specifically, we detect an account compromised super-quickly, we reach out automatically. And so, you know, a customer has got some kind of contact from us usually within a couple of hours. It’s not having the effect that we need it to. Customers are still being surprised a month later by a large bill. And so, we’re digging into how much of that is because they never saw the contact, they didn’t know what to do with the contact.

Corey: It got buried with all the other, “Hey, we saw you spun up an S3 bucket. Have you heard of what S3 is?” Again, that’s all valuable, but you have 300-some-odd services. If you start doing that for every service, you’re going to hit mail sending limits for Gmail.

James: Exactly. It’s not just enough that we detect those and notify customers; we have to reduce the size of the surprise. It’s one thing to spend 100 bucks a month on average, and then suddenly find that your spend has jumped $250 because you reused the password somewhere and somebody got ahold of it and it’s cryptocurrency-mining your account. It’s a whole different ballgame to spend 100 bucks a month and then at the end of the month discover that your bill is suddenly $2,000 or $20,000. And so, that’s something that I really wanted to make some progress on this year.

Corey: I’ve really enjoyed our conversation. If people want to learn more about how you view these things, how you’re approaching some of these problems, or potentially are just the right kind of warped to consider joining up, where’s the best place for them to go?

James: They should drop me an email at jamesg@amazon.com. That is the most direct way to get hold of me, and I promise I will get back to you. I try to stay on top of my email as much as possible. But that will come straight to me, and I’m always happy to talk to folks about the space, talk to folks about opportunities in this team, opportunities across AWS, or just hear what’s not working, make sure that it’s something that we’re aware of and looking at.

Corey: Throughout Amazon, but particularly within Commerce Platform, I’ve always appreciated the response of, whenever I report something, no matter how ridiculous it is—and I assure you there’s an awful lot of ridiculousness in my bug reports—the response has always been the same: “Tell me more. Help me understand what it is you’re trying to achieve—even if it is ridiculous—so we can look at this and see what is actually going on.” Every Amazonian team has been great about that or you’re not at Amazon very long, but you folks have taken that to an otherworldly level. I just want to thank you for doing that.

James: I appreciate you for calling that out. We try, you know, we really do. We take listening to our customers very seriously because, at the end of the day, that’s what makes us better, and that’s how we make sure we’re in it for the long haul.

Corey: Thanks once again for being so generous with your time. I really appreciate it.

James: Yeah, thanks for having me on. I’ve enjoyed it.

Corey: James Greenfield, VP of Commerce Platform at AWS. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry comment—possibly on YouTube as well—about how you aren’t actually giving this five-stars at all; you have taken three trillions of a star off of the rating.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Priyanka

Priyanka Vergadia is currently a Staff Developer Advocate at Google Cloud where she works with enterprises to build and architect their cloud platforms. She enjoys building engaging technical content and continuously experiments with new ways to tell stories and solve business problems using Google Cloud tools. You can check out some of the stories that she has created for the developer community on the Google Cloud Platform Youtube channel. These include "Deconstructing Chatbots", "Get Cooking in Cloud", "Pub/Sub Made Easy" and more. ..

Links Referenced:

  • LinkedIn: https://www.linkedin.com/in/pvergadia/
  • Twitter: https://twitter.com/pvergadia
  • Priyanka's book: https://www.amazon.com/Visualizing-Google-Cloud-Illustrated-References/dp/1119816327

TranscriptAnnouncer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Finding skilled DevOps engineers is a pain in the neck! And if you need to deploy a secure and compliant application to AWS, forgettaboutit! But that’s where DuploCloud can help. Their comprehensive no-code/low-code software platform guarantees a secure and compliant infrastructure in as little as two weeks, while automating the full DevSecOps lifestyle. Get started with DevOps-as-a-Service from DuploCloud so that your cloud configurations are done right the first time. Tell them I sent you and your first two months are free. To learn more visit: snark.cloud/duplo. Thats’s snark.cloud/D-U-P-L-O-C-L-O-U-D.

Corey: What if there were a single place to get an inventory of what you're running in the cloud that wasn't "the monthly bill?" Further, what if there were a way to compare that inventory to what you were already managing via Terraform, Pulumi, or CloudFormation, but then automatically add the missing unmanaged or drifted parts to it? And what if there were a policy engine to immediately flag and remediate a wide variety of misconfigurations? Well, stop dreaming and start doing; visit snark.cloud/firefly to learn more.

Corey: Welcome to Screaming in the Cloud, I’m Corey Quinn. Periodically, I get the privilege of speaking to people who work in varying aspects of some would call it developer evangelism, some would call it developer advocacy, developer relations is a commonly accepted term, and I of course call it devrelopers because I enjoy annoying absolutely everyone by giving things terrible names. My guest today is Priyanka Vergadia, who is a staff developer advocate at Google Cloud. Priyanka, thank you for joining me.

Priyanka: Thank you so much for having me. Corey. I’m so excited to be your developer—what did you call it again?

Corey: Devreloper. Yes indeed.

Priyanka: Devreloper. That is the term I’m going to be using from now on. I am a devreloper. Anyway.

Corey: Excellent.

Priyanka: Yeah.

Corey: I’m starting to spread this out so that eventually we’re going to form a giant, insufferable army of people who pronounce it that way, and it’s going to be great.

Priyanka: It’s going to be awesome. [laugh].

Corey: One of the challenges, even as I alluded to different titles within this space, everyone has a slightly different definition of where the role starts and stops, just in terms of its function, let alone the myriad ways that can be expressed. In the before times, I knew a number of folks in the developer advocacy space who were more or less worldwide experts in accumulating airline miles and racking up status and going from conference to conference to conference to more or less talk about things that had a tenuous at best connection to where they worked. Great. Other folks have done things in very different ways. Some people write extensively, blog posts and the rest, others build things a sample code, et cetera, et cetera.

It seems like every time I talk to someone in the space, they have found some new and exciting way of carrying the message of what their company does to arguably a very cynical customer group. Where do you start and stop with your devrelopment?

Priyanka: Yeah. So, that is such—like, all the devrelopers have their own style that they have either adopted or learned over time that works for them. When I started, I think about three years ago, I did go to conferences, did those events, give talks, all of that, but I was also—my actual introduction to DevRel [laugh] was with videos. I started creating my first series was deconstructing chatbots, and I was very interested in learning more about chatbots. So, I was like, you know what, I’m just going to teach everybody, and learn.

So like, learn and teach at the same time was my motto, and that’s kind of how I got started into, like, okay, I’m going to create a few videos to learn this and teach it. And during the process I was like, “I want to do this more.” And that’s kind of transitioned, my move from being in front of customers, which I still end up doing, but I was doing more of just, you know, working with customers extensively to get their deployments done. This was a segue for me to, you know, think back, sit back and think about what’s working and what I personally enjoy doing more, and that’s what got me into creating videos. And it’s like, okay, I’m going to become a devreloper now.

And that’s kind of how the whole, like, journey started. And for me, like you were pointing out earlier—should I just stop because I’ve been talking too long? [laugh].

Corey: No, keep going. Please, [unintelligible 00:04:10] it’s fine.

Priyanka: [laugh]. For me, I started—I found my, I would say, in the last two years—it was all before the pandemic, we were all either writing blogs or doing videos or going to conferences, so it was, you know, the pandemic kind of brought us to a point where it’s like, “Okay, let’s think about—we can’t meet each other; let’s think about other ways to communicate and how can we make it creative and exciting?”

Corey: And the old way started breaking down, too, where it’s, “Yay, I’m going to watch an online conference.” “What is it?” “Oh, it’s like a crappy Zoom only you don’t have to pretend to pay attention in the same way.” And as a presenter, then you’ve got to modify what you’re
doing to understand that people’s attention spans are shorter, distraction is always a browser tab away, and unlike a physical event, people don’t feel the same sense of shame of getting up from the front row and weaving in front of 300 people, and not watching the rest of your talk. I mean, don’t get me wrong, I’ll still do it, but I’ll feel bad about it.

Now it’s, “Oh, nope, I’m sitting here in my own little… hovel, I’m just going to do and watch whatever I want to do.” So, you’ve got to—it forces you to up your game, and it—

Priyanka: Yep.

Corey: Still doesn’t quite have the same impact.

Priyanka: Yeah. Or just switch off the camera, if you’re like me, and just—uh, shut off the camera, go away or do something else. And, yeah, it’s very easy to do that. So, it’s not the same, which is why it prompted, I think all of us DevRel people to think about new ways to connect, which is for me that way to connect is art and visual aspects, to kind of bring that—because that—we are all whether we accept it or not or like it or not, we’re all visual learners, so that’s kind of how I think when it comes to creating content is visually appealing, and that’s when people can dive in. [laugh].

Corey: I am in the, I guess opposite side of the universe from you, where I acknowledge and agree with everything you’re saying that people are visual creatures inherently, but I have effectively zero ability in that direction. My medium has always been playing games with words and language. And over time, I had the effectively significantly belated realization that wait a minute, just because I’m not good at a thing doesn’t mean that other people might not be good at that thing, and I don’t have to do every last part of it myself. Suddenly, I didn’t have to do my own crappy graphic design because you can pay people who are worlds better than I’ll ever be, and so on and so forth. I don’t edit my own podcast audio because I’m bad at that, too.

But talking about things is a different story, writing about things, building things is where I tend to see a lot of what I do tend to resonate. But I admit I bias for the things that I enjoy doing and the way that I enjoy consuming things. You do as well because relatively recently, as of time of this recording, you have done what I don’t believe anyone actually wants to do. You wrote a book. Now, everyone wants to have written a book, but no one actually wants to write a book.

Priyanka: So, true. [laugh].

Corey: But it’s not like most technical books. Tell me about it.

Priyanka: Yeah, I actually never thought I would write a book. If you asked me two years ago—three years ago, I would say, I would have never thought that I would write a book because I am not a text person. So, I don’t like to read a lot of texts because it zones out. So, for me, when I started creating some of these sketches, and sharing it on social media and in blogs and things like that, and gotten the attention that it has gotten from people, that’s when I was like, okay, ding, ding, ding. I think I can do a visual book with these images.

And this was like, halfway through, I’d already created, like, 30 sketches at this point. And I was like, “Okay, maybe I can turn this into a book,” which would be interesting for me because I like doing art-type things along with teaching, and it’s not text because I wanted to do this in a very unique way. So yeah, that’s kind of how it ended up happening.

Corey: I have a keen appreciation for people who approach things with a different point of view. One of your colleagues, Forrest Brazeal, took a somewhat similar approach in the in his book, The Read Aloud Cloud, where it was illustrated, and everything he did was in rhyme, which is a constant source of envy for me, where it’s, “Mmm, I’ve got to find a way to one-up him again.” And it’s… he is inexorable, as far as just continuing to self-improve. So, all right, we’re going to find a way to wind up defeating that. With you, it’s way easier.

I read a book, like, wow, this is gorgeous and well-written that it’s attractive to look at, and I will never be able to do any of those things. That’s all you. It doesn’t feel like we’re trying to stand at the same spot in the universe in quite the same way. Nothing but love for Forrest. Let’s be clear. I am teasing. I consider him a friend.

Priyanka: He is amazing. Well honestly, like, I actually got to know Forrest when I decided to do this book. Wiley, who’s the publisher, sent me Forrest’s book, and he said, “You should look at this book because the idea that you are presenting to me, we could lay it out in this format.” Like, in the, you know, physical format. So, he sent me that book. And that’s how I know Forrest, honestly.

So, I told him that—this is a little story that I told him after. But anyway, yeah. I—the—[sigh]—I was going to make a point about the vid—the aspect of creating images, like, honestly, like, I designed the aspects of, like, how you layout information in the sketches, I studied a bunch of stuff to come up with, how do I make it precise and things like that. But there’s no way this book was possible without some design help. Like, I can’t possibly do the entire thing unless I have, like, five years. [laugh]. So—

Corey: Right on top of all of this, you do presumptively have a day job as well—and while—

Priyanka: Exactly.

Corey: This is definitely related. “I’m just going to go write a book.” “Oh, is it a dissertation?” “No, it’s going to look more like a children’s book than that,” is what they’re going to hear. And it’s yeah, I’m predicting some problems with the performance evaluation process at large companies when you start down those paths.

Priyanka: Exactly. So, I ended up, like, showing all these numbers, like, of the blog views and reads and social media, the presence of some of these images that were going wider. And in the GCPSketchnote GitHub repo got a huge number of stars. And it was like, everybody could see that writing a book would be amazing. From that point on, I was just like, I don’t think I can scale that.

So, when I was drawing—this is an example—when I drew my first sketch, it took me an entire weekend to just draw one sketch, which is what—I was only doing that the entire weekend—like, assume, like, 16 hours of work, just drawing the one sketch. So, if I went with that pace, this book was not possible. So, you know, after I had the idea laid out, had the process in place, I got some design help, which made it—which expedited the process much, much faster. [laugh].

Corey: There’s a lot to be said, for doing something that you enjoy. Do you do live sketchnoting during conference talks as well, or do you tend to not do it while someone is talking at a reasonably fast clip, and well, in 45 minutes, this had better be done, so let’s go. I’ve seen people who can do that, and I just marvel in awe at what they do.

Priyanka: I don’t do live. I don’t do live sketching. For me, paper and pen is a better medium so that’s just the medium that I like to work with. So, when the talk is happening, I’m actually taking notes on a pen and a paper. And then after, I can sketch it out, faster in a fast way.

Like, I did one sketchnote for Next 2020, I think, and that was done, like, a day after Next was over so I could take all the bits and pieces that were important and put it into that sketch. But I can’t do it live. That’s just one of the things I haven’t figured out yet. [laugh].

Corey: For me, I was always writing my email newsletter, so it was relatively rapid turnaround, and Twitter was interesting for me. I finally cracked the nut on how to express myself in a way that worked. The challenge that I ran into then was okay, there are thoughts I occasionally have that don’t lend themselves to then 140—now 280—characters, so I should probably start writing long-form. And then I want to start writing 1000 to 1500-word blog posts every week that goes out. And that forced me to become a better writer across the board. And then it became about one-upping myself, sort of, live-tweeting conference talks.

And the personal secret of why I do that is I’m ADHD in a bottle. Someone gets on stage—you say you zone out when you read a giant quantity of data; you prefer something more visual, more interactive. For me, I’m the opposite, where when someone gets on stage and starts talking, it’s, “Okay, get to—yes, you’re doing the intro of what a cloud might be. I get that point. This is supposed to be a more advanced talk. Can we speed it up a bit?”

And doing the live-tweeting about it, but not just relating what is said, but by making a joke about it, it’s how I keep myself engaged and from zoning out. Because let’s face it, this industry is extraordinarily boring, if you don’t bring a little bit of light to it.

Priyanka: Yeah, that is—

Corey: And that how to continue and how to do that was hard, and it took me time to get there.

Priyanka: Yeah. Yeah, no, I totally agree. Like, that’s exactly why I got into, like, training videos and sketches. Like, and videos and also. Like, I come up with, like, fake examples of companies that may or may not exist.

Like, I made up a dog shoe making company that ships out shoes when you need them and then return them and there’s a size and stuff, like, you have to come up with interesting things to make the content interesting because otherwise, this can get boring pretty quickly, which is going back to your example of, “Speed it up; get to the point.” [laugh].

Corey: This episode is sponsored in parts by our friend EnterpriseDB. EnterpriseDB has been powering enterprise applications with
PostgreSQL for 15 years. And now EnterpriseDB has you covered wherever you deploy PostgreSQL on premises, private cloud, and they just announced a fully managed service on AWS and Azure called BigAnimal, all one word.

Don't leave managing your database to your cloud vendor because they're too busy launching another half dozen manage databases to focus on any one of them that they didn't build themselves. Instead, work with the experts over at EnterpriseDB. They can save you time and money, they can even help you migrate legacy applications, including Oracle, to the cloud.

To learn more, try BigAnimal for free. Go to biganimal.com/snark, and tell them Corey sent you.

Corey: It’s always just fun to start experimenting with it, too, because all right, once I was done learn learning how to live-tweet other people’s talk and mostly get it correct because someone says something, I have three to five seconds to come up with what I want to talk about and maybe grab a picture and then move on to the next thing. And it’s easy to get that wrong and say things you don’t necessarily intend to and get taken the wrong way. I’ve mostly gotten past that. And—I’m not saying I’m always right, but I better than I used to be. And then it was okay, “How do I top this?”

And I started live-tweeting conference talks that I was giving live, which is always fun, but being able to pre-write some tweets at certain times, have certain webhooks in your slide deck and whatnot that fire these things off. And again, I’m not saying that he this is recommended or even a good idea, but it definitely wasn’t boring. And—

Priyanka: Yeah.

Corey: And continue to find ways to make the same type of material new and interesting is one of the challenges because the stuff is
complex.

Priyanka: Also bite-size, right? Like, it’s—I think Twitter is, like, the [unintelligible 00:15:54] words are obviously limiting, but it also forces you to think about it in bite-size, right? Like, okay, if I have a blog post then I’m summarizing it, how would I do it in two sentences? It forces me to think about it that way, which makes it very applicable to the time span that we have now, right, which is maybe, like, 30 seconds, you can have somebody on [unintelligible 00:16:18]

Corey: Attention is a rare and precious commodity.

Priyanka: Yeah. Yeah.

Corey: People who [unintelligible 00:16:21] engagement, I think that’s the wrong metric to go after because that inspires a whole bunch of terrible incentives, whereas finding something that is interesting, and a way to bring light to it and have a perspective on it that makes people think about it differently. For me, it’s been humor, but that’s my own approach to things. Your direction, it seems to be telling a story through visual arts. And that is something we don’t see nearly as much of.

Priyanka: Yeah. I think it’s also because it’s something that you—you know, like, I grew up drawing and painting. I was drawing since I was three years old, so that’s my way of thinking. Like, I don’t—I was talking to another devreloper the other day, and we were talking about—

Corey: It’s catching on. I love it.

Priyanka: —[laugh]. Two different ways of how we think. So, for me, when I design a piece of content, I have my visuals first, and then he was talking about when he designs his content, he has his bullet points and a blog post first. So, it’s like, two very different ways of approaching this similar thing. And then from that, from the images or the deck that I’m building up, I would come up with the narrative and stuff like that.

My thinking starts with images and narrative of tying, like, the images together. But it’s, that is the whole, like, fun of being in DevRel, right? Like, you are your own personality, and bringing whatever your personality, like you mentioned, humor and your case, art in my case, in somebody else’s case, it could be totally different thing, right? So, yeah.

Corey: Now, please correct me if I’m wrong on this, but an area of emphasis for you has been data analytics as well as Kubernetes, more or less things that are traditionally considered to be much more back-end if you’re looking at a spectrum of all things technology. Is that directionally accurate, or am I dramatically is understanding a lot of what you’re saying?

Priyanka: No, that’s very much accurate. I like to—I tend to be on the infrastructure back-and creating pipeline, creating easier processes, sort of person, not much into front-end. I dabble into it, but don’t enjoy it. [laugh].

Corey: This makes you something of a unicorn, in the sense of there are a tremendous number of devreloper types in the front-end slash JavaScript world because their entire career is focused on making things look visually appealing. That is what front-end is. I know this because I am rubbish at it. My idea of a well-designed interface that everyone looks at and smiles at [unintelligible 00:19:12] of command-line arguments when you’re writing a script for something. And it’s on a green screen, and sometimes I’ll have someone helped me coordinate to come up with a better color palette for the way that I’m looking at my terminal on my Mac. Real exciting times over here, I assure you.

So, the folks who are working in that space and they have beautifully designed slides, yeah, you tend to expect that. I gave a talk years ago at the front-end conference in Zurich, and I was speaking in the afternoon. And I went there and every presentation, slides were beautiful. And this was before I was working here and had a graphic designer on retainer to make my slides look not horrible. It was black Helvetica text on a white background, and I’m looking at this and I’m feeling ashamed that it’s—okay, I have two hours to fix this. What do I do?

I did the only thing I could think of; I changed Helvetica text to Comic Sans because if it’s going to look terrible and it’s going to be a designer thing that puts them off, you may as well go all-in. And that was a recurring meme at the time. I’ve since learned that there is an argument—I don’t know if it’s true or not—that Comic Sans is easier to read for folks with dyslexia, for example. And that’s fine. I don’t know if that’s accurate or not, but I stopped making jokes about it just because if people—even if it’s not true, and people believe that it’s, “Are you being unintentionally crappy to people?” It’s, “Well, I sure hope not. I’m rarely intentionally crappy. But when I do, I don’t want to be mistaken for not being.” It’s, save it up and use it when it counts.

Priyanka: Yeah, yeah. I’ve—yeah, I think, when it comes to these big events—and like front-end for me is—I would think, like, I actually thought that I would be great at front-end because I have interest in art and stuff. I do make things that [crosstalk 00:20:57]—

Corey: That’s my naive assumption, too. I’m learning as you speak here. Please continue[.

Priyanka: Yeah. And I was just—I thought that I would be and I have tried it, and I only like it to an extent, to present my idea. But I don’t like to go in deeper and, like, make my CSS pretty or make this—make it look pretty. I am very much intrigued by all the back-end stuff, and most of my experience, over the past ten years in Cloud has been in the back-end stuff, mainly just because I love APIs, I love—like, you know, as long as I can connect, or the idea of creating a demo or something that involves a bunch of APIs and a back-end, to present an idea in a front-end, I would work on that front-end. But otherwise, I’m not going to choose to do it. [laugh]. Which I found interesting for myself as well. It’s a realization. [laugh].

Corey: Every time I try and do something with front-end, it doesn’t matter the framework, I find myself more confused at the end than I was when I started. There’s something I don’t get. And anytime I see someone on Twitter, for example, talking about how a front-end is easier or somehow less than, I read that and I can’t help myself. It’s, “You ridiculous clown. You have no idea what you’re talking about.”

I don’t believe that I’m bad at all of the things under engineering—just most of them—and I think I pick things up reasonably quickly. It is a mystery that does not align with this, and if it’s easy for you, you don’t recognize—arguably—a skill that you have, but not everyone does, by a landslide. And that’s a human nature thing, too. It’s if it was easy for me, it’s obviously easy for everyone. If something’s hard for me, no one would understand how this works and the people that do are wizards from the future.

Priyanka: Yep. So true.

Corey: It never works that way.

Priyanka: Yeah. It never works that way. At least we have this in common, that you don’t like to work on front-ends. [laugh].

Corey: There’s that too. And I think that no matter where you fall on the spectrum of technology, I would argue that something that we all share in common is, it doesn’t matter how far we are down in the course of our entire career, from the very beginning to the very end, it is always a consistent, constant process of being humbled and made to feel like a fool by things you are supposedly professionally good at. And oh my stars, I’ve just learned to finally give up and embrace it. It’s like, “So, what’s going to make me feel dumb today?”

Priyanka: Exactly.

Corey: It’s the learn in public approach, which is important.

Priyanka: It’s so important. Especially, like, if you're thinking about it, like that’s the part of DevRel that makes it so exciting, too, right? Like, just
learning a new thing today and sharing it with you. Like, I’m not claiming that I’m an expert, but hey, let’s talk about it. And sure, I might end up looking dumb one day, I might end up looking smart the other day, but that’s not the point. The point is, I end up learning every day, right? And that’s the most important part, which is why I love this particular job, which is—what did we call it—devreloper.

Corey: Devreloping. And as a part of that, you’re talking to people constantly, be it people in the community and ecosystem, people who—you say you’ve talk to customers, but you also talk to these other folks. I would challenge you on that, where when you’re at a company like Google
Cloud, increasingly everyone in the community in the ecosystem is in one way or another, indistinguishable from being your customer; it all starts to converge at some point. All major cloud providers have that luxury, to be perfectly honest. What do you see in the ecosystem that people are struggling with as you talk to them?

And again, any one person is going to have a problem or bone to pick with some particular service or implementation, and okay, great. What I’m always interested in is what is the broad sweep of things? Because when I hear someone complaining that a given service from a given cloud provider is terrible. Okay, great. Everyone has an opinion. When I started to hear that four or five, six times, it’s okay, there’s something afoot here, and now I’m curious as to what it is. What patterns are you seeing emerge these days?

Priyanka: Yeah. I think more and more patterns along the lines of how can you make it automated? How can you make anything automated, right? Like, from machine learning’s perspective, how do I not need ML skills to build an ML model? Like, how can we get there faster, right?

Same for, like, in the infrastructure side, the serverless… aspect? How can you make it easy for me so I can just build an application and just deploy it so it becomes your problem to run it and not mine?

Corey: Oh, the—you are preaching to the choir on that. I feel like all of these services that talk about, “This is how you build and train a machine learning model,” yadda, yadda, it’s for an awful lot of the use cases out there, it’s exposing implementation details about which I could not possibly care less. It’s the, I want an API that I throw something at—like, be it a picture—and then I want to get a response of, “Yes, it’s a hot dog,” or, “That’s disgusting,” or whatever it is that it decides that it wants to say, great because that’s the business outcome I’m after, and I do not care what wizardry happens on the back-end, I don’t care if it’s people who are underpaid and working extremely quickly by hand to do it, as long as it’s from a business perspective, it hits a certain level of performance, reliability, et cetera. And then price, of course, yeah.

And that is not to say I’m in favor of exploiting people, let’s be clear here because I’m pretty sure most of these are not actually humans on the back-end, but okay. I just want that as the outcome that I think people are after, and so much of the conversation around how to build and train models and all misses the point because there are companies out there that need that, absolutely, there are, but there are a lot more that need the outcome, not the focus on this. And let’s face it, an awful lot of businesses that would benefit from this don’t have the budget to hire the team of incredibly expensive people it takes to effectively leverage these things because I have an awful lot of observations about people in machine learning space, one of them is absolutely not that, “Wow, I bet those people are inexpensive for me to hire.” It doesn’t work that way.

Priyanka: It doesn’t. Yeah. And so, yeah. I think the future of, like, the whole cloud space, like, when it started, we started with how can I run my server not in my basement, but somewhere else, right? Now, we are at a different stage where we have a different sets of problems and requirements for businesses, right?

And that’s where I see it growing. It’s like, how can I make this automated fast, not my problem? How can I make it not my problem is, like, the biggest [laugh] biggest, I think, theme that we are seeing, whether it’s infrastructure, data science, data analytics, in all of these spaces.

Corey: I get a lot of interesting feedback for my comparative takes on the various cloud providers, and one thing that I’ve said for a while about Google Cloud has been that its developer experience is unparalleled compared to basically anything else on the market. It makes things
just work, and that’s important because a bad developer experience has the unfortunate expression—at least for me—of, “Oh, this isn’t working the way I want it to. I must be dumb.” No, it’s a bad user experience for you. What I am seeing emerge as well from Google Cloud is an incredible emphasis—and I do think they’re aligned here—on storytelling, and doing so effectively.

You’re there communicating visually; Forrest is there, basically trying to be the me of Google Cloud—which is what I assume he’s doing; he would argue everything about that and he’d be right to do it, but that’s what I’m calling it because this is my show; he can come on and argue with me himself if he takes issue with it. But I love the emphasis on storytelling and unifying solutions and the rest, as opposed to throwing everything at the wall to see what sticks to it. I think there’s more intention being put into an awful lot of not just what you’re building, but how you’re talking about it, now it’s integrated with the other things that you’re building. That’s no small thing.

Priyanka: Yeah. That is so hard, especially when you know the cloud space; like, hundreds of products, they all have their unique requirement to solve a problem, but nobody cares, right? Like, as a consumer, I shouldn’t have to care that there are 127 products or whatever. It doesn’t matter to me as a consumer or customer, all that matters is whether I can solve my business problem with a set of your tools, right? So, that’s exactly why, like, we have this team that I work in that I’m a part of, which has an entire focus on storytelling.

We do YouTube videos with storytelling, we do art like this, I’ve also dabbled into comics a little bit. And we continue to go back to the drawing board with how else we can tell these stories. I know—I mentioned this to Forrest—I’m working on a song as well, which I have never done
before, and [laugh] I think I’m going to butcher it. I kind of have it ready for, like, six months but never released it, right, because I’m just too scared to do that. [laugh] but anyway.

Corey: Ship and then turn the internet off for a week and it’ll be gone regardless, by the time you come back. Problem solved until the reporters start calling, and then you have problems.

Priyanka: I might have to just do that, and be, like, you know what world? Keep saying whatever you want to say, I’m not here. [laugh]. But anyway, going back to that point of storytelling, and it’s so—I think we have weaved it into the process. And it’s going really well, and now we are investing more in, like, R&D and doing more of how we can tell stories in different ways.

Corey: I have to say, I’m a big fan of the way that you’re approaching this. If people want to learn more about what you’re up to—and arguably, as I argue they should get a copy of your book because it is glorious—where’s the best place to find you?

Priyanka: Thank you. Okay, so LinkedIn and Twitter are my platforms that I check every single day, so you can message me, connect with me, I am available as—my handle is pvergadia. I don’t know if they have [crosstalk 00:31:11]—

Corey: Oh, this is all going in the [show notes 00:31:13] you need not worry.

Priyanka: Okay, perfect. So yeah, I don’t have to spell it because my last name is hard. [laugh]. So, you’ll find it in the show notes. But yeah, you can connect with me there. And you will find at the top of both of my profiles, the link to order the book, so you can do it there.

Corey: Excellent. And I’ve already done so, and I’m just waiting for it to arrive. So, this is—it’s going to be an exciting read if nothing else. One of these days, I’d have to actually live-tweet a reading thereof. We’ll see how that plays out.

Priyanka: That would be amazing.

Corey: Be careful what you wish for. Some of the snark could be a little too cutting; we have to be cautious of that.

Priyanka: [laugh]. I’m always scared of your tweets. Like, do I want to read this or not? [laugh].

Corey: If nothing else, it at least tries to be funny. So, there is that.

Priyanka: Yes. Yes, for sure.

Corey: I really—

Priyanka: No, I’m excited. I’m excited for when you get a chance to read it and just tweet whatever you feel like, from, you know, all the bits and pieces that I’ve brought together. So, I would love to get your take. [laugh].

Corey: Oh, you will, one way or another. That’s one of those non-optional things. It’s one of the fun parts of dealing with me. It’s, “Aw crap. That shitposter is back again.” Like the kid outside of your yard just from across the street, staring at your house and pointing and it’s, “Oh, dear. Here we go.” Throwing stones.

Priyanka: [laugh]. I’m excited either way. [laugh].

Corey: He’s got a platypus with him this time. What’s going on? It happens. We deal with what we have to. Thank you so much for being so generous with your time. I appreciate it.

Priyanka: Thank you so much for having me. It was amazing. You are a celebrity, and I wanted to be, you know, a part of your show for a long
time, so I’m glad we’re able to make it work.

Corey: You are welcome back anytime.

Priyanka: I will. [laugh].

Corey: An absolute pleasure to talk with you. Thanks again.

Priyanka: Thank you.

Corey: Priyanka Vergadia staff developer—but you call it developer advocate—at Google Cloud. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on whatever platform you’re using to listen to this thing, whereas if you’ve hated it, please do the exact same thing, making sure to hit the like and subscribe buttons on the YouTubes because that’s where it is. But if you did hate it, also leave an insulting, angry comment but not using words. I want you to draw a picture telling me exactly what you didn’t like about this episode.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Amy

Amy Tobey has worked in tech for more than 20 years at companies of every size, working with everything from kernel code to user interfaces. These days she spends her time building an innovative Site Reliability Engineering program at Equinix, where she is a principal engineer. When she's not working, she can be found with her nose in a book, watching anime with her son, making noise with electronics, or doing yoga poses in the sun.

Links Referenced:

  • Equinix Metal: https://metal.equinix.com
  • Personal Twitter: https://twitter.com/MissAmyTobey
  • Personal Blog: https://tobert.github.io/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Vultr. Optimized cloud compute plans have landed at Vultr to deliver lightning-fast processing power, courtesy of third-gen AMD EPYC processors without the IO or hardware limitations of a traditional multi-tenant cloud server. Starting at just 28 bucks a month, users can deploy general-purpose, CPU, memory, or storage optimized cloud instances in more than 20 locations across five continents. Without looking, I know that once again, Antarctica has gotten the short end of the stick. Launch your Vultr optimized compute instance in 60 seconds or less on your choice of included operating systems, or bring your own. It’s time to ditch convoluted and unpredictable giant tech company billing practices and say goodbye to noisy neighbors and egregious egress forever. Vultr delivers the power of the cloud with none of the bloat. “Screaming in the Cloud” listeners can try Vultr for free today with a $150 in credit when they visit getvultr.com/screaming. That’s G-E-T-V-U-L-T-R dot com slash screaming. My thanks to them for sponsoring this ridiculous podcast.

Corey: Finding skilled DevOps engineers is a pain in the neck! And if you need to deploy a secure and compliant application to AWS, forgettaboutit! But that’s where DuploCloud can help. Their comprehensive no-code/low-code software platform guarantees a secure and compliant infrastructure in as little as two weeks, while automating the full DevSecOps lifestyle. Get started with DevOps-as-a-Service from DuploCloud so that your cloud configurations are done right the first time. Tell them I sent you and your first two months are free. To learn more visit: snark.cloud/duplo. Thats’s snark.cloud/D-U-P-L-O-C-L-O-U-D.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Every once in a while I catch up with someone that it feels like I’ve known for ages, and I realize somehow I have never been able to line up getting them on this show as a guest. Today is just one of those days. And my guest is Amy Tobey who has been someone I’ve been talking to for ages, even in the before-times, if you can remember such a thing. Today, she’s a Senior Principal Engineer at Equinix. Amy, thank you for finally giving in to my endless wheedling.

Amy: Thanks for having me. You mentioned the before-times. Like, I remember it was, like, right before the pandemic we had beers in San Francisco wasn’t it? There was Ian there—

Corey: Yeah, I—

Amy: —and a couple other people. It was a really great time. And then—

Corey: I vaguely remember beer. Yeah. And then—

Amy: And then the world ended.

Corey: Oh, my God. Yes. It’s still March of 2020, right?

Amy: As far as I know. Like, I haven’t checked in a couple years.

Corey: So, you do an awful lot. And it’s always a difficult question to ask someone, so can you encapsulate your entire existence in a paragraph? It’s—

Amy: [sigh].

Corey: —awful, so I’d like to give a bit more structure to it. Let’s start with the introduction: You are a Senior Principal Engineer. We know it’s high level because of all the adjectives that get put in there, and none of those adjectives are ‘associate’ or ‘beginner’ or ‘junior,’ or all the other diminutives that companies like to play games with to justify paying people less. And you’re at Equinix, which is a company that is a bit unlike most of the, shall we say, traditional cloud providers. What do you do over there and both as a company, as a person?

Amy: So, as a company Equinix, what most people know about is that we have a whole bunch of data centers all over the world. I think we have the most of any company. And what we do is we lease out space in that data center, and then we have a number of other products that people don’t know as well, which one is Equinix Metal, which is what I specifically work on, where we rent you bare-metal servers. None of that fancy stuff that you get any other clouds on top of it, there’s things you can get that are… partner things that you can add-on, like, you know, storage and other things like that, but we just deliver you bare-metal servers with really great networking. So, what I work on is the reliability of that whole system. All of the things that go into provisioning the servers, making them come up, making sure that they get delivered to the server, make sure the API works right, all of that stuff.

Corey: So, you’re on the Equinix cloud side of the world more so than you are on the building data centers by the sweat of your brow, as they say?

Amy: Correct. Yeah, yeah. Software side.

Corey: Excellent. I spent some time in data centers in the early part of my career before cloud ate that. That was sort of cotemporaneous with the discovery that I’m the hardware destruction bunny, and I should go to great pains to keep my aura from anything expensive and important, like, you know, the SAN. So—

Amy: Right, yeah.

Corey: Companies moving out of data centers, and me getting out was a great thing.

Amy: But the thing about SANs though, is, like, it might not be you. They’re just kind of cursed from the start, right? They just always were kind
of fussy and easy to break.

Corey: Oh, yeah. I used to think—and I kid you not—that I had a limited upside to my career in tech because I sometimes got sloppy and I was fairly slow at crimping ethernet cables.

Amy: [laugh].

Corey: That is very similar to growing up in third grade when it became apparent that I was going to have problems in my career because my handwriting was sloppy. Yeah, it turns out the future doesn’t look like we predicted it would.

Amy: Oh, gosh. Are we going to talk about, like, neurological development now or… [laugh] okay, that’s a thing I struggle with, too right, is I started typing as soon as they would let—in fact, before they would let me. I remember in high school, I had teachers who would grade me down for typing a paper out. They want me to handwrite it and I would go, “Cool. Go ahead and take a grade off because if I handwrite it, you’re going to take two grades off my handwriting, so I’m cool with this deal.”

Corey: Yeah, it was pretty easy early on. I don’t know when the actual shift was, but it became more and more apparent that more and more things are moving towards a world where you could type. And I was almost five when I started working on that stuff, and that really wound up changing a lot of aspects of how I started seeing things. One thing I think you’re probably fairly well known for is incidents. I want to be clear when I say that you are not the root cause as—“So, why are things broken?” “It’s Amy again. What’s she gotten into this time?” Great.

Amy: [laugh]. But it does happen, but not all the time.

Corey: Exa—it’s a learning experience.

Amy: Right.

Corey: You’ve also been deeply involved with SREcon and a number of—a lot of aspects of what I will term—and please don’t yell at me for this—SRE culture—

Amy: Yeah.

Corey: Which is sometimes a challenging thing to wind up describing or putting a definition around. The one that I’ve always been somewhat partial to is, “SRE is DevOps, except you worked at Google for a while.” I don’t know how necessarily accurate that is, but it does rile people up.

Amy: Yeah, it does. Dave Stanke actually did a really great talk at SREcon San Francisco just a couple weeks ago, about the DORA report. And the new DORA report, they split SRE out into its own function and kind of is pushing against that old model, which actually comes from Liz Fong-Jones—I think it’s from her, or older—about, like, class SRE implements DevOps, which is kind of this idea that, like, SREs make DevOps happen. Things have evolved, right, since then. Things have evolved since Google released those books, and we’re all just figured out what works and what doesn’t a little bit.

And so, it’s not that we’re implementing DevOps so much. In fact, it’s that ops stuff that kind of holds us back from the really high impact work that SREs, I think, should be doing, that aren’t just, like, fixing the problems, the symptoms down at the bottom layer, right? Like what we did as sysadmins 20 years ago. You know, we’d go and a lot of people are SREs that came out of the sysadmin world and still think in that mode, where it’s like, “Well, I set up the systems, and when things break, I go and I fix them.” And, “Why did the developers keep writing crappy code? Why do I have to always getting up in the middle of the night because this thing crashed?”

And it turns out that the work we need to do to make things more reliable, there’s a ceiling to how far away the platform can take us, right? Like, we can have the best platform in the world with redundancy, and, you know, nine-way replicated data storage and all this crazy stuff, and still if we put crappy software on top, it’s going to be unreliable. So, how do we make less crappy software? And for most of my career, people would be, like, “Well, you should test it.” And so, we started doing that, and we still have crappy software, so what’s going on here? We still have incidents.

So, we write more tests, and we still have incidents. We had a QA group, we still have incidents. We send the developers to training, and we still have incidents. So like, what is the thing we need to do to make things more reliable? And it turns out, most of it is culture work.

Corey: My perspective on this stems from being a grumpy old sysadmin. And at some point, I started calling myself a systems engineer or DevOps or production engineer, or SRE. It was all from my point of view, the same job, but you know, if you call yourself a sysadmin, you’re just asking for a 40% pay cut off the top.

Amy: [laugh].

Corey: But I still tended to view the world through that lens. I tended to be very good at Linux systems internals, for example, understanding system calls and the rest, but increasingly, as the DevOps wave or SRE wave, or Google-isation of the internet wound up being more and more of a thing, I found myself increasingly in job interviews, where, “Great, now, can you go wind up implementing a sorting algorithm on the whiteboard?” “What on earth? No.” Like, my lingua franca is shitty Bash, and no one tends to write that without a bunch of tab completions and quick checking with manpages—die.net or whatnot—on the fly as you go down that path.

And it was awful, and I felt… like my skill set was increasingly eroding. And it wasn’t honestly until I started this place where I really got into writing a fair bit of code to do different things because it felt like an orthogonal skill set, but the fullness of time, it seems like it’s not. And it’s a reskilling. And it made me wonder, does this mean that the areas of technology that I focused on early in my career, was that all a waste? And the answer is not really. Sometimes, sure, in that I don’t spend nearly as much time worrying about inodes—for example—as I once did. But every once in a while, I’ll run into something and I looked like a wizard from the future, but instead, I’m a wizard from the past.

Amy: Yeah, I find that a lot in my work, now. Sometimes things I did 20 years ago, come back, and it’s like, oh, yeah, I remember I did all that threading work in 2002 in Perl, and I learned everything the very, very, very hard way. And then, you know, this January, did some threading work to fix some stability issues, and all of it came flooding back, right? Just that the experiences really, more than the code or the learning or the text and stuff; more just the, like, this feels like threads [BLEEP]-ery. Is a diagnostic thing that sometimes we have to say.

And then people are like, “Can you prove it?” And I’m like, “Not really,” because it’s literally thread [BLEEP]-ery. Like, the definition of it is that there’s weird stuff happening that we can’t figure out why it’s happening. There’s something acting in the system that isn’t synchronized, that isn’t connected to other things, that’s happening out of order from what we expect, and if we had a clear signal, we would just fix it, but we don’t. We just have, like, weird stuff happening over here and then over there and over there and over there.

And, like, that tells me there’s just something happening at that layer and then have to go and dig into that right, and like, just basically charge through. My colleagues are like, “Well, maybe you should look at this, and go look at the database,” the things that they’re used to looking at and that their experiences inform, whereas then I bring that ancient toiling through the threading mines experiences back and go, “Oh, yeah. So, let’s go find where this is happening, where people are doing dangerous things with threads, and see if we can spot something.” But that came from that experience.

Corey: And there’s so much that just repeats itself. And history rhymes. The challenge is that, do you have 20 years of experience, or do you have one year of experience repeated 20 times? And as the tide rises, doing the same task by hand, it really is just a matter of time before your full-time job winds up being something a piece of software does. An easy example is, “Oh, what’s your job?” “I manually place containers onto specific hosts.” “Well, I’ve got news for you, and you’re not going to like it at all.”

Amy: Yeah, yeah. I think that we share a little bit. I’m allergic to repeated work. I don’t know if allergic is the right word, but you know, if I sit and I do something once, fine. Like, I’ll just crank it out, you know, it’s this form, or it's a datafile I got to write and I’ll—fine I’ll type it in and do the manual labor.

The second time, the difficulty goes up by ten, right? Like, just mentally, just to do it, be like, I’ve already done this once. Doing it again is anathema to everything that I am. And then sometimes I’ll get through it, but after that, like, writing a program is so much easier because it’s like exponential, almost, growth in difficulty. You know, the third time I have to do the same thing that’s like just typing the same stuff—like, look over here, read this thing and type it over here—I’m out; I can’t do it. You know, I got to find a way to automate. And I don’t know, maybe normal people aren’t driven to live this way, but it’s kept me from getting stuck in those spots, too.

Corey: It was weird because I spent a lot of time as a consultant going from place to place and it led to some weird changes. For example, “Oh, thank God, I don’t have to think about that whole messaging queue thing.” Sure enough, next engagement, it’s message queue time. Fantastic. I found that repeating myself drove me nuts, but you also have to be very sensitive not to wind up, you know, stealing IP from the people that you’re working with.

Amy: Right.

Corey: But what I loved about the sysadmin side of the world is that the vast majority of stuff that I’ve taken with me, lives in my shell config. And what I mean by that is I’m not—there’s nothing in there is proprietary, but when you have a weird problem with trying to figure out the best way to figure out which Ruby process is stealing all the CPU, great, turns out that you can chain seven or eight different shell commands together through a bunch of pipes. I don’t want to remember that forever. So, that’s the sort of thing I would wind up committing as I learned it. I don’t remember what company I picked that up at, but it was one of those things that was super helpful.

I have a sarcastic—it’s a one-liner, except no sane editor setting is going to show it in any less than three—of a whole bunch of Perl, piped into du, piped into the rest, that tells you one of the largest consumers of files in a given part of the system. And it rates them with stars and it winds up doing some neat stuff. I would never sit down and reinvent something like that today, but the fact that it’s there means that I can do all kinds of neat tricks when I need to. It’s making sure that as you move through your career, on some level, you’re picking up skills that are repeatable and applicable beyond one company.

Amy: Skills and tooling—

Corey: Yeah.

Amy: —right? Like, you just described the tool. Another SREcon talk was John Allspaw and Dr. Richard Cook talking about above the line; below the line. And they started with these metaphors about tools, right, showing all the different kinds of hammers.

And if you’re a blacksmith, a lot of times you craft specialized hammers for very specific jobs. And that’s one of the properties of a tool that they were trying to get people to think about, right, is that tools get crafted to the job. And what you just described as a bespoke tool that you had created on the fly, that kind of floated under the radar of intellectual property. [laugh].

So, let’s not tell the security or IP people right? Like, because there’s probably billions and billions of dollars of technically, like, made-up IP value—I’m doing air quotes with my fingers—you know, that’s just basically people’s shell profiles. And my God, the Emacs automation that people have done. If you’ve ever really seen somebody who’s amazing at Emacs and is 10, 20, 30, maybe 40 years of experience encoded in their emacs settings, it’s a wonder to behold. Like, I look at it and I go, “Man, I wish I could do that.”

It’s like listening to a really great guitar player and be like, “Wow, I wish I could play like them.” You see them just flying through stuff. But all that IP in there is both that person’s collection of wisdom and experience and working with that code, but also encodes that stuff like you described, right? It’s just all these little systems tricks and little fiddly commands and things we don’t want to remember and so we encode them into our toolset.

Corey: Oh, yeah. Anything I wound up taking, I always would share it with people internally, too. I’d mention, “Yeah, I’m keeping this in my shell files.” Because I disclosed it, which solves a lot of the problem. And also, none of it was even close to proprietary or anything like that. I’m sorry, but the way that you wind up figuring out how much of a disk is being eaten up and where in a more pleasing way, is not a competitive advantage. It just isn’t.

Amy: It isn’t to you or me, but, you know, back in the beginning of our careers, people thought it was worth money and should be proprietary. You know, like, oh, that disk-checking script as a competitive advantage for our company because there are only a few of us doing this work. Like, it was actually being able to, like, manage your—[laugh] actually manage your servers was a competitive advantage. Now, it’s kind of commodity.

Corey: Let’s also be clear that the world has moved on. I wound up buying a DaisyDisk a while back for Mac, which I love. It is a fantastic, pretty effective, “Where’s all the stuff on your disk going?” And it does a scan and you can drive and collect things and delete them when trying to clean things out. I was using it the other day, so it’s top of mind at the moment.

But it’s way more polished than that crappy Perl three-liner. And I see both sides, truly I do. The trick also, for those wondering [unintelligible 00:15:45], like, “Where is the line?” It’s super easy. Disclose it, what you’re doing, in those scenarios in the event someone is no because they believe that finding the right man page section for something is somehow proprietary.

Great. When you go home that evening in a completely separate environment, build it yourself from scratch to solve the problem, reimplement it and save that. And you’re done. There are lots of ways to do this. Don’t steal from your employer, but your employer employs you; they don’t own you and the way that you think about these problems.

Every person I’ve met who has had a career that’s longer than 20 minutes has a giant doc somewhere on some system of all of the scripts that they wound up putting together, all of the one-liners, the notes on, “Next time you see this, this is the thing to check.”

Amy: Yeah, the cheat sheet or the notebook with all the little commands, or again the Emacs config, sometimes for some people, or shell
profiles. Yeah.

Corey: Here’s the awk one-liner that I put that automatically spits out from an Apache log file what—the httpd log file that just tells me what are the most frequent talkers, and what are the—

Amy: You should probably let go of that one. You know, like, I think that one’s lifetime is kind of past, Corey. Maybe you—

Corey: I just have to get it working with Nginx, and we’re good to go.

Amy: Oh, yeah, there you go. [laugh].

Corey: Or S3 access logs. Perish the thought. But yeah, like, what are the five most high-volume talkers, and what are those relative to each other? Huh, that one thing seems super crappy and it’s coming from Russia. But that’s—hmm, one starts to wonder; maybe it’s time to dig back in.

So, one of the things that I have found is that a lot of the people talking about SRE seem to have descended from an ivory tower somewhere. And they’re talking about how some of the best-in-class companies out there, renowned for their technical cultures—at least externally—are doing these things. But there’s a lot more folks who are not there. And honestly, I consider myself one of those people who is not there. I was a competent engineer, but never a terrific one.

And looking at the way this was described, I often came away thinking, “Okay, it was the purpose of this conference talk just to reinforce how smart people are, and how I’m not,” and/or, “There are the 18 cultural changes you need to make to your company, and then you can do something kind of like we were just talking about on stage.” It feels like there’s a combination of problems here. One is making this stuff more accessible to folks who are not themselves in those environments, and two, how to drive cultural change as an individual contributor if that’s even possible. And I’m going to go out on a limb and guess you have thoughts on both aspects of that, and probably some more hit me, please.

Amy: So, the ivory tower, right. Let’s just be straight up, like, the ivory tower is Google. I mean, that’s where it started. And we get it from the other large companies that, you know, want to do conference talks about what this stuff means and what it does. What I’ve kind of come around to in the last couple of years is that those talks don’t really reach the vast majority of engineers, they don’t really apply to a large swath of the enterprise especially, which is, like, where a lot of the—the bulk of our industry sits, right? We spend a lot of time talking about the darlings out here on the West Coast in high tech culture and startups and so on.

But, like, we were talking about before we started the show, right, like, the interior of even just America, is filled with all these, like, insurance and banks and all of these companies that are cranking out tons of code and servers and stuff, and they’re trying to figure out the same problems. But they’re structured in companies where their tech arm is still, in most cases, considered a cost center, often is bundled under finance, for—that’s a whole show of itself about that historical blunder. And so, the tech culture is tend to be very, very different from what we experience in—what do we call it anymore? Like, I don’t even want to say West Coast anymore because we’ve gone remote, but, like, high tech culture we’ll say. And so, like, thinking about how to make SRE and all this stuff more accessible comes down to, like, thinking about who those engineers are that are sitting at the computers, writing all the code that runs our banks, all the code that makes sure that—I’m trying to think of examples that are more enterprise-y right?

Or shoot buying clothes online. You go to Macy’s for example. They have a whole bunch of servers that run their online store and stuff. They have internal IT-ish people who keep all this stuff running and write that code and probably integrating open-source stuff much like we all do. But when you go to try to put in a reliability program that’s based on the current SRE models, like SLOs; you put in SLOs and you start doing, like, this incident management program that’s, like, you know, you have a form you fill out after every incident, and then you [unintelligible 00:20:25] retros.

And it turns out that those things are very high-level skills, skills and capabilities in an organization. And so, when you have this kind of IT mindset or the enterprise mindset, bringing the culture together to make those things work often doesn’t happen. Because, you know, they’ll go with the prescriptive model and say, like, okay, we’re going to implement SLOs, we’re going to start measuring SLIs on all of the services, and we’re going to hold you accountable for meeting those targets. If you just do that, right, you’re just doing more gatekeeping and policing of your tech environment. My bet is, reliability almost never improves in those cases.

And that’s been my experience, too, and why I get charged up about this is, if you just go slam in these practices, people end up miserable, the practices then become tarnished because people experienced the worst version of them. And then—

Corey: And with the remote explosion as well, it turns out that changing jobs basically means their company sends you a different Mac, and the next Monday, you wind up signing into a different Slack team.

Amy: Yeah, so the culture really matters, right? You can’t cover it over with foosball tables and great lunch. You actually have to deliver tools that developers want to use and you have to deliver a software engineering culture that brings out the best in developers instead of demanding the best from developers. I think that’s a fundamental business shift that’s kind of happening. If I’m putting on my wizard hat and looking into the future and dreaming about what might change in the world, right, is that there’s kind of a change in how we do leadership and how we do business that’s shifting more towards that model where we look at what people are capable of and we trust in our people, and we get more out of them, the knowledge work model.

If we want more knowledge work, we need people to be happy and to feel engaged in their community. And suddenly we start to see these kind of generational, bigger-pie kind of things start to happen. But how do we get there? It’s not SLOs. It maybe it’s a little bit starting with incidents. That’s where I’ve had the most success, and you asked me about that. So, getting practical, incident management is probably—

Corey: Right. Well, as I see it, the problem with SLOs across the board is it feels like it’s a very insular community so far, and communicating it to engineers seems to be the focus of where the community has been, but from my understanding of it, you absolutely need buy-in at significantly high executive levels, to at the very least by you air cover while you’re doing these things and making these changes, but also to help drive that cultural shift. None of this is something I have the slightest clue how to do, let’s be very clear. If I knew how to change a company’s culture, I’d have a different job.

Amy: Yeah. [laugh]. The biggest omission in the Google SRE books was [Ers 00:22:58]. There was a guy at Google named Ers who owns availability for Google, and when anything is, like, in dispute and bubbles up the management team, it goes to Ers, and he says, “Thou shalt…” right? Makes the call. And that’s why it works, right?

Like, it’s not just that one person, but that system of management where the whole leadership team—there’s a large, very well-funded team with a lot of power in the organization that can drive availability, and they can say, this is how you’re going to do metrics for your service, and this is the system that you’re in. And it’s kind of, yeah, sure it works for them because they have all the organizational support in place. What I was saying to my team just the other day—because we’re in the middle of our SLO rollout—is that really, I think an SLO program isn’t [clear throat] about the engineers at all until late in the game. At the beginning of the game, it’s really about getting the leadership team on board to say, “Hey, we want to put in SLIs and SLOs to start to understand the functioning of our software system.” But if they don’t have that curiosity in the first place, that desire to understand how well their teams are doing, how healthy their teams are, don’t do it. It’s not going to work. It’s just going to make everyone miserable.

Corey: It feels like it’s one of those difficult to sell problems as well, in that it requires some tooling changes, absolutely. It requires cultural change and buy-in and whatnot, but in order for that to happen, there has to be a painful problem that a company recognizes and is willing to pay to make go away. The problem with stuff like this is that once you pay, there’s a lot of extra work that goes on top of it as well, that does not have a perception—rightly or wrongly—of contributing to feature velocity, of hitting the next milestone. It’s, “Really? So, we’re going to be spending how much money to make engineers happier? They should get paid an awful lot and they’re still complaining and never seem happy. Why do I care if they’re happy other than the pure mercenary perspective of otherwise they’ll quit?” I’m not saying that it’s not worth pursuing; it’s not a worthy goal. I am saying that it becomes a very difficult thing to wind up selling as a product.

Amy: Well, as a product for sure, right? Because—[sigh] gosh, I have friends in the space who work on these tools. And I want to be careful.

Corey: Of course. Nothing but love for all of those people, let’s be very clear.

Amy: But a lot of them, you know, they’re pulling metrics from existing monitoring systems, they are doing some interesting math on them, but what you get at the end is a nice service catalog and dashboard, which are things we’ve been trying to land as products in this industry for as long as I can remember, and—

Corey: “We’ve got it this time, though. This time we’ll crack the nut.” Yeah. Get off the island, Gilligan.

Amy: And then the other, like, risky thing, right, is the other part that makes me uncomfortable about SLOs, and why I will often tell folks that I talk to out in the industry that are asking me about this, like, one-on-one, “Should I do it here?” And it’s like, you can bring the tool in, and if you have a management team that’s just looking to have metrics to drive productivity, instead of you know, trying to drive better knowledge work, what you get is just a fancier version of more Taylorism, right, which is basically scientific management, this idea that we can, like, drive workers to maximum efficiency by measuring random things about them and driving those numbers. It turns out, that doesn’t really work very well, even in industrial scale, it just happened to work because, you know, we have a bloody enough society that we pushed people into it. But the reality is, if you implement SLOs badly, you get more really bad Taylorism that’s bad for you developers. And my suspicion is that you will get worse availability out of it than you would if you just didn’t do it at all.

Corey: This episode is sponsored by our friends at Revelo. Revelo is the Spanish word of the day, and its spelled R-E-V-E-L-O. It means “I reveal.” Now, have you tried to hire an engineer lately? I assure you it is significantly harder than it sounds. One of the things that Revelo has recognized is something I’ve been talking about for a while, specifically that while talent is evenly distributed, opportunity is absolutely not. They’re exposing a new talent pool to, basically, those of us without a presence in Latin America via their platform. It’s the largest tech talent marketplace in Latin America with over a million engineers in their network, which includes—but isn’t limited to—talent in Mexico, Costa Rica, Brazil, and Argentina. Now, not only do they wind up spreading all of their talent on English ability, as well as you know, their engineering skills, but they go significantly beyond that. Some of the folks on their platform are hands down the most talented engineers that I’ve ever spoken to. Let’s also not forget that Latin America has high time zone overlap with what we have here in the United States, so you can hire full-time remote engineers who share most of the workday as your team. It’s an end-to-end talent service, so you can find and hire engineers in Central and South America without having to worry about, frankly, the colossal pain of cross-border payroll and benefits and compliance because Revelo handles all of it. If you’re hiring engineers, check out revelo.io/screaming to get 20% off your first three months. That’s R-E-V-E-L-O dot I-O slash screaming.

Corey: That is part of the problem is, in some cases, to drive some of these improvements, you have to go backwards to move forwards. And it’s one of those, “Great, so we spent all this effort and money in the rest of now things are worse?” No, not necessarily, but suddenly are aware of things that were slipping through the cracks previously.

Amy: Yeah. Yeah.

Corey: Like, the most realistic thing about first The Phoenix Project and then The Unicorn Project, both by Gene Kim, has been the fact that companies have these problems and actively cared enough to change it. In my experience, that feels a little on the rare side.

Amy: Yeah, and I think that’s actually the key, right? It's for the culture change, and for, like, if you really looking to be, like, do I want to work at
this company? Am I investing my myself in here? Is look at the leadership team and be, like, do these people actually give a crap? Are they looking just to punt another number down the road?

That’s the real question, right? Like, the technology and stuff, at the point where I’m at in my career, I just don’t care that much anymore. [laugh]. Just… fine, use Kubernetes, use Postgres, [unintelligible 00:27:30], I don’t care. I just don’t. Like, Oracle, I might have to ask, you know, go to finance and be like, “Hey, can we spend 20 million for a database?” But like, nobody really asks for that anymore, so. [laugh].

Corey: As one does. I will say that I mostly agree with you, but a technology that I found myself getting excited about, given the time of the recording on this is… fun, I spent a bit of time yesterday—from when we’re recording this—teaching myself just enough Go to wind up being together a binary that I needed to do something actively ridiculous for my camera here. And I found myself coming away deeply impressed by a lot of things about it, how prescriptive it was for one, how self-contained for another. And after spending far too many years of my life writing shitty Perl, and shitty Bash, and worse Python, et cetera, et cetera, the prescriptiveness was great. The fact that it wound up giving me something I could just run, I could cross-compile for anything I need to run it on, and it just worked. It’s been a while since I found a technology that got me this interested in exploring further.

Amy: Go is great for that. You mentioned one of my two favorite features of Go. One is usually when a program compiles—at least the way I code in Go—it usually works. I’ve been working with Go since about 0.9, like, just a little bit before it was released as 1.0, and that’s what I’ve noticed over the years of working with it is that most of the time, if you have a pretty good data structure design and you get the code to compile, usually it’s going to work, unless you’re doing weird stuff.

The other thing I really love about Go and that maybe you’ll discover over time is the malleability of it. And the reason why I think about that more than probably most folks is that I work on other people’s code most of the time. And maybe this is something that you probably run into with your business, too, right, where you’re working on other people’s infrastructure. And the way that we encode business rules and things in the languages, in our programming language or our config syntax and stuff has a huge impact on folks like us and how quickly we can come into a situation, assess, figure out what’s going on, figure out where things are laid out, and start making changes with confidence.

Corey: Forget other people for a minute they’re looking at what I built out three or four years ago here, myself, like, I look at past me, it’s like, “What was that rat bastard thinking? This is awful.” And it’s—forget other people’s code; hell is your own code, on some level, too, once it’s slipped out of the mental stack and you have to re-explore it and, “Oh, well thank God I defensively wound up not including any comments whatsoever explaining what the living hell this thing was.” It’s terrible. But you’re right, the other people’s shell scripts are finicky and odd.

I started poking around for help when I got stuck on something, by looking at GitHub, and a few bit of searching here and there. Even these large, complex, well-used projects started making sense to me in a way that I very rarely find. It’s, “What the hell is that thing?” is my most common refrain when I’m looking at other people’s code, and Go for whatever reason avoids that, I think because it is so prescriptive about formatting, about how things should be done, about the vision that it has. Maybe I’m romanticizing it and I’ll hate it and a week from now, and I want to go back and remove this recording, but.

Amy: The size of the language helps a lot.

Corey: Yeah.

Amy: But probably my favorite. It’s more of a convention, which actually funny the way I’m going to talk about this because the two languages I work on the most right now are Ruby and Go. And I don’t feel like two languages could really be more different.

Syntax-wise, they share some things, but really, like, the mental models are so very, very different. Ruby is all the way in on object-oriented programming, and, like, the actual real kind of object-oriented with messaging and stuff, and, like, the whole language kind of springs from that. And it kind of requires you to understand all of these concepts very deeply to be effective in large programs. So, what I find is, when I approach Ruby codebase, I have to load all this crap into my head and remember, “Okay, so yeah, there’s this convention, when you do this kind of thing in Ruby”—or especially Ruby on Rails is even worse because they go deep into convention over configuration. But what that’s code for is, this code is accessible to people who have a lot of free cognitive capacity to load all this convention into their heads and keep it in their heads so that the code looks pretty, right?

And so, that’s the trade-off as you said, okay, my developers have to be these people with all these spare brain cycles to understand, like, why I would put the code here in this place versus this place? And all these, like, things that are in the code, like, very compact, dense concepts. And then you go to something like Go, which is, like, “Nah, we’re not going to do Lambdas. Nah”—[laugh]—“We’re not doing all this fancy stuff.” So, everything is there on the page.

This drives some people crazy, right, is that there’s all this boilerplate, boilerplate, boilerplate. But the reality is, I can read most Go files from top to the bottom and understand what the hell it’s doing, whereas I can go sometimes look at, like, a Ruby thing, or sometimes Python and e—Perl is just [unintelligible 00:32:19] all the time, right, it’s there’s so much indirection. And it just be, like, “What the [BLEEP] is going on? This is so dense. I’m going to have to sit down and write it out in longhand so I can understand what the developer was even doing here.” And—

Corey: Well, that’s why I got the Mac Studio; for when I’m not doing A/V stuff with it, that means that I’ll have one core that I can use for, you know, front-end processing and the rest, and the other 19 cores can be put to work failing to build Nokogiri in Ruby yet again.

Amy: [laugh].

Corey: I remember the travails of working with Ruby, and the problem—I have similar problems with Python, specifically in that—I don’t know if I’m special like this—it feels like it’s a SRE DevOps style of working, but I am grabbing random crap off a GitHub constantly and running it, like, small scripts other people have built. And let’s be clear, I run them on my test AWS account that has nothing important because I’m not a fool that I read most of it before I run it, but I also—it wants a different version of Python every single time. It wants a whole bunch of other things, too. And okay, so I use ASDF as my version manager for these things, which for whatever reason, does not work for the way that I think about this ergonomically. Okay, great.

And I wind up with detritus scattered throughout my system. It’s, “Hey, can you make this reproducible on my machine?” “Almost certainly not, but thank you for asking.” It’s like ‘Step 17: Master the Wolf’ level of instructions.

Amy: And I think Docker generally… papers over the worst of it, right, is when we built all this stuff in the aughts, you know, [CPAN 00:33:45]—

Corey: Dev containers and VS Code are very nice.

Amy: Yeah, yeah. You know, like, we had CPAN back in the day, I was doing chroots, I think in, like, ’04 or ’05, you know, to solve this problem, right, which is basically I just—screw it; I will compile an entire distro into a directory with a Perl and all of its dependencies so that I can isolate it from the other things I want to run on this machine and not screw up and not have these interactions. And I think that’s kind of what you’re talking about is, like, the old model, when we deployed servers, there was one of us sitting there and then we’d log into the server and be like, I’m going to install the Perl. You know, I’ll compile it into, like, [/app/perl 558 00:34:21] whatever, and then I’ll CPAN all this stuff in, and I’ll give it over to the developer, tell them to set their shebang to that and everything just works. And now we’re in a mode where it’s like, okay, you got to set up a thousand of those. “Okay, well, I’ll make a tarball.” [laugh]. But it’s still like we had to just—

Corey: DevOps, but [unintelligible 00:34:37] dev closer to ops. You’re interrelating all the time. Yeah, then Docker comes along, and add dev is, like, “Well, here’s the container. Good luck, asshole.” And it feels like it’s been cast into your yard to worry about.

Amy: Yeah, well, I mean, that’s just kind of business, or just—

Corey: Yeah. Yeah.

Amy: I’m not sure if it’s business or capitalism or something like that, but just the idea that, you know, if I can hand off the shitty work to some other poor schlub, why wouldn’t I? I mean, that’s most folks, right? Like, just be like, “Well”—

Corey: Which is fair.

Amy: —“I got it working. Like, my part is done, I did what I was supposed to do.” And now there’s a lot of folks out there, that’s how they work, right? “I hit done. I’m done. I shipped it. Sure. It’s an old [unintelligible 00:35:16] Ubuntu. Sure, there’s a bunch of shell scripts that rip through things. Sure”—you know, like, I’ve worked on repos where there’s hundreds of things that need to be addressed.

Corey: And passing to someone else is fine. I’m thrilled to do it. Where I run into problems with it is where people assume that well, my part was the hard part and anything you schlubs do is easy. I don’t—

Amy: Well, that’s the underclass. Yeah. That’s—

Corey: Forget engineering for a second; I throw things to the people over in the finance group here at The Duckbill Group because those people are wizards at solving for this thing. And it’s—

Amy: Well, that’s how we want to do things.

Corey: Yeah, specialization works.

Amy: But we have this—it’s probably more cultural. I don’t want to pick, like, capitalism to beat on because this is really, like, human cultural thing, and it’s not even really particularly Western. Is the idea that, like, “If I have an underclass, why would I give a shit what their experience
is?” And this is why I say, like, ops teams, like, get out of here because most ops teams, the extant ops teams are still called ops, and a lot of them have been renamed SRE—but they still do the same job—are an underclass. And I don’t mean that those people are below us. People are treated as an underclass, and they shouldn’t be. Absolutely not.

Corey: Yes.

Amy: Because the idea is that, like, well, I’m a fancy person who writes code at my ivory tower, and then it all flows down, and those people, just faceless people, do the deployment stuff that’s beneath me. That attitude is the most toxic thing, I think, in tech orgs to address. Like, if you’re trying to be like, “Well, our liability is bad, we have security problems, people won’t fix their code.” And go look around and you will find people that are treated as an underclass that are given codes thrown over the wall at them and then they just have to toil through and make it work. I’ve worked on that a number of times in my career.

And I think just like saying, underclass, right, or caste system, is what I found is the most effective way to get people actually thinking about what the hell is going on here. Because most people are just, like, “Well, that’s just the way things are. It’s just how we’ve always done it. The developers write to code, then give it to the sysadmins. The sysadmins deploy the code. Isn’t that how it always works?”

Corey: You’d really like to hope, wouldn’t you?

Amy: [laugh]. Not me. [laugh].

Corey: Again, the way I see it is, in theory—in theory—sysadmins, ops, or that should not exist. People should theoretically be able to write code as developers that just works, the end. And write it correct the first time and never have to change it again. Yeah. There’s a reason that I always like to call staging environments in places I work ‘theory’ because it works in theory, but not in production, and that is fundamentally the—like, that entire job role is the difference between theory and practice.

Amy: Yeah, yeah. Well, I think that’s the problem with it. We’re already so disconnected from the physical world, right? Like, you and I right now are talking over multiple strands of glass and digital transcodings and things right now, right? Like, we are detached from the physical reality.

You mentioned earlier working in data centers, right? The thing I miss about it is, like, the physicality of it. Like, actually, like, I held a server in my arms and put it in the rack and slid it into the rails. I plugged into power myself; I pushed the power button myself. There’s a server there. I physically touched it.

Developers who don’t work in production, we talked about empathy and stuff, but really, I think the big problem is when they work out in their idea space and just writing code, they write the unit tests, if we’re very lucky, they’ll write a functional test, and then they hand that wad off to some poor ops group. They’re detached from the reality of operations. It’s not even about accountability; it’s about experience. The ability to see all of the weird crap we deal with, right? You know, like, “Well, we pushed the code to that server, but there were three bit flips, so we had to do it again. And then the other server, the disk failed. And on the other server…” You know? [laugh].

It’s just, there’s all this weird crap that happens, these systems are so complex that they’re always doing something weird. And if you’re a developer that just spends all day in your IDE, you don’t get to see that. And I can’t really be mad at those folks, as individuals, for not understanding our world. I figure out how to help them, and the best thing we’ve come up with so far is, like, well, we start giving this—some responsibility in a production environment so that they can learn that. People do that, again, is another one that can be done wrong, where it turns into kind of a forced empathy.

I actually really hate that mode, where it’s like, “We’re forcing all the developers online whether they like it or not. On-call whether they like it or not because they have to learn this.” And it’s like, you know, maybe slow your roll a little buddy because the stuff is actually hard to learn. Again, minimizing how hard ops work is. “Oh, we’ll just put the developers on it. They’ll figure it out, right? They’re software engineers. They’re probably smarter than you sysadmins.” Is the unstated thing when we do that, right? When we throw them in the pit and be like, “Yeah, they’ll get it.” [laugh].

Corey: And that was my problem [unintelligible 00:39:49] the interview stuff. It was in the write code on a whiteboard. It’s, “Look, I understood how the system fundamentally worked under the hood.” Being able to power my way through to get to an outcome even in language I don’t know, was sort of part and parcel of the job. But this idea of doing it in artificially constrained environment, in a language I’m not super familiar with, off the top of my head, it took me years to get to a point of being able to do it with a Bash script because who ever starts with an empty editor and starts getting to work in a lot of these scenarios? Especially in an ops role where we’re not building something from scratch.

Amy: That’s the interesting thing, right? In the majority of tech work today—maybe 20 years ago, we did it more because we were literally building the internet we have today. But today, most of the engineers out there working—most of us working stiffs—are working on stuff that already exists. We’re making small incremental changes, which is great that’s what we’re doing. And we’re dealing with old code.

Corey: We’re gluing APIs together, and that’s fine. Ugh. I really want to thank you for taking so much time to talk to me about how you see all these things. If people want to learn more about what you’re up to, where’s the best place to find you?

Amy: I’m on Twitter every once in a while as @MissAmyTobey, M-I-S-S-A-M-Y-T-O-B-E-Y. I have a blog I don’t write on enough. And there’s a couple things on the Equinix Metal blog that I’ve written, so if you’re looking for that. Otherwise, mainly Twitter.

Corey: And those links will of course be in the [show notes 00:41:08]. Thank you so much for your time. I appreciate it.

Amy: I had fun. Thank you.

Corey: As did I. Amy Tobey, Senior Principal Engineer at Equinix. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, or on the YouTubes, smash the like and subscribe buttons, as the kids say. Whereas if you’ve hated this episode, same thing, five-star review all the platforms, smash the buttons, but also include an angry comment telling me that you’re about to wind up subpoenaing a copy of my shell script because you’re convinced that your intellectual property and secrets are buried within.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Tomasz

Tomasz is a Frontend Engineer at Stedi, Co-Founder/Head of React at Cloudash, egghead.io instructor with over 200 lessons published, a tech speaker, an AWS Community Hero and a lifelong learner.

Links Referenced:

  • Cloudash: https://cloudash.dev/
  • Twitter: https://twitter.com/tlakomy

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Honeycomb. When production is running slow, it’s hard to know where problems originate. Is it your application code, users, or the underlying systems? I’ve got five bucks on DNS, personally. Why scroll through endless dashboards while dealing with alert floods, going from tool to tool to tool that you employ, guessing at which puzzle pieces matter? Context switching and tool sprawl are slowly killing both your team and your business. You should care more about one of those than the other; which one is up to you. Drop the separate pillars and enter a world of getting one unified understanding of the one thing driving your business: production. With Honeycomb, you guess less and know more. Try it for free at honeycomb.io/screaminginthecloud. Observability: it’s more than just hipster monitoring.

Corey: This episode is sponsored in part by our friends at ChaosSearch. You could run Elasticsearch or Elastic Cloud—or OpenSearch as they’re calling it now—or a self-hosted ELK stack. But why? ChaosSearch gives you the same API you’ve come to know and tolerate, along with unlimited data retention and no data movement. Just throw your data into S3 and proceed from there as you would expect. This is great for IT operations folks, for app performance monitoring, cybersecurity. If you’re using Elasticsearch, consider not running Elasticsearch. They’re also available now in the AWS marketplace if you’d prefer not to go direct and have half of whatever you pay them count towards your EDB commitment. Discover what companies like Equifax, Armor Security, and Blackboard already have. To learn more, visit chaossearch.io and tell them I sent you just so you can see them facepalm, yet again.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. It’s always a pleasure to talk to people who ask the bold questions. One of those great bold questions is, what if CloudWatch’s web page didn’t suck? It’s a good question. It’s one I ask myself all the time.

And then I stumbled across a product that wound up solving this for me, and I’m a happy customer. To be clear, they’re not sponsoring anything that I do, nor should they. It’s one of those bootstrapped, exciting software projects called Cloudash. Today, I’m joined by the Head of React at Cloudash, Tomasz Łakomy. Tomasz, thank you for joining me.

Tomasz: It’s a pleasure to be here.

Corey: So, where did this entire idea come from? Because I sit and I get upset every time I have to go into the CloudWatch dashboard because first, something’s broken. In an ideal scenario, I don’t have to care about monitoring or observability or anything like that. But then it’s quickly overshadowed by the fact that this interface is terrible. And the reason I know it’s terrible is that every time I’m in there, I feel dumb.

My belief is—for the longest time, I thought that was a problem with me. But no, invariably, when you wind up working with something and consistently finding it a bad—you don’t know enough to solve for it, it’s not you. It is, in fact, the signs of a poorly designed experience, start to finish. “You should be smarter to use this tool,” is very rarely correct. And there are a bunch of observability tools and monitoring tools for serverless things that have made sense over the years and made this easier, but one of the most—and please don’t take this the wrong way—stripped down, bare essentials of just the facts, style of presentation is Cloudash. It’s why I continue to pay for it every month with a smile on my face. How did you get here from there?

Tomasz: Yeah that’s a good question. I would say that. Cloudash was born out of desire for simple things to be simple. So, as you mentioned, Cloudash is basically the monitoring and troubleshooting tool for serverless applications, made for serverless developers because I am very much into serverless space, as is Maciej Winnicki, who is the another half of Cloudash team. And, you know, the whole premise of serverless was things are going to be simpler, right?

So, you know, you have a bunch of code, you’re going to dump it into a Lambda function, and that’s it. You don’t have to care about servers, you don’t have to care about, you know, provisioning stuff, you don’t have to care about maintenance, and so on. And that is not exactly true because why PagerDuty still continues to be [unintelligible 00:02:56] business even in serverless spaces. So, you will get paged every now and then. The problem is—what we kind of found is once you have an incident—you know, PagerDuty always tends to call it in the middle of the night; it’s never, like, 11 a.m. during the workday; it’s always the middle of the night.

Corey: And no one’s ever happy when it calls them either. It’s, “Ah, hell.” Whatever it rings, it’s yeah, the original Call of Duty. PagerDuty hooked up to Nagios. I am old enough to remember those days.

Tomasz: [unintelligible 00:03:24] then business, like, imagine paying for something that’s going to wake you up in the middle of the night. It
doesn’t make sense. In any case—

Corey: “So, why do you pay for that product? Because it’s really going to piss me off.” “Okay, well… does that sound like a good business to you? Well, AWS seems to think so. No one’s happy working with that stuff.” “Fair. Fair enough.”

Tomasz: So, in any case, like we’ve established an [unintelligible 00:03:43]. So you wake up, you go to AWS console because you saw a notification that this-and-this API has, you know, this threshold was above it, something was above the threshold. And then you go to the CloudWatch console. And then you see, okay, those are the logs, those are the metrics. I’m going to copy this request ID. I’m going to go over here. I’m going to go to X-Ray.

And again, it’s 3 a.m. so you don’t exactly remember what do you investigate; you have, like, ten minutes. And this is a problem. Like, we’ve kind of identified that it’s not simple to do these kinds of things, too—it’s not simple to open something and have an understanding, okay, what exactly is happening in my serverless app at this very moment? Like, what’s going on?

So, we’ve built that. So, Cloudash is a desktop app; it lives on your machine, which is a single pane of glass. It’s a single pane of glass view into your serverless system. So, if you are using CloudFormation in order to provision something, when you open Cloudash, you’re going to see, you know, all of the metrics, all the Lambda functions, all of the API Gateways that you have provisioned. As of yesterday, API Gateway is no longer cool because they did launch the direct integration, so you have—you can call Lambda functions with [crosstalk 00:04:57]—

Corey: Yeah, it’s the one they released, and then rolled back and somehow never said a word—because that’s an AWS messaging story, and then some—right around re:Invent last year. And another quarter goes by and out it goes.

Tomasz: It’s out yesterday.

Corey: Yeah, it’s terrific. I love that thing. The only downside to it is, ah, you have to use one of their—you have to use their domain; no custom domain support. Really? Well, you can hook up CloudFront to it, but the pricing model that way makes it more expensive than API Gateway.

Okay, so I could use Cloudflare in front of it, and then it becomes free, so I bought a domain just for that purpose. That’s right, my serverl—my direct Lambda URLs now live behind the glorious domain of cheapass.cloud because of course. They are. It’s a day-one product from AWS, so of course, it’s not feature-complete.

But one of the things I like about the serverless model, and it’s also a challenge when it comes to troubleshooting stuff is that it’s very much set it and forget it style because serverless in many cases, at least the way that I tend to use it, is back-office stuff, its back-end things, it’s processing on things that are not necessarily always direct front and center. So, these things can run on their own for years until finally, you find a strange bug in a new use case, or you want to go and change something. And then it’s how the hell did this ever work? And it’s still working, kind of, but what fool built this? Of course, it was me; it’s always me.

But what happened here? You’re basically excavating your own legacy code, trying to understand what’s going on. And so, you’re already upset then. Cloudash makes this easier to find the things, to navigate through a whole bunch of different accounts. And there are a bunch of decisions that you made while building the app that are so clearly correct, that I get actively annoyed when others don’t because oh, it looks at your AWS configuration file in your user home directory. Great, awesome. It’s a desktop app, but it still consults that file. Yay, integration between ClickOps and the terminal. Wonderful.

But ah, use SSO for a lot of stuff, so that’s going to fix your little red wagon. I click on that app, and suddenly, bam, a browser opens asking me to log in and authenticate, allow the request. It works, and then suddenly, it goes back to doing exactly what you’d expect it to. It’s really nice. The affordances behind this are glorious.

Tomasz: Like I said, one of our kind of design goals when building Cloudash was to make simple things simple again. The whole purpose is to make sure that you can get into the root cause of an issue within, like, five minutes, if not less. And this is kind of the app that you’re going to tend to open whenever that—as I said, because some of the systems can be around for, like, ages, literally without any incident whatsoever, then the data is going to change because somebody [unintelligible 00:07:30] got that the year is 2020 and off you go, we have an incident.

But what’s important about Cloudash is that we don’t send logs anywhere. And that’s kind of important because you don’t pay for [PUT 00:07:42] metric API because we are not sending those logs anywhere. If you install Cloudash on your machine, we are not going to get your logs from the last ten years, put them in into a system, charge you for that, just so you are able to, you know, find out what happened in this particular hour, like, two weeks ago. We genuinely don’t care about your logs; we have enough of our own logs at work to, you know, to analyze, to investigate, and so on; we are not storing them anywhere.

In fact, you know, whatever happens on your machine stays on the machine. And that is partially why this is a desktop app. Because we don’t want to handle your credentials. We don’t—absolutely, we don’t want you to give us any of your credentials or access keys, you know, whatever. We don’t want that.

So, that is why you install Cloudash, it’s going to run on your machine, it’s going to use your local credentials. So, it’s… effectively, you could say that this is a much more streamlined and much more laser-focused browser or like, an eye into AWS systems, which live on the serverless side of things.

Corey: I got to deal with it in a bit of an interesting way, recently. I have a detector in my company’s production AWS org, to detect when ClickOps is afoot. Now, I’m a big proponent of ClickOps, but I also want to know what’s going on, so I have a whole thing that [runs detects 00:09:04] when people are doing things in the console versus via API. And it alerts on certain subsets of them. I had to build a special case for the user agent string coming out of Cloudash because no, no, this is an app, this is not technically ClickOps—it is also read-only, which is neither here nor there, to my understanding.

But it was, “Oh yeah, this is effectively an Electron app.” It just wraps, effectively, a browser and presents that as an application. And cool. From my perspective, that’s an implementation detail. It feels like a native app—because it is—and I can suddenly see the things I care about in a way that is much more straightforward without having to have four different browser tabs open where, okay, here’s the CloudTrail log for this thing, here’s the metrics next to it. Oh, those are two separate windows already, and so on and so forth. It just makes hunting down to the obnoxious problems so much nicer.

It’s also, you’re one of those rare products where if I don’t use it for a month, I don’t get the bill at the end of the month and think, “Ooh, that’s going to—did I waste the money?” It’s no, nice. I had a whole month where I didn’t have to mess with this. It’s great.

Tomasz: Exactly. I feel like, you know, it’s one of those systems where, as you said, we send you an email at the end of every month that we’re going to charge you X dollars for the month—by the way, we have fixed pricing and then you can cancel anytime—and it’s like one of those things that, you know, I didn’t have to open this up for a month. This is awesome because I didn’t have any incidents. But I know whenever again, PagerDuty is going to decide, “Hey, dude, wake up. You know, if slept for three hours. That is definitely long enough,” then you know that; you know, this app is there and you can use that.

We very much care about, you know, building this stuff, not only for our customers, but we also use that on a daily basis. In fact, I… every single time that I have to—I want to investigate something in, like, our serverless systems at Stedi because everything that we do at work, at Stedi, since this incident serverless paradigm. So, I tend to open Cloudash, like, 95% of the time whenever I want to investigate something. And whenever I am not able to do something in Cloudash, this goes, like, straight to the top of our, you know, issue lists or backlog or whatever you want to call it. Because we want to make this product, not only awesome, you know, for customers to buy a [unintelligible 00:11:22] or whatever, but we also want to be able to use that on a daily basis.

And so far, I think we’ve kind of succeeded. But then again, we have quite a long way to go because we have more ideas, than we have the time, definitely, so we have to kind of prioritize what exactly we’re going to build. So, [unintelligible 00:11:39] integrations with alarms. So, for instance, we want to be able to see the alarms directly in the Cloudash UI. Secondly, integration with logs insights, and many other ideas. I could probably talk for hours about what we want to build.

Corey: I also want to point out that this is still your side gig. You are by day a front-end engineer over at Stedi, which has a borderline disturbing number of engineers with side gigs, generally in the serverless space, doing interesting things like this. Dynobase is another example, a DynamoDB desktop client; very similar in some respects. I pay for that too. Honestly, for a company in Stedi’s space, which is designed as basically a giant API for deep, large enterprise business stuff, there’s an awful lot of stuff for small-scale coming out of that.

Like, I wind up throwing a disturbing amount of money in the general direction of Stedi for not being their customer. But there’s something about the culture that you folks have built over there that’s just phenomenal.

Tomasz: Yeah. For the record, you know, having a side gig is another part of interview process at Stedi. You don’t have to have [laugh] a side project, but yeah, you’re absolutely right, you know, the amount of kind of side projects, and you know, some of those are monetized, as you mentioned, you know, Cloudash and Dynobase and others. Some of those—because for instance, you talked to Aidan, I think a couple of weeks ago about his shenanigans, whenever you know, AWS is going to announce something he gets in and try to [unintelligible 00:13:06] this in the most amusing ways possible. Yeah, I mean, I could probably talk for ages about why Stedi is by far the best company I’ve ever worked at, but I’m going to say this: that this is the most talented group of people I’ve ever met, and myself, honestly.

And, you know, the fact that I think we are the second largest, kind of, group of AWS experts outside of AWS because the density of AWS Heroes, or ex-AWS employees, or people who have been doing cloud stuff for years, is frankly, massive, I tend to learn something new about cloud every single day. And not only because of the Last Week in AWS but also from our Slack.

Corey: This episode is sponsored by our friends at Oracle Cloud. Counting the pennies, but still dreaming of deploying apps instead of “Hello, World” demos? Allow me to introduce you to Oracle’s Always Free tier. It provides over 20 free services and infrastructure, networking, databases, observability, management, and security. And—let me be clear here—it’s actually free. There’s no surprise billing until you intentionally and proactively upgrade your account. This means you can provision a virtual machine instance or spin up an autonomous database that manages itself, all while gaining the networking, load balancing, and storage resources that somehow never quite make it into most free tiers needed to support the application that you want to build. With Always Free, you can do things like run small-scale applications or do proof-of-concept testing without spending a dime. You know that I always like to put asterisks next to the word free? This is actually free, no asterisk. Start now. Visit snark.cloud/oci-free that’s snark.cloud/oci-free.

Corey: There’s something to be said for having colleagues that you learn from. I have never enjoyed environments where I did not actively feel like the dumbest person in the room. That’s why I love what I do now. I inherently am. I have to talk about so many different things, that whenever I talk to a subject matter expert, it is a certainty that they know more about the thing than I do, with the admitted and depressing exception of course of the AWS bill because it turns out the reason I had to start becoming the expert in that was because there weren’t any. And here we are now.

I want to talk as well about some of—your interaction outside of work with AWS. For example, you’ve been an Egghead instructor for a while with over 200 lessons that you published. You’re an AWS Community Hero, which means you have the notable distinction of volunteering for a for-profit company—good work—no, the community is very important. It’s helping each other make sense of the nonsense coming out of there. You’ve been involved within the ecosystem for a very long time. What is it about, I guess—the thing I’m wondering about myself sometimes—what is it about the AWS universe that drew you in, and what keeps you here?

Tomasz: So, give you some context, I’ve started, you know, learning about the cloud and AWS back in early-2019. So, fun fact: Maciej Winnicki—again, the co-founder of Cloudash—was my manager at the time. So, we were—I mean, the company I used to work for at the time, OLX Group, we are in the middle of cloud transformation, so to speak. So, going from, you know, on-premises to AWS. And I was, you know, hired as a senior front-end engineer doing, you know, all kinds of front-end stuff, but I wanted to grow, I wanted to learn more.

So, the idea was, okay, maybe you can get AWS Certified because, you know, it’s one of those corporate goals that you have to have something to put that checkbox next to it. So, you know, getting certified, there you go, you have a checkbox. And off you go. So, I started, you know, diving in, and I saw this whole ocean of things that, you know, I was not entirely aware of. To be fair, at the time I knew about this S3, I knew that you can put a file in an S3 bucket and then you can access it from the internet. This is, like, the [unintelligible 00:16:02] idea of my AWS experiences.

Corey: Ideally, intentionally, but one wonders sometimes.

Tomasz: Yeah, exactly. That is why you always put stuff as public, right? Because you didn’t have to worry about who [unintelligible 00:16:12] [laugh] public [unintelligible 00:16:15]. No, I’m kidding, of course. But still, I think what’s [unintelligible 00:16:20] to AWS is what—because it is this endless ocean of things to learn and things to play with, and, you know, things to teach.

I do enjoy teaching. As you said, I have quite a lot of, you know, content, videos, blog posts, conference talks, and a bunch of other stuff, and I do that for two reasons. You know, first of all, I tend to learn the best by teaching, so it helps me very much, kind of like, solidify my own knowledge. Whenever I record—like, I have two courses about CDK, you know, when I was recording those, I definitely—that kind of solidify my, you know, ideas about CDK, I get to play with all those technologies.

And secondly, you know, it’s helpful for others. And, you know, people have opinions about certificates, and so on and so forth, but I think that for somebody who’s trying to get into either the tech industry or, you know, cloud stuff in general, being certified helps massively. And I’ve heard stories about people who are basically managed to double or triple their salaries by going into tech, you know, with some of those certificates. That is why I strongly believe, by the way, that those certificates should be free. Like, if you can pass the exam, you shouldn’t have to worry about this $150 of the fee.

Corey: I wrote a blog post a while back, “The Dumbest Dollars a Cloud Provider Can Make,” and it’s charging for training and certification because if someone’s going to invest that kind of time in learning your platform, you’re going to try and make $150 bucks off them? Which in some cases, is going to put people off from even beginning that process. “What cloud provider I’m not going to build a project on?” Obviously, the one I know how to work with and have a familiarity with, in almost every case. And the things you learn in your spare time as an independent learner when you get a job, you tend to think about your work the same way. It matters. It’s an early on-ramp that pays off down the road and the term of years.

I used to be very anti-cert personally because it felt like I was jumping through hoops, and paying, in some cases, for the privilege. I had a CCNA for a while from Cisco. There were a couple of smaller companies, SaltStack, for example, that I got various certifications from at different times. And that was sort of cheating because I helped write the software, but that’s neither here nor there. It’s the—and I do have a standing AWS cert that I get a different one every time—mine is about to expire—because it gets me access to lounges at physical events, which is the dumbest of all reasons to get certs, but here you go. I view it as the $150 lounge pass with a really weird entrance questionnaire.

But in my case it certs don’t add anything to what I do. I am not the common case. I am not early in my career. Because as you progress through your career, things—there needs to be a piece of paper that says you know things, and early on degree or certifications are great at that. In the time it becomes your own list of experience on your resume or CV or LinkedIn or God knows what. Polywork if you’re doing it the right way these days.

And it shows a history of projects that are similar in scope and scale and impact to the kinds of problems that your prospective employer is going to have to solve themselves. Because the best answer to hear—especially in the ops world—when there’s a problem is, “Oh, I’ve seen this before. Here’s how you fix it.” As opposed to, “Well, I don’t know. Let me do some research.”

There’s value to that. And I don’t begrudge anyone getting certs… to a point. At least that’s where I sit on it. At some point when you have 25 certs, it’s when you actually do any work? Because it’s taking the tests and learning all of these things, which in many ways does boil down to
trivia, it stands in counterbalance to a lot of these things.

Tomasz: Yeah. I mean, I definitely, totally agree. I remember, you know, going from zero to—maybe not Hero; I’m not talking about AWS Hero—but going from zero to be certified, there was the Solutions Architect Associate. I think it took me, like, 200 hours. I am not the, you know, the brightest, you know, the sharpest tool in the shed, so it probably took me, kind of, somewhat more.

I think it’s doable in, like, 100 hours, but I tend to over-prepare for stuff, so I didn’t actually take the actual exam until I was able to pass the sample exams with, like, 90% pass, just to be extra sure that I’m actually going to pass it. But still, I think that, you know, at some point, you probably should focus on, you know, getting into the actual stuff because I hold two certificates, you know, one of those is going to expire, and I’m not entirely sure if I want to go through the process again. But still, if AWS were to introduce, like, a serverless specialty exam, I would be more than happy to have that. I genuinely enjoy, kind of, serverless, and you know, the fact that I would be able to solidify my knowledge, I have this kind of established path of the things that I should learn about in order to get this particular certificate, I think this could be interesting. But I am not probably going to chase all the 12 certificates.

Maybe if AWS IQ was available in Poland, maybe that would change because I do know that with IQ, those certs do matter. But as of
[unintelligible 00:21:26] now, I’m quite happy with my certs that I have right now.

Corey: Part of the problem, too, is the more you work with these things, the harder it becomes to pass the exams, which sounds weird and counterintuitive, but let me use myself as an example. When I got the cloud practitioner cert, which I believe has lapsed since then, and I got one of the new associate-level betas—I’ll keep moving up the stack until I start failing exams. But I got a question wrong on the cloud practitioner because it was, “How long does it take to restore an RDS database from a snapshot backup?” And I gave the honest answer of what I’ve seen rather than what it says in the book, and that honest answer can be measured in days or hours. Yeah.

And no, that’s not the correct answer. Yeah, but it is the real one. Similarly, a lot of the questions get around trivia, syntax of which of these is the correct argument, and which ones did we make up? It’s, I can explain in some level of detail, virtually every one of AWS has 300 some-odd services to you. Ask me about any of them, I could tell you what it is, how it works, how it’s supposed to work and make a dumb joke about it. Fine, whatever.

You’ll forgive me if I went down that path, instead of memorizing what is the actual syntax of this YAML construct inside of a CloudFormation template? Yeah, I can get the answer to that question in the real world, with about ten seconds of Googling and we move on. That’s the way most of us learn. It’s not cramming trivia into our heads. There’s something broken about the way that we do certifications, and tech interviews in many cases as well.

I look back at some of the questions I used to ask people for Linux sysadmin-style jobs, and I don’t remember the answer to a lot of these things. I could definitely get back into it, but if I went through one of these interviews now, I wouldn’t get the job. One would argue I shouldn’t because of my personality, but that’s neither here nor there.

Tomasz: [laugh]. I mean, that’s why you use CDK, so you’d have to remember random YAML comments. And if you [unintelligible 00:23:26] you don’t have YAML anymore. [unintelligible 00:23:27].

Corey: Yes, you’re quite the CDK fanboy, apparently.

Tomasz: I do like CDK, yes. I don’t like, you know, mental overhead, I don’t like context switching, and the way we kind of work at Stedi is everything is written in TypeScript. So, I am a front-end engineer, so I do stuff in the front-end line in TypeScript, all of our Lambda functions are written in TypeScript, and our [unintelligible 00:23:48] is written in TypeScript. So, I can, you know, open up my Visual Studio Code and jump between all of those files, and the language stays the same, the syntax stays the same, the tools stay the same. And I think this is one of the benefits of CDK that is kind of hard to replicate otherwise.

And, you know, people have many opinions about the best to deploy infrastructure in the cloud, you know? The best infrastructure-as-code tool is the one that you use at work or in your private projects, right? Because some people enjoy ClickOps like you do; people—

Corey: Oh yeah.

Tomasz: Enjoy CloudFormation by hand, which I don’t; people are very much into Terraform or Serverless Framework. I’m very much into CDK.

Corey: Or the SAM CLI, like, three or four more, and I use—

Tomasz: Oh, yeah. [unintelligible 00:24:33]—

Corey: —all of these things in various ways in some of my [monstrous 00:24:35] projects to keep up on all these things. I did an exploration with the CDK. Incidentally, I think you just answered why I don’t like it.

Tomasz: Because?

Corey: Because it is very clear that TypeScript is a first-class citizen with the CDK. My language of choice is shitty bash because, grumpy old sysadmin; it happens. And increasingly, that is switching over to terrible Python because I’m very bad at that. And the problem that I run into as I was experimenting with this is, it feels like the Python support is not fully baked, most people who are using the CDK are using a flavor of JavaScript and, let’s be very clear here, the every time I have tried to explore front-end, I have come away more confused than I was when I started, part of me really thinks I should be learning some JavaScript just because of its versatility and utility to a whole bunch of different problems. But it does not work the way I think, on some level, that it should because of my own biases and experiences. So, if you’re not a JavaScript person, I think that you have a much rockier road with the CDK.

Tomasz: I agree. Like I said, I tend to talk about my own experiences and my kind of thoughts about stuff. I’m not going to say that, you know, this tool or that tool is the best tool ever because nothing like that exists. Apart from jQuery, which is the best thing that ever happened to the web since, you know, baked bread, honestly. But you are right about CDK, to the best of my knowledge, kind of, all the other languages that are supported by CDK are effectively transpiled down from TypeScript. So it’s, like, first of all, it is written in TypeScript, and then kind of the Python, all of the other languages… kind of come second.

You know, and afterwards, I tend to enjoy CDK because as I said, I use TypeScript on a daily basis. And you know, with regards to front-end, you mentioned that you are, every single time you is that you end up being more confused. It never goes away. I’ve been doing front-end stuff for years, and it’s, you know, kind of exactly the same. Fun story, I actually joined Cloudash because, well, Maciej started working on Cloudash alone, and after quite some time, he was so frustrated with the modern front-end landscape that he asked me, “Dude, you need to help me. Like, I genuinely need some help. I am tired of React. I am tired of React hooks. This is way too complex. I want to go back to doing back-end stuff. I want to go back, you know, thinking about how we’re going to integrate with all those APIs. I don’t want to do UI stuff anymore.”

Which was kind of like an interesting shift because I remember at the very beginning of my career, where people were talking about front-end—you know, “Front-end is not real programming. Front-end is, you know, it’s easy, it’s simple. I can learn CSS in an hour.” And the amount of people who say that CSS is easy, and are good at CSS is exactly zero. Literally, nobody who’s actually good at CSS says that, you know, CSS, or front-end, or anything like that is easy because it’s not. It’s incredibly complex. It’s getting probably more and more complex because the expectations of our front-end UIs [unintelligible 00:27:44].

Corey: It’s challenging, it is difficult, and one of the things I find most admirable about you is not even your technical achievements, it’s the fact that you’re teaching other people to do this. In fact, this gets to the last point I want to cover on our conversation today. When I was bouncing topic ideas off of you, one of the points you brought up that I’m like, “Oh, we’re keeping that and saving that for the end,” is why—to your words—why speaking at tech events gets easier, but never easy. Let’s dive into that. Tell me more about it.

Tomasz: Basically, I’ve accidentally kickstarted my career by speaking at meetups which later turned into conferences, which later turned into me publishing courses online, which later turned into me becoming an AWS Hero, and here we are, you know, talking to each other. I do enjoy, you know, going out in public and speaking and being on stage. I think, you know, if somebody has, kind of, the heart, the ability to do that, I do strongly recommend, you know, giving it a shot, not only to give, like, an honestly life-changing experience because the first time you go in front of hundreds of people, this is definitely, you know, something that’s going to shake you, while at the same time acknowledging that this is absolutely, definitely not for everyone. But if you are able to do that, I think this is definitely worth your time. But as you said—by quoting me—that it gets easier, so every single time you go on stage, talk at a meetup or at a conference or online conferences—which I’m not exactly a fan of, for the record—it’s—

Corey: It’s too much like work, too much like meetings. There’s nothing different about it.

Tomasz: Yeah, exactly. Like, there’s no journey. There’s no adventure in online conferences. I know that, of course, you know, given all of that, you know, we had to kind of switch to online conferences for quite some time where I think we are pretending that Covid is not a thing anymore, so we, you know, we’re effectively going back, but kind of the point I wanted to make is that I am a somewhat experienced public speaker—I’d like to say that because I’ve been doing that for years—but I’ve been, you know, talking to people who actually get paid to speak at the conferences, to actually kind of do that for a living, and they all say the same thing. It gets simpler, it gets easier, but it’s never freaking easy, you know, to go out there, and you know, to share whatever you’ve learned.

Corey: I’m one of those people. I am a paid public speaker fairly often, even ignoring the podcast side, and I’ve spoken on conference stages a couple hundred times at least. And it does get easier but never easy. That’s a great way of framing it. You… I get nervous before every talk I give.

There are I think two talks I’ve given that I did not have an adrenaline hit and nervous energy before I went onstage, and both of those were duds. Because I think that it’s part of the process, at least for me. And it’s like, “Oh, how do you wind up not being scared for before you go on stage?” You don’t. You really don’t.

But if that appeals to you and you enjoy the adrenaline rush of the rest, do it. If you’re one of those people who’ve used public speaking as, “I would prefer death over that,” people are more scared of public speaking their death, in some cases, great. There are so many ways to build audiences and to reach people that fine, if you don’t like doing it on stage, don’t force yourself to. I’d say try it once; see how it feels meetups are great for this.

Tomasz: Yeah. Meetups are basically the best way to get started. I’m yet to meet a meetup, either, you know, offline or online, who is not looking for speakers. It’s always quite the opposite, you know? I was, you know, co-organizing a meetup in my city here in Poznań, Poland, and the story always goes like this: “Okay, we have a date. We have a venue. Where are the speakers?” And then you know, the tumbleweed is going to roll across the road and, “Oh, crap, we don’t have any speakers.” So, we’re going to try to find some, reach out to people. “Hey, I know that you did this fantastic project at your workplace. Come to us, talk about this.” “No, I don’t want to. You know, I’m not an expert. I am, you know, I have on the 50 years of experience as an engineer. This is not enough.” Like I said, I do strongly recommend it, but as you said, if you’re more scared of public speaking than, like, literally dying, maybe this is not for you.

Corey: Yeah. It comes down to stretching your limits, finding yourself interesting. I find that there are lots of great engineers out there. The ones that I find myself drawn to are the ones who aren’t just great at building something, but at storytelling around the thing that they are built of, yes, you build something awesome, but you have to convince me to care about it. You have to show me the thing that got you excited about this.

And if you can’t inspire that excitement in other people, okay. Are you really excited about it? Or what is the story here? And again, it’s a different skill set. It is not for everyone, but it is absolutely a significant career accelerator if it’s leveraged right.

Tomasz: [crosstalk 00:32:45].

Corey: [crosstalk 00:32:46] on it.

Tomasz: Yeah, absolutely. I think that we don’t talk enough about, kind of, the overlap between engineering and marketing. In the good sense of marketing, not the shady kind of marketing. The kind of marketing that you do for yourself in order to elevate yourself, your projects, your successes to others. Because, you know, try as you might, but if you are kind of like sitting in the corner of an office, you know, just jamming on your keyboard 40 hours per week, you’re not exactly likely to be promoted because nobody’s going to actively reach out to you to find out about your, you know, recent successes and so on.

Which at the same time, I’m not saying that you should go @channel in Slack every single time you push a commit to the main branch, but there’s definitely, you know, a way of being, kind of, kind to yourself by letting others know that, “Okay, I’m here. I do exist, I have, you know, those particular skills that you may be interested about. And I’m able to tell a story which is, you know, convincing.” So it’s, you know, you can tell a story on stage, but you can also tell your story to your customers by building a future that they’re going to use. [unintelligible 00:33:50].

Corey: I really want to thank you for taking the time to speak with me today. If people want to learn more, where’s the best place to find you?

Tomasz: So, the best place to find me is on Twitter. So, my Twitter handle is @tlakomy. So, it’s T-L-A-K-O-M-Y. I’m assuming this is going to be in the [show notes 00:34:06] as well.

Corey: Oh, it absolutely is. You beat me to it.

Tomasz: [laugh]. So, you can find Cloudash at cloudash.dev. You can probably also find my email, but don’t email me because I’m terrible, absolutely terrible at email, so the best way to kind of reach out to me is via my Twitter DMs. I’m slightly less bad at those.

Corey: Excellent. And we will, of course, put links to that in the [show notes 00:34:29]. Thank you so much for being so generous with your time. I appreciate it.

Tomasz: Thank you. Thank you for having me.

Corey: Tomasz Łakomy, Head of React at Cloudash. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, and if you’re on the YouTubes, smash the like and subscribe button, as the kids say. Whereas if you’ve hated this episode, please do the exact same thing—five-star reviews smash the buttons—but this time also leave an insulting and angry comment written in the form of a CloudWatch log entry that no one is ever able to find in the native interface.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Michael

Michael is the creator of IT automation platforms Cobbler and Ansible, the latter allegedly used by ~60% of the Fortune 500, and at one time one of the top 10 contributed to projects on GitHub.

Links Referenced:

  • Speaking Tech: https://michaeldehaan.substack.com/
  • michaeldehaan.net: https://michaeldehaan.net
  • Twitter: https://twitter.com/laserllama

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored by our friends at Revelo. Revelo is the Spanish word of the day, and its spelled R-E-V-E-L-O. It means “I reveal.” Now, have you tried to hire an engineer lately? I assure you it is significantly harder than it sounds. One of the things that Revelo has recognized is something I’ve been talking about for a while, specifically that while talent is evenly distributed, opportunity is absolutely not. They’re exposing a new talent pool to, basically, those of us without a presence in Latin America via their platform. It’s the largest tech talent marketplace in Latin America with over a million engineers in their network, which includes—but isn’t limited to—talent in Mexico, Costa Rica, Brazil, and Argentina. Now, not only do they wind up spreading all of their talent on English ability, as well as you know, their engineering skills, but they go significantly beyond that. Some of the folks on their platform are hands down the most talented engineers that I’ve ever spoken to. Let’s also not forget that Latin America has high time zone overlap with what we have here in the United States, so you can hire full-time remote engineers who share most of the workday as your team. It’s an end-to-end talent service, so you can find and hire engineers in Central and South America without having to worry about, frankly, the colossal pain of cross-border payroll and benefits and compliance because Revelo handles all of it. If you’re hiring engineers, check out revelo.io/screaming to get 20% off your first three months. That’s R-E-V-E-L-O dot I-O slash screaming.

Corey: This episode is sponsored in part by LaunchDarkly. Take a look at what it takes to get your code into production. I’m going to just guess that it’s awful because it’s always awful. No one loves their deployment process. What if launching new features didn’t require you to do a full-on code and possibly infrastructure deploy? What if you could test on a small subset of users and then roll it back immediately if results aren’t what you expect? LaunchDarkly does exactly this. To learn more, visit launchdarkly.com and tell them Corey sent you, and watch for the wince.

Corey: Once upon a time, Docker came out and change an entire industry forever. But believe it or not, for many of you, this predates your involvement in the space. There was a time where we had to manage computer systems ourselves with our hands—kind of—like in the prehistoric days, chiseling bits onto disk and whatnot. It was an area crying out for automation, as we started using more and more computers to run various websites. “Oh, that’s a big website. It needs three servers now.” Et cetera.

The times have changed rather significantly. One of the formative voices in that era was Michael DeHaan, who’s joining me today, originally one of the—or if not the creator of Cobbler, and later—for which you became better known—Ansible. First, thanks for joining me.

Michael: Thank you for having me. You’re also making me feel very, very old there. So, uh, yes.

Corey: I hear you. I keep telling people, I’m in my mid-30s, and my wife gets incensed because I’m turning 40 in July. But still. I go for the idea of yeah, the middle is expanding all the time, but it’s always disturbing talking to people who are in our sector, who are younger than some of the code that we’re using, which is just bizarre to me. We’re all standing on the backs of giants. Like it or not, one of them’s you.

Michael: Oh, well, thank you. Thank you very much. Yeah, I was, like, talking to some undergrads, I was doing a little bit of stuff helping out my alma mater for a little bit, and teaching somebody the REST lecture. I was like, “In another year, REST is going to be older than everybody in the room.” And then I was just kind of… scared.

Corey: Yeah. It’s been a wild ride for basically everyone who’s been around long enough if you don’t fall off the teeter-totter and wind up breaking a limb somewhere. So, back in the bad old days, before cloud, when everything was no longer things back then were constrained by how much room you had on your credit card like they are today with cloud, but instead by things like how much space you had in the data center, what kind of purchase order you could ram through your various accounting departments. And one of the big problems you have is, great. So, finally—never on time—Dell has shipped out a whole bunch of servers—or HP or Supermicro or whoever—and the remote hands—which is always distinct from smart hands, which says something very insulting, but they seem to be good about it—would put them into racks for you.

And great, so you’d walk in and see all of these brand new servers with nothing on them. How do we go ahead and configure these things? And by hand was how most of us started, and that means, oh, great, we’re going to screw things up and not do them all quite the same, and it’s just a treasure and a joy. Cobbler was something that you came up with that revolutionized how provisioning of bare-metal systems worked. Tell me about it.

Michael: Yeah, um, so it’s basically just glue. So, the story of how I came up with that is I was working for the Emerging Technologies Group at Red Hat, and I just joined. And they were like, “We have to have a solution to install Xen and KVM virtual machines.” So obviously, everybody’s familiar with, like, EC2 and things now, but this was about people running non-VMware virtualization themselves. So, that was part of the problem, but in order to make that interesting, we really needed to have some automation around bare-metal installs.

And that’s PXE boot. So, it’s TFTP and DHCP protocol and all that kind of boring stuff. And there was glue that existed, but it was usually humans would have to click on buttons to—like Red Hat had system-config-netboot, but what really happened was sysadmins all wrote their own automation at, like, every single company. And the idea that I had, and it was sort of cemented by the fact that, like, my boss, a really good guy left for another company and I didn’t have a boss for, like, a couple years, was like, I’m just going to make IRC my boss, and let’s get all these admins together and build a tool we can share, right?

So, that was a really good experience, and it’s just basically gluing all that stuff together to fully automate an install over a network so that when a system comes on, you can either pick it out from a menu; or maybe you’ve already got the MAC address and you can just say, “When you see this MAC address, go install this operating system.” And there’s a kickstart file, or a preseed in the case of Debian, that says, “When you’re booting up through the installer, basically, here’s just the answers and go do these things.” And that install processes a lot slower than what we’re used to, but for a bare-metal machine, that’s a pretty good way to do it.

Corey: Yeah, it got to a point where you could walk through and just turn on all the servers in a rack and go out to lunch, come back, they would all be configured and ready to go. And it sounds relatively basic the way we’re talking about it now, but there were some gnarly cases. Like, “When I’ve rebooted the database server, why did it wipe itself and reprovision?” And it’s, “Oh, dear.” And you have to make sure that things are—that there’s a safety built into these things.

And you also don’t want to have to wind up plugging in a keyboard and monitor to all of these individual machines one-by-one to hit yes and acknowledge the thing. And it was a colossal pain in the ass. That’s one of the things that cloud has freed us from.

Michael: Yeah, definitely. And one of the nice things about the whole cloud environment is like, if you want to experiment with those ideas, like, I want to set up some DHCP or DNS, I don’t have to have this massive lab and all the electricity and costs. But like, if I want to play with a load balancer, I can just get one. That kind of gives the experience of playing with all these data center technologies to everybody, which is pretty cool.

Corey: On some level, you can almost view the history of all these things as speeding things up. With a well-tuned Cobbler install, it still took multiple minutes, in some cases, tens of minutes to go from machine you’re powering on to getting it provisioned and ready to go. Virtual machines dropped that down to minutes. And cloud, of course, accelerated that a bit. But then you wind up with things like Docker and it gets down to less than a second. It’s the meantime to dopamine.

But in between the world of containers and bare-metal, there was another project—again, the one you’re best known for—Ansible. Tell me about that because I have opinions on this whole space.

Michael: [laugh]. Yeah. So, how Ansible got started—well, I guess configuration management is pretty old, so the people writing their own scripts, CFEngine came out, Puppet was a much better CFEngine. I was working at a company and I kind of wanted another open-source project because I enjoyed the Cobbler experience. So, I started Ansible on the side, kind of based on some frustrations around Puppet but also the desire to unify Capistrano kind of logic, which was like, “How do I push out my apps onto these servers that are already running,” with Puppet-style logic was like, “Is this computer’s firewall configured directly? And is the time set correctly?”

And you can obviously use that to install apps, but there’s some places where that blurred together where a lot of people are using two different tools. And there’s some prior art that I worked on called Funk, which I wrote with Seth Vidal and Adrian Likins at Red Hat, which was, like, 50% of the Ansible idea, and we just never built the config management layer on top. So, the idea was make something really, really simple that just uses SSH, which was controversial at the time because people thought it, like, wouldn’t scale, because I was having trouble with setting up Puppet security because, like, it had DNS or timing issues or whatever.

Corey: Yeah. Let’s dive in a bit to what config management is first because it turns out that not everyone was living in the trenches in quite the same way that we were. I was a traveling trainer for Puppet for a summer once, and the best descriptor I found to explain it to people who are not in this space was, “All right, let’s say that you go and you buy a new computer. What do you do? Well, you’re going to install the applications you’d like to use, you’re going to set up your own user account, you’re going to set your password correctly, you’re going to set up preferences, copy some files over so you have the stuff you care about. Great. Now, imagine you need to do that to a thousand computers and they all need to be the same. How do you do that?” Well, that is the world of configuration management.

And there was sort of a bifurcation there, where there was the idea of, first, we’re going to have configuration management that just describes what the system should look like, and that’s going to run on a schedule or whatnot, and then you’re going to have the other side of it, which is the idea of remote execution, of I want to run an arbitrary command on this server, or this set of servers, or all the servers, depending upon what it is. And depending on where you started on the side of that world, you wound up wanting things from the other side of that space. With Puppet, for example, is very oriented configuration management and the question became, well, can you use this for remote execution with arbitrary commands? And they wound up doing some work with Mcollective, which was a very complicated and expensive way to say, “No, not really.” There was a need for things that needed to hang out in that space.

The two that really stuck out from that era were Ansible, which had its wild runaway success, and the one that I was smacking around for a bit, SaltStack, which never saw anywhere approaching that level of popularity.

Michael: Yeah, sure. I mean, I think that you hit it pretty much exactly right. And it’s hard to say what makes certain things take off, but I think, like, the just SSH approach was interesting because, well for one, everybody’s running it. But there was this belief that this would not scale. And I tried to optimize the heck out of that because I liked performance, but it turns out that wasn’t really a business problem because if you can imagine you just wrote this little bit of automation, and you’re going to run it against your entire infrastructure and you’ve got 30,000 machines, do you want that to—if you were to, like, run an update command on 30,000 machines at once, you’re going to DDoS something. Definitely, right?

Corey: Yeah. Suddenly you have 30,000 machines all talk to the same things at the same times. And you want to do them in batches or smear it across.

Michael: Right, so because that was there, like, you just add batch support in Ansible and things are fine, right? People want to target little small groups of things. So, like, that whole story wasn’t true, and I think it was just a matter of testing this belief that everybody thought that we needed to have this whole network of things. And honestly, Salt’s idea of using a message bus is great, but we took a little bit different approach with YAML because we have YAML variables in it, but they had something that compiled down to YAML. And I think those are some differences in the dialect and some things other people preferred, but—

Corey: And they use Jinja, at one point to wind up making it effectively Turing complete; you could wind up having this ridicu—like, loop flow control and loops and the rest. And it was an interesting exposure to things, but yikes, at some l—at the same time.

Michael: If you use all the language features in anything you can make something complicated, and too complicated. And I was like, I wanted automation to look like grocery lists. And when I started out, I said, “Hey, if anybody is doing this all day, for a day job, I will have failed.” And it clearly shows you that I have because there are people that are doing that all day. And the goal was, let me concentrate on dev and ops and my other things and keep this really, really simple.

And some people just, like, get really, really into that automation technology, which is—in my opinion—why some of the earlier stuff was really popular because sysadmin were bored, so they see something new and it’s kind of like a Java developer finding Perl for the first time. They’re like, “I’m going to use all these things.” And they have all their little widgets, and it gets, like, really complicated.

Corey: The thing that I always found interesting and terrifying at the same time about Ansible was the fact that you did ride on top of SSH, which is great because every company already had a way of controlling access by SSH to IT systems; everyone uses it, so it has an awful lot of eyes on the security protocol on the rest. The thing that I found terrifying in the early days was that more or less every ops person would wind up checking this out onto their laptop or whatnot, so whenever they wanted to run something, they would just run it from their laptop over a VPN or whatnot from wherever they happen to be, and you wind up with a dueling banjos type of circumstance where people were often not doing it from a centralized place. And in time, best practices emerged where, okay, that is going to be the command and control server where that runs at, and you log into it. And then you start guarding that with CI/CD flows and the rest. And like anything else, it wound up building some operational approaches to it.

Michael: Yeah. Like, I kind of think that created a problem that allowed us to sell a product, right, which was good. If you knew what you were doing, you could use Jenkins completely and you’d be fine, right, if you had some level of discipline and access control, and you wanted to wire that up. And if you think about cloud, this whole, like, shadow IT idea of, “I just want to do this thing, therefore I’m going to get an Amazon account,” it’s kind of the same thing. It’s like, “I want to use this config management, but it’s not approved. Who can stop me?” Right?

And that kind of probably got us in the door in few accounts that way. But yeah, it did definitely create the problem where multiple people could be running things at the same time. So yeah, I mean, that’s true.

Corey: And the idea of, “Hey, maybe I should be controlling these things in Git,” or some other form of version control was sort of one of those evolutionary ideas that, oh, we could treat this like code. And the early days of DevOps, that was a controversial thing. These days, you say you’re not doing it and people look at you very strangely. And things were going reasonably well in that direction for a while. Then this whole Docker thing showed up, where, well, what if instead of having these long-lived servers where you have to install updates and run patches and maintain a whole user list on them, instead you had this immutable infrastructure that every time there was a change, you would just go ahead and deploy a brand new set of servers?

And you could do this in the olden days with virtual machines and whatnot; it just took a long time to push things out, so do I really want to roll the entire fleet for a two-line config change? Probably not, so we’re going to batch it up, or maybe do this hybrid model. With Docker, it takes less than a second to wind up provisioning the—switching over to the new container series and you’re done; you can keep going with that. That really solved a lot of these problems.

But there were companies that, like, the entire configuration management space, who suddenly found themselves in a really weird position. Some of them tried to fight the tide forever and say, “Oh, this is terrible because it means we don’t have a business model anymore.” But you can only fight the future for so long. And I think today, we’d be hard-pressed to say that Docker hasn’t won, on some level.

Michael: I mean, I think it has, like, the technology has won. But I guess the interesting thing is, config management now seems to be trying to pivot towards networking where I think the tool hasn’t ever been designed for networking, so it’s kind of a round peg, square hole. But it’s all people have that unless they’re buying something. Or, like, deploying the undercloud because, like, people are still running essentially clouds on top of clouds to get their Kubernetes deployments going and those are monstrous. Or maybe to deploy a data layer; like, I know Kafka has gotten off of ZooKeeper, but the Kafka-ZooKeeper thing—and I don’t remember ZooKeeper [unintelligible 00:14:37] require [unintelligible 00:14:38] or not, but managing those sort of long, persistent implications, it still has a little bit of a place where it exists.

But I mean, I think the whole immutable systems idea is theoretically completely great. I never was really happy with the whole Docker development workflow, and I think it does create a problem where people don’t know what they’re deploying and you kind of encourage that to where they could be deploying different versions of libraries, or—and that’s kind of just a problem of the whole microservices thing in general where, “Did somebody change this?” And then I was working very briefly at one company where we essentially built a whole dashboard to detect service versions and what version of the base image everybody was on, and all these other things, and it can get out of hand, too. So, it’s kind of like trading some problems for other problems, I think to me. But in general, containerization is good. I just wished the management glue around it was easy, right?

Corey: I wound up giving a talk at a conference a while back, 2015 or so, called, “Heresy in the Church of Docker,” and it was a throwaway five-minute lightning talk, and someone approached me afterwards with, “Hey, can you give the full version of that at ContainerCon?” “There’s a full version? Yes. Yes, I can.” And it talked about a number of problems with the management layer and the rest.

Now, Kubernetes absolutely solves virtually every problem that I identified with it, but when you look at the other side of it, getting Kubernetes rolled out is effectively you get to cosplay being a cloud provider yourself. It is incredibly complicated, and of course, we’re right back to managing it all with YAML.

Michael: Right. And I think that’s an interesting point, too, is I don’t know who’s exactly responsible for, like, the YAML explosion. And I like it as a data format; it’s really good for humans. Cobbler originally used it more of an internal storage, which I think was a mistake because, like, even—I was trying to avoid setting up a database at the time, so—because I knew if I had to require setting up a database in 2007 or 2008, I’d get way less users, so it used flat files.

A lot of the YAML dialects people are developing now are very, very nested and they requires, like, loading a webpage, for the Docks, like, all the time and reading what’s valid here, what’s valid there. I think people learn the wrong lesson from Ansible’s YAML usage, right? It was supposed to be, like, YAML’s good for things that are grocery lists. And there’s a lot of places where I didn’t do a good job. But when you see methods taking 15 parameters and you have to constantly have the reference up, maybe that’s a sign that you should do something else.

Corey: At least you saved us, on some level, from having to do this all in XML. But still, there are wrong ways and more wrong ways to do it. I don’t think anyone could ever agree on the right way to approach these things.

Michael: Yeah. I mean, and YAML, at the time was a good answer because I knew I didn’t want to write and maintain a parser as, like, a guy that was running a project. We had a lot of awesome contributors, but if I had to also maintain a DSL, not only does that mean that I have to write the code for this thing—which I, you know, observed slowing down some other projects—but also that I’d have to explain it to people. Looking kind of like Bash was not a bad thing. Not having to know and learn something, so you can kind of feel really effective in about 15 minutes or something like that.

Corey: One of the things that I find really interesting about you personally is that you were starting off in a bare-metal world; Ansible was sort of wherever you wanted to run it. Great, as long as there are systems that can receive these things, we’re great. And now the world has changed, and for better or worse, configuration management slash remote execution is not the problem it once was and it is not a best practice way of solving a lot of those problems either. But you aren’t spending your time basically remembering the glory years. You’re actively moving forward doing some fairly interesting stuff. How would you describe what you’re into these days?

Michael: I tried to create a few projects to, like, kind of build other systems management things for the same audience for a while. I was building a build server and a new—trying to do some next-gen config stuff. And I saw people weren’t interested. But I like having conversations with people, right, and I think one of the lessons from Ansible was how to explain highly technical things to technical audiences and cut out a lot of the marketing goo and all that; how to get people excited about an idea and make a community be really authentic. So, I’ve been writing about that for really, it’s—rebooted blog is only a couple of weeks old. But also kind of trying to do some—helping out companies with some, like, basic marketing kind of stuff, right?

There’s just this pattern that everybody has where every website starts with this little basic slogan and two buttons and then there’s a bunch of adjectives, but it doesn’t say anything. So, how can you have really good documentation, and how can you explain an idea? Because, like, really, the reason you’re in it is not just to sell stuff, but it’s to help people and to see them get excited about your ideas. And there’s just, like, we’re not doing a good job in this, like, world where there’s thousands upon thousands of applications, all competing at once to, like—how do you rise above that?

Corey: And that’s always the hard part is at some point, this does become your identity and you become known for a thing. And when you start branching out from that thing, you bring the expertise from that area that you were in, but you start applying it to new things. I feel like so many companies get focused—and people get focused—on assuming that their audience is just like them, where they’re coming in with the exact same biases, the exact same experiences. And given that basically no one was as deep in the weeds as you were when it came to configuration management, that meant that you were spending time in that side of the world, not in other pursuits which aligned in some ways more directly with people developing other things. So, I suspect this might be one of the weird things we have in common when we show up and see something new.

And a company is really excited. It’s like, it’s basically a few people talking [unintelligible 00:20:12] that both founders are technical. And they’re super excited about something they can’t quite articulate. And it’s this, “Slow down. Tell me exactly what it is your product does.” And that’s a hard thing to do because my default response was always the if I don’t understand that is clearly the way in which I am deficient somehow. But marketing is really about clear communication and there’s not that much of it in our space, at least not for early-stage companies.

Michael: Yeah, I don’t know why that is. I mean, I think there’s this belief in that there’s, like, this buyer audience where there’s some vice president that’s going to buy your stuff if you drop the right buzzwords. And 15 years ago, like, you had to say ‘synergy,’ and now you say ‘time to value’ or ‘total cost of ownership’ or something. And I don’t think that’s true. I mean, I think people use products that they like and that they need to be shown them to try them out.

So like, why can’t your webpage have a diagram and a screenshot instead of this, like, picture of a couple of people drinking coffee around a computer, right? It’s basic stuff. But I agree with you, I kind of feel dumb when I’m looking at all these tech products that I should be excited about, and, like, the way that we get there, as we ask questions. And the way that I’ve actually figured out what some of these things do is usually having to ask questions from someone who uses them that I randomly find on my diminishing circle of friends, right? And that’s kind of busted.

So, Ansible definitely had a lot of privilege in the way that it was launched in the sense that I launched it off Cobbler list and Cobbler list started off of [ET Management Tools 00:21:34] which was a company list. But people can do things like meetup groups really easily, they can give talks, they can get their blogs reblogged, and, you know, hope for some Hacker News or Reddit juice or whatever. But in order to get that to happen, you have to be able to talk to engineers that really want to know what you’re doing, and they should be excited about it. So, learn to talk to them.

Corey: You have to speak their language but without going so deep in the weeds that the only people that understand it are the folks who are never going to use your product because they want to build it themselves. It’s a delicate balance to strike.

Michael: And it’s a difficult thing to do, too, when you know about it. So, when I was, like, developing all the Ansible docs, I’ve told people many times—and I hope it’s true—that I, like, spent, like, 40% of my time just on the website and the docs, and whenever I heard somebody complain, I tried to fix it. But the idea was like, you can lose somebody really fast, but you kind of have to forget what you know about the product. So, the worst person to sometimes look at that as the person that built it. So, you have to forget what you know, and try to see, like, what questions they’re asking, what do they need to find out? How do they want to learn something?

And for me, I want to see a lot of pictures. A lot of people write a bunch of giant walls of text, or worse for me is when there’s just these little pithy expressions and I don’t know what they mean, right? And everybody’s, like, kind of doing that these days.

Corey: This episode is sponsored in part by our friends at ChaosSearch. You could run Elasticsearch or Elastic Cloud—or OpenSearch as they’re calling it now—or a self-hosted ELK stack. But why? ChaosSearch gives you the same API you’ve come to know and tolerate, along with unlimited data retention and no data movement. Just throw your data into S3 and proceed from there as you would expect. This is great for IT operations folks, for app performance monitoring, cybersecurity. If you’re using Elasticsearch, consider not running Elasticsearch. They’re also available now in the AWS marketplace if you’d prefer not to go direct and have half of whatever you pay them count towards your EDB commitment. Discover what companies like Equifax, Armor Security, and Blackboard already have. To learn more, visit chaossearch.io and tell them I sent you just so you can see them facepalm, yet again.48]

Corey: One thing that I’ve really found myself enjoying recently has been your substack-based newsletter, Speaking Techis what you call it. And I didn’t quite know what to expect when I signed up for it, but it’s been a few weeks now, and you are more or less hitting across the board on a bunch of different things, ranging from engineering design patterns, to a teardown of random company’s entire website from a marketing and messaging perspective—which I just adore personally; like that is very aligned with how I see the world—

Michael: There’s more of that coming.

Corey: Yeah, [unintelligible 00:23:17] a bunch of other stuff. Let’s talk about, for example, the idea of those teardowns. I always found that I have to be somewhat careful in how I talk about it when I’m doing a tweet thread or something like that because you are talking about people’s work, let’s be clear here, and I tend to be a lot kinder to small, early-stage companies than I am to, you know, $1.6 trillion companies who really should have solved for this by now, on some level. But so much of it misses the mark of great, here’s the way that I think about these things. Here’s the way that I don’t understand what the hell you’re telling me.

An easy example of this for me, at least I’m curious to get your thoughts on it, I tend to almost always just skim what they’re saying, great. Let’s look at the pricing page because I find that speaks to people in a way that very often companies forget that they’re speaking to customers.

Michael: Yeah, for sure. I always tried to find the product page lately, and then, like, the product page now is, like, a regurgitation of the homepage. But it’s what you said earlier. I think I try to stay nice to everybody, but it’s good to show people how to understand things by counterexample, to some extent, right? Like, oh, I’ve got some stuff coming out—I don’t know when this is actually going to get published—but next week, where I was like just taking random snippets of home pages, and like, “What’s everybody doing with the header these days?”

And there’s just, like, ridiculous amounts of copying going on. But it’s not just for, like, people’s companies because everybody listening here isn’t going to have a company. If you have a project and you wanted to get it noticed, right, I think, like, in the early days, the projects that I paid attention to and got excited about were often the ones that spend time on their website and their messaging and their experience. So, everybody kind of understands you have to write a good readme now but some of, like, the early Ruby crowd, for instance, did awesome, awesome web pages. They know how to pick out fonts, and I still don’t know how to pick out fonts. But—

Corey: I ask someone good at those things. That’s how I pick ‘em.

Michael: Yeah, yeah. That’s not my job; get somebody that’s good at that. But all that matters, right? So, if you do invest a little bit in not promoting yourself, not promoting your company, but trying to help people and communicate to them, you can build that audience around your thing and it makes it a lot more interesting.

Corey: There’s so many great tools out there that I find on GitHub that other people have to either point me to or I find it when I’m looking at it from a code-first perspective, just trying to find a particular example of the library being used, where they do such a terrible job of describing the problem that they solve, and it doesn’t take much; it takes a paragraph or two at most. Or the idea that, “Oh, yeah, here’s a way to do this thing. Just go ahead and get your credential file somewhere else.” Great. Could you maybe link to an example of how to do that?

It’s the basic stuff; assume that someone who isn’t you might possibly want to use this. And I’m not even slightly suggesting that you wind up talking your way through how to do all of that. Just link to somewhere that has a good write-up of it and call it good. Just don’t get in the way of people’s first-time user experiences.

Michael: Yeah, for sure. And—

Corey: For some reason, that’s a radical thought.

Michael: Yeah, I think one of the things the industry has—well, not the industry; it’s not their problem to solve, but, like, we don’t really have a way for people to find what’s cool and interesting anymore. So, various people have their own little lists on GitHub or whatever, but there’s just so many people posting on the one or two forums people read and it goes by in a day. So, it’s really, really hard to get attention. Even your own circle of followers isn’t really logging into Twitter or anything, or LinkedIn. Or there’s all the congratulations for your five years of Acme Corp kind of posts, and it’s really, really hard to get attention.

And I feel for everybody, so like, if somebody like GitHub or Microsoft is listening, and you wanted to build, like, a dashboard of here’s the cool 15 projects for the week, kind of thing where everybody would see it, and start spotlight some of these really cool new things, that would be awesome, right?

Corey: Whenever you see those roundups, that was things like Kubernetes and Docker. And great, I don’t think those projects need the help in the same way.

Michael: No, no, they don’t. It’s like maybe somebody’s cool data thing, or a cool visualization, or the other thing that’s—it’s completely random, but I used to write fun graphics programs for fun or games and libraries. And I don’t see that anymore, right? Maybe if you find it, you can look for it, but the things that get people excited about programming. Maybe they have no commercial value at all, but the way that people discover stuff is getting so consolidated is about Docker and Kubernetes. And everyone’s talking about these three things, and if you’re not Google or you’re not Facebook, it’s really—or Amazon, obviously—it’s hard to get attention.

Corey: Open-source on some level has changed from a community perspective. And part of it is because once upon a time, you could start with the very low-level stuff and build something, get it up and working. And that’s where things like [Cobbler 00:27:44] and Ansible came out of. Now it’s, “Click the button and use the thing everyone else is using. And if you’re not doing that, what are you doing over there?”

So, the idea of getting started tinkering with computers are built on top of so many frameworks and other things. And that’s always been the case, but now it’s much more apparent in some ways. “Okay, I’m going to go ahead and build out my first HTML file and serve it out using something in Node.” “Great, what is those NPM stuff that’s scrolling past?” It’s like, “The devil. That is the devil’s own language you are seeing scroll past. And you don’t need to worry about that; just pretend it’s not there.”

But back when I was learning all this stuff, we’re paying attention to things scrolling past, like, you know, compilation messages and the Linux boot story as it wound up scrolling past. Terrible story; the protagonist was unreliable, but all right. And you start learning how these things work when you start scratching at the things that you’re just sort of hand-waving and glossing over. These days, it feels like every time I use a modern project, that’s everything.

Michael: I mean, it is. And like what, React has, like, 2000 dependencies, right? So, how do you ever feel like you understand it? Or when recruiters are asking for ten years at Amazon. And then—or we find a library that it can only explain itself by being like this other library and requiring these other five.

And you read one of those, and it becomes, like, this… tree of knowledge that you have no way of possibly understanding. So, we’ve just built these stacks upon stacks upon stacks of things. And I tend to think I kind of believe in minimalism. And like, wouldn’t it be cool if we just burned this all and start—you know, we burn the forest and let something new regrow. But we tend to not do that. We just—now running a cloud on top of a cloud, and our JavaScript is thousands of miles high.

Corey: I really wish that there were better paths for getting started. Like, I used to think that the right way to wind up learning how all this stuff work is to do what I did: Start off as, you know, the grumpy sysadmin type, and then—or help desk—and then work your way up and the rest. Those jobs aren’t there anymore, and it doesn’t leave people in a productive environment. “Oh, you want to build a computer game. Great. For an iPhone? Terrific.” Where do you go to get started on that? It’s a hard thing to do.

And people don’t care at that scale, nor should they necessarily, on how to run your own servers. Back in the day when you wanted to have a blog on the internet, you were either reduced to using LiveJournal or MySpace, or you were running your own web server and had to learn how to make sure that it didn’t become an attack platform. There was a learning curve that was fairly steep. Now, there are so many different paths to go down, you don’t really need to know how any of these things even work.

Michael: Yeah, I think, like, one of the—I don’t know whether DevOps means anything as a topic or not, but one of the original pieces around that movement was systems administrators learning to code things and really starting to enjoy it, whether that was Python or Ruby, and so on. And now it feels like we’re gluing all the things together, but that’s happening in App Dev as well, right? The number of people that can build a really, really good library from the ground up, like, something that has C bindings, that’s a really, really small crowd. And most of it, what we’re doing is gluing together other people’s libraries and compensating for the flaws and bugs in them, and duct tape and error handling or whatever. And it feels like programming has changed a lot because of this—and it’s good if you want to get an idea up quickly, no doubt. But it’s a different experience.

Corey: The problem I always ran into was the similar problems I had with doing Debian packaging. It was always the, oh, great, there’s going to be four or five different guides on how to do it—same story with RPM—and they’re all going to be assuming different things, and you can crossover between them without realizing it. And then you just do something monstrous that kind of works until an actual Debian developer shoves you aside like you were a hazard to everyone around you. Let me do it for you. And there we go.

It’s basically, get people to do work for you by being really bad at it. And I don’t love that pattern, but I’m still reminded of that because there are so many different ways to achieve any outcome that, okay, I want to run a ridiculous Hotdog or Not Hotdog style website out there. Great. I can upload things. Well, Docker or serverless? What provider do I want to put it on? And oh, by the way, a lot of those decisions very early on are one-way doors that you don’t realize you’re crossing through, as well as not knowing what the nuances of all of those things are. And that’s dangerous.

Michael: I think people are also learning the vendor as well, right? Some people get really engrossed in whether it’s Amazon, or Google, or HashiCorp, or somebody’s API, and you spend so much of your brain cells just learning how these people’s systems work versus, like, general programming practices or whatever.

Corey: I make it a point to build something on other cloud providers that aren’t Amazon every now and then, just because I don’t want to wind up effectively embracing a monoculture.

Michael: Yeah, for sure. I mean, I think that’s kind of the trend I see with people looking just at the Kubernetes stuff, or whatever, it’s that I don’t think it necessarily existed in web dev; there seems to be a lot of—still a lot of creativity and different frameworks there, but people are kind of… what’s popular? What gets me my next job, and that kind of thing. Whereas before it was… I wasn’t necessarily a sysadmin; I kind of stumbled into building admin tools. I kind of made hammers not houses or whatever, but basically, everybody was kind of building their own tools and deciding what they wanted. Now, like, people that are wanting to make money or deciding what people want for them. And it’s kind of not always the simplest, easiest thing.

Corey: So, many open-source projects now are—for example, one that I was dealing with recently was the AWS CLI. Great, like, I’m thrilled to throw in issues and challenges here, but I’m not going to spend significant time writing code against it because, one, it’s basically impossible to get these things accepted when all the maintainers work at Amazon, and two, is it really an open-source project in the way that you and I think about community and the rest, but it’s basically sole purpose is to funnel money to Amazon faster. Like, that isn’t really a community ethos I feel comfortable getting behind to be perfectly honest. They’re a big company; they can afford to pay people to build these things out, full time.

Michael: Yeah. And GitHub, I mean, we all mostly, I think, appreciate the fact that we can host the Git repo and it’s performant and everything, and we don’t have blazing unicorns quite as often or whatever they used to have, but it kind of changed the whole open-source culture because we used to talk about things on mailing lists, like, what should this be, and there was like an upfront conversation, or it might happen on IRC. And now people are used to just saying, “I’ve got a problem. Fix it for me.” Or they’re throwing code over the wall and it might not be the code or feature that you wanted because they’re not really part of your thing.

So before, people would get really engrossed with, like, just a couple of projects, and if they were working on them as kind of like a collective of people working against different organizations, we’d talk about things, and they kind of know what was going on. And now it’s very easy to get a patch that you don’t want and you’re, like, “Oh, can you change all of these things?” And then somebody’s, like, now they’re offended because now they have to do all this extra work, whereas that conversation didn’t happen. And GitHub could absolutely remodel themselves to encourage those kinds of conversations and communities, but part of the death of open-source and the fact that now it’s, “Give me free code,” is because of that kind of absence of the—because we’re looking at that is, like, the front of a community versus, like, a conversation.

Corey: I really want to appreciate your taking so much time out of your day to basically reminisce about some of these things. But on a forward-looking basis, if people want to learn more about how you see things, where’s the best place to find you?

Michael: Yeah. So, if you’re interested in my blog, it’s pretty random, but it’s michaeldehaan.substack.com. I run a small emerging
consultancy thing off of michaeldehaan.net. And that’s basically it. My Twitter is laserllama if you want to follow that. Yeah, thank you very much for having me. Great conversation. Definitely making all this technology feel old and busted, but maybe there’s still some merit in going back—

Corey: Old and busted because it wasn’t built this year? Great—

Michael: Yes.

Corey: —yes, its legacy, which is a condescending engineering term for ‘it makes money.’ Yeah, there’s an entire universe of stuff out there. There are companies that are still toying with virtualization: “Is this something we get on board with?” There’s nothing inherently wrong with that. I find that judging what a bunch of startups are doing or ‘company started today’ is a poor frame of reference to look at what you should do with your 200-year-old insurance company.

Michael: Yeah, like, [unintelligible 00:35:53] software engineering is just ridiculously new. Like, if you compare it to, like, bridge-building, or even electrical engineering, right? The industry doesn’t know what it’s doing and it’s kind of stumbling around trying to escape local maxima and things like that.

Corey: I will, of course, put links to where to find you into the [show notes 00:36:09]. Thanks again for being so generous with your time. It’s appreciated.

Michael: Yeah, thank you very much.

Corey: Michael DeHaan, founder of Cobbler, Ansible, and oh, so much more than that. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice—and/or smash the like and subscribe buttons on the YouTubes—whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, smash the buttons as mentioned, and leave a loud, angry comment explaining what you hated about it that I will then summarily reject because it wasn’t properly formatted YAML.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Avery

wvdial, bup, sshuttle, netselect, popularity-contest, redo, gfblip, GFiber, and now @Tailscale doing WireGuard mesh. Top search result for "epic treatise."

Links Referenced:

  • Webpage: https://tailscale.com
  • Tailscale Twitter: https://twitter.com/tailscale
  • Personal Twitter: https://twitter.com/apenwarr

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by LaunchDarkly. Take a look at what it takes to get your code into production. I’m going to just guess that it’s awful because it’s always awful. No one loves their deployment process. What if launching new features didn’t require you to do a full-on code and possibly infrastructure deploy? What if you could test on a small subset of users and then roll it back immediately if results aren’t what you expect? LaunchDarkly does exactly this. To learn more, visit launchdarkly.com and tell them Corey sent you, and watch for the wince.

Corey: This episode is sponsored by our friends at Revelo. Revelo is the Spanish word of the day, and its spelled R-E-V-E-L-O. It means “I reveal.” Now, have you tried to hire an engineer lately? I assure you it is significantly harder than it sounds. One of the things that Revelo has recognized is something I’ve been talking about for a while, specifically that while talent is evenly distributed, opportunity is absolutely not. They’re exposing a new talent pool to, basically, those of us without a presence in Latin America via their platform. It’s the largest tech talent marketplace in Latin America with over a million engineers in their network, which includes—but isn’t limited to—talent in Mexico, Costa Rica, Brazil, and Argentina. Now, not only do they wind up spreading all of their talent on English ability, as well as you know, their engineering skills, but they go significantly beyond that. Some of the folks on their platform are hands down the most talented engineers that I’ve ever spoken to. Let’s also not forget that Latin America has high time zone overlap with what we have here in the United States, so you can hire full-time remote engineers who share most of the workday as your team. It’s an end-to-end talent service, so you can find and hire engineers in Central and South America without having to worry about, frankly, the colossal pain of cross-border payroll and benefits and compliance because Revelo handles all of it. If you’re hiring engineers, check out revelo.io/screaming to get 20% off your first three months. That’s R-E-V-E-L-O dot I-O slash screaming.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Generally, at the start of these shows, I mention something about money. When I have a promoted guest, which means that they are sponsoring this episode, I talk about that. This is not that moment. There’s no money changing hands here.

And in fact, I’m about to talk about a product that I am a huge fan of, but I’m, also as of this recording, not paying for. So, one might think I’m the product, but no. Let’s actually start by talking about money. My guest today is Avery Pennarun, the CEO of Tailscale, and as of today, being the day that this goes out, you folks have just raised $100 million in a Series B. First, thank you for joining me, followed immediately by congratulations.

Avery: It’s great to be here, and thank you. It’s an exciting announcement that I hope we don’t end up spending too much time talking about because money is a lot more boring than technology. But yeah, we are very happy, both to be here and to be making the announcement.

Corey: Yeah. CRV and Insight Partners are the lead investors on the round. And it’s great to see because I’ve been using Tailscale for a while now. And it is a transformative experience for the way that I think about these things. A while back, I wrote a Lambda layer that lets Lambda functions take advantage of it, but in fairness, I did write it, so anyone looking at that should—“Haha, that’s why you’re not a developer full-time. You’re bad at it.” Yes, I am.

But I can’t stop raving about how useful Tailscale is, with the counterpoint that it’s also very difficult to explain to people who are not—at least in my experience—broken in a very particular way, as I am. What is Tailscale? And what does it do?

Avery: Right. Well, I mean, first of all, one of the things I really like about Tailscale and what we built is that, you know, even if you’re not a super great developer—like you just described yourself—you can get excited about it, you can use it for things, you can build on top of it, and contribute back without having to understand every single little detail of what it does, right? Tailscale is something that a lot of people get excited about without having to know how it works; they just know what it gives them, right? The answer to what Tailscale is, is sort of… it can be hard to explain to people who don’t know about the kinds of problems that it solves, but the super short answer is it connects all of your devices and virtual machines and containers to each other, wherever they are, without going through an intermediary, right? So, it minimizes latency and it maximizes throughput, and it minimizes pain. And it sounds like that should be hard, but you can get it all done in, like, five minutes.

Corey: I have been using it for a while now. Originally, I was using it and federating through it I believe, via Google. I rebuilt and tore down the entire network in about five minutes, instead started federating through GitHub. Nowadays, you apparently changed your position on that identity and you use third-party SSL sources, as well as retaining user information and login stuff yourselves, which is just, it’s almost starved for choice, on some level. But I am such a fan of the product that if you’ll forgive me if I talk for about a minute or so on how I use it and my experience of it.

Avery: Go for it.

Corey: So, I wind up firing up Tailscale, and I have a network that from any of my devices, I can talk to any other. I have a couple of EC2 machines hanging out in AWS, I have a Raspberry Pi that I use as a DNS server sitting in the other room, I have my iPad, I have my iPhone, I have my laptop, I have my desktop, I have a VM sitting over in Google Cloud, I have a different VM sitting over an Oracle Cloud. And all of these things can talk to each other directly over a secured network. I can override DNS and talk to these things just by the machine name, I can talk to them via the address that winds up being passed out to them through this. It is transformative. It works on IPv4, IPv6, if I’m on a network without IPv6 access using Tailscale, suddenly I can.

I can emerge from almost any other node on this network. And adding a new device to this is effectively opening a link in a browser on either that device or a different one, clicking approve once I log in, and it’s done. That is my experience of it, so far. Is that directionally correct as far as how you think about the product? Because again, I use DNS TXT records as a database for God’s sake. I am probably not the world’s foremost technical authority on the proper use of things.

Avery: Right. Yeah. I mean, that’s a good description of what it does. I think it actually—it’s weird, right? It’s hard to get across in words just how simple it is, right?

That one-minute description used a bunch of technical-sounding terminology that probably the listeners to your podcast will understand. But, like, the average tech person doesn’t need to know any of those things in order to use Tailscale, right? You download it from the app store on your phone and your laptop. And you install Tailscale on both from the App Store. You log into your Google account or your GitHub account, and that’s it. Those two devices are tied together in time and space; they can see each other. You can access a web server that you’re running on your laptop from your phone without doing anything else, right?

And then you can start a VM in AWS and you load Tailscale in there, and now that’s part of your network. And so, there’s—you don’t need to know what IPv4 and IPv6 even are. You don’t need to know what DNS even is. It just, you know, the magic sort of comes together. We do a ton of stuff behind the scenes to make that magic work. But it’s this —one thing that one customer said to us one time is, like, “It makes the internet work the way you thought the internet worked until you learned how the internet worked.” If that makes sense.

Corey: Right. It basically works on duct tape and toothpicks all spit together, and it’s amazing that it works at all. I mean, this is going to sound relatively banal, but the way that I’ve used Tailscale the most is on my phone or on my iPad or on my Mac. I will connect to the Tailscale network by default, and when that is done, it passes out my pi-hole’s IP address as the custom DNS server for the entire network. So, I don’t see a whole bunch of ads, not just in browser, but in apps and the rest.

And every once in a while when something is broken because an ad server is apparently critical to something, great, I turn off the VPN on that device, use the natural stuff. My experience of the internet gets worse as a result and the thing starts working again, then I turn it back on. It is more or less the thing that I use as a very strange-looking ad blocker, in some respects, that I can toggle on and off with the click of a button. But it’s magic, it is effectively magic. From the device side, it’s open up an app and toggle a switch, or it is grab from the menu bar on a Mac, there’s an application that runs and just click the connect button or the disconnect button.

There is no MFA every time you connect. There is no type in a username and password. There is no lengthy handshake. I hit connect and it is connected by the time I have moved the mouse back from the menu bar to the application I was working in. Whenever I show this to someone who uses a corporate VPN, they don’t believe me.

Avery: Right. Yeah, exactly. It’s hard to believe. It's like, “Hey, did anything actually happen here?” Because we removed you know, for example, it doesn’t by default catch all your traffic, it only catches the traffic to your private network, so it’s safe to leave it on all the time because it’s not interfering with what you’re doing.

What you’re describing is using Pi-Hole, which is a Raspberry Pi-based DNS server that is an ad blocker, most people using Pi-Hole have one at home, so when they’re at home they get ads blocked, but when they leave home they don’t get their ads blocked. If you add Tailscale to that, you can use your Pi-Hole even when you’re not at home, and it sort of makes it that much more useful. I think an important difference from, say, other services that you can use an adblocker or a privacy VPN is that we never see your traffic, right? Tailscale creates a private network between you and all your personal devices, and that private network is private even from us, right? We help you connect the devices to each other, but when your traffic goes to Pi-Hole, it’s your Pi-Hole. It’s not our adblocker. It’s your adblocker, right, so we never see what traffic you’re going to, we never see what DNS names you're looking up because it was just never made available to us, right?

Corey: Right. But did you do—the level of visibility you have into my network is fascinating in a variety of different ways, but it is also equally fascinating—one of those ways—is that how limited it is. You know what devices I have, the last time they’ve connected, the version of Tailscale they’re running, an IP address on it, and you also wind up seeing what services are advertised and available on those networks if I decide to enable that. Which is great for things like development; I’m going to be doing development in a local dev sense on an EC2 instance somewhere. And well, I don’t want to set up a tunnel with SSH to wind up having to proxy traffic over there just so I can wind up hitting some high port that I bound to, and I certainly don’t want to expose that to the general internet; that is a worst practice for all these things.

And Tailscale magically makes this go away. I haven’t done this in much depth yet with a variety of my team members, but when you start working on this with teams who are doing development work, someone can have something running on their laptop and just seamlessly share it with their colleagues. It’s transformative, especially in an area where very often that colleague is not sitting in the same room getting the greasy fingerprints on your laptop screen.

Avery: Yep. Yeah, exactly. So, you mentioned the services list which you have to specifically opt into, and the reason we did that is that, you know, the list of devices and hostnames and IP addresses, we have to collect because that’s how the service works, right? You send us the information about your devices, and then we send the public keys for those devices to the other devices. We can’t get out of collecting that, whereas the services list is purely an interesting add-on feature, and we decided that we didn’t want to collect that by default because it would make people nervous about their privacy.

So, if you want that feature, you click it on; if you don’t want it, don’t turn it on, you can still share services with people inside your network; they just need to know that those services exist. You send them the URL or whatever and it’ll work, but it doesn’t show up as a list of things that we can see in that case. But yeah, sharing stuff between your coworkers is definitely… is a major use case for Tailscale and dev and infrastructure teams in particular. Like, you can—designers, for example, run a test version of the website on their laptop, and then they say, “Hey, visit this URL on my laptop.” And you don’t have to be in the same office, you can both be sitting in different cafes in different cities. Tailscale will make it so that the connection between those two computers still works, even if they’re both behind firewalls, even if they’re both behind different NATs, and so on.

Corey: One of the things that astounded me the most; I am reluctant to completely trust things that are new that touch the network. Early on in my career, I made network engineering mistake 101, which is making a change to the firewall in your data center without having another way in. And the drive across town or calling remote hands to get them to let you back in and when you locked things out. Because you folks are building these things on a pretty consistent clip; there are a lot of updates and releases across all of the platforms. And invariably, I find myself on some devices version behind or so, just because of the pace of innovation. “Oh, great. We’re updating the VPN client. Cool. So, I’m going to expect this thing to drop and I’m going to have to go in and jigger it to get it working again.”

That has never happened. I have finally given in to, I guess, the iron test of this, and I have closed SSH from the internet to most of these nodes. In fact, some of them sit —the Pi-Hole sitting at home, if you’re not on my home network, there is no outside way in without breaking in. It is absolutely one of those things that disappears into the background in a way that I was extraordinarily surprised to find.

Avery: Right. Well, that is something—I mean, I’m old and grumpy, I guess, is sort of the beginning part of all this, right? I’ve seen all this annoying stuff that happens with software. And, you know, and many of us, in fact, at Tailscale are old and grumpy, and we just didn’t want to repeat those same things. So, first of all, network stuff to an even stronger degree than virtually any other kind of product, if your network stops working, everything stops working, right, so it’s number one priority that Tailscale has to not mess up your network.

Because if it does, you instantly lose faith. There’s kind of like—Tailscale gives you this magical feeling when you first install it, but that feeling of magic goes away very quickly the first time it screws something up and you can’t connect when you really need to. So, we put a huge amount of work into making sure that you can connect when you really need to. We have a lot of automated tests. One of our policies that I think is almost unheard of is that we intend to never deprecate support for older versions of the Tailscale client.

And to this day, we’re about three years into Tailscale, we’ve never deprecated an old client that anybody is using. So eventually, people—though in fact hard to believe, but eventually, people do stop using some old versions, so those ones don’t work anymore, necessarily. But any version of Tailscale that is in use today is going to keep working as long as anybody is using it. We have a very, very, very strong backwards compatibility policy. Because the worst thing that I can imagine is having some Raspberry Pi sitting out in the void somewhere that I haven’t looked at for two years, that whoops, Tailscale broke it, and now I can’t connect to it, and now I have to go drive down there and fix it, right? It would be just insultingly terrible for that to happen.

And we just make sure that doesn’t happen. Another thing that people get excited about is, like, on a Debian system or whatever, if you’ve got the Debian package installed, you can do an apt-get upgrade. Tailscale upgrades and even your SSH session doesn’t drop. Every now and then people [comment and was like 00:14:13] —

Corey: That was the weirdest part. I was expecting it to go away or hang for a long period of time. And sure, I guess it might drop a packet or so, I’ve never bothered to look because it is so seamless.

Avery: Right. Yeah, exactly. It’s just, like, “Wait. Did anything even happen?” It’s like, “Yes”—

Corey: Right—

Avery: —“Something happened. We upgraded it out from underneath you.”

Corey: —my next thing is [crosstalk 00:14:28]—yeah, I grep Tailscale on the process table. Like, okay, is this just a stale thing that’s existing [unintelligible 00:14:34] to bounce it? No, it has just been started. It was so seamless under the hood that it was amazing. There is something that is—a lot of things have been very deeply right on this.

Something else that I think is worth pointing out is that if any company had the brainpower there to roll their own crypto, it would be you folks, but you don’t. You’re riding on top of WireGuard, an open-source project that does full-mesh VPNs with terrible user interfaces.

Avery: Yep. So, you know, I guess disclosure. Back in 1997 when I started my first startup, I was not smart enough to not roll my own crypto. And therefore the VPN I wrote at the time definitely had giant security holes. It was also not that popular, so nobody found them. But I, you know eventually I found [crosstalk 00:15:21]—

Corey: “Except a bank, which I really shouldn’t disclose.” Kidding, I’m kidding. But yeah.

Avery: [laugh]. No, no, no. The bank never used that software. [laugh]. But yeah. Nowadays, I’ve been through a lot, and I… I would not describe myself as a security expert. Although people often describe me as a security expert. I don’t know what that means. But I am enough of an expert to know that I should not be rolling my own crypto. And the people who invented WireGuard, it’s one of the—I feel like I’m overstating things, but I’m not—it’s one of the biggest leaps forward in cryptography, in probably the history of computing. Now, it builds on a series of things that are part of the same leap forward, right? It’s built on the protocol that Signal uses called the Noise Protocol, right? Signal and Noise are built on the Ed25519 curve, made by —or popularized by Dan Bernstein who’s a major cryptographer in this area. Sometimes popular, sometimes—

Corey: Oh, djb.

Avery: —not popular. Yeah, exactly.

Corey: He also, near and dear to my heart, wrote djbdns, which was a well-known, widely deployed DNS server, by which I of course mean database. Please, continue.

Avery: Yep. [laugh]. I’ve been a huge fan of basically everything djb has ever made in the history of—

Corey: Oh, you’re a qmail person. I am on the postfix side of [unintelligible 00:16:37].

Avery: Yep. Well, my first startup back in 1997, we made Linux-based server appliances for small businesses. And we use qmail, we use djbdns, we used a couple of other djb products. And you know, for the history of that product—you know, leaving aside my VPN that was a security hole—the djb stuff never had a single problem. That company was eventually acquired by IBM.

One of the first things IBM did is, like, “Whoa, djb has a super-weird software license. We can’t be doing this. Let’s replace it with software that has a decent license.” So, they dropped out djbdns and started using BIND. Within a week, there was a security hole in BIND that affected all of these appliances that they now controlled, right?

So, djb is a very big-brained, super genius in security, whatever you might think of his personality. And it’s sort of like was the basis for this revolution in cryptography that WireGuard has sort of brought to the networking world. And it’s hard to overstate. Just, like, the number of lines of code, there’s something like 100 times less code to implement WireGuard than to implement IPsec. Like, that is very hard to believe, but it is actually the case.

And that made it something really powerful to build on top of. Like, it’s super hard for somebody like me to screw up the security of a WireGuard deployment, where it’s very easy to screw up the security of an IPsec deployment.

Corey: This episode is sponsored by our friends at Oracle Cloud. Counting the pennies, but still dreaming of deploying apps instead of “Hello, World” demos? Allow me to introduce you to Oracle’s Always Free tier. It provides over 20 free services and infrastructure, networking, databases, observability, management, and security. And—let me be clear here—it’s actually free. There’s no surprise billing until you intentionally and proactively upgrade your account. This means you can provision a virtual machine instance or spin up an autonomous database that manages itself, all while gaining the networking, load balancing, and storage resources that somehow never quite make it into most free tiers needed to support the application that you want to build. With Always Free, you can do things like run small-scale applications or do proof-of-concept testing without spending a dime. You know that I always like to put asterisks next to the word free? This is actually free, no asterisk. Start now. Visit snark.cloud/oci-free that’s snark.cloud/oci-free.

Corey: I just want to call something out as well, that when I say that you folks definitely have the intellectual firepower to roll your own crypto should you choose to do so, but you chose not to, if anything, I’m understating it. To be clear, one of the blog posts you had somewhat recently out was how you are maintaining what is effectively your own fork of the Go programming language. Which is one of those things when someone hears that it’s like, “I’m sorry, can you say that again? Because I am almost certain I misunderstood something.” What is the high-level version of that?

Avery: Well, there's, I think, two important points there. One of them is that yes, we did fork the Go programming language; it’s supposed to be a temporary fork because it allows us to do some experiments with the go back-end. And the primary reason we were able to do that is because we employ a couple of people who used to be on the core Go team. And that was not because we went out looking for people who used to be on the core Go team, that’s just how it worked out. But because we do, it’s easier for them to fork Go than it would be for the average person, and in many ways, it’s easier for them to get their job done by just continuing to work on the codebase they’ve already worked on.

But the second point is actually, as compilers go, the Go compiler is probably the very easiest one I’ve ever seen to be able to fork and edit. Like it’s super-clear code, you’re just editing Go code, which is already pretty easy. But they really put a ton of work into making it readable and understandable. So, like, average people actually can fork the Go compiler and not be completely bamboozled by how difficult everything is, right? Compared to, like, GCC where just building the thing is something that takes you weeks to learn how to do, right, Go is just, like, you run this script and build your compiler [unintelligible 00:19:35]—

Corey: Yeah. Let me clear this quarter on my schedule so I can go ahead and do that. Yeah, no, thank you.

Avery: Yeah. I’ve built copies of GCC and it’s absolutely nightmarish, right? And built people’s forks of GCC for special embedded processors and stuff. And this is, like, a f—this is a career that you can specialize in, building GCC, right? There are people that do this, right? And the Go compiler, it’s really—

Corey: Well, it’s 40 years of load-bearing technical debt.

Avery: Yeah. Yeah. But the Go compiler. It’s very nice; it’s just a program that’s written in Go, that compiles under Go, and then you end up with one binary, right? And as long as you have that binary, everything just works, right? And so, it’s actually surprisingly easy to fork Go. I don’t want to—you know, I wouldn’t put that on the same level of difficulty as, like, not screwing up cryptography, if you’re trying to do it yourself. [crosstalk 00:20:16]

Corey: [crosstalk 00:20:16] their own crypto algorithm that they themselves can’t defeat. Yeah, it turns out that basically, breaking crypto is a team sport. Who knew?

Avery: Yeah. Exactly. Generally, with security, you have this problem a lot, right? It’s a lot harder to build a system that nobody can break into, than it is to break into a random system, right? Because you know, the job of securing something against everybody is much harder than the job of finding something you can break into.

Corey: So, I did have a question about something you said earlier, where one of the use cases—one of the design goals—is not to have a breaking change to a point where an old device cannot still connect to the private network. But you do have a key expiry for devices where a device needs to relog in, and it can be anywhere between 3 and 180 as I look at it. I don’t know if some of the more enterprise-y options have longer options that they can set, but what happen—how do you not have to drive out to the back of beyond to re-authenticate that Raspberry Pi every six months?

Avery: Ah. So, this is something, it’s at the policy layer, and we have not finished refining this to perfection, I would say, right now. What we do have though, if your key does expire, there’s a button in the admin panel to say, like, boost this device for a little bit longer. Sort of unexpire it for another 30 minutes—I don’t remember what the—how much time it is—then you can SSH into the device and do a proper key refresh on it without actually having to drive out there. Now, we did for one version, accidentally break the key reactivation feature so that if the client noticed it’s key is expired, it actually disconnected from the Tailscale network altogether and then didn’t receive the message to, like, “Hey, could you please increase the length of your key?” That was fixable by power cycling it, which you could often get somebody to do without driving all the way out there. But we fixed that, so now that—

Corey: “Have you tried turning it off and back on again,” is still a surprisingly effective way of troubleshooting something.

Avery: Yeah, exactly. So, that wasn’t—I mean, it was kind of annoying for some people. But yeah, the reason we use, by default, every key always expires is because unlimited time credentials are one of the worst security holes that people don’t really acknowledge. Because technically, it’ll never be the, like—you know, it’ll never show up as the highest severity security hole that you have an unlimited time credential sitting in your home directory, but it is something that—well, I can tell a story. There is a company that I heard about that had you know—SSH keys are typically unlimited time credentials; the easiest way to do it is you run ssh-keygen, it puts something in your home directory, you copy the public key to all the devices you want to be able to log into, and then you never think about it again.

So, this is a company that, of course, every developer in their company had done this; they had a production network with a bunch of SSH keys in it. Some not very ethical employee worked there, had keys in their production systems, and eventually got fired. Now, of course, this company had good processes in place, they went through all the devices and took out this person’s public key from all the devices. What they didn’t know is that during lunch one day, this person had gone around to all their coworkers' workstations that hadn’t been locked, downloaded the private keys for those people on his—

Corey: Oh no.

Avery: —computer before he got fired. And so, shortly after he got fired, their entire production network got wiped out. Now, they didn’t have enough forensics at the time to know how it all got wiped out, so they spent some time putting it all back in place, this time with forensics. About a month later—they rebuilt everything from scratch, all new public keys and everything. You couldn’t possibly have any backdoors in this system, right?

And then a month later, it all got wiped out again. This time, the forensics revealed and, like, it was one of the existing employees, coming from a different country, that had gotten into their private production network and wiped everything out. How did that happen? It was because this person had years earlier, downloaded all their public—or private keys when he wandered around through the office. You can fix this problem instantly, by just expiring your keys and forcing your rotation periodically, right?

SSH doesn’t make that very easy. You can with SSH setup, SSH certificate authentication, which is a huge ordeal to get configured, but once it’s working, it solves this particular problem, right? Tailscale [crosstalk 00:24:19]—

Corey: On Mac and iOS, there is a slight improvement to this that I’m a big fan of because I agree with you. I am lousy at rotating my keys, but there’s an open-source project called Secretive that I use on the Mac that stores the private key in the Secure Enclave, which the Mac will not let out of it. And I have to use Touch ID to authenticate every time I want to connect to something. Which can get annoying from time to time, but there is no way for someone to copy that off. Historically, I would—

Avery: That’s true.

Corey: Have a passphrase that was also tied to the key so if someone grabbed it off the disk, it still theoretically would not be usable. And that was—but again, that is an absolute vector that needs to be addressed and thought about. Key rotation is huge.

Avery: And you have to go through this effort to sort it all out, right? So Tailscale, we just have this policy: We don’t do unlimited length credentials; we do key rotation for everything, and we just sort of set different time limits for this rotation depending on how picky you want to be about it. But any key expiry is much, much better than no key expiry. Even if you set it to a six-month key expiry, you still have at least it’s only the six-month window that somebody could theoretically reuse your keys. And we can also rotate keys behind the scenes and so on.

So, in the SSH case, the way people use Tailscale, you stopped opening the SSH port to the world. You’re only SSH when you’re
connected over Tailscale. The fact that your Tailscale keys rotate and expire over time is what protects your SSH session. So, you could keep using static SSH keys that never expire—don’t try to figure out all this other complicated stuff, right—and you’re still protected from these private SSH, like, unlimited length keys. Now, that said, for servers, Tailscale does have a button where you can say, like, “Please stop expiring the key.” This is a server, nobody’s ever going to get physical access to the machine.

The only thing we could do with the private key for this machine is allow other people to SSH into it, which is not very dangerous, right? It’s pretty much, like, somebody stealing your SSH authorized keys file; like, it doesn’t really matter. And for that case, you turn off the expiry altogether. But expiring keys is intended for use by, like, devices that employees are actually holding in their hands where if it expires, it’s no big deal, you push the login button and it refreshes.

Corey: There’s something that is very nice about dealing with something that is just so sensible. I mean, we’ve all—at least in the olden days of running sysadmin stuff, we had this problem we would generate—or purchase back in those days—SSL certificates and, great, they expire to a year or so at the end of the year, people forget, and then it would expire you to run around fixing this. And the default knee-jerk response was that was awful. Let’s get the next one for five years so we didn’t have to think about it that long.

And it’s always a wildcard and so it gets put all over the place, and you wind up with these problems. One of the things that Let’s Encrypt has done super well is forcing a rotation every 90 days so you know where it is. It’s just often enough you want to automate it. And ACM, the AWS certificate manager that they use, takes a slightly different approach. It doesn’t give you the private key; it embeds it in other places so they can handle the rotation themselves.

And they start screaming in your email if they can’t verify that it’s time for renewal long before it hits. It’s different approaches to the problem, but yeah, five years out, how should I know all the places the certificate has wound up in that intervening time? Most of the people who did it aren’t there anymore. And one day, surprise, a website breaks, either because its SSL cert isn’t working, or one of the back-end services it depends on suddenly doesn’t have that working. It’s become a mess, so having a forced modernity to these things is important.

Avery: Right. It’s forced modernity, and it’s just basically, it’s all behind the scenes. Like, you don’t even think about the fact that Tailscale gave you a key because that is not relevant to your day-to-day life, right? You logged in, something happened, all these devices ended up on your network. What actually happened is that public and private keys—you know, a private key was generated, the public keys were distributed properly, things are getting rotated, but you don’t have to care about all that stuff.

So, it’s fun that Tailscale is what we call secure by default, right? People love to use it because it’s easier, it makes their life easier, but security teams like it because actually, it changes the default security posture from, like, “Ugh, I’m going to have to tell everybody to please stop doing these five things because it always creates security holes,” to like, “Whoa, the thing that they’re going to do most naturally is actually going to be safe.” Right? I really like that about it. You’re not thinking about certificates, but their certificates are getting rotated exactly as they should be.

Corey: There’s just something so nice about computers doing the heavy lifting for us. It’s one of the weird things about Tailscale is it falls into a very strange spot where there is effectively zero maintenance burden on me, but I still use it to toggle it on or off in scenarios often enough to remember that it’s there and that I’m using it. It is the perfect sweet spot of being somewhat close to top of mind, but never in a sense that is, “Oh, I got to deal with this freaking thing again.” It never feels that way. Logging into it, it has long-lived sessions at the browser, so it isn’t one of those, ah, you have to go back to GitHub and re-authenticate and do all these other dog-and-pony show things. It just works. It is damn near a consumer-level of ease-of-use, start to finish. The hard part, of course, is how on earth you explain this to someone [laugh] without a background in this space.

Avery: Yeah, exactly. It’s something we ask ourselves sometimes is, like, well, you know, Tailscale is great for developers right now. It is easy enough to use, even for consumers, but, like, how would you explain it to consumers and find a good use case for consumers? And it’s something that I think we are going to do eventually, but it hasn’t been, up until now, a super high priority for us just because developers are this sort of like the core audience that we haven’t even finished building a great product that does everything that they want, yet. There is one little feature in Tailscale that’s the beginning of something that's consumer-friendly; it’s called Taildrop.

I don’t know if you’ve seen this one. You can turn it on, and basically, it acts like AirDrop in Apple products, except you don’t need to care about physical proximity and it works with every kind of device, not just Apple devices, right? So, you can add it as—it shows up in the share pane on your Mac OS or Windows or iOS device. You can use it from Linux, you just use it to send files of any type, and it sends them point to point not through a cloud provider so that we never see a copy of the file. It only goes between your devices over your encrypted network. So, that’s something that consumers kind of like.

Corey: Feels like Tailprint for Bonjour could wind up being another aspect of this as well. And I’m still hoping for something almost Ansible-like where run the following command, whether it’s pre-approved or not, on a following subset of things. In my case, for example, it’s, I would love it if it would just automatically, when I press the button, update Tailscale across all of the nodes that support it, namely the Linux boxes. I don’t think you can trigger an App Store update from within a sandboxed app on iOS, but I’ve been—

Avery: Right.

Corey: Surprised before. Yeah. But it’s nice to be able to do some things.

Avery: Yeah. This is one of those—yeah, we get that request a lot for, like, can you push a button to auto-update Tailscale? It makes me really sad that we get this request because the need for this is a sign that all of the OS vendors have completely botched software updates, right? Like, the OS should be the thing, updating your software on a good schedule based on a set of rules, and it shouldn’t be the job of every single application to provide their own software update. It’s actually a massive, embarrassing, security hole that software can even update itself, right?

Because if it can update itself, then you know, imagine someone breaks into the production services of a company that is offering a particular program. They put malware into a version of the software, they put it into the software update server, and then they trigger everything in the network to push the software update to those devices. Now, you’ve got malware installed on all your devices, right? It’s very strange that people asked for this as a feature. [laugh].

Tailscale currently does not have that feature; it doesn’t push software updates on its own. But it’s such a popular feature that I think we’re going to have to implement it because everybody wants this because Windows, for example, is simply just never going to automatically update your software for you. We have to have these weird-super admin rights on your machine so that we can push software updates because nobody else will. I feel really weird about that. You know, the security world should be protesting this more.

But instead, they’re like, asking, can you please put this feature in because I’ve got a checklist in my compliance thing that says, “Is all your software up-to-date?” I don’t have a checklist item that says, “Does any of my software have super-admin rights that they shouldn’t have?” Right? It’s sort of, I guess, the next level of supply-chain management is the big word. Nobody—there is no supply chain management for software.

Corey: There isn’t, for better or worse. I wish there were, but there simply is not. Ugh. Next year, maybe. We hope.

Avery: Yep. So, you have to trust your vendors, fundamentally, which I guess will always be true. That’s true for Tailscale as well,
right? Whether or not we include the software update pushing. If you’re installing a VPN product provided by a vendor, you have to trust that we’re going to put the right stuff into the software.

And the best—the only thing I can really do is just be honest about these issues and say, “Well, look, we try our best. We definitely try not to implement features that are going to turn into security holes for you.” And I think we do a lot better than most vendors do in that area. But it’s very hard to be perfect because nobody knows how to do software supply chain well.

Corey: Ugh. I hear you. I that’s the nice thing, too. Honestly, the big reason I know I need to update these things and the reason I want to do it’s actually you. Because whenever I log in and look at my devices in the Tailscale thing, there’s a little icon next to the one that there’s an update available here.

And you have fixed a lot of the niceties on this, like, ah, there’s an update available for the iOS version. It’s, “Really? Because it’s not available in the Apple Store yet,” as I sit there spamming the thing. That stopped happening. There’s a lot of just very nice quality-of-life improvements that are easy to miss.

Avery: Yep, yeah, that’s kind of weird. We actually went a little overboard on the update available notifications for a while because there’s always this trade-off, right? Like I said, we have a policy of never breaking old versions, so when people see the update available notification, they kind of panic. It’s, like, “Oh no, I better install the update, before Talescale cuts me off.” And, like, well, we’re not actually ever going to cut you off, so you shouldn’t have to worry about that stuff.

But on the other hand, you’re not going to get the latest features and bug fixes unless you’re running the latest version, so when people email us saying, “Hey, I’m using Tailscale from six months ago, and I have this problem,” the first thing our support team does is say, “Well, can you please try the latest one, and does the problem go away?” Because it’s kind of inefficient debugging six-month-old software. So, one way we were trying to, like, minimize that cost is, like, hey, we could just tell people there’s a new version available and then maybe they’ll update it themselves. But that resulted in people panicking. Like, oh, no, I need to install the software really, really soon because I can’t afford to break my network.

Corey: Right.

Avery: And because our system is based on WireGuard and this is —you know, I’ll probably jinx it by saying this but, like, we’ve never had an actual security hole that we’ve had to issue a Tailscale update to resolve, right? People see the update available thing and, like, “Oh, no, I bet there’s a whole bunch of vulnerabilities that they fixed.” It’s like, “Well, no.” WireGuard has also never had a vulnerability, right? [laugh] it’s… yeah, it’s, you know, sooner or later there probably will be one, and when there is one, we’ll probably have to make the, you know, update notification in red or something instead of just the little icon on the admin panel. But yeah, it’s—

Corey: [laugh].

Avery: —we try [crosstalk 00:35:23]—

Corey: Nice job on jinxing it, by the way, I appreciate that.

Avery: Yeah I know. I mean, I try to try my best. [laugh]. But I’ve actually been surprised. It’s very much like my experience with all the djb stuff we used in the past.

Like, when we were using qmail and djbdns for years, there was never once a security hole, right? It’s very interesting that it is possible to design software that never once has a security hole. And nobody does that, right? I mean, I would say I’m not as smart as djb; our software is probably, you know, not going to be as one hundred percent perfect as that, but we try really, really hard to aim for that as a goal.

Corey: Yeah. I really want to thank you for taking the time to speak with me about everything Tailscale is up to. And again, congratulations on your Series B. If people want to learn more, where should they go?

Avery: I guess, tailscale.com is the place. We also have @tailscale in Twitter. My own personal Twitter is @apenwarr, which you probably won’t be able to spell unless you Google for me or something—

Corey: But it’s in the [show notes 00:36:19], which makes this even easier.

Avery: It is? Ah, there you go. So yeah, there’s lots of information. But the number one thing I tell people is, like, look, it is a lot easier to get started than you think it is. Even after you’ve heard it 100 times, nobody ever believes how easy it is to get started. Just go to
the App Store, download the app, log into your account, and you’re already done, right? Try that and you don’t even have to read anything.

Corey: I would tear you apart for that statement if it weren’t—if it were slightly less true than it is, but it is transformative. Give it a try. It’s a strong endorsement from me. Thank you so much for your time. I appreciate it.

Avery: Thank you, too. Great talking to you, and talk next time.

Corey: Indeed. Avery Pennarun, CEO of Tailscale. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this show, please leave a five-star review on your podcast platform of choice, and smash the like and subscribe buttons, whereas if you’ve hated it, same thing—five-star review, smash the buttons—and also leave an angry bitter comment about how you are smart enough to roll your own crypto, so you don’t understand why other people wouldn’t do it.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Yoav

Yoav is a security veteran recognized on Microsoft Security Response Center’s Most Valuable Research List (BlackHat 2019). Prior to joining Orca Security, he was a Unit 8200 researcher and team leader, a chief architect at Hyperwise Security, and a security architect at Check Point Software Technologies. Yoav enjoys hunting for Linux and Windows vulnerabilities in his spare time.

Links Referenced:

  • Orca Security: https://orca.security
  • Twitter: https://twitter.com/yoavalon

TranscriptAnnouncer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Vultr. Optimized cloud compute plans have landed at Vultr to deliver lightning fast processing power, courtesy of third gen AMD EPYC processors without the IO, or hardware limitations, of a traditional multi-tenant cloud server. Starting at just 28 bucks a month, users can deploy general purpose, CPU, memory, or storage optimized cloud instances in more than 20 locations across five continents. Without looking, I know that once again, Antarctica has gotten the short end of the stick. Launch your Vultr optimized compute instance in 60 seconds or less on your choice of included operating systems, or bring your own. It's time to ditch convoluted and unpredictable giant tech company billing practices, and say goodbye to noisy neighbors and egregious egress forever. Vultr delivers the power of the cloud with none of the bloat. "Screaming in the Cloud" listeners can try Vultr for free today with a $150 in credit when they visit getvultr.com/screaming. That's G E T V U L T R.com/screaming. My thanks to them for sponsoring this ridiculous podcast.

Corey: Finding skilled DevOps engineers is a pain in the neck! And if you need to deploy a secure and compliant application to AWS, forgettaboutit! But that’s where DuploCloud can help. Their comprehensive no-code/low-code software platform guarantees a secure and compliant infrastructure in as little as two weeks, while automating the full DevSecOps lifestyle. Get started with DevOps-as-a-Service from DuploCloud so that your cloud configurations are done right the first time. Tell them I sent you and your first two months are free. To learn more visit: snark.cloud/duplocloud. Thats’s snark.cloud/D-U-P-L-O-C-L-O-U-D.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Periodically, I would say that I enjoy dealing with cloud platform security issues, except I really don’t. It’s sort of forced upon me to deal with much like a dead dog is cast into their neighbor’s yard for someone else to have to worry about. Well, invariably, it seems like it’s my yard.

And I’m only on the periphery of these things. Someone who’s much more in the trenches in the wide world of cloud security is joining me today. Yoav Alon is the CTO at Orca Security. Yoav, thank you for taking the time to join me today and suffer the slings and arrows I’ll no doubt be hurling your way.

Yoav: Thank you, Corey, for having me. I've been a longtime listener, and it’s an honor to be here.

Corey: I still am periodically surprised that anyone listens to these things. Because it’s unlike a newsletter where everyone will hit reply and give me a piece of their mind. People generally don’t wind up sending me letters about things that they hear on the podcast, so whenever I talk to somebody listens to it as, “Oh. Oh, right, I did turn the microphone on. Awesome.” So, it’s always just a little on the surreal side.

But we’re not here to talk necessarily about podcasting, or the modern version of an AM radio show. Let’s start at the very beginning. What is Orca Security, and why would folks potentially care about what it is you do?

Yoav: So, Orca Security is a cloud security company, and our vision is very simple. Given a customer’s cloud environment, we want to detect all the risks in it and implement mechanisms to prevent it from occurring. And while it sounds trivial, before Orca, it wasn’t really possible. You will have to install multiple tools and aggregate them and do a lot of manual work, and it was messy. And we wanted to change that, so we had, like, three guiding principles.

We call it seamless, so I want to detect all the risks in your environment without friction, which is our speak for fighting with your peers. We also want to detect everything so you don’t have to install, like, a tool for each issue: A tool for vulnerabilities, a tool for misconfigurations, and for sensitive data, IAM roles, and such. And we put a very high priority on context, which means telling you what’s important, what’s not. So, for example, S3 bucket open to the internet is important if it has sensitive data, not if it’s a, I don’t know, static website.

Corey: Exactly. I have a few that I’d like to get screamed at in my AWS account, like, “This is an open S3 bucket and it’s terrible.” I look at it the name is assets.lastweekinaws.com. Gee, I wonder if that’s something that’s designed to be a static hosted website.

Increasingly, I’ve been slapping CloudFront in front of those things just to make the broken warning light go away. I feel like it’s an underhanded way of driving CloudFront adoption some days, but not may not be the most charitable interpretation thereof. Orca has been top-of-mind for a lot of folks in the security community lately because let’s be clear here, dealing with security problems in cloud providers from a vendor perspective is an increasingly crowded—and clouded—space. Just because there’s so much—there’s investment pouring into it, everyone has a slightly different take on the problem, and it becomes somewhat challenging to stand out from the pack. You didn’t really stand out from the pack so much as leaped to the front of it and more or less have become the de facto name in a very short period of time, specifically—at least from my world—when you wound up having some very interesting announcements about vulnerabilities within AWS itself. You will almost certainly do a better job of relating the story, so please, what did you folks find?

Yoav: So, back in September of 2021, two of my researchers, Yanir Tsarimi and Tzah Pahima, each one of them within a relatively short span of time from each other, found a vulnerability in AWS. Tzah found a vulnerability in CloudFormation which we named BreakingFormation and Yanir found a vulnerability in AWS Glue, which we named SuperGlue. We’re not the best copywriters, but anyway—

Corey: No naming things is hard. Ask any Amazonian.

Yoav: Yes. [laugh]. So, I’ll start with BreakingFormation which caught the eyes of many. It was an XXE SSRF, which is jargon to say that we were able to read files and execute HTTP requests and read potentially sensitive data from CloudFormation servers. This one was mitigated within 26 hours by AWS, so—

Corey: That was mitigated globally.

Yoav: Yes, globally, which I’ve never seen such quick turnaround anywhere. It was an amazing security feat to see.

Corey: Particularly in light of the fact that AWS does a lot of things very right when it comes to, you know, designing cloud infrastructure. Imagine that, they’ve had 15 years of experience and basically built the idea of cloud, in some respects, at the scale that hyperscalers operate at. And one of their core tenets has always been that there’s a hard separation between regions. There are remarkably few global services, and those are treated with the utmost of care and delicacy. To the point where when something like that breaks as an issue that spans more than one region, it is headline-making news in many cases.

So it’s, they almost never wind up deploying things to all regions at the same time. That can be irksome when we’re talking about things like I want a feature that solves a problem that I have, and I have to wait months for it to hit a region that I have resources living within, but for security, stuff like this, I am surprised that going from, “This is the problem,” to, “It has been mitigated,” took place within 26 hours. I know it sounds like a long time to folks who are not deep in the space, but that is superhero speed.

Yoav: A small correction, it’s 26 hours for, like, the main regions. And it took three to four days to propagate to all regions. But still, it’s speed of lighting in for security space.

Corey: When this came out, I was speaking to a number of journalists on background about trying to wrap their head around this, and they said that, “Oh yeah, and security is always, like, the top priority for AWS, second only to uptime and reliability.” And… and I understand the perception, but I disagree with it in the sense of the nightmare scenario—that every time I mention to a security person watching the blood drain from their face is awesome—but the idea that take IAM, which as Werner said in his keynote, processes—was it 500 million or was it 500 billion requests a second, some ludicrous number—imagine fails open where everything suddenly becomes permitted. I have to imagine in that scenario, they would physically rip the power cables out of the data centers in order to stop things from going out. And that is the right move. Fortunately, I am extremely optimistic that will remain a hypothetical because that is nightmare fuel right there.

But Amazon says that security is job zero. And my cynical interpretation is that well, it wasn’t, but they forgot security, decided to bolt it on to the end, like everyone else does, and they just didn’t want to renumber all their slides, so instead of making it point one, they just put another slide in front of it and called the job zero. I’m sure that isn’t how it worked, but for those of us who procrastinate and building slide decks for talks, it has a certain resonance to it. That was one issue. The other seemed a little bit more pernicious focusing on Glue, which is their ETL-as-a-Service… service. One of them I suppose. Tell me more about it.

Yoav: So, one of the things that we found when we found the BreakingFormation when we reported the vulnerability, it led us to do a quick Google search, which led us back to the Glue service. It had references to Glue, and we started looking around it. And what we were able to do with the vulnerability is given a specific feature in Glue, which we don’t disclose at the moment, we were able to effectively take control over the account which hosts the Glue service in us-east-1. And having this control allowed us to essentially be able to impersonate the Glue service. So, every role in AWS that has a trust to the Glue service, we were able to effectively assume a role into it in any account in AWS. So, this was more critical a vulnerability in its effect.

Corey: I think on some level, the game of security has changed because for a lot of us who basically don’t have much in the way of sensitive data living in AWS—and let’s be clear, I take confidentiality extremely seriously. Our clients on the consulting side view their AWS bills themselves as extremely confidential information that Amazon stuffs into a PDF and emails every month. But still. If there’s going to be a leak, we absolutely do not want it to come from us, and that is something that we take extraordinarily seriously. But compared to other jobs I’ve had in the past, no one will die if that information gets out.

It is not the sort of thing that is going to ruin people’s lives, which is very often something that can happen in some data breaches. But in my world, one of the bad cases of a breach of someone getting access to my account is they could spin up a bunch of containers on the 17 different services that AWS offers that can run containers and mine cryptocurrency with it. And the damage to me then becomes a surprise bill. Okay, great. I can live with that.

Something that’s a lot scarier to a lot of companies with, you know, serious problems is, yep, fine, cost us money, whatever, but our access to our data is the one thing that is going to absolutely be the thing that cannot happen. So, from that perspective alone, something like Glue being able to do that is a lot more terrifying than subverting CloudFormation and being able to spin up additional resources or potentially take resources down. Is that how you folks see it too, or is—I’m sure there’s nuance I’m missing.

Yoav: So yeah, the access to data is top-of-mind for everyone. It’s a bit scary to think about it. I have to mention, again, the quick turnaround time for AWS, which almost immediately issued a patch. It was a very fast one and they mitigated, again, the issue completely within days. About your comment about data.

Data is king these days, there is nothing like data, and it has all the properties of everything that we care about. It’s expensive to store, it’s expensive to move, and it’s very expensive if it leaks. So, I think a lot of people were more alarmed about the Glue vulnerability than the CloudFormation vulnerability. And they’re right in doing so.

Corey: I do want to call out that AWS did a lot of things right in this area. Their security posture is very clearly built around defense-in-depth. The fact that they were able to disclose—after some prodding—that they checked the CloudTrail logs for the service itself, dating back to the time the service launched, and verified that there had never been an exploit of this, that is phenomenal, as opposed to the usual milquetoast statements that companies have. We have no evidence of it, which can mean that we did the same thing and we looked through all the logs in it’s great, but it can also mean that, “Oh, yeah, we probably should have logs, shouldn’t we? But let’s take a backlog item for that.” And that’s just terrifying on some level.

It becomes a clear example—a shining beacon for some of us in some cases—of doing things right from that perspective. There are other sides to it, though. As a customer, it was frustrating in the extreme to—and I mean, no offense by this—to learn about this from you rather than from the provider themselves. They wound up putting up a security notification many hours after your blog post went up, which I would also just like to point out—and we spoke about it at the time and it was a pure coincidence—but there was something that was just chef’s-kiss perfect about you announcing this on Andy Jassy’s birthday. That was just very well done.

Yoav: So, we didn’t know about Andy’s birthday. And it was—

Corey: Well, I see only one of us has a company calendar with notable executive birthdays splattered all over it.

Yoav: Yes. And it was also published around the time that AWS CISO was announced, which was also a coincidence because the date was chosen a lot of time in advance. So, we genuinely didn’t know.

Corey: Communicating around these things is always challenging because on the one hand, I can absolutely understand the cloud providers’ position on this. We had a vulnerability disclosed to us. We did our diligence and our research because we do an awful lot of things correctly and everyone is going to have vulnerabilities, let’s be serious here. I’m not sitting here shaking my fist, angry at AWS’s security model. It works, and I am very much a fan of what they do.

And I can definitely understand then, going through all of that there was no customer impact, they’ve proven it. What value is there to them telling anyone about it, I get that. Conversely, you’re a security company attempting to stand out in a very crowded market, and it is very clear that announcing things like this demonstrates a familiarity with cloud that goes beyond the common. I radically changed my position on how I thought about Orca based upon these discoveries. It went from, “Orca who,” other than the fact that you folks have sponsored various publications in the past—thanks for that—but okay, a security company. Great to, “Oh, that’s Orca. We should absolutely talk to them about a thing that we’re seeing.” It has been transformative for what I perceive to be your public reputation in the cloud security space.

So, those two things are at odds: The cloud provider doesn’t want to talk about anything and the security company absolutely wants to demonstrate a conversational fluency with what is going on in the world of cloud. And that feels like it’s got to be a very delicate balancing act to wind up coming up with answers that satisfy all parties.

Yoav: So, I just want to underline something. We don’t do what we do in order to make a marketing stand. It’s a byproduct of our work, but it’s not the goal. For the Orca Security Research Pod, which it’s the team at Orca which does this kind of research, our mission statement is to make cloud security better for everyone. Not just Orca customers; for everyone.

And you get to hear about the more shiny things like big headline vulnerabilities, but we also have very sensible blog posts explaining how to do things, how to configure things and give you more in-depth understanding into security features that the cloud providers themselves provide, which are great, and advance the state of the cloud security. I would say that having a cloud vulnerability is sort of one of those things, which makes me happy to be a cloud customer. On the one side, we had a very big vulnerability with very big impact, and the ability to access a lot of customers' data is conceptually terrifying. The flip side is that everything was mitigated by the cloud providers in warp speed compared to everything else we’ve seen in all other elements of security. And you get to sleep better knowing that it happened—so no platform is infallible—but still the cloud provider do work for you, and you’ll get a lot of added value from that.

Corey: You’ve made a few points when this first came out, and I want to address them. The first is, when I reached out to you with a, “Wow, great work.” You effectively instantly came back with, “Oh, it wasn’t me. It was members of my team.” So, let’s start there. Who was it that found these things? I’m a huge believer giving people credit for the things that they do.

The joy of being in a leadership position is if the company screws up, yeah, you take responsibility for that, whether the company does something great, yeah, you want to pass praise onto the people who actually—please don’t take this the wrong way—did the work. And not that leadership is not work, it absolutely is, but it’s a different kind of work.

Yoav: So, I am a security researcher, and I am very mindful for the effort and skill it requires to find vulnerabilities and actually do a full circle on them. And the first thing I’ll mention is Tzah Pahima, which found the BreakingFormation vulnerability and the vulnerability in CloudFormation, and Yanir Tsarimi, which found the AutoWarp vulnerability, which is the Azure vulnerability that we have not mentioned, and the Glue vulnerability, dubbed SuperGlue. Both of them are phenomenal researcher, world-class, and I’m very honored to work with them every day. It’s one of my joys.

Corey: Couchbase Capella Database-as-a-Service is flexible, full-featured and fully managed with built in access via key-value, SQL, and full-text search. Flexible JSON documents aligned to your applications and workloads. Build faster with blazing fast in-memory performance and automated replication and scaling while reducing cost. Capella has the best price performance of any fully managed document database. Visit couchbase.com/screaminginthecloud to try Capella today for free and be up and running in three minutes with no credit card required. Couchbase Capella: make your data sing.

Corey: It’s very clear that you have built an extraordinary team for people who are able to focus on vulnerability research. Which, on some level, is very interesting because you are not branded as it were as a vulnerability research company. This is not something that is your core competency; it’s not a thing that you wind up selling directly that I’m aware of. You are selling a security platform offering. So, on the one hand, it makes perfect sense that you would have a division internally that works on this, but it’s also very noteworthy, I think, that is not the core description of what it is that you do.

It is a means by which you get to the outcome you deliver for customers, not the thing that you are selling directly to them. I just find that an interesting nuance.

Yoav: Yes, it is. And I would elaborate and say that research informs the product, and the product informs research. And we get to have this fun dance where we learn new things by doing research. We [unintelligible 00:18:08] the product, and we use the customers to teach us things that we didn’t know. So, it’s one of those happy synergies.

Corey: I want to also highlight a second thing that you have mentioned and been very, I guess, on message about since news of this stuff first broke. And because it’s easy to look at this and sensationalize aspects of it, where, “See? The cloud providers security model is terrible. You shouldn’t use them. Back to data centers we go.” Is basically the line taken by an awful lot of folks trying to sell data center things.

That is not particularly helpful for the way that the world is going. And you’ve said, “Yeah, you should absolutely continue to be in cloud. Do not disrupt your cloud plan as a result.” And let’s be clear, none of the rest of us are going to find and mitigate these things with anything near the rigor or rapidity that the cloud providers can and do demonstrate.

Yoav: I totally agree. And I would say that the AWS security folks are doing a phenomenal job. I can name a few, but they’re all great. And I think that the cloud is by far a much safer alternative than on-prem. I’ve never seen issues in my on-prem environment which were critical and fixed in such a high velocity and such a massive scale.

And you always get the incremental improvements of someone really thinking about all the ins and outs of how to do security, how to do security in the cloud, how to make it faster, more reliable, without a business interruptions. It’s just phenomenal to see and phenomenal to witness how far we’ve come in such a relatively short time as an industry.

Corey: AWS in particular, has a reputation for being very good at security. I would argue that, from my perspective, Google is almost certainly slightly better at their security approach than AWS is, but to be clear, both of them are significantly further along the path than I am going to be. So great, fantastic. You also have found something interesting over in the world of Azure, and that honestly feels like a different class of vulnerability. To my understanding, the Azure vulnerability that you recently found was you could get credential material for other customers simply by asking for it on a random high port. Which is one of those—I’m almost positive I’m
misunderstanding something here. I hope. Please?

Yoav: I’m not sure you’re misunderstanding. So, I would just emphasize that the vulnerability again, was found by Yanir Tsarimi. And what he found was, he used a service called Azure Automation which enables you essentially to run a Python script on various events and schedules. And he opened the python script and he tried different ports. And one of the high ports he found, essentially gave him his credentials. And he said, “Oh, wait. That’s a really odd port for an HTTP server. Let’s try, I don’t know, a few ports on either way.” And he started getting credentials from other customers. Which was very surprising to us.

Corey: That is understating it by a couple orders of magnitude. Yes, like, “Huh. That seems sub-optimal,” is sort of like the corporate messaging approved thing. At the time you discover that—I’m certain it was a three-minute-long blistering string of profanity in no fewer than four languages.

Yoav: I said to him that this is, like, a dishonorable bug because he worked very little to find it. So it was, from start to finish, the entire research took less than two hours, which, in my mind, is not enough for this kind of vulnerability. You have to work a lot harder to get it. So.

Corey: Yeah, exactly. My perception is that when there are security issues that I have stumbled over—for example, I gave a talk at re:Invent about it in the before times, one of them was an overly broad permission in a managed IAM policy for SageMaker. Okay, great. That was something that obviously was not good, but it also was more of a privilege escalation style of approach. It wasn’t, “Oh, by the way, here’s the keys to everything.”

That is the type of vulnerability I have come to expect, by and large, from cloud providers. We’re just going to give you access credentials for other customers is one of those areas that… it bugs me on a visceral level, not because I’m necessarily exposed personally, but because it more or less shores up so many of the arguments that I have spent the last eight years having with folks are like, “Oh, you can’t go to cloud. Your data should live on your own stuff. It’s more secure that way.” And we were finally it feels like starting to turn a cultural corner on these things.

And then something like that happens, and it—almost have those naysayers become vindicated for it. And it’s… it almost feels, on some level, and I don’t mean to be overly unkind on this, but it’s like, you are absolutely going to be in a better security position with the cloud providers. Except to Azure. And perhaps that is unfair, but it seems like Azure’s level of security rigor is nowhere near that of the other two. Is that generally how you’re seeing things?

Yoav: I would say that they have seen more security issues than most other cloud providers. And they also have a very strong culture of report things to us, and we’re very streamlined into patching those and giving credit where credit’s due. And they give out bounties, which is an incentives for more research to happen on those platforms. So, I wouldn’t say this categorically, but I would say that the optics are not very good. Generally, the cloud providers are much safer than on-prem because you only hear very seldom on security issues in the cloud.

You hear literally every other day on issues happening to on-prem environments all over the place. And people just say they expect it to be this way. Most of the time, it’s not even a headline. Like, “Company X affected with cryptocurrency or whatever.” It happens every single day, and multiple times a day, breaches which are massively bigger. And people who don’t want to be in the cloud will find every reason not to be the cloud. Let us have fun.

Corey: One of the interesting parts about this is that so many breaches that are on-prem are just never discovered because no one knows what the heck’s running in an environment. And the breaches that we hear about are just the ones that someone had at least enough wherewithal to find out that, “Huh. That shouldn’t be the way that it is. Let’s dig deeper.” And that’s a bad day for everyone. I mean, no one enjoys those conversations and those moments.

And let’s be clear, I am surprisingly optimistic about the future of Azure Security. It’s like, “All right, you have a magic wand. What would you do to fix it?” It’s, “Well, I’d probably, you know, hire Charlie Bell and get out of his way,” is not a bad answer as far as how these things go. But it takes time to reform a culture, to wind up building in security as a foundational principle. It’s not something you can slap on after the fact.

And perhaps this is unfair. But Microsoft has 30 years of history now of getting the world accustomed to oh, yeah, just periodically, terrible vulnerabilities are going to be discovered in your desktop software. And every once a month on Tuesdays, we’re going to roll out a whole bunch of patches, and here you go. Make sure you turn on security updates, yadda, yadda, yadda. That doesn’t fly in the cloud. It’s like, “Oh, yeah, here’s this month’s list of security problems on your cloud provider.” That’s one of those things that, like, the record-scratch, freeze-frame moment of wait, what are we doing here, exactly?

Yoav: So, I would say that they also have a very long history of making those turnarounds. Bill Gates famously did his speech where security comes first, and they have done a very, very long journey and turn around the company from doing things a lot quicker and a lot safer. It doesn’t mean they’re perfect; everyone will have bugs, and Azure will have more people finding bugs into it in the near future, but security is a journey, and they’ve not started from zero. They’re doing a lot of work. I would say it’s going to take time.

Corey: The last topic I want to explore a little bit is—and again, please don’t take this as anyway being insulting or disparaging to your company, but I am actively annoyed that you exist. By which I mean that if I go into my AWS account, and I want to configure it to be secure. Great. It’s not a matter of turning on the security service, it’s turning on the dozen or so security services that then round up to something like GuardDuty that then, in turn, rounds up to something like Security Hub. And you look at not only the sheer number of these services and the level of complexity inherent to them, but then the bill comes in and you do some quick math and realize that getting breached would have been less expensive than what you’re spending on all of these things.

And somehow—the fact that it’s complex, I understand; computers are like that. The fact that there is—[audio break 00:27:03] a great messaging story that's cohesive around this, I come to accept that because it’s AWS; talking is not their strong suit. Basically declining to comment is. But the thing that galls me is that they are selling these services and not inexpensively either, so it almost feels, on some level like, shouldn’t this on some of the built into the offerings that you folks are giving us?

And don’t get me wrong, I’m glad that you exist because bringing order to a lot of that chaos is incredibly important. But I can’t
shake the feeling that this should be a foundational part of any cloud offering. I’m guessing you might have a slightly different opinion than mine. I don’t think you show up at the office every morning, “I hate that we exist.”

Yoav: No. And I’ll add a bit of context and nuance. So, for every other company than cloud providers, we expect them to be very good at most things, but not exceptional at everything. I’ll give the Redshift example. Redshift is a pretty good offering, but Snowflake is a much better offering for a much wider range of—

Corey: And there’s a reason we’re about to become Snowflake customers ourselves.

Yoav: So, yeah. And there are a few other examples of that. A security company, a company that is focused solely on your security
will be much better suited to help you, in a lot of cases more than the platform. And we work actively with AWS, Azure, and GCP requesting new features, helping us find places where we can shed more light and be more proactive. And we help to advance the conversation and make it a lot more actionable and improve from year to year. It’s one of those collaborations. I think the cloud providers can do anything, but they can’t do everything. And they do a very good job at security; it doesn’t mean they’re perfect.

Corey: As you folks are doing an excellent job of demonstrating. Again, I’m glad you folks exist; I’m very glad that you are publishing the research that you are. It’s doing a lot to bring a lot I guess a lot of the undue credit that I was giving AWS for years of, “No, no, it’s not that they don’t have vulnerabilities like everyone else does. It just that they don’t ever talk about them.” And they’re operationalizing of security response is phenomenal to watch.

It’s one of those things where I think you’ve succeeded and what you said earlier that you were looking to achieve, which is elevating the state of cloud security for everyone, not just Orca customers.

Yoav: Thank you.

Corey: Thank you. I really appreciate your taking the time out of your day to speak with me. If people want to learn more, where’s the best place they can go to do that?

Yoav: So, we have our website at orca.security. And you can reach me out on Twitter. My handle is at @yoavalon, which is @-Y-O-A-V-A-L-O-N.

Corey: And we will of course put links to that in the [show notes 00:29:44]. Thanks so much for your time. I appreciate it.

Yoav: Thank you, Corey.

Corey: Yoav Alon, Chief Technology Officer at Orca Security. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, or of course on YouTube, smash the like and subscribe buttons because that’s what they do on that platform. Whereas if you’ve hated this podcast, please do the exact same thing, five-star review, smash the like and subscribe buttons on YouTube, but also leave an angry comment that includes a link that is both suspicious and frightening, and when we click on it, suddenly our phones will all begin mining cryptocurrency.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Kate

Kate Holterhoff, an industry analyst with RedMonk, has a background in frontend engineering, academic research, and technical communication. Kate comes to RedMonk from the digital marketing sector and brings with her expertise in frontend engineering, QA, accessibility, and scrum best practices.

Before pursuing a career in the tech industry Kate taught writing and communication courses at several East Coast universities. She earned a PhD from Carnegie Mellon in 2016 and was awarded a postdoctoral fellowship (2016-2018) at Georgia Tech, where she is currently an affiliated researcher.

Links:

  • RedMonk: https://redmonk.com/
  • Visual Haggard: https://visualhaggard.org
  • Twitter: https://twitter.com/kateholterhoff

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Couchbase Capella Database-as-a-Service is flexible, full-featured, and fully managed with built-in access via key-value, SQL, and full-text search. Flexible JSON documents aligned to your applications and workloads. Build faster with blazing fast in-memory performance and automated replication and scaling while reducing cost. Capella has the best price-performance of any fully managed document database. Visit couchbase.com/screaminginthecloud to try Capella today for free and be up and running in three minutes with no credit card required. Couchbase Capella: Make your data sing.

Corey: This episode is sponsored by our friends at Revelo. Revelo is the Spanish word of the day, and its spelled R-E-V-E-L-O. It means, “I reveal.” Now, have you tried to hire an engineer lately? I assure you it is significantly harder than it sounds. One of the things that Revelo has recognized is something I’ve been talking about for a while, specifically that while talent is evenly distributed, opportunity is absolutely not. They’re exposing a new talent pool to, basically, those of us without a presence in Latin America via their platform. It’s the largest tech talent marketplace in Latin America with over a million engineers in their network, which includes—but isn’t limited to—talent in Mexico, Costa Rica, Brazil, and Argentina. Now, not only do they wind up spreading all of their talent on English ability, as well as you know, their engineering skills, but they go significantly beyond that. Some of the folks on their platform are hands down the most talented engineers that I’ve ever spoken to. Let’s also not forget that Latin America has high time zone overlap with what we have here in the United States, so you can hire full-time remote engineers who share most of the workday as your team. It’s an end-to-end talent service, so you can find and hire engineers in Central and South America without having to worry about, frankly, the colossal pain of cross-border payroll and benefits and compliance because Revelo handles all of it. If you’re hiring engineers, check out revelo.io/screaming to get 20% off your first three months. That’s R-E-V-E-L-O dot I-O slash screaming.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Every once in a while on the Twitters, I see a glorious notification. Now, doesn’t happen often, but when it does, I have all well, atwitter, if you’ll pardon the term. They have brought someone new in over at RedMonk.

RedMonk has been a longtime friend of the show. They’re one of the only companies that can say that about and not immediately get a cease-and-desist for having said that. And their most recent hire is joining me today. Kate Holterhoff is a newly minted analyst over at RedMonk. Kate, thank you for joining me.

Kate: It’s great to be here.

Corey: One of the things that’s always interesting about RedMonk is how many different directions you folks seem to go in all at once. It seems that I keep crossing paths with you folks almost constantly: When I’m talking to clients, when I’m talking to folks in the industry. And it could easily be assumed that you folks are 20, 30, 40 people, but to my understanding, there are not quite that many of you there.

Kate: That is very true. Yes. I am the fifth analyst on a team of seven. And yeah, brought on the first of the year, and I’m thrilled to be here. I actually, I would say, recruited by one of my friends at Georgia Tech, Kelly Fitzpatrick, who I taught technical communication with when we were both postdocs in their Brittain Fellowship program.

Corey: So, you obviously came out of an academic background. Is this your first excursion to industry?

Kate: No, actually. After getting my PhD in literary and cultural studies at Carnegie Mellon in 2016, I moved to Atlanta and took a
postdoc at Georgia Tech. And after that was kind of winding down, I decided to make the jump to industry. So, my first position out of that was at a digital marketing agency in Atlanta. And I was a frontend engineer for several years.

Towards the end of my tenure there, I moved into doing more of their production engineering and QA work. Although it was deeply tied to my frontend work, so we spent a lot of time looking at how the web sites look at different media queries, making sure that there were no odd break points. So, it certainly was an organic move there as their team expanded.

Corey: You spent significant amounts of time in the academic landscape. When you start talking about, “Well, I took on a postdoc position,” that’s usually the sign of not your first year on a college campus in most cases. I mean, again, with an eighth grade
education, I’m not really the person to ask, but I sit here in awe as people who are steeped in academia wind up going about the magic that, from where I sit, they tend to do. What was it that made you decide that I really enjoy the field that I’ve gotten a doctorate in. You just recently published a book in that is—or at least tangentially related to this space.

But you decide, “You know what I really want to do now? That’s right, frontend engineering. I want to spend, more or less, 40-some-odd hours a week slowly going mad because CSS, and I can’t quite get that thing to line up the way that I want it to.” Now, at least that’s my experience with it, for folks who are, you know, competent at it, I presume that’s a bit of a different story.

Kate: Yes. I considered naming my blog at RedMonk, “How to Center a Div.” So yes, that is certainly an ongoing issue, I think, for anyone in [unintelligible 00:06:15] any, you know, practitioners. So, I guess my story probably began in 2013, the real move into technology. So, getting a PhD, of course, takes a very, very long time.

So, I started at Carnegie Mellon in 2009, and in 2013, I started a digital archive called Visual Haggard. And it’s a Ruby on Rails site; you can visit it at visualhaggard.org. And it is a digital archive of illustrations that were created to accompany a 19th century writer, H. Rider Haggard.

And I became very interested in all the illustrations that had been created to accompany both the serialization of his fictions, but also the later novelizations. And it’s kind of like how we have all these different movie adaptations of, like, Spider Man that come out every couple of years. These illustrations were just very iterative. And generally, this fellowship that I saw really only focused on, you know, the first illustrations that, you know, came out. So, this was a sort of response to that: How can we use technology to showcase all the different types of illustrations and how maybe different artists would interpret that literature differently?

And so, that drove me into a discipline called the digital humanities, which really sort of, you know, focuses on that question, which is, you know, how to computers help us to understand the humanities better? And so, that incorporates not only the arts, but also literature, philosophy, you know, new media. But it’s an extremely broad subject, and it’s evolving, as you can imagine, as the things that technology can do expands. So, I became interested in this subject and really was drawn to the sort of archival aspects of this. Which wasn’t really my training; I think that’s something that, you know, you think of librarians as being more focused on, but I became acquainted with all these, you know, very obscure editions.

But in any event, it also taught me how to [laugh] use technology, I really—I was involved in the [RDF 00:08:08] export for [laugh] incorporating the site on Nines, which is sort of a larger agglomeration of 19th century archives. And I was just really drawn to a lot of the new things that we could do. So, I began to use it more in my teaching. So, not only did I—and of course as I taught communication courses at Carnegie Mellon, and then I moved to teaching them at Georgia Tech, you can imagine I had many students who were engineers, and they were very interested in these sorts of questions as well. So, the move felt very organic to me, but I think any academic that you speak to, their identity is very tied up in their sort of, you know, academic standing.

And so, the idea of jumping ship, of not being labeled an academic anymore is kind of terrifying. But I, you know, ultimately opted to do it. It certainly was, yeah, but you know, what [laugh] what I learned is that there’s the status called an affiliated researcher. So, I didn’t necessarily have to be a professor or someone on the tenure track in order to continue doing research.

Corey: Was it hard for you?

Kate: So, the book project, which is titled Illustration in Fin-de-Siècle Transatlantic Romance Fiction, and has a chapter devoted to
H. Rider Haggard, I wrote it, while really not even being an instructor or sort of traditional academic. I had access to the library through this affiliated researcher status, which I maintained by keeping a relationship with the folks at Georgia Tech, and was able to do all my research while you know, having a job in industry. And I think what a lot of academics need to do is think about what it is about academia that they really value. Is it the teaching?

Because in industry, we spend a lot of time teaching [laugh]. Sharing our knowledge is something that’s extremely important. Is that the research? As an analyst, I get to do research all the time, which is really fun for me. And then, you know, is it really just kind of focusing on historical aspects? And that was also important to me.

So, you know, this status allowed me to keep all the best parts of being an academic while kind of sloughing off the [laugh] parts that weren’t so good, which is, um, say the fact that 80% of courses in the university are taught now by adjuncts or folks who are not on the tenure track line. Which is, you know, pretty shocking, you know. The academy is going through some… troubles right now, and hiring issues are—they need to be acknowledged, and I think folks who are considering getting a PhD need to look for other career paths beyond just through modeling it on their advisors, or, you know, in order to become, ostensibly, a professor themselves.

Corey: I don’t know if I’ve told the story before in public, but I briefly explored the possibility of getting a PhD myself, which is interesting given that I’d have to… there’s some prerequisites I’d probably have to nail first, like, get a formal GED might be, like, step one, before proceeding on. And strangely enough for me, it was not the higher level, I guess, contribution to a body of knowledge in a particular direction. I mean, cloud economics being sort of an easy direction for me [laugh] to go in, given that I eat, sleep, live, and breathe it, but rather the academic rigor around so much of it. And the incentives feel very different, which to be clear, is a good thing. My entire career path has always been focused on not starving to death, and how do we turn this problem into money, whereas academia has always seemed to be focused on knowledge for the sake of knowledge without much, if any, thought toward the practical application slash monetization thereof? Is that a fair characterization from where you sit? I’m trying not to actively be insulting, but it’s possible I may be unintentionally so.

Kate: No, I think you’re right on. And so yeah, like, the book that I published, I probably won’t see any remuneration for that. There is very little—I’m actually [laugh] not even sure what the contract says, but I don’t intend to make any money with this. Professors, even those who have reached the height of their career, unless they’re, you know, on specific paths, don’t make a lot of money, those in the humanities, especially. You don’t do this to become wealthy.

And the Visual Haggard archive, I don’t—you know, everything is under a Creative Commons license. I don’t make money from people, you know, finding images that they’re looking for to reproduce, say, on a t-shirt or something. So yeah, I suspect you do it for the love. I always explained it as having a sort of existential anxiety of, like, trying to, you know, cheat death. I think it was Umberto Eco who said that in order to live forever, you have to have a child and a book.

And at this point, I have two children and a book now, so I can just, you know, die and my, you know, [laugh] my legacy lives on. But I do feel like the reasons that folks go into upper higher education vary, and so I wouldn’t want to speak for everyone. But for me, yeah, it is not a place to make money, it’s a place to establish sort of more intangible benefits.

Corey: This episode is sponsored in part by our friends at ChaosSearch. You could run Elasticsearch or Elastic Cloud—or OpenSearch as they’re calling it now—or a self-hosted ELK stack. But why? ChaosSearch gives you the same API you’ve come to know and tolerate, along with unlimited data retention and no data movement. Just throw your data into S3 and proceed from there as you would expect. This is great for IT operations folks, for app performance monitoring, cybersecurity. If you’re using Elasticsearch, consider not running Elasticsearch. They’re also available now in the AWS marketplace if you’d prefer not to go direct and have half of whatever you pay them count towards your EDB commitment. Discover what companies like Klarna, Equifax, Armor Security, and Blackboard already have. To learn more, visit chaossearch.io and tell them I sent you just so you can see them facepalm, yet again.

Corey: I guess one of the weird things from where I sit is looking at the broad sweep of industry and what I know of RedMonks perspective, you mentioned that as a postdoc, you taught technical communication. Then you went to go to frontend engineering, which in many respects is about effectively, technically—highly technical and communicating with the end-user. And now you are an analyst at RedMonk. And seeing what I have seen of your organization in the larger ecosystem, teaching technical communication is a terrific descriptor of what it is you folks actually do. So, from a certain point of view, I would argue that you’re still pursuing the path that you are on in some respects. Is that even slightly close to the way that you view things, or am I just more or less ineffectively grasping at straws, as I am wont to do?

Kate: No, I feel like there is a continuous thread. So, even before I got my PhD, I got a—one of my bachelor’s degrees was in art. So, I used to paint murals; I was very interested in public art. And so, it you know, it feels to me that there is this thread that goes from an interest in the arts and how the public can access them to, you know, doing web development that’s focused on the visual aspects, you know, how are these things responsive? What is it that actually makes the DOM communicate in this visual way? You know, how are cascading style sheets,allowing us to do these sorts of marvelous things?

You know, I could talk about my favorite, you know, selectors and things. [laugh]. Because I will defend CSS. I actually don’t hate it, although we use SASS if it matters. But you know, that I think there’s a lot to be said for the way that the web looks today rather than, you know, 20 years ago.

So there, it feels very natural to me to have moved from an interest in illustration to trying to, you know, work in a more frontend way, but then ultimately [laugh] move from that into doing, sort of, QA, which is, like, well, let’s take a look at how we’re communicating visually and see if we can improve that to, you know, look for things that maybe aren’t coming across as well as they could. Which really forced me to work in the interactive team more with the UI/UX folks who are, you know, obviously telling the designers where to put the buttons and, you know, how to structure the, you know, the text blocks in relationship to the images and things like that. So, it feels natural to me, although it might not seem so on the outside. You know, in the process, I really I guess, acquired a love of that entire area.

And I think what’s great about working at RedMonk now is that I get to see how these technologies are evolving. So, you know, I actually just spun up a site on [unintelligible 00:16:27] not long ago. And, I mean, it is so cool. I mean, you know, coming from a background where we were working with, you know, jQuery, [laugh] things have really evolved. You know, it’s exciting. And I think we’re seeing the, [like, as 00:16:39] the full stack approach to this.

Corey: I used to volunteer for the jQuery infrastructure team and help run jquery.net, once upon a time.

Kate: Ohh.

Corey: I assume that is probably why it is no longer in vogue. Like, oh, Corey was too close to it got his stink all over the thing. Let’s find something better immediately, which honestly, not the worst approach in the world to take.

Kate: I’m so impressed. I had no idea.

Corey: It was mostly—because again, I was bad at frontend; always have been, but I know how to make computers run—kind of—and on the backend side of things and the infrastructure piece of it. It’s like I tend to—at least at the time—break the world into more or less three sets: You had the ops types, think of database admins and the rest; you had the backend engineers, people who wrote code that made things talk to each other from an API perspective, and you had frontend folks who took all of the nonsense and had this innovative idea that, “Huh, maybe a green screen glowing text terminal isn’t the pinnacle of user experience that we might possibly think about, and start turning it into something that a human being can use.”

And whatever I hear folks from one of those constituencies start talking disparagingly about the others, it’s… yeah, go walk a mile in their shoes and then tell me how you feel. A couple years ago, I took a two week break to, all right, it’s time for me to learn JavaScript. And by the end of the two weeks period, I was more confused than I was when I began. And it’s just a very different way of thinking than I have become accustomed to working with. So, from where I sit, people who work on that stuff successfully are effectively just this side of wizards.

I think that there’s—I feel the same way about database types. That’s an area I never go into either because I’m terrible at that, and the stakes over their company-killing proportions in a way that I took down a web server usually doesn’t.

Kate: Yeah, I think that’s often the motto, well, at least at my last company, which was like, “It’s just a website. No one will die.”
[laugh].

Corey: Honestly, I find that the people who have really have the best attitude about that tend to be, strangely enough, military veterans because it’s, “The site is down. How are you so calm?” It’s, “Well, no one’s shooting at me and no one’s going to die? It’s fine.” Like, “We’re all going to go home to our families tonight. It’ll work out.” It having perspective is important.

Kate: Yeah. It is interesting how the impetus—I mean, going back to your question about, you know, making money at this field, you know, how that kind of factors in, I guess, frontend does tend to have a more relaxed attitude than say, yeah, if you drop a table or something. But at the same time, you know, compared to academia, it did feel a little bit more [laugh] like, “Okay, well, this—you know, we’ve got the project manager that is breathing down our neck. They got to send them something, you know, what’s going on here?” So, yeah, it does become a little bit more, I don’t know, these things ramped up a little bit, and the importance, you know, varies by, you know, whatever part of it you’re working on.

It’s interesting, as an analyst, I don’t hear the terms backend and frontend as much, and that was really how my team was divided, you know? It was really, kind of, opaque when you walked in. Started the job, I was like, “Okay, well, is this something that the frontend should be dealing with or the backend? You know, what’s going on?” And then, you know, ultimately, I was like, “Oh, no, I know exactly what this is.”

And then anyone who came on later, I was like, “No, no, no. We talk to the backend folks for this sort of problem.” So, I don’t know if that’s also something that’s falling out of vogue, but that was, you know, the backend handled all the DevOps aspects as well, and so, you know, anything with our virtual boxes and, you know, trying to get things running and, you know, access to our… yeah, the servers, you know, all of that was kind of handled by backend. But yeah, I worked with some really fantastic frontend, folks. They were just—I feel like they we could bet had been better categorized as full stack. And many of them have CS degrees and they chose to go into frontend. So, you know, it’s a—I have no patience for, you know—

Corey: Oh God, you mean you chose this instead of it being something that happened to you in a horrible accident one of these days?

Kate: [laugh]. Exactly.

Corey: And that’s not restricted to frontend; that’s working with computers, in my experience.

Kate: [laugh].

Corey: Like, oh, God, it’s hard to remember I chose this at one point. Now, it feels almost like I’m not suited for anything else. You have a clear ability to effectively communicate technical concepts. If not, you more or less wasted most of your academic career,
let’s be very clear. Then you decided that you’re going to go and be an engineer for a while, and you did that.

Why RedMonk? Why was that the next step because with that combination of skills, the world is very much your oyster. What made you look at RedMonk and say, “Yes, this is where I should work?” And let me be very clear. There are days I have strongly considered, like, if I weren’t doing this, where would I be? And yeah, I would probably annoy RedMonk into actively blocking me on all social media or hiring me. There’s no third option there. So, I agree wholeheartedly with the decision. What was it that made it for you?

Kate: I mean, it was certainly not just one thing. One of the parts of academia that I really enjoyed was the ability to go to conferences and just travel and really get to meet people. And so, that was something that seemed to be a big part of it [unintelligible 00:21:27] so that’s kind of the part that maybe doesn’t get mentioned so much. And then especially in the Covid era, you know, we’re not doing as much traveling, as you’re well aware.

Corey: We’re spending all of our time having these conversations via screen.

Kate: You know, I do enjoy that.

Corey: Yeah. Like in the before times, probably one out of every eight episodes or so of this show was recorded in person.

Kate: Wow.

Corey: Now, it’s, “I don’t know. I don’t really know if I want to go across town.” It’s a—honestly, I’ve become a bit of a shut-in here. But you get it down to a science. But you lose something by doing it.

Kate: That’s true.

Corey: There’s a lack of high bandwidth communication.

Kate: And many of my academic friends, when they would go to conferences, they would just kind of hide in their hotel room until they had to present. And I was the kind of person that was down in the bar hanging out. So, to me, it [laugh] felt very natural. But in terms of the intellectual parts, in all seriousness, I think the ability to pull apart arguments is something that I just truly enjoy. So, when I was teaching, which of course was how—was why they paid me to be an academic, you know, I loved when I could sit in a classroom and I would ask a question. You know, I kind of come up with these questions ahead of time.

And the students would say something totally unexpected, and then I’d have another one, say something totally out of the blue as well. And I get to take them and say, “You’re both right. Here’s how we combine them, and here’s how we’re going to move forward.” Sort of, the ability to take an argument and sort of mold it into something constructive, I think can be very useful, both in, you know, meeting with clients who maybe are, you know, coming at things a little bit differently than then maybe we would recommend in order to, you know, help them to reach developers, the practitioners, but also, you know, moderating panels is something that a lot of my colleagues do. I mean, that’s a big part of the job, too, is, you know, speaking and… well, not only doing sort of keynote talks, which my colleague Rachel is doing that at, I think, a [GlueCon 00:23:14] this year.

And then—but also, you know, just in video format, you know, to having multiple presenters and, kind of, taking their ideas and making something out of that sort of forwards the argument. I think that’s a lot of fun. I like to think I do an okay job at it. And I certainly have a lot of experience with it. And then just finally, you know, listening to argument [unintelligible 00:23:30] a big part of the job is going to briefings where clients explain what their product does, and we listen and try to give them feedback about how to reach the developer audience, and, you know, just trying to work on that communication aspect.

And I think what I would like to push is more of the visual part of this. So, I think a lot of times, people don’t always think through the icons that they include, or the illustrations, or the just the stock photos. And I find those so fascinating. [laugh]. I know, that’s not always the most—the part that everyone wants to focus on, but to me, the visuals of these pitches are truly interesting. They really, kind of, maybe say things that they don’t intend always, and that also can really make concrete ideas that are, especially with some of this really complex technology, it can really help potential buyers to understand what it accomplishes better.

Corey: Some of the endless engagements I’ve been on that I enjoy the most have been around talking to vendors who are making things. And it starts off invariably as, “Yeah, we want to go ahead and tell the world about this thing that we’ve done.” And my perspective has always been just a subtle frame shift. It’s like, “Yeah, let me save some time. No one cares. Absolutely no one cares. You’re in love with the technical thing that you built, and the only people who are going to love it as much as you do are either wanting to work where you, or they’re going to go build their own and they’re not going to be your customer. So, don’t talk about you. No one cares about you. Talk about the pain that you solve. Talk about the painful thing that you’re target customer is struggling with that you make disappear.”

And I didn’t think that would be, A, as revelatory as it turned out to be, and B, a lesson that I had to learn myself. When I was starting o—when I was doing some product development here where I once again fell into the easy trap of assuming if I know something, everyone must know it, therefore, it’s easy, whereas if I don’t know something, it’s very hard, and no one could possibly wrap their head around it. And we all come from different places, and meeting people wherever they are in their journey, it’s a delicate lesson to learn. I never understood what analysts did until I started being an analyst myself, and I’ve got to level with you, I spent six months of doing those types of engagements feeling like a giant fraud. I’m just a loudmouth with an opinion, what is what does that mean?

Well, in many ways, it means analyst. Because it’s having an opinion is in so many ways, what customers are really after. Raw data, you can find that a thousand different ways, but finding someone who could talk on what something means, that’s harder. And I think that we don’t teach anything approaching that in most of our STEM curriculum.

Kate: Yeah, I think that’s really on point. Yeah, I mean, especially when some of these briefings are so mired in acronyms, and sort of assumed specialization. I know I spend a lot of time just thinking about what it is that confuses me about their pitch, more so than what, you know, is actually coming through. So, I think actually, one of the tools that we use—writing instructors; my past life—was thinking like someone with an eighth grade education. So, I actually think that your reference to having [laugh] you know, that’s sort of chestnut, that can actually be useful because you say, “If I, you know, took my slide deck and showed it to a bunch of eighth graders, would they understand what it is that I’m saying?”

You know, maybe you don’t want them to get the technical details, but what problem does it solve? If they don’t understand that, you’re not doing a good enough job. And so that, to me, is [laugh] actually something that a lot of folks need to hear. That yeah, these vendors because they’re just so deep in it, they’re so in the weeds, that they can’t maybe see how someone who’s just looking for a database, or a platform, or whatever, they actually need this sort of simplified and yet broad enough explanation for what it is that they’re actually trying to do what service they actually provide.

Corey: From where I sit, one of the hardest things is just reaching people in the right way. And I’m putting out a one to two-thousand word blog post every week because I apparently hate myself. And that was a constant struggle for me when I started doing that a year or two ago. And what has worked for me that really get me moving down that is, instead of trying to teach everyone all the things, I pick an individual—and it varies from week to week—that I think about and I want to explain something to that person. And then I wind up directing what it is I’m about to tell—what it is I’m writing—to that person.

Sometimes they’re a complete layperson. Other times they are fairly advanced in a particular area of technology. And the responses to these things differ, but it’s always—I always learn something from the feedback that I get. And if nothing else, is one of those ways to become a better writer. While I would start by writing. Just do it, don’t whine—don’t worry about getting it perfect; just go out there and power through things.

At least, that’s my approach. And I’m talking about the burden of writing a thousand words a week. You wrote an actual book. My belief is that, the more people I’ve talked to who’ve done that, no one actually wants to write a book; people want to have written a book, and that definitely resonates with me. I am tempted to just slap a bunch of these—

Kate: Yeah.

Corey: —blogs posts together and call it a book one of these days as an anthology. But it feels like it’s cheating. If I ever decide to go down that path, I want to do it right.

Kate: I guess, I come at it from the perspective of I don’t know what I think until I write it down. So, it helps me to formulate ideas better. I also feel like my strength is in rereading things and trying to edit them down to really get to the kernel of what it is I think. And a lot of times how I begin a chapter or a blog post or whatever is not where it should begin, that maybe I’m somewhere in the middle, maybe this is a conclusion. There’s something magical, in my view, that [laugh] happens when you write, that you are able to pause and take a little bit more time and maybe come up with a better word for what it is that you’re trying to communicate.

I also am—I benefit from readers. So, for instance, in my book, I have one chapter that really focuses on Harper’s Weekly, which is an American newspaper. I’m not an Americanist; I don’t have a deep knowledge of that, so what I did is I revise that chapter and send it to American periodicals and got feedback from their readers. Super useful. In terms of my blog at RedMonk, anytime I publish something, you can bet that at least one founder and probably at least one other analyst has read it through and giving me some extremely incisive feedback. It never is just from my mind. It’s something that is collaboration.

And I am grateful to anyone who takes the time to read my writing because, you know, all of us have so much time, of course. It really helps me to understand what it is that I’m trying to dig into. So, for instance, I’ve been writing a series for RedMonk on certifications, which makes a lot of sense; I’ve come from an academic background, here it is, you know, I’m seeing all these tech certifications. And so, it’s interesting to me to see similarities and differences and what sort of issues that we’re seeing come up with them. So, for instance, I just wrote about the vendor-specific versus vendor-neutral certifications. What are the advantages of getting a certification from the CN/CF versus from say, VMware and—

Corey: Oh, I have opinions, on all of [those 00:30:44]—

Kate: I—

Corey: —and most of them are terrible.

Kate: —I’m sure you do. [laugh]. It came naturally out of the job, you know, sitting through briefings and, kind of, seeing these things evolve, and the questions that I have from a long history of teaching, but. I think it also suggests the collaborative aspect of this, of coming to my colleagues—you know, I’ve been here before, for what, four months?—and saying, you know, “Is this normal? Like, what are we seeing here? Let me write a little bit about what I think is going on with certifications, and then you tell me, you know, what it is that you’ve seen with your years and years of expertise,” right?

So, Stephen O’Grady’s been doing this for longer than he really likes to admit, right? So, this is grateful to have such well-established colleagues that can help me on that journey. But, you know, to kind of spiral back to your original question, I think that writing to me is an exploration, it’s something that helps me to get to something a little more, I guess, meaningful than just where I began. You know, just the questions that I have, I can kind of dig down and find some substance there. I would encourage you to take any one of your blog posts and think about maybe where they—or using the jumping off points for your eventual book, which I will be looking for on newsstands any day now.

Corey: I am looking forward to seeing how you continue to evolve your coverage area, as well as reading more of your writings around these things. I am—they always say that the cobblers children have no shoes, and I am having an ongoing war with the RedMonk RSS feed because I’ve been subscribed to it three times now, and I’m still not seeing everything that comes through, such as your posts. Time for me to go and yell at some people over on your end about how these things work because it is such good content. And every time RedMonk puts something out, it doesn’t matter who over there has written it, I wind up reading it with this sense of envy, in that I wish I had written something like this. It is always an experience, and your writing is absolutely no exception to that. You fit in well over there.

Kate: It means a lot to me. Thank you. [laugh].

Corey: No, thank you. I want to thank you for spending so much time talking to me about things that I feel like I’m still not quite smart enough to wrap my head around, but that’s all right. If people want to learn more, where’s the best place to find you?

Kate: Certainly Twitter. So, my Twitter handle is just my name, @kateholterhoff. And I don’t post as often as maybe I should, but I try to maintain an ongoing presence there.

Corey: And we will of course, put a link to that in the [show notes 00:33:04].

Kate: Thank you.

Corey: Thank you so much for your time. I appreciate it. Kate Holterhoff, analyst at RedMonk. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice—or if you’re on YouTube, smash that like and subscribe button—whereas if you’ve hated this podcast, please do the exact same thing—five-star review, smashed buttons—but then leave an angry, incoherent comment, and it’s going to be extremely incoherent because you never learned to properly, technically communicate.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor
recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Venkat

Venkat Venkataramani is CEO and co-founder of Rockset. In his role, Venkat helps organizations build, grow and compete with data by making real-time analytics accessible to developers and data teams everywhere. Prior to founding Rockset in 2016, he was an Engineering Director for the Facebook infrastructure team that managed online data services for 1.5 billion users. These systems scaled 1000x during Venkat's eight years at Facebook, serving five billion queries per second at single-digit millisecond latency and five 9's of reliability. Venkat and his team also created and contributed to many noted data technologies and open-source projects, including Facebook's TAO distributed data store, RocksDB, Memcached, MySQL, MongoRocks, and others. Prior to Facebook, Venkat worked on tools to make the Oracle database easier to manage. He has a master’s in computer science from the University of Wisconsin-Madison, and bachelor’s in computer science from the National Institute of Technology, Tiruchirappalli.

Links Referenced:

  • Company website: https://rockset.com
  • Company blog: https://rockset.com/blog

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored by our friends at Revelo. Revelo is the Spanish word of the day, and its spelled R-E-V-E-L-O. It means “I reveal.” Now, have you tried to hire an engineer lately? I assure you it is significantly harder than it sounds. One of the things that Revelo has recognized is something I’ve been talking about for a while, specifically that while talent is evenly distributed, opportunity is absolutely not. They’re exposing a new talent pool to, basically, those of us without a presence in Latin America via their platform. It’s the largest tech talent marketplace in Latin America with over a million engineers in their network, which includes—but isn’t limited to—talent in Mexico, Costa Rica, Brazil, and Argentina. Now, not only do they wind up spreading all of their talent on English ability, as well as you know, their engineering skills, but they go significantly beyond that. Some of the folks on their platform are hands down the most talented engineers that I’ve ever spoken to. Let’s also not forget that Latin America has high time zone overlap with what we have here in the United States, so you can hire full-time remote engineers who share most of the workday as your team. It’s an end-to-end talent service, so you can find and hire engineers in Central and South America without having to worry about, frankly, the colossal pain of cross-border payroll and benefits and compliance because Revelo handles all of it. If you’re hiring engineers, check out revelo.io/screaming to get 20% off your first three months. That’s R-E-V-E-L-O dot I-O slash screaming.

Corey: This episode is sponsored in part by LaunchDarkly. Take a look at what it takes to get your code into production. I’m going to just guess that it’s awful because it’s always awful. No one loves their deployment process. What if launching new features didn’t require you to do a full-on code and possibly infrastructure deploy? What if you could test on a small subset of users and then roll it back immediately if results aren’t what you expect? LaunchDarkly does exactly this. To learn more, visit launchdarkly.com and tell them Corey sent you, and watch for the wince.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Today’s promoted guest episode is one of those questions I really like to ask because it can often come across as incredibly, well, direct, which is one of the things I love doing. In this case, the question that I am asking is, when you look around at the list of colossal blunders that people make in the course of careers in technology and the rest, it’s one of the most common is, “Oh, yeah. I don’t like the way that this thing works, so I’m going to build my own database.” That is the siren call to engineers, and it is often the prelude to horrifying disasters. Today, my guest is Venkat Venkataramani, co-founder and CEO at Rockset. Venkat, thank you for joining me.

Venkat: Thanks for having me, Corey. It’s a pleasure to be here.

Corey: So, it is easy for me to sit here in my beautiful ivory tower that is crumbling down around me and use my favorite slash the best database imaginable, which is TXT records shoved into Route 53. Now, there are certainly better databases than that for most use cases. Almost anything really, to be honest with you, because that is a terrifying pattern; good joke, terrible practice. What is Rockset as we look at the broad landscape of things that store data?

Venkat: Rockset is a real-time analytics platform built for the cloud. Let me break that down a little bit, right? I think it’s a very good question when you say does the world really need another database? Don’t we have enough already? SQL databases, NoSQL databases, warehouses, and lake houses now.

So, if you really break it down, the first digital transformation that happened in the ’80s was when people actually retired pen and paper records and started using a relational database to actually manage their business records and what have you instead of ledgers and books and what have you. And that was the first digital transformation. That was—and Oracle called the rows in a table ‘records’ for a reason. They’re called records to this date. And then, you know, 20 years later, when all businesses were doing system of record and transactions and transactional databases, then analytics was born, right?

This was, like, the whole reason why I wanted to make better data-driven business decisions, and BI was born, warehouses and data lakes started becoming more and more mainstream. And there was really a second category of database management systems because the first category it was very good at to be a system of record, but not really good at complex analytics that businesses are asking to be able to guide their decisions. Fast-forward 20 years from then, the nature of applications are changing. The world is going from batch to real-time, your data never stops coming, advent of Apache Kafka and technologies like that, 5G, IoTs, data is coming from all sorts of nooks and corners within an enterprise, and now customers in enterprises are acquiring the data in real-time at a scale that the world has never seen before.

Now, how do you get analytics out of that? And then if you look at the database market—entire market—there are still only two large categories of databases: OLTP databases for transaction processing, and warehouses and data lakes for batch analytics. Now suddenly, you need the speed of OLTP at the scale of batch, right, in terms of, like, complexity of compute, complexity of storage. So, that is really why we thought the data management space needs that third leg, and we call it real-time analytics platform or real-time analytics processing. And this is where the data never stops coming; the queries never stopped coming.

You need the speed and the scale, and it’s about time we innovate and solve the problem well because in 2015, 2016, when I was researching for this, every company that was looking to solve build applications that were real-time applications was building a custom Rube Goldberg machine of sorts. And it was insanely complex, it was insanely expensive. Fast-forward now, you can build a real-time application in a matter of hours with the simplicity of the cloud using Rockset.

Corey: There’s a lot to be said that the way we used to do things after the first transformation and we got into the world of batch processing, where—in the days of punch cards, which was a bit before my time and I believe yours as well—where they would drop them off and then the next day, or two days, they would come back later after the run, they would get the results only to figure out syntax error because you put the wrong card first or something like that. And it was maddening. In time, that got better, but still, nightly runs have become a thing to the point where even now, by default, if you wind up looking at the typical timing of a default Linux install, for example, you see that in the middle of the night is when a bunch of things will rotate when various cleanup jobs get done, et cetera, et cetera. And that seemed like a weird direction to go in. One of the most famous Google April Fools Day jokes was when they put out their white paper on MapReduce.

And then Yahoo fell for it hook, line, and sinker, built out Hadoop, and we’ve been stuck with this idea of performing these big query jobs on top of existing giant piles of data, where ideally, you can measure it with a wall clock; in practice, you often measure the calendar in some cases. And as the world continues to evolve, being able to do streaming processing and understand in real-time what is going on, is unlocking different approaches, at least by all accounts. Do you have an example you can give me of a problem that real-time analytics solves for a customer? Because I can sit here and talk all day about how things might theoretically work, but I have to get out of my Route 53-based ivory tower over here, what are customers seeing?

Venkat: That’s a great question. And I want one hundred percent agree. I think Google did build MapReduce, and I think it’s a very nice continuation of what happened there and what is happening in the world now. And built MapReduce and they quickly realized re-indexing the whole world [laugh] every night, as the size of the internet is exploding is a bad idea. And you know how Google index is now? They do real-time indexing.

That is how they index the wor—you know, web. And they look for the changes that are happening in the internet, and they only index the changes. And that is exactly the same principle behind—one of the core principles behind Rockset’s real-time analytics platform. So, what is the customer story? So, let me give you one of my favorite ones.

So, the world’s number one or number two buy now, pay later company, they have hundreds of millions of users, they have 300,000-plus merchants, they operate in, like, maybe 100-plus countries, so many different payment methods, you can imagine the complexity. At any given point in time, some part of the product is broken, well, Apple Pay stopped working in Switzerland for this e-commerce merchant. Oh God, like, we got to first detect that. Forget even debugging and figuring out what happened and having an incident response team. So, what did they do as they scale the number of payments processed in the system across the world—it’s, like, in millions; first, it was millions in the day, and there was millions in an hour—so like everybody else, they built a batch-based system.

So, they would accumulate all these payment records, and every six hours—so initially, it was a day, and then afterwards, you know, you try to see how far I can push it, and they couldn’t push it beyond every six hours. Every six hours, some batch job would come and process through all the payments that happened, have some statistical models to detect, hey, here are some of the things that you might want to double-click and follow up on. And as they were scaling, the batch job that they will kick off every six hours was starting to take more than six hours. So, you can see how the story goes. Now, fast-forward, they came to us and say—it’s almost like Rockset has, like, a big red button that says, “Real-time this.”

And then they kind of like, “Can you make this real-time? Because not only that we are losing millions of potential revenue dollars in a year because something stops working and we’re not processing payments, and we don’t find out about that up to, like, three hours later, five hours later, six hours later, but our merchants are also very unhappy. We are also not able to protect our customers’ business because that is all we are about.” And so fast-forward, they use Rockset, and simply using SQL now they have all the metrics and statistical computation that they want to do, happens in real-time, that are accurate up to the second. All of their anomaly detectors run every minute and the anomaly detectors take, like, hundreds of milliseconds to run.

And so, now they’ve cut down the business observability, I would say. It’s not metrics and machine observability is actually the—you know, they have now business observability in real-time. And that not only actually saves them a lot of potential revenue loss from downtimes, that’s also allowing them to build a better product and give their customers a better experience because they are now telling their merchants and their customers that something is not working in some part of your e-commerce footprint before even the customers notice that something is wrong. And that allows them to build a better product and a better customer experience than their competitors. So, this is a very real-world example of why companies and enterprises are moving from batch to real-time.

Corey: With the stories that you, and frankly, a lot of other data analytics companies tend to fall back on all the time has been stories of the ones you’re telling, where you’re talking about the largest buy now, pay later lender, for example. These are companies operating at massive scale who have tremendous existing transaction volume, and they’re built out already. That’s great, but then I wanted to try to cut to the truth of some of these things. And when I visit your pricing page at Rockset, it doesn’t have what I would expect if that were the only use case. And what that would be is, “Great. Call here to conta—open up a sales quote, and we’ll talk to you et cetera, et cetera, et cetera.”

And the answer then is, “Okay, I know it’s going to have at least two commas in it, ideally, not three, but okay, great.” Instead, you have a free tier where it’s, “Hey, we’ll give you a pile of credits, here’s some limits on our free account, et cetera, et cetera.” Great. That is awesome. So, it tells me that there is a use case here for folks who have not already, on some level, made a good show of starting the process of conquering the world.

Rather, someone with an idea some evening at two in the morning can wind up diving in and getting started. What is the Twitter for Pets, in my garage, spare-time side project story for using something like Rockset? What problem will I have as I wind up building those things out, when I don’t have any user traffic or data yet, but I want to, you know for once in my life, do the smart thing in advance rather than building an impressive tower of technical debt?

Venkat: That is the first thing we built, by the way. When we finish our product, the first thing we built was self-service. The first thing we built was a free forever tier, which has certain limits because somebody has to pay the bill, right? And then we also have compute instances that are very, very affordable that cost you, like, approximately $1 a day. And so, we built all of that because real-time analytics is not a need that only, like, the large-scale companies have. And I’ll give you a very, very simple example.

Let’s say you’re building a game, it’s a mobile game. You can use Amazon DynamoDB and use AWS Lambdas and have a serverless stack and, like, you’re really only paying… you’re kind of keeping your footprint very, very small, and you’re able to build a very lively game and see if it gets [wider 00:12:16], and it’s growing. And once it grows, you can have all the big company scaling problems. But in the early days, you’re just getting started. Now, if you think about DynamoDB and Lambdas and whatnot, you can build almost every part of the game except probably the leaderboard.

So, how do I build a leaderboard when thousands of people are playing and all of their individual gameplays and scores and everything is just another simple record in DynamoDB. It’s all serverless. But DynamoDB doesn’t give me a SQL SELECT *, order by score, limit 100, distinct by the same player. No, this is a analytical question, and it has to be updated in real-time, otherwise, you really don’t have this thing where I just finished playing. I go to the leaderboard, and within a second or two, if it doesn’t update, you kind of lose people along the way. So, this is one of actually a very popular use case, when the scale is much smaller, which is, like, Rockset augments NoSQL database like a Dynamo or a Mongo where you can continue to use that for—or even a Postgres or MySQL for that case where you can use that as your system of record and keep it small, but cover all of your compute-heavy and analytical parts of your application with Rockset.

So, it’s almost like kind of a CQRS pattern where you use your OLTP database as your system of record, you connect Rockset to it, and so—Rockset comes in with built-in connectors, by the way, so you don’t have to write a single line of code for your inserts and updates and deletes in your transactional database to get reflected in Rockset within one to two seconds. And so now, all of a sudden you have a fully indexed, fast SQL replica of your transactional database that on which you can do all sorts of analytical queries and that’s fully isolated with your transactional database. So, this is the pattern that I’m talking about. The mobile leaderboard is an example of that pattern where it comes in very handy. But you can imagine almost everybody building some kind of an application has certain parts of it that is very analytical in nature. And by augmenting your transactional database with Rockset, you can have your cake and eat it too.

Corey: One of the challenges I think that at least I’ve run into when it comes to working with data—and let’s be clear, I tend to deal with data in relatively small volumes, mostly. The stuff that’s significantly large, like, oh, I don’t know, AWS bills from large organizations, the format of those is mostly predefined. When I’m building something out, we’re using, I don’t know, DynamoDB or being dangerous with SQLite or whatnot, invariably I find that even at small-scale, I paint myself into a corner by data model design or how I wind up structuring access or the rest, and the thing that I’m doing that makes perfect sense today winds up being incredibly challenging to change later. And I still, in production and have a DynamoDB table that has the word ‘test’ in its name because of course I do.

It’s not a great place to find yourself in some cases. And I’m curious as to what you’ve seen, as you’ve been building this out and watching customers, especially ones who already had significant datasets as they move to you. Do you have any guidance around how to avoid falling down that particular well?

Venkat: I will say a lot of the complexity in this world is by solving the right problems using the wrong tool, or by solving the right problem on the wrong part of the stack. I’ll unpack this a little bit, right? So, when your patterns change, your application is getting more complex, it is demanding more things, that doesn’t necessarily mean the first part of the application you build—and let’s say DynamoDB was your solution for that—was the wrong choice. That is the right choice, but now you’re expanded the scope of your application and the demand that you have on your backend transactional database. And now you have to ask the question, now in the expanded scope, which ones are still more of the same category of things on why I chose Dynamo and which ones are actually not at all?

And so, instead of going and abusing the GSIs and other really complex and expensive indexing options and whatnot, that Dynamo, you know, has built, and has all sorts of limitations, instead of that, what do I really need and what is the best tool for the job, right? What is the best system for that? And how do I augment? And how do I manage these things? And this goes to the first thing I said, which is, like, this tremendous complexity when you start to build a Rube Goldberg machine of sorts.

Okay, now, I’m going to start making changes to Dynamo. Oh, God, like, how do I pick up all of those things and not miss a single record? Now, replicate that to another second system that is going to be search-centric or reporting-centric, and do I have to rethink this once in a while? Do I have to build and manage these pipelines? And suddenly, instead of going from one system to two system, you actually end up going from one system to, like, four different things that with all the pipes and tubes going into the middle.

And so, this is what we really observed. And so, when you come in to Rockset and you point us at your DynamoDB table, you don’t write a single line of code, and Rockset will automatically scan your Dynamo tables, move that into Rockset, and in real-time, your changes, insert, updates, deletes to Dynamo will be reflected in Rockset. And this is all using Dynamo Streams API, Dynamo Scan API, and whatnot, behind the scenes. And this just gives you an example of if you use the right tool for the job here, when suddenly your application is demanding analytical queries on Dynamo, and you do the right research and find the right tool, your complexity doesn’t explode at all, and you can still, again, continue to use Dynamo for what it is very, very good at while augmenting that with a system built for analytics with full-featured SQL and other capabilities that I can talk about, for the parts of your application for which Dynamo is not a good fit. And so, if you use the right tool for the job, you should be in very good place.

The other thing is part about this wrong part of the stack. I’ll give a very kind of naive example, and then maybe you can extrapolate that to, like, other patterns on how people could—you know, accidental complexities the worst. So, let’s just say you need to implement access control on your data. Let’s say the best place to implement access control is at the database level, just happens to be that is the right thing. But this database that I picked, doesn’t really have role-based access control or what have you, it doesn’t really give me all the security features to be able to protect the data the way I want it.

So, then what I’m going to do is, I’m going to go look at all the places that is actually having business logic and querying the database and I’m going to put a whole bunch of permission management and roles and privileges, and you can just see how that will be so error-prone, so hard to maintain, and it will be impossible to scale. And this is what is the worst form of accidental complexity because if you had just looked at it that one week or two weeks, how do I get something out, or the database I picked doesn’t have it, and then the two weeks, you feel like you made some progress by, kind of like, putting some duct tape if conditions on all the access paths. But now, [laugh] you’ve just painted yourself into a really, really bad corner.

And so, this is another variation of the same problem where you end up solving the right problems in the wrong part of the stack, and that just introduces tremendous amount of accidental complexity. And so, I think yeah, both of these are the common pitfalls that I think people make. I think it’s easy to avoid them. I would say there’s so much research, there’s so much content, and if you know how to search for these things, they’re available in the internet. It’s a beautiful place. [laugh]. But I guess you have to know how to search for these things. But in my experience, these are the two common pitfalls a lot of people fall into and paint themselves in a corner.

Corey: Couchbase Capella Database-as-a-Service is flexible, full-featured and fully managed with built in access via key-value, SQL, and full-text search. Flexible JSON documents aligned to your applications and workloads. Build faster with blazing fast in-memory performance and automated replication and scaling while reducing cost. Capella has the best price performance of any fully managed document database. Visit couchbase.com/screaminginthecloud to try Capella today for free and be up and running in three minutes with no credit card required. Couchbase Capella: make your data sing.

Corey: A question I have, though, that is an extension is this—and I want to give some flavor to it—but why is there a market for real-time analytics? And what I mean by that is, early on in my tenure of fixing horrifying AWS bills, I saw a giant pile of money being hurled over at effectively a MapReduce cluster for Elastic MapReduce. Great. Okay, well, stream-processing is kind of a thing; what about migrating to that? Well, that was a complete non-starter because it wasn’t just the job running on those things; there were downstream jobs, and with their own downstream jobs. There were thousands of business processes tied to that thing.

And similarly, the idea of real-time analytics, we don’t have any use for that because of, oh I don’t know, I only wind up pulling these reports on a once-a-week basis, and that’s fine, so what do I need that updated for in real-time if I’m looking at them once a week? In practice, the answer is often something aligned with the, “Well, yeah, but you had a real-time updating dashboard, you would find that more useful than those reports.” But people’s expectations and business processes have shaped themselves around constraints that now can be removed, but how do you get them to see that? How do you get them to buy in on that? And then how do you untangle that enormous pile of previous constraint into something that leverages the technology that’s now available for a brighter future?

Venkat: I think [unintelligible 00:21:40] a really good question, who are the people moving to real-time analytics? What do they see? And why can they do it with other tech? Like, you know, as you say… EMR, you know, it’s just MapReduce; can’t I just run it in sort of every twenty-four hours, every six hours, every hour? How about every five minutes? It doesn’t work that way.

Corey: How about I spin up a whole bunch of parallel clusters on different timescales so I constantly—

Venkat: [laugh].

Corey: Have a new report coming in. It’s real-time, except—

Venkat: Exactly.

Corey: You’re constantly putting out new ones, but they’re just six hours delayed every time.

Venkat: Exactly. So, you don’t really want to do this. And so, let me unpack it one at a time, right? I mean, we talked about a very good example of a business team which is building business observability at the buy now, pay later company. That’s a very clear value-prop on why they want to go from batch to real-time because it saves their company tremendous losses—potential losses—and also allows them to build a better product.

So, it could be a marketing operations team looking to get more real-time observability to see what campaigns are working well today and how do I double down and make sure my ad budget for the day is put to good use? I don’t have to mention security operations, you know, needing real-time. Don’t tell me I got owned three days ago. Tell me—[laugh] somebody is, you know, breaking glass and might be, you know, entering into your house right now. And tell me then and not three days later, you know—

Corey: “Yeah, what alert system do you have for security intrusion?” “So, I read the front page of_The New York Times_ every morning and waiting to see my company’s name.” Yeah, there probably are better ways to reduce that cycle time.

Venkat: Exactly, right. And so, that is really the need, right? Like, I think more and more business teams are saying, “I need operational intelligence and not business intelligence.” Don’t make me play Monday morning quarterback.

My favorite analogy is it’s the middle of the third quarter. I’m six points down. A couple of people, star players in my team and my opponent’s team are injured, but there’s some in offense, some in defense. What plays do I do and how do I play the game slightly differently to change the outcome of the game and win this game as opposed to losing by six points. So, that I think is kind of really what is driving businesses.

You know, I want to be more agile, I want to be more nimble, and take, kind of, being data-driven decision-making to another level. So that, I think, is the real force in play. So, now the real question is, why can they do it already? Because if you go ask a hundred people, “Do you want fast analytics on real-time data or slow analytics on stale data?” How many people are going to say give me slow and stale? Zero, right? Exactly zero people.

So, but then why hasn’t it happened yet? I think it goes back to the world only has seen two kinds of databases: Transaction processing systems, built for system of record, don’t lose my data kind of systems; and then batch analytics, you know, all these warehouses and data lakes. And so, in real-time analytics use cases, the data never stops coming, so you have to actually need a system that is running 24/7. And then what happens is, as soon as you build a real-time dashboard, like this example that you gave, which is, like, I just want all of these dashboards to automatically update all the time, immediately people respond, says, “But I’m not going to be like Clockwork Orange, you know, toothpicks in my eyelids and be staring at this 24/7. Can you do something to alert or detect some anomalies and tap on my shoulder when something off is going on?”

And so, now what happens is somebody’s actually—a program more than a person—is actually actively monitoring all of these metrics and graphs and doing some analysis, and only bringing this to your attention when you really need to because something is off, right? So, then suddenly what happens is you went from, accumulate all the data and run a batch report to [unintelligible 00:25:16], like, the data never stops coming, the queries never stopped coming, I never stop asking questions; it’s just a programmatic way of asking those things. And at that point, you have a data app. This is not a analytics dashboard report anymore. You have a full-fledged application.

In fact, that application is harder to build and scale than any application you’ve ever built before [laugh] because in those situations, again, you don’t have this torrent of data coming in all the time and complex analytical questions you’re asking on the data 24/7, you know? And so, that I think is really why real-time analytics platform has to be built as almost a third leg. So, this is what we call data apps, which is when your data never stops coming and your queries never stop coming. So, this is really, I think, what is pushing all the expensive EMR clusters or misusing your warehouse, misusing your data lakes. At the end of the day, is what is I think blowing up your Snowflake bills, is what blowing up your warehouse builds because you somehow accidentally use the wrong tool for the job [laugh] going back to the one that we just talked about.

You accidentally say, “Oh, God, like, I just need some real-time.” With enough thrust, pigs can fly. Is that a good idea? Probably not, right? And so, I don’t want to be building a data app on my warehouse just because I can. You should probably use the best tool for the job, and really use something that was built ground up for it.

And I’ll give you one technical insight about how real-time analytics platforms are different than warehouses.

Corey: Please. I’m here for this.

Venkat: Yes. So really, if you think about warehouses and data lakes, I call them storage-optimized systems. I’ve been building databases all my life, so if I have to really build a database that is for batch analytics, you just break down all of your expenses in terms of let’s say, compute and storage. What I’m burning 24/7 is storage. Compute comes and goes when I’m doing a batch data load, or I’m running—an analyst who logs in and tries to run some queries.

But what I’m actually burning 24/7 is storage, so I want to compress the heck out of the data, and I want to store it in very cheap media. I want to store it—and I want to make the storage as cheap as possible, so I want to optimize the heck out of the storage use. And I want to make computation on that possible but not efficient. I can shuffle things around and make the analysis possible, but I’m not trying to be compute-efficient. And we just talked about how, as soon as you get into real-time analytics, you very quickly get into the data app business. You’re not building a real-time dashboard anymore, you’re actually building your application.

So, as soon as you get into that, what happens is you start burning both storage and compute 24/7. And we all know, relatively, [laugh] compute and RAM is about a hundred to a thousand times more expensive than storage in the grand scheme of things. And
so, if you actually go and look at your Snowflake bill, if you go look at your warehouse bill—BigQuery, no matter what—I bet the computational part of it is about 90 to 95% of the bill and not the storage. And then, if you again, break down, okay, who’s spending all the compute, and you’ll very quickly narrow down all these real-time-y and data app-y use cases where you can never turn off the compute on your warehouse or your BigQuery, and those are the ones that are blowing up your costs and complexity. And on the Rockset side, we are actually not storage-optimized; we’re compute-optimized.

So, we index all the data as it comes in. And so, the storage actually goes slightly higher because the, you know, we stored the data and also the indexes of those data automatically, but we usually fold the computational cost to a quarter of what a typical warehouse needs. So, the TCO for our customers goes down by two to four folds, you know? It goes down by half or even to a quarter of what they used to spend. Even though their storage cost goes up in net, that is a very, very small fraction of their spend.

And so really, I think, good real-time analytics platforms are all compute-optimized and not storage-optimized, and that is what allows them to be a lot more efficient at being the backend for these data applications.

Corey: As someone who spends a lot of time staring into the depths of AWS bills, I think that people also lose sight of the reality that it doesn’t matter what you’re spending on AWS; it invariably pales in comparison to what you’re spending on people to work with these things. The reason to go to cloud is not because it is the cheapest possible way to get computers to do things; it’s because it’s a capability story. It’s about unlocking capacity and capabilities you do not have otherwise. And that dramatically increases your feature velocity and it lets you to achieve things faster, sooner, with better results. And unlocking a capability is always going to be more interesting to a company than saving money on it. When a company cares first, last, and always about just save money, make the bill lower, the end, it’s usually a company in decline. Or alternately, something very strange is going on over there.

Venkat: I agree with that. One of our favorite customers told us that Rockset took their six-month roadmap and shrunk it to a single afternoon. And their supply chain SaaS backend for heavy construction, 80% of concrete that are being delivered and tracked in North America follows through their platform, and Rockset powers all of their real-time analytics and reporting. And before Rockset, what did they have? They had built a beautiful serverless stack using DynamoDB, even have AWS Lambdas and what-have-you.

And why did they have to do all serverless? Because the entire team was two people. [laugh]. And maybe a third person once in a while, they’ll get, so 2.5. Brilliant people, like, you know, really pioneers of building an entire data stack on AWS in a serverless fashion; no pipes, no ETL.

And then they were like, oh God, finally, I have to do something because my business demands and my customers are demanding real-time reporting on all of these concrete trucks and aggregate trucks delivering stuff. And real-time reporting is the name of the game for them, and so how do I power this? So, I have to build a whole bunch of pipes, deliver it to, like, some Elasticsearch or some kind of like a cluster that I had to keep up in real-time. And this will take me a couple of months, that will take me a couple of months. They came into Rockset on a Thursday, built their MVP over the weekend, and they had the first working version of their product the following Tuesday.

So—and then, you know, there was no turning back at that point, not a single line of code was written. You know, you just go and create an account with Rockset, point us at your Dynamo, and then off you go. You know, you can use start using SQL and go start building your real-time application. So again, I think the tremendous value, I think a lot of customers like us, and a lot of customers love us. And if you really ask them what is one thing about Rockset that you really like, I think it’ll come back to the same thing, which is, you gave me a lot of time back.

What I thought would take six months is now a week. What I thought would be three weeks, we got that in a day. And that allows me to focus on my business. I want to spend more time with my stakeholders, you know, my CPO, my sales teams, and see what they need to grow our business and succeed, and not build yet another data pipeline and have data pipelines and other things coming out of my nose, you know? So, at the end of the day, the simplicity aspects of it is very, very important for real-time analytics because, you know, we can’t really realize our vision for real-time being the new default in every enterprise for whenever analytics concern without making it very, very simple and accessible to everybody.

And so, that continues to be one of our core thing. And I think you’re absolutely right when you say the biggest expense is actually the people and the time and the energy they have to spend. And not having to stand up a huge data ops team that is building and managing all of these things, is probably the number one reason why customers really, really like working with our product.

Corey: I want to thank you for taking so much time to talk me through what you’re working on these days. If people want to learn more, where’s the best place to find you?

Venkat: We are Rockset, I’ll spell it out for your listeners ROCKSET—rock set—rockset.com. You can go there, you can start a free trial. There is a blog, rockset.com/blog has a prolific blog that is very active. We have all sorts of stories there, and you know engineers talking about how they implemented certain things, to customer case studies.

So, if you’re really interested in this space, that’s one on space to follow and watch. If you’re interested in giving this a spin, you know, you can go to rockset.com and start a free trial. If you want to talk to someone, there is, like, a ‘Request Demo’ button there; you click it and one of our solutions people or somebody that is more familiar with Rockset would get in touch with you and you can have a conversation with them.

Corey: Excellent. And links to that will of course go in the [show notes 00:34:20]. Thank you so much for your time today. I appreciate it.

Venkat: Thanks, Corey. It was great.

Corey: Venkat Venkataramani, co-founder and CEO at Rockset. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an insulting crappy comment that I will immediately see show up on my real-time dashboard.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Chris

Chris Harris is Vice President, Global Field Engineering at Couchbase, a provider of a leading modern database for enterprise applications that 30% of the Fortune 100 depend on. With almost 20 years of technical field and professional services experience at early-stage, open source and growth technology companies, Chris held leadership roles at Cloudera, Hortonworks, MongoDB and others before joining Couchbase.

Links Referenced:

  • couchbase.com: https://couchbase.com
  • LinkedIn: https://www.linkedin.com/in/chris-harris-5451953/
  • Twitter: https://twitter.com/cj_harris5

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Couchbase Capella Database-as-a-Service is flexible, full-featured and fully managed with built in access via key-value, SQL, and full-text search. Flexible JSON documents aligned to your applications and workloads. Build faster with blazing fast in-memory performance and automated replication and scaling while reducing cost. Capella has the best price performance of any fully managed document database. Visit couchbase.com/screaminginthecloud to try Capella today for free and be up and running in three minutes with no credit card required. Couchbase Capella: make your data sing.

Corey: This episode is sponsored in part by LaunchDarkly. Take a look at what it takes to get your code into production. I’m going to just guess that it’s awful because it’s always awful. No one loves their deployment process. What if launching new features didn’t require you to do a full-on code and possibly infrastructure deploy? What if you could test on a small subset of users and then roll it back immediately if results aren’t what you expect? LaunchDarkly does exactly this. To learn more, visit launchdarkly.com and tell them Corey sent you, and watch for the wince.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. One of the stranger parts of running this show is when I have a promoted guest episode like this one, where someone comes on, and great, “Oh, where do you work?” And the answer is a database company. Well, great, unless it’s Route 53, it’s clearly not the best database in the world, but let’s talk about how you’re making a strong showing for number two.

It sounds like it’s this whole ridiculous, negging nonsense or whatever the kids are calling it these days, but that’s not how it's intended. Today’s promoted guest is Chris Harris, who’s the Vice President of Global Field Engineering at Couchbase. Chris, thank you for joining me and I really hope I got it right and that Couchbase is a database company or that makes no sense whatsoever.

Chris: It’s great to be on the show, and thank you for the invitation. I’m looking forward to it. Yeah, we’re a database company. That’s exactly what we do.

Corey: I always find it interesting when companies start pivoting from a thing that they were and, “What do you do?” “We build databases.” [unintelligible 00:01:29] getting out of that space it’s, “What do you do?” “We’re a finance company.” And then there’s a period of time in which they start reframing what they do. It’s, “We’re a data platform.” Or, “We’re now a tech company.”

Really? Because I don’t get that sense in any meaningful perspective. Couchbase was founded as a database company. You went public last year—congratulations on that—and now you continue to say, “Yes, we’re a database company,” rather than an everything trying to eat the world all at the same time, mostly ineffectively, company. So, what kind of database are you folks?

Chris: So, if you look at the database world, you can see—I’ve been in the space for quite some time now, a good few years, and I’ve had the privilege, if you like, of being at other database companies, been in the analytics space, and I’m here at Couchbase. But if you look at the history over the last—let’s just not go back all the way that far, but let’s go back to, like, ten years ago, everybody was building their applications on traditional relational databases. And what you saw is that the Oracle and MySQL, as traditional databases of the world. And then… probably at the time, we realized that with, talking ten years ago where we had this demand for high throughput of data, next generation of applications will be built, and then people realized the traditional database architectures weren’t going to cut it, if you like. And it spawned this industry.

You know, a big NoSQL market was created. And you have document databases, and then you have graph databases, and then you have analytics databases, and you have search databases, and then you have every sort of database you could possibly think of type database that’s out there in the world.

Corey: You have so many kinds you need to keep track of it all inside of the database.

Chris: That’s what you have to do, right? [laugh]. But the interesting thing is it became different types of database. And even see this in many of the code providers today, right, that you have multiple different types of databases no matter what you’re trying to do, right? So, we kind of went—Couchbase kind of took a step back and went, okay, we were originally a cache, right, this is where we came from, and then kind of built that into a document database, and then kind of went to the market and went hold on here, rather than it being let’s call it a noSQL versus SQL discussion, why can’t it just be a database, right?

Why can’t you have a SQLite interface on top of a modern architecture? Why can you do that, right? Why can’t you have the flexibility and architectural [unintelligible 00:04:16] of a JSON-based database with the interface of—with SQL, and then analytics built on top of that, right? So, why can’t you have the power of SQL on the next generation architecture? So, that’s kind of where we fit in the world.

Corey: When we talk about origin stories and where things come from, well, let’s start with you. I guess the impolite version of the question is, “Why on earth would you be in a space like this for so long?” But you’ve been on a lot of interesting places doing somewhat similar things. You were at Cloudera, you were at Hortonworks until you apparently heard a who or whatnot, you were at MongoDB, you were at VMware, you were at Red Hat. And that’s going reverse chronologically, but it’s clear that you’re very focused on a particular expression of a particular problem. Why are you the way that you are? Only pretend that’s a polite question.

Chris: “Why am I the way that I am?” Well, first of all, I love technology, right? That’s the key. And I think many of us in the industry would definitely say that, right? I started off in core engineering, building—I know some people today wouldn’t probably remember this, but when you had Chip and Pin where your credit card and you have to type it in and put in a pin number, that was created originally in the UK, and then went off and built e-commerce websites for retailers.

Well, that then turned into—was a common theme that I kept seeing is that lot of the technology that we’re using was open-source technology. And that kind of got me into the open-source movement, if you like, and I was lucky enough to then join Red Hat when they built middleware frameworks, so got into that space there. And then did a lot of innovation in the middleware space. Went to SpringSource and we did some great work there in the Java Development Framework space. But what became interesting is that—you still see it today—like, in this innovation happening in that middleware space and there’s some great innovation happening, right?

There’s all this stuff with Lambda and serverless architecture that’s out there, but they always came back to, we’ve got the database, this thing that is in the architecture if it goes down, you’re stuffed, right? This is where the core value of your company is sitting. So, then that got me interested to see what innovation was happening in this space. And as I say, I got into this field in the early stages of NoSQL, where there was that spawn of new database technologies being created. And then from there, it was like, “Okay, let’s get into what was happening in the analytic space.”

Again, I’m still in the Hortonworks, and Cloudera space, that’s all open-source. But it came down to this is different types of databases that were required different types of skills. And then I started talking to the team here, who was like, “How can we take as great innovation and leverage the skills I already have?” And I thought that was an interesting point.

Corey: In the interest of full disclosure, I tend to take the exact inverse approach to the way that you did. When I was going through the worlds of systems administrator, than rebadged as DevOps, or SRE, or systems engineer, production engineer, whatever we’re calling ourselves this week, I was always focused primarily on stateless things like web servers, or whatnot because it turns out that—this should be no surprise to longtime listeners of this show—but I’m really bad with computers. And most other things, too; I just brute force my way through it. And that’s hilarious when you keep taking down web servers you can push a button and recreate. When you do that to a database or anything that’s stateful, it leaves a mark.

And if you do it the wrong way, just well enough, you might not have a company anymore, so your DR plan starts to look a lot more like updating your resume. So, I always tried to shy away from things that played to my specific weaknesses that would, you know, follow me around like a stink. You, on the other hand, apparently sound—how to frame it—you know, good at things, and in a way that I never was. So you’re—ah, you see a problem, you’re running towards it trying to help fix it; I’m trying to how do I keep myself away from making the problem worse is my first approach. It seems like you have definitely been focused on not just data themselves—I mean at some level, [if it was a 00:08:55] pure data problem, it feels like we’d be talking a lot more about storage, but rather how to wind up organizing that data, how to wind up presenting that data, and the relationship that data has to other things
that are going on. I’m not speaking in the sense of a traditional relational database, necessarily, but the idea of how that data
empowers businesses and enables them to do different things. Is that directionally a fair synopsis of how you see it?

Chris: I think the [unintelligible 00:09:21] thing is what I would agree with. What makes it really interesting to me is what we enable people to do with that data, and being able to build, kind of, really fascinating innovation applications that are affecting their underlying businesses, right, from it could be health care, it could be airlines, financial services, some really high, interesting use cases that people are doing that are leveraging the database to be able to drive that level of innovation. Because it’s very difficult; I can build some sophisticated application, but if I can’t get the performance out of my database, I have a pretty poor experience to my users in today’s world. Because, fortunately or unfortunately, people aren’t very patient, right? If you have a website that doesn’t return very quickly, a customer’s gone like minutes ago. You literally got to instantly respond to someone. That’s a challenging problem.

Corey: It absolutely is. Something that I found as I’ve talked to a bunch of different companies operating in different ways is the requirements on data stores are generally very different depending upon primarily latency and performance. There’s only so long people that are going to watch the spinning circle of doom on a website spin before they realize they’re going to go somewhere that has its act together. Conversely, for a lot of business intelligence and analytics queries, there are an awful lot of stories where the thing that people care about is that we actually have to have the results of this query by noon on Thursday. And there are very different use cases for that, and some companies seem to be focused very much on, “We’re going to solve both of those use cases extremes and everything in between with the same product offering,” and others tend to say, “Okay, this is the area of the market we’re going to focus on.”

You could also say that this is an expression of the larger industry question of do I want, more or less a one-size-fits-most database that’s general-purpose, or do I want very specific purpose-built databases based upon the use case and the problem? Where do you find yourself on that spectrum?

Chris: I find myself on that spectrum is that there’s—if you want to describe it at a high level and we can break it down, there’s operational-type databases, where I’d say Couchbase fits where you’re talking about, I’ve just built an application; I’m talking to the live user, right, this is what I care about, and when I’m talking about speed and performance here, I’m talking about something that returned within milliseconds of response time. I’m playing an online game, or I’m doing online betting on a sports game. That has to be pretty much instant, right? If we’re playing multiplayer games and you’re doing something, then I want to be able to see what you’re doing straight away, right? People don’t expect it to lay there.

If you’re looking at streaming—people do this with Couchbase—streaming the Olympic Games or Super Bowl in the US, and you want to be able to be there, that whole profile management of that user has to be instant, have that stream to you has to be instant. People use telephone calls and use Couchbase to do, behind the scenes, profile management, right, so they know who you are who’s making that call. That’s an operational database problem. That’s not a traditional analytical problem, right? So, there’s a whole other space in the database world for analytics, right, which is bringing all the data together into one place, and I’ll help you do data science, AI, machine learning, be able to crunch and compute large volumes of data. If I get back to you, rather than a week in an hour, that’s great, but that’s not operational. That’s analytical.

Corey: In data center environments, it’s an argument to be made for going in a bunch of different directions; we’re going to use a bunch of different data stores to store all these things. Because, generally speaking, the marginal cost of moving data from one of your data storage systems to another one of your data storage systems, one rack row over is fairly small, whereas in cloud, effectively, there are no real capacity constraints anymore until you can get the bill, but that’s the entire problem where a lot of the transfer for these things is metered per gigabyte. So, there’s a increased desire on a lot of architectural pressures, to wind up making sure that where the data lives, it stays. And whatever it is that you do with that data, it should be able to operate on that data in a way that fits your performance characteristic requirements in the place that it currently is. And on the one hand, I can definitely see that driving a lot of decisions people have made.

The counterargument is that it feels a little weird when the cost constraints of how the cloud providers—mostly you, AWS—have decided to build these things out. And that, in turn, is shaping your entire approach to not just your architecture, but your systems design of how data winds up working its way through your lifecycle. It’s frustrating, on some level, especially given that they themselves offer something like 15 distinct managed databases offerings but more announced all the time. It becomes very difficult not only to disambiguate between all of them but to afford moving data from one to the next.

Chris: The affordability is an interesting discussion, right, because you can look at it from a billing perspective and go absolutely, there’s a challenge associated to that. Then is a question of where is my data because it’s spread across all these different services; that’s another challenge. And then you have the challenge of, okay, the cost associated to having developers build applications against all these different types of services because they all require different APIs and different ways of programming. So that’s, there’s a cost associated everywhere.

Corey: Oh, by far and away, the most expensive part of your AWS, or any cloud spend, is not the infrastructure itself; it’s the payroll expense associated with the people working on it. People always cost more than infrastructure. If not, something very strange is going on.

Chris: But then you look at it, and you go, okay, if that’s the case, I kind of use the analogy, right, that it’s like a car, where everyone is talking these days about the electric car [that’s going 00:16:05] on that path, right? Now, I should be able—if I was getting an electric car—think of it now, I actually have one—that I can get in the car and I can drive it like any other car. I know what a steering wheel is, I know where my pedals work, it looks and feels like a normal car. But architecturally it’s fundamentally different how it operates. So, why can you apply that same thing, that same analogy to a database, right?

So, why can’t I have the ability from an operational perspective, [unintelligible 00:16:42] talk about operational databases, not necessarily, I don’t know, full-blown analytical databases, but operationally being able to say I can store the data in an enterprise database; I can use that to leverage my SQL skills like I have before, and also use it to have a document store under operational analytics, to eventing, to full-stack search, key things that people want to do operationally, but keeping the data together in one database, like an iPhone. I want a database to have these capabilities; I don’t want to have all these different types of devices that are everywhere. I want, you know, my iPhone to be able to go to have the capabilities that I’m using. Or my car, to feel like I’m driving a car; doesn’t matter if the underlying architecture of the engine changed. That’s great, I want the benefit, but I want to be able to drive it in the same way that I’ve driven any other car out there. And that’s kind of trying to solve multiple problems that because you’re trying to solve the issue of skills.

Corey: It’s one of the hard challenges out there, and I think your car analogy can even be extended a bit further because in the early days of the automobile, you were more or less taking some significant risk by driving a car if you weren’t also mechanically inclined and to fix it yourself. And in time, we’ve sort of seen that continue to evolve where they mostly work, and now they work really reliably. And then you take it even a step beyond that, and all right now I’m just going to pay a car service so someone else has to deal with the car and a driver, and I don’t have to deal with any of that aspect. And it feels like there are certain parallels, similar to that, toward the end of last year, 2021, you folks, more or less moved away from you can have it in any color you want, as long as you run it yourself—more or less—into offering a fully-managed database-as-a-service cloud option called Capella, which, on the ads for this show, I periodically sing because if you didn’t want me to do that, you would not have named it Capella. Now, what was it that inspired you folks to say, “Hm, we could actually offer this as a managed service ourselves?”

It’s definitely a direction a lot of companies have gone in, but usually, they have to wait to be forced into it by—let’s be serious for a second here—Amazon launching the Amazon Basics version of whatever it is themselves and, “Okay, well, they validated our market for us. Let’s explore it.”

Chris: If you look at that, you go Couchbase has been around for a good few years now selling, as you point out, high-performance databases to large-scale enterprises, on real mission-critical, people call it tier-zero type applications, high-performance applications. And these are some of the most fascinating, most innovative type of applications that I’ve been involved with through my career. Now, how can we take that capability, provide it to the mass market if you’d like, to be able to give it to people that don’t need to have a large number of people out there managing their own infrastructure, being able to understand how to finely tune that underlying infrastructure to get the level of performance that you need from high-performance databases. Now, there are use cases for doing that, so it’s not one or the other. It’s not that you have to go all-in.

There are particular companies out there that, for the economics reasons, for the use case reasons that are running today on-premise, and there’s a rational reason for why they do that, right? But for a lot of people out there, whether they’re leveraging the cloud, there’s an opportunity here to take the power of the database, allow us to then manage it for people, take away that complexity of it, but being able to give them the power so they can leverage their skills, take advantage of Couchbase far easier than ever have been able to in the past. It’s opened up a bigger market for us, to summarize your question.

Corey: This episode is sponsored by our friends at Oracle Cloud. Counting the pennies, but still dreaming of deploying apps instead
of “Hello, World” demos? Allow me to introduce you to Oracle’s Always Free tier. It provides over 20 free services and infrastructure, networking, databases, observability, management, and security. And—let me be clear here—it’s actually free. There’s no surprise billing until you intentionally and proactively upgrade your account. This means you can provision a virtual machine instance or spin up an autonomous database that manages itself, all while gaining the networking, load balancing, and storage resources that somehow never quite make it into most free tiers needed to support the application that you want to build. With Always Free, you can do things like run small-scale applications or do proof-of-concept testing without spending a dime. You know that I always like to put asterisks next to the word free? This is actually free, no asterisk. Start now. Visit snark.cloud/oci-free that’s snark.cloud/oci-free.

Corey: One way that I tend to evaluate where a given vendor sees themselves—and it’s sort of an odd thing to do, but given that I do fix AWS bills for a living, it probably makes sense—I wind up pulling up the website, I ignore the baseline stuff of the, “This is what Gartner says,” and here’s a giant series of scrolls. I just go for the hamburger menu and I look for, “All right, where’s the pricing information?” Because pricing speaks a lot. And there are two things I generally try to find. One is, is there a free trial that I can basically click and get started working with?

Because invariably, I’m trying to beat my head off of a problem at two in the morning, and if it’s, “Oh, talk to a salesperson,” well as a hobbyist, or as an engineer who does not have signing authority for things, but it’s talk to sales, I realize, “Oh, yeah. One, I probably can’t afford it. Two, it’s going to be a week or so before I can actually make progress on this, and I’m hoping to get something up by sunrise, and it’s probably not for me.” Conversely, the enterprise tier should always have a, “Call for details,” because that is a signal to large enterprise procurement departments and buyers and the rest were it’s, “Oh, we will never accept default terms. We always want them customized. We also don’t believe in signing any contract without at least two commas in it.”

Great. So, being able to speak to both ends of the market is one of those critical things that you folks absolutely nail that. What I like is the fact that if someone has a problem that they’re experimenting with at two in the morning, they can get started with your database-as-a-service platform—Capella; or however you want to sing it—and they don’t need to wind up talking to you folks directly, first. There’s no long-term commitments, there’s no [unintelligible 00:22:39] of the infrastructure themselves. There’s no getting hounded for the rest of their days over making a purchase for something that didn’t pan out.

To me, that’s always been the real innovation and breakthrough of cloud is that I can spend a few hours some evening kicking around an idea, and if it doesn’t work, I can turn it off and spend 17 cents on the process, whereas if it does work, I can keep scaling up without at some point having to replace all of the Raspberry Pi’s and popsicle sticks, I build things with real enterprise-grade stuff. There’s a real accessibility and democratization that is entered into it. So, I’m always excited when I see companies that are embracing that model. Because, yeah, I’m a grumpy old sysadmin because it’s not like there’s a second kind of sysadmin, but—and I have a particular exposure and experience level with these things that I can’t expect modern developers to work on. They have an idea, they want to launch something, and they just need a database to throw things against and put data into, and ideally get it back again when they query later. And that empowers them to move forward.

They’re not in this because they really want to run virtual machines themselves and get those set up and secured and patched and hardened, and then install the software on top of it, and, “Why is it not working? Oh, security groups, how you vex me again. I’ll just open you to the entire world,” and so on. And we know where that path leads. So, it’s nice to see that there is an accessible option there.

Conversely, if you come at this with an approach if we are only available in our hosted cloud environment, well now those big enterprise companies that have, you know, compliance concerns are going to have some thoughts for you, none of them particularly pleasant in some cases. So, I like the fact that you’re able to expand your offering to encompass different user personas without also, I don’t know, turning what has historically been a database into now it’s an LDAP server, and trying to eat the world, piece by piece, component by component.

Chris: It’s interesting that you say that because I think there’s a number of things that you’re touching on that were to me, if you look at us as a company in particularly this space, there’s a lot of focus around the community and the open-source community. And I think there’s an element of how do you make it accessible to people as a community as a whole? And then you kind of go down the path of, “Okay, let’s allow people,”—as a developer, let’s think of it this way, right, the ultimate thing they want to do, and you touched upon it there, is they want to build an application. They get passionate about building the application or maybe even in the weekend, and they got this funky idea that they’re going to literally knock some code out.

And I remember my fond memories of being an application engineer of being able to sit down for hours just been able to put my ideas into code and watch it execute. The last thing that I want to do is get to the point where I get the database and go, oh, here we go. This is going to take me a bunch of hours, now, and I’m going to set it all up and do other stuff. And I almost literally want to be able to click a few buttons—

Corey: You know what I want to do tonight? Feel really dumb as I tackle a problem I don’t fully understand. Gr—I’d love smacking into walls that point out my own ignorance. It’s discouraging as hell. I’m right there with you.

Chris: Yeah, you don’t want to do that, right? So, you almost want to make like the database disappear for people, right? You want to be able to just say, like, “Here’s your command. Off you go. Bring the data back. Bring it back in full. Allow it to scale.”

Because you want that developer to have that experience of not breaking their flow. And what do you want them to be able to be so excited about the application and innovation that they’ve built, that they want to go and show that teammates? They want to say, “Look at the great thing I built over the weekend. Look at this, this is amazing.” Right?

And then be able to get all the teammates pretty excited about what they built in a way in which they can try it out really easily, right? They can take this little thing that they built into the database, click some buttons, and off we go, right? And now your development team is super excited about some of the great innovation that you have. But you also have to have the reverse. You have to have the architecturally sound, so then when you get to the architect, if you like, who is looking at the bigger picture of what’s the future going to look like? Is it the right technology? Is this something that we can bring into the organization? And you know, this is a cool bit of application you just built me, but you know, is this realistic that I can deploy this thing?

And this is where you start going back into it still has to have high performance, the security has to be there, the scalability has to be there so that I can potentially—I can start small and grow this thing horizontally as I see the requirements coming. There are different set of requirements architecturally, so we’re looking at—you know, as a company, our key focus is how do you drive that developer community so that you give the people the freedom to build the next generation of applications in the simplest way [unintelligible 00:27:35], say with free trials, click some buttons, have the database up in minutes, but also then being able to have that capability in the underlying database to take it to the architect. That’s what our core focus is every day.

Corey: I agree with everything that you’re saying. You’re making an awful lot of great points, but for me, the proof in the pudding is the second thing that I tend to look at on your website after the pricing page, and that is your list of customers. Because it’s always interesting when someone talks about how they’re revolutionizing everything, and this is the way to go, and everyone who’s anyone is doing these things. And then you look at their customer page and either they don’t have one, which is telling, or the customers on that page are terrifying in that, “Wow, that sounds like a whole bunch of fly-by-night startups whose primary industry is scamming people.”

You have a bunch of household blue-chip names as well as a bunch of newer companies that are very clearly not what people think of as legacy—you know, that condescending engineering term that means it makes money. It’s across the board, it is broad-spectrum, and it is companies that absolutely know exactly what it is that they’re doing when it comes to these things. That to me is far more convincing than almost anything else that can be said because it’s—look, you can come on and talk to me about anything you want about your product, and I can dismiss it and, “Yeah, whatever. Great.” But when I start talking to customers, as I did prior to recording this episode, and seeing how they talk about you folks, that to me is what reaffirms that, okay, this is actually something that has legs and is solving real customer problems.

Because early stage, it’s, “We have this idea for this company we’re going to build that it’s going to be great.” “Awesome. Go talk to more customers.” That is a default, safe piece of advice generically you can give to anyone. And it’s easy to give and hard to take.

I’ve been saying this for years, and I still screwed it up and we started trying to launch a SaaS product here called DuckTools. Yeah, it turns out that we didn’t talk to enough customers first about what they’re actually trying to achieve, and we assumed we knew the answers. It’s an easy mistake to make. What I really appreciate is—about a Couchbase in particular—is not just the fact that you have all of these customer references, but the fact that each one talks about what the value to the business is not just in terms of, “Oh yes, now we can query data and there was no way for us to do that before.” Of course, people have found ways to do that since business started.

Instead, it’s much more about this is how it made it more efficient, more optimal, how it unlocked possibilities and capabilities for us. That alone tells me that there definitely is significant value that you’re delivering to customers. In my own business, whenever I think I’ve seen it all, I have to do is talk to one more customer and learn something new. What have you seen in recent memory, from a customer, that surprised you about how they’re using Couchbase?

Chris: You look at that, and you can see—I could probably talk for hours on different types of customers, but it’s the ones that you can literally see in your life and you can reflect to, right? So, if you taken one of the biggest airlines that are out there today, they’re completely changing, kind of, the whole experience. And our whole experience of and how do I get feedback? Because Couchbase’s customers, [unintelligible 00:31:01] customer, right, is what they’re thinking about, right? They’re an airline.

So, these passengers; fine. But how many times have you got on a plane, and you see all these people, literally, there’s obviously the passengers, and then there’s the cabin crew, and then there’s the people on the ground, and then there’s the pilot, and for the sake of the discussion, the staff that are there are literally passing paper back and forth to each other. And surely there a better way to do this. And for someone who likes to solve complex technical problems, you go, “Wow, this is going to be a bit of a challenge.” Because if you want to collect feedback from an aeroplane in the air, [laugh] right, and you want to connect that to the ground data that people are having in terms of maintenance data, you want to do that across the world, in multiple different time zones, that’s pretty tricky problem to try to go solve, right?

So therefore, how do you get a database that is able to work remotely and on what people would call the edge; let’s just call it in this case in a device that’s literally a cabin crew member is carrying around with them that’s not connected because there is no connection because I’m in the middle of the air. But I want to pair it with the other cabin crew members that are around, right, in flight, and then when I land, I want to sync that data backup to the maintenance people. So, you need a database that’s able to operate on a device with no connection, and then being able to synchronize backup to a cloud database that is then collecting data from all the other flights around the world.

Corey: Synchronization sounds super easy until you actually try and do it, and then, “Oh, wow.” It’s like, you could cut to pieces by the edge cases.

Chris: And then people go, “Well, there’s no problem. There’s internet everywhere these days.” Yeah, sure there is. [laugh]. You get disconnected all of the time.

Corey: Not to name names. This is very evocative, an earlier episode of this show I had with Tyler Slove, who’s a senior manager over at United Airlines, about specifically how they’re approaching a lot of their own reimagining and the rest. It’s a fascinating use case, and as someone who’s a bit of a travel geek himself—you know, in the before times—that’s always an area of intense interest because it’s… I’m sorry, I’m still a little boy at heart; it’s magic to me. You get on a plane, you go somewhere else, close the doors, it opens it up, and you’re on the other side of the world. And now there’s internet on it? Oh, my God, who would have imagined such a thing?

Chris: Uh-huh. But that’s changing the experience for people. It’s just really fascinating.

Corey: Completely. And it’s empowering and unlocking that experience you’re talking about of being able to sync between the crews, about handling all this stuff behind the scenes. Everyone loves to complain about airlines because no one knows really how to run the massive logistical part of an airline. But the WiFi was a little bit slow or the food was cold; well, that’s something I know how to complain about Twitter.

Chris: [laugh].

Corey: It becomes this idea of almost a bikeshed problem expression, where it’s, “Oh, yeah. I’m just going to complain about things I can wrap my head around.” Yeah.

Chris: I was talking to somebody recently, and they were—swapping topics a little bit—and they were like, so—they were talking about innovation on some new web application that they built. And I literally have to explain them, and I said, “Well, if you think of it, the underlying whole technology stack that’s behind this for high-scale e-commerce, it’s sophisticated, right, because people will literally walk away from a page, an application, a mobile app, if they don’t get an instant response time. And that request has to literally travel, physically, quite a fair amount of distance, talk to multiple different types of technology, answer to that question, then come back to you instantly.” The sheer amount of technology that’s involved here of moving that data around is a complicated architectural problem to fix. A database only plays a small part of that. You can’t be the slowest player in the party.

Corey: No. And that is always the challenge is that when you’re looking at different use cases, there’s always a constraint, and how that constraint winds up manifesting in different ways, if it’s not the thing that’s slowing things down, it’s also not where the attention goes. If you have a single thing like, the database for example, slowing things down, everyone cares about improving databases, people focusing on, “Well, we’re going to improve the JavaScript load time on the website,” that’s not the problem. Find the bottleneck and focus on it. And although I’m generally a fan of picking a database and using that as a general-purpose thing until it makes sense not too—much like I am cloud providers—

[audio break 00:35:54]

Corey: —journey personally, where’s the best place to find you?

Chris: Clearly, if you want to find more about Couchbase, you can obviously go to couchbase.com. You kindly pointed out you can go and look at the trial for Capella and try out the tech. You’re more than welcome to do that as a free trial.

If you want to contact me particularly, you find me on LinkedIn; I’m Chris Harris at Couchbase. You’ll find me [unintelligible 00:36:26] with Chris Harris in general and probably find lots of them. In the UK, Chris Harris is a famous racing driver. That’s not me; it’s someone else. So, find me on LinkedIn, I’m sure it won’t be that difficult to find what you find. Or you can find me on Twitter.

Corey: And we will of course, but links to all of that into the [show notes 00:36:43]. I really want to thank you for being so generous with your time today. It’s always appreciated to talk to people who actually know what they’re doing.

Chris: You’re more than welcome. It’s been great to be on the show. Thanks, Corey.

Corey: Chris Harris, Vice President of Global Field Engineering at Couchbase. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry comment. I’m going to wind up using all of those angry comments, at one point, as a database.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Tom

Tom enjoys being a bridge between people and technology. When he's not thinking about ways to make enterprise demos less boring, Tom enjoys spending time with his wife and dogs, reading, and gaming with friends.

Links Referenced:

  • LaunchDarkly: https://launchdarkly.com
  • Heidi Waterhouse Twitter: https://twitter.com/wiredferret

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Couchbase Capella Database-as-a-Service is flexible, full-featured and fully managed with built in access via key-value, SQL, and full-text search. Flexible JSON documents aligned to your applications and workloads. Build faster with blazing fast in-memory performance and automated replication and scaling while reducing cost. Capella has the best price performance of any fully managed document database. Visit couchbase.com/screaminginthecloud to try Capella today for free and be up and running in three minutes with no credit card required. Couchbase Capella: make your data sing.

Corey: This episode is sponsored by our friends at Revelo. Revelo is the Spanish word of the day, and its spelled R-E-V-E-L-O. It means “I reveal.” Now, have you tried to hire an engineer lately? I assure you it is significantly harder than it sounds. One of the things that Revelo has recognized is something I’ve been talking about for a while, specifically that while talent is evenly distributed, opportunity is absolutely not. They’re exposing a new talent pool to, basically, those of us without a presence in Latin America via their platform. It’s the largest tech talent marketplace in Latin America with over a million engineers in their network, which includes—but isn’t limited to—talent in Mexico, Costa Rica, Brazil, and Argentina. Now, not only do they wind up spreading all of their talent on English ability, as well as you know, their engineering skills, but they go significantly beyond that. Some of the folks on their platform are hands down the most talented engineers that I’ve ever spoken to. Let’s also not forget that Latin America has high time zone overlap with what we have here in the United States, so you can hire full-time remote engineers who share most of the workday as your team. It’s an end-to-end talent service, so you can find and hire engineers in Central and South America without having to worry about, frankly, the colossal pain of cross-border payroll and benefits and compliance because Revelo handles all of it. If you’re hiring engineers, check out revelo.io/screaming to get 20% off your first three months. That’s R-E-V-E-L-O dot I-O slash screaming.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Today’s promoted episode is brought to us by our friends at LaunchDarkly. And it’s always interesting when there’s a promoted guest episode because they generally tend to send someone who has a story to tell in different ways.

Sometimes they send me customers of theirs. Other times they send me executives. And for this episode, they have sent me Tom Totenberg, who’s a senior solutions engineer at LaunchDarkly. Tom, thank you for drawing the short straw. It’s appreciated.

Tom: [laugh]. Anytime. Thank you so much for having me, Corey.

Corey: So, you’re a senior solutions engineer, which in many different companies is interpreted differently, but one of the recurring themes tends to pop up is often that is a different way of saying sales engineer because if you say sales, everyone hisses and recoils when you enter the conversation. Is that your experience or do you see your role radically differently?

Tom: Well, I used to be one of those people who did recoil when I heard the word sales. I was raised in a family where you didn’t talk about finances, you know? That’s considered to be faux pas, and when you hear the word sales, you immediately think of a car lot. But what I came to realize is that, especially when we talk about cloud software or any sort of community where you start to run into the same people at conferences over and over and over again, turns out the good salespeople are the ones who actually try to form relationships and try to solve problems. And I realized that oh, I like to work with those people. It’s pretty exciting. It’s nice to be aspirational about what people can do and bring in the technical chops to see if you can actually make it happen. So, that’s where I fit in.

Corey: The way that I’ve always approached it has been rather different. Because before I got into tech, I worked in sales a bunch of times and coming up from the—I guess, clawing your way up doing telesales was a polite way of describing—back in the days before there were strong regulations against it, calling people at dinner to sell them credit cards. And what’s worse is I was surprisingly effective at it for a kid who, like, you grew up in a family where we didn’t talk about money. And it’s easy to judge an industry by its worst examples. Another one of these would be recruiting, for example.

When everyone talks about how terrible third-party recruiters are because they’re referring to the ridiculous spray-and-pray model of just blasting out emails to everything that hold still long enough that meets a keyword. And yeah, I’ve also met some recruiters that are transformative as far as the conversations you have with them go. But some of that with sales. It’s, “Oh, well, you can’t be any fun to talk to because I had a really bad experience buying a used car once and my credit was in the toilet.”

Tom: Yeah, exactly. And you know, I have a similar experience with recruiters coming to LaunchDarkly. So, not even talking about the product; I was a skeptic, I was happy where I was, but then as I started talking to more and more people here, I’m assuming you’ve read the book Accelerate; you probably had a hand in influencing part of it.

Corey: I can neither confirm nor deny because stealing glory is something I only do very intentionally.

Tom: Oh okay, excellent. Well, I will intentionally let you have some of that glory for you then. But as I was reading that book, it reminded me again of part of why I joined LaunchDarkly. I was a skeptic, and they convinced me through everyone that I talked to just what a nice place it is, and the great culture, it’s safe to fail, it’s safe to try stuff and build stuff. And then if it fails, that’s okay. This is the place where that can happen, and we want to be able to continue to grow and try something new.

That’s again, getting back to the solutions engineer, sales engineer part of it, how can we effectively convey this message and teach people about what it is that we do—LaunchDarkly or not—in a way that makes them excited to see the possibilities of it? So yeah, it’s really great when you get to work with those type of people, and it absolutely shouldn’t be influenced by the worst of them. Sometimes you need to find the right ones to give you a chance and get in the door to start having those conversations so you can make good decisions on your own, not just try to buy whatever someone’s—whatever their initiative is or whatever their priority is, right?

Corey: Once upon a time when I first discovered LaunchDarkly, it was pretty easy to describe what you folks did. Feature flags. For longtime listeners of the show, and I mean very longtime listeners of the show, your colleague Heidi Waterhouse was guest number one. So, I’ve been talking to you folks about a variety of different things in a variety of different ways. But yeah, “LaunchDarkly. Oh, you do feature flags.”

And over time that message has changed somewhat into something I have a little bit of difficulty to be perfectly honest with you in pinning down. At the moment we’re recording this, if I pull up launchdarkly.com, it says, “Fundamentally change how you deliver software. Innovate faster, deploy fearlessly, and make each release a masterpiece.”

And I look at the last release I pushed out, which wound up basically fixing a couple of typos there, and it’s like, “Well, shit. Is it going to make me sign my work because I’m kind of embarrassed by a lot of it.” So, it’s aspirational, I get it, but it also somehow [occludes 00:05:32] a little bit of meaning. What is it you’d say it is you do here.

Tom: Oh, Office Space. Wonderful. Good reference. And also, to take about 30 seconds back, Heidi Waterhouse, what a wonderful human. wiredferret on Twitter. Please, please go look her up. She’s got just always such wonderful things to say. So—

Corey: If you don’t like Heidi Waterhouse, it is a near certainty it is because you’ve not yet met her. She’s wonderful.

Tom: Exactly. Yes, she is. So, what is it we’d say we do here? Well, when people think about feature flags—or at this point now, ‘feature management,’ which is a broader scope—that’s the term that we’re using now, it’s really talking about that last bit of software delivery, the last mile, the last leg, whatever your—you know, when you’re pushing the button, and it’s going to production. So, you know, a feature flag, if you ask someone five or ten years ago, they might say, oh, it’s a fancy if statement controlled by a config file or controlled by a database.

But with a sort of modern architecture, with global delivery, instant response time or fraction of a second response time, it’s a lot more fundamental than that. That’s why the word fundamental is there: Because it comes down to psychological safety. It comes down to feeling good about your life every day. So, whether it is that you’re fixing a couple typos, or if you’re radically changing some backend functionality, and trying out some new sort of search algorithm, a new API route that you’re not sure if it’s going to work at scale, honestly, you shouldn’t have to stay up at night, you shouldn’t have to think about deploying on a weekend because you should be able to deploy half-baked code to production safely, you should be able to do all of that. And that’s honestly what we’re all about.

Now, there’s some extra elements to it: Feedback loops, experimentation, metrics to make sure that your releases are doing well and doing what you anticipated that they would do, but really, that’s what it comes down to is just feeling good about your work and making sure that if there is a fire, it’s a small fire, and the entire audience isn’t going to get part of the splash zone, right? We’re making it just a little safer. Does that answer your question? Is that what you’re getting at? Or am I still just speaking in the lingo?

Corey: That gets it a lot closer. One of the breakthrough moments—of course I picked it up from one of Heidi’s talks—is feature flag seems like a front end developer thing, yadda, yadda, yadda. And she said historically, yeah, in some ways, in some cases, that’s how it started. But think about it this way. Think about separating out configuration from your deploy process. And what would that mean? What would that entail?

And I look at my current things that I have put out there, and there is no staging environment, my feature branches main, and what would that change? In my case, basically nothing. But that’s okay. Because I’m an irresponsible lunatic who should not be allowed near anything expensive, which is why I’m better at stateless things because I know better than to take my aura near things like databases.

Tom: Yeah. So, I don’t know how old you are Corey. But back—

Corey: I’m in my mid-30s, which—

Tom: Hey—

Corey: —enrages my spouse who’s slightly older. Because I’m turning 40 in July, but it’s like, during the pandemic, as it has for many of us, the middle has expanded.

Tom: There you go. Right. Exa—[laugh] exactly. Can neither confirm nor deny. You can only see me from about the mid-torso up, so, you know, you’re not going to see whether I’ve expanded.

But when we were in school doing group projects, we didn’t have Google Docs. We couldn’t see what other people were working on. You’d say, “Hey, we’ve got to write this paper. Corey, you take the first section, I’ll take the second section, and we’ll go and write and we’ll try to squish it back together afterward.” And it’s always a huge pain in the ass, right? It’s terrible. Nobody likes group projects.

And so the old method of Gitflow, where we’re creating these feature branches and trying to squish them back later, and you work on that, and you work on this thing, and we can’t see what each other are doing, it all comes down to context switching. It is time away from work that you care about, time away from exciting or productive work that you actually get to see what you’re doing and put it into production, try it out. Nobody wants to deal with all the extra administrative overhead. And so yeah, for you, when you’ve got your own trunk-based development—you know, it’s all just main—that’s okay. When we’re talking about teams of 40, 50, 100, 1000 suddenly becomes a really big deal if you were to start to split off and get away from trunk-based development because there’s so much extra work that goes into trying to squish all that work back together, right? So, nobody wants to do all the extra stuff that surrounds getting software out there.

Corey: It’s toil. It feels consistently like it is never standardized so you always have to wind up rolling your own CI/CD thing for whatever it is. And forget between jobs; between different repositories and building things out, it’s, “Oh, great. I get to reinvent the wheel some more.” It’s frustrating.

Tom: [laugh]. It’s either that or find somebody else’s wheel that they put together and see if you can figure out where all those spokes lead off to. “Is this secure? I don’t know.”

Corey: How much stuff do you have running in your personal stuff that has more or less been copied around for a decade or so? During the pandemic, I finally decided, all right, you know what I’m doing? That’s right, being productive. We should fix that. I’m going to go ahead and redo my shell config—my zshrc—from scratch because, you know, 15 years of technical debt later, a lot of the things I used to really need it to do don’t really apply anymore.

Let’s make it prettier, and let’s make it faster. And that was great and all, but just looking through it, it was almost like going back in time for weird shell aliases that I don’t need anymore. It’s, well, that was super handy when I ran a Ruby production environment, but I haven’t done that in seven years, and I haven’t been in this specific scenario that one existed for since 2011. So maybe, maybe I can turn that one off.

Tom: Yeah, maybe. Maybe we can get rid of that one. I mean, when’s the last time you ran npm install on something you were going to try out here and paid attention to the warnings that came up afterward? “Hey, this one’s deprecated. That one’s deprecated.” Well, let’s see if it works first, and then we’ll worry about that later.

Corey: Exactly. Security problems? Whatever. It’s a Lambda function. What do I care?

Tom: Yeah, it’s fine. [laugh]. Exactly. Yeah. So, a lot of this is hypothetical for someone in my position, too, because I didn’t ever get formal training as a software developer. I can copy and paste from Stack Overflow with the best of them and there’s all sorts of resources out there, but really the people that we’re talking to are the ones who actually live that day in, day out.

And so I try to step into their shoes and try to feel that pain. But it’s tough. Like, you have to be able to speak both languages and try to relate to people to see what are they actually running into, and is that something that we can help with? I don’t know.

Corey: The way that I tend to think about these things—and maybe it’s accurate, and maybe it’s not—it’s just, no one shows up hoping to do a terrible job at work today, but we are constrained by a whole bunch of things that are imposed upon us. In some of the more mature environments, some of that is processes there for damn good reasons. “Well, why can’t I just push everything I come up with to production?” “It’s because we’re a bank, genius. How about you think a little bit before you open your mouth?”

Other times, it’s because well, I have to go and fight with the CI/CD system, and I’m just going to go ahead and patch this one-line change into production. Better processes, better structure have made that a lot more… they’ve made it a lot easier to be able to do things the right way. But I would say we’re nowhere near its final form, yet. There’s so much yak-shaving that has to go into building out anything that it’s frustrating, on some level, just all of the stuff you have to do, just to get the scaffolding in place to write nonsense. I mean, back when they announced Lambda functions it was, “In the future, the only code you’ll write is business logic.”

Yeah, well, I use a crap-ton of Lambda here and it feels like most of the code I write is gluing all of the weird formats and interchanges together in different APIs. Not a lot of business logic in that; and awful lot of JSON finickiness.

Tom: Yeah, I’m with you. And especially at scale, I still have a hard time wrapping my mind around how all of that extra translation is possibly going to give the same sort of performance and same sort of long-term usability, as opposed to something that just natively speaks the same language end-to-end. So yeah, I agree, there’s still some evolution, some standardization that still needs to happen because otherwise we’re going to end up with a lot of cruft at various points in the code to, just like you said, translate and make sure we’re speaking the same language.

Getting back to process though, I spent a good chunk of my career working with companies that are, I would say, a little more conservative, and talking to things like automotive companies, or medical device manufacturers. Very security-conscious, compliant places. And so agile is a four-letter word for them, right, [laugh] where we’re going faster automatically means we’re being dangerous because what would the change control board say? And so there’s absolutely a mental shift that needs to happen on the business side. And developers are fighting this cultural battle, just to try to say, hey, it’s better if we can make small iterative changes, there is less risk if we can make small, more iterative changes, and convincing people who have never been exposed to software or know the ins and outs of what it takes to get something from my laptop to the cloud or production or you know, wherever, then that’s a battle that needs to be fought before you can even start thinking about the tooling. Living in the Midwest, there’s still a lot of people having that conversation.

Corey: So, you are clearly deep in the weeds of building and deploying things into production. You’re clearly deep into the world of explaining various solutions to different folks, and clearly you have the obvious background for this. You majored in music. Specifically, you got a master’s in it. So, other than the obvious parallel of you continue to sing for your supper, how do you get from there to here?

Tom: Luck and [laugh]. Natural curiosity. Corey, right now you are sitting on the desk that is also housing my PC gaming computer, right? I’ve been building computers just to play video games since I was a teenager. And that natural curiosity really came in handy because when I—like many people—realize that oh, no, the career choice that I made when I was 18 ended up being not the career choice that I wanted to pursue for the rest of my life, you have to be able to make a pivot, right, and start to apply some of the knowledge that you got towards some other industries.

So, like many folks who are now solutions engineers, there’s no degree for solutions engineering, you can’t go to school for it; everyone comes from somewhere else. And so in my case, that just happened to be music theory, which was all pedagogy and teaching and breaking down big complex pieces of music into one node at a time, doing analysis, figuring out what’s going on underneath the hood. And all of those are transferable skills that go over to software, right? You open up some giant wall of spaghetti code and you have to start following the path and breaking it down because every piece is easy one note at a time, every bit of code—in theory—is easy one line at a time, or one function at a time, one variable at a time. You can continue to break it down further and further, right?

So, it’s all just taking the transferable skills that you may not see how they get transferred, but then bringing them over to share your unique perspective, because of your background, to wherever it is you’re going. In my case, it was tech support, then training, and then solutions engineering.

Corey: There’s a lot to be said for blending different disciplines. I think that there was, uh, the naughts at least, and possibly into the teens, there was a bias for hiring people who look alike. And no, I’m not referring to the folks who are the white dudes you and I clearly present as but the people with a similar background of, “Oh, you went to these specific schools”—as long as they’re Stanford—“And you majored in a narrow list of things”—as long as they’re all computer science. And then you wind up going into the following type of role because this is the pedigree we expect and everything, soup to nuts, is aligned around that background and experience. Where you would find people who would be working in the industry for ten years, and they would bomb the interview because it turns out that most of us don’t spend our days implementing quicksort on whiteboards or doing other algorithmic-based problems.

We’re mostly pushing pixels around a screen hoping to make ourselves slightly happier than we were. Here we are. And that becomes a strange world; it becomes a really, really weird moment, and I don’t know what the answer is for fixing any of that.

Tom: Yeah, well, if you’re not already familiar with a quote, you should be, which is that—and I’m going to paraphrase here—but, “Diverse backgrounds lead to diversity in thought,” right? And that presents additional opportunities, additional angles to solve whatever problems you’re encountering. And so you’re right, you know, we shouldn’t be looking for people who have the specific background that we are looking for. How it’s described in Accelerate? Can you tell that I read it recently?

Which we should be looking for capabilities, right? Are you capable? Do you have the capacity to do the problem-solving, the logic? And of course, some education or experience to prove that, but are you the sort of person who will be able to tackle this challenge? It doesn’t matter, right, if you’ve handled that specific thing before because if you’ve handled that specific thing before, you’re probably going to implement it the same way, again, even if that’s not the appropriate solution, this time.

So, scrap that and say, let’s find the right people, let’s find people who can come up with creative solutions to the problems that we’re facing. Think about ways to approach it that haven’t been done before. Of course don’t throw out everything with the—you know, the bathwater out with a baby or whatever that is, but come in with some fresh perspectives and get it done.

Corey: I really wish that there was more of an acceptance for that. I think we’re getting there. I really do, but it takes time. And it does pay dividends. I mean, that’s something I want to talk to you about.

I love the sound of my own voice. I wouldn’t have two podcasts if I didn’t. The counterargument, though, is that there’s an awful lot of things that get, you know, challenging, especially when, unlike in a conference setting, it’s most people consider it rude to get up and walk out halfway through. When we’re talking and presenting information to people during a pandemic situation, well, that changes a lot. What do you do to retain people’s interest?

Tom: Sure. So, Covid really did a number on anyone who needs to present information or teach. I mean, just ask the millions of elementary, middle school, and high schoolers out there, even the college kids. Everyone who’s still getting their education suddenly had to switch to remote learning.

Same thing in the professional world. If you are doing trainings, if you’re doing implementation, if you’re doing demos, if you’re trying to convey information to a new audience, it is so easy to get distracted at the computer. I know this firsthand. I’m one of those people where if I’m sitting in an airport lobby and there’s a TV on my eyes are glued to that screen. That’s me. I have a hard time looking away.

And the same thing happens to anyone who’s on the receiving end of any sort of information sharing, right? You got Slack blowing you up, you’ve got email that’s pinging you, and that’s bound to be more interesting than whatever the person on the screen is saying. And so I felt that very acutely in my job. And there’s a couple of good strategies around it, right, which is, we need to be able to make things interactive. We shouldn’t be monologuing like I am doing to you right now, Corey.

We shouldn’t be [laugh] just going off on tangents that are completely irrelevant to whoever’s listening. And there’s ways to make it more interactive. I don’t know if you are familiar, or how much you’ve watched Twitch, but in my mind, the same sorts of techniques, the same sorts of interactivity that Twitch streamers are doing, we should absolutely be bringing that to the business world. If they can keep the attention of 12-year-olds for hours at a time, why can we not capture the attention of business professionals for an hour-long meeting, right? There’s all sorts of techniques and learnings that we can do there.

Corey: The problem I keep running into is, if you go stumbling down that pathway into the Twitch streaming model, I found it awkward the few experiments I’ve made with it because unless I have a whole presentation ready to go and I’m monologuing the whole time, the interactive part with the delay built in and a lot of ‘um’ and ‘ah’ and waiting and not really knowing how it’s going to play out and going seat of the pants, it gets a little challenging in some respects.

Tom: Yeah, that’s fair. Sometimes it can be challenging. It’s risky, but it’s also higher reward. Because if you are monologuing the entire time, who’s to say that halfway through the content that you are presenting is content that they want to actually hear, right? Obviously, we need to start from some sort of fundamental place and set the stage, say this is the agenda, but at some point, we need to get feedback—similar to software development—we need to know if the direction that we’re going is the direction they also want to go.

Otherwise, we start diverging at minute 10 and by minute 60, we have presented nothing at all that they actually want to see or want to learn about. So, it’s so critical to get that sort of feedback and be able to incorporate it in some way, right? Whether that way is something that you’re prepared to directly address. Or if it’s something that says, “Hey, we’re not on the same page. Let’s make sure this is actually a good use of time instead of [laugh] me pretending and listening to myself talk and not taking you into account.” That’s critical, right? And that is just as important, even if it feels worse in the moment.

Corey: This episode is sponsored in part by our friends at ChaosSearch. You could run Elasticsearch or Elastic Cloud—or OpenSearch as they’re calling it now—or a self-hosted ELK stack. But why? ChaosSearch gives you the same API you’ve come to know and tolerate, along with unlimited data retention and no data movement. Just throw your data into S3 and proceed from there as you would expect. This is great for IT operations folks, for app performance monitoring, cybersecurity. If you’re using Elasticsearch, consider not running Elasticsearch. They’re also available now in the AWS marketplace if you’d prefer not to go direct and have half of whatever you pay them count towards your EDB commitment. Discover what companies like Equifax, Armor Security, and Blackboard already have. To learn more, visit chaossearch.io and tell them I sent you just so you can see them facepalm, yet again.

Corey: From where I sit, one of the many, many, many problems confronting us is that there’s this belief that everyone is like we are. I think that’s something fundamental, where we all learn in different ways. I have never been, for example—this sounds heretical sitting here saying it, but why not—I’m not a big podcast person; I don’t listen to them very often, just because it’s such a different way of consuming information. I think there are strong accessibility reasons for there to be transcripts of podcasts. That’s why every 300-and-however-many-odd episodes that this one winds up being the sequence in, every single one of them has a transcript attached to it done by a human.

And there’s a reason for that. Not just the accessibility wins which are obvious, but the fact that I can absorb that information way more quickly if I need to review something, or consume that. And I assume other people are like me, they’re not. Other people prefer to listen to things than to read them, or to watch a video instead of listening, or to build something themselves, or to go through a formal curriculum in order to learn something. I mean, I’m sitting here with an eighth-grade education, myself. I take a different view to how I go about learning things.

And it works for me, but assuming that other people learn the same way that I do will be awesome for a small minority of people and disastrous for everyone else. So, maybe—just a thought here—we shouldn’t pattern society after what works for me.

Tom: Absolutely. There is a multiple intelligence theory out there, something they teach you when you’re going to be a teacher, which is that people learn in different ways. You don’t judge a fish by its ability to climb a tree. We all learn in different ways and getting back to what we were talking about presenting effectively, there needs to be multiple approaches to how those people can consume information. I know we’re not recording video, but for everyone listening to this, I am waving my hands all over the place because I am a highly visual learner, but you must be able to accept that other people are relying more on the auditory experience, other people need to be able to read that—like you said with the accessibility—or even get their hands on it and interact with it in some way.

Whether that is Ctrl-F-ing your way through the transcript—or Command-F I’m sorry, Mac users [laugh]; I am also on a Mac—but we need to make sure that the information is ready to be consumed in some way that allows people to be successful. It’s ridiculous to think that everyone is wired to be able to sit in front of a computer or in a little cubicle for eight hours a day, five days a week, and be able to retain concentration and productivity that entire time. Absolutely not. We should be recording everything, allowing people to come back and consume it in small chunks, consume it in different formats, consume it in the way that is most effective to them. And the onus for that is on the person presenting, it is not on the consumer.

Corey: I make it a point to make what I am doing accessible to the people I am trying to reach, not to me. And sometimes I’m slacking, for example, we’re not recording video today, so whenever it looks like I’m not paying attention to you and staring off to the side, like, oh, God, he’s boring. No. I have the same thing mirrored on both of my screens; I just prefer to look at the thing that is large and easy to read, rather than the teleprompter, which is a nine-inch screen that is about four feet in front of my face. It’s one of those easier for me type of things.

On video, it looks completely off, so I don’t do it, but I’m oh good, I get to take the luxury of not having to be presentable on camera in quite the same way. But when I’m doing a video scenario, I absolutely make it a point to not do that because it is off-putting to the people I’m trying to reach. In this case, I’m not trying to reach you; I already have. This is a promoted guest episode you’re trying to reach the audience, and I believe from what I can tell, you’re succeeding, so please keep at it.

Tom: Oh, you bet. Well, thank you. You know this already, but this is the very first podcast I’ve ever been a guest on. So, thank you also for making it such a welcoming place. For what it’s worth, I was not offended and didn’t think you weren’t listening. Obviously, we’re having a great time here.

But yeah, it’s something that especially in the software space, people need to be aware of because everyone’s job is—[laugh]. Whether you like it or not, here’s a controversial statement: Everyone’s job is sales. Are you selling your good ideas for your product, to your boss, to your product manager? Are you able to communicate with marketing to effectively say, “Hey, this is what, in tech support, I’m seeing. This is what people are coming to me with. This is what they care about.”

You are always selling your own performance to your boss, to your customers, to other departments where you work, to your spouse, to everybody you interact with. We’re all selling ourselves all the time. And all of that is really just communication. It’s really just making sure you’re able to meet people where they are and, effectively, bridge your point of view with theirs to make sure that we’re on the same page and, you know, we’re able to communicate well. That’s so especially important now that we’re all remote.

Corey: Just so you don’t think this is too friendly of a place, let’s go ahead and finish out the episode with a personal attack. Before you wound up working at LaunchDarkly. You were at Perforce. What’s up with that? I mean, that seems like an awfully big company to cater to its single customer, who is of course J. Paul Reed.

Tom: [laugh]. Yeah. Well, Perforce is a wonderful place. I have nothing but love for Perforce, but it is a very different landscape than LaunchDarkly, certainly. When I joined Perforce, I was supporting product called Helix ALM, which, they’re still headquartered—Perforce is headquartered here in Minneapolis. I just saw some Perforce folks last week. It truly is a great place, and it is the place that introduced me to so many DevOps concepts.

But that’s a fair statement. Perforce has been around for a while. It has grown by acquisition over the past several years, and they are putting together new offerings by mixing old offerings together in a way that satisfies more modern needs, things like virtual production, and game development, and trying to package this up in a way that you can then have a game development environment in a box, right? So, there’s a lot of things to be said for that, but it very much is a different landscape than a smaller cloud-native company. Which it’s its own learning curve, let me tell you, but truly, yeah, to your Perforce, there’s a lot more complexity to the products themselves because they’ve been around for a little bit longer.

Solid, solid products, but there’s a lot going on there. And it’s a lot harder to learn them right upfront. As opposed to something like LaunchDarkly, which seems simple on the surface and you can get started with some of the easy concepts in implementation in, like, an hour, but then as you start digging deeper, whoof, suddenly, there’s a lot more complexity hidden underneath the surface than just in terms of how this is set up, and some of those edge cases.

Corey: I have to say for the backstory, for those who are unfamiliar, is I live about four miles away from J. Paul Reed, who is a known entity in reliability engineering, in the DevOps space, has been for a long time. So, to meet him, of course I had to fly to Israel. And he was keynoting DevOpsDays Tel Aviv. And I had not encountered him before, and it was this is awesome, I loved his talk, it was fun.

And then I gave a talk a little while later called, “Terrible Ideas in Git.” And he’s sitting there just glaring at me, holding his water bottle that is a branded Perforce thing, and it’s like, “Do you work there?” He’s like, “No. I just love Perforce.” It’s like, “Congratulations. Having used it, I think you might be the only one.”

I kid. I kid. It was great and a lot of different things. It was not quite what I needed when I needed it to but that’s okay. It’s gotten better and everyone else is not me, as we’ve discussed; people have different use cases. And that started a very long-running joke that J. Paul Reed is the entirety of the Perforce customer base.

Tom: [laugh]. Yeah. And to your point, there’s definitely use cases—you’re talking about Perforce Version Control or Helix Core.

Corey: Back in those days, I don’t believe it was differentiated.

Tom: It was just called Perforce. Exactly right. But yeah, as Perforce has gotten bigger, now there’s different product lines; you name it. But yeah, some of those modern scalable problems, being able to handle giant binary files, being able to do automatic edge replication for globally distributed teams so that when your team in APAC comes online, they’re not having to spend the first two hours of their day just getting the most recent changes from the team in the Americas and Europe. Those are problems that Perforce is absolutely solving that are out there, but it’s not problems that everybody faces and you know, there’s just like everybody else, we’re navigating the landscape and trying to find out where the product actually fits and how it needs to evolve.

Corey: And I really do wish you well on it. I think there’s going to be an awful lot of—

Tom: Mm-hm.

Corey: —future stories where there is this integration. And you’d say, “Oh, well, what are you wishing me well for? I don’t work there anymore.” But yeah, but isn’t that kind of we’re talking about, on some level, of building out things that are easy, that are more streamlined, that are opinionated in the right ways, I suppose. And honestly, that’s the thing that I found so compelling about LaunchDarkly. I have a hard time imagining I would build anything for production use that didn’t feature it these days if I were, you know, better at computers?

Tom: Sure. Yeah. [laugh]. Well, we do have our opinions on how some things should work, right? Where the data is exposed because with any feature flagging system or feature management—LaunchDarkly included—you’ve got a set of rules, i.e. who should see this, where is it turned on? Where is it turned off? Who in your audience or user base should be able to see these features? That’s the rules engine side of it.

And on the other side, you’ve got the context to decide, well, you know, I’m Corey, I’m logging in, I’m in my mid-30s. And I know all this information about Corey, and those rules need to then be able to determine whether something should be on or off or which experience Corey gets. So, we are very opinionated over the architecture, right, and where that evaluation actually happens and how that data is exposed or where that’s exposed. Because those two halves need to meet and both halves have the potential to be extremely sensitive. If I’m targeting based off of a list of 10,000 of my premium users’ email addresses, I should not be exposing that list of 10,000 email addresses to a web browser or a mobile phone.

That’s highly insecure. And inefficient; that’s a large amount of text to send, over 10,000 email addresses. And so when we’re thinking about things like page load times, and people being able to push F12 to inspect the page, absolutely not, we shouldn’t be exposing that there. At the same time, it’s a scary prospect to say, “Hey, I’m going to send personal information about Corey over to some third-party service, some edge worker that’s going to decide whether Corey should see a feature or not.” So, there’s definitely architectural considerations of different use cases, but that’s something that we think through all the time and make sure is secure.

There’s a reason—I’m going to put on my sales engineer hat here—which is to say that there is a reason that the Center for Medicare and Medicaid Services is our sponsor for FedRAMP moderate certification, in process right now, expected to be completed mid-2022. I don’t know. But anybody who is unfamiliar with that, if you’ve ever had to go through high trust certification, you know, any of these compliances to make your regulators happy, you know that FedRAMP is so incredibly stringent. And that comes down to evaluating where are we exposing the data? Who gets to see that? Is security built in and innate into the architecture? Is that something that’s been thought through?

I have went so far afield from the original point that you made, but I agree, right? We’ve got to be opinionated about some things while still providing the freedom to use it in a way that is actually useful to you and [laugh] and we’re not, you know, putting up guardrails, that mean that you’ve got such a narrow set of use cases.

Corey: I’d like to hope—maybe I’m wrong on this—that it gets easier the more that we wind up doing these things because I don’t think that it necessarily has been easy enough for an awful lot of us.

Tom: When you say ‘it,’ what do you mean?

Corey: All of it. That’s the best part, I suppose the easy parts of working on computers, which I guess might be typing if you learn it early enough.

Tom: Sure. [laugh] yeah. Mario Teaches Typing, or Starcraft taught me how to type quickly. You can’t type slowly or else your expansion is going to get destroyed. No, so for someone who got their formal education in music or for someone with an eighth-grade education, I agree there needs to be resources out there.

And there are. Not every single StackOverflow post with a question that’s been asked has the response, “That’s a dumb question.” There are some out there. There’s definitely a community or a group of folks who think that there is a correct way to do things and that if you’re asking a question, that it’s a dumb question. It really isn’t. It’s getting back to the diverse backgrounds and diverse schools of thought that are coming in.

We don’t know where someone is coming from that led them to that question without the context, and so we need to continue providing resources to folks to make it easy to self-enable and continue abstracting away the machine code parts of it in friendlier and friendlier ways. I love that there are services like Squarespace out there now, right, that allow anybody to make a website. You don’t have to have a degree in computer science to spin something up and share it with the world on the web. We’re going to continue to see that type of abstraction, that type of on-ramp for folks, and I’m excited to be part of it.

Corey: I really look forward to it. I’m curious to see what happens next for you, especially as you continue—‘you’ being the corporate ‘you’ here; that’s like the understood ‘you’ are the royal ‘you.’ This is the corporate ‘you’—continue to refine the story of what it is LaunchDarkly does, where you start, where you stop, and how that winds up playing out.

Tom: Yeah, you bet. Well, in the meantime, I’m going to continue to play with things like GitHub Copilot, see how much I can autofill, and see which paths that takes me down?

Corey: Oh, I’ve been using it for a while. It’s great. Just tab-complete my entire life. It’s
amazing.

Tom: Oh, yeah. Absolutely.

Corey: [unintelligible 00:36:08] other people’s secrets start working, great, that makes my AWS bill way lower when I use someone else’s keys. But that’s neither here nor there.

Tom: Yeah, exactly. That’s a next step of doing that npm install or, you know, bringing in somebody else’s [laugh] tools that they’ve already made. Yeah, just a couple weeks ago, I was playing around with it, and I typed in two lines: I imported the LaunchDarkly SDK and the configuration for the LaunchDarkly SDK, and then I just let it autofill, whatever it wanted. It came out with about 100 lines of something or other. [laugh]. And not all of it made sense, but hey, I saw where the thought process was. It was pretty cool to see.

Corey: I really want to thank you for spending as much time and energy as you have talking about how you see the world and where you folks are going. If people want to learn more. Where’s the best place to find you?

Tom: At launchdarkly.com. Of course, any other various different booths, DevOpsDays, we’re at re:Invent, we’re at QCon right now. We’re at all sorts of places, so come stop by, say hi, get a demo. Maybe we’ll talk.

Corey: Excellent. We will be tossing links to that into the [show notes 00:37:09]. Thanks so much for your time. I really appreciate it.

Tom: Corey, Thank you.

Corey: Tom Totenberg, senior solutions engineer at LaunchDarkly. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry and insulting comment, and then I’ll sing it to you.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Casey

Casey spends his days leveraging AWS to help organizations improve the speed at which they deliver software. With a background in software development, he has spent the past 20 years architecting, building, and supporting software systems for organizations ranging from startups to Fortune 500 enterprises.

Links Referenced:

  • “17 Ways to Run Containers in AWS”: https://www.lastweekinaws.com/blog/the-17-ways-to-run-containers-on-aws/
  • “17 More Ways to Run Containers on AWS”: https://www.lastweekinaws.com/blog/17-more-ways-to-run-containers-on-aws/
  • kubernetestheeasyway.com: https://kubernetestheeasyway.com
  • snark.cloud/quinntainers: https://snark.cloud/quinntainers
  • ECS Chargeback: https://github.com/gaggle-net/ecs-chargeback
  • twitter.com/nektos: https://twitter.com/nektos

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored by our friends at Revelo. Revelo is the Spanish word of the day, and its spelled R-E-V-E-L-O. It means “I reveal.” Now, have you tried to hire an engineer lately? I assure you it is significantly harder than it sounds. One of the things that Revelo has recognized is something I’ve been talking about for a while, specifically that while talent is evenly distributed, opportunity is absolutely not. They’re exposing a new talent pool to, basically, those of us without a presence in Latin America via their platform. It’s the largest tech talent marketplace in Latin America with over a million engineers in their network, which includes—but isn’t limited to—talent in Mexico, Costa Rica, Brazil, and Argentina. Now, not only do they wind up spreading all of their talent on English ability, as well as you know, their engineering skills, but they go significantly beyond that. Some of the folks on their platform are hands down the most talented engineers that I’ve ever spoken to. Let’s also not forget that Latin America has high time zone overlap with what we have here in the United States, so you can hire full-time remote engineers who share most of the workday as your team. It’s an end-to-end talent service, so you can find and hire engineers in Central and South America without having to worry about, frankly, the colossal pain of cross-border payroll and benefits and compliance because Revelo handles all of it. If you’re hiring engineers, check out revelo.io/screaming to get 20% off your first three months. That’s R-E-V-E-L-O dot I-O slash screaming.

Corey: Couchbase Capella Database-as-a-Service is flexible, full-featured and fully managed with built in access via key-value, SQL, and full-text search. Flexible JSON documents aligned to your applications and workloads. Build faster with blazing fast in-memory performance and automated replication and scaling while reducing cost. Capella has the best price performance of any fully managed document database. Visit couchbase.com/screaminginthecloud to try Capella today for free and be up and running in three minutes with no credit card required. Couchbase Capella: make your data sing.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. My guest today is someone that I had the pleasure of meeting at re:Invent last year, but we’ll get to that story in a minute. Casey Lee is the CTO with a company called Gaggle, which is—as they frame it—saving lives. Now, that seems to be a relatively common position that an awful lot of different tech companies take. “We’re saving lives here.” It’s, “You show banner ads and some of them are attack platforms for JavaScript malware. Let’s be serious here.” Casey, thank you for joining me, and what makes the statement that Gaggle saves lives not patently ridiculous?

Casey: Sure. Thanks, Corey. Thanks for having me on the show. So Gaggle, we’re ed-tech company. We sell software to school districts, and school districts use our software to help protect their students while the students use the school-issued Google or Microsoft accounts.

So, we’re looking for signs of bullying, harassment, self-harm, and potentially suicide from K-12 students while they’re using these platforms. They will take the thoughts, concerns, emotions they’re struggling with and write them in their school-issued accounts. We detect that and then we notify the school districts, and they get the students the help they need before they can do any permanent damage to themselves. We protect about 6 million students throughout the US. We ingest a lot of content.

Last school year, over 6 billion files, about the equal number of emails ingested. We’re looking for concerning content and then we have humans review the stuff that our machine learning algorithms detect and flag. About 40 million items had to go in front of humans last year, resulted in about 20,000 what we call PSSes. These are Possible Student Situations where students are talking about harming themselves or harming others. And that resulted in what we like to track as lives saved. 1400 incidents last school year where a student was dealing with suicide ideation, they were planning to take their own lives. We detect that and get them help within minutes before they can act on that. That’s what Gaggle has been doing. We’re using tech, solving tech problems, and also saving lives as we do it.

Corey: It’s easy to lob a criticism at some of the things you’re alluding to, the idea of oh, you’re using machine learning on student data for young kids, yadda, yadda, yadda. Look at the outcome, look at the privacy controls you have in place, and look at the outcomes you’re driving to. Now, I don’t necessarily trust the number of school administrations not to become heavy-handed and overbearing with it, but let’s be clear, that’s not the intent. That is not what the success stories you have alluded to. I’ve got to say I’m a fan, so thanks for doing what you’re doing. I don’t say that very often to people who work in tech companies.

Casey: Cool. Thanks, Corey.

Corey: But let’s rewind a bit because you and I had passed like ships in the night on Twitter for a while, but last year at re:Invent something odd happened. First, my business partner procrastinated at getting his ticket—that’s not the odd part; he does that a lot—but then suddenly ticket sales slammed shut and none were to be had anywhere. You reached out with a, “Hey, I have a spare ticket because someone can’t go. Let me get it to you.” And I said, “Terrific. Let me pay you for the ticket and take you to dinner.”

You said, “Yes on the dinner, but I’d rather you just look at my AWS bill and don’t worry about the cost of the ticket.” “All right,” said I. I know a deal when I see one. We grabbed dinner at the Venetian. I said, “Bust out your laptop.” And you said, “Oh, I was kidding.” And I said, “Great. I wasn’t. Bust it out.”

And you went from laughing to taking notes in about the usual time that happens when I start looking at these things. But how was your recollection of that? I always tend to romanticize some of these things. Like, “And then everyone’s restaurant just turned, stopped, and clapped the entire time.” Maybe that part didn’t happen.

Casey: Everything was right up until the clapping part. That was a really cool experience. I appreciate you walking through that with me. Yeah, we’ve got lots of opportunity to save on our AWS bill here at Gaggle, and in that little bit of time that we had together, I think I walked away with no more than a dozen ideas for where to shave some costs. The most obvious one, the first thing that you keyed in on, is we had RIs coming due that weren’t really well-optimized and you steered me towards savings plans. We put that in place and we’re able to apply those savings plans not just to our EC2 instances but also to our serverless spend as well.

So, that was a very worthwhile and cost-effective dinner for us. The thing that was most surprising though, Corey, was your approach. Your approach to how to review our bill was not what I thought at all.

Corey: Well, what did you expect my approach was going to be? Because this always is of interest to me. Like, do you expect me to, like, whip a portable machine learning rig out of my backpack full of GPUs or something?

Casey: I didn’t know if you had, like, some secret tool you were going to hit, or if nothing else, I thought you were going to go for the Cost Explorer. I spend a lot of time in Cost Explorer, that’s my go-to tool, and you wanted nothing to do with Cost Exp—I think I was actually pulling up Cost Explorer for you and you said, “I’m not interested. Take me to the bills.” So, we went right to the billing dashboard, you started opening up the invoices, and I thought to myself, “I don’t remember the last time I looked at an AWS invoice.” I just, it’s noise; it’s not something that I pay attention to.

And I learned something, that you get a real quick view of both the cost and the usage. And that’s what you were keyed in on, right? And you were looking at things relative to each other. “Okay, I have no idea about Gaggle or what they do, but normally, for a company that’s spending x amount of dollars in EC2, why is your data transfer cost the way it is? Is that high or low?” So, you’re looking for kind of relative numbers, but it was really cool watching you slice and dice that bill through the dashboard there.

Corey: There are a few things I tie together there. Part of it is that this is sort of a surprising thing that people don’t think about but start with big numbers first, rather than going alphabetically because I don’t really care about your $6 Alexa for Business spend. I care a bit more about the $6 million, or whatever it happens to be at EC2—I’m pulling numbers completely out of the ether, let’s be clear; I don’t recall what the exact magnitude of your bill is and it’s not relevant to the conversation.

And then you see that and it’s like, “Huh. Okay, you’re spending $6 million on EC2. Why are you spending 400 bucks on S3? Seems to me that those two should be a little closer aligned. What’s the deal here? Oh, God, you’re using eight petabytes of EBS volumes. Oh, dear.”

And just, it tends to lead to interesting stuff. Break it down by region, service, and use case—or usage type, rather—is what shows up on those exploded bills, and that’s where I tend to start. It also is one of the easiest things to wind up having someone throw into a PDF and email my way if I’m not doing it in a restaurant with, you know, people clapping standing around.

Casey: [laugh]. Right.

Corey: I also want to highlight that you’ve been using AWS for a long time. You’re a Container Hero; you are not bad at understanding the nuances and depths of AWS, so I take praise from you around this stuff as valuing it very highly. This stuff is not intuitive, it is deeply nuanced, and you have a business outcome you are working towards that invariably is not oriented day in day out around, “How do I get these services for less money than I’m currently paying?” But that is how I see the world and I tend to live in a very different space just based on the nature of what I do. It’s sort of a case study and the advantage of specialization. But I know remarkably little about containers, which is how we wound up reconnecting about a week or so before we did this recording.

Casey: Yeah. I saw your tweet; you were trying to run some workload—container workload—and I could hear the frustration on the other end of Twitter when you were shaking your fist at—

Corey: I should not tweet angrily, and I did in this case. And, eh, every time I do I regret it. But it played well with the people, so that does help. I believe my exact comment was, “‘me: I’ve got this container. Run it, please.’ ‘Google Cloud: Run. You got it, boss.’ AWS has 17 ways to run containers and they all suck.”

And that’s painting with an overly broad brush, let’s be clear, but that was at the tail end of two or three days of work trying to solve a very specific, very common, business problem, that I was just beating my head off of a wall again and again and again. And it took less than half an hour from start to finish with Google Cloud Run and I didn’t have to think about it anymore. And it’s one of those moments where you look at this and realize that the future is here, we just don’t see it in certain ways. And you took exception to this. So please, let’s dive in because 280 characters of text after half a bottle of wine is not the best context to have a nuanced discussion that leaves friendships intact the following morning.

Casey: Nice. Well, I just want to make sure I understand the use case first because I was trying to read between the lines on what you needed, but let me take a guess. My guess is you got your source code in GitHub, you have a Docker file, and you want to be able to take that repo from GitHub and just have it continuously deployed somewhere in Run. And you don’t want to have headaches with it; you just want to push more changes up to GitHub, Docker Build runs and updates some service somewhere. Am I right so far?

Corey: Ish, but think a little further up the stack. It was in service of this show. So, this show, as people who are listening to this are probably aware by this point, periodically has sponsors, which we love: We thank them for participating in the ongoing support of this show, which empowers conversations like this. Sometimes a sponsor will come to us with, “Oh, and here’s the URL we want to give people.” And it’s, “First, you misspelled your company name from the common English word; there are three sublevels within the domain, and then you have a complex UTM tagging tracking co—yeah, you realize people are driving to work when they’re listening to this?”

So, I’ve built a while back a link shortener, snark.cloud because is it the shortest thing in the world? Not really, but it’s easily understandable when I say that, and people hear it for what it is. And that’s been running for a long time as an S3 bucket with full of redirects, behind CloudFront. So, I wind up adding a zero-byte object with a redirect parameter on it, and it just works.

Now, the challenge that I have here as a business is that I am increasingly prolific these days. So, anything that I am not directly required to be doing, I probably shouldn’t necessarily be the one to do it. And care and feeding of those redirect links is a prime example of this. So, I went hunting, and the things that I was looking for were, obviously, do the redirect. Now, if you pull up GitHub, there are hundreds of solutions here.

There are AWS blog posts. One that I really liked and almost got working was Eric Johnson’s three-part blog post on how to do it serverlessly, with API Gateway, and DynamoDB, no Lambdas required. I really liked aspects of what that was, but it was complex, I kept smacking into weird challenges as I went, and front end is just baffling to me. Because I needed a front end app for people to be able to use here; I need to be able to secure that because it turns out that if you just have a, anyone who stumbles across the URL can redirect things to other places, well, you’ve just empowered a whole bunch of spam email, and you’re going to find that service abused, and everyone starts blocking it, and then you have trouble. Nothing lasts the first encounter with jerks.

And I was getting more and more frustrated, and then I found something by a Twitter engineer on GitHub, with a few creative search terms, who used to work at Google Cloud. And what it uses as a client is it doesn’t build any kind of custom web app. Instead, as a database, it uses not S3 objects, not Route 53—the ideal database—but a Google sheet, which sounds ridiculous, but every business user here knows how to use that.

Casey: Sure.

Corey: And it looks for the two columns. The first one is the slug after the snark.cloud, and the second is the long URL. And it has a TTL of five seconds on cache, so make a change to that spreadsheet, five seconds later, it’s live. Everyone gets it, I don’t have to build anything new, I just put it somewhere around the relevant people can access it, I gave him a tutorial and a giant warning on it, and everyone gets that. And it just works well. It was, “Click here to deploy. Follow the steps.”

And the documentation was a little, eh, okay, I had to undo it once and redo it again. Getting the domain registered was getting—ported over took a bit of time, and there were some weird SSL errors as the certificates were set up, but once all of that was done, it just worked. And I tested the heck out of it, and cold starts are relatively low, and the entire thing fits within the free tier. And it is reminiscent of the magic that I first saw when I started working with some of the cloud providers services, years ago. It’s been a long time since I had that level of delight with something, especially after three days of frustration. It’s one of the, “This is a great service. Why are people not shouting about this from the rooftops?” That was my perspective. And I put it out on Twitter and oh, Lord, did I get comments. What was your take on it?

Casey: Well, so my take was, when you’re evaluating a platform to use for running your applications, how fast it can get you to Hello World is not necessarily the best way to go. I just assumed you’re wrong. I assumed of the 17 ways AWS has to run containers, Corey just doesn’t understand. And so I went after it. And I said, “Okay, let me see if I can find a way that solves his use case, as I understand it, through a quick tweet.”

And so I tried to App Runner; I saw that App Runner does not meet your needs because you have to somehow get your Docker image pushed up to a repo. App Runner can take an image that’s already been pushed up and deployed for you or it can build from source but neither of those were the way I understood your use case.

Corey: Having used App Runner before via the Copilot CLI, it is the closest as best I can tell to achieving what I want. But also let’s be clear that I don’t believe there’s a free tier; there needs to be a load balancer in front of it, so you’re starting with 15 bucks a month for this thing. Which is not the end of the world. Had I known at the beginning that all of this was going to be there, I would have just signed up for a bit.ly account and called it good. But here we are.

Casey: Yeah. I tried Copilot. Copilot is a great developer experience, but it also is just pulling together tons of—I mean just trying to do a Copilot service deploy, VPCs are being created and tons IAM roles are being created, code pipelines, there’s just so much going on. I was like 20 minutes into it, and I said, “Yeah, this is not fitting the bill for what Corey was looking for.” Plus, it doesn’t solve my the way I understood your use case, which is you don’t want to worry about builds, you just want to push code and have new Docker images get built for you.

Corey: Well, honestly, let’s be clear here, once it’s up and running, I don’t want to ever have to touch the silly thing again.

Casey: Right.

Corey: And that’s so far has been the case, after I forked the repo and made a couple of changes to it that I wanted to see. One of them was to render the entire thing case insensitive because I get that one wrong a lot, and the other is I wanted to change the permanent 301 redirect to a temporary 302 redirect because occasionally, sponsors will want to change where it goes in the fullness of time. And that is just fine, but I want to be able to support that and not have to deal with old cached data. So, getting that up and running was a bit of a challenge. But the way that it worked, was following the instructions in the GitHub repo.

The developer environment had spun up in the Google’s Cloud Shell was just spectacular. It prompted me for a few things and it told me step by step what to do. This is the sort of thing I could have given a basically non-technical user, and they would have had success with it.

Casey: So, I tried it as well. I said, “Well, okay, if I’m going to respond to Corey here and challenge him on this, I need to try Cloud Run.” I had no experience with Cloud Run. I had a small example repo that loosely mapped what I understood you were trying to do. Within five minutes, I had Cloud Run working.

And I was surprised anytime I pushed a new change, within 45 seconds the change was built and deployed. So, here’s my conclusion, Corey. Google Cloud Run is great for your use case, and AWS doesn’t have the perfect answer. But here’s my challenge to you. I think that you just proved why there’s 17 different ways to run containers on AWS, is because there’s that many different types of users that have different needs and you just happen to be number 18 that hasn’t gotten the right attention yet from AWS.

Corey: Well, let’s be clear, like, my gag about 17 ways to run containers on AWS was largely a joke, and it went around the internet three times. So, I wrote a list of them on the blog post of “17 Ways to Run Containers in AWS” and people liked it. And then a few months later, I wrote “17 More Ways to Run Containers on AWS” listing 17 additional services that all run containers.

And my favorite email that I think I’ve ever received in feedback was from a salty AWS employee, saying that one of them didn’t really count because of some esoteric reason. And it turns out that when I’m trying to make a point of you have a sarcastic number of ways to run containers, pointing out that well, one of them isn’t quite valid, doesn’t really shatter the argument, let’s be very clear here. So, I appreciate the feedback, I always do. And it’s partially snark, but there is an element of truth to it in that customers don’t want to run containers, by and large. That is what they do in service of a business goal.

And they want their application to run which is in turn to serve as the business goal that continues to abstract out into, “Remain a going concern via the current position the company stakes out.” In your case, it is saving lives; in my case, it is fixing horrifying AWS bills and making fun of Amazon at the same time, and in most other places, there are somewhat more prosaic answers to that. But containers are simply an implementation detail, to some extent—to my way of thinking—of getting to that point. An important one [unintelligible 00:18:20], let’s be clear, I was very anti-container for a long time. I wrote a talk, “Heresy in the Church of Docker” that then was accepted at ContainerCon. It’s like, “Oh, boy, I’m not going to leave here alive.”

And the honest answer is many years later, that Kubernetes solves almost all the criticisms that I had with the downside of well, first, you have to learn Kubernetes, and that continues to be mind-bogglingly complex from where I sit. There’s a reason that I’ve registered kubernetestheeasyway.com and repointed it to ECS, Amazon’s container service that is not requiring you to cosplay as a cloud provider yourself. But even ECS has a number of challenges to it, I want to be very clear here. There are no silver bullets
in this.

And you’re completely correct in that I have a large, complex environment, and the application is nuanced, and I’m willing to invest a few weeks in setting up the baseline underlying infrastructure on AWS with some of these services, ideally not all of them at once because that’s something a lunatic would do, but getting them up and running. The other side of it, though, is that if I am trying to evaluate a cloud provider’s handling of containers and how this stuff works, the reason that everyone starts with a Hello World-style example is that it delivers ideally, the meantime to dopamine. There’s a reason that Hello World doesn’t have 18 different dependencies across a bunch of different databases and message queues and all the other complicated parts of running a modern application. Because you just want to see how it works out of the gate. And if getting that baseline empty container that just returns the string ‘Hello World’ is that complicated and requires that much work, my takeaway is not that this user experience is going to get better once I’d make the application itself more complicated.

So, I find that off-putting. My approach has always been find something that I can get the easy, minimum viable thing up and running on, and then as I expand know that you’ll be there to catch me as my needs intensify and become ever more complex. But if I can’t get the baseline thing up and running, I’m unlikely to be super enthused about continuing to beat my head against the wall like, “Well, I’ll just make it more complex. That’ll solve the problem.” Because it often does not. That’s my position.

Casey: Yeah, I agree that dopamine hit is valuable in getting attached to want to invest into whatever tech stack you’re using. The challenge is your second part of that. Your second part is will it grow with me and scale with me and support the complex edge cases that I have? And the problem I’ve seen is a lot of organizations will start with something that’s very easy to get started with and then quickly outgrow it, and then come up with all sorts of weird Rube Goldberg-type solutions. Because they jumped all in before seeing—I’ve got kind of an example of that.

I’m happy to announce that there’s now 18 ways to run containers on AWS. Because in your use case, in the spirit of AWS customer obsession, I hear your use case, I’ve created an open-source project that I want to share called Quinntainers—

Corey: Oh, no.

Casey: —and it solves—yes. Quinntainers is live and is ready for the world. So, now we’ve got 18 ways to run containers. And if you have Corey’s use case of, “Hey, here’s my container. Run it for me,” now we’ve got a one command that you can run to get things going for you. I can share a link for you and you could check it out. This is a [unintelligible 00:21:38]—

Corey: Oh, we’re putting that in the [show notes 00:21:37], for sure. In fact, if you go to snark.cloud/quinntainers, you’ll find it.

Casey: You’ll find it. There you go. The idea here was this: There is a real use case that you had, and I looked at AWS does not have an out-of-the-box simple solution for you. I agree with that. And Google Cloud Run does.

Well, the answer would have been from AWS, “Well, then here, we need to make that solution.” And so that’s what this was, was a way to demonstrate that it is a solvable problem. AWS has all the right primitives, just that use case hadn’t been covered. So, how does Quinntainers work? Real straightforward: It’s a command-line—it’s an NPM tool.

You just run a [MPX 00:22:17] Quinntainer, it sets up a GitHub action role in your AWS account, it then creates a GitHub action workflow in your repo, and then uses the Quinntainer GitHub action—reusable action—that creates the image for you; every time you push to the branch, pushes it up to ECR, and then automatically pushes up that new version of the image to App Runner for you. So, now it’s using App Runner under the covers, but it’s providing that nice developer experience that you are getting out of Cloud Run. Look, is container really the right way to go with running containers? No, I’m not making that point at all. But the point is it is a—

Corey: It might very well be.

Casey: Well, if you want to show a good Hello World experience, Quinntainer’s the best because within 30 seconds, your app is now set up to continuously deliver containers into AWS for your very specific use case. The problem is, it’s not going to grow for you. I mean that it was something I did over the weekend just for fun; it’s not something that would ever be worthy of hitching up a real production workload to. So, the point there is, you can build frameworks and tools that are very good at getting that initial dopamine hit, but then are not going to be there for you unnecessarily as you mature and get more complex.

Corey: And yet, I’ve tilted a couple of times at the windmill of integrating GitHub actions in anything remotely resembling a programmatic way with AWS services, as far as instance roles go. Are you using permanent credentials for this as stored secrets or are you doing the [OICD 00:23:50][00:23:50] handoff?

Casey: OIDC. So, what happens is the tool creates the IAM role for you with the trust policy on GitHub’s OIDC provider, sets all that up for you in your account, locks it down so that just your repo and your main branch is able to push or is able to assume the role, the role is set up just to allow deployments to App Runner and ECR repository. And then that’s it. At that point, it’s out of your way. And you’re just git push, and couple minutes later, your updates are now running an App Runner for you.

Corey: This episode is sponsored in part by our friends at Vultr. Optimized cloud compute plans have landed at Vultr to deliver lightning fast processing power, courtesy of third gen AMD EPYC processors without the IO, or hardware limitations, of a traditional multi-tenant cloud server. Starting at just 28 bucks a month, users can deploy general purpose, CPU, memory, or storage optimized cloud instances in more than 20 locations across five continents. Without looking, I know that once again, Antarctica has gotten the short end of the stick. Launch your Vultr optimized compute instance in 60 seconds or less on your choice of included operating systems, or bring your own. It's time to ditch convoluted and unpredictable giant tech company billing practices, and say goodbye to noisy neighbors and egregious egress forever.

Vultr delivers the power of the cloud with none of the bloat. "Screaming in the Cloud"
listeners can try Vultr for free today with a $150 in credit when they visit getvultr.com/screaming. That's G E T V U L T R.com/screaming. My thanks to them for sponsoring this ridiculous podcast.

Corey: Don’t undersell what you’ve just built. This is something that—is this what I would use for a large-scale production deployment, obviously not, but it has streamlined and made incredibly accessible things that previously have been very complex for folks to get up and running. One of the most disturbing themes behind some of the feedback I got was, at one point I said, “Well, have you tried running a Docker container on Lambda?” Because now it supports containers as a packaging format. And I said no because I spent a few weeks getting Lambda up and running back when it first came out and I’ve basically been copying and pasting what I got working ever since the way most of us do.

And response is, “Oh, that explains a lot.” With the implication being that I’m just a fool. Maybe, but let’s be clear, I am never the only person in the room who doesn’t know how to do something; I’m just loud about what I don’t know. And the failure mode of a bad user experience is that a customer feels dumb. And that’s not okay because this stuff is complicated, and when a user has a bad time, it’s a bug.

I learned that in 2012. From Jordan Sissel the creator of LogStash. He has been an inspiration to me for the last ten years. And that’s something I try to live by that if a user has a bad time, something needs to get fixed. Maybe it’s the tool itself, maybe it’s the documentation, maybe it’s the way that GitHub repo’s readme is structured in a way that just makes it accessible.

Because I am not a trailblazer in most things, nor do I intend to be. I’m not the world’s best engineer by a landslide. Just look at my code and you’d argue the fact that I’m an engineer at all. But if it’s bad and it works, how bad is it? Is sort of the other side of it.

So, my problem is that there needs to be a couple of things. Ignore for a second the aspect of making it the right answer to get something out of the door. The fact that I want to take this container and just run it, and you and I both reach for App Runner as the default AWS service that does this because I’ve been swimming in the AWS waters a while and you’re a frickin AWS Container Hero, where it is expected that you know what most of these things do. For someone who shows up on the containers webpage—which by the way lists, I believe 15 ways to run containers on mobile and 19 ways to run containers on non-mobile, which is just fascinating in its own right—and it’s overwhelming, it’s confusing, and it’s not something that makes it is abundantly clear what the golden path is. First, get it up and working, get it running, then you can add nuance and flavor and the rest, and I think that’s something that’s gotten overlooked in our mad rush to pretend that we’re all Google engineers, circa 2012.

Casey: Mmm. I think people get stressed out when they tried to run containers in AWS because they think, “What is that golden path?” You said golden path. And my advice to people is there is no golden path. And the great thing about AWS is they do continue to invest in the solutions they come up with. I’m still bitter about Google Reader.

Corey: As am I.

Casey: Yeah. I built so much time getting my perfect set of RSS feeds and then I had to find somewhere else to—with AWS, the different offerings that are available for running containers, those are there intentionally, it’s not by accident. They’re there to solve specific problems, so the trick is finding what works best for you and don’t feel like one is better than the other is going to get more attention than others. And they each have different use cases.

And I approach it this way. I’ve seen a couple of different people do some great flowcharts—I think Forrest did one, Vlad did one—on ways to make the decision on how to run your containers. And I break it down to three questions. I ask people first of all, where are you going to run these workloads? If someone says, “It has to be in the data center,” okay, cool, then ECS Anywhere or EKS Anywhere and we’ll figure out if Kubernetes is needed.

If they need specific requirements, so if they say, “No, we can run in the cloud, but we need privileged mode for containers,” or, “We need EBS volumes,” or, “We want really small container sizes,” like, less than a quarter-VCP or less than half a gig of RAM—or if you have custom log requirements, Fargate is not going to work for you, so you’re going to run on EC2. Otherwise, run it on Fargate. But that’s the first question. Figure out where are you going to run your containers. That leads to the second question: What’s your control plane?

But those are different, sort of related but different questions. And I only see six options there. That’s App Runner for your control plane, LightSail for your control plane, Rosa if you’re invested in OpenShift already, EKS either if you have Momentum and Kubernetes or you have a bunch of engineers that have a bunch of experience with Kubernetes—if you don’t have either, don’t choose it—or ECS. The last option Elastic Beanstalk, but let’s leave that as a—if you’re not currently invested in Elastic Beanstalk don’t start today. But I look at those as okay, so I—first question, where am I going to run my containers? Second question, what do I want to use for my control plane? And there’s different pros and cons of each of those.

And then the third question, how do I want to manage them? What tools do I want to use for managing deployment? All those other tools like Copilot or App2Container or Proton, those aren’t my control plane; those aren’t where I run my containers; that’s how I manage, deploy, and orchestrate all the different containers. So, I look at it as those three questions. But I don’t know, what do you think of that, Corey?

Corey: I think you’re onto something. I think that is a terrific way of exploring that question. I would argue that setting up a framework like that—one or very similar—is what the AWS containers page should be, just coming from the perspective of what is the neophyte customer experience. On some level, you almost need a slide of have choose your level of experience ranging from, “What’s a container?” To, “I named my kid Kubernetes because I make terrible life decisions,” and anywhere in between.

Casey: Sure. Yeah, well, and I think that really dictates the control plane level. So, for example, LightSail, where does LightSail fit? To me, the value of LightSail is the simplicity. I’m looking at a monthly pricing: Seven bucks a month for a container.

I don’t know how [unintelligible 00:30:23] works, but I can think in terms of monthly pricing. And it’s tailored towards a console user, someone just wants to click in, point to an image. That’s a very specific user, there’s thousands of customers that are very happy with that experience, and they use it. App Runner presents that scale to zero. That’s one of the big selling points I see with App Runner. Likewise, with Google Cloud Run. I’ve got that scale to zero. I can’t do that with ECS, or EKS, or any of the other platforms. So, if you’ve got something that has a ton of idle time, I’d really be looking at those. I would argue that I think I did the math, Google Cloud Run is about 30% more expensive than App Runner.

Corey: Yeah, if you disregard the free tier, I think that’s have it—running persistently at all times throughout the month, the drop-out cold starts would cost something like 40 some odd bucks a month or something like that. Don’t quote me on it. Again and to be clear, I wound up doing this very congratulatory and complimentary tweet about them on I think it was Thursday, and then they immediately apparently took one look at this and said, “Holy shit. Corey’s saying nice things about us. What do we do? What do we do?” Panic.

And the next morning, they raised prices on a bunch of cloud offerings. Whew, that’ll fix it. Like—

Casey: [laugh].

Corey: Di-, did you miss the direction you’re going on here? No, that’s the exact opposite of what you should be doing. But here we are. Interestingly enough, to tie our two conversation threads together, when I look at an AWS bill, unless you’re using Fargate, I can’t tell whether you’re using Kubernetes or not because EKS is a small charge. And almost every case for the control plane, or Fargate under it.

Everything else just manifests as EC2 spend. From the perspective of the cloud provider. If you’re running a Kubernetes cluster, it is a single-tenant application that can have some very funky behaviors like cross-AZ chatter back and fourth because there’s no internal mechanism to say talk to the free thing, rather than the two cents a gigabyte thing. It winds up spinning up and down in a bunch of different ways, and the behavior patterns, because of how placement works are not necessarily deterministic, depending upon workload. And that becomes something that people find odd when, “Okay, we look at our bill for a week, what can you say?”

“Well, first question. Are you running Kubernetes at all?” And they’re like, “Who invited these clowns?” Understand, we’re not prying into your workloads for a variety of excellent legal and contractual reasons, here. We are looking at how they behave, and for specific workloads, once we have a conversation engineering team, yeah, we’re going to dive in, but it is not at all intuitive from the outside to make any determination whether you’re running containers, or whether you’re running VMs that you just haven’t done anything with in 20 years, or what exactly is going on. And that’s just an artifact of the billing system.

Casey: We ran into this challenge in Gaggle. We don’t use EKS, we use ECS, but we have some shared clusters, lots of EC2 spend, hard to figure out which team is creating the services that’s running that up. We actually ended up creating a tool—we open-sourced it—ECS Chargeback, and what it does is it looks at the CPU memory reservations for each task definition, and then prorates the overall charge of the ECS cluster, and then creates metrics in Datadog to give us a breakdown of cost per ECS service. And it also measures what we like to refer to as waste, right? Because if you’re reserving four gigs of memory, but your utilization never goes over two gigs, we’re paying for that reservation, but you’re underutilizing.

So, we’re able to also show which services have the highest degree of waste, not just utilization, so it helps us go after it. But this is a hard problem. I’d be curious, how do you approach these shared ECS resources and slicing and dicing those bills?

Corey: Everyone has a different approach, too. This there is no unifiable, correct answer. A previous show guest, Peter Hamilton, over at Remind had done something very similar, open-sourced a bunch of these things. Understanding what your spend is important on this, and it comes down to getting at the actual business concern because in some cases, effectively dead reckoning is enough. You take a look at the cluster that is really hard to attribute because it’s a shared service. Great. It is 5% of your bill.

First pass, why don’t we just agree that it is a third for Service A, two-thirds for Service B, and we’ll call it mostly good at that point? That can be enough in a lot of cases. With scale [laugh] you’re just sort of hand-waving over many millions of dollars a year there. How about we get into some more depth? And then you start instrumenting and reporting to something, be it CloudWatch, be a Datadog, be it something else, and understanding what the use case is.

In some cases, customers have broken apart shared clusters for that specific reason. I don’t think that’s necessarily the best approach from an engineering perspective, but again, this is not purely an engineering decision. It comes down to serving the business need. And if you’re taking up partial credits on that cluster, for a tax credit for R&D for example, you want that position to be extraordinarily defensible, and spending a few extra dollars to ensure that it is the right business decision. I mean, again, we’re pure advisory; we advise customers on what we would do in their position, but people often mistake that to be we’re going to go for the lowest possible price—bad idea, or that we’re going to wind up doing this from a purely engineering-centric point of view.

It’s, be aware of that in almost every case, with some very notable weird exceptions, the AWS Bill costs significantly less than the payroll expense that you have of people working on the AWS environment in various ways. People are more expensive, so the idea of, well, you can save a whole bunch of engineering effort by spending a bit more on your cloud, yeah, let’s go ahead and do that.

Casey: Yeah, good point.

Corey: The real mark of someone who’s senior enough is their answer to almost any question is, “It depends.” And I feel I’ve fallen into that trap as well. Much as I’d love to sit here and say, “Oh, it’s really simple. You do X, Y, and Z.” Yeah… honestly, my answer, the simple answer, is I think that we orchestrate a cyber-bullying campaign against AWS through the AWS wishlist hashtag, we get people to harass their account managers with repeated requests for, “Hey, could you go ahead and [dip 00:36:19] that thing in—they give that a plus-one for me, whatever internal system you’re using?”

Just because this is a problem we’re seeing more and more. Given that it’s an unbounded growth problem, we’re going to see it more and more for the foreseeable future. So, I wish I had a better answer for you, but yeah, that’s stuff’s super hard is honest, but it’s also not the most useful answer for most of us.

Casey: I’d love feedback from anyone from you or your team on that tool that we created. I can share link after the fact. ECS Chargeback is what we call it.

Corey: Excellent. I will follow up with you separately on that. That is always worth diving into. I’m curious to see new and exciting approaches to this. Just be aware that we have an obnoxious talent sometimes for seeing these things and, “Well, what about”—and asking about some weird corner edge case that either invalidates the entire thing, or you’re like, “Who on earth would ever have a problem like that?” And the answer is always, “The next customer.”

Casey: Yeah.

Corey: For a bounded problem space of the AWS bill. Every time I think I’ve seen it all, I just have to talk to one more customer.

Casey: Mmm. Cool.

Corey: In fact, the way that we approached your teardown in the restaurant is how we launched our first pass approach. Because there’s value in something like that is different than the value of a six to eight-week-long, deep-dive engagement to every nook and cranny. And—

Casey: Yeah, for sure. It was valuable to us.

Corey: Yeah, having someone come in to just spend a day with your team, diving into it up one side and down the other, it seems like a weird thing, like, “How much good could you possibly do in a day?” And the answer in some cases is—we had a Honeycomb saying that in a couple of days of something like this, we wound up blowing 10% off their entire operating budget for the company, it led to an increased valuation, Liz Fong-Jones says that—on multiple occasions—that the company would not be what it was without our efforts on their bill, which is just incredibly gratifying to hear. It’s easy to get lost in the idea of well, it’s the AWS bill. It’s just making big companies spend a little bit less to another big company. And that’s not exactly, you know, saving the lives of K through 12 students here.

Casey: It’s opening up opportunities.

Corey: Yeah. It’s about optimizing for the win for everyone. Because now AWS gets a lot more money from Honeycomb than they would if Honeycomb had not continued on their trajectory. It’s, you can charge customers a lot right now, or you can charge them a little bit over time and grow with them in a partnership context. I’ve always opted for the second model rather than the first.

Casey: Right on.

Corey: But here we are. I want to thank you for taking so much time out of well, several days now to argue with me on Twitter, which is always appreciated, particularly when it’s, you know, constructive—thanks for that—

Casey: Yeah.

Corey: For helping me get my business partner to re:Invent, although then he got me that horrible puzzle of 1000 pieces for the Cloud-Native Computing Foundation landscape and now I don’t ever want to see him again—so you know, that happens—and of course, spending the time to write Quinntainers, which is going to be at snark.cloud/quinntainers as soon as we’re done with this recording. Then I’m going to kick the tires and send some pull requests.

Casey: Right on. Yeah, thanks for having me. I appreciate you starting the conversation. I would just conclude with I think that yes, there are a lot of ways to run containers in AWS; don’t let it stress you out. They’re there for intention, they’re there by design. Understand them.

I would also encourage people to go a little deeper, especially if you got a significantly large workload. You got to get your hands dirty. As a matter of fact, there’s a hands-on lab that a company called Liatrio does. They call it their Night Lab; it’s a one-day free, hands-on, you run legacy monolithic job applications on Kubernetes, gives you first-hand experience on how to—gets all the way up into observability and doing things like Canary deployments. It’s a great, great lab.

But you got to do something like that to really get your hands dirty and understand how these things work. So, don’t sweat it; there’s not one right way. There’s a way that will probably work best for each user, and just take the time and understand the ways to make sure you’re applying the one that’s going to give you the most runway for your workload.

Corey: I will definitely dig into that myself. But I think you’re right, I think you have nailed a point that is, again, a nuanced one and challenging to put in a rage tweet. But the services don’t exist in a vacuum. They’re not there because, despite the joke, someone wants to get promoted. It’s because there are customer needs that are going on that, and this is another way of meeting those needs.

I think there could be better guidance, but I also understand that there are a lot of nuanced perspectives here and that… hell is someone else’s workflow—

Casey: [laugh].

Corey: —and there’s always value in broadening your perspective a bit on those things. If people want to learn more about you and how you see the world, where’s the best place to find you?

Casey: Probably on Twitter: twitter.com/nektos, N-E-K-T-O-S.

Corey: That might be the first time Twitter has been described as a best place for anything. But—

Casey: [laugh].

Corey: Thank you once again, for your time. It is always appreciated.

Casey: Thanks, Corey.

Corey: Casey Lee, CTO at Gaggle and AWS Container Hero. And apparently writing code
in anger to invalidate my points, which is always appreciated. Please do more of that, folks. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, or the YouTube comments, which is always a great place to go reading, whereas if you’ve hated this podcast, please leave a five-star review in the usual places and an angry comment telling me that I’m completely wrong, and then launching your own open-source tool to point out exactly what I’ve gotten wrong this time.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Scott

Cloud security historian.

Developed flaws.cloud, CloudMapper, and Parliament.

Founding team for fwd:cloudsec

Links:

  • Block: https://block.xyz/
  • Twitter: https://twitter.com/0xdabbad00

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Vultr. Optimized cloud compute plans have landed at Vultr to deliver lightning fast processing power, courtesy of third gen AMD EPYC processors without the IO, or hardware limitations, of a traditional multi-tenant cloud server. Starting at just 28 bucks a month, users can deploy general purpose, CPU, memory, or storage optimized cloud instances in more than 20 locations across five continents. Without looking, I know that once again, Antarctica has gotten the short end of the stick. Launch your Vultr optimized compute instance in 60 seconds or less on your choice of included operating systems, or bring your own. It's time to ditch convoluted and unpredictable giant tech company billing practices, and say goodbye to noisy neighbors and egregious egress forever. Vultr delivers the power of the cloud with none of the bloat. "Screaming in the Cloud" listeners can try Vultr for free today with a $150 in credit when they visit getvultr.com/screaming. That's G E T V U L T R.com/screaming. My thanks to them for sponsoring this ridiculous podcast.

Corey: Couchbase Capella Database-as-a-Service is flexible, full-featured and fully managed with built in access via key-value, SQL, and full-text search. Flexible JSON documents aligned to your applications and workloads. Build faster with blazing fast in-memory performance and automated replication and scaling while reducing cost. Capella has the best price performance of any fully managed document database. Visit couchbase.com/screaminginthecloud to try Capella today for free and be up and running in three minutes with no credit card required. Couchbase Capella: make your data sing.

Corey: Welcome to Screaming in the Cloud, I’m Corey Quinn. I am joined by a returning guest with a bit of a different job. Scott Piper was formerly an independent security researcher—basically the independent security researcher in the AWS space—but now he’s a Principal Engineer over at Block. Scott, welcome back.

Scott: Thanks for having me, again, Corey.

Corey: So, you’ve taken a corporate job, and when that happened, I have to confess, I was slightly discouraged because oh, now it’s going to be like one of those stories of when someone you know goes to work at Apple because no one knows anyone at Apple; we just used to know people who went there and then we kind of lost touch because it’s a very insular thing. Not the Block slash Square slash whatever they’re calling themselves this week has that reputation. But InfoSec is always a very nuanced space and companies that have large footprints and, you know, handle financial transaction processing generally don’t encourage loud voices that attract attention around anything that isn’t directly aligned with the core mission of the company. But you’re still as public and prolific as ever. Was that a difficult balance for you to strike?

Scott: So, when I was considering employment options, that was something that I made
clear to any companies that I was talking to, that this is something that probably will and should continue because a lot of my value to these companies is because I’m able to have discussions, able to impact change because of that public persona. So yeah, so I think that it was something that they were aware of, and a risk that they took. [laugh]. But yeah, it’s been useful.

Corey: This is the sort of conversation I would have expected to have with, “Yeah, things seem to be continuing the same, and I haven’t rocked any boats, yet and they haven’t fired me, knock on wood.” Except that recently you’ve launched yet something else that I am personally a fan of. Now, before we get into the specifics of what it is you’re up to these days, I should call out that since your last appearance on this show, I have really leaned into the Thursday newsletter podcast duo of Last Week in AWS: Security Edition. Rounding up what happened the previous week—yes, it was the previous week, and it comes out on Thursdays—because, you know, timing and publication, things are hard, computers, you know how it is—aimed at a target audience that is very much not you: People who have to care about security, but are not immersed in the space. It’s a, “All right, what now? What do I have to pay attention to?”

Because there’s a lot of noise in this space, there’s a lot of vendor-captured stuff out there. There’s very little that is for people who work in security but don’t have the word security anywhere near their job title. And I have to confess that one of my easy shortcuts is, “Oh, it’s a pretty thin issue this week,” which is not inherently a bad thing, let’s be clear, it’s not yay, the three things you need to care about in security then eight more of filler; that’s not what we’re about. But I always want to make sure I didn’t miss something meaningful, and one of my default publication steps is, “What’s Scott been tweeting about this week?” Just to make sure that I didn’t miss something that I really should be talking about.

And every single time I pull up your Twitter feed, I find myself learning something, whether it’s a new concept, or whether it is a nuance on an existing thing I was already aware of. So first, thank you for all the work that you do as a member of the community, despite having a, “Regular corporate job,” quote-unquote, you’re still very present. It’s appreciated.

Scott: Thank you. Yeah. And I mean, that newsletter is great for people that don’t want to be spending multiple hours per day trolling through Twitter and reading that. So, it provides, also, something great for the community to not have to spend all that time on Twitter like I do [laugh], unfortunately.

Corey: It also strives—sort of—to be something approaching an upbeat position of not quite as cynical and sarcastic as the Monday issue. I try to be not just this is the thing that happened, but go a little bit into and this is why it matters. This is how to think about it. This thing that Amazon put out is nonsense, however, here’s the kernel hidden within it that might lead to something, such as thinking about how you do sign-on, or how to think about protecting MFA devices, or stuff like that you normally care about a lot right after you really should have cared about it but didn’t at all. So, it’s just the idea of aiming in a slightly different audience.

Scott: Yeah definitely. And it provides value that it does, it takes some delay so that you can read what everybody has written, how they’ve responded to the different news outtakes, you’re not just including the hot takes. For example, as of this morning, there’s a certain incident with an authentication provider, and it’s not really clear if there was actually a breach or not. And so it’s valuable to take a moment to understand what happened, get all the voices to have expressed their points, so you can summarize those issues.

Corey: An internal term that we’ve used to describe the position here is that I am prolific but I also have things to do as a part of my job that do not involve sitting there hitting refresh on Twitter like mad all the time. The idea is to have the best take not the first take—

Scott: Exactly.

Corey: And if that means that I lose a bunch of eyeballs and early ad impressions in the middle of the night and whatnot, well, great. I don’t sell ad impressions anyway, so what does it matter? It winds up lending itself to a more thoughtful analysis of figuring out, in the sober light of day, is this a nothing-burger or is this enormous? With that SSO issue that you’re alluding to—[cough] Okta—sorry, something caught in my throat there—very clearly, something is going on, but if I had written next week’s newsletter last night while it was still very unclear, it would have been a very different tone than the one that I would have written this morning after their public statement, and even still a certainly different tone that it would take a couple of days once more information is almost certain to come to light. And that is something that is, I think, underappreciated in certainly on Twitter, where an old tweet—there’s nothing worse than an old tweet unless you’re using it to drag someone for something—that, “Well, we have different perspectives on that nowadays. It’s not 2018 anymore.” Right. Okay, cool.

Scott: Yep. [laugh].

Corey: But something that you’ve done has been a bit of a pivot lately. Historically, you have been right there in my sweet spot of needling cloud providers for their transgressions in various ways. Cool, right there with you. We could co-author a book on the subject. But lately, you’ve started a community list of [IMSDv2 00:07:04] abuses.

Now, first, we should talk about what IMSDv2 is. It’s the name that it clearly came from Amazon because that’s a name only a cloud provider bad at naming things could possibly love. What is it?

Scott: So, it’s the Instance Metadata Service, Version Two. If there’s a version two, you can imagine there was a version one at some point. And the version two—

Corey: And there’s a version two because Amazon prod—the first one was terrible, but they don’t turn anything off, ever, so this is the way and the light and the future; we’re going to leave that old thing around until your great-grandchild dies of old age.

Scott: Exactly, yeah. So, when EC2s first came out, and IAM roles first came out, you wanted to give your EC2s the ability to use AWS privileges, so this is how those EC2s are getting access to their credentials that they can use. And the way in which this was originally done was there’s this magic IP address, this 169.254.169.254 IP address, which is very important for security on AWS because if anything can access that magic IP address from an EC2 instance, you can steal their credentials of that EC2, and therefore basically become that EC2 instance, in terms of what it can do in the AWS environment.

And so in 2019, there was a large breach of Capital One that was related to this. And so as a result of that—I think that AWS probably had this new version, probably, in the works for a while, but I think that motivated their faster release of this new version, and so IMDSv2 changed how you would obtain these credentials. So, you basically—instead of making a single GET request to this IP address, now you had to make multiple requests, they were now PUT request instead of a GET request, there was a challenge and response, there’s the hop limit. So, there’s all these various things that are going to make it harder and basically mitigate a lot of the different types of vulnerabilities that previously would be used in order to obtain these credentials. The problem, though, is that IMDSv1 still exists on EC2s, unless you as a customer are enforcing IMDSv2.

And so, in order to do this in a large environment, it’s difficult—theoretically, it’s a simple thing; all you should have to do is update your SDK and now you’re able to make use of the latest version. And if you’re using any version of the SDK that was released in the past over two years, you already should be using IMDSv2 there, but you have to enforce it. And so that’s where the problem is. And what was most problematic to me is now that I work for a company, we have run into the problem that there are some vendor solutions that we use that weren’t allowing us to enforce IMDSv2 across all of our different accounts. And this is something I’ve heard from a number of other customers as well.

And so I decided to create this list with vendors that I’ve had to deal with, vendors that other customers have had to deal with, in order to basically try and solve this problem once and for all. It’s been multiple years now and a lot of these vendors, unfortunately, were also security vendors. And so that makes the conversation a little bit easier, to basically put them on this wall-of-shame and say, “You’re a security vendor and you’re not allowing your customers to enforce best practices of security.”

Corey: I want to call on a couple of things around that. Originally the metadata service was used for a number of other things—still is—beyond credentials. It is not the credential service as envisioned by a lot of folks. The way that—also we’ll find those credentials empty until there’s an EC2 instance role, and those credentials will both be scoped what that instance does and automatically rotated in the fullness of time so they’re not long-lived credentials that once you have them, they will last forever. This is, of course, a best practice and something you should be leveraging, but scope those credentials down, or you wind up with one of the ways that was chained together in the Capital One breach a few years ago.

It’s also worth noting that service would have been more useful earlier in time with a few functions. For example, you can use the metadata service to retrieve the instance tags about the EC2 instance. When I requested it in 2015, it was not possible. But they had released it in January of this year, 2022, long after we have all come up with workarounds for this, where we could have used that to set the hostname internally on the system, if you’re looking for something basic and easy. It would have been something then you could have used to automatically self-register with DNS without having to jump through a whole bunch of hoops to do it manually.

And you look at this, and it’s wow, that’s a whole lot of crappy tooling I can just throw into the trash heap of history you don’t need anymore. But the IMSDv2, you’re right, makes it a lot harder, there has to be a conversation, not just something you can sort of bankshot something off of to get access to it. And it’s a terrific mitigation. What I’ve liked about your list of more or less shaming companies for doing this is, on the one hand, you have companies who take themselves off of the list as soon as it’s up there. It’s, “Oh, we love when people talk about us. Wait, what’s that? They’re saying something unkind? On the internet?” And they’ll fix it, which honestly is better than I expected.

And then every once in a while you’ll see something that’s horrifying of, “Oh, yeah, we’re
not vulnerable to that at all because we tell you to create permanent long-lived credentials, store them on disk and we’ll use those instead.” And it’s… that is, like, guaranteeing that no one is going to break down your door by making your walls out of tissue paper. Don’t do that. Like, that has gone so far around the band that has come back around again. So, hopefully that got fixed.

Scott: And I think you pointed out a couple of things I want to talk about with this is that, one, it has actually been very successful in terms of getting large vendors to make changes. Currently, of the seven vendors that have ever been listed there, are three of them have already made fixes and have been removed from the list. And the list has only been up for about a month. And so, in terms of getting enterprise solution vendors to make changes within, like, just a few weeks is very surprising to me. And these are things that people have been asking for for years now, and so it had motivated them a lot there.

And the other thing that I want to point out is people have looked at the success that it’s had and considered maybe we should make wall-of-shame lists, for all the things that we want. And I want to point out that there are some things about this problem, the IMDSv2 specifically, that make it work for having this wall-of-shame list like this. One of them is that not supporting or not allowing customers to enforce IMDSv2 is basically always bad. There is not a use case where you can make a claim—

Corey: There is no nuance where that, in this case, is the thing to do, like having an open S3 bucket: There are use cases where that is very much something you want to do, but it’s the uncommon case.

Scott: Exactly. That I think is an important thing. Another thing is it’s not just putting up a list, you know, like that is what people are seeing publicly, but behind the scenes, there’s a lot of other things that are happening. One, I am communicating with various customers, customers that are reporting this issue to me, in order to try to better understand what’s happening there, so that I can then relay that information to the company. So, I’m not just putting up the list; I’m also, behind the scenes, having conversations with these different companies to try to get timelines from them, to try to make sure that they are aware of the problem, they are aware that they’re on this list, how to get off the list. So, there’s that conversation happening.

There’s also the conversation that I’m happening with AWS in order to make various requests that AWS improve this for customers, to make this easier. And this is something that is public on that repo. I have my list of requests to AWS so that people can relay that to their own TAMs at AWS to basically say these are things we want as well. And so this includes things like, “I want an AWS account to have the ability to default to always be enforcing IMDSv2.” You know, so as an example, when you create an EC2 through the web console—which people can say, oh, you should always be using Infrastructure as Code; the reality is many folks are using the web console to create EC2s to do other changes.

And when you create an EC2 in the web console, by default, it’s going to allow IMDSv1 still. And so my request to AWS is, you should allow me to just default enforce IMDSv2. Also, the web console does not give you visibility into which EC2s are enforcing it and which ones are not. And also, you do not have the ability in the web console to enforce it. You cannot click on an EC2 and say, “Please enforce it now.”

So, it’s all these various, like, minor changes that I’m requesting AWS to do.

Corey: It has to be done at instance creation time.

Scott: Exactly. And so there is an API that you can make in order to change it afterwards, but that’s only an API so you have to use the CLI or some other mechanism; you can’t do it in the web console. But the other thing that I’m requesting AWS do is if security is a priority for AWS and they have all these other partners that are security companies, that they should be requiring their partners to also be enforcing this in their various products. So, if a partner is basically not allowing your AWS customers to enforce security best practices, then perhaps that partnership should be revoked in some way. And so that’s a more aggressive thing that I’m asking AWS to do, but I think is reasonable.

Corey: I’d also like them to get all of their own first-party services to support this, too.

Scott: That’s true as well. So, AWS is currently on the list. And so, they have one service, Data Pipelines, which if you are an AWS customer and you are using that service, you are not going to be able to enforce IMDSv2 in your environment. So, AWS themselves, unfortunately, is not allowing customers to enforce this. And then AWS themselves in their own production servers, we have seen indications that they do not enforce IMDSv2 on their own production servers.

So, the best practice that they are telling customers to follow, they unfortunately are not following it themselves. And so the way in which we saw this was Orca is a security company that ended up finding this issue with AWS—and there’s a lot of questions in terms of what all exactly they found—but they had this post that they called “Breaking Formation” in which they were somehow able to find—basically exploit to some degree—and again, it’s unclear exactly what they were able to exploit here—but they were able to exploit AWS production servers that are responsible for the CloudFormation service. And in their blog post, they had a screenshot which showed that those production servers are not enforcing IMDSv2. And so AWS themselves is struggling with this as well, as are many customers. So, it’s something that, you know, I put together this list of requests in hopes that AWS can make it easier for not only customers but also themselves to be able to enforce it.

Corey: There are a lot of different things that we wish companies did differently, particularly if that company is AWS. Why is this the particular windmill that you’ve decided to tilt at given—let’s say—it’s not exactly slim pickins out there as far as changes that we wish companies would make? Obviously, you mentioned at one point, there is no drawback to enabling this, but a lot could be said for other aspects as well. Why is this one so important?

Scott: So, in part, I personally have some, I guess, history with this [laugh], basically, IMDSv2, and so we can discuss this. This is back when Capital One had their breach in 2019, there was this Senator, Senator Ron Wyden, who sent this email over to AWS, to Steve Schmidt, who was the CISO at the time there and still is the CISO, and he basically—

Corey: Now, he’s head of security for all of Amazon.

Scott: Yeah, yeah.

Corey: CJ is now the AWS CISO. And he has the good sense to hide.

Scott: Yeah. [laugh]. So, at the time, this Senator Ron Wyden had send over this email—and obviously it’s not Senator Ron Wyden himself, you know, it’s one of his, like, technical people on staff that is able to give him this information—and he sends this email to AWS saying, “Hey, this metadata service played a role in this very significant breach. Why hasn’t this been fixed?” And Steve Schmidt responded, and because it’s communications between a senator, I guess it has to become public.

So, Steve Schmidt responds, saying that, “Hey, we never knew that this was an issue before,” is essentially what he responds with. And that irked me because I had reported this to AWS previously, as had many other people. So, there was a conference presentation by this guy Andrés Riancho at BlackHat, I believe in 2014, and he had presented previously in 2013, so it was a known issue; it had been around for a while. But I took the time to actually report it to AWS Security. So, I went through the correct channel of making sure that AWS was aware of a security concern, as a security researcher—so reporting it through that correct channel there—and provided Senator Ron Wyden with all this information.

And so, then he then requested that the FTC begin a federal investigation into AWS, related to basically not following the best practices that security researchers have recommended. So, that was, kind of like, my early, I guess, involvement with this issue. So, it’s something that I’ve been interested in for a while to make sure that this is resolved completely at some point.

Corey: This episode is sponsored by our friends at Oracle Cloud. Counting the pennies, but still dreaming of deploying apps instead of “Hello, World” demos? Allow me to introduce you to Oracle’s Always Free tier. It provides over 20 free services and infrastructure, networking, databases, observability, management, and security. And—let me be clear here—it’s actually free. There’s no surprise billing until you intentionally and proactively upgrade your account. This means you can provision a virtual machine instance or spin up an autonomous database that manages itself, all while gaining the networking, load balancing, and storage resources that somehow never quite make it into most free tiers needed to support the application that you want to build. With Always Free, you can do things like run small-scale applications or do proof-of-concept testing without spending a dime. You know that I always like to put asterisks next to the word free? This is actually free, no asterisk. Start now. Visit snark.cloud/oci-free that’s snark.cloud/oci-free.

Corey: It’s always fun watching where people come from, as far as the security problems that they call out. There was, I believe in the cloud security forum Slack, a thread of recently about what security issues are top-of-mind and that should be fixed as a baseline expectation. In fact, let me dig it out because that is one of those things that I think is well worth having the conversation properly on this.

Good examples of risky, insecure defaults in AWS. And people are talking about IMDSv1, and they’re talking about all kinds of other in-depth things, and my contribution to it was, “If I go and I spin up an AWS account, until I go out of my way, I’m operating as root in that account. That seems bad.” And a few responses to that were oh, the basically facepalming, “Oh, of course.” I wish that there were an easy way to get AWS SSO as the default because it is the right answer for so many different things. It solves so many painful problems that otherwise you’re going to wind up stuck with.

And this stuff is hard and confusing; when people are starting out with this for the first time, they’re not approaching this from, “All right, how do I be extremely secure?” They want to get some work done. For fun a year ago, I spun up a test account—unattached to any organization—and because account aliases are globally unique, I somehow came up with the account ‘shitposting’ because that’s pretty much what I use it for. The actual reason I wanted that was I wanted something completely unattached from any other account that I could easily take screenshots from at any point, and the worst case scenario is okay, I’ve exposed some credential of my own in an account that has no privileged access to anything; I just have to apologize for all the Bitcoin mining now. And honestly, I think AWS would love that marketing campaign; they’d see my face on a billboard looking horrified. It’ll be great.

But I turned on every security service as I went because, of course, security is the most important thing. And there were so many to turn on, and the bill was approaching 50 bucks a month for an empty account. And it’s. It starts to feel a little weird and more
than a little wrong.

Scott: [laugh]. Yeah, my personal concern in terms of default security features is really that problem of the cost controls, I think that that still is a big issue that AWS does not have cost controls such that when a student wants to try and use AWS for the very first time and somehow they spin up large EC2 instance, or they just you know, end up creating an access key and that access key gets leaked and somehow their account gets compromised and used for Bitcoin mining, now they’re stuck with that large AWS bill. For a student who has no budget, is in debt, and now is suddenly being, you know, hit with multiple thousands of dollars on their bill, that I think is very problematic, and that is something that I wish AWS would change as a default is basically, if you are creating AWS account for the very first time, have some type of—I don’t know how this would look, but maybe just be able to say, like, I don’t ever want this AWS account to spend more than $100 per month, and I’m okay if you end up destroying all my data in the account because I have no money and money is more important to me than whatever data I may store in here.

Corey: Make an answer to that question mandatory, just as putting a credit card in is mandatory. Because there are two extremes here. It’s more or less the same problem of AWS not knowing who its customers are beyond an AWS account, but there’s a spectrum somewhere between I’m a student who wants to learn how the cloud works, and my approach to security is very much the same. Don’t let randos spin up resources in my account, and I don’t ever want to be charged. If that means you turn off my “Hello World” blog post, okay, great.

On the other end, it’s this is Netflix. And this is our, you know, eight-millionth account that we’re spending up to do a thing and what do you mean you’re applying service quotas to it? I thought we had an understanding?—everything is a service quota, let’s be clear—

Scott: Yep.

Corey: —or a company that’s about to run a Superbowl ad. Yeah, there’s going to be a lot of traffic there. Don’t touch it. Just make it work. We don’t care what it costs.

Understanding where you fall on the cost perspective—as well as a security point of view of, “We’re a bank, which means forget security best practices, we have compliance obligations that cannot be altered in this account and here’s what they are.” There has to be a way that is easy and approachable for people to wind up moving that slider to whatever position best represents them. Because there are accounts where I never want to be charged a thing. And that’s an important thing because—and I’ve been talking about this for a while because I’m convinced it’s a matter of time—that poor kid who wound up trading on margin at Robinhood, woke up saw that he was seven-hundred-and-some-odd grand in debt and killed himself. When it all settled out, I think he turned something like a $30,000 profit when all was said and done, which just serves to make it worse.

I can see a scenario in which that happens, and part of the contributors to it are that we used to see that the surprise bill for compromised accounts was 10, 15, 20 grand. Now, they’re 70 to 90 because there are more regions, more services to run containers—because of course there are—and the payoff is such that the people exploiting this have gotten very practiced and very operationalized at spinning up those resources quickly, and they cost a lot very quickly. I mean, the third use case that they’re not aiming at yet is people like me, where it’s, oh, you have a free account that sandboxed; I want to get the high score on the free tier because all their fraud is attuned to you making money. With me, it’s nope, just going to run up the store to embarrass Amazon. That’s not a common exploit vector, but I’m very much here.

Scott: [laugh]. Yep. And that also is the thing though: The Denial of Wallet attack is also a concern on AWS, as well, where you’ve written a blog post about this, how if you are able to make use of data transfer in different ways, you can run up very high multi-million dollar bills in people’s AWS accounts and even AWS’s own protections and defenses against trying to look for cost spikes and things like that is delayed by multiple hours. And so you can still end up spending a lot of money in people’s accounts, or one thing that’s wild is an S3 object locking; that feature, the whole purpose behind it is to ensure data can never be deleted. It exists for various compliance reasons, so even AWS themselves cannot delete certain data.

So, if an attacker is able to abuse that functionality in somebody’s account, they can end up locking data such that for the next 100 years, it can never be deleted and you’re going to have to pay for that for the next 100 years inside your account. The only way of not paying for that anymore is to move everything that you have in an AWS account to a new account, and then ask AWS to delete that account, which is not going to be reasonable under most circumstances.

Corey: Yeah, alternatively, it’s one of those scenarios where well, the only other option is to start physically ripping hard drives out of racks in a bunch of different data centers. It’s wild to me. It’s such an attack surface that honestly I believe for the longest time that AWS Security is otherworldly good. And as we start seeing from these breaches, no, what really is otherworldly good is their ability to apply pressure to people not to go public with things they discover that they then wind up keeping quiet because once this whole Orca stuff came out, we started digging, and Aidan Steele found some stuff where you could just get unfiltered, raw outputs of CloudTrail events by setting up a couple of rules in weird ways.

And that was a giant problem, and it was never disclosed publicly. I don’t know if any of my events were impacted; I can’t trust that they would have told me if they were. And for the first time, I’m looking at things like confidential computing, which are designed around well, what if you don’t trust your cloud provider? Historically, I guess I was naive because my approach was, “Well, then you shouldn’t be using the cloud.” Now it’s, “Well, that’s actually kind of a good point.”

Because it’s not that I don’t trust my cloud provider to necessarily do what they’re telling me. I just don’t trust them to tell me what they’re doing. And that’s part of it. The, “Well, we found an issue, but you can’t prove we had an issue, so we’re going to say nothing.” And when it comes to light—because it always does—it erodes trust in a big way. And trust is everything in cloud.

Scott: Yeah. And so with some of the breaches that have come out, I created another GitHub repo to start tracking all the different security incidents that I could find for the three cloud providers, Azure, GCP, and AWS. And so on there, I started listing not only some of the blog posts from security companies that had been able to exploit vulnerabilities in the cloud providers, but also just anything else that I felt was a security mistake in some way. And so there’s a number of things I tried to avoid on there. Like, I tried to avoid listing something that’s kind of like a business decision, for example, services that get released that don’t have CloudTrail support. That’s a security concern to me, but that’s kind of a business decision that they decided to release a service before it supported all that functionality.

So, I tried to start listing off all those different things in order to also keep track of you know, is there a security provider that’s worse than the others? Are there any type of common patterns that I can see? And so I tried to look through some of those different things. And that’s been interesting because also I really only focus on AWS, and so I haven’t really known what all has been happening with GCP and Azure. And that was interesting because there’s been two issues that have happened on AWS where the exact same issue happened on the other cloud providers. And so that tells me, that’s concerning to me because that tells me tht—

Corey: Because those are not discovered at the same time let’s be clear.

Scott: Yeah. These were, like, over a year apart. And so basically, somebody had found something on GCP, and then a year-plus later, somebody else found the exact same issue on AWS. And then similarly, there was an issue with Azure and then a year-plus later, same issue on AWS. And that’s concerning because that tells me that AWS may not be monitoring what are the security issues that are impacting other cloud providers, and therefore checking whether or not they happen to themselves?

That’s something that you would expect a mature security team to be doing is to be monitoring what are public incidents that are happening to my competitors, and am I impacted similarly? Or what can I do to try and identify those issues, fix them, make sure they never happen? All those types of steps in terms of security maturity. And that’s something that then I’m a little concerned of that we’ve seen those issues happen before. There’s also, on AWS specifically, they have had a number of issues related to their IAM-managed policies that keep cropping up.

And so they have had a number of incidents where they were releasing policies that shouldn’t have been released in some way. And that’s concerning that showed that they don’t really have a change management process that you would expect. Usually, you would expect a company to be having GitHub PRs and approval processes and things like that, in order to make sure that there’s a second set of eyes on something before it gets released.

Corey: Particularly things of this level of sensitivity. This is not—like, I was making fun of them a day or two ago for having broken the copyright footer and not updating them since 2020 because instead of the ‘copyright’ symbol, they used an ‘at’ symbol. Minor stuff, but like that’s fun to needle people about, but it doesn’t actually matter for anything.

Scott: Yeah.

Corey: Security matters and mistakes show.

Scott: Yeah. And so there had been some examples where they released a policy that was called, like, ‘cheese puffs something’ and it’s like, okay, that’s clearly, like, an internal service of some sort. But I’d called them out and, like, I’d sent an email to AWS Security being like, “Hey, you need to make sure that you have change management processes on your IAM policies because one day you’re going to do something that is bad.” And one day they did. They made a change to the read-only access policy, and that basically—they removed every single privilege, somebody had ended up, you know, internally, removed every single privileges to the read-only access policy and replaced it with a whole bunch of write privileges for, I think, the Cassandra service.

And so, that was like, clearly they’ve made a mistake that they should have made sure they were correcting because you know, they had these previous incidents. Another kind of similar one was in December, there was a support policy where they had added S3 GetObject to that policy, and that was concerning in terms of have they just given all of their support employees access to everybody’s content in their S3 buckets? And so AWS made some statements saying that there were other controls in place there so it wouldn’t have been possible. But it’s those types of things that [crosstalk 00:33:17]—

Corey: Originally, those statements were made on Twitter, let’s be clear here.

Scott: Yes. Yeah. [laugh].

Corey: And I feel like there’s a—while I deeply appreciate how accessible a lot of their senior people are, I cannot point the executive leadership team at a client to some tweets that someone made. That is not a public statement of record that works on this.

Scott: Exactly.

Corey: They’re learning. We’ll get there sooner or later, I presume. I want to thank you for taking the time to speak with me, as always, I’ll throw links to these repos into the [show notes 00:33:46], but if they want to know more what you have to say, where’s the best place to find you?

Scott: So, my Twitter, which, unfortunately, is a handle written in hex, but it’s—‘dabbadoo’ is how you would pronounce it, but it’s probably easiest to see a link for it. So, that’s probably the main place to look for me.

Corey: That’s why my old Twitter handle was my amateur radio callsign. I don’t use that one anymore. It’s just easier. And I think that’s the right answer. Besides, given what you do, it’s easy enough if people want your attention. They screw up badly enough, you’ll come to them.

Scott: Yep. [laugh].

Corey: Scott, I really appreciate your time. Thanks again.

Scott: Thank you.

Corey: Scott Piper, Principal Engineer at Block and, more or less, roving security troubadour for lack of a better term. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice or a comment on the YouTubes saying that this episode is completely invalid because you wind up using the old version of the metadata service and you’ve never had a problem. That you know of.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Molly

Molly White is a software engineer and team lead. She's also a longtime Wikipedia editor and advocate for free and open knowledge, and has more recently become an outspoken critic of cryptocurrencies and web3 more broadly.

Links:

  • web3isgoinggreat.com: https://web3isgoinggreat.com
  • lasttweetinaws.com: https://lasttweetinaws.com
  • mollywhite.net: https://mollywhite.net
  • @molly0xFFF: https://twitter.com/molly0xFFF
  • @web3isgreat: https://twitter.com/web3isgreat
  • ponzl: http://ponzl.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Vultr. Optimized cloud compute plans have landed at Vultr to deliver lightning fast processing power, courtesy of third gen AMD EPYC processors without the IO, or hardware limitations, of a traditional multi-tenant cloud server. Starting at just 28 bucks a month, users can deploy general purpose, CPU, memory, or storage optimized cloud instances in more than 20 locations across five continents. Without looking, I know that once again, Antarctica has gotten the short end of the stick. Launch your Vultr optimized compute instance in 60 seconds or less on your choice of included operating systems, or bring your own. It's time to ditch convoluted and unpredictable giant tech company billing practices, and say goodbye to noisy neighbors and egregious egress forever. Vultr delivers the power of the cloud with none of the bloat. "Screaming in the Cloud" listeners can try Vultr for free today with a $150 in credit when they visit getvultr.com/screaming. That's G E T V U L T R.com/screaming. My thanks to them for sponsoring this ridiculous podcast.

Corey: This episode is sponsored by our friends at Revelo. Revelo is the Spanish word of the day, and its spelled R-E-V-E-L-O. It means “I reveal.” Now, have you tried to hire an engineer lately? I assure you it is significantly harder than it sounds. One of the things that Revelo has recognized is something I’ve been talking about for a while, specifically that while talent is evenly distributed, opportunity is absolutely not. They’re exposing a new talent pool to, basically, those of us without a presence in Latin America via their platform. It’s the largest tech talent marketplace in Latin America with over a million engineers in their network, which includes—but isn’t limited to—talent in Mexico, Costa Rica, Brazil, and Argentina. Now, not only do they wind up spreading all of their talent on English ability, as well as you know, their engineering skills, but they go significantly beyond that. Some of the folks on their platform are hands down the most talented engineers that I’ve ever spoken to. Let’s also not forget that Latin America has high time zone overlap with what we have here in the United States, so you can hire full-time remote engineers who share most of the workday as your team. It’s an end-to-end talent service, so you can find and hire engineers in Central and South America without having to worry about, frankly, the colossal pain of cross-border payroll and benefits and compliance because Revelo handles all of it. If you’re hiring engineers, check out revelo.io/screaming to get 20% off your first three months. That’s R-E-V-E-L-O dot I-O slash screaming.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. For a while now I have resisted the siren song of doing an episode covering the wide world of Web 3. So, if you’re deep into that space, you can rejoice because it’s finally time to change that. Now, the other side of that, for at least some of you, is that my guest today is Molly White, who’s a software engineer. But more notable as of recent days for running a collection of interesting stories coming out of the world of Web 3, at web3isgoinggreat.com. Molly, thank you for joining me.

Molly: Thanks for having me.

Corey: So, by day, you’re a software engineer, which means you’re already predisposed to writing things that humans find very difficult to understand. And now you’re—in your spare time apparently—writing about Web 3, which is a topic that humans find very difficult to understand. For some reason, you have a flair for telling stories about this basically impenetrable to outsiders space in a way that makes it look, first off simpler to understand, and secondly—let’s be clear—patently ridiculous. How did you find your way into this part of the world?

Molly: Well, [sigh] I think as a software engineer, it’s a little hard to avoid the Web 3 thing. You hear about it from your colleagues or the people on tech Twitter, or, you know, you see it in the news. And it’s—

Corey: Or people behind you at Starbucks who won’t stop talking, et cetera.

Molly: Yeah, they sneak right up on you. And so people, you know, when you hear about something that’s supposed to be the future of the web, you know, if you’re a web, software engineer, I think it’s sort of natural to try to figure out, “Oh, what’s this thing? You know, I need to learn more about this.” And that’s sort of how I got into it. You know, I tell the story about how I’ve known about cryptocurrencies for a long time—you know, Bitcoin has been around for a while now—and I was just extremely uninterested [laugh] in them for a very long time.

But with this sort of rebrand, recently, I’ve sort of been forced, I think, to pay a little more attention to it since it seems to be one of those things that people have to engage with whether they want to or not. Or at least people hope that is what the Web 3 thing will become. So.

Corey: I come from a background of being a grumpy old Unix administrator, and I’ve been around long enough to see the inevitable cycle where this shiny, exciting new technology of today is the legacy garbage I have to support in production tomorrow. And this breeds more than a little bit of cynicism, where whenever someone says, “We have this new thing,” it’s, “All right. Let me get out the checklist. What happens when jerks get involved? How is it going to break? How am I going to hate this thing? How is it going to completely ruin my week?”

And people building technologies—this is probably no surprise—don’t generally like questions like that. And I get it because they’re trying to do something creative and build something that solves a problem that sometimes they can—they’re the only ones who can define. But also tends to be this sort of love with the technology where, “I see nothing wrong with the technology I’ve built whatsoever.” It’s, “Yeah, you probably wouldn’t.”

And that’s okay because that doesn’t bound itself to cryptocurrency or blockchain stuff—

Molly: Yeah, absolutely.

Corey: It runs the gamut from databases to messaging protocols to someone’s sketched-out version of an iPhone they wish that someone would build, et cetera, et cetera. No technology is perfect. There’s an ancient place on the internet I used to hang out that had the motto of, “All hardware sucks, all software sucks.” Varying degrees and to different levels, but they all suck. And they’re not wrong.

I love aspects of the Web 3 community: Their optimism, for example, is something that I find inspiring; their ability to stay on message is also incredibly—honestly, it’s admirable. I just wish the message were slightly different. What is your take on how all of this stuff is—I guess, not just the technology itself, but the hype train that accompanies it?

Molly: Yeah. Well, first of all, I agree with a lot of what you just said. I think optimists have a great role to play in the software world, and I think cynics also do, and I sort of wish there was a little more balance in this particular technology. I also agree with you the community and sort of the people who—a lot of the people who are trying to get involved in this stuff, I really admire and I think are really passionate, and, like, really smart and, you know, have the right motivations. But it’s a little frustrating sometimes to see that the optimism can turn into a very aggressive, sort of almost protectiveness around the technology where they are almost unwilling to, you know, examine whether or not there might be flaws behind the product that they’re hoping to build.

And that’s where I get really worried because I think in order to build software responsibly, you need to be open to the skepticism and the criticism and the questions, and it overwhelmingly has felt like the sort of Web 3 community has not been, which I find really worrisome.

Corey: Spare me from the cascade of, “Do your own research,” whenever you say something negative. Like, it seems to be a pervasive ail of our society, where it’s, “You’re just going to believe what people tell you?” “What, you mean, legitimate experts who’ve been looking at”—

Molly: [laugh]. Yeah, the experts?

Corey: —“The space for decades?” Yeah, it’s like, what research am I going to do on YouTube in 20 minutes that is going to outweigh that? It’s not do your own research; it’s carefully curate the bias of the media your consuming until you come around to my worldview, and that’s not the same thing as research.

Molly: I agree. Yeah. And it’s weird how we see that same thing cropping up in, like, Covid-19 conspiracies and QAnon, and then it’s like, “And also crypto.” Okay, that’s a little weird.

Corey: It’s very odd watching the rise of this. Blockchain is an interesting technology, absolutely, and this recent extension into non-fungible tokens, or NFTs, I—the first time I saw it was relatively recently and it very quickly became impossible to avoid, if for no other reason than I keep getting tagged by brand new empty Twitter accounts doing replies of, “Tag three people to wind up getting airdropped on this,” and I sure do love the fact that Twitter can find no way whatsoever to stop this from happening. Lovely. And it’s, “Okay, looking at this, what is this?” Like, “Look at how much money these things sell for.”

And I have extensive background in finance, so I can spot it from a mile away, like, “Oh, yeah, that. That’s a money-laundering scam.” Like, wait, that doesn’t seem fair. Like, you tell me everyone involved in NFTs is a money launderer? Oh, absolutely not; that’s a terrible money laundering scam.

You need to have people who are not money launderers, otherwise, the entire thing gets shut down and everyone gets arrested. You need to have people who are not themselves criminals, basically interspersed and the dominant party in this, so the rest of us can money launder. And it’s like there’s nothing new under the sun, and the idea that regulators are somehow complete naive fools does not usually pay dividends. And people have a long time to reflect on this in federal prison.

Molly: Right. Yeah. And I think we’ve been seeing this sort of trailing regulation coming in a little bit. You know, there’s a lag between when someone does something blatantly criminal and then when the, you know, the US Attorney’s Office announcements come out a year or two later saying that, “Oh, and we just charged this person with”—you know–“Fraud or whatever it is.” I sort of, every once in a while, I read some of those announcements from, you know, the Department of Justice or the various other groups and, you know, they’ll describe what someone has done in the crypto space or, you know, financial fraud, and it’s like, whoo, boy that looks similar to a lot of these projects that are just launching now. I wonder if anyone’s getting a little bit uncomfortable reading these. [laugh].

Corey: This thing that I understand as well that I am, I am a cynic and I have been basically down on emerging technologies a lot, which means I’ve been wrong an awful lot. In 2006 2007, I thought virtualization was a very niche thing that was only going to be suitable for a couple of weird workloads because, like, how many underutilized computers could there really be in the world? Yeah, I was wrong. Then I said that cloud was absolutely not going to take on, and through about 2012, I was very anti-cloud because, “Oh, you’re going to trust your stuff on someone else’s computer and give your uptime to them and your—their security over to them. It’ll never catch on.” Yeah, I was wrong there, too.

I thought containers were ridiculous. And they are in some ways, but they’re also the way the world works. I’m actually very bullish on serverless, which means it’s not going to succeed in the market because I’m invariably—

Molly: Oh, no. [laugh].

Corey: —wrong in this. Exactly. But people are saying, “Well, what makes Web 3 or any of these blockchain technologies any different than all the other things that I was wrong about?” And my feeling around this is that at least I could understand what those other technologies, what the problem they were setting out to solve was. It continues to shift depending upon the narrative line that people are pushing on this.

And I also remember the response I got every case previously, which was, “You’ll see it sooner or later. It’ll be fine.” There was never this urgency baked into it of, “You have to get in now, or you’re going to lose out and be poor forever.” And I was extraordinarily gullible growing up, so when I see this, it’s like, okay, when you’re trying to pressure me into doing something, it’s because you’re deriving some benefit if I do. And I’m very cynical these days, perhaps unfairly so.

When you grow up being constantly you may have to fall for pranks and whatnot because you have no sense of guile, then great. You sort of have an overreaction the other direction. It’s mostly served me well, but I look at this and okay, ignoring the entire bubble of that ecosystem—and I’m hoping you have an answer for this—what is the real-world problem that I, as an individual or as a business, have that Web 3 solves for me?

Molly: Well, it’s been great for ransomware. So if—

Corey: Oh, yes.

Molly: —you’re doing that—[laugh]. Yeah, no, I have a very similar feeling to this around, you know—

Corey: Ransomware is ridic—there’s already ways to do that. It’s like, we’re going to take your data and we’re not going to give it back to you unless you pay us enormous piles of money. Yeah, that’s called cloud egress charges. It’s been done, and it’s a lot less computationally intensive. I’m mostly kidding, but not entirely.

And yeah, yeah, it’s so much easier now to wind up extorting money from people through this thing. Yeah, I. Don’t find that often to be a feature, and, frankly, people who do I don’t really want them within a thousand miles of me.

Molly: Right. Yeah. And I mean, I think a lot of the problems—you’ll even see people saying this, you know, with very thin veils of legitimacy, but a lot of people are basically saying, “Well, we want to do something that we can’t do with traditional money because it’s regulated.” [laugh]. And so seeing that, it’s like, “Really?” You’re like, you know, the regulations—

Corey: What are you trying to do over there, buddy? And yes, I admit, there are certain things that I find obnoxious about the way that in, you know, non-crypto society that we deal with money and the challenges we have with it. Some of the fees attached to things, some of the way that it takes, “Wow, we can send messages at the speed of thought in real-time, but it still takes how long for a payment to clear through these systems?” I get it; there are reasons they are the things are the way that they are. But it all mostly works, to be clear.

Are there opportunities for improvement? Absolutely. Do I think that the way to do that is to basically come up with an entirely new form of money? I—maybe if you’re starting from scratch, but I kind of have a hard time accepting that it’s going to work that way for everyone.

Molly: Right. And I also think there’s sort of this pervasive issue with a lot of the projects in Web 3, where they are actually trying to solve very real problems, very serious problems, and you know, the fact that there’s, you know, unfairness in the banking system, or that there’s fees, or that, you know, there are people who are making an enormous amount of money off of people just trying to send small amounts of money, like, I get that, and I get that you might want to solve those problems. But overwhelmingly, it seems like there’s sort of this opinion of like, okay, so we have this bad thing now. We have this different thing here, so this different thing has to be better than this bad thing. And it’s like, no, no, no, wait, hang on. It’s possible to, like, replace a bad thing with something that’s worse, and I think we need to consider that what we’re trying to do here looks a lot like that.

And so you know, people are talking about, you know, “Oh, well bank the unbanked.” They don’t have access to banking and so we’ll fix that with blockchains. And it’s like, no, I think what we’ll do actually with the blockchain is we’ll probably end up scamming the unbanked because this place is totally unregulated and regulations actually protect people a lot of the time. You know, so that I think that’s really worrisome, the sort of just idea that we have something different and so it’s better.

Corey: The thing that always catches my eye when people talk about this: “Oh, it’s the new form of money. It’s going to solve all of the social injustice problems.” Okay, maybe I stumbled upon this secret hidden community of altruists that are out there, but generally speaking, looking at the broad sweep of human behavior, you can make a few observations, and one of them is that the rich generally do not desire company. And the idea of, oh, this is going to magically fix systemic inequality, I don’t know that that’s necessarily true.

Molly: Right.

Corey: And a lot of the pr—say what you will about problems with the existing financial regulations that are out there if I screw up, and I accidentally wind up doing a wire transfer of rent or for buying a car to the wrong account, there are established ways that gets reversed, and between large institutions, it’s basically a phone call, a letter, and it gets done within a day or so. Whereas with crypto, it’s [sings] doesn’t it suck to be you? Di di. And that just becomes… well, is there any recourse? None. That doesn’t strike me as a feature, to be honest, that strikes me as a bug.

Molly: Right. And I was actually—it is interesting. I was recently rereading the Bitcoin white paper because it’s one of those things that people are constantly like, “Well, read the Bitcoin white paper and you’ll totally understand it all.” And it’s interesting how in the Bitcoin white paper, they talk about how this new system will prevent fraud, but if you look at what they’re talking about as fraud, they’re talking about people illegitimately reversing transactions. So like, you know, take the example you buy something on Amazon, you receive whatever item it is, and then you do a chargeback. And you then you have your item, and you haven’t paid for it.

That is, like, the one thing that this, you know, person who came up with Bitcoin is describing as fraud. And that’s, like, the one thing that is hoping to be prevented. And it’s like, I kind of get the idea that, like, [laugh] at some point, you know, someone’s scammed, Satoshi in this way, and it’s just, like, this is what came from it. But it’s such an odd perception that is, like, the only kind of fraud and, like, that is always a bad thing to be able to reverse a transaction. I find that really fascinating because it’s just like, that’s actually a really good thing a lot of the time.

Corey: This episode is sponsored by our friends at Oracle Cloud. Counting the pennies, but still dreaming of deploying apps instead of “Hello, World” demos? Allow me to introduce you to Oracle’s Always Free tier. It provides over 20 free services and infrastructure, networking, databases, observability, management, and security. And—let me be clear here—it’s actually free. There’s no surprise billing until you intentionally and proactively upgrade your account. This means you can provision a virtual machine instance or spin up an autonomous database that manages itself, all while gaining the networking, load balancing, and storage resources that somehow never quite make it into most free tiers needed to support the application that you want to build. With Always Free, you can do things like run small-scale applications or do proof-of-concept testing without spending a dime. You know that I always like to put asterisks next to the word free? This is actually free, no asterisk. Start now. Visit snark.cloud/oci-free that’s snark.cloud/oci-free.

Corey: But I will defend the Web 3 community, which I know is somewhat surprising because again, my feelings on this stuff are nuanced. But—

Molly: Sure.

Corey: Everyone says they have this massive problem with InfoSec and the rest, and I don’t believe that that is necessarily true. I do not believe that the people writing the code that powers these blockschain—or however the pluralize is improperly—are somehow much worse developers than everyone else. But the incentives are radically different because if I screw up on some of my Lambda functions, great, you can get access to I don’t know the API tokens for my lasttweetinaws.com Twitter client. Okay, great. Now, you can spam Twitter. It’s not that interesting to people and it’s not considered high value.

Whereas yeah, if I wind up breaking through this little thing, I can wind up getting, what, $200 million? Yeah, suddenly, it’s probably worth spending significant time on security reviews. So, I do think that folks are being a little unfairly maligned there just because the way that they’re approaching this it does not match the rigor that is take—and care that is taken to systems that in the fiat finance world—as they love to call it—that wind up [unintelligible 00:20:10] money, there’s oversight, there is planning, there is testing, there are entire teams of people doing nothing other than InfoSec review, rather than, “Well, it’s on GitHub; my job is done.”

Molly: Yeah. Yeah, I’ve heard people refer to it as self-paying bug bounties before where the bounty is, you know, the money that you can pull out of these exchanges or whatever project you actually, you know, are able to exploit. And I think that’s very accurate. And I think you’re right, you know, I think that there’s nothing that—I mean, there, I’m sure they are particularly incompetent developers in Web 3 as there are particularly incompetent developers in any sector, but I do think that you’re right, that it’s just an enormous incentive to find any small bug. And I think also part of it is that a lot of the concepts that people are working with are extremely difficult to, sort of, wrap your mind around.

You know, this is a little bit of a tangent, but a lot of the attacks, we just saw three attacks in one day that all relied on something called a flash loan exploit. And trying to wrap my head around what a flash loan is just like, it doesn’t jibe with, like, current financial systems and so it’s really hard to, sort of, comprehend, and I think it’s probably hard for developers to code against because it’s just a very different way of thinking about loans. You know, like a flash loan is basically a loan that you take out the loan, and you pay it back in one transaction, which, in real life has no purpose, right? There’s no reason you would go to a bank, borrow $10,000, and then immediately give them those $10,000 back. But you know, [crosstalk 00:21:47]—

Corey: Financial equivalent of a managed NAT gateway that winds up just transferring for every cent that’s passed through it. I’ve seen stuff historically, before they fixed bugs like this, in credit card reward systems, where basically you can just cycle spend through and it doesn’t do anything other than suddenly starts cranking your point balance into the stratosphere, so you could save up your poi—frequent flier miles to go to space or whatnot.

Molly: Yeah, exactly. Right. And I think, you know, that’s the, sort of, same idea here. You know, people use these flash loans for all sorts of weird, you know, yield farming and just sort of bananas stuff. And, you know, I think trying to code against a lot of stuff, you have to really understand those things very well, and not necessarily just be a good developer, but also understand the economics behind it and the incentives that people are, you know, chasing. And that’s tough. [laugh].

Corey: I will say that you are far from alone in criticizing crypto, but I’ve patterned a lot of my own cynicism and trepidation around the space after the way that you engage with it. And what I mean by that is not that I build hilarious websites about these things that chronicle its shortcomings, but rather that you don’t personalize it, you don’t take the step that so many folks do and say, “Oh, this person is now going to work at a crypto company, therefore, they’re a sellout. Therefore, they’re out to scam people. Therefore, they’re just the devil incarnate.”

And it’s no, I don’t believe that either. I’m curious to hear their reasons for it. They don’t owe me an explanation and I’m certainly not going to harass them on Twitter about these things, but the idea that someone is somehow now working for a company that engages in this stuff, and therefore they are now to be written off as a human being is something that I just find distasteful in the extreme. And I’ve never once seen you cross that line.

Molly: Yeah, I also really disagree with that. Which, you know, may be controversial to some of my fellow skeptics, but I think we can agree to disagree on that. I don’t think that it is, you know—I think that people have very good reasons for going to work for companies that I don’t necessarily personally agree with, you know? And I think there are a lot of examples of people who work for companies in spaces that are, you know, questionable. A lot of the big social networks have done things that are pretty horrifying when, you know, look at the recent exposés around Facebook or, you know, all those things.

You know, there are people who work for defense companies, which I don’t necessarily agree with, you know, those kinds of things. And I think everyone has to do their own, sort of, personal math around what makes sense for them, where their ethics lie. You know, a lot of these companies I will say pay a lot of money, and I can’t necessarily fault someone for needing to pay the bills, right, even if it means working for a company that I think is maybe not the best. But I think there’s—

Corey: Yeah, I used to give people who worked with Facebook tremendous amounts of crap. I don’t do that anymore. I was wrong. I’m not going to personally harangue people for where they work. You never know someone’s individual situation. I—

Molly: Exactly.

Corey: I’m not apologizing for the company; I want to do no business with them, but I will no longer be going after people individually because they work there. Because until you walk a mile in someone’s shoes, you don’t know what’s going on there.

Molly: Right. And I also think there’s just not much point to it, right? Like, if we want to hold Facebook to account, for example, I don’t think going after some software engineer or customer support rep or whatever is going to make any difference, other than making their life particularly unpleasant. And, you know, that I think applies to the Web 3 crypto space as well. You know, I will absolutely dunk on someone who I think is, you know, malicious and scammy and taking advantage of people, and I will say the same things about companies that are doing that, but I do think that there are very well-intentioned people who are working in this space for a ton of different reasons.

It’s, you know, personal curiosity; some people just aren’t convinced yet that—you know, some people think this could do a lot of good and that they, you know, should engage in the space in good faith and, you know, go work for these companies, and try to make sure that the companies are pushing towards good. You know, I don’t personally think that there’s much that can be done there; I think that’s a tough angle, but I respect people for trying. And I think there’s also a huge amount of just, uh—I think a lot of the hate or the vitriol that has been targeted at these people who are going to work for crypto companies is very selective, in some ways, you know? You see a woman going to work for a crypto company or person of color going to work for a crypto company and their replies look markedly different from the white guy who goes to work for a crypto company. It’s all, “Congratulations,” and, “Oh, he’s going to be so rich,” and all that kind of stuff, and there’s not so much you know, hand-wringing of hands, whether or not they are a scammer or all that kind of thing. It’s like, you’re only allowed to, you know, go and get that bag or whatever if you’re a white guy; everyone else is held to a different standard.

Corey: For people who look like me, the bars on the floor, let’s be very clear here. It’s, “Good for you. Go after it. Go and get it.” And there is a systemic problem, on some level, that I think that we have not really grappled with as a society, which is that even well-paid software engineers still feel the pressure that in order to be prosperous and guarantee financial security for you and your family, you now also to be a part-time trader in various ways, and invest, which very often is misused in place of what it actually is, which is speculation or gambling.

Molly: Gambling. Yeah. [laugh].

Corey: That is the way to prosperity. Because we have survivorship biases; no one likes to trumpet their failures. It’s the same problem we see with tech conferences: People get up and talk about, “This is the thing we built and it’s awesome.” And you talk to people who work there. It’s like, “Yeah, I don’t recall that project going anywhere near that well.”

And, yeah, it we all tell these aspirational, heroic stories of what we’ve done, and we trumpet the things we’re proud of. And it just, it isn’t sustainable. It isn’t something that I think we’ve spent a lot of time on. And this is software engineers were talking about. Remember, once growing up, at least, there was the idea that you could—this was this wild, subversive idea that you could be a schoolteacher in a city and actually live in the city in which you taught. Now, that is basically a fantasy. And we see that across the board. That’s not great for anything.

Molly: Yeah. And I actually blame, you know, economic circumstance for a lot of the crypto hype. You know, there are a lot of people who are in tough spots right now, you know? The pandemic has certainly had a huge impact on some people, especially people working in, you know, service jobs and things like that. And so people are in, you know, pretty dire straits as a result of that.

There’s also enormous student loans, the housing market is bonkers, you know? There’s so many things that are really making people suffer financially. And so then when crypto comes along, and people start talking about 60,000% APY and all this, you know, you’re going to triple your money, you’re going to buy this bored ape NFT at $100 and it’s going to be $500,000 next year. People fall for that because it’s enormously appealing, right? And I think there’s a lot of blame to be placed on the media for, sort of, buying into a lot of that.

There’s been a lot of very credulous reporting, I would say, on some of the people who claim to make a ton of money off of these things. And so people, you know, that, when they see, you know, CNBC, for example, will highlight these, you know, people who were just scraping by, they were going paycheck to paycheck, they put $50 into a project and now they’re millionaires, you know? And people see that because there’s no point for CNBC to talk about the person who invested their, you know, their rent payment into a crypto project and then couldn’t pay rent because they lost it all, you know? Or the person who took out margin loans and is now in debt to these various companies that are lending people money to gamble on crypto. Those are not the stories that make the headlines and so people get a very skewed view of how many people are actually making a ton of money in this space and how many people are actually losing a lot of money in this space. And I put a lot of blame on various media companies for that.

Corey: Well, take Twitter as an example. Yeah, I would classify them in many respects as a media company. Imagine for a second that if every time someone tweeted something about AWS, like, “Well, I got surprised at my AWS bill,” or, “Huh, I’m having some trouble with AWS Lambda.” Suddenly, 15 bots all replied and quote tweets and the rest saying, “Ah, this person helped me out. Talk to them,” or fake accounts with a, “Here’s our support forum. Please fill this out.” It goes to a Google Doc.

It seems like the easiest thing in the world to automatically wind up detecting and blocking—just because it is clearly keyword triggered, it is very obvious when it happens, and somehow, it just keeps persisting. It makes you wonder, on some level—like, counts as engagement and users, so it makes the numbers that we report on earnings go up, so I guess we’re going to keep doing it. It just feels like, on some level, Twitter has empowered a lot of this in a way that most normal places would not.

Which, of course, brings us to the other project you’ve been involved with for a very long time: Wikipedia. Now, it seems like a weird thing to say, “Oh, yes. You’ve been an editor on Wikipedia.” Yeah, so is basically everyone at some point because it turns out, it’s a couple of clicks away. Your something a little more than that, but I don’t pretend to understand the Wikipedia [instructor 00:31:31]. Tell me about that.

Molly: Yeah. Yeah. So, I’m a Wikipedia editor. I’m a prolific, I guess, Wikipedia editor, you might say. But I’ve been actively editing Wikipedia for over a decade now. I’m also a member of the, sort of, administrative group on that project, and I’ve also served a couple of terms on what’s called the arbitration committee, which helps mediate disputes among community members. So yeah, I’m pretty involved. [laugh].

Corey: How much of your involvement in that community has bled over into your, frankly, amazing coverage of Web 3?

Molly: An enormous amount. [laugh]. I think you can very—I think, if you look at the entries on Web 3 is Going Great, you can kind of see the Wikipedia voice in them. It’s a little hard for me to escape that sort of style of writing because I’ve just been doing it for so long, and it’s the majority of the writing I do. So, you definitely see that a lot.

And, you know, I’ve had a couple of people say things, you know, like, you know, “How do you cover stuff in such a, you know, detached way?” And it’s like, “Oh, well, I write encyclopedias in my spare time.” [laugh]. There’s obviously a lot more sarcasm and sort of personal bias in the Web 3 is Going Great project, which is why I started it because I can’t do that on Wikipedia, and I won’t do that on Wikipedia. But that’s where a lot of it comes from, is that sort of that instinct, I think, that you develop as a Wikipedia editor to sort of research and chronicle and record and share what you’re seeing. It’s hard to escape.

Corey: I do want to call attention, though, to other long-form writing that you do because, “Wikipedia, who wrote this?” Well, the answer is always lots of people. But if you go to mollywhite.net and look at your long-form writing, it’s pretty easy to understand who wrote this.

It’s not nearly as clinical and encyclopedaic as you might expect, from your description just now. It’s very approachable, very engaging, writing that reflects on topics in a way that only long-form can and Twitter generally cannot. And it’s great, you could read this and not realize that you’re deeply involved in the Wikipedia part of it, right up until the point you get to the end. And then you see the extensive list of references at the bottom of the page because apparently footnotes and citation is a habit that you can’t get away from there, but it’s—

Molly: I can’t help myself. [laugh].

Corey: Nowhere near as dry and clinical as you’re implying.

Molly: Yeah, that’s true. I do take more of an essay approach in my long-form
writing. One thing I’ve really loved about Web 3 is Going just Great is that you sort of don’t need to necessarily know, like, what’s a blockchain, and what’s an NFT, and what’s, you know, distributed, you know, database or whatever, before you start reading it. It’s sort of approachable, you can read one or two entries, and then you can go do whatever else and you don’t have to do this, sort of, deep dive. But it also lacks, I think, that ability to go a little deeper into some of the problems or some of the really huge issues I see with the underlying technologies because it’s, you know, it’s very much of one hit, and then you move on to the next thing.

So, I’ve started blogging a little bit on the side, I guess, to sort of go into a little more depth on some of my concerns, just both as, like, a technologist but also just as someone who’s been on the internet for a long time and who’s been a member of communities. You know, the Wikimedia community is very similar to a lot of the communities that you’re seeing crop up in these Web 3 projects, especially to do with the DAOs. And so, I sort of have over a decade of experience in it community like that, and I’m watching a lot of these new DAOs, you know, who say they’re coming up with this brand new model, you know, they’ve invented this new form of governance. And I’m watching them, it’s like, “Oh, you’re about to step on that same rake we stepped on 15 years ago.”

And that concerns me a lot, especially because you know, the Wikimedia
community, there’s harm that can be done by—for sure, and it has happened, but there’s not really financial harm that happens with a Wikipedia editor, you know? You’re not buying to engage with the Wikimedia community. I certainly hope not because you’re being scammed. But with these DAOs, you know, you’re paying to engage with a community that is not taking lessons that it could be taking from Wikimedia, from co-ops, from mutual organizations, you know? They could be looking at history a little bit more, I think.

Corey: Tech does this across the board. It’s, “We’re in San Francisco. We’re going to reinvent and disrupt industry x.” Okay, fine, great. Maybe it works. Maybe it doesn’t. Godspeed. “And while we’re at it, we’re going to reinvent other things, too, that we think the world gets wrong, like, how to interview people.” And the common thing on Twitter is no one knows how to interview engineers properly; it can’t be solved.

And yes, yes it can. There were multi-decade studies conducted in places like GM, Coca-Cola, et cetera, on how to lead to positive outcomes while interviewing and what to do, and whenever you bring that up, Twitter gets very angry about that because, “No, that’s different. That’s a different time and a different era, and the world works differently, now.” And, “Great, okay, keep disrupting things, but you can save a lot of time by having a conversation or two with people who’ve walked that road before? You don’t have to go it alone.”

Molly: Right. Yeah, I think that’s a huge thing. Kelsey Hightower has done a lot of conversations around that, around how—he’s doing incredible work talking about, you know, blockchains, and crypto and stuff—and he’s talked a lot about how it looks like a lot of these projects are sort of reliving history around—you know, he has a very technological approach to it, so he talks about, you know, the sort of security things that are not being considered and the, you know, the various infrastructure sides of things that are just sort of being reinvented without any sort of consideration to the past lessons. I think that’s just a very classic—you’re right, it’s a very classic, like, Silicon Valley way of doing things. There’s sort of the running joke about how people reinvent buses every couple of years.

You know, Uber is like, we’re going to make a service where a bunch of people can all get in a car together and drive someplace. It’s like, “Oh, right, yeah. A public bus system.” We have those. I think there’s very similar comparisons to draw in Web 3.

Corey: There absolutely are. And I want to thank you for being so generous with your time. If people want to learn more, where can they find you?

Molly: You can find me on Twitter. I am both @molly0xFFF, and also @web3isgreat on Twitter. And then there’s my website, web3isgoinggreat.com, and my other website, mollywhite.net. I’ll be at all of those.

Corey: Yes. And for fun, I wound up pointing a domain over to you, over to your site as well, ponzl—P-O-N-Z-E-L dot com. It’s like P-O-N-Z-I except the I is an L because the crypto people can never seem to quite take the L. But there you have it. It’s not a Ponzi scheme; it’s something different.

Molly: It’s a Ponzl scheme.

Corey: Thank you so much for being generous with your time. I appreciate it.

Molly: Thank you for having me.

Corey: Molly White, Web 3 chronicler of our time, and software engineer. I’m
Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry comment telling me that no, crypto is not reinventing a bus because a bus can only run over someone once.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Johnny

Johnny was born in Cleveland, OH and graduated from the University of Toledo with a Bachelor's in Computer Science Engineering. He began his career as a software engineer focused on embedded device protocols and systems engineering. Eventually he realized that Program Management worked better with the grain of his brain, so he took his career in that direction.

In 2019, he was hired by Google Cloud to serve as a Communications Lead on their incident management teams. Most recently, he joined Waymo in November 2021 as a Technical Program Manager, acting as an anti-entropy agent for the self-driving car company's offboard infrastructure teams.

Outside his day job, Johnny enjoys mountain biking, playing piano and trumpet, personal finance, coaching, and studying complex systems. He currently lives in Sunnyvale, CA with his wife Emily, and is expecting their first child in April 2022!

Links:

  • Original Twitter thread: https://twitter.com/QuinnyPig/status/1436129343399346184
  • Personal website: https://jmpod.com
  • LinkedIn: https://www.linkedin.com/in/jmpod
  • Twitter: https://twitter.com/gratitudeisfree/
  • Instagram: https://www.instagram.com/gratitudeisfree/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Couchbase Capella Database-as-a-Service is flexible, full-featured and fully managed with built in access via key-value, SQL, and full-text search. Flexible JSON documents aligned to your applications and workloads. Build faster with blazing fast in-memory performance and automated replication and scaling while reducing cost. Capella has the best price performance of any fully managed document database. Visit couchbase.com/screaminginthecloud to try Capella today for free and be up and running in three minutes with no credit card required. Couchbase Capella: make your data sing.

Corey: This episode is sponsored in part by LaunchDarkly. Take a look at what it takes to get your code into production. I’m going to just guess that it’s awful because it’s always awful. No one loves their deployment process. What if launching new features didn’t require you to do a full-on code and possibly infrastructure deploy? What if you could test on a small subset of users and then roll it back immediately if results aren’t what you expect? LaunchDarkly does exactly this. To learn more, visit launchdarkly.com and tell them Corey sent you, and watch for the wince.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Every once in a while I get feedback from people who I’ve encountered who are impacted in various ways. Most of it is feedback delivered of the kind you might expect, like, “Unsubscribe me from this newsletter,” or, “Block,” or sometimes bricks thrown through my window. But occasionally, I get some truly horrifying feedback, and far and away one of the most horrifying things I can ever be told is, “So, I was reading one of your tweet threads and it changed the course of my career.”

It’s like, “Oh, dear,” because nothing good is going to happen after something like that. It’s, “Yeah, they were going to name
something terrible here at AWS, so I ran over my boss in the parking lot,” is sort of what I’m expecting to hear. But I got that exact feedback about life-changing tweet threads from today’s guest. We’ll get into what that tweet thread was a little bit, but let’s first let the other person talk for a minute. Johnny Podhradsky is a technical program manager at Waymo. Specifically, of Offboard Infrastructure. Johnny, thanks for suffering through a long, painful introduction, as well as, more or less, the slings and arrows that invariably come with being on the show.

Johnny: Thanks, Corey. I’m grateful to be here.

Corey: So, first things first. I always like to find out what people actually do for a living that is usually a source of entertainment, if nothing else. You are a technical program manager—or TPM as they say in tech companies—of Offboard Infrastructure. I’m assuming because Waymo, is at least theoretically, a self-driving car company, ‘offboard’ means things that are not on the vehicle themselves.

Johnny: That’s exactly right. Yeah.

Corey: Fantastic. Now, ask the dumb question because I’m still not sure I have an answer after however many years in this industry. What does a technical program manager do?

Johnny: [laugh]. I get that question a lot. Often people try to distinguish between what’s a technical program manager do versus what does a product manager do.

Corey: Or a project manager, too, because there’s a lot of different ways it can express itself, and I’m a PM, and it’s, “Oh, wonderful. That’s like four different acronyms I can disambiguate into and I’m probably going to get it wrong.”

Johnny: And to make it even more confusing, it varies company by company. So, just focus in on specifically what I do as a technical program manager, I’m an anti-entropy agent, right? I make sure things stay on track, specifically embedded into technical teams. So, I have a degree in engineering; I’m able to speak fluently about technology. And the entire idea, the entire purpose of my existence is to make sure that things don’t fall apart. So, I’m keeping track of people and resources; I’m keeping track of overall timelines; risks and mitigations for programs that are ongoing, whether they’re small with just a few people or cross-org, cross-functional teams; serving as an unblocker and making sure that all the dependencies that exist between the various tasks in the teams are addressed ahead of time so that we know what needs to be done when.

Corey: It’s one of those useful almost glue functions, it feels like that is, “Well, what have you actually built? Point at the thing you’ve constructed yourself from your hands on your keyboard?” And it’s hard to do and it’s very nebulous, when you’re not directly able to point to a website, for example. “Yeah, you see that button in the corner? I made that button.” Great.

Like, that’s the visceral thing that people can wrap their heads around. Project and program management feels to me like one of those areas that, in theory, you don’t need those people to be a part of building anything, but in practice you very much do. Another example of this—from my own history, of course—is operations because in theory, you just have developers write code correctly the first time and then they leave it where it is and it never needs to be updated again, and there’s no reason to have operations folks. Yeah. As they say, the difference between theory and practice is that in theory, there is none.

Johnny: I’ll buy that. Yeah, when it comes to actual, I mean, digital, but physical deliverables and things that you can show that you’ve done, there are standards that you can have with documentation, like Gantt charts and risk registers and all that sort of thing, but it is very much a glue role. It is very much a gentle nudge to get things done. And it really revolves around the transparency and making sure that the people who are invested in the success of whatever it is that you’re doing program-wise are aware of what’s going on as far ahead of time as possible. That’s why I like to consider it sort of an anti-entropy role because things will just naturally go off the rails if no one is there to help guide them.

I mean, that doesn’t happen in every situation, of course, but having someone dedicated to the role of making sure that things are moving according to a good rhythm is a critical role. And it just so happens that that is sort of the way the grain of my brain works and I discovered that throughout the course of my career.

Corey: So, let’s get back to the reason you originally reached out to me. I think that is always an interesting topic to explore because whenever someone says, “Wow, your tweet really helped me with my career,” I get worried. Because as I said before, I am one of the absolute best in the world at getting myself fired from jobs, so when it comes to being a good employee, mostly my value is as a counter-example of advice I’ll give [unintelligible 00:05:49] job interviews. For example, when they say something condescending and rude, insult them right back because A, it’s funny, and that plays well on Twitter. And B, interviews are always two-way streets, and if they’re going to treat you like crap, you don’t want to work there anyway, so you may as well have some fun with it. But a lot of what I say doesn’t really lend itself to the kind of outcomes that lead to happy employment scenarios. So, I’ve got to ask, what the hell did I say?

Johnny: Yeah, it was kind of serendipitous. I’m in a number of Slack communities, one of them being the Cleveland Tech Slack—if you’re in Cleveland or around Cleveland, I highly recommend it—and someone just randomly posted this thread right in the middle of me interviewing at Waymo. So, previously before Waymo, I was at Google, and I loved my job. I loved the team that I was on, I loved the—I mean, I was still very much in the honeymoon phase of Silicon Valley. I had moved to Silicon Valley from Cleveland in 2019 with my then fiance.

And so I was just, you know, bright-eyed and bushy-tailed, and everything was just incredible to me; why would I ever consider leaving this? So, I had an interview at Waymo and I ended up getting an offer and I just didn’t know whether I should take it. Because I loved where I was at and I really enjoyed the opportunities, so it was just, you know, ten out of ten. One of the things that I was thinking about then was, you know, I kept thinking back to our first team dinner where our teammates were sharing their stories of their careers. And my mentor, Ted, had mentioned how he had worked on the iPhone at Apple and was in the same room with Steve Jobs.

And me being a Cleveland boy, just it sounded like, “Whoa.” My eyes got really big like dinner plates. And it’s just like, “I’m sitting at a table with people who have done these things with these people.” And I was wondering, like, what did that mean for my career? And so where did I want to take my career and have those kinds of stories? So fast-forwarding, you know, I was interviewing at Waymo; I ended up getting the offer. And I was just on the fence; I couldn’t decide if that was the way I wanted to go, if I really wanted to leave my amazing job at Google.

Corey: What was holding you back on that? Was it a sense of well you want to be disloyal to the existing team? You were thriving in the role you’re in? Was it the risk of well, I don’t know how I’ll do in a different company solving different problems? What was it that was holding you back?

Johnny: It was all of those. When you do an apples-to-apples comparison, you don’t really know what you’re getting into when you’re going to a new company, and that’s part of why your thread was so critical in making my decision. Just to say exactly what you said in the tweet, “So, an anonymous Twitter person DM’ed me this morning with a scenario. Quote, ‘I work at a large cloud company that makes inscrutable naming decisions, and I have an offer elsewhere for 35% more. Should I take it?’” to which you said, “Oh, good heavens, yes. A thread.”

What followed is a number of questions that you asked exactly like you just asked now and your short answers to them. And they were just so on point and so quick, and it was so serendipitous for me to see that because this ended up being the tipping point that made me decide that, yes, this is the direction that I want to go. And you know, I’m—let’s see, I started in November, so five months into the role. It was more than I ever expected; it’s harder than I ever expected, but I’m growing so much, I’m getting a ton of eustress, if you’re familiar with that concept of the positive stress that makes your muscles grow. And just wanted to give back to you and in thanks and gratitude for being that tipping point. And that thread definitely led me down this path, so thank you for that.

Corey: It’s interesting because so far as of this recording, there are no two podcast episodes that came out of that thread because, to be clear, this was the thread-summary of a half-hour conversation I had with the person who messaged me about whether or not she should take the role. Because her manager had gone to bat for her to give her a raise and… yeah, she wanted to be loyal and show thanks for that. Which I get, but the counterpoint to that is okay, you turn down the offer out of loyalty. Great. A month goes by.

Now, your manager tells you that he or she is leaving to go work at a different company. Well, that opportunity is gone. Now, what? When it comes to career management, you can’t love a company because the company can’t ever love you back. And I got some pushback on that from Brian Hall, the VP of Product Marketing at Google Cloud—something about Google seems to be inspiring feedback on this one—because he spent something like 20 years at Microsoft and learned how to work within an organization, and then transfer jobs a couple of times to Amazon, they tried to non-compete lawsuit him on the way out—because, I don’t know, his PowerPoints were just that amazing or something, or they’re never going to replace his ability to name services badly—who knows why.

But he took the other position on this. And I’m not saying that my way is always right, it is provably not, as a self-described terrible employee, but it really is interesting that that’s the thing that resonated the most. I take a very mercenary approach to my career and I’m not convinced that’s at all the best way, but when someone dangles a significant opportunity in front of you, I always take the view that it’s better to explore and learn something about yourself if it appeals and the rest of the stars tend to align. And there’s a certain reluctance to go out and try new things, but it’s not like you’re leaving your family. It’s not like you’re selling out people who’ve come to depend on you.

Employment is fundamentally a business transaction and the company is never going to be able to have any sort of feeling for you, so you shouldn’t necessarily have this sense of loyalty, and oh, it’d be it would leave the team in the lurch if I left. That is the company’s problem to deal with. No one is irreplaceable.

Johnny: Yeah, and a lot of times when you were talking there, you talked about ‘the company, the company,’ but really, it’s the people that you’re working with that—and that was really what was weighing on me the most. I found myself in the same position. I had just recently gotten promoted. You know, my manager, and my team had gone to bat for me a lot, and so it’s hard for me to walk away. But it was ultimately the strong relationships that I had built with the team and my managers over time that allowed me to make this step because as a program manager, I’m always thinking that anything I work on needs to survive multiple generations of stakeholders.

So, everything that I do on a day-to-day basis has a breadcrumb trail, so that, hey, if I were to get hit by a bus tomorrow, someone with minimal amount of effort, can pick that up and move forward. And I’ve actually built that mindset into my entire career. Walking away from a role, you know, it’ll always leave a gap, it’ll always be challenging for the people and the teams around you, especially if you, you know, have a great affection for them, but by setting myself up to exit and still being there, since you know, Waymo is within the Alphabet companies and I can still talk with my old team, it wasn’t like I was completely leaving; I was kind of still there if I needed to be, if they needed help or needed to find something. But I can definitely see what how that would be challenging moving to a totally different company. But yeah, it’s really important that if you’re thinking about exiting, you have a good exit plan. And I’m all about planning as a program manager, and that just helped kind of grease the wheels a little bit.

Corey: I want to call it my own bias. You’re right, I use the term team and company interchangeably because that’s been my entire career. I, right now, have 12 employees here at The Duckbill Group and it is indistinguishable for me to make any meaningful distinction between team and company. Personally, I’m also not allowed to leave the company, given that I own it, and it looks really bad to the rest of the team if I decide, yeah, I’m going to go do something else now. People don’t like playing games with their future.

You’re on the exact opposite end of a very wide spectrum. It’s not that Google slash Alphabet is a big company, but you went from working on cloud computing to self-driving cars and you didn’t leave the company, you’re still at the same place as far as the benefits, the tenure, the organization, the name on the paycheck in all likelihood, and a bunch of other niceties as well. It almost presents is looking a little bit more like a transfer than it does leaving for a brand new job slash company.

Johnny: It definitely was a soft landing to go from Google to Waymo. There were a lot of risks—again, talking about risks and mitigations—that I was concerned about that we’re just kind of alleviated by the fact that okay, you can keep your same health care plan and various other things. So, that made it a soft landing for me. But yeah, it really was just making sure that the thing that I was working on at Google was able to be carried forward by the team and the people that I really enjoyed working with. So.

Corey: As you went through all of this, you said that you were in Ohio before you wound up taking the job at Google—

Johnny: Yeah, Cleveland [crosstalk 00:14:22].

Corey: —and one of the best parts about Ohio [unintelligible 00:14:22] family and spending time there is you get to leave at some point. And—

Johnny: [laugh].

Corey: There was a large part of that of, great. I felt the same way growing up in Maine, let’s be very clear here, where when I came to California, it was going to this storied place out of legend. And that was wild. And once your worldview expands, it feels very
hard to go back again. At least for me.

It took me years to really internalize that if this particular job or this particular path didn’t work out, my failure mode—if you want to call it that—was not and then I return to Maine with my tail between my legs and go back to the relatively dead end retail fast food job that I was working before, comparatively. No. It’s like, you go in a different direction; you apply the skill set; you have the stamp of validation on you. I mean, you have something working for you that I never did, which is the legitimacy of a household name on your resume. Whereas you look at mine, it’s just basically a collection of, “Who are they again?” And, “You make that company up?”

Which, fine, whatever. There’s a bias in tech—particularly—towards big company names because that’s a stamp of approval. You’ve already got that. The world is very much your oyster when it comes to solving the type of problem that you’ve been aimed at. I’m used to thinking about this from a almost purely technical point of view.

It’s like I’m here to write some javascript—badly—and I can write bad JavaScript for you or I can write bad JavaScript for that company across the street, and everyone knows what it is that they’re going to get from you: Technical debt. Whereas when you’re a technical program manager, that is something that you said varies from between company to company. And you hear founders talking about, “Oh yeah, our first engineering hire, we’re going to bring in a VP of engineering; we’re going to bring in a whole bunch of engineers; it’s going to be great.” You very rarely hear people talk about how excited they are like, “Oh yeah, employee number three is going to be a technical program manager, and we’re going to just blow the doors off of folks.” Which haven’t been through the growth process myself, yeah, we really should have had a technical program manager analog far sooner; it would have helped us blow the doors off of competition. And great, the things we learn, but only in hindsight.

Articulating the value of what a software engineer does is relatively straightforward, even for folks who aren’t great salespeople for their own work. Being a TPM inherently requires, on some level, a verification that your understanding and the person that you’re talking to are communicating about the same thing. Like, if you wind up having to solve code on a whiteboard, maybe that is part of your conception of it—I mean, you work at Google, probably—but for most companies, it’s yeah, my ability to write shitty JavaScript is not the determining factor of success in a TPM role. How do you go about even broaching that conversation?

Johnny: So, part of the way that program managers can be successful is through anticipating what’s coming next and understanding not only the patterns that were implanted over time, but also thinking ahead. And this actually kind of takes me back to why I learned program management in the first place. Pretty early in my life, I started feeling a great deal of anxiety, especially thinking towards future situations, or, you know, even in the present moment. I mean, we’ve all been through it right? Right before the big test, you’re feeling anxious; maybe talking to your crush—or before you talk to your crush—you’re feeling this anticipatory anxiety; in hindsight replaying that interview that you just went through.

For me, I was kind of like, constantly stuck in this future-state mode about being anxious about what’s coming next, and that combined with ADHD—which is something that I also have—is kind of a wicked combination. And we can talk about that separately, but once I started understanding what program management did and how program management allowed businesses to keep things on track, I realized that there was a parallel into my own life there. The skill of program management actually became my defense against the crippling anxiety that I felt anticipating future events. And it’s really become kind of the primary lens by which I understand and synthesize the world around me. And I know that sounds kind of weird, but with ADHD, I have a tendency to either being total diffuse mode and just working on nothing in particular, and letting my attention take me, or being in hyperfocus mode. And when you’re hyper-focused and anxious, it can be a deadly combination, right?

So, what I learned was taking that hyperfocus and taking that idea of program management and figuring out what it takes to get from here to there. I’m a strong believer in go as far as you can see, and when you get there, you’ll see further. And this skill of program management kind of becomes the stepwise function by which I get to that later point, very much like you were saying with coming to Waymo: You never know what you’re going to get until you get there. Well, now I see further and in hindsight, it was the right decision. So, the concept of program management is bringing structure, is bringing order, is bringing hierarchy to the chaos and uncertainty that we all naturally navigate in whatever we’re doing and trying to transmute that into some kind of transparent order and rhythm, not only for my own benefit to reduce my overall anxiety, but also for the benefit of everyone else who’s interested in what’s going on. Does that answer your question?

Corey: No, it absolutely does. Dealing with ADHD has been sort of what I’ve been struggling with my entire life. I was lucky and got diagnosed very early, but I always thought it was an aspect of business, but in many respects, it’s not just about owning a business; it’s about any aspect of your career, where the hardest thing you’re ever going to have to do, on some level, is learn to understand and handle your own psychology where there are so many aspects of how things happening can impact us internally. I can’t control what event happens next, of people yelling at me on Twitter, or I get a cease and desist from Amazon after they finally realized five years in, “You’re not nearly as funny as we thought you were. Stop it.”

Great. I can deal with those things, but the question is how I’m going to handle what happens in that type of eventuality? It’s, am I going to spiral into a bitter depression? Am I going to laugh it off and keep going on things that are clearly working? Am I going to do something else? And so much of it comes from—at least in my experience—the ability to think through what’s going on in a somewhat dispassionate way, and not internalize all of it to a point where you freeze. It’s way easier said than done, I want to be very clear on this.

Johnny: That’s absolutely right. Stepping back, seeing the forest for the trees. I’ve recently become fascinated with systems thinking. You know, I’m in Silicon Valley, so I might as well start looking into a complex adaptive systems—

Corey: Oh, no.

Johnny: —[crosstalk 00:21:09] buzzword. We don’t have to go down that thread because I’m very much an amateur when it comes to it, but what it does is it forces you to look at the connections between the components rather than the reductionism approach of let’s look at this component, let’s look at this component… instead, it forces you to step back and see the system as a whole. And so when you’re responding to you just got a cease and desist, you know, of course you’re going to feel depression, of course you’re going to feel anxiety, and understanding all those as part of the system of experiencing that situation, it lets you kind of step back and say, okay, it’s normal to be feeling this, it’s normal to be feeling that. How can I harness these and structure my approach so that I can get to some further point where I not only know what I can do, and what options are available to me, but I have a clear path forward and strategy for how I want to approach this.

Corey: How long have you been in your career at this point?

Johnny: So, I graduated college in 2009. And I worked at my first company for about ten years from 2005, so I guess you could say 17 years, plus or minus, if you don’t count internships.

Corey: Looking back, it’s easy to look at where we are at any given point in our career and feel that, oh, well, here’s where I started, and here’s where I am now, and here are the steps I took along the way where there’s a sense of plodding inevitability to it. But there never is because when you’re in the moment, in the eternal now that we live in, it’s there are millions of things you could do next. If you were to be able to go back to your to talk to yourself at the beginning of your career, what would you do differently? What advice would you give yourself that would have really helped out early on?

Johnny: You know, I think the thing that gave me the most leverage in my career was—as I move forward—is seeking out communities of like-minded, positive people. On the surface, that sounds a little shallow; of course, you would want to seek out communities, but what I’ve observed is that the self-organizing communities that pop up around technologies, or ideas, or roles, their communities of people who want to help you succeed. And I think, you know, one of the ways I reached out to you and was able to contact you was through one of these communities, right? So, you know, I talked a little bit the Cleveland Tech Slack earlier; most people aren’t familiar with what mediums are even available. There’s Discord, there’s forums, there’s Slack, there’s probably other areas that I’m not aware of, where you can find people who will help you find that next step in your career.

Actually [laugh] I got my first taste of community in online video games, so—

Corey: Oh no.

Johnny: —playing World of Warcraft back in 2003, you know you would have a guild—I was, gosh, how old was I in 2003, basically, early-20s and, you know, you’d have a guild of 40 people trying to coordinate all over one single voice chat server. And there was various groups and subdivisions, and so that was almost a project management exercise in itself. That’s where I first learned project management. By the way, I have a sneaking suspicion that the roles that we play and that we are have an affinity for in video games mirror the roles that were best suited to play in life. So, I find myself playing a support class in League of Legends or a priest in World of Warcraft or Lord of the Rings Online. I’m always that support person, the glue that helps keep things moving. And surprise, that’s exactly what I do for my career. And it works perfectly. So.

Corey: The accountant I keep playing gets eaten by goblins constantly, but, you know—

Johnny: [laugh].

Corey: —that’s the joy that I suppose.

Johnny: So, pretty early on, I developed this skill of creating friendships, and those friendships, in turn opened me up to these new communities. So, if I were to give one piece of advice to my early self, it would be to put more emphasis on finding and seeking out the communities that consists of people who are interested in the things that you’re interested in, but also are willing to help you get to where you want to go. How do you succeed? Well, you find someone who is doing what you want and you talk to them. About it and you figure out how to get to where you’re at from where you’re at.

And maybe they can’t help you, maybe they can help you but, you know, we have a unique ability to crowdsource our questions, whether it’s on Reddit, whether it’s on Slack or Discord, and just say, “Hey, I’m thinking about this thing. Does anyone have any thoughts?” You’re immediately—you know, if you ask the question correctly—given five or six different opinions, and then you can kind of meld and understand, okay, here are the options. Again, going back to what we were saying about how do you even decide what the next steps are? You can crowdsource that now, and so the one piece of advice that I would give is to seek out
communities of like-minded positive people.

Corey: This episode is sponsored in part by our friends at Vultr. Optimized cloud compute plans have landed at Vultr to deliver lightning fast processing power, courtesy of third gen AMD EPYC processors without the IO, or hardware limitations, of a traditional multi-tenant cloud server. Starting at just 28 bucks a month, users can deploy general purpose, CPU, memory, or storage optimized cloud instances in more than 20 locations across five continents. Without looking, I know that once again, Antarctica has gotten the short end of the stick. Launch your Vultr optimized compute instance in 60 seconds or less on your choice of included operating systems, or bring your own. It's time to ditch convoluted and unpredictable giant tech company billing practices, and say goodbye to noisy neighbors and egregious egress forever. Vultr delivers the power of the cloud with none of the bloat. "Screaming in the Cloud" listeners can try Vultr for free today with a $150 in credit when they visit getvultr.com/morning. That's G E T V U L T R.com/morning. My thanks to them for sponsoring this ridiculous podcast.

Corey: And I think the positivity is important. There’s a lot as particularly in tech, that breeds a certain cynicism that breeds a contempt almost. And Lord knows, I’m not one to judge; I revel in a lot of that when it comes to making fun of companies’ ridiculous marketing and some of the nonsense we have to deal with, but it has to be tempered. You can’t do what some of the communities I started out with did. IRC, learn how to configure Debian or FreeBSD, where it was generally, “Oh, great, someone else joined? Let’s see what this dumbass wants.”

It doesn’t work that way. It’s like just waiting for someone to ask a question so you can sink the knives in is not helpful. Punch up, not down. And making people feel welcomed and valued, even if they don’t understand the local behavioral norms quite yet is super important. I’m increasingly discovering, as I suspect you are as well, that I’m older than I thought were when I talk to folks who are just starting their careers about here’s how to manage a career, here’s how to think about this, I am veering dangerously close to giving actively harmful advice, if I’m not extraordinarily careful because the path that I walked is very much closed.

It is a different world; there are different paths; there’s a different societal understanding of technology and its place in the world. There’s a—what worked for me does absolutely not work the same way for folks who aren’t wildly over-represented. And I increasingly have to back off lest I wind up giving the, I guess, career Boomer advice style of irrelevant and actively harmful stuff. How are you thinking about that?

Johnny: So, I guess that kind of gets into the underpinnings of what I think it takes to be successful, right, and how do you find success in any aspect of your career? And—

Corey: And what is success?

Johnny: It differs for every person—yeah, what is success? And we were talking just before the show about how every person experiences not only what is success, but what does success mean and what do you believe the key is differently. For me—and this is pretty on—brand with where I am in my career and what I do—is I think the key to success is preparation. And it really ties into finding those communities and asking those questions, right?

There’s three key aspects to it, right? First is understanding how you learn. Everyone learns differently, and so knowing how you learn—and you know, college and school is kind of meant to kind of eke that out; it’s how best do you learn? How best can you succeed with these tasks that we give you, study for this test, learn these concepts? If you can understand how you learn, that’s the first step in preparing correctly, right, building your personal knowledge systems around that, taking notes, ordered hierarchy, structured thinking, that sort of thing.

Knowledge management is a good field, if you ever have some time to figure out what you want to do with your external hard drive of your whiteboard like I have back behind me here. The second aspect is just mastering how to seek out information, right? So, how do you prepare? Well, you have to understand how to seek out information. You mentioned, you know, positive communities versus potentially cynical or toxic communities. Their opinions are still very valid.

They might be jaded and they might provide a cynical opinion, but you still need to encompass that within the spectrum of your understanding of the world, right, because they have something that happened to them, or they have some experience that still is very valid from their perspective. So, seeking out information, understanding the people and the tools at your disposal, the communities that you can go to knowing how to discern the signal from the noise. And again, that’s really where your thread that really helped me—because you nailed a bunch of the questions that I just wasn’t entirely sure on in that Twitter thread, and when I went through that, it hit some of the major points that I was just uncertain on, and you just gave very clear, albeit, you know,
somewhat tongue in cheek cynical advice, to say like, don’t worry about the company, worry about yourself. And that really was
helping me get to that next step.

And then lastly, how do you prepare? And this is the one I always struggle with. It’s calibrating your confidence barometer. What
does that even mean? How can you calibrate your own barometer of your confidence? It’s a knowingness; it’s knowing what to expect.

And so for example, when I was getting into Google, I had no idea what to expect in terms of the interviews. So, what’s the first thing I do? I go out and I ask a bunch of people, people who know people who are at Google people who are at Google, what do I expect? What should I prepare for? What communities should I join? What books should I read? What YouTube videos should I
watch?

I ended up finding a book called Cracking the PM Interview by Gayle—I think her name is Laakmann McDowell. There’s a Cracking the Coding Interview as well. That ended up being, like, exactly what I needed, and going through that cover-to-cover got me into Google, amongst other things, and talking with the community. So, calibrating your confidence parameter, that knowingness of, I know that I’m ready enough for this. There will always be things that catch you by surprise, but knowing that you’re ready and having that preparation and that internal knowingness not only increases your confidence, but it also increases your ability to operate improvisationally when you’re in the moment.

And in fact, that’s exactly what I went through for this podcast. I have a little document in front of me where I just jotted my notes down last night, I was thinking through, what do I want to cover? What do I want to say? How can I respond to the questions that he’s going to ask me? He might ask me, you know, a curveball, but I have some thoughts that are structured, I’m prepared for this so that no matter what happens, I’ll be okay. And again, that really gets down to that essence of philosophy of program management that I have. No matter what happens, I’ll be okay; no matter what happens, we’ll be okay. And believing in that and having a level of knowingness—[laugh].

Corey: I am not a planner at all. For me, my confidence comes from the fact that I can’t predict what’s going to happen so I don’t even try. Instead, what I do is I focus on preparing myself to be effectively dynamic enough that whatever curveball comes my way, I can twist myself in a knot and catch it, which drives people to distraction when they’re trying to plan a panel that I’m going to be on. “Okay, so we’re going to ask this, what’s your answer going to be?” I have absolutely no idea until I find the words coming out of my mouth.

And if I try and do a rehearsal, I’ll make completely different points, and that really bothers folks. It’s, I don’t know; I’m not here to read a script. I’m here to tell stories, which is great for, you know, improv panel activity and challenging if you’re trying to get a software project off the ground. So, you know, there are different strengths that call us in different ways.

Johnny: Exactly. I mean, the flip side of preparation is improvisation. And you know, I spent ten years as a jazz musician playing trumpet in a swing band back in Cleveland before I moved out here. And that really helped me understand how to think
improvisationally, right? They give you the chords, the underlying structure by which you can operate, and then you can kind of choose your own path through there.

And sometimes it’s good, sometimes it’s bad, you learn over time, you come up with libraries of ideas to pull out of your head at any given time. So, there is an aspect of preparation to improvisation. And I think if you, I would encourage you to think about it more; I bet you do more planning than you think you do; maybe you just don’t call it that.

Corey: No, I have people for that now.

Johnny: [laugh]. “I have people for that.”

Corey: I am very deliberately offloading that. Honestly, that was part of the challenge I had psychologically of running my own place. If I were just a little better at following a list or planning things in advance, all these people around me wouldn’t have to do all this extra work to clean up my mess. Instead, it’s okay, let it go. Just let it go and instead, focus on the thing that I can do this differentiated. That was my path. I don’t know how well it works for others, and again, I’m swimming in privilege when I say it.

One last topic I want to get into, I think it might be part of the reason that you and I are talking so much about the future, the next generation, and the rest is we’re recording this on March 9th. I don’t know the date this is going to air, but there’s a decent chance that will be after April 22nd, where you and your wife Emily are expecting your first child. So congratulations, even though I’m a little early. I definitely want to get that in there.

Johnny: Thank you.

Corey: Have you found that since you realized you were expecting a child—with an arrival date, which is generally more accurate than most Amazon order dates—that you find yourself thinking a lot more about the future and how you’re going to wind up encapsulating some of the lessons you picked up along the way for, I guess, the next generation of your family?

Johnny: Yeah. I mean, everyone who finds himself in this situation, finds himself somewhere between panic and bliss, right? There’s some balance that I have to find there. And fortunately, my wife Emily, and I have a very strong rapport when it comes to how I think and how she thinks, and so we’re able to—you know, our emotional intelligence is very high; we talk about that sort of thing a lot. And we try to plan for the future as best we can, knowing that things will go off the rails as soon as you know, what’s the old saying about the best laid plans and how, you know, every plan is—

Corey: Man plans and God laughs.

Johnny: Yeah, or goes awry as soon as the first shot is fired, et cetera. Thinking more than five years out is still pretty challenging for me, but thinking within the first five years, we can already sketch out some plans. I already have some ideas of where we want to go and what we want to do and how we want this new child, this being, to experience the world and how we want to impart the things and the wisdom that we’ve learned and experiences and skills that we’ve developed—Emily and I—to this new child, realizing that I have no idea what’s coming and I have no idea what to expect because I just really haven’t had much exposure to babies or children at all in my life, so I’m just kind of rolling the dice here and trusting that it’ll all work out really well. And again, going back to communities, the communities that I’m in, there are parenting channels, there are friends and family that I can talk to. So, I have everything that I need in terms of knowledge.

Now, I just need to go through the experience, right? So, I’m definitely thinking a lot about the future. In fact, I’ve got a—I don’t know if you can see it here—quarterly plan for my life up here on the wall that I [unintelligible 00:35:33]. It’s just something that I can glance at every so often, and there it is, right, there: ‘Q1 2022: Kid.’

Corey: How long has that ‘Q1 2022: Kid’ been on the board? Like oh, since 2014? Like that is remarkably good planning.

Johnny: Mid-2021.

Corey: Okay, fair enough.

Johnny: No joking: Mid-2021.

Corey: [laugh].

Johnny: Yeah, just even having that up there and writing a sticky note and slapping it on there for, like, a hey, here’s what I think,
some of them fall off, some of them don’t fall off, but I’ll tell you what, more than more often than not, it actually ends up working and happening and being realized, no matter what it is. Because just having it there and glancing at it every so often is that repetition, it keeps it on my mind. It’s like, hey, I should probably think about that. The next thing you know, it’s done. And then I can take it off and put it in my binder of accomplishments.

Corey: I am about five years ahead of you on that particular path that you’re on because five years ago, I was expecting my first child. And I don’t want to spoil the surprise entirely, but I will Nostradamus this prediction here, five years from now, when you go back and listen to or watch this episode and listen to yourself talk about how you’re planning to parent and your hopes and your dreams, you are going to, in a fit of rage, attempt to build a time machine to travel back to what is now the present day for us, in order to slap yourself unconscious for how naive you are being [laugh] because that is—I’m hearing my words coming out of your mouth in a bunch of different ways, and oh my God, I was—it’s the common parent story you all these hopes and dreams and aspirations for kids and then they hand you a tiny little baby and suddenly it becomes viscerally real in a different way where, “It’s going to be a little while until I can teach you to do a job interview, isn’t it?” And other things start wind up happening to, like—

Johnny: [laugh]. Right.

Corey: —what do I do? I’ve never held a baby before. How do I not drop it and kill it? And later in time they learn to talk. They talk an awful lot, and then it’s like, how do I give them a bath without drowning them in the process? Not because I’m bad at it, but just because I’m at my wit’s end because I haven’t slept in three days.

Parenting is one of the hardest things you’ll ever do and everyone has opinions on it. And it’s gratifying to know that the world continues to go on even in these after-times where things have gotten fairly dark. It’s nice to see that flash of optimism and remember walking down at myself. It’s exciting times for you. Congratulations.

Johnny: Yeah. Thank you. It’s a beautiful thing. And I’m self-aware and I have a knowingness of my naivete, right? And that’s part of the fun.

And the whole idea of it is an explorative journey. I have no idea what to expect, but I have a good support system; my wife is incredible. She has an early childhood education degree, so that’s going to be really useful. Yeah. And so kind of going back to that concept of preparation.

And I don’t feel a lot of anxiety about it because I am feeling like I have the knowledge, the community, the friends, the family in place so that no matter what happens, I’ll be able to maneuver through it. And I can ask, and I can get help. Yeah, so that’s where my head is at with that. [laugh].

Corey: We’ll be checking back in once you’re up to your elbows and diapers and I assure you, you’ll be lucky if it stops your elbows.

Johnny: [laugh].

Corey: I really want to thank you for taking the time to talk to me about your own journey and, I guess, a variety of different things; hard to encapsulate it all at once. If people want to learn more or chat with you, where’s the best place to find you?

Johnny: Yeah, thanks for asking. So, I have a website jmpod.com, JM Pod. My middle name is Michael. So, John Michael Podhradsky. jmpod.com. That links to my blog, there’s links to LinkedIn, Twitter, Instagram. I’m most active on Instagram.

I’m always looking to connect with and just chat with new people, people who want a new perspective, people who are interesting or want to share their stories with me. Coaching is something that I thought of doing in the long-term. It’s not on the plate right now
because I’m focused on my current career, but that’s something that I’m very interested in doing, so you know, happy to field that questions or if anyone wants to reach out and hey, what communities can I look for or where should I be looking for communities, I’m happy to help with that as well.

Corey: I will, of course, put a link to that in the [show notes 00:39:39]. Thanks again for your time. I really appreciate it.

Johnny: Yeah, this was a fantastic experience. It’s the first podcast I’ve done, I’m hoping it went well, and I really appreciate that you even asked me to do this. It was a surprise. My eyes went like dinner plates when you said, “Hey, why don’t you come join me?” And I said, “Absolutely. That sounds like a fantastic idea.” So, thank you again, Corey. I really appreciate spending time with you and looking forward to doing it again sometime in the future. With a baby in the background, screaming. [laugh].

Corey: Oh, yes. They do eventually sleep; you won’t believe it for the first three months, but they do eventually pass out. Johnny Podhradsky, technical program manager of Offboard Infrastructure at Waymo. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry comment telling me exactly which tweet of mine you followed for advice and it did not in fact help your career one iota.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About “Matty”

Matt Stratton is a Staff Developer Advocate at Pulumi, founder and co-host of the popular Arrested DevOps podcast, and the global chair of the DevOpsDays set of conferences.

Matt has over 20 years of experience in IT operations and is a sought-after speaker internationally, presenting at Agile, DevOps, and cloud engineering focused events worldwide. Demonstrating his keen insight into the changing landscape of technology, he recently changed his license plate from DEVOPS to KUBECTL.

He lives in Chicago and has three awesome kids, whom he loves just a little bit more than he loves Diet Coke. Matt is the keeper of the Thought Leaderboard for the DevOps Party Games online game show and you can find him on Twitter at @mattstratton.

Links Referenced

  • Pulumi: https://www.pulumi.com/
  • Arrested DevOps: https://www.arresteddevops.com/
  • 8bits.tv: https://8bits.tv
  • Twitter: https://twitter.com/mattstratton
  • LinkedIn: https://www.linkedin.com/in/mattstratton/
  • speaking.mattstratton.com: https://speaking.mattstratton.com
  • twitch.tv/Pulumi: https://twitch.tv/Pulumi
  • 8bit.tv: https://8bit.tv
  • duckbillgroup.com: https://duckbillgroup.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Vultr. Spelled V-U-L-T-R because they’re all about helping save money, including on things like, you know, vowels. So, what they do is they are a cloud provider that provides surprisingly high performance cloud compute at a price that—while sure they claim its better than AWS pricing—and when they say that they mean it is less money. Sure, I don’t dispute that but what I find interesting is that it’s predictable. They tell you in advance on a monthly basis what it’s going to going to cost. They have a bunch of advanced networking features. They have nineteen global locations and scale things elastically. Not to be confused with openly, because apparently elastic and open can mean the same thing sometimes. They have had over a million users. Deployments take less that sixty seconds across twelve pre-selected operating systems. Or, if you’re one of those nutters like me, you can bring your own ISO and install basically any operating system you want. Starting with pricing as low as $2.50 a month for Vultr cloud compute they have plans for developers and businesses of all sizes, except maybe Amazon, who stubbornly insists on having something to scale all on their own. Try Vultr today for free by visiting: vultr.com/screaming, and you’ll receive a $100 in credit. Thats V-U-L-T-R.com slash screaming.

Corey: Couchbase Capella Database-as-a-Service is flexible, full-featured and fully managed with built in access via key-value, SQL, and full-text search. Flexible JSON documents aligned to your applications and workloads. Build faster with blazing fast in-memory performance and automated replication and scaling while reducing cost. Capella has the best price performance of any fully managed document database. Visit couchbase.com/screaminginthecloud to try Capella today for free and be up and running in three minutes with no credit card required. Couchbase Capella: make your data sing.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Returning today for yet another round on the Screaming in the Cloud podcast is my dear friend, and hopefully yours as well, Matty Stratton. Since the last time we spoke, you’ve changed jobs, Mattie; you’re now a staff developer advocate at Pulumi. I don’t believe you were the last time you were on this show, but memory escapes me.

Matty: You know, I was just wondering that myself, and I guess we’ll have to go back to the archives.

Corey: Yes, but that sounds like work, so we’re going to roll with it anyway.

Matty: Everyone who’s listening, go do the homework for us. And, like, just tweet and let us know what my job was last time.

Corey: And yell at us if we get it wrong, of course.

Matty: Yell at us if we get it right.

Corey: In the interest of being, well, I guess a little on the judgey side—because why not I tend to be good at that.

Matty: I was hoping to be on the judgey side on this show.

Corey: Oh, absolutely. You have a very strange career trajectory, in that—the companies you work for and how that winds up going back and forth. But when we first met, you were at Chef; and Chef, great company. And after that it was PagerDuty; great company.

Matty: [laugh].

Corey: And then it was IBM Hat, which I—was it Red Hat, was it IBM side?

Matty: For me, it was Red Hat.

Corey: So, it went from Chef, which is great, and a company that was doing a lot of things on the container side of the world became a thing and mutable infrastructure did sort of change Chef’s business model. And then you went to PagerDuty, the wake-you-up-in-the-middle-of-the-night service named after some legacy technologies. And should be very direct in the popular consciousness, IBM views pagers as newfangled technology in some circles, in some areas, so it feels like you were traveling back in time a bit, again and again and again. On the federal side as well which, for excellent reasons, is not usually the absolute bow wave of innovation because you don’t usually want your government doing that in some ways. And now you’ve leapfrogged into Pulumi, which is sort of the bleeding edge of the modern way we think about provisioning cloud infrastructure.

It feels like it’s a very interesting trajectory. Now, this is speaking as a complete outsider, I’m going to assume that’s not how you view basically any characterization of any of those companies I’ve just named. How do you view it?

Matty: You know, I don’t know that I necessarily disagree with the way that you’ve put everything, but there’s some nuance and some interesting stuff when it comes to that. So, I’m going to specifically talk about the Red Hat thing; why did I leave PagerDuty? And one of the interesting things is, I actually had an offer from Pulumi at the time that I took the job from Red Hat. So, it actually took me a year to come and work at Pulumi. And the little bit of the short answer is Red Hat backed up a big truck of money. And we all have a price.

Corey: Yeah, the dulcet tones of a dump truck full of gold bricks emptying itself into your backyard, it’s hard to say no to.

Matty: The reason that I want to bring that up is that has nothing to do with specifically Red Hat the company versus other companies. It was the role. It was a sales-oriented role, so if you don’t know, sales gets paid a lot of money and there’s good reason. One of the reasons—again, if you don’t work in sales, you don’t necessarily know this—is, the last day of the quarter, you will have your VP of sales talking, he’ll be like, “Corey, you are amazing. I love you. Look at this big deal you brought in.” Twenty-four hours later, “What have you done for me lately?”

Corey: Mm-hm.

Matty: That didn’t matter, right? And I remember the CEO of PagerDuty—so Jen Tejada—at one of the sales kickoffs I was at, she said—you know, because salespeople, like, you might know this, like, the top sales reps in the company, they go on trips, they have all this stuff—and Jen said, you know, “I’ve got engineers here that are like, well, I don’t understand.” It’s like, “How come the salespeople get to go to Bermuda or do whatever?” And she’s like, “Would you like your paycheck to change every quarter based upon specifically what you did and have the stress of what have you done all this stuff? No? Okay, cool. Then you can keep”—you know, there’s a trade-off. So, the point of that was—

Corey: And as your paycheck gets smaller, you’re getting closer and closer to losing your job because a salesperson needs to perform to keep. It’s very feast or famine. It’s a heck of a role, and I have nothing but respect for people who can do it.

Matty: And people can do it well. And I do feel like a lot of people don’t understand how sales works, especially in a larger organization, and I think it’s really important. So, one of the things that was interesting is we’ve all—I shouldn’t say all, but many of us have worked in jobs that have some form of variable compensation, some kind of annual bonus. So, let’s say for example, at x company I’m working at, they’re like, “Mattie, your bonus is equal to 10% of your paycheck.” Well, the most it could be, generally speaking, it’s like, let’s say that your bonus would be, I’m just going to make up a number and say it’s a $10,000 bonus.

That’s the most it could be, and that’s if everything is amazing. Maybe I’ll get a little more. Now, your commission, your what they call your on-target earnings and sales, they’ll tell you a number and they’ll say, “Okay, Corey, you’re on-target earnings are, say $200,000.” And you’re like, “Oh.” But whatever.

The thing is, if you’re only getting you’re on-target earnings, you probably are needing to look for another job. So, you remember, like, we hear it differently, those of us that have done bonuses in a non-sales way. We’re like, “But that’s not a lot.” You’re like, “No, but what they tell you your commission is, it’s actually… it better end up being more or else you have trouble.” Anyway, point is—

Corey: And in some cases, it could be a significant multiple of that number as well, for top performers.

Matty: Absolutely.

Corey: The upside is always interesting, and calculating out the nuances of the sales plan is always a challenge, speaking as a business owner. It is a very specific field that has a bunch of nuance to it. Something I learned very early on is that if you manage salespeople as if they were engineers, or manage engineers as if they were salespeople, you are going to have an absolutely terrible time.

Matty: I think one of the things that, along those lines, I’ve have had conversations with people who work in different parts of technology, different parts of the business, who their long-term desire is to be a CEO, and I’m like, you really should go spend some time working in sales because most CEOs—again, this is blunt, but it’s true—if you think about it, what is the area of the business that they pay the most attention to? And I don’t mean, they don’t care about the other stuff, but who is the person on the executive team that the CEO is mostly joined at the hip with, and it’s your chief revenue officer, it’s your head of sales because you have to understand that, you have to understand pipeline, how that—you have to understand a lot of things as a CEO, but if you don’t know how sales works—it doesn’t mean know how to sell but know the ideas behind it. I mean, you should know how to sell, but you know what I mean?

Corey: Yeah, I think every CEO is selling. It is a sales job, whether that is selling the company to prospective employees, whether it is selling strategic partnerships, whether it’s being brought in to help close strategic deals, et cetera, you’re always selling in that role.

Matty: That’s a very good point. I should rephrase that, where I wasn’t saying you don’t need to know—

Corey: CEO who has no idea how to sell [unintelligible 00:07:42] the fundamentals of—like, you put them in a meeting, and they wind up saying the wrong thing and pooching the deal, yeah, they’re not CEO for very long.

Matty: It’s not just knowing how to sell, it’s understanding how a sales process works. That’s sort of the thing.

Corey: I’ll take it one step further beyond that, and that is that I believe that every professional is working in sales and is selling something, but not everyone’s aware of it“. Well, I’m an engineer, and I don’t do any sort of sales work.” Well, I hear about that from folks who are—“I have all these great ideas, but none of them ever get implemented.” Well, you’re not doing an effective job of selling the idea. “I keep getting put up for promotion and not getting it,” or, “I’m not doing well in job interviews.” Or, “I’m trying to get a raise and it just isn’t working for me.” And every job has elements of sales to it. I’d argue a lot of facets of modern life have sales elements to it.

Matty: They do and I think the reason that people get hung out—I agree with you; I could not agree with you more. I have a talk I used to give called “The Five Love Languages of DevOps” but it was really a talk about effecting organizational change, and you have to be a salesperson, right? But I think we have this—and this is a much larger topic because it comes into how people always want to distance themselves from sales—we have this thing in our head that when we think of sales, we think of tricky people. Shysters, right? Someone that’s trying to, like, pull a fast one on us, like the used car salesperson thing.

And I’m like, that’s not most salespeople. Like, salespeople want you to—because when we talk about learning how to sell, it’s not learning how to trick somebody. It’s actually learning about how to—I mean, here’s the biggest thing. You want to know—we talk about DevOps all the time and stuff like that, you know, and empathy. You want to know one of the most important skills of a salesperson is? Freaking empathy.

Because you need to be able to understand what your prospect—and that’s if you’ve, you know, there’s the book, The Challenger Sale, which like all business books can be summarized in a blog post, right, so you can just go read the blog post about The Challenger Sale; that’ll tell you everything you need to know, but a good salesperson that’s a challenger-style salesperson knows the customer better than they know themselves and knows there problems they might have that they’re not aware of. And it’s not because they’re smarter; they have a different perspective. So, the same thing is true. So, to Corey’s point, we’re always selling. And even whether it’s figuratively, like, conceptually—but I used to say when I was a Chef I said, the two best sales—most effective salespeople at Chef were Adam Jacob, the founder, and Nathan Harvey, the VP of community.

Sales engineers are powerful because a customer will tell things to a sales engineer they won’t tell the rep because they think the rep is trying to take advantage of them, which isn’t true. Most important conversations that happen are on the walk from the front desk to the conference room. How many conversations would I have with the SRE, or whatever, who was the one who came to get me from reception, and we’re just walking to the conference room. I learned so much there than in any other discovery session? You know, and then you use that to be—

Corey: And there’s not such thing as an easy sale either. And I think that gets overlooked a lot. Like, here at The Duckbill Group, if you bring us in on a consulting engagement to fix your AWS bill, you will turn a profit on that engagement. That has always been true. And we are quite literally selling money.

It is effectively one of the easiest possible sales you can make; it is incredibly easy to calculate out what the ROI looks like on any of these things, and it’s great, and we still have a full-on enterprise sales force because that is what it takes to wind up getting deals done when you’re selling business-to-business. These are not selling t-shirts to the masses. It is a nuanced field, and honestly, when I’m interviewing people, one of the easiest ways for me to discount someone as a potential hire is that they start talking smack about sales because it is clear, first, they lack empathy, and secondly, they don’t understand what sales does.

Matty: One of the things that I think people who are not connected with it don’t understand that again, back to Corey’s point about because selling is hard, and selling internally is hard. So, this is the thing. So, you can have a champion inside your prospect who’s, like, “I’m all about hiring Duckbill.” But they have to convince other people. So, what are salespeople really good at doing? They’re really good at helping you build your business case to be able to get your thing that you want.

Corey: How to turn your champion into an effective advocate for the thing that’s going to make their job easier because they’re not the person that signs off on it.

Matty: And they’re not the expert. Like, this used to happen when I was at Chef and I would have a customer who was like, “Okay.” They go and buy a bunch of licenses, and they’re like, “Well, it didn’t get deployed.” And we’re like, “Well, how can we help you?” And they’re like, “Well, no, it’s just internal stuff. We got to convince people or whatever.”

And I was like, “So, what you need to do is what you’re telling me, what you need to do is sell Chef, right?” “Uh-huh.” There is nobody on this planet better at selling Chef than Chef. So, that’s where that comes in because again, that’s how everybody wins. So anyway, I went there because I was getting paid like a salesperson.

Also, I one thing I wanted to touch on. So, you’re right, usually, public sector is not seen as the most cutting edge. One of the things that’s interesting at Red Hat, especially on the sales side—and friends of mine who are working on the commercial side may disagree with this, but it’s generally not been true—what they call NAPS, so the North America Public Sector, I used to say I was a NAPS specialist, which sounded awesome. Because that was my title, I was NAPS specialist; I specialized in NAPS—is actually—

Corey: Your status in the internal messaging system should always be sleeping at that point, why not?

Matty: Sleeping. Yeah. But it’s sort of known that actually the kind of emergent tech group and sales inside of the public sector, inside Red Hat, is very innovative compared to other ones. So, a lot of stuff was created there. So, it was we were doing something around a transformation office that wasn’t being done in the same way anywhere else, so it was very exciting.

So, I—also was the opportunity to go and work with people like Andrew Clay Shafer and John Willis and people that were—you know, it was all the people I was going to get to work with. So, that got me excited to be there. And then Covid happened, and I got news for you. Like, my job was to have challenging conversations with people about how they should do work differently. It’s pretty easy to tune somebody out on the Zoom, it’s a lot harder to tune somebody out when they’re challenging you in a room.

So, it was very hard to do this job during Covid, so our team really kind of disbanded towards the end of the year. I was really on the fence to join in the first place, and the person who was referring me to come work on the team who wanted to convince me said, you know, “What’s holding you back?” And I said, “Well, it’s not”—I said, “I really like developer advocacy. I like DevRel. That’s not this job.” And he said, “Hey. Come try this for a year, and… if it turns out you didn’t like it or wasn’t for you, then go back and do DevRel.”

And so that’s sort of what happened. And I have seen though I am much happier in a smaller organization that’s creating—you know, like, I like to feel my impact. I think everybody should spend some time in a large org because if you’re going to be working with other people—right, you know what I mean—especially if you’re a vendor, if you work on the vendor side like I do and stuff, Corey, you and I’ve talked before about background and doing developer advocacy, and I always say that, like, I do DevRel on easy mode because it’s very easy for me to have empathy for my prospects and community because I did the job for 20 years. It’s not impossible to be effective doing this job if you haven’t literally done it. It’s just that much harder. So, I [crosstalk 00:15:04]—

Corey: It’s a lot harder. And there’s a credibility question and the rest. Yeah.

Matty: I do this on easy mode. I can sit there and I can say, “Yes, I feel your pain. I literally did it for 20 years.”

Corey: And you’re at a point, too, let’s be clear here, that you have a gravitas to you. I use you as my default example when I talk about, like, the expression of DevRel in that if you—like back when you were at PagerDuty, which I guess dates the reference a bit, but it was, okay. If you sit down and say you’re doing on-call wrong, now I’ve been around this industry at that point 15 years or so, and I’m pretty sure I’m not. But if you’re going to say that you have already got my attention in a constructive way, not in a, “Well, let me just tear this apart.” It’s, no, no. I’m about to learn something by whatever it is you’re about to say. And it’s very hard to have that level of credibility without having done the role.

Matty: That’s true. Without doing it in that way. I mean, this is [crosstalk 00:15:59]—

Corey: In the practitioner way of practicing the thing for which you are advocating. Like, someone telling me that I’m doing on-call wrong, who has never themselves been in a role where they themselves were on call is a little lacking in the authenticity department. It’s not impossible and it can’t be overcome.

Matty: And you have to do it in a different way, right?

Corey: Yes.

Matty: And this goes back to another thing that I say a lot—my pithy Stratton quote is, “DevRel contains multitudes,” right? So, this is one of the things that we ran into, like, when we’re building out our advocacy team at PagerDuty, it was seeing sort of my boss was an amazing dude and everything like that. I love him, but like, we don’t scale horizontally. Our team was made up of enough of different kinds of people that, like, the way that I was able to do it because I had a certain experience, you couldn’t expect that out of another one of my teammates because they actually had a different way of doing it that was just as effective, but in a different way because they have a different background, they have a different—so that’s—

Corey: And there’s so many ways to do DevRel. Oh, yeah. Like, I’m going to call it my own
bias here where when I think about DevRel, I think about it through a lens of the way I approach things, and when I give conference talks, of how I present myself, and the rest. And my approach would absolutely be aligned with what I just described, “So, you’re doing AWS billing wrong.” And based upon who I am, and what I do, I can make that claim with some credibility.

If I were relatively new to the industry and giving a talk about AWS billing, I would not lead that way because it does not present nearly as well, and it’s going to call into question a whole bunch of skepticism. I would instead approach it as, “Here are some interesting facets about AWS billing that you may or may not be aware of.” There are different ways to approach it. Let’s also be clear that it’s not just conference talks; it can be blog posts, it can be documentation, it can be writing sample code, it can be Twitter, it could be TikTok of all things. There are so many ways to communicate with an audience, and your audience is wherever you happen to find them.

Ideally, not in line at the Starbucks harassing the poor person in front who’s just trying to order their coffee, but you know, as long as it’s all consensual, talk to people who are interested in this stuff, wherever they happen to be.

Matty: I think that’s a really important statement you said there towards the end, which is meet people where they are, whether that’s where you want them to be or not. And this comes up, it’s interesting because one of the things—I’m a big believer in repurposing of content, and that’s just partially because of effectiveness, but it’s like, hey, if I give a talk, I should make that a blog post, I should make it a video, I should do a code example. And it’s not so much because then I can hit all my OKRs with my boss.—I mean, that’s part of it, right?—but not everybody likes the same kind of content.

You know, there are people who really like videos, and there are people who are like, “I don’t want to learn from a video at all.” And there’s two ways you can approach that. One is you can say, “You’re wrong. Videos are better. You should watch all my videos.” And take a guess about how well that’s going to work with them getting your information or say, “I’ll meet you where you are.”

And I learned this even well before doing DevRel when I just thought about internal communication at an organization I was at when I was at Apartments.com and I was like, how do we get information? And you can’t just say, like, well, we have this email we send to everybody. Well, everybody doesn’t read email, right? So, it could be, maybe some people like RSS feeds, they want to capture it there. And the example I always gave was the most effective way that I ever saw that information was communicated inside our organization was signs in the restroom.

Corey: Oh, yeah. That’s a well-renowned way of doing it. That I think that Google pioneered this for a while. They had these all these things up about interesting things going on inside the—

Matty: Oh—

Corey: —company, about the way some systems worked—

Matty: —I was at Google office and using the restroom, and I was standing there, and right in front of me with a whole good practice on cross-site scripting vulnerabilities. I guarantee they probably sent that email to everybody, it’s probably been in meetings, and the people who saw it, [unintelligible 00:19:53] they saw it in the restroom.

Corey: Now, of course, I’m sure they probably sell ads on those sheets, but okay.

Matty: Yeah. You know, a little bit of that. When I was at Apartments.com, the floor that I worked on, the main restroom I used was a shared restroom with another office, which meant corporate never put anything up in there, and there was actually a fair amount of stuff that I didn’t know about because I ignored it everywhere else and [unintelligible 00:20:14] anyway. So, the point is, back—if you will do work in person, which who is doing that anymore and why bother?—your most effective way to communicate. So, if you can figure out how to do DevRel in signs in a restroom at a conference—ohh, conferences should sell sponsorship of restroom signs.

Corey: The jokes write themselves and almost certainly violate the code of conduct of at least four different [unintelligible 00:20:38], but it works. It works.

Matty: [laugh]. We’ll take those to Twitter.

Corey: You’ve been around the industry for a while. You are one of the cohosts of the Arrested DevOps podcast; you’ve been instrumental in organizing a number of DevOps Days… or Devs-Ops days, however you want to mis-pluralize that is fine by me; roll with it. Ant—

Matty: We argue more about the capitalization than the pluralization.

Corey: Very fair. I want to talk to you a little bit of how the DevOps movement slash community slash role has evolved. For a long time now, it’s been, “Great. So, where are the DevOps people sitting?” And then when you hear the shouted response of, “It’s not a job. It’s a culture,” good work. You found them. Now, you can go talk to them and all. What has changed over the past few years in the world of DevOps?

Matty: So, I am fond of saying you can’t buy DevOps, but I can sell it to you.

Corey: Oh, absolutely. You’re an exemplary DevOps salesman.

Matty: Yeah. So, what happened? When we think back across the decade-plus, you know, back since 2009, one of the things I think that’s interesting is, when we look at things like DevSecOps, or the other portmanteaus that are being created. It’s a little bit like that meme, right, with the astronaut: “Wait. You mean, it’s been DevSecOps all along?” You know, it’s, “Yes, always has.”

That’s the thing. Like, for those who don’t know, Andrew Clay Shafer is best known as coining the term. And I love Andrew, but wow, is it the worst name in the world for what
we’re talking about. Because it makes us all think that it’s only about development and operations. And it’s always been about cross-functional across all of those things. And if it helps us to give it a different name, great.

Corey: It’s replacing dysfunction with cross-function.

Matty: Yes. There we go. That’s DevOps right there. That’s the best definition of DevOps I’ve heard. You heard it here.

Corey: That one coins a phrase, in case you wondered.

Matty: So, we still use the term CALMS to say what is about: It’s about Culture, Automation, Lean, Measurement, and Sharing. That’s held up for a reason. For something that was scrawled on a napkin in 2010, there’s a reason we still talk that way. It sounds like we talk about culture more than anything else, and it’s not because it’s more important. It’s because it’s the one that we have to scream from the rooftops.

You don’t have to convince engineers to play with automation tools; they’re going to do it. That’s fine, right? So, they’re all equal. Now, that said, what’s changed is we have definitely found DevOps to feel a lot more that it’s about automation. It’s about the technology. We’ve veered away from the people to your statement about, like, “Oh, it’s a culture, not a ti”—well, it’s all of these things.

Corey: This episode is sponsored by our friends at Oracle Cloud. Counting the pennies, but still dreaming of deploying apps instead of “Hello, World” demos? Allow me to introduce you to Oracle’s Always Free tier. It provides over 20 free services and infrastructure, networking, databases, observability, management, and security. And—let me be clear here—it’s actually free. There’s no surprise billing until you intentionally and proactively upgrade your account. This means you can provision a virtual machine instance or spin up an autonomous database that manages itself, all while gaining the networking, load balancing, and storage resources that somehow never quite make it into most free tiers needed to support the application that you want to build. With Always Free, you can do things like run small-scale applications or do proof-of-concept testing without spending a dime. You know that I always like to put asterisks next to the word free? This is actually free, no asterisk. Start now. Visit snark.cloud/oci-free that’s snark.cloud/oci-free.

Corey: Well, one thing I do want to call out because the whole point of having you on the show, of course, is to embarrass you with proof-positive, for example, that you are in fact, a good person at heart despite, you know, your dubious friendship with people like me, is we both used to be adamant about the idea of DevOps is not a role, not a job title, and we both stopped, but for different reasons. The reason that I stopped was that I took a job as the director of DevOps at a company because I was trying to solve about five or six different things that were important for me to negotiate for, and job title did not make the cut of impactful changes. You had a far less self-serving reason for no longer picking that particular fight. What was it?

Matty: [laugh]. I do want to call out one of my favorite jokes which is not supposed to be gatekeeping, but it’s making fun of Corey so it’s okay—

Corey: Hmm.

Matty: —Nathan Harvey said years ago, and it was actually I think, intended as a shot at our friend Pete Cheslock, who also has had the title of director of DevOps, which said, “The only DevOps tool is a person that calls themselves director of DevOps.”

Corey: Oh, absolutely. It’s super lucrative. I was really insulted by that and cried all the way to the bank.

Matty: Uh-huh. Now, I’ll tell you there’s two reasons that I’ve changed my tune on—you know, I used to say it’s not a tool, title, or team. I still will agree that it’s not a tool. The title and team—and the reason for that is twofold, and neither of which are self-serving other than I don’t want people to think I’m a jerk. The first reason that deviated me from a little bit was again, to go back to your friend and mine, Pete Cheslock, he gave a talk, I don’t remember where it was, but he made the point where he said, “You look at it, the title ‘DevOps engineer’ is a 30 to 35% pay bump, so it’s like, I don’t care what you call yourself. Go get paid.” So, that’s that.

Corey: Yes.

Matty: So, first of all, I was like cool—

Corey: J. Paul Reed did a whole talk-pay thing that shined a light on that.

Matty: Absolutely. The one that I think is more empathetic and probably was… is maybe a little more important—or equally so—Ian Coldwater has pointed out before, and this really resonated with me, is that when we get on Twitter and are like, “Oh, my God. DevOps engineer is not a real title, blah, blah, blah.” The people that hear that are the people who have that title. They did not give themselves that title. It’s very exclusionary, and all that will happen out of that is it doesn’t eff—

Corey: “I’m going to go quit my job and not be able to make rent this month.” “Why?” “Because Twitter said that my job title was bad.”

Matty: Yeah.

Corey: All the reasons to quit a job, I promise you job title is not one of them. Unless it is something horrifying, as into the territory of discriminating or belittling. There are always exceptions to every rule, but by and large, “That’s a ridiculous job title,” is not the reason to quit a job. Says the self-proclaimed chief cloud economist.

Matty: Totally yeah. I mean, like, you know what is very similar? There’s a meme about, like, every time people want to make fun of a political figure or something and they’ll make fun of them being overweight, or any kind of thing, and the meme is like, the only people who hear that are your friends that have a similar condition, not the actual person you’re making fun of, so all you’re doing there is hurting people who… so that’s a similar thing.

Now, I will say—and I think you and I might disagree about this a little bit, so that’ll be fine—

Corey: I hope so.

Matty: So, when I hear—and actually the title doesn’t do this, for me; it’s actually very specifically a DevOps team. When people say, “We have a DevOps team.” This is not a perfect analogy when I say it’s a code smell; I call it an organizational smell. And what I mean by that—it’s not as bad as a code smell—what it does is it makes me ask more questions. If it’s relevant to me to ask questions. It might be none of my damn business. If you tweet that I’m on the DevOps team, I’m not going to come into your mentions and start questioning your existence, but—

Corey: Oh please, I have way better personal attacks than that.

Matty: Oh, yeah. But if I’m working with you and we’re working on that, or we’re having a conversation, and it comes up that you have a team called DevOps Team, I’m going to ask questions because that could be, okay or it could be, [sigh] I want to use the word dangerous lightly; it’s not, but like, counter-effective. And the reason for that is if the DevOps team is the one who does all your automation and you haven’t really enabled other squads and all you’ve done is move a silo around, doesn’t make you a bad person, but that’s not the most effective way you could be. So, it makes me start to ask questions, right? But sometimes DevOps teams are people who lead in the organization, they are empowerment teams, maybe they run dojo, maybe they are subject matter experts that help.

As long as there are good bridges still being built, it’s not bad, right? So, it just—again, it raises questions. It’s not inherently wrong. I am sure that… Pulumi where wo—actually, many of the tools I’ve worked with have been called DevOps tools; I will still tell you there’s no tool that gives you DevOps, right? You can’t—

Corey: But when other people—like, read as ‘buyers’—refer to you as the ‘DevOps tool company,’ well, you can be right or you can make a sale, in some cases.

Matty: [laugh]. Yeah, I’m not going to tell you—

Corey: On some level, you have to meet people where they are, and this is a part of that. I say that in full sincerity. Same story with the idea of culture. I hear this question all the time, “How do we wind up making all of our engineers aware of AWS billing issues?” And to a point, you should have understanding that when you turn something on it runs forever, bigger things cost more than smaller things, but the knowledge fits on an index card.

You shouldn’t have every engineer wanting to—or needing to—become deep experts in this space. Having a centralized team that specializes in that, at a sufficient level of org size and maturity, makes an awful lot of sense, and they can float around. But yeah, having the AWS bill team, in some cases is the right answer and others it’s the complete wrong answer, and it really does depend. I think the way that we solve this problem, authoritatively, is a way that neither you nor I can argue with it because the only source for authoritative DevOps answers is from the source itself, and that is, of course, Emily Freeman, whose treatise on the subject, DevOps for Dummies, despite the weird title, is absolutely fantastic work that gives insight into all of this. And are you prepared to tell her she’s wrong? Because I’m certainly not.

Matty: Well, there are plenty of people who will. As we know.

Corey: Yes. And we call them shitheads if we’re being perfectly honest with you.

Matty: Yeah. [laugh].

Corey: The internet what a ple—no, Emily is an absolute treasure in the space and I’m continuing to watch her meteoric rise with nothing other than pure admiration. It is just spectacular to see her succeed.

Matty: I could not agree more. This is something I struggle with a little bit. I don’t think Emily would mind me saying it this way. This is the thing where you don’t want to sound condescending, but I always love when I look at people and it’s not—it’s going to come off a little bit about, like, “I knew them when,” and it’s not like I was a Corey Quinn fan before he went pop, but I love to see and remember where we all came from, and it’s true of myself and it’s true of other people, but that’s one of my favorite things is I love to see my friends succeed.

Corey, I love to see what you’ve done. Like, I think back to when we knew each other. I’m not saying you weren’t successful, but it’s funny, this [unintelligible 00:30:08] sounds a little condescending to be like, oh, I’m so proud of you, but I am. And I’m impressed. It’s great to see.

And Emily’s another example. Like, I remember when I first met Emily, and not like I was any big deal, either, but it’s like, everybody comes from somewhere, right? Like Jacquie Grindrod who just recently left Hashi, I remember when she started to get into DevRel and I was talking to her because she’s like, “I may be thinking I want to do this thing.” And you look and you see these people. And it’s not supposed to be like, “Oh, I remember when you were like the cute little baby DevRel.” It’s not like that.

And it’s like, it’s just impressive to see—and not even impressive. It’s you like to see people who do good work and have a good heart and want to help people grow and be successful. And I’ll tell you something, here—we’re going to get real for a second—you can be jealous of them. It’s okay. And I’m going to be honest, there are times that—Emily and Corey are both good friends of mine, and there are times that I’m like, “Wow. I’m a little jealous of you. Sometimes I’m a lot jealous of you. Sometimes I’m not at all.” So, I’m telling everybody, it’s okay to be jealous. [laugh].

Corey: I agree with the sentiment that I changed the word ‘envious’ because envy is one of those, like—

Matty: Okay.

Corey: —“I want that, too,” whereas jealousy is a lot more a shade of, “I want to have it and I don’t want them to.” And I don’t believe that’s the direction you’re heading in. [laugh].

Matty: No. Thank you. No, you’re exactly right. Envy is the better one yeah because it’s never—

Corey: Now, I recently learned the distinction there by getting very wrong and saying things I didn’t intend to imply, which is why I bring it up. Again, let my mistake be something others can learn from. Sometimes the best purpose I can serve in this industry is as a counter-example.

Matty: Example. I was going to say, you know, just for everybody, I remember at the beginning, you know, Corey said, “Maybe we’ll learn something.” I’m like, I guess that’s what we learned [laugh] is the difference between envy and jealousy.

Corey: Yeah.

Matty: [unintelligible 00:31:50] gotta say, you know, it took us half an hour to get there. But you know.

Corey: No no. And I appreciate your friendship throughout the years. Like, you were one of those people that has been something of a guiding star, where it’s, sometimes I get it right, sometimes I get it wrong, and you’ve always been someone who has been very willing to share which side of the divide you think I’m on with anything that I’ve done. And for lack of a better term, you knew me before I basically bought ink by the barrel. And back when I was just the conference speaker that had to follow one of your ridiculous talks, like, “Oh, God. Those are big shoes to fill. I’d better learn how to give a conference talk.” So, most of what I become is your fault. But I do want to thank you for your guidance over the years on these things.

Matty: Can we tell the real story about how I claim ownership of The Duckbill Group?

Corey: By all means, take it away.

Matty: Oh, okay. So, [[laugh]] I honestly still think that I should have a part ownership in
The Duckbill Group because for those of you who don’t know, Corey mentioned that I had worked at PagerDuty, and actually that job came down between the two of us and Corey didn’t get it. And then went and started his own company and became famous and amazing. So really, it’s because of me is what I’m trying to get at. I—

Corey: To be fair, they made the right hire. Which one of us do you think makes the better employee, let’s be very clear?

Matty: [laugh].

Corey: And yeah, I am thrilled to deal in you in on ownership of The Duckbill Group because the way we’re structured, you cannot have ownership without also assuming liability. So yeah—

Matty: [laugh].

Corey: I would love to dump legal responsibility for my shenanigans on someone else.
Come on in. Yeah, there’s always a cutting edge to everything else. But no, you’re right. I always wonder what would have happened if that decision had gone differently.

And I’m very glad it played out the way that it did. You were the right hire for the company in a way that I never would have been. But I would have given it a good try for a while before they begrudgingly had to fire me or I sensed the axe was coming and left on my own. That is the nature of me as an employee. You have a very different perspective because you’re good at things that I’m terrible at.

Matty: And vice versa. It was interesting. You just talked about, like, how would things go different? So I—yesterday—just recorded—I don’t know when it’s going to come out—I was on a podcast called 8 Bits—so it’s 8bits.tv—and it’s really a show about people’s journey through tech.

And what was interesting that came out of that conversation was, first of all, how much of how I got to where I am is because of spite. Which you’re going to have to go back and listen to the episode to hear the whole story of all the spite. But we did talk about, like, those junction points that happen that seem innocuous. And it’s like, I made this one choice that wasn’t even necessarily a choice and you follow all the forking logic that gets you to, Corey, you and I are sitting here on a podcast right now. How many decisions that weren’t even decisions? There’s the alternate universe where this doesn’t happen where this doesn’t exist, right?

Corey: It’s weird how this stuff all works. Years before I’d met either one of you, you videotaped my wife’s law school musical and burned it to CD. We found that out when you were here over dinner one night.

Matty: That was my favorite thing.

Corey: It was surreal.

Matty: Yeah, I was at dinner with Corey and his wife and we got into a conversation about that she had gone to law school in Chicago. And I was like, “Oh, funny thing. Like, I produced the video of the law school mu”—and she was like, “Wait, what was that?” And I couldn’t even remember. I had to, like, dig back into, like, an old blog post. And was that and then yeah, and Bethany, like—

Corey: She walks into the other room and comes back with a DVD that you burned, your handwriting on it.

Matty: Yeah.

Corey: Yeah.

Matty: Yeah, pretty much. Yeah. The world is small. Be nice to everybody.

Corey: It never hurts. I want to thank you for taking time out of your day to basically tell
stories once again. It’s always good to talk to you. If people want to learn more about who you are, what you’re up to, where’s the best place they can find you.

Matty: So, really the best place is Twitter. You know, so I’m at @mattstratton on Twitter. If you’re not a Twitter person, that’s okay. LinkedIn is not great for fi—I don’t always remember to post stuff there. If you want to know about upcoming, you know, so if you go to speaking.mattstratton.com, that has all my previous talks, my upcoming talks, and things as hopefully we’ll have more and more of that.

And yeah, and every week, I stream on twitch.tv/Pulumi on Thursdays. And it’s not webinars, it’s not slick demos, it’s just me screwing around and sometimes having fun people on, and sometimes just proving how little I know about coding. So yeah, good times. Thank you for having me on, again, Corey. It’s always fun.

Corey: Of course. Links to all that’s going on in the [show notes 00:36:20]. And as always, it’s a pleasure.

Matty: Also, I will say, Corey, I’ll give you the link to that 8bit.tv, if you want to put that in the [show notes 00:36:28]—

Corey: Oh, of course, we will.

Matty: —if people want to go and find that. Because I think it’s similar, connected to what we talked about.

Corey: Good. I look forward to listening to it myself. Mattie Stratton, staff developer advocate at Pulumi. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with a long angry comment detailing that DevOps is in fact a role and here’s what it means, and then go ahead and describe a sysadmin.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Joe

Joe Onisick is a polarizing technologist with nearly 25 years’ experience architecting, building, operating complex IT systems and advising customers on the same. Onisick’s passion is marrying technology to a customer’s real-time business challenges and leading them through the entirety of the adoption curve. Onisick is a Principal and co-founder of Transformation Continuum (transformationcontinuum.com), and founder of Define the Cloud (definethecloud.net).

Links:

  • transformation CONTINUUM: https://transformationcontinuum.com/
  • Twitter: https://twitter.com/JoeOnisick

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored by our friends at Revelo. Revelo is the Spanish word of the day, and its spelled R E V E L O. It means; I reveal. Now, have you tried to hire an engineer lately? I assure you it is significantly harder than it sounds. One of the things that Revelo has recognized as something I've been talking about for a while, specifically that while talent is evenly distributed opportunity is absolutely not. They're exposing a new talent pool to, basically, those of us without a presence in Latin America via their platform. It's the largest tech talent marketplace in Latin America with over a million engineers in their network, which includes, but isn't limited to, talent in Mexico, Costa Rica, Brazil, and Argentina. Now, not only do they wind up spreading all of their talent on English ability, as well as , you know, their engineering skills, but they go significantly beyond that. Some of the folks on their platform are hands down the most talented engineers that I've ever spoken to. Let's also not forget that Latin America has high time zone overlap with what we have here in the United States. So, you can hire full-time remote engineers who share most of the workday as your team. It's an end-to-end talent service. So, you can find and hire engineers in Central and South America without having to worry about, frankly, the colossal pain of cross border payroll and benefits and compliance because Revelo handles all of it. If you're hiring engineers, check out revelo.io/screaming to get 20% off your first three months. That's R E V E L O.io/screaming.

Corey: This episode is sponsored in part by our friends at Vultr. Spelled V-U-L-T-R because they’re all about helping save money, including on things like, you know, vowels. So, what they do is they are a cloud provider that provides surprisingly high performance cloud compute at a price that—while sure they claim its better than AWS pricing—and when they say that they mean it is less money. Sure, I don’t dispute that but what I find interesting is that it’s predictable. They tell you in advance on a monthly basis what it’s going to going to cost. They have a bunch of advanced networking features. They have nineteen global locations and scale things elastically. Not to be confused with openly, because apparently elastic and open can mean the same thing sometimes. They have had over a million users. Deployments take less that sixty seconds across twelve pre-selected operating systems. Or, if you’re one of those nutters like me, you can bring your own ISO and install basically any operating system you want. Starting with pricing as low as $2.50 a month for Vultr cloud compute they have plans for developers and businesses of all sizes, except maybe Amazon, who stubbornly insists on having something to scale all on their own. Try Vultr today for free by visiting: vultr.com/screaming, and you’ll receive a $100 in credit. Thats V-U-L-T-R.com slash screaming.

Corey: Welcome to Screaming in the Cloud, I’m Corey Quinn. My guest today is someone I’ve really admired from afar for a while just because he’s a study in contrast. By day, he is a transformation—effectively—expert. He’s a principal at his own consultancy that focuses on helping companies achieve their digital transformation. Very forward-looking, very high-level modern technology. But he also wound up effectively leaving Silicon Valley to go live in the middle of the woods. It’s not usually a common combination. Joe Onisick is the principal at transformation CONTINUUM. Joe, thank you for joining me and suffering my fairly ignorant questions.

Joe: Corey, thanks a lot for having me and the brilliant intro there.

Corey: [laugh]. So, I stumbled across you on Twitter of all places, which is where I spend my work time, my free time, my spare time, et cetera. When people say, “Where are you dialing in from?” I say, “Oh, Twitter.” And that usually gets a laugh, but it’s also a little
unfortunately true.

And your pinned tweet thread talks about how you weren’t particularly happy with your life, where things weren’t serving you and you decided it was time to make a change. It’s the kind of thing that I think an awful lot of people flirt with the idea of, but you actually went ahead and did it. What happened.?

Joe: So, I did a whole series of things. I think the big thing I tried to do was not bite off everything at once. So, the first thing I did was quit drinking. I was a—you know, which it says in the tweet and I’m pretty public about I was an extremely heavy alcoholic. So, I cut that out because I wasn’t happy with it.

And you know, the whole idea was I thought it was keeping me happy and it wasn’t. So, got rid of that to see how things were and then just started a series of changes, which has, I think, gotten more extreme over time.

Corey: Well, one of the early tweets in the thread was one of your coworkers at the time was planning to climb I think it was Kilimanjaro, and your position was, well, that’s not something I would normally do. May I join you? If that’s how it starts, it seems like well, that seems pretty far on most people’s extreme scale.

Joe: Yeah, that was an interesting one. The idea of starting in a rainforest and ending on a glacier up 20,000 feet was not of any interest to me at all, but it seemed like a life experience I wanted to put under my belt.

Corey: I’m assuming that you’re probably glad you did it because you don’t meet too many people who are like, “Oh, yeah. I climbed a mountain. It sucked. I never wish I hadn’t done it.” It feels almost like it’s writing a book, on some level where no one wants to write a book; they want to have written a book. Is climbing a mountain similar to that, or does it go in a bit of a different direction?

Joe: I think it was very similar to that. We did a ten-day track, but you can do it much shorter. So, we spent about seven days acclimatizing around the mountain and hiking around the mountain. So, it was more a little up and down, but more level. So, the first 15,000 feet was actually pretty enjoyable. It’s the summit day where you go from 15,000 to 20,000, that is—it’s just sheer misery, especially if it’s not something you do every day.

Corey: I thought I had a rough time whenever I visit my in-laws who live in Colorado Springs, and it’s great hanging out in their house and whatnot, and I run up the stairs and I get winded and it’s “Wow, what a tubby piece of crap I am. How did this happen?” It’s like, “Oh right, we’re at 9000 feet; the air is a lot thinner here.” So, I basically spend the entire trip out there, trying to move as little as possible as opposed to at home where I sit in front of my computer attempting to move as little as possible. But it hits in a different way.

You quit your job in Silicon Valley as a part of this journey of—was it a journey of discovery? Was it just a series of changes? How do you contextualize it? How do you describe it?

Joe: I’m trying to learn how to be whoever I am would be the way I’d describe it. I’ve spent my entire life being someone I thought I was supposed to be, and I never stopped to think who I am. So, a lot of this is just trying everything to see what fits.

Corey: And then you make one of the classic blunders as you do this; you decide, “You know, I’m not going to work a traditional job anymore. I’m going to start a consultancy.” That is truly the path of fools, speaking as someone who did exactly that. And looking back at it, it was one of the best things I’ve ever done for sure, but if I had known how much work it was going to be and all of the ins and outs and ups and downs in the managing of my own psychology, I’m not sure I would have the courage to get started.

Joe: Yeah, that's a great way to say it. I look back—my favorite example is one of my mentors started a couple of companies. His wife has had several exits. I mean, he’s just a wealth of knowledge of tech: Tech the industry, and starting companies, and when I brought the idea to him, he asked, “So, you’re thinking of starting a consultancy?” And I said, “Yes.” He goes, “I have one word of advice.” And I waited for him to reply, “Don’t.”

Corey: When you said that to people in my experience, they think, “Oh, they’re trying to hoard all the wealth and happiness for themselves.” It’s yes, that is what I’m trying to do. I view consulting as a zero sum game. There’s only enough room for one of us. Yeah, it never works that way.

It’s just such an up and down thing and when I talk to folks who work at big tech companies and they are asking, “Oh, you know, I want to become an independent consultant because I’m tired of my job and my company and the rest,” don’t do that. It’s going to be a few lean years and it’s going to take an awful lot of trying. And honestly, the hardest part of all of it, at least in tech—this is, to be clear, not a sympathetic problem—is at any point, you can walk away and say, “The hell with this,” and within a week, wind up getting a salaried job somewhere very comfortable, where you don’t have to deal with all the hard parts of running a business and it pays three times your first year’s revenue. And it’s so much easier to go down that path. Fortunately for me, that wasn’t really on the table because I’m an insufferable jackass who, my personality shines through and it turns out, this is not a desirable component in most workplaces.

Joe: I think we share that. I think I’ve made myself fully unemployable now, so I don’t have that parachute, which makes the consulting a little easier.

Corey: You also have an additional challenge that, for better or worse, I don’t, which is I fix the horrifying AWS bill, which means that I could demonstrate ROI with, more or less, basic arithmetic when people say, “So, why should I bring you in?” It is one of the easiest enterprise sales—not that there’s an easy enterprise sale—that’s possible because it’s, “What are you selling?” “Money.” The end. You advise on digital transformation, which is inherently a sticky concept itself. What is it that you do for companies?

Joe: So, I’d say we started out with probably the stupidest business model you could ever come up with. We decided we were going to address Fortune 100 technology companies at the same time as addressing the largest value-added resellers in the world, and at the same time, driving adoption services on behalf of them for their customers. So, we have three customer bases: The end-user of technology, the reseller of technology, and the vendor of technology, and we’re helping them all adapt to the transformations happening in the industry. So, off the bat, we were already crazy because everyone would tell you pick a segment and focus, right? Not just technology vendors, but a specific hardware or software.

But to create the value chain we do of getting their products to market and making sure they fit that market, we have to have visibility into all ends of the spectrum. So, we tackled the hard challenge to be able to be successful with what we wanted to do.

Corey: It sounds an awful lot like you are taking a more… I’ll use the term ‘honest’—I think honesty is the right word here—a more honest approach to getting companies to their desired outcomes. There are a lot of folks who specialize in, “Digital transformation,” quote-unquote, and that’s very much a thin veneer over, “So, what do you really do?” It’s, “Oh, we do cloud migration, specifically into this one cloud vendor.” And that journey of their digital transformation generally involves writing a very large and very specific check to a third-party company. And that’s the end of it, and it’s rinse, repeat, go all-in. You have an established track record of very much not doing that. Was that something that you did originally, or was that how the practice wound up evolving?

Joe: So, I’ve kind of worked in all components of it. I’ve built giant channel practices within some of the world’s largest VARs; I’ve worked on the—or started my career on the end-user side and then I got kind of drafted into the vendor side for a while. So, I’ve got exposure to all of it. I think the honesty piece has been—to a fault, integrity is a thing that for me, right? It’s a trigger. I always tell people, I’m opinionated; you’re going to get my opinions, but you’ll never get anyone else’s opinions. So, they might be subject to change, but they’re always mine.

Corey: There’s an idea of you could buy my attention, but not my opinion, and that has been something of a guiding star for what I do just because people look at it and say, “Oh, that’s this bold moral stance, and that’s just inspirational,” and no. Absolutely not. It’s that I suck at biting my tongue. When I look at something and I find it ridiculous, I can only go so long without, more or less, asking why the emperor is prancing around naked in front of everyone. And contrary to popular opinion, in corporate life, this is not a particularly valuable skill, in fact, just the opposite.

But it does lend itself to a certain perspective on the larger industry. When you talk to companies who are looking for digital transformation, how does that conversation go? It seems like, for better or worse, it is a nebulous problem, and companies are generally not the looking for things via Google ads, for example? “Yes, hello. I’d like to buy one digital transformation, please.”

Joe: Yeah, so it starts in several different ways. A lot of our business starts with a vendor with a new product that they know fits the market and fits where things are going, but they can’t get it to move, right? They can’t get it to sell, they can’t get customers to adopt it, they can’t get sales teams to understand it. And so we come in and try and fit it into the bigger picture while tying it to what people already understand and know.

You can call it, like, chunking learning, right? I’m not going to be able to learn astrophysics if I don’t have a baseline in math. So, we try and tie the future to today so that people can grasp and understand it. And the same ends up at the opposite end of the spectrum: You can’t go in and talk to a laggard customer about how machine learning and AI is going to transform their business operations if they’re still wondering how to manage what they’ve got today.

Corey: There’s an underappreciated skill in meeting customers where they are, and very often that can express itself as a perception of being condescending in some cases, and I think that’s where a lot of people get it wrong. The hallmark of a terrible junior consultant is to walk in and say, “Oh, what moron built this?” Invariably to said, quote-unquote, “Moron.” People don’t show up at work hoping to do a crappy job today. There’s a reason that things exist the way that they do.

Yeah, maybe it’s because they just didn’t know any better, but maybe there’s a constraint or context you don’t have. And generally in my experience, failing to respect that context is just the kiss of death because, think, it’s the only thing that separates software from being able to do your entire job.

Joe: Yeah, and it’s a lost art, right? It’s one of the things I do and love doing is training engineers how to be consultants, or salespeople how to be consultants, and it tends to be a lost art. We have these products or solutions that we’re positioning or that are our favorites and we try to shoehorn them in every hole. One of my favorite examples was, I was asked to go into a California government agency and buy them and sell them SDN, they wanted to know why they needed to adopt SDN. And instead of coming in and preaching SDN, which was what I was theoretically getting paid to do, I started asking some questions and immediately realized these people don’t want Software-Defined Networking at all.

They want, you know, to be on the command line whenever they can, and not have to touch the gear other than that. So, I started to dig a little more and eventually find out, they hired a new CTO, and that CTO had SDN-ified their last network, and so they thought it was going to get shoved down their throat. And they were trying to figure out how to get around that. And so instead of selling an SDN, I gave them the 15 reasons why their operation wouldn’t benefit from it and found another problem to solve for them.

Corey: There’s really something to be said for having the courage to deviate from the engagement plan. I find that there’s a certain type of consultancy that as soon as they realize the facts on the ground are not as described or things have changed, they keep trying to get back on track for the thing that they believe they’re there to do. But I’ve always viewed it as being there to help customers, and sometimes that means that it’s a bit different than what you expected. There are times I have actively advised customers to spend more on AWS. It’s, yeah, you could not have backups for those incredibly important things over there, but I [wouldn’t 00:13:12] generally recommend that. And I always get these strange looks. And it evolved my business practice a bit away from, for example, guaranteeing that I’d achieved a certain level of savings just because that it got people focused on the wrong outcome.

Joe: Absolutely. I draw some analogies, I do some woodworking as a hobby, and occasionally I’ll go out and buy a tool like a router or a bandsaw because I want that tool, and then I design projects around that tool. That’s great for a hobby when you have some spare income to blow. That’s a terrible way to run an IT operation.

Corey: That’s a lot of fun as a hobby, but if you’re a professional carpenter, that’s probably the wrong [laugh] direction to take things in. It’s a different approach to things. Your background is fascinating, and I would argue makes you incredibly well-suited for the role you’re in. You’ve been a principal engineer, you’ve been a CTO, you’ve been a VP of Sales and Marketing, you’ve sort of done, more or less, every major business function out there. The one I don’t see on your background listed is accounting and finance, but yeah, turns out you run a business, you learn real quick how at least the important moving parts there are.

What was it that made you decide to take that background, that eclectic group of skills and say, yep, consultancy, first off, and then it’s going to be aimed at solving these expensive existential questions that companies are wrestling with? Because it turns out the world increasingly runs on computers and that’s not something a lot of our customers are great at out of the gate.

Joe: So, some of it happened just by opportunity and chance. My first sales engineering role pulled me out of the customer side, and when the hiring manager called me to interview me and explain a sales engineer role, I told him, you know, “This isn’t for me. I don’t want to sell.” And then he ended up calling back the next day and explaining this training certification knowledge and growth path he put me on, and I changed my mind real quick because he was going to invest in me. So, some of it started by accident, then I realized the value in the diversity of knowledge.

I mean the human brain is a pattern-matching machine. The more data sets it has to match patterns on, the more powerful it gets, so the more diverse my job roles and the more diverse my education, my reading, my study become, the more I can help any given job I have by finding parallels to other things I’ve experienced.

Corey: You started your consultancy right around the time of the pandemic if memory serves, and that the running gag has been for a while now—it’s one of those haha, only serious type of jokes—is the global pandemic has done more to accelerate your company’s digital transformation than your last ten CIOs combined. And there’s something to be said for necessity forcing the issue in some cases. How have you seen it evolving?

Joe: So yeah, the pandemic definitely accelerated digital transformation and in fact, it was part of our first-year revenue success was that. There were some challenges that came with it. Large companies didn’t know what the financial market would look like, so they locked down spending and budgets quite a bit, so you got some good and bad there. But I think it accelerated a lot of things.

I think the maybe the disappointing part to me is that a lot of the things that the pandemic accelerated, were things that should have been happening anyway: Expanding remote work, building out better hybrid models to be able to secure SaaS, Infrastructure as a Service, and on-premises properties together, those types of things. They were things that we should have been doing, but nobody was forced to until the ‘oh, crap’ happened.

Corey: It’s one of those areas that is always felt like companies approach strangely. I’ve worked for a number of large companies over the course of my career who effectively decided to one day wake up, plant a flag in the ground and declare it we’re not a finance company—or whatever it is that they did—we’re a tech company. And in practice, I find that the execution of that vision doesn’t tend to extend much further beyond just putting a sign on the wall. Is that something you’ve seen and is a common trope, or do I just have really interesting luck in picking employers?

Joe: No, I think we see that a lot. I think we see a lot of large, intelligent organizations see a shift happening in the world and they decide they have to address that or do that, right? You saw a lot of this in the early days of cloud. They didn’t figure out a business problem or financial problem to move to cloud; they just saw all their peers doing it, so they put a stake in the sand and said, “We’re going to cloud.” And I think that’s a bad way to design the business operations. If your core isn’t a tech company, then, “What do you mean by that?” would be the first question I ask.

Corey: One thing I want to talk about because I don’t get to see it very often. I am almost always brought in to companies when they’re already running in the cloud—specifically, AWS since that is where I start and stop professionally these days—and they’re already there, and surprise, it costs money. You’re there earlier than I am; you are helping them get there in the first place. I’ve viewed for a while the idea that moving to cloud to save money is a losing proposition. If you ask me in good faith to say, “All right, in five years, will we make money or lose money on this journey?”

It really comes down to what answer do you want because I can make an extremely strong good-faith argument in either direction, but my honest opinion is that it’s a capability story, not a cost savings play. That is how I’ve come to view it, but given that I’m viewing it after the fact, and I’m only seeing a very specific example of it, I’m curious to know how you see it.

Joe: I would not recommend to a client to move to the cloud for the purpose of saving cost. If there’s something else leading it, scalability, elasticity, operational flexibility, whatever you’re looking at, that should be the primary goal. If you can also build it to save some costs, that’s fantastic. And there’s really two reasons I look at that. One is, IT should be a business enabler if you’re doing it right, and if you have something enabling your business driving revenue, why would you want to starve it of funding? Why would cost be your primary goal—cost savings?

And the second piece is, in my life, I always find that the success of a decision is 20% making the right decision and 80% making it the right decision after it’s made, right? It’s the effort afterwards to make it work that’s going to show you whether you’re getting the cost savings or not. It’s not easy to jump to cloud and create the new operational model that’s going to be the cheaper operational model, so if you’re not willing to do that work, once you’re in cloud, you’re not going to save money on it.

Corey: This episode is sponsored by our friends at Oracle Cloud. Counting the pennies, but still dreaming of deploying apps instead of “Hello, World” demos? Allow me to introduce you to Oracle’s Always Free tier. It provides over 20 free services and infrastructure, networking, databases, observability, management, and security. And—let me be clear here—it’s actually free. There’s no surprise billing until you intentionally and proactively upgrade your account. This means you can provision a virtual machine instance or spin up an autonomous database that manages itself, all while gaining the networking, load balancing, and storage resources that somehow never quite make it into most free tiers needed to support the application that you want to build. With Always Free, you can do things like run small-scale applications or do proof-of-concept testing without spending a dime. You know that I always like to put asterisks next to the word free? This is actually free, no asterisk. Start now. Visit snark.cloud/oci-free that’s snark.cloud/oci-free.

Corey: You have a, I would say, unpopular opinion on taking multi-cloud as an action item in the direction to go in. The reason I don’t call it that unpopular is because it echoes a lot of my own thinking on these things, and Lord knows, I have suffered the slings and arrows over the years for advocating such a thing, but what is your position on adopting multiple clouds?

Joe: So, if I was going to put it in the least objective possible terms, it would be, I want to be single architecture—single cloud in this case—unless. Right? I should be architecting for the simplest environment, I can build given my requirements. And so when I see clients try and jump into multi-cloud because it’s the buzzword or it’s something that a vendor is trying to sell them, multi-cloud is not a solution, it’s a necessity, in some cases.

Corey: My perspective has been to pick a provider—I don’t care which one—go all in until you have a reason to do something different. Multi-cloud is, in my experience, something that happens to you rather than something that is an intentional choice. But where your data winds up living is fundamentally where everything else is going to wind up centering around as well. The old-school procurement story of not wanting to be tied to one particular vendor because they’re going to soak you is a good piece of advice and I apply it in almost every IT decision, except when it comes to cloud. Because the pattern is different, the model is different, the way the discounting works is radically different.

And maybe that’s just because I haven’t done a lot of this work in traditional IT, but is this also the wrong approach, going back to the world of data centers and networking vendors and server vendors and the like, or is it really a different world?

Joe: No, I think it’s very much the same world. I’m religious about standardization wherever possible because it reduces the operational friction across the board that gets ignored in a lot of these costs. And that operational friction can end up in headcount and salary and cost that you see, but it also ends up in frustration for those teams, complexity of what you do, and another form of lock-in that prevents you from modernizing that infrastructure. So, anywhere you can find a standard single vendor that works—and it’s going to have some caveats, like everything—I would. And that’s not to say you should always standardize on everything; it’s standardize in less.

Corey: One of the things that I tend to see as far as a multi-cloud pattern that just doesn’t work is in no small part, very much an intentional choice—I believe—on the part of the cloud providers, where inbound data transfer is free; outbound costs an awful lot of money. And that, if for nothing other than basic economics has acted as a brake on the adoption of those patterns, in many cases. Is that something that you experience as these companies are moving to cloud is something that they need to become accustomed to? Is that something they know going in and they just intrinsically accept it? How does that awareness play out?

Joe: So, I think you’re hitting on the biggest problem of multi-cloud is how do I get access to the data sitting in one cloud? Every cloud provider wants to give you cheap storage because once your data is there, you’re going to use their compute, their bandwidth, everything else. And so when I am working with a client that is looking at multi-cloud, the first thing we want to solve for is, where’s the demilitarized zone we can put your data that can serve it effectively to any cloud you’re using? Because most of the time, your apps aren’t going to work in isolation. And that tends to be a solvable problem, but one of the harder problems to solve, and one of the things I don’t see a lot of people thinking of first when they start to put apps in different clouds.

Corey: For me, when I was advising—lightly—on Cloud migrations and digital transformations as such, the problem wasn’t the technology or even the budget or the rest, it was the growing awareness that people were going to have to think about things in a different context. Tying it back to economics, for example, when you ask someone who’s in a data center and looking to move to cloud, “Okay, great. How much data per month are your app servers sending and receiving to the database servers?” And the answer? “Why on earth would I have to know that? Why would I care?”

And it’s oh, you’re very much about to care. There’s a reason I’m asking this. It’s a cultural transformation, much more than it is a technical one, in my experience. Do you find that that comes as a surprise to folks or by the time that they get serious enough about digital transformation to bring someone like you in that they’ve already checked the basic boxes?

Joe: I think we’ve improved a lot over time. I mean, I think there were great horror stories of they’re ready to flip the switch on a cloud migration, and then they talked to the CFO who has no desire to deal with an OpEx model, or something to that effect, right? So, I think we’ve moved a lot past that. But I think people are still very naive about the overall dependencies, the data transfer. I used to say you can ask any given customer how many applications they have, and if they can give you a ballpark, that’s amazing. So, to know what the dependencies are, what the data transfer rates [crosstalk 00:24:49]—

Corey: [crosstalk 00:24:49] start counting on it, and it’s like it’s one of those, “Yeah, don’t bother giving me specific count; just give me breadbox sizing. Are we talking dozens, hundreds, thousands, millions? At least give me an order of magnitude here.”

Joe: Right. And if you don’t know how many apps you have, how do you know how they communicate and how much data they transfer, and the rest? And oh, by the way,
figuring that all out is an expensive exercise.

Corey: Very often, I tend to view hybrid as something that no one intends to do, but they get there almost by accident where they start migrating some workloads, and it goes super well, then they realize, “Huh, I have a mainframe over there and there is no AWS/400 I can migrate it to, so we’re going to give up, call it hybrid, plant the flag, declare victory, and the end; we’re a hybrid now.” I feel like that is in many cases, what a multi-cloud… pattern might evolve to be. I think we’re still early enough in the cycle that moving from all-in on Cloud Provider A to all-in on Cloud Provider B isn’t an exercise most companies have undertaken. But it feels like that might be something that gives rise to a multi-cloud world, just because that is the pattern that people fall into turns out to be more of a trap than anything.

Joe: Yeah, I think we’re always more willing to spend $10 a month for eternity than $100 right now on a problem. So, we get this idea of we’re not going to take that legacy, monolithic app and re-architect it for the cloud; we’re going to leave it and run in a hybrid model. Over time you’re over-engineering; over time, you’re spending more money; over time, you’re not solving the problem. One of the things that, you know, here on my ranch I try and do is never do band-aid fixes because as soon as I go put a bandaid on something, it’s going to stay there until it breaks on me again. If you’re not going to fix it right the first time, you’re going to have challenges with it all the time.

Corey: It’s the idea of buying the best tool that you can find on this, when you buy the most expensive–or best tool—which is often the most expensive—it’s one of those you cry once, whereas if you’ve buy the crappy tool, every time you use it, it irritates you, but you can’t justify replacing it. It’s the same model. One thing that I keep smacking into, it on some level, makes me feel like a bit of a fraud because I’m here talking to companies about their AWS bill, where it starts where it stops, but regardless of how big or how small that bill is, it is always dwarfed by payroll expenses. And the hard part of cloud migrations and modernization is not, “Well, how do we move all the applications from the data center into the cloud?” Compared to, “We have 5000 employees who are working in the on-prem environment and know how that works, and cloud is something they find in the sky when they go outside once in a while. How do we get those people upskilled?” That seems to be the challenge of the age, right now. I am bounded to only the computery bits, as far as what I tend to explore. You’re not. How does staff upskilling and staff expertise point of impacting your work?

Joe: That’s a huge point, right? Your operational costs around your staff, staff tooling, and operations are always far exceeding any of your infrastructure costs, cloud or not. And I think one of the biggest hindrances I see to that is companies have this fear that if they train people and upskill them that they’re going to lose them. And, you know, I take a pretty hard stance on that, if you’re that worried about losing your people because you’re training them a little bit that, maybe you should fix your culture or your paychecks, or both. That’s a huge hindrance to it.

You have to train your people because they’re costing you more not knowing what you need to know. If they do leave, that happens, that’s business, that’s how things work. It’s more expensive to you over time to not be investing in the knowledge they need. And wherever you can carry your existing staff forward, you’re going to save a lot money over hiring that new staff, especially in this current market.

Corey: There is a reality as well—and I want to challenge you on this one a little bit—that if you have a team of people who are working in your data centers on various things, and let’s say their market rate is $60,000 a year—to pick a number arbitrarily—upskilling them to cloud-first is hard. And I want to be clear, not everyone either has the capacity or the desire to, “All right, I’m going to basically become a cloud developer now.” But for the folks who do and are able to make that transition, they’re making $60,000 a year but they’ve just learned a new skill that has a going market rate of perhaps $120,000 in that market. On some level it’s a well, I could go work somewhere else and double my pay. It’s you’d have to convince me that there was a strong compelling reason for them not to do it. If they were asking me for advice, like, why wouldn’t you? That’s one of those obvious type of answers in most scenarios. How do you square that circle?

Joe: There’s going to be some risk involved either way, so I’m not trying to shy away from that. But I think if you have people that generally like their job and what they do, people tend to not want to switch jobs as much. We all experience inertia and complacency, right, at some level. I think the second piece is, using the numbers you’re using as an example, if I’m making 60 today, and you train me for a $120,000 job, and somewhere along that line, when I showed the aptitude and have the skillset, you bump me from 60 to 80 or 90 without me asking, you just bought a level of loyalty for $30,000 a year cheaper than you would have bought my replacement. And that doesn’t mean I’m going to stay forever, but I’m really going to like where I’m at when I get a giant bump without coming into your office and demanding it.

Corey: I think that there’s a misunderstanding across a lot of sectors of the economy that employment is not strictly about the numbers. And I know that because in my 20s, I was in crippling credit card debt, and every career decision I made was around what had the biggest number on the paycheck. And there’s nothing inherently wrong with that approach, but it also didn’t serve me super well, in some scenarios. If I’m chasing—even now—the thing that pays me the absolute most money, yeah, it turns out that running a boutique consultancy is not the answer to that question. I could do a lot of things that are considerable more ethically dubious; I’d be miserable, but it would make more money in some respects.

Employees are in a very much a similar boat. It’s yeah, I could go make 10% more somewhere else, but I like what I’m working on. I like the people. I like the culture, I like the baseline level of respect the company has for me, and I like the fact that it’s not just empty words when they say that they invest in their people. And I think that is one of those things that really hits and convinces people that, yeah, is this place perfect? No, no place is, but that’s why I stay. And that counts for an awful lot and I think that gets overlooked.

Joe: I agree completely. And I think, you know, I want to be careful because there’s a level of money that shifts at, right? At some point, you got to pay the bills, you got to pay off the loans, you got to pay the mortgage. And so the more money to get to that level is extremely important. And probably the most important thing in your career choice. Once you hit comfort and normalcy—

Corey: Oh, yeah. Going from between 30,000 and 40,000 is very different than debating between 170 and 180. It’s a percentage thing, and there are certain steps at which point it is a dramatic lifestyle improvement. At other points, that same amount of money is more or less, it looks suspiciously like a rounding error. And it also depends on people’s individual situations, too. I want to be very clear, this is not in defense of underpaying people in any respect. I’m a huge fan of charge market rate and get more money if you possibly can.

Joe: Absolutely. And I think it’s a combination of those things. And you have to remember, it’s going to be different to different individuals, right? A single person with no intent on a family might be one hundred percent okay, with 80 hour weeks for the right money because they don’t have a whole lot of other commitments, right? Whereas it’s somebody else in a different set of boats is going to care more about a four-day work week or the rest.

So, I think two things would help companies maintain the talent, especially in a market like this, and that’s having a rounded out package that includes great salaries along with benefits, and probably providing some choice so that the individual can get what they’re really looking for within the big picture of the benefits package.

Corey: I really appreciate your spending the time to talk with me about all this today. If people want to learn more about what you’re up to and how you think about these and many other things, where’s the best place to find you?

Joe: I’d say so transformationcontinuum.com is probably the best place. I’m on Twitter, but I’ll warn you I’m a bit of a porcupine, so I’m not for everybody’s tastes.

Corey: A lot of that going around on this [laugh] conversation today. Thank you again for your time. I really do appreciate it.

Joe: This was fantastic. Thank you, Corey.

Corey: Joe Onisick, principal at transformation CONTINUUM. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry, insulting comment that I will only accept if you send it from 20,000 feet above sea level.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Ashleigh

Ashleigh Early is a passionate advocate for sales people and through her consulting, coaching, and The Other Side of Sales, she is devoted to making B2B sales culture more inclusive so anyone can thrive. Over the past ten years Ashleigh has led, built, re-built, and consulted for 2 unicorns, 3 acquisitions, 1 abject failure and every step in between. She is also the Head of Sales at the Duckbill Group! You can find Ashleigh on Twitter @AshleighatWork and more about the Other Side of Sales at Othersideofsales.com

Links:

  • Twitter: https://twitter.com/ashleighatwork
  • LinkedIn: https://www.linkedin.com/in/ashleighearly

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Couchbase Capella Database-as-a-Service is flexible, full-featured and fully managed with built in access via key-value, SQL, and full-text search. Flexible JSON documents aligned to your applications and workloads. Build faster with blazing fast in-memory performance and automated replication and scaling while reducing cost. Capella has the best price performance of any fully managed document database. Visit couchbase.com/screaminginthecloud to try Capella today for free and be up and running in three minutes with no credit card required. Couchbase Capella: make your data sing.

Corey: Today’s episode is brought to you in part by our friends at MinIO the high-performance Kubernetes native object store that’s built for the multi-cloud, creating a consistent data storage layer for your public cloud instances, your private cloud instances, and even your edge instances, depending upon what the heck you’re defining those as, which depends probably on where you work. It’s getting that unified is one of the greatest challenges facing developers and architects today. It requires S3 compatibility, enterprise-grade security and resiliency, the speed to run any workload, and the footprint to run anywhere, and that’s exactly what MinIO offers. With superb read speeds in excess of 360 gigs and 100 megabyte binary that doesn’t eat all the data you’ve gotten on the system, it’s exactly what you’ve been looking for. Check it out today at min.io/download, and see for yourself. That’s min.io/download, and be sure to tell them that I sent you.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. My guest today does something that I, sort of, dabbled around the fringes of once upon a time, but then realized I wasn’t particularly good at it and got the hell out of it and went screaming into clouds instead. Ashleigh Early is the Head of Sales here at The Duckbill Group. Ashleigh, thank you for joining me.

Ashleigh: Thanks for coming on and running, screaming from my chosen profession [laugh]. You’re definitely not the only one.

Corey: Well, let’s be clear here; there are two ways that can go because sure, I used to dabble around in sales when I was, basically, trying to figure how to not starve to death. But I also used to run things; it’s basically a smart team. I was managing people and realized I was bad at that, too. So, really, that’s, sort of, an open-ended direction. We can go either side and…

But, let’s go with sales. That seems like a more interesting way for this to play out. So, you’ve been here for—what is it now—it feels like ages, but my awareness for the passing of time in the middle of a global panini is relatively not great.

Ashleigh: Yeah. I think we’re at day—what is it—1,053 of March 2020? So, time is irrelevant; it’s a construct; I don’t know. But, technically, by the Gregorian Calendar, I think I’m at six months.

Corey: It’s very odd to me, at least the way that I contextualized doing this. Back when I started what became The Duckbill Group, I was an independent consultant. It was, more or less, working people I knew through my network who had a very specific, very expensive problem: The AWS bill is too high. And I figured, this is genius. It is the easiest possible sale in the world and one of the only scenarios where I can provably demonstrate ROI to a point where, “Bring me in; you will inherently save money.”

And all of that is true, but one of things I learned very quickly was that, even with the easiest sale of, “Hi. I’d like to sell you this bag of money,” there is no such thing as an easy enterprise sale. There is nuance to it. There is a lot of difficulty to it. And I was left with the, I guess, driving question—after my first few months of playing this game—of, “How on earth does anyone make money in this space?”

The reason I persisted was, basically, a bunch of people did favors for me, but they didn’t owe me at all. It was, “Oh, great. I’ll give them the price quote.” And they’re, like, “Oh, yeah.” So cool, they turned around and quoted that to their boss at triple the rate because, “Don’t slit your own throat on this.” They were right. And not for nothing, it turns out when you’re selling advice, charging more for it makes it likelier to succeed as a project.

But, I had no idea what I was doing. And, like most engineers on Twitter, I look at something I don’t understand deeply myself, and figure, “Oh. Well, it’s not engineering, therefore, it’s easy.” Yeah, it turns out that running a business is humbling across a whole bunch of different axes.

Ashleigh: I wouldn’t even say, it’s not running a business; it’s working with humans. Working with humans is humbling. If you’re working with a machine or even something as simple as, like, you know, you’re making a product. It’s follow a recipe; it’s okay. Follow the instructions. I do A, then B, then C, then D, unless you don’t enjoy using the instructions because you don’t enjoy using instructions. But you still follow a set general process; you build a thing that comes out correctly.

The moment that process is, talk to this person, and then Person A, then Person B, then Person C, then Person D, then Back to Person A, then Person D, and then finally to Person E, everything goes to heck in a handbasket. That’s what really makes it interesting. And for those of us who are of a certain disposition, we find that fascinating and enthralling. If you’re of another disposition, that’s hell on earth [laugh]. So, it’s a very—yeah, it’s a very interesting thing.

Corey: Back when I was independent, and people tried to sell me things—and yeah, sometimes it worked. It was always interesting going through various intake funnels and the rest. And, like, “Well, what role do you hold in the organization? Do you influence the decision? Do you make the decision? How many people need to be involved in the rest?”

And I was looking around going, “How many people do you think fit in my home office here? Let’s be serious.” I mean, there are times I escalated to the Chihuahua because she’s unpleasant and annoying and basically, sometimes so are people. But that’s a separate topic for later. But it became a very different story back as the organizational distance between the people that needed to sign off on a sale increased.

Ashleigh: Mm-hm. Absolutely. And you might have felt me squirm when you described those questions because one of my biggest pet peeves is when people take sales terminology and directly use that with clients. Just like if you’re an engineer and you’re describing what you do, you’re not going to go home and explain to your dad in technical jargon what exactly; you’re going to tell him broad strokes. And if they’re interested, go deeper and deeper; technical, more technical.

I hate when salespeople use sales jargon, like, “What’s your role in the organization? Are you the decision-maker?” Don’t—mmm. There are better ways to deal with that. So, that’s just a sign of poor training. It’s not the sales rep’s fault; it’s his company’s fault—their company’s fault. But that’s a different thing.

It’s fascinating to me, kind of, watching this—what you said spoke of two things there. One is poor training, and two, of a lack of awareness of the situation and a lack of just doing a little bit of pre-work. Like, you do five seconds of research on Corey Quinn, you can realize that the company is ten to 15 people tops. So, it makes sense to ask a question around, “Hey, do you need anyone else to sign off before we can move forward with this project?”

That tells me if I need to get someone for technical, for budget, for whatever, but asking if you’re a decision-maker, or if you’re influencing, or if you’re doing initial research, like, that’s using sales terminology, not actually getting to the root of the problem and immediately making it very clear, you didn’t do any actual research in advance, which is not—in modern selling—not okay.

Corey: My business partner, Mike, has a CEO job title, and he’ll get a whole bunch of cold outreach constantly all day, every day. I conducted a two-week experiment where in front of my Chief Cloud Economist job title, I put ‘CTO/’ just to see what would happen, and sure enough, I started getting outreach left, right, up, down, and sideways. Not just for things that a CTO figure might theoretically wind up needing to buy, but also, job opportunities for a skill set that I haven’t dusted off in a decade.

So, okay. Once people can have something that hits their filters when you’re searching for very specific titles, then you wind up getting a lot more outreach. But if you create a job title that no one sensible would ever pick for themselves, suddenly a lot of that tends to go by the wayside. It shined a light on how frustratingly dreary a lot of the sales prospecting work really can be from—

Ashleigh: Oh, yeah.

Corey: —just from the side of someone who gets it. Now, I’m not exaggerating when I say that I did work in sales once upon a time. Not great at it, but one of the first white-collar-style jobs that I had was telemarketing, of all things. And I was spectacular at it because I was fortunate enough to be working on a co-branded affinity credit card that was great, and I had the opportunity to position it as a benefit of an existing membership or something else people already had. I was consistently top-ten out of 400 people on a shift, and it was great.

But it was also something that was very time-limited, and if you’re having an off day, everything winds up crumbling. And, eventually, I drifted off and started doing different things. But I’ve never forgotten those days. And that’s why it just grinds my gears both to see crappy sales stuff happening, and two, watching people on Twitter—particularly—taking various sales-prospect outreach for a drag. And it’s—

Ashleigh: Oh, God. Yeah.

Corey: —you know, not everyone is swimming in the ocean of privilege that some of the rest of us are. And understand that you’re just making yourself look like a jerk when you’re talking to someone who is relatively early-career and didn’t happen to google you deeply enough before sending you an email that you find insulting. That bugs me a fair bit.

Ashleigh: And I think part of that is just a lack of humanity and understanding. Like, there’s—I mean, I get it; I’m the first person to be jumping on Twitter and [unintelligible 00:08:41] when something goes down, or something’s not working, and saying, you know—I’m the first one to get angry and start complaining. Don’t get me wrong. However, what I think a lot of people—it’s really easy to dehumanize something you don’t see very often, or you’re not involved in directly. And I find it real interesting you mentioned you worked in, you know, doing telemarketing.

I lasted literally two weeks in telemarketing. I full-on rage-quit. It was a college job. I worked in my college donations center. I lasted two weeks, and I fully walked out on a shift. I was, like, “Screw this; I’m never doing anything like that ever again. I hate this.”

But what I hated about it was I hated the lack of connection. I was, like, I’m not just going to read some scripts and get yelled at for having too much banter. Like, I’m getting money; what do you care? I’m getting more money than other people. Maybe they’re not making as many calls, but I’m getting just as much, so why do you care how I do this?

But what really gets me is you have to remember—and I think a lot of people don’t understand how, kind of, most large, modern sales organizations work. And just really quickly giving you a very, very generic explanation, the way a lot of organizations work is they employ something called SDRs or Sales Development Reps. That title can be permeated in a million different ways. There’s ADRs, MDRs, BDRs, whatever. But basically, it’s their job to do nothing but scour the internet using, sometimes, actual, like, scripts.

Sometimes they use LinkedIn; sometimes they have—they purchase databases. So, for example, like, you might change your title on LinkedIn, but it’s not changing in the database. Just trust me Corey, they have you flagged as a CTO. Sorry. What [crosstalk 00:10:16].

Corey: My personal favorite is when I get cold outreach asking me on the phone call about whether we have any needs for whatever it is they happen to be selling at—and then they name a company that I left in 2012. I don’t know how often that database has been sold and resold and sold onwards, yet again. And it’s just, I work in tech. What do you think the odds are that I’m still in the same job I was ten years ago? And I get that it happens, but at some point, it just becomes almost laughable.

Ashleigh: Yeah. If you work in a company—that when in doubt—I tell every sales, kind of, every company team that I work with—do not use those vendors. Ninety percent of them are not very good; they’re using old databases; they don’t update. You’re better off paying for a database that is subscription-based because then, literally, you’ve got an SLA on data quality, and you can flag and get things fixed. The number one sales-data provider, I happen to know for a fact, I actually earned, I think, almost $10,000 in donations to a charity in—what was this—this was 2015 because I went through and did a scrub of are RCRM versus I think, LinkedIn or something else, and I flagged everything that wasn’t accurate and sent it back to them.

And they happened to have a promotion where for every—where you could do a flag that wasn’t accurate because they were no longer at the company. They would donate a buck to charity, and I think I sent them, like, 10,000 or something. [unintelligible 00:11:36] I was like, “None of these are accurate.” And they’re, like, you know? And they sent me this great email, like, “Thank you for telling us; we really appreciate it.”

I didn’t even know they were doing this promotion. They thought I’d be saving up for it. And I was, like, “No, I just happened to run this analysis and thought you’d want to know.” So, subscriptions—

Corey: You know, it turns out computers are really fast at things.

Ashleigh: Yeah, and I was very proud I figured out how to run a script. I was, like, “Yay. Look at me; I wrote a macro.” This was very exciting for—the first—God, the first five or so years of my sales career, I’ve consistently called myself a dumb salesperson because I was working in really super-technical products. I worked for Arista Networks, FireEye, Bromium, you know, PernixData. I was working in some pretty reasonably hard tech, and I’d always, kind of, introduced myself, I definitely talked about my technical aptitude because I have a degree in political science and opera. These are not technical fields, and yet here I am every day, talking about, you know, tech [crosstalk 00:12:25].

Corey: Well, if the election doesn’t pan out the way you want, why don’t you sing about it? Why not? You can tie all these things together.

Ashleigh: You can. And, honestly, there have several points—I’ve done a whole other shows on, like, how those two, seemingly, completely disparate things have actually been some of the greatest gifts to my career. And most notably, I think, is the fact that I have my degree in political science as a Bachelor of Science, which means I have a BS in BS, which is incredibly relevant to my career in a lot of different ways.

Corey: This episode is sponsored by our friends at Oracle Cloud. Counting the pennies, but still dreaming of deploying apps instead of “Hello, World” demos? Allow me to introduce you to Oracle’s Always Free tier. It provides over 20 free services and infrastructure, networking, databases, observability, management, and security. And—let me be clear here—it’s actually free. There’s no surprise billing until you intentionally and proactively upgrade your account. This means you can provision a virtual machine instance or spin up an autonomous database that manages itself, all while gaining the networking, load balancing, and storage resources that somehow never quite make it into most free tiers needed to support the application that you want to build. With Always Free, you can do things like run small-scale applications or do proof-of-concept testing without spending a dime. You know that I always like to put asterisks next to the word free? This is actually free, no asterisk. Start now. Visit snark.cloud/oci-free that’s snark.cloud/oci-free.

Ashleigh: Yeah, so wrapping up, kind of, how modern-skills organizations work, most companies’ employees can be called BDRs, and they’re typically people who have less than five years of sales experience. They, rightly or wrongly, tend to be people in their early-20s who have very little training. Most people get SDRs on phones within a week, which means—

Corey: These are the people that are doing the cold outreach?

Ashleigh: —they’ve gotten maybe five or six hours of product training. Hmm? Sorry.

Corey: These are the people who are doing the cold outreach?

Ashleigh: These are the people who are doing the cold outreach. So, their whole job is just to get appointments for account execs. Account execs make it—again; tons of different names, but these are the closers. They’ll run you through the sales cycle. They typically have between five and thirty years of experience.

But they’re the ones depending on how big your company is. [unintelligible 00:13:35] the bigger your company, typically the more experience your sales rep’s going to have in terms of managing most separate deal cycles. But what ends up happening is you end up with this SDR organization—this is where I’ve spent most of my career is helping people build healthy sales-development organizations. In terms of this churn-and-burn culture where you’ve got people coming in and basically flaming out because they go on Twitter or—heaven forbid—Reddit and get sales advice from these loud-mouthed, terrible people, who are telling them to do things that didn’t work ten years ago, but they then go try it; they send it out, and then their prospects suddenly blasting them on Twitter.

It’s not that rep’s fault that they got no training in the first place, they got no support, they just had to figure it out because that’s the culture. It’s the company’s fault. And a lot of times, people don’t—there was a big push against this last year, I think, within the sales community against other sales leaders doing it, but now, it’s starting to spread out. Like, I have no problem dragging someone for a really terrible email. Anonymize the company; anonymize the email. And, if you want to give feedback, give it to them directly. And you can also say, “I’m going to post this, but it’s not coming back to you.” And tell them, like—

Corey: Whenever I get outreach from—

Ashleigh: “Get out of that terrible company.”

Corey: Yeah. Whenever I get outreach from AWS for a sales motion or for recruiting or whatnot. I always anonymize the heck out of the rep. It’s funny to me because it’s, “Don’t you know who I am?” It is humorous, on some level. And it’s clear that is a numbers game, and they’re trying to do a bunch of different things, but a cursory google of my name would show it. It’s just amusing.

I want to be clear that whenever I do that, I don’t think the rep has done anything wrong. They’re doing exactly what they should. I just find it very funny that, “Wait, me? Work at an AWS? The bookstore?” It seems like it would be a—yeah. Yeah, the juxtaposition is just hilarious to me. They’ve done nothing wrong, and that’s okay. It’s a hard racket.

I remember—at least they have the benefit over my first enterprise sales job where I was selling tape drives into the AS/400 market, competing against IBM on price. That was in the days of “No one ever gets fired for buying IBMs.” So, yeah. The place you want to save money on is definitely the backup system that’s going to save all of your systems. I made one sale in my time there—and apparently set a company record because it wasn’t specifically aimed at the AS/400—and I did the math on that and realized, “Huh, I’d have to do two of these a month in order to beat the draw against commission structure that they had.”

So, I said, “To hell with this,” and I quit. The CEO was very much a sales pro, and, “Well, you need to figure out whether you’re a salesperson or not.” Even back then, I had an attitude problem, but it was, “Yeah, I think that—oh, I know that I am. It’s just a question is am I going to be a salesperson here?” And the answer is, “No.” It [laugh]—

Ashleigh: Yeah.

Corey: It’s a two-way street.

Ashleigh: It is. And I say this all the time to people who—I work with a lot of salespeople now who are, like, “I don’t think sales is for me. I don’t know, I need [unintelligible 00:16:24]. The past three companies didn’t work.” The answer isn’t, “Is sales for you?”

The answer is, “Are you selling the right thing at the right place?” And one of the things we’ve learned from the ‘Great Recession’ and the ‘Great Reshuffling’ in everything is there’s no reason to stay at a terrible company, and there’s no reason to stay at a company where you’re not really passionate and understand what you’re selling. I joked about, you know, I talked down about myself for the first bit of my career. Doesn’t mean I didn’t—like, I might not understand exactly how heuristics work, but I understand what heuristics are. Just don’t ask me to design any of them.

You know, like, you have to understand and you have to be really excited about it. And that’s what modern sales is. And so, yes, you’re going to get a ton of the outreach because that’s how people—it still works. That’s why we all still get Nigerian prince emails. Somebody, somewhere, still clicks those things, sadly. And that gets me really angry.

Corey: It’s a pure numbers game.

Ashleigh: Exactly. Ninety percent if enterprise B2B sales is not that anymore. Even the companies that are using BDRs—which is most of them—are now moving to what’s called ‘account-based selling’. We’re using hyper-personalized messaging. You’re probably noticing videos are popping up more.

I’m a huge fan of video. I think it’s a great way to force personalization. It’s, like, “Hi. Corey, I see you. I’m talking to you. I’ve done my research. I know what you’re doing at The Duckbill Group and here’s how I think we can help. If that’s not the case, no worries. Let me know; I’ll leave you alone.” That’s what selling should be.

Corey: I have yet to receive one of those, but I’m sure it’ll happen now that I’ve mentioned that and put that out into the universe.

Ashleigh: Probably.

Corey: What always drove me nuts—and maybe this is unfair—but when I’m trying to use a product, probably something SaaS-based—and I see this a lot—where, first, if you aren’t letting me self-serve and get off with the free tier and just start testing something, well, that’s already a ding against you because usually I’m figuring this out at 2 o’clock in the morning when I can’t sleep, and I want to work on something. I don’t want to wait for a sales cycle, and I have to slow things down. Cool. But at some point, for sophisticated customers, you absolutely need to have a sales conversation. But, okay, great. Usually, I encounter this more with lead magnets or other things designed to get my contact info.

But what drives me up a wall, when they start demanding information that is very clearly trying to classify me in their sales funnel, on some level. I’ll give you my name, my company, and my work email address—although I would think that from my work email address, you could probably figure out where I work and the rest—but then there are other questions. How big is your company? What is your functional role within the company? And where are you geographically?

Well, that’s an interesting question. Why does that matter in 2022? Well, very often leads get circulated out to people based upon geography. And I get it, but it also frustrates me, just because I don’t want to have to deal with classifying and sorting myself out for what is going to be a very brief conversation [laugh] with a salesperson. Because if the product works, great, I’m going to buy. If it doesn’t work, I’m going to get frustrated and not want to hear from you forever.

Which gets to my big question for you—and please don’t take the question as anything other than the joking spirit in which it’s intended—but why are so many salespeople profoundly annoying?

Ashleigh: I would—uh, hmm.

Corey: Sales processes is probably the better way to frame it because—

Ashleigh: I was going to say, “Yeah, it’s not the people; it’s the process.” So—

Corey: —it’s not the individual’s fault, as we’ve talked about it.

Ashleigh: —yeah, I was going to say, I was, like, “Okay, I think it’s less the people; more of the processes.” And processes that will make [crosstalk 00:19:37]—

Corey: Yeah. It expresses itself as the same person showing up again and again. But that is not—

Ashleigh: Totally.

Corey: —their fault. That is the process by which they are being measured at as a part of their job. And it’s unfair to blame them for that. But the expression is, “This person’s annoying the hell out of me, what gives?”

Ashleigh: “Oh, my gosh. Why does she keep [unintelligible 00:19:51] my inbox? Leave me alone. Just let me freaking test it.” I said, “I needed two weeks. Just let me have the two weeks to freaking test the thing. I will get back to you.” [unintelligible 00:19:58] yeah, no, I know.

And even since moving into leadership several years ago, same thing. I’m like, “Okay, no.” I’ve gotten to the point where I’ve had several conversations with salespeople. I’m like, “I know the game. I know what you’re trying to do. I respect it. Leave me alone. I promise I will get back to you, just lea”—I have literally said this to people. And the weird thing is most salespeople respect that. We really respect the transparency on that.

Now, the trick is what you’re talking about with lead capture and stuff like this, again, it comes down to company’s design and it comes down to companies who value the buyer experience and customer journey, and companies who don’t. And this, I think, is actually more driven by—in my humble opinion—our slightly over-reliance on venture capital, which is all about for a gathering of as much data as possible, figuring out how to monetize it, and move from there. In their mind, personal experience and emotion doesn’t really factor into that equation very much, so you end up with these buyer journeys that are less about the buyer and more about getting them from click to purchase as efficiently as possible in terms of company resources, which includes salespeople time. So, as to why you have to fill out all those things, that just to me reeks of a company that maybe doesn’t really understand the client experience and probably is going to have a pretty, mmm, support program as well, which means the product had better be really freaking good for me to buy it.

Corey: To be clear, at The Duckbill Group, we do not have a two-in-the-morning click here and get you onboarded. Turns out that we have yet to really see the value in building a shopping cart system, where you can buy, “One consulting please,” and call it good. We’re not quite at the level of productizing our offering yet and having conversations is a necessary part of what we do. But that also aligns with our customer expectation where there is not a general expectation in this industry that you can buy a full-on bespoke consulting engagement without talking to a human being. That, honestly, if someone trying to sell someone such a thing, I would be terrified.

Ashleigh: Yeah, run screaming. Good Lord. No, exactly. And that’s one of the reasons I love working with this team and I love this problem is because this isn’t a quick, you know, download, install, and save, you know, save ten percent on your AWS bill by installing Duckbill Group. It ain’t that simple. If it were that simple, like, AWS wouldn’t have the market cap it does.

So, that’s one of the things I love. I love really meaty problems that don’t have clean answers, and specifically have answers that look slightly different for everybody. I love those sort of problems. I’ve done the highly prioritized stuff: Click here, buy, get it on the free tier, and then it’s all about up-sale, cross-sale as needed. Been there, done that; that’s fun, and that’s a whole different bucket of challenges, but what we’re dealing with every single day on the consulting’s of The Duckbill Group is far more nuanced and far more exciting because we’re also seeing some truly incredible architecture designs. Like, companies who are really on the bleeding edge of what they’re doing. And it’s just really fun—

Corey: Cost and architecture are the same thing in the Cloud.

Ashleigh: —[crosstalk 00:22:59] that little—

Corey: It’s a blast to see it.

Ashleigh: It’s so much fun. It’s, it’s, it’s… the world’s best jigsaw puzzle because it covers, like, every single continent and all these different nuances, and you got to think about a ‘ephemerality,’ which is my new favorite word. So…

Corey: It’s fun because you are building a sales team here, which opens up a few interesting avenues for me. For one, I don’t have to manage and yell at individual salespeople in the same way. For example, we talk about it being a process and not a person thing. We’re launching some outbound sales work and basically, having the person to talk to about that process—namely you—means that I don’t need to be hovering over people’s shoulders the way I felt that I once did, as far as what are we sending people? These passive-aggressive drip campaigns of, “Clearly, you don’t mind lighting money on fire. If that changes, please let me know.”

It’s email eight in a sequence. It’s no. This stuff has an implicit ‘Love, Mike and Corey’ at the bottom of everything that comes out of this company, and it represents us on some respect. And let’s be clear, we have a savvy, sophisticated, and more-attractive-than-the-average audience listening to all of these shows. And they’ll eat me alive if we start doing stuff like that—

Ashleigh: Oh, yeah.

Corey: —not to mention that I find it not particularly respectful of their time and who they are. It doesn’t work, so we have to be very conscious of that. The fact that I never had to explain that concept in any depth to you made bringing you in one of the easiest decisions we’ve ever made.

Ashleigh: Well, I think it helped—I think in one of my interviews I went off on the ‘alligator email,’ which is this infamous email we’ve all gotten, which is basically, like, you know, “Hi. I haven’t heard from you yet, so I want to know which one of these three scenarios has happened to you. One, you’re not interested in my product but didn’t have the balls to email me and say that you’re not interested. Two, you’re no longer in this position, in which case, you’re not going to read this email anyway. Or three, you’re being chased by an alligator, and I should call animal control because you need help.” This email was—

Corey: He, he, he, hilarious.

Ashleigh: Ugh. And there’s variations of it. And I’ve seen variations of it that are very well done and are on brand and work with the company. I’ve seen variations that could be legitimately, I think, great humor. And that’s great.

Humor in emails and humor in sales is fantastic. I have to shout out my friend, Jon Selig up in Canada, who actually, literally, does workshops on how sales teams can integrate humor into their prospecting. It’s freaking brilliant. But—

Corey: Near and dear to my heart.

Ashleigh: —if you’re not actually trained in that stuff, don’t do it. Don’t do the alligator email. But I think I went off on that during one of our interviews just because I was just sick of seeing these things. And what kills me, again, it comes back to the beginning, is people who have no training, no experience coming in—I mean, it really kills me, too, because there’s a real concerted effort in the sales community to get more diverse people into sales to, kind of, kill the sales bro just by washing them out, basically. And so, we’re recruiting hard with veterans, with black and other racial minority groups, LGBTQ communities, all sorts of things, and indigenous peoples.

And so, we’re bringing people that also are maybe a little bit more mature, a little bit older, have families they’re supporting, and we’re throwing them in a role with no support and very little training. And then they wash out, and we wonder why. It’s, like, well, maybe because you didn’t—it’s, like, when I explain this to other people who aren’t in sales, like, “Really, imagine coming in to being hired for a coding job, being told you’re going to be trained on, you know, Ruby on Rails or C# or whatever it is we’re currently using”—my reference is probably super outdated—but then, being given a book, and that’s it. And told, “Learn it. And by the way, your first project is due in a month.” That’s what we’re doing in sales—

Corey: For a lot of folks, that’s how we learned in the engineering spaces, but let’s be clear, the people who do well in that, generally have tailwinds of privilege at their back. They don’t have headwinds of, “You suck at this.” It was, you’re-born-on-third-you-didn’t-hit-a-triple school-of-thought. It’s—

Ashleigh: Yeah.

Corey: —the idea of building an onboarding pipeline, of making this stuff more accessible to people earlier on is incredibly important. One of my, I guess, awakening moments as we were building this company was it turns out that if you manage salespeople as if they were engineers, it doesn’t go super well. Whereas, if you manage engineers like they’re salespeople, they quit—rage quit—cry, and call you out as being an abusive manager.

One of the best descriptions I ever heard from an advisor was that salespeople are sharks. But that’s not intended to be unkind. It is simply a facet of their nature. They enjoy the hunt; they enjoy chasing things down, and they like playing games. Whereas, as soon as you start playing games with your engineers on how much money they’re going to make this week, that turns out to be a very negative thing. It’s a different mindset. It’s about motivating people as whatever befits what it is that they want to be doing.

Ashleigh: It is. And the other thing is it’s a cultural conditioning. So, it’s really interesting to say, you know, “People,” you know, “Playing games.” We do enjoy—there’s definitely some enjoyment of the competition; there’s the thrill of the hunt, absolutely, but at the same time, you want your salespeople to quit? Screw with their money.

You screw with their money; we will bail so fast it’ll make your head spin. So, it’s like, people think, “Oh, we love this.” No, it’s really more—think of it as we are gamblers.

Corey: Yeah. To be clear when I say, “Playing games with money,” I’m talking about the idea of, “Sell to a company in this profile this quarter, and we’ll throw a $5,000 bonus your way,” or something like that. It is if the business wants to see something, great, make it worth the sales team’s while to pursue it, or don’t be surprised when no one really cares that much about those things—

Ashleigh: Exactly.

Corey: It’s all upside. It is not about, “He, he. And if you don’t sell to this weird thing that I can’t really describe effectively to you, we’re going to cut your bet—” Yeah, that goes over like a lead balloon. As it should. My belief is that compensation should always go up, not down.

Ashleigh: Yeah. No, it should. Aside from that, here’s a fun stat—I believe this came out of Forrester, it might’ve been out of [Topel 00:28:54]; I apologize, I don’t remember exactly who said this, but a recent study found that less than 68 percent of sales reps make their quota every month. So, imagine that where if you’re—we have this thing called OTE, which is On Target Earnings. So, if you have this number you’re supposed to take home every month, only 68 percent of sales reps actually do that every month.

So, that means we live with this number as our target, but we’re living and budgeting anywhere from 30 to 50 percent below that. And then hoping and doing the work that goes in there. That’s what we’ve been conditioned to accept, and that’s why you end up with sales reps that use terms like ‘shark’ and are aggressive and are in your face and can get—[unintelligible 00:29:30]—

Corey: I didn’t realize it was pejorative.

Ashleigh: I know. No. But here’s the thing too, but somebody called it ‘commission breath,’ which I love. It’s, like, you can smell commission breath coming off us when we’re desperate. You totally can. It’s because of this antiquated way of building commissions.

And this is something that I—this was really obvious to me, and apparently, I was a little bit ahead of the curve. When I started designing comp plans, everyone told me, “You want to design a comp plan? Tie it to what you want them to do very specifically.” So, if you want them to move a pen, design a comp plan that they get a buck when they put the pen from the heel of your hand to the tips of your fingers. Then they get a buck. And then they can do that repeatedly. That’s literally how I was taught design comp plans.

In my head, that meant that I need to design it in such a way that it’s doable for my team because I don’t want my team worrying about how they’re going to put food on the table while they’re talking to a client because they’re going get commission breath and it’ll piss off the client. That’s not a good client experience; that’s not going to lead to good performance. Apparently—

Corey: Yeah. My concern as a business owner has nothing to do with salespeople making too much money. In fact, I am never happier than I am than paying out commissions. The concern, then, therefore has to become the, “Okay, great. How do I keep the salespeople from being inadvertently incentivized to sell something for $10 that costs me $12 to fulfill?”

It’s a question of what behaviors do you incentivize that align what they’re motivated by with what the company needs. And very often getting that wrong—which happens from time to time—is not viewed as a learning experience that it should be. But instead, “They’re just out to screw us.” And I’ve seen so many company owners get so annoyed whenever their salespeople outperform. But what did you expect? That is the positive outcome. As opposed to what? The underperforming sales rep that can’t close a deal? Please.

Ashleigh: Well, no. And let’s think about this too, especially if it’s tied to commission and you’re paying out commission. It’s, like, okay, commission is always some, sort of, percentage—depending on a lot of things—but some sort of percent of what they’re bringing in. If you design a comp plan that has you paying out more in commission than the sales that were earned to bring it in, that’s on you; you screwed up. And you need to either be honest and say, “I screwed up; I can’t pay this,” and know that you’re going to lose some sales reps, but you won’t lose as many as if you just refuse to pay it.

But, honestly, and I’m not even kidding, I know people. I’ve worked at a company that I happen to know did this. That literally fired people because they didn’t have the money to pay out the commission. And because they fired them before the commission was due to be paid out, then that person no longer had a legal claim to it. That’s common. So, the commission goes both ways.

Corey: To be clear, we’ve never done that, but I also would say that if we had, that’s a screaming red flag for our consultancy, given the nature of what it is that we do here. It turns out that when we’re building out comp plans, we model out various scenarios. Like, what is the worst way that this could wind up unfolding? And, okay, some of our early drafts it’s, yeah, it turns out that we would not be able to pay salaries because we wound up giving all of that in commission to people with uncapped upside. Okay, great.

But we’re also not going to cap people’s commissions because that winds up being a freaking problem, so how do we wind up motivating in a way that continues to grow and continues to incentivize the behaviors we want? And it turns out it’s super complicated which why we brought you in. It’s easier.

Ashleigh: Yeah, it’s a pain. But the other side of this too, I think, is there is another force at play here, which is finance. A lot of traditional finance modeling is built around that 50 to 70 percent of people hit commission. So, if all of the sudden, you design a comp plan such of a way that a hundred percent of the team is hitting commission, finance loses their shit. So, you have to make sure that when you’re designing these things, one of the things I learned, I learned the hard way—this is how I learned that not everyone does it this way—I built my first comp plan; my team’s hitting it.

My team’s overperforming, not a ton, but we’re doing really well. All of the sudden, I’m getting called to Finance and getting raked over the coals. And they’re like, “What did you do?” I’m like, “What do you mean what did I do? I designed a comp plan; we’re hitting goal. Why are you mad?” “Well, we only had this much budgeted for commission.”

And I was, like, “That’s not my fault.” “Well, that’s what historic performance was.” “Okay, well that’s not what we’re going to do going forward. We’re going to do this.” And they’re like, “Oh, well, you need to notify us if you’re going to change it like that.” And I was, like, “Wait a minute. You modeled so that my team would not hit OTE?” “Yes.” “That’s how you’ve always done this?” “Yes.” “Okay. Well, that’s not what we’re going to do going forward, and if that’s a problem, I’ll go find a door.” Because, no.

Especially when we’re talking about people who are living in extremely expensive areas. I spent most of my career living and working in San Francisco, managing teams of people who made less than six figures. And that’s rough when you’re paying two grand in rent every month. And 60 percent of your pay is commission. Like, no. You need to know that money’s coming.

So, I talk about modern sales a lot because that’s what I’m trying to use because there’s Glengarry Glen Ross, kind of, Wolf of Wall Street school, which is not how anyone behaves anymore, and if you’re in an environment that’s like that or treats your salespeople like that? Please leave. And then you’ve got modern sales, which is all about, “Okay, let’s figure out how we can set up our salespeople to be the best people they can be to give our clients the best experience they can.” That’s where you get top performance out of, and that’s where you never run into the terrible emails with the alligators, and the, “Clearly you like lighting piles of money on fire.” That’s where you don’t get emails to Corey Quinn asking him if he’s interested in coming to work for AWS, the book company.

It’s by incentivizing the people and creating good humans where they can really thrive as salespeople and as people in general. The rest comes with time. But, it’s this whole, new way of looking at things. And it’s big, and it’s scary, and it costs more upfront, but you get more on the back end every single time.

Corey: Not that you care about this an awful lot, but you have your own podcast that talks about this, The Other Side of Sales. What inspired you to decide, not just to build sales teams through a different lens, but also to, “You know what? I’m going to go out and talk into microphones through the internet from time to time.” Which, let’s be clear, it takes a little bit of a certain warped perspective. I say this myself, having done this far too often.

Ashleigh: Yeah. No, it’s a fun little origin story. So, I’m a huge Star Trek geek; obsessive. And I was listening to a Star Trek podcast run by a couple of guys who are a little bit embarrassed to run a Star Trek podcast, called The Greatest Generation. Definitely not safe for work, but a really good podcast if you’re into Star Trek at all.

And they always do, kind of, letters at the end of the shows. And one of the letters at the end of the show one day was, “Hey, I was really inspired by you guys and I started my own podcast on this random thing that I am super excited about.” And I’m literally driving in the car with my husband, and I’m, like, “Huh. I don’t know why I’m not listening to sales podcasts. I listen to enough of these other random ones.” Jumped online, pulled up a list of sales podcasts, and I think I went through three or four articles of, like, every sales podcast that was big. And this was, like, January of 2019.

Corey: “By Broseph McBrowerson, but Everyone Calls Him ‘Browie.’” Yeah.

Ashleigh: Literally, there was, Conversations with Women in Sales with the late, great—with the amazing Lori Richardson, who’s now with it, but she took over for a mentor of mine who passed in 2020, sadly. But there was that, and then there was one other that was hosted by a husband-and-wife team. And that was it out of, like, 30 podcasts. And [laugh] so it was this moment of, like, epiphany of, like, “I can start my own podcast,” and, “Oh, I probably need to,” because, literally, no one looks or sounds like someone who I would actually want to hang out with ever, or do business with, in a lot of cases. And that’s really changed. I’m so grateful.

But really, what it came down to was I didn’t feel there was a podcast for me. There wasn’t a podcast I could listen to about sales that could help me, that I felt like I identified with. So, I was, like, “All right, fine. I’ll start my own.” I called up a friend, and she was, literally, going through the same thing at the same time, so we said, “Screw it. We’ll do our own.”

We went full Bender from Futurama. We’re like, “Just screw it; we’ll have our own podcast… with liquor… and heels… and honest conversations that happens to us every day,” and random stuff. It’s a lot of fun. And we’ve gone through a few iterations and it’s been a long journey. We’re about to hit our hundredth episode, which is really exciting.

But yeah, we’re—The Other Side of Sales is on a mission to make B2B sales culture truly inclusive so everyone can thrive, so, our conversations are all interviews with amazing sales pros who are trying to do amazing things and who are 90—I think are over 90 percent—are from a minority background, which is really exciting to, kind of, try and shift that conversation from Broseph McBrowerson. Our original tagline was the ‘anti-sales bro’ podcast, but we thought that was a little too antagonistic. So…

Corey: Yeah, being a little too antagonistic is, generally, my failure mode, so I hear you on that. I really want to thank you for taking so much time out of your day to speak with me. Because—well, not that I should thank you. It’s one of those, I should really turn around and say, “Wait a minute. Why aren’t you selling things? Why are you still talking to me?” But no—

Ashleigh: No, I’m waiting for you to say, “Back to work.”

Corey: Do appreciate your—exactly. I think that’s a different podcast. Thank you so much for your time. If people want to learn more, where’s the best place to find you?

Ashleigh: Well, definitely please go check out duckbillgroup.com. We would love to talk to with you about anything to do with your AWS bill. Got a ton of resources on there around how to get that managed and sorted.

If you’re interested in connecting with me you can always hit me up at—I’m on Twitter @ashleighatwork, which is another deep-cut Star Trek reference, or you can hit me up at LinkedIn. Just search Ashleigh Early. My name is spelled a little weird because I’m a little weird. It’s A-S-H-L-E-I-G-H, and then Early, like ‘early in the morning.’

Corey: And links to all of that will wind up in the [show notes 00:39:11]. Thanks so much for your time. It’s appreciated.

Ashleigh: This has been fun; we’ll do it again soon.

AndIf your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Tim

Timothy William Bray is a Canadian software developer, environmentalist, political activist and one of the co-authors of the original XML specification. He worked for Amazon Web Services from December 2014 until May 2020 when he quit due to concerns over the terminating of whistleblowers. Previously he has been employed by Google, Sun Microsystemsand Digital Equipment Corporation (DEC). Bray has also founded or co-founded several start-ups such as Antarctica Systems.

Links Referenced:

  • Textuality Services: https://www.textuality.com/
  • laugh]. So, the impetus for having this conversation is, you had a [blog post: https://www.tbray.org/ongoing/When/202x/2022/01/30/Cloud-Lock-In
  • @timbray: https://twitter.com/timbray
  • tbray.org: https://tbray.org
  • duckbillgroup.com: https://duckbillgroup.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Vultr. Spelled V-U-L-T-R because they’re all about helping save money, including on things like, you know, vowels. So, what they do is they are a cloud provider that provides surprisingly high performance cloud compute at a price that—while sure they claim its better than AWS pricing—and when they say that they mean it is less money. Sure, I don’t dispute that but what I find interesting is that it’s predictable. They tell you in advance on a monthly basis what it’s going to going to cost. They have a bunch of advanced networking features. They have nineteen global locations and scale things elastically. Not to be confused with openly, because apparently elastic and open can mean the same thing sometimes. They have had over a million users. Deployments take less that sixty seconds across twelve pre-selected operating systems. Or, if you’re one of those nutters like me, you can bring your own ISO and install basically any operating system you want. Starting with pricing as low as $2.50 a month for Vultr cloud compute they have plans for developers and businesses of all sizes, except maybe Amazon, who stubbornly insists on having something to scale all on their own. Try Vultr today for free by visiting: vultr.com/screaming, and you’ll receive a $100 in credit. Thats V-U-L-T-R.com slash screaming.

Corey: Couchbase Capella Database-as-a-Service is flexible, full-featured and fully managed with built in access via key-value, SQL, and full-text search. Flexible JSON documents aligned to your applications and workloads. Build faster with blazing fast in-memory performance and automated replication and scaling while reducing cost. Capella has the best price performance of any fully managed document database. Visit couchbase.com/screaminginthecloud to try Capella today for free and be up and running in three minutes with no credit card required. Couchbase Capella: make your data sing.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. My guest today has been on a year or two ago, but today, we’re going in a bit of a different direction. Tim Bray is a principal at Textuality Services.

Once upon a time, he was a Distinguished Engineer slash VP at AWS, but let’s be clear, he isn’t solely focused on one company; he also used to work at Google. Also, there is scuttlebutt that he might have had something to do, at one point, with the creation of God’s true language, XML. Tim, thank you for coming back on the show and suffering my slings and arrows.

Tim: Oh, you’re just fine. Glad to be here.

Corey: [laugh]. So, the impetus for having this conversation is, you had a blog post somewhat recently—by which I mean, January of 2022—where you talked about lock-in and multi-cloud, two subjects near and dear to my heart, mostly because I have what I thought was a fairly countercultural opinion. You seem to have a very closely aligned perspective on this. But let’s not get too far ahead of ourselves. Where did this blog posts come from?

Tim: Well, I advised a couple of companies and one of them happens to be using GCP and the other happens to be using AWS and I get involved in a lot of industry conversations, and I noticed that multi-cloud is a buzzword. If you go and type multi-cloud into Google, you get, like, a page of people saying, “We will solve your multi-cloud problems. Come to us and you will be multi-cloud.” And I was not sure what to think, so I started writing to find out what I would think. And I think it’s not complicated anymore. I think the multi-cloud is a reality in most companies. I think that many mainstream, non-startup companies are really worried about cloud lock-in, and that’s not entirely unreasonable. So, it’s a reasonable thing to think about and it’s a reasonable thing to try and find the right balance between avoiding lock-in and not slowing yourself down. And the issues were interesting. What was surprising is that I published that blog piece saying what I thought were some kind of controversial things, and I got no pushback. Which was, you know, why I started talking to you and saying, “Corey, you know, does nobody disagree with this? Do you disagree with this? Maybe we should have a talk and see if this is just the new conventional wisdom.”

Corey: There’s nothing worse than almost trying to pick a fight, but no one actually winds up taking you up on the opportunity. That always feels a little off. Let’s break it down into two issues because I would argue that they are intertwined, but not necessarily the same thing. Let’s start with multi-cloud because it turns out that there’s just enough nuance to—at least where I sit on this position—that whenever I tweet about it, I wind up getting wildly misinterpreted. Do you find that as well?

Tim: Not so much. It’s not a subject I have really had too much to say about, but it does mean lots of different things. And so it’s not totally surprising that that happens. I mean, some people think when you say multi-cloud, you mean, “Well, I’m going to take my strategic application, and I’m going to run it in parallel on AWS and GCP because that way, I’ll be more resilient and other good things will happen.” And then there’s another thing, which is that, “Well, you know, as my company grows, I’m naturally going to be using lots of different technologies and that might include more than one cloud.” So, there’s a whole spectrum of things that multi-cloud could mean. So, I guess when we talk about it, we probably owe it to our audiences to be clear what we’re talking about.

Corey: Let’s be clear, from my perspective, the common definition of multi-cloud is whatever the person talking is trying to sell you at that point in time is, of course, what multi-cloud is. If it’s a third-party dashboard, for example, “Oh, yeah, you want to be able to look at all of your cloud usage on a single pane of glass.” If it’s a certain—well, I guess, certain not a given cloud provider, well, they understand if you go all-in on a cloud provider, it’s probably not going to be them so they’re, of course, going to talk about multi-cloud. And if it’s AWS, where they are the 8000-pound gorilla in the space, “Oh, yeah, multi-clouds, terrible. Put everything on AWS. The end.” It seems that most people who talk about this have a very self-serving motivation that they can’t entirely escape. That bias does reflect itself.

Tim: That’s true. When I joined AWS, which was around 2014, the PR line was a very hard line. “Well, multi-cloud that’s not something you should invest in.” And I’ve noticed that the conversation online has become much softer. And I think one reason for that is that going all-in on a single cloud is at least possible when you’re a startup, but if you’re a big company, you know, a insurance company, a tire manufacturer, that kind of thing, you’re going to be multi-cloud, for the same reason that they already have COBOL on the mainframe and Java on the old Sun boxes, and Mongo running somewhere else, and five different programming languages.

And that’s just the way big companies are, it’s a consequence of M&A, it’s a consequence of research projects that succeeded, one kind or another. I mean, lots of big companies have been trying to get rid of COBOL for decades, literally, [laugh] and not succeeding and doing that. So—

Corey: It’s ‘legacy’ which is, of course, the condescending engineering term for, “It makes money.”

Tim: And works. And so I don’t think it’s realistic to, as a matter of principle, not be multi-cloud.

Corey: Let’s define our terms a little more closely because very often, people like to pull strange gotchas out of the air. Because when I talk about this, I’m talking about—like, when I speak about it off the cuff, I’m thinking in terms of where do I run my containers? Where do I run my virtual machines? Where does my database live? But you can also move in a bunch of different directions. Where do my Git repositories live? What Office suite am I using? What am I using for my CRM? Et cetera, et cetera? Where do you draw the boundary lines because it’s very easy to talk past each other if we’re not careful here?

Tim: Right. And, you know, let’s grant that if you’re a mainstream enterprise, you’re running your Office automation on Microsoft, and they’re twisting your arm to use the cloud version, so you probably are. And if you have any sense at all, you’re not running your own Exchange Server, so let’s assume that you’re using Microsoft Azure for that. And you’re running Salesforce, and that means you’re on Salesforce’s cloud. And a lot of other Software-as-a-Service offerings might be on AWS or Azure or GCP; they don’t even tell you.

So, I think probably the crucial issue that we should focus our conversation on is my own apps, my own software that is my core competence that I actually use to run the core of my business. And typically, that’s the only place where a company would and should invest serious engineering resources to build software. And that’s where the question comes, where should that software that I’m going to build run? And should it run on just one cloud, or—

Corey: I found that when I gave a conference talk on this, in the before times, I had to have a ever lengthier section about, “I’m speaking in the general sense; there are specific cases where it does make sense for you to go in a multi-cloud direction.” And when I’m talking about multi-cloud, I’m not necessarily talking about Workload A lives on Azure and Workload B lives on AWS, through mergers, or weird corporate approaches, or shadow IT that—surprise—that’s not revenue-bearing. Well, I guess we have to live with it. There are a lot of different divisions doing different things and you’re going to see that a fair bit. And I’m not convinced that’s a terrible idea as such. I’m talking about the single workload that we’re going to spread across two or more clouds, intentionally.

Tim: That’s probably not a good idea. I just can’t see that being a good idea, simply because you get into a problem of just terminology and semantics. You know, the different providers mean different things by the word ‘region’ and the word ‘instance,’ and things like that. And then there’s the people problem. I mean, I don’t think I personally know anybody who would claim to be able to build and deploy an application on AWS and also on GCP. I’m sure some people exist, but I don’t know any of them.

Corey: Well, Forrest Brazeal was deep in the AWS weeds and now he’s the head of content at Google Cloud. I will credit him that he probably has learned to smack an API around over there.

Tim: But you know, you’re going to have a hard time hiring a person like that.

Corey: Yeah. You can count these people almost as individuals.

Tim: And that’s a big problem. And you know, in a lot of cases, it’s clearly the case that our profession is talent-starved—I mean, the whole world is talent-starved at the moment, but our profession in particular—and a lot of the decisions about what you can build and what you can do are highly contingent on who you can hire. And you can’t hire a multi-cloud expert, well, you should not deploy, [laugh] you know, a multi-cloud application.

Now, having said that, I just want to dot this i here and say that it can be made to kind of work. I’ve got this one company I advise—I wrote about it in the blog piece—that used to be on AWS and switched over to GCP. I don’t even know why; this happened before I joined them. And they have a lot of applications and then they have some integrations with third-party partners which they implemented with AWS Lambda functions. So, when they moved over to GCP, they didn’t stop doing that.

So, this mission-critical latency-sensitive application of theirs runs on GCP that calls out to AWS to make calls into their partners’ APIs and so on. And works fine. Solid as a rock, reliable, low latency. And so I talked to a person I know who knows over on the AWS side, and they said, “Oh, yeah sure, you know, we talked to those guys. Lots of people do that. We make sure, you know, the connections are low latency and solid.” So, technically speaking, it can be done. But for a variety of business reasons—maybe the most important one being expertise and who you can hire—it’s probably just not a good idea.

Corey: One of the areas where I think is an exception case is if you are a SaaS provider. Let’s pick a big easy example: Snowflake, where they are a data warehouse. They’ve got to run their data warehousing application in all of the major clouds because that is where their customers are. And it turns out that if you’re going to send a few petabytes into a data warehouse, you really don’t want to be paying cloud egress rates to do it because it turns out, you can just bootstrap a second company for that much money.

Tim: Well, Zoom would be another example, obviously.

Corey: Oh, yeah. Anything that’s heavy on data transfer is going to be a strange one. And there’s being close to customers; gaming companies are another good example on this where a lot of the game servers themselves will be spread across a bunch of different providers, just purely based on latency metrics around what is close to certain customer clusters.

Tim: I can’t disagree with that. You know, I wonder how large a segment that is, of people who are, I think you’re talking about core technology companies. Now, of the potential customers of the cloud providers, how many of them are core technology companies, like the kind we’re talking about, who have such a need, and how many people who just are people who just want to run their manufacturing and product design and stuff. And for those, buying into a particular cloud is probably a perfectly sensible choice.

Corey: I’ve also seen regulatory stories about this. I haven’t been able to track them down specifically, but there is a pervasive belief that one interpretation of UK banking regulations stipulates that you have to be able to get back up and running within 30 days on a different cloud provider entirely. And also, they have the regulatory requirement that I believe the data remain in-country. So, that’s a little odd. And honestly, when it comes to best practices and how you should architect things, I’m going to take a distinct backseat to legal requirements imposed upon you by your regulator. But let’s be clear here, I’m not advising people to go and tell their auditors that they’re wrong on these things.

Tim: I had not heard that story, but you know, it sounds plausible. So, I wonder if that is actually in effect, which is to say, could a huge British banking company, in fact do that? Could they in fact, decamp from Azure and move over to GCP or AWS in 30 days? Boy.

Corey: That is what one bank I spoke to over there was insistent on. A second bank I spoke to in that same jurisdiction had never heard of such a thing, so I feel like a lot of this is subject to auditor interpretation. Again, I am not an expert in this space. I do not pretend to be—I know I’m that rarest of all breeds: A white guy with a microphone in tech who admits he doesn’t know something. But here we are.

Tim: Yeah, I mean, I imagine it could be plausible if you didn’t use any higher-level services, and you just, you know, rented instances and were careful about which version of Linux you ran and we’re just running a bunch of Java code, which actually, you know, describes the workload of a lot of financial institutions. So, it should be a matter of getting… all the right instances configured and the JVM configured and launched. I mean, there are no… architecturally terrifying barriers to doing that. Of course, to do that, it would mean you would have to avoid using any of the higher-level services that are particular to any cloud provider and basically just treat them as people you rent boxes from, which is probably not a good choice for other business reasons.

Corey: Which can also include things as seemingly low-level is load balancers, just based upon different provisioning modes, failure modes, and the rest. You’re probably going to have a more consistent experience running HAProxy or nginx yourself to do it. But Tim, I have it on good authority that this is the old way of thinking, and that Kubernetes solves all of it. And through the power of containers and powers combining and whatnot, that frees us from being beholden to any given provider and our workloads are now all free as birds.

Tim: Well, I will go as far as saying that if you are in the position of trying to be portable, probably using containers is a smart thing to do because that’s a more tractable level of abstraction that does give you some insulation from, you know, which version of Linux you’re running and things like that. The proposition that configuring and running Kubernetes is easier than configuring and running [laugh] JVM on Linux [laugh] is unsupported by any evidence I’ve seen. So, I’m dubious of the proposition that operating at the Kubernetes-level at the [unintelligible 00:14:42] level, you know, there’s good reasons why some people want to do that, but I’m dubious of the proposition that really makes you more portable in an essential way.

Corey: Well, you’re also not the target market for Kubernetes. You have worked at multiple cloud providers and I feel like the real advantage of Kubernetes is people who happen to want to protect that they do so they can act as a sort of a cosplay of being their own cloud provider by running all the intricacies of Kubernetes. I’m halfway kidding, but there is an uncomfortable element of truth to that to some of the conversations I’ve had with some of its more, shall we say, fanatical adherents.

Tim: Well, I think you and I are neither of us huge fans of Kubernetes, but my reasons are maybe a little different. Kubernetes does some really useful things. It really, really does. It allows you to take n VMs, and pack m different applications onto them in a way that takes reasonably good advantage of the processing power they have. And it allows you to have different things running in one place with different IP addresses.

It sounds straightforward, but that turns out to be really helpful in a lot of ways. So, I’m actually kind of sympathetic with what Kubernetes is trying to be. My big gripe with it is that I think that good technology should make easy things easy and difficult things possible, and I think Kubernetes fails the first test there. I think the complexity that it involves is out of balance with the benefits you get. There’s a lot of really, really smart people who disagree with me, so this is not a hill I’m going to die on.

Corey: This is very much one of those areas where reasonable people can disagree. I find the complexity to be overwhelming; it has to collapse. At this point, it’s finding someone who can competently run Kubernetes in production is a bit hard to do and they tend to be extremely expensive. You aren’t going to find a team of those people at every company that wants to do things like this, and they’re certainly not going to be able to find it in their budget in many cases. So, it’s a challenging thing to do.

Tim: Well, that’s true. And another thing is that once you step onto the Kubernetes slope, you start looking about Istio and Envoy and [fabric 00:16:48] technology. And we’re talking about extreme complexity squared at that point. But you know, here’s the thing is, back in 2018 I think it was, in his keynote, Werner said that the big goal is that all the code you ever write should be application logic that delivers business value, which you know rep—

Corey: Didn’t CGI say the same thing? Didn’t—like, isn’t there, like, a long history dating back longer than I believe either of us have been alive have, “With this, all you’re going to write is business logic.” That was the Java promise. That was the Google App Engine promise. Again, and again, we’ve had that carrot dangled in front of us, and it feels like the reality with Lambda is, the only code you will write is not necessarily business logic, it’s getting the thing to speak to the other service you’re trying to get it to talk to because a lot of these integrations are super finicky. At least back when I started learning how this stuff worked, they were.

Tim: People understand where the pain points are and are indeed working on them. But I think we can agree that if you believe in that as a goal—which I still do; I mean, we may not have got there, but it’s still a worthwhile goal to work on. We can agree that wrangling Istio configurations is not such a thing; it’s not [laugh] directly value-adding business logic. To the extent that you can do that, I think serverless provides a plausible way forward. Now, you can be all cynical about, “Well, I still have trouble making my Lambda to talk to my other thing.” But you know, I’ve done that, and I’ve also deployed JVM on bare metal kind of thing.

You know what? I’d rather do things at the Lambda level. I really rather would. Because capacity forecasting is a horribly difficult thing, we’re all terrible at it, and the penalties for being wrong are really bad. If you under-specify your capacity, your customers have a lousy experience, and if you over-specify it, and you have an architecture that makes you configure for peak load, you’re going to spend bucket-loads of money that you don’t need to.

Corey: “But you’re then putting your availability in the cloud providers’ hands.” “Yeah, you already were. Now, we’re just being explicit about acknowledging that.”

Tim: Yeah. Yeah, absolutely. And that’s highly relevant to the current discussion because if you use the higher-level serverless function if you decide, okay, I’m going to go with Lambda and Dynamo and EventBridge and that kind of thing, well, that’s not portable at all. I mean, APIs are totally idiosyncratic for AWS and GCP’s equivalent, and Azure’s—what do they call it? Permanent functions or something-a-rather functions. So yeah, that’s part of the trade-off you have to think about. If you’re going to do that, you’re definitely not going to be multi-cloud in that application.

Corey: And in many cases, one of the stated goals for going multi-cloud is that you can avoid the downtime of a single provider. People love to point at the big AWS outages or, “See? They were down for half a day.” And there is a societal question of what happens when everyone is down for half a day at the same time, but in most cases, what I’m seeing, your instead of getting rid of a single point of failure, introducing a second one. If either one of them is down your applications down, so you’ve doubled your outage surface area.

On the rare occasions where you’re able to map your dependencies appropriately, great. Are your third-party critical providers all doing the same? If you’re an e-commerce site and Stripe processes your payments, well, they’re public about being all-in on AWS. So, if you can’t process payments, does it really matter that your website stays up? It becomes an interesting question. And those are the ones that you know about, let alone the third, fourth-order dependencies that are almost impossible to map unless everyone is as diligent as you are. It’s a heavy, heavy lift.

Tim: I’m going to push back a little bit. Now, for example, this company I’m advising that running GCP and calling out to Lambda is in that position; either GCP or Lambda goes off the air. On the other hand, if you’ve got somebody like Zoom, they’re probably running parallel full stacks on the different cloud providers. And if you’re doing that, then you can at least plausibly claim that you’re in a good place because if Dynamo has an outage—and everything relies on Dynamo—then you shift your load over to GCP or Oracle [laugh] and you’re still on the air.

Corey: Yeah, but what is up as well because Zoom loves to sign me out on my desktop whenever I log into it on my laptop, and vice versa, and I wonder if that authentication and login system is also replicated full-stack to everywhere it goes, and what the fencing on that looks like, and how the communication between all those things works? I wouldn’t doubt that it’s possible that they’ve solved for this, but I also wonder how thoroughly they’ve really tested all of the, too. Not because I question them any; just because this stuff is super intricate as you start tracing it down into the nitty-gritty levels of the madness that consumes all these abstractions.

Tim: Well, right, that’s a conventional wisdom that is really wise and true, which is that if you have software that is alleged to do something like allow you to get going on another cloud, unless you’ve tested it within the last three weeks, it’s not going to work when you need it.

Corey: Oh, it’s like a DR exercise: The next commit you make breaks it. Once you have the thing working again, it sits around as a binder, and it’s a best guess. And let’s be serious, a lot of these DR exercises presume that you’re able to, for example, change DNS records on the fly, or be able to get a virtual machine provisioned in less than 45 minutes—because when there’s an actual outage, surprise, everyone’s trying to do the same things—there’s a lot of stuff in there that gets really wonky at weird levels.

Tim: A related similar exercise, which is people who want to be on AWS but want to be multi-region. It’s actually, you know, a fairly similar kind of problem. If I need to be able to fail out of us-east-1—well, God help you, because if you need to everybody else needs
to as well—but you know, would that work?

Corey: Before you go multi-cloud go multi-region first. Tell me how easy it is because then you have full-feature parity—presumably—between everything; it should just be a walk in the park. Send me a postcard once you get that set up and I’ll eat a bunch of words. And it turns out, basically, no one does.

Tim: Mm-hm.

Corey: Another area of lock-in around a lot of this stuff, and I think that makes it very hard to go multi-cloud is the security model of how does that interface with various aspects. In many cases, I’m seeing people doing full-on network overlays. They don’t have to worry about the different security group models and VPCs and all the rest. They can just treat everything as a node sitting on the internet, and the only thing it talks to is an overlay network. Which is terrible, but that seems to be one of the only ways people are able to build things that span multiple providers with any degree of success.

Tim: Well, that is painful because, much as we all like to scoff and so on, in the degree of complexity you get into there, it is the case that your typical public cloud provider can do security better than you can. They just can. It’s a fact of life. And if you’re using a public cloud provider and not taking advantage of their security offerings, infrastructure, that’s probably dumb. But if you really want to be multi-cloud, you kind of have to, as you said.

In particular, this gets back to the problem of expertise because it’s hard enough to hire somebody who really understands IAM deeply and how to get that working properly, try and find somebody who can understand that level of thing on two different cloud providers at once. Oh, gosh.

Corey: This episode is sponsored in part by LaunchDarkly. Take a look at what it takes to get your code into production. I’m going to just guess that it’s awful because it’s always awful. No one loves their deployment process. What if launching new features didn’t require you to do a full-on code and possibly infrastructure deploy? What if you could test on a small subset of users and then roll it back immediately if results aren’t what you expect? LaunchDarkly does exactly this. To learn more, visit launchdarkly.com
and tell them Corey sent you, and watch for the wince.

Corey: Another point you made in your blog post was the idea of lock-in, of people being worried that going all-in on a provider was setting them up to be, I think Oracle is the term that was tossed around where once you’re dependent on a provider, what’s to stop them from cranking the pricing knobs until you squeal?

Tim: Nothing. And I think that is a perfectly sane thing to worry about. Now, in the short term, based on my personal experience working with, you know, AWS leadership, I think that it’s probably not a big short-term risk. AWS is clearly aware that most of the growth is still in front of them. You know, the amount of all of it that’s on the cloud is still pretty small and so the thing to worry about right now is growth.

And they are really, really genuinely, sincerely focused on customer success and will bend over backwards to deal with the customers problems as they are. And I’ve seen places where people have negotiated a huge multi-year enterprise agreement based on Reserved Instances or something like that, and then realize, oh, wait, we need to switch our whole technology stack, but you’ve got us by the RIs and AWS will say, “No, no, it’s okay. We’ll tear that up and rewrite it and get you where you need to go.” So, in the short term, between now and 2025, would I worry about my cloud provider doing that?
Probably not so much.

But let’s go a little further out. Let’s say it’s, you know, 2030 or something like that, and at that point, you know, Andy Jassy decided to be a full-time sports mogul, and Satya Narayana has gone off to be a recreational sailboat owner or something like that, and private equity operators come in and take very significant stakes in the public cloud providers, and get a lot of their guys on the board, and you have a very different dynamic. And you have something that starts to feel like Oracle where their priority isn’t, you know, optimizing for growth and customer success; their priority is optimizing for a quarterly bottom line, and—

Corey: Revenue extraction becomes the goal.

Tim: That’s absolutely right. And this is not a hypothetical scenario; it’s happened. Most large companies do not control the amount of money they spend per year to have desktop software that works. They pay whatever Microsoft’s going to say they pay because they don’t have a choice. And a lot of companies are in the same situation with their database.

They don’t get to budget, their database budget. Oracle comes in and says, “Here’s what you’re going to pay,” and that’s what you pay. You really don’t want to be in a situation with your cloud, and that’s why I think it’s perfectly reasonable for somebody who is doing cloud transition at a major financial or manufacturing or service provider company to have an eye to this. You know, let’s not completely ignore the lock-in issue.

Corey: There is a significant scale with enterprise deals and contracts. There is almost always a contractual provision that says if you’re going to raise a price with any cloud provider, there’s a fixed period of time of notice you must give before it happens. I feel like the first mover there winds up getting soaked because everyone is going to panic and migrate in other directions. I mean, Google tried it with Google Maps for their API, and not quite Google Cloud, but also scared the bejesus out of a whole bunch of people who were, “Wait. Is this a harbinger of things to come?”

Tim: Well, not in the short term, I don’t think. And I think you know, Google Maps [is absurdly 00:26:36] underpriced. That’s hellishly expensive service. And it’s supposed to pay for itself by, you know, advertising on maps. I don’t know about that.

I would see that as the exception rather than the rule. I think that it’s reasonable to expect cloud prices, nominally at least, to go on decreasing for at least the short term, maybe even the medium term. But that’s—can’t go on forever.

Corey: It also feels to me, like having looked at an awful lot of AWS environments that if there were to be some sort of regulatory action or some really weird outage for a year that meant that AWS could not onboard a single new customer, their revenue year-over-year would continue to increase purely by organic growth because there is no forcing function that turns the thing off when you’re done using it. In fact, they can migrate things around to hardware that works, they can continue building you for the things sitting there idle. And there is no governance path on that. So, on some level, winding up doing a price increase is going to cause a massive company focus on fixing a lot of that. It feels on some level like it is drawing attention to a thing that they don’t really want to draw attention to from a purely revenue extraction story.

When CentOS back-walked their ten-year support line two years, suddenly—and with an idea that it would drive [unintelligible 00:27:56] adoption. Well, suddenly, a lot of people looked at their environment, saw they had old [unintelligible 00:28:00] they weren’t using. And massively short-sighted, massively irritated a whole bunch of people who needed that in the short term, but by the renewal, we’re going to be on to Ubuntu or something else. It feels like it’s going to backfire massively, and I’d like to imagine the strategist of whoever takes the reins of these companies is going to be smarter than that. But here we are.

Tim: Here we are. And you know it’s interesting you should mention regulatory action. At the moment, there are only three credible public cloud providers. It’s not obvious the Google’s really in it for the long haul, as last time I checked, they were claiming to maybe be breaking even on it. That’s not a good number, you know? You’d like there to be more than that.

And if it goes on like that, eventually, some politician is going to say, “Oh, maybe they should be regulated like public utilities,” because they kind of are right? And I would think that anybody who did get into Oracle-izing would be—you know, accelerate that happening. Having said that, we do live in the atmosphere of 21st-century capitalism, and growth is the God that must be worshiped at all costs. Who knows. It’s a cloudy future. Hard to see.

Corey: It really is. I also want to be clear, on some level, that with Google’s current position, if they weren’t taking a small loss at least, on these things, I would worry. Like, wait, you’re trying to catch AWS and you don’t have anything better to invest that money into than just well time to start taking profits from it. So, I can see both sides of that one.

Tim: Right. And as I keep saying, I’ve already said once during this slot, you know, the total cloud spend in the world is probably on the order of one or two-hundred billion per annum, and global IT is in multiple trillions. So, [laugh] there’s a lot more space for growth. Years and years worth of it.

Corey: Yeah. The challenge, too, is that people are worried about this long-term strategic point of view. So, one thing you talked about in your blog post is the idea of using hosted open-source solutions. Like, instead of using Kinesis, you’d wind up using Kafka or instead of using DynamoDB you use their managed Cassandra service—or as I think of it Amazon Basics Cassandra—and effectively going down the path of letting them manage this thing, but you then have a theoretical Exodus path. Where do you land on that?

Tim: I think that speaks to a lot of people’s concerns, and I’ve had conversations with really smart people about that who like that idea. Now, to be realistic, it doesn’t make migration easy because you’ve still got all the CI and CD and monitoring and management and scaling and alarms and alerts and paging and et cetera, et cetera, et cetera, wrapped around it. So, it’s not as though you could just pick up your managed Kafka off AWS and drop a huge installation onto GCP easily. But at least, you know, your data plan APIs are the same, so a lot of your code would probably still run okay. So, it’s a plausible path forward. And when people say, “I want to do that,” well, it does mean that you can’t go all serverless. But it’s not a totally insane path forward.

Corey: So, one last point in your blog post that I think a lot of people think about only after they get bitten by it is the idea of data gravity. I alluded earlier in our conversation to data egress charges, but my experience has been that where your data lives is effectively where the rest of your cloud usage tends to aggregate. How do you see it?

Tim: Well, it’s a real issue, but I think it might perhaps be a little overblown. People throw the term petabytes around, and people don’t realize how big a petabyte is. A petabyte is just an insanely huge amount of data, and the notion of transmitting one over the internet is terrifying. And there are lots of enterprises that have multiple petabytes around, and so they think, “Well, you know, it would take me 26 years to transmit that, so I can’t.”

And they might be wrong. The internet’s getting faster all time. Did you notice? I’ve been able to move some—for purely personal projects—insane amounts of data, and it gets there a lot faster than you did. Secondly, in the case of AWS Snowmobile, we have an existence proof that you can do exabyte-ish scale data transfers in the time it takes to drive a truck across the country.

Corey: Inbound only. Snowmobiles are not—at least according to public examples—are valid for Exodus.

Tim: But you know, this is kind of place where regulatory action might come into play if what the people were doing was seen to be abusive. I mean, there’s an existence proof you can do this thing. But here’s another point. So, I suppose you have, like, 15 petabytes—that’s an insane amount of data—displayed in your corporate application. So, are you actually using that to run the application, or is a huge proportion of that stuff just logs and data gathered of various kinds that’s being used in analytics applications and AI models and so on?

Do you actually need all that data to actually run your app? And could you in fact, just pick up the stuff you need for your app, move it to a different cloud provider from there and leave your analytics on the first one? Not a totally insane idea.

Corey: It’s not a terrible idea at all. It comes down to the idea as well of when you’re trying to run a query against a bunch of that data, do you need all the data to transit or just the results of that query, as well? It’s a question of, can you move the compute closer to the data as opposed to the data to where the compute lives?

Tim: Well, you know and a lot of those people who have those huge data pools have it sitting on S3, and a lot of it migrated off into Glacier, so it’s not as if you could get at it in milliseconds anyhow. I just ask myself, “How much data can anybody actually use in a day? In the course of satisfying some transaction requests from a customer?” And I think
it’s not petabyte. It just isn’t.

Now, there are—okay, there are exceptions. There’s the intelligence community, there’s the oil drilling community, there are some communities who genuinely will use insanely huge seas of data on a routine basis, but you know, I think that’s kind of a corner case, so before you shake your head and say, “Ah, they’ll never move because the data gravity,” you know… you need to prove that to me and I might be a little bit skeptical.

Corey: And I think that is probably a very fair request. Just tell me what it is you’re going to be doing here to validate the idea that is in your head because the most interesting lies I’ve found customers tell isn’t intentionally to me or anyone else; it’s to themselves. The narrative of what they think they’re doing from the early days takes root, and never mind the fact that, yeah, it turns out that now that you’ve scaled out, maybe development isn’t 80% of your cloud bill anymore. You learn things and your understanding of what you’re doing has to evolve with the evolution of the applications.

Tim: Yep. It’s a fun time to be around. I mean, it’s so great; right at the moment lock-in just isn’t that big an issue. And let’s be clear—I’m sure you’ll agree with me on this, Corey—is if you’re a startup and you’re trying to grow and scale and prove you’ve got a viable business, and show that you have exponential growth and so on, don’t think about lock-in; just don’t go near it. Pick a cloud provider, pick whichever cloud provider your CTO already knows how to use, and just go all-in on them, and use all their most advanced features and be serverless if you can. It’s the only sane way forward. You’re short of time, you’re short of money, you need growth.

Corey: “Well, what if you need to move strategically in five years?” You should be so lucky. Great. Deal with it then. Or, “Well, what if we want to sell to retail as our primary market and they hate AWS?”

Well, go all-in on a provider; probably not that one. Pick a different provider and go all in. I do not care which cloud any given company picks. Go with what’s right for you, but then go all in because until you have a compelling reason to do otherwise, you’re going to spend more time solving global problems locally.

Tim: That’s right. And we’ve never actually said this probably because it’s something that both you and I know at the core of our being, but it probably needs to be said that being multi-cloud is expensive, right? Because the nouns and verbs that describe what clouds do are different in Google-land and AWS-land; they’re just different. And it’s hard to think about those things. And you lose the capability of using the advanced serverless stuff. There are a whole bunch of costs to being multi-cloud.

Now, maybe if you’re existentially afraid of lock-in, you don’t care. But for I think most normal people, ugh, it’s expensive.

Corey: Pay now or pay later, you will pay. Wouldn’t you ideally like to see that dollar go as far as possible? I’m right there with you because it’s not just the actual infrastructure costs that’s expensive, it costs something far more dear and expensive, and that is the cognitive expense of having to think about both of these things, not just how each cloud provider works, but how each one breaks. You’ve done this stuff longer than I have; I don’t think that either of us trust a system that we don’t understand the failure cases for and how it’s going to degrade. It's, “Oh, right. You built something new and awesome. Awesome. How does it fall over? What direction is it going to hit, so what side should I not stand on?” It’s based on an understanding of what you’re about to blow holes in.

Tim: That’s right. And you know, I think particularly if you’re using AWS heavily, you know that there are some things that you might as well bet your business on because, you know, if they’re down, so is the rest of the world, and who cares? And, other things, eh, maybe a little chance here. So, understanding failure modes, understanding your stuff, you know, the cost of sharp edges, understanding manageability issues. It’s not obvious.

Corey: It’s really not. Tim, I want to thank you for taking the time to go through this, frankly, excellent post with me. If people want to learn more about how you see things, and I guess how you view the world, where’s the best place to find you?

Tim: I’m on Twitter, just @timbray T-I-M-B-R-A-Y. And my blog is at tbray.org, and that’s where that piece you were just talking about is, and that’s kind of my online presence.

Corey: And we will, of course, put links to it in the [show notes 00:37:42]. Thanks so much for being so generous with your time. It’s always a pleasure to talk to you.

Tim: Well, it’s always fun to talk to somebody who has shared passions, and we clearly do.

Corey: Indeed. Tim Bray principal at Textuality Services. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry comment that you then need to take to all of the other podcast platforms out there purely for redundancy, so you don’t get locked into one of them.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Peter

Peter's spent more than a decade building scalable and robust systems at startups across adtech and edtech. At Remind, where he's VP of Technology, Peter pushes for building a sustainable tech company with mature software engineering. He lives in Southern California and enjoys spending time at the beach with his family.

Links:

  • Redis: https://redis.com/
  • Remind: https://www.remind.com/
  • Remind Engineering Blog: https://engineering.remind.com
  • LinkedIn: https://www.linkedin.com/in/hamiltop
  • Email: peterh@remind101.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Today’s episode is brought to you in part by our friends at MinIO the high-performance Kubernetes native object store that’s built for the multi-cloud, creating a consistent data storage layer for your public cloud instances, your private cloud instances, and even your edge instances, depending upon what the heck you’re defining those as, which depends probably on where you work. It’s getting that unified is one of the greatest challenges facing developers and architects today. It requires S3 compatibility, enterprise-grade security and resiliency, the speed to run any workload, and the footprint to run anywhere, and that’s exactly what MinIO offers. With superb read speeds in excess of 360 gigs and 100 megabyte binary that doesn’t eat all the data you’ve gotten on the system, it’s exactly what you’ve been looking for. Check it out today at min.io/download, and see for yourself. That’s min.io/download, and be sure to tell them that I sent you.

Corey: This episode is sponsored in part by our friends at Vultr. Spelled V-U-L-T-R because they’re all about helping save money, including on things like, you know, vowels. So, what they do is they are a cloud provider that provides surprisingly high performance cloud compute at a price that—while sure they claim its better than AWS pricing—and when they say that they mean it is less money. Sure, I don’t dispute that but what I find interesting is that it’s predictable. They tell you in advance on a monthly basis what it’s going to going to cost. They have a bunch of advanced networking features. They have nineteen global locations and scale things elastically. Not to be confused with openly, because apparently elastic and open can mean the same thing sometimes. They have had over a million users. Deployments take less that sixty seconds across twelve pre-selected operating systems. Or, if you’re one of those nutters like me, you can bring your own ISO and install basically any operating system you want. Starting with pricing as low as $2.50 a month for Vultr cloud compute they have plans for developers and businesses of all sizes, except maybe Amazon, who stubbornly insists on having something to scale all on their own. Try Vultr today for free by visiting: vultr.com/screaming, and you’ll receive a $100 in credit. Thats V-U-L-T-R.com slash screaming.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn and this is a fun episode. It is a promoted episode, which means that our friends at Redis have gone ahead and sponsored this entire episode. I asked them, “Great, who are you going to send me from, generally, your executive suite?” And they said, “Nah. You already know what we’re going to say. We want you to talk to one of our customers.” And so here we are. My guest today is Peter Hamilton, VP of Technology at Remind. Peter, thank you for joining me.

Peter: Thanks, Corey. Excited to be here.

Corey: It’s always interesting when I get to talk to people on promoted guest episodes when they’re a customer of the sponsor because to be clear, you do not work for Redis. This is one of those stories you enjoy telling, but you don’t personally have a stake in whether people love Redis, hate Redis, adopt that or not, which is exactly what I try and do on these shows. There’s an authenticity to people who have in-the-trenches experience who aren’t themselves trying to sell the thing because that is their entire job in this world.

Peter: Yeah. You just presented three or four different opinions and I guarantee we felt all at the different times.

Corey: [laugh]. So, let’s start at the very beginning. What does Remind do?

Peter: So, Remind is a messaging tool for education, largely K through 12. We support about 30 million active users across the country, over 2 million teachers, making sure that every student has, you know, equal opportunities to succeed and that we can facilitate as much learning as possible.

Corey: When you say messaging that could mean a bunch of different things to a bunch of different people. Once on a lark, I wound up sitting down—this was years ago, so I’m sure the number is a woeful underestimate now—of how many AWS services I could use to send a message from me to you. And this is without going into the lunacy territory of, “Well, I can tag a thing and then mail it to you like a Snowball Edge or something.” No, this is using them as intended, I think I got 15 or 16 of them. When you say messaging, what does that mean to you?

Peter: So, for us, it’s about communication to the end-user. We will do everything we can to deliver whatever message a teacher or district administrator has to the user. We go through SMS, text messaging, we go through Apple and Google’s push services, we go through email, we go through voice call, really pulling out all the stops we can to make sure that these important messages get out.

Corey: And I can only imagine some of the regulatory pressure you almost certainly experience. It feels like it’s not quite to HIPAA levels, where ohh, there’s a private cause of action if any of this stuff gets out, but people are inherently sensitive about communications involving their children. I always sort of knew this in a general sense, and then I had kids myself, and oh, yeah, suddenly I really care about those sorts of things.

Peter: Yeah. One of the big challenges, you can build great systems that do the correct thing, but at the end of the day, we’re relying on a teacher choosing the right recipient when they send a message. And so we’ve had to build a lot of processes and controls in place, so that we can, kind of, satisfy two conflicting needs: One is to provide a clear audit log because that’s an important thing for districts to know if something does happen, that we have clear communication; and the other is to also be able to jump in and intervene when something inappropriate or mistaken is sent out to the wrong people.

Corey: Remind has always been one of those companies that has a somewhat exalted reputation in the AWS space. You folks have been early adopters of a bunch of different services—which let’s be clear, in the responsible way, not the, “Well, they said it on stage; time to go ahead and put everything they just listed into production because we for some Godforsaken reason, view it as a todo list.”—but you’ve been thoughtful about how you approach things, and you have been around as a company for a while. But you’ve also been making a significant push toward being cloud-native by certain definitions of that term. So, I know this sounds like a college entrance essay, but what does cloud-native mean to you?

Peter: So, one of the big gaps—if you take an application that was written to be deployed in a traditional data center environment and just drop it in the cloud, what you’re going to get is a flaky data center.

Corey: Well, that’s unfair. It’s also going to be extremely expensive.

Peter: [laugh]. Sorry, an expensive, flaky data set.

Corey: There we go. There we go.

Peter: What we’ve really looked at–and a lot of this goes back to our history in the earlier days; we ran a top of Heroku and it was kind of the early days what they call the Twelve-Factor Application—but making aggressive decisions about how you structure your architecture and application so that you fit in with some of the cloud tools that are available and that you fit in, you know, with the operating models that are out there.

Corey: When you say an aggressive decision, what sort of thing are you talking about? Because when I think of being aggressive with an approach to things like AWS, it usually involves Twitter, and I’m guessing that is not the direction you intend that to go.

Peter: No, I think if you look at Twitter or Netflix or some of these players that, quite frankly, have defined what AWS is to us today through their usage patterns, not quite that.

Corey: Oh, I mean using Twitter to yell at them explicitly about things—

Peter: Oh.

Corey: —because I don’t do passive-aggressive; I just do aggressive.

Peter: Got it. No, I think in our case, it’s been plotting a very narrow path that allows us to avoid some of the bigger pitfalls. We have our sponsor here, Redis. Talk a little bit about our usage of Redis and how that’s helped us in some of these cases. One of the pitfalls you’ll find with pulling a non-cloud-native application and put it in the cloud is state is hard to manage.

If you put state on all your machines and machines go down, networks fail, all those things, you now no longer have access to that state and we start to see a lot of problems. One of the decisions we’ve made is try to put as much data as we can into data stores like Redis or Postgres or something, in order to decouple our hardware from the state we’re trying to manage and provide for users so that we’re more resilient to those sorts of failures.

Corey: I get the sense from the way that we’re having this conversation, when you talk about Redis, you mean actual Redis itself, not ElastiCache for Redis, or as to I’m tending to increasingly think about AWS’s services, Amazon Basics for Redis.

Peter: Yeah. I mean, Amazon has launched a number of products. They have their ElastiCache, they have their new MemoryDB, there’s a lot different ways to use this. We’ve relied pretty heavily on Redis, previously known as Redis Labs, and their enterprise product in their cloud, in order to take care of our most important data—which we just don’t want to manage ourselves—trying to manage that on our own using something like ElastiCache, there’s so many pitfalls, so many ways that we can lose that data. This data is important to us. By having it in a trusted place and managed by a great ops team, like they have at Redis, we’re able to then lean in on the other aspects of cloud data to really get as much value as we can out of AWS.

Corey: I am curious. As I said you’ve had a reputation as a company for a while in the AWS space of doing an awful lot of really interesting things. I mean, you have a robust GitHub presence, you have a whole bunch of tools that have come out Remind that are great, I’ve linked to a number of them over the years in the newsletter. You are clearly not afraid, culturally, to get your hands dirty and build things yourself, but you are using Redis Enterprise as opposed to open-source Redis. What drove that decision? I have to assume it’s not, “Wait. You mean, I can get it for free as an open-source project? Why didn’t someone tell me?” What brought you to that decision?

Peter: Yeah, a big part of this is what we could call operating leverage. Building a great set of tools that allow you to get more value out of AWS is a little different story than babysitting servers all day and making sure they stay up. So, if you look through, most of our contributions in open-source space have really been around here’s how to expand upon these foundational pieces from AWS; here’s how to more efficiently launch a suite of servers into an auto-scaling group; here’s, you know, our troposphere and other pieces there. This was all before Amazon CDK product, but really, it was, here’s how we can more effectively use CloudFormation to capture our Infrastructure as Code. And so we are not afraid in any way to invest in our tooling and invest in some of those things, but when we look at the trade-off of directly managing stateful services and dealing with all the uncertainty that comes, we feel our time is better spent working on our product and delivering value to our users and relying on partners like Redis in order to provide that stability we need.

Corey: You raise a good point. An awful lot of the tools that you’ve put out there are the best, from my perspective, approach to working with AWS services. And that is a relatively thin layer built on top of them with an eye toward making the user experience more polished, but not being so heavily opinionated that as soon as the service goes in a different direction, the tool becomes completely useless. You just decide to make it a bit easier to wind up working with specific environment variables or profiles, rather than what appears to be the AWS UX approach of, “Oh, now type in your access key, your secret key and your session token, and we’ve disabled copy and paste. Go, have fun.” You’ve really done a lot of quality of life improvements, more so than you have this is the entire system of how we do deploys, start to finish. It’s opinionated and sort of a, like, a take on what Netflix, did once upon a time, with Asgard. It really feels like it’s just the right level of abstraction.

Peter: We did a pretty good job. I will say, you know, years later, we felt that we got it wrong a couple times. It’s been really interesting to see that, that there are times when we say, “Oh, we could take these three or four services and wrap it up into this new concept of an application.” And over time, we just have to start poking holes in that new layer and we start to see we would have been better served by sticking with as thin a layer as possible that enables us, rather than trying to get these higher-level pieces.

Corey: It’s remarkably refreshing to hear you say that just because so many people love to tell the story on podcasts, or on conference stages, or whatever format they have of, “This is what we built.” And it is an aspirationally superficial story about this. They don’t talk about that, “Well, firstly, without these three wrong paths first.” It’s always a, “Oh, yes, obviously, we are smart people and we only make the correct decision.”

And I remember in the before times sitting in conference talks, watching people talk about great things they’d done, and I’ll turn next to the person next to me and say, “Wow, I wish I could be involved in a project like that.” And they’ll say, “Yeah, so do I.” And it turns out they work at the company the speaker is from. Because all of these things tend to be the most positive story. Do you have an example of something that you have done in your production environment that going back, “Yeah, in hindsight, I would have done that completely differently.”

Peter: Yeah. So, coming from Heroku moving into AWS, we had a great open-source project called Empire, which kind of bridge that gap between them, but used Amazon’s ECS in order to launch applications. It was actually command-line compatible with the Heroku command when it first launched. So, a very big commitment there. And at the time—I mean, this comes back to the point I think you and I were talking
about earlier, where architecture, costs, infrastructure, they’re all interlinked.

And I’m a big fan of Conway’s Law, which says that an organization’s structure needs to match its architecture. And so six, seven years ago, we’re heavy growth-based company and we are interns running around, doing all the things, and we wanted to have really strict guardrails and a narrow set of things that our development team could do. And so we built a pretty constrained: You will launch, you will have one Docker image per ECS service, it can only do these specific things. And this allowed our development team to focus on pretty buttons on the screen and user engagement and experiments and whatnot, but as we’ve evolved as a company, as we built out a more robust business, we’ve started to track revenue and costs of goods sold more aggressively, we’ve seen, there’s a lot of inefficient things that come out of it.

One particular example was we used PgBouncer for our connection pooling to our Postgres application. In the traditional model, we had an auto-scaling group for a PgBouncer, and then our auto-scaling groups for the other applications would connect to it. And we saw additional latency, we saw additional cost, and we eventually kind of twirl that down and packaged that PgBouncer alongside the applications that needed it. And this was a configuration that wasn’t available on our first pass; it was something we intentionally did not provide to our development team, and we had to unwind that. And when we did, we saw better performance, we saw better cost efficiency, all sorts of benefits that we care a lot about now that we didn’t care about as much, many years ago.

Corey: It sounds like you’re describing some semblance of an internal platform, where instead of letting all your engineers effectively, “Well, here’s the console. Ideally, you use some form of Infrastructure as Code. Good luck. Have fun.” You effectively gate access to that. Is that something that you’re still doing or have you taken a different approach?

Peter: So, our primary gate is our Infrastructure as Code repository. If you want to make a meaningful change, you open up a PR, got to go through code review, you need people to sign off on it. Anything that’s not there may not exist tomorrow. There’s no guarantees. And we’ve gone around, occasionally just shut random servers down that people spun up in our account.

And sometimes people will be grumpy about it, but you really need to enforce that culture that we have to go through the correct channels and we have to have this cohesive platform, as you said, to support our development efforts.

Corey: So, you’re a messaging service in education. So, whenever I do a little bit of digging into backstories of companies and what has made, I guess, an impression, you look for certain things and explicit dates are one of them, where on March 13th of 2020, your business changed just a smidgen. What happened other than the obvious, we never went outside for two years?

Peter: [laugh]. So, if we roll back a week—you know, that’s March 13th, so if we roll back a week, we’re looking at March 6th. On that day, we sent out about 60 million messages over all of our different mediums: Text, email, push notifications. On March 13th that was 100 million, and then, a few weeks later on March 30th, that was 177 million. And so our traffic effectively tripled over the course of those three weeks. And yeah, that’s quite a ride, let me tell you.

Corey: The opinion that a lot of folks have who have not gotten to play in sophisticated distributed systems is, “Well, what’s the hard part there you have an auto-scaling group. Just spin up three times the number of servers in that fleet and problem solved. What’s challenging?” A lot, but what did you find that the pressure points were?

Peter: So, I love that example, that your auto-scaling group will just work. By default, Amazon’s auto-scaling groups only support 1000 backends. So, when your auto-scaling group goes from 400 backends to 1200, things break, [laugh] and not in ways that you would have expected. You start to learn things about how database systems provided by Amazon have limits other than CPU and memory. And they’re clearly laid out that there’s network bandwidth limits and things you have to worry about.

We had a pretty small team at that time and we’d gotten this cadence where every Monday morning, we would wake up at 4 a.m. Pacific because as part of the pandemic, our traffic shifted, so our East Coast users would be most active in the morning rather than the afternoon. And so at about 7 a.m. on the east coast is when everyone came online. And we had our Monday morning crew there and just looking to see where the next pain point was going to be.

And we’d have Monday, walk through it all, Monday afternoon, we’d meet together, we come up with our three or four hypotheses on what will break, if our traffic doubles again, and we’d spend the rest of that next week addressing those the best we could and repeat for the next Monday. And we did this for three, four or five weeks in a row, and finally, it stabilized. But yeah, it’s all the small little things, the things you don’t know about, the limits in places you don’t recognize that just catch up to you. And you need to have a team that can move fast and adapt quickly.

Corey: You’ve been using Redis for six, seven years, something along those lines, as an enterprise offering. You’ve been working with the same vendor who provides this managed service for a while now. What are the fruits of that relationship? What is the value that you see by continuing to have a long-term relationship with vendors? Because let’s be serious, most of us don’t stay in jobs that long, let alone work with the same vendor.

Peter: Yeah. So, coming back to the March 2020 story, many of our vendors started to see some issues here that various services weren’t scaled properly. We made a lot of phone calls to a lot of vendors in working with them, and I… very impressed with how Redis Labs at the time was able to respond. We hopped on a call, they said, “Here’s what we think we need to do, we’ll go ahead and do this. We’ll sort this out in a few weeks and figure out what this means for your contract. We’re here to help and support in this pandemic because we recognize how this is affecting everyone around the world.”

And so I think when you get in those deeper relationships, those long-term relationships, it is so helpful to have that trust, to have a little bit of that give when you need it in times of crisis, and that they’re there and willing to jump in right away.

Corey: There’s a lot to be said for having those working relationships before you need them. So often, I think that a lot of engineering teams just don’t talk to their vendors to a point where they may as well be strangers. But you’ll see this most notably because—at least I feel it most acutely—with AWS service teams. They’ll do a whole kickoff when the enterprise support deal is signed, three years go passed, and both the AWS team and the customer’s team have completely rotated since then, and they may as well be strangers. Being able to have that relationship to fall back on in those really weird really, honestly, high-stress moments has been one of those things where I didn’t see the value myself until the first time I went through a hairy situation where I found that that was useful.

And now it’s oh, I—I now bias instead for, “Oh, I can fit to the free tier of this service. No, no, I’m going to pay and become a paying customer.” I’d rather be a customer that can have that relationship and pick up the phone than someone whining at people in a forum somewhere of, “Hey, I’m a free user, and I’m having some problems with production.” Just never felt right to me.

Peter: Yeah, there’s nothing worse than calling your account rep and being told, “Oh, I’m not your account rep anymore.” Somehow you missed the email, you missed who it was. Prior to Covid, you know—and we saw this many, many years ago—one of the things about Remind is every back-to-school season, our traffic 10Xes in about three weeks. And so we’re used to emergencies happening and unforeseen things happening. And we plan through our year and try to do capacity planning and everything, but we been around the block a couple of times.

And so we have a pretty strong culture now leaning in hard with our support reps. We have them in our Slack channels. Our AWS team, we meet with often. Redis Labs, we have them on Slack as well. We’re constantly talking about databases that may or may not be performing as we expect them, too. They’re an extension of our team, we have an incident; we get paged. If it’s related to one of the services, we hit them in Slack immediately and have them start checking on the back end while we’re checking on our side. So.

Corey: One of the biggest takeaways I wish more companies would have is that when you are dependent upon another company to effectively run your production infrastructure, they are no longer your vendor, they’re your partner, whether you want them to be or not. And approaching it with that perspective really pays dividends down the road.

Peter: Yeah. One of the cases you get when you’ve been at a company for a long time and been in relationship for a long time is growing together is always an interesting approach. And seeing, sometimes there’s some painful points; sometimes you’re on an old legacy version of their product that you were literally the last customer on, and you got to work with them to move off of. But you were there six years ago when they’re just starting out, and they’ve seen how you grow, and you’ve seen how they’ve grown, and you’ve kind of been able to marry that experience together in a meaningful way.

Corey: This episode is sponsored by our friends at Oracle Cloud. Counting the pennies, but still dreaming of deploying apps instead of “Hello, World” demos? Allow me to introduce you to Oracle’s Always Free tier. It provides over 20 free services and infrastructure, networking, databases, observability, management, and security. And—let me be clear here—it’s actually free. There’s no surprise billing until you intentionally and proactively upgrade your account. This means you can provision a virtual machine instance or spin up an autonomous database that manages itself, all while gaining the networking, load balancing, and storage resources that somehow never quite make it into most free tiers needed to support the application that you want to build. With Always Free, you can do things like run small-scale applications or do proof-of-concept testing without spending a dime. You know that I always like to put asterisks next to the word free? This is actually free, no asterisk. Start now. Visit snark.cloud/oci-free that’s snark.cloud/oci-free.

Corey: Redis is, these days, of data platform back once upon a time, I viewed it as more of a caching layer. And I admit that the capabilities of the platform has significantly advanced since those days when I viewed it purely through lens of cache. But one of the interesting parts is that neither one of those use cases, in my mind, blends particularly well with heavy use of Spot Fleets, but you’re doing exactly that. What are your folks doing over there?

Peter: [laugh]. Yeah, so as I mentioned earlier, coming back to some of the Twelve-Factor App design, we heavily rely on Redis as sort of a distributed heap. One of our challenges of delivering all these messages is every single message has its in-flight state: Here’s the content, here’s who we sent it to, we wait for them to respond. On a traditional application, you might have one big server that stores it all in-memory, and you get the incoming requests, and you match things up. By moving all that state to Redis, all of our workers, all of our application
servers, we know they can disappear at any point in time.

We use Amazon’s Spot Instances and their Spot Fleet for all of our production traffic. Every single web service, every single worker that we have runs on this infrastructure, and we would not be able to do that if we didn’t have a reliable and robust place to store this data that is in-flight and currently being accessed. So, we’ll have a couple hundred gigs of data at any point in time in a Redis Database, just representing in-flight work that’s happening on various machines.

Corey: It’s really neat seeing Spot Fleets being used as something more than a theoretical possibility. It’s something I’ve always been very interested in, obviously, given the potential cost savings; they approach cheap is free in some cases. But it turns out—we talked earlier about the idea of being cloud-native versus the rickety, expensive data center in the cloud, and an awful lot of applications are simply not built in a way that yeah, we’re just going to randomly turn off a subset of your systems, ideally, with two minutes of notice, but all right, have fun with that. And a lot of times, it just becomes a complete non-starter, even for stateless workloads, just based upon how all of these things are configured. It is really interesting to watch a company that has an awful lot of responsibility that you’ve been entrusted with who embraces that mindset. It’s a lot more rare than you’d think.

Peter: Yeah. And again, you know, sometimes, we overbuild things, and sometimes we go down paths that may have been a little excessive, but it really comes down to your architecture. You know, it’s not just having everything running on Spot. It’s making effective use of SQS and other queueing products at Amazon to provide checkpointing abilities, and so you know that should you lose an instance, you’re only going to lose a few seconds of productive work on that particular workload and be able to kick off where you left off.

It’s properly using auto-scaling groups. From the financial side, there’s all sorts of weird quirks you’ll see. You know, the Spot market has a wonderful set of dynamics where the big instances are much, much cheaper per CPU than the small ones are on the Spot market. And so structuring things in a way that you can colocate different workloads onto the same hosts and hedge against the host going down by spreading across multiple availability zones. I think there’s definitely a point where having enough workload, having enough scale allows you to take advantage of these things, but it all comes down to the architecture and design that really enables it.

Corey: So, you’ve been using Redis for longer than I think many of our listeners have been in tech.

Peter: [laugh].

Corey: And the key distinguishing points for me between someone who is an advocate for a technology and someone who’s a zealot—or a pure critic—is they can identify use cases for which is great and use cases for which it is not likely to be a great experience. In your time with Redis, what have you found that it’s been great at and what are some areas that you would encourage people to consider more carefully before diving into it?

Peter: So, we like to joke that five, six years ago, most of our development process was, “I’ve hit a problem. Can I use Redis to solve that problem?” And so we’ve tried every solution possible with Redis. We’ve done all the things. We have number of very complicated Lua scripts that are managing different keys in an atomic way.

Some of these have been more successful than others, for sure. Right now, our biggest philosophy is, if it is data we need quickly, and it is data that is important to us, we put it in Enterprise Redis, the cloud product from Redis. Other use cases, there’s a dozen things that you can use for a cache, Redis is great for cache, memcache does a decent job as well; you’re not going to see a meaningful difference between those sorts of products. Where we’ve struggled a little bit has been when we have essentially relational data that we need fast access to. And we’re still trying to find a clear path forward here because you can do it and you can have atomic updates and you can kind of simulate some of the ACID characteristics you would have in a relational database, but it adds a lot of complexity.

And that’s a lot of overhead to our team as we’re continuing to develop these products, to extend them, to fix any bugs you might have in there. And so we’re kind of recalibrating a bit, and some of those workloads are moving to other data stores where they’re more appropriate. But at the end of the day, it’s data that we need fast, and it’s data that’s important, we’re sticking with what we got here because it’s been working pretty well.

Corey: It sounds almost like you started off with the mindset of one database for a bunch of different use cases and you’re starting to differentiate into purpose-built databases for certain things. Or is that not entirely accurate?

Peter: There’s a little bit of that. And I think coming back to some of our tooling, as we kind of jumped on a bit of the microservice bandwagon, we would see, here’s a small service that only has a small amount of data that needs to be stored. It wouldn’t make sense to bring up a RDS instance, or an Aurora instance, for that, you know, in Postgres. Let’s just store it in an easystore like Redis. And some of those cases have been great, some of them have been a little problematic.

And so as we’ve invested in our tooling to make all our databases accessible and make it less of a weird trade-off between what the product needs, what we can do right now, and what we want to do long-term, and reduce that friction, we’ve been able to be much more deliberate about the data source that we choose in each case.

Corey: It’s very clear that you’re speaking with a voice of experience on this where this is not something that you just woke up and figured out. One last area I want to go into with you is when I asked you what is you care about primarily as an engineering leader and as you look at serving your customers well, you effectively had a dual answer, almost off the cuff, of stability and security. I find the two of those things are deeply intertwined in most of the conversations I have, but they’re rarely called out explicitly in quite the way that you do. Talk to me about that.

Peter: Yeah, so in our wild journey, stability has always been a challenge. And we’ve alway—you know, been an early startup mode, where you’re constantly pushing what can we ship? How quickly can we ship it? And in our particular space, we feel that this communication that we foster between teachers and students and their parents is incredibly important, and is a thing that we take very, very seriously. And so, a couple years ago, we were trying to create this balance and create not just a language that we could talk about on a podcasts like this, but really recognizing that framing these concepts to our company internally: To our engineers to help them to think as they’re building a feature, what are the things they should think about, what are the concerns beyond the product spec; to work with our marketing and sales team to help them to understand why we’re making these investments that may not get particular feature out by X date but it’s still a worthwhile investment.

So, from the security side, we’ve really focused on building out robust practices and robust controls that don’t necessarily lock us into a particular standard, like PCI compliance or things like that, but really focusing on the maturity of our company and, you know, our culture as we go forward. And so we’re in a place now we are ISO 27001; we’re heading into our third year. We leaned in hard our disaster recovery processes, we’ve leaned in hard on our bug bounties, pen tests, kind of, found this incremental approach that, you know, day one, I remember we turned on our bug bounty and it was a scary day as the reports kept coming in. But we take on one thing at a time and continue to build on it and make it an essential part of how we build systems.

Corey: It really has to be built in. It feels like security is not something could be slapped on as an afterthought, however much companies try to do that. Especially, again, as we started this episode with, you’re dealing with communication with people’s kids. That is something that people have remarkably little sense of humor around. And rightfully so.

Seeing that there is as much if not more care taken around security than there is stability is generally the sign of a well-run organization. If
there’s a security lapse, I expect certain vendors to rip the power out of their data centers rather than run in an insecure fashion. And your job done correctly—which clearly you have gotten to—means that you never have to make that decision because you’ve approached this the right way from the beginning. Nothing’s perfect, but there’s always the idea of actually caring about it being the first step.

Peter: Yeah. And the other side of that was talking about stability, and again, it’s avoiding the either/or situation. We can work in as well along those two—stability and security—we work in our cost of goods sold and our operating leverage in other aspects of our business. And every single one of them, it’s our co-number one priorities are stability and security. And if it costs us a bit more money, if it takes our dev team a little longer, there’s not a choice at that point. We’re doing the correct thing.

Corey: Saving money is almost never the primary objective of any company that you really want to be dealing with unless something bizarre is going on.

Peter: Yeah. Our philosophy on, you know, any cost reduction has been this should have zero negative impact on our stability. If we do not feel we can safely do this, we won’t. And coming back to the Spot Instance piece, that was a journey for us. And you know, we tested the waters a bit and we got to a point, we worked very closely with Amazon’s team, and we came to that conclusion that we can safely do this. And we’ve been doing it for over a year and seen no adverse effects.

Corey: Yeah. And a lot of shops I’ve talked to folks about well, when we go and do a consulting project, it’s, “Okay. There’s a lot of things that could have been done before we got here. Why hasn’t any of that been addressed?” And the answer is, “Well. We tried to save money once and it caused an outage and then we weren’t allowed to save money anymore. And here we are.” And I absolutely get that perspective. It’s a hard balance to strike. It always is.

Peter: Yeah. The other aspect where stability and security kind of intertwine is you can think about security as InfoSec in our systems and locking things down, but at the end of the day, why are we doing all that? It’s for the benefit of our users. And Remind, as a communication platform, and safety and security of our users is as dependent on us being up and available so that teachers can reach out to parents with important communication. And things like attendance, things like natural disasters, or lockdowns, or any of the number of difficult situations schools find themselves in. This is part of why we take that stewardship that we have so seriously is that being up and protecting a user’s data just has such a huge impact on education in this country.

Corey: It’s always interesting to talk to folks who insists they’re making the world a better place. And it’s, “What do you do?” “We’re improving ad relevance.” I mean, “Okay, great, good for you.” You’re serving a need that I would I would not shy away from classifying what you do, fundamentally, as critical infrastructure, and that is always a good conversation to have. It’s nice being able to talk to folks who are doing things that you can unequivocally look at and say, “This is a good thing.”

Peter: Yeah. And around 80% of public schools in the US are using Remind in some capacity. And so we’re not a product that’s used in a few civic regions. All across the board. One of my favorite things about working in Remind is meeting people and telling them where I work, and they recognize it.

They say, “Oh, I have that app, I use that app. I love it.” And I spent years and ads before this, and you know, I’ve been there and no one ever told me they were glad to see an ad. That’s never the case. And it’s been quite a rewarding experience coming in every day, and as you said, being part of this critical infrastructure. That’s a special thing.

Corey: I look forward to installing the app myself as my eldest prepares to enter public school in the fall. So, now at least I’ll have a hotline of exactly where to complain when I didn’t get the attendance message because, you know, there’s no customer quite like a whiny customer.

Peter: They’re still customers. [laugh]. Happy to have them.

Corey: True. We tend to be. I want to thank you for taking so much time out of your day to speak with me. If people want to learn more about what you’re up to, where’s the best place to find you?

Peter: So, from an engineering perspective at Remind, we have our blog, engineering.remind.com. If you want to reach out to me directly. I’m on LinkedIn; good place to find me or you can just reach out over email directly, peterh@remind101.com.

Corey: And we will put all of that into the show notes. Thank you so much for your time. I appreciate it.

Peter: Thanks, Corey.

Corey: Peter Hamilton, VP of Technology at Remind. This has been a promoted episode brought to us by our friends at Redis, and I’m Cloud Economist Corey Quinn. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an angry and insulting comment that you will then hope that Remind sends out to 20 million students all at once.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Philip

Philip Griffiths is VP Global Business Development and regularly speaks at events from DevOps to IoT to Cyber Security. Prior to this, he worked for Atos IT Services in various roles working with C-suit executives to realise their digital transformation. He lives in Cambridge with his wife and two daughters.

Links:

  • NetFoundry: https://netfoundry.io/
  • Blog article: https://netfoundry.io/demystifying-the-magic-of-zero-trust-with-my-daughter-and-opensource/
  • netfoundry.io/screaminginthecloud: https://netfoundry.io/screaminginthecloud

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Today’s episode is brought to you in part by our friends at MinIO the high-performance Kubernetes native object store that’s built for the multi-cloud, creating a consistent data storage layer for your public cloud instances, your private cloud instances, and even your edge instances, depending upon what the heck you’re defining those as, which depends probably on where you work. It’s getting that unified is one of the greatest challenges facing developers and architects today. It requires S3 compatibility, enterprise-grade security and resiliency, the speed to run any workload, and the footprint to run anywhere, and that’s exactly what MinIO offers. With superb read speeds in excess of 360 gigs and 100 megabyte binary that doesn’t eat all the data you’ve gotten on the system, it’s exactly what you’ve been looking for. Check it out today at min.io/download, and see for yourself. That’s min.io/download, and be sure to tell them that I sent you.

Corey: This episode is sponsored by our friends at Oracle Cloud. Counting the pennies, but still dreaming of deploying apps instead of “Hello, World” demos? Allow me to introduce you to Oracle’s Always Free tier. It provides over 20 free services and infrastructure, networking, databases, observability, management, and security. And—let me be clear here—it’s actually free. There’s no surprise billing until you intentionally and proactively upgrade your account. This means you can provision a virtual machine instance or spin up an autonomous database that manages itself, all while gaining the networking, load balancing, and storage resources that somehow never quite make it into most free tiers needed to support the application that you want to build. With Always Free, you can do things like run small-scale applications or do proof-of-concept testing without spending a dime. You know that I always like to put asterisks next to the word free? This is actually free, no asterisk. Start now. Visit snark.cloud/oci-free that’s snark.cloud/oci-free.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Today’s promoted episode is about a topic that is near and dear to my heart. In the AWS universe, we have seen over time that the networking has gotten more and more capable going from EC2 Classic to the world of VPC network to a whole bunch of other things. But with that capability comes a stupendous amount of complexity, to the point where the easy answer to, “Do you understand how networking works within AWS?” Is, of course, no, “I don’t.”

I’m joined today by Philip Griffiths, who’s the Head of Business Development at NetFoundry. Philip, thank you for joining me.

Philip: Pleasure to be here, Corey.

Corey: So, NetFoundry has what I would argue to be one of the most intriguing-slash-differentiated approaches to handling that ever-increasing complexity around the networking story, not just in AWS, but a number of different cloud providers, and between them, and that approach is to ignore it completely. Have I nailed the salient approach here with that, I guess we’ll call it a flippant statement.

Philip: Yeah, I’d probably say so. It’s the interesting thing where a lot of people say cloud networking is hard, and from our perspective, it should just be super easy, you should be able to provision it in a few minutes with only outbound ports, and set up your policy so that malicious actors can’t get inside it. It should be that easy, and programmable, and it’s a shame that the current world is not.

Corey: One of the hard problems has always been in, I guess, security, which is the thing that everyone pretends to care about right up front, but in practice, often winds up bolting it on after the fact because, “We care about security,” is sort of the trademark phrase of things that we see, usually an email announcing a data breach when it was very clear that companies did not care about security. It’s not just me complaining about how complex the network stack is, but by what directly flows from that. If you aren’t able to fit all of that into your head as far as what’s going on from a security perspective, the odds of misconfiguration creep in and you don’t really become aware of what your risk exposure is. I’m really partial to the idea of just avoiding it entirely. Is NetFoundry, effectively, a network overlay? Is it something that goes a bit beyond that? Effectively, where do you folks start and where do you stop?

Philip: Yes, that is precisely correct. We are a network overlay that’s been built on the principles of zero trust. What is very unique is the ability to be able to start it wherever you want. So yes, you can deploy it from the AWS Marketplace in a few minutes into your VPC or into your operating system, but we also have the ability to actually put it directly into the application stack itself, which has some very interesting complications. What I find as the most interesting starting point is the oxymoron of secure networking.

There are no secure networks. It’s not possible. Networks are designed to share information and taking it to first principles, you can only isolate networks. And this is why we had the thought process for if we’re going to put our overlay network into stuff and make it secure, we have to start at the application level because then we can actually just isolate it to an application communicating into an application, which has profound implications.

Corey: The network part is relatively straightforward. I imagine it just becomes, more or less, what resembles a fairly flat network where everything internal is allowed to talk to each other, and then, in turn, this winds up effectively elevating what should be allowed to talk to what and on what ports and whatnot into something that’s a lot closer to the application logic, and transcends whatever provider it happens to be traversing.

Philip: Yeah, correct. Following the principles of zero trust, we utilize strong embedded identity as a function of what the endpoints are, what the source and destination is. And therefore you build up your policies and services to say what should communicate to what on the basis that the default the least privileged: Absolutely nothing. Your underlay then, the only thing you need is commodity internet with outbound ports. The whole concept of north-south, east-west, if you’re app-embedded, you don’t even need public DNS; you don’t even need DNS at all. Naming conventions go out the window; you don’t need to conform to the standards. You know, you could say, “I want to hit Jenkins.” You go to Jenkins because that can be done.

Corey: I would approach this entire endeavor with a fair bit of suspicion and no small amount of alarm if it were something that you had developed internally, as far as, “Well, we’re just going to replace what amounts to your entire network stack and just go ahead and trust us. It’s fine.” But you didn’t do that. You’re riding on top of the OpenZiti open-source project. And that basically assuages a whole raft of concerns I would have if something like this were proprietary, and people who know what they’re doing—who, let’s be clear, aren’t me—were not able to inspect it and say, “Okay, this passes muster”—as they have done—or alternately, “No, this is terrifyingly dangerous for a variety of excellent reasons.”

And it really feels like a lot of the zero-trust stories that we see these days that are taking advantage of either a network overlay approach or shifting authentication into a different layer, have all taken a somewhat similar tack. I used to think it was a good idea; now I’m starting to suspect it might very well be the only viable model. Do you find that that’s accurate, or was this a subject of some contention when you were starting out?

Philip: So, there’s two very interesting [sigh] thoughts that came to me as you were saying that. The number one is yes, we drove forward with OpenZiti because we’ve seen open-source just completely dominate the industry and everything new that’s been built. If you want to deploy an application, you’re building on Linux. And in fact, you’re probably [laugh] also running on Kubernetes if you’re building new. And our objective was to be able to turn OpenZiti into you know, the open-source, zero-trust private network and equivalent where it’s just standard: You’ll bake your application with Ziti, by design.

It will become a check function that people say you have to comply to. When I look at other vendors and how they look at zero-trust, I broadly see a few things that dishearten me. And again, it’s a big market, a lot of people—everyone says they’re zero-trust nowadays—but I broadly categorize it into a few ways. You have people who are effectively acting as a proxy and they’re adding authentication as a way to check what people should have access to. And they may give access to the whole network, they may do granular; it varies between them. In fact, I’ve just written a blog on this where I effectively call that no-magic zero trust. It’s a blog conceptualized within Harry Potter and [unintelligible 00:07:36] a conversation with my daughter.

Corey: Yeah, any way to tell a story that beats the traditional enterprise voice is very much appreciated over in this corner of the world.

Philip: [laugh]. Yeah, exactly. You have a second tier, which is what I like to think as semi-magical. And that’s where you start saying, I am going to use a software-defined perimeter. So, that it’s first packet authenticate, or outbound-only based upon embedded identity. And in my eyes, this is basically an invisibility cloak.

You then have app-embedded or magical zero-trust. And this is where you’re putting the invisibility cloak inside your application, but you’re also giving it a port key so that when it needs to connect to something else on the other side of the world, it just happens; it’s transparent. And broadly speaking, I think it’s very good that the whole world, including the US government is taking zero-trust incredibly importantly, but the distribution of how people tackle a problem is wildly different. There are some zero-trust solutions, which going in the right direction, but fundamentally, if you’re putting it in front of your—
I won’t name a vendor, but there was a vendor who in December, they released a report
that said in 90 seconds, common vulnerabilities are exploited something like 96% of the time. 24-hours, 100%.

A few days later, they had a 9.8 CVE on their zero-trust VPN concentrator with a public IP, to which I thought, “If you’re not patching that immediately, you’ve got problems if someone is coming into your network.”

Corey: Absolutely. We just completed our annual security awareness training here, and so much of it just… it really made my skin crawl, there was an entire module on how to effectively detect phishing emails, and I got to tell you, if they ever start running spellcheck on their some of their [spear-phishing 00:09:23] campaigns, then we’re all doomed because that was what the entire training was here. My position is, is that okay, if someone in your company clicks a bad link and it destroys the company’s infrastructure, maybe it’s the person who’s clicking the link that is not necessarily the critical failure point here. Great, if someone compromises an employee workstation, there should be a way to contain the blast radius, they should not now be inside the walls and able to traverse into whatever it is that they want. There should be additional barriers, and zero trust—though it has become, as you say, a catch-all term—seems to be a serious way of looking at this type of compromise and this sort of mitigation against that sort of behavior.

Philip: Definitely. And I think that leads itself to, if you’re using the correct zero-trust solution, you’re able to close [unintelligible 00:10:12] ports, great, you’ve now massively reduced your attack surface. But what if someone does get a phishing injection of ransomware or something to their endpoint or into their servers? The two things that I like to think about is that if you’re creating your overlay network so that the only communication from your server is outbound into the public IPs of your private overlay, then effectively even if the ransomware gets in there, it can’t then connect to its command and control module to then go through the kill cycle to other activities. The other is that if you then look at it [instead 00:10:46] of on the server-side, but actually on the client-side, if someone infects my Mac laptop with ransomware, we use this internal application called Mattermost.

And it’s basically Slack, but open-source. If my Mattermost is Ziti-fied, even I’ve got ransomware on my device, it can’t side-channel attack into Mattermost because you would actually have to break into the Mattermost application and somehow get that Mattermost application to make a compromised query or whatever to get past the system. So really, when I look at zero-trust, it’s not about saying, “We’re secure. Job done. You know, fire the security department because we don’t need them anymore.” It’s all about saying—

Corey: Box check. Hand it off to the auditor.

Philip: [laugh]. Exactly. It’s more about saying the cost of attack, the cost of compromised is increased, ideally, to the point where the malicious actors don’t have a return on investment. Because if they don’t have a return on investment, they will find something else that’s not your applications and your systems to try and compromise.

Corey: I want to make sure that I’m contextualizing this properly because we’re talking—I think—about what almost looks like two different worlds here. There’s the, this is how things wind up working in the ecosystem as far as your server environment goes in a cloud provider, but then we’re also talking about what goes on in your corporate network of people who are using laptops, which is increasingly being done from home these days. Where do you folks start? Where do you stop? Do you transcend into the corporate network as well, or is this primarily viewed as a production utility?

Philip: We do. One of our original design principles with OpenZiti was for it to be a platform rather than a point solution. So, we designed it from the ground up to be able to support any IP packets, TCP, UDP, et cetera, whether you’re doing, client-server, server-server, machine-server, server-initiated, client-initiated, yadda, yadda, yadda. So effectively, the same technology can be applied to many different use cases, depending on where you want to use it. We’ve been doing work recently to handle, let’s call them the hard use cases.

Probably one of the hardest ones out there is VoIP. There is a playbook that is currently taking place where the VoIP-managed service provider gets DDoSed by malicious actors; the playbook is to move it onto a CDN so that you move the attack surface and you get respite for a few hours. And there’s not really any way to solve it because blocking DDoS attacks at layer 3, layer 4 is incredibly difficult unless you can make your PBX dark. And I’ve seen a couple of our OpenZitiE engineers making calls from one device to another without going through the PBX by doing that over OpenZiti, and being able to solve some of the challenges that’s normally associated with VoIP. Again, it was really one of our design principles: How can we make the platform is so flexible that we can do X, Y, Zed today; we’re able to build it, again to become a standard, because it can handle anything.

Corey: One of the big questions that people are going to have going into this is, and this may sound surprising is a little bit less about technical risk of things like encryption and the rest and a lot more around the idea of okay, does this mean that what you are building becomes a central point of business risk? In other words, if the NetFoundry SaaS installation and wherever they happen to be using as their primary winds up going down, does that mean suddenly nothing can talk to one another? Because it turns out that, you know, computers are not particularly useful in 2022 if they aren’t able to talk to other computers, by and large. “The network is the computer,” as was famously stated. What is the failure mode in the event that you experience technical interruption?

Philip: We have this internal sessions, which we call Ziti Kitchens, where our engineering team that are creating Ziti educate on stuff that they’re building. And one of them in the Ziti Kitchen was around HA, HS, et cetera, and all of the functions that we’ve built in so that you have redundancy and availability within the different components. Because effectively it’s an overlay network, so we’ve designed it to be a mesh overlay network. You can setup with one point of failure, but then simultaneously, you can very easily set up to have no points of failure because it can have that redundancy and the overlay has its own mechanisms to do things like smart routing and calculation of underlying costs.

That cost in that instance would be, well, AWS has gone down, so the latency to send a packet or flow over it is incredibly high, therefore I’m going to avoid that route and send the traffic to another location. I always remember this Ziti Kitchen episode because the underlying technology that does it is called Terminators—Ziti has these things called Terminators—some of the slide there was this little heads over the Terminator with the red eyes, you know, the silver exoskeleton, which always made me laugh.

Corey: It’s helpful to have things that fail out of band as opposed to—think of the traditional history in security before everything was branded with zero-trust as a prerequisite for exhibiting at RSA; before that was firewalls was the story, and the question always was, if a firewall fails, do you want it to fail open or fail closed? And believe it or not, there are legitimate answers in both directions; depends on context and what you’re doing. There are some things for example, IAM in a cloud world where you absolutely never want to fail open, full stop. You would rather someone bodily rip the power cable out the back of the data center rather than let that happen. With something like this, where nothing is able to talk to one another if the entire system goes down, yeah, you want to have the control system that you folks run to be out of band, that is almost always the right answer.

As I look at the various case studies that you have on your website and the serious companies that are using what you have built, do you find that they are primarily centralizing around individual cloud providers? Are you seeing that they’re using this as a expression of multi-cloud because I can definitely see a story where oh, it helps bring two cloud providers from a networking and security perspective onto the same page, but I can also see, even within one cloud provider, the idea that, hey, I don’t have to play around with your ridiculous nonsense? What use cases are you seeing emerge among
your customers?

Philip: Definitely, the multi-cloud challenge is one that we’re seeing as a emerging trend. We do a lot of work with Oracle and, you know, their stated position is multi-cloud is a fact. In fact for them, if we make the secure networking easier, we can bring workloads into our cloud quicker [unintelligible 00:17:21] the main driver between our partnership. We recently did a blog talking about Superclouds and the advent of organizations like Snowflake and HashiCorp and Confluence and Databricks basically building value and business applications which abstracts away the underlying complexity. But you get into the problem of the standard shared security model, where the customer has to deal with DNS and VPNs and MPLS and AWS Private Endpoint or Azure Private Link or whatever they call it, and you have to assemble this Frankenstein of stuff just to enable a VM to communicate to another VM.

And the posit of our blog—in fact, we use that exact quote—John Gage—“The computer is the network.” If you can put a network inside the application, you’ve now given your supercloud superpowers because [unintelligible 00:18:13] natively—I mean, this is very
marketing term, but, “Develop once; deploy anywhere,” and be multi-cloud-native.

Corey: The idea of being able to adapt to emerging usage patterns without full-on redeploy is handy. What I also would like to highlight, too, is that you are, of course, a network overlay and that is something that is fairly well understood and people have seen it, but your preferred adoption model goes up a couple of steps beyond that into altering the way that the application thinks about these things. And you offer an SDK that ranges from single line of code implementation to I think up to 20, so it’s not a massive rewrite of the application, but it does require modification of the stack. What does that buy you, for lack of a better term? Because once you have the application becomes aware of what is effectively its own, “Special network,” quote-unquote, its work to wind up modifying existing applications around something like this. What’s the payoff?

Philip: So, there’s three broad ones that immediately come to my mind. Number one is the highest security that effectively—your private network is inside the app, so you have to somehow break into the app and that can be incredibly complicated, particularly run the app in something like a confidential compute enclave; you can now have a distributed confidential system.

The second is what you’re getting in programmability. You’re able to effectively operate in a fully—even, you know, you get to a GitOps environment. We’re currently working on documentation which says, “Hey, you can do all this stuff in GitOps and then it’ll go into your CI/CD and that’ll talk to the APIs.” And it’ll effectively do everything in a completely programmable manner so that you can treat your private networks as cattle rather than as pets.

The third is transparency. You used the words earlier of bolt-on networking because that’s how we always think about networking security: We bolt it on. As a user, we have to jump through the VPN hoop, we have to go through the bastion, we have to interact with the network. If your private network’s inside the application, then you interact with the application. I can have a mobile application on my device and I have no idea that it’s part of a private network and that the API is private and the malicious actors can’t get to it. I just interact with the application. That is it.

That is what no one else has the ability to do and where OpenZiti has its most power because then you get rid of the constant tug of war between the security team that want to lock everything down and the users and the developers who want to move fast and give a great experience. You can effectively have your cake and eat it.

Corey: The challenge, of course, with rolling a lot of these things out in a way that becomes highly programmable is that unlocks a bunch of capability, but the double-edged sword there is always one of complexity. I mean, we take a look at the way that AWS networking has progressed, and they finally rolled out the VPC Reachability Analyzer, so when two things can’t talk to each other, well, you run this thing and it tells you exactly why, which is super handy. And then just as a way of twisting the knife a little bit, every time you run it, they charge at ten cents for the privilege, which doesn’t actually matter in the context of what anyone is being compensated for, until and unless you build this into something programmatic, but it stings a little bit. And the idea of being able to program these things to abstract away a lot of that complexity is incredibly compelling, except for the part where now it feels like it really increases developer burden on a lot of these things. Have you found that to be true? Do you find that it is sort of like a sliding scale? What has the customer experience been around this?

Philip: I would say a sliding scale. You know, we had one organization who they started with the OpenZiti Tunnelers, and then we convinced them to use the SDK and [unintelligible 00:21:51], “Oh, this was super easy.” And now they just run OpenZiti on themselves. But then they’ve also said at some point, we’ll use the NetFoundry platform, which effectively gives us a SaaS experience in consuming that. One of the huge focus—well, we’ve got a few big focuses for product development, but one of the really big areas is really giving more visibility and monitoring so that rather than people having to react to configuration problems or things which they need to fix in order to ensure your perfect network overlay, instead, those things are being seen and automatically dealt with human-in-the-loop if you want it, in order to remove that burden.

Because ultimately, if you can get the network to a point where as long as you’ve got underlay and you’ve set your policy, the overlay is going to work, it’s going to be secure, and it’s going to give you the uptime you need, that is the Nirvana that we all have to strive for.

Corey: This episode is sponsored in part by our friends at Vultr. Spelled V-U-L-T-R because they’re all about helping save money, including on things like, you know, vowels. So, what they do is they are a cloud provider that provides surprisingly high performance cloud compute at a price that—while sure they claim its better than AWS pricing—and when they say that they mean it is less money. Sure, I don’t dispute that but what I find interesting is that it’s predictable. They tell you in advance on a monthly basis what it’s going to going to cost. They have a bunch of advanced networking features. They have nineteen global locations and scale things elastically. Not to be confused with openly, because apparently elastic and open can mean the same thing sometimes. They have had over a million users. Deployments take less that sixty seconds across twelve pre-selected operating systems. Or, if you’re one of those nutters like me, you can bring your own ISO and install basically any operating system you want. Starting with pricing as low as $2.50 a month for Vultr cloud compute they have plans for developers and businesses of all sizes, except maybe Amazon, who stubbornly insists on having something to scale all on their own. Try Vultr today for free by visiting: vultr.com/screaming, and you’ll receive a $100 in credit. Thats V-U-L-T-R.com slash screaming.

Corey: A common criticism of things that shall we say abstract away the network is a fairly common predictable failure mode. I’ve been making fun of Kubernetes on this particular point for years, and I’m annoyed that at the time that we’re recording this, that is still accurate. But from the cloud providers’ perspective, when you run Kubernetes, it looks like one big really strangely behaved single-tenant application. And Kubernetes itself is generally not aware of zone affinity, so it could just as easily wind up tossing traffic to the node next to it at zero cost or across an availability zone at two cents per gigabyte, or, God forbid across the internet at nine cents a gigabyte and counting depending upon how it works. And the application-side has absolutely no conception of this.

How does OpenZiti address this in the real world because it’s one of those things where it almost doesn’t matter what you folks charge on top of it, but instead oh wow, this winds up being so hellaciously expensive that we can’t use it regardless of whatever benefit it provides just because it becomes a non-starter.

Philip: So, when we built the overlay and the mesh, we did it from the perspective of making it as programmable and self-driven as possible. So, with the whole Terminator strategies that was mentioned earlier, it gives you the ability to start putting logic into how you want packets to flow. Today, it does it on a calculation of end-to-end latency and chooses and reroutes traffic in order to give that information. But there’s no reason that you couldn’t hook it up into understanding what is the numerical in monetary cost for sending a packet along a certain path. Or even what is my application performance monitoring tool saying? Because what that says versus what the network believes could be different things. And effectively you can ingest that information to make your smart routing decisions so all of that logic can exist within the overlay that operates for you.

Corey: I will say that really harkens back, on some level, to what I was experimenting with back when I got my CCNA many years ago where there’s an idea of routing protocols have built into the idea of the cost of a link. I will freely admit slash confess that at the time of the low-cost link, I assumed this was about what was congested or what would wind up having, theoretically, some transit versus peering agreement. It never occurred to me that I’d have to think about those things in a local network and have to calculate in the Byzantine pricing models of cloud providers. But I’ve seen examples of folks who are using OpenZiti, and NetFoundry alike, to wind up building in these costing models so that yeah, ideally, it just keeps everything local, but of that path degrades then yes, we would prefer to go over an expensive link than to basically have TCP terminate on the floor until everything comes back up. It sort of feels like there’s an awful lot of logic you can bake into that goes well beyond what routing protocols are capable of, just by virtue of exposing that programmability.

Well, for this customer because they’re on the pre—on the extreme tier, then we want to have the expensive fallback; for low-tier customers, we might want to have them just have an outage until things end. And it really comes down to letting business decisions express themselves in terms of application behavior while in degraded state. I love that idea.

Philip: Yeah, I understand. We don’t do it today, but there will be a point in the future—I strongly believe—that we’ll be able to say, hey, I’ll give you an SLA on the internet. Because we’ll have such path diversity and visibility of how the internet operates that we’ll be able to say within certain risk parameters of what we can deliver. But then you
can take it to other logical extremes. You could say, “Hey, I want to build a green overlay.
I want to make sure that I’m using Arm instances and in data centers of renewable energy so that my network is green.”

Or you can say on a GDPR-compliant overlay so that my data stays within a certain country. You start being able to say—you know, really start dreaming up what are the different policies that I can apply to this because you’re applying a central policy to then what is in the distributed system.

Corey: One last topic I want to cover before we call it an episode is that you are, effectively, a SaaS company that is built on top of an open-source project. And that has been an interesting path for a lot of companies that early on, figured that if they wrote the software, a lot of the contributors who are doing the lion’s share of contribution, that they were clearly the best people to run it. And Amazon’s approach towards operational excellence—as they called it—wound up causing some challenges when they launched the Amazon Basics version of that service. I feel like there are some natural defenses built into OpenZiti to keep it from suffering that fate, but I’m very curious to get your take on it.

Philip: Fundamentally, our take is that—in fact, our mission is to take what was previously impossible and turn it into a standard. And the only way you can really create standards is to have a open-source that is adopted by the wider community and that ecosystems get built around and into. And that means giving an OpenZiti to absolutely everyone so that they can use it, they can innovate on top of it. We all know that very few people actually want to host their own infrastructure, so we assume a large percentage of people will come and go, “Hey, NetFounder, you provide us the hosting, you provide us the SaaS capability so we don’t have to do that ourselves.” But fundamentally in the knowledge that there’s something bigger because it’s not just us maintaining this project; there’s a bunch of people who are doing pull requests and find out cool, fun ways to build further value on what we can build ourselves.

We believe the recent history is littered with examples of the new world built on open-source. And fundamentally, we think that’s really the only way to be able to change an industry so profoundly as we intend to.

Corey: I would also argue that, to be very direct—and I can probably get away with saying this in a way that I suspect you might not be able to—but if AWS had it in their character to simplify things and make it a lot easier for people to work with in a networking sense, what’s stopping them? They didn’t need to wait for an open-source company to wind up coming out of nowhere and demonstrating the value of this. Customers have been asking it for years. I think that at this point, this is something that is unlikely to ever wind up being integrated into a cloud provider’s primary offering. Until and unless the entire industry shifts, at which point we’re having a radically different conversation very far down the road.

Philip: Yeah, potentially because it opens the interesting thing that if you make it so easy for someone to take their data out, do they use your cloud less? There are some cloud providers that will lean into that because they do see more clouds in the future and others that won’t. I see it more myself that as those kind of things happen, it’ll be done on a product-by-product basis. For example, we’re talking to an organization, and [unintelligible 00:29:49] like, “Oh, could you Ziti-fy our JDBC driver so that when users access our database, they don’t have to use a VPN?” [unintelligible 00:29:55], “Yeah. We’ve already done that with JDBC. We called it ZDBC.”

So, we’ll just, instead of using the general industry one—probably the Oracle one or something because that’s kind of standard—we’ll take your one that you’ve created for yourself and be able to solve that problem for you.

Corey: I really want to thank you for taking the time to speak with me today. If people want to learn more, where’s the best place to find you?

Philip: Best place to go to is netfoundry.io/screaminginthecloud. From there, anyone can grab some free Ziggy swag. Ziggy’s our little open-source mascot, cute little piece of pasta with many different outfits. Little sass as well. And you can find further information both on OpenZiti and NetFoundry.

Corey: And we will put links to both of those in the [show notes 00:30:40]. Thanks so much for taking the time to speak with me today. I really appreciate it.

Philip: It’s a pleasure. Thanks, Corey.

Corey: Philip Griffiths, Head of Business Development at NetFoundry. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry comment telling me exactly why I’m wrong about AWS’s VPC complexity, and that comment will get moderated and I won’t get to read it until you pay me ten cents to tell you how it got moderated.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Nipun

Nipun Agarwal is a Senior Vice President, MySQL HeatWave and Advanced Development, Oracle. His interests include distributed data processing, machine learning, cloud technologies and security. Nipun was part of the Oracle Database team where he introduced a number of new features. He has been awarded over 170 patents.

Links:

  • Oracle: https://www.oracle.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Vultr. Spelled V-U-L-T-R because they’re all about helping save money, including on things like, you know, vowels. So, what they do is they are a cloud provider that provides surprisingly high performance cloud compute at a price that—while sure they claim its better than AWS pricing—and when they say that they mean it is less money. Sure, I don’t dispute that but what I find interesting is that it’s predictable. They tell you in advance on a monthly basis what it’s going to going to cost. They have a bunch of advanced networking features. They have nineteen global locations and scale things elastically. Not to be confused with openly, because apparently elastic and open can mean the same thing sometimes. They have had over a million users. Deployments take less that sixty seconds across twelve pre-selected operating systems. Or, if you’re one of those nutters like me, you can bring your own ISO and install basically any operating system you want. Starting with pricing as low as $2.50 a month for Vultr cloud compute they have plans for developers and businesses of all sizes, except maybe Amazon, who stubbornly insists on having something to scale all on their own. Try Vultr today for free by visiting: vultr.com/screaming, and you’ll receive a $100 in credit. Thats V-U-L-T-R.com slash screaming.

Corey: Couchbase Capella Database-as-a-Service is flexible, full-featured and fully managed with built in access via key-value, SQL, and full-text search. Flexible JSON documents aligned to your applications and workloads. Build faster with blazing fast in-memory performance and automated replication and scaling while reducing cost. Capella has the best price performance of any fully managed document database. Visit couchbase.com/screaminginthecloud to try Capella today for free and be up and running in three minutes with no credit card required. Couchbase Capella: make your data sing.

Corey: Welcome to Screaming in the Cloud, I’m Corey Quinn. Today’s promoted episode is a returning guest with a slight difference. When last we spoke, Nipun Agarwal was a VP over at Oracle, but now—that’s right. When people stay in a company long enough and perform well, they wind up getting additional adjectives in lieu of other things—Nipun, you’re now a Senior VP over at Oracle. Congratulations, I think, unless that just means you’ve gotten older. Welcome back.

Nipun: Thank you, Corey.

Corey: So, now that you’re at SVP level, I can ask some of the harder questions that we didn’t necessarily—seem fair to get into the last time we spoke, such as what is an Oracle, and what might they do these days? For folks who have, I don’t know, been living in a cave for 40 years.

Nipun: Corey, glad to be back on your show. And since the last time we spoke, we have had, like, you know, a lot of enhancements and innovations, and I’ll be happy to describe those in detail whenever is a good time.

Corey: Absolutely so you’ve been focused on MySQL for a very long time. And you’ve been using it so long, I really should be calling it YourSQL, but that’s neither here nor there. And you’ve also been focusing on HeatWave, which is effectively MySQL with then some—I’m just going to cheat and call it magic that is layered on top of it. That is probably a terrible descriptor of what it actually does, but understand I’m coming from a perspective where I firmly believe the best database in the world is Amazon Route 53, which is a DNS server, so people look at that and say, ‘well, that’s not really what it’s designed to do,’ which really sounds like a ‘them’ problem. And fair enough. We’re going to invert it here. So, why is HeatWave a terrible DNS server? What is it exactly?

Nipun: So, MySQL is the most popular database in the world—it’s the most popular open-source database in the world—lots of people use it. All the major cloud vendors, they take the MySQL database, and either as is or, like, you know, with some enhancements, they offer a managed service, whether it’s Amazon, Azure, Google, pretty much all the major cloud vendors. Now, MySQL has been designed and optimized for transaction processing, so it does a great job for transaction processing. But when customers need to run complex queries or when they need to run analytics, customers would have to take the data out of the MySQL database into some other database for running analytics.

Corey: Let me make sure I understand your terms properly. When you say ‘transactional,’ you’re talking about I’m shopping for underpants on a website. I go ahead and make a purchase; that’s considered a transaction, and a database change reflecting my purchase makes sense. From an analytics perspective, you’re like, “All right, let’s see who bought underpants during this time period.” It’s effectively, usually, a small individual record versus now we’re going to start doing deep dives into effectively a lot of those records in aggregate, is that directionally correct, or is my understanding more than a little flawed about things beyond DNS?

Nipun: Right. What you describe is very accurate. That transaction processing is about point queries making frequent changes, whereas when we talk about analytics, it typically involves scanning a much larger amount of data to get the results, and aggregations is a very good example of that.

Corey: So historically, that seems that people have used very different tooling for different sides of those. Ideally—I admit, back in the bad old days when I was a systems administrator, we were running MySQL a fair bit, and we had the primary database, which was the thing that handled all of the live transactions and the rest, and whenever we ran business reporting queries on it, it’s like, “Huh, why is the website super slow?” And it didn’t seem to work very well. Now, back then, at the scale we were operating at the solution was, “Ah, we’re going to use a replica, and then we’re going to basically beat the crap out of the replica for our reporting queries.” And if that gets a little slow and bogged down, who cares? Well, just other people running reporting queries; people can still buy underpants.

So, that was the way that we handled it back then. This was a decade ago. Data sets have gotten significantly larger since then, and apparently, my way of viewing it is, as they say, quaint when they’re trying not to be actively insulting. The right way to do it these days is to have completely separate systems that wind up handling those queries with different user interfaces by and large. That is, to my understanding, the rise of ‘Big Data,’ and you can hear the initial caps in Big Data with people talk about it like that.

Nipun: Correct. So, what you describe is absolutely correct that people would extract the data out of databases, take it to specialized databases, which are [apt 00:05:11] for running decision-making analytic processing. But the downside is that a people need to express the logic and write code to extract this data, and then customers end up with these two different databases. They got to keep the data in sync, they got to move the data periodically. So, there are a lot of, like, you know, issues in terms of having to manage two different databases, one for transaction processing, one for analytics.

What we have done with HeatWave is to enhance the MySQL database service in the Oracle Cloud so that now the single MySQL database is optimized both for transaction processing as well as analytics. So, now you have a single database. And whether you want to run point queries or these aggregate queries, you can do it on the same data. So, the data remains as is. You’re bringing richness of computation, richness in query processing, to the customers.

Corey: One of the truisms of cloud is that it forces a reevaluation, in many cases, of things that people historically hadn’t had to think about it. A classic example when I was consulting on cloud migrations, was building up costing models, as you might imagine. And my customers would ask me questions, such is, “Great. So, what’s this going to cost us?” And I would come back with, “Well, okay, how many gigabytes in a given month does transfer between this database and that other database, you know, in the machine sitting right next to it?” And their response started off with a, “Why on earth do you think we would know that?” Followed by, “Wait, why do we need to know
that?” Followed by, “Oh, God. It costs us to do what?”

And very quickly an architectural pattern has emerged within cloud of—you know, people experience this the second time, they plan for it. And as a result, whatever database is the most cost-effective is the one that data is already in because moving data from point to point is inherently an expensive proposition. Depending on where the second point is, it can be an extortionately expensive proposition. Which means that very often, we’ll start to see patterns that are, I guess, sacrificing one side of the database interaction model or the other, that transactions are going to be a little slower because you need to have it in the same place you’re going to be running large scale analytics on, or alternately, analytics are going to be super crappy, just because you have to wind up querying systems during downtimes and low periods. It just becomes a giant mess, regardless of whether it’s bad in one way, bad in another, or just expensive, it hasn’t worked for people. And my sense is that that is what HeatWave is directly aimed.

Nipun: Yes. Indeed. So, there are multiple reasons why HeatWave is being so successful. One is the case that okay, customers need a single database, instead of having multiple. The second thing is, there is absolutely no change required to MySQL applications, so the MySQL applications or MySQL compatible applications work as-is with this query [unintelligible 00:08:10] HeatWave without any change.

But the third reason why this is so popular is that HeatWave has been designed from the ground up for scalability, performance, and optimized for the underlying gear, which is the underlying cloud platform. As a result, it offers a very good price-performance compared to any of the service we have run against. So, not only is it providing the benefits of having a single database, no change to the application, but also it is extremely fast and low price. And that’s because a lot of technology innovations we did, like, almost like, over a decade to build this, scale our system for analytic processing, which has been optimized for the underlying cloud [commodity 00:08:55] gear.

Corey: So, help me understand. Is HeatWave a, effectively, reengineering of MySQL? Is it a completely separate layer that exists distinct from an existing MySQL database? Or is it something else entirely?

Nipun: So, we started off designing HeatWave separately as something ground up, which came out of many years of research and advanced developing. And once we knew that we could scale up HeatWave for analytic processing, and it is very well optimized for the underlying hardware and such. Then we did the work of enhancing the MySQL database so that it can be integrated, right? So yes, it started off as a standalone effort from the ground up so that we didn’t have to, you know, [live 00:09:38] any constraints of any existing codebase, so we could design it and optimize it right from the ground up to be the best possible. But then we integrated this thing with the MySQL database so that the customers can use it without requiring any change to the application in terms of the semantics or any new syntax, right? So, there’s absolutely no new syntax and no change to the semantics for existing MySQL applications. So, it gives you best of both worlds.

Corey: So, this has frequently been described in the context of a competitor to very—again, forgive the Amazonian focus; that’s where I spend most of my time, usually complaining about things—but it’s been positioned in some ways as a competitor to things such as RDS or Aurora, as well as Redshift, or Snowflake if we’re stepping slightly outside that ecosystem. The challenge that I keep running into, very often, is that when I talk to customers using those systems—and yes, those systems invariably show up on the bill as one of the big numbers, regardless of how you slice it—it feels like their use case for each of those is very different, it feels very much like half of those are aimed at purely transactional and half of them are aimed at the data warehousing story, the large amounts of data for analytics queries. And my default knee-jerk reaction, whenever someone says, “Ah, we built a thing that does both of those super well,” it’s, “Yeah, I’ve heard this before, it was the HP multifunction printer where it does three things, none of them well.” And no one has a multifunction printer that they liked for the longest time—because it’s moving parts and computers and the devil in equal measure—and it’s okay, so you’re trying to build something that stands between two worlds, but it’s easy to come away with the conclusion, as a result, that it’s not the best of breed for either use case, but rather a series of trade-offs or compromises that are made to enable both use cases. I get the sense that that is not your impression of what you’ve built.

Nipun: Correct. And I’ll give you a data point for that. In the data point is—

Corey: Yay. Data I love that. As opposed to your opinion is bad because my opinion is good. No, no, coming with data is a great approach. Please continue.

Nipun: [laugh]. In terms of the customers who are using or adopting MySQL HeatWave, one of the largest segments of the customers who are migrating their production workloads from other databases or other services and coming to HeatWave are AWS customers who are migrating their production workloads from RDS or Aurora and are going production with MySQL HeatWave. So, the fact that the customers are doing that is an evidence that there is some value to it. And the reasons they are doing it is absolutely no change to their application, it is faster, it is cheaper. Now, in addition, what they find is that many of these customers were moving their data from Aurora or RDS into Redshift or Snowflake for analytics. They don’t need to do that, right, and that’s an additional savings they get.

But we have a lot of evidence that existing customers have MySQL-based services—definitely AWS, but even on other clouds—and Aurora are migrating, and that’s very encouraging for us that, hey, we should be doing something right for customers to want to migrate their workloads to MySQL HeatWave.

Corey: You had a couple of announcements coming out about what’s new and what’s coming to HeatWave, and one of the ones that we’re talking about today is the idea of elasticity. Something you just said reminds me of a couple years ago when Amazon had relatively recently brought out Aurora and they said much the same thing of, “Oh, it’s super-elastic. You don’t have to take it down to make it bigger.” And it’s great. Well, you just talked about people removing data as they migrate somewhere else, and the question I had at the time was, “Okay, great. So, that’s how the database embiggens. That’s great. How does it emsmallen? Does that wind up having that same elastic property?”

And the response was very defensive, “Well, why would someone ever do that? Data only gets bigger.” And it’s, yeah, well, you haven’t worked with me in production where I accidentally drop a table now and again, and data does get smaller. And the answer for the longest time there was elasticity and auto-scaling was basically unidirectional because that’s what customers are asking for. Right. So, I have to ask, when you say elasticity around HeatWave, is that unidirectional, or does it mean that oh, now there’s less data, so we’re going to go back down again.

Nipun: It is bidirectional, so customers can upsize or they can downsize. Now, I have to say that HeatWave is a highly scalable system. And what that means is that as customers add more nodes to the cluster, the performance of the system improves almost linearly with the number of nodes which have been added. So, as a result, we have a lot of customers who start with a cluster size of certain number, and based on the workloads, they either add nodes or they reduce the number of nodes, right? So, it’s a very common operation; people want to scale up and scale down.

And with the real-time elasticity feature we have introduced, customers can do either operation and with absolutely no downtime. There’s absolutely no time when the cluster is not available for queries or for DMLs, right? So, while the resize operation is going on, this cluster is fully available and customers can upsize to a number of nodes and downsize to any number of nodes.

Corey: As it scales in or scales out, is that effectively doing its own internal sharding and rebalancing of data under the hood, invisible to customers? Is there something else going on? Like, how does this work?

Nipun: Right. So, take the example that customer has, say, four nodes and they want to add two more notes. There are couple of interesting properties over here. We have a technical super-partitioning, by which we know exactly which are the blocks of data which have to be populated to the new nodes which have been added. However, one of the key design points of our elasticity is that there is no data movement between the nodes.

So, all the data which has to be populated in the new nodes which are being added is fetched from the object store, the [OCI 00:15:50] object store. As a result, the existing cluster of four nodes is working as is, queries are working as is, without any degradation in performance. When the data has been populated to these additional nodes, the system then starts having the queries execute on the larger cluster. So, the smaller cluster is available all the time, then the larger clusters available, so from a user’s perspective, they see absolutely no downtime. And since there is no data movement happening from the initial four nodes, there is no degradation of the existing queries which will be running on the older cluster.

Corey: It’s 2022 and you’re announcing enhancements to a technology, so of course, it is a given that you are now talking as well about machine learning. Now, in a general sense, whenever someone says that my immediate instinctive reaction is to check my wallet in case someone is in the middle of picking my pocket because it seems like it winds up in some very weird places. What is machine learning and its applicability to HeatWave? Because generally speaking, when I look at things you can use machine learning for the answer is often finding signal from noise in large datasets and, of course, the ever-popular bias laundering. But I get the sense that neither one of those is quite what you’re talking about here. What monstrosity have you built?

Nipun: With MySQL HeatWave, customers are bringing in more data from either consolidating multiple MySQL databases into one, bringing workloads from other database into MySQL, but the volume of data which now customers are putting into MySQL HeatWave is growing because they want to run transaction processing, analytics all together in one database. Now, as the size of the data is growing, we are finding that many customers want to extract the data or currently need to extract the data out of the MySQL database to run machine-learning processing. So, some of the very large customers of MySQL HeatWave have been using HeatWave very successfully for transaction processing and analytics, but they had to extract the data out to some other ecosystem, to some other service for machine-learning processing. With the announcement we have made, which is HeatWave ML, we are now providing in-database support for machine learning, meaning that customers of MySQL HeatWave can do training, inference, as well as explanations, all inside MySQL HeatWave, without the data or the model ever having to leave MySQL.

And this is something which is fairly unique. Apart from the Oracle database, I’m not aware of any other database, which provides in-database machine-learning capabilities, and certainly not as rich, right, which is very efficient training, inference, and explanations. And all models which are created by HeatWave ML inside MySQL HeatWave can be explained, which is a pretty important capability which enterprise customers like to have.

Corey: This episode is sponsored in part by our friends at Sysdig. Sysdig is the solution for securing DevOps. They have a blog post that went up recently about how an insecure AWS Lambda function could be used as a pivot point to get access into your environment. They’ve also gone deep in-depth with a bunch of other approaches to how DevOps and security are inextricably linked. To learn more, visit sysdig.com and tell them I sent you. That’s S-Y-S-D-I-G dot com. My thanks to them for their continued support of this ridiculous nonsense.

Corey: What does this wind up empowering customers to do? Give an example or two, just because it’s easy to talk about this stuff in the abstract as far as, “Oh, it would theoretically let someone do X, Y or Z.” But the problem I found, generally speaking, in the world of machine learning is that it is challenging to articulate it in a way that people hear the story and think, “Hey, that looks like something I might want to do.” As opposed to the common stories are, “Well, if you have a world-spanning data set and want to do this, this, and this”—like, “Well, I don’t. And I don’t and I don’t and I don’t, so what value is it to me?” What capabilities does it unlock?

Nipun: Right. So, with the introduction of HeatWave, what we had said is that customers don’t need multiple databases: One for transaction processing, one for analytics; they can do both transactional processing and analytics with one database, right? That’s what we started off with. Now, the same thing holds true for machine learning. Current customers of most databases need to extract data out of the database for doing machine learning.

And we are saying, “Hey, that’s not [unintelligible 00:19:45], analytics, mixed workloads, or machine learning. Your data can all be inside MySQL, MySQL HeatWave, and you can do all the processing with that service.” Now, the kinds of capabilities customers like to have for machine learning, training as the most important one. And training is a very time-consuming operation. And typically when customers do training and they’re using some other service, it’s time-consuming and it is very expensive as well.

One of the very interesting properties here is that when you’re running machine learning inside HeatWave, you don’t need to provision any additional cluster, or you don’t need to have any custom gear. This machine-learning training is happening on the same cluster which the user has provisioned for analytics or for transaction processing. So, on the same hardware, on the same cluster, now they can run machine-learning processing. So, the kind of use case which you’re asking is when customers have this data—and I’ll walk you through an example. Take the case of credit card, right?

If a bank wants to determine whether they want to, like, deny someone a credit card or approve it, it’s based on some characteristics. Many of the times, people use a rule-based mechanism, but now with data-driven approaches, people want to look at a lot of data and the system makes a recommendation that yes, this person is appropriate for, like, you know, granting the loan or not. And this is something for which customers—or, like, the enterprises want to have rich models which accurately provide a characterization of the data so that they can make the right predictions. So, training is very important because you want to get the training be done right on the data because it influences the quality of the predictions which are being made. And once a prediction is made, there may be reasons, like, there could be regulatory compliance reasons because of which the enterprise may need to offer an explanation that why was the credit card denied, just to kind of make sure that there wasn’t any bias or unfairness.

And that’s where machine-learning explanation capabilities are also very helpful. So, this is an example: when someone goes for apply for a credit card, whether it’s rejected or approved. Another example is that when someone is making a call, like a marketing team is making a call, and the system want to predict that will a call lead to a successful outcome or not. That’s another example. So, machine learning is being used very—now—extensively, and one of the advantages of a database is a database is where there’s a lot of data, so it’s a very, very good opportunity to harness this data using machine learning. Because machine learning is really tied to the richness of data and to the amount of data someone has.

Corey: That makes a lot of sense. So, it’s… it definitely shines a light at a, if not the easy answer for a lot of those questions, a directions that are people are going to have a better time of mapping to their specific use cases. One that I think is easier for everyone to map to a specific use case is another component of what you folks are announcing which is cost reduction, which is, to be direct, not something people generally think of Oracle as the first example of. A company that’s like, “Ah, that’s the thing that’s going to cost me less money.” And to be clear, I have no problem with that. I pride myself on absolutely not being the least expensive answer to basically anything. But it is an interesting direction to go in. There are a few ways you can wind up saving folks money. Which path have you folks taken?

Nipun: Now, there are multiple ways in which we can reduce the cost for the customer. So, one thing to realize it is MySQL customers are very cost-sensitive. And in the previous benchmarks and results we have shown, we have shown that, you know, compared to other vendors, we are significantly faster—that HeatWave significantly faster and significantly cheaper. So, we have class of customers come to us saying, “Hey, you know what? Can you trade-off some performance for even lower cost?”

And the way we have done is the following: We have doubled the amount of data which can be processed on a HeatWave node. So, HeatWave is an in-memory system, so the size of the cluster depends upon the amount of data which is being processed. And it depends upon the amount of data which can be processed per node. So, if you double the amount of data that can be processed per node, it means that now customers need a cluster half the size compared to what they were doing in the past, which reduces their cost by half. Now, please note, when they’re running on a cluster half the size, the amount of time it takes to run the same query will double.

So, what it means is, the system is providing the same price-performance because half the cost, double the time. But it’s a choice that customers have. If they still want to get the same performance [unintelligible 00:24:38] earlier, they can continue to run on the larger cluster, but now they have a choice. So, in a way, we are providing an even lower entry point for customers. That’s the first part of cost savings.

Corey: And that makes sense because with a lot of the workloads you see where it’s nice to be able to run analytics on the same type of data, you don’t need the same level of responsiveness on a lot of those queries either, where it’s, “So, we’re trying to get an answer to this giant analytics query.” “Okay, so great. How quickly do you need it working?” When transactions are measured in fractions of a second, the answer to analytics queries is, “Well, Tuesday would be nice. We’d like it by Tuesday if you can find a way to pull that off.”

So, there’s no reason to pay for near-line-rate speeds if you don’t need it for a lot of those queries, which is absolutely going to be an interesting option for folks. Now, you said there was a second aspect as well.

Nipun: Yes. And the second aspect is, again, for analytics, right? Customers want to run the queries, they want to run it occasionally, they don’t want to run it all the time, so what we are now introducing is a feature called ‘Pause and Resume.’ And what it does is that if you’re not using the cluster, you can pause and the system makes a copy of the data and all the metadata associated with the data in a backup, and when the user wants, they can resume and, like, you know, fetch the data, which is still in the in-memory presentation and all the metadata associated with Autopilot. And just resume, right? So, this is another way by which customers when they’re not using the cluster for some duration time, they can pause it, and for the duration they pause it, they’re not being charged.

Corey: I am a big believer of the number one step of cloud economics is like, “Oh, should I buy it some reservations or lock into long-term contract?” “No. You should turn things off when you’re not using them.” And people look at you strange, and say, “What? You can turn things off?” And yes, you absolutely can, which makes people feel better about generally not doing it.

But again, customer behaviors are usually ones that makes sense in their context. I just look at from a billing perspective, and it seems a little weird. I like the option, particularly for things that are either non-production or only going to be relevant to production during certain time windows, there are a number of areas where that begins to make an awful lot of sense, and people would do it if it didn’t require backing up the database, destroying the cluster, then re-provisioning the database restoring the cluster. And, yeah, people don’t generally have weeks to spend on spin-up and spin-down.

Nipun: Yes, in fact, that’s a very, very good observation, Corey. I want to say that many of our customers who are running their production workloads and HeatWave, they also have a test environment. And exactly on the lines of what you said, that they want to have a copy of the data in the test environment, should something bad happen, but they don’t want the cluster on all the time. They just wanted for some duration of time and for them, this pause and resume will be a very good idea. And also like, you know, save them money. So, something which we have seen with many of our customers.

Corey: The last component of your announcement is one that I approach with a significant amount of skepticism because every time I start drifting in this direction, one thing is for certain: It’s that I’m going to get yelled at on the internet. I’m referring, of course, to benchmarking. Now, Oracle historically has been a company that prefers people not benchmark and publish results of those benchmarks, backdating into the mists of history. And the argument has always been that people don’t generally tend to benchmark database workloads appropriately, due to a series of misunderstandings, and let’s be clear, this stuff is complicated. And a number of companies in the space love to talk about their benchmarks are great, and when you look into it, it’s okay, those numbers are great.

And you sort of know that the benchmarks that didn’t perform so well are not the ones that they’re talking about. And then their competitor immediately winds up chiming in, where it’s, “Ah, they’re doing it wrong because when you do these other benchmarks, our solution winds up being better.” And it winds up in a nerd slap-fight that no one, even the participants, particularly enjoy. What makes your benchmarks interesting is that you talk through not just what the benchmark results are—because, of course, that’s the entire point—you’re also putting the benchmark methodology and tooling up on GitHub where people can grab it and run it themselves, and see for yourself is the entire approach. That is—how do I put this politely—that is atypical of large companies in general and Oracle in particular. What changed?

Nipun: Right. So, there are three things over here, Corey, right? The first thing is, as we talked about, MySQL is the most popular open-source database in the world. Pretty much all cloud vendors, they have some version of MySQL which they’re offering as a managed service, and in many cases, they’re enhancing MySQL and then offering their service. So, in the context of MySQL, it becomes very important for us to give the opportunity to our customers, for them to compare which service is better for their needs.

So, is more important in the context of MySQL, since everyone is offering it and some of them have derivatives, that we provide some mechanism for people to compare. So, that’s the task for having a benchmark. That’s the first point. Second thing is when you want to compare the performance or the cost of these, like, you know, various flavors, instead of us coming with our own, say, workloads which you see from customers, it’s good to have a well-published benchmark, a well-understood benchmark, so that people can say, “Okay, you know what? Based on TPC-H, what is the performance?” Or, “On [TPC DS 00:29:54], what is the performance?”

In some cases, when a benchmark isn’t available, what we have done is for machine learning, we have used a bunch of open datasets and based on those open datasets, we are publishing the benchmarks to say, “Hey, we are so much faster or so much cheaper.”

And then the third aspect is in terms of why we are making them all available in GitHub or open-source. That these benchmarks are a starting point, but customers will have workloads which are different from these benchmarks, so we want to provide the opportunity for the customers to first look at what is our methodology, what have we used to come up with these numbers so they can reproduce them, but, B, if their workloads are different, they can enhance or augment these benchmarks in the way they would like, and then run them to see how they compare, right? So, we want to be fully transparent about what we have done, how we have done, and let customers decide on their own which is going to be the best platform from a cost perspective, from a performance perspective. So, this is the reason why we have chosen to benchmark and GitHub, like, make available all over scripts in the open-source.

Corey: One of the things I think I admire the most about that is I’ve always viewed benchmarks as being borderline worthless because I do not care in the slightest how your system performs on hand-selected ratings on sample data that you provide, whereas I care everything for how the system performs with my workloads and my data sets. So, unless I am talking to someone who is effectively a neutral third-party benchmark source, in which case they are immediately attacked for being shills for one company or another, and sometimes both or neither at the same time because people are terrible, but seeing how it runs on my workloads and with my constraints is the important and valuable thing. And this is the easiest I can ever see it being for getting a good representative feel for exactly how different offerings are going to perform under the specific conditions that my production environment lives within. Because it’s me we’re talking about the specific conditions of my production environment are, of course, terrifying.

Nipun: Right. So, I want to point out, yes, one is the fact that we have made these benchmarks methodology, like, you know, very transparent, but the second aspect of that is what we talked about last time, which is MySQL Autopilot, right? This is machine-learning-based automation, data-driven-based automation. So, we are very actively working on making it easy for customers to not have to do any configuration changes or optimizations; that the system determines, based on the queries, based on the workloads, how to best tune the system, right?

So, we are working in both angles: One is to make the system more intelligent, so that based on the workload, the system can optimize for the users workload, and then, B, making our approach very transparent so that customers can compare for themselves. So, we are very, very aware of this, and again, for MySQL customers, for many of these open-source customers, simplicity is very important and we are working hard to make it simpler and transparent to our users.

Corey: I really want to thank you for taking me on a tour of what you’re announcing today. Now, so let me ask one of the forbidden questions: What’s on the roadmap? What’s coming that customers can look forward to?

Nipun: So, one of the things which we are working on is that there has been a very good reception of the HeatWave capabilities we have introduced, so MySQL HeatWave is one of the fastest-growing services in the Oracle Cloud. But there has been a lot of interest in customers who have been asking us to provide similar capabilities on AWS. So, this is something which we are working on; it’s in the
roadmap. And please stay tuned for more news on this.

Corey: You can bet that I will. I really want to thank you for taking the time out of your day to basically suffer my slings and arrows, and also spend time teaching what amounts to a remedial database course to a moron. But thank you once again for being as generous with your time as you always are.

Nipun: Well, thank you, Corey. It’s always a pleasure to come and talk to the show. Thank you again, for the opportunity.

Corey: Always. Nipun Agarwal, SVP at Oracle in charge of MySQL, YouSQL, and HeatWave. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice and explain how databases always fail your personal benchmark of doing a SELECT on a terabyte of data at once.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Rick
I lead the developer relations team for strategic accounts at MongoDB. My responsibilities include defining technical standards for the global strategic accounts team and consulting with the largest customers and opportunities for the business. My role spans technology sectors and as part of my engagements I routinely provide guidance on industry best practices, technology transformation, distributed systems implementation, cloud migration, and more. I led the architecture and design effort at Amazon for migrating thousands of relational workloads from RDBMS to NoSQL and built the center of excellence team responsible for defining the best practices and design patterns used today by thousands of Amazon internal service teams and AWS customers. I currently operate as the technical leader for our global strategic account teams to build the market for MongoDB technology by facilitating center of excellence capabilities within our customer organizations through training, evangelism, and direct design consultation activities.

30+ years of software and IT expertise.

9 patents in Cloud Virtualization, Complex Event Processing, Root Cause Analysis, Microprocessor Architecture, and NoSQL Database technology.

Links:

  • MongoDB: https://www.mongodb.com/
  • Twitter: https://twitter.com/houlihan_rick

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: The company 0x4447 builds products to increase standardization and security in AWS organizations. They do this with automated pipelines that use well-structured projects to create secure, easy-to-maintain and fail-tolerant solutions, one of which is their VPN product built on top of the popular OpenVPN project which has no license restrictions; you are only limited by the network card in the instance. To learn more visit: snark.cloud/deployandgo

Corey: This episode is sponsored by our friends at Oracle Cloud. Counting the pennies, but still dreaming of deploying apps instead of “Hello, World” demos? Allow me to introduce you to Oracle’s Always Free tier. It provides over 20 free services and infrastructure, networking, databases, observability, management, and security. And—let me be clear here—it’s actually free. There’s no surprise billing until you intentionally and proactively upgrade your account. This means you can provision a virtual machine instance or spin up an autonomous database that manages itself, all while gaining the networking, load balancing, and storage resources that somehow never quite make it into most free tiers needed to support the application that you want to build. With Always Free, you can do things like run small-scale applications or do proof-of-concept testing without spending a dime. You know that I always like to put asterisks next to the word free? This is actually free, no asterisk. Start now. Visit snark.cloud/oci-free that’s snark.cloud/oci-free.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. A year or two before the pandemic hit, I went on a magical journey to a mythical place called Australia. I know, I was shocked as anyone to figure out that this was in fact real. And while I was there, I gave the opening keynote at a conference that was called Latency Conf, which is great because there’s a heck of a timezone shift, and I imagine that’s what it’s talking about.

The closing keynote was delivered by someone I hadn’t really heard of before, and he started talking about single table design with respect to DynamoDB, which, okay, great; let’s see what he’s got to say. And the talk started off engaging and entertaining and a high-level overview and then got deeper and deeper and deeper and I felt, “Can I please be excused? My brain is full.” That talk was delivered by Rick Houlihan, who now is the Director of Developer Relations for Strategic Accounts over at MongoDB, and I’m fortunate enough to be able to get him here to more or less break down some of what he was saying back then, catch up with what he’s been up to, and more or less suffer my slings and arrows. Rick, thank you for joining me.

Rick: Great. Thanks, Corey. I really appreciate—you brought back some memories, you know, trip down memory lane there. And actually, interestingly enough, that was the world’s introduction to single table design was that. That was my dry-run rehearsal for re:Invent 2018 is where I delivered that talk, and it has become since the most positive—

Corey: This was two weeks before re:Invent, which was just a great thing. I’d been invited to go; why not? I figured I’d see a couple of clients I had out in that direction. And I learned things like Australia is a big place. So, doing a one-week trip, including Sydney, Melbourne, and Perth. Don’t do that.

Rick: I had no idea that it took so long to fly from one side to the other, right? I mean, that’s a long plane [laugh] [crosstalk 00:02:15]—

Corey: Oh, yeah. And you were working at AWS at the time—

Rick: Absolutely.

Corey: —so I can only assume that they basically stuffed you into a dog kennel and threw you underneath the seating area, given their travel policy?

Rick: Well, you know, I have the—[clear throat] actually at the time, they just upgraded the policy to allow the intermediate seating, right? So, if you wanted to get the—

Corey: Ohhh—

Rick: I know—

Corey: Big spender. Big spender.

Rick: Yes, yes. I can get a little bit extra legroom, so I didn’t have my knees shoved into some of these back. But it was good.

Corey: So, let’s talk about, I guess… we’ll call it the elephant in the room. You were at MongoDB, where you were a big proponent of the whole no-SQL side of the world. Then you went to go work at AWS and you carried the good word of DynamoDB far and wide. It made an impression; I built my entire newsletter pipeline production system on top of DynamoDB. It has the same data in three different tables because I’m not good at listening or at computers.

But now you’re back at Mongo. And it’s easy to jump to the conclusion of, “Oh, you’re just shilling for whoever it is that happens to sign your paycheck.” And at this point, are you—what’s the authenticity story? But I’ve been paying attention to what you’ve been saying, and I think that’s a bad take because you have been saying the same things all along since before you were on the Dynamo side of it. I do some research for this show, and you’ve been advocating for outcomes and the right ways to do things. How do you view it?

Rick: That’s basically the story here, right? I’ve always been a proponent of NoSQL. You know, what I took—the knowledge—it was interesting, the knowledge I took from MongoDB evolved as I went to AWS and I delivered, you know, thousands of applications and deployed workloads that I’d never even imagined I would have my hands on before I went there. I mean, honestly, what a great place it was to cut your teeth on data modeling at scale, right? I mean, that’s the—there is no greater scale.

That’s when you learn where things break. And honestly, a lot of the lessons I took from MongoDB, well, when I applied them at scale at AWS, they worked with varying levels of success, and we had to evolve those into the sets of design patterns, which I started to propose for DynamoDB customers, which had been highly effective. I still believe in all those patterns. I would never tell somebody that they need to drop everything and run to MongoDB, but, you know, again, all those patterns apply to MongoDB, too, right? A very—a lot—I wouldn’t say all of them, but many of them, right?

So, I’m a proponent of NoSQL. And I think we talked before the call a little bit about, you know, if I was out there hocking relational technology right now and saying RDBMS is the future, then everybody who criticizes anything I say, I would absolutely have to, you know, say that there’s some validity there. But I’m not saying anything different I’ve ever said. MongoDB announced Serverless, if you remember, in July, and that was a big turning point for me because the API that we offer, the developer experience for MongoDB is unmatched, and this is what I talk to people now. And it’s the patterns that I’ve always proposed, I still model data the same way, I don’t do it any different, and I’ve always said, if you go back to my earlier sessions on NoSQL, it’s all the same.

It doesn’t matter if it’s MongoDB, DynamoDB, or any other technology. I’ve always shown people how to model their data and NoSQL and I don’t care what database you’re using, I’ve actually helped MongoDB customers do their job better over the years as well. So.

Corey: Oh, yeah. And looking back at some of your early talks as well, you passed my test for, “Is this person a shill?” Because you wound up in those talks, addressing head-on when is a relational model the right thing to do? And then you put the answers up on a slide, and this—and what—it didn’t distill down to, “If you’re a fool.”

Rick: [laugh].

Corey: Because there are use cases where if you don’t [unintelligible 00:05:48] your access patterns, if you have certain constraints and requirements, then yeah. That you have always been an advocate for doing the right thing for the workload. And in my experience, for my use cases, when I looked at MongoDB previously, it was not a fit for me. It was very much a you run this on an instance basis, you have to handle all this stuff. Like three—you kno, keeping it in triplicate in three different DynamoDB tables, my newsletter production pipeline now, including backups and the rest, of DynamoDB portion has climbed to the princely sum of $1.30 a month, give or take.

Rick: A month. Yes, exactly.

Corey: So, there’s no answer for that there. Now that Mongo Serverless is coming out into the world, oh, okay, this starts to be a lot more compelling. It starts to be a lot more flexible.

Rick: I was just going to say, for your use case there, Corey, you’re probably looking at the very similar pricing experience now, with MongoDB Serverless. Especially when you look at the pricing model, it’s very close to the on-demand table model. It actually has discounted tiering above it, which I haven’t really broken it down yet against a provision capacity model, but you know, there’s a lot of complexity in DynamoDB pricing. And they’re working on this, they’ll get better at it as well, but right now you have on-demand, you have provisioned throughput, you have [clear throat] reserved capacity allocations. And, you know, there’s a time and place for all of those, but it puts the—again, it’s just complexity, right?

This is the problem that I’ve always had with DynamoDB. I just wish that we’d spent more time on improving the developer experience, right, enhancing the API, implementing some of these features that, you know, help. Let’s make single table design a first-class citizen of the DynamoDB API. Right now it’s a red—it’s a—I don’t want to say redheaded stepchild, I have two [laugh] I have two redhead children and my wife is redhead, but yeah. [laugh].

Corey: [laugh]. That’s—it’s—

Rick: That’s the way it’s treated, right? It’s treated like a stepchild. You know, it’s like, come on, we’re fully funding the solutions within our own umbrella that are competing with ourselves, and at the same time, we’re letting the DynamoDB API languish while our competitors are moving ahead. And eventually, it just becomes, you know, okay, guys, I want to work with the best tooling on the market, and that’s really what it came down to. As long as DynamoDB was the king of serverless, yes, absolutely; best tooling on the market.

And they still are [clear throat] the leader, right? There’s no doubt that DynamoDB is ahead in the serverless landscape, that the MongoDB solution is in its nascency. It’s going to be here, it’s going to be great, that’s part of what I’m here for. And that’s again, getting back to why did you make the move, I want to be part of this, right? That’s really what it comes down to.

Corey: One of the things that I know that was my own bias has always been that if I’m looking at something like—that I’m looking at my customer environments to see what’s there, I can see DynamoDB because it has its own line item in the bill. MongoDB is generally either buried in marketplace charges, or it’s running on a bunch of EC2 instances, or it just shows up as data transfer. So, it’s not as top-of-mind for the way that I view things in… through the lens of you know, billing. So, that does inform my perception, but I also know that when I’m talking to large-scale companies about what they’re doing, when they’re going all-in on AWS, a large number of them still choose things like Mongo. When I’ve asked them why that is, sometimes you get the answer of, “Oh, legacy. It’s what we built on before.” Cool—

Rick: Sure.

Corey: —great. Other times, it’s a, “We’re not planning to leave, but if we ever wanted to go somewhere else, it’s nice to not have to reimagine the entire data architecture and change the integration points start to finish because migrations are hard enough without that.” And there is validity to the idea of a strategic exodus being possible, even if it’s not something you’re actively building for all the time, which I generally advise people not to do.

Rick: Yeah. There’s a couple things that have occurred over the last, you know, couple of years that have changed the enterprise CIO and CTO's assessment of risk, right? Risk is the number one decision factor in a CTOs portfolio and a CIO’s, you know, decision-making process, right? What is the risk? What is the impact of that risk? Do I need to mitigate that risk, or do I accept that risk? Okay?

So, right now, what you’ve seen is with Covid, people have realized that you know, on-prem infrastructure is a risk, right? It used to be an asset; now it’s a risk. Those personnel that have to run that on-prem infrastructure, hey, what happens when they’re not available? The infrastructure is at risk. Okay.

So, offloading that to cloud providers is the natural solution. Great. So, what happens when you offload to a cloud provider and IAD goes down, or you know, us-east-1 goes down—we call it IAD or we used to call it IAD internally at AWS when I was there because, you know, the regions were named by airport codes, but it’s us-east-1—how many times has us-east-1 had problems? Do you want to really be the guy that every time us-east-1 goes down, you’re in trouble? What happens when people in us-east-1 have trouble? Where do they go?

Corey: Down generally speaking.

Rick: [crosstalk 00:10:37]—well, if they’re well-architected, right, if they’re well-architected, what do they do? They go to us-west-2. How much infrastructure is us-west-2 have? So, if everybody in us-east-1 is well-architected, then they all go to us-west-2. What happens in us-west-2? And I guarantee you—and I’ve been warning about this at AWS for years, there’s a cascade failure coming, and it’s going to be coming because we’re well-architecting everybody to failover from our largest region to our smaller regions.

And those smaller regions, they cannot take the load and nobody’s doing any of that planning, so, you know, sooner or later, what you’re going to see is dominoes fall, okay? [clear throat]. And it’s not just going to be us-east-1, it’s going to be us-east-1 failed, and the rollover caused a cascade failure in us-west-2, which caused a cascade—

Corey: Because everyone’s failing over during—

Rick: That’s right. That’s right.

Corey: —this event the same way. And also—again, not to dunk on them unnecessarily, but when—

Rick: No, I’m not dunking.

Corey: —us-east-1 goes, down a lot of the control plane services freeze up—

Rick: Oh, of course they do.

Corey: —like [unintelligible 00:11:25].

Rick: Exactly. Oh, we not single point of failure, right? Uh-huh, exactly. There you go, Route 53, now—and that actually surprised me is DynamoDB instead of Route 53 is your primary database. So, I’m actually must have had some impact on you—

Corey: To move one workload off of Dynamo to Route 53 [crosstalk 00:11:39] issue number because I have to practice what I preach.

Rick: That’s right. Exactly.

Corey: It was weird; they the thing slower and little bit less, uh—

Rick: [laugh]. I love it when [crosstalk 00:11:45]—yeah, yeah—

Corey: —and a little bit [crosstalk 00:11:45] cache-y. But yeah.

Rick: —sure. Okay, I can understand that. [laugh].

Corey: But it made the architecture diagram a little bit more head-scratching, and really, that’s what it’s all about. Getting a high score.

Rick: Right. So, if you think about your data, right, I mean, would you rather be running on an infrastructure that’s tied to a cloud provider that could experience these kinds of regional failures and cascade failures, or would you rather have your data infrastructure go across cloud providers so that when provider has problems, you can just go ahead and switch the light bulb over on the other one and ramp right back up, right? You know? And honestly, you’re running active, active configurations and that kind of, [clear throat] you know, deployment, you know, design, and you’re never going to go down. You’re always going—

Corey: The challenge I’ve had—

Rick: —to be the one that stays up.

Corey: The theory is sound, but the challenge I’ve had in production with trying these things is that one, the thing that winds up handling the failover piece is often causes more outage than the underlying stuff itself.

Rick: Well, sure. Yeah.

Corey: Two, when you’re building something to run a workload to run in multiple cloud providers, you’re forced to use a lot of—

Rick: Lowest common denominator?

Corey: Lowest common denominator stuff. Yeah.

Rick: Yeah, yeah totally. I hear that all the time.

Corey: Unless you’re actively running it in both places, it looks like a DR Plan, which doesn’t survive the next commit to the codebase. It’s the—

Rick: I totally buy that. You’re talking about the stack, stack duplication, all that kind of—that’s an overhead and complexity, I don’t worry about at the data layer, right?

Corey: Oh, yeah.

Rick: The data layer—

Corey: If you’re talking about—

Rick: —[crosstalk 00:12:58]

Corey: —[crosstalk 00:12:58] data layer, oh, everything you’re saying makes perfect sense.

Rick: Makes perfect sense, right? And honestly, you know, let’s put it this way: If this is what you want to do—

Corey: What do you mean identity management and security handover working differently? Oh, that’s a different team’s problem. Oh, I miss those days.

Rick: Yeah, you know, totally right. It’s not ideal. But you know, I mean, honestly, it’s not a deal that somebody wants to manage themselves, is moving that data around. The data is the lock-in. The data is the thing that ties you to—

Corey: And the cost of moving it around in some cases, too.

Rick: That’s exactly right. You know, so you know, having infrastructure that spans providers and spans both on-prem and cloud, potentially, you know, that can span multiple on-prem locations, man, I mean, that’s just that’s power. And MongoDB provides that; I mean, DynamoDB can’t. And that’s really one of the biggest limitations that it will always have, right? And we talked about, and I still believe in the power of global tables, and multi-region deployments, and everything, it’s all real.

But these types of scenarios, I think this is the next generation of failure that the cloud providers are not really prepared for, they haven’t experienced it, they don’t know what it’s even going to look like, and I don’t think you want to be tied to a single provider when these things start happening, right, if you have a large amount of infrastructure deployed someplace. It just seems like [clear throat] that’s a risk that you’re running at these days, and you can mitigate that risk somewhat by going with a MongoDB Atlas. I agree, all those other considerations. But you know, I also heard—it’s a lot of fun, too, right? There’s a lot of fun in that, right?

Because if you think about it, I can deploy technologies in ways on any cloud provider, they’re going to be cloud provider agnostic, right? I can use, you know, containerized technologies, Kubernetes, I can use—hell, I’m not even afraid to use Lambda functions, and just, you know, put a wrapper around that code and deploy it both as a Lambda or a Cloud Function in GCP. The code’s almost the same in many cases, right? What it’s doing with the data, you can code this stuff in a way—I used to do it all the time—you abstract the data layer, right? Create a DAL. How about a CAL? A cloud [laugh] cloud access layer, right, you know? [laugh].

Corey: I wish, on some level, we could go down some of these paths. And someone asked me once a while back of, “Well, you seem to have a lot of opinions on this. Do you think you could build a better cloud than AWS?” And my answer—

Rick: Hell yes.

Corey: —look them a bit by surprise of, “Absolutely. Step one, I want similar resources, so give me $20 billion to spend”—

Rick: I was going to say, right?

Corey: —”then I’m going to hire the smart people.” Not that we’re somehow smarter or better or anything else than the people who built AWS originally, but now—

Rick: We have all those lessons learned.

Corey: —we have fifteen years of experience to fall back on.

Rick: Exactly.

Corey: “Oh. I wouldn’t make that mistake again.”

Rick: Exactly. Don’t need to worry about that. Yeah exactly.

Corey: You can’t just turn off a cloud service and relaunch it with a completely different interface and API and the rest.

Rick: People who criticize, you know, services like DynamoDB, like—and other AWS services—look, these things are like any kind of retooling of the services, it’s like rebuilding the engine on the airplane while it’s flying.

Corey: Oh, yeah.

Rick: And you have to do it with a level of service assurance that—I mean, come on. DynamoDB provides four nines out of the box, right? Five nines if you turn on global tables. And they’re doing this at the same time as they have pipeline releases dropping regularly, right? So, you can imagine what kind of, you know, unit testing goes on there, what kind of Canary deployments are happening.

It’s just, it’s an amazing infrastructure that they maintain, incredibly complex, you know? In some ways, these are lessons that we need to learn in MongoDB if we’re going to be successful operating a shared backplane serverless, you know, processing fabric. We have to look at what DynamoDB does right. And we need to build our own infrastructure that mirrors those things, right? And in some ways, these things are there, in some ways, they’re working on, in some ways, we got a long ways to go.

But you know, I mean, it’s this is the exciting part of that journey for me. Now, in my case, I focus on strategic accounts, right? Strategic accounts are big, you know, they’re the potential to be our whale customers, right? These are probably not customers who would be all that interested in serverless, right? They’re customers that would be more interested in provisioned infrastructure because they’re the people that I talked to when I was at DynamoDB; I would be talking to customers who are interested in like, reserved capacity allocations, right? If you’re talking about—

Corey: Yeah, I wanted to ask you about that. You’re developer advocacy—which I get—for strategic accounts.

Rick: Right.

Corey: And I’m trying to wrap my head around—

Rick: Why [crosstalk 00:17:19]—

Corey: [crosstalk 00:17:19] strategic accounts are the big ones, potential spend lots of stuff. Why do they need special developer advocacy?

Rick: [laugh]. Well, yeah, it’s funny because, you know, one of the reasons why it started talking to Mark Porter about this, you know, was the fact that, you know, the overlap is really around [clear throat] the engagements that I ran when I was doing the Amazon retail migration, right? When Amazon retail started to move to NoSQL, we deprecated 3000 Oracle server instances, we moved a large percentage of those workloads to NoSQL. The vast majority probably just were lift-and-shift into RDS and whatnot because they were too small, too old, not worth upgrading whatnot, but every single tier, what we call tier-one service, right, every money-making service was redesigned and redeployed on DynamoDB, right? So, we’re talking about 25,000 developers that we had to ramp. This is back four years ago; now we have, like, 75,000.

But back then we had 25,000 global developers, we had [clear throat] a technology shift, a fundamental paradigm shift between relational modeling and NoSQL modeling, and the whole entire organization needed to get up to speed, right? So, it was about creating a center of excellence, it was about operating as an office of the CTO within the organization to drive this technology into the DNA of our company. And so that exercise was actually incredibly informative, educational, in that process of executing a technology transformation in a major enterprise. And this is something that we want to reproduce. And it’s actually what I did for Dynamo as well, really more than
anything.

Yes, I was on Twitter, I was on Twitch, I did a lot of these things that were kind of developer advocate, you know, activities, but my primary job at AWS was working with large strategic customers, enabling their teams, you know, teaching them how to model their data in NoSQL, and helping them cross the chasm, right, from relational. And that is advocacy, right? The way I do it is I use their workloads. [clear throat]. I use their—the customers, you know, project teams themselves, I break down their models, I break down their access patterns when I leave, essentially—with the whole day of design reviews, we’ll walk through 12 or 15 workloads, and when I leave these guys have an idea: How would I do it if I wanted to use NoSQL, right?

Give them enough breadcrumbs so that they can actually say, “Okay, if I want to take it to the next step, I can do it without calling up and say, ‘Hey, can we get a professional services team in here?’” right? So, it’s kind of developer advocacy and it’s kind of not, right? We’re kind of recognizing that these are whales, these are customers with internal resources that are so huge, they could suck our Developer’s Advocacy Team in and chew it up, right? So, what we’re trying to do is form a focus team that can hit hard and move the needle inside the accounts. That’s what I’m doing. Essentially, it’s the same work I did for [clear throat] AWS for DynamoDB. I’m just doing it for, you know—they traded for a new quarterback. Let’s put it that way. [laugh].

Corey: This episode is sponsored in part by our friends at Sysdig. Sysdig is the solution for securing DevOps. They have a blog post that went up recently about how an insecure AWS Lambda function could be used as a pivot point to get access into your environment. They’ve also gone deep in-depth with a bunch of other approaches to how DevOps and security are inextricably linked. To learn more, visit sysdig.com and tell them I sent you. That’s S-Y-S-D-I-G dot com. My thanks to them for their continued support of this ridiculous nonsense.

Corey: So, one thing that I find appealing about the approach maps to what I do in the world of cloud economics, where I—like, in my own environment, our AWS bill is creeping up again—we have 14 AWS accounts—and that’s a little over $900 a month now. Which, yeah, big money, big money.

Rick: [laugh].

Corey: In the context of running a company, that no one notices or cares. And our customers spend hundreds of millions a year, pretty commonly. So, I see the stuff in the big accounts and I see the stuff in the tiny account here. Honestly, the more interesting stuff is generally in on the smaller side of the scale, just because you’re not going to have a misconfiguration costing a third of your bill when a third of your bill is $80 million a year. So—

Rick: That’s correct. If you do then that’s a real problem, right?

Corey: Oh yeah.

Rick: [laugh].

Corey: It’s very much a two opposite ends of a very broad spectrum. And advice for folks in one of those situations is often disastrous to folks on the other side of that.

Rick: That’s right. That’s right. I mean, at some scale, managing granularity hurts you, right? The overhead of trying to keep your costs, you know, it—but at the same time, it’s just different, a different measure of cost. There’s a different granularity that you’re looking at, right? I mean, things below a certain, you know, level stop becoming important when, you know, the budget start to get a certain scale or a certain size, right? Theoretically—

Corey: Yeah, for there’s certain workloads, things that I care about with my dollar-a-month Dynamo spend, if I were to move that to Mongo Serverless, great, but my considerations are radically different than a company that is spending millions a month on their database structure.

Rick: That’s right. Really, that’s what it comes down to.

Corey: Yeah, we don’t care about the pennies. We care about is it going to work? How do we back it up? What’s the replication factor?

Rick: And that—but also, it’s more than that. It’s, you know, for me, from my perspective, it really comes down to that, you know, companies are spending millions of dollars a year in database services. These are companies that are spending ten times that, five times that, in you know, in developers, you know, expense, right? Building services, maintaining the code that runs—that the services run.

You know, the biggest problem I had with MongoDB is the level of code complexity. It’s a cut after cut after cut, right? And the way I kind of describe the experience—and other people have described it to me; I didn’t come up with this analogy. I had a customer tell me this as they were leaving DynamoDB—“DynamoDB is death by a thousand cuts. You love it, you start using it, you find a little problem, you start fixing it. You start fixing it. You start fixing—you come up with a pattern. Talk to Rick, he’ll come up with something. He’ll tell you how to do that.” Okay?

And you know, how many customers did I would do this with? You know, and it’s honestly, they’re 15-minute phone calls for me, but every single one of those 15-minute phone calls turns into eight hours of developer time writing the code, debugging it, deploying it over and over again, it’s making sure it’s going the way it’s [crosstalk 00:23:02]—

Corey: Have another 15-minute call with Rick, et cetera, et cetera. Yeah.

Rick: Another 15—exactly. And it’s like okay, that’s you know—eventually, they just get tired of it, right? And I actually had a customer that tell me—a big customer—tell me flat out, “Yeah, you proved that the DynamoDB can support our workload and it’ll probably do it cheaper, but I don’t have a half-a-dozen Ricks on my team, right? I don’t have any Ricks on my team. I can’t be getting you in here every single time we have to do a complex data model overhaul, right?”

And this was—granted, it was one of the more complex implementations that I’ve ever done. In order to make it work. I had to overload the fricking table with multiple access patterns on the partition key, something I never done in my life. I made it work, but it was just—honestly, that was an exercise to me that taught me something. If I have to do this, it’s unnatural, okay?

And that’s—[laugh] you know what I mean? And honestly, there’s API improvements that we could have done to make that less of a problem. It’s not like we haven’t known since the last, I don’t know, I joined the company that a thousand WCUs per storage partition was pretty small. Okay? We’ve kind of known that for I don’t know, since DynamoDB, was invented. As matter of fact is, from what I know, talking to people who were around back then, that was a huge bone of contention back in the day, right? A thousand WCUs, ten
gigabytes, there were a lot of the PEs on the team that were going, “No way. No way. That’s way too small.” And then there were other
people that were like, “Nah, nobody’s ever going to need more than that.” And you know, a lot of this was based on the analysis of
[crosstalk 00:24:28]—

Corey: Oh, nothing ever survives first contact from—

Rick: Of course.

Corey: —customer, particularly a customer who is not themselves deeply familiar with what’s happening under the hood. Like, I had this problem back when I was traveling trainer for Puppet for a while. It was, “Great. Well, Puppet is obviously a piece of crap because everyone I talked to has problems with it.” So, I was one of the early developers behind SaltStack—

Rick: Oh nice.

Corey: —and, “Ah, this is going to be a thing of beauty and it’ll be awesome.” And that lasted until the first time I saw what somebody’s done with it in the wild. It was, “Oh, okay, that’s an [unintelligible 00:25:00] choice.”

Rick: Okay, that’s how—“Yeah, I never thought about that,” right? Happy path. We all love the happy path, right? As we’re working with technologies, we figure out how we like to use it, we all use it that way. Of course, you can solve any problem you want the way that you’d like to solve it. But as soon as someone else takes that clay, they mold a different statue and you go, “Oh, I didn’t realize it could look like that.” Right, exactly.

Corey: So, here’s one for you that I’ve been—I still struggle with this from time to time, but why would I, if I’m building something out—well, first off, why on earth would I do that? I have people for that who are good at things—but if I’m building something out and it has a database layer, why would someone choose NoSQL over—

Rick: Oh, sure.

Corey: —over SQL?

Rick: [crosstalk 00:25:38] question.

Corey: —and let me be clear here—and I’m coming at this from the perspective of someone who, basically me a few years ago, who has no real understanding of what databases are. So, my mental model of a database is Microsoft Excel, where I can fire up a [unintelligible 00:25:51] table of these things—

Rick: Sure. [laugh]. Hey, well then, you know what? Then you should love NoSQL because that’s kind of the best analogy of what is NoSQL. It’s like a spreadsheet, right? Whereas a relational database is like a bunch of spreadsheets, each with their own types of rows, right? So—[laugh].

Corey: Oh, my mind was blown with relational stuff [unintelligible 00:26:07] wait, you could have multiple tables? It’s, “What do you think relational meant there, buddy?” My map of NoSQL was always key and value, and that was it. And that’s all it can be. And sure, for some things, that’s what I use, but not everything.

Rick: That’s right. So, you know, the bottom line is, when you think about the relational database, it all goes back to, you know, the first paper ever written on the relational model, Edgar Codd—and I can’t remember the exact title, but he wrote the distributed model, the data model for distributed systems, something like that. He discussed, you know, the concept of normalization, the power of normalization, why you would want this. And the reason why we wanted this, why he thought this was important, this actually kind of demonstrates how—boy, they used to write killer abstracts to papers, right? It’s like the very first sentence, this is why I’m write in this paper. You read the first sentence, you know: “Future users of modern computer systems must have a way to be able to ask questions of the data without knowing how to write code.”

I mean, I don’t know if those were the words, but that was basically what he said, that was why he invented the normalized data model. Because, you know, with the hierarchical management systems at the time, everyone had to know everything about the data in order to be able to get any answers, right? And he was like, “No, I want to be able to just write a question and have the system answer that.” Now, at the time, a lot of people felt like that’s great, and they agreed with his normalized model—it was elegant—but they all believe that the CPU overhead at the time was way too high, right? To generate these views of data on the fly, no freaking way. Storage is expensive. But it ain’t that expensive, right?

Well, this little thing called Moore’s Law, right? Moore’s Law balanced his checkbook for, like, 40 years, 50 years, it balanced the relational database checkbook, okay? So, as the CPUs got faster and faster, crunching, the data became less and less of a problem, okay? And so we crunched bigger and bigger data sets, we got very, very happy with this. Up until about 2014.

At 2014, a really interesting thing happened. If you look at the top 500, which is the supercomputers, the top 500 supercomputing clusters around the world, and you look at their performance increases year-to-year after 2014, it went off a cliff. No longer beating Moore’s Law. Ever since, they’ve been—and per-core performance, you know, CPU, you know, instructions executed per second, everything. It’s just flattening. Those curves are flattening. Moore’s Law is broken.

Now, you’ll get people argue about it, but the reality is, if it wasn’t broken, the top 500 would still be cruising away. They’re not. Okay? So, what this is telling us is that the relational database is losing its horsepower. Okay?

Why is it happening? Because, you know, gate length has an absolute minimum, it’s called zero, right? We can’t have a logic gate that’s the—with negative distance, right? [laugh]. So, you know, these things—but storage, storage, hey, it just keeps on getting cheaper and cheaper, right?

We’re going the other way with storage, right? It’s gigabytes, it’s terabytes, it’s petabytes, you know, with CPU, we’re going smaller and smaller and smaller, and the fab cost is increasing. There’s just—it’s going to take a next-generation CPU technology to get back on track with Moore’s Law.

Corey: Well, here’s the challenge. Everything you’re saying makes perfect sense from where your perspective is. I reiterate, you are working with strategic accounts, which means ‘big.’ When I’m building something out in the evenings because I want to see if something is possible, performance considerations and that sort of characteristic does not factor into it. When I’m a very small-scale, I care about cost to some extent—sure, whatever—but the far more expensive aspect of it, in the ways that matter have been what is the expensive—what—the big expensive piece is—

Rick: We’ve talked about it.

Corey: —engineering time—

Rick: That’s what we just talked about, right?

Corey: —where it’s, “What I’m I familiar with?”

Rick: As a developer, right, why would I use MongoDB over DynamoDB? Because the developer experience [crosstalk 00:29:33]—

Corey: Exactly. Sure, down the road there are performance characteristics and yeah, at the time I have this super-large, scaled-out, complex workload, yeah, but most workloads will not get to that.

Rick: Will not ever get there. Ever get there. [crosstalk 00:29:45]—

Corey: Yeah, so optimizing for [crosstalk 00:29:45], how’s it going to work when I’m Facebook-scale? It’s—

Rick: So, first of—no, exactly, Facebook scale is irrelevant here. What I’m talking about is actually a cost ratchet that’s going to lever on midsize workloads soon, right? Within the next four to five years, you’re going to see mid-level workloads start to suffer from significant performance cost deficiencies compared to NoSQL workloads running on the same. Now you—hell, you see it right now, but you don’t really experience it, like you said, until you get to scale, right? But in midsize workloads, [clear throat] that’s going to start showing up, right? This cost overhead cannot go away.

Now, the other thing here that you got to understand is, just because it’s new technology doesn’t make it harder to use. Just because you don’t know how to use something, right, doesn’t mean that it’s more difficult. And NoSQL databases are not more difficult than the relational database. I can express every single relationship in a NoSQL database that I express in a relational database. If you think about the modern OLTP applications, we’ve done the analysis, ad nauseum: 70% of access patterns are for a single object, a single row of data from a single table; another 20% are for a row of datas—a range of rows from a single table. Okay, that leaves only 10% of your access patterns involve any kind of complex table traversal or entity traversals. Okay?

And most of those are simple one-to-many hierarchies. So, let’s put those into perspective here: 99% of the access patterns in an OLTP application can be modeled without denormalization in a single table. Because single table doesn’t require—just because I put all the objects in one place doesn’t mean that it’s denormalized. Denormalized requires strong redundancies in the stored set. Duplication of data. Okay?

Edgar Codd himself said that the normalized data model does not depend on storage, that they are irrelevant. I could put all the objects in the same document. As long as there’s no duplication of data, there’s no denormalization. I know, I can see your head going, “Wow,” but it’s true, right? Because as long as I can clearly express the relationships of the data without strong redundancies, it is a normalized data model.

That’s what most people don’t understand. NoSQL does not require denormalization. That’s a decision you make, and it usually happens when you have many-to-many relationships; then we need to start duplicating the data.

Corey: In many cases, at least my own experience—because again, I am bad at computers—I find that the data model is not something that is sat out—that you sit down and consciously plan very often. It’s rather something—

Rick: Oh yeah.

Corey: —happens to you instead. I mean—

Rick: That’s right. [laugh].

Corey: —realistically, like, using DynamoDB for this is aspirational. I just checked, and if I look at—so I started this newsletter back in March of 2017. I spun up this DynamoDB table that backs it, and I know it’s the one that’s in production because it has the word ‘test’ in its name, because of course it does. And I’m looking into it, and it has 8700 items in it now and it’s 3.7 megabytes. It’s—

Rick: Sure, oh boy. Nothing, right?

Corey: —not for nothing, this could have been just as easily and probably less complex for my level of understanding at the time, a CSV file that I—

Rick: Right. Exactly, right.

Corey: —grabbed from a Lambda out of S3, do the thing to it, and then put it back.

Rick: [unintelligible 00:32:45]. Right.

Corey: And then from a performance and perspective side on my side, it would make no discernible difference.

Rick: That’s right because you’re not making high-velocity requests against the small object. It’s just a single request every now and then.
S3 performance would probably—you might even be less. It might even cost you less to use S3.

Corey: Right. And 30 to 100 of the latest ones are the only things that are ever looked at in any given week, the rest of it is mostly deadstock that could be transitioned out elsewhere.

Rick: Exactly.

Corey: But again, like, now that they have their lower cost infrequent access storage, then great. It’s not item level; it’s table levels, so
what’s the point? I can knock that $1.30 a month down to, what, $1.10?

Rick: Oh well, yeah, no, I mean, again, Corey for those small workloads, you know what? It’s like, go with what you know. But the reality is, look, as a developer, we should always want to know more, and we should always want to know new things, and we should always be aware of where the industry is headed. And honestly, I’ve heard through—I’m an old, old school, relational guy, okay, I cut my teeth on—oh, God, I don’t even know what version of MS SQL Server it was, but when I was, you know, interviewing at MongoDB. I was talking to Dan Pasette, about the old Enterprise Manager, where we did the schema designer and all this, and we were reminiscing about, you know, back in the day, right?

Yeah, you know, reality of things are is that if you don’t get tuned into the new tooling, then you’re going to get left behind sooner or later. And I know a lot of people who that has happened to over the years. There’s a reason why I’m 56 years old and still relevant in tech, okay? [laugh].

Corey: Mainframes, right? I kid.

Rick: Yes, mainframes.

Corey: I kid. You’re not that much older than I am, let’s be clear here.

Rick: You know what? I worked on them, okay? And some of my peers, they never stopped, right? They just kind of stayed there.

Corey: I’m still waiting for AWS/400. We don’t see them yet, but hope springs eternal.

Rick: I love it. I love that. But no, one of the things that you just said that I think it hit me really, it’s like the data model isn’t something you think about. The data model is something that just happens, right? And you know what, that is a problem because this is exactly what developers today think. They think know the relational database, but they don’t.

You talk to any DBA out there who’s coming in after the fact and cleaned up all the crappy SQL that people like me wrote, okay? I mean, honestly, I wrote some stuff in the day that I thought, “This is perfect. There’s no way that could be anything better than this,” right? Nice derived table joins insi—and you know what? Then here comes the DBA when the server is running at 90% CPU and 100% percent memory utilization and page swapping like crazy, and you’re saying we got to start sharding the dataset.

And you know, my director of engineering at the time said, “No, no, no. What we need is somebody to come in and clean up our SQL.” I said, “What do you mean? I wrote that SQL.” He’s like, “Like I said, we need someone to come and clean up our SQL.”

I said, “Okay, fine.” We brought the guy in. 1500 bucks an hour, we paid this guy, I was like, “There’s no way that this guy is going to be
worth that.” A day and a half later, our servers are running at 50% CPU and 20% memory utilization. And we’re thinking about, you know, canceling orders for additional hardware. And this was back in the day before cloud.

So, you know, developers think they know what they’re doing. [clear throat]. They don’t know what they’re doing when it comes to the database. And don’t think just because it’s a relational database and they can hack it easier that it’s better, right? Yeah, it’s, there’s no substitute for knowing what you’re doing; that’s what it comes down to.

So, you know, if you’re going to use a relational database, then learn it. And honestly, it’s a hell of a lot more complicated to learn a
relational database and do it well than it is to learn how to model your data in NoSQL. So, if you sit two developers down, and you say, “You learn NoSQL, you learn relational,” two months later, this guy is still going to be studying. This guy’s going to be writing code for seven weeks. Okay? [laugh]. So, you know, that’s what it comes down to. You want to go fast, use NoSQL and you won’t have any
problems.

Corey: I think that’s a good place to leave it. If people want to learn more about how you view these things, where’s the best place to find you?

Rick: You know, always hit me up on Twitter, right? I mean, @houlihan_rick, that’s my—underbar rick, that’s my Twitter handle. And you know, I apologize to folks who have hit me up on Twitter and gotten no response. My Twitter as you probably have as well, my message request box is about 3000 deep.

So, you know, every now and then I’ll start going in there and I’ll dig through, and I’ll reply to somebody who actually hit me up three months ago if I get that far down the queue. It is a Last In, First Out, right? I try to keep things as current as possible. [laugh].

Corey: [crosstalk 00:36:51]. My DMs are a trash fire. Apologies as well. And we will, of course, put links to it in the [show notes 00:36:55].

Rick: Absolutely.

Corey: Thank you so much for your time. I really do appreciate it. It’s always an education talking to you about this stuff.

Rick: I really appreciate being on the show. Thanks a lot. Look forward to seeing where things go.

Corey: Likewise.

Rick: All right.

Corey: Rick Houlihan Director of Developer Relations, Strategic Accounts at MongoDB. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an upset comment talking about how we didn’t go into the proper and purest expression of NoSQL non-relational data, DNS TXT records.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Rachel

Rachel leads product and technical marketing for Chronosphere. Previously, Rachel wore lots of marketing hats at CloudHealth (acquired by VMware), and before that, she led product marketing for cloud-integrated storage at NetApp. She also spent many years as an analyst at Forrester Research. Outside of work, Rachel tries to keep up with her young son and hyper-active dog, and when she has time, enjoys crafting and eating out at local restaurants in Boston where she’s based.

Links:

  • Chronosphere: https://chronosphere.io
  • Twitter: https://twitter.com/RachelDines
  • Email: rachel@chronosphere.io

TranscriptAnnouncer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: The company 0x4447 builds products to increase standardization and security in AWS organizations. They do this with automated pipelines that use well-structured projects to create secure, easy-to-maintain and fail-tolerant solutions, one of which is their VPN product built on top of the popular OpenVPN project which has no license restrictions; you are only limited by the network card in the instance. To learn more visit: snark.cloud/deployandgo

Corey: Couchbase Capella Database-as-a-Service is flexible, full-featured and fully managed with built in access via key-value, SQL, and full-text search. Flexible JSON documents aligned to your applications and workloads. Build faster with blazing fast in-memory performance and automated replication and scaling while reducing cost. Capella has the best price performance of any fully managed document database. Visit couchbase.com/screaminginthecloud to try Capella today for free and be up and running in three minutes with no credit card required. Couchbase Capella: make your data sing.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. A repeat guest joins me today, and instead of talking about where she works, instead we’re going to talk about how she got there. Rachel Dines is the Head of Product and Technical Marketing at Chronosphere. Rachel, thank you for joining me.

Rachel: Thanks, Corey. It’s great to be here again.

Corey: So, back in the early days of me getting started, well, I guess all this nonsense, I was an independent consultant working in the world of cloud cost management and you were over at CloudHealth, which was effectively the 800-pound gorilla in that space. I’ve gotten louder, and of course, that means noisier as well. You wound up going through the acquisition by VMware at CloudHealth, and now you’re over at Chronosphere. We’re going to get to all of that, but I’d rather start at the beginning, which, you know, when you’re telling stories seems like a reasonable place to start. Your first job out of school, to my understanding, was as an analyst at Forrester is that correct?

Rachel: It was yeah. Actually, I started as a research associate at Forrester and eventually became an analyst. But yes, it was Forrester. And when I was leaving school—you know, I studied art history and computer science, which is a great combination, makes a ton of sense—I can explain it another time—and I really wanted to go work at the equivalent of FAANG back then, which was just Google. I really wanted to go work at Google.

And I did the whole song-and-dance interview there and did not get the job. Best thing that’s ever happened to me because the next day a Forrester recruiter called. I didn’t know what Forrester was—once again, I was right out of college—I said, “This sounds kind of interesting. I’ll check it out.” Seven years later, I was a principal analyst covering, you know, cloud-to-cloud resiliency and backup to the cloud and cloud storage. And that was an amazing start to my career, that really, I’m credited a lot of the things I’ve learned and done since then on that start at Forrester.

Corey: Well, I’ll admit this: I was disturbingly far into my 30s before I started to realize what it is that Forrester and its endless brethren did. I’m almost certain you can tell that story better than I can, so what is it that Forrester does? What is its place in the ecosystem?

Rachel: Forrester is one of the two or three biggest industry analyst firms. So, the people that work there—the analysts there—are basically paid to be, like, big thinkers and strategists and analysts, right? There’s a reason it’s called that. And so the way that we spent all of our time was, you know, talking to interesting large, typically enterprise IT, and I was in the infrastructure and operations group, so I was speaking to infrastructure, ops, precursors to DevOps—DevOps wasn’t really a thing back in ye olden times, but we’re speaking to them and learning their best practices and publishing reports about the technology, the people and the process that they dealt with. And so you know, over a course of a year, I would talk to hundreds of different large enterprises, the infrastructure and ops leaders at everyone from, like, American Express to Johnson & Johnson to Monsanto, learn from them, write research and reports, and also do
things like inquiries and speaking engagements and that kind of stuff.

So, the idea of industry analysts is that they’re neutral, they’re objective. You can go to them for advice, and they can tell you, you know, these are the shortlist of vendors you should consider and this is what you should look for in a solution.

Corey: I love the idea of what that role is, but it took me a while as a condescending engineer to really wrap my head around it because I viewed it as oh, it’s just for a cover your ass exercise so that when a big company makes a decision, they don’t get yelled at later, and they said, “Well, it seemed like the right thing to do. You can’t blame us.” And that is an overwhelmingly cynical perspective. But the way it was explained to me, it really was put into context—of all things—by way of using the AWS bill as a lens. There’s a whole bunch of tools and scripts and whatnot on GitHub that will tell you different things about your AWS environment, and if I run them in my environment, yeah, they work super well.

I run them in a client environment and the thing explodes because it’s not designed to work at a scale of 10,000 instances in a single availability zone. It’s not designed to do backing off so it doesn’t exhaust rate limits across the board. It requires a rethinking at that scale. When you’re talking about enterprise-scale, a lot of the Twitter zeitgeist, as it were, about what tools work well and what tools don’t for various startups, they fail to cross over into the bowels of a regulated entity that has a bunch of other governance and management concerns that don’t really apply. So, there’s this idea of okay, now that we’re a large, going entity with serious revenue behind this, and migrating to any of these things is a substantial lift. What is the right answer? And that is sort of how I see the role of these companies in the ecosystem playing out. Is that directionally correct?

Rachel: I would definitely agree that that is directionally correct. And it was the direction that it was going when I was there at Forrester. And by the way, I’ve been gone from there for, I think, eight-plus years. So, you know, it’s definitely evolved it this space—

Corey: A lifetime in tech.

Rachel: Literally feels like a lifetime. Towards the end of my time there was when we were starting to get briefings from this bookstore company—you might have heard of them—um, Amazon?

Corey: Barnes and Noble.

Rachel: Yes. And Barnes and Noble. Yes. So, we’re starting to get briefings from Amazon, you know, about Amazon Web Services, and S3 had just been introduced. And I got really excited about Netflix and chaos engineering—this was 2012, right?—and so I did a bunch of research on chaos engineering and tried to figure out how it could apply to the enterprises.

And I would, like, bring it to Capital One, and they were like, “Ya crazy.” Turns out I think I was just a little bit ahead of my time, and I’m seeing a lot more of the industry analysts now today looking at like, “Okay, well, yeah, what is Uber doing? Like, what is Netflix doing?” And figure out how that can translate to the enterprise. And it’s not a one-to-one, right, just because the people and the structures and the process is so different, so the technology can’t just, like, make the leap on its own. But yes, I would definitely agree with that, but it hasn’t necessarily always been that way.

Corey: Oh, yeah. Like, these days, we’re seeing serverless adoption on some levels being driven by enterprises. I mean, Liberty Mutual is doing stuff there that is really at the avant-garde that startups are learning from. It’s really neat to see that being turned on its head because you always see these big enterprises saying, “We’re like a startup,” but you never see a startup saying, “We’re like a big enterprise.” Because that’s evocative of something that isn’t generally compelling.

“Well, what does that mean, exactly? You take forever to do expense reports, and then you get super finicky about it, and you have so much bureaucracy?” No, no, no, it’s, “Now, that we’re process bound, it’s that we understand data sovereignty and things like that.” But you didn’t stay there forever. You at some point decided, okay, talking to people who are working in this industry is all well and good, but time for you to go work in that industry yourself. And you went to, I believe, NetApp by way of Riverbed.

Rachel: Yes, yeah. So, I left Forrester and I went over to Riverbed to work on their cloud storage solution as a product marketing. And I had an amazing six months at Riverbed, but I happened to join, unfortunately, right around the time they were being taken private, and they ended up divesting their storage product line off to NetApp. And they divested some of their other product lines to some other companies as part of the whole deal going private. So, it was a short stint at Riverbed, although I’ve met some people that I’ve stayed in touch with and are still my friends, you know, many years later.

And so, yeah, ended up over at NetApp. And it wasn’t necessarily what I had initially planned for, but it was a really fun opportunity to take a cloud-integrated storage product—so it was an appliance that people put in their data centers; you could send backups to it, and it shipped those backups on the back end to S3 and then to Glacier when that came out—trying to make that successful in a company that was really not overly associated with cloud. That was a really fun process and a fun journey. And now I look at NetApp and where they are today, and they’ve acquired Spot and they’ve acquired CloudCheckr, and they’re, like, really going all-in in public cloud. And I like to think, like, “Hey, I was in the early days of that.” But yeah, so that was an interesting time in my life for multiple reasons.

Corey: Yeah, Spot was a fascinating product, and I was surprised to see it go to NetApp. It was one of those acquisitions that didn’t make a whole lot of sense to me at the time. NetApp has always been one of those companies I hold in relatively high regard. Back when I was coming up in the industry, a bit before the 2012s or so, it was routinely ranked as the number one tech employer on a whole bunch of surveys. And I don’t think these were the kinds of surveys you can just buy your way to the top of.

People who worked there seemed genuinely happy, the technology was fantastic, and it was, for example, the one use case in which I would run a database where its data store lived on a network file system. I kept whining at the EFS people over at AWS for years that well, EFS is great and all but it’s no NetApp. Then they released NetApps on tap on FSX as a first-party service, in which case, okay, thank you. You have now solved every last reservation I have around this. Onward.

And I still hold the system in high regard. But it has, on some level, seen an erosion. We’re no longer in a world where I am hurling big money—or medium money by enterprise standards—off to NetApp for their filers. It instead is something that the cloud providers are providing, and last time I checked, no matter how much I spend on AWS they wouldn’t let me shove a NetApp filer into us-east-1 without asking some very uncomfortable questions.

Rachel: Yeah. The whole storage industry is changing really quickly, and more of the traditional on-premises storage vendors have needed to adapt or… not, you know, be very successful. I think that NetApp’s done a nice job of adapting in recent years. But I’d been in storage and backup for my entire career at that point, and I was like, I need to get out. I’m done with storage. I’m done with backup. I’m done with disaster recovery. I had that time; I want to go try something totally new.

And that was how I ended up leaving NetApp and joining CloudHealth. Because I’d never really done the startup thing. I done a medium-sized company at Riverbed; I’d done a pretty big company at NetApp. I’ve always been an entrepreneur at heart. I started my first business on the playground in second grade, and it was reselling sticks of gum. Like, I would go use my allowance to buy a big pack of gum, and then I sold the sticks individually for ten cents apiece, making a killer margin. And it was a subscription, actually. [laugh].

Corey: Administrations generally—at least public schools—generally tend to turn a—have a dim view of those things, as I recall from my misspent youth.

Rachel: Yeah. I was shut down pretty quickly, but it was a brilliant business model. It was—so you had to join the club to even be able to buy into getting the sticks of gum. I was, you know, all over the subscription business [laugh] back then.

Corey: And area I want to explore here is you mentioned that you double-majored. One of those majors was computer science—art history was sort of set aside for the moment, it doesn’t really align with either direction here—then you served as a research associate turned analyst, and then you went into product marketing, which is an interesting direction to go in. Why’d you do it?

Rachel: You know, product marketing and industry analysts are there’s a lot of synergy; there’s a lot of things that are in common between those two. And in fact, when you see people moving back and forth from the analyst world to the vendor side, a lot of the time it is to product marketing or product management. I mean, product marketing, our whole job is to take really complex technical concepts and relate them back to business concepts and make them make sense of the broader world and tell a narrative around it. That’s a lot of what an analyst is doing too. So, you know, analysts are writing, they’re giving public talks, they’re coming up with big ideas; that’s what a great product marketer is doing also.

So, for me, that shift was actually very natural. And by the way, like, when I graduated from school, I knew I was never going to code for a living. I had learned all I was going to learn and I knew it wasn’t for me. Huge props, like, you know, all the people that do code for a living, I knew I couldn’t do it. I wasn’t cut out for it.

Corey: I found somewhat similar discoveries on my own journey. I can configure things for a living, it’s fun, but I still need to work with people, past a certain point. I know I’ve talked about this before on some of these shows, but for me, when starting out independently, I sort of assumed at some level, I was going to shut it down, and well, and then I’ll go back to being an SRE or managing an ops team. And it was only somewhat recently that I had the revelation that if everything that I’m building here collapses out from under me or gets acquired or whatnot and I have to go get a real job again, I’ll almost certainly be doing something in the marketing space as opposed to the engineering space. And that was an interesting adjustment to my self-image as I went through it.

Because I’ve built everything that I’ve been doing up until this point, aligned at… a certain level of technical delivery and building things as an engineer, admittedly a mediocre one. And it took me a fair bit of time to get, I guess, over the idea of myself in that context of, “Wow, you’re not really an engineer. Are you a tech worker?” Kind of. And I sort of find myself existing in the in-between spaces.

Did you have similar reticence when you went down the marketing path or was it something that you had, I guess, a more mature view of it [laugh] than I did and said, “Yeah, I see the value immediately,” whereas I had to basically be dragged there kicking and screaming?

Rachel: Well, first of all, Corey, congratulations for coming to terms with the fact that you are a marketer. I saw it in you from the minute I met you, and I think I’ve known you since before you were famous. That’s my claim to fame is that I knew you before you were famous. But for me personally, no, I didn’t actually have that stigma. But that does exist in this industry.

I mean, I think people are—think they look down on marketing as kind of like ugh, you know, “The product sells itself. The product markets itself. We don’t need that.” But when you’re on the inside, you know you can have an amazing product and if you don’t position it well and if you don’t message it well, it’s never going to succeed.

Corey: Our consulting [sub-projects 00:14:31] are basically if you bring us in, you will turn a profit on the engaging. We are selling what basically [unintelligible 00:14:37] money. It is one of the easiest ROI calculations. And it still requires a significant amount of work on positioning even on the sales process alone. There’s no such thing as an easy enterprise sale.

And you’re right, in fact, I think the first time we met, I was still running a DevOps team at a company and I was deploying the product that you were doing marketing for. And that was quite the experience. Honestly, it was one of the—please don’t take this the wrong way at all—but you were at CloudHealth at the time and the entire point was that it was effectively positioned in such a way of, right, this winds up solving a lot of the problems that we have in the AWS bill. And looking at how some of those things were working, it was this is an annoying, obnoxious problem that I wish I could pay to make someone else’s problem, just to make it go away. Well, that indirectly led to exactly where we are now.

And it’s really been an interesting ride, just seeing how that whole thing has evolved. How did you wind up finding yourself at CloudHealth? Because after VMware, you said it was time to go to a startup. And it’s interesting because I look at where you’ve been now, and CloudHealth itself gets dwarfed by VMware, which is sort of the exact opposite of a startup, due to the acquisition. But CloudHealth
was independent for years while you were there.

Rachel: Yeah, it was. I was at CloudHealth for about three-plus years before we were acquired. You know, how did I end up there? It’s… it’s all hazy. I was looking at a lot of startups, I was looking for, like, you know, a Series B company, about 50 people, I wanted something in the public cloud space, but not storage—if I could get away from storage that was the dream—and I met the folks from CloudHealth, and obviously, I hadn’t heard about—I didn’t know about cloud cost management or cloud governance or FinOps, like, none of those were things back then, but I was I just was really attracted to the vision of the founders.

The founders were, you know, Joe Kinsella and Dan Phillips and Dave Eicher, and I was like, “Hey, they’ve built startups before. They’ve got a great idea.” Joe had felt this pain when he was a customer of AWS in the early days, and so I was like—

Corey: As have we all.

Rachel: Right?

Corey: I don’t think you’ll find anyone in this space who hasn’t been a customer in that situation and realized just how painful and maddening the whole space is.

Rachel: Exactly, yeah. And he was an early customer back in, I think, 2014, 2015. So yeah, I met the team, I really believed in their vision, and I jumped in. And it was really amazing journey, and I got to build a pretty big team over time. By the time we were acquired a couple of years later, I think we were maybe three or 400 people. And actually, fun story. We were acquired the same week my son was born, so that was an exciting experience. A lot of change happened in my life all at once.

But during the time there, I got to, you know, work with some really, really cool large cloud-scale organizations. And that was during that time that I started to learn more about Kubernetes and Mesos at the time, and started on the journey that led me to where I am now. But that was one of the happiest accidents, similar to the happy accident of, like, how did I end up at Forrester? Well, I didn’t get the job at Google. [laugh]. How did I end up at CloudHealth? I got connected with the founders and their story was really inspiring.

Corey: Couchbase Capella Database-as-a-Service is flexible, full-featured and fully managed with built in access via key-value, SQL, and full-text search. Flexible JSON documents aligned to your applications and workloads. Build faster with blazing fast in-memory performance and automated replication and scaling while reducing cost. Capella has the best price performance of any fully managed document database. Visit couchbase.com/screaminginthecloud to try Capella today for free and be up and running in three minutes with no credit card required. Couchbase Capella: make your data sing.

Corey: It’s amusing to me the idea that, oh, you’re at NetApp if you want to go do something that is absolutely not storage. Great. So, you go work at CloudHealth. You’re like, “All right. Things are great.” Now, to take a big sip of scalding hot coffee and see just how big AWS billing data could possibly be. Yeah, oops, you’re a storage company all over again.

Some of our, honestly, our largest bills these days are RDS, Athena, and of course, S3 for all of the bills storage we wind up doing for our customers. And it is… it is not small. And that has become sort of an eye-opener for me just the fact that this is, on some level, a big data
problem.

Rachel: Yeah.

Corey: And how do you wind up even understanding all the data that lives in just the outputs of the billing system? Which I feel is sort of a good setup for the next question of after the acquisition, you stayed at VMware for a while and then matriculated out to where you are now where you’re the Head of Product and Technical Marketing at Chronosphere, which is in the observability space. How did you get there from cloud bills?

Rachel: Yeah. So, it all makes sense when I piece it together in my mind. So, when I was at CloudHealth, one of the big, big pain points I was seeing from a lot of our customers was the growth in their monitoring bills. Like, they would be like, “Okay, thanks. You helped us, you know, with our EC2 reservations, and we did right-sizing, and you help with this. But, like, can you help with our Datadog bill? Like, can you help with our New Relic bill?”

And that was becoming the next biggest line item for them. And in some cases, they were spending more on monitoring and APM and like, what we now call some things observability, they were spending more on that than they were on their public cloud, which is just bananas. So, I would see them making really kind of bizarre and sometimes they’d have to make choices that were really not the best choices. Like, “I guess we’re not going to monitor the lab anymore. We’re just going to uninstall the agents because we can’t pay this anymore.”

Corey: Going down from full observability into sampling. I remember that. The New Relic shuffle is what I believe we call it at the time. Let’s be clear, they have since fixed a lot of their pricing challenges, but it was the idea of great suddenly we’re doing a lot more staging environments, and they come knocking asking for more money but it’s a—I don’t need that level of visibility in the pre-prod environments, I guess. I hate doing it that way because then you have a divergence between pre-prod and actual prod. But it was economically just a challenge. Yeah, because again, when it comes to cloud, architecture and cost are really one and the same.

Rachel: Exactly. And it’s not so much that, like—sure, you know, you can fix the pricing model, but there’s still the underlying issue of it’s not black and white, right? My pre-prod data is not the same value as my prod data, so I shouldn’t have to treat it the same way, shouldn’t have to pay for it the same way. So, seeing that trend on the one hand, and then, on the other hand, 2017, 2018, I started working on the container cost allocation products at CloudHealth, and we were—you know, this was even before that, maybe 2017, we were arguing about, like, Mesos and Kubernetes and which one was going to be, and I got kind of—got very interested in that world.

And so once again, as I was getting to the point where I was ready to leave CloudHealth, I was like, okay, there’s two key things I’m seeing in the market. One is people need a change in their monitoring and observability; what they’re doing now isn’t working. And two, cloud-native is coming up, coming fast, and it’s going to really disrupt this market. So, I went looking for someone that was at the intersection of the two. And that’s when I met the team at Chronosphere, and just immediately hit it off with the founders in a similar way to where I hit it off with the founders that CloudHealth. At Chronosphere, the founders had felt pain—

Corey: Team is so important in these things.

Rachel: It’s really the only thing to me. Like, you spend so much time at work. You need to love who you work with. You need to love your—not love them, but, you know, you need to work with people that you enjoy working with and people that you learn from.

Corey: You don’t have to love all your coworkers, and at best you can get away with just being civil with them, but it’s so much nicer when you can have a productive, working relationship. And that is very far from we’re going to go hang out, have beers after work because that leads to a monoculture. But the ability to really enjoy the people that you work with is so important and I wish that more folks paid attention to that.

Rachel: Yeah, that’s so important to me. And so I met the team, the team was fantastic, just incredibly smart and dedicated people. And then the technology, it makes sense. We like to joke that we’re not just taking the box—the observability box—and writing Kubernetes in Crayon on the outside. It was built from the ground up for cloud-native, right?

So, it’s built for this speed, containers coming and going all the time, for the scale, just how much more metrics and observability data that containers emit, the interdependencies between all of your microservices and your containers, like, all of that stuff. When you combine it makes the older… let’s call them legacy. It’s crazy to call, like, some of these SaaS solutions legacy but they really are; they weren’t built for cloud-native, they were built for VMs and a more traditional cloud infrastructure, and they’re starting to fall over. So, that’s how I got involved. It’s actually, as we record, it’s my one-year anniversary at Chronosphere. Which is, it’s been a really wild year. We’ve grown a lot.

Corey: Congratulations. I usually celebrate those by having a surprise meeting with my boss and someone I’ve never met before from HR. They don’t offer your coffee. They have the manila envelope of doom in front of them and hold on, it’s going to be a wild meeting. But on the plus side, you get to leave work early today.

Rachel: So, good thing you run in your own business now, Corey.

Corey: Yeah, it’s way harder for me to wind up getting surprise-fired. I see it coming [laugh]—

Rachel: [laugh].

Corey: —aways away now, and it looks like an economic industry trend.

Rachel: [sigh]. Oh, man. Well, anyhow.

Corey: Selfishly, I have to ask. You spent a lot of time working in cloud cost, to a point where I learned an awful lot from you as I was exploring the space and learning as I went. And, on some level, for me at least, it’s become an aspect of my identity, for better or worse. What was it like for you to leave and go into an orthogonal space? And sure, there’s significant overlap, but it’s a very different problem aimed at different buyers, and honestly, I think it is a more exciting problem that you are in now, from a business strategic perspective because there’s a limited amount of what you can cut off that goes up theoretically to a hundred percent of the cloud bill. But getting better observability means you can accelerate your feature velocity and that turns into something rather significant rather quickly. But what was it like?

Rachel: It’s uncomfortable, for sure. And I tend to do this to myself. I get a little bit itchy the same way I wanted to get out of storage. It’s not because there’s anything wrong with storage; I just wanted to go try something different. I tend to, I guess, do this to myself every five years ago, I make a slightly orthogonal switch in the space that I’m in.

And I think it’s because I love learning something new. The jumping into something new and having the fresh eyes is so terrifying, but it’s also really fun. And so it was really hard to leave cloud cost management. I mean, I got to Chronosphere and I was like, “Show me the cloud bill.” And I was like, “Do we have Reserved Instances?” Like, “Are we doing Committed Use Discounts with Google?”

I just needed to know. And then that helped. Okay, I got a look at the cloud bill. I felt a little better. I made a few optimizations and then I got back to my actual job which was, you know, running product marketing for Chronosphere. And I still love to jump in and just make just a little recommendation here and there. Like, “Oh, I noticed the costs are creeping up on this. Did we consider this?”

Corey: Oh, I still get a kick out of that where I was talking to an Amazonian whose side project was 110 bucks a month, and he’s like, yeah, I don’t think you could do much over here. It’s like, “Mmm, I’ll bet you a drink I can.”—

Rachel: Challenge accepted.

Corey: —it’s like, “All right. You’re on.” Cut it to 40 bucks. And he’s like, “How did you do that?” It’s because I know what I’m doing and this pattern repeats.

And it’s, are the architectural misconfigurations bounded by contacts that turn into so much. And I still maintain that I can look at the AWS bill for most environments for last month and have a pretty good idea, based upon nothing other than that, what’s going on in the environment. It turns out that maybe that’s a relatively crappy observability system when all is said and done, but it tells an awful lot. I can definitely see the appeal of wanting to get away from purely cost-driven or cost-side information and into things that give a lot more context into how things are behaving, how they’re performing. I think there’s been something of an industry rebrand away from monitoring, alerting, and trending over time to calling it observability.

And I know that people are going to have angry opinions about that—and it’s imperative that you not email me—but it all is getting down to the same thing of is my site up or down? Or in larger distributed systems, how down is it? And I still think we’re learning an awful lot. I cringe at the early days of Nagios when that was what I was depending upon to tell me whether my site was up or not. And oh, yeah, turns out that when the Nagios server goes down, you have some other problems you need to think about. It became this iterative, piling up on and piling up on and piling up on until you can get sort of good at it.

But the entire ecosystem around understanding what’s going on in your application has just exploded since the last time I was really
running production sites of any scale, in anger. So, it really would be a different world today.

Rachel: It’s changing so fast and that’s part of what makes it really exciting. And the other big thing that I love about this is, like, this is a must-have. This is not table stakes. This is not optional. Like, a great observability solution is the difference between conquering a market or being overrun.

If you look at what our founders—our founders at Chronosphere came from Uber, right? They ran the observability team at Uber. And they truly believe—and I believe them, too—that this was a competitive advantage for them. The fact that you could go to Uber and it’s always up and it’s always running and you know you’re not going to have an issue, that became an advantage to them that helped them conquer new markets. We do the same thing for our customers.

Corey: The entire idea around how these things are talked about in terms of downtime and the rest is just sort of ludicrous, on some level, because we take specific cases as industry truths. Like, I still remember, when Amazon was down one day when I was trying to buy a pair of underwear. And by that theory, it was—great, I hit a 404 page and a picture of a dog. Well, according to a lot of these industry truisms, then, well, one day a week for that entire rotation of underpants, I should have just been not wearing any. But no here in reality, I went back an hour later and bought underpants.

Now, counterpoint: If every third time I wound up trying to check out at Amazon, I wound up hitting that error page, I would spend a lot more money at Target. There is a point at which repeated downtime comes at a cost. But one-offs for some businesses are just fine. Counterpoint with if Uber is down when you’re trying to get a ride, well, that ride [unintelligible 00:28:36] may very well be lost for them and there is a definitive cost. No one’s going to go back and click on an ad as well, for example, and Amazon is increasingly an advertising company.

So, there’s a lot of nuance to it. I think we can generally say that across the board, in most cases, downtime bad. But as far as how much that is and what form that looks like and what impact that has on your company, it really becomes situationally dependent.

Rachel: I’m just going to gloss over the fact that you buy your underwear on Amazon and really not make any commentary on that. But I mean—

Corey: They sell everything there. And the problem, of course, is the crappy counterfeit underwear under the Amazon Basics brand that they ripped off from the good underwear brands. But that’s a whole ‘nother kettle of wax for a different podcast.

Rachel: Yep. Once again, not making any commentary on your—on that. Sorry, I lost my train of thought. I work in my dining room. My husband, my dog are all just—welcome to pandemic life here.

Corey: No, it’s fair. They live there. We don’t, as a general rule.

Rachel: [laugh]. Very true. Yeah. You’re not usually in my dining room, all of you but—oh, so uptime downtime, also not such a simple conversation, right? It’s not like all of Amazon is down or all of DoorDash is down. It might just be one individual service or one individual region or something that is—

Corey: One service in one subset of one availability zone. And this is the problem. People complain about the Amazon status page, but if every time something was down, it reflected there, you’d see a never ending sea of red, and that would absolutely erode confidence in the platform. Counterpoint when things are down for you and it’s not red. It’s maddening. And there’s no good answer.

Rachel: No. There’s no good answer. There’s no good answer. And the [laugh] yeah, the Amazon status page. And this is something I—bringing me back to my Forrester days, availability and resiliency in the cloud was one of the areas I focused on.

And, you know, this was once again, early days of public cloud, but remember when Netflix went down on Christmas Eve, and—God, what year was this? Maybe… 2012, and that was the worst possible time they could have had downtime because so many people are with their families watching their Doctor Who Christmas Specials, which is what I was trying to watch at the time.

Corey: Yeah, now you can’t watch it. You have to actually talk to those people, and none of us can stand them. And oh, dear Lord, yeah—

Rachel: What a nightmare.

Corey: —brutal for the family dynamic. Observability is one of those things as well that unlike you know, the AWS bill, it’s very easy to explain to people who are not deep in the space where it’s, “Oh, great. Okay. So, you have a website. It goes well. Then you want—it gets slow, so you put it on two computers. Great. Now, it puts on five computers. Now, it’s on 100 computers, half on the East Coast, half on the West Coast. Two of those computers are down. How do you tell?”

And it turns in—like, they start to understand the idea of understanding what’s going on in a complex system. “All right, how many people work at your company?” “2000,” “Great. Three laptops are broken. How do you figure out which ones are broken?” If you’re one of the people with a broken laptop, how do you figure out whether it’s your laptop or the entire system? And it lends itself really well to analogies, whereas if I’m not careful when I describe what I do, people think I can get them a better deal on underpants. No, not that kind of Amazon bill. I’m sorry.

Rachel: [laugh]. Yeah, or they started to think that you’re some kind of accountant or a tax advisor, but.

Corey: Which I prefer, as opposed to people at neighborhood block parties thinking that I’m the computer guy because then it’s, “Oh, I’m having trouble with the printer.” It’s, “Great. Have you tried [laugh] throwing away and buying a new one? That’s what I do.”

Rachel: This is a huge problem I have in my life of everyone thinking I’m going to fix all of their computer and cloud things. And I come from a big tech family. My whole family is in tech, yet somehow I’m the one at family gatherings doing, “Did you turn it off and turn it back on again?” Like, somehow that’s become my job.

Corey: People get really annoyed when you say that and even more annoyed when it fixes the problem.

Rachel: Usually does. So, the thread I wanted to pick back up on though before I got distracted by my husband and dog wandering around—at least my son is not in the room with us because he’d have a lot to say—is that the standard industry definition of observability—so once again, people are going to write to us, I’m sure; they can write to me, not you, Corey, about observability, it’s just the latest buzzword. It’s just monitoring, or you know—

Corey: It’s hipster monitoring.

Rachel: Hipster monitoring. That’s what you like to call it. I don’t really care what we call it. The important thing is it gets us through three phases, right? The first is knowing that something is wrong. If you don’t know what’s wrong, how are you supposed to ever go fix it, right? So, you need to know that those three laptops are broken.

The next thing is you need to know how bad is it? Like, if those three laptops are broken is the CEO, the COO, and the CRO, that’s real bad. If it’s three, you know, random peons in marketing, maybe not so bad. So, you need to triage, you need to understand roughly, like, the order of magnitude of it, and then you need to fix it. [laugh].

Once you fix it, you can go back and then say, all right, what was the root cause of this? How do we make sure this doesn’t happen again? So, the way you go through that cycle, you’re going to use metrics, you might use logs, you might use traces, but that’s not the definition of observability. Observability is all about getting through that, know, then triage, then fix it, then understand.

Corey: I really want to thank you for taking the time to speak with me today. If people do want to learn more, give you their unfiltered opinions, where’s the best place to find you?

Rachel: Well, you can find me on Twitter, I’m @RachelDines. You can also email me, rachel@chronosphere.io. I hope I don’t regret giving out that email address. That’s a good way you can come and argue with me about what is observability. I will not be giving advice on cloud bills. For that, you should go to Corey. But yeah, that’s a good way to get in touch.

Corey: Thank you so much for your time. I really appreciate it.

Rachel: Yeah, thank you.

Corey: Rachel Dines, Head of Product and Technical Marketing at Chronosphere. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, and castigate me with an angry comment telling me that I really should have followed the thread between the obvious link between art history and AWS billing, which is almost certainly a more disturbing Caravaggio.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Clint

Clint is the CEO and a co-founder at Cribl, a company focused on making observability viable for any organization, giving customers visibility and control over their data while maximizing value from existing tools.

Prior to co-founding Cribl, Clint spent two decades leading product management and IT operations at technology and software companies, including Splunk and Cricket Communications. As a former practitioner, he has deep expertise in network issues, database administration, and security operations.

Links:

  • Cribl: https://cribl.io/
  • Cribl.io: https://cribl.io
  • Docs.cribl.io: https://docs.cribl.io
  • Sandbox.cribl.io: https://sandbox.cribl.io

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Today’s episode is brought to you in part by our friends at MinIO the high-performance Kubernetes native object store that’s built for the multi-cloud, creating a consistent data storage layer for your public cloud instances, your private cloud instances, and even your edge instances, depending upon what the heck you’re defining those as, which depends probably on where you work. It’s getting that unified is one of the greatest challenges facing developers and architects today. It requires S3 compatibility, enterprise-grade security and resiliency, the speed to run any workload, and the footprint to run anywhere, and that’s exactly what MinIO offers. With superb read speeds in excess of 360 gigs and 100 megabyte binary that doesn’t eat all the data you’ve gotten on the system, it’s exactly what you’ve been looking for. Check it out today at min.io/download, and see for yourself. That’s min.io/download, and be sure to tell them that I sent you.

Corey: This episode is sponsored in part by our friends at Sysdig. Sysdig is the solution for securing DevOps. They have a blog post that went up recently about how an insecure AWS Lambda function could be used as a pivot point to get access into your environment. They’ve also gone deep in-depth with a bunch of other approaches to how DevOps and security are inextricably linked. To learn more, visit sysdig.com and tell them I sent you. That’s S-Y-S-D-I-G dot com. My thanks to them for their continued support of this ridiculous nonsense.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I have a repeat guest joining me on this promoted episode. Clint Sharp is the CEO and co-founder of Cribl. Clint, thanks for joining me.

Clint: Hey, Corey, nice to be back.

Corey: I was super excited when you gave me the premise for this recording because you said you had some news to talk about, and I was really excited that oh, great, they’re finally going to buy a vowel so that people look at their name and understand how to pronounce it. And no, that’s nowhere near forward-looking enough. It’s instead it’s some, I guess, I don’t know, some product announcement or something. But you know, hope springs eternal. What have you got for us today?

Clint: Well, one of the reasons I love talking to your audiences because product announcements actually matter to this audience. It’s super interesting, as you get into starting a company, you’re such, like, a product person, you’re like, “Oh, I have this new set of things that’s really going to make your life better.” And then you go out to, like, the general media, and you’re like, “Hey, I have this product.” And they’re like, “I don’t care. What product? Do you have a funding announcement? Do you have something big in the market that—you know, do you have a new executive? Do you”—it’s like, “No, but, like, these features, like these things, that we—the way we make our lives better for our customers. Isn’t that interesting?” “No.”

Corey: Real depressing once you—“Do you have a security breach to announce?” It’s, “No. God no. Why would I wind up being that excited about it?” “Well, I don’t know. I’d be that excited about it.” And yeah, the stuff that mainstream media wants to write about in the context of tech companies is exactly the sort of thing that tech companies absolutely do not want to be written about for. But fortunately, that is neither here nor there.

Clint: Yeah, they want the thing that gets the clicks.

Corey: Exactly. You built a product that absolutely resonates in its target market and outside of that market. It’s one of those, what is that thing, again? If you could give us a light refresher on what Cribl is and does, you’ll probably do a better job of it than I will. We hope.

Clint: We’d love to. Yeah, so we are an observability company, fundamentally. I think one of the interesting things to talk about when it comes to observability is that observability and security are merging. And so I like to say observability and include security people. If you’re a security person, and you don’t feel included by the word observability, sorry.

We also include you; you’re under our tent here. So, we sell to technology professionals, we help make their lives better. And we do that today through a flagship product called LogStream—which is part of this announcement, we’re actually renaming to Stream. In some ways, we’re dropping logs—and we are a pipeline company. So, we help you take all of your existing agents, all of your existing data that’s moving, and we help you process that data in the stream to control costs and to send it multiple places.

And it sounds kind of silly, but one of the biggest problems that we end up solving for a lot of our enterprises is, “Hey, I’ve got, like, this old Syslog feed coming off of my firewalls”—like, you remember those things, right? Palo Alto firewalls, ASA firewalls—“I actually get that thing to multiple places because, hey, I want to get that data into another security solution. I want to get that data into a data lake. How do I do that?” Well, in today’s world, that actually turns out is sort of a neglected set of features, like, the vendors who provide you logging solutions, being able to reshape that data, filter that data, control costs, wasn’t necessarily at the top of their priority list.

It wasn’t nefarious. It wasn’t like people are like, “Oh, I’m going to make sure that they can’t process this data before it comes into my solution.” It’s more just, like, “I’ll get around to it eventually.” And the eventually never actually comes. And so our streaming product helps people do that today.

And the big announcement that we’re making this week is that we’re extending that same processing technology down to the endpoint with a new product we’re calling Cribl Edge. And so we’re taking our existing best-in-class management technology, and we’re turning it into an agent. And that seems kind of interesting because… I think everybody sort of assumed that the agent is dead. Okay, well, we’ve been building agents for a decade or two decades. Isn’t everything exactly the same as it was before?

But we really saw kind of a dearth of innovation in that area in terms of being able to manage your agents, being able to understand what data is available to be collected, being able to auto-discover the data that needs to be able to be collected, turning those agents into interactive troubleshooting experiences so that we can, kind of, replicate the ability to zoom into a remote endpoint and replicate that Linux command line experience that we’re not supposed to be getting anymore because we’re not supposed to SSH into boxes anymore. Well, how do I replicate that? How do I see how much disk is on this given endpoint if I can’t SSH into that box? And so Cribl Edge is a rethink about making this rich, interactive experience on top of all of these agents that become this really massive distributed system that we can process data all the way out at where the data is being emitted.

And so that means that now we don’t nec—if you want to process that data in the stream, okay, great, but if you want to process that data at its origination point, we can actually provide you cheaper cost because now you’re using a lot of that capacity that’s sitting out there on your endpoints that isn’t really being used today anyway—the average utilization of a Kubernetes cluster is like 30%—

Corey: It’s that high. I’m sort of surprised.

Clint: Right? I know. So, Datadog puts out the survey every year, which I think is really interesting, and that’s a number that always surprised me is just that people are already paying for this capacity, right? It’s sitting there, it’s on their AWS bill already, and with that average utilization, a lot of the stuff that we’re doing in other clusters, or while we’re moving that data can actually just be done right there where the data is being emitted. And also, if we’re doing things like filtering, we can lower egress charges, there’s lots of really, really good goodness that we can do by pushing that processing further closer to its origination point.

Corey: You know, the timing of this episode is somewhat apt because as of the time that we’re recording this, I spent most of yesterday troubleshooting and fixing my home wireless network, which is a whole Ubiquity-managed thing. And the controller was one of their all-in-one box things that kept more or less power cycling for no apparent reason. How do I figure out why it’s doing that? Well, I’m used to, these days, doing everything in a cloud environment where you can instrument things pretty easily, where things start and where things stop is well understood. Finally, I just gave up and used a controller that’s sitting on an EC2 instance somewhere, and now great, now I can get useful telemetry out of it because now it’s stuff I know how to deal with.

It also, turns out that surprise, my EC2 instance is not magically restarting itself due to heat issues. What a concept. So, I have a newfound appreciation for the fact that oh, yeah, not everything lives in a cloud provider’s regions. Who knew? This is a revelation that I think is going to be somewhat surprising for folks who’ve been building startups and believe that anything that’s older than 18 months doesn’t exist.

But there’s a lot of data centers out there, there are a lot of agents living all kinds of different places. And workloads continue to surprise me even now, just looking at my own client base. It’s a very diverse world when we’re talking about whether things are on-prem or whether they’re in cloud environments.

Clint: Well, also, there’s a lot of agents on every endpoint period, just due to the fact that security guys want an agent, the observability guys want an agent, the logging people want an agent. And then suddenly, I’m, you know, I’m looking at every endpoint—cloud, on-prem, whatever—and there’s 8, 10 agents sitting there. And so I think a lot of the opportunity that we saw was, we can unify the data collection for metric type of data. So, we have some really cool defaults. [unintelligible 00:07:30] this is one of the things where I think people don’t focus much on, kind of, the end-user experience. Like, let’s have reasonable defaults.

Let’s have the thing turn on, and actually, most people’s needs are set without tweaking any knobs or buttons, and no diving into YAML files and looking at documentation and trying to figure out exactly the way I need to configure this thing. Let’s collect metric data, let’s collect log data, let’s do it all from one central place with one agent that can send that data to multiple places. And I can send it to Grafana Cloud, if I want to; I can send it to Logz.io, I can send it to Splunk, I can send it to Elasticsearch, I can send it to AWS’s new Elasticsearch-y the thing that we don’t know what they’re going to call it yet after the lawsuit. Any of those can be done right from the endpoint from, like, a rich graphical experience where I think that there’s a really a desire now for people to kind of jump into these configuration files where really a lot of these users, this is a part-time job, and so hey, if I need to go set up data collection, do I want to learn about this detailed YAML file configuration that I’m only going to do once or twice, or should I be able to do it in an easy, intuitive way, where I can just sit down in front of the product, get my job done and move on without having to go learn some sort of new configuration language?

Corey: Once upon a time, I saw an early circa 2012, 2013 talk from Jordan Sissel, who is the creator of Logstash, and he talked a lot about how challenging it was to wind up parsing all of the variety of log files out there. Even something is relatively straightforward—wink, wink, nudge, nudge—as timestamps was an absolute monstrosity. And a lot of people have been talking in recent years about OpenTelemetry being the lingua franca that everything speaks so that is the wave of the future, but I’ve got a level with you, looking around, it feels like these people are living in a very different reality than the one that I appear to have stumbled into because the conversations people are having about how great it is sound amazing, but nothing that I’m looking at—granted from a very particular point of view—seems to be embracing it or supporting it. Is that just because I’m hanging out in the wrong places, or is it still a great idea whose time has yet to come, or something else?

Clint: So, I think a couple things. One is every conversation I have about OpenTelemetry is always, “Will be.” It’s always in the future. And there’s certainly a lot of interest. We see this from customer after customer, they’re very interested in OpenTelemetry and what the OpenTelemetry strategy is, but as an example OpenTelemetry logging is not yet finalized specification; they believe that they’re still six months to a year out. It seems to be perpetually six months to a year out there.

They are finalized for metrics and they are finalized for tracing. Where we see OpenTelemetry tends to be with companies like Honeycomb, companies like Datadog with their tracing product, or Lightstep. So, for tracing, we see OpenTelemetry adoption. But tracing adoption is also not that high either, relative to just general metrics of logs.

Corey: Yeah, the tracing implementations that I’ve seen, for example, Epsagon did this super well, where it would take a look at your Lambdas Function built into an application, and ah, we’re going to go ahead and instrument this automatically using layers or extensions for you. And life was good because suddenly you got very detailed breakdowns of exactly how data was flowing in the course of a transaction through 15 Lambdas Function. Great. With everything else I’ve seen, it’s, “Oh, you have to instrument all these things by hand.” Let me shortcut that for you: That means no one’s going to do it. They never are.

It’s anytime you have to do that undifferentiated heavy lifting of making sure that you put the finicky code just so into your application’s logic, it’s a shorthand for it’s only going to happen when you have no other choice. And I think that trying to surface that burden to the developer, instead of building it into the platform so they don’t have to think about it is inherently the wrong move.

Clint: I think there’s a strong belief in Silicon Valley that—similar to, like, Hollywood—that the biggest export Silicon Valley is going to have is culture. And so that’s going to be this culture of, like, developer supporting their stuff in production. I’m telling you, I sell to banks and governments and telcos and I don’t see that culture prevailing. I see a application developed by Accenture that’s operated by Tata. That’s a lot of inertia to overcome and a lot of regulation to overcome as well, and so, like, we can say that, hey, separation of duties isn’t really a thing and developers should be able to support all their own stuff in production.

I don’t see that happening. It may happen. It’ll certainly happen more than zero. And tracing is predicated on the whole idea that the developer is scratching their own itch. Like that I am in production and troubleshooting this and so I need this high-fidelity trace-level information to understand what’s going on with this one user’s experience, but that doesn’t tend to be in the enterprise, how things are actually troubleshot.

And so I think that more than anything is the headwind that slowing down distributed tracing adoption. It’s because you’re putting the onus on solving the problem on a developer who never ends up using the distributed tracing solution to begin with because there’s another operations department over there that’s actually operating the thing on a day-to-day basis.

Corey: Having come from one of those operations departments myself, the way that I would always fix things was—you know, in the era that I was operating it made sense—you’d SSH into a box and kick the tires, poke around, see what’s going on, look at the logs locally, look at the behaviors, the way you’d expect it to these days, that is considered a screamingly bad anti-pattern and it’s something that companies try their damnedest to avoid doing at all. When did that change? And what is the replacement for that? Because every time I asked people for the sorts of data that I would get from that sort of exploration when they’re trying to track something down, I’m more or less met with blank stares.

Clint: Yeah. Well, I think that’s a huge hole and one of the things that we’re actually trying to do with our new product. And I think the… how do I replicate that Linux command line experience? So, for example, something as simple, like, we’d like to think that these nodes are all ephemeral, but there’s still a disk, whether it’s virtual or not; that thing sometimes fills up, so how do I even do the simple thing like df -kh and see how much disk is there if I don’t already have all the metrics collected that I needed, or I need to go dive deep into an application and understand what that application is doing or seeing, what files it’s opening, or what log files it’s writing even?

Let’s give some good examples. Like, how do I even know what files an application is running? Actually, all that information is all there; we can go discover that. And so some of the things that we’re doing with Edge is trying to make this rich, interactive experience where you can actually teleport into the end node and see all the processes that are running and get a view that looks like top and be able to see how much disk is there and how much disk is being consumed. And really kind of replicating that whole troubleshooting experience that we used to get from the Linux command line, but now instead, it’s a tightly controlled experience where you’re not actually getting an arbitrary shell, where I could do anything that could give me root level access, or exploit holes in various pieces of software, but really trying to replicate getting you that high fidelity information because you don’t need any of that information until you need it.

And I think that’s part of the problem that’s hard with shipping all this data to some centralized platform and getting every metric and every log and moving all that data is the data is worthless until it isn’t worthless anymore. And so why do we even move it? Why don’t we provide a better experience for getting at the data at the time that we need to be able to get at the data. Or the other thing that we get to change fundamentally is if we have the edge available to us, we have way more capacity. I can store a lot of information in a few kilobytes of RAM on every node, but if I bring thousands of nodes into one central place, now I need a massive amount of RAM and a massive amount of cardinality when really what I need is the ability to actually go interrogate what’s running out there.

Corey: The thing that frustrates me the most is the way that I go back and find my old debug statements, which is, you know, I print out whatever it is that the current status is and so I can figure out where something’s breaking.

Clint: [Got here 00:15:08].

Corey: Yeah. I do it within AWS Lambda functions, and that’s great. And I go back and I remove them later when I notice how expensive CloudWatch logs are getting because at 50 cents per gigabyte of ingest on those things, and you have that Lambda function firing off a fair bit, that starts to add up when you’ve been excessively wordy with your print statements. It sounds ridiculous, but okay, then you’re storing it somewhere. If I want to take that log data and have something else consume it, that’s nine cents a gigabyte to get it out of AWS and then you’re going to want to move it again from wherever it is over there—potentially to a third system, because why not?—and it seems like the entire purpose of this log data is to sit there and be moved around because every time it gets moved, it winds up somehow costing me yet more money. Why do we do this?

Clint: I mean, it’s a great question because one of the things that I think we decided 15 years ago was that the reason to move this data was because that data may go poof. So, it was on a, you know, back in my day, it was an HP DL360 1U rackmount server that I threw in there, and it had raid zero discs and so if that thing went dead, well, we didn’t care, we’d replace it with another one. But if we wanted to find out why it went dead, we wanted to make sure that the data had moved before the thing went dead. But now that DL360 is a VM.

Corey: Yeah, or a container that is going to be gone in 20 minutes. So yeah, you don’t want to store it locally on that container. But discs are also a fair bit more durable than they once were, as well. And S3 talks about its 11 nines of durability. That’s great and all but most of my application logs don’t need that. So, I’m still trying to figure out where we went wrong.

Clint: Well, I think it was right for the time. And I think now that we have durable storage at the edge where that blob storage has already replicated three times and we can reattach—if that box crashes, we can reattach new compute to that same block storage. Actually, AWS has some cool features now, you can actually attach multiple VMs to the same block store. So, we could actually even have logs being written by one VM, but processed by another VM. And so there are new primitives available to us in the cloud, which we should be going back and re-questioning all of the things that we did ten to 15 years ago and all the practices that we had because they may not be relevant anymore, but we just never stopped to ask why.

Corey: Yeah, multi-attach was rolled out with their IO2 volumes, which are spendy but great. And they do warn you that you need a file system that actively supports that and applications that are aware of it. But cool, they have specific use cases that they’re clearly imagining this for. But ten years ago, we were building things out, and, “Ooh, EBS, how do I wind up attaching that from multiple instances?” The answer was, “Ohh, don’t do that.”

And that shaped all of our perspectives on these things. Now suddenly, you can. Is that, “Ohh don’t do that,” gut visceral reaction still valid? People don’t tend to go back and re-examine the why behind certain best practices until long after those best practices are now actively harmful.

Clint: And that’s really what we’re trying to do is to say, hey, should we move log data anymore if it’s at a durable place at the edge? Should we move metric data at all? Like, hey, we have these big TSDBs that have huge cardinality challenges, but if I just had all that information sitting in RAM at the original endpoint, I can store a lot of information and barely even touch the free RAM that’s already sitting out there at that endpoint. So, how to get out that data? Like, how to make that a rich user experience so that we can query it?

We have to build some software to do this, but we can start to question from first principles, hey, things are different now. Maybe we can actually revisit a lot of these architectural assumptions, drive cost down, give more capability than we actually had before for fundamentally cheaper. And that’s kind of what Cribl does is we’re looking at software is to say, “Man, like, let’s question everything and let’s go back to first principles.” “Why do we want this information?” “Well, I need to troubleshoot stuff.” “Okay, well, if I need to troubleshoot stuff, well, how do I do that?” “Well, today we move it, but do we have to? Do we have to move that data?” “No, we could probably give you an experience where you can dive right into that endpoint and get really, really high fidelity data without having to pay to move that and store it forever.” Because also, like, telemetry information, it’s basically worthless after 24 hours, like, if I’m moving that and paying to store it, then now I’m paying for something I’m never going to read back.

Corey: This episode is sponsored in part by our friends at Vultr. Spelled V-U-L-T-R because they’re all about helping save money, including on things like, you know, vowels. So, what they do is they are a cloud provider that provides surprisingly high performance cloud compute at a price that—while sure they claim its better than AWS pricing—and when they say that they mean it is less money. Sure, I don’t dispute that but what I find interesting is that it’s predictable. They tell you in advance on a monthly basis what it’s going to going to cost. They have a bunch of advanced networking features. They have nineteen global locations and scale things elastically. Not to be confused with openly, because apparently elastic and open can mean the same thing sometimes. They have had over a million users. Deployments take less that sixty seconds across twelve pre-selected operating systems. Or, if you’re one of those nutters like me, you can bring your own ISO and install basically any operating system you want. Starting with pricing as low as $2.50 a month for Vultr cloud compute they have plans for developers and businesses of all sizes, except maybe Amazon, who stubbornly insists on having something to scale all on their own. Try Vultr today for free by visiting: vultr.com/screaming, and you’ll receive a $100 in credit. Thats V-U-L-T-R.com slash screaming.

Corey: And worse, you wind up figuring out, okay, I’m going to store all that data going back to 2012, and it’s petabytes upon petabytes. And great, how do I actually search for a thing? Well, I have to use some other expensive thing of compute that’s going to start diving through all of that because the way I set up my partitioning, it isn’t aligned with anything looking at, like, recency or based upon time period, so right every time I want to look at what happened 20 minutes ago, I’m looking at what happened 20 years ago. And that just gets incredibly expensive, not just to maintain but to query and the rest. Now, to be clear, yes, this is an anti-pattern. It isn’t how things should be set up. But how should they be set up? And it is the collective the answer to that right now actually what’s best, or is it still harkening back to old patterns that no longer apply?

Clint: Well, the future is here, it’s just unevenly distributed. So there’s, you know, I think an important point about us or how we think about building software is with this customer is first attitude and fundamentally bringing them choice. Because the reality is that doing things the old way may be the right decision for you. You may have compliance requirements to say—there’s a lot of financial services institutions, for example, like, they have to keep every byte of data written on any endpoint for seven years. And so we have to accommodate their requirements.

Like, is that the right requirement? Well, I don’t know. The regulator wrote it that way, so therefore, I have to do it. Whether it’s the right thing or the wrong thing for the business, I have no choice. And their decisions are just as right as the person who says this data is worthless and should all just be thrown away.

We really want to be able to go and say, like, hey, what decision is right? We’re going to give you the option to do it this way, we’re going to give you the option to do it this way. Now, the hard part—and that when it comes down to, like, marketing, it’s like you want to have this really simple message, like, “This is the one true path.” And a lot of vendors are this way, “There’s this new wonderful, right, true path that we are going to take you on, and follow along behind me.” But the reality is, enterprise worlds are gritty and ugly, and they’re full of old technology and new technology.

And they need to be able to support getting data off the mainframe the same way as they’re doing a brand new containerized microservices application. In fact, that brand new containerized microservices application is probably talking to the mainframe through some API. And so all of that has to work at once.

Corey: Oh, yeah. And it’s all of our payment data is in our PCI environment that PCI needs to have every byte logged. Great. Why is three-quarters of your infrastructure considered the PCI environment? Maybe you can constrain that at some point and suddenly save a whole bunch of effort, time, money, and regulatory drag on this.

But as you go through that journey, you need to not only have a tool that will work when you get there but a tool that will work where you are today. And a lot of companies miss that mark, too. It’s, “Oh, once you modernize and become the serverless success story of the decade, then our product is going to be right for you.” “Great. We’ll send you a postcard if we ever get there and then you can follow up with us.”

Alternately, it’s well, “Yeah, we’re this is how we are today, but we have a visions of a brighter tomorrow.” You’ve got to be able to meet people where they are at any point of that journey. One of the things I’ve always respected about Cribl has been the way that you very fluidly tell both sides of that story.

Clint: And it’s not their fault.

Corey: Yeah.

Clint: Most of the people who pick a job, they pick the job because, like—look, I live in Kansas City, Missouri, and there’s this data processing company that works primarily on mainframes, it’s right down the road. And they gave me a job and it pays me $150,000 a year, and I got a big house and things are great. And I’m a sysadmin sitting there. I don’t get to play with the new technology. Like, that customer is just as an applicable customer, we want to help them exactly the same as the new Silicon Valley hip kid who’s working at you know, a venture-backed startup, they’re doing everything natively in the cloud. Those are all right decisions, depending on where you happen to find yourself, and we want to support you with our products, no matter where you find yourself on the technology spectrum.

Corey: Speaking of old and new, and the trends of the industry, when you first set up this recording, you mentioned, “Oh, yeah, we should make it a point to maybe talk about the acquisition,” at which point I sprayed coffee across my iMac. Thanks for that. Turns out it wasn’t your acquisition we were talking about so much as it is the—at the time we record this—-the yet-to-close rumored acquisition of Splunk by Cisco.

Clint: I think it’s both interesting and positive for some people, and sad for others. I think Cisco is obviously a phenomenal company. They run the networking world. The fact that they’ve been moving into observability—they bought companies like AppDynamics, and we were talking about Epsagon before the show, they bought—ServiceNow, just bought Lightstep recently. There’s a lot of acquisitions in this space.

I think that when it comes to something like Splunk, Splunk is a fast-growing company by compared to Cisco. And so for them, this is something that they think that they can put into their distribution channel, and what Cisco knows how to do is to sell things like they’re very good at putting things through their existing sales force and really amplifying the sales of that particular thing that they have just acquired. That being said, I think for a company that was as innovative as Splunk, I do find it a bit sad with the idea that it’s going to become part of this much larger behemoth and not really probably driving the observability and security industry forward anymore because I don’t think anybody really looks at Cisco as a company that’s driving things—not to slam them or anything, but I don’t really see them as driving the industry forward.

Corey: Somewhere along the way, they got stuck and I don’t know how to reconcile that because they were a phenomenally fast-paced innovative company, briefly the most valuable company in the world during the dotcom bubble. And then they just sort of stalled out somewhere and, on some level, not to talk smack about it, but it feels like the level of innovation we’ve seen from Splunk has curtailed over the past half-decade or so. And selling to Cisco feels almost like a tacit admission that they are effectively out of ideas. And maybe that’s unfair.

Clint: I mean, we can look at the track record of what’s been shipped over the last five years from Splunk. And again they’re a partner, their customers are great, I think they still have the best log indexing engine on the market. That was their core product and what has made them the majority of their money. But there’s not been a lot new. And I think objectively we can look at that without throwing stones and say like, “Well, what net-new? You bought SignalFX. Like, good for you guys like that seems to be going well. You’ve launched your observability suite based off of these acquisitions.” But organic product-wise, there’s not a lot coming out of the factory.

Corey: I’ll take it a bit further-slash-sadder, we take a look at some great companies that were acquired—OpenDNS, Duo Security, SignalFX, as you mentioned, Epsagon, ThousandEyes—and once they’ve gotten acquired by Cisco, they all more or less seem to be frozen in time, like they’re trapped in amber, which leads us up to the natural dinosaur analogy that I’ll probably make in a less formal setting. It just feels like once a company is bought by Cisco, their velocity peters out, a lot of their staff leaves, and what you see is what you get. And I don’t know if that’s accurate, I’m just not looking in the right places, but every time I talk to folks in the industry about this, I get a lot of knowing nods that are tied to it. So, whether or not that’s true or not, that is very clearly, at least in some corners of the market, the active perception.

Clint: There’s a very real fact that if you look even at very large companies, innovation is driven from a core set of a handful of people. And when those people start to leave, the innovation really stops. It’s those people who think about things back from first principles—like why are we doing things? What different can we do?—and they’re the type of drivers that drive change.

So, Frank Slootman wrote a book recently called Amp it Up that I’ve been reading over the last weekend, and he talks—has this article that was on LinkedIn a while back called “Drivers vs. Passengers” and he’s always looking for drivers. And those drivers tend to not find themselves as happy in bigger companies and they tend to head for the exits. And so then you end up with the people who are a lot of the passenger type of people, the people who are like—they’ll carry it forward, they’ll continue to scale it, the business will continue to grow at whatever rate it’s going to grow, but you’re probably not going to see a lot of the net-new stuff. And I’ll put it in comparison to a company like Datadog who I have a vast amount of respect for I think they’re incredibly innovative company, and I think they continue to innovate.

Still driven by the founders, the people who created the original product are still there driving the vision, driving forward innovation. And that’s what tends to move the envelope is the people who have the moral authority inside of an even larger organization to say, “Get behind me. We’re going in this direction. We’re going to go take that hill. We’re going to go make things better for our customers.” And when you start to lose those handful of really critical contributors, that’s where you start to see the innovation dry up.

Corey: Where do you see the acquisitions coming from? Is it just at some point people shove money at these companies that got acquired that is beyond the wildest dreams of avarice? Is it that they believe that they’ll be able to execute better on their mission and they were independently? These are still smart, driven, people who have built something and I don’t know that they necessarily see an acquisition as, “Well, time to give up and coast for a while and then I’ll leave.” But maybe it is. I’ve never found myself in that situation, so I can’t speak for sure.

Clint: You kind of I think, have to look at the business and then whoever’s running the business at that time—and I sit in the CEO chair—so you have to look at the business and say, “What do we have inside the house here?” Like, “What more can we do?” If we think that there’s the next billion-dollar, multi-billion-dollar product sitting here, even just in our heads, but maybe in the factory and being worked on, then we should absolutely not sell because the value is still there and we’re going to grow the company much faster as an independent entity than we would you know, inside of a larger organization. But if you’re the board of directors and you’re looking around and saying like, hey look, like, I don’t see another billion-dollar line of bus—at this scale, right, if your Splunk scale, right? I don’t see another billion-dollar line of business sitting here, we could probably go acquire it, we could try to add it in, but you know, in the case of something like a Splunk, I think part of—you know, they’re looking for a new CEO right now, so now they have to go find a new leader who’s going to come in, re-energize and, kind of, reboot that.

But that’s the options that they’re considering, right? They’re like, “Do I find a new CEO who’s going to reinvigorate things and be able to attract the type of talent that’s going to lead us to the next billion-dollar line of business that we can either build inside or we can acquire and bring in-house? Or is the right path for me just to say, ‘Okay, well, you know, somebody like Cisco’s interested?’” or the other path that you may see them go down to something like Silver Lake, so Silver Lake put a billion dollars into the company last year. And so they may be looking at and say, “Okay, well, we really need to do some restructuring here and we want to do it outside the eyes of the public market. We want to be able to change pricing model, we want to be able to really do this without having to worry about the stock price’s massive volatility because we’re making big changes.”

And so I would say there’s probably two big options there considering. Like, do we sell to Cisco, do we sell to Silver Lake, or do we really take another run at this? And those are difficult decisions for the stewards of the business and I think it’s a different decision if you’re the steward of the business that created the business versus the steward of the business for whom this is—the I’ve been here for five years and I may be here for five years more. For somebody like me, a company like Cribl is literally the thing I plan to leave on this earth.

Corey: Yeah. Do you have that sense of personal attachment to it? On some level, The Duckbill Group, that’s exactly what I’m staring at where it’s great. Someone wants to buy the Last Week in AWS media side of the house.

Great. Okay. What is that really, beyond me? Because so much of it’s been shaped by my personality. There's an audience, sure, but it’s a skeptical audience, one that doesn’t generally tend to respond well to mass market, generic advertisements, so monetizing that is not going to go super well.

“All right, we’re going to start doing data mining on people.” Well, that’s explicitly against the terms of service people signed up for, so good luck with that. So, much starts becoming bizarre and strange when you start looking at building something with the idea of, oh, in three years, I’m going to unload this puppy and make it someone else’s problem. The argument is that by building something with an eye toward selling it, you build a better-structured business, but it also means you potentially make trade-offs that are best not made. I’m not sure there’s a right answer here.

Clint: In my spare time, I do some investments, angel investments, and that sort of thing, and that’s always a red flag for me when I meet a founder who’s like, “In three to five years, I plan to sell it to these people.” If you don’t have a vision for how you’re fundamentally going to alter the marketplace and our perception of everything else, you’re not dreaming big enough. And that to me doesn’t look like a great investment. It doesn’t look like the—how do you attract employees in that way? Like, “Okay, our goal is to work really hard for the next three years so that we will be attractive to this other bigger thing.” They may be thinking it on the inside as an available option, but if you think that’s your default option when starting a company, I don’t think you’re going to end up with the outcome is truly what you’re hoping for.

Corey: Oh, yeah. In my case, the only acquisition story I see is some large company buying us just largely to shut me up. But—

Clint: [laugh].

Corey: —that turns out to be kind of expensive, so all right. I also don’t think it serve any of them nearly as well as they think it would.

Clint: Well, you’ll just become somebody else on Twitter. [laugh].

Corey: Yeah, “Time to change my name again. Here we go.” So, if people want to go and learn more about a Cribl Edge, where can they do that?

Clint: Yeah, cribl.io. And then if you’re more of a technical person, and you’d like to understand the specifics, docs.cribl.io. That’s where I always go when I’m checking out a vendor; just skip past the main page and go straight to the docs. So, check that out.

And then also, if you’re wanting to play with the product, we make online available education called Sandboxes, at sandbox.cribl.io, where you can go spin up your own version of the product, walk through some interactive tutorials, and get a view on how it might work for you.

Corey: Such a great pattern, at least for the way that I think about these things. You can have flashy videos, you can have great screenshots, you can have documentation that is the finest thing on this earth, but let me play with it; let me kick the tires on it, even with a sample data set. Because until I can do that, I’m not really going to understand where the product starts and where it stops. That is the right answer from where I sit. Again, I understand that everyone’s different, not everyone thinks like I do—thankfully—but for me, that’s the best way I’ve ever learned something.

Clint: I love to get my hands on the product, and in fact, I’m always a little bit suspicious of any company when I go to their webpage and I can’t either sign up for the product or I can’t get to the documentation, and I have to talk to somebody in order to learn. That’s pretty much I’m immediately going to the next person in that market to go look for somebody who will let me.

Corey: [laugh]. Thank you again for taking so much time to speak with me. I appreciate it. As always, it’s a pleasure.

Clint: Thanks, Corey. Always enjoy talking to you.

Corey: Clint Sharp, CEO and co-founder of Cribl. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry comment. And when you hit submit, be sure to follow it up with exactly how many distinct and disparate logging systems that obnoxious comment had to pass through on your end of things.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Alex

Alex holds a Ph.D. in Computer Science and Engineering from UC San Diego, and has spent over a decade building high-performance, robust data management and processing systems. As an early member of a couple fast-growing startups, he’s had the opportunity to wear a lot of different hats, serving at various times as an individual contributor, tech lead, manager, and executive. Prior to joining the Duckbill Group, Alex spent a few years as a freelance data engineering consultant, helping his clients build, manage and maintain their data infrastructure. He lives in Los Angeles, CA.

Links:

  • Twitter: https://twitter.com/alexras/
  • Personal page: https://alexras.info
  • Old Consulting website with blog: https://bitsondisk.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: The company 0x4447 builds products to increase standardization and security in AWS organizations. They do this with automated pipelines that use well-structured projects to create secure, easy-to-maintain and fail-tolerant solutions, one of which is their VPN product built on top of the popular OpenVPN project which has no license restrictions; you are only limited by the network card in the instance. To learn more visit: snark.cloud/deployandgo

Corey: Today’s episode is brought to you in part by our friends at MinIO the high-performance Kubernetes native object store that’s built for the multi-cloud, creating a consistent data storage layer for your public cloud instances, your private cloud instances, and even your edge instances, depending upon what the heck you’re defining those as, which depends probably on where you work. It’s getting that unified is one of the greatest challenges facing developers and architects today. It requires S3 compatibility, enterprise-grade security and resiliency, the speed to run any workload, and the footprint to run anywhere, and that’s exactly what MinIO offers. With superb read speeds in excess of 360 gigs and 100 megabyte binary that doesn’t eat all the data you’ve gotten on the system, it’s exactly what you’ve been looking for. Check it out today at min.io/download, and see for yourself. That’s min.io/download, and be sure to tell them that I sent you.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’m the chief cloud economist at The Duckbill Group, which people are generally aware of. Today, I’m joined by our most recent principal cloud economist, Alex Rasmussen. Alex, thank you for joining me today, it is a pleasure to talk to you, as if we aren’t talking to each other constantly, now that you work here.

Alex: Thanks, Corey. It’s great being here.

Corey: So, I followed a more, I’d say traditional path for a cloud economist, but given that I basically had to invent the job myself, the more common path because imagine that you start building a role from scratch and the people you wind up looking for initially look a lot like you. And that is grumpy sysadmin, historically, turned into something, kind of begrudgingly, that looks like an SRE, which I still maintain are the same thing, but it is imperative people not email me about that. Yes, I know, you work at Google. But instead, what I found during my tenure as a sysadmin, is that I was working with certain things an awful lot, like web servers, and other things almost never, like databases and data warehouses. Because if you screw up a web server, we all have a good laugh, the site’s down for a couple of minutes, life goes on, you have a shame trophy on your desk if that’s your corporate culture, things continue.

Mess up the data severely enough, and you don’t have a company anymore. So, I was always told to keep my aura away from the expensive spendy things that power a company. You are sort of the first of a cloud economist subtype that doesn’t resemble that. Before you worked here, you were effectively an independent consultant working on data engineering. Before that, you had a couple of jobs, but you had gotten a PhD in computer science, which means, first, you are probably one of the people in this world most qualified to pass some crappy job interview of solving a sorting algorithm on a whiteboard, but how did you get here from where you were?

Alex: Great question. So, I like to joke that I kind of went to school until somebody told me that I had to stop. And I took that and went and started—or didn’t start, but I was an early engineer at a startup and then was an executive at another early-stage one, and did a little bit of everything. And went freelance, did that for a couple of years, and worked with all kinds of different companies—vast majority of those being startups—helping them with data infrastructure problems. I’ve done a little bit of everything throughout my career.

I’ve been, you know, IC, manager, manager, manager, IT guy, everything in between. I think on the data side of things, it just sort of happened, to be honest with you, it kind of started with the stuff that I did for my dissertation and parlayed that into a job back when the big data wave was starting to kind of truly crest. And I’ve been working on data infrastructure, basically my entire career. So, it wasn’t necessarily something that was intentional. I’ve just been kind of taking the opportunity that makes the most sense for me it kind of every juncture. And my career path has been a little bit strange, both by academic and industrial standards. But I like where I’m at and I gained something really valuable from each of those experiences. So.

Corey: It’s been an interesting area of I won’t say weakness here, but it’s definitely been a bit of a challenge when we look at an AWS environment and even talking about a typical AWS customer without thinking of any of them in particular, I can already tell you a few things are likely to be true. For example, the number one most expensive line item in their bill is going to be EC2, and compute is the thing that powers it. Now, maybe that is they’re running a bunch of instances the old-fashioned way. Maybe they’re running Kubernetes but that’s how it shows up. There’s a lot of things that could be, and we look at what rounds that out.

Now, the next item down should almost certainly not be data transfer and if so we should have a conversation, but data in one form or another is very often going to be number two. And that can mean a bunch of different things, historically. It could mean, “Oh, you have a whole bunch of stuff in S3. Let’s talk about access patterns. Let’s talk about lifecycle policies. Let’s talk about making sure the really important stuff is backed up somewhere. Maybe you want to spend more on that particular aspect of it.”

If it’s on EBS volumes, that’s interesting and definitely worth looking into and trying to understand the context of what’s going on. Periodically we’ll see a whole bunch of additional charges that speak to some of that EC2 charge in the form of EMR, AWS’s Elastic MapReduce, which charges a per-hour instance charge, but also charges you for the instances that are running under the hood and under the EC2 line item. So, there’s a lot of data lifecycle stuff, there’s a lot of data ecosystem stories, that historically we’ve consulted out with experts in that particular space. And that’s great, but we were starting to have to drag those people in on more and more engagements as we saw them. And we realized that was really something we had to build out as a core competency for ourselves.

And we started out not intending to hire for someone with that specialty, but the more we talked to you, the more it became clear that this was a very real and very growing need that we and our customers have. How closely it is what you’re doing now as far as AWS bill analysis and data pattern deep-dive align with what you were doing as a freelance consultant in the space?

Alex: A lot more than you might expect. You know, I think that increasingly, what you’re seeing now is that a company’s core differentiator is its data, right, how much of it they have, what they do with it. And so, you know, to your point, I think when you look at any company’s cloud spend, it’s going to be pretty heavy on the data side in terms of, like, where have you put it? What are you doing to process it? Where is it going once it’s been processed? And then how is that—

Corey: And data transfer is a very important first word in that two-word sequence.

Alex: Oh, sure is. And so I think that, like, in a lot of ways, the way that a customer’s cloud architecture looks and the way that their bill looks kind of as a consequence of that is kind of a reification in a way of the way that the data flows from one place to another and what’s done with it at each step along the way. I think what complicates this is that companies that have been around for a little while have lived through this kind of very amorphous, kind of, polyglot way that we’re approaching data. You know, back when I was first getting started in the big data days, it was MapReduce, MapReduce, MapReduce, right? And we quickly [crosstalk 00:07:29]—

Corey: Oh, yes. The MapReduce white paper out of Google, a beautiful April Fool’s Day prank that the folks at Yahoo fell for hook, line, and sinker. They wrote Hadoop, and now we’re all stuck with that pattern. Great gag, they really should have clarified they were kidding. Here we are.

Alex: Exactly. So—

Corey: I mostly kid.

Alex: No, for sure. But I think especially when it comes to data, we tend to over-index on what the large companies do and then quickly realize that we’ve made a mistake and correct backwards, right? So, there was this big push toward MapReduce for everything until people realize that it was just a pain in the neck to operate and to build. And so then we moved into Spark, so kind of up-leveled a little bit. And then there was this kind of explosion of NoSQL and NewSQL databases that hit the market.

And MongoDB inexplicably won that war and now we’re kind of in this world where everything is cloud data warehouse, right? And now we’re trying to wrestle with, like, is it actually a good idea to put everything in one warehouse and have SQL be the lingua franca on top of it? But it’s all changing so rapidly. And when you come into a customer that’s been around for 10 or 15 years, and has, you know, been in the cloud for a substantial—

Corey: Yeah, one of those ancient customers. That is—

Alex: I know, right?

Corey: —basically old enough to almost get a driver’s license? Oh, yeah.

Alex: Right. It’s one of those things where it’s like, “Ah, yes, in startup years, you’re, like, a hundred years old,” right? But still, you know, I think you see this, kind of—I wouldn’t call it a graveyard of failed experiments, right, but it’s a collection of, like, “Well, we tried this, and it kind of worked and we’re keeping it around because the cost of moving this stuff around—the kind of data gravity, so to speak—is high enough that we’re not going to bother transitioning it over.” But then you get into this situation where you have to bend over backwards to integrate anything with anything else. And we’re still kind of in the early days of fixing that.

Corey: And the AWS bill pattern that we see all the time across the board of those experiments were not successful and do not need to exist, but there’s no context into that. The person that set them up left five years ago, the jobs are still running on time. What’s happening with them? Well, we could stop them and see who screams, but very often, that’s not the right answer either.

Alex: And I think there’s also something to note there, too, which is like, getting rid of data is very scary, right? I mean, if you resize a Kubernetes cluster from 15 nodes to 10, nobody’s going to look at you sideways. But if you go, “Hey, we’re just going to drop these tables.” The immediate reaction that you get, particularly from your data science team more often than not is, “Oh, God, what if we need that?” And so the conversation never really happens, and that causes this kind of snowball of data debt that persists in some cases for many, many years.

Corey: Yeah, in some cases, what I found has been successful on those big unknown questions is don’t delete the data, but restrict access to it for a few weeks and see what happens. Look into it a bit and make sure that it’s not like, “Oh, cool. We just did for a month, and now we don’t need that data. Let’s get rid of it.” And then another month goes by it’s like, “So, time to report quarterly earnings. Where’s the data?”

Oh, dear, that’s not going to go well, for anyone. And understanding what’s happening, the idea of cloning a petabyte of data so you can run an experiment on it. And okay, turns out the experiment wasn’t needed. Do we still need to keep all of that?

Alex: Yeah.

Corey: The underlying platform advancements have been helpful toward this as well, a petabyte of data now in Glacier Deep Archive cost the princely sum of a thousand bucks a month, which is pretty close to the idea of why would I ever delete data ever again? I can get it back within a day if I need it, so let’s just put it there instead.

Alex: Right. You know, funny story. When I was in graduate school, we were dealing with, you know, 100 terabyte datasets on the regular that we had to generate every time because we only had 200 terabytes of raw storage. [laugh]. And this was before cloud was yet
mature enough that we could get the kind of performance numbers that we wanted off of it.

And we would end up having to delete the input data to make room for the output data. [laugh]. And thankfully, we don’t need to do that anymore. But there are a lot of, kind of, anti-patterns that arise from that too, right? If data is easy to keep around forever, it stays around forever.

And if it’s easy to, let’s say, run a SQL command against your Snowflake instance that scans 20 terabytes of data, you’re just going to do it, and the exposure of that to you is so minimal that you can end up causing a whole bunch of problems for yourself by the fact that you don’t have to deal with stuff at that low-level of abstraction anymore.

Corey: It’s always fun watching how this stuff manifests—because I’m dipping a toe into it from time to time—the easy, naive answer that we could give every customer but we don’t is, “Huh. So, you have a whole bunch of EMR stuff? Well, you know, if you migrate that into something else, you’ll save a whole bunch of money on that.” With no regard for the 500 jobs that run against that EMR cluster on a consistent basis that form is a key part of business process. “Yeah, if you could just do the entire flow of how data is operated with throughout your entire business that would be swell because you can save tens of thousands of dollars a month on that.” Yeah, how about we don’t suggest things that are just absolute buffoonery.

Alex: Well, and it’s like, you know, you hit on a good point. Like, one of my least favorite words in the English language is the word ‘just.’ And you know, I spent a few years as a freelance data consultant, and you know, a lot of what I would hear sometimes from customers is, “Well, why don’t we ‘just’ deprecate X?”

Corey: “Why don’t we just—” “I’m going to stop you there because there is no ‘just.’”

Alex: Exactly.

Corey: There’s always context that we cannot have as outsiders.

Alex: Precisely. Precisely. And digging into that really is—it’s the fun part of the job, but it’s also the hard part of the job.

Corey: Before we created The Duckbill Group, which was really when I took Mike Julian on as business partner and CEO and formed the entity, I had something in common with you; I was freelancing for a couple of years beforehand. Now, I know why I wound up deciding, all right, we’re going to turn this into a company, but what was it that I guess made you decide to, you know, freelancing is all well and good, but it’s time to get something that looks a lot more like a quote-unquote, “Traditional job.”

Alex: So, I think, on one level, I went freelance because I wasn’t exactly sure what I wanted to do next. And I knew what I was good at. I knew what I had a lot of experience at, and I thought, “Well, I can just go out and kind of find a bunch of people that are willing to hire me to do what I’m good at doing, and then maybe eventually I’ll find one of them that I like enough that I’ll go and work for them. Or maybe I’ll come up with some kind of a business model that I can repeat enough times that I don’t have to worry that I wake up tomorrow and all of my clients are gone and then I have to go live in a van down by the river.”

And I think when I heard about the opening at The Duckbill Group, I had been thinking for a little while about well, this has been going fine for a long time, but effectively what I’ve been doing is I’ve been you know, a staff-level data engineer for hire. And do I want to do something more than that, you know? Do I want to do something more comp—perhaps more sophisticated or more complex than that? And I rapidly came to the conclusion that in order to do that, I would have to have sales and marketing, and I would have to, you know, spend a lot of my time bringing in business. And that’s just not something that I have really any experience in or I’m any good at.

And, you know, I also recognize that, you know, I’m a relatively small fish in a relatively large pond, and if I wanted to get the kind of like, large scale people, the like the big, you know, Fortune 1000 company kind of customers, they may not pay attention to somebody like me. And so I think that ultimately, what I saw with The Duckbill Group was, number one, a group of people that were strongly aligned to the way that I wanted to keep doing this sort of work, right? Cultural alignment was really strong, good people, but also, you know, you folks have a thing that you figured out, and that puts you 10 to 15 steps ahead of where I was. And I was kind of staring down the barrel that, I’m like, am I going to have to take six months not doing client work so that I can figure out how to make this business sustain? And, you know, I think that ultimately, like, I just looked at it, and I said, this just makes sense to me, like, as a next step. And so here we all are.

Corey: This episode is sponsored by our friends at Oracle Cloud. Counting the pennies, but still dreaming of deploying apps instead of “Hello, World” demos? Allow me to introduce you to Oracle’s Always Free tier. It provides over 20 free services and infrastructure, networking, databases, observability, management, and security. And—let me be clear here—it’s actually free. There’s no surprise billing until you intentionally and proactively upgrade your account. This means you can provision a virtual machine instance or spin up an autonomous database that manages itself, all while gaining the networking, load balancing, and storage resources that somehow never quite make it into most free tiers needed to support the application that you want to build. With Always Free, you can do things like run small-scale applications or do proof-of-concept testing without spending a dime. You know that I always like to put asterisks next to the word free? This is actually free, no asterisk. Start now. Visit snark.cloud/oci-free that’s snark.cloud/oci-free.

Corey: It’s always fun seeing how people perceive what we’ve done from the outside. Like, “Oh, yeah, you just stumbled right onto the thing that works, and you’ve just been going, like, gangbusters ever since.” Then you come aboard, it’s like, “Here, look at this pile of things that didn’t pan out over here.” And it’s, you get to see how the sausage is made in a way that we talk about from time to time externally, but surprisingly, most of our marketing efforts aren’t really focused on, “And here’s this other time we screwed up as well.” And we’re honest about it, but it’s not sort of the thing that we promote as the core message of what we do and who we are.

A question I like to ask people during job interviews, and I definitely asked you this, and I’ll ask you now, which is going to probably throw some folks for a loop because who talks to their current employees like this? But what’s next for you? When it comes time for you to leave the Duckbill Group, what do you want to do after this job?

Alex: That’s a great question. So, I mean, as we’ve mentioned before, you know, my career trajectory has been very weird and circuitous. And, you know, I would be lying to you if I said that I had absolute certainty about what the rest of that looks like. I’ve learned a few things about myself in the course of my career, such as it is. In my kind of warm, gooey center, I build stuff. Like, that is what gives me joy, it is what makes me excited to wake up in the morning.

I love looking at big, complicated things, breaking them down into pieces, and figuring out how to make the pieces work in a way that makes sense. And, you know, I’ve spent a long time in the data ecosystem. I don’t know, necessarily, if that’s something that I’m going to do forever. I’m not necessarily pigeonholing myself into that part of the space just yet, but as long as I get to kind of wake up in the morning, and say, “I’m going to go and build things and it’s not going to actively make the world any worse,” I’m happy with that. And so that’s really—you know, might go back to freelancing, might go and join another group, another company, big small, who knows. I’m kind of leaving that up to the winds of destiny, so to speak.

Corey: One thing that I have found incredi—sorry. Let me just address that first. Like that—

Alex: Sure.

Corey: —is the right way to think about it. My belief has always been that you don’t necessarily have, like, the ten-year plan, or the five-year plan or whatever it is because that’s where you’re going to go so much as it gives you direction and forces you to keep moving so you don’t wind up sitting in the same place for five years with one year of experience repeated five times. It helps you remember the bigger picture. Because I’ve always despised this fiction that we see in job interviews where average tenure in our industry is 18 to 36 months, give or take, but somehow during the interviews, we all talk like this is now your forever job, and after 25 years, you’ll retire. And yeah, let’s be a little more realistic than that.

My question is always what is next and how can we align in a way that helps you get to what’s coming? That’s the purpose behind the question, and that’s—the only way to make that not just a drippingly insincere question is to mean it and to continue to focus on it from time to time of, great. What are you learning what’s next? Now, at the time of this recording, you’ve been here, I believe three weeks if I’m not mistaken?

Alex: I’ve—this is week two for me at time of recording.

Corey: Excellent. Yes, my grasp of time is sort of hazy at the best of times. I have a—I do a lot of things.

Alex: For sure.

Corey: But yeah, it has been an eye-opening experience for me, not because, “Oh, wow, we have an employee.” Yeah, we’ve done that a few times before. But rather because of your background, you are asking different questions than we typically get during onboarding. I had a blog post go out recently—or will be by the time this airs—about a question that you asked about, “Wow, onboarding into our internal account structure for AWS is way more polished than I’ve ever seen it before. Is that something you built in-house? What is that?”

And great. Oh, terrific, I’d forgotten that this is kind of a novel thing. No. What we’re using is AWS’s SSO offering, which is such a well-built, polished product that I can only assume that it’s under NDA because Amazonians don’t talk about it ever. But it’s great.

It has a couple of annoyances, but beyond that, it’s something that I’m a big fan of, but I’d forgotten how transformative that is, compared to the usual approach of all right, here’s your username, here’s a password you’re going to have to change, here are your IAM credentials to store on disk forever. It’s the ability to look at what we’re doing through the eyes of someone who is clearly deep into the technical weeds, but not as exposed to all of the minutiae of the 300-some-odd AWS services is really a refreshing thing for all of us, just because it helps us realize what it’s like to see some of this stuff for the first time, as well as gives me content ideas because if it’s new to you, I promise you are not the only person who’s seeing it that way. And if you don’t really understand something well enough to explain it, I would argue you don’t really understand the thing, so it forces me to get more awareness around exactly how different facets work. It’s been an absolutely fantastic experience so far, from my perspective.

Alex: Thank you. Right back at you. I mean, spending so many years working with startups, my kind of level of expected sophistication is, “I’m going to write your password on the back of a napkin. I have fifteen other things to do. Go figure it out.” And so you know, it’s always nice to see—particularly players like AWS that are such 800-pound gorillas—going in and trying to uplevel that experience in a way that feels like—because I mean, like, look, AWS could keep us with the, “Here’s a CSV with your username and password. Good luck, have fun.” And you know, they would still make—

Corey: And they’re going to have to because so much automation is built around that—

Alex: Oh yeah—

Corey: In so many places.

Alex: —so much.

Corey: It’s always net-additive, they never turn anything off, which is increasingly an operational burden.

Alex: Yeah, absolutely. Absolutely. But yeah, it’s nice to see them up-level this in a way that feels like they’re paying attention to their customers’ pain. And that’s always nice to see.

Corey: So, we met a few years ago—in the before times—at a mixer that we wound up throwing—slash meetup. It was in Southern California for some AWS event or another. You’ve been aware of who we are and what we do for a while now, so I’m very curious to know—and the joy of having these conversations is that I don’t actually know what the answer is going to be, so this may never see the light of day if it goes to weird—

Alex: [laugh].

Corey: —in the wrong direction, but—no I’m kidding. What has been, I guess, the biggest points of dissonance or surprises based upon your perception of who we are and what we do externally, versus joining and seeing how the sausage is made?

Alex: You know, I think the first thing is—um, well, how to put this. I think that a lot of what I was expecting, given how much work you all do and how big—well, ‘you all;’ we do—and how big the list of clients is and how it gets bigger every day, I was expecting this to be, like, this very hyper put together, like, every little detail has been figured out kind of engagement where I would have to figure out how you all do this. And coming in and realizing that a lot of it is just having a lot of in-depth knowledge born from experience of a bunch of stuff inside of this ecosystem, and then the rest of it is kind of free jazz, is kind of encouraging. Because as someone that was you know, as a freelancer, right, who do you see, right? You see people who have big public presences or people who are giant firms, right?

On the GCP side, SADA Systems is a great example. They’re another local company for me here in Los Angeles, and—

Corey: Oh, yes. [unintelligible 00:24:48] Miles has been a recurring guest on the show.

Alex: Yeah. And he’s great. And, like, they have this enormous company that’s got, like, all these different specializations and they’re basically kind of like the middleman for GCP on a lot of things. And, like, you see that, and then you kind of see the individual people that are like, “Yeah, you know, I’m not really going to tell you that I only have two clients and that if both of them go away, I’m screwed, but, like, I only have two clients, and if both of them go away, I’m screwed.” And so, you know, I think honestly seeing that, like, what you’ve built so far and what I hope to help you continue to build is, you know, you’ve got just enough structure around the thing so that it makes sense, and the rest of it, you’re kind of admitting that no plan ever survives contact with the client, right, and that everybody’s going to be different than that everybody’s problems are going to be different.

And that you can’t just go in and say, “Here’s a dashboard, here’s a calculator, have fun, give me my money,” right? Because that feels like—in optimization spaces of any kind, be that cloud, or data or whatever, there’s this, kind of, push toward, how do I automate myself out of a job, and the realization that you can’t for something like this, and that ultimately, like, you’re just going to have to go with what you know, is something that I kind of had a suspicion was the case, but this really made it clear to me that, like, oh, this is actually a reasonable way of going about this.

Corey: We thought otherwise at one point. We thought that this was something could be easily addressed their software. We launched our DuckTools SaaS platform in beta and two months later, did the—our incredible journey has come to an end, and took it off of a public offering. Because it doesn’t lend itself to solving these problems in software in any reasonable way. I am ever more convinced over time that the idea of being able to solve cloud cost optimization with software at VC-scale is a red herring.

And yeah, it just isn’t going to work because it’s one size fits some. Our customers are, by definition, exceptional in many respects, and understanding the context behind why things are the way that they are mean that we can only go so far with process because then it becomes a let’s have a conversation and let’s be human. Otherwise, we try to overly codify the process, and congratulations, we just now look like really crappy software, but expensive because it’s all people doing it. It doesn’t work that way. We have tools internally that help smooth over a lot of those edges, but by and large, people who are capable of performing at especially at the principal level for a cloud economics role, inherently are going to find themselves stifled by too much process because they need to have the freedom to dig into the areas that are relevant to the customer.

It’s why we can’t recraft all of our statements of work in ways that tend to shy away from explicitly defined deliverables. Because we deliver an outcome, but it’s going to depend entirely, in most cases, up on what we discover along the way. Maybe a full-on report isn’t the best way of presenting the data in the way that we see it. Maybe it’s a small proof of concept script or something like that. Maybe it’s,
I don’t know, an interpretive dance in front of the company’s board.

Alex: [laugh]. Right.

Corey: I’m open to exploring opportunities. But it comes down to what is right for the customer. There’s a reason we only ever charge a fixed fee for these things, and it’s because at that point, great, we’re giving you the advice that we’d implement ourselves. We have no partnerships with any vendor in the space just to avoid bias or the perception of same. It’s important that we are the authoritative source around these things.

Honestly, the thing that surprised me the most about all this is how true to that vision we’ve stayed as we’ve as we flushed out what works, what doesn’t. And we can distantly fail to go out of business every month. I am ecstatic about that. I expected this to wind up cratering into a mountain four months after I went freelance. Not yet.

Alex: Well, I mean, I think there’s another aspect of this too, right? Because I’ve spent a lot of my career working inside of venture capital-backed companies. And there’s a lot of positive things to be said about having ready access to that kind of cash, but it does something to your business the second you take it. And I’ve been in a couple of situations where, like, once you actually have that big bucket of money, the incentive is grow, right? Hire more people get more customers, go, go, go, go, go.

And sometimes what you’ll find is that you’ll spend the time and the money on an initiative and it’s clearly not working. And you just kind of have to keep doubling down because now you’ve got customers that are using this thing and now you have to maintain it, and before you know it, you’ve got this albatross hanging around your neck. And like one of the things that I really respect about the way that Duckbill Group is is handling this by not taking outside cash is, like, it frees you up to make these kinds of bets, and then two months later say, “Well, that didn’t work,” and try something else. And you know, that’s very difficult to do once you have to go and convince someone with, you know, money flowing out of their ears, that that’s the right thing to do.

Corey: We have to be intentional about what we’re doing. One of the benefits of bringing you aboard is that one, it does improve our capacity for handling more engagements at the same time, but it also improves the quality of the engagements that we are delivering. Instead of basically doing a round-robin assignment policy we can—

Alex: Right.

Corey: —we consult with each other; we talk about specific areas in which we have specific expertise. You get dragged into a lot of data
portions of existing engagements, and the rest of us get pulled into other areas in which you might not be as strong. For example, “What
are all of these ridiculous services? I can’t make heads or tails have the ridiculous naming side of it.” Surprise, that’s not a you problem.

It comes down to being able to work collaboratively and let each other shine in a way that doesn’t mean we load people up with work. We’re very strict about having a 40-hour or less work week, just because we’re not rushing for an exit. We want to enjoy our time working, we want to enjoy what we’re doing, and then we want to go home and don’t think about work until it’s time to come back and think about these things. Like, it’s a lifestyle company, but that lifestyle doesn’t need to be run, run, run, run, run all the time, and it doesn’t need to be something that people barely tolerate.

Alex: Yeah. And I think that, you know, especially coming from being an army of one in a lot of engagements, it is really refreshing to be able to—see because, you know, I’m fortunate enough, I have friends in the industry that I can go and say like, “I have no idea how to make heads or tails of X.” And you know, I can get help that way, but ultimately, like, the only other outlet that I have here is the customer and they’re not bringing me in if they have those answers readily to hand. And so being able to bounce stuff off of other people inside of an organization like this has been really refreshing.

Corey: One of the things I’ve appreciated about your tenure here so far is the questions that you ask are pitched at the perfect level, by which I mean, it is never something you could answer with a three-second visit to Google, but it’s also not something that you’ve spent three days spinning your wheels on trying to understand. You do a bit of digging; it’s a little unclear, especially since there are multiple paths to go down, and then you flag it for clarification. And there’s really so much to be said for that. Really, when we’re looking for markers of seniority in the interview process, it’s admitting you don’t know something, but then also talking about how you would go about getting the answer. And it’s—because no one has all this stuff in their head. I spend a disturbing amount of time looking at search engines and trying to reformulate queries and to get answers that make sense.

I don’t have the entirety of AWS shoved into my head. Yet. I’m sure there’s something at re:Invent that’s going to be scary and horrifying that will claim to do it and basically have a poor user interface, but all right. When that comes, we’ll reevaluate then because this industry is always changing.

Alex: For sure. For sure. And I think it’s, it’s worth pointing out that, like, one of the things that having done this for a long time gives you is this kind of scaffolding in your head that you can hang things over. We’re like, you don’t need to have every single AWS service memorized, but if you’ve got that scaffold in your head going, “Oh, like, this thing sounds like it hangs over this part of the mental scaffold, and I’ve seen other things that do that, so I wonder if it does this and this and this,” right? And that’s a lot of it, honestly.

Because especially, like, when I was solely in the data space, there’s a new data wareho—or a new, like, data catalog system coming out every other week. You know, there are a thousand different things that claim to do MLOps, right? And whenever, like, someone comes to me and says, “Do you have experience with such and such?” And the answer was usually, “Well if you hum a few bars, I can fake it.” And, you know, that tends to help a great deal.

Corey: Yeah. “No, but I’ll find out and get back to you,” the right answer. Making it up and being wrong is the best way to get rejected from an environment. That’s not just consulting; that’s employment, too. If 95% of the time, you give the right answer, but that one time and 20 you’re going to just make it up, well, I have to validate the other 19 because I never know when someone’s faking it or not. There’s that level of earned trust that’s important.

Alex: Well, yeah. And you’re being brought in to be the expert in the room. That doesn’t necessarily mean that you are the all-seeing, all-knowing oracle of knowledge but, like, if you say a thing, people are just going to believe you. And so, you know, it’s beholden on you—

Corey: If not, we have a different problem.

Alex: Well, yeah, exactly. Hopefully, right? But yeah, I mean, it’s beholden on you to be honest with your customer at a certain point, I think.

Corey: I really want to thank you for taking the time out of your day to got with me about this. And I would love to have you back on in a couple of months once you’re fully up to speed and spinning at the proper RPMs and see what’s happened then. I—

Alex: Thank you. I’d—

Corey: —really appreciate—

Alex: —love to.

Corey: —your time where’s the best place for people to learn more about you if they haven’t heard your name before?

Alex: Well, let’s see. I am @alexras on Twitter, A-L-E-X-R-A-S. My personal website is alexras.info.

I’ve done some writing on data stuff, including a pretty big collection of blog posts on the data side of the AWS ecosystem that are still on my consulting page, bitsondisk.com. Other than that—I mean, yeah, Twitter is probably the best place to find me, so if you want to talk more about any weird, nerd data stuff, then please feel free to reach out there.

Corey: And links to that will, of course, be in the [show notes 00:35:57]. Thanks again for your time. I really appreciate it.

Alex: Thank you. It’s been a pleasure.

Corey: Alex Rasmussen, principal cloud economist here at The Duckbill Group. I am Corey Quinn, cloud economist to the stars, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry, insulting comment that
you then submit to three other podcast platforms just to make sure you have a backup copy of that particular piece of data.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Aidan

Aidan is an AWS enthusiast, due in no small part to sharing initials with the cloud. He's been writing software for over 20 years and getting paid to do it for the last 10. He's still not sure what he wants to be when he grows up.

Links:

  • Stedi: https://www.stedi.com/
  • GitHub: https://github.com/aidansteele
  • Blog posts: https://awsteele.com/
  • Ipv6-ghost-ship: https://github.com/aidansteele/ipv6-ghost-ship
  • Twitter: https://twitter.com/__steele

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Couchbase Capella Database-as-a-Service is flexible, full-featured and fully managed with built in access via key-value, SQL, and full-text search. Flexible JSON documents aligned to your applications and workloads. Build faster with blazing fast in-memory performance and automated replication and scaling while reducing cost. Capella has the best price performance of any fully managed document database. Visit couchbase.com/screaminginthecloud to try Capella today for free and be up and running in three minutes with no credit card required. Couchbase Capella: make your data sing.

Corey: Today’s episode is brought to you in part by our friends at MinIO the high-performance Kubernetes native object store that’s built for the multi-cloud, creating a consistent data storage layer for your public cloud instances, your private cloud instances, and even your edge instances, depending upon what the heck you’re defining those as, which depends probably on where you work. It’s getting that unified is one of the greatest challenges facing developers and architects today. It requires S3 compatibility, enterprise-grade security and resiliency, the speed to run any workload, and the footprint to run anywhere, and that’s exactly what MinIO offers. With superb read speeds in excess of 360 gigs and 100 megabyte binary that doesn’t eat all the data you’ve gotten on the system, it’s exactly what you’ve been looking for. Check it out today at min.io/download, and see for yourself. That’s min.io/download, and be sure to tell them that I sent you.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’m joined this week by someone who is honestly, feels like they’re after my own heart. Aidan Steele by day is a serverless engineer at Stedi, but by night, he is an absolute treasure and a delight because not only does he write awesome third-party tooling and blog posts and whatnot around the AWS ecosystem, but he turns them into the most glorious, intricate, and technical shit posts that I think I’ve ever seen. Aidan, thank you for joining me.

Aidan: Hi, Corey, thanks for having me. It’s an honor to be here. Hopefully, we get to talk some AWS, and maybe also talk some nonsense as well.

Corey: I would argue that in many ways, those things are one in the same. And one of the things I always appreciated about how you approach things is, you definitely seem to share that particular ethos with me. And there’s been a lot of interesting content coming out from you in recent days. The thing that really wound up showing up on my radar in a big way was back at the start of January—2022, for those listening to this in the glorious future—about using IPv6 to use multi-factor auth, which it is so… I don’t even have the adjectives to throw at this because, first it is ridiculous, two, it is effective, and three, it is just who thinks like that? What is this and what did you—what monstrosity have you built?

Aidan: So, what did I end up calling it? I think it was ipv6-ghost-ship. And I think I called it that because I’d recently watched, oh, what was that series that was recently on Apple TV? Uh, the Isaac Asimov—

Corey: If it’s not Paw Patrol, I have no idea what it is because I have a four-year-old who is very insistent about these things. It is not so much a TV show as it is a way of life. My life is terrible. Please put me out of my misery.

Aidan: Well, at least it’s not Bluey. That’s the one I usually hear about. That’s Australia’s greatest export. But it was one of the plot devices was a ship that would teleport around the place, and you could never predict where it was next. And so no one could access it. And I thought, “Oh, what about if I use the IPv6 address space?”

Corey: Oh, Foundation?

Aidan: That’s the one. Foundation. That’s how the name came about. The idea, honestly, it was because I saw—when was it?—sometime last year, AWS added support for those IP address prefixes. IPv4 prefixes were small; very useful and important, but IPv6 with more than 2 trillion IP addresses, per instance, I thought there’s got to be fun to be had there.

Corey: 281 trillion, I believe is the—

Aidan: 281 trillion.

Corey: Yeah. It is sarcastically large space. And that also has effectively, I would say in InfoSec sense, killed port scanning, the idea I’m going to scan the IP range and see what’s there, just because that takes such a tremendous amount of time. Now here, in reality, you also wind up with people using compromised resources, and yeah, it turns out, I can absolutely scan trillions upon trillions of IP addresses as long as I’m using your AWS account and associated credit card in which to do it. But here in the real world, it is not an easily discoverable problem space.

Aidan: Yeah. I made it as a novelty, really. I was looking for a reason to learn more about IPv6 and subnetting because it’s the term I’d heard, a thing I didn’t really understand, and the way I learn things is by trying to build them, realizing I have no idea what I’m doing, googling the error messages, reluctantly looking at the documentation, and then repeating until I’ve built something. And yeah, and then I built it, published it, and seemed to be pretty popular. It struck a chord. People retweeted it. It tickled your fancy. I think it spoke something in all of us who are trying not to take our jobs too seriously, you know, know we can have a little fun with this ludicrous tech that we get to play with.

Corey: The idea being, you take the multi-factor auth code that your thing generates, and that is the last series of octets for the IP address you wind up going towards and that is such a large problem space that you’re not going to find it in time, so whatever it is automatically connect to that particular IP address because that’s the only one that’s going to be listening for a 30 to 60-second span for the connection to be established. It is a great idea because SSH doesn’t support this stuff natively. There’s no good two-factor auth approach for this. And I love it. I’d be scared to death to run this in production for something that actually matters.

And we also start caring a lot more about how accurate are the clocks on those instances, all of a sudden. But, oh, I just love the concept so much because it hits on the ethos of—I think—what so much of the cloud does were these really are fundamental building blocks that we can use to build incredible, awe-inspiring things that are globe-spanning, and also ridiculousness. And there’s so much value of being able to do the same thing, sometimes at the same time.

Aidan: Yeah, it’s interesting, you mentioned, like, never using in prod, and I guess when I was building it, I thought, you know, that would be apparent. Like, “Yes, this is very neat, but surely no one’s going to use it.” And I did see someone raised an issue on the GitHub project which was talking about clock skew. And I mentioned—

Corey: Here at the bank where I’m running this in production, we’re—

Aidan: [laugh].

Corey: —having some trouble with the clock. Yeah, it’s—

Aidan: You know, I mentioned that the underlying 2FA library did account for clock scheme 30 seconds either way, but it made me realize, I might need to put a disclaimer on the project. While the code is probably reasonably sound, I personally wouldn’t run it in production, and it was more meant to be a piece of performance art or something to tickle one’s fancy and to move on, not to roll it out. But I don’t know, different strokes for different folks.

Corey: I have gotten a lot better about calling out my ridiculous shitpost things when I do them. And the thing that really drove that home for me was talking about using DNS TXT records to store information about what server a virtual machine lives on—or container or whatnot—thus using Route 53 is a database. And that was a great gag, and then someone did a Reddit post of “This seems like a really good idea, so I’m going to start doing it, and I’m having these questions.”

And at that point is like, “Okay, I’ve got a break character at that point.” And is, yeah, “Hi. That’s my joke. Don’t do it because X, Y, and Z are your failure modes, there are better tools for it. So yeah, there are ways you can do this with DNS, but it’s not generally a great idea, and there are some risk factors to it. And okay, A, B, and C are the things you don’t want to do, so let’s instead do it in a halfway intelligent way because it’s only funny if everyone’s laughing. Otherwise, we fall into this trap of people take you seriously and they feel bad as a result when it doesn’t work in production. So, calling it out as this is a joke tends to put a lot of that aside. It also keeps people from feeling left out.

Aidan: Yeah. I realized that because the next novelty project I did a few days later—not sure if you caught it—it was a Rick Roll over ICMPv6 packets, where if you had run ping six to a certain IP range, it would return the lyrics to music’s greatest treasure. So, I think that was hopefully a bit more self-evident that this should never be taken seriously. Who knows, I’m sure someone will find a use for it in prod.

Corey: And I was looking through this, this is great. I love some of the stuff that you’re doing because it’s just fantastic. And I started digging a bit more to things you had done. And at that point, it was whoa, whoa, whoa, wait a minute. Back in 2020, you found an example of an issue with AWS’s security model where CloudTrail would just start—if asked nicely—spewing other people’s credential sets and CloudTrail events and whatnot into your account.

And, A, that’s kind of a problem. B, it was something that didn’t make that big of a splash when it came out—I don’t even think I linked to it at the time—and, C, it was examples of after the recent revelations around CloudFormation and Glue that the fine folks at Orca Security found out. That wasn’t a one-off because you’d done this a year beforehand. We have now an established track record of cross-account data sharing and, potentially, exploits, and I’m looking at this and I got to level with you I felt incredibly naive because I had assumed that since we hadn’t heard of this stuff in any real big sense that it simply didn’t happen.

So, when we heard about Azure; obviously, it’s because Azure is complete clown shoes and the excellent people that AWS would never make these sorts of mistakes. Except we now have evidence that they absolutely did and didn’t talk about it publicly. And I’ve got a level with you. I feel more than a little bit foolish, betrayed, naive for all this. What’s your take on it?

Aidan: Yeah, so just to clarify, it wasn’t actually in your account. It was the new AWS custom resource execution model was you would upload a Lambda function that would run in an Amazon-managed account. And so that immediately set off my spidey sense because executing code in someone else’s account seems fraught with peril. And so—

Corey: Yeah, you can do all kinds of horrifying things there, like, use it to run containers.

Aidan: Yeah. [laugh]. Thankfully, I didn’t do anything that egregious. I stayed inside the Lambda function, but I look—I poked around at
what credentials have had, and it would use CloudWatch to reinvoke itself and CloudWatch kept recording CloudTrail. And I won’t go into all the details, but it ended up being that you could see credentials being recorded in CloudTrail in that account, and I could, sort of,
funnel them out of there.

When I found this, I was a little scared, and I don’t think I’d reported an issue to AWS before, so I didn’t want to go too far and do anything that could be considered malicious. So, I didn’t actively seek out other people’s credentials.

Corey: Yeah, as a general rule, it’s best once you discover things like that to do the right thing and report it, not proceed to, you know, inadvertently commit felonies.

Aidan: Yeah. Especially because it was my first time. I felt better safe than sorry. So, I didn’t see other credentials, but I had no reason to believe that, I wouldn’t see it if I kept looking. I reported it to Amazon. Their security team was incredibly professional, made me feel very
comfortable reporting it, and let me know when, you know, they’d remediated it, which was a matter of days later.

But afterwards, it left me feeling a little surprised because I was able to publish about it, and a few people responded, you know, the sorts of people who pay close attention to the industry, but Amazon didn’t publish anything as far as I was aware. And it changed the way I felt about AWS security, because like you, I sort of felt that AWS, more or less had a pretty perfect track record. They would have advisories about possible [Zen 00:12:04] exploits, and so on. But they’d never published anything about potential for compromise. And it makes me wonder how many of the things might have been reported in the past where either the third-party researcher either didn’t end up publishing, or they published and it just disappeared into the blogosphere, and I hadn’t seen it.

Corey: They have a big earn trust principle over there, and I think that they always focus on the trust portion of it, but I think what got overlooked is the earn. When people are giving you trust that you haven’t earned, on some level, the right thing to do is to call it out and be transparent around these things. Yes, I know, Wall Street’s going to be annoyed and headlines, et cetera, et cetera, but I had always had the impression that had there been a cross-account vulnerability or a breach of some sort, they would communicate this and they would have their executives go on a speaking tour about it to explain how defense-in-depth mitigated some of it, and/or lessons learned, and/or what else we can learn. But it turns out that wasn’t was happening at all. And I feel like they have been given trust that was unearned and now I am not happy with it.

I suddenly have a lot more of a, I guess, skeptical position toward them as a result, and I have very little tolerance left for what has previously been a staple of the AWS security discussions, which is an executive getting on stage for a while and droning on about the shared responsibility model with the very strong implication that “Oh, yeah, we’re fine. It’s all on your side of the fence that things are going to break.” Yeah, turns out, that’s not so true. Just you know, about the things on your side of the fence in a way that you don’t about the things that are on theirs.

Aidan: Yeah, it’s an interesting one. Like, I think about it and I think, “Well, they never made an explicit promise that they would publish these things,” so, on one hand, I say to myself, “Oh, maybe that’s on me for making that assumption.” But, I don’t know, I feel like the way we felt was justified. Maybe naive in hindsight, but then, you know, I guess… I’m still not sure how to feel because of, like, I think about recent issues and how a couple of AWS Distinguished Engineers jumped on Twitter, and to their credit were extremely proactive in engaging with the community.

But is that enough? It might be enough for say, to set my mind at ease or your mind at ease because we are, [laugh] to put it mildly, highly engaged, perhaps a little too engaged in the AWS space, but Twitter’s very ephemeral. Very few of AWS’s customers—

Corey: Yeah, I can’t link to tweets by distinguished engineers to present to an executive leadership team as an official statement from Amazon. I just can’t.

Aidan: Yeah. Yeah.

Corey: And so the lesson we can take from this is okay, so “Well, we never actually said this.” “So, let me get this straight. You’re content to basically let people assume whatever they want until they ask you an explicit question around these things. Really? Is that the lesson you want me to take from this? Because I have a whole bunch of very explicit questions that I will be asking you going forward, if that is in fact, your position. And you are not going to like the fact that I’m asking these questions.”

Even if the answer is a hard no, people who did not have this context are going to wonder why are people asking those questions? It’s a massive footgun here for them if that is the position that they intend to have. I want to be clear as well; this is also a messaging problem.
It is not in any way, a condemnation of their excellent folks working on the security implementation themselves. This stuff is hard and those people are all-stars. I want to be very clear on this. It is purely around the messaging and positioning of the security posture.

Aidan: Yeah, yeah. That’s a good clarification because like you, my understanding that the service teams are doing a really stellar, above-average job, industry-wide, and the AWS Security Response Teams, I have absolute faith in them. It is a matter of messaging. And I guess what particularly brings it to front-of-mind is, it was earlier this month, or maybe it was last month, I received an email from a company called Sourcegraph. They do code search.

I’m not even a customer of theirs yet, you know? I’m on a free trial, and I got an email that—I’m paraphrasing here—was something to the effect of, we discovered that it was possible for your code to appear in other customers’ code search results. It was discovered by one of our own engineers. We found that the circumstances hadn’t cropped up, but we wanted to tell you that it was possible. It didn’t happen, and we’re working on making sure it won’t happen again.

And I think about how radically different that is where they didn’t have a third-party researcher forcing their hand; they could have very easily swept under the rug, but they were so proactive that, honestly, that’s probably what’s going to tipped me over to the edge into me becoming a customer. I mean, other than them having a great product. But yeah, it’s a big contrast. It’s how I like to see other companies work, especially Amazon.

Corey: This episode is sponsored in part by our friends at Sysdig. Sysdig is the solution for securing DevOps. They have a blog post that went up recently about how an insecure AWS Lambda function could be used as a pivot point to get access into your environment. They’ve also gone deep in-depth with a bunch of other approaches to how DevOps and security are inextricably linked. To learn more, visit sysdig.com and tell them I sent you. That’s S-Y-S-D-I-G dot com. My thanks to them for their continued support of this ridiculous nonsense.

Corey: The two companies that I can think of that have had security problems have been CircleCI and Travis CI. Circle had an incredibly transparent early-on blog post, they engaged with customers on the forums, and they did super well. Travis basically denied, stonewalled for ages, and now the only people who use Travis are there because they haven’t found a good way to get off of it yet. It is effectively DOA. And I don’t think those two things are unrelated.

Aidan: Yeah. No, that’s a great point. Because you know, I’ve been in this industry long enough. You have to know that humans write code and humans make mistakes—I know I’ve made more than my fair share—and I’m not going to write off the company for making a mistake. It’s entirely in their response. And yeah, you’re right. That’s why Circle is still a trustworthy business that should earn people’s business and why Travis—why I recommend everyone move away from.

Corey: Yeah, I like Orca Security as a company and as a product, but at the moment, I am not their customer. I am AWS’s customer. So, why the hell am I hearing it from Orca and not AWS when this happens?

Aidan: Yeah, yeah. It’s… not great. On one hand, I’m glad I’m not in charge of finding a solution to this because I don’t have the skills or the expertise to manage that communication. Because like I think you said in the past, there’s a lot of different audiences that they have to communicate with. They have to communicate with the stock market, they have to communicate with execs, they have to communicate with developers, and each of those audiences demands a different level of detail, a different focus. And it’s tricky. And how do you manage that? But, I don’t know, I feel like you have an obligation to when people place that level of trust in you.

Corey: It’s just a matter of doing right by your customers, on some level.

Aidan: Yeah.

Corey: How long have you been working on an AWS-side environments? Clearly, this is not like, “Well, it’s year two,” because if so I’m going to feel remarkably behind.

Aidan: [laugh]. So, I’ve been writing code in some capacity or another for 20 years. It took about five years to get anyone to pay me to do so. But yeah, I guess the start of my professional career—and by ‘professional,’ I want to use it in strictest term, means getting paid for money; not that I [laugh] am necessarily a professional—coincided with the launch of AWS. So, I don’t hadn’t experienced with the before times of data centers, never had to think about direct connect, but it means I have been using AWS since sometime in 2008.

I was just looking at my bill earlier, I saw that my first bill was for $70. It was—I was using a C1xLarge, which was 80 cents an hour, and it had eight-core CPUs. And to put that in context at the time—

Corey: Eight vCPUs, technically I believe.

Aidan: An it basically is—

Corey: —or were they using [eCPU 00:20:31] model back then?

Aidan: Yeah, no, that was vCPUs. But to me, that was extraordinary. You know, I was somewhere just after high school. It was—the Netflix Prize was around. If you’re not sure what that was, it was Netflix had this open competition where they said anyone who could improve upon their movie recommendation algorithm could win a million dollars.

And obviously being a teenager, I had a massive ego and [laugh] no self-doubt, so I thought I could win this, but I just don’t have enough CPUs or RAM on my laptop. And so when EC2 launched, and I could pay 80 cents an hour, rather than signing up for a 12-month contract with a colocation company, it was just a dream come true. I was able to run my terrible algorithms, but I could run them eight times faster. Unfortunately and obviously, I didn’t win because it turns out, I’m not a world-class statistician. But—

Corey: Common mistake. I make that mistake myself all the time.

Aidan: [laugh]. Yeah. I mean, you know, I think I was probably 19 at the time, so I had—my ego did make me think I was one, but it turned out not to be so. But I think that was what really blew my mind was that me, a nobody, could create an account with Amazon and get access to these incredibly powerful machines for less than the dollar. And so I was hooked.

Since then, I’ve worked at companies that are AWS customers since then. I’ve worked at places that have zero EC2 service, worked at places that have had thousands, and places in between. And it’s got to a point, actually, where, I guess, my career is so entwined with AWS that one, my initials are actually AWS, but also—and this might sound ridiculous, and it’s probably just a sign of my privilege—that I wouldn’t consider working somewhere that used another cloud. Not—

Corey: No, I think that’s absolutely the right approach.

Aidan: Yeah.

Corey: I had a Twitter thread on this somewhat recently, and I’m going to turn it into a blog post because I got some pushback. If I were looking at doing something and I would come into the industry right now, my first choice would be Google Cloud because its developer experience is excellent. But I’m not coming to this without any experience. I have spent a decade or so learning not just how it was works, but also how it breaks, understanding the failure mode and what that’s going to look like and what it’s good at and what it’s not. That’s the valuable stuff for running things in a serious way.

Aidan: Yeah. It’s an interesting one. And I mean, for better or worse, AWS is big. I’m sure you will know much better than I do the exact numbers, but if a junior developer came to me and said, “Which cloud should I learn, or should I learn all of them?” I mean, you’re right, Google Cloud does have a better developer experience, especially for new developers, but when I think about the sheer number of jobs that are available for developers, I feel like I would be doing them a disservice by not suggesting AWS, at least in Australia. It seems they’ve got such a huge footprint that you’ll always be able to find a job working as an AWS-familiar engineer. It seems like that would be less the case with Google Cloud or Azure.

Corey: Again, I am not sitting here, suggesting that anyone should, “Oh, clouds are insecure. We’re going to run our own stuff in our own data centers.” That is ridiculous in this era. They are still going to do a better job of security than any of us will individually, let’s be clear here. And it empowers and unlocks an awful lot of stuff.

But with their privileged position as these hyperscale providers that are the default choice for building things, I think comes with a significant level of responsibility that I am displeased to discover that they’ve been abdicating. And I don’t love that.

Aidan: Yeah, it’s an interesting one, right, because, like you’re saying, they have access and the expertise that people doing it themselves will never match. So, you know, I’m never going to hesitate to recommend people use AWS on account security because your company’s security posture will almost always be better for using AWS and following their guidelines, and so on. But yeah, like you say, with great power comes significant responsibility to earn trust and retain that trust by admitting and publicizing when mistakes are made.

Corey: One last topic I want to get into with you is one that you and I have talked about very briefly elsewhere, that I feel like you and I are both relatively up-to-date on AWS intricacies. I think that we are both better than the average bear working with the platform. But I know that I feel this way, and I suspect you do too that VPCs have gotten confusing as hell. Is that just me? Am I a secret moron that no one bothered to ever tell me this, and I should update my own self-awareness?

Aidan: [laugh]. Yeah, it’s… I mean, that’s been the story of my career with AWS. When I started, VPCs didn’t exist. It was EC2 Classic—well, I guess at the time, it was just EC2—and it was simple. You launched an instance and you had an IP address.

And then along came VPCs, and I think at the time, I thought something to the effect of “This seems like needless complexity. I’m not going to bother learning this. It will never be relevant.” In the end that wasn’t true. I worked in much large deployments when VPCs made fantastic sense made a lot of things possible, but I still didn’t go into the weeds.

Since then, AWS has announced that EC2 Classic will be retired; an end of an era. I’m not personally still running anything in EC2 Classic, and I think they’ve done an incredible job of maintain support for this long, but VPC complexity has certainly been growing year-on-year since then. I recently was using the AWS console—like we all do and no one ever admits to—to edit a VPC subnet route table. And I clicked the drop-down box for a target, and I was overwhelmed by the number of options. There were NAT gateways, internet gateways, carrier gateways, I think there was a thing called a wavelength gateway, ENI, and… I [laugh] I think I was surprised because I just scroll through the list, and I thought, “Wow, that is a lot of different options. Why is that?”

Especially because it’s not so relevant to me. But I realized a big thing of what AWS has been doing lately is trying to make themselves available to people who haven’t used the cloud yet. And they have these complicated networking needs, and it seems like they’re trying to—reasonably successfully—make anything possible. But with that comes, you know, additional complexity.

Corey: I appreciate that the capacity is there, but there has to be an abstraction model for getting rid of some of this complexity because otherwise, the failure mode is you wind up with this amazingly capable thing that can build marvels, but you also need to basically have a PhD in some of these things to wind up tying it all together. And if you bring someone else in to do it, then you have no idea how to run the thing. You’re effectively a golden retriever trying to fly a space shuttle.

Aidan: Yeah. It’s interesting, like, clearly, they must be acutely aware of this because they have default VPCs, and for many use cases, that’s all people should need. But as soon as you want, say a private subnet, then you need to either modify that default VPC or create a new one, and it’s sort of going from 0 to 100 complexity extremely quickly because, you know, you need to create route tables to everyone’s favorite net gateways, and it feels like the on-ramp needs to be not so steep. Not sure what the solution is, I hope they find one.

Corey: As do I. I really want to thank you for taking the time to speak with me about so many of these things. If people want to learn more about what you’re up to, where’s the best place to find you?

Aidan: Twitter’s the best place. On Twitter, my username is @__Steele, which is S-T-E-E-L-E. From there, that’s where I’ll either—I’ll at least speculate on the latest releases or link to some of the silly things I put on GitHub. Sometimes they’re not so silly things. But yeah, that’s where I can be found. And I’d love to chat to anyone about AWS. It’s something I can geek out about all day, every day.

Corey: And we will certainly include links to that in the [show notes 00:29:50]. Thank you so much for taking the time to speak with me
today. I really appreciate it.

Aidan: Well, thank you so much for having me. It’s been an absolute delight.

Corey: Aidan Steele, serverless engineer at Stedi, and shit poster extraordinaire. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an immediate request to correct the record about what I’m not fully understanding about AWS’s piss-weak security communications.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About David

David is an AWS expert who likes to design and build scalable solutions that are fully automated and take care of themselves. Now he is focusing on selling his own products on the AWS Marketplace.

Links:

  • 0x4447: https://0x4447.com/
  • Products page: https://products.0x4447.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Today’s episode is brought to you in part by our friends at MinIO the high-performance Kubernetes native object store that’s built for the multi-cloud, creating a consistent data storage layer for your public cloud instances, your private cloud instances, and even your edge instances, depending upon what the heck you’re defining those as, which depends probably on where you work. It’s getting that unified is one of the greatest challenges facing developers and architects today. It requires S3 compatibility, enterprise-grade security and resiliency, the speed to run any workload, and the footprint to run anywhere, and that’s exactly what MinIO offers. With superb read speeds in excess of 360 gigs and 100 megabyte binary that doesn’t eat all the data you’ve gotten on the system, it’s exactly what you’ve been looking for. Check it out today at min.io/download, and see for yourself. That’s min.io/download, and be sure to tell them that I sent you.

Corey: This episode is sponsored in part by our friends at Sysdig. Sysdig is the solution for securing DevOps. They have a blog post that went up recently about how an insecure AWS Lambda function could be used as a pivot point to get access into your environment. They’ve also gone deep in-depth with a bunch of other approaches to how DevOps and security are inextricably linked. To learn more, visit sysdig.com and tell them I sent you. That’s S-Y-S-D-I-G dot com. My thanks to them for their continued support of this ridiculous nonsense.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Today’s promoted episode is brought to us by 0x4447. And my guest today is David Gatti, their CEO. David, thank you for taking the time to speak with me today.

David: Thank you for getting me on the show.

Corey: One of the things that I find fascinating about what you do and where you come from is that for the last five years, you’ve been running an independent company that I would classify based upon our conversations as pretty close to a consultancy. However, you’ve gone down the path that I didn’t when I set up my own consultancy, and started actually selling software—not just software: Solutions—as a packaged thing that you can wind up doling out to various customers, whereas I just went with the very high touch approach of, “Oh, let me come in and have a whole series of conversations with people.” Your scale is a heck of a lot more. So, do you view yourself these days as a software company, as a consultancy, or something else entirely?

David: So, right now, I did put aside the consultancy because yeah, one thing that I realized, it’s possible but it’s very hard to scale, it’s also hard to find people at the same level. So yeah, the scalability of the business is quite hard, whereas with software sold on the AWS Marketplace, that is much easier to scale than what I was doing before, and that’s why I decided to take a break from consulting and focusing one hundred percent on the products that I sell on the AWS Marketplace to see how this goes and how it actually works, and can a business be built around it.

Corey: The common wisdom that I’ve encountered is that consulting, especially when you’re doing it yourself, is one of those things that is terrific when you find yourself in the position that I originally did of your employer showing up and, “Knock, knock,” “Who’s there?” “Not you anymore. Get out.” And there’s a somewhat, in my case, limited runway as far as how long I’ve got before I have to go find another job. With consulting, you can effectively go out and start talking to people, and provided that you can land a project, it starts throwing off revenue, basically immediately, whereas building software, building packages, things that you end up selling to people, it’s almost like a real estate business on some level, where you have to take a lot of investment up front to wind up building the thing, where—because no one is, generally speaking, going to pay you spec work to go ahead and build something for 18 months and come back and hope that it works.

David: Right.

Corey: I also bias towards the services because I’m bad at writing code. You, on the other hand, write things that seem to actually work, which is another refreshing difference.

David: Yes. So, I did that, but now I have a guy that is just a Linux expert. So, you were saying that there is a high investment in the beginning, but what actually—in my case what happened, I’ve been selling these products for the past three years basically as a hobby. So, when I was doing AWS consulting, I was seeing, like, a company has a problem, a repeating problem, so I was just creating a product, putting it on the Marketplace, and then sending it to them. So basically, they had a situation where I can manage those projects to update when there’s a need to do an update, and there was always a standardization behind that, right?

So, if they had, you know, five SFTP servers, and there was a need to make an update, I was making the update on my image, putting it on the Marketplace, and then updating all those servers in one go in a much quicker fashion then managing them one by one, right? And so I had this thing for three years. So now, when I started doing this full-time, I have a little bit of a leap on what’s going on. So, I already had a bunch of clients that are using their products, so that actually helped me not to have to wait three years before I saw any revenue coming in.

Corey: I always thought that the challenge behind building something like this was that well, you needed to actually be conversant in a programming language; that was the thing that you needed to package and build these things. But I take a look at what you have on the AWS Marketplace—and I will throw a link to this in the [show notes 00:04:39]—but you offer right now four different offerings: A Rsyslog server, a Samba server, VPN server, and an SFTP server, and every one of those four things, back in my DevOps days, I built and implemented on AWS, generally either from scratch or from something in the Marketplace—and I’ll get to that in a bit—that didn’t really meet a variety of needs. And every single time I built these things, it drove me up a wall because I had to do this without, like, solving a global problem locally, myself, to meet some pile of needs, then I had to worry about the maintenance of the thing, making sure that the care and feeding continued to work. And it just wasn’t—it didn’t work for me in the way that I wanted it to. It never occurred to me that I really could have just solved this whole thing once, [unintelligible 00:05:28] it on the Marketplace, and then just gone and grabbed the thing.

David: Exactly. So, that was my exact thinking here. Especially when your work with the client, this [unintelligible 00:05:38] was also great [idea 00:05:39] because when you work with clients, they want to do things as fast as possible, right? So, can they say, “I need an SFTP server?” Of course, it takes, you know, half a day to set up something, but then they scream at you and say, like, “Hey, do the next thing. Do the next thing. Do the next thing.” And you never end up configuring the server that you’re making a reliable way, sometimes you misconfigure it because, oh I forgot this option, and now everybody on the internet can access the server itself.

Corey: Wait, screw up a server config? That doesn’t sound like something I would do.

David: Well, of course not.

Corey: Yeah, no one [unintelligible 00:06:08] they’re going to until oops.

David: Yes. You’re amazing and you’re perfect, of course, but I’m not. And I was seeing, like, oh, you know, in the middle of the night, oh, I
forgot this option. I forgot this. I forgot that.

And so there was never a, basically, one place when the configuration just correct, right? And that was something that sparked my idea when I realized the Marketplace exists. It’s like, oh, wait a moment, I can spend few weeks to do it, right, put it there and never worry about it again. And so if when a client says like, “Hey, I need this,” I can deploy it literally, in less than one minute. You have any of those products that actually I’m selling up and running, right?

And of course, the VPN is going to be a little bit slower because it needs to generate all the certificates at the beginning, but for example, the SFTP one is just poof, you’re deployment with our CloudFormation file, provide username and password, and you’re up and running. And I see, for example, this thing with clients, which sometimes it’s funny, when there’s two clients that they use the SFTP server only once a day for one hour. So, every day is like one new instance created, then one instance removed, and one instance created and one instance removed. And so it keeps on going like that.

Corey: The thing that always drove me nuts about building these things out was first I had to go and find something on those rare occasions where I used the Marketplace. Again, I wasn’t really working in the same modern Marketplace that we think of today when we talk about the AWS Marketplace. It was very early on, the only way that it would deliver software was via, “Here’s an AMI, grab the thing, and go ahead and deploy it, and it’s going to have an additional hourly cost on. It the end.” And more or less the whole Henry Ford approach of, “Oh, you can get it in any color you want, as long as it’s black.”

So, back in those days, I would spin up an OpenVPN server—and I did this at several companies—I would go and find the thing on the Marketplace from I think it was the OpenVPN company behind the project. Great, I grabbed the thing, it had no additional cost through the Marketplace. I then had to go and get a custom license file from the vendor themselves, load the thing in, then start provisioning users. And this had no integration that I could discern with anything else we had going on, so all of this stuff was built through the web config on this thing, there was no facility for backing the thing up—certificate, material, et cetera, et cetera—so if something happened to that instance or that image, or we had to go through a DR exercise, well, time to reprovision everyone by hand again. And it was annoying because the money didn’t matter. At a company scale, it really doesn’t for something like this unless you’re into the usurious ranges. It does not matter.

It’s the, I want to manage this simply and effectively in a way that makes sense, and in many cases in a way that is congruent with our on-prem environment. So, “Oh, there’s a custom AWS service that offers something kind of like this. Use that instead.” It’s, yeah, I don’t like the idea, personally, of having to use a higher-level managed service that I’m very often going to need the most, right when things are getting wonky during an outage scenario. I want something that I understand and can work with.

And I’ve always liked, even if I have all the latest whiz-bang accesses into an environment, in production environments, I spin up something like this anyway, just to give myself a backdoor in the event that everything else breaks. And I really like how you’ve structured your VPN server as far as backing up its config, sharing its configs, you can scale it to more than one instance—what a ridiculous concept that is—and so on and so forth.

David: So, it’s not more than one—I mean, yes, you can deploy to more than one time, but the thing that—because again, when you were saying, like, companies don’t care about the cost, right? It’s more about how annoying it is to use and set up, right? And so I’m one of those people that when I, for example, see things like I’ve been playing with servers since the ’90s, right, and I was keeping rebuilding and recreating everything every single time from scratch.

And, yeah, it was always painful. It took always a lot of time. For example, our server took six months to set up the right way. And also the pricing [unintelligible 00:10:11] the competition has is quite aggravating, I will say. Like, it’s very hard to scale above a certain point, especially for the midsize companies.

And the goal with the Marketplace is also, like, make it as simple as possible. Because AWS itself doesn’t make it easy to be on the Marketplace, and it’s almost, like, crazy how hard it is. So, for anybody who will like to—who might think, like, “Oh, I would like to try this AWS Marketplace thing,” I would say should do it, but be super patient. You cannot rush it because it’s going to take you on average six months to understand how even the process of uploading anything and updating it and managing it is going to take it because their website that they’ve built has nothing to do with the console and it’s a completely custom solution that is very clunky and still very old-fashioned, how you have to manage it.

Corey: Tell me more about that. I’ve never gone through the process of putting something up on the Marketplace. To my understanding, you need to be an AWS partner in order to use the Marketplace, correct?

David: No you don’t have to.

Corey: Okay.

David: No. Thankfully not. I hope it’s not going to do this thing is not going to change. [crosstalk 00:11:20]—

Corey: Yeah. I wound up manifesting it into existence by saying that. Yeah. If you’re on the Marketplace team listening to this, don’t do
that, please. I really don’t want to get yelled at and have made things worse for people.

David: Don’t give them ideas. [laugh]. Okay?

Corey: Exactly.

David: No, it’s anybody can do it. But yeah, how to add a new product. So, the process is you have to build an AMI first. And then you have to submit the AMI to AWS by first creating a special AMI role—sorry, I always get confused AMI, [IAM 00:11:51], I never—IAM is
users. Okay.

Corey: I think we have a few more acronyms that use most of the same letters. I think that’s the right answer here.

David: [laugh]. So, either IAM or AMI, whichever is responsible for roles, you have to create a special role to give AWS access to your AMI. Then you submit the image to AWS providing the role that they have to use. They scan it and they do simple checks to make sure that you don’t for example, have SSH enabled with regular users, do some regular scanning to make sure that you’re not using an image from ten years ago, right, of Linux. And once you pass that, you are able to actually create your first product.

Then you have to write your title, description provide, for example, the ports that needs to be open, the URLs to separate resources, the pricing page, which takes on average one hour to fill up because let’s say that you have 20 instances that you support, and for every instance, you have to write the price for that instance per one hour. Then if you want to have a discount of let’s say 20%—because you can set it by the hour, or someone can pay you for the full year. And so for the full year, you might have a discount. So, you have to have also the price per hour discounted by the amount of percentage that you want, and then you have to repeat it 40 times. Because there is no way to upload that.

Corey: That feels like the internal AWS billing system in some respects. “Well, if it’s good enough for us it good enough for our customers.”
And—

David: [laugh]. Exactly.

Corey: —now, I have empathy for the folks in the billing system internally; their job is very hard, but that doesn’t mean that it’s okay to wind up exposing those sharp edges to folks who are, you know, paying customers of these things.

David: Right. And it’d be a simple thing like being able to import the CSV file with just two columns and that would be perfect. But no, you have to do it by hand. There is no other way. So hopefully—

Corey: Or someone has to. Welcome to the crappiest internship of your life.

David: Exactly.

Corey: It feels like bringing people into data entry for stuff like that is cheating.

David: Exactly. So, you do that and then I don’t remember exactly what the other steps are to a new creating a completely new product because I did that three years ago, and so now, I’m been just updating those products, but yeah, then they have to review your submission, and once everything is okay, then your product is on the Marketplace, and you can—are already accept everything. If you, for example, want to have the image also available in some specific regions that are not the default ones, you have to enable this by hand. I don’t remember anymore how, but it’s not obvious.

Corey: And you have to keep redoing this every time they launch a new region as well, I would imagine.

David: So, they say that you can have enabled the option to automatically add it, but it still won’t work. Well, it will work, but… let’s say, so in my case, I’m using CloudFormation. I gave a complimentary CloudFormation file where if you want to deploy my product, you go to the documentation page, you click the orange button, and you basically provide the parameters, and you click next, next, next and the product is deployed within a few minutes.

And in that CloudFormation file, I have a map of every AMI in every region. Okay? So, if they add a new region and they automatically add the AMI there, then if you don’t get notified that there is a new region, you don’t know that you have to update the CloudFormation file, and then someone might say, like, “Hey, David, why this product is not deployed in this region.” It’s like, “Oops. I didn’t know that they have to update the CloudFormation file with a new region.” Right?

Corey: Yeah, I’m a big believer in ClickOps, the idea of doing things in the console, but everything you’re talking about sounds like a fraught enough process that I’m guessing you have some form of automation that helps you with a lot of this.

David: Yeah. So, I hate repeating anything more than once, so everything in my book is automated as much as possible. The documentation, for example, how I structure it, there is a section that tells you how to deploy it by just using CloudFormation file and clicking next, next, next, next until you have it. And then there’s also the option if you want to deploy manually because you don’t trust what the CloudFormation file is doing, right? Of course, you can see the source file if you wanted to, but sometimes people are a little bit wary about big CloudFormation files.

In any case, I have this option, but they have this option as a separate thing. So, AWS has an option where you could add a CloudFormation file that goes with your product. The problem is to be able to submit a CloudFormation file natively so they will take care of it requires you to get Microsoft Office 365. Because they give you an Excel file that has, I think, a few thousand columns. And for example, numbers under [unintelligible 00:16:40], when you export, you save the final—or sorry, you export it, it will cut around 500 columns. So, you miss, like, two-thirds of what AWS will likely to send you. And why they do that, I have no idea. I don’t know if they still do it after three years, but when I was doing it, they told me like, “Hey, this is the file. Fill it by hand.”

Corey: About that time period, that was exactly how they did large-scale corporate discounts on custom contracts is that they would edit the AWS bill in Excel, or if not, the next closest thing to it because there were periodically errors that looked an awful lot like someone typo-ing something by hand.

David: What—

Corey: Computers are generally bad doing that, and it took an extra couple of weeks to get those bills, which is right around the speed of human.

David: Wow.

Corey: I see none of those problems anymore, which tells me, that’s right, someone finally upgraded off of Microsoft Excel to the new level. Probably Airtable.

David: [laugh]. Maybe. So, I don’t know if that process is still there, but what they did, like, then I realized, oh, wait a moment, I can just have a CloudFormation file in S3 bucket publicly available and just use that instead of going through that process. Because I didn’t want to pay on a yearly basis for a product that I’m going to use literally once a year. That didn’t make any sense to me and so I decided I’m going to do it this way. That’s why, yeah, if they add on a new region, I have to go out and update my own CloudFormation file because I maintain that myself, whereas they would maintain it for me, I guess.

Corey: The way that I see all of the nuts and bolts of the engineering parts of getting all these things up and running on the Marketplace, it feels like it is finicky; it is sharp edges that AWS is basically known for in many respects, but without the impetus of making that meaningfully better, just because there’s such an overriding business reason, that—it’s not like there’s a good competitor for something like this. So, if you want to sell things to AWS people in most frictionless way possible, it reflects on the AWS bill, causes discounting, counts for their spend commitments, and the rest, it’s really the AWS Marketplace is the only game in town for a lot of that.

David: Right. So, I don’t know if they don’t do it because they don’t have enough competition or pressure because to me when I first started doing this AWS Marketplace, it felt to me like more Amazon than AWS, right? It feels more like an Amazon team was behind it and not people from AWS itself. It felt like completely something different. Not to mention, yeah, the console that they provide is something completely custom that has nothing to do with the typical AWS console.

Corey: I’ve heard stories about the underpants store division’s seller tools as well; very similar to the experience you’re describing.

David: Mmm. And also the support is different. So, it’s not connected to the AWS console one. The good thing about it, it’s free, but it’s also only by email. And so yeah, it’s a very weird, clunky situation where I mean, I’m someone that, I guess, loves the pain of AWS. [laugh].

I don’t know if that’s a good thing or a bad thing. But when I started, I decided, you know what, I’m going to figure it out, and once I do, I’m going to feel happy that I was able to. Maybe that’s their goal: It’s to give us purpose in life. So, maybe that’s the goal of AWS. I don’t know.

Corey: There are times I really wonder about that where it feels like it could be so much more than it is, but it’s not. And, again, my experience with it is very similar to what you’ve described, where it’s buying an AMI, the end. But now they’re talking about selling SaaS subscriptions on it, they’re talking about selling professional services—in some cases—on it. And effectively, it almost feels like it’s trying to become the Marketplace through which all IT transacting starts to happen. And the tailwind that sort of is giving energy to a lot of those efforts is, if you have a multimillion-dollar spend commitment with AWS in return for discounting, you have to make sure you spend enough within the timeframe, 50% of all spend on the AWS Marketplace counts toward that.

Now, other cloud providers, it’s 100% of spend, but you know, AWS is nothing if not very tight with the dollar. So okay, fine, whatever. There’s a reason for companies to go down that path. Talk to me a little bit about the business aspect of it because for me, it seems like the clear win, in the absence of anything else is—especially at larger companies—they already have a business relationship with AWS. The value to someone selling software on the Marketplace feels like it would be, first and foremost, an end-run around companies procurement departments.

It’s just oh, someone has to click a button and they’re up and running, as opposed to going through the entire onboarding and contracting and all the rest, manual way. Other than the technical challenges of getting things up and running on it, how have you found that it works as far as getting in front of additional customers, as far as driving adoption? You could theoretically have—I imagine—have not gone down the Marketplace route at all and just sold this directly on your website, click here to buy a license file the way that a lot of stuff I used to as well, and would have cut out a lot of the painful building an AMI and putting it into the Marketplace story. What’s the value to being in the Marketplace?

David: Yeah, so in the beginning, the value was basically that it’s on the Marketplace, as I was saying, I was using it with pre-existing clients, so it was easy for me because I knew AWS images were there. So, it was easy to just click my own CloudFormation file and tell the client after one minute, “Hey, it’s up and running. You have a bunch of profiles for your VPN. Enjoy and have fun.” Right?

That experience, once you have it on the Marketplace, it’s nice because it just works. And you don’t have to do much work. Then I realized that AWS, in the search bar in the console, when you were typing, for example, you know, you type EC2, S3, CloudFormation, to find the service, what they were doing originally is when you were typing in the search bar, you were getting the services of AWS, and then when there was nothing left, they were showing the results of the Marketplace, which was basically amazing because you have primetime in the console with your product, you had to do zero marketing, and you get every week, took new clients that are using our product. And the trend was growing pretty, pretty well.

And that was a proposition that is just amazing. Like, nobody has that because you can have Fortune 500 companies using our product without doing anything. It just—is it simple to deploy? Yes. Does it provide value? Is the price great? And people were just using them. Fast forward now; what happened is AWS changed the console. And instead of showing, after the services, the Marketplace, like, now they show the sub-section of the services, they show the results from the blog, the articles, videos, whatever, I don’t even know what they’ve put there—

Corey: Originally, you could search my name in that search bar, and it would pop up a profile of me they did for re:Inforce in the security blog.

David: [laugh]. There you go.

Corey: “Meet Corey Quinn. A ‘cloud economist’—scare quotes and all—who does not work here. And it was glorious. Now, they’ve changed the algorithm so it pops up. “Oh, you want Corey Quinn, you must mean IoT Core.” So, that blog post is still there, but it’s below the fold because of course they give precedence to a service that they have that nobody uses or understands. Because, Amazon.

David: Yeah, of course. And so that was awful because suddenly I realized that, oh, I’m getting less and less new clients because you know, after six months, one year, people are shutting off their things because they’re finished using them, and I will not getting new ones. But at that time, I was doing [AWS 00:24:06] consulting, so it’s like, oh, maybe it was a glitch in the Matrix, whatever. I got lucky.

But then after a few months, I realized, wait a moment. When I was working in AWS, I realized that the console results changed, and I went like, oh, that’s what happened and that’s why I’m getting less clients, right? So, in the beginning, that was a great thing and that’s why I’m actually paying you to promote my business and my products because now there is no way to put the products in front of customers because AWS took it away. And so that’s why I decided to actually go full-force on this to make sure that I promote as much as possible because that one cool feature that AWS was providing, they took it away for whatever reason because blog posts are more important than their partners, [laugh] I guess.

Corey: Well, it depends on the partner and the tier of partner, and it feels like it’s a matter—to be clear, full disclosure: I am not an AWS partner; I’m not partnered with any vendor in this space, for either real or perceived conflict of interest issues, so I don’t have a particular horse in the race. But back when there were a small number of partners, the network really worked. Now, there are tens of thousands of partners, and well, what winds up being surfaced? Customers, as a result seem to be caring less about various partner statuses, unless they’re trying to check a box on some contractual requirement. Instead, they just want the problem solved, and it’s becoming increasingly challenging to differentiate just by the nature of how this works.

I don’t believe, in 2022, that you could build almost anything, and put it on the AWS Marketplace in isolation and expect that to suddenly drive adoption by the fact that you’re there. It feels, to me, at least on the other side of the fence, that the Marketplace experience is all about, you go there and you look for the name of the thing that you already know that you want because you’ve heard about it from other means, and then you just click it and you go, and that’s the end of it. It’s a procurement story; it’s not a discoverability story.

David: Right. And yeah, so that’s sort of a bit disappointing, and I even made a post on Reddit about it to just bring this up to AWS itself to say, like, “Hey, UI change is pretty severe.” Because I mean, they get a percentage of every hour, the products are running, so basically they shoot themselves in the foot by making less money because now they’re getting less products are being shown to potential customers. So, yeah, that’s a disappointing thing.

When it comes to also you ask what other way there is to show their products to potential customers, so there is an option where AWS can help you out. And when I talked to them, I think last year, they said that if you reach $2 million in sales a year, then they will basically show you around other potential customers, right? Which is a little bit disappointing because especially if you’re a small company like mine, it’s pretty hard to get to that $2 million in a meaningful time. And if once you reach that point, you might go like, “Hmm, how is this going to help me if you now show me in front of other people?” So yeah.

And of course, I understand them in a sense that if they show a product from the Marketplace to a big company and the product turns out to be of poor quality, then of course the client is going to tell AWS like, “Why you’re showing us something that just doesn’t do its job?” Right? But it’d be nice to have a [unintelligible 00:27:24] when you say, “Okay, you’re starting out. After a few years, so we can show you to this midsize clients.” You don’t have to go to, immediately, Fortune 500 companies. That doesn’t make any sense, right?

Corey: And I still—even the companies that are at that level, I’ve talked to them about how they’ve grown their business, and not a single one has ever credited anything AWS did to help them grow. Other than, “Well, they threw re:Invent, so we spent extortionate piles of money and set up a booth there, and the fact that we were allowed in the building to talk to people was helpful, I guess.” But it’s all through their own works on this, I’m not convinced, to be very direct with you, that AWS knows how to effectively drive sales and adoption of things on their own Marketplace. That is an increasing source of concern.

David: Right. And then there’s no plan of what to do with a company that is starting on the Marketplace, once it’s a few—or it’s already a few years and established in the Marketplace and a big one. Yeah, they don’t have any way to go about it, which is a bit disappointing. But again, I like a challenge. I like the misery of AWS, so I’m just doing it. [laugh].

Corey: No, I hear you. Would you recommend other people in your position explore selling on the Marketplace, given the challenges and advantages both that you’ve experienced?

David: So, if you were to start from scratch, it will take you, like, three years—maybe not three years, but it’s not something that should be the primary revenue source of the business if you want to go into the AWS Marketplace situation because you have to have enough capital to do enough marketing to see if you can get in front of people. If you already do some consulting like me, where I did some stuff on the side, and then realized, oh, people are using it, people like it, they get some feedback, the want new features, like, “Oh, maybe I can start growing this bigger and bigger, right?” It’s not something that’s going to happen immediately. And especially the updating process that happens, it can get quite stressful because when you make an update—so you have a version of a product that’s working and running, right? Now, you make an update and you have to spend at least a week or even sometimes two weeks to test that out to make sure that you didn’t miss anything because you don’t want people to update something and it stops working right?

Corey: You can’t break customer experiences on these things.

David: Yeah. No.

Corey: It becomes a nightmare.

David: Because especially you don’t know if, literally, a Fortune 500 company is using your product or, like, a tiny company that has only ten employees, right?

Corey: Your update broke the file server with a VPN means it’s unlikely that they’re going to come back anytime soon, too.

David: Right.

Corey: You’re also depending on AWS, in some respects, to steward the relationship because you’re you don’t have direct contact with your buyers.

David: No. So, that’s important thing. They don’t give you access to the contacts; they give you access to the company information. So, I actually do have Fortune 500 companies using my products, but yeah, there’s no way to get in touch with them. The only thing that you get is the company name, the address, the domain that they used to create an email. So, at least you can get a sense of, like, who this
company is.

But yeah, there is no way to get in touch if there is a problem. So, the only way that you can notify the customer that there’s a new update is when you make an update, there is a text area that you can say what’s new, what did you change, right? And that’s the only communication that you get with the client. So if, for example, you do a big mistake, [laugh], you basically have that just little text box, and hopefully, someone reads it. But you know, AWS is known for sending 20 emails a week for every account that you open. Good luck getting through that noise.

Corey: Hope that you don’t miss the important ones as you go through. No—

David: Exactly.

Corey: —I hear you. These are problems that I think are on AWS’s plate to solve. Hopefully, someone over there is listening to this and will at least reach out with a bit of a better story. I really want to thank you for taking the time to speak with me today. We’ll include links, of course, to this in the [show notes 00:31:09]. Where else can people find you?

David: They can find us basically on the product page of what we sell. So, we have products.0x4447.com/. That’s where, basically, we keep all our products. We keep updating the page to provide more information about those products, how to get in touch with us, we provide training, demos, anything that you want. It’s very easy to get in touch with us instead of—sometimes when it comes to AWS. So yeah, we are out there, pretty easy to find us. The domain—the company name is so unique that you either get our website or—

Corey: Easy to find on Google.

David: Yeah, so we’re basically—the hex editor. And that’s basically it. [laugh].

Corey: Excellent. Well, we’ll definitely put links to that in the show [notes 00:31:50]. Thank you so much for taking the time to speak with me today. I really appreciate it.

David: Thank you very much.

Corey: David Gatti, CEO of 0x4447. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry comment that makes sure to mention exactly how long you’ve been working on the AWS Marketplace team.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Seth

Seth Vargo is an engineer at Google. Previously he worked at HashiCorp, Chef Software, CustomInk, and some Pittsburgh-based startups. He is the author of Learning Chef and is passionate about reducing inequality in technology. When he is not writing, working on open source, teaching, or speaking at conferences, Seth advises non-profits.

Links:

  • Twitter: https://twitter.com/sethvargo

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: The company 0x4447 builds products to increase standardization and security in AWS organizations. They do this with automated pipelines that use well-structured projects to create secure, easy-to-maintain and fail-tolerant solutions, one of which is their VPN product built on top of the popular OpenVPN project which has no license restrictions; you are only limited by the network card in the instance.

Corey: Couchbase Capella Database-as-a-Service is flexible, full-featured and fully managed with built in access via key-value, SQL, and full-text search. Flexible JSON documents aligned to your applications and workloads. Build faster with blazing fast in-memory performance and automated replication and scaling while reducing cost. Capella has the best price performance of any fully managed document database. Visit couchbase.com/screaminginthecloud to try Capella today for free and be up and running in three minutes with no credit card required. Couchbase Capella: make your data sing.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I have a return guest today, though it barely feels like it qualifies because Seth Vargo was guest number three on this podcast. I’ve had a couple of folks on since then, and for better or worse, I’m no longer quite as scared of the microphone as I was back in those early days. Seth, thank you for joining me.

Seth: Yeah, thank you so much for having me back, Corey. Really excited to figure out whatever we’re talking about today.

Corey: Well, let’s start there because last time we spoke, you were if memory serves a developer advocate at Google Cloud.

Seth: Correct.

Corey: And you’ve changed jobs, but not companies—but kind of companies because, welcome to large environments—but over the past few years, you have remained at Google. You are no longer at Google Cloud and you’re no longer a developer advocate. In fact, your title is simply ‘Engineer at Google.’ And what you’ve been focusing on, to my understanding, is helping Alphabet companies, namely—you know, the Alphabet, always in parentheses in journalistic styles, Google’s parent company because no one thinks of it in terms of Alphabet—is—you’re effectively helping companies within the conglomerate umbrella securely and privately consume public cloud.

Seth: Yes, that is correct. So, I used to work in what we call the Cloud PA—PA stands for product area. Other product areas are like Chrome and Android—and I moved to the Core PA where I’m helping lead and run an initiative that, like you said, is to help Alphabet companies to, you know, securely and privately use public cloud services.

Corey: So, I am going to go out on a limb because my position on multi-cloud has always been pick a cloud—I don’t particularly care which one—but pick one and focus on that. I’m going to go out on a limb and presume that given that you are not at Google Cloud anymore, but you are at Google, you probably have a slight preference as far as which public cloud these various companies within the umbrella should be consuming.

Seth: Yeah. I mean, obviously, I think most viewers will think the answer is GCP. And if you said GCP, you would be, like, 95% correct.

Corey: Well, you’d also be slightly less than that correct, because they’re doing a whole rebrand and calling it Google Cloud in public, as opposed to GCP. You really don’t work for the same org anymore. You’re not up-to-date on the very latest messaging talking points.

Seth: I missed—ugh, there’s so many TLAs that you lose all your TLAs over time.

Corey: Oh, yes.

Seth: So, Google Cloud would be, like, 95% correct. But what you have to really understand is, Google has its own, you know, cloud—we didn’t call it a cloud at the time, you might call it on-prem or legacy infrastructure, if you will—primarily built on a scheduling system called Borg, which is like Kubernetes version zero. And a lot of the Alphabet companies have workloads that run onboard. So, we’re actually talking about hybrid cloud here, which, you know, you may not think of Google is like a hybrid cloud customer, but a workload that runs on our production infrastructure called Borg that needs to interact with a workload that runs on Google Cloud, that is hybrid cloud, it’s no different than a customer who has their own data center that needs peering to a public cloud provider, you know, whether that’s Google Cloud, or AWS, or Azure.

I think the other thing is if you look at, like, the regulatory space, particularly a lot of the Alphabet companies operate in, say, like healthcare, or finance, or FinTech, where certain countries and certain jurisdictions have regulations around, like, you must be multi-cloud. You know, some people might say that means you have to run, you know, the same instance of the same app across clouds, or some people say your data can be here, but your workloads can be over there. That’s to be interpreted, but you know, I would say 95% of GCP, but there is a—or sorry, 95% is Google Cloud—

Corey: There we go.

Seth: But there is a small percentage that is definitely going to be other cloud providers and hybrid cloud as well.

Corey: My position on multi-cloud has often—people like to throw it in my face of, “See you gave this general guidance, and therefore whenever you say something that goes against it, you’re a giant phony.” And it’s yeah, Twitter doesn’t do so well with the nuance. My position of pick a provider and go all-in is intended as general guidance for the common case. There are exceptions to this and any individual company or customer is going to have more context than that general guidance will. So, if you say you need to be in multiple clouds for certain reasons, you’re probably correct.

If you say you need to be in multiple clouds because your regulator demands it, you are certainly correct. I am not arguing against that in any way. I do want to disclaim my one of my biases here as well, and that is specifically that if I were building a startup today and I were not me—by which I mean having spent ten years in the AWS ecosystem learning, not just how it works, but how it breaks because that’s important in production, and you know, also having a bunch of service owners at AWS on speed dial—and I, were approaching this from the naive, I need to pick a cloud, which one would I go with, my bias is for Google Cloud. And the reason behind that is the developer experience is spectacular as the primary but not only perspective on that. So, I am curious to know that as you’re helping what are effectively internal customers move to Google Cloud, is their interaction with Google Cloud as a platform the same as it would be if I as a random outside customer, were using Google Cloud? Is there a bunch of internal backchannels? “Oh, you get the good kind of internal Google Cloud that most of us don’t get access to?” Or something else?

Seth: Yeah, so that’s a great question. So first, you know, thank you for the kind words on the developer experience—

Corey: They were honest words, to be clear. Let me be very direct with you, if I thought your developer experience was trash, I might not say it outright in their effort not to be, you know, actively antagonistic to someone I’m having on the show right now, but I would not say it if I didn’t believe it.

Seth: Yeah. And I totally—I know you, I’ve known you for many years. I totally believe you. But I do thank you for saying that because that was the team that I was on before this was largely responsible for that across the platform. But back to your original question around, like, what does the support experience look like? So, it’s a little bit of both.

So, Alphabet companies, they get a technical account manager, very similar to how, you know, reasonable-sized spend customer would get a technical account manager. That account manager has access to the Cloud support channels. So, all that looks the same. I think we’re things look a little bit different is because myself and some of our other leads came from Cloud, you know, I generally don’t like this phrase, but we know people. So, we tend not to go directly to Cloud when we can, right?

We want Alphabet companies to really behave and act as if they were an external entity, but we’re able to help the technical account manager navigate the support process a little bit better by saying like, “You need to ask for this person,” right? You need to say these words to get in front of the right person to get this ticket assigned to the right person. So, the process is still the same, but we’re able to leverage our pre-existing knowledge with Cloud. The same way, if you had a [unintelligible 00:07:45] or an ex-Googler who worked for your company, would be able to kind of help move that support process along a little bit faster.

Corey: I am quite sincere when I say that this is a problem that goes far beyond simply Google. A disturbing portion of my job as a cloud economist helping my clients consists of nothing other than introducing Amazonians to one another. And these are hard problems at scale. I work at a company with a dozen people in it. And it turns out that yeah, it’s pretty easy to navigate who’s responsible for what. When you have a hyperscale-size company in the trillion-dollar range, a lot of that breaks down super quickly.

Seth: And there’s just a lot of churn at all levels of the organization. And, you know, we talked about this when I first joined the show, like, I switched roles, I used to be in Cloud, and now I’m in what we call Core. I still get people who are reaching out to me, at Google and externally, who are saying, “Oh, can you answer this question? Hey, how do I do this?” And I, you know, I’ve gradually over the past couple of months, you know, convinced people that I don’t work on that anymore, and I try to be helpful where I can, but the—

Corey: You use the old name and everything. They’re eventually going to learn, right?

Seth: I know. They’ll be like, “What do you call this? GCP? Okay, great. We don’t need you anymore.” But it’s true, right? Like, there’s people leave the organization, people join the organization, there’s reorgs, there’s strategic changes, people, you know, switch roles
within the org, and all of that leads to complexity with, you know, navigating, what is the size of a small nation, in some cases.

Corey: Your line in your biography says that you enable Alphabet companies to securely and privately consume public cloud. Now, that would make perfect sense and I would really have no further questions based on what we’ve already said, except for the words securely and privately, and I want to dive into that, first. Let’s work backwards with the second one first. What is ‘privately’ mean in this context?

Seth: So, privately means, like, privacy-preserving for both the Alphabet company and the users or customers that they have. So, when we look at that from the perspective of the Alphabet company, that means protecting their data from the eyes of the cloud provider. So, that’s things like customer-managed encryption keys, you know, bring-your-own-encryption, that’s making sure that you have things like, actually, transparency so that if at any point the cloud provider is accessing your data, even for a legitimate purpose, like submitting a support ticket or something—or diagnosing a support ticket, that you have visibility into that. Then the privacy-preserving side on the Alphabet company’s customers is about providing that same level of visibility to their customers as well as making sure that any data that they’re storing is, you know, private, it’s not accessible to certain parties, it’s following whether it’s like, you know, actual legislation around how long data can be persisted, things like GDPR, or if it’s just a general, like, data retention, insider risk management, all of that comes into this idea of, like, building a private system or privacy-preserving system.

Corey: Let’s be very clear that my position on it is that Google’s relationship with privacy has been somewhat challenged, in due to no small part to the sheer scale of how large Google has grown. And let’s be clear, I believe firmly that at certain points of scale, yeah, you deserve elevated levels of scrutiny. That is how we want society to function, by and large. And there are times where it feels a little odd on the cloud side. For example, as the time is recording, somewhat recently, there was a bug in some of the copyright detection stuff where Google Drive would start flagging files as having copyright challenges if they contained just the character ‘1’ in them.

Which, okay, clearly a bug, but it was a bit of a reminder for some folks that wait, but that’s right, Google does tend to scan these things. Well, when you have a bunch of end-user customers and in the ways that Google does, that stuff is baked in and it shapes how you wind up seeing things. From Amazon’s perspective, historically, they basically sold books and then later underpants. And doing e-commerce transactions was basically the extent of their data work with customers. They weren’t really running large-scale, file sharing systems and abilities—in collaboration suites, at least not that really had any of those pesky things called customers.

So, that is not built into their approach and their needs in the same way. To be clear, I am sympathetic to the problems, but it’s also… it’s a challenging problem, especially as you continue to evolve and move things into cloud, you absolutely must be able to trust your cloud
provider, or you should not be working on that cloud provider, has been my approach.

Seth: Yeah, I mean, there’s certainly things that you can do to mitigate. But in general, like, there is some level of trust, forget the data, on the availability side, right? Like when the cloud provider says, “This is our SLA.” And you agree to that SLA, like, yeah, you get money back if they mess it up, but ultimately, you’re trusting them to adhere to that SLA, right? And you get recompense if they fail to do so, but that’s still, like, trust—trust is far more than just on the privacy side, right? It’s on… the promise on the roadmap, it’s on privacy, it’s on the SLA, right?

Corey: Yeah. And you see that concern expressed more articulately from enterprise customers, when there’s a matter of trusting companies to do what they say, such as the continued investment that Alphabet slash Google is making in Google Cloud. It’s easy to take the approach of well, you’ve turned off a bunch of consumer services, so therefore, you’re going to turn off the cloud at some point, too. No, let me be very clear, for the record, I do not believe that you are going to one day flip a switch and turn off Google Cloud. And neither do your customers.

Instead, the approach, the way that enterprises express this, it’s not about you flipping the switch and turning it off—that’s what contracts are for—their question, and they enshrine this in contracts, in some cases, in the event, not that you turn it off, but that you fail to appropriately continue to invest in the platform. Because at enterprise scale, this is how things tend to die. It is not through flipping a switch, in most cases, it’s through, “We’re just going to basically mothball it, keep it more or less exactly as it is until it slowly fades into irrelevance for a long period of time.” And when you’re providing the infrastructure to run things for serious institutions, that part isn’t okay. And credit where due, I have seen every indication that Google means it when they say this is an area of strategic and continued ongoing focus for us as a company.

Seth: Yeah, I mean, Google is heavily investing in cloud. I mean, this is a brand new group that I’m working in and we’re trying to get Alphabet companies onto cloud, so obviously there’s some very high-level top-down executive support for this. I will say that the—a hundred percent agree with everything you’re saying—the traditional enterprise approach of build this Java app—because let’s be honest, it’s always Java—build this Java app, compile it into a JAR and run it forever is becoming problematic. We saw this recently with, like, the log4j—

Corey: Yeah, to be in a container. What the hell?

Seth: [laugh].

Corey: I’m kidding. I’m kidding. Please don’t send me email, whatever you do.

Seth: What’s a container? I’m just kidding. Like, the idea of, like, software rotting is very real and it’s becoming more and more of a risk to security, to privacy, to public cloud providers, to enterprises, where when you see something like log4j happen and you can’t answer the question, like, do we have any code that uses that? Like, if getting the answer to that question takes you six weeks, [sigh] boy like, a lot of stuff can happen in six weeks while that particular thing is exploited. And you know, kind of gets into software supply chain a little bit, but I do agree that, like, secure, private, and stable APIs are super important, and it’s an area where Google is investing. At the same time, I think the industry is moving, the enterprise industry is moving away a little bit from set-it-and-forget-it as a strategy.

Corey: I want to talk about the security portion as well as far as securely consuming public cloud goes. And let me start off with a disclaimer here because I don’t want people to misconstrue what I’m about to say. If you are migrating to one of the big three cloud providers, their security will be better than anything you will be able to achieve as a company yourself. Not you personally because Google is a bit of an asterisk to that statement, given what you have been doing and have been doing since the ’90s in your on-prem world with Borg and the rest, but my philosophy on the relative positioning of the security of cloud providers relative to one another has changed. I spent four months beating the crap out of Azure forever having an issue where there was control plane access and then really saying nothing about it.

And after I wound up finding—the day after I put out a blog post on that topic because I was tired of the lack of response, it came out that right at the same time AWS had a very similar problem and had not said anything themselves. And they went back and forth, apparently waiting to wind up doing a release until this happened, Orca Security wound up putting one out there, and it was frustrating on a couple of levels. First, the people at both of these companies who work in security are stars. There is no argument, no bones about that. Problems are going to happen, things are going to occur as a result, and the only saving grace then is the transparency and communication around it, and there was none of it from them.

I’m also more than a little bit irked that my friends at AWS were aware of this, basically watched me drag Azure for four months knowing that they’d done the same thing and never bothered to say a word. But okay, that’s a choice. I’ve been saying for a while that of the big three, Google’s security posture is the most impressive. And it used to be a slight difference. Like, you nosed ahead of AWS in that respect, not by a huge margin, but by a bit.

I don’t think it’s nearly as close these days, in my mind, and talking to other large companies about these things, and people who are paid to worry about these things all day long, I am very far from alone in that perspective. So, I guess my question for you is, as you look at moving the workload securely to Google Cloud, it feels like security is baked into everything that all aspects of your company have done. Why is that a specific area of focus? Or is that how it gets baked into everything you folks do?

Seth: So, you kind of like set up the answer for this perfectly. I swear we didn’t talk about this extensively beforehand.

Corey: You didn’t know any of that was coming, by the way, just to be very clear here. I don’t sit here and feed, “All right, I’m going to say this. And here’s the right res—” No, this is an impromptu, more or less ad hoc show every time I do it.

Seth: Yeah. And I’m going to preface this by saying, like, I don’t want this to sound, like, egotistical, but I have never found a company that has as rigorous security and privacy policies, reviews, and procedures as Google.

Corey: I thought I had and I was wrong.

Seth: Yeah. And—

Corey: And I have a lot of apologizing to people to do as a result of that.

Seth: And honestly, every time I interact with our internal security engineering teams, or our IP protection teams, I’m that Nathan Fillion meme, where he’s like, what—you know, like, “Okay, I get it. I get it.” Right?

Corey: And then facepalm it, uh, I should say some—I can’t—yeah. Oh, yeah.

Seth: The reason that it’s hard for Alphabet companies to securely and privately move to cloud specifically for security, is because Alphabet’s stance is so much more rigorous than anyone else in the industry, to the point where, in some cases, even our own cloud provider doesn’t meet the bar for what we require for an internal workload. And that’s really what it comes down to is, like, the reason that Google is the most secure cloud is because our bar is so high that sometimes we can’t even meet it.

Corey: I have to assume that the correct answer on this is that you then wind up talking to those product teams and figure out how to get them to a point where they can support that bar because the alternative is effectively, it’s like, “Oh, yeah, this is Google Cloud and it’s absolutely right for multinational banks to use, but you know, not Google workloads. That stuff’s important.” And I don’t think that is necessarily how you folks tend to view these things.

Seth: So, it’s a bidirectional stream, right? So, a lot of it is working with a product management team to figure out where we can add these additional security properties into the system—I should say, tri-directional. The second area is where the policy is so specific to Google that Google should actually build its own layer on top of it that adds the security because it’s not generally applicable to even big, huge cloud customers. And then the third area is Google’s a very big company. Sometimes we didn’t write stuff down, and sometimes we have policies where no one can really articulate where that policy came from.

And something that’s new with this approach that we’re taking now is, like, we’re actually trying to figure out where that policy came from, and get at the impetus of what it was trying to protect against and make sure that it’s still applicable. And I don’t know if you’ve ever worked with governments or you know, large companies, right, they have this spreadsheet of hundreds of thousands of lines—

Corey: You are basically describing my client list. Please continue.

Seth: I mean, like, sometimes they have to use an Access database because they exhaust the number of rows in an Excel spreadsheet. And it’s just checklist upon checklist upon checklist. And that’s not how Google does security, right? Security is a very all-encompassing, kind of, 360 type of thing. But we do have policies that are difficult to articulate what they’re actually protecting against, and we are constantly re-evaluating those, and saying, like, “This made sense on Borg. Does it actually make sense on Cloud?” And in some cases, it may not. We get the same protections using, say, a GCP-native service, and we can omit that requirement for this particular workload.

Corey: This episode is sponsored by our friends at Oracle Cloud. Counting the pennies, but still dreaming of deploying apps instead of “Hello, World” demos? Allow me to introduce you to Oracle’s Always Free tier. It provides over 20 free services and infrastructure, networking, databases, observability, management, and security. And—let me be clear here—it’s actually free. There’s no surprise billing until you intentionally and proactively upgrade your account. This means you can provision a virtual machine instance or spin up an autonomous database that manages itself, all while gaining the networking, load balancing, and storage resources that somehow never quite make it into most free tiers needed to support the application that you want to build. With Always Free, you can do things like run small-scale applications or do proof-of-concept testing without spending a dime. You know that I always like to put asterisks next to the
word free? This is actually free, no asterisk. Start now. Visit snark.cloud/oci-free that’s snark.cloud/oci-free.

Corey: I think that when it comes to things like policies that are intelligently crafted around security, you folks—and to be fair, the AWS security engineers as well—have been doing it right in that, okay, we’re going to build a security control to make sure that a thing can’t happen. That’s not enough. Then there’s the defense-in-depth. Okay, let’s say that control fails for some variety of ways. Here are the other things we’re going to do to prevent cross-account access, for example.

And that in turn, winds up continuing to feed on itself and build into a culture of assuming that you can always continue to invest in security. How far is enough? Well, for most folks, they haven’t gone far enough yet.

Seth: Another way to put this is like, how well do you want to sleep at night? You know, there’s folks on the Google security engineering team who are so smart, and they work on, like, our offensive security team, so their full-time job is to try to hack Google and then figure out how to prevent that. And, you know, so I’ve read some of the reports and some of the ways they think and I’m like, “How do you… how do you pick up a mobile phone and go to like, any website confidently knowing what you know?” Right? [laugh] and like, how do you—

Corey: Who said anything about confidently? Yeah.

Seth: Yeah. Yeah. How do you use self-checkout at a supermarket and, like, not just, like, wear your entire full-body tinfoil hat suit? But you know, I think the bigger risk is not knowing what the risks are. And this is a lot what we’re seeing in software supply chain, too, is a lot of security is around threat modeling and not checklists. But we tend to, like, gravitate toward checklists because they’re concrete.

But you really have to ask yourself, like, do I need the same security properties on my static blog website that is stored on an S3 bucket or a GCS bucket that’s public to the internet, that I do on my credit card processing service? And a lot of times we don’t treat those differently, we don’t apply a different threat model to them, and then everything has to have the same level of security.

Corey: And then everything is in-scope for whatever it is you’re trying to defend against. And that is a short path to madness.

Seth: Yes. Yes. Your static HTML files and your GCS bucket are in scope for SOC 1 and 2 because you didn’t have a way to say they weren’t.

Corey: Yeah. You’ve also done some—again, the nice thing about being at a company for a while—from what I can tell, given that I’ve never done until I started this place—is you move around and work on different projects. You were involved as well, personally, in the exposure notifications project, the joint collaboration thing between a number of companies in the somewhat early days of the pandemic that all of our phones talk to one another and anonymously and in a privacy-preserving way, let us know that hey, by the way, someone you were in close contact with has tested positive for Covid 19 in the previous fixed period of time. What did do you do over there?

Seth: Yeah, so the exposure notifications project was a joint effort, primarily between Apple and Google to use Android and iOS devices to help stop the spread of Covid or reduce the spread of Covid as much as possible. The idea being because the incubation period is roughly 14 days, at least pre-Omicron, if we could tell you hey, you might have been exposed and get you to stay at home for three or four days, self-isolate, we could dramatically reduce the spread of Covid. And we know from some of the studies that have come out of, like, the UK and European region that, like, the technology actually reduced the spread of cases by, like, fourteen-hundred percent in some cases. I was one of the tech leads for the server-side. So, the way the system works is it uses the low-energy Bluetooth on iOS and Android devices to basically broadcast random IDs.

So, I know this is Screaming into the Cloud, but if we can just quickly Screaming into the Void as a rebrand—

Corey: Oh, yeah.

Seth: —that’s basically what’s happening. [laugh]. You’re generating these random identifiers, and just, like, yelling them, and there’s other phones out there who are listening. And they collect these we’ll call RPIs—or Rolling Indicators. They have no data in them.

They’re like literally, like, a UUID or 32 bytes of random data, they aren’t at all, like, associated with your device or your person. So, then what happens is, like, let’s say you’re in a supermarket, you’re near someone for, you know, every so often, and your phones exchange these IDs. If you then test positive, those IDs go up to a centralized server, the server again, also has no idea who you are, so the whole thing is privacy-preserving, end-to-end, then the server basically bundles all of what we call the TEKs, or the Temporary Exposure Keys—into a tarball that go up onto a CDN, and then every night, all of the devices that are participating in EN download this into a local key match. So, at no point does the server ever know that you were in a supermarket with someone else, only your phone knows that you came in contact with this TEK in the past 14 days—or 21 days in some jurisdictions—and it’ll generate an exposure notification or an exposure alert, which says, like, “Hey, in the past 14 days, you’ve come in contact with someone who’s confirmed positive for Covid.” And then there’s guidance kind of varies by state and by health jurisdiction of, like, self-isolate, or go get tested, or whatever. But the idea—

Corey: Or go to the bar in some places, apparently.

Seth: Oh. Yeah. The server itself is actually—there’s a verification component because ideally, like, we don’t want people to just be like, oh, I’m Covid positive, and then like, all their friends get an alert, right? There needs to be some kind of verification mechanism where you either have a positive test, or you have a clinician or a physician who issues you code that you can put into your app so you can then release your keys. And then there’s the actual key server component, which I kind of already described.

So, it’s a pretty complex system and actually is entirely serverless. So, the whole thing, including all, like, background job processing, it was designed to be serverless from the beginning. Total greenfield project, right, like, nothing like this exists, so we’re really fortunate there. We made some fun and interesting design decisions to keep costs down while, you know, abusing slash using some of the features of serverless like auto-scaling and, you know, being able to fan out across multiple regions and things like that—

Corey: And using DNS as a database. My personal favorite approach to things?

Seth: We don’t use DNS as a database. We do use Postgres—

Corey: A missed opportunity.

Seth: —a real database. But we do use DNS, just not for storing information.

Corey: So, one question I have for you is that you’ve been at Google for a while and you’ve done an awful lot of things there, but previously, you’ve also done things that don’t really directly aligne any of this stuff going on there. You were at HashiCorp and you were at Chef, neither of whom, to my understanding are technologies that Google makes extensive use of internally for their own stuff. It seems like—and even when you’re at Google, you have been continually reinventing what it is that you do. I find that admirable because very often, when you see people at a company for a protracted period of time, they sort of get more or less pigeonholed into the role that looks fairly similar from year-to-year. You’ve been incredibly dynamic. Was it intentional and how do you do it?

Seth: So, I have a diagnosed medical condition called Career-DHD. I’m just kidding, but I do. I get bored, and it’s actually something that I’m really forward with my managers about. I’ve always been very straight with my managers and the people I work with it, like, 8 to 12 months from now, I will be doing something different. It will be different.

Corey: I wish I’d figured that out earlier on. In my case, the way that I wound up solving for that is I’ve got to come in, I’m going to solve a interesting problem. When I’m done with that, the consulting engagement is over and then I’m going to go away and everyone knows the score going in. Works out way better than, and then I’m going to go cause problems on purpose in other people’s parts of the org because I see problems there. That was where I always went off the rails.

Seth: [laugh]. Yeah, I mean, I don’t take a dissimilar approach. You know, I try to find high-priority, strategic things that also align with my interest. And it’s important to me that there’s things that I can provide and things that I can learn. I never like to be the smartest person in the room because you shouldn’t be in that room anymore; there’s no one for you to learn from. And it’s great to share knowledge, but—

Corey: I’m not convinced I’m the smartest person in the room right now, despite the fact that right now I’m the only person in the room that I’m sitting in.

Seth: I mean, that Minecraft store is pretty intelligent.

Corey: I saw Chihuahua wandering around here, too, a—

Seth: [laugh].

Corey: —minute ago, so there is that.

Seth: But, you know, I think from, like, a career advice standpoint, I tell everyone, you should interview somewhere else at least once a year. You never know what’s out there, and worst-case scenario, you kept your interview skills up to date.

Corey: Keeping those skills in tune is so critically important just because it’s a unique skill set that, for many folks, does not have a whole lot of applicability in their day-to-day job. So, if you suddenly have to find a new job, great, you’re rusty at this, it’s been years, and you’re trying to remember, like, okay, when someone asks you what you’re looking for in your next job, they’re not trying to pick a fight. Don’t respond as if they were. Like, the basic stuff. It’s a skill, like anything else.

Seth: Yeah. And, like, the common questions like, you know, “What do you want to do with your life?” Or like, “What accomplishment are you most proud of?” Like, having those not prepared, but like knowing in general what you want to say from those is very important when you’re thinking about interviewing for other jobs. But even in a big company, like the transfer process is, pretty similar for, like, applying externally to other roles; like sometimes there’s interviews—

Corey: Do they make you code on whiteboards to solve algorithm problems?

Seth: Not me. But—

Corey: Good.

Seth: —in general—

Corey: Google has evolved its interview process since the last time I went through that particular brand of corporate hazing. Good, good, good.

Seth: Yeah. The interview process has definitely been refactored a lot, especially with Covid and remote, but also just trying to be accessible to folks. I know one of the big changes Google has made is we no longer require, like, eight congruent hours of your time. You can split interviews out over multiple days, which has been really accommodating for folks that have, you know, already have a full-time job or have family obligations at home that don’t let them just, like, take eight hours away and devote a hundred percent of their time to interviews. So, I think that is, you know, not a whole lot of positive things that come out of Covid, but the flexibility with, like, interviewing has enabled more people to participate in the interview process that otherwise would not have been able to do so.

Corey: And there’s something to be said, for making this more accessible to folks who come from backgrounds that don’t all look identical. It’s incredibly important.

Seth: Yep.

Corey: One thing that I definitely want to make sure we get to before the end of this is something you’ve been talking about that’s a bit orthogonal, but maybe not entirely so, which is software supply chain security. That has been a common thread of discussion in some circles for a while. What is it, for those who are unfamiliar, like me sometimes, and what does it imply?

Seth: Yeah, so I mean, in the past year—but if you look back, you’ll find more cases of it—. We live in a world where no company—Google, Amazon, the US government—writes every line of code that they run. And even if you do, right, even if you could find a company that doesn’t rely on any external dependencies, what language are they using? Did they write that language? Okay, let’s say hypothetically, you write every single line of code and you wrote your own language, and only your employees contribute to that language.

What operating system are you running on? Because I guarantee you, Linus probably contributed to it, or Gates contributed to it, and they don’t work for you. But let’s say you wrote your own operating system, right—so we’re getting into, like, crazy Google things now, right? Like, only Google would write their own programming language and their own operating system, right? Who manufactured your CPU, right? Like, did you actually—

Corey: There’s always dependencies all the way down. We see this sometimes with companies talk about oh, yeah, we’re going to go to multiple clouds or a different clouds so that we don’t get impacted if there’s another AWS outage in us-east-1. Cool, great. Power to you, but are you sure your payment providers not going to go down? Are they taking a dependency on us-east-1?

Great, let’s say that they’re not. Are you sure that their vendors who are in the critical path are also not taking critical and core dependencies on that? And are you sure that they’re aware of who all of those critical dependencies and those vendors are, and so on and so forth? It is a vast interconnected web. This is a problem. Dependency sprawl is real and I don’t think that there’s a good way to get to the bottom of it, particularly across company boundaries like that.

Seth: Yeah. And this is where if you look at the non-software supply chain, like, if you look at construction, right? If you’re working with a reputable construction agency, they’re actually able to tell you, given a granite countertop or, you know, a quartz countertop, from what beach and what lot on what date the grains of sand in that countertop came from. That is a reality of that industry that is natural. You think about, like, automotive, like, VIN, the Vehicle Identification Numbers, like, they tell you exactly what manufacturer, and then there’s records that show you exactly what human being on the line put that particular part in that machine.

And we don’t have that in software today. Like, we have some, you know, bastardized versions of, like, Software Bills of Material, or SBOM, but the simple fact of the matter is like because software has grown organically and because this wasn’t ingrained in software from the beginning like it was from, you know, traditional manufacturing, you’re going to have an insecure software supply chain for most of my life. Now, what does that actually mean, right—insecure has this negative connotation—it means that you need to make sure that you’re aware of everything that you’re depending on—which is kind of what you were saying is, like, both the technical dependencies and the process or the people dependencies—and you need to have a rigorous process for how you’re going to respond to these incidents. And I think log4j was a really good eye-opening moment for folks when they realized that they didn’t have a way to make a large-scale dependency update across their entire fleet of applications.

Corey: Because who has to do that on a consistent basis? It happens rarely, but when it happens, it’s super important.

Seth: But I do think that more and more, we’re going to see it happened more and more frequently. And ideally, you know, my opinion is that we’re going to get to a point where this is inescapable, but ideally, we get to the point where it’s like, “Oh, okay, this dependency is vulnerable. I have a playbook. I follow the playbook. Everything is patched in 30 minutes or less, and I can move on with my life.” And it’s not a six-week fire drill with people working late and, you know, going super crazy, trying to mitigate these issues.

You know, there’s a lot of work happening in this space. We have, like, SLSA, which is an open standard—SLSA—for how you declare, kind of like, your software bill of materials and things like binary authorization and attestations. There’s, like, Sigstore, there’s Chainguard, there’s some companies evolving in this space. Every time I talk to GitHub, I tell them, I’m like, “Hey, if this VP and that VP, like, talked together and, like, worked on something, you could do something amazing in this space.” But I think it’s going to be quite a while until we get to a point where we can say the software supply chain is secure.

Because like I was saying at the beginning, like, until you manufacture your own CPU, like, you’re dependent on Intel and AMD. And until you write your own programming language, you’re dependent on Ruby, Python, Go, whatever it might be. And until you take no dependencies on some external system—which by the way, might be a bad business decision, like, if someone did the work for you already in an open-source ecosystem, it’s probably a better business decision to evaluate and use that than to build it yourself. Until we have the analysis on that supply chain, and we can in a dashboard, or the click of a button, or the run of a command, very easily see the security status of our supply chain—software supply chain—and determine if a particular vulnerability is or is not relevant, I think we’re still going to be in this firefighting mode for at least another couple of years.

Corey: And I want to say you’re wrong, but I know you’re not. And that’s what, I guess, keeps a lot of us awake at night for unfortunate reasons. Seth, I really want to thank you for taking the time to speak with me. If people want to learn more, where’s the best place to find you?

Seth: I’m on Twitter. You can find me at—

Corey: I’m sorry to hear that. So, am I. It’s the experience.

Seth: Yeah, you can find me at @sethvargo. If you say mean and hateful things to me, I actually exercise this finger, and you can click the block button real fast. But yeah, I mean, my DMs are open. If you have any questions, comments, complaints, concerns, you can throw the complaints away and come to me for everything else.

Corey: Thank you so much for being so generous with your time. I really appreciate it.

Seth: Yeah, thanks for having me. It’s always a pleasure.

Corey: Seth Vargo, engineer at Google. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an angry comment asking how dare I malign the good name of the other cloud provider that isn’t Google that also just so coincidentally happens to employ you.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Brooke

Brooke is the Head of Enablement - AI/ML and Data at Blackbook.ai, an Australian based consulting firm and AWS Partner. Brooke has degrees in Mathematics and Data Engineering and they specialise in developing technically robust solutions that help “non-data people” harness the power of AI for their industry, and communicate this effectively.

Outside of their 'day job', Brooke speaks at Data, AI, Software Engineering, UX and Business conferences and events to Australian and international audiences, and has guest lectured at the University of Queensland Business School and Griffith University. Brooke is proudly a volunteer member of the Queensland National Science Week Committee, and is always on the lookout for new ways to promote STEM pathways to young people, especially young women and members of the LGBTIQA+ community from regional Australia.

Links:

  • Blackbook: https://blackbook.ai/
  • Twitter: https://twitter.com/brooke_jamieson
  • TikTok: https://www.tiktok.com/@brookebytes
  • LinkedIn: https://www.linkedin.com/in/brookejamieson/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: The company 0x4447 builds products to increase standardization and security in AWS organizations. They do this with automated pipelines that use well-structured projects to create secure, easy-to-maintain and fail-tolerant solutions, one of which is their VPN product built on top of the popular OpenVPN project which has no license restrictions; you are only limited by the network card in the instance.

Corey: This episode is sponsored in part by our friends at Sysdig. Sysdig is the solution for securing DevOps. They have a blog post that went up recently about how an insecure AWS Lambda function could be used as a pivot point to get access into your environment. They’ve also gone deep in-depth with a bunch of other approaches to how DevOps and security are inextricably linked. To learn more, visit sysdig.com and tell them I sent you. That’s S-Y-S-D-I-G dot com. My thanks to them for their continued support of this ridiculous nonsense.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. As my 30s draw to a close, I am basically beating myself up emotionally by making myself feel tremendously, tremendously old. And there’s no better way to do that than to go on TikTok where it pops up with, “Hey if you were born before 2004”—and then I just closed the video because it’s ridiculous. It’s more or less of a means of self-flagellation.

But there are good parts to it. One of those good parts is I get to talk to people who I don’t generally encounter in other areas of the giant cloud ecosystem, and my guest today is a shining example of someone who has been very prolific on TikTok but for some reason or other, hadn’t really come across my radar previously. Brooke Jamieson is the Head of Enablement of AI and machine learning at Blackbook. Brooke, thank you for joining me today.

Brooke: Thanks so much for having me. Welcome to 6 a.m. in Brisbane. [laugh].

Corey: It was right before the pandemic that I did my first trip to Australia, discovered that was a real place. Like, “Oh, yeah. You’re going to go to give a talk in Perth. What, are you taking a connection through Narnia?” No, no, it turns out it’s a real place, unlike New Zealand.

Brooke: Oh, yeah. New Zealand’s fake.

Corey: [laugh].

Brooke: I booked a conference in Portugal soon, and it’s going to take me 31 hours to get there from here. So. [laugh].

Corey: I remember the days of international travel. Hopefully for me, they’ll come back again, sooner or later.

Brooke: Fingers crossed.

Corey: What really struck my notice about a lot of your content is the way that you fold multiple things together. First and foremost, you talk an awful lot about machine learning, data engineering, et cetera, and you are the second person that I’ve encountered that really makes me think that there is something to all of this. The first being Emily Freeman, which I’ve discussed on the show previously, and on Twitter, and shouting from the rooftops because she works at AWS and is able to tell the story, which basically, I think makes her a heretic compared to most folks over in that org. But there’s something about making incredibly complex things easily accessible, which is hard enough in its own right, but you also managed to do it basically via short-form video on TikTok. How did you discover all this?

Brooke: Yeah, I have a very strange resume. [laugh]. It is sort of a layered Venn diagram is the way I normally talk about it if I’m doing a conference talk or something. So, I studied pure maths at university the first time, and then I went back and studied data engineering after. But then I also worked in fashion as a model internationally, and then I’ve also worked in things like user experience, doing lots of behavioral science, and everything even design-related around that.

And then I’ve also done lots more work into cloud and AI and everything that happens. So overall, it’s just being about educating people on this. Most of my role now is educating executives and showing them how they were lied to at various conferences so that they can actually make an informed decision. Because if I go to talk to a board, I know when I leave, they’re going to have a conversation about what we talked about without me in the room, and I think executives keep making terrible decisions because they can’t have that conversation as a group. They don’t know what to do when the tour guide isn’t there anymore because they don’t have a shared vocabulary or a framework to talk about what they might like to do, or what they might like to prioritize to do first, things like that.

So, so much of what I do is just really helping people to understand, conceptually from a high level what they’re actually trying to do, so that then they can deliver on that rather than thinking, oh, I just really saw this cool model of a specific AI thing at a conference, and it was a cool animated slide. And I would like to purchase exactly one of those for my company, thank you.

Corey: It’s odd because you don’t have a quote-unquote, “Traditional”—if there is such a thing—DevRel role: You’re not an advocate, you’re not an evangelist. And none of your content and talks that I have seen have been actively selling any product, but they very much been selling ideas and concepts. And it really strikes me that you have threaded the needle beautifully as far as understanding the assignment. You’re trying to cause a shift in the audience, get them to see things in a way that they don’t already without trying to push a particular product or a particular solution. How much of that was happy accident and how much of that was something you set out to do intentionally?

Brooke: First, thanks so much. Second of all, I think this comes from studying maths. So, the number one skill you get from doing a pure maths degree is you have a toolbox with you, and then there’s a number of things in that toolbox. There’s different ways you can solve problems, and usually, there’s a few different ways you can solve a given problem, but you just open up your toolbox that grows over time, and you can see what you can use in there to solve a problem. So, that’s really how I’ve continued to exist, even working in user experience roles as well, just like what elements do we have to even work with here?

And I brought that with me into the cloud as well because I think the really big thing with actually selling tech products is being confident enough to know that there are a number of things you can actually use instead of your product, but if you’re confident enough in the product you have, it will be the obvious solution anyway, so instead, I just get people thinking about what they actually need it for, how they could use it solving a problem and give them ideas on how to apply it. And you would know this: In cloud, there’s always ten million different ways to do something. [laugh]. And it’s just, instead of getting them to think—because then you just get stuck in a thought vortex about, “This one or this one?” Or, “What am I doing,” but instead latch on to an idea of what you’re trying to achieve, and then work out the most optimal way to do that for your underlying infrastructure as well. And even the training of staff that you have, is really important.

Corey: There’s a definite idea around selling—like, I think it’s called ‘solution selling.’ I don’t know; I don’t have a background in this stuff. I’ve basically stumbled into it. But periodically, I’ll have folks come on this show, and I’ll chat with them, “So, what is the outcome we’re looking to have in the audience here?” Because again, telling a story with no real target in mind doesn’t always go super well. And, “Oh, I want people to sign up for my product.” “Okay, how do you envision them doing that?”

And their story is to sit there and pitch the whole time, and it’s, yeah, that’s going to be a really bad show, and I don’t want to put that out. Instead, if you’re active in a particular space, my approach has always been to talk about the painful problem that you solve and allude to what you do and a bit of how you do it. If you make the audience marinate in the painful problem, the folks who are experiencing that are going to sit up and self-select of, “Ooh, that sounds a lot like the problems we have. If they’re talking about this, they might have some ideas and solutions.” It’s a glimpse and a hook into reaching out to find out more.

And to be clear, that’s not the purpose of this show, but if someone wants to pitch a particular product or service, that’s the way to do it because the other stuff just doesn’t work. Giving away free t-shirts, for example, okay, you’ll get a bunch of people clicking links and whatnot, but you’re also effectively talking to people who are super willing to spend time filling out forms and talking to people to get a free t-shirt. I don’t know that for many products, that’s the best way to get qualified leads in.

Brooke: Yeah, it’s tricky. And I think it’s just because everyone’s doing what everyone told them to do. I love reading really terrible sales books. I started when I was younger, just because I could see people trying to use these tactics on me; and I just wanted to know everything there was to know about what it’s like to be a used car salesman in the middle of [laugh] America. And so I’ve read all of these things, and lots of the strategies in them, they only work if you’re in a very specific area that they’re actually working in, and no one’s getting to the problem of how do you actually like to be sold to? How can you improve the experience?

And overall, for consulting, usually, it’s someone—the best end game is someone has seen you around doing other things, and then they come back and they’re like, “I’ve got a really weird problem. I didn’t even know if this is what you can do. Can you help me with this?” And that is—the best client to have, they’re the best—they’re so open to ideas, they trust you because they’ve seen you do good work over time. And you would have seen this so many times, it’s about someone just come to you with a really strange problem, and it may or may not even be what you’ve actually helped them with.

Corey: Help me understand a bit what you do as a Head of Enablement? Because I’ve heard the term a few different ways, always at different companies. As far as day job goes, where do you start? Where do you stop?

Brooke: Yeah, it’s a very fake-sounding job title.

Corey: [unintelligible 00:08:53]—“Oh, what are you?” “Oh, I’m an enabler.” Like, effectively standing behind someone who’s debating relapsing into something, like, “Do it. Do it. Do it.” Now, I don’t imagine that’s what you do. But then again, AI and ML is a weird space. Maybe it is.

Brooke: Just when my friends are online shopping, and they’re not sure if they should buy something. I’m the one messaging them saying, “Yes, get it.” That’s me. So no, what I do is I—there’s really technical people in our teams, we’ve got about 150 consultants across Australia, and then there’s very non-technical business executives who have a problem. And if you don’t have a good conduit between those two groups, the business won’t get what they need, and the technical people won’t have the actual brief they need to solve the problem.

Because so many times people will come to us with what they think is a problem, but it’s actually a symptom, not the root cause, so you just need a really good understanding of overall how businesses work, how business processes work, as well, and then also just really good user experience, information architecture knowledge to go through that. But then all of that would only work if I also had the technical underpinnings so I can then make sure we have everything we need and then communicate that to the development team to make sure that everyone’s getting what they need from it. Lots of places, my job doesn’t exist in a lot of companies, and that’s because they just try to mash [laugh] those two groups together with varying levels of success.

Corey: Or it’s sales enablement of, “Here’s the pitch deck you use. I’m going to build slides all day,” et cetera. “Here’s what the engineers are going to babble about. When they use this phrase, go ahead and repeat this talking point and they’ll shut up and go away,” is often how it manifests. And I don’t get that sense from you at all.

I’m going to call you out slightly on this one. The way you just describe it like, “Well, there are some very technical people, and there are some non-technical people.” And you didn’t actually put yourself into either one of those categories, but let’s call out a bit of background on you. You have a degree in mathematics, but that wasn’t enough, so you decided to go a little more technical than that; you also have a degree in data engineering. If you’re listening to this, please don’t take this the wrong way—

Brooke: Definitely take it the wrong way. [laugh].

Corey: —but you do not present as someone who is first and foremost like, “Code speaks. Code is everything,” the stereotypical technical person who gets lost in their absolute love of the technology to the exclusion of all else. You speak in a way that makes this stuff accessible. Never once in watching any of your content, have I come away feeling dumb as a result, and that’s an incredibly rare thing. But make no mistake, you are profoundly technical on these things.

Brooke: Yeah, making people not feel worried is my number one marketable skill when talking to executives because executives make bad decisions when they don’t know how to have that conversation. But all of that is because they’ve been rising up in their organization for 20 to 30 years, and they didn’t ask questions early on when tech was new, and then it’s gotten to a point where they feel like they can’t ask questions [laugh] anymore because they’re the one in charge, and they’re too nervous to admit they don’t understand something. So, much of what I do that is successful when talking to executives is just really making sure that I’m never out to try and look like the smart one. So, I’m not ever just flexing technical knowledge to make people think that I am the God Almighty of all things tech.

I don’t care about that, so it’s mostly about how can I make people really comfortable with something that they’ve been too scared to ask about probably for quite some time? So that then they can make an informed choice on that front and so they can actually be empowered by that knowledge that they now have. They probably were too scared to ask it the whole time. But it’s just a way of getting through to them. And then you get so much trust from that as well, just because, as well, I’m always very confident to tell people if they’ve been given the wrong information by other parties, I will absolutely tell them immediately, or if they just don’t know how to give success metrics for project, so they end up just forgetting that false negatives or false positives can exist. [laugh].

So, educating them, even on accuracy and recall measures and things like that, as well and doing it in a way where they don’t ever feel threatened is the number one key to success that no one ever tells you about as a thing because no one wants to even admit that people could possibly be threatened by this.

Corey: A lot of the content that you wind up building is aimed around career advice, particularly for folks early on in their careers. And the reason I bring that up is that you are alluding to something that I see when I interview folks all the time—I went through it myself—where there was a time you’re going through a technical interview, and you get the flop sweat where I don’t know the answer to this question. And there are a few things you can do: You can give up and shut down, which okay, that is in many cases are natural inclination, but not particularly helpful in those environments; you can bluff your way through the answer, which I generally don’t advise because when an interviewer is asking you a technical question, it’s a reasonable guess that they know the right answer; but the mark of seniority that it took me a distressingly long time to learn this is I just sometimes laugh, I say, “I have absolutely no idea, but if I had to guess…” and then I’ll speculate wildly. And that, in my experience, is the mark of the kind of person you generally want to have on your team. And there are elements of what you just said, threaded throughout that entire approach of not making people feel less than.

On the other side of that interview table, when I’m sitting there as a candidate. I hated those interviews where someone sits there and tries to prove they’re the smartest person in the room. Yeah, I too, am the smartest person in the room when I wrote the interview questions. But for me, it’s a given Tuesday; for the person I’m interviewing, it’s determining the next stage of their career. There’s a power imbalance there.

Brooke: Yeah. And this has always happened to me in job interviews as well. I have a very polarizing resume just because it’s not traditional. I didn’t do software engineering at university and then work as a software engineer. I just haven’t gone through that linear pathway, so there’s lots of people just trying to either figure me out or get to a gotcha moment where they can really just work out what’s actually happening.

I have no interest in it. I’m happy to say that I don’t know something. And being able to openly say that you don’t know something is so helpful, as you’re saying. It’s the number one skill I wish people would learn because it’s fine to not [laugh] understand things.

Corey: A couple of times, I was the first DevOps hire in startup that was basically being interviewed by a bunch of engineers. And I went through a lot of those interviews and took the job only a couple of times, and one of the key differentiators for me was when they sat down and looked sort of sheepish and asked me a question of the form, “Look, I know how to interview a software engineer, but I sort of get the sense that you’re not going to do that super well.” Yeah, surprise; I’m not a software engineer. “What is the best way to interview you to really expose where you start and where you stop?” Which I think is such a great question, if you don’t know.

Now, in my world, the way that—now that I’m on the other side of the table, I bring in experts to help me evaluate people, otherwise I run the very real risk of hiring the person that sounds the most confident, and that doesn’t generally end well in technical spaces. And really figure out what it is that makes people shine. Everyone talks about how to pass the technical interview, but there’s very little discussion on the other side of it, which is what kind of training do most of us—are most of us given to effectively conduct a technical interview?

Brooke: Yeah, and not even just interview skills, but leadership skills. Number one thing I always talk to you when I’m talking to university students is I let them know that probably the manager they will end up having, if they work in tech, probably has no leadership training or management training of any sort. So, [laugh] if you are just assuming that they will be, like, really just a straight, always making the best management decision or always doing something the most perfect way, they probably have no idea what they’re doing as well. And that’s really important for people to go into jobs and interviews knowing, is that it’s fine if the other person doesn’t know as well. Do they want you to win? I think that’s the number one thing I always am left with after interviews.

And even when I was interviewing—I worked as a marketing manager for a while, so even when I was interviewing for marketing jobs, you could tell whether the person on the other side of the table wanted you to do well or not. And especially when you’re looking for an early career job, regardless of any other factor, if someone wants you to do well, that is a good job for you to have. It will just mean so much more to your momentum throughout your career.

Corey: I’m a big believer in even if you decide not to continue with a particular candidate, my objective has always been that I want them to think well of the company, I want them to consider reapplying down the road when their skill set changes or what we’re looking for changes. And I want them to walk away from the experience with a, “That was a very fair and honest experience. I might recommend applying there to other people I know.” And we’ve had some people come through that way, so we’re definitely succeeding. Whereas I went through the Google SRE interview twice—the second time, I think, was in 2015—and I swore midway through the process that even if they offered me the job, I wouldn’t accept it because they didn’t want to work with a place where they were going to treat people like that.

Full disclosure, I did not get the offer because I’m bad at solving, you know, coding challenges in a Google Doc. Who knew. But it was one of those, I will not put myself through that again. So yeah, it turns out now I’ve made myself completely unemployable by anyone, so problem solved. “Oh, yeah, I’m never going to put myself through one of those job interviews,” says man who made himself completely un-interviewable, any job ever, ever again. I’ll have to change my name and enter witness protection if I want to [laugh] enter the industry after the nonsense I’m pulling.

Brooke: Just wear a mustache. They’ll never know.

Corey: Oh, yeah, you joke but I—some of me wonders on that one. So, I am curious as to your adventures with TikTok. I know I started the show talking about that, but it’s still a weird format for me. I thought I got weird comments on Twitter. Oh, no, no, no, not compared to some of the people responding to things on the TikToks.

And it’s a different format, it’s a different audience, it feels like, but there’s still a strong appetite for career discussions and for technical discussions as well. How did you stumble on the platform? And how did you figure out what you would be talking about there?

Brooke: Yeah, I put off making a TikTok for so long. So, I worked as a model internationally before my current job—I did it during and after uni—and so I have a very fashion Instagram that’s very polished… like, I put thought into the outfits that I’m wearing in photos, which means I just haven’t posted a lot lately because I just don’t have the energy. So, the idea of going on TikTok to do something that is very quick is horrible to me [laugh] as a thought. Also, I hid my fashion past, I was closeted for a long time in tech, just because it was actively negatively impacting career prospects. But one of the best gifts about moving to a leadership team in a management space is that people don’t care about that as much anymore, which is really good.

So, just it was a big move for me in terms of bringing my closeted to past back into what I’m actually doing in tech, just to get more people aware of the opportunities that are out there. Because there’s so many people during Covid that wanted to work from home and they wanted to transition to a job that would allow them to work remotely with benefits and security and everything that goes along with that, and tech is a really good industry to get that in.

Corey: Oh, there are millions of jobs now that didn’t exist two years ago that empower full remote, either within a given country or globally. Just, do you have an internet connection wherever you happen to be? I mean, we have people here who are excited to go and do all kinds of traveling, and we have people who have—this has been challenging for them—but, on the paperwork on our side, just fill up the forms, but we’ve had to effectively open tax accounts with different states as they relocate during the course of the pandemic. And power to them; that’s what administrative teams are for. But it’s really nice to be able to empower stuff like that because for the longest time—I live in San Francisco, and it felt like the narrative was, “We are a disruptive industry that is changing the face of the world. And we are applying that disruption by taking a job that can be done from literally anywhere and creating a land crunch in eight square miles in an earthquake zone.”

It really didn’t seem like it was the most forward-thinking type of event. And I’m hoping—in fact, we’re seeing evidence of it—that this is going to be one of the lasting changes of the pandemic. People don’t want to go back into the crappy offices.

Brooke: Yeah. And especially in Australia, as well. So, when you’re talking about San Francisco being crunched into eight miles, Australia is like that, but the whole country. So, it’s a very large space, but there’s only a few capital cities dotted around that I think more than 90-something percent of people live in those big centers.

Corey: Yeah. Yeah, I made that mistake by taking my week down there and visit—and giving talks in Sydney, Melbourne, and Perth. And it
was, well, why not? Like, I’m flying all the way over there. How far apart could it be?

It’s, “What do you mean, there’s that many time zones? And the flight is how many hours to go from one side to the other?” Yeah, on
professional advice for people who are considering doing that: Don’t.

Brooke: Yeah, it’s not for the faint-hearted. But it also means that there’s so many people that don’t live in capital cities that could now work remotely for tech companies. Or I know friends that they originally lived in a capital city, and then have gone to move into regional centers. And as someone that grew up in a regional center, that’s so important to be able to spread the tech ecosystem out further. It’s 1000 kilometers from where I live now, which I don’t know what that is in miles in freedom units, but probably it’s about a ten-hour drive if you’re driving there; if you’re driving without stopping.

So, it’s a really long way away. And that’s just—it’s not, like, something you can just drive a few times over for a meetup. There’s just nothing around for quite a long time. So, being able to disperse technical knowledge throughout the country is something that’s really important to me, especially just because it’s opening up futures for more diverse groups, even the people that are using tech in the vast majority of geographically distributed Australia are completely ignored from making that tech. And that’s something that’s really a growing issue that is getting fixed as there are more opportunities to move remote jobs there. But people don’t even know that these jobs exist, so it’s just about getting out into the regions to show people that it’s possible.

Corey: This episode is sponsored in part by our friends at Vultr. Spelled V-U-L-T-R because they’re all about helping save money, including on things like, you know, vowels. So, what they do is they are a cloud provider that provides surprisingly high performance cloud compute at a price that—while sure they claim its better than AWS pricing—and when they say that they mean it is less money. Sure, I don’t dispute that but what I find interesting is that it’s predictable. They tell you in advance on a monthly basis what it’s going to going to cost. They have a bunch of advanced networking features. They have nineteen global locations and scale things elastically. Not to be confused with openly, because apparently elastic and open can mean the same thing sometimes. They have had over a million users. Deployments take less that sixty seconds across twelve pre-selected operating systems. Or, if you’re one of those nutters like me, you can bring your own ISO and install basically any operating system you want. Starting with pricing as low as $2.50 a month for Vultr cloud compute they have plans for developers and businesses of all sizes, except maybe Amazon, who stubbornly insists on having something to scale all on their own. Try Vultr today for free by visiting: vultr.com/screaming, and you’ll receive a $100 in credit. Thats V-U-L-T-R.com slash screaming.

Corey: You were mentioning that you are about to embark on a 31-hour travel nightmare nonsense thing to go to Portugal to give a talk. What is the talk you’re giving, and what’s the venue?

Brooke: It’s for NDC Porto. And so, I’ve done NDC in Sydney, virtually, twice—

Corey: NDC is… I’m sorry?

Brooke: It’s a really big conference. I think it started as… Norwegian something? Norwegian Developers Conference, someone will roast
me in the comments of this.

Corey: Well, not on this show. Generally, we get a pretty awesome audience base compared to, you know, the TikTok. So, I’m sure we’ll excerpt parts of this for the TikToks, and then oh, then all hell is going to break loose.

Brooke: [laugh]. That we’ll find them. Yeah, it’s a big conference series. So, they have them in Oslo, Copenhagen, Sydney, London, Melbourne, and Porto as well. So, it’s quite a big—I think it’s very eurocentric. I don’t know if it would have be in any of the US audiences yet. But it’s a really wholesome group.

Last time I did the conference in Sydney, the segment before me was someone showing their pet llamas on camera. So… love that. [laugh]. But my talk is just about enterprise applications of AI and machine learning. So, it’s mostly the same sessions that I give to executives; I just give it to software engineers, and then tell them about how I talk to executives while I’m doing it to show why it works.

Corey: You also give periodic talks at universities as well. You have been very prolific on the speaking circuit. What’s the common thread that winds up tying all of these disparate audiences together?

Brooke: People ask me and I say yes. Um—[laugh].

Corey: Hey, there we go.

Brooke: Yeah. No, it’s mostly about I just want to make this an easier pathway for other people. If you can see this art here—it’s backwards, probably, but it says, “Be who you needed when you were younger.” I made this, and it’s just how I go through my tech life. When I talk to high school students and university students, no one’s ever honest with them about what it’s actually like to have a job because everyone is just telling them how fantastic it is and how everyone will think it’s so fantastic that they’re a graduate of that institution, and they will get their dream job, and they’ll ride home on a unicorn and everything will be perfect.

And no adults are ever honest to people because everyone wants something from them. So, it is an absolute immense position of privilege to be able to go in and say, “Here is unfortunate realities of what you’re about to step into.” Because my parents, neither of them worked in office, my mom teaches children with disabilities—so she’s retired now—and my dad is a telephone technician, so like, I didn’t know anyone working in office growing up, it wasn’t part of what I did. So, I can’t tell people that this is what networking actually is. That will be someone who you are in the room with right now who is extremely wealthy, and their parents own something, and they will sail through life. You need to work much harder than them. [laugh].

And just being able to have these actual conversations with students—because it’s so valuable—and it’s guidance that I wish I had earlier on and that you can actually, if you’re aware of what’s happening in these systems, you can hedge against it. But it’s just, I think it’s doing students a disservice to not be honest to them, so I take a lot of pride in doing that. [laugh].

Corey: I like doing that, but I’m also worried that I am going to send the wrong message if I do. Because let’s be honest, I can get away with an awful lot of stuff based upon my perceived position in the industry, the fact that I am clearly self-employed—when you own the company, it turns out you can get away with a lot—and also I’m 15 to 20 years into my career, whereas if I pulled a lot of this nonsense fresh out of school in my first job, I would have been fired. No ‘would have’ about it; I was fired and didn’t even pull half of the jokes that I pull now. So, when I give interview advice on TikTok, like here’s how you pick a fight with the interviewer. Yeah, if someone actually does that, it’s not going to go well, so I live in fear of effectively giving the kind of advice that is actively harmful. If I’m going to do that, I at least try to put a disclaimer into it. But we’ll see.

Brooke: Yeah. And even just showing people that it is possible for an interviewer to not want you to do well. So, many people are not aware of that because the only idea they have of interviews is that… been something they’ve been told at whatever school or boot camp they went through that someone really wants them to succeed and will help them to develop their journey. That’s not… it’s not normal, so being able to actually decipher what is and isn’t happening there is a really good skill. And people just aren’t aware that things like that are possible, or even I don’t know, as a… [unintelligible 00:27:33] person in STEM, I have a lot of sage advice to give to people about what it is and isn’t like in reality.

And that’s where my history of very strange jobs comes into play as well. So, I worked at a car parts store for four years growing up,
selling people different types of filters and fuel filters and sound systems for their car.

Corey: And blinker fluid, depending on how sketchy the numbers look that month. Of course, of course.

Brooke: Yeah. But people would come into the store and just ask for a man straightaway, or call up, and then I was the only one that had any physics education in the store, so some days if [my brother 00:28:08] wasn’t working, so they would call up and ask for what type of resistance they needed for their car stereo, and I would tell them, but they would [unintelligible 00:28:17] put a man on. And then eventually they would be like, “I don’t know. I have to ask Brooke.” And put me back on the phone, and I would just pretend to have never heard the start of it.

But it’s just, if you don’t have a diverse background of jobs you’ve had, or different service jobs, it gives you more structure about how to actually talk about what it is like to work in tech because some bits are much worse than they appear, and some things are actually a lot better than they appear as well. It’s just depending on who’s talking about it in the media at a given day.

Corey: And I’ve said it before—it’s always worth repeating—this is what privilege looks like because it’s easy for me to sit here and say, “Look, the stuff that I built, the company I’ve put together, the reputation for myself that I wound up establishing, well, I had to do it all myself. None of it was handed to me.” And that is true. However, I didn’t have to fight against bullshit like that. I didn’t have a headwind of people telling me that I was somehow unqualified or didn’t belong in the place that I was in.

When I made a pronouncement, even when it was wrong, it was presumed accurate until proven otherwise. So, there’s a lot of stuff around this that just contributes to a terrible toxic environment. That is what privilege is, and you can’t set that aside, you can’t turn that away. And we all have privilege in different ways, but it’s often considered to be controversial. I don’t see it that way at all. It’s one of those, “You were born on third base; you didn’t hit a triple.”

Brooke: And it’s just about what you do with it, as well. There are some people who are immensely privileged and then they just do nothing to help anyone else. They don’t let the ladder down for anyone else after them, so—

Corey: “Send the elevator back down,” is what Stephen O’Grady over at RedMonk said, and is a phrase that’s stuck with me for years now. And it’s the perfect expression of it. It’s as opposed to folks who wind up pulling up the rope behind them, “Well, screw you. I got mine.” No thanks.

Brooke: Yeah.

Corey: That’s not how I want to be remembered.

Brooke: And that’s why it’s so important to talk to people about this early in their career as well because these people will become a manager probably. So, then say, “Hey, when you are inevitably a manager, [laugh] you are in a position of power now. Here are things you can actively do that will be who you needed when you were younger.” That’s what will actually help people, too. So, being able to really specifically say, once you are in an organization, you have the opportunity to make change, especially in graduate roles in organizations, I noticed they get so much bandwidth to actively make decisions because higher-ups are just so excited that there’s someone young working there.

So, being able to go through and look at what they are actually doing. And people trust you, then they trust your opinion, and they’ll trust your opinion. Especially on issues like sustainability, everyone’s just, “Oh, who’s a child we can ask?” So, being able to then give them an
answer that’s helpful. Or say even, “I didn’t know about this. Maybe you should ask someone that this affects.”

Being able to then hand the microphone to someone else is a skill that is never actively taught to anyone. So, I think that’s what’s really—it’s a slow part of diversity and inclusion changing over time, but it’s a really important part of actively modeling that behavior of what it looks like to do a decent job.

Corey: Brooke, I really want to thank you for taking the time to speak with me today. If people want to learn more, where’s the best place to find you?

Brooke: I’m on Twitter as @brooke_jamieson; I’m on TikTok as BrookeBytes, and I’m on LinkedIn is probably the best place to—I check
that inbox the most. And my name is just Brooke Jamieson, which will be in the show notes.

Corey: And we will, of course, put links to all of that in the [show notes 00:31:38]. Thanks again for your time. I really appreciate it.

Brooke: Thanks so much for having me.

Corey: Brooke Jamieson, Head of Enablement for AI, ML, and data at Blackbook. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with a long, rambling, angry comment that says this is not the content that you expected, you were not happy with it at all, and if I really wanted to have these conversations, I should have instead first demonstrated both of our technical suitability by solving algorithm problems on a whiteboard.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Liz

Liz Rice is Chief Open Source Officer with cloud native networking and security specialists Isovalent, creators of the Cilium eBPF-based networking project. She is chair of the CNCF's Technical Oversight Committee, and was Co-Chair of KubeCon + CloudNativeCon in 2018. She is also the author of Container Security, published by O'Reilly.

She has a wealth of software development, team, and product management experience from working on network protocols and distributed systems, and in digital technology sectors such as VOD, music, and VoIP. When not writing code, or talking about it, Liz loves riding bikes in places with better weather than her native London, and competing in virtual races on Zwift.

Links:

  • Isovalent: https://isovalent.com/
  • Container Security: https://www.amazon.com/Container-Security-Fundamental-Containerized-Applications/dp/1492056707/
  • Twitter: https://twitter.com/lizrice
  • GitHub: https://github.com/lizrice
  • Cilium and eBPF Slack: http://slack.cilium.io/
  • CNCF Slack: https://cloud-native.slack.com/join/shared_invite/zt-11yzivnzq-hs12vUAYFZmnqE3r7ILz9A

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Today’s episode is brought to you in part by our friends at MinIO the high-performance Kubernetes native object store that’s built for the multi-cloud, creating a consistent data storage layer for your public cloud instances, your private cloud instances, and even your edge instances, depending upon what the heck you’re defining those as, which depends probably on where you work. It’s getting that unified is one of the greatest challenges facing developers and architects today. It requires S3 compatibility, enterprise-grade security and resiliency, the speed to run any workload, and the footprint to run anywhere, and that’s exactly what MinIO offers. With superb read speeds in excess of 360 gigs and 100 megabyte binary that doesn’t eat all the data you’ve gotten on the system, it’s exactly what you’ve been looking for. Check it out today at min.io/download, and see for yourself. That’s min.io/download, and be sure to tell them that I sent you.

Corey: This episode is sponsored in part by our friends at Sysdig. Sysdig is the solution for securing DevOps. They have a blog post that went up recently about how an insecure AWS Lambda function could be used as a pivot point to get access into your environment. They’ve also gone deep in-depth with a bunch of other approaches to how DevOps and security are inextricably linked. To learn more, visit sysdig.com and tell them I sent you. That’s S-Y-S-D-I-G dot com. My thanks to them for their continued support of this ridiculous nonsense.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. One of the interesting things about hanging out in the cloud ecosystem as long as I have and as, I guess, closely tied to Amazon as I have been, is that you learned that you never quite are able to pronounce things the way that people pronounce them internally. In-house pronunciations are always a thing. My guest today is Liz Rice, the Chief Open Source Officer at Isovalent, and they’re responsible for, among other things, the Cilium open-source project, which is around eBPF, which I can only assume is internally pronounced as ‘Ehbehpf’. Liz, thank you for joining me today and suffering my pronunciation slings and arrows.

Liz: I have never heard ‘Ehbehpf’ before, but I may have to adopt it. That’s great.

Corey: You also are currently—in a term that is winding down if I’m not misunderstanding—you were the co-chair of KubeCon and CloudNativeCon at the CNCF, and you are also currently on the technical oversight committee for the foundation.

Liz: Yeah, yeah. I’m currently the chair, in fact, of the technical oversight committee.

Corey: And now that Amazon has joined, I assumed that they had taken their horrible pronunciation habits, like calling AMIs ‘Ah-mies’ and whatnot, and started spreading them throughout the ecosystem with wild abandon.

Liz: Are we going to have to start calling CNCF ‘Ka’Nff’ or something?

Corey: Exactly. They’re very frugal, by which I mean they never buy a vowel. So yeah, it tends to be an ongoing challenge. Joking and all the rest aside, let’s start, I guess, at the macro view. The CNCF does an awful lot of stuff, where if you look at the CNCF landscape, for example, like, I think some of my jokes on the internet go a bit too far, but you look at this thing and last time I checked, there were something like four or 500 different players in various spaces.

And it’s a very useful diagram, don’t get me wrong by any stretch of the imagination, but it also is one of those things that is so staggeringly vast that I’ve got a level with you on this one, given my old, ancient sysadmin roots, “The hell with it. I’m going to run some VMs in a three-tiered architecture just like grandma and grandpa used to do,” and call it good. Not really how the industry is evolved, but it’s overwhelming.

Liz: But that might be the right solution for your use case so, you know, don’t knock it if it works.

Corey: Oh, yeah. If it’s a terrible architecture and it works, is it really that terrible of an architecture? One wonders.

Liz: Yeah, yeah. I mean, I’m definitely not one of those people who thinks, you know, every solution has the same—you know, is solved by the same hammer, you know, all problems are not the same nail. So, I am a big fan of a lot of the CNCF projects, but that doesn’t mean to say I think those are the only ways to deploy software. You know, there are plenty of things like Lambda are a really great example of something that is super useful and very applicable for lots of applications and for lots of development teams. Not necessarily the right solution for everything. And for other people, they need all the bells and whistles that something like Kubernetes gives them. You know, horses for courses.

Corey: It’s very easy for me to make fun of just about any company or service or product, but the thing that always makes me set that aside and get down to brass tacks has been, “Okay, great. You can build whatever you want. You can tell whatever glorious marketing narrative you wish to craft, but let’s talk to a real customer because once we do that, then if you’re solving a problem that someone is having in the wild, okay, now it’s no longer just this theoretical exercise and PowerPoint. Now, let’s actually figure out how things work when the rubber meets the road.”

So, let’s start, I guess, with… I’ll leave it to you. Isovalent are the creators of the Cilium eBPF-based networking project.

Liz: Yeah.

Corey: And eBPF is the part of that I think I’m the most familiar with having heard the term. Would you rather start on the company side or on the eBPF side?

Liz: Oh, I don’t mind. Let’s—why don’t we start with eBPF? Yeah.

Corey: Cool. So easy, ridiculous question. I know that it’s extremely important because Brendan Gregg periodically gets on stage and tells amazing stories about this; the last time he did stuff like that, I went stumbling down into the rabbit hole of DTrace, and I have never fully regretted doing that, nor completely forgiven him. What is eBPF?

Liz: So, it stands for extended Berkeley Packet Filter, and we can pretty much just throw away those words because it’s not terribly
helpful. What eBPF allows you to do is to run custom programs inside the kernel. So, we can trigger these programs to run, maybe because a network packet arrived, or because a particular function within the kernel has been called, or a tracepoint has been hit. There are tons of places you can attach these programs to, or events you can attach programs to.

And when that event happens, you can run your custom code. And that can change the behavior of the kernel, which is, you know, great power and great responsibility, but incredibly powerful. So Brendan, for example, has done a ton of really great pioneering work showing how you can attach these eBPF programs to events, use that to collect metrics, and lo and behold, you have amazing visibility into what’s happening in your system. And he’s built tons of different tools for observing everything from, I don’t know, memory use to file opens to—there’s just endless, dozens and dozens of tools that Brendan, I think, was probably the first to build. And now this sort of new generations of eBPF-based tooling that are kind of taking that legacy, turning them into maybe more, going to say user-friendly interfaces, you know, with GUIs, and hooking them up to metrics platforms, and in the case of Cilium, using it for networking and hooking
it into Kubernetes identities, and making the information about network flows meaningful in the context of Kubernetes, where things like IP addresses are ephemeral and not very useful for very long; I mean, they just change at any moment.

Corey: I guess I’m trying to figure out what part of the stack this winds up applying to because you talk about, at least to my mind, it sounds like a few different levels all at once: You talk about running code inside of the kernel, which is really close to the hardware—it’s oh, great. It’s adventures in assembly is almost what I’m hearing here—but then you also talk about using this with GUIs, for example, and operating on individual packets to run custom programs. When you talk about running custom programs, are we talking things that are a bit closer to, “Oh, modify this one field of that packet and then call it good,” or are you talking, “Now, we launch Microsoft Word.”

Liz: Much more the former category. So yeah, let’s inspect this packet and maybe change it a bit, or send it to a different—you know, maybe it was going to go to one interface, but we’re going to send it to a different interface; maybe we’re going to modify that packet; maybe we’re going to throw the packet on the floor because we don’t—there’s really great security use cases for inspecting packets and saying, “This is a bad packet, I do not want to see this packet, I’m just going to discard it.” And there’s some, what they call ‘Packet of Death’ vulnerabilities that have been mitigated in that way. And the real beauty of it is you just load these programs dynamically. So, you can change the kernel or on the fly and affect that behavior, just immediately have an effect.

If there are processes already running, they get instrumented immediately. So, maybe you run a BPF program to spot when a file is opened. New processes, existing processes, containerized processes, it doesn’t matter; they’ll all be detected by your program if it’s observing file open events.

Corey: Is this primarily used from a security perspective? Is it used for—what are the common use cases for something like this?

Liz: There’s three main buckets, I would say: Networking, observability, and security. And in Cilium, we’re kind of involved in some aspects of all those three things, and there are plenty of other projects that are also focusing on one or other of those aspects.

Corey: This is where when, I guess, the challenge I run into the whole CNCF landscape is, it’s like, I think the danger is when I started down this path that I’m on now, I realized that, “Oh, I have to learn what all the different AWS services do.” This was widely regarded as a mistake. They are not Pokémon; I do not need to catch them all. The CNCF landscape applies very similarly in that respect. What is the real-world problem space for which eBPF and/or things like Cilium that leverage eBPF—because eBPF does sound fairly low-level—that turn this into something that solves a problem people have? In other words, what is the problem that Cilium should be the go-to answer for when someone says, “I have this thing that hurts.”

Liz: So, at one level, Cilium is a networking solution. So, it’s Kubernetes CNI. You plug it in to provide connectivity between your applications that are running in pods. Those pods have to talk to each other somehow and Cilium will connect those pods together for you in a very efficient way. One of the really interesting things about eBPF and networking is we can bypass some of the networking stack.

So, if we are running in containers, we’re running our applications in containers in pods, and those pods usually will have their own networking namespace. And that means they’ve got their own networking stack. So, a packet that arrives on your machine has to go through the networking stack on that host machine, go across a virtual interface into your pod, and then go through the networking stack in that pod. And that’s kind of inefficient. But with eBPF, we can look at the packet the moment it’s come into the kernel—in fact in some cases, if you have the right networking interfaces, you can do it while it’s still on the network interface card—so you look at that packet and say, “Well, I know what pod that’s destined for, I can just send it straight there.” I don’t have to go through the whole networking stack in the kernel because I already know exactly where it’s going. And that has some real performance improvements.

Corey: That makes sense. In my explorations—we’ll call it—with Kubernetes, it feels like the universe—at least at the time I went looking into it—was, “Step One, here’s how to wind up launching Kubernetes to run a blog.” Which is a bit like using a chainsaw to wind up cutting a sandwich. Okay, massively overpowered but I get the basic idea, like, “Okay, what’s project Step Two?” It’s like, “Oh, great. Go build Google.”

Liz: [laugh].

Corey: Okay, great. It feels like there’s some intermediary steps that have been sort of glossed over here. And at the small-scale that I kicked the tires on, things like networking performance never even entered the equation; it was more about get the thing up and running. But yeah, at scale, when you start seeing huge numbers of containers being orchestrated across a wide variety of hosts that has serious repercussions and explains an awful lot. Is this the sort of thing that gets leveraged by cloud providers themselves, is it something that gets built in mostly on-prem environments, or is it something that rides in, almost, user-land for most of these use cases that customers coming to bringing to those environments? I’m sorry, users, not customers. I’m too used to the Amazonian phrasing of everyone as a customer. No, no, they are users in an open-source project.

Liz: [laugh]. Yeah, so if you’re using GKE, the GKE Dataplane V2 is using Cilium. Alibaba Cloud uses Cilium. AWS is using Cilium for EKS Anywhere. So, these are really, I think, great signals that it’s super scalable.

And it’s also not just about the connectivity, but also about being able to see your network flows and debug them. Because, like you say, that day one, your blog is up and running, and day two, you’ve got some DNS issue that you need to debug, and how are you going to do that? And because Cilium is working with Kubernetes, so it knows about the individual pods, and it’s aware of the IP addresses for those pods, and it can map those to, you know, what’s the pod, what service is that pod involved with. And we have a component of Cilium called Hubble that gives you the flows, the network flows, between services. So, you know, we’ve probably all seen diagrams showing Service A talking to Service B, Service C, some external connectivity, and Hubble can show you those flows between services and the outside world, regardless of how the IP addresses may be changing underneath you, and aggregating network flows into those services that make sense to a human who’s looking at a Kubernetes deployment.

Corey: A running gag that I’ve had is that one of the drawbacks and appeals of Kubernetes, all at once, is that it lets you cosplay as a cloud provider, even if you don’t happen to work for one of them. And there’s a bit of truth to it, but let’s be serious here, despite what a lot of the cloud providers would wish us to believe via a bunch of marketing, there’s a tremendous number of data center environments out there, hybrid environments, and companies that are in those environments are not somehow laggards, or left behind technologically, or struggling to digitally transform. Believe it or not—I know it’s not a common narrative—but large companies generally don’t employ people who lack critical thinking skills and strategic insight. There’s usually a reason that things are the way that they are and when you don’t understand that my default approach is that, oh context that gets missing, so I want to preface this with the idea there is nothing wrong in those environments. But in a purely cloud-native environment—which means that I’m very proud about having no single points of failure as I have everything routing to a single credit card that pays the cloud providers—great. What is the story for Cilium if I’m using, effectively, the managed Kubernetes options that Name Any Cloud Provider will provide for me these days? Is it at that point no longer for me or is it something that instead expresses itself in ways I’m not seeing, yet?

Liz: Yeah, so I think, as an open-source project—and it is the only CNI that’s at incubation level or beyond, so you know, it’s CNCF-supported networking solution; you can use it out of the box, you can use it for your tiny blog application if you’ve decided to run that on Kubernetes, you can do so—things start to get much more interesting at scale. I mean, that… continuum between you know, there are people purely on managed services, there are people who are purely in the cloud, hybrid cloud is a real thing, and there are plenty of businesses who have good reasons to have some things in their own data centers, something’s in the public cloud, things distributed around the world, so they need connectivity between those. And Cilium will solve a lot of those problems for you in the open-source, but also, if you’re telco scale and you have things like BGP networks between your data centers, then that’s where the paid versions of Cilium, the enterprise versions of Cilium, can help you out. And, as Isovalent, that’s our business model to have, like—we fully support or we contribute a lot of resources into the open-source Cilium, and we want that to be the best networking solution for anybody, but if you are an enterprise who wants those extra bells and whistles, and the kind of scale that, you know, a telco, or a massive retailer, or a large media organization, or name your vertical, then we have solutions for that as well. And I think it was one of the really interesting things about the eBPF side of it is that, you know, we’re not bound to just Kubernetes, you know? We run in the kernel, and it just so happens that we have that Kubernetes interface for allocating IP addresses to endpoints that happened to be pods. But—

Corey: So, back to my crappy pile of VMs—because the hell with all this newfangled container nonsense—I can still benefit from something like Cilium?

Liz: Exactly, yeah. And there’s plenty of people using it for just load-balancing, which, why not have an eBPF-based high-performance load balancer?

Corey: Hang on, that’s taking me a second to work my way through. What is the programming language for eBPF? It is something custom?

Liz: Right. So, when you load your BPF program into the kernel, it’s in the form of eBPF bytecode. There are people who write an eBPF bytecode by hand; I am not one of those people.

Corey: There are people who used to be able to write Sendmail configs without running through the M four preprocessor, and I don’t understand those people either.

Liz: [laugh]. So, our choices are—well, it has to be a language that can be compiled into that bytecode, and at the moment, there are two options: C, and more recently, Rust. So, the C code, I’m much more familiar with writing BPF code in C, it’s slightly limited. So, because these BPF programs have to be safe to run, they go through a verification process which checks that you’re not going to crash the kernel, that you’re not going to end up in some hardware loop, and basically make your machine completely unresponsive, we also have to know that BPF programs, you know, they’ll only access memory that they’re supposed to and that they can’t mess up other processes. So, there’s this BPF verification step that checks for example that you always check that a pointer isn’t nil before you dereference it.

And if you try and use a pointer in your C code, it might compile perfectly, but when you come to load it into the kernel, it gets rejected because you forgot to check that it was non-null before.

Corey: You try and run it, the whole thing segfaults, you see the word ‘fault’ there and well, I guess blameless just went out the window there.

Liz: [laugh]. Well, this is the thing: You cannot segfault in the kernel, you know, or at least that’s a bad [day 00:19:11]. [laugh].

Corey: You say that, but I’m very bad with computers, let’s be clear here. There’s always a way to misuse things horribly enough.

Liz: It’s a challenge. It’s pretty easy to segfault if you’re writing a kernel module. But maybe we should put that out as a challenge for the listener, to try to write something that crashes the kernel from within an eBPF because there’s a lot of very smart people.

Corey: Right now the blood just drained from anyone who’s listening, in the kernel space or the InfoSec space, I imagine.

Liz: Exactly. Some of my colleagues at Isovalent are thinking, “Oh, no. What’s she brought on here?” [laugh].

Corey: What have you done? Please correct me if I’m misunderstanding this. So, eBPF is a very low-level tool that requires certain amounts of braining in order [laugh] to use appropriately. That can be a heavy lift for a lot of us who don’t live in those spaces. Cilium distills this down into something that is all a lot more usable and understandable for folks, and then beyond that, you wind up with Isovalent, that winds up effectively productizing and packaging this into something that becomes a lot more closer to turnkey. Is that directionally accurate?

Liz: Yes, I would say that’s true. And there are also some other intermediate steps, like the CLI tools that Brendan Gregg did, where you can—I mean, a CLI is still fairly low-level, but it’s not as low-level as writing the eBPF code yourself. And you can be quite in-dep—you know, if you know what things you want to observe in the kernel, you don’t necessarily have to know how to write the eBPF code to do it, but if you’ve got these fairly low-level tools to do it. You’re absolutely right that very few people will need to write their own… BPF code to run in the kernel.

Corey: Let’s move below the surface level of awareness; the same way that most of us don’t need to know how to compile our own kernel in this day and age.

Liz: Exactly.

Corey: A few people very much do, but because of their hard work, the rest of us do not.

Liz: Exactly. And for most of us, we just take the kernel for granted. You know, most people writing applications, it doesn’t really matter if—they’re just using abstractions that do things like open files for them, or create network connections, or write messages to the screen, you don’t need to know exactly how that’s accomplished through the kernel. Unless you want to get into the details of how to observe it with eBPF or something like that.

Corey: I’m much happier not knowing some of the details. I did a deep dive once into Linux system kernel internals, based on an incredibly well-written but also obnoxiously slash suspiciously thick O’Reilly book, Linux Systems Internalsand it was one of those, like, halfway through, “Can I please be excused? My brain is full.” It’s one of those things that I don’t use most of it on a day-to-day basis, but it’s solidified by understanding of what the computer is actually doing in a way that I will always be grateful for.

Liz: Mmm, and there are tens of millions of lines of code in the Linux kernel, so anyone who can internalize any of that is basically a superhero. [laugh].

Corey: I have nothing but respect for people who can pull that off.

Corey: Couchbase Capella Database-as-a-Service is flexible, full-featured and fully managed with built in access via key-value, SQL, and full-text search. Flexible JSON documents aligned to your applications and workloads. Build faster with blazing fast in-memory performance and automated replication and scaling while reducing cost. Capella has the best price performance of any fully managed document database. Visit couchbase.com/screaminginthecloud to try Capella today for free and be up and running in three minutes with no credit card required. Couchbase Capella: make your data sing.

In your day job, quote-unquote—which is sort of a weird thing to say, given that you are working at an open-source company; in fact, you are the Chief Open Source Officer, so what you’re doing in the community, what you’re exploring on the open-source project side of things, it is all interrelated. I tend to have trouble myself figuring out where my job starts and stops most weeks; I’m sympathetic to it. What inspired you folks to launch a company that is, “Ah, we’re going to be in the open-source space?” Especially during a time when there’s been a lot of pushback, in some respects, about the evolution of open-source and the rise of large cloud providers, where is open-source a viable strategy or a tactic to get to an outcome that is pleasing for all parties?

Liz: Mmm. So, I wasn’t there at the beginning, for the Isovalent journey, and Cilium has been around for five or six years, now, at this point. I very strongly believe in open-source as an effective way of developing technology—good technology—and getting really good feedback and, kind of, optimizing the speed at which you can innovate. But I think it’s very important that businesses don’t think—if you’re giving away your code, you cannot also sell your code; you have to have some other thing that adds value. Maybe that’s some extra code, like in the Isovalent example, the enterprise-related enhancements that we have that aren’t part of the open-source distribution.

There’s plenty of other ways that people can add value to open-source. They can do training, they can do managed services, there’s all sorts of different—support was the classic example. But I think it’s extremely important that businesses don’t just expect that I can write a bunch of open-source code, and somehow magically, through building up a whole load of users, I will find a way to monetize that.

Corey: A bunch of nerds will build my product for me on nights and weekends. Yeah, that’s a bit of an outmoded way of thinking about these things.

Liz: Yeah exactly. And I think it’s not like everybody has perfect ability to predict the future and you might start a business—

Corey: And I have a lot of sympathy for companies who originally started with the idea of, “Well, we are the project leads. We know this code the best, therefore we are the best people in the world to run this as a service.” The rise of the hyperscale cloud providers has called that into significant question. And I feel for them because it’s difficult to completely pivot your business model when you’re already a publicly-traded company. That’s a very fraught and challenging thing to do. It means that you’re left with a bunch of options, none of
them great.

Cilium as a project is not that old, neither is Isovalent, but it’s new enough in the iterative process, that you were able to avoid that particular pitfall. Instead, you’re looking at some level of making this understandable and useful to humans, almost the point where it disappears from their level of awareness that they need to think about. There’s huge value in something like that. Do you think that there is a future in which projects and companies built upon projects that follow this model are similarly going to be having challenges with hyperscale cloud providers, or other emergent threats to the ecosystem—sorry, ‘threat’ is an unfair and unkind word here—but changes to the ecosystem, as we see the world evolving in ways that most of us did not foresee?

Liz: Yeah, we’ve certainly seen some examples in the last year or two, I guess, of companies that maybe didn’t anticipate, and who necessarily has a crystal ball to anticipate how cloud providers might use their software? And I think in some cases, the cloud providers has not always been the most generous or most community-minded in their approach to how they’ve done that. But I think for a company, like Isovalent, our strong point is talent. It would be extremely rare to find the level of expertise in, you know, what is a pretty specialized area. You know, the people at Isovalent who are working on Cilium are also working on eBPF itself, and that level of expertise is, I think, pretty unrivaled.

So, we’re in such a new space with eBPF, we’ve only in the last year or so, got to the point where pretty much everyone is running a kernel that’s new enough to use eBPF. Startups do have a kind of agility that I think gives them an advantage, which I hope we’ll be able to capitalize on. I think sometimes when businesses get upset about their code being used, they probably could have anticipated it. You know, if it’s open-source, people will use your software, and you have to think of that.

Corey: “What do you mean you’re using the thing we gave away for free and you’re not paying us to use it?”

Liz: Yeah.

Corey: “Uh, did you hear what you just said?” Some of this was predictable, let’s be fair.

Liz: Yeah, and I think you really have to, as a responsible business, think about, well, what does happen if they use all the open-source code? You know, is that a problem? And as far as we’re concerned, everybody using Cilium is a fantastic… thing. We fully welcome everyone using Cilium as their data plane because the vast majority of them would use that open-source code, and that would be great, but there will be people who need that extra features and the expertise that I think we’re in a unique position to provide. So, I joined Isovalent just about a year ago, and I did that because I believe in the technology, I believe in the company, I believe in, you know, the foundations that it has in open-source.

It’s a very much an open-source first organization, which I love, and that resonates with me and how I think we can be successful. So, you know, I don’t have that crystal ball. I hope I’m right, we’ll find out. We should do this again, you know, a couple of years and see how that’s panning out. [laugh].

Corey: I’ll book out the date now.

Liz: [laugh].

Corey: Looking back at our conversation just now, you talked about open-source, and business strategy and how that’s going to be evolving. We talked about the company, we talked about an incredibly in-depth, technical product that honestly goes significantly beyond my current level of technical awareness. And at no point in any of those aspects of the conversation did you talk about it in a way that I did not understand, nor did you come off in any way as condescending. In fact, you wrote an O’Reilly book on Container Security that’s written very much the same way. How did you learn to do that? Because it is, frankly, an incredibly rare skill.

Liz: Oh, thank you. Yeah, I think I have never been a fan of jargon. I’ve never liked it when people use a complicated acronym, or really early days in my career, there was a bit of a running joke about how everything was TLAs. And you think, well, I understand why we use an acronym to shorten things, but I don’t think we need to assume that everybody knows what everything stands for. Why can’t we explain things in simple language? Why can’t we just use ordinary terms?

And I found that really resonates. You know, if I’m doing a presentation or if I’m writing something, using straightforward language and explaining things, making sure that people understand the, kind of, fundamentals that I’m going to build my explanation on. I just think that has a—it results in people understanding, and that’s my whole point. I’m not trying to explain something to—you know, my goal is that they understand it, not that they’ve been blown away by some kind of magic. I want them to go away going, “Ah, now I understand how this bit fits with that bit,” or, “How this works.” You know?

Corey: The reason I bring it up is that it’s an incredibly undervalued skill because when people see it, they don’t often recognize it for what it is. Because when people don’t have that skill—which is common—people just write it off as oh, that person’s a bad communicator. Which I think is a little unfair. Being able to explain complex things simply is one of the most valuable yet undervalued skills that I’ve found in this entire space.

Liz: Yeah, I think people sometimes have this sort of wrong idea that vocabulary and complicated terms are somehow inherently smarter. And if you use complicated words, you sound smarter. And I just don’t think that’s accessible, and I don’t think it’s true. And sometimes I find myself listening to someone, and they’re using complicated terms or analogies that are really obscure, and I’m thinking, but could you explain that to me in words of one syllable? I don’t think you could. I think you’re… hiding—not you [laugh]. You know, people—

Corey: Yeah. No, no, that’s fair. I’ll take the accusation as [unintelligible 00:31:24] as I can get it.

Liz: [laugh]. But I think people hide behind complex words because they don’t really understand them sometimes. And yeah, I would rather people understood what I’m saying.

Corey: To me—I’ve done it through conference talks, but the way I generally learn things is by building something with them. But the way I really learn to understand something is I give a conference talk on it because, okay, great. I can now explain Git—which was one of my early technical talks—to folks who built Git. Great. Now, how about I explain it to someone who is not immersed in the space whatsoever? And if I can make it that accessible, great, then I’ve succeeded. It’s a lot harder than it looks.

Liz: Yeah, exactly. And one of the reasons why I enjoy building a talk is because I know I’ve got a pretty good understanding of this, but by the time I’ve got this talk nailed, I will know this. I might have forgotten it in six months time, you know, but [laugh] while I’m giving that talk, I will have a really good understanding of that because the way I want to put together a talk, I don’t want to put anything in a talk that I don’t feel I could explain. And that means I have to understand how it works.

Corey: It’s funny, this whole don’t give talks about things you don’t understand seems like there’s really a nouveau concept, but here we are, we’re [working on it 00:32:40].

Liz: I mean, I have committed to doing talks that I don’t fully understand, knowing that—you know, with the confidence that I can find out between now and the [crosstalk 00:32:48]—

Corey: I believe that’s called a forcing function.

Liz: Yes. [laugh].

Corey: It’s one of those very high-risk stories, like, “Either I’m going to learn this in the next three months, or else I am going to have some serious egg on my face.”

Liz: Yeah, exactly, definitely a forcing function. [laugh].

Corey: I really want to thank you for taking so much time to speak with me today. If people want to learn more, where can they find you?

Liz: So, I am online pretty much everywhere as lizrice, and I am on Twitter. I’m on GitHub. And if you want to come and hang out, I am on the Cilium and eBPF Slack, and also the CNCF Slack. Yeah. So, come say hello.

Corey: There. We will put links to all of that in the [show notes 00:33:28]. Thank you so much for your time. I appreciate it.

Liz: Pleasure.

Corey: Liz Rice, Chief Open Source Officer at Isovalent. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an angry comment containing an eBPF program that on every packet fires off a Lambda function. Yes, it will be extortionately expensive; almost half as much money as a Managed NAT Gateway.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Emily

Emily is an Android engineer by day, but makes tech jokes and satires videos by night. She lives in San Francisco with two ridiculously fluffy dogs.

Links:

  • Uber: https://eng.uber.com/
  • Blog: https://www.emilykager.com/
  • Twitter: https://twitter.com/EmilyKager
  • TikTok: https://www.tiktok.com/@shmemmmy

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Couchbase Capella Database-as-a-Service is flexible, full-featured and fully managed with built in access via key-value, SQL, and full-text search. Flexible JSON documents aligned to your applications and workloads. Build faster with blazing fast in-memory performance and automated replication and scaling while reducing cost. Capella has the best price performance of any fully managed document database. Visit couchbase.com/screaminginthecloud to try Capella today for free and be up and running in three minutes with no credit card required. Couchbase Capella: make your data sing.

Corey: This episode is sponsored by our friends at Oracle HeatWave is a new high-performance query accelerator for the Oracle MySQL Database Service, although I insist on calling it “my squirrel.” While MySQL has long been the worlds most popular open source database, shifting from transacting to analytics required way too much overhead and, ya know, work. With HeatWave you can run your OLAP and OLTP—don’t ask me to pronounce those acronyms again—workloads directly from your MySQL database and eliminate the time-consuming data movement and integration work, while also performing 1100X faster than Amazon Aurora and 2.5X faster than Amazon Redshift, at a third of the cost. My thanks again to Oracle Cloud for sponsoring this ridiculous nonsense.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Today’s episode is a little bit off of the beaten path because, you know, normally we talk to folks doing things in the world of cloud. What is cloud, you ask? Great question. Whatever someone’s trying to sell you that day happens to be cloud.

But it usually looks like SaaS products, Platform as a Service products, Infrastructure as a Service products, with ridiculous names because no one ever really thought what that might look like to pronounce out loud. But today, we’re going in a completely different
direction. My guest is Emily Kager, a senior Android engineer at a small scrappy startup called Uber. Emily, thank you for joining me.

Emily: Thanks for having me.

Corey: So, I’m going to outright come out and say it I know remarkably little about, I don’t even want to say the mobile ecosystem in general, but even Android specifically because I fell down the iPhone hole a long time ago, and platform lock-in is a very real thing. Whenever you start talking about technical things, that generally tends to sail completely past me. You’re talking about things like Promises and whatnot. And it’s like, oh, that sounds suspiciously close to JavaScript, a language that I cannot make sense of to save my life. And it’s clear you know an awful lot about what you’re doing. It’s also clear, I don’t know, a whole heck of a lot about that side of the universe.

Emily: Well, that’s good because I don’t know much about the cloud.

Corey: Exactly. Which sounds like well, we don’t have a whole lot of points of commonality to have a show on, except for this small little thing, where recently, I decided in an attempt to recapture my lost youth and instead wound up feeling older than I ever have before, I joined the TikToks and started making small videos that I would consider humorous, but almost no one else will. And okay, great. I give it a hearty, sensible chuckle and move on, and then I start scrolling to see what else is out there. And I started encountering you, kind of a lot.

And oh, my God, this is content that it’s relatable, it is educational, dare I say, and most of all, it’s engaging without being overbearing. And this is a new type of content creation that I hadn’t really spent a lot of time with before. So, I want to talk to you about that.

Emily: Awesome. I want to apologize for having to see my face as you’re just scrolling throughout your day, but happy to chat about it. [laugh].

Corey: No, no, it’s—compared to some of the things I wind up on the TikTok algorithm, it is ridiculous. I think it’s about 80% confident that I’m a lesbian for some Godforsaken reason. Which hey, power to the people. I don’t think I qualify, but you know, that’s just how it works. And what I found really interesting about it, what does tie it back to the world of cloud, is that a recurring theme of this show has been, since the beginning, where does the next generation of cloud-engineering-type come from?

Because I’ve been in this space, almost 20 years, and it turns out that my path of working to help desk until you realize that you like the computers, but not so much being screamed at by the general public, then go find a unicorn job somewhere you can bluff your way into because the technical interviewer is out sick that day, and so on and so forth, isn’t really a path that is A) repeatable by a whole lot of people, and B) something that exists anymore. So, how do people who are just entering the workforce now or transitioning into tech from other fields learn about this stuff? And we’ve had a bunch of people talking about approaches to educating people on these sorts of things, but I don’t think I’ve ever spoken to someone who’s been as effective at it in minute or less long videos as you are.

Emily: That’s super kind. Yeah, I think there’s actually a whole discussion and joke set on TikTok of people’s parents suggesting why don’t you just go slide your resume under the CEOs door? Like, why don’t you just go get a job [laugh] that way? I think the realities of—what year are we in? 2022? [laugh]—

Corey: All year long, I’m told.

Emily: Yeah, [laugh] yeah. Yeah. I think that’s not going to be the reality anymore, right? You can’t just go shake hands with the CEO and work your way up from the mailroom and yeah, that’s not the way anymore. So yeah, I think I, you know, started just putting some feelers out, making educational content mostly about my own experiences as a change career person in the tech world.

I have some, I would say interesting perspectives on how to enter the industry, you know, either through undergrad or after undergrad, so. And it’s done really well. I think people are really interested in tech is a career at this point. Like, it’s kind of well known that they’re good jobs, well paid, and, you know, pretty, like, good work-life balance, most of the time. So yeah, the youth are interested.

Corey: It’s something that offers a path forward that lends itself to folks with less traditional backgrounds. For example, you have a master’s degree; I have an eighth-grade education on paper. And, yes, I’m proof-positive that it is possible to get into this space and, by some definitions, excel in it without having a degree, but let’s also be clear, here, I have the winds of privilege at my back, and I was stupendously lucky. It is harder to do without the credential than it is with the credential.

Emily: Yep.

Corey: But the credential is not required in the same way that it is if I want to be a surgeon. Yeah, you’re going to spend a lot of time in either school or prison with that approach. So, you have really two paths there; one is preferable over the other. Tech, it feels like there’s always more than one way to get in. And there’s always, it seems, as many stories as there are people out there about how they wound up approaching their own path to it. What was yours?

Emily: Yeah. First of all, it’s funny, you mentioned surgeons because I actually just today saw on my ‘For You’ page some surgeons sharing, you know, their own suturing techniques. And I think it’s a really interesting platform even, you know, within different fields and different subsets to kind of share information and keep up to date and connect with people in your own industry. So, beyond learning how to get into [laugh] an industry, it can also be helpful for other things. But sorry, I completely forgot the original question. How—what was my path? Is that what the question was?

Corey: Yeah. How did you get here is always a good question. It’s the origin stories that we sometimes tell, sometimes we wind up occluding aspects of it. But I find it’s helpful to tell these stories just because, if nothing else, it reaffirms to folks who are watching or listening or reading depending on how they want to consume this, that when they feel like well, I tried to get a credential and didn’t succeed, or I applied for a job and didn’t get it, there are other paths. There is not only one way to get there.

Emily: Yeah. And I think it’s also super important to talk about failures that we’ve had, right? So, when I was in undergrad, I was studying neuroscience and I was pre-med. And I thought I wanted to go to med school, kind of decided halfway through, I was only lukewarm about it, and I don’t think med school is the type of thing that you want to feel lukewarm about as you’re [laugh] approaching, you know, hundreds of thousands of dollars of debt and a ten-plus year commitment to schooling and whatever else, right? So yeah, I felt very lukewarm about the whole thing.

Both my parents were doctors, so I just didn’t really have exposure to many other careers or job options. I’m from a pretty, like, rural area, so tech had never really [laugh] occurred to me either. So yeah, then I decided to just take a year off after undergrad, felt super lost. I think when you’re 22, everything feels so important, [laugh] and you look at everyone else who already has their first job at 22, and I was like, “Wow, I’m a huge failure. I’m never going to have a job.” Which is, you know, hilarious looking back because 22-year-olds are so young. And yeah, just decided to take a year off. I worked at a nonprofit. I hated it, hated the work. Decided, like I, you know, can never do this forever.

Corey: I can’t do nonprofit stuff. I’m going to do for-profit stuff. And it turns out that most—when you say nonprofit, it doesn’t mean what I thought. It ap—usually means, you know, something that’s dedicated to a charitable cause, not, you know, a VC-backed company that doesn’t know how to make any money.

Emily: Yeah. I mean, it could still be very corporate at nonprofit. After that, actually—

Corey: Oh, yes. Money is the root of all good as well as evil.

Emily: Yeah. And I actually had a task at the nonprofit where I was sorting a ton of things in spreadsheets. And I was like, wow, it’d be easy if there was just, like, some program I could write to, like, do this. So, I actually reached out to my brother, who was a computer science nerd—affectionately—and he helped me write some, like, Excel macros, and I was like, “This is so cool.” And I ended up taking a free course, CS50, which is great, by the way, great course, super high quality from Harvard, totally free to take online.

And really liked it, so I did something a little crazy and decided to just dive right in. [laugh]. And I applied to a post-bacc program to kind of take all the courses that a CS undergrad would have taken just after. And that post-bacc turned into a master’s program.

Corey: And here you are now on the other side of having done it. If—sort of the dangerous questions: If you had known then what you know now, would you have gone down the same path, or would you have done something different to get into the space?

Emily: Yeah, I mean, I think it’s hard once you’ve kind of made it, to be like, “I would change all this.” I think I would probably try more things in undergrad. That would be the real answer to that. It obviously would have been a lot easier and more time-efficient if I didn’t have to go back to school and do something. But that being said, I don’t think that getting a post-bacc or a Master's is the only way into tech; it was just my path.

And I try not to… I try not to promote other paths that I don’t really know much about independently, right? So—on me. So—but plenty of people are successful going through boot camps or self-teaching, even, I think they’re just much more difficult paths because the reality is, like, having a degree is still definitely an easier path when you show up to an interview and you can just kind of show your piece of paper, which, for better or worse, that’s the reality sometimes.

Corey: My wife’s a corporate attorney, so I’ve been law adjacent for over a decade now, and one of the things that always struck me about that field is the big law approach is you go to a top-tier law school, you wind up putting your nose to the grindstone for all three years, and you hope to get an offer at one of the big law firms. And they all keep their salaries in lockstep. I think right now they’re all—they just upgraded again to $235,000 a year starting. And if you don’t get one of those rare, prestigious jobs at a number of select firms, it’s almost a bimodal distribution where you’re making somewhere between 60 and $80,000 a year to start somewhere else. It is the one path to make big money in law as you’re fresh out of school, and there are no real do-overs in most cases.

So, it’s easy to apply that type of thinking to tech, and it’s just not true. Talking to folks who have this dream of working at Google and they finally go through the interview process. And it turns out that oh no, they froze when asked to solve Fizz Buzz, or invert a binary tree on a whiteboard, or whatever ridiculous brainteaser question they’re being asked, and, “Oh, no, my life is over.” And it’s, you know, you can go to, I don’t know, Stripe, two blocks down the street and try again. And if that doesn’t work, Microsoft, or Amazon, or go down the entire list of tech companies you’ve heard of and haven’t heard of, and they all compensate directionally the same way. It’s not a one-shot, ‘this is it’ moment in the same way. And I—

Emily: Yeah.

Corey: —I think that’s a unique thing to tech right now.

Emily: Yeah, definitely. And I think a lot of kids—I say kids, but really, like, you know, 18 to 20-year-olds—

Corey: Oh, believe me, after being on TikTok for a couple of weeks, let me say that every one of you are children, to my perspective. I am
now Grandpa Quinn over here.

Emily: [laugh]. I’ll take it. Yeah, but a lot of them have reached out like, “I didn’t get hired at FAANG right out of school. Is my life over? Is my career over?” And I’ve never worked at a FAANG. [laugh]. I’m pretty happy. I definitely think I have a successful career, and I almost think I’m better for not having gone right into it, you know?

I think it can be great for some people. There’s great, you know… definitely great salaries, great mentorship options, but it’s not the only option. And I think maybe tech is unique in that way, but there’s just so many good companies to work at, and so many great opportunities, you really don’t need to go to the name brand in the same way that maybe you would have to in law. It’s funny you say that because my partner is also a lawyer [laugh] and [crosstalk 00:13:00]—

Corey: Oh, dear. We should start a support group of our own, on some level.

Emily: I know, yeah. He just went through the whole big law recruiting thing. So, I know much about that. [laugh].

Corey: It’s always an experience. The way that I have found across the board as well is there’s also a shared, I guess, esprit de corps
almost across the industry. I mean, you are on the Android side of the world, and I historically was on the DevOps side of the universe, although now mocking cloud services—but not the way test engineers say when they use the term ‘mocking’—is what I do. But there are
shared experiences that tie us together, and that’s part of what I found so interesting about a lot of your content.

Because yes, there is some of the deep dive stuff into Android and, cool, sails right over my head—I hear the whistling sound vaguely as it goes over—but then there’s other stories about things that are unique—that are, I guess, a shared experience. For me, one of the things that tied all of tech together, regardless of where in the ecosystem you fit in, is a shared sense of being utterly intimidated to hell by the miracle of Git, where it’s like, Git’s entire superpower is making you feel dumb. Doesn’t matter who you are, from someone who doesn’t know what Git is all the way to Linus himself. Someone is go—at some point, you’re going to look at it and wonder, “What the hell is going on?” It’s just a question of how far you get along the path before it changes your understanding of the universe.

And I wound up starting to give talks, in the before times, at front-end conferences about this, which you want to talk about dispiriting things. I would build slides like, you know, a DevOps person would: Black Helvetica text on a white slide. Everyone else has these
beautifully pristine, great slides. I have 20 minutes to go.

How can I fix it? Change the font to Comic Sans because if you’re going to have something that looks crappy, make it look like it was intentionally so.

Emily: And did it work?

Corey: Oh, it worked swimmingly. It was fantastic. I like the idea of being able to reach people in different areas, no matter where they are in their journey, and one of the things that appeals to me about TikTok in general in your content in particular, is it seems like we have something of a shared perspective on, getting people’s attention is required in order to teach them something, and I think we both use the same vehicle for that, which is humor.

Emily: Yeah, I would agree. I think the other interesting thing I just wanted to touch on; you were talking about is, we don’t really know too much about each other’s fields in tech. And I think when you’re talking to a younger audience, maybe who you want to get interested in tech, it’s really hard to communicate all the different avenues into tech that they can take. And this is something that I’m still struggling with because I know my experience as an Android developer, a mobile developer, I probably medium I understand, you know, back end development, but I don’t think I could explain to a college student why or what even is, [laugh] you know, cloud development and how they could get involved in that, or all these other fields that I just really don’t know much about. And I think that’s kind of what ties a lot of people in tech together as well, right? Because we know our little corners of the world, and you have to start to get comfortable with the things that you don’t know. And I think that’s really hard to explain to [laugh] the younger generation as you’re trying to get them excited about things.

Corey: Oh, yeah. And the reality, too, of what we tell people and how the world works is radically different. Like, I want to learn a technology that will absolutely last for an entire career and then some, and I want to be able to be employed anytime, anywhere, at any company. The easy slam dunk answer that I think will not change in either of our lifetimes is Microsoft Excel. It powers the world.

People think I’m kidding, but it is the IDE of back-office processes and communications. If Excel were to go away or even worse,
Microsoft were to change Excel’s interface, people would be storming Redmond by noon.

Emily: Yeah, I believe it. Yeah, you know, it’s interesting, right? Like, it’s hard to tell people—because people will tell to me, “Well, do you have to keep learning things?” And I’m like, “Yeah. You got to keep learning things, like, all the time.”

But I don’t think that should be, you know, a deterrent from the career; it’s just a reality. But to try to manage, like, the fears a lot of people have coming into tech and also encouraging them to still, you know, try it, go after it, I think that’s something I struggle with when I’m creating my content for—towards, like, younger people. [laugh].

Corey: Today’s episode is brought to you in part by our friends at MinIO the high-performance Kubernetes native object store that’s built for the multi-cloud, creating a consistent data storage layer for your public cloud instances, your private cloud instances, and even your edge instances, depending upon what the heck you’re defining those as, which depends probably on where you work. It’s getting that unified is one of the greatest challenges facing developers and architects today. It requires S3 compatibility, enterprise-grade security and resiliency, the speed to run any workload, and the footprint to run anywhere, and that’s exactly what MinIO offers. With superb read speeds in excess of 360 gigs and 100 megabyte binary that doesn’t eat all the data you’ve gotten on the system, it’s exactly what you’ve been looking for. Check it out today at min.io/download, and see for yourself. That’s min.io/download, and be sure to tell them that I sent you.

Corey: Something I found on Twitter is that among other things that Twitter has going on for it, it doesn’t do nuance, it does, effectively, things that are black and white, yes or no, it’s always a binary in many respects. And one of those is that, like, should—like, is passion or requirement for working in tech. And there’s the, “Yes, you absolutely have to be passionate for this and power through it.” And the answer, “No, you don’t need to be passionate about it’s okay to do it for the money and not kill yourself working 20 hours a day.” And from my perspective, I take a more moderate stance, which is how you get both sides of that argument to hate you, but it’s, I don’t think you need to have this all-consuming drive for tech, but I do think you need to like it.

Emily: Oh yeah.

Corey: I think you need to enjoy what you’re doing or it’s going to feel like unmitigated toil and misery, and you will not be happy in the space. And if you’re not happy, really is the rest of it all worth it?

Emily: I think that applies to most careers, though, right? Like that—definitely, when I was looking to switch careers, that was the main thing I was looking for. Number one was like, you know, pretty solid salary. And number two was, do I just not hate it? [laugh]. And I think if you’re doing anything and you hate it, you’re going to be miserable, right?

Like, even if you’re doing it to make a paycheck if you actually hate every single day when you wake up in the morning and you dread, you know, going to bed because the next morning, you have to wake up and do it again, like, you’re going to be miserable. But I do think, yeah, like, to your point, there’s a middle ground in all this, right? You don’t have to dream about tech, but I think you do have to realize that, yeah, if you’re going to be in this industry for decades, you’re going to have to be able to learn and be interested enough in things that, you know, learning isn’t a huge slog either. So.

Corey: I’ve never understood the folks who don’t want to learn as they go through their career because it just seems like a recipe to do the same thing every year for 40 years, and then you retire with what 40 years of experience—one year experience repeated 40 times. It’s a… any technology or any disruption change happens, and suddenly you’re in a very uncomfortable situation when we’re talking about knowledge workers.

Emily: Yeah, I think people—you know, I think we talk a lot about, like, imposter syndrome in our industry right? So, I think people already feel like maybe, “I don’t know anything so why would I put myself out there and learn new things?” I mean, I definitely sometimes struggle with this where I’m like, “I’m very comfortable [laugh] in, like, what I do day-to-day. I know what I’m doing.” So yeah, when you have to learn, like, a totally new language or new architecture, whatever, it can feel very overwhelming to be like, wow, I actually am, you know, super stupid. [laugh]. But it’s just new things, right? You’re learning new things, and—

Corey: Like, “Find the imposter. Oh, no, it’s me.” Yes, it’s a consistent problem.

Emily: But it’s a really powerful thing to acknowledge that you can feel stupid and you can ask questions and you can be new to something, and that’s, like, totally valid. And I started taking a new language course a year or two ago, and showing up every day and speaking a new language and feeling like an idiot, it was actually super empowering because everyone in the class is doing it, you know? We didn’t know the language and we were just, you know, talking gibberish to each other, and that’s fine. We were learning.

Corey: The emotional highs and lows are also—they hit quickly. I have never felt smarter or dumber in a two-minute span of each other than when working on technology. It’s one of those, “I will never understand how this works—oh my God, it works. I’m a genius. Just kidding. It doesn’t work. Nevermind. Forget everything I just said.” It’s a real emotional roller coaster.

Emily: [laugh]. There’s only two ends of the spectrum, right? Like, there’s no middle ground in this situation. It’s, “I’m a genius,” or, “I should quit and never work on technology ever again.”

Corey: So, I’ve been experimenting on TikTok a bit and you’ve been on it significantly longer. You have, as of this recording, something in the direction of 65,000 followers on the TikToks. I have a bit more than that on the Twitters, which only took me a brief 14 years to do. So, great. I’ve noticed that as I wind up—as you hit certain inflection points on Twitter, your experience definitely changes, when—as far as just, like, the unfortunate comments coming out of the woodwork.

Like, I was making fun of LinkedIn at some point, and then there was some troll comment in the comments, and I looked at who the commenter was and it was the official LinkedIn brand account. And okay, well, that’s novel, but all right. I’d like to add them to my professional network on TikTok. So, there we go. But have you noticed inflection points as well, in your—experience changes on the platform as you continue to grow?

Emily: Yeah. I think—I saw something once that Twitter is only fun if you have less than, like, [laugh] 5000 followers or something. So, I think we both surpassed that a while ago. And yeah, I think it can be a very interesting experience as you start to gain followers. And to be honest, like, I’m on both platforms, just to kind of make content.

It’s a very, like, creative outlet for me. I don’t necessarily care that much about how many followers I have. But it is an interesting progression to see, like, you know, you get a little bit of engagement, and it’s usually, like, a back and forth; you’re kind of like actually connecting to people, and then as you kind of surpass maybe five or ten-thousand followers, there’s all these people who come in who you don’t know who they are, they don’t know who you are, they make assumptions about you, they are saying really mean things that I think just because you have, like, a high follower account that they’re like, “I can say whatever I want to this person.” And it’s definitely an interesting change. I think over the years—because I’ve been fairly public for a number of years now—you kind of get more immune to it. I’m sure you feel the same way, but you’re like, whatever, just kind of brush off a lot of these things. But—

Corey: Oh, yeah. You become more of a persona to people than an actual person.

Emily: Yeah.

Corey: And that is—

Emily: Yeah.

Corey: —people forget that—you know, everyone yells at you about, “That was an unkind thing, express more empathy all the ti”—I mean, you get that all the time when you get—when set a slight foot wrong. And they’re right—don’t think I’m saying otherwise—but they’re not expressing a lot of empathy for you at the same time, either. So, it’s one of those you have to disengage and disconnect on certain levels and just start to ignore it. But it’s been a wild ride.

Emily: I used to wonder, I used to see, like, accounts that have you know, 50, 60,000 followers on Twitter back when I was a smaller account, and they didn’t—they never tweeted, and I was like, “How’d they get so many followers? They never tweet.” And now I understand. It’s that they gained that many followers and then they left. [laugh]. They’re done.

Corey: [unintelligible 00:23:18] like, “This platform sucks now.” And it’s—a lot of folks, like, “Oh, Twitter’s not as good as it used to be.” It’s like, well hang on. Has the platform itself changed or has your exposure to it changed? And it’s a question that doesn’t really have a great answer or way to find out, but it’s… it’s been a—it’s an ongoing struggle for folks. And I do have empathy for that. I try to avoid getting involved in pile-ons wherever possible.

Emily: Yeah. That’s been a new change for me, too. I think a lot of my early brand on Twitter—as dumb as that word is—was, you know, kind of finding, like, misogynists in tech and really, like, creating a pile-on on them. And, you know, I think there is a space for calling out bad behavior in the industry, but you want to be careful because really, there are other people on the other side of the screen. And unless someone’s really implying—like, unless they’re really intending ill intent, you know, I think I’ve kind of now moved less towards that type of [laugh] pile-on. It is fun though. That’s the thing. It’s fun.

Corey: Plus the algorithm rewards engagement. Say horrifying things and get a bunch of attention and more followers. But you don’t necessarily want to participate in that.

Emily: Yeah, exactly. And that’s the other thing I realized that if someone is really saying something stupid, me bringing attention to it is only going to amplify it more. So. Especially as you gain followers and you have more of an audience to whatever you quote, tweet, or retweet, or comment on, right? So.

Corey: As I look at, like, the sheer amount of content that you’ve put out—it’s weird because if someone asked me this question, I don’t know that I would have a good answer, but I am curious. You are consistently exploring new boundaries in terms of the humor, the content, the topics, the rest. How do you come up with it?

Emily: This is going to be a really unsatisfying answer. [laugh]. I don’t know. [laugh]. I’m a runner, and a lot of times when I’m running I don’t use headphones. A lot of people say I’m sociopathic because I just am by myself in the world, and—this is such, like, a weird answer—but yeah, I just kind of—I’m thinking about things, usually I’m like digesting my day, things that happened, things that were annoying.

And to be honest, I think it’s pretty easy to identify things that are relatable, right? So, a lot of the gripes that all engineers have, right? So, you’re like, “Wow, it was really annoying that I had to make a ticket in Jira today.” And you can kind of think about how is it annoying, and how can I make this funny and relatable to someone else? So—and to be hon—like, when I had, you know, a group of coworkers that I worked really closely in my last job, I would just send them the jokes, and then if they thought it was funny, I would just, like, post it on Twitter.

And that’s kind of… you know, it’s just, like, the basic chit-chat that you do. But now we’re all remote, so I found an outlet through Twitter and TikTok, where I would just express all my, you know, stupid engineering jokes to the world. [laugh]. Whether they want it or not.

Corey: Something I found is that—and it always has frustrated me, and I figured, one day, I too, would figure out how to solve for this. And no. There are things I will tweet out that I think are screamingly funny and hilarious, and no one cares. Conversely, I’ll jot off something right before I dive into a meeting, and I’ll come back and find out it’s gone around the internet three times. And there seems to be no rhyme or reason to it, other than that my sense of humor is not quite dialed into exactly where most folks in this industries are. It’s close enough that could be overlooked, but I still feel like the best jokes go unappreciated.

Emily: Oh, I agree. I mean, I send jokes by friends all the time that I’m like, “I’m posting this,” and it gets, like, you know, 20 likes. And I don’t even care. I think, you know—I think that’s the—you know can—you start to learn as a content creator that you’re like, “I’m going to put out the content that I want to put out and hope other people find it funny, but at the end of the day, I don’t really care.” So, I’m laughing at my own jokes. I’ll admit that. So. I think they’re funny. My—

Corey: [crosstalk 00:26:58]—

Emily: —[crosstalk 00:26:58] funny, too.

Corey: —for me because if—I’m keeping myself engaged, otherwise it gets boring, and I lose interest in the sound of my own voice, which is just a terrible sin for me. So, it’s—I have to keep it engaging or I’ll lose interest.

Emily: Yeah, exactly.

Corey: Do you find when you’re trying to put together content, that—for TikTok, for example—that you’ve come up with something that, “Huh, this doesn’t really fit the video format. Maybe it’s more of a blog post or something else.” Do you find that one content venue feeds another? Do you reuse content across multiple platforms? And if so—

Emily: Yeah.

Corey: —what have you learned from all that?

Emily: That’s an interesting question. I think—I do maintain a blog, but I don’t post so often on it, and I find that the—for the more serious content I’m making that’s not jokes, right? I think TikTok just really hits a different audience. Like, people don’t find my blog, it’s not discoverable, maybe they’re not checking it, and I think definitely the younger audience prefers to consume things in video content. And a lot of my content is also aimed towards people who maybe are exploring tech who don’t work in tech yet, and so to really hit them, they probably aren’t following me and they probably don’t know who I am, they probably don’t even know what to look for in my blog.

So, for example, I have a blog post all about how I transitioned into tech, blah, blah, blah, and people still ask me all the time on TikTok, “How did you transition into tech? How did you”—I’m like, “It’s in my blog.” On my—like, you know, linked my bio. But you still have to just kind of—I think, like, I tend to just recreate the content into the different platforms. And it can be a bit tedious, but I try to keep my blog up to date with, like, different stories of things that have happened to me. But these days, I mostly just post on TikTok, to be honest. [laugh].

Corey: I had the same problem, but content reuse saved me. I started writing a long-form blog post of roughly 1000 to 1500 words every week, then reading it into a microphone. It became the AWS Morning Brief podcast and emailing out to the newsletter as well. So, it’s one piece of content used three different times, which was awesome, but then there’s the other side of it, which is, I need to come up with an interesting idea or concept or something to talk about for 1000 words every week, like clockwork. And one of the things that made this way easier is a tip I got from Scott Hanselman that I have been passing on whenever it seems appropriate—like in this conversation—which is if you find yourself explaining something a third time, turn it into a blog post because then you’ll just be able to link people to the thing that you wrote where you go into significantly more depth around what you’re talking about than you can in a two-tweet exchange, and that in turn, gives you a place to dump that stuff out.

And I found that has worked super well for me because once I’ve written it and gotten it out, I also often find I stopped making the same reference all the time because now I’ve said it, I’ve said my piece. Now, I can move on and come up with a second analogy, or a new joke or something.

Emily: Yeah. I’ve also found that um—that’s a great idea from Scott; he’s also great on the TikToks [laugh]—

Corey: Oh, yes he is.

Emily: —[crosstalk 00:29:45] [laugh]. Building his account. Yeah, I think another interesting thing is, specifically on TikTok and Twitter because it’s more of a conversation between you and your community, I tend to get a lot of ideas just from people asking me questions, right? So, in the comments of something, it could be related to the video I just made and it really helps me expand upon, you know, what I was just saying and maybe answer a follow-up question in a different video. Or maybe it’s just a totally unrelated question.

So, someone finds, you know, one of my comedy videos and is like, “Hey, you work in tech. Like, what is that like in San Francisco?” Right? So, I think I’ve found a ton of inspiration just from community people and really what they’re asking for, right? Because at the end of the day, you want to make content that people actually care about and want to know the answers to.

Corey: Yeah, seems like that does help. If it’s, “How do I wind up building a following or getting a lot of traffic or the rest?” And it’s Lord knows, once you have a website that has a certain amount of Google juice, you just get besieged by random requests from basically every channel. “Hey, I saw this great article linked to a back issue of the newsletter talking about this thing. Would you mind including my link to it, this would help your readers.” And it’s just it’s a pure SEO scam.

And it’s yeah, I don’t—my approach to SEO has been this, again, ancient, old-timey idea of I’m going to write compelling original content that ideally other people find valuable and then assume that the rest is going to take care of itself. Because, on some level, that is what all these algorithms are trying to do is surface the useful stuff. I feel like as long as you hold to that, you’re not going to go too far wrong.

Emily: No, that’s true. Also, something funny about reusing content is sometimes I’ll post a joke on Twitter, and if it does well, I’ll make it into a video format. And you know, sometimes I change the format of the joke around, whatever. But I—a couple times this happened—I’ll post something on Twitter, and then, like, a day or two later, I’ll make a TikTok about it, and a lot of people will come in and be like, “I already saw this joke on Twitter.” And they won’t know it’s from me, so they’re basically accusing me of joke stealing when really I’m just content-raising is what I should tell them. But it is funny. [laugh].

Corey: That’s happened me a couple times on Twitter. People are like, “Hey, that’s a stolen joke.” And then they’ll google it and they’ll dig it out. Like, “Here’s the original—oh, wait, you said it two years ago.” “Yeah. No one liked it then, so here we are.” “If you liked it then, why didn’t you blow it up like you did now?” So.

Emily: They remembered it from two years ago, but they didn’t remember it was yours. [laugh].

Corey: At some level, I feel like I could almost loop my Twitter account and just let it continue to play out again for the next seven years, and other than the live-streaming stuff and the live-tweeting various events, I feel like it would do fairly well, but who knows.

Emily: Yeah. Yeah. But at the end of the day, I think there’s also a finite amount of funny tech jokes, and we’re all just kind of recycling each other’s jokes at some point. So, I don’t get too offended by that. I’m like, “Sure. We all made the same joke about NFTs. Great.” Like, I don’t care. [laugh].

Corey: I really want to thank you for taking the time to speak with me today.

Emily: [crosstalk 00:32:36] been fun.

Corey: If people want to learn more and appreciate some of that awesome content, where’s the best place to find you?

Emily: Yeah, I’m on the Twitters and the TikToks, just like you.

Corey: Excellent. And we will, of course, put links to that in the [show notes 00:32:45].

Emily: Had a great time. Thank you so much for having me again.

Corey: No, thank you for coming. Emily Kager, senior Android engineer at Uber. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an angry comment that links to a TikTok video of you ranting for a solid minute, but because computers and phones alike are very hard, you’re using the wrong camera, and we just get that video of your floor.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Kelly

KellyAnn Fitzpatrick is a Senior Industry Analyst at RedMonk, the developer-focused industry analyst firm. Having previously worked as a QA analyst, test & release manager, and tech writer, she has experience with containers, CI/CD, testing frameworks, documentation, and training. She has also taught technical communication to computer science majors at the Georgia Institute of Technology as a Brittain Postdoctoral Fellow.

Holding a Ph.D. in English from the University at Albany and a B.A. in English and Medieval Studies from the University of Notre Dame, KellyAnn’s side projects include teaching, speaking, and writing about medievalism (the ways that post-medieval societies reimagine or appropriate the Middle Ages), and running to/from donut shops.

Links:

  • RedMonk: https://redmonk.com/
  • Twitter: https://twitter.com/drkellyannfitz

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Today’s episode is brought to you in part by our friends at MinIO the high-performance Kubernetes native object store that’s built for the multi-cloud, creating a consistent data storage layer for your public cloud instances, your private cloud instances, and even your edge instances, depending upon what the heck you’re defining those as, which depends probably on where you work. It’s getting that unified is one of the greatest challenges facing developers and architects today. It requires S3 compatibility, enterprise-grade security and resiliency, the speed to run any workload, and the footprint to run anywhere, and that’s exactly what MinIO offers. With superb read speeds in excess of 360 gigs and 100 megabyte binary that doesn’t eat all the data you’ve gotten on the system, it’s exactly what you’ve been looking for. Check it out today at min.io/download, and see for yourself. That’s min.io/download, and be sure to tell them that I sent you.

Corey: This episode is sponsored by our friends at Oracle HeatWave is a new high-performance query accelerator for the Oracle MySQL Database Service, although I insist on calling it “my squirrel.” While MySQL has long been the worlds most popular open source database, shifting from transacting to analytics required way too much overhead and, ya know, work. With HeatWave you can run your OLAP and OLTP—don’t ask me to pronounce those acronyms again—workloads directly from your MySQL database and eliminate the time-consuming data movement and integration work, while also performing 1100X faster than Amazon Aurora and 2.5X faster than Amazon Redshift, at a third of the cost. My thanks again to Oracle Cloud for sponsoring this ridiculous nonsense.

Corey: Welcome to Screaming in the Cloud, I’m Corey Quinn. It’s always a good day when I get to sit down and have a chat with someone who works over at our friends at RedMonk. Today is no exception because after trying for, well, an embarrassingly long time, my whining and pleading has finally borne fruit, and I’m joined by Kelly Fitzpatrick, who’s a senior industry analyst at RedMonk. Kelly, thank you for, I guess, finally giving in to my always polite, but remarkably persistent requests to show up on the show.

Kelly: Great, thanks for having me. It’s great to finally be on the show.

Corey: So, let’s start at the very beginning because I am always shockingly offended whenever it happens, but some people don’t actually know what RedMonk is. What is it you’d say it is that you folks do?

Kelly: Oh, I love this question. Because it’s like, “What do you do,” versus, “What are you?” And that’s a very big difference. And I’m going to start with maybe what we are. So, we are a developer-focused industry analyst firm. You put all those things, kind of, together.

And in terms of what we do, it means that we follow tech trends. And that’s something that many industry analysts do, but our perspective is really interested in developers specifically and then practitioners more broadly. So, it’s not just, “Okay, these are things that are happening in tech that you care about if you’re a CIO,” but what tech things affect developers in terms of how they’re building software and why they want to build software and where they’re building software?

Corey: So, backing it up slightly because it turns out that I don’t know the answer to this either. What exactly is an industry analyst firm? And the reason I bring this up is I’ve been invited to industry analyst events, and that is entirely your colleague, James Governor’s, fault because he took me out for lunch at I think it was Google Next a few years ago and said, “Oh, you’re definitely an analyst.” “Okay, cool. Well, I don’t think I am. Why should I be an analyst?”

“Oh, because companies have analyst budgets.” “Oh, you said, analyst”—protip: Never get in the way of people trying to pay you to do things. But I still feel like I don’t know what an analyst is, in this sense. Which means I’m about to get a whole bunch of refund requests when this thing airs.

Kelly: I should hope not. But industry analysts, one of the jokes that we have around RedMonk is how do we explain to our families what an industry analyst is? And I think even Steve and James, who are RedMonk’s founders, they’ve been doing this for quite a long time, like, much longer than they ever want to admit that they do, and they still are like, “Okay, how do I explain this to my parents?” Or you know, anyone else who’s asking, and partly, it’s almost like a very—a term that you’ll see in the tech industry, but outside of it doesn’t really have that much, kind of, currency in the same way that you can tell someone that you’re like, maybe a business analyst or something like that, or any of those, almost like spy-like versions of analyst. I think was it The Hunt for Red October, the actual hero of that is an analyst, but not the type of analyst that I am in any way, shape or form.

But you know, industry analyst firms, specifically, it’s like we keep up on what tech is out there. People engage with us because they want to know what to buy for the things that they’re doing and the things that they’re building, or how to better create and sell the stuff that they are building to people who build software. So, in our case, it’s like, all right, what type of tools are developers using? And where does this particular tool that our company is building fit into that? And how do you talk about that with developers in a way that makes sense to them?

Corey: On some level, what I imagine your approach to this stuff is aligns somewhat with my own. Before you became an industry analyst, which I’m still not entirely sure I know what that is—I’m sorry, not your fault; just so many expressions of it out there—before you wound up down that path, you were a QA manager; you wound up effectively finding interesting bugs in software, documentation, et cetera. And, on some level, that’s, I think, what has made me even somewhat useful in the space is I’ll go ahead and try and build something out of something that a vendor has released, and huh, the documentation says it should work this way, but I try it and it breaks and it fails. And the response is always invariably the same, which is, “That’s interesting,” which is engineering-speak for, “What the hell is that?” I have this knack for stumbling over weird issues, and I feel like that aligns with what makes for a successful QA person. Is that directionally correct, or am I dramatically misunderstanding things and I’m just accident-prone?

Kelly: [laugh]. No, I think that makes a lot of sense. And especially coming from QA where it’s like, not just making sure that something works, but making sure that something doesn’t break if you try to break it in different ways, the things that are not necessarily the expected, you know, behaviors, that type of mindset, I think, for me translated very easily to, kind of, being an analyst. Because it’s about asking questions; it’s about not just taking the word of your developers that this software works, but going and seeing if it actually does and kind of getting your hands dirty, and in some cases, trying to figure out where certain problems or who broke the build, or why did the build break is always kind of super fun mystery that I love doing—not really, but, like, everyone kind of has to do it—and I think that translates to the analyst world where it’s like, what pieces of these systems, or tech stacks, or just the way information is being conveyed about them is working or is not, and in what ways can people kind of maybe see things a different way that the people who are building or writing about these things did not anticipate?

Corey: From my position, and this is one of the reasons I sort of started down this whole path is if I’m trying to build something with a product or a platform—or basically anything, it doesn’t really matter what—and the user experience is bad, or there are bugs that get in my way, my default response—even now—is not, “Oh, this thing’s a piece of crap that’s nowhere near ready for primetime use,” but instead, it’s, “Oh, I’m not smart enough to figure out how to use it.” It becomes a reflection on the user, and they feel bad as a result. And I don’t like that for anyone, for any product because it doesn’t serve the product well, it certainly doesn’t serve the human being trying to use it and failing well, and from a pure business perspective, it certainly doesn’t serve the ability to solve a business problem in any meaningful respect. So, that has been one of the reasons that I’ve been tilting at that particular windmill for as long as I have.

Kelly: I think that makes sense because you can have the theoretically best, most innovative, going to change everyone’s lives for the better, product in the world, but if nobody can use it, it’s not going to change the world.

Corey: As you take a look at your time at RedMonk, which has been, I believe, four years, give or take?

Kelly: We’re going to say three to four.

Corey: Three to four? Because you’ve been promoted twice in your time there, let’s be very clear, and this is clearly a—

Kelly: That’s a very, very astute observation on your part.

Corey: It is a meteoric rise. And what makes that also fascinating from my perspective, is that despite being a company that is, I believe, 19 years old, you aren’t exactly a giant company that throws bodies at problems. I believe you have seven full-time employees, two of whom have been hired in the last quarter.

Kelly: That’s true. So, seven full-time employees and five analysts. So, we have—of that it’s five analysts, and we only added a fifth analyst the beginning of this year, with Dr. Kate Holterhoff. [unintelligible 00:08:09], kind of, bring her on the team.

So, we had been operating with, like, kind of, six full-time employees. We were like, “We need some more resources in this area.” And we heard another analyst, which if you talk about, okay, we hired one more, but when you’re talking about hiring one more and adding that to a team of, like, four analysts, it’s such a big difference, just in terms of, kind of, resources. And I think your observation about you ca—we don’t just throw bodies at problems is kind of correct. That is absolutely not the way we go about things at all.

Corey: At a company that is taking the same model that The Duckbill Group does—by which I mean not raising a bunch of outside money is, as best I can tell—that means that you have to fall back on this ancient business model known as making more money than it costs to run the place every month, you don’t get to do this massive scaled out hiring thing. So, bringing on multiple employees at a relatively low turnover company means that suddenly you’re onboarding not just one new person, but two. What has that been like? Because to be very clear, if you’re hiring 20 engineers or whatnot, okay, great, and you’re having significant turnover, yeah, onboarding two folks is not that big of a deal, but this is a significant percentage of your team.

Kelly: It is. And so for us—and Kate started at the beginning of this year, so she’s only been here for a bit—but in terms of onboarding another analyst, this is something where I haven’t done before, but, like, my colleagues have, whereas the other new member of our team, Morgan Harris, who is our Account Engagement Manager, and she is amazing, and has also, like, very interesting background and client success in, like, fashion, which is, you know, awesome when I’m trying to figure out what [unintelligible 00:09:48] fit I need to do, we have someone in-house who can actually give me advice on that. But that’s not something that we have onboarded for that role very much in the past, so bringing on someone where they’re the only person in their role and, like, having to begin to learn the role. And then also to bring in another analyst where we have a little bit more experience onboarding analysts, it takes a lot of patience for everybody involved. And the thing I love about RedMonk and the people that I get to work with is that they actually have that patience and we function very well as, like, a team.

And because of that, I think things that could really have thrown us off course, like losing an account engagement or onboarding one and then onboarding a new analyst, like, over the holidays, during a pandemic, and everything else that is happening, it’s going much more smoothly than it could have otherwise.

Corey: These are abnormal times, to be sure. It’s one of those things where it’s, we’re a couple years into a pandemic now, and I still feel like we haven’t really solved most of the problems that this has laid bare, which kind of makes me despair of ever really figuring out what that’s going to look like down the road.

Kelly: Yeah, absolutely. And there is very much the sense that, “Okay, we should be kind of back to normal, going to in-person conferences.” And then you get to an in-person conference, and then they all move back to virtual or, as in your case, you go to an in-person conference and then you have to sequester yourself away from your family for a couple of weeks to make sure that you’re not bringing something home.

Corey: So, I have to ask. You have been quoted as saying that 2022—for those listening, that is this year—is the year of documentation. You’re onboarding two new people into a company that does not see significant turnover, which means that invariably, “Oh, it’s been a while since we’ve updated the documentation. Whoops-a-doozy,” is a pretty common experience there. How much of your assertion that this is the year of documentation comes down to the, “Huh. Our onboarding stuff is really out of date,” versus a larger thing that you’re seeing in the industry?

Kelly: That is a great question because you never know what your documentation is like until you have someone new, kind of, come in with fresh eyes, has a perspective not only on, “Okay, I have no idea what this means,” or, “This is not where I thought it would be,” or, “This, you know, system is not working in any… in any way similar to anything I have ever seen in any other part of my, like, kind of, working career.” So, that’s where you really see what kind of gaps you have, but then you also kind of get to see which parts are working out really well. And not to spend, kind of, too much on that, but one of the best things that my coworkers did for me when I started was, Rachel Stephens had kept a log of, like, all the questions that she had as a new analyst. And she just, like, gave that to me with some advice on different things, like, in a spreadsheet, which I think is—I love spreadsheets so much and so does Rachel. And I think I might love spreadsheets more than Rachel at this point, even though she actually has a hat that says, “Spreadsheets.”

But when Kate started, it was fascinating to go through that and see what parts of that were either no longer relevant because the entire world had changed, or because the industry had advanced, or because there’s all these new things you need to know now that we're not on the list of things that you needed to know three years ago. And then what other, even, topics belong down on that kind of list of things to know. So, I think documentation is always a good, like, check-in for things like that.

But going back to, like, your larger question. So, documentation is important, not just because we happened to be onboarding, but a lot of people, I think once they no longer could be in the office with people and rely on that kind of face-to-face conversations to smooth over things began, I think, to realize how essential documentation was to just their everyday to day, kind of, working lives. So, I think that’s something that we’ve definitely seen from the pandemic. But then there are certainly other signals in the software industry-specific, which we can go into or not depending on your level of interest.

Corey: Well, something that I see that I have never been a huge fan of in corporate life—and it feels like it is very much a broad spectrum—has been that on one side of the coin, you have this idea that everything we do is bespoke and we just hire smart people and get out of their way. Yeah, that’s more uncontrolled anarchy than it is a repeatable company process around anything. And the other extreme is this tendency that companies have, particularly the large, somewhat slow-moving companies, to attempt to codify absolutely everything. It almost feels like it derives from the what I believe to be mistaken belief that with enough process, eventually you can arrive at the promised land where you don’t have to have intelligent, dynamic people working behind things, you can basically distill it down to follow the script and push the buttons in the proper order, and any conceivable outcome is going to be achieved. I don’t know if that’s accurate, but that’s always how it felt when you start getting too deeply mired in documentation-slash-process as almost religion.

Kelly: And I think—you know, I agree. There has to be something between, “All right, we don’t document anything and it’s not necessary and we don’t need it.” And then—

Corey: “We might get raided by the FBI. We want nothing written down.” At which point it’s like, what do you do here? Yeah.

Kelly: Yeah. Leave no evidence, leave no paper trail of anything like that. And going too far into thinking that processes is absolutely everything, and that absolutely anyone can be plugged into any given role and things will be equally successful, or that we’ll just be automated away or become just these, kind of, automatons. And I think that balance, it’s important to think about that because while documentation is important, and you know, I will say 2022, I think we’re going to hear more and more about it, we see it more as an increasingly valuable thing in tech, you can’t solve everything with documentation. You can use it as the, kind of, duct tape and baling wire for some of the things that your company is doing, but throwing documentation at it is not going to fix things in the same way that throwing engineers at a problem is not going to fix it either. Or most problems. I mean, there are some that you can just throw engineers at.

Corey: Well, there’s a company wiki, also known as where documentation goes to die.

Kelly: It is. And those, like, internal wikis, as horrible as they can be in terms of that’s where knowledge goes to die as well, places that have nothing like that, it can be even more chaotic than places that are relying on the, kind of, company internal wiki.

Corey: So, delving into a bit of a different topic here, before you were in the QA universe, you were what distills down to an academic. And I know that sometimes that can be interpreted as a personal attack in some quarters; I assure you, despite my own eighth grade level of education, that is not how this is intended at all. Your undergraduate degree was in medieval history—or medieval studies and your PhD was in English. So, a couple of questions around that. One, when we talk about medieval studies, are we talking about writing analyst reports about Netscape Navigator, or are we talking things a bit later in the sweep of history than that?

Kelly: I appreciate the Netscape Navigator reference. I get that reference.

Corey: Well, yeah. Medieval studies; you have to.

Kelly: Medieval studies, when you—where we study the internet in the 1990s, basically. I completely lost the line of questioning that you’re asking because I was just so taken by the Netscape Navigator reference.

Corey: Well, thank you. Started off with the medieval studies history. So, medieval studies of things dating back to, I guess, before we had reasonably recorded records in a consistent way. And also Twitter. But I’m wondering how much of that lends itself to what you do as an analyst.

Kelly: Quite a bit. And as much as I want to say, it’s all Monty Python references all the time, it isn’t. But the disciplinary rigor that you have to pick up as a medievalist or as anyone who’s getting any kind of PhD ever, you know, for the most part, that very much easily translated to being an analyst. And even more so tech culture is, in so many ways, like, enamored—there’s these pop culture medieval-isms that a lot of people who move in technical circles appreciate. And that kind of overlap for me was kind of fascinating.

So, when I started, like, working in tech, the fact that I was like writing a dissertation on Lord of the Rings was this little interesting thing that my coworkers could, like, kind of latch on to and talk about with me, that had nothing to do with tech and that had nothing to do with
the seemingly scary parts of being an academic.

Corey: This episode is sponsored in part by our friends at Vultr. Spelled V-U-L-T-R because they’re all about helping save money,
including on things like, you know, vowels. So, what they do is they are a cloud provider that provides surprisingly high performance cloud compute at a price that—while sure they claim its better than AWS pricing—and when they say that they mean it is less money. Sure, I don’t dispute that but what I find interesting is that it’s predictable. They tell you in advance on a monthly basis what it’s going to going to cost. They have a bunch of advanced networking features. They have nineteen global locations and scale things elastically. Not to be confused with openly, because apparently elastic and open can mean the same thing sometimes. They have had over a million users. Deployments take less that sixty seconds across twelve pre-selected operating systems. Or, if you’re one of those nutters like me, you can bring your own ISO and install basically any operating system you want. Starting with pricing as low as $2.50 a month for Vultr cloud compute they have plans for developers and businesses of all sizes, except maybe Amazon, who stubbornly insists on having something to scale all on their own. Try Vultr today for free by visiting: vultr.com/screaming, and you’ll receive a $100 in credit. Thats V-U-L-T-R.com slash screaming.

Corey: I want to talk a little bit about the idea of academic rigor because to my understanding, in the academic world, the publication process is… I don’t want to say it’s arduous. But if people subjected my blog post anything approaching this, I would never write another one as long as I lived. How does that differ? Because a lot of what I write is off-the-cuff stuff—and I’m not just including tweets, but also tweets—whereas academic literature winds up in peer-reviewed journals and effectively expands the boundaries of our collective societal knowledge as we know it. And it does deserve a different level of scrutiny, let’s be clear. But how do you find that shifts given that you are writing full-on industry analyst reports, which is something that we almost never do on our side, just honestly, due to my own peccadilloes?

Kelly: You should write some industry reports. They’re so fun. They’re very fun.

Corey: I am so bad at writing the long-form stuff. And we’ve done one or two previously, and each time my business partner had to basically hold my nose to the grindstone by force to get me to ship, on some level.

Kelly: And also, I feel like you might be underselling the amount of writing talent it takes to tweet.

Corey: It depends. You can get a lot more trouble tweeting than you can in academia most of the time. Every Twitter person is Reviewer 2. It becomes this whole great thing of, “Well, did you consider this edge corner case nuance?” It’s, “I’ve got to say, in 208 any characters, not really. Kind of ran out of space.”

Kelly: Yeah, there’s no space at all. And it’s not what that was intended. But going back to your original question about, like, you know, academic publishing and that type of process, I don’t miss it. And I have actually published some academic pieces since I became an analyst. So, my book finally came out after I had started as—it came out the end of 2019 and I had already been at RedMonk for a year.

It’s an academic book; it has nothing to do with being an industry analyst. And I had an essay come out in another collection around the same time. So, I’ve had that come out, but the thing is, the cycle for that started about a year earlier. So, the timeframe for getting things out in, especially the humanities, can be very arduous and frustrating because you’re kind of like, “I wrote this thing. I want it to actually appear somewhere that people can read it or use it or rip it apart if that’s what they’re going to do.”

And then the jokes that you hear on Twitter about Reviewer 2 are often real. A lot of academic publishing is done in, like, usually, like, a double-blind process where you don’t know who’s reviewing you and the reviewers don’t know who you are. I’ve been a reviewer, too, so I’ve been on that side of it. And—

Corey: Which why you run into the common trope of people—

Kelly: Yes.

Corey: —suggesting, “Oh, you don’t know what you’re talking about. You should read this work by someone else,” who is in fact, the author they are reviewing.

Kelly: Absolutely. That I think happens even when people do know who [laugh] who’s stuff they’re reviewing. Because it happens on Twitter all the time.

Corey: Like, “Well, have you gotten to the next step beyond where you have a reviewer saying you should wind up looking at the work cited by”—and then they name-check themselves? Have we reached that level of petty yet, or has that still yet to be explored?

Kelly: That is definitely something that happens in academic publishing. In academic circles, there can be these, like, frenemy relations among people that you know, especially if you are in a subfield that is very tiny. You tend to know everybody who is in that subfield, and there’s, like, a lot of infighting. And it does not feel that far from tech, sometimes. [unintelligible 00:21:52] you could look at the whole tech industry, and you look at the little areas that people specialize in, and there are these communities around these specializations that—you can see some of them on Twitter.

Clearly, not all of them exist in the Twitterverse, but in some ways, I think that translated over nicely of, like, the year-long publication and, like, double peer-review process is not something that I have to deal with as much now, and it’s certainly something that I don’t miss.

Corey: You spent extensive amounts of time studying the past, and presumably dragons as well because, you know, it’s impossible to separate medieval studies from dragons in my mind because basically, I am a giant child who lives through fantasy novels when it comes to exploring that kind of past. And do you wind up seeing any lessons we can take from the things you have studied to our current industry? That is sort of a strange question, but they say that history doesn’t repeat, but it rhymes, and I’m curious to how far back that goes. Because most people are citing, you know, 1980s business studies. This goes centuries before that.

Kelly: I think the thing that maybe stands out for me the most the way that you framed that is, when we look at the past and we think of something like the Middle Ages, we will often use that term and be like, “Okay, here’s this thing that actually existed, right?” Here’s, like, this 500 years of history, and this is where the Middle Ages began, and here’s where it ended, and this is what it was like, and this is what the people were like. And we look at that as the some type of self-evident thing that exists when in reality, it’s a concept that we created, that people who lived in later ages created this concept, but then it becomes something that has real currency and, really, weight in terms of, like, how we talk about the world.

So, someone will say, you know, I like that film. It was very medieval. And it’ll be a complete fantasy that has nothing to do with Middle Ages but has a whole bunch of these tropes and signals that we translate as the Middle Ages. I feel like the tech industry has a great capacity to do that as well, to kind of fold in along with things that we tend to think of as being very scientific and very logical but to take a concept and then just kind of begin to act as if it is an actual thing when it’s something that people are trying to make a thing.

Corey: Tech has a lot of challenges around the refusing to learn from history aspect in some areas, too. One of the most common examples I’ve heard of—or at least one that resonated the most with me—is hiring, where tech loves to say, “No one really knows how to hire effectively and well.” And that is provably not true. Ford and GM and Coca-Cola have run multi-decade studies on how to do this. They’ve gotten it down to a science.

But very often, we look at that in tech and we’re trying to invent everything from first principles. And I think, on some level, part of that comes out as, “Well, I wouldn’t do so well in that type of interview scenario, therefore, it sucks.” And I feel like we’re too willing in some cases to fail to heed the lessons that others have painstakingly learned, so we go ahead and experiment on our own and try and reinvent things that maybe we should not be innovating around if we’re small, scrappy, and trying to one area of the industry. Maybe going back to how we hire human beings should not be one of those areas of innovation that you spend all your time on as a company.

Kelly: I think for some companies, I think it depends on how you’re hiring now. It’s like, if your hiring practices are horrible, like, you probably do need to change them. But to your point, like, spending all of your energy on how are we hiring, can be counterproductive. Am I allowed to ask you a question?

Corey: Oh, by all means. Mostly, the questions people ask me is, “What the hell is wrong with you?” But that’s fine, I’m used to that one, too. Bonus points if you have a different one.

Kelly: Like, your hiring processes at Duckbill Group. Because you’ve hired, you know, folks recently. How do you describe that? Like, what points of that you think… are working really well?

Corey: The things that have worked out well for us have been being very transparent at the beginning around things like comp, what the job looks like, where it starts, where it stops, what we expect from people, what we do not expect from people, so there are no surprises down that path. We explain how many rounds of interviews there are, who they’ll be meeting with at each stage. If we wind up declining to continue with a candidate in a particular cycle, anything past the initial blind resume submission, we will tell them; we don’t ghost people. Full stop. Originally, we wanted to wind up responding to every applicant with a, “Sorry, we’re not going to proceed,” if the resume was a colossal mismatch. For example, we’re hiring for a cloud economist, and we have people with PhDs in economics, and… that’s it. They have not read the job description.

And then when you started doing that people would argue with us on a constant basis, and it just became a soul-sucking time sink. So, it’s unfortunate, but that’s the reality of it. But once we’ve had a conversation with you, doing that is the right answer. We try and move relatively quickly. We’re honest with folks because we believe that an interview is very much a two-way street.

And even if we declined to proceed—or you declined to proceed with us; either way—that you should still think well enough of us that you would recommend us to people for whom it might be a fit. And if we treat you like crap, you’re never going to do that. Not to mention, I just don’t like making people feel like crap as a general rule. So, that stuff that has all come out of hiring studies.

So, has the idea of a standardized interview. We don’t have an arbitrary question list that we wind up smacking people with from a variety of different angles. And if you drew the lucky questions, you’ll do fine. We also don’t set this up as pass-fail, we tend to presume that by the time you’ve been around the industry for as long as generally is expected for years of experience for the role, we’re not going to suddenly unmask you as not knowing how computers work through our ridiculous series of trivia questions. We don’t ask those.

We also make the interview look a lot like what the job is, which is apparently a weird thing. It’s in a lot of tech companies it’s, “Go and solve whiteboard algorithms for us.” And then, “Great. Now, what’s the job?” “It’s going to be moving around some CSS nonsense.”

It’s like, first that is very different, and secondly, it’s way harder to move CSS than to implement quicksort, for most folks. At least for me. So, it’s… yeah, it just doesn’t measure the right things. That’s our approach. I’m not saying we cracked it by any means to be very clear here. This is just what we have found that sucks the least.

Kelly: Yeah, I think the, ‘we’re not going to do obscure whiteboarding exercises’ is probably one of the key things. I think some people are still very attached those personal reasons. And I think the other thing I liked about what you said, is to make the interview as similar to the job as you can, which based on my own getting hired process at RedMonk and then to some levels of being involved in hiring our, kind of, new hires, I really like that. And I think that for me, the process will like, okay, you submit your application. There’d be—I think I’d to do a writing sample.

But then it was like, you get on a call and you talk to Steve. And then you get on a call and you talk to James. And talking to people is my job. Like for the most part. I write things, but it’s mostly talking to people, which you may not believe by the level of articulate, articulate-ness, I am stumbling my way through in this sentence.

And then the transparency angle, I think it’s something that most companies are not—may not be able to approach hiring in such a transparent way for whatever reason, but at least the motion towards being transparent about things like salaries, as opposed to that horrible salary negotiation part where that can be a nightmare for people, especially if there’s this code of silence around what your coworkers or potential coworkers are making.

Corey: We learned we were underpaying our clouds economists, so we wound up adjusting the rate advertised; at the same time we wound up improving the comp for existing team because, “Yeah, we’re just going to make you apply again to be paid a fair wage for what you do,” no. Not how we play these games.

Kelly: Yeah, which is, you know, one of the things that we’re seeing in the industry now. Of course, the term ‘The Great Resignation’ is out there. But with that comes, you know, people going to new places partly because that’s how they can get, like, the salary increase or whatever it is they want for among other reasons.

Corey: Some of the employees who have left have been our staunchest advocates, both for new applicants as well as new clients. There’s something to be said for treating people as you mean to go on. My business partner, I’ve been clear that we aspire for this to be a 20, 25-year company, and you don’t do that by burning bridges.

Kelly: Yeah. Or just assuming that your folks are going to stay for three years and move on, which tends to be the kind of the lifespan of where people stay.

Corey: Well, if they do, that’s fine because it is expected. I don’t want people to wind up feeling that they owe us anything. If it no longer makes sense for them to be here because they’re not fulfilled or whatnot—this has happened to us before we’ve tried to change their mind, talked to them about what they wanted, and okay, we can’t offer what you’re after. How can we help you move on? That’s the way it works.

And like, the one thing we don’t do in interviews—and this is something I very much picked up from the RedMonk culture as well—is we do a lot of writing here, so there’s a writing sample of here’s a list of theoretical findings for an AWS bill—if we’re talking about a cloud economist role—great. Now, the next round is people are going to talk to you about that, and we’re going to roleplay as if we were a client. But let’s be clear, I won’t tolerate abusive behavior from clients to our team, I will fire a client if it happens. So, we’re not going to wind up bullying the applicant and smacking ‘em around on stuff—or smacking them around to be clear. That was an ‘em not a him, let’s be clear.

It’s a problem of not wanting to even set the baseline expectation that you just have to sit there and take it when clients decide to go down unfortunate paths. And I believe it’s happened all of maybe once in our five-and-a-half-year history. So, why would you ever sit around and basically have a bunch of people chip away at an applicant’s self-confidence? By virtue of being in the room and having the conversation, they are clearly baseline competent at a number of things. Now, it’s just a question of fit and whether their expression of skills is what we’re doing right now as a company.

At least that’s how I see it. And I think that there is a lot of alignment here, not just between our two companies, but between the kinds of companies I look at and can actively recommend that people go and talk to.

Kelly: Yeah. I think that emphasis on, it’s not just about what a company is doing—like, what is their business, you know, how they’re making money—but how they’re treating people, like, on their way in and on the way out. I don’t think you can oversell how important that is.

Corey: Culture is what you wind up with instead of what you intend. And I think that’s something that winds up getting lost a fair bit.

Kelly: Yeah, culture is definitely not something you can just go buy, right? [laugh], where you can, like—this is what our culture will be.

Corey: No, no. But if there is, “Culture-in-a-box. Like, you may not be able to buy it, but I would love to sell it to you,” seems to be the watchwords of a number of different companies out there. Kelly, I really want to thank you for taking the time to speak with me today. If people want to learn more, where can they find you?

Kelly: They can find me on Twitter at @drkellyannfitz, that’s D-R-K-E-L-L-Y-A-N-N-F-I-T-Z—I apologize for having such a long Twitter handle—or my RedMonk work and of my colleagues, you can find that at redmonk.com.

Corey: And we will, of course, include links to that in the [show notes 00:33:14]. Thank you so much for your time. I appreciate it.

Kelly: Thanks for having me.

Corey: Kelly Fitzpatrick, senior industry analyst at RedMonk. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry comment telling me how terrible this was and that we should go listen to Reviewer 2’s podcast instead.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Anna

Anna has nearly ten years of experience researching and advising organizations on cloud adoption with a focus on security best practices. As a Gartner Analyst, Anna spent six years helping more than 500 enterprises with vulnerability management, security monitoring, and DevSecOps initiatives. Anna's research and talks have been used to transform organizations' IT strategies and her research agenda helped to shape markets. Anna is the Director of Thought Leadership at Sysdig, using her deep understanding of the security industry to help IT professionals succeed in their cloud-native journey.

Anna holds a PhD in Materials Engineering from the University of Michigan, where she developed computational methods to study solar cells and rechargeable batteries.

How do I adapt my security practices for the cloud-native world?

How do I select and deploy appropriate tools and processes to address business needs?

How do I make sense of new technology trends like threat deception, machine learning, and containers?

Links:

  • Sysdig: https://sysdig.com/
  • “2022 Cloud-Native Security and Usage Report”: https://sysdig.com/2022-cloud-native-security-and-usage-report/
  • Twitter: https://twitter.com/aabelak
  • LinkedIn: https://www.linkedin.com/in/aabelak/
  • Email: anna.belak@sysdig.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Today’s episode is brought to you in part by our friends at MinIO the high-performance Kubernetes native object store that’s built for the multi-cloud, creating a consistent data storage layer for your public cloud instances, your private cloud instances, and even your edge instances, depending upon what the heck you’re defining those as, which depends probably on where you work. It’s getting that unified is one of the greatest challenges facing developers and architects today. It requires S3 compatibility, enterprise-grade security and resiliency, the speed to run any workload, and the footprint to run anywhere, and that’s exactly what MinIO offers. With superb read speeds in excess of 360 gigs and 100 megabyte binary that doesn’t eat all the data you’ve gotten on the system, it’s exactly what you’ve been looking for. Check it out today at min.io/download, and see for yourself. That’s min.io/download, and be sure to tell them that I sent you.

Corey: This episode is sponsored by our friends at Oracle HeatWave is a new high-performance query accelerator for the Oracle MySQL Database Service, although I insist on calling it “my squirrel.” While MySQL has long been the worlds most popular open source database, shifting from transacting to analytics required way too much overhead and, ya know, work. With HeatWave you can run your OLAP and OLTP—don’t ask me to pronounce those acronyms again—workloads directly from your MySQL database and eliminate the time-consuming data movement and integration work, while also performing 1100X faster than Amazon Aurora and 2.5X faster than Amazon Redshift, at a third of the cost. My thanks again to Oracle Cloud for sponsoring this ridiculous nonsense.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Once upon a time, I went to a conference talk at, basically, a user meetup. This was in the before times, when that wasn’t quite as much of a deadly risk because of a pandemic, and mostly a deadly risk due to me shooting my mouth off when it wasn’t particularly appreciated.

At that talk, I wound up seeing a new open-source project that was presented to me, and it was called Sysdig. I wasn’t quite sure on what it did at the time and I didn’t know what it would be turning into, but here we are now, what is it, five years later. Well, it’s turned into something rather interesting. This is a promoted episode brought to us by our friends at Sysdig and my guest today is their Director of Thought Leadership, Anna Belak. Anna, thank you for joining me.

Anna: Hi, Corey. I’m very happy to be here. I’m a big fan.

Corey: Oh, dear. So, let’s start at the beginning. Well, we’ll start with the title: Director of Thought Leadership. That is a lofty title, it sounds like you sit on the council of the Lords of Thought somewhere. Where does your job start and stop?

Anna: I command the Council of the Lords of thought, actually. [laugh].

Corey: Supply chain issues mean the robe wasn’t available. I get it, I get it.

Anna: There is a robe. I’m just not wearing it right now. So, the shortest way to describe the role is probably something that reports into engineering, interestingly, and it deals with product and marketing in a way that is half evangelism and half product strategy. I just didn’t feel like being called any of those other things, so they were like, “Director of Thought Leadership you are.” And I was like, “That sounds awesome.”

Corey: You know, it’s one of those titles that people generally don’t see a whole lot of, so if nothing else, I always liked those job titles that cause people to sit up and take notice as opposed to something that just people fall asleep by the time you get halfway through it because, in lieu of a promotion, people give you additional adjectives in your title. And we’re going to go with it. So, before you wound up at Sysdig, you were at Gartner for a number of years.

Anna: That’s right, I spent about six years at Gartner, and there half the time I covered containers, Kubernetes, and DevOps from an infrastructure perspective, and half the time I spent covering security operations, actually, not specifically with respect to containers, or cloud, but broadly. And so my favorite thing is security operations, as it relates to containers and cloud-native workloads, which is kind of how I ended up here.

Corey: I wouldn’t call that my favorite thing. It’s certainly something that is near and dear to the top of mind, but that’s not because I like it, let’s put it [laugh] that way. It’s one of those areas where getting it wrong is catastrophic. Back in 2017, when I went to that meetup in San Francisco, Sysdig seemed really interesting to me because it looked like it tied together a whole bunch of different diagnostic tools, LSOF, strace, and the rest. Honestly—and I mean no slight to the folks who built out this particular tool—it felt like DTrace, only it understood the value of being accessible to its users without basically getting a doctorate in something.

I like the idea, and it felt like it was very much aimed at an in-depth performance analysis story or an observability play. But today, it seems that you folks have instead gone in much more of a direction of DevSecOps, if the people listening to this, and you, will pardon the term. How did that happen? What was that product evolution like?

Anna: Yeah, I think that’s a fair assessment, actually. And again, no disrespect to DTrace of which I’m also a fan. So, we certainly started out in the container observability space, essentially because this whole Docker Kubernetes thing was exploding in popularity—I mean, before it was exploding, it was just kind of like, peaking out—and very quickly, our founder Loris, who is the co-founder of Wireshark, was like, “Hey, there’s a visibility issue here. We can’t see inside these things with the tools that we have that are built for host instrumentation, so I’m going to make a thing.” And he made a thing, and it was an awesome thing that was open-sourced.

And then ultimately, what happened is, the ecosystem of containers and communities evolved, and more and more people started to adopt it. And so more people needed kind of a more, let’s say, hefty, serious tool for observability, and then what followed was another tool for security because what we actually discovered was the data that we’re able to collect from the system with Sysdig is incredibly useful for noticing security problems. So, that caused us to kind of expand into that space. And today we are very much a tool that still has an observability component that is quite popular, has a security component which is it’s fairly broad: We cover CSPM use cases, we cover [CIEM 00:05:04] use cases, and we are very, kind of let’s say, very strong and very serious about our detection response and runtime security use cases, which come from that pedigree of the original Sysdig as well.

Corey: You can get a fairly accurate picture of what the future of technology looks like by taking a look at what my opinion of something is, and then doing the exact opposite of that. I was a big believer that virtualization, “Complete flash in the pan; who’s going to use that?” Public cloud, “Are you out of your tree? No one’s going to trust other companies with their holy of holies.” And I also spent a lot of time crapping on containers and not actually getting into them.

Instead, I leapfrogged over into the serverless land, which I was a big fan of, which of course means that it’s going to be doomed sooner or later. My security position has also somewhat followed similar tracks where, back when you’re running virtual machines that tend to be persistent, you really have to care about security because you are running full-on systems that are persistent, and they run all kinds of different services simultaneously. Looking at Lambda Functions, for example, in the modern serverless world, I always find a lot of the tooling and services and offerings around security for that are a little overblown. They have a defined narrow input, they have a defined output, there usually aren’t omnibus functions shoved in here where they have all kinds of different code paths. And it just doesn’t have the same attack surface, so it often feels like it’s trying to sell me something I don’t need. Security in the container world is one of those areas I never had to deal with in anger, as a direct result. So, I have to ask, how bad is it?

Anna: Well, I have some data to share with you, but I’ll start by saying that I maybe was the opposite of you, so we’ll see which one of us wins this one. I was an instant container fangirl from the minute I discovered them. But I crapped out—

Corey: The industry shows you were right on that one. I think the jury [laugh] is pretty much in on this one.

Anna: Oh, I will take it. But I did crap on Lambda Functions pretty hard. I was like, “Serverless? This is dumb. Like, how are we ever going to make that work?” So, it seems to be catching on a little bit, at least it. It does seem like serverless is playing the function of, like, the glue between bits, so that does actually make a lot of sense. In retrospect, I don’t know that we’re going to have—

Corey: Well, it feels like it started off with a whole bunch of constraints around it, and over time, they’ve continued to relax those constraints. It used to be, “How do I package this?” It’s, “Oh, simple. You just spent four days learning about all the ins and outs of this,” and now it’s, “Oh, yeah. You just give it a Docker file?” “Oh. Well, that seems easier. I could have just been stubborn and waited.” Hindsight.

Anna: Yeah, exactly. So, containers as they are today, I think are definitely much more usable than they were five-plus years ago. There are—again there’s a lot of commercial support around these things, right? So, if you’re, you know, like, a big enterprise client, then you don’t really have time to fool around in open-source, you can go in, buy yourself a thing, and they’ll come with support, and somebody will hold your hand as you figure it out, and it’s actually quite, quite pleasant. Whether or not that has really gone mainstream or whether or not we’ve built out the entire operational ecosystem around it in a, let’s say, safe and functional way remains to be seen. So, I’ll share some data from our report, which is actually kind of the key thing I want to talk about.

Corey: Yeah, I wanted to get into that. You wound up publishing this somewhat recently, and I regret that as of the time of this recording, I have not yet had time to go into it in-depth, and of course eviscerate it in my typical style on Twitter—although that may have been rectified by the time that this show airs, to be very clear—but it’s the Sysdig “2022 Cloud-Native Security and Usage Report”.

Anna: Please at me when you Twitter-shred it. [laugh].

Corey: Oh, when I read through and screenshot it, and I’d make what observations that I imagine are witty. But I’m looking forward to it; I’ve done that periodically with the Flexera, “State of the Cloud” report for last few years, and every once in a while, whatever there’s a, “We’ve done a piece of thought leadership, and written a report,” it’s, “Oh, great. Let’s make fun of it.” That’s basically my default position on things. I am not a popular man, as you might imagine. But not having had the chance to go through it in-depth, what did this attempt to figure out when the study was built, and what did you learn that you found surprising?

Anna: Yeah, so the first thing I want to point out because it’s actually quite important is that this report is not a survey. This is actual data from our actual back end. So, we’re a SaaS provider, we collect data for our customers, we completely anonymize it, and then we show in aggregate what in fact we see them doing or not doing. Because we think this is a pretty good indicator of what’s actually happening versus asking people for their opinion, which is, you know, their opinion.

Corey: Oh, I love that. My favorite lies that people tell are the lies they don’t realize that they’re telling. It’s, I’ll do an AWS bill analysis and, “Great. So, tell me about all these instances you have running over in Frankfurt.” “Oh, we don’t have anything there.”

I believe you’re being sincere when you say this, however, the data does show otherwise, and yay, now we’re in a security incident.

Anna: Exactly.

Corey: I’m a big believer of going to the actual source for things like this where it’s possible.

Anna: Exactly. So, I’ll tell you my biggest takeaway from the whole thing probably was that I was surprised by the lack of… surprise. And I work in cloud-native security, so I’m kind of hoping every single day that people will start adopting these modern patterns of, like, discarding images, and deploying new ones when they found a vulnerability, and making ephemeral systems that don’t run for a long time like a virtual machine in disguise, and so on. And it appears that that’s just not really happening.

Corey: Yeah, it’s always been fun, more than a little entertaining, when I wind up taking a look at the aspirational plans that companies have. “Great, so when are you going to do”—“Oh, we’re going to get to that after the next sprint.” “Cool.” And then I just set a reminder and I go back a year later, and, “How’s that coming?” “Oh, yeah. We’re going to get to that next sprint.”

It’s the big lie that we always tell ourselves that right after we finished this current project, then we’re going to suddenly start doing smart things, making the right decisions, and the rest. Security, cost, and a few other things all tend to fall on the side of, you can spend infinite money and infinite time on these things, but it doesn’t advance what your business is doing, but if you do none of those things, you don’t really have a business anymore. So, it’s always a challenge to get it prioritized by the strategic folks.

Anna: Exactly. You’re exactly right because what people ultimately do is they prioritize business needs, right? They are prioritizing whatever makes them money or creates the trinkets their selling faster or whatever it is, right? The interesting thing, though, is if you think about who our customers would be, like, who the people in this dataset are, they are all companies who are probably more or less born in the cloud or at least have some arm that is born in the cloud, and they are building software, right? So, they’re not really just your average enterprises you might see in a Gartner client base which is more broad; they are software companies.

And for software companies, delivering software faster is the most important thing, right, and then delivering secure software faster, should be the most important thing, but it’s kind of like the other thing that we talk about and don’t do. And that’s actually what we found. We found that people do deliver software faster because of containers and cloud, but they don’t necessarily deliver secure software faster because as is one of our data points, 75% of containers that run in production have critical or high vulnerabilities that have a patch available. So, they could have been fixed but they weren’t fixed. And people ask why, right? And why, well because it’s hard; because it takes time; because something else took priority; because I’ve accepted the risk. You know, lots of reasons why.

Corey: One of the big challenges, I think, is that I can walk up and down the expo hall at the RSA Conference, which until somewhat recently, you were not allowed to present that or exhibit at unless you had the word ‘firewall’ in your talk title, or wound up having certain amounts of FUD splattered across your banners at the show floor. It feels like there are 12 products—give or take—for sale there, but there are hundreds of booths because those products have different names, different messaging, and the rest, but it all feels like it distills down to basically the same general categories. And I can buy all of those things. And it costs an enormous pile of money, and at the end of it, it doesn’t actually move the needle on what my business is doing. At least not in a positive direction, you know? We just set a giant pile of money on fire to make sure that we’re secure.

Well, great. Security is never an absolute, and on top of that, there’s always the question of what are we trying to achieve as a business. As a goal—from a strategic perspective—security often looks a lot like, “Please let’s not have a data breach that we have to report to people.” And ideally, if we have a lapse, we find out about it through a vector that is other than the front page of The New York Times. That feels like it’s a challenging thing to get prioritized in a lot of these companies. And you have found in your report that there are significant challenges, of course, but also that some companies in some workloads are in fact getting it right.

Anna: Right, exactly. So, I’m very much in line with your thinking about this RSA shopping spree, and the reality of that situation is that even if we were to assume that all of the products you bought at the RSA shopping center were the best of breed, the most amazing, fantastic, perfect in every way, you would still have to somehow build a program on top of them. You have to have a process, you have to have people who are bought into that process, who are skilled enough to execute on that process, and who are more or less in agreement with the people next door to them who are stuck using one of the 12 trinkets you bought, but not the one that you’re using. So, I think that struggle persists into the cloud and may actually be worse in the cloud because now, not only are we having to create a processor on all these tools so that we can actually do something useful with them, but the platform in which we’re operating is fundamentally different than what a lot of us learned on, right?

So, the priorities in cloud are different; the way that infrastructure is built is a little different, like, you have to program a YAML file to make yourself an instance, and that’s kind of not how we are used to doing it necessarily, right? So, there are lots of challenges in terms of skills gap, and then there’s just this eternal challenge of, like, how do we put the right steps into place so that everybody who’s involved doesn’t have to suffer, right, and that the thing that comes out at the end is not garbage. So, our approach to it is to try to give people all the pieces they need within a certain scope, so again, we’re talking about people developing software in a cloud-native world, we’re focused kind of on containers and cloud workloads even though it’s not necessarily containers. So that’s, like, our sandbox, right? But whoever you are, right, the idea is that you need to look to the left—because we say ‘shift left’—but then you kind of have to follow that thread all the way to the right.

And I actually think that the thing that people most often neglect is the thing on the right, right? They maybe check for compliance, you know, they check configurations, they check for vulnerabilities, they check, blah, blah, blah, all this checking and testing. They release their beautiful baby into the world, and they’re like, okay, I wash my hands of it. It’s fine. [laugh]. Right but—

Corey: It has successfully been hurled over the fence. It is the best kind of problem, now: Someone else’s.

Anna: It’s gone. Yeah. But it’s someone else’s—the attacker community, right, who are now, like, “Oh, delicious. A new target.” And like, that’s the point at which the fun starts for a lot of those folks who are on the offensive side. So, if you don’t have any way to manage that thing’s security as it’s running, you’re kind of like missing the most important piece, right? [laugh].

Corey: One of the challenges that I tend to see with a lot of programmatic analysis of this is that it doesn’t necessarily take into account any of the context because it can’t. If I have, for example, a containerized workload that’s entire job is to take an image from S3, run some analysis or transformation on it then output the results of that to some data store, and that’s all it’s allowed to talk to you, it can’t ever talk to the internet, having a system that starts shrieking about, “Ah, there’s a vulnerability in one of the libraries that was used to build that container; fix it, fix it, fix it,” doesn’t feel like it’s necessarily something that adds significant value to what I do. I mean, I see this all the time with very purpose-built Lambda Functions that I have doing one thing and one thing only. “Ah, but one of the dependencies in the JSON processing library could turn into something horrifying.” “Yeah, except the only JSON it’s dealing with is what DynamoDB returns. The only thing in there is what I’ve put in there.”

That is not a realistic vector of things for me to defend against. The challenge then becomes when everything is screaming that it’s an emergency when you know, due to context, that it’s not, people just start ignoring everything, including the, “Oh, and by the way, the building is on fire,” as one of—like, on page five, that’s just a small addendum there. How do you view that?

Anna: The noise insecurity problem, I think, is ancient and forever. So, it was always bad, right, but in cloud—at least some containers—you would think it should be less bad, right, because if we actually followed these sort of cloud-native philosophy, of creating very purp—actually it’s called the Unix philosophy from, like, I don’t know, before I was born—creating things that are fairly purposeful, like, they do one thing—like you’re saying—and then they disappear, then it’s much easier to know what they’re able to do, right, because they’re only able to do what we’ve told them, they’re able to do. So, if this thing is enabled to make one kind of network connection, like, I’m not really concerned about all the other network connections it could be making because it can’t, right? So, that should make it easier for us to understand what the attack surface actually is. Unfortunately, it’s fairly difficult to codify and productize the discovery of that, and the enrichment of the vulnerability information or the configuration information with that.

That is something we are definitely focusing on as a vendor. There are other folks in the industry that are also working on this kind of thing. But you’re exactly right, the prioritization of not just a vulnerability, but a vulnerability is a good example. Like, it’s a vulnerability, right? Maybe it’s a critical or maybe it’s not.

First of all, is it exposed to the outside world somehow? Like, can we actually talk to this system? Is it mitigated, right? Maybe there’s some other controls in place that is mitigating that vulnerability. So, if you look at all this context, at the end of the day, the question isn’t really, like, how many of these things can I ignore? The question is at the very least, which are the most important things that I actually can’t ignore? So, like you’re saying, like, the buildings on fire, I need to know, and if it’s just, like, a smoldering situation, maybe that’s not so bad. But I really need to know about the fire.

Corey: This episode is sponsored in part by LaunchDarkly. Take a look at what it takes to get your code into production. I’m going to just guess that it’s awful because it’s always awful. No one loves their deployment process. What if launching new features didn’t require you to do a full-on code and possibly infrastructure deploy? What if you could test on a small subset of users and then roll it back immediately if results aren’t what you expect? LaunchDarkly does exactly this. To learn more, visit launchdarkly.com and tell them Corey sent you, and watch for the wince.

Corey: It always becomes a challenge of prioritization, and that has been one of those things that I think, on some level, might almost cut against a tool that works at the level that Sysdig does. I mean, something that you found in your report, but I feel like, on some level, is one of those broadly known, or at least unconsciously understood things is, you can look into a lot of these tools that give incredibly insightful depth and explore all kinds of neat, far-future, bleeding edge, absolute front of the world, deep-dive security posture defenses, but then you have a bunch of open S3 buckets that have all of your company’s database backups living in them. It feels like there’s a lot of walk before you can run. And then that, on some level, leads to the wow, we can’t even secure our S3 buckets; what’s the point of doing anything beyond that? It’s easy to, on some level, almost despair, want to give up, for some folks that I’ve spoken to. Do you find that is a common thing or am I just talking to people who are just sad all the time?

Anna: I think a lot of security people are sad all the time. So, the despair is real, but I do think that we all end up in the same solution, right? The solution is defense in depth, the solution is layer control, so the reality is if you don’t bother with the basic security hygiene of keeping your buckets closed, and like not giving admin access to every random person and thing, right? If you don’t bother with those things, then, like, you’re right, you could have all the tools in the world and you could have the most advanced tools in the world, and you’re just kind of wasting your time and money.

But the flipside of that is, people will always make mistakes, right? So, even if you are, quote-unquote, “Doing everything right,” we’re all human, and things happen, and somebody will leave a bucket open on accident, or somebody will misconfigure some server somewhere, allowing it to make a connection it shouldn’t, right? And so if you actually have built out a full pipeline that covers you from end-to-end, both pre-deployment, and at runtime, and for vulnerabilities, and misconfigurations, and for all of these things, then you kind of have checks along the way so that this problem doesn’t make it too far. And if it does make it too far and somebody actually does try to exploit you, you will at least see that attack before they’ve ruined everything completely.

Corey: One thing I think Sysdig gets very right that I wish this was not worthy of commenting on, but of course, we live in the worst timeline, so of course it is, is that when I pull up the website, it does not market itself through the whole fear, uncertainty, and doubt nonsense. It doesn’t have the scary pictures of, “Do you know what’s happening in your environment right now?” Or the terrifying statistics that show that we’re all about to die and whatnot. Instead, it talks about the value that it offers its customers. For example, I believe its opening story is, “Run with confidence.” Like, great, you actually have some reassurance that it is not as bad as it could be. That is, on the one hand, a very uplifting message and two, super rare. Why is it that so much of the security industry resorts to just some of the absolute worst storytelling tactics in order to drive sales?

Anna: That is a huge compliment, Corey, and thank you. We try very hard to be kind of cool in our marketing.

Corey: It shows. I’m tired of the 1990s era story of, “Do you know where the hackers are?” And of course, someone’s wearing, like, a ski mask and typing with gloves on—which is always how I break into things; I don’t know about you—but all right, we have the scary clip art of the hacker person, and it just doesn’t go anywhere positive.

Anna: Yeah. I mean, I think there certainly was a trend for a while have this FUD approach. And it’s still prevalent in the industry, in some circles more than others. But at the end of the day, Cloud is hard and security is hard, and we don’t really want to add to the suffering; we would like to add to the solution, right? So, I don’t think people don’t know that security is hard and that hackers are out there.

And you know, there’s, like, ransomware on the news every single day. It’s not exactly difficult to tell that there’s a challenge there, so for us to have to go and, like, exacerbate this fear is almost condescending, I feel, which is kind of why we don’t. Like, we know people have problems, and they know that they need to solve them. I think the challenge really is just making sure that A) can folks know where to start and how to build a sane roadmap for themselves? Because there are many, many, many things to work on, right?

We were talking about context before, right? Like, so we actually try to gather this context and help people. You made a comment about how having a lot of telemetry might actually be a little bit counterproductive because, like, there’s too much data, what do I do well—

Corey: Here’s the 8000 findings we found that you fail—great. Yeah. Congratulations, you’re effectively the Nessus report as a company. Great. Here you go.

Anna: Everything is over.

Corey: Yeah.

Anna: Well, no shit, Nessus, you know. Nessus did its thing. All right. [laugh].

Corey: Oh, Nessus was fantastic. Nessus was—for those who are unaware, Nessus was an open-source scanner made by the folks at Tenable, and what was great about it was that you could run it against an environment, it would spit out all the things that it found. Now, one of the challenges, of course, is that you could white-label this and slap whatever logo you wanted on the top, and there were a lot of ‘security consultancies’ that use the term incredibly… lightly, that would just run a Nessus report, drop off the thick print out. “Here’s the 800 things you need to fix. Pay me.” And wander on off into the sunset.

And when you have 800 things you need to fix, you fix none of them. And they would just sit there and atrophy on the shelf. Not to say that all those things weren’t valid findings, but you know, the whole, you’re using an esoteric, slightly deprecated TLS algorithm on one of your back-end services, versus your Elasticsearch database does not have a password set. Like, there are different levels of concern here. And that is the problem.

Anna: Yeah. That is in fact one of the problems we’re aggressively trying to solve, right? So, because we see so much of the data, we’re actually able to piece together a lot of context to gives you a sense of risk, right? So, instead of showing all the data to the customer—the customer can see it if they want; like, it’s all in there, you can look at it—one of the things we’re really trying to do is collect enough information about the finding or the event or the vulnerability or whatever, so we can kind of tell you what to do.

For example, one you can do this is super basic, but if you’re looking at a specific vulnerability, like, let’s say it’s like Log4j or whatever, you type it in, and you can see all your systems affected by this thing, right? Then you can, in the same tool, like, click to the other tab, and you can see events associated with this vulnerability. So, if you can see the systems that the vulnerability is on and you can see there’s weird activity on those systems, right? So, if you’re trying to triage some weird thing in your environment, during the Log4j disaster, it’s very easy for you to be like, “Huh. Okay, these are the relevant systems. This is the vulnerability. Like, here’s all that I know about this stuff.”

So, we kind of try to simplify as much as possible—my design team uses the word ‘easify,’ which I love; it’s a great word—to easify, the experience of the end-user so that they can get to whatever it is they’re trying to do today. Like, what can I do today to make my company more secure as quickly as possible? So, that is sort of our goal. And all this huge wealth of information we gather, we try to package for the users in a way that is, in fact, digestible. And not just like, “Here’s a deluge of suffering,” like, “Look.” [laugh]. You know?

Corey: This is definitely complicated in the environment I tend to operate in which is almost purely AWS. How much more complex is get when people start looking into the multi-cloud story, or hybrid environments where they have data center is talking to things within AWS? Because then it’s not just the expanded footprint, but the entire security model works slightly differently in all of those different environments as well, and it feels like that is not a terrific strategy.

Anna: Yeah, this is tough. My feelings on multi-cloud are mostly negative, actually.

Corey: Oh, thank goodness. It’s not just me.

Anna: I was going to say that, like, multi-cloud is not a strategy; it’s just something that happens to you.

Corey: Same with hybrid. No one plans to do hybrid. They start doing a cloud migration, realize halfway through some things are really hard to move, give up, plant the flag, declare victory, and now it’s called hybrid.

Anna: Basically. But my position—and again, as an analyst, you kind of, I think, end up in this position, you just have a lot of sympathy for the poor people who are just trying to get these stupid systems to run. And so I kind of understand that, like, nothing’s ideal, and we’re just going to have to work with it. So multi-cloud, I think is one of those things where it’s not really ideal, we just have to work with it. There’s certainly advantages to it, like, there’s presumably some level of mythical redundancy or whatever. I don’t know.

But the reality is that if you’re trying to secure a pile of junk in Azure and a pile of junk in AWS, like, it’d be nice if you had, like, one tool that told you what to do with both piles of junk, and sometimes we do do that. And in fact, it’s very difficult to do that if you’re not a third-party tool because if you’re AWS, you don’t have much incentive to, like, tell people how to secure Azure, right? So, any tool in the category of, like, third-party CSPM—Gartner calls them CWPP—kind of, cloud security is attempting to span those clouds because they always have to be relevant, otherwise, like, what’s the point, right?

Corey: Well, I would argue cynically there’s also the VC model, where, “Oh, great. If we cover multiple cloud providers, that doubles or triples our potential addressable market.” And, okay, great, I don’t have those constraints, which is why I tend to focus on one cloud provider where I tend to see the problems I know how to solve as opposed to trying to conquer the world. I guess I have my bias on that one.

Anna: Fair. But there’s—I think the barrier to entry is lower as a security vendor, right? Especially if you’re doing things like CSPMs. Take an example. So, if you’re looking at compliance requirements, right, if your team understands, like, what it means to be compliant with PCI, you know, like, [line three 00:28:14] or whatever, you can apply that to Azure and Amazon fairly trivially, and be like, “Okay, well, here’s how I check in Azure, and here’s how I check in Amazon,” right?

So, it’s not very difficult to, I think, engineer that once you understand the basic premise of what you’re trying to accomplish. It does become complicated as you’re trying to deal with more and more different cloud services. Again, if you’re kind of trying to be a cloud security company, you almost have no choice. Like, you have to either say, “I’m only doing this for AWS,” which is kind of a weird thing to do because they’re kind of doing their own half-baked thing already, or I have to do this for everybody. And so most default to doing it for everybody.

Whether they do it equally well, for everybody, I don’t know. From our perspective, like, there’s clearly a roadmap, so we have done one of them first and then one of them second and one of the third, and so I guarantee you that we’re better in some than others. So, I think you’re going to have pluses and minuses no matter what you do, but ultimately what you’re looking for is coverage of the tool’s capabilities, and whether or not you have a program that is going to leverage that tool, right? And then you can check the boxes of like, “Okay. Does it do the AWS thing? Does it do this other AWS thing? Does it do this Azure thing?”

Corey: I really appreciate your taking the time out of your day to speak with me. We’re going to throw a link to the report itself in the [show notes 00:29:23], but other than that, if people want to learn more about how you view these things, where’s the best place to find you?

Anna: I am—rarely—but on Twitter at @aabelak. I am also on LinkedIn like everybody else, and in the worst case, you could find me by email, at anna.belak@sysdig.com.

Corey: And we will of course put links to that in the [show notes 00:29:44]. Thank you so much for taking the time to speak with me today. I appreciate it.

Anna: Thanks for having me, Corey. It’s been fun.

Corey: Anna Belak, Director of Thought Leadership at Sysdig. I’m Cloud Economist Corey Quinn and this is streaming on the cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an angry comment telling me not only why this entire approach to security is awful and doomed to fail, but also what booth number I can find you at this year’s RSA Conference.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Lynn

Cloud Architect who codes, Angel Investor

Links:

  • Lynn Langit Consulting: https://lynnlangit.com/
  • Groove Capital: https://www.groovecap.com/groove-capital-minnesotas-first-check-fund
  • Twitter: https://twitter.com/lynnlangit
  • GitHub: https://github.com/lynnlangit

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Today’s episode is brought to you in part by our friends at MinIO the high-performance Kubernetes native object store that’s built for the multi-cloud, creating a consistent data storage layer for your public cloud instances, your private cloud instances, and even your edge instances, depending upon what the heck you’re defining those as, which depends probably on where you work. It’s getting that unified is one of the greatest challenges facing developers and architects today. It requires S3 compatibility, enterprise-grade security and resiliency, the speed to run any workload, and the footprint to run anywhere, and that’s exactly what MinIO offers. With superb read speeds in excess of 360 gigs and 100 megabyte binary that doesn’t eat all the data you’ve gotten on the system, it’s exactly what you’ve been looking for. Check it out today at min.io/download, and see for yourself. That’s min.io/download, and be sure to tell them that I sent you.

Corey: This episode is sponsored by our friends at Oracle HeatWave is a new high-performance query accelerator for the Oracle MySQL Database Service, although I insist on calling it “my squirrel.” While MySQL has long been the worlds most popular open source database, shifting from transacting to analytics required way too much overhead and, ya know, work. With HeatWave you can run your OLAP and OLTP—don’t ask me to pronounce those acronyms again—workloads directly from your MySQL database and eliminate the time-consuming data movement and integration work, while also performing 1100X faster than Amazon Aurora and 2.5X faster than Amazon Redshift, at a third of the cost. My thanks again to Oracle Cloud for sponsoring this ridiculous nonsense.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. So, I’ve been doing this podcast for a little while now—by my understanding, this is episode 300 and something—but back when the very first episode aired, I had pre-recorded the first twelve episodes. Episode number ten was with Lynn Langit who is, among many other things, the CEO of Lynn Langit Consulting, she is also the first person to achieve the AWS Community Hero and equivalent designations at all three of the primary tier-one hyperscale cloud providers, which I can’t even wrap my head around what it takes to get that at one of those companies. Lynn, thank you so much for agreeing to come back now that I’m no longer scared of the microphone.

Lynn: Well, thank you for having me. It’s great to be back, Corey.

Corey: So, it’s been a few years now since we really sat down and caught up. And what an interesting few years it’s been. There’s been a whole minor global pandemic thing that wound up hitting us from unexpected and unpleasant places. There’s been a significant, I would say, not revolution but evolution in how adoption of cloud services has been proceeding. The types of problems that customers are encountering, the conversational discourse has moved significantly away from, “Should we be using cloud?” Into, “Okay, we obviously should be using Cloud. How should we be using it?” And the industry keeps on churning. Sure there’s still rough parts, there are still ridiculous aspects of it, but what have you been up to?

Lynn: Well, as you might remember, I have an independent consultancy where I do really what my customers need. I work across different clouds, which keeps it interesting and fun, but I’ve had a focus over the past few years in supporting bioinformatics research. Before the pandemic, it was mostly cancer research. Since the pandemic, it’s been all Covid, all the time.

Corey: All Covid, all the time sort of has been the unofficial theme of this. And it’s weird. I know, we’re in 2022, now, but it still feels like on some level, it’s like, “Man, this is March 2020; it’s still dragging on, on some level.” There have been a number of stories in the world that is, let’s say medicine-adjacent, more so than—we’re all sort of medicine adjacent these days, but there’s been a lot of refocusing away from things like cancer research into Covid and similar pandemic respiratory diseases. Do you think that there’s a longer-term story where we’re going to start seeing progress stall on things that were previously areas of focus—in your case cancer—in favor of reducing infectious disease, or is it really one of those ‘rising tide lifts all boats’ type of scenarios?

Lynn: Yeah, it’s the latter. It’s been really interesting. Without getting too much into the details, you know, you think of genomic research for drug discovery, you know, we started with this idea of different DNA sequencing cohorts. So, like people from the—you know, that started from the United States, people that started from Africa, you know, different cohort as a normative to evaluate the effectiveness of diseases, what was an area of research already was to go down to the level of what’s called single-cell RNA. So, look at the expression of the genomics by cell area, so by the different parts of your body.

Well, this is similar to what has been done to understand the impact and the efficacy of potential Covid drugs. So, this whole single-cell RNA mapping cohorts of what is normal for different types of populations has resulted in this data explosion that I’ve never seen before. And I see it as positive for the impact of human health. However, it really drives the need for adoption to the cloud. These research facilities are running out of space if they’re still working on-prem.

Corey: I spend an awful lot of time thinking about data and its storage from a primarily cost-focused perspective, for obvious reasons, and that is nuanced and intricate and requires, sort of, an end-to-end lifecycle policy. There’s this idea of, ideally, you would delete old data you don’t need anymore, but failing that you, maybe aspirationally, don’t need 500 copies of the same thing lying around. Maybe there are ways to fix that. And that’s all within one cloud ecosystem. You work across all of the clouds. How do you keep it all straight in your head trying to figure out things around lifecycles, things around just understanding the capabilities of the various platforms? Because I got to say, from my perspective, it’s challenging enough only bounding it to one.

Lynn: Yeah, it’s the constant problem. The big clients I had over this past year were not on Amazon, they were on other platforms. So, it seems like it sort of goes in cycles. And what I’ll sometimes need to do is hire subcontractors that have been working on those platforms because you can’t, I mean, you can’t even know one platform, much less all of them to the level of complexity in order to implement. One thing that is kind of interesting though, in bioinformatics is—and different than the other domains—is when you talk about data, it’s a function of time first and cost second.

So, they will run on less computational resources, so that they can, for example, not overspend their research grant, and wait longer for the results. And this has been really an interesting shift in my work because I used to work with FinTech and ad tech, where it’s all about, get it out there fast. And we don’t really care how much it costs, we just want it super fast. So, this continuum of time or money shifts by vertical. And that’s been something that—I don’t know, it’s kind of obvious, in hindsight, but I didn’t really expect until I got into the different domains.

Corey: It’s always been fascinating to me watching how different organizations and different organization types wind up have interacting with cost. I mean, I’ve been saying for a while now that cost and architecture are the same thing when it comes to cloud. What are your trade-offs? What are your constraints? In many venture-backed companies, it’s when you have a giant pile of other people’s money raring to go, and it’s a spend it and hit your milestone if you want to get another round of funding, or this has been an incredible journey Medium post in the making, then, yeah, okay, go ahead and make the result happen faster. Save money is not the first, second or third order of business as far as what you’re trying to achieve.

In academia, where everything’s grant powered. And it’s a question of, we need to be able to deliver, and we need to be able to show results and be able to go and play the game and understand the cultural context we’re operating in, and ideally get another grant next year, it completely shifts the balance of what needs to be prioritized and when. And I don’t think there’s been a lot of discussion around that because most cloud cost discussions inherently center around industry.

Lynn: They do and they focus on the industries where they’re willing to spend most. So, most of the reference examples are, they always prioritize for time and money is sort of unlimited. I’ll give you an example—this was from a few years back—some work I did with a research group in Australia, and again, it was a genomics example. They were running on-prem, and to do a single query, it took them 500 hours. And I was just like, “Are you kidding me?”

And they’re like, “Hey, cloud lady, what can you do?” Right? So, we gave two solutions, and the first solution was kind of a more of a lift-and-shift kind of a solution because they didn’t know anything about cloud. And it took a few hours. The second solution was what was in our opinion, super elegant, it was one of the earliest data lakes, it took minutes.

Well, it was a big hit to the ego that they adopted… the easier solution. But again, it’s a learning because another dimension about cloud architecture is usability. The FinTechs are like, “We’re going to get it really done fast; we’ll hire who we need to hire.” The biotechs, they can’t afford to hire who they need to hire because there all being hired by the FinTechs. So, you have these different dimensions you need to optimize for that aren’t really obvious if you just work in the industries that optimize for time.

Corey: And the thing that always gets overlooked is that in most environments, the people working on things are more expensive than the infrastructure themselves. And back when Lambda and all the serverless joy came out, my first iteration of lastweekinaws.com website was powered entirely by Lambda functions, S3, and other assorted bits of nonsense. Today, it’s on WordPress.

And it’s not because I think that is somehow the superior architecture from a purely technologist point of view, but because I have to find other people who aren’t me or one of the other six people in the world at the time who could stuff all that into their head and work on it effectively, should be able to make changes to the website. That is not something I need to be focusing on. There’s something to be said for going to where there’s a significant talent pool, rather than pushing the frontiers of innovation in areas that don’t directly benefit whatever it is your organization is targeting.

Lynn: Yeah, it’s really interesting, when Covid hit back in 2020—kind of an interesting little story here—one of my clients is the Broad Institute at MIT and Harvard—they’re a well-known research organization for, you know, cancer genomic datasets—they were tasked with pivoting their labs so that they could provide Covid testing capability. And I was a long-term contractor with them, so they brought me in for an architectural cloud consultant. I said, “This clearly is a serverless. I know you guys haven’t done this before, but this is going to be burstable, you don’t know how big this is going to need to go.” And then just to make life interesting, in the middle of the build of that, I was one of the first people in Minnesota to get Covid, so I actually wasn’t able to go and complete it, nor was I able to get a test because there weren’t tests.

I mean, you know, I can’t make this stuff up. I was in the ER saying, “Okay, is this the end of me, or can I go back and get you some tests?” [laugh]. So, it’s really kind of two things—kind of a weird story. And also, life situations will cause change, and so the Broad did launch that pipeline, and it was serving up to 10% of the Covid tests in the United States.

But they had never done anything serverlessly or had considered it before because they didn’t need to have that amount of change. It was really, again, a big thing when I came into human health. Prior to that, I was doing all serverless all the time. You know, I came into human health, and they were saying, “Okay, we’re going to have massive VMs.” And I was like, “No…” but you know, you have to meet the client where they are.

Corey: I think it’s the easiest thing in the world, particularly as a junior consultant—because you do not see senior consultants doing this ever, you know, after the first time—to walk into an environment, look around and have zero context into what’s going on—because you’re a consultant; you haven’t been there and say, “This is ridiculous. What fool built this?” Invariably, to said fool. Now, most people don’t show up in the morning hoping to do a terrible job at work today, so there are constraints that you are certainly not seeing. And maybe it was an offering wasn’t available that maybe they weren’t aware of it. Maybe there was a constraint that you’re not seeing.

But the best case is you’re right and you just made them feel terrible, which is not generally a great way to land more consulting projects. It’s always frustrating to me because even looking at a bill and having a pretty good idea of what’s going on, I always frame it as, “Can you help me understand why this is the case? Had you considered this, or is that not an option?” As opposed to categorically saying, well, this is not the way to do it. Because once you’re wrong when you’re delivering expertise, it takes a lot to build that back, if it’s even possible.

Lynn: Well, again, from human health because, you know, they were consuming the vendor information, they thought they wanted to learn how to use Kubernetes, but what they really needed to learn was how to do archiving to reduce their storage costs.

Corey: Yes. Kubernetes is a terrific solution for a bunch of problems and create several orders of magnitude more somewhere along the way. My somewhat accurate, somewhat snarky observation is that Kubernetes is great if your primary problem is you want to pretend you work at Google but didn’t pass their technical screen. I don’t really want to cosplay as a cloud provider myself, most days. That said, there are use cases for which it makes sense, but context is everything, and generally speaking, I don’t tend to follow a hype trend to figure out whether or not it’s going to solve my particular problem.

Lynn: Well, here’s the soundbite: “Kubernetes is today’s Hadoop.”

Corey: Oh, there are people who are not going to like that. I made a tweet, I think—

Lynn: Tough.

Corey: —three years ago now—

Lynn: It’s true. [laugh].

Corey: Oh, yeah. Tweet three years ago or so that said, “Hot take: In five years, nobody’s going to care about Kubernetes.” And I think I have a year or two left on that prediction. And what I said at the time was that not that it’s going to go away and not be anywhere—because enterprises do not move that quickly—but it’s no longer going to be the sort of thing that everyone is concerned about at a very high level. The Linux kernel has a bunch of aspects to it that we used to have to care about a fair bit. Now, a few people really, really need to care about those things; because of those folks’ hard work, the rest of us don’t have to think about it at all. And that is the nature of technology, in the fullness of time.

Lynn: Well, another way to think about it is Kubernetes is a C++. Certain people are going to be experts in it and need to, and that’s valid, right, but what percentage of developers code in C++. Like, ten? Five? You know, it’s kind of analogous, right?

So, it’s one of the signatures of my consultancy. You know, I’m this pragmatic midwesterner, and I love to say, “Look,”—like you said—“If you think you need this, you really need to understand the actual cost of it because it’s non-trivial on all clouds.” And I get to say that because I’m independent. You know, they’re doing solid work to abstract it into a higher-level implementation, but when I hear a customer say, “I need Kubernetes,” the burden of proof is on them [laugh] before I’m going to build that.

Corey: Speaking of hype-driven emerging technologies, you are arguably one of the few people on the planet I can have this conversation with, and I do not mean that as an insult other people operating in this space. For context, a couple of years ago, AWS launched Brakets—which they spelled Braket without a C because it’s Amazon and spelling is hard, presumably; I know, I know, there’s a reason behind it—and it is their service that enables you to get access to quantum computers the same way we get access to any other AWS service: Through a somewhat janky console and some APIs. And, okay, quantum computing. We’ve heard a lot about it forever; it always seemed a bit like science fiction and it was never really clearly articulated what kind of value it can solve for us.

So, “Aha, now it’s here. I don’t need to go and build or buy a quantum computer somewhere else.” And I tried using the Quickstart, and it turns out that the Hello World tutorial for quantum computing—at least to my mind—is basically an application for a PhD program at Berkeley. And I am not that type of academic for better or worse, so I kept smacking my head off of that and realizing, okay, whatever this is, is clearly not for me. You have been doing some deep dives in the quantum computing space, but as we’ve just mentioned, your day job is not, to my understanding, a college professor. You are a consultant, you run your own consultancy, solving data problems, particularly towards bioinformatics. What is the deal—to the layperson—of quantum computing these days?

Lynn: Well, yeah, like you, I was introduced years ago and tried to read the books, and I didn’t have the math and just, you know, saw it as a curiosity. Last year, I picked up a book from O’Reilly called Practical Quantum Computing, which of course, because the name was attractive to me. I read it, felt like I was getting a little bit more knowledge, implemented a learning JavaScript library with a browser-based editor—so zero-install—and it was a simulator, you couldn’t run it on actual QPUs. So, I decided to see if there’s any other interest in my tech community, and I got about five other developers and we ran a 15-week long book club because we all just wanted to move forward with our knowledge. Because there is this fundamental difference in the information you can get from a qubit versus a bit because a qubit can basically be, like, a globe, and so it has a superposition, and so you can have all the different mathematical points on the globe, versus a bit is on or off.

I mean, that’s intuitive, like, “Hey, I could get more information out of that.” So, the potential usages—it’s always been tech that leads the way—is on figuring out of what are called NP-hard or computationally complex problems, and, again, this is at the edge of my knowledge, but this is where bioinformatics is. I think of it in an oversimplified way, as [N by N by N by N, all by all by all 00:16:49]. We want to see all possible combinations of all possible inputs. So, for example, we can figure out which Covid drug we should try—which set of drugs we should try—and we want that as fast as possible.

So, I wanted to see, okay, you know, where’s this at? Plus, like you said, Amazon introduced Braket; when Amazon introduces something, then there’s some customers somewhere that are using it. I mean, that’s—you know, kind of pay attention to it now. So, as I was doing this book club, I investigated all the different cloud vendors and captured all that learning in a GitHub, and just recently recorded a LinkedIn Learning course. Which again, in the learning ladder is, if this is, you know, Hello World and this is actual implementation, it’s like right here.

But right here doesn’t exist. Like, there’s nothing there, so I tried to make something to say, okay, the Amazon Braket example, how does that actually work? What is a Hadamard Gate? Why do you care? What is amplification? How do you measure it? Like, what would you do with that? And so, you know, I tried to interpret some academic papers and do that learning layer in the middle to help move people towards productivity. Am I fully there? No. Did I move further? I hope so. Do you want to come along with me? Great.

Corey: You’ve done something, though, that I don’t think anyone else yet has when I had conversations with them about quantum computing, which is we all are shaped by our own needs and our own experiences when we interact with a cloud provider. To me, I, perhaps foolishly, took Amazon seriously when they called it Amazon Web Services. “Oh, okay. Clearly, this is going to be things to help me build websites and website accessories, more or less.” So, it’s always odd to me when I’ll see something like oh, and here’s our IoT solution that winds up powering a fleet of 10,000 robots, and I’m looking around my website going, “I don’t really have a problem that could be solved by the 10,000 robots. I have a bunch that could be made a lot worse.”

But it feels like it’s this orthogonal thing that is removed. But some areas, it’s okay. I can see the points of commonality and how you get there from here, and if I think really hard, I can do that with IoT stuff. For example, iRobot is a cloud-connected robot that talks to something that looks like a website and vacuums my house. Whereas with quantum computing, it always felt very isolated, very much an island as far as being connected to anything else that I can recognize. Bioinformatics research, as you describe it, well, yeah, I can see you get the bioinformatics research from web services. And now I can see how you can get to quantum computing through the bioinformatics side of things.

Lynn: Well, the other thing that really was useful for me, I am doing TensorFlow, finally. Took me a few years, but for neural networks. And so I am using, with some of my bioinformatics clients, acceleration with GPUs and TPUs, if I happen to be on Google because it’s a known thing that when you’re training a neural network, again, similar you have complexity, so you have a specialized chip, where you can offload some of the linear algebra onto that chip. So, you split the classic and the tensor portion, if you will, and you do computation on both sides. And so it’s not a huge leap to say, “Well, I’m not going to use a GPU, I’m going to use a QPU,” because you split. And that’s the way it actually works.

There’s actually a really interesting paper I put in my GitHub. It is a QCNN, and it is—that’s a Quantum Convolutional Neural Network that is used to analyze images of breast cancer. Because again, on the image, you can think of the pixels as what’s called a tensor, which is just vectors in multiple dimensions, you need the [all by all by all 00:20:17] again; that’s really how it goes in my head. You know, you have the globe of the qubit and you want to get the all possible combinations faster, so that you can analyze all combinations in the, in this case, the image. And they found, not only was it faster, it was more accurate. And that’s why I am interested in this.

Corey: Couchbase Capella Database-as-a-Service is flexible, full-featured and fully managed with built in access via key-value, SQL, and full-text search. Flexible JSON documents aligned to your applications and workloads. Build faster with blazing fast in-memory performance and automated replication and scaling while reducing cost. Capella has the best price performance of any fully managed document database. Visit couchbase.com/screaminginthecloud to try Capella today for free and be up and running in three minutes with no credit card required. Couchbase Capella: make your data sing.

Corey: The neat part is that this might be one of the first clear-cut stories where, “What could I use a quantum computer for?” And the answer isn’t something that’s forward-looking or theoretical. I mean, the obvious gag when you said reading about Practical Quantum Computing is that book is probably in pre-release, I would assume.

Lynn: [laugh].

Corey: But it’s a hard thing to solve for, and I do have the awareness that I am not an academic, academia has never been my friend, so I bias heavily for, “Well, can we use this to solve real-world problems slash make money?”—because industry—and academia focuses, ideally and aspirationally on the expansion of the limits of human knowledge. And sometimes it’s okay to do those things without an immediate, “Well, how can I turn a profit on it next quarter?” What a dismal, bleak society we have if that’s all that we wind up focusing on any given point in time.

Lynn: Yeah, that’s for sure.

Corey: Which, of course, sets us up for one other thing that’s a relatively recent change for you. You now have mentioned in your bio, which I believe is new since the last time we spoke, that you are an angel investor. And that is something that I recently found being applied to me as well after I made an investment in a startup that I was very excited about. I talked about in the show previously; it’s called Byte Check. But honestly, I didn’t realize that what I was doing was called angel investing until I read the press release because ‘strategic angel’ are two words that no one ever applies to me, particularly in that order. What happened? What are you doing these days?

Lynn: Well, I live in Minneapolis. So—and I moved there in 2019, so you know, my 2020 story is first I had Covid, got over that, and then I was there during the tragedy of George Floyd. So, I wanted to understand more about what were the root causes, and what I could do to make an impact in the recovery of my city. And I was really surprised to find that Minnesota is one of the most charitable states in the United States, it ranks one or two, but yet we have in the Twin Cities of Minneapolis and St. Paul, we have really unacceptable income inequality and poverty. So, something’s not working.

I’m a pretty charitable person; I always allocate a certain percentage of my money to charity, but I said, “I want to accelerate this.” So, at the same time, there was a new angel investment fund launched, it’s called Groove Capital, that was going to focus on women-owned and BIPOC businesses. And I thought, “Hmm, this seems good.”

Now, I was super intimidated because I lived in California for so many years, and check sizes in California, you just add a zero. And I thought, you know, “I don’t have generational wealth. This is my own money.” You know, I’m well-compensated, but I’m not loaded.

Corey: Yeah there’s a common trope right now that oh, angel investor is a polite way of saying I am rich—

Lynn: Right.

Corey: —but I rent my home at this point, living in San Francisco. It is, I am not exactly sitting here diving into a money bin out back, Scrooge McDuck-style either.

Lynn: Right. Well, I mean, you know, I’ll just be transparent about it. Like everybody else, or many people, I moved out of California because of the cost of doing business there and reduced my cost of living by 40% move into the Midwest, which is awesome. So anyway, I joined this fund, and it’s been just fantastic because I’ve listened to deals on my own and felt just like a complete, like, I don’t know what I’m doing. But I’m taking advantage—

Corey: How do you evaluate an idea that someone has that’s early-stage, barely better in some cases than back-of-an-envelope scrawlings?

Lynn: For sure, right. But what I found through the fund is I can contribute both money and time because, you know, I did this cloud expertise, and in addition to writing checks for a couple companies that I really believe in, for example, I got all these companies on the X cloud company for startups program. Because that wasn’t just a known thing in my ecosystem. I was like, “Why are you paying a cloud bill? You could be on the startup program for the first year.”

So, I’m impacting these new businesses with both my experience and my dollars, and I just really love it. I just really, really love it. And you know, the reasons I want to talk about it is because more people who have expertise in tech should do this because you can really, really be impactful. One of the companies that I invested in is called TurnSignl. They are coming to Los Angeles.

It was three attorneys and one of their brothers is a police officer. They wanted to de-escalate situations that happen with traffic stops. So, it’s a mobile app, where you push a button and you’re connected to an attorney. And they do training for the community and police officers, and the idea to record the conversation and to get an attorney involved to de-escalate and get everybody home safely. And that was my first investment and I’m—it’s going national, and I’m like, really, really—the kind of things I want to do you know.

Corey: It is simultaneously such a terrific idea and such a stunning indictment of the society that makes something like that necessary.

Lynn: Well, you know, we have to find practical solutions. We have to find ways forward.

Corey: Oh, please. Don’t interpret anything I’m saying a shade on that. It’s like, “Well, I wish the world were differently.” Yeah, I think most people do. But you have to deal for better or worse with the hand that you’re dealt, and this is, for better or worse, at the time of recording this, the society that we have, and finding the best path forward is often not easy.

But it beats just sitting here complaining about everything every day, and not doing anything to be part of that change. The surprising thing I learned as I went through it was that in many cases, the value of individual angel investors is not the check that they’re writing, that’s basically just almost a formality, on some level. It is the expertise, it is the insight into particular markets, and the rest. The part of what you’re saying that surprises me that I hadn’t really considered, but of course, it must exist, is the idea of angel funds. Is this generally run by an existing VC firm? Is it a group of like-minded friends who decide, ah, we’re going to just basically do the investing equivalent of a giving circle where everyone puts some money in the pot and then that decides where to go? How is it structured?

Lynn: Yeah, the way ours worked is you do pay a fee—it’s a small fee—to be part of it, and then they have people who vet deals for you. And then what I really like about it is the community aspect because just like in tech, when you’re learning something new in tech, you have community, same thing here. We have a Slack, we have a website for each deal, we have in-person meetups when Covid situation allows, and we have chosen to start by investing in Minnesota, although we’re going to, in fund two we’re going to invest in Upper Midwest. And for example, here’s something I would have never known. There’s an angel tax credit Minnesota, that for certain businesses, you can get a 25% tax credit. Which hey, do good, be good, get good. I would have never known about that, I would have never known how to do it. All my investments so far have qualified. Fantastic. My money goes further.

Corey: Yeah, it’s about well, what are you talking about worrying about taxes? That there’s about to be doing something good? Yeah, great. If you believe in a cause, take advantage of the tax code as written—I am not advocating tax fraud; pay every cent that you owe, let’s be serious here. They have no sense of humor about that—

Lynn: [laugh].

Corey: —and take advantage of that. That means you have additional money to do good with. I wish that more people had an awareness around that particular school of thought.

Lynn: Well, make your money go further, make your money effective.

Corey: Oh yes.

Lynn: Because like it or not, we run on money. We run on money. And so be smart, from everything where you shop to how you spend. That’s how we’re going to make change.

Corey: One last area I want to explore with you is that for a long time you’ve been working on, effectively, data pipelines and similar things in that space, tied to your consulting work. You are clearly skilled across all of the various cloud providers and even tieing into the expertise side of what you’re doing as an angel investor, you’ve always been a staunch advocate for, I guess we’ll call it doing security the right way. And I’ve always been tangentially related to security throughout the course of my career. And somewhat recently, I launched another day of my newsletter focused on security within AWS, for folks who are not themselves in the security space of what do you need to know. But so much of it comes down to the do the easy thing now, the right way to do it before you wind up having to do a whole bunch of damage control. And you’ve been advocating for that since before it was trendy to do so. I imagine you’re still somewhat passionate about that perspective.

Lynn: Well, I always like to say, you know, Werner Vogels doesn’t talk anything about tech; he just talks about, “Please use our security.” And I don’t blame him. I mean, you know, I joke that I am an AWS Community Hero because I made a bunch of YouTube videos about securing buckets. And that was, like, seven years ago and I just had a financial client, literally in November, and their buckets, you know, was made public because it was easy for the developer. I’m like, “Ugh, can we just do our foundations?”

I don’t know why it is not seen as a valuable skill. I mean, I’ve made craploads of money because people come after they have an incident, but you know, I wish we would be better. And I’m worried because as we start to get more and more of our health information in these big repositories—granted, we have some laws; yay, good—but it’s just not valued like coding up a new feature with node or something. And why not? I don’t understand.

So, I make all these educational resources: I make courses, I have GitHub repos, I have videos. You know, just do it. Plus the people who learned security. I mean, we are always in demand. I’m not a security professional, but I always do security kind of like as a courtesy. And people are like, “Oh, you know, you’re great. Oh, my friend needs you.” Dah-dah-dah… I mean, you’ll be working forever.

Corey: It feels like it’s aligned with cost in that it is almost a reactive function. You can spend all your time on it, but it’s not going to advance the state of your org further toward its stated goals. You’ve got to do it, but there’s also never really any ‘done’ there. It’s just easier for me on the cost side because I can very easily quantify the return on investment, whereas with security, it’s much more nebulous. And, of course, you wind up with the vendor—I’m going to call it what it is, in some cases—nonsense that is in this space, where, “Oh, you’re completely doomed, unless you buy their particular product.” You know, walk up or down the aisle at RSA a few times and your shopping cart is full. And great, are you more secure? You’re a lot more complex, but does this get you to a better outcome?

And it’s, I am so continually frustrated by all of these fancy whiz-bang solutions that are sort of going around the easy stuff—not easy, but it’s the baseline level of things: Secure your S3 buckets, or—for users themselves—it’s use a password manager that has a strong password on it, use it for everything, use MFA for the important things that you need to use, make sure your email is secure, don’t click random nonsense. There’s a whole separate pile of things. If I can click the wrong link in an email and it destroys my company, maybe it’s not me clicking that link in the email that’s the root problem here. Maybe there’s an entire security model revisitation that’s due. But I’m sorry, I will rant like a loon about the dismal state of security these days, if you let me, and you absolutely should not.

Lynn: Well, I would just entreat the audience, basic threat modeling is not complicated. It’s like cost modeling. It’s just a basic of having successful business on the cloud.

Corey: [sigh]. I wish the world work differently than it does, and yet here we are. Lynne, I really want to thank you for taking the time to come on the show a second time. If people want to learn more about what you’re up to and talk to you about anything we’ve discussed, what’s the best way to find you?

Lynn: So, if you can’t find me, you’re not looking. I have an internet-easy name. But two places that I’m pretty active: Twitter—just my name, @lynnlangit—and go to my GitHub. In particular, I have a learning cloud kind of meta-repository that has over 100 links to mostly free things on every cloud and just use them. Have at it, learn, be a practitioner, use the cloud more effectively.

Corey: And we will, of course, put links to that in the [show notes 00:32:25]. Thanks so much for coming back on. I really appreciate it.

Lynn: Thanks for having me. It’s been fun.

Corey: Lynn Langit, CEO of Lynn Langit Consulting, and oh so much more. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry comment talking about how security really isn’t that important, and right before you submit that comment accidentally type your banking password into the form, too.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Sean

Sean is a senior software engineer at TheZebra, working to build developer experience tooling with a focus on application stability and scalability. Over the past seven years, they have helped create software and proprietary platforms that help teams understand and better their own work.

Links:

  • TheZebra: https://www.thezebra.com/
  • Twitter: https://twitter.com/sc_codeUM
  • LinkedIn: https://www.linkedin.com/in/sean-corbett-574a5321/
  • Email: scorbett@thezebra.com

Transcript

Sean: Hello, and welcome to Screaming in the Cloud with your host, Chief cloud economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Today’s episode is brought to you in part by our friends at MinIO the high-performance Kubernetes native object store that’s built for the multi-cloud, creating a consistent data storage layer for your public cloud instances, your private cloud instances, and even your edge instances, depending upon what the heck you’re defining those as, which depends probably on where you work. It’s getting that unified is one of the greatest challenges facing developers and architects today. It requires S3 compatibility, enterprise-grade security and resiliency, the speed to run any workload, and the footprint to run anywhere, and that’s exactly what MinIO offers. With superb read speeds in excess of 360 gigs and 100 megabyte binary that doesn’t eat all the data you’ve gotten on the system, it’s exactly what you’ve been looking for. Check it out today at min.io/download, and see for yourself. That’s min.io/download, and be sure to tell them that I sent you.

Corey: This episode is sponsored in part by our friends at Sysdig. Sysdig is the solution for securing DevOps. They have a blog post that went up recently about how an insecure AWS Lambda function could be used as a pivot point to get access into your environment. They’ve also gone deep in-depth with a bunch of other approaches to how DevOps and security are inextricably linked. To learn more, visit sysdig.com and tell them I sent you. That’s S-Y-S-D-I-G dot com. My thanks to them for their continued support of this ridiculous nonsense.

Corey: Welcome to Screaming in the Cloud, I’m Corey Quinn. An awful lot of companies out they’re calling themselves unicorns, which is odd because if you look at the root ‘uni,’ it means one, but they’re sure a lot of them out there. Conversely, my guest today works at a company called TheZebra with the singular definite article being the key differentiator here, and frankly, I’m a big fan of being that specific. My guest is Senior Software Development Engineer in Test, Sean Corbett. Sean, thank you for taking the time to join me today, and more or less suffer the slings and arrows, I will no doubt be hurling your direction.

Sean: Thank you very much for having me here.

Corey: So, you’ve been a great Twitter follow for a while: You’re clearly deeply technically skilled; you also have a soul, you’re strong on the empathy point, and that is an embarrassing lack in large swaths of our industry. I’m going to talk about that right now because I’m sure it comes through the way it does when you talk about virtually anything else. Instead, you are a Software Development Engineer in Test or SDET. I believe you are the only person I’m aware of in my orbit who uses that title, so I have to ask—and please don’t view this as me in any way criticizing you; it’s mostly my own ignorance speaking—what is that?

Sean: So, what is a Software Development Engineer in Test? If you look back—I believe it was Microsoft originally came up with the title, and what it stems from was they needed software development engineers who particularly specialized in creating automation frameworks for testing stuff at scale. And that was over a decade ago, I believe. Microsoft has since stopped using the term, but it persists in areas in the industry.

And what is an SDET today? Well, I think we’re going to find out it’s a strange mixture of things. SDET today is not just someone that creates automated frameworks or writes tests, or any of those things. An SDET is the strange amalgamation of everything from full-stack to DevOps to even some product management to even a little bit machine-learning engineer; it’s a truly strange field that, at least for me, has allowed me to basically embrace almost every other discipline and area of the current modern engineering around, to some degree. So, it’s fun, is what it is. [laugh].

Corey: This sounds similar in some respects to oh, I think back to a role that I had in 2008, 2009, where there was an entire department that was termed QA or Quality Assurance, and they were sort of the next step. You know, development would build something and start, and then deploy it to a test environment or staging environment, and then QA would climb all over this, sometimes with automation—which was still in the early days, back in that era—and sometimes by clicking the button, and going through scripts, and making sure that the website looked okay. Is that aligned with what you’re doing, or is that a bit of a different branch?

Sean: That is a little bit of a different branch from me. The way I would put it is QA and QA departments are an interesting artifact that I think, in particular, newer orgs still feel like they might need one, and what you quickly realize today, particularly with modern development and this, kind of, DevOps focus is that having that centralized QA department doesn’t really work. So, SDETs absolutely can do all those things: They can climb over a test environment with automation, they can click the buttons, they can tell you everything’s good, they can check the boxes for you if you want, but if that is what you’re using your SDETs for you are, frankly, missing out because I guarantee you, the people that you’ve hired as SDETs have a lot more skills than that, and not utilizing those to your advantage is missing out on a lot of potential benefit, both in terms of not just quality—which is this fantastic concept that dates all the way back to—gives people a lot of weird feelings [laugh] to be frank, and product.

Corey: So, one of the challenges I’ve always had is people talk about test-driven development, which sounds like a beautiful idea in theory, and in practice is something people—you know, just like using the AWS console, and then lying about it forms this heart and soul of ClickOps—we claim to be using test-driven development but we don’t seem to be the reality of software development. And again, no judgment on these; things are hard. I built out a, more or less, piecing together a whole bunch of toothpicks and string to come up with my newsletter production pipeline. And that’s about 29 Lambdas Function, behind about 5 APIs Gateway, and that was all kinds of ridiculous nonsense.

And I can deploy each of the six or so microservices that do this, independently. And I sometimes even do continuous build or slash continuous deploy to it because integration would imply I have tests, which is why I bring the topic up. And more often than not—because I’m very bad at computers—I will even have syntax errors, make it into this thing, and I push the button and suddenly it doesn’t work. It’s the iterative guess-and-check model that goes on here. So, I introduced regressions, a fair bit at the time, and the reason that I’m being so blase about this is that I am the only customer of this system, which means that I’m not out there making people’s lives harder, no one is paying me money to use this thing, no one else is being put out by it. It’s just me smacking into a wall and feeling dumb all the time.

And when I talk to people about the idea of building tests. And it’s like, “Oh, you should have unit tests and integration tests and all the rest.” And I did some research into the topics, and a lot of it sounds like what people were talking about 10 to 15 years ago in the world of tests. And again, to be clear, I’ve implemented none of these things because I am irresponsible and bad at computers. But what has changed over the last five or ten years? Because it feels like the overall high level as I understood it from intro to testing 101 in the world of Python, the first 18 chapters are about dependency manager—because of course they are; it’s Python—then the rest of it just seems to be the concepts that we’ve never really gotten away from. What’s new, what’s exciting, what's emerging in your space?

Sean: There’s definitely some emerging and exciting stuff in the space. There’s everything from, like, what Applitools does with using machine learning to do visual regressions—that’s a huge advantage, a huge time saver, so you don’t have to look pixel by pixel, and waste your time doing it—to things like our team at TheZebra is working on, which is, for example, a framework that utilizes Directed Acrylic Graph workflows that’s written GoLang—the prototype is—and it allows you to work with these tests, rather than just as kind of these blasé scripts that you either keep in a monorepo, or maybe possibly in each individual services’ repo, and just run them all together clumsily in this, kind of, packaged product, into this distributed resource that lets you think about tests as these, kind of, user flows and experiences and to dip between things like API layer, where you might, for example, say introduce regression [unintelligible 00:07:48] calling to a third-party resource, and something goes wrong, you can orchestrate that workflow as a whole. Rather than just having to write a script after script after script after script to cover all these test cases, you can focus on well, I’m going to create this block that represents this general action, can accept a general payload that conforms to this spec, and I’m going to orchestrate these general actions, maybe modify the payload of it, but I can recall those actions with a slightly different payload and not have to write script after script after script after script.

But the problem is that, like you’ve noticed, a lot of test tooling doesn’t embrace those, kind of, modern practices and ideas. It’s still very much the, your tests, you—particularly integration tests do this—will exist in one place, a monorepo, they will have all the resources there, they’ll be packaged together, you will run them after the fact, after a deploy, on an environment. And it makes it so that all these testing tools are very reactive, they don’t encourage a lot of experimentation, and they make it at times very difficult to experiment, in particular because the more tests you add, the more chaotic that code and that framework gets, and the harder it gets to run in a CI/CD environment, the longer it takes. Whereas if you have something like this graph tool that we’re building, these things just become data. You can store them in a database, for the love of God. You can apply modern DevOps practices, you can implement things like Jaeger.

Corey: I don’t think it’s ever used or anything in the database. Great, then you can use anything itself as a database, which is my entire schtick, so great.

Sean: Exactly.

Corey: That’s right, that means the entire world can indeed be reduced to TXT records in DNS, which I maintain is the… the holiest of all databases. I’m sorry, please, continue.

Sean: No, nonono, that’s true. The thing that has always driven me is this idea that why are we still just, kind of, spitting out code to test things in a way that is very prescriptive and very reactive? And so, the exciting things in test come from places like Applitools and places like the—oh, I forget. It was at a Test Days conference, where they talked about—they developed this test framework that was able to auto generate the models, and then it was so good at auto generating those models for test, they’d actually ended up auto generating the models for the actual product. [laugh]. I think it used a degree of machine learning to do so. It was for a flashcard site. A friend of mine, Jacob Evans on Twitter always likes to talk about it.

These are where the exciting things lay is where people are starting to break out of that very reactive, prescriptive, kind of, test philosophy of, like I like to say, checking the boxes to, “Let’s stop checking boxes and let’s create, like insight tooling. Let’s get ahead of the curve. What is the system actively doing? Let’s check in. What data do we have? What is the system doing right at this moment? How
ahead of the curve can we get with what we’re actually using to test?”

Corey: One question I have is the cultural changes because back in those early days where things were handed off from the developers to the QA team, and then ideally to where I was sitting over in operations—lots of handoffs; not a lot of integrations there—QA was not popular on the development side of the world, specifically because their entire perception was that of, “Oh, they’re just the critics. They’re going to wind up doing the thing I just worked hard on and telling me what’s wrong with it.” And it becomes a ‘Department of No,’ on some level. One of the, I think, benefits of test automation is that suddenly you’re blaming a computer for things, which is, “Yep. You are a developer. Good work.” But the idea of putting people almost in the line of fire of being either actually or perceived as the person who’s the blocker, how has that evolved? And I’m really hoping the answer is that it has.

Sean: In some places, yes, in some places, no. I think it’s always, there’s a little bit more nuance than just yes, it’s all changed, it’s all better, or just no, we’re still back in QA are quote-unquote, “The bad guys,” and all that stuff. The perception that QA are the critics and are there to block a great idea from seeing fruition and to block you from that promotion definitely still persists. And it also persists a lot in terms of a number of other attitudes that get directed towards QA folks, in terms of the fact that our skill sets are limited to writing stuff like automation tooling for test frameworks and stuff like that, or that we only know how to use things like—okay, well, they know how to use Selenium and all this other stuff, but they don’t know how to work a database, they don’t know how an app [unintelligible 00:12:07] up, they don’t all the work that I put in. That’s really not the case. More and more so, folks I’m seeing in test have actually a lot of other engineers experience to back that up.

And so the places where I do see it moving forward is actually like TheZebra, it’s much more of a collaborative environment where the engineers are working together with the teams that they’re embedded in or with the SDETs to build things and help things that help engineers get ahead of the curve. So, the way I propose it to folks is, “We’re going to make sure you know and see exactly what you wrote in terms of the code, and that you can take full [confidence 00:12:44] on that so when you walk up to your manager for your one-on-one, you can go like, ‘I did this. And it’s great. And here’s what I know what it does, and this is where it goes, and this is how it affects everything else, and my test person helped me see all this, and that’s awesome.’” It’s this transition of QA and product as these adversarial relationships to recognizing that there’s no real differentiator at all there when you stop with that reactive mindset in test. Instead of trying to just catch things you’re trying to get ahead of the curve and focus on insight and that sort of thing.

Corey: This episode is sponsored in part by our friends at Vultr. Spelled V-U-L-T-R because they’re all about helping save money, including on things like, you know, vowels. So, what they do is they are a cloud provider that provides surprisingly high performance cloud compute at a price that—while sure they claim its better than AWS pricing—and when they say that they mean it is less money. Sure, I don’t dispute that but what I find interesting is that it’s predictable. They tell you in advance on a monthly basis what it’s going to going to cost. They have a bunch of advanced networking features. They have nineteen global locations and scale things elastically. Not to be confused with openly, because apparently elastic and open can mean the same thing sometimes. They have had over a million users. Deployments take less that sixty seconds across twelve pre-selected operating systems. Or, if you’re one of those nutters like me, you can bring your own ISO and install basically any operating system you want. Starting with pricing as low as $2.50 a month for Vultr cloud compute they have plans for developers and businesses of all sizes, except maybe Amazon, who stubbornly insists on having something to scale all on their own. Try Vultr today for free by visiting: vultr.com/screaming, and you’ll receive a $100 in credit. Thats V-U-L-T-R.com slash screaming.

Corey: One of my questions is, I guess, the terminology around a lot of this. If you tell me you’re an SDE, I know that oh, you’re a Software Development Engineer. If you tell me you’re a DBA, I know oh, great, you’re a Database Administrator. If you told me you’re an SRE, I know oh, okay, great. You worked at Google.

But what I’m trying to figure out is I don’t see SDET, at least in the waters that I tend to swim in, as a title, really, other than you. Is that a relatively new emerging title? Is it one that has historically been very industry or segment-specific, or you’re doing what I did, which is, “I don’t know what to call myself, so I described myself as a Cloud Economist,” two words no one can define. Cloud being a bunch of other people’s computers, and economist meaning claiming to know everything about money, but dresses like a flood victim. So, no one knows what I am when I make it up, and then people start giving actual job titles to people that are Cloud Economists now, and I’m starting to wonder, oh dear Lord, have I started the thing? What is, I guess, the history and positioning of SDET as a job title slash acronym?

Sean: So SDET, like I was saying, it came from Microsoft, I believe, back in the double-ohs.

Corey: Mmm.

Sean: And other companies caught on. I think Google actually [unintelligible 00:14:33] as well. And it’s hung on certain places, particularly places that feel like they need a concentrated quality department. That’s where you usually will see places that have that title of SDET. It is increasingly less common because the idea of having centralized quality—like I said before, particularly with the modern, kind of, DevOps-focused development, Agile, and all that sort of thing, it becomes much, much more difficult.

If you have a waterfall type of development cycle, it’s a lot easier to have a central singular quality department, and then you can have SDET stuff [unintelligible 00:15:08], that gets a lot easier when you have Agile and you have that, kind of, regular integration and you have, particularly, DevOps [unintelligible 00:15:14] cycle, it becomes increasingly difficult, so a lot of places that have been moving away from that. It is definitely a strange title, but it is not entirely rare. If you want to peek, put a SDET on your LinkedIn for about two weeks and see how many offers come in, or how many folks in your inbox you get. It is absolutely in demand. People want engineers to write these test frameworks, but that’s an entirely different point; that gets down to the point of the fact that people want people in these roles because a lot of test tooling, frankly, sucks.

Corey: It’s interesting you talk about that as a validation of it. I get remarkably few outreaches on LinkedIn, either for recruiting, which almost never happens or for trying to sell me something which happens once every week or so. My business partner has a CEO title, and he winds up getting people trying to sell him things four times a day by lunchtime, and occasionally people reaching out of, “Hey, I don’t know much about your company, but if it’s not going well, do you want to come work on something completely unrelated?” Great. And it’s odd because both he and I have similar settings where neither of us have the ‘looking for work’ box checked on LinkedIn because it turns out that does send a message to your staff who are depending on their job still being here next month, and that isn’t overly positive because we’re not on the market.

But changing just titles and how we describe what we do and how we do it absolutely has a bearing as to how that is perceived by others. And increasingly, I’m spending more of my time focusing less on the technical substance of things and more about how what they do is being communicated. Because increasingly, what I’m finding about the world of enterprise technology and enterprise cloud and all of this murky industry in which we swim, is that the technology is great—anything can be made to work; mostly—but so few companies are doing an effective job of telling the story. And we see it with not just an engineering-land; in most in all parts of the business. People are not storytelling about what they do, about the outcomes they drive, and we’re falling back to labels and buzzwords and acronyms and the rest.

Where do you stand on this? I know we’ve spoken briefly before about how this is one of those things that you’re paying attention to as well, so I know that we’re not—I’m not completely off base here. What’s your take on it?

Sean: I definitely look at the labels and things of that sort. It’s one of those things where humans like to group and aggregate things. Our brains like that degree of organization, and I’m going to say something that is very stereotypical here: This is helped a lot by social media which depends on things like hashtags and ability to group massive amounts of information is largely facilitated. And I don’t know if it’s caused by it, but it certainly aggravates the situation.

We like being able to group things with few words. But as you said before, that doesn’t help us. So, in a particular case, with something like a SDET title, yeah, that does absolutely send a signal, and it doesn’t necessarily send the right one in terms of the person that you’re talking to, you might have vastly different capabilities from the next SDET that you talk to. And it’s were putting up a story of impact-driven, kind of, that classic way of focusing on not just the labels, but what was actually done and who had helped and who had enabled and the impact of it, that is key. The trick is trying to balance that with this increasing focus on the cut-down presentation.

You and I’ve talked about this before, too, where you can only say so much on something like a LinkedIn profile before people just turn off their brains and they walk away to the next person. Or you can only put so much on your resume before people go, “Okay, ten pages, I’m done.” And it’s just one of those things where… the trick I find that test people increasingly have is there was a very certain label applied to us that was rooted in one particular company’s needs, and we have spent the better part of over a decade trying to escape and redefine that, and it’s incredibly challenging. And a lot of it comes down to folks like, for example, Angie Jones, who simply, just through pure action and being very open about exactly what they’re doing, change that narrative just by showing. That form of storytelling is show it, don't say it, you know? Rather than saying, “Oh, well, I bring into all this,” they just show it, and they bring it forward that way.

Corey: I think you hit on something there with the idea of social media, where there is validity to the idea of being able to describe something concisely. “What’s your elevator pitch?” Is a common question in business. “What is the problem you solve? What would someone use you for?”

And if your answer to that requires you sabotage the elevator for 45 minutes in order to deliver your message, it’s not going to work. With some products, especially very early-stage products where the only people who are working on them are the technical people building them, they have a lot of passion for the space, but they aren’t—haven’t quite gotten the messaging down to be able to articulate it. People’s attention spans aren’t great, by and large, so there’s a, if it doesn’t fit in a tweet, it’s boring and crappy is sort of the takeaway here. And yeah, you’re never going to encapsulate volume and nuance and shading into a tweet, but the baseline description of, “So, what do you do?” If it doesn’t fit in a tweet, keep workshopping it, to some extent.

And it’s odd because I do think you’re right, it leads to very yes or no, binary decisions about almost anything, someone is good or trash. There’s no, people are complicated, depending upon what aspect we’re talking about. And same story with companies. Companies are incredibly complex, but that tends to distill down in the Twitter ecosystem to, “Engineers are smart and executives are buffoons.” And anytime a company does something, clearly, it’s a giant mistake.

Well, contrary to popular opinion, Global Fortune 2000 companies do not tend to hire people who are not highly capable at the thing they’re doing. They have context and nuance and constraints that are not visible from the outside. So, that is one of the frustrating parts to me. So, labels are helpful as far as explaining what someone is and where they fit in the ecosystem. For example, yeah, if you describe yourself as an SDET, I know that we’re talking about testing to some extent; you’re not about to show up and start talking to me extensively about, oh, I don’t know, how you market observability products.

It at least gives a direction and bounding to the context. The challenge I always had, why I picked a title that no one else had, was that what I do is complicated, and if once people have a label that they think encompasses where you start and where you stop, they stop listening, in some cases. What’s been your experience, given that you do have a title that is not as widely traveled as a number of the more commonly used ones?

Sean: Definitely that experience. I think that I’ve absolutely worked at places where—the thing is, though, and I do want to cite this, that when folks do end up just turning off once they have that nice little snippet that they think encompasses who you are—because increasingly nowadays, we like to attach what you do to who you are—and it makes a certain degree of sense, absolutely, but it’s very hard to encompass those sorts of things, and let alone, kind of, closely nestle them together when you have, you know, 280 characters.

Yes, folks like to do that to folks like SDETs. There’s a definite mindset of, ‘stay in your lane,’ in certain shops. I will say that it’s not to the benefit of those shops, and it creates and often aggravates an adversarial relationship that is to the detriment of both, particularly today where the ability to spin up a rival product of reasonable quality and scale has never been easier, slowing yourself down with arbitrary delineations that are meant to relegate and overly-define folks, not necessarily for the actual convenience of your business, but for the convenience of your person, that is a very dangerous move. A previous company that I worked at almost lost a significant amount of their market share because they actively antagonized the SDET team to the point where several key members left. And it left them completely unable to cover areas of product with scalable automation tooling and other things. And it’s a very complex product.

And it almost cost them their position in the industry, potentially, the entire company as a whole got very close to that point. And that’s one of the things we have to be careful of when it comes to applying these labels, is that when you apply a label to encompass someone, yes, you affect them, but it also we’ll come back and affect you because when you apply that label to someone, you are immediately confining your relationship with that person. And that relationship is a two-way street. If you apply a label that closes off other roads of communication or potential collaboration or work or creativity or those sorts of things, that is your decision and you will have to accept those consequences.

Corey: I’ve gotten the sense that a lot of folks, as they describe what they do and how they do it, they are often thinking longer-term; their careers often trend toward the thing that happens to them rather than a thing that winds up being actively managed. And… like, one of my favorite interview questions whenever I’m looking to bring someone in, it’s always, “Yeah, ignore this job we’re talking about. Magically you get it or you don’t; whatever. That’s not relevant right now. What’s your next job? What’s the one after that? What is the trajectory here?”

And it’s always fun to me to see people’s responses to it. Often it’s, “I have no idea,” versus the, “Oh, I want to do this, and this is the thing I’m interested in working with you for because I think it’ll shore up this, this, and this.” And like, those are two extreme ends of the spectrum. There’s no wrong answer, but it’s helpful, I find, just to ask the question in the final round interview that I’m a part of, just to, I guess sort of like, boost them a bit into a longer-term picture view, as opposed to next week, next month, next year. Because if what you’re doing doesn’t bring you closer to what you want to be doing in the job after the next one, then I think you’re looking at it wrong, in some cases.

And I guess I’ll turn the question on to you. If you look at what you’re doing now, ignore whatever you do next, what’s your role after that? Like, where are you aiming at?

Sean: Ignoring the next position… which is interesting because I always—part of how I learned to operate, kind of in my earlier years was focus on the next two weeks because the longer you go out from that window, the more things you can’t control, [laugh] and the harder it is to actually make an effective plan. But for me, the real goal is I want to be in any position that enables the hard work we do in building these things to make people’s lives easier, better, give them access to additional information, maybe it’s joy in terms of, like, a content platform, maybe it’s something that helps other developers do what they do, something like Honeycomb, for example, just that little bit of extra insight to help them work a little bit better. And that’s, for me, where I want to be, is building things that make the hard work we do to create these tools, these products easier. So, for me, that would look a lot like an internal tooling team of some sort, something that helps with developer efficiency, with workflow.

One of the reasons—and it’s funny because I got to asked this recently: “Why are you still even in test? You know what reputation this field has”—wrongly deserved, maybe so—“Why are you still in test?” My response was, “Because”—and maybe with a degree of hubris, stubbornly so—“I want to make things better for test.” There are a lot of issues we’re facing, not just in terms of tooling, but in terms of processes, and how we think about solving problems, and like I said before, that kind of reactive nature, it sort of ends up kind of being an ouroboros, eating its own tail. Reactive tools generate reactive engineers, that then create more reactive tools, and it becomes this ouroboros eating itself.

Where I want to be in terms of this is creating things that change that, push us forward in that direction. So, I think that internal tooling team is a fantastic place to do that, but frankly, any place where I could do that at any level would be fantastic.

Corey: It’s nice to see the things that you care about involve a lot more about around things like impact, as opposed to raw technologies and the rest. And again, I’m not passing judgment on anyone who chooses to focus on technology or different areas of these things. It’s just, it’s nice to see folks who are deeply technical themselves, raising their head a little bit above it and saying, “All right, here’s the impact I want to have.” It’s great, and lots of folks do, but I’m always frustrated when I find myself talking to folks who think that the code ultimately speaks; code is the arbiter. Like, you see this with some of the smart contract stuff, too.

It’s the, “All right, if you believe that’s going to solve all the problems, I have a simple challenge to you, and then I will never criticize you again: Go to small claims court for a morning, four hours and watch all the disputes that wind up going through there, and ask yourselves how many of those a smart contract would have solved?”

Every time I bring that point up to someone, they never come back and say, “This is still a good idea.” Maybe I’m a little too anti-computer, a little bit too human these days. But again, most of cloud economics, in my experience, is psychology more than it is math.

Sean: I think it’s really the truth. And I think that [unintelligible 00:29:06] that I really want to seize on for a second because code and technology as this ultimate arbiter, we’ve become fascinated with it, not necessarily to our benefit. One of the things you will often see me—to take a line from Game of Thrones—whinging about [laugh] is we are overly focused on utilizing technology, whether code or anything else, to solve what are fundamentally human problems. These are problems that are rooted in human tendencies, habits, characters, psychology—as you were saying—that require human interaction and influence, as uncomfortable as that may be to quote-unquote, “Solve.”

And the reality of it is, is that the more that we insist upon, trying to use technology to solve those problems—things like cases of equity in terms of generational wealth and things of that sort, things like helping people communicate issues with one another within a software development engineering team—the more we will create complexity and additional problems, and the more we will fracture people’s focus and ability to stay focused on what the underlying cause of the problem is, which is something human. And just as a side note, the fundamental idea that code is this ultimate arbiter of truth is terrible because if code was the ultimate arbiter of truth, I wouldn’t have a job, Corey. [laugh]. I would be out of business so fast.

Corey: Oh, yeah, it’s great. It’s—ugh, I—it feels like that’s a naive perspective that people tend to have early in their career, and Lord knows I did. Everything was so straightforward and simple, back when I was in that era, whereas the older I get, the more the world is shades of nuance.

Sean: There are cases where technology can help, but I tend to find those a very specific class of solutions, and even then they can only assist a human with maybe providing some additional context. This is an idea from a Seeking SRE book that I love to reference—I think it’s, like, the first chapter—the Chief of Netflix SRE, I think it is, he talks about this is this, solving problems is this thing of relaying context, establishing context—and he focused a lot less on the technology side, a lot more of the human side, and brings in, like, “The technology can help this because it can give you a little bit better insight of how to communicate context, but context is valuable, but you’re still going to have to do some talking at the end of the day and establish these human relationships.” And I think that technology can help with a very specific class of insight or context issues, but I would like to reemphasize that is a very specific class, and very specific sort, and most of the human problems we’re trying to solve the technology don’t fall in there.

Corey: I think that’s probably a great place for us to call it an episode. I really appreciate the way you view these things. I think that you are one of the most empathetic people that I find myself talking to on an ongoing basis. If people want to learn more, where’s the best place to find you?

Sean: You can find me on Twitter at S-C—underscore—code, capital U, capital M. That’s probably the best place to find me. I’m most frequently on there.

Corey: We will, of course, include links to that in the [show notes 00:32:37].

Sean: And then, of course, my LinkedIn is not a bad place to reach out. So, you can probably find me there, Sean Corbett, working at TheZebra. And as always, you can reach me at scorbett@thezebra.com. That is my work email; feel free to email me there if you have any questions.

Corey: And we will, of course, put links to all of that in the [show notes 00:33:00]. Sean, thank you so much for taking the time to speak with me today. I really appreciate it.

Sean: Thank you.

Corey: Sean Corbett, Senior Software Development Engineer in Test at TheZebra—because there’s only one. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry ranting comment about how absolutely code speaks, and it is the ultimate arbiter of truth, and oh wait, what’s that the FBI is at the door make some inquiries about your recent online behavior.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Tyler

Lifelong learner, passionate coach, obsessed with continuous improvement, avid solver of people puzzles.

Links:

  • United Airlines: https://www.united.com/
  • LinkedIn: https://www.linkedin.com/in/tylerslove/

TranscriptAnnouncer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Couchbase Capella database as a service is flexible, full-featured, and fully managed with built-in access via Key-Value SQL, and full-text search. Flexible JSON documents align to your applications and workloads. Build faster with blazing fast in-memory performance and automated replication and scaling, while reducing costs. Capella has the best price-performance of any fully managed document database. Visit couchbase.com/ScreamingintheCloud to try Capella today for free, and be up and running in 3 minutes. No credit card required. Couchbase Capella make your data sing.

Corey: This episode is sponsored by our friends at Oracle HeatWave is a new high-performance query accelerator for the Oracle MySQL Database Service, although I insist on calling it “my squirrel.” While MySQL has long been the worlds most popular open source database, shifting from transacting to analytics required way too much overhead and, ya know, work. With HeatWave you can run your OLAP and OLTP—don’t ask me to pronounce those acronyms again—workloads directly from your MySQL database and eliminate the time-consuming data movement and integration work, while also performing 1100X faster than Amazon Aurora and 2.5X faster than Amazon Redshift, at a third of the cost. My thanks again to Oracle Cloud for sponsoring this ridiculous nonsense.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Calling this show Screaming in the Cloud has been pretty… easy most of the time because that’s mostly what I do: I shake my fist and I yell at clouds. And most companies are okay with that. Today’s guest is likely a little bit on the other side of that because when I’m screaming at clouds, it’s often out the window, when I’m in a plane.

Today, I’m joined by Tyler Slove, who’s a Senior Manager in the Enterprise Cloud and DevOps Group at United Airlines, a company I spend way too much time dealing with when we’re not in the midst of a global pandemic. Tyler, thank you for joining me.

Tyler: Yeah. Thanks for the invite, Corey. Really excited to be here.

Corey: So, I want to talk a little bit about, first, how glad I am to finally talk to you because airlines are kind of like computers—and particularly cloud—where when you first see it, it is magic; it is transformative, it’s endless possibilities, the power of flight slash instant provisioning of computer resources. Okay, so not everyone is going to find those quite the same way. What’s novel today is commonplace tomorrow, and then you get annoyed because your plane is 20 minutes late as it hurls you through the sky to the other side of the planet with the miracle of flight while you’re on the internet the whole way. And it’s one of those problems where it is sort of definitionally, a thankless job. It is either in the background that just empowers things, or everyone’s yelling at you on Twitter. So, given that you work with both sides of that, how do you find that commonality to play out in your world?

Tyler: Yeah, it’s an interesting thought, and I hadn’t necessarily connected the dots before. Because I, like you, are just as frustrated when that flight is, like, 20 minutes delayed. It’s like, “Oh, I wanted to be—[laugh]—where I wanted to be at that time.” And, you know, when you think about it, it’s actually an ongoing joke I have with one of my mentors. Like, airlines should not work; when you think about the maintenance, the aircraft, the crews, the weather, legal stuff, like, it’s amazing how complex they are, and it’s something that’s kept me interested for, you know, the first three years that I’ve been here.

But it is similar, actually, to being in an operational role, right? You do everything right, everything’s resilient, you roll through an Amazon, like, region-specific issue without any blips, and no one reaches out to you. But you know, you have one issue, and then it’s you’re getting out of bed at three in the morning, and everyone’s got a big retrospective about why you didn’t do something that could have resulted in that not happening. And I can see the parallel.

Corey: We all tend to have blind spots, and I more or less had my idea of big enterprise technology fixed a while back. And it occurred to me a few years ago that this is probably no longer accurate because I’m sitting here thinking of, well, United Airlines—with whom I do extortionately large amount of travel, let’s be very clear here; we’re talking I think I did 140,000 miles domestically flown in 2019, the last year that was even close to normal. Protip: Don’t fly that much. It really winds up doing a number on your internal clock and having any semblance of life. But I’m sitting there thinking that it’s old-school technology; there’s a mainframe that powers all of this, and all of the staff checking me in are using these ancient Unix green screens has always been my assumption.

And that thought occurred to me as I’m staring at my iPhone, checking in automatically in the mobile app—that was very modern and working at the same time—and the penny finally dropped for me of this is probably not accurate, how I’m envisioning the technology on the back end working. And there have been announcements that United is moving an awful lot of its systems to AWS specifically. What is that—I don’t want to call it modernization because that sends the wrong undertone or subtext to it, but what has that cloud transformation been like?

Tyler: So, it’s the marrying together of those two things without the time that you would potentially want to just rewrite the functionality that the mainframes that have gotten us to do the amount of you know flights and revenue that we do, and that are rock solid, like, we don’t get the chance to shut that thing down for three months and rebuild it—or what would be, realistically, more like three years. So, it’s how do we build a—

Corey: Yeah, it’s a heck of a delay notice to put on the airport flight thing: “Flight delayed?” “Oh, when is it rescheduled to?” “2025.” Yeah, turns out that doesn’t usually happen.

Tyler: Yeah, and so we’ve got to do it at the same time. And there’s, you know, analogies of, like, changing the tire while you’re driving or changing the engine on the jet while it’s flying. And we’ve actually—it’s felt like that, but it’s been in an exciting way. So, we really are able to decouple the front end from the back end or some of the core systems and then, piece-by-piece, modernize them, and do them in a way that is safe and responsible, given you know, the amount of folks that are relying on us to get to where they want to go every day.

So yeah, it’s been challenging for sure, but it’s also the right thing to do. It’s the direction we need to go where we can focus more of our engineering talent, which is scarce or limited, you know, we would rather have folks invested in improving the user experience instead of—what we have is a world-class data center, but you know, the number of people that are focused on making that what it is, I would much rather see that happen—or that investment be put into a higher up the value chain.

Corey: It’s also, on some level, on a baseline trying to understand how it all fits together. You look at the challenges that an airline has, you have challenges with labor, with press, with you know, the big problem of the logistics of not just the scheduling and the rest of making sure that everything flows throughout an enormous what is effectively logistics network, but also the, you know, the minor detail of keeping the planes in the sky when they’re supposed to be in the sky. And it feels like on some other you flip through the list of concerns a company has, and technology in the computer sense feels like it’s going to be, like, chapter 47 of that giant book. Obviously, that’s not true because technology is an empowering story. It is not just the booking system; it controls, more or less, everything.

At some level, I’d like to make fun of big companies saying, “Oh, we’re not a”—insert whatever the company really does here—“We’re a tech company.” But without technology, I don’t think you, at this point, have much of an airline. How do you see yourselves in the broader sense? Are you increasingly a tech company?

Tyler: We are increasingly a tech company. I think we’re… we’re seen as partners with the VPs of the different functional areas, right? It’s not a separation of the business and IT the way that maybe we would have thought about it five or ten years ago. It’s, both of us can’t be successful without each other, and the functions have come to trust that we will spend the time we need to understand the problems that they’re solving, and we’ll bring different perspectives, we’re going to bring technical solutions, but we’re also going to bring, you know, potentially system or flow changes and business process improvements. And that takes some getting—that right a few times and building up that trust and spending the time you need to, like, go past, “Oh, here’s a set of user stories. Just do them.” Of, like, “What are we trying to solve here? Could we just remove this process? Do we even need to do this thing anymore?” And once you prove yourself, I’ve never felt like we’ve been put in a backroom or seen as a lower priority. We’re working on the same stuff together, and we win or lose together.

Corey: I know a lot about the airline industry because I go to tech conferences, and when I’m at tech conferences, invariably the speaker—who’s usually J. Paul Reed, but not always—decides to talk about computers, and incident response, and the rest through the lens of the airline industry, which for some reason has always been one of those neck and neck things that are just completely inseparable for those types of talks. And they talk about airline incidents, and very often it’s not even, like, the horrifying headline-making stuff, but things like two aircraft passed closer to one another than they should have, and the NTSB does a full investigation. And they talk about how, “Oh, this is exactly the sort of thing you should do whenever there’s a computer-related issue.” And I am curious, given that you do in fact have those investigations with the plane-facing stuff, how much of that culture carries over into the, “Hmm. We took a systems outage on the computer side.” And how much of that is similar versus how much of this is just conference-ware.

Tyler: It’s actually quite similar; that part of our culture permeates through. And we’re actually looking at what’s the right level of time to spend to get to the root cause when sometimes it’s hard to explain in computers. Or there’s so many variables that it’s going to take us, you know, weeks or dozens of hours to really get there. But yeah, after any significant incident, we’re religious about having a follow-up problem review where we get all the information that we need, and we, kind of, are expected to figure out exactly—like, replay what happened, step-by-step, and what were the controls that were in place to avoid such a thing, and were those complied with or not, et cetera. And earlier at my time in United, definitely was frustrated with how—I’m like, “I just need to get back to delivery. We’ve got this—this sprint is ending, and I can’t spend four hours doing this.”

Like, that was a… what was seen as, like, a one-time event. And I don’t think that all the things that culminated in that are going to happen again, and I’ve done a few things that I feel are going to mitigate the risk moving forward, but actually, I’ve changed my perspective on this now. So, we are forcing—or not even forcing; we’re simulating major incidents and then doing that type of a problem review so that we can learn ahead of time and we can make it a heck of a lot more fun [laugh] and open and transparent conversation. So hey, me or someone from my team gets behind the curtain and, like, creates some simulation of a major issue in one of our pre-production environments, and then the team that’s responsible for the operations and whatnot of that response.

And we look at what alerts went off? What alerts do we expect to go off that didn’t? What was maybe a leading indicator that we aren’t yet looking at? And kind of so we’re calling that a game day, and we took that, you know, from—AWS has influenced our thinking on that, or they contributed to it. And it’s a really good way to build those relationships, when there’s not a lot on the line, you’re not coming around what could be a customer-impacting negative experience, which is, you know, really what drives us to do good work is to make sure that never happens.

And it does happen, but you know, we’re getting more and more resilient. And this is a way to turn that on its head and be able to take the positive of that, and get the spirit, and get people to collaborate better because they—like, “Hey, I did that fun thing together. Now, when we’re in the heat of it, we’re going to collaborate better, we’re going to be, kind of, more open with the information we’re sharing because we understand each other’s people and their intentions, and you know, where someone’s coming from.” So, yeah, we were pretty excited about that.

Corey: I have to admit I’m a little on the envious side about how your timing has worked out. Because back in 2008, when the cloud was still a new thing and some of the early adopters were diving in, the experience really sucked. I mean, this was before CloudFormation and other ways of managing systems. And by migrating over the last few years, so many of those sharp edges have been smoothed, and established patterns and processes, and understanding of how cloud interplays with enterprise IT has evolved dramatically. What has been your experience migrating to AWS? What’s worked well and what hasn’t?

Tyler: Yeah, so the migration itself has been very deliberate. So, we were focused on AWS from the beginning, and it was—we believe that they’re a leader, that they’re going to give us what we need, but also we didn’t want to fragment our engineers across multiple platforms and have them have to pick a team. Like, “Am I going to choose to learn how to build stuff in AWS, or GCP?” So, from just a transformation, and to get everybody on the same page, and upskill the organization, we’re focused on AWS. And there’s definitely, like, some learning curve, or moving into an environment where there used to be a centralized team that handled a lot of stuff for you and made it magic—like, as an engineer; I just have to make sure that my app builds, and then I can send it to someone, and they’re going to deploy it, and it’s going to work and then you know, we… shifting the responsibility to, okay, we actually believe that if—we could do that; we could just have the same function that did that in the on-prem world, do that for you in the cloud world, but our belief is that we come up with better software when the engineer understands and can control the entire workload and that it’s like, “Hey, I can configure my app to take advantage of this particular portion of the underlying infrastructure.”

And that became very clear with, like, Lambda or things like that, where it’s… you know, there’s only so many configurations, and it doesn’t make sense to try to get someone else to do that for you. So, there’s mindset changes that had to happen. There’s also just, like, proving it out. Like, is this going to be more reliable than our data center, which is extremely reliable? And there have been issues in the cloud, like, where we have something running parallel, and we have a cloud issue and it didn’t impact on-prem.

So, how do we learn from that? And then how do we kind of continue on and figure out, how do we build resilient workloads in the cloud? How do we make sure that we cover our bases on not just getting it running, but like, getting it running the right way, and then doing the testing that we need to do—like I mentioned earlier on the game days—to really be confident in it so that we can ultimately move away from needing to have any sort of backup in the data center.

Corey: I was poking around in an AWS account recently, and it looked like there were seven different ways of managing the systems that have been brought to bear in that account, and different design philosophies, competing approaches. And the sad part is that this was my personal AWS account. No one else has ever built anything in that account except for me. And if I have that problem as one person—admittedly a strange person—I can’t imagine what the governance story around something like AWS looks like for an organization that has thousands of people working in your IT org. How do you wind up managing the way to build things appropriately?

I can’t fathom—even though I am a fan of ClickOps—just letting everyone loose with admin rights in the AWS console. There has to be some form of gating approach. Is that done through patterns? Is that done through some sort of internal platform that abstracts away for folks? How are you managing this?

Tyler: Yeah, so this is one of the things that led to a learning curve at the beginning, but I think it’s worthwhile. And I can’t take credit for this because it was a decision that happened before I came, but we’re all-in on infrastructure as code. So, we’re not extremely prescriptive about what that means across the entire enterprise, but you cannot deploy anything into an environment, like, higher than a development area without it being defined as CloudFormation and promoted through. And that allows us consistency, auditability, [laugh] and a lot of other things.

So, that was kind of phase one, and that’s been—I believe—in place since we started in the cloud. Like, maybe there were some pocket accounts and some things that existed before, but once we were all-in, and it was, kind of, official that’s been in place. And I’m glad we held to that because there’s been a lot of, like, “Oh, just remove that. Let people build stuff through the console because they need to move fast.” And we’re like, “Yes, that would move them fast right now, but the level of inconsistency would be extremely risky to be able to handle that, and handle production incidents if you don’t have a pre-prod environment to test the patch that you’re trying to put in on the fly, that manages hundreds of orders a second.”

So, we started with CloudFormation. We were kind of all-in on CloudFormation, and then over the last year or so—maybe a little bit longer—it’s become apparent that CloudFormation has some limitations. And it can be also intimidating to have to, in excruciating detail, like, define every single parameter of every resource you’re trying to create. And—

Corey: It’s wordy. It’s YAML or JSON, whichever one you hate the most, invariably, is the one you’re dealing with today. And yeah, it has its limitations.

Tyler: Yeah. And then they’re sharing that happens, right? So, it’s like, I’ve got someone that I go to lunch with, that’s like, “Oh, I just built this solution. It’s all in CloudFormation.” They send it over, and then I’m looking at, it’s like, “Can I reuse this? Which parameters here are things that I should change for my app, and which ones are there because security mandated it, or it’s part of, like, a corporate compliance thing, or other reasons why?”

So, what we are really excited about in the last few months, we’ve really invested in CDK constructs and being able to define. You know, as my small team, we have visibility and strong, like, partnerships with our cloud engineering group, with our security groups, and whatnot, and we can say, “Hey, if you want to build an ECS cluster, like, this is a good, known way to start.” And you can just provide, like, X number of parameters that are meaningful to you, and you can inherit all the rest. And you’re going to get our logging standards, you’re going to get our security standards, all that, like, more or less built-in. And we also can version that.

So, we can know, hey, this person built off the CDK App 1.1, and then we have some sort of security change, right? So say, now we want to install some other agent on all these things. And it’s like, “Okay, all the ones that were deployed on 1.1, we need to move it from 1.1 to 1.2.”

And we can test what that upgrade path looks like in a lab environment, and then we can, you know, release it and have, you know, 30 different app teams all consume that update in a relatively self-service manner that means we don’t have to do it one by one. And then, yeah, it just gives us the ability to respond to stuff as quickly as we need to in the current environment.

Corey: Today’s episode is brought to you in part by our friends at MinIO the high-performance Kubernetes native object store that’s built for the multi-cloud, creating a consistent data storage layer for your public cloud instances, your private cloud instances, and even your edge instances, depending upon what the heck you’re defining those as, which depends probably on where you work. It’s getting that unified is one of the greatest challenges facing developers and architects today. It requires S3 compatibility, enterprise-grade security and resiliency, the speed to run any workload, and the footprint to run anywhere, and that’s exactly what MinIO offers. With superb read speeds in excess of 360 gigs and 100 megabyte binary that doesn’t eat all the data you’ve gotten on the system, it’s exactly what you’ve been looking for. Check it out today at min.io/download, and see for yourself. That’s min.io/download, and be sure to tell them that I sent you.

Corey: It’s a constant challenge and it’s really neat seeing the adoption of things like the CDK, which I’ve always sort of mentally put on the same stack as, “Oh, yeah, this is something that scrappy tiny startups use.” But you’re the exact opposite of that. The fact that you’re using it and finding success with it says a lot. I think you’re also right there with the most nimble, advanced, tiniest of startups in the world, and you’re still trying to figure out how to contextualize this into the broader lifecycle and understand the long-term architectural implications of how this stuff works. If it helps anything, I can assure you, you are very far from alone.

If anyone else is feeling that way, exactly the same position. And if you’re out there saying, “Oh, yeah. We’ve solved this. This is how we do it.” Find a second person to agree with you. But then come talk to me. Because everyone solves it locally; no one solves that globally. It’s a hard problem.

Tyler: Yeah. We’ve had this vision of, like, a vending machine for stuff. And then we’ve tried that in different ways and templates, and we think that this is the right pattern.

Corey: Yeah, every time AWS builds a vending machine for accounts and whatnot, it’s like the worst kind of vending machine; the kind that eats all your money.

Tyler: Service catalog. Yeah.

Corey: Yeah. It becomes a disaster. So, I want to talk about a couple of other things as well. When we started talking a year or so ago, you were a team lead. Today, you are a senior manager, and it turns out that, unlike when you start your own company and can invent your own made-up title, like, Cloud Economist, those words mean things. So first, congratulations on the promotion, how’d it come about?

Tyler: Thank you. Yeah, it came about—I guess, I really have always been passionate about people leadership, but I know that in order to properly lead and, like, have the context, and you need to know what it’s like to do these hard things that my team is solving, and be responsible for those, kind of, as an individual. So, you know, I’ve been spending the last, like, five or so years as an individual contributor, kind of learning how all this stuff works, and then learning from a lot of different managers. You know, I’ve been really lucky to have some people that, kind of, took me under their wing, coached me, and is just, like, the person that puts the wind in your sails, but like, not in a… not in a fake way, but like actually sees you and puts you into situations that are going to force you to grow and have your back if something goes wrong. And I kind of saw that and I wanted to be that for someone else.

So, you know, it’s… yeah, it was something that I kind of put my hat in the ring, and a position came and I was tapped to step up and do it. But it was initially for a very small team, right, so a three-person team. But it’s since expanded to be six or seven over the next month or so.

Corey: One of the things that I found always interesting slash admirable about you is we travel in somewhat similar circles. We both have pitched in from time to time as mentors in Forrest Brazeal’s cloud resume challenge, and it’s nice to see people who are working at established companies who are very busy with their day jobs, also taking the time out of the day to help, effectively, what is the next generation of cloud engineer find their way within this industry. How did you get onto that track?

Tyler: Yeah, so I guess it’s, you got to send the elevator back down. I have the experience of, kind of, being on the edge of, like—I was on the waitlist for my university, I had to—also was on the waitlist for my first job as a rotational program, and there was always kind of this, like, I had to claw for it, I had to prove myself, and also had to—I was the first in my family to pursue opportunities like this. And I got the itch for it, then I also see there’s so much potential in folks. And like, even looking at my parents as examples, right? My father’s an auto mechanic, and he’s probably one of the smartest people I know, but didn’t really… have the opportunity to get into technology. [unintelligible 00:22:44] kind of in a blue-collar job.

But I just feel like there’s so much untapped potential, and I am passionate about helping people at least, like, understand what opportunities are available to them. And not just assume that if you don’t have an example of someone who’s a software engineer in your life, or a sibling, or a parent, like, that’s outside of your reach.

Corey: I love the phrase, ‘send the elevator back down’ because it’s true. I feel like the only reason that anyone that you have ever heard of in tech, who you have any modicum of respect for—and I include both of us on that list as well, but basically everyone else in the industry, too—the only reason all of us are here in the roles that we’re in is that at some point, someone did a favor for us that they didn’t have to, but they did. And it’s almost impossible to pay that back, so instead, I’ve stopped trying. I instead try to do those favors in a forward-looking way for other people whenever I can. And there’s a lot to be said for expressing that through a way of helping people find their way and see what happens.

Because let’s face it, the industry that you and I came up in doesn’t really exist in the same way. There is no fleet of help desk positions out there the way there was when I first started getting exposed to technology, that would get me into this direction, so people have to come through alternate paths. And some people try and express that through advice that no longer applies for a world long gone. I try and at least keep up with what’s going on in this space.

Tyler: Yeah, absolutely. It’s a dynamic environment for sure, and when I look at just how challenging it is to try to, like, find a senior cloud engineer, and then looking at, okay, is what we’re doing here, like, really rocket science? Does it require ten years of experience? And I think the answer is no, like, we’ve got a small enough group here, we know what we’re doing, and everyone’s passionate about bringing other people up and, like, finding their strengths, giving them a problem, not giving them the answer to the problem, and kind of strategically building to bigger, bigger things until the next day, you know—or before you know it, they’re able to solve problems that you would have previously thought, like, “Oh, that’s something that I have to get my hands on.” And it’s just so powerful to see that and to be part of that. So, that’s kind of the approach we’re taking.

Corey: It refreshing to see. So, many companies are requiring that they hire senior talent, and they can’t take junior talent because, “Oh, that person would take six months to come up to speed in this environment. We want to hit the ground running.” And the job req has been open for nine months. At some point, building talent becomes the best slash only way forward.

I’m still at a scale now where I’m not in a position be able to do that, just because we are dropping principal consultants into dynamic strange situations, and that is a terrible environment for a junior, but as you scale past a certain point—I don’t really know what that point is, but yes, United Airlines has scaled past that point—bringing folks up, taking interns, making interns job offers, and continuing to expand what is happening, I think, on some level, one of the big hiring challenges for United and other similarly situated companies has been that, oh, the technology must be ancient caribou-era of trekking across the tundra level of development. But we just talked about using the CDK, and pattern design for things. The public perception and the reality are incredibly divergent.

Tyler: Yeah. Maybe I’m strange in this regard. But since college, I’ve worked only in very, very large organizations. And seeing the satisfaction that you have, or you can get from working with those systems, and being able to churn out a modern customer experience, or modernizing the system for operational efficiency, just it’s very satisfying to me to be in that environment. I know that it probably scares other people away.

But it’s just the scale; it’s hard to get that scale somewhere more—I don’t know, I guess, like, younger, newer because you don’t have years of legacy. But I don’t necessarily see that as a bad thing. Like, years of success and technology that’s supported that success that you need to figure out how to handle.

Corey: One last question that I have for you harkens back to something that I said earlier, where I congratulated you on your promotion to management. It’s not really a promotion, at least not the way that I think it should be thought about. Because it’s very much an orthogonal skill. You were a great engineer and architect building things yourself. And now you manage a team where if you’re diving into fix things by hand, you are misunderstanding the role in many respects, suddenly, your toolkit is no longer doing the thing yourself, but rather delegating the thing to be done and making sure that it gets done and your primary slash only toolkit to do all of that is hiring and developing talent. How have you negotiated that transition? Do you still find yourself itching to dive in and fix the work yourself? Are you better at letting go than I was for a long time? Where do you find yourself on that?

Tyler: Yeah, so that the inclination is still there, but I’ve learned to, like, recognize it and let it go. But I also have told my team members, like, 90% of the time, I’m going to give you all the latitude in the world, and I’m going to spend all my time helping you understand the problem that we’re facing as I understand it, and the potential roadblocks, and then there may be some times where I’m going to be like, “I really want it done this way.” And I ask them to give me that… give me that ability. I have yet to really break that one out. But that’s the only way that you can scale, and you get so much satisfaction about over… empowering someone to solve a hard challenge, and then seeing that they did it in a way different than you did it, and they did it better. [laugh].

And that’s a little bit of an ego hit, but you’re like, that’s what it’s about. And then they can build that confidence and then take on larger challenges. And that’s what gets me out of bed in the morning; that’s what gets me excited is working with people who just really want to do good work. And I can help put the right challenges in front of them, help shield them from stuff that’s not adding value, but like, asking for their time, connecting them with others that is going to kind of get that wind in their sails, and just get out of their way.

And then once the success is there, do everything I can to get that out and make sure that people know the good work that we’re doing. Because as much as you can say your work speaks for itself, in a huge organization, it’s not so much the case. Like, good work often goes unacknowledged if there’s not someone if you’re—like, promoting that. And most individuals aren’t comfortable—myself included—promoting my own work. Like, I wouldn’t do that, but I’m more than happy to promote the work of someone on my team.

Corey: On some level, as managers, you get recognized and evaluated based upon the performance of your team, not the things that you personally achieve. And that has always been a difficult transition. I got to level with you; I never handled it super well. It sounds like you are way better suited for the role than I ever was.

Tyler: Well, it’s early on, but yeah, I’m very excited.

Corey: If I really want to evaluate a manager, all I have to do is really talk to their team, more often than not, and you start to see things when you probe properly. I really want to thank you for taking so much time out of your day to speak with me. If people want to learn more about what you’re up to and how you see things, where can they find you?

Tyler: I’m probably most active on LinkedIn. So, just tylerslove at LinkedIn.

Corey: We’ll be sure to add that to both the [show notes 00:29:58], as well as I will add you to my professional network on LinkedIn, which I believe is the catchphrase that they’re using. Thanks so much for your time. I appreciate it.

Tyler: All right. Thanks, Corey.

Corey: Tyler Slove, Senior Manager for Enterprise Cloud and DevOps at United Airlines. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud, of the usual kind. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry comment disavowing all of this newfangled technology we’ve been talking about and that’s why you only travel via steamship.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Randall

Randall Hunt, VP of Cloud Strategy and Solutions at Caylent, is a technology leader, investor, and hands-on-keyboard coder based in Los Angeles, CA. Previously, Randall led software and developer relations teams at Facebook, SpaceX, AWS, MongoDB, and NASA. Randall spends most of his time listening to customers, building demos, writing blog posts, and mentoring junior engineers. Python and C++ are his favorite programming languages, but he begrudgingly admits that Javascript rules the world. Outside of work, Randall loves to read science fiction, advise startups, travel, and ski.

Links:

  • Caylent.com: https://caylent.com/
  • Twitter: https://twitter.com/jrhunt
  • Riot Games Talk: https://youtu.be/oGK-ojM7ZMc
  • James Hamilton Talk: https://youtu.be/uj7Ting6Ckk

TranscriptAnnouncer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Sysdig. Sysdig is the solution for securing DevOps. They have a blog post that went up recently about how an insecure AWS Lambda function could be used as a pivot point to get access into your environment. They’ve also gone deep in-depth with a bunch of other approaches to how DevOps and security are inextricably linked. To learn more, visit sysdig.com and tell them I sent you. That’s S-Y-S-D-I-G dot com. My thanks to them for their continued support of this ridiculous nonsense.

Corey: It seems like there is a new security breach every day. Are you confident that an old SSH key or a shared admin account isn’t going to come back and bite you? If not, check out Teleport. Teleport is the easiest, most secure way to access all of your infrastructure. The open source Teleport Access Plane consolidates everything you need for secure access to your Linux and Windows servers—and I assure you there is no third option there. Kubernetes clusters, databases, and internal applications like AWS Management Console, Yankins, GitLab, Grafana, Jupyter Notebooks, and more. Teleport’s unique approach is not only more secure, it also improves developer productivity. To learn more visit: goteleport.com. And no, that is not me telling you to go away, it is: goteleport.com.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. About a year ago, from the time of this recording, I had Randall Hunt on this podcast and we had a great conversation. He worked elsewhere, did different things, and midway through the recording, there was a riot slash coup attempt at the US Capitol. Yeah.

So, talking to Randall was the best thing that happened to me that day. And I’m hoping that this recording is a lot less eventful. Randall, thank you for joining me once again.

Randall: It’s great to see you, buddy. It’s been a long time.

Corey: It really has.

Randall: Well, I guess we saw each other at re:Invent.

Corey: We did, but that was—re:Invent as a separate, otherworldly place called Las Vegas. But since then, you’ve taken a new role. You are now the VP of Cloud Strategy and Solutions at Caylent. And the first reaction I had to that was, “What the hell is a Caylent? Let’s find out.”

So, I pulled up the website, and it was—you're an AWS partner, what I was able to figure out, but you didn’t lead with that, which is a great thing because, “We’re an AWS partner” is the least effective marketing strategy I can imagine. You are doing consulting on the implementation side the way that I would approach doing consulting implementation if I were down that path. Which I’m very much not, I’m pure advisory around one problem. But you talk about solutions, you talk about outcomes for your customers, you don’t try to be all things to all people. You’re Randall Hunt; you have a lot of options when it comes to what you do for careers. How did you wind up at Caylent?

Randall: Well, you know, I was doing a startup for a little while, and unfortunately, you know, I lost some people in my family. And I was just, like, a little mentally burnt out, so I took a break. And I had already bought my re:Invent ticket and everything. So, then I was like, “Okay, well, I’ll go to re:Invent; I’ll see everybody and try and avoid getting Covid.” So, I was masked up the whole time.

And while I was there, I ran into this group of folks who are on Caylent. And I did some research on them, and then we had some meetings. And I had already been kind of chatting with a bunch of different AWS partners, consulting partners, big and small. And none of them really stood out to me. I'm not trying to diss on any of these other partners because I think they’re all pretty amazing in what they do, but then a lot of them are just kind of… the same shop, pushing out the same code. They don’t have this operational excellence.

Corey: Swap one partner for another in many cases, and there’s not a lot of difference perceivable from the customer side of the story. And I know you’re going to be shocked by this, but I’m not a huge fan of the way that AWS talks about these things, with their messaging. Imagine that? Like, and sure enough in the partner program, AWS continues what it does with services, and gives things bad names. In this case, it’s a ‘competency.’ If we used to work together and someone reaches out for a reference check, and I say, “Randall? Oh, yeah. He was competent.”

Randall: [laugh].

Corey: That has a lot of implications that aren’t necessarily positive. It feels almost begrudging when they frame it that way. And it’s just odd.

Randall: The way that I look at it—so I don’t know if you’ve ever been through this program, but in order to achieve those competencies, you have to demonstrate. So, in order to be able to list them on your little partner card, right, or in the marketplace or whatever, you have to be able to go and say, “These are five customers where we delivered.” And then AWS will go and talk to those customers and ask for your satisfaction scores from those customers. You have to explain which services you use, what the initial set was, and then what the outcome was. And so it’s a big matrix that they make you fill out to accomplish each of these, and you have to have real-world customer examples.

So, I like that there’s that verification for people to know, but I don’t think that AWS does a great job of explaining what that means. Like, what goes into getting a competency. And I don’t know how to explain it quickly.

Corey: Same here. When I look at those partner cards on various websites—in some cases above the fold on the landing page—they list out all the different competencies, and it’s, on some level, if I know what all of those things are, and what they imply, and how that works, for a lot of problems I don’t need a partner at that point because at that point I’m deep enough in the weeds to do a lot of it myself. To be clear, I have the exact opposite outlier type that most companies probably should not emulate. One of our marketing approaches here at The Duckbill Group has been we are not AWS partners, as a selling point of all things. We’re not partnering with any company in the space, just due to real or perceived conflicts of interest.

We also do one very specific, very expensive problem in an advisory sense, and that is it. If we were doing implementation, and we lead with, “Oh, yeah. We’re not AWS partners,” it doesn’t go so well. I once was talking to somebody wanted me to do a security assessment there, and, “All right, it’s not what we do, but”—this was early days, and I gave the talk, and it turns out every talking point I’ve got for what works well in the costing space makes me look deranged when I’m talking about another space, it’s like, “Oh, yeah. We’re doing security stuff. Yeah, but we’re no AWS partners, and we’re not part of any vendor in this space.”

“That sounds actively dangerous and harmful. What the hell is the matter with you people?” Because security is a big space, and you need to work closely with cloud providers when doing security things there. The messaging doesn’t [laugh] land quite the same way. That’s why I don’t do other kinds of consulting these days.

Randall: Yeah. But to your question of what the hell is a Caylent?

Corey: [laugh].

Randall: So, Caylent’s name is derivative of Caylus, which is a God from Roman mythology. And I think it’s the root of the word celestial. But I just looked up the etymology, and I can’t confirm that. But let’s be real, you know—

Corey: Well, hang on a second because we look at Athena, which is AWS’s service named after the Greek Goddess of spending money on cloud services—

Randall: [laugh].

Corey: We have Kubernetes, which is the Greek God of cosplaying as a Google engineer. And I’m not a huge fan of either of those things, so why am I going to like Caylent any better?

Randall: Oh, it’s Roman, not Greek.

Corey: Ah, that would do it.

Randall: [laugh]. No, I—beyond meeting the team there, and then reading through some of their case studies and projects when I was at re:Invent, in my first day at the company, I just went around and I just spoke with as many of the engineers as I could. And I was blown away at some of the cool stuff that they’re working on and some of the talent. And here’s the thing is, Caylent was pretty small last year, you know? I think they were at 30 people sometime last year, and now they’re—it’s, you know, 400% growth, almost.

And they’ve done some really, really cool important work during Covid. For companies like eMed. They’ve done some work for, you know, all of these other firms. But between you and me—let’s get down to business, which is, you know I love space.

Corey: Oh, yeah. To be clear, when we’re talking about space, that can mean a bunch of different things. Like, “Honestly, don’t be near me,” could be how I interpret that.

Randall: You know how I love space, and rockets, and orbital mechanics, and satellites, and these sorts of things. And SciFi. And Caylent’s whole branding scheme is around this little guy called the [Caylien 00:07:28]. It’s our little mascot, a little alien dude, and it is kind of our whole branding persona. And everything else that we do is rockets. And we don’t have onboarding, we have launch plans. That whole branding, it seems silly but—

Corey: It’s very evocative, the Roman mythology, I think that’s a great direction to go in. I realized that for the start of this episode, I forgot to give folks who are not familiar with you a bit of backstory. You’ve done a lot of things: You worked at NASA, and then you were at MongoDB, and you were a boomerang at AWS—your second time there is where I wound up meeting you—in between, you decided to work at a little company called SpaceX. So, yeah, space is kind of a thing for you.

And then you were in a few different roles at AWS, and that’s where I encountered you. And you had a way of talking to people on stage, or in a variety of different contacts, and building up proofs of concept, where you made a lot of the technical hard things look easy without being condescending to anyone, in the event that the rest of us mere mortals found them a little trickier to do. You did a great job of not just talking about what the service did, but about what problem it solves, and thus by extension, why I should care. And it was really neat to watch you just break things down like that in a way that makes sense. Now that you’re over at Caylent as the VP of Cloud Strategy, the two things I see are, on the strength side, you have an ability to articulate the why behind what customers, and companies, and technologists are doing.

The caution I have, and I’m curious about how you’re challenging that is, your default goto explain things in many cases is to write some code that demonstrates the thing that you’re talking about. Great engineer; as a VP, depending on how that expresses itself, that could be something that poses a bit of a challenge. How do you view it?

Randall: You gave me some good advice on this. I don’t know if you remember, but you said, “Randall, if you’re in management, you got to make sure you’re not just an engineer with an inflated title.” You know, “You have to lead. Good leaders aren’t passive; they’re active.” And I kind of took that to heart.

I’m never going to stop coding, I’m never going to be hands-off keyboard, but one of the things that I’ve been focusing on lately, as opposed to doing pure implementation, is what is the Caylent culture and the Caylent way of doing things, and how can we onboard junior talent and get them to learn as much as they can about the cloud so we can cover the cost of their certification and things like that, but how do we make it so we’re not just teaching them the things they need to be successful in the role, but the things they need to be successful in their career, even as they leaves Caylent. You know, even beyond Caylent. When you’re hiring somebody, when you’re evaluating a cloud engineer, if they have Caylent on their resume, I want that to be a very strong signal for hiring managers where they’re like, “Oh, I know, Caylent does amazing work, so we’re going to definitely put this person in for an interview.” And then I’ve been an independent consultant many, many times. So, I’ve done work just off on the side, like, implementation and stuff for probably hundreds of companies over the last decade-plus, but what I haven’t done is really worked with a consulting firm before.

I have this interesting dilemma that I’m trying to evaluate right now, which is, you work with a very broad set of customers who have a very broad set of values and principles and ways of doing things. And you, as a consultant, are not able to just prescriptively come in and say, “This is how you should do it.” You know, we’re not McKinsey; we don’t come in and talk to the board and say, “You have to restructure the whole company.” That’s not what we do. What we do is we build things and we help with DevOps.

And so I’ve been playing around with this, so let me workshop it on you and you tell me what you think. It’s—

Corey: Hit me.

Randall: At Caylent, we work within the customer’s values, but we strive to be ambassadors of our Caylent culture. “Always be on the lookout for values, ideas, tools, and practices that our customers have that would work well here at Caylent. And these are our principles unless you know better ones.” I don’t know if you know that phrase, by the way. It’s an old Amazon thing.

Corey: Oh, yeah. I remember that quite a bit. It’s included in most of their tenet descriptions of, “These are ours unless you know better ones.” They don’t say that about the leadership principle because—

Randall: Right.

Corey: —it’s like, “These are leadership principles unless you know better ones.” Yes, several. But that’s beside the point. The idea of being able to—being about to always learn and the rest. You also hit on something that applies to my entire philosophy of employment.

Something we do in this industry is we tend to stay in jobs for, I don’t know, ideally, two to five years in most cases, and then we move on. But magically, during the interview process, we all pretend that this is your forever job, and suddenly, this is the place that’s going to change all of it, and you’re going to be here for 25 years and retire with a gold pocket watch and a pension. And most people don’t have either of those things in this century, so it’s a little bit of an unrealistic fantasy. Something I like to ask our candidates during the interview process is always, “Great. Ignore this job. Ignore it entirely. What’s the job after this one? Where are you going?”

Because if you don’t plan these things, your career becomes what happens to you instead. And even if what you plan changes, that’s great. It keeps you moving, from doing the same thing year after year after year after year. Early in my career, I worked with someone had been at the company for seven years, but it was time for him to go and he couldn’t for the life and remember what he did years two through four, which—

Randall: Yeah.

Corey: —you may as well not have been there.

Randall: There’s a really good quote from the CEO of GitLab that… says, “At GitLab, we hire people on trajectory, not on pedigree.” And I love that. And—you know, I never finished college, so the fact that I’ve been able to get the opportunities I’ve been able to get without a college degree, and without a fancy name on my resume—

Corey: We are exactly the same on that, but hang on a second; you have a lot of fancy names on your resume, so slow your roll there, Speed Racer.

Randall: Okay. Okay, well, [laugh] but that’s after, right? Like, I think once you land one, the rest don’t matter. But I—

Corey: I still never have. The most impressive thing on my resume is, honestly, The Duckbill Group.

Randall: Well, I think that’s pretty impressive now, right?

Corey: Oh, it is—

Randall: [laugh].

Corey: —we’re pretty good at what we do. But it doesn’t have the household recognition that you know, SpaceX does. Yet.

Randall: Yet. [laugh]. I’m really loving building things and working with customers, but you’re totally right. As you move into leadership, it’s not your job to write code day in and day out. I know a couple people. So, Elliot Horowitz, who used to be the CTO over at MongoDB, he would still code all the time. And I’d love to be able to find a way to keep my hands-on keyboard skills sharp, but continue to have the larger impact that you can have in leadership for a larger number of people.

Corey: I have the same problem because my consulting clients, it’s pure advisory. I don’t write production code for a variety of excellent reasons, including that I’m bad at it. And with managing the team here, as soon as I step in and start writing the code myself, in front of—instead of someone else whose core function it is, well, that causes a bunch of problems culturally as well as the problem of I’m suddenly in the critical path, and there’s probably something more impactful I could be and should be working on.

So, my answer, in all seriousness, has been shitposting. When I build ridiculous things that—you helped out architecturally with one of them: The stop.lying.cloud status page replacement for AWS.

Randall: Oh, yeah, that you were regenerating every time? I remember that.

Corey: Yeah. I wrote a whole blog post about that. Like, I have a Twitter client that I wrote the first version of, and then paid someone to make better: lasttweetinaws.com, that’s out there for a bunch of things.

My production pipeline for the newsletter. And the reason I build a lot of these things myself is that it keeps me touching the technology so I don’t become a talking head. But if I decide I don’t want to touch code this week, nothing is not happening for the business as a direct result of that. Plus, you know, it’s nice to have a small-scale environment that I can take screenshots of without worrying about it. And oh, heavens, I’m suddenly sharing data that shouldn’t be shared publicly. So, I find a way to still bring it in and tie it in without it being the core function of my role. That may help. It may not.

Randall: No, it does help. There’s this person in our industry, Charity Majors. I've been reading some of her blog posts about engineering management and how that all kind of shakes out, and I’ve tried to take as much lessons from that as I can. Because, right, you know, being in leadership is fairly new for me, I don’t know if I’m good at this, I might suck at it. And by the way, if Cayliens are listening and you see me screw up, just shoot me a message on Slack, anytime, day or night. It’s like, “Hey, Randall, you screwed this up.” Just let me know because—

Corey: Or call it out on Twitter; that’s more entertaining. I kid. I kid. That’s what’s known as a career-limiting move in most places. Not because Randall’s going to take any objection to it, but because it’s—people can see the things that you write, and it’s one of those, “Oh, you’re just going to call down your own internal company leadership in public?” Even if it’s a gag or something people don’t have the context on that. It does not look good to folks who lack the context. I’ve learned as I’ve iterated forward that appearances count for an awful lot on things like that. I’m sorry, please continue.

Randall: And the other thing that I’ve discovered is that you can have an outsized impact by focusing on education within your own company. So, one of my primary functions is to just stay on top of AWS news. So—

Corey: Yeah. Me too.

Randall: Exactly, right? So, literally every RSS feed from AWS, I watched every single re:Invent video. So it’s, like, 19 days' worth of video. And obviously, you know, I put it on double-speed, and I would skip through a bunch of things. But I go, and I review everything, and I try and create context with the people who are moving and shaking things at AWS and building cool stuff.

And my realization is that I need to work to grow my network and connect with people who have accomplished very impressive things in business. And by leveraging that network and learning about the challenges they faced, it becomes a compression algorithm for experience. And I know that’s an uncommon, unpopular opinion, that most people will say there is no compression algorithm for experience, but I think taking lessons learned and leveraging them within your own organization is probably one of the most important things you can do.

Corey: I would agree with you, but I also going to take it a step further. “There’s no compression algorithm for experience.” It sounds pithy, but it’s one of the most moronic things I’ve heard in recent memory because of course there is. We all stand—

Randall: It’s called machine learning. [laugh].

Corey: —on the shoulders of giants. We can hire consultancies, you can hire staff who have solved similar problems before, you can buy a product that bakes all of that experience into it. And, yeah, you can absolutely find ways of compressing experience. I feel like anytime a big cloud company that charges per gigabyte tells you that there’s no compression algorithm for anything, it’s because, “Ah, I see what’s going on here. You’re trying to basically gouge customers. Got it.”

Randall: I want to come back to that in one second, right, because I do want to talk about cloud networking because I have so many thoughts on this, and AWS did some cool stuff. But there’s one other thing that I’ve been thinking about a lot lately, and one of the hardest things that I found in business is to not slow down as your organization grows. It becomes really easy to introduce excuses for going slower or to introduce processes that create bottlenecks. And my whole focus right now is—Caylent’s in this hyper-growth period: We’re hiring a lot, we’re growing a lot, we have so many inbound customers that we want to be able to build cool stuff for. And help them out with their DevOps culture, and help them get moved into the 21st century, right?

How do we grow without just completely becoming bureaucratic, you know? I want people to be a manager of one and be able to be autonomous and feel empowered to go and do things on behalf of customers, but you also have to focus on security and compliance and the checkboxes that your customers want you to have and that your customers need to be able to trust you. And so I’m really looking for good ideas on how to, like, not slow down as we grow.

Corey: Today’s episode is brought to you in part by our friends at MinIO the high-performance Kubernetes native object store that’s built for the multi-cloud, creating a consistent data storage layer for your public cloud instances, your private cloud instances, and even your edge instances, depending upon what the heck you’re defining those as, which depends probably on where you work. It’s getting that unified is one of the greatest challenges facing developers and architects today. It requires S3 compatibility, enterprise-grade security and resiliency, the speed to run any workload, and the footprint to run anywhere, and that’s exactly what MinIO offers. With superb read speeds in excess of 360 gigs and 100 megabyte binary that doesn’t eat all the data you’ve gotten on the system, it’s exactly what you’ve been looking for. Check it out today at min.io/download, and see for yourself. That’s min.io/download, and be sure to tell them that I sent you.

Corey: That’s always an interesting challenge because slowing down is an inherent… side effect of maturity, on some level, and people look, “Well, look at AWS. They do all kinds of super quickly.” Yeah, they release new things from small teams very quickly, but look at the pace of change that comes to foundational services like SQS or S3, like the things that are foundational to all of that? And yeah, you don’t want to iterate on that super quickly and change constantly because people depend on the behaviors on the, in some cases, the bugs, and any change you make is going to disrupt someone’s workflow. So, there’s always a bit of a balance there.

I want to talk specifically about how you view AWS because people ask me the same thing all the time, and you stand in a somewhat similar position. You worked there, I never have, but you have been critical of things that AWS has done, rightfully so. I very rarely find myself disagreeing with you. You’re also a huge fan of things that they do, which I am as well. And I want to be very clear for anyone who questions this, you work for a large partner now, and there are always going to be constraints, real or imagined, around what you can say about a company with whom a good portion of your business flows through.

But I have never once known you to shill for something you don’t believe in. I think your position on this is the same as mine, which is—

Randall: A hundred percent.

Corey: I don’t need to say every thought that flits through my head about something, but I will not lie to my audience—or to other people, or my customers, or anyone else for that matter—about something, regardless of what people want me to do. I’ve turned down sponsorships on that basis. You can buy my attention, but not my opinion, and I’ve always got a very strong sense of that same behavior from you.

Randall: You’re totally right there. I mean—

Corey: [unintelligible 00:21:39] disagree with that. Like, “No, no, I’m a hell of a shill. What are you—thanks for not seeing it though.” Come on, of course you’re going to agree with that.

Randall: So, when I was at AWS, I did have to shill a little bit because they have some pretty intense PR guidelines. But—

Corey: Rule number one: Never say anything at any time proactively. But okay. Please continue.

Randall: No, no, I think they’ve relaxed it over the years. Because—so Amazon had very strict PR, and then when AWS was kind of coming up, like, a lot of those PR rules were kind of copy-pasted into AWS. And it took a while for the culture of AWS, which is very much engineering-focused, to filter up into PR. So, I think modern-day AWS PR is actually a lot more relaxed than it was, say in, like, 2014. And that’s how we have Senior Principal and distinguished engineers on Twitter who are able to share really cool details about services with us.

And I love that. You know, Colm’s threads are great to read. And then, you know, there are a bunch of people that I follow, who all have cool details and deep-dives into things, Matt Wilson as well. And so when you talk about being authentic and not just reiterating the information that comes from AWS, I have this balance that I have to play that I was honestly not good at earlier on in my career—maybe it was just a maturity thing—where I would say every thought that came through my head. I wouldn’t take a beat and think about, you know, how can I say this in a way that’s actionable for the team, as opposed to just pure criticism?

And now, I am fully committed to being as authentic as possible. So, when a service stinks, I say it. I am very much down on Timestream right now. For what it’s worth, I have not tried it this month, but you know, I keep trying to use Timestream, right, and I keep running into issues. That these lifecycle policies, they don’t actually move things in the timely manner that you expect them to.

And, you know, there’s this idea that AWS has around purpose-built databases and they’re trying to shove all of these different workloads into different databases, but a lot of times—you know, DynamoDB can be your core data processing engine, and everything else can flow from that. Or you can even use MongoDB. But throwing in Timestream and MemoryDB and all these other things on top of it, it becomes less and less differentiated. And a lot of these workloads are getting served by other native services, like cloud-native services.

And anyway, that’s a whole tangent, but basically, I wanted to say, you can expect me to continue to be very opinionated about AWS services, and I think that’s one of the reasons that customers want us there is we will advise you on the full spectrum of compute, right? We’re not going to say, “Oh, you have to go serverless.” There are still some workloads that are not well served by serverless. There’s still some stuff that just doesn’t work well with serverless. And then there’ll be other workloads where EKS is where you want to be running things, you know? Maybe you do need Kubernetes.

I used to go on Twitter all the time, and I would say things like, “You don’t need Kubernetes. You don’t need Kubernetes. Like, you only need Kubernetes at this scale. Like, you’re not there yet. Calm down.” That’s changed. So, these days, I think Kubernetes is way easier to deal with and it’s a lot more mature, so I don’t shy away from recommending Kubernetes these days.

Corey: What is your take on, I guess, some of the more interesting global infrastructure stuff that they’re doing lately because I’ve been having some challenges, on some level, building some multi-region stuff, and increasingly, it’s felt to me like a lot of the region expansions and the rest have been for very specific folks, in very specific places, with very specific—often regulatory—constraints. These aren’t designed to the point where anyone would want to use more than two or three in any applications deployment. And I know this because when I try to do it, the [SAs 00:25:13] look at me like I’m something of a loon.

Randall: So, there were two really cool launches from AWS, this year at re—or last year at re:Invent. There was Cloud WAN or Cloud Wide Area Network, and there was SiteLink, and there was also VPC Access Analyzer. But when we talk about AWS’s global infrastructure, I like going back to James Hamilton’s talk. I don’t remember if it’s 2017 or 2018, but it was, “Tuesday Night Live” with AWS or something, and it walked through what a region is. And so the AWS Cloud these days is 26 regions, there are eight more on the way, and then there’s something like 30 local zones.

And I think that AWS is focused on getting closer to their customers, creating better peering relationships with different telecom providers, creating more edge locations, creating more regional caches, is transformative for what can be delivered. I play video games, so Riot Games gave a cool talk at re:Invent about how they use a mix of Outposts, and edge locations, and local zones to be able to get their Valorant gamers on to—Valorant is this first-person shooter game—get those gamers on the most local server that minimizes latency and pain for them. And that’s the kind of future that I want to see us build towards, and that’s something—I’m still incredibly bullish on AWS. I know Azure and Google are making improvements, and great for them for doing that because it raises all of us up to compete, but the thing that AWS has done that separates them from a lot of the other clouds is they have enabled workloads that literally would not have been possible without fundamental investment in global infrastructure. I’m talking things like undersea cables, I’m talking things like net-new applied photonics for fiber: There’s researchers at AWS whose sole job is to figure out how to fit more stuff into fiber.

So, James Hamilton did this talk, right, and he broke down what an availability zone is—and there are 84 availability zones in all now—and he walked through an availability zone is not a single data center; an availability zone typically comprises multiple data centers that are separated from each other with different infrastructure and stuff. And then he broke down, like, the largest AWS availability zone is 14 data centers. And all new regions, by the way, have three availability zones, and those availability zones, they’re meaningfully separated, more than a mile but less than an issue than when, like, speed of light effects come in. And that’s where you can build services like Aurora, where you have this shared storage layer on top of a data engine. And that’s how you can build FSx for Lustre, and EBS, and EFS.

And, like, all of these services are things that are really only possible at scale. And Peter DeSantis talked a lot about this in his keynote, by the way, about the advantages of aggregate workload monitoring. I think AWS’s ability to innovate from first principles is probably unparalleled in our global economy right now. That’s not to say they will always be there, and that’s not to say that they’re always going to be that level of innovation, but for the last ten years, they’ve shown again and again that they can just go gangbusters and release new stuff. I mean, we have 400-gigabit-per-second networking now. Like, what the heck?

Corey: And we still charge two cents per gigabyte when we throw that amount of capacity from one availability zone in a region to another. Which, of course I’m still salty about. Remember, my role is economics, so I have a different perspective on these things.

Randall: Well, I like that Cloudflare and Google and Microsoft—and even Oracle, by the way; I don’t know—at some point we should talk about Oracle Cloud because I used to be really down on them, but now that I’ve played around with it more, they’re like coming up, you know? They’re getting better and better.

Corey: I am very impressed by a lot of stuff that Oracle Cloud is doing. With the disclaimer that they periodically sponsor this podcast. I think they’re still doing that. That’s the fun thing is that I have an editorial firewall. But I’m not saying this because they’re paying me to say this; I’m saying it because I experimented with it.

I was really looking forward to just crapping all over it. And it was good. And… “Who is this really? Like, did someone just slapping Oracle sticker on something pleas”—no. It’s actually nice. But yeah, we should dive into that at some point.

Randall: I want to say one more thing on global infrastructure, and I know we don’t have a lot of time left, but even 800-gigabit-per-second networking on the Trainium instances, by the way now. Which is just mind-blowing.

So, the fact that AWS has redone two-inch conduits—and I have this picture that I took at re:Invent that I can share with you later, if you want—of all their different fiber and, like, networking and switches and stuff. In aggregate, one of their regions has 5000 terabits of capacity. 5000 terabits. It’s 388 unique fiber paths. It’s just—it’s absolutely fascinating, and it’s a scale that enables the modern economy and the modern world.

Like the app we’re using to record this podcast, all of these things rely on AWS global infrastructure backbone, and that’s why I think they charge what they charge for, you know, these networking services. They’re recouping the cost of that fundamental investment. But now, last year they announced 100 gigabytes free for S3 and non-CloudFront services, and then one terabyte per month for free from CloudFront. So, that’s a huge improvement. It’s a little late, but I mean, they got it done.

Corey: I do want to the point of transparency and honesty, the app that we’re using to record this does, in fact, use Google Cloud. But again—

Randall: Oh.

Corey: —it’s—yeah, again, it’s one of the big ones, regardless. You can always tell which one is it, and not, “No, I’m running this myself on a Raspberry Pi.” Yeah. There’s a lot that goes into these things. Honestly, I think the big winners in all this are those of us who are building things on top of these technologies—

Randall: Yes.

Corey: —because I can just build the ridiculous thing I want to and deploy it worldwide without signing $20 million of contracts first.

Randall: Yeah. And going back to your point about multi-region stuff, I think that’s getting better and better over time. There’s some missteps. So like, let’s take DynamoDB global tables, for instance—

Corey: Which is not in every region, so it’s basically this point, hemispherical tables.

Randall: Well, even so, it’s good enough, right? Like, it gives you the controls that you need to be able to slide that shared responsibility model and that shared cost model in the way that you need to. Or shared availability model. What is frustrating though, is that while this global availability is getting better and better from a software perspective, it’s getting harder and harder from a code perspective. So, actually writing the code to take advantage of some of this global infrastructure is imperfect. And Forrest Brazeal, from Google Cloud, he spoke a little bit about this recently, and we had a cool Twitter discussion.

Corey: Fantastic. I’m a big fan of Forrest. I’m glad that he found a place to land. I’m sad that it’s not in the AWS ecosystem, but here we are.

Randall: I mean, I’ll follow that man anywhere. He’s the Tom [Lehrer 00:31:48] of cloud. Just glad he’s still around to keep making some cool stuff.

Corey: I don’t want to know what I am of cloud, ever. Don’t tell me. Talk about it amongst yourselves, but don’t tell me. Randall, I want to thank you for taking the time to speak with me. It is always a pleasure. If people want to learn more about what you’re up to, where can they find you these days?

Randall: caylent.com. I’m going to be writing a bunch of AWS blog posts on there, so go there. Also go to Twitter, @jrhunt on Twitter.

And if you need help building your cloud-native apps and some DevOps consulting, or just a general 30-minute phone call to understand what you should do, reach out to me; reach out to Caylent. We’re happy to help. We love taking these conversations and learning what you’re building.

Corey: And we will, of course, put links to that in the [show notes 00:32:30]. Thank you so much for taking the time to speak with me. I appreciate it.

Randall: Thank you for having me on. It’s great to see you.

Corey: Until the next time. Randall Hunt, VP of Cloud Strategy and Solutions at Caylent. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice and an angry comment complaining about the differences between Greek and Roman mythology, and the best mythology is the stuff you have on your website about how easy it is to use your company, which is called Corporate Mythology.

Randall: I love it.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Jason

Jason Frazier is a Software Engineering Manager at Ekata, a Mastercard Company. Jason’s team is responsible for developing and maintaining Ekata’s product APIs. Previously, as a developer, Jason led the investigation and migration of Ekata’s Identity Graph from AWS Elasticache to Redis Enterprise Redis on Flash, which brought an average savings of $300,000/yr.

Links:

  • Ekata: https://ekata.com/
  • Email: jason.frazier@ekata.com
  • LinkedIn: https://www.linkedin.com/in/jasonfrazier56

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Sysdig. Sysdig is the solution for securing DevOps. They have a blog post that went up recently about how an insecure AWS Lambda function could be used as a pivot point to get access into your environment. They’ve also gone deep in-depth with a bunch of other approaches to how DevOps and security are inextricably linked. To learn more, visit sysdig.com and tell them I sent you. That’s S-Y-S-D-I-G dot com. My thanks to them for their continued support of this ridiculous nonsense.

Corey: Today’s episode is brought to you in part by our friends at MinIO the high-performance Kubernetes native object store that’s built for the multi-cloud, creating a consistent data storage layer for your public cloud instances, your private cloud instances, and even your edge instances, depending upon what the heck you’re defining those as, which depends probably on where you work. It’s getting that unified is one of the greatest challenges facing developers and architects today. It requires S3 compatibility, enterprise-grade security and resiliency, the speed to run any workload, and the footprint to run anywhere, and that’s exactly what MinIO offers. With superb read speeds in excess of 360 gigs and 100 megabyte binary that doesn’t eat all the data you’ve gotten on the system, it’s exactly what you’ve been looking for. Check it out today at min.io/download, and see for yourself. That’s min.io/download, and be sure to tell them that I sent you.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. This one is a bit fun because it’s a promoted episode sponsored by our friends at Redis, but my guest does not work at Redis, nor has he ever. Jason Frazier is a Software Engineering Manager at Ekata, a Mastercard company, which I feel, like, that should have some sort of, like, music backstopping into it just because, you know, large companies always have that magic sheen on it. Jason, thank you for taking the time to speak with me today.

Jason: Yeah. Thanks for inviting me. Happy to be here.

Corey: So, other than the obvious assumption, based upon the fact that Redis is kind enough to be sponsoring this episode, I’m going to assume that you’re a Redis customer at this point. But I’m sure we’ll get there. Before we do, what is Ekata? What do you folks do?

Jason: So, the whole idea behind Ekata is—I mean, if you go to our website, our mission statement is, “We want to be the global leader in online identity verification.” What that really means is, in more increasingly digital world, when anyone can put anything they want into any text field they want, especially when purchasing anything online—

Corey: You really think people do that? Just go on the internet and tell lies?

Jason: I know. It’s shocking to think that someone could lie about who they are online. But that’s sort of what we’re trying to solve specifically in the payment space. Like, I want to buy a new pair of shoes online, and I enter in some information. Am I really the person that I say I am when I’m trying to buy those shoes? To prevent fraudulent transactions. That’s really one of the basis that our company goes on is trying to reduce fraud globally.

Corey: That’s fascinating just from the perspective of you take a look at cloud vendors at the space that I tend to hang out with, and a lot of their identity verification of, is this person who they claim to be, in fact, is put back onto the payment providers. Take Oracle Cloud, which I periodically beat up but also really enjoy aspects of their platform on, where you get to their always free tier, you have to provide a credit card. Now, they’ll never charge you anything until you affirmatively upgrade the account, but—“So, what do you do need my card for?” “Ah, identity and fraud verification.” So, it feels like the way that everyone else handles this is, “Ah, we’ll make it the payment networks’ problem.” Well, you’re now owned by Mastercard, so I sort of assume you are what the payment networks, in turn, use to solve that problem.

Jason: Yeah, so basically, one of our flagship products and things that we return is sort of like a score, from 0 to 400, on how confident we are that this person is who they are. And it’s really about helping merchants help determine whether they should either approve, or deny, or forward on a transaction to, like, a manual review agent. As well as there’s also another use case that’s even more popular, which is just, like, account creation. As you can imagine, there’s lots of bots on everyone’s [laugh] favorite app or website and things like that, or customers offer a promotion, like, “Sign up and get $10.”

Well, I could probably get $10,000 if I make a thousand random accounts, and then I’ll sign up with them. But, like, make sure that those accounts are legitimate accounts, that’ll prevent, like, that sort of promo abuse and things like that. So, it’s also not just transactions. It’s also, like, account openings and stuff, make sure that you actually have real people on your platform.

Corey: The thing that always annoyed me was the way that companies decide, oh, we’re going to go ahead and solve that problem with a CAPTCHA on it. It’s, “No, no, I don’t want to solve machine learning puzzles for Google for free in order to sign up for something. I am the customer here; you’re getting it wrong somewhere.” So, I assume, given the fact that I buy an awful lot of stuff online, but I don’t recall ever seeing anything branded with Ekata that you do this behind the scenes; it is not something that requires human interaction, by which I mean, friction.

Jason: Yeah, for sure. Yeah, yeah. It’s behind the scenes. That’s exactly what I was about to segue to is friction, is trying to provide a frictionless experience for users. In the US, it’s not as common, but when you go into Europe or anything like that, it’s fairly common to get confirmations on transactions and things like that.

You may have to, I don’t know text—or get a code text or enter that online to basically say, like, “Yes, I actually received this.” But, like, helping—and the reason companies do that is for that, like, extra bit of security and assurance that that’s actually legitimate. And obviously, companies would like to prefer not to have to do that because, I don’t know, if I’m trying to buy something, this website makes me do something extra, the site doesn’t make me do anything extra, I’m probably going to go with that one because it’s just more convenient for me because there’s less friction there.

Corey: You’re obviously limited in how much you can say about this, just because it’s here’s a list of all the things we care about means that great, you’ve given me a roadmap, too, of things to wind up looking at. But you have an example or two of the sort of the data that
you wind up analyzing to figure out the likelihood that I’m a human versus a robot.

Jason: Yeah, for sure. I mean, it’s fairly common across most payment forms. So, things like you enter in your first name, your last name, your address, your phone number, your email address. Those are all identity elements that we look at. We have two data stores: We have our Identity Graph and our Identity Network.

The Identity Graph is what you would probably think of it, if you think of a web of a person and their identity, like, you have a name that’s linked to a telephone, and that name is also linked to an address. But that address used to have previous people living there, so on and so forth. So, the various what we just call identity elements are the various things we look at. It’s fairly common on any payment form, I’m sure, like, if you buy something on Amazon versus eBay or whatever, you’re probably going to be asked, what’s your name? What’s your address? What’s your email address? What’s your telephone?

Corey: It’s one of the most obnoxious parts of buying things online from websites I haven’t been to before. It’s one of the genius ideas behind Apple Pay and the other centralized payment systems. Oh, yeah. They already know who you are. Just click the button, it’s done.

Jason: Yeah, even something as small as that. I mean, it gets a little bit easier with, like, form autocompletes and stuff like, oh, just type J and it’ll just autocomplete everything for me. That’s not the worst of the world, but it is still some amount of annoyance and friction. [laugh].

Corey: So, as I look through all this, it seems like one of the key things you’re trying to do since it’s in line with someone waiting while something is spinning in their browser, that this needs to be quick. It also strikes me that this is likely not something that you’re going to hit the same people trying to identify all the time—if so, that is its own sign of fraud—so it doesn’t really seem like something can be heavily cached. Yet you’re using Redis, which tells me that your conception of how you’re using it might be different than the mental space that I put Redis into what I’m thinking about where this ridiculous architecture diagram is the Redis part going to go?

Jason: Yeah, I mean, like, whenever anyone says Redis, thinks of Redis, I mean, even before we went down this path, you always think of, oh, I need a cache, I’ll just stuff in Redis. Just use Redis as a cache here and there. I don’t know, some small—I don’t know, a few tens, hundreds gigabytes, maybe—cache, spin that up, and you’re good. But we actually use Redis as our primary data store for our Identity Graph, specifically for the speed that we can get. Because if you’re trying to look for a person, like, let’s say you’re buying something for your brother, how do we know if that’s true or not? Because you have this name, you’re trying to send it to a different address, like, how does that make sense? But how do we get from Corey to an address? Like, oh, maybe used to live with your brother?

Corey: It’s funny, you pick that as your example; my brother just moved to Dublin, so it’s the whole problem of how do I get this from me to someone, different country, different names, et cetera? And yeah, how do you wind up mapping that to figure out the likelihood that it is either credit card fraud, or somebody actually trying to be, you know, a decent brother for once in my life?

Jason: [laugh]. So, I mean, how it works is how you imagine you start at some entry point, which would probably be your name, start there and say, “Can we match this to this person’s address that you believe you’re sending to?” And we can say, “Oh, you have a person-person relationship, like he’s your brother.” So, it maps to him, which we can then get his address and say, “Oh, here’s that address. That matches what you’re trying to send it to. Hey, this makes sense because you have a legitimate reason to be sending something there. You’re not just sending it to some random address out in the middle of nowhere, for no reason.”

Corey: Or the drop-shipping scams, or brushing scams, or one of—that’s the thing is every time you think you’ve seen it all, all you have to do is look at fraud. That’s where the real innovation seems to be happening, [laugh] no matter how you slice it.

Jason: Yeah, it’s quite an interesting space. I always like to say it’s one of those things where if you had the human element in it, it’s not super easy, but it’s like, generally easy to tell, like, okay, that makes sense, or, oh, no, that’s just complete garbage. But trying to do it at scale very fast in, like, a general case becomes an actual substantially harder problem. [laugh]. It’s one of those things that people can probably do fairly well—I mean, that’s why we still have manual reviews and things like that—but trying to do it automatically or just with computers is much more difficult. [laugh].

Corey: Yeah, “Hee hee, I scammed a company out of 20 bucks is not the problem you’re trying to avoid for.” It’s the, “Okay, I just did that ten million times and now we have a different problem.”

Jason: Yeah, exactly. I mean, one of the biggest losses for a lot of companies is, like, fraudulent transactions and chargebacks. Usually, in the case on, like, e-commerce companies—or even especially like nowadays where, as you can imagine, more people are moving to a more online world and doing shopping online and things like that, so as more people move to online shopping, some companies are always going to get some amount of chargebacks on fraudulent transactions. But when it happens at scale, that’s when you start seeing many losses because not only are you issuing a chargeback, you probably sent out some products, that you’re now out some physical product as well. So, it’s almost kind of like a double-whammy. [laugh].

Corey: So, as I look through all this, I tended to always view Redis in terms of, more or less, a key-value store. Is that still accurate? Is that how you wind up working with it? Or has it evolved significantly past them to the point where you can now do relational queries against it?

Jason: Yeah, so we do use Redis as a key-value store because, like, Redis is just a traditional key-value store, very fast lookups. When we first started building out Identity Graph, as you can imagine, you’re trying to model people to telephones to addresses; your first thought is, “Hey, this sounds a whole lot like a graph.” That’s sort of what we did quite a few years ago is, let’s just put it in some graph database. But as time went on and as it became much more important to have lower and lower latency, we really started thinking about, like, we don’t really need all the nice and shiny things that, like, a graph database or some sort of graph technology really offers you. All we really need to do is I need to get from point A to point B, and that’s it.

Corey: Yeah, [unintelligible 00:10:35] graph database, what’s the first thing I need to do? Well, spend six weeks in school trying to figure out exactly what the hell of graph database is because they’re challenging to wrap your head around at the best of times. Then it just always seemed overpowered for a lot of—I don’t want to say simple use cases; what you’re doing is not simple, but it doesn’t seem to be leveraging the higher-order advantages that graph database tends to offer.

Jason: Yeah, it added a lot of complexity in the system, and [laugh] me and one of our senior principal engineers who’s been here for a long time, we always have a joke: If you search our GitHub repository for… we’ll say kindly-worded commit messages, you can see a very large correlation of those types of commit messages to all the commits to try and use a graph database from multiple years ago. It was not fun to work with, just added too much complexity, and we just didn’t need all that shiny stuff. So, that’s how we really just took a step back. Like, we really need to do it this way. We ended up effectively flattening the entire graph into an adjacency list.

So, a key is basically some UUID to an entity. So, Corey, you’d have some UUID associated with you and the value would be whatever your information would be, as well as other UUIDs to links to the other entities. So, from that first retrieval, I can now unpack it, and, “Oh, now I have a whole bunch of other UUIDs I can then query on to get that information, which will then have more IDs associated with it,” is more or less sort of how we do our graph traversal and query this in our graph queries.

Corey: One of the fun things about doing this sort of interview dance on the podcast as long as I have is you start to pick up what people are saying by virtue of what they don’t say. Earlier, you wound up mentioning that we often use Redis for things like tens, or hundreds of gigabytes, which sort of leaves in my mind the strong implication that you’re talking about something significantly larger than that. Can you disclose the scale of data we’re talking about her?

Jason: Yeah. So, we use Redis as our primary data store for our Identity Graph, and also for—soon to be for our Identity Network, which is our other database. But specifically for our Identity Graph, scale we’re talking about, we do have some compression added on there, but
if you say uncompressed, it’s about 12 terabytes of data that’s compressed, with replication into about four.

Corey: That’s a relatively decent compression factor, given that I imagine we’re not talking about huge datasets.

Jason: Yeah, so this is actually basically driven directly by cost: If you need to store less data, then you need less memory, therefore, you need to pay for less.

Corey: So, our users once again have shored up my longtime argument that when it comes to cloud, cost and architecture are in fact the same thing. Please, continue by all means.

Jason: I would be lying if I said that we didn’t do weekly slash monthly reviews of costs. Where are we spending costs in AWS? How can
we improve costs? How can we cut down on costs? How can you store less—

Corey: You are singing my song.

Jason: It is a [laugh] it is a constant discussion. But yeah, so we use Zstandard compression, which was developed at Facebook, and it’s a dictionary-based compression. And the reason we went for this is—I mean like if I say I want to compress, like, a Word document down, like, you can get very, very, very high level of compression. It exists. It’s not that interesting, everyone does it all the time.

But with this we’re talking about—so in that, basically, four or so terabytes of compressed data that we have, it’s something around four to four-and-a-half billion keys and values, and so in that we’re talking about each key-value only really having anywhere between 50 and 100 bytes. So, we’re not compressing very large pieces of information. We’re compressing very small 50 to 100 byte JSON values that we have give UUID keys and JSON strings stored as values. So, we’re compressing these 50 to 100 byte JSON strings with around 70, 80% compression. I mean, that’s using Zstandard with a custom dictionary, which probably gave us the biggest cost savings of all, if you can [unintelligible 00:14:32] your dataset size by 60, 70%, that’s huge. [laugh].

Corey: Did you start off doing this on top of Redis, or was this an evolution that eventually got you there?

Jason: It was an evolution over time. We were formally Whitepages. I mean, Whitepages started back in the late-90s. It really just started off as a—we just—

Corey: You were a very early adopter of Redis [laugh]. Yeah, at that point, like, “We got a time machine and started using it before it existed.” Always a fun story. Recruiters seem to want that all the time.

Jason: Yeah. So, when we first started, I mean, we didn’t have that much data. It was basically just one provider that gave us some amount of data, so it was kind of just a—we just need to start something quick, get something going. And so, I mean, we just did what most people do just do the simplest thing: Just stuff it all in a Postgres database and call it good. Yeah, it was slow, but hey, it was back a long time ago, people were kind of okay with a little bit—

Corey: The world moved a bit slower back then.

Jason: Everything was a bit slower, no one really minded too much, the scale wasn’t that large. But business requirements always change over time and they evolve, and so to meet those ever-evolving business requirements, we move from Postgres, and where a lot of the fun commit messages that I mentioned earlier can be found is when we started working with Cassandra and Titan. That was before my time before I had started, but from what I understand, that was a very fun time. But then from there, that’s when we really kind of just took a step back and just said, like, “There’s so much stuff that we just don’t need here. Let’s really think about this, and let’s try to optimize a bit more.”

Like, we know our use case, why not optimize for our use case? And that’s how we ended up with the flattened graph storage stuffing into Redis. Because everyone thought of Redis as a cache, but everyone also knows that—why is it a cache? Because it’s fast. [laugh]. We need something that’s very fast.

Corey: I still conceptualize it as an in-memory data store, just because when I turned on disk persistence model back in 2011, give or take, it suddenly started slamming the entire data store to a halt for about three seconds every time it did it. It was, “What’s this piece of crap here?” And it was, “Oh, yeah. Turns out there was a regression on Zen, which is what AWS is used as a hypervisor back then.” And, “Oh, yeah.”

So, fork became an expensive call, it took forever to wind up running. So oh, the obvious lesson we take from this is, oh, yeah, Redis is not designed to be used with disk persistence. Wrong lesson to take from the behavior, but did cement, in my mind at least, the idea that this is something that we tend to use only as an in-memory store. It’s clear that the technology has evolved, and in fact, I’m super glad that Redis threw you my direction to talk to you about this stuff because until talking to you, I was still—I got to admit—sort of in the position of thinking of it still as an in-memory data store because the fact that Redis says otherwise because they’re envisioning it being something else, well okay, marketers going to market. You’re a customer; it’s a lot harder for me to talk smack about your approach to this thing, when I see you doing it for, let’s be serious here, what is a very important use case. If identity verification starts failing open and everyone claims to be who they say they are, that’s something is visible from orbit when it comes to the macroeconomic effect.

Jason: Yeah, exactly. It’s actually funny because before we move to primarily just using Redis, before going to fully Redis, we did still use Redis. But we used ElastiCache, we had it loaded into ElastiCache, but we also had it loaded into DynamoDB as sort of a, I don’t want this to fail because we weren’t comfortable with actually using Redis as a primary database. So, we used to use ElastiCache with a fallback to DynamoDB, just in that off chance, which, you know, sometimes it happens, sometimes it didn’t. But that’s when we basically just went searching for new technologies, and that’s actually how we landed on Redis on Flash, which is a kind of breaks the whole idea of Redis as an in-memory database to where it’s Redis, but it’s not just an in-memory database, you also have flashback storage.

Corey: So, you’ll forgive me if I combine my day job with this side project of mine, where I fixed the horrifying AWS bills for large companies. My bias, as a result, is to look at infrastructure environments primarily through the lens of AWS bill. And oh, great, go ahead and use an enterprise offering that someone else runs because, sure, it might cost more money, but it’s not showing up on the AWS bill, therefore, my job is done. Yeah, it turns out that doesn’t actually work or the answer to every AWS billing problem is to migrate to Azure to GCP. Turns out that doesn’t actually solve the problem that you would expect.

But you’re obviously an enterprise customer of Redis. Does that data live in your AWS account? Is it something using as their managed service and throwing over the wall so it shows up as data transfer on your side? How is that implemented? I know they’ve got a few different models.

Jason: There’s a couple of aspects onto how we’re actually bill. I mean, so like, when you have ElastiCache, you’re just billed for your, I don’t know, whatever nodes using, cache dot, like, r5 or whatever they are… [unintelligible 00:19:12]

Corey: I wish most people were using things that modern. But please, continue.

Jason: But yeah, so you basically just build for whatever last cache nodes you have, you have your hourly rate, I don’t know, maybe you might reserve them. But with Redis Enterprise, the way that we’re billed is there’s two aspects. One is, well, the contract that we signed
that basically allows us to use their technology [unintelligible 00:19:31] with a managed service, a managed solution. So, there’s some amount that we pay them directly within some contract, as well as the actual nodes themselves that exist in the cluster. And so basically the way that this is set up, is we effectively have a sub-account within our AWS account that Redis Labs has—or not Redis Labs; Redis Enterprise—has access to, which they deploy directly into, and effectively using VPC peering; that’s how we allow our applications to talk directly to it.

So, we’re built directly—or so the actual nodes of the cluster, which are i3.8x, I believe, on they basically just run EC2 instances. All of those instances, those exist on our bill. Like, we get billed for them; we pay for them. It’s just basically some sub-account that they have access to that they can deploy into. So, we get billed for the instances of the cluster as well as whatever we pay for our enterprise contract. So, there’s sort of two aspects to the actual billing of it.

Corey: This episode is sponsored in part by our friends at Vultr. Spelled V-U-L-T-R because they’re all about helping save money, including on things like, you know, vowels. So, what they do is they are a cloud provider that provides surprisingly high performance cloud compute at a price that—while sure they claim its better than AWS pricing—and when they say that they mean it is less money. Sure, I don’t dispute that but what I find interesting is that it’s predictable. They tell you in advance on a monthly basis what it’s going to going to cost. They have a bunch of advanced networking features. They have nineteen global locations and scale things elastically. Not to be confused with openly, because apparently elastic and open can mean the same thing sometimes. They have had over a million users. Deployments take less that sixty seconds across twelve pre-selected operating systems. Or, if you’re one of those nutters like me, you can bring your own ISO and install basically any operating system you want. Starting with pricing as low as $2.50 a month for Vultr cloud compute they have plans for developers and businesses of all sizes, except maybe Amazon, who stubbornly insists on having something to scale all on their own. Try Vultr today for free by visiting: vultr.com/screaming, and you’ll receive a $100 in credit. Thats V-U-L-T-R.com slash screaming.

Corey: So, it’s easy to sit here as an engineer—and believe me, having been one for most of my career, I fall subject to this bias all the time—where it’s, “Oh, you’re going to charge me a management fee to run this thing? Oh, that’s ridiculous. I can do it myself instead,” because, at least when I was learning in my dorm room, it was always a “Well, my time is free, but money is hard to come by.” And shaking off that perspective as my career continued to evolve was always a bit of a challenge for me. Do you ever find yourself or your team drifting toward the direction of, “Well, what we’re paying for Redis Enterprise for? We could just run it ourselves with the open-source version and save whatever it is that they’re charging on top of that?”

Jason: Before we landed on Redis on Flash, we had that same thought, like, “Why don’t we just run our own Redis?” And the decision to that is, well, managing such a large cluster that’s so important to the function of our business, like, you effectively would have needed to hire someone full time to just sit there and stare at the cluster the whole time just to operate it, maintain it, make sure things are running smoothly. And it’s something that we made a decision that, no, we’re going to go with a managed solution. It’s not easy to manage and maintain clusters of that size, especially when they’re so important to business continuity. [laugh]. From our eyes, it was just not worth the investment for us to try and manage it ourselves and go with the fully managed solution.

Corey: But even when we talk about it, it’s one of those well—it’s—everyone talks about, like, the wrong side of it first, the oh, it’s easier if things are down if we wind up being able to say, “Oh, we have a ticket open,” rather than, “I’m on the support forum and waiting for people to get back to me.” Like, there’s a defensibility perspective. We all just sort of, like sidestep past the real truth of it of, yeah, the people who are best in the world running and building these things are right now working on the problem when there is one.

Jason: Yeah, they’re the best in the world at trying to solve what’s going on. [laugh].

Corey: Yeah, because that is what we’re paying them to do. Oh, right. People don’t always volunteer for for-profit entities. I keep forgetting that part of it.

Jason: Yeah, I mean, we’ve had some very, very fun production outages that just randomly happened because to our knowledge, we would just like—I would, like… “I have no idea what’s going on.” And, you know, working with their support team, their DevOps team, honestly, it was a good, like, one-week troubleshooting. When we were validating the technology, we accidentally halted the database for seemingly no reason, and we couldn’t possibly figure out what’s going on. We kept talking to—we were talking to their DevOps team. They’re saying, “Oh, we see all these writes going on for some reason.” We’re like, “We’re not sending any writes. Why is there writes?”

And that was the whole back and forth for almost a week, trying to figure out what the heck was going on, and it happened to be, like, a very subtle case, in terms of, like, the how the keys and values are actually stored between RAM and flash and how it might swap in and out of flash. And like, all the way down to that level where I want to say we probably talked to their DevOps team at least two to three times, like, “Could you just explain this to me?” Like, “Sure,” like, “Why does this happen? I didn’t know this was a thing.” So, on and so forth. Like, there’s definitely some things that are fairly difficult to try and debug, which definitely helps having that enterprise-level solution.

Corey: Well, that’s the most valuable thing in any sort of operational experience where, okay, I can read the documentation and all the other things, and it tells me how it works. Great. The real value of whether I trust something in production is whether or not I know how it breaks where it’s—

Jason: Yeah.

Corey: —okay—because the one thing you want to hear when you’re calling someone up is, “Oh, yeah. We’ve seen this before. This is what you do to fix it.” The worst thing in the world is, “Oh, that’s interesting. We’ve never seen that before.” Because then oh, dear Lord, we’re off in the mists of trying to figure out what’s going on here, while production is down.

Jason: Yeah kind of like, “What is this database do, like, in terms of what do we do?” Like, I mean, this is what we store our Identity Graph in. This has the graph of people’s information. If we’re trying to do identity verification for transactions or anything, for any of our products, I mean, we need to be able to query this database. It needs to be up.

We have a certain requirement in terms of uptime, where we want it at least, like, four nines of uptime. So, we also want a solution that, hey, even if it wants to break, don’t break that bad. [laugh]. There’s a difference between, “Oh, a node failed and okay, like, we’re good in 10, 20 seconds,” versus, “Oh, node failed. You lost data. You need to start reloading your dataset, or you can’t query this anymore.” [laugh]. There’s a very large difference between those two.

Corey: A little bit, yeah. That’s also a great story to drive things across. Like, “Really? What is this going to cost us if we pay for the enterprise version? Great. Is it going to be more than some extortionately large number because if we’re down for three hours in the course of a year, that’s we owe our customers back for not being able to deliver, so it seems to me this is kind of a no-brainer for things like that.”

Jason: Yeah, exactly. And, like, that’s part of the reason—I mean, a lot of the things we do at Ekata, we usually go with enterprise-level for a lot of things we do. And it’s really for that support factor in helping reduce any potential downtime for what we have because, well, if we don’t consider ourselves comfortable or expert-level in that subject, I mean, then yeah, if it goes down, that’s terrible for our customers. I mean, it’s needed for literally every single query that comes through us.

Corey: I did want to ask you, but you keep talking about, “The database” and, “The cluster.” That seems like you have a single database or a single cluster that winds up being responsible for all of this. That feels like the blast radius of that thing going down must be enormous. Have you done any research into breaking that out into smaller databases? What is it that’s driven you toward this architectural pattern?

Jason: Yeah, so for right now, so we have actually three regions were deployed into. We have a copy of it in us-west in AWS, we have one an eu-central-1, and we also have one, an ap-southeast-1. So, we have a complete copy of this database in three separate regions, as well as we’re spread across all the available availability zones for that region. So, we try and be as multi-AZ as we can within a specific region. So, we have thought about breaking it down, but having high availability, having multiple replication factors, having also, you know, it stored in multiple data centers, provides us at least a good level of comfortability.

Specifically, in our US cluster, we actually have two. We literally also—with a lot of the cost savings that we got, we actually have two. We have one that literally sits idle 24/7 that we just call our backup and our standby where it’s ready to go at a moment’s notice. Thankfully, we haven’t had to use it since I want to say its creation about a year-and-a-half ago, but it sits there in that doomsday scenario: “Oh, my gosh, this cluster literally cannot function anymore. Something crazy catastrophic happened,” and we can basically hot swap back into another production-ready cluster as needed, if needed.

Because the really important thing is that if we broke it up into two separate databases if one of them goes down, that could still fail your entire query. Because what if that’s the database that held your address? We can still query you, but we’re going to try and get your address and well, there, your traversal just died because you can no longer get that. So, even trying to break it up doesn’t really help us too much. We can still fail the entire traversal query.

Corey: Yeah, which makes an awful lot of sense. Again, to be clear, you’ve obviously put thought into this goes way beyond the me hearing something in passing and saying, “Hey, you considered this thing?” Let’s be very clear here. That is the sign of a terrible junior consultant. “Well, it sounds like what you built sucked. Did you consider building something that didn’t suck?” “Oh, thanks, Professor. Really appreciate your pointing that out.” It’s one of those useful things.

Jason: It’s like, “Oh, wow, we’ve been doing this for, I don’t know, many, many years.” It’s like, “Oh, wow, yeah. I haven’t thought about that one yet.” [laugh].

Corey: So, it sounds like you’re relatively happy with how Redis has worked out for you as the primary data store. If you were doing it all
again from scratch, would you make the same technology selection there or would you go in a different direction?

Jason: Yeah, I think I’d make the same decision. I mean, we’ve been using Redis on Flash for at this point three, maybe coming up to four years at this point. There’s a reason we keep renewing our contract and just keep continuing with them is because, to us, it just fits our use case so well, and we very much choose to continue going with this direction in this technology.

Corey: What would you have them change as far as feature enhancements and new options being enabled there? Because remember, asking them right now in front of an audience like this puts them in a situation where they cannot possibly refuse. Please, how would you improve Redis from where it is now?

Jason: I like how you think. That’s [laugh] a [fair way to 00:28:42] to describe it. There’s a couple of things for optimizations that can always be done. And, like, specifically with, like, Redis on Flash, there’s some issue we had with storing as binary keys that to my knowledge hasn’t necessarily been completed yet that basically prevents us from storing as binary, which has some amount of benefit because well, binary keys require less memory to store. When you’re talking about 4 billion keys, even if you’re just saving 20 bytes of key, like you’re talking about potentially hundreds of gigabytes of savings once you—

Corey: It adds up with the [crosstalk 00:29:13].

Jason: Yeah, it adds up pretty quick. [laugh]. So, that’s probably one of the big things that we’ve been in contact with them about fixing that hasn’t gotten there yet. The other thing is, like, there’s a couple of, like, random… gotchas that we had to learn along the way. It does add a little bit of complexity in our loading process.

Effectively, when you first write a value into the database it’ll write to RAM, but then once it gets flushed to flash, the database effectively asks itself, “Does this value already exist in flash?” Because once it’s first written, it’s just written to RAM, it isn’t written to backing flash. And if it says, “No it’s not,” the database then does a write to write it into Flash and then evict it out of RAM. That sounds pretty innocent, but if it already exists in flash when you read it, it says, “Hey, I need to evict this does it already exist in Flash?” “Yep.” “Okay, just chuck it away. It already exists, we’re good.”

It sounds pretty nice, but this is where we accidentally halted our database is once we started putting a huge amount of load on the cluster, our general throughput on peak day is somewhere in the order of 160 to 200,000 Redis operations per second. So, you’re starting to think of, hey, you might be evicting 100,000 values per second into Flash, you’re talking about added 100,000 operate or write operations per second into your cluster, and that accidentally halted our database. So, the way we actually go around this is once we write our data store, we actually basically read the whole thing once because if you read every single key, you pretty much guarantee to cycle everything into Flash, so it doesn’t have to do any of those writes. For right now, there is no option to basically say that, if I write—for our use case, we do very little writes except for upfront, so it’d be super nice for our use case, if we can say, “Hey, our write operations, no, I want you to actually do a full write-through to flash.” Because, you know, that would effectively cut our entire database prep in half. We no longer had to do that read to cycle everything through. Those are probably the two big things, and one of the biggest gotchas that we ran into [laugh] that maybe it isn’t, so known.

Corey: I really want to thank you for taking the time to speak with me today. If people want to learn more, where can they find you? And I
will also theorize wildly, that if you’re like basically every other company out there right now, you’re probably hiring on your team, too.

Jason: Yeah, I very much am hiring; I’m actually hiring quite a lot right now. [laugh]. So, they can reach me, my email is simply jason.frazier@ekata.com. I unfortunately, don’t have a Twitter handle. Or you can find me on LinkedIn. I’m pretty sure most people have LinkedIn nowadays.

But yeah, and also feel free to reach out if you’re also interested in learning more or opportunities, like I said, I’m hiring quite extensively. I’m specifically the team that builds our actual product APIs that we offer to customers, so a lot of the sort of latency optimizations that we do usually are kind of through my team, in coordination with all the other teams, since we need to build a new API with this requirement. How do we get that requirement? [laugh]. Like, let’s go start exploring.

Corey: Excellent. I will, of course, throw a link to that in the [show notes 00:32:10] as well. I want to thank you for spending the time to speak with me today. I really do appreciate it.

Jason: Yeah. I appreciate you having me on. It’s been a good chat.

Corey: Likewise. I’m sure we will cross paths in the future, especially as we stumble through the wide world of, you know, data stores in AWS, and this ecosystem keeps getting bigger, but somehow feels smaller all the time.

Jason: Yeah, exactly. You know, we’ll still be where we are hopefully, approving all of your transactions as they go through, make sure that you don’t run into any friction.

Corey: Thank you once again, for speaking to me, I really appreciate it.

Jason: No problem. Thanks again for having me.

Corey: Jason Frazier, Software Engineering Manager at Ekata. This has been a promoted episode brought to us by our friends at Redis. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry, insulting comment telling me that Enterprise Redis is ridiculous because you could build it yourself on a Raspberry Pi in only eight short months.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Wayne

Professionally, I'm a Vice President at Amazon Web Services (AWS) where I lead a set of businesses delivering cloud infrastructure services. In 2013, I founded and continue to lead the AWS Boston regional development center. I'm an always-curious entrepreneur who is passionate about building innovative teams and businesses that deliver highly disruptive value to customers. I love engaging people who build and deliver customer-obsessed solutions, as well as customers wanting to realize value from those solutions. I hold over 40 patents in distributed and highly-available computer systems, digital video processing, and file systems. Personally, I'm a proud dad to great people, I love to cook and grow things, it relaxes and grounds me, and I cherish finding adventure in the ordinary as well as the extraordinary.

Links:

  • LinkedIn: https://www.linkedin.com/in/wayneduso/
  • Twitter: https://twitter.com/wayneduso

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Sysdig. Sysdig is the solution for securing DevOps. They have a blog post that went up recently about how an insecure AWS Lambda function could be used as a pivot point to get access into your environment. They’ve also gone deep in-depth with a bunch of other approaches to how DevOps and security are inextricably linked. To learn more, visit sysdig.com and tell them I sent you. That’s S-Y-S-D-I-G dot com. My thanks to them for their continued support of this ridiculous nonsense.

Corey: This episode is sponsored in part by our friends at Vultr. Spelled V-U-L-T-R because they’re all about helping save money, including on things like, you know, vowels. So, what they do is they are a cloud provider that provides surprisingly high performance cloud compute at a price that—while sure they claim its better than AWS pricing—and when they say that they mean it is less money. Sure, I don’t dispute that but what I find interesting is that it’s predictable. They tell you in advance on a monthly basis what it’s going to going to cost. They have a bunch of advanced networking features. They have nineteen global locations and scale things elastically. Not to be confused with openly, because apparently elastic and open can mean the same thing sometimes. They have had over a million users. Deployments take less that sixty seconds across twelve pre-selected operating systems. Or, if you’re one of those nutters like me, you can bring your own ISO and install basically any operating system you want. Starting with pricing as low as $2.50 a month for Vultr cloud compute they have plans for developers and businesses of all sizes, except maybe Amazon, who stubbornly insists on having something to scale all on their own. Try Vultr today for free by visiting: vultr.com/screaming, and you’ll receive a $100 in credit. Thats v-u-l-t-r.com slash screaming.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Way back in the winter of 2017, I went to my first re:Invent, and at that re:Invent, or shortly before, they announced EFS, Elastic File System, which is basically a NetApp in the cloud, killing some of my stuffing a NetApp into us-east-1 jokes. And my initial, considered, reaction—because I’d been writing the newsletter for about six months at that
point—was, “What a piece of crap.” And the general manager of the product said, “Hey, can we meet at re:Invent?”

And showed up with, as I recall, a couple of engineers who had no neck, and I thought, “Oh, great, this is how I die.” Instead of what I expected, what happened was he asked a bunch of questions and took notes. And one by one, every issue that I had with a service—and it was lengthy—wound up getting knocked out in the next couple of years. And today, it’s one of my absolute favorite services and I use it daily. That was an early but lasting impression of how a lot of my interactions with AWS were to go. And I’m very glad today to have that former GM and now VP of Engineering at AWS, Wayne Duso here to suffer more of my slings and arrows. Wayne, thank you for joining me.

Wayne: Corey, it’s always a pleasure. Thank you.

Corey: It really was a transformative moment for me, just because I was still finding my own voice in this. Because my big fear then was, “No one’s going to read it; no one’s going to listen. I’m going to starve to death and have to get a real job and get fired some more.” And it was just about get out there at any cost. I didn’t have the reach that I did, then, and I was a lot more cynical, and in some ways, directly opposed to service launches.

I try not to do that anymore; don’t always succeed. But I never got the sense that you took any of that feedback personally. And in some cases, you set me straight when I was wrong on things, and in others, you’re not only listened, you agreed with me and took steps to make it better. This is not me shining you on. This is very much the course of EFS.

Wayne: Yeah, Corey, you know, your feedback was super important. It was early days, and we don’t pretend that we get everything right the first time. You know, we don’t. You write about how we don’t get it right the first time quite often. And it’s appreciated because, you know, as engineers, as technologists, as leaders, you’re always looking for input.

Input is one of the most valuable things we can get. You know, not often people don’t provide us input. And when I met you and you had your long scroll of input for us, for me it was a treasure trove. But it really was a treasure trove. And to this day, whether it’s you or, you know, the next you, those scrolls are still treasure troves because it means that somebody has a different view than what I currently have.

Corey: It stuck with me because suddenly it came crashing home to me that a lot of the criticisms that I was lobbying against AWS weren’t just me shouting into the void. They were being heard. And also, it really was one of those reaffirmations back then; it’s like, “Oh, yeah, making fun of Amazon Web Services; they’re a division of a company that’s however many trillions of dollars it was at the time,” and, “That’s not a person. There are no people there. It’s just a big company.”

And the dawning, creeping realization was that companies are made of people. And instead of just saying, “This is crap.” I’ve got to be able to articulate why that is. And I guess the one enduring criticism that I had about EFS that still holds true on some level, is—it has nothing to do with the capabilities of the service, but more to do with the pattern of if I’m building something to work in the cloud on day one using what is, effectively, NFS that attaches to a bunch of different instances simultaneously—or now containers or Lambdas Function—great, that feels like it’s not the direction that cloud architecture has historically gone, but now that the capability is there, that’s starting to shift, and its own right once again. So, it’s just… it was a service that I didn’t fully understand then—I’m not sure that I still do—but I definitely see the use of it and the utility of it. And by just sitting here and doing effectively nothing, it gets better with time. And if there’s a better story about cloud, I’m not sure what it is.

Wayne: You make it sound like a wine. If you, you know, if you do a proper job mixing those grapes, eventually you put in a bottle, it gets better and better as time goes on, but.

Corey: And eventually becomes vinegar, and we can list, like, three or four of those services, too, but let’s be charitable here.

Wayne: [laugh]. We’re not going to go there. So, thank you for popping the cork on this bottle of wine and drinking it on a regular basis. Files are everywhere. You know, when I came to the company in 2012, and I was handed a question, “How should we build? What should we build?”—in terms of a file offering—I asked why in the same way that I’m sure you asked why when we launched. And the truth of the matter is that doesn’t matter what type of operating system you’re on or what you’re doing files are a fundamental abstraction.

And every operating system—every language—has an interface, which is file-based. It’s a really good abstraction. And you know, objects are great abstraction blocks are great abstraction, like, they all exist for a reason. They all serve a purpose. If they didn’t, they would have gone away.

And files have, you know, been around for at least 50 years, and honestly, they’ll be around for at least another 50, and probably beyond. So, it was really important to build a capability that customers—builders—could use straight out of the tools that they use every day without having to go through any transformation. It was important.

Corey: When I say that EFS is a contender out of maybe four others for my favorite service, that people think that it’s because I love files or I love storage, and okay, fine, but that’s not the real reason; that’s sort of irrelevant. I mean, I like SSL certificates too, but ACM isn’t on that list either. What I found valuable about it, what I love about it as an exemplar service is that if I go through the console wizard to spin up an EC2 instance—which of course I do because the ultimate form of managing cloud resources is using the console and then lying about it; it’s called ClickOps—and as I go through that process, adding an EFS file system to it is built into that wizard and it just works. When you spin up an EFS volume, it automatically configures AWS Backup to begin taking backup snapshots of what’s going on in there.

And there’s a bunch of niceties built into that. Recently, at re:Invent, there were additional changes rolled out—or pre-re:Invent—rolled out to EFS that enable automatic data tiering where, “Oh, that file hasn’t been touched in a while; we’re going to move that to the infrequent access tier and charge you less for it.” And this all happens transparently in the background. It is a service that gets better without requiring active customer interaction, and you’re sort of swimming against the stream of the typical AWS service approach by talking to other service teams and integrating into them in a way that is transparent to the customer. And I’m looking at this and I’m tapping the screen going, “More like this, please.” How did you get there?

Wayne: I’m a lucky guy in many ways. You know, I use this term inside of teams on a regular basis, which is, “We stand on the shoulders of giants.” And EFS was not the first service to launch at AWS; we know that… it wasn’t SQS, it was S3. And—

Corey: We’re talking beta or general availability—

Wayne: [laugh].

Corey: Those are the two teams that will wind up fighting to the death, and I just stay out of it.

Wayne: I know. I just wanted to beat you on that one. And so many great services, whether it’s Dynamo or S3 or EBS, came before the services that I built for customers—or my teams have built for customers. And so I get to stand on their shoulders and look far out onto the horizon; and also, I get to look backwards at mistakes or issues that we may have made, you know, earlier. And if I make the same mistakes, shame on me. I should be able to ensure that what we produce takes from the lessons that are already been learned.

So, for EFS as an example, I make sure that I look at all the lessons that came from EBS or S3 or Dynamo in terms of their scale and their usability and so on and so forth and say, “How do we make sure that we’re as good if not better for our customers?” And at some point, somebody will stand on the shoulders of EFS and they’ll look forward and say, “Oh, we got to do better than what they did in their first couple of years.” And I look forward to that.

Corey: I want to be clear that your portfolio has expanded significantly. It turns out that you were the GM and then now you’re a VP of Engineering. And so what changed? “Oh, same job, just different title.” And that is very much not true.

You own a bunch of other services as well—largely storage-oriented—including the snow family, which means that you are the person to, basically, harangue into—that I get to a point being able to check that item off my lifelong bucket list of beeping the horn on a snowmobile. It’s going to happen someday; I just don’t know how or when, but I feel like you’re someone who can help me set that up.

Wayne: But I live in a part of the country where we need snowmobiles, so it behooves me. [laugh].

Corey: Oh, yes. Growing up in New England, I remember those days, too. It’s, “Okay. Don’t plow the street. That’s fine. I’ll just cross-country ski to school.”

Wayne: I have stories but we won’t go there.

Corey: Oh, yes.

Wayne: You know, Corey, in my tenure, which is now a roughly a little over nine years at AWS, I’ve had the opportunity to meet with a lot of enterprise customers, I have a very rich enterprise background base on my career, so it was very natural for me to engage cohorts of customers, storage administrators, IT administrators, application administrators, network administrators, all of which have a rich history in success in the enterprise. And in speaking with them, it became incredibly clear that there were a series of enterprise-based services that needed to be built for cloud-scale, the cloud model of consumption, that just didn’t exist. And I’m not a big fan for creating services for the sake of creating services, just like you are not a big fan of us doing that, so I wanted to make sure that everything we built was filling a need that could not be filled any other way. And in that nine years, my teams and I have built roughly ten services covering storage, edge compute, edge storage, and data services like data protection, data movement, data management. And each of these services has deep roots in those enterprise cohorts.

Corey: One thing that’s become obvious to anyone who pays attention in the space has been that when we look at the sweep of history when it comes to technology, it’s that the tide always rises. When I first started playing around with technology, firewall engineer was a specific job. Now, any network administrators expected to be able to handle that just in due course. And when you’re talking about storage, the same thing happens. Now, there have been people who spent 30 years of their career as storage admins or working on specific technologies, and now that cloud is disrupting so many aspects of this, there’s a tendency that technologists have to push back against anything that disrupts the thing that they work on because they equate their identity to the thing that they build.

And let me be very clear here. I don’t think most people are immune to this. I know I’m not. It’s one of the reasons I tend to get so irritated about things that are disruptors to the way I thought about doing things. Why do you think I hate containers so much? It’s because well, that’s something that sounds like a different paradigm, and I’m not good at that, but I’m very good at this older thing.

And letting go, like, I’ve done that a few times and changed my focus throughout my career, but it’s always been challenging to do it. When I tell the story. “Oh, yes. I saw this thing and then I pivoted.” “Yeah, it wasn’t that easy.” There was angst and pain tied to it, and the constant awareness that I’m going to become a dinosaur if I don’t change my area of focus—and by the way, I’m 26 years old at the time—it was a hard thing to do.

Storage is such a important thing for so many companies in so many environments, that it’s such a nuanced deep area, that there is significant disruption happening there. How have you found those conversations to go with the folks at your customers who probably identify themselves as, “I hate the cloud. We should build more data centers.” Which is not actually what they’re saying, but it’s how it’s expressed.

Wayne: Yeah. Evolutionary change is something that the folks that you’re talking about understand, they embrace evolutionary change; they’re curious people, they love to learn; they’re passionate about what they do, and they’re passionate about their customers. It’s the revolutionary change that is hard. And they often view the cloud consumption model versus the traditional IT consumption model of receiving equipment, setting it up, putting on the floor, so on and so forth, they see that as an evolution, where they saw cloud as a
revolution. And they don’t want to necessarily embrace such a big change.

It’s part of my job and the job of my peers to have those folks, the storage administrators, the application administrators understand that it’s not as revolutionary as it may seem. And I want to give you an example of how we’ve done that. We’ve talked at length about EFS; EFS has a sibling and it’s known as FSx. And that really stands for File System X, which is any file system. What does that really mean?

Corey: I figured that one out a—what was it? It was something like three years after FSx launched, two years, something like that—I thought FSx was some storage term I’d never heard before. And nope, I was over-complicating it. It turns out—surprise—I don’t know everything either.

Wayne: Yeah, well, we were wrestling really hard with what we should name FSx. Look, I’ll tell you that story in a second, but what FSx is, is it’s a sibling service to EFS. EFS was built to be a cloud-native set-and-forget super-simple file system on AWS. Now, with that, there are 30 years of features that we did not implement when we launched EFS because the vast majority of those aren’t needed by everyone, so we didn’t complicate the solution.

That being said, if you think about storage administrators and application administrators in the enterprise, today, they use very specific file systems and the characteristics of those file systems, whether it’s performance, or durability, or—more importantly—management APIs, Snap APIs, replication APIs, all sorts of various data management capabilities, data service capabilities. FSx is a service that was designed to bring those file systems to AWS as fully managed offerings so they could be consumed in a cloud-native fashion. You’ve seen the launch of four of those to date. The first one we launched was FSs for Windows Server so that customers that used Windows servers on-prem could lift and shift their applications their workloads to a like offering on AWS. It’s bit-compatible.

The second was Lustre, we talked to a whole bunch of customers, amazing use cases, you know, FMI that is helping have cancer treatments that are specific to individual patients, they needed to be able to run high-performance workloads on AWS. Lustre was a great solution for them, but everybody knew that Lustre was a little hard to run on-prem. It took a lot of energy, took people to keep it up and running. But we took all of that away from them and provided them a fully managed Lustre offering, which they just create the file system, load the data into the file system, and go, and they worry about nothing after that.

Those were the first two. With the success of those, we heard customers coming to us—same customers say, “Hey, listen. I got a floor full of NetApps. I would love to run my NetApp-based workloads and applications on AWS. Can you help us?”

And we partnered with NetApp—it was a very deep partnership for two years—to create a bit-compatible like-for-like, no excuses, NetApp offering on FSx. And then we quickly followed it a couple months later, with a ZFS offering, which is FSx for OpenZFS because for the folks that aren’t running NetApp, there’s a whole bunch of on-pram that are running ZFS-based file systems, and they rely on the data management APIs of that file system. So, you know, for this cohort, we took what seemed to be a revolutionary change, and we turned it into an evolutionary change. And now our job is to go speak to all of those folks and have them understand that they can make that switch without losing the 20 years, the 30 years of expertise, without losing the trust of their customers, the folks that are running on their storage because they can guarantee them that what they have today is what they will have when they move. Super important.

Corey: I think that it does come down to trust that it does come down as well to the challenge that you’re in when it comes to storage. As we look at the broad sweep of where cloud business is coming from. It’s easy to sit here and pretend that we’re all somehow—we’re all doing net new; we’re building out this thing in a garage somewhere. “I have an idea.” “What is it?” “Twitter for Pets.” “Is that a good idea?” “Absolutely not, but I just raised 200 million in seed funding, so we’re going to do it anyway.”

A lot of this stuff is existing legacy workloads. And legacy is, of course, a condescending engineering term for, “It makes money.” But that is what you have to be able to address because to move your workload to cloud is a heavy lift. It becomes borderline impossible with oh and to do that, you’re going to have to completely re-architect the entire data layer, because that thing you’re doing right now, not so much. On tap in the cloud as an FSx NetApp offering was transformative for me.

Because back when I was in data center land, NetApp was the only way I would willingly run NFS in production. And I ran production MySQL databases on top of it. Professional advice, if you’re listening to this elsewhere, don’t do that. But it can be done, and it was amazing. And WAFL remains one of my favorite file systems ever.

And I kept joking that I just wish I could get it in AWS, but they won’t let me shove a filer into us-east-1. I know because I’ve asked. Well, you went ahead and did that for me. So thanks.

Wayne: You’re welcome. It’s our pleasure. [laugh].

Corey: I really hope that is still relevant to places I worked back in 2007, but let’s face it, “This is a temporary fix,” and 20 years later, it’s still load-bearing in production. So, here we are.

Wayne: You know, you’ll never hear me use the word ‘legacy.’ And in fact, within the company, we’ll be in rooms and people use the word legacy—it’s a very popular word people love to use it—and what I often will tell them is that if an application workload is important to a customer and it is helping run their business. There is nothing legacy about it. It is current, it is real, and we need to think about it in the way they think about it, not in a way, which has us in any way deprecate its importance, the customers will deprecate its importance if that’s important to them, so our job is to embrace what they have and help them, if they so choose, to bring those workloads to AWS without change. I’m going to give you an example. I’m not sure I can use the customer’s name because I often lose track of who I can talk about in public and who I can’t. So, I [crosstalk 00:20:32]—

Corey: I long ago stopped paying attention to the details of NDAs, which sounds like a terrifying statement to make until you realize the way I do that is I just assume every conversation I have is confidential until I can affirmatively prove otherwise. And I’ll ask people sometimes like, “Hey, remember the thing you said to me? Can I quote that in public?” And they look at me strangely, and say, “It was on a podcast and we put it out to the world. Yes. Yes you can.” Oh, good.

Wayne: [laugh]. Well, in this particular case, I can tell you the country and I ca—and actually continent in this case. It was a customer down in Australia, and they ran into a situation where they needed to move out of their data centers, and they needed a place to go and AWS was already provided to them. And they moved 40 Windows-based applications to AWS over a weekend. I feel bad for the folks; hope they get a few weeks off—at least a few days off after that event.

But they were able to move almost their entire business to AWS and FSx for windows over a weekend because the experience simply was, create the file system, spin up the application, mount the file system and go. And it worked exactly as it had on-prem. Exactly as it had on-prem. When I hear stories like that, as a builder, as an engineer, as a human, I am incredibly happy. It is a day worth living when you know that you’ve helped a customer.

Corey: When you come in and look at an architecture like that, see it for the first time. And you look at like, “Hey, you’re using the cloud like it’s a data center. What’s the deal here?” And if you say that in a condescending, insulting way, they’re sitting around talking about what an amazing achievement something like that is—and let’s be clear; having done those projects, it’s an amazing achievement—but to have someone come in with no context, “Oh, you should have done this in a more cloud-native way.” It’s, “Thank you, Seymour. Yes, if we were building this stuff bespoke today, greenfield, we would have radically different constraints and radically different capabilities, but we don’t have a time machine so instead we have to move forward and we can’t just burn everything down every 18 months and start over from scratch.” There’s value there.

Wayne: Yeah. And they have the ability over time, Corey, to—I’m using this term lightly—to modernize. Because I don’t really know what that means; I’m a big fan of mid-century modern furniture, which is no longer modern. [laugh]. But it is lovely.

Corey: A glimpse of the future that didn’t happen. An alternate path, a speculative fiction expressed through furniture. I hear you.

Wayne: But what I’d like to make sure people understand is that they can move today. They can utilize these capabilities and these services today and move, and then they can take their time and how they want to evolve those applications based on the needs of their customers, based on needs of their business. I do not want to slow people down. The entire intent of the portfolio that I’ve had the privilege to build and provide to customers over these years, its intention is to enable customers to move as quickly as they want, and to then take whatever time they need to evolve that, based on your business needs. It’s really that simple.

Corey: Today’s episode is brought to you in part by our friends at MinIO the high-performance Kubernetes native object store that’s built for the multi-cloud, creating a consistent data storage layer for your public cloud instances, your private cloud instances, and even your edge instances, depending upon what the heck you’re defining those as, which depends probably on where you work. It’s getting that unified is one of the greatest challenges facing developers and architects today. It requires S3 compatibility, enterprise-grade security and resiliency, the speed to run any workload, and the footprint to run anywhere, and that’s exactly what MinIO offers. With superb read speeds in excess of 360 gigs and 100 megabyte binary that doesn’t eat all the data you’ve gotten on the system, it’s exactly what you’ve been looking for. Check it out today at min.io/download, and see for yourself. That’s min.io/download, and be sure to tell them that I sent you.

Corey: There’s a lot that you said that I want to dive into, but that the piece that I want to focus on specifically is you said you’ll never use the word legacy. And I’m going to challenge you on that because in one of our conversations that we had, I asked you what product you enjoyed building the most? And your answer was, “Leaders.” And that gets into a different form of legacy. And yes, I’m playing semantic games here, but it’s the question there of what are you going to be remembered for, because God willing, none of us are building technological solutions that are still going to be at least in common use in 40 years from now, but I don’t necessarily know that the same can be said of people.

Tell me more about this because you no longer oversee a product; you oversee a bunch of products. And something I learned the hard way is that when you become management and later different forms of management, you can do very little of the work yourself, if any. Your only tool is delegation, and that means you need to have the right people to delegate to which means hiring and handling leaders is effectively your IDE for lack of a better term. Talk to me about that.

Wayne: It’s a really important point. And one of the mental frameworks I put in place to myself a handful of years back, which guided me to AWS in fact, is ‘who,’ ‘how,’ and ‘what.’ Who I work with is most important to me, how I do my work come second, and what I work on comes a distant third. And the reason for that is so simple to me now—it wasn’t when I was a young engineer where ‘what’ was the most important thing on the planet; you know, what I worked on define me. And it took me years to understand that what I work on doesn’t define me, it’s who I work with and how I do that work that defines who I am. You know, a really, really lousy day with great people is still a good day. And a really great day with people you don’t want to be around is still not a great day.

Corey: I think that a lot of companies, maybe all companies, tend to get a somewhat unfair reputation once they hit a certain point of size and scale. There have been all kinds of exposés in various places about how a company X—and it doesn’t matter who X is; it can be Amazon, it can be any company people have heard of—is a terrible place to work. But I was talking to a friend at AWS recently who is coming back from parental leave, and they told me that, “So, get this. AWS changed their policy while I was out on parental leave. I have another month of it to take later in the year, and I’m really looking forward to that.” It’s like, “That’s amazing.” As opposed to the constant stories of oh, this thing is awful and this thing is terrible.

It’s very hard to see from the outside what a company is actually like. And I apply that to me. I only have the faintest glimmer of what it’s like to actually work at AWS. But talking to the people that I get to talk to and seeing how things are built, is it perfect? No. No place is. But I hear stories about different approaches to leadership and different teams, and, on some level, I become intensely grateful that I never tried to work on any of those teams because I would have given everything I was building up to have the chance to be a part of those groups, and then where would the industry be? Certainly without my snark, and that—we would all be the poorer for it.

Wayne: That’s true.

Corey: But there are always pockets of amazing things, same way there are pockets of terrible things. AWS at its—what—however many tens or hundreds of thousands of employees it has now doesn’t have a corporate culture. It has hundreds of thousands of corporate cultures because it is a team-by-team structure there. And culture—there are—from a principles perspective, has to flow from the top, but then you have management and how they run their organizations and their teams, and the blast radius attached to that is tremendous. And that’s the sort of thing that always scared the hell out of me when I was debating. Do I stop being an independent contractor—independent consultant and start hiring people here full-time?

Because the biggest blocker I had was that, yeah, if I say something wrong and get smacked off the earth by AWS, okay, great. I had it coming. Asking people to risk the next phase of their careers on me saying or doing the right thing was a hell of a responsibility. That’s the stuff that keeps me up at night. Not that I’m going to have a wrong take or something. It’s the letting down the people who are depending on me to get things right.

Wayne: Yeah. So, leadership is a, it’s a lot of fun; it’s a lot of responsibility as well. And so to your question, the thing that I will build—going back to the who, how, and what—the thing that I will build will the most important—the last thing I build will the most important—is the leaders I leave behind. Those leaders will go on for decades and build more leaders. They’ll build more products, they’ll build more businesses, they’ll do great things, but my job is to ensure that I build the right leaders that ensure that the culture that we do have at AWS—or wherever else they decide to go in the fullness of time—they will carry with them the best elements of the AWS or Amazon culture, through AWS or into other companies.

I’ll use the word: That will be my legacy. And right now, I look at doing that, not just with folks I bring into management roles, but I look at that in terms of the young engineers, the young data scientists, the young DevOps folks that we bring into the company, and say, “Where can I find incredibly passionate and curious learners, regardless of their background, regardless of where they come from, who can, in the fullness of the next five to seven years, become those leaders?” At that point, we will have a company that continues to represent the people, continue to represent the customers and the communities we serve. We will understand our customer. That to me is what it means to be customer-obsessed is to understand your customer. And you can understand your customer best by being a representative population of who they are.

Corey: One thing that we often decry in this industry is the pox of short-term thinking and over-emphasis on next quarter’s numbers. Think longer-term. But when people say that, they’re often thinking in three to five years out. You’re talking about something that spans decades and that is borderline unheard of in corporate life.

Wayne: Yeah, well. Thanks. [laugh].

Corey: [laugh]. I mean, I had to think about this stuff a fair bit myself recently because let’s face it, I have rendered myself completely unemployable. I went into this knowing it was, like, burn the boats behind me because this, for better or worse, is the last job I will ever have. And—maybe because I’m going to get killed in two months; we don’t know–but it’s—I’m very good at antagonizing people with a lot of bigger spite budget than I have—but that is how I have to think about these things. In the context of something at your scope and scale—because if your storage service or services have a bad day, so to all of your customers. That is something that weighs on the mind of every Amazonian I’ve ever spoken to, including someone who’s joined their first job out of school, two weeks ago.

It’s a culture of not fear, but awareness of the weight of responsibility. And some folks obviously carry that better than others, but it is one of those things when I start talking to people who are new Amazon employees, and I’m starting to be able to categorize them mentally how they talk about certain things of, “You’re going to go far at this company,” versus the, “Let’s see how this plays out.” You can pick up on some of these things, in the fullness of time.

Wayne: Having this level of responsibility, knowing that everything from entertainment to critical life care is dependent on the decisions you make and on the product you build and operate is a great responsibility. I can speak personally, it motivates me every day. You know, I can probably pick three days out of the near ten years I’ve been building at AWS where I just wanted the day off. [laugh]. I didn’t want that responsibility.

Now, of course, we have the leadership I just talked about that runs the services every day, who was there to make sure that on those three days that I didn’t show up, everything was going to be fine. And it, of course, was fine. But doing what we do is hard work. But I’ve never found anything more rewarding. And just speaking to a customer—a single customer—who’s had a great day, or a customer hasn’t had a great day, but you’ve done the right thing and there trust in you is strong, they’re unhappy with you, but their trust in you is strong. That is also a great day.

Corey: So, much of what happens in cloud is thankless. We started off this episode talking about how the EFS integration into other services. It is just this thing that I remark upon about how wonderful and great it is, but for most customers, like, “Okay.” They just click it and it’s great. When it’s not there, they find it obnoxious and challenging, but when it’s there, just, “Okay, this is working as it should,” and move on. So, much of the work and challenge and victories are unsung, and I try not to pass those things idly by and just take the time to appreciate the things that I see.

Because as I said, companies are made of people, and there’s a tremendous amount of innovation and improvement that’s going on constantly. I sometimes say—probably not enough—that 90, to 95% of what AWS does is excellent. That missing gap to get to 100 is both frustrating and, honestly, rife with opportunities for hilarity. So yeah, I don’t want to spend all my time talking how great all this stuff is—I should just work in marketing if that’s the case—I want to talk about what it takes to go from great to perfect, and you’re never going to get there, but that’s okay; it means I’ve going to have a job for as long as I want one.

Wayne: You know, there’s a term that we often use, and that is for our customers, we need to operate our services so that they’re indistinguishable from perfect. That’s a tall bar.

Corey: Mistakes show.

Wayne: Yeah. But it is our responsibility. And frankly, as builders and strong owners, it comes with a lot of fun. It’s really hard, but it comes with a lot of fun. Like, being able to do that… you know, this conversation we’re having right now is likely traveling over some piece of the cloud… everything that was enabled during… the pandemic, which has been very challenging for a lot of people—for everybody—for your haircut for a while it was very challenging for.

Corey: Oh, that was one of the hardest parts of the pandemic. I expect to find that the list of pandemic-related fallout on Wikipedia someday.

Wayne: [laugh]. And, you know, what we do enabled a large part of people’s personal lives, business lives, the economy to continue at some level of normalcy, which if it did not exist, it would have been a very different place. And we’re very grateful for being able to be able to do that. We’re very proud that we’ve been able to do that at the scale we do during the last couple of years. It’s taught us a lot of lessons. It’s been fantastic.

Corey: I’m very light-hearted about some of the workloads I put on AWS. I have a Last Tweet in AWS Twitter client that basically does shitposting threads in long-form. And it’s great. And if it winds up breaking, then there’s no big deal. I don’t care about that.

But I don’t have to care about that; you do have to care about that because just because I don’t take one of my workload seriously, you as AWS can never have that context. And that’s why you take every weird blip, everything, “I don’t understand this,” or, “This hasn’t lived up to my expectations on it,” as if it were a life-critical system because for some customers they are. And I have always appreciated how—and been bemused at times—by how deadly seriously AWS folks take my complaints about, “Yeah, my shitposting app isn’t posting shit quite the way I want it to.” [unintelligible 00:36:05] anything but the utmost of professionalism and respect when I’m talking about service gaps and challenges I’m having building and deploying things. And I feel a little bad at times, just because I’m making people care so seriously about things that don’t actually matter for crap. But I’ve always appreciated it.

Wayne: Our frame of reference is the millisecond and the penny. We worry about every penny you spend, and we want to make sure you—

Corey: Oh, there are times that, in some cases, for some services—I believe it was trillionths of a cent is how a couple of them have a granularity in the billing system. So, all you care about far smaller denominations than pennies.

Wayne: Well, I might be a little generous in what I’m saying right now, but for the frame of reference for the listener, you know, we don’t think about things at the quarter and the million dollars. That’s not our frame of reference. We think about things in terms of the millisecond a penny. And you know, yes, you’re right, we now think more in microseconds than we do milliseconds, and we do think in fractions of pennies, not pennies. But it’s a frame of reference.

What matters is the details. And there is no workload is unimportant. If it’s your workload, it’s just as important to us as any other workload. Even if it does poke snark at us. We’re okay with that.

Corey: And sometimes there is the idea of a memento mori, or someone yelling at the emperor that they have no clothes. The problem is in the story of the emperor not having any clothes, the kid would at least occasionally shut up once in a while, and I never seem to so
there is that part of it too.

Wayne: Springs hope eternal.

Corey: Exactly. Wayne, I want to thank you for taking so much time to speak with me today. If people want to learn more about what you’re up to, and how you view these things, where can they find you?

Wayne: Well, the two places they can find me most often, Corey, not at my desk or at a local bar; they can find me on LinkedIn, and it’s Wayne Duso. There’s only two of us on LinkedIn, so find the guy who wears glasses that has no hair, and that’s me. And the second place they can find me is on Twitter, which my handle is not obfuscated at all. It’s @wayneduso.

Corey: And we will, of course, put links to both of those in the [show notes 00:38:05]. Thank you so much for your time today. I really appreciate it.

Wayne: Corey, it’s always a pleasure. And I’m looking forward to sharing maybe a drink at that bar with you soon.

Corey: I look forward to it. I can’t wait to go back out to bars again. Oh, my God. [sigh]. Wayne Duso VP of Engineering at AWS. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice telling me that I’m completely wrong when it comes to using EFS for things that I should instead be using a better storage system that’s more cloud-native. Like Route 53 [unintelligible 00:38:48].

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Ana

Ana Visneski is the founder of Merewif, a crisis communications and management consulting firm. She is a veteran of the U.S. Coast Guard where she was a first responder to major disasters from Hurricane Katrina to the BP Oil Spill, and various other incidents. After the USCG, Ana moved on to a whole new disaster that needed an experienced crisis operator - running Launch Operations for AWS. Following that she was the global lead for AWS Disaster Response, overseeing deploying AWS technology response to natural disasters and overseeing the response to COVID. She has a Master of Communication Digital Media and a Master of Communication in Networks from the University of Washington, where she currently teaching Crisis Communications.

Links:

  • Mirewif: https://www.themerewif.com/
  • Oracle HeatWave: https://www.oracle.com/mysql/heatwave/
  • Twitter: https://twitter.com/acvisneski
  • The—T-H-E—merewif—M-E-R-E-W-I-F dot com: https://www.themerewif.com/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Vultr. Spelled V-U-L-T-R because they’re all about helping save money, including on things like, you know, vowels. So, what they do is they are a cloud provider that provides surprisingly high performance cloud compute at a price that—while sure they claim its better than AWS pricing—and when they say that they mean it is less money. Sure, I don’t dispute that but what I find interesting is that it’s predictable. They tell you in advance on a monthly basis what it’s going to going to cost. They have a bunch of advanced networking features. They have nineteen global locations and scale things elastically. Not to be confused with openly, because apparently elastic and open can mean the same thing sometimes. They have had over a million users. Deployments take less that sixty seconds across twelve pre-selected operating systems. Or, if you’re one of those nutters like me, you can bring your own ISO and install basically any operating system you want. Starting with pricing as low as $2.50 a month for Vultr cloud compute they have plans for developers and businesses of all sizes, except maybe Amazon, who stubbornly insists on having something to scale all on their own. Try Vultr today for free by visiting: vultr.com/screaming, and you’ll receive a $100 in credit. Thats v-u-l-t-r.com slash screaming.

Corey: This episode is sponsored in part by our friends at Sysdig. Sysdig is the solution for securing DevOps. They have a blog post that went up recently about how an insecure AWS Lambda function could be used as a pivot point to get access into your environment. They’ve also gone deep in-depth with a bunch of other approaches to how DevOps and security are inextricably linked. To learn more, visit sysdig.com and tell them I sent you. That’s S-Y-S-D-I-G dot com. My thanks to them for their continued support of this ridiculous nonsense.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. My guest today has been on this show before, generally at a previous point in her career where she was making a transition. That time, she was leaving AWS, as happens to awesome people a fair bit of the time—more than it potentially should—and going to work at H2O.ai, a company that does some sort of machine learning thing that I can’t be bothered to remember offhand. I talked to her again, as she has just left that company to start her own thing. Ana Visneski is the Chief Chaos Coordinator at Mirewif. Ana, thank you for joining me yet again.

Ana: Oh, I mean, how could I not when you’re the one who got me to get off my butt and actually start my own company?

Corey: What’s fun is that your company is a crisis communications firm, and first that’s definitely useful for me because I do put the ‘crisis’ in ‘crisis comms,’ let’s not kid ourselves.

Ana: You’re not wrong. [laugh].

Corey: But I’m also your first customer.

Ana: Mm-hm.

Corey: And you’re in one of the harder niches to get people to stand up and say, “Yeah. Oh, yeah. Can I get a testimonial on this?” “Absolutely not. We hired you because we did something horrible.”

And that’s not really how I tend to view crisis comms. I mean, it’s sort of a similar problem to what I had when I started The Duckbill Group of, “Hey, can I use you as a testimonial about your horrifying AWS bill?” “No.” And I understand how it looks, which is not the reality of it. And in time, I found ways to get people to slap their logo on their website. But I want to be the first logo, the fact that I have a platypus associated is just a nice bonus.

Ana: Absolutely. You will be the first logo when I finally get around adding logos. The interesting thing is, that it’s not just crisis comms that I’m doing with the company. I also do threat assessment, violence assessment, so risk analysis, basically, on if you have an employee that might be a risk or, for some of my video game or gaming companies, if you have someone in your fan organization that is a potential risk.

I also do crisis management planning. So, I will put together an operational plan—similar to what I built when I was at AWS—a top to bottom, this is how you run a crisis to make sure your people don’t burn out, make sure your leadership is aware of what’s going on and gets the proper daily briefings, that sort of thing. And then lastly, I’ve actually been doing some consulting with governments on their disaster response technology needs. So, there’s a lot of different aspects to it.

Corey: Yeah, to be very clear, none of those things are things that I have roped you in for. I don’t have employees that I’m looking there with, “Oh, if they blow their stack this is going to be a disaster.” Like that is not the nature of the work we’re doing together. What we’re doing is more along the lines of, “Okay, great. I have a bad tweet that blows up. How do I handle this without, ‘All right, pass me that shovel. We’re digging this puppy deeper. Now, okay. Holes dug nice and deep. Let’s work on the edging details a little bit.’”

Ana: [laugh]. Yep.

Corey: It’s the, “How do I avoid making things worse in moments of crisis?” And we’re building plans for things that I hope to never need around things like data breaches, like, the stuff that every business should have a plan for. Because when disaster strikes, as it tends to in various ways, I don’t want to be sitting here flipping through the Yellow Pages for, “I’ve messed up.” Like, I don’t know what section that would be in. Having a plan ready to go is important.

Ana: I would say it’s actually critical.

Corey: Yeah.

Ana: So, that’s the thing is, unfortunately—and as Covid taught a lot of people—having that plan in place before things go wrong before the shit hits the fan, is what’s going to save you or not. It’ll save you millions of dollars, it’ll save your employees, and it could potentially save lives. And so what I think a lot of companies have finally figured out is, “Oh, wait. We weren’t ready for Covid. We actually need to be ready for the next thing.”

But I also teach crisis communications for the communication leadership program for the University of Washington; it’s a graduate program. You’ve been a guest speaker there. You were one of the favorite guest speakers. And there I tell them all the time is that you have to plan. The two critical things before anything even starts is planning and trust.

If you don’t have plans in place on how you’re going to do things, you’re going to have people running around like chickens with their heads cut off going, “Oh, what do we do?” And someone’s going to do something that makes it worse—inevitably—with the best of intentions. And then the other thing is, if your audience, if your customers don’t trust you to be doing the right thing in the first place, then no amount of planning is going to help from that deficit.

Corey: It also, in my experience working with you, comes down to avoiding putting your foot in your mouth with the best of intentions.

Ana: Yes.

Corey: Heaven forbid if you have an employee pass and tweeting out something like, “We are heartbroken to announce the loss of our dear friend and colleague, [Shtephen 00:06:45]. Also, we’re hiring.” Like, make sure you don’t wind up coming across as the worst example of humanity. It’s the basic stuff.

Ana: Even more than just that basic of don’t put your I’m hiring—because you saw that tweet that was going around with, “So-and-so has passed. Please mourn off the clock.” Whether that was a joke or not—and it’s up for debate if it was real or not, like—

Corey: We’ve all known people who would have said such a thing and it would not have been a joke.

Ana: Exactly. But the other thing is, it’s not even just that. It’s knowing the timelines for notifications. So, for example, there should be at least a 24-hour next-of-kin notification window, where if someone is passed, the friends and family grieving can be notified. The last thing you want is a friend of Shtephen to find out that he died because you tweeted about it. That is traumatizing.

So, you actually have to have a plan in place of, you’ve received notification from Shtephen’s wife that he has passed. Obviously, you’re going to be offering her your support. Say, “Hey, here’s the things we can offer you to help.” You have, you know, your package of, like, here’s the ways we can help you. But then you also say, “Can you let me know when it’s appropriate for me to tell the other employees?”

Because the moment you start telling employees—this recently happened; a friend of mine in the Coast Guard passed, and unfortunately, some others found out about his passing because someone posted about it on Facebook. That is not the way you need to find out. So, it’s not even the blatantly obvious things, like, “Oh, hey, don’t post about hiring,” it’s also just the order in which you notify so that things don’t leak.

Corey: I didn’t even know Shtephen was married. I mean, what kind of—

Ana: [laugh].

Corey: —crappy employer am I here? Yeah, it’s the human side of it.

Ana: Mm-hm.

Corey: And that’s one of the things I’ve always admired about you. It’s—and again, when I started doing all these nonsense things, I had a circle of friends that I could run things past of, “Hey, is this tweet a bridge too far?” And in time, I needed to rely on those people a little bit less because it turns out that I have a pretty good eye for what’s going to make people feel bad. And that’s really the only thing I care about is if it makes someone feel bad, then I’m not thrilled with the tweet most of the time.

And I figured out where that line lies. And then I got loud and big enough on Twitter where I started having to think about it again, where, all right, I know it’s not mean, but I’m going to hear about it. Is the juice worth the squeeze? And the reason I like working with you on things like that is I’ve grown well past the point where I’m comfortable asking people to volunteer for basically what amounts to something of my own brand-building exercise. Paying people for advice has always been something that I’m a big fan of, and now I’m able to do that and have a professional way.

And I don’t think you’ve ever once been wrong. There are times you’ve given guidance that I have not followed, but that’s what you see anytime you’re talking about someone a downside, risk side of the business. That’s the entire function of an attorney for a business is to identify risk. If you start letting attorneys, for example, my wife, great attorney, great wife, wound up—

Ana: And very tolerant human being. [laugh].

Corey: Oh, extraordinarily—living saint. But she wound up editing a proposal that I was going to send out—back when I was independent—once. And I looked at it and she’s like, “Oh, well that could go wrong, and that could go wrong and no, we’re going to change that and the rest.” It’s like, this is—I understand where you’re coming from, but this is a sales document. And it was for a proposal, it was something like $7,000 back then.

It’s like, worst-case scenario, I’m a nice person, I will fall over myself apologizing and give them a full refund. The end. That sort of caps my downside risk here, if they want to be obnoxious and go to court, well, I’ve been doing this for three months, I guess I’m shutting down the LLC because that’s been sued into oblivion. I’m getting a real job. Like that was the risk mitigation there.

She’s used to doing risk analysis for a company with 250,000 employees, and yeah, they have more to lose than I do in those things, so I get it. But you don’t generally have lawyers on your sales team that are proactively over-promising things, for obvious reasons. At least—because there’s no way to get a salesperson disbarred. I’ve checked.

Ana: Of course you did. When I’m teaching class, one of the other things I do is I actually have some lawyers come in and talk. And the reason is, I learned this one when I was in the Coast Guard, and I was running District Eight. So, it’s basically the entire Gulf Coast and all the way up the Mississippi to the Canadian border. So, all of the units contained in that area, I was in charge of their media relations, their community relations.

And this was, like, right after Katrina. I learned pretty quickly that having a very good relationship with my lawyer—so the head of legal—it made us a one-two punch that was unbeatable because I could look at it from the human empathy, communication, subtext aspect, and he’d look at it from the legal aspect, and the two of us would be like, “Okay, you can do this legally, but here’s the impact of it if you say it this way, or if you do this.” Or, “Ehh, don’t do this one, legally.” Like, it’s just a great thing. But risk analysis, from my perspective versus a lawyer’s, are slightly different.

I do, of course, talk to lawyers, obviously, a lot, and look at the legal side of stuff. But a lot of what I’m looking at is perception, subtext, potential pitfalls. You and I’ve had many conversations, and you know me well enough to know that most of the time I’m giving you guidance, but if I see one more, I’m like, “Absolutely not. Do not do that.” I will lean into it so heavily, and be like, “Corey, here’s the eight ways this is going to go badly for you. You’re going to end up in The Times for bad stuff.”

Corey: And you say that so infrequently that I definitely pay attention when you do. I don’t always listen, I mean, [crosstalk 00:12:14] I wound up posting that Andy Jassy birthday video. But you know—

Ana: I helped with that video, though. [laugh].

Corey: —you were instrumental behind that video. Thank you for that.

Ana: You’re welcome. But that’s—so what’s fun about working with you, and different than my other clients is there are these moments where I get to also express my weird sense of humor, you know, where it’s just like, calling Jeff Bezos, a space cowboy. Those moments of getting to find—help you with that line. Because I have that same sense of humor line and I don’t get to express it a lot with my other clients because most of them are very, very serious bidness. And not to say your business isn’t serious, but you yourself are almost—

Corey: But we do have fun with it.

Ana: —never serious. Exactly, exactly. And that one, like, I really enjoy that aspect of it. But with a lot of the other stuff, it is incredibly serious. And like the risk analysis that your wife does, versus the risk analysis type I do, I’m actually looking at emotional stuff.

So, when we’re talking about acts of violence, for example, acts of violence are, almost to a one, about power. So, what I do is I actually sit and look at okay, this person is lashing out. What power dynamic has them wanting to lash out? So like, if you look at a lot of the school shootings, it’s about kids who feel bullied, they want to regain power by showing they have power or the guys who write their manifesto about hating women, et cetera, et cetera. So, it’s always about a power dynamic.

So, it’s not about, is it legal to go in and shoot the office? It’s clearly not. But has the system taught them that they can push the line far enough that this sort of behavior, they might get famous for it? Or might get away with it? And then how do you mitigate that particular power dynamic? And so that gets real tricky. And luckily, with you, I have not had to deal with that one.

Corey: For better or worse, I come out from a good place to place a good intention. I’m trying to imagine if I just said, “To hell with it,” and decided to just take off the gloves and be a complete bully every time I felt like it. I could do some damage at this point. But… no.

Ana: You could, but the thing is remember what I said at the very beginning: It’s about trust. What has made you so very successful, what has made you so good at what you do is you’re very intentional and very careful. Not to say you’re not a pain in the ass. I will agree with some—

Corey: And I do get wrong. Let’s be clear. I’m no saint.

Ana: Oh, no, no, no. No. You’ve gotten stuff wrong, but you immediately apologize for it. So, when I’m talking about this from a space of trust, it’s not that you’re not obnoxious; you totally can be.

Corey: Extraordinarily so.

Ana: You can totally be a snarky pain in the ass. Like I said, your wife is a saint. And sometimes—like, we were talking about recently, backing off on mocking people for working for Facebook because you and I both saw what it did to Chloe. And it’s just not cool to do that to someone who’s making a career choice, whether we agree with it or not. I personally have companies I would never work for. You and I have discussed contracts—not with you, but contracts I wouldn’t take. Me personally, it’s in my contract, I will not defend someone who is a sexual harasser or sexual assaulter. Like, I won’t defend them. If they do #MeToo stuff—

Corey: Mm-hm. The way that we’ve codified that—

Ana: —I won’t do it.

Corey: —here is generally speaking—and this is a truism, I would encourage everyone in business to consider is, if you don’t respect a client’s business, you probably should not take their money. And—

Ana: [laugh].

Corey: —that leads to a lot of things.

Ana: Yeah. I wish that was more common. [laugh].

Corey: Yeah. It’s—and again, I’ve never once shamed a company for this. I have declined to work with a number of companies in different capacities. And I’ve never been very open about this because I don’t want companies to be listening to this and think, “Ohh, we sell ads. He might not want to work with us, so we’re not going to reach out.” First, I will never mention, name, or drag anyone publicly.

Ana: Oh, yeah. Same.

Corey: Secondly, there’s no such thing as any saint in these industries.

Ana: Oh, no.

Corey: I’m not talking about, “Oh, you display ads to people? [tsking noise].” No, I’m talking about, “You make landmines.” Let’s be clear here. This is a whole other side of the universe. And I still never drag the companies that I declined to work with, in public, for having the temerity to reach out. Just seems like it’s the wrong incentive structure if I start down that path.

Ana: I was just talking to a client that I firmly believe we’re at a pivot point in the way businesses are run. I was calling 2022 the Year of Transparency. And the reason I’m saying that is because in the last couple years with people working from home, with Covid, with Black Lives Matter, with all the stuff that’s been going on in the world, and then, like, Activision Blizzard, and the lawsuits, and pay disparity, and Paizo unionizing—Paizo is a tabletop company that makes Pathfinder RPG—

Corey: Mmm.

Ana: —you know, all these companies. So, we’re starting to see the game industry see unionization, we’re seeing Starbucks employees want to unionize. People are not going to accept, “No comment,” anymore. They’re not going to accept, “We’re just not going to answer this.” And I can already see your brain ticking on who you’re about to—I know where you’re thinking.

But my point is, when I’ve been talking to some of them, “I’m like, you have to be prepared that the old-school mentality of people not sharing their pay, like, not sharing how much they make compared to the person sitting next to them, that’s gone.” People share that information now. There are companies where they are having spreadsheets. Now, one thing I did like about AWS was I always knew, like, my peers and I were encouraged if we want—my manager was awesome—my first manager was like, “If you guys want to talk about what you’re making, go ahead.” And I was able to find out that because I had the masters, and more experience, and all this other stuff, I was actually—in my level group—the highest-paid one, even though I was the only woman at first. That’s pretty cool to know.

Corey: That’s the kind of story that never makes the rounds.

Ana: Well, and the thing is, we’re not going to see people accepting obfuscation anymore. I think that’s done. It’s too easy to share information now for companies to think that their dirty laundry isn’t going to come out, to think that they can lie and get away with stuff. As you know, we’ve talked about this a bit, I’m actually working on a book with a comic book artist—I didn’t get his permission to say his name, so I’m not going to say it yet—and it’s literally a picture book on how to not screw things up in today’s digital media age when it comes to how you communicate with people. It’s called Oh, Noes: A Picture Book for Execs. [laugh]. Um, but you know, you got to focus on the fact that people aren’t going to accept obfuscation and lies anymore. They’re not going to accept, “Oh, we’re the company. We’ve got your best interests at heart.” It’s not how it works anymore.

Corey: That’s what I see in this entire industry, where there’s this idea that we’re not going to say anything, we’re just going to do our thing and not comment on any of these things. Which, okay, it’s a strategy. But customers and the community and loud obnoxious—

Ana: They talk to each other.

Corey: —people on Twitter are going to comment in your absence. And that becomes a problem.

Ana: Have you seen the movie—what is it?—John Tucker Must Die?

Corey: I have not.

Ana: It’s a movie about three girls at the same high school who find out the guy is dating all three of them, and how they plot to destroy him. And every time I see one of these things happen where a big tech company—or any company—doesn’t say anything, but then their customers start talking to each other going, “Wait a second,” I always think of that movie. And it’s like, you can’t think that people aren’t going to talk to each other anymore.

Especially once you get huge. When you’re looking at these big, big companies, people want to take you down. Like, they’re over this idea of monopolization and this idea that you can do things and there’s no accountability. So yeah, I’ve been calling this the Year of Transparency because I think we’re going to see huge shifts in what is and isn’t okay to hide from your customers. Trust is your most valuable asset. And it can be lost in seconds.

Corey: It’s the easiest thing in the world to get, and it’s incredibly easy to lose it, and almost impossible to regain it once you’ve lost it.

Ana: Yes. And I think my students get sick of me saying this because I say it every week: “Trust is easy to get if you do it right, but you got to do it right.” You actually have to be honest, you have to, you know—and I’m not saying share secrets. But you can be—like, a good example with AWS is, they do great COEs after they have a big splat. You know, 2017, when they had a service disruption, and the latest ones, like, they do a good COE. Being able to rely on that sort of thing is critical.

Corey: For me, it’s one of the things that we do here just because of the sensitive information with which we are entrusted, and the way that we operate in the industry, we hold ourselves to a bar that is pretty similar to what you’ll see in regulated industries and the rest. I periodically disclose all of my investments, which is nowhere near as interesting as most people would think.

Ana: [laugh].

Corey: I make it clear exactly where my interests are. This is the reason we have no partners with any company in this space, just because it is the perception of conflict of interest is huge. I mean, half our consulting business is doing contract negotiation on behalf of customers, with AWS directly. As soon as it comes out that we have a back channel deal with someone, everyone’s going to question what’s going on. It’s easier never to enter into those engagements rather than having to try and back-walk it later. No. Does that leave opportunities on the table? Sometimes. But I think this is the better long-term play if I can think beyond next quarter's numbers.

Ana: Yeah, absolutely. And that’s, like, similar for me is that I have to be mindful of not taking contracts with companies that are in conflict with each other. And I don’t mean conflict like they’re at war, but like, where my working with each of them puts me in a position where there could be questions on who my loyalties are to.

Corey: On the sponsorship side of our business, we refuse to do anything that even looks like an exclusivity contract, of, “All right. None of our direct competitors will be allowed to sponsor for a fixed period of ti”—sure, if you buy out the ads you don’t want them to take, I guess, sure. But you don’t get editorial control, either. It’s the same approach: You can buy my attention, but never my opinion. Paying me does not make me say nicer things about you, directly.

It does force me to look more closely into what your company does, and no one’s purely good or purely evil. I will talk more about what I see, good and bad. That is the nature of what you get with me, and that is something that I don’t think a number of folks realize, out of that ecosystem.

Ana: Well, there’s a level of professional maturity that goes with taking criticism. And when you have worked on something for a very, very, very long time, and it is your baby and you’re getting criticized, it can be natural to have an emotional response. And that’s something that, as a crisis communicator, I look at. Are the attacks coming in—and attacks, or commentary, or negative press—is it coming in, in an emotional way, like, what’s happening is there’s been a nerve hit because there’s an emotional investment in whatever’s going on? Or is it an impact of concern over finances, concern over jobs? So, there’s different reasons why people will react and things. And that’s one of the things I have to always keep in mind when I’m looking at stuff. As you well know. We’ve had many conversations about this. [laugh].

Corey: This episode is sponsored by our friends at Oracle HeatWave is a new high-performance query accelerator for the Oracle MySQL Database Service, although I insist on calling it, “My squirrel.” While MySQL has long been the worlds most popular open-source database, shifting from transacting to analytics required way too much overhead and, ya know, work. With HeatWave you can run your OLAP and OLTP—don’t ask me to pronounce those acronyms again—workloads directly from your MySQL database and eliminate the time-consuming data movement and integration work, while also performing 1100X faster than Amazon Aurora and 2.5X faster than Amazon Redshift, at a third of the cost. My thanks again to Oracle Cloud for sponsoring this ridiculous nonsense.

Corey: One of the things that I always admired about you—and I have never once incidentally tried to change this in any way—but you have never leaked confidential information to me about anyone or anything. And to be clear, I have never asked. Back when you’re running launch operations at AWS, I don’t want to know things that are coming out if I can avoid it because then it gets very challenging for me to remember what I can talk about versus what I can’t. My insight into AWS product roadmaps is not much better than anyone else in the industry. I just pay attention and I have a knack for being able to see what’s coming.

But because of the perception that I have the inside track, I don’t break the news; I don’t create the news; I just talk about what other people have already written about publicly. It’s safer that way for me, and I’ve always appreciated your ability to respect confidentiality because for stuff like this, it matters more than anything else.

Ana: Absolutely. My confidentiality is huge thing. I just don’t talk about stuff. And in fact, like, my husband doesn’t even know who half my clients are. He knows the number of clients I have, but he doesn’t know who I’m working with. And that’s because, you know, I don’t need him to know. And it’s a confidentiality thing. And you know, spouse, you’re my husband, I have ten clients. And that’s what I’ll say. You know? He knows about you, obviously.

Corey: Well, I should hope so. He’s lovely. I was at your wedding, lovely though it was.

Ana: That is true.

Corey: So, one other thing that you’re in the process of launching as we speak is apparently your own podcast. Loathe though I am to drive people to the competition, tell me about it.

Ana: [laugh]. It’s not actually competition, and we do have to give you credit for the name. One of your superpowers is giving really funny, punny names to just about anything. Next time we get a pet, I’m going to be like, “I want a pun that goes around this. What can I name the dog?”

Corey: And how long did it take me to name your podcast?

Ana: God, like, two minutes. It’s so annoying because I’d been—

Corey: It took that long?

Ana: You let me finish typing.

Corey: Yeah, that was nice of me, I thought.

Ana: Yeah, you let me finish typing, and then you’re like—okay, so not even two minutes. Like, a minute.

Corey: It’s not that I’m that good at naming things. It’s just that I’ve never worked at AWS, and people who are so bad at it, that when some—they just encounter someone who’s average with these things, we look like wizards from the future.

Ana: So yeah, we’re launching a podcast called [Disasterpiece Theater 00:26:15]. And it’s actually a podcast where we’re going to have subject matter experts from NASA, from medical fields, cybersecurity folks, we’re going to actually have a shark expert so we can talk about The Meg and how that works. But the whole point of the podcast is taking pop culture movies—so like Jurassic Park, Alien: Covenant, all of these—and talking about how they’d actually work in the real world. How would Alien: Covenant have gone down if these people were trained the way people on ships are trained now? Or The Meg, what would you actually do if you had a shark in that scene where there’s hundreds of thousands of people in the water? How would that actually go? I mean, the shark wouldn’t be three miles long, but same concept.

So, it’s going to be a lot of fun, just kind of going through. One of my favorite guests is my dad. My dad’s going to be talking to us about the movie 2012. My dad’s a naval architect marine engineer. And he and I had the most fascinating conversation after watching that movie on how ships like that would actually be built, and what would happen, and what would have to happen, and the different rules and regulations that would have to change, and how you would actually—like, and his pet peeves with lazy things the writers did. So, it’s going to be a lot of fun. We’re doing it as a short run to see how it sticks. It’ll be eight episodes to start, and then if there’s a desire for more, we’ll do a second season.

Corey: I’m really looking forward to seeing how it comes out. I’d ask, “What’s going on? You’re starting a company and this side project for funsies? What’s the point?” But I started this podcast show not too long after I started what became The Duckbill Group. So yeah.

Ana: What’s funny is, this has all kind of cascaded in weird ways because next month, the company’s—Merewif’s been around a year next month.

Corey: Wow. Hard to believe.

Ana: Which is totally crazy to think about. But I only—I was doing it as a side gig while I was at H2O until October—end of September. So, it’s only been full-time since October. The podcast idea—

Corey: Why now? Why now, though? What drove you—

Ana: [laugh].

Corey: —you went from giant company to start-up to launching your own thing. And you’re launching your own thing in the same way that I launched my own company, which I’m going to shorthand to ‘the dumb way,’ which is right now there is so much constipated capital sloshing around the VC ecosystem, and we both started companies that are absolutely never going to be a VC-scale opportunity because, you know, what can you do with $4 billion in investment? Oh, something monstrous, for damn sure. But there’s no—there’s no good answer to that. But we’re never going to be the VC-scale opportunity.

Ana: [laugh]. I dread to think what you would figure out what to do with, like, $400 million. It’s terrifying.

Corey: Oh, the video would be ridiculous. We’re talking, like, Pixar quality…ridiculousness, making fun of various things in this industry, on a lark.

Ana: Oh, I can imagine. I can imagine. I can imagine a weekly game show with you, too, where you brought in engineers from the different services and ask them random questions, kind of like Jeopardy, but with, like, the floor dropping out underneath them. Then they just get replaced with the next engineer or whatever. Like, they get an answer wrong; they drop through the floor; the next one slides in.

Corey: I like that, yeah. That has legs.

Ana: And this is why you and I are not allowed to come up with ideas together.

Corey: Yeah this is—

Ana: Anyway.

Corey: —what we do to break on us from time to time. Yeah.

Ana: [laugh]. So, the timing was a couple things. One, I’ve wanted to do this company since I got out of the Coast Guard. Like, it’s something I wanted to do, but I needed to get more private experience. Because up until 2016, all of my experience was public sector. It was military, it was Coast Guard.

So, while I’d worked with—in disasters, I had been side-by-side with BP dealing with that disaster and all sorts of stuff, I didn’t actually have the experience myself. And I kept going, “Oh, well. I’ll do it eventually. I’ll do it eventually. I’m not ready yet. I’m not ready yet.”

And then literally, you and one other person were like, “No, I literally need you to set this up right now because I need your help with something and I need an official way to pay you.”

Corey: It seemed like the right thing to do. Yeah. Yeah.

Ana: So, it was, like, “Okay.” And the name Merewif actually means ‘sea witch’ or ‘siren’ in Old English. Little known fact: I majored in English specializing in medieval and ancient literature when I was in college.

Corey: That explains your depth of insight into the AWS documentation.

Ana: [laugh]. Yeah. I can read like nobody’s business. And so, in traditional stories, a lot of times, the hero will go to a witch or a sea witch for advice, or for knowledge, or for medicines, or whatever. So, it kind of tied together the fact that I was in the Coast Guard—so I’ve always been around oceans—my Old English, Middle English background.

And yeah, it just—the name made sense to me. So, it was like, “Well, I have a name now. Let’s just do it.” And so I did it. And then as the year went on, I started getting a lot of interest, different friends in the industry found out what I was doing, or they found out through a friend, an alumni classmate of mine pinged me going, “Hey, this company really needs your help. Can I do an introduction?” I said, “Okay.”

And so it started taking off. And so by September, I was like, “Well, if I can get a couple things lined up, I’m going to have too much to do with the job I love, which is Merewif, to stay at a day job that I’m like, ‘Ehh. It’s a job.’” And it’s been incredible. Like, it’s busy. It sometimes means waking up at two in the morning to see what you’re up to.

Corey: It happens, sometimes. To be clear, that is out of your own choice. The beautiful thing about my business is that it’s strictly a business hours problem.

Ana: Yes, except I knew that the video was launching today, and I wanted to take one more scrub on it to make sure that [laugh] there wasn’t anything over the line.

Corey: Yeah. We go right up to it, but try not to cross it.

Ana: Yes. And so—and that’s the killer thing is, like, I’m loving every day. Like, it’s crazy. It’s different things. I do hate being my own finance department. But you know.

Corey: Fractional CFOs are one of our first strategic hires that we made here, and it was a bit of a stretch, and it’s a, “We think we can afford it because, Dan”—who’s been a guest on this show—“As our CFO says we can, and that’s sort of his job, so all right. Let’s see what happens.” And sure, it’s great way to fail if he’s not good at his job, but he was right. And it has been an absolute Godsend just for the things I don’t have to worry about that have been taken off of my head, are—it’s like not having to plan a wedding anymore. That level of relief.

Ana: [laugh].Oh, yeah. Covid messed up my wedding, too. So, that ended up being in our backyard. But you know, at the end of the day, every day I’m doing work that I’ve spent my whole career becoming really good at, and becoming an expert at, and being able to talk with [countries 00:32:30] that can’t necessarily afford to hire someone like me full time, but to be able to walk them through, “All right, here’s the cloud technologies that are available for you, but you’re also going to want to have, for example, a snowball edge in your area because you’re going to lose connectivity.” And, “Oh, hey, talk to the guys over at Project OWL.”

It’s a cool one if you haven’t looked at it. They’re basically these floating little—they look like little ducks; well, the original versions of them did—and they basically allow—they’re WiFi repeaters in some ways, where they float. So, if you disperse them in an area where disasters happen, even if it flooded, it’s going to keep that wireless network up and available in that entire area, for everyone who’s impacted. Which is a huge problem in the last mile. So, getting to do this stuff that I love anyway, it was just time.

And I’m loving teaching at the UW. I’m back at the program I actually graduated from. And this will be of no shock to you, at some point in the near future, I’m going to be applying to do my PhD. It’s been a goal of mine since I was little to be the first PhD in my family. Were weirdly competitive about very strange things.

Corey: I will be extremely disappointed if your dissertation does not feature the word ‘shitposting,’ and of course, a link to something that cites my work.

Ana: Actually shitposting could end up in there because what I really want to study is the impact of emerging technologies, including social media and things like that, and how they’re impacting the ability of responders to have a common operating picture. So, it’s clouding the ability. So, a common operating picture is how the Coast Guard and the Fish and Wildlife and the local fire department all know what’s going on when a disaster happens, right? That’s great, but they now all have separate systems. And if you think the local fire department or the local fisheries guys have the same level of security as, say, the Coast Guard does on their systems, they don’t.

So, how do you get them into the same common operating picture? And then what happens if it’s a hurricane, and you have people tweeting pictures of the hurricane, and they’re not even in the area from the hurricane? So, you have all this additional noise, you have all these additional security needs that weren’t there, say, during Katrina, when we were doing everything by, like—no joke—a lot of faxing and text messaging and driving things back and forth. How do you deal with that? So yeah, that’s actually what I’m looking at doing.

So yeah, shitposting might end up in there as a what do you do when you’re in a disaster and you have shitposting cluttering up your mess? So yeah, that’s what I’m hoping to do at some point. But I’ve got so much work right now with Merewif that, right now, I don’t have time to get the PhD. [laugh]. So.

Corey: Industry and academia tend to be a little on the different side. And for what it’s worth, like, there are a lot of companies doing PR, crisis comms work, et cetera, et cetera. The reason that there was really—this was one of those no-bid contracts because you understand this industry in a way that few people do. You’ve worked within it, you understand the dynamics within it, as well as adjacent industries like gaming, for example. Having someone who understands the moving parts of an industry, who the major players are and how that all fits together, it’s something that you can’t take some random comms firm off the street and expect them to understand it in the evolving way that social media, among others, has really shifted the entire narrative. So, I don’t know of anyone else who’s doing it the way that you do. They’re certainly not talking about it the same.

Ana: Way. There are a few firms that do something similar, but they’re bigger and they have a lot of people and they’re not as specialized as I am. So, they have an idea of it, but they’re not necessarily from that industry. Or, you know, I’ve been playing video games since I was—what—ten. And I’ve been very involved. I do panels about women in the military, and how we’re represented in video games and comic books, I do those quite often.

Actually, real quick, that reminds me back to the PhD thing.

Corey: Of course.

Ana: The other reason I want to get the PhD is because, as a woman, having that extra boost of not only have I been doing this for—oh God, almost 20 years; that makes me feel really old—almost 20 years, but I also have a PhD in this specific technique. In order to get this PhD, I have to convince a university to let me combine an IT PhD, like, either an information technology or an IS tech—like, a science PhD and a communication PhD into one. There is no school that quite offers what I want, so I’m going to actually have to combine them. But I will say that one of the other reasons I really want to do it, other than the fact that I get to look at my little brother—who you know—and go, “Pttht, I got it first,” is because as a woman, it does give me one more way to keep the door open that my male counterparts don’t necessarily need. And as you know, in this industry, that’s a lot. I mean, it’s not easy being a younger-looking blue-haired woman who’s like, “Hi, I know my shit.”

Corey: Meanwhile, I am presumed competent in a way that people who aren’t over-represented are not. And when I say something, it is presumed true, as opposed to being nibbled to death by ducks with, “Well, can you back up that assertion?” Because sometimes, no. I’m speculating, but I am presumed to be right as a default.

Ana: Yep.

Corey: And people love to say that, “Oh, yeah, privilege isn’t really a thing.” Let’s be very clear here. I did have to build a lot of the stuff that’s here. None of this was handed to me. But I didn’t have a headwind at fighting against me every step of the way the I would have if I didn’t look like this.

Ana: One of the things I’ve joked about a lot is my being a veteran, has actually helped me with some of those headwinds because there are assumptions made about my personality—[laugh] the fact that I’m blunt, the fact that I—

Corey: No.

Ana: —tend to be very straightforward. And I believe my very first meeting with Ariel Kelman when he was a VP at Amazon—at AWS—was, in one of the meetings, the very first one was the words, “Are you shitting me?” Came out of my mouth over something. [laugh]. Could help it; just came out of my mouth.

I am very good at filtering when I need to, but in that moment, whoof, I couldn’t have. So, being a veteran does help a bit because there’s some personality assumptions that other women deal with the, “Oh, she’s a bitch.” With me. It’s, “Oh, she’s scary because she was a veteran.” I’m like, “All right. [laugh]. Cool. We’ll lean into that. We will tell you this has been my personality since I was five. We’ll let you think it was the Coast Guard that made me this way.”

Corey: You joined early. Got it.

Ana: Oh, totally. Joined at five. Well, my dad was Coast Guard, so let’s just count that. I grew up in the Coast Guard.

Corey: I just never grew up. It was easier.

Ana: You know, when they’re going to let you drive a 378-foot ship, you kind of have to grow up a little bit.

Corey: One would hope anyway.

Ana: [laugh]. Well, and I mean, you know, there’s the other factor is that, you know—actually, in my AWS interview, I think I scared my Bar Raiser by telling one of these stories—there were times where when I made a decision, someone could get killed if I was wrong.

Corey: So, that does happen at Amazon scale, but less frequently than it does in the armed services.

Ana: Well, yeah. I mean, there it’s you’re literally being dumb and leaving people in place in front of a tornado, which I’m not going to get into. I’m very—

Corey: Or a power bus is—a safety isn’t put on and someone gets electrocuted. But it’s always small-scale stuff, not—it’s not as common.

Ana: Yeah. And when you’re doing—like, I was a search and rescue controller, and I had to know the area I was operating in the winds, the potential risks, what type of vessels were in that area, and then we had a computer software called SAROPS that helped me search. But, like, growing up in an industry where if I screwed up someone could die gives you a completely different perspective on a lot of things.

Corey: Compared to that, there is no stress in the computer industry. There really isn’t.

Ana: I used to joke at launch when people were freaking out—and I told Ariel this once and I thought he was going to snort his coffee—but we were sitting there and people were like, “Oh, my gosh,” for re:Invent I was like, “Is the building flooding?” “No.” “Is it on fire?” “No.” “Is anyone shooting at us?” “No.” “Okay, cool. Chill out. [laugh]. It’ll be okay.”

Corey: Yeah, “You can weather some mean tweets. I promise. It’ll be okay.”

Ana: “Deep breaths.” But you know, at the same time, on the empathy scale is understanding that not everyone has that experience. So, that’s the other thing that’s critical to understand as a crisis communicator or as a leader of any kind, is that the stresses and crazy things I’ve been through have made me who I am. The stresses and crazy things you’ve been through have made you who you are, right? Well, what you find—what will trigger your brain to go, “This is fight or flight. Oh, my gosh, this is terrifying. Oh, gosh, I could”—you know, for some of these people at re:Invent, “Oh, my gosh, I could lose my job. If I lose my job, I can’t feed my family.”

So, even though I don’t panic because I’m like, “Meh, no one’s shooting at me. Cool.” Understanding that for the person next to them, they could physically be having that response of fight or flight is a critical part of leadership and crisis comms. You know, I think too often people are like, “Oh, my hardship beats your hardship.” Well, yeah.

Not everyone has been in 60-foot seas where they literally bounce off bulkheads and pass a mushroom through their nose because, by the way, you can get that seasick. But it’s true. And if you look at some of the younger people you’re hiring, what they consider as, “Oh, my gosh, this could be a problem.” You’re like, “Well, okay. We’re going to be okay. Take a breath.”

Corey: Perspective is one of those things that comes with experience, for better or worse.

Ana: [laugh]. Yeah, right?

Corey: So, I want to thank you for taking so much time to speak with me today.

Ana: Oh, absolutely.

Corey: If people want to learn more, where can they find you?

Ana: So, I am @acvisneski, on Twitter. And also, my webpage is the—T-H-E—merewif—M-E-R-E-W-I-F dot com. Those are the two best places.

Corey: And we’ll put them in the [show notes 00:42:00], of course.

Ana: Awesome.

Corey: Thank you so much for joining me today. I really appreciate it.

Ana: Oh, happy to. It’s always fun.

Corey: It really is. Ana Visneski, Chief Chaos Coordinator at the Merewif. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an angry comment telling me that this was the worst possible way to find out that Shtephen was no longer with us.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Yiftach

Yiftach is an experienced technologist, having held leadership engineering and product roles in diverse fields from application acceleration, cloud computing and software-as-a-service (SaaS), to broadband networks and metro networks. He was the founder, president and CTO of Crescendo Networks (acquired by F5, NASDAQ:FFIV), the vice president of software development at Native Networks (acquired by Alcatel, NASDAQ: ALU) and part of the founding team at ECI Telecom broadband division, where he served as vice president of software engineering.

Yiftach holds a Bachelor of Science in Mathematics and Computer Science and has completed studies for Master of Science in Computer Science at Tel-Aviv University.

Links:

  • Redis, Inc.: https://redis.com/
  • Redis open source project: https://redis.io
  • LinkedIn: https://www.linkedin.com/in/yiftachshoolman/
  • Twitter: https://twitter.com/yiftachsh

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Rising Cloud, which I hadn’t heard of before, but they’re doing something vaguely interesting here. They are using AI, which is usually where my eyes glaze over and I lose attention, but they’re using it to help developers be more efficient by reducing repetitive tasks. So, the idea being that you can run stateless things without having to worry about scaling, placement, et cetera, and the rest. They claim significant cost savings, and they’re able to wind up taking what you’re running as it is, in AWS, with no changes, and run it inside of their data centers that span multiple regions. I’m somewhat skeptical, but their customers seem to really like them, so that’s one of those areas where I really have a hard time being too snarky about it because when you solve a customer’s problem, and they get out there in public and say, “We’re solving a problem,” it’s very hard to snark about that. Multus Medical, Construx.ai, and Stax have seen significant results by using them, and it’s worth exploring. So, if you’re looking for a smarter, faster, cheaper alternative to EC2, Lambda, or batch, consider checking them out. Visit risingcloud.com/benefits. That’s risingcloud.com/benefits, and be sure to tell them that I said you because watching people wince when you mention my name is one of the guilty pleasures of listening to this podcast.

Corey: This episode is sponsored in part by our friends at Sysdig. Sysdig is the solution for securing DevOps. They have a blog post that went up recently about how an insecure AWS Lambda function could be used as a pivot point to get access into your environment. They’ve also gone deep in-depth with a bunch of other approaches to how DevOps and security are inextricably linked. To learn more, visit sysdig.com and tell them I sent you. That’s S-Y-S-D-I-G dot com. My thanks to them for their continued support of this ridiculous nonsense.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. This promoted episode is brought to us by a company that I would have had to introduce differently until toward the end of last year. Today, they’re Redis, but for a while they’ve been Redis Labs, here to talk with me about that and oh, so much more is their co-founder and CT, Yiftach Shoolman. Yiftach, thank you for joining me.

Yiftach: Hi, Corey. Nice to be a guest of you. This is a very interesting podcast, and I often happen to hear it.

Corey: I’m always surprised when people tell me that they listen to this because unlike a newsletter or being obnoxious on Twitter, I don’t wind up getting a whole lot of feedback from people via email or whatnot. My operating theory has been that it’s like a—when I send an email out, people will get that, “Oh, an email. I know how to send one of those.” And they’ll fire something back. But podcasts are almost like a radio show, and who calls into radio shows? Well, lunatics, generally, and if I give feedback, I’ll feel like a lunatic.

So, I get very little email response on stuff like this. But when I talk to people, they mention the show. It’s, “Oh, right. Good. I did remember to turn the microphone on. People are out there listening.” Thank you.

So you, back in August of 2021, the company that formerly known as Redis Labs, became known as Redis. What caused the name change? And sure, is a small change as opposed to, you know, completely rebranding a company like Square to Block, but what was it that really drove that, I guess, rebrand?

Yiftach: Yeah, a great question. And by way, if you look at our history, we started the company under the name of Garantia Data, which is a terrible name. [laugh]. And initially, what we wanted to do is to accelerate databases with both technologies like memcached, and Redis. Eventually, we built a solution for both, and we found out that Redis is much more used by people. That was back in 2011.

So, in 2021, we finally decided to say let’s unify the brand because, you know, as a contributors to Redis from day one, and creator of Redis is also part of the company, Salvatore Sanfilippo. We believed that we should not confuse the market with multiple messages about Redis. Redis is not just the cache and we don’t want people to definitely interpret this. Redis is more than a cache, it’s actually, if you look at our customer, like, 66% of them are using it as a real-time database. And we wanted to unify everyone around this naming to avoid different interpretation. So, that was the motivation, maybe.

Corey: It’s interesting you talk about, I guess, the evolution of the use cases for Redis. Back in 2011, I was using Redis in an AWS environment, and, “Ah, disk persistence, we’re going to turn that on.” And it didn’t go so well back in those days because I found that the entire web app that we were using would periodically just slam to a halt for about three seconds whenever Redis wound up doing its disk persistent stuff, diving in far deeper than I really had any right to be doing, I figured out this was a regression in the Xen hypervisor and Xen kernel that AWS was using back then around the fork call. Not Redis’s fault, to be very clear. But I looked at this and figured, “Ah. I know how to fix this.”

And that’s right. We badgered AWS into migrating to Nitro years later and not using Xen anymore, and that solve that particular problem. But this was early on in my technical career. It sort of led to the impression of, “Oh, Redis. That’s a cache, I should never try and use it as anything approaching a database.” Today, that guidance no longer holds, you are positioning yourself as a data platform. What did that dawning awareness look like? How did you get to where you are from where Redis was once envisioned in the industry: Primarily as a cache?

Yiftach: Yeah, very good question. So, I think we should look at this problem from the application perspective, or from the user perspective. Sounds like a marketing term, but we all know we are in the age of real-time. Like, you expect everything to be instantly. You don’t want to wait, no one wants to wait, especially after Covid and everything’s that brought to the you know, online services.

And the expectation today from a real-time application is to be able to reply less than 100 milliseconds in order to feel the real-time. When I say 100 milliseconds, from the time you click the button until you get the first byte of the response. Now, if you do the math, you can see that, like, 50% of this goes to the network and 50% of this goes to the data center. And inside the data center, in order to complete the transaction in less than 50 milliseconds, you need a database that replies in no time, like, less than a millisecond. And today, I must say, only Redis can guarantee that.

If you use Redis as a cache, every transaction—or there is a potential at least—that not all the information will be in Redis when the transaction is happening and you need to bring it probably from the main database, and you need to processing it, and you need to update Redis about it. And this takes a while. And eventually, it will help the end-user experience. And just to mention, if you look at our support tickets, like, I would say the majority of them is, why Redis replies—why Redis latency grew from 0.25 millisecond to 0.5 millisecond because there is a multiplier effect for the end-user. So, I hoping I managed to answer what are the new challenges that we see today in the market.

Corey: Tell me a little bit more about the need for latency around things like that. Because as we look at modern web apps across the board, people are often accessing them through mobile devices, which, you know, we look at this spinning circle of regret as it winds up loading a site or whatnot, it takes a few seconds. So, the idea of oh, that the database call has to complete in millisecond or less time seems a little odd viewed purely from a perspective of, “Really? Because I spent a lot of time waiting for the webpage to load.” What drives that latency requirement?

Yiftach: First of all, I agree with you. A lot of time, you know, application were not built for it then. This is why I think we still have an opportunity to improve existing application. But for those applications that weren’t built for real-time, for instance, in the gaming space, it is clear that if you delay your reaction for your avatar, in more than two frame, like, I mean, 60 millisecond, the experience is very bad, and customers are not happy with this. Or, in transaction scoring example, when you swipe the card, you want the card issuer to approve or not approve it immediately. You don’t want to wait. [unintelligible 00:07:19] is another example.

But in addition to that there are systems like mobility as a service, like the Ubers of the world, or the Airbnb of the world. Or any e-commerce site. In order to be able to reply in second, they need to process behind the scene, thousand, and sometime millions of operations per second in order to get to the right decision. Yeah? You need to match between riders and drivers. Yeah, and you need to match between guests and free room in the hotel. And you need to see that the inventory is up-to-date with the shoppers.

And all these takes a lot of transactions and a lot of processing behind the scene in order just to reply in second in a consistent manner. And this is why that this is useful in all these application. And by the way, just a note, you know, we recently look at how many operations per second actually happening in our cloud environment, and I must tell you that I was surprised to see that we have over one thousand clusters or databases with the speed of between 1 million to 10 million operation per second. And over 150 databases with over 10 million operations per second, which is huge. And if you ask yourself how come, this is exactly the reason. This is exactly the reason. For every user interaction, usually you need to do a lot of interaction with your data.

Corey: That kind of transaction volume, it would never occur to me to try and get that on a single database. It would, “All right, let’s talk about sharding for days and trying to break things out.” But it’s odd because a lot of the constraints that I was used to in my formative years, back when I was building software—badly—are very much no longer the case. The orders of magnitude are different. And things that used to require incredibly expensive, dedicated hardware now just require, “Oh yeah, you can click the button and get one of those things in the cloud, and it’s dirt cheap.”

And it’s been a very strange journey. Speaking of clicking buttons, and getting things available in the cloud, Redis has been a thing, and its rise has almost perfectly tracked the rise of the cloud itself. There’s of course the Redis open-source project, which has been around for ages and is what you’re based on top of. And then obviously AWS wind up launching—“Ah, we’re going to wind up ‘collaborating’”—and the quotes should be visible from orbit on that—“With Redis by launching ElasticCache for Redis.” And they say, “Oh, no, no, it’s not competition. It’s validating your market.”

And then last year, they looked at you folks again, like, “Ah, we’re launching a second service: MemoryDB in this case.” It’s like Redis, except bad. And I look at this, and I figure what is their story this time? It’s like, “Oh, we’re going to validate the shit out of your market now.” It’s, on some level, you should be flattered having multiple services launched trying to compete slash offer the same types of things.

Yet Redis is not losing money year-over-year. By all accounts, you folks are absolutely killing it in the market. What is it like to work both in
cloud environments and with the cloud vendors themselves?

Yiftach: This is a very good question. And we use the term frenemy, like, they’re our friend, but sometimes they are our enemy. We try to collaborate and compete them fairly. And, you know, AWS is just one example. I think that the other cloud took a different approach.

Like with GCP, we are fully integrated in the console, what is called, “Third-party first-class service.” You click the button through the GCP console and then you’re redirected to our cloud, Redis Enterprise cloud. With Azure even, we took a one step further and we provide a fully integrated solution, which is managed by Azure, Azure Cache for Redis, and we are the enterprise tier. But we are also cooperating with AWS. We cooperating on the marketplace, and we cooperate in other activities, including the open-source itself.

Now, to tell you that we do not have, you know, a competition in the market, the competition is good. And I think MemoryDB is a validation of your first question, like, how can you use Redis [more than occasion 00:11:33], and I encourage users to test the differences between these services and to decide what fits to their use case. I promise you my perspective, at least, that we provide a better solution. We believe that any real-time use case should eventually be served by Redis, and you don’t need to create another database for that, and you don’t need to create another caching layer; you have everything in a single data platform, including all the popular models, data models, not only key-value, but also JSON, and search, and graph, and time-series… and probably AI, and vector embedding, and other stuff. And this is our approach.

Corey: Now, I know it’s unpopular in AWS circles to point this out, but I have it on good authority that they are not the only large-scale cloud provider on the planet. And in fact, if I go to the Google Cloud Console, they will sell me Redis as well, but it’s through a partner affinity as a first-party offering in the console called Redis Enterprise. And it just seems like a very different interaction model, as in, their approach is they’re going to build their own databases that solve for a wide variety of problems, some of them common and some of them ridiculous, but if you want actual Redis or something that looks like Redis, their solution is, “Oh, well, why don’t we just give you Redis directly, instead of making a crappy store-brand equivalent of Redis?” It just seems like a very different go to market approach. Have you seen significant uptake of Redis as a product, through partnering with Google Cloud in that way?

Yiftach: I would do answer this politely and say that I can no more say that the big cloud momentum is only on AWS. [laugh]. We see a lot of momentum in other clouds in terms of growth. And I would challenge the AWS guys to think differently about partnership with ISV. I’m not saying that they’re not partnering with us, but I think the partnerships that we have with other clouds are more… closer. Yeah. It’s like there is less friction. And it’s up to them, you know? It’s up to any cloud vendor to decide the approach they wants to take in this market. And it’s good.

Corey: It’s a common refrain that I hear is that AWS is where we see the eight-hundred-pound gorilla in the space, it’s very hard to deny that. But it also has been very tricky to wind up working with them in a partnership sense. Partnering is not really a language that Amazon speaks super well, kind of like, you know, toddlers and sharing. Because today, they aren’t competing directly with you, but what about tomorrow? And it’s such a large distributed company that in many cases, your account manager or your partner manager there doesn’t realize that they’re launching a competitor until 12 hours before it launches. And that’s—yeah, feels great. It just feels very odd.

That said, you are a partner with AWS and you are seeing significant adoption via the AWS Marketplace, and the reason I know that is because I see it in my own customer accounts on the consulting side, I’m starting to see more and more spend via the marketplace, partially due to offset spend commitments that they’ve made contractually with AWS, but also, privately I tend to believe a lot of that is also an end-run around their own internal procurement department, who, “Oh, you want some Redis. Great. Give me nine months, and then find three other vendors to give me competitive bids on it.” And yeah, that’s not how the world works anymore. Have you found that the marketplace is a discovery mechanism for customers to come to Redis, or are they mostly going into the marketplace saying, “I need Redis. I want it to be Redis Enterprise, from Redis, but this is the way I’m going to procure it.”

Yiftach: My [unintelligible 00:15:17], you know, there are people that are seeing differently, that marketplace is how to be discovered through the marketplace. I still see it, I still see it as a billing mechanism for us, right? I mean, AWS helping us in sell. I mean, their sell are also sell partner and we have quite a few deals with them. And this mechanism works very nicely, I must say.

And I know that all the marketplaces are trying to change it, for years. That customer whenever they look at something, they will go through the marketplace and find it there, but it’s hard for us to see the momentum there. First of all, we don’t have the metrics on the marketplace; we cannot say it works, it doesn’t works. What we do see that works is that when we own the customer and when the customer is ascertaining how to pay, through the credit card or through the wire, they usually prefer to pay through the commit from the cloud, whether it is AWS, GCP, or Azure. And for that, we help them to do the transaction seamlessly.

So, for me, the marketplace, the number one reason for that is to use your existing commit with the cloud provider and to pay for ourselves. That said, I must say that [with disregard 00:16:33] [laugh] AWS should improve something because not the entire deal is committed. It’s like 50% or 60%, don’t remember the exact number. But in other clouds when ISVs are interacting with them, the entire
deal is credited for the commit, which is a big difference.

Corey: I do point out, this is an increasing trend that I’m seeing across the board. For those who are unaware, when you have a large-scale commitment to spend a certain dollar amount per year on AWS Marketplace spend counts at a 50% rate. So, 50 cents of every dollar you spend to the marketplace counts toward your commit. And once upon a time, this was something that was advertised by AWS enterprise sales teams, as, “Ah. This is a benefit.”

And they’re talking about moving things over that at the time are great, you can move that $10,000 a year thing there. And it’s, “You have a $50 million annual commit. You’re getting awfully excited about knocking $5,000 off of it.” Now, as we see that pattern starting to gain momentum, we’re talking millions a year toward a commit, and that is the game changer that they were talking about. It just looks ridiculous at the smaller-scale.

Yiftach: Yeah. I agree. I agree. But anyway, I think this initiative—and I’m sure that AWS will change it one day because the other cloud, they decided not to play this game. They decided to give the entire—you know, whatever you pay for ISVs, it will be credited with your commit.

Corey: We’re all biased by our own experiences, so I have a certain point of view based upon what I see in my customer profile, but my customers don’t necessarily look like the bulk of your customers. Your website talks a lot about Redis being available in all cloud providers, in on-prem environments, the hybrid and multi-cloud stories. Do you see significant uptake of workloads that span multiple clouds, or are they individual workloads that are on multiple providers? Like for example Workload A lives on Azure, Workload B lives on GCP? Or is it just Workload A is splattered across all four cloud providers?

Yiftach: Did the majority of the workloads is splitted between application and each of them use different cloud. But we started to see more and more use cases in which you want to use the same data sets across cloud, and between hybrid and cloud, and we provide this solution as well. I don’t want to promote myself so much because you worried me at the beginning, but we create these products that is called Active-Active Redis that is based on CRDT, Conflict-free Replicated Data Type. But in a nutshell, it allows you to run across multiple clouds, or multiple region in the same cloud, or hybrid environment with the speed the of Redis while guaranteeing that eventually all your rights will be converged to the same value, and while maintaining the speed of Redis. So, I would say quite a few customers have found it very attractive for them, and very easy to migrate between clouds or between hybrid to the cloud because in this approach of Active-Active, you don’t need the single cut-off.

A single cut-off is very complex process when you want to move a workload from one cloud to another. Think about it, it is not only data; you want to make sure that the whole entire application works. It never works in one shot and you need to return back, and if you don’t have the data with you, you’re stuck. So, that mechanism really helps. But the bigger picture, like you mentioned, we see a lot of [unintelligible 00:20:12] distribution need, like, to maintain the five nines availability and to be closer to the user to guarantee the real-time. Send dataset deployment across multiple clouds, and I agree, we see a growth there, but it is still not the mainstream, I would say.

Corey: I think that my position on multi-cloud has been misconstrued in a variety of corners, which is doubtless my fault for failing to explain it properly. My belief has been when you’re building something on day-one, greenfield pickup provider—I don’t care which one—go all in. But I also am not a big fan of potentially closing off strategic or theoretical changes down the road. And if you’re picking, let’s say, DynamoDB, or Cloud Spanner, or Cosmos DB, and that is the core of your application, moving a workload from Cloud A to Cloud B is already very hard. If you have to redo the entire interface model for how it talks to his data store and the expectations built into that over a number of years, it becomes basically impossible.

So, I’m a believer in going all-in but only to a certain point, in some cases, and for some workloads. I mean, I done a lot of work with DynamoDB, myself for my newsletter production pipeline, just because if I can’t use AWS anymore, I don’t really need to write Last Week in AWS. I have a hard time envisioning a scenario in which I need to go cross-cloud but still talk about the existing thing. But my use case is not other folks’ use case. So, I’m a big believer in the theoretical exodus, just because not doing that in many corporate environments becomes a lot less defensible. And Redis seems like a way to go in that direction.

Yiftach: Yeah. Totally with you. I think that this is a very important—and by the way, it is not… to say that multi-cloud is wrong, but it allows you to migrate workload from one cloud to another, once you decide to do it. And it’s put you in a position as a consumer—no one wants—why no one likes [unintelligible 00:22:14]. You know, because of the pricing model [laugh], okay, right?

You don’t want to repeat this story, again with AWS, and with any of them. So, you want to provide enough choices, and in order to do that, you need to build your application on infrastructures that can be migrated from one cloud to another and will not be, you know, reliant on single cloud database that no one else has, I think it’s clear.

Corey: This episode is sponsored in part by our friends at Vultr. Spelled V-U-L-T-R because they’re all about helping save money, including on things like, you know, vowels. So, what they do is they are a cloud provider that provides surprisingly high performance cloud compute at a price that—while sure they claim its better than AWS pricing—and when they say that they mean it is less money. Sure, I don’t dispute that but what I find interesting is that it’s predictable. They tell you in advance on a monthly basis what it’s going to going to cost. They have a bunch of advanced networking features. They have nineteen global locations and scale things elastically. Not to be confused with openly, because apparently elastic and open can mean the same thing sometimes. They have had over a million users. Deployments take less that sixty seconds across twelve pre-selected operating systems. Or, if you’re one of those nutters like me, you can bring your own ISO and install basically any operating system you want. Starting with pricing as low as $2.50 a month for Vultr cloud compute they have plans for developers and businesses of all sizes, except maybe Amazon, who stubbornly insists on having something to scale all on their own. Try Vultr today for free by visiting: vultr.com/screaming, and you’ll receive a $100 in credit. Thats v-u-l-t-r.com slash screaming.

Corey: Well, going greenfield story of building something out today, “I’m going to go back to my desk after this and go ahead and start building out a new application.” And great, I want to use Redis because this has been a great conversation, and it’s been super compelling. I am probably not going to go to redis.com and sign up for an enterprise Redis agreement to start building out.

It’s much likelier that I would start down the open-source path because it turns out that I believe ‘docker pull redis’ is pretty close to—or ‘docker run redis latest’ or whatever it is, or however it is you want to get Redis—I have no judgment here—is going to get you the open-source tool super well. What is the nature of your relationship between the open-source Redis and the enterprise Redis that runs on money?

Yiftach: So, first of all, we are, like, the number one contributor to the Redis open-source. So, I would say 98% of the code of Redis contributed by our team. Including the creator of Redis, Salvatore Sanfilippo, was part of our team. Salvatore has stepped back in, like—when was it? Like, one-and-a-half, almost two years ago because the project became, like, a monster, and he said, “Listen, this is too much. I worked, like, 10 years or 11 years. I want to rest a bit.”

And the way we built the core team around Redis, we said we will allocate three people from the company according to their contribution. So, the leaders—the number two after Salvatore in terms of contribution, I mean, significant contribution, not typo and stuff [laugh] like this. And we also decided to make it, like, a community-driven project, and we invited people from other places, including AWS, Madelyn, and Zhao Zhao from Alibaba.

And this is based on the past contribution to Redis, not because they are from famous cloud providers. And I think it works very well. We have a committee which is driven by consensus, and this is how we agree what we put in the open-source and what we do not. But in addition to the pure open-source, we also invested a lot in what we call Source Available. Source Available is a new approach that, I think, we were the first who started it, back in 2018, when we wanted to have a mechanism to be able to monetize the company.

And what we did by then, we added all the modules which are extensions to the latest open-source that allow you to do the model, like JSON and search and graph and time series and AI and many others with Redis under the Source Available license. That mean you can use it like BSD; you can change everything, without copyleft, you don’t need to contribute back. But there is one restriction. You cannot create a service or a product that compete directly with us. And so far, it works very well, and you can launch Docker containers with search, and with JSON—or with all the modules combined; we also having this—and get the experience from day zero.

We also made sure that all your clients are now working very well with these modules, and we even created the object mapping client for each of the major language. So, we can easily integrate it with Spring, in Django, and Node.js platform, et cetera. This is called when OM .NET, OM Java, OM Node.js, OM Python, et cetera, very nicely. You don’t need to know all the commands associated. We just speak [unintelligible 00:26:22] level with Redis and get all the functionality.

Corey: It’s a hard problem to solve for, and whenever we start talking about license changes for things like this, it becomes a passionate conversation. People with all sorts of different perspectives and assumptions baked in—and a remembrance of yesteryear—all have different thoughts on coulda, woulda, shoulda, et cetera, et cetera. But let’s be very clear, running a business is hard. And building a business on top of an open-source model is way harder. Now, if we were launching an open-source company today in 2022, there are different things we would do; we would look at it very differently. But over a decade ago, it didn’t work that way. If you were to be looking
at how open-source companies can continue to thrive in a time of cloud, what guidance do you have for him?

Yiftach: This is a great question, and I must say that the every month or every few weeks, I have a discussion with a new team of founders that want to create an open-source, and they asked me what is my opinion here. And I would say, today, that we and other ISV, we built a system for you to decide what you want to freely open-source, and take into account that if this goes very well, the cloud provider will pick it up and will make a service out of it. Because this is the way they work. And the way for you to protect yourself is to have different types of licenses, like we did. Like you can decide about Source Available and restrict it to the minimum.

By the way, I think that Source Available is much better than AGPL with the copyleft and everything that it’s provide. So, AGPL is a pure open-source, but it has so many other implications that people just don’t want to touch it. So, it’s available, you can do whatever you want, you just cannot create a competing product. And of course, if there are some code that you want to close, use closed-source. So, I would say think very seriously about your licensing model. This is my advice. It’s not to say that open-source is not great. I truly believe that it helps you to get the adoption; there are a lot of other benefits that open-source creates.

Corey: Historically, it feels that open-source was one of those things that people wanted the upside of the community, and the adoption, and getting people to work. Especially on a shoestring budget, and people can go in and fix these things. Like, that’s the naive approach of, “Oh, it just because we get a whole bunch of free, unpaid labor.” Okay, fine, whatever. It also engenders trust and lets people iterate rapidly and build things to suit their use cases, and you learn so much more about the use cases as you continue to grow.

But on the other side of it, there’s always the Docker problem where they gave away the thing that added stupendous value. If they hadn’t gone open-source with Docker, it never would have gotten the adoption that it did, but by going open-source, they wound up, effectively, being forced more or less than to say, “Okay, we’re going to give away this awesome thing and then sell services around it.” And that’s not really a venture-scaled business, in most cases. It’s a hard market.

Yiftach: And the [gate 00:29:26] should never be the cloud. Because people, like you mentioned, people doesn’t start with the cloud. They start to develop with on the laptop or somewhere with Docker or whatever. And this is where Source Available can shine because it allows you to do the same thing like open-source—and be very clear, please do not confuse your user. Tells them that this is Source Available; they should know in advance, so they will be not surprise later on when they move to the production stage.

Then if they have some question, legal questions, for Redis, we’re able to answer, yeah. And if they don’t, they need to deal with the implication of this. And so far, we found it suitable to most of the users. Of course, there will be always open-source gurus.

Corey: If there’s one thing people have on the internet, it’s opinions.

Yiftach: Yeah. I challenge the open-source gurus to change their mindset because the world has changed. You know, we cannot treat the open-source like we used to treat it there in 2008 or early-90s. It is a different world. And you want companies like Redis, you want people to invest in open-source. And we need to somehow survive, right? We need to create a business. So, I challenge these [OSI 00:30:38] committees to think differently. I hope they will, one day.

Corey: One last topic that I want to cover is the explosion of AI—artificial intelligence—or machine-learning, or bias-laundering, depending upon who you ask. It feels in many ways like a marketing slogan, and I viewed it as more or less selling pickaxes into a digital gold rush on the part of the cloud providers, until somewhat recently, when I started encountering a bunch of customers who are in fact using it for very interesting and very appropriate use cases. Now, I’m seeing a bunch of databases that are touting their machine-learning capabilities on top of the existing database offerings. What’s Redis’s story around that?

Yiftach: Great question. Great question. So, I think today, I have two story, which are related to the real-time AI, yeah, we are in the real-time world. One of them is what we call the online feature store. Just to explain the audience what is a feature store, usually, when you do inferencing, you need to enhance the transaction with some data, in order to get the right quality.

Where do you store this data? So, for online transaction, usually you want to store it in Redis because you don’t want to delay your application whenever you do inferencing. So, the way it works, you get a transaction, you bring the features, you combine them together, sends them to inferencing, and then whatever you want to do with the results. One of the things that we did with Redis, we combine AI inferencing inside with this, and we allow you to do that in one API call, which makes the real-time much, much faster. You can decide to use Redis just as a [unintelligible 00:32:16] feature store; this is also great.

The other aspect of AI is vector embedding. Just to make sure that we are all aligned with vector embedding term, so vector embedding allows you to provide a context for your text, for your video, for your image in just 128-byte, or floating point. It really depends on the quality of vector. And think about is that tomorrow, every profile in your database will have a vector that explain the context of the product, the context of the user, everything, like, in one single object in your profile.

So, Redis has it. So, what can you do once you have it? For instance, you can search where are the similar vector—this is called vector similarity search—for recommendation engines, and for many, many, many others implications. And you would like to combine it with metadata, like, not only bring me all the similar context, but also, you know, some information about the visitor, like the age, like the
height, like where does the person live? So, it’s not only vector similarity search, it’s search with vector similarity search.

Now, the question could be asked, do we want to create a totally different database just for this vector similarity search, and then I will make it fast as Redis because you need everything to run in real-time? And this is why I encourage people to look at what they have in Redis. And again, I don’t want to be marketeer here, but they don’t think that the single-feature deployment require a new database. And we added this capability because we do see the need to support it in real-time. I hope my answer was not too long.

Corey: No, no, it’s the right answer because the story that you’re telling about this is not about how smart you are; it’s not about hype-driven stuff. You’re talking about using this to meet actual customer needs. And if you tell me that, “Oh, we built this thing because we’re smart,” yeah, I can needle you to death on that and make fun of you until I’m blue in the face. But when you say, “I’m going to go ahead and do this because our customers have this pain,” well, that’s a lot harder for me to criticize because, yeah, you have to meet your customers where they are; that’s the right thing to do. So, this is the kind of story that is slowly but surely changing my view on the idea of machine-learning across the board.

Yiftach: I’m happy that you like it. We like it as well. And we see a lot of traction today. Vector similarity search is becoming, like, a real thing. And also features store.

Corey: I want to thank you so much for taking the time to speak with me today. If people want to learn more, where can they find you?

Yiftach: Ah, I think first of all, you can go to redis.io or redis.com and look for our solution. And I’m available on LinkedIn and Twitter, and you can find me.

Corey: And we will of course put links to all of that in the [show notes 00:35:10]. Thank you so much for your time today. I appreciate it.

Yiftach: Thank you, Corey. It was very nice conversation. I really enjoy it. Thank you very much.

Corey: Thank you. You as well. Yiftach Shoolman, CTO and co-founder at Redis. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with a long rambling angry comment about open-source licensing that no one is going to be bothered to read.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Laura

Laura leads the research program at Jeli.io. She has a Master’s degree in Human Factors & Systems Safety and a PhD in Cognitive Systems Engineering. Her doctoral work focused on distributed incident response practices in DevOps teams responsible for critical digital services. She was a researcher with the SNAFU Catchers Consortium from 2017-2020 and her research interests lie in resilience engineering, coordination design and enabling adaptive capacity across distributed work teams. As a backcountry skier and alpine climber, she also studies cognition & resilient performance in high risk, high consequence mountain environments.

Links:

  • Howie: The Post-Incident Guide: https://www.jeli.io/howie-the-post-incident-guide/
  • Jeli: https://www.jeli.io
  • Twitter: https://twitter.com/lauramdmaguire

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Today’s episode is brought to you in part by our friends at MinIO the high-performance Kubernetes native object store that’s built for the multi-cloud, creating a consistent data storage layer for your public cloud instances, your private cloud instances, and even your edge instances, depending upon what the heck you’re defining those as, which depends probably on where you work. It’s getting that unified is one of the greatest challenges facing developers and architects today. It requires S3 compatibility, enterprise-grade security and resiliency, the speed to run any workload, and the footprint to run anywhere, and that’s exactly what MinIO offers. With superb read speeds in excess of 360 gigs and 100 megabyte binary that doesn’t eat all the data you’ve gotten on the system, it’s exactly what you’ve been looking for. Check it out today at min.io/download, and see for yourself. That’s min.io/download, and be sure to tell them that I sent you.

Corey: This episode is sponsored in part by our friends at Sysdig. Sysdig is the solution for securing DevOps. They have a blog post that went up recently about how an insecure AWS Lambda function could be used as a pivot point to get access into your environment. They’ve also gone deep in-depth with a bunch of other approaches to how DevOps and security are inextricably linked. To learn more, visit sysdig.com and tell them I sent you. That’s S-Y-S-D-I-G dot com. My thanks to them for their continued support of this ridiculous nonsense.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. One of the things that’s always been a treasure and a joy in working in production environments is things breaking. What do you do after the fact? How do you respond to that incident?

Now, very often in my experience, you dive directly into the next incident because no one has time to actually fix the problems but just spend their entire careers firefighting. It turns out that there are apparently alternate ways. My guest today is Laura Maguire who leads the research program at Jeli, and her doctoral work focused on distributed incident response in DevOps teams responsible for critical digital services. Laura, thank you for joining me.

Laura: Happy to be here, Corey, thanks for having me.

Corey: I’m still just trying to wrap my head around the idea of there being a critical digital service, as someone whose primary output is, let’s be honest, shitposting. But that’s right, people do use the internet for things that are a bit more serious than making jokes that are at least funny only to me. So, what got you down this path? How did you get to be the person that you are in the industry and standing in the position you hold?

Laura: Yeah, I have had a long circuitous route to get to where I am today, but one of the common threads is about safety and risk and how do people manage safety and risk? I started off in natural resource industries, in mountain safety, trying to understand how do we stop things from crashing, from breaking, from exploding, from catching fire, and how do we help support the people in those environments? And when I went back to do my PhD, I was tossed into the world of software engineers. And at first I thought, now, what do firefighters, pilots, you know, emergency room physicians have to do with software engineers and risk in software engineering? And it turns out, there’s actually a lot, there’s a lot in common between the types of people who handle real-time failures that have widespread consequences and the folks who run continuous deployment environments.

And so one of the things that the pandemic did for us is it made it immediately apparent that digital service delivery is a critical function in society. Initially, we’d been thinking about these kinds of things as being financial markets, as being availability of electronic health records, communication systems for disaster recovery, and now we’re seeing things like communication and collaboration systems for schools, for businesses, this helps keep society functioning.

Corey: What makes part of this field so interesting is that the evolution in the space where, back when I first started my career about a decade-and-a-half ago, there was a very real concern in my first Linux admin gig when I accidentally deleted some of the data from the data warehouse that, “Oh, I don’t have a job anymore.” And I remember being surprised and grateful that I still did because, “Oh, you just learned something. You going to do it again?” “No. Well, not like that exactly, but probably some other way, yeah.”

And we have evolved so far beyond that now, to the point where when that doesn’t happen after an incident, it becomes almost noteworthy in its own right and it blows up on social media. So, the Overton window of what is acceptable disaster response and incident management, and how we learn from those things has dramatically shifted even in the relatively brief window of 15 years. And we’re starting to see now almost a next-generation approach to this. One thing that you were, I believe the principal author behind is Howie: The Post-Incident Guide, which is a thing that you have up on jeli.io—that’s J-E-L-I dot I-O—talking about how to run post-incident investigations. What made you decide to write something like this?

Laura: Yeah, so what you described at the beginning there about this kind of shift from blameless—blameful-type approaches to incident response to thinking more broadly about the system of work, thinking about what does it mean to operate in continuous deployment environments is really fundamental. Because working in these kinds of worlds, we don’t have an established knowledge base about how these systems work, about how they break because they’re continuously changing, the knowledge, the expertise required to manage them is continuously changing. And so that shift towards a blameless or blame-aware post-incident review is really important because it creates this environment where we can actually share knowledge, share expertise, and distribute more of our understandings of how these systems work and how they break. So that, kind of, led us to create the Howie Guide—the how we got here post-incident guide. And it was largely because companies were kind of coming from this position of, we find the person who did the thing that broke the system and then we can all rest easy and move forward. And so it was really a way to provide some foundation, introduce some ideas from the resilience engineering literature, which has been around for, you know, the last 30 or 40 years—

Corey: It’s kind of amazing, on some level, how tech as an industry has always tried to reinvent things from first principles. I mean, we figured out long before we started caring about computers in the way we do that when there was an incident, the right response to get the learnings from it for things like airline crashes—always a perennial favorite topic in this space for conference talks—is to make sure that everyone can report what happened in a safe way that’s non-accusatory, but even in the early-2010s, I was still working in environments where the last person to break production or break the bill had the shame trophy hanging out on their desk, and it would stay there until the next person broke it. And it was just a weird, perverse incentive where it’s, “Oh if I broke something, I should hide it.”

That is absolutely the most dangerous approach because when things are broken, yes, it’s generally a bad thing, so you may as well find the silver lining in it from my point of view and figure out, okay, what have we learned about our systems as a result of the way that these things break? And sometimes the things that we learn are, in fact, not that deep, or there’s not a whole lot of learnings about it, such as when the entire county loses power, computers don’t work so well. Oh, okay. Great, we have learned that. More often, though, there seem to be deeper learnings.

And I guess what I’m trying to understand is, I have a relatively naive approach on what the idea of incident response should look like, but it’s basically based on the last time I touched things that were production-looking, which was six or seven years ago. What is the current state of the art that the advanced leaders in the space as they start to really look at how to dive into this? Because I’m reasonably certain it’s not still the, “Oh, you know, you can learn things when your computers break.” What is pushing the envelope these days?

Laura: Yeah, so it’s kind of interesting. You brought up incident response because incident response and incident analysis are the, sort of like, what do we learn from those things are very tightly coupled. What we can see when we look at someone responding in real-time to a failure is, it’s difficult to detect all of the signals; they don’t pop up and wave a little flag and say, like, “I am what’s broken.” There’s multiple compounding and interacting factors. So, there’s difficulty in the detection phase; diagnosis is always challenging because of how the systems are interrelated, and then the repair is never straightforward.

But when we stop and look at these kinds of things after the fact, of really common theme emerges, and that it’s not necessarily about a specific technical skill set or understanding about the system, it’s about the shared, distributed understanding of that. And so to put that in plain speak, it’s what do you know that’s important to the problem? What do I know that’s important to the problem? And then how do we collectively work together to extract that specific knowledge and expertise, and put that into practice when we’re under time pressure, when there’s a lot of uncertainty, when we’ve got the VP DMing us and being like, “When’s the system going to be back up?” and Twitter’s exploding with unhappy customers?

So, when we think about the cutting edge of what’s really interesting and relevant, I think organizations are starting to understand that it’s how do we coordinate and we collaborate effectively? And so using incident analysis as a way to recognize not only the technical aspects of what went wrong but the social aspects of that as well. And the teamwork aspects of that is really driving some innovation in this space.

Corey: It seems to me, on some level, that the increasing sophistication of what environments look like is also potentially driving some of these things. I mean, again, when you have three web servers and one of them’s broken, okay, it’s a problem; we should definitely jump on that and fix it. But now you have thousands of containers running hundreds of microservices for some Godforsaken reason because what we decided this thing that solves the problem of 500 engineers working on the same repository is a political problem, so now we’re going to use microservices for everything because, you know, people. Great. But then it becomes this really difficult to identify problem of what is actually broken?

And past a certain point of scale, it’s no longer a question of, “Is it broken?” so much as, “How broken is it at any given point in time?” And getting real-time observability into what’s going on does pose more than a little bit of a challenge.

Laura: Yeah, absolutely. So, the more complexity that you have in the system, the more diversity of knowledge and skill sets that you have. One person is never going to know everything about the system, obviously, and so you need kind of variability in what people know, how current that knowledge is, you need some people who have legacy knowledge, you have some people who have bleeding edge, my fingers were on the keyboard just moments ago, I did the last deploy, that kind of variability in whose knowledge and skill sets you have to be able to bring to bear to the problem in front of you. One of the really interesting aspects, when you step back and you start to look really carefully about how people work in these kinds of incidents, is you have folks that are jumping, get things done, probe a lot of things, they look at a lot of different areas trying to gather information about what’s happening, and then you have people who sit back and they kind of take a bit of a broader view, and they’re trying to understand where are people trying to find information? Where might our systems not be showing us what’s going on?

And so it takes this combination of people working in the problem directly and people working on the problem more broadly to be able to get a better sense of how it’s broken, how widespread is that problem, what are the implications, what might repair actually look like in this specific context?

Corey: Do you suspect that this might be what gives rise, sometimes, to it seems middle management’s perennial quest to build the single pane of glass dashboard of, “Wow, it looks like you’re poking around through 15 disparate systems trying to figure out what’s going on. Why don’t we put that all on one page?” It’s a, “Great, let’s go tilt at that windmill some more.” It feels like it’s very aligned with what you’re saying. And I just, I don’t know where the pattern comes from; I just know I see it all the time, and it drives me up a wall.

Laura: Yeah, I would call that pattern pretty common across many different domains that work in very complex, adaptive environments. And that is—like, it’s an oversimplification. We want the world to be less messy, less unstructured, less ad hoc than it often is when you’re working at the cutting edge of whatever kind of technology or whatever kind of operating environment you’re in. There are things that we can know about the problems that we are going to face, and we can defend against those kinds of failure modes effectively, but to your point, these are very largely unstructured problem spaces when you start to have multiple interacting failures happening concurrently. And so Ashby, who back in 1956 started talking about, sort of, control systems really hammered this point home when he was talking about, if you have a world where there’s a lot of variability—in this case, how things are going to break—you need a lot of variability in how you’re going to cope with those potential types of failures.

And so part of it is, yes, trying to find the right dashboard or the right set of metrics that are going to tell us about the system performance, but part of it is also giving the responders the ability to, in real-time, figure out what kinds of things they’re going to need to address the problem. So, there’s this tension between wanting to structure unstructured problems—put those all in a single pane of glass—and what most folks who work at the frontlines of these kinds of worlds know is, it’s actually my ability to be flexible and to be able to adapt and to be able to search very quickly to gather the information and the people that I need, that are what’s really going to help me to address those hard problems.

Corey: Something I’ve noticed for my entire career, and I don’t know if it’s just unfounded arrogance, and I’m very much on the wrong side of the Dunning-Kruger curve here, but it always struck me that the corporate response to any form of outage has is generally trending toward oh, we need a process around this, where it seems like the entire idea is that every time a thing happens, there should be a documented process and a runbook on how to perform every given task, with the ultimate milestone on the hill that everyone’s striving for is, ah, with enough process and enough runbooks, we can then eventually get rid of all the people who know all this stuff works, and basically staff at up with people who’d know how to follow a script and run push the button when told to buy the instruction manual. And that’s always rankled, as someone who got into this space because I enjoy creative thinking, I enjoy looking at the relationships between things. Cost and architecture are the same thing; that’s how I got into this. It’s not due to an undying love of spreadsheets on my part. That’s my business partner’s problem.

But it’s this idea of being able to play with the puzzle, and the more you document things with process, the more you become reliant on those things. On some level, it feels like it ossifies things to the point where change is no longer easily attainable. Is that actually what happens, or am I just wildly overstating the case? Either as possible. Or a third option, too. You’re the expert; I’m just here asking ridiculous questions.

Laura: Yeah, well, I think it’s a balance between needing some structure, needing some guidelines around expected actions to take place. This is for a number of reasons. One, we talked about earlier about how we need multiple diverse perspectives. So, you’re going to have people from different teams, from different roles in the organization, from different levels of knowledge, participating in an incident response. And so because of that, you need some form of script, some kind of process that creates some predictability, creates some common ground around how is this thing going to go, what kinds of tools do we have at our disposal to be able to either find out what’s going on, fix what’s going on, get the right kinds of authority to be able to take certain kinds of actions.

So, you need some degree of process around that, but I agree with you that too much process and the idea that we can actually apply operational procedures to these kinds of environments is completely counterproductive. And what it ends up doing is it ends up, kind of, saying, “Well, you didn’t follow those rules and that’s why the incident went the way it did,” as opposed to saying, “Oh, these rules actually didn’t apply in ways that really matter, given the problem that was faced, and there was no latitude to be able to adapt in real-time or to be able to improvise, to be creative in how you’re thinking about the problem.” And so you’ve really kind of put the responders into a bit of a box, and not given them productive avenues to, kind of, move forward from. So, having worked in a lot of very highly regulated environments, I recognize there’s value in having prescription, but it’s also about enabling performance and enabling adaptive performance in real-time when you’re working at the speeds and the scales that we are in this kind of world.

Corey: This episode is sponsored by our friends at Oracle HeatWave is a new high-performance query accelerator for the Oracle MySQL Database Service, although I insist on calling it “my squirrel.” While MySQL has long been the worlds most popular open source database, shifting from transacting to analytics required way too much overhead and, ya know, work. With HeatWave you can run your OLAP and OLTP—don’t ask me to pronounce those acronyms again—workloads directly from your MySQL database and eliminate the time-consuming data movement and integration work, while also performing 1100X faster than Amazon Aurora and 2.5X faster than Amazon Redshift, at a third of the cost. My thanks again to Oracle Cloud for sponsoring this ridiculous nonsense.

Corey: Yeah, and let’s be fair, here; I am setting up something of a false dichotomy. I’m not suggesting that the answer is oh, you either are mired in process, or it is the complete Wild West. If you start a new role and, “Great. How do I get started? What’s the onboarding process?” Like, “Step one, write those docs for us.”

Or how many times have we seen the pattern where day-one onboarding is, “Well, here’s the GitHub repo, and there’s some docs there. And update it as you go because this stuff is constantly in motion.” That’s a terrible first-time experience for a lot of folks, so there has to be something that starts people off in the right direction, a sort of a quick guide to this is what’s going on in the environment, and here are some directions for exploration. But also, you aren’t going to be able to get that to a level of granularity where it’s going to be anything other than woefully out of date in most environments without resorting to draconian measures. I feel like—

Laura: Yeah.

Corey: —the answer is somewhere in the middle, and where that lives depends upon whether you’re running Twitter for Pets or a nuclear reactor control system.

Laura: Yeah. And it brings us to a really important point of organizational life, which is that we are always operating under constraints. We are always managing trade-offs in this space. It’s very acute when you’re in an incident and you’re like, “Do I bring the system back up but I still don’t know what’s wrong or do I leave it down a little bit longer and I can collect more information about the nature of the problem that I’m facing?”

But more chronic is the fact that organizations are always facing this need to build the next thing, not focus on what just happened. You talked about the next incident starting and jumping in before we can actually really digest what just happened with the last incident; these kinds of pressures and constraints are a very normal part of organizational life, and we are balancing those trade-offs between time spent on one thing versus another as being innovating, learning, creating change within our environment. The reason why it’s important to surface that is that it helps change the conversation when we’re doing any kind of post-incident learning session.

It’s like, oh, it allows us to surface things that we typically can’t say in a meeting. “Well, I wasn’t able to do that because I know that team has a code freeze going on right now.” Or, “We don’t have the right type of, like, service agreement to get our vendor on the phone, so we had to sit and wait for the ticket to get dealt with.” Those kinds of things are very real limiters to how people can act during incidents, and yet, don’t typically get brought up because they’re just kind of chronic, everyday things that people deal with.

Corey: As you look across the industry, what do you think that organizations are getting, I guess, it’s the most wrong when it comes to these things today? Because most people are no longer in the era of, “All right. Who’s the last person to touch it? Well, they’re fired.” But I also don’t think that they’re necessarily living the envisioned reality that you described in the Howie Guide, as well as the areas of research you’re exploring. What’s the most common failure mode?

Laura: Hmm. I got to tweak that a little bit to make it less about the failure mode and more about the challenges that I see organizations facing because there are many failure modes, but some common issues that we see companies facing is they’re like, “Okay, we buy into this idea that we should start looking at the system, that we should start looking beyond the technical thing that broke and more broadly at how did different aspects of our system interact.” And I mean, both people as a part of the system, I mean processes part of the system, as well as the software itself. And so that’s a big part of why we wrote the Howie Guide, is because companies are struggling with that gap between, “Okay, we’re not entirely sure what this means to our organization, but we’re willing to take steps to get there.” But there’s a big gap between recognizing that and jumping into the academic literature that’s been around for many, many years from other kinds of high-risk, high-consequence type domains.

So, I think some of the challenges they face is actually operationalizing some of these ideas, particularly when they already have processes and practices in place. There’s ideas that are very common throughout an organization that take a long time to shift people’s thinking around, the implicit biases or orientations towards a problem that we as individuals have, all of those kinds of things take time. You mentioned the Overton window, and that’s a great example of it is intolerable in some organizations to have a discussion about what do people know and not know about different aspects of the system because there’s an assumption that if you’re the engineer responsible for that, you should know everything. So, those challenges, I think, are quite limiting to helping organizations move forward. Unfortunately, we see not a lot of time being put into really understanding how an incident was handled, and so typically, reviews get done on the side of the desk, they get done with a minimal amount of effort, and then the learnings that come out of them are quite shallow.

Corey: Is there a maturity model, where it makes sense to begin investing in this, whereas if you’ve do it too quickly, you’re not really going to be able to ship your MVP and see what happens; if you go too late, you have a globe-spanning service that winds up being down all the time so no one trusts it. What is the sweet spot for really started to care about incident response? In other words, how do people know that it’s time to start taking this stuff more seriously?

Laura: Ah. Well… you have kids?

Corey: Oh, yes. One and four. Oh yeah.

Laura: Right—

Corey: Demons. Little demons whom I love very much.

Laura: [laugh]. They look angelic, Corey. I don’t know what you’re talking about. Would you not teach them how to learn or not teach them about the world until they started school?

Corey: No, but it would also be considered child abuse at this age to teach them about the AWS bill. So, there is a spectrum as far as what is appropriate learnings at what stage.

Laura: Yeah, absolutely. So, that’s a really good point is that depending on where you are at in your operation, you might not have the resources to be able to launch full-scale investigations. You may not have the complexity within your system, within your teams, and you don’t have the legacy to, sort of, draw through, to pull through, that requires large-scale investigations with multiple investigators. That’s really why we were trying to make the Howie Guide very applicable to a broad range of organizations is, here are the tools, here are the techniques that we know can help you understand more about the environment that you’re operating in, the people that you’re working with, so that you can level up over time, you can draw more and more techniques and resources to be able to go deeper on those kinds of things over time. It might be appropriate at an early stage to say, hey, let’s do these really informally, let’s pull the team together, talk about how things got set up, why choices were made to use the kinds of components that we use, and talk a little bit more about why someone made a decision they did.

That might be low-risk when you’re small because y’all know each other, largely you know the decisions, those conversations can be more frank. As you get larger, as more people you don’t know are on those types of calls, you might need to handle them differently so that people have psychological safety, to be able to share what they knew and what they didn’t know at the time. It can be a graduated process over time, but we’ve also seen very small, early-stage companies really treat this seriously right from the get-go. At Jeli, I mean, one of our core fundamentals is learning, right, and so we do, we spend time on sharing with each other, “Oh, my mental model about this was X. Is that the same as what you have?” “No.” And then we can kind of parse what’s going on between those kinds of things. So, I think it really is an orientation towards learning that is appropriate any size or scale.

Corey: I really want to thank you for taking the time to speak with me today. If people want to learn more about what you’re up to, how you view these things and possibly improve their own position on these areas, where can they find you?

Laura: So, we have a lot of content on jeli.io. I am also on Twitter at—

Corey: Oh, that’s always a mistake.

Laura: [laugh]. @lauramdmaguire. And I love to talk about this stuff. I love to hear how people are interpreting, kind of, some of the ideas that are in the resilience engineering space. Should I say, “Tweet at me,” or is that dangerous, Corey?

Corey: It depends. I find that the listeners to this show are all far more attractive than the average, and good people, through and through. At least that’s what I tell the sponsors. So yeah, it should be just fine. And we will of course include links to those in the [show notes 00:27:11].

Laura: Sounds good.

Corey: Thank you so much for your time. I really appreciate it.

Laura: Thank you. It’s been a pleasure.

Corey: Laura Maguire, researcher at Jeli. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this
podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please give a five-star review on your podcast platform of choice along with an angry, insulting comment that I will read just as soon as I get them all to display on my single-pane-of-glass dashboard.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Serena

Serena is a Network Engineer who specializes in Data Center Compute and Virtualization. She has degrees in Computer Information Systems with a concentration on networking and information security and is currently pursuing a master’s in Data Center Systems Engineering. She is most known for her content on TikTok and Twitter as Shenetworks. Serena’s content focuses on networking and security for beginners which has included popular videos on bug bounties, switch spoofing, VLAN hoping, and passing the Security+ certification in 24 hours.

Links:

  • Cisco cert Discord study group:https://discord.com/invite/uXQ8yWnN8a
  • Beacons:https://beacons.page/shenetworks
  • TikTok:https://www.tiktok.com/@shenetworks
  • sysengineer’s TikTok:https://www.tiktok.com/@sysengineer
  • Twitter:https://twitter.com/notshenetworks

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Sysdig. Sysdig is the solution for securing DevOps. They have a blog post that went up recently about how an insecure AWS Lambda function could be used as a pivot point to get access into your environment. They’ve also gone deep in-depth with a bunch of other approaches to how DevOps and security are inextricably linked. To learn more, visit sysdig.com and tell them I sent you. That’s S-Y-S-D-I-G dot com. My thanks to them for their continued support of this ridiculous nonsense.

Corey: Today’s episode is brought to you in part by our friends at MinIO the high-performance Kubernetes native object store that’s built for the multi-cloud, creating a consistent data storage layer for your public cloud instances, your private cloud instances, and even your edge instances, depending upon what the heck you’re defining those as, which depends probably on where you work. It’s getting that unified is one of the greatest challenges facing developers and architects today. It requires S3 compatibility, enterprise-grade security and resiliency, the speed to run any workload, and the footprint to run anywhere, and that’s exactly what MinIO offers. With superb read speeds in excess of 360 gigs and 100 megabyte binary that doesn’t eat all the data you’ve gotten on the system, it’s exactly what you’ve been looking for. Check it out today at min.io/download, and see for yourself. That’s min.io/download, and be sure to tell them that I sent you.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Today’s guest was on relatively recently, but it turns out that when I have people on the show to talk about things, invariably I tend to continue talking to them about things and that leads down really interesting rabbit holes. Today is a stranger rabbit hole than most. Joining me once again is @SheNetworks or Serena [DiPenti 00:00:51]. Thanks for coming back and subjecting yourself to, basically, my nonsense all over again in the same month.

Serena: Thanks for having me back. Excited.

Corey: So, you have a, I think study group is the term that you’re using. I don’t know how to describe it in a way that doesn’t make me sound ridiculous and describing and speaking with my hands and the rest. It’s a Discord, as the kids of today tend to use. There are some private channels on an existing Discord group, and we’ll get to the mechanics of that in a second. But it’s a study group for various Cisco certifications, which it’s been a while since I had one; my CCNA is something I took back in 2009. I’ve checked, it’s expired to the point where they can’t even look it up anymore to figure out who I might have been, once upon a time. What is this group and where did it come from?

Serena: Yeah, so the Discord itself is kind of a collective of a bunch of people that are creators on TikTok. And it’s just, like, a cool place to connect, especially people from TikTok join, people from Twitter join, they want to interact, you know, a great place to get resources if you’re early in your career. I—you know, new year, new me resolution was [laugh] I wanted to start studying for the CCNP a little bit, and I’ve been doing it pretty loosely for a while, but I kind of was like, all right, time to actually sit down and dedicate some real time to this. And I put on Twitter, you know, if anybody else was interested—I know there’s other various study groups out there and things like that, but I was just like, hey, you know, it was anyone interested and a study group and I got really good response. Of course, a lot of people are at the CCNA level, so I made a channel for CCNA and CCNP, so whatever level you’re at, you can come in and ask questions. It’s really great.

Corey: One thing that irked me when I first joined, as well, there’s no CCENT which was sort of the entry-level Cisco cert, the first half of the CCNA, and I did a bit of googling before shooting my mouth off. And it turns out that Cisco sunset that cert a while back, so CCNA is now the entry-level cert, as I understand it.

Serena: Yeah. So, when I did my CCNA, I did the C-C-E-N-T—the CCENT, and then the ICND2, and that’s how I got my CCNA. And then I went and got the Data Center CCNA, which was two exams… two? Or maybe it was just one. I can’t remember fully. But they basically got rid of all of their CCNAs and created one new one that’s just the CCNA Enterprise.

Corey: What I found worked out for me when I was going through the process of getting the CCNA—the CCENT, I forget how at the time, came along for the ride. And it was the CCENT, the baseline stuff that really added value to my entire career. That piece of advice that I would give anyone in the technical space is when your hand-waving over a thing you don’t really understand. Maybe stop doing that one afternoon when you don’t have anything else going on, dig into it.

For me, it was always, “What the hell is a subnet mask?” “I don’t know. It’s the thing that I put the right numbers in, the box stops turning gray and will turn black and let me click the button; life goes on.” Figuring out what that meant and how it was calculated was interesting and it made me understand what’s going on at a deeper level. Which means that invariably when things break as—they’re computers; they break—I could have a better understanding of the holistic system and ideally have a better chance of getting to an outcome of fixing it.

So, I’m not sitting here suggesting that anyone who wants to, “Oh, you want to work in the cloud and go and build things out on top of AWS or GCP. Great, go and get a Cisco certification is the first stop along your journey.” But understanding how the network works is absolutely going to serve you well for the rest of your technical career because not a lot has changed in the networking sense over the past 13 years since I sat the certification exam. It turns out that the TCP handshake still works the same way: Badly.

Serena: [laugh]. Yeah, and to your point, the troubleshooting part is really where you need that depth of knowledge, right? And that’s typically when it’s crunch time and things are gone awry. And you really need to have an understanding of okay, is it the subnet mask? And the quicker that you can identify that outage, that problem, the quicker you get a resolution. And you do need depth of knowledge for that, and understanding that kind of underlying infrastructure is so helpful.

Corey: And that was always the useful part of the certification—and the exam that went along with it—to me was, “Okay, with a subnet mask of whatever you’re talking about here, great. How many usable IP addresses are there in the network?” And yeah, that’s the kind of thing that we really care about.

The stuff that drove me nuts was the other half of it, where it’s the, “Ah, what is the proper syntactical command on the Cisco command line to display this thing?” And it’s, “First, I can probably look that up or tab-complete it or whatnot. Secondly, I get it’s a Cisco exam, but this is a world where interoperability is very much a thing and it is incredibly likely that the thing I need to find that out on is not going to ultimately be a Cisco device, once I’m working in enterprise.”

Serena: Yeah, I do have similar feedback when it comes to that because right now, I’ve been trying to do kind of a chapter a day out of the Cisco Press book, and that’s my main source of studying right now. I like to read a lot, so reading is usually my main method of studying, I guess. But I’m in a chapter right now that’s, like, 100 pages of just hardware specifics. And we’re talking about, like, PCIe cards and VICs and the different models and which ports are unified and you can configure for Fibre Channel, and which are uplink on the different generations. And I’m like, “Ohh.”

I hate that. It’s my least favorite part of studying because for that, I mean, I always just pull up the documentation. And it’s like, “Okay, here’s the ports that can be, you know, configured as Fibre Channel over Ethernet or Fibre Channel,” or whatever. Remembering it off the top of my head, which model, which year, which ports, I’m not great with that. And I don’t think it’s, honestly, that valuable when it comes to certification exams because you really should be using the documentation when you are doing those types of configurations between hardware and generations and compatibility.

Corey: We sort of see the same thing in the development space, where, okay, the job we’re hiring you to do is to work on some front end work and change how things are rendered, but when we’re doing the job interview for that role, oh, now we have an empty whiteboard, we want you to write syntactically valid code that will implement some sorting algorithm or whatnot, while some condescending jerk sits there. And, “Nope, that’s not it,” in the background in a high-pressure environment because for that jackwagon, it’s any given Thursday, but for you, it determines the next phase of your career. And I hated that stuff. Whereas in the real world, I’m not going to be implementing an algorithm like that in any realistic sense; I’ll be using the one built into whatever language I’m using. It’s important from a computer science perspective to know it, but from a day-to-day job environment, not so much.

And I can’t recall the last time that I had to fix a technical issue where I did not have the internet as a resource while I was fixing that issue, even when it’s the internet is down because it turns out without the network, I just have a whole bunch of expensive space heaters here, great, my phone still worked. I could check, “Oh, what is the command to get back into that firewall?” That it turns out, I just locked myself out of by—yeah, it turns out when you close a port and you’re using that port, mistakes show.

Serena: Yeah, I agree with that. And I mean, that goes into the much broader conversation of technical interviews because even as a network engineer, one time I had a whiteboard technical interview where they were asking, like, routing questions, but I didn’t have access to any equipment, and so it was just basically asking them questions. And I’m a very visual person, so for me to not be able to, like, kind of put my hands on something and, like, run some commands and look over it myself. I did so horribly in that interview, and I left feeling just, like—I left feeling really bad about myself, honestly, because I had done so bad. And for me, I was assuming they were using some routing protocol. And they’re like, “No, it’s actually all statically configured.” And I was like, I would be able to know that if I could run commands and, like, actually look. But it was so bad.

Corey: Right. And it’s stressful working in front of people. I know that whatever I’m typing in front of an audience, I don’t do it, but it feels like what I did first is, all right, let me put my mittens on, and then I—because I can’t type to save my life, and I look incompetent across five different levels at that point. And yeah, it’s these contrived problems. One of the things I like about the study group is when there’s a question that is, I guess, not the answer, I would expect, it’s okay, we can talk about that. Give me more context behind why.

I thought it was this. Clearly, I’m missing something—or the bot is broken—so what is going on here? Help me understand why this is the way that it is? And back when I was learning how this stuff all worked, I went through originally a class at a community college and then finished it up with apparently with sort of a brain dump style boot camp, which I didn’t really realize was a thing until after the fact. It was just memorization of these things.

Which okay, great. I could memorize my way through some things I would never use again like EIGRP, one of Cisco’s proprietary routing protocols that I’ve never heard of anyone using in the real world before, but I’m sure it’s a thing and they’re trying to push it. Great. I can skate past that well enough to hang a cert, but it didn’t feel like the way to learn it because there was no context. It was just the rote memorization.

Serena: Mm-hm. Yeah, and that is very difficult. I’m a big fan of theory, so you know, when we’re talking about VIC cards, I was going through each generation, and which you would use for a blade or a rack server, whatever. I think that your time is better spent understanding what a VIC card is, why it’s important, maybe, like, the history, and all that instead of being, like, “This version isn’t compatible with this UCS blade server,” or whatever. Because I am studying for the Data Center flavor of the CCNP right now, so it’s a little bit of a different path. I think most people take the enterprise, that’s the more traditional route, switch, IOS. Mine’s more UCS, Nexus, HyperFlex type questions.

Corey: One thing that I always appreciate is, for example, take subnet mask [crosstalk 00:10:57] calculations. Yeah, I can figure that out on a whiteboard now. But here in the real world, everyone uses a subnet calculator. It’s the way that things work. And there’s a lot of discussion back and forth about things like that, without talking about the real-world implications, such as, if you’re building out two subnets inside of a larger range, don’t put them right next to each other because if you need to expand the network later, you’re in a world of pain compared to if you had given them some significant breathing room.

And okay, great. You probably don’t need to use all the [10.0.0.0/8 00:11:30] network in your small-scale environment, and even some larger-scale ones you’re hard-pressed to use all those things.

It’s just the real-world experience, and you understand that you don’t want to do that. The second time. The first time you do it because why not? It’s easy to remember for humans. And then you run into weird issues with oh, well, why would I ever have more than 254 servers sitting in a subnet—or 253, whatever the number is these days, don’t yell at me—great.

What about containers running on top of those things? Oh, right, the worst answer to so many architectural patterns, we’ll throw some containers at it. And you’re back into those problems.

Serena: Yeah.

Corey: It’s the real-world scars you get.

Serena: Yeah. And I think that there is such a difference between when you’re studying and learning versus—and taking certifications or tests—than in the real world. And that was very discouraging for me when I was first learning because I would take these exams—and we had a Cisco academy where I went to college—and I would take these exams, and my professor was just known for her very difficult test, so I think her advanced routing course, maybe only 30% of the people who took it passed it their first try. And so I would take these exams, I’d walk away being like, “I don’t know anything. I’m never going to be a good network engineer, I’m never going to be able to get a job or anything,” because I couldn’t regurgitate which show command was showing me errors on a switch, right?

And then now in the real world, I’m like, okay, relieved because I was like, I can look this up, like, I can take my time. And then you know, with getting your hands on—I mean, you learned so much within your first year; that is probably more than I learned in all four years of school. But saying that, it was really great for me to have that base of all of that underlying networking and already kind of understanding the terminology alone is such a big… barrier, I would say, like, just being able to sit in a room and listen to these conversations and understand what’s going on. That’s half the battle in the beginning. [laugh].

Corey: I have never heard anyone be prouder of being bad at their job than a professor saying, “I have a 30% pass rate.” Isn’t your whole ethos of that role to be someone who teaches people how to do a thing? So, if two-thirds of your class is not learning that thing, it doesn’t mean you’re a hard grader, it means you’re bad at conveying the concept and/or testing for understanding of the thing that you’ve just taught them. If you’re a teacher listening to this, please don’t email me until you fix your problem first.

Serena: [laugh]. See, and… she would come in and say on the first day class—I took multiple classes with her and she was like, “If you read everything in the book, and pay attention to all the slides, you’re still going to fail.” She wanted you to really go above and beyond, and commit and run all these labs and do all these things, and in college, I hated it. I was so resentful and angry because it really did make me feel bad. But at the same time, there was one point someone had asked her a question, and she was like, “Why don’t you ask Serena? She has the highest grade in the class.”

And I was shocked because I had, like, a C in the [laugh] class. And I was like, “Me? I’m the one that has the highest grade in the class?” And I would definitely do things a little bit differently if I were teaching that course because it, I think, turned off a lot of people into the field. But me passing those grades, I mean, I really could have probably taken the CCNP right when I was done with those courses and passed with flying colors. But I didn’t have the money to take the CCNP exams until much later when I had a job. And now it’s like so much has changed. The exams have changed. I’m in Data Center now. So, a little bit different. But yeah. [laugh].

Corey: I never understood the idea of charging for certs. If people are spending the time and energy to learn about your company’s specific technology well enough to take the exam, they’re probably going to want to use it in their career as they move forward, so charging a few 100 bucks to sit the test has never struck me as a good idea. And the cloud companies do the exact same things as well. And every company that attains some level of success launches a certification exam, but then they charge a few 100 bucks for it, which… does that money really matter because either you’re an engineer, and your company is going to be paying for it, or you’re making engineering money these days, and it’s just an irritant, but it feels to me like the people that really get disadvantaged by that are the early learners, the students, the folks who are planning to have a career in this, but a few 100 bucks becomes a barrier.

Serena: Oh, it’s a huge barrier. I mean, it was a big barrier for me. I didn’t have money to go to college, so I took out student loans. I worked my way through college and constantly had a job, which then was difficult because my grades suffered because I didn’t have the same amount of time.

Corey: You did have the highest grade in class, I recall.

Serena: [laugh]. For that one course. For the one course. [laugh]. But I didn’t have the same amount of time in a day to study as some of my classmates who didn’t have to have a job in college.

But then also, I couldn’t afford $300 to take one exam out of the three that you needed at the time for the CCNP. And that’s when I was early in my career. The CCNA, too, like, I didn’t have the money to take that exam either. And I think a lot of people are in that position because they are trying to better their knowledge. They’re trying to achieve a new job.

That’s what those certifications are geared towards, right? And so putting that $300—I mean, that person might be working a minimum wage job, and they’re trying to get out of that minimum wage job into a higher—paying tech job. And $300 is a lot of money. It is a lot of money. My rent in college was $300. That’s a whole month’s rent for me, right, to put it in perspective. So yeah.

Corey: Yeah. We’ll be throwing a bunch of credit codes your way for folks who are learning and [unintelligible 00:17:10] the financial burden because it’s important that people be able to not have money being the obstacle to learning a technical field. I am curious, though, as to the genesis of this whole Discord because I heard you talking about it, I joined, but there are a lot of other people talking about different things. Most notably and importantly, there’s an Ohio slander channel—

Serena: [laugh].

Corey: —in there, which is just spot-on perfect from where I sit. But it’s not just you, and it’s not just networking stuff. It’s a systems engineering Slack. Where did it come from?

Serena: Yeah so sysengineer, my friend [Chris Lynd 00:17:43]—she’s also a TikTok creator—and she set up her own Discord server, which I have kind of like inserted myself into. It’s very hard to run your own server, right, so it’s kind of more of a collective at this point. But she’s sysengineer on TikTok, and so her server is just sysengineer. And there’s a lot of memes, right? Because we have a lot of, like, Gen Z—I mean, who doesn’t love a good meme? And Chris Lynd, sysengineer, is from Ohio, I’m from Ohio. So, the Ohio slander thing is kind of funny because we’re just like always talking crap about Ohio. [laugh].

Corey: Which it deserves, let’s be very clear here. I have family in Ohio, myself. Every time I visited them, my favorite part was leaving Ohio. I mean, data transfer between AWS regions, the least expensive one is the one cent instead of two cents between Ohio and Virginia
because even data wants to get out of Ohio.

Serena: It was like, 11 of the astronauts are from Ohio. And it was like, “What about Ohio makes me want to leave the Earth?” [laugh].

Corey: Yeah, “How far can I get from Ohio, the absolute furthest place away?” “Well, here’s the furthest place on earth.” “Not far enough.” I know, if you’re from Ohio, I know you’re going to be very upset. You’re going to be listening to this and angrily riding your horse to Pennsylvania to send an angry email my way, but that’s okay. You’ll get there eventually.

Serena: But yeah, there’s a lot of memes and stuff from TikTok. It’s funny because we love to joke; we love to keep it light-hearted; we want to attract people who are younger, a lot of the memes come from TikTok. And so it’s a fun, good time. And there’s developers on there, there’s tons of people that work other jobs that aren’t systems engineering, or network engineering. So, we have a bunch of different opportunities and channels for other people to kind of ask questions and connect with other people in the field. Especially with everyone being remote for the most part now, and Covid, you don’t have a ton of social interaction, so it’s a good place to go get some social interaction.

Corey: This episode is sponsored by our friends at Oracle HeatWave is a new high-performance query accelerator for the Oracle MySQL Database Service, although I insist on calling it “my squirrel.” While MySQL has long been the worlds most popular open source database, shifting from transacting to analytics required way too much overhead and, ya know, work. With HeatWave you can run your OLAP and OLTP—don’t ask me to pronounce those acronyms again—workloads directly from your MySQL database and eliminate the time-consuming data movement and integration work, while also performing 1100X faster than Amazon Aurora and 2.5X faster than Amazon Redshift, at a third of the cost. My thanks again to Oracle Cloud for sponsoring this ridiculous nonsense.

Corey: It’s also great because when I was early in my career, I was a traveling consultant, and periodically I would find myself, well, working 40 hours a week and then in a hotel room for the rest of it. That’s sort of depressing; I would go to local meetups. I’ll never forget going to one Linux user group meeting. In this town, apparently, Linux wasn’t really a thing, so the big conversational topic is how to sneak Linux into your Windows job. And I’m sitting around here going, “I don’t know if that’s necessarily the best way to go about it.”

But I checked; there were no reasonable Linux jobs in that community. So, all of their focus in these user groups was about doing it as a side project, as this aspirational thing. And I’m sitting here visiting from out of town, I’m thinking, “Well, I have a job in the Linux environment. And how did I find it? I just went online and looked for jobs that had the word Linux in the title, and there you go.”

That option is not open to everyone in every geography, so being able to get exposed to folks who aren’t all in your neighborhood is one of the big benefits I found online forums like this.

Serena: Yeah. One of the things that I think was positive that came out of Covid is, if you are in a smaller region—one of the reasons I left Ohio was because of a lack of jobs, right. And because there was more opportunity in other areas. And now I wouldn’t have had to move. Not that say—I mean, I would have probably moved out of Ohio anyway.

But if you don’t want to, if your whole family’s there now, you’re luckily not really stuck with just the jobs that are in your local area. There’s tons of remote jobs now. I think that’s fantastic, and like I said, one of the positive things that did come out of Covid.

Corey: The thing that I don’t fully understand is folks who are working for remote companies—we’re a distributed companies outside The Duckbill Group, and we pay the same for a role, regardless of where on—or off—the planet you happen to be sitting, just because the value you’re adding makes zero difference to me based upon where you happen to be. And there are a number of companies out there who are being very particular about well, where are you geographically because then we need to adjust your comp so you’re appropriate for that market. And it’s, really? Is the work you’re doing this month materially different than the work you’re doing next month, as far as value goes, based upon where you’re sitting? I don’t buy it. But it’s also challenging at giant companies to wind up paying the same across the board for all of your staff in one fell swoop.

Serena: I think it’s particularly bad. I had seen some companies that were basically saying if they’re already employed and already getting some salary, and then, like, if you move, we’re going to lower your salary. And I was like, it just to me seems so greedy, especially coming from these massive companies that charge huge profits, that you’re going to be concerned over a ten, twenty, thirty-thousand dollar difference, right? And it’s like, it just seems greedy to me because it’s like, well, you had no problem paying that while I was living there, but now it’s a problem that I move closer to family or something like that? I luckily was not in that position, but it would have put a distaste in my mouth towards that company, I think, as an employee in that position.

Corey: We want to know where people are for tax purposes, we have this whole thing about not committing tax fraud, but aside from that, we don’t care where you happen to be. We’ve had people take a month in Costa Rica, for example. Great. Have fun. Let us know what you think. As long as you have internet there and you make the scheduled meetings you’ve committed to make, great.

But that’s part of the benefit of having a company has been distributed since before the pandemic. What I really have sympathy for is folks who had built companies that depended on an in-office culture, and suddenly you’re forced into remote during a very stressful time.

Serena: Mm-hm. Yeah. Luckily, I mean, most of my jobs are very easily remote, but I can see that. I don’t know. The whole—I don’t ever want to work in an office again, personally. It’s just not for me. I have done really well transitioning to work from home and still keeping up with all my coworkers, and reaching out to them, having meetings.

I think, at this point, after two years in, companies are going to have a really hard time justifying to their employees, like, oh, we have to be back in office. And it’s like, well, why? Is productivity down? Are we not as profitable? Like, what happened within these last two years that is making you think, like, we need to go back into the office? And they don’t really have anything besides, “Culture?” And it’s like, yeah, you’re going to need to do more than that. [laugh].

Corey: It’s important for us to see our co-workers from time to time, and once it’s safe to do so we’re going to be doing quarterly meetups in various places, but that’s also… it’s not every day.

Serena: Right.

Corey: The technology problems, I have less sympathy for it now than I did at the start of the pandemic, where network engineers were basically calling the data center and, “Yeah, can you go reboot the VPN concentrator?” “Uh, okay. Which server is that? Probably the one that’s glowing white-hot right now.” Because they aren’t designed for the entire company to be using it simultaneously all the time. Two years later, we have mostly fixed those problems.

Serena: Yeah, yeah. Two years later, it’s like, okay, you’re going to really have to convince me to go back into the office. [laugh]. And I like the flexibility. Like, I really do. If I want to move, I can move. If I want to, like you said, go to Costa Rica for a month, I could do that. But there’s a lot of options, flexibility. I’ve been having a great time work from home.

Corey: And I’ve been having a lot of fun exploring the bounds of this new Discord group, and I’ll throw a link to it in the [show notes 00:24:49] because anyone who wants to show up and can validate that their human being is welcome to join until they turn into a jerk which is basically the [audio break 00:24:57] the community these days, let’s be clear, but I found there are a couple of Discord bots—and yeah, it’s all the same thing now—that ask test questions, and you can give an answer and it tells you in a DM whether you got it right or not, which is always fun when the bot is broken, and you’re sitting there going well, that doesn’t make much sense. But what other stuff has been built into this? For those of us who spend all of our time in Slack these days, what is the advantage of the Discord way of doing things?

Serena: I guess for me, I’m not, like, a huge Discord person. This is really the only one that I participate in. I’m in a couple of my friends Discord as well, but there’s a lot of stickers that are customizable, that relate back to memes a lot of the times. But yeah, the bot that you had mentioned is a great feature that Discord has where @terranovatech, who’s also another TikTok content creator—his name’s Anthony—he created from Python a practice question bot for CCNA and CCNP. And so, uploaded some questions to those.

The bot is in beta guys, so you know, just like, [laugh] be aware of that. We are trying to constantly improve it and add new features. I have been adding a ton of questions for [D core 00:26:05] as I go through my book studying; I’ll, you know, create practice questions. And that’s typically a part of my normal studying routine, is creating practice questions that I can then go back to after I’ve read something to solidify it in my mind. And you know, you can use those questions, too, you can suggest questions. If you’re like, “Hey, I was doing studying and I think this would be a cool question to add to the Discord bot.” We can do that as well. And so that’s great. I love that feature.

Corey: One last question before we wind up calling it an episode. Recently, you have caused a bit of TikTok controversy, for lack of a better term. And sure enough, we’ve had people swing in from all over the planet that chime in and yell at you in the comments. What’s going on there?

Serena: Okay. Yeah, so that’s not unusual for me to cause some TikTok drama in the tech space. Okay, so there’s a TikTok trend right now where it’s a song and the song lyrics are, “You look so dumb right now.” Okay? And the other videos, like, if you click the sound, you can see, like, some of the videos will say, like, “They told me I needed to rotate my tires, but they rotate every time I drive.”

And someone was like, “My girlfriend said she needs new foundation, but our house is just fine.” And so in the background, you hear the song that says, like, “You look so dumb right now.” So, it’s just, like, a funny… funny joke. I did it, and I was like, I knew some people were going to miss the joke. And I said, you know, “When they say you need a backup, but you use RAID.” [laugh]. And so the sound is, “You look so dumb right now.”

And I was definitely expecting people to miss the joke. And so I even tweeted at the same time, I was like, “I posted a new video, like, about that joke.” And so I was like, “Be prepared for the comments.” Because I knew even someone would be, like, she’s just backtracking now. Like, she just is embarrassed. But I was like, “It’s the joke guys.”

I even put in the caption #thisisajoke. And, like, 90% of people that commented on it just completely missed that joke and were very upset that I made that—that I said that.

Corey: Anyone who believes RAID is a backup only has to make one mistake deleting the wrong thing or overwriting something important before they realize that is very much not the case. And if you’ve been in tech for longer than about 20 minutes, you probably made a mistake like that at one point. It’s not one of those things that could reasonably be expected that someone would take seriously. But yet, here we are with entire legions of people with no sense of humor.

Serena: Yeah, it ended up in, like, Facebook groups and stuff, too, where these people thought I was being serious. And in the comments, I started making more jokes because someone’s like, well, what if your data center catches on fire? And I was like, “Well, don’t have a fire at your data center. Like, I don’t understand. Obviously.” And so I just tried to, like, you know, make more jokes back to, kind of, keep it up
and people were very upset. [laugh].

Corey: That’s why you’re not allowed to smoke in them. Problem solved. Where would the fire come from? Yeah.

Serena: There was, like, someone was like, “Well, what if you get ransomware?” And I was like, “We have Norton.” Like, what—[laugh] like, just, like, making the most red—and I was trying to really go outlandish with some of them because they’re like, “RAID is not a replacement for cold storage.” And I was like, “Well, we have a lot of fans, so our RAID is very cold.” [laugh]. And, like, just kept it going. Some people were not happy.

Corey: I love that. They just keep doubling down on the dumb. The problem is some people are lifelong experts at it, and they’re always going to beat you with experience when you try it. It’s…

Serena: [laugh]. Yeah.

Corey: Honestly, the hardest thing to learn, one it was valuable, least from my perspective, is learning when to just ignore the comments and keep going.

Serena: Yeah. I definitely get some that I ignore. I mean, if they’re, like, overly mean, I’ll block somebody or something like that. You know,
for someone just missing a joke, it’s like, “Okay, whatever.” But yeah, some people—even after they’re like, “Hey, man. This is just a joke.” They’re like, “Well, this isn’t a funny joke.” And I was like, “I will never make a joke about RAID as a backup again. I promise.” [laugh].

Corey: No, you already told that joke. There are better ones you can explore.

Serena: Yeah. For sure.

Corey: So, if people want to come and hang out in this Discord, what’s the best way for them to find it? We’ll put it in the [show notes 00:30:05], but sometimes people listen rather than read.

Serena: Yeah, I think if you even just Google ‘sysengineer Discord’ it should come up like that; it’s on the Google returned searches. It’s a link in my Beacons on my TikTok. It’s in a link in sysengineer’s TikTok. So, there’s a couple different places that you can find and join.

Corey: And of course, in the [show notes 00:30:27] for this podcast, as well.

Serena: And the [show notes 00:30:30] of this podcast, of course. [laugh].

Corey: Thank you so much for taking the time to talk to me about all this. If people want to follow you beyond just the Discord, where’s the best place for them to find you?

Serena: So, I’m @SheNetworks on TikTok and then I’m @notshenetworks on Twitter. So, you can find me in both of those locations.

Corey: Fantastic. Thanks so much for taking the time to speak with me today. I appreciate it.

Serena: Thanks for having me on.

Corey: Serena DiPenti, network engineer and of course@SheNetworks on the internet. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry comment telling me which RAID level makes the best backup.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About AB

AB Periasamy is the co-founder and CEO of MinIO, an open source provider of high performance, object storage software. In addition to this role, AB is an active investor and advisor to a wide range of technology companies, from H2O.ai and Manetu where he serves on the board to advisor or investor roles with Humio, Isovalent, Starburst, Yugabyte, Tetrate, Postman, Storj, Procurify, and Helpshift. Successful exits include Gitter.im (Gitlab), Treasure Data (ARM) and Fastor (SMART).

AB co-founded Gluster in 2005 to commoditize scalable storage systems. As CTO, he was the primary architect and strategist for the development of the Gluster file system, a pioneer in software defined storage. After the company was acquired by Red Hat in 2011, AB joined Red Hat’s Office of the CTO. Prior to Gluster, AB was CTO of California Digital Corporation, where his work led to scaling of the commodity cluster computing to supercomputing class performance. His work there resulted in the development of Lawrence Livermore Laboratory’s “Thunder” code, which, at the time was the second fastest in the world.

AB holds a Computer Science Engineering degree from Annamalai University, Tamil Nadu, India.

AB is one of the leading proponents and thinkers on the subject of open source software - articulating the difference between the philosophy and business model. An active contributor to a number of open source projects, he is a board member of India's Free Software Foundation.

Links:

  • MinIO: https://min.io/
  • Twitter: https://twitter.com/abperiasamy
  • MinIO Slack channel: https://minio.slack.com/join/shared_invite/zt-11qsphhj7-HpmNOaIh14LHGrmndrhocA
  • LinkedIn: https://www.linkedin.com/in/abperiasamy/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Sysdig. Sysdig is the solution for securing DevOps. They have a blog post that went up recently about how an insecure AWS Lambda function could be used as a pivot point to get access into your environment. They’ve also gone deep in-depth with a bunch of other approaches to how DevOps and security are inextricably linked. To learn more, visit sysdig.com and tell them I sent you. That’s S-Y-S-D-I-G dot com. My thanks to them for their continued support of this ridiculous nonsense.

Corey: This episode is sponsored in part by our friends at Rising Cloud, which I hadn’t heard of before, but they’re doing something vaguely interesting here. They are using AI, which is usually where my eyes glaze over and I lose attention, but they’re using it to help developers be more efficient by reducing repetitive tasks. So, the idea being that you can run stateless things without having to worry about scaling, placement, et cetera, and the rest. They claim significant cost savings, and they’re able to wind up taking what you’re running as it is, in AWS, with no changes, and run it inside of their data centers that span multiple regions. I’m somewhat skeptical, but their customers seem to really like them, so that’s one of those areas where I really have a hard time being too snarky about it because when you solve a customer’s problem, and they get out there in public and say, “We’re solving a problem,” it’s very hard to snark about that. Multus Medical, Construx.ai, and Stax have seen significant results by using them, and it’s worth exploring. So, if you’re looking for a smarter, faster, cheaper alternative to EC2, Lambda, or batch, consider checking them out. Visit risingcloud.com/benefits. That’s risingcloud.com/benefits, and be sure to tell them that I said you because watching people wince when you mention my name is one of the guilty pleasures of listening to this podcast.in a silo

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’m joined this week by someone who’s doing something a bit off the beaten path when we talk about cloud. I’ve often said that S3 is sort of a modern wonder of the world. It was the first AWS service brought into general availability. Today’s promoted guest is the co-founder and CEO of MinIO, Anand Babu Periasamy, or AB as he often goes, depending upon who’s talking to him. Thank you so much for taking the time to speak with me today.

AB: It’s wonderful to be here, Corey. Thank you for having me.

Corey: So, I want to start with the obvious thing, where you take a look at what is the cloud and you can talk about AWS’s ridiculous high-level managed services, like Amazon Chime. Great, we all see how that plays out. And those are the higher-level offerings, ideally aimed at problems customers have, but then they also have the baseline building blocks services, and it’s hard to think of a more baseline building block than an object store. That’s something every cloud provider has, regardless of how many scare quotes there are around the word cloud; everyone offers the object store. And your solution is to look at this and say, “Ah, that’s a market ripe for disruption. We’re going to build through an open-source community software that emulates an object store.” I would be sitting here, more or less poking fun at the idea except for the fact that you’re a billion-dollar company now.

AB: Yeah.

Corey: How did you get here?

AB: So, when we started, right, we did not actually think about cloud that way, right? “Cloud, it’s a hot trend, and let’s go disrupt is like that. It will lead to a lot of opportunity.” Certainly, it’s true, it lead to the M&S, right, but that’s not how we looked at it, right? It’s a bad idea to build startups for M&A.

When we looked at the problem, when we got back into this—my previous background, some may not know that it’s actually a distributed file system background in the open-source space.

Corey: Yeah, you were one of the co-founders of Gluster—

AB: Yeah.

Corey: —which I have only begrudgingly forgiven you. But please continue.

AB: [laugh]. And back then we got the idea right, but the timing was wrong. And I had—while the data was beginning to grow at a crazy rate, end of the day, GlusterFS has to still look like an FS, it has to look like a file system like NetApp or EMC, and it was hugely limiting what we can do with it. The biggest problem for me was legacy systems. I have to build a modern system that is compatible with a legacy architecture, you cannot innovate.

And that is where when Amazon introduced S3, back then, like, when S3 came, cloud was not big at all, right? When I look at it, the most important message of the cloud was Amazon basically threw everything that is legacy. It’s not [iSCSI 00:03:21] as a Service; it’s not even FTP as a Service, right? They came up with a simple, RESTful API to store your blobs, whether it’s JavaScript, Android, iOS, or [AAML 00:03:30] application, or even Snowflake-type application.

Corey: Oh, we spent ten years rewriting our apps to speak object store, and then they released EFS, which is NFS in the cloud. It’s—

AB: Yeah.

Corey: —I didn’t realize I could have just been stubborn and waited, and the whole problem would solve itself. But here we are. You’re quite right.

AB: Yeah. And even EFS and EBS are more for legacy stock can come in, buy some time, but that’s not how you should stay on AWS, right? When Amazon did that, for me, that was the opportunity. I saw that… while world is going to continue to produce lots and lots of data, if I built a brand around that, I’m not going to go wrong.

The problem is data at scale. And what do I do there? The opportunity I saw was, Amazon solved one of the largest problems for a long time. All the legacy systems, legacy protocols, they convinced the industry, throw them away and then start all over from scratch with the new API. While it’s not compatible, it’s not standard, it is ridiculously simple compared to anything else.

No fstabs, no [unintelligible 00:04:27], no [root 00:04:28], nothing, right? From any application anywhere you can access was a big deal. When I saw that, I was like, “Thank you Amazon.” And I also knew Amazon would convince the industry that rewriting their application is going to be better and faster and cheaper than retrofitting legacy applications.

Corey: I wonder how much that’s retconned because talking to some of the people involved in the early days, they were not at all convinced they [laugh] would be able to convince the industry to do this.

AB: Actually, if you talk to the analyst reporters, the IDC’s, Gartner’s of the world to the enterprise IT, the VMware community, they would say, “Hell no.” But if you talk to the actual application developers, data infrastructure, data architects, the actual consumers of data, for them, it was so obvious. They actually did not know how to write an fstab. The iSCSI and NFS, you can’t even access across the internet, and the modern applications, they ran across the globe, in JavaScript, and all kinds of apps on the device. From [Snap 00:05:21] to Snowflake, today is built on object store. It was more natural for the applications team, but not from the infrastructure team. So, who you asked that mattered.

But nevertheless, Amazon convinced the rest of the world, and our bet was that if this is going to be the future, then this is also our opportunity. S3 is going to be limited because it only runs inside AWS. Bulk of the world’s data is produced everywhere and only a tiny fraction will go to AWS. And where will the rest of the data go? Not SAN, NAS, HDFS, or other blob store, Azure Blob, or GCS; it’s not going to be fragmented. And if we built a better object store, lightweight, faster, simpler, but fully compatible with S3 API, we can sweep and consolidate the market. And that’s what happened.

Corey: And there is a lot of validity to that. We take a look across the industry, when we look at various standards—I mean, one of the big problems with multi-cloud in many respects is the APIs are not quite similar enough. And worse, the failure patterns are very different, of I don’t just need to know how the load balancer works, I need to know how it breaks so I can detect and plan for that. And then you’ve got the whole identity problem as well, where you’re trying to manage across different frames of reference as you go between providers, and leads to a bit of a mess. What is it that makes MinIO something that has been not just something that has endured since it was created, but clearly been thriving?

AB: The real reason, actually is not the multi-cloud compatibility, all that, right? Like, while today, it is a big deal for the users because the deployments have grown into 10-plus petabytes, and now the infrastructure team is taking it over and consolidating across the enterprise, so now they are talking about which key management server for storing the encrypted keys, which key management server should I talk to? Look at AWS, Google, or Azure, everyone has their own proprietary API. Outside they, have [YAML2 00:07:18], HashiCorp Vault, and, like, there is no standard here. It is supposed to be a [KMIP 00:07:23] standard, but in reality, it is not. Even different versions of Vault, there are incompatibilities for us.

That is where—like from Key Management Server, Identity Management Server, right, like, everything that you speak around, how do you talk to different ecosystem? That, actually, MinIO provides connectors; having the large ecosystem support and large community, we are able to address all that. Once you bring MinIO into your application stack like you would bring Elasticsearch or MongoDB or anything else as a container, your application stack is just a Kubernetes YAML file, and you roll it out on any cloud, it becomes easier for them, they’re able to go to any cloud they want. But the real reason why it succeeded was not that. They actually wrote their applications as containers on Minikube, then they will push it on a CI/CD environment.

They never wrote code on EC2 or ECS writing objects on S3, and they don’t like the idea of [past 00:08:15], where someone is telling you just—like you saw Google App Engine never took off, right? They liked the idea, here are my building blocks. And then I would stitch them together and build my application. We were part of their application development since early days, and when the application matured, it was hard to remove. It is very much like Microsoft Windows when it grew, even though the desktop was Microsoft Windows Server was NetWare, NetWare lost the game, right?

We got the ecosystem, and it was actually developer productivity, convenience, that really helped. The simplicity of MinIO, today, they are arguing that deploying MinIO inside AWS is easier through their YAML and containers than going to AWS Console and figuring out how to do it.

Corey: As you take a look at how customers are adopting this, it’s clear that there is some shift in this because I could see the story for something like MinIO making an awful lot of sense in a data center environment because otherwise, it’s, “Great. I need to make this app work with my SAN as well as an object store.” And that’s sort of a non-starter for obvious reasons. But now you’re available through cloud marketplaces directly.

AB: Yeah.

Corey: How are you seeing adoption patterns and interactions from customers changing as the industry continues to evolve?

AB: Yeah, actually, that is how my thinking was when I started. If you are inside AWS, I would myself tell them that why don’t use AWS S3? And it made a lot of sense if it’s on a colo or your own infrastructure, then there is an object store. It even made a lot of sense if you are deploying on Google Cloud, Azure, Alibaba Cloud, Oracle Cloud, it made a lot of sense because you wanted an S3 compatible object store. Inside AWS, why would you do it, if there is AWS S3?

Nowadays, I hear funny arguments, too. They like, “Oh, I didn’t know that I could use S3. Is S3 MinIO compatible?” Because they will be like, “It came along with the GitLab or GitHub Enterprise, a part of the application stack.” They didn’t even know that they could actually switch it over.

And otherwise, most of the time, they developed it on MinIO, now they are too lazy to switch over. That also happens. But the real reason that why it became serious for me—I ignored that the public cloud commercialization; I encouraged the community adoption. And it grew to more than a million instances, like across the cloud, like small and large, but when they start talking about paying us serious dollars, then I took it seriously. And then when I start asking them, why would you guys do it, then I got to know the real reason why they wanted to do was they want to be detached from the cloud infrastructure provider.

They want to look at cloud as CPU network and drive as a service. And running their own enterprise IT was more expensive than adopting public cloud, it was productivity for them, reducing the infrastructure, people cost was a lot. It made economic sense.

Corey: Oh, people always cost more the infrastructure itself does.

AB: Exactly right. 70, 80%, like, goes into people, right? And enterprise IT is too slow. They cannot innovate fast, and all of those problems. But what I found was for us, while we actually build the community and customers, if you’re on AWS, if you’re running MinIO on EBS, EBS is three times more expensive than S3.

Corey: Or a single copy of it, too, where if you’re trying to go multi-AZ and you have the replication traffic, and not to mention you have to over-provision it, which is a bit of a different story as well. So, like, it winds up being something on the order of 30 times more expensive, in many cases, to do it right. So, I’m looking at this going, the economics of running this purely by itself in AWS don’t make sense to me—long experience teaches me the next question of, “What am I missing?” Not, “That’s ridiculous and you’re doing it wrong.” There’s clearly something I’m not getting. What am I missing?

AB: I was telling them until we made some changes, right—because we saw a couple of things happen. I was initially like, [unintelligible 00:12:00] does not make 30 copies. It makes, like, 1.4x, 1.6x.

But still, the underlying block storage is not only three times more expensive than S3, it’s also slow. It’s a network storage. Trying to put an object store on top of it, another, like, software-defined SAN, like EBS made no sense to me. Smaller deployments, it’s okay, but you should never scale that on EBS. So, it did not make economic sense. I would never take it seriously because it would never help them grow to scale.

But what changed in recent times? Amazon saw that this was not only a problem for MinIO-type players. Every database out there today, every modern database, even the message queues like Kafka, they all have gone scale-out. And they all depend on local block store and putting a scale-out distributed database, data processing engines on top of EBS would not scale. And Amazon introduced storage optimized instances. Essentially, that reduced to bet—the data infrastructure guy, data engineer, or application developer asking IT, “I want a SuperMicro, or Dell server, or even virtual machines.” That’s too slow, too inefficient.

They can provision these storage machines on demand, and then I can do it through Kubernetes. These two changes, all the public cloud players now adopted Kubernetes as the standard, and they have to stick to the Kubernetes API standard. If they are incompatible, they won’t get adopted. And storage optimized that is local drives, these are machines, like, [I3 EN 00:13:23], like, 24 drives, they have SSDs, and fast network—like, 25-gigabit 200-gigabit type network—availability of these machines, like, what typically would run any database, HDFS cluster, MinIO, all of them, those machines are now available just like any other EC2 instance.

They are efficient. You can actually put MinIO side by side to S3 and still be price competitive. And Amazon wants to—like, just like their retail marketplace, they want to compete and be open. They have enabled it. In that sense, Amazon is actually helping us. And it turned out that now I can help customers build multiple petabyte infrastructure on Amazon and still stay efficient, still stay price competitive.

Corey: I would have said for a long time that if you were to ask me to build out the lingua franca of all the different cloud providers into a common API, the S3 API would be one of them. Now, you are building this out, multi-cloud, you’re in all three of the major cloud marketplaces, and the way that you do that and do those deployments seems like it is the modern multi-cloud API of Kubernetes. When you first started building this, Kubernetes was very early on. What was the evolution of getting there? Or were you one of the first early-adoption customers in a Kubernetes space?

AB: So, when we started, there was no Kubernetes. But we saw the problem was very clear. And there was containers, and then came Docker Compose and Swarm. Then there was Mesos, Cloud Foundry, you name it, right? Like, there was many solutions all the way up to even VMware trying to get into that space.

And what did we do? Early on, I couldn’t choose. I couldn’t—it’s not in our hands, right, who is going to be the winner, so we just simply embrace everybody. It was also tiring that to allow implement native connectors to all of them different orchestration, like Pivotal Cloud Foundry alone, they have their own standard open service broker that’s only popular inside their system. Go outside elsewhere,
everybody was incompatible.

And outside that, even, Chef Ansible Puppet scripts, too. We just simply embraced everybody until the dust settle down. When it settled down, clearly a declarative model of Kubernetes became easier. Also Kubernetes developers understood the community well. And coming from Borg, I think they understood the right architecture. And also written in Go, unlike Java, right?

It actually matters, these minute new details resonating with the infrastructure community. It took off, and then that helped us immensely. Now, it’s not only Kubernetes is popular, it has become the standard, from VMware to OpenShift to all the public cloud providers, GKS, AKS, EKS, whatever, right—GKE. All of them now are basically Kubernetes standard. It made not only our life easier, it made every other [ISV 00:16:11], other open-source project, everybody now can finally write one code that can be operated portably.

It is a big shift. It is not because we chose; we just watched all this, we were riding along the way. And then because we resonated with the infrastructure community, modern infrastructure is dominated by open-source. We were also the leading open-source object store, and as Kubernetes community adopted us, we were naturally embraced by the community.

Corey: Back when AWS first launched with S3 as its first offering, there were a bunch of folks who were super excited, but object stores didn’t make a lot of sense to them intrinsically, so they looked into this and, “Ah, I can build a file system and users base on top of S3.” And the reaction was, “Holy God don’t do that.” And the way that AWS decided to discourage that behavior is a per request charge, which for most workloads is fine, whatever, but there are some that causes a significant burden. With running something like MinIO in a self-hosted way, suddenly that costing doesn’t exist in the same way. Does that open the door again to so now I can use it as a file system again, in which case that just seems like using the local file system, only with extra steps?

AB: Yeah.

Corey: Do you see patterns that are emerging with customers' use of MinIO that you would not see with the quote-unquote, “Provider’s” quote-unquote, “Native” object storage option, or do the patterns mostly look the same?

AB: Yeah, if you took an application that ran on file and block and brought it over to object storage, that makes sense. But something that is competing with object store or a layer below object store, that is—end of the day that drives our block devices, you have a block interface, right—trying to bring SAN or NAS on top of object store is actually a step backwards. They completely missed the message that Amazon told that if you brought a file system interface on top of object store, you missed the point, that you are now bringing the legacy things that Amazon intentionally removed from the infrastructure. Trying to bring them on top doesn’t make it any better. If you are arguing from a compatibility some legacy applications, sure, but writing a file system on top of object store will never be better than NetApp, EMC, like EMC Isilon, or anything else. Or even GlusterFS, right?

But if you want a file system, I always tell the community, they ask us, “Why don’t you add an FS option and do a multi-protocol system?” I tell them that the whole point of S3 is to remove all those legacy APIs. If I added POSIX, then I’ll be a mediocre object storage and a terrible file system. I would never do that. But why not write a FUSE file system, right? Like, S3Fs is there.

In fact, initially, for legacy compatibility, we wrote MinFS and I had to hide it. We actually archived the repository because immediately people started using it. Even simple things like end of the day, can I use Unix [Coreutils 00:19:03] like [cp, ls 00:19:04], like, all these tools I’m familiar with? If it’s not file system object storage that S3 [CMD 00:19:08] or AWS CLI is, like, to bloatware. And it’s not really Unix-like feeling.

Then what I told them, “I’ll give you a BusyBox like a single static binary, and it will give you all the Unix tools that works for local filesystem as well as object store.” That’s where the [MC tool 00:19:23] came; it gives you all the Unix-like programmability, all the core
tool that’s object storage compatible, speaks native object store. But if I have to make object store look like a file system so UNIX tools would run, it would not only be inefficient, Unix tools never scaled for this kind of capacity.

So, it would be a bad idea to take step backwards and bring legacy stuff back inside. For some very small case, if there are simple POSIX calls using [ObjectiveFs 00:19:49], S3Fs, and few, for legacy compatibility reasons makes sense, but in general, I would tell the community don’t bring file and block. If you want file and block, leave those on virtual machines and leave that infrastructure in a silo and gradually phase them out.

Corey: This episode is sponsored in part by our friends at Vultr. Spelled V-U-L-T-R because they’re all about helping save money, including on things like, you know, vowels. So, what they do is they are a cloud provider that provides surprisingly high performance cloud compute at a price that—while sure they claim its better than AWS pricing—and when they say that they mean it is less money. Sure, I don’t dispute that but what I find interesting is that it’s predictable. They tell you in advance on a monthly basis what it’s going to going to cost. They have a bunch of advanced networking features. They have nineteen global locations and scale things elastically. Not to be confused with openly, because apparently elastic and open can mean the same thing sometimes. They have had over a million users. Deployments take less that sixty seconds across twelve pre-selected operating systems. Or, if you’re one of those nutters like me, you can bring your own ISO and install basically any operating system you want. Starting with pricing as low as $2.50 a month for Vultr cloud compute they have plans for developers and businesses of all sizes, except maybe Amazon, who stubbornly insists on having something to scale all on their own. Try Vultr today for free by visiting: vultr.com/screaming, and you’ll receive a $100 in credit. Thats v-u-l-t-r.com slash screaming.

Corey: So, my big problem, when I look at what S3 has done is in it’s name because of course, naming is hard. It’s, “Simple Storage Service.” The problem I have is with the word simple because over time, S3 has gotten more and more complex under the hood. It automatically tiers data the way that customers want. And integrated with things like Athena, you can now query it directly, whenever of an object appears, you can wind up automatically firing off Lambda functions and the rest.

And this is increasingly looking a lot less like a place to just dump my unstructured data, and increasingly, a lot like this is sort of a database, in some respects. Now, understand my favorite database is Route 53; I have a long and storied history of misusing services as databases. Is this one of those scenarios, or is there some legitimacy to the idea of turning this into a database?

AB: Actually, there is now S3 Select API that if you’re storing unstructured data like CSV, JSON, Parquet, without downloading even a compressed CSV, you can actually send a SQL query into the system. IN MinIO particularly the S3 Select is [CMD 00:21:16] optimized. We can load, like, every 64k worth of CSV lines into registers and do CMD operations. It’s the fastest SQL filter out there. Now, bringing these kinds of capabilities, we are just a little bit away from a database; should we do database? I would tell definitely no.

The very strength of S3 API is to actually limit all the mutations, right? Particularly if you look at database, they’re dealing with metadata, and querying; the biggest value they bring is indexing the metadata. But if I’m dealing with that, then I’m dealing with really small block lots of mutations, the separation of objects storage should be dealing with persistence and not mutations. Mutations are [AWS 00:21:57] problem. Separation of database work function and persistence function is where object storage got the storage right.

Otherwise, it will, they will make the mistake of doing POSIX-like behavior, and then not only bringing back all those capabilities, doing IOPS intensive workloads across the HTTP, it wouldn’t make sense, right? So, object storage got the API right. But now should it be a database? So, it definitely should not be a database. In fact, I actually hate the idea of Amazon yielding to the file system developers and giving a [file three 00:22:29] hierarchical namespace so they can write nice file managers.

That was a terrible idea. Writing a hierarchical namespace that’s also sorted, now puts tax on how the metadata is indexed and organized. The Amazon should have left the core API very simple and told them to solve these problems outside the object store. Many application developers don’t need. Amazon was trying to satisfy everybody’s need. Saying no to some of these file system-type, file manager-type users, what should have been the right way.

But nevertheless, adding those capabilities, eventually, now you can see, S3 is no longer simple. And we had to keep that compatibility, and I hate that part. I actually don’t mind compatibility, but then doing all the wrong things that Amazon is adding, now I have to add because it’s compatible. I kind of hate that, right?

But now going to a database would be pushing it to the whole new level. Here is the simple reason why that’s a bad idea. The right way to do database—in fact, the database industry is already going in the right direction. Unstructured data, the key-value or graph, different types of data, you cannot possibly solve all that even in a single database. They are trying to be multimodal database; even they are struggling with it.

You can never be a Redis, Cassandra, like, a SQL all-in-one. They tried to say that but in reality, that you will never be better than any one of those focused database solutions out there. Trying to bring that into object store will be a mistake. Instead, let the databases focus on query language implementation and query computation, and leave the persistence to object store. So, object store can still focus on storing your database segments, the table segments, but the index is still in the memory of the database.

Even the index can be snapshotted once in a while to object store, but use objects store for persistence and database for query is the right architecture. And almost all the modern databases now, from Elasticsearch to [unintelligible 00:24:21] to even Kafka, like, message queue. They all have gone that route. Even Microsoft SQL Server, Teradata, Vertica, name it, Splunk, they all have gone object storage route, too. Snowflake itself is a prime example, BigQuery and all of them.

That’s the right way. Databases can never be consolidated. There will be many different kinds of databases. Let them specialize on GraphQL or Graph API, or key-value, or SQL. Let them handle the indexing and persistence, they cannot handle petabytes of data. That [unintelligible 00:24:51] to object store is how the industry is shaping up, and it is going in the right direction.

Corey: One of the ways I learned the most about various services is by talking to customers. Every time I think I’ve seen something, this is amazing. This service is something I completely understand. All I have to do is talk to one more customer. And when I was doing a bill analysis project a couple of years ago, I looked into a customer’s account and saw a bucket with okay, that has 280 billion objects in it—and wait was that billion with a B?

And I asked them, “So, what’s going on over there?” And there’s, “Well, we built our own columnar database on top of S3. This may not have been the best approach.” It’s, “I’m going to stop you there. With no further context, it was not, but please continue.”

It’s the sort of thing that would never have occurred to me to even try, do you tend to see similar—I would say they’re anti-patterns, except somehow they’re made to work—in some of your customer environments, as they are using the service in ways that are very different than ways encouraged or even allowed by the native object store options?

AB: Yeah, when I first started seeing the database-type workloads coming on to MinIO, I was surprised, too. That was exactly my reaction. In fact, they were storing these 256k, sometimes 64k table segments because they need to index it, right, and the table segments were anywhere between 64k to 2MB. And when they started writing table segments, it was more often [IOPS-type 00:26:22] I/O pattern, then a throughput-type pattern. Throughput is an easier problem to solve, and MinIO always saturated these 100-gigabyte NVMe-type drives, they were I/O intensive, throughput optimized.

When I started seeing the database workloads, I had to optimize for small-object workloads, too. We actually did all that because eventually I got convinced the right way to build a database was to actually leave the persistence out of database; they made actually a compelling argument. If historically, I thought metadata and data, data to be very big and coming to object store make sense. Metadata should be stored in a database, and that’s only index page. Take any book, the index pages are only few, database can continue to run adjacent to object store, it’s a clean architecture.

But why would you put database itself on object store? When I saw a transactional database like MySQL, changing the [InnoDB 00:27:14] to [RocksDB 00:27:15], and making changes at that layer to write the SS tables [unintelligible 00:27:19] to MinIO, and then I was like, where do you store the memory, the journal? They said, “That will go to Kafka.” And I was like—I thought that was insane when it started. But it continued to grow and grow.

Nowadays, I see most of the databases have gone to object store, but their argument is, the databases also saw explosive growth in data. And they couldn’t scale the persistence part. That is where they realized that they still got very good at the indexing part that object storage would never give. There is no API to do sophisticated query of the data. You cannot peek inside the data, you can just do streaming read and write.

And that is where the databases were still necessary. But databases were also growing in data. One thing that triggered this was the use case moved from data that was generated by people to now data generated by machines. Machines means applications, all kinds of devices. Now, it’s like between seven billion people to a trillion devices is how the industry is changing. And this led to lots of machine-generated, semi-structured, structured data at giant scale, coming into database. The databases need to handle scale. There was no other way to solve this problem other than leaving the—[unintelligible 00:28:31] if you looking at columnar data, most of them are machine-generated data, where else would you store? If they tried to build their own object storage embedded into the database, it would make database mentally complicated. Let them focus on what they are good at: Indexing and mutations. Pull the data table segments which are immutable, mutate in memory, and then commit them back give the right mix. What you saw what’s the fastest step that happened, we saw that consistently across. Now, it is actually the standard.

Corey: So, you started working on this in 2014, and here we are—what is it—eight years later now, and you’ve just announced a Series B of $100 million dollars on a billion-dollar valuation. So, it turns out this is not just one of those things people are using for test labs; there is significant momentum behind using this. How did you get there from—because everything you’re saying makes an awful lot of sense, but it feels, at least from where I sit, to be a little bit of a niche. It’s a bit of an edge case that is not the common case. Obviously, I missing something because your investors are not the types of sophisticated investors who see something ridiculous and, “Yep. That’s the thing we’re going to go for.” There right more than they’re not.

AB: Yeah. The reason for that was the saw what we were set to do. In fact, these are—if you see the lead investor, Intel, they watched us grow. They came into Series A and they saw, everyday, how we operated and grew. They believed in our message.

And it was actually not about object store, right? Object storage was a means for us to get into the market. When we started, our idea was, ten years from now, what will be a big problem? A lot of times, it’s hard to see the future, but if you zoom out, it’s hidden in plain sight.

These are simple trends. Every major trend pointed to world producing more data. No one would argue with that. If I solved one important problem that everybody is suffering, I won’t go wrong. And when you solve the problem, it’s about building a product with fine craftsmanship, attention to details, connecting with the user, all of that standard stuff.

But I picked object storage as the problem because the industry was fragmented across many different data stores, and I knew that won’t be the case ten years from now. Applications are not going to adopt different APIs across different clouds, S3 to GCS to Azure Blob to HDFS to everything is incompatible. I saw that if I built a data store for persistence, industry will consolidate around S3 API. Amazon S3, when we started, it looked like they were the giant, there was only one cloud industry, it believed mono-cloud. Almost everyone was talking to me like AWS will be the world’s data center.

I certainly see that possibility, Amazon is capable of doing it, but my bet was the other way, that AWS S3 will be one of many solutions, but not—if it’s all incompatible, it’s not going to work, industry will consolidate. Our bet was, if world is producing so much data, if you build an object store that is S3 compatible, but ended up as the leading data store of the world and owned the application ecosystem, you cannot go wrong. We kept our heads low and focused on the first six years on massive adoption, build the ecosystem to a scale where we can say now our ecosystem is equal or larger than Amazon, then we are in business. We didn’t focus on commercialization; we focused on convincing the industry that this is the right technology for them to use. Once they are convinced, once you solve business problems, making money is not hard because they are already sold, they are in love with the product, then convincing them to pay is not a big deal because data is so critical, central part of their business.

We didn’t worry about commercialization, we worried about adoption. And once we got the adoption, now customers are coming to us and they’re like, “I don’t want open-source license violation. I don’t want data breach or data loss.” They are trying to sell to me, and it’s an easy relationship game. And it’s about long-term partnership with customers.

And so the business started growing, accelerating. That was the reason that now is the time to fill up the gas tank and investors were quite excited about the commercial traction as well. And all the intangible, right, how big we grew in the last few years.

Corey: It really is an interesting segment, that has always been something that I’ve mostly ignored, like, “Oh, you want to run your own? Okay, great.” I get it; some people want to cosplay as cloud providers themselves. Awesome. There’s clearly a lot more to it than that, and I’m really interested to see what the future holds for you folks.

AB: Yeah, I’m excited. I think end of the day, if I solve real problems, every organization is moving from compute technology-centric to data-centric, and they’re all looking at data warehouse, data lake, and whatever name they give data infrastructure. Data is now the centerpiece. Software is a commodity. That’s how they are looking at it. And it is translating to each of these large organizations—actually, even the mid, even startups nowadays have petabytes of data—and I see a huge potential here. The timing is perfect for us.

Corey: I’m really excited to see this continue to grow. And I want to thank you for taking so much time to speak with me today. If people want to learn more, where can they find you?

AB: I’m always on the community, right. Twitter and, like, I think the Slack channel, it’s quite easy to reach out to me. LinkedIn. I’m always excited to talk to our users or community.

Corey: And we will of course put links to this in the [show notes 00:33:58]. Thank you so much for your time. I really appreciate it.

AB: Again, wonderful to be here, Corey.

Corey: Anand Babu Periasamy, CEO and co-founder of MinIO. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with what starts out as an angry comment but eventually turns into you, in your position on the S3 product team, writing a thank you note to MinIO for helping validate your market.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Natalie

I'm interested in solving human problems through technology (she/her). Share your screen (or I'll share mine) and we'll figure this out!

Links:

  • Netlify: https://www.netlify.com/
  • Twitter: https://twitter.com/codeFreedomRitr

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Sysdig. Sysdig is the solution for securing DevOps. They have a blog post that went up recently about how an insecure AWS Lambda function could be used as a pivot point to get access into your environment. They’ve also gone deep in-depth with a bunch of other approaches to how DevOps and security are inextricably linked. To learn more, visit sysdig.com and tell them I sent you. That’s S-Y-S-D-I-G dot com. My thanks to them for their continued support of this ridiculous nonsense.

Corey: It seems like there is a new security breach every day. Are you confident that an old SSH key or a shared admin account isn’t going to come back and bite you? If not, check out Teleport. Teleport is the easiest, most secure way to access all of your infrastructure. The open source Teleport Access Plane consolidates everything you need for secure access to your Linux and Windows servers—and I assure you there is no third option there. Kubernetes clusters, databases, and internal applications like AWS Management Console, Yankins, GitLab, Grafana, Jupyter Notebooks, and more. Teleport’s unique approach is not only more secure, it also improves developer productivity. To learn more visit: goteleport.com. And no, that is not me telling you to go away, it is: goteleport.com.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. A recurring theme of this show has been where does the next generation of cloud engineer come from because of the road that a lot of us walked is closed, and a lot of the jobs that some of us took no longer exist in any meaningful form. There are a bunch of answers around oh, we’re going to get people right out of school from computer science programs into this space, but that doesn’t always solve some of the answers. Here to talk to me today is someone who took a different path. Natalie Davis is a software engineer at Netlify, and she entered tech by changing careers from another industry. Natalie, how are you? Thank you for joining me.

Natalie: I’m really good, Corey. Thanks for having me. I’m very excited to be here and kind of share my experiences.

Corey: So, you have entered tech within the last few years. You went to a boot camp, you spent a year as an engineer at a different company, and now you’re at Netlify, one of those companies that, at least for some of us was one of those things you vaguely hear about in the background, sort of a buzz, and the buzz gets louder and louder and louder, and no seems that every time I turn around, I’m tripping over Netlify. In good ways, to be clear.

Natalie: I mean, that’s definitely good news for me. [laugh]. Yeah, Netlify is a company I first grew familiar with while I was in boot camp. It was the first place I ever hosted a website, a nice little to-do app. And now a couple of years later, here I am, in the guts of it.

Corey: So, what were you doing before you decided, “You know what? I’m going to enter tech.” Because if you stand back and you look at it, like that seems like a great culture with no problems whatsoever inherent to it in any way, shape, or form. That’s where I want to be. Honestly, I find myself in tech these days, in spite of a lot of things rather than because of it. But again, I am cynical, jaded, again, old and grumpy because you don’t get to be a Unix sysadmin without being old and grumpy by somewhere around week three.

Natalie: So, that’s something I actually find very interesting. Because I came to tech after having existed in another industry—and I’ll talk about that in a moment—for about 15 years, I don’t find tech as toxic as people who have always been in tech find it. There are problems in tech, but we’re talking about those problems; we’re trying to come up with solutions. Whereas in retail, where I spent the first 15 years of my career, no one’s talking about those problems. And they exist, and they exist on an amplified level because not only are people being treated horribly, not only are people consistently being profiled and discriminated against, but they’re doing it for $10 an hour, so there’s not even the incentive of at least I get to live well. So, I always push back just a little bit on that, tech is so toxic.

Corey: That is a fantastic approach. I hadn’t considered it from that perspective. I mean, I sit here in something of an ivory tower. My clients tend to be big companies doing things in a B2B level, whether I’m talking about media sponsorships or consulting projects. The one time a year that I deal with the quote-unquote, “General public,” or a B2C type of thing is my annual charity t-shirt fundraiser.

And I have remarked before on this show that those $35 t-shirts cause more customer service headaches for me than the entire rest of the year put together because you sell someone $100,000 consulting project, and you’re responsible adults, and you can have conversations and figure out how to move forward, but when someone spends $35 on a shirt—for charity, I will point out—and it doesn’t show up, or it’s the wrong size or something, they have opinions, and they will in some cases put you on blast. But even in that sense, it’s not the quote-unquote, “General public,” it’s people in this industry, by and large, who are themselves working professionals, not people walking into a retail store and deciding the best way to get what they want is to basically abuse the staff.

Natalie: Yeah, yeah. I noticed that even within retail. I spent most of my retail career in better or luxury retail, but there was one year that I worked in an outlet—and I won’t name them—but that was the worst experience of my life. People calling corporate on me over 40 cent discounts. It was just unbelievable. [laugh].

Corey: It’s a different era, so coming from that, you look at tech and your perspective then is that you see that it has challenges in it, but it’s, “Oh, compared to what I used to deal with, this is nothing.”

Natalie: Correct. Although I did know that there were challenges in tech, but I viewed it more from a standpoint of how tech was impacting communities like mine. And that was part of what drew me to tech because obviously, there weren’t enough people like me in the room, and that meant that there was room for someone like me to enter the room and shake some tables. So, that was part of why I wanted to come to tech.

Corey: This is evocative of other conversations I’ve had, generally with people in the midst of an outage, where everyone’s running around with their hair on fire because the computers aren’t working, and there’s one person sitting there who’s just, you would think it is any random Tuesday, and at people ask them, “How on earth are you so calm?” And their answer is, “Oh, I’m a veteran. No one’s shooting at me. The computers don’t work. I know everyone here is going to go home to their families tonight. This isn’t stress. You haven’t seen stress.”

I have seen shades of that from folks who have transitioned into this industry from, honestly, industries that treat people far worse. So, that’s an area I haven’t considered. I’d like the direction, I like the angle you have on this. This is sort of a strange follow-up to that, but what inspired you to enter tech from retail? I mean, the easy answer is you look around, you’re like, “Okay, I’ve had enough of this, I’m going to go learn how tech works.” It’s never that easy.

Natalie: Yeah, it definitely wasn’t that easy. So, I married a wonderful man who is a firefighter. My brother-in-law works with non-traditional students at the high school age, his wife is a nurse. So, I’m surrounded by these people who actually have careers, who actually are doing things that they’re passionate about. And that wasn’t a part of my life before marrying into this family.

So, it kind of woke something up in me like, hey, I don’t just have to work for a living; I can work for a passion. And no, no one dreams of labor, sure. Like, one day, I’ll win the lotto and I won’t have to do anything except be a professional student, which would be my ideal path, but it did awaken the possibility that even people in my life can go have these passions. So, then I started thinking, “Well, what can I do aside from retail, without incurring another $100,000 worth of college debt?” And then I started—I jumped on Twitter. Following tech accounts now, and—

Corey: Oh, geez, you are a glutton for punishment. It’s one of those, “All right. So, I don’t think the industry is that bad. I’m going to prove it by going on Twitter.” Okay, let’s scrap it on that one.

Natalie: But around this time was the time where there was an article about automatic hand dryers and how they weren’t recognizing black hands as hands. And I think maybe there was something about an automated self-driving car—that’s what I’m looking for—that wasn’t recognizing black people as people in the same way that it was recognizing others. And I’ve always been a fighter. I’ve always been a rebel. You might not be able to tell it now I seem to have grown up quite a bit, and you know, I’m more conservative with the way I respond to the issues that I see in the world.

If I’m going to pursue my passion, it needs to be me fighting for something that’s important to me. Tech, okay, cool. Then there’s this thing about tech where, sure you can go the CS degree route, and I think that’s a great route. I don’t think it’s the right route for everybody. There’s almost like this Wild West aspect where if you can build, that’s it. If you can do the job, you can do the job.

And I didn’t think that it was going to be easy, but I know I’ve got grit, I know, I’ve got determination. I know if I set my mind to a thing, I can do a thing. And I liked that you could come in and just be able to do the work, and that would be enough. So, I jumped in a boot camp.

Corey: Would you recommend boot camps as a way for people to break into tech? The reason I asked i—I’m not talking about any particular boot camp here—

Natalie: Sure.

Corey: —but I’m interested in what is the common guidance for folks who find themselves in similar situations and decide that, “You know what? I think that I want to go deal with tech because tech does have its problems, but people aren’t literally spitting on you, most days, or throwing drinks at you and, let’s be very direct because there’s a taboo against talking about this sometimes the pay is a lot better in tech than it is in most other industries.” And we all like to—

Natalie: Oh yeah.

Corey: —dance around the fact that, “Oh, compensation. No, no, no. You should do it because you love it.” It’s, yeah, being able to do what you love is one of those privileges that comes along with having money and making money doing the thing that you love. If the thing that you love is getting screamed at on Black Friday by hordes of people, great. You’re still going to not necessarily be able to afford the same trappings of a life that you can by having something that compensates better.

Natalie: Thank you for bringing that up because I certainly should have mentioned that the pay was attractive to me in the industry as well. Like, I thought only doctors and lawyers made six figures or better. I didn’t realize I could get there.

Corey: I’ve always had the baseline assumption that everyone is in tech to some degree for the money. Whenever I meet someone who’s like, “No, I’m in tech and I’m not doing it for the money.” I like to follow up with that because sometimes they’re right. “Really? So, what do you do?” Like, “Oh, yeah, I work for this nonprofit doing tech stuff.” “Okay. I believe you when you say that.” When I work for one of the FAANG big tech companies, and people are, “Oh, yeah, I’m here because I love the work.” [pause] “Really? Like, you’re out there making the world a better place by improving ad conversion rates? Okay.”

Like, we all tell ourselves lies to get through the day, and I’m also not suggesting by any means that money is a bad motivator for anything. The thing that always irked me is when people don’t acknowledge, yeah, part of the reason I’m in this industry is because it pays riches beyond the wildest dreams of avarice that I had growing up. I never expected to find myself in a situation where I’m making, as you say, lawyer and doctor money. Honestly, I look around and I’m still astounded that the things that I do on computers—badly, may I point out—is valued by anyone. Yet, here we are.

Natalie: I wholeheartedly agree. Every time that direct deposit hits my account, my mind is just blown. Like, “You all know I was just putzing around on my computer all week, right? And like, this is what I get? Cool. Cool.” But to get back to your question is, boot camp—I’m sorry, I don’t remember exactly how you phrased it.

Corey: No, no, the question I really have is, is boot camp the common case recommendation now for folks who want to break in? Are there better slash alternate paths—if you had to do it all again—that you might have pursued?

Natalie: I have to say, people reach out to me for advice: How did you do what you did, they never liked what I have to say because I’m going to start with, you have to understand who you are. You have to understand what works for you. I know that I’m incredibly capable, and I learn quite well, but I need structure in order to do so because if you leave me to my own devices, I will get lost in the weeds of something that does not matter much, but it’s quite interesting. And now I’ve spent a month learning about event handlers, but I don’t know how to do anything else. So, for me, boot camp provided both the structure and the baked-in community that I need it because no one in my life is in tech; no one can talk to me about these things. I needed a group of people who I could share the struggle that learning to code is. Because my God, that was a struggle. I’ve done a lot of hard things in my life, and I don’t think many of them had me doubting my abilities the way learning to code did.

Corey: There’s always that constant ebb and flow of it, where you—it’s a rush, like, “I am a genius,” and then something doesn’t work it, “Oh, I’m a fool. Why didn’t anyone bother to tell me this at any point in my life?” And it’s the constant, almost swing between highs and lows on a constant basis. There’s a support group for that in tech, it’s called everyone, and we made it the bar.

Natalie: [laugh]. Yeah, I haven’t stopped experiencing that since I’ve gotten—although I’ve gotten much better with dealing with the emotions that come along with that.

Corey: Yes, sometimes I find going for a walk and calming down helps because if I keep staring at this thing, I’m going to say something unfortunate, possibly on Twitter, and no one wants that.

Natalie: Well, I kind of want it. It’s fun to watch. [laugh].

Corey: Yeah, but it’s tied to my name, and that’s the challenge.

Natalie: Ah, yes, yes. So yeah, I mean, there are people out there who have gone the self-taught route, and oh, my goodness, those people are so inspiring and amazing to me because I don’t think I could have pulled it off that way. I think something else you have to think about is the support system you have. I don’t know that I would have been able to dedicate myself the way I did in boot camp if I didn’t have my husband, who was able to kind of shoulder the financial burden on our family, while I was just living in this office for 14 hours a day. And that’s unfortunate, and I think that’s something that I hope gets addressed by someone. I don’t know who; I don’t have the solution.

But yeah, it took a certain level of privilege for me to pour myself in the way that I did. So, that’s something that you have to think about, what kind of time do you have to dedicate? Now, when you’re thinking about that, also understand that it’s a marathon, not a race, right? It doesn’t matter if Billy did it in a year, if it takes you five years to get there, that’s how long it took you to get there. But once you’re there, you’re there.

Corey: There are certain one-way doors that people pass through. Another common one that we see a lot of in the industry is the idea of going from engineer to management. Once you have crossed through that door and become a manager, you can go back to being an engineer and then back to being a manager, but crossing into the management realm the first time is one of those things that is not clearly defined in many places. And every time you talk to somebody like, “How do you break that barrier?” And the answer is, “Oh. I was in the right place at the right time, and I got lucky,” is generally the common answer to it.

I keep looking for ways to systematically get there, and that was interesting to me because I wanted to be a manager very much back in the first part of the 2010s. And I put myself in weird roles chasing that, and I think I wanted to do it for the right reasons, namely, to inspire and to be the manager I wished I’d always had. And it turns out I was really bad at it on a variety of different levels. And okay, this is not for me. I decided to go in a bit of a different direction, even now, the entire company rolls up the reporting chain that does not include me. I have a business partner who handles that. No one has to report to me on a weekly basis, which is really something we should put on our careers page as a benefit to help attract people.

Natalie: [laugh]. Absolutely. I mean, I’m thinking about that, and like, what does my next five years look like? Do I want to go into management role? I’ve got a ton of leadership experience in retail.

It’s not a direct translation, but of course, there are some transferable skills there. But also, it is beautiful to be an individual contributor, to not have to follow up with a team of 12 to see where they’re at and what they’re working on. So, I still haven’t decided where I want to go.

Corey: When I have the privilege of talking to high-level executives about the hardest part on their journey, very often the story they say is that—especially if they started off in the engineering world, where, “Yeah, I love what I do, my job is great, but…” and then they pause a minute, and, “Back in the before times, it was easier.” [unintelligible 00:16:13] you’re like, “Oh, here. Let me buy you eight drinks.” And then they get really honest. And they say the hard part really is that you don’t get to do anything yourself.

Your only tool to solve all of these problems is delegation. So, you’ve got to build and manage and maintain and develop the team, and then you have to give them context and basically let them go and hope that they can deliver the thing that you need when you need it delivered. And for a lot of us who are used to working on the computer of, I push the button and the computer does what I say—you know, aspirationally, after you wind up fixing it eight times in a row, only to figure out that comma should have been a semicolon. Great—and then you’re, “Oh, yeah. Okay, that makes sense.”

It is hard for folks in an engineering sense to often let go and that leads to things like micromanagement, and the failure mode of a boss who shows up and basically winds up writing code and reverting your commits in the middle of the night and they’re treating main as their feature branch. And yeah, we’ve all seen those weird patterns there. It’s a hard, hard thing to do. You’ve been management in a retail role. Do you aspire to manage people in the tech industry as your career in this zany place evolves?

Natalie: I just haven’t decided, I think in some ways, it makes a lot of sense. I did enjoy mentoring and coaching and helping people level up. That was kind of my specialty. I got a lot of people promoted, and that felt good to see them kind of take off and fly. But I am kind of in love with the, how do I make this thing do what I want it to do.

That digging in and the mystery and the following the trail and console logging 6000 different variables, and then finally, finally, finally, it works, and I don’t know if I want to give that up. Honestly, the thing that pushed me into management and retail, initially, was I can make a lot more money in management than I can as a sales associate. And with that incentive kind of removed—and sure I can make more money as a manager, but money ceases to be the same kind of motivator once your needs are met. Like, I’m in a good place, I don’t have to worry. So, now I have to think about, do I really want to go back to not being able to do the work—because I found it difficult even in retail not to just jump in and make the sale because I know how to make a sale and I can see where you’re going wrong. And I’ve got to let you fail, but then I’ve lost the sale.

So, I don’t know that I want to give up the individual contributor role. But I’m very open. I feel like in this stage of my career, anything is possible. I’m just kind of exploring what’s out there and seeing where it leads.

Corey: This episode is sponsored by our friends at Oracle HeatWave is a new high-performance query accelerator for the Oracle MySQL Database Service, although I insist on calling it “my squirrel.” While MySQL has long been the worlds most popular open source database, shifting from transacting to analytics required way too much overhead and, ya know, work. With HeatWave you can run your OLAP and OLTP—don’t ask me to pronounce those acronyms again—workloads directly from your MySQL database and eliminate the time-consuming data movement and integration work, while also performing 1100X faster than Amazon Aurora and 2.5X faster than Amazon Redshift, at a third of the cost. My thanks again to Oracle Cloud for sponsoring this ridiculous nonsense.

Corey: Very often there’s this mistaken belief that, “All right, I’ve been an engineer, so now I need to be a manager to get promoted.” And they’re orthogonal skills. Whenever I looked at management roles, and the requirements are well, there’s going to be a coding on the whiteboard component to the interview, it’s, “What exactly do you think a manager does here?” Or the, “Oh, yeah. You’re going to be half managing the team and half participating in the team’s work.” It’s great. Those are two jobs. Which one would you rather I fail at?

Because let’s be very realistic here. There’s also a bias, it’s linked to ageism, for sure in this industry, but you look at someone who’s in their 40s or 50s, or 60s or whatever it happens to be, who’s an individual contributor, and you look at them, and there’s a lot of people that see that either overtly or subtly think that oh, yeah, they got lost somewhere along the way. They have gone in a different direction, they missed some opportunities. And I don’t think that’s necessarily fair. I think that it fails to acknowledge exactly what you’re talking about, that there’s a love and a passion behind some of the things you get to deal with and some things you don’t have to deal with when you’re working as an engineer versus working as management.

From my perspective, I’d argue everyone should at least do a stint in management at some point or another just because I have a lot more empathy for those quote-unquote, “Crappy managers” that I had back in the early part of my career, now that I’ve been on the other side of that table. It’s like, I used to be like, “Why would that person fire me?” And now looking at it from that perspective, it’s, “Why did that person wait three whole months to fire me?” It’s one of those areas where I see it now with the broader context.

And it’s strange, I’ve always said I’m a terrible employee, but I would be a much better one now as a result. So, I learned the lesson just in time for it to be completely useless to me, personally, but if I can pass that on to people, that’s why I have a microphone.

Natalie: Absolutely, yeah. There’s a lot of tension, especially when you’re kind of middle-level management because you’re trying to make your people happy, but then you’ve got these demands coming from the top, and they don’t want what your people want at all. And that’s
difficult.

Corey: That was my failure when I would—I failed to manage up completely. I was obstinate as an employee and got myself fired a lot and figured as a manager, I’m going to do exactly the same thing because it’ll work great now.

Natalie: [laugh].

Corey: Yeah, turns out it doesn’t work that way at all for anyone.

Natalie: But I think there’s something else interesting in that perspective in that I came to tech at what is considered a late age. I joined boot camp, I think maybe… I was 38 when I joined boot camp.

Corey: Understand, some people say, “I came to tech late—I was 14 years old—compared to some folks.” And it’s like this whole, “Oh, if you weren’t in the cradle with a keyboard in your hand, you’re too late for this.” And that is some bullshit.

Natalie: I laughed so much. I want to see more people like me join late because I can tell you, I haven’t had the typical boot camp experience. I’ve been extremely fortunate in that I have had a community that’s really supportive of me, but within a week of telling Twitter I was officially looking for work, I had three interviews with three different companies lined up. And that happened because I had previous experience, both in life and in the industry, so I understood how important it was to build my network and what that looked like, and kind of did that consistently throughout the whole time that I was in boot camp. If I had come at the age of 20, or 14, I wouldn’t have had those skills that—kind of—made it relatively—not relatively. That’s easy. That was an easy journey. I’m still blown away, and I pinch myself almost every day to think about the fairy tale entry I’ve had into tech.

But again, it happened because I came at an older age because I had those life skills. So please, if you’re out there and thinking you’re too old, you have to stop listening to people who haven’t lived enough life to understand how life works. You have to understand who you are, understand what your skills are, and then understand that tech is thirsty for those skills.

Corey: I wish that this were a more common approach. At some level, I feel like there are headwinds against people moving into tech later into their career, gatekeeping, and whatnot. And I used to think that it was this, “Oh because, you know, people just want to hire more folks that look like them.” And I’m increasingly realizing that is actually the more benevolent answer; I suspect, there’s at least some element as well, where when someone is new to their career, they’re in their early-20s, fresh out of school, they are not nearly as cynical, they are not as good at drawing boundaries. So, they’ll work for magic equity at a startup that might one day possibly turn into something, earning significantly below market rate salaries, and they’ll be putting in 80 hours a week because they’re building something.

You only do that once or twice in most people’s careers before they realize, wait a minute, that’s kind of a scam. Or they’ll have an exit and the founder buys a yacht and they get enough to buy a used Toyota. And it’s, “Hmm. Seems like that was an awful lot of late nights, weekends, a time away from my family that I could have been spending doing more productive things.” And they work out what it is by the hour that I put in, and it’s like fractions of a penny by the time they’re all done. And it’s, “Yeah, that was ill-advised.”

Natalie: Yeah.

Corey: There’s a cynicism that comes to it, where folks who are further along in their career or come into this industry, from other careers as well, have a lot better understanding of the dynamics of interpersonal relationships in the workplace, as well as understanding that when something smells off, it very well might be off. And early in your career, you just think, “Oh, this is just how it is. This is what workplaces must be. Why didn’t anyone ever tell me that?” To me at least, that’s why mentorship, especially mentorship from people in other companies at times and career growth is just such a critical thing.

Because I used to do the exact same thing till someone took me aside and said, “You know, you just did that thing today at 4:45 and your coworker came up with an emergency it has to be pushed out? Yeah. Watch what happens someone does it to me next.” And he did—great. Because I wasn’t able to get to it—“Okay, when did you first find out about this? When does it need to get done? Why didn’t you mention this earlier because I’m packing up to go home now? Well, I guess it’s not going to get done. I will do it tomorrow instead.”

And that’s not being a jerk; that’s drawing boundaries. And that was transformative to me because I used to think that my job was to just do whatever my boss said, regardless of the rest. Like, call my then fiance, “Oh, sorry. I’m not able to be there for dinner tonight because I’ve got to do this emergency at work.” That’s not an emergency. It’s really not.

Natalie: Yeah.

Corey: Basic stuff like that, but it’s the thing you only learned by working in the workforce and having a career for a period of time because it’s so different than what the public education system is, coming up through it, where it’s basically, comply, obey, et cetera. You aren’t really going to have much luck drawing boundaries when you don’t do your homework at night.

Natalie: Absolutely. I mean, two of the things that you just said that I love is, when you come to it after having lived a bit of life, you absolutely are able to suss out certain things, and kind of sense, “Ooh, that’s not good, and I don’t want to pursue this any longer.” I’ve been really fortunate not to experience a ton of things that a lot of people experience, regardless of race, gender, age, there are just some parts of tech that—I don’t want to say allegedly; that can be toxic because I don’t want to invalidate anyone’s experience. But because I’ve lived so much life, and so much of my career was understanding people, that the moment I started to see those signs, I just kind of separated myself from affiliation with that person, or that group, or that entity, and kind of pursued what I knew would work for me.

And then mentorship, and especially mentorship outside of your company. I’ve got great mentors at my company, but I’ve got at least three mentors who all work at different places who had just—I wouldn’t be here without them. They’re my place to go when, hey, is this normal? Because I didn’t have any experience in the tech industry. And I’d run everything by them.

I don’t always do what they tell me to do. Sometimes I get their advice, I listen to it, I think about how it might apply in my life, and then I just tuck it in my back pocket and do what I intended to do in the first place.

Corey: One of the things people get wrong about mentorship is that it has to be mentee-led, not mentor-led. And again, it’s never expected whenever you’re asking someone for advice that you’re going to do exactly what they say, but if you’re going to go to all the trouble of taking someone’s time, you should at least consider what they say. And it may not apply; it may be completely wrong. Every once in a while, we rotate through paid advisors at our company where we have people come in for time to advise us, and sometimes some of those valuable advisors we have, we never did a single thing that they tell us to do, but listening to them and how they articulate and how they clear it out. It’s, “Okay, we strongly agree with aspects of this, but here’s why it is a complete non-starter for us.”

And that is valuable, even though from their perspective, “You never take my advice.” And it’s not that, like, “Well, we think your advice is garbage.” No, it’s well reasoned, and it’s nuanced, but it’s not quite right because of the following reasons. That’s something that I think gets lost on.

Natalie: Yeah, yeah, I would agree with that. And I think you made a really good point. You have to consider the advice if this is someone whom you’ve come to ask how you might handle a certain situation, and they take the time to give their insight, you have to consider that. If you don’t consider it, why are you wasting everyone’s time?

Corey: One last question I want to get into before we call this an episode. It is abundantly clear that you are a net add to virtually any team that you find yourself on based upon a variety of things that you’ve evinced during this episode. Why did you choose to work at Netlify? And let’s be clear, that is not casting shade at Netlify.

Natalie: [laugh].

Corey: Like, “You can work anywhere. Why are you at that crap hole?” No, I have a bunch of friends that Netlify and every story I have heard about that company has been positive. So, great. Why are you there?

Natalie: For me, it’s always going to start with people. I was happy at Foxtrot, my first employer. I was growing there, I was doing well. I liked everyone I worked with. But when Cassidy slides in your DMs and you have a chance to work directly with her and learn from her, you have to explore that opportunity.

So, that’s what at least led me to having the conversation. And then the way I was treated by everyone through the interview process. No one was trying to trip me up, no one was asking me ridiculous questions. And they were actively fighting to make sure that I came in at a pay rate that made sense, and that I was trusted and given responsibility. And I have to say, once I got there, I found out that I had taken the wrong role.

I asked questions about what I was doing. I joined as part of the DX team and my role was to be a template engineer. So, I asked some questions: How much of my role would be coding? Because I knew I couldn’t stray too far from the keyboard at this stage of my career. And I got answers, but I didn’t know the right questions to ask.

When I heard I was—be coding, I thought that meant like how I do now. I work on a product team with a PM and a designer, and they cut issues for me. But what happened in DX is it was much more self-directed, and the work was very different over there. It’s incredibly important work. It’s valuable work, but it didn’t line up with my skill set.

So, having that conversation with Cassidy, and then going on to have that conversation with my VP of engineer, a woman named Dana, and having the safety to have those conversations to say, “Hey, I know I just got here. This isn’t right for me. I owe more to the DX team and I owe more to myself.” And to be well-received, and to immediately begin to have conversations with engineering managers to find out the right place for me, made me incredibly happy that I chose Netlify, and it kind of reinforced the things they were telling me in the interview process were real.

Corey: The fact that you were able to make that transition within the first six months of working at a company and not transition to a different company, either by your choice or not, speaks volumes about how Netlify approaches engineering talent, and its business, and human beings.

Natalie: I agree one hundred percent because they could have very easily told me, “Hey, you were hired to do this role. You didn’t interview for a product team role, you’re welcome to continue to do the work that you were hired to do or move on.” But they didn’t do that. No one—in fact, they encouraged me to find the right place for myself.

Corey: We talked a minute ago about the one of the values of mentors being able to normalize, is this normal or is this not? Let me just say from what I’ve seen for almost 20 years in this industry, that is not normal. That is an outlier in one of the most exceptional ways possible, and it is a great story to hear.

Natalie: I tell you, I’ve had an absolutely termed entrance into tech. But also it goes back to, like, when I was in the interview process, I wasn’t really focusing on, like, what I would be doing as much as who would I be doing it with and getting a feel for both Cassidy and Jason. And I was one hundred percent confident that at the end of the day, what they wanted was to bring me into the company and for me to do work that fulfills me.

Corey: And it sounds like you’ve got there.

Natalie: Absolutely. I’m very happy with the things I’m learning. This codebase is huge. I’m digging in. It’s amazing. I couldn’t ask for more in life right now.

Corey: I want to thank you for being so generous with your time to talk with me today. If people want to learn more, where can they find you?

Natalie: I am on Twitter. My username is @codeFreedomRitr, but that’s spelled C-O-D-E-F-R-E-E-D-O-M-R-I-T-R.

Corey: Excellent. That is some startup to your word spelling there. That is fantastic. You could raise a $20 million seed round on that alone.

Natalie: [laugh]. I mean, can I count that as, like, an endorsement? Can I—

Corey: Oh, absolutely. Yeah. I have strong opinions on the naming of various things. No, well done. Thank you so much for speaking with me today. I really appreciate it.

Natalie: Thank you for having me, Corey. This has been a lovely experience.

Corey: Natalie Davis, software engineer at Netlify. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an angry comment that you are then going to send to corporate and demand your 40 cents back.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Nancy

Nancy Wang is a global product and technical leader at Amazon Web Services, where she leads P&L, product, engineering, and design for its data protection and governance businesses. Prior to Amazon, she led SaaS product development at Rubrik, the fastest-growing enterprise software unicorn and built healthdata.gov for the U.S. Department of Health and Human Services. Passionate about advancing more women into technical roles, Nancy is the founder & CEO of Advancing Women in Tech, a global 501(c)(3) nonprofit with 16,000+ members worldwide.

Nancy is an angel investor in data security and compliance companies, and an LP with several seed- and growth-stage funds such as Operator Collective and IVP. She earned a degree in computer science from the University of Pennsylvania.

Links:

  • https://coursera.org/awit
  • Advancing Women in Technology: https://www.advancingwomenintech.org
  • LinkedIn: https://www.linkedin.com/in/wangnancy/
  • Advancing Women in Technology LinkedIn: https://www.linkedin.com/company/advancingwomenintech/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Sysdig. Sysdig is the solution for securing DevOps. They have a blog post that went up recently about how an insecure AWS Lambda function could be used as a pivot point to get access into your environment. They’ve also gone deep in-depth with a bunch of other approaches to how DevOps and security are inextricably linked. To learn more, visit sysdig.com and tell them I sent you. That’s S-Y-S-D-I-G dot com. My thanks to them for their continued support of this ridiculous nonsense.

Corey: This episode is sponsored in part by our friends at Rising Cloud, which I hadn’t heard of before, but they’re doing something vaguely interesting here. They are using AI, which is usually where my eyes glaze over and I lose attention, but they’re using it to help developers be more efficient by reducing repetitive tasks. So, the idea being that you can run stateless things without having to worry about scaling, placement, et cetera, and the rest. They claim significant cost savings, and they’re able to wind up taking what you’re running as it is, in AWS, with no changes, and run it inside of their data centers that span multiple regions. I’m somewhat skeptical, but their customers seem to really like them, so that’s one of those areas where I really have a hard time being too snarky about it because when you solve a customer’s problem, and they get out there in public and say, “We’re solving a problem,” it’s very hard to snark about that. Multus Medical, Construx.ai, and Stax have seen significant results by using them, and it’s worth exploring. So, if you’re looking for a smarter, faster, cheaper alternative to EC2, Lambda, or batch, consider checking them out. Visit risingcloud.com/benefits. That’s risingcloud.com/benefits, and be sure to tell them that I said you because watching people wince when you mention my name is one of the guilty pleasures of listening to this podcast.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’ve said repeatedly on this show—and I stand by it—that absolutely nobody cares about backups. Because they don’t. They do care tremendously about restores, usually right after they really should have been caring about backups.

My guest today has more informed opinions on these things than I do, just because I’m bad at computers. But Nancy Wang is someone else entirely. She is AWS’s general manager of the AWS Backup service, and heads the Data Protection Team. Nancy, thank you for tolerating me, I appreciate it.

Nancy: Hey, no worries because you know, when I heard you say I don’t care about backups, I knew I had to come on the show and correct you. [laugh].

Corey: It’s the sort of thing where there’s no one is fanatical as a convert. And every grumpy old sysadmin that is in my cohort either cares a lot about backups or just doesn’t even think about it at all. And the question is—the only thing that separates those two groups is have you lost data yet? And once you’ve lost data and you feel like a heel, you realize, “Wow, this was eminently preventable. What can I do differently to fix this?”

And that’s when people start preaching the virtues of backups, and you know, this novel ridiculous idea of testing the backups you’ve made to make sure that it isn’t just—yeah, it says it’s completing correctly, but if you haven’t restored it, you don’t really know.

Nancy: Yeah. I mean, that’s so true, right? And that’s why when we’re thinking about our holistic data protection strategy, it’s less so about, “Hey, make sure that you take backups”—which is albeit a very important part of the data protection hygiene—but is making sure that you can regularly test the things that you’re backing up to make sure that, frankly, when you happen to be in a disaster scenario, or someone fat fingers a restore process, that you have good known bits to restore from.

Corey: So, people will be forgiven for not, potentially, understanding what AWS Backup is, where it starts and where it stops. I mean, let’s be clear, this is sort of the price you as a company get to pay for having 300-some-odd services; not everyone is conversant with every single one of them. I know, I’m as offended as anyone at that fact, but apparently other people have lives. So, what is AWS Backup?

Nancy: So, on that note, Corey, I do have to say that I’m probably at a more of an advantage in terms of my name being very descriptive and what it does versus, maybe, Athena or Redshift where it’s very clear, hey, we do backups. But actually, if you parse apart the product—and this is why the team itself is called data protection—there are various axes to think about what we do, right? So, to help illustrate, perhaps if you think about axes one as in, what are the different types of application data that we protect, right? There’s obviously database data, there’s going to be file system data, there’s various storage platform data, right? And those are comprised by AWS services that I’m sure you all are very familiar with, love dearly, like RDS, EBS, with EC2, VMs, et cetera, but also, more recently, we added S3, which we’ll get to that in just a bit, but because I’d love to talk about, you know, how folks think about S3 and why you might want to back it up, right? So, that’s axis number
one.

Now, if we turn to axis number two, it’s about the different platforms where these application data might reside. So there’s, of course, in-cloud, and that’s the place where most people are familiar with and why they might choose to seek out a first party native data protection provider like AWS Backup. And by the way, we just extended our support to on-premises as well, starting with VMware, which is a thing that a lot of backup admins were super excited to hear about, and all those vExperts out there.

And of course, the final axis is we think about how we make sure that we not just protect your data, but we are also able to give you tools like compliance reporting, which we announced in August at re:Inforce, via our CISO, Stephen Schmidt, about, “Hey, once you take your backups, are you monitoring continuously the resource configurations of the application data that you’re protecting?” Are your backup plans architected to meet RPO requirements that your organization needs to meet? Are they being, for example, retained for the right amount of times? Is it seven years or is it a month? Many different organizations have widely varying RPO requirements, so making sure that all of that is captured, monitored, and also reportable so when, hey, those, that auditor decides to knock on your door, you have a report ready to say, “Hey, I’m in compliance. And by the way, I’m proactively thinking about how my organization can meet evolving regulations.”

Corey: Please tell me you’re familiar with AWS Audit Manager, which is, to my understanding, aimed at solving exactly this problem. If the answer is no, this would admittedly not be the first time there I found, “Oh, wow. We have a complete service duplicate hanging out somewhere at AWS.” “Oh, good. How do we make it run in containers?” Being the next obvious question there.

Nancy: Sure. Which is actually a great lead-in to, again, another descriptive name of an AWS service, which is AWS Backup Audit Manager. So, if you recall from the re:Inforce keynote, it was one of the slides that was highlighted. The reason being, I’m a firm believer of a managed solution. Because look, we all know that AWS is great at building, I would say, tools or building blocks, or primitives to design end-to-end solutions.

Corey: It’s the Lego approach to cloud services. “What can I build with this?” “You’re only constrained by your imagination.” “Okay, but what can I build?” “Here to talk about that is someone from Netflix.”

Great. I want to build Twitter for Pets, which I guess now has to stream video? Yeah, it becomes a very different story. The higher-level service offerings are generally not a common area that AWS has excelled in, but this seems to be a notable exception.

Nancy: That’s actually where my background is, right? So, previous to AWS, I worked at a not-so-small startup anymore, called Rubrik, down in Silicon Valley, where we spent a lot of time thinking about what is the end-to-end solution for customers. How can customers simply deploy with one click, make sure that they can create policies that are repeatable, that are automated, and go off when you want them to, and make sure that you have reporting, at the end of the day. So, that’s really what we focused on, right?

But I digress, Corey. To your question about AWS Audit Manager, the name of the service within AWS Backup that handles compliance reporting, and auditing is called AWS Audit Manager, and we certainly didn’t pick that name by fluke. The reason being, we wanted AWS Backup, from that managed solution point of view, to be the single central platform where customers come to create data protection policies, where they come to execute those data protection policies, in backup plans, store their backups in encrypted backup vaults, and have the ability to restore them when they want, and finally, report on them. So, it is that single platform.

Now, with that said, if, for example, you wanted that reporting to come from AWS Audit Manager, which is a service that does a lot of reporting across many AWS services, you also have that ability. So, depending on what user persona you might be, whether you’re from the central compliance office or you’re a member of the data protection team within an organization, you might choose to use that functionality separately. And that’s the flexibility that my team strived to provide.

Corey: One of the most interesting things about AWS Backup is that I did not affirmatively go out of my way to use your service. I did not—to my recollection—wind up saying, “Oh, time to learn about this new thing, and set it up, and be very diligent about it.” But sure enough, I find it showing up on the AWS inventory—which is of course, the bill. And I look at this in a random account I use for various, you know, shitposting extravaganzas, and sure enough, it’s last—so far, this month, it is—I’m recording this near the end of the month—it charged me $3.40 to backup 70 gigs of data.

Which is first, like on the one hand, there is an argument of, “Now, wait a minute. I didn’t opt into this. What gives?” The other side of it though, is how dare you make sure that my data isn’t going to be lost, not through your negligence, but through my own, when I get sloppy with an rm -rf. And because I’ve been using ZFS a fair bit, and it is integrated extraordinarily tightly with that service. It goes super well.

It works out when setting this up, unless you go out of your way to disable it, it will set up a backup plan. And first, that is not generally aligned with how AWS thinks about things, which you across the board, generally the philosophy I’ve gotten is, “Oh, you want to do this thing? That’s a different service team. Do it yourself.” But also, it’s one of those areas that is the least controversial. If you have to make a decision one way or
another, yeah, it’s opt people into backups. Was that as hard to get approved as I would suspect it would be, or was that sort of a no-brainer?

Nancy: Hopefully you can let me know what your account number is, Corey, so I can make sure it doesn’t get marked for fraud—A—but B, going into, you know, our philosophy on protecting data: So, EFS actually was one of our first AWS services that was supported by the AWS Backup service, which is actually quite a fascinating story in itself because the service [AWS Backup] only launched in 2019. Now, AWS has been around for much, much longer than that—

Corey: And it feels even three times longer than that. But yes.

Nancy: [laugh]. Exactly, right. So, as a central data protection platform for the AWS overall cloud platform, it’s quite interesting that from a managed solution perspective, the service is not yet, you know, four years old. We’re barely embarking on our third year together. So, with that said, why we started with EFS and a few other services is we wanted to cover the most commonly used stateful data stores for AWS Cloud, EFS being one of them, as the first cloud-native—as Wayne Duso would say—Elastic File System in the cloud.

And so what we did is a deeper level integration, what we call our “data plane integration.” So, what does that mean? Customers protecting EFS file systems have the ability to not just restore their entire file system as a file system volume, but also have the ability to specify individual files, folders, that they want to restore from. And so, file level recovery, super, super important. And it’s something that we also want to bring for other file systems down the road as well.

And so, to your question, Corey, a common design principle that we think about is, how do we make sure that customers are protected? Obviously, in a world where we cannot yet use AI to transcribe every part of a customer’s intent when they’re looking to protect their data, the closest that we can get is, “Hey, you create a file system. We assume that you want it protected, unless you tell us you don’t want to.” And so for certain resources, like EFS, where we have a deeper level integration to our own data plane, we can then say, “Once you create a file system will opt you automatically into AWS Backup protection until you tell us to stop.” And from there, you have all the goodness that comes with AWS Backup, such as file-level restore, such as for example now, WORM [write-once-read-many] lock, which disables the ability to mutate backups from anyone, even someone with admin access.

Corey: So, a big announcement in your area at re:Invent, was AWS Backup support for S3. Allow me to set up an intentionally insulting straw man argument here. S3 has vaunted 11 nines of durability, which I think exceeds the likelihood the gravity is going to continue to function. So, are they lying by having AWS Backups supporting it now, or are you just basically selling us something we don’t need? Which is it?

Nancy: Well, you know, Corey, judging by the hundreds of customers who have been filling up my inbox—and that’s why I actually ended up creating a special email alias for the S3 preview—so what we launched at re:Invent was a public preview of the ability to start baking in S3 backup protection—or bucket protection—into their existing data protection workflows, right? And so judging by the hundreds of customers, many of them in highly regulated industries, and FinServ, in healthcare, as well as in the US government, I would say that I think they find it pretty important, and we’re not just peddling things they don’t need. So, I’m getting ahead of myself. We’re actually—we should probably start the conversation—is a deeper dive into how we think about data protection on AWS.

And so there’s two really core schools of thought, right? One is, you know, focused on data durability, which in itself is a function of technology. So, to your point of 11 nines, right? That is very much true, and that’s why S3 increasingly becomes the platform of choice, now, for all of customer’s, you know, analytics information, and other stateful stores that they want to keep an S3 buckets for applications, right? But second of all—and this is a part where AWS Backup wants to focus on—is that concept of data resiliency, which itself is a function of external factors. Because, for example, human errors, such as fat-fingering, or miscellaneous entries, could impact for example, how you can access information that’s stored in your S3 bucket, or unfortunately, sometimes what we’ve heard is accidentally deleting an S3 bucket or certain objects in your S3 bucket.

Corey: This speaks to the idea of that RAID is not a backup. Sure, you want to make sure a drive failure doesn’t lose your data, but you also want to make sure that you overwriting a file that was super important doesn’t happen either and RAID, nor data durability and S3, are going to save you from that.

Nancy: Yeah. Because for example, we have built in—and this is actually very core to not just AWS Backup, but really how we think about data protection on AWS—is again, that separation of control. So, I encourage you to try to delete, let’s say, an EBS volume that is protected by AWS Backup, from the EBS console. You’ll likely find a very glaring error in your face that says, “You do not have sufficient privileges to do so.” And the reason we actually make such a separation of control, or our role-based access control—RBAC—so core to our product design is so that, for example, whoever creates that primary volume should not be the same person that deletes it, unless they do happen to be the same person with two different roles.

And that prevents, for example, unintended mutations. That also enables the data protection administrator to have the ability to, let’s say, do cross-region copies: Having your S3 bucket or objects stored in another region, in another account, that can be completely locked down to anyone, even those with administrator access, right? So, like I said, before, all the platform goodness, AWS Backup, such as version control, WORM locks, having multiple copies of those backups, as well as different protection domains, that’s what customers look for when they come to this service.

And to your point, especially even with highly durable platforms like S3, there’s still external factors that you simply can’t control for all the time, right? And having that peace of mind, having that protection that you know is on 24/7, hey, that keeps businesses up, right? And that keeps consumers like you and me able to enjoy all the goodness that those businesses offer.

Corey: This episode is sponsored by our friends at Oracle HeatWave is a new high-performance query accelerator for the Oracle MySQL Database Service, although I insist on calling it “my squirrel.” While MySQL has long been the worlds most popular open source database, shifting from transacting to analytics required way too much overhead and, ya know, work. With HeatWave you can run your OLAP and OLTP—don’t ask me to pronounce those acronyms again—workloads directly from your MySQL database and eliminate the time-consuming data movement and integration work, while also performing 1100X faster than Amazon Aurora and 2.5X faster than Amazon Redshift, at a third of the cost. My thanks again to Oracle Cloud for sponsoring this ridiculous nonsense.

Corey: I agree wholeheartedly with everything that you’re saying. I had a consulting client where it’s coming in optimize the AWS bill, and, “Wow, that sure is a lot of petabytes over in that S3 infrequent access bucket. How about you change the Infrequent Access-One Zone?” “Oh, no, no, no. We lose this data, it basically ends a division of the company.” “Cool. Do you have multi-factor delete turned on?” “No.” “Do you have versioning turned on?” “No.” “Okay. This is why I call it cost optimization, not cost cutting. You should be backing that up somewhere because there is far likelier—by several orders of magnitude—that you or someone on your team intentionally—unlikely—or by accident—very likely, as someone who’s extremely accident prone with computers, from my own perspective because I am—is going to accidentally cause data loss there. So yeah, spend more money and back that up.”

And they started doing that. So, it’s always nice when your recommendations get accepted. But yeah, if data is that important, you absolutely need to have a strategy around that. What I love so far about what I’ve seen from AWS Backup is—and please don’t take this in any way as criticism on it—is that it’s so brainless. It just works. Because people don’t think about backups until it’s too late to have thought about backups.

Nancy: Yeah, don’t worry, I don’t take that as offense, Corey, otherwise I wouldn’t be on the show. Absolutely, right? My motto is set it and forget it, right? Just as I want to make it super simple for our mission, for customers to understand our mission, as well as, frankly, the engineers who build the service to understand our mission, it is, “We protect our customers’ data on AWS. How? With set-it-and-forget-it data protection policies.”

And we try to configure these policies to be fairly comprehensive. You can set everything from, like I mentioned, warm lock, where you want your backup copies created to: Which regions? Which accounts, for example? Which user role do you want to use with these data protection policies? Which services do you want to protect?

And even recently, we created the selection ability—or as we call it, AWS Backup Select—so you can include, exclude different resources, even when you have the common union of tags specified on your backup plan. So, the reason we went this comprehensive is so that once you configure a data protection policy, you can really rest assured that, hey, I’ve done everything in my power to make sure that these resources, this application data that is so critical to my business, is being protected. And oh, by the way, I can see these backups—or as we call in our lexicon, Recovery Points—directly in my console, in my account.

Corey: And there’s tremendous value to doing that. That is the sort of thing that customers like to see. This is—if you have to move up the stack somewhere, this feels like the place to begin doing it, just because it’s so critical to the rest of it. We all have side projects as well. Like, for example, I wind up making insulting parody music videos for people’s birthdays when they’re not expecting it. You have 80 hours of training content on Coursera. What is that about? Because I don’t think it’s all about backups.

Nancy: No. Although at some point, we should probably get AWS Backup as one of the modules in AWS certification. But I digress. The reason why training is so important to me is one of the ways, actually, that folks find me online is through my presence in the nonprofit world. So, I’m
the founder and CEO of a 501(c)(3) organization that’s called Advancing Women in Technology, or AWIT, or A-W-I-T for short.

The mission of AWIT is really to get more women leaders into visible, into senior tech leadership roles, so frankly and from a selfish perspective, I’m not the only woman in a room many of the times when decisions are being made, right? And that’s not just, you know, I’m talking about my current role, but in various roles that I’ve had throughout the tech industry. So, where does that start? And there’s a lot of different amazing organizations that focus on the early career, beginning in the pipeline, which is super important because it is important to get women, underrepresented groups in the door so that they can advance and they can accelerate their careers to becoming leaders, but the areas where AWIT focus is actually in that mid-career.

Because once folks, and especially women and underrepresented groups are in the door 10 to 15 years, they’re maybe in their first managerial role, or they’re in their first leadership role, that’s the core time when you want to retain that population, where you want to advance that population, so that in the next, I would say, generation—or hopefully it doesn’t even take that long; next 5, 10 years—we see a much more representative leadership room, or board table, right? So, that’s really where that goal starts. And so, why do we have 80 hours of training content because part of advancing your career and accelerating your career is having the right skills. Of course having a right network is also very important, and that’s something else that we preach, but upskilling yourself, constantly learning about new technologies—I mean, the tech world changes by the minute, right, and so being familiar with new technologies, new frameworks, new ways of thinking about product problems, is really what we focus on. So, we were the first to create the Real-World Product Management Specialization, which you can check out on Coursera. You’ll see my mug shot in a lot of those videos.

But actually, also of those of some of the best and brightest underrepresented leaders in the industry, such as Sandy Carter, Mai-Lan Tomsen Bukovec, Sabrina Farmer, I mean, the list goes on and on. Including, you know, personal friend who created Coffee Meets Bagel. So hey, for all those connections made out there on that platform, you know, she’s also a woman CEO, and used to be a product manager at Amazon.

Corey: A dear friend met his partner on Coffee Meets Bagel. I hear good things.

Nancy: Oh, awesome.

Corey: Fortunately, I was married before it launched, so I’ve never used the service myself. If I were a reference customer now, that would raise questions.

Nancy: [laugh]. Well, let’s just say I’m not on the platform, either, so I can’t verify or deny that you have a profile. Yeah. So, just having those underrepresented groups and individuals, really stellar rock stars, role models that we would all consider to be super inspirational, as speakers, as instructors on the courses have given so many folks the inspiration, the encouragement that they need to upskill themselves. And so yes, now educated over 20,000 learners worldwide using those courses.

And I still receive just amazing notes from them on a daily basis, all over LinkedIn about how they’ve managed to get promotions from taking these courses, or how they’ve managed to get jobs in FAANG tech companies as a result of taking these courses. And really, that’s the impact that I want to make is one to n, being able to impact a global audience, upskilling a global audience. And so again, in the future, and not so distant future, the leadership room gets so much more representative.

Corey: And to complete the trifecta of interesting things you do, you are also an early angel investor and a limited partner in a number of startups. Tell me a little bit about that. It’s odd to—at least in my experience—to see folks who are heavily involved in the nonprofit space, the corporate space at a giant tech company, and doing investment all at the same time. It seems like that is not a particularly common combination, at least in the circles in which I travel.

Nancy: You could also probably blame it on my extreme ADHD. That’s probably very true. Don’t worry, I try to control it, most of the time.

Corey: I’ve been struggling to control my own my entire life, which probably explains a lot about why I do the things that I do. I hear you.

Nancy: It makes sense, right? From one to another. It honestly makes me better at my job. And I’ll explain why. So, if you look at some of the new or joint marketing campaigns that AWS Backup or data protection team has done this past year with various startups—namely Open Raven; there’ll be others we’re working with in the new year—being able to just get some of that inspiration from founders, so thinking about how can we have a better together story?

You specialize in, let’s say with the case of Open Raven, in data visibility and let’s say scanning S3 buckets for vulnerabilities, for different content. And hey, we specialize in data recovery process, or then that data protection policy creation process. How do we come together to form a really awesome solution for our highly regulated customers, or compliance-minded customers? That’s the story that I love to tell, and frankly, I just get so inspired from talking to startup founders. The reason why I have also advised a few venture capitalists—namely Felicis Ventures—on, for example, their investment thesis is I just see so much potential in this environment, right?

And there’s really that adage, where it’s big enough sandbox for a lot of players. Just like, for example, how Snowflake and Redshift have managed to coexist together on the AWS platform, there’s a lot of just goodness, too, that exists between the data security world, how they customers think about securing their data, to the data protection world because, hey, you can’t protect what you can’t see, so you need to be make sure that you have that data visibility angle, along with that protection angle, along with that recovery angle. And hey, all of this needs to be within your data perimeter, within a secure zone, right? How do you securitize your data? So, all of that really comes together in this melding world.

And of course, there’s also adjacent themes such as, well, once you protect your data, how can you also make sure that the quality of your data is high? And that’s where pretty interesting startups in the data observability space, such as Monte Carlo, have come up. Which is, “Hey, I need to rely on my business data to make important decisions that affect my customers, so how can I make sure that what’s ever coming out of my data lake or data warehouse is correct, it truly reflects the state of the business?” So, all of that is converging, and that’s why, you know, it’s just super exciting to be a part of this space, to not only create net new, I would say greenfield opportunities on the AWS platform, but also use this as an opportunity to partner with startup CEOs and various startups in the data space, data infrastructure space, to create more use cases, more solutions for customers who otherwise we’d have to rely on either custom scripts, or simply not having any solutions in this space at all.

Corey: There’s something to be said for doing the—how do I frame this?—the boring work that’s always behind the scenes, that is never top of mind. People don’t get excited about things like data protection, about compliance, about cost optimization, about making sure that the fire insurance is paid up on the building before you wind up insulting execs at big companies, et cetera, et cetera. And that—but it is incredibly important—in my case, especially that last one—just because if you don’t get that done, there’s massive risk, and managing that risk is important. It’s nice to see that it’s not just the shiny features that are getting the attention. It’s the stuff of, “Okay, how do we do this safely and securely?” That is the area that I think is not being particularly well served these days, so it’s honestly refreshing to see someone focusing on
that as an area of active investment.

Nancy: I mean, absolutely. Perhaps one data point I should also share, because I do get questions asked of, “What gets you so excited about compliance, about audit?” Well, I used to work for the US government. So, if that tells you anything—and I used to hold an active secret clearance—that hopefully explains some things about why I’m passionate about the areas I am. But, that’s really where, you know, back to your comment that you made on the core tenet or the ethos of the AWS Backup service, which is, “Set it, forget it, make it super simple,” is I want to design systems or solutions that enable customers to focus on developing applications, working on building business logic, whereas we will create the comprehensive data protection policies that protect your data.

And especially in the world of ever evolving cyber attacks where the attackers are getting more and more sophisticated, they have more backdoor methods that go undetected for many months, as was the case in attacks over the past recent years, or in the case of pesky ransomware attacks, where certain insurance companies have even stopped paying ransoms, right, and you’re wondering, “Well, how do I get my data back?” This is the world that we live in. And so, you know, yes, there might be ever-evolving more, I would say, sophisticated ways to detect vulnerabilities, or attacks, or do pattern matching between known attack patterns, but really what remains core and should be core to a lot of companies’ recovery strategies, as per the NIST cybersecurity framework, is actually having a good way to restore. And that goes back to something that you mentioned at the beginning of this recording, Corey, which is making sure that you’re regularly testing your backups because as you said, no one cares that you’re taking backups, but people do care about the ability to restore. So, having known good bits that exist in a secure vault, that exists maybe in some air gap account or region, where you know that it’s going to be there for you, that it’s restorable is going to be super key.

And we’re already seeing that trend in a lot of customers that I speak with. And by the way, these aren’t just customers in highly regulated industries. They’re really customers that now are increasingly relying on data to make business decisions. Just like, for example, there’s that adage that says, you know, “Software is eating the world,” well, now most businesses are data-driven businesses, and so data is core to their business mission. And so protecting that, it should also be core to their business mission.

Corey: I really wish that were the case a bit more than it is.

Nancy: True that. So, I would have to say, “Hear, hear.” And this is actually what makes my job so, just, fun frankly, is that I get to have these conversations with thought leaders at various different companies, who are my clients or customers of AWS. And these are different, I would say, leaders, ranging from IT leaders, to compliance leaders, to CISOs who I have these conversations with. And oftentimes it does start with this very, I would say, innocuous question, which is, “Well, why should I think about protecting my data?” And then we’re able to go into, “Well, this is how you think about tiering your data, this is how you think about different SLAs that you might have for your data, and then finally, this is how you would think about architecting a data protection solution into your environment.”

Corey: Nancy, I want to thank you for taking some time out of your day to speak with me. If people want to learn more about what you’re up to and how you’re viewing these things, where can they find you?

Nancy: Feel free to connect with me on LinkedIn, whether you have a service that you desperately want AWS Backup to protect—yes, I get a lot of those tweets or LinkedIn posts—absolutely happy to consider them and to prioritize them on the future roadmap. Or if you want to give me a feedback about your experience, more than happy to take those as well. Also, if you’re a startup founder and you have a brilliant new idea, and data infrastructure, always happy to grab coffee or drinks and hear about those ideas.

And lastly, if you’re looking to upskill yourself either product management or cloud tech skills, find us on Coursera at https://www.coursera.org/awit, or on LinkedIn as Advancing Women in Technology. Either way, whether you fit into one or more or all of these buckets, I’d love to hear from you.

Corey: And we will, of course, put links to that in the [show notes 00:32:36]. Thank you so much for speaking with me today. I really
appreciate it.

Nancy: Well, thank you, Corey. It’s always a pleasure, and I’ll see you very soon in person in SF.

Corey: I look forward to it. Nancy Wang, General Manager of AWS Backup and AWS Data Protection. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an insulting comment that I will then delete because it wasn’t backed up.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Rachel

Rachel Kelly is a Senior Engineer at Fastly in Infrastructure, and is a proud career-switcher over to tech as of about eight years ago. She lives in the Pacific Northwest and spends her time thinking about crafts, cycling, leadership, and ditching Google. Previously, she worked at Bright.md wrestling Ansible and Terraform into shape, and before then, a couple years at Puppet. You can reach Rachel on twitter @wholemilk, or at hello@rkode.com.

Links:

  • Fastly: https://www.fastly.com
  • SeaGL: https://seagl.org
  • Twitter: https://twitter.com/wholemilk

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by LaunchDarkly. Take a look at what it takes to get your code into production. I’m going to just guess that it’s awful because it’s always awful. No one loves their deployment process. What if launching new features didn’t require you to do a full-on code and possibly infrastructure deploy? What if you could test on a small subset of users and then roll it back immediately if results aren’t what you expect? LaunchDarkly does exactly this. To learn more, visit launchdarkly.com and tell them Corey sent you, and watch for the wince.

Corey: It seems like there is a new security breach every day. Are you confident that an old SSH key or a shared admin account isn’t going to come back and bite you? If not, check out Teleport. Teleport is the easiest, most secure way to access all of your infrastructure. The open source Teleport Access Plane consolidates everything you need for secure access to your Linux and Windows servers—and I assure you there is no third option there. Kubernetes clusters, databases, and internal applications like AWS Management Console, Yankins, GitLab, Grafana, Jupyter Notebooks, and more. Teleport’s unique approach is not only more secure, it also improves developer productivity. To learn more visit: goteleport.com. And no, that is not me telling you to go away, it is: goteleport.com.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. A periodic subject that comes up from folks desperate to sell people things is this idea of cloud repatriation, where people have put their entire business in the cloud decided, “Mmm, not so much. I’ll build some data centers and move it there.” It’s an inspiring story if you’re selling things for data centers, but it’s not something we’re seeing widespread evidence of, and I maintain that.

Today, we’re going to talk about that, only completely different. My guest today is Rachel Kelly, senior infrastructure engineer at Fastly. And no, Fastly has not done a cloud repatriation of which I am aware. But Rachel, you’ve done a career repatriation. You went from working with
AWS in your previous company to working in bare metal. First, welcome to the show, and thank you for joining me.

Rachel: Thanks, Corey. Super happy to be here.

Corey: Now, let’s talk about why you would do such a thing. It feels almost like you’re Benjamin Button-ing here.

Rachel: Yeah, a bit. The normal flow has been to go from sort of a sysadmin level, where you’re managing servers fairly directly, to an operational level, where you are managing entire swathes of servers to entire data centers and so forth. But I went from managing just the SaaS web app to managing enormous groups of servers in data centers all over the world. And I did that because the provisioning of the web app, even on AWS, was absolutely my favorite part. What I’ve always wanted to get better with is the Linux and networking side of how our internet runs, and at Fastly, we are responsible for such a huge percentage of traffic all over the world. We have enormous customers who rely on us to deliver that data. And I get to be part of the group of people that puts those enormous groups of servers into production.

Corey: I started my career in the more traditional way of starting out in data centers, building things out, and then finally scampering off into a world of cloud. And you learn things going through the data center side of the world that don’t necessarily command the same levels of attention in the cloud environment because you don’t have to think about these things. Networking is a great example. During the Great Recession, there was a salary freeze. I was not super thrilled in my job, but I couldn’t find another one, so I spent the year learning how networks worked, and it made me a better systems administrator as a direct result of this. Same story with file systems, not necessarily because I did extensive amounts of work with their innards, but because every sysadmin interview under the sun asked the same questions about how inodes work, how journaling works, et cetera, and you have to be able to pass the trivia-based hazing process in order to get a job
when you’ve just been fired from your last one.

So, that became where I was focusing on these things. And now looking at a world of cloud, feels like we don’t really need that in any meaningful sense. I mean, a couple people need to know it, but by and large no one has to think about it. So, is that just a bunch of useless knowledge that is taking up valuable space in your brain that could be used for other stuff or do you think that there’s a valid story for folks who are working in purely cloud environments to still learn how the things underlying these concepts work?

Rachel: First of all, I think that there is so much that we can do with less particular networking knowledge than we’ve ever needed in the past, thanks so much in part to AWS and all of their hangers-on. But yes, there are still people who need this networking knowledge. And once you have that kind of knowledge, once you’re able to see how the routes talk to each other, and how your firewalls actually work, and how to abstract out these larger networks and determining your subnetting and everything, you can utilize that really beautifully, even in something like VPC on AWS. Without that kind of knowledge, like, you can still get quite a bit done—which I think is a testament to the power of abstraction in AWS—but I mean, boy oh boy, what you can do once you have some of that knowledge.

Corey: I’m not allowed in the AWS data centers because I’m very bad at dodging bullets, but I find the knowledge is still useful because it helps me reason about things. When I know what—at least in a traditional environment—it’s doing, I know what AWS is emulating, and I can safely assume that I haven’t discovered some bug in their network stack for almost anything reasonable that I’d be working on other than maybe their documentation explaining it. So, when I start reasoning about it from that perspective, things make a lot more sense. And that’s always been helpful. The argument historically has been when you’re hiring—at least in the earlier days of cloud—well, I’m trying to hire, but it’s hard to find cloud talent, so the story was always, “Oh, don’t worry. If you’ve worked in a data center, we’ll teach you the cloudy pieces because it’s the natural evolution of things.” And there’s a whole cottage industry of people training for exactly that use case. Because you are who you are, and doing what you do, how do you find hiring works when you’re going the exact opposite direction?

Rachel: Oh, my gosh, it’s so interesting. In my area, we are trying to build these huge groups of servers based on bare metal. Do we hire sysadmins? Maybe. Do we hire ops folks? Maybe. Do we hire network engineers? Also, maybe.

There are so many angles that we need to be aware of when pulling new talent into our area. And I think it’s fascinating what all of these different, largely, like, non-programmer types have to contribute to the provisioning process. We need someone with expertise in security, and quality, and networking, and file systems, and everything else between those items. And it’s really exciting seeing what people can add to our process.

Corey: There’s so much in there that I love, but at the part I’m going to focus on is you’re talking about new hires as being additive. And that is valuable. It can lead to some pretty toxic and shitty behaviors, where it’s, “We want to make sure everyone we hire is schmucks we’ve hired now.” Like, no, that is not what we’re talking about. But culture is something you get whether you want it or not, and I firmly believe teams are atomic, when you bring someone new in or let someone go, you haven’t changed the team, you have a new team, in many respects, and that dynamic becomes incredibly important.

The idea of hiring people for strength has always been what I look for, as opposed to absence of weakness, where it’s okay, I’m going to ask you a whole bunch of questions around all the different aspects of computing; I’m going to find the area you’re bad at, and we just beat the snot out of you on that. It’s, yeah, if I want to join a fraternity, I would.

Rachel: [laugh]. Yeah, when I was job seeking, I wound up in interviews at places where their method of interviewing was very much hazing. “Well, let’s see, I haven’t read your resume. It says that you’ve set up a few things with Nginx. Do you know about this particular command in Nginx?” It’s like, “Well, geez, I could look it up and figure it out, but that’s not the point of this job.”

I mean, we work together collaboratively every day, and if that doesn’t sound familiar to you, I’m going to leave this interview. But yes, I mean, everybody’s additive. There was another gal who joined at the same time that I did at Fastly, and we both have a very operational background. And we were additive to the very strong networking and data center engineers who were already on the team. And as far as I can tell, the team changed overnight when we joined.

It is now our role—both this other gal’s and mine—to work so much on the automation piece of our build process, which has been focused on lightly in some areas, but that we can bring that with—even just shell scripting, we are able to enhance that process by so much. And I just fantasize about the day that we can get someone in who is directly on our team and focused on security, or directly on our team and focused on testing. The heights we could soar to with that kind of in-department knowledge, where we’re still focused on creating these builds, it’s just so exciting to think about.

Corey: It is and it’s easy to look at data centers as the way things used to be but not the future at all, but CDNs are increasingly becoming something very different than they used to be. And I admit I’m a little stodgy; I tend to fight the tide. There’s value in having something that is serving static assets close to your customer. There’s value to the CDN, in following the telco story, of aspiring to be more than just the quote-unquote, “Dumb pipe,” because that’s a commodity; you want to add differentiated value. But I’m also leery to wind up putting things that look like business logic into the edge at this stage.

And I’m starting to feel like I might be wrong as far as the way that the world views these things. But I like the idea that if a CDN takes an outage—which is not common, but it does happen—that I should be able to seamlessly—well, “Seamlessly”—failover to a different CDN within an hour or so. But if there’s significant business logic in your CDN, you’ve got to either have that replicated in near real-time between the two providers, or your migration is now measured with a calendar instead of a stopwatch.

Rachel: Yeah, absolutely. I mean, that’s an incredibly hard problem. We want to be able to really provide that uptime. And we don’t really have outages. Everybody remembers—well, listeners of this show will probably remember, the Fastly outage, but—

Corey: The Fastly outage, and that’s the—

Rachel: The—

Corey: —best part is the fact that I’m talking about ‘the’ and everyone knows the one I’m talking about, that says something.

Rachel: Yeah. In June of this year, we had an outage for 45 minutes, and it was just an incredible and beautiful effort on the engineering side to get us back up as quick as possible. There were a handful of naysayers, certainly, in the outage, but we fixed it real fast. One thing that I loved was your tweet about it in June, when our outage happened. “The fact that Fastly was able to detect, identify, and remediate this clearly complex problem as quickly as they did may be one of the most technically impressive things I’ve seen in years.” I appreciated that so much. So, many folks internal to Fastly appreciated that point of view so much because the answer to should I have a backup CDN? Like, yeah, maybe, and it is complicated because you have so much logic on the edge right there, but really, the answer is, we really do a good job of staying up. And that cannot be the full picture for any company that needs just a ton of HA, but that is what we’d really like to present, we really want you to be able to trust us. And I feel like we have demonstrated that.

Corey: I would argue from where I sit you absolutely have. If this were a three times a week situation, it wouldn’t matter, no one would care because no one’s going to trust the CDN that breaks like that.

Rachel: Right.

Corey: It gets to the idea of utility computing. And that means different things to different people, but to me, what that says is that when I use an actual utility, like water or electricity, when I turn the faucet or flip a switch, I don’t wonder if it’s going to work or not. Of course, now I have IoT light switches, so I absolutely wonder if it’s going to work or not, but going to the water story, yeah, I turn on the faucet, if something doesn’t happen, or the water comes out a different color than expecting, I have immediate concerns. And that is extraordinarily atypical and I can talk about that one time it happened. It’s not that every third time I go and wash my hands, the water catches fire because there’s fracking nearby, or something. Or it’s poisonous because I live in Flint. It is just a thing that works.

No one is going to sit here and have a business problem and say, “You know what I really need? I really need a local point of presence close to my users so that the static asset can be served more quickly and efficiently to this.” No, the business problem is, “Our website is slow, so people aren’t using it.” It’s how do you speak to things like that? And how do you make working with it either programmatically or through a console—because surprise, business users generally don’t interact with things via APIs—how do you make that straightforward? How do you make that accessible, and Fastly does—

Rachel: Oh gosh.

Corey: —a bang-up job on this.

Rachel: I think that Fastly has done a good job on it. How that has happened, I simply cannot tell you whatsoever. I am so far from support and marketing. I know that those folks work their tails off and really are focused on selling the story of you need your assets to be more easily delivered to the people who want to consume it. No, and you would never use that as a soundbite for Fastly because it [laugh] it sounds like a robot said it.

Corey: It’s always—I was gonna interesting, but I’m also going to go with strange—the ability to, for whatever reason, build out a large scaling infrastructure business like this—CDNs are one of those businesses where you’re not going to come up with this in your garage and a cloud provider tonight and be ready to deploy in a couple of weeks. It takes time to get these facilities out there. It takes tremendous capital investment. But I want to switch a little bit because I know that you’re a believer in this in the same way that I am. As much fun as it is to talk smack about cloud providers, I think it’s impossible to effectively understate just how transformative the idea of being able to prototype things via a cloud provider is.

Yeah, it’s not going to be all businesses, I’m not going to build a manufacturing company on a cloud provider overnight in my spare time, but I can build the bones of a SaaS app and see if it works or not without having to buy infrastructure or entering into long-term contracts. I just need a credit card and then I’ll use a free tier that’s going to lie to me and then hit me with a surprise $60,000 bill. But yeah, you know, the thought is there.

Rachel: The thought is there. I think that if you know a little bit what you’re doing with a not even terribly clever operations engineer to get into AWS with you, you can prototype that for pretty cheaply. If you’re not spending all this money on transfer fees and whatever else. If you really just want this small mock up of hey, does this work? Can it be reached from the network? Again, getting your networking knowledge in will only serve you, even in this setting, even though we’re in the modern era.

I mean, I think it’s incredible, and I think it’s responsible for the total democratization of the modern internet as we know it. Yes, there are other cloud providers, but AWS is who brought this to everybody. Their support for when you run into a jam is some of the most technical and capable of any support organization I’ve ever interfaced with. And at my previous role we did all the time because, you know, the internet gets complicated, if you can imagine that. And I just think that’s phenomenal.

On AWS, I want something where I’m hooking up some VPC to this Redis Database over here to a few EC2 instances with backups going over here, and some extremely restricted amount of dummy data flowing from all of those objects. And there’s nothing like that. [laugh].

Corey: Oh, yeah. And part of the reason behind this, as it turns out, is architectural. The billing system aspires to an eight-hour consistency model, in which case, I spin up something and it shows up in the bill eight hours later. In practice, this can take multiple days. But it’s never going to get fixed until the business decides, all right, you can set up a free tier account with the following limits on it, and to get past these, you have to affirmatively upgrade your account so we can start charging you and we automatically going turn things off or let you stop adding storage to it or whatnot, whenever you cross these limits.

Well today, you can do whatever you want for the first eight hours. And the way to fix this is, cool, Amazon eats it. Whenever their billing system doesn’t catch something, they eat the free tier. And given how much they love money, and trimming margins, and the rest, suddenly you have an incentive because if someone screws up royally and gets that $60,000 bill before the billing system can clamp down on it, okay, great. I would rather the $1.6 trillion company eat that bill than the poor schmoo sitting in their dorm room halfway around the world.

Rachel: That’s such a good point. Some schmo in their dorm room. How many kids have been bitten by this that we don’t hear about because people become ashamed of “Stupid mistakes” like that—that was big air quotes, for those of you at home. It’s not a stupid mistake.

Corey: People think I’m kidding when I say this, but Robinhood had a tragic story, right? A 19-year-old was day-trading, saw on the app that he had lost $900,000—which turned out not to be true once things settled—and killed himself. And that is tragic. It is not a question of if, it’s a question of when someone sees this, reads that you’re on the hook for it, support takes a few days to respond, they see their life flashing before their eyes because in many cases, that is more money than people in some of these places will expect to earn in a year, and does something horribly tragic. And at that point, there’s a bell that has been rung that cannot be unrung.

Of all the things I want to fix, yeah, I complain and I whine about an awful lot of stuff, but this is the one that has the most tragic consequences. No story for a human is going to end in tragedy because of the usurious pricing for Managed NAT Gateway data transfer, but a surprise bill that we know support is going to wipe over something like that, that is going to break people. And that’s not okay.

Rachel: No, it’s not okay. I think that you write very well about that topic in particular, and I really would love to see some changes take place. I
know that Amazon knows their business better than to need to rely on some Adore Me-style subscription model that you can’t figure out how to get out of. Like, have some faith in your products or don’t sell it.

Corey: I really, really wish that more companies saw it that way. And the hell of it is the best shining example is a recurring sponsor of this show: Oracle Cloud. Oracle is, let’s be honest, they’re Oracle; that’s less a brand than a warning label in many cases, but I’ve often said the Oracle Cloud biggest challenge is the word Oracle at the front of it—

Rachel: Absolutely.

Corey: —because their service offering is legitimate, their free tier is actually free—I’ve been running some fairly beefy stuff there for over a year, and have never been charged a dime for it. And it’s not because I’m special; it’s because I haven’t taken the affirmative upgrade-my-account step. And their data transfer pricing is great. Within the confines of those things, yeah, it’s terrific. I can’t speak to what it looks like a super large-scale for a cloud-native app, yet, but that’s going to change; people are starting to take them a lot more seriously.

And I’ve got to say, in previous years in the re:Invent keynotes, they’ve made fun and kicked at Oracle a fair bit, which no one has any sympathy for. Now, I don’t think that would lend the same way, just among people who have decided to suspend disbelief long enough and kick the tires in the Oracle free tier. It’s like, well, yeah, you can say a lot of negative things about Oracle—and I have a list of them—but you know, what I never got with Oracle: A surprise bill. And its Oracle we’re talking about, where surprise billing is the entire reason that they—

Rachel: It’s the model.

Corey: —are a company.

Rachel: Yeah. [laugh].

Corey: That is the model. And in this case, they are nailing it. And I’ve often said that you can buy my attention, but not my opinion. Long before they sponsored this show, I was talking, like, this about this particular offering. “Oh, so you’re saying we should migrate everything to Oracle databases?” “Good, Lord, no. Not without talking with someone who’s been down that path.” And almost everyone who has will scream at you about it. It’s a separate model. It’s a separate division. It’s a separate way of thinking about things. And I’m a big fan of that.

Rachel: Oh, that’s great. There have been ruinous results of Oracle’s decisions and acquisitions in our industry, and yet, this does appear to be a slice of the market that they have given autonomy to the people running it. And I feel like that’s really the key. I know just a hair about the product process—the new product introduction process at Amazon in general, And therefore, I actually do have a bit of faith that they will fix this. It’s just a huge problem, and when Oracle is eating your lunch, I mean, I just—you really have some things to reconsider.

Corey: This episode is sponsored in part by our friends at Rising Cloud, which I hadn’t heard of before, but they’re doing something vaguely interesting here. They are using AI, which is usually where my eyes glaze over and I lose attention, but they’re using it to help developers be more efficient by reducing repetitive tasks. So, the idea being that you can run stateless things without having to worry about scaling, placement, et cetera, and the rest. They claim significant cost savings, and they’re able to wind up taking what you’re running as it is, in AWS, with no changes, and run it inside of their data centers that span multiple regions. I’m somewhat skeptical, but their customers seem to really like them, so that’s one of those areas where I really have a hard time being too snarky about it because when you solve a customer’s problem, and they get out there in public and say, “We’re solving a problem,” it’s very hard to snark about that. Multus Medical, Construx.ai, and Stax have seen significant results by using them, and it’s worth exploring. So, if you’re looking for a smarter, faster, cheaper alternative to EC2, Lambda, or batch, consider checking them out. Visit risingcloud.com/benefits. That’s risingcloud.com/benefits, and be sure to tell them that I said you because watching people wince when you mention my name is one of the guilty pleasures of listening to this podcast.

Corey: I am an Amazon fan. I think that given the talent, and the insight, and the drive that they have there—not to mention the fact that they’re a $1.6 trillion company—if they want to do something, it will get done. And there are very few bounds I would put on it. Which means that everything that Amazon does, is, on some level, a choice. There are very few things they could not achieve with concerted effort if they cared enough.

Corey: I want to also tell a story about you for a change, because why not? Back in 2018, I was just really getting to have an audience, and the rest, and I found myself at the replay party at re:Invent. And it was a weird moment for me because I’d finished most of my speaking stuff, I had hung out with my meetups and my friends and the rest, and I’m wandering around the party—

Rachel: Your DevOps stand-up, as I recall.

Corey: That’s what it w—that’s what it was. Yeah, my DevOps stand-up, cloud comedy, whatever you want to call it. And I’m walking around, and it’s isolating and weird after something like that—back in the before times, at least—and when people know me as a character, more or less, but not as a person, and it’s isolating, and it’s lonely, and it’s—again, you don’t feel great after four days in Las Vegas, and it’s dark, and it’s hard to tell who’s who we ran into each other and just started walking around and having a conversation outside because apparently 4000 decibels as a little much for volume for both of us. And it was just great finding someone who I can talk to as a human being. There’s not enough of that in different ways. Because remember, back then, I was an independent consultant I didn’t have colleagues to hang out with. It was—

Rachel: Oh, that was pre-Duckbill.

Corey: That was when I was still the Quinn Advisory Group.

Rachel: Oh, very good. Okay. Yes, I do remember that.

Corey: The Duckbill Group was formed about a month-and-a-half after that as memory serves.

Rachel: Oh, okay. Cool.

Corey: But yeah, same problem. It’s, how do I build this? How do I turn this into something was a separate problem that hadn’t quite—hadn’t come up with an answer yet. So, I’m an independent consultant, wandering around, feeling lonely. My clients are all off doing their own things because it turns out that I’m great at representing clients in meetings with Amazon execs, but lousy at representing them on the dance floor.

So, it was just the empathy that exuded from you was just phenomenal. And I don’t know ever thank you for just how refreshing it was to be able to just step back from the show for a minute and be a person. So thanks.

Rachel: Oh, likewise. I remember I had gotten in touch with you beforehand as well to say, like, “I’m going to be at re:Invent. I don’t know any women who will be there. Can you please introduce me to some?” And you introduce me to some lovely people who, along with you, really helped me navigate my first re:Invent in a huge way, which was—you think it’s going to be overwhelming, multiply that by ten or a hundred. That is how much information is coming at you all the time when you are at re:Invent.

So, to go to this funny party where there was like some EDM DJ, who I think was, like, well-known or something in 2018, be like, [laugh] that’s really not my thing. But I want to bum around this party, I do want to see what’s going on, and if I can touch base with anybody else that I have met during this conference. And I remember we, kind of like, stuck close to each other. And that was so—that was, it was so human. And I appreciated that so much from you as well.

I was sent by my company—as anybody who goes to [OSCON 00:31:03] or re:Invent are, if they pay full freight [laugh]—it was so lovely to just have a buddy to bum around with and make fun of things, and talk shop, and everything in between.

Corey: I do want to give one small tip, something buried in there that I think is just something I’ve been doing extensively for a while, but I haven’t really ever called it out, or at least not recently—and I’ll do a tweet thread about this after we’re done recording—the counterpoint that I want to that I want to point out is that introductions are great, but every person I introduced you to, I had your permission to give their email address to them, and I reached out to them independently in every case and said, “Hey, someone would like”—once I was had your permission to reference you—“She would like to talk to other folks who don’t look like me who are going to re:Invent. May I introduce you?” The idea of a double opt-in introduction goes so far. And I’m talking about this for folks who aren’t me. In my case, fine. If some rando wants to introduce me to some other rando, knock yourself out. There is very little showing up in my inbox that I am not going to have some way of handling. But not everyone thinks about things that way, and it just shows a baseline level of human respect.

Rachel: Yeah, absolutely. I actually just did that this morning. I’m sure all of us get these calls a few times a year: “I’m thinking about switching to tech because the money’s there, the stability is there, the job market is there, and I have been underpaid and treated poorly for a long time,” or whatever variation on that story that I know we all are aware of. And I talked with him for a while last night, and then I put him in touch with the dual opt-in emails with someone in the field that he’s looking at, exactly, and a recruiter friend of mine to help give more perspective on the industry as a whole. And with both of those people, I asked permission to introduce them to the friend of mine who had reached out to me, and both of them responded right away because when you are fielding questions like these all day, you become familiar with the kindest way to do that.

And I really love being able to use my network in that way. Yes, I know a person at X, and yes, I would love to introduce you to Y. And I will make sure that everybody agrees and knows that this is coming, and I’m not just taken by surprise. Where I do get those emails and I understand that etiquette is something to learn, it isn’t directly common-sense sometimes. And then you sit down and you think about it, or someone says to you like, “I really need you to give me a heads up before giving my contact information to someone that I don’t know.”

Corey: It happens. It’s about being accessible. It’s about making the industry better than it is. And on that topic, I have one more area I want to delve into before we call it a show, and that is you are on the program committee for SeaGL, the Seattle GNU/Linux conference.

Rachel: That’s right.

Corey: I have fond memories of that conference, once upon a time. I gave a keynote a few years ago back when I was, you know, able to go places without it being a deadly risk, and much more involved in the community side of the world when it comes to conferences. I’ve unfortunately pulled back from a lot of it, just due to demands on my time. But great conference. Enjoyed a lot of the conversations once you, sort of, steered around the true believers around some areas of things, to the point where it subverts, you know, being civil to people. But it was a good conference. There was a lot to recommend it.

Rachel: SeaGL is a beautiful little conference. It is community-focused. We don’t let sponsors get on stage. We really restrict how much the people giving us money are able to dictate what we do. What we do is create a platform for people to discuss open-source in a human way, I would say.

I think in our earlier days, we had a lot of focus on software freedom at all costs, and that has softened in the name of humans and social justice in a way that I feel very proud of. I have been the program chair for three years now, and it’s just wonderful seeing the trends that come up every year. Our conference is Friday and Saturday, November 5th and 6th, so I hope that by the time you hear this, you will still have an opportunity to go to that; I’m not sure. Some of the themes this year have just been so interesting. It’s all about—and this will be very interesting to a particular subset of people, and maybe not to everybody—but about open-source governance, and how do we maintain the soul and the purpose of an open-source project, while keeping people housed and fed who are working on these things, and to not sign over all the rights of a given project to our corporate overlords and such.

So, there’s a number of talks that are going to be talking about that. A few years ago, the trend that I was really excited about that I personally gave a talk about as well, is how to start owning and managing your own data entirely. I gave a talk on trying to get off Google, which is Herculean and close to impossible. And I understand that, and that’s frustrating. But you know, we see these trends where we’re trying to help our community protect itself and remain open at the same time in a technical and open-source context. And it’s just an exciting and lovely organization and event each year. This is our second year being virtual. I was shocked by how good our virtual experience was last year. And I have high hopes for this year, too. So, I hope you can come check it out.

Corey: I would highly recommend it though I believe this will be airing after the show goes out.

Rachel: Ah darn.

Corey: But there’s always next year.

Rachel: That’s right. And they’re all recorded as well, all the talks will be recorded. The publication date on those might be a little bit after but
yes, they will all be up.

Corey: But we will of course include links to that in the [show notes 00:37:13] because there’s always next year.

Rachel: That’s right.

Corey: I want to thank you so much for taking the time to speak with me. If people want to learn more, where can they find you?

Rachel: I think probably the best place is on Twitter. That is @wholemilk on Twitter. Like, the dairy product by the gallon that’s me.

Corey: And that link to that will go in the [show notes 00:37:33] as well. Thank you so much for taking the time to speak with me today. I really appreciate it.

Rachel: Thank you, Corey. This has been great.

Corey: It really has. Rachel Kelly, senior infrastructure engineer at Fastly. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with a comment telling me that you should absolutely shove your business logic fully into the CDN, then wind up not being able to edit the comment because it’s locked to a single CDN.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Deirdré

For over 35 years, Deirdré Straughan has been helping technologies grow and thrive through marketing and community. Her product experience spans consumer apps and devices, cloud services and technologies, and kernel features. Her toolkit includes words, websites, blogs, communities, events, video, social, marketing, and more. She has written and edited technical books and blog posts, filmed and produced videos, and organized meetups, conferences, and conference talks. She just started a new gig heading up open source community at Intel. You can find her @deirdres on Twitter, and she also shares her opinions on beginningwithi.com

Links:

  • “Marketing Your Tech Talent”: https://youtu.be/9pGSIE7grSs
  • Personal Webpage: https://beginningwithi.com
  • Twitter: https://twitter.com/deirdres

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by LaunchDarkly. Take a look at what it takes to get your code into production. I’m going to just guess that it’s awful because it’s always awful. No one loves their deployment process. What if launching new features didn’t require you to do a full-on code and possibly infrastructure deploy? What if you could test on a small subset of users and then roll it back immediately if results aren’t what you expect? LaunchDarkly does exactly this. To learn more, visit launchdarkly.com and tell them Corey sent you, and watch for the wince.

Corey: This episode is sponsored in part by our friends at Rising Cloud, which I hadn’t heard of before, but they’re doing something vaguely interesting here. They are using AI, which is usually where my eyes glaze over and I lose attention, but they’re using it to help developers be more efficient by reducing repetitive tasks. So, the idea being that you can run stateless things without having to worry about scaling, placement, et cetera, and the rest. They claim significant cost savings, and they’re able to wind up taking what you’re running as it is, in AWS, with no changes, and run it inside of their data centers that span multiple regions. I’m somewhat skeptical, but their customers seem to really like them, so that’s one of those areas where I really have a hard time being too snarky about it because when you solve a customer’s problem, and they get out there in public and say, “We’re solving a problem,” it’s very hard to snark about that. Multus Medical, Construx.ai, and Stax have seen significant results by using them, and it’s worth exploring. So, if you’re looking for a smarter, faster, cheaper alternative to EC2, Lambda, or batch, consider checking them out. Visit risingcloud.com/benefits. That’s risingcloud.com/benefits, and be sure to tell them that I said you because watching people wince when you mention my name is one of the guilty pleasures of listening to this podcast.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. One of the best parts about running this podcast has been that I can go through old notes of conferences I’ve went to, and the people whose talks I’ve seen, the folks who have done interesting things that back when I had no idea what I was doing—as if I do now—and these are people I deeply admire. And now I have an excuse to reach out to them and drag them onto this show to basically tell them that until they blush. And today is no exception for that. Deirdré Straughan has had a career that has spanned three decades, I believe, if I’m remembering correctly.

Deirdré: A bit more, even.

Corey: Indeed. And you’ve been in I want to say marketing, but I’m scared to frame it that way, not because that’s not what you’ve been doing, but because so few people do marketing to technical audiences well, that the way you do it is so otherworldly good compared to what is out there that it almost certainly gives the wrong impression. So, first things first. Thank you for joining me.

Deirdré: Very happy to. Thank you for having me. It’s always a delight to talk with you.

Corey: So, what is it you’d say it is you do, exactly? Because I’m doing a very weak job of explaining it in a way that is easy for folks who have never heard of you before—which is a failing—to contextualize?

Deirdré: Um, well, there’s one—you know, I was until recently working for AWS, and one of the—went to an internal conference once at which they said—it was a marketing conference, and they said, “As the marketing organization, our job is to educate.” Now, you can discuss whether or not we think AWS does that well, but I deeply agree with that statement, that as marketers, our job is to educate people. You know, the classical marketing is to educate people about the benefits of your product. You know, “Here’s why ours is better.” The Kathy Sierra approach to that, which I think is very, very wise is, don’t market your product by telling people how wonderful the product is. Tell them how they can kick ass with it.

Corey: How do you wind up disambiguating between that and, let’s just say it’s almost a trope at this point where someone will talk about something, be it a product, be it an entire Web3 thing, whatever, and when someone comes back and says, “Well, I don’t think that’s a great idea.” The response is, “Oh, no, no. You just need to be educated properly about it.” Or, “Do your own research.” That sort of thing. And that is to be clear, not anything I’ve ever seen you say, do, or imply. But that almost feels like the wrong direction to take that in, of educating folks.

Deirdré: Well, yeah, I mean, the way it’s used in those terms, it sounds condescending. In my earliest, earlier part of my career, I was dealing with consumer software. So, this was in the early days of CD recording. We were among the pioneering CD recording products, and the idea was to make it—my Italian boss saw this market coming because he was doing recording CDs as a service, like, you were a law firm that needed to store a lot of data, and he would cut a CD for you, and you would store that. And you know, this was on a refrigerator-sized thing with a command-line interface, very difficult to use, very easy to waste these $100 blank CDs.

But he was following the market, and he saw that there was going to be these half-height CD-ROM drives. And he said, “Well, what we need to go with that is software that is actually usable by the consumer.” And that’s what we did; we created that software. And so in that case, there were things the customer still had to know about CDR, but my approach was that, you know, I do the documentation, I have to explain this stuff, but I should have to explain less and less. More and more of that should be driven into the interface and just be so obvious and intuitive that nobody ever has to read a manual. So, education can be any of those things. Your software can be educating the customer while they’re using it.

Corey: I wish that were one of those things we could point out and say, “Well, yeah, years later, it’s blindingly obvious to everyone.” Except for the part where it’s not, where every once in a while on Twitter, I will go and try a new service some cloud company launches, or something else I’ve heard about, and I will, effectively, screenshot and then live tweet my experiences with it. And very often—I’ll get accused of people saying, “Ahh, you’re pretending to be dumb and not understanding that’s how that interface works.” No, I’m not. It turns out that the failure mode of bad interfaces and of not getting this right is not that people look at it and say, “Ah, that product is crap.” It’s that, “Oh, I’m dumb, and no one ever told me about it.”

That’s why I’m so adamant about this. Because if I’m looking at an interface and I get something wrong, it is extremely unlikely that I’m the only person who ever has. And it goes beyond interfaces, it goes out to marketing as well with poor messaging around a product—when I say marketing, I’m talking the traditional sense of telling a story, and here’s a press release. “Great. You’ve told me what it does, you told me about big customers and the rest, but you haven’t told me what painful problem do I have that it solves? And why should I care about it?” Almost like that’s the foregone conclusion.

No, no. We’re much more interested in making sure that they get the company name and history right in the ‘About Us’ at the bottom of the
press release. And it’s missing the forest for the trees, in many respects. It’s—

Deirdré: Yeah.

Corey: —some level—it suffers from a similar problem of sales, where you have an entire field that is judged based upon some of the worst examples out there. And on the technical side of the world—and again, all these roles are technical, but the more traditional, ‘I write code for a living’ types, there’s almost a condescension or a dismissiveness that is brought toward people who work in sales, or in marketing, or honestly, anything that doesn’t spend all their time staring into an IDE for a living. You know, the people who get to do something that makes them happy, as opposed to this misery that the coder types that we sometimes find ourselves trapped into. How have you seen that?

Deirdré: Yeah. And it’s also a condescension towards customers.

Corey: Oh absolutely.

Deirdré: I have seen so many engineers who will, you know, throw something out there and say, “This is the most beautiful, sexy, amazing thing I’ve ever done.” And there have been a few occasions when I’ve looked at it and gone, you know, “Yes, I can see how from a technical point of view, that’s beautiful and amazing and sexy, but no customer is ever going to use it.” Either because they don’t need it or because they won’t understand it. There’s no way in that context to have that make sense. And so yeah, you can do beautiful, brilliant engineering, but if you never sell it and no one ever uses it, what’s the point?

Corey: One am I of the ways that I’ve always found to tell a story that resonates—and it sometimes takes people by surprise when they’re doing a sponsorship or something I do, or whatnot, and they’re sitting there talking about how awesome everything is, and hey, let’s do a webinar together. And it’s cool, we can do that, but I’d rather talk to one of your customers because you can say anything you want about your product, and I can sit here and make fun of it because I have deep-seated personality problems, and that’s great. But when a customer says, “I have this problem, and this is the thing that I pay money for to fix that problem,” it is much harder for people to dismiss that because you’re voting with your dollars. You’re not saying this because if your product succeeds, you get to go buy a car or something. Now, someone instead is saying this because, “I had a painful point, and not only am I willing to pay money to make this painful thing go away, but then I want to go out in public and talk about that.”

That is an incredibly hard thing to refute, bordering on the impossible, in some circumstances. That’s what always moved me. If you have a customer telling stories about how great something is, I will listen. If you have your own internal employees talking about great something is, I have some snark for you.

Deirdré: And that is another thing AWS gets right, is they—

Corey: Oh, very much so.

Deirdré: —work very hard to get the customer in front of the audience. Although, with a new technology service, et cetera, there was a point before you may have those customers in which the other kind of talk, where you have a highly technical engineer speaking to a highly
technical audience and saying, “Here’s our shiny new thing and here’s what you can do with it,” then you get the customers who will come along later and say, “Yes, we did thing with the shiny new thing, and it was great.” An engineer talking about what they did is not always to be overlooked.

Corey: Your career trajectory has been fascinating to me in a variety of different ways. You were at Sun Microsystems. And I guess personally, I just hope that when you decide to write your memoirs, you title it, The Sun Also Crashes. You know, it’s such a great title; I haven’t seen anything use it yet, and I hope I live to see someone doing that.

And then you were at Oracle for ten months—wonder how that happened? For those who are unaware, there was an acquisition story—and then you went to spend three-and-a-half years running educational programs and community at Joyent, back before. Community architect—which is what you were at the time—was really a thing. Community was just the people that showed up to talk about the technology that you’ve done. You were one of the first people that I can think of in this industry when I’ve been paying attention, who treated it as something more than that. How do you get there?

Deirdré: So, my early career, I was living in Italy because I was married to an Italian at the time, and I had already been working in tech before I left the United States, and enjoyed it and wanted to continue it. But there was not much happening in tech in Italy then. And I just got very, very lucky; I fell in with this Italian software entrepreneur—absolute madman—and he was extremely unusual in Italy in those days. He was basically doing a Silicon Valley-style software startup in Milan. And self-funded, partly funded by his wealthy girlfriend. You know, we were small, scrappy, all of that. And so he decided that he could make better software to do CD recording, as these CD-ROM drives were becoming cheaper, and he could foresee that there would be a consumer market for them.

Corey: What era was this? Because I remember—

Deirdré: This—

Corey: —back when I was in school, basically when I was failing out of college, burning a bunch of CDRs to play there, and every single tool I ever used was crap. You’re right. This was a problem.

Deirdré: So, we started on that software in, ohh, ’91.

Corey: Yeah.

Deirdré: Yeah. His goal was, “I’m going to make the leading CD recording software for the Windows market.” Hired a bunch of smart engineers, of which there are plenty in Italy, and started building this thing. I had done a project for him, documenting another OCR—Optical Character Recognition—product, and he said, “How would you like to write a book together about CD recording?” And it’s like, “Okay, sure.”

So, we wrote this book, and, you know, it was like, basically, me reading and him explaining to me the various color book specs from Philips and Sony that explain, you know, right down to the pits and lands, how CD recording works, and then me translating it into layman’s terms. And so the book got published in January of 1993 by Random House. It’s one of the first books, if not the first book in the world to actually be published with a CD included.

Corey: Oh, so you’re ultimately the person who’s responsible—indirectly—for hey, you could send CDs out, and then the sea of AOL mailers showing up—basically the mini-frisbee plague that lasted a decade or so, for the rest of us?

Deirdré: Yeah. And this was all marketing. For him, the whole idea of writing a book was a marketing ploy because on the CD, we included a trial version of the software. And that was all he wanted to put on there, but I thought, “Well, let’s take this a step further.” This was—I had been also doing a little bit of work in journalism, just to scrape by in Italy.

I was actually an Italian computer journalist, and I was getting sent to conferences, including the launch of Adobe PDF. Like, they sent me to Scotland to learn about PDFs. Like, “Okay.” But then it wasn’t quite ready at the time, so I ended up using FrameMaker instead. But I made an entire hypertext version of that book and put it on that CD, which was launched in early ’93 when the internet was barely becoming a thing.

So, we launched the book, sold the book. Turned out the CD had been manufactured wrong and did not work.

Corey: Oh, dear.

Deirdré: And I was just dying. And the publisher said, “Well, you know, if you can get ahold of the readers, the people”—you know, because they were getting complaints—they said, “If you can reach the readers somehow and let them know, there’s a number they can call and we’ll send them a replacement disk.” We had put our CompuServe email address in the book. It’s like, “Hey, we’d love to hear from you. Write to us at”—

Corey: Weren’t those the long string of numbers as a username.

Deirdré: Yeah.

Corey: Yeah.

Deirdré: Mm-hm. You could reach it via external email at the time, I believe. And we didn’t really expect that many people would bother. But, you know, because there was this problem, we were getting a lot of contacts. And so I was like, I was determined I was going to solve this situation, and I was interacting with them.

And those were my first experiences with interacting with customers, especially online. You know, and we did have a solution; we were able to defuse the situation and get it fixed, but, you know, so that was when I realized it was very powerful because I could communicate very quickly with people anywhere in the world, and—quickly over whatever the modem speed was [laugh] at that time, you know, 1800 baud or something. And so I got intr—I had already been using CompuServe when I was in college, and so I was interested in how do you communicate with people in this new medium.

And I started applying that to my work. And then I went and applied it everywhere. It’s like, “Okay, well, there’s this new thing coming, you know, called the internet. Well, how can I use that?” Publishing a paper manual seems kind of stupid in this day and age, so I can update them much more quickly if I have it on a website.

So, by that time, the company had been acquired by Adaptec. Adaptec had a website, which was mostly about their cables and things, and so I just, kind of, made a section of the website. It was like, “Here is all about CDR.” And it got to where it was driving 70% of the traffic to Adaptec, even though our products were a small percentage of the revenue. And at the same time, I was interacting with customers on the Usenet and by email.

Corey: And then later, mailing lists, and the rest. And now it—we take it for granted, but it used to be that so much of this was unidirectional, where at an absolute high level, the best you could hope for in some cases is, “I really have something to say to this author. I’m going to write a letter and mail it to the publisher and hope that they forward it.” And you never really know if it’s going to wind up landing or not? Now it’s, “I’m going to jump on Twitter and tell this person what I think.”

And whether that’s a good or bad change, it has changed the world. And it’s no longer unidirectional where your customers just silent masses anymore, regardless of what you wind up doing or selling. And I sell consulting services. Yeah, I deal with customers a lot; we have high bandwidth conversations, but I also do an annual charity t-shirt drive and I get a lot of feedback and a lot of challenges with deliveries in the rest toward the end of the year. And that is something else. We have to do it. It’s not what it used to be just mail a self-addressed stamped envelope to somewhere, and hope for the best. And we’ll blame the post office if it doesn’t work. The world changed, and it’s strange that happens in your own lifetime.

Deirdré: Yeah. And there were people who saw it coming, early on. I became aware of The Cluetrain Manifesto because a customer wrote to me and said, I think you’re the best example I see out there of people actually living this. And The Cluetrain Manifesto said, “The internet is going to change how companies interact with customers. You are going to have to be part of a conversation, rather than just, we talk to you and tell you what’s what.” And I was already embracing that.

And then it has had profound implications. It’s, in some ways, a democratization of companies and their products because people can suddenly be very vociferous about what they think about your product and what they want improved, and features they’d like added, and so forth. And I never said the customer is always right, but the customer should always be treated politely. And so I just developed this—it was me, but it was a persona which was true to me, where I am out here, I’m interacting with people, I am extremely forthcoming and honest—

Corey: That you are, which is always appreciated, to be clear. I have a keen appreciation for folks who I know beyond the shadow of a doubt will tell me where I stand with them. I’ve never been a fan of folks who will, “I can’t stand that guy. Oh, great, here he comes. Hi.” No.

There is something very refreshing about the way that you approach honesty, and that you have always had that. And it manifests in different forms. You are one of those people where if you say something in public, be it in writing, be it on stage, be it in your work, you believe it. There has never been a shadow of doubt in my mind that someone could pay you to say something or advocate for something in which you do not believe.

Deirdré: Thanks. Yeah, it’s just partly because I’ve never been good at lying. It just makes me so deeply uncomfortable that I can’t do it. [laugh].

Corey: That’s what a good liar would say, let’s be very clear here. Like, what’s the old joke? Like, “If you can only be good at one thing, be good at lying because then you’re good at everything.” No.

Deirdré: [laugh].

Corey: It’s a terrible way to go through life.

Deirdré: Yeah. And the earn trust thing was part of my… portfolio from very early on. Which was hilarious because in those days, as now, there were people whose knee-jerk reaction was, if you’re out here representing a company, you automatically must be lying to me, or about to lie to me, or have lied to me. But because I had been so out there and so honest, I had dozens of supporters who would pile in and say, “No, no, no. That’s not who she is.” And so it was, yeah, it was interesting. I had my trolls but I also had lots of defenders.

Corey: The real thing that I’ve seen as well sometimes is when someone is accused of something like that, people will chime in—look, like, I get this myself. People like you. I don’t generally have that problem—but people will chime in with, like, “I don’t like Corey, but no, he’s generally right about these things.” That’s, okay, great. It’s like, the backhanded compliment. And I’ll take what I can get.

I want to fast-forward in time a little bit from the era of mailing books with CDs in them, and then having to talk to people via other ways to get them in CompuServe to 2013 when you gave a talk at one of—no, I’m not going to say, ‘one of.’ It is the best community conference of which I am aware. Monktoberfest as put on by our friends at RedMonk. It was called “Marketing Your Tech Talent” and it’s one of those videos it’s worth the watch. If you’re listening to this, and you haven’t seen it, you absolutely should fix that. Tell me about it. Where did the talk come from?

Deirdré: As you can see in the talk, it was stuff I had been doing. It actually started earlier than that. When I joined Sun Microsystems as a contractor in 2007, my remit was to try to get Sun engineers to communicate. Like, Sun had done this big push around blogging, they’d encourage everybody to open up your own blog. Here’s our blogging platform, you can say whatever you want.

And there were, like, 3000 blogs, about half of which were just moribund; they had put out one or two posts, and then nothing ever again. And for some reason—I don’t know who decided—but they decided that engineers had goals around this and engineering teams had to start producing content in this way, which was a strange idea. So, I was brought on. It’s, like, you know, “Help these engineers communicate. Help them with blogging, and somehow find a way to get them doing it.”

And so I did a whole bunch of things from, like, running competitions to just going and talking to people. But we finally got to where Dan Maslowski, who was the manager who hired me in, he said, “Well, we’ve got this conference. It was the SNIA, the Storage Networking Industries Association Conference. We're a big sponsor, we’ve got, like, ten talks. And why don’t you just go—you know, I’m going to buy you a video camera, go record this thing.”

And I’d used a video camera a little bit, but, you know, it’s like, never in this context, so it’s like, okay, let’s figure out, you know, what kind of mic do I need? And so I went off to the conference with my video blogging rig, and videoed all those talks. And then the idea was like, “Okay, we’ll put them up on”—you know, Sun had its own video channels and things—“We’ll put it out there, and this information will then be available to more people; it’ll help the engineers communicate what they’re doing.”

And the funny part was, I run into with Sun, the professional video people wanted nothing to do with it. Like, “Your stuff is not high enough quality. You don’t meet our branding guidelines. You cannot put this on the Sun channels.” Okay, fine. So, I started putting it on YouTube, which in those days meant splitting it into ten-minute segments because that was all they would give you. [laugh]. And so it was like, everything I was doing was guerilla marketing because I was always in the teeth on somebody in the corporation who wanted to—it’s like, “Oh, we’re not going to put out video unless it can be slickly produced in the studio, and we’re only going to do that for VPs, not for engineers.”

Corey: Oh, yeah. The little people, as it were. This talk, in many ways—I don’t know if ever told you this story or not—but it did shape how I approached building out my entire approach: The sponsorship side of the business that I have, how I approach communicating with people. And it’s where in many ways, the newsletter has taken its ethos. One of the things that you mentioned in that talk was, first, you were actually the first time that I ever saw someone explicitly comparing the technical talent slash DevRel—which is not a term I would call it, but all right—to the Hollywood model, where you have this idea that there’s an agent that winds up handling these folks that are freelancers. They are named talent. They’re the ones that have the draw; that’s what people want, so we have to develop this.

Okay, what why is it important to develop this? Because you absolutely need to have your technical people writing technical content, not folks who are divorced from that entire side of the world because it doesn’t resonate, it doesn’t land. This is I think, what DevRel was sort of been turned into; it’s, what it DevRel? Well, it’s special marketing because engineers need special handling to handle these things. No, I think it’s everyone needs to be marketed to in a way that has authenticity that meets them where they are, and that’s a little harder to do with people who spend their lives writing code than it would be someone who is it was at a more accessible profession.

But I don’t think that a lot of it’s being done right. This was the first encouragement that I’d gotten early on that maybe I am onto something here because here’s someone I deeply respect saying a lot of the same things—from a slightly different angle; like I was never doing this as part of a large technology company—but it was still, there’s something here. And for better or worse. I think I’ve demonstrated by now that there is some validity there. But back then it was transformational.

Deirdré: Well, thank you.

Corey: It still kind of is in many respects. This is all new to someone.

Deirdré: Yeah. I felt, you know, I’d been putting engineers in front of the public and found it was powerful, and engineers want to hear from other engineers. And especially for companies like Sun and Oracle and Joyent, we’re selling technology to other technologists. So, there’s a
limited market for white papers because VPs and CEOs want to read those, but really, your main market is other technologists and that’s who you need to talk to and talk to them in their own way, in their own language. They weren’t even comfortable with slickly produced videos. Neither being on the camera nor watching it.

Corey: Yeah, at some point, it was like, “I look too good.” It’s like, “Oh, yeah. It’s—oh, you’re going to do a whole video production thing? Great.” “Okay. [unintelligible 00:24:13] the makeup artists coming in.” Like, “What do you mean makeup?” And it’s—

Deirdré: Oh, it was worse at Sun. We wasted so much money because you would get an engineer and put him in the studio under all these lights with these great big cameras, and they would just freeze.

Corey: Mmm.

Deirdré: And it’s like, you know, “Well, hurry up, hurry up. We’ve got half an hour of studio time. Get your thing; say it.” And, [frantic noise]. You know, whereas I would take them in some back conference room and just set up a camera and be sitting in a chair opposite. It’s like, “Relax. Tell me what you want to tell me. If we have to do ten takes, it’s fine.” Yeah, video quality wasn’t great, but the content was great.

Corey: It seems like there is a new security breach every day. Are you confident that an old SSH key or a shared admin account isn’t going to come back and bite you? If not, check out Teleport. Teleport is the easiest, most secure way to access all of your infrastructure. The open source Teleport Access Plane consolidates everything you need for secure access to your Linux and Windows servers—and I assure you there is no third option there. Kubernetes clusters, databases, and internal applications like AWS Management Console, Yankins, GitLab, Grafana, Jupyter Notebooks, and more. Teleport’s unique approach is not only more secure, it also improves developer productivity. To learn more visit: goteleport.com. And no, that is not me telling you to go away, it is: goteleport.com.

Corey: Speaking of content, one more topic I want to cover a little bit here is you recently left your job at AWS. And even if you had not told me that, I would have known because your blog has undergone something of a renaissance—beginningwithi.com for those who want to follow along, and of course, we’ll put links to this in the [show notes 00:25:08]—you’ve been suddenly talking about a lot of different things. And I want to be clear, I don’t recall any of these posts being one of those, “I just left a company, I’m going to set them on fire now.”

It’s been about a variety of different topics, though, that have been very top-of-mind for folks. You talk about things like equal work for equal pay. You talk about remote work versus cost of commuting a fair bit. And as of this recording, you most recently wound up talking specifically about problematic employers in tech. But what you’re talking about is also something that this happened during the days of the Sun acquisition through Oracle.

So, people are thinking, like, “Wait a minute, is she subtweeting what happened today”—no. These things rhyme and they repeat. I’m super thrilled whenever I see this in my RSS reader, just because it is so… they oh, good. I get I’m going to read something now that I’m going to enjoy, so let me put this in distraction-free mode and really dig into it. Because your writing is a joy.

What is it that has inspired you to bring that back to life? Is it just to having a whole bunch of free time, and well, I’m not writing marketing stocks anymore, so I guess I’m going to write blog posts instead.

Deirdré: My blog, if you looked at our calendar, over the years, it sort of comes and goes depending what else is going on in my life. I actually was starting to do a little bit more writing, and I even did a few little TikTok videos before I quit AWS. I’m starting to think about some of the more ancient history parts of my career. It’s partly just because of what’s been going on in the world. [Brendan 00:26:35] and I moved to Australia a year ago, and it was something that had been planned for a long time.

We did not actually expect that we would be able to move our jobs the way we did. And then, you know, with pandemic, everything changed; that actually accelerated our departure timeline because we’ve been planning initially to let our son stay in school in California, through until he finished elementary, but then he wasn’t in school, so there seems no point, whereas in Australia, he could be in a classroom. And so, you know, the whole world is changing, and the working world is changing, but also, we all started working from home. I’ve been working from home—mostly—since 1993. And I was working very remotely because I was working from Italy for a California company.

And because I was one of the first people doing it, the people in California did not know what to make of me. And I would get people who would just completely ignore any emails I sent. It was like as if I did not exist because they had never seen me in person. So, I would just go to California four times a year and spend a few weeks, and then I would get the face time, and after that it was easy to interact any way I needed to.

Corey: It feels like it’s almost the worst kind of remote because you have most people at office, and then you have a few outliers, and that tends to, in my experience at least, lead to a really weird team dynamics where you have almost a second class of folks who aren’t taken nearly as seriously. It’s why when we started our company here, it was everyone is going to be remote all the time. We were distributed. There is no central office because as soon as you do, that’s where things are disastrous. My business partner and I live a couple states apart.

Deirdré: Yeah. And I think that’s the fairest way to do it. In companies that have already existed, where they do have headquarters, and you know, there’s that—

Corey: Yeah, you can’t suddenly sell your office space, and all 300,000 employees [laugh] are now working from home. That’s a harder thing, too.

Deirdré: Yeah. But I think it’s interesting that the argument is being framed as like, “Oh, people work better in the office, people learn more in the office.” And we’ve even had the argument trotted out here that people should be forced back to the office because the businesses in the central business district depend on that. It’s like—

Corey: Mmm.

Deirdré: —well, what about the businesses that have since, you know in the meantime sprung up in the more suburban centers? Now, you’ve got some thriving little cafes out there now? Are we supposed to just screw them over? It’s ultimately people making economic arguments that have nothing to do with the well-being of employees. And the pandemic at least has—I think, a lot of people have come to realize that life is just too short to put up with a lot of bullshit, and by and large, commuting is bullshit. [laugh].

Corey: It’s a waste of time, it’s not great for the environment, there’s—yeah, and again, I’m not sitting here saying the entire world should do a particular thing. I don’t think that there’s one-size-fits-everyone solutions possible in this space. Some companies, it makes sense for the people involved to be in the same room. In some cases, it’s not even optional. For others, there’s no value to it, but getting there is hard.

And again, different places need to figure out what’s right for them. But it’s also the world is changing, and trying to pretend that it hasn’t, it just feels regressive, and I don’t think that’s going to align with where the industry and where people are going. Especially in full remote situations we’ve had the global pandemic, some wit on Twitter recently opined that it’s never been easier for a company to change jobs. You just have to wait for the different the new laptop to show up, and then you just join a different Zoom link, and you’re in your new job. It’s like, “You know, you’re not that far from wrong here.”

Deirdré: [laugh]. Yep.

Corey: There’s no, like, “Well, where’s the office? What’s the”—no. It is, my day-to-day looks remarkably similar, regardless of where I work.

Deirdré: Yeah.

Corey: That means something.

Deirdré: I was one of the early beneficiaries as well of this work-life balance, that I could take my kid to school in the morning, and then work, and then pick her up from school in the afternoon and spend time with her. And then California would be waking up for meetings, so after dinner, I’d be having meetings. Yeah, sometimes it was pain, but it was workable, and it gave me more flexibility, you know, whereas the times I had to commute to an office… tended to be hellish. I think part of the reason the blog has had a lot more activities I’ve just been in sort of a more reflective phase. I’ve gotten to this very privileged position where I suddenly realized, I actually have enough money to retire on, I have a husband who is extremely supportive of whatever I want to do, and I’m in a country that has a public health care system, if it doesn’t completely crumble under COVID in the next few weeks.

Corey: Hopefully, we’ll get this published before that happens.

Deirdré: Yes. And so I don’t have to work. It’s like, up to this point in my career, I have always desperately needed that next job. I don’t think I have ever been in the position of having competing offers. You know, there’s people who talk about, you know, you can always go find a better offer. It’s like, no, when you’re a weirdo like me and you’re a middle-aged woman, is not that easy.

Corey: People saying that invariably—“So, what is your formal job?” Like, “Oh, SDE3.” Like, okay, great. So, that means that they’re are mul—not just, they don’t probably need to hire you; they need to hire so many of you that they need to start segregating them with Roman numerals. Great.

Maybe that doesn’t apply to everyone. Maybe that particular skill set right now is having its moment in the sun, but there’s a lot of other folks who don’t neatly fit into those boxes. There’s something to be said for empathy. Because this is my lived experience does not mean it is yours. And trying to walk a mile in someone else’s shoes is almost increasingly—especially in the world of social media—a bit of a lost skill.

Deirdré: [laugh]. I mean, it’s partly that recruiters are not always the sharpest tools in the shed, and/or they’re very young, very new to it all. It’s just people like to go for what’s easy. And like, for example, me at the moment, it’s easy to put me in that product marketing manager box. It’s like, “Oh, I need somebody to fill that slot. You look like that person. Let’s talk.” Whereas before, people would just look at my resume and go, “I don’t know what she is.”

Corey: I really think the fact that you’ve never had competing offers just shows an extreme lack of vision from a number of companies around what marketing effectively to a technical audience can really be. It’s nice to see that what you have been advocating for and doing the work for, for your entire career is really coming into its own now.

Deirdré: Yeah. We’ll see what happens next. It’s been interesting. Yeah, I’ve never had so much attention from recruiters as when I got AWS on my resume. And then even more once it said, product marketing manager because, you know, “Okay. You've got the FAANG and you’ve got a title we recognize. Let’s talk to you.”

Corey: Exactly. That’s, “Oh, yay. You fit in that box, finally.” Because it’s always been one of those. Yeah, like, “What is it you actually do?” There’s a reason that I’ve built what I do now into the last job I’ll ever have. Because I don’t even know where to begin describing me to what I do and how I do it. Even at cocktail parties, there’s nothing I can say that doesn’t sound completely surreal. “I make fun of Amazon for a living.” It’s true, but it also sounds psychotic, and here we are. It’s—

Deirdré: Well, it’s absolutely brilliant marketing, and it’s working very well for you. So [laugh].

Corey: The realization that I had was that if this whole thing collapsed and I had to get a job again, what would I be doing? It probably isn’t engineering. It’s almost certainly much more closely aligned with marketing. I just hope I never have to find out because, honestly, I’m having way too much fun.

Deirdré: Yeah. And that’s another thing I think is changing. I think more and more of us are realizing working for other people has its limitations. You know, it can be fun, it can be exciting, depending on the company, and the team, and so on. But you’re very much beholden to the culture of the company, or the team, or whatever.

I grew up in Asia, as a child, of American expats. So, I’m what is called a third culture kid, which means I’m not totally American, even though my parents were. I’m not—you know, I grew up in Thailand, but I’m not Thai. I grew up in India, but I’m not Indian. You’re something in between.

And your tribe is actually other people like you, even if they don’t share the specific countries. Like, one of my best friends in Milan was a woman who had grown up in Brazil and France. It’s like, you know, no countries in common, but we understood that experience. And something I’ve been meaning to write about for a long time is that third culture kids tend to be really good at adapting to any culture, which can include corporate cultures.

So, every time I go into a new company, I’m treating that as a new cultural experience. It’s like, Ericsson was fascinating. It’s this very old Swedish telecom, with this wild old history, and a footprint in something like 190 countries. That makes it amazingly unique and fascinating. The thing I tripped over was I did not know anything about Swedish culture because they give cultural training to the people who are actually going to be moving to Sweden.

Corey: But not the people working elsewhere, even though you’re at a—

Deirdré: Yeah.

Corey: Yeah, it’s like, well, dealing with New Yorkers is sort of its own skill, or dealing with Israelis, which is great; they have great folks, but it’s a fun culture of management by screaming, in my experience, back when I had family living out there. It was great.

Deirdré: One of my favorite people at AWS is Israeli. [laugh].

Corey: Exactly. And it’s, you have to understand some cultural context here. And now to—even if you’re not sitting in the same place. Yeah, we’re getting better as an industry, bit by bit, brick by brick. I just hope that will wind up getting there within my lifetime, at least.

I really want to thank you for taking the time to come on the show. If people want to learn more, where can they find you?

Deirdré: Oh. Well, as you said, my website beginningwithi.com, and I am on Twitter as @deirdres. That’s D-E-I-R-D-R-E-S. [laugh]. So.

Corey: And we will, of course, include links to that in the [show notes 00:36:23].

Deirdré: So yeah, I’m pretty out there, pretty easy to find, and happy to chat with people.

Corey: Which I highly recommend. Thank you again, for being so generous with your time, not just now, but over the course of your entire career.

Deirdré: Well, I’m at a point where sometimes I can help people, and I really like to do that. The reason I ever aspired to high corporate office—which I’ve now clearly I’m not ever going to make—was because I wanted to be in a position to make a difference. And so, even if all the difference I’m making is a small one, it’s still important to me to try to do that.

Corey: Thank you again. I really do appreciate your time.

Deirdré: Okay. Well, it was great talking to you. As always.

Corey: Likewise. Deirdré Straughan, currently gloriously unemployed. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an angry insulting comment that you mailed to me on a CDR that doesn’t read.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Andrew

I create free cloud certification courses and somehow still make money.

Links:

  • ExamPro Training, Inc.: https://www.exampro.co/
  • PolyWork: https://www.polywork.com/andrewbrown
  • LinkedIn: https://www.linkedin.com/in/andrew-wc-brown
  • Twitter: https://twitter.com/andrewbrown

Transcript

Andrew: Hello, and welcome to Screaming in the Cloud with your host, Chief cloud economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Redis, the company behind the incredibly popular open source database that is not the bind DNS server. If you’re tired of managing open source Redis on your own, or you’re using one of the vanilla cloud caching services, these folks have you covered with the go to manage Redis service for global caching and primary database capabilities; Redis Enterprise. To learn more and deploy not only a cache but a single operational data platform for one Redis experience, visit redis.com/hero. Thats r-e-d-i-s.com/hero. And my thanks to my friends at Redis for sponsoring my ridiculous non-sense.

Corey: This episode is sponsored in part by our friends at Rising Cloud, which I hadn’t heard of before, but they’re doing something vaguely interesting here. They are using AI, which is usually where my eyes glaze over and I lose attention, but they’re using it to help developers be more efficient by reducing repetitive tasks. So, the idea being that you can run stateless things without having to worry about scaling, placement, et cetera, and the rest. They claim significant cost savings, and they’re able to wind up taking what you’re running as it is in AWS with no changes, and run it inside of their data centers that span multiple regions. I’m somewhat skeptical, but their customers seem to really like them, so that’s one of those areas where I really have a hard time being too snarky about it because when you solve a customer’s problem and they get out there in public and say, “We’re solving a problem,” it’s very hard to snark about that. Multus Medical, Construx.ai and Stax have seen significant results by using them. And it’s worth exploring. So, if you’re looking for a smarter, faster, cheaper alternative to EC2, Lambda, or batch, consider checking them out. Visit risingcloud.com/benefits. That’s risingcloud.com/benefits, and be sure to tell them that I said you because watching people wince when you mention my name is one of the guilty pleasures of listening to this podcast.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. My guest today is… well, he’s challenging to describe. He’s the co-founder and cloud instructor at ExamPro Training, Inc. but everyone knows him better as Andrew Brown because he does so many different things in the AWS ecosystem that it’s sometimes challenging—at least for me—to wind up keeping track of them all. Andrew, thanks for joining.

Andrew: Hey, thanks for having me on the show, Corey.

Corey: How do I even begin describing you? You’re an AWS Community Hero and have been for almost two years, I believe; you’ve done a whole bunch of work as far as training videos; you’re, I think, responsible for #100daysofcloud; you recently started showing up on my TikTok feed because I’m pretending that I am 20 years younger than I am and hanging out on TikTok with the kids, and now I feel extremely old. And obviously, you’re popping up an awful lot of places.

Andrew: Oh, yeah. A few other places like PolyWork, which is an alternative to LinkedIn, so that’s a space that I’m starting to build up on there as well. Active in Discord, Slack channels. I’m just kind of everywhere. There’s some kind of internet obsession here. My wife gets really mad and says, “Hey, maybe tone down the social media.” But I really enjoy it. So.

Corey: You’re one of those folks where I have this challenge of I wind up having a bunch of different AWS community Slacks and cloud community, Slacks and Discords and the past, and we DM on Twitter sometimes. And I’m constantly trying to figure out where was that conversational thread that I had with you? And tracking it down is an increasingly large search problem. I really wish that—forget the unified messaging platform. I want a unified search platform for all the different messaging channels that I’m using to talk to people.

Andrew: Yeah, it’s very hard to keep up with all the channels for myself there. But somehow I do seem to manage it, but just with a bit less sleep than most others.

Corey: Oh, yeah. It’s like trying to figure out, like, “All right, he said something really useful. What was that? Was that a Twitter DM? Was it on that Slack channel? Was it that Discord? No, it was on that brick that he threw through my window with a note tied to it. There we go.”

That’s always the baseline stuff of figuring out where things are. So, as I mentioned in the beginning, you are the co-founder and cloud instructor at ExamPro, which is interesting because unlike most of the community stuff that you do and are known for, you don’t generally talk about that an awful lot. What’s the deal there?

Andrew: Yeah, I think a lot of people give me a hard time because they say, Andrew, you should really be promoting yourself more and trying to make more sales, but that’s not why I’m out here doing what I’m doing. Of course, I do have a for-profit business called ExamPro, where we create cloud certification study courses for things like AWS, Azure, GCP, Terraform, Kubernetes, but you know, that money just goes to fuel what I really want to do, is just to do community activities to help people change their lives. And I just decided to do that via cloud because that’s my domain expertise. At least that’s what I say because I’ve learned up on in the last four or five years. I’m hoping that there’s some kind of impact I can make doing that.

Corey: I take a somewhat similar approach. I mean, at The Duckbill Group, we fixed the horrifying AWS bill, but I’ve always found that’s not generally a problem that people tend to advertise having. On Twitter, like, “Oh, man, my AWS bill is killing me this month. I’ve got to do something about it,” and you check where they work, and it’s like a Fortune 50. It’s, yeah, that moves markets and no one talks about that.

So, my approach was always, be out there, be present in the community, talk about this stuff, and the people who genuinely have billing problems will eventually find their way to me. That was always my approach because turning everything I do into a sales pitch doesn’t work. It just erodes confidence, it reminds people of the used mattress salesman, and I just don’t want to be that person in that community. My approach has always been if I can help someone with a 15-minute call or whatnot, yeah, let’s jump on a phone call. I’m not interested in nickel-and-diming folks.

Andrew: Yeah. I think that if you’re out there doing a lot of hard work, and a lot of it, it becomes undeniable the value you’re putting out there, and then people just will want to give you money, right? And for me, I just feel really bad about taking anybody’s money, and so even when there’s some kind of benefit—like my courses, I could charge for access for them, but I always feel I have to give something in terms of taking somebody’s money, but I would never ask anyone to give me their money. So, it’s bizarre. [laugh] so.

Corey: I had a whole bunch of people a year or so after I started asking, like, “I really find your content helpful. Can I buy you a cup of coffee or something?” And it’s, I don’t know how to charge people a dollar figure that doesn’t have a comma in it because it’s easy for me to ask a company for money; that is the currency of effort, work, et cetera, that companies are accustomed to. People view money very differently, and if I ask you personally for money versus your company for money, it’s a very different flow. So, my solution to it was to build the annual charity t-shirt drive, where it’s, great, spend 35 bucks or whatever on a snarky t-shirt once a year for ten days and all proceeds go to benefit a nonprofit that is, sort of, assuaged that.

But one of my business philosophies has always been, “Work for free before you work for cheap.” And dealing with individuals and whatnot, I do not charge them for things. It’s, “Oh, can you—I need some advice in my career. Can I pay you to give me some advice?” “No, but you can jump on a Zoom call with me.” Please, the reason I exist at all is because people who didn’t have any reason to did me favors, once upon a time, and I feel obligated to pay that forward.

Andrew: And I appreciate, you know, there are people out there that you know, do need to charge for their time. Like—

Corey: Oh. Oh, yes.

Andrew: —I won’t judge anybody that wants to. But you know, for me, it’s just I can’t do it because of the way I was raised. Like, my grandfather was very involved in the community. Like, he was recognized by the city for all of his volunteer work, and doing volunteer work was, like, mandatory for me as a kid. Like, every weekend, and so for me, it’s just like, I can’t imagine trying to take people’s money.

Which is not a great thing, but it turns out that the community is very supportive, and they will come beat you down with a stick, to give you money to make sure you keep doing what you’re doing. But you know, I could be making lots of money, but it’s just not my priority, so I’ve avoided any kind of funding so like, you know, I don’t become a money-driven company, and I will see how long that lasts, but hopefully, a lot longer.

Corey: I wish you well. And again, you’re right; no shade to anyone who winds up charging for their time to individuals. I get it. I just always had challenges with it, so I decided not to do it. The only time I find myself begrudging people who do that are someone who picked something up six months ago and decided, oh, I’m going to build some video course on how to do this thing. The end. And charge a bunch of money for it and put myself out as an expert in that space.

And you look at what the content they’re putting out is, and one, it’s inaccurate, which just drives me up a wall, and two, there’s a lack of awareness that teaching is its own skill. In some areas, I know how to teach certain things, and in other areas, I’m a complete disaster at it. Public speaking is a great example. A lot of what I do on the public speaking stage is something that comes to me somewhat naturally. So, can you teach me to be a good public speaker? Not really, it’s like, well, you gave that talk and it was bad. Could you try giving it only make it good? Like, that is not a helpful coaching statement, so I stay out of that mess.

Andrew: Yeah, I mean, it’s really challenging to know, if you feel like you’re authority enough to put something out there. And there’s been a few courses where I didn’t feel like I was the most knowledgeable, but I produced those courses, and they had done extremely well. But as I was going through the course, I was just like, “Yeah, I don’t know how any this stuff works, but this is my best guess translating from here.” And so you know, at least for my content, people have seen me as, like, the lens of AWS on top of other platforms, right? So, I might not know—I’m not an expert in Azure, but I’ve made a lot of Azure content, and I just translate that over and I talk about the frustrations around, like, using scale sets compared to AWS auto-scaling groups, and that seems to really help people get through the motions of it.

I know if I pass, at least they’ll pass, but by no means do I ever feel like an expert. Like, right now I’m doing, like, Kubernetes. Like, I have no idea how I’m doing it, but I have, like, help with three other people. And so I’ll just be honest about it and say, “Hey, yeah, I’m learning this as well, but at least I know I passed, so you know, you can pass, too.” Whatever that’s worth.

Corey: Oh, yeah. Back when I was starting out, I felt like a bit of a fraud because I didn’t know everything about the AWS billing system and
how it worked and all the different things people can do with it, and things they can ask. And now, five years later, when the industry basically acknowledges I’m an expert, I feel like a fraud because I couldn’t possibly understand everything about the AWS billing system and how it works. It’s one of those things where the more you learn, the more you realize that there is yet to learn. I’m better equipped these days to find the answers to the things I need to know, but I’m still learning things every day. If I ever get to a point of complete and total understanding of a given topic, I’m wrong. You can always go deeper.

Andrew: Yeah, I mean, by no means am I even an expert in AWS, though people seem to think that I am just because I have a lot of confidence in there and I produce a lot of content. But that’s a lot different from making a course than implementing stuff. And I do implement stuff, but you know, it’s just at the scale that I’m doing that. So, just food for thought for people there.

Corey: Oh, yeah. Whatever, I implement something. It’s great. In my previous engineering life, I would work on large-scale systems, so I know how a thing that works in your test environment is going to blow up in a production scale environment. And I bring those lessons, written on my bones the painful way, through outages, to the way that I build things now.

But the stuff that I’m building is mostly to keep my head in the game, as opposed to solving an explicit business need. Could I theoretically build a podcast transcription system on top of Transcribe or something like that for these episodes? Yeah. But I’ve been paying a person to do this for many years to do it themselves; they know the terms of art, they know how this stuff works, and they’re building a glossary as they go, and understanding the nuances of what I say and how I say it. And that is the better business outcome; that’s the answer. And if it’s production facing, I probably shouldn’t be tinkering with it too much, just based upon where the—I don’t want to be the bottleneck for the business functioning.

Andrew: I’ve been spending so much time doing the same thing over and over again, but for different cloud providers, and the more I do, the less I want to go deep on these things because I just feel like I’m dumping all this information I’m going to forget, and that I have those broad strokes, and when I need to go deep dive, I have that confidence. So, I’d really prefer people were to build up confidence in saying, “Yes, I think I can do this.” As opposed to being like, “Oh, I have proof that I know every single feature in AWS Systems Manager.” Just because, like, our platform, ExamPro, like, I built it with my co-founder, and it’s a quite a system. And so I’m going well, that’s all I need to know.

And I talk to other CTOs, and there’s only so much you need to know. And so I don’t know if there’s, like, a shift between—or difference between, like, application development where, let’s say you’re doing React and using Vercel and stuff like that, where you have to have super deep knowledge for that technical stack, whereas cloud is so broad or diverse that maybe just having confidence and hypothesizing the work that you can do and seeing what the outcome is a bit different, right? Not having to prove one hundred percent that you know it inside and out on day one, but have the confidence.

Corey: And there’s a lot of validity to that and a lot of value to it. It’s the magic word I always found in interviewing, on both sides of the interview table, has always been someone who’s unsure about something start with, “I’m not sure, but if I had to guess,” and then say whatever it is you were going to say. Because if you get it right, wow, you’re really good at figuring this out, and your understanding is pretty decent. If you’re wrong, well, you’ve shown them how you think but you’ve also called them out because you’re allowed to be wrong; you’re not allowed to be authoritatively wrong. Because once that happens, I can’t trust anything you say.

Andrew: Yeah. In terms of, like, how do cloud certifications help you for your career path? I mean, I find that they’re really well structured, and they give you a goal to work towards. So, like, passing that exam is your motivation to make sure that you complete it. Do employers care? It depends. I would say mostly no. I mean, for me, like, when I’m hiring, I actually do care about certifications because we make certification courses but—

Corey: In your case, you’re a very specific expression of this that is not typical.

Andrew: Yeah. And there are some, like, cases where, like, if you work for a larger cloud consultancy, you’re expected to have a professional certification so that customers feel secure in your ability to execute. But it’s not like they were trying to hire you with that requirement, right? And so I hope that people realize that and that they look at showing that practical skills, by building up cloud projects. And so that’s usually a strong pairing I’ll have, which is like, “Great. Get the certifications to help you just have a structured journey, and then do a Cloud project to prove that you can do what you say you can do.”

Corey: One area where I’ve seen certifications act as an interesting proxy for knowledge is when you have a company that has 5000 folks who work in IT in varying ways, and, “All right. We’re doing a big old cloud migration.” The certification program, in many respects, seems to act as a bit of a proxy for gauging where people are on upskilling, how much they have to learn, where they are in that journey. And at that scale, it begins to make some sense to me. Where do you stand on that?

Andrew: Yeah. I mean, it’s hard because it really depends on how those paths are built. So, when you look at the AWS certification roadmap, they have the Certified Cloud Practitioner, they have three associates, two professionals, and a bunch of specialties. And I think that you might think, “Well, oh, solutions architect must be very popular.” But I think that’s because AWS decided to make the most popular, the most generic one called that, and so you might think that’s what’s most popular.

But what they probably should have done is renamed that Solution Architect to be a Cloud Engineer because very few people become Solutions Architect. Like that’s more… if there’s Junior Solutions Architect, I don’t know where they are, but Solutions Architect is more of, like, a senior role where you have strong communications, pre-sales, obviously, the role is going to vary based on what companies decide a Solution Architect is—

Corey: Oh, absolutely take a solutions architect, give him a crash course in finance, and we call them a cloud economist.

Andrew: Sure. You just add modifiers there, and they’re something else. And so I really think that they should have named that one as the cloud engineer, and they should have extracted it out as its own tier. So, you’d have the Fundamental, the Certified Cloud Practitioner, then the Cloud Engineer, and then you could say, “Look, now you could do developer or the sysops.” And so you’re creating this path where you have a better trajectory to see where people really want to go.

But the problem is, a lot of people come in and they just do the solutions architect, and then they don’t even touch the other two because they say, well, I got an associate, so I’ll move on the next one. So, I think there’s some structuring there that comes into play. You look at Azure, they’ve really, really caught up to AWS, and may I might even say surpass them in terms of the quality and the way they market them and how they construct their certifications. There’s things I don’t like about them, but they have, like, all these fundamental certifications. Like, you have Azure Fundamentals, Data Fundamentals, AI Fundamentals, there’s a Security Fundamentals.

And to me, that’s a lot more valuable than going over to an associate. And so I did all those, and you know, I still think, like, should I go translate those over for AWS because you have to wait for a specialty before you pick up security. And they say, like, it’s intertwined with all the certifications, but, really isn’t. Like—and I feel like that would be a lot better for AWS. But that’s just my personal opinion. So.

Corey: My experience with AWS certifications has been somewhat minimal. I got the Cloud Practitioner a few years ago, under the working theory of I wanted to get into the certified lounge at some of the events because sometimes I needed to charge things and grab a cup of coffee. I viewed it as a lounge pass with a really strange entrance questionnaire. And in my case, yeah, I passed it relatively easily; if not, I would have some questions about how much I actually know about these things. As I recall, I got one question wrong because I was honest, instead of going by the book answer for, “How long does it take to restore an RDS database from a snapshot?”

I’ve had some edge cases there that give the wrong answer, except that’s what happened. And then I wound up having that expire and lapse. And okay, now I’ll do it—it was in beta at the time, but I got the sysops associate cert to go with it. And that had a whole bunch of trivia thrown into it, like, “Which of these is the proper syntax for this thing?” And that’s the kind of question that’s always bothered me because when I’m trying to figure things like that out, I have entire internet at my fingertips. Understanding the exact syntax, or command-line option, or flag that needs to do a thing is a five-second Google search away in most cases. But measuring for people’s ability to memorize and retain that has always struck me as a relatively poor proxy for knowledge.

Andrew: It’s hard across the board. Like Azure, AWS, GCP, they all have different approaches—like, Terraform, all of them, they’re all different. And you know, when you go to interview process, you have to kind of extract where the value is. And I would think that the majority of the industry, you know, don’t have best practices when hiring, there’s, like, a superficial—AWS is like, “Oh, if you do well, in STAR program format, you must speak a communicator.” Like, well, I’m dyslexic, so that stuff is not easy for me, and I will never do well in that.

So like, a lot of companies hinge on those kinds of components. And I mean, I’m sure it doesn’t matter; if you have a certain scale, you’re going to have attrition. There’s no perfect system. But when you look at these certifications, and you say, “Well, how much do they match up with the job?” Well, they don’t, right? It’s just Jeopardy.

But you know, I still think there’s value for yourself in terms of being able to internalize it. I still think that does prove that you have done something. But taking the AWS certification is not the same as taking Andrew Brown’s course. So, like, my certified cloud practitioner was built after I did GCP, Oracle Cloud, Azure Fundamentals, a bunch of other Azure fundamental certifications, cloud-native stuff, and then I brought it over because was missing, right? So like, if you went through my course, and that I had a qualifier, then I could attest to say, like, you are of this skill level, right?

But it really depends on what that testament is and whether somebody even cares about what my opinion of, like, your skillset is. But I can’t imagine like, when you have a security incident, there’s going to be a pop-up that shows you multiple-choice answer to remediate the security incident. Now, we might get there at some point, right, with all the cloud automation, but we’re not there yet.

Corey: It’s been sort of thing we’ve been chasing and never quite get there. I wish. I hope I live to see it truly I do. My belief is also that the value of a certification changes depending upon what career stage someone is at. Regardless of what level you are at, a hiring manager or a company is looking for more or less a piece of paper that attests that they’re to solve the problem that they are hiring to solve.

And entry-level, that is often a degree or a certification or something like that in the space that shows you have at least the baseline fundamentals slash know how to learn things. After a few years, I feel like that starts to shift into okay, you’ve worked in various places solving similar problems on your resume that the type that we have—because the most valuable thing you can hear when you ask someone, “How would we solve this problem?” Is, “Well, the last time I solved it, here’s what we learned.” Great. That’s experience. There’s no compression algorithm for experience? Yes, there is: Hiring people with experience.

Then, at some level, you wind up at the very far side of people who are late-career in many cases where the piece of paper that shows that they know what they’re doing is have you tried googling their name and looking at the Wikipedia article that spits out, how they built fundamental parts of a system like that. I think that certifications are one of those things that bias for early-career folks. And of course, partners when there are other business reasons to get it. But as people grow in seniority, I feel like the need for those begins to fall off. Do you agree? Disagree? You’re much closer to this industry in that aspect of it than I am.

Andrew: The more senior you are, and if you have big names under your resume there, no one’s going to care if you have certification, right? When I was looking to switch careers—I used to have a consultancy, and I was just tired of building another failed startup for somebody that was willing to pay me. And I’m like—I was not very nice about it. I was like, “Your startup’s not going to work out. You really shouldn’t be building this.” And they still give me the money and it would fail, and I’d move on to the next one. It was very frustrating.

So, closed up shop on that. And I said, “Okay, I got to reenter the market.” I don’t have a computer science degree, I don’t have big names on my resume, and Toronto is a very competitive market. And so I was feeling friction because people were not valuing my projects. I had, like, full-stack projects, I would show them.

And they said, “No, no. Just do these, like, CompSci algorithms and stuff like that.” And so I went, “Okay, well, I really don’t want to be doing that. I don’t want to spend all my time learning algorithms just so I can get a job to prove that I already have the knowledge I have.” And so I saw a big opportunity in cloud, and I thought certifications would be the proof to say, “I can do these things.”

And when I actually ended up going for the interviews, I didn’t even have certifications and I was getting those opportunities because the certifications helped me prove it, but nobody cared about the certifications, even then, and that was, like, 2017. But not to say, like, they didn’t help me, but it wasn’t the fact that people went, “Oh, you have a certification. We’ll get you this job.”

Corey: Yeah. When I’m talking to consulting clients, I’ve never once been asked, “Well, do you have the certifications?” Or, “Are you an AWS partner?” In my case, no, neither of those things. The reason that we know what we’re doing is because we’ve done this before. It’s the expertise approach.

I question whether that would still be true if we were saying, “Oh, yeah, and we’re going to drop a dozen engineers on who are going to build things out of your environment.” “Well, are they certified?” is a logical question to ask when you’re bringing in an external service provider? Or is this just a bunch of people you found somewhere on Upwork or whatnot, and you’re throwing them at it with no quality control? Like, what is the baseline level experience? That’s a fair question. People are putting big levels of trust when they bring people in.

Andrew: I mean, I could see that as a factor of some clients caring, just because like, when I used to work in startups, I knew customers where it’s like their second startup, and they’re flush with a lot of money, and they’re deciding who they want to partner with, and they’re literally looking at what level of SSL certificate they purchased, right? Like now, obviously, they’re all free and they’re very easy to get to get; there was one point where you had different tiers—as if you would know—and they would look and they would say—

Corey: Extended validation certs attend your browser bar green. Remember those?

Andrew: Right. Yeah, yeah, yeah. It was just like that, and they’re like, “We should partner with them because they were able to afford that and we know, like…” whatever, whatever, right? So, you know, there is that kind of thought process for people at an executive level. I’m not saying it’s widespread, but I’ve seen it.

When you talk to people that are in cloud consultancy, like solutions architects, they always tell me they’re driven to go get those professional certifications [unintelligible 00:22:19] their customers matter. I don’t know if the customers care or not, but they seem to think so. So, I don’t know if it’s just more driven by those people because it’s an expectation because everyone else has it, or it’s like a package of things, like, you know, like the green bar in the certifications, SOC 2 compliance, things like that, that kind of wrap it up and say, “Okay, as a package, this looks really good.” So, more of an expectation, but not necessarily matters, it’s just superficial; I’m not sure.

Corey: This episode is sponsored by our friends at Oracle HeatWave is a new high-performance accelerator for the Oracle MySQL Database Service. Although I insist on calling it “my squirrel.” While MySQL has long been the worlds most popular open source database, shifting from transacting to analytics required way too much overhead and, ya know, work. With HeatWave you can run your OLTP and OLAP, don’t ask me to ever say those acronyms again, workloads directly from your MySQL database and eliminate the time consuming data movement and integration work, while also performing 1100X faster than Amazon Aurora, and 2.5X faster than Amazon Redshift, at a third of the cost. My thanks again to Oracle Cloud for sponsoring this ridiculous nonsense.

Corey: You’ve been building out certifications for multiple cloud providers, so I’m curious to get your take on something that Forrest Brazeal, who’s now head of content over at Google Cloud, has been talking about lately, the idea that as an engineer is advised to learn more than one cloud provider; even if you have one as a primary, learning how another one works makes you a better engineer. Now, setting aside entirely the idea that well, yeah, if I worked at Google, I probably be saying something fairly similar.

Andrew: Yeah.

Corey: Do you think there’s validity to the idea that most people should be broad across multiple providers, or do you think specialization on one is the right path?

Andrew: Sure. Just to contextualize for our listeners, Google Cloud is highly, highly promoting multi-cloud workloads, and one of their flagship products is—well, they say it’s a flagship product—is Anthos. And they put a lot of money—I don’t know that was subsidized, but they put a lot of money in it because they really want to push multi-cloud, right? And so when we say Forrest works in Google Cloud, it should be no
surprise that he’s promoting it.

But I don’t work for Google, and I can tell you, like, learning multi-cloud is, like, way more valuable than just staying in one vertical. It just
opened my eyes. When I went from AWS to Azure, it was just like, “Oh, I’m missing out on so much in the industry.” And it really just made me such a more well-rounded person. And I went over to Google Cloud, and it was just like… because you’re learning the same thing in different variations, and then you’re also poly-filling for things that you will never touch.

Or like, I shouldn’t say you never touch, but you would never touch if you just stayed in that vertical when you’re learning. So, in the industry, Azure Active Directory is, like, widespread, but if you just stayed in your little AWS box, you’re not going to notice it on that learning path, right? And so a lot of times, I tell people, “Go get your CLF-C01 and then go get your AZ-900 or AZ-104.” Again, I don’t care if people go and sit the exams. I want them to go learn the content because it is a large eye-opener.

A lot of people are against multi-cloud from a learning perspective because say, it’s too much to learn all at the same time. But a lot of people I don’t think have actually gone across the cloud, right? So, they’re sitting from their chair, only staying in one vertical saying, “Well, you can’t learn them all at the same time.” And I’m going, “I see a way that you could teach them all at the same time.” And I might be the first person that will do it.

Corey: And the principles do convey as well. It’s, “Oh, well I know how SNS works on AWS, so I would never be able to understand how Google Pub/Sub works.” Those are functionally identical; I don’t know that is actually true. It’s just different to interface points and different guarantees, but fine. You at least understand the part that it plays.

I’ve built things out on Google Cloud somewhat recently, and for me, every time I do, it’s a refreshing eye-opener to oh, this is what developer experience in the cloud could be. And for a lot of customers, it is. But staying too far within the bounds of one ecosystem does lend itself to a
loss of perspective, if you’re not careful. I agree with that.

Andrew: Yeah. Well, I mean, just the paint more of a picture of differences, like, Google Cloud has a lot about digital transformation. They just updated their—I’m not happy that they changed it, but I’m fine that they did that, but they updated their Google Digital Cloud Leader Exam Guide this month, and it like is one hundred percent all about digital transformation. So, they love talking about digital transformation, and those kind of concepts there. They are really good at defining migration strategies, like, at a high level.

Over to Azure, they have their own cloud adoption framework, and it’s so detailed, in terms of, like, execution, where you go over to AWS and they have, like, the worst cloud adoption framework. It’s just the laziest thing I’ve ever seen produced in my life compared to out of all the providers in that space. I didn’t know about zero-trust model until I start using Azure because Azure has Active Directory, and you can do risk-based policy procedures over there. So, you know, like, if you don’t go over to these places, you’re not going to get covered other places, so you’re just going to be missing information till you get the job and, you know, that job has that information requiring you to know it.

Corey: I would say that for someone early career—and I don’t know where this falls on the list of career advice ranging from, “That is genius,” to, “Okay, Boomer,” but I would argue that figuring out what companies in your geographic area, or the companies that you have connections with what they’re using for a cloud provider, I would bias for learning one enough to get hired there and from there, letting what you learn next be dictated by the environment you find yourself in. Because especially larger companies, there’s always something that lives in a different provider. My default worst practice is multi-cloud. And I don’t say that because multi-cloud doesn’t exist, and I’m not saying it because it’s a bad idea, but this idea of one workload—to me—that runs across multiple providers is generally a challenge. What I see a lot more, done intelligently, is, “Okay, we’re going to use this provider for some things, this other provider for other things, and this third provider for yet more things.” And every company does that.

If not, there’s something very strange going on. Even Amazon uses—if not Office 365, at least exchange to run their email systems instead of Amazon WorkMail because—

Andrew: Yeah.

Corey: Let’s be serious. That tells me a lot. But I don’t generally find myself in a scenario where I want to build this application that is anything more than Hello World, where I want it to run seamlessly and flawlessly across two different cloud providers. That’s an awful lot of work that I struggle to identify significant value for most workloads.

Andrew: I don’t want to think about securing, like, multiple workloads, and that’s I think a lot of friction for a lot of companies are ingress-egress costs, which I’m sure you might have some knowledge on there about the ingress-egress costs across providers.

Corey: Oh, a little bit, yeah.

Andrew: A little bit, probably.

Corey: Oh, throwing data between clouds is always expensive.

Andrew: Sure. So, I mean, like, I call multi-cloud using multiple providers, but not in tandem. Cross-cloud is when you want to use something like Anthos or Azure Arc or something like that where you extend your data plane or control pla—whatever the plane is, whatever plane across all the providers. But you know, in practice, I don’t think many people are doing cross-cloud; they’re doing multi-cloud, like, “I use AWS to run my primary workloads, and then I use Microsoft Office Suite, and so we happen to use Azure Active Directory, or, you know, run particular VM
machines, like Windows machines for our accounting.” You know?

So, it’s a mixed bag, but I do think that using more than one thing is becoming more popular just because you want to use the best in breed no matter where you are. So like, I love BigQuery. BigQuery is amazing. So, like, I ingest a lot of our data from, you know, third-party services right into that. I could be doing that in Redshift, which is expensive; I could be doing that in Azure Synapse, which is also expensive. I mean, there’s a serverless thing. I don’t really get serverless. So, I think that, you know, people are doing multi-cloud.

Corey: Yeah. I would agree. I tend to do things like that myself, and whenever I see it generally makes sense. This is my general guidance. When I talk to individuals who say, “Well, we’re running multi-cloud like this.” And my response is, “Great. You’re probably right.”

Because I’m talking in the general sense, someone building something out on day one where they don’t know, like, “Everyone’s saying multi-cloud. Should I do that?” No, I don’t believe you should. Now, if your company has done that intentionally, rather than by accident, there’s almost certainly a reason and context that I do not have. “Well, we have to run our SaaS application in multiple cloud providers because that’s where our customers are.” “Yeah, you should probably do that.” But your marketing, your billing systems, your back-end reconciliation stuff generally does not live across all of those providers. It lives in one. That’s the sort of thing I’m talking about. I think we’re in violent agreement here.

Andrew: Oh, sure, yeah. I mean, Kubernetes obviously is becoming very popular because people believe that they’ll have a lot more mobility, Whereas when you use all the different managed—and I’m still learning Kubernetes myself from the next certification I have coming out, like, study course—but, you know, like, those managed services have all different kind of kinks that are completely different. And so, you know, it’s not going to be a smooth process. And you’re still leveraging, like, for key things like your database, you’re not going to be running that in Kubernetes Cluster. You’re going to be using a managed service.

And so, those have their own kind of expectations in terms of configuration. So, I don’t know, it’s tricky to say what to do, but I think that, you know, if you have a need for it, and you don’t have a security concern—like, usually it’s security or cost, right, for multi-cloud.

Corey: For me, at least, the lock-in has always been twofold that people don’t talk about. More—less lock-in than buy-in. One is the security model where IAM is super fraught and challenging and tricky, and trying to map a security model to multiple providers is super hard. Then on top of that, you also have the buy-in story of a bunch of engineers who are very good at one cloud provider, and that skill set is not in less demand now than it was a year ago. So okay, you’re going to start over and learn a new cloud provider is often something that a lot of engineers won’t want to countenance.

If your team is dead set against it, there’s going to be some friction there and there’s going to be a challenge. I mean, for me at least, to say that someone knows a cloud provider is not the naive approach of, “Oh yeah, they know how it works across the board.” They know how it breaks. For me, one of the most valuable reasons to run something on AWS is I know what a failure mode looks like, I know how it degrades, I know how to find out what’s going on when I see that degradation. That to me is a very hard barrier to overcome. Alternately, it’s entirely possible that I’m just old.

Andrew: Oh, I think we’re starting to see some wins all over the place in terms of being able to learn one thing and bring it other places, like OpenTelemetry, which I believe is a cloud-native Kubernetes… CNCF. I can’t remember what it stands for. It’s like Linux Foundation, but for cloud-native. And so OpenTelemetry is just a standardized way of handling your logs, metrics, and traces, right? And so maybe CloudWatch will be the 1.0 of observability in AWS, and then maybe OpenTelemetry will become more of the standard, right, and so maybe we might see more managed services like Prometheus and Grafa—well, obviously, AWS has a managed Prometheus, but other things like that. So, maybe some of those things will melt away. But yeah, it’s hard to say what approach to take.

Corey: Yeah, I’m wondering, on some level, whether what the things we’re talking about today, how well that’s going to map forward. Because the industry is constantly changing. The guidance I would give about should you be in cloud five years ago would have been a nuanced, “Mmm, depends. Maybe for yes, maybe for no. Here’s the story.” It’s a lot less hedge-y and a lot less edge case-y these days when I answer that question. So, I wonder in five years from now when we look back at this podcast episode, how well this discussion about what the future looks like, and certifications, and multi-cloud, how well that’s going to reflect?

Andrew: Well, when we look at, like, Kubernetes or Web3, we’re just seeing kind of like the standardized boilerplate way of doing a bunch of things, right, all over the place. This distributed way of, like, having this generic API across the board. And how well that will take, I have no idea, but we do see a large split between, like, serverless and cloud-natives. So, it’s like, what direction? Or we’ll just have both? Probably just have both, right?

Corey: [Like that 00:33:08]. I hope so. It’s been a wild industry ride, and I’m really curious to see what changes as we wind up continuing to grow. But we’ll see. That’s the nice thing about this is, worst case, if oh, turns out that we were wrong on this whole cloud thing, and everyone starts exodusing back to data centers, well, okay. That’s the nice thing about being a small company. It doesn’t take either of us that long to address the reality we see in the industry.

Andrew: Well, that or these cloud service providers are just going to get better at offering those services within carrier hotels, or data centers, or on your on-premise under your desk, right? So… I don’t know, we’ll see. It’s hard to say what the future will be, but I do believe that cloud is sticking around in one form or another. And it basically is, like, an essential skill or table stakes for anybody that’s in the industry. I mean, of course, not everywhere, but like, mostly, I would say. So.

Corey: Andrew, I want to thank you for taking the time to speak with me today. If people want to learn more about your opinions, how you view these things, et cetera. Where can they find you?

Andrew: You know, I think the best place to find me right now is Twitter. So, if you go to twitter.com/andrewbrown—all lowercase, no spaces, no underscores, no hyphens—you’ll find me there. I’m so surprised I was able to get that handle. It’s like the only place where I have my
handle.

Corey: And we will of course put links to that in the [show notes 00:34:25]. Thanks so much for taking the time to speak with me today. I really
appreciate it.

Andrew: Well, thanks for having me on the show.

Corey: Andrew Brown, co-founder and cloud instructor at ExamPro Training and so much more. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an angry comment telling me that I do not understand certifications at all because you’re an accountant, and certifications matter more in that industry.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Shir

Shir Tamari is the Head of Research of Wiz, the cloud security company. He is an experienced security and technology researcher specializing in vulnerability research and practical hacking. In the past, he served as a consultant to a variety of security companies in the fields of research, development and product.

About Sagi

Sagi Tzadik is a security researcher in the Wiz Research Team. Sagi specializes in research and exploitation of web applications vulnerabilities, as well as network security and protocols. He is also a Game-Hacking and Reverse-Engineering enthusiast.

About Nir

Nir Ohfeld is a security researcher from Israel. Nir currently does cloud-related security research at Wiz. Nir specializes in the exploitation of web applications, application security and in finding vulnerabilities in complex high-level systems.

Links:

  • Wiz: https://www.wiz.io
  • Cloud CVE Slack channel: https://cloud-cve-db.slack.com/join/shared_invite/zt-y38smqmo-V~d4hEr_stQErVCNx1OkMA
  • Wiz Blog: https://wiz.io/blog
  • Twitter: https://twitter.com/wiz_io

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Redis, the company behind the incredibly popular open source database that is not the bind DNS server. If you’re tired of managing open source Redis on your own, or you’re using one of the vanilla cloud caching services, these folks have you covered with the go to manage Redis service for global caching and primary database capabilities; Redis Enterprise. To learn more and deploy not only a cache but a single operational data platform for one Redis experience, visit redis.com/hero. Thats r-e-d-i-s.com/hero. And my thanks to my friends at Redis for sponsoring my ridiculous non-sense.

Corey: This episode is sponsored in part by our friends at Rising Cloud, which I hadn’t heard of before, but they’re doing something vaguely interesting here. They are using AI, which is usually where my eyes glaze over and I lose attention, but they’re using it to help developers be more efficient by reducing repetitive tasks. So, the idea being that you can run stateless things without having to worry about scaling, placement, et cetera, and the rest. They claim significant cost savings, and they’re able to wind up taking what you’re running as it is in AWS with no changes, and run it inside of their data centers that span multiple regions. I’m somewhat skeptical, but their customers seem to really like them, so that’s one of those areas where I really have a hard time being too snarky about it because when you solve a customer’s problem and they get out there in public and say, “We’re solving a problem,” it’s very hard to snark about that. Multus Medical, Construx.ai and Stax have seen significant results by using them. And it’s worth exploring. So, if you’re looking for a smarter, faster, cheaper alternative to EC2, Lambda, or batch, consider checking them out. Visit risingcloud.com/benefits. That’s risingcloud.com/benefits, and be sure to tell them that I said you because watching people wince when you mention my name is one of the guilty pleasures of listening to this podcast.

Corey: Welcome to Screaming in the Cloud, I’m Corey Quinn. One of the joyful parts of working with cloud computing is that you get to put a whole lot of things you don’t want to deal with onto the shoulders of the cloud provider you’re doing business with—or cloud providers as the case may be, if you fallen down the multi-cloud well. One of those things is often significant aspects of security. And that’s great, right, until it isn’t. Today, I’m joined by not one guest, but rather three coming to us from Wiz, which I originally started off believing was, oh, it’s a small cybersecurity research group. But they’re far more than that. Thank you for joining me, and could you please introduce yourself?

Shir: Yes, thank you, Corey. My name is Shir, Shir Tamari. I lead the security research team at Wiz. I working in the company for the past year. I’m working with these two nice teammates.

Nir: Hi, my name is Nir Ohfield,. I’m a security researcher at the Wiz research team. I’ve also been working for the Wiz research team for the last year. And yeah.

Sagi: I’m Sagi, Sagi Tzadik. I also work for the Wiz research team for the last six months.

Corey: I want to thank you for joining me. You folks really burst onto the scene earlier this year, when I suddenly started seeing your name come up an awful lot. And it brought me back to my childhood where there was an electronics store called Nobody Beats the Wiz. It was more or less a version of Fry’s on a different coast, and they went out of business and oh, good. We’re going back in time. And suddenly it felt like I was going back in time in a different light because you had a number of high profile vulnerabilities that you had discovered, specifically in the realm of Microsoft Azure. The two that leap to mind the most readily for me are ChaosDB and the OMIGOD exploits. There was a third as well, but why don’t you tell me, in your own words, what it is that you discovered and how that played out?

Shir: We, sort of, found the vulnerabilities in Microsoft Azure. We did report multiple vulnerabilities also in GCP, and AWS. We had multiple vulnerabilities in AWS [unintelligible 00:02:42] cross-account. It was a cross-account access to other tenants; it just was much less severe than the ChaosDB vulnerability that we will speak on more later. And a both we’ve present in Blackhat in Vegas in [unintelligible 00:02:56]. So, we do a lot of research. You mentioned that we have a third one. Which one did you refer to?

Corey: That’s a good question because you had the I want to say it was called as Azurescape, and you’re doing a fantastic job with branding a number of your different vulnerabilities, but there’s also, once you started reporting this, a lot of other research started coming out as well from other folks. And I confess, a lot of it sort of flowed together and been very hard to disambiguate, is this a systemic problem; is this, effectively, a whole bunch of people piling on now that their attention is being drawn somewhere; or something else? Because you’ve come out with an awful lot of research in a short period of time.

Shir: Yeah, we had a lot of good research in the past year. It’s a [unintelligible 00:03:36] mention Azurecape was actually found by a very good researcher in Palo Also. And… do you remember his name?

Sagi: No, I can’t recall his name is.

Corey: Yeah, they came out of unit 42 as I recall, their cybersecurity division. Every tech company out there seems to have some sort of security research division these days. What I think is, sort of, interesting is that to my understanding, you were founded, first and foremost, as a security company. You’re not doing this as an ancillary to selling something else like a firewall, or, effectively, you’re an ad comp—an ad tech company like Google, we you’re launching Project Zero. You are first and foremost aimed at this type of problem.

Shir: Yes. Wiz is not just a small research company. It’s actually pretty big company with over 200 employees. And the purpose of this product is a cloud security suite that provides [unintelligible 00:04:26] scanning capabilities in order to find risks in cloud environments. And the research team is a very small group. We are [unintelligible 00:04:35] researchers.

We have multiple responsibilities. Our first responsibility is to find risks in cloud environments: It could be misconfigurations, it could be vulnerabilities in libraries, in software, and we add those findings and the patterns we discover to the product in order to protect our customers, and to allow them for new risks. Our second responsibility is also to do a community research where we research everyone vulnerabilities in public products and cloud providers, and we share our findings with the cloud providers, then also with the community to make the cloud more secure.

Corey: I can’t shake the feeling that if there weren’t folks doing this sort of research and shining a light on what it is that the cloud providers are doing, if they were to discover these things at all, they would very quietly, effectively, fix it in the background and never breathe a word of it in public. I like the approach that you’re taking as far as dragging it, kicking and screaming, into the daylight, but I also have to imagine that probably doesn’t win you a whole lot of friends at the company that you’re focusing on at any given point in time. Because whenever you talk to a company about a security issue, it seems like the first thing they’re concerned about is, “Okay, how do we wind up spinning this or making sure that we minimize the reputational damage?” And then there’s a secondary reaction of, “Oh, and how do we protect our customers? But mostly, how do we avoid looking bad as a result?” And I feel like that’s an artifact of corporate culture these days. But it feels like the relationship has got to be somewhat interesting to navigate from your perspective.

Shir: So, once we found a vulnerability and we discuss it with the vendor, okay, first, I will mention that most cloud providers have a bug bounty program where they encourage researchers to find vulnerabilities and to discover new security threats. And all of them, as a public disclosure, [unintelligible 00:06:29] program will researchers are welcome and get safe harbor, you know, where the disclosure vulnerabilities. And I think it’s, like, common interest, both for customers, but for researchers, and the cloud providers to know about those vulnerabilities, to mitigate it down. And we do believe that sometimes cloud providors does resolve and mitigate vulnerabilities behind the scenes, and we know—we don’t know for sure, but—I don’t know about everything, but just by the vulnerabilities that we find, we assume that there is much more of them that we never heard about. And this is something that we believe needs to be changed in the industry.

Cloud providers should be more transparent, they should show more information about the result vulnerabilities. Definitely when a customer data was accessible, or where it was at risk, or at possible risk. And this is actually—it’s something that we actually trying to change in the industry. We have a community and, like, innovative community. It’s like an initiative that we try to collect, we opened a Slack channel called the Cloud CVE, and we try to invite as much people as we can that concern about cloud’s vulnerabilities, in order to make a change in the industry, and to assist cloud providers, or to convince cloud providers to be more transparent, to enumerate cloud vulnerabilities so they have an identifier just, like cloud CVE, like a CVE, and to make the cloud more protected and more transparent customers.

Corey: The thing that really took me aback by so much of what you found is that we’ve become relatively accustomed to a few patterns over the past 15 to 20 years. For example, we’re used to, “Oh, this piece of software you run on your desktop has a horrible flaw. Great.” Or this thing you run in your data center, same story; patch, patch, patch, patch patch. That’s great.

But there was always the sense that these were the sorts of things that were sort of normal, but the cloud providers were on top of things, where they were effectively living up to their side of the shared responsibility bargain. And that whenever you wound up getting breached, for whatever reason—like in the AWS world, where oh, you wound up losing a bunch of customer data because you had an open S3 bucket? Well, yeah, that’s not really something you can hang super effectively around the neck of the cloud provider, given that you’re the one that misconfigured that. But what was so striking about what you found with both of the vulnerabilities that we’re talking about today, the customer could have done everything absolutely correctly from the beginning and still had their data exposed. And that feels like it’s something relatively new in the world of cloud service providers.

Is this something that’s been going on for a while and we’re just now shining a light on it? Have I just missed a bunch of interesting news stories where the clouds have—“Oh, yeah, by the way, people, we periodically have to go in and drag people out of our cloud control plane because oops-a-doozy, someone got in there again with the squirrels,” or is this something that is new?

Shir: So, we do see an history other cases where probability [unintelligible 00:09:31] has disclosed vulnerabilities in the cloud infrastructure itself. There was only few, and usually, it was—the research was conducted by independent researchers. And I don’t think it had such an impact, like ChaosDB, which allowed [cross-system 00:09:51] access to databases of other customers, which was a huge case. And so if it wasn’t a big story, so most people will not hear about it. And also, independent researchers usually don’t have the back that we have here in Wiz.

We have a funding, we have the marketing division that help us to get coverage with reporters, who make sure to make—if it’s a big story, we make sure that other people will hear about it. And I believe that in most bug bounty programs where independent researchers find vulnerabilities, usually they more care about the bounty than the aftereffect of stopping the vulnerability, sharing it with the community. Usually also, independent [unintelligible 00:10:32] usually share the findings with the research community. And the research community is relatively small to the IT community. So, it is new, but it’s not that new.

There was some events back in history, [unintelligible 00:10:46] similar vulnerabilities. So, I think that one of the points here is that everyone makes a mistake. You can find bugs which affected mostly, as you mentioned previously, this software that you installed on your desktop has bugs and you need to patch it, but in the case of cloud providers, when they make mistakes, when they introduce bugs to the service, it affects all of their customers. And this is something that we should think about. So, mistakes that are being made by cloud providers have a lot of impact regarding their customers.

Corey: Yeah. It’s not a story of you misconfigured, your company’s SAN, so you’re the one that was responsible for a data breach. It’s suddenly, you’re misconfiguring everyone’s SAN simultaneously. It’s the sheer scale and scope of what it is that they’ve done. And—

Shir: Yeah, exactly.

Corey: —I’m definitely on board with that. But the stuff I’ve seen in the past, from cloud providers—AWS, primarily, since that is admittedly where I tend to focus most of my time and energy—has been privilege escalation style stuff, where, okay, if you assign some users at your company—or wherever—access to this managed IAM policy, well, they’ll have suddenly have access to things that go beyond the scope of that. And that’s not good, let’s be very clear on that, but it is a bit different between that and oh, by the way, suddenly, someone in another company that has no relationship established with you at all can suddenly rummage through your data that you’re storing in Cosmos DB, their managed database offering. That’s the thing to me that I think was the big head-turning aspect of this, not just for me, but for a number of folks I’ve spoken to, in financial services, in government, in a bunch of environments where data privacy is not optional in the same way that it is when, you know, you’re running a social media for pets app.

Nir: [laugh]. Yeah, but the thing is, that until the publication of ChaosDB, no one ever heard about the [unintelligible 00:12:40] data tampering in any cloud providers. Meaning maybe in six months, you can see a similar vulnerabilities in other cloud providers that maybe other security research groups find. So yeah, so Azure was maybe the first, but we don’t think they will be the last.

Shir: Yes. And also, when we do the community research, it is very important to us to take big targets. We enjoy the research. One day, the research will be challenging and we want to do something that it was new and great, so we always put a very big targets. To actually find vulnerability in the infrastructure of the cloud provider, it was very challenging for us.

When didn’t came ChaosDB by that; we actually found it by mistake. But now we think actively that this is our next goal is to find vulnerabilities in the infrastructure and not just vulnerabilities that affect only the—vulnerabilities within the account itself, like [unintelligible 00:13:32] or bad scoped policies that affects only one account.

Corey: That seems to be the transformative angle that you don’t see nearly as much in existing studies around vulnerabilities in this space. It’s always the, “Oh, no. We could have gotten breached by those people across the hallway from us in our company,” as opposed to folks on the other side of the planet. And that is, I guess, sort of the scary thing. What has also been interesting to me, and you obviously have more experience with this than I do, but I have a hard time envisioning that, for example, AWS, having a vulnerability like this and not immediately swinging into disaster firefighting mode, sending their security execs on a six month speaking tour to explain what happened, how it got there, all of the steps that they’re taking to remediate this, but Azure published a blog post explaining this in relatively minor detail: Here are the mitigations you need to take, and as far as I can tell, then they sort of washed their hands of the whole thing and have enthusiastically begun saying absolutely nothing since.

And that I have learned is sort of fairly typical for Microsoft, and has been for a while, where they just don’t talk about these things when it arises. Does that match your experience? Is this something that you find that is common when a large company winds up being, effectively, embarrassed about their security architecture, or is this something that is unique to Microsoft tends to approach these things?

Shir: I would say in general, we really like the Microsoft MSRC team. The group in Microsoft that’s responsible for handling vulnerabilities, and I think it’s like the security division inside Microsoft, MSRC. So, we have a really good relationship and we had really good time working with them. They’re real professionals, they take our findings very seriously. I can tell that in the ChaosDB incident, they didn’t plan to publish a blog post, and they did that after the story got a lot of attention.

So, I’m looking at a PR team, and I have no idea out there decide stuff and what is their strategy, but as I mentioned earlier, we believe that there is much more cloud vulnerabilities that we never heard of, and it should change; they should publish more.

Nir: It’s also worth mentioning that Microsoft acted really quick on this vulnerability and took it very seriously. They issued the fix in less than 48 hours. They were very transparent in the entire procedure, and we had multiple teams meeting with them. The entire experience was pretty positive with each of the vulnerability we’ve ever reported to Microsoft.

Sagi: So, it’s really nice working with the guys that are responsible for security, but regarding PR, I agree that they should have posted more information regarding this incident.

Corey: The thing that I found interesting about this, and I’ve seen aspects of it before, but never this strongly is, I was watching for, I guess, what I would call just general shittiness, for lack of a better term, from the other providers doing a happy dance of, “Aha, we’re better than you are,” and I saw none of that. Because when I started talking to people in some depth at this at other companies, the immediate response—not just AWS, to be clear—has been no, no, you have to understand, this is not good for anyone because this effectively winds up giving fuel to the slow-burning fire of folks who are pulling the, “See, I told you the cloud wasn’t secure.” And now the enterprise groundhog sees that shadow and we get six more years of building data centers instead of going to the cloud. So, there’s no one in the cloud space who’s happy with this kind of revelation and this type of vulnerability. My question for you is given that you are security researchers, which means you are generally cynical and pessimistic about almost everything technological, if you’re like most of the folks in that space that I’ve spent time with, is going with cloud the wrong answer? Should people be building their own data centers out? Should they continue to be going on this full cloud direction? I mean, what can they do if everything’s on fire and terrible all the time?

Shir: So, I think that there is a trade-off when you embrace the cloud. On one hand, you get the fastest deployment times, and a good scalability regarding your infrastructure, but on the other end, when there is a security vulnerability in the cloud provider, you are immediately affected. But it is worth mentioning that the security teams or the cloud providers are doing extremely good job. Most likely, they are going to patch the vulnerability faster than it would have been patched in on-premise environment. And it’s good that you have them working for you.

And once the vulnerability is mitigated—depends on the vulnerability but in the case of ChaosDB—when the vulnerability was mitigated on Microsoft’s end, and it was mitigated completely. No one else could have exploited after the mitigated it once. Yes, it’s also good to mention that the cloud provides organization and companies a lot of security features, [unintelligible 00:18:34] I want to say security features, I would say, it provides a lot of tooling that helps security. The option to have one interface, like one API to control all of my devices, to get visibility to all of my servers, to enforce policies very easily, it’s much more secure than on-premise environments, where there is usually a big mess, a lot of vendors.

Because the power was in the on-prem, the power was on the user, so the user had a lot of options. Usually used many types of software, many types of hardware, it’s really hard to mitigate the software vulnerability in on-prem environments. It’s really helped to get the visibility. And the cloud provides a lot of security, like, a good aspects, and in my opinion, moving to the cloud for most organization would be a more secure choice than remain on-premise, unless you have a very, very small on-prem environment.

Corey: This episode is sponsored by our friends at Oracle HeatWave is a new high-performance accelerator for the Oracle MySQL Database Service. Although I insist on calling it “my squirrel.” While MySQL has long been the worlds most popular open source database, shifting from transacting to analytics required way too much overhead and, ya know, work. With HeatWave you can run your OLTP and OLAP, don’t ask me to ever say those acronyms again, workloads directly from your MySQL database and eliminate the time consuming data movement and integration work, while also performing 1100X faster than Amazon Aurora, and 2.5X faster than Amazon Redshift, at a third of the cost. My thanks again to Oracle Cloud for sponsoring this ridiculous nonsense.

Corey: The challenge I keep running into is that—and this is sort of probably the worst of all possible reasons to go with cloud, but let’s face it, when us-east-1 recently took an outage and basically broke a decent swath of the internet, a lot of companies were impacted, but they didn’t see their names in the headlines; it was all about Amazon’s outage. There’s a certain value when a cloud provider takes an outage or a security breach, that the headlines screaming about it are about the provider, not about you and your company as a customer of that provider. Is that something that you’re seeing manifest across the industry? Is that an unhealthy way to think about it? Because it feels almost like it’s cheating in a way. It’s, “Yeah, we had a security problem, but so did the entire internet, so it’s okay.”

Nir: So, I think that if there would be evidence that these kind of vulnerabilities were exploited while disclosure, then you wouldn’t see headlines of companies, shouting in the headlines. But in the case of the us reporting the vulnerabilities prior to anyone exploiting them, results in nowhere a company showing up in the headlines. I think it’s a slightly different situation than an outage.

Shir: Yeah, but also, when one big provider have an outage or a breach, so usually, the customers will think it’s out of my responsibility. I mean, it’s bad; my data has been leaked, but what can I do? I think it’s very easy for most people to forgive companies [unintelligible 00:21:11]. I mean, you know what, it’s just not my area. So, maybe I’m not answer that into that. [laugh].

Corey: No, no, it’s very fair. The challenge I have, as a customer of all of these providers, to be honest, is that a lot of the ways that the breach investigations are worded of, “We have seen no evidence that this has been exploited.” Okay, that simultaneously covers the two very different use cases of, “We have pored through our exhaustive audit logs and validated that no one has done this particular thing in this particular way,” but it also covers the use case, “Of, hey, we learned we should probably be logging things, but we have no evidence that anything was exploited.” Having worked with these providers at scale, my gut impression is that they do in fact, have fairly detailed logs of who’s doing what and where. Would you agree with that assessment, or do you find that you tend to encounter logging and analysis gaps as you find these exploits?

Shir: We don’t really know. Usually when—I mean, ChaosDB scenario, we got access to a Jupyter Notebook. And from the Jupyter Notebook, we continued to another internal services. And we—nobody stopped us. Nobody—we expected an email, like—

Corey: “Whatcha doing over there, buddy?”

Shir: Yeah. “Please stop doing that, and we’re investigating you.” And we didn’t get any. And also, we don’t really know if they monitor it or not. I can tell from my technical background that logging so many environments, it’s hard.

And when you do decide to log all these events, you need to decide what to log. For example, if I have a database, a managed database, do I log all the queries that customers run? It’s too much. If I have an HTTP application—a managed HTTP application—do I save all the access logs, like all the requests? And if so, what will be the retention time? For how long?

We believe that it’s very challenging on the cloud provider side, but it just an assumption. And doing the discussion with Microsoft, the didn’t disclose any, like, scenarios they had with logging. They do mention that they’re [unintelligible 00:23:26] viewing the logs and searching to see if someone exploited this vulnerability before we disclosed it. Maybe someone discovered before we did. But they told us they didn’t find anything.

Corey: One last area I’d love to discuss with you before we call it an episode is that it’s easy to view Wiz through the lens of, “Oh, we just go out and find vulnerabilities here and there, and we make companies feel embarrassed—rightfully so—for the things that they do.” But a little digging shows that you’ve been around for a little over a year as a publicly known entity, and during that time, you’ve raised $600 million in funding, which is basically like what in the world is your pitch deck where you show up to investors and your slides are just, like, copies of their emails, and you read them to them?

[laugh]

I mean, on some level, it seems like that is a… as-, astounding amount of money to raise in a short period of time. But I’ve also done a little bit of digging, and to be clear, I do not believe that you have an extortion-based business model, which is a good thing. You’re building something very interesting that does in-depth analysis of cloud workloads, and I think it’s got an awful lot of promise. How does the vulnerability research that you do tie into that larger platform, other than, let’s be honest, some spectacularly effective marketing.

Sagi: Specifically in the ChaosDB vulnerability, we were actually not looking for a vulnerability in the cloud service providers. We were originally looking for common misconfigurations that our customers can make when they set up their Cosmos DB accounts, so that our product will be able to alert our customers regarding such misconfigurations. And then we went to the Azure portal and started to enable all of the features that Cosmos DB has to offer, and when we enabled enough features, we noticed some feature that could be vulnerable, and we started digging into it. And we ended up finding ChaosDB.

But our original work was to try and find misconfigurations that our customers can make in order to protect them and not to find a vulnerability in the [CSP 00:25:31]. This was just, like, a byproduct of this research.

Shir: Yes. There is, as I mentioned earlier, our main responsibility is to add a little security rist content to the product, to help customers to find new security risks in their environment. As you mentioned, like, the escalation possibilities within cloud accounts, and bad scoped policies, and many other security risks that are in the cloud area. And also, we are a very small team inside a big company, so most of the company, they are doing heavy [unintelligible 00:26:06] and talk with customers, they understand the risks, they understand the market, what the needs for tomorrow, and maybe we are well known for our vulnerabilities, but it just a very small part of the company.

Corey: On some level, it says wonderful things about your product, and also terrifying things from different perspectives of, “Oh, yeah, we found one of the worst cloud breaches in years by accident,” as opposed to actively going in trying to find the thing that has basically put you on the global map of awareness around these things. Because there a lot of security companies out there doing different things. In fact, go to RSA, and you’ll see basically 12 companies that just repeated over and over and over with different names and different brandings, and they’re all selling some kind of firewall. This is something actively different because everyone can tell beautiful pictures with slides and whatnot, and the corporate buzzwords. You’re one of those companies that actually did something meaningful, and it felt almost like a proof of concept. On some level, the fact that you weren’t actively looking for it is kind of an amazing testament for the product itself.

Shir: Yeah. We actually used the product in the beginning, in order to overview our own environment, and what is the most common services we use. In order—and we usually we mix this information with our product managers, know to understand what customers use and what products and services we need to research in order to bring value to the product.

Sagi: Yeah, so the reason we chose to research Cosmos DB was that, we found that a lot of our Azure customers are using Cosmos DB on their production environments, and we wanted to add mitigations for common misconfigurations to our product in order to protect our customers.

Nir: Yeah, the same goes with our other research, like OMIGOD, where we’ve seen that there is a excessive amount of [unintelligible 00:27:56] installations in an Azure environment, and it raised our [laugh] it raised our attention, and then found this vulnerability. It’s mostly, like, popularity-guided research. [laugh].

Shir: Yeah. And also [unintelligible 00:28:11] mention that maybe we find vulnerabilities by accident, but the service, we are doing vulnerability itself for the past ten years, and even more. So, we are very professional and this is what we do, and this is what we like to do. And we came skilled to the [crosstalk 00:28:25].

Corey: It really is neat to see, just because every other security tool that I’ve looked at in recent memory tells you the same stuff. It’s the same problem you see in the AWS billing space that I live in. Everyone says, “Oh, we can find these inactive instances that could be right-sized.” Great, because everyone’s dealing with the same data. It’s the security stuff is no different. “Hey, this S3 bucket is open.” Yes, it’s a public web server. Please stop waking me up at two in the morning about it. It’s there by design.

But it goes back and forth with the same stuff just presented differently. This is one of the first truly novel things I’ve seen in ages. If nothing else, you convince me to kick the tires on it, and see what kind of horrifying things I can learn about my own environments with it.

Shir: Yeah, you should. [laugh]. Let’s poke [unintelligible 00:29:13].

[laugh].

Corey: I want to thank you so much for taking the time to speak with me today. If people want to learn more about the research you’re up to and the things that you find interesting, where can they find you all?

Shir: Most of our publication—I mean, all of our publications are under the Wiz, which is wiz.io/blog, and people can read all of our research. Just today we are announcing a new one, so feel free to go and read there. And they also feel free to approach us on Twitter, the service, we have a Twitter account. We are open for, like, messages. Just send us a message.

Corey: And we will certainly put links to all of that in the [show notes 00:29:49]. Shir, Sagi, Nir, thank you so much for joining me today. I really appreciate your time.

Shir: Thank you.

Sagi: Thank you.

Nir: Thank you much.

Shir: It was very fun. Yeah.

Corey: This has been Screaming in the Cloud. I’m Cloud Economist Corey Quinn and thank you for listening. If you’ve enjoyed this podcast,
please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry insulting comment from someone else’s account.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Pete

I enjoy improving companies organizational structures, providing insight into building and growing autonomous high functioning, high performing technical teams. I'm fascinated by the dynamics of high performance, and take great pride in building and supporting those teams. I also enjoy the intricacies of Systems Architecture, Design, and Implementation work. I like to use modern tools to solve difficult technology problems. I'm most excited by Automation, Observability, Data Engineering.

I'm a product minded technologist. For the last 20 years working from Internet Service Providers and Hosting Companies to modern SaaS hosted on Cloud providers. I like to understand how people use the products that I build, and I like to build things that last a long time.

I consider product needs, business requirements, and technical capabilities when building products or planning new features. I work to understand the user and how and why they consume a service. All of our actions can impact many different ways, and I enjoy understanding how services, product teams, and business units work. I like to find ways to take one team's success and apply it more broadly, leveling up the entire business.

I like to get things done. I'm not too fond of unnecessary processes that slow down progress. I like iterative improvements, bringing new features into users' hands as quickly as possible, even if they are tiny changes. I want to share what I learn—both internal to a company and external to a broader community.

I enjoy the business side of technology as much as the technical side. I went back to school and received my MBA to understand the language of business. I enjoyed my product and finance classes the most. I like to understand the financial impact of product decisions. I don't like waste (in time or money), and I also believe premature optimization is the root of all evil.

Links:

  • Last Tweet in AWS: https://lasttweetinaws.com
  • Twitter: https://twitter.com/petecheslock
  • LinkedIn: https://www.linkedin.com/in/petecheslock/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part byLaunchDarkly. Take a look at what it takes to get your code into production. I’m going to just guess that it’s awful because it’s always awful. No one loves their deployment process. What if launching new features didn’t require you to do a full-on code and possibly infrastructure deploy? What if you could test on a small subset of users and then roll it back immediately if results aren’t what you expect? LaunchDarkly does exactly this. To learn more, visitlaunchdarkly.com and tell them Corey sent you, and watch for the wince.

Corey: This episode is sponsored in part by our friends at Redis, the company behind the incredibly popular open source database that is not the bind DNS server. If you’re tired of managing open source Redis on your own, or you’re using one of the vanilla cloud caching services, these folks have you covered with the go to manage Redis service for global caching and primary database capabilities; Redis Enterprise. To learn more and deploy not only a cache but a single operational data platform for one Redis experience, visit redis.com/hero. Thats r-e-d-i-s.com/hero. And my thanks to my friends at Redis for sponsoring my ridiculous non-sense.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I am joined—as is tradition, for a post re:Invent wrap up, a month or so later, once everything is time to settle—by my friend and yours, Pete Cheslock. Pete, how are you?

Pete: Hi, I’m doing fantastic. New year; new me. That’s what I’m going with.

Corey: That’s the problem. I keep hoping for that, but every time I turn around, it’s still me. And you know, honestly, I wouldn’t wish that on anyone.

Pete: Exactly. [laugh]. I wouldn’t wish you on me either. But somehow I keep coming back for this.

Corey: So, in two-thousand twenty—or twenty-twenty, as the children say—re:Invent was fully virtual. And that felt weird. Then re:Invent 2021 was a hybrid event which, let’s be serious here, is not really those things. They had a crappy online thing and then a differently crappy thing in person. But it didn’t feel real to me because you weren’t there.

That is part of the re:Invent tradition. There’s a midnight madness thing, there’s a keynote where they announce a bunch of nonsense, and then Pete and I go and have brunch on the last day of re:Invent and decompress, and more or less talk smack about everything that crosses our minds. And you weren’t there this year. I had to backfill you with Tim Banks. You know, the person that I backfield you with here at The Duckbill Group as a principal cloud economist.

Pete: You know, you got a great upgrade in hot takes, I feel like, with Tim.

Corey: And other ways, too, but it’s rude of me to say that to you directly. So yeah, his hot takes are spectacular. He was going to be doing this with me, except you cannot mess with tradition. You really can’t.

Pete: Yeah. I’m trying to think how many—is this third year? It’s at least three.

Corey: Third or fourth.

Pete: Yeah, it’s at least three. Yeah, it was, I don’t want to say I was sad to not be there because, with everything going on, it’s still weird out there. But I am always—I’m just that weird person who actually likes re:Invent, but not for I feel like the reasons people think. Again, I’m such an extroverted-type person, that it’s so great to have this, like, serendipity to re:Invent. The people that you run into and the conversations that you have, and prior—like in 2019, I think was a great example because that was the last one I had gone to—you know, having so many conversations so quickly because everyone is there, right? It’s like this magnet that attracts technologists, and venture capital, and product builders, and all this other stuff. And it’s all compressed into, like, you know, that five-day span, I think is the biggest part that makes so great.

Corey: The fear in people’s eyes when they see me. And it was fun; I had a pair of masks with me. One of them was a standard mask, and no one recognizes anyone because, masks, and the other was a printout of my ridiculous face, which was horrifyingly uncanny, but also made it very easy for people to identify me. And depending upon how social I was feeling, I would wear one or the other, and it worked flawlessly. That was worth doing. They really managed to thread the needle, as well, before Omicron hit, but after the horrors of last year. So, [unintelligible 00:03:00]—

Pete: It really—

Corey: —if it were going on right now, it would not be going on right now.

Pete: Yeah. I talk about really—yeah—really just hitting it timing-wise. Like, not that they could have planned for any of this, but like, as things were kind of not too crazy and before they got all crazy again, it feels like wow, like, you know, they really couldn’t have done the event at any other time. And it’s like, purely due to luck. I mean, absolute one hundred percent.

Corey: That’s the amazing power of frugality. Because the reason is then is it’s the week after Thanksgiving every year when everything is dirt cheap. And, you know, if there’s one thing that I one-point-seve—sorry, their stock’s in the toilet—a $1.6 trillion company is very concerned about, it is saving money at every opportunity.

Pete: Well, the one thing that was most curious about—so I was at the first re:Invent in-what—2012 I think it was, and there was—it was quaint, right?—there was 4000 people there, I want to say. It was in the thousands of people. Now granted, still a big conference, but it was in the Sands Convention Center. It was in that giant room, the same number of people, were you know, people’s booths were like tables, like, eight-by-ten tables, right? [laugh].

It had almost a DevOpsDays feel to it. And I was kind of curious if this one had any of those feelings. Like, did it evoke it being more quaint and personable, or was it just as soulless as it probably has been in recent years?

Corey: This was fairly soulless because they reduced the footprint of the event. They dropped from two expo halls down to one, they cut the number of venues, but they still had what felt like 20,000 people or something there. It was still crowded, it was still packed. And I’ve done some diligent follow-ups afterwards, and there have been very few cases of Covid that came out of it. I quarantined for a week in a hotel, so I don’t come back and kill my young kids for the wrong reasons.

And that went—that was sort of like the worst part of it on some level, where it’s like great. Now I could sit alone at a hotel and do some catch-up and all the rest, but all right I’d kind of like to go home. I’m not used to being on the road that much.

Pete: Yeah, I think we’re all a little bit out of practice. You know, I haven’t been on a plane in years. I mean, the travel I’ve done more recently has been in my car from point A to point B. Like, direct, you know, thing. Actually, a good friend of mine who’s not in technology at all had to travel for business, and, you know, he also has young kids who are under five, so he when he got back, he actually hid in a room in their house and quarantine himself in the room. But they—I thought, this is kind of funny—they never told the kids he was home. Because they knew that like—

Corey: So, they just thought the house was haunted?

Pete: [laugh].

Corey: Like, “Don’t go in the west wing,” sort of level of nonsense. That is kind of amazing.

Pete: Honestly, like, we were hanging out with the family because they’re our neighbors. And it was like, “Oh, yeah, like, he’s in the guest room right now.” Kids have no idea. [laugh]. I’m like, “Oh, my God.” I’m like, I can’t even imagine. Yeah.

Corey: So, let’s talk a little bit about the releases of re:Invent. And I’m going to lead up with something that may seem uncharitable, but I don’t think it necessarily is. There weren’t the usual torrent of new releases for ridiculous nonsense in the same way that there have been previously. There was no, this service talks to satellites in space. I mean, sure, there was some IoT stuff to manage fleets of cars, and giant piles of robots, and cool, I don’t have those particular problems; I’m trying to run a website over here.

So okay, great. There were enhancements to a number of different services that were in many cases appreciated, in other cases, irrelevant. Werner said in his keynote, that it was about focusing on primitives this year. And, “Why do we have so many services? It’s because you asked for it… as customers.”

Pete: [laugh]. Yeah, you asked for it.

Corey: What have you been asking for, Pete? Because I know what I’ve been asking for and it wasn’t that. [laugh].

Pete: It’s amazing to see a company continually say yes to everything, and somehow, despite their best efforts, be successful at doing it. No other company could do that. Imagine any other software technology business out there that just builds everything the customers ask for.
Like from a product management business standpoint, that is, like, rule 101 is, “Listen to your customers, but don’t say yes to everything.” Like,
you can’t do everything.

Corey: Most companies can’t navigate the transition between offering the same software in the Cloud and on a customer facility. So, it’s like, “Ooh, an on-prem version, I don’t know, that almost broke the company the last time we tried it.” Whereas you have Amazon whose product
strategy is, “Yes,” being able to put together a whole bunch of things. I also will challenge the assertion that it’s the primitives that customers want. They don’t want to build a data center out of popsicle sticks themselves. They want to get something that solves a problem.

And this has been a long-term realization for me. I used to work at Media Temple as a senior systems engineer running WordPress at extremely large scale. My websites now run on WordPress, and I have the good sense to pay WP Engine to handle it for me, instead of doing it myself because it’s not the most productive use of my time. I want things higher up the stack. I assure you I pay more to WP Engine than it would cost me to run these things myself from an infrastructure point of view, but not in terms of my time.

What I see sometimes as the worst of all worlds is that AWS is trying to charge for that value-added pricing without adding the value that goes along with it because you still got to build a lot of this stuff yourself. It’s still a very janky experience, you’re reduced to googling random blog posts to figure out how this thing is supposed to work, and the best documentation comes from externally. Whereas with a company that’s built around offering solutions like this, great. In the fullness of time, I really suspect that if this doesn’t change, their customers are going to just be those people who build solutions out of these things. And let those companies capture the up-the-stack margin. Which I have no problem with. But they do because Amazon is a company that lies awake at night actively worrying that someone, somewhere, who isn’t them might possibly be making money somehow.

Pete: I think MongoDB is a perfect example of—like, look at their stock price over the last whatever, years. Like, they, I feel like everyone called for the death of MongoDB every time Amazon came out with their new things, yet, they’re still a multi-billion dollar company because I can just—give me an API endpoint and you scale the database. There’s is—

Corey: Look at all the high-profile hires that Mongo was making out of AWS, and I can’t shake the feeling they’re sitting there going, “Yeah, who’s losing important things out of production now?” It’s, everyone is exodus-ing there. I did one of those ridiculous graphics of the naming all the people that went over there, and in—with the hurricane evacuation traffic picture, and there’s one car going the other way that I just labeled with, “Re:Invent sponsorship check,” because yeah, they have a top tier sponsorship and it was great. I’ve got to say I’ve been pretty down on MongoDB for a while, for a variety of excellent reasons based upon, more or less, how they treated customers who were in pain. And I’d mostly written it off.

I don’t do that anymore. Not because I inherently believe the technology has changed, though I’m told it has, but by the number of people who I deeply respect who are going over there and telling me, no, no, this is good. Congratulations. I have often said you cannot buy authenticity, and I don’t think that they are, but the people who are working there, I do not believe that these people are, “Yeah, well, you bought my opinion. You can buy their attention, not their opinion.” If someone changes their opinion, based upon where they work, I kind of question everything they’re telling me is, like, “Oh, you’re just here to sell something you don’t believe in? Welcome aboard.”

Pete: Right. Yeah, there’s an interview question I like to ask, which is, “What’s something that you used to believe in very strongly that you’ve more recently changed your mind on?” And out of politeness because usually throws people back a little bit, and they’re like, “Oh, wow. Like, let me think about that.” And I’m like, “Okay, while you think about that I want to give you mine.”

Which is in the past, my strongly held belief was we had to run everything ourselves. “You own your availability,” was the line. “No, I’m not buying Datadog. I can build my own metric stack just fine, thank you very much.” Like, “No, I’m not going to use these outsourced load balancers or databases because I need to own my availability.”

And what I realized is that all of those decisions lead to actually delivering and focusing on things that were not the core product. And so now, like, I’ve really flipped 180, that, if any—anything that you’re building that does not directly relate to the core product, i.e. How your business makes money, should one hundred percent be outsourced to an expert that is better than you. Mongo knows how to run Mongo better than
you.

Corey: “What does your company do?” “Oh, we handle expense reports.” “Oh, what are you working on this month?” “I’m building a load balancer.” It’s like that doesn’t add the value. Don’t do that.

Pete: Right. Exactly. And so it’s so interesting, I think, to hear Werner say that, you know, we’re just building primitives, and you asked for this. And I think that concept maybe would work years ago, when you had a lot of builders who needed tools, but I don’t think we have any, like, we don’t have as many builders as before. Like, I think we have people who need more complete solutions. And that’s probably why all these businesses are being super successful against Amazon.

Corey: I’m wondering if it comes down to a cloud economic story, specifically that my cloud bill is always going to be variable and it’s difficult to predict, whereas if I just use EC2 instances, and I build load balancers or whatnot, myself, well, yeah, it’s a lot more work, but I can predict accurately what my staff compensation costs are more effectively, that I can predict what a CapEx charge would be or what the AWS bill is going to be. I’m wondering if that might in some way shape it?

Pete: Well, I feel like the how people get better in managing their costs, right, you’ll eventually move to a world where, like, “Yep, okay, first, we turned off waste,” right? Like, step one is waste. Step two is, like, understanding your spend better to optimize but, like, step three, like, the galaxy brain meme of Amazon cost stuff is all, like, unit economics stuff, where trying to better understand the actual cost deliver an actual feature. And yeah, I think that actually gets really hard when you give—kind of spread your product across, like, a slew of services that have varying levels of costs, varying levels of tagging, so you can attribute it. Like, it’s really hard. Honestly, it’s pretty easy if I have 1000 EC2 servers with very specific tags, I can very easily figure out what it costs to deliver product. But if I have—

Corey: Yeah, if I have Corey build it, I know what Corey is going to cost, and I know how many servers he’s going to use. Great, if I have Pete it, Pete’s good at things, it’ll cut that server bill in half because he actually knows how to wind up being efficient with things. Okay, great. You can start calculating things out that way. I don’t think that’s an intentional choice that companies are making, but I feel like that might be a natural outgrowth of it.

Pete: Yeah. And there’s still I think a lot of the, like, old school mentality of, like, the, “Not invented here,” the, “We have to own our availability.” You can still own your availability by using these other vendors. And honestly, it’s really heartening to see so many companies realize that and realize that I don’t need to get everything from Amazon. And honestly, like, in some things, like I look at a cloud Amazon bill, and I think to myself, it would be easier if you just did everything from Amazon versus having these ten other vendors, but those ten other vendors are going to be a lot better at running the product that they build, right, that as a service, then you probably will be running it yourself. Or even
Amazon’s, like, you know, interpretation of that product.

Corey: A few other things that came out that I thought were interesting, at least the direction they’re going in. The changes to S3 intelligent tiering are great, with instant retrieval on Glacier. I feel like that honestly was—they talk a good story, but I feel like that was competitive response to Google offering the same thing. That smacks of a large company with its use case saying, “You got two choices here.” And they’re like, “Well, okay. Crap. We’re going to build it then.”

Or alternately, they’re looking at the changes that they’re making to intelligent tiering, they’re now shifting that to being the default that as far as recommendations go. There are a couple of drawbacks to it, but not many, and it’s getting easier now to not have the mental overhead of trying to figure out exactly what your lifecycle policies are. Yeah, there are some corner cases where, okay, if I adjust this just so, then I could save 10% on that monitoring fee or whatnot. Yeah, but look how much work that’s going to take you to curate and make sure that you’re not doing something silly. That feels like it is such an in the margins issue. It’s like, “How much data you’re storing?” “Four exabytes.” Okay, yeah. You probably want some people doing exactly that, but that’s not most of us.

Pete: Right. Well, there’s absolutely savings to be had. Like, if I had an exabyte of data on S3—which there are a lot of people who have that level of data—then it would make sense for me to have an engineering team whose sole purpose is purely an optimizing our data lifecycle for that data. Until a point, right? Until you’ve optimized the 80%, basically. You optimize the first 80, that’s probably, air-quote, “Easy.” The last 20 is going to be incredibly hard, maybe you never even do that.

But at lower levels of scale, I don’t think the economics actually work out to have a team managing your data lifecycle of S3. But the fact that now AWS can largely do it for you in the background—now, there’s so many things you have to think about and, like, you know, understand even what your data is there because, like, not all data is the same. And since S3 is basically like a big giant database you can query, you got to really think about some of that stuff. But honestly, what I—I don’t know if—I have no idea if this is even be worked on, but what I would love to see—you know, hashtag #AWSwishlist—is, now we have countless tiers of EBS volumes, EBS volumes that can be dynamically modified without touching, you know, the physical host. Meaning with an API call, you can change from the gp2 to gp3, or io whatever, right?

Corey: Or back again if it doesn’t pan out.

Pete: Or back again, right? And so for companies with large amounts of spend, you know, economics makes sense that you should have a team that is analyzing your volumes usage and modifying that daily, right? Like, you could modify that daily, and I don’t know if there’s anyone out there that’s actually doing it at that level. And they probably should. Like, if you got millions of dollars in EBS, like, there’s legit savings that you’re probably leaving on the table without doing that. But that’s what I’m waiting for Amazon to do for me, right? I want intelligent tiering for EBS because if you’re telling me I can API call and you’ll move my data and make that better, make that [crosstalk 00:17:46] better [crosstalk 00:17:47]—

Corey: Yeah it could be like their auto-scaling for DynamoDB, for example. Gives you the capacity you need 20 minutes after you needed it. But fine, whatever because if I can schedule stuff like that, great, I know what time of day, the runs are going to kick off that beat up the disks. I know when end-of-month reporting fires off. I know what my usage pattern is going to be, by and large.

Yeah, part of the problem too, is that I look at this stuff, and I get excited about it with the intelligent tiering… at The Duckbill Group we’ve got a few hundred S3 buckets lurking around. I’m thinking, “All right, I’ve got to go through and do some changes on this and implement all of that.” Our S3 bill’s something like 50 bucks a month or something ridiculous like that. It’s a no, that really isn’t a thing. Like, I have a screenshot bucket that I have an app installed—I think called Dropshare—that hooks up to anytime I drag—I hit a shortcut, I drag with the mouse to select whatever I want and boom, it’s up there and the URL is not copied to my clipboard, I can paste that wherever I want.

And I’m thinking like, yeah, there’s no cleanup on that. There’s no lifecycle policy that’s turning into anything. I should really go back and age some of it out and do the rest and start doing some lifecycle management. It—I’ve been using this thing for years and I think it’s now a whopping, what, 20 cents a month for that bucket. It’s—I just don’t—

Pete: [laugh].

Corey: —I just don’t care, other than voice in the back of my mind, “That’s an unbounded growth problem.” Cool. When it hits 20 bucks a month, then I’ll consider it. But until then I just don’t. It does not matter.

Pete: Yeah, I think yeah, scale changes everything. Start adding some zeros and percentages turned into meaningful numbers. And honestly, back on the EBS thing, the one thing that really changed my perspective of EBS, in general, is—especially coming from the early days, right? One terabyte volume, it was a hard drive in a thing. It was a virtual LUN on a SAN somewhere, probably.

Nowadays, and even, like, many years after those original EBS volumes, like all the limits you get in EBS, those are actually artificial limits, right? If you’re like, “My EBS volume is too slow,” it’s not because, like, the hard drive it’s on is too slow. That’s an artificial limit that is likely put in place due to your volume choice. And so, like, once you realize that in your head, then your concept of how you store data on EBS should change dramatically.

Corey: Oh, AWS had a blog post recently talking about, like, with io2 and the limits and everything, and there was architecture thinking, okay. “So, let’s say this is insufficient and the quarter-million IOPS a second that you’re able to get is not there.” And I’m sitting there thinking, “That is just ludicrous data volume and data interactivity model.” And it’s one of those, like, I’m sitting here trying to think about, like, I haven’t had to deal with a problem like that decade, just because it’s, “Huh. Turns out getting these one thing that’s super fast is kind of expensive.” If you paralyze it out, that’s usually the right answer, and that’s how the internet is mostly evolved. But there are use cases for which that doesn’t work, and I’m excited to see it. I don’t want to pay for it in my view, but it’s nice to see it.

Pete: Yeah, it’s kind of fun to go into the Amazon calculator and price out one of the, like, io2 volumes and, like, maxed out. It’s like, I don’t know, like $50,000 a month or a hun—like, it’s some just absolutely absurd number. But the beauty of it is that if you needed that value for an hour to run some intensive data processing task, you can have it for an hour and then just kill it when you’re done, right? Like, that is what is most impressive.

Corey: I copied 130 gigs of data to an EFS volume, which was—[unintelligible 00:21:05] EFS has gone from “This is a piece of junk,” to one of my favorite services. It really is, just because of its utility and different ways of doing things. I didn’t have the foresight, just use a second EFS volume for this. So, I was unzipping a whole bunch of small files onto it. Great.

It took a long time for me to go through it. All right, now that I’m done with that I want to clean all this up. My answer was to ultimately spin up a compute node and wind up running a whole bunch of—like, 400, simultaneous rm-rf on that long thing. And it was just, like, this feels foolish and dumb, but here we are. And I’m looking at the stats on it because the instance was—all right, at that point, the load average [on the instance 00:21:41] was like 200, or something like that, and the EFS volume was like, “Ohh, wow, you’re really churning on this. I’m now at, like, 5% of the limit.” Like, okay, great. It turns out I’m really bad at computers.

Pete: Yeah, well, that’s really the trick is, like, yeah, sure, you can have a quarter-million IOPS per second, but, like, what’s going to break before you even hit that limit? Probably many other things.

Corey: Oh, yeah. Like, feels like on some level if something gets to that point, it a misconfiguration somewhere. But honestly, that’s the thing I find weirdest about the world in which we live is that at a small-scale—if I have a bill in my $5 a month shitposting account, great. If I screw something up and cost myself a couple hundred bucks in misconfiguration it’s going to stand out. At large scale, it doesn’t matter if—you’re spending $50 million a year or $500 million a year on AWS and someone leaks your creds, and someone spins up a whole bunch of Bitcoin miners somewhere else, you’re going to see that on your bill until they’re mining basically all the Bitcoin. It just gets lost in the background.

Pete: I’m waiting for those—I’m actually waiting for the next level of them to get smarter because maybe you have, like, an aggressive tagging system and you’re monitoring for untagged instances, but the move here would be, first get the creds and query for, like, the most used tags and start applying those tags to your Bitcoin mining instances. My God, it’ll take—

Corey: Just clone a bunch of tags. Congratulations, you now have a second BI Elasticsearch cluster that you’re running yourself. Good work.

Pete: Yeah. Yeah, that people won’t find that until someone comes along after the fact that. Like, “Why do we have two have these things?” And you’re like—[laugh].

Corey: “Must be a DR thing.”

Pete: It’s maxed-out CPU. Yeah, exactly.

Corey: [laugh].

Pete: Oh, the terrible ideas—please, please, hackers don’t take are terrible ideas.

Corey: I had a, kind of, whole thing I did on Twitter years ago, talking about how I would wind up using the AWS Marketplace for an embezzlement scheme. Namely, I would just wind up spinning up something that had, like, a five-cent an hour charge or whatnot on just, like, basically rebadge the CentOS Community AMI or whatnot. Great. And then write a blog post, not attached to me, that explains how to do a thing that I’m going to be doing in production in a week or two anyway. Like, “How to build an auto-scaling group,” and reference that AMI.

Then if it ever comes out, like, “Wow, why are we having all these marketplace charges on this?” “I just followed the blog post like it said here.” And it’s like, “Oh, okay. You’re a dumbass. The end.”

That’s the way to do it. A month goes by and suddenly it came out that someone had done something similarly. They wound up rebadging these community things on the marketplace and charging big money for it, and I’m sitting there going like that was a joke. It wasn’t a how-to. But yeah, every time I make these jokes, I worry someone’s going to do it.

Pete: “Welcome to large-scale fraud with Corey Quinn.”

Corey: Oh, yeah, it’s fraud at scale is really the important thing here.

Corey: This episode is sponsored by our friends at Oracle HeatWave is a new high-performance accelerator for the Oracle MySQL Database Service. Although I insist on calling it “my squirrel.” While MySQL has long been the worlds most popular open source database, shifting from transacting to analytics required way too much overhead and, ya know, work. With HeatWave you can run your OLTP and OLAP, don’t ask me to ever say those acronyms again, workloads directly from your MySQL database and eliminate the time consuming data movement and integration work, while also performing 1100X faster than Amazon Aurora, and 2.5X faster than Amazon Redshift, at a third of the cost. My thanks again to Oracle Cloud for sponsoring this ridiculous nonsense.

Corey: I still remember a year ago now at re:Invent 2021 was it, or was it 2020? Whatever they came out with, I want to say it wasn’t gp3, or maybe it was, regardless, there was a new EBS volume type that came out that you were playing with to see how it worked and you experimented with it—

Pete: Oh, yes.

Corey: —and the next morning, you looked at the—I checked Slack and you’re like well, my experiments yesterday cost us $5,000. And at first, like, the—my response is instructive on this because, first, it was, “Oh, my God. What’s going to happen now?” And it’s like, first, hang on a
second.

First off, that seems suspect but assume it’s real. I assumed it was real at the outset. It’s “Oh, right. This is not my personal $5-a-month toybox account. We are a company; we can absolutely pay that.” Because it’s like, I could absolutely reach out, call it a favor. “I made a mistake, and I need a favor on the bill, please,” to AWS.

And I would never live it down, let’s be clear. For a $7,000 mistake, I would almost certainly eat it. As opposed to having to prostrate myself like that in front of Amazon. I’m like, no, no, no. I want one of those like—if it’s like, “Okay, you’re going to, like, set back the company roadmap by six months if you have to pay this. Do you want to do it?” Like, [groans] “Fine, I’ll eat some crow.”

But okay. And then followed immediately by, wow, if Pete of all people can mess this up, customers are going to be doomed here. We should figure out what happened. And I’m doing the math. Like, Pete, “What did you actually do?” And you’re sitting there and you’re saying, “Well, I had like a 20 gig volume that I did this.” And I’m doing the numbers, and it’s like—

Pete: Something’s wrong.

Corey: “How sure are you when you say ‘gigabyte,’ that you were—that actually means what you think it did? Like, were you off by a lot? Like, did you mean exabytes?” Like, what’s the deal here?

Pete: Like, multiple factors.

Corey: Yeah. How much—“How many IOPS did you give that thing, buddy?” And it turned out what happened was that when they launched this, they had mispriced it in the system by a factor of a million. So, it was fun. I think by the end of it, all of your experimentation was somewhere between five to seven cents. Which—

Pete: Yeah. It was a—

Corey: Which is why you don’t work here anymore because no one cost me seven cents of money to give to Amazon—

Pete: How dare you?

Corey: —on my watch. Get out.

Pete: How dare you, sir?

Corey: Exactly.

Pete: Yeah, that [laugh] was amazing to see, as someone who has done—definitely maid screw-ups that have cost real money—you know, S3 list requests are always a fun one at scale—but that one was supremely fun to see the—

Corey: That was a scary one because another one they’d done previously was they had messed up Lightsail pricing, where people would log in, and, like, “Okay, so what is my Lightsail instance going to cost?” And I swear to you, this is true, it was saying—this was back in 2017 or so—the answer was, like, “$4.3 billion.” Because when you see that you just start laughing because you know it’s a mistake. You know, that they’re not going to actually demand that you spend $4.3 billion for a single instance—unless it’s running SAP—and great.

It’s just, it’s a laugh. It’s clearly a mispriced, and it’s clearly a bug that’s going to get—it’s going to get fixed. I just spun up this new EBS volume that no one fully understands yet and it cost me thousands of dollars. That’s the sort of thing that no, no, I could actually see that happening. There are instances now that cost something like 100 bucks an hour or whatnot to run. I can see spinning up the wrong thing by mistake and getting bitten by it. There’s a bunch of fun configuration mistakes you can make that will, “Hee, hee, hee. Why can I see that bill spike from orbit?” And that’s the scary thing.

Pete: Well, it’s the original CI and CD problem of the per-hour billing, right? That was super common of, like, yeah, like, an i3, you know, 16XL server is pretty cheap per hour, but if you’re charged per hour and you spin up a bunch for five minutes. Like, it—you will be shocked [laugh] by what you see there. So—

Corey: Yeah. Mistakes will show. And I get it. It’s also people as individuals are very different psychologically than companies are. With companies it’s one of those, “Great we’re optimizing to bring in more revenue and we don’t really care about saving money at all costs.”

Whereas people generally have something that looks a lot like a fixed income in the form of a salary or whatnot, so it’s it is easier for us to cut spend than it is for us to go out and make more money. Like, I don’t want to get a second job, or pitch my boss on stuff, and yeah. So, all and all, routing out the rest of what happened at re:Invent, they—this is the problem is that they have a bunch of minor things like SageMaker Inference Recommender. Yeah, I don’t care. Anything—

Pete: [laugh].

Corey: —[crosstalk 00:28:47] SageMaker I mostly tend to ignore, for safety. I did like the way they described Amplify Studio because they made it sound like a WYSIWYG drag and drop, build a React app. It’s not it. It basically—you can do that in Figma and then it can hook it up to some things in some cases. It’s not what I want it to be, which is Honeycode, except good. But we’ll get there some year. Maybe.

Pete: There’s a lot of stuff that was—you know, it’s the classic, like, preview, which sure, like, from a product standpoint, it’s great. You know, they have a level of scale where they can say, “Here’s this thing we’re building,” which could be just a twinkle in a product managers, call it preview, and get thousands of people who would be happy to test it out and give you feedback, and it’s a, it’s great that you have that capability. But I often look at so much stuff and, like, that’s really cool, but, like, can I, can I have it now? Right? Like—or you can’t even get into the preview plan, even though, like, you have that specific problem. And it’s largely just because either, like, your scale isn’t big enough, or you don’t have a good enough relationship with your account manager, or I don’t know, countless other reasons.

Corey: The thing that really throws me, too, is the pre-announcements that come a year or so in advance, like, the Outpost smaller ones are finally available, but it feels like when they do too many pre-announcements or no big marquee service announcements, as much as they talk about, “We’re getting back to fundamentals,” no, you have a bunch of teams that blew the deadline. That’s really what it is; let’s not call it anything else. Another one that I think is causing trouble for folks—I’m fortunate in that I don’t do much work with Oracle databases, or Microsoft SQL databases—but they extended RDS Custom to Microsoft SQL at the [unintelligible 00:30:27] SQL server at re:Invent this year, which means this comes down to things I actually use, we’re going to have a problem because historically, the lesson has always been if I want to run my own databases and tweak everything, I do it on top of an EC2 instance. If I want to managed database, relational database service, great, I use RDS. RDS Custom basically gives you root into the RDS instance. Which means among other things, yes, you can now use RDS to run containers.

But it lets you do a lot of things that are right in between. So, how do you position this? When should I use RDS Custom? Can you give me an easy answer to that question? And they used a lot of words to say, no, they cannot. It’s basically completely blowing apart the messaging and positioning of both of those services in some unfortunate ways. We’ll learn as we go.

Pete: Yeah. Honestly, it’s like why, like, why would I use this? Or how would I use this? And this is I think, fundamentally, what’s hard when you just say yes to everything. It’s like, they in many cases, I don’t think, like, I don’t want to say they don’t understand why they’re doing this, but if it’s not like there’s a visionary who’s like, this fits into this multi-year roadmap.

That roadmap is largely—if that roadmap is largely generated by the customers asking for it, then it’s not like, oh, we’re building towards this Northstar of RDS being whatever. You might say that, but your roadmap’s probably getting moved all over the place because, you know, this company that pays you a billion dollars a year is saying, “I would give you $2 billion a year for all of my Oracle databases, but I need this specific thing.” I can’t imagine a scenario that they would say, “Oh, well, we’re building towards this Northstar, and that’s not on the way there.” Right? They’d be like, “New Northstar. Another billion dollars, please.”

Corey: Yep. Probably the worst release of re:Invent, from my perspective, is RUM, Real User Monitoring, for CloudWatch. And I, to be clear, I wrote a shitposting Twitter threading client called Last Tweet in AWS. Go to lasttweetinaws.com. You can all use it. It’s free; I just built this for my own purposes. And I’ve instrumented it with RUM. Now, Real User Monitoring is something that a lot of monitoring vendors use, and also CloudWatch now. And what that is, is it embeds a listener into the JavaScript that runs on client load, and it winds up looking at what’s going on loading times, et cetera, so you can see when users are unhappy. I have no problem with this. Other than that, you know, liking users? What’s up with that?

Pete: Crazy.

Corey: But then, okay, now, what this does is unlike every other RUM tool out there, which charges per session, meaning I am going to be… doing a web page load, it charges per data item, which includes HTTP errors, or JavaScript errors, et cetera. Which means that if you have a high transaction volume site and suddenly your CDN takes a nap like Fastly did for an hour last year, suddenly your bill is stratospheric for this because errors abound and cascade, and you can have thousands of errors on a single page load for these things, and it is going to be visible from orbit, at least with a per session basis thing, when you start to go viral, you understand that, “Okay, this is probably going to cost me some more on these things, and oops, I guess I should write less compelling content.” Fine. This is one of those one misconfiguration away and you are wailing and gnashing teeth. Now, this is a new service. I believe that they will waive these surprise bills in the event that things like that happen. But it’s going to take a while and you’re going to be worrying the whole time if you’ve rolled this out naively. So it’s—

Pete: Well and—

Corey: —I just don’t like the pricing.

Pete: —how many people will actively avoid that service, right? And honestly, choose a competitor because the competitor could be—the competitor could be five times more expensive, right, on face value, but it’s the certainty of it. It’s the uncertainty of what Amazon will charge you. Like, no one wants a surprise bill. “Well, a vendor is saying that they’ll give us this contract for $10,000. I’m going to pay $10,000, even though RUM might be a fraction of that price.”

It’s honestly, a lot of these, like, product analytics tools and monitoring tools, you’ll often see they price be a, like, you know, MAU, Monthly Active User, you know, or some sort of user-based pricing, like, the number of people coming to your site. You know, and I feel like at least then, if you are trying to optimize for lots of users on your site, and more users means more revenue, then you know, if your spend is going up, but your revenue is also going up, that’s a win-win. But if it’s like someone—you know, your third-party vendor dies and you’re spewing out errors, or someone, you know, upgraded something and it spews out errors. That no one would normally see; that’s the thing. Like, unless you’re popping open that JavaScript console, you’re not seeing any of those errors, yet somehow it’s like directly impacting your bottom line? Like that doesn’t feel [crosstalk 00:35:06].

Corey: Well, there is something vaguely Machiavellian about that. Like, “How do I get my developers to care about errors on consoles?” Like, how about we make it extortionately expensive for them not to. It’s, “Oh, all right, then. Here we go.”

Pete: And then talk about now you’re in a scenario where you’re working on things that don’t directly impact the product. You’re basically just sweeping up the floor and then trying to remove errors that maybe don’t actually affect it and they’re not actually an error.

Corey: Yeah. I really do wonder what the right answer is going to be. We’ll find out. Again, we live, we learn. But it’s also, how long does it take a service that has bad pricing at launch, or an unfortunate story around it to outrun that reputation?

People are still scared of Glacier because of its original restore pricing, which was non-deterministic for any sensible human being, and in
some cases lead to I’m used to spending 20 to 30 bucks a month on this. Why was I just charged two grand?

Pete: Right.

Corey: Scare people like that, they don’t come back.

Pete: I’m trying to actually remember which service it is that basically gave you an estimate, right? Like, turn it on for a month, and it would give you an estimate of how much this was going to cost you when billing started.

Corey: It was either Detective or GuardDuty.

Pete: Yeah, it was—yeah, that’s exactly right. It was one of those two. And honestly, that was unbelievably refreshing to see. You know, like, listen, you have the data, Amazon. You know what this is going to cost me, so when I, like, don’t make me spend all this time to go and figure out the cost. If you have all this data already, just tell me, right?

And if I look at it and go, “Yeah, wow. Like, turning this on in my environment is going to cost me X dollars. Like, yeah, that’s a trade-off I want to make, I’ll spend that.” But you know, with some of the—and that—a little bit of a worry on some of the intelligent tiering on S3 is that the recommendation is likely going to be everything goes to intelligent tiering first, right? It’s the gp3 story. Put everything on gp3, then move it to the proper volume, move it to an sc or an st or an io. Like, gp3 is where you start. And I wonder if that’s going to be [crosstalk 00:37:08].

Corey: Except I went through a wizard yesterday to launch an EC2 instance and its default on the free tier gp2.

Pete: Yeah. Interesting.

Corey: Which does not thrill me. I also still don’t understand for the life of me why in some regions, the free tier is a t2 instance, when t3 is available.

Pete: They’re uh… my guess is that they’ve got some free t—they got a bunch of t2s lying around. [laugh].

Corey: Well, one of the most notable announcements at re:Invent that most people didn’t pay attention to is their ability now to run legacy instance types on top of Nitro, which really speaks to what’s going on behind the scenes of we can get rid of all that old hardware and emulate the old m1 on modern equipment. So, because—you can still have that legacy, ancient instance, but now you’re going—now we’re able to wind up greening our data centers, which is part of their big sustainability push, with their ‘Sustainability Pillar’ for the well-architected framework. They’re talking more about what the green choices in cloud are. Which is super handy, not just because of the economic impact because we could use this pretty directly to reverse engineer their various margins on a per-service or per-offering basis. Which I’m not sure they’re aware of yet, but oh, they’re going to be.

And that really winds up being a win for the planet, obviously, but also something that is—that I guess puts a little bit of choice on customers. The challenge I’ve got is, with my serverless stuff that I build out, if I spend—the Google search I make to figure out what the most economic, most sustainable way to do that is, is going to have a bigger carbon impact on the app itself. That seems to be something that is important at scale, but if you’re not at scale, it’s one of those, don’t worry about it. Because let’s face it, the cloud providers—all of them—are going to have a better sustainability story than you are running this in your own data centers, or on a Raspberry Pi that’s always plugged into the wall.

Pete: Yeah, I mean, you got to remember, Amazon builds their own power plants to power their data centers. Like, that’s the level they play, right? There, their economies of scale are so entirely—they’re so entirely different than anything that you could possibly even imagine. So, it’s something that, like, I’m sure people will want to choose for. But, you know, if I would honestly say, like, if we really cared about our computing costs and the carbon footprint of it, I would love to actually know the carbon footprint of all of the JavaScript trackers that when I go to various news sites, and it loads, you know, the whatever thousands of trackers and tracking the all over, like, what is the carbon impact of some of those choices that I actually could control, like, as a either a consumer or business person?

Corey: I really hope that it turns into something that makes a meaningful difference, and it’s not just greenwashing. But we’ll see. In the fullness of time, we’re going to figure that out. Oh, they’re also launching some mainframe stuff. They—like that’s great.

Pete: Yeah, those are still a thing.

Corey: I don’t deal with a lot of customers that are doing things with that in any meaningful sense. There is no AWS/400, so all right.

Pete: [laugh]. Yeah, I think honestly, like, I did talk to a friend of mine who’s in a big old enterprise and has a mainframe, and they’re actually replacing their mainframe with Lambda. Like they’re peeling off—which is, like, a great move—taking the monolith, right, and peeling off the individual components of what it can do into these discrete Lambda functions. Which I thought was really fascinating. Again, it’s a five-year-long journey to do something like that. And not everyone wants to wait five years, especially if their support’s about to run out for that giant box in the, you know, giant warehouse.

Corey: The thing that I also noticed—and this is probably the—I guess, one of the—talk about swing and a miss on pricing—they have a—what is it?—there’s a VPC IP Address Manager, which tracks the the IP addresses assigned to your VPCs that are allocated versus not, and it’s 20 cents a month per IP address. It’s like, “Okay. So, you’re competing against a Google Sheet or an Excel spreadsheet”—which is what people are using for these things now—“Only you’re making it extortionately expensive?”

Pete: What kind of value does that provide for 20—I mean, like, again—

Corey: I think Infoblox or someone like that offers it where they become more cost-effective as soon as you hit 500 IP addresses. And it’s just
—like, this is what I’m talking about. I know it does not cost AWS that kind of money to store an IP address. You can store that in a Route 53 TXT record for less money, for God’s sake. And that’s one of those, like, “Ah, we could extract some value pricing here.”

Like, I don’t know if it’s a good product or not. Given its pricing, I don’t give a shit because it’s going to be too expensive for anything beyond trivial usage. So, it’s a swing and a miss from that perspective. It’s just, looking at that, I laugh, and I don’t look at it again.

Pete: See I feel—

Corey: I’m not usually price sensitive. I want to be clear on that. It’s just, that is just Looney Tunes, clown shoes pricing.

Pete: Yeah. It’s honestly, like, in many cases, I think the thing that I have seen, you know, in the past few years is, in many cases, it can honestly feel like Amazon is nickel-and-diming their customers in so many ways. You know, the explosion of making it easy to create multiple Amazon accounts has a direct impact to waste in the cloud because there’s a lot of stuff you have to have her account. And the more accounts you have, those costs grow exponentially as you have these different places. Like, you kind of lose out on the economies of scale when you have a smaller number of accounts.

And yeah, it’s hard to optimize for that. Like, if you’re trying to reduce your spend, it’s challenging to say, “Well, by making a change here, we’ll save, you know, $10,000 in this account.” “That doesn’t seem like a lot when we’re spending millions.” “Well, hold on a second. You’ll save $10,000 per account, and you have 500 accounts,” or, “You have 1000 accounts,” or something like that.

Or almost cost avoidance of this cost is growing unbounded in all of your accounts. It’s tiny right now. So, like, now would be the time you want to do something with it. But like, again, for a lot of companies that have adopted the practice of endless Amazon accounts, they’ve almost gone, like, it’s the classic, like, you know, I’ve got 8000 GitHub repositories for my source code. Like, that feels just as bad as having
one GitHub repository for your repo. I don’t know what the balance is there, but anytime these different types of services come out, it feels like, “Oh, wow. Like, I’m going to get nickeled and dimed for it.”

Corey: This ties into the re:Post launch, which is a rebranding of their forums, where, okay, great, it was a little crufty and it need modernize, but it still ties your identity to an IAM account, or the root email address for an Amazon account, which is great. This is completely worthless because as soon as I change jobs, I lose my identity, my history, the rest, on this forum. I’m not using it. It shows that there’s a lack of awareness that everyone is going to have multiple accounts with which they interact, and that people are going to deal with the platform longer than any individual account will. It’s just a continual swing and a miss on things like that.

And it gets back to the billing question of, “Okay. When I spin up an account, do I want them to just continue billing me—because don’t turn this off; this is important—or do I want there to be a hard boundary where if you’re about to charge me, turn it off. Turn off the thing that’s about to cost me money.” And people hem and haw like this is an insurmountable problem, but I think the way to solve it is, let me specify that intent when I provision the account. Where it’s, “This is a production account for a bank. I really don’t want you turning it off.” Versus, “I’m a student learner who thinks that a Managed NAT Gateway might be a good thing. Yeah, I want you to turn off my demo Hello World app that will teach me what’s going on, rather than surprising me with a five-figure bill at the end of the month.”

Pete: Yeah. It shouldn’t be that hard. I mean, but again, I guess everything’s hard at scale.

Corey: Oh, yeah. Oh yeah.

Pete: But still, I feel like every time I log into Cost Explorer and I look at—and this is years it’s still not fixed. Not that it’s even possible to fix—but on the first day of the month, you look at Cost Explorer, and look at what Amazon is estimating your monthly bill is going to be. It’s like because of your, you know—

Corey: Your support fees, and your RI purchases, and savings plans purchases.

Pete: [laugh]. All those things happened, right? First of the month, and it’s like, yeah, “Your bill’s going to be $800,000 this year.” And it’s like, “Shouldn’t be, like, $1,000?” Like, you know, it’s the little things like that, that always—

Corey: The one-off charges, like, “Oh, your Route 53 zone,” and all the stuff that gets charged on a monthly cadence, which fine, whatever. I
mean, I’m okay with it, but it’s also the, like, be careful when that happen—I feel like there’s a way to make that user experience less jarring.

Pete: Yeah because that problem—I mean, in my scenario, companies that I’ve worked at, there’s been multiple times that a non-technical person will look at that data and go into immediate freakout mode, right? And that’s never something that you want to have happen because now that’s just adding a lot of stress and anxiety into a company that is—with inaccurate data. Like, the data—like, the answer you’re giving someone is just wrong. Perhaps you shouldn’t even give it to them if it’s that wrong. [laugh].

Corey: Yeah, I’m looking forward to seeing what happens this coming year. We’re already seeing promising stuff. They—give people a timeline on how long in advance these things record—late last night, AWS released a new console experience. When you log into the AWS console now, there’s a new beta thing. And I gave it some grief on Twitter because I’m still me, but like the direction it’s going. It lets you customize your view with widgets and whatnot.

And until they start selling widgets on marketplace or having sponsored widgets, you can’t remove I like it, which is no guarantee at some point. But it shows things like, I can move the cost stuff, I can move the outage stuff up around, I can have the things that are going on in my account—but who I am means I can shift this around. If I’m a finance manager, cool. I can remove all the stuff that’s like, “Hey, you want to get started spinning up an EC2 instance?” “Absolutely not. Do I want to get told, like, how to get certified? Probably not. Do I want to know what the current bill is and whether—and my list of favorites that I’ve pinned, whatever services there? Yeah, absolutely do.” This is starting to get there.

Pete: Yeah, I wonder if it really is a way to start almost hedging on organizations having a wider group of people accessing AWS. I mean, in previous companies, I absolutely gave access to the console for tools like QuickSight, for tools like Athena, for the DataBrew stuff, the Glue DataBrew. Giving, you know, non-technical people access to be able to do these, like, you know, UI ETL tasks, you know, a wider group of a company is getting access into Amazon. So, I think anything that Amazon does to improve that experience for, you know, the non-SREs, like the people who would traditionally log in, like, that is an investment definitely worth making.

Corey: “Well, what could non-engineering types possibly be doing in the AWS console?” “I don’t know, jackhole, maybe paying the bill? Just a thought here.” It’s the, there are people who look at these things from a variety of different places, and you have such sprawl in the AWS world that there are different personas by a landslide. If I’m building Twitter for Pets, you probably don’t want to be pitching your mainframe migration services to me the same way that you would if I were a 200-year-old insurance company.

Pete: Yeah, exactly. And the number of those products are going to grow, the number of personas are going to grow, and, yeah, they’ll have to do something that they want to actually, you know, maintain that experience so that every person can have, kind of, the experience that they want, and not be distracted, you know? “Oh, what’s this? Let me go test this out.” And it’s like, you know, one-time charge for $10,000 because, like, that’s how it’s charged. You know, that’s not an experience that people like.

Corey: No. They really don’t. Pete, I want to thank you for spending the time to chat with me again, as is our tradition. I’m hoping we can do it in person this year, when we go at the end of 2022, to re:Invent again. Or that no one goes in person. But this hybrid nonsense is for the birds.

Pete: Yeah. I very much would love to get back to another one, and yeah, like, I think there could be an interesting kind of merging here of our annual re:Invent recap slash live brunch, you know, stream you know, hot takes after a long week. [laugh].

Corey: Oh, yeah. The real way that you know that it’s a good joke is when one of us says something, the other one sprays scrambled eggs out of their nose. Yeah, that’s the way to do it.

Pete: Exactly. Exactly.

Corey: Pete, thank you so much. If people want to learn more about what you’re up to—hopefully, you know, come back. We miss you, but you’re unaffiliated, you’re a startup advisor. Where can people find you to learn more, if they for some unforgivable reason don’t know who or what a Pete Cheslock is?

Pete: Yeah. I think the easiest place to find me is always on Twitter. I’m just at @petecheslock. My DMs are always open and I’m always down to expand my network and chat with folks.

And yeah, right, now, I’m just, as I jokingly say, professionally unaffiliated. I do some startup advisory work and have been largely just kind of—honestly checking out the state of the economy. Like, there’s a lot of really interesting companies out there, and some interesting problems to solve. And, you know, trying to spend some of my time learning more about what companies are up to nowadays. So yeah, if you got some interesting problems, you know, you can follow my Twitter or go to LinkedIn if you want some great, you know, business hot takes about, you know, shitposting basically.

Corey: Same thing. Pete, thanks so much for joining me, I appreciate it.

Pete: Thanks for having me.

Corey: Pete Cheslock, startup advisor, professionally unaffiliated, and recurring re:Invent analyst pal of mine. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry comment calling me a jackass because do I know how long it took you personally to price CloudWatch RUM?

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Maciej

Maciej Winnicki is a serverless enthusiast with over 6 years of experience in writing software with no servers whatsoever. Serverless Engineer at Stedi, Cloudash Founder, ex-Engineering Manager, and one of the early employees at Serverless Inc.

Links:

  • Cloudash: https://cloudash.dev
  • Maciej Winnicki Twitter: https://twitter.com/mthenw
  • Tomasz Łakomy Twitter: https://twitter.com/tlakomy
  • Cloudash email: hello@cloudash.dev

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part byLaunchDarkly. Take a look at what it takes to get your code into production. I’m going to just guess that it’s awful because it’s always awful. No one loves their deployment process. What if launching new features didn’t require you to do a full-on code and possibly infrastructure deploy? What if you could test on a small subset of users and then roll it back immediately if results aren’t what you expect? LaunchDarkly does exactly this. To learn more, visitlaunchdarkly.com and tell them Corey sent you, and watch for the wince.

Corey: This episode is sponsored in part by our friends at Rising Cloud, which I hadn’t heard of before, but they’re doing something vaguely interesting here. They are using AI, which is usually where my eyes glaze over and I lose attention, but they’re using it to help developers be more efficient by reducing repetitive tasks. So, the idea being that you can run stateless things without having to worry about scaling, placement, et cetera, and the rest. They claim significant cost savings, and they’re able to wind up taking what you’re running as it is in AWS with no changes, and run it inside of their data centers that span multiple regions. I’m somewhat skeptical, but their customers seem to really like them, so that’s one of those areas where I really have a hard time being too snarky about it because when you solve a customer’s problem and they get out there in public and say, “We’re solving a problem,” it’s very hard to snark about that. Multus Medical, Construx.ai and Stax have seen significant results by using them. And it’s worth exploring. So, if you’re looking for a smarter, faster, cheaper alternative to EC2, Lambda, or batch, consider checking them out. Visit risingcloud.com/benefits. That’s risingcloud.com/benefits, and be sure to tell them that I said you because watching people wince when you mention my name is one of the guilty pleasures of listening to this podcast.

Corey: Welcome to Screaming in the Cloud. I’m Cloud Economist Corey Quinn. And my guest today is Maciej Winnicki, who is the founder of Cloudash. Now, before I dive into the intricacies of what that is, I’m going to just stake out a position that one of the biggest painful parts of working with AWS in any meaningful sense, particularly in a serverless microservices way, is figuring out what the hell’s going on in the environment. There’s a bunch of tools offered to do this and they’re all—yeee, they aspire to mediocrity. Maciej, thank you for joining me today.

Corey: Welcome to Screaming in the Cloud. I’m Cloud Economist Corey Quinn. And my guest today is Maciej Winnicki, who is the founder of Cloudash. Now, before I dive into the intricacies of what that is, I’m going to just stake out a position that one of the biggest painful parts of working with AWS in any meaningful sense, particularly in a serverless microservices way, is figuring out what the hell’s going on in the environment. There’s a bunch of tools offered to do this and they’re all—yeee, they aspire to mediocrity. Maciej, thank you for joining me today.

Maciej: Thank you for having me.

Corey: So, I turned out to have accidentally blown up Cloudash, sort of before you were really ready for the attention. You, I think, tweeted about it or put it on Hacker News or something; I stumbled over it because it turns out that anything that vaguely touches cloud winds up in my filters because of awesome technology, and personality defects on my part. And I tweeted about it as I set it up and got the thing running, and apparently this led to a surge of attention on this thing that you’ve built. So, let me start off with an apology. Oops, I didn’t realize it was supposed to be a quiet launch.

Maciej: I actually thank you for that. Like, that was great. And we get a lot of attention from your tweet thread, actually because at the end, that was the most critical part. At the end of the twitter, you wrote that you’re staying as a customer, so we have it on our website and this is perfect. But actually, as you said, that’s correct.

Our marketing strategy for releasing Cloudash was to post it on LinkedIn. I know this is not, kind of, the best strategy, but that was our plan. Like, it was like, hey, like, me and my friend, Tomasz, who’s also working on Cloudash, we thought like, let’s just post it on LinkedIn and we’ll see how it goes. And accidentally, I’m receiving a notification from Twitter, “Hey, Corey started tweeting about it.” And I was like, “Oh, my God, I’m having a heart attack.” But then I read the, you know—

Corey: Oops.

Maciej: [laugh]. Yeah. I read the, kind of, conclusion, and I was super happy. And again, thank you for that because this is actually when Cloudash kind of started rolling as a product and as a, kind of, business. So yeah, that was great.

Corey: To give a little backstory and context here is, I write a whole bunch of serverless nonsense. I build API’s Gateway, I hook them up to Lambda’s Function, and then it sort of kind of works. Ish. From there, okay, I would try and track down what was going on because in a microservices land, everything becomes a murder mystery; you’re trying to figure out what’s broken, and things have exploded. And I became a paying customer of IOpipe. And then New Relic bought them. Well, crap.

Then I became a paying customer of Epsagon. And they got acquired by Cisco, at which point I immediately congratulated the founders, who I know on a social basis, and then closed my account because I wanted to get out before Cisco ruins it because, Cisco. Then it was, what am I
going to use next? And right around that time is when I stumbled across Cloudash. And it takes a different approach than any other entity in the space that I’ve seen because you are a native Mac desktop app. I believe your Mac only, but you seem to be Electron, so I could be way
off base on that.

Maciej: So, we’re Linux as well right now and soon we’ll be Windows as well. But yeah, so, right now is Mac OS and Linux. Yeah, that’s correct.
So, our approach is a little bit different.

So, let me start by saying what’s Cloudash? Like, Cloudash is a desktop app for, kind of, monitoring and troubleshooting serverless architectures services, like, serverless stuff in general. And the approach that we took is a little bit different because we are not web-based, we’re desktop-based. And there’s a couple of advantages of that approach. The first one is that, like, you don’t need to share your data with us because we’re not, kind of, downloading your metrics and logs to our back end and to process them, et cetera, et cetera. We are just using the credentials, the AWS profiles that you have defined on your computer, so nothing goes out of your AWS account.

And I think this is, like, considering, like, from the security perspective, this is very crucial. You don’t need to create a role that you give us access to or anything like that. You just use the stuff that you have on your desktop, and everything stays on your AWS account. So, nothing—we don’t download it, we don’t process it, we don’t do anything from that. And that’s one approach—well, that’s the one advantage. The other advantage is, like, kind of, onboarding, as I kind of mentioned because we’re using the AWS profiles that you have defined in your computer.

Corey: Well, you’re doing significantly more than that because I have a lot of different accounts configured different ways, and when I go to one of them that uses SSO, it automatically fires me off to the SSO login page if I haven’t logged in that day for a 12 hour session—

Maciej: Yes.

Corey: —for things that have credentials stored locally, it uses those; and for things that are using role-chaining to use assuming roles from the things I have credentials for, and the things that I just do role assumption in, and it works flawlessly. It just works the way that most of my command-line tools do. I’ve never seen a desktop app that does this.

Maciej: Yeah. So, we put a lot of effort into making sure that this works great because we know that, like, no one will use Cloudash if there’s—like, not no one, but like, we’re targeting, like, serverless teams, maybe, in enterprise companies, or serverless teams working on some startups. And in most cases, those teams or those engineers, they use SSO, or at least MFA, right? So, we have it covered. And as you said, like, it should be the onboarding part is really easy because you just pick your AWS profile, you just pick region, and just pick, right now, a CloudFormation stack because we get the information about your service based on CloudFormation stack. So yeah, we put a lot of effort in making sure that this works without any issues.

Corey: There are some challenges to it that I saw while setting it up, and that’s also sort of the nature of the fact you are, in fact, integrating with CloudWatch. For example, it’s region specific. Well, what if I want to have an app that’s multi-region? Well, you’re going to have a bad time because doing [laugh] anything multi-region in AWS means you’re going to have a bad time that gets particularly obnoxious and EC2 get to when you’re doing something like Lambda@Edge, where, oh, where are the logs live; that’s going to be in a CloudFront distribution in whatever region it winds up being accessed from. So, it comes down to what distribution endpoint or point of presence did that particular request go through, and it becomes this giant game of whack-a-mole. It’s frustrating, and it’s obnoxious, and it’s also in no way your fault.

Maciej: Yeah, I mean, we are at the beginning. Right now, it’s the most straightforward, kind of pe—how people think about stacks of serverless. They’re think in terms of regions because I think for us, regions, or replicated stacks, or things like that are not really popular yet. Maybe they will become—like, this is how AWS works as a whole, so it’s not surprising that we’re kind of following this path. I think my point is that our main goal, the ultimate goal, is to make monitoring, as I said, the troubleshooting serverless app as simple as possible.

So, once we will hear from our customers, from our users that, “Hey, we would like to get a little bit better experience around regions,” we will definitely implement that because why not, right? And I think the whole point of Cloudash—and maybe we can go more deep into that later—is that we want to bring context into your metrics and logs. If you’re seeing a, for example, X-Ray trace ID in your logs, you should be able with one click just see that the trace. It’s not yet implemented in Cloudash, but we are having it in the backlog. But my point is that, like, there should be some journey when you’re debugging stuff, and you shouldn’t be just, like, left alone having, like, 20 tabs, Cloudash tabs open and trying to figure out where I was—like, where’s the Lambda? Where’s the API Gateway logs? Where are the CloudFront logs? And how I can kind of connect all of that? Because that’s—it’s an issue right now.

Corey: Even what you’ve done so far is incredibly helpful compared to the baseline experience that folks will often have, where I can define a service that is comprised of a number of different functions—I have one set up right now that has seven functions in it—I grab any one of those things, and I can set how far the lookback is, when I look at that function, ranging from 5 minutes to 30 days. And it shows me at the top the metrics of invocations, the duration that the function runs for, and the number of errors. And then, in the same pane down below it, it shows the CloudWatch logs. So, “Oh, okay, great. I can drag and zoom into a specific timeframe, and I see just the things inside of that.”

And I know this sounds like well, what’s the hard part here? Yeah, except nothing else does it in an easy-to-use, discoverable way that just sort of hangs out here. Honestly, the biggest win for me is that I don’t have to log in to the browser, navigate through some ridiculous other thing to track down what I’m talking about. It hangs out on my desktop all the time, and whether it’s open or not, whenever I fire it up, it just works, basically, and I don’t have to think about it. It reduces the friction from, “This thing is broken,” to, “Let me see what the logs say.”

Very often I can go from not having it open at all to staring at the logs and having to wait a minute because there’s some latency before the event happens and it hits CloudWatch logs itself. I’m pretty impressed with it, and I’ve been keeping an eye on what this thing is costing me. It is effectively nothing in terms of CloudWatch retrieval charges. Because it’s not sitting there sucking all this data up all the time, for everything that’s running. Like, we’ve all seen the monitoring system that winds up costing you more than it costs more than they charge you ancillary fees. This doesn’t do that.

I also—while we’re talking about money, I want to make very clear—because disclaiming the direction the money flows in is always important—you haven’t paid me a dime, ever, to my understanding. I am a paying customer at full price for this service, and I have been since I discovered it. And that is very much an intentional choice. You did not sponsor this podcast, you are not paying me to say nice things. We’re talking because I legitimately adore this thing that you’ve built, and I want it to exist.

Maciej: That’s correct. And again, thank you for that. [laugh].

Corey: It’s true. You can buy my attention, but not my opinion. Now, to be clear, when I did that tweet thread, I did get the sense that this was something that you had built as sort of a side project, as a labor of love. It does not have VC behind it, of which I’m aware, and that’s always going to, on some level, shade how I approach a service and how critical I’m going to be on it. Just because it’s, yeah, if you’ve raised a couple 100 million dollars and your user experience is trash, I’m going to call that out.

But if this is something where you just soft launched, yeah, I’m not going to be a jerk about weird usability bugs here. I might call it out as “Ooh, this is an area for improvement,” but not, “What jackwagon thought of this?” I am trying to be a kinder, gentler Corey in the new year. But at the same time, I also want to be very clear that there’s room for improvement on everything. What surprised me the most about this is how well you nailed the user experience despite not having a full team of people doing UX research.

Maciej: That was definitely a priority. So, maybe a little bit of history. So, I started working on Cloudash, I think it was April… 2019. I think? Yeah. It’s 2021 right now. Or we’re 2022. [unintelligible 00:11:33].

Corey: Yeah. 2022, now. I—

Maciej: I’m sorry. [laugh].

Corey: —I’ve been screwing that up every time I write the dates myself, I’m with you.

Maciej: [laugh]. Okay, so I started working on Cloudash, in 2020, April 2020.

Corey: There we go.

Maciej: So, after eight months, I released some beta, like, free; you could download it from GitHub. Like, you can still download on GitHub, but at that time, there was no license, you didn’t have to buy a license to run it. So, it was, like, very early, like, 0.3 version that was working, but sort of, like, [unintelligible 00:12:00] working. There were some bugs.

And that was the first time that I tweeted about it on Twitter. It gets some attention, but, like, some people started using it. I get some feedback, very initial feedback. And I was like, every time I open Cloudash, I get the sense that, like, this is useful. I’m talking about my own tool, but like, [laugh] that’s the thing.

So, further in the history. So, I’m kind of service engineer by my own. I am a software engineer, I started focusing on serverless, in, like, 2015, 2016. I was working for Serverless Inc. as an early employee.

I was then working as an engineering manager for a couple of companies. I work as an engineering manager right now at Stedi; we’re also, like, fully serverless. So I, kind of, trying to fix my own issues with serverless, or trying to improve the whole experience around serverless in AWS. So, that’s the main purpose why we’re building Cloudash: Because we want to improve the experience. And one use case I’m often mentioning is that, let’s say that you’re kind of on duty. Like, so in the middle of night PagerDuty is calling you, so you need to figure out what’s going on with your Lambda or API Gateway.

Corey: Yes. PagerDuty, the original [Call of Duty: Nagios 00:13:04]. “It’s two in the morning; who is it?” “It’s PagerDuty. Wake up, jackass.” Yeah. We all had those moments.

Maciej: Exactly. So, the PagerDuty is calling you and you’re, kind of, in the middle of night, you’re not sure what’s going on. So, the kind of thing that we want to optimize is from waking up into understanding what’s going on with your serverless stuff should be minimized. And that’s the purpose of Cloudash as well. So, you should just run one tool, and you should immediately see what’s going on. And that’s the purpose.

And probably with one or two clicks, you should see the logs responsible, for example, in your Lambda. Again, like that’s exactly what we want to cover, that was the initial thing that we want to cover, to kind of minimize the time you spent on troubleshooting serverless apps. Because as we all know, kind of, the longer it’s down, the less money you make, et cetera, et cetera, et cetera.

Corey: This episode is sponsored by our friends at Oracle Cloud. Counting the pennies, but still dreaming of deploying apps instead of "Hello, World" demos? Allow me to introduce you to Oracle's Always Free tier. It provides over 20 free services and infrastructure, networking, databases, observability, management, and security. And—let me be clear here—it's actually free. There's no surprise billing until you intentionally and proactively upgrade your account. This means you can provision a virtual machine instance or spin up an autonomous database that manages itself all while gaining the networking load, balancing and storage resources that somehow never quite make it into most free tiers needed to support the application that you want to build. With Always Free, you can do things like run small scale applications or do proof-of-concept testing without spending a dime. You know that I always like to put asterisks next to the word free. This is actually free, no asterisk. Start now. Visit snark.cloud/oci-free that's snark.cloud/oci-free.

Corey: One of the things that I appreciate about this is that I have something like five different microservices now that power my newsletter production pipeline every week. And periodically, I’ll make a change and something breaks because testing is something that I should really get around to one of these days, but when I’m the only customer, cool. Doesn’t really matter until suddenly I’m trying to write something and it doesn’t work. Great. Time to go diving in, and always I’m never in my best frame of mind for that because I’m thinking about writing for humans not writing for computers. And that becomes a challenge.

And okay, how do I get to the figuring out exactly what is broken this time? Regression testing: It really should be a thing more than it has been for me.

Maciej: You should write those tests. [laugh].

Corey: Yeah. And then I fire this up, and okay, great. Which sub-service is it? Great. Okay, what happened in the last five minutes on that service? Oh, okay, it says it failed successfully in the logs. Okay, that’s on me. I can’t really blame you for that. But all right.

And then it’s a matter of adding more [print or 00:14:54] debug statements, and understanding what the hell is going on, mostly that I’m bad at programming. And then it just sort of works from there. It’s a lot easier to, I guess, to reason about this from my perspective than it is to go through the CloudWatch dashboards, where it’s okay, here’s a whole bunch of metrics on different graphs, most of which you don’t actually care about—as opposed to unified view that you offer—and then “Oh, you want to look at logs, that’s a whole separate sub-service. That’s a different service team, obviously, so go open that up in another browser.” And I’m sitting here going, “I don’t know who designed this, but are there any windows in their house? My God.”

It’s just the saddest thing I can possibly experience when I’m in the middle of trying to troubleshoot. Let’s be clear, when I’m troubleshooting, I am in no mood to be charitable to anyone or anything, so that’s probably unfair to those teams. But by the same token, it’s intensely frustrating when I keep smacking into limitations that get in my way while I’m just trying to get the thing up and running again.

Maciej: As you mentioned about UX that, like, we’ve spent a lot of time thinking about the UX, trying different approaches, trying to understand which metrics are the most important. And as we all know, kind of, serverless simplifies a lot of stuff, and there’s, like, way less metrics that you need to look into when something is happening, but we want to make sure that the stuff that we show—which is duration errors, and p95—are probably the most important in most cases, so like, covering most of this stuff. So sorry, I didn’t mention that before; it was very important from the very beginning. And also, like, literally, I spent a lot of time, like, working on the colors, which sounds funny, [laugh] but I wanted to get them right. We’re not yet working on dark mode, but maybe soon.

Anyways, the visual part, it’s always close to my heart, so we spent a lot of time going back to what just said. So, definitely the experience around using CloudWatch right now, and CloudWatch logs, CloudWatch metrics, is not really tailored for any specific use case because they have to be generic, right? Because AWS has, like, I don’t know, like, 300, or whatever number of services, probably half of them producing logs—maybe not half, maybe—

Corey: We shouldn’t name a number because they’ll release five more between now and when this publishes in 20 minutes.

Maciej: [laugh]. So, CloudWatch has to be generic. What we want to do with Cloudash is to take those generic tools—because we use, of course, CloudWatch logs, CloudWatch metrics, we fetch data from them—but make the visual part more tailored for specific use case—in our case, it’s the serverless use case—and make sure that it’s really, kind of—it shows only the stuff that you need to see, not everything else. So again, like that’s the main purpose. And then one more thing, we—like this is also some kind of measurement of success, we want to reduce number of tabs that you need to have open in your browser when you’re dealing with CloudWatch. So, we tried to put most important stuff in one view so you don’t need to flip between tabs, as you usually do when try to under some kind of broader scope, or broader context of your, you know, error in Lambda.

Corey: What inspired you to do this as a desktop application? Because a lot of companies are doing similar things, as SaaS, as webapps. And I have to—as someone who yourself—you’re a self-described serverless engineer—it seems to me that building a webapp is sort of like the common description use case of a lot of serverless stuff. And you’re sitting here saying, “Nope, it’s desktop app time.” Which again, I’m super glad you did. It’s exactly what I was looking for. How do you get here?

Maciej: I’d been thinking about both kinds of types of apps. So like, definitely webapp was the initial idea how to build something, it was the webapp. Because as you said, like, that’s the default mode. Like, we are thinking webapp; like, let’s build a webapp because I’m an engineer, right? There is some inspiration coming from Dynobase, which was made by a friend [unintelligible 00:18:55] who also lives in Poland—I didn’t mention that; we’re based in [Poznań 00:18:58], Poland.

And when I started thinking about it, there’s a lot of benefits of using this approach. The biggest benefit, as I mentioned, is security; and the second benefit is just most, like, cost-effective because we don’t need to run in the backend, right? We don’t need to download all your metrics, all your logs. We I think, like, let’s think about it, like, from the perspective. Listen, so everyone in the company to start working, they
have to download all of your stuff from your AWS account. Like, that sounds insane because you don’t need all of that stuff elsewhere.

Corey: Store multiple copies of it. Yeah I, generally when I’m looking at this, I care about the last five to ten minutes.

Maciej: Exactly.

Corey: I don’t—

Maciej: Exactly.

Corey: —really care what happened three-and-a-half years ago on this function. Almost always. But occasionally I want to look back at, “Oh, this has been breaking. How long has it been that way?” But I already have that in the AWS environment unless I’ve done the right thing and turned on, you know, log expiry.

Maciej: Exactly. So, this is a lot of, like, I don’t want to be, like, you know, mean to anyone but like, that’s a lot of waste. Like, that’s a lot of waste of compute power because you need to download it; of cost because you need to get this data out of AWS, which you need to pay for, you know, get metric data and stuff like this. So, you need to—

Corey: And almost all of its—what is it? Write once, read never. Because it’s, you don’t generally look at these things.

Maciej: Yeah, yeah. Exactly.

Corey: And so much of this, too, for every invocation I have, even though it’s low traffic stuff, it’s the start with a request ID and what version is running, it tells me ‘latest.’ Helpful. A single line of comment in this case says ‘200.’ Why it says that, I couldn’t tell you. And then it says ‘End request ID.’ The end.

Now, there’s no way to turn that off unless you disabled the ability to write to CloudWatch logs in the function, but ingest on that cost 50 cents a gigabyte, so okay, I guess that’s AWS’s money-making scam of the year. Good for them. But there’s so much of that, it’s like looking at—like, when things are working, it’s like looking at a low traffic site that’s behind a load balancer, where there’s a whole—you have gigabytes, in some cases, of load balancer—of web server logs on the thing that’s sitting in your auto-scaling group. And those logs are just load balancer health checks. 98% of it is just that.

Same type of problem here, I don’t care about that, I don’t want to pay to store it, I certainly don’t want to pay to store it twice. I get it, that makes an awful lot of sense. It also makes your security job a hell of a lot easier because you’re not sitting on a whole bunch of confidential data from other people. Because, “Well, it’s just logs. What could possibly be confidential in there?” “Oh, my sweet summer child, have you seen some of the crap people put in logs?”

Maciej: I’ve seen many things in logs. I don’t want to mention them. But anyways—and also, you know, like, usually when you gave access to your AWS account, it can ruin you. You know, like, there might be a lot of—like, you need to really trust the company to give access to your AWS account. Of course, in most cases, the roles are scoped to, you know, only CloudWatch stuff, actions, et cetera, et cetera, but you know, like, there are some situations in which something may not be properly provisioned. And then you give access to everything.

Corey: And you can get an awful lot of data you wouldn’t necessarily want out of that stuff. Give me just the PDF printout of last month’s bill for a lot of environments, and I can tell you disturbing levels of detail about what your architecture is, just because when you—you can infer an awful lot.

Maciej: Yeah.

Corey: Yeah, I hear you. It makes your security story super straightforward.

Maciej: Yeah, exactly. So, I think just repeat my, like, the some inspiration. And then when I started thinking about Cloudash, like, definitely one of the inspiration was Dynobase, from the, kind of, GUI for, like, more powerful UI for DynamoDB. So, if you’re interested in that stuff, you can also check this out.

Corey: Oh, yeah, I’ve been a big fan of that, too. That’ll be a separate discussion on a different episode, for sure.

Maciej: [laugh]. Yeah.

Corey: But looking at all of this, looking at the approach of, the only real concern—well, not even a concern. The only real challenge I have with it for my use case is that when I’m on the road, the only thing that I bring with me for a computer is my iPad Pro. I’m not suggesting by any means that you should build this as a new an iPad app; that strikes me as, like, 15 levels of obnoxious. But it does mean that sometimes I still have to go diving into the CloudWatch console when I’m not home. Which, you know, without this, without Cloudash, that’s what I was doing originally anyway.

Maciej: You’re the only person that requested that. And we will put that into backlog, and we will get to that at some point. [laugh].

Corey: No, no, no. Smart question is to offer me a specific enterprise tier pricing—.

Maciej: Oh, okay. [laugh].

Corey: —that is eye-poppingly high. It’s like, “Hey, if you want a subsidize feature development, we’re thrilled to empower that.” But—

Maciej: [laugh]. Yeah, yeah. To be honest, I like that would be hard to write [unintelligible 00:23:33] implement as iPad app, or iPhone app, or whatever because then, like, what’s the story behind? Like, how can I get the credentials, right? It’s not possible.

Corey: Yeah, you’d have to have some fun with that. There are a couple of ways I can think of offhand, but then that turns into a sandboxing issue, and it becomes something where you have to store credentials locally, regardless, even if they’re ephemeral. And that’s not great. Maybe turn it into a webapp someday or something. Who knows.

What I also appreciate is that we had a conversation when you first launched, and I wound up basically going on a Zoom call with you and more or less tearing apart everything you’ve built—and ideally constructive way—but looking at a lot of the things you’ve changed in your website, you listened to an awful lot of feedback. You doubled your pricing, for example. Used to be ten bucks a month; now you’re twenty.
Great. I’m a big believer in charging more.

You absolutely add that kind of value because it’s, “Well, twenty bucks a month for a desktop app. That sounds crappy.” It's, “Yeah, jackwagon,
what’s your time worth?” I was spending seven bucks a month in serverless charges, and 120 or 130 a month for Epsagon, and I was thrilled to pieces to be doing it because the value I got from being able to quickly diagnose what the hell was going on far outstripped what the actual cost of doing these things. Don’t fall into the trap of assuming that well, I shouldn’t pay for software. I can just do it myself. Your time is never free. People think it is, but it’s not.

Maciej: That’s true. The original price of $9.99, I think that was the price was the launch promo. After some time, we’ve decided—and after adding more features: API Gateway support—we’ve decided that this is, like, solving way more problems, so like, you should probably pay a little bit more for that. But you’re kind of lucky because you subscribed to it when it was 9.99, and this will be your kind of prize for the end of, you know—

Corey: Well, I’m going to argue with you after the show to raise the price on mine, just because it’s true. It’s the—you want to support the things that you want to exist in the world. I also like the fact that you offered an annual plan because I will go weeks without ever opening the app. And that doesn’t mean it isn’t adding value. It’s that oh, yeah, I will need that now that I’m hitting these issues again.

And if I’m paying on a monthly basis, and it shows up with a, “Oh, you got charged again.” “Well, I didn’t use it this month; I should cancel.” And [unintelligible 00:25:44] to an awful lot of subscriber churn. But in the course of a year, if I don’t have at least one instance in which case, wow, that ten minute span justified the entire $200 annual price tag, then, yeah, you built the wrong thing or it’s not for me, but I can think of three incidents so far since I started using it in the past four months that have led to that being worth everything you will charge me a year, and then some, just because it made it so clear what was breaking.

Maciej: So, in that regard, we are also thinking about the team licenses, that’s definitely on the roadmap. There will be some changes to that. And we definitely working on more and more features. And if we’re—like, the roadmap is mostly about supporting more and more AWS services, so right now it’s Lambda, API Gateway, we’re definitely thinking about SQS, SNS, to get some sense how your messages are going through, probably something, like, DynamoDB metrics. And this is all kind of serverless, but why not going wider? Like, why not going to Fargate? Like, Fargate is theoretically serverless, but you know, like, it’s serverless on—

Corey: It’s serverless with a giant asterisk next to it.

Maciej: Yeah, [laugh] exactly. So, but why not? Like, it’s exactly the same thing in terms of, there is some user flow, there is some user journey, when you want to debug something. You want to go from API Gateway, maybe to the container to see, I don’t know, like, DynamoDB metric or something like that, so it should be all easy. And this is definitely something.

Later, why not EC2 metrics? Like, it would be a little bit harder. But I’m just saying, like, first thing here is that you are not, like, at this point, we are serverless, but once we cover serverless, why not going wider? Why not supporting more and more services and just making sure that all those use cases are correctly modeled with the UI and UX, et cetera?

Corey: That’s going to be an interesting challenge, just because that feels like what a lot of the SaaS monitoring and observability tooling is done. And then you fire this thing up, and it looks an awful lot like the AWS console. And it’s, “Yeah, I just want to look at this one application that doesn’t use any of the rest of those things.” Again, I have full faith and confidence in your ability to pull this off. You clearly have done that well based upon what we’ve seen so far. I just wonder how you’re going to wind up tackling that challenge when you get there.

Maciej: And maybe not EC2. Maybe I went too far. [laugh].

Corey: Yeah, honestly, even EC2-land, it feels like that is more or less a solved problem. If you want to treat it as a bunch of EC2, you can use Nagios. It’s fine.

Maciej: Yeah, totally.

Corey: There are tools that have solved that problem. But not much that I’ve seen has solved the serverless piece the way that I want it solved. You have.

Maciej: So, it’s definitely a long road to make sure that the serverless—and by serverless, I mean serverless how AWS understands serverless, so including Fargate, for example. So, there’s a lot of stuff that we can improve. It’s a lot of stuff that can make easier with Cloudash than it is with CloudWatch, just staying inside serverless, it will take us a lot of time to make sure that is all correct. And correctly modeled, correctly designed, et cetera. So yeah, I went too far with EC2 sorry.

Corey: Exactly. That’s okay. We all go too far with EC2, I assure you.

Maciej: Sorry everyone using EC2 instances. [laugh].

Corey: If people want to kick the tires on it, where can they find it?

Maciej: They can find it on cloudash.dev.

Corey: One D in the middle. That one throws me sometimes.

Maciej: One D. Actually, after talking to you, we have a double-D domain as well, so we can also try ‘Clouddash’ with double-D. [laugh].

Corey: Excellent, excellent. Okay, that is fantastic. Because I keep trying to put the double-D in when I’m typing it in my search tool on my desktop, and it doesn’t show up. And it’s like, “What the—oh, right.” But yeah, we’ll get there one of these days.

Maciej: Only the domain. It’s only the domain. You will be redirected to single-D.

Corey: Exactly.

Maciej: [laugh].

Corey: We’ll have to expand later; I’ll finance the feature request there. It’ll go well. If people want to learn more about what you have to think about these things, where else can they find you?

Maciej: On Twitter, and my Twitter handle is @mthenw. M-then-W, which is M-T-H—mthenw. And my co-founder @tlakomy. You can probably add that to [show notes 00:29:35]. [laugh].

Corey: Oh, I certainly will. It’s fine, yeah. Here’s a whole bunch of letters. I hear you. My Twitter handle used to be my amateur radio callsign. It turns out most people don’t think like that. And yeah, it’s become an iterative learning process. Thank you so much for taking the time to speak with me today and for building this thing. I really appreciate both of them.

Maciej: Thank you for having me here. I encourage everyone to visit cloudash.dev, if you have any feature requests, any questions just send us an email at hello@cloudash.dev, or just go to GitHub repository in the issues; just create an issue, describe what you want and we can talk about it.

We are always happy to help. The main purpose, the ultimate goal of Cloudash is to make the serverless engineer’s life easier, on very high level. And on a little bit lower level, just to make, you know, troubleshooting and debugging serverless apps easier.

Corey: Well, from my perspective, you’ve succeeded.

Maciej: Thank you.

Corey: Thank you. Maciej Winnicki, founder of Cloudash. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry comment telling me exactly why I’m wrong for using an iPad do these things, but not being able to send it because you didn’t find a good way to store the credentials.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Matt

Matt is an AWS DevTools Hero, Serverless Architect, Author and conference speaker.

He is focused on creating the right environment for empowered teams to rapidly deliver business value in a well-architected, sustainable and serverless-first way.

You can usually find him sharing reusable, well architected, serverless patterns over at cdkpatterns.com or behind the scenes bringing CDK Day to life.

Links:

  • AWS CDK Patterns: https://cdkpatterns.com
  • The CDK Book: https://thecdkbook.com
  • CDK Day: https://www.cdkday.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: It seems like there is a new security breach every day. Are you confident that an old SSH key, or a shared admin account, isn’t going to come back and bite you? If not, check out Teleport. Teleport is the easiest, most secure way to access all of your infrastructure. The open source Teleport Access Plane consolidates everything you need for secure access to your Linux and Windows servers—and I assure you there is no third option there. Kubernetes clusters, databases, and internal applications like AWS Management Console, Yankins, GitLab, Grafana, Jupyter Notebooks, and more. Teleport’s unique approach is not only more secure, it also improves developer productivity. To learn more visit: goteleport.com. And not, that is not me telling you to go away, it is: goteleport.com.

Corey: This episode is sponsored in part by our friends at Rising Cloud, which I hadn’t heard of before, but they’re doing something vaguely interesting here. They are using AI, which is usually where my eyes glaze over and I lose attention, but they’re using it to help developers be more efficient by reducing repetitive tasks. So, the idea being that you can run stateless things without having to worry about scaling, placement, et cetera, and the rest. They claim significant cost savings, and they’re able to wind up taking what you’re running as it is in AWS with no changes, and run it inside of their data centers that span multiple regions. I’m somewhat skeptical, but their customers seem to really like them, so that’s one of those areas where I really have a hard time being too snarky about it because when you solve a customer’s problem and they get out there in public and say, “We’re solving a problem,” it’s very hard to snark about that. Multus Medical, Construx.ai and Stax have seen significant results by using them. And it’s worth exploring. So, if you’re looking for a smarter, faster, cheaper alternative to EC2, Lambda, or batch, consider checking them out. Visit risingcloud.com/benefits. That’s risingcloud.com/benefits, and be sure to tell them that I said you because watching people wince when you mention my name is one of the guilty pleasures of listening to this podcast.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’m joined today by Matt Coulter, who is a Technical Architect at Liberty Mutual. You may have had the privilege of seeing him on the keynote stage at re:Invent last year—in Las Vegas or remotely—that last year of course being 2021. But if you make better choices than the two of us did, and found yourself not there, take the chance to go and watch that keynote. It’s really worth seeing.

Matt, first, thank you for joining me. I’m sorry, I don’t have 20,000 people here in the audience to clap this time. They’re here, but they’re all remote as opposed to sitting in the room behind me because you know, social distancing.

Matt: And this left earphone, I just have some applause going, just permanently, just to keep me going. [laugh].

Corey: That’s sort of my own internal laugh track going on. It’s basically whatever I say is hilarious, to that. So yeah, doesn’t really matter what I say, how I say it, my jokes are all for me. It’s fine. So, what was it like being on stage in front of that many people? It’s always been a wild experience to watch and for folks who haven’t spent time on the speaking circuit, I don’t think that there’s any real conception of what that’s like. Is this like giving a talk at work, where I just walk on stage randomly, whatever I happened to be wearing? And, oh, here’s a microphone, I’m going to say words. What is the process there?

Matt: It’s completely different. For context for everyone, before the pandemic, I would have pretty regularly talked in front of, I don’t know, maybe one, two hundred people in Liberty, in Belfast. So, I used to be able to just, sort of, walk in front of them, and lean against the pillar, and use my clicker, and click through, but the process for actually presenting something as big as a keynote and re:Invent is so different. For starters, you think that when you walk onto the stage, you’ll actually be able to see the audience, but the way the lights are set up, you can pretty much see about one row of people, and they’re not the front row, so anybody I knew, I couldn’t actually see.

And yeah, you can only see, sort of like, the from the void, and then you have your screens, so you’ve six sets of screens that tell you your notes as well as what slides you’re on, you know, so you can pivot. But other than that, I mean, it feels like you’re just talking to yourself outside of whenever people, thankfully, applause. It’s such a long process to get there.

Corey: I’ve always said that there are a few different transition stages as the audience size increases, but for me, the final stage is more or less anything above 750 people. Because as you say, you aren’t able to see that many beyond that point, and it doesn’t really change anything meaningfully. The most common example that you see in the wild is jokes that work super well with a small group of people fall completely flat to large audiences. It’s why so much corporate numerous cheesy because yeah, everyone in the rehearsals is sitting there laughing and the joke kills, but now you’ve got 5000 people sitting in a room and that joke just sounds strained and forced because there’s no longer a conversation, and no one has the shared context that—the humor has to change. So, in some cases when you’re telling a story about what you’re going to say on stage, during a rehearsal, they’re going to say, “Well, that joke sounds really corny and lame.” It’s, “Yeah, wait until you see it in front of an audience. It will land very differently.” And I’m usually right on that.

I would also advise, you know, doing what you do and having something important and useful to say, as opposed to just going up there to tell jokes the whole time. I wanted to talk about that because you talked about how you’re using various CDK and other serverless style patterns in your work at Liberty Mutual.

Matt: Yeah. So, we’ve been using CDK pretty extensively since it was, sort of, Q3 2019. At that point, it was new. Like, it had just gone GA at the time, just came out of dev preview. And we’ve been using CDK from the perspective of we want to be building serverless-first, well-architected apps, and ideally we want to be building them on AWS.

Now, the thing is, we have 5000 people in our IT organization, so there’s sort of a couple of ways you can take to try and get those people onto the cloud: You can either go the route of being, like, there is one true path to architecture, this is our architecture and everything you want to build can fit into that square box; or you can go the other approach and try and have the golden path where you say this is the paved road that is really easy to do, but if you want to differentiate from that route, that’s okay. But what you need to do is feed back into the golden path if that works. Then everybody can improve. And that’s where we’ve started been using CDK. So, what you heard me talk about was the software accelerator, and it’s sort of a different approach.

It’s where anybody can build a pattern and then share it so that everybody else can rapidly, you know, just reuse it. And what that means is effectively you can, instead of having to have hundreds of people on a central team, you can actually just crowdsource, and sort of decentralize the function. And if things are good, then a small team can actually come in and audit them, so to speak, and check that it’s well-architected, and doesn’t have flaws, and drive things that way.

Corey: I have to confess that I view the CDK as sort of a third stage automation approach, and it’s one that I haven’t done much work with myself. The first stage is clicking around in the console; the second is using CloudFormation or Terraform; the third stage is what we’re talking about here is CDK or Pulumi, or something like that. And then you ascend to the final fourth stage, which is what I use, which is clicking around in the AWS console, but then you lie to people about it. ClickOps is poised to take over the world. But that’s okay. You haven’t gotten that far yet. Instead, you’re on the CDK side. What advantages does CDK offer that effectively CloudFormation or something like it doesn’t?

Matt: So, first off, for ClickOps in Liberty, we actually have the AWS console as read-only in all of our accounts, except for sandbox. So, you can ClickOps in sandbox to learn, but if you want to do something real, unfortunately, it’s going to fail you. So.—

Corey: I love that pattern. I think I might steal that.

Matt: [laugh]. So, originally, we went heavy on CloudFormation, which is why CDK worked well for us. And because we’ve actually—it’s been a long journey. I mean, we’ve been deploying—2014, I think it was, we first started deploying to AWS, and we’ve used everything from Terraform, to you name it. We’ve built our own tools, believe it or not, that are basically CDK.

And the thing about CloudFormation is, it’s brilliant, but it’s also incredibly verbose and long because you need to specify absolutely everything that you want to deploy, and every piece of configuration. And that’s fine if you’re just deploying a side project, but if you’re in an enterprise that has responsibilities to protect user data, and you can’t just deploy anything, they end up thousands and thousands and thousands of lines long. And then we have amazing guardrails, so if you tried to deploy a CloudFormation template with a flaw in it, we can either just fix it, or reject the deploy. But CloudFormation is not known to be the fastest to deploy, so you end up in this developer cycle, where you build this template by hand, and then it goes through that CloudFormation deploy, and then you get the failure message that it didn’t deploy because of some compliance thing, and developers just got frustrated, and were like, sod this. [laugh].

I’m not deploying to AWS. Back the on-prem. And that’s where CDK was a bit different because it allowed us to actually build abstractions with all of our guardrails baked in, so that it just looked like a standard class, for developers, like, developers already know Java, Python, TypeScript, the languages off CDK, and so we were able to just make it easy by saying, “You want API Gateway? There’s an API Gateway class. You want, I don’t know, an EC2 instance? There you go.” And that way, developers could focus on the thing they wanted, instead of all of the compliance stuff that they needed to care about every time they wanted to deploy.

Corey: Personally, I keep lobbying AWS to add my preferred language, which is crappy shell scripting, but for some reason they haven’t really been quick to add that one in. The thing that I think surprises me, on some level—though, perhaps it shouldn’t—is not just the adoption of serverless that you’re driving at Liberty Mutual, but the way that you’re interacting with that feels very futuristic, for lack of a better term. And please don’t think that I’m in any way describing this in a way that’s designed to be insulting, but I do a bunch of serverless nonsense on Twitter for Pets. That’s not an exaggeration. twitterforpets.com has a bunch of serverless stuff behind it because you know, I have personality defects.

But no one cares about that static site that’s been a slide dump a couple of times for me, and a running joke. You’re at Liberty Mutual; you’re an insurance company. When people wind up talking about big enterprise institutions, you’re sort of a shorthand example of exactly what they’re talking about. It’s easy to contextualize or think of that as being very risk averse—for obvious reasons; you are an insurance company—as well as wanting to move relatively slowly with respect to technological advancement because mistakes are going to have drastic consequences to all of your customers, people’s lives, et cetera, as opposed to tweets or—barks—not showing up appropriately at the right time. How did you get to the, I guess, advanced architectural philosophy that you clearly have been embracing as a company, while having to be respectful of the risk inherent that comes with change, especially in large, complex environments?

Matt: Yeah, it’s funny because so for everyone, we were talking before this recording started about, I’ve been with Liberty since 2011. So, I’ve seen a lot of change in the length of time I’ve been here. And I’ve built everything from IBM applications right the way through to the modern serverless apps. But the interesting thing is, the journey to where we are today definitely started eight or nine years ago, at a minimum because there was something identified in the leadership that they said, “Listen, we’re all about our customers. And that means we don’t want to be wasting millions of dollars, and thousands of hours, and big trains of people to build software that does stuff. We want to focus on why are we building a piece of software, and how quickly can we get there? If you focus on those two things you’re doing all right.”

And that’s why starting from the early days, we focused on things like, okay, everything needs to go through CI/CD pipelines. You need to have your infrastructure as code. And even if you’re deploying on-prem, you’re still going to be using the same standards that we use to deploy to AWS today. So, we had years and years and years of just baking good development practices into the company. And then whenever we started to move to AWS, the question became, do we want to just deploy the same thing or do we want to take full advantage of what the cloud has to offer? And I think because we were primed and because the leadership had the right direction, you know, we were just sitting there ready to say, “Okay, serverless seems like a way we can rapidly help our customers.” And that’s what we’ve done.

Corey: A lot of the arguments against serverless—and let’s be clear, they rhyme with the previous arguments against cloud that lots of people used to make; including me, let’s be clear here. I’m usually wrong when I try to predict the future. “Well, you’re putting your availability in someone else’s hands,” was the argument about cloud. Yeah, it turns out the clouds are better at keeping things up than we are as individual companies.

Then with serverless, it’s the, “Well, if they’re handling all that stuff for you on their side, when they’re down, you’re down. That’s an unacceptable business risk, so we’re going to be cloud-agnostic and multi-cloud, and that means everything we build serverlessly needs to work in multiple environments, including in our on-prem environment.” And from the way that we’re talking about servers and things that you’re building, I don’t believe that is technically possible, unless some of the stuff you’re building is ridiculous. How did you come to accept that risk organizationally?

Matt: These are the conversations that we’re all having. Sort of, I’d say once a week, we all have a multi-cloud discussion—and I really liked the article you wrote, it was maybe last year, maybe the year before—but multi-cloud to me is about taking the best capabilities that are out there and bringing them together. So, you know, like, Azure [ID 00:12:47] or whatever, things from the other clouds that they’re good at, and using those rather than thinking, “Can I build a workload that I can simultaneously pay all of the price to run across all of the clouds, all of the time, so that if one’s down, theoretically, I might have an outage?” So, the way we’ve looked at it is we embraced really early the well-architected framework from AWS. And it talks about things like you need to have multi-region availability, you need to have your backups in place, you need to have things like circuit breakers in place for if third-party goes down, and we’ve just tried to build really resilient architectures as best as we can on AWS. And do you know what I think, if [laugh] it AWS is not—I know at re:Invent, there it went down extraordinarily often compared to normal, but in general—

Corey: We were all tired of re:Invent; their us-east-1 was feeling the exact same way.

Matt: Yeah, so that’s—it deserved a break. But, like, if somebody can’t buy insurance for an hour, once a year, [laugh] I think we’re okay with it versus spending millions to protect that one hour.

Corey: And people make assumptions based on this where, okay, we had this problem with us-east-1 that froze things like the global Route 53 control planes; you couldn’t change DNS for seven hours. And I highlighted that as, yeah, this is a problem, and it’s something to severely consider, but I will bet you anything you’d care to name that there is an incredibly motivated team at AWS, actively fixing that as we speak. And by—I don’t know how long it takes to untangle all of those dependencies, but I promise they’re going to be untangled in relatively short order versus running data centers myself, when I discover a key underlying dependency I didn’t realize was there, well, we need to break that. That’s never going to happen because we’re trying to do things as a company, and it’s just not the most important thing for us as a going concern. With AWS, their durability and reliability is the most important thing, arguably compared to security.

Would you rather be down or insecure? I feel like they pick down—I would hope in most cases they would pick down—but they don’t want to do either one. That is something they are drastically incentivized to fix. And I’m never going to be able to fix things like that and I don’t imagine that you folks would be able to either.

Matt: Yeah, so, two things. The first thing is the important stuff, like, for us, that’s claims. We want to make sure at any point in time, if you need to make a claim you can because that is why we’re here. And we can do that with people whether or not the machines are up or down. So, that’s why, like, you always have a process—a manual process—that the business can operate, irrespective of whether the cloud is still working.

And that’s why we’re able to say if you can’t buy insurance in that hour, it’s okay. But the other thing is, we did used to have a lot of data centers, and I have to say, the people who ran those were amazing—I think half the staff now work for AWS—but there was this story that I heard where there was an app that used to go down at the same time every day, and nobody could work out why. And it was because someone was coming in to clean the room at that time, and they unplugged the server to plug in a vacuum, and then we’re cleaning the room, and then plugging it back in again. And that’s the kind of thing that just happens when you manage people, and you manage a building, and manage a premises. Whereas if you’ve heard that happened that AWS, I mean, that would be front page news.

Corey: Oh, it absolutely would. There’s also—as you say, if it’s the sales function, if people aren’t able to buy insurance for an hour, when us-east-1 went down, the headlines were all screaming about AWS taking an outage, and some of the more notable customers were listed as examples of this, but the story was that, “AWS has massive outage,” not, “Your particular company is bad at technology.” There’s sort of a reputational risk mitigation by going with one of these centralized things. And again, as you’re alluding to, what you’re doing is not life-critical as far as the sales process and getting people to sign up. If an outage meant that suddenly a bunch of customers were no longer insured, that’s a very different problem. But that’s not your failure mode.

Matt: Exactly. And that’s where, like, you got to look at what your business is, and what you’re specifically doing, but for 99.99999% of businesses out there, I’m pretty sure you can be down for the tiny window that AWS is down per year, and it will be okay, as long as you plan for it.

Corey: So, one thing that really surprised me about the entirety of what you’ve done at Liberty Mutual is that you’re a big enterprise company, and you can take a look at any enterprise company, and say that they have dueling mottos, which is, “I am not going to comment on that,” or, “That’s not funny.” Like, the safe mode for any large concern is to say nothing at all. But a lot of folks—not just you—at Liberty have been extremely vocal about the work that you’re doing, how you view these things, and I almost want to call it advocacy or evangelism for the CDK. I’m slightly embarrassed to admit that for a little while there, I thought you were an AWS employee in their DevRel program because you were such an advocate in such strong ways for the CDK itself.

And that is not something I expected. Usually you see the most vocal folks working in environments that, let’s be honest, tend to play a little bit fast and loose with things like formal corporate communications. Liberty doesn’t and yet, there you folks are telling these great stories. Was that hard to win over as a culture, or am I just misunderstanding how corporate life is these days?

Matt: No, I mean, so it was different, right? There was a point in time where, I think, we all just sort of decided that—I mean, we’re really good at what we do from an engineering perspective, and we wanted to make sure that, given the messaging we were given, those 5000 teck employees in Liberty Mutual, if you consider the difference in broadcasting to 5000 versus going external, it may sound like there’s millions, billions of people in the world, but in reality, the difference in messaging is not that much. So, to me what I thought, like, whenever I started anyway—it’s not, like, we had a meeting and all decided at the same time—but whenever I started, it was a case of, instead of me just posting on all the internal channels—because I’ve been doing this for years—it’s just at that moment, I thought, I could just start saying these things externally and still bring them internally because all you’ve done is widened the audience; you haven’t actually made it shallower. And that meant that whenever I was having the internal conversations, nothing actually changed except for it meant external people, like all their Heroes—like Jeremy Daly—could comment on these things, and then I could bring that in internally. So, it almost helped the reverse takeover of the enterprise to change the culture because I didn’t change that much except for change the audience of who I was talking to.

Corey: This episode is sponsored by our friends at Oracle HeatWave is a new high-performance accelerator for the Oracle MySQL Database Service. Although I insist on calling it “my squirrel.” While MySQL has long been the worlds most popular open source database, shifting from transacting to analytics required way too much overhead and, ya know, work. With HeatWave you can run your OLTP and OLAP, don’t ask me to ever say those acronyms again, workloads directly from your MySQL database and eliminate the time consuming data movement and integration work, while also performing 1100X faster than Amazon Aurora, and 2.5X faster than Amazon Redshift, at a third of the cost. My thanks again to Oracle Cloud for sponsoring this ridiculous nonsense.

Corey: One thing that you’ve done that I want to say is admirable, and I stumbled across it when I was doing some work myself over the break, and only right before this recording did I discover that it was you is the cdkpatterns.com website. Specifically what I love about it is that it publishes a bunch of different patterns of ways to do things. This deviates from a lot of tutorials on, “Here’s how to build this one very specific thing,” and instead talks about, “Here’s the architecture design; here’s what the baseline pattern for that looks like.” It’s more than a template, but less than a, “Oh, this is a messaging app for dogs and I’m trying to build a messaging app for cats.” It’s very generalized, but very direct, and I really, really like that model of demo.

Matt: Thank you. So, watching some of your Twitter threads where you experiment with new—

Corey: Uh oh. People read those. That’s a problem.

Matt: I know. So, whatever you experiment with a new piece of AWS to you, I’ve always wondered what it would be like to be your enabling architect. Because technically, my job in Liberty is, I meant to try and stay ahead of everybody and try and ease the on-ramp to these things. So, if I was your enabling architect, I would be looking at it going, “I should really have a pattern for this.” So that whenever you want to pick up that new service the patterns in cdkpatterns.com, there’s 24, 25 of them right there, but internally, there’s way more than dozens now.

The goal is, the pattern is the least amount to code for you to learn a concept. And then that way, you can not only see how something works, but you can maybe pick up one of the pieces of the well-architected framework while you’re there: All of it’s unit tested, all of it is proper, you know, like, commented code. The idea is to not be crap, but not be gold-plated either. I’m currently in the process of upgrading that all to V2 as well. So, that [unintelligible 00:21:32].

Corey: You mentioned a phrase just now: “Enabling architect.” I have to say this one that has not crossed my desk before. Is that an internal term you use? Is that an enterprise concept I’ve somehow managed to avoid? Is that an AWS job role? What is that?

Matt: I’ve just started saying [laugh] it’s my job over the past couple of years. That—I don’t know, patent pending? But the idea to me is—

Corey: No, it’s evocative. I love the term, I’d love to learn more.

Matt: Yeah, because you can sort of take two approaches to your architecture: You can take the traditional approach, which is the ‘house of
no’ almost, where it’s like, “This is the architecture. How dare you want to deviate. This is what we have decided. If you want to change it, here’s the Architecture Council and go through enterprise architecture as people imagine it.” But as people might work out quite quickly, whenever they meet me, the whole, like, long conversational meetings are not for me. What I want to do is teach engineers how to help themselves, so that’s why I see myself as enabling.

And what I’ve been doing is using techniques like Wardley Mapping, which is where you can go out and you can actually take all the components of people’s architecture and you can draw them on a map for—it’s a map of how close they are to the customer, as well as how cutting edge the tech is, or how aligned to our strategic direction it is. So, you can actually map out all of the teams, and—there’s 160, 170 engineers in Belfast and Dublin, and I can actually go in and say, “Oh, that piece of your architecture would be better if it was evolved to this. Well, I have a pattern for that,” or, “I don’t have a pattern for that, but you know what? I’ll build one and let’s talk about it next week.” And that’s always trying to be ahead, instead of people coming to me and I have to say no.

Corey: AWS Proton was designed to do something vaguely similar, where you could set out architectural patterns of—like, the two examples that they gave—I don’t know if it’s in general availability yet or still in public preview, but the ones that they gave were to build a REST API with Lambda, and building something-or-other with Fargate. And the idea was that you could basically fork those, or publish them inside of your own environment of, “Oh, you want a REST API; go ahead and do this.” It feels like their vision is a lot more prescriptive than what yours is.

Matt: Yeah. I talked to them quite a lot about Proton, actually because, as always, there’s different methodologies and different ways of doing things. And as I showed externally, we have our software accelerator, which is kind of our take on Proton, and it’s very open. Anybody can contribute; anybody can consume. And then that way, it means that you don’t necessarily have one central team, you can have—think of it more like an SRE function for all of the patterns, rather than… the Proton way is you’ve separate teams that are your DevOps teams that set up your patterns and then separate team that’s consumer, and they have different permissions, different rights to do different things. If you use a Proton pattern, anytime an update is made to that pattern, it auto-deploys your infrastructure.

Corey: I can see that breaking an awful lot.

Matt: [laugh]. Yeah. So, the idea is sort of if you’re a consumer, I assume you [unintelligible 00:24:35] be going to change that infrastructure. You can, they’ve built in an escape hatch, but the whole concept of it is there’s a central team that looks to what the best configuration for that is. So, I think Proton has so much potential, I just think they need to loosen some of the boundaries for it to work for us, and that’s the feedback I’ve given them directly as well.

Corey: One thing that I want to take a step beyond this is, you care about this? More than most do. I mean, people will work with computers, yes. We get paid for that. Then they’ll go and give talks about things. You’re doing that as well. They’ll launch a website occasionally, like, cdkpatterns.com, which you have. And then you just sort of decide to go for the absolute hardest thing in the world, and you’re one of four authors of a book on this. Tell me more.

Matt: Yeah. So, this is something that there’s a few of us have been talking since one of the first CDK Days, where we’re friends, so there’s AWS Heroes. There’s Thorsten Höger, Matt Bonig, Sathyajith Bhat, and myself, came together—it was sometime in the summer last year—and said, “Okay. We want to write a book, but how do we do this?” Because, you know, we weren’t authors before this point; we’d never done it before. We weren’t even sure if we should go to a publisher, or if we should self-publish.

Corey: I argue that no one wants to write a book. They want to have written a book, and every first-time author I’ve ever spoken to at the end has said, “Why on earth would anyone want to do this a second time?” But people do it.

Matt: Yeah. And that’s we talked to Alex DeBrie, actually, about his book, the amazing Dynamodb Book. And it was his advice, told us to self-publish. And he gave us his starter template that he used for his book, which took so much of the pain out because all we had to do was then work out how we were going to work together. And I will say, I write quite a lot of stuff in general for people, but writing a book is completely different because once it’s out there, it’s out there. And if it’s wrong, it’s wrong. You got to release a new version and be like, “Listen, I got that wrong.” So, it did take quite a lot of effort from the group to pull it together. But now that we have it, I want to—I don’t have a printed copy because it’s only PDF at the minute, but I want a copy just put here [laugh] in, like, the frame. Because it’s… it’s what we all want.

Corey: Yeah, I want you to do that through almost a traditional publisher, selfishly, because O’Reilly just released the AWS Cookbook, and I had a great review quote on the back talking about the value added. I would love to argue that they use one of mine for The CDK Book—and then of course they would reject it immediately—of, “I don’t know why you do all this. Using the console and lying about it is way easier.” But yeah,
obviously not the direction you’re trying to take the book in. But again, the industry is not quite ready for the lying version of ClickOps.

It’s really neat to just see how willing you are to—how to frame this?—to give of yourself and your time and what you’ve done so freely. I sometimes make a joke—that arguably isn’t that funny—that, “Oh, AWS Hero. That means that you basically volunteer for a $1.6 trillion company.”

But that’s not actually what you’re doing. What you’re doing is having figured out all the sharp edges and hacked your way through the jungle to get to something that is functional, you’re a trailblazer. You’re trying to save other people who are working with that same thing from difficult experiences on their own, having to all thrash and find our own way. And not everyone is diligent and as willing to continue to persist on these things. Is that a somewhat fair assessment how you see the Hero role?

Matt: Yeah. I mean, no two Heroes are the same, from what I’ve judged, I haven’t met every Hero yet because pandemic, so Vegas was the first time [I met most 00:28:12], but from my perspective, I mean, in the past, whatever number of years I’ve been coding, I’ve always been doing the same thing. Somebody always has to go out and be the first person to try the thing and work out what the value is, and where it’ll work for us more work for us. The only difference with the external and public piece is that last 5%, which it’s a very different thing to do, but I personally, I like even having conversations like this where I get to meet people that I’ve never met before.

Corey: You sort of discovered the entire secret of why I have an interview podcast.

Matt: [laugh]. Yeah because this is what I get out of it, just getting to meet other people and have new experiences. But I will say there’s
Heroes out there doing very different things. You’ve got, like, Hiro—as in Hiro, H-I-R-O—actually started AWS Newbies and she’s taught—ah, it’s hundreds of thousands of people how to actually just start with AWS, through a course designed for people who weren’t coders before. That kind of thing is next-level compared to anything I’ve ever done because you know, they have actually built a product and just given it away. I think that’s amazing.

Corey: At some level, building a product and giving it away sounds like, “You know, I want to never be lonely again.” Well, that’ll work because you’re always going to get support tickets. There’s an interesting narrative around how to wind up effectively managing the community, and users, and demands, based on open-source maintainers, that we’re all wrestling with as an industry, particularly in the wake of that whole log4j nonsense that we’ve been tilting at that windmill, and that’s going to be with us for a while. One last thing I want to talk about before we wind up calling this an episode is, you are one of the organizers of CDK Day. What is that?

Matt: Yeah, so CDK Day, it’s a complete community-organized conference. The past two have been worldwide, fully virtual just because of the situation we’re in. And I mean, they’ve been pretty popular. I think we had about 5000 people attended the last one, and the idea is, it’s a full day of the community just telling their stories of how they liked or disliked using the CDK. So, it’s not a marketing event; it’s not a sales event; we actually run the whole event on a budget of exactly $0. But yeah, it’s just a day of fun to bring the community together and learn a few things. And, you know, if you leave it thinking CDK is not for you, I’m okay with that as much as if you just make a few friends while you’re there.

Corey: This is the first time I’d realized that it wasn’t a formal AWS event. I almost feel like that’s the tagline that you should have under it. It’s—because it sounds like the CDK Day, again, like, it’s this evangelism pure, “This is why it’s great and why you should use it.” But I love conferences that embrace critical views. I built one of the first talks I ever built out that did anything beyond small user groups was “Heresy in the Church of Docker.”

Then they asked me to give that at ContainerCon, which was incredibly flattering. And I don’t think they made that mistake a second time, but it was great to just be willing to see some group of folks that are deeply invested in the technology, but also very open to hearing criticism. I think that’s the difference between someone who is writing a nuanced critique versus someone who’s just [pure-on 00:31:18] zealotry. “But the CDK is the answer to every technical problem you’ve got.” Well, I start to question the wisdom of how applicable it really is, and how objective you are. I’ve never gotten that vibe from you.

Matt: No, and that’s the thing. So, I mean, as we’ve worked out in this conversation, I don’t work for AWS, so it’s not my product. I mean, if it succeeds or if it fails, it doesn’t impact my livelihood. I mean, there are people on the team who would be sad for, but the point is, my end goal is always the same. I want people to be enabled to rapidly deliver their software to help their customers.

If that’s CDK, perfect, but CDK is not for everyone. I mean, there are other options available in the market. And if, even, ClickOps is the way to go for you, I am happy for you. But if it’s a case of we can have a conversation, and I can help you get closer to where you need to be with some other tool, that’s where I want to be. I just want to help people.

Corey: And if I can do anything to help along that axis, please don’t hesitate to let me know. I really want to thank you for taking the time to speak with me and being so generous, not just with your time for this podcast, but all the time you spend helping the rest of us figure out which end is up, as we continue to find that the way we manage environments evolves.

Matt: Yeah. And, listen, just thank you for having me on today because I’ve been reading your tweets for two years, so I’m just starstruck at this moment to even be talking to you. So, thank you.

Corey: No, no. I understand that, but don’t worry, I put my pants on two legs at a time, just like everyone else. That’s right, the thought leader on Twitter, you have to jump into your pants. That’s the rule. Thanks again so much. I look forward to having a further conversation with you about this stuff as I continue to explore, well honestly, what feels like a brand new paradigm for how we manage code.

Matt: Yeah. Reach out if you need any help.

Corey: I certainly will. You’ll regret asking. Matt [Coulter 00:33:06], Technical Architect at Liberty Mutual. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, write an angry comment, then click the submit button, but lie and say you hit the submit button via an API call.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Miles

As Chief Technology Officer at SADA, Miles Ward leads SADA’s cloud strategy and solutions capabilities. His remit includes delivering next-generation solutions to challenges in big data and analytics, application migration, infrastructure automation, and cost optimization; reinforcing our engineering culture; and engaging with customers on their most complex and ambitious plans around Google Cloud.

Previously, Miles served as Director and Global Lead for Solutions at Google Cloud. He founded the Google Cloud’s Solutions Architecture practice, launched hundreds of solutions, built Style-Detection and Hummus AI APIs, built CloudHero, designed the pricing and TCO calculators, and helped thousands of customers like Twitter who migrated the world’s largest Hadoop cluster to public cloud and Audi USA who re-platformed to k8s before it was out of alpha, and helped Banco Itau design the intercloud architecture for the bank of the future.

Before Google, Miles helped build the AWS Solutions Architecture team. He wrote the first AWS Well-Architected framework, proposed Trusted Advisor and the Snowmobile, invented GameDay, worked as a core part of the Obama for America 2012 “tech” team, helped NASA stream the Curiosity Mars Rover landing, and rebooted Skype in a pinch.

Earning his Bachelor of Science in Rhetoric and Media Studies from Willamette University, Miles is a three-time technology startup entrepreneur who also plays a mean electric sousaphone.

Links:

  • SADA.com: https://sada.com
  • Twitter: https://twitter.com/milesward
  • Email: miles@sada.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: It seems like there is a new security breach every day. Are you confident that an old SSH key, or a shared admin account, isn’t going to come back and bite you? If not, check out Teleport. Teleport is the easiest, most secure way to access all of your infrastructure. The open source Teleport Access Plane consolidates everything you need for secure access to your Linux and Windows servers—and I assure you there is no third option there. Kubernetes clusters, databases, and internal applications like AWS Management Console, Yankins, GitLab, Grafana, Jupyter Notebooks, and more. Teleport’s unique approach is not only more secure, it also improves developer productivity. To learn more visit: goteleport.com. And not, that is not me telling you to go away, it is: goteleport.com.

Corey: This episode is sponsored in part by our friends at Redis, the company behind the incredibly popular open source database that is not the bind DNS server. If you’re tired of managing open source Redis on your own, or you’re using one of the vanilla cloud caching services, these folks have you covered with the go to manage Redis service for global caching and primary database capabilities; Redis Enterprise. To learn more and deploy not only a cache but a single operational data platform for one Redis experience, visit redis.com/hero. Thats r-e-d-i-s.com/hero. And my thanks to my friends at Redis for sponsoring my ridiculous non-sense.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I am joined today, once again by my friend and yours, Miles Ward, who’s the CTO at SADA. However, he is, as I think of him, the closest thing the Google Cloud world has to Corey Quinn. Now, let’s be clear, not the music and dancing part that is Forrest Brazeal, but Forrest works at Google Cloud, whereas Miles is a reasonably salty third-party. Miles, thank you for coming back and letting me subject you to that introduction.

Miles: Corey, I appreciate that introduction. I am happy to provide substantial salt. It is easy, as I play brass instruments that produce my spit in high volumes. It’s the most disgusting part of any possible introduction. For the folks in the audience, I am surrounded by a collection of giant sousaphones, tubas, trombones, baritones, marching baritones, trumpets, and pocket trumpets.

So, Forrest threw down the gauntlet and was like, I can play a keyboard, and sing, and look cute at the same time. And so I decided to fail at all three. We put out a new song just a bit ago that’s, like, us thanking all of our customers and partners, covering Kool & the Gang “Celebration,” and I neither look good, [laugh] play piano, or smiling, or [capturing 00:01:46] any of the notes; I just play the bass part, it’s all I got to do.

Corey: So, one thing that I didn’t get to talk a lot about because it’s not quite in my universe, for one, and for another, it is during the pre
re:Invent—pre:Invent, my nonsense thing—run up, which is Google Cloud Next.

Miles: Yes.

Corey: And my gag a few years ago is that I’m not saying that Google is more interested in what they’re building and what they’re shipping, but even their conference is called Next. Buh dum, hiss.

Miles: [laugh].

Corey: So, I didn’t really get to spend a lot of attention on the Google Cloud releases that came out this year, but given that SADA is in fact the, I believe, largest Google Cloud partner on the internet, and thus the world—

Miles: [unintelligible 00:02:27] new year, three years in a row back, baby.

Corey: Fantastic. I assume someone’s watch got stuck or something. But good work. So, you have that bias in the way that I have a bias, which is your business is focused around Google Cloud the way that mine is focused on AWS, but neither of us is particularly beholden to that given company. I mean, you do have the not getting fired as partner, but that’s a bit of a heavy lift; I don’t think I can mouth off well enough to get you there.

So, we have a position of relative independence. So, you were tracking Google Next, the same way that I track re:Invent. Well, not quite the same way I track re:Invent; there are some significant differences. What happened at Cloud Next 2021, that the worst of us should be paying attention to?

Miles: Sure. I presented 10% of the material at the first re:Invent. There are 55 sessions; I did six. And so I have been at Cloud events for a really long time and really excited about Google’s willingness to dive into demos in a way that I think they have been a little shy about. Kelsey Hightower is the kind of notable deep exception to that. Historically, he’s been ready to dive into the, kind of, heavy hands-on piece but—

Corey: Wait, those were demos? [Thought 00:03:39] was just playing Tetris on stage for the love of it.

Miles: [laugh]. No. And he really codes all that stuff up, him and the whole team.

Corey: Oh, absol—I’m sorry. If I ever grow up, I wish to be Kelsey Hightower.

Miles: [laugh]. You and me both. So, he had kind of led the charge. We did a couple of fun little demos while I was there, but they’ve really gotten a lot further into that, and I think are doing a better job of packaging the benefits to not just developers, but also operators and data scientists and the broader roles in the cloud ecosystem from the new features that are being launched. And I think, different than the in-person events where there’s 10, 20,000, 40,000 people in the audience paying attention, I think they have to work double-hard to capture attention and get engineers to tune in to what’s being launched.

But if you squint and look close, there are some, I think, very interesting trends that sit in the back of some of the very first launches in what I think are going to be whole veins of launches from Google over the course of the next several years that we are working really hard to track along with and make sure we’re extracting maximum value from for our customers.

Corey: So, what was it that they announced that is worth paying attention to? Now, through the cacophony of noise, one announcement that [I want to note 00:04:49] was tied to Next was the announcement that GME group, I believe, is going to be putting their futures exchange core trading systems on Google Cloud. At which point that to me—and I know people are going to yell at me, and I don’t even slightly care—that is the last nail in the coffin of the idea that well, Google is going to turn this off in a couple years. Sorry, no. That is not a thing that’s going to happen. Worst case, they might just stop investing it as aggressively as they are now, but even that would be just a clown-shoes move that I have a hard time envisioning.

Miles: Yeah, you’re talking now over a dozen, over ten year, over a billion-dollar commitments. So, you’ve got to just really, really hate your stock price if you’re going to decide to vaporize that much shareholder value, right? I mean, we think that, in Google, stock price is a material fraction of the recognition of the growth trajectory for cloud, which is now basically just third place behind YouTube. And I think you can do the curve math, it’s not like it’s going to take long.

Corey: Right. That requires effectively ejecting Thomas Kurian as the head of Google Cloud and replacing him with the former SVP of Bad Decisions at Yahoo.

Miles: [laugh]. Sure. Google has no shyness about continuing to rotate leadership. I was there through three heads of Google Cloud, so I don’t expect that Thomas will be the last although I think he may well go down in history as having been the best. The level of rotation to the focuses that I think are most critical, getting enterprise customers happy, successful, committed, building macroscale systems, in systems that are critical to the core of the business on GCP has grown at an incredible rate under his stewardship. So, I think he’s doing a great job.

Corey: He gets a lot of criticism—often from Googlers—when I wind up getting the real talk from them, which is, “Can you tell me what you really think?” Their answer is, “No,” I’m like, “Okay, next question. Can I go out and buy you eight beers and then”— and it’s like, “Yeah.” And the answer that I get pretty commonly is that he’s brought too much Oracle into Google. And okay, that sounds like a bad thing because, you know, Oracle, but let’s be clear here, but what are you talking about specifically? And what they say distills down to engineers are no longer the end-all be-all of everything that Google Cloud. Engineers don’t get to make sales decisions, or marketing decisions, or in some cases, product decisions. And that is not how Google has historically been run, and they don’t like the change. I get it, but engineering is not the only hard thing in the world and it’s not the only business area that builds value, let’s be clear on this. So, I think that the things that they don’t like are in fact, what Google absolutely needs.

Miles: I think, one, the man is exceptionally intimidating and intentionally just hyper, hyper attentive to his business. So, one of my best employees, Brad [Svee 00:07:44], he worked together with me to lay out what was the book of our whole department, my team of 86 people there. What are we about? What do we do? And like I wanted this as like a memoriam to teach new hires as got brought in. So, this is, like, 38 pages of detail about our process, our hiring method, our promotional approach, all of it. I showed that to my new boss who had come in at the time, and he thought some of the pictures looked good. When we showed it to TK, he read every paragraph. I watched him highlight the paragraphs as he went through, and he read it twice as fast as I can read the thing. I think he does that to everybody’s documents, everywhere. So, there’s a level of just manual rigor that he’s brought to the practice that was certainly not there before that. So, that alone, it can be intimidating for folks, but I think people that are high performance find that very attractive.

Corey: Well, from my perspective, he is clearly head and shoulders above Adam Selipsky, and Scott Guthrie—the respective heads of AWS and Azure—for one key reason: He is the only one of those three people who follows me on Twitter. And—

Miles: [laugh].

Corey: —honestly, that is how I evaluate vendors.

Miles: That’s the thing. That’s the only measure, yep. I’ve worked on for a long time with Selipsky, and I think that it will be interesting to see whether Adam’s approach to capital allocation—where he really, I think, thinks of himself as the manager of thousands of startups, as opposed to a manager of a global business—whether that’s a more efficient process for creating value for customers, then, where I think TK is absolutely trying to build a much more unified, much more singular platform. And a bunch of the launches really speak to that, right? So, one of the product announcements that I think is critical is this idea of the global distributed cloud, Google Distributed Cloud.

We started with Kubernetes. And then you layer on to that, okay, we’ll take care of Kubernetes for you; we call that Anthos. We’ll build a bunch of structural controls and features into Anthos to make it so that you can really deal with stuff in a global way. Okay, what does that look like further? How do we get out into edge environments? Out into diverse hardware? How do we partner up with everybody to make sure that, kind of like comparing Apple’s approach to Google’s approach, you have an Android ecosystem of Kubernetes providers instead of just one place you can buy an outpost. That’s generally the idea of GDC. I think that’s a spot where you’re going to watch Google actually leverage the muscle that it already built in understanding open-source dynamics and understanding collaboration between companies as opposed to feeling like it's got to be built here. We’ve got to sell it here. It’s got to have our brand on it.

Corey: I think that there’s a stupendous and extreme story that is still unfolding over at Google Cloud. Now, re:Invent this year, they wound up talking all about how what they were rolling out was a focus on improving primitives. And they’re right. I love their managed database service that they launched because it didn’t exist.

Miles: Yeah Werner’s slide, “It’s primitives, not frameworks.” I was like, I think customers want solutions, not frameworks or primitives. [laugh]. What’s your plan?

Corey: Yeah. However, I take a different perspective on all of this, which is that is a terrific spin on the big headline launches all missed the re:Invent timeline, and… oops, so now we’re just going to talk about these other things instead. And that’s great, but then they start talking about industrial IOT, and mainframe migrations, and the idea of private 5G, and running fleets of robots. And it’s—

Miles: Yeah, that’s a cool product.

Corey: Which one? I’m sorry, they’re all very different things.

Miles: Private 5G.

Corey: Yeah, if someone someday will explain to me how it differs from Wavelength, but that’s neither here nor there. You’re right, they’re all interesting, but none of them are actually doing the thing that I do, which is build websites, [unintelligible 00:11:31] looking for web services, it kind of says it in the name. And it feels like it’s very much broadening into everything, and it’s very difficult for me to identify—and if I have trouble that I guarantee you customers do—of, which services are for me and which are very much not? In some cases, the only answer to that is to check the pricing. I thought Kendra, their corporate information search thing was for me, then it’s 7500 bucks a month to get started with that thing, and that is, “I can hire an internal corporate librarian to just go and hunt through our Google Drive.” Great.

Miles: Yeah.

Corey: So, there are—or our Dropbox, or our Slack. We have, like, five different information repositories, and this is how corporate nonsense starts, let me assure you.

Miles: Yes. We call that luxury SaaS, you must enjoy your dozens of overlapping bills for, you know, what Workspace gives you as a single flat rate.

Corey: Well, we have [unintelligible 00:12:22] a lot of this stuff, too. Google Drive is great, but we use Dropbox for holding anything that touches our customer’s billing information, just because I—to be clear, I do not distrust Google, but it also seems a little weird to put the confidential billing information for one of their competitors on there to thing if a customer were to ask about it. So, it’s the, like, I don’t believe anyone’s doing anything nefarious, but let’s go ahead and just make sure, in this case.

Miles: Go further man. Vimeo runs on GCP. You think YouTube doesn’t want to look at Vimeo stats? Like they run everything on GCP, so they have to have arrived at a position of trust somehow. Oh, I know how it’s called encryption. You’ve heard of encryption before? It’s the best.

Corey: Oh, yes. I love these rumors that crop up every now and again that Amazon is going to start scanning all of its customer content, somehow. It’s first, do you have any idea how many compute resources that would take and to if they can actually do that and access something you’re storing in there, against their attestations to the contrary, then that’s your story because one of them just makes them look bad, the other one utterly destroys their entire business.

Miles: Yeah.

Corey: I think that that’s the one that gets the better clicks. So no, they’re not doing that.

Miles: No, they’re not doing that. Another product launch that I thought was super interesting that describes, let’s call it second place—the third place will be the one where we get off into the technical deep end—but there’s a whole set of coordinated work they’re calling Cortex. So, let’s imagine you go to a customer, they say, “I want to understand what’s happening with my business.” You go, “Great.” So, you use SAP, right? So, you’re a big corporate shop, and that's your infrastructure of choice. There are a bunch of different options at that layer.

When you set up SAP, one of the advantages that something like that has is they have, kind of, pre-built configurations for roughly your business, but whatever behaviors SAP doesn’t do, right, say, data warehousing, advanced analytics, regression and projection and stuff like that, maybe that’s somewhat outside of the core wheelhouse for SAP, you would expect like, oh okay, I’ll bolt on BigQuery. I’ll build that stuff over there. We’ll stream the data between the two. Yeah, I’m off to the races, but the BigQuery side of the house doesn’t have this like bitching menu that says, “You’re a retailer, and so you probably want to see these 75 KPIs, and you probably want to chew up your SKUs in exactly this way. And here’s some presets that make it so that this is operable out of the box.”

So, they are doing the three way combination: Consultancies plus ISVs plus Google products, and doing all the pre-work configuration to go out to a customer and go I know what you probably just want. Why don’t I just give you the whole thing so that it does the stuff that you want? That I think—if that’s the very first one, this little triangle between SAP, and Big Query, and a bunch of consultancies like mine, you have to imagine they go a lot further with that a lot faster, right? I mean, what does that look like when they do it with Epic, when they go do it with Go just generally, when they go do it with Apache? I’ve heard of that software, right? Like, there’s no reason not to bundle up what the obvious choices are for a bunch of these combinations.

Corey: The idea of moving up the stack and offering full on solutions, that’s what customers actually want. “Well, here’s a bunch of things you can do to wind up wiring together to build a solution,” is, “Cool. Then I’m going to go hire a company who’s already done that is going to sell it to me at a significant markup because I just don’t care.” I pay way more to WP Engine than I would to just run WordPress myself on top of AWS or Google Cloud. In fact, it is on Google Cloud, but okay.

Miles: You and me both, man. WP Engine is the best. I—

Corey: It’s great because—

Miles: You’re welcome. I designed a bunch of the hosting on the back of that.

Corey: Oh, yeah. But it’s also the—I—well, it costs a little bit more that way. Yeah, but guess what’s not—guess what’s more expensive than that bill, is my time spent doing the care and feeding of this stuff. I like giving money to experts and making it their problem.

Miles: Yeah. I heard it said best, Lego is an incredible business. I love their product, and you can build almost any toy with it. And they have not displaced all other plastic toy makers.

Corey: Right.

Miles: Some kids just want to buy a little car. [laugh].

Corey: Oh, yeah, you can build anything you want out of Lego bricks, which are great, which absolutely explains why they are a reference AWS customer.

Miles: Yeah, they’re great. But they didn’t beat all other toy companies worldwide, and eliminate the rest of that market because they had the better primitive, right? These other solutions are just as valuable, just as interesting, tend to have much bigger markets. Lego is not the largest toy manufacturer in the world. They are not in the top five of toy manufacturers in the world, right?

Like, so chasing that thread, and getting all the way down into the spots where I think many of the cloud providers on their own, internally, had been very uncomfortable. Like, you got to go all the way to building this stuff that they need for that division, inside of that company, in that geo, in that industry? That’s maybe, like, a little too far afield. I think Google has a natural advantage in its more partner-oriented approach to create these combinations that lower the cost to them and to customers to getting out of that solution quick.

Corey: So, getting into the weeds of Google Next, I suppose, rather than a whole bunch of things that don’t seem to apply to anyone except the four or five companies that really could use it, what things did Google release that make the lives of people building, you know, web apps better?

Miles: This is the one. So, I’m at Amazon, hanging out as a part of the team that built up the infrastructure for the Obama campaign in 2012, and there are a bunch of Googlers there, and we are fighting with databases. We are fighting so hard, in fact, with RDS that I think we are the only ones that [Raju 00:17:51] has ever allowed to SSH into our RDS instances to screw with them.

Corey: Until now, with the advent of RDS Custom, meaning that you can actually get in as root; where that hell that lands between RDS and EC2 is ridiculous. I just know that RDS can now run containers.

Miles: Yeah. I know how many things we did in there that were good for us, and how many things we did in there that were bad for us. And I have to imagine, this is not a feature that they really ought to let everybody have, myself included. But I will say that what all of the Googlers that I talk to, you know, at the first blush, were I’m the evil Amazon guy in to, sort of, distract them and make them build a system that, you know, was very reliable and ended up winning an election was that they had a better database, and they had Spanner, and they didn’t understand why this whole thing wasn’t sitting on Spanner. So, we looked, and I read the white paper, and then I got all drooly, and I was like, yes, that is a much better database than everybody else’s database, and I don’t understand why everybody else isn’t on it. Oh, there’s that one reason, but you’ve heard of it: No other software works with it, anywhere in the world, right? It’s utterly proprietary to Google. Yes, they were kind—

Corey: Oh, you want to migrate it off somewhere else, or a fraction of it? Great. Step one, redo your data architecture.

Miles: Yeah, take all of my software everywhere, rewrite every bit of it. And, oh all those commercial applications? Yeah, forget all those, you got, too. Right? It was very much where Google was eight years ago. So, for me, it was immensely meaningful to see the launch at Next where they described what they are building—and have now built; we have alpha access to it—a Postgres layer for Spanner.

Corey: Is that effectively you have to treat it as Postgres at all times, or is it multimodal access?

Miles: You can get in and tickle it like Spanner, if you want to tickle it like Spanner. And in reality, Spanner is ANSI SQL compliant; you’re still writing SQL, you just don’t have to talk to it like a REST endpoint, or a GRPC endpoint, or something; you can, you know, have like a—

Corey: So, similar to Azure’s Cosmos DB, on some level, except for the part where you can apparently look at other customers’ data in that thing?

Miles: [laugh]. Exactly. Yeah, you will not have a sweeping discovery of incredible security violations in the structure Spanner, in that it is the control system that Google uses to place every ad, and so it does not suck. You can’t put a trillion-dollar business on top of a database and not have it be safe. That’s kind of a thing.

Corey: The thing that I find is the most interesting area of tech right now is there’s been this rise of distributed databases. Yugabyte—or You-ji-byte—Pla-netScale—or PlanetScale, depending on how you pronounce these things.

Miles: [laugh]. Yeah, why, why is G such an adversarial consonant? I don’t understand why we’ve all gotten to this place.

Corey: Oh, yeah. But at the same time, it’s—so you take a look at all these—and they all are speaking Postgres; it is pretty clear that ‘Postgres-squeal’ is the thing that is taking over the world as far as databases go. If I were building something from scratch that used—

Miles: For folks in the back, that’s PostgreSQL, for the rest of us, it’s okay, it’s going to be, all right.

Corey: Same difference. But yeah, it’s the thing that is eating the world. Although recently, I’ve got to say, MongoDB is absolutely stepping up in a bunch of really interesting ways.

Miles: I mean, I think the 4.0 release, I’m the guy who wrote the MongoDB on AWS Best Practices white paper, and I would grab a lot of customer’s and—

Corey: They have to change it since then of, step one: Do not use DocumentDB; if you want to use Mongo, use Mongo.

Miles: Yeah, that’s right. No, there were a lot of customers I was on the phone with where Mongo had summarily vaporized their data, and I think they have made huge strides in structural reliability over the course of—you know, especially this 4.0 launch, but the last couple of years, for sure.

Corey: And with all the people they’ve been hiring from AWS, it’s one of those, “Well, we’ll look at this now who’s losing important things from production?”

Miles: [laugh]. Right? So, maybe there’s only actually five humans who know how to do operations, and we just sort of keep moving around these different companies.

Corey: That’s sort of my assumption on these things. But Postgres, for those who are not looking to depart from the relational model, is eating the world. And—

Miles: There’s this, like, basic emotional thing. My buddy Martin, who set up MySQL, and took it public, and then promptly got it gobbled up by the Oracle people, like, there was a bet there that said, hey, there’s going to be a real open database, and then squish, like, the man came and got it. And so like, if you’re going to be an independent, open-source software developer, I think you’re probably not pushing your pull requests to our friends at Oracle, that seems weird. So instead, I think Postgres has gobbled up the best minds on that stuff.

And it works. It’s reliable, it’s consistent, and it’s functional in all these different, sort of, reapplications and subdivisions, right? I mean, you have to sort of squint real hard, but down there in the guts of Redshift, that’s Postgres, right? Like, there’s Postgres behind all sorts of stuff. So, as an interface layer, I’m not as interested about how it manages to be successful at bossing around hardware and getting people the zeros and ones that they ask for back in a timely manner.

I’m interested in it as a compatibility standard, right? If I have software that says, “I need to have Postgres under here and then it all will work,” that creates this layer of interop that a bunch of other products can use. So, folks like PlanetScale, and Yugabyte can say, “No, no, no, it’s cool. We talk Postgres; that’ll make it so your application works right. You can bring a SQL alchemy and plug it into this, or whatever your interface layer looks like.”

That’s the spot where, if I can trade what is a fairly limited global distribution, global transactional management on literally ridiculously unlimited scalability and zero operations, I can handle the hard parts of running a database over to somebody else, but I get my layer, and my software talks to it, I think that’s a huge step.

Corey: This episode is sponsored in part by my friends at Cloud Academy. Something special just for you folks. If you missed their offer on Black Friday or Cyber Monday or whatever day of the week doing sales it is—good news! They’ve opened up their Black Friday promotion for a very limited time. Same deal, $100 off a yearly plan, $249 a year for the highest quality cloud and tech skills content. Nobody else can get this because they have a assured me this not going to last for much longer. Go to CloudAcademy.com, hit the "start free trial" button on the homepage, and use the Promo code cloud at checkout. That’s c-l-o-u-d, like loud, what I am, with a “C” in front of it. It's a free trial, so you'll get 7 days to try it out to make sure it's really a good fit for you, nothing to lose except your ignorance about cloud. My thanks again for sponsoring my ridiculous nonsense.

Corey: I think that there’s a strong movement toward building out on something like this. If it works, just because—well, I’m not multiregion today, but I can easily see a world in which I’d want to be. So, great. How do you approach the decision between—once this comes out of alpha; let’s be clear. Let’s turn this into something that actually ships, and no, Google that does not mean slapping a beta label on it for five years is the answer here; you actually have to stand behind this thing—but once it goes GA—

Miles: GA is a good thing.

Corey: Yeah. How do you decide between using that, or PlanetScale? Or Yugabyte?

Miles: Or Cockroach or or SingleStore, right? I mean, there’s a zillion of them that sit in this market. I think the core of the decision making for me is in every team you’re looking at what skills do you bring to bear and what problem that you’re off to go solve for customers? Do the nuances of these products make it easier to solve? So, I think there are some products that the nature of what you’re building isn’t all that dependent on one part of the application talking to another one, or an event happening someplace else mattering to an event over here. But some applications, that’s, like, utterly critical, like, totally, totally necessary.

So, we worked with a bunch of like Forex exchange trading desks that literally turn off 12 hours out of the day because they can only keep it consistent in one geographical location right near the main exchanges in New York. So, that’s a place where I go, “Would you like to trade all day?” And they go, “Yes, but I can’t because databases.” So, “Awesome. Let’s call the folks on the Spanner side. They can solve that problem.”

I go, “Would you like to trade all day and rewrite all your software?” And they go, “No.” And I go, “Oh, okay. What about trade all day, but not rewrite all your software?” There we go. Now, we’ve got a solution to that kind of problem.

So like, we built this crazy game, like, totally other end of the ecosystem with the Dragon Ball Z people, hysterical; your like—you literally play like Rock, Paper, Scissors with your phone, and if you get a rock, I throw a fireball, and you get a paper, then I throw a punch, and we figure out who wins. But they can play these games like Europe versus Japan, thousands of people on each side, real-time, and it works.

Corey: So, let’s be clear, I have lobbied a consistent criticism at Google for a while now, which is the Google Cloud global control plane. So, you wind up with things like global service outages from time to time, you wind up with this thing is now broken for everyone everywhere. And that, for a lot of these use cases, is a problem. And I said that AWS’s approach to regional isolation is the right way to do it. And I do stand by that assessment, except for the part where it turns out there’s a lot of control plane stuff that winds up single tracking through us-east-1, as we learned in the great us-east-1 outage of 2021.

Miles: Yeah, when I see customers move from data center to AWS, what they expect is a higher count of outages that lasts less time. That’s the trade off, right? There’s going to be more weird spurious stuff, and maybe—maybe—if they’re lucky, that outage will be over there at some other region they’re not using. I see almost exactly the same promise happening to folks that come from AWS—and in particular from Azure—over onto GCP, which is, there will be probably a higher frequency of outages at a per product level, right? So, like sometimes, like, some weird product takes a screw sideways, where there is structural interdependence between quite a few products—we actually published a whole internal structural map of like, you know, it turns out that Cloud SQL runs on top of GCE not on GKE, so you can expect if GKE goes sideways, Cloud SQL is probably not going to go sideways; the two aren’t dependent on each other.

Corey: You take the status page and Amazon FreeRTOS in a region is having an outage today or something like that. You’re like, “Oh, no. That’s terrible. First, let me go look up what the hell that is.” And I’m not using it? Absolutely not. Great. As hyperscalers, well, hyperscale, they’re always things that are broken in different ways, in different locations, and if you had a truly accurate status page, it would all be red all the time, or varying shades of red, which is not helpful. So, I understand the challenge there, but very often, it’s a partition that is you are not exposed to, or the way that you’ve architected things, ideally, means it doesn’t really matter. And that is a good thing. So, raw outage counts don’t solve that. I also maintain that if I were to run in a single region of AWS or even a single AZ, in all likelihood, I will have a significantly better uptime across the board than I would if I ran it myself. Because—

Miles: Oh, for sure.

Corey: —it is—

Miles: For sure they’re way better at ops than you are. Me, right?

Corey: Of course.

Miles: Right? Like, ridiculous.

Corey: And they got that way, by learning. Like, I think in 2022, it is unlikely that there’s going to be an outage in an AWS availability zone by someone tripping over a power cable, whereas I have actually done that. So, there’s a—to be clear in a data center, not an AWS facility; that would not have flown. So, there is the better idea of of going in that direction. But the things like Route 53 is control plane single-tracking through the us-east-1, if you can’t make DNS changes in an outage scenario, you may as well not have a DR plan, for most use cases.

Miles: To be really clear, it was a part of the internal documentation on the AWS side that we would share with customers to be absolutely explicit with them. It’s not just that there are mistakes and accidents which we try to limit to AZs, but no, go further, that we may intentionally cause outages to AZs if that’s what allows us to keep broader service health higher, right? They are not just a blast radius because you, oops, pulled the pin on the grenade; they can actually intentionally step on the off button. And that’s different than the way Google operates. They think of each of the AZs, and each of the regions, and the global system as an always-on, all the time environment, and they do not have systems where one gets, sort of, sacrificed for the benefit of the rest, right, or they will intentionally plan to take a system offline.

There is no planned downtime in the SLA, where the SLAs from my friends at Amazon and Azure are explicit to, if they choose to, they decide to take it offline, they can. Now, that’s—I don’t know, I kind of want the contract that has the other thing where you don’t get that.

Corey: I don’t know what the right answer is for a lot of these things. I think multi-cloud is dumb. I think that the idea of having this workload that you’re going to seamlessly deploy to two providers in case of an outage, well guess what? The orchestration between those two providers is going to cause you more outages than you would take just sticking on one. And in most cases, unless you are able to have complete duplication of not just functionality but capacity between those two, congratulations, you’ve now just doubled your number of single points of failure, you made the problem actively worse and more expensive. Good job.

Miles: I wrote an article about this, and I think it’s important to differentiate between dumb and terrifyingly shockingly expensive, right? So, I have a bunch of customers who I would characterize as rich, as like, shockingly rich, as producing businesses that have 80-plus percent gross margins. And for them, the costs associated with this stuff are utterly rational, and they take on that work, and they are seeing benefits, or they wouldn’t be doing it.

Corey: Of course.

Miles: So, I think their trajectory in technology—you know, this is a quote from a Google engineer—it’s just like, “Oh, you want to see what the future looks like? Hang out with rich people.” I went into houses when I was a little kid that had whole-home automation. I couldn’t afford them; my mom was cleaning house there, but now my house, I can use my phone to turn on the lights. Like—

Corey: You know, unless us-east-1 is having a problem.

Miles: Hey, and then no Roomba for you, right? Like utterly offline. So—

Corey: Roomba has now failed to room.

Miles: Conveniently, my lights are Philips Hue, and that’s on Google, so that baby works. But it is definitely a spot where the barrier of entry and the level of complexity required is going down over time. And it is definitely a horrible choice for 99% of the companies that are out there right now. But next year, it’ll be 98. And the year after that, it’ll probably be 97. [laugh].

And if I go inside of Amazon’s data centers, there’s not one manufacturer of hard drives, there’s a bunch. So, that got so easy that now, of course you use more than one; you got to do—that’s just like, sort of, a natural thing, right? These technologies, it’ll move over time. We just aren’t there yet for the vast, vast majority of workloads.

Corey: I hope that in the future, this stuff becomes easier, but data transfer fees are going to continue to be a concern—

Miles: Just—[makes explosion noise]—

Corey: Oh, man—

Miles: —like, right in the face.

Corey: —especially with the Cambrian explosion of data because the data science folks have successfully convinced the entire industry that there’s value in those mode balancer logs in 2012. Okay, great. We’re never deleting anything again, but now you’ve got to replicate all of that stuff because no one has a decent handle on lifecycle management and won’t for the foreseeable future. Great, to multiple providers so that you can work on these things? Like, that is incredibly expensive.

Miles: Yeah. Cool tech, from this announcement at Next that I think is very applicable, and recognized the level of like, utter technical mastery—and security mastery to our earlier conversation—that something like this requires, the product is called BigQuery Omni, what Omni allows
you to do is go into the Google Cloud Console, go to BigQuery, say I want to do analysis on this data that’s in S3, or in Azure Blob Storage, Google will spin up an account on your behalf on Amazon and Azure, and run the compute there for you, bring the result back. So, just transfer the answers, not the raw data that you just scanned, and no work on your part, no management, no crapola. So, there’s like—that’s multi-cloud. If I’ve got—I can do a join between a bunch of rows that are in real BigQuery over on GCP side and rows that are over there in S3. The cross-eyedness of getting something like that to work is mind blowing.

Corey: To give this a little more context, just because it gets difficult to reason about these things, I can either have data that is in a private subnet in AWS that traverses their horribly priced Managed NAT Gateways, and then goes out to the internet and sent there once, for the same cost as I could take that same data and store it in S3 in their standard tier for just shy of six full months. That’s a little imbalanced, if we’re being direct here. And then when you add in things like intelligent tiering and archive access classes, that becomes something that… there’s no contest there. It’s, if we’re talking about things that are now approaching exabyte scale, that’s one of those, “Yeah, do you want us to pay by a credit card?”—get serious. You can’t at that scale anyway—“Invoice billing, or do we just, like, drive a dump truck full of gold bricks and drop them off in Seattle?”

Miles: Sure. Same trajectory, on the multi-cloud thing. So, like a partner of ours, PacketFabric, you know, if you’re a big, big company, you go out and you call Amazon and you buy 100 gigabit interconnect on—I think they call theirs Direct Connect, and then you hook that up to the Google one that’s called Dedicated Interconnect. And voila, the price goes from twelve cents a gig down to two cents a gig; everybody’s much happier. But Jesus, you pay the upfront for that, you got to set the thing up, it takes days to get deployed, and now you’re culpable for the whole pipe if you don’t use it up. Like, there are charges that are static over the course of the month.

So, PacketFabric just buys one of those and lets you rent a slice of it you need. And I think they’ve got an incredible product. We’re working with them on a whole bunch of different projects. But I also expect—like, there’s no reason the cloud providers shouldn’t be working hard to vend that kind of solution over time. If a hundred gigabit is where it is now, what does it look like when I get to ten gigabit? When I get to one gigabit? When I get to half gigabit? You know, utility price that for us so that we get to rational pricing.

I think there’s a bunch of baked-in business and cost logic that is a part of the pricing system, where egress is the source of all of the funding at Amazon for internal networking, right? I don’t pay anything for the switches that connect to this machine to that machine, in region. It’s not like those things are cheap or free; they have to be there. But the funding for that comes from egress. So, I think you’re going to end up seeing a different model where you’ll maybe have different approaches to egress pricing, but you’ll be paying like an in-system networking fee.

And I think folks will be surprised at how big that fee likely is because of the cost of the level of networking infrastructure that the providers deploy, right? I mean, like, I don’t know, if you’ve gone and tried to buy a 40 port, 40 gig switch anytime recently. It’s not like they’re those little, you know, blue Netgear ones for 90 bucks.

Corey: Exactly. It becomes this, [sigh] I don’t know, I keep thinking that’s not the right answer, but part of it also is like, well, you know, for things that I really need local and don’t want to worry about if the internet’s melting today, I kind of just want to get, like, some kind of Raspberry Pi shoved under my desk for some reason.

Miles: Yeah. I think there is a lot where as more and more businesses bet bigger and bigger slices of the farm on this kind of thing, I think it’s Jassy’s line that you’re, you know, the fat in the margin in your business is my opportunity. Like, there’s a whole ecosystem of partners and competitors that are hunting all of those opportunities. I think that pressure can only be good for customers.

Corey: Miles, thank you for taking the time to speak with me. If people want to learn more about you, what you’re up to, your bad opinions, your ridiculous company, et cetera—

Miles: [laugh].

Corey: —where can they find you?

Miles: Well, it’s really easy to spell: SADA.com, S-A-D-A dot com. I’m Miles Ward, it’s @milesward on Twitter; you don’t have to do too hard of a math. It’s miles@sada.com, if you want to send me an email. It’s real straightforward. So, eager to reach out, happy to help. We’ve got a bunch of engineers that like helping people move from Amazon to GCP. So, let us know.

Corey: Excellent. And we will, of course, put links to this in the [show notes 00:37:17] because that’s how we roll.

Miles: Yay.

Corey: Thanks so much for being so generous with your time, and I look forward to seeing what comes out next year from these various cloud companies.

Miles: Oh, I know some of them already, and they’re good. Oh, they’re super good.

Corey: This is why I don’t do predictions because like, the stuff that I know about, like, for example, I was I was aware of the Graviton 3 was coming—

Miles: Sure.

Corey: —and it turns out that if your—guess what’s going to come up and you don’t name Graviton 3, it’s like, “Are you simple? Did you not see that one coming?” It’s like—or if I don’t know it’s coming and I make that guess—which is not the hardest thing in the world—someone would think I knew and leaked. There’s no benefit to doing predictions.

Miles: No. It’s very tough, very happy to do predictions in private, for customers. [laugh].

Corey: Absolutely. Thanks again for your time. I appreciate it.

Miles: Cheers.

Corey: Myles Ward, CTO at SADA. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice and be very angry in your opinion when you write that obnoxious comment, but then it’s going to get lost because it’s using MySQL instead of Postgres.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Levi

Levi's passion lies in helping others learn to cloud better.

Links:

  • Jamf: https://www.jamf.com
  • Twitter: https://twitter.com/levi_mccormick

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: It seems like there is a new security breach every day. Are you confident that an old SSH key, or a shared admin account, isn’t going to come back and bite you? If not, check out Teleport. Teleport is the easiest, most secure way to access all of your infrastructure. The open-source Teleport Access Plane consolidates everything you need for secure access to your Linux and Windows servers, and I assure you there is no third option there. Kubernetes clusters, databases, and internal applications like AWS Management Console, Yankins, GitLab, Grafana, Jupyter Notebooks, and more. Teleport’s unique approach is not only more secure, it also improves developer productivity. To learn more visit: goteleport.com. And not, that is not me telling you to go away, it is: goteleport.com.

Corey: This episode is sponsored in part by our friends at Rising Cloud, which I hadn’t heard of before, but they’re doing something vaguely interesting here. They are using AI, which is usually where my eyes glaze over and I lose attention, but they’re using it to help developers be more efficient by reducing repetitive tasks. So, the idea being that you can run stateless things without having to worry about scaling, placement, et cetera, and the rest. They claim significant cost savings, and they’re able to wind up taking what you’re running as it is in AWS with no changes, and run it inside of their data centers that span multiple regions. I’m somewhat skeptical, but their customers seem to really like them, so that’s one of those areas where I really have a hard time being too snarky about it because when you solve a customer’s problem and they get out there in public and say, “We’re solving a problem,” it’s very hard to snark about that. Multus Medical, Construx.ai and Stax have seen significant results by using them. And it’s worth exploring. So, if you’re looking for a smarter, faster, cheaper alternative to EC2, Lambda, or batch, consider checking them out. Visit risingcloud.com/benefits. That’s risingcloud.com/benefits, and be sure to tell them that I said you because watching people wince when you mention my name is one of the guilty pleasures of listening to this podcast.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I am known-slash-renowned-slash-reviled for my creative pronunciations of various technologies, company names, et cetera. Kubernetes, for example, and other things that get people angry on the internet. The nice thing about today’s guest is that he works at a company where there is no possible way for me to make it more ridiculous than it sounds because Levi McCormick is a cloud architect at Jamf. I know Jamf sounds like I’m trying to pronounce letters that are designed to be silent, but no, no, it’s four letters: J-A-M-F. Jamf. Levi, thanks for joining me.

Levi: Thanks for having me. I’m super excited.

Corey: Exactly. Also professional advice for anyone listening: Making fun of company names is hilarious; making fun of people’s names makes you a jerk. Try and remember that. People sometimes blur that distinction.

So, very high level, you’re a cloud architect. Now, I remember the days of enterprise architects where their IDEs were basically whiteboards,
and it was a whole bunch of people sitting in a room. They call it an ivory tower, but I’ve been in those rooms; I assure you there is nothing elevated about this. It’s usually a dank sub-basement somewhere. What do you do, exactly?

Levi: Well, I am part of the enterprise architecture team at Jamf. My roles include looking at our use of cloud; making sure that we’re using our resources to the greatest efficacy possible; coordinating between many teams, many products, many architectures; trying to make sure that we’re using best practices; bringing them from the teams that develop them and learn them, socializing them to other teams; and just trying to keep a handle on this wild ride that we’re on.

Corey: So, what I find fun is that Jamf has been around for a long time. I believe it is not your first name. I want to say Casper was originally?

Levi: I believe so, yeah.

Corey: We’re Jamf customers. You’re not sponsoring this episode or anything, to the best of my knowledge. So, this is not something I’m trying to shill the company, but we’re a customer; we use you to basically ensure that all of our company MacBooks, and laptops, et cetera, et cetera, are basically ensured that there’s disk encryption turned on, that people have a password, and that screensaver is turned on, basically to mean that if someone gets their laptop stolen, it’s a, “Oh, I have to spend more money with Apple,” and not, “Time to sound the data breach alarm,” for reasons that should be blindingly obvious. And it’s great not just at the box check, but also fixing the real problem of I [laugh] don’t want to lose data that is sensitive for obvious reasons. I always thought of this is sort of a thing that worked on the laptops. Why do you have a cloud team?

Levi: Many reasons. First of all, we started in the business of providing the software that customers would run in their own data centers, in their own locations. Sometime in about 2015, we decided that we are properly equipped to run this better than other people, and we started
to provide that as a service. People would move in, migrate their services into the cloud, or we would bring people into the cloud to start with.

Device management isn’t the only thing that we do. We provide some SSO-type services, we recently acquired a company called Wandera, which does endpoint security and a VPN-like experience for traffic. So, there’s a lot of cloud powering all of those things.

Corey: Are you able to disclose whether you’re focusing mostly on AWS, on Azure, on Google Cloud, or are you pretending a cloud with something like IBM?

Levi: All of the above, I believe.

Corey: Excellent. That tells you it’s a real enterprise, in seriousness. It’s the—we talk about the idea of going all in on one providers being a general best practice of good place to start. I believe that. And then there are exceptions, and as companies grow and accumulate technical debt, that also is load-bearing and generates money, you wind up with this weird architectural series of anti-patterns, and when you draw it on a whiteboard of, “Here’s our architecture,” the junior consultant comes in and says, “What moron built this?” Usually two said quote-unquote, “Moron,” and then they’ve just pooched the entire engagement.

Yeah, most people don’t show up in the morning hoping to do a terrible job today, unless they work at Facebook. So, there are reasons things are the way they are; they’re constraints that shape these things. Yeah, if people were going to be able to shut down the company for two years and rebuild everything from scratch from the ground up, it would look wildly different. But you can’t do that most of the time.

Levi: Yeah. Those things are load bearing, right? You can’t just stop traffic one day, and re-architect it with the golden image of what it should have been. We’ve gone through a series of acquisitions, and those architectures are disparate across the different acquired products. So, you have to be able to leverage lessons from all of them, bring them together and try and just slowly, incrementally march towards a better future state.

Corey: As we take a look at the challenges we see The Duckbill Group over on my side of the world, where we talk to customers, it’s I think it is surprising to folks to learn that cloud economics as I see it is—well, first, cost and architecture the same thing, which inherently makes sense, but there’s a lot more psychology that goes into it than math. People often assume I spend most of my time staring into spreadsheets. I assure you that would not go super well. But it has to do with the psychological elements of what it is that people are wrestling with, of their understanding of the environment has not kept pace with reality, and APIs tend to, you know, tell truths.

It’s always interesting to me to see the lies that customers tell, not intentionally, but the reality of it of, “Okay, what about those big instances you’re running in Australia?” “Oh, we don’t have any instances in Australia.” “Look, I understand that you are saying that in good faith, however…” and now we’re in a security incident mode and it becomes a whole different story. People’s understanding always trails. What do you spend the bulk of your time doing? Is it building things? Is it talking to people? Is it trying to more or less herd cats in certain directions? What’s the day-to-day?

Levi: I would say it varies week-to-week. Depends on if we have a new product rolling out. I spend a lot of my time looking at architectural diagrams, reference architectures from AWS. The majority of the work I do is in AWS and that’s where my expertise lies. I haven’t found it financially incentivized to really branch out into any of the other clouds in terms of expertise, but I spend a lot of my time developing solutions, socializing them, getting them in front of teams, and then educating.

We have a wide range of skills internally in terms of what people know or what they’ve been exposed to. I’d say a lot of engineers want to learn the cloud and they want to get opportunities to work on it, and their day-to-day work may not bring them those opportunities as often as they’d like. So, a good portion of my time is spent educating, guiding, joining people’s sprints, joining in their stand-ups, and just kind of talking through, like, how they should approach a problem.

Corey: Whenever you work at a big company, you invariably wind up with—well, microservices becomes the right answer, not because of the technical reasons; because of the people reason, the way that you get a whole bunch of people moving in roughly the same direction. You are a large scale company; who owns services in your idealized view of the world? Is it, “Well, I wrote something and it’s five o’clock. Off to production with it. Talk to you in two days, if everything—if we still have a company left because I didn’t double-check what I just wrote.”

Do you think that the people who are building services necessarily should be the ones supporting it? Like, in other words, Amazon’s approach of having the software engineers being responsible for the ones running it in production from an ops perspective. Is that the direction you trend towards, or do you tend to be from my side of the world—which is grumpy sysadmin—where people—developers hurl applications into your yard for you to worry about?

Levi: I would say, I’m an extremist in the view of supporting the Amazon perspective. I really like you build it, you run it, you own it, you architect it, all of it. I think the other teams in the organization should exist to support and enable those paths. So, if you have platform teams are a really common thing you see hired right now, I think those platforms should be built to enable the company’s perspective on operating infrastructure or services, and then those service teams on top of that should be enabled to—and empowered to make the decisions on how they want to build a service, how they want to provide it. Ultimately, the buck should stop with them.

You can get into other operational teams, you could have a systems operation team, but I think there should be an explicit contract between a service team, what they build, and what they hand off, you know, you could hand off, like, a tier one level response, you know, you can do playbooks, you could do, you know, minimal alert, response, routing, that kind of stuff with a team, but I think that even that team should have a really strong contract with, like, here’s what our team provides, here’s how you engage with our team, here’s how you will transition services to our team.

Corey: The challenge with doing that, in some shops, has been that if you decide to roll out a, you build it, you own it, approach that has not been there since the beginning, you wind up with a lot of pushback from engineers who until now really enjoyed their 5:30 p.m. quitting time, or whenever it was they wound up knocking off work. And they started pushing back, like, “Working out of hours? That’s inhumane.” And the DevOps team would be sitting there going, “We’re right here. How dare you? Like, what do you think our job is?” And it’s a, “Yes, but you’re not people.” And then it leads to this whole back and forth acrimonious—we’ll charitably call it a debate. How do you drive that philosophy?

Levi: It’s a challenge. I’ve seen many teams fracture, fall apart, disperse, if you will, under the transition of going through, like, an extreme service ownership. I think you balance it out with the carrot of you also get to determine your own future, right? You get to determine the programming language you use, you get to determine the underlying technologies that you use. Again, there’s a contract: You have to meet this list of security concerns, you need to meet these operational concerns, and how you do that is up to you.

Corey: When you take a look across various teams—let’s bound this to the industry because I don’t necessarily want you to wind up answering tough questions at work the day this episode airs—what do you see the biggest blockers to achieving, I guess, a functional cultural service ownership?

Levi: It comes down to people’s identity. They’ve established their own identity, “As I am X,” right? I’m a operations engineer. I’m a developer, I’m an engineer. And getting people to kind of branch out of that really fixed mindset is hard, and that, to me, is the major blocker to people assuming ownership.

I’ve seen people make the transition from, “I’m just an engineer. I just want to write code.” I hate those lines. That frustrates me so much: “I just want to write code.” Transitioning into that, like, ownership of, “I had an idea. I built the platform or the service. It’s a huge hit.” Or you know, “Lots of people are using it.” Like, seeing people go through that transformation become empowered, become fulfilled, I think is great.

Corey: I didn’t really expect to get called out quite like this, but you’re absolutely right. I was against the idea, back when I was a sysadmin type because I didn’t know how to code. And if you have developers supporting all of the stuff that they’ve built, then what does that mean for me? It feels like my job is evaporating. I don’t know how to write code.

Well, then I started learning how to write code incredibly badly. And then wow, it turns out, everyone does this. And here we are. But it’s—I don’t build applications, for obvious reasons. I’m bad at it, but I found another way to proceed in the wide world that we live in of high technology.

But yeah, it was hard because this idea of my sense of identity being tied to the thing that I did, it really was an evolve-or-die dinosaur kind of moment because I started seeing this philosophy across the board. You take a look, even now at modern SRE is, or modern DevOps folks, or modern sysadmins, what they’re doing looks a lot less like logging into Linux systems and tinkering on the command line a lot more like running and building distributed applications. Sure, this application that you’re rolling out is the one that orchestrates everything there, but you’re still running this in the same way the software engineers do, which is, interestingly.

Levi: And that doesn’t mean a team has to be only software engineers. Your service team can be multiple disciplines. It should be multiple disciplines. I’ve seen a traditional ops team broken apart, and those individuals distributed into the services that they were chiefly skilled in supporting in the past, as the ops team, as we transitioned those roles from one of the worst on-call rotations I’ve ever seen—you know, 13 to 14 alerts a night—transitioning those out to those service teams, training them up on the operations, building the playbooks. That was their role. Their role wasn’t necessarily to write software, day one.

Corey: I quit a job after six weeks because of that style of, I guess, mismanagement. Their approach was that, oh, we’re going to have our monitoring system live in AWS because one of our VPs really likes AWS—let’s be clear, this was 2008, 2009 era—latency was a little challenging there. And [unintelligible 00:17:04] he really liked Big Brother, which was—not to—now before that became a TV show and at rest, it was a monitoring system—but network latency was always a weird thing in AWS in those days, so instead, he insisted we set up three of them. And whenever—if we just got one page, it was fine. But if we got three, then we had to jump in. And two was always undefined.

And they turned this off from I think, 10 p.m. to 6 a.m. every night, just so the person I call could sleep. And I’m looking at this, like, this might be the worst thing I’ve ever seen in my life. This was before they released the Managed NAT Gateway, so possibly it was.

Levi: And then the flood, right, when you would get—

Corey: Oh, God this was the days, too—

Levi: Yeah.

Corey: —when you were—if you weren’t careful, you’d set this up to page you on the phone with a text message and great, now it takes time for my cell provider to wind up funneling out the sudden onslaught of 4000 text messages. No thanks.

Levi: If your monitoring system doesn’t have the ability to say, you know, the alert flood, funnel them into one alert, or just pause all alerts, while—because we know there’s an incident; you know, us-east-1 is down, right? We know this; we don’t need to get 500 text messages to each engineer that’s on call.

Corey: Well, my philosophy at that point was no, I’m going to instead take a step beyond. If I’m not empowered to fix this thing that is waking me up—and sometimes that’s the monitoring system, and sometimes it’s the underlying application—I’m not on call.

Levi: Yes, exactly. And that’s why I like the model of extre—you know, the service ownership: Because those alerts should go to the people—the pain should be felt by the people who are empowered to fix it. It should not land anywhere else. Otherwise, that creates misaligned incentives and nothing gets better.

Corey: Yeah. But in large distributed systems, very often the person is on call more or less turns into a traffic router.

Levi: Right. That’s unfair to them.

Corey: That’s never fun—yeah, that’s unfair, and it’s not fun, either, and there’s no great answer when you’ve all these different contributory
factors.

Levi: And how hard is it to keep the team staffed up?

Corey: Oh, yeah. It’s a, “Hey, you want a really miserable job one week out of every however many there are in the cycle?” Eh, people don’t like that.

Levi: Exactly.

Corey: This episode is sponsored by our friends at Oracle HeatWave, a new high-performance accelerator for the Oracle MySQL Database Service, although I insist on calling it, “My squirrel.” While MySQL has long been the world's most popular open source database, shifting from transacting to analytics required way too much overhead and, you know, work. With HeatWave you can run your OLAP and OLTP—don’t ask me to ever say those acronyms again—workloads directly from your MySQL database and eliminate the time consuming data movement and integration work, while also performing 1100X faster than Amazon Aurora, and 2.5X faster than Amazon Redshift, at a third of the cost. My thanks again to Oracle Cloud for sponsoring this ridiculous nonsense.

Corey: So, I’ve been tracking what you’re up to for little while now—you’re always a blast to talk with—what is this whole Cloud Builder thing that you were talking about for a bit, and then I haven’t seen much about it.

Levi: Ah, so at the beginning of the pandemic, our mutual friend, Forrest Brazeal, released the Cloud Resume Challenge. I looked at that, and I thought, this is a fantastic idea. I’ve seen lots of people going through it. I recommend the people I mentor go through it. Great way to pick up
a couple cloud skills here and there, tell an interesting story in an interview, right? It’s a great prep.

I intended the Cloud Builder Challenge to be a natural kind of progression from that Resume Challenge to the Builder Challenge where you get operational experience. Again, back to that, kind of, extreme service ownership mentality, here’s a project where you can build, really modeled on the Amazon GameDays from re:Invent, you build a service, we’ll send you traffic, you process those payloads, do some matching, some sorting, some really light processing on these payloads, and then send it back to us, score some points, we’ll build a public dashboard, people can high five each other, they can razz each other, kind of competition they want to do. Really low, low pressure, but just a fun way to get more operational experience in an area where there is really no downside. You know, playing like that at work, bad idea, right?

Corey: Generally, yes. [crosstalk 00:21:28] production, we used to have one of those environments; oops-a-doozy.

Levi: Yeah. I don’t see enough opportunities for people to gain that experience in a way that reflects a real workload. You can go out and you can find all kinds of Hello Worlds, you can find all kinds of—like, for front end development, there are tons of activity activities and things you can do to learn the skills, but for the middleware, the back end engineers, there’s just not enough playgrounds out there. Now, standing up a Hello World app, you know, you’ve got your infrastructures code template, you’ve got your pre-written code, you deploy it, congratulations. But now what, right?

And I intended this challenge to be kind of a series of increasingly more difficult waves, if you will, or levels. I really had a whole gamification aspect to it. So, it would get harder, it would get bigger, more traffic, you know, all of those things, to really put people through what it would be like to receive your, “Post got slash-dotted today,” or those kinds of things where people don’t get an opportunity to deal with large amounts of traffic, or variable payloads, that kind of stuff.

Corey: I love the idea. Where is it?

Levi: It is sitting in a bunch of repos, and I am afraid to deploy it. [laugh].

Corey: What is it that scares you about it specifically?

Levi: The thing that specifically scares me is encouraging early career developers to go out there, deploy this thing, start playing with it, and then incur a huge cloud bill.

Corey: Because they failed to secure something or other reasons behind that?

Levi: There are many ways that this could happen, yeah. You could accidentally push your access key, secret key up into a public repo. Now, you’ve got, you know, Bitcoin miners or Monero miners running in your environment. You forget to shut things off, right? That’s a really common thing.

I went through a SageMaker demo from AWS a couple years ago. Half the room of intelligent, skilled engineers forgot to shut off the SageMaker instances. And everybody ran out of the $25 of credit they had from the demo—

Corey: In about ten minutes. Yeah.

Levi: In about ten minutes, yeah. And we had to issue all kinds of requests for credits and back and forth. But granted, AWS was accommodating to all of those people, but it was still a lot of stress.

Corey: But it was also slow. They’re very slow on that, which is fair. Like, if someone’s production environment is down, I can see why you care more about that than you do about someone with, “Ah, I did something wrong and lost money.” The counterpoint to that is that for early career folks, that money is everything. We remember earlier this year, that tragic story from the Robinhood customer who committed suicide after getting a notification that he was $730,000 in debt. Turns out it wasn’t even accurate; he didn’t owe anything when all was said and done.

I can see a scenario in which that happens in the AWS world because of their lack of firm price controls on a free tier account. I don’t know what the answer on this is. I’m even okay with a, “Cool you will—this is a special kind of account that we will turn you off at above certain levels.” Fine. Even if you hard cap at the 20 or 50 bucks, yeah, it’s going to annoy some people, but no one is going to do something truly tragic over that. And I can’t believe that Oracle Cloud of all companies is the best shining example of this because you have to affirmatively upgrade your account before they’ll charge you a dime. It’s the right answer.

Levi: It is. And I don’t know if you’ve ever looked at—well, I’m sure you’d have. You’ve probably looked at the solutions provided by AWS for monitoring costs in your accounts, preventing additional spend. Like, the automation to shut things down, right, it’s oftentimes more engineering work to make it so that your systems will shut down automatically when you reach a certain billing threshold than the actual applications that are in place there.

Corey: And I don’t for the life of me understand why things are the way that they are. But here we go. It’s a—[sigh] it just becomes this perpetual strange world. I wish things were better than they are, but they’re not.

Levi: It makes me terribly sad. I mean, I think AWS is an incredible product, I think the ecosystem is great, and the community is phenomenal; everyone is super supportive, and it makes me really sad to be hesitant to recommend people dive into it on their own dime.

Corey: Yeah. And that is a—[sigh] I don’t know how you fix that or square that circle. Because I don’t want to wind up, I really do not want to wind up, I guess, having to give people all these caveats, and then someone posts about a big bill problem on the internet, and all the comments are, “Oh, you should have set up budgets on that.” Yeah, that’s thing still a day behind. So okay, great, instead of having an enormous bill at the end of the month, you just have a really big one two days later.

I don’t think that’s the right answer. I really don’t. And I don’t know how to fix this, but, you know, I’m not the one here who’s a $1.7 trillion company, either, that can probably find a way to fix this. I assure you, the bulk of that money is not coming from a bunch of small accounts that forgot to turn something off or got exploited.

Levi: I haven’t done my 2021 taxes yet, but I’m pretty sure I’m not there either.

Corey: The world in which we live.

Levi: [laugh]. I would love this challenge. I would love to put it out there. If I could, on behalf of, you know, early career people who want to learn—if I could issue credits, if I could spin up sandboxes and say, like, “Here’s an account, I know you’re going to be safe. I have put in a $50 limit.” Right?

Corey: Yeah.

Levi: “You can’t spend more than $50,” like, if I had that control or that power, I would do this in a heartbeat. I’m passionate about getting people these opportunities to play, you know, especially if it’s fun, right? If we can make this thing enjoyable, if we can gamify it, we can play around, I think that’d be great. The experience, though, would be a significant amount of engineering on my side, and then a huge amount of outreach, and that to me makes me really sad.

Corey: I would love to be able to do something like that myself with a, “Look, if you get a bill, they will waive it, or I will cover it.” But then you
wind up with the whole problem of people not operating in good faith as well. Like, “All right, I’m going to mine a bunch of Bitcoin and claim someone else did it.” Or whatnot. And it’s just… like, there are problems with doing this, and the whole structure doesn’t lend itself to that
working super well.

Levi: Exactly. I often say, you know, I face a lot of people who want to talk about mining cryptocurrency in the cloud because I’m a cloud
architect, right? That’s a really common conversation I have with people. And I remind them, like, it’s not economical unless you’re not paying for it.

Corey: Yeah, it’s perfectly economical on someone else’s account.

Levi: Exactly.

Corey: I don’t know why people do things the way that they do, but here we are. So, re:Invent. What did you find that was interesting, promising there, promising but not there yet, et cetera? What was your takeaway from it? Since you had the good sense not to be there in person?

Levi: [laugh]. To me, the biggest letdown was Amplify Studio.

Corey: I thought it was just me. Thank you. I just assumed it was something I wasn’t getting from the explanation that they gave. Because what I heard was, “You can drag and drop, basically, a front end web app together and then tie it together with APIs on the back end.” Which is exactly what I want, like Retool does; that’s what I want only I want it to be native. I don’t think it’s that.

Levi: Right. I want the experience I already have of operating the cloud, knowing the security posture, knowing the way that my users access it, knowing that it’s backed by Amazon, and all of their progressively improving services, right? You say it all the time. Your service running on Amazon is better today than it was two years ago. It was better than it was five years ago. I want that experience. But I don’t think Amplify Studio delivered.

Corey: I wish it had. And maybe it will, in the fullness of time. Again, AWS services do not get worse as they age they get better.

Levi: Some gets stale, though.

Corey: Yeah. The worst case scenario is they sit there and don’t ever improve.

Levi: Right. I thought the releases from S3 in terms of, like, the intelligent tiering, were phenomenal. I would love to see everybody turn on intelligent tiering with instant access. Those things to me were showing me that they’re thinking about the problem the right way. I think we’re missing a story of, like, how do we go from where we’re at today—you know, if I’ve got trillions of objects in storage, how do I transition into that new world where I get the tiering automatically? I’m sure we’ll see blog posts about people telling us; that’s what the community is great for.

Corey: Yeah, they explain these things in a way that the official docs for some reason fail to.

Levi: Right. And why don’t—

Corey: Then again, it’s also—I think—I think it’s because the people that are building these things are too close to the thing themselves. They don’t know what it’s like to look at it through fresh eyes.

Levi: Exactly. They’re often starting from a blank slate, or from a greenfield perspective. There’s not enough thought—or maybe there’s a lot of thought to it, but there’s not enough communication coming out of Amazon, like, here’s how you transition. We saw that with Control Tower, we saw that with some of the releases around API Gateway. There’s no story for transitioning from existing services to these new offerings. And I would love to see—and maybe Amazon needs a re:Invent Echo, where it’s like, okay, here’s all the new releases from re:Invent and here’s how you apply them to existing infrastructure, existing environments.

Corey: So, what’s next for you? What are you looking at that’s exciting and fun, and something that you want to spend your time chasing?

Levi: I spend a lot of my time following AWS releases, looking at the new things coming out. I spend a lot of energy thinking about how do we bring new engineers into the space. I've worked with a lot of operations teams—those people who run playbooks, they hop on machines, they do the old sysadmin work, right—I want to bring those people into the modern world of cloud. I want them to have the skills, the empowerment to know what’s available in terms of services and in terms of capabilities, and then start to ask, “Why are we not doing it that way?” Or start looking at making plans for how do we get there.

Corey: Levi, I really want to thank you for taking the time to speak with me. If people want to learn more. Where can they find you?

Levi: I’m on Twitter. My Twitter handle is @levi_mccormick. Reach out, I’m always willing to help people. I mentor people, I guide people, so if you reach out, I will respond. That’s a passion of mine, and I truly love it.

Corey: And we’ll of course, include a link to that in the [show notes 00:32:28]. Thank you so much for being so generous with your time. I appreciate it.

Levi: Thanks, Corey. It’s been awesome.

Corey: Levi McCormick, cloud architect at Jamf. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with a comment telling me that service ownership is overrated because you are the storage person, and by God, you will die as that storage person, potentially in poverty.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Aaron

I am a Cloud Focused Product Management and Technical Product Ownership Consultant. I have worked on several Cloud Products & Services including resale, management & governance, cost optimisation, platform management, SaaS, PaaS. I am also recognised as a AWS Community Builder due to my work building cloud communities cross-government in the UK over the last 3 years.

I have extensive commercial experience dealing with Cloud Service Providers including AWS, Azure, GCP & UKCloud. I was the Single Point of Contact for Cloud at the UK Home Office and was the business representative for the Home Office's £120m contract with AWS. I have been involved in contract negotiation, supplier relationship management & financial planning such as business cases & cost management.

I run a IT Consultancy called Embue, specialising in Agile, Cloud & DevOps consulting, coaching and training.

Links:

  • Twitter: https://twitter.com/AaronBoothUK
  • LinkedIn: https://www.linkedin.com/in/aaronboothuk/
  • Embue: https://embue.co.uk
  • Publicgood.cloud: https://publicgood.cloud

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: It seems like there is a new security breach every day. Are you confident that an old SSH key, or a shared admin account, isn’t going to come back and bite you? If not, check out Teleport. Teleport is the easiest, most secure way to access all of your infrastructure. The open-source Teleport Access Plane consolidates everything you need for secure access to your Linux and Windows servers, and I assure you there is no third option there. Kubernetes clusters, databases, and internal applications like AWS Management Console, Yankins, GitLab, Grafana, Jupyter Notebooks, and more. Teleport’s unique approach is not only more secure, it also improves developer productivity. To learn more visit: goteleport.com. And not, that is not me telling you to go away, it is: goteleport.com.

Corey: This episode is sponsored in part by our friends at Rising Cloud, which I hadn’t heard of before, but they’re doing something vaguely interesting here. They are using AI, which is usually where my eyes glaze over and I lose attention, but they’re using it to help developers be more efficient by reducing repetitive tasks. So, the idea being that you can run stateless things without having to worry about scaling, placement, et cetera, and the rest. They claim significant cost savings, and they’re able to wind up taking what you’re running as it is in AWS with no changes, and run it inside of their data centers that span multiple regions. I’m somewhat skeptical, but their customers seem to really like them, so that’s one of those areas where I really have a hard time being too snarky about it because when you solve a customer’s problem and they get out there in public and say, “We’re solving a problem,” it’s very hard to snark about that. Multus Medical, Construx.ai and Stax have seen significant results by using them. And it’s worth exploring. So, if you’re looking for a smarter, faster, cheaper alternative to EC2, Lambda, or batch, consider checking them out. Visit risingcloud.com/benefits. That’s risingcloud.com/benefits, and be sure to tell them that I said you because watching people wince when you mention my name is one of the guilty pleasures of listening to this podcast.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. So, when I went to re:Invent last year, I discovered a whole bunch of things I honestly was a little surprised to discover. One of those things is my guest today, Aaron Booth, who’s a cloud consultant with an emphasis on sustainability. Now, you see a number of consultants at things like re:Invent, but what made Aaron interesting was that this was apparently his first time visiting the United States, and he started with not just Las Vegas, but Las Vegas to attend re:Invent. Aaron, thank you for joining me, and honestly, I’m a little surprised you survived.

Aaron: Yeah, I think one of the things about going to Las Vegas or Nevada is no one really prepared me for how dry it was. I ended up walking out of re:Invent with my fingers, like, bleeding, and everything else. And there was so much about America that I didn’t expect, but that was one thing I wish somebody had warned me about. But yeah, it was my first time in the US, first time at re:Invent, and I really enjoyed it. It was probably the best investment in myself and my business that I think I’ve done so far.

Corey: It’s always strange to look at a place that you live and realize, oh, yeah, this is far away for someone else. What would their experience be of coming and learning about the culture we have here? And then you go to Las Vegas, and it’s easy to forget there are people who live there. And even the people who live there do not live on the strip, in the casinos, at loud, obnoxious cloud conferences. So, it feels like it’s one of those ideas of oh, I’m going to go to a movie for the first time and then watching something surreal, like Memento or whatnot, that leaves everyone very confused. Like, “Is this what movies are like?” “Well, this one, but no others are quite like that.” And I feel that way about Las Vegas and re:Invent, simultaneously.

Aaron: I mean, talking about movies, before it came to the US and before I came to Vegas, I was like, “Oh, how can I prepare myself for this trip?” I ended up watching Fear and Loathing in Las Vegas. And I don’t know if you ever seen it, with Johnny Depp, but it’s probably not the best representation, or the most modern representation what Vegas would be like. And I think halfway through the conference, went down to Fremont Street in the old downtown. And they have this massive, kind of, free block screen in the sky that is lit up and doing all these animations. And you’re just thinking, “What world am I on?” And it kind of is interesting as well, from a point of view of, we’re at this tech conference; it’s in Vegas; what is the reason for that? And there’s obviously lots of different things. We want people to have fun, but you know, it is an interesting place to put 30,000 people, especially during a pandemic.

Corey: It really is. I imagine it’s going to have to stay there because in a couple more years, you’re going to need a three block long screen just to list all of the various services that AWS offers because they don’t believe in turning anything off. Now, it would be remiss for me not to ask you, what was announced at re:Invent that got you the most, let’s call it excited, I guess? What got you enthusiastic? What are you happy to start working with more?

Aaron: I think from my perspective, there’s a few different announcements. The first one that comes to mind is the stuff of AWS Amplify Studio, and that’s taken this, kind of, no-code Figma designs and turn into a working front end. And it’s really interesting for me to think about, okay, what is the point of cloud? Why are we moving forward in the world, especially in technology? And, you know, abstracting a lot of stuff we worry about today to simple drag-and-drop tools is probably going to be the next big thing for most of the world.

You know, we’ve come from a privileged position in the West where we follow technology along the whole of the journey, where now we have an opportunity to open this out to many more regions, and many more AWS customers, for example. But for me, as a small business owner—I’ve run multiple businesses—there’s a lot of effort you put into, okay, I need to set up a business, and a website, and newsletter, or whatever else. But the more you can just turn that into, “I’ve got an idea, and I can give it to people with one click,” you’ll enable a lot more business and a lot more future customers as well.

Corey: I was very excited about that one, too, just from a perspective of I want to drag and drop something together to make a fairly crappy web app, that sounds like the thing that I could use to do that. No, that feels a lot more like what Honeycode is trying to be, as opposed to the Amplify side of the world, which is still very focused on React. Which, okay, that makes sense. There’s a lot of front end developers out there, and if you’re trying to get into tech today and are asking what language should I learn, I would be very hard-pressed to advise you pick anything that isn’t JavaScript because it is front end, it is back end, it runs slash eats the world. And I’ve just never understood it. It does not work the way that I think about computers because I’m old and grumpy. I have high hopes of where it might go, but so far I’m looking at it’s [sigh] it’s not what I want it to be, yet. And maybe that’s just because I’m weird.

Aaron: Well, I mean, you know, you mentioned part of the problem really is two different competing AWS services themselves, which with a business like AWS and their product strategy being the word, “Yes,” you know, you’re never really going to get a lot of focus or forward direction with certain products. And hopefully, there’ll be the next, no-code tool announced in re:Invent in a few years’ time, which is exactly what we’re looking for, and gives startup founders or small businesses drag-and-drop tools. But for now, there’s going to be a lot of competing services.

Corey: There’s so much out there that it’s almost impossible to wind up contextualizing re:Invent as a single event. It feels like it’s too easy to step back and say, “Oh, okay. I’m here to build websites”—is what we’re talking about now in the context of Amplify—and then they start talking about mainframes. And then they start talking about RoboRunner to control 10,000 robots at once. And I’m looking around going, “I don’t have problems that feel a lot like that. What’s the deal?”

Aaron: I think even just, like you said in perspective of re:Invent is like, when you go to an event like this, that you can’t experience everything and you probably have a very specific focus of, you know, what am I here to do. And I was really surprised—again, my first time at a big tech conference, as well as Vegas and the US is, how important it was just to meet people and how valuable that was. First time I met you, and you know, going from somebody who’s probably very likely interacted with you on Twitter before the event to being on this podcast and having a great conversation now is kind of crazy to think that the value you can get out of it. I mean, in terms of over services, and areas of re:Invent that I found interesting was the announcement of the new sustainability pillar, as part of the well-architected framework. You know, I’ve tried to use that before in previous workplaces, and it has been useful. You know, I’m hoping it is more useful in the future, and the cynical part of me worries about whether the whole point of putting this as part of a well-architected framework review where the customer is supposed to do it is Amazon passing the buck for sustainability. But it’s an interesting way forward for what we care about.

Corey: An interesting quirk of re:Invent—to me—has always been that despite there being tens of thousands of people there are always a few folks that you wind up running into again and again and again throughout the week. One year for me it was Ben Kehoe; this trip it was you where we kept finding ourselves at the same events, we kept finding ourselves at the same restaurants, and we had three or four meals together as a result, and it was a blast talking to you. And I was definitely noticing that sustainability was a topic that you kept going back to a bunch of different ways. I mean previously, before starting your current consulting company, you did a lot of work in the government—specifically the UK Government, for those who are having trouble connecting the fact this is the first time in America to the other thing. Like, “Wow, you can be far away and work for the government?” It’s like, we have more than one on this planet, as it turns out.

Yes, it was a fun series of conversations, and I am honestly a little less cynical about the idea of the sustainability pillar, in no small part due to the conversations that we had together. I initially had the cynical perspective of here’s how to make your cloud infrastructure more sustainable. It’s, isn’t that really a “you” problem? You’re the cloud provider. I can’t control how you get energy on the markets, how you wind up handling heat issues, how you address water issues from your data center outflows, et cetera. It seems to me that the only thing I can really do is use the services you give me, and then it becomes a “you” problem. You have a more nuanced take on it.

Aaron: I think there’s a log of different things to think about when it comes to sustainability. One of the main ones is, from my perspective, you know, I worked at the UK Home Office in the UK, and we’d been using cloud for about six or seven years. And just looking at how we use clouds as an enterprise organization, one of the things I really started to see was these different generations of cloud and you’ve got aspects of legacy infrastructure, almost, that we lifted-and-shifted in the early days, versus maybe stuff would run on serverless now. And you know, that’s one element, from a customer is how you control your energy usage is actually the use of servers, how efficient your code is, and there’s definitely a difference between stringing together EC2 and S3 buckets compared to using serverless or Lambda functions.

Corey: There’s also a question of scale. When I’m trying to build something out of Lambda functions, and okay, which region is the most cost effective way to run this thing? The Google search for that will have a larger climate impact than any decision I can make at the scale that I operate at. Whereas if you’re a company running tens of thousands of instances at any given point in time and your massive scale, then yeah, the choices you make are going to have significant impact. I think that a problem AWS has always struggled with has been articulating who needs to care about what, when.

If you go down the best practices for security and governance and follow the white papers, they put out as a one-person startup trying to build an idea this evening, just to see if it’s viable, you’re never going to get anywhere. If you ignore all those things, and now you’re about to go public as a bank, you’re going to have a bad time, but at what point do you have to start caring about these different things in different ways? And I don’t think we know the answer yet, from a sustainability perspective.

Aaron: I think it’s interesting in some senses, that sustainability is only just enter the conversation when it comes to stuff we care about in businesses and enterprises. You know, we all know about risk registers, and security reviews, and all those things, but sustainability, while we’ve, kind of, maybe said nice public statements, and put things on our website, it’s not really been a thing that’s, okay, this is how we’re going to run our business, and the thing we care about as number one. You know, Amazon always says security is job zero, but maybe one day someone will be saying sustainability is our job zero. And especially when it comes down to, sort of, you know, the ethics of running a business and how you want that to be run, whether it is going to be a capitalistic VC-funded venture to extract wealth from citizens and become a billionaire versus creating something that’s a bit more circular, and gives back as sustainability might be a key element of what you care about when you make decisions.

Corey: The challenge that I find as well is, I don’t know how you can talk about the relative sustainability impact of various cloud services within the AWS umbrella without, effectively, AWS explaining to you what their margins are on different services, in many respects. Power usage is the primary driver of this and that determines the cost of running things. It is very clear that it is less expensive and more efficient to run more modern hardware than older hardware, so we start seeing, okay, wow, if I start seeing those breakdowns, what does that say about the margin on some of these products and services? And I don’t think they want to give that level of transparency into their business, just because as soon as someone finds out just how profitable Managed NAT gateways are, my God, everything explodes.

Aaron: I think it’s interesting from a cloud provider or hyperscaler perspective, as well, is, you know, what is your USP? And I think Amazon is definitely not saying sustainability is their USP right now, and I think you know, there are other cloud providers, like Azure for example, who basically can provide you a Power BI plugin; if you just log in with your Cloud account details, it will show you a sustainability dashboard and give you more of this information that you might be looking for, whereas Amazon currently doesn’t offer anything like that automated. And even having conversations with your account team or trying to get hold of the right person, Amazon isn’t going to go anywhere at the moment, just because maybe that’s the reason why we don’t want to talk about it: It’s too sensitive. I’m sure that’ll change because of the public statements they’ve made at re:Invent now and previously of, you know, where they’re going in terms of energy usage. They want to be carbon neutral by 2025, so maybe it’ll change to next re:Invent, we’ll get the AWS Sustainability Explorer add-on for [unintelligible 00:15:23] or 12—

Corey: Oh no.

Aaron: —tools to do the same thing [laugh].

Corey: In the Google Cloud Console, you click around, and there are green leafs next to some services and some regions, and it’s, on the one hand, okay, I appreciate the attention that is coming from. On the other hand, it feels like you’re shaming me for putting things in a region that I’ve already built things out in when there weren’t these green leafs here, and I don’t know that I necessarily want to have that conversation with my entire team because we can’t necessarily migrate at this point. And let’s also be clear, here, I cannot fathom a scenario in which running your own data centers is ever going to be more climate-friendly than picking a hyperscaler.

Aaron: And I think that’s sort of, you know, we all might think about is, at the end of the day, if your sustainability strategy for your business is to go all-in-on cloud, and bet horse on AWS or another cloud provider, then, at the end of the day, that’s going to be viable. I know, from the, sort of, hands-on stuff I’ve done with our own data centers, you can never get it as efficient as what some of these cloud providers are doing. And I mean, look at Microsoft. The fact that they’re putting some of their data centers under the sea to use that as a cooling mechanism, and kind of all the interesting things that they’re able to do because they can invest at scale, you’re never going to be able to do that with the cupboard beyond the desks in your local office to make it more efficient or sustainable.

Corey: There are definite parallels between Cloud economics and sustainability because as mentioned, I worship at the altar of Our Lady of Turn that Shit Off because that’s important. If you don’t have a workload running and it doesn’t exist, it has no climate impact. Mostly. I’m sure there are corner cases. But that does lead to the question then of okay, what is the climate sustainability impact, for example, of storing a petabyte of data and EBS versus in S3?

And that has architectural impact as well, and there’s also questions of how often does it move because when you move it, Lord knows there is nothing more dear than the price of data transfer for data movement. And in order to answer those questions, they’re going to start talking a lot more about their architecture. I believe that is why Peter DeSantis’s keynote talked so much about—finally—the admission of what we sort of known for ages now that they use erasure coding to make S3 as durable yet inexpensive, as it is. That was super interesting. Without that disclosure, it would have been pretty clear as soon as they start publishing sustainability numbers around things like that.

Aaron: And I think is really interesting, you know, when you look at your business and make decisions like that. I think the first thing to start with is do you need that data at all? What’s a petabyte of data are going to do? Unless it’s for serious compliance reasons for, you know, the sector or the business that you’re doing, the rest of it is, you know, you’ve got to wonder how long is that relevant for. And you know, even as individuals, we could delete junk mail and take things off our internal emails, it’s the same thing of businesses, what you’re doing with this data.

But it is interesting, when you look at some of the specific services, even just the tiering of S3, for example, put that into Glacier instead of keeping it on S3 general. And I think you’ve talked about this before, I think cost the same to transfer something in and out of Glacier as just to hold it for a month. So, at the end of the day, you’ve got to make these decisions in the right way, and you know, with the right goals in mind, and if you’re not able to make these decisions or you need help, then that’s where, you know, people like us come in to help you do this.

Corey: There’s also the idea of—when I was growing up, the thing they always told us about being responsible was, “Oh, turn out the lights when you’re not in the room.” Great. Well, cloud economics starts to get in that direction, too. If you have a job that fires off once a day at two in the morning and it stops at four in the morning, you should not be running those instances the other 22 hours of the day. What’s the deal here?

And that becomes an interesting expiratory area just as far as starting to wonder, okay, so you’re telling me that if I’m environmentally friendly, I’m also going to save money? Let’s be clear people, in many cases—in a corporate sense—care about sustainability only insofar as that don’t get yelled out about it. But when it comes to saving money, well, now you’ve got the power of self-interest working for you. And if you can dress them both up and do the exact same things and have two reasons to do it. That feels like it could in some respects, be an accelerator towards achieving both outcomes.

Aaron: Definitely. I think, you know, at the end of the day, we all want to work on things that are going to hopefully make the world a better place. And if you use that as a way of motivating, not just yourself as a business, but the workforce and the people that you want to work for you, then that is a really great goal as well. And I think you just got to look at companies that are in this world and not doing very great things that maybe they end up paying more for engineers. I think I read an interesting article the other day about Facebook is basically offering almost double or 150 percent of over salaries because it feels like a black mark on the soul to work for that company. And if there is anything—maybe it’s not greenwashing per se, but if you can just make your business a better place, then that could be something that you can hopefully attract other like-minded people with.

Corey: This episode is sponsored by our friends at Oracle Cloud. Counting the pennies, but still dreaming of deploying apps instead of, “Hello World” demos? Allow me to introduce you to Oracle’s Always Free tier. It provides over 20 free services and infrastructure, networking, databases, observability, management, and security. And let me be clear here, it’s actually free. There’s no surprise billing until you intentionally and proactively upgrade your account. This means you can provision a virtual machine instance or spin up an autonomous database that manages itself all while gaining the networking, load balancing, and storage resources that somehow never quite make it into most free tiers needed to support the application that you want to build. With Always Free, you can do things like run small-scale applications, or do proof-of-concept testing without spending a dime. You know that I always like to put asterisks next to the word free. This is actually free, no asterisk. Start now. Visit snark.cloud/oci-free that’s snark.cloud/oci-free.

Corey: One would really like to hope that the challenge, of course, is getting there in such a way that it, well, I guess makes sense, is probably the best way to frame it. These are still early days, and we don’t know how things are going to wind up… I guess, it playing out. I have hopes, I have theories, but I just don’t know.

Aaron: I mean, even looking at Cloud as a concept, how long we’ve all worked with this now ranges probably from fifteen to five, and for me the last six years, but you got to think looking at the outages at the end of last year at Amazon, that [unintelligible 00:21:57], very close to re:Invent, that impacted a lot of different workloads, not just if you were hosted in us-west or east-1, but actually for a lot of the regional services that actually were [laugh]… discovered to be kind of integral to these regions. You know, one AZ going down can impact single-sign-on logins around the world. And let’s see what Amazon looks like in ten years’ time as well because it could be very different.

Corey: Do you find that as you talk to folks, both in government and in private sector, that there is a legitimate interest in the sustainability story? Or is it the self-serving cynical perspective that I’ve painted?

Aaron: I mean, a lot of my experience is biased towards the public sector, so I’ll start with that. In terms of the public sector, over the last few years, especially in the UK, there’s been a lot more focus on sustainability as part of your business cases and your project plans for when
you’re making new services or building new things. And one of the things they’ve recently asked every government department in the UK to do is come up with a sustainability strategy for their technology. And that’s been something that a lot of people have been working on as part of something called the One Gov Cloud Strategy Working Groups—which in the UK, we do love an abbreviation, so [laugh] a bit of a long name—but I think there’s definitely more of an interest in it.

In terms of the private sector, I’m not too sure if that’s something that people are prioritizing. A lot of the focus I kind of come across as either, we want to focus on enterprise customers, so we’re going to offer migration professional services, or you’re a new business and you’re starting to go up and already spending a couple a hundred pounds, or thousands of pounds a month. And at that scale, it’s probably not going to be something you need to worry about right now.

Corey: I want to talk a little bit about how you got into tech in the first place because you told me elements of this story, and I generally find them to be—how do I put this?—they strain the bounds of credulity. So, how did you wind up in this ridiculous industry?

Aaron: I mean, hoping as I explain them, you don’t just think I’m a liar. I have got a Scouse accent, so you’re probably predisposed towards it. But my journey into tech was quite weird, I guess, in the sense that when I was 16—I was, again, like I said, born in Liverpool and didn’t really know what I wanted to do in the world, and had no idea what the hell to do. So, I was at college, and kind of what happened to me there is I joined, like, an entrepreneurship club and was like, “Okay, I’ll start my own business and do something interesting.” And I went to a conference at college, and there was a panel with Richard Branson and other few of business leaders, and I stood up and asked the question said, you know, “I’m 16. I want to start a business. Where can I get money to start a business?”

And the panel answered with kind of a couple of different things, but one of them was, “Get a job.” The other one was, “Get money off your parents.” And I was kind of like, “Oh, a bit weird. I’ve got a job already. You know, I would ask my parents put their own benefits.”

And asked the woman with the microphone, “Can I say something back?” And she said, “No.” So, being… a young person, I guess, and just I stood back up and said, you know, “You’re in Liverpool. You’ve kind of come to one of the poorest cities in some sense in the UK, and you kind of—I’ve already got a job. What can I really do?”

And that’s when Richard Branson turned round and said, “Well, what is it you want to do?” And I said, “I make really good cheesecakes and I want to sell them to people.” And after that sort of exchange, he said he’d give me the money. So, he gave me 200 pounds to start my own business. And that was just, kind of like, this whirlwind of what the hell’s going on here?

But for me, it’s one of those moments in my life, which I think back on, and honestly, it’s like one of these ten [left 00:25:15] moments of, you know, I didn’t stand back up and say something, if I didn’t join the entrepreneurship club, like, I just wouldn’t be in the position I am right now. And it was also weird in the sense that I said at the start of the story, I didn’t know what I wanted to do in my life. This was the first time that anyone had ever said to me, “I trust you to do something, and here’s 200 pounds to do it.” And it was such a small thing, and a small moment that basically got me to where I am today. And kind of a condensed version of that is, you know, after that event, I started volunteering for a charity who—a, sort of, magazine launch, and then applied for the civil service and progressed through six to eight years of the civil service.

And it was because of that moment, and that experience, and that confidence boost, where I was like, “Oh, I actually can do something with my life.” And I think tech, and I think a lot of people talk about this is, it can be a bit of a crazy whirlwind, and to go from that background into, you know, working with great people and earning great money is a bit of a crazy thing sometimes.

Corey: Is there another path that you might have gone down instead and completely missed out on, for lack of a better term—and not missed out. You probably would have been far happier not working in tech; I know I would have been—but as far as trying to figure out, like, what does the road not taken look like for you?

Aaron: I’m not too sure, really. And at the time, I was working in a club. I was like 16, 17 years old, working in a nightclub in Liverpool for five pounds an hour, and was doing that while I was studying, and that was almost like, what was in my mind at the time. When it came to the end of college, I was applying for universities, I got in on, like, a second backup course, and that was the only thing to do was food science. And it was like, I can't imagine coming out of university three years after that, studying something that’s not really that relevant to a lot of industries, and trying to find a good job. It could have just been that I was working in a supermarket for minimum wage after I came out for uni trying to find what I wanted to do in the world. And, yeah, I’m really glad that I kind of ended up where I am now.

Corey: As you take a look at what you want your career to be about in the broad sweep of things, what is it that drives you? What is it that makes you, for example, decide to spend the previous portion of career working in public service? That is a very, shall we say, atypical path—I say, as someone who lives in San Francisco and is surrounded by people who want to make the world a better place, but all those paths just
coincidentally would result in them also becoming billionaires along the way.

Aaron: I mean, it is interesting. You know, one of the things that worked for the civil service for so long, is the fact that I did want to do more than just make somebody else more money. And you know, there are not really a lot of ways you can do that and make a good wage for yourself. And I think early on in your career, working for somewhere like the civil service or federal government can be a little bit of that opportunity. And especially with some of the government’s focus on tech these days, and investments—you know, I joined through an apprenticeship scheme and then progressed on to a digital leadership scheme, you know, they were guided schemes to help me become a better leader and improve my skills.

And I think I would have probably not gone to the same position if I just got the tech job or my first engineering job somewhere else. I think, if I was to look at the future and where do I want to go, what do I care about? And, you know, you ask me, sort of, this question at re:Invent, and it took me a few days to really figure out, but one of the things when I talk about making the world a better place is thinking about how you can start businesses that give back to people in local areas, or kind of solve problems and kind of keep itself running a bit like a trust does, [laugh], if only that keeping rich people running. And a lot of the time, like, you’ve highlighted is coincidentally these things that we try and solve whether it’s, like, a new app or a new thing that does something seems to either be making money for VCs, reinventing things that we already have, or just trying to make people billionaires rather than trying to make everyone rise up and—high tide rise all ships, is the saying. And there are a few people that do this, a few CEOs who take salaries the same as everyone else in the business. And I think that’s hopefully you know, as I grow my own business and work on different things in the future, is how can I just help people live better lives?

Corey: It’s a big question, and it’s odd in that I don’t find that most people asking it tend to find themselves going toward government work so much as they do NGOs, and nonprofits, and things that are very focused on specific things.

Aaron: And it can be frustrating in some sense is that, you know, you look at the landscape of NGOs, and charities, and go, “Why are they involved in solving this problem?” You know, one of the big problems we have in the UK is the use of food banks where people who don’t have enough money, whether they receive benefits or not, have to go and get food which is donated just by people of the UK and people who donate to these charities. You know, at the end of the day, I’m really interested in government, and public sector work, and potentially one day, being a bit more involved in policy elements of that, is how can we solve these problems with broad brushstrokes, whether it’s technology advancements, or kind of policy decisions? And one of the interesting things that I got close to a few times, but I don’t think we’ve ever really solved is stuff like how can we use Agile to build policy?

How can we iterate on what that policy might look like, get customers or citizens of countries involved in those conversations, and measure outcomes, and see whether it’s successful afterwards. And a lot of the time, policies and decisions are just things that come out of politicians minds, and it’d be interesting to see how we can solve some of these problems in the world with stuff like Agile methodologies or tech practices.

Corey: So, it’s easy to sit and talk about these things in the grand sweep of how the world could be or how it should look, but for those of us who think in more, I guess, tactical terms, what’s a good first step?

Aaron: I think from my point of view, and you know, meeting so many people at re:Invent, and just have my eyes opened of these great conversations we can have a great people and get things changed, one of the things that I’m looking at starting next year is a podcast and a newsletter, around the use of public cloud for public good. And when I say that, it does cover elements of sustainability, but it is other stuff like how do we use Cloud to deliver things in the public sector and NGOs and charities? And I think having more conversations like that would be really interesting. Obviously, that’s just the start of a conversation, and I’m sure when I speak to more people in the future, more opportunities and more things might come out of it. But I’d just love to speak to more people about stuff like this.

Corey: I want to thank you for spending so much time to speak with me today about… well, the wide variety of things, and of course, spending as much time as you did chatting with me at re:Invent in person. If people want to learn more, where can they find you?

Aaron: So yep, got a few social media handles on Twitter, I’m @AaronBoothUK. On LinkedIn is the same, forward slash aaronboothuk, and I’ve also got the website for my consultancy, which is embue.co.uk—E-M-B-U-E dot co dot uk. And for the newsletter, it’s publicgood.cloud.

Corey: And we will, of course, include links to that in the [show notes 00:32:11]. Thank you so much for taking the time to speak with me. I really do appreciate it.

Aaron: Thank you so much for having me.

Corey: Aaron Booth, cloud consultant with an emphasis on sustainability. I’m Cloud Economist Corey Quinn with an emphasis on optimizing bills. And this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry comment that you will then kickstart the coal-burning generator under your desk to wind up posting.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Ell

Ell, former SysAdmin, cloud builder, podcaster, and container advocate, has always been a security enthusiast. This enthusiasm and driven curiosity have helped her become an active member of the InfoSec community, leading her to explore the exciting world of Genetic Software Mapping at Intezer.

Links:

  • Intezer: https://www.intezer.com
  • Twitter: https://twitter.com/Ell_o_Punk

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: It seems like there is a new security breach every day. Are you confident that an old SSH key, or a shared admin account, isn’t going to come back and bite you? If not, check out Teleport. Teleport is the easiest, most secure way to access all of your infrastructure. The open source Teleport Access Plane consolidates everything you need for secure access to your Linux and Windows servers—and I assure you there is no third option there. Kubernetes clusters, databases, and internal applications like AWS Management Console, Yankins, GitLab, Grafana, Jupyter Notebooks, and more. Teleport’s unique approach is not only more secure, it also improves developer productivity. To learn more visit: goteleport.com. And not, that is not me telling you to go away, it is: goteleport.com.

Corey: This episode is sponsored by our friends at Oracle Cloud. Counting the pennies, but still dreaming of deploying apps instead of "Hello, World" demos? Allow me to introduce you to Oracle's Always Free tier. It provides over 20 free services and infrastructure, networking, databases, observability, management, and security. And—let me be clear here—it's actually free. There's no surprise billing until you intentionally and proactively upgrade your account. This means you can provision a virtual machine instance or spin up an autonomous database that manages itself all while gaining the networking load, balancing and storage resources that somehow never quite make it into most free tiers needed to support the application that you want to build. With Always Free, you can do things like run small scale applications or do proof-of-concept testing without spending a dime. You know that I always like to put asterisks next to the word free. This is actually free, no asterisk. Start now. Visit snark.cloud/oci-free that's snark.cloud/oci-free.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. If there’s one thing we love doing in the world of cloud, it’s forgetting security until the very end, going back and bolting it on as if we intended to do it that way all along. That’s why AWS says security is job zero because they didn’t want to remember all of their slides once they realized they forgot security. Here to talk with me about that today is Ell Marquez, security research advocate at Intezer. Ell, thank you for joining me.

Ell: Of course.

Corey: So, what does a security research advocate do, for lack of a better question, I suppose? Because honestly, you look at that, it’s like, security research advocate, it seems, would advocate for doing security research. That seems like a good thing to do. I agree, but there’s probably a bit more nuance to it, then I can pick up just by the [unintelligible 00:01:17] reading of the title.

Ell: You know, we have all of these white papers that you end up getting, the pen test reports that are dropped on your desk that nobody ever gets to, they become low priority, my job is to actually advocate that you do something with the information that you get. And part of that just involves translating that into plain English, so anyone can go with it.

Corey: I’ve got to say, if you want to give the secrets of the universe and make sure that no one ever reads them, make sure that it has a whole bunch of academic-style citations at the beginning, and ideally put it behind some academic paywall, and it feels like people will claim to have read it but never actually read the thing.

Ell: Don’t forget charts.

Corey: Oh yes, with the charts. In varying shades of blue. Apparently that’s the only color you’re allowed to do some of these charts in; despite having a full universe of color palettes out there, we’re just going to put it in varying shades of corporate blue and hope that people read it.

Ell: Yep, that sounds about security there. [laugh].

Corey: So, how much of, I guess, modern security research these days is coming out of academia versus coming out of industry?

Ell: In my experience in, you know, research I’ve done in researching researchers, it all really revolves around actual practitioners these days, people who are on the front lines, you know, monitoring their honey pots, and actually reporting back on what they’re seeing, not just theoretical.

Corey: Which I guess brings us to the question of, I wind up watching all of the keynotes that all the big cloud providers put on and they simultaneously pat me on the head and tell me that their side of security is just fine with their shared responsibility model and the rest, whereas all of the breaches I’m ever going to deal with and the only way anyone can ever see my data is if I make a mistake in configuring something. And honestly, does that really sound like something I would do? Probably not, but let’s face it, they claim that they are more or less infallible. How accurate is that?

Ell: I wish that I could find the original person that said this, but I’ve heard it so many times. And it’s actually the ‘cloud irresponsibility model.’ We have this blind faith that if we’re paying somebody for it, it’s going to be done correctly. I think you may have seen this with billing. How many people are paying for redundant security services with a cloud provider?

Corey: I’ve once—well, more than once have noticed that if you were to configure every AWS security service that they have and enable it in your account, that the resulting bill would be larger than the cost of the data breach it was preventing. So, on some level, there is a point at which it just becomes ridiculous and it’s not necessarily worth pursuing further. I honestly used to think that the shared responsibility model story was a sales pitch, and then I grew ever more cynical. And now my position on it is that it’s because if you get breached, it’s your fault is what they’re trying to say. But if you say it outright to someone who just got breached, they’re probably not going to give you money anymore. So, you need to wrap that in this whole involved 45-minute presentation with slides, and charts, and images and the rest because people can’t refute one of those quite the way that they can a—it’s in a tweet sentence of, “It’s your fault.”

Ell: I kind of have to agree with them in the end that it is your fault. Like, the buck stops with you, regardless. You are the one that chose to trust that cloud provider was going to do everything because your security team might make a mistake, but the cloud provider is made up of humans as well who can make just as many mistakes. At the end of the day, I don’t care what cloud provider you used; I care that my data was compromised.

Corey: One of the things that irks me the most is when I read about a data breach from a vendor that I had either trusted knowingly with my data or worse, never trusted but they somehow scraped it somewhere and then lost it, and they said, “Oh, a third-party contractor that we hired.” It’s, “Yeah, look, I’m doing business with you, ideally, not the people that you choose to do business with in turn. I didn’t select that contractor. You did, you can pass out the work and delegate that. You cannot delegate the responsibility.” So no, Verizon, when you talk about having a third-party contractor have a data breach of customer data, you lost the data by not vetting your contractors appropriately.

Ell: Let’s go back in time to hopefully something everybody remembers: Target. Target being compromised because of their HVAC provider. Yet how many people—you know this is being recorded in the holiday season—are still shopping at Target right now? I don’t know if people forget or they just don’t care.

Corey: A year later, their stock price was higher than it was before the breach. Sure they had a complete turnover of their C-suite at that point; their CSO and CEO were forced out as a result, but life went on. And they continue to remain a going concern despite quite literally having a bull’s eye painted on the building. You’d think that would be a metaphor for security issues. But no, no, that is something they actually do.

Ell: You know, when you talk about, you know, the CEO being let go or, you know, being run out, but what part did he honestly have to do with it? They’re talking about, oh, well, they made the decisions and they were responsible. What because they got that, you know, list of just 8000 papers with the charts on it?

Corey: As I take a look at a lot of the previous issues that we’ve seen with I’ve been doing my whole S3 Bucket Negligence Awards for a while, but once I actually had a bucket engraved and sent to a company years ago, the Pokémon Company, based upon a story that I read in the Wall Street Journal, how they declined to do business with a prospective vendor because going through their onboarding process, they noticed among other things, insufficient security controls around a whole bunch of things including S3 buckets, and it’s holy crap, a company actually making a meaningful decision based upon security. And say what you will about the Pokémon Company, their audience is—at least theoretically—children and occasionally adults who believe they’re children—great, not here to shame—but they understand that this is not
something you can afford to be lax in and they kiboshed the entire deal. They didn’t name the vendor, obviously, but that really took me aback. It was such a rarity to see that, and it’s why I unfortunately haven't had to make a bucket like that since. I wish I did. I wish more
companies did things like this. But no it’s just a matter of, well, we claim to do the right thing, and we checked all the boxes and called it good, and oops, these things happen.

Ell: Yes, but even when it goes that way, who actually remembers what happened, and did you ever follow up if there were any consequences to not going, “Okay, third-party. You screwed up, we’re out. We’re not using you.” I can’t name a single time that happened.

Corey: Over at The Duckbill Group, we have large enterprise customers. We have to be respectful and careful with their data, let’s be very clear here. We have all of their AWS billing data going back for some fixed period of time. And it worries me what happens if that data gets breached. Now, sure, I’ve done the standard PR crisis comms thing, I have statements and actions prepared to go in the event that it happens, but I’m also taking great pains to make sure it doesn’t.

It’s the idea of okay, let’s make sure that we wind up keeping these things not just distinct from the outside world, but distinct from individual clients so we’re not mixing and matching any of this stuff. It’s one of those areas where if we wind up having a breach, it’s not because we didn’t follow the baseline building blocks of doing this right. It’s something that goes far beyond what we would typically expect to see in an environment like this. This, of course, sets aside the fact that while a breach like that would be embarrassing, it isn’t actually material to anyone’s business. This is not to say that I’m not taking it seriously because we have contractual provisions that we will not disclose a lot of this stuff, but it does not mean the end of someone’s business if this stuff were to go public in the same way that, for example, back when I worked at Grindr many years ago, in the event that someone’s data had been leaked there, people could theoretically been killed. There’s a spectrum of consequences here, but it still seems like you just do the basic block-and-tackling to make sure that this stuff isn’t publicly exposed, then you start worrying about the more advanced stuff. But with all these breaches, it seems like people don’t even do that.

Ell: You have Tesla, right, who’s working on going to Mars, sending people there who had their S3 buckets compromised. At that point, if we’ve got this technology, just giant there, I think we’re safe to do that whole, “Hey, assume breach, assume compromise.” But when I say that, it drives me up the wall how many people just go, “Okay, well, there’s nothing we can do. We should just assume that there’s going to be an issue,” and just have this mentality where they give up. No, that gives you a starting point to work from, but that’s not the way it’s being seen.

Corey: One of the things that I’ve started doing as I built up my new laptop recently has been all right, how do I work with this in such a way that I don’t have credentials that are going to grant access to things in any long-lived way ever residing on disk? And so that meant with AWS, I started using SSO to log into a bunch of things. It goes through a website, and then it gives a token and the rest that lasts for 12 hours. Great.

Okay, SSH keys, how do I handle that? Historically, I would have them encrypted with a passphrase, but then I found for Mac OS an app called Secretive that stores it in the Secure Enclave. I have to either type in a password or prove it with a biometric Touch ID nonsense every time something tries to access the key. It’s slightly annoying when I’m checking out five or six Git repos at once, but it also means that nothing that I happen to have compromised in a browser or whatnot is going to be able to just grab the keys, send it off somewhere, and then I’ll never realize that I’ve been compromised throughout. It’s the idea of at least theoretically defense in depth because it’s me, it’s my personal electronics, in all likelihood, that are going to be compromised, more so than it is configured, locked-down S3 buckets, managed properly. And if not me, someone else in my company who has access to these things.

Ell: I’m going to give you the best advice you’re ever going to get, and people are going to go, “Duh,” but it’s happening right now: Don’t get complacent, don’t get lazy, how many of us are, “Okay, we’re just going to put the key over here for a second.” Or, “We’re just going to do this for a minute,” and then we forget. I recently, you know, did some research into Emotet and—you know, the new virus and the group behind it—you know how they got caught? When they were raided, everything was in plain text. They forgot to use their VPN for a while, all the files that they’d gotten no encryption. These were the people that that’s what they were looking for, but you get lazy.

Corey: I’ve started treating at least the security credential side of doing weird things, even one off bash scripts, as if they were in production. I stuff the credentials into something like AWS’s parameter store, and then just have a one line snippet of code that retrieves them at runtime to wind up retrieving those. Would it be easier to just slap it in there in the code? Absolutely, of course it would. But I also look at my newsletter production pipeline, and I count the number of DynamoDB tables that are in active use that are labeled Test or Dev, and I realized, huh, I’m actually kind of bad at taking something that was in Dev and getting it ready for production. Very often, I just throw a load at it and call it good. So, if I never get complacent around things like that, it’s a lot harder for me to get yelled at for checking secrets into Git, for example.

Ell: Probably not the first time that you’ve heard this but, Corey, I’m going to have to go with you’re abnormal because that is not what we’re seeing in a day-to-day production environment.

Corey: Oh, of course not. And the reason I do this is because I was a grumpy old sysadmin for so long, and have gotten burned in so many weird ways of messing things up. And once it’s in Git, it’s eternal—we all know that—and I don’t ever want to be in a scenario where I open-source something and surprise, surprise, come to find out in the first two days of doing something, I had something on disk. It’s just better not to go down that path if at all possible.

Ell: Being a former sysad as well, I must say, what you’re able to do within your environment, your computer is almost impossible within a corporate environment. Because as a sysad, I’m looking at, “What did the devs do again? Oh, man, what’s the security team going to do?” And you’re stuck in the middle trying to figure out how to solve a problem and then manage it through that entire environment.

Corey: I never really understood intrinsically the value of things like single-sign-on, until I wound up starting this company. Because first, it was just me for a few years. And yeah, I can manage my developer environments and my AWS environments in such a way that if they get compromised, it’s not going to be through basic, “Oops, I forgot that’s how computers work,” type of moment. It’s going to be at least something a little bit more difficult, I would imagine. Because if you—all right, if you managed to wind up getting my keys and the passphrase, and in some cases, the MFA device, great, good, congratulations, you’ve done something novel and probably deserve the data.

Whereas as soon as I started bringing other people in who themselves were engineers, I sort of still felt the same way. Okay, we’re all responsible adults here, and by and large, since I wasn’t working with junior people, that held true. And then I started bringing in people who did not come from a deeply computer-y technical background, doing things like finance, and doing things like sales, and doing things like marketing, all of which are themselves deeply technical in their own way, but data privacy and data security are not really something that aligns with that. So, it got into the weeds of, “How do I make sure that people are doing responsible things on their work computers like turning on disk encryption, and forcing a screensaver, and a password and the rest.” And forcing them to at least do some responsible things like having 1Password for everyone was great until I realized a couple people weren’t even using it for something, and oh dear. It becomes a much more difficult problem at scale when you have to deal with people who, you know, have actual work to do rather than sitting around trying to defend the technology against any threat they can imagine.

Ell: In what you just said though, there is one flaw is we tend to focus on, like you said, marketing and finance and all these organizations who—don’t get phished, don’t click on this link. But we kind of give the just the openness that your security team, your sysads, your developers, they’re going to know best practices. And then we focus on Windows because that’s what the researchers are doing. And then we focus on Windows because that’s what marketing is using, that’s what finance is using. So, what there’s no way to compromise a Mac or Linux box? That’s a huge, huge open area that you’re allowing for attackers.

Corey: Let’s be very clear here. We don’t have any Windows boxes—of which I’m aware—in the company. And yeah, the technical folk we have brought in, most of them I’d worked—or at least the early folks—I’d worked with previously. And we had a shared understanding of security. At least we all said the right things.

But yeah, as you—right, as you grow, as you scale, this becomes a big deal. And it’s, I also think there’s something intrinsically flawed about a model where the entire instruction set is, it all falls on you to not click the link or you’re going to doom us all. Maybe if someone can click a link and doom us all, the problem is not with them; it’s the fact that we suck at building secure systems that respect defense in depth.

Ell: Something that we do wrong, though, is we split it up. We have endpoint protection when we’re talking about, you know, our Windows boxes, our Linux boxes, our Mac boxes. And then we have server-side and cloud security. Those connect. Think about, there’s a piece of malware called EvilGNOME. You go in on a Linux box, you have access to my camera, keylogging, and watching exactly what I’m doing. I’m your sysad. I then cat out your SSH keys, I go into your box, they now have the password, but we don’t look for that. We just assume that those two aren’t really that connected, and if we monitor our network and we monitor these devices, we’ll be fine. But we don’t connect the two pieces.

Corey: One thing that I did at a consulting client back in 2012, or so that really raised eyebrows whenever I told people about it was that we wound up going to some considerable trouble building a allow list within Squid—a proxy server that those of us in Linux-land are all too familiar with in some cases—so everything in production could only talk to the outside world via that proxy; it was not allowed to establish any outbound connections other than through that proxy. So, it was at that point only allowed to talk to specify update servers, specified third-party APIs and the rest, so at least in theory, I haven’t checked back on them since, I don’t imagine that the log4yay nonsense that we’ve seen recently would necessarily work there. I mean, sure, you have the arbitrary execution of code—that’s bad—but reaching out to random endpoints on the internet would not have worked from within that environment. And I liked that model, but oh my God, was it a pain in the butt to set up properly because it turns out, even in 2012, just to update a Linux system reasonably, there’s a fair number of things it needs to connect to, from time-to-time, once you have all the things like New Relic instrumentation in, and the app repository you’re talking to, and whatever container source you’re using, and, and, and. Then you wind up looking at challenges like, oh, I don’t know, if you’re looking at an AWS-style environment, like most modern things are, okay, we’re only going to allow it to talk to AWS endpoints. Well, that’s kind of the entire internet now. The goalposts move, the rules change, the game marches on.

Ell: On an even simpler point, with that you’re assuming only outbound traffic through those devices. Are they not connected to anything within the internal network? Is there no way for an attacker to pivot between systems? I pivot over to that, I get the information, and I make an outbound connection on something that’s not configured that way.

Corey: We had—you’re allowed to talk outbound to the management subnet, which was on its own VLAN, and that could make established connections into other things, but nothing else was allowed to connect into that. There was some defense in depth and some thought put into this. I didn’t come up with most of this to be clear, it was—this was smart people sitting around. And yeah, if I sit here and think about this for a while, of course there’s going to be ways to do it. This was also back in the days of doing it in physical data centers, so you could have a pretty good idea of what was connect to the outside world just by looking at where the cables went. But there was also always the question of how does this–does this do what I think it’s doing or what have I overlooked? Security’s job is never done.

Ell: Or what was misconfigured in the last update. It’s an assumption that everything goes correctly.

Corey: Oh, there is that. I want to talk though, about the things I had to worry about back then, it seems like in many cases get kicked upstairs to the cloud providers that we’re using these days. But then we see things like Azurescape where security researchers were able to gain access to the Azure control plane where customers using Cosmos DB—Azure’s managed database service, one of them—could suddenly have their data accessed by another customer. And Azure is doing its clam up thing and not talking about this publicly other than a brief disclosure, but how is this even possible from security architecture point of view? It makes me wonder if it hadn’t been disclosed publicly by the researcher, would they have ever said something? Most assuredly not.

Ell: I’ve worked with several researchers, in Intezer and outside of Intezer, and the amount of frustration that I see within reasonable disclosure, it just blows my mind. You have somebody threatening to sue the researcher if they bring it out. You have a company going, “Okay, well, we’ve only had six weeks. Give us three more weeks.” And next thing we know, it’s six months.

There is just this pushback about what we can actually bring out to the public on why they’re vulnerable in organizations. So, we’re put in this catch-22 as researchers. At what point is my responsibility to the public, and at what point is my responsibility to protect myself, to keep myself from getting sued personally, to keep my company from going down? How can we win when we have small research groups and these
massive cloud providers?

Corey: This episode is sponsored in part by something new. Cloud Academy is a training platform built on two primary goals. Having the highest quality content in tech and cloud skills, and building a good community the is rich and full of IT and engineering professionals. You wouldn’t think those things go together, but sometimes they do. Its both useful for individuals and large enterprises, but here's what makes it new. I don’t use that term lightly. Cloud Academy invites you to showcase just how good your AWS skills are. For the next four weeks you’ll have a chance to prove yourself. Compete in four unique lab challenges, where they’ll be awarding more than $2000 in cash and prizes. I’m not kidding, first place is a thousand bucks. Pre-register for the first challenge now, one that I picked out myself on Amazon SNS image resizing, by visiting cloudacademy.com/corey. C-O-R-E-Y. That’s cloudacademy.com/corey. We’re gonna have some fun with this one!

Corey: For a while, I was relatively confident that we had things like Google’s Project Zero, but then they started softening their disclosure timelines and the rest, and it was, we had the full disclosure security distribution list that has been shuttered to my understanding. Increasingly, it’s become risky to—yourself—to wind up publishing something that has not been patched and blessed by the providers and the rest. For better or worse, I don’t have those problems, just because I’m posting about funny implications of the bill. Yeah, worst case, AWS is temporarily embarrassed, and they can wind up giving credits to people who were affected and be mad at me for a while, but there’s no lasting harm in the way that there is with well, people were just able to look at your data for six months, and that’s our bad oops-a-doozy. Especially given the assertions that all of these providers have made to governments, to banks, to tax authorities, to all kinds of environments where security really, really matters.

Ell: The last statistic that I heard, and it was earlier this year, that it takes over 200 days for compromise even to be detected. How long is it going to take for them to backtrack, figure out how it got in, have they already patched those systems and that vulnerability is gone, but they managed to establish persistence somehow, the layers that go into actually doing your digital forensics only delay the amount of time that any of that is going to come out where that they have some information to present to you. We keep going, “Oh, we found this vulnerability. We’re working on patches. We have it fixed.” But does every single vendor already have it pitched? Do they know how it actually interacted within one customer’s environment that allowed that breach to happen? It’s just ridiculous to think that’s actually occurring, and every company is now protected because that patch came out.

Corey: As I take a look at how companies respond to these things, you’re right, the number one concern most of them have is image control, if I’m being honest with you. It’s the reputational management of we are still good at security, even though we’ve had a lapse here. Like, every breach notification starts out with, “Your security is important to us.” Well, clearly not that important because look at the email you had to send. And it’s almost taken on aspects of a comedy piece where it [grips 00:23:10] with corporate insincerity. On some level, when you tell a company that they have a massive security vulnerability, their first questions are not about the data privacy; it’s about how do we spend this to make ourselves come out of this with the least damage possible. And I understand it, but it’s still crappy.

Ell: Us tech folk talk to each other. When we have security and developers speaking to each other, we’re a lot more honest than when we’re talking to the public, right? We don’t try to hold that PR umbrella over ourselves. I was recently on a panel speaking with developers, head SRE folk—what was there? I think there was a CISO on there—and one of the developers just honestly came out and said, “At the end, my job is to say, ‘How much is that breach going to cost, versus how much money will the company lose if I don’t make that deployment?’” The first thing that you notice there is that whole how much money you’ll lose? The second part is why is the developer the one looking at the breach?

Corey: Yeah. The work flows downward. One of the most depressing aspects to me of the CISO role is that it seems like the job is to delegate everything, sign binding contracts in your name, and eventually get fired when there’s a breach and your replacement comes in to sign different papers. All the work gets delegated, none of the responsibility does, ideally—unless you’re SolarWinds and try and blame it on an intern; I mean, I wish I had an ablative intern or two around here to wind up a casting blame they don’t deserve on them. But that’s a separate argument—there is no responsibility-taking as I look at this. And that’s really a depressing commentary on the state of the world.

Ell: You say there’s no responsibility taken, but there is a lot of blame assigned. I love the concept of post-mortems to why that breach happened, but the only people in the room are the security team because they had that much control over anything. Companies as a whole need a scapegoat, and more and more, security teams are being blamed for every single compromised as more and more responsibility, more and more privileges, and visibility into what’s going on is being taken away from them. Those two just don’t balance. And I think it’s causing a lot of just complacency and almost giving up from our security teams.

Corey: To be clear, when we talk about blameless post-mortems for things like this, I agree with it wholeheartedly within the walls of a company. However, externally as someone whose data has been taken in some of these breaches, oh, I absolutely blame the company. As I should, especially when it’s something like well, we have inadvertently leaked your browsing history. Why were you collecting that in the first place? Is sort of the next logical question.

I don’t believe that my ISP needs that to serve me better. But now you have Verizon sending out emails recently—as of this recording—saying that unless anyone opts out, all the lines in our cell account are going to wind up being data mined effectively, so they can better target advertisements and understand us better. It’s no, I absolutely do not want you to be doing that on my phone. Are you out of your mind? There are a few things in this world that we consider more private than our browsing histories. We ask the internet things we wouldn’t ask our doctors in many cases, and that is no small thing as far as the level of trust that we place in our ISPs that they are now apparently playing fast and loose with.

Ell: I’m going to take this step back because you do a lot of work with cloud providers. Do you think that we actually know what information is being collected about our companies and what we have configured internally and externally by the cloud provider?

Corey: That’s a good question. I’ve seen this before, where people will give me the PDF exploded view of last month’s AWS bill, and they’ll laugh because what information can I possibly get out of that. It just shows spend on services. But I could do that to start sketching out a pretty good idea of what their architecture looks like from that alone. There’s an awful lot of value in the metadata.

Now, I want to be clear, I do not believe on any provider—except possibly Azure because who knows at this point—that if you encrypt the data, using their encryption facilities—with AWS, I know it’s KMS, for example—I do not believe that they can arbitrarily decrypt it and then scan for whatever it is they’re looking for. I do not believe that they are doing that because as soon as something like that comes out, it puts the lie to a whole bunch of different audit attestations that they’ve made and brings the entire empire crumbling down. I don’t think they’re going to get any useful data from that. However, if I’m trying to build something like Amazon Prime Video, and I can just look at the bill from the Netflix account. Well, that tells me an awful lot about things that they might be doing internally; it’s highly suggestive. Could that be used to give them an unfair advantage? Absolutely.

I had a tweet a while back that I don’t believe that Google’s Gmail division is scanning inboxes for things that look like AWS invoices to target their sales teams, but I sure would feel better if they would assure me that was the case. No one was able to ever assure me of that. It’s I don’t mean to be sitting here slinging mud, but at the same time, it’s given that when you don’t explicitly say you’re not doing something as a company, there’s a great chance you might be doing it, that’s the sort of stuff that worries me, it’s a bunch of unfair dirty trick style stuff.

Ell: Maybe I’m just cynical, or maybe I just focus on these topics too much, but after giving a presentation on cloud security, I had two groups, both, you know, from three letter government agencies, come up to me and say, “How do I have these conversations with the cloud provider?” In the conversation, they say, “We’ve contacted them several times; we want to look at this data; we want to see what they’ve collected, and we get ghosted, or we end up talking to attorneys. And despite over a year of communication, we’ve yet to be able to sit down with them.”

Corey: Now, that’s an interesting story. I would love to have someone come to me with that problem. I don’t know how I would solve that yet. But I have a couple ideas.

Ell: Hey, maybe they’re listening, and they’ll reach out to you. But—

Corey: You know, if you’re having that problem of trying to understand what your cloud provider is doing, please talk to me. I would love to go a little more in depth on that conversation, under an NDA or six.

Ell: I was at a loss because the presentation that I was giving was literally about the compromise of managed service providers, whether that be an outsourced security group, whether that be your cloud provider, we’re seeing attack groups going after these tar—think about how juicy they are. Why do I need to compromise your account or your company if I can compromise that managed service provider and have access to 15 companies?

Corey: Oh, yeah. It’s why would someone spend time trying to break into my NetApp when they could break into S3 and get access to everyone’s data, theoretically? It’s a centralization of security model risk.

Ell: Yeah, it seems to so many people as just this crazy idea. It’s so far out there. We don’t need to worry about it. I mean, we’ve talked about how Azure Functions has been compromised. We talked about all of these cloud services that people are specifically going after and being able to make traction in these attacks.

It’s not just this crazy idea. It’s something that’s happening now, and with the progress that attackers are making, criminal groups are making, this is going to happen pretty soon.

Corey: Sometimes when I’m out for a meal with someone who works with AWS in the security org, there’ll be an appetizer where, “Oh, there’s two of you. I’m going to bring three of them,” because I guess waitstaff love to watch people fight like that. And whenever I want the third one, all I have to do is say, “Can you imagine a day in which, just imagine hypothetically, IAM failed open and allowed every request to go through regardless of everything else?” Suddenly, they look sick, lose their appetite, and I get the third one. But it’s at least reassuring to know that even the idea of that is that disgusting to them, and it’s not the, “Oh, that happened three weeks ago, but don’t tell anyone.” Like, there’s none of that going on.

I do believe that the people working on these systems at the cloud providers are doing amazingly good work. I believe they are doing far better than I would be able to do in trying to manage all those things myself, by a landslide. But nothing is ever perfect. And it makes me wonder that if and when there are vulnerabilities, as we’ve already seen—clearly—with Azure, how forthcoming and transparent would they really be? And that’s the thing that keeps me up at night.

Ell: I keep going back during this talk, but just the interaction with the people there and the crowd was just so eye-opening. And I don’t want to be that person, but I keep getting to these moments of, “I told you so.” And I’m not going to go into SolarWinds. Lord, that has been covered, but shortly after that, we saw the same group going through and trying to—I’m not sure if they successfully did it, but they were targeting networks for cloud computing providers. How many companies focused outside of that compromise at that moment to see what it was going to build out to?

Corey: That’s the terrifying thing is if you can compromise a cloud service provider at this point, it’s well, you could sell that exploit on the dark web to someone. Yeah, that is a—if you can get a remote code execution be able to look into any random Cloud account, there’s almost no amount of money that is enough for something like that. You could think of the insider trading potential of just compromising Slack. A single company, but everyone talks about everything there, and Slack retains data in perpetuity. Think at the sheer M&A discussions you could come up with? Think of what you could figure out with a sort of a God’s eye view of something like that, and then realize that they run on AWS, as do an awful lot of other companies. The damage would be incalculable.

Ell: I am not an attacker, nor do I play one on TV, but let’s just, kind of, build this out. If I was to compromise a cloud provider, the first thing I would do is lay low. I don’t want them to know that I’m there. The next thing I would do is start getting into company environments and scanning them. That way I can see where the vulnerabilities are, I can compromise them that way, and not give out the fact that I came in through that cloud provider. Look, I’m just me sitting here. I’m not a nation state. I’m not somebody who is paid to do this from nine to five, I can only imagine what they would come up with.

Corey: It really feels like this is no longer a concern just for those folks who manage have gotten on the bad side of some country’s secret service. It seems like APTs, Advanced Persistent Threats, are now theoretically something almost anyone has to worry about.

Ell: Let me just set the record straight right now on what I think we need to move away from: The whole APTs are nation states. Not anymore. And APT is anyone who has advanced tactics, anyone who’s going to be persistent—because you know what, it’s not that they’re targeting you, it’s that they know that they eventually can get in. And of course, they’re a threat to you. When I was researching my work into Advanced Persistent Threats, we had a group named TNT that said, “Okay, you know what? We’re done.”

So, I contacted them and I said, “Here’s what I’m presenting on you. Would you mind reviewing it and tell me if I’m right?” They came back and said, “You know what? We’re not in APT because we target open Docker API ports. That’s how easy it is.” So, these big attack groups are not even having to rely on advanced methods anymore. The line onto what that is just completely blurring.

Corey: That’s the scariest part to me is we take a look at this across the board. And the things I have to worry about are no longer things that are solely within my arena of control. They used to be, back when it was in my data center, but now increasingly, I have to extend trust to a whole bunch of different places. Because we’re not building anything ourselves. We have all kinds of third-party dependencies, and we have to trust that they’re doing the right things as they go, too, and making sure that they’re bound so that the monitoring agent that I’m using can’t compromise my entire environment. It’s really a good time to be professionally paranoid.

Ell: And who is actually responsible for all this? Did you know that 70% of the vulnerabilities on our systems right now are on the application level? Yet security teams have to protect it? That doesn’t make sense to me at all. And yet, developers can pull in any third-party repository that they need in order to make that application work because hey, we’re on a deadline. That function needs to come out.

Corey: Ell, I want to thank you for taking the time to speak with me. If people want to learn more about how you see the world and what kind of security research you’re advocating for, where can they find you?

Ell: I live on Twitter to the point where I’m almost embarrassed to say, but you can find me at @Ell_o_Punk.

Corey: Excellent. And we will wind up putting a link to that in the [show notes 00:35:37], as we always do. Thanks so much again for your time. I appreciate it.

Ell: Always. I’d be happy to come again. [laugh].

Corey: Ell Marquez, security research advocate at Intezer. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an angry comment that ends in a link that begs me to click it that somehow it looks simultaneously suspicious and frightening.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Serena

Serena is a Network Engineer who specializes in Data Center Compute and Virtualization. She has degrees in Computer Information Systems with a concentration on networking and information security and is currently pursuing a master’s in Data Center Systems Engineering. She is most known for her content on TikTok and Twitter as Shenetworks. Serena’s content focuses on networking and security for beginners which has included popular videos on bug bounties, switch spoofing, VLAN hoping, and passing the Security+ certification in 24 hours.

Links:

  • TikTok: https://www.tiktok.com/@shenetworks
  • Twitter: https://twitter.com/notshenetworks?ref_src=twsrc%5Egoogle%7Ctwcamp%5Eserp%7Ctwgr%5Eauthor

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: It seems like there is a new security breach every day. Are you confident that an old SSH key, or a shared admin account, isn’t going to come back and bite you? If not, check out Teleport. Teleport is the easiest, most secure way to access all of your infrastructure. The open source Teleport Access Plane consolidates everything you need for secure access to your Linux and Windows servers—and I assure you there is no third option there. Kubernetes clusters, databases, and internal applications like AWS Management Console, Yankins, GitLab, Grafana, Jupyter Notebooks, and more. Teleport’s unique approach is not only more secure, it also improves developer productivity. To learn more visit: goteleport.com. And not, that is not me telling you to go away, it is: goteleport.com.

Corey: This episode is sponsored in part by our friends at Redis, the company behind the incredibly popular open source database that is not the bind DNS server. If you’re tired of managing open source Redis on your own, or you’re using one of the vanilla cloud caching services, these folks have you covered with the go to manage Redis service for global caching and primary database capabilities; Redis Enterprise. To learn more and deploy not only a cache but a single operational data platform for one Redis experience, visit redis.com/hero. Thats r-e-d-i-s.com/hero. And my thanks to my friends at Redis for sponsoring my ridiculous non-sense.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Once upon a time, I was a grumpy Unix systems administrator—because it’s not like there’s a second kind of Unix systems administrator—then I decided it was time to get better at the networking piece, so I got a CCNA one year. Did this make me a competent network engineer? Absolutely not. But it made me a slightly better systems person.

My guest today is coming from the other side of the world, specifically someone who is, in fact, good at the networking things. Serena—or @SheNetworks as you might know her from TikTok or @notshenetworks from the Twitters—thank you for joining me, I appreciate your time.

Serena: Yeah, thanks for inviting me on.

Corey: So, at a very high level, you are a network engineer, and you specialize in data center compute and virtualization, which is fun because I remember doing a lot of that once upon a time before I went basically all in on Cloud consulting, and then sort of forgot that data centers existed. That’s still a thing that’s still going well, and there are computers out there that don’t belong to what are the three biggest tech companies in the world?

Serena: Yeah. Shockingly, there’s still a ton of data centers out there, still a lot of private hosting, and a lot of the environments that we see are mixed environment; they will have some cloud, some on-prem. But yes, data centers are still relevant. [laugh].

Corey: On some level, it feels like once you get into the world of cloud, you don’t have to really think about networking anymore. You know, until there’s a big outage, and suddenly everyone had think about the networks. But it also feels like it is abstractions piled upon abstractions in the cloud infrastructure space. How much of what happens in data centers these days maps to what happens in these hyperscaler provider environments?

Serena: That’s a good question. I think—so I have two CCNAs; I’m very familiar with networking, I’m very familiar with virtualization, and I went and got my AWS certification because as we’re talking about a lot of cloud things happening now, it’s big, it’s good to know about it. And underlying infrastructure under the cloud is all the data centers that I work with, all the networking things that I work with. So, it maps very well to me. I thought I had, like, a really easy time studying for my AWS certification because a lot of the concepts just had, like, a different fancy name for AWS versus just what you know, as, like, NAT, or, you know, DNS, different things like that.

Corey: Of course, NAT used to be a thing that was—everyone would yell at you, “It’s not security,” even though there are—I would argue there are security elements tied into it. But honestly, that feels like one of the best ways to pick fights with people who are way better at this than I am. Nowadays, of course, I just view NAT through a lens of, “Yeah, I totally want to pay an extra four-and-a-half cents per gigabyte passing through a managed NAT gateway,” which remains, of course, my nemesis. The intersection of security, networking, and billing leads to basically just being very angry all the time.

Serena: Yeah. You come into the field, like, so ready to go, and then sometimes you do get beat down. But it’s worth it, I think. I really like what I do.

Corey: And what you do is something of an anomaly because most people who focus on this world of data center networking and the security aspects thereof, and the virtualization stuff, are all—how do I put it politely?—old, grumpy and unpleasant. I mean, I guess I’m not going to put it politely because I’m just going to be honest with it. Because I’m one of those people, let’s be clear here. Instead, you are creating a whole bunch of content on Twitter and on TikTok, where I’ve got to say that the union set in the Venn diagram between TikTok and deep-dive networking and cybersecurity is basically you. How did you get there?

Serena: That’s a really good question. To your first point, the, you know, old grumpy, kind of, stereotype, those are honestly some of my favorite people, truly, because I don’t know what it is, but I just vibe with them in a work environment so well. And it’s funny, you know, when I got my first job out of college, I was definitely the youngest person on my team by far. And we would all go out to lunch, I would mess with all of them, we’d all play pranks on each other. Just integrating into the teams was always super easy for me, which I’m really lucky that—not everybody has that experience, especially in their first job; things are a little rough.

But it’s always great. Like, I love the diversity in tech. And to your second point, how did I end up here, right, with this kind of intersection from this networking world to TikTok? People are always confused. Like, how did that happen? How are you finding followers on TikTok that are interested in networking?

And I’m just as shocked honestly. [laugh]. I started making this content this time last year, and… you know, at first I was like, nobody wants to learn about DNS on TikTok. This is where people dance and play pranks and all this stuff.

Corey: And if there’s dancing when it comes to DNS, at some point, something has gone other hilarious or terrifyingly. That again, I use it as a database, so who am I to talk?

Serena: [laugh]. Yeah, but it’s been fun. I am shocked. But there’s such a wide variety of people now using TikTok and it’s growing so quickly. Early on in my TikTok career, I had messages and emails from people who are vice presidents at major Fortune 100 companies asking me, you know, if I’d be interested in working there or, you know, something like that, and I was just—I was so shocked because there was a company that was a Fortune 100, and one of their VPs joined one of my Lives, and was asking me questions, just about, like, my background career, and then they sent me a follow up email [laugh] to be like, “Hey.”

So, I was like, “Did I just get interviewed on my Live on TikTok?” And that they always, like, cracked me up. And at that point, I knew I was like, okay, this is something different; like, this is interesting. Because, you know, at the end of the day, you see the views and the numbers and the followers, but you don’t have, really, faces to put to them or names, and you don’t really know where a lot of these people are from, so you don’t know who’s seeing it. And a lot of times, I think I made the assumption that they are younger kids. Which is true, but there are also a lot of very seasoned professionals that have been in this field for a very long time that also follow me, and comment on my videos, and add great input and things like that.

Corey: There’s a giant misunderstanding, I think across the industry, that the executives at the big serious companies, you know, the ones whose mottos may as well be, “That’s not funny,” have no personality themselves as people and that they live their entire lives in this corporate bubble where they talk to their kids primarily via I don’t know, Microsoft Teams, or WebEx, or something else equally sad. And in practice, that just doesn’t work that way. They’re human beings, too. And granted, you have to present in certain ways in certain rooms, but the idea that, oh, you’re only going to reach developers with attitude problems by having a personality of being on modern platforms. I mean, it’s an easy mistake to make.

I know this because I spent years making it myself with the nonsense that I do until suddenly people are reaching out and it’s, “Huh. You sure did use a lot of high-level strategic terms for a developer.” And you start digging into it, and it’s like, “Oh, you’re your chief operating officer to giant company. I bet your code is terrible.” Is it? It’s like, “Yeah. Turns out, maybe I’m not looking at that through the right lens.” Meeting people where they are with engaging content is important, and I think that a lot of folks completely miss that bus.

Serena: Yeah, I agree. And this is a small field, right, so it gets kind of nerve wracking sometimes because sometimes you say things and it’s so easy to be like, this is how I joke with my friends. But I’m still somewhat in a professional capacity because of me associating with my career, right? And then when my videos reach a million, half-a-million views, when we think about how many people are actually in this field that would be interested in viewing that content, you realize, oh, wow. Like, this is a huge mixed bag of people, which does include very high level executives, all the way to people that are in high school that are just interested in learning more. So, it’s definitely been interesting to figure that out along the way. [laugh]. But yeah, they will have regular personalities. They all like TikTok too. If they don’t, they’re lying. [laugh].

Corey: I used to be very down on the whole TikTok thing, but I started experimenting with it. And yeah, it turns out I have a face for radio and, you know, the social graces for Twitter. So, it’s not really my cup of tea, but I enjoy watching it. I found that I’m not really a video person, but something about the TikTok format means I’m just going to start scrolling. And oh, dear, it’s been six hours and my phone battery died. Thank God, or I’d still be there. There’s something very captivating about it and I really like the format.

The problem I always had with looking at a lot of the deeply technical content out there is so many companies are out there producing this and selling this. And that’s fine. Like, money is not the end all, be all [of this 00:09:40]. I’m about to spend weeks of my life on something, the fact that it cost me 30 or 50 bucks or whatnot is really not economic thing I should be concerning myself with. But it all feels like it’s classroom stuff. It’s if you give people an option, are you going to go to a college lecture or are you going to go to a comedy show? Does the idea of, I want to be entertained. If you can teach me something while entertaining me, that feels like the winning combination, and you’ve absolutely nailed that.

Serena: I think a lot of these companies that are producing content, hold themselves back a lot. And that is why they’re not successful, right? Because there’s so many stipulations, and there’s teams of people, and boardrooms of approvals, and all these things, and me, all I’m doing—I record all my TikToks on my iPhone, and I just use in-app editing. I spend a lot of time kind of researching, right, maybe I will experiment with different formats, but the best format that’s worked for me is just being authentic, kind of, not having that corporate vibe, right? And also not really expecting anything in return.

So, a lot of times, corporations are putting out content because they obviously want to drive traffic to their websites, and different things like that, but the companies that do the best are the ones that are just putting out content for free, and really not necessarily expecting anything in return. And they also give themselves so much more leeway into the type of content that they create because they’re not thinking about the numbers at the end of it, right? You just got to put stuff out there and people will see it. For me, I just put stuff out there, I don’t need to wait for someone to approve my TikTok for me to push it out and have this content there. So, that is a big difference.

And I’ve learned that through working with sponsors where they’ll send you a giant list of talking points they want you to say and I’m like, “You guys know this is a 60-second video, right?” It needs to be really small. You need to, like, really learn how to get the really important stuff out there because the rest of the smaller stuff doesn’t matter as much. Like, sell them on one big thing, and that really makes a difference.

Corey: Oh, very much so. I see that sometimes with this show where people will reach out and ask about sponsoring, and they’ll want to have a URL that I read into the microphone, and it’s with UTM tracking parameters and the rest. And it’s, like, “I appreciate where you’re coming from and your intention here, however, that is not generally how this format works, so let’s talk about this and the outcome.” And again, it’s a brave new world out there. Yeah, if you’re used to buying display ads in various places, that is exactly what you do.

For some reason, there’s this corporate mentality toward we’re going to spend $25 million on a billboard saturation campaign, and not really give any thought about what we’re actually going to say now that we have all of that visual real estate to get people’s attention with. It’s, there’s not enough focus on the message itself, and I think that is a giant lost opportunity. Enterprise marketing doesn’t have to be boring, it can be a lot of fun.

Serena: I agree. And I think podcasting was the last, probably, big area that people budgeted for marketing, right? So, you have your traditional TV commercials and there was YouTube, and—you know, TV commercials, billboards, newspapers, then there’s YouTube, and then podcasts, I would say, probably came a little bit later, as far as these companies look at for marketing potential. And now TikTok is so new and a lot of these marketing companies have no idea how to be successful on it because it’s just so different. It’s Gen Z, the humor is different.

It’s kind of like [laugh] the wild west on social media where things are just, like, crazy, and you have to fight the algorithm because on TikTok it’s, if you don’t like it, you just scroll within three seconds. The attention span is so short. So, you really have to capture people’s attention within those first three seconds. Versus a podcast, you have the whole, let’s say, first 20 minutes to get people, kind of, interested before you can be like, oh, hey, and here’s my sponsor. So, it’s very different versus TikTok, they’ll just, like, oh, scroll. So, [laugh] you have to get creative and think differently.

Corey: Many moons ago, when I was getting my CCNA, I worked at a company where we wound up getting a core switches for the data center, which was at the time, something like 65 grand. Great. And then we rented—because we had configured it in our office—and then a couple of us had to rent a commercial van, which I think ran something like $30,000 itself to transport this thing 20 miles to the data center, and I’m sitting there going, like, “Wow, the switch is worth way more than the van that’s sitting within. Also were really shitty movers and that doesn’t seem like the best idea for anything.” But I just think they remember that, and it left an impression on me.

What I like about cloud with what I do is I can take a credit card and then spend less than $10 on AWS—or theoretically, Azure, or Google Cloud or, you know, $2 million on IBM because oops-a-doozy, but fine—and I wind up coming out the other side of that with having done some interesting disaster stuff. You are teaching people about how this stuff works, but in a data center world, it seems to me that the startup costs of, “Oh, I’m going to buy this random router or switch to wind up doing some demonstration stuff for,” it feels like the startup costs of getting hands on that equipment would be out of reach for an awful lot of people. Am I just completely out of touch with how that world works?

Serena: No, you’re right, you’re one hundred percent, right. It is difficult. So, in college, my undergraduate degree is computer information systems, and they had a Cisco Networking Academy. And so we had old switches, old layer 3 switches, and then we had some routers, and
this is all stuff that was EOL, donated equipment, right? And this is going to—

Corey: It breaks down you’re bidding against very faraway places with no budget on eBay for replacements. Oh, yes.

Serena: Yeah, exactly. And it was a lot of IOS stuff, right? And so when I was in college, I had no idea that NX-OS existed, which is the data center Nexus version operating system for their switches and things. And so when I got to my first job and saw NX-OS, I was like, “Oh, crap, [laugh] like, what is this?” Right?

Because I honestly didn’t even know. I graduated and did not know that existed. And I didn’t know a lot of the stuff that I was working on at my first shop existed. And I really had to rely on, kind of, the fundamentals. And they are transferable, right? That’s why it’s good to kind of get into—like, I know what these routing protocols are. I know, layer 2, I know this cabling, so let me just learn these command differences and things like that.

And once you get into a production environment in general, out of a lab, it hits the fan. Like, everything you feel like you’ve learned is gone almost because there’s so many layers and now all of a sudden, you have these firewalls, when before you were just trying to get, like, your routing neighborships to establish [laugh] and you weren’t worried about rules on a firewall somewhere. And [crosstalk 00:16:39]—

Corey: “Oh, and by the way, in this environment, that link that you’re working on goes down, every minute it’s down, here is the number of commas in the amount of money that we’re losing, and yes, that’s a plural.” It’s, “Okay, so I guess I’m going to double-check everything I run first.” Yeah, it’s that caution that gives people a bit of credence there. [unintelligible 00:16:58] do these things in a, more or less, cowboy style in these environments, at least not for very long. Because you can break individual servers; that’s fine, but if you break the network suddenly, you may as well not have the computers.

Serena: Yeah. It can be paralyzing, truly. It can be very overwhelming your first networking job. Especially for me, I was just dealing with outages constantly because I worked for a vendor, and I was [laugh] like, I was just scared, you know? Because I would get these cases and it would be a hospital outage.

And I’m like, “I just graduated college. Like, what do you want from me?” You know, and back to your original point, it is difficult in a data center space because the equipment’s so expensive. So, a lot of people ask, “Do you have a home lab?” And one—there’s a couple of reasons I don’t really have a significant home lab. One, I move so much.

Corey: Oh, and in the spare room basically is always 90 degrees and sounds like a jet engine taking off.

Serena: Yeah.

Corey: Yeah, it’s one of those, I should probably find a different place where I don’t live, to have that equipment. Yeah.

Serena: Yeah. And I have access, like, remotely to all the lab equipment that I really need. So, I don’t personally have one, but a lot of things that I do work with are so expensive, that I’m like, I can’t afford to put this data center equipment in my house. That doesn’t make any sense.

And there is luckily now a lot of virtual labs that you can do. There’s some sandboxes by Cisco and other vendors, where you can kind of get a little bit of hands-on experience. A lot of it relates to their certifications. You can rent racks, but that gets pretty pricey, too. So, it is difficult, and sometimes that’s why a lot of these jobs, I think I have a lot of people who are looking for entry-level work, and it’s hard to get into a specifically a data center space.

And aside from racking, stacking, working in a data center—maybe a NOC—if you want to get into the actual,s I’m configuring Nexus switches, I’m configuring, you know, Palo Alto firewalls, it can be difficult because it’s hard to get to that point, there’s not a clear path.

Corey: What is the entry path these days? I entered tech by working on a help desk, and those aren’t really the jobs that they once were, in a lot of different ways. So, I’ve stopped talking to entry-level folks with the position of, “Oh, yeah, this is what you should do because that’s what I did.” It turns into, like, “Okay, Boomer. Great job. Tell me a little bit more, though, about what the Great War was like, first.” No, we aren’t going to go down that path. It’s just I don’t know what the entry-level point is for someone who’s legitimately interested in these things these days.

Serena: Nobody does. It’s crazy. And you’re right at the, “Okay, Boomer,” thing. See, networking was one of those… things that just got pushed onto people in, just, a general IT department, right? So, that’s when everything was like, “Okay, we need to get on the internet, so, you know, hey, you handle some of the computer stuff. It’s your job now. Good luck. Figure it out.”

And so, people started doing that and they kind of just got pushed into it, and then as the internet grew, as our capabilities grew, then the job became, like, a little bit more specialized. And now we have, you know, dedicated network engineers, we have people running data centers. But that’s not necessarily a viable path now for people just because there’s so much to it now. There’s cloud, there’s security risks, there’s data center, wireless, pho—I mean, you can be an engineer just for phones, right? So, it’s a little bit difficult for, especially, the younger people coming in, and the people that I talk to, and figuring out, well, how do I get to what you’re doing?

And the way that I did is I went and got a four-year degree and then joined a new college graduate program at a Fortune 100 company. Which is a great path, I highly recommend it to anybody that can do it, but it’s also not available for everybody, right, because not everybody has the means to get a four-year education, nor do you necessarily need one to do what I do. So, everybody’s kind of has this different path, and it’s very confusing for people who are aspiring network engineers, or aspiring cloud engineers, even.

Corey: This episode is sponsored by our friends at Oracle HeatWave is a new high-performance accelerator for the Oracle MySQL Database Service. Although I insist on calling it “my squirrel.” While MySQL has long been the worlds most popular open source database, shifting from transacting to analytics required way too much overhead and, ya know, work. With HeatWave you can run your OLTP and OLAP, don’t ask me to ever say those acronyms again, workloads directly from your MySQL database and eliminate the time consuming data movement and integration work, while also performing 1100X faster than Amazon Aurora, and 2.5X faster than Amazon Redshift, at a third of the cost. My thanks again to Oracle Cloud for sponsoring this ridiculous nonsense.

Corey: The narrative the cloud companies have been pushing for a while—like, and I’m in that space deeply enough that I haven’t really thought to go super deep into questioning this—is that well, the future is all cloud, the data center is basically this legacy thing that the tide is slowly eroding, in the fullness of time, because everything will one day be cloud. Do you think that’s accurate?

Serena: I don’t. I really don’t think that’s accurate. Don’t get me wrong, I think that the cloud is here to stay, and a lot of people are going to be using it. And it’s going to be—and it currently is a huge part of our lives. Like, as we’ve seen recently with a few of the AWS outages, when it goes down and goes down hard because everything’s so centralized.

And people like to think, like, oh, you know, we have all this redundancy, yadda, yadda. That has not protected us so far, [laugh] like, from these major outages, right? And a lot of places that I see—especially when you’re looking at public sector—is a hybrid, where you do have data center on-prem and you have cloud. And I think that, personally, is the best way to go. Unless, you know, maybe you’re a fast growing startup and AWS or Azure makes a lot of sense to you.

And it does. There’s great use cases for that, right? But they’re—not only aside from the whole cloud shift, there’s another shift of, you know, making our data centers eco-friendly, too, and workload optimization. So, maybe the price point that you’re looking for, what’s going to save your business the most money, is doing that hybrid. So, I’m going to store a lot of my private documents on site, I’m going to have this as a backup disaster recovery, but we’re also going to operate in the cloud. I don’t think that the data centers as we know them are going to go extinct. [laugh]. I think they will be around.

Corey: Well, AWS finally made their Outpost—the smaller ones; read as servers that run AWS services on in your facility—available a year after announcing them. And I looked at it like, oh, wow, these things are 600 bucks a month. Which is not nothing, but certainly something I could afford to wind up exploring and doing some content. But okay, first, it’s a three-year commitment. So, that’s 20 grand or so. Okay, not ideal, but fine.

That would effectively almost double my AWS bill, but that’s not the hardest part because, oh, and to get one of these, you have to have enterprise support. And when I pointed this out to some Amazonian friends, their response was, “Well, what’s the problem on this?” Yeah, enterprise support starts at $15,000 a month minimum, and that means that people aren’t going to pick these up to do proof of concept work. They’re going to do it when they already have a significant infrastructure out there, and I think that’s leaving an awful lot of money on the table by making people jump through sales hoops, and getting proof of concept credits, and doing all the other stuff for this. It’s just ship me a box for a few weeks and let me kick the tires on in my environment and see if it works or doesn’t work.

Worst case, I’ll ship it back to you. Worst, worst case, I lose the thing, and then you charge me whatever it costs to replace this. But it still feels like they are really doing the whole, “Oh, it’s only big legacy companies that have on-premises stuff.” I don’t like that narrative.

Serena: I don’t either. And I honestly think it’s a bad idea, right, because if you do put all of your eggs in the AWS basket and they have all the power, that’s not going to give us a lot of bargaining, right? That’s not going to give people a lot of—because they’ll know. They know how hard it is to get off of AWS at that point: They know it’s costly, it takes manpower, it takes knowledge, right? And I think that it is in people’s best interest to kind of have that mixed environment. Just for long-term, I’m just very wary of centralizing everything in one area. I think it’s a bad idea. [laugh]. I think that we need to be prepared for ourselves, and that means also relying a little bit on ourselves. We can’t just, in my opinion, put everything in the AWS basket. [laugh].

Corey: Not very long anyway. It just doesn’t seem to work.

Serena: Right. And it’s a great product.

Corey: Oh, it absolutely is, but—

Serena: There’s so many positive things about using cloud. Because I’m not the type of person that likes to, kind of, talk crap about any vendor. I think everybody has their pros, cons, flaws, whatever. It’s really about what works best for your environment, and that’s part of being a network engineer or an architect is evaluating your environment and figuring out what is going to be the best for you, right? There’s no one size fits all, unfortunately.

Corey: Yeah. And AWS is uniformly excellent, let’s be very clear. Okay, not—maybe not uniformly. Some services are significantly better than others, but I have an opinion piece in the information—paywalled, unfortunately, but I’m working on i—the general thesis that AWS has gotten too big to fail, in that when it’s not—like, first, they are going to have better uptime than you or I will running our own data centers, across the board.

They are very good at keeping things up, but when they do go down, it’s not just your company or my company anymore having an outage, it is a significant portion of, you know, the global economy, and that is an awful lot of systemic concentrated risk. I’m not suggesting they did anything wrong, as far as how they sold these things—though, some people will want to argue with that—but it’s the, “What does this mean?” Are we ready to reckon with that as a society that whenever us-east-1 has a bad day, so does the stock market? Is that something we’re really prepared to accept or wrangle with? Or worse than that, there are life-critical services now. Does that mean that we’re going to accept there is some number of people who will die when there’s an outage of a data center? And that’s new territory for me. I have not worked in environments where it was life or death consequential. At least not directly.

Serena: Yeah, I have. So, I have definitely worked in those environments, right, and it’s very scary, and especially when it’s outside of your control. So, if you are relying, or just waiting on AWS to get back up, you don’t have the control to get in there and start fixing things yourself, which is my instinct, right? Like, I immediately want to get hands-on. I put my troubleshooting hat on, like, let’s figure this out, let me look through logs, let me do this.

And you don’t have that option with AWS when it’s a significant outage that’s impacting multiple people, it’s not some configuration internally to you, right?And that’s scary. It’s a scary place to be. And I think that we need to really consider the cascading effects that will happen, which a lot of these outages that are kind of starting to show us, right? And luckily, there hasn’t been anything major catastrophic, but we do need to really consider life when we’re talking about, you know, hospitals, 911 systems, all of these critical infrastructures that are going to be cloud managed, and out of our control, and centralized.

So, you know, you lose one 911 system, okay, well, you can do a backup, right? You may be able to route all your calls to the city over because their 911 systems are up and running. Well, what if there’s are out now, too, because you’re both hosted on AWS?

Corey: Or you’re, “Ah, we’re going to diversify and we’re going to have this other one on a different cloud provider.” That’s great, but there’s a critical third-party dependency that’s right back to the thing you’re trying to avoid. And there you go again.

Serena: Yep. And that’s dependency hell, right? [laugh].

Corey: Oh, yeah. And I don’t know how we get away from that.

Serena: Yeah.

Corey: Like, we don’t want everyone writing all their own stuff from scratch, like starting with assembly, move up the stack. But here we are.

Serena: Right. And it’s funny because these AWS outages specifically effects—or cloud outages, right? I feel like I’m picking on them. I’m not
trying to—sorry, AWS, but [laugh] don’t come for me.

But you know, explaining to my mom, why her Ring doorbell is not working and her Roomba stopped working when that outage happened, right, she’s like, “Why is this not—it won’t connect.” Like, “I don’t understand.” She’s like, “What’s AWS?” And then to tell my mom that the company that she buys her socks from, like, that she goes online and, like, buys on Amazon is the company that also is hosting her Roomba, you know, services, her Ring services, it’s so interesting to have those conversations. And a lot of people who aren’t in our field don’t understand that. They don’t understand cloud, they don’t understand on-prem versus, you know, hosted by a third-party. So, it’s interesting to watch that kind of unfold now because it’s very new. It’s very new territory.

Corey: And one last question before we wind up calling it an episode. It is remarkably clear in talking to you that you are in no way, shape, or form, junior. You are not a beginner. You know exactly how this stuff works in significant depth. Your content that you put out is aimed at beginners. I do something very similar. So, to be very clear, this is not a criticism in the slightest, but I am curious as to why that’s the direction you went in.

Serena: I think there’s a few reasons. Well, I might have this knowledge, right? I still consider myself very junior in my career, very early in my career. There’s so many things that I don’t know and I recognize that. When you’re first starting out, you might have this kind of inflated sense of knowledge where you’re like—like, me, I was like, “Oh, yeah. I know all about OSPF and running on IOS and the command line,” until I figured out there was an NX-OS and I’m like, “Oh crap, what else do I not know about?” Right? [laugh].

Corey: Oh, by the way, that never goes away. I feel exactly the same way 20 years into my career, now. I still have absolutely no idea what I’m doing. So smile, nod, and get used to it is the only insight I’ve got there. But please, go on.

Serena: And even on Twitter sometimes, I’m reading people’s stuff, and I’m like, “How did you get into these obscure protocols and all these things?” And, you know, I just kind of dive deeper into there. But I think the big reason that I create a lot of my content for beginners is because I remember so well how it was at the beginning, learning about subnetting, and that IOS—[laugh]—[unintelligible 00:30:52] learning about subnetting, and all of the different models that we have, right? And I was overwhelmed, and I was stressed out, and it just seems so… just, like, a giant mountain to climb. It seems so daunting in the beginning, for me it did because there’s so much, right?

And it felt like everybody was so far ahead of me. And I don’t want other people to really feel like that. Like, I don’t want people to be turned off from networking because they feel like the bar is too high, that we’re not letting enough new people enter because we’re discouraging them from the beginning by saying, “Oh, well, you’re going to have to know all this. And let me throw this certification book at you.” And they’re big. Like, my certification books—and these are massive. And this is for one half of the CCNA.

Corey: For those who aren’t, like, on the video call—it’s not being recorded video-wise—she’s holding a book that you could use to kill a mid-sized dog by accident if it falls off a table. It looks like a phonebook with a hardcover on it.

Serena: Yeah. [laugh]. It’s huge, right? And there are thousands of pages, and we just give this to somebody and say, like, “Here you go. Make sure you remember all this.” And this is all new information.

Corey: And does it still cover things like EIGRP? Like Cisco's proprietary routing protocols that I’ve never once seen in the wild?

Serena: Yeah. So, sometimes you will have to learn that, and they’ve changed it recently, too. They update their certification exam. So, you will learn about some legacy protocols because sometimes you do run into them.

Corey: Oh, yes. That’s when I have the good sense to pay professionals who know what they’re doing.

Serena: [laugh]. Yeah. Exactly. So yeah, you do run into those sometimes. But it feels so daunting for new people, and I totally recognize that. And by nature of TikTok I, especially when I first start making content, I assume that most of the people on there are going to be people who are younger, who are interested in this career.

And as you know, in tech in general, especially networking, security, cloud, there’s a massive shortage of people, and how are we solving that, right? And my contribution to helping solve that is by getting people interested. And now I have people that DM me and say, “I passed my [Network+ 00:33:01],” or, “I just took the CCNA,” or, “This has been helping me with my class so much.” And that is like, okay, this is great.

Like, that’s exactly what I want. I want to help the pipeline, I want to get more people interested and help a diverse group of people get interested in tech and say, “Hey, like, this is, you know, where I came from. And I did it; you can do it; let’s do it together,” type situation.

Corey: I really want to thank you for being so generous with your time. If people want to learn more, as they absolutely should, where can they find you?

Serena: I am on TikTok as @SheNetworks. I am on Twitter as @notshenetworks because somebody else—

Corey: That is very confusing.

Serena: [laugh]. I know. Well, my initial thing was like, I didn’t really use Twitter that much, and I would just like—I kind of used it as, like, a backchannel to my TikTok, right, where I would just, like, “Hey, I’m going to go live,” or do this. And then my Twitter, kind of, got a little out of control [laugh] and out of my hands. And so—

Corey: It does that sometimes.

Serena: Yeah. I had no idea there would be so much interest. And it surprises me every day. So, it’s exciting though. I really love all the people that I’ve met, and I feel like I fit in, and I’ve met so many good friends that it’s been great. But yeah, so @notshenetworks on Twitter because somebody had shenetworks and it was a joke. And [laugh] so if you want to find me there, you could also find me there.

Corey: And we will, of course, put links to that in the [show notes 00:34:20]. Thank you so much for taking the time to speak with me today. I really do appreciate it.

Serena: Thank you for having me. This has been great. [laugh].

Corey: Serena, also known as @SheNetworks, networking content creator to the stars. I’m cloud economist, Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice and then a long, angry, rambling comment about how the network isn’t that important that you’re then not going to be able to submit because the network isn’t working.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Anil

Anil Dash is the CEO of Glitch, the friendly developer community where coders collaborate to create and share millions of web apps. He is a recognized advocate for more ethical tech through his work as an entrepreneur and writer. He serves as a board member for organizations like the Electronic Frontier Foundation, the leading nonprofit defending digital privacy and expression, Data & Society Research Institute, which researches the cutting edge of tech's impact on society, and The Markup, the nonprofit investigative newsroom that pushes for tech accountability. Dash was an advisor to the Obama White House’s Office of Digital Strategy, served for a decade on the board of Stack Overflow, the world’s largest community for coders, and today advises key startups and non-profits including the Lower East Side Girls Club, Medium, The Human Utility, DonorsChoose and Project Include.

As a writer and artist, Dash has been a contributing editor and monthly columnist for Wired, written for publications like The Atlantic and Businessweek, co-created one of the first implementations of the blockchain technology now known as NFTs, had his works exhibited in the New Museum of Contemporary Art, and collaborated with Hamilton creator Lin-Manuel Miranda on one of the most popular Spotify playlists of 2018. Dash has also been a keynote speaker and guest in a broad range of media ranging from the Obama Foundation Summit to SXSW to Desus and Mero's late-night show.

Links:

  • Glitch: https://glitch.com
  • Web.dev: https://web.dev
  • Glitch Twitter: https://twitter.com/glitch
  • Anil Dash Twitter: https://twitter.com/anildash

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: It seems like there is a new security breach every day. Are you confident that an old SSH key, or a shared admin account, isn’t going to come back and bite you? If not, check out Teleport. Teleport is the easiest, most secure way to access all of your infrastructure. The open source Teleport Access Plane consolidates everything you need for secure access to your Linux and Windows servers—and I assure you there is no third option there. Kubernetes clusters, databases, and internal applications like AWS Management Console, Yankins, GitLab, Grafana, Jupyter Notebooks, and more. Teleport’s unique approach is not only more secure, it also improves developer productivity. To learn more visit: goteleport.com. And not, that is not me telling you to go away, it is: goteleport.com.Corey: It seems like there is a new security breach every day. Are you confident that an old SSH key, or a shared admin account, isn’t going to come back and bite you? If not, check out Teleport. Teleport is the easiest, most secure way to access all of your infrastructure. The open source Teleport Access Plane consolidates everything you need for secure access to your Linux and Windows servers—and I assure you there is no third option there. Kubernetes clusters, databases, and internal applications like AWS Management Console, Yankins, GitLab, Grafana, Jupyter Notebooks, and more. Teleport’s unique approach is not only more secure, it also improves developer productivity. To learn more visit: goteleport.com. And not, that is not me telling you to go away, it is: goteleport.com.

Corey: This episode is sponsored in part by our friends at Redis, the company behind the incredibly popular open source database that is not the bind DNS server. If you’re tired of managing open source Redis on your own, or you’re using one of the vanilla cloud caching services, these folks have you covered with the go to manage Redis service for global caching and primary database capabilities; Redis Enterprise. To learn more and deploy not only a cache but a single operational data platform for one Redis experience, visit redis.com/hero. Thats r-e-d-i-s.com/hero. And my thanks to my friends at Redis for sponsoring my ridiculous non-sense.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Today’s guest is a little bit off the beaten path from the cloud infrastructure types I generally drag, kicking and screaming, onto the show. If we take a look at the ecosystem and where it’s going, it’s clear that in the future, not everyone who wants to build a business, or a tool, or even an application is going to necessarily spring fully-formed into the world from the forehead of some God, knowing how to code. And oh, “I’m going to go to a boot camp for four months to learn how to do it first,” is increasingly untenable. I don’t know if you would call it low-code or not. But that’s how it feels. My guest today is Anil Dash, CEO of Glitch. Anil, thank you for joining me.

Anil: Thanks so much for having me.

Corey: So, let’s get the important stuff out of the way first, since I have a long-standing history of mispronouncing the company Twitch as ‘Twetch,’ I should probably do the same thing here. So, what is Gletch? And what does it do?

Anil: Glitch is, at its simplest, a tool that lets you build a full-stack app in your web browser in about 30 seconds. And, you know, for your community, your audience, it’s also this ability to create and deploy code instantly on a full-stack server with no concern for deploy, or DevOps, or provisioning a container, or any of those sort of concerns. And what it is for the users is, honestly, a community. They’re like, “I looked at this app that was on Glitch; I thought it was cool; I could do what we call [remixing 00:02:03].” Which is to kind of fork that app, a running app, make a couple edits, and all of a sudden live at a real URL on the web, my app is running with exactly what I built. And that’s something that has been—I think, just captured a lot of people’s imagination to now where they’ve built over 12 or 15 million apps on the platform.

Corey: You describe it somewhat differently than I would, and given that I tend to assume that people who create and run successful businesses don’t generally tend to do it without thought, I’m not quite, I guess, insufferable enough to figure out, “Oh, well, I thought about this for ten seconds, therefore I’ve solved a business problem that you have been needling at for years.” But when I look at Glitch, I would describe it as something different than the way that you describe it. I would call it a web-based IDE for low-code applications and whatnot, and you never talk about it that way. Everything I can see there describes it talks about friendly creators, and community tied to it. Why is that?

Anil: You’re not wrong from the conventional technologist’s point of view. I—sufficient vintage; I was coding in Visual Basic back in the ’90s and if you squint, you can see that influence on Glitch today. And so I don’t reject that description, but part of it is about the audience we’re speaking to, which is sort of a next generation of creators. And I think importantly, that’s not just age, right, but that could be demographic, that can be just sort of culturally, wherever you’re at. And what we look at is who’s making the most interesting stuff on the internet and in the industry, and they tend to be grounded in broader culture, whether they’re on, you know, Instagram, or TikTok, or, you know, whatever kind of influencer, you want to point at—YouTube.

And those folks, they think of themselves as creators first and they think of themselves as participating in the community first and then the tool sort of follow. And I think one of the things that’s really striking is, if you look at—we’ll take YouTube as an example because everyone’s pretty familiar with it—they have a YouTube Creator Studio. And it is a very rich and deep tool. It does more than, you know, you would have had iMovie, or Final Cut Pro doing, you know, 10 or 15 years ago, incredibly advanced stuff. And those [unintelligible 00:04:07] use it every day, but nobody goes to YouTube and says, “This is a cloud-based nonlinear editor for video production, and we target cinematographers.” And if they did, they would actually narrow their audience and they would limit what their impact is on the world.

And so similarly, I think we look at that for Glitch where the social object, the central thing that people organize around a Glitch is an app, not code. And that’s this really kind of deep and profound idea, which is that everybody can understand an app. Everybody has an idea for an app. You know, even the person who’s, “Ah, I’m not technical,” or, “I’m not really into technology,” they’re like, “But you know what? If I could make an app, I would make this.”

And so we think a lot about that creative impulse. And the funny thing is, that is a common thread between somebody that literally just got on the internet for the first time and somebody who has been doing cloud deploys for as long as there’s been a cloud to deploy to, or somebody has been coding for decades. No matter who you are, you have that place that is starting from what’s the experience I want to build, the app I want to build? And so I think that’s where there’s that framing. But it’s also been really useful, in that if you’re trying to make a better IDE in the cloud and a better text editor, and there are multiple trillion-dollar companies that [laugh] are creating products in that category, I don’t think you’re going to win. On the other hand, if you say, “This is more fun, and cooler, and has a better design, and feels better,” I think we could absolutely win in a walk away compared to trillion-dollar companies trying to be cool.

Corey: I think that this is an area that has a few players in it could definitely stand to benefit by having more there. My big fear is not that AWS is going to launch stuff in your space and drive you out of business; I think that is a somewhat naive approach. I’m more concerned that they’re going to try to launch something in your space, give it a dumb name, fail that market and appropriately, not understand who it’s for and set the entire idea back five years. That is, in some cases, it seems like their modus operandi for an awful lot of new markets.

Anil: Yeah, I mean, that’s not an uncommon problem in any category that’s sort of community driven. So, you know, back in the day, I worked on building blogging tools at the beginning of this, sort of, social media era, and we worried about that a lot. We had built some of the first early tools, Movable Type, and TypePad, and these were what were used to launch, like, Gawker and Huffington Post and all the, sort of, big early sites. And we had been doing it a couple years—and then at that time, major player—AOL came in, and they launched their own AOL blog service, and we were, you know, quaking in our boots. I remember just being kind of like, pit in your stomach, “Oh, my gosh. This is going to devastate the category.”

And as it turns out, people were smart, and they have taste, and they can tell. And the domain that we’re in is not one that is about raw computing power or raw resources that you can bring to bear so much as it is about can you get people to connect together, collaborate together, and feel like they’re in a place where they want to make something and they want to share it with other people? And I mean, we’ve never done a single bit of advertising for Glitch. There’s never been any paid acquisition. There’s never done any of those things. And we go up against, broadly in the space, people that have billboards and they buy out all the ads of the airport and, you know, all the other kind of things we see—

Corey: And they do the typical enterprise thing where they spend untold millions in acquiring the real estate to advertise on, and then about 50 cents on the message, from the looks of it. It’s, wow, you go to all this trouble and expense to get something in front of me, and after all of that to get my attention, you don’t have anything interesting to say?

Anil: Right.

Corey: [crosstalk 00:07:40] inverse of that.

Anil: [crosstalk 00:07:41] it doesn’t work.

Corey: Yeah. Oh, yeah. It’s brand awareness. I love that game. Ugh.

Anil: I was a CIO, and not once in my life did I ever make a purchasing decision based on who was sponsoring a golf tournament. It never happened, right? Like, I never made a call on a database platform because of a poster that was up at, you know, San Jose Airport. And so I think that’s this thing that developers in particular, have really good BS filters, and you can sort of see through.

Corey: What I have heard about the airport advertising space—and I but a humble cloud economist; I don’t know if this is necessarily accurate or not—but if you have a company like Accenture, for example, that advertises on airport billboards, they don’t even bother to list their website. If you go to their website, it turns out that there’s no shopping cart function. I cannot add ‘one consulting’ to my cart and make a purchase.

Anil: “Ten pounds of consult, please.”

Corey: Right? I feel like the primary purpose there might very well be that when someone presents to your board and says, “All right, we’ve had this conversation with Accenture.” The response is not, “Who?” It’s a brand awareness play, on some level. That said, you say you don’t do a bunch traditional advertising, but honestly, I feel like you advertise—more successfully—than I do at The Duckbill Group, just by virtue of having a personality running the company, in your case.

Now, your platform is for the moment, slightly larger than mine, but that’s okay,k I have ambition and a tenuous grasp of reality and I’m absolutely going to get there one of these days. But there is something to be said for someone who has a track record of doing interesting things and saying interesting things, pulling a, “This is what I do and this is how I do it.” It almost becomes a personality-led marketing effort to some degree, doesn’t it?

Anil: I’m a little mindful of that, right, where I think—so a little bit of context and history: Glitch as a company is actually 20 years old. The product is only a few years old, but we were formerly called Fog Creek Software, co-founded by Joel Spolsky who a lot of folks will know from back in the day as Joel on Software blog, was extremely influential. And that company, under leadership of Joel and his co-founder Michael Pryor spun out Stack Overflow, they spun out Trello. He had created, you know, countless products over the years so, like, their technical and business acumen is off the charts.

And you know, I was on the board of Stack Overflow from, really, those first days and until just recently when they sold, and you know, you get this insight into not just how do you build a developer community that is incredibly valuable, but also has a place in the ecosystem that is unique and persists over time. And I think that’s something that was very, very instructive. And so when it came in to lead Glitch I, we had already been a company with a, sort of, visible founder. Joel was as well known as a programmer as it got in the world?

Corey: Oh, yes.

Anil: And my public visibility is different, right? I, you know, I was a working coder for many years, but I don’t think that’s what people see me on social media has. And so I think, I’ve been very mindful where, like, I’m thrilled to use the platform I have to amplify what was created on a Glitch. But what I note is it’s always, “This person made this thing. This person made this app and it had this impact, and it got these results, or made this difference for them.”

And that’s such a different thing than—I don’t ever talk about, “We added syntax highlighting in the IDE and the editor in the browser.” It’s just never it right. And I think there are people that—I love that work. I mean, I love having that conversation with our team, but I think that’s sort of the difference is my enthusiasm is, like, people are making stuff and it’s cool. And that sort of is my lens on the whole world.

You know, somebody makes whatever a great song, a great film, like, these are all things that are exciting. And the Glitch community’s creations sort of feel that way. And also, we have other visible people on the team. I think of our sort of Head of Community, Jenn Schiffer, who’s a very well known developer and her right. And you know, tons of people have read her writing and seen her talks over the years.

And she and I talk about this stuff; I think she sort of feels the same way, which is, she’s like, “If I were, you know, being hired by some cloud platform to show the latest primitives that they’ve deployed behind an API,” she’s like, “I’d be miserable. Like, I don’t want to do that in the world.” And I sort of feel the same way. But if you say, “This person who never imagined they would make an app that would have this kind of impact.” And they’re going to, I think of just, like, the last couple of weeks, some of the apps we’ve seen where people are—it could be [unintelligible 00:11:53]. It could be like, “We made a Slack bot that finally gets this reporting into the right channel [laugh] inside our company, but it was easy enough that I could do it myself without asking somebody to create it even though I’m not technically an engineer.” Like, that’s
incredible.

The other extreme, we have people that are PhDs working on machine learning that are like, “At the end of the day, I don’t want to be responsible for managing and deploying. [laugh]. I go home, and so the fact that I can do this in create is really great.” I think that energy, I mean, I feel the same way. I still build stuff all the time, and I think that’s something where, like, you can’t fake that and also, it’s bigger than any one person or one public persona or social media profile, or whatever. I think there’s this bigger idea. And I mean, to that point, there are millions of developers on Glitch and they’ve created well over ten million apps. I am not a humble person, but very clearly, that’s not me, you know? [laugh].

Corey: I have the same challenge to it’s, effectively, I have now a 12 employee company and about that again contractors for various specialized functions, and the common perception, I think, is that mostly I do all the stuff that we talk about in public, and the other 11 folks sort of sit around and clap as I do it. Yeah, that is only four of those people’s jobs as it turns out. There are more people doing work here. It’s challenging, on some level, to get away from the myth of the founder who is the person who has the grand vision and does all the work and sees all these things.

Anil: This industry loves the myth of the great man, or the solo legend, or the person in their bedroom is a genius, the lone genius, and it’s a lie. It’s a lie every time. And I think one of the things that we can do, especially in the work at Glitch, but I think just in my work overall with my whole career is to dismantle that myth. I think that would be incredibly valuable. It just would do a service for everybody.

But I mean, that’s why Glitch is the way it is. It’s a collaboration platform. Our reference points are, you know, we look at Visual Studio and what have you, but we also look at Google Docs. Why is it that people love to just send a link to somebody and say, “Let’s edit this thing together and knock out a, you know, a memo together or whatever.” I think that idea we’re going to collaborate together, you know, we saw that—like, I think of Figma, which is a tool that I love. You know, I knew Dylan when he was a teenager and watching him build that company has been so inspiring, not least because design was always supposed to be collaborative.

And then you think about we’re all collaborating together in design every day. We’re all collaborating together and writing in Google Docs—or whatever we use—every day. And then coding is still this kind of single-player game. Maybe at best, you throw something over the wall with a pull request, but for the most part, it doesn’t feel like you’re in there with somebody. Certainly doesn’t feel like you’re creating together in the same way that when you’re jamming on these other creative tools does. And so I think that’s what’s been liberating for a lot of people is to feel like it’s nice to have company when you’re making something.

Corey: Periodically, I’ll talk to people in the AWS ecosystem who for some reason appear to believe that Jeff Barr builds a lot of these services himself then writes blog posts about them. And it’s, Amazon does not break out how many of its 1.2 million or so employees work at AWS, but I’m guessing it’s more than five people. So yeah, Jeff probably only wrote a dozen of those services himself; the rest are—

Anil: That’s right. Yeah.

Corey: —done by service teams and the rest. It’s easy to condense this stuff and I’m as guilty of it as anyone. To my mind, a big company is one that has 200 people in it. That is not apparently something the world agrees with.

Anil: Yeah, it’s impossible to fathom an organization of hundreds of thousands or a million-plus people, right? Like, our brains just aren’t wired to do it. And I think so we reduce things to any given Jeff, whether that’s Barr or Bezos, whoever you want to point to.

Corey: At one point, I think they had something like more men named Jeff on their board than they did women, which—

Anil: Yeah. Mm-hm.

Corey: —all right, cool. They’ve fixed that and now they have a Dave problem.

Anil: Yeah [unintelligible 00:15:37] say that my entire career has been trying to weave out of that dynamic, whether it was a Dave, a Mike, or a Jeff. But I think that broader sort of challenge is this—that is related to the idea of there being this lone genius. And I think if we can sort of say, well, creation always happens in community. It always happens influenced by other things. It is always—I mean, this is why we talk about it in Glitch.

When you make an app, you don’t start from a blank slate, you start from a working app that’s already on the platform and you’re remix it. And there was a little bit of a ego resistance by some devs years ago when they first encountered that because [unintelligible 00:16:14] like, “No, no, no, I need a blank page, you know, because I have this brilliant idea that nobody’s ever thought of before.” And I’m like, “You know, the odds are you'll probably start from something pretty close to something that’s built before.” And that enabler of, “There’s nothing new under the sun, and you’re probably remixing somebody else’s thoughts,” I think that sort of changed the tenor of the community. And I think that’s something where like, I just see that across the industry.

When people are open, collaborative, like even today, a great example is web browsers. The folks making web browsers at Google, Apple, Mozilla are pretty collaborative. They actually do share ideas together. I mean, I get a window into that because they actually all use Glitch to do test cases on different bugs and stuff for them, but you see, one Glitch project will add in folks from Mozilla and folks from Apple and folks from the Chrome team and Google, and they’re like working together and you’re, like—you kind of let down the pretense of there being this secret genius that’s only in this one organization, this one group of people, and you’re able to make something great, and the web is greater than all of them. And the proof, you know, for us is that Glitch is not a new idea. Heroku wanted to do what we’re doing, you know, a dozen years ago.

Corey: Yeah, everyone wants to build Heroku except the company that acquired Heroku, and here we are. And now it’s—I was waiting for the next step and it just seemed like it never happened.

Anil: But you know when I talked to those folks, they were like, “Well, we didn’t have Docker, and we didn’t have containerization, and on the client side, we didn’t have modern browsers that could do this kind of editing experience, all this kind of thing.” So, they let their editor go by the wayside and became mostly deploy platform. And—but people forget, for the first year or two Heroku had an in-browser editor, and an IDE and, you know, was constrained by the tech at the time. And I think that’s something where I’m like, we look at that history, we look at, also, like I said, these browser manufacturers working together were able to get us to a point where we can make something better.

Corey: This episode is sponsored by our friends at Oracle HeatWave is a new high-performance accelerator for the Oracle MySQL Database Service. Although I insist on calling it “my squirrel.” While MySQL has long been the worlds most popular open source database, shifting from transacting to analytics required way too much overhead and, ya know, work. With HeatWave you can run your OLTP and OLAP, don’t ask me to ever say those acronyms again, workloads directly from your MySQL database and eliminate the time consuming data movement and integration work, while also performing 1100X faster than Amazon Aurora, and 2.5X faster than Amazon Redshift, at a third of the cost. My thanks again to Oracle Cloud for sponsoring this ridiculous nonsense.

Corey: I do have a question for you about the nuts and bolts behind the scenes of Glitch and how it works. If I want to remix something on Glitch, I click the button, a couple seconds later it’s there and ready for me to start kicking the tires on, which tells me a few things. One, it is certainly not using CloudFormation to provision it because I didn’t have time to go and grab a quick snack and take a six hour nap. So, it apparently is running on computers somewhere. I have it on good authority that this is not just run by people who are very fast at assembling packets by hand. What does the infrastructure look like?

Anil: It’s on AWS. Our first year-plus of prototyping while we were sort of in beta and early stages of Glitch was getting that time to remix to be acceptable. We still wish it were faster; I mean, that’s always the way but, you know, when we started, it was like, yeah, you did sit there for a minute and watch your cursor spin. I mean, what’s happening behind the scenes, we’re provisioning a new container, standing up a full stack, bringing over the code from the Git repo on the previous project, like, we’re doing a lot of work, lift behind the scenes, and we went through every possible permutation of what could make that experience be good enough. So, when we start talking about prototyping, we’re at five-plus, almost six years ago when we started building the early versions of what became Glitch, and at that time, we were fairly far along in maturity with Docker, but there was not a clear answer about the use case that we’re building for.

So, we experimented with Docker Swarm. We went pretty far down that road; we spent a good bit of time there, it failed in ways that were both painful and slow to fix. So, that was great. I don’t recommend that. In fairness, we have a very unusual use case, right? So, Glitch now, if you talk about ten million containers on Glitch, no two of those apps are the same and nobody builds an orchestration infrastructure assuming that every single machine is a unique snowflake.

Corey: Yeah, massively multi-tenant is not really a thing that people know.

Anil: No. And also from a security posture Glitch—if you look at it as a security expert—it is a platform allowing anonymous users to execute
arbitrary code at scale. That’s what we do. That’s our job. And so [laugh], you know, so your threat model is very different. It’s very different.

I mean, literally, like, you can go to Glitch and build an app, running a full-stack app, without even logging in. And the reason we enable that is because we see kids in classrooms, they’re learning to code for the first time, they want to be able to remix a project and they don’t even have an email address. And so that was about enabling something different, right? And then, similarly, you know, we explored Kubernetes—because of course you do; it’s the default choice here—and some of the optimizations, again, if you go back several years ago, being able to suspend a project and then quickly sort of rehydrate it off disk into a running app was not a common use case, and so it was not optimized. And so we couldn’t offer that experience because what we do with Glitch is, if you haven’t used an app in five minutes, and you’re not a paid member, who put that app to sleep. And that’s just a reasonable—

Corey: Uh, “Put the app to sleep,” as in toddler, or, “Put the app to sleep,” as an ill puppy.

Anil: [laugh]. Hopefully, the former, but when we were at our worst and scaling the ladder. But that is that thing; it’s like we had that moment that everybody does, which is that, “Oh, no. This worked.” That was a really scary moment where we started seeing app creation ramping up, and number of edits that people were making in those apps, you know, ramping up, which meant deploys for us ramping up because we automatically deploy as you edit on Glitch. And so, you know, we had that moment where just—well, as a startup, you always hope things go up into the right, and then they do and then you’re not sleeping for a long time. And we’ve been able to get it back under control.

Corey: Like, “Oh, no, I’m not succeeding.” Followed immediately by, “Oh, no, I’m succeeding.” And it’s a good problem to have.

Anil: Exactly. Right, right, right. The only thing worse than failing is succeeding sometimes, in terms of stress levels. And organizationally, you go through so much; technically, you go through so much. You know, we were very fortunate to have such thoughtful technical staff to navigate these things.

But it was not obvious, and it was not a sort of this is what you do off the shelf. And our architecture was very different because people had looked at—like, I look at one of our inspirations was CodePen, which is a great platform and the community love them. And their front end developers are, you know, always showing off, “Here’s this cool CSS thing I figured out, and it’s there.” But for the most part, they’re publishing static content, so architecturally, they look almost more like a content management system than an app-running platform. And so we couldn’t learn anything from them about our scaling our architecture.

We could learn from them on community, and they’ve been an inspiration there, but I think that’s been very, very different. And then, conversely, if we looked at the Herokus of the world, or all those sort of easy deploy, I think Amazon has half a dozen different, like, “This will be easier,” kind of deploy tools. And we looked at those, and they were code-centric not app-centric. And that led to fundamentally different assumptions in user experience and optimization.

And so, you know, we had to chart our own path and I think it was really only the last year or so that we were able to sort of turn the corner and have high degree of confidence about, we know what people build on Glitch and we know how to support and scale it. And that unlocked this, sort of, wave of creativity where there are things that people want to create on the internet but it had become too hard to do so. And the canonical example I think I was—those of us are old enough to remember FTPing up a website—

Corey: Oh, yes.

Anil: —right—to Geocities, or whatever your shared web host was, we remember how easy that was and how much creativity was enabled by
that.

Corey: Yes, “How easy it was,” quote-unquote, for those of us who spent years trying to figure out passive versus active versus ‘what is going on?’ As far as FTP transfers. And it turns out that we found ways to solve for that, mostly, but it became something a bit different and a bit weird. But here we are.

Anil: Yeah, there was definitely an adjustment period, but at some point, if you’d made an HTML page in notepad on your computer, and you could, you know, hurl it at a server somewhere, it would kind of run. And when you realize, you look at the coding boot camps, or even just to, like, teach kids to code efforts, and they’re like, “Day three. Now, you’ve gotten VS Code and GitHub configured. We can start to make something.” And you’re like, “The whole magic of this thing getting it to light up. You put it in your web browser, you’re like, ‘That’s me. I made this.’” you know, north star for us was almost, like, you go from zero to hello world in a minute. That’s huge.

Corey: I started participating one of those boot camps a while back to help. Like, the first thing I changed about the curriculum was, “Yeah, we’re not spending time teaching people how to use VI in, at that point, the 2010s.” It was, that was a fun bit of hazing for those of us who were becoming Unix admins and knew that wherever we’d go, we’d find VI on a server, but here in the real world, there are better options for that.

Anil: This is rank cruelty.

Corey: Yeah, I mean, I still use it because 20 years of muscle memory doesn’t go away overnight, but I don’t inflict that on others.

Anil: Yeah. Well, we saw the contrast. Like, we worked with, there’s a group called Mouse here in New York City that creates the computer science curriculum for the public schools in the City of New York. And there’s a million kids in public school in New York City, right, and they all go through at least some of this CS education. [unintelligible 00:24:49] saw a lot of work, a lot of folks in the tech community here did. It was fantastic.

And yet they were still doing this sort of very conceptual, theoretical. Here’s how a professional developer would set up their environment. Quote-unquote, “Professional.” And I’m like, you know what really sparks kids’ interests? If you tell them, “You can make a page and it’ll be live and you can send it to your friend. And you can do it right now.”

And once you’ve sparked that creative impulse, you can’t stop them from doing the rest. And I think what was wild was kids followed down that path. Some of the more advanced kids got to high school and realized they want to experiment with, like, AI and ML, right? And they started playing with TensorFlow. And, you know, there’s collaboration features in Glitch where you can do real-time editing and a code with this. And they went in the forum and they were asking questions, that kind of stuff. And the people answering their questions were the TensorFlow team at Google. [laugh]. Right?

Corey: I remember those days back when everything seemed smaller and more compact, [unintelligible 00:25:42] but almost felt like a balkanization of community—

Anil: Yeah.

Corey: —where now it’s oh, have you joined that Slack team, and I’m looking at this and my machine is screaming for more RAM. It’s, like, well, it has 128 gigs in it. Shouldn’t that be enough? Not for Slack.

Anil: Not for chat. No, no, no. Chat is demanding.

Corey: Oh, yeah, that and Chrome are basically trying to out-ram each other. But if you remember the days of volunteering as network staff on Freenode when you could basically gather everyone for a given project in the entire stack on the same IRC network. And that doesn’t happen anymore.

Anil: And there’s something magic about that, right? It’s like now the conversations are closed off in a Slack or Discord or what have you, but to have a sort of open forum where people can talk about this stuff, what’s wild about that is, for a beginner, a teenage creator who’s learning this stuff, the idea that the people who made the AI, I can talk to, they’re alive still, you know what I mean? Like, yeah, they’re not even that old. But [laugh]. They think of this is something that’s been carved in stone for 100 years.

And so it’s so inspiring to them. And then conversely, talking to the TensorFlow team, they made these JavaScript examples, like, tensorflow.js was so accessible, you know? And they’re like, “This is the most heartwarming thing. Like, we think about all these enterprise use cases or whatever. But like, kids wanting to make stuff, like recognize their friends’ photo, and all the vision stuff they’re doing around [unintelligible 00:26:54] out there,” like, “We didn’t know this is why we do it until we saw this is why we do it.”

And that part about connecting the creative impulse from both, like, the most experienced, advanced coders at the most august tech companies that exist, as well as the most rank beginners in public schools, who might not even have a computer at home, saying that’s there—if you put those two things together, and both of those are saying, “I’m a coder; I’m able to create; I can make something on the internet, and I can share it with somebody and be inspired by it,” like, that is… that’s as good as it gets.

Corey: There’s something magic in being able to reach out to people who built this stuff. And honestly—you shouldn’t feel this way, but you do—when I was talking to the folks who wrote the things I was working on, it really inspires you to ask better questions. Like when I’m talking to Dr. Venema, the author of Postfix and I’m trying to figure out how this thing works, well, I know for a fact that I will not be smarter than he is at
basically anything in that entire universe, and maybe most beyond that, as well, however, I still want to ask a question in such a way that doesn’t make me sound like a colossal dumbass. So, it really inspires you—

Anil: It motivates you.

Corey: Oh, yeah. It inspires you to raise your question bar up a bit, of, “I am trying to do x. I expect y to happen. Instead, z is happening as opposed to what I find the documentation that”—oh, as I read the documentation, discover exactly what I messed up, and then I delete the whole email. It’s amazing how many of those things you never send because when constructing a question the right way, you can help yourself.

Anil: Rubber ducking against your heroes.

Corey: Exactly.

Anil: I mean, early in my career, I’d gone through sort of licensing mishap on a project that later became open-source, and sort of stepped it in and as you do, and unprompted, I got an advice email from Dan Bricklin, who invented the spreadsheet, he invented VisiCalc, and he had advice and he was right. And it was… it was unreal. I was like, this guy’s one of my heroes. I grew up reading about his work, and not only is he, like, a living, breathing person, he’s somebody that can have the kindness to reach out and say, “Yeah, you know, have you tried this? This might work.”

And it’s, this isn’t, like, a guy who made an app. This is the guy who made the app for which the phrase killer app was invented, right? And, you know, we’ve since become friends and I think a lot of his inspiration and his work. And I think it’s one of the things it’s like, again, if you tell somebody starting out, the people who invented the fundamental tools of the digital era, are still active, still building stuff, still have advice to share, and you can connect with them, it feels like a cheat code. It feels like a superpower, right? It feels like this impossible thing.

And I think about like, even for me, the early days of the web, view source, which is still buried in our browser somewhere. And you can see the code that makes the page, it felt like getting away with something. “You mean, I can just look under the hood and see how they made this page and then I can do it too?” I think we forget how radical that is—[unintelligible 00:29:48] radical open-source in general is—and you see it when, like, you talk to young creators. I think—you know, I mean, Glitch obviously is used every day by, like, people at Microsoft and Google and the New York Timesor whatever, like, you know, the most down-the-road, enterprise developers, but I think a lot about the new creators and the people who are learning, and what they tell me a lot is the, like, “Oh, so I made this app, but what do I have to do to put it on the internet?”

I’m like, “It already is.” Like, as soon as you create it, that URL was live, it all works. And their, like, “But isn’t there, like, an app store I have to ask? Isn’t there somebody I have to get permission to publish this from? Doesn’t somebody have to approve it?”

And you realize they’ve grown up with whether it was the app stores on their phones, or the cartridges in their Nintendo or, you know, whatever it was, they had always had this constraint on technology. It wasn’t something you make; it’s something that is given to you, you know, handed down from on high. And I think that’s the part that animates me and the whole team, the community, is this idea of, like, I geek out about our infrastructure. I love that we’re doing deploys constantly, so fast, all the time, and I love that we’ve taken the complexity away, but the end of the day, the reason why we do it, is you can have somebody just sort of saying, I didn’t realize there was a place I could just make something put it in front of, maybe, millions of people all over the world and I don’t have to ask anybody permission and my idea can matter as much as the thing that’s made by the trillion-dollar company.

Corey: It’s really neat to see, I guess, the sense of spirit and soul that arises from a smaller, more, shall we say, soulful company. No disparagement meant toward my friends at AWS and other places. It’s just, there’s something that you lose when you get to a certain point of scale. Like, I don’t ever have to have a meeting internally and discuss things, like, “Well, does this thing that we’re toying with doing violate antitrust law?” That is never been on my roadmap of things I have to even give the slightest crap about.

Anil: Right, right? You know, “What does the investor relations person at a retirement fund think about the feature that we shipped?” Is not a question that we have to answer. There’s this joy in also having community that sort of has come along with us, right? So, we talk a lot internally about, like, how do we make sure Glitch stays weird? And, you know, the community sort of supports that.

Like, there’s no reason logically that our logo should be the emoji of two fish. But that kind of stuff of just, like, it just is. We don’t question it anymore. I think that we’re very lucky. But also that we are part of an ecosystem. I also am very grateful where, like… yeah, that folks at Google use Glitch as part of their daily work when they’re explaining a new feature in Chrome.

Like, if you go to web.dev and their dev portal teaches devs how to code, all the embedded examples go to these Glitch apps that are running, showing running code is incredible. When we see the Stripe team building examples of, like, “Do you want to use this new payment API that we made? Well, we have a Glitch for you.” And literally every day, they ship one that sort of goes and says, “Well, if you just want to use this new Stripe feature, you just remix this thing and it’s instantly running on Glitch.”

I mean, those things are incredible. So like, I’m very grateful that the biggest companies and most influential companies in the industry have embraced it. So, I don’t—yeah, I don’t disparage them at all, but I think that ability to connect to the person who’d be like, “I just want to do payments. I’ve never heard of Stripe.”

Corey: Oh yeah.

Anil: And we have this every day. They come into Glitch, and they’re just like, I just wanted to take credit cards. I didn’t know there’s a tool to do that.

Corey: “I was going to build it myself,” and everyone shrieks, “No, no. Don’t do that. My God.” Yeah. Use one of their competitors, fine,k but building it yourself is something a lunatic would do.

Anil: Exactly. Right, right. And I think we forget that there’s only so much attention people can pay, there’s only so much knowledge they have.

Corey: Everything we say is new to someone. That’s why I always go back to assuming no one’s ever heard of me, and explain the basics of what I do and how I do it, periodically. It’s, no one has done all the mandatory reading. Who knew?

Anil: And it’s such a healthy exercise to, right, because I think we always have that kind of beginner’s mindset about what Glitch is. And in fairness, I understand why. Like, there have been very experienced developers that have said, “Well, Glitch looks too colorful. It looks like a toy.” And that we made a very intentional choice at masking—like, we’re doing the work under the hood.

And you can drop down into a terminal and you can do—you can run whatever build script you want. You can do all that stuff on Glitch, but that’s not what we put up front and I think that’s this philosophy about the role of the technology versus the people in the ecosystem.

Corey: I want to thank you for taking so much time out of your day to, I guess, explain what Glitch is and how you view it. If people want to learn more about it, about your opinions, et cetera. Where can they find you?

Anil: Sure. glitch.com is easiest place, and hopefully that’s a something you can go and a minute later, you’ll have a new app that you built that you want to share. And, you know, we’re pretty active on all social media, you know, Twitter especially with Glitch: @glitch. I’m on as @anildash.

And one of the things I love is I get to talk to folks like you and learn from the community, and as often as not, that’s where most of the inspiration comes from is just sort of being out in all the various channels, talking to people. It’s wild to be 20-plus years into this and still never get tired of that.

Corey: It’s why I love this podcast. Every time I talk to someone, I learn something new. It’s hard to remain too ignorant after you have enough people who’ve shared wisdom with you as long as you can retain it.

Anil: That’s right.

Corey: Thank you so much for taking the time to speak with me.

Anil: So, glad to be here.

Corey: Anil Dash, CEO of Gletch—or Glitch as he insists on calling it. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry comment telling me how your small team at AWS is going to crush Glitch into the dirt just as soon as they find a name that’s dumb enough for the service.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Dan

Dan is CISO and VP of Cybersecurity for Shipt, a Target subsidiary. He worked previously as a Distinguished Engineer on Target’s cloud infrastructure. He served as CTO for Joe Biden’s 2020 Presidential campaign. Prior to that Dan worked with the Hillary for America tech team through the Groundwork, and contributed as a founding developer on Spinnaker while at Netflix. Dan is an O’Reilly published author and avid public speaker.

Links:

  • Shipt: https://www.shipt.com/
  • Twitter: https://twitter.com/danveloper
  • LinkedIn: https://www.linkedin.com/in/danveloper

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: It seems like there is a new security breach every day. Are you confident that an old SSH key, or a shared admin account, isn’t going to come back and bite you? If not, check out Teleport. Teleport is the easiest, most secure way to access all of your infrastructure. The open source Teleport Access Plane consolidates everything you need for secure access to your Linux and Windows servers—and I assure you there is no third option there. Kubernetes clusters, databases, and internal applications like AWS Management Console, Yankins, GitLab, Grafana, Jupyter Notebooks, and more. Teleport’s unique approach is not only more secure, it also improves developer productivity. To learn more visit: goteleport.com. And not, that is not me telling you to go away, it is: goteleport.com.

Corey: Writing ad copy to fit into a 30 second slot is hard, but if anyone can do it the folks at Quali can. Just like their Torque infrastructure automation platform can deliver complex application environments anytime, anywhere, in just seconds instead of hours, days or weeks. Visit Qtorque.io today and learn how you can spin up application environments in about the same amount of time it took you to listen to this ad.

Corey: Welcome to Screaming in the Cloud, I’m Corey Quinn. Sometimes I talk to people who are involved in working on the nonprofit slash political side of the world. Other times I talk to folks who are deep in the throes of commercial businesses, and I obviously personally spend more of my time on one of those sides of the world than I do the other. But today’s guest is a little bit different, Dan Woods is the CISO and VP of Cybersecurity at Shipt, a division of Target where he’s worked for a fair number of years, but took some time off for his side project, the side hustle as the kids call it, as the CTO for the Biden campaign. Dan, thank you for joining me.

Dan: Yeah. Thank you, Corey. Happy to be here.

Corey: So, you have an interesting track record as far as your career goes, you’ve been at Target for a long time. You were a distinguished engineer—not to be confused with ‘extinguished engineer,’ which is just someone who is finally—the fire has gone out. And from there you went from being a distinguished engineer to a VP slash CISO, which generally looks a lot less engineer-like, and a lot more, at least in my experience, of sitting in a whole lot of executive-level meetings, managing teams, et cetera. Was that, in fact, an individual contributor—or IC—move into a management track, or am I just misunderstanding this because these are commonly overloaded terms in our industry?

Dan: Yeah, yeah, no, that’s exactly right. So, IC to leadership, two distinct tracks, distinct career paths. It was something that I’ve spent a number of years thinking about and more or less working toward and making sure that it was the right path for me to go. The interesting thing about the break that I took in the middle of Target when I was CTO for the campaign is that that was a leadership role, right. I led the team. I managed the team.

I did performance reviews and all of that kind of managerial stuff, but I also sat down and did a lot of tech. So, it was kind of like a mix of being a senior executive, but also still continuing to be a distinguished engineer. So, then the natural path out of that for me was to make a decision about do I continue to be an individual contributor or do I go into a leadership track? And I felt like for a number of reasons that my interests more aligned with being on the leadership side of the world, and so that’s how I’ve ended up where I am.

Corey: And correct me if I’m wrong because generally speaking political campaigns are not usually my target customers given the fact that they’re turning the entire AWS environment off in a few months—win or lose—and yeah, that is, in fact, remains the best way to save money on your AWS bill; it’s hard for me to beat that. But at that point most of the people you’re working with are in large part volunteers I would imagine.

So, managing in a traditional sense of, “Well, we’re going to have your next quarterly review.” Well, your candidate might not be in the race then, and what we’re going to put you on a PIP, and what exactly you’re going to stop letting me volunteer here? You’re going to dock them pay—you’re not paying me for this. It becomes an interesting management challenge I would imagine just because the people you’re working with are passionate and volunteering, and a lot of traditional management and career advice doesn’t necessarily map one-to-one I would have to assume.

Dan: That is the best way that I’ve heard it described yet. I try to explain this to folks sometimes and it’s kind of difficult to get that message across that like there is sort of a base level organization that exists, right. There were full-time employees who were a part of the tech team, really great group of folks especially from very early on willing to join the campaign and be a part of what it was that we were doing.

And then there was this whole ecosystem of folks who just wanted to volunteer, folks who wanted to be a part of it but didn’t want to leave their 9:00 to 5:00 who wanted to come in. One of the most difficult things about—we rely on volunteers very heavily in the political space, and very grateful for all the folks who step up and volunteer with organizations that they feel passionate about. In fact, one of the best little tidbits of wisdom the President imparted to me at one point, we were having dinner at his house very early on in the campaign, and he said, “The greatest gift that you can give somebody is your time.” And I think that’s so incredibly true. So, the folks who volunteer, it’s really important, really grateful that they’re all there.

In particular, how it becomes difficult, is that you need somebody to manage the volunteers, right, who are there. You need somebody to come up with work and check in that work is getting done because while it’s great that folks want to volunteer five, ten hours a week, or whatever it is that they can put in, we also have very real things that need to get done, and they need to get done in a timely manner.

So, we had a lot of difficulty especially early on in the campaign utilizing the volunteers to the extent that we could because we were such a small and scrappy team and because everybody who was working on the campaign at the time had a lot of responsibilities that they needed to see through on their own. And so getting into this, it’s quite literally a full-time job having to sit down and follow up with volunteers and make sure that they have the appropriate amount of work and make sure that we’ve set up our environment appropriately so that volunteers can come and go and all of that kind of stuff, so yeah.

Corey: It’s always an interesting joy looking at the swath of architectural decisions and how they came to be. I talked on a previous episode with Jackie Singh, who was, I believe, after your tenure as CISO, she was involved on the InfoSec side of things, and she was curious as to your thought process or rationale with a lot of the initial architectural decisions that she talked about on her episode which I’m sure she didn’t intend it this way, but I am going to blatantly miscategorize as, “Justify yourself. What were you thinking?” Usually it takes years for that kind of, “I don’t understand what’s going on here so I’m playing data center archeologist or cloud spelunker.” This was a very short window. How did decisions get made architecturally as far as what you’re going to run things on? It’s been disclosed that you were on AWS, for example. Was that a hard decision?

Dan: No, not at all. Not at all. We started out the campaign—I in particular I was one of the first employees hired onto the campaign and the idea all along was that we’re not going to be clever, right? We’re basically just going to develop what needs to be developed. And the idea with that was that a lot of the code that we were going to sit down and write or a lot of the infrastructure that we were going to build was going to be glue, it not AWS Glue, right, ideally, but just glue that would bind data streams together, right?

So, data movement, vendor A produces a CSV file for you and it needs to end up in a bucket somewhere. So, somebody needs to write the code to make that happen, or you need to find a sufficient vendor who can make that happen. There’s a lot more vendors today believe it or not than there were two years ago that are doing much better in that kind of space, but two years ago we had the constraints of time and money.

Our idea was that the code that we were going to write was going to be for those purposes. What it actually turned into is that in other areas of the business—and I will call it a business because we had formalized roadmaps and different departments working on different things—but in other areas of the business where we didn’t have enough money to purchase a solution, we had the ability to go and write software.

The interesting thing about this group of technologists who came together especially early on in the campaign to build out the tech team most of them came from an enterprise software development background, right? So, we had the know-how of how to build things at scale and how to do continuous delivery and continuous deployment, and how to operate a cloud-native environment, and how to build applications for that world.

So, we ended up doing things like writing an API for managing our donor vetting pipeline, right? And that turned into a complex system of Lambda functions and continuous delivery for a variety of different services that facilitated that pipeline. We also built an architecture for our mobile app which there were plenty of companies that wanted to sell us a mobile app and we just couldn’t afford it so we ended up writing the mobile app ourselves.

So, after some point in time, what we said was we actually have a fairly robust and complex software infrastructure. We have a number of microservices that are doing various things to facilitate the operation of the business, and something that we need to do is we need to spend a little bit of time and make sure that we’re building this in a cohesive way, right? And what part of that means was that, for example, we had to take a step back and say, “Okay, we need to have a unified identity service.” We can’t have a different identity—or we can’t have every single individual service creating its own identity. We need to have—

Corey: I really wish you could pass that lesson out on some of the AWS service teams.

Dan: [laugh]. Yes, I know. I know. Yeah. So, we went through—

Corey: So, there were some questionable choices you made in there, like you started that with the beginning of, “Well, we had no time which is fine and no budget. So, we chose AWS.” It’s like, “Oh, that looks like the exact opposite direction of a great decision, given, you know, my view on it.” Stepping past that entirely, you are also dealing with challenges that I don’t think map very well to things that exist in the corporate world. For example, you said you had to build a donor vetting pipeline.

It’s in the corporate world I didn’t have it. It’s one of those, “Why in the world would I get in the way of people trying to give me money?” And the obvious answer in your case is, federal law, and it turns out that the best outcome generally does not involve serving prison time. So, you have to address these things in ways that don’t necessarily have a one-to-one analog in other spaces.

Dan: That’s true. That’s true. Yes, correct to the federal law thing. Our more pressing reason to do this kind of thing was that we made a commitment very early on in the campaign that we wouldn’t take money from executives of the gas and oil industry, for example. There were another bunch of other commitments that were made, but it was inconceivable for us to have enough people that could possibly go manually through those filings. So, for us to be able to build an automated system for doing that meant that we were literally saving thousands of human hours and still getting a beneficial result out of it.

Corey: And everything you do is subject to intense scrutiny by folks who are willing to make hay out of anything. If it had leaked at the time, I would have absolutely done some ridiculous nonsense thing about, “Ah, clearly looking at this AWS bill. Joe Biden’s supports managed NAT gateway data processing pricing.” And it’s absolutely not, but that doesn’t stop people from making hay about this because headlines are going to be headlines.

And do you have to also deal with the interesting aspect—industrial espionage is always kind of a thing, but by and large most companies don’t have to worry that effectively half of the population is diametrically opposed to the thing it is that they’re trying to do to the point where they might very well try to get insiders there to start leaking things out. Everything you do has to be built with optics in mind, working under tight constraints, and it seems like an almost insurmountable challenge except for the fact where you actually pulled it off.

Dan: Yeah. Yeah. Yeah. We kept saying that the tech was not the story, right, and we wanted to do everything within our power to keep the conversation on the candidate and not on emails or AWS bills or any of that kind of stuff. And so we were very intentional about a lot of the decisions that we ended up making with the idea that if the optics are bad, we pull away from the primary mission of what it is that we’re trying to do.

Corey: So, what was it that qualified you to be the CTO of a—at the time very fledgling and uncertain campaign, given that you were coming from a role where you were a distinguished engineer, which is not nothing, let’s be clear, but it’s an executive-level of role rather than a hands-on level of role as CTO. And then if we go back in time, you were one of the founding developers of Spinnaker over at Netflix.

And I have a lot of thoughts about Netflix technology and a lot of thoughts about Spinnaker as well, and none of those thoughts are, “This seems like a reasonable architecture I should roll out for a presidential campaign.” So, please, don’t take this as the insult that probably sounds like, but why were you the CTO that got tapped?

Dan: Great question. And I think in some ways, right place, right time. But in other ways probably needs to speak a little bit to the journey of how I’ve gotten anywhere in my career. So, going back to Netflix, yeah, so I worked in Netflix. I had the opportunity to work with a lot of incredibly bright and talented folks there. One of the people in particular who I met there and became friends with was Corey Bertram who worked on the core SRE team.

Corey left Netflix to go off and at the time he was just like, “I’m going to go do a political startup.” The interesting thing about Netflix at the time—this was 2013, so, this was just after the Obama for America ’12 campaign. And a bunch of folks from OFA world came and worked at Netflix and a variety of other organizations in the Bay Area. Corey was not one of those people but we were very well-connected with folks in that world, and Corey said he was going off to do a political startup, and so after my non-mutual departure from Netflix, I was talking to Corey and he said, “Hey, why don’t you come over and help us figure out how to do continuous delivery over on the political startup.” That political startup turned into the groundwork which turned into essentially the tech platform for the Hillary for America campaign.

So, I had the opportunity working for the groundwork to work very closely with the folks in the technology organization at HFA. And that got me more exposure to what that world is and more connections into that space. And the groundwork was run by Corey, but was the CEO or head—I don’t even know what he called himself, was Michael Slaby, who was President Obama’s CTO in 2008 and had a bigger technical role in the 2012 campaign.

And so, for his involvement in HFA ’16 meant that he was a person who was very well connected for the 2020 campaign. And when we were out at a political conference in late 2018 and he said, “Hey, I think that Vice President Biden is going to run. Do you have any interest in talking with his team?” And I said, “Yes, absolutely. Please introduce me.”

And I had a couple of conversations with Greg Schultz who was the campaign manager and we just hit it off. And it was a really great fit. Greg was an excellent leader. He was a real visionary, exactly the person that President Biden needed. And he brought me in to set up the tech operation and get everything to where we ultimately won the primary and won the election after that.

Corey: And then, as all things do, it ended and the question then becomes, “Great, what’s next?” And the answer for you was apparently, “Okay, I’m going to go back to Target-ish.” Although now you’re the CISO of a Target subsidiary, Shipt and Target’s relationship is—again, I imagine I have that correct as far as you are in fact a subsidiary of Target, so it wasn’t exactly a new company, but rather a transition into the previous organization you were in a different role.

Dan: Yeah, correct. Yeah, it’s a different department inside of Target, but my paycheck still come from Target. [laugh].

Corey: So, what was it that inspired you to go into the CISO role? Because obviously security is everyone’s job, which is what everyone says, which is why we get away with treating it like it’s nobody’s job because shared responsibilities tend to work out that way.

Dan: Yeah.

Corey: And you’ve done an awful lot of stuff that was not historically deeply security-centric although there’s always an element passing through it. Now, going into a CISO role as someone without a deep InfoSec background that I’m aware of, what drove that? How did that work?

Dan: You know, I think the most correct answer is that security has always been in my blood. I think like most people who started out—

Corey: There are medications for that now.

Dan: Yeah, [laugh] good. I might need them. [laugh]. I think like most folks who are kind of my era who started seriously getting into software development and computer system administration in the late ‘90s, early thousands, cybersecurity it wasn’t called cybersecurity at the time. It wasn’t even called InfoSec, right, it was just called, I don’t know, dabbling or something. But that was a gateway for getting into Linux system administration, network engineering, so forth and so on.

And for a short period of time I became—when I was getting my RHCE certification way back in the day, I became pretty entrenched in network security and that was a really big focus area that I spent a lot of time on and I got whatever the supplemental network security
certification from Red Hat was at the time. And then I realized pretty quickly that the world isn’t going to need box operators for very long, and this was just before the DevOps revolution had really come around and more and more things were automated.

So, we were still doing hand deployments. I was still dropping WAR files onto a file system and restarting Apache. That was our deployment process. And I saw the writing on the wall and I said, “If I don’t dedicate myself to becoming first and foremost a software engineer, then I’m not going to have a very good time in technology here.” So, I jumped out of that and I got into software development, and so that’s where my software engineering career evolved out of.

So, when I was CTO for the campaign, I like to tell people that I was a hundred percent of CTO, I was a hundred percent a CIO, and I was a hundred percent of CISO for the first 514 days of the campaign or whatever it was. So, I was 300 percent doing all of the top-level technology jobs for the campaign, but cybersecurity was without a doubt the one that we would drop everything for every single time.

And that was by necessity; we were constantly under attack on the campaign. And a lot of my headspace during that period of time was dedicated to how do we make sure that we’re doing things in the most secure way? So, when I left—when I came back into Target and I came back in as a distinguished engineer there were some areas that they were hoping that I could contribute positively and help move a couple of things along.

The idea always the whole time was going to be for me to jump into a leadership position. And I got a call one day from Rich Agostino who’s the CISO for Target and he said, “Hey, Shipt needs a cybersecurity operation built out and you’re looking for a leadership role. Would you be interested in doing this?” And believe it or not, I had missed the world of cybersecurity so much that when the opportunity came up I said, “Yes, absolutely. I’ll dive in head first.” And so that was the path for getting there.

Corey: This episode is sponsored by our friends at Oracle HeatWave is a new high-performance accelerator for the Oracle MySQL Database Service. Although I insist on calling it “my squirrel.” While MySQL has long been the worlds most popular open source database, shifting from transacting to analytics required way too much overhead and, ya know, work. With HeatWave you can run your OLTP and OLAP, don’t ask me to ever say those acronyms again, workloads directly from your MySQL database and eliminate the time consuming data movement and integration work, while also performing 1100X faster than Amazon Aurora, and 2.5X faster than Amazon Redshift, at a third of the cost. My thanks again to Oracle Cloud for sponsoring this ridiculous nonsense.

Corey: My take to cybersecurity space is, a little, I think, different than most people’s journeys through it. The reason I started a Thursday edition of the Last Week in AWS newsletter is the security happenings in the AWS ecosystem for folks who don’t have the word security in their job titles because I used to dabble in that space a fair bit. The problem I found is that is as you move up the ladder to executives that our directors, VPs, and CISOs, the language changes significantly.

And it almost becomes a dialect of corporate-speak that I find borderline impenetrable, versus the real world terminology we’re talking about when, “Okay, let’s make sure that we rotate credentials on a reasonable expected basis where it makes sense,” et cetera et cetera. It almost becomes much more of a box-checking compliance exercise slash layering on as much as you possibly can that for plausible deniability for the inevitable breach that one day hits and instead of actually driving towards better outcomes.

And I understand that’s a cynical, strange perspective, but I started talking to people about this, and I’m very far from alone in that, which is why people are subscribing to that newsletter and that’s the corner of the market I wanted to start speaking to. So, given that you’ve been an engineer practitioner trying to build things and now a security executive as well, is my assessment of the further higher up you go the entire messaging and purpose change, or is that just someone who’s been in the trenches for too long and hasn’t been on that side of the world, and I have a certain lack of perspective that would make this all very clear. Which I freely accept, if that’s the case.

Dan: No, I think that you’re right for a lot of organizations. I think that that’s a hundred percent true, and it is exactly as you described: a box-checking exercise for a lot of organizations. Something that’s important to remember about Target is—Target was the subject of a data breach in 2012, and that was before there were data breaches every single day, right.

Now, we look at a data breach and we say that’s just going to happen, right, that’s the cost of doing business. But back in 2012 it was really a very big story and it was a very big deal, and there was quite a bit of activity in the Target technology world after that breach. So, it reshaped the culture quite literally, new executives were brought in, but there’s this whole world of folks inside of Target who have never forgotten that, right, and work day-in and day-out to make sure that we don’t have another breach.

So, security at Target is a main centrally thought about kind of thing. So, it’s very much something that is a part of the way that people operate inside of Target. So, coming over to Shipt, obviously, Shipt is—it is a subsidiary. It is a part of Target, but it doesn’t have that long history and hasn’t had that same kind of experience. The biggest thing that we really needed at Shipt is first and foremost to get the program established, right. So, I’m three or four months onto the job now and we’ve tripled the team size. I’ve been—

Corey: And you’ve stayed out of the headlines, which is basically the biggest and most accurate breach indicator I’ve found so far.

Dan: So far so good. Well, but the thing that we want to do though is to be able to bring that same kind of focus of importance that Target has on cybersecurity into the world of engineering at Shipt. And it’s not just a compliance game, and it’s not just a thing where we’re just trying to say that we have it. We’re actually trying to make sure that as we go forward we’ve got all these best practices from an organization that’s been through the bad stuff that we can adopt into our day-to-day and kind of get it done.

When we talk about it at an executive level, obviously we’re not talking about the penetration tests done by the red team the earlier day, right. We’re not calling any of that stuff out in particular. But we do try to summarize it in a way that makes it clear that the thing that we’re trying to do is build a security-minded culture and not just check some boxes and make sure that we have the appropriate titles in the appropriate places so that our insurance rates go down, right. We’re actually trying to keep people safe.

Corey: There’s a lot to be said for that. With the Target breach back in—I want to say 2012, was it?

Dan: 2012. Yep.

Corey: Again, it was a wake-up call and the argument that I’ve always seen is that everyone is vulnerable—just depends on how much work it’s going to take to get there. And for, credit where due, there was a complete rotation in the executive levels which whether that’s fair or not, I—people have different opinions on it; my belief has always been you own the responsibility, regardless of who’s doing the work.

And there’s no one as fanatical as a convert, on some level, and you’ve clearly been doing a lot of things in the right direction. The thing that always surprises me is that when I wind up seeing these surveys in the industry that—what is it? 65% of companies say that they would be vulnerable to a breach, and everybody said, “Oh, we should definitely look at those companies.” My argument is, “Hang on a sec. I want to talk to the 35% who say, ‘oh, we’re impenetrable.’” because, spoiler, you are not.

No one is. Just the question of how heavy is the lift and how much work is it going to take to get there? I do know that mouthing off in public about how perfect the security of anything is, is the best way to more or less climb to the top of a mountain during a thunderstorm, a hold up a giant metal rod, and curse the name of God. It doesn’t lead to positive outcomes, basically ever. In turn, this also leads to companies not talking about security openly.

I find that in many cases it is easier for me to get people to talk about their AWS bills than their InfoSec posture. And I do believe, incidentally, those two things are not entirely unrelated, but how do you view it? It was surprisingly easy to get Shipt’s CISO to have a conversation with me here on this podcast. It is significantly more challenging in most other companies.

Dan: Well, in fairness, you’ve been asking me for about two-and-a-half years pretty regularly [laugh] to come.

Corey: And I always say I will stop bothering you if you want. You said, “No, no. Ask me again in a few months. Ask me again, after the election. Ask me again after—I don’t know, like, the one-day delivery thing gets sorted out.” Whatever it happens to be. And that’s fine. I follow up religiously, and eventually I can wear people down by being polite yet persistent.

Dan: So, persistence on you is actually to credit here. No, I think to your question though, I think that there’s a good balance. There’s a good balance in being open about what it is that you’re trying to do versus over-sharing areas that maybe you’re less proficient in, right. So, it wouldn’t make a lot of sense for me to come on here and tell you the areas that we need to develop into security. But on the other side of things, I am very happy to come in and talk to you about how our incident response plan is evolving, right, and what our plan looks like for doing all of that kind of stuff.

Some of the best security practitioners who I’ve worked with in the world will tell you that you’re not going to prevent a breach from a motivated attacker, and your job as CISO is to make sure that your response is appropriate, right, more so than anything. So, our incident response areas where today we’re dedicating quite a bit of effort to build up our proficiency, and that’s a very important aspect of the cybersecurity program that we’re trying to build here.

Corey: And unlike the early days of a campaign, you still have to be ultra-conscious about security, but now you have the luxury of actually being able to hire security staff because it turns out that, “Please come volunteer here,” is not presumably Shipt’s hiring pitch.

Dan: That’s correct. Yeah, exactly. We have a lot of buy-in from the rest of leadership to build out this program. Shipt’s history with cybersecurity is one where there were a couple of folks who did a remarkably good job for just being two or three of them for a really long period of time who ran the cybersecurity operation very much was not a part of the engineering culture at Shipt, but there still was coverage.

Those folks left earlier in the year, all of them, simultaneously, unfortunately. And that’s sort of how the position became open to me in the first place. But it also meant that I was quite literally starting with next to nothing, right. And from that standpoint it made it feel a lot like the early days of the campaign because I was having to build a team from scratch and having to get people motivated to come and work on this thing that had kind of an unknown future roadmap associated with it and all of that kind of stuff.

But we’ve been very privileged to—because we have that leadership support we’re able to pay market rates and actually hire qualified and capable and competent engineers and engineering leaders to help build out the aspects of this program that we need. And like I said, we’ve managed to—we weren’t exactly at zero when I walked in the door. So, when I say we were able to quadruple the team, it doesn’t mean that we just added four zeros there, [laugh] but we’ve got a little bit over a dozen people focusing on all areas of security for the business that we can think of. And that’s just going to continue to grow. So, it’s exciting; it’s a challenge. But having the support of the entire organization behind something like this really, really helps a lot.

Corey: I know we’re running out of time for a lot of the interview, but one more question I want to ask you about is, when you’re the CISO for a nationally known politician who is running for the highest office, the risk inherent to getting it wrong is massive. This is one of those mistakes will show indelibly for the rest of, well, one would argue US history, you could arguably say that there will be consequences that go that far out.

On the other side of it, once you’re done on the campaign you’re now the CISO at Shipt. And I am not in any way insinuating that the security of your customers, and your partners, and your data across the board is important. But it does not seem to me from the outside that it has the same, “If we get this wrong there are repercussions that will extend into my grandchildren’s time.” How do you find that your ability to care as deeply about this has changed, if it has?

Dan: My stress levels are a lot lower I’ll say that, but—

Corey: You can always spot the veterans on an SRE team because—when I say veterans I mean veterans from the armed forces because, “No one’s shooting at me. We can’t serve ads right now. I’m really not going to run around and scream like, ‘My hair’s on fire,’ because this is nothing compared to what stress can look like.” And yeah there’s always a worst stressor, but, on some level, it feels like it would be an asset. And again this is not to suggest you don’t take security seriously. I want to be very clear on that point.

Dan: Yeah, yeah, no. The important challenge of the role is building this out in a way that we have coverage over all the areas that we really need, right, and that is actually the kind of stuff that I enjoy quite a bit. I enjoy starting a program. I enjoy seeing a program come to fruition. I enjoy helping other people build their careers out, and so I have a number of folks who are at earlier at points in their career who I’m very happy that we have them on our team because I can see them grow and I can see them understand and set up what the next thing for them to do is.

And so when I look at the day-to-day here, I was motivated on the campaign by that reality of like there is some quite literal life or death stuff that is going to happen here. And that’s a really strong presser to make sure that you’re doing all the right stuff at the right time. In this case, my motivation is different because I actually enjoy building this kind of stuff out and making sure that we’re doing all the right stuff and not having the stress of, like, this could be the end of the world if we get this wrong.

Means that I can spend time focusing on making sure that the program is coming together as it should, and getting joy from seeing the program come together is where a lot of that motivation is coming from today. So, it’s just different, right? It’s a different thing, but at the end of the day it’s very rewarding and I’m enjoying it and can see this continuing on for quite some time.

Corey: And I look forward to ideally getting you back in another two-and-a-half years after I began badgering you in two hours in order to come back on the show. If—

Dan: [laugh].

Corey: —people want to hear more about what you’re up to, how you view about these things, potentially consider working with you, where can they find you?

Dan: Best place although I’ve not been as active because it has been very busy the last couple of months, but find me on Twitter, @danveloper, find me on LinkedIn. Those—you know, I posted a couple of blog posts about the technology choices that we made on the campaign that I think folks find interesting, and periodically I’ll share out my thoughts on Twitter about whatever the most current thing is, Kubernetes or AWS about to go down or something along those lines. So, yeah, that’s the best way. And I tweet out all the jobs and post all the jobs that we’re hiring for on LinkedIn and all of that kind of stuff. So, usual social channels. Just not Facebook.

Corey: Amen to that. And I will of course include links to those things in the [show notes 00:37:29]. Thank you so much for taking the time to speak with me. I appreciate it.

Dan: Thank you, Corey.

Corey: Dan Woods, CISO and VP of Cybersecurity at Shipt, also formerly of the Biden campaign because wherever he goes he clearly paints a target on his back. I’m Cloud Economist, Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast please leave a five-star review on your podcast platform of choice along with an incoherent rant that is no doubt tied to either politics or the alternate form of politics: Spinnaker.

Dan: [laugh].

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Julia

Julia Ferraioli calls herself an Open Source Archaeologist, focusing on sustainability, tooling, and research. Her background includes research in machine learning, robotics, HCI, and accessibility. Julia finds energy in developing creative demos, creating beautiful documents, and rainbow sprinkles. She’s also a fierce supporter of LaTeX, the Oxford comma, and small pull requests.

Links:

  • Open Source Stories: https://www.opensourcestories.org

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: It seems like there is a new security breach every day. Are you confident that an old SSH key, or a shared admin account, isn’t going to come back and bite you? If not, check out Teleport. Teleport is the easiest, most secure way to access all of your infrastructure. The open source Teleport Access Plane consolidates everything you need for secure access to your Linux and Windows servers—and I assure you there is no third option there. Kubernetes clusters, databases, and internal applications like AWS Management Console, Yankins, GitLab, Grafana, Jupyter Notebooks, and more. Teleport’s unique approach is not only more secure, it also improves developer productivity. To learn more visit: goteleport.com. And not, that is not me telling you to go away, it is: goteleport.com.

Corey: This episode is sponsored in part by our friends at Redis, the company behind the incredibly popular open source database that is not the bind DNS server. If you’re tired of managing open source Redis on your own, or you’re using one of the vanilla cloud caching services, these folks have you covered with the go to manage Redis service for global caching and primary database capabilities; Redis Enterprise. To learn more and deploy not only a cache but a single operational data platform for one Redis experience, visit redis.com/hero. Thats r-e-d-i-s.com/hero. And my thanks to my friends at Redis for sponsoring my ridiculous non-sense.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. My guest today is someone I have been very politely badgering to come on the show for a while, ever since I saw her speak a couple years ago in the Before Times, at Monktoberfest. As I’ve said before, anytime the RedMonk folks are involved in something, it is something you probably want to be involved in. That is my new guiding star philosophy when it comes to conferences, Twitter threads, opinions, breakfast cereals, you name it. Please welcome Julia Ferraioli, the co-founder of Open Source Stories, Julia, thank you for joining me today.

Julia: Thank you for having me. And I definitely agree on the RedMonk side of things. They are fantastic folk.

Corey: They’re a small company, which is sort of interesting to me from a perspective of just how outsized their impact on this entire industry is. But it’s, I’ve had as many of them as they will let me have on the show. They are welcome to come back whatever they want, just because they—every single one of them, though they’re very different from one another, make everyone around them better with their presence. And that’s just a hard thing to see. I didn’t mean to turn this into a love letter to RedMonk, but here we are.

Julia: I don’t mind it. They have the ability to amplify the goodness that they see, anything from their survey designs to just how they interact online. It’s wonderful to see.

Corey: Speaking of amplifications, you are the co-founder of Open Source Stories, the idea of telling the—to my understanding—the stories behind open source. Like this is sort of like—what is it, Behind the Music, only in this case it’s Behind the Code? I mean, how do you envision this?

Julia: Oh, I like that framing. So, Open Source Stories is a project that myself and Amanda Casari founded not that terribly long ago because when we were doing research about how to model open source and open source ecosystems, we realized that a lot of the research papers that have been published about open source are pulled mostly from GitHub Archive, which is this repository of GitHub data. It could be the actual Git commit history as well as the activity streams from GitHub as well, but that doesn’t capture a lot of the nuances behind open source, things like the narratives, how communities interact, where communication is happening, et cetera. All of these things can happen outside of the hosting platform. So, we launched this project to help tell these stories of the people and events and scenarios behind the open source projects that really power our industry.

Corey: I’m going to get letters for this one, I’m sure of it, but I’ve been involved in the open source ecosystem for a while and I’ve noticed that there’s been a recurring theme among various projects, particularly the more passionate folks working on them, where they talk an awful lot but they aren’t very good at telling stories at the same time. And nowhere is this more evident than when we look at what passes for a lot of these projects’ documentation. One of the transformative talks that I went to was Jordan Sissel’s years and years ago, at the Southern California Linux Expo. And it was a talk about LogStash, which doesn’t actually matter because the part that he’d said that really resonated with me, that his whole theme of his talk was around, was if a new user has a bad time, it’s a bug. And the idea that, “Oh, you didn’t read the documentation properly.”

When about I started working with Linux, in some IRC chat rooms, the standard response to someone asked for help was to assume that they’re an idiot, begin immediately accosting them with RTFM, for Read the Frickin’ Manual, and then look for ways that you could turn this back around on them and make it their fault. And I looked at this and at the time, it’s like, “Wow, these are people that are mean to other people,” and I was a small, angry teenager; it’s like, “This is my jam. Here I am.” And yeah, many decades later, I’m looking at this and I feel a sense of shame because that’s not the energy I want to put into the world. A lot of those communities have evolved and grown and what used to be the area and arena for hobbyists is now powering trillion-dollar companies.

Julia: Absolutely. I like the whole, “If the user has a bad experience, that’s a bug,” because it absolutely is. And I feel like a lot of these projects haven’t invested nearly as much into the user experience as they have into polishing the code. And the attitude that that kind of perpetuates throughout the project about how you treat your users, it’s pervasive and it really sets up the types of features that you develop, the contributors that you encouraged to commit to the project, and it just creates a—to put it minorly—less than welcoming environment for users, contributors, maintainers alike. And we don’t really need that sort of hostility, especially when we’re talking about projects that underpin the foundations, in some cases, of the internet.

Corey: When we look at what open source is, I mean, I shortcut to thinking in terms of the context through which I’ve always approached it, which was generally code, or in my sad, particular story, back in the olden days on good freenode, when that was where a lot of this discourse happened, I was network staff and helping a bunch of different communities get channels set up through a Byzantine process. Because of course there was a Byzantine process; it was an open source community, and if there’s one thing we love in open source, it is pretending to be lawyers when we’re not. And we’re sort of cargo-culting what we think process and procedure often look like. So yeah, there was a bunch of nonsensical paperwork happening there, but it was mostly about helping folks collaborate and communicate. But I’ve first and foremost, think in terms of code and in terms of community. What is open source to you?

Julia: Well, I entered open source in the Sourceforge days, when all you had to do was go and download some code from the internet and hit the right download button, make sure not to hit one of the extraneous ones. And all you need for that is for the code to be under the right license. And to an extent that’s what’s true today for open source. At the heart of it, this minimum criteria for what constitutes open source is, “Okay, does it comply with the open source definition that the Open Source Initiative puts forth?” Now, I understand that not everybody necessarily agrees with the Open Source definition, but it’s useful as a shortcut for how we think about the basic requirements. But what I find when people are talking about open source online is that they have these very different models. You’ll hear from people that, “Okay, well, if it doesn’t have a standard governance model, it’s not really open source.”

Corey: The ‘No True Scotsman’ argument.

Julia: Yeah. So, I find that we’ve got these different expectations for what open source is, and that leads to us talking past each other or discounting different types of open source when what we really need to do is come up with better language, a better vocabulary, for how to talk about these things. So, for example, I used to work in developer relations, and in developer relations one of the big things that you do is release sample code. Now, oftentimes, I’m not looking for that sample code to be picked up by a bunch of different developers and incorporated as a library into their project—

Corey: [laugh]. Well, that’s your error in that case because congratulations, that’s running in production at a bank somewhere, now.

Julia: Oh, I know. And that has definitely happened with my code, and I’m ashamed to say that. [laugh]. But generally speaking, you’re not looking to build a huge community around sample code, right?

Corey: You say that, but that again, Stack Overflow, it was—

Julia: Okay.

Corey: —[unintelligible 00:09:22] done rather well. So, there’s that.

Julia: Well yes, that is true, but when you release code on Stack Overflow, or GitHub, or in a Jest, or just on your blog, the thing that allows the bank to come in and incorporate that into their own application, or to even just learn from it, is the fact that it is open source. Now, it doesn’t have a lot of the things that a community like Python or Kubernetes has, but it is still open source; it just has a different purpose than those communities and those ecosystems.

Corey: So, I think it is challenging right now to talk about open source as if it were the same type of thing that it was back in the ’90s, and the naughts—and even the teens—where it’s a bunch of, more or less either hobbyists or people are perceived to be hobbyists. Sure, an awful lot of them are making commits from their redhat.com email address, but okay. And some of these people are increasingly being paid to work places, but then you see almost—I don’t necessarily agree with the framing of The New York Times article by Daisuke Wakabayashi—who’s a previous guest on the show—of Amazon strip-mining open source, but they definitely are in there—and other companies as well—are sort of appropriating it, or subverting it, or turning it into something that it was not previously, for lack of a better term. What’s your take on that?

Julia: Oh, that’s a hard one. From a fundamentals perspective, that is absolutely within their rights under the definition of open source, and in some cases, the spirit of open source as well.

Corey: Oh, and I would argue with someone who said that they should be constrained from doing this as far as a matter of legalities, or rights, or ridiculous Looney Tunes license changes.

Julia: Well, there are definitely folks who are trying to make that the case.

Corey: Yeah. Oh, yeah. I’m on the position of, they’re within their rights to do it, but it’s time for a good old fashioned public shunning as a result.

Julia: I’m not sure I agree. I think that it is a natural consequence of how open source has gained in popularity and, in some cases, it’s a testament to open source’s success. Now, does it pose some serious challenges for the open source community and open source ecosystem? Absolutely because this is a new way of using open source that was unanticipated, and in fact, could be characterized as a Black Swan event in [open source-ware 00:12:18].

Corey: The fundamental attribution error that I see, back at the very beginning, was that what we wrote the software, therefore, we are the best in the world at running it, therefore, if there’s going to be a managed service, clearly ours will be the best. Amazon’s core strength has apparently been operational excellence as they like to call it; my position on that is a little bit less of tying into the mystery, a little bit more of they’re really fast and getting paged and fixing things in a hurry before customers notice. So okay, great, but it’s column A, column B, whatever. The bigger concern I have with Amazon as its product strategy is, “Yes.” If it were just a way to run EC2 instances or virtual machines, then sure, that’s great.

And every open source project should, on some level, see some validation of its market through a lens of, “Oh, we’re getting some competition. That’s great.” The challenge I see is that in the line of competitors, Amazon is at or near the front all the time on basically everything. And it’s if they would pick a lane to stay in, great.

Google is a good example of this. There are things that Google very strongly considers in its wheelhouse, but for other things, they partner with the open source-based company in question to create a managed service partner offering and that’s great. Amazon pulls a, “Nope. We’re just going to build this out as first-party. The end.”

And they compete with everyone, including themselves on almost every axis. And that’s where it just gets into a, “Leave some oxygen for the rest of us.” I mean, it feels like they lie awake at night worrying that someone who isn’t them somehow making money somewhere. That is, I think, on some level, more of the Black Swan event than someone else deciding that they can host a particular open source project more effectively. But that’s where I stand. And again, this is just me as an enthusiastic and obnoxious observer. You’re operating in this space. What do you think? That’s the important part of the story.

Julia: Well, I mean, you definitely have a point, Amazon—or AWS, maybe not necessarily Amazon—takes on different technologies far and wide, so they’re not limiting themselves to a space. But that said, I think it comes down less to what is possible with open source and what is okay under the guise of open source, and what is good for the open source ecosystem. And when you fork a project, you do have to understand that you are bifurcating the open source ecosystem. And that can lead to sustainability problems down the road. So, I think the jury is still out on whether forking a project, running it as a managed service—as Amazon is doing with some of the open source projects—if that’s going to come back to bite them just from a developer community standpoint because you’re going to have people committing to one or the other, but possibly not both.

Corey: I think this is why Amazon—I know, they’re very annoyed by their perception in the open source ecosystem, but you take a look at other large tech companies, and almost all of them have a few notable open source projects that started life there. For example, we have—I think Cassandra came out of Facebook, but don’t quote me on that; Kubernetes came out of Google, a fact for which they steadfastly refused to apologize, so far; and so on, and so forth. But Amazon’s open source initiatives have been, “We’ve open sourced this thing that is basically only used at Amazon.” Or, my personal favorite, we’ve put all of our documentation up on GitHub so that you can write a corrections to it yourself from the community, which I’m hearing as, “Please, volunteer for a $1.6 trillion company so that they don’t have to improve their documentation by hiring expensive people internally.”

You can sort of guess my position on that. It seems like they have not launched anything that has a deep heart within Amazon that is broadly adopted outside of their walls. My question for you is, do you believe that having that level of adoption externally is required for a healthy open source project?

Julia: Again, I think it goes back to the goals of why you’re open-sourcing something. I don’t believe that it’s necessarily required for the open source project to be quality and be usable, but if your goal is adoption or if your goal is to get ideas and best practices out there, then yeah, you do need that engagement by the broader community, you do need the contributors. But there are a lot of cases where open-sourcing technology is more for the validation, rather than the adoption of the tech. So, it really depends.

Corey: I’d say the most cynical reason I’ve seen to open source things comes from Netflix, where they have a recurring pattern of open-sourcing something, there are two or three commits, and then it basically sits there unattended. What I firmly believe is happening is that a senior engineer at Netflix is working on the thing and they’re about to change jobs, so they open source the project so that they can change jobs and then pick up where they left off with an internal fork, I view it as a game of, basically, they’re passing themselves a football as they run across the street. And people laugh when I say that, but I’ve also had people over drinks say, “You are closer than you might think, sometimes.” Which on some level is terrifying. Feels like life is imitating art, but here we go.

Julia: That definitely happens, and I have seen it [laugh] as well. People want to essentially use open source to exfiltrate IP.

Corey: Yeah. Only doing it legitimate way as opposed to the, “Please don’t—hope they don’t find that USB stick I’ve hidden in my sock on my last day.”

Julia: Yes. And this is why open source offices have a challenging job in helping facilitate the release of open source software. So, it is hard to ascertain when that is happening.

Corey: Yeah, no company is ever going to have a big statement that is going to be anything other than, honestly, marketing speak when it comes time to explain why they’re doing a certain thing. It’s, “Oh, yeah, we’re open-sourcing this so we don’t get sued in three years by this other company that might prove to be a competitive threat.” Or, “We’re open-sourcing this as a hiring and recruiting technique.” I mean, I would argue, it wasn’t open source, but one of the best approaches that I’ve seen from that perspective came out of Google, I’m firmly convinced to this day that App Engine was run not by their SRE team, but by their recruiting arm, “Because if you can build a great app on App Engine, well, this is, kind of like, how we think about things inside of Google; come and work here,” either via acqui-hiring or a just outright interview funnel. Maybe that’s too cynical, too, but again, that leads to the question of is it really open source when it has these deep ties to specific platforms?

Here’s an open source tool that presumes you’re running on top of AWS. Well, great, sure it’s built by the community and anyone can access these things, but without paying per second to a cloud provider, probably the referenced cloud provider they’re developing this against, it’s not going to get very far. So, it’s a nuanced argument, and there are shades of that nuance to every aspect of it. And if there’s one thing that Twitter is terrible at is capturing nuance in 280 characters. And even in the, “All right, this is my nuanced take on open source in this thread, I will tweet, one of 5,712.” Great. That’s not really the forum for that either. And people lose sight of nuance. It’s a sticky, delicate thing, and it feels like a lot of the open source community has been enthusiastically agreeing with each other—sometimes violently so—but they’re not sharing a common language in which to do it.

Julia: Yeah. And in terms of the purposes of open source projects, it is okay for them to have different ones as long as they’re telegraphing those purposes to their users and the people who are looking at the projects for their own use. But whether it’s open source? I think it’s okay for that to be the baseline and then build out the vocabulary of the types of projects that you want from there, based on those expectations. Yes, this particular technology only works with this cloud provider. That’s open source that facilitates and accelerates development with that cloud provider.

Corey: This episode is sponsored by our friends at Oracle Cloud. Counting the pennies, but still dreaming of deploying apps instead of "Hello, World" demos? Allow me to introduce you to Oracle's Always Free tier. It provides over 20 free services and infrastructure, networking, databases, observability, management, and security. And—let me be clear here—it's actually free. There's no surprise billing until you intentionally and proactively upgrade your account. This means you can provision a virtual machine instance or spin up an autonomous database that manages itself all while gaining the networking load, balancing and storage resources that somehow never quite make it into most free tiers needed to support the application that you want to build. With Always Free, you can do things like run small scale applications or do proof-of-concept testing without spending a dime. You know that I always like to put asterisks next to the word free. This is actually free, no asterisk. Start now. Visit snark.cloud/oci-free that's snark.cloud/oci-free.

Corey: I always try and stay away from explicit value judgments on a lot of these things because it’s nuanced, and no one who doesn’t work at Facebook wakes up expecting to do terrible things today. We’re all trying to do the best we can with the constraints are operating within. The challenge is that when you’re at a company like an AWS, or a Google, or a Microsoft, or one of these giant companies, the same pressures that the rest of the quote-unquote “mere mortals” in ecosystem have to contend with are very different. But talking to people who work at these big companies, they have meetings and review processes that here at my twelve-person company, I don’t even have to consider.

Easy example of that: Never once have I put something out into the world and had a single discussion about is this going to get us in trouble with respect to antitrust? That has never been on my radar as far as things I have to care about. Even at my previous job at a highly regulated financial company, where you could argue that they are approaching monopoly status in some areas of the market organically, with passive investing being what it is, great, their open source discussions were always much more aligned with what licenses are we willing to accept legal risk for using internally? Because there are things that are—like IP is why we have a business in many respects, so anything that touches that theoretically means we’d have to disclose how the entire system, how the rest of it works, is not allowed to be used here. And there are reviews and processes and compliance requirements for that.

I get that concern, and at a certain point of scale, you’re negligent if you don’t have a function that looks at it through that lens. But I look back to the early days of just puttering around with, “I want to do a thing and I found this project somewhere that people are excited about,” in the pre-GitHub days, I can download it off as Sourceforge or whatnot and I can make it work. And but it doesn’t do this one thing I want to do, “Hey, the code’s available. Can I fix it myself? Absolutely not. I’m crap at writing code. But I can talk to people and piece it together from wisdom that they offer.” And it turns into something awful until finally it gets enough traction that someone who knows what they’re doing looks at it and refactors and it makes it good.

And that’s the open source community I recognize and that I see from my early developmental period. I don’t recognize what we see in ecosystem today through that same lens of, “Okay, go online. Be nice to people”—well, that’s new—“See how this thing works. And oh, if I’m having a problem, I’m probably not the only person who’s having a problem like this.” You have to get really good at using Google more than you do at writing code in some respects. But at that point, it’s almost entirely a copy-and-paste, except that’s not technical enough for the open source world. So instead, we have to learn the 500 arcane subcommands to Git in order to get it out there. But it works. Ish.

Julia: I think that community is still out there. I really do. I think that it is harder to find and it’s not necessarily where you might tend to look, but those projects are still there. They’re still running. They might be a little less high-profile than a lot of the ones that are getting a lot of attention right now, but they are still there.

Corey: On some level, it feels like the blame for this lies—at least partially—at the feat of Slack and its success because it used to be that you had IRC, that was how folks communicated. And I remember the early days of that and things like Jabber or internal servers, grea—or internal IRC servers at companies—great, you’d have engineering all talking on that, and oh, you want to have someone in finance or marketing join that thing? Yeah, the short answer is, that won’t be happening. But you can try and delude yourself and set it up with a special client and the rest.

Slack removed all of that friction, but it’s balkanized to the point where every once in a while, I have to go through and remove a bunch of Slack channels slash workspaces slash whatever we’re calling them this week from my desktop client because it’s basically eating all the RAM like it’s trying to be Google Chrome. And then it’s great, but there’s no universal federated thing the way that there was with IRC where I just pop in a different channel for a different project. And IRC is still there and it comes back to life whenever Slack takes an outage. And then Slack gets fixed, it sort of bleeds off again. But I don’t want to be in 500 different Slack workspaces, one for every open source project that I’m using, and there’s no coherent sense of identity and community anymore the way there once was. And I feel like I’m old man yelling at the passing of time at this. But you’re right, open source to me was always much more about community than it was about code.

Julia: Yeah, and I think that we do not talk about the impact of tools for open source that we use. Because you’re right; with IRC, it was unified. You could pretty much guarantee that projects of a certain size were present there. And with Slack, you have to sign up for yet another account, not quite yet sure why I can’t find the right channels that I need to join in Slack. So, there’s a lot of navigation and a lot of prerequisite knowledge that you need to have in order to be productive.

And then you’ve got other tools being used for communication by other communities like, I believe Gitter is a major one as well. Then you have to make sure that you’re up-to-date with all of these different interfaces, Discord, everything. And the sociological implication of that shouldn’t be underestimated. What are you going to do if you find a project that uses a communication tool that you just really don’t want to use or don’t want to sign up for yet another account? Maybe you pass on by and you find one that works within your existing set of tools. There aren’t a lack of open source projects to join right now. You can be choosy. And we don’t yet know what the impact is of that.

Corey: It’s challenging. There’s no good answer that I found that solves all of these things. It’s become so balkanized, on some level, that every project out there that I see—and there are some small ones that are incredibly foundational to, basically, civilization as we know it, but it’s not working right because it’s you have to figure out where they are and what the community norms are because they change from project to project, and there are so many different things. And, like, you can go into NPM and install some relatively trivial thing that does command-line string processing, or whatnot, and it installs 40 different dependencies. And there’s a problem and you want to figure out exactly how that works, and et cetera, et cetera, et cetera.

Julia: Absolutely. With NPM specifically, or Node specifically, it is interesting that the development model kind of encourages this obscurity, an obfuscation of a functionality. So, it is hard to go in, debug an issue, go to the specific community, understand how they work, contribute a patch, just to fix something that is, you know, five levels up. It gets confusing for developers. It can contribute to longer-term bugs that we see propagate throughout the system. It is not an easy problem to solve, and I have a lot of sympathy for newcomers to the open source ecosystem because it is so hard to navigate. And I think that’s an as yet unsolved problem that we need to address.

Corey: So, what was it that inspired you to create Open Source Stories? I mean, I love the direction you’re taking this in; I love the way you’re thinking about [audio break 00:29:38]. Where did it come from? What started this?

Julia: Well, when Amanda and I were going back and doing research around—you know, aside from the code for an open source project, where are the different entry points? Where are the different interaction points between projects, ecosystems, and the industry? And we did a couple of interviews, just very organic interviews, with some subject matter experts in Node, in Python, in Go. And there was a point where we stopped—or at least I stopped taking notes because I was just so fascinated by the narrative that our interviewee was putting forth and was talking about. And what we wanted was for it to not just be this meeting between a few people, we wanted to be able to share that with anyone. And so one of the things that really inspired us was StoryCorps, which allows you to record, much like we’re doing today, 40 minutes worth of interactions between one to three people.

Corey: Oh, we’re going to cut it down to five minutes at most. Like, one question; one answer. Boom, we’re done.

Julia: [laugh].

Corey: I kid, I kid.

Julia: But it’s really about facilitating the sharing of knowledge and sharing of these oral histories. Because as you’re doing research into interactions in specific open source communities, you’ll get articles, you’ll get changelogs, all of that good stuff, but you won’t get the nuance that we’ve been talking about over the course of this podcast. You lose the story behind the story, right? How are decisions made? How are people thinking about the interactions with their users? What are the turning points for a project? What are those conversations between the maintainers that changed the entire game?

Those are the sorts of stories that we’re hoping to capture because they’re important for history, for knowledge sharing, for learning from our past, and making decisions for the future. And so that’s really what we wanted to capture. And we wanted to capture the narratives behind the people that don’t necessarily show up in the codebase, too: Talking about the designers, the product managers, the marketers behind open source that make it successful. Because there’s so much more than code.

Corey: Oh, my God, yes. It’s… how do I put this politely without getting letters? Well, I guess I’ll take a stab at it and see how it plays out. I look at so much of the brilliant code that has been written, and the documentation is abhorrent, and the design of the site, and the icon, and the interface, it looks like a joke that I put on Twitter trying to be funny. It’s, the code is important, don’t get me wrong, but there’s so much more to it than that.

And we see this in the industry, too, where companies have gone out of business, trying to get their codebase just right. It’s, yeah, you can launch code that is really, really bad, but if you have product-market fit, it is survivable. I’ve heard stories in the early days of Twitter that we saw the fail whale all the time because it was an abhorrent monstrosity, to the point it became a running joke. But it turns out, when you hit product-market fit, you can afford really good engineers to come in and fix a lot of that stuff. That stuff is more important than the quality of the code, and that is something that I think that we have a collective industry-wide delusion about. And it’s a blind spot for us.

Julia: Yeah. I think we get wrapped up in the cleverness of the tech, and I’ve fallen prey to this, too. I get so involved in how I’m solving the problem and forget about the actual problem that I’m trying to solve, right? It’s not necessarily about the how, but about the what. And without your fantastic tech writers, designers, usability experts, your open source project is going to be your open source project. It’s not going to necessarily get that wide adoption, if that is indeed your goal for the technology that you’re releasing.

So, it really is about making sure that as we’re launching and working on these open source projects and ecosystems, that we are inviting people to the table that have these other unique skills that goes beyond that code and speaks to what makes the project different and unique.

Corey: I really want to say how much I appreciate your taking the time to talk to me about this. If people want to get involved themselves, how do they do that? Because I have a hard time accepting that you’re doing something called Open Source Stories that eschews community involvement.

Julia: Yeah. So, we absolutely would love more folks to get involved. I have been primarily the person working on the site, so we can always use contributors to the site itself, but we also want more storytellers and facilitators. And so if you go to opensourcestories.org, we’ve got a page specifically designed to facilitate contributions. So, check that out, and we look forward to hearing from anyone who wants to participate.

Corey: And we will, of course, include links to that in the show notes. Thank you so much for taking the time to speak with me today. I really appreciate it.

Julia: Thanks for having me.

Corey: Julia Ferraioli, co-founder of Open Source Stories. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry comment, calling me a fool because I did not bother to RTFM first.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Jake

Technical Lead by day at the Met Office in the UK, leading a team of software developers delivering services for the UK. By night, gamer and fitness instructor, attempting to get a home cinema and gaming setup whilst coralling 3 cats, 2 rabbits, 2 fish tanks, and my wonderful girlfriend.

Links:

  • Met Office: https://www.metoffice.gov.uk
  • Twitter: https://twitter.com/jakehendy

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: It seems like there is a new security breach every day. Are you confident that an old SSH key, or a shared admin account, isn’t going to come back and bite you? If not, check out Teleport. Teleport is the easiest, most secure way to access all of your infrastructure. The open source Teleport Access Plane consolidates everything you need for secure access to your Linux and Windows servers—and I assure you there is no third option there. Kubernetes clusters, databases, and internal applications like AWS Management Console, Yankins, GitLab, Grafana, Jupyter Notebooks, and more. Teleport’s unique approach is not only more secure, it also improves developer productivity. To learn more visit: goteleport.com. And not, that is not me telling you to go away, it is: goteleport.com.

Corey: This episode is sponsored in part by our friends at Redis, the company behind the incredibly popular open source database that is not the bind DNS server. If you’re tired of managing open source Redis on your own, or you’re using one of the vanilla cloud caching services, these folks have you covered with the go to manage Redis service for global caching and primary database capabilities; Redis Enterprise. To learn more and deploy not only a cache but a single operational data platform for one Redis experience, visit redis.com/hero. Thats r-e-d-i-s.com/hero. And my thanks to my friends at Redis for sponsoring my ridiculous non-sense.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. It’s often said that the sun never sets on the British Empire, but it’s often very cloudy and hard to see the sun because many parts of it are dreary and overcast. Here to talk today about how we can predict those things in advance—in theory—is Jake Hendy, Tech Lead at the Met Office. Jake, thanks for joining me.

Jake: Hey, Corey, it’s lovely to be here. Thanks for inviting me on.

Corey: There’s a common misconception that its startups in San Francisco or the culture thereof, if you can even elevate it to being a culture above something you’d find in a petri dish, that is where cloud stuff happens, where the computer stuff is done. And I’ve always liked cutting against that. There are governments that are doing interesting things with Cloud; there are large companies and ‘move fast and break things’ is the exact opposite of what you generally want from institutions that date back centuries. What’s it like working on Cloud, something that for all intents and purposes didn’t exist 20 years ago, in the context of a government office?

Jake: As you can imagine, it was a bit of a foray into cloud for us when it first came around. We weren’t one of the first people to jump. The Met Office, we’ve got our own data centers, which we’ve proudly sit on that contains supercomputers and mainframes as well as a plethora of x86 hardware. So, we didn’t move fast at the start, but nowadays, we don’t move at breakneck speeds, but we like to take advantage of those managed services. It gets out of the way of managing things for us.

Corey: Let’s back up a second because I tend to be stereotypically American in many ways. What is the Met Office?

Jake: What is the Met Office? The Met Office is the UK’s National Meteorological Service. And what does that mean? We do a lot of things though with meteorology, from weather forecasting and climate research from our Hadley Centre—which is world-renowned—down to observations, collections, and partnerships around the world. So, if you’ve been on a plane over Europe, the Middle East, Africa, over parts of Asia, that plane took off because the Met Office provided a forecast for that plane. There’s a whole range of things we can talk about there, if you want Corey, of what the Met Office actually does.

Corey: Well, let’s ask some of the baseline questions. You think of a weather office in a particular country as, oh okay, it tracks the weather in the area of operations for that particular country. Are you looking at weather on a global basis, on a somewhat local basis, or—as mentioned—since due to a long many-century history it turns out that there are UK Commonwealth territories scattered around the globe, where do you start? Where do you stop?

Jake: We don’t start and we don’t stop. The Met Office is very much a 24/7 operation. So, we’ve got a 24/7 operation center with staff constantly manning it, doing all sorts of things. So, we’ve got a defense, we work heavily with our defense colleagues from UK armed forces to NATO partners; we’ve got aviation, as mentioned; we’ve got marine shipping from—most of the listeners in the UK will have heard of the shipping forecast at one point or another. And we’ve got private sector as well, from transport, to energy, supermarkets, and more. We have a very heavy UK focus, for obvious reasons, but our remit goes wide. You can actually go and see some of our model data is actually on Amazon Open Data. We’ve got MOGREPS, which is our ensemble forecast, as well as global models and UK models, with a 24-hour time lag, but feel free to go and have a play. And you can see the wide variety of data that we produce in just those few models.

Corey: Yeah, just pulling up your website now; looking at where I am here in San Francisco, it gives me a detailed hour-by-hour forecast. There are only two problems I see with it. The first is that it’s using Celsius units, which I—

Jake: [laugh].

Corey: —as a matter of policy, don’t believe in because in this country, we don’t really use things that make sense in measuring context. And also, I don’t believe it’s a real weather site because it’s not absolutely festooned with advertisements for nonsense, which is apparently—I wasn’t aware—a thing that you could have on the internet. I thought that showing weather data automatically meant that you had to attempt to cater to the lowest common denominator at all times.

Jake: That’s an interesting point there. So, the Met Office is owned and operated by Her Majesty’s Government. We are a Trading Fund with the Department for Business, Energy and Industrial Strategy. But what does that mean it’s a Trading Fund?k it means that we’re funded by public money. So, that’s called the Public Weather Service.

But we also offer a more commercial venture. So, depending on what extensions you’ve got going on in your browser, there are actually adverts that do run on our website, and we do this to help recover some of the cost. So, the Public Weather Service has to recover some of that. And then lots of things are funded by the Public Weather Service, from observations, to public forecasting. But then there are more those commercial ventures such as the energy markets that have more paid products, and things like that as well. So, maybe not that many adverts, but definitely more usable.

Corey: Yeah, I disabled the ad blocker, and I’m reloading it and I’m not seeing any here. Maybe I’m just considered to be such a poor ad targeting prospect at this point that people have just given up in despair. Honestly, people giving up on me in despair is kind of my entire shtick.

Jake: We focus heavily on user-centered design, so I was fortunate in their previous team to work in our digital area, consumer digital, which looked after our web and mobile channels. And I can heartily say that there are a lot of changes, had a lot of heavy research into them. Not just internal, getting [unintelligible 00:06:09] and having a look at it, but what does this is actually mean for members of the? Public sending people out doing guerrilla public testing, standing outside Tescos—which is one of our large superstores here—and saying, “Hey, what do you think of this?” And then you’d get a variety of opinions, and then features would be adjusted, tweaked, and so on.

Corey: So, you folks have been a relatively early adopter, especially in an institutional context. And by institution, I mean, one of those things that feels like it is as permanent as the stones in a castle, on some level, something that’s lasted more than 20 years here in California, what a concept. And part of me wonders, were you one of the first UK government offices to use the cloud, and is that because you do weather and someone was very confused by what Cloud meant?

Jake: [laugh]. I think we were possibly one of the first; I couldn’t say if we were the first. Over in the UK, we’ve got a very capable network of government agencies doing some wonderful, and very cloud things. And the Government Digital Service was an initiative set up—uh, I can’t remember, and I—unfortunately I can’t remember the name of the report that caused its creation, but they had a big hand in doing design and cloud-first deployments. In the Met Office, we didn’t take a, “Ah, screw it. Let’s jump in,” we took a measured step into the cloud waters.

Like I said, we’ve been running supercomputers since the ’50s, and mainframes as well, and x86. I mean, we’ve been around for 100 years, so we constantly adapt, and engage, and iterate, and improve. But we don’t just jump in and take a risk because like you said, we are an institution; we have to provide services for the public. It’s not something that you can just ignore. These are services that protect life and property, both at home and abroad.

Corey: You have provided a case study historically to AWS, about your use cases of what you use, back in 2014. It was, oh, you’re a heavy user of EC2, and looking at the clock, and oh, it’s 2014. Surprise. But you’ve also focused on other services as well. I believe you personally provided a bit of a case study slash story of round your use of Pinpoint of all things, which is a wrapper around SES, their email service, in the hopes of making it a little bit more, I guess, understandable slash fully-featured for contacting people, but in my experience is a great sales device to drive business to its competitors.

What’s it been like working, I guess, both simultaneously with the tried and true, tested yadda, yadda, yadda, EC2 RDS style stuff, but then looking at what else you’re deep into Lambda, and DynamoDB, and SQS sort of stands between both worlds give it was the first service in beta, but it also is a very modern way of thinking about services. How do you contextualize all of that? Because AWS has product strategies, clearly, “Yes.” And they build anything for anyone is more or less what it seems. How do you think about the ecosystem of services that are available and apply it to problems that you’re working on?

Jake: So, in my personal opinion, I think the Met Office is one of a very small handfuls of companies around the world that could use every Amazon service that’s offered, even things like Ground Station. But on my first day in the office, I went and sat at my desk and was talking to my new colleagues, and I looked to the left and he said, “Oh, yeah, that’s a satellite dish collecting data from a satellite passing overhead.” So, we very much pick the best tool for the job. So, we have systems which do heavy number crunching, and very intense things, we’ll go for EC2.

We have systems that store data that needs relationships and all sorts of things. Fine, we’ll go RDS. In my space, we have over a billion observations a year coming through the system I lead on SurfaceNet. So, do we need RDS? No. What about if we use something like S3 and Glue and Athena to run queries against this?

We’re very fortunate that we can pick the best tool for the job, and we pride ourselves on getting the most out of our tools and getting the most value for money. Because like I said, we’re funded by the taxpayer; the taxpayer wants value for money, and we are taxpayers ourselves. We don’t want to see our money being wasted when we got a hundred size auto-scaling group, when we could do it with Lambda instead.

Corey: It’s fascinating talking about some of the forward-looking stuff, and oh, serverless and throw everything at Cloud and be all in on cloud. Cloud, cloud, cloud. Cloud is the future. But earlier this year, there was a press release where the Met Office and Microsoft are going to be joining forces to build the world’s, and I quote, “Most powerful weather and climate forecasting supercomputer.” The government—your government, to be clear—is investing over a billion pounds in the project.

It is slated to be online and running by the middle of next year, 2022, which for a government project as I contextualize them feels like it’s underwear-on-outside-the-pants superhero speed. But that, I guess, is what happens when you start looking at these public-private partnerships in some respects. How do you contextualize that? What is the story behind, oh, we’re—you’re clearly investing heavily in cloud, but you’re also building your own custom enormous supercomputer rather than just waiting for AWS to drop one at re:Invent. What is the decision-making process look like? What is the strategy behind it?

Jake: Oh. [laugh]. So—I’ll have to be careful here—supercomputing is something that we’ve been doing for a long time, since the ’50s, and we’ve grown with that. When the Met Office moved offices from Bracknell in 2002, 2003, we run two supercomputers for operational resilience, at that point [unintelligible 00:12:06] building in the new building; it was ready, and they were like, “Okay, let’s move a supercomputer.” So, it came hurtling down the motorway, plugged in, and congrats, we’ve now got two supercomputers running again. We’re very fortunate—

Corey: We had one. It got lonely. We wanted to make it a friend. Yeah, I get it.

Jake: Yeah. It’s long distance; it works. And the Met Office is actually very good at running projects. We’ve done many supercomputers over the years, and supercomputing our models, we run some very intense models, and we have more demands. We know we can do better.

We know there’s the observations in my group we collect, there’s the science that’s continually improving and iterating and getting better, and our limit isn’t poor optimizations or poorly written code. They’re scientists running some fantastic code; we have a team who go and optimize these models, and you know, in one release, they may knock down a model runtime by four minutes. And you think, okay, that’s four minutes, but for example, if that’s four minutes across 400 nodes, all of a sudden you’ve now got 400 nodes that have then got four minutes more of compute. That could be more research, that could be a different model run. You know, we’re very good at running these things, and we’re very fortunate with very technically capable to understand the difference between a workload that belongs on AWS, a workload that belongs on a supercomputer.

And you know, a supercomputer has many benefits, which the cloud providers… are getting into, you know, we have a high performance clusters on Amazon and Azure, or with, you know, InfiniBand networking. But sometimes you really can’t beat a hunking great big ton of metal and super water-cooling, sat in a data center somewhere, backed by—we’re very fortunate to have one hundred percent renewable energy for the supercomputer, which is—if you look at any of the power requirements for a supercomputer is phenomenal, so we’re throwing that credentials behind it for climate change as well. You can’t beat a supercomputer sometimes.

Corey: This episode is sponsored by our friends at Oracle HeatWave is a new high-performance accelerator for the Oracle MySQL Database Service. Although I insist on calling it “my squirrel.” While MySQL has long been the worlds most popular open source database, shifting from transacting to analytics required way too much overhead and, ya know, work. With HeatWave you can run your OLTP and OLAP, don’t ask me to ever say those acronyms again, workloads directly from your MySQL database and eliminate the time consuming data movement and integration work, while also performing 1100X faster than Amazon Aurora, and 2.5X faster than Amazon Redshift, at a third of the cost. My thanks again to Oracle Cloud for sponsoring this ridiculous nonsense.

Corey: I’m somewhat fortunate in the despite living in a world of web apps, these days, my business partner used to work at the Department of Energy at Oak Ridge National Lab, helping with the care and feeding of the supercomputer clusters that they had out there. And you’re absolutely right; that matches my understanding with the idea that there are certain workloads you’re not going to be able to beat just having this enormous purpose-built cluster sitting there ready to go. Or even if you can, certainly not economically. I have friends who are in the batch side of the world, the HPC side of the world over in the AWS organizations, and they keep—“Hey, look at this. This thing’s amazing.”

But so much of what they’re talking about seems to distill down to, “I have this one-off giant compute task that needs to get done.” Yes, you’re right. If I need to calculate the weather one time, then okay, I can make an argument for going with cloud but you’re doing this on what appears to be a pretty consistent basis. You’re not just assuming—as best I can tell that, “And starting next Wednesday, it will be sunny forever. The end.”

Jake: I’m sure many people would love it if we could do weather on-demand.

Corey: Oh, yes. [unintelligible 00:15:09] going to reserved instance weather. That would be great. Like, “All right. I’d like to schedule some rain, please.” It really seems like it’s one of those areas that is one of the most commonly accepted in science fiction without any real understanding of just what it would take to do something like that. Even understanding and predicting the weather is something that is beyond an awful lot of our current capabilities.

Jake: This is exactly it. So, the Met Office is world-renowned for its research capabilities and those really in-depth, very powerful models that we run. So, I mentioned earlier, something called MOGREPS, which is the Met Office’s ensemble-based models. And what do we mean by ensembles? You may see in the documentation it’s got 18 members.

What does that mean? It means that we actually run a simulation 18 times, and we tweak the starting parameters based on these real world inputs. And then you have a number of members that iterate through and supercomputer runs all of them. And we have deterministic models, which have one set of inputs. And you know, it’s not just, as you say, one time; these models must run.

There are a number of models we do, models on sea state as well, and they’ve all got to run, so we generally tend to run our supercomputers at top capacity. It’s not often you get to go on a supercomputer and there’ll be some space for your job to execute right this minute. And there’s all the setup as well, so it’s not just okay, the supercomputer is ready to go, but there’s all the things that go into it, like, those observations, whether it’s from the surface, whether it’s from satellite data passing overhead, we have our own lightning network, as well. We have many things, like a radar network that we own, and operate. We collaborate with the environment agency for rainfall. And all these things they feed into these models.

Okay, now we produce a model, and now it’s got to go out. So, it’s got to come off the supercomputer, it’s got to be processed, maybe the grid that we run the models on needs to be reprojected because different people feed maps in different ways. Then there’s got to be cut up because not every customer wants to know what the weather is everywhere. They’ve got a bit they care about. And of course, these models aren’t small; you know, they can be terabytes, so there’s also a case of customers might not want to download terabytes; that might cost them a lot. They might only be able to process gigabytes an hour.

But then there’s other products that we do processing on, so weather models, it might take 40 minutes to over an hour for a model to run. Okay, that’s great. You might have missed the first step. Okay, well, we can enrich it with other data that’s come in, things like nowcasting, where we do very short runs for the next six-hour forecast. There’s a whole number of things that run in the office. And we don’t have a choice; they run operationally 24/7, around the clock.

I mentioned to you before we started recording, we had an incident of ‘Beast from the East’ a number of years back. Some of your listeners may remember this; in the UK, we had a front come in from the east and the UK was blanketed with snow. It was a real severe event. We pretty much kept most of our services running. We worked really hard to make sure that they continued working.

And personally I say, perhaps when you go shopping for Black Friday, you might go to a retailer and it’s got a queue system up because, you know, it mimics that queue thing when you’re outside a store, like in Times Square, and it’s raining, be like oh, I might get a deal a minute. I think possibly in the Met Office, we have almost the inverse problem. If the weather’s benign, we’re still there. People rely on us to go, “Yeah, okay. I can go out and have fun.” When the weather’s bad, we don’t have a choice. We have to be there because everybody wants us to be there, but we need to be there. It’s not a case of this is an optional service.

Corey: People often forget that yeah, we are living in a world in which, especially with climate change doing what it’s doing, if you get this wrong, people can very easily die. That is not something to take lightly. It’s not just about can I go outside and play a pickup game of basketball today?

Jake: Exactly. So, you know, operationally, we have something called the National Severe Weather Warning Service, where we issue guidance and alerts across the UK, based on severe weather. And there’s a number of different weather types that we issued guidance for. And the severity of that goes from yellow to amber to red. And these are manually generated products, so there’s the chief meteorologist who’s on shift, and he approves these.

And these warnings don’t just go out to the members of the public. They go out to Cabinet Office, they go out to first responders, they go out to a number of people who are interested in the weather and have a responsibility. But the other side is that we don’t issue a weather warning willy-nilly. It’s a measured, calculated decision by our very capable operations team. And once that weather system has passed, the weather story has changed, we’ll review it. We go back and we say what could we have done differently?

Could the models have predicted this earlier? Could we have new data which would have picked up on this? Some of our next generation products that are in beta, would they have spotted this earlier? There’s a lot of service review that continually goes on because like I said, we are the best, and we need to stay the best. People rely on us.

Corey: So, here’s a question that probably betrays my own ignorance, and that’s okay, that’s what I’m here to do. When I was a kid, I distinctly remember—first, this is not the era wish the world was black and white; I’m a child of the ’80s, let’s be clear here, so this is not old-timey nonsense quite as much, but distinctly remember that it was a running gag how unreliable the weather report always was, and it was a bit hit or miss, like, “Well, the paper says it’s going to be sunny today, but we’re going to pack an umbrella because we know how this works.” It feels, and I could be way off base on this, but it really feels like weather forecasting has gotten significantly more accurate since I was a kid. Is that just nostalgia, and I remember my parents complaining about it, or has there been a qualitative improvement in the accuracy of weather forecasting?

Jake: I wish I could tell you all the scientific improvements that we’ve made, but there’s many groups of scientists in the office who I would more than happily shift that responsibility over to, but quite simply, yes. We have a lot of partners we work with around the world—the National Weather Service, DWD in Germany, Meteo France, just to name but a few; there are many—and we all collaborate with data. We all iterate. You know, the American Meteorological Society holds a conference every year, which we attend. And there have been absolutely leaping changes in forecast quality and accuracy over the years.

And that’s why we continually upgrade our supercomputers. Like I said, yeah, there’s research and stuff, but we’re pulling in all this science and Meteorology is generally very chaotic systems. We’re still discovering many things around how the climate works and how the weather systems work. And we’re going to use them to help improve quality of life, early warnings, actually, we can say, oh, in three days time, it’s going to be sunny at the beach. Be great if you could know that seven days in advance. It would be great if you knew that 14 days in advance.

I mean, we might not do that because at the moment, we might have an idea, but there’s also the case of understanding, you know, it’s a probability-based decision. And people say, “Oh, it’s not going to rain.” But actually, it’s a case of, well, we said there’s a 20% probability is going to rain. That doesn’t mean it’s not going to, but it’s saying, “Two times out of ten, at this time it’s going to rain.” But of course, if you go out 14 days, that’s a long lead time, and you know, you talk about chaos theory, and the butterfly moves and flaps its wings, and all of a sudden a [cake 00:22:50] changes color from green to pink or something like that, some other location in the world.

These are real systems that have real impacts, so we have to balance out the science of pure numbers, but what do people do with it? And what can people do with it, as well? So, that’s why we talk about having timely data as well. People say, “Well, you could run these simulations and all your products take longer to process them and generate them,” but for example, in SurfaceNet, we have five minutes to process an observation once it comes in. We could spend hours fine-tuning that observation to make it perfect, but it needs to be useful.

Corey: As you take a look throughout all of the things that AWS is doing—and sure, not all of these are going to necessarily apply directly to empowering the accuracy of weather forecasts, let’s be clear here—but you have expressed personal interest in for example, IoT, a bunch of the serverless nonsense we’re seeing out there. What excites you the most? What has you the most enthusiastic about what the future the cloud might hold? Because unlike almost everyone else I talk to in this space, you are not selling anything. You don’t have a position—that I’m aware of—that oh, yeah, I super want to see this particular thing win the industry because that means you get to buy a boat.

You work for the Met Office; you know that in some cases, oh, that boat is not going to have a great time in that part of the world anyway. I don’t need one. So, you’re a little bit more objective than most people. I have pushing a corporate story. What excites you? Where do you see the future of this industry going in ways that are neat?

Jake: Different parts of the office will tell you different things, you know. We worked with Google DeepMind on AI and machine learning. We work with many partners on AI and machine learning, we use it internally, as well. On a personal level, I like quality of life improvements and things that just make my life as both the developer fun and interesting. So, CDK was a big thing.

I was a CloudFormation wizard—still hate writing YAML—but the CDK came along and it was [unintelligible 00:24:52] people wouldn’t say, but that wasn’t, like, know when Lambda launched back in, what, 2013? 2014? No, but it made our lives easier. It meant that actually, we didn’t have to worry about, okay, how do we do templating with YAML? Do we have to run some pre-processes or something?

It meant that we could invest a little bit of time upfront on CDK and migrating everything over, and then that freed us up to actually doing things that we need for what we call the business or the organization, delivering value, you know? It’s great playing with tech but, you know, I need to deliver value. And I think, what was it, in the Google SRE book, they limit the things they do, toiling of manual tasks that don’t really contribute anything, they’re more like keeping the lights on. Let’s get rid of that. Let’s focus on delivering value.

It’s why Lambda is so great. I could patch an EC2, I can automate it, you know, you got AWS Systems Manager Patch Manager, or… whatever its name is, they can go and manage all those patches for you. Why when I can do it in a Lambda and I don’t need to worry about it?

Corey: So, one last question that I have for you is that you’re a tech lead. It’s easy for folks to fall into the trap of assuming, “Oh, you’re a government. It’s like an enterprise only bigger, slower, and way, way, way busier.” How many hundreds of thousands of engineers are working at the Met Office along with you?

Jake: So, you can have a look at our public report and you can see the number of staff we have. I think there’s about 1800 staff that work at the Met Office. And that includes our account manage, that includes our scientists, that includes HR and legal. And I’d say there’s probably less than 300 people who work in technology, as we call it, which is managing our IT estate, managing our Linux estate, managing our storage area networks because, funnily enough, managing petabytes of data is not an easy thing. You know, managing a supercomputer, a mainframe.

There really aren’t that many people here at the office, but we do so much great stuff. So, as a technical lead, I’m not just a leader of services, but I lead a team of people. I'm responsible for them, for empowering them, and helping them to develop their own careers and their own training. So, it’s me and a team of four that look after SurfaceNet. And it’s not just SurfaceNet; we’ve got other systems we look after that SurfaceNet produces data for. Sending messages around the world on the World Meteorological Organization’s global telecommunications system. What a mouthful. But you know, these messages go all around the world. And some people might say, “Well, I got a huge team for that.” Well, [unintelligible 00:27:27]. We have other teams that help us—I say, help us—in their own right, they transmit that data. But we’re really—I personally wouldn’t say we were huge, but boy, do we pack a punch.

Corey: Can I just say on a personal note, it’s so great to talk to someone who’s focusing on building out these environments and solving these problems for a higher purpose slash calling than—and I will get letters for this—than showing ads to people on the internet. I really want to thank you for taking time out of your day to speak with me. If people want to learn more about what you’re up to, how you do it, potentially consider maybe joining you if they are eligible to work at the Met Office, where can they find you?

Jake: Yeah, so you do have to be a resident in the UK, but www.metoffice.gov.uk is our home on the internet. You can find me on Twitter at @jakehendy, and I could absolutely chew Corey’s ear off for many more hours about many of the wonderful services that the Met Office provides. But I can tell he’s got something more interesting to do. So, uh [crosstalk 00:28:29]—

Corey: Oh, you’d be surprised. It’s loads of fun to—no, it’s always fun to talk to people who are just in different areas that I don’t get to work with very often. It turns out that most of my customers are not focused on telling you what the weather is going to do. And that’s fine; it takes all kinds. It’s just neat to have this conversation with a different area of the industry. Thank you so much for being so generous with your time. I appreciate it.

Jake: Thank you very much for inviting me on. I guess if we get some good feedback, I’ll have to come on and I will have to chew your ear off after all.

Corey: Don’t offer if you’re not serious.

Jake: Oh, I am.

Corey: Jake Hendy, Tech Lead at the Met Office. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with a comment yelling at one or both of us for having the temerity to rain on your parade.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Frank

Frank Chen is a maker. He develops products and leads software engineering teams with a background in behavior design, engineering leadership, systems reliability engineering, and resiliency research. At Slack, Frank focuses on making engineers' lives simpler, more pleasant, and more productive, in the Developer Productivity group. At Palantir, Frank has worked with customers in healthcare, finance, government, energy and consumer packaged goods to solve their hardest problems by transforming how they use data. At Amazon, Frank led a front-end team and infrastructure team to launch AWS WorkDocs, the first secure multi-platform service of its kind for enterprise customers. At Sandia National Labs, Frank researched resiliency and complexity analysis tooling with the Grid Resiliency group. He received a M.S. in Computer Science focused in Human-Computer Interaction from Stanford. Frank's thesis studied how the design / psychology of exergaming interventions might produce efficacious health outcomes. With the Stanford Prevention Research Center, Frank developed health interventions rooted in behavioral theory to create new behaviors through mobile phones. He prototyped early builds of Tiny Habits with BJ Fogg and worked in the Persuasive Technology Lab. He received a B.S. in Computer Science from UCLA. Frank researched networked systems and image processing with the Center for embedded Networked Systems. With the Rand Corporation, he built research systems to support group decision-making.

Links:

  • Slack: https://slack.com
  • “Infrastructure Observability for Changing the Spend Curve”: https://slack.engineering/infrastructure-observability-for-changing-the-spend-curve/
  • “Right Sizing Your Instances Is Nonsense”: https://www.lastweekinaws.com/blog/right-sizing-your-instances-is-nonsense/
  • Personal webpage: https://frankc.net
  • Twitter: @frankc

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: It seems like there is a new security breach every day. Are you confident that an old SSH key, or a shared admin account, isn’t going to come back and bite you? If not, check out Teleport. Teleport is the easiest, most secure way to access all of your infrastructure. The open source Teleport Access Plane consolidates everything you need for secure access to your Linux and Windows servers—and I assure you there is no third option there. Kubernetes clusters, databases, and internal applications like AWS Management Console, Yankins, GitLab, Grafana, Jupyter Notebooks, and more. Teleport’s unique approach is not only more secure, it also improves developer productivity. To learn more visit: goteleport.com. And not, that is not me telling you to go away, it is: goteleport.com.

Corey: This episode is sponsored by our friends at Oracle Cloud. Counting the pennies, but still dreaming of deploying apps instead of "Hello, World" demos? Allow me to introduce you to Oracle's Always Free tier. It provides over 20 free services and infrastructure, networking, databases, observability, management, and security. And—let me be clear here—it's actually free. There's no surprise billing until you intentionally and proactively upgrade your account. This means you can provision a virtual machine instance or spin up an autonomous database that manages itself all while gaining the networking load, balancing and storage resources that somehow never quite make it into most free tiers needed to support the application that you want to build. With Always Free, you can do things like run small scale applications or do proof-of-concept testing without spending a dime. You know that I always like to put asterisks next to the word free. This is actually free, no asterisk. Start now. Visit snark.cloud/oci-free that's snark.cloud/oci-free.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Several people are undoubtedly angrily typing, and part of the reason they can do that, and the fact that I know that is because we’re all using Slack. My guest today is Frank Chen, senior staff software engineer at Slack. So, I guess, sort of… [sales force 00:00:53]. Frank, thanks for joining me.

Frank: Hey, Corey, I have been a longtime listener and follower, and just really delighted to be here.

Corey: It’s one of the weird things about doing a podcast is that for better or worse, people don’t respond to it in the same way that they do writing a newsletter, for example, because you receive an email, and, “Oh, well, I know how to write an email. I can hit reply and send an email back and give that jackwagon a piece of my mind,” and people often do. But with podcasts, I feel like it’s much more closely attuned to the idea of an AM radio talk show. And who calls into a radio talk show? Lunatics, and most people don’t self-describe as lunatics, so they don’t want to do that.

But then when I catch up with people one-on-one or at events in person, I find out that a lot more people listen to this show than I thought they did. Because I don’t trust podcast statistics because lies, damn lies, and analytics are sort of how I view this world. So, you’ve worked at a bunch of different companies. You’re at Slack now, which, of course, upsets some people because, “Slack is ruining the way that people come and talk to me in the office.” Or it’s making it easier for employees to collaborate internally in ways their employers wish they wouldn’t. But that’s neither here nor there.

Before this, you were at Palantir, and before this, you’re at Amazon, working on Amazon WorkDocs of all things, which is supposedly rumored to have at least one customer somewhere, but I’ve never seen them. Before that you were at Sandia National Labs, and you’ve gotten a master’s in computer science from Stanford. You’ve done a lot of things and everything you’ve done, on some level, seems like the recurring theme is someone on Twitter will be unhappy at you for a career choice you’ve made. But what is the common thread—in seriousness—between the different places that you’ve been?

Frank: One thing that’s been a driver for where I work is finding amazing people to work with and building something that I believe is valuable and fun to keep doing. The thing that brought me to Slack is I became my own Slack admin, [laugh] when I met a girl and we moved in together into a small apartment in Brooklyn. And she had a cat that, you know, is a sweetheart, but also just doesn’t know how to be social. Yes, you covered that with ‘cat.’ Part of moving it together, I became my own Slack admin and discovered well, we can build a series of home automations to better train and inform our little command center for when the cat lies about being fed, or not fed, clipping his nails, and discovering and tracking bad behaviors. In a lot of ways this was like the human side of a lot of the data work that I had been doing at my previous role. And it was like a fun way to use the same frameworks that I use at work to better train and be a cat caretaker.

Corey: Now, at some point, you know that some product manager at Amazon is listening to this and immediately sketching notes because their product strategy is, “Yes,” and this is going to be productized and shipping in two years as Amazon Prime Meow. But until then we’ll enjoy the originality of having a Slack bot more or less control the home automation slash making your house seem haunted for anyone who didn’t write the code themselves. There's an idea of solving real world problems that I definitely understand. I mean, and again, it might not even be a fair question entirely. Just because I am… for better or worse, staggering through my world, and trying—and failing most days—to tell a narrative that, “Oh, why did I start my tech career at a university, and then spend time in ad tech, and then spend time in consulting, and then FinTech, and the rest?” And the answer is, “Oh, I get fired an awful lot, and that sucked.”

So, instead of going down that particular rabbit hole of a mess, I went in other directions. I started finding things that would pay me and pay me more money because I was in debt at the time. But that was the narrative thread that was the, “I have rent to pay and they have computers that aren’t behaving properly.” And that’s what dictated the shape of my career for a long time. It’s only in retrospect that I started to identify some of the things that aligns with it. But it’s easy to look at it with the shine of hindsight and not realize that no, no, that’s sort of retconning what happened in the past.

Frank: Yeah, I have a mentor and my former adviser had this way of describing, building out the jankiest prototype you can to prove out an idea. And this manifested in his class in building out paper prototypes, or really, really janky ideas for what helping people through technology might look like. And I feel like it a lot of ways, even when those prototypes fail, like, in a career or some half baked tech prototype I put together, it might succeed and great, we could keep building upon that, but when it fails, you actually discover, “Oh, this is one way that I didn’t succeed.” And even in doing so, you discover things about yourself, your way of building, and maybe a little bit about your infrastructure, or whatever it is that you build on a day-to-day basis. And wrapping that back to the original question, it’s like, well, we think we’re human beings, right, we’re static, but in a lot of ways we’re human becomings. We think we know what the future might look like with our careers, what we’re building on a day-to-day basis, and what we’re building a year from now, but oftentimes, things change if we discover things about ourselves, the people we work with, and ultimately, the things that we put out into the world.

Corey: Obviously, I’ve been aware of who Slack is, for a long time; I’ve been a paying customer for years because it basically is IRC with reaction gifs, and not having to teach someone how to sign into IRC when they work in accounting. So, the user experience alone solved the problem.

Frank: And you’ve actually worked with us in the past before. [laugh]. Slack, it’s the Searchable Log for all Content and Knowledge; I think that backronym, that’s how it works. And I was delighted when I had mentioned your jokes and you’re trolling [a folk 00:07:00] on Twitter and on your podcast to my former engineering manager, Chris Merrill, who was like, oh, you should search the Slack. Corey actually worked with us and he put together a lot of cool tooling and ideas for us to think about.

Corey: Careful. If we talk too much, or what I did when I was at Slack years ago, someone’s going to start looking into some of the old commits and whatnot and start demanding an apology, and we don’t want that. It’s, “Wow, you’re right. You are a terrible engineer.” “Told you.” There’s a reason I don’t do that anymore.

Frank: I think that’s all of us. [laugh]. An early career mentor of mine, he was like, “Hey, Frank, listen. You think you’re building perfect software at any point in time? No, you’re building future tech debt.” And yeah, we should put much more emphasis on interfaces and ideas we’re putting out because the implementation is going to change over time, and likely your current implementation is shit. And that is, okay.

Corey: That’s the beautiful part about this is that things grow and things evolve. And it’s interesting working with companies, and as a consultant, I tend to build my projects in such a way that I start on day one and people know that I’m leaving with usually a very short window because I don’t want to build a forever job for myself; I don’t want to show up and start charging by the hour or by the day, if I can possibly avoid it. Because then it turns into eternal projects that never end because I’m billing and nothing’s ever done. No, no, I like charging fixed fee and then getting out at a predetermined outcome, but then you get to hear about what happens with companies as they move on.

This combines with the fact that I have a persistent alert for my name, usually because I’m looking for various ineffective character assassination from enterprise marketing types because you know, I dish it out, I should certainly be able to take it. But I found a blog post on the Slack engineering blog that mentioned my name, and it’s, “Aw, crap. Are they coming after me for a refund?” No, it was not. It was you writing a fairly sizable post. Tell me more about that.

Frank: Yeah, I’m part of an organization called Developer Productivity. And our goal is to help folk at Slack deliver services to their customers, where we build, test, and release high quality software. And a lot of our time is spent thinking about internal tooling and making infrastructure bets. As engineers, right, it’s like, we have this idea for what the world looks like, we have this idea for what our infrastructure looks like, but what we discover using a set of techniques around observability of just asking questions—advanced questions, basic questions, and hell, even dumb questions—we discover hey, the things that we think our computers are doing aren’t actually doing what they say they’re doing. And the question is like, great. Now, what? How can we ask better questions? How can we better tune, change, and equip engineers with tooling so that they can do better work to make Slack customers have simple, pleasant, and productive experiences?

Corey: And I have to say that there’s a lot that Slack does that is incredibly helpful. I don’t know that I’m necessarily completely bought into the idea that all work should happen in Slack. It’s, well, on some level, I—like people like to debate the ‘should people work from home? Should people all work in an office?’ Discussion.

And, on some level, it seems if you look at people who are constantly fighting that debate online, it’s, “Do you ever do work at all?” on some level. But I’m not here to besmirch others; I’m here to talk about, on some level, what you alluded to in your blog post. But I want to start with a disclaimer that Slack as far as companies go is not small, and if you take a look around, most companies are using Slack whether they know it or not. The list of side-channel Slack groups people have tend to extend massively.

I look and I pare it down every once in a while, whenever I cross 40 signed-in Slacks on my desktop. It is where people talk for a wide variety of different reasons, and they all do different things. But if you’re sitting here listening to this and you have a $2,000 a month AWS bill, this is not for you. You will spend orders of magnitude more money trying to optimize a small cost. Once you’re at significant points of scale, and you have scaled out to the point where you begin to have some ability to predict over months or years, that’s what a lot of this stuff starts to weigh in.

So, talk to me a bit about how you wound up—and let me quote directly from the article, which is titled, “Infrastructure Observability for Changing the Spend Curve,” and I will, of course, throw a link to this in the [show notes 00:11:38]. But you talk in this about knocking, I believe it was orders of magnitude off of various cost areas within your bill.

Frank: Yeah. The article itself describes three big-ish projects, where we are able to change the curve of the number of tests that we run, and a change in how much it costs to run any single test.

Corey: When you say test, are you talking CI/CD infrastructure test or code test, to make sure it goes out, or are you talking something higher up the stack, as far as, “Huh, let’s see how some users respond when, I don’t know, we send four notifications on every message instead of the usual one,” to give a ridiculous example?

Frank: Yeah, this is in the CI/CD pipelines. And one of these projects was around borrowing some concepts from data engineering: oversubscription and planning your capacity to have access capacity at peak, where at peak, your engineers might have a 5% degradation in performance, while still maintaining high resiliency and reliability of your tests in order to oversubscribe, either CPU or memory and keep throughput on the overall system stable and consistent and fast enough. I think, with spend in developer productivity, I think, both, like, the metrics you’re trying to move and why you’re optimizing for it at any given time are, like, this, like, calculus. Or it’s like, more art than science in that there’s no one right answer, right? It’s like, oh, yeah—very naively—like, yeah, let’s throw the biggest machines most expensive machines we can at any given problem. But that doesn’t solve the crux of your problem. It’s like, “Hey, what are the things in your system doing?” And what is the right guess to capitalize around how much to spend on your CI/CD [unintelligible 00:13:39] is oftentimes not precise, nor is this blog article meant to be prescriptive.

Corey: Yeah, it depends entirely on what you’re doing and how because it’s, on some level, well, we can save a whole bunch of money if we slow all of our CI/CD runs down by 20 minutes. Yeah, but then you have a bunch of engineers sitting idle and I promise you, that costs a hell of a lot more than your cloud bill is going to be. The payroll is almost always a larger expense than your infrastructure costs, and if it’s not, you should seriously consider firing at least part of your data science team, but you didn’t hear it from me.

Frank: Yeah. And part of the exploration on profiling and performance and resiliency was, like, around interrogating what the boundaries and what the constraints were for our CI/CD pipelines. Because Slack has grown in engineering and in the number of tests we were running on a month-to-month basis; for a while from 2017 to mid 2020, we were growing about 10% month-over-month in test suite execution numbers. Which means on a given year, we doubled almost two times, which is quite a bit of strain on internal resources and a lot of dependent services where—and internal systems, we oftentimes have more complexity and less understood changes in what dependencies your infrastructure might be using, what business logic your internal services are using to communicate with one another than you do your production.

And so, by, like, performing a series of curiosity-driven development, we’re able to both answer, at that point in time, what our customers internally were doing, and start to put together ideas for eliminating some bottlenecks, and hell, even adding bottlenecks with circuit breakers where you keep the overall throughput of your system stable, while deferring or canceling work that otherwise might have overloaded dependencies.

Corey: There’s a lot to be said for understanding what the optimization opportunities are, in an environment and understanding what it is you’re attempting to achieve. Having those test for something like Slack makes an awful lot of sense because let’s be very clear here, when you’re building an application that acts as something people use to do expense reports—to cite one of my previous job examples—it turns out you can be down for a week and a majority of your customers will never know or care. With Slack, it doesn’t work that way. Everyone more or less has a continuous monitor that they’re typing into for a good portion of the day—angrily or otherwise—and as soon as it misses anything, people know. And if there’s one thing that I love, on some level, seeing change when I know that Slack is having a blip, even if I’m not using Slack that day for anything in particular, because Twitter explodes about it. “Slack is down. I’m now going to tweet some stuff to my colleagues.” All right. You do you, I suppose.

And credit where due, Slack doesn’t go down nearly as often as it used to because as you tend to figure out how these things work, operational maturity increases through a bunch of tests. Fixing things like durability, reliability, uptime, et cetera, should always, to some extent, take precedence priority-wise over let’s save some money. Because yeah, you could turn everything off and save all the money, but then you don’t have a business anymore. It’s focused on where to cut, where to optimize in the right way, and ideally as you go, find some of the areas in which, oh, I’m paying AWS a tax for just going about my business. And I could have flipped a switch at any point and saved—“How much money? Oh, my God, that’s more than I’ll make in my lifetime.”

Frank: Yeah, and one thing I talk about a little bit is distributed tracing as one of the drivers for helping us understand what’s happening inside of our systems. Where it helps you figure out and it’s like this… [best word 00:17:24] to describe how you ask questions of deployed code? And there a lot of ways it’s helped us understand existing bottlenecks and identify opportunities for performance or resiliency gains because your past janky Band-Aids become more and more obvious when you can interrogate and ask questions around what is it performing like it used to? Or what has changed recently?

Corey: This episode is sponsored in part by something new. Cloud Academy is a training platform built on two primary goals. Having the highest quality content in tech and cloud skills, and building a good community the is rich and full of IT and engineering professionals. You wouldn’t think those things go together, but sometimes they do. Its both useful for individuals and large enterprises, but here's what makes it new. I don’t use that term lightly. Cloud Academy invites you to showcase just how good your AWS skills are. For the next four weeks you’ll have a chance to prove yourself. Compete in four unique lab challenges, where they’ll be awarding more than $2000 in cash and prizes. I’m not kidding, first place is a thousand bucks. Pre-register for the first challenge now, one that I picked out myself on Amazon SNS image resizing, by visiting cloudacademy.com/corey. C-O-R-E-Y. That’s cloudacademy.com/corey. We’re gonna have some fun with this one!

Corey: It’s also worth pointing out that as systems grow organically, that it is almost impossible for any one person to have it all in their head anymore. I saw one of the most overly complicated architecture flow trees that I think I’ve seen in recent memory, and it was on the Slack engineering blog about how something was architected, but it wasn’t the Slack app itself; it was simply the [decision tree for ‘Should we send a notification?’ 00:18:17] and it is more complicated than almost anything I’ve written, except maybe my newsletter content publication pipeline. It is massive. And I’ll throw a link to that in the [show notes 00:18:31] as well, just because it is well worth people taking a look at.

But there is so much complexity at scale for doing the right thing, and it’s necessary because if I’m talking to you on Slack right now and getting notifications every time you reply on my phone, it’s not going to take too long before I turn off notifications everywhere, and then I don’t notice that Slack is there, and it just becomes useless and I use something else. Ideally, something better—which is hard to come by—moderately worse, like, email or completely worse, like, Microsoft Teams.

Frank: I tell all my close collaborators about this. I typically set myself away on Slack because I like to make time for deep, focused work. And that’s very hard with a constant stream of notifications. How people use Slack and how people notify others on Slack is, like, not incumbent on the software itself, but it’s a reflection of the work culture that you’re in. The expectation for an email-driven culture is, like, oh, yeah, you should be reading your email all the time and be able to respond within 30 minutes. Peace, I have friends that are lawyers, [laugh] and that is the expectation at all times of day.

Corey: I married one of those. Oh, yeah, people get very salty. And she works with a global team spread everywhere, to the point where she wakes up and there’s just a whole flurry of angry people that have tried to reach her in the middle of the night. Like, “Why were you sleeping at 2 a.m.? It’s daytime here.” And yeah, time zones. Not everyone understands how they work, from my estimation.

Frank: [laugh]. That’s funny. My sweetheart is a former attorney. On our first international date, we spent an entire day-and-a-half hopping between WiFi spots in Prague so that she could answer a five minute question from a partner about standard deviations.

Corey: So, one thing that you link to that really is what drew my notice to this—because, again, if you talk about AWS cost optimization, I’m probably going to stumble over it, but if you mention my name, that’s sort of a nice accelerator—and you linked to my article called Why “Right Sizing Your Instances Is Nonsense.” And that is a little overblown, to some extent, but so many folks talk about it in the cost optimization space because you can get a bunch of metrics and do these things programmatically, and somewhat without observability into what’s going on because, “Well, I can see how busy the computers are and if it’s not busy, we could use smaller computers. Problem solved,” versus, the things that require a fair bit of insight into what is that thing doing exactly because it leads you into places of oh, turn off that idle fleet that’s not doing anything is all labeled ‘backup,’ where you’re going to have three seconds of notice before it gets all the traffic.

There’s an idea of sometimes things are the way they are for a reason. And it’s also not easy for a lot of things—think databases—to seamlessly just restart the thing and have it scale back up and run on a different instance class. That takes weeks of planning and it’s hard. So, I find that people tend to reach for it where it doesn’t often make sense. At your level of scale and operational maturity, of course, you should optimize what instance classes things are using and what sizes they are, especially since that stuff changes over time as far as what AWS has made available. But it’s not the sort of thing that I suggest as being the first easy thing to go for. It’s just what people think is easy because it requires no judgment and computers can do it. At least that’s their opinion.

Frank: I feel like you probably have a lot more experience than me, and talked about war stories, but I recall working with customers where they want to lift-and-shift on-prem hardware to VMs on-prem. I’m like, “It’s not going to be as simple as you’re making it out to be.” Whereas, like, the trend today is probably oh, yeah, we’re going to shift on-prem VMs to AWS, or hell, like, let’s go two levels deeper and just run everything on Kubernetes. Similar workloads, right? It’s not going to be a huge challenge. Or [laugh] everything serverless.

Corey: Spare me from that entire school of thought, my God.

Frank: [laugh].

Corey: Yeah, but it’s fun, too, because this came out a month ago, and you’re talking about using—an example you gave was a c5.9xlarge instance. Great. Well, the c6i is out now as well, so are people going to look at that someday and think, “Oh, wow. That’s incredibly quaint.”

It’s, you wrote this a month ago, and it’s already out of date, as far as what a lot of the modern story instances are. From my perspective, one of the best things that AWS has done in this space has been to get away from the reserved instance story and over into savings plans, where it’s, “I know, I’m going to run some compute—maybe it’s Fargate, maybe it’s EC2; let’s be serious, it’s definitely going to be EC2—but I don’t want to tie myself to specific instance types for the next three years.” Great, well, I’m just going to commit to spending some money on AWS for the next three years because if I decide today to move off of it, it’s going to take me at least that long to get everything out. So okay, then that becomes something a lot more palatable for an awful lot of folks.

Frank: One thing you brought up in the article I linked to is instance types. You think upgrading to the newest instance type will solve all your challenges, but oftentimes it’s not obvious that it won’t all the time, and in fact, you might even see degraded resiliency and degraded performance because different packages that your software relies upon might not be optimized for the given kernel or CPU type that you’re running against. And ultimately, you go back to just asking really basic questions and performing some end-to-end benchmarking so that you can at least get a sense for what your customers are doing today, and maybe make a guess for what they’re going to do tomorrow.

Corey: I have to ask because I’m always interested in what it is that gives rise to blog posts like this—which, that’s easy; it’s someone had to do a project on these things, and while we learn things that would probably apply to other folks—like, you’re solving what is effectively a global problem locally when you go down this path. It’s part of the reason I have a consulting business is things I learned at one company apply almost identically to another company, even though that they’re in completely separate industries and parts of the world because AWS billing is, for better or worse, a bounded problem space despite their best efforts to, you know, use quantum computers to fix that. What was it that gave rise to looking at the CI/CD system from an optimization point of view?

Frank: So internally, I initially started writing a white paper about, hey, here’s a simple question that we can answer, you know, without too much effort. Let’s transition all of our C3 instances to C5 instances, and that could have been the one and done. But by thinking about it a little more and kind of drawing out, while we can actually borrow a model for oversubscription from another field, we could potentially decrease our spend by quite a bit. That eventually [laugh] evolved into a 70 page white paper—no joke—that my former engineering manager said, “Frank, no one’s going to [BLEEP] read this.” [laugh].

Corey: Always. Always, always. Like, here’s a whole bunch of academically research and the rest. It’s like, “Great. Which of these two buttons do I press?” is really the question people are getting at. And while it’s great to have the research and the academic stuff, it’s also a, “Great we’re trying to achieve an outcome which, what is the choice?” But it’s nice to know that people are doing actual research on the back end, instead, “Eh, my gut tells me to take the path on the left because why not? Left is better; right’s tricky friend.”

Frank: Yeah. And it was like, “Oh, yeah. I accidentally wrote a really long thing because there was, like, a lot of variables to test.” I think we had spun up 16-plus auto-scaling groups. And ran something like the cross-section of a couple of representative test suites against them, as well as configurations for a number of executors per instance.

And about a year ago, I translated that into a ten page blog article that when I read through, I really didn’t enjoy. [laugh]. And that template blog article is ultimately, like, about a page in the article you’re reading today. And the actual kick in the butt to get this out the door was about four months ago. I spoke at o11ycon rescources which you’re a part of.

And it was a vendor conference by Honeycomb, and it was just so fun to share some of the things we’ve been doing with distributed tracing, and how we were able to solve internal problems using a relatively simple idea of asking questions about what was running. And the entire team there was wonderful in coaching and just helping me think through what questions people might have of this work. And that was, again, former academic. The last time I spoke at a conference was about a decade earlier, and it was just so fun to be part of this community of people trying to all solve the same set of problems, just in their own unique ways.

Corey: One of the things I loved about working with Honeycomb was the fact that whenever I asked them a question, they have instrumented their own stuff, so they could tell me extremely quickly what something was doing, how it was doing it, and what the overall impact on this was. It’s very rare to find a client that is anywhere near that level of awareness into what’s going on in their infrastructure.

Frank: Yeah, and that blog article, right, it’s like, here’s our current perspective, and here’s, like, the current set of projects we’re able to make to get to this result. And we think we know what we want to do, but if you were to ask that same question, “What are we doing for our spend a year from now?” the answer might be very different. Probably similar in some ways, but probably different.

Corey: Well, there are some principles that we’ll never get away from. It’s, “Is no one using the thing? Turn that shit off.” That’s one of those tried and true things. “Oh, it’s the third copy of that multiple petabyte of data thing? Maybe delete it or stuff in a deep archive.” It’s maybe move data less between various places. Maybe log things fewer times, given that you’re paying 50 cents per gigabyte ingest, in some cases. Et cetera, et cetera, et cetera. There’s a lot to consider as far as the general principles go, but the specifics, well, that’s where it gets into the weeds. And at your scale, yeah, having people focus on this internally with the context and nuance to it is absolutely worth doing. Having a small team devoted to this at large companies will pay for itself, I promise. Now, I go in and advise in these scenarios, but past a certain point, this can’t just be one person’s part-time gig anymore.

Frank: I’m kind of curious about that. How do you think about working with a company and then deprecating yourself, and allowing your tools and, like, the frameworks you put into place to continue, like, thrive?

Corey: We’re advisory only. We make no changes to production.

Frank: Or I don’t know if that’s the right word, deprecate. I think… that’s my own word. [laugh].

Corey: No, no, it’s fair. It’s a—what we do is we go in and we are advisory. It’s less of a cost engagement, more of an architecture engagement because in cloud, cost and architecture are the same thing. We look at what’s going on, we look at the constraints of why we’ve been brought in, and we identify things that companies can do and the associated cost savings associated with that, and let them make their own decision. Because it’s, if I come in and say, “Hey, you could save a bunch of money by migrating this whole subsystem to serverless.”

Great, I sound like a lunatic evangelist because yeah, 18 months of work during which time the team doing that is not advancing the state of the business any further so it’s never going to happen. So, why even suggest it? Just look at things that are within the bounds of possibility. Counterpoint: when a client says, “A full re-architecture is on the table,” well, okay, that changes the nature of what we’re suggesting. But we’re trying to get away from what a lot of tooling does, which is, “Great. Here’s 700 things you can adjust and you’ll do none of them.” We come back with a, “Here’s three or four things you can do that’ll blow 20% off the bill. Then let’s see where you stand.” The other half of it, of course, is large scale enterprise contract negotiation, that’s a bit of a horse of a different color. I want to thank you so much for taking the time to speak with me today. I really do appreciate it. If folks want to hear more about what you’re up to, and how you think about these things. Where can they find you?

Frank: You can find me at frankc.net. Or at me at @FrankC on Twitter.

Corey: Oh, inviting people to yell at you at Twitter. That’s never a great plan. Yeash. Good luck. Thanks again. We’ve absolutely got to talk more about this in-depth because I think this is one of those areas that you have the folks above a certain point of scale, talk about these things semi-constantly and live in the space, whereas folks who are in relatively small-scale environments are listening to this and thinking that they’ve got to do this.

And no. No, you do not want to spend millions of dollars of engineering effort to optimize a bill that’s 80 grand a year, I promise. It’s focus on the thing that’s right for your business. At a certain point of scale, this becomes that. But thank you so much for being so generous with your time. I appreciate it.

Frank: Thank you so much, Corey.

Corey: Frank Chen, senior staff software engineer at Slack. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an angry comment that seems to completely miss the fact that Microsoft Teams is free because it sucks.

Frank: [laugh].

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Robert

R2 advocates for Liquibase customers and provides technical architecture leadership. Prior to co-founding Datical (now Liquibase), Robert was a Director at the Austin Technology Incubator. Robert co-founded Phurnace Software in 2005. He invented and created the flagship product, Phurnace Deliver, which provides middleware infrastructure management to multiple Fortune 500 companies.

Links:

  • Liquibase: https://www.liquibase.com
  • Liquibase Community: https://www.liquibase.org
  • Liquibase AWS Marketplace: https://aws.amazon.com/marketplace/seller-profile?id=7e70900d-dcb2-4ef6-adab-f64590f4a967
  • Github: https://github.com/liquibase
  • Twitter: https://twitter.com/liquibase

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: It seems like there is a new security breach every day. Are you confident that an old SSH key, or a shared admin account, isn’t going to come back and bite you? If not, check out Teleport. Teleport is the easiest, most secure way to access all of your infrastructure. The open source Teleport Access Plane consolidates everything you need for secure access to your Linux and Windows servers—and I assure you there is no third option there. Kubernetes clusters, databases, and internal applications like AWS Management Console, Yankins, GitLab, Grafana, Jupyter Notebooks, and more. Teleport’s unique approach is not only more secure, it also improves developer productivity. To learn more visit: goteleport.com. And not, that is not me telling you to go away, it is: goteleport.com.

Corey: You know how Git works right?

Announcer: Sorta, kinda, not really. Please ask someone else.

Corey: That's all of us. Git is how we build things, and Netlify is one of the best ways I’ve found to build those things quickly for the web. Netlify’s Git-based workflows mean you don’t have to play slap-and-tickle with integrating arcane nonsense and web hooks, which are themselves about as well understood as Git. Give them a try and see what folks ranging from my fake Twitter for Pets startup, to global Fortune 2000 companies are raving about. If you end up talking to them—because you don’t have to; they get why self-service is important—but if you do, be sure to tell them that I sent you and watch all of the blood drain from their faces instantly. You can find them in the AWS marketplace or at www.netlify.com. N-E-T-L-I-F-Y dot com.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. This is a promoted episode. What does that mean in practice? Well, it means the company who provides the guest has paid to turn this into a discussion that’s much more aligned with the company than it is the individual.

Sometimes it works, Sometimes it doesn’t, but the key part of that story is I get paid. Why am I bringing this up? Because today’s guest is someone I met in person at Monktoberfest, which is the RedMonk conference in Portland, Maine, one of the only reasons to go to Maine, speaking as someone who grew up there. And I spoke there, I met my guest today, and eventually it turned into this, proving that I am the envy of developer advocates everywhere because now I can directly tie me attending one conference to making a fixed sum of money, and right now they’re all screaming and tearing off their headphones and closing this episode. But for those of you who are sticking around, thank you. My guest today is the CTO and co-founder of Liquibase. Please welcome Robert Reeves. Robert, thank you for joining me, and suffering the slings and arrows I’m about to hurled directly into your arse, as a warning shot.

Robert: [laugh]. Man. Thanks for having me. Corey, I’ve been looking forward to this for a while. I love hanging out with you.

Corey: One of the things I love about the Monktoberfest conference, and frankly, anything that RedMonk gets up to is, forget what’s on stage, which is uniformly excellent; forget the people at RedMonk who are wonderful and I aspire to do more work with them in different ways; they’re great, but the people that they attract are invariably interesting, they are invariably incredibly diverse in terms of not just demographics, but interests and proclivities. It’s just a wonderful group of people, and every time I get the opportunity to spend time with those folks I do, and I’ve never once regretted it because I get to meet people like you. Snark and cynicism about sponsoring this nonsense aside—for which I do thank you—you’ve been a fascinating person to talk to you because you’re better at a lot of the database-facing things than I am, so I shortcut to instead of forming my own opinions, I just skate off of yours in some cases. You’re going to get letters now.

Robert: Well, look, it’s an occupational hazard, right? Releasing software, it’s hard so you have to learn these platforms, and part of it includes the database. But I tell you, you’re spot on about Monktoberfest. I left that conference so motivated. Really opened my eyes, certainly injecting empathy into what I do on a day-to-day basis, but it spurred me to action.

And there’s a lot of programs that we’ve started at Liquibase that the germination for that seed came from Monktoberfest. And certainly, you know, we were bummed out that it’s been canceled two years in a row, but we can’t wait to get back and sponsor it. No end of love and affection for that team. They’re also really smart and right about a hundred percent of the time.

Corey: That’s the most amazing part is that they have opinions that generally tend to mirror my own—which, you know—

Robert: [laugh].

Corey: —confirmation bias is awesome, but they almost never get it wrong. And that is one of the impressive things is when I do it, I’m shooting from the hip and I already have an apology half-written and ready to go, whereas when dealing with them, they do research on this and they don’t have the ‘I’m a loud, abrasive shitpostter on Twitter’ defense to fall back on to defend opinions. And if they do, I’ve never seen them do it. They’re right, and the fact that I am as aligned with them as I am, you’d think that one of us was cribbing from the other. I assure you that’s not the case.

But every time Steve O’Grady or Rachel Stephens, or Kelly—I forget her last name; my apologies is all Twitter, but she studied medieval history, I remember that—or James Governor writes something, I’m uniformly looking at this and I feel a sense of dismay, been, “Dammit. I should have written this. It’s so well written and it makes such a salient point.” I really envy their ability to be so consistently on point.

Robert: Well, they’re the only analysts we pay money to. So, we vote with our dollars with that one. [laugh].

Corey: Yeah. I’m only an analyst when people have analyst budget. Other than that, I’m whatever the hell you describe me. So, let’s talk about that thing you’re here to show. You know, that little side project thing you found and are the CTO of.

I wasn’t super familiar with what Liquibase does until I looked into it and then had this—I got to say, it really pissed me off because I’m looking at it, and it’s how did I not know that this existed back when the exact problems that you solve are the things I was careening headlong into? I was actively annoyed. You’re also an open-source project, which means that you’re effectively making all of your money by giving things away and hoping for gratitude to come back on you in the fullness of time, right?

Robert: Well, yeah. There’s two things there. They’re open-source component, but also, where was this when I was struggling with this problem? So, for the folks that don’t know, what Liquibase does is automate database schema change. So, if you need to update a database—I don’t care what it is—as part of your application deployment, we can help.

Instead of writing a ticket or manually executing a SQL script, or generating a bunch of docs in a NoSQL database, you can have Liquibase help you out with that. And so I was at a conference years ago, at the booth, doing my booth thing, and a managing director of a very large bank came to me, like, “Hey, what do you do?” And saw what we did and got angry, started yelling at me. “Where were you three years ago when I was struggling with this problem?” Like, spitting mad. [laugh]. And I was like, “Dude, we just started”—this was a while ago—it was like, “We just started the company two years ago. We got here as soon as we could.”

But I struggled with this problem when I was a release manager. And so I’ve been doing this for years and years and years—I don’t even want to talk about how long—getting bits from dev to test to production, and the database was always, always, always the bottleneck, whether it was things didn’t run the same in test as they did, eventually in production, environments weren’t in sync. It’s just really hard. And we’ve automated so much stuff, we’ve automated application deployment, lowercase a compiled bits; we’re building things with containers, so everything’s in that container. It’s not a J2EE app anymore—yay—but we haven’t done a damn thing for the database.

And what this means is that we have a whole part of our industry, all of our database professionals, that are frankly struggling. I always say we don’t sell software Liquibase. We sell piano recitals, date nights, happy hours, all the stuff you want to do but you can’t because you’re stuck dealing with the database. And that’s what we do at Liquibase.

Corey: Well, you’re talking about database people. That’s not how I even do it. I would never call myself that, for very good reason because you know, Route 53 remains the only database I use. But the problem I always had was that, “Great. I’m doing a deployment. Oh, I’m going to put out some changes to some web servers. Okay, what’s my rollback?” “Well, we have this other commit we can use.” “Oh, we’re going to be making a database schema change. What’s your rollback strategy,” “Oh, I’ve updated my resume and made sure that any personal files I had on my work laptop been backed up somewhere else when I immediately leave the company when we can’t roll back.” Because there’s not really going to be a company anymore at that point.

It’s one of those everyone sort of holds their breath and winces when it comes to anything that resembles a schema change—or an ALTER TABLE as we used to call it—because that is the mistakes will show territory and you can hope and plan for things in pre-prod environments, but it’s always scary. It’s always terrifying because production is not like other things. That’s why I always call my staging environment ‘theory’ because things work in theory but not in production. So, it’s how do you avoid the mess of winding up just creating disasters when you’re dealing with the reality of your production environments? So, let’s back up here. How do you do it? Because it sounds like something people would love to sell me but doesn’t exist.

Robert: [laugh]. Well, it’s real simple. We have a file, we call it the change log. And this is a ledger. So, databases need to be evolved. You can’t drop everything and recreate it from scratch, so you have to apply changes sequentially.

And so what Liquibase will do is it connects to the database, and it says, “Hey, what version are you?” It looks at the change log, and we’ll see, ehh, “There’s ten change sets”—that’s what components of a change log, we call them change sets—“There’s ten change sets in there and the database is telling me that only five had been executed.” “Oh, great. Well, I’ll execute these other five.” Or it asks the database, “Hey, how many have been executed?” And it says, “Ten.”

And we’ve got a couple of meta tables that we have in the database, real simple, ANSI SQL compliant, that store the changes that happen to the database. So, if it’s a net new database, say you’re running a Docker container with the database in it on your local machine, it’s empty, you would run Liquibase, and it says, “Oh, hey. It’s got that, you know, new database smell. I can run everything.”

And so the interesting thing happens when you start pointing it at an environment that you haven’t updated in a while. So, dev and test typically are going to have a lot of releases. And so there’s going to be little tiny incremental changes, but when it’s time to go to production, Liquibase will catch it up. And so we speak SQL to the database, if it’s a NoSQL database, we’ll speak their API and make the changes requested. And that’s it. It’s very simple in how it works.

The real complex stuff is when we go a couple of inches deeper, when we start doing things like, well, reverse engineering of your database. How can I get a change log of an existing database? Because nobody starts out using Liquibase for a project. You always do it later.

Corey: No, no. It’s one of those things where when you’re doing a project to see if it works, it’s one of those, “Great, I’ll run a database in some local Docker container or something just to prove that it works.” And, “Todo: fix this later.” And yeah, that todo becomes load-bearing.

Robert: [laugh]. That’s scary. And so, you know, we can help, like, reverse engineering an entire database schema, no problem. We also have things called quality checks. So sure, you can test your Liquibase change against an empty database and it will tell you if it’s syntactically correct—you’ll get an error if you need to fix something—but it doesn’t enforce things like corporate standards. “Tables start with T underscore.” “Do not create a foreign key unless those columns have an ID already applied.” And that’s what our quality checks does. We used to call it rules, but nobody likes rules, so we call it quality checks now.

Corey: How do you avoid the trap of enumerating all the bad things you’ve seen happen because at some point, it feels like that’s what leads to process ossification at large companies where, “Oh, we had this bad thing happen once, like, a disk filled up, so now we have a check that makes sure that all the disks are at least 20, empty.” Et cetera. Great. But you keep stacking those you have thousands and thousands and thousands of those, and even a one-line code change then has to pass through so many different tests to validate that this isn’t going to cause the failure mode that happened that one time in a unicorn circumstance. How do you avoid the bloat and the creep of stuff like that?

Robert: Well, let’s look at what we’ve learned from automated testing. We certainly want more and more tests. Look, DevOp’s algorithm is, “All right, we had a problem here.” [laugh]. Or SRE algorithm, I should say. “We had a problem here. What happened? What are we going to change in the future to make sure this doesn’t happen?” Typically, that involves a new standard.

Now, ossification occurs when a person has to enforce that standard. And what we should do is seek to have automation, have the machine do it for us. Have the humans come up and identify the problem, find a creative way to look for the issue, and then let the machine enforce it. Ossification happens in large organizations when it’s people that are responsible, not the machine. The machines are great at running these things over and over again, and they’re never hung over, day after Super Bowl Sunday, their kid doesn’t get sick, they don’t get sick. But we want humans to look at the things that we need that creative energy, that brain power on. And then the rote drudgery, hand that off to the machine.

Corey: Drudgery seems like sort of a job description for a lot of us who spend time doing operation stuff.

Robert: [laugh].

Corey: It’s drudgery and it’s boring, punctuated by moments of sheer terror. On some level, you’re more or less taking some of the adrenaline high of this job away from people. And you know, when it comes to databases, I’m kind of okay with that as it turns out.

Robert: Yeah. Oh, yeah, we want no surprises in database-land. And that is why over the past several decades—can I say several decades since 1979?

Corey: Oh, you can s—it’s many decades, I’m sorry to burst your bubble on that.

Robert: [laugh]. Thank you, Corey. Thank you.

Corey: Five, if we’re being honest. Go ahead.

Robert: So, it has evolved over these many decades where change is the enemy of stability. And so we don’t want change, and we want to lock these things down. And our database professionals have become changed from sentinels of data into traffic cops and TSA. And as we all know, some things slip through those. Sometimes we speed, sometimes things get snuck through TSA.

And so what we need to do is create a system where it’s not the people that are in charge of that; that we can set these policies and have our database professionals do more valuable things, instead of that adrenaline rush of, “Oh, my God,” how about we get the rush of solving a problem and saving the company millions of dollars? How about that rush? How about the rush of taking our old, busted on-prem databases and figure out a way to scale these up in the cloud, and also provide quick dev and test environments for our developer and test friends? These are exciting things. These are more fun, I would argue.

Corey: You have a list of reference customers on your website that are awesome. In fact, we share a reference customer in the form of Ticketmaster. And I don’t think that they will get too upset if I mention that based upon my work with them, at no point was I left with the impression that they played fast and loose with databases. This was something that they take very seriously because for any company that, you know, sells tickets to things you kind of need an authoritative record of who’s bought what, or suddenly you don’t really have a ticket-selling business anymore. You also reference customers in the form of UPS, which is important; banks in a variety of different places.

Yeah, this is stuff that matters. And you support—from the looks of it—every database people can name except for Route 53. You’ve got RDS, you’ve got Redshift, you’ve got Postgres-squeal, you’ve got Oracle, Snowflake, Google’s Cloud Spanner—lest people think that it winds up being just something from a legacy perspective—Cassandra, et cetera, et cetera, et cetera, CockroachDB. I could go on because you have multiple pages of these things, SAP HANA—whatever the hell that’s supposed to be—Yugabyte, and so on, and so forth. And it’s like, some of these, like, ‘now you’re just making up animals’ territory.

Robert: Well, that goes back to open-source, you know, you were talking about that earlier. There is no way in hell we could have brought out support for all these database platforms without us being open-source. That is where the community aligns their goals and works to a common end. So, I’ll give you an example. So, case in point, recently, let me see Yugabyte, CockroachDB, AWS Redshift, and Google Cloud Spanner.

So, these are four folks that reached out to us and said, either A) “Hey, we want Liquibase to support our database,” or B) “We want you to improve the support that’s already there.” And so we have what we call—which is a super creative name—the Liquibase test harness, which is just genius because it’s an automated way of running a whole suite of tests against an arbitrary database. And that helped us partner with these database vendors very quickly and to identify gaps. And so there’s certain things that AWS Redshift—certain objects—that AWS Redshift doesn’t support, for all the right reasons. Because it’s data warehouse.

Okay, great. And so we didn’t have to run those tests. But there were other tests that we had to run, so we create a new test for them. They actually wrote some of those tests. Our friends at Yugabyte, CockroachDB, Cloud Spanner, they wrote these extensions and they came to us and partnered with us.

The only way this works is with open-source, by being open, by being transparent, and aligning what we want out of life. And so what our friends—our database friends—wanted was they wanted more tooling for their platform. We wanted to support their platform. So, by teaming up, we help the most important person, [laugh] the most important person, and that’s the customer. That’s it. It was not about, “Oh, money,” and all this other stuff. It was, “This makes our customers' lives easier. So, let’s do it. Oop, no brainer.”

Corey: There’s something to be said for making people’s lives easier. I do want to talk about that open-source versus commercial divide. If I Google Liquibase—which, you know, I don’t know how typing addresses in browsers works anymore because search engines are so fast—I just type in Liquibase. And the first thing it spits me out to is liquibase.org, which is the Community open-source version. And there’s a link there to the Pro paid version and whatnot. And I was just scrolling idly through the comparison chart to see, “Oh, so ‘Community’ is just code for shitty and you’re holding back advanced features.” But it really doesn’t look that way. What’s the deal here?

Robert: Oh, no. So, Liquibase open-source project started in 2006 and Liquibase the company, the commercial entity, started after that, 2012; 2014, first deal. And so, for—Nathan Voxland started this, and Nathan was struggling. He was working at a company, and he had to have his application—of course—you know, early 2000s, J2EE—support SQL Server and Oracle and he was struggling with it. And so he open-sourced it and added more and more databases.

Certainly, as open-source databases grew, obviously he added those: MySQL, Postgres. But we’re never going to undo that stuff. There’s rollback for free in Liquibase, we’re not going to be [laugh] we’re not going to be jerks and either A) pull features out or, B) even worse, make Stephen O’Grady’s life awful by changing the license [laugh] so he has to write about it. He loves writing about open-source license changes. We’re Apache 2.0 and so you can do whatever you want with it.

And we believe that the things that make sense for a paying customer, which is database-specific objects, that makes sense. But Liquibase Community, the open-source stuff, that is built so you can go to any database. So, if you have a change log that runs against Oracle, it should be able to run against SQL Server, or MySQL, or Postgres, as long as you don’t use platform-specific data types and those sorts of things. And so that’s what Community is about. Community is about being able to support any database with the same change log. Pro is about helping you get to that next level of DevOps Nirvana, of reaching those four metrics that Dr. Forsgren tells us are really important.

Corey: Oh, yes. You can argue with Nicole Forsgren, but then you’re wrong. So, why would you ever do that?

Robert: Yeah. Yeah. [laugh]. It’s just—it’s a sucker’s bet. Don’t do it. There’s a reason why she’s got a PhD in CS.

Corey: She has been a recurring guest on this show, and I only wish she would come back more often. You and I are fun to talk to, don’t get me wrong. We want unbridled intellect that is couched in just a scintillating wit, and someone is great to talk to. Sorry, we’re both outclassed.

Robert: Yeah, you get entertained with us; you learn with her.

Corey: Exactly. And you’re still entertained while doing it is the best part.

Robert: [laugh]. That’s the difference between Community and Pro. Look, at the end of the day, if you’re an individual developer just trying to solve a problem and get done and away from the computer and go spend time with your friends and family, yeah, go use Liquibase Community. If it’s something that you think can improve the rest of the organization by teaming up and taking advantage of the collaboration features? Yes, sure, let us know. We’re happy to help.

Corey: Now, if people wanted to become an attorney, but law school was too expensive, out of reach, too much time, et cetera, but they did have a Twitter account, very often, they’ll find that they can scratch that itch by arguing online about open-source licenses. So, I want to be very clear—because those people are odious when they email me—that you are licensed under the Apache License. That is a bonafide OSI approved open-source license. It is not everyone except big cloud companies, or service providers, which basically are people dancing around—they mean Amazon. So, let’s be clear. One, are you worried about Amazon launching a competitive service with a dumb name? And/or have you really been validated as a product if AWS hasn’t attempted and failed to launch a competitor?

Robert: [laugh]. Well, I mean, we do have a very large corporation that has embedded Liquibase into one of their flagship products, and that is Oracle. They have embedded Liquibase in SQLcl. We’re tickled pink because that means that, one, yes, it does validate Liquibase is the right way to do it, but it also means more people are getting help. Now, for Oracle users, if you’re just an Oracle shop, great, have fun. We think it’s a great solution. But there’s not a lot of those.

And so we believe that if you have Liquibase, whether it’s open-source or the Pro version, then you’re going to be able to support all the databases, and I think that’s more important than being tied to a single cloud. Also—this is just my opinion and take it for what it’s worth—but if Amazon wanted to do this, well, they’re not the only game in town. So, somebody else is going to want to do it, too. And, you know, I would argue even with Amazon’s backing that Liquibase is a little stronger brand than anything they would come out with.

Corey: This episode is sponsored by our friends at Oracle HeatWave is a new high-performance accelerator for the Oracle MySQL Database Service. Although I insist on calling it “my squirrel.” While MySQL has long been the worlds most popular open source database, shifting from transacting to analytics required way too much overhead and, ya know, work. With HeatWave you can run your OLTP and OLAP, don’t ask me to ever say those acronyms again, workloads directly from your MySQL database and eliminate the time consuming data movement and integration work, while also performing 1100X faster than Amazon Aurora, and 2.5X faster than Amazon Redshift, at a third of the cost. My thanks again to Oracle Cloud for sponsoring this ridiculous nonsense.

Corey: So, I want to call out though, that on some level, they have already competed with you because one of database that you do not support is DynamoDB. Let’s ignore the Route 53 stuff because, okay. But the reason behind that, having worked with it myself, is that, “Oh, how do you do a schema change in DynamoDB?” The answer is that you don’t because it doesn’t do schemas for one—it is schemaless, which is kind of the point of it—as well as oh, you want to change the primary, or the partition, or the sort key index? Great. You need a new table because those things are immutable.

So, they’ve solved this Gordian Knot just like Alexander the Great did by cutting through it. Like, “Oh, how do you wind up doing this?” “You don’t do this. The end.” And that is certainly an approach, but there are scenarios where those were first, NoSQL is not a acceptable answer for some workloads.

I know Rick [Horahan 00:26:16] is going to yell at me for that as soon as he hears me, but okay. But there are some for which a relational database is kind of a thing, and you need that. So, Dynamo isn’t fit for everything. But there are other workloads where, okay, I’m going to just switch over. I’m going to basically dump all the data and add it to a new table. I can’t necessarily afford to do that with anything less than maybe, you know, 20 milliseconds of downtime between table one and table two. And they’re obnoxious and difficult ways to do it, but for everything else, you do kind of need to make ALTER TABLE changes from time to time as you go through the build and release process.

Robert: Yeah. Well, we certainly have plans for DynamoDB support. We are working our way through all the NoSQLs. Started with Mongo, and—

Corey: Well, back that out a second then for me because there’s something I’m clearly not grasping because it’s my understanding, DynamoDB is schemaless. You can put whatever you want into various arbitrary fields. How would Liquibase work with something like that?

Robert: Well, that’s something I struggled with. I had the same question. Like, “Dude, really, we’re a schema change tool. Why would we work with a schemaless database?” And so what happened was a soon-to-be friend of ours in Europe had reached out to me and said, “I built an extension for MongoDB in Liquibase. Can we open-source this, and can y’all take care of the care and feeding of this?” And I said, “Absolutely. What does it do?” [laugh].

And so I looked at it and it turns out that it focuses on collections and generating data for test. So, you’re right about schemaless because these are just documents and we’re not going to go through every single document and change the structure, we’re just going to have the application create a new doc and the new format. Maybe there’s a conversion log logic built into the app, who knows. But it’s the database professionals that have to apply these collections—you know, indices; that’s what they call them in Mongo-land: collections. And so being able to apply these across all environments—dev, test, production—and have consistency, that’s important.

Now, what was really interesting is that this came from MasterCard. So, this engineer had a consulting business and worked for MasterCard. And they had a problem, and they said, “Hey, can you fix this with Liquibase?” And he said, “Sure, no problem.” And he built it.

So, that’s why if you go to the MongoDB—the liquibase-mongodb repository in our Liquibase org, you’ll see that MasterCard has the copyright on all that code. Still Apache 2.0. But for me, that was the validation we needed to start expanding to other things: Dynamo, Couch. And same—

Corey: Oh, yeah. For a lot of contributors, there’s a contributor license process you can go through, assign copyright. For everything else, there’s MasterCard.

Robert: Yeah. Well, we don’t do that. Look, you know, we certainly have a code of conduct with our community, but we don’t have a signing copyright and that kind of stuff. Because that’s baked into Apache 2.0. So, why would I want to take somebody’s ability to get credit and magical internet points and increase the rep by taking that away? That’s just rude.

Corey: The problem I keep smacking myself into is just looking at how the entire database space across the board goes, it feels like it’s built on lock-in, it’s built on it is super finicky to work with, and it generally feels like, okay, great. You take something like Postgres-squeal or whatever it is you want to run your database on, yeah, you could theoretically move it a bunch of other places, but moving databases is really hard. Back when I was at my last, “Real job,” quote-unquote, years ago, we were late to the game; we migrated the entire site from EC2 Classic into a VPC, and the biggest pain in the ass with all of that was the RDS instance. Because we had to quiesce the database so it would stop taking writes; we would then do snapshot it, shut it down, and then restore a new database from that RDS snapshot.

How long does it take, at least in those days? That is left as an experiment for the reader. So, we booked a four hour maintenance window under the fear that would not be enough. It completed in 45 minutes. So okay, there’s that. Sparked the thing up and everything else was tested and good to go. And yay. Okay.

It took a tremendous amount of planning, a tremendous amount of work, and that wasn’t moving it very far. It is the only time I’ve done a late-night deploy, where not a single thing went wrong. Until I was on the way home and the Uber driver sideswiped a city vehicle. So, there we go—

Robert: [laugh].

Corey: —that’s the one. But everything else was flawless on this because we planned these things out. But imagine moving to a different provider. Oh, forget it. Or imagine moving to a different database engine? That’s good. Tell another one.

Robert: Well, those are the problems that we want our database professionals to solve. We do not want them to be like janitors at an elementary school, cleaning up developer throw-up with sawdust. The issue that you’re describing, that’s a one time event. This is something that doesn’t happen very often. You need hands on the keyboard, you want people there to look for problems.

If you can take these database releases away from those folks and automate them safely—you can have safety and speed—then that frees up their time to do these other herculean tasks, these other feats of strength that they’re far better at. There is no silver bullet panacea for database issues. All we’re trying to do is take about 70% of DBAs time and free it up to do the fun stuff that you described. There are people that really enjoy that, and we want to free up their time so they can do that. Moving to another platform, going from the data center to the cloud, these sorts of things, this is what we want a human on; we don’t want them updating a column three times in a row because dev couldn’t get it right. Let’s just give them the keys and make sure they stay in their lane.

Corey: There’s something glorious about being able to do that. I wish that there were more commonly appreciated ways of addressing those pains, rather than, “Oh, we’re going to sell you something big and enterprise-y and it’s going to add a bunch of process and not work out super well for you.” You integrate with existing CI/CD systems reasonably well, as best I can tell because the nice thing about CI/CD—and by nice I mean awful—is that there is no consensus. Every pipeline you see, in a release engineering process inherently becomes this beautiful bespoke unicorn.

Robert: Mm-hm. Yeah. And we have to. We have to integrate with whatever CI/CD they have in place. And we do not want customers to just run Liquibase by itself. We want them to integrate it with whatever is driving that application deployment.

We’re Switzerland when it comes to databases, and CI/CD. And I certainly have my favorite of those, and it’s primarily based on who bought me drinks at the last conference, but we cannot go into somebody’s house and start rearranging the furniture. That’s just rude. If they’re deploying the app a certain way, what we tell that customer is, “Hey, we’re just going to have that CI/CD tool call Liquibase to update the database. This should be an atomic unit of deployment.” And it should be hidden from the person that pushes that shiny button or the automation that does it.

Corey: I wish that one day that you could automate all of the button pushing, but the thing that always annoyed me in release engineering was the, “Oh, and here’s where we stop to have a human press the button.” And I get it. That stuff’s scary for some folks, but at the same time, this is the nature of reality. So, you’re not going to be able to technology your way around people. At least not successfully and not for very long.

Robert: It’s about trust. You have to earn that database professional’s trust because if something goes wrong, blaming Liquibase doesn’t go very far. In that company, they’re going to want a person [laugh] who has a badge to—with a throat to choke. And so I’ve seen this pattern over and over again.

And this happened at our first customer. Major, major, big, big, big bank, and this was on the consumer side. They were doing their first production push, and they wanted us ready. Not on the call, but ready if there was an issue they needed to escalate and get us to help them out. And so my VP of Engineering and me, we took it. Great. Got VP of engineering and CTO. Right on.

And so Kevin and I, we stayed home, stayed sober [laugh], you know—a lot of places to party in Austin; we fought that temptation—and so we stayed and I’m texting with Kevin, back and forth. “Did you get a call?” “No, I didn’t get a call.” It was Friday night. Saturday rolls around. Sunday. “Did you get a—what’s going on?” [laugh].

Monday, we’re like, “Hey. Everything, okay? Did you push to the next weekend?” They’re like, “Oh, no. We did. It went great. We forgot to tell you.” [laugh]. But here’s what happened. The DBAs push the Liquibase ‘make it go’ button, and then they said, “Uh-Oh.” And we’re like, “What do you mean, uh-oh?” They said, “Well, something went wrong.” “Well, what went wrong?” “Well, it was too fast.” [laugh]. Something—no way. And so they went through the whole thing—

Corey: That was my downtime when I supposed to be compiling.

Robert: Yeah. So, they went through the whole thing to verify every single change set. Okay, so that was weekend one. And then they go to weekend two, they do it the same thing. All right, all right. Building trust.

By week four, they called a meeting with the release team. And they said, “Hey, process change. We’re no longer going to be on these calls. You are going to push the Liquibase button. Now, if you want to integrate it with your CI/CD, go right ahead, but that’s not my problem.” Dev—or, the release team is tier one; dev is tier two; we—DBAs—are tier three support, but we’ll call you because we’ll know something went wrong. And to this day, it’s all automated.

And so you have to earn trust to get people to give that up. Once they have trust and you really—it’s based on empathy. You have to understand how terrible [laugh] they are sometimes treated, and to actively take care of them, realize the problems they’re struggling with, and when you earn that trust, then and only then will they allow automation. But it’s hard, but it’s something you got to do.

Corey: You mentioned something a minute ago that I want to focus on a little bit more closely, specifically that you’re in Austin. Seems like that’s a popular choice lately. You’ve got companies that are relocating their headquarters there, presumably for tax purposes. Oracle’s there, Tesla’s there. Great. I mean, from my perspective, terrific because it gets a number of notably annoying CEOs out of my backyard. But what’s going on? Why is Austin on this meteoric rise and how’d it get there?

Robert: Well, a lot of folks—overnight success, 40 years in the making, I guess. But what a lot of people don’t realize is that, one, we had a pretty vibrant tech hub prior to all this. It all started with MCC, Microcomputer Consortium, which in the ’80s, we were afraid of the Japanese taking over and so we decided to get a bunch of companies together, and Admiral Bobby Inman who was director planted it in Austin. And that’s where it started. You certainly have other folks that have a huge impact, obviously, Michael Dell, Austin Ventures, a whole host of folks that have really leaned in on tech in Austin, but it actually started before that.

So, there was a time where Willie Nelson was in Nashville and was just fed up with RCA Records. They would not release his albums because he wanted to change his sound. And so he had some nice friends at Atlantic Records that said, “Willie, we got this. Go to New York, use our studio, cut an album, we’ll fix it up.” And so he cut an album called Shotgun Willie, famous for having “Whiskey River” which is what he uses to open and close every show.

But that album sucked as far as sales. It’s a good album, I like it. But it didn’t sell except for one place in America: in Austin, Texas. It sold more copies in Austin than anywhere else. And so Willie was like, “I need to go check this out.”

And so he shows up in Austin and sees a bunch of rednecks and hippies hanging out together, really geeking out on music. It was a great vibe. And then he calls, you know, Kris, and Waylon, and Merle, and say, “Come on down.” And so what happened here was a bunch of people really wanted to geek out on this new type of country music, outlaw country. And it started a pattern where people just geek out on stuff they really like.

So, same thing with Austin film. You got Robert Rodriguez, you got Richard Linklater, and Slackers, his first movie, that’s why I moved to Austin. And I got a job at Les Amis—a coffee shop that’s closed—because it had three scenes in that. There was a whole scene of people that just really wanted to make different types of films. And we see that with software, we see that with film, we see it with fashion.

And it just seems that Austin is the place where if you’re really into something, you’re going to find somebody here that really wants to get into it with you, whether it’s board gaming, D&D, noise punk, whatever. And that’s really comforting. I think it’s the community that’s just welcoming. And I just hope that we can continue that creativity, that sense of community, and that we don’t have large corporations that are coming in and just taking from the system. I hope they inject more.

I think Oracle’s done a really good job; their new headquarters is gorgeous, they’ve done some really good things with the city, doing a land swap, I think it was forty acres for nine acres. They coughed up forty for nine. And it was nine acres the city wasn’t even using. Great. So, I think they’re being good citizens. I think Tesla’s been pretty cool with building that factory where it is. I hope more come. I hope they catch what is ever in the water and the breakfast tacos in Austin.

Corey: [laugh]. I certainly look forward to this pandemic ending; I can come over and find out for myself. I’m looking forward to it. I always enjoyed my time there, I just wish I got to spend more of it.

Robert: How many folks from Duckbill Group are in Austin now?

Corey: One at the moment. Tim Banks. And the challenge, of course, is that if you look across the board, there really aren’t that many places that have more than one employee. For example, our operations person, Megan, is here in San Francisco and so is Jesse DeRose, our manager of cloud economics. But my business partner is in Portland; we have people scattered all over the country.

It’s kind of fun having a fully-distributed company. We started this way, back when that was easy. And because all right, travel is easy; we’ll just go and visit whenever we need to. But there’s no central office, which I think is sort of the dangerous part of full remote because then you have this idea of second-class citizens hanging out in one part of the country and then they go out to lunch together and that’s where the real decisions get made. And then you get caught up to speed. It definitely fosters a writing culture.

Robert: Yeah. When we went to remote work, our lease was up. We just didn’t renew. And now we have expanded hiring outside of Austin, we have folks in the Ukraine, Poland, Brazil, more and more coming. We even have folks that are moving out of Austin to places like Minnesota and Virginia, moving back home where their family is located.

And that is wonderful. But we are getting together as a company in January. We’re also going to, instead of having an office, we’re calling it a ‘Liquibase Lounge.’ So, there’s a number of retail places that didn’t survive, and so we’re going to take one of those spots and just make a little hangout place so that people can come in. And we also want to open it up for the community as well.

But it’s very important—and we learned this from our friends at GitLab and their culture. We really studied how they do it, how they’ve been successful, and it is an awareness of those lunch meetings where the decisions are made. And it is saying, “Nope, this is great we’ve had this conversation. We need to have this conversation again. Let’s bring other people in.” And that’s how we’re doing at Liquibase, and so far it seems to work.

Corey: I’m looking forward to seeing what happens, once this whole pandemic ends, and how things continue to thrive. We’re long past due for a startup center that isn’t San Francisco. The whole thing is based on the idea of disruption. “Oh, we’re disruptive.” “Yes, we’re so disruptive, we’ve taken a job that can be done from literally anywhere with internet access and created a land crunch in eight square miles, located in an earthquake zone.” Genius, simply genius.

Robert: It’s a shame that we had to have such a tragedy to happen to fix that.

Corey: Isn’t that the truth?

Robert: It really is. But the toothpaste is out of the tube. You ain’t putting that back in. But my bet on the next Tech Hub: Kansas City. That town is cool, it has one hundred percent Google Fiber all throughout, great university. Kauffman Fellows, I believe, is based there, so VC folks are trained there. I believe so; I hope I’m not wrong with that. I know Kauffman Foundation is there. But look, there’s something happening in that town. And so if you’re a buy low, sell high kind of person, come check us out in Austin. I’m not trying to dissuade anybody from moving to Austin; I’m not one of those people. But if the housing prices [laugh] you don’t like them, check out Kansas City, and get that two-gig fiber for peanuts. Well, $75 worth of peanuts.

Corey: Robert, I want to thank you for taking the time to speak with me so extensively about Liquibase, about how awesome RedMonk is, about Austin and so many other topics. If people want to learn more, where can they find you?

Robert: Well, I think the best place to find us right now is in AWS Marketplace. So—

Corey: Now, hand on a second. When you say the best place for anything being the AWS Marketplace, I’m naturally a little suspicious. Tell me more.

Robert: [laugh]. Well, best is, you know, it’s—[laugh].

Corey: It is a place that is there and people can find you through it. All right, then.

Robert: I have a list. I have a list. But the first one I’m going to mention is AWS Marketplace. And so that’s a really easy way, especially if you’re taking advantage of the EDP, Enterprise Discount Program. That’s helpful. Burn down those dollars, get a discount, et cetera, et cetera. Now, of course, you can go to liquibase.com, download a trial. Or you can find us on Github, github.com/liquibase. Of course, talking smack to us on Twitter is always appreciated.

Corey: And we will, of course, include links to that in the [show notes 00:46:37]. Robert Reeves, CTO and co-founder of Liquibase. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice along with an angry comment complaining about how Liquibase doesn’t support your database engine of choice, which will quickly be rendered obsolete by the open-source community.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Danielle

Danielle Baskin is a serial entrepreneur and multimedia artist whose work has been featured in The New York Times, The Guardian, NPR, The New Yorker, WSJ, and more. She's also the CEO of Dialup, a globally acclaimed voice-chat app.

Links:

  • Dialup: https://dialup.com
  • Twitter: https://twitter.com/djbaskin
  • Cofounder Quest: https://cofounder.quest
  • Personal Website: https://daniellebaskin.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: It seems like there is a new security breach every day. Are you confident that an old SSH key, or a shared admin account, isn’t going to come back and bite you? If not, check out Teleport. Teleport is the easiest, most secure way to access all of your infrastructure. The open source Teleport Access Plane consolidates everything you need for secure access to your Linux and Windows servers—and I assure you there is no third option there. Kubernetes clusters, databases, and internal applications like AWS Management Console, Yankins, GitLab, Grafana, Jupyter Notebooks, and more. Teleport’s unique approach is not only more secure, it also improves developer productivity. To learn more visit: goteleport.com. And not, that is not me telling you to go away, it is: goteleport.com.

Corey: You know how Git works right?

Announcer: Sorta, kinda, not really. Please ask someone else.

Corey: That's all of us. Git is how we build things, and Netlify is one of the best ways I’ve found to build those things quickly for the web. Netlify’s Git-based workflows mean you don’t have to play slap-and-tickle with integrating arcane nonsense and web hooks, which are themselves about as well understood as Git. Give them a try and see what folks ranging from my fake Twitter for Pets startup, to global Fortune 2000 companies are raving about. If you end up talking to them—because you don’t have to; they get why self-service is important—but if you do, be sure to tell them that I sent you and watch all of the blood drain from their faces instantly. You can find them in the AWS marketplace or at www.netlify.com. N-E-T-L-I-F-Y dot com.

Corey: Welcome to Screaming in the Cloud, I’m Corey Quinn. It’s always fun when I get the opportunity to talk to people whose work inspires me, and makes me reflect more deeply upon how I go about doing things in various ways. Now, for folks who have been following my journey for a while, it’s pretty clear that humor plays a big part in this, but that is not something that I usually talk about with respect to whose humor inspires me.

Today that’s going to change a little bit. My guest is Danielle Baskin, who among so many other things is the CEO of a company called Dialup, but more notably is renowned for pulling a bunch of—I don’t know if we’d call them pranks. I don’t know if we would call them performance art. I don’t know if we would call them shitposting in real life, but they are all amazing. Danielle, thank you so much for joining. How do you describe what it is that you do?

Danielle: Thanks for having me. Yeah, I’ve used a few different terms. I’ve called it situation design. I’ve called it serious jokes. I have called what I do business art, but all the things you said, shitposting IRL, that’s part of it too.

Corey: It’s been an absolute pleasure to just watch what you’ve done since I first became aware of you, which our mutual friend, Chloe Condon first pointed me in your general direction with, “Hey, Corey, you think you’re funny? You should watch what Danielle is doing.” That’s not how she framed it, but that’s what I took from it because I’m incredibly egotistical, which is now basically a brand slash core personality trait. There you have it.

And I encountered you for the first time in person—I believe only time to date—at I believe it was Oracle OpenWorld on the expo floor. She had been talking about you a couple of days before, and I saw someone who could only be you because you were dressed as a seer to be at Oracle OpenWorld. The joke should be clear to folks but we’ll explain it later for the folks who are—might need to replay that a bit. I staggered up to you with, “Hey, are you Chloe’s friend?”

Let me give listeners here some advice through counterexample. Don’t do that. It makes you look like a sketchy person who has no clue how social graces work. No one has any context and as soon as you said, “No,” I realized, “Oh, I came across as a loon.” I am going to say, “Never mind. My mistake,” and walk away like a sensible person will after bungling an introduction like that. I’m not usually that inartful about these things. I don’t know what the hell happened, but it happens often when we meet people that we consider celebrities, and sorry, for some of us that’s you.

Danielle: [laugh] yeah, also in fairness to you I was probably fully immersed in character being my wizard self, and so I was not there to, you know, be pulled back to reality. For some context, I was at Oracle OpenWorld because I made a thing called same exact name, oracleopenworld.org, but it’s a divination conference for oracles, for fortune-tellers, for wizards, for seers, and it happened at the exact same place in time, so there was a whole crew of people dressed up with capes, and robes, and tall pointy hats doing tarot readings and practicing our divination skills.

Corey: Now, I could wind up applying about two dozen different adjectives to Oracle, but playful is absolutely not one of them. I would not ever accuse Oracle, or frankly any large company of that scale of having anything even remotely resembling a sense of humor. As someone who does have to factor in the not that remote possibility of getting kicked out of events that I attend, how do you handle that and not find yourself arrested?

Danielle: Oh, we were kicked out every single time.

Corey: Oh, good good good.

Danielle: I’ve done this for four years. The first year we were kicked out just because we didn’t have badges. I made up our own conference lanyard; of course, there’s security issues with that. We were pushed out onto the sidewalk, but I wanted to be inside the conference and closer to the building.

The next year I did a two-layer conference badge, so I put the real one underneath the fake one so that if security went up to us we had the right to be there. What sort of happened—so, like, the first year we got kicked out was because we were all distributed; maybe there was like 20 of us. Sometimes we were together. Sometimes we were having our own adventures. My friend Brian decided do a séance for the Deloitte team.

Corey: Well, that’s Deloitte-ful. Tell me more.

Danielle: [laugh]. Brian has never done a séance before, but he is a good improv actor and also a spiritual person, so this is, like, perfect for him. As the Deloitte team if they wanted to do a séance they were, like, sure because I think they didn’t have anything going—I mean, people are bored at this conference.

Corey: Oh, of course, they are.

Danielle: Especially if your boss flew you there to stand at your booth and you’ve been saying the same thing over and over again; you’re looking for something interesting. So, he grabs the pillows from a lounge area and little tea light candles and makes a whole circle so that the team can sit down.

He’s wearing a bright rainbow cape and he stands in the middle and he could have a booming voice if he wants to. So, he just starts riffing and going—he just goes into séance mode, and this was enough to trigger security noticing that something really weird was happening. And when they went—

Corey: They come over and say, “What the hell is this?” The answer was “Kubernetes.”

Danielle: I had said everyone can blame—if you get in trouble just blame me just say, “I’m doing this with my friend, Danielle,” and have them talk to me. I wanted more people to come and be wizards. I don’t want them to worry about it, so I will take all of the issues on me. He said that he should talk to his manager, Danielle, or I don’t know.

He said something that made it seem we were all part of a company. Which then makes it seem like our whole project was secret guerilla marketing for something. And we didn’t pay for booth. We were not selling anything. We were just trolling. Or not troll—I mean, we were having our own divination summit. We were genuinely—

Corey: You were virally marketing is the right answer and from my perspective—

Danielle: Yeah, no, I wasn’t doing viral marketing. They think anything that’s unusual and getting people’s attention has the ultimate goal of selling something, which it’s not a philosophy I live by.

Corey: No, it feels like the weird counter-intuitive thing here is the way to get the blessing of everyone from this would’ve—the only step you missed was charging Deloitte for doing it at their booth because it attracts attention.

Danielle: Oh, sure. Oracle should have been paying us a lot of money for entertaining people. Actually, genuinely I had some real heart-to-heart conversations with people who wanted to have a tarot reading about how should they talk to their boss about not listening to them. This is something magical that happens when you are dressed up in costume and you are acting really weird people feel they can say anything because you’re acting way more unusual than them, so it sort of takes away people’s barriers. So, people are very honest with me about their situation.

People had questions about their family. Anyway, I was in the middle of a heart-to-heart tarot reading, and security at Oracle was alerted to find anyone with a cape. Find the wizards and kick them out because they didn’t pay to be here. There’s some weird marketing thing happen.

Corey: “Find and eject the wizards,” is probably the most surreal thing that they have been told that year.

Danielle: Oh, yeah. And they didn’t know why. The message why I did not transmit to all the security, but they were just told to find us. Two guards with their walkie-talkies in their uniforms went up to me and they had to escort me off the premises. Which means we had to walk through the conference together and I asked them, “Why?” They’re like, “We don’t know. We were just told to find you.”

Corey: Imagine them trying to find you stopping and asking people, “Excuse me, have you seen the wizard?”

Danielle: Exactly.

Corey: It is hard to be taken seriously when asking questions like that.

Danielle: Totally, totally. So yeah, unfortunately, we had to leave and that has consistently happened because I’ve done it four times. The final year I went, there was a message before the event even started that you’re not allowed to wear a cape.

Corey: The fact that you can have actual changes made to company policy for large-scale, incredibly expensive events like that is a sign that you’ve made it.

Danielle: It doesn’t even point to any particular incident. Yeah, it’s cool to have this sort of lore. When I asked in the last year I went, “I asked why can’t we wear a cape?” And one of the event organizer security, I don’t know what her role was. She said, “There was an incident the previous year.” Which she was talking about me and my friends.

Corey: Of course, but that is the best part of it.

Danielle: It’s just lore than something once happened with these, like, dark spirits that tried to mess up the Oracle conference with their magic.

Corey: Times change and events evolve. Years ago I attended an AWS Summit with a large protest sign that said on it AMI has three syllables, and it got a bit of an eyebrow raise from people at the door, but okay, great. Then people started protesting those events for one of the very many reasons people have to protest Amazon, and they keep piling more on that pile all the time which is neither here nor there.

I realized, okay, I can’t do that anymore because regardless of what the sign says I will get tackled at the door for trying to bring something like that in, and I don’t try and actively disrupt keynotes. So okay, it’s time to move on and not get myself viewed through certain lenses that are unhelpful, but it’s always a question of moving on and try to top what I did previous years. Weren’t you also at Dreamforce wearing pajamas?

Danielle: I did a few things at Dreamforce. One year I literally set up a tent. They spend millions of dollars on beautiful fake trees and rocks, and also Dreamforce gets taken over every time the event occurs. I did a few things. I thought I should make it seem like this is real nature so I brought camping gear and a tent and just brought a hiking backpack in.

Set it up in the middle of the conference floor laying by the waterfall, but there were people in suits networking around me that did not ask me any questions. I just stayed in the tent, but then I decided to list it on Airbnb. So, inside my tent, I was making an Airbnb listing telling people that they could stay at Dreamforce and explore the beautiful nature there, but it took an hour-and-a-half to get kicked out.

Corey: The emails that you must have back and forth with places like Airbnb’s customer support line and the rest have got to be legendary at this point.

Danielle: [laugh] I get interesting cease-and-desists. I wish there was more dialogue. With Airbnb I just got my listing taken down and I couldn’t talk to a human, and even when I got kicked out of Dreamforce they wanted me to leave immediately. I totally snuck in; I didn’t have a badge or anything. So, I guess they’re in the right for that. The second year at Dreamforce I wore a ghillie suit so I hid. So, I stayed a little bit after the conference ended by hiding as a bush.

Corey: That is both amazing and probably terrifying for the worker that encountered you while trying to clean up.

Danielle: Oh, I mean often employees—like it depends. Some people find my pranks really delightful because it shakes up their day. Security guards also find this amusing. There’s some type of organizer that absolutely hates my pranks.

Corey: There’s something to be said for self-selecting your own audience. One question that I—sure you get; if I get it I know you get it—where it’s difficult for people to sometimes draw the line between the fun whimsical things that you do as pranks and the actual things that you do. A great example of this is something you’ve been doing for, I think, four years now, the decruiter.

Danielle: Yeah. The decruiter a service that’s the opposite of a recruiter so it is—

Corey: At the first re:Invent AWS had a slide that was apparently he made the night before or something and they misspelled security as decurity. From that perspective, what’s a decruiter?

Danielle: Yes, I love decurity as a way to talk about infiltrating a space, like, “No I’m a decurity officer.” Yeah, decruiter is basically a service where you talk to us to find out if you should quit your job. Instead of finding out if you should work at a place or figuring out what opportunities there are, we discuss the unemployed life—or the inbet—like, being self-employed, between jobs, switching careers, it’s a whole spectrum but there’s a few recruiters and we’re all like very experienced not having an employer or working for a company. And so, we ask people about how would you spend your free time. What’s your financial situation? Are you able to afford leaving? It gets pretty personal, but it’s highly specific therapy, but we also don’t have a high acceptance rate. I’ve only decruited like 15% people that I’ve talked to.

Corey: Most of them realize that, oh, there’s a lot of things I would have to do if I didn’t have a job and I’m just going to stay where I am?

Danielle: Yeah. Well, I think a lot of people think that as soon as they leave their job a lot of other things in their life will magically transform, or they’ll finally be able to do their creative project they’ve always wanted to do. This is true some percentage of the time, but I always encourage people to do things outside of work and not seek in their whole fulfillment through their job.

There’s plenty of time where you can explore other ideas and even overlap them to make sure that like when you quit you have things lined up. A lot of people don’t know how to answer, “If you suddenly left tomorrow and could just float for three months, what would you do?” If people give me a good answer—and this is similar to an actual job interview I was like, “Why are you excited about working this company?”

If people give me a good answer, that’s a conversation. A lot of people have no idea, but they’re just stuck in a situation where there’s things they could do in their outside of work life that would make them feel happier. That’s why it’s sort of like therapy, but there’s a lot of internal company issues that I talk about. A common reason that people want to leave is that they love their role, they love the company’s mission, but they do not like their manager, but their manager is really good friends with the CEO and they absolutely can’t say anything. This is so common.

Corey: They always say people they’ll quit jobs they quit managers and there is something to be said for that.

Danielle: Yes, it’s scary for people to speak up or who do you write a letter to? How do you secretly talk with your team about it? Are you the only one feeling that way? Typically the people that are the most nervous about saying anything are kind of young either in their early 20s and they feel like they can’t say anything.

I encourage them to come up with a strategy for making change within their corporation but sometimes it’s not worth it. If there’s tons of other opportunities for them it’s not worth them fixing their company.

Corey: It’s also I think not incumbent upon people to fix their entire corporate culture unless they’re at a somewhat higher executive level. That’s a fun thing. The derecruiter.com we’ll definitely throw a link to that in the [show notes 00:15:49] and I’ll start driving people to it when they ask me for advice on these things. Then you decided, okay, that’s fun.

You’re one of those people I feel has a bit of the same alignment that I do which is, why do one thing when I could do a bunch of things? And you decided, ah, you’re going to do a startup. What is the best thing that you can do that really can capitalize on emerging cultural trends? That’s right. Getting millennial to make phone calls to each other. Tell me about that story.

Danielle: Yeah, and it’s not just millennials, though I’m millennial. So, a lot of millennials use Dialup. I mean, Dialup started as a project where basically me and a friend set up a robocall between ourselves. So, like a bot would call our phones and if we would pick up we’d both be connected, but neither of us was actually calling each other. So, it was a way to just always be catching up with each other.

So, many friends asked me if they could join the robocalls. That was sort of the seat of Dialup is getting serendipitous phone calls throughout the day that connect you to a person that you might know or might want to meet. Because there’s overlap of interest or overlap of someone you know. It grew from me and 20 friends to now 31,000 people who are actively using it all over the world and these conversations can be really incredible.

Sometimes people stay on the phone for four hours. People have flown out to meet each other. I get notes every day of how a call has impacted someone one. So, that’s what I’m up to now, but I’m trying to do more interesting things with voice technology. I just like realized, oh, the voice as a medium it just transports you to other worlds. You have space to imagine.

I mean, people listening to this podcast right now they’re not seeing us, but they probably are imagining us, what our rooms look like, what we look like. They’re imagining the stories that we’re telling them without the distraction of video. I want to do more interesting things with intimate audio—not broadcast stuff. Not Clubhouse or Spaces or anything like that, but just more interesting ways to connect people in one-on-ones.

Corey: Something I’ve noticed is that the voice has a power that text does not. It makes it easier to remember that there’s a human on the other side of things. It is far easier for me to send off an incendiary tweet at someone than it is for me to call them up and then berate them, not really my style.

The more three-dimensional someone becomes in various capacities and the higher bandwidth the communication takes on, I think the easier it is to remember that most people who don’t work at Facebook wake up in the morning hoping to do a good job today. Extending empathy to the rest of the world, that’s an important thing.

Danielle: Yeah, for sure. It’s incredible that humans can detect emotional qualities in a voice call. It’s hard to describe why, but people can detect pauses and little mutters. You can sort of know when someone’s laughing or when someone’s listening even though you’re missing all of the visual cues.

Corey: This episode is sponsored by our friends at Oracle Cloud. Counting the pennies, but still dreaming of deploying apps instead of "Hello, World" demos? Allow me to introduce you to Oracle's Always Free tier. It provides over 20 free services and infrastructure, networking, databases, observability, management, and security. And—let me be clear here—it's actually free. There's no surprise billing until you intentionally and proactively upgrade your account. This means you can provision a virtual machine instance or spin up an autonomous database that manages itself all while gaining the networking load, balancing and storage resources that somehow never quite make it into most free tiers needed to support the application that you want to build. With Always Free, you can do things like run small scale applications or do proof-of-concept testing without spending a dime. You know that I always like to put asterisks next to the word free. This is actually free, no asterisk. Start now. Visit snark.cloud/oci-free that's snark.cloud/oci-free.

Corey: Taking a glance at dialup.com, it appears to be a completely free service. You mentioned that it has 30,000 folks involved. Are you taking the VC model of we’re going to get a whole bunch of users first and then figure out how to make money later? Sometimes it works super well. Other times it basically becomes Docker retold.

Danielle: I’ve been thinking about this a lot and I swing back and forth. Right now Dialup is its own thing, connecting strangers. It’s free though I do have some paying clients because I do serendipitous one-on-ones within organizations. I’ve got a secret B2B page, and so that is a little bit of revenue. Right now I’m trying to sort of expand beyond Dialup and make a new thing, in which case I am leaning more towards building a sustainable and profitable company rather than do the raise-VC-money-until-you-die model.

Corey: I think it’s long past time to disrupt the trope of starving artist. What about well-paid artist? It seems like that would inspire and empower people to create a lot more art when they’re not worrying about freezing to death. To that end or presumably to that end you are in the process of looking for a co-founder in what is arguably the most Danielle Baskin possible way. How are you doing it?

Danielle: Oh, yeah. I could have done a regular LinkedIn post linking to a Google Doc, but that is not my style, and as a self-employed person I can’t reach out to old coworkers and be like, “Oh, you’re on my team a few years ago. What are you up to now?” So, I’m sort of under-networked and I thought I should make a game that sort of explains what I’m doing, but have people discover the game in an interesting way. So, I bought a bunch of floppy discs—I have a floppy disc dealer outside of LA.

Corey: For those who are not millennials and are in fact younger than that—and of course let’s not forget Gen X, the Baby Boom Generation, the Silent Generation which I can only assume is comprised entirely of people who represent big companies from a PR point of view because they never comment on anything. What is a floppy disc for someone who was born in, I don’t know, 2005?

Danielle: Oh, a floppy disk is how you would run software on your computer.

Corey: Yeah, a USB stick with no capacity you can wreck with a magnet.

Danielle: Yes, it’s like a flat wide USB stick, but it only contains—

Corey: 1.44 megabytes on the three-and-a-half-inch version.

Danielle: I think some of them then went up to 2.88.

Corey: Ohh.

Danielle: You can’t even fit a picture—a modern picture. You could do a super low-resolution pixel art.

Corey: This picture of grandma has a whopping eight pixels in it. Oh, okay, great. I guess.

Danielle: Yeah. More complex software would be eight floppy disks that you have to insert disk A, insert disk B.

Corey: Anti-piracy warnings in that day of ‘don’t copy that floppy.’ It was a seminal thing for a long time.

Danielle: I have it in my game; it says ‘don’t make illegal copies of this game.’ My game is not literally on the floppy disc. All floppy discs come with pretty interesting artwork on the label. There’s a little space for a sticker, and because I have hundreds of floppy disks, I sort of looked at—I had a ton of design inspiration.

So, I made floppy discs in the aesthetic of the other ones that say Cofounder Quest—like it’s this game—and it leads you to a website. I scattered these in strategic places around the bay area, and I also mailed some to people outside of the bay area. If you stumble across this in person or on the internet, it leads you to this adventure game that’s around seven minutes to play.

It really explains what I want to do with Dialup, and explains me, and explains my aesthetic, and the sort of playful experiences that I’m into without telling you. So, you get to really experience it. At the end, it basically leads you to a job description and tells you to reach out to me if you’re interested.

Corey: I was independent for years and I finally decided to take on a business partner. As it turns out, Mike Julian, who’s the CEO of The Duckbill Group and I go back ten years, he's my best friend. I kept correcting him. He introduced me as his friend. I said, “No, Mike, your best friend.” Then I got him on audio at one point saying, “Oh, Corey Quinn? He’s my best friend.” I have that on my soundboard and I play it every time he gets uppity. That’s the sort of nonsense it’s important in a co-founder relationship. It is a marriage in some respects.

Danielle: Oh, for sure.

Corey: It’s a business entity. Each one of you can destroy the other financially in different ways. You have to have shared values. The idea of speed-dating your way through finding some random co-founder as a job application, on some level, has always struck me as a little dissonant. I like the approach you’re taking of this is who I am and how I go about things. If this aligns then we should talk, and if you don’t like this you’re not going to like any of the rest of this.

Danielle: For sure. I’m definitely self-selecting with who would actually reach out after playing. I also understand. I’m not going to find a co-founder in a few weeks. I’m just starting conversations with people and then seeing who I should continue talking to or seeing if we could do a mini-project together.

Yeah, it’s weird. It’s a very intense relationship. That’s why people do end up becoming co-founders with someone that they already know who’s a friend. It’s possible I already know my co-founder and they’ve been in front of me this whole time. I think these sorts of moments happen, but I also think that it’s cool to totally expand your network and meet someone who maybe has an overlap in spirit, but is someone that you would’ve never otherwise met. That there could be this great overlap or convergence there. I wanted to cast a very wide net with who this would reach, but it’s still going to be a multi-month-long process or longer.

Corey: It’s not these one-off projects that are the most interesting part to me. It is the sheer variety and consistency of this. During the pandemic I believe you wound up having the verified checkmark badges for houses and fill out this form if you want one and for folks in San Francisco. Absolutely, of course, I filled that out. I read a fairly bad take news article on it of a bunch of people fell for this prank.

No, absolutely not. If people are familiar with your work then they know exactly what they’re getting into with something like this and you support the kinds of things you want to see more of in the world. I didn’t fall for anything. I wanted to see where it led and that’s how I feel on everything you do.

Danielle: Yeah, you appreciated the joke.

Corey: Yeah.

Danielle: Yeah, I think people who are familiar with my work understand that I take jokes very seriously. So, it’s not simply—like, usually it’s not just a website that’s like, huh, this was a trick. It’s more of an ongoing theater piece. So, I actually did go through all of the applicants for the Blue Check Homes. Oh, for some context, I made a website where you could apply to have a blue verified badge and a plaster crest put on your house if you are a dignified authentic person that lives in the house.

So, I’m interviewing—I narrowed it down to 50 people from all the applicants and I’m going through and interviewing people with a committee. I’m recording all of the interviews because I think this will make an interesting mini-documentary. I’m actually making one in installing one, but I’m documenting all of it.

When I started it—for a lot of projects I don’t have the ending planned yet. I like the sort of joke to unfold on the internet in real-time, and then figure out what the next thing I should do from there is and continue the project in a sort of curious exploratory mindset as opposed to just saying, “All right, the joke is done.”

Corey: What is your process for coming up with this stuff? Because for me the most intimidating thing I ever see in the course of a week is not the inevitable cease and desist I get from every large cloud company for everything I do. Rather an empty page where it’s all right time for me to write a humorous blog post, or start drafting the bones of a Twitter thread, or start writing my resignation and if I don’t come with an idea by the end of it, I’ll submit it. Where does the creative process start from with you?

Danielle: Yeah. I rarely have creative brainstorming sessions. I’m a person who thinks of a million bad ideas and then there’s one good one. My mind leaps to a ton of ideas. I rarely write down ideas. I don’t do any sort of—you might imagine I’m in a room of whiteboards and post-it notes, workshopping things and doing creative brainstorm sessions, but I don’t.

I think I act upon the things that I feel just extremely excited about and feel like I must do this immediately. It’s hard to explain, but with a lot of my ideas, I just feel this surge of energy. I have to do this because no one else will do it and it’s funny at this moment. If I don’t feel that way I kind of don’t do anything and see if the idea keeps reemerging. With a lot of ideas I may be thought of it a year ago and it just kept resurfacing, but I don’t really force myself to churn out creative projects if that makes sense. People have told me that my work reminds them of Mischief. It’s like as a company that puts out a prank on a Tuesday every two weeks.

Corey: Not familiar with them, but there have been a whole bunch of flash mob groups, and other folks who affected just wind up being professional pranksters, which I love the concept.

Danielle: Yeah, yeah, for sure. I do churn out a lot of pranks and I even have my own prank calendar. I’m not strict with my own deadlines and I also think timing is important. So, you might think of a good idea, but then it’s just the spirit of the zeitgeist doesn’t want you to do it that week. I improvise the things that I want to launch. I mostly do things that I just feel are rich in something I could explore.

Like, with Cofounder Quest I was always on the fence about it because it feels to me annoying to tell people you’re trying to hire someone or to put yourself out there and be pitching your startup. So, I was kind of nervous about that, but I also thought if I leave a floppy disk in the park, and then put a picture on the internet it’ll lead to something—there’s something that it will lead to.

It might lead to finding a co-founder. It might lead to meeting interesting people, but also I’ve never built an interactive game with audio and so I was interested in learning that, but yeah, I tend to land on ideas that I think are rich in terms of things I could learn. Things that I could turn into more immersive theater and things that keep resurfacing as opposed to keeping myself on a strict schedule of creative ideas if that makes sense.

Corey: It makes a lot of sense. It’s one of those things that it is not commonly understood for those of us who came up in the nose of the grindstone 40 hours a week, have a work ethic. Even if you’re not busy look busy. Sometimes work looks a lot more like getting up and going to a coffee shop and meeting some stranger from the internet than it does sitting down churning out code.

Danielle: For sure. I think that it is important to continue being in conversations with people. I think good ideas emerge while you’re in the middle of talking, and you realize your own limitations and ideas when you have to explain things to other people. While something you’re very clear in your head as soon as there’s a person you don’t know and they ask you, “What are you working on?” You realize, oh, there’s so many gaps. It made perfect sense to me, but there’s a lot of gaps. So yeah, I think it’s important to stay in dialogue and also have to explain yourself to new people instead of just sort of making ideas in a vacuum.

Corey: I want to thank you for being so generous with your time and talking to me about all the various things you have going on. If people want to follow along and learn more about what you’re up to, where can they find you?

Danielle: I post a lot of my projects on Twitter. So, I’m @djbaskin. If you want to play Cofounder Quest, it’s cofounder.quest. That is an actual domain. I also have a website daniellebaskin.com, which has a lot of my projects, many of which we didn’t discuss. I also do, similar to Oracle OpenWorld, I like to host popup events that involve lots of people trolling. So, if you want to get involved in anything you see I’m always happy to bring more wizards on board.

Corey: We will, of course, put links to that in the [show notes 00:31:10]. Danielle, thank you so much for taking the time to speak with me today.

Danielle: Oh yeah, thanks for having me. It was great talking with you.

Corey: Danielle Baskin, CEO of Dialup, and oh so very much more. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast please leave a five-star review on your podcast platform of choice, along with a long rambling comment applying to be the co-host of this podcast, viewing it of course as a podcasting call.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Sahir

Sahir is responsible for product strategy across the MongoDB portfolio. He joined MongoDB in 2016 as SVP, Cloud Products & GTM to lead MongoDB’s cloud products and go-to-market strategy ahead of the launch of Atlas and helped grow the cloud business from zero to over $150 million annually. Sahir joined MongoDB from Sumo Logic, an SaaS machine-data analytics company, where he managed platform, pricing, packaging and technology partnerships. Before Sumo Logic, Sahir was the Director of Cloud Management Strategy & Evangelism at VMware, where he launched VMware’s first organically developed SaaS management product and helped grow the management tools business to over $1B in revenue. Earlier in his career, Sahir held a variety of technical and sales-focused roles at DynamicOps, BMC Software, and BladeLogic.

Links:

  • MongoDB: https://www.mongodb.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Redis, the company behind the incredibly popular open source database that is not the bind DNS server. If you’re tired of managing open source Redis on your own, or you’re using one of the vanilla cloud caching services, these folks have you covered with the go to manage Redis service for global caching and primary database capabilities; Redis Enterprise. Set up a meeting with a Redis expert during re:Invent, and you’ll not only learn how you can become a Redis hero, but also have a chance to win some fun and exciting prizes. To learn more and deploy not only a cache but a single operational data platform for one Redis experience, visit redis.com/hero. Thats r-e-d-i-s.com/hero. And my thanks to my friends at Redis for sponsoring my ridiculous nonsense.

Corey: Are you building cloud applications with a distributed team? Check out Teleport, an open source identity-aware access proxy for cloud resources. Teleport provides secure access to anything running somewhere behind NAT: SSH servers, Kubernetes clusters, internal web apps and databases. Teleport gives engineers superpowers!

  • Get access to everything via single sign-on with multi-factor.
  • List and see all SSH servers, kubernetes clusters or databases available to you.
  • Get instant access to them all using tools you already have.

Teleport ensures best security practices like role-based access, preventing data exfiltration, providing visibility and ensuring compliance. And best of all, Teleport is open source and a pleasure to use.

Download Teleport at https://goteleport.com. That’s goteleport.com.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. For the first in-person recording in ages we have a promoted guest joining us from MongoDB. Sahir Azam is the chief product officer. Thank you so much for joining me.

Sahir: Thank you for having me, Corey. It’s really exciting to be able to talk and actually meet in person.

Corey: I know it feels a little scandalous these days, when we’re in a position of meeting people in person, like you’re almost like you’re doing something wrong somehow. So, MongoDB has been a staple of the internet for a long time. It’s oh good; another database to keep track of. What do you do these days? What is MongoDB in this ecosystem?

Sahir: That’s a great question. I think we’re fortunate that MongoDB has been very popular for a very long time. We’re seeing, you know, massive adoption grow across the globe and the massive developer community is sort of adopting the technology. What I would bring across is today MongoDB is really one of the leading cloud database companies in the world. The majority of the company’s business comes from our cloud service; we partner very heavily with AWS and other cloud providers on making sure we have global availability of that. That’s our flagship product.

And we’ve invested really heavily in the last I would say, five or six years, and really extending the capabilities of the product to not just be the, sort of, database for modern web scale applications, but also to be able to handle mission-critical use cases across every vertical, you know, enterprises to startups, and doing so in a way that really empowers a general purpose strategy for any app they want to build.

Corey: You’re talking about general purpose which, I guess, leads to the obvious question that AWS has been pushing for a while, the idea of purpose-built databases, which makes sense from a certain point of view, and then they, of course, take that way beyond the bounds of normalcy. I don’t know what the job is for someone whose role is to disambiguate between the 20 different databases that they offer by, who knows, probably the end of this year. And I don’t know what that looks like. What’s your take on that whole idea of a different database for every problem slash every customer slash every employee slash API request.

Sahir: What we see is customers clearly moved to the cloud because they want to be able to move faster, innovate faster, be more competitive in whatever market or business or organization they’re in. And certainly, I think the days of a single vendor database to rule all use cases are gone. We’re not [laugh] by any means supportive of that. However, the idea that you would have 15 different databases that need to be rationalized, integrated, scripted together, frankly, may be interesting for technical teams who want to cobble together, you know, a bespoke architecture. But when we look at it from, sort of a skills, repeatability, cost, simplicity, perspective of architecture, we’re seeing these, sort of like, almost like Rube Goldbergian sort of architectures.

And in a large organization that wants to adopt the cloud en mass, the idea of every development team coming up with their own architecture and spending all of that time and duplication and integration of work is a distraction from ultimately their core mission, which is driving more capability and differentiation in the application for their end customer. So, to be blunt, we actually think the idea of having 15 different databases, ‘the right tool for the job’ is the wrong approach. We think that there’s certain key technologies that most organizations will use for 70, 80% of use cases, and then use the niche technologies for where they really need specialized solutions for particular needs.

Corey: So, if you’re starting off with a general-purpose database then, what is the divergence point at which point—like in my case, eventually I have to admit that using TXT records in Route 53 as a database starts to fall down for certain use cases. Not many, mind you, but one or two here and there. At what point when you’re sticking with a general-purpose database does migrating to something else—what’s the tipping point there?

Sahir: Yeah, I think what we see is if you have a general-purpose database that hits the majority of your needs, oftentimes, especially with a microservices kind of modern architecture, it’s not necessarily replacing your general-purpose database with a completely different solution, it may be augmenting it. So, you may have a particular need for, I don’t know deep graph capabilities, for example, for a particular traversal use case. Maybe you augment that with a specialized solution for that. But the idea is that there’s a certain set of velocity you can enable an organization by building skill set and consolidation around a technology provider that gives much more repeatability, security, less data duplication, and ultimately focuses your organization in teams on innovation as opposed to plumbing and that’s where the 15 different databases been cobbled together may be interesting, but it’s not really focusing on innovation, it’s focusing more on the technology problems that you solved.

Corey: So, we’re recording this on site in Las Vegas, as re:Invent, thankfully and finally, draws to a close. How was your conference?

Sahir: It’s been fantastic. And to be clear, we are huge fans and partners of AWS. This is one of our most exciting conferences we sponsor. We go big, [laugh] we throw a party, we have a huge presence, we have hundreds of customer meetings. So, although I’m a little ragged, as you can probably tell from my voice from many meetings and conversations and drinks with friends, it’s actually been a really great week.

Corey: It is one of those things where having taken a year off, you forget so much of it, where it’s, “Oh, I can definitely walk between those two hotels,” and then you sort of curse the name of God as you wind up going down that path. It was a relief, honestly, to not see, for example, another managed database service being launched that I can recall in that flurry of announcements, did you catch any?

Sahir: I didn’t catch any new particular database services that at least caught my eye. Granted, I’ve been in meetings most of the time, however, we’re really excited about a lot of the infrastructure innovation. You know, I just happened to have a meeting with the compute teams on the Amazon side and what they’re doing with, you know, Wavelength, and Local Zones, and new hardware, and chips with Graviton, it’s all stuff we’re really excited about. So, it is always interesting to see the innovation coming out of AWS.

Corey: You mentioned that you are a partner with AWS, and I get it, but AWS is also one of those companies whose product strategy is ‘yes.’ And they a couple years ago launched their DocumentDB, in parentheses with MongoDB compatibility, which they say, “Oh, customers were demanding this,” but no, no, they weren’t. I’ve been talking to customers; what they wanted was actual MongoDB. The couple of folks I’m talking to who are using it are using it for one reason and one reason only, and that is replication traffic between AZs on native AWS services is free; everyone else must pay. So, there’s some sub-offering in many respects that is largely MongoDB compatible to a point. Okay, but… how do you wind up, I guess, addressing the idea of continuing to partner with a company that is also heavily advantaging its own first party services, even when those are not the thing that best serves customers.

Sahir: Yeah, I’ve been in technology for a while, and you know, the idea of working with major platform players in the context of being, in our case, a customer, a partner, and a competitor is something we’re more than comfortable with, you know, and any organization at our scale and size is navigating those same dynamics. And I think on the outside, it’s very easy to pay way more attention to the competitive dynamics of oh, you run in AWS but you compete with them, but the reality is, honestly, there’s a lot more collaboration, both on the engineering side but also in the field. Like, we go jointly work with customers, getting them onto our platform, way more often than I think the world sees. And that’s a really positive relationship. And we value that and we’re investing heavily on our side to make sure you know, we’re good partners in that sense.

The nuances of DocumentDB versus the real MongoDB, the reality of the situation is yes, if you want the minimal MongoDB experience for, you know, a narrow percentage of our functionality, you can get that from that technology, but that’s not really what customers want. Customers choose MongoDB for the breadth of capabilities that we have, and in particular, in the last few years, it’s not just the NoSQL query capability of Mongo, we’ve integrated rich aggregation capabilities for analytics, transactional guarantees, a globally distributed architecture that scales horizontally and across regions much further than anything a relational architecture can accomplish. And we’ve integrated other domains of data, so things like full text, search, analytics, mobile synchronization are all baked into our Atlas platform. So, to be honest, when customers compare the two on the merits of the technology, we’re more than happy to be competitors with AWS.

Corey: No, I think that everyone competes with AWS, including its own product teams amongst each other because, you know, that’s how you, I guess, innovate more rapidly. What do I know? I don’t run a hyperscale platform. Thankfully.

If I go and pull up your website, it’s mongodb.com. It is natural for me to assume that you make a database, but then I start reading; after the big text and the logo, it says that you are an application data platform. Tell me more about that.

Sahir: Yeah, and this has been a relatively new area of focus for us over the last couple of years. You know, I think many people know MongoDB as a non-relational modern database. Clearly, that’s our core product. I think in general, we have a lot of capabilities in the database that many customers are unaware of in terms of transactional guarantees and schema management and others, so that’s kind of all within the core database. But over the last few years, we’ve both built and acquired technology, things like Realm, that allows for mobile synchronization; event-driven architectures; APIs to be created on your data Easily; Atlas data lake, which allows for data transformation and analytics to be done using the same API as the core Mongo database; as I mentioned a couple minutes ago, things like search, where we actually allow customers to remove the need for a separate search engine for their application and make it really seamless operationally, and from the developer experience standpoint.

And you know, there’s no real term in the industry for that, so we kind of describe ourselves as an application data platform because really, what we’re trying to do is simplify the data architecture for applications, so you don’t need ten different niche database technologies to be able to build a powerful, modern, scalable application; you can build it in a unified way with an amazing developer experience that allows your teams to focus on differentiation and competitiveness as opposed to plumbing together the data infrastructure.

Corey: So, when I hear platform, I think about a number of different things that may or may not be accurate, but the first thing that I think is, “Oh. There’s code running on this then, as sort of part of an ecosystem.” Effectively is their code running on the data platform that you built today that wasn’t written by people at MongoDB?

Sahir: Yes, but it’s typically the customer’s code as part of their application. So, you know, I’ll give you a couple of simple examples. We provide SDKs to be able to build web and mobile applications. We handle the synchronization of data from the client and front end of an application back to the back end seamlessly through our Realm platform. So, we’re certainly, in that case, operating some of the business logic, or extending beyond sort of just the back end data.

Similarly, a lot of what we focus on is modern event-driven architectures with MongoDB. So, to make it easier to create reactive applications, trigger off of changes in your data, we built functions and triggers natively in the platform. Now, we’re not trying to be a full-on application hosting platform; that’s not our business, our business is a data platform, but we really invest in making sure that platform is open, accessible, provides APIs, and functional capabilities make it very easy to integrate into any application our customers want to build.

Corey: It seems like a lot of different companies now are trying to, for lack of a better term, get some of the love that Snowflake has been getting for, “Oh, their data cloud is great.” But when you take a step back and talk to people about, “So, what do you think about Mongo?” The invariable response you’re going to get every time is, “Oh, you mean the database?” Like, “No, no. The character from the Princess Bride. Yes, the database.” How do you view that?

Sahir: Yeah, it’s easy to look at all the data landscape through a simple lens, but the reality is, there’s many sub markets within the database and data market overall. And for MongoDB we’re, frankly, an operational data company. And we’re not focused on data warehousing, although you can use MongoDB for various analytical capabilities. We’re focused on helping organizations build amazing software, and leveraging data as an enabler for great customer experiences, for digital transformation initiatives, for solving healthcare problems, or [unintelligible 00:12:51] problems in the government, or whatever it might be. We’re not really focused on selling customers’—or platforms of data from—not the customers’ data, but other—allowing people to monetize their data. We’re focused on their applications and developers building those experiences.

Corey: Yeah. So, you’re if you were selling customers’ data, you just rebrand as FacebookDB and be done with it, or MetaDB now—

Sahir: MetaDB?

Corey: Yeah. As far as the general Zeitgeist around Mongo goes, back when I was first hearing about it, in I don’t know, I want to say the first half of the 2010s, the running gag was, “Oh, Mongo. It’s Snapchat for databases,” with the gag being that it lost production data was unsafe for a bunch of things. To be clear, based upon my Route 53 comments, I am not a database expert by any stretch of the imagination. Now, the most common thing in my experience that loses production data is me being allowed near it. But what was the story? What gave rise to that narrative?

Sahir: Yeah, I think that—thank you for bringing that up. I mean, to be clear, you know, if a database doesn’t keep your data safe, consistent, and guaranteed, the rest of the functionality doesn’t matter, and we take that extremely seriously at MongoDB. Now, you know, MongoDB, has been around a long time, and for better or worse—I think there’s, frankly, good things and bad things about this—the database exploded in popularity extremely fast, partially because it was so easy to use for developers and it was also very different than the traditional relational database models. And so I think in many ways, customer’s expectation of where the technology was compared to where we were from a maturity standpoint, combined with running an operating it the same way as a traditional system, which was, frankly, wrong for a distributed database caused, unfortunately, some situations where customers stubbed their toes and, you know, we weren’t able to get to them and help them as easily as we could. Thankfully, you know, none of those issues fundamentally are, like, foundational problems. You know, we’ve matured the product for many, many years, you know, we work with 30,000-plus customers worldwide on mission-critical applications. I just want to make sure that everyone understands that, like, we take any issue that has to do with data loss or data corruption, as sort of the foundational [P zero 00:14:56] problem we always have to solve.

Corey: I tend to form a lot of my opinions based upon very little on what, you know, sorry to say it, execs say and a lot more about what I see. There was a whole buzz going around on Twitter that HSBC was moving a whole bunch of its databases over to Mongo. And everyone was saying, “Oh, they’re going to lose all their data.” But I’ve done work with a fair number of financial services companies, and of all the people I talk to, they’re pretty far on one end of that spectrum of, “How cool are we with losing data?” So, voting with a testimonial and a wallet like that—because let’s be clear, getting financial services companies to reference anything for anyone anywhere is like pulling teeth—that says a lot more than any, I guess, PR talking points could.

Sahir: Yeah, I appreciate you saying that. I mean, we’re very fortunate to have a very broad customer base, everything from the world’s largest gaming companies to the world’s largest established banks, the world’s most fastest growing fintechs, to health care organizations distributing vaccines with technologies built on Mongo. Like, you name it, there’s a use case in any vertical, as mission critical as you can think, built on our technology. So, customers absolutely don’t take our word for granted. [laugh]. They go, you know, get comfortable with a new database technology over a span of years, but we’ve really hit sort of mainstream adoption for the majority of organizations. You mentioned financial services, but it’s really any vertical globally, you know, we can count on our customer list.

Corey: How do you, I guess, for lack of a better term, monetize what it is you do when you’re one of the open-source—and yes, if you’re an open-source zealot who wants to complain about licensing, it’s imperative that you do not email me—but you are available for free—for certain definitions of free; I know, I know—that I can get started with a two o’clock in the morning and start running it myself in my environment. What is the tipping point that causes people to say, “Well, that was a good run. Now, I’m going to pay you folks to run it for me.”

Sahir: Yeah, so there’s two different sides to that, first and foremost, the majority of our engineering investment for our business goes in our core database, and our core database is free. And the way we actually, you know, survive and make money as a business, so we can keep innovating, you know, on top of the billion dollars of investment we’ve put in our technology over the years is, for customers who are self-managing in their own data center, we provide a set of management tools, enterprise security integrations, and others that are commercially licensed to be able to manage MongoDB for mission-critical applications in production, that’s a product called Enterprise Advanced. It’s typically used for large enterprise accounts in their own data centers. The flagship product for the company these days, the fastest growing part of the business is a product we call Atlas—or platform we call Atlas. That’s a cloud data service.

So, you know, you can go onto our website, sign up with our free tier, swipe a credit card, all consumption-based, available in every AWS region, as well as Azure and GCP, has the ability to run databases across AWS, Azure, and GCP, which is quite unique to us. And that, like any cloud data technology, is then used in conjunction with a bunch of other application components in the cloud, and customers pay us for the consumption of that database and how much they use.

Corey: This episode is sponsored by our friends at Oracle Cloud. Counting the pennies, but still dreaming of deploying apps instead of "Hello, World" demos? Allow me to introduce you to Oracle's Always Free tier. It provides over 20 free services and infrastructure, networking, databases, observability, management, and security. And—let me be clear here—it's actually free. There's no surprise billing until you intentionally and proactively upgrade your account. This means you can provision a virtual machine instance or spin up an autonomous database that manages itself all while gaining the networking load, balancing and storage resources that somehow never quite make it into most free tiers needed to support the application that you want to build. With Always Free, you can do things like run small scale applications or do proof-of-concept testing without spending a dime. You know that I always like to put asterisks next to the word free. This is actually free, no asterisk. Start now. Visit snark.cloud/oci-free that's snark.cloud/oci-free.

Corey: I want to zero in a little bit on something you just said, where you can have data shared between all three of the primary hyperscalers. That sounds like a story that people like to tell a lot, but you would know far better than I: how common is that use case?

Sahir: It’s definitely one from a strategic standpoint, especially in large enterprises, that’s really important. Now, to your point, the actual usage of cross-cloud databases is still very early, but the fact that customers know that we can go, in three minutes, spin up a database cluster that allows them to either migrate, or span data across multiple regions from multiple providers for high availability, or extend their data to another cloud for analytics purposes or whatnot, is something that it almost is like science fiction to them, but it’s crucial as a capability I know they will need in the future.

Now, to our surprise, we’ve seen more real production adoption of it probably sooner than we would have expected, and there’s kind of three key use cases that come into play. One—you know, for example, I was with a challenger bank from Latin America yesterday; they need high availability in Latin America. In the countries they’re in, no single infrastructure cloud provider has multiple regions. They need to span across multiple regions. They mix and match cloud providers, in their case AWS being their primary, and they have a secondary cloud provider, in their case GCP, for high availability.

But it’s also regulatorily-driven because the banking SEC regulations in that country state that they need to be able to show portability because they don’t want concentration risk of their banking sector to be on a single cloud provider or single cloud provider’s region. So, we see that in multiple countries happening right now. That’s one use case.

The other tends to be geographic reach. So, we work with a very large international gaming company, majority of their use cases happen to be run out of the US. They happen to have a spike of customers using their game [unintelligible 00:19:58] gamers using it in Taiwan; their cloud provider of choice didn’t have a region in Taiwan, but they were able to seamlessly extend a replica into a different cloud to serve low-latency performance in that country. That’s the second.

And then the third, which is a little bit more emerging is kind of the analytic-style use case where you may have your operational data running in a particular cloud provider, but you want to leverage the best of every cloud provider’s, newest, fanciest services on top of your data. So, isn’t it great if you can just hit a couple clicks, we’ll extend your data and keep it in sync in near real time, and allow you to plumb into some new service from another cloud provider.

Corey: In an ideal world with all things being equal, this is a wonderful vision. There’s been a lot of noise made—a fair bit of it by me, let’s be fair—around the data egress pricing for—it’s easy to beat up on AWS because they are the largest cloud provider and it’s not particularly close, but they all do it. Does that serve as a brake on that particular pattern?

Sahir: Thankfully, for a database like ours and various mechanisms we use, it’s not a barrier to entry. It’s certainly a cost component to enabling this capability, for sure. We absolutely would love to see the industry be more open and use less of egress fees as a way to wall people into a particular cloud providers. We certainly have that belief, and would push that notion and continually do in the industry. But it hasn’t been a barrier to adoption because it’s not the major cost component of operating a multi-cloud database.

Corey: Well, [then you start 00:21:27] doing this whole circular replication thing, at which point, wow. It just goes round and round and round and lives on the network all the time. I’m told that’s what a storage area network is because I’m about as good at storage as I am at databases. As you look at Atlas, since you are in all of the major hyperscalers, is the experience different in any way, depending upon which provider you’re running in?

Sahir: By and large, it’s pretty consistent. However, what we are not doing is building to the lowest common denominator. If there’s a service integration that our customers on AWS want, and that service doesn’t integrate, it doesn’t exist on another cloud provider, or vice versa, we’re not going to stop ourselves from building a great customer experience and integration point. And the same thing goes for infrastructure; if there’s some infrastructure innovation that delivers price, performance, great value for our customers and it’s only on a single cloud, we’re not going to stop ourselves from delivering that value to customers. So, there’s a line there, you know, we want to provide a great experience, portability across the cloud providers, consistency where it makes sense, but we are not going to water down our experience on a particular cloud provider if customers are asking for some native capabilities.

Corey: It always feels like a strange challenge historically to wind up—at least in large, regulated environments—getting a new vendor in. Originally an end run around this was using the AWS Marketplace or whatever marketplace you were using at any given cloud provider. Then procurement caught on and in some cases banned in the Marketplace outright and now, the Marketplace is sort of reformed, in some ways, to being a tool for procurement to use. Have you seen significant uptake of your offering through the various cloud marketplaces?

Sahir: We do work with all the cloud marketplaces. In fact, we just made an announcement with AWS that we’re going to be implementing the pay-as-you-go marketplace model for self-service as well on AWS. So, it is definitely a driver for our business. It tends to be used most heavily when we’re selling with the, you know, sales teams from the cloud providers, and customers want to benefit from a single bill, benefit from, you know, drawing down on their large commitments that they might have with any given cloud providers. So, it drives really good alignment between the customer, us as a third-party on AWS or Azure GCP, and the infrastructure cloud provider. And so we’re all aligned on a motion. So, in that sense, it’s definitely been helpful, but it’s largely been a procurement and fulfillment sort of value proposition to drive that alignment, I’d say, by and large today.

Corey: I don’t know if you’re able to answer this without revealing anything confidential, so please feel free not to, but as you look across the total landscape—since I would say that you have a fairly reasonable snapshot of the industry as a whole—am I right when I say that AWS is the behemoth in the space, or is it a closer horse race than most people would believe, based upon your perspective?

Sahir: I think in general, for sure AWS is the market share leader. It would be crazy to say anything otherwise. They innovated this model, you know, the amount of innovation happening at AWS is incredible, you know, and we’re benefiting from it as a customer as well. However, we do believe it’s a multi-cloud future. I mean, look at the growth of Azure. You know, we’re seeing Google show up in large enterprises across the globe as well.

And even beyond the three American clouds, you know, we work heavily with Alibaba and Tencent in mainland China, which is a completely different market than Western world. So, I do think the trend over time will be a more heterogeneous, more multi-cloud world—which I’m biased; that does favor MongoDB, but that’s the trend we’re seeing—but that doesn’t mean that AWS won’t continue to still be a leader and a very strong player in that market.

Corey: I want to talk a little bit about Jepsen. And for those who are unaware, jepsen.io is run by Kyle Kingsbury. Kyle is wonderful, and he’s also nuts. If you followed him back when he was on Twitter, you’ve also certainly seen them.

But beyond that, he is the de facto resource I go to when it comes to consistency testing and stress testing of databases. I’m a little annoyed he hasn’t taken on Route 53 yet, but hope does spring eternal. He’s evaluated Mongo a number of times, and his conclusions, as always are mixed sometimes, shall we say, incendiary, but they always seem relatively fair. What is your experience been, working with him? And do you share my opinion of him as being a neutral and fair arbiter of these things?

Sahir: I do. I think he’s got real expertise and credibility in beating up distributed database systems and finding the edges of where they don’t live up to what we all hope they do, right? Whether it’s us or anyone else, just to be clear. And so anytime Kyle finds some flaw in MongoDB, we take it seriously, we add it to our test suite, [laugh] we remediate, and I think we have a pretty good history of that. And in fact, we’ve actually worked with Kyle to welcome him beating up our database on multiple occasions, too, so it’s not an adversarial relationship at all.

Corey: I have to ask, since you are a more modern generation of database, then many from the previous century, but there’s always been a significant, shall we say… concern, when I wind up looking at it [it again in 00:26:33] any given database, and I look in the terms and conditions and, like, “Oh, it’s a great database. We’re by far the best. Whatever you do, do not publish benchmarks.” What’s going on with that?

Sahir: I think benchmarks can be spun in any direction you want, by any vendor. And it’s not just database technology. I’ve been in IT for a while, and you know, that applies to any technology. So, we absolutely do not shy away from our performance or benchmark or comparisons to any technology. We just think that, you know, vendors benchmarking technologies for their—are doing so largely to only make their own technologies look good versus competition.

Corey: I tend to be somewhat skeptical of the various benchmark stuff. I remember repeatedly oh, I’ll wind up running whatever it is—I think it’s Geek Speed—on my various devices to see oh, how snappy and performant is it going to be? But then I’m sitting there opening Microsoft Word and watching the beach ball spin, and spin, and spin, and it turns out, don’t care about benchmarks in a real-world use case in many scenarios.

Sahir: Yeah, it’s kind of a good analogy, right? I mean, performance of an application, sure, the database at the heart of it is a crucial component, but there’s many more aspects of it that have to do with the overall real world performance than just some raw benchmark results for any database, right? It’s the way you model your data, the way the rest of the architecture of the application interacts and hangs together with the database, many, many layers of complexity. So, I don’t always think those benchmarks are indicative of how real world performance will look, but at the same time, I’m very confident in MongoDB’s performance comparatively to our peers, so it’s not something we’re afraid of.

Corey: As you take a look at where you’ve been and where you are now, what’s next? Where are you going? Because I have a hard time believing that, “Yep, we’re deciding it’s feature complete and we’re just going to sell this until the end of time exactly as is, we’re laying off our entire engineering team and we’re going to be doing support from our yacht, parked comfortably in international waters.” That’s a slightly different company. What’s the plan?

Sahir: So, [laugh] you’re—we are not parking anything, anytime soon. We are continuing to invest heavily in the innovation of the technology, and really, it’s two reasons: you know, one, we’re seeing an acceleration of adoption of MongoDB, either with any customers that have used us for a long time, but for more important and more use cases, but also just broader adoption globally as more and more developers learn to code, they’re choosing Mongo as the place to start, increasingly. And so that’s really exciting for us, and we need to keep up with those customer demands and that roadmap of asks that they have.

And at the same time, customer requirements are increasing as more and more organizations are software-first organizations, the requirements of what they demand from us continually increase, which requires continual innovation in our architecture and our functionality to keep up with those and stay ahead of those customer requirements. So, what you’ll see from us is, one, making sure we can build the best modern database we can. That’s the core of what we do; everything we do now especially is cloud first, so working closely with our cloud partners on that. And even though we’re very fortunate to be a high-performance, high-growth company with a very pervasive open technology, we’re still in a giant market that has a lot of legacy technologies powering old applications. So, [laugh] you know, we have a long, long runway to become a long-standing major player in this market.

And then we’re going to continue this vision of an application data platform, which is really just about simplifying the capabilities and data architecture for organizations and developers so they can focus on building their application and less on the plumbing.

Corey: I want to thank you so much for taking the time to speak with me today. If people want to learn more, where can they go?

Sahir: Clearly, you can go to mongodb.com. You can also reach out to us on our community sites: our own or on any of the public sites that you would typically find developers hanging out. We always have folks from our teams or our champions program of advocates worldwide helping out our customers and users. And I just want to thank you, Corey, for having me. I’ve followed you online for a while; it’s great to finally be able to meet in person.

Corey: Uh-oh. It’s disturbing having realized some of the things I’ve said on Twitter and realizing I’m now within range to get punched in the face. But, you know, we take what we can get. Thank you so much for taking the time to speak with me. I appreciate it.

Sahir: My pleasure.

Corey: Sahir Azam, Chief Product Officer at MongoDB. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry comment telling me that is not the reason that AWS is building many new databases. Tell me which one you’re building and why it solves a problem other than getting you the promotion you probably don’t deserve.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Sam

A 25-year veteran of the Silicon Valley and Seattle technology scenes, Sam Ramji led Kubernetes and DevOps product management for Google Cloud, founded the Cloud Foundry foundation, has helped build two multi-billion dollar markets (API Management at Apigee and Enterprise Service Bus at BEA Systems) and redefined Microsoft’s open source and Linux strategy from “extinguish” to “embrace”.

He is nerdy about open source, platform economics, middleware, and cloud computing with emphasis on developer experience and enterprise software. He is an advisor to multiple companies including Dell Technologies, Accenture, Observable, Fletch, Orbit, OSS Capital, and the Linux Foundation.

Sam received his B.S. in Cognitive Science from UC San Diego, the home of transdisciplinary innovation, in 1994 and is still excited about artificial intelligence, neuroscience, and cognitive psychology.

Links:

  • DataStax: https://www.datastax.com
  • Sam Ramji Twitter: https://twitter.com/sramji
  • Open||Source||Data: https://www.datastax.com/resources/podcast/open-source-data
  • Screaming in the Cloud Episode 243 with Craig McLuckie: https://www.lastweekinaws.com/podcast/screaming-in-the-cloud/innovating-in-the-cloud-with-craig-mcluckie/
  • Screaming in the Cloud Episode 261 with Jason Warner: https://www.lastweekinaws.com/podcast/screaming-in-the-cloud/what-github-can-give-to-microsoft-with-jason-warner/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Redis, the company behind the incredibly popular open source database that is not the bind DNS server. If you’re tired of managing open source Redis on your own, or you’re using one of the vanilla cloud caching services, these folks have you covered with the go to manage Redis service for global caching and primary database capabilities; Redis Enterprise. Set up a meeting with a Redis expert during re:Invent, and you’ll not only learn how you can become a Redis hero, but also have a chance to win some fun and exciting prizes. To learn more and deploy not only a cache but a single operational data platform for one Redis experience, visit redis.com/hero. Thats r-e-d-i-s.com/hero. And my thanks to my friends at Redis for sponsoring my ridiculous non-sense.

Corey: Are you building cloud applications with a distributed team? Check out Teleport, an open source identity-aware access proxy for cloud resources. Teleport provides secure access to anything running somewhere behind NAT: SSH servers, Kubernetes clusters, internal web apps and databases. Teleport gives engineers superpowers!

  • Get access to everything via single sign-on with multi-factor.
  • List and see all SSH servers, kubernetes clusters or databases available to you.
  • Get instant access to them all using tools you already have.

Teleport ensures best security practices like role-based access, preventing data exfiltration, providing visibility and ensuring compliance. And best of all, Teleport is open source and a pleasure to use.

Download Teleport at https://goteleport.com. That’s goteleport.com.

Corey: Welcome to Screaming in the Cloud, I’m Cloud Economist Corey Quinn, and recurring effort that this show goes to is to showcase people in their best light. Today’s guest has done an awful lot: he led Kubernetes and DevOps Product Management for Google Cloud; he founded the Cloud Foundry Foundation; he set open-source strategy for Microsoft in the naughts; he advises companies including Dell, Accenture, the Linux Foundation; and tying all of that together, it’s hard to present a lot of that in a great light because given my own proclivities, that sounds an awful lot like a personal attack. Sam Ramji is the Chief Strategy Officer at DataStax. Sam, thank you for joining me, and it’s weird when your resume starts to read like, “Oh, I hate all of these things.”

Sam: [laugh]. It’s weird, but it’s true. And it’s the only life I could have lived apparently because here I am. Corey, it’s a thrill to meet you. I've been an admirer of your public speaking, and public tweeting, and your writing for a long time.

Corey: Well, thank you. The hard part is getting over the voice saying don’t do it because it turns out that there’s no real other side of public shutting up, which is something that I was never good at anyway, so I figured I’d lean into it. And again, I mean, that the sense of where you have been historically in terms of your career not, “Look what you’ve done,” which is a subtext that I could be accused of throwing in sometimes.

Sam: I used to hear that a lot from my parents, actually.

Corey: Oh, yeah. That was my name growing up. But you’ve done a lot of things, and you’ve transitioned from notable company making significant impact on the industry, to the next one, to the next one. And you’ve been in high-flying roles, doing lots of really interesting stuff. What’s the common thread between all those things?

Sam: I’m an intensely curious person, and the thing that I’m most curious about is distributed cognition. And that might not be obvious from what you see is kind of the… Lego blocks of my career, but I studied cognitive science in college when that was not really something that was super well known. So, I graduated from UC San Diego in ’94 doing neuroscience, artificial intelligence, and psychology. And because I just couldn’t stop thinking about thinking; I was just fascinated with how it worked.

So, then I wanted to build software systems that would help people learn. And then I wanted to build distributed software systems. And then I wanted to learn how to work with people who were thinking about building the distributed software systems. So, you end up kind of going up this curve of, like, complexity about how do we think? How do we think alone? How do we learn to think? How do we think together?

And that’s the directed path through my software engineering career, into management, into middleware at BEA, into open-source at Microsoft because that’s an amazing demonstration of distributed cognition, how, you know, at the time in 2007, I think, Sourceforge had 100,000 open-source projects, which was, like, mind boggling. Some of them even worked together, but all of them represented these groups of people, flung around the world, collaborating on something that was just fundamentally useful, that they were curious about. Kind of did the same thing into APIs because APIs are an even better way to reuse for some cases than having the source code—at Apigee. And kept growing up through that into, how are we building larger-scale thinking systems like Cloud Foundry, which took me into Google and Kubernetes, and then some applications of that in Autodesk and now DataStax. So, I love building companies. I love helping people build companies because I think business is distributed cognition. So, those businesses that build distributed systems, for me, are the most fascinating.

Corey: You were basically handed a heck of a challenge as far as, “Well, help set open-source strategy,” back at Microsoft, in the days where that was a punchline. And credit where due, I have to look at Microsoft of today, and it’s not a joke, you can have your arguments about them, but again in those days, a lot of us built our entire personality on hating Microsoft. Some folks never quite evolved beyond that, but it’s a new ballgame and it’s very clear that the Microsoft of yesteryear and the Microsoft of today are not completely congruent. What was it like at that point understanding that as you’re working with open-source communities, you’re doing that from a place of employment with a company that was widely reviled in the space.

Sam: It was not lost on me. The irony, of course, was that—

Corey: Well, thank God because otherwise the question where you would have been, “What do you mean they didn’t like us?”

Sam: [laugh].

Corey: Which, on some levels, like, yeah, that’s about the level of awareness I would have expected in that era, but contrary to popular opinion, execs at these companies are not generally oblivious.

Sam: Yeah, well, if I’d been clever as a creative humorist, I would have given you that answer instead of my serious answer, but for some reason, my role in life is always to be the straight guy. I used to have Slashdot as my homepage, right? I love when I’d see some conspiracy theory about, you know, Bill Gates dressed up as the Borg, taking over the world. My first startup, actually in ’97, was crushed by Microsoft. They copied our product, copied the marketing, and bundled it into Office, so I had lots of reasons to dislike Microsoft.

But in 2004, I was recruited into their venture capital team, which I couldn’t believe. It was really a place that they were like, “Hey, we could do better at helping startups succeed, so we’re going to evangelize their success—if they’re building with Microsoft technologies—to VCs, to enterprises, we’ll help you get your first big enterprise deal.” I was like, “Man, if I had this a few years ago, I might not be working.” So, let’s go try to pay it forward.

I ended up in open-source by accident. I started going to these conferences on Software as a Service. This is back in 2005 when people were just starting to light up, like, Silicon Valley Forum with, you know, the CEO of Demandware would talk, right? We’d hear all these different ways of building a new business, and they all kept talking about their tech stack was Linux, Apache, MySQL, and PHP. I went to one eight-hour conference, and Microsoft technologies were mentioned for about 12 seconds in two separate chunks. So, six seconds, he was like, “Oh, and also we really like Microsoft SQL Server for our data layer.”

Corey: Oh, Microsoft SQL Server was fantastic. And I know that’s a weird thing for people to hear me say, just because I’ve been renowned recently for using Route 53 as the primary data store for everything that I can. But there was nothing quite like that as far as having multiple write nodes, being able to handle sharding effectively. It was expensive, and you would take a bath on the price come audit time, but people were not rolling it out unaware of those things. This was a trade off that they were making.

Oracle has a similar story with databases. It’s yeah, people love to talk smack about Oracle and its business practices for a variety of excellent reasons, at least in the database space that hasn’t quite made it to cloud yet—knock on wood—but people weren’t deploying it because they thought Oracle was warm and cuddly as a vendor; they did it because they can tolerate the rest of it because their stuff works.

Sam: That’s so well said, and people don’t give them the credit that’s due. Like, when they built hypergrowth in their business, like… they had a great product; it really worked. They made it expensive, and they made a lot of money on it, and I think that was why you saw MySQL so successful and why, if you were looking for a spec that worked, that you could talk through through an open driver like ODBC or JDBC or whatever, you could swap to Microsoft SQL Server. But I walked out of that and came back to the VC team and said, “Microsoft has a huge problem. This is a massive market wave that’s coming. We’re not doing anything in it. They use a little bit of SQL Server, but there’s nothing else in your tech stack that they want, or like, or can afford because they don’t know if their businesses are going to succeed or not. And they’re going to go out of business trying to figure out how much licensing costs they would pay to you in order to consider using your software. They can’t even start there. They have to start with open-source. So, if you’re going to deal with SaaS, you’re going to have to have open-source, and get it right.”

So, I worked with some folks in the industry, wrote a ten-page paper, sent it up to Bill Gates for Think Week. Didn’t hear much back. Bought a new strategy to the head of developer platform evangelism, Sanjay Parthasarathy who suggested that the idea of discounting software to zero for startups, with the hope that they would end up doing really well with it in the future as a Software as a Service company; it was dead on arrival. Dumb idea; bring it back; that actually became BizSpark, the most popular program in Microsoft partner history.

And then about three months later, I got a call from this guy, Bill Hilf. And he said, “Hey, this is Bill Hilf. I do open-source at Microsoft. I work with Bill Gates. He sent me your paper. I really like it. Would you consider coming up and having conversation with me because I want you to think about running open-source technology strategy for the company.” And at this time I’m, like, 33 or 34. And I’m like, “Who me? You’ve got to be joking.” And he goes, “Oh, and also, you’ll be responsible for doing quarterly deep technical briefings with Bill… Gates.” I was like, “You must be kidding.” And so of course I had to check it out. One thing led to another and all of a sudden, with not a lot of history in the open-source community but coming in it with a strategist’s eye and with a technologist's eye, saying, “This is a problem we got to solve. How do we get after this pragmatically?” And the rest is history, as they say.

Corey: I have to say that you are the Chief Strategy Officer at DataStax, and I pull up your website quickly here and a lot of what I tell earlier stage companies is effectively more or less what you have already done. You haven’t named yourself after the open-source project that underlies the bones of what you have built so you’re not going to wind up in the same glorious challenges that, for example, Elastic or MongoDB have in some ways. You have a pricing page that speaks both to the reality of, “It’s two in the morning. I’m trying to get something up and running and I want you the hell out of my way. Just give me something that I can work with a reasonable free tier and don’t make me talk to a salesperson.” But also, your enterprise tier is, “Click here to talk to a human being,” which is speaking enterprise slash procurement slash, oh, there will be contract negotiation on these things.

It’s being able to serve different ends of your market depending upon who it is that encounters you without being off-putting to any of those. And it’s deceptively challenging for companies to pull off or get right. So clearly, you’ve learned lessons by doing this. That was the big problem with Microsoft for the longest time. It’s, if I want to use some Microsoft stuff, once you were able to download things from the internet, it changed slightly, but even then it was one of those, “What exactly am I committing to here as far as signing up for this? And am I giving them audit rights into my environment? Is the BSA about to come out of nowhere and hit me with a surprise audit and find out that various folks throughout the company have installed this somewhere and now I owe more than the company’s worth?” That was always the haunting fear that companies had back then.

These days, I like the approach that companies are taking with the SaaS offering: you pay for usage. On some level, I’d prefer it slightly differently in a pay-per-seat model because at least then you can predict the pricing, but no one is getting surprise submarined with this type of thing on an audit basis, and then they owe damages and payment in arrears and someone has them over a barrel. It’s just, “Oh. The bill this month was higher than we expected.” I like that model I think the industry does, too.

Sam: I think that’s super well said. As I used to joke at BEA Systems, nothing says ‘I love you’ to a customer like an audit, right? That’s kind of a one-time use strategy. If you’re going to go audit licenses to get your revenue in place, you might be inducing some churn there. It’s a huge fix for the structural problem in pricing that I think package software had, right?

When we looked at Microsoft software versus open-source software, and particularly Windows versus Linux, you would have a structure where sales reps were really compensated to sell as much as possible upfront so they could get the best possible commission on what might be used perpetually. But then if you think about it, like, the boxes in a curve, right, if you do that calculus approximation of a smooth curve, a perpetual software license is a huge box and there’s an enormous amount of waste in there. And customers figured out so as soon as you can go to a pay-per-use or pay-as-you-go, you start to smooth that curve, and now what you get is what you deserve, right, as opposed to getting filled with way more cost than you expect. So, I think this model is really super well understood now. Kind of the long run the high point of open-source meets, cloud, meets Software as a Service, you look at what companies like MongoDB, and Confluent, and Elastic, and Databricks are doing. And they’ve really established a very good path through the jungle of how to succeed as a software company. So, it’s still difficult to implement, but there are really world-class guides right now.

Corey: Moving beyond where Microsoft was back in the naughts, you were then hired as a VP over at Google. And in that era, the fact that you were hired as a VP at Google is fascinating. They preferred to grow those internally, generally from engineering. So, first question, when you were being hired as a VP in the product org, did they make you solve algorithms on a whiteboard to get there?

Sam: [laugh]. They did not. I did have somewhat of an advantage [because they 00:13:36] could see me working pretty closely as the CEO of the Cloud Foundry Foundation. I’d worked closely with Craig McLuckie who notably brought Kubernetes to the world along with Joe Beda, and with Eric Brewer, and a number of others.

And he was my champion at Google. He was like, “Look, you know, we need him doing Kubernetes. Let’s bring Sam in to do that.” So, that was helpful. I also wrote a [laugh] 2000-word strategy document, just to get some thoughts out of my head. And I said, “Hey, if you like this, great. If you don’t throw it away.” So, the interviews were actually very much not solving problems in a whiteboard. There were super collaborative, really excellent conversations. It was slow—

Corey: Let’s be clear, Craig McLuckie’s most notable achievement was being a guest on this podcast back in Episode 243. But I’ll say that this is a close second.

Sam: [laugh]. You’re not wrong. And of course now with Heptio and their acquisition by VMware.

Corey: Ehh, they’re making money beyond the wildest dreams of avarice, that’s all well and good, but an invite to this podcast, that’s where it’s at.

Sam: Well, he should really come on again, he can double down and beat everybody. That can be his landmark achievement, a two-timer on
Screaming in [the] Cloud.

Corey: You were at Google; you were at Microsoft. These are the big titans of their era, in some respect—not to imply that there has beens; they’re bigger than ever—but it’s also a more crowded field in some ways. I guess completing the trifecta would be Amazon, but you’ve had the good judgment never to work there, directly of course. Now they’re clearly in your market. You’re at DataStax, which is among other things, built on Apache Cassandra, and they launched their own Cassandra service named Keyspaces because no one really knows why or how they name things.

And of course, looking under the hood at the pricing model, it’s pretty clear that it really is just DynamoDB wearing some Groucho Marx classes with a slight upcharge for API level compatibility. Great. So, I don’t see it a lot in the real world and that’s fine, but I’m curious as to your take on looking at all three of those companies at different eras. There was always the threat in the open-source world that they are going to come in and crush you. You said earlier that Microsoft crushed your first startup.

Google is an interesting competitor in some respects; people don’t really have that concern about them. And your job as a Chief Strategy Officer at Amazon is taken over by a Post-it Note that simply says ‘yes’ on it because there’s nothing they’re not going to do, or try, and experiment with. So, from your perspective, if you look at the titans, who is it that you see as the largest competitive threat these days, if that’s even a thing?

Sam: If you think about Sun Tzu and the Art of War, right—a lot of strategy comes from what we’ve learned from military environments—fighting a symmetric war, right, using the same weapons and the same army against a symmetric opponent, but having 1/100th of the personnel and 1/100th of the money is not a good plan.

Corey: “We’re going to lose money, going to be outcompeted; we’ll make it up in volume. Oh, by the way, we’re also slower than they are.”

Sam: [laugh]. So, you know, trying to come after AWS, or Microsoft, or Google as an independent software company, pound-for-pound, face-to-face, right, full-frontal assault is psychotic. What you have to do, I think, at this point is to understand that these are each companies that are much like we thought about Linux, and you know, Macintosh, and Windows as operating systems. They’re now the operating systems of the planet. So, that creates some economies of scale, some efficiencies for them. And for us. Look at how cheap object storage is now, right? So, there’s never been a better time in human history to create a database company because we can take the storage out of the database and hand it over to Amazon, or Google, or Microsoft to handle it with 13 nines of durability on a constantly falling cost basis.

So, that’s super interesting. So, you have to prosecute the structure of the world as it is, based on where the giants are and where they’ll be in the future. Then you have to turn around and say, like, “What can they never sell?”

So, Amazon can never sell something that is standalone, right? They’re a parts factory and if you buy into the Amazon-first strategy of cloud computing—which we did at Autodesk when I was VP of cloud platform there—everything is a primitive that works inside Amazon, but they’re not going to build things that don’t work outside of the Amazon primitives. So, your company has to be built on the idea that there’s a set of people who value something that is purpose-built for a particular use case that you can start to broaden out, it’s really helpful if they would like it to be something that can help them escape a really valuable asset away from the center of gravity that is a cloud. And that’s why data is super interesting. Nobody wakes up in the morning and says, “Boy, I had such a great conversation with Oracle over the last 20 years beating me up on licensing. Let me go find a cloud vendor and dump all of my data in that so they can beat me up for the next 20 years.” Nobody says that.

Corey: It’s the idea of data portability that drives decision-making, which makes people, of course, feel better about not actually moving in anywhere. But the fact that they’re not locked in strategically, in a way that requires a full software re-architecture and data model rewrite is compelling. I’m a big believer in convincing people to make decisions that look a lot like that.

Sam: Right. And so that’s the key, right? So, when I was at Autodesk, we went from our 100 million dollar, you know, committed spend with 19% discount on the big three services to, like—we started realize when we’re going to burn through that, we were spending $60 million or so a year on 20% annual growth as the cloud part of the business grew. Thought, “Okay, let’s renegotiate. Let’s go and do a $250 million deal. I’m sure they’ll give us a much better discount than 19%.” Short story is they came back and said, “You know, we’re going to take you from an already generous 19% to an outstanding 22%.” We thought, “Wait a minute, we already talked to Intuit. They’re getting a 40% discount on a $400 million spend.”

So, you know, math is hard, but, like, 40% minus 22% is 18% times $250 million is a lot of money. So, we thought, “What is going on here?” And we realized we just had no credible threat of leaving, and Intuit did because they had built a cross-cloud capable architecture. And we had not. So, now stepping back into the kind of the world that we’re living in 2021, if you’re an independent software company, especially if you have the unreasonable advantage of being an open-source software company, you have got to be doing your customers good by giving them cross-cloud capability. It could be simply like the Amdahl coffee cup that Amdahl reps used to put as landmines for the IBM reps, later—I can tell you that story if you want—even if it’s only a way to save money for your customer by using your software, when it gets up to tens and hundreds of million dollars, that’s a really big deal.

But they also know that data is super important, so the option value of being able to move if they have to, that they have to be able to pull that stick, instead of saying, “Nice doggy,” we have to be on their side, right? So, there’s almost a detente that we have to create now, as cloud vendors, working in a world that’s invented and operated by the giants.

Corey: This episode is sponsored by our friends at Oracle HeatWave is a new high-performance accelerator for the Oracle MySQL Database Service. Although I insist on calling it “my squirrel.” While MySQL has long been the worlds most popular open source database, shifting from transacting to analytics required way too much overhead and, ya know, work. With HeatWave you can run your OLTP and OLAP, don’t ask me to ever say those acronyms again, workloads directly from your MySQL database and eliminate the time consuming data movement and integration work, while also performing 1100X faster than Amazon Aurora, and 2.5X faster than Amazon Redshift, at a third of the cost. My thanks again to Oracle Cloud for sponsoring this ridiculous nonsense.

Corey: When we look across the, I guess, the ecosystem as it’s currently unfolding, a recurring challenge that I have to the existing incumbent cloud providers is they’re great at offering the bricks that you can use to build things, but if I’m starting a company today, I’m not going to look at building it myself out of, “Ooh, I’m going to take a bunch of EC2 instances, or Lambda functions, or popsicles and string and turn it into this thing.” I’m going to want to tie together things that are way higher level. In my own case, now I wind up paying for Retool, which is, effectively, yeah, it runs on some containers somewhere, presumably, I think in Azure, but don’t quote me on that. And that’s great. Could I build my own thing like that?

Absolutely not. I would rather pay someone to tie it together. Same story. Instead of building my own CRM by running some open-source software on an EC2 instance, I wind up paying for Salesforce or Pipedrive or something in that space. And so on, and so forth.

And a lot of these companies that I’m doing business with aren’t themselves running on top of AWS. But for web hosting, for example; if I look at the reference architecture for a WordPress site, AWS’s diagram looks like a punchline. It is incredibly overcomplicated. And I say this as someone who ran large WordPress installations at Media Temple many years ago. Now, I have the good sense to pay WP Engine. And on a monthly basis, I give them money and they make the website work.

Sure, under the hood, it’s running on top of GCP or AWS somewhere. But I don’t have to think about it; I don’t have to build this stuff together and think about the backups and the failover strategy and the rest. The website just works. And that is increasingly the direction that business is going; things commoditize over time. And AWS in particular has done a terrible job, in my experience, of differentiating what it is they’re doing in the language that their customers speak.

They’re great at selling things to existing infrastructure engineers, but folks who are building something from scratch aren’t usually in that cohort. It’s a longer story with time and, “Well, we’re great at being able to sell EC2 instances by the gallon.” Great. Are you capable of going to a small doctor’s office somewhere in the American Midwest and offering them an end-to-end solution for managing patient data? Of course not. You can offer them a bunch of things they can tie together to something that will suffice if they all happen to be software engineers, but that’s not the opportunity.

So instead, other companies are building those solutions on top of AWS, capturing the margin. And if there’s one thing guaranteed to keep Amazon execs awake at night, it’s the idea of someone who isn’t them making money somehow somewhere, so I know that’s got to rankle them, but they do not speak that language. At all. Longer-term, I only see that as a more and more significant crutch. A long enough timeframe here, we’re talking about them becoming the Centurylinks of the world, the tier one backbone provider that everyone uses, but no one really thinks about because they’re not a household name.

Sam: That is a really thoughtful perspective. I think the diseconomies of scale that you’re pointing to start to creep in, right? Because when you have to sell compute units by the gallon, right, you can’t care if it’s a gallon of milk, [laugh] or a gallon of oil, or you know, a gallon of poison. You just have to keep moving it through. So, the shift that I think they’re going to end up having to make pragmatically, and you start to see some signs of it, like, you know, they hired but could not retain Matt [Acey 00:23:48]. He did an amazing job of bringing them to some pragmatic realization that they need to partner with open-source, but more broadly, when I think about Microsoft in the 2000s as they were starting to learn their open-source lessons, we were also being able to pull on Microsoft’s deep competency and partners. So, most people didn’t do the math on this. I was part of the field governance council so I understood exactly how the Microsoft business worked to the level that I was capable. When they had $65 billion in revenue, they produced $24 billion in profit through an ecosystem that generated $450 billion in revenue. So, for every dollar Microsoft made, it was $8 to partners. It was a fundamentally platform-shaped business, and that was how they’re able to get into doctors offices in the Midwest, and kind of fit the curve that you’re describing of all of those longtail opportunities that require so much care and that are complex to prosecute. These solved for their diseconomies of scale by having 1.2 million partner companies. So, will Amazon figure that out and will they hire, right, enough people who’ve done this before from Microsoft to become world-class in partnering, that’s kind of an exercise left to the [laugh] reader, right? Where will that go over time? But I don’t see another better mathematical model for dealing with the diseconomies of scale you have when you’re one of the very largest providers on the planet.

Corey: The hardest problem as I look at this is, at some point, you hit a point of scale where smaller things look a lot less interesting. I get that all the time when people say, “Oh, you fix AWS bills, aren’t you missing out by not targeting Google bills and Azure bills as well?” And it’s, yeah. I’m not VC-backed. It turns out that if I limit the customer base that I can effectively service to only AWS customers, yeah turns out, I’m not going to starve anytime soon. Who knew? I don’t need to conquer the world and that feels increasingly antiquated, at least going by the stories everyone loves to tell.

Sam: Yeah, it’s interesting to see how cloud makes strange bedfellows, right? We started seeing this in, like, 2014, 2015, weird partnerships that you’re like, “There’s no way this would happen.” But the cloud economics which go back to utilization, rather than what it used to be, which was software lock-in, just changed who people were willing to hang out with. And now you see companies like Databricks going, you know, we do an amazing amount of business, effectively competing with Amazon, selling Spark services on top of predominantly Amazon infrastructure, and everybody seems happy with it. So, there’s some hint of a new sensibility of what the future of partnering will be. We used to call it coopetition a long time ago, which is kind of a terrible word, but at least it shows that there’s some nuance in you can’t compete with everybody because it’s just too hard.

Corey: I wish there were better ways of articulating these things because it seems from the all the outside world, you have companies like Amazon and Microsoft and Google who go and build out partner networks because they need that external accessibility into various customer profiles that they can’t speak to super well themselves, but they’re also coming out with things that wind up competing directly or indirectly,
with all of those partners at the same time. And I don’t get it. I wish that there were smarter ways to do it.

Sam: It is hard to even talk about it, right? One of the things that I think we’ve learned from philosophy is if we don’t have a word for it, we can’t be intelligent about it. So, there’s a missing semantics here for being able to describe the complexity of where are you partnering? Where are you competing? Where are you differentiating? In an ecosystem, which is moving and changing.

I tend to look at the tools of game theory for this, which is to look at things as either, you know, nonzero-sum games or zero-sum games. And if it’s a nonzero-sum game, which I think are the most interesting ones, can you make it a positive sum game? And who can you play positive-sum games with? An organization as big as Amazon, or as big as Microsoft, or even as big as Google isn’t ever completely coherent with itself. So, thinking about this as an independent software company, it doesn’t matter if part of one of these hyperscalers has a part of their business that competes with your entire business because your business probably drives utilization of a completely different resource in their company that you can partner within them against them, effectively. Right?

For example, Cassandra is an amazingly powerful but demanding workload on Kubernetes. So, there’s a lot of Cassandra on EKS. You grow a lot of workload, and EKS business does super well. Does that prevent us from working with Amazon because they have Dynamo or because they have Keyspaces? Absolutely not, right?

So, this is when those companies get so big that they are almost their own forest, right, of complexity, you can kind of get in, hang out, do well, and pretty much never see the competitive product, unless you’re explicitly looking for it, which I think is a huge danger for us as independent software companies. And I would say this to anybody doing strategy for an organization like this, which is, don’t obsess over the tiny part of their business that competes with yours, and do not pay attention to any of the marketing that they put out that looks competitive with what you have. Because if you can’t figure out how to make a better product and sell it better to your customers as a single purpose corporation, you have bigger problems.

Corey: I want to change gears slightly to something that’s probably a fair bit more insulting, but that’s okay. We’re going to roll with it. That seems to be the theme of this episode. You have been, in effect, a CIO a number of times at different companies. And if we take a look at the typical CIO tenure, industry-wide, it’s not long; it approaches the territory from an executive perspective of, “Be sure not to buy green bananas. You might not be here by the time they ripen.” And I’m wondering what it is that drives that and how you make a mark in a relatively short time frame when you’re providing inputs and deciding on strategy, and those decisions may not bear fruit for years.

Sam: CIO used to—we used say it stood for ‘Career Is Over’ because the tenure is so short. I think there’s a couple of reasons why it’s so short. And I think there’s a way I believe you can have impact in a short amount of time. I think the reason that it’s been short is because people aren’t sure what they want the CIO role to be.

Do they want it to be a glorified finance person who’s got a lot of data processing experience, but now really has got, you know, maybe even an MBA in finance, but is not focusing on value creation? Do they want it to be somebody who’s all-singing, all-dancing Chief Data Officer with a CTO background who did something amazing and solved a really hard problem? The definition of success is difficult. Often CIOs now also have security under them, which is literally a job I would never ever want to have. Do security for a public corporation? Good Lord, that’s a way to lose most of your life. You’re the only executive other than the CEO that the board wants to hear from. Every sing—

Corey: You don’t sleep; you wait, in those scenarios. And oh, yeah, people joke about ablative CSOs in those scenarios. Yeah, after SolarWinds, you try and get an ablative intern instead, but those don’t work as well. It’s a matter of waiting for an inevitability. One of the things I think is misunderstood about management broadly, is that you are delegating work, but not the responsibility. The responsibility rests
with you.

So, when companies have these statements blaming some third-party contractor, it’s no, no, no. I’m dealing with you. You were the one that gave my data to some sketchy randos. It is your responsibility that data has now been compromised. And people don’t want to hear that, but it’s true.

Sam: I think that’s absolutely right. So, you have this high risk, medium reward, very fungible job definition, right? If you ask all of the CIO’s peers what their job is, they’ll probably all tell you something different that represents their wish list. The thing that I learned at Autodesk, I was only there for 15 months, but we established a fundamental transformation of the work of how cloud platform is done at the company that’s still in place a couple years later.

You have to realize that you’re a change agent, right? You’re actually being hired to bring in the bulk of all the different biases and experiences you have to solve a problem that is not working, right? So, when I got to Autodesk, they didn’t even know what their uptime was. It took three months to teach the team how to measure the uptime. Turned out the uptime was 97.7% for the cloud, for the world’s largest engineering software company.

That is 200 hours a year of unplanned downtime, right? That is not good. So, a complete overhaul [laugh] was needed. Understanding that as a change agent, your half-life is 12 to 18 months, you have to measure success not on tenure, but on your ability to take good care of the patient, right? It’s going to be a lot of pain, you’re going to work super hard, you’re going to have to build trust with everyone, and then people are still going to hate you at the end. That is something you just have to kind of take on.

As a friend of mine, Jason Warner joined Redpoint Ventures recently, he said this when he was the CTO of GitHub: “No one is a villain in their own story.” So, you realize, going into a big organization, people are going to make you a villain, but you still have to do incredibly thoughtful, careful work, that’s going to take care of them for a long time to come. And those are the kinds of CIOs that I can relate to very well.

Corey: Jason is great. You’re name-dropping all the guests we’ve had. My God, keep going. It’s a hard thing to rationalize and wrap heads around. It’s one of those areas where you will not be measured during your tenure in the role, in some respects. And, of course, that leads to the cynical perspective as well, where well, someone’s not going to be here long and if they say, “Yeah, we’re just going to keep being stewards of the change that’s already underway,” well, that doesn’t look great, so quick, time to do a cloud migration, or a cloud repatriation, or time to roll something else out. A bit of a different story.

Sam: One of the biggest challenges is how do you get the hearts and the minds of the people who are in the organization when they are no fools, and their expectation is like, “Hey, this company’s been around for decades, and we go through cloud leaders or CIOs, like Wendy’s goes through hamburgers.” They could just cloud-wash, right, or change-wash all their language. They could use the new language to describe the old thing because all they have to do is get through the performance review and outwait you. So, there’s always going to be a level of
defection because it’s hard to change; it’s hard to think about new things.

So, the most important thing is how do you get into people’s hearts and minds and enable them to believe that the best thing they could do for their career is to come along with the change? And I think that was what we ended up getting right in the Autodesk cloud transformation. And that requires endless optimism, and there’s no room for cynicism because the cynicism is going to creep in around the edges. So, what I found on the job is, you just have to get up every morning and believe everything is possible and transmit that belief to everybody.

So, if it seems naive or ingenuous, I think that doesn’t matter as long as you can move people’s hearts in each conversation towards, like, “Oh, this person cares about me. They care about a good outcome from me. I should listen a little bit more and maybe make a 1% change in what I’m doing.” Because 1% compounded daily for a year, you can actually get something done in the lifetime of a CIO.

Corey: And I think that’s probably a great place to leave it. If people want to learn more about what you’re up to, how you think about these things, how you view the world, where can they find you?

Sam: You can find me on Twitter, I’m @sramji, S-R-A-M-J-I, and I have a podcast that I host called Open||Source||Datawhere I invite innovators, data nerds, computational networking nerds to hang out and explain to me, a software programmer, what is the big world of open-source data all about, what’s happening with machine learning, and what would it be like if you could put data in a container, just like you could put code in a container, and how might the world change? So, that’s Open||Source||Data podcast.

Corey: And we’ll of course include links to that in the [show notes 00:35:58]. Thanks so much for your time. I appreciate it.

Sam: Corey, it’s been a privilege. Thank you so much for having me.

Corey: Likewise. Sam Ramji, Chief Strategy Officer at DataStax. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with a comment telling me exactly which item in Sam’s background that I made fun of is the place that you work at.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and
we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Aparna

Aparna Sinha is Director of Product for Kubernetes and Anthos at Google Cloud. Her teams are focused on transforming the way we work through innovation in platforms. Before Anthos and Kubernetes, Aparna worked on the Android platform. She joined Google from NetApp where she was Director of Product for storage automation and private cloud. Prior to NetApp, Aparna was a leader in McKinsey and Company’s business transformation office working with CXOs on IT strategy, pricing, and M&A. Aparna holds a PhD in Electrical Engineering from Stanford and has authored several technical publications. She serves on the Governing Board of the Cloud Native Computing Foundation (CNCF).

Links:

  • DevOps Research Report: https://www.devops-research.com/research.html
  • Twitter: https://twitter.com/apbhatnagar

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Redis, the company behind the incredibly popular open source database that is not the bind DNS server. If you’re tired of managing open source Redis on your own, or you’re using one of the vanilla cloud caching services, these folks have you covered with the go to manage Redis service for global caching and primary database capabilities; Redis Enterprise. Set up a meeting with a Redis expert during re:Invent, and you’ll not only learn how you can become a Redis hero, but also have a chance to win some fun and exciting prizes. To learn more and deploy not only a cache but a single operational data platform for one Redis experience, visit redis.com/hero. Thats r-e-d-i-s.com/hero. And my thanks to my friends at Redis for sponsoring my ridiculous non-sense.

Corey: You know how Git works right?

Announcer: Sorta, kinda, not really. Please ask someone else.

Corey: That's all of us. Git is how we build things, and Netlify is one of the best ways I’ve found to build those things quickly for the web. Netlify’s Git-based workflows mean you don’t have to play slap-and-tickle with integrating arcane nonsense and web hooks, which are themselves about as well understood as Git. Give them a try and see what folks ranging from my fake Twitter for Pets startup, to global Fortune 2000 companies are raving about. If you end up talking to them—because you don’t have to; they get why self-service is important—but if you do, be sure to tell them that I sent you and watch all of the blood drain from their faces instantly. You can find them in the AWS marketplace or at www.netlify.com. N-E-T-L-I-F-Y dot com.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. We have a bunch of conversations on this show covering a wide gamut of different topics, things that I find personally interesting, usually, and also things I’m noticing in the industry. Fresh on the heels of Google Next, we get to ideally have conversations about both of those things. Today, I’m speaking with the Director of Product Management at Google Cloud, Aparna Sinha. Aparna, thank you so much for joining me today. I appreciate it.

Aparna: Thank you, Corey. It’s a pleasure to be here.

Corey: So, Director of Product Management is one of those interesting titles. We’ve had a repeat guest here, Director of Outbound Product Management Richard Seroter, which is great. I assume—as I told him—outbound products are the ones that are about to be discontinued. He’s been there a year and somehow has failed the discontinue a single thing, so okay, I’m sure that’s going to show up on his review. What do you do? The products aren’t outbound; they’re just products, and you’re managing them, but that doesn’t tell me much. Titles are always strange.

Aparna: Yeah, sure. Richard is one of my favorite people, by the way. I work closely with him. I am the Director of Product for Developer Platform. That’s Google Cloud’s developer platform.

It includes many different products—actually, 30-Plus products—but the primary pieces are usually when a developer comes to Google Cloud, the pieces that they interact with, like our command-line interface, like our Cloud Shell, and all of the SDK pieces that go behind it, and then also our DevOps tooling. So, as you’re writing the application in the IDE and as you’re deploying it into production, that’s all part of the developer platform. And then I also run our serverless platform, which is one of the most developer-friendly capabilities from a compute perspective. It’s also integrated into many different services within GCP. So, behind the title, that’s really what I work on.

Corey: Okay, so you’re, I guess, in part responsible for well, I guess, a disappointment of mine a few years ago. I have a habit on Twitter—because I’m a terrible person—of periodically spinning up a new account on various cloud providers and kicking the tires and then live-tweeting the experience, and I was really set to dunk on Google Cloud; I turned this into a whole blog post. And I came away impressed, where the developer experience was pretty close to seamless for getting up and running. It was head and shoulders above what I’ve seen from other cloud providers, and on the one hand, I want to congratulate you and on the other, it doesn’t seem like that’s that high of a bar, to be perfectly honest with you because it seems that companies get stuck in their own ways and presuppose that everyone using the product is the same as the people building the product. Google Cloud has been and remains a shining example of great developer experience across the board.

If I were starting something net new and did not have deep experience with an existing cloud provider—which let’s face it, the most valuable thing about the cloud is knowing how it’s going to break because everything breaks—I would be hard-pressed to not pick GCP, if not as the choice, at least a strong number two. So, how did that come to be? I take a look at a lot of Google’s consumer apps and, “This is a great user experience,” isn’t really something I find myself saying all that often. Google Cloud is sort of its own universe. What happened?

Aparna: Well, thank you, first of all, for the praise. We are very humble about it, actually. I think that we’re grateful if our developers find the experience to be seamless. It is something that we measure all the time. That may be one of the reasons why you found it to be better than other places. We are continuously trying to improve the time to value for developers, how long it takes them to perform certain actions. And so what you measure is what you improve, right? If you don’t measure it, you don’t improve it. That’s one of our SRE principles.

Corey: I wish. I’ve been measuring certain things for years, and they don’t seem to be improving at all. It’s like, “Wow, my code is still terrible, but I’m counting the bugs and the number isn’t getting smaller.” Turns out there might be additional steps required.

Aparna: Yes, you know, we measure it, we look at it, we take active OKRs to improve these things, especially usability. Usability is extremely important for certainly the developer platform, for my group; that’s something that’s extremely important. I would say, stepping back, you said it’s not that common to find a good user experience in the cloud, I think in general—you know, and I’ve spent the majority of my career, if not all of my career, working on enterprise software. Enterprise software is not always designed in the most user-friendly way; it’s not something that people always think about. Some of the enterprise software I’ve used has been really pretty… pretty bad. Just a list of things.

Corey: Oh, yeah. And it seems like their entire philosophy—I did a bit of a dive into this, and I think it was Stripe’s Patrick McKenzie who wound up pointing this out originally, though; but the internet is big and people always share and reshare ideas—the actual customer for enterprise software is very often procurement or a business unit that is very organizationally distant from the person who’s using it. And I think in a world of a cloud platform, that is no longer true. Yeah, there’s a strategic decision of what Cloud do we use, but let’s be serious, that decision often comes into play long after there’s already been a shadow IT slash groundswell uprising. The sales process starts to look an awful lot less like, “Pick our cloud,” and a lot more like, “You’ve already picked our cloud. How about we formalize the relationship?”

And developer experience with platforms is incredibly important and I’m glad to see that this is a—well, it’s bittersweet to me. I am glad to see that this is something that Google is focusing on, and I’m disappointed to admit that it’s a differentiator.

Aparna: It is a differentiator. It is extremely important. At Google, there are a couple of reasons why this is part of our DNA, and it is actually related to the fact that we are also a consumer products company. We have a very strong user experience team, a very strong measurements-oriented—they measure everything, and they design everything, and they run focus groups. So, we have an extraordinary usability team, and it’s actually one of the groups that—just like every other group—is fungible; you can move between consumer and cloud. There’s no difference in terms of your training and skill set.

And so, I know you said that you’re not super impressed with our consumer products, but I think that the practice behind treating the user as king, treating the user as the most important part of your development, is something that we bring over into cloud. And it’s just a part of how we do development, and I think that’s part of the reason why our products are usable. Again, I shy away from taking any really high credit on these things because I think I always have a very high bar. I want them to be delightful, super delightful, but we do have good usability scores on some of the pieces. I think our command line, I think, is quite good. I think—there’s always improvements, by the way, Corey—but I think that there are certain things that are delightful.

And a lot of thought goes into it and a lot of multi-functional—meaning across product—user experience and engineering. We have end-developer relations. We have, sort of this four-way communication about—you know, with friction logs and with lots of trials and lots of discussion and measurements, is how we improve the user experience. And I would love to see that in more enterprise software. I think that my experience in the industry is that the user is becoming more important, generally, even in enterprise software, probably because of the migration to cloud.

You can’t ignore the user anymore. This shouldn’t be all about procurement. Anybody can procure a cloud service. It’s really about how easily and how quickly can they get to what they want to do as a user, which I think also the definition of what a developer is changing and I think that’s one of the most exciting things about our work is that the developer can be anybody; it can be my kids, and it can be anyone across the world. And our goal is to reach those people and to make it easy for them.

Corey: If I had to bet on a company not understanding that distinction, on some level, Google’s reputation lends itself to that where, oh, great. It’s like, I’m a little old to go back to school and join a fraternity and be hazed there, so the second option was, oh, I’ll get an interview to be an SRE at Google where, “Oh, great, you’ve done interesting things, but can you invert a binary tree on a whiteboard?” “No, I cannot. Let’s save time and admit that.” So, the concern that I would have had—you just directly contradicted—was the idea that you see at some companies where there’s the expectation that all developers are like their developers.

Google, for better or worse, has a high technical bar for hiring. A number of companies do not have a similar bar along similar axes, and they’re looking for different skill sets to achieve different outcomes, and that’s fine. To be clear, I am not saying that, oh, the engineers at Google are
all excellent and the engineers all at a bank are all crap. Far from it.

That is not true in either direction, but there are differences as far as how they concern themselves with software development, how they frame a lot of these things. And I am surprised that Google is not automatically assuming that developers are the type of developers that you have at Google. Where did that mindset shift come from?

Aparna: Oh, absolutely not. I think we would be in trouble if we did that. I studied electrical engineering in school. This would be like assuming that the top of the class is kind of like the kind of people that we want to reach, and it’s just absolutely not. Like I said, I want to reach total beginners, I want to reach people who are non-developers with our developer platform.

That’s our explicit goal, and so we view developers as individuals with a range of superpowers that they’ve gained throughout their lives, professionally and personally, and people who are always on a path to learn new things, and we want to make it easy for them. We don’t treat them as bodies in an employment relationship with some organization, or people with certain minimum bar degrees, or whatever it is. As far as interviewing goes, Corey, in product management, which is the practice that I’m part of, we actually look for, in the interview, that the candidate is not thinking about themselves; they’re not imposing themselves on the user base.

So, can you think outside of yourself? Can you think of the user base? And are you inquisitive? Are you curious? Do you observe? And how well do you observe differences and diversity, and how well are you able to grasp what might be needed by a particular segment? How well are you able to segment the user base?

That’s what we look for, certainly in product management, and I’m quite sure also in user experience. You’re right, on engineering, of course, we’re looking for technical skills, and so on, but that’s not how we design our products, that’s not how we design the usability of our products.

Corey: “If you people were just a little bit smarter slash more like me, then this would work a lot better,” is a common trope. Which brings us, of course, to the current state of serverless. I tend to view serverless as largely a failed initiative so far. And to be clear, I’m viewing this from an AWS-centric lens; that is the… we’ll be charitable and call it pool in which I swim. And they announced Lambda in 2015; that’s great. “The only code you will ever write in the future is business logic.” Yeah, I might have heard that one before about 15 other technologies dating back to the 60s, but okay.

And the expectation was that it was going to take off and set the world on fire. You just needed to learn the constraints of how this worked. And there were a bunch of them, and they were obnoxious, and it didn’t have a learning curve so much as a learning cliff. And nowadays, we do see it everywhere, but it’s also in small doses. It’s mostly used as digital spackle to plaster over the gaps between various AWS services.

What I’m not seeing across the board is a radical mindset shift in the way that developers are engaging with cloud platforms that would be heralded by widespread adoption of serverless principles. That said, we are on the heels here of Google Cloud Next, and that you had a bunch of serverless announcements, I’m going to go out on a limb and guess you might not agree with my dismal take on the serverless side of the world?

Aparna: Well, I think this is a great question because despite the fact that I like not to be wishy-washy about anything, I actually both agree and disagree [laugh] with what you said. And that’s funny.

Corey: Well, that’s why we’re talking about this here instead of on Twitter where two contradictory things can’t possibly both be true. Wow, imagine that; nuance, it doesn’t fit 280 characters. Please, continue.

Aparna: So, what I agree with is that—I agree with you that the former definition of serverless and the constrained way that we are conditioned thinking about serverless is not as expansive as originally hoped, from an adoption perspective. And I think that at Google, serverless is just no longer about only event-driven programming or microservices; it’s about running complex workloads at scale while still preserving the delightful developer experience. And this is where the connection to the developer experience comes in. Because the developer experience, in my mind, it’s about time to value. How quickly can I achieve the outcome that I need for my business?

And what are the things that get in the way of that? Well, setting up infrastructure gets in the way of that, having to scale infrastructure gets in the way of that, having to debug pieces that aren’t actually related to the outcome that you’re trying to get to gets in the way of that. And the beauty of serverless, it’s all in how you define serverless: what does this name actually mean? If serverless only means functions and event-driven applications, then yes, actually, it has a better developer experience, but it is not expansive, and then it is limited, and it’s trapped in its skin the way that you mentioned it. [laugh].

Corey: And it doesn’t lend itself very well to legacy applications—legacy, of course, being condescending engineering-speak for ‘it makes money.’ But yeah, that’s the stuff that powers the world. We’re not going to be redoing all those things as serverless-powered microservices anytime soon, in most cases.

Aparna: At Google Cloud, we are redefining serverless. And so what we are taking from Serverless is the delightful user experience and the fact that you don’t have to manage the infrastructure, and what we’re putting in the serverless is essentially serverless containers. And this is the big revolution in serverless, is that serverless—at least a Google Cloud with serverless containers and our Cloud Run offering—is able to run much bigger varieties of applications and we are seeing large enterprises running legacy applications, like you say, on Cloud Run, which is serverless from a developer experience perspective. There’s no cluster, there is no server, there’s no VM, there’s nothing for you to set up from a scaling perspective. And it essentially scales infinitely.

And it is very developer-focused; it’s meant for the developer, not for the operator or the infrastructure admin. In reality in enterprise, there is very much a segmentation of roles. And even in smaller companies, there’s a segmentation of roles even within the same person. Like, they may have to do some infrastructure work and they may do some development work. And what serverless—at least in the context of Google Cloud—does, is it removes the infrastructure work and maximizes the development work so that you can focus on your application and you can get to that end result, that business value that you’re trying to achieve.

And with Cloud Run, what we’ve done is we’ve preserved that—and I would say, actually, arguably improved that because we’ve done usability studies that show that we’re 22 points above every other serverless offering from a usability perspective. So, it’s super important to me that anybody can use this service. Anybody. Maybe even not a developer can use this service. And that’s where our focus is.

And then what we’ve done underneath is we’ve removed many of the restrictions that are traditionally associated with serverless. So, it doesn’t have to be event-driven, it is not only a particular set of languages or a particular set of runtimes. It is not only stateless applications, and it’s not only request-based billing, it’s not only short-running jobs. These are the kinds of things that we have removed and I think we’ve just redefined serverless.

Corey: [unintelligible 00:17:05], on some level, the idea of short-lived functions with a maximum cap feels like a lazy answer to one of the hard problems in computer science, the halting problem. For those not familiar, my layman’s understanding of it is, “Okay, you have a program that’s running in a loop. How do you deterministically say that it is done executing?” And the functional answer to that is, “Oh, after 15 minutes, it’s done. We’re killing it.” Which I guess is an answer, but probably not one that’s going to get anyone a PhD.

It becomes very prescriptive and it leads to really weird patterns trying to work around some of those limitations. And historically, yeah, by working within the constraints of the platform, it works super well. What interests me about Cloud Run is that it doesn’t seem to have many of those constraints in quite the same way. It’s, “Can you shove whatever monstrosity you’ve got into a container? You can’t? Well, okay, there are ways to get there.”

Full disclosure, I was very anti-container; the industry has yet again proven to me that I cannot predict the future. Here we are. “Great, can you shove a container in and hand it to some other place to run it where”—spoiler, people will argue with me on this and they are wrong—“Google engineers are better at running infrastructure to run containers than you are.” Full stop. That is the truism of how this works; economies of scale.

I love the idea of being able to take something, throw it over a wall, and not have to think about the rest of it. But everything that I’m thinking about in this context looks certain ways and it’s the type of application that I’m working on or that I’m looking at most recently. What are you seeing in Cloud Run as far as interesting customer use cases? What are people doing with it that you didn’t expect them to?

Aparna: Yeah, I think this is a great time to ask that question because with the pandemic last year—I guess we’re still in the pandemic, but with the pandemic, we had developers all over the world become much more important and much more empowered, just because there wasn’t really much of an operations team, there wasn’t really as much coordination even possible. And so we saw a lot of customers, a lot of developers moving to cloud, and they were looking for the easiest thing that they could use to build their applications. And as a result, serverless and Cloud Run in particular, became extremely popular; I would say hockey stick in terms of usage.

And we’re seeing everything under the sun. ecobee—this is a home automation company that makes smart thermostats—they’re using Cloud Run to launch a new camera product with multi-factor authentication and security built-in, and they had a very tight launch timeline. They were able to very quickly meet that need. Another company—and you talk about, you know, sort of brick and mortar—IKEA, which you and I all like to shop [laugh] at, particularly doing the—

Corey: Oh, I love building something from 500 spare parts, badly. It’s like basically bringing my AWS architecture experience into my living room. It’s great. Please continue.

Aparna: Yeah, it’s like, yeah—

Corey: The Swedish puzzle manufacturer.

Aparna: Yes. They’re a great company, and I think it just in the downturn and the lockdown, it was actually a very dicey time, very tricky time, particularly for retailers. Of course, everybody was refurbishing their home or [laugh], you know, improving their home environment and their furniture. And IKEA started using serverless containers along with serverless analytics—so with BigQuery, and Cloud Run, and Cloud Functions—and one of the things they did is that they were able to cut their inventory refresh rate from more than three hours to less than three minutes. This meant that when you were going to drive up and do some curbside pickup, you know the order that you placed was actually in stock, which was fantastic for CSAT and everything.

But that’s the technical piece that they were able to do. When I spoke with them, the other thing that they were able to do with the Cloud Run and Cloud Functions is that they were able to improve the work-life balance of their engineers, which I thought was maybe the biggest accomplishment. Because the platform, they said, was so easy for them to use and so easy for them to accomplish what they needed to accomplish, that they had a better [laugh] better life. And I think that’s very meaningful.

In other companies, MediaMarktSaturn, we’ve talked about them before; I don’t know if I’ve spoken to you about them, but we’ve certainly talked about them publicly. They’re a retailer in EMEA, and because of their use of Cloud Run, and they were able to combine the speed of serverless with the flexibility of containers, and their development team was able to go eight times faster while handling 145% increase in digital channel traffic. Again, there are a lot more digital channel traffic during COVID. And perhaps my favorite example is the COVID-19 exposure notifications work that we did with Apple.

Corey: An unfortunate example, but a useful one. I—

Aparna: Yes.

Corey: —we all—I think we all wish it wasn’t necessary, but here’s the world in which we live. Please, tell me more.

Aparna: I have so many friends in engineering and mathematics and these technical fields, and they’re always looking at ways that technology can solve these problems. And I think especially something like the pandemic which is so difficult to track, so difficult with the time that it takes for this virus to incubate and so on, so difficult to track these exposures, using the smartphone, using Bluetooth, to have a record of who has it and who they’ve been in contact with, I think really interesting engineering problem, really interesting human problem. So, we were able to work on that, and of course, when you need a platform that’s going to be easy to use, that’s going to be something that you can put into production quickly, you’re going to use Cloud Run. So, they used Cloud Run, and they also used Cloud Run for Anthos, which is the more hybrid version, for the on-prem piece. And so both of those were used in conjunction to back all of the services that were used in the notifications work.

So, those are some of the examples. I think net-net, it’s that I think usability, especially in enterprise software is extremely important, and I think that’s the direction in which software development is going.

Corey: Are you building cloud applications with a distributed team? Check out Teleport, an open source identity-aware access proxy for cloud resources. Teleport provides secure access to anything running somewhere behind NAT: SSH servers, Kubernetes clusters, internal web apps and databases. Teleport gives engineers superpowers!

  • Get access to everything via single sign-on with multi-factor.
  • List and see all SSH servers, kubernetes clusters or databases available to you.
  • Get instant access to them all using tools you already have.

Teleport ensures best security practices like role-based access, preventing data exfiltration, providing visibility and ensuring compliance. And best of all, Teleport is open source and a pleasure to use.

Download Teleport at https://goteleport.com. That’s goteleport.com.

Corey: It’s easy for me to watch folks—like you—in keynotes at events—like Cloud Next—talk about things and say, “This is how the world is building things, and this is what the future looks like.” And I can sit there and pick to pieces all day, every day. It basically what I do because of deep-seated personality problems with me. It’s very different to say that about a customer who has then taken that thing and built it into something that is transformative and solves a very real problem that they have. I may not relate to that problem that they have, but I do not believe that customers are going to have certain problems, find solutions like this and fix them, and the wrong in how they’re approaching these things.

No one sees the constraints that shape things; no one shows up in the morning hoping to do a crap job today unless you know you’re the VP of Integrity at Facebook or something. But there’s a very real sense of companies have a bunch of different drivers, and having a tool or a service or a platform that solves it for them, you’d better be very sure before you step up and start saying, “No, you’re doing it wrong.” In earlier years, I did not see a whole lot of customer involvement with Cloud Next. It was always a, “Well, a bunch of Googlers are going to tell me how this stuff works, and they’ll talk about theoretical things.”

That’s not the case anymore. You have a whole bunch of highly respectable reference customers out there doing a whole lot of really interesting things. And more to the point, they’re willing to go on record talking about this. And I’m not talking about fun startups that are, “Great, it’s Twitter, only for pets.” Great. I’m talking banks, companies where mistakes are going to show and leave a mark. It’s really hard to reconcile what I’m seeing with Google Cloud in 2021 than what I was seeing in, let’s say, five or six years ago. What drove that change?

Aparna: Yes, Corey, I think you’re definitely correct about that. There’s no doubt about it that we have a number of really tremendous customers, we really tremendous enterprise references and so on. I run the Google Cloud Developer Platform, and for me, the developers that I work with and the developers that this platform serves are the inspiration for what we do. And in the last six or seven years that I’ve worked in Google Cloud, that has always been the case. So, nothing has changed from my perspective, in that regard.

If anything, what has changed is that we have far more users, we have been growing exponentially, and we have many more large enterprise customers, but in terms of my journey, I started with the Kubernetes open-source project, I was one of the very early people on that, and I was working with a lot of developers, in that case, in the open-source community, a lot of them became GKE customers, and it just grew. And now we have so many [laugh] customers and so many developers, and we have developed this platform with them. We are very much—it’s been a matter of co-innovation, especially on Kubernetes. It has been very much, “Okay, you tell us,” and it’s a need-based relationship, you know? Something is not working, we are there and we fix it.

Going back to 2017 or whenever it was that Pokemon Go was running on GKE, that was a moment when we realized, “Oh, this platform needs to scale. Okay, let’s get at it.” And that’s where, Corey, it really helps to have great engineers. For all the pros and cons, I think that’s where you want those super-sharp, super-driven, super-intelligent folks because they can make things like that happen, they can make it happen in less than a week, so that—they can make it happen over a Saturday so that Pokemon Go can go live in Japan and everybody can be playing that game. And that’s what inspires me.

And that’s a game, but we have a lot of customers that are running health applications. We have a customer that’s running ambulances on the platform. And so this is life-threatening stuff; we have to take that very seriously, and we have to be listening to them and working with them. But I’m inspired, and I think that our roadmap, and the products, and the features that we build are inspired by what they are building on the platform. And they’re combining all kinds of different things. They’re taking our machine learning capabilities, they’re taking our analytics capabilities, they’re taking our Maps API, and they’re combining it with Cloud Run, they’re combining it with GKE. Often they’re using both of those.

And they’re running new services. We’ve got a customer in Indonesia that's running in a food delivery service; I’ve got customers that are analyzing the cornfields in the middle of the country to improve crop yield. So, that’s the kind of inspiring work, and each of those core, each of those users are coming back to us and saying, “Oh, you know, I need a different type of”—it’s very detailed, like, “I need a different type of file system that gives me greater speed or better performance.” We just had a gaming company that was running on GKE that we really won out over a different cloud in terms of performance improvements that we were able to provide on the container startup times. It was just a significant performance improvement. We’ll probably publish it in the coming few months.

That’s the kind of thing that drives it, and I’m very glad that I have a strong engineering team in Google Cloud, and I’m very glad that we have these amazing customers that are trying to do these amazing things, and that they’re directly engaging with us and telling us what they need from us because that’s what we’re here for.

Corey: To that end, one more area I want to go into before we call this a show, you’ve had Cloud Build for a little while, and that’s great. Now, at—hot off the presses, you wound up effectively taking that one step further with Cloud Deploy. And I am still mostly someone with terrible build and release practices that people would be ashamed of, struggle to understand the differentiation between what I would do with Cloud Build and what I would do with Cloud Deploy. I understand they’re both serverless. I understand that they are things that large companies care about. What is the story there?

Aparna: Yeah, it’s a journey. As you start to use containers—and these days, like you said, Corey, containers, a lot of people are using them—then you start to have a lot of microservices, and one of the benefits of container usage is that it’s really quick to release new versions. You can have different versions of your application, you can test them out, you can roll them out. And so these DevOps practices, they become much more attainable, much more reachable. And we just put out the, I think, the seventh version of the DevOps Research Report—the DORA report—that shows that customers that follow best practices, they achieve their results two times better in terms of business outcomes, and so on.

And there’s many metrics that show that this kind of thing is important. But I think the most important thing I learned during the pandemic, as we were coming out of the pandemic, is a lot of—and you mentioned enterprises—large banks, large companies’ CIOs and CEOs who basically were not prepared for the lockdown, not prepared for the fact that people aren’t going to be going into branches, they came to Google Cloud and they said that, “I wish that I had implemented DevOps practices. I wish that I had implemented the capability to roll out changes frequently because I need that now. I need to be able to experiment with a new banking application that’s mobile-only. I need to be able to experiment with curbside delivery. And I’m much more dependent on the software than I used to be. And I wish that I had put those DevOps practices.”

And so the beginning of 2021, all our conversations were with customers, especially those, you know you said ‘legacy,’ I don’t think that’s the right word, but the traditional companies that have been around for hundreds of years, all of them, they said, “Software is much more important. Yes, if I’m not a software company, at least a large division of my group is now a software group, and I want to put the DevOps practices into play because I know that I need that and that’s a better way of working.”

By the way, there’s a security aspect to that I’d like to come back to because it’s really important—especially in banking, financial services, and public sector—as you move to a more agile DevOps workflow, to have security built into that. So, let me come back to that. But with regard to Cloud Build and Cloud Deploy is something I’ve been wanting to bring into market for a couple of years. And we’ve been talking about it, we’ve been working on it actively for more than a year on my team. And I’m very, very excited about this service because what it does is it allows you to essentially put this practice, this DevOps practice into play whereas your artifacts are built and stored in the artifact repository, they can then automatically be deployed into your runtime—which is GKE Cloud Run—in the future, you can deploy them, and you can set how you want to deploy them.

Do you want to deploy them to a particular environment that you want to designate the test environment, the environment to which your developers have access in a certain way? Like, it’s a test environment, so they can make a lot of changes. And then when do you want to graduate from test to staging, and when do you want to graduate to production and do that gradual rollout? Those are some of the things that Cloud Deploy does.

And I think it’s high time because how do you manage microservices at scale? How do you really take advantage of container-based development is through this type of tooling. And that’s what Cloud Deploy does. It’s just the beginning of that, but it’s a delightful product. I’ve been playing around with it; I love it, and we’ve seen just tremendous reception from our users.

Corey: I’m looking forward to kicking the tires on it myself. I want to circle back to talk about the security aspect of it. Increasingly, I’m spending more of my attention looking at cloud security because everyone else has, too, and some of us have jobs that don’t include the word security but need to care about it. That’s why I have a Thursday edition of my newsletter, now, talking specifically about that. What is the story around security these days from your perspective?

And again, it’s a huge overall topic, and let’s be clear here, I’m not asking, “What does Google Cloud think about security?” That would fill an encyclopedia. What is your take on it? And where do you want to talk about this in the context of Cloud Deploy?

Aparna: Yeah, so I think about security from the perspective of the Google Cloud Developer Platform, and specifically from the perspective of the developer. And like you said, security is not often in the title of anybody in the developer organization, so how do we make it seamless? How do we make it such that security is something that is not going to catch you as you’re doing your development? That’s the critical piece. And at the same time, one of the things we saw during 2020 and 2021 is just the number of cyberattacks just went through the roof. I think there was a 400 to 600% increase in the number of software supply chain attacks. These are attacks where some malicious hacker has come in and inserted some malicious code into your software. [laugh]. Your software, Corey. You know, you the unsuspecting developer is—

Corey: Well, it used to be my software; now there’s some debate about that.

Aparna: Right. That’s true because most software is using open-source dependencies; and these open-source dependencies, they have a pretty intricate web of dependencies that they are themselves using. So, it’s a transitive problem where you’re using a language like Python, or whatever language you’re using. And there’s a number of—

Corey: Crappy bash by default. But yes.

Aparna: Well, it was actually a bash script vulnerability, I think, in the Codecov breach that happened, I think it was, in earlier this year, where a malicious bash script was injected into the build system, in fact, of Codecov. And there are all these new attack vectors that are specifically targeting developers. And whether it’s nation-states or whoever it is that’s causing some of these attacks, it’s a problem that is of national and international magnitude. And so I’m really excited that we have the expertise in Google Cloud and beyond Google Cloud.

Google, it’s a very security-conscious company. This company is a very security-conscious company. [laugh]. And we have built a lot of tooling internally to avoid those kinds of attacks, so what we’ve done with Cloud Build, and what we’re going to do with Cloud Deploy, we’re building in the capability for code to be signed, for artifacts to be signed with cryptographic keys, and for that signing, that attestation—we call it an attestation—that attestation to be checked at various points along the software supply chain. So, as you’re writing code, as you’re submitting the code, as you’re building the containers, as you’re storing the containers, and then finally as you’re deploying them into whatever environment you’re deploying them, we check these keys, and we make sure that the software that is going through the system is actually what you intended and that there isn’t this malicious code injection that’s taking place.

And also, we scan the software, we scan the code, we scan the artifacts to check for vulnerabilities, known vulnerabilities as well as unknown vulnerabilities. Known vulnerabilities from a Google perspective; so Google’s always a little bit ahead, I would say, in terms of knowing what the vulnerabilities are out there because we do work so much on software across operating systems and programming languages, just across the full gamut of software in the industry, we work on it, and we are constantly securing software. So, we check for those vulnerabilities, we alert you, we help to remediate those vulnerabilities.

Those are the type of things that we’re doing. And it’s all in service of certainly keeping enterprise developers secure, but also just longtail an average, everybody, helping them to be secure so that they don’t get hacked and their companies don’t get hacked.

Corey: It’s nice to see people talking about this stuff, who is not directly a security vendor. But by which I mean, you’re not using this as the fear, uncertainty, and doubt angle to sell a given service that, “We have to talk about this exploit because otherwise, no one will ever buy this.” Something like Cloud Deploy is very much aligned with a best practices approach to release engineering. It’s not, strictly speaking, a security product, but being able to wrap things that are very security-centric around it is valuable.

Now, sponsors are always going to do interesting things at various expo halls, and oh, yeah, saw the same product warmed over. This is very much not that, and I don’t interpret anything you’re saying is trying to sell something via the fear, uncertainty, and doubt model. There are a lot of different areas that I will be skeptical hearing about from different companies; I do take security words from Google extremely seriously because, let’s be clear, in the past 20 however many years it has been, you have established a clear track record for caring about these things.

Aparna: Yeah. And I have to go back to my initial mission statement, which is to help developers accelerate time to value. And one of the things that will certainly get in the way of accelerating time to value is security breaches, by the nature of them. If you are not running a supply chain that is secure, then it is very difficult for you to empower your developers to do those releases frequently and to update the software frequently because what if the update has an issue? What if the update has a security vulnerability?

That’s why it’s really important to have a toolchain that prevents against that, that checks for those things, that logs those things so that there’s an audit trail available, and that has the capability for your security team to set policies to avoid those kinds of things. I think that’s how you get speed. You get with security built in, and that’s extremely important to developers and especially cloud developers.

Corey: I want to thank you for taking the time to speak to me about all the things that you’ve been working on and how you view this industry unfolding. If people want to learn more about what you’re up to, and how you think about these things, where can they find you?

Aparna: Well, Corey, I’m available on Twitter, and that may be one of the best ways to reach me. I’m also available at various customer events that we are having, most of them are online now. And so I’ll provide you more details on that and I can be reached that way.

Corey: Excellent. I will, of course, include links to that in the [show notes 00:38:43]. Thank you so much for being so generous with your time. I appreciate it.

Aparna: Thank you so much. I greatly enjoyed speaking with you.

Corey: Aparna Sinha, Director of Product Management at Google Cloud. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. And that sentence needed the word ‘cloud’ about four more times in it. And if you’ve enjoyed this episode, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with a loud angry comment telling me that I just don’t understand serverless well enough.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Dan

After earning his CPA in New York, Dan dedicated his early career to education, helping to build eight schools across two continents. Three of those schools make up the charter school network Coney Island Prep, a 160-person, $20mm+ organization where Dan served as both CFO and COO.

He has served as CFO of many fast-growth start-ups, is a recurring guest lecturer at Harvard Graduate School of Education, and is an avid adventurer and musician. But most importantly, he's a dedicated husband and an enamored father who, at this point, knows the lyrics to each and every Raffi song ever created.

Links:

  • Duckbillgroup.com: https://duckbillgroup.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Redis, the company behind the incredibly popular open source database that is not the bind DNS server. If you’re tired of managing open source Redis on your own, or you’re using one of the vanilla cloud caching services, these folks have you covered with the go to manage Redis service for global caching and primary database capabilities; Redis Enterprise. Set up a meeting with a Redis expert during re:Invent, and you’ll not only learn how you can become a Redis hero, but also have a chance to win some fun and exciting prizes. To learn more and deploy not only a cache but a single operational data platform for one Redis experience, visit redis.com/hero. Thats r-e-d-i-s.com/hero. And my thanks to my friends at Redis for sponsoring my ridiculous non-sense.

Corey: This episode is sponsored in part by Honeycomb. When production is running slow, it's hard to know where problems originate: is it your application code, users, or the underlying systems? I’ve got five bucks on DNS, personally. Why scroll through endless dashboards, while dealing with alert floods, going from tool to tool to tool that you employ, guessing at which puzzle pieces matter? Context switching and tool sprawl are slowly killing both your team and your business. You should care more about one of those than the other, which one is up to you. Drop the separate pillars and enter a world of getting one unified understanding of the one thing driving your business: production. With Honeycomb, you guess less and know more. Try it for free at Honeycomb.io/screaminginthecloud. Observability, it’s more than just hipster monitoring.

Corey: Welcome to Screaming in the Cloud, I’m Corey Quinn. A common myth that has sort of permeated the entire ecosystem is, when you associate a single person with a company, they’re the only person is really there. It turns out that with AWS, Jeff Barr is writing an awful lot of blog posts, but I have it on good authority that there are at least three services he didn’t personally create himself. Similarly here at The Duckbill Group, there are a lot of folks who aren’t necessarily in the public eye as much as I am because they don’t have the overriding personality flaws that I do. One of those people is my guest today. Dan Shapiro is our CFO, and I guess the closest thing you could consider to us having adult supervision. Dan, thanks for joining me.

Dan: Thanks for having me, and I appreciate you considering me an adult, even though I don’t always feel that way.

Corey: I always assume there’s someone else out there who knows what’s going on and how life works, and one day they’re going to bother to explain it to me because the alternative is that none of us know and we’re making it up as we go along, and I’m not sure I can wrap my head around that fear.

Dan: Yeah, my job is to allow you guys to have the vision, some wild ideas, try to put data to ideas, tell you if I think it’s a good idea. I think, you know, I’m the balloon string to your guys’ balloon is the way I think about it.

Corey: It’s funny, given that we fix the AWS bill and deal with customers, who are in many cases themselves working in finance, that for the first year or so that we were doing this as The Duckbill Group, and never back when I was independent did we really have anyone in a CFO position. What we were doing made sense to me from a very simplistic naive point of view. Okay, well, how is the company doing? Well, I have an app that tells me that. No, it’s not QuickBooks, it’s I’m going to pull up the banking app and see how much money is currently sitting in the company accounts.

And that would inform things like can we hire new people, can we afford this piece of equipment, or embark on this project? And it was sort of Fisher-Price finance, from our perspective. And when you came in, you were very good at this in that you did not make us feel like the naive hayseeds we very clearly were. You were excellent at hiding your contempt, which is great. But before we dive into the specifics, let’s ask the big question that I didn’t know the answer to back when we first started working with you, which is, what does a CFO do?

Dan: CFO does a lot of tactical things, but ultimately, I think that they have two main 30,000-foot view functions. One is to safeguard the assets of the company, and two is to ensure that those assets are employed in the best way possible for the best outcomes that the company is after. And that doesn’t always necessarily have to be financial; it could be operational successes. But whatever the goals of the company are, the CFO’s role is to utilize the assets to help approximate those goals as best as you can, get the most out of your resources. So, I think that, you know, you play a little bit of defense in trying to make sure that you’re protected and you play a little bit of offense in trying to make sure that when you put your chips on the table, that they’re in the right place.

Corey: So, please take this in the spirit of which it’s intended, but if I’m starting out my career today, I can attend a boot camp, learn how the ins-and-outs of software programming work, get a job somewhere, and if I manage my career right, in two or three years I’ll get a job somewhere as a senior engineer. So, from that perspective, are you more or less an accountant with title inflation? What’s going on? What is the difference between finance and accounting?

Dan: Yeah, so to define those two words, I think of the delineator being time, right? So, today backwards is accounting. It’s a historical record of everything that has happened, well categorized. And I think of finance as today forward. It is the strategy and planning and potentially analysis of what you think is going to happen in the future.

So, those two buckets are delineated by time in my mind. Am I an accountant? Yes. Do I think that the best financial people have an accounting background? I do. I think accounting is the alphabet of finance and I think it’s hard to really fully understand or be effective without, you know, at least some baseline accounting information. So yeah, I think accounting is super important. It’s a codebase to an engineer, except there’s really only one code.

Corey: I will say that once you started working with us, it enabled me to understand things in a way that I hadn’t before. And on some level, I feel… a little embarrassed, I’m in my mid 30s, here—which annoys the hell out of my wife because I’m 39—but I’ve gotten this far in my career without understanding a lot of the basic literacy of corporate finance. Now, let’s be clear, I’m very good at handling personal finance aspects to my life. I can quote chapter and verse in some cases, from ERISA requirements around 401(k)s and how much I’m allowed to put in in a given year. And that’s great, but business finance looks very different than personal finance in a few key ways. And I’m curious as to what your impression is of those key differences because I have thoughts, but I don’t generally invite experts in areas onto this show and then explain their field to them.

Dan: [laugh]. I will go into that but I am curious what you think. You know, or what you perceive in your experience in working with me, what you’ve seen—maybe one or two differences that you’ve seen between personal finance and corporate finance. I certainly feel like there’s differences, but I actually don’t feel like there’s a tremendous differences, I just think it’s much greater scale with a lot more people involved. But I’m curious what your experience has been, and then I can kind of jump off of that.

Corey: Sure. From my perspective, it is—and this is probably heresy, but I tend to view finance in many ways as being more psychological than it is mathematical. And if I were to pick a random person off the street, and give them a choice of you need to come up with a thousand bucks: you can either make another $1,000 or you can save another $1,000—and let’s be clear, I’m talking about someone who is in a technical-style career track; I understand the margins, this becomes a very different thing in either direction and I want to be sensitive to that—but if I ask most people, their answer is almost universally going to be to save the money because if you look at folks who are in a salary job, if they want to make more money, they need to either come up with a side project and start moonlighting, they need to petition their boss for a raise, they need to do a few other things here and there and okay, it becomes a pain. And well, I haven’t updated my resume in years, I’m not great at interviewing for jobs, I don’t really want to leave and upset the applecart. Whereas saving a thousand bucks when you’re making six figures a year is not necessarily that hard. Eat out less, cancel Netflix, et cetera, et cetera, and you can get there by saving.

Companies philosophically are the exact opposite. Because there’s a theoretical upper bound of one hundred percent of a company’s AWS bill that I can cut, either by moving them to another provider—which is cheating—or flying to Seattle and taking hostages—which is not particularly something we’d like to do, and it doesn’t generally work more than once. And that’s fine, but they can earn a multiple of that by launching the right feature or product to the right market at the right time, faster. That’s always going to be more compelling for a company because they’re after growth, not protecting the baseline that they have, in most cases. When that shifts, they tend to be companies in decline.

Dan: Yeah, that’s a great point. And I think to add onto that, when you’re running a company, almost every company is going to have something like 60 to 70% of their expenses tied up in their people. So, cutting discretionary spending at a corporate level is not always going to yield the impact that you might need it to. And if you actually need to cut expenses, you’re talking about very serious decisions about reducing your workforce, which happens in big steps, right? You can’t just find, you know, 5k here, 10k there; you’re talking about, you know, reducing your org chart, which is a very serious decision that a lot of companies take very seriously, which is great.

Versus there’s a lot of avenues to increasing revenue, you know, new product streams, growing your team, investing in your team, raising prices, follow-on sales with these existing customers. You know, there’s just a lot more opportunity to generate revenue because companies, by definition, have a lot more opportunities to create revenue than just, you know, like you said, somebody who has a salary job.

Corey: One of the things that really woke me up to what our customers are going through is I take a look at the finances of The Duckbill Group—as pointed out and categorized and [aligned 00:09:09] by you, which first, thank you, it’s extremely helpful, and two, it helps me contextualize things. You raised a flag on things being out of bounds, where they had been historically. At one point, our AWS bill was a little over $2,000 a month, and your question was, “What’s up with this?” It is not the sort of thing that was going to make or break the company; our payroll is six figures a month and compared to that the AWS bill is irrelevant.

But given what we do, and given that it’s always a good idea to be good financial stewards of the money that has been entrusted to us, “Great, let’s take a look at that.” And it was dropped down to 700, 800 bucks. I think it [crested 00:09:47] at $950 last month, as of the time of this recording because I was doing some fun experiments. And sure it’s annoying in that I want to keep it as low as possible just given the nature of what we do, but from a business perspective, it does not fundamentally matter. Now, that is not the case for most of our clients, it matters; it’s a large expense, but payroll is still bigger.

In many cases, depending on the company, real estate is bigger. And Netflix, one of the larger AWS customers, has publicly stated that their biggest expense is content. Yeah, they’ve built a bunch of studios, they have to get rights to all the content that they stream, they pay people through the nose, and their AWS bill is reportedly massive—it would have to be given what they do—but it’s never the number one driving focus. And, on some level, it feels like a company deals with two classes of problem. There’s the side that we’re on of cost control, risk mitigation, et cetera—insurance hangs out here, too—and the other is speeding up time to market or growth and expansion. Our class of problem is a good diligence thing, but it isn’t usually ‘summon the board in the middle of the night because of an opportunity’ territory. If I look
at those as the two great problems, it is more lucrative as a business to be targeting the former.

Dan: Yeah. I mean, so just to go back to that AWS example for a second, right, it’s contradictory, right, because on the one hand, are you really going to win the day by shaving, you know, $500 off your AWS bill a month if you’re a $3 million company? You know, is that going to move the needle? Not necessarily.

But I believe in hygiene and habits, and I believe as you’re a growing company, that it’s a lot harder to put in good financial process and good financial hygiene, the bigger you get. And so if you start doing these things early, you know, if you have a quarterly review of your subscriptions, your SaaS products, you have some kind of budget versus actual process in place where you have—even if it’s wrong, right? I mean, you just take a stab at what you think, you track your actuals, and at least now you have some data to rally around from a financial perspective, as opposed to saying, “Oh, look, we spent X on Y. That’s interesting.” Or, “Does that feel right? Or”—

And so you have a baseline and you have a measuring stick, and then you have a process in place for figuring out what happened, and then how to better guess in the future. And so yeah, I think that hygiene is important and I think the hard part—and I think it was exhibited by you guys early on—is when you’re starting a company, this is the last thing that you want to focus on, or even think you need to focus on because it’s not on fire, right? Finance is never on fire until it’s absolutely on fire, and it’s going to kill you. And I think a lot of founders sweep finance and accounting to the back of the closet because a client is upset, or an employee quit, or their code broke, and those things are on fire today, and if they don’t get fixed, the company can’t move forward.

And so finance, it doesn’t matter, doesn’t matter, doesn’t matter, doesn’t matter; it kills your company. So, I think it’s not that the founders don’t know about it, aren’t smart enough for it, don’t care about it. It’s just that it’s not on fire, and when you’re starting a company, all you can do is pay attention to the things that are.

Corey: It’s easy to look back and beat myself up for my lack of understanding or lack of focus on things. The similarly I talked to technical people all the time where the first DevOps hire into an environment. It’s been application engineers or developers building the environment so far, and it’s easy to look at it in a condescending way of, “Oh, our entire field knows not to build things this way. What’s wrong with you fools?” Well, it worked well enough to get to a point where they could afford to hire you, so perhaps show some respect when you’re in an environment like that.

This is something that most business folk, like I don’t know, a CFO intrinsically understands, but many engineers still don’t seem to have fully wrap their heads around just because it’s easy to over-index on the area that you’re focusing in. You almost certainly view companies, start to finish, through a lens of finance. A lot of engineering folks I spend time with—and used to be one of, myself—view it through a lens of well, what’s their stack look like? What’s their technical debt? How are they actually architected?

Which is interesting, but usually not the indicator of whether a company will succeed or fail. There are other broader focuses that are important to look at, and being able to view our own company through a lens of something other than dealing with just the technology or just the sales aspect of it, or, “All right, time for me to go shitposting again because that’s our substitute for marketing.” Instead, we’re focusing on what the larger picture is through a finance lens. It introduces a level of rigor that I hadn’t expected, though, clearly, we do still put our own stamp on it. I suspect we are probably the only company you have ever worked with that has an explicit line item in the budget labeled ‘Spite’ for example, from which we make our ridiculous parody videos and other assorted nonsense, for those who are unfamiliar with the Spite Budget.

Dan: That was a new one for me. I’ve seen, I think, every other chart of account line item before and Duckbill Group was the first one that I had to type in ‘Spite Budget’ for. I didn’t even really understand what it meant, and then as I, you know, got to know you a little bit better, I knew exactly what it meant. And it is—

Corey: Yeah, this is what passes humor with his set.

Dan: [laugh].

Corey: Got it. Okay, I’ll smile, nod, and continue to keep the finances in order.

Dan: Yeah. So, when you leave, I just cross it out and I write marketing. But [laugh] it is what makes you you and what makes this company successful. But in all seriousness, it’s an outstanding financial investment because our business is—like every business—is driven by eyeballs. And whether you’re a media company or you’re any other company, you need people to know about you.

You need people to want to believe that your services are valuable, and want to engage in your services. And the Spite Budget brings eyeballs. And the other thing is, I think it’s always wrapped up in farce, but there’s always an underlying truth to it which is why people find it compelling. I think if you were out there actually, slinging shit to us, you know, [laugh] you’re parlance, I think you wouldn’t get 100 Twitter followers. But I think because the undercurrent is truth, you don’t give yourself credit for this but there’s an immense amount of knowledge behind the truth that people find it compelling. So yeah, from a serious financial perspective, I think it’s one of the best investments that we make. [laugh].

Corey: It’s definitely a lot of fun. And it always bothered me when I would walk around re:Invent or travel for client trips and whatnot to see billboard ads in airports because I’ve priced out what those things cost, and for those who are wondering, it’s not a small number. And okay, so you’re going to go to all of the expense of buying out ads in multiple airports doing a country wide brand saturation campaign, and with all of that space, and all of those millions of dollars you’re spending on this to get your message in front of folks, you fill it with something that is so anodyne, something that is so… I guess, droll, that no one notices or cares. You wind up getting all of these eyeballs and then have nothing interesting to say. And to me, that’s the cardinal sin.

I can build an audience, and the way I built it is by having something interesting to say. But take a look at any brand awareness billboard for a large consultancy. I mean, Accenture had one years ago in airports that I still haven’t stopped making fun of them for, even more so since they blocked me on Twitter. And it said, “The new isn’t on its way. It’s here now. New, applied now.” So yeah, I guess they’re bringing the new and it’s… okay, they spent how many millions of dollars on that campaign and then wound up effectively building the tagline and the phrasing just by, I don’t know, having some analyst in some junior role bang their head off a keyboard a few times? I don’t get it. I just don’t get it.

Dan: Yeah, without being a marketing expert, I think that’s the product of risk tolerance being a directly inverse relationship to the height of an org chart. I think the higher you grow in your decision-making capability and potency, the more responsibility you have on your shoulders, the less willing you are to do anything that’s interesting, fun, risky, and that’s why we have a lot of the marketing that we have today. We’re lucky to be a small company, we’re lucky to be a small company that has creative founders. I think a lot of founders for their marketing go, they use a lot of external folks, and you know, they—sometimes it works, but a lot of times they lose the actual, like, heart and soul of the company because they just aren’t in it, they don’t understand it. And I think it’s fun, a lot of the stuff that, you know, we get to do.

But I think we’re fortunate that we are in a unique spot in the sense that—you know, we talk about this a lot—we’re a company that has a services arm and we’re a company that has a media arm. And where many companies in services have to think about marketing and sales in a very straight up and down way, we get to play in our media space, which is a revenue-generating arm of our business, in a lot of experimental ways that other companies are spending thousands, millions of dollars on marketing expenses, we get to do it and make money doing it, via media. And it’s really this incredible, bifurcated business where one serves the other and vice versa. And, you know, it’s fascinating.

Corey: I still don’t pretend to understand how we got to the place that we did. When you look back, it’s easy to see a sense of plodding inevitability, where, “Oh, yeah. You did this, that led to that, and it led to this, and here you are now.” But at the time you’re making the decisions, you are throwing darts blindfolded. Professional advice for those in bars, don’t do that. They do not find it nearly as amusing as it sounds.

Dan: I think that’s how it feels, but I don’t actually believe that story that you like to tell. You and Mike are—[laugh] you enjoy the self-deprecation but you’re very bright guys and I think you are always fiscally responsible. You just didn’t have the language that I have to show how you’re fiscally responsible. And I think you really were conservative founders in the fact that you made very short-term decisions and always made conservative short-term decisions, and that puts you in a good place. I think what I have brought to the table is an ability to look a little bit further out and think about, you know, okay, it’s not just next month, right? It’s not just two months from now. What are we building here? What are the long-term goals, and what are the intermediary goals that will be true if we’re on the right path?

And that’s some planning, that’s using historicals, that’s using some good analysis tools to try and get there, but it’s also a matter of time. When you’re a founder of a company, you can’t spend all of your time on finance and accounting. It’s just the reality. You have 8 million things to do, and so you do exactly what you guys did, which is you look at your bank account, you try and make a decision in a vacuum. But you don’t have time to really grind over the details, or the strategy of the next 6, 12, 18 months because you’ve got people to hire, and you’ve got clients to appease, and you’ve got work to do that will generate the revenue.

And so there’s a whole world of fractional CFOs, some who are good, some are not good. There’s a world of bookkeepers out there that are competent. And I think, even though it’s not a great use of cash, or revenue-generating use of cash, which I think a lot of startups, you know, are reticent to go down that path, I think having good hygiene and good strategy on your financial front early on is necessary for the future planning of a lot of these startups. And also probably peace of mind, right? I mean, I’m not sure—

Corey: Oh, I sleep way better now than I did before.

Dan: [laugh]. Yeah, I hope so. I mean, I think just the not knowing creates such an undertone of stress that, you know, may or may not be recognized by founders. And I think being able to look at a document and say, “Okay, this makes sense to me, I know what the future probably holds.” And I hope it allows you to think about other things than money.

Corey: This episode is sponsored by our friends at Oracle Cloud. Counting the pennies, but still dreaming of deploying apps instead of "Hello, World" demos? Allow me to introduce you to Oracle's Always Free tier. It provides over 20 free services and infrastructure, networking, databases, observability, management, and security. And—let me be clear here—it's actually free. There's no surprise billing until you intentionally and proactively upgrade your account. This means you can provision a virtual machine instance or spin up an autonomous database that manages itself all while gaining the networking load, balancing and storage resources that somehow never quite make it into most free tiers needed to support the application that you want to build. With Always Free, you can do things like run small scale applications or do proof-of-concept testing without spending a dime. You know that I always like to put asterisks next to the word free. This is actually free, no asterisk. Start now. Visit snark.cloud/oci-free that's snark.cloud/oci-free.

Corey: One challenge we had was that a lot of the advice for first-time founders—since there’s a lot of those in our space—is not applicable to
us. The way that Mike and I built this place, we provided the initial investment personally; we’re the only investors in this place, we don’t have outside funding, we don’t have investors, and we’ve given no equity away, so it is all us. And the way that most companies in tech tend to work is I would go and debase myself in front of a bunch of investors, and some VC would absolutely bite on my Twitter for Pets idea and give me $50 million to go and build the prototype.

And at that point, I’m not making money. I have $50 million in the bank that I am burning through every month as I hire, as I build the MVP, et cetera, et cetera, and there’s a ticking clock–it’s called a runway in our space—and if we don’t wind up hitting a certain milestone of being able to raise more money by the time that we’ve crossed a certain point, it’s time to begin an orderly shutdown of the company. And that’s a ticking clock hanging over the head of many founders. In our case, I was looking at it like that originally, but we are profitable, we are actively closing deals, and one of the things you taught me, for example, was when we have accounts receivable, money that is owed to us that has not yet been paid, we can count that and use that as in many ways an asset because it is. It also helps the fact that we are in a market where businesses do not generally decline to pay us money that they owe us.

This is very much an understood thing in business, but my question was always, “Well, yeah, but what if they don’t pay? What does that mean for us?” And that’s really been a non-issue, which at first I was really grateful for and thought, “Wow, we have great customers.” It turns out that this is expected. This is like running a store and being super ecstatic because most of the people in your store aren’t shoplifting.

It’s, yes, that is the baseline expectation for people doing business in the modern era. That was an eye opener. But it also kept me up at night because it was, “Oh, if suddenly this money in the bank is all that we’re going to get in, then we only have enough runway to go two, three, four months and then we have to shut the company down.” Yeah, but we’re profitable every month, so what’s the concern here? It doesn’t work that way.

Dan: Right. We look at two key metrics as a company every single week, right? We look at a metric called months of cash.

Corey: Which is pretty self-descriptive. It’s in the title.

Dan: Yeah, [laugh] just, you take your bank balance and you divide it by your, you know, average monthly burn, you know, total expenses out each month. And that will give you some number. You know, we really like that number to be at least two, closer to three, but I think something like having, like, six months of cash—now this is for a company that has not raised a bunch of money, right? So, if you’ve raised $5 million, your months of cash is going to be well more than that because you’re investing in your—

Corey: And the goal should be to lower that months of cash because it’s designed to be used to hire and grow, not sit there, and you should not be turning a profit on that interest; you should be spending through it at a reasonable clip.

Dan: Right. And for companies who have not raised money, right, even if you have a month’s of cash, you know, five, six, that’s probably not right either, right? I mean, there’s probably opportunities for investment where you want to get your chips on the table for growth, right? Because you don’t always have to be after hypergrowth. But typically, if a company’s not growing, it’s declining, and that’s not a great thing, right? So, some amount of growth is always important for a healthy company.

So, months of cash is one metric that we look at every week. And the other is what I call an adjusted quick ratio. But essentially, it’s taking current assets over your current liabilities. Now, many startups are going to be debt free, short-term debt free, and so what we do is we try and talk about cash plus accounts receivable—so invoices that are out that have not been paid—and that gives you your numerator, over any outflows that you should experience over the next 30 or so days, right, we think about that as, like, our current liabilities. Because there’s no real debt.

In larger corporate finance, you have big, big balance sheets, quick ratios are more formulaic corporate formula, but in this case, we look at cash plus accounts receivable, over 30 days of liabilities. And that gives us a number that we’re tracking constantly. And we have our own benchmarks for internally for what that number is, but that’s a really great tracking of liquidity; that’s more than just what’s the cash in the bank, right? You’re looking at your cash, you’re looking at your receivables, and you’re looking at your upcoming payments, you know, whether that’s payroll or big vendor payments or any other, you know, non-recurring expenses that you know are coming down the pike.

Corey: Yeah, a lot of things can impact that number. And that means, oh, that number is out of kilter. It’s time to dig a little bit into why, but it’s not itself a diagnostic, it’s just an indicator that gives a snapshot of overall financial health.

Dan: That’s right. There are also ratios that you can track on a weekly basis and not just wait month-over-month because you can actually start to see trend lines up and down and start to ask questions about, “Hey, what happened this week?” “Oh, we sent out a big invoice, you know, we just haven’t been paid on it. That’s why we saw our AQR go up.” Or, “Hey, we just took in a big invoice for redoing the website. We owe $50,000 to x vendor and that’s why the denominator of that equation went up. And so our AQR went down a little bit this week, but we expect it to climb back up in the coming weeks.”

And so it’s a really good way to track liquidity. On a longer-term, we’re talking about things like margin analysis, right? Gross margins is an important thing we talk about. You know, revenue, minus the cost of goods to deliver that revenue, want to make sure that we’re delivering products with efficiency. And, you know, we talk about net margin or EBITDA, or, you know, however, you want to describe the bottom line, and that’s obviously your profitability.

So, those are sort of the things we rally around internally. We try and look at them at least weekly, or monthly, depending on the metric. And this just goes back to hygiene and habits. You don’t have to do a tremendous amount of work here. And you also—I think, people can get wrapped up in esoteric formulas that they’ve read in books, right, but I think if you’re looking at cash, you’re looking at cash and receivables over expenses, you’re looking at your margins, and you’re looking at your bottom line, those four metrics tracked well, you know, weekly or monthly over time, should really protect you, and should also give you great insights about your business.

Corey: This is another example of viewing personal and company finances through a different lens. If I were to come home and tell my spouse that I had just dropped $50,000 on a website for example, the conversation would not be pleasant immediately afterwards. But it’s not personal money. Conversely, an awful lot of business owners that I’ve heard stories about get into trouble by treating the company as their personal piggy bank, which Mike and I are very clear never to do. If you’re listening to this from tax authority, I want to emphasize, we keep our noses remarkably clean.

And again, it comes down to one of those, I don’t want to lose sleep at night over things. That’s why you’re here, Dan. I don’t want to lose sleep over the idea that we’re going to run out of money, or that I’m going to get audited and my great fraud will be discovered. It’s no, I don’t want to get audited because it’s a pain in the neck, and I have to wind up providing a whole bunch of receipts and whatnot, but everything’s legitimate. That’s the type of fear and paperwork thing that I just want to be able to avoid, by and large.

Dan: Yeah, we do keep clean books, and there’s probably things that we could be expensing that we’re not, right? More than the tax gray area that exists there is the difficulty with partners and equity—and I don’t mean equity on the balance sheet; I mean, like, fairness amongst partners, right? So, if one partner is going out more often than the other for dinners, is that really fair for that person to get their meals paid for more than the partner? Vice versa, if one person lives in a city and the other person doesn’t and their car runs through the business. You know, how do you deal with those differences with multiple owners of a business? And I think those complexities are sometimes more difficult to deal with in the tax gray area of what is deductible and what is not?

Corey: Oh, yeah. That’s why I’ve always been such a big fan of making sure that when you have a business partner that your values are aligned. Mike and I are incredibly dissimilar in a bunch of different ways. He loves spreadsheets, I love shitposting on Twitter. But when it comes to values in how we approach these things that was laid out at the very beginning in our partnership agreement.

And it’s never been something that we’ve had cause to even visit, like, “Well, let’s see what the partnership agreement says.” Just because it makes sense. It’s, yeah, Mike doesn’t spend tens of thousands of dollars a year on travel—in non-pandemic years—just because he’s not constantly out there doing the things that I have been doing, historically. Now, post-pandemic that might change it, if he winds up traveling a lot and I’m sitting at home, I’m not sitting there upset because he gets to expense dinners out. It doesn’t work that way. Whenever you start seeing that type of breakdown that feels like it’s a proxy for something that’s deeper.

Dan: Yeah, I would agree with that. I think also just not like a marriage, you know, it comes down to communication and expectations. I mean, like any management in company, right? I mean, you have to be clear about what is expected of your partner. Ideally, you’re measuring those things, and you’re making sure that everybody’s in line with those expectations, but yeah, I would say you and Mike have always been aligned in that sense, and I think you’re right, the instance—or examples where it goes south, it’s not because someone spent $600 more than the other, it’s because there’s a lot more going on there.

Corey: Yeah, it’s the proxy for things. So, at some point it’s, “Okay, why do you have a $4 million yacht that just parked in your driveway? And yeah, what’s the deal here?” One last area I want to cover because this is probably relevant to folks who I imagine most of our listeners are, who are employees elsewhere. I always was extremely bothered by the fact that I had remarkably little financial wiggle room when it came to expensing things.

Like, “Let me get this straight. I have root in production, but you’re going to question me over a $15 expense?” That always sat very strangely. And I carry that with me to the point where even now at the time of this recording, we’re going through one of our periodic exercises that you put us through of, yeah, here’s a list of all the recurring expenses on your credit card. Is this something that we still need? And how do I
categorize it if so? If not, let’s cancel it.

And it’s weird because instinctively, I hear that, even now, as you saying, what’s up with this $9.99 a month thing? And I’m sitting here going,
“Wait a minute.” Like, I’m trying to justify, like, “Is that really worth $10 a month to me? Should I—wait a minute, I don’t work for you.”

Dan: Right. [laugh].

Corey: It was a bit of a different moment here. Like, I felt like I’m being called to the carpet by my boss, but that is absolutely not what’s going on. So, my question for you, as someone who has a storied career in finance, is that what’s being asked in most companies by management when they question expenses? Is that about, “Help me understand what this is,” or is it in fact what I’m hearing, a, “Justify why you think a $10 monthly expense is worthwhile?”

Dan: I think it really depends. I mean, in my case, in our case, I am genuinely trying to understand how to categorize most things so that when I do my reporting, I can tell you where we spent our money. So, you know, if I see a charge that I don’t know what it is and I ask you for clarity on it, it is a hundred percent so that I understand is this COGS? Was this part of delivering our revenue? Is this, you know, a sales and marketing expense that at the end of the year, I could say, “Hey, we spent x percent on sales and marketing for y revenue,” and I want that to be an accurate number.

If you’re talking about big corporate bureaucracy and they have expense policies and a lot of rigorous rules that were derived from the head of HR who doesn’t know the 500 people that this rule is going to impact, I think there’s probably a different motivation for those questions. And I think it’s tough; it’s not to say that you shouldn’t have policies, you need to have policies, but you know, the bigger you get, the harder it is to have the right policies for the right people and you end up making blanket decisions to try and get to an average lau that nobody’s probably happy with, instead of empowering your people and trusting your people, that they’re going to spend their money on things that are good for the company. At Duckbill, I know, it’s not even a question. Everybody here has a corporate card, and I don’t think I’ve once seen you question somebody about an expense that they had. The only questions they get are from me in terms of how to use it for accounting purpose, but the underlying current there is we hired these people; we trust these people; they’re our employees and we’re all swimming in the same direction, so if they spent money on something, I would assume it’s because it’s good for the company.

And I think that goes a long way. I would always prefer to have better tracking and reporting metrics to spot the bad apples than to have a
policy that turns the good apples, the 95% of good apples into sour apples. Not to have a terrible analogy on this recording. But I [laugh] would much rather convey to my employees that I trust them and then do a little bit of extra work on the back end to ask questions about, you know, the expenses I think are a little funny, versus writing a policy that says, “Here’s x. Here’s what you can spend on y.” And now everyone is dissatisfied. Someone was telling me just because someone wore shorts to the office, don’t write a policy, “No shorts in the office.” Just talk to that person and say, “Hey, pants might be a better choice.”

Corey: Right. You should not need a list of policies that are organically built every time someone does something they shouldn’t. And at some level, it’s one of those ideas where collective punishment never works. I wound up getting an email from my boss once to the entire team of, “We have to start at nine or there’s going to be problems.” And I’m sweating bullets because I came in at 9:03, a couple of days this week.

Meanwhile, the person next to me has slowly started drifting from coming in at 10:45 to 11:30. And it’s one of those I think I’m going to get fired. No, no, no, it’s really you have one particular person and you’re just being chickenshit as far as approaching them and telling them that there’s an issue.

Dan: That’s right. Yep. Totally agree.

Corey: Dan, thank you so much for taking the time to speak with me today. If people want to learn more, where can they find you?

Dan: I don’t have a huge footprint on the interwebs, but I’m always here at Duckbill, you know, serving the company, and if anybody has specific Duckbill questions, they can always reach out and I’m happy to converse. But I really appreciate you having me, and happy to talk Duckbill anytime.

Corey: No, it’s appreciated. Dan Shapiro, CFO at The Duckbill Group. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry comment talking about how as an accountant, you are annoyed by this episode because you will never be depreciated in your own lifetime.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Ivan

Ivan Pepelnjak, CCIE#1354 Emeritus, is an independent network architect, blogger, and webinar author at ipSpace.net. He's been designing and implementing large-scale service provider and enterprise networks as well as teaching and writing books about advanced internetworking technologies since 1990.

https://www.ipspace.net/About_Ivan_Pepelnjak

Links:

  • ipSpace.net: https://ipspace.net

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by my friends at ThinkstCanary. Most companies find out way too late that they’ve been breached. ThinksCanary changes this and I love how they do it. Deploy canaries and canary tokens in minutes and then forget about them. What's great is the attackers tip their hand by touching them, giving you one alert, when it matters. I use it myself and I only remember this when I get the weekly update with a “we’re still here, so you’re aware” from them. It’s glorious! There is zero admin overhead to this, there are effectively no false positives unless I do something foolish. Canaries are deployed and loved on all seven continents. You can check out what people are saying at canary.love. And, their Kub config canary token is new and completely free as well. You can do an awful lot without paying them a dime, which is one of the things I love about them. It is useful stuff and not an, “ohh, I wish I had money.” It is speculator! Take a look; that’s canary.love because it's genuinely rare to find a security product that people talk about in terms of love. It really is a unique thing to see. Canary.love. Thank you to ThinkstCanary for their support of my ridiculous, ridiculous non-sense.

Corey: Developers are responsible for more than ever these days. Not just the code they write, but also the containers and cloud infrastructure their apps run on. And a big part of that responsibility is app security — from code to cloud.

That’s where Snyk comes in. Snyk is a frictionless security platform that meets developers where they are, finding and fixing vulnerabilities right from the CLI, IDEs, repos, and pipelines. And Snyk integrates seamlessly with AWS offerings like CodePipeline, EKS, ECR, etc., etc., etc., you get the picture! Deploy on AWS. Secure with Snyk. Learn more at snyk.io/scream. That’s S-N-Y-K-dot-I-O/scream. Because they have not yet purchased a vowel.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I have an interesting and storied career path. I dabbled in security engineering slash InfoSec for a while before I realized that being crappy to people in the community wasn’t really my thing; I was a grumpy Unix systems administrator because it’s not like there’s a second kind of those out there; and I dabbled ever so briefly in the wide world of network administration slash network engineering slash plugging the computers in to make them talk to one another, ideally correctly. But I was always a dabbler. When it comes time to have deep conversations about networking, I immediately tag out and look to an expert. My guest today is one such person. Ivan Pepelnjak is oh so many things. He’s a CCIE emeritus, and well, let’s start there. Ivan, welcome to the show.

Ivan: Thanks for having me. And oh, by the way, I have to tell people that I was a VAX/VMS administrator in those days.

Corey: Oh, yes the VAX/VMS world was fascinating. I talked—

Ivan: Yes.

Corey: —to a company that was finally emulating them on physical cards because that was the only way to get them there. Do you refer to them as VAXen, or VAXes, or how did you wind up referring—

Ivan: VAXes.

Corey: VAXes. Okay, I was on the other side of that with the inappropriately pluralizing anything that ends with an X with an en—‘boxen’ and the rest. And that’s why I had no friends for many years.

Ivan: You do know what the first VAX was, right?

Corey: I do not.

Ivan: It was a Swedish Hoover company.

Corey: Ooh.

Ivan: And they had a trademark dispute with Digital over the name, and then they settled that.

Corey: You describe yourself in your bio as a CCIE Emeritus, and you give the number—which is low—number 1354. Now, I’ve talked about certifications on this show in the context of the modern era, and whether it makes sense to get cloud certifications or not. But this is from a different time. Understand that for many listeners, these stories might be older than you are in some cases, and that’s okay. But Cisco at one point, believe it or not, was a shining beacon of the industry, the kind of place that people wanted to work at, and their certification path was no joke.

I got my CCNA from them—Cisco Certified Network Administrator—and that was basically a byproduct of learning how networks worked. There are several more tiers beyond that, culminating in the CCIE, which stands for Cisco Certified Internetworking Expert, or am I misremembering?

Ivan: No, no, that’s it.

Corey: Perfect. And that was known as the doctorate of networking in many circles for many years. Back in those days, if you had a CCIE, you are guaranteed to be making an awful lot of money at basically any company you wanted to because you knew how networking—

Ivan: In the US.

Corey: —worked. Well, in the US. True. There’s always the interesting stories of working in places that are trying to go with the lowest bidder for networking gear, and you wind up spending weeks on end trying to figure out why things are breaking intermittently, and only to find out at the end that someone saved 20 bucks by buying cheap patch cables. I digress, and I still have the scars from those.

But it was fascinating in those days because there was a lab component of getting those tests. There were constant rumors that in the middle of the night, during the two-day certification exam, they would come in and mess with the lab and things you’d set up—

Ivan: That’s totally true.

Corey: —you’d have to fix it the following day. That is true?

Ivan: Yeah. So, in the good old days, when the lab was still physical, they would even turn the connectors around so that they would look like they would be plugged in, but obviously there was no signal coming through. And they would mess up the jumpers on the line cards and all that stuff. So, when you got your broken lab, you really had to work hard, you know, from the physical layer, from the jumpers, and they would mess up your config and everything else. It was, you know, the real deal. The thing you would experience in real world with, uh, underqualified technicians putting stuff together. Let’s put it this way.

Corey: I don’t wish to besmirch our brethren working in the data centers, but having worked with folks who did some hilariously awful things with cabling, and how having been one of those people myself from time to time, it’s hard to have sympathy when you just spent hours chasing it down. But to be clear, the CCIE is one of those things where in a certain era, if you’re trying to have an argument on the internet with someone about how networks work and their responses, “Well, I’m a CCIE.” Yeah, the conversation was over at that point. I’m not one to appeal to authority on stuff like that very often, but it’s the equivalent of arguing about medicine with a practicing doctor. It’s the same type of story; it is someone where if they’re wrong, it’s going to be in the very fringes or the nuances, back in this era. Today, I cannot speak to the quality of CCIEs. I’m not attempting to besmirch any of them. But I’m also not endorsing that certification the way I once did.

Ivan: Yeah, well, I totally agree with you. When this became, you know, a mass certification, the reason it became a mass certification is because reseller discounts are tied to reseller status, which is tied to the number of CCIEs they have, it became, you know, this, well, still high-end, but commodity that you simply had to get to remain employed because your employer needed the extra two point discount.

Corey: It used to be that the prerequisite for getting the certification was beyond other certifications was, you spent five or six years working on things.

Ivan: Well, that was what gave you the experience you needed because in those days, there were no boot camps. Today, you have [crosstalk 00:06:06]—

Corey: Now, there’s boot camp [crosstalk 00:06:07] things where it’s we’re going to train you for four straight weeks of nothing but this, teach to the test, and okay.

Ivan: Yeah. No, it’s even worse, there were rumors that some of these boot camps in some parts of the world that shall remain unnamed, were actually teaching you how to type in the commands from the actual lab.

Corey: Even better.

Ivan: Yeah. You don’t have to think. You don’t have to remember. You just have to type in the commands you’ve learned. You’re done.

Corey: There’s an arc to the value of a certification. It comes out; no one knows what the hell it is. And suddenly it’s, great, you can use that to really identify what’s great and what isn’t. And then it goes at some point down into the point where it becomes commoditized and you need it for partner requirements and the rest. And at that point, it is no longer something that is a reliable signal of anything other than that someone spent some time and/or money.

Ivan: Well, are you talking about bachelor degree now?

Corey: What—no, I don’t have one of those either. I have—

Ivan: [laugh].

Corey: —an eighth grade education because I’m about as good of an academic as it probably sounds like I am. But the thing that really differentiated in my world, the difference between what I was doing in the network engineering sense, and the things that folks like you who were actually, you know, professionals rather than enthusiastic amateurs took into account was that I was always working inside of the LAN—Local Area Network—inside of a data center. Cool, everything here inside the cage, I can make a talk to each other, I can screw up the switching fabric, et cetera, et cetera. I didn’t deal with any of the WAN—Wide Area Network—think ‘internet’ in some cases. And at that point, we’re talking about things like BGP, or OSPF in some parts of the world, or RIP. Or RIPv2 if you make terrible life choices.

But BGP is the routing protocol that more or less powers the internet. At the time of this recording, we’re a couple weeks past a BGP… kerfuffle that took Facebook down for a number of hours, during which time the internet was terrific. I wish they could do that more often, in fact; it was almost like a holiday. It was fantastic. I took my elderly relatives out and got them vaccinated. It was glorious.

Now, we’re back to having Facebook and, terrific. The problem I have whenever something like this happens is there’s a whole bunch of crappy explainers out there of, “What is BGP and how might it work?” And people have angry opinions about all of these things. So instead, I prefer to talk to you. Given that you are a networking trainer, you have taught people about these things, you have written books, you have operated large—scale environments—

Ivan: I even developed a BGP course for Cisco.

Corey: You taught it for Cisco, of all places—

Ivan: Yeah. [laugh].

Corey: —back when that was impressive, and awesome and not a has-been. It’s honestly, I feel like I could go there and still wind up going back in time, and still, it’s the same Cisco in some respects: ‘evolve or die dinosaur,’ and they got frozen in amber. But let’s start at the very beginning. What is BGP?

Ivan: Well, you know, when the internet was young, they figured out that we aren’t all friends on the internet anymore. And I want to control what I tell you, and you want to control what you tell me. And furthermore, I want to control what I believe from what you’re telling me. So, we needed a protocol that would implement policy, where I could say, “I will only announce my customers to you, but not what I’ve heard from Verizon.” And you will do the same.

And then I would say, “Well, but I don’t want to hear about that customer of yours because he’s also my customer.” So, we need some sort of policy. And so they invented a protocol where you will tell me what you have, I will tell you what I have and then we would both choose what we want to believe and follow those paths to forward traffic. And so BGP was born.

Corey: On some level, it seems like it’s this faraway thing to people like me because I have a residential internet connection and I am not generally allowed to make my own BGP announcements to the greater world. Even when I was working in data centers, very often the BGP was handled by our upstream provider, or very occasionally by a router they would drop in with the easiest maintenance instructions in the world for me of, “Step one, make sure it has power. Step two, never touch it. Step three, we’d prefer if you don’t even look at it and remain at least 20 feet away to keep from bringing your aura near anything we care about.” And that’s basically how you should do with me in the context of hardware. So, it was always this arcane magic thing.

Ivan: Well, it’s not. You know, it’s like power transmission: when you know enough about it, it stops being magic. It’s technology, it’s a bit more complicated than some other stuff. It’s way less complicated than some other stuff, like quantum physics, but still, it’s so rarely used that it gets this aura of being mysterious. And then of course, everyone starts getting their opinion, particularly the graduates of the Facebook Academy.

And yes, it is true that usually BGP would be used between service providers, so whenever, you know, we are big enough to need policy, if you just need one uplink, there is no policy there. You either use the uplink or you don’t use the uplink. If you want to have two different links to two different points of presence or to two different service providers, then you’re already in the policy land. Do I prefer one provider over the other? Do I want to announce some things to one provider but other things to the other? Do I want to take local customers from both providers because I want to, you know, have lower latency because they are local customers? Or do I want to use one solely as the backup link because I paid so little for that link that I know it’s shitty.

So, you need all that policy stuff, and to do that, you really need BGP. There is no other routing protocol in the world where you could implement that sort of policy because everything else is concerned mostly with, let’s figure out as fast as possible, what is reachable and how to get there. And BGP is like, “Hey, slow down. There’s policy.”

Corey: Yeah. In the context of someone whose primary interaction with networks is their home internet, where there’s a single cable coming in from the outside world, you plug it into a device, maybe yours, maybe ISPs, maybe we don’t care. That’s sort of the end of it. But think in terms of large interchanges, where there are multiple redundant networks to get from here to somewhere else; which one should traffic go down at any given point in time? Which networks are reachable on the other end of various distant links? That’s the sort of problem that BGP is very good at addressing and what it was built for. If you’re running BGP internally, in a small network, consider not doing exactly that.

Ivan: Well, I’ve seen two use cases—well, three use cases for people running BGP internally.

Corey: Okay, this I want to hear because I was always told, “No touch ‘em.” But you know, I’m about to learn something. That’s why I’m talking to you.

Ivan: The first one was multinationals who needed policy.

Corey: Yes. Many multi-site environments, large-scale companies that have redundant links, they’re trying to run full mesh in some cases, or partial mesh where—between a bunch of facilities.

Ivan: In this case, it was multiple continents and really expensive transcontinental links. And it was, I don’t want to go from Europe to Sydney over US; I want to go over Middle East. And to implement that type of policy, you have to split, you know, the whole network into regions, and then each region is what BGP calls an autonomous system, so that it gets its stack, its autonomous system number and then you can do policy on that saying, “Well, I will not announce Asian routes to Europe through US, or I will make them less preferred so that if the Middle East region goes down, I can still reach Asia through US but preferably, I will not go there.”

The second one is yet again, large networks where they had too many prefixes for something like OSPF to carry, and so their OSPF was breaking down and the only way to solve that was to go to something that was designed to scale better, which was BGP.

And third one is if you want to implement some of the stuff that was designed for service providers, initially, like, VPNs, layer two or layer three, then BGP becomes this kitchen sink protocol. You know, it’s like using Route 53 as a database; we’re using BGP to carry any information anyone ever wants to carry around. I’m just waiting for someone to design JSON in BGP RFC and then we are, you know… where we need to be.

Corey: I feel on some level, like, BGP gets relatively unfair criticism because the only time it really intrudes on the general awareness is when something has happened and it breaks. This is sort of the quintessential network or systems—or, honestly, computer—type of issue. It’s either invisible, or you’re getting screamed at because something isn’t working. It’s almost like a utility. On some level. When you turn on a faucet, you don’t wonder whether water is going to come out this time, but if it doesn’t, there’s hell to pay.

Ivan: Unless it’s brown.

Corey: Well, there is that. Let’s stay away from that particular direction; there’s a beautiful metaphor, probably involving IBM, if we do. So, the challenge, too, when you look at it is that it’s this weird, esoteric thing that isn’t super well understood. And as soon as it breaks, everyone wants to know more about it. And then in full on charging to the wrong side of the Dunning-Kruger curve, it’s, “Well, that doesn’t sound hard. Why are they so bad at it? I would be able to run this better than they could.” I assure you, you can’t. This stuff is complicated; it is nuanced; it’s difficult. But the common question is, why is this so fragile and able to easily break? I’m going to turn that around. How is it that something that is this esoteric and touches so many different things works as well as it does?

Ivan: Yeah, it’s a miracle, particularly considering how crappy the things are configured around the world.

Corey: There have been periodic outages of sites when some ISP sends out a bad BGP announcement and their upstream doesn’t suppress it because hey, you misconfigured things, and suddenly half the internet believes oh, YouTube now lives in this tiny place halfway around the world rather than where it is currently being Anycasted from.

Ivan: Called Pakistan, to be precise.

Corey: Exact—there was an actual incident there; we are not dunking on Pakistan as an example of a faraway place. No, no, an Pakistani ISP wound up doing exactly this and taking YouTube down for an afternoon a while back. It’s a common problem.

Ivan: Yeah, the problem was that they tried to stop local users accessing YouTube. And they figured out that, you know, YouTube, is announcing this prefix and if they would announce to more specific prefixes, then you know, they would attract the traffic and the local users wouldn’t be able to reach YouTube. Perfect. But that leaked.

Corey: If you wind up saying that, all right, the entire internet is available on this interface, and a small network of 256 nodes available on the second interface, the most specific route always wins. That’s why the default route or route of last resort is the entire internet. And if you don’t know where to send it, throw it down this direction. That is usually, in most home environments, the gateway that then hands it up to your ISP, where they inspect it and do all kinds of fun things to sell ads to you, and then eventually get it to where it’s going.

This gets complicated at these higher levels. And I have sympathy for the technical aspects of what happened at Facebook; no sympathy whatsoever for the company itself because they basically do far more harm than they do good and I’ve been very upfront about that. But I want to talk to you as well about something that—people are going to be convinced I’m taking this in my database direction, but I assure you I’m not—DNS. What is the relationship between BGP and DNS? Which sounds like a strange question, sometimes.

Ivan: There is none.

Corey: Excellent.

Ivan: It’s just that different large-scale properties decided to implement the global load-balancing global optimal access to their servers in different ways. So, Cloudflare is a typical example of someone who is doing Anycast, they are announcing the same networks, the same prefixes, from hundreds locations around the world. So, BGP will take care that you always get to the close Cloudflare [unintelligible 00:18:46]. And that’s it. That’s how they work. No magic. Facebook didn’t believe in the power of Anycast when they started designing their service. So, what they’re doing is they have DNS servers around the world, and the DNS servers serve the local region, if you wish. And that DNS server then decides what facebook.com really stands for. So, if you query for facebook.com, you’ll get a different answer in Europe than in US.

Corey: Just a slight diversion on what Anycast is. If I ping Google’s public resolver 8.8.8.8—easy to remember—from my computer right now,
the packet gets there and back in about five milliseconds.

Wherever you are listening to this, if you were to try that same thing you’d see something roughly similar. Now, one of two things is happening; either Google has found a way to break the laws of physics and get traffic to a central point faster than light for the 8.8.8.8 that I’m talking to and the one that you are talking to are not in fact the same computer.

Ivan: Well, by the way, it’s 13 milliseconds for me. And between you and me, it’s 200 millisecond. So yes, they are cheating.

Corey: Just a little bit. Or unless they tunneled through the earth rather than having to bounce it off of satellites and through cables.

Ivan: No, even that wouldn’t work.

Corey: That’s what the quantum computers are for. I always wondered. Now, we know.

Ivan: Yeah. They’re entangling the replies in advance, and that’s how it works. Yeah, you’re right.

Corey: Please continue. I just wanted to clarify that point because I got that one hilariously wrong once upon a time and was extremely confused for about six months.

Ivan: Yeah. It’s something that no one ever thinks about unless, you know, you’re really running large-scale DNS because honestly, root DNS servers were Anycasted for ages. You think they’re like 12 different root DNS servers; in reality, there are, like, 300 instances hidden behind those 12 addresses.

Corey: And fun trivia fact; the reason there are 12 addresses is because any more than that would no longer fit within the 512 byte limit of a UDP packet without truncating.

Ivan: Thanks for that. I didn’t know that.

Corey: Of course. Now, EDNS extensions that you go out with a larger [unintelligible 00:21:03], but you can’t guarantee that’s going to hit. And what happens when you receive a UDP packet—when you receive a DNS result with a truncate flag set on the UDP packet? It is left to the client. It can either use the partial result, or it can try and re-establish over a TCP connection.

That is one of those weird trivia questions they love to ask in sysadmin interviews, but it’s yeah, fundamentally, if you’re doing something that requires the root nameservers, you don’t really want to start going down those arcane paths; you want it to just be something that fits in a single packet not require a whole bunch of computational overhead.

Ivan: Yeah, and even within those 300 instances, there are multiple servers listening to the same IP address and… incoming packets are just sprayed across those servers, and whichever one gets the packet replies to it. And because it’s UDP, it’s one packet in one packet out. Problem solved. It all works. People thought that this doesn’t work for TCP because, you know, you need a whole session, so you need to establish the session, you send the request, you get the reply, there are acknowledgements, all that stuff.

Turns out that there is almost never two ways to get to a certain destination across the internet from you. So, people thought that, you know, this wouldn’t work because half of your packets will end in San Francisco, and half of the packets will end in San Jose, for example. Doesn’t work that way.

Corey: Why not?

Ivan: Well, because the global Internet is so diverse that you almost never get two equal cost paths to two different destinations because it would be San Francisco and San Jose announcing 8.8.8.8 and it would be a miracle if you would be sitting just in the middle so that the first packet would go to San Francisco, the second one would go to San Jose, and you know, back and forth. That never happens. That’s why Cloudflare makes it work by analysing the same prefix throughout the world.

Corey: So, I just learned something new about how routing announcements work, an aspect of BGP, and you a few minutes ago learned something about the UDP size limit and the root name servers. BGP and DNS are two of the oldest protocols in existence. You and I are also decades into our careers. If someone is starting out their career today, working in a cloud-y environment, there are very few network-centric roles because cloud providers handle a lot of this for us. Given these protocols are so foundational to what goes on and they’re as old as they are, are we as an industry slash sector slash engineers losing the skills to effectively deploy and manage these things?

Ivan: Yes. The same problem that you have in any other sufficiently developed technology area. How many people can build power lines? How many people can write a compiler? How many people can design a new CPU? How many people can design a new motherboard?

I mean, when I was 18 years old, I was wire wrapping my own motherboard, with 8-bit processor. You can’t do that today. You know, as the technology is evolving and maturing, it’s no longer fun, it’s no longer sexy, it stops being a hobby, and so it bifurcates into users and people who know about stuff. And it’s really hard to bridge the gap from one to the other. So, in the end, you have, like, this 20 [graybeard 00:24:36] people who know everything about the technology, and the youngsters have no idea. And when these people die, don’t ask me [laugh] how we’ll get any further on.

Corey: This episode is sponsored by our friends at CloudAcademy. That’s right, they have a different lab challenge up for you called, “Code Red: Repair an AWS Environment with a Linux Bastion Host.” What does it do? Well, its going to assess your ability to troubleshoot AWS networking and security issues in a production like environment. Well, kind of, its not quite like production because some exec is not standing over your shoulder, wetting themselves while screaming. But..ya know, you can pretend in fact I’m reasonably certain you can retain someone specifically for that purpose should you so choose. If you are the first prize winner who completes all four challenges with the fastest time, you’ll win a thousand bucks. If you haven’t started yet you can still complete all four challenges between now and December 3rd to be eligible for the grand prize. There's only a few days left until the whole thing ends, so I would get on it now. Visit cloudacademy.com/corey. That’s cloudacademy.com/C-O-R-E-Y, for god’s sake don’t drop the “E” that drives me nuts, and thank you again to Cloud Academy for not only promoting my ridiculous non sense but for continuing to help teach people how to work in this ridiculous environment.

Corey: On some level, it feels like it’s a bit of a down the stack analogy for what happened to me early in my career. My first systems administration job was running a large-scale email system. So, it was a hobby that I was interested in. I basically bluffed my way into working at a university for a year—thanks, Chapman; I appreciate that [laugh]—and it was great, but it was also pretty clear to me that with the rise of things like hosted email, Gmail, and whatnot, it was not going to be the future of what the present day at that point looked like, which was most large companies needed an email administrator. Those jobs were dwindling.

Now, if you want to be an email systems administrator, there are maybe a dozen companies or so that can really use that skill set and everyone else just outsources that said, at those companies like Google and Microsoft, there are some incredibly gifted email administrators who are phenomenal at understanding every nuance of this. Do you think that is what we’re going to see in the world of running BGP at large scale, where a few companies really need to know how this stuff works and everyone else just sort of smiles, nods and rolls with it?

Ivan: Absolutely. We’re already there. Because, you know, if I’m an end customer, and I need BGP because I have to uplinks to two ISPs, that’s really easy. I mean, there are a few tricks you should follow and hopefully, some of the guardrails will be built into network operating systems so that you will really have to configure explicitly that you want to leak [unintelligible 00:26:15] between Verizon and AT&T, which is great fun if you have too low-speed links to both of them and now you’re becoming transit between the two, which did happen to Verizon; that’s why I’m mentioning them. Sorry, guys.

Anyway, if you are a small guy and you just need two uplinks, and maybe do a bit of policy, that’s easy and that’s achievable, let’s say with some Google and paste, and throwing spaghetti at the wall and seeing what sticks. On the other hand, what the large-scale providers—like for example Facebook because we were talking about them—are doing is, like, light years away. It’s like comparing me turning on the light bulb and someone running, you know, nuclear reactor.

Corey: Yeah, you kind of want the experts running some aspects on that. Honestly, in my case, you probably want someone more competent flipping the light switch, too. But that’s why I have IoT devices here that power my lights, it on the one hand, keeps me from hurting myself on the other leads to a nice seasonal feel because my house is freaking haunted.

Ivan: So, coming back to Facebook, they have these DNS servers all around the world and they don’t want everyone else to freak out when one of these DNS servers goes away. So, that’s why they’re using the same IP address for all the DNS servers sitting anywhere in the world. So, the name server for facebook.com is the same worldwide. But it’s different machines and they will give you different answers when you ask, “Where is facebook.com?”

I will get a European answer, you will get a US answer, someone in Asia will get whatever. And so they’re using BGP to advertise the DNS servers to the world so that everyone gets to the closest DNS server. And now it doesn’t make sense, right, for the DNS server to say, “Hey, come to European Facebook,” if European Facebook tends to be down. So, if their DNS server discovers that it cannot reach the servers in the data center, it stops advertising itself with BGP.

Why would BGP? Because that’s the only thing it can do. That’s the only protocol where I can tell you, “Hey, I know about this prefix. You really should send the traffic to me.” And that’s what happened to Facebook.

They bricked their backbone—whatever they did; they never told—and so their DNS server said, “Gee, I can’t reach the data center. I better stop announcing that I’m a DNS server because obviously I am disconnected from the rest of Facebook.” And that happens to all DNS servers because, you know, the backbone was bricked. And so they just, you know, [unintelligible 00:29:03] from the internet, they've stopped advertising themselves, and so we thought that there was no DNS server for Facebook. Because no DNS server was able to reach their core, and so all DNS servers were like, “Gee, I better get off this because, you know, I have no clue what’s going on.”

So, everything was working fine. Everything was there. It’s just that they didn’t want to talk to us because they couldn’t reach the backend servers. And of course, people blamed DNS first because the DNS servers weren’t working. Of course they weren’t. And then they blame the BGP because it must be BGP if it isn’t DNS. But it’s like, you know, you’re blaming headache and muscle cramps and high fever, but in fact you have flu.

Corey: For almost any other company that wasn’t Facebook, this would have been a less severe outage just because most companies are interdependent on each other companies to run infrastructure. When Facebook itself has evolved the way that it has, everything that they use internally runs on the same systems, so they wound up almost with a bootstrapping problem. An example of this in more prosaic terms are okay, the data center had a power outage. Okay, now I need to power up all the systems again and the physical servers I’m trying to turn on need to talk to a DNS server to finish booting but the DNS server is a VM that lives on those physical servers. Uh-oh. Now, I’m in trouble. That is a overly simplified and real example of what Facebook encountered trying to get back into this, to my understanding.

Ivan: Yes, so it was worse than that. It looks like, you know, even out-of-band management access didn’t work, which to me would suggest that out-of-band management was using authentication servers that were down. People couldn’t even log to Zoom because Zoom was using single-sign-on based on facebook.com, and facebook.com was down so they couldn’t even make Zoom calls or open Google Docs or whatever. There were rumors that there was a certain hardware tool with a rotating blade that was used to get into a data center and unbrick
a box. But those rumors were vehemently denied, so who knows?

Corey: The idea of having someone trying to physically break into a data center in order to power things back up is hilarious, but it does lead to an interesting question, which is in this world of cloud computing, there are a lot of people in the physical data centers themselves, but they don’t have access, in most cases to log into any of the boxes. One of the most naive things I see all the time is, “Oh well, the cloud provider can read all of your data.” No, they can’t. These things are audited. And yeah, theoretically, if they’re lying outright, and somehow have
falsified all of the third-party audit stuff that has been reported and are willing to completely destroy their business when it gets out—and I assure you, it would—yeah, theoretically, that’s there. There is an element of trust here. But I’ve had to answer a couple of journalists questions recently of, “Oh, is AWS going to start scanning all customer content?” No, they physically cannot do it because there are many ways you can configure things where they cannot see it. And that’s exactly what we want.

Ivan: Yeah, like a disk encryption.

Corey: Exactly. Disk encryption, KMS on some level, using—rolling your own, et cetera, et cetera. They use a lot of the same systems we do. The point being, though, is that people in the data centers do not even have logging rights to any of these nodes for the physical machines, in some cases, let alone the customer tenants on top of those things. So, on some level, you wind up with people building these systems that run on top of these computers, and they’ve never set foot in one of the data centers.

That seems ridiculous to me as someone who came up visiting data centers because I had to know where things were when they were working so I could put them back that way when they broke later. But that’s not necessary anymore.

Ivan: Yeah. And that’s the problem that Facebook was facing with that outage because you start believing that certain systems will always work. And when those systems break down, you’re totally cut off. And then—oh, there was an article in ACM Queue long while ago where they were discussing, you know, the results of simulated failures, not real ones, and there were hilarious things like phone directory was offline because it wasn’t on UPS and so they didn’t know whom to call. Or alerts couldn’t be diverted to a different data center because the management station for alert configuration was offline because it wasn’t on UPS.

Or, you know the one, right, where in New York, they placed the gas pump in the basement, and the diesel generators were on the top floor, and the hurricane came in and they had to carry gas manually, all the way up to the top floor because the gas pump in the basement just stopped working. It was flooded. So, they did everything right, just the fuel wouldn’t come to the diesel generators.

Corey: It’s always the stuff that is under the hood on these things that you can’t make sense of. One of the biggest things I did when I was evaluating data center sites was I’d get a one-line diagram—which is an electrical layout of the entire facility—great. I talked to the folks running it. Now, let’s take a walk and tour it. Hmmm, okay. You show four transformers on your one-line diagram. I see two transformers and two empty concrete pads. It’s an aspirational one-line diagram. It’s a joke that makes it a one-liner diagram and it’s not very funny. So it’s, okay if I can’t trust you for those little things, that’s a problem.

Ivan: Yeah, well, I have another funny story like that. We had two power feeds coming into the house plus the diesel generator, and it was, you know, the properly tested every month diesel generator. And then they were doing some maintenance and they told us in advance that they will cut both power feeds at 2 a.m. on a Sunday morning.

And guess what? The diesel generator didn’t start. Half an hour later UPS was empty, we were totally dead in water with quadruple redundancy because you can’t get someone it’s 2 a.m. on a Sunday morning to press that button on the diesel generator. In half an hour.

Corey: That is unfortunate.

Ivan: Yeah, but that’s how the world works. [laugh].

Corey: So, it’s been fantastic reminding myself of some of the things I’ve forgotten because let’s be clear, in working with cloud, a lot of this stuff is completely abstracted away. I don’t have to care about most of these things anymore. Now, there’s a small team of people that AWS who very much has to care; if they don’t, I will say mean things to them on Twitter, if I let my HugOps position slip up just a smidgen. But they do such a good job at this that we don’t have problems like this, almost ever, to the point where when it does happen, it’s noteworthy. It’s been fun talking to you about this just because it’s a trip down a memory lane that is a lot more aligned with the things that are there and we tend not to think about them. It’s almost a How it’s Made episode.

Ivan: Yeah. And don’t be so relaxed regarding the cloud networking because, you know, if you don’t go full serverless with nothing on-premises, you know what protocol you’re running between on-premises and the cloud on direct connect? It’s called BGP.

Corey: Ah. You know, I did not know that. I’ve done some ridiculous IPsec pairings over those things, and was extremely unhappy for a while afterwards, but I never got to the BGP piece of it. Makes sense.

Ivan: Yeah, even over IPsec if you want to have any dynamic failover, or multiple sites, or anything, it’s [BP 00:36:56].

Corey: I really want to thank you for taking the time to go through all this with me. If people want to learn more about how you view these things, learn more things from you, as I’d strongly recommend they should if they’re even slightly interested by the conversation we’ve had, where can they find you?

Ivan: Well, just go to ipspace.net and start exploring. There’s the blog with thousands of blog entries, some of them snarkier than others. Then there are, like, 200 webinars, short snippets of a few hours of—

Corey: It’s like a one man version of re:Invent. My God.

Ivan: Yeah, sort of. But I’ve been working on this for ten years, and they do it every year, so I can’t produce the content at their speed. And then there are three different full-blown courses. Some of them are just, you know, the materials from the webinars, plus guest speakers plus hands-on exercises, plus I personally review all the stuff people submit, and they cover data centers, and automation, and public clouds.

Corey: Fantastic. And we will, of course, put links to that into the [show notes 00:38:01]. Thank you so much for being so generous with your time. I appreciate it.

Ivan: Oh, it’s been such a huge pleasure. It’s always great talking with you. Thank you.

Corey: It really is. Thank you once again. Ivan Pepelnjak network architect and oh so much more. CCIE #1354 Emeritus. And read the bio; it’s well worth it. I am Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice and a comment formatted as a RIPv2 announcement.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and
we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Clinton

Clinton Herget is Principal Solutions Engineer at Snyk, where he focuses on helping our large enterprise and public sector clients on their journey to DevSecOps. A seasoned technologist, Clinton spent his 15+ year career prior to Snyk as a web software engineer, DevOps consultant, cloud solutions architect, and technical director in the systems integrator space, leading client delivery of complex agile technology solutions. Clinton is passionate about empowering software engineers and is a frequent conference speaker, developer advocate, and everything-as-code evangelist.

Links:

  • Try Snyk for free today at: https://snyk.co/Screaming-in-the-Cloud

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by my friends at ThinkstCanary. Most companies find out way too late that they’ve been breached. ThinksCanary changes this and I love how they do it. Deploy canaries and canary tokens in minutes and then forget about them. What's great is the attackers tip their hand by touching them, giving you one alert, when it matters. I use it myself and I only remember this when I get the weekly update with a “we’re still here, so you’re aware” from them. It’s glorious! There is zero admin overhead to this, there are effectively no false positives unless I do something foolish. Canaries are deployed and loved on all seven continents. You can check out what people are saying at canary.love. And, their Kub config canary token is new and completely free as well. You can do an awful lot without paying them a dime, which is one of the things I love about them. It is useful stuff and not an, “ohh, I wish I had money.” It is speculator! Take a look; that’s canary.love because it's genuinely rare to find a security product that people talk about in terms of love. It really is a unique thing to see. Canary.love. Thank you to ThinkstCanary for their support of my ridiculous, ridiculous non-sense.

Corey: Writing ad copy to fit into a 30 second slot is hard, but if anyone can do it the folks at Quali can. Just like their Torque infrastructure automation platform can deliver complex application environments anytime, anywhere, in just seconds instead of hours, days or weeks. Visit Qtorque.io today and learn how you can spin up application environments in about the same amount of time it took you to listen to this ad.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. This promoted episode features Clinton Herget, who’s a principal solutions engineer at Snyk. Or ‘Snick.’ Or ‘Cynic.’ Clinton, thank you for joining me, how the heck do I pronounce your company’s name?

Clinton: That is always a great place to start, Corey, and we like to say it is ‘sneak’ as in sneaking around or a pair of sneakers. Now, our colleagues in the UK do like to say ‘Snick,’ but that is because they speak incorrectly. We will accept it; it is still wrong. As long as you’re not saying ‘Sink’ because it really has nothing to do with plumbing and we prefer to avoid that association.

Corey: Generally speaking, I try not to tell other people how to run their business, but I will make an exception here because I can’t take it anymore. According to CrunchBase, your company has raised $1.4 billion. Buy a vowel for God’s sake. How much could it possibly cost for a single letter that clarifies all of this? My God.

Clinton: Yeah, but then we wouldn’t spend the first 20 minutes of every sales conversation talking about how to pronounce the company name and we would need to fill that with content. So, I think we’re just going to stay the course from here on out.

Corey: I like that. So, you’re a principal solutions engineer. First, what does that do? And secondly, I’ve known an awful lot of folks who I would consider problem engineers, but they never self-describe that way. It’s always solutions-oriented?

Clinton: Well, it’s because I worked for Snyk, and we’re not a problems company, Corey, we’re a solutions company.

Corey: I like that.

Clinton: It’s an interesting role, right, because I work with some of our biggest customers, a lot of our strategic partners here in North America, and I’m kind of the evangelist that comes out and says, “Hey, here’s what sucks about being a developer. Here’s how we could maybe be better.” And I want to connect with other engineers to say, “Look, I share your pain, there might be an easier way, if you, you know, give me a few minutes here to talk about Snyk.”

Corey: So, I’ve seen Snyk around for a while. I’ve had a few friends who worked there almost since the beginning and they talk about this thing—this was before, I believe, you had the Dobermann logo back in the early days—and I keep periodically seeing you folks in a variety of different contexts and different places. Often I’ll be installing something from Docker Hub, for example, and it will mention that, oh, there’s a Snyk scan thing that has happened on the command line, which is interesting because I, to the best of my knowledge, don’t pay Docker for things that I do because, “No, I’m going to build it myself out of popsicle sticks,” is sort of my entire engineering ethos. But I keep seeing you in different cases where as best I am aware, I have never paid you folks for services. What is it you do as a company because you’re one of those folks that I just keep seeing again and again and again, but I can’t actually put my finger on what it is you do.

Clinton: Yeah, you know, most people aren’t aware that popsicle sticks are actually a CNCF graduated project. So, you know, that’s that—

Corey: Oh, and they’re load-bearing in almost every piece of significant technical debt over the last 50 years.

Clinton: Absolutely. Look at your bill of materials; it’s there. Well, here’s where I can drop in the other fun fact about Snyk’s name, it’s actually an acronym, right, stands for So, Now You Know. So, now you know that much, at least. Popsicle sticks, key component to any containerized infrastructure. Look, Snyk is a developer security company, right? And people hear that and go, “I’m sorry, what? I’m a developer; I don’t give a shit about security.” Or, “I’m a security person”—

Corey: Usually they don’t say that out loud as often as you would hope, but it’s like, “That’s not true. I say that I care about security an awful lot.” It’s like, “Yeah, you say that. Therein lies the rub.”

Clinton: Until you get a couple of drinks in them at the party at re:Invent and then the real stuff comes out, right? No, Snyk is always been historically committed to the open-source community. We want to help open-source developers every bit as much as, you know, we’re helping the engineers at our top-tier customers. And that’s because fundamentally, open-source is inextricably linked to the way software is developed today, right? There is nobody not using open-source.

And so we, sort of, have to be supporting those communities at the same time. And that fundamentally is where the innovation is happening. And you know, my sales guys hate when I say this, right, but you can get an amazing amount of value out of Snyk by using the freemium solution, using the open-source tooling that we’ve put out in the community, you get full access to our vulnerability database, which is updated every day, and if you’re working on public projects, that’s going to be free forever, right? We’re fundamentally committed to making that work. If you’re an enterprise that happens to have money to spend, I guess we’ll take that too, right, but my job is really talking to developers and figuring out, you know, how can we reduce the amount of pain in your life through better security tooling?

Corey: The challenging part is that your business, although I confess is significantly larger than my business, we’re sort of on some level solving the same problem. And that sounds odd to say because I focus on fixing AWS bills and you’re focused on improving developer security. But I’m moving up about six levels to the idea that there are only two big problems in the world of technology, in the world of
companies for that matter. And the problem that we’re solving is the worst one of the two. And that is reducing risk exposure.

It is about eliminating downside. It’s cost optimization, it’s security tooling, it is insurance, et cetera, et cetera, et cetera. And the other problem, the one that I’ve always found, that is the thing that will get people actually excited rather than something they feel obligated to do is speeding up time to market, improving feature velocity, being able to deliver the right things sooner. That’s the problem companies are biasing towards investing in extremely heavily. They’ll convene the board to come up with an answer there.

That said, you stray closer into that problem space than most security companies that I’m aware of just because you do in fact, speed up the developer process. It let people move faster, but do it safely at least is my general understanding. If I’m completely wrong on this, and, “Nope, we are purely risk mitigation, then this is going to look fairly silly, but it wouldn’t be the first time I put my foot in my mouth.”

Clinton: Yeah, Corey, it sounds like you really read the first three words of the website, right? “Develop fast. Stay secure.” And I think that fundamentally gets at the traditional alignment, where security equals slow, right, because risk mitigation is all about preventing problematic things from going into production. But only doing that as a stop gate at the end of the process, right, by essentially saying we assume all developers are bad and want to do bad things, and so we’re going to put up this big gate and generate an 1100 page PDF, and then throw it back to them and say, “Now, go figure out all of the bad things you did and how to fix them. And by the way, you’re already overshooting your delivery target.” Right? So, there’s no way to win in that traditional model unless you’re empowering developers earlier with the right context they need to actually write more secure code to begin with, rather than remediating after the fact when those fixes are actually most expensive.

Corey: It’s the idea of the people who want to slow down and protect things and not break are on the operation side of the world, and then you have developers who want to ship things. And you have that natural tension, so we’re going to smash them together and call it DevOps, which at least if nothing else, leads to interesting stories on stages. Whether it actually leads to lasting cultural transformation is another thing entirely. And then someone said, “Well, what about security?” And the answer is, “We have a security department?” And the answer is, “Yeah, you know, those grumpy people that say no all the time whenever we ask if we could do anything.” “Oh, that security department. I ignore them and go around them instead.” And it’s, “All right, well, we need help on that so we’re going to smash them in, too.” Welcome to DevSecOps, which is basically buzzword-driven cultural development. And here we are. But there is something to be said for you can no longer be the Department of No. I would argue that you couldn’t do that successfully previously, but at least now we’re a little more aware of it.

Clinton: I think you could certainly do that when you were deploying software a couple times a year, right? Because you could build in all of the time to very expensively and time consumingly fix things after the fact, right? We’re no longer in that world. I think when you’re deploying every few seconds or a few minutes, what you need is tooling that, first of all, runs at that speed, that gives developers insights into what risk are they bringing on board with that application once it will be deployed, but then also give them the context they actually need to fix things, right? I mean, regardless of where those vulnerabilities are found, it still ultimately is a line of code that has to be written by a developer and committed and pushed through a pipeline to make it back into production.

And that’s true, whether we’re talking about application security and proprietary code, we’re talking about vulnerabilities in open-source, vulnerabilities in the container, infrastructure as code. I mean, it used to be that a network vulnerability was fixed by somebody going into the data center, unplugging a Cat 5 cable and plugging it in somewhere else, right? I mean, that was the definition of network security. It was a hardware problem. Now, networking is software-defined. I mean [laugh]—

Corey: Oh, the firewall I trust is basically a wire cutter. Yeah, cut through the entire cable, and that is the only secure firewall. And it’s like, oh, no, no, there are side-channel attacks. It’s not completely going to solve things for you. Yeah.

Clinton: You know, without naming names, there are certainly vendors in the security space that still consider mitigation to be shutting down
access to a workload, right. Like, let’s remediate by taking this off of the internet and allowing it to no longer be accessible.

Corey: I don’t think it’s come from a security standpoint, but that does feel like it’s a disturbing proportion of Google’s product strategy.

Clinton: [laugh]. Absolutely. But you know, I do think maybe we can take the forward-looking step of saying there are ways to fix issues while keeping applications online at the same time. For example, by arming engineers with the security intelligence they need when they’re making decisions about what goes into those applications. Because those wire cutters now, that’s a line in a YAML file, right?

That’s a Kubernetes deployment, that’s a CloudFormation template, and that is living in code in the same repo with everything else, with all of the other logic. And so it’s fundamentally indistinguishable at the point where all security is really now developer security, except the security tooling available doesn’t speak to the developer, it doesn’t integrate into their workflow, it doesn’t enable them to make remediations, it’s still slapping them on the wrist. And this is why I think when you talk about—to invoke one of the most overused buzzwords in the security industry—when you talk about shifting left, that’s really only half the story. I mean, if you’re taking a traditional solution that’s designed to slow things down, and shifting that into the developer workflow, you’re just slowing them down earlier, right? You’re not enabling them with better decision-making capacity so they can say, “Oh, I now understand the risks that I’m bringing on board by not sanitizing a string before I dump it into a SQL, you know, query. But now I understand that better because Snyk is giving me that information at the right time when I don’t have to context switch out of it, which is, as I’m writing that line of code to begin with.”

Corey: When I look at your website—and I’m really, really hoping that your marketing folks don’t turn me into a liar on this one between the time we have recorded this and the time it sees the light of day in a week or so—it’s notable because you are a security vendor, but you almost wouldn’t know that from your website. And that is a compliment because at no point, start to finish, on the landing page at snyk.io do I
see anything that codes to, “Hackers are coming to kill you. Give us money immediately to protect yourself.”

You’re not slinging FUD. You’re talking entirely about how to improve velocity. The closest it gets to even mentioning security stuff is, “Ship on time with peace of mind.” That is as close as it gets to talking about security stuff. There is no fear based on this, and you don’t treat people like children and say, “Security is extremely important.” “Thank you, Professor, I really appreciate that helpful tip.”

Clinton: Yeah, you know, again, I think we take the very controversial approach that developers are not bad people who want to make applications less secure, right? And I think again, when you go into that 40-year trajectory of that constant tension between the engineering and the security sides of the house, it really involves certain perceptions about what those other people are like: security are bad and want to shut everything down; developers are, you know, wild cowboys who don’t care about standardization and are just introducing a bunch of risk, right? Where Snyk comes in is fundamentally saying, “Hey, we can actually all live together in a world where we recognize there’s pain on both sides?” And look, Corey, I’m coming to you after essentially waking up every day for 20 years and writing code of some kind or other, and I can tell you, developers are already scared enough, man. It is a fearful and anxiety ridden experience to know that you’re not completely in command of what happens to that application once it leaves your IDE, right?

You know at some point you’re going to get that PDF dumped on you; you’re going to have a build block, you’re going to have a bug report come in from a very important customer at three o’clock in the morning and you’re going to have to do something about it. I think every software engineer in the world carries that fear around with them. They don’t have to be told you have the capacity to do bad stuff here and you should be better at it. What they need is somebody to tell them here’s how to do things better, right? Here’s not necessarily even why a cross-site scripting attack is dangerous—although we can certainly educate you on that as well—but here’s what you need to do to remediate it. Here’s how other developers have fixed that in applications that look like yours.

And if you get that intelligence at the right point, then it becomes truly—to go back to your original question—it becomes about solutions rather than about problems, right? The last thing we ever want to do is adopt that traditional approach of saying, “You did a bad thing. It’s your fault. You have to go figure out what to do. And then by the way, you have to do all the refactoring on top of that because we didn’t tell you you did the bad thing until three weeks later when that traditional SaaS tool finally finished running.”

Corey: Exactly. It’s a question of how much can you reduce that feedback loop? If I get pinged 60 seconds after I commit code that there’s a problem with it, great. I still have that in my head. Mostly. I hope. But if it’s six months later it’s, “Who even wrote this?” And I pull up git blame and, “Ah, crap, it was me. What was I possibly thinking back then?” It’s about being able to move rapidly and fix things, I guess, as early in the process as possible, the whole shift-left movement. That’s important. That’s valuable.

Clinton: Yeah, the context switching is so expensive, right, because the minute you switch away from that file, you’re reading some documentation. You’re out of that world. Most of the developer’s time is spent getting into and out of different contexts. Once you’re in there, I mean, you could rattle off 40 lines of code in a sitting and actually clear a ticket and you feel really good about yourself, right? The next day, when that comes back from QA saying you did something wrong here, that’s the painful part of having to get back in.

And by the time you’ve already done that, you’ve doubled the amount of time you’ve spent on that feature. So, it’s all about integrating the right intelligence in the right context at the right time, and doing so in such a way that we’re not throwing around blame, that we’re not saying, “You should have known better.” We’re saying, “We want to help you do this better because, you know, ultimately, you’re going to write another SQL query. That’s okay. We hope that maybe this will inspire you to sanitize those strings properly, and we’re going to give you some suggestions on how to do that.”

Corey: Yeah. Developer time is way more expensive than the infrastructure. That is, I think, a little understood facet of how this works from an engineering perspective because an awful lot of us came up in this industry considering our time to be free. Because we were doing this as a hobby in some cases, it was. When I was in my dorm room back many years ago, as I was basically in the process of being expelled from boarding school, it was very clearly my time was not worth a whole hell of a lot to anyone at that point.

Speaking of expensive things, I want to talk for a minute about your pricing. And what I like about this is, let me be clear here. I am a big fan of taking shortcuts wherever I can, and one of the shortcuts I love doing—and I don’t know if I’ve talked about it on this show before—is when I’m talking to a company and I need to figure out do they know what they’re doing or are they clowns, I cheat and I go to the pricing page. And there are two big things that I look for, and you have them both.

The first is that over on the far left side of the spectrum, it’s do you have a free option? And yes, you do. And, “Click here to get started immediately.” Great because it’s three in the morning, I need to get something done, I’m under a deadline, I do not have time for a conversation with sales, and as an engineer, I absolutely don’t want to deal with that type of sales process because it feels weird to go and ask my boss to go ahead and sign off on something because I feel like my spending authority is capped at $20. Now that I have a little more context, I understand exactly why [laugh] my spending authority was capped at $20 back when I was an engineer.

Clinton: Yeah, exactly right. And so it’s not only that commitment to ensuring every software engineer in the world can have access to Snyk immediately by making one click because, you know, ultimately, we’re committed to that community, right? There’s 3 million developers using Snyk currently. That’s about 10% of all engineers in the world. We’re very proud of that number.

We expect that to continue to grow and I think it shows that there is need out there, right? And if we can enable every engineer who’s up at 3 a.m. faced with some security prospect to say, you know, it is as simple as getting a free account and getting a vulnerability report, getting the remediation advice, being able to sleep easier. I think we’re successful as a company, regardless of what the bottom line is. But when you look at how to scale that into the enterprise, the way security solutions are priced, I mean, it’s like throwing a bunch of wet noodles at the wall and seeing what sticks, right?

Corey: Yes. And that’s the other piece of your pricing that I like is a lot of people are going to be listening to that, what I’m saying right now about, “Oh, well, we have a free tier. Why do you think we’re clowns?” It’s, “Ah. Because the other end is just as important if not more so, which is there has to be an enterprise tier, and the price for that has got to be, ‘Click here to have a conversation.’” And the reason behind that is if you work in procurement, which is very often who’s going to be reaching out on something like this, you are going to need custom contracts; you are going to want a long-term enterprise deal, and if the top tier is X dollars per thing that’s already there, it reeks of unsophisticated vendor to a buyer in that position, and it makes the people a big blue chip companies think, “Oh, they don’t know how to deal with someone at our scale.” Pricing his messaging, and I think people lose sight of that. You absolutely say the right things on both ends. I look at this, and there’s nothing I would change or improve about your pricing page, which to be honest, is really rare.

Clinton: I’m not sure all of our sales leaders would agree with you there, but I will pass that feedback along. Well, and the other thing I would add to that is, what everyone who’s in a pricing conversation wants is predictability about what is this going to be in the future, right? And so we base our pricing on how many developers are in your organization, right? That’s probably a number you know; that’s probably a number that you can predict over time. We’re not going to say, “How many CPUs are we using, right? What’s the footprint of the cloud resources we’re deploying to scan your stuff?” These are all things that you have very little control over and there is alchemy there that introduces a financial risk into that situation. And we’re all about risk mitigation at scale, right?

Corey: You don’t pop up halfway through a cycle of, “Oh, you’ve gone on a hiring spree. Time to go ahead and pay us a bunch more money you didn’t plan for or budget for.” I’ve had vendors pop up a quarter after I signed a deal—repeatedly—and it drives me up a wall because back in my engineering days, it was, great, now I have to spend time on this that I hadn’t planned for; I have to go to my boss and ask for more money, never a great conversation, and as a cherry on top, I get to look like I don’t know how to manage vendors for crap. It’s just everyone is angry about those conversations. And even the salespeople reaching out had the decency to act a little sheepish about having to have that
conversation with me.

Clinton: The best ones do, at least. Well, and on top of that, you know, maybe that tool has been capped so that now your bills are breaking because you went one over your cap, right? So, I—

Corey: Yeah. I love it. When I fail in production. That’s my favorite thing. It’s like, “All right, we’re going to wind up not scanning for security stuff anymore. And if you go five beyond your cap, we’re going to start introducing vulnerabilities.” It’s, “That’s awesome. Just, great plan.” But I’m kidding. I’m kidding. I want to be very clear, I have never heard a whisper of an actual vendor doing that, on purpose anyway.

Clinton: Exactly. Right. And you know, look. We want to make it as easy as possible, and that’s why, for example, we’re on AWS Marketplace. You can use your existing EDP program to, you know, buy Snyk, just as—

Corey: At 50% of your spend on Snyk then winds up counting toward your spend commit, which is always an interesting approach that some people are like, “Ooh. So, we can wind up transferring the money that we’re spending on a vendor to count toward our commit?” But in many cases, it’s how much are you spending on other third-party vendors in this space because you’re getting excited about a few tens of thousands in most cases, and you have a $50 million annual [laugh] commit. What are you doing there, buddy? That’s like trying to become a millionaire via credit card points. It doesn’t usually pan out that way.

Clinton: Fair enough. Yeah. And then look, we’re very proud of that partnership with Amazon. And look if hey, if they can lock some of our customers into $15 million a year spend contracts, we’ll take a few pennies on that, right?

Corey: Oh, yeah, as a vendor, you’d be silly not too. It makes sense. But you’re doing significantly more than that. As of this week being re:Invent week, you are—well, tell me about it.

Clinton: Yeah, Corey, we are thrilled to announce this week that AWS is now integrating with Snyk’s vulnerability database within Amazon Inspector. And this is going to bring the best-of-breed security intelligence with a curated vulnerability database, including all of our proprietary research around things like exploit maturity, reachability, vulnerable conditions, social trends on vulnerabilities, all available within Amazon Inspector to any developer utilizing it. We also have an AWS code pipeline integration that makes it easy for anyone utilizing AWS for your CI/CD to get immediate feedback on vulnerabilities in your applications as they move through that pipeline. And remember, we’re never just going to say, “We’ve identified a vulnerability. Now, you need to figure out what to do with it.” We’re always going to integrate the remediation advice because our audience at the end of the day is the developer whose job it is to make the fix and who has such a wide variety of responsibility these days, the best we can do is say to them, not just, “We found something wrong,” but, “Here’s the solution that we think you should implement to get that secure code back out into production.”

Corey: This episode is sponsored by our friends at CloudAcademy. That’s right, they have a different lab challenge up for you called, “Code Red: Repair an AWS Environment with a Linux Bastion Host.” What does it do? Well, its going to assess your ability to troubleshoot AWS networking and security issues in a production like environment. Well, kind of, its not quite like production because some exec is not standing over your shoulder, wetting themselves while screaming. But..ya know, you can pretend in fact I’m reasonably certain you can retain someone specifically for that purpose should you so choose. If you are the first prize winner who completes all four challenges with the fastest time, you’ll win a thousand bucks. If you haven’t started yet you can still complete all four challenges between now and December 3rd to be eligible for the grand prize. There's only a few days left until the whole thing ends, so I would get on it now. Visit cloudacademy.com/corey. That’s cloudacademy.com/C-O-R-E-Y, for god’s sake don’t drop the “E” that drives me nuts, and thank you again to Cloud Academy for not only promoting my ridiculous non sense but for continuing to help teach people how to work in this ridiculous environment.

Corey: First, congratulations. It’s neat to have a first-party integration like that with an AWS service, as opposed to, you know, their somewhat storied approach of, “Hey, it’s an open-source project. We’re just going to implement something that’s API compatible ourselves, and irritate people.” Now, to be clear, my problem is not that you should expect to build anything and not face competition. My concern is a little bit more along the lines of, “Huh. Why is that same company always the first in line to compete with something.” Which is neither here nor there.

Security is also one of those areas where I think competition is important. You want it continual background level of investment in the space because this stuff is super important. What I like about Snyk and a number of companies in this space is I know exactly where you stand. Let’s contrast that for a second with AWS. You’re integrating with Inspector, which is a great service, but you’re not, I don’t believe, integrating with their other security services such as [big breath in] Amazon Detective, the Audit Manager—if you want to consider that one of them—Amazon Macie, AWS Firewall Manager, AWS Shield, the Network Firewall, IoT Device Defender, CloudTrail, Config.

Amazon Inspector is in one you’re there, but not really Security Hub, or GuardDuty, or IAM itself. And I look at all of these services—I mean, IAM is free, of course, but the rest are very much not—and I do some basic arithmetic and I’m starting to realize that if I can figure all the various AWS security services together and what that’s going to cost me, it turns out the answer is more than the data breach. So, on some level, it’s one of those—at what point is it so confusing and it starts to look like a cross-sell deal between all of the different services, and turn them all on because you could ever have too much security, we still have to ship things eventually. And their security messaging has been extraordinarily confused for a long time. At some level, the fact that you are now integrating with them on the Inspector side means that for the first time, I think I understand what Inspector does now, which is more than a little messed up. But here we are.

Clinton: Indeed. Well, the first thing I would say on that is, you know, stay tuned. As we move into the new year. I think you’re going to see a lot more announcements both, you know, on the AWS side, but also kind of industry-wide and terms of integration with Snyk. That Vulnerability Database feed also, as you mentioned earlier, in use in Docker Hub, so anyone with Containers and Docker Hub can get advantage by scanning with our Snyk container tool.

We have other integrations with Red Hat, for example. And there are actually many other companies utilizing that DB feed to, again, get access to that best in breed vulnerability data. When you talk about that model of, you know, being outcompeted on the security front, I think that’s more difficult to do when you’re actually talking about data, right? Like tooling, on some level—and I might get in trouble for saying this—but tooling is commodity, right? Somebody tomorrow is going to come out with a better tool to do a thing a little bit faster in a little bit more intuitive way. What can’t be easily replicated is the data and intelligence behind that, right? And so that’s why—

Corey: Yeah, the secret sauce that makes you folks work is not the fact of, “Ah, we can fire off or catch a web hook, and then run the following command against the codebase.” That is—sure it’s handy and it’s useful and you’re good at that, but that is not the reason that people become your customer.

Clinton: Exactly right. Look, there’s a lot of tools that can resolve the dependency tree within your open-source application, right? We can do that as well. We leverage a lot of open-source to do that, you know, we’re very open with that. As I mentioned earlier, a lot of Snyk tooling is available on GitHub, you can see how it works, that code is public.

Really the value we’re providing is in that curated security research that our dedicated team is working on day in and day out and verifying public security data that’s out in CVEs. Is this actually accurate? Do we agree with the severity rating? Might there be other factors that could modify that severity rating? What happens when you are scanning an application that might have some vulnerable conditions versus others? Don’t you want to prioritize those vulnerabilities differently? What happens at runtime, right? If you’re deploying an application to an EC2 instance with an OpenSSH ingress into your security group, that’s going to make certain vulnerabilities a lot bigger risk than if you’ve got your IAC configured correctly, right? So, the really the overall mission of Snyk as we move into this broader, kind of, ASPM application, you know, security posture management space, is to say, how many different signals across the SDLC can we combine in intuitive ways for the developer to understand that risk at the right time with the right context and armed with the remediation advice to make a better decision as they’re writing their code, you know, rather than after the fact? If I could sum it all up, kind of, that’s the vision of where we are both today and ultimately where we’re going.

Corey: There also needs to be an understanding of who the customer is. If I go through the launch wizard and spin up in a brand new account, my first EC2 instance, and I spin up an instance by going through the wizard, the first thing it does is yell at me. Because, “Ah, that SSH port is open to the world.” Which you need to get into it, once it’s there. So, it sets that up for me and yells at me all in the same breath. And it’s, this is not a promising start; I kind of need that to get into it.

Conversely, if you’re not someone learning this stuff for the first time, and you’re, oh I don’t know, a production engineer at a bank, you care quite a bit differently in that use case about things like OpenSSH groups, it’s security posture, et cetera, et cetera. An awful lot of the tooling is, “Ah, you’re failing this benchmark, and this benchmark, and this benchmark,” from CIS and the rest of all these rules of, oh, you’re not encrypting your data at rest. Well, it’s in an AWS data center environment. Yeah, if someone could break in and steal the drives from multiple facilities and somehow recombine them together and get out alive, yeah, that’s really not my threat model.

But it’s easy to turn it on and check a box and make an auditor go away. But that’s not where I would spend the bulk of my energies if I’m trying to improve my security posture. And it turns into rote checklists super easily. The thing I’ve always appreciated about the stuff that you’re tooling in the open-source world has highlighted is it’s not nonsense. And I really can’t understate just how valuable that is.

Clinton: Absolutely. And that comes from a combination of signals across that SDLC, from the open-source, from the container, from the proprietary code, from the IAC, but then also what’s happening at runtime, right? Like, how are those containers actually deployed onto EKS? What ports are open? What running binaries are on the container that might influence, you know, what packages you choose to upgrade, versus not?

All of that matters, and what—you know, the issue I think now is getting that visibility to the developer at the right time so that they can make it actionable. And the thing about infrastructure as code, that I think that’s really interesting and not super well understood is a lot of those defaults are really insecure. And developers have no idea, right? Like, they might not be aware that if you don’t define that encryption for your S3 bucket, it’ll happily deploy unencrypted, right? Yes, that’s a compliance problem, but that’s also potentially exacerbator have other vulnerabilities that might be in that application.

But you only see those when you can combine and have a single pane of glass that gives you the runtime signaling plus everything that’s happening in the application, armed with the correct information to actually remediate that at the time, and say, “Don’t you think you wanted to add, you know, AES encryption to this bucket? Don’t you think you wanted to close down port 22?” And also, combine that with your internal business logic, right? Like maybe for an internal only application that never transits beyond your VPC perimeter, sure, it’s fine to have port 22 open, right? There’s just going to be people within your zero-trust environment authenticating to it. But for your production web application, that might be a different story.

Corey: There are other concerns, too. For example, I’m sitting here complaining about the idea of encrypting at rest in an AWS environment, but if you’ve signed customer contracts that state that you’re doing it, you’d better freaking do it, as opposed to, “Well, I know what the actual security risk is and it’s no big deal.” Yeah, don’t make that decision. If you are contractually obligated to do a thing. Don’t YOLO it; do what you say you’re going to do. That’s that whole integrity thing.

Clinton: Oh, sure. And look in a battle between security and compliance. Compliance always wins, right? But from a developer perspective, I don’t know that we on the front lines writing code actually differentiate, right? That certainly is a matter for the people defining the policies and, you know, creating their gating mechanisms in CI to figure out.

What I want to know as a developer is, is my build going to succeed, right? Or am I going to get shut down and get the nastygram that says, you know, “We couldn’t launch this for x, y, and z reason.” Now, everybody on my team hates me, my lead dev is on me, now there’s a bunch of merge conflicts because my branch is behind. I want to get that out into production, but in order to do that, I need information on how are all these signals going to be compiled together in a way that, you know, creates that red light or green light on the risk dashboard later on. But up until I think, you know, relatively recently, I don’t have visibility into that except to launch the commit, you know, start the build and see what happens, and then I have that context-switching problem, right, because it’s hours or days later, that I finally get that signal back.

So yes, I think we have a compliance story to tell from the Snyk perspective as well. A lot of those same issues, you know, we’re detecting, especially with regard to infrastructure as code, but it ultimately is up to various parts of the organization to work together and say, “What balance do we want to strike between security and velocity,” right? Understanding that those are not mutually opposed. What we need is tooling and more importantly a culture that takes both into account and allows us to develop securely and fast at the same time.

Corey: I want to thank you so much for taking the time to speak with me about all this. If people want to learn more, where can they find you? And for God’s sake, please don’t say in your booth at re:Invent.

Clinton: [laugh]. I will not be at re:Invent this year. I’ve had a little bit too much of the Vegas Strip here recently.

Corey: No, I hear you. Right now, the people going are those whose employers find them expendable, which is why I’m there.

Clinton: I wouldn’t say that Corey. I think you’ll do great, and you know, just make sure to bank all your vacation for a couple weeks after. Look, come to snyk.io start a conversation, but more importantly, just start using it, right?

I don’t want to give you the sales pitch; I want you to see the value in the tooling, and the easiest way to do that as an engineer is just to start using it. And if there is value there, you want to bring it to your enterprise. I would love to have that conversation and move forward. But engineer to engineer, like, figure out if this is going to work for you: does it make your life easier? Does it reduce the pain and anxiety you feel before making that commit into the production branch? And if so, then yeah, we’d love to talk.

Corey: I will, of course, put links to that in the [show notes 00:33:22]. Thank you so much for speaking to me today. I really appreciate it.

Clinton: Thank you, Corey. Glad to do it.

Corey: Clinton Herget, principal solutions engineer at Snyk. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry comment yelling at Snyk about how they’re a terrible company because they continually refuse to patronize your side business down at the Vowel Emporium.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Brian

Brian is an accomplished dealmaker with experience ranging from developer platforms to mobile services. Before InfluxData, Brian led business development at Twilio. Joining at just thirty-five employees, he built over 150 partnerships globally from the company’s infancy through its IPO in 2016. He led the company’s international expansion, hiring its first teams in Europe, Asia, and Latin America. Prior to Twilio Brian was VP of Business Development at Clearwire and held management roles at Amp’d Mobile, Kivera, and PlaceWare.

Links:

  • InfluxData: https://www.influxdata.com

TranscriptAnnouncer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by my friends at ThinkstCanary. Most companies find out way too late that they’ve been breached. ThinksCanary changes this and I love how they do it. Deploy canaries and canary tokens in minutes and then forget about them. What's great is the attackers tip their hand by touching them, giving you one alert, when it matters. I use it myself and I only remember this when I get the weekly update with a “we’re still here, so you’re aware” from them. It’s glorious! There is zero admin overhead to this, there are effectively no false positives unless I do something foolish. Canaries are deployed and loved on all seven continents. You can check out what people are saying at canary.love. And, their Kub config canary token is new and completely free as well. You can do an awful lot without paying them a dime, which is one of the things I love about them. It is useful stuff and not an, “ohh, I wish I had money.” It is speculator! Take a look; that’s canary.love because it's genuinely rare to find a security product that people talk about in terms of love. It really is a unique thing to see. Canary.love. Thank you to ThinkstCanary for their support of my ridiculous, ridiculous nonsense.

Corey: Writing ad copy to fit into a 30 second slot is hard, but if anyone can do it the folks at Quali can. Just like their Torque infrastructure automation platform can deliver complex application environments anytime, anywhere, in just seconds instead of hours, days or weeks. Visit Qtorque.io today and learn how you can spin up application environments in about the same amount of time it took you to listen to this ad.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. This promoted guest episode is brought to us by our friends at InfluxData. And my guest is titled as the Chief Marketing Officer at InfluxData, and I don’t even care because his bio has something absolutely fascinating that I want to address instead. Brian Mullen is an accomplished dealmaker is how the bio starts. And so many of us spend time negotiating deals, but so few people describe ourselves in that way. First, Brian, thank you for joining us. And secondly, what’s up with that?

Brian: [laugh]. Well, thanks, Corey, very excited to be here. And yes, dealmaker; I guess that would be apropos. How did I get into marketing? Well, a lot of my career is spent in business development, and so I think that’s where the dealmaker part comes from.

Several different roles, including my first role at Influx—when I joined Influx—was in business development and partnerships. And so, prior to coming to Influx, I spent many years building out the business development team at Twilio, growing that up, and we did a lot of deals with carriers, with Cloud partners, with all kinds of different partners; you name it, we worked with them. And then moving into Influx, joined in an BD capacity here and had a couple different roles that eventually evolved to Chief Marketing Officer. But that’s where the dealmaker comes from. I like to do deals, it’s always nice to have one on the side in whatever capacity you’re working in, it’s nice to have a deal or two working on the side. It kind of keeps you fresh.

Corey: It’s fun because people think, “Oh, a deal. You’re thinking of mergers and acquisitions, and how hard could that be? You just show up with a bag of money and give it to people and then you have a deal closed.” And oh, if only it were that simple. Every client engagement we have on the consulting side has been a negotiation back and forth, and the idea is to ideally get everyone to the point where they’re happy, but honestly, if everyone’s slightly unhappy but can live with the result, we’ll take that too.

And as people go through their own careers it’s, you’re always trying to make a deal in some form: when you try to get a project approved, or you’re trying to get resources thrown at something—by which I generally mean money, not people, though people, too—it’s something that isn’t necessarily clearly understood or discussed very often, despite the fact that half of what I do is negotiating with AWS on behalf of clients for better contractual terms. The thing that I think takes people by surprise the most is that dealmaking is almost never about pounding the table, being angry, and walking out, like you read the world’s worst guide to buying a car or something. It’s about finding the win for everyone. At least that’s the way I’ve always approached it.

Brian: That’s a good point. And actually that wording that you described of finding a win for everybody, that’s how I always thought about it. I think about it as first of all, you’re trying to understand what the other party—and it could be an individual, it could be a company, it could be a group of companies, sometimes—you’re trying to understand what their goals are, what their agenda is and see how that matches with your own; sometimes they’re opposing, sometimes they’re overlapping. And then everyone has to have some perceived win in a deal. And it’s not competitively; it’s more like you just have to have value, that is kind of what the win is – having value in that deal.

And so that’s the way I always approached it. And doing deals, whether you’re in BD or sales, or if you’re working with vendors and you’re in a different functional role, sometimes it’s not even commercial, it’s just about aligning resources, perhaps. Our deal might be that you and I are both going to put a collective effort into building something or taking something to market. In another scenario might be like, I’m going to pay for this service that you’re delivering, or vice versa. Or we’re going to go and bring two revenue-generating products together and take them to market. Whatever it might be, it doesn’t matter so much what the mechanics are of the deal, but it’s usually about aligning those agendas and in having someone get utility, get value on the other side.

Corey: I think that people lose sight of the fact as well, that when you’re talking about a service provider—and let’s be clear, InfluxData has launched a cloud platform that we’ll talk about in a minute—this is not the one-off transactional relationship; once the deal is signed, you’ve got to work with these people. When they host parts of your production infrastructure, whether you want to admit it or not they’re your partner more so than they are your vendor. It has to be an ongoing relationship that people are, if they at least aren’t thrilled with it, can at least be happy enough to live with, otherwise it just winds up with this growing sense of resentment and it just sort of leads nowhere.

Brian: Yeah, there really is no deal moment. Yes, people sign agreements with companies, but that’s just the very beginning. Your relationship evolves from there. We’re delivering a product, we’re delivering this platform that handles time-series data to our customers, and we’re asking them to trust us with their product that they’re taking out to market. They’re asking us to handle their data and to deliver service to them that they’re turning into their production applications. And so it’s a big responsibility. And so we care about the relationship with our customers to continue that.

Corey: So, I first really became aware of time-series data a few years back during a re:Invent keynote when they pre-announced Timestream, which took entirely too long to come to market. Okay, great. So, you’re talking about time-series data. Can you explain what that means in simple terms? And I learned over the next eight minutes that they were talking about it, that no, no, they couldn’t. I wound up more confused by the end of the announcement than I was at the beginning.

So, assuming that I have the same respect for databases as you would expect for someone whose favorite data store is Route 53—because you can misuse it as a beautiful database—what is time-series data and why does it matter in 2021?

Brian: Sure, it's a good question. And I was there in that audience as well that day. So, we think of time-series data as really any type of data that’s stamped in time, in some way. It could be every hour, every minute, every second, every half second, whatever. But more specifically, it’s any type of data that is generated by some source—and that could be a sensor sources within systems or an actual application—and these things change over time, and then therefore, stamped in time in some way.

They can come at different frequencies, like I said, from nanoseconds to seconds, or minutes and hours, but the most important thing is that they usually trigger a workflow, trigger some sort of action. And so that’s really what our platform is about. It allows people to handle this type of data and then work with it from there in their applications, trigger new workflows, et cetera. Because the historical context of what happens is super important.

And when we talk about sources, it could be really many things. It could be in physical spaces, and we have a lot of IoT types of customers and use cases. And those are things like devices and sensors on the factory floor, out in the field, it’s on a vehicle. It’s even in space, believe it or not. There are customers that are using us on satellites.

And then it can also be sources from within software, applications, and infrastructure, things like VMs, and containers, and microservices, all emitting time-series data. And it could be applications like crypto, or financial, or stock market, agricultural type of applications that are themselves as applications emitting data. So, you think about all these sources that are out there from the physical world to the virtual world, and they’re all generating time-series data, and our platform is really specially designed to handle that kind of data. And we can get into some details of what exactly that means, but that’s really why we’re here. That’s what time-series is all about.

Corey: And this is the inherent challenge I think we’re seeing across the entire industry slash ecosystem. I mean, this is airing during re:Invent week, but at the time we are recording this, we have not yet seen the Tuesday keynote that Adam Selipsky will take to the stage, and no doubt, render the stat I’m about to throw at you completely obsolete. But depending on how you count them, there’s somewhere between 13 and 15 managed database or database-like services today that AWS offers. And they never turn things off and they’re always releasing new things, supposedly on behalf of customers; in practice because someone somewhere wants to get promoted by launching a new service; good for them. Godspeed.

If we look into the uncertain future, at some point, someone’s job is going to be disambiguating between the 40 different managed database services that AWS offers and picking the one that works. What differentiates time-series from—let’s just start with an easy one—something like MySQL or Postgres—or ‘Postgres-squeal’ is how I insist on pronouncing that one. Let’s stay away from things like Neptune because no one knows what a social graph database is and I assure you, you almost certainly don’t need one. Where does something like Influx work in a way that, “Huh. Running this on MySQL is really starting to suck.”

Brian: When and why is it time to consider a specialized tool. And in fact, that’s actually what we see a lot with our customers is coming to us around that time when a time-series is a problem to solve for them is reaching the point where they really need a specialized tool that’s kind of built for that. And so one way to look at that is really just to think about time-series in general as a type of data. It’s rapidly rising. It’s the fastest growing data category out there right now.

And the reason for that is it’s being driven by two big macro trends. One is the explosion of all these applications and services running in the cloud. They’re expanding horizontally, they’re running in more regions, they’re in many cases running on multiple clouds, and so it’s just getting big—the workloads are getting bigger and bigger. And those are emitting time-series data. And then simultaneously, you have this growth of all these devices and sensors that are coming online out in the real world: batteries, and temperature gauges, and all kinds of stuff, both new and old, that is coming online, and those sources are generating a lot of time-series data.

So typically, we’re in a moment now, where a lot of developers are faced with this massive growth of time-series data. And if you think about some data set that you have, that you’re putting into some kind of traditional database, now add the component of time as a multiplier by all the data you have. Instead of that one data, that one metric, you’re now looking at doing that every one second in perpetuity. And so it’s just an order of magnitude more data that you’re dealing with. And then you also have this notion of—when you have that magnitude of data, you have fidelity, you’re taking a lot of it in at the same time, I mean, very quickly, so you have batch or stream data coming in at super high volume, and you may need that for a few minutes or a few hours or days, but maybe you don’t need it for months and years.

And so you’d maybe dropped down to kind of a lower fidelity for the longer-term. But you really have this toggling back and forth of the high fidelity and low fidelity, all coming at you at pretty high volume. And so typically what happens is, is when the workloads get big enough, the legacy tools, they’re just not equipped to do it. And a developer—if they have a small set of time-series they’re dealing with, what is the first thing they’re going to do? They’re going to look around and be like, “Hey, what do I have here? Oh, I’ve got Mongo over here. I’ve got Splunk, or I’ve got this old relational database, I can put it in.”

And that’s typically what they’ll do, and that works fine until it doesn’t. And then that’s when they come around looking for a specialized tool. So, we really sit in Influx and, frankly, other time-series products really do sit at that point where people are considering a specialized tool just because the workload has gotten such that it requires that.

Corey: Yeah. Taking a look at most of the offerings in the space; anything that winds up charging anything more than a very tiny fraction of a penny—from what you’re describing—is going to quickly become non-economical, where it’s, “Oh, we’re going to charge you”—like using S3: every, I think, 1000 writes cost a penny—“Oh, we’re just going to use S3 for this.” Well, at some of these data volumes, that means that your request charge on S3 is very quickly going to become the largest single line item in your bill, which is nothing short of impressive in a lot of cases, but it also probably means that you’ve taken a very specific tool—like an iPad—and tried to use it as something else—like a hammer—and no one’s particularly happy with that outcome.

Brian: Yeah. First of all, having usage-based pricing is really important. We think about it as allowing people to have the full version of the product without a major commitment, and be using it in test scenarios and then later in the very early production scenarios. But as a principle, it’s important for people that just signed up two hours ago using your product are basically using the same full product that the biggest customers that you have are using that are paying many, many thousands or tens of thousands per month. And so the way to do that is to offer usage-based pricing and not force people to commit to something before they’re ready to do it.

And so there’s ways to unlock lower pricing, and we, like a lot of companies, offer annual pricing and we have a sales team that worked with folks to basically draw down their unit costs on the use of the platform once they kind of get comfortable with their workload. So, there’s definitely avenues to get lower price, and we’re believers in that. And we also want to, from a product development perspective, try to make the product more efficient. And so we basically are trying to drive down the costs through efficiencies in the product: make it run faster, make queries take less time, and also ship products on top of it that require developers to write less code themselves, kind of, do more of the work for them.

Corey: One of the things I find particularly compelling about what you’ve done is it is an open-source project. If I want to go ahead and run some time-series experiments myself, I can spin it up anywhere I want and run it however I see fit. Now, at some point, if I’m doing this for anything more than, “Oh, let’s see how I can misuse this today,” I probably want to at least consider letting someone who’s better at running these things than I am take it over. And as I’m looking through your customer list, the thing that strikes me is how none of these things are quite like the other. We’re talking about companies like Hulu is probably not using it the same way as Capital One is, at least I certainly hope not. You have Texas Instruments; you also have Adobe. And it sort of runs an entire gamut of none of these companies quite look alike; I have to imagine their use cases are also somewhat varied, too.

Brian: Yeah, that’s right. And we really do see as a platform, and with time-series being the common problem that people are looking to solve, we see this pretty broad set of use cases and customer types. And we have some more traditional customers like the Cisco’s and the IBM’s of the world, and then some relatively new folks like Tesla and Hulu and others that are a little bit more recent. But they’re all trying to solve the same fundamental problem with time-series, which is “How can I handle it in an efficient way and make use of it meaningfully in my applications and services?”

And we were talking earlier about having some sources of time-series data being in, kind of a virtual space, like in infrastructure and software, and then some being in physical space, like in devices and sensors out in the real world. So, we have breadth in that way, too. We have folks who are building big software observability infrastructure solutions on us, and we also have people that are pulling data off of the devices on a solar panel that’s sitting on a house in the emerging world, right? So, you have basically these two far ends of the spectrum, but all using this specialized tool to handle the time-series data that they’re generating.

Corey: It seems to me that for most of these use cases and the way you describe it, it’s more about the overall shape of the data when we’re talking about time-series more so than it is any particular data point in isolation. Is that accurate, or are there cases where that is very much not the case?

Brian: I think that’s accurate. What people are mostly trying to understand is context for what’s happening. And so it’s not necessarily—to your point—not searching for one specific data point or moment, but it’s really understanding context for some general state that has changed or some trend that has emerged, whatever that might be, and then making sense of that, and then taking action on that. And taking an action could mean a couple of different things, too. It could be in an observability sense, where somebody in an operator type of mode where they’re looking at dashboards and paying attention to infrastructure that’s running and then need to take some sort of action based on that. It also, in many cases, is automated in some way: it’s either some series of automated responses to some state that is reached that is visible in the data, or is actually kicking off some new series of tasks or actions inside of an application based on what is occurring and shown by the time-series data.

Corey: You know what doesn’t add to your AWS bill? Free developer security from Snyk. Snyk is a frictionless security platform that meets developers where they are, finding and fixing vulnerabilities right from the CLI, IDEs, repos, and pipelines. And Snyk integrates seamlessly with AWS offerings like CodePipeline, EKS, ECR, and oh so much more.

Secure with Snyk and save some loot. Learn more at snyk.io/scream. That’s S-N-Y-K-dot-I-O/scream

Corey: So, we’ve talked about, you have an open-source product, which is the sort of thing that most people listening to this should have a vague idea of, “Oh, that means I can go on GitHub and download it and start using it, if it’s not already in my package manager.” Great. You also have the enterprise offering, which is more or less, I presume, a supported distribution of this—for lack of a better term—that you then wind up providing blessed configurations thereof and helping run support for that—for companies that want to run it on-prem. Is that directionally accurate, or am I grossly mischaracterizing [laugh] what your enterprise offering is?

Brian: Directionally accurate, of course. You could have a great job in marketing. I really think you could.

Corey: Oh, you know, I would argue, on some level, I probably do. The challenge I have is that I keep conflating marketing with spectacle and that leads down to really unfortunate, weird places. But one additional area, which is relatively recent since the last time I spoke with Paul—one of the cofounders of your company—on this show is InfluxDB Cloud, which is one of those, “Oh, let me see if I look—if I’m right.” And sure enough, yeah, you wind up managing the infrastructure for us and it becomes a pay-per consumption model the way that most cloud service providers do, without the really obnoxious hidden 15 levels of billing dimensions.

Brian: Yes, we are trying to bring the transparency back. But yes, you’re correct. We have open-source and we have—it’s very popular—we have over 500,000-plus instances of that deployed globally today in the community. And that’s typically very common for developers to get started using the open-source, easily recognizable, it’s been out for a long time, and so many people start the journey there.

And then we have InfluxDB Enterprise, which it’s actually a clustered version of InfluxDB open-source. So, it allows you to basically handle in an environment that you want to manage yourself, you manage a cluster and scale it out and handle ever-increasing workloads and have things like redundancy and replication, et cetera. But that’s really specifically for people who want to deploy and operate the software themselves, which is a good set of people; we have a lot of folks who have done that. But one of the areas that’s a little bit more recent is InfluxDB Cloud, which is really, for folks who don’t want to have anything to do with the management; they really just want to use it as a service, send their data in—

Corey: Yeah, give me an API endpoint, and I want you to worry about the care, and the feeding, and the waking up at two in the morning when a disk starts filling up. Yeah, that is the best kind of problem from my perspective: someone else’s.

Brian: Exactly. That’s our job. And increasingly, we’ve seen folks gravitate to that. We’ve got a lot of folks have signed up on this product since it launched in 2019, and it’s really increasingly where they begin their journey, maybe not even going to the open-source just going directly to this because it’s relatively simple to get started.

It’s priced based on usage. People pay for three vectors: they have the amount of data in; they have number of queries made against the platform; and then storage, how much data you have and for how long. And depending on the use case, some people keep it around for relatively short time, like a few days or a couple of weeks. Other folks have it for many, many months and potentially years in some places. So, you really have that option.

But I would say the three products are really about how you want to run it. Do you care about running the, kind of, underlying infrastructure and managing it or do you just want to hit an endpoint, as you said.

Corey: You launched this, I want to say in 2019, which feels about directionally right. And I know it was after Timestream was announced, so I just want to say first, how kind and selfless it was of you to validate AWS’s market, which is, you know how they always like to clarify and define what they’re doing when they decide to enter every single market anywhere to compete with everyone. It turns out, I don’t get the sense that they like it quite [laugh] as much being on the other side of that particular divide, but that’s the best kind of problem, too: again, someone else’s.

Brian: Yeah, I think that’s really true.

Corey: The challenge that I have is that it seems like a weird direction to go in as a company, though it is clearly based upon a number of press releases you have made about the success and market traction that you found, it feels, on some level, like it is falling into an older version of an open-source trap of assuming that, “Well, we wrote the software therefore we are the best people you could pick to run it.” That was what a lot of companies did; it turns out that AWS has this operational excellence, as they call it, and what the rest of us call burning through people and making them wake up in the middle of the night to fix things before it becomes customer-visible. But from the outside, there’s no difference. It seems, however, that you have built something that is clearly resonating, and in a big way, in a way that—I’ve got to be direct with you—the AWS time-series service that they are offering has not been finding success.

Brian: Thank you for saying that, and we feel pretty excited about the success we’ve had even being in the same market as Amazon. And Amazon does a phenomenal job at running products at scale, and the breadth that they have in their product lineup is pretty impressive, especially when they roll out new stuff at AWS re:Invent every year. But we’ve been able to find some pretty good success with our approach, and it’s based on a couple of things. So, one is being the company that actually develops and still deploys the open-source is really important. People gravitate to that.

Our roots as a company are open-source, we’ve been a part of and fostered this community over many, many years, and there’s a certain trust in the direction that we’re taking the company. And Paul, our founder who you mentioned, he’s been front and center with that community, pretty deeply engaged for many, many years. I think that carries a lot of weight. At least that’s the way we think about it. But then as far as commercial products go, we really think about it as going to where our customers are, going to where developers are. And that could mean the language that they prefer, the language of preference for them. And that could [crosstalk 00:22:25]—

Corey: Oh, and it’s very clear; it seems that most database companies that I talk to—again, without naming names—tend to focus on the top-down sale, but I’ve never worked in an environment where the database that will be used was dictated by anyone other than the application developers who are the closest to the technical requirements for the workload. I’ve never understood this model of, “Oh, we’re going to talk to the C suite because we believe that they’re going to pick a database vendor based upon who has box seats this season.” I’ve never gotten that and that probably means I’m a terrible enterprise marketer, on some level. But unlike almost every other player in the database space, I’ve never struggled to understand what the hell your messaging has meant, other than the technical bits that I just don’t have quite enough neurons to bang together to create sparks to fully understand. It is very clearly targeted at a builder rather than someone who’s more or less spending their entire life in meetings. Which, oh, God, that’s me.

Brian: [laugh]. Yes, it’s very much the case. We are focused on the developer. And that developer is a builder of an application or service that is seeing the light of day, it’s going out and being used by their own end-users and end-customers.

And so we care about going to where those developers are, and that could mean going and making your product easily used in the language and tool that customer cares about. So, if you’re a Python developer, it’s important for us to have tools and make it easy for Python developers. We have client libraries for Python, for example. It also means going to the cloud where your customers are. And this is something that differentiates us as well, when you start looking at what the other cloud providers are offering, in that data—like it or not—has gravity. And so somebody that has built their whole stack on AWS and sure they care about using a service that is going to receive their data, and that also being in AWS, but—

Corey: It has to live where the customers are, especially with data egress charges being what they are, too.

Brian: Exactly.

Corey: And data gravity is real. The cloud provider people pick is the one where their data lives because of that particular inflection in the market.

Brian: Absolutely true. And so that’s great if you’re only going after people who are on AWS, but what about Google Cloud and what about Microsoft Azure? There are a lot of developers that are building on those platforms as well, and that’s one of the reasons we want to go there as well. So, InfluxDB Cloud is a multi-cloud offering, and it’s equal experience and capability and pricing on each of the three major clouds. You can buy directly from us; you can put it on any of your cloud bills in one of those marketplaces, and to us that’s like a really, really fundamental point is to bring your product and make it as easy to use on those platforms and in those languages, and in those realms and use cases where people are already working.

Corey: I’m a big believer in multi-cloud for the use case you just defined. Because I know I’m going to get letters if I don’t say this based upon my public multi-cloud is a dumb default worst practice for most folks—because it is, on a workload-by-workload basis—but you’re building a service that has to be close to where your customers are and for that specific thing, yeah, it makes an awful lot of sense for you to have a presence across all the different providers. Now, here’s the $64,000 question for you: is the experience as an InfluxDB Cloud customer meaningfully different between different providers?

Brian: It’s not. We actually pride ourselves on it being the same. Using InfluxDB, you sign up for InfluxDB Cloud, you come in, you set up your account, create your organization, and then you choose which underlying cloud provider you want your account to be provisioned in. And so it actually comes as a secondary choice; it’s not something that is gated in the beginning, and that allows us to deliver a uniform experience across the board. And you may in a future use case, maybe somebody wants to have part of what they’re building data living in AWS and maybe part of it living in Azure, I mean, that could be a scenario as well.

However, typically what we’ve seen—and you’ve probably seen this as well—is most developers are—and organizations—are building mostly on one cloud. I don’t see a lot of multi-cloud in that organization. But we ourselves need to be multi-cloud in order to go to where those people are working. And so that’s the distinction. It’s for us as a company that delivers product to those people, it’s important for us to go where they are, whereas they themselves are not necessarily running on all three cloud products; they’re probably running on one platform.

Corey: Yeah. On a workload-by-workload basis, that’s what generally makes sense. Anytime you have someone who has a particular workload that needs to be in multiple providers, okay, great, you’re going to put that out there, but their backend systems, their billing, their marketing, all the rest, is not going to go down that path for a variety of excellent reasons, mostly that it is a colossal pain, and a bunch of, more or less, solving the same problems over and over, rather than the whole point of cloud being to make it someone else’s. I want to thank you for taking so much time to speak to me about how you’re viewing the evolution of the market, how you’re seeing your move into cloud, and how you’re effectively targeting folks who can actually care about the implementation details of a database rather than, honestly, suits. If people want to learn more, where can they find you?

Brian: They can go to our website; it’s the easiest place to go. So, influxdata.com. You can read all about InfluxDB, it’s a pretty easy sign up to get underway. So, I recommend that people get their hands dirty with the product. That’s the easiest way to understand what it’s all about.

Corey: And if you do end up doing that, please tell them I sent you because the involuntary flinch whenever people mention my name to vendors is one of my favorite parts of being me. Brian, thank you so much for being so generous with your time. I appreciate it.

Brian: Thanks so much for having us on. It was great.

Corey: Brian Mullen, Chief Marketing Officer—and dealmaker—at InfluxData. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with a long, angry comment telling me that you work on the Timestream service team, and your product is the best. It’s found huge success, but I’ve just never met any of your customers and I can’t because they all live in Canada.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Thomas

Thomas Hazel is Founder, CTO, and Chief Scientist of ChaosSearch. He is a serial entrepreneur at the forefront of communication, virtualization, and database technology and the inventor of ChaosSearch's patented IP. Thomas has also patented several other technologies in the areas of distributed algorithms, virtualization and database science. He holds a Bachelor of Science in Computer Science from University of New Hampshire, Hall of Fame Alumni Inductee, and founded both student & professional chapters of the Association for Computing Machinery (ACM).

Links:

  • ChaosSearch: https://www.chaossearch.io

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by my friends at ThinkstCanary. Most companies find out way too late that they’ve been breached. ThinksCanary changes this and I love how they do it. Deploy canaries and canary tokens in minutes and then forget about them. What's great is the attackers tip their hand by touching them, giving you one alert, when it matters. I use it myself and I only remember this when I get the weekly update with a “we’re still here, so you’re aware” from them. It’s glorious! There is zero admin overhead to this, there are effectively no false positives unless I do something foolish. Canaries are deployed and loved on all seven continents. You can check out what people are saying at canary.love. And, their Kub config canary token is new and completely free as well. You can do an awful lot without paying them a dime, which is one of the things I love about them. It is useful stuff and not an, “ohh, I wish I had money.” It is speculator! Take a look; that’s canary.love because it's genuinely rare to find a security product that people talk about in terms of love. It really is a unique thing to see. Canary.love. Thank you to ThinkstCanary for their support of my ridiculous, ridiculous non-sense.

Corey: This episode is sponsored in part by our friends at Vultr. Spelled V-U-L-T-R because they’re all about helping save money, including on things like, you know, vowels. So, what they do is they are a cloud provider that provides surprisingly high performance cloud compute at a price that—while sure they claim its better than AWS pricing—and when they say that they mean it is less money. Sure, I don’t dispute that but what I find interesting is that it’s predictable. They tell you in advance on a monthly basis what it’s going to going to cost. They have a bunch of advanced networking features. They have nineteen global locations and scale things elastically. Not to be confused with openly, because apparently elastic and open can mean the same thing sometimes. They have had over a million users. Deployments take less that sixty seconds across twelve pre-selected operating systems. Or, if you’re one of those nutters like me, you can bring your own ISO and install basically any operating system you want. Starting with pricing as low as $2.50 a month for Vultr cloud compute they have plans for developers and businesses of all sizes, except maybe Amazon, who stubbornly insists on having something to scale all on their own. Try Vultr today for free by visiting: vultr.com/screaming, and you’ll receive a $100 in credit. Thats v-u-l-t-r.com slash screaming.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. This promoted episode is brought to us by our friends at ChaosSearch.

We’ve been working with them for a long time; they’ve sponsored a bunch of our nonsense, and it turns out that we’ve been talking about them to our clients since long before they were a sponsor because it actually does what it says on the tin. Here to talk to us about that in a few minutes is Thomas Hazel, ChaosSearch’s CTO and founder. First, Thomas, nice to talk to you again, and as always, thanks for humoring me.

Thomas: [laugh]. Hi, Corey. Always great to talk to you. And I enjoy these conversations that sometimes go up and down, left and right, but I look forward to all the fun we’re going to have.

Corey: So, my understanding of ChaosSearch is probably a few years old because it turns out, I don’t spend a whole lot of time meticulously studying your company’s roadmap in the same way that you presumably do. When last we checked in with what the service did-slash-does, you are effectively solving the problem of data movement and querying that data. The idea behind data warehouses is generally something that’s shoved onto us by cloud providers where, “Hey, this data is going to be valuable to you someday.” Data science teams are big proponents of this because when you’re storing that much data, their salaries look relatively reasonable by comparison. And the ChaosSearch vision was, instead of copying all this data out of an object store and storing it on expensive disks, and replicating it, et cetera, what if we queried it in place in a somewhat intelligent manner?

So, you take the data and you store it, in this case, in S3 or equivalent, and then just query it there, rather than having to move it around all over the place, which of course, then incurs data transfer fees, you’re storing it multiple times, and it’s never in quite the format that you want it. That was the breakthrough revelation, you were Elasticsearch—now OpenSearch—API compatible, which was great. And that was, sort of, a state of the art a year or two ago. Is that generally correct?

Thomas: No, you nailed our mission statement. No, you’re exactly right. You know, the value of cloud object stores, S3, the elasticity, the durability, all these wonderful things, the problem was you couldn’t get any value out of it, and you had to move it out to these siloed solutions, as you indicated. So, you know, our mission was exactly that, transformed customers’ cloud storage into an analytical database, a multi-model analytical database, where our first use case was search and log analytics, replacing the ELK stack and also replacing the data pipeline, the schema management, et cetera. We automate the entire step, raw data to insights.

Corey: It’s funny we’re having this conversation today. Earlier, today, I was trying to get rid of a relatively paltry 200 gigs or so of small files on an EFS volume—you know, Amazon’s version of NFS; it’s like an NFS volume except you’re paying Amazon for the privilege—great. And it turns out that it’s a whole bunch of operations across a network on a whole bunch of tiny files, so I had to spin up other instances that were not getting backed by spot terminations, and just firing up a whole bunch of threads. So, now the load average on that box is approaching 300, but it’s plowing through, getting rid of that data finally.

And I’m looking at this saying this is a quarter of a terabyte. Data warehouses are in the petabyte range. Oh, I begin to see aspects of the problem. Even searching that kind of data using traditional tooling starts to break down, which is sort of the revelation that Google had 20-some-odd years ago, and other folks have since solved for, but this is the first time I’ve had significant data that wasn’t just easily searched with a grep. For those of you in the Unix world who understand what that means, condolences. We’re having a support group meeting at the bar.

Thomas: Yeah. And you know, I always thought, what if you could make cloud object storage like S3 high performance and really transform it into a database? And so that warehouse capability, that’s great. We like that. However to manage it, to scale it, to configure it, to get the data into that, was the problem.

That was the promise of a data lake, right? This simple in, and then this arbitrary schema on read generic out. The problem next came, it became swampy, it was really hard, and that promise was not delivered. And so what we’re trying to do is get all the benefits of the data lake: simple in, so many services naturally stream to cloud storage. Shoot, I would say every one of our customers are putting their data in cloud storage because their data pipeline to their warehousing solution or Elasticsearch may go down and they’re worried they'll lose the data.

So, what we say is what if you just said activate that data lake and get that ELK use case, get that BI use case without that data movement, as you indicated, without that ETL-ing, without that data pipeline that you’re worried is going to fall over. So, that vision has been Chaos. Now, we haven’t talked in, you know, a few years, but this idea that we’re growing beyond what we are just going after logs, we’re going into new use cases, new opportunities, and I’m looking forward to discussing with you.

Corey: It’s a great answer that—though I have to call out that I am right there with you as far as inappropriately using things as databases. I know that someone is going to come back and say, “Oh, S3 is a database. You’re dancing around it. Isn’t that what Athena is?” Which is named, of course, after the Greek Goddess of spending money on AWS? And that is a fair question, but to my understanding, there’s a schema story behind that does not apply to what you’re doing.

Thomas: Yeah, and that is so crucial is that we like the relational access. The time-cost complexity to get it into that, as you mentioned, scaled access, I mean, it could take weeks, months to test it, to configure it, to provision it, and imagine if you got it wrong; you got to redo it again. And so our unique service removes all that data pipeline schema management. And because of our innovation because of our service, you do all schema definition, on the fly, virtually, what we call views on your index data, that you can publish an elastic index pattern for that consumption, or a relational table for that consumption. And that’s kind of leading the witness into things that we’re coming out with this quarter into 2022.

Corey: I have to deal with a little bit of, I guess, a shame here because yeah, I’m doing exactly what you just described. I’m using Athena to wind up querying our customers' Cost and Usage Reports, and we spend a couple hundred bucks a month on AWS Glue to wind up massaging those into the way that they expect it to be. And it’s great. Ish. We hook it up to Tableau and can make those queries from it, and all right, it’s great.

It just, burrr goes the money printer, and we somehow get access and insight to a lot of valuable data. But even that is knowing exactly what the format is going to look like. Ish. I mean, Cost and Usage Reports from Amazon are sort of aspirational when it comes to schema sometimes, but here we are. And that’s been all well and good.

But now the idea of log files, even looking at the base case of sending logs from an application, great. Nginx, or Apache, or [unintelligible 00:07:24], or any of the various web servers out there all tend to use different logging formats just to describe the same exact things, start spreading that across custom in-house applications and getting signal from that is almost impossible. “Oh,” people say, “So, we’ll use a structured data format.” Now, you’re putting log and structuring requirements on application developers who don’t care in the first place, and now you have a mess on your hands.

Thomas: And it really is a mess. And that challenge is, it’s so problematic. And schemas changing. You know, we have customers and one reasons why they go with us is their log data is changing; they didn’t expect it. Well, in your data pipeline, and your Athena database, that breaks. That brings the system down.

And so our system uniquely detects that and manages that for you and then you can pick and choose how you want to export in these views dynamically. So, you know, it’s really not rocket science, but the problem is, a lot of the technology that we’re using is designed for static, fixed thinking. And then to scale it is problematic and time-consuming. So, you know, Glue is a great idea, but it has a lot of sharp [pebbles 00:08:26]. Athena is a great idea but also has a lot of problems.

And so that data pipeline, you know, it’s not for digitally native, active, new use cases, new workloads coming up hourly, daily. You think about this long-term; so a lot of that data prep pipelining is something we address so uniquely, but really where the customer cares is the value of that data, right? And so if you’re spending toils trying to get the data into a database, you’re not answering the questions, whether it’s for security, for performance, for your business needs. That’s the problem. And you know, that agility, that time-to-value is where we’re very uniquely coming in because we start where your data is raw and we automate the process all the way through.

Corey: So, when I look at the things that I have stuffed into S3, they generally fall into a couple of categories. There are a bunch of logs for things I never asked for nor particularly wanted, but AWS is aggressive about that, first routing through CloudTrail so you can get charged 50-cent per gigabyte ingested. Awesome. And of course, large static assets, images I have done something to enter colloquially now known as shitposts, which is great. Other than logs, what could you possibly be storing in S3 that lends itself to, effectively, the type of analysis that you built around this?

Thomas: Well, our first use case was the classic log use cases, app logs, web service logs. I mean, CloudTrail, it’s famous; we had customers that gave up on elastic, and definitely gave up on relational where you can do a couple changes and your permutation of attributes for CloudTrail is going to put you to your knees. And people just say, “I give up.” Same thing with Kubernetes logs. And so it’s the classic—whether it’s CSV, where it’s JSON, where it’s log types, we auto-discover all that.

We also allow you, if you want to override that and change the parsing capabilities through a UI wizard, we do discover what’s in your buckets. That term data swamp, and not knowing what’s in your bucket, we do a facility that will index that data, actually create a report for you for knowing what’s in. Now, if you have text data, if you have log data, if you have BI data, we can bring it all together, but the real pain is at the scale. So classically, app logs, system logs, many devices sending IoT-type streams is where we really come in—Kubernetes—where they’re dealing with terabytes of data per day, and managing an ELK cluster at that scale. Particularly on a Black Friday.

Shoot, some of our customers like—Klarna is one of them; credit card payment—they’re ramping up for Black Friday, and one of the reasons why they chose us is our ability to scale when maybe you’re doing a terabyte or two a day and then it goes up to twenty, twenty-five. How do you test that scale? How do you manage that scale? And so for us, the data streams are, traditionally with our customers, the well-known log types, at least in the log use cases. And the challenge is scaling it, is getting access to it, and that’s where we come in.

Corey: I will say the last time you were on the show a couple of years ago, you were talking about the initial logging use case and you were speaking, in many cases aspirationally, about where things were going. What a difference a couple years is made. Instead of talking about what hypothetical customers might want, or what—might be able to do, you’re just able to name-drop them off the top of your head, you have scaled to approximately ten times the number of employees you had back then. You’ve—

Thomas: Yep. Yep.

Corey: —raised, I think, a total of—what, 50 million?—since then.

Thomas: Uh, 60 now. Yeah.

Corey: Oh, 60? Fantastic.

Thomas: Yeah, yeah.

Corey: Congrats. And of course, how do you do it? By sponsoring Last Week in AWS, as everyone should. I’m taking clear credit for that every time someone announces around, that’s the game. But no, there is validity to it because telling fun stories and sponsoring exciting things like this only carry you so far. At some point, customers have to say, yeah, this is solving a pain that I have; I’m willing to pay you money to solve it.

And you’ve clearly gotten to a point where you are addressing the needs of those customers at a pretty fascinating clip. It’s bittersweet from my perspective because it seems like the majority of your customers have not come from my nonsense anymore. They’re finding you through word of mouth, they’re finding through more traditional—read as boring—ad campaigns, et cetera, et cetera. But you’ve built a brand that extends beyond just me. I’m no longer viewed as the de facto ombudsperson for any issue someone might have with ChaosSearch on Twitters. It’s kind of, “Aww, the company grew up. What happened there?”

Thomas: No, [laugh] listen, this you were great. We reached out to you to tell our story, and I got to be honest. A lot of people came by, said, “I
heard something on Corey Quinn’s podcasts,” or et cetera. And it came a long way now. Now, we have, you know, companies like Equifax, multi-cloud—Amazon and Google.

They love the data lake philosophy, the centralized, where use cases are now available within days, not weeks and months. Whether it’s logs and BI. Correlating across all those data streams, it’s huge. We mentioned Klarna, [APM Performance 00:13:19], and, you know, we have Armor for SIEM, and Blackboard for [Observers 00:13:24].

So, it’s funny—yeah, it’s funny, when I first was talking to you, I was like, “What if? What if we had this customer, that customer?” And we were building the capabilities, but now that we have it, now that we have customers, yeah, I guess, maybe we’ve grown up a little bit. But hey, listen to you’re always near and dear to our heart because we remember, you know, when you stop[ed by our booth at re:Invent several times. And we’re coming to re:Invent this year, and I believe you are as well.

Corey: Oh, yeah. But people listening to this, it’s if they’re listening the day it’s released, this will be during re:Invent. So, by all means, come by the ChaosSearch booth, and see what they have to say. For once they have people who aren’t me who are going to be telling stories about these things. And it’s fun. Like, I joke, it’s nothing but positive here.

It’s interesting from where I sit seeing the parallels here. For example, we have both had—how we say—adult supervision come in. You have a CEO, Ed, who came over from IBM Storage. I have Mike Julian, whose first love language is of course spreadsheets. And it’s great, on some level, realizing that, wow, this company has eclipsed my ability to manage these things myself and put my hands-on everything. And eventually, you have to start letting go. It’s a weird growth stage, and it’s a heck of a transition. But—

Thomas: No, I love it. You know, I mean, I think when we were talking, we were maybe 15 employees. Now, we’re pushing 100. We brought on Ed Walsh, who’s an amazing CEO. It’s funny, I told him about this idea, I invented this technology roughly eight years ago, and he’s like, “I love it. Let’s do it.” And I wasn’t ready to do it.

So, you know, five, six years ago, I started the company always knowing that, you know, I’d give him a call once we got the plane up in the air. And it’s been great to have him here because the next level up, right, of execution and growth and business development and sales and marketing. So, you’re exactly right. I mean, we were a young pup several years ago, when we were talking to you and, you know, we’re a little bit older, a little bit wiser. But no, it’s great to have Ed here. And just the leadership in general; we’ve grown immensely.

Corey: Now, we are recording this in advance of re:Invent, so there’s always the question of, “Wow, are we going to look really silly based upon what is being announced when this airs?” Because it’s very hard to predict some things that AWS does. And let’s be clear, I always stay away from predictions, just because first, I have a bit of a knack for being right. But also, when I’m right, people will think, “Oh, Corey must have known about that and is leaking,” whereas if I get it wrong, I just look like a fool. There’s no win for me if I start doing the predictive dance on stuff like that.

But I have to level with you, I have been somewhat surprised that, at least as of this recording, AWS has not moved more in your direction because storing data in S3 is kind of their whole thing, and querying that data through something that isn’t Athena has been a bit of a reach for them that they’re slowly starting to wrap their heads around. But their UltraWarm nonsense—which is just, okay, great naming there—what is the point of continually having a model where oh, yeah, we’re going to just age it out, the stuff that isn’t actively being used into S3, rather than coming up with a way to query it there. Because you’ve done exactly that, and please don’t take this as anything other than a statement of fact, they have better access to what S3 is doing than you do. You’re forced to deal with this thing entirely from a public API standpoint, which is fine. They can theoretically change the behavior of aspects of S3 to unlock these use cases if they chose to do so. And they haven’t. Why is it that you’re the only folks that are doing this?

Thomas: No, it’s a great question, and I’ll give them props for continuing to push the data lake [unintelligible 00:17:09] to the cloud providers’ S3 because it was really where I saw the world. Lakes, I believe in. I love them. They love them. However, they promote the move the data out to get access, and it seems so counterintuitive on why wouldn’t you leave it in and put these services, make them more intelligent? So, it’s funny, I’ve trademark ‘Smart Object Storage,’ I actually trademarked—I think you [laugh] were a part of this—‘UltraHot,’ right? Because why would you want UltraWarm when you can have UltraHot?

And the reason, I feel, is that if you’re using Parquet for Athena [unintelligible 00:17:40] store, or Lucene for Elasticsearch, these two index technologies were not designed for cloud storage, for real-time streaming off of cloud storage. So, the trick is, you have to build UltraWarm, get it off of what they consider cold S3 into a more warmer memory or SSD type access. What we did, what the invention I created was, that first read is hot. That first read is fast.

Snowflake is a good example. They give you a ten terabyte demo example, and if you have a big instance and you do that first query, maybe several orders or groups, it could take an hour to warm up. The second query is fast. Well, what if the first query is in seconds as well? And that’s where we really spent the last five, six years building out the tech and the vision behind this because I like to say you go to a doctor and say, “Hey, Doc, every single time I move my arm, it hurts.” And the doctor says, “Well, don’t move your arm.”

It’s things like that, to your point, it’s like, why wouldn’t they? I would argue, one, you have to believe it’s possible—we’re proving that it is—and two, you have to have the technology to do it. Not just the index, but the architecture. So, I believe they will go this direction. You know, little birdies always say that all these companies understand this need.

Shoot, Snowflake is trying to be lake-y; Databricks is trying to really bring this warehouse lake concept. But you still do all the pipelining; you still have to do all the data management the way that you don’t want to do. It’s not a lake. And so my argument is that it’s innovation on why. Now, they have money; they have time, but, you know, we have a big head start.

Corey: I remembered last year at re:Invent they released a, shall we say, significant change to S3 that it enabled read after write consistency, which is awesome, for again, those of us in the business of misusing things as databases. But for some folks, the majority of folks I would say, it was a, “I don’t know what that means and therefore I don’t care.” And that’s fine. I have no issue with that. There are other folks, some of my customers for example, who are suddenly, “Wait a minute. This means I can sunset this entire janky sidecar metadata system that is designed to make sure that we are consistent in our use of S3 because it now does it automatically under the hood?” And that’s awesome. Does that change mean anything for ChaosSearch?

Thomas: It doesn’t because of our architecture. We’re append-only, write-once scenario, so a lot of update-in-place viewpoints. My viewpoint is that if you’re seeing S3 as the database and you need that type of consistency, it make sense of why you’d want it, but because of our distributive fabric, our stateless architecture, our append-only nature, it really doesn’t affect us.

Now, I talked to the S3 team, I said, “Please if you’re coming up with this feature, it better not be slower.” I want S3 to be fast, right? And they said, “No, no. It won’t affect performance.” I’m like, “Okay. Let’s keep that up.”

And so to us, any type of S3 capability, we’ll take advantage of it if benefits us, whether it’s consistency as you indicated, performance, functionality. But we really keep the constructs of S3 access to really limited features: list, put, get. [roll-on 00:20:49] policies to give us read-only access to your data, and a location to write our indices into your account, and then are distributed fabric, our service, acts as those indices and query them or searches them to resolve whatever analytics you need. So, we made it pretty simple, and that is allowed us to make it high performance.

Corey: I’ll take it a step further because you want to talk about changes since the last time we spoke, it used to be that this was on top of S3, you can store your data anywhere you want, as long as it’s S3 in the customer’s account. Now, you’re also supporting one-click integration with Google Cloud’s object storage, which, great. That does mean though, that you’re not dependent upon provider-specific implementations of things like a consistency model for how you’ve built things. It really does use the lowest common denominator—to my understanding—of object stores. Is that something that you’re seeing broad adoption of, or is this one of those areas where, well, you have one customer on a different provider, but almost everything lives on the primary? I’m curious what you’re seeing for adoption models across multiple providers?

Thomas: It’s a great question. We built an architecture purposely to be cloud-agnostic. I mean, we use compute in a containerized way, we use object storage in a very simple construct—put, get, list—and we went over to Google because that made sense, right? We have customers on both sides. I would say Amazon is the gorilla, but Google’s trying to get there and growing.

We had a big customer, Equifax, that’s on both Amazon and Google, but we offer the same service. To be frank, it looks like the exact same product. And it should, right? Whether it’s Amazon Cloud, or Google Cloud, multi-select and I want to choose either one and get the other one. I would say that different business types are using each one, but our bulk of the business isn’t Amazon, but we just this summer released our SaaS offerings, so it’s growing.

And you know, it’s funny, you never know where it comes from. So, we have one customer—actually DigitalRiver—as one of our customers on Amazon for logs, but we’re growing in working together to do a BI on GCP or on Google. And so it’s kind of funny; they have two departments on two different clouds with two different use cases. And so do they want unification? I’m not sure, but they definitely have their BI on Google and their operations in Amazon. It’s interesting.

Corey: You know its important to me that people learn how to use the cloud effectively. Thats why I’m so glad that Cloud Academy is sponsoring my ridiculous non-sense. They’re a great way to build in demand tech skills the way that, well personally, I learn best which I learn by doing not by reading. They have live cloud labs that you can run in real environments that aren’t going to blow up your own bill—I can’t stress how important that is. Visit cloudacademy.com/corey. Thats C-O-R-E-Y, don’t drop the “E.” Use Corey as a promo-code as well. You’re going to get a bunch of discounts on it with a lifetime deal—the price will not go up. It is limited time, they assured me this is not one of those things that is going to wind up being a rug pull scenario, oh no no. Talk to them, tell me what you think. Visit: cloudacademy.com/corey, C-O-R-E-Y and tell them that I sent you!

Corey: I know that I’m going to get letters for this. So, let me just call it out right now. Because I’ve been a big advocate of pick a provider—I care not which one—and go all-in on it. And I’m sitting here congratulating you on extending to another provider, and people are going to say, “Ah, you’re being inconsistent.”

No. I’m suggesting that you as a provider have to meet your customers where they are because if someone is sitting in GCP and your entire approach is, “Step one, migrate those four petabytes of data right on over here to AWS,” they’re going to call you that jackhole that you would be by making that suggestion and go immediately for option B, which is literally anything that is not ChaosSearch, just based upon that core misunderstanding of their business constraints. That is the way to think about these things. For a vendor position that you are in as an ISV—Independent Software Vendor for those not up on the lingo of this ridiculous industry—you have to meet customers where they are. And it’s the right move.

Thomas: Well, you just said it. Imagine moving terabytes and petabytes of data.

Corey: It sounds terrific if I’m a salesperson for one of these companies working on commission, but for the rest of us, it sounds awful.

Thomas: We really are a data fabric across clouds, within clouds. We’re going to go where the data is and we’re going to provide access to
where that data lives. Our whole philosophy is the no-movement movement, right? Don’t move your data. Leave it where it is and provide access at scale.

And so you may have services in Google that naturally stream to GCS; let’s do it there. Imagine moving that amount of data over to Amazon to analyze it, and vice versa. 2020, we’re going to be in Azure. They’re a totally different type of business, users, and personas, but you’re getting asked, “Can you support Azure?” And the answer is, “Yes,” and, “We will in 2022.”

So, to us, if you have cloud storage, if you have compute, and it’s a big enough business opportunity in the market, we’re there. We’re going there. When we first started, we were talking to MinIO—remember that open-source, object storage platform?—We’ve run on our laptops, we run—this [unintelligible 00:25:04] Dr. Seuss thing—“We run over here; we run over there; we run everywhere.”

But the honest truth is, you’re going to go with the big cloud providers where the business opportunity is, and offer the same solution because the same solution is valued everywhere: simple in; value out; cost-effective; long retention; flexibility. That sounds so basic, but you mentioned this all the time with our Rube Goldberg, Amazon diagrams we see time and time again. It’s like, if you looked at that and you were from an alien planet, you’d be like, “These people don’t know what they’re doing. Why is it so complicated?” And the simple answer is, I don’t know why people think it’s complicated.

To your point about Amazon, why won’t they do it? I don’t know, but if they did, things would be different. And being honest, I think people are catching on. We do talk to Amazon and others. They see the need, but they also have to build it; they have to invent technology to address it. And using Parquet and Lucene are not the answer.

Corey: Yeah, it’s too much of a demand on the producers of that data rather than the consumer. And yeah, I would love to be able to go upstream to application developers and demand they do things in certain ways. It turns out as a consultant, you have zero authority to do that. As a DevOps team member, you have limited ability to influence it, but it turns out that being the ‘department of no’ quickly turns into being the ‘department of unemployment insurance’ because no one wants to work with you. And collaboration—contrary to what people wish to believe—is a key part of working in a modern workplace.

Thomas: Absolutely. And it’s funny, the demands of IT are getting harder; the actual getting the employees to build out the solutions are getting harder. And so a lot of that time is in the pipeline, is the prep, is the schema, the sharding, and et cetera, et cetera, et cetera. My viewpoint is that should be automated away. More and more databases are being autotune, right?

This whole knobs and this and that, to me, Glue is a means to an end. I mean, let’s get rid of it. Why can’t Athena know what to do? Why can’t object storage be Athena and vice versa? I mean, to me, it seems like all this moving through all these services, the classic Amazon viewpoint, even their diagrams of having this centralized repository of S3, move it all out to your services, get results, put it back in, then take it back out again, move it around, it just doesn’t make much sense. And so to us, I love S3, love the service. I think it’s brilliant—Amazon’s first service, right?—but from there get a little smarter. That’s where ChaosSearch comes in.

Corey: I would argue that S3 is in fact, a modern miracle. And one of those companies saying, “Oh, we have an object store; it’s S3 compatible.” It’s like, “Yeah. We have S3 at home.” Look at S3 at home, and it’s just basically a series of failing Raspberry Pis.

But you have this whole ecosystem of things that have built up and sprung up around S3. It is wildly understated just how scalable and massive it is. There was an academic paper recently that won an award on how they use automated reasoning to validate what is going on in the S3 environment, and they talked about hundreds of petabytes in some cases. And folks are saying, ah, S3 is hundreds of petabytes. Yeah, I have clients storing hundreds of petabytes.

There are larger companies out there. Steve Schmidt, Amazon’s CISO, was recently at a Splunk keynote where he mentioned that in security info alone, AWS itself generates 500 petabytes a day that then gets reduced down to a bunch of stuff, and some of it gets loaded into Splunk. I think. I couldn’t really hear the second half of that sentence because of the sound of all of the Splunk salespeople in that room becoming excited so quickly you could hear it.

Thomas: [laugh]. I love it. If I could be so bold, those S3 team, they’re gods. They are amazing. They created such an amazing service, and when I started playing with S3 now, I guess, 2006 or 7, I mean, we were using for a repository, URL access to get images, I was doing a virtualization [unintelligible 00:29:05] at the time—

Corey: Oh, the first time I played with it, “This seems ridiculous and kind of dumb. Why would anyone use this?” Yeah, yeah. It turns out I’m
really bad at predicting the future. Another reason I don’t do the prediction thing.

Thomas: Yeah. And when I started this company officially, five, six years ago, I was thinking about S3 and I was thinking about HDFS not being a good answer. And I said, “I think S3 will actually achieve the goals and performance we need.” It’s a distributed file system. You can run parallel puts and parallel gets. And the performance that I was seeing when the data was a certain way, certain size, “Wait, you can get high performance.”

And you know, when I first turned on the engine, now four or five years ago, I was like, “Wow. This is going to work. We’re off to the races.” And now obviously, we’re more than just an idea when we first talked to you. We’re a service.

We deliver benefits to our customers both in logs. And shoot, this quarter alone we’re coming out with new features not just in the logs, which I’ll talk about second, but in a direct SQL access. But you know, one thing that you hear time and time again, we talked about it—JSON, CloudTrail, and Kubernetes; this is a real nightmare, and so one thing that we’ve come out with this quarter is the ability to virtually flatten. Now, you heard time and time again, where, “Okay. I’m going to pick and choose my data because my database can’t handle whether it’s elastic, or say, relational.” And all of a sudden, “Shoot, I don’t have that. I got to reindex that.”

And so what we’ve done is we’ve created a index technology that we’re always planning to come out with that indexes the JSON raw blob, but in the data refinery have, post-index you can select how to unflatten it. Why is that important? Because all that tooling, whether it’s elastic or SQL, is now available. You don’t have to change anything. Why is Snowflake and BigQuery has these proprietary JSON APIs that none of these tools know how to use to get access to the data?

Or you pick and choose. And so when you have a CloudTrail, and you need to know what’s going on, if you picked wrong, you’re in trouble. So, this new feature we’re calling ‘Virtual Flattening’—or I don’t know what we’re—we have to work with the marketing team on it. And we’re also bringing—this is where I get kind of excited where the elastic world, the ELK world, we’re bringing correlations into Elasticsearch. And like, how do you do that? They don’t have the APIs?

Well, our data refinery, again, has the ability to correlate index patterns into one view. A view is an index pattern, so all those same constructs that you had in Kibana, or Grafana, or Elastic API still work. And so, no more denormalizing, no more trying to hodgepodge query over here, query over there. You’re actually going to have correlations in Elastic, natively. And we’re excited about that.

And one more push on the future, Q4 into 2022; we have been given early access to S3 SQL access. And, you know, as I mentioned, correlations in Elastic, but we’re going full in on publishing our [TPCH 00:31:56] report, we’re excited about publishing those numbers, as well as not just giving early access, but going GA in the first of the year, next year.

Corey: I look forward to it. This is also, I guess, it’s impossible to have a conversation with you, even now, where you’re not still forward-looking about what comes next. Which is natural; that is how we get excited about the things that we’re building. But so much less of what you’re doing now in our conversations have focused around what’s coming, as opposed to the neat stuff you’re already doing. I had to double-check when we were talking just now about oh, yeah, is that Google cloud object store support still something that is roadmapped, or is that out in the real world?

No, it’s very much here in the real world, available today. You can use it. Go click the button, have fun. It’s neat to see at least some evidence that not all roadmaps are wishes and pixie dust. The things that you were talking to me about years ago are established parts of ChaosSearch now. It hasn’t been just, sort of, frozen in amber for years, or months, or these giant periods of time. Because, again, there’s—yeah, don’t sell me vaporware; I know how this works. The things you have promised have come to fruition. It’s nice to see that.

Thomas: No, I appreciate it. We talked a little while ago, now a few years ago, and it was a bit of aspirational, right? We had a lot to do, we had more to do. But now when we have big customers using our product, solving their problems, whether it’s security, performance, operation, again—at scale, right? The real pain is, sure you have a small ELK cluster or small Athena use case, but when you’re dealing with terabytes to petabytes, trillions of rows, right—billions—when you were dealing trillions, billions are now small. Millions don’t even exist, right?

And you’re graduating from computer science in college and you say the word, “Trillion,” they’re like, “Nah. No one does that.” And like you were saying, people do petabytes and exabytes. That’s the world we’re living in, and that’s something that we really went hard at because these are challenging data problems and this is where we feel we uniquely sit. And again, we don’t have to break the bank while doing it.

Corey: Oh, yeah. Or at least as of this recording, there’s a meme going around, again, from an old internal Google Video, of, “I just want to serve five terabytes of traffic,” and it’s an internal Google discussion of, “I don’t know how to count that low.” And, yeah.

Thomas: [laugh].

Corey: But there’s also value in being able to address things at much larger volume. I would love to see better responsiveness options around things like Deep Archive because the idea of being able to query that—even if you can wait a day or two—becomes really interesting just from the perspective of, at that point, current cost for one petabyte of data in Glacier Deep Archive is 1000 bucks a month. That is ‘why would I ever delete data again?’ Pricing.

Thomas: Yeah. You said it. And what’s interesting about our technology is unlike, let’s say Lucene, when you index it, it could be 3, 4, or 5x the raw size, our representation is smaller than gzip. So, it is a full representation, so why don’t you store it efficiently long-term in S3? Oh, by the way, with the Glacier; we support Glacier too.

And so, I mean, it’s amazing the cost of data with cloud storage is dramatic, and if you can make it hot and activated, that’s the real promise of a data lake. And, you know, it’s funny, we use our own service to run our SaaS—we log our own data, we monitor, we alert, have dashboards—and I can’t tell you how cheap our service is to ourselves, right? Because it’s so cost-effective for long-tail, not just, oh, a few weeks; we store a whole year’s worth of our operational data so we can go back in time to debug something or figure something out. And a lot of that’s savings. Actually, huge savings is cloud storage with a distributed elastic compute fabric that is serverless. These are things that seem so obvious now, but if you have SSDs, and you’re moving things around, you know, a team of IT professionals trying to manage it, it’s not cheap.

Corey: Oh, yeah, that’s the story. It’s like, “Step one, start paying for using things in cloud.” “Okay, great. When do I stop paying?” “That’s the neat part. You don’t.” And it continues to grow and build.

And again, this is the thing I learned running a business that focuses on this, the people working on this, in almost every case, are more expensive than the infrastructure they’re working on. And that’s fine. I’d rather pay people than technologies. And it does help reaffirm, on some level, that—people don’t like this reminder—but you have to generate more value than you cost. So, when you’re sitting there spending all your time trying to avoid saving money on, “Oh, I’ve listened to ChaosSearch talk about what they do a few times. I can probably build my own and roll it at home.”

It’s, I’ve seen the kind of work that you folks have put into this—again, you have something like 100 employees now; it is not just you building this—my belief has always been that if you can buy something that gets you 90, 95% of where you are, great. Buy it, and then yell at whoever selling it to you for the rest of it, and that’ll get you a lot further than, “We’re going to do this ourselves from first principles.” Which is great for a weekend project for just something that you have a passion for, but in production mistakes show. I’ve always been a big proponent of buying wherever you can. It’s cheaper, which sounds weird, but it’s true.

Thomas: And we do the same thing. We have single-sign-on support; we didn’t build that ourselves, we use a service now. Auth0 is one of our providers now that owns that [crosstalk 00:37:12]—

Corey: Oh, you didn’t roll your own authentication layer? Why ever not? Next, you’re going to tell me that you didn’t roll your own payment gateway when you wound up charging people on your website to sign up?

Thomas: You got it. And so, I mean, do what you do well. Focus on what you do well. If you’re repeating what everyone seems to do over and over again, time, costs, complexity, and… service, it makes sense. You know, I’m not trying to build storage; I’m using storage. I’m using a great, wonderful service, cloud object storage.

Use whats works, whats works well, and do what you do well. And what we do well is make cloud object storage analytical and fast. So, call us up and we’ll take away that 2 a.m. call you have when your cluster falls down, or you have a new workload that you are going to go to the—I don’t know, the beach house, and now the weekend shot, right? Spin it up, stream it in. We’ll take over.

Corey: Yeah. So, if you’re listening to this and you happen to be at re:Invent, which is sort of an open question: why would you be at re:Invent while listening to a podcast? And then I remember how long the shuttle lines are likely to be, and yeah. So, if you’re at re:Invent, make it on down to the show floor, visit the ChaosSearch booth, tell them I sent you, watch for the wince, that’s always worth doing. Thomas, if people have better decision-making capability than the two of us do, where can they find you if they’re not in Las Vegas this week?

Thomas: So, you find us online chaossearch.io. We have so much material, videos, use cases, testimonials. You can reach out to us, get a free trial. We have a self-service experience where connect to your S3 bucket and you’re up and running within five minutes.

So, definitely chaossearch.io. Reach out if you want a hand-held, white-glove experience POV. If you have those type of needs, we can do that with you as well. But we booth on re:Invent and I don’t know the booth number, but I’m sure either we’ve assigned it or we’ll find it out.

Corey: Don’t worry. This year, it is a low enough attendance rate that I’m projecting that you will not be as hard to find in recent years. For example, there’s only one expo hall this year. What a concept. If only it hadn’t taken a deadly pandemic to get us here.

Thomas: Yeah. But you know, we’ll have the ability to demonstrate Chaos at the booth, and really, within a few minutes, you’ll say, “Wow. How come I never heard of doing it this way?” Because it just makes so much sense on why you do it this way versus the merry-go-round of data movement, and transformation, and schema management, let alone all the sharding that I know is a nightmare, more often than not.

Corey: And we’ll, of course, put links to that in the [show notes 00:39:40]. Thomas, thank you so much for taking the time to speak with me today. As always, it’s appreciated.

Thomas: Corey, thank you. Let’s do this again.

Corey: We absolutely will. Thomas Hazel, CTO and Founder of ChaosSearch. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast episode, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this episode, please leave a five-star review on your podcast platform of choice along with an angry comment because I have dared to besmirch the honor of your homebrewed object store, running on top of some trusty and reliable Raspberries Pie.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and
we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Rachel

Rachel Stephens is a Senior Analyst with RedMonk, a developer-focused industry analyst firm. RedMonk focuses on how practitioners drive technological adoption. Her research covers a broad range of developer and infrastructure products, with a particular focus on emerging growth technologies and markets. (But not crypto. Please don't talk to her about NFTs.)

Before joining RedMonk, Rachel worked as a database administrator and financial analyst. Rachel holds an MBA from Colorado State University and a BA in Finance from the University of Colorado.

Links:

  • RedMonk: https://redmonk.com/
  • Great analysis: https://redmonk.com/rstephens/2021/09/30/a-new-strategy-r2/
  • “Convergent Evolution of CDNs and Clouds”: https://redmonk.com/sogrady/2020/06/10/convergent-evolution-cdns-cloud/
  • “Everything is Securities Fraud?”: https://cafe.com/stay-tuned/everything-is-securities-fraud-with-matt-levine/
  • Twitter: https://twitter.com/rstephensme

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by my friends at ThinkstCanary. Most companies find out way too late that they’ve been breached. ThinkstCanary changes this and I love how they do it. Deploy canaries and canary tokens in minutes and then forget about them. What's great is the attackers tip their hand by touching them, giving you one alert, when it matters. I use it myself and I only remember this when I get the weekly update with a “we’re still here, so you’re aware” from them. It’s glorious! There is zero admin overhead to this, there are effectively no false positives unless I do something foolish. Canaries are deployed and loved on all seven continents. You can check out what people are saying at canary.love. And, their Kub config canary token is new and completely free as well. You can do an awful lot without paying them a dime, which is one of the things I love about them. It is useful stuff and not an, “ohh, I wish I had money.” It is speculator! Take a look; that’s canary.love because it's genuinely rare to find a security product that people talk about in terms of love. It really is a unique thing to see. Canary.love. Thank you to ThinkstCanary for their support of my ridiculous, ridiculous non-sense.

Corey: This episode is sponsored in part by our friends at Vultr. Spelled V-U-L-T-R because they’re all about helping save money, including on things like, you know, vowels. So, what they do is they are a cloud provider that provides surprisingly high performance cloud compute at a price that—while sure they claim its better than AWS pricing—and when they say that they mean it is less money. Sure, I don’t dispute that but what I find interesting is that it’s predictable. They tell you in advance on a monthly basis what it’s going to going to cost. They have a bunch of advanced networking features. They have nineteen global locations and scale things elastically. Not to be confused with openly, because apparently elastic and open can mean the same thing sometimes. They have had over a million users. Deployments take less that sixty seconds across twelve pre-selected operating systems. Or, if you’re one of those nutters like me, you can bring your own ISO and install basically any operating system you want. Starting with pricing as low as $2.50 a month for Vultr cloud compute they have plans for developers and businesses of all sizes, except maybe Amazon, who stubbornly insists on having something to scale all on their own. Try Vultr today for free by visiting: vultr.com/screaming, and you’ll receive a $100 in credit. Thats v-u-l-t-r.com slash screaming.

Corey: Welcome to Screaming in the Cloud, I’m Corey Quinn. The last time I spoke to Rachel Stephens over at RedMonkwas in December of 2019. Well, on this podcast anyway; we might have exchanged conversational tidbits here and there at some point since then. But really, if we look around the world there’s nothing that’s materially different than it was today from December in 2019, except, oh, that’s right everything. Rachel Stephens, you’re still a senior analyst at RedMonk, which hey, in this day and age, longevity at a company is something that is almost enough to occasion comment on its own. Thanks for coming back for another round, I appreciate it.

Rachel: Oh, I’m so happy to be here, and it’s exciting to talk about the state of the world a few years later than the last time we talked. But yeah, it’s been a hell of a couple years.

Corey: Really has, but rather than rehashing pandemic stuff because I feel like unless people have been living in a cave for the last couple of years—because we’ve all been living in caves for the last couple of years—they know what’s up with that. What’s new in your world? What has changed for you aside from all of this in the past couple of years working in one of the most thankless of all jobs, an analyst in the cloud computing industry?

Rachel: Well, the job stuff is all excellent and I’ve had wonderful time working at RedMonk. So, RedMonk overall is an analyst firm that is focused on helping people understand technology trends, particularly from the view of the developer or the practitioner. So, helping to understand how the people who are using technologies are actually driving their overall adoption. And so there has been all kinds of interesting things that have happened in that score in the last couple of years. We’ve seen a lot of interesting trends, lots of fun things to look at in the space and it’s been a lot.

On a personal side, like a week into lockdown I found out that I was pregnant, so I went through all of locking down and the heart of the pandemic pregnant. I had my maternity leave earlier this year and came back and so excited to be back in. But it’s also just been a lot to catch up on in the space as you come back from leave which I’m sure you are well familiar with.

Corey: Yes, I did the same thing, slightly differently timed. My second daughter Josephine was born at the end of September. When did your kiddo arrive on the scene to a world of masked strangers?

Rachel: So, I have an older daughter who just turned four, and then my youngest is coming up on his first birthday. He was born in December.

Corey: Excellent. It sounds like our kids are basically the same age, in both directions. And from my perspective, at least looking back, what advice would I give someone for having a baby in a pandemic? It distills down to ‘don’t,’ just because it changes so much, it’s no longer a trivial thing to have a grandparent come out and spend time with the kid. It’s the constant… drumbeat of is this over? Is this not over?

And that manifested a bunch of different ways. And I’m glad that I got the opportunity to take some time off to spend time with my family during that timeframe, but at the same time, it would’ve great if there were options such as not being stuck at home with every rambunctious—at the time—three-year-old as I went through that entire joy of having the kid.

Rachel: Yeah. No, for the longest time, my thing was like, okay, like, there’s no amount of money you could pay me to go back to middle school. I would never do it. And my new high bar there is no amount of money that you could pay me to go back to April 2020. That was the hardest month of my entire life was getting through that, like, first trimester, both parents at home, toddler at home, nowhere to go, no one to help. That was a [BLEEP] hard month. [laugh] that was bad.

Corey: Oh, my God, yes, and we don’t talk about this because we’re basically communicating with people on social media, and everyone feels bad looking at social media because they’re comparing their blooper reel with everyone else’s highlights. And it feels odd on some level to complain about things like that. And let’s be very clear as a man, I wind up in society getting lauded for even deigning to mention that I have children, whereas when mothers wind up talking about anything even slightly negative it’s, “Oh, you sound like a bad mom.” And it is just one of the most abhorrent things out there in the world, I suppose. It’s a strange inverted thing but one of the things that surprised me the most when I was expecting my first kid was looking at the different parenting forums, and the difference in tone was palpable where on the dad forums, everyone is super supportive and you got this dude it’s great. You’re fine. You’re doing your best.

Sure, these the occasional, “I gave my toddler beer and now people yell at me,” and it’s, “What is wrong with you asshole?” But everyone else is mostly sane and doing their best. Whereas a lot of the ‘mommy’ forums seem to bias more toward being relatively dismissive other people’s parenting choices. And I understand I’m stereotyping wildly, not all forums, not all people, et cetera, et cetera, but it really was an interesting window into an area that as a stereotypical white man world, I don’t see a lot of places I hang out with that are traditionally male that are overwhelmingly supportive in quite the same way. It was really an eye-opening experience for me.

Rachel: I think you hit on some really important trends. One of the things that I have struggled with is—so I came into RedMonk—I’ve had a Twitter account forever, and it was always just, like, my personal Twitter account until I started working at RedMonk five years ago. And then all of a sudden I’m tweeting in technical and work capacities as well. And finding that balance initially was always a challenge.

But then finding that balance again after having kids was very different because I would always—it was, kind of, mix of my life and also what I’m seeing in the industry and what I’m working on and this mix of things. And once you started tweeting about kids, it very much changes the potential perception that people have of who you are, what you’re doing. I know this is just a mommy blogger kind of thing.

You have to be really cognizant of that balance and making sure that you continue to put yourself in a place where you can still be your authentic self, but you really as a mom in the workspace and especially in tech have to be cognizant of not leaning too far into that. Because it can really damage your credibility with some audiences which is a super unfortunate thing, but also something I’ve learned just, like, I have to
be really careful about how much mom stuff I share on Twitter.

Corey: It’s bizarre to me that we have to shade aspects of ourselves like this. And I don’t know what the answer is. It’s a weird thing that I never thought about before until suddenly I find that, oh, I’m a parent. I guess I should actually pay attention to this thing now. And it’s one of those once you see it you can’t unsee it things and it becomes strange and interesting and also more than a little sad in some respects.

Rachel: I think there are some signs that we are getting to a better place, but it’s a hard road for parents I think, and moms, in particular, working moms, all kinds of challenges out there. But anyways, it’s one of those ones that is nice because love having my kids as a break, but sometimes Mondays come and it’s such a relief to come back to work after a weekend with kids. Kids are a lot of work. And so it has brought elements of joy to my personal life, but it has also brought renewed elements of joy to my work life as really being able to lean into that side of myself. So, it’s been a good year.

Corey: Now that I have a second kid, I’m keenly aware of why parents are always very reluctant to wind up—the good parents at least to say, “Oh yeah, I have X number of kids, but that one’s my favorite.” And I understand now why my mom always said with my brother and I that, “I
can’t stand either one of you.” And I get that now. Looking at the children of cloud services it’s like, which one is my favorite?

Well, I can’t stand any of them, but the one that I hate the most is the Managed NAT Gateway because of its horrible pricing. In fact, anything
involving bandwidth pricing in this industry tends to be horrifying, annoying, ill-behaved, and very hard to discipline. Which is why I think it’s probably time we talk a little bit about egress charges in cloud providers.

You had a great analysis of Cloudflare’s R2, which is named after a robot in Star Wars and is apparently also the name of their S3 competitor, once it launches. Again, this is a pre-announcement, yeah, I could write blog posts that claim anything; the proof is really going to be in the pudding. Tell me more about, I guess, what you noticed from that announcement and what drove you to, “Ah, I have thoughts on this?”

Rachel: So, I think it’s an interesting announcement for several reasons. I think one of them is that it makes their existing offering really compelling when you start to add in that object store to something like the CDN, or to their edge functions which is called their Workers platform. And so, if you start to combine some of those functionalities together with a better object storage story, it can make their existing offering a lot more compelling which I think is an interesting aspect of this.

I think one of the aspects that is probably gotten a lot more of the traction though is their lack of egress pricing. So, I think that’s really what took everyone’s imaginations by storm is what does the world look like when we are not charging egress pricing on object storage?

Corey: What I find interesting is that when this came out first, a lot of AWS fans got very defensive over it, which I found very odd because their egress charges are indefensible from my point of view. And their response was, “Well, if you look at how a lot of the data access patterns work this isn’t as big of a deal as it looks like,” and you’re right. If I have a whole bunch of objects living in an object store, and a whole bunch of people each grab one of those objects this won’t help me in any meaningful sense.

But if I have one object that a bunch of people grab, well, suddenly we’re having a different conversation. And on some level, it turns into an interesting question of what differentiates this with their existing CDN-style approach. From my perspective, this is where the object actually lives rather than just a cache that is going to expire. And that is transformative in a bunch of different ways, but my, I guess, admittedly overstated analysis for some use cases was okay, I store a petabyte in AWS and use it with and without this thing. Great, the answer came out to something like 51 or $52,000 in egress charges versus zero on Cloudflare. That’s an interesting perspective to take. And the orders of magnitude in difference are eye-popping assuming that it works as advertised, which is always the caveat.

Rachel: Yeah. I remember there was a RedMonk conversation with one of the cloud vendors set us up with a client conversation that want to, kind of, showcase their products kind of thing. And it was a movie studio and they walked us through what they architecturally have to do when they drop a trailer. If you think about that thing from this use case where all of a sudden you have videos that are all going out globally at the same time, and everybody wants to watch it and you’re serving it over and over, that’s a super interesting and compelling use case and very different from a cost perspective.

Corey: You’ll notice the video streaming services all do business with something that is not AWS for what they stream to end-users from. Netflix has its own Open Connect project that effectively acts as their own homebuilt CDN that they partner with providers to put in their various environments. There are a bunch of providers that focus specifically on this. But if you do the math for the Netflix story at retail pricing—let’s be clear at large scale, no one pays retail pricing for anything, but okay—even assuming that you’re within hailing distance of the same universe as retail pricing; you don’t have to watch too many hours of Netflix before the data egress charges cost more than you’re paying a month than subscription. And I have it on good authority—read as from their annual reports—that a much larger expense for Netflix than their cloud and technology and R&D expenses is their content expenses.

They’re making a lot of original content. They’re licensing an awful lot of content, and that’s way more expensive than providing it to folks. They have to have a better economic model. They need to be able to make a profit of some kind on streaming things to people. And with the way that all the major cloud providers wind up pricing this stuff, it’s not tenable. There has to be a better answer.

Rachel: So, Netflix calls to mind an interesting antidote that has gone around the industry which is who can become each other faster? Can HBO become Netflix faster, or can Netflix become HBO faster? So, can you build out that technology infrastructure side, or can you build out that content side? And I think what you’re talking to with their content costs speaks to that story in terms of where people are investing and trying to actually make dents in their strategic outlays.

I think a similar concept is actually at play when we talk about cloud and CDN. We do have this interesting piece from my coworker, Steve O’Grady, and he called it “Convergent Evolution of CDNs and Clouds.” And they originally evolved along separate paths where CDNs were designed to do this edge-caching scenario, and they had the core compute and all of the things that go around it happening in the cloud.

And I think we’ve seen in recent years both of them starting to grow towards each other where CDNs are starting to look a lot more cloud-like, and we’re seeing clouds trying to look more CDN-like. And I think this announcement in particular is very interesting when you think about what’s happening in the CDN space and what it actually means for where CDNs are headed.

Corey: It’s an interesting model in that if we take a look at all of the existing cloud providers they had some other business that funded the incredible expense outlay that it took to build them. For example, Amazon was a company that started off selling books and soon expanded to selling everything else, and then expanded to putting ads in all of their search portals, including in AWS and eroding customer trust.

Google wound up basically making all of their money by showing people ads and also killing Google Reader. And of course, Microsoft has been a software company for a lot longer than they’ve been a cloud provider, and given their security lapses in Azure recently is the question of whether they’ll continue to be taken seriously as a cloud provider.

But what makes Cloudflare interesting from this approach is they start it from the outside in of building out the edge before building regions or anything like that. And for a lot of use cases that works super well, in theory. In practice, well, we’ve never seen it before. I’m curious to see how it goes. Obviously, they’re telling great stories about how they envision this working out in the future. I don’t know how accurate it’s going to be—show, don’t tell—but I can at least acknowledge that the possibility is definitely there.

Rachel: I think there’s a lot of unanswered questions at this point, like, will you be able to have zero egress fees, and edge-like latencies, and global distribution, and have that all make sense and actually perform the way that the customer expects? I think that’s still to be seen. I think one of the things that we have watched with interest is this rise of—I think for lack of a better word techno-nationalism where we are starting
to see enclaves of where people want technology to be residing, where they want things to be sourced from, all of these interesting things.

And so having this global network of storage flies in the face of some of those trends where people are building more and more enclaves of we’re going to go big and global. I think that’s interesting and I think data residency in this global world will be an interesting question.

Corey: It also gets into the idea of what is the data that’s going to live there. Because the idea of data residency, yes, that is important, but where that generally tends to matter the most is things like databases or customer information. Not the thing that we’re putting out on the internet for anyone who wants to, to be able to download, which has historically been where CDNs are aiming things.

Yes, of course, they can restrict it to people with logins and the rest, but that type of object storage in my experience is not usually subject to heavy regulation around data residency. We’ll see because I get the sense that this is the direction Cloudflare is attempting to go in, and it’s really interesting to see how it works. I’m curious to know what their stories are around, okay, you have a global network. That’s great. Can I stipulate which areas my data can live within or not?

At some point, it’s going to need to happen if they want to look at regulated entities, but not everyone has to start with that either. So, it really just depends on what their game plan is on this. I like the fact that they’re willing to do this. I like the fact they’re willing to be as transparent as they are about their contempt for AWS’s egress fees. And yes, of course, they’re a competitor.

They’re going to wind up smacking competition like this, but I find it refreshing because there is no defense for what they’re saying, their math is right. Their approach to what customers experience from AWS in terms of egress fees is correct. And all of the defensiveness at, well, you know, no one pays retail price for this, yeah, but they see it on the website when they’re doing back-of-the-envelope math, and they’re not going to engage with you under the expectation that you’re going to give them a 98% discount.

So, figure out what the story is. And it’s like beating my head against the wall. I also want to be fair. These networks are very hard to build, and there’s a tremendous amount of investment. The AWS network is clearly magic in some respects just because having worked in data centers myself, the things that I see that I’m able to do between various EC2 instances at full network line rate would not have been possible in the data centers that I worked within.

So, there’s something going on that is magic and that’s great. And I understand that it’s expensive, but they’ve done a terrible job of messaging that. It just feels like, oh, bandwidth in is free because, you know, that’s how it works. Sending it out, ooh, that’s going to cost you X and their entire positioning and philosophy around it just feels unnecessary.

Rachel: That’s super interesting. And I think that also speaks to one of the questions that is still an open concern for what happens to Cloudflare if this is wildly successful. Which, based off of people’s excitement levels at this point, it’s seems like it’s very potentially going to be successful. And what does this mean for the level of investment that they’re going to have to make in their own infrastructure and network and order to actually be able to serve all of this?

Corey: The thing that I find curious is that in a couple of comment threads on Hacker News and on Twitter, Cloudflare’s CEO, Matthew Prince—who’s always been extremely accessible as far as executives of giant cloud companies go—has said that at their scale and by which they he’s referring to Cloudflare, and he says, “I assume that Amazon can probably get at least as decent economics on bandwidth pricing as we can,” which is a gross understatement because Amazon will spend years fighting over 50 cents.

Great, but what’s interesting is that he refers to bandwidth at that scale as being much closer to a fixed cost than something that’s a marginal cost for everything that a customer uses. The way that companies buy and sell bandwidth back and forth is complex, but he’s right. It is effectively a fixed monthly fee for a link and you can use as much or as little of that link as you want. 95th percentile billing aficionados, please don’t email me.

But by and large, that’s the way to think about it. You pay for the size of the pipe, not how much water flows through it. And as long as you can keep the links going without saturating them to the point where more data can’t fit through at a reasonable amount of time, your cost don’t change. So, yeah, if there’s a bunch of excitement they’ll have to expand the links, but that’s generally a fixed cost as opposed to a marginal cost per gigabyte.

That’s not how they think about it. There’s a whole translation layer that’s an economic model. And according to their public filings, they have something like a 77% gross margin which tells me that, okay, they are not in fact losing money on bandwidth even now where they generally don’t charge on a metered basis until you’re on the Cloudflare Enterprise Plan.

Rachel: Yeah. I think it’s going to just be really interesting to watch. I’m definitely interested to see what happens as they open this up, and like, 11 9s of availability feels like a lot of availabilities. It’s just the engineering of this, the economics of this it feels like there’s a lot of open questions that I’m excited to watch.

Corey: You’re onto my favorite part of this. So, the idea of 11 9s because it sounds ludicrous. That is well within the boundaries of probability of things such as, yeah, it is likely that gravity is going to stop working than it is that’s going to lose data. How can you guarantee that? Generally speaking, although S3 has always been extremely tight-lipped about how it works under the hood, other systems have not been.

And it looks an awful lot like the idea of Reed-Solomon erasure coding, where for those of us who spent time downloading large files of questionable legality due to copyright law and whatnot off of Usenet, they had the idea of parity files where they’d take these giant media files up—they’re Linux ISOs; of course they are—and you’d slice them into a bunch of pieces and then generate parity files as well.

So, you would wind up downloading the let’s say 80 RAR files and, oh, three of them were corrupt, each parity file could wind up swapping in so as long as you had enough that added up to 80, any of those could wind up restoring the data that had been corrupted. That is almost certainly what is happening at the large object storage scale. Which is great, we’re going to break this thing into a whole bunch of chunks. Let’s say here is a file you’ve uploaded or an object.

We’re going to break this into a hundred chunks—let’s say arbitrarily—and any 80 of those chunks can be used to reconstitute the entire file. And then you start looking at where you place them and okay, what are the odds of simultaneous drive failure in these however many locations? And that’s how you get that astronomical number. It doesn’t mean what people think of does. The S3 offers 11 9s of durability on their storage classes, including the One Zone storage class.

Which is a single availability zone instead of something that’s an entire region, which means that they’re not calculating disaster recovery failure scenarios into that durability number. Which is fascinating because it’s far, like, you’re going to have all the buildings within the same office park burn down than it is all of the buildings within a hundred square miles burn down, but those numbers remain the same.

There’s a lot of assumptions baked into that and it makes for an impressive talking point. I just hear it as, oh yeah, you’re a real object-store. That’s how I see it. There’s a lot that’s yet to be explained or understood. And I think that I’m going to be going up one side and down the other as soon as this exists in the real world and I’m looking forward to seeing it. I’m just a little skeptical because it has been preannounced.

The important part for me is even the idea that they can announce something like this and not be sued for securities fraud tells me that it is at least theoretically economically possible that they could be telling the truth on this. And that alone speaks volumes to just how out-of-bounds it tends to be in the context of giant cloud customers.

Rachel: I mean, if you read Matt Levine, “Everything is Securities Fraud?“ so, I don’t know how much we want to get excited about that.

Corey: Absolutely. A huge fan of his work.

Corey: You know its important to me that people learn how to use the cloud effectively. Thats why I’m so glad that Cloud Academy is sponsoring my ridiculous non-sense. They’re a great way to build in demand tech skills the way that, well personally, I learn best which I learn by doing not by reading. They have live cloud labs that you can run in real environments that aren’t going to blow up your own bill—I can’t stress how important that is. Visit cloudacademy.com/corey. Thats C-O-R-E-Y, don’t drop the “E.” Use Corey as a promo-code as well. You’re going to get a bunch of discounts on it with a lifetime deal—the price will not go up. It is limited time, they assured me this is not one of those things that is going to wind up being a rug pull scenario, oh no no. Talk to them, tell me what you think. Visit: cloudacademy.com/corey, C-O-R-E-Y and tell them that I sent you!

Corey: So, we’ve talked a fair bit about what data egress looks like. What else have you been focusing on? What have you found that is fun, and exciting, and catches your eye in this incredibly broad industry lately?

Rachel: Oh, there’s all kinds of exciting things. One of the pieces of research that’s been on my back-burner, usually I do it early summer, and it is—due to a variety of factors—still in my pipeline, but I always do a piece of research about base infrastructure pricing. And it’s an annual piece of looking about what are all of the cloud providers doing in regards to their pricing on that core aspect of compute, and storage, and memory.

And what does that look like over time, and what does that look like across providers? And it is absolutely impossible to get an apples-to-apples comparison over time and across providers. It just can’t actually be done. But we do our best [laugh] and then caveat the hell out of it from there. But that’s the piece of research that’s most on my backlog right now and one that I’m working on.

Corey: I think that there’s a lot of question around the idea of what is the cost of a compute unit—or something like that—between providers? The idea of if I have this configuration will cost me more on cloud provider a or cloud provider B, my pet working theory is that whenever people ask for analyses like that—or a number of others, to be perfectly frank with you—what they’re really looking for is confirmation bias to go in the direction that they wanted to go in already. I have yet to see a single scenario where people are trying to decide between cloud providers and they say, “That one because it’s going to be 10% less.” I haven’t seen it. That said I am, of course, at a very particular area of the industry. Have you seen it?

Rachel: I have not seen it. I think users find it interesting because it’s always interesting to look at trends over time. And in particular, with this analysis, it’s interesting to watch the number of providers narrow and then widen back out because we’ve been doing this since 2012. So, we used to have [unintelligible 00:26:24] and HPE used to be in there. So, like, we used to—CenturyLink. We used to have this broader list of cloud providers that we considered that would narrow down to this doesn’t really count anymore.

And now why do you need to back out? It’s like, okay, Oracle Cloud you’re in, Alibaba, Tencent, like, let’s look at you. And so, like, it’s interesting to just watch the providers in the mix shift over time which I think is interesting. And I think one of just the broad trends that is interesting is early years of this, there was steep competition on price, and that leveled off for solid three, four years.

We’ve seen some degree of competitiveness reemerge with competitors like Oracle in particular. So, those broad-brush trends are interesting. The specifics of the pricing if you’re doing 10% difference kind of things I think you’re missing the point of the analysis largely, but it’s interesting to look at what’s happening in the industry overall.

Corey: If you were to ask me to set up a simple web app, if there is such a thing, and tell you in advance what it was going to cost to host, and I can get it accurate within 20%, I am on fire in terms of both analysis and often dumb luck just because it is so difficult to answer the question. Getting back to our earlier conversational topic, let’s say I put CloudFront, Amazon’s sorry excuse for a CDN, in front of it which is probably the closest competitor they have to Cloudflare as a CDN, what’ll it cost me per gigabyte? Well, that’s a fascinating question. The answer comes down to where are you visiting it from? Depending where on the planet, people who are viewing my website, or using my web app are sitting, the cost per gigabyte will vary between eight-and-a-half cents—retail pricing—and fourteen cents. That’s a fairly wide margin and there’s no way to predict that in advance for most use cases. It’s the big open-ended question.

And people build out their environments and they want to know they’re making a rational decision and that their provider is not charging three times more than their competitor is for the exact same thing, but as long as it’s within a certain level of confidence interval, that makes sense.

Rachel: Yeah, and I think the other thing that’s interesting about this analysis and one of the reasons that it’s a frustrating analysis for me, in particular, is that I feel like that base compute is actually not where most of the cloud providers are actually competing anymore. So, like, it was definitely the interesting story early in cloud.

I think very clearly not the focus area for most of us now. It has moved up an abstraction layer. It’s moved to manage services. It has moved to other areas of their product portfolio. So, it’s still useful. It’s good to know. But I think that the broader portfolio of the cloud providers is definitely more the story than this individual price point.

Corey: That is an interesting story because I believe it, and it speaks to the aspirational version of where a lot of companies see themselves going. And then in practice, I see companies talking like this constantly, and then I look in their environment and say, “Okay, you’re basically spending 70% of your entire cloud bill on EC2 instances, running—it’s a bunch of VMs that sit there.”

And as much as they love to talk about the future and how other things are being considered and how their—use of machine learning in the rest, and Kubernetes, of course, a lot of this stuff all distills down to, yeah, it runs in software. It sits on top of EC2 instances and that’s what you get billed for. At re:Invent it’s always interesting and sad at the same time that they don’t give EC2 nearly enough attention or stage time because it’s not interesting, despite it being a majority of AWS bill.

Rachel: I think that’s a fantastic point, well made.

Corey: I’ll take it even one step further—and this is one where I think is almost a messaging failure on some level—Google Cloud offers sustained use discounts which apparently they don’t know how to talk about appropriately, but it’s genius. The way this works is if you run a VM for more than in a certain number of hours in a month, the entire month is now charged for that VM at a less than retail rate because you’ve been using it in a sustained way.

All you have to do to capture that is don’t turn it off. You know, what everyone’s doing already. And sure if you commit to usage on it you get a deeper discount, but what I like about this is if you buy some reserved instances is or you buy some committed use discounting, great, you’ll save more money, but okay, here’s a $20 million buy. You should click the button on, people are terrified to click at that button because I don’t usually get to approve dollar figure spend with multiple commas in them. That’s kind of scary. So, people hem and they haw and they wait six months. This is maybe not as superior mathematically, but it’s definitely an easier sell psychologically, and they just don’t talk about it.

It’s what people say they care about when people actually do are worlds apart. And the thing that continually astounds me because I didn’t
expect it, but it’s obvious in hindsight that when it comes to cloud economics it’s more about psychology than it is about math.

Rachel: I think one of the things that, having come from the finance world into the analyst world, and so I definitely have a particular point of view, but one of the things that was hardest for me when I worked in finance was not the absolute dollar amount of anything but the variability of it. So, if I knew what to expect I could work with that and we could make it work. It was when things varied in unexpected ways that it was a lot more challenging.

And so I think one of the things that when people talked about, like, this shift to cloud and the move to cloud, and everyone is like, “Oh, we’re moving things from the balance sheet to the income statement.” And everyone talked about that like it’s a big deal. For some parts of the organization that is a big deal, but for a lot of the organization, the shift that matters is the shift from a fixed cost to a variable cost because that lack of predictability makes a lot of people’s jobs, a lot more difficult.

Corey: The thing that I always find fun is a thought exercise is okay, let’s take a look at any given cloud company’s cloud bill for the last 18 to 36 months and add all of that up. Great, take that big giant number and add 20% to it. If you could magically go back in time and offer that larger number to them as here’s your cloud bill and all of your usage for the next 18 to 36 months. Here you go. Buy this instead.

And the cloud providers laugh at me and they say, “Who in the world would agree to that deal?” And my answer is, “Almost everyone.” Because at the company’s scale it’s not like the individual developer response of, “Oh, my God, I just spent how much money? I’ve got to eat this month.” Companies are used to absorbing those things. It’s fine. It’s just a, “We didn’t predict this. We didn’t plan for this. What does this do to our projections, our budget, et cetera?”

If you can offer them certainty and find some way to do it, they will jump at that. Most of my projects are not about make the bill lower, even though that is what is believed, in some cases by people working on these projects internally at these companies. It’s about making it understandable. It’s about making it predictable, it’s about understanding when you see a big spike one month. What project drove that?

Spoiler, it’s almost always the data science team because that’s what they do, but that’s neither here nor there. Please don’t send me letters. But yeah, it’s about understanding what is going on, and that understanding and being able to predict it is super hard when you’re looking at usage-based pricing.

Rachel: Exactly.

Corey: I want to thank you for taking so much time to speak with me. If people want to hear more about your thoughts, your observations, et cetera, where can they find you?

Rachel: Probably the easiest way to get in touch with me is on Twitter, which is @rstephensme that’s R-S-T-E-P-H-E-N-S-M-E.

Corey: And we will, of course, put links to that in the [show notes 00:34:08]. Thank you so much for your time. I appreciate it.

Rachel: Thanks for having me. This was great.

Corey: Rachel Stephens, senior analyst at RedMonk. I’m Cloud Economist, Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an angry comment, angrily defending your least favorite child, which is some horrifying cloud service you have launched during the pandemic.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Tim

Tim’s tech career spans over 20 years through various sectors. Tim’s initial journey into tech started as a US Marine. Later, he left government contracting for the private sector, working both in large corporate environments and in small startups. While working in the private sector, he honed his skills in systems administration and operations for large Unix-based datastores.

Today, Tim leverages his years in operations, DevOps, and Site Reliability Engineering to advise and consult with clients in his current role. Tim is also a father of five children, as well as a competitive Brazilian Jiu-Jitsu practitioner. Currently, he is the reigning American National and 3-time Pan American Brazilian Jiu-Jitsu champion in his division.

Transcript

Corey: Hello, and welcome to Screaming in the Cloud with your host, Chief cloud economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Vultr. Spelled V-U-L-T-R because they’re all about helping save money, including on things like, you know, vowels. So, what they do is they are a cloud provider that provides surprisingly high performance cloud compute at a price that—while sure they claim its better than AWS pricing—and when they say that they mean it is less money. Sure, I don’t dispute that but what I find interesting is that it’s predictable. They tell you in advance on a monthly basis what it’s going to going to cost. They have a bunch of advanced networking features. They have nineteen global locations and scale things elastically. Not to be confused with openly, because apparently elastic and open can mean the same thing sometimes. They have had over a million users. Deployments take less that sixty seconds across twelve pre-selected operating systems. Or, if you’re one of those nutters like me, you can bring your own ISO and install basically any operating system you want. Starting with pricing as low as $2.50 a month for Vultr cloud compute they have plans for developers and businesses of all sizes, except maybe Amazon, who stubbornly insists on having something to scale all on their own. Try Vultr today for free by visiting: vultr.com/screaming, and you’ll receive a $100 in credit. Thats v-u-l-t-r.com slash screaming.

Corey: This episode is sponsored in part by something new. Cloud Academy is a training platform built on two primary goals. Having the highest quality content in tech and cloud skills, and building a good community the is rich and full of IT and engineering professionals. You wouldn’t think those things go together, but sometimes they do. Its both useful for individuals and large enterprises, but here's what makes it new. I don’t use that term lightly. Cloud Academy invites you to showcase just how good your AWS skills are. For the next four weeks you’ll have a chance to prove yourself. Compete in four unique lab challenges, where they’ll be awarding more than $2000 in cash and prizes. I’m not kidding, first place is a thousand bucks. Pre-register for the first challenge now, one that I picked out myself on Amazon SNS image resizing, by visiting cloudacademy.com/corey. C-O-R-E-Y. That’s cloudacademy.com/corey. We’re gonna have some fun with this one!

Corey: Welcome to Screaming in the Cloud. I am Cloud Economist Corey Quinn joined by Principal Cloud Economist here at The Duckbill Group Tim Banks. Tim, how are you?

Tim: I’m doing great, Corey. How about yourself?

Corey: I am tickled pink that we are able to record this not for the usual reasons you would expect, but because of the glorious pun in calling this our Banksgiving episode. I have a hard and fast rule of, I don’t play pun games or make jokes about people’s names because that can be an incredibly offensive thing. “And oh, you’re making jokes about my name? I’ve never heard that one before.” It’s not that I can’t do it—I play games with language all the time—but it makes people feel crappy. So, when you suggested this out of the blue, it was yes, we’re doing it. But I want to be clear, I did not inflict this on you. This is your own choice; arguably a poor one. We’re going to find out.

Tim: 1000% my idea.

Corey: So, this is your show. It’s a holiday week. So, what do you want to do with our Banksgiving episode?

Tim: I want to give thanks for the folks who don’t normally get acknowledged through the year. Like you know, we do a lot of thanking the rock stars, we do a lot of thanking the big names, right, we also do a lot of, you know, some snarky jabs at some folks. Deservingly—not folks, but groups and stuff like that; some folks deserve it, and we won’t be giving them thanks—but some orgs and some groups and stuff like that. And I do think with that all said, we should acknowledge and thank the folks that we normally don’t get to, folks who’ve done some great contributions this year, folks who have helped us, helped the industry, and help services that go unsung, I think a great one that you brought up, it’s not the engineers, right? It’s the people that make sure we get paid. Because I don’t work for charity. And I don’t know about you, Corey. I haven’t seen the books yet, but I’m pretty sure none of us here do and so how do we get paid? Like I don’t know.

Corey: Oh, sure you have. We had a show on a somewhat simplified P&L during the all hands meeting because, you know, transparency matters. But you’re right, those are numbers there and none of that is what we could have charged but didn’t because we decided to do more
volunteer work for AWS. If we were going to go down that path, we would just be Community Heroes and be done with it.

Tim: That’s true. But you know, it’s like, I do my thing and then, you know, I get a paycheck every now and then. And so, as far as I know, I think most of that happens because of Dan.

Corey: Dan is a perfect example. He’s been a guest on this show, I don’t know it has as aired at the time that this goes out because I don’t have to think about that, which is kind of the point. Dan’s our CFO and makes sure that a lot of the financial trains keep running on time. But let’s also be clear, the fact that I can make predictions about what the business is going to be doing by a metric other than how much cash is in the bank account at this very moment really freed up some opportunity for us. It turned into adult supervision for folks who, when I started this place and then Mike joined, and it was very much not an area that either one of us was super familiar with. Which is odd given what we do here, but we learned quickly.

The understanding not just how these things work—which we had an academic understanding of—but why it mattered and how that applies to real life. Finance is one of those great organizations that doesn’t get a lot of attention or respect outside of finance itself. Because it’s, “Oh, well they just control the money. How hard could it be?” Really, really hard.

Tim: It really is. And when we dig into some of these things and some of the math that goes and some of what the concerns are that, you know, a lot of engineers don’t really have a good grasp on, and it’s eye opening to understand some of the concerns. At least some of the concerns at least from an engineering aspect. And I really don’t give much consideration day to day about the things that go on behind the scenes to make sure that I get paid.

But you look at this throughout the industry, like, how many of the folks that we work with, how many folks out there doing this great work for the industry, do they know who their payroll person is? Do they know who their accountant team is? Do they know who their CFO or the other people out there that are doing the work and making sure the lights stay on, that people get paid and all the other things that happen, right? You know, people take that for granted. And it’s a huge work and those people really don’t get the appreciation that I think they deserve. And I think it’s about time we did that.

Corey: It’s often surprising to me how many people that I encounter, once they learn that there are 12 employees here, automatically assume that it’s you, me, and maybe occasionally Mike doing all the work, and the other nine people just sort of sit here and clap when I tell a funny joke, and… well, yes, that is, of course, a job duty, but that’s not the entire purpose of why people are here.

Natalie in marketing is a great example. “Well, Corey, I thought you did the marketing. You go and post on Twitter and that’s where business comes from.” Well, kind of. But let’s be clear, when I do that, and people go to the website to figure out what the hell I’m talking about.

Well, that website has words on it. I didn’t put those words on that site. It directs people to contact us forms, and there are automations behind that that make sure they go to the proper place because back before I started this place and I was independent, people would email me asking for help with their bill and I would just never respond to them. It’s the baseline adult supervision level of competence that I keep aspiring to. We have a sales team that does fantastic work.

And that often is one of those things that’ll get engineering hackles up, but they’re not out there cold-calling people to bug them about AWS bills. It’s when someone reaches out saying we have a problem with our AWS spend, can you help us? The answer is invariably, “Let’s talk about that.” It’s a consultative discussion about why do you care about the bill, what does success look like, how do you know this will be a success, et cetera, et cetera, et cetera, that make sure that we’re aimed at the right part of the problem. That’s incredibly challenging work and I am grateful beyond words, I don’t have to be involved with the day-in, day-out of any of those things.

Tim: I think even beyond just that handling, like, the contracts and the NDAs, and the various assets that have to be exchanged just to get us virtually on site, I’ve [unintelligible 00:06:46] a couple of these things, I’m glad it’s not my job. It is, for me, overwhelmingly difficult for me to really get a grasp and all that kind of stuff. And I am grateful that we do have a staff that does that. You’ve heard me, you see me, you know, kind of like, sales need to do better, and a lot of times I do but I do want to make sure we are appreciating them for the work that they do to make sure that we have work to do. Their contribution cannot be underestimated.

Corey: And I think that’s something that we could all be a little more thankful for in the industry. And I see this on Twitter sometimes, and it’s probably my least favorite genre of tweet, where someone will wind up screenshotting some naive recruiter outreach to them, and just start basically putting the poor person on blast. I assure you, I occasionally get notices like that. The most recent example of that was, I got an email to my work email address from an associate account exec at AWS asking what projects I have going on, how my work in the cloud is going, and I can talk to them about if I want to help with cost optimization of my AWS spend and the rest. And at first, it’s one of those, I could ruin this person’s entire month, but I don’t want to be that person.

And I did a little LinkedIn stalking and it turns out, this looks like this person’s first job that they’ve been in for three months. And I’ve worked in jobs like that very early in my career; it is a numbers game. When you’re trying to reach out to 1000 people a month or whatnot, you aren’t sitting there googling what every one of them is, does, et cetera. It’s something that I’ve learned, that is annoying, sure. But I’m in an incredibly privileged position here and dunking on someone who’s doing what they are told by an existing sales apparatus and crapping on them is not fair.

That is not the same thing as these passive-aggressive [shit-tier 00:08:38] drip campaigns of, “I feel like I’m starting to stalk you.” Then don’t send the message, jackhole. It’s about empathy and not crapping on people who are trying to find their own path in this ridiculous industry.

Tim: I think you brought up recruiters, and, you know, we here at The Duckbill Group are currently recruiting for a senior cloud economist and we don’t actually have a recruiter on staff. So, we’re going through various ways to find this work and it has really made me appreciate the work that recruiters in the past that I’ve worked with have done. Some of the ones out there are doing really fantastic work, especially sourcing good candidates, vetting good candidates, making sure that the job descriptions are inclusive, making sure that the whole recruitment process is as smooth as it can be. And it can’t always be. Having to deal with all the spinning plates of getting interviews with folks who have production workloads, it is pretty impressive to me to see how a lot of these folks get—pull it off and it just seems so smooth. Again, like having to actually wade through some of this stuff, it’s given me a true appreciation for the work that good recruiters do.

Corey: We don’t have automated systems that disqualify folks based on keyword matches—I’ve never been a fan of that—but we do get applicants that are completely unsuitable. We’ve had a few come in that are actual economists who clearly did not read the job description; they’re spraying their resume everywhere. And the answer is you smile, you decline it and you move on. That is the price you pay of attempting to hire people. You don’t put them on blast, you don’t go and yell at an entire ecosystem of people because looking for jobs sucks. It’s hard work.

Back when I was in my employee days, I worked harder finding new jobs than I often did in the jobs themselves. This may be related to why I get fired as much, but I had to be good at finding new work. I am, for better or worse, in a situation where I don’t have to do that anymore because once again, we have people here who do the various moving parts. Plus, let’s be clear here, if I’m out there interviewing at other
companies for jobs, I feel like that sends a message to you and the rest of the team that isn’t terrific.

Tim: We might bring that up. [laugh].

Corey: “Why are you interviewing for a job over there?” It’s like, “Because they have free doughnuts in the office. Later, jackholes.” It—I don’t think that is necessarily the culture we’re building here.

Tim: No, no, it’s not. Specially—you know, we’re more of a cinnamon roll culture anyways.

Corey: No. In my case, it’s one of those, “Corey, why are you interviewing for a job at AWS?” And the answer is, “Oh, it’s going to be an amazing shitpost. Just wait and watch.”

Tim: [laugh]. Now, speaking of AWS, I have to absolutely shout out to Emily Freeman over there who has done some fantastic work this year. It’s great when you see a person get matched up with the right environment with the right team in the right role, and Emily has just been hitting out of the park ever since he got there, so I’m super, super happy to see her there.

Corey: Every time I get to collaborate with her on something, I come away from the experience even more impressed. It’s one of those phenomenal collaborations. I just—I love working with her. She’s human, she’s empathetic, she gets it. She remains, as of this recording, the only person who has ever given a talk that I have heard on ML Ops, and come away with a better impression of that space and thinking maybe it’s not complete nonsense.

And that is not just because it’s Emily, so I—because—I’m predisposed to believe her, though I am, it’s because of how she frames it, how she views these things, and let’s be clear, the content that she says. And that in turn makes me question my preconceptions on this, and that is why she has that I will listen and pay attention when she speaks. So yeah, if Emily’s going to try and make a point, there’s always going to be something behind it. Her authenticity is unimpeachable.

Tim: Absolutely. I do take my hat’s off to everyone who’s been doing DevRel and evangelism and those type of roles during pandemics. And we just, you know, as the past few months, I’ve started back to in-person events. But the folks who’ve been out there finding new way to do
those jobs, finding a way to [crosstalk 00:12:50]—

Corey: Oh, staff at re:Invent next week. Oh, my God.

Tim: Yeah. Those folks, I don’t know how they’re being rewarded for their work, but I can assure you, they probably need to be [unintelligible 00:12:57] better than they are. So, if you are staff at re:Invent, and you see Corey and I, next week when we’re there—if you’re listening to this in time—we would love to shake your hand, elbow bump you, whatever it is you’re comfortable with, and laud you for the work you’re doing. Because it is not easy work under the best of circumstances, and we are certainly not under the best of circumstances.

Corey: I also want to call out specific thanks to a group that might take some people aback. But that group is AWS marketing, which given how much grief I give them seems like an odd thing for me to say, but let’s be clear, I don’t have any giant companies whose ability to continue as a going concern is dependent upon my keeping systems up and running. AWS does. They have to market and tell stories to everyone because that is generally who their customers are: they round to everyone. And an awful lot of those companies have unofficial mottos of, “That’s not funny.” I’m amazed that they can say anything at all, given how incredibly varied their customer base is, I could get away with saying whatever I want solely because I just don’t care. They have to care.

Tim: They do. And it’s not only that they have to care, they’re in a difficult situation. It’s like, you know, they—every company that sizes is, you know, they are image conscious, and they have things that say what like, “Look, this is the deal. This is the scenario. This is how it went down, but you can still maintain your faith and confidence in us.” And people do when AWS services, they have problems, if anything comes out like that, it does make the news and the reason it doesn’t make the news is because it is so rare. And when they can remind us of that in a very effective way, like, I appreciate that. You know, people say if anything happens to S3, everybody knows because everyone depends on it and that’s for good reason.

Corey: And let’s not forget that I run The Duckbill Group. You know, the company we work for. I have the Last Week in AWS newsletter and blog. I have my aggressive shitposting Twitter feed. I host the AWS Morning Brief podcast, and I host this Screaming in the Cloud. And it’s challenging for me to figure out how to message all of those things because when people ask what you do, they don’t want to hear a litany that goes on for 25 seconds, they want a sentence.

I feel like I’ve spread in too many directions and I want to narrow that down. And where do I drive people to and that was a bit of a marketing challenge that Natalie in our marketing department really cut through super well. Now, pretend I work in AWS. The way that I check this based upon a public list of parameters they stub into Systems Manager Parameter Store, there are right now 291 services that they offer. That is well beyond any one person’s ability to keep in their head. I can talk incredibly convincingly now about AWS services that don’t exist and people who work in AWS on messaging, marketing, engineering, et cetera, will not call me out on it because who can provably say that ‘AWS Strangle Pony’ isn’t a real service.

Tim: I do want to call out the DevOps—shout out I should say, the DevOps term community for AWS Infinidash because that was just so well done, and AWS took that with just the right amount of tongue in cheek, and a wink and a nod and let us have our fun. And that was a good time. It was a great exercise in improv.

Corey: That was Joe Nash out of Twilio who just absolutely nailed it with his tweet, “I am convinced that a small and dedicated group of Twitter devs could tweet hot takes about a completely made up AWS product—I don’t know AWS Infinidash or something—and it would appear as a requirement on job specs within a week.” And he was right.

Tim: [laugh]. Speaking of Twitter, I want to shout out Twitter as a company or whoever does a product management over there for Twitter Spaces. I remember when Twitter Spaces first came out, everyone was dubious of its effect, of it’s impact. They were calling it, you know, a Periscope clone or whatever it was, and there was a lot of sneering and snarking at it. But Twitter Spaces has become very, very effective in having good conversations in the group and the community of folks that have just open questions, and then to speak to folks that they probably wouldn’t only get to speak to about this questions and get answers, and have really helpful, uplifting and difficult conversations that you wouldn’t otherwise really have a medium for. And I’m super, super happy that whoever that product manager was, hats off to you, my
friend.

Corey: One group you’re never going to hear me say a negative word about is AWS support. Also, their training and certification group. I know that are technically different orgs, but it often doesn’t feel that way. Their job is basically impossible. They have to teach people—even on the support side, you’re still teaching people—how to use all of these different varied services in different ways, and you have to do it in the face of what can only really be described as abuse from a number of folks on Twitter.

When someone is having trouble with an AWS service, they can turn into shitheads, I’ve got to be honest with you. And berating the poor schmuck who has to handle the AWS support Twitter feed, or answer your insulting ticket or whatnot, they are not empowered to actually fix the underlying problem with a service. They are effectively a traffic router to get the message to someone who can, in a format that is understood internally. And I want to be very clear that if you insult people who are in customer service roles and blame them for it, you’re just being a jerk.

Tim: No, it really is because I’m pretty sure a significant amount of your listeners and people initially started off working in tech support, or customer service, or help desk or something like that, and you really do become the dumping ground for the customers’ frustrations because you are the only person they get to talk to. And you have to not only take that, but you have to try and do the emotional labor behind soothing them as well as fixing the actual problem. And it’s really, really difficult. I feel like the people who have that in their background are some of the best consultants, some of the best DevRel folks, and the best at talking to people because they’re used to being able to get some technical details out of folks who may not be very technical, who may be under emotional distress, and certainly in high stress situations. So yeah, AWS support, really anybody who has support, especially paid support—phone or chat otherwise—hats off again. That is a service that is thankless, it is a service that is almost always underpaid, and is almost always under appreciated.

Corey: This episode is sponsored by our friends at Oracle HeatWave is a new high-performance accelerator for the Oracle MySQL Database Service. Although I insist on calling it “my squirrel.” While MySQL has long been the worlds most popular open source database, shifting from transacting to analytics required way too much overhead and, ya know, work. With HeatWave you can run your OLTP and OLAP, don’t ask me to ever say those acronyms again, workloads directly from your MySQL database and eliminate the time consuming data movement and integration work, while also performing 1100X faster than Amazon Aurora, and 2.5X faster than Amazon Redshift, at a third of the cost. My thanks again to Oracle Cloud for sponsoring this ridiculous nonsense.

Corey: I’ll take another team that’s similar to that respect: Commerce Platform. That is the team that runs all of AWS billing. And you would be surprised that I’m thanking them, but no, it’s not the cynical approach of, “Thanks for making it so complicated so I could have a business.” No, I would love it if it were so simple that I had to go find something else to do because the problem was that easy for customers to solve. That is the ideal and I hope, sincerely, that we can get there.

But everything that happens in AWS has to be metered and understood as far as who has done what, and charge people appropriately for it. It is also generally invisible; people don’t understand anything approaching the scale of that, and what makes it worst of all, is that if suddenly what they were doing broke and customers weren’t built for their usage, not a single one of them would complain about it because, “All right, I’ll take it.” It’s a thankless job that is incredibly key and central to making the cloud work at all, but it’s a hard job.

Tim: It really is. And is a lot of black magic and voodoo to really try and understand how this thing works. There’s no simple way to explain it. I imagine if they were going to give you the index overview of how it works with a 10,000 feet, that alone would be, like, a 300 page document. It is a gigantic moving beast.

And it is one of those things where scale will show all the flaws. And no one has scale I think like AWS does. So, the folks that have to work and maintain that are just really, again, they’re under appreciated for all that they do. I also think that—you know, you talk about the same thing in other orgs, as we talked about the folks that handle the billing and stuff like that, but you mentioned AWS, and I was thinking the other day how it’s really awesome that I’ve got my AWS driver. I have the same, like, group of three or four folks that do all my deliveries for AWS.

And they have been inundated over this past year-and-a-half with more and more and more stuff. And yet, I’ve still managed—my stuff is always put down nicely on my doorstep. It’s never thrown, it’s not damaged. I’m not saying it’s never been damaged, but it’s not damaged, like, maybe FedEx I’ve [laugh] had or some other delivery services where it’s just, kind of, carelessly done. They still maintain efficiency, they maintain professionalism [unintelligible 00:21:45] talking to folks.

What they’ve had to do at their scale and at that the amount of stuff they’ve had to do for deliveries over this past year-and-a-half has just been incredible. So, I want to extend it also to, like, the folks who are working in the distribution centers. Like, a lot of us here talk about AWS as if that’s Amazon, but in essence, it is those folks that are working those more thankless and invisible jobs in the warehouses and fulfillment centers, under really bad conditions sometimes, who’s still plug away at it. I’m glad that Amazon is at least saying they’re making efforts to improve the conditions there and improve the pay there, things like that, but those folks have enabled a lot of us to work during this pandemic
with a lot of conveniences that they themselves would never be able to enjoy.

Corey: Yeah. It’s bad for society, but I’m glad it exists, obviously. The thing is, I would love it if things showed up a little more slowly if it meant that people could be treated humanely along the process. That said, I don’t have any conception of what it takes to run a company with 1.2 million people.

I have learned that as you start managing groups and managing managers of groups, it’s counterintuitive, but so much of what you do is no longer you doing the actual work. It is solely through influence and delegation. You own all of the responsibility but no direct put-finger-on-problem capability of contributing to the fix. It takes time at that scale, which is why I think one of the dumbest series of questions from, again, another group that deserves a fair bit of credit which is journalists because this stuff is hard, but a naive question I hear a lot is, “Well, okay. It’s been 100 days. What has Adam Selipsky slash Andy Jassy changed completely about the company?”

It’s, yeah, it’s a $1.6 trillion company. They are not going to suddenly grab the steering wheel and yank. It’s going to take years for shifts that they do to start manifesting in serious ways that are externally visible. That is how big companies work. You don’t want to see a complete change in direction from large blue chip companies that run things. Like, again, everyone’s production infrastructure. You want it to be predictable, you want it to be boring, and you want shifts to be gradual course corrections, not vast swings.

Tim: I mean, Amazon is a company with a population of a medium to medium-large sized city and a market cap of the GDP of several countries. So, it is not a plucky startup; it is not this small little tech company. It is a vast enterprise that’s distributed all over the world with a lot of folks doing a lot of different jobs. You cannot, as you said, steer that ship quickly.

Corey: I grew up in Maine and Amazon has roughly the same number employees as live in Maine. It is hard to contextualize how all of that works. There are people who work there that even now don’t always know who Andy Jassy is. Okay, fine, but I’m not talking about don’t know him on site or whatever. I’m saying they do not recognize the name. That’s a very big company.

Tim: “Andy who?”

Corey: Exactly. “Oh, is that the guy that Corey makes fun of all the time?” Like, there we go. That’s what I tend to live for.

Tim: I thought that was Werner.

Corey: It’s sort of every one, though I want to be clear, I make it a very key point. I do not make fun of people personally because it—even if they’re crap, which I do not believe to be the case in any of the names we’ve mentioned so far, they have friends and family who love and care about them. You don’t want someone to go on the internet and Google their parent’s name or something, and then just see people crapping all over. That’s got to hurt. Let people be people. And, on some level, when you become the CEO of a company of that scale, you’re stepping out of reality and into the pages of legend slash history, at some point. 200 years from now, people will read about you in history books, that’s a wild concept.

Tim: It is I think you mentioned something important that we would be remiss—especially Duckbill Group—to mention is that we’re very thankful for our families, partners, et cetera, for putting up with us, pets, everybody. As part of our jobs, we invite strangers from the internet into our homes virtually to see behind us what is going on, and for those of us that have kids, that involves a lot of patience on their part, a lot of patients on our partners’ parts, and other folks that are doing those kind of nurturing roles. You know, our pets who want to play with us are sitting there and not able to. It has not been easy for all of us, even though we’re a remote company, but to work under these conditions that we have been over the past year-and-a-half. And I think that goes for a lot of the folks in industry where now all of a sudden, you’ve been occupying a room in the house or space in the house for some 18-plus months, where before you’re always at work or something like that. And that’s been a hell of an adjustment. And so we talk about that for us folks that are here pontificating on podcasts, or banging out code, but the adjustments and the things our families have had to go through and do to tolerate us being there cannot be overstated how important that is.

Corey: Anyone else that’s on your list of people to thank? And this is the problem because you’re always going to forget people. I mean, the podcast production crew: the folks that turn our ramblings into a podcast, the editing, the transcription, all of it; the folks that HumblePod are
just amazing. The fact that I don’t have to worry about any of this stuff as if by magic, means that you’re sort of insulated from it. But it’s amazing to watch that happen.

Tim: You know, honestly, I super want to thank just all the folks that take the time to interact with us. We do this job and Corey shitposts, and I shitpost and we talk, but we really do this and rely on the folks that do take the time to DM us, or tweet us, or mention us in the thread, or reach out in any way to ask us questions, or have a discussion with us on something we said, those folks encourage us, they keep us accountable, and they give us opportunities to learn to be better. And so I’m grateful for that. It would be—this role, this job, the thing we do where we’re viewable and seen by the public would be a lot less pleasant if it wasn’t for y’all. So, it’s too many to name, but I do appreciate you.

Corey: Well, thank you, I do my best. I find this stuff to be so boring if you couldn’t have fun with it. And so many people can’t have fun with it, so it feels like I found a cheat code for making enterprise software solutions interesting. Which even saying that out loud sounds like I’m shitposting. But here we are.

Tim: Here we are. And of course, my thanks to you, Corey, for reaching out to me one day and saying, “Hey, what are you doing? Would you want to come interview with us at The Duckbill Group?”

Corey: And it was great because, like, “Well, I did leave AWS within the last 18 months, so there might be a non-compete issue.” Like, “Oh, please, I hope so. Oh, please, oh, please, oh, please. I would love to pick that fight publicly.” But sadly, no one is quite foolish enough to take me up on it.

Don’t worry. That’s enough of a sappy episode, I think. I am convinced that our next encounter on this podcast will be our usual aggressive self. But every once in a while it’s nice to break the act and express honest and heartfelt appreciation. I’m really looking forward to next week with all of the various announcements that are coming out.

I know people have worked extremely hard on them, and I want them to know that despite the fact that I will be making fun of everything that they have done, there’s a tremendous amount of respect that goes into it. The fact that I can make fun of the stuff that you’ve done without any fear that I’m punching down somehow because, you know it is at least above a baseline level of good speaks volumes. There are providers I absolutely do not have that confidence towards them.

Tim: [laugh]. Yeah, AWS, as the enterprise level service provider is an easy target for a lot of stuff. The people that work there are not. They do great work. They’ve got amazing people in all kinds of roles there. And they’re often unseen for the stuff they do. So yeah, for all the folks who have contributed to what we’re going to partake in at re:Invent—and it’s a lot and I understand from having worked there, the pressure that’s put on you for this—I’m super stoked about it and I’m grateful.

Corey: Same here. If I didn’t like this company, I would not have devoted years to making fun of it. Because that requires a diagnosis, not a newsletter, podcast, or shitposting Twitter feed. Tim, thank you so much for, I guess, giving me the impetus and, of course, the amazing name of the show to wind up just saying thank you, which I think is something that we could all stand to do just a little bit more of.

Tim: My pleasure, Corey. I’m glad we could run with this. I’m, as always, happy to be on Screaming in the Cloud with you. I think now I get a vest and a sleeve. Is that how that works now?

Corey: Exactly. Once you get on five episodes, then you end up getting the dinner jacket, just, like, hosting SNL. Same story. More on that to come in the new year. Thanks, Tim. I appreciate it.

Tim: Thank you, Corey.

Corey: Tim Banks, principal cloud economist here at The Duckbill Group. I am, of course, Corey Quinn, and thank you for listening.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Stephanie

Stephanie Wong is an award-winning speaker, engineer, pageant queen, and hip hop medalist. She is a leader at Google with a mission to blend storytelling and technology to create remarkable developer content. At Google, she's created over 400 videos, blogs, courses, and podcasts that have helped developers globally. You might recognize her as the host of the GCP Podcast. Stephanie is active in her community, fiercely supporting women in tech and mentoring students.

Links:

  • Personal Website: https://stephrwong.com
  • Twitter: https://twitter.com/stephr_wong

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Vultr. Spelled V-U-L-T-R because they’re all about helping save money, including on things like, you know, vowels. So, what they do is they are a cloud provider that provides surprisingly high performance cloud compute at a price that—while sure they claim its better than AWS pricing—and when they say that they mean it is less money. Sure, I don’t dispute that but what I find interesting is that it’s predictable. They tell you in advance on a monthly basis what it’s going to going to cost. They have a bunch of advanced networking features. They have nineteen global locations and scale things elastically. Not to be confused with openly, because apparently elastic and open can mean the same thing sometimes. They have had over a million users. Deployments take less that sixty seconds across twelve pre-selected operating systems. Or, if you’re one of those nutters like me, you can bring your own ISO and install basically any operating system you want. Starting with pricing as low as $2.50 a month for Vultr cloud compute they have plans for developers and businesses of all sizes, except maybe Amazon, who stubbornly insists on having something to scale all on their own. Try Vultr today for free by visiting: vultr.com/screaming, and you’ll receive a $100 in credit. Thats v-u-l-t-r.com slash screaming.

Corey: This episode is sponsored by our friends at Oracle Cloud. Counting the pennies, but still dreaming of deploying apps instead of "Hello, World" demos? Allow me to introduce you to Oracle's Always Free tier. It provides over 20 free services and infrastructure, networking, databases, observability, management, and security. And—let me be clear here—it's actually free. There's no surprise billing until you intentionally and proactively upgrade your account. This means you can provision a virtual machine instance or spin up an autonomous database that manages itself all while gaining the networking load, balancing and storage resources that somehow never quite make it into most free tiers needed to support the application that you want to build. With Always Free, you can do things like run small scale applications or do proof-of-concept testing without spending a dime. You know that I always like to put asterisks next to the word free. This is actually free, no asterisk. Start now. Visit snark.cloud/oci-free that's snark.cloud/oci-free.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. One of the things that makes me a little weird in the universe is that I do an awful lot of… let’s just call it technology explanation slash exploration in public, and turning it into a bit of a brand-style engagement play. What makes this a little on the weird side is that I don’t work for a big company, which grants me a tremendous latitude. I have a whole lot of freedom that lets me be all kinds of different things, and I can’t get fired, which is something I’m really good at.

Inversely, my guest today is doing something remarkably similar, except she does work for a big company and could theoretically be fired if they were foolish enough to do so. But I don’t believe that they are. Stephanie Wong is the head of developer engagement at Google. Stephanie, thank you for volunteering to suffer my slings and arrows about all of this.

Stephanie: [laugh]. Thanks so much for having me today, Corey.

Corey: So, at a very high level, you’re the head of developer engagement, which is a term that I haven’t seen a whole lot of. Where does that start and where does that stop?

Stephanie: Yeah, so I will say that it’s a self-proclaimed title a bit because of the nuance of what I do. I would say at its heart, I am still a part of developer relations. If you’ve heard of developer advocacy or developer evangelist, I would say this slight difference in shade of what I do is that I focus on scalable content creation and becoming a central figure for our developer audiences to engage and enlighten them with content that, frankly, is remarkable, and that they’d want to share and learn about our technology.

Corey: Your bio is fascinating in that it doesn’t start with the professional things that most people do with, “This is my title and this is my company,” is usually the first sentence people put in. Yours is, “Stephanie Wong is an award-winning speaker, engineer, pageant queen, and hip hop medalist.” Which is both surprising and more than a little bit refreshing because when I read a bio like that my immediate instinctive reaction is, “Oh, thank God. It’s a real person for a change.” I like the idea of bringing the other aspects of what you are other than, “This is what goes on in an IDE, the end,” to your audience.

Stephanie: That is exactly the goal that I had when creating that bio because I truly believe in bringing more interdisciplinary and varied backgrounds to technology. I, myself have gone through a very unconventional path to get to where I am today and I think in large part, my background has had a lot to do with my successes, my failures, and really just who I am in tech as an uninhibited and honest, credible person today.

Corey: I think that there’s a lack of understanding, broadly, in our industry about just how important credibility and authenticity are and even the source of where they come from. There are a lot of folks who are in the DevRel space—devrelopers, as I insist upon calling them, over their protests—where, on some level, the argument is, what is developer relations? “Oh, you work in marketing, but they’re scared to tell you,” has been my gag on that one for a while. But they speak from a position of, “I know what’s what because I have been in the trenches, working on these large-scale environments as an engineer for the last”—fill in the blank, however long it may have been—“And therefore because I have done things, I am going to tell you how it is.” You explicitly call out that you don’t come from the traditional, purely technical background. Where did you come from? It’s unlikely that you’ve sprung fully-formed from the forehead of some god, but again, I’m not entirely sure how Google finds and creates the folks that it winds up advancing, so maybe you did.

Stephanie: Well, to tell you the truth. We’ve all come from divine creatures. And that’s where Google sources all employees. So. You know. But—[laugh].

Corey: Oh, absolutely. “We climbed to the top of Olympus and then steal fire from the gods.” “It’s like, isn’t that the origin story of Prometheus?” “Yeah, possibly.” But what is your background? Where did you come from?

Stephanie: So, I have grown up, actually, in Silicon Valley, which is a little bit ironic because I didn’t go to school for computer science or really had the interest in becoming an engineer in school. I really had no idea.

Corey: Even been more ironic than that because most of Silicon Valley appears to never have grown up at all.

Stephanie: [laugh]. So, true. Maybe there’s a little bit of that with me, too. Everybody has a bit of Peter Pan syndrome here, right? Yeah, I had no idea what I wanted to do in school and I just knew that I had an interest in communicating with one another, and I ended up majoring in communication studies.

I thought I wanted to go into the entertainment industry and go into production, which is very different and ended up doing internships at Warner Brothers Records, a YouTube channel for dance—I’m a dancer—and I ended up finding a minor in digital humanities, which is sort of this interdisciplinary minor that combines technology and the humanities space, including literature, history, et cetera. So, that’s where I got my start in technology, getting an introduction to information systems and doing analytics, studying social media for certain events around the world. And it wasn’t until after school that I realized that I could work in enterprise technology when I got an offer to be a sales engineer. Now, that being said, I had no idea what sales engineering was. I just knew it had something to do with enterprise technology and communications, and I thought it was a good fit for my background.

Corey: The thing that I find so interesting about that is that it breaks the mold of what people expect, when, “If someone’s going to talk to me about technology—especially coming from a”—it’s weird; it’s one of the biggest companies on the planet, and people still on some level equate Google with the startup-y mentality of being built in someone’s garage. That’s an awfully big garage these days, if that’s even slightly close to true, which it isn’t. But there’s this idea of, “Oh, you have to go to Stanford. You have to get a degree in computer science. And then you have to go and do this, this, this, this, and this.”

And it’s easy to look dismissively at what you’re doing. “Communications? Well, all that would teach you to do is communicate to people clearly and effectively. What possible good is that in tech?” As we look around the landscape and figure out exactly why that is so necessary in tech, and also so lacking?

Stephanie: Exactly. I do think it’s an underrated skill in tech. Maybe it’s not so much anymore, but I definitely think that it has been in the past. And even for developers, engineers, data scientists, other technical practitioner, especially as a person in DevRel, I think it’s such a valuable skill to be able to communicate complex topics simply and understandably to a wide variety of audiences.

Corey: The big question that I have for you because I’ve talked to an awful lot of folks who are very concerned about the way that they approach developer relations, where—they’ll have ratios, for example—where I know someone and he insists that he give one deeply technical talk for every four talks that are not deeply technical, just because he feels the need to re-establish and shore up his technical bona fides. Now, if there’s one thing that people on the internet love, it is correcting people on things that are small trivia aspect, or trying to pull out the card that, “Oh, I’ve worked on this system for longer than you’ve worked on this system, therefore, you should defer to me.” Do you find that you face headwinds for not having the quote-unquote, “Traditional” engineering technical background?

Stephanie: I will say that I do a bit. And I did, I would say when I first joined DevRel, and I don’t know if it was much more so that it was being imposed on me or if it was being self-imposed, something that I felt like I needed to prove to gain credibility, not just in my organization, but in the industry at large. And it wasn’t until two or three years into it, that I realized that I had a niche myself. It was to create stories with my content that could communicate these concepts to developers just as effectively. And yes, I can still prove that I can go into an hour-long or a 45-minute-long tech talk or a webinar about a topic, but I can also easily create a five to ten-minute video that communicates concepts and inspires audiences just the same, and more importantly, be able to point to resources, code labs, tutorials, GitHub repos, that can allow the audience to be hands-on themselves, too. So really, I think that it was over time that I gained more experience and realized that my skill sets are valuable in a different way, and it’s okay to have a different background as long as you bring something to the table.

Corey: And I think that it’s indisputable that you do. The concept of yours that I’ve encountered from time to time has always been insightful, it is always been extremely illuminating, and—you wouldn’t think of this as worthy of occasion and comment, but I feel it needs to be said anyway—at no point in any of your content did I feel like I was being approached in a condescending way, where at every point it was always about uplifting people to a level of understanding, rather than doing the, “Well, I’m smarter than you and you couldn’t possibly understand the things that I’ve been to.” It is relatable, it is engaging, and you add a very human face to what is admittedly an area of industry that is lacking in a fair bit of human element.

Stephanie: Yeah, and I think that’s the thing that many folks DevRel continue to underline is the idea of empathy, empathizing with your audiences, empathizing with the developers, the engineers, the data engineers, whoever it is that you’re creating content for, it’s being in their shoes. But for me, I may not have been in those shoes for years, like many other folks historically have been in for DevRel, but I want to at least go through the journey of learning a new piece of technology. For example, if I’m learning a new platform on Google Cloud, going through the steps of creating a demo, or walking through a tutorial, and then candidly explaining that experience to my audience, or creating a video about it. I really just reject the idea of having ego in tech and I would love to broaden the opportunity for folks who came from a different background like myself. I really want to just represent the new world of technology where it wasn’t full of people who may have had the privilege to start coding at a very early age, in their garages.

Corey: Yeah, privilege of, in many respects, also that privilege means, “Yes, I had the privilege of not having to have friends and deal with learning to interact with other human beings, which is what empowered me to build this company and have no social skills whatsoever.” It’s not the aspirational narrative that we sometimes are asked to believe. You are similar in some respects to a number of things that I do—by which I mean, you do it professionally and well and I do it as basically performance shitpost art—but you’re on Twitter, you make videos, you do podcasts, you write long-form and short-form as well. You are sort of all across the content creation spectrum. Which of those things do you prefer to do? Which ones of those are things you find a little bit more… “Well, I have to do it, but it’s not my favorite?” Or do you just tend to view it as content is content; you just look at different media to tell your story?

Stephanie: Well, I will say any form of content is queen—I’m not going to say king, but—[laugh] content is king, content is queen, it doesn’t matter.

Corey: Content is a baroness as it turns out.

Stephanie: [laugh]. There we go. I have to say, so given my background, I mentioned I was into production and entertainment before, so I’ve always had a gravitation towards video content. I love tinkering with cameras. Actually, as I got started out at Google Cloud, I was creating scrappy content using webcams and my own audio equipment, and doing my own research, and finding lounges and game rooms to do that, and we would just upload it to our own YouTube channel, which probably wasn’t allowed at the time, but hey, we got by with it.

And eventually, I got approached by DevRel to start doing it officially on the channel and I was given budget to do it in-studio. And so that was sort of my stepping stone to doing this full-time eventually, which I never foresaw for myself. And so yeah, I have this huge interest in—I’m really engaged with video content, but once I started expanding and realizing that I could repurpose that content for podcasting, I could repurpose it for blogs, then you start to realize that you can shard content and expand your reach exponentially with this. So, that’s when I really started to become more active on social media and leverage it to build not just content for Google Cloud, but build my own brand in
tech.

Corey: That is the inescapable truth of DevRel done right is that as you continue doing it, in time, in your slice of the industry, it is extremely likely that your personal brand eclipses the brand of the company that you represent. And it’s in many ways a test of corporate character—if it makes sense—as do how they react to that. I’ve worked in roles before I started this place where I was starting to dabble with speaking a lot, and there was always a lot of insecurity that I picked up of, “Well, it feels like you’re building your personal brand, not advancing the company here, and we as a company do not see the value in you doing that.” Direct quote from the last boss I had. And, well, that partially explains why I’m here, I suppose.

But there’s insecurity there. I’d see the exact opposite coming out of Google, especially in recent times. There’s something almost seems to be a renaissance in Google Cloud, and I’m not sure where it came from. But if I look at it across the board, and you had taken all the labels off of everything, and you had given me a bunch of characteristics about different companies, I would never have guessed that you were describing Google when you’re talking about Google Cloud. And perhaps that’s unfair, but perceptions shape reality.

Stephanie: Yeah, I find that interesting because I think traditionally in DevRel, we’ve also hired folks for their domain expertise and their brand, depending on what you’re representing, whether it’s in the Kubernetes space or Python client library that you’re supporting. But it seems like, yes, in my case, I’ve organically started to build my brand while at Google, and Google has been just so spectacular in supporting that for me. But yeah, it’s a fine line that I think many people have to walk. It’s like, do you want to continue to build your own brand and have that carry forth no matter what company you stay at, or if you decide to leave? Or can you do it hand-in-hand with the company that you’re at? For me, I think I can do it hand-in-hand with Google Cloud.

Corey: It’s taken me a long time to wrap my head around what appears to be a contradiction when I look at Google Cloud, and I think I’ve mostly figured it out. In the industry, there is a perception that Google as an entity is condescending and sneering toward every other company out there because, “You’re Google, you know how to do all these great, amazing things that are global-spanning, and over here at Twitter for Pets, we suck doing these things.” So, Google is always way smarter and way better at this than we could ever hope to be. But that is completely opposed to my personal experiences talking with Google employees. Across the board, I would say that you all are self-effacing to a fault.

And I mean that in the sense of having such a limited ego, in some cases, that it’s, “Well, I don’t want to go out there and do a whole video on this. It’s not about me, it’s about the technology,” are things that I’ve had people who work at Google say to me. And I appreciate the sentiment; it’s great, but that also feels like it’s an aloofness. It also fails to humanize what it is that you’re doing. And you are a, I’ve got to say, a breath of fresh air when it comes to a lot of that because your stories are not just, “Here’s how you do a thing. It’s awesome. And this is all the intricacies of the API.”

And yeah, you get there, but you also contextualize that in a, “Here’s why it matters. Here’s the problem that solves. Here is the type of customer’s problem that this is great for,” rather than starting with YAML and working your way up. It’s going the other way, of, “We want to sell some underpants,” or whatever it is the customer is trying to do today. And that is the way that I think is one of the best ways to drive adoption of what’s going on because if you get people interested and excited about something—at least in my experience—they’re going to figure out how the API works. Badly in many cases, but works. But if you start on the API stuff, it becomes a solution looking for a problem. I like your approach to this.

Stephanie: Thank you. Yeah, I appreciate that. I think also something that I’ve continued to focus on is to tell stories across products, and it doesn’t necessarily mean within just Google Cloud’s ecosystem, but across the industry as well. I think we need to, even at Google, tell a better story across our product space and tie in what developers are currently using. And I think the other thing that I’m trying to work on, too, is contextualizing our products and our launches not just across the industry, but within our product strategy. Where does this tie in? Why does it matter? What is our forward-looking strategy from here? When we’re talking about our new data cloud products or analytics, [unintelligible 00:17:21], how does this tie into our API strategy?

Corey: And that’s the biggest challenge, I think, in the AI space. My argument has been for a while—in fact, I wrote a blog post on it earlier this year—that AI and machine learning is a marvelously executed scam because it’s being pushed by cloud providers and the things that you definitely need to do a machine learning experiment are a bunch of compute and a whole bunch of data that has to be stored on something, and wouldn’t you know it, y’all sell that by the pound. So, it feels, from a cynical perspective, which I excel at espousing, that approach becomes one of you’re effectively selling digital pickaxes into a gold rush. Because I see a lot of stories about machine learning how to do very interesting things that are either highly, highly use-case-specific, which great, that would work well, for me too, if I ever wind up with, you know, a petabyte of people’s transaction logs from purchasing coffee at my national chain across the country. Okay, that works for one company, but how many companies look like that?

And on the other side of it, “It’s oh, here’s how we can do a whole bunch of things,” and you peel back the covers a bit, and it looks like, “Oh, but you really taught me here is bias laundering?” And, okay. I think that there’s a definite lack around AI and machine learning of telling stories about how this actually matters, what sorts of things people can do with it that aren’t incredibly—how do I put this?—niche or a problem in search of a solution?

Stephanie: Yeah, I find that there are a couple approaches to creating content around AI and other technologies, too, but one of them being inspirational content, right? Do you want to create something that tells the story of how I created a model that can predict what kind of bakery item this is? And we’re going to do it by actually showcasing us creating the outcome. So, that’s one that’s more like, okay. I don’t know how relatable or how appropriate it is for an enterprise use case, but it’s inspirational for new developers or next gen developers in the AI space, and I think that can really help a company’s brand, too.

The other being highly niche for the financial services industry, detecting financial fraud, for example, and that’s more industry-focused. I found that they both do well, in different contexts. It really depends on the channel that you’re going to display it on. Do you want it to be viral? It really depends on what you’re measuring your content for. I’m curious from you, Corey, what you’ve seen across, as a consumer of content?

Corey: What’s interesting, at least in my world, is that there seems to be, given that what I’m focusing on first and foremost is the AWS ecosystem, it’s not that I know it the best—I do—but at this point, it’s basically Stockholm Syndrome where it’s… with any technology platform when you’ve worked with it long enough, you effectively have the most valuable of skill sets around it, which is not knowing how it works, but knowing how it doesn’t, knowing what the failure mode is going to look like and how you can work around that and detect it is incredibly helpful. Whereas when you’re trying something new, you have to wait until it breaks to find the sharp edges on it. So, there’s almost a lock-in through, “We failed you enough times,” story past a certain point. But paying attention to that ecosystem, I find it very disjointed. I find that there are still events that happen and I only find out when the event is starting because someone tweets about it, and for someone who follows 40 different official AWS RSS feeds, to be surprised by something like that tells me, okay, there’s not a whole lot of cohesive content strategy here, that is at least making it easy for folks to consume the things that they want, especially in my case where even the very niche nature of what I do, my interest is everything.

I have a whole bunch of different filters that look for various keywords and the rest, and of course, I have helpful folks who email me things constantly—please keep it up; I’m a big fan—worst case, I’d rather read something twice than nothing. So, it’s helpful to see all of that and understand the different marketing channels, different personas, and the way that content approaches, but I still find things that slip through the cracks every time. The thing that I’ve learned—and it felt really weird when I started doing it—was, I will tell the same stories repeatedly in different forums, or even the same forum. I could basically read you a Twitter thread from a year ago, word-for-word, and it would blow up
bigger than it did the first time. Just because no one reads everything.

Stephanie: Exactly.

Corey: And I’ve already told my origin story. You’re always new to someone. I’ve given talks internally at Amazon at various times, and I’m sort of loud and obnoxious, but the first question I love to ask is, “Raise your hand if you’ve never heard of me until today.” And invariably, over three-quarters of the room raises their hand every single time, which okay, great. I think that’s awesome, but it teaches me that I cannot ever expect someone to have, quote-unquote, “Done the reading.”

Stephanie: I think the same can be said about the content that I create for the company. You can’t assume that people, A) have seen my tweets already or, B) understand this product, even if I’ve talked about it five times in the past. But yes, I agree. I think that you definitely need to have a content strategy and how you format your content to be more problem-solution-oriented.

And so the way that I create content is that I let them fall into three general buckets. One being that it could be termed definition: talking about the basics, laying the foundation of a product, defining terms around a topic. Like, what is App Engine, or Kubeflow 101, or talking about Pub/Sub 101.

The second being best practices. So, outlining and explaining the best practices around a topic, how do you design your infrastructure for scale and reliability.

And the third being diagnosis: investigating; exploring potential issues, as you said; using scripts; Stackdriver logging, et cetera. And so I just kind of start from there as a starting point. And then I generally follow a very, very effective model. I’m sure you’re aware of it, but it’s called the five point argument model, where you are essentially telling a story to create a compelling narrative for your audience, regardless of the topic or what bucket that topic falls into.

So, you’re introducing the problem, you’re sort of rising into a point where the climax is the solution. And that’s all to build trust with your audience. And as it falls back down, you’re giving the results in the conclusion, and that’s to inspire action from your audience. So, regardless of what you end up talking about this problem-solution model—I’ve found at least—has been highly effective. And then in terms of sharing it out, over and over again, over the span of two months, that’s how you get the views that you want.

Corey: This episode is sponsored in part by something new. Cloud Academy is a training platform built on two primary goals. Having the highest quality content in tech and cloud skills, and building a good community the is rich and full of IT and engineering professionals. You wouldn’t think those things go together, but sometimes they do. Its both useful for individuals and large enterprises, but here's what makes it new. I don’t use that term lightly. Cloud Academy invites you to showcase just how good your AWS skills are. For the next four weeks you’ll have a chance to prove yourself. Compete in four unique lab challenges, where they’ll be awarding more than $2000 in cash and prizes. I’m not kidding, first place is a thousand bucks. Pre-register for the first challenge now, one that I picked out myself on Amazon SNS image resizing, by visiting cloudacademy.com/corey. C-O-R-E-Y. That’s cloudacademy.com/corey. We’re gonna have some fun with this one!

Corey: See, that’s a key difference right there. I don’t do anything regular in terms of video as part of my content. And I do it from time to time, but you know, getting gussied up and whatnot is easier than just talking into a microphone. As I record this, it’s Friday, I’m wearing a Hawaiian shirt, and I look exactly like the middle-aged dad that I am. And for me at least, a big breakthrough moment was realizing that my audience and I are not always the same.

Weird confession for someone in my position: I don’t generally listen to podcasts. And the reason behind that is I read very quickly, and even if I speed up a podcast, I’m not going to be able to consume the information nearly as quickly as I could by reading it. That, amongst other reasons, is one of the reasons that every episode of this show has a full transcript attached to it. But I’m not my audience. Other people prefer to learn by listening and there’s certainly nothing wrong with that.

My other podcast, the AWS Morning Brief, is the spoken word version of the stuff that I put out in my newsletter every week. And that is—it’s just a different area for people to consume the content because that’s what works for them. I’m not one to judge. The hard part for me was getting over that hump of assuming the audience was like me.

Stephanie: Yeah. And I think the other key part of is just mainly consistency. It’s putting out the content consistently in different formats because everybody—like you said—has a different learning style. I myself do. I enjoy visual styles.

I also enjoy listening to podcasts at 2x speed. [laugh]. So, that’s my style. But yeah, consistency is one of the key things in building content, and building an audience, and making sure that you are valuable to your audience. I mean, social media, at the end of the day is about the people that follow you.

It’s not about yourself. It should never be about yourself. It’s about the value that you provide. Especially as somebody who’s in DevRel in this position for a larger company, it’s really about providing value.

Corey: What are the breakthrough moments that I had relatively early in my speaking career—and I think it’s clear just from what you’ve
already said that you’ve had a similar revelation at times—I gave a talk, that was really one of my first talks that went semi-big called, “Terrible Ideas in Git.” It was basically, learn how to use Git via anti-pattern. What it secretly was, was under the hood, I felt it was time I learned Git a bit better than I did, so I pitched it and I got a talk accepted. So well, that’s what we call a forcing function. By the time I give that talk, I’d better be [laugh] able to have built a talk that do this intelligently, and we’re going to hope for the best.

It worked, but the first version of that talk I gave was super deep into the plumbing of Git. And I’m sure that if any of the Git maintainers were in the audience, they would have found it great, but there aren’t that many folks out there. I redid the talk and instead approached it from a position of, “You have no idea what Git is. Maybe you’ve heard of it, but that’s as far as it goes.” And then it gets a little deeper there.

And I found that making the subject more accessible as opposed to deeper into the weeds of it is almost always the right decision from a content perspective. Because at some level, when you are deep enough into the weeds, the only way you’re going to wind up fixing something or having a problem that you run into get resolved, isn’t by listening to a podcast or a conference talk; it’s by talking to the people who built the thing because at that level, those are the only people who can hang at that level of depth. That stops being fodder for conference talks unless you turn it into an after-action report of here’s this really weird thing I learned.

Stephanie: Yeah. And you know, to be honest, the one of the most successful pieces of content I’ve created was about data center security. I visited a data center and I essentially unveiled what our security protocols were. And that wasn’t a deeply technical video, but it was fun and engaging and easily understood by the masses. And that’s what actually ended up resulting in the highest number of views.

On top of that, I’m now creating a video about our subsea fiber optic cables. Finding that having to interview experts from a number of different teams across engineering and our strategic negotiators, it was like a monolith of information that I had to take in. And trying to format that into a five-minute story, I realized that bringing it up a layer of abstraction to help folks understand this at a wider level was actually beneficial. And I think it’ll turn into a great piece of content. I’m still working on it now. So, [laugh] we’ll see how it turns out.

Corey: I’m a big fan of watching people learn and helping them get started. The thing that I think gets lost a lot is it’s easy to assume that if I look back in time at myself when I was first starting my professional career two decades ago, that I was exactly like I am now, only slightly more athletic and can walk up a staircase without getting winded. That’s never true. It never has been true. I’ve learned a lot about not just technology but people as I go, and looking at folks are entering the workforce today through the same lens of, “Well, that’s not how I would handle that situation.” Yeah, no kidding. I have two decades of battering my head against the sharp edges and leaving dents in things to inform that opinion.

No, when I was that age, I would have handled it way worse than whatever it is I’m critiquing at the time. But it’s important to me that we wind up building those pathways and building those bridges so that people coming into the space, first, have a clear path to get here, and secondly, have a better time than I ever did. Where does the next generation of talent come from has been a recurring question and a recurring theme on the show.

Stephanie: Yeah. And that’s exactly why I’ve been such a fierce supporter of women in tech, and also, again, encouraging a broader community to become a part of technology. Because, as I said, I think we’re in the midst of a new era of technology, of people from all these different backgrounds in places that historically have had more remote access to technology, now having the ability to become developers at an early age. So, with my content, that’s what I’m hoping to drive to make this information more easily accessible. Even if you don’t want to become a Google Cloud engineer, that’s totally fine, but if I can help you understand some of the foundational concepts of cloud, then I’ve done my job well.

And then, even with women who are already trying to break into technology or wanting to become a part of it, then I want to be a mentor for them, with my experience not having a technical background and saying yes to opportunities that challenged me and continuing to build my own luck between hard work and new opportunities.

Corey: I can’t wait to see how this winds up manifesting as we see understandings of what we’re offering to customers in different areas in different ways—both in terms of content and terms of technology—how that starts to evolve and shift. I feel like we’re at a bit of an inflection point now, where today if I graduate from school and I want to start a business, I have to either find a technical co-founder or I have to go to a boot camp and learn how to code in order to build something. I think that if we can remove that from the equation and move up the stack, sure, you’re not going to be able to build the next Google or Pinterest or whatnot from effectively Visual Basic for Interfaces, but you can build an MVP and you can then continue to iterate forward and turn it into something larger down the road. The other part of it, too, is that moving up the stack into more polished solutions rather than here’s a bunch of building blocks for platforms, “So, if you want a service to tell you whether there’s a picture of a hot dog or not, here’s a service that does exactly that.” As opposed to, “Oh, here are the 15 different services, you can bolt together and pay for each one of them and tie it together to something that might possibly work, and if it breaks, you have no
idea where to start looking, but here you go.” A packaged solution that solves business problems.

Things move up the stack; they do constantly. The fact is that I started my career working in data centers and now I don’t go to them at all because—spoiler—Google, and Amazon, and people who are not IBM Cloud can absolutely run those things better than I can. And there’s no differentiated value for me in solving those global problems locally. I’d rather let the experts handle stuff like that while I focus on interesting problems that actually affect my business outcome. There’s a reason that instead of running all the nonsense for lastweekinaws.com myself because I’ve worked in large-scale WordPress hosting companies, instead I pay WP Engine to handle it for me, and they, in turn, hosted on top of Google Cloud, but it doesn’t matter to me because it’s all just a managed service that I pay for. Because me running the website itself adds no value, compared to the shitpost I put on the website, which is where the value derives from. For certain odd values of value.

Stephanie: [laugh]. Well, two things there is that I think we actually had a demo created on Google Cloud that did detect hot dogs or not hot dogs using our Vision API, years in the past. So, thanks for reminding me of that one.

Corey: Of course.

Stephanie: But yeah, I mean, I completely agree with that. I mean, this is constantly a topic in conversation with my team members, and with clients. It’s about higher level of abstractions. I just did a video series with our fellow, Eric Brewer, who helped build cloud infrastructure here at Google over the past ten decades. And I asked him what he thought the future of cloud would be in the next ten years, and he mentioned, “It’s going to be these higher levels of abstraction, building platforms on top of platforms like Kubernetes, and having more services like Cloud run serverless technologies, et cetera.”

But at the same time, I think the value of cloud will continue to be providing optionality for developers to have more opinionated services, services like GKE Autopilot, et cetera, that essentially take away the management of infrastructure or nodes that people don’t really want to deal with at the end of the day because it’s not going to be a competitive differentiator for developers. They want to focus on building software and focusing on keeping their services up and running. And so yeah, I think the future is going to be that, giving developers flexibility and freedom, and still delivering the best-of-breed technology. If it’s covering something like security, that’s something that should be baked in as much as possible.

Corey: You’re absolutely right, first off. I’m also looking beyond it where I want to be able to build a website that is effectively Twitter, only for pets—because that is just a harebrained enough idea to probably raise a $20 million seed round these days—and I just want to be able to have the barks—those are like tweets, only surprisingly less offensive and racist—and have them just be stored somewhere, ideally presumably under the hood somewhere, it’s going to be on computers, but whether it’s in containers, or whether it’s serverless, or however is working is the sort of thing that, “Wow, that seems like an awful lot of nonsense that is not central nor core to my business succeeding or failing.” I would say failing, obviously, except you can lose money at scale with the magic of things like SoftBank. Here we are.

And as that continues to grow and scale, sure, at some point I’m going to have bespoke enough needs and a large enough scale where I do have to think about those things, but building the MVP just so I can swindle some VCs is not the sort of thing where I should have to go to that depth. There really should be a golden-path guardrail-style thing that I can effectively drag and drop my way into the next big scam. And that is, I think, the missing piece. And I think that we’re not quite ready technologically to get there yet, but I can’t shake the feeling and the hope that’s where technology is going.

Stephanie: Yeah. I think it’s where technology is heading, but I think part of the equation is the adoption by our industry, right? Industry adoption of cloud services and whether they’re ready to adopt services that are that drag-and-drop, as you say. One thing that I’ve also been talking a lot about is this idea of service-oriented networking where if you have a service or API-driven environment and you simply want to bring it to cloud—almost a plug-and-play there—you don’t really want to deal with a lot of the networking infrastructure, and it’d be great to do something like PrivateLink on AWS, or Private Service Connect on Google Cloud.

While those conversations are happening with customers, I’m finding that it’s like trying to cross the Grand Canyon. Many enterprise customers are like, “That sounds great, but we have a really complex network topology that we’ve been sitting on for the past 25 years. Do you really expect that we’re going to transition over to something like that?” So, I think it’s about providing stepping stones for our customers until they can be ready to adopt a new model.

Corey: Yeah. And of course, the part that never gets said out loud but is nonetheless true and at least as big of a deal, “And we have a whole team of people who’ve built their entire identity around that network because that is what they work on, and they have been ignoring cloud forever, and if we just uplift everything into a cloud where you folks handle that, sure, it’s better for the business outcome, but where does that leave them?” So, they’ve been here for 25 years, and they will spend every scrap of political capital they’ve managed to accumulate to torpedo a cloud migration. So, any FUD they can find, any horse-trading they can do, anything they can do to obstruct the success of a cloud initiative, they’re going to do because people are people, and there is no real plan to mitigate that. There’s also the fact that unless there’s a clear business value story about a feature velocity increase or opening up new markets, there’s also not an incentive to do things to save money. That is never going to be the number one priority in almost any case short of financial disaster at a company because everything they’re doing is building out increasing revenue, rather than optimizing what they’re already doing.

So, there’s a whole bunch of political challenges. Honestly, moving the computer stuff from on-premises data centers into a cloud provider is the easiest part of a cloud migration compared to all of the people that are involved.

Stephanie: Yeah. Yeah, we talked about serverless and all the nice benefits of it, but unless you are more a digitally-born, next-gen developer, it may be a higher burden for you to undertake that migration. That’s why we always [laugh] are talking about encouraging people to start with newer surfaces.

Corey: Oh, yeah. And that’s the trick, too, is if you’re trying to learn a new cloud platform these days—first, if you’re trying to pick one, I’d be hard-pressed to suggest anything other than Google Cloud, with the possible exception of DigitalOcean, just because the new user experience is so spectacularly good. That was my first real, I guess, part of paying attention to Google Cloud a few years ago, where I was, “All right, I’m going to kick the tires on this and see how terrible this interface is because it’s a Google product.” And it was breathtakingly good, which I did not expect. And getting out of the way to empower someone who’s new to the platform to do something relatively quickly and straightforwardly is huge. And sure, there’s always room to prove, but that is the right area to focus on. It’s clear that the right energy was spent in the right places.

Stephanie: Yeah. I will say a story that we don’t tell quite as well as we should is the One Google story. And I’m not talking about just between Workspace and Google Cloud, but our identity access management and knowing your Google account, which everybody knows. It’s not like Microsoft, where you’re forced to make an account, or it’s not like AWS where you had a billion accounts and you hate them all.

Corey: Oh, my God, I dread logging into the AWS console every time because it is such a pain in the ass. I go to cloud.google.com sometimes to check something, it’s like, “Oh, right. I have to dig out my credentials.” And, “Where’s my YubiKey?” And get it. Like, “Oh. I’m already log—oh. Oh, right. That’s right. Google knows how identity works, and they don’t actively hate their customers. Okay.” And it’s always a breath of fresh air. Though I will say that by far and away, the worst login experience I’ve seen yet is, of course, Azure.

Stephanie: [laugh]. That’s exactly right. It’s Google account. It’s yours. It’s personal. It’s like an Apple iCloud account. It’s one click, you’re in, and you have access to all the applications. You know, so it’s the same underlying identity structure with Workspace and Gmail, and it’s the same org structure, too, across Workspace and Google Cloud. So, it’s not just this disingenuous financial bundle between GCP and Workspace; it’s really strategic. And it’s kind of like the idea of low code or no code. And it looks like that’s what the future of cloud will be. It’s not just by VMs from us.

Corey: Yeah. And there are customers who want to buy VMs and that’s great. Speed up what they’re doing; don’t get in the way of people giving you their money, but if you’re starting something net-new, there’s probably better ways to do it. So, I want to thank you for taking as much time as you have to wind up going through how you think about, well, the art of storytelling in the world of engineering. If people want to learn more about who you are, what you’re up to, and how you approach things, where can they find you?

Stephanie: Yeah, so you can head to stephrwong.com where you can see my work and also get in touch with me if you want to collaborate on any content. I’m always, always, always open to that. And my Twitter is @stephr_wong.

Corey: And we will, of course, put links to that in the [show notes 00:40:03]. Thank you so much for taking the time to speak with me.

Stephanie: Thanks so much.

Corey: Stephanie Wong, head of developer engagement at Google Cloud. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry comment telling me that the only way to get into tech these days is, in fact, to graduate with a degree from Stanford, and I can take it from you because you work in their admissions office.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and
we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Brian

I lead the Google Cloud Product and Industry Marketing team. We’re focused on accelerating the growth of Google Cloud by establishing thought leadership, increasing demand and usage, enabling our sales teams and partners to tell our product stories with excellence, and helping our customers be the best advocates for us.

Before joining Google, I spent over 25 years in product marketing or engineering in different forms. I started my career at Microsoft and had a very non-traditional path for 20 years. I worked in every product division except for cloud. I did marketing, product management, and engineering roles. And, early on, I was the first speech writer for Steve Ballmer and worked on Bill Gates’ speeches too. My last role was building up the Microsoft Surface business from scratch and as VP of the hardware businesses. After Microsoft, I spent a year as CEO at a hardware startup called Doppler Labs, where we made a run at transforming hearing, and then two years as VP at Amazon Web Services leading product marketing, developer advocacy, and a bunch more marketing teams.

I have three kids still at home, Barty, Noli, and Alder, who are all named after trees in different ways. My wife Edie and I met right at the beginning of our first year at Yale University, where I studied math, econ, and philosophy and was the captain of the Swim and Dive team my senior year. Edie has a PhD in forestry and runs a sustainability and forestry consulting firm she started, that is aptly named “Three Trees Consulting”. We love the outdoors, tennis, running, and adventures in my 1986 Volkswagen Van, which is my first and only car, that I can’t bring myself to get rid of.

Links:

  • Twitter: https://twitter.com/IsForAt
  • LinkedIn: https://www.linkedin.com/in/brhall/
  • Episode 10: https://www.lastweekinaws.com/podcast/screaming-in-the-cloud/episode-10-education-is-not-ready-for-teacherless/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Redis, the company behind the incredibly popular open source database that is not the bind DNS server. If you’re tired of managing open source Redis on your own, or you’re using one of the vanilla cloud caching services, these folks have you covered with the go to manage Redis service for global caching and primary database capabilities; Redis Enterprise. Set up a meeting with a Redis expert during re:Invent, and you’ll not only learn how you can become a Redis hero, but also have a chance to win some fun and exciting prizes. To learn more and deploy not only a cache but a single operational data platform for one Redis experience, visit redis.com/hero. Thats r-e-d-i-s.com/hero. And my thanks to my friends at Redis for sponsoring my ridiculous non-sense.

Corey: Writing ad copy to fit into a 30 second slot is hard, but if anyone can do it the folks at Quali can. Just like their Torque infrastructure automation platform can deliver complex application environments anytime, anywhere, in just seconds instead of hours, days or weeks. Visit Qtorque.io today and learn how you can spin up application environments in about the same amount of time it took you to listen to this ad.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’m joined today by a special guest that I’ve been, honestly, antagonizing for years now. Once upon a time, he spent 20 years at Microsoft, then he wound up leaving—as occasionally people do, I’m told—and going to AWS, where according to an incredibly ill-considered affidavit filed in a court case, he mostly focused on working on PowerPoint slides. AWS is famously not a PowerPoint company, and apparently, you can’t change culture. Now, he’s the VP of Product and Industry Marketing at Google Cloud. Brian Hall, thank you for joining me.

Brian: Hi, Corey. It’s good to be here.

Corey: I hope you’re thinking that after we’re done with our conversation. Now, unlike most conversations that I tend to have with folks who are, honestly, VP level at large cloud companies that I enjoy needling, we’re not going to talk about that today because instead, I’d rather focus on a minor disagreement we got into on Twitter—and I mean that in the truest sense of disagreement, as opposed to the loud, angry, mutual blocking, threatening to bomb people’s houses, et cetera, nonsense that appears to be what substitutes for modern discourse—about, oh, a month or so ago from the time we’re recording this. Specifically, we talked about, I’m in favor of job-hopping to advance people’s career, and you, as we just mentioned, spent 20 years at Microsoft and take something of the opposite position. Let’s talk about that. Where do you stand on the idea?

Brian: I stand in the position that people should optimize for where they are going to grow the most. And frankly, the disagreement was less about job-hopping because I’m going to explain how I job-hopped at Microsoft effectively.

Corey: Excellent. That is the reason I’m asking you rather than poorly stating your position and stuffing you like some sort of Christmas turkey straw-man thing.

Brian: And I would argue that for many people, changing jobs is the best thing that you can do, and I’m often an advocate for changing jobs even before sometimes people think they should do it. What I mostly disagreed with you on is simply following the money on your next job. What you said is if a—and I’m going to get it somewhat wrong—but if a company is willing to pay you $40,000 more, or some percentage more, you should take that job now.

Corey: Gotcha.

Brian: And I don’t think that’s always the case, and that’s what we’re talking about.

Corey: This is the inherent problem with Twitter is that first, I tend to write my Twitter threads extemporaneously without a whole lot of thought being put into things—kind of like I live my entire life, but that’s neither here nor there—

Brian: I was going to say, that comes across quite clearly.

Corey: Excellent. And 280 characters lacks nuance. And I definitely want to have this discussion; this is not just a story where you and I beat heads and not come to an agreement on this. I think it’s that we fundamentally do agree on the vast majority of this, I just want to make sure that we have this conversation in a way, in a forum that doesn’t lend itself to basically empowering the worst aspects of my own nature. Read as, not Twitter.

Brian: Great. Let’s do that.

Corey: So, my position is, and I was contextualizing this from someone who had reached out who was early in their career, they had spent a couple of years at AWS and they were entertaining an offer elsewhere for significantly more money. And this person, I believe I can—I believe it’s okay for me to say this: she—was very concerned that, “I don’t want to look like I’m job-hopping, and I don’t dislike my team. My manager is great. I feel disloyal for leaving. What should I do?”

Which first, I just want to say how touched I am that someone who is early in their career and not from a wildly overrepresented demographic like you and I felt a sense of safety and security in reaching out to ask me that question. I really wish more people would take that kind of initiative. It’s hard to inspire, but here we are. And my take to her was, “Oh, my God. Take the money.” That was where this thread started because when I have conversations with people about those things, it becomes top of mind, and I think, “Hmm, maybe there’s a one-to-many story that becomes something that is actionable and useful.”

Brian: Okay, so I’m going to give two takes on this. I’ll start with my career because I was in a similar position as she was, at one point in my career. My background, I lucked into a job at Microsoft as an intern in 1995, and then did another internship in ’96 and then started full time on the Internet Explorer team. And about a year-and-a-half into that job, I—we had merged with the Windows ’98 team and I got the opportunity to work on Bill Gates’s speech for the Windows ’98 launch event. And I—after that was right when Steve Ballmer became president of Microsoft and he started doing a lot more speeches and asked to have someone to help him with speeches.

And Chris Capossela, who’s now the CMO at Microsoft, said, “Hey, Brian. You interested in doing this for Steve?” And my first reaction was, well, even inside Microsoft, if I move, it will be disloyal. Because my manager’s manager, they’ve given me great opportunities, they’re continuing to challenge me, I’m learning a bunch, and they advised not doing it.

Corey: It seems to me like you were in a—how to put this?—not to besmirch the career you have wrought with the sweat of your brow and the toil of your back, but in many ways, you were—in a lot of ways—you were in the right place at the right time, riding a rocket ship, and built opportunities internally and talked to folks there, and built the relationships that enabled you to thrive inside of a company’s ecosystem. Is that directionally correct?

Brian: For sure. Yet, there’s also, big companies are teams of teams, and loyalty is more often with the team and the people that you work with than the 401k plan. And in this case, you know, I was getting this pressure that says, “Hey, Brian. You’re going to get all these opportunities. You’re doing great doing what you’re doing.”

And I eventually had the luck to ask the question, “Hey, if I go there and do this role”—and by the way, nobody had done it before, and so part of their argument was, “You’re young, Steve’s… Steve. Like, you could be a fantastic ball of flames.” And I said, “Okay, if [laugh] let’s say that happens. Can I come back? Can I come back to the job I was doing before?”

And they were like, “Yeah, of course. You’re good at what you do.” To me, which was, “Okay, great. Then I’m gone. I might as well go try this.” And of course, when I started at Microsoft, I was 20, 21, and I thought I’d be there for two or three years and then I’d end up going back to school or somewhere else. But inside Microsoft, what kept happening as I just kept getting new opportunities to do something else that I’d learned a bunch from, and I ultimately kind of created this mentality for how I thought about next job of, “Am I going to get more opportunities if I am able to be successful in this new job?” Really focused on optionality and the ability to do work that I want to do and have more choices to do that.

Corey: You are also on a I almost want to call it a meteoric trajectory. In some ways. You effectively went from—what was your first role there? It was—

Brian: The lowest level of college hire you can do at Microsoft, effectively.

Corey: Yeah. All the way on up to at the end of it the Corporate VP for Microsoft Devices. It seems to me that despite the fact that you spent 20 years there, you wound up having a bunch of different jobs and an entire career trajectory internal to the organization, which is, let’s be clear, markedly different from some of the folks I’ve interviewed at various times, in my career as an employer and as a technical interviewer at a consulting company, where they’d been somewhere for 15 years, and they had one year of experience that they repeated 15 times. And it was one of the more difficult things that I encountered is that some folks did not take ownership of their career and focus on driving it forward.

Brian: Yeah, that, I had the opposite experience, and that is what kept me there that long. After I would finish a job, I would say, “Okay, what do I want to learn how to do next, and what is a challenge that would be most interesting?” And initially, I had to get really lucky, honestly, to be able to get these. And I did the work, but I had to have the opportunity, and that took luck. But after I had a track record of saying, “Hey, I can jump from being a product marketer to being a speechwriter; I can do speechwriting and then go do product management; I can move from product management into engineering management.”

I can do that between different businesses and product types, you build the ability to say, “Hey, I can learn that if you give me the chance.” And it, frankly, was the unique combination of experiences I had by having tried to do these other things that gave me the opportunity to have a fast trajectory within the company.

Corey: I think it’s also probably fair to say that Microsoft was a company that, in its dealings with you, is operating in good faith. And that is a great thing to find when you see it, but I’m cynical; I admit that. I see a lot of stories where people give and sacrifice for the good of the company, but that sacrifice is never reciprocated. And we’ve all heard the story of folks who will put their nose to the grindstone to ship something on time, only to be rewarded with a layoff at the end, and stories like that resonate.

And my argument has always been that you can’t love a company because the company can’t love you back. And when you’re looking at do I make a career move or do I stay, my argument is that is the best time to be self-interested.

Brian: Yeah, I don’t think—companies are there for the company, and certainly having a culture that supports people that wants to create opportunity, having a manager that is there truly to make you better and to give you opportunity, that all can happen, but it’s within a company and you have to do the work in order to try and get into that environment. Like, I worked hard to have managers who would support my growth, would give me the bandwidth and leash early on to not be perfect at what I’m doing, and that always helped me. But you get to go pick them in a company like that, or in the industry in general, you get—just like when a manager is hiring you, you also get to understand, hey, is this a person I want to work for?

But I want to come back to the main point that I wanted to make. When I changed jobs, I did it because I wanted to learn something new and I thought that would have value for me in the medium-term and long-term, versus how do I go max cash in what I’m already good at?

Corey: Yes.

Brian: And that’s the root of what we were disagreeing with on Twitter. I have seen many people who are good at something, and then another company says, “Hey, I want you to do that same thing in a worse environment, and we’ll pay you more.”

Corey: Excellence is always situational. Someone who is showered in accolades at one company gets fired at a different company. And it’s not because they suddenly started sucking; it’s because the tools and resources that they needed to succeed were present in one environment and not the other. And that varies from person to person; when someone doesn’t work out of the company, I don’t have a default assumption that there’s something inherently wrong with them.

Of course, I look at my own career and the sheer, staggeringly high number of times I got fired, and I’m starting to think, “Huh. The only consistent factor in all of these things is me. Nah, couldn’t be my problem. I just worked for terrible places, for terrible people. That’s got to be the way it works.” My own peace of mind. I get it. That is how it feels sometimes and it’s easy to dismiss that in different ways. I don’t want to let my own bias color this too heavily.

Brian: So, here are the mistakes that I’ve seen made: “I’m really good at something; this other company will pay me to do just that.” You move to do it, you get paid more, but you have less impact, you don’t work with as strong of people, and you don’t have a next step to learn more. Was that a good decision? Maybe. If you need the money now, yes, but you’re a little bit trading short-term money for medium-and long-term money where you’re paid for what you know; that’s the best thing in this industry. We’re paid for what we know, which means as you’re doing a job, you can build the ability to get paid more by knowing more, by learning more, by doing things that stretch you in ways that you don’t already know.

Corey: In 2006, I bluffed my way through a technical interview and got a job as a Unix systems administrator for a university that was paying $65,000 a year, and I had no idea what I was going to do with all of that money. It was more money than I could imagine at that point. My previous high watermark, working for an ethically challenged company in a sales role at a target comp of 55, and I was nowhere near it. So okay, let’s go somewhere else and see what happens. And after I’d been there a month or two, my boss sits me down and said, “So”—it’s our annual compensation adjustment time—“Congratulations. You now make $68,000.”

And it’s just, “Oh, my God. This is great. Why would I ever leave?” So, I stayed there a year and I was relatively happy, insofar as I’m ever happy in a job. And then a corporate company came calling and said, “Hey, would you consider working here?”

“Well, I’m happy here and I’m reasonably well compensated. Why on earth would I do that?” And the answer was, “Well, we’ll pay you $90,000 if you do.” It’s like, “All right. I guess I’m going to go and see what the world holds.”

And six weeks later, they let me go. And then I got another job that also paid $90,000 and I stayed there for two years. And I started the process of seeing what my engagement with the work world look like. And it was a story of getting let go periodically, of continuing to claw my way up and, credit where due, in my 20s I was in crippling credit card debt because I made a bunch of poor decisions, so I biased early on for more money at almost any cost. At some point that has to stop because there’s always a bigger paycheck somewhere if you’re willing to go and do something else.

And I’m not begrudging anyone who pursues that, but at some point, it ceases to make a difference. Getting a raise from $68,000 to $90,000 was life-changing for me. Now, getting a $30,000 raise? Sure, it’d be nice; I’m not turning my nose up at it, don’t get me wrong, but it’s also not something that moves the needle on my lifestyle.

Brian: Yeah. And there are a lot of those dimensions. There’s the lifestyle dimension, there’s the learning dimension, there’s the guaranteed pay dimension, there’s the potential paid dimension, there is the who I get to work with, just pure enjoyment dimension, and they all matter. And people should recognize that job moves should consider all of these.

And you don’t have to have the same framework over time as well. I’ve had times where I really just wanted to bear down and figure something out. And I did one job at Microsoft for basically six years. It changed in terms of scope of things that I was marketing, and which division I was in, and then which division I was in, and then which division I was in—because Microsoft loves a good reorg—but I basically did the same job for six years at one point, and it was very conscious. I was trying to get really good at how do I manage a team system at scale. And I didn’t want to leave that until I had figured that out. I look back and I think that’s one of the best career decisions I ever made, but it was for reasons that would have been really hard to explain to a lot of people.

Corey: Let’s also be very clear here that you and I are well-off white dudes in tech. Our failure mode is pretty much a board seat and a book deal. In fact, if—

Brian: [laugh].

Corey: —I’m not mistaken, you are on the board of something relatively recently. What was that?

Brian: United Way of King County. It’s a wonderful nonprofit in the Seattle area.

Corey: Excellent. And I look forward to reading your book, whenever that winds up dropping. I’m sure it’ll be only the very spiciest of takes. For folks who are earlier in their career and who also don’t have the winds of privilege at their backs the way that you and I do, this also presents radically differently. And I’ve spoken to a number of folks who are not wildly over-represented about this topic, in the wake of that Twitter explosion.

And what I heard was interesting in that having a manager who has your back counts for an awful lot and is something that is going to absolutely hold you to a particular company, even when it might make sense on paper for you to leave. And I think that there’s something strong there. My counterargument is okay, so you turn down the offer, a month goes past and your manager gives notice because they’re going to go somewhere else. What then? It’s one of those things where you owe your employer a duty of confidentiality, you owe them a responsibility to do your best work, to conduct yourself in an ethical manner, but I don’t believe you owe them loyalty in the sense of advancing their interests ahead of what’s best for you and your career arc.

And what’s right for any given person is, of course, a nuanced and challenging thing. For some folks, yeah, going out somewhere else for more money doesn’t really change anything and is not what they should optimize for. For other folks, it’s everything. And I don’t think either of those takes is necessarily wrong. I think it comes down to it depends on who you are, and what your situation is, and what’s right for you.

Brian: Yeah. I totally agree. For early in career, in particular, I have been a part of—I grew up in the early versions of the campus hiring program at Microsoft, and then hired 500-plus, probably, people into my teams who were from that.

Corey: You also do the same thing at AWS if I’m not mistaken. You launched their first college hiring program that I recall seeing, or at least that’s what scuttlebutt has it.

Brian: Yes. You’re well-connected, Corey. We started something called the Product Marketing Leadership Development Program when I was in AWS marketing. And then one year, we hired 20 people out of college into my organization. And it was not easy to do because it meant using, quote-unquote, “Tenured headcount” in order to do it. There wasn’t some special dispensation because they were less paid or anything, and in a world where headcount is a unit of work, effectively.

And then I’m at Google now, in the Google Cloud division, and we have a wonderful program that I think is really well done, called the Associate Product Marketing Manager Program, APMM. And what I’d say is for the people early in career, if you get the opportunity to have a manager who’s super supportive, in a system that is built to try and grow you, it’s a wonderful opportunity. And by ‘system built to grow you,’ it really is, do you have the support to get taught what you need to get taught on the job? Are you getting new opportunities to learn new things and do new things at a rapid clip? Are you shipping things into the market such that you can see the response and learn from that response, versus just getting people’s internal opinions, and then are people stretching roles in order to make them amenable for someone early in career?

And if you’re in a system that gives you that opportunity—like let’s take your example earlier. A person who has a manager who’s greatly supportive of them and they feel like they’re learning a lot, that manager leaves, if that system is right, there’s another manager, or there’s an opportunity to put your hand up and say, “Hey, I think I need a new place,” and that will be supported.

Corey: This episode is sponsored by our friends at Oracle Cloud. Counting the pennies, but still dreaming of deploying apps instead of "Hello, World" demos? Allow me to introduce you to Oracle's Always Free tier. It provides over 20 free services and infrastructure, networking, databases, observability, management, and security. And—let me be clear here—it's actually free. There's no surprise billing until you intentionally and proactively upgrade your account. This means you can provision a virtual machine instance or spin up an autonomous database that manages itself all while gaining the networking load, balancing and storage resources that somehow never quite make it into most free tiers needed to support the application that you want to build. With Always Free, you can do things like run small scale applications or do proof-of-concept testing without spending a dime. You know that I always like to put asterisks next to the word free. This is actually free, no asterisk. Start now. Visit snark.cloud/oci-free that's snark.cloud/oci-free.

Corey: I have a history of mostly working in small companies, to the point where I consider a big company to be one that has more than 200 employees, so, the idea of radically transitioning and changing teams has never really been much on the table as I look at my career trajectory and my career arc. I have seen that I’ve gotten significant 30% raises by changing jobs. I am hard-pressed to identify almost anyone who has gotten that kind of raise in a single year by remaining at a company.

Brian: One hundred percent. Like, I know of people who have, but it—

Corey: It happens, but it’s—

Brian: —is very rare.

Corey: —it’s very rare.

Brian: It’s, it’s, it’s almost the, the, um, the example that proves the point. I getting that totally wrong. But yes, it’s very rare, but it does happen. And I think if you get that far out of whack, yes. You should… you should go reset, especially if the other attributes are fine and you don’t feel like you’re just going to get mercenary pay.

What I always try and advise people is, in the bigger companies, you want to be a good deal. You don’t want to be a great deal or a bad deal. Where a great deal is you’re getting significantly underpaid, a bad deal is, “Uh oh. We hired this person to [laugh] senior,” or, “We promoted them too early,” because then the system is not there to help you, honestly, in the grand scheme of things. A good deal means, “Hey, I feel like I’m getting better work from this person for what we are giving them than what the next clear alternative would be. Let’s support them and help them grow.” Because at some level, part of your compensation is getting your company to create opportunities for you to grow. And part of the reason people go to a manager is they know they’ll give them that compensation.

Corey: I am learning this the interesting way, as we wind up hiring and building out our, currently, nine-person company. It’s challenging for us to build those opportunities while bootstrapped, but it is incumbent upon us, you’re right. That is a role of management is how do you identify growth opportunities for people, ideally, while remaining at the company, but sometimes that means that helping them land somewhere else is the right path for their next growth step.

Brian: Well, that brings up a word for managers. What you pay your employees—and I’m talking big company here, not people like yourself, Corey, where you have to decide whether you reinvesting money or putting in an individual.

Corey: Oh, yes—

Brian: But at big companies—

Corey: —a lot of things that apply when you own a company are radically departed from—

Brian: Totally.

Corey: —what is—

Brian: Totally.

Corey: —common guidance.

Brian: Totally. At a big company, managers, you get zero credit for how much your employees get paid, what their raise is, whether they get promoted or not in the grand scheme of things. That is the company running their system. Yes, you helped and the like, but it’s—like, when people tell me, “Hey, Brian, thank you for supporting my promotion.” My answer is always, “Thank you for having earned it. It’s my job to go get credit where credit is due.” And that’s not a big part of my job, and I honestly believe that.

Where you do get credit with people, where you do show that you’re a good manager is when you have the conversations with them that are harder for other people to have, but actually make them better; when you encourage them in the right way so that they grow faster; when you treat them fairly as a human being, and mostly when you do the thing that seems like it’s against your own interest.

Corey: That resonates. The moments of my career as a manager that I’m proud of stuff are the ones that I would call borderline subversive: telling a candidate to take the competing offer because they’re going to have a better time somewhere else is one of those. But my philosophy ties back to the idea of job-hopping, where I’m going to know these people for longer than either of us are going to remain in our current role, on some level. I am curious what your approach is, given that you are now at the, I guess, other end for folks who are just starting out. How do you go about getting people into Cloud marketing? And, on some level, wouldn’t you consider that being a form of abuse?

Brian: [laugh]. It depends on whether they get to work with you or not, Corey.

Corey: There is that.

Brian: I won’t tell you which one’s abuse or not. So first, getting people into cloud marketing is getting people who do not have deeply technical backgrounds in most cases, oftentimes fantastic—people who are fantastic at understanding other people and communicating really well, and it gives them an opportunity to be in tech in one of the fastest-growing, fastest-changing spaces in the world. And so to go to a psych major, a marketing major, an American studies major, a history major, who can understand complex things and then communicate really well, and say, “Hey, I have an opportunity for you to join the fastest growing space in technology,” is often compelling.

But their question kind of is, “Hey, will I be able to do it?” And the answer has to be, “Hey, we have a program that helps you learn, and we have a set of managers who know how to teach, and we create opportunities for you to learn on the job, and we’re invested in you for more than a short period of time.” With that case, I’ve been able to hire and grow and work with, in some cases, people for over 15 years now that I worked with at Microsoft. I’m still in touch with many of the people from the Product Marketing Leadership Development Program at AWS. And we have a fantastic set of APMMs at Google, and it creates a wonderful opportunity for them.

Increasingly, we’re also seeing that it is one of the best ways to find people from many backgrounds. We don’t just show up at the big CompSci schools. We’re getting some wonderful, wonderful people from all the states in the nation, from the historically black colleges and universities, from majors that tend to represent very different groups than the traditional tech audiences. And so it’s been a great source of broadening our talent pool, too.

Corey: There’s a lot to be said for having people who’ve been down this path and seeing the failure modes, reaching out to make so that the next generation—for lack of a better term—has an easier time than we did. The term I’ve heard for the concept is ‘send the elevator back down,’ which is important. I think it’s—otherwise we wind up with a whole industry that looks an awful lot like it did 20 years ago, and that’s not ideal for anyone. The paths that you and I walked are closed, so sitting here telling people they should do what we did has very strong, ‘Okay, Boomer’ energy to it.

Brian: [laugh].

Corey: There are different paths, and the world and industry are changing radically.

Brian: Absolutely. And my—like, the biggest thing that I’d say here is—and again, just coming back to the one thing we disagreed on—look at the bigger picture and own your career. I would never say that isn’t the case, but the bigger picture means not just what you’re getting paid tomorrow, but are you learning more? What new options is it creating for you? And when I speak options, I mean, will you have more jobs that you can do that excite you after you do that job? And those things matter in addition to the pay.

Corey: I would agree with that. Money is not everything, but it’s also not nothing.

Brian: Absolutely.

Corey: I will say though you spent 20 years at Microsoft. I have no doubt that you are incredibly adept at managing your career, at managing corporate politics, at advancing your career and your objectives and your goals and your aspirations within Microsoft, but how does that translate to companies that have radically different corporate cultures? We see this all the time with founders who are ex-Google or ex-Microsoft, and suddenly it turns out that the things that empower them to thrive in the large corporate environment doesn’t really work when you’re a five-person startup, and you don’t have an entire team devoted to that one thing that needs to get done.

Brian: So, after Microsoft, I went to a company called Doppler Labs for a year. It was a pretty well-funded startup that made smart earbuds—this was before AirPods had even come out—and I was really nervous about the going from big company to startup thing, and I actually found that move pretty easy. I’ve always been kind of a hands-on, do-it-yourself, get down in the details manager, and that’s served me well. And so getting into a startup and saying, “Hey, I get to just do stuff,” was almost more fun. And so after that—we ended up folding, but it was a wonderful ride; that’s a much longer conversation—when I got to Amazon and I was in AWS—and by the way, the one division I never worked at Microsoft was Azure or its predecessor server and tools—and so part of the allure of AWS was not only was it another trillion-dollar company in my backwater hometown, but it was also cloud computing, was the space that I didn’t know well.

And they knew that I knew the discipline of product marketing and a bunch of other things quite well, and so I got that opportunity. But I did realize about four months in, “Oh, crap. Part of the reason that I was really successful at Microsoft is I knew how everything worked.” I knew where things have been tried and failed, I knew who to go ask about how to do things, and I knew none of that at Amazon. And it is a—a lot of what allows you to move fast, make good decisions, and frankly, be politically accepted, is understanding all that context that nobody can just tell you. So, I will say there is a cost in terms of your productivity and what you’re able to get done when you move from a place that you’re good at to a place that you’re not good at yet.

Corey: Way back in episode 10 of this podcast—as we get suspiciously close to 300 as best I can tell—I had Lynn Langit get on as a guest. And she was in the Microsoft MVP program, the AWS Hero program, and the Google Expert program. All three at once—

Brian: Lynn is fantastic.

Corey: It really is.

Brian: Lynn is fantastic.

Corey: I can only assume that you listened to that podcast and decided, huh, all three, huh? I can beat that. And decided that—

Brian: [laugh].

Corey: —instead of being in the volunteer to do work for enormous multinational companies group, you said, “No, no, no. I’m going to be a VP in all three of those.” And here we are. Now that you are at Google, you have checked all three boxes. What is the next mountain to climb for you?

Brian: I have no clue. I have no clue. And honestly—again, I don’t know how much of this is privilege versus by being forward-looking. I’ve honestly never known where the heck I was going to go in my career. I’ve just said, “Hey, let’s have a journey, and let’s optimize for doing something you want to do that is going to create more opportunities for you to do something you want to do.”

And so even when I left Microsoft, I was in a great position. I ran the Surface business, and HoloLens, and a whole bunch of other stuff that was really fun, but I also woke up one day and realized, “Oh, my gosh. I’ve been at Microsoft for 20 years. If I stay here for the next job, I’m earning the right to get another job at Microsoft, more so than anything else, and there’s a big world out there that I want to explore a bit.” And so I did the startup; it was fun, I then thought I’d do another startup, but I didn’t want to commute to San Francisco, which I had done.

And then I found most of the really, really interesting startups in Seattle were cloud-related and I had this opportunity to learn about cloud from, arguably, one of the best with AWS. And then when I left AWS, I left not knowing what I was going to do, and I kind of thought, “Okay, now I’m going to do another cloud-oriented startup.” And Google came, and I realized I had this opportunity to learn from another company. But I don’t know what’s next. And what I’m going to do is try and do this job as best I can, get it to the point where I feel like I’ve done a job, and then I’ll look at what excites me looking forward.

Corey: And we will, of course, hold on to this so we can use it for your performance review, whenever that day comes.

Brian: [laugh].

Corey: I want to thank you for taking so much time to speak with me today. If people care more about what you have to say, perhaps you’re hiring, et cetera, et cetera, where can they find you?

Brian: Twitter, IsForAt: I-S-F-O-R-A-T. I’m certainly on Twitter. And if you want to connect professionally, I’m happy to do that on LinkedIn.

Corey: And we will, of course, put links to those things in the [show notes 00:36:03]. Thank you so much for being so generous with your time. I appreciate it. I know you have a busy week of, presumably, attempting to give terrible names to various cloud services.

Brian: Thank you, Corey. Appreciate you having me.

Corey: Indeed. Brian Hall, VP of Product and Industry Marketing at Google Cloud. I am Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an insulting comment in the form of a PowerPoint deck.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Micheal Benedict

Micheal Benedict leads Engineering Productivity at Pinterest. He and his team focus on developer experience, building tools and platforms for over a thousand engineers to effectively code, build, deploy and operate workloads on the cloud. Mr. Benedict has also built Infrastructure and Cloud Governance programs at Pinterest and previously, at Twitter -- focussed on managing cloud vendor relationships, infrastructure budget management, cloud migration, capacity forecasting and planning and cloud cost attribution (chargeback).

Links:

  • Pinterest: https://www.pinterest.com
  • Twitter: https://twitter.com/micheal
  • LinkedIn: https://www.linkedin.com/in/michealb/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: You know how git works right?

Announcer: Sorta, kinda, not really Please ask someone else!

Corey: Thats all of us. Git is how we build things, and Netlify is one of the best way I’ve found to build those things quickly for the web. Netlify’s git based workflows mean you don't have to play slap and tickle with integrating arcane non-sense and web hooks, which are themselves about as well understood as git. Give them a try and see what folks ranging from my fake Twitter for pets startup, to global fortune 2000 companies are raving about. If you end up talking to them, because you don't have to, they get why self service is important—but if you do, be sure to tell them that I sent you and watch all of the blood drain from their faces instantly. You can find them in the AWS marketplace or at www.netlify.com. N-E-T-L-I-F-Y.com

Corey: This episode is sponsored in part by our friends at Vultr. Spelled V-U-L-T-R because they’re all about helping save money, including on things like, you know, vowels. So, what they do is they are a cloud provider that provides surprisingly high performance cloud compute at a price that—while sure they claim its better than AWS pricing—and when they say that they mean it is less money. Sure, I don’t dispute that but what I find interesting is that it’s predictable. They tell you in advance on a monthly basis what it’s going to going to cost. They have a bunch of advanced networking features. They have nineteen global locations and scale things elastically. Not to be confused with openly, because apparently elastic and open can mean the same thing sometimes. They have had over a million users. Deployments take less that sixty seconds across twelve pre-selected operating systems. Or, if you’re one of those nutters like me, you can bring your own ISO and install basically any operating system you want. Starting with pricing as low as $2.50 a month for Vultr cloud compute they have plans for developers and businesses of all sizes, except maybe Amazon, who stubbornly insists on having something to scale all on their own. Try Vultr today for free by visiting: vultr.com/screaming, and you’ll receive a $100 in credit. Thats v-u-l-t-r.com slash screaming.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Sometimes when I have conversations with guests here, we run long. Really long. And then we wind up deciding it was such a good conversation, and there’s still so much more to say that we schedule a follow-up, and that’s what happened today. Please welcome back Micheal Benedict, who is, as of the last time we spoke and presumably still now, the head of engineering productivity at Pinterest. Micheal, how are you?

Micheal: I’m doing great, and thanks for that introduction, Corey. Thankfully, yes, I am still the head of engineering productivity; I’m really glad to speak more about it today.

Corey: The last time that we spoke, we went up one side and down the other of large-scale environments running on AWS and billing aspects thereof, et cetera, et cetera. I want to stay away from that this time and instead focus on the rest of engineering productivity, which is always an interesting and possibly loaded term. So, what is productivity engineering? It sounds almost like it’s an internal dev tools team, or is it something more?

Micheal: Well, thanks for asking because I get this question asked a lot of times. So, for one, our primary job is to enable every developer, at least at our company, to do their best work. And we want to do this by providing them a fast, safe, and a reliable path to take any idea into production without ever worrying about the infrastructure. As you clearly know, learning anything about how AWS works—or any public cloud provider works—is a ton of investment, and we do want our product engineers, our mobile engineers, and all the other folks to be focused on delivering amazing experiences to our Pinners. So, we could be doing some of the hard work in providing those abstractions for them in such way, and taking away the pain of managing infrastructure.

Corey: The challenge, of course, that I’ve seen is that a lot of companies take the approach of, “Ah. We’re going to make AWS available to all of our engineers in it’s raw, unfiltered form.” And that lasts until the first bill shows up. And then it’s, “Okay. We’re going to start building some guardrails around that.” Which makes a lot of sense. There then tends to be a move towards internal platforms that effectively wrap cloud services.

And for a while now, I’ve been generally down on the concept and publicly so in the general sense. That said, what I say that applies as a best practice or something that most people should consider does tend to fall apart when we talk about specific use cases. You folks are an extremely large environment; how do you view it? First off, do you do internal platforms like that? And secondly, would you recommend that other companies do the same thing?

Micheal: I think that’s such a great question because every company evolves with its own pace of development. And I wouldn’t say Pinterest by itself had a developer productivity or an engineering productivity organization from the get-go. I think this happens when you start realizing that your core engineers who are working on product are now spending a certain fraction of time—which starts ballooning pretty fast—in managing the underlying systems and the infrastructure. And at that point in time, it’s probably a good question to ask, how can I reduce the friction in those people’s lives such that they could be focused more on the product. And, kind of, centralize or provide some sort of common abstractions through a central team which can take away all that pain.

So, that is generally a good guiding principle to think about when your engineers are spending at least 30% of their time on operating the systems rather than building capabilities, that’s probably a good time to revisit and see whether a central team would make sense to take away some of that. And just simple examples, right? This includes upgrading OS on your EC2 machines, or just trying to make sure you’re patching all the right versions on your next big Kubernetes cluster you’re running for serving x number of users. The moment you start seeing that, you want to start thinking about, if there is a central team who could take away that pain, what are the things they could be investing on to help up-level every other engineer within your organization. And I think that’s one of the best ways to be thinking about it.

And it was also a guiding principle for us within Pinterest to view what investments we could make in these central teams which can up-level each and every different type of engineer in the company as well. And just an example on that could be your mobile engineer would have very different expectations from your backend engineer who was working on certain aspects of code in your product. And it is truly important to understand where you want to centralize capabilities, which both these types of engineers could use, or you want to divest and have unique capabilities where it’s going to make them productive. There’s no one-size-fits-all solution for this, but I’m happy to talk about what we have at Pinterest, which has been reasonably working well. But I do think there’s a lot more improvements we could be doing.

Corey: Yeah, but let’s also be clear that, as you’ve mentioned, you are heavily biased towards EC2 instances for a lot of what you do. If we look at the AWS console and we see hundreds of different services now, and it’s easy to sit here and say, “Oh, internal platforms are terrible because all of those services are going to be enhanced in various ways and you’re never going to be able to keep up with feature parity.” Yeah, but if you can wrap something like EC2 in an internal platform wrapper, that begins to be a different story because sure, someone’s going to go and try something new with a different AWS service, they’re going to need direct access. But the EC2 product across the board generally does not evolve in leaps and bounds with transformative changes overnight. Let’s also not forget that at a company with the scale that Pinterest operates at, “Hey, AWS just dusted off a new feature and docs are still rolling out, and it’s not in CloudFormation yet, but we’re going to roll it out to production,” probably seems like the wrong direction to go in, I would assume.

Micheal: And yes, I think that brings one of the key guardrails, I think, which these groups provide. So, when we start thinking about what teams, centralized teams like engineering productivity, developer tools, developer platforms actually do is they help with a couple of things. The top three are: they can help pave a path for the most common use cases. Like to your point, provisioning EC2 does take a set of steps, all the time. If you’re going to have a thousand people doing that every time they’re building a new service or trying to expand capacity playing with their launch templates, those are things you can start streamlining and making it simple by some wrapper because you want to address those 80% use cases which are usually common, and you can have a wrapper or could just automate that. And that’s one of the key things: can you provide a paved path for those use cases?

The second thing is, can you do that by having the right guardrails in place? How often have you heard the story that, “I just clicked a button and that now spun up, like, a thousand-plus instances.” And now you have to juggle between trying to stop them or do something about it.

Corey: Back in 2013, you folks were still focusing on this fair bit. I remember because Jeremy Carroll, who I believe was your first SRE there once upon a time, wound up doing a whole series of talks around how Pinterest approached doing an AMI Factory. And back in those days, the challenges were, “Okay. We have the baseline AMI, and that’s great, but we also want to do deployments of things and we don’t really want to do a new deploy of an entire fleet of EC2 instances for a single line of config change, so how do we wind up weighing off of when you bake a new AMI versus when you just change something that has—in what is deployed to them?” And it was really a complicated problem back then.

I’m not convinced it’s not still a complicated problem, but the answers are a lot more cohesive. And making sure that every team—when you’re talking about a company as large as Pinterest with that many teams—is doing things in the same way, seems like it’s critically important otherwise you wind up with a whole bunch of unique-looking instances that each have to be managed by hand as opposed to something that can be reasoned around collectively.

Micheal: Yep. And that last part you mentioned is extremely crucial as well because like I said, our audience or our customers are just not the engineers; we do work with our product managers and business partners as well because at times, we have to tie or change our architecture based on certain cost optimizations which would make sense, like you just articulated. We don’t want to have all the instance types. It does not add much value to a developer unless they’re explicitly seeking a high-memory instance or a [GP-based instance in a 00:10:25] certain way. So, we can then work with our business partners to make sure that we’re committing to only a certain type of instances, and how we can abstract our tools to only give you that. For example, our deployment system, Teletraan which is an open-source system, actually condenses down all these instance types to a couple of categories like high-compute, high-memory—and you’ve probably seen that in many of the new cloud providers as well—so people don’t have to learn or know the underlying instance type.

When we moved from c3 to c5, it was just called as a high-compute system, so the next time someone provisioned a new service or deployed it using our system, they would just select high-compute as the de facto instance type and we would just automatically provision a C5 for them. So, that just reduces the extra complexity or the cognitive overhead individuals would have to go through in learning each instance type, what is the base AMI that comes on it, what are the different configurations that need to go in terms of setting up your AZ-scaling properties. We give them a good reasonable set of defaults to get started with, and then they can then work on optimizing or making changes to it.

Corey: Ignoring entirely your mispronunciation of AMI, which is, of course, three syllables—and that is a petty hill upon which I will die—it occurs to me the more I work with AWS in various ways, the easier it gets. And I used to think in some respects, it was because the platform was so—it was improving so dramatically around me. But no, in many cases, it’s because the first time you write some CloudFormation by hand, it’s a nightmare and you keep smacking into weird issues. But the second or third time, it’s super easy because you just copy the thing you’ve already built and change the relevant bits around. And that was the learning curve that I went through playing around with a lot of these things.

When you start looking at this from a large-scale environment where it’s not just about upskilling the people that you have to understand how these things integrate in AWS land, but also the consistent onboarding of engineers at a fairly progressive clip is, great, you effectively have to start doing trainings on all these things, and there’s a lot of knobs and dials that can blow up and hurt people. At some point, building the guardrails or building the environment in which you are getting all the stuff abstracted away from where the application engineers have to think about this at all, it eventually reaches a tipping point where it starts to feel like it’s no longer optional if you want to continue growing as a company because you don’t have the luxury of spending six months of onboarding before you let someone touch the thing they were hired to build.

Micheal: And you will see that many companies very often have very similar programming practices like you just described. Even I learned that the same way: you have a base template, you just copy-paste it and start from there on. And no one goes through the bootstrapping process manually anymore; you want to—I think we call it cargo-culting, but in general, just get something to bootstrap and start from there. But one of the things we learned in sort of the hard way is that can also lead to, kind of, you pushing, you know, not great practices because people don’t know what is a blessed version of a good template or what actually would make sense. So, some of those things, we have been working on.

And this is where centralized teams like engineering productivity are really helpful is we provide you with the blessed or the canonical way to do certain things. Case in point example is a CI/CD pipeline or delivery of software services. We have invested enough in experimenting on what works with some of the more nuanced use cases at Pinterest, in helping generate, sort of, a canonical version which would cover 80% of the use cases. Someone could just go and try to build a service and they could just use the same canonical pipeline without learning much or making changes to it. This also reduces that cargo-culting nature which I called, rather than copying it from unknown sources and trying to like—again, it may cause havoc to our systems, so we can avoid a lot of that because of these practices.

Corey: So, let’s step a little bit beyond AWS—I know I hate doing it, too—but I’m going to assume that your remit is broader than, oh, AWS whisperer-slash-Wrangler. So, tell me a little bit more about what it is that your day-to-day looks like if there is anything that could be said not to focus purely around AWS whispering.

Micheal: So, one of the challenges—and I want to talk about this a bit more—is our environments have become extremely complex over time. And it’s the nature of, like, rising entropy. Like, we’ve just noticed that there’s two things: we have a diverse set of customer base, and these include everyone trying to do different workloads or work service types. What that essentially translates into is that we realized that our solution may not fit all of them. For example, what works for a machine-learning engineer in terms of iterating on building a model and delivering a model is not the same as someone working on a long-running service and trying to deploy that. The same would apply for someone trying to operate a Kafka system.

And that has made, I think, definitely our job a bit challenging in trying to assess where do you actually draw the line on the abstraction? What is the right layer of abstraction across your local development experience, across when you move over to staging your code in a PR model and getting feedback and subsequently actually releasing it to production? Because this changes dramatically based on what is the workload type you’re working on. And we feel like that has been one of the biggest challenges where I know I spent my day-to-day and my team does too, in trying to help provide some of the right solutions for these individuals. There’s—very often we’ll also get asked from individuals trying to do a very nuanced thing.

Of late, we have been talking about thinking about how you operate functions, like provide Functions as a Service within the company? It just put us in a difficult spot at times because we have to ask the hard question, “Is this required?” I know the industry is doing it; it’s definitely there. I personally believe, yes, it could be a future, but is that absolutely important? Is that going to benefit Pinterest in any formal way if we invest on some core abstractions?

And those are difficult conversations to have because we have exciting engineers coming in trying to do amazing things; it puts us in a hard spot, as well, as to sometimes saying graciously, no. I know many companies deal with it when they have these centralized teams, but I think it’s part of that job. Like when you say it’s day-to-day, I would say I’m probably saying no a couple of times in that day.

Corey: Let’s pretend for the sake of argument that I am, tomorrow morning, starting another company—Twitter for Pets—and over the next ten years, it grows to be larger than Pinterest in terms of infrastructure, probably not revenue because it turns out pets are not the lucrative source of ad revenue that I was hoping it would be but, you know, directionally the same thing. It seems to me that building out this sort of function with this sort of approach to things is dramatically early as far as optimizations go when it’s just me puttering around on something. I’m always cognizant of the wrong people taking the wrong message when we’re talking about things that happen like this at scale. When does having an engineering productivity group begin to make sense?

Micheal: I mentioned this earlier; like, yeah, there is definitely not a right answer, but we can start small. For example, this group actually started more as a delivery team. You know, when we started, we realized that we had different ways of deploying services or software at Pinterest, so we first gathered together to figure out, okay, what are the different ways and can we start simplifying that part? And that’s where it started expanding. Okay, we are doing button-based deployments right now we have thousand-plus microservices, and we are seeing more incidents than we wanted to because anything where there’s a human involved means there’s a potential gap for error. I myself was involved in a SEV 0 incident, and I will be honest; we ended up deploying a Hello World application in one of our production fleet. Not the thing I wanted to be associated with my name, but, you know—

Corey: And you were suddenly saying hello to the world, in fact—

Micheal: [laugh].

Corey: —and oops-a-doozy.

Micheal: Yeah. So—and that really prompted us to rethink how we need to enable guardrails to do safe production rollouts. And that’s how those conversations start ballooning out.

Corey: And the healthy correct way. We’ve all broken production in various ways, and it’s—you correctly are identifying, I believe, the direction you’re heading in where this is a process problem and a tooling problem; it is not that you are secretly crap and should never have been allowed near anything in production. I mean, that’s my excuse for me, but in your case, this is a common thing where it’s, if someone can unintentionally cause issues like that, there needs to be better processes and procedures as the organization matures.

Micheal: Yep. And that’s kind of like always the route or the starting point for these discussions. And it starts growing from there on because, okay, you’ve helped improve the deploy process but now we’re seeing insane amount of slowness, say on the build processes, or even post-deploy, there’s, like, issues on how we monitor and look into data.

And that I think forces these conversations, okay, where do we have these bespoke tools available? What are people doing today? And you have to ask those hard questions, like what can we actually remove from here? The goal is not to introduce yet another new system. Many a times, to be honest bash just gets the job done. [laugh].

Personally, I’m okay with that as long as it’s consistent and people, you know, are able to contribute to it and you have good practices in validating it, if it works, we should go for it rather than introducing yet another YAML [laugh] and some of that other aspects of doing that work. And that’s what we encourage as well. That’s how I think a lot of this starts connecting together in terms of, okay, now this is becoming a productivity group; they’re focused on certain challenges where investing probably one person here may up-level a few other engineers who don’t have to do that on a day-to-day basis. And I think that’s one of the key items for, especially, folks who are running mid-sized companies to realize and start investing in these type of teams to really up-level, sort of, the rest of the engineering.

Corey: You’ve been doing this for a fair while. If you were to go back and start over again on day one—which is always a terrifying question, on some level—what would you have done differently about building out this function as Pinterest continued to scale out?

Micheal: Well, first, I must acknowledge that this was just not me, and there’s, like, ton of people involved in helping make this happen.

Corey: No, that’s fair. We’ll blame them for the missteps; that is—

Micheal: [laugh].

Corey: —just fine with me. I kid. I kid.

Micheal: I think, definitely the nuances. If I look back, all the decisions that were made then at that point in time, there was a decision made to move to Phabricator, which was back then a great open-source code management system where with the current information at that point in time. And I’m not—I think it’s very hard to always look back and say, “Oh, we could have chosen x at one point in time.” And I think in reality, that’s how engineering organizations always evolve, that you have to make do with the information you have right now to make a decision that works for you over a couple of years.

And I’ll give you a small example of this. There was a time when Pinterest was actually on GitHub Enterprise—this was like circa 2013, I would say—and it really served as well for, like, five-plus years. Only then at certain point, we realized that it’s hard to hire PHP engineers to support a tool like that, and we had to rethink what is the ROI and the investments we’ve made here? Can we ever map up or match back to one of the offerings in the industry today? And that’s when you make decisions that, okay, at this point in time, it’s clear that business continuity talks, you know, and it’s hard to operate a system, which is, at this moment not supported, and then you make a call about making a shift or moving.

And I think that’s the key item. I don’t think there’s anything dramatically I would have changed since the start. Perhaps definitely investing a bit more individuals into the group and going from there. But that said, I’m really, sort of, at least proud of the fact that usually these teams are extremely lean and small, and they always have an outsized impact, especially when they’re working with other engineers, other [opinionated 00:22:13] engineers for what it’s worth.

This episode is sponsored by our friends at Oracle Cloud. Counting the pennies, but still dreaming of deploying apps instead of "Hello, World" demos? Allow me to introduce you to Oracle's Always Free tier. It provides over 20 free services and infrastructure, networking databases, observability, management, and security.

And - let me be clear here - it's actually free. There's no surprise billing until you intentionally and proactively upgrade your account. This means you can provision a virtual machine instance or spin up an autonomous database that manages itself all while gaining the networking load, balancing and storage resources that somehow never quite make it into most free tiers needed to support the application that you want to build.

With Always Free you can do things like run small scale applications, or do proof of concept testing without spending a dime. You know that I always like to put asterisks next to the word free. This is actually free. No asterisk. Start now. Visit https://snark.cloud/oci-free that's https://snark.cloud/oci-free.

Corey: Most folks show up intending to do good today, and you make the best decision at the time with the context and constraints that you
have, but my question I think is less around, “Well, what were the biggest mistakes you made?” But more to do with the idea of, based upon what you’ve learned and as you have shown—as you’ve shined light on these dark areas, as you have been exploring it, has anything jumped out at you that is, “Oh, yeah. Now, that I know—if I had known then what I know now, I would definitely have made this other decision.” Ideally, something that applies a little more globally than specific within Pinterest, just because the whole idea, aspirationally, is that people might learn something from our conversation. At least I will, if nothing else.

Micheal: No, I think that’s a great question. And I think the three things that jump to me, top of mind. I think technology is means to an end unless it gives you a competitive edge. And it’s really hard to figure out at what point in time what technology and why we adopted it, it’s going to make the biggest difference. Humans always tend to have a bias towards aligning towards where we want to go. So, that’s the first one in my mind.

The second one is, and we spoke about this last time, embrace your cloud provider as much as possible. You’d want to avoid taking on operational burden which is not going to add value to the business. If there is something you see your operating which can be offloaded—because your provider can, trust me, do a way better job than you or your team of few can ever do—embrace that as soon as possible. It’s better that way because then it frees up your time to focus on the most important thing, which I’ve realized over time is—I really think teams like ours are actually—we’re probably the most value as a glue to all the different experiences a software engineer would go through as part of
their SDLC lifecycle.

If we can simplify someone’s life by giving them a clear view as to where their commit or the work is in this grand scheme of rolling out and giving them the right amount of data to take action when something goes wrong, trust me, they will love you for what you’re doing because you’re saving them ton of time. Many times, we don’t realize that when we publish 11 different ways for you to go and check to just get your basic validation of work done. We tend to so much focus on the technological aspect of what the tool does, rather than the experience of it, and I’ve realized, if you can bridge the experience, especially for teams like ours, people really don’t even need to know whether you’re running Kubernetes or any of those solutions behind the scenes. And I think that’s one of the biggest takeaways I have.

Corey: I want to double down on something you said about the fact that you are not going to be able to run these services as effectively as your provider can. And relatively recently—in fact, since the first time we spoke—AWS has released a investment report in Virginia. And from 2011 through 2020, they have invested in building AWS data centers there, $35 billion. I promise almost no company that employs people listening to this that are not themselves a cloud provider is going to make that kind of investment in running these things themselves.

Now, do cloud providers have sharp edges? Yes, absolutely. That is what my entire career is about, unfortunately. But you’re not going to do a better job of running things more sustainably, more reliably, et cetera, et cetera. But there are other problems with this—and that’s what I want to start exploring here—where in the olden days, when I ran things in data centers and they went down a lot more as a result, sometimes when there were outages, I would have the CEO of the company just standing there nervous worrying over my shoulder as I frantically typed to fix things.

Spoiler: my typing accuracy did not improve by having someone looming over me. Now, when there’s an outage that your cloud provider takes, in many cases the thing that you are doing to fix it is reloading the status page and waiting for an update because it is completely out of your hands. Is that something that you’ve had to encounter? Because you can push buttons and turn dials when things are broken and you control it, but in an AWS—or other cloud provider—outage, all you can really do is wait unless you have a DR plan that is large-scale and effective enough that you won’t feel foolish or have wasted a huge amount of time and energy migrating off and then—because then it gets repaired in ten minutes. How do you approach that, from your perspective? I guess, the expectation management piece?

Micheal: It’s definitely I know something which keeps a lot of folks within infrastructure up at night because, like you just said, at times we can feel extremely powerless when we obviously don’t have direct control—or visibility at times, as well—on what’s happening. One of the things we have realized over time as part of running on our cloud provider for over a decade now, it forces us to rethink a bit on our priority workflows, what we want our Pinners to always have access to, what they need to see, what is not important or critical. Because it puts into perspective, even for the infrastructure teams, is to what is the most important thing we should always have it available and running, what is okay to be in a degraded state, until what time, right? So, it actually forces us to define SLOs and availability criteria within the team where we can broadcast that to the larger audience including the executives. So, none of this comes as a surprise at that point.

I mean, it’s not the answer, probably, you’re looking for because is there’s nothing we can do except set expectations clearly on what we can do and how when you think about the business when these things do happen. So, I know people may have I have a different view on this; I’m definitely curious to hear as well, but I know at Pinterest at least we have converged on our priority workflows. When something goes out, how do we jump in to provide a degraded experience? We have very clear run books to do that, and especially when it’s a SEV 0, we do have clear processes in place on how often we need to update our entire company on where things are. And especially this is where your partnership with the cloud provider is going to be a big, big boon because you really want to know or have visibility, at the minimum some predictability on when things can get resolved, and how you want to work with them on some creative solutions. This is outside the DR strategy, obviously; you should still be focused on a DR strategy, but these are just simple things we’ve learned over time on how to just make it predictable for individuals within the company, so not everyone is freaking out.

Corey: Yeah, from my perspective, I think the big things that I found that have worked, in my experience—mostly by getting them wrong the first time—is explain that someone else running the infrastructure when they take an outage; there’s not much we can do. And no, it’s not the sort of thing where picking up the phone and screaming at someone is going to help us, is the sort of thing that is best to communicate to
executive stakeholders when things are running well, not in the middle of that incident.

Then when things break, it’s one of those, “Great, you’re an exec. You know what your job is? Literally anything other than standing in the middle of the engineering floor, making everyone freak out even more. We’ll have a discussion later about what the contributing factors were when you demand that we fire someone because of an outage. Then we’re going to have a long and hard talk about what kind of culture you’re trying to build here again?” But there are no perfect answers here.

It’s easy to sit here in the silver light of day with things working correctly and say, “Oh, yeah. This is how outages should be handled.” But then when it goes down, we’re all basically an inch away at best from running around with our hair on fire, screaming, “Fix it, fix it, fix it, fix it, now.” And I am empathetic to that. There’s a reason but I fix AWS bills for a living, and one of those big reasons is that it’s a strictly business-hours problem and I don’t have to run production infrastructure that faces anything that people care about, which is kind of amazing and freeing for someone who spent too many years on call.

Micheal: Absolutely. And one of the things is that this is not only with the cloud provider, I think in today’s nature of how our businesses are set up, there’s probably tons of other APIs you are using or you’re working with you may not be aware of. And we ended up finding that the hard way as well. There were a certain set of APIs or services we were using in the critical path which we were not aware of. When these outages happen, that’s when you find that out.

So, you’re not only beholden to your provider at that point in time; you have to have those SLO expectations set with your other SaaS providers as well, other folks you’re working with. Because I don’t think that’s going to change; it’s probably only going to get complicated with all the different types of tools you’re using. And then that’s a trade-off you need to really think about. An example here is just like—you know, like I said, we moved in the past from GitHub to Phabricator—I didn’t close the loop on that because we’re moving back to GitHub right now [laugh] and that’s one of the key projects I’m working with. Yeah, it’s circle of life.

But the thing is, we did a very strong evaluation here because we felt like, “Okay, there’s a probability that GitHub can go down and that means people will be not productive for that couple of hours. What do we do then?” And we had to put a plan together to how we can mitigate that part and really build that confidence with the engineering teams, internally. And it’s not the best solution out there; the other solution was just run our own, but how is that going to make any other difference because we do have libraries being pulled out of GitHub and so many other aspects of our systems which are unknowingly dependent on it anyways. So, you have to still mitigate those issues at some point in your entire SDLC process.

So, that was just one example I shared, but it’s not always on the cloud provider; I think there are just many aspects of—at least today how businesses are run, you’re dependent; you have critical dependencies, probably, on some SaaS provider you haven’t really vetted or evaluated. You will find out when they go down.

Corey: So, I don’t think I’ve told this story before, but before I started this place, I was doing a fair bit of consulting work for other companies. And I was doing a project at Pinterest years ago. And this was one of the best things I’ve ever experienced at a company site, let alone a client site, where I was there early in the morning, eight o’clock or so, so you know, engineers love to show up at the crack of 11:30. But so I was working a little early; it was great. And suddenly my SSH session that I was using to remote into something or other hung.

And it’s tap up, tap enter a couple of times, tap it a couple more. It was hung hard. “What’s the—” and then someone gently taps me on the shoulder. So, I take the headphones off. It was someone from corporate IT was coming around saying, “Hey, there’s a slight problem with our corporate firewall that we’re fixing. Here’s a MiFi device just for you that you can tether to get back online and get worked on until the firewall gets back.”

And it was incredible, just the level of just being on top of things, and the focus on keeping the people who were building things and doing expensive engineering work that was awesome—and also me—productive during that time frame was just something I hadn’t really seen before. It really made me think about the value of where do you remove bottlenecks from people getting their jobs done? It was—it remains one of the most impressive things I’ve seen.

Micheal: That is great. And as you were telling me that I did look up our [laugh] internal system to see whether a user called Corey Quinn existed, and I should confirm this with you. I do see entries over here, a couple of commits, but this was 2015. Was that the time you were around, or is this before that even?

Corey: That would have been around then, yes. I didn’t start this place until late 2016.

Micheal: I do see your commits, like, from 2015, and I—

Corey: And they’re probably terrible, I have no doubt. There’s a reason I don’t read code for a living anymore.

Micheal: Okay, I do see a lot of GIFs—and I hope it’s pronounced as GIF—okay, this is cool. We should definitely have a chat about this separately, Corey?

Corey: Oh, yeah. “Would you explain this code?” “Absolutely not. I wrote it. Of course, I have no idea what it does. That’s the rule. That’s the way code always works.”

Micheal: Oh, you are an honorary Pinterest engineer at this point, and you have—yes—contributed to our API service and a couple of Puppet profiles I see over here.

Corey: Oh, yes—

Micheal: [Amazing 00:36:11]. [laugh].

Corey: You don’t wind up thinking that’s a risk factor that should be disclosed. I kid. I kid. It’s, I made a joke about this when VMware acquired SaltStack and I did some analytics and found that 60 some odd lines of code I had written, way back when that were still in the current version of what was being shipped. And they thought, “Wait, is this actually a risk?”

And no, I am making a joke. The joke is, is my code is bad. Fortunately, there are smart people around me who review these things. This is why code review is so important. But there was a lot to admire when I was there doing various things at Pinterest. It was a fun environment to work in, the level of professionalism was phenomenal, and I was just a big fan of a lot of the automation stuff.

Phabricator was great. I love working with it, and, “Great, I’m going to use this to the next place I go.” And I did and then it was—I looked at what it took to get it up and running, and oh, yeah, I can see why GitHub is so popular these days. But it was neat. It was interesting seeing that type of environment up close.

Micheal: That is great to hear. You know, this is what I enjoy, like, hearing some of these war stories. I am surprised; you seem to have committed way more than I’ve ever done in my [laugh] duration here at Pinterest. I do managing for a living, but then again—Corey, the good news is your code is still running on production. And we—

Corey: Oh dear.

Micheal: —haven’t—[laugh]. We haven’t removed or made any changes to it, so that’s pretty amazing. And thank you for all your contributions.

Corey: Oh, please, you don’t have to thank me. I was paid, it was fine. That’s the value of—

Micheal: [laugh].

Corey: —[work 00:37:38] for hire. It’s kind of amazing. And the best part about consultants is, is when we’re done with a project, we get the hell out everyone’s happy about it.

More happy when it’s me that’s leaving because of obvious personality-related reasons. But it was just an interesting company from start to finish. I remember one other time, I wound up opening a ticket about having a slight challenge with a flickering on my then Apple-branded display that everyone was using before they discontinued those. And I expected there to be, “Oh, okay. You’re a consultant. Great. How did we not put you in the closet with a printer next to that thing, breathing the toner?” Like most consulting clients tend to do, and sure enough, three minutes later, I’m getting that tap on the shoulder again; they have a whole replacement monitor. “Can you go grab a cup of coffee? We’ll run the cable for it. It’ll just be about five minutes.” I started to feel actively bad about requesting things because I did a lot of consulting work for a lot of different companies, and not to be unkind, but treating consultants and contractors super well is not something that a lot of companies optimize for. I can’t necessarily blame them for that. It just really stood out.

Micheal: Yep, I do hope we are keeping up with that right now because I know our team definitely has a lot of consultants working with us as well. And it’s always amazing to see; we do want to treat them as FTs. It doesn’t even matter at that point because we’re all individuals and we’re trying to work towards common goals. Like you just said, I think I personally have learned a few items as well from some of these folks. Which is again, I think speaks to how we want to work and create a culture of, like, we’re all engineers; we want to be solving problems together, and as you were doing it, we want to do it in such a way that it’s still fun, and we’re not having the restrictions of titles or roles and other pieces. But I think I digressed. It was really fun to see your commits though, I do want to track this at some point before we move completely over to GitHub, at least keep this as a record, for what it’s worth.

Corey: Yeah basically look at this graffiti in the codebase of, “A shit-poster was here,” and here I am. And that tends to be, on some level, the mark we live on the universe. What’s always terrifying is looking at things I did 15 years ago in my first Linux admin job. Can I still ping the thing that I built there? Yes, I can. And how is that even possible? That should not have outlived me; honestly, it should never have seen the light of day in production, but here we are. And you never know how long that temporary kluge you put together is going to last.

Micheal: You know, one of the things I was recalling, I was talking to someone in my team about this topic as well. We always talk about 10x engineers. I don’t know what your thoughts are on that, but the fact that you just mentioned you built something; it still pings. And there’s a bunch of things, in my mind, when you are writing code or you’re working on some projects, the fact that it can outlast you and live on, I think that’s a big, big contribution. And secondly, if your code can actually help up-level, like, ten other people, I think you’ve really made the mark of 10x engineer at that point.

Corey: Yeah, the idea of the superhuman engineer is always been a strange and dangerous one. If for nothing else, from where I sit, excellence is inherently situational. Like we just talked about someone at Pinterest: is potentially going to be able to have that kind of impact specifically because—to my worldview—that there’s enough process and things around there that empower them to succeed. Then if you were to take that engineer and drop them into a five-person startup where none of those things exist, they might very well flounder. It’s why I’m always a little suspicious of this is a startup founded by engineers from Google or Facebook, or wherever it is.

It’s, yeah, and what aspects of that culture do you think are one-to-one matches with the small scrappy startup in the garage? Right, I predicting some challenges here. Excellence is always situational. An amazing employee at one company can get fired at a second one for lack of performance, and that does not mean that there’s anything wrong with them and it does not mean that they are a fraud. It means that what they needed to be successful was present in one of those shops, but not the other.

Micheal: This is so true. And I really appreciate you bringing this up because whenever we discuss any form of performance management, that is a—in my view personally—I think that’s an incorrect term to be using. It is really at that point in time, either you have outlived the
environment you are in, or the environment is going in a different direction where I think your current skill set probably could be best used in the environment where it’s going to work. And I know it’s very fuzzy at that point, but like you said, yes, excellence really means you don’t want to tie it to the number of commits you have pushed out, or any specific aspect of your deliverables or how you work.

Corey: There are no easy answers to any of these things, and it’s always situational. It’s why I think people are sometimes surprised when I will make comments about the general case of how things should be, then I talk to a specific environment where they do the exact opposite, and I don’t yell at them for it. It’s there—in a general sense, I have some guidance, but they are usually reasons things are the way they are, and I’m interested in hearing them out. Everything’s situational, the worst consultant in the world is the one that shows up, has no idea what’s going on, and then asked, “What moron set this up?” Invariably, two said, quote-unquote, “Moron.” And the engagement doesn’t go super well from there. It’s, “Okay, why is this the way that it is? What constraints shaped it? What was the context behind the problem you were trying to solve?” And, “Well, why didn’t you use this AWS service?” “Because it didn’t exist for another three years when we were building that thing,” is
a—

Micheal: Yes.

Corey: —common answer.

Micheal: Yes, you should definitely appreciate that of all the decisions that have been made in past. People tend to always forget why they were made. You’re absolutely right; what worked back then will probably not work now, or vice versa, and it’s always situational. So, I think I can go on about this for hours, but I think you hit that to the point, Corey.

Corey: Yeah, I do my best. I want to thank you for taking another block of time out of your day to wind up talking with me about various aspects of what it takes to effectively achieve better levels of engineering productivity at large companies, with many teams, working on
shared codebases. If people want to learn more about what you’re up to, where can they find you?

Micheal: I’m definitely on Twitter. So, please note that I’m spelled M-I-C-H-E-A-L on Twitter. So, you can definitely read on to my tweets there. But otherwise, you can always reach out to me on LinkedIn, too.

Corey: Fantastic and we will, of course, include a link to that in the [show notes 00:44:02]. Thanks once again for your time. I appreciate it.

Micheal: Thanks a lot, Corey.

Corey: Micheal Benedict, head of engineering productivity at Pinterest. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with a comment telling me that you work at Pinterest, have looked at the codebase, and would very much like a refund and an apology.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Guang Ming

Guang Ming Whitley was elected to Mount Pleasant Town Council in 2017 and resides in Old Mount Pleasant with her husband, four children, and a dog.

She earned a B.S. in Chemical Engineering from the University of Southern California and a J.D. from the University of Chicago Law School, where she was a member of Law Review and a moot court semi-finalist. After completing her law degree, Guang Ming taught at the University of Chicago and practiced intellectual property law in Los Angeles. She then retired from active practice to serve as Chief Operating Officer of the Whitley Household. In 2020, she cofounded Lattice Climbers, a company dedicated to teaching soft and life skills to young adults.

Guang Ming is also President of the Girls State Alumnae Foundation and attended the American Legion Auxiliary Girls State in 1996, where she was elected governor. She has volunteered with the ALA Girls State program in a variety of capacities since 2000.

Links:

  • Lattice Climbers: https://www.latticeclimbers.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Honeycomb. When production is running slow, it's hard to know where problems originate: is it your application code, users, or the underlying systems? I’ve got five bucks on DNS, personally. Why scroll through endless dashboards, while dealing with alert floods, going from tool to tool to tool that you employ, guessing at which puzzle pieces matter? Context switching and tool sprawl are slowly killing both your team and your business. You should care more about one of those than the other, which one is up to you. Drop the separate pillars and enter a world of getting one unified understanding of the one thing driving your business: production. With Honeycomb, you guess less and know more. Try it for free at Honeycomb.io/screaminginthecloud. Observability, it’s more than just hipster monitoring.

Corey: You know how git works right?

Announcer: Sorta, kinda, not really Please ask someone else!

Corey: Thats all of us. Git is how we build things, and Netlify is one of the best way I’ve found to build those things quickly for the web. Netlify’s git based workflows mean you don't have to play slap and tickle with integrating arcane non-sense and web hooks, which are themselves about as well understood as git. Give them a try and see what folks ranging from my fake Twitter for pets startup, to global fortune 2000 companies are raving about. If you end up talking to them, because you don't have to, they get why self service is important—but if you do, be sure to tell them that I sent you and watch all of the blood drain from their faces instantly. You can find them in the AWS marketplace or at www.netlify.com. N-E-T-L-I-F-Y.com

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Sometimes people like to ask me what this show is really about and my answer has always been, “The business of cloud,” which is intentionally overbroad; really gives me an excuse to talk about anything that strikes my fancy at a given time. A recurring theme has always been, “Where does the next generation of folks working on cloud come from?”

That’s not strictly bounded to engineers; that goes throughout the entire ecosystem. There are a lot of jobs that are important to the functioning of businesses that don’t require a whole bunch of typing into a text editor and being mad about YAML all day long. Today, my guest is Guang Ming Whitley. Guang Ming, thank you for joining me, I’ll let you tell the story. Who are you exactly?

Guang Ming: Oh, my goodness. That’s a tough question. Well, I am someone who has lived my life in a series of segments. I started off as an engineer—a chemical engineer—then went off to law school, taught for a year—

Corey: Well, let’s interject as well. That is how I got looped into this whole nonsense; you were law school classmates with my spouse. And whatever you’re in town, she gets very excited at the chance to see you, and we finally got to meet not that long ago, had a great conversation. It was, “Oh, my God, you need to come on the podcast.” Which is neither here nor there. Please, continue.

Guang Ming: So, then I had a segment as a stay-at-home mother. I started having babies and I had a lot of them. I had one daughter, then a son, and then I had identical twin boys. And once I started having them in litters, we decided that it was time to stop. So, four kids in and about a decade as a stay-at-home mom, during which time I wrote some books.

And then ran for office back in 2017. And then in 2020, was working with someone just, kind of, over coffee, just having, you know, conversation, and we came up with the idea to start a business, and Lattice Climbers was born out of that.

Corey: And Lattice Climbers is what I think we’re going to be talking about the most today because there’s an entire episode baked into every one of those steps. Maybe not every one of them would fit on a cloud-oriented podcast, but there’s a lot of interesting backstory there and it resonates with me because my entire life has been lived in phases as well. And the more I talk to people, the more I start to realize that maybe I’m not that bizarre. People go through stages and they’d love to retcon what the story was at the time and make it all look like there’s a common thread and narrative running through, but when we’re going through it, it feels—to me at least—like I’ve been careening from thing to thing to thing without ever really having an end goal in mind. But in hindsight, looking back, it just seems like it was inevitable that I would go from where I was to here. It never feels that way at the time for me.

Guang Ming: Well, I think for me, where I’ve ended up with Lattice Climbers has felt sort of inevitable because one of the through-lines of all my segments that I’ve gone through is a program called Girls State, and it is one that I have volunteered with. It’s sponsored by the American Legion Auxiliary and it’s a government simulation program. Over the course of one week, you simulate city, county, and state government. And it’s all about civic engagement and education of young women, and empowerment. So, it’s such a fantastic program.

And I love it, but one of the things that I’ve seen with this program is, as the young women come through the program, some of them have skills, and some of them don’t have skills. And there’s elements that are missing, and that’s something that I want to try to help with, with Lattice Climbers.

Corey: So, what is Lattice Climbers in a nutshell? It’s still very early days, which is fine, terrific; the fact that you care enough about a problem that is clearly plaguing not just our industry, but arguably our entire society is worth exploring in-depth. And with the understanding that the narrative may very well shift as times go on what is Lattice Climbers today?

Guang Ming: So, Lattice Climbers steps into the gap between formal education and the skills necessary to actually adult at life, to survive in the real world.

Corey: That is an area that is of intense interest to me. For listeners who may not have listened to every single episode here, my academic background is checkered, to put it politely. On paper, I have an eighth-grade education and no one can take that away from me. I was expelled from two boarding schools in high school, I wound up getting a diploma from a homeschooling organization that years later I discovered was not accredited, then I failed out of college. But again, no one can take that eighth-grade education away from me.

But also look at me. I am a white dude in tech where my failure mode is a board seat and a book deal somewhere, and there are winds of privilege at my back when I do that. What also has been a strong contributing factor is that when I was 12 years old, my dad sat me down and had a long conversation with me about how to handle a job interview, what a job interview was—because when I was 12, I had no idea—and what they’re looking to gain from asking you these questions, and why they’re asking you the things that they do, what answers they’re looking for, and the purpose behind the meeting that you’re in. And that more than almost anything else as a single moment in my childhood shaped the reason that I became moderately successful [laugh] in my career, depending on what phase of my career we’re talking about. That’s stuff is super important and they don’t teach it formally in any program I’ve ever seen. How do you approach it?

Guang Ming: So, what we do is we have an intake quiz that assesses your skill gaps, sort of like a self-assessment, and then it gives you a customized curriculum just meant to fill your specific skill gaps. So professionalism, where we cover things like interview skills, behavior at events, table manners, those kinds of things. Financial literacy, we have little mantras like, “Credit cards are not free money,” [laugh] which some people never learned that. And then there’s different tracks, so depending on whether or not you’re college-bound, or vocational school, or military-bound, you can pick a different track for that and receive two-minute lessons, sort of the gems, distilled down. And there’s little animations; we try to keep it as brief and information-packed as possible.

Corey: Would it be fair to categorize this as more or less micro-lessons in how to adult?

Guang Ming: Exactly. That is exactly what we’re trying to do.

Corey: I somewhat recently read one of the best stories I’ve ever heard about teaching students in middle school about financial literacy. And invariably, the financial literacy courses are all sponsored by financial institutions, and that’s great. So, what happened was, someone from the bank came in and spoke to the students and then took them all to the bank and had them all open a bank account and deposit $5 into it. Great. A couple of years go by and it earns interest—not much because $5—the bank was then acquired and acquired again and eventually became rolled into Wells Fargo, and had a small balance fee, which then of course wiped out all of these accounts.

And I don’t think that there is any better lesson in the way the financial system works—in some ways—than that. And yes, that’s cynical, but that idea of, if you are sort of toward the bottom, this system is basically stacked against you in a bunch of different ways. Look, I’m not here to rail against capitalism or society as it stands, but understanding that basic concept is foundational to realizing that maybe the credit card company isn’t always your friend with your very best interests in mind.

Guang Ming: Mm-hm. And we tried to explain that, too, you that when you get a credit limit, that is based on what your ability to pay the minimum balance every month. They don’t care if you can pay it off. They care about making that interest off of you, and I think that’s something that children and young adults need to understand.

Corey: It feels like it ties into the idea of thinking critically. The problem with that is the root of that entire financial literacy anecdote that I came out with just now, is that the financial literacy program was developed and promoted by financial institutions. What I like is that I checked your website very briefly, and given the significant absence of a pile of disclosures at the bottom, I don’t believe you’re a bank.

Guang Ming: We are not a bank, and we are not sponsored by a bank. We want to provide practical real-life advice that is useful, and in digestible chunks.

Corey: A while back, before I wound up starting down the path that I’m on now, I basically yelled at people for fun on the internet. I know, imagine that. I was the moderator of two particular subreddits: personal finance—which, great, I spent my 20s in crippling debt; there’s no one as passionate about that stuff as someone who has been converted. Great. And the other was the legal advice subreddit, which is probably horrifying to people like you who are actual attorneys.

But it turns out that an awful lot of what I was doing in both of those subreddits was giving life advice to people on how to function in society. On the legal side of it, “You can’t sue a dog.” “Okay, you are not going to be able to go down to the police station and explain your way out of troubles. Get an attorney.” It’s baseline-level stuff.

“Oh, you’ve been given a contract that seems unreasonable, but they’d say that you need you to sign it.” “Yeah. How about don’t do that without having someone review it?” It’s not actually legal advice. It is how to function in society as an adult, but that’s a less catchy subreddit title as it turns out.

Guang Ming: Well, it’s all about raising your awareness level. So, I have a friend and she tells this story, MBA grad and spent her first six months on the job wearing sneakers every day to work—they were cute, fashionable sneakers, but they were sneakers—and they were not part of appropriate business attire for the work environment she was in because she just was oblivious to that as being an issue. And it took someone who was more senior to finally sit her down and say, “You shouldn’t do this. You need to wear appropriate shoes to work.” And she was mortified but learned from that experience.

So, what if you never had to have that? What if you never had to have that sit-down conversation with someone correcting you? What if you had a little, sort of, pocket guide that gave you that level of awareness? It’s like, “Take a look at your office. See what people are wearing. You can’t wear what the CEO is wearing because you’re not the CEO”—I mean, unless you are the CEO, then you can wear whatever you want, but if you’re just an underling at the company, if you’re just starting out, you need to understand what the company culture is and you need to conform to that culture. Unfortunately, that’s just, like… the truth of the matter.

Corey: The common wisdom is, “Oh, if you don’t know how to dress or how to behave in a certain scenario, reach out to one of your mentors and ask them for advice.” Not everyone has one of those things. I get some crap sometimes through it, but one of the big reasons I have open DMs on Twitter is specifically so people can message me and ask me questions about the industry generally, life in general; I’m always willing to talk to folks who are trying to figure things out. That’s important. Since a disproportionate number of the listeners to this show do work in tech and the idea of having a dress code is ridiculous, yeah, in a lot of tech culture at t-shirt and jeans is just fine, but in other cases, it’s not.

And, for example, I’ll get on stage wearing a full bespoke three-piece suit and give a talk. And it’s fun. It’s hilarious. It plays with people’s expectations, but it’s important to understand I view that more as costuming than I do how I believe someone should necessarily dress in that environment. I am, for better or worse, a very distinctive personality in this space, and using me as a blueprint for someone who is starting out their career is going to lead to disaster.

Yes, I’m mouthy and I make fun of big companies because that’s my thing. I also got fired an awful lot in—

Guang Ming: [laugh].

Corey: —my career, and those two things are not entirely unrelated, let’s be very clear here. There’s a lot that we can learn through observation, but dialing it in and figuring out what the expectations, are important.

Guang Ming: Well, I think a lot of young adults—one of the things we focus on, as well, is the importance of mentoring and finding good mentors. And then you being the kind of person that a mentor would want to mentor. Because I think there’s a lot of formal mentoring in work environments, and those don’t always work as well as the organic relationships. So, we want to be that mentor that you never knew that you needed, the mentor that you wish you always had, to give you all that baseline information so that when you do meet with your substantive mentor, they can truly help you in ways that we cannot with our scalable mentoring micro-lessons.

Corey: I have to ask, what is your revenue model? Because if this turns into charging kids money to learning these things, that has a giant exploitative flashing warning sign around it.

Guang Ming: So, what we’re planning to do is work with school districts and with nonprofits, and do sort of like a B2B model where we pilot with the school district, we pilot with the technical college, and give them an opportunity to add 30 to 50 students, work with the program. And if they find it something valuable, they find that it’s a value-add and it’s helping their students land jobs and have a better career, I think that then they’ll use our program for their full technical school.

Corey: I’m done a fair number of mentorships in the course of my career. I helped administer and run the LOPSA 00:13:43—or League Of Professional System Administrators—mentorship program for a couple of years. The reason that I have a career at all is that people did favors for me, and you can never repay that; you can only pay it forward. So, I had a number of people assigned to me through that program and through other areas as well, and what I’ve learned is that the success of a mentorship is almost entirely on the person seeking guidance: how diligent are they about following up, about going and asking great questions? Because otherwise, if someone comes and says, “Hey, can you mentor me?”—they never frame it quite like that, but that’s fine; the terminology is always squishy here.

Like, “Hey, can you give me advice on things?” “Sure.” And then they don’t ask any questions. Well, if I just butt in with unsolicited advice, that’s not helping them in a mentoring capacity; that’s being a dude on Twitter. So, I’m trying to figure out the way of solving for that, and I don’t know if there is an answer. What’s your take?

Guang Ming: I think that for many young people, there is a baseline level of information that they need, that almost any mentor can give, but it takes up a lot of time to get to that point. So, for example, I had a young woman reach out to me, and she wanted to get a foot in the door in the legal world and wanted some advice. And I couldn’t. It was like, pulling teeth. I couldn’t get her to say a word about herself. And our conversation lasted less than five minutes because I couldn’t get her to speak about herself.

And I almost let it end at that. But then I circled back with her a week later, and called her and said, “You know, I’m going to connect you to someone because I want to help you in your journey. But I need you to think before you get to that conversation about who you are, what you want, where you’re going, what’s your story. You know, I know just from the person who connected us that you’re the first in your family to go to college. Speak to that.” And just really tried to help her understand that she needed to craft a narrative around herself. And I think a lot of young adults don’t know how to craft that narrative.

Corey: The problem that I see when I look at this systemically is that all of this stuff seems like it’s very bespoke. It’s [spreading an 00:15:45] opportunity, but it is incumbent upon folks to learn about it for themselves. One of the most foundational memories of my ill-fated academic career was in public school for my first sophomore year of high school, where the US history teacher said, all right. Today, we’re not doing our traditional stuff, what I’m about to do is not in the curriculum. Please feel free to complain to your parents and then have them take it to the school board.

And what he did was he passed at a flyer where each one of us had different numbers on it, and it was a, “You are a family of x number of people; you made this much money last year.” And then he passed out 1040-EZ forms. And he taught us how to file a tax return in the course of that 45-minute session. And it was, instead of learning a series of whitewashed facts about American history, I was learning how to function as an adult in society. And the fact that he had to do this almost as a subversive thing as opposed to being an accepted part of the curriculum is just mind-boggling to me. I see what you’re doing is important and valuable, but it also in some level kind of feels like a band-aid over a massive societal failing. Is that accurate or am I missing something?

Guang Ming: No, I think that certain school districts are trying to do this, they’re trying to integrate financial literacy into calculus. Some schools will even offer a course, but the course isn’t an AP course; it doesn’t give you special credit, and so students don’t take it, or it’s viewed as a less valuable course even though it’s probably the most valuable course. And there’s also a level of embarrassment. Like, for certain things, we cover personal hygiene. The importance of brushing your teeth every day, and taking a shower, and wearing deodorant.

Which is something you wouldn’t necessarily think you would need to teach someone, but wait till you’re in certain work environments, and that is actually something that people need to know that they’re bothering their coworkers by this lack. That can be really embarrassing.

With Lattice Climbers, you can do this in the privacy of your own home, you can do it in your bedroom, you can do it wherever you are, and you can get these little lessons and not feel embarrassed. Or sometimes you are afraid to ask a question because you feel dumb asking it. When we did a pilot with 17-to 19-year-olds, the favorite video was actually making an appointment. Just giving tips on how to gather the appropriate documentation you would need to, say for example, make a doctor’s appointment, and sample scripts—we have downloadables that go along with—sample scripts of how a conversation would potentially run if you were to call.

Corey: The way you describe this and the problem you’re solving, I have a hard time seeing this as the business opportunity that becomes a $60 billion company because to do that, you would have to do something that is abjectly terrifying. So apparently, becoming rich beyond the wildest dreams of avarice is not the reason that you’re doing this. What made you decide that this was a problem you wanted to address?

Guang Ming: So, I am the daughter of an immigrant and a first-generation college student. And there were so many things that my parents just didn’t know to teach me. They were very focused on academics and there was no focus on anything outside of book smarts. So, when I had my first college interview, my mom took me to the Fashion for Price Boutique next to the Drug Emporium in the strip mall near our home and bought me an interview suit—we didn’t have a ton of money—and the interview suit involved zebra print zippers and a very short skirt. And that is what I wore to my Harvard interview. The one [laugh] school I didn’t get into.

And not only that, not only was I dressed wholly inappropriately, I also was a deer in the headlights. I had never done a mock interview, I had never done anything that would help prepare me for this situation. And I look back at that 17-year-old and I think, “How can I help her? How can I help people like her who don’t have the social or cultural capital to know these things, to know how to move in the world that they want to be in desperately? How do I help them overcome that obstacle?” And that is how Lattice Climbers was born.

Corey: The idea of having an experience like that as being necessary to forge this is—it’s moving. It’s the sort of thing that you hear about other people—you [unintelligible 00:20:04] secondhand cringing from hearing that sort of story, at least I do. And I can definitely understand not wanting other folks to have to go through this. We talk about hilarious interview mistakes that we’ve made, that we’ve had candidates make, and in some cases, most of the ones that I like to talk about are the folks who are—let’s [unintelligible 00:20:23] here—25 years into their career or so, where they really should know better. Because making fun of some naive kid who’d never been in an interview scenario before is just being shitty, let’s be clear.

At some point, though, you should learn how to comport yourself in a working environment that makes sense. But without having mentorships or guidance like that, it feels like a lot of people have stories like this. I think what makes your story different than most of them is that you’re willing to talk about it in public. Most of us bury those things down the memory hole, I would think.

Guang Ming: Yes. I very much own the zebra print story, and it is something that I share when I speak at Girls State. I speak at Girls State just about every year to the young women, and I talk a lot about some of these things that we go over in Lattice Climbers to just try to impart, even in a six-minute speech, some of the key nuggets that I want them to take away with them, as they move through life.

Corey: Tell me a little more about Girls State. I’ve heard the term a couple of times, but know remarkably little about it because, for better or worse, my daughters are still at a point where—I regret this constantly—I have to know entirely too much about the Paw Patrol.

Guang Ming: So, Girls Day is a program sponsored by the American Legion Auxiliary. It is a week-long civic engagement program that simulates government over the course of one week. And for California Girls State, it is one girl from each sponsored high school, with about 540 young women—and of course, we’ve had to be virtual for the past couple of years, but they’ve done it in a webinar virtual sessions—and the program is all about women empowerment and encouraging civic engagement. And one of the things that has really impacted Lattice Climbers has been my observations in Girls State as a counselor for over 20 years. Because we work with young women from all different backgrounds, whether their parents are migrant workers in Modesto or doctors in Big Sur, there are gaps that these young women have that differ based on their backgrounds.

And what I’m hoping to do with Lattice Climbers is fill those gaps and help them avoid these missteps and increase their trajectory as they climb the lattice. And that is one thing that we do is we don’t talk about climbing the ladder because a ladder implies that there is one pathway to the top, there’s room for only one there. We approach it as you’re climbing a lattice: we’re all in it together and there are infinite paths to success.

Corey: All of these things that you talk about are challenging at the best of times, and these are very clearly not the best of times. One of the reasons that, to date, we at The Duckbill Group have not hired junior folks is because in a full-remote environment—and to be clear, even without the pandemic The Duckbill Group has been full-remote since its inception—I don’t know that’s necessarily the best way to expose someone new to the workforce. It feels to me like there’s not a lot of examples around there. There’s a requirement to be a lot more self-directed, and it’s likely, for example, that someone will get stuck and spin on something for a while rather than asking for help because they don’t want to appear like they don’t know what they’re doing and inadvertently make things worse. Do you think that remote as we move forward is going to be an increasing burden on folks like this, or—which I’m perfectly willing to accept—am I completely wrong, and that in fact having a full-remote environment like this is in fact a terrific opportunity for folks new to the workforce?

Guang Ming: No, I think full-remote is an issue. I think that it takes so much more emotional energy to connect through a video than it does to connect in person. And there’s also the lack of organic interactions. There are so many mentorships that develop just from walking down the hall and running into someone over coffee, or at the wat—I mean literally at the watercooler and having the opportunity to chat with someone about something non-work-related that can then evolve into a mentoring relationship. And there is just a lack of that.

All these young people entering the work environment, they can wear pajamas all day and lay in bed with their laptop on their laps and work, and they may love that, but I think that if you want to work in a professional office environment, you need to understand appropriate attire, you need to understand appropriate behavior at events. I think that, especially if you’re from certain backgrounds and you’ve never been around an open buffet before, it can be very tempting to just pile that plate as high as you can with crab legs or, you know, shrimp cocktail. And it’s not appropriate in that setting. And so we cover those—

Corey: Wait. It’s not?

Guang Ming: [laugh]. Well, it depe—if you’re the CEO of the company, Corey, you can do whatever you want.

Corey: No, no. That’s my business partner. I am just the chief cloud economist because it’s not professional to put the word shit-poster on a business card. Or so they tell me.

Guang Ming: [laugh].

Corey: In my experience, the worst of all worlds, though, is not the full-remote; it’s not the in-office; it’s the hybrid scenario where you have some people that are in an office together working and then you have folks who are remote, and regardless of what your intentions are, it is almost impossible to avoid having a striated structure where the in-person folks collaborate in different ways and make decisions informally to which remote folks are not privy. And it’s not to do with cliques or anything like that, but the watercooler discussions, or, “I’m going to go grab lunch. Do you want to come with me?” Type of engagement stories. And I can’t shake the feeling that remote really needs to be all or nothing, at least within the bounds of a team, if not company-wide.

Guang Ming: I think that a hybrid version could work if there was a concerted effort to include the remote individuals if there was a scheduled Zoom happy hour. So, one of the things that happened during the COVID times is there’s a group that I’m a part of, and we just had a happy hour on Saturday nights. At 8 p.m. everyone just kind of logged on and hung out for a period of time. And it was really good to connect with people in that casual environment; there wasn’t always pressure to speak, there wasn’t always pressure to perform, it was just being together and having that togetherness. So, I think that in a work environment, you could create opportunities for that. And then also, I think, bringing people into the office for specific meetings and things that are important like that. And then potentially—I don’t like assigning mentors, but I think you almost have to assign mentors when you have remote workforce.

Corey: This episode is sponsored in part by something new. Cloud Academy is a training platform built on two primary goals. Having the highest quality content in tech and cloud skills, and building a good community the is rich and full of IT and engineering professionals. You wouldn’t think those things go together, but sometimes they do. Its both useful for individuals and large enterprises, but here's what makes it new. I don’t use that term lightly. Cloud Academy invites you to showcase just how good your AWS skills are. For the next four weeks you’ll have a chance to prove yourself. Compete in four unique lab challenges, where they’ll be awarding more than $2000 in cash and prizes. I’m not kidding, first place is a thousand bucks. Pre-register for the first challenge now, one that I picked out myself on Amazon SNS image resizing, by visiting cloudacademy.com/corey. C-O-R-E-Y. That’s cloudacademy.com/corey. We’re gonna have some fun with this one!

Corey: The challenge also becomes one of, great for junior folks, that makes an awful lot of sense. Hire someone with 15 years of experience, and, “Oh, we’re going to assign you a mentor here.” And they’re like, “Oh, really. So, that’s what condescending means. I was always looking for a perfect example.”

Guang Ming: [laugh].

Corey: That’s a delicate balance to strike in my experience.

Guang Ming: Oh, very true. Very true. I was thinking more of, like, the young adults starting out in their career because our focus is really early career, as well as young adults in high school.

Corey: And that’s, I guess, my question for you next is why is that the target age range that you think is best served by this? Now, having, again, spent too much time gazing into the mess that is the Paw Patrol, I understand why preschoolers are not the target market for this, but my approach has generally been targeting folks who are entering the workforce. Although let’s be very clear, a large part of that is because I generally don’t appreciate the optics of going and hanging out at the local high school trying to talk to kids.

Guang Ming: Well, I think that high school is really where it starts. This is the age at which brains are starting to develop a little bit more; they’re starting to have more social awareness. This is where beginnings of your network are important. And I think that the sooner that we can convey to young people that they’re only as strong as their networks, the better. If they can understand that it’s their teachers, their coaches, the parents of their friends are all the beginnings of their network.

That’s how you get internships, that’s how you get a leg up is through these connections because if you’re just a resume floating out there, your chances of getting looked at—and we all know how the world works—well, we should all know how the world works, which is it’s all about your connections that helps you launch to the next thing.

Corey: That’s the thing that I think is understated in this is that we wind up telling students a whole bunch of things that are well-intentioned lies. The, “Oh, put your nose to the grindstone and work hard, and one day you will surely be promoted.” Now, I get flack when I say this sometimes from folks who’ve been at the same company for 15 years and demonstrated growth trajectory internally, but that’s the exception, not the rule. Big moves generally look a lot like transitioning between companies.

“Oh, you don’t want to be a job hopper. It looks bad on the resume.” Yeah, you know who says that? People who don’t want you to quit your job because you’re unhappy because then they have to backfill you, or people who are trying to recruit you in and want to make sure that you when you show up at this new job, you stay there for a while. It’s self-serving.

Yeah, there’s going to be some questions about it in the interview process, but you should have an answer ready to go for it. It’s the interview skills piece of it and make sure that you don’t inadvertently torpedo your own candidacy with conversations like that. And this is stuff that I find that is—it’s not just the newer generation that we’re talking about here; people well into their careers still haven’t cracked a lot of these codes, mostly because, for better or worse, it turns out that people aren’t nearly as cynical [laugh] about things as I am.

Guang Ming: Well, and we also cover things like how to leave a job professionally. Because as we live in a world where you’re not going to go work for one company for the next 30 years, or where you shouldn’t go work for the same company for 30 years necessarily, but there are stories out there of people just ghosting on the job, ghosting on job interviews, and that burns bridges. And everyone you meet is a potential connection in your network as you climb the lattice, and so you need to preserve those relationships moving forward because you never know who you help out along the way, or who helps you out along the way, you never know how that connection is going to play out later on in life.

Corey: That’s the trick is that it’s talking to people and being friendly with them. And there are ways to do networking properly in my world, and there are ways not to. And, “Oh, I should talk to you because down the road you might be useful to me,” is just cynical and terrible. I hate the pattern.

Whereas, I like keeping in touch with people because I find them interesting. My default assumption has always been that I’m going to be talking to someone for longer than either one of us is going to be doing whatever it is we’re currently doing, and trying to treat relationships as transactional is a mistake. But that’s what networking is often interpreted as.

Guang Ming: It’s so true. And people can tell. They can tell when you’re being fake. They can tell when you’re being transactional. They can tell when you just are waiting for the ask.

I think it actually is really hard to be genuine and natural for some people that comes across as transactional, and one of the ways that we talked about avoiding that is through just an ongoing relationship. So, you don’t only reach out to the person when you have an ask, you reach out to the person quarterly. And you can have a spreadsheet—almost—about it, and of the people that you want to contact and maintain cont—and even if it’s just a text message that says, “Hey, this is what I’m up to. Hope all is well with you.” And even if they don’t respond, or just it’s a one-word answer, you’ve at least had that touchpoint with them over the course of time.

Corey: There’s often a criticism levied at folks who are advocating for networking, that it is a lot harder when you’re an introvert or when you are neurodivergent, in certain ways. To be clear, I’ve neurodivergent in ways that do not directly negatively impact my ability to socialize with folks; it just means they think I’m a jerk. But there are folks who definitely have different expressions of different divergences. And that’s fine. How do you view the networking aspect for folks who do not work nearly as well interpersonally?

Guang Ming: That’s so hard because interpersonal skills are something that is so necessary, and I think that unfortunately, there are people who get by one hundred percent on their social skills. Like, their people skills are all they need to move forward in the world. And I think that you have to work at it, and you have to study how to behave in those situations. It’s almost like—so for example, my husband is an introvert, but he was also an actor in college. And when he goes into these situations, it’s almost like putting on a show.

Like you talked about putting on your three-piece suit. There is the extrovert persona that he wears in these environments, and then he takes it off when he gets home. And I think that you almost have to create that persona for yourself. And you can acknowledge that you’re neurodivergent, and you can acknowledge that you’re an introvert, and I think that’s way more acceptable these days than it used to be. And there are lots of people that are in the world that are neurodivergent and are introverts, and so I think it’s completely fine to be that way.

Corey: I’ve never had a good answer for folks who ask those questions, just because it is so different from my lived experience that I don’t have an answer that’s worth listening to, and I try very hard to stay in my lane. I don’t ever want that to be interpreted as it’s not important because it very much is.

So, one last question I have for you is I love, love, love your zebra print suit story, but it’s also back when you were applying to school, back in early career, which you are very clearly not now; it’s decades old. Do you have any other similar stories from folks that you’ve been working through, either at Lattice Climbers or through Girls State, that illustrate this in a somewhat more modern era?

Guang Ming: Oh, absolutely. So, there was a young woman at Girls State; we were all in a room and they were talking about colleges, the girls were talking about colleges. And this one young woman remains silent during this conversation. And so I approached her. I said, “Well, what about you? What are your plans for college?” And she shared that she wasn’t going to go to college because her parents didn’t go to college, they didn’t have a lot of money, and she just didn’t think that college was in the cards for her.

And I disabused her of that notion. I told her absolutely not. You’re here at Girls State, which means you’re the top girl from your high school. You absolutely should go to college, and I told her that there were so many paths to college. You could go to community college and then transfer; you could go to technical college; there’s so many different options.

And there’s so many scholarships out there, especially for low-income individuals. Well, we became friends on social media, and about three years after Girls State—because they attend the summer after their junior years—I received a message from her, sort of, out of the blue, and she let me know that that conversation that we had changed her life because she had gone to community college—she had taken my advice; she had gone to community college and had just been accepted as a transfer student to UC Berkeley. And that story just makes me tear up every time I think about it. And that one conversation had that huge impact on her life, and I’m hoping that through Lattice Climbers and our little lessons, that we can have that kind of impact on young lives, that we can help them avoid these missteps that could have huge impacts on their trajectory, and we could help them increase their trajectory on the lattice.

Corey: It’s similar in some respects to the folks I talk to who are building products for the cloud industry. It’s, “Yes, yes, of course. You’re always going to have a story about how it works for you. That’s fine. Let’s talk about your customers.” Like, “Find me a customer, someone else in the world who has a story like this that really demonstrates the value you provide.”

And I love the fact that it is so easy for you to come up with these things off the top of your head, even when you weren’t necessarily expecting the question. So, you’re onto something. This is a clear problem, and it’s not going away anytime soon, and it’s largely underserved because there’s no opportunity to invest venture capital into it and make a ridiculous return on that investment because there’s not money in solving it that I can see—and apparently, most the industry can see—compared to another Twitter for Pets app.

Guang Ming: [laugh]. Well, there is not that much money in them there hills because no one owns the problem, and because no one owns the problem, it’s very hard to find people willing to pay to solve the problem. But that doesn’t mean that the problem isn’t there and that doesn’t mean that it doesn’t need to be solved. And I actually think that companies should have an incentive to do it because it will help with employee retention, it will help with employee performance if they do invest in their workers, and in high school students who, the sooner that they know these things, the better it will be for their long-term careers.

Corey: And if nothing else, I think that’s the lesson to take away from this for the young folk—the youth, as it were—that this is the single greatest thing I look at and credit my professional trajectory has been in learning to handle expectations in corporate environments. And sure, I have fun with them and I play games with them, but you have to know the rules before you can break them in this context. And there are business meetings in which I assure you, you would question whether it was the same person. And that’s what it comes down to, I think, on some level is, if you know how to handle a job interview, you will always be able to find something to put food on the table. Conversely, if you’re terrific at any number of different things, but absolutely cannot handle the dynamics of a job interview, you are going to struggle to find work anywhere until you find someone willing to alter their corporate process just in order to bring you aboard.

It’s a skill that you need to be at least conversant with. And what makes it even worse, as it’s a skill that you only really get to practice when you’re looking for jobs. I want to thank you for taking the time to speak with me so much about all this stuff. If people want to learn more about what you’re up to and how you’re approaching it, where can they find you?

Guang Ming: So, we are at latticeclimbers.com. And we are currently in waitlist mode, so you can sign up on our waitlist and get more information about when we’re ready to launch. We are working with some nonprofits and some school districts on some pilot programs, and we’re hoping to have that going, hopefully by the end of the year.

Corey: And we will, of course, put a link to that in the [show notes 00:38:07]. Thank you so much for taking the time to speak with me today. I really appreciate it.

Guang Ming: Well, thank you so much for having me. I really appreciate the opportunity to share what we’re doing with Lattice Climbers, and I just hope, like I said, if I can get one person to not wear zebra print to that Harvard interview, [laugh] then I will view Lattice Climbers as a success.

Corey: [laugh]. Excellent. Thank you so much. Once again, Guang Ming Whitley, co-founder of Lattice Climbers. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, along with a rambling comment explaining why we’re wrong and that a zebra-print suit for a college interview is in fact a best practice.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Matthew

Matthew Prince is co-founder and CEO of Cloudflare. Cloudflare’s mission is to help build a better Internet. Today the company runs one of the world's largest networks, which spans more than 200 cities in over 100 countries. Matthew is a World Economic Forum Technology Pioneer, a member of the Council on Foreign Relations, winner of the 2011 Tech Fellow Award, and serves on the Board of Advisors for the Center for Information Technology and Privacy Law. Matthew holds an MBA from Harvard Business School where he was a George F. Baker Scholar and awarded the Dubilier Prize for Entrepreneurship. He is a member of the Illinois Bar, and earned his J.D. from the University of Chicago and B.A. in English Literature and Computer Science from Trinity College. He’s also the co-creator of Project Honey Pot, the largest community of webmasters tracking online fraud and abuse.

Links:

  • Cloudflare: https://www.cloudflare.com
  • Blog post: https://blog.cloudflare.com/aws-egregious-egress/
  • Bandwidth Alliance: https://www.cloudflare.com/bandwidth-alliance/
  • Announcement of R2: https://blog.cloudflare.com/introducing-r2-object-storage/
  • Blog.cloudflare.com: https://blog.cloudflare.com
  • Duckbillgroup.com: https://duckbillgroup.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Writing ad copy to fit into a 30 second slot is hard, but if anyone can do it the folks at Quali can. Just like their Torque infrastructure automation platform can deliver complex application environments anytime, anywhere, in just seconds instead of hours, days or weeks. Visit Qtorque.io today and learn how you can spin up application environments in about the same amount of time it took you to listen to this ad.

Corey: This episode is sponsored in part by Honeycomb. When production is running slow, it's hard to know where problems originate: is it your application code, users, or the underlying systems? I’ve got five bucks on DNS, personally. Why scroll through endless dashboards, while dealing with alert floods, going from tool to tool to tool that you employ, guessing at which puzzle pieces matter? Context switching and tool sprawl are slowly killing both your team and your business. You should care more about one of those than the other, which one is up to you. Drop the separate pillars and enter a world of getting one unified understanding of the one thing driving your business: production. With Honeycomb, you guess less and know more. Try it for free at Honeycomb.io/screaminginthecloud. Observability, it’s more than just hipster monitoring.

Corey: Welcome to Screaming in the Cloud, I’m Corey Quinn. Today, my guest is someone I feel a certain kinship with, if for no other reason than I spend the bulk of my time antagonizing AWS incredibly publicly. And my guest periodically descends into the gutter with me to do the same sort of things. The difference is that I’m a loudmouth with a Twitter account and Matthew Prince is the co-founder and CEO of Cloudflare, which is, of course, publicly traded. Matthew, thank you for deigning to speak with me today. I really appreciate it.

Matthew: Corey, it’s my pleasure, and appreciate you having me on.

Corey: So, I’m mostly being facetious here, but not entirely, in that you have very publicly and repeatedly called out some of the same things I love calling out, which is AWS’s frankly egregious egress pricing. In fact, that was a title of a blog post that you folks put out, and it was so well done I’m ashamed I didn’t come up with it myself years ago. But it’s something that is resonating with a large number of people in very specific circumstances as far as what their company does. Talk to me a little bit about that. Cloudflare is a CDN company and increasingly looking like something beyond that. Where do you stand on this? What got you on this path?

Matthew: I was actually searching through really old emails to find something the other day, and I found a message from all the way back in 2009, so actually even before Michelle and I had come up with a name for Cloudflare. We were really just trying to understand the pricing on public clouds and breaking it all down. How much does the compute cost? How much does storage cost? How much does bandwidth cost?

And we kept running the numbers over and over and over again, and the storage and compute costs actually seemed relatively reasonable and you could understand it, but the economics behind the bandwidth just made no sense. It was clear that as bandwidth usage grew and you got scale that your costs eventually effectively went to zero. And I think it was that insight that led to us starting Cloudflare. And the self-service plans at Cloudflare have always been unlimited bandwidth, and from the beginning, we didn’t charge for bandwidth. People told us at the time we were crazy to not do that, but I think that that realization, that over time and at scale, bandwidth costs do go to zero is really core to who Cloudflare is.

Cloudflare launched a little over 11 years ago now, and as we’ve watched the various public clouds and AWS in particular just really over that same 11 years not only not follow the natural price of bandwidth down, but really hold their costs steady. At some point, we’ve got a lot of mutual customers and it’s a complaint that we hear from our mutual customers all the time, and we decided that we should do something about it. And so that started four years ago, when we launched the Bandwidth Alliance, and worked with almost all the major public clouds with the exception of Amazon, to say that if someone is sending traffic from a public cloud network to Cloudflare’s network, we’re not going to charge them for the bandwidth. It’s going across a piece of fiber optic cable that yeah, there’s some cost to put it in place and maybe there’s some maintenance costs associated with it, but there’s not—

Corey: And the equipment at the end costs money, but it’s not cloud cost; it just cost on a per second, every hour of your lifetime basis. It’s a capital expense that is amortized across a number of years et cetera, et cetera.

Matthew: And it’s a fixed cost. It’s not a variable cost. You put that fiber optic cable and you use a port on a router on each side. There’s cost associated with that, but it’s relatively de minimis. And so we said, “If it’s not costing us anything and it’s not costing a cloud provider anything, why are we charging customers for that?”

And I think it’s an argument that resonated with almost every other provider that was out there. And so Google discounts traffic when it’s sent to us, Microsoft discounts traffic when it’s sent to us, and we just announced that Oracle has joined this discounting their traffic, which was already some of the most cost-effective bandwidth from any cloud provider.

Corey: Oh, yeah. Oracle’s fantastic. As you were announced, I believe today, the fact that they’re joining the Bandwidth Alliance is both fascinating and also, on some level, “Okay. It doesn’t matter as much because their retail starting cost is 10% of Amazon’s.” You have to start pushing an awful lot of traffic relative to what you would do AWS before it starts to show up. It’s great to see.

Matthew: And the fact that they’re taking that down to effectively zero if you’re using us is even better, right? And I think it again just illustrates how Amazon’s really alone in this at being so egregious in how they do that. And it’s, when we’ve done the math to calculate what their markups are, it’s almost 80 times what reasonable assumptions on what their wholesale costs are. And so we really do believe in fighting for our customers and being customer-centric, and this seems like a place where—again, Amazon provides an incredible service and so many things, but the data transfer costs are just completely outrageous. And I’m glad that you’re calling them out on it, and I’m glad we’re calling them out on it and I think increasingly they look isolated and very anti-customer.

Corey: What’s interesting to me is that ingress to AWS at all the large public tier-one cloud providers is free. Which has led, I think, to the assumption—real or not—that bandwidth doesn’t actually cost anything, whereas going outbound, all I can assume is that one day, some Amazon VP was watching a rerun of Meet the Parents and they got to the line where Ben Stiller says, “Oh, you can milk anything with nipples,” and said, “Holy crap. Our customers all have nipples; we can milk them with egress charges.” And here we are. As much as I think the cloud empowers some amazing stuff, the egress charges are very much an Achilles heel to a point where it starts to look like people won’t even
consider public cloud for certain workloads based upon that.

People talk about how Netflix is a great representation of the ideal AWS customers. Yeah, but they don’t stream a single byte to customers from AWS. They have their own CDN called Open Connect that they put all around the internet, specifically for that use case because it would bankrupt them otherwise.

Matthew: If you’re a small customer, bandwidth does cost something because you have to pay someone to do the work of interconnecting with all of the various networks that are out there. If you start to be, though, a large customer—like a Cloudflare, like an AWS, like an Azure—that is sending serious traffic to the internet, then it starts to actually be in the interest of ISPs to directly interconnect with you, and the costs of your bandwidth over time will approach zero. And that’s the just economic reality of how bandwidth pricing works. I think that the confusion, to some extent, comes from all of us having bought our own home internet connection. And I think that the fact that you get more bandwidth up in most internet connections, and you get down, people think that there’s some physics, which is associated with that.

And there are; that turns out just to be the legacy of the cable system that was really designed to send pictures down to your—

Corey: It wasn’t really a listening post. Yeah.

Matthew: Right. And so they have dedicated less capacity for up and again, in-home network connections, that makes a ton of sense, but that’s not how internet connections work globally. In fact, you pay—you get a symmetric connection. And so if they can demonstrate that it’s free to take the traffic in, we can’t figure out any reason that’s not simply about customer lock-in; why you would charge to take data out, but you wouldn’t charge to put it in. Because actually cost more from writing data to a disk, it costs more than reading it from a disk.

And so by all reasonable accounts, if they were actually charging based on what their costs were, they would charge for ingress but they want to charge for egress. But the approach that we’ve taken is to say, “For standard bandwidth, we just aren’t going to charge for it.” And we do charge for if you use our premium routing services, which is something called Argo, but even then it’s relatively cheap compared with what is just standard kind of internet connectivity that’s out there. And as we see more of the clouds like Microsoft and Google and Oracle show that this is a place where they can be much more customer-centric and customer-friendly, over time I’m hopeful that will put pressure on Amazon and they will eliminate their egress fees.

Corey: People also tend to assume that when I talk about this, that I’m somehow complaining about the level of discounting or whatnot, and they yell at me and say, “Oh, well, you should know by now, Corey, that no one at significant scale pays retail pricing.” “Thanks, professor. I appreciate that, but four years ago, or so I sat down with a startup founder who was sketching out the idea for a live video streaming service and said, ‘There’s something wrong with my math because if I built this on AWS—which he knew very well, incidentally—it looks like it would cost me at our scale of where we’re hoping to hit $65,000 a minute.’” And I checked and yep, sure enough, his math was not wrong, so he obviously did not build his proof of concept on top of AWS. And the last time I checked, they had raised several 100 million dollars in a bunch of different funding rounds.

That is a company now that will not be on AWS because it was never an option. I want to talk as well about your announcement of R2, which is just spectacular. It is—please correct me if I get any of this wrong—it’s an object store that lives in your existing distributed-points-of-presence-slash-data-centers-slash-colo-slash-a-bunch-of-computers-in-fancy-warehouse-rooms-with-the-lights-are-always-on-And-it’s-always-cold-and-noisy. And people can store data there—

Matthew: [crosstalk 00:10:23] aisles it’s cold; in the other aisles, it’s hot. But yes.

Corey: Exactly. But it turns out when you lurk around to the hot aisle, that’s not where all the buttons are and the things you’re able to plug into, so it’s freeze or sweat, and there’s never a good answer. But it’s an object store that costs a fair bit less than retail pricing for Amazon S3, or most other object stores out there. Which, okay, great. That’s always good to see competition in the storage space, but specifically, you’re not charging any data transfer costs whatsoever for doing this. First, where did this come from?

Matthew: So, we needed it ourselves. I think all of the great products at Cloudflare start with an internal need. If you look at why do we build our zero-trust solutions? It’s because we said we needed a security solution that was fast and reliable and secure to protect our employees as they were going out and using the internet.

Why did we build Cloudflare Workers? Because we needed a very flexible compute platform where we could build systems ourselves. And that’s not unique to us. I mean, why did Amazon build AWS? They built it because they needed those tools in order to continue to grow and expand as quickly as possible.

And in fact, I think if you look at the products that Google makes that are really great, it ends up being the ones that Google’s employees use themselves. Gmail started as Caribou once upon a time, which was their internal email system. And so we needed an object store and the sometimes belligerent CEO of Cloudflare insisted that our team couldn’t use any of the public cloud object stores. And so we had to build it.

That was the start of it and we’ve been using it internally for products over time. It powers, for example, Cloudflare Images, it powers a lot of our streaming video services, and it works great. And at some point, we said, “Can we take this and make it available to everyone?” The question that you’ve asked on Twitter, and I think a lot of people reasonably ask us, “What’s the catch?”

Corey: Well, in my defense, I think it’s fair. There was an example that I gave of, “Okay, I’m going to go ahead and keep—because it’s new, I don’t trust new object stores. Great. I’m going to do the same experiment twice, keep one the pure AWS story and the other, I’m just going to add Cloudflare R2 to the mix so that I have to transfer out of AWS once.” For a one gigabyte file that gets shared out for a petabyte’s worth of bandwidth, on AWS it costs roughly $52,000 to do that. If I go with the R2 solution, it cost me 13 cents, all of which except for a penny-and-a-half are AWS charges. And that just feels—when you’re looking at that big of a gap, it’s easy to look at that and think, “Okay, someone is trying to swindle me somewhere. And when you can’t spot the sucker, it’s probably me. What’s the catch?”

Matthew: I guess it’s not really a catch; it’s an explanation. We have been able to drive our bandwidth costs down low enough that in that particular use case, we have to store the file, and that, again, that—there’s a hard disk in there and we replicate it to make sure that it’s available so it’s not just one hard disk, but it’s multiple hard disks in various places, but that amortized over time, isn’t that big a cost. And then bandwidth is effectively zero. And so if we can do that, then that’s great.

Maybe a different way of framing the question is like, “Why would we do that?” And I think what we see is that there is an opportunity for customers to be able to use the best of various cloud providers and hook the different parts together. So, people talk about multi-cloud all the time, and for a while, the way that I think people thought about that was you take the exact same workload and you run it in Azure and AWS. That turns out not to be—I mean, maybe some people do that, but it’s super rare and it’s incredibly hard.

Corey: It has been a recurring theme of most things I say where, by default, that is one of the dumbest things I can imagine.

Matthew: Yeah, that isn’t good. But what people do want to do is they want to say, “Listen, there’s some really great services that Amazon provides; we want to use those. And there’s some really great services that Azure provides, and we want to use those. And Google’s got some great machine learning, and so does IBM. And I want to sort of mix and match the various pieces together.”

And the challenge in doing that is the egress fees. If everyone just had a detente and said there's going to be no egress fees for us to be able to hook these various [pits 00:14:48] together, then you would be able to take advantage of a lot of the different technologies and we would actually get stronger applications. And so the vision of what we’re trying to build is how can we be the fabric that can stitch the various cloud providers together so that you can do that. And when we looked at that, and we said, “Okay, what’s the path to getting there?” The big place where there’s the just meatiest cost on egress fees is object stores.

And so if you could have a centralized object store, and you can say then from that object go use whatever the best service is at Amazon, go use whatever the best service is at Google, go use whatever the best service is at Azure, that then allows, I think, actually people to take advantage of the cloud in a way which is what people really should mean when they talk about multi-cloud. Which is, there should be competition on the various features themselves, and you should be able to pick and choose the best of all of the different bits. And I think we as consumers then benefit from that. And so when we’re looking at how we can strategically enable that future, building an object store was a real key part of that, and that’s part of what we’re doing. Now, how do we make money off of that? Well, there’s a little bit off the storage, and again, even [laugh]—

Corey: Well, that is the Amazonian answer there. It’s like, “Your margin is my opportunity,” is a famous Bezos quote, and I figure you’re sitting there saying, “Ah, it would cost $52,000 to do that in Amazon. Ah, we can make a penny-and-a-half.” That’s very Amazonian, you could probably get hired over there with that philosophy.

Matthew: Yeah. And this is a commodity service, just [laugh] storing data. If you look across the history of what Cloudflare has done, in 2014, we made encryption free because it’s absurd to pay for math, right? I mean, it’s just crazy right?

Corey: Or to pay for security as a value-add. No, that should be baked into whatever you’re doing, in an ideal world.

Matthew: Domain registration. Like, it’s writing something down in a ledger. It’s a commodity; of course it should go to whatever the absolute cost is. On the other hand, there are things that we do that aren’t commodities where we are able to better protect people because we see so much traffic, and we’ve built the machine learning models, and we’ve done those things, and so we charge for those things. So commodities, we think over time, go to effectively, whatever their cost is, and then the value is in the actual intelligent services that are on top of it.

But an object store is a commodity and so we should be trying to drive that pricing down. And in the case of bandwidth, it’s effectively free for us. And so if we can be that fabric that connects the different class together, I think that makes sense is a strategy for us and that’s why R2 made a ton of sense for us to build and to launch.

Corey: There seems to be a lack of ability for lots of folks, at least on the internet to imagine a use case other than theirs. I cheated by being a consultant, I get to borrow other people’s use cases at a high degree of turnover. But the question I saw raised was, “Well, how many workloads really do that much egress from static objects that don’t change? Doesn’t sound like there’d be a whole lot of them.” And it’s, “Oh, my sweet summer child. Sure, your app doesn’t do a lot of that, but let me introduce it to my friends who are hosting videos on their website, for example, or large images that get accessed a whole bunch of times; things that are written once and then read forever by the internet.”

Matthew: And we sit in a position where because of the role that Cloudflare plays where we sit in front of a number of these different cloud providers, we could actually look at the use cases and the data, and then build products in order to solve that. And that’s why we started with Workers; that’s why we then built the KV store that was on top of that; we built object-store next. And so you can see as we’re sort of marching through these things, it is very much being informed by the data that we actually see from real customers. And one of the things that I really like about R2 is in exactly the example that you gave where you can keep everything in S3; you can set R2 in front of it and put it in slurp mode, and effectively it just—as those objects get pulled out, it starts storing them there. And so the migration path is super easy; you don’t have to actually change anything about your application and will cut your bills substantially.

And so I think that’s the right thing to enable a multi-cloud world where, again, it’s not you’re running the exact same workload in different places, but you get to take advantage of the really great tack that all of these companies are building and use that. And then the companies will compete on building that tech well. So, it’s not just about how do I get the data in and then kind of underinvest in all of the different services that I provide. It’s how can we make sure that on a service-by-service basis, you actually are having real competition over time. And again, I think that’s the right thing for customers, and absolutely R2 might not be the right thing for every use case that’s out there, but I think that it wi—enabling more competition is going to make the cloud better for everyone.

Corey: Oh, yeah. It’s always fun hearing it from Amazonians. It’s, “You have a service that talks to satellites in orbit. You really think that’s a general-purpose thing that every company out there has to deal with?” No. Well, not yet, anyway.

It also just feels to me like their transfer approach is antithetical to almost every other aspect of how they have built their cloud. Amazonians have told me repeatedly—I believe them—that their network is effectively magic. The fact that you can get near line rate between any two points without melting various [unintelligible 00:20:14], which shows that there was significant thought, work, effort, planning, technology, et cetera, put into the network. And I don’t dispute that. But if I’m trying to build a workload and put it inside of AWS, I can control how it performs tied to budget; I can have a lot of RAM for things that are memory intensive, or I can have a little RAM; I can have great CPU performance or terrible CPU performance.

The challenge with data transfer is it is uniformly great. “I want to get that data over there super quickly.” Yeah, awesome. I’m fine paying a premium for that. But I have this pile of data right here. I want to get it over there, ideally by Tuesday. There’s no good way to do that, even with their Snowball—or Snow Family devices—when you fill them with data and send them into AWS, yeah, that’s great. Then you just pay for the use of the device.

Use them to send data out of AWS, they tack on an additional per-gigabyte fee for getting the data out. You’re training as a lawyer, you went to the same law school that my wife did, the University of Chicago, which, oh, interesting stories down that path. But if we look at this, my argument is that the way to do an end-run around this is to sue Amazon for something, and then demand access to the data you have living in their environment during discovery. Make them give it to you for free, though, they’d probably find a way to charge it there, too. It’s just a complete lack of vision and lack of awareness because it feels like they’re milking a cash cow until it dies.

Matthew: Yeah, they probably would charge for it and you’d also have to pay a lot of lawyers. So, I’m not sure that’s the cost [crosstalk 00:21:44]—

Corey: Its only works above certain volumes, I figure.

Matthew: I do think that if your pricing strategy is designed to lock people in to prevent competition, then that does create other challenges. And there are certainly some University of Chicago law professors out there that have spent their careers arguing why antitrust laws don’t make any sense, but I think that this is definitely one of those areas where you can see very clearly that customers are actually being harmed by the pricing strategy that’s there. And the pricing strategy is not tied in any way to the underlying costs which are associated with that. And so I do think that, especially as you see other providers in the space—like Oracle—taking their bandwidth costs to effectively zero, that’s the sort of thing that I think will have regulators start to scratch their heads. If tomorrow, AWS took egress costs to zero, and as a result, R2 was not as advantaged as it is today against them, you know, I think there are a lot of people who would say, “Oh, they showed Cloudflare.” I would do a happy dance because that’s the best thing [thing they can do 00:22:52] for our customers.

Corey: Our long-term goals, it sounds like, are relatively aligned. People think that I want to see AWS reign ascendant; people also say I want to see them burning and crashing into the sea, and neither one of those are true. What I want is, I want someone in a few years from now to be doing a startup and trying to figure out which cloud provider they should pick, and I want that to be a hard decision. Ideally, if you wind up reducing data transfer fees enough, it doesn’t even have to be only one. There are stories that starts to turn into an actual realistic multi-cloud story that isn’t, at its face, ridiculous. But right now, you have to pick a horse and ride it, for a variety of reasons. And I don’t like that.

Matthew: It’s entirely egress-based. And again, I think that customers are better off if they are able to pick who is the best service at any time. And that is what encourages innovation. And over time, that’s even what’s good for the various cloud providers because it’s what keeps them being valuable and keeps their customers thinking that they’re building something which is magical and that they aren’t trapped in the decision that they made, which is when we talk to a lot of the customers today, they feel that way. And it’s I think part of why something like R2 and something like the Bandwidth Alliance has gotten so much attention because it really touches a nerve on what’s frustrating customers today. And if tomorrow Amazon announced that they were eliminating egress fees and going head-to-head with R2, again, I think that’s a wonderful outcome. And one that I think is unlikely, but I would celebrate it if it happened.

Corey: This episode is sponsored by our friends at Oracle Cloud. Counting the pennies, but still dreaming of deploying apps instead of "Hello, World" demos? Allow me to introduce you to Oracle's Always Free tier. It provides over 20 free services and infrastructure, networking databases, observability, management, and security.

And - let me be clear here - it's actually free. There's no surprise billing until you intentionally and proactively upgrade your account. This means you can provision a virtual machine instance or spin up an autonomous database that manages itself all while gaining the networking load, balancing and storage resources that somehow never quite make it into most free tiers needed to support the application that you want to build.

With Always Free you can do things like run small scale applications, or do proof of concept testing without spending a dime. You know that I always like to put asterisks next to the word free. This is actually free. No asterisk. Start now. Visit https://snark.cloud/oci-free that's https://snark.cloud/oci-free.

Corey: My favorite is people who don’t do research on this stuff. They wind up saying, “Oh, yeah. Cloudflare is saying that bandwidth is a fixed cost. Of course not. They must be losing their shirt on this.”

You are a publicly-traded company. Your gross margins are 76% or 77%, depending upon whether we’re talking about GAAP or non-GAAP. Point being, you are clearly not selling this at a loss and hoping to make it up in volume. That’s what a VC-backed company does. Is something that is real and as accurate.

I want to, on some level, I guess, low-key apologize because I keep viewing Cloudflare through a lens that is increasingly inaccurate, which is as a CDN. But you’ve had Cloudflare Workers for a while, effectively Functions as a Service that run at the edge, which has this magic aura around it, that do various things, which is fascinating to me. You’re launching R2; it feels like you are in some ways aiming at becoming a cloud provider, but instead of taking the traditional approach of building it from the region’s outward, you’re building it from the outward in. Is that a fair characterization?

Matthew: I think that’s right. I think fundamentally what Cloudflare is, is a network. And I remember early on in the pandemic, we did a series of fireside chats with people we thought we could learn from. And so was everyone from Andre Iguodala, the basketball player, to Mark Cuban, the entrepreneur, to we had a [unintelligible 00:25:56] governor and all kinds of things. And we these were just internal on off the record.

And I got to do one with Eric Schmidt, the former CEO of Google. And I said, “You know, Eric, one of the things that we struggle with is describing what is Cloudflare.” And without hesitation, he said, “Oh, that’s easy. You’re the network I plug into and don’t have to worry about anything else.” And I think that’s better than I could say it, myself, and I think that’s what it is that we fundamentally are: we’re the network that fits together.

Now, it turns out that in the process of being that network and enabling that network, we are going to build things like R2, which start to be an object store and starts to sort of step into some of the cloud provider space. And Workers is really just a way of programming that network in order to do that, but it turns out that there are a bunch of workloads that if you move them into the network itself, make sense—not going to be every workload, but a lot of workloads that makes sense there. And again, I think that you can actually be very bullish on all of the big public cloud providers and bullish on Cloudflare at the same time because what we want to do is enable the ability for people to mix and match, and change, and be the fabric that connects all of those things together. And so over time, if Amazon says, “We’re going to drop egress fees,” it may be that R2 isn’t a product that exists—I don’t think they’re going to do that, so I think it’s something that is going to be successful for us and get a lot of new users to us—but fundamentally, I think that where the traditional public clouds think of themselves as the place you put data and you process data, I think we think of ourselves as the place you move data. And that’s somewhat different.

That then translates into it as we’re building out the different pieces, where it does feel like we’re building from the outside in. And it may be that over time, that put versus move distinction becomes narrower and narrower as we build more and more services like R2, and durable objects, and KV, and we’re working on a database, and all those things. And it could be that we converge in a similar place.

Corey: One thing I really appreciate about your vision because it is so atypical these days, is that you aren’t trying to build the multifunction printer of companies. You are not trying to be all things to all people in every scenario. Which is impossible to do, but companies are still trying their level best to do it. You are staking out the bounds of where you were willing to start and where you’re willing to stop, in a variety of different ways. I would be—how do I put it?—surprised if you at some point in the next five years come out with, “And this is our own database that we have built out that directly competes with the following open-source project that we basically have implemented their API and gone down that particular path.” It does not sound like it is in your core wheelhouse at that point. You don’t need—to my understanding—to write your own database engine in order to do what you do.

Matthew: Maybe. I mean, we actually are kind of working on a database because—

Corey: Oh, no, here we go again.

Matthew: [laugh]—and yeah—in a couple of different ways. So, the first way is, we want to make sure that if you’re using Workers, you can connect to whatever database you want to use anywhere in the world. And that’s something that’s coming and we’ll be there. At the same time, the challenge of distributed computing turns out not to be the computing, it turns out to be the data and figuring out how to—CAP theorem is real, right? Consistency, Availability, and Partition tolerance; you can pick any two out of the three, but you can’t get all three.

And so you there’s always going to be some trade-off that’s there. And so we don’t see a lot of good examples. There’s some really cool companies that are working on things in the space, but we don’t see a lot of really good examples of who has built a database that can be run on a distributed workload system, like Cloudflare to it do well. And so our team internally needs that, and so we’re trying to figure out how to build it for ourselves, and I would imagine that after we build it for ourselves—if it works the way we expect it will—that that will then be something that we open up.

Our motivation and the way we think about products is we need to build the tools for our own team. Our team itself is customer zero, and then some of those things are very specific to us, but every once in a while, when there are functions that makes sense for others, then we’ll build them as well. And that does maybe risk being the multifunction printer, but again, I think that because the customer for that starts with ourselves, that’s how we think about it. And if there’s someone else’s making a great tool, we’ll use that. But in this case, we don’t see anyone that’s built a multi-tenant, globally-distributed, ACID-compliant relational database.

Corey: I can’t let it pass on challenge. Sure they have, and you’re running it yourself. DNS: the finest database in the world. You stuff whatever you want to text records, and now you have taken a finely crafted wrench and turned it into a barely acceptable hammer, which is what I love about doing that terrible approach. Yeah, relational is not going to quite work that way. But—

Matthew: Yes. That’s a fancy key-value store, right? So—and we’ve had that for a long time. As we’re trying to build those things up, the good news is that, again, we’ve run data at scale for quite some time and proven that we can do it efficiently and reliably.

Corey: There’s a lot that can be said about building the things you need to deliver your product to customers. And maybe a database is a poor example here, but I don’t see that your motivation in this space is to step into something completely outside your areas of expertise solely because there’s money to be made over there. Well, yeah, fortune passes everywhere. The question is, which are you best positioned to wind up delivering an actual transformative solution to that space, and what parts of it are just rent-seeking where it’s okay, we’re going to go and wherever the money is, we’re chasing that down.

Matthew: Yeah, we’re still a for-profit business, and we’ve been able to grow revenue well, but I think it is that what motivates us and what drives us comes back to our mission, which is how do you help build a better internet? And you can look at every single thing that we’ve done, and we try to be very long-term-oriented. So, for instance, when we in 2014 made encryption free, the number one reason at the time, when people upgraded for the free version of our service, the paid version of our service is they got encryption for that. And so it was super scary to say, “Hey, we’re going to take the biggest feature and give it away for free,” but it was clearly the direction of history and we wanted to be on the right side of history. And we considered it a bug that the internet wasn’t built in an encrypted way from the beginning.

So, of course, that was going to head that direction. And so I think that we and then subsequently Let’s Encrypt, and a bunch of others have said, it’s absurd that you’re charging for math. And again, I think that’s a good example of how we think about products. And we want to continue to disrupt ourselves and take the things that once upon a time were reserved for our customers that spend $10 million-plus with us, and we want to keep pushing those things down because, over time, the real opportunity is if you do right by customers, there will be plenty of ways that you can earn some of their budget. And again, we think that is the long-term winning strategy.

Corey: I would agree with this. You’re not out there making sneakers and selling them because you see people spend a lot of money on that; you’re delivering value for customers. I say this as one of your paying customers. I have zero problem paying you every month like clockwork, and it is the least cloud-like experience because I know exactly what the bill is going to be in advance, which is apparently not how things should be done in this industry, yadda, yadda, yadda. It is a refreshingly delightful experience every time.

The few times I’ve had challenges with the service, it has almost always been a—I’ll call it a documentation gap, where the way it was explained in the formal documentation was not how I conceptualize things, which, again, explaining what these complex things are to folks who are not steeped in certain areas of them is always going to be a challenge. But I cannot think back to a single customer service failure I’ve had with you folks. I can’t look back at any point where you have failed me as a customer, which is a strange thing to say, given how incredibly efficient I am at stumbling over weird bugs.

Matthew: Terrific to have you as a customer. We are hardly perfect and we make mistakes, but one of the things I think that we try to do and one of the core values of Cloudflare is transparency. If I think about, like, the original sins of tech, a lot of it is this bizarre secrecy which pervades the entire industry. When we make mistakes, we talk about them, and we explain them. When there’s an error, we don’t throw up a white page; we put up a page that has our logo on it because we want to own it.

And that sometimes gets blowback because you’re in front of it, but again, I think it’s the right thing to do for customers. And it’s and I think it’s incredibly important. One of the things that’s interesting is you mentioned that you know what your bill is going to be. If you go back and look at the history of hosting on the internet, in the early days of internet hosting, it looks a lot like AWS.

Corey: Oh, 95th percentile transit billing; go for one five minutes segment over and boom, your bill explodes. Oh, I remember those days. Unkindly.

Matthew: And it was super complicated. And then what happened is the hosting world switched from this incredibly complicated billing to much more simplified, predictable, unlimited bandwidth with maybe some asterisks, but largely that was in place. And then it’s strange that Amazon came along and then has brought us back to the more complicated world that’s out there. I would have predicted that that’s a sine wave—

Corey: It has to be. I mean—

Matthew: —and it’s going to go back and forth over time. But I would have predicted that we would be more in the direction of coming back toward simplify, everything included. And again, I think that’s how we’ve priced our things from the beginning. I’m surprised that it has held on as long as it has, but I do think that there’s going to be an opportunity for—and I don’t think Amazon will be the leader here, but I think there will be an opportunity for one of the big clouds.

And again, I think Oracle is probably doing this the best of any of them right now—to say, “How can we go away from that complexity? How can we make bills predictable? How can we not nickel and dime everything, but allow you to actually forecast and budget?” And it just seems like that’s the natural arc of history, and we will head back toward that. And, again, I think we’ve done our part to push that along. And I’m excited that other cloud providers seem to be thinking about that now as well.

Corey: Oh, yeah. What I do with fixing AWS bills is the same thing folks were doing in the 70s and 80s with long-distance bills for companies. We’re definitely hitting that sine wave. I know that if I were at AWS in a leadership role, I would be actively embarrassed that the company that is delivering a better customer experience around financial things is Oracle of all companies, given their history of audits and surprising people and the rest. It is ridiculous to me.

One last topic that I want to cover with you before we call it an episode is, back in college, you had a thesis that you have done an excellent job of effectively eliminating from the internet. And the theme of this, to my understanding, was that the internet is a fad. And I am so aligned with that because I’m someone who has said for years that emerging technologies are fads. I’ve said it about cloud, about virtualization, about containers. And I just skipped Kubernetes. And now I’m all-in on serverless, which means, of course it’s going to fail because I’m always wrong on these things. But tell me about that.

Matthew: When I was seven years old in 1980, my grandmother gave me an Apple ][+ computer for Christmas. And I took to it like a just absolute duck to water and did things that made me very popular in junior high school, like going to computer camp. And my mom used to sign up for continuing education classes at the local university in computer science, and basically sneak me in, and I’d do all the homework and all that. And I remember when I got to college, there was a small group of students that would come around and help other students set their computer up, and I had it all set up and was involved. And so, got pretty deeply involved in the computer science program at college.

And then I remember there was a group of three other students—so they were four of us—and they wanted to start an online digital magazine. And at the time, this was pre-web, or right in the early days of the web; it was sort of nineteen… ninety-three. And we built it originally on old Apple technology called HyperCard. And we used to email out the old HyperCard stacks. And the HyperCard stacks kept getting bigger and bigger and bigger, and we’d send them out to the school so [laugh] that we—so we kept crashing the mail servers.

But the college loved this, so they kept buying bigger and bigger mail servers. But they were—at some point, they said, “This won’t scale. You got to switch technologies.” And they introduced us to two different groups. One was a printer company based out in San Francisco that had this technology called PDF. And I was a really big fan of PDF. I thought PDF was the future, it was definitely going to be how everything got published.

And then the other was this group of dorky graduate students at the University of Illinois that had this thing called a browser, which was super flaky, and crashed all the time, and didn’t work. And so of the four of us, I was the one who voted for PDF and the other three were like, “Actually, I think this HTML thing is going to be a hit.” And we built this. We won an award from Wired—which was only a print magazine at the time—that called us the first online-only weekly publication. And it was such a struggle to get anyone to write for it because browsers sucked and, you know, trying to get students on campus, but no one on campus cared.

We would get these emails from the other side of the world, where I remember really clearly is this—in broken English—email from Japan saying, “I love the magazine. Please keep writing more for the magazine.” And I remember thinking at the time, “Why do I care if someone in Japan is reading this if the girl down the hall who I have a crush on isn’t?” Which is obviously what motivates dorky college students like myself. And at that same time, you saw all of this internet explosion.

I remember the moment when Netscape went public and just blew through all the expectations. And it was right around the time I was getting ready to graduate for college, and I was kind of just burned out on the entire thing. And I thought, “If I can’t even get anyone to write for this dopey magazine and yet we’re winning awards, like, this stuff has to all just be complete garbage.” And so wrote a thesis on—ehh, it was not a very good [laugh] thesis. It’s—but one of the things I said was that largely the internet was a fad, and that if it wasn’t, that it had some real risks because if you enabled everyone to connect with whatever their weird interests and hobbies were, that you would very quickly fall to the lowest common denominator. And predicted some things that haven’t come true. I thought for sure that you would have both a liberal and conservative search engine. And it’s a miracle to this day, I think that doesn’t exist.

Corey: Now, that you said it, of course, it’s going to.

Matthew: Well, I don’t know I’ve… [sigh] we’ll see. But it is pretty amazing that Google has been able to, again, thread that line and stay largely apolitical. I’m surprised there aren’t more national search engines; the fact that it only Russia and China have national search engines and France and Germany don’t is just strange to me. It seems like if you’re controlling the source of truth and how people find it, that seems like something that governments would try and take over. There are some things that in retrospect, look pretty wise, but there were a lot more things that looked really, really stupid. And so I think at some level, I had to build Cloudflare to atone for that stupidity all those years ago.

Corey: There’s something to be said for looking back and saying, “Yeah, I had an opinion, and with the light of new information, I am changing my opinion.” For some reason, in some circles, it feels like that gets interpreted as a sign of weakness, but I couldn’t disagree more, it’s, “Well, I had an opinion based upon what I saw at the time. Turns out, I was wrong, and here we are.” I really wish more people were capable of doing that.

Matthew: It’s one of the things we test for in hiring. And I think the characteristic that describes people who can do that well is really empathy. The understanding that the experiences that you have lead you to have a unique set of insights, but they also create a unique set of blind spots. And it’s rare that you find people that are able to do that. And whenever you do—whenever we do we hire them.

Corey: To that end, as far as hiring and similar topics go, if people want to learn more about how you view things, and how you see the world, and what you’re releasing—maybe even potentially work with you—where can they find you?

Matthew: [laugh]. So, the joke, sometimes, internal at Cloudflare is that Cloudflare is a blogging company that runs this global network just to have something to write about. So, I think we’re unlike most corporate blogs, which are—if our corporate blog were typical, we’d have articles on, like, “Here are the top six reasons you need a fast website,” which would just be, you know, shoot me. But instead, I think we write about the things that are going on online and our unique view into them. And we have a core value of transparency, so we talk about that. So, if you’re interested in Cloudflare, I’d encourage you to—especially if you’re of the sort of geekier variety—to check out blog.cloudflare.com, and I think that’s a good place to learn about us. And I still write for that occasionally.

Corey: You’re one of the only non-AWS corporate blogs that I pay attention to, for that exact reason. It is not, “Oh, yay. More content marketing by folks who just feel the need to hit a quota as opposed to talking about something valuable and interesting.” So, it’s appreciated.

Matthew: The secret to it was we realized at some point that the purpose of the blog wasn’t to attract customers, it was to attract potential employees. And it turns out, if you sort of change that focus, then you talk to people like their peers, and it turns out then that the content that you create is much more authentic. And that turns out to be a great way to attract customers as well.

Corey: I want to thank you for taking so much time out of your day to speak with me. I really appreciate it.

Matthew: Thanks for all you’re doing. And we’re very aligned, and keep fighting the good fight. And someday, again, we’ll eliminate cloud egress fees, and we can share a beer when we do.

Corey: I will absolutely be there for it. Matthew, Prince, CEO, and co-founder of Cloudflare. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with a rambling comment explaining that while data packets into a cloud provider are cheap and crappy, the ones being sent to the internet are beautiful, bespoke, unicorn snowflakes, so of course they cost money.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Richard

He’s also an instructor at Pluralsight, a frequent public speaker, and the author of multiple books on software design and development. Richard maintains a regularly updated blog (seroter.com) on topics of architecture and solution design and can be found on Twitter as @rseroter.

Links:

  • Twitter: https://twitter.com/rseroter
  • LinkedIn: https://www.linkedin.com/in/seroter
  • Seroter.com: https://seroter.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Vultr. Spelled V-U-L-T-R because they’re all about helping save money, including on things like, you know, vowels. So, what they do is they are a cloud provider that provides surprisingly high performance cloud compute at a price that—while sure they claim its better than AWS pricing—and when they say that they mean it is less money. Sure, I don’t dispute that but what I find interesting is that it’s predictable. They tell you in advance on a monthly basis what it’s going to going to cost. They have a bunch of advanced networking features. They have nineteen global locations and scale things elastically. Not to be confused with openly, because apparently elastic and open can mean the same thing sometimes. They have had over a million users. Deployments take less that sixty seconds across twelve pre-selected operating systems. Or, if you’re one of those nutters like me, you can bring your own ISO and install basically any operating system you want. Starting with pricing as low as $2.50 a month for Vultr cloud compute they have plans for developers and businesses of all sizes, except maybe Amazon, who stubbornly insists on having something to scale all on their own. Try Vultr today for free by visiting: vultr.com/screaming, and you’ll receive a $100 in credit. Thats v-u-l-t-r.com slash screaming.

Corey: You know how git works right?

Announcer: Sorta, kinda, not really Please ask someone else!

Corey: Thats all of us. Git is how we build things, and Netlify is one of the best way I’ve found to build those things quickly for the web. Netlify’s git based workflows mean you don't have to play slap and tickle with integrating arcane non-sense and web hooks, which are themselves about as well understood as git. Give them a try and see what folks ranging from my fake Twitter for pets startup, to global fortune 2000 companies are raving about. If you end up talking to them, because you don't have to, they get why self service is important—but if you do, be sure to tell them that I sent you and watch all of the blood drain from their faces instantly. You can find them in the AWS marketplace or at www.netlify.com. N-E-T-L-I-F-Y.com

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Once upon a time back in the days of VH1, which was like MTV except it played music videos, would have a show that was, “Where are they now?” Looking at former celebrities. I will not use the term washed up because that’s going to be insulting to my guest.

Richard Seroter is a returning guest here on Screaming in the Cloud. We spoke to him a year ago when he was brand new in his role at Google as director of outbound product management. At that point, he basically had stars in his eyes and was aspirational around everything he wanted to achieve. And now it’s a year later and he has clearly failed because it’s Google. So, outbound products are clearly the things that they are going to be deprecating, and in the past year, I am unaware of a single Google Cloud product that has been outright deprecated. Richard, thank you for joining me, and what do you have to say for yourself?

Richard: Yeah, “Where are they now?” I feel like I’m the Leif Garrett of cloud here, joining you. So yes, I’m still here, I’m still alive. A little grayer after twelve months in, but happy to be here chatting cloud, chatting whatever else with you.

Corey: I joke a little bit about, “Oh, Google winds up killing things.” And let’s be clear, your consumer division which, you know, Google is prone to that. And understanding a company’s org chart is a challenge. A year or two ago, I was of the opinion that I didn’t need to know anything about Google Cloud because it would probably be deprecated before I really had to know about it. My opinion has evolved considerably based upon a number of things I’m seeing from Google.

Let’s be clear here, I’m not saying this to shine you on or anything like that; it’s instead that I’ve seen some interesting things coming out of Google that I consider to be the right moves. One example of that is publicly signing multiple ten-year deals with very large, serious institutions like Deutsche Bank, and others. Okay, you don’t generally sign contracts with companies of that scale and intend not to live up to them. You’re hiring Forrest Brazeal as your head of content for Google Cloud, which is not something you should do lightly, and not something that is a short-term play in any respect. And the customer experience has continued to improve; Google Cloud products have not gotten worse, and I’m seeing in my own customer conversations that discussions about Google Cloud have become significantly less dismissive than they were over the past year. Please go ahead and claim credit for all of that.

Richard: Yeah. I mean, the changes a year ago when I joined. So, Thomas Kurian has made a huge impact on some of that. You saw us launch the enterprise APIs thing a while back, which was, “Hey, here’s, for the most part, every one of our products that has a fixed API. We’re not going to deprecate it without a year’s notice, whatever it is. We’re not going to make certain types of changes.” Maybe that feels like, “Well, you should have had that before.” All right, all we can do is improve things moving forward. So, I think that was a good change.

Corey: Oh, I agree. I think that was a great thing to do. You had something like 80-some-odd percent coverage of Google Cloud services, and great, that’s going to only increase with time, I can imagine. But I got a little pushback from a few Googlers for not being more congratulatory towards them for doing this, and look, it’s a great thing. Don’t get me wrong, but you don’t exactly get a whole lot of bonus points and kudos and positive press coverage—not that I’m press—for doing the thing you should have been doing [laugh] all along.

It’s, “This is great. This is necessary.” And it demonstrates a clear awareness that there was—rightly or wrongly—a perception issue around the platform’s longevity and that you’ve gone significantly out of your way to wind up addressing that in ways that go far beyond just yelling at people on Twitter they don’t understand the true philosophy of Google Cloud, which is the right thing to do.

Richard: Yeah, I mean, as you mentioned, look, the consumer side is very experimental in a lot of cases. I still mourn Google Reader. Like, those things don’t matter—

Corey: As do we all.

Richard: Of course. So, I get that. Google Cloud—and of course we have the same cultural thing, but at the same time, there’s a lifecycle management that’s different in Google Cloud. We do not deprecate products that much. You know, enterprises make decade-long bets. I can’t be swap—changing databases or just turning off messaging things. Instead, we’re building a core set of things and making them better.

So, I like the fact that we have a pretty stable portfolio that keeps getting a little bit bigger. Not crazy bigger; I like that we’re not just throwing everything out there saying, “Rock on.” We have some opinions. But I think that’s been a positive trend, customers seem to like that we’re making these long-term bets. We’re not going anywhere for a long time and our earnings quarter after quarter shows it—boy, this will actually be a profitable business pretty soon.

Corey: Oh, yeah. People love to make hay, and by people, I stretch the term slightly and talk about, “Investment analysts say that Google Cloud is terrible because at your last annual report you’re losing something like $5 billion a year on Google Cloud.” And everyone looked at me strangely, when I said, “No, this is terrific. What that means is that they’re investing in the platform.” Because let’s be clear, folks at Google tend to be intelligent, by and large, or at least intelligent enough that they’re not going to start selling cloud services for less than it costs to run them.

So yeah, it is clearly an investment in the platform and growth of it. The only way it should be turning a profit at this point is if there’s no more room to invest that money back into growing the platform, given your market position. I think that’s a terrific thing, and I’m not worried at all about it losing money. I don’t think anyone should be.

Richard: Yeah, I mean, strategically, look, this doesn’t have to be the same type of moneymaker that even some other clouds have to be to their portfolio. Look, this is an important part, but you look at those ten-year deals that we’ve been signing: when you look at Univision, that’s a YouTube partnership; you look at Ford that had to do with Android Auto; you look at these others, this is where us being also a consumer and enterprise SaaS company is interesting because this isn’t just who’s cranking out the best IaaS. I mean, that can be boring stuff over time. It’s like, who’s actually doing the stuff that maybe makes a traditional company more interesting because they partner on some of those SaaS services. So, those are the sorts of deals and those sorts of arrangements where cloud needs to be awesome, and successful, and make money, doesn’t need to be the biggest revenue generator for Google.

Corey: So, when we first started talking, you were newly minted as a director of outbound product management. And now, you are not the only one, there are apparently 60 of you there, and I’m no closer to understanding what the role encompasses. What is your remit? Where do you start? Where do you stop?

Richard: Yeah, that’s a good question. So, there’s outbound product management teams, mostly associated with the portfolio area. So network, storage, AI, analytics, database, compute, application modernization-y sort of stuff—which is what I cover—containers, dev tools, serverless. Basically, I am helping make sure the market understands the product and the product understands the market. And not to be totally glib, but a lot of that is, we are amplification.

I’m amplifying product out to market, analysts, field people, partners: “Do you understand this thing? Can I help you put this in context?” But then really importantly, I’m trying to help make sure we’re also amplifying the market back to our product teams. You’re getting real customer feedback: “Do you know what that analyst thinks? Have you heard what happened in the competitive space?”

And so sometimes companies seem to miss that, and PMs poke their head up when I’m about to plan a product or I’m about to launch a product because I need some feedback. But keeping that constant pulse on the market, on customers, on what’s going on, I think that can be a secret weapon. I’m not sure everybody does that.

Corey: Spending as much time as I do on bills, admittedly AWS bills, but this is a pattern that tends to unfold across every provider I’ve seen. The keynotes are chock-full of awesome managed service announcements, things that are effectively turnkey at further up the stack levels, but the bills invariably look a lot more like, yeah, we spend a bit of money on that and then we run 10,000 virtual instances in a particular environment and we just treat it like it’s an extension of our data center. And that’s not exciting; that’s not fun, quote-unquote, but it’s absolutely what customers are doing and I’m not going to sit here and tell them that they’re wrong for doing it. That is the hallmark of a terrible consultant of, “I don’t understand why you’re doing what you’re doing, so it must be foolish.” How about you stop and gain some context into
why customers do the things that they do?

Richard: No, I send around a goofy newsletter every week to a thousand or two people, just on things I’m learning from the field, from customers, trying to make sure we’re just thinking bigger. A couple of weeks ago, I wrote an idea about modernization is awesome, and I love when people upgrade their software. By the way, most people migration is a heck of a lot easier than if I can just get this into your cloud, yeah love that; that’s not the most interesting thing, to move VMs around, but most people in their budget, don’t have time to rewrite every Java app to go. Everybody’s not changing .NET framework to .NET core.

Like, who do I think everybody is? No, I just need to try to get some incremental value first. Yes, then hopefully I’ll swap out my self-managed SQL database for a Spanner or a managed service. Of course, I want all of that, but this idea that I can turn my line of business loan processing app into a thousand functions overnight is goofy. So, how are we instead thinking more pragmatically about migration, and then modernizing some of it? But even that sort of mindset, look, Google thinks about innovation modernization first. So, also just trying to help us take a step back and go, “Gosh, what is the normal path? Well, it’s a lot of migration first, some modernization, and then there’s some steady-state work there.”

Corey: One of the things that surprised me the most about Google Cloud in the market, across the board, has been the enthusiastic uptake for enterprise workloads. And by enterprise workloads, I’m talking about things like SAP HANA is doing a whole bunch of deployments there; we’re talking Big Iron-style enterprise-y things that, let’s be honest, countervene most of the philosophy that Google has always held and espoused publicly, at least on conference stages, about how software should be built. And I thought that would cut against them and make it very difficult for you folks to gain headway in that market and I could not have been more wrong. I’m talking to large enterprises who are enthusiastically talking about Google Cloud. I’ve got a level with you, compared to a year or two ago, I don’t recognize the place.

Richard: Mmm. I mean, some of that, honestly, in the conversations I have, and whatever I do a handful of customer calls every week, I think folks still want something familiar, but you’re looking for maybe a further step on some of it. And that means, like, yes, is everybody going to offer VMs? Yeah, of course. Is everyone going to have MySQL? Obviously.

But if I’m an enterprise and I’m doing these generational bets, can I cheat a little bit, and maybe if I partner with a more of an innovation partner versus maybe just the easy next step, am I buying some more relevance for the long-term? So, am I getting into environment that has some really cool native zero-trust stuff? Am I getting into environment with global backend services and I’m not just stitching together a bunch of regional stuff? How can I cheat by using a more innovation vendor versus just lifting and shifting to what feels like hosted software in another cloud? I’m seeing more of that because these migrations are tough; nobody should be just randomly switching clouds. That’s insane.

So, can I make, maybe, one of these big bets with somebody who feels like they might actually even improve my business as a whole because I can work with Google Pay and improve how I do mobile payments, or I could do something here with Android? Or, heck, all my developers are using Angular and Flutter; aren’t I going to get some benefit from working with Google? So, we’re seeing that, kind of, add-on effect of, “Maybe this is a place not just to host my VMs, but to take a generational leap.”

Corey: And I think that you’re positioning yourselves in a way to do it. Again, talk about things that you wouldn’t have expected to come out of Google of all places, but your console experience has been first-rate and has been for a while. The developer experience is awesome; I don’t need to learn the intricacies of 12 different services for what I’m trying to do just in order to get something basic up and running. I can stop all the random little billing things in my experimental project with a single click, which that admittedly has a confirm, which you kind of want. But it lets you reason about these things.

It lets you get started building something, and there’s a consistency and cohesiveness to the console that, again, I am not a graphic designer, by any stretch of the imagination. My most commonly used user interface is a green-screen shell prompt, and then I’m using Vim to wind up writing something horrifying, ideally in Python, but more often in YAML. And that has been my experience, but just clicking around the console, it’s clear that there was significant thought put into the design, the user experience, and the way of approaching folks who are starting to look very different, from a user persona perspective.

Richard: I can—I mean, I love our user research team; they’re actually fun to hang out with and watch what they do, but you have to remember, Google as a company, I don’t know, cloud is the first thing we had to sell. Did have to sell Gmail. I remember 15 years ago, people were waiting for invites. And who buys Maps or who buys YouTube? For the most part, we’ve had to build things that were naturally interesting and easy-to-use because otherwise, you would just switch to anything else because everything was free.

So, some of that does infuse Google Cloud, “Let’s just make this really easy to use. And let’s just make sure that, maybe, you don’t hate yourself when you’re done jumping into a shell from the middle of the console.” It’s like, that should be really easy to do—or upgrade a database, or make changes to things. So, I think some of the things we’ve learned from the consumer good side, have made their way to how we think of UX and design because maybe this stuff shouldn’t be terrible.

Corey: There’s a trope going around, where I wound up talking about the next million cloud customers. And I’m going to have to write a sequel to it because it turns out that I’ve made a fundamental error, in that I’ve accepted the narrative that all of the large cloud vendors are pushing, to the point where I heard from so many folks I just accepted it unthinkingly and uncritically, and that’s not what I should be doing. And we’ll get to what I was wrong about in a minute, but the thinking goes that the next big growth area is large enterprises, specifically around corporate IT. And those are folks who are used to managing things in a GUI environment—which is fine—and clicking around in web apps. Now, it’s easy to sit here on our high horse and say, “Oh, you should learn to write code,” or YAML, which is basically code. Cool.

As an individual, I agree, someone should because as soon as they do that, they are now able to go out and take that skill to a more lucrative role. The company then has to backfill someone into the role that they just got promoted out of, and the company still has that dependency. And you cannot succeed in that market with a philosophy of, “Oh, you built something in the console. Now, throw it away and do it right.” Because that is maddening to that user persona. Rightfully so.

I’m not that user persona and I find it maddening when I have to keep tripping over that particular thing. How did that come to be, from your perspective? First, do you think that is where the next million cloud customers come from? And have I adequately captured that user persona, or am I completely often the weeds somewhere?

Richard: I mean, I shared your post internally when that one came out because that resonated with me of how we were thinking about it. Again, it’s easy to think about the cloud-native operators, it’s Spotify doing something amazing, or this team at Twitter doing something, or whatever. And it’s not even to be disparaging. Like, look, I spent five years in enterprise IT and I was surrounded by operators who had to run
dozen different systems; they weren’t dedicated to just this thing or that. So, what are the tools that make my life easy?

A lot of software just comes with UIs for quick install and upgrades, and how does that logic translate to this cloud world? I think that stuff does matter. How are you meeting these people a little better where they are? I think the hard part that we will always have in every cloud provider is—I think you’ve said this in different forums, but how do I not sometimes rub the data center on my cloud or vice versa? I also don’t want to change the experience so much where I degrade it over the long term, I’ve actually somehow done something worse.

So, can I meet those people where they are? Can we pull some of those experiences in, but not accidentally do something that kind of messes up the cloud experience? I mean, that’s a fine line to walk. Does that make sense to you? Do you see where there’s a… I don’t know, you could accidentally cater to a certain audience too much, and change the experience for the worse?

Corey: Yes, and no. My philosophy on it is that you have to meet customers where they are, but only to a point. At some point, what they’re asking for becomes actively harmful or disadvantageous to wind up providing for them. “I want you to run my data center for me,” is on some level what some cloud environments look like, and I’m not going to sit here and tell people they’re inherently wrong for that. Their big reason for moving to the cloud was because they keep screwing up replacing failed hard drives in their data center, so we’re going to put it in the cloud.

Is it more expensive that way? Well, sure in terms of actual cash outlay, it almost certainly is, but they’re also not going down every month when a drive fails, so once the value of that? It’s a capability story. That becomes interesting to me, and I think that trying to sit here in isolation, and say that, “Oh, this application is not how we would build it at Google.” And it’s, “Yeah, you’re Google. They are insert an entire universe of different industries that look nothing whatsoever like Google.” The constraints are different, the resources are different, and—

Richard: Sure.

Corey: —their approach to problem-solving are different. When you built out Google, and even when you’re building out Google Cloud, look at some of the oldest craftiest stuff you have in your entire all of Google environment, and then remember that there are companies out there that are hundreds of years old. It’s a different order of magnitude as far as era, as far as understanding of what’s in the environment, and that’s okay. It’s a very broad and very diverse world.

Richard: Yeah. I mean, that’s, again, why I’ve been thinking more about migration than even some of the modernization piece. Should you bring your network architecture from on-prem to the cloud? I mean, I think most cases, no. But I understand sometimes that edge firewall, internal trust model you had on-prem, okay, trying to replicate that.

So, yeah, like you say, I want to meet people where they are. Can we at least find some strategic leverage points to upgrade aspects of things as you get to a cloud, to save you from yourself in some places because all of a sudden, you have ten regions and you only had one data center before. So, many more rooms for mistakes. Where are the right guardrails? We’re probably more opinionated than others at Google
Cloud.

I don’t really apologize for that completely, but I understand. I mean, I think we’ve loosened up a lot more than maybe people [laugh] would have thought a few years ago, from being hyper-opinionated on how you run software.

Corey: I will actually push back a bit on the idea that you should not replicate your on-premises data center in your cloud environment. Sure, are there more optimal ways to do it that are arguably more secure? Absolutely. But a common failure mode in moving from data center to cloud is, “All right, we’re going to start embracing this entirely new cloud networking paradigm.” And it is confusing, and your team that knows how the data center network works really well are suddenly in way over their heads, and they’re inadvertently exposing things they don’t intend to or causing issues.

The hard part is always people, not technology. So, when I glance at an environment and see things like that, perfect example, are there more optimal ways to do it? Oh, from a technology perspective, absolutely. How many engineers are working on that? What’s their skill set? What’s their position on all this? What else are they working on? Because you’re never going to find a team of folks who are world-class experts in every cloud? It doesn’t work that way.

Richard: No doubt. No doubt, you’re right. There’s areas where we have to at least have something that’s going to look similar, let you replicate aspects of it. I think it’s—it’ll just be interesting to watch, and I have enough conversations with customers who do ask, “Hey, where are the places we should make certain changes as we evolve?” And maybe they are tactical, and they’re not going to be the big strategic redesign their entire thing. But it is good to see people not just trying to shovel everything from one place to the next.

Corey: This episode is sponsored in part by something new. Cloud Academy is a training platform built on two primary goals. Having the highest quality content in tech and cloud skills, and building a good community the is rich and full of IT and engineering professionals. You wouldn’t think those things go together, but sometimes they do. Its both useful for individuals and large enterprises, but here's what makes it new. I don’t use that term lightly. Cloud Academy invites you to showcase just how good your AWS skills are. For the next four weeks you’ll have a chance to prove yourself. Compete in four unique lab challenges, where they’ll be awarding more than $2000 in cash and prizes. I’m
not kidding, first place is a thousand bucks. Pre-register for the first challenge now, one that I picked out myself on Amazon SNS image
resizing, by visiting cloudacademy.com/corey. C-O-R-E-Y. That’s cloudacademy.com/corey. We’re gonna have some fun with this one!

Corey: Now, to follow up on what I was saying earlier, what I think I’ve gotten wrong by accepting the industry talking points on is that the next million cloud customers are big enterprises moving from data centers into the cloud. There’s money there, don’t get me wrong, but there is a larger opportunity in empowering the creation of companies in your environment. And this is what certain large competitors of yours get very wrong, where it’s we’re going to launch a whole bunch of different services that you get to build yourself from popsicle sticks. Great. That is not useful.

But companies that are trying to do interesting things, or people who want to found companies to do interesting things, want something that looks a lot more turnkey. If you are going to be building cloud offerings, that for example, are terrific building blocks for SaaS companies, then it behooves you to do actual investments, rather than just a generic credit offer, into spurring the creation of those types of companies. If you want to build a company that does payroll systems, in a SaaS, cloud way, “Partner with us. Do it here. We will give you a bunch of credits. We will introduce you to your first ten prospective customers.”

And effectively actually invest in a company success, as opposed to pitch-deck invest, which is, “Yeah, we’ll give you some discounting and some credits, and that’s our quote-unquote, ‘investment.’” actually be there with them as a partner. And that’s going to take years for folks to wrap their heads around, but I feel like that is the opportunity that is significantly larger, even than the embedded existing IT space because rather than fighting each other for slices of the pie, I’m much more interested in expanding that pie overall. One of my favorite questions to get asked because I think it is so profoundly missing the point is, “Do you think it’s possible for Google to go from number three to number two,” or whatever the number happens to be at some point, and my honest, considered answer is, “Who gives a shit?” Because number three, or number five, or number twelve—it doesn’t matter to me—is still how many hundreds of billions of dollars in the fullness of time. Let’s be real for a minute here; the total addressable market is expanding faster than any cloud or clouds are going to be able to capture all of.

Richard: Yeah. Hey, look, whoever who’ll be more profitable solving user problems, I really don’t care about the final revenue number. I can be the number one cloud tomorrow by making Google Cloud free. What’s the point? That’s not a sustainable business. So, if you’re just going for who can deploy the most VCPUs or who can deploy the most whatever, there’s ways to game that. I want to make sure we are just uniquely solving problems better than anybody else.

Corey: Sorry, forgive me. I just sort of zoned out for a second there because I’m just so taken aback and shocked by the idea of someone working at a large cloud provider who expresses a philosophy that isn’t lying awake at night fretting over the possibility of someone who isn’t them as making money somewhere.

Richard: [laugh]. I mean, your idea there, it’ll be interesting to watch, kind of, the maker’s approach of are you enabling that next round of startups, the next round of people who want to take—I mean, honestly, I like the things we’re doing building block-wise, even with our AI: we’re not just handing you a vision API, we’re giving you a loan processing AI that can process certain types of docs, that more packaged version of AI. Same with healthcare, same with whatever. I can imagine certain startups or a company idea going, “Hey, maybe I could disrupt or serve a new market.”

I always love what Square did. They’ve disrupted emerging markets, small merchants here in North America, wherever, where I didn’t need a big expensive point of sale system. You just gave me the nice, right building blocks to disrupt and run my business. Maybe Google Cloud can continue to provide better building blocks, but I do like your idea of actually investment zones, getting part of this. Maybe the next million
users are founders and it’s not just getting into some of these companies with, frankly, 10, 20, 30,000 people in IT.

I think there’s still plenty of room in these big enterprises to unlock many more of those companies, much more of their business. But to your point, there’s a giant market here that we’re not all grabbing yet. For crying out loud, there’s tons of opportunity out here. This is not zero-sum.

Corey: Take it a step further beyond that, and today, if you have someone who’s enterprising, early on in their career, maybe they just got out of school, maybe they have just left their job and are ready to snap, or they have some severance money that they want to throw into something. Great. What do they want to do if they have an idea for a company? Well today, that answer looks a lot like, well, time to go to a boot camp and learn to code for six months so you can build a badly done MVP well enough to get off the ground and get some outside investment, and then go from there. Well, what if we cut that part out entirely?

What if there were building blocks of I don’t need to know or care that there’s a database behind it, or what a database looks like. Picture Visual Basic in a web browser for building apps, and just take this bit of information I give you and store it and give it back to me later. Sure, you’re going to have some significant challenges in the architecture or something like that as it goes from this thing that I’m talking about as an MVP to something planet-scale—like a Spotify for example—but that’s not most businesses, and that’s okay. Get out of the way and let people innovate and iterate on what it is they’re doing more rapidly, and make it more accessible to teach people. That becomes huge; that gets the infrastructure bits that cloud providers excel at out of the way, and all it really takes is packaging those things into a golden path of what a given company of a particular profile should be doing, if—unless they have reason to deviate from it—and instead of having this giant paradox of choice issue, it’s, “Oh, okay, I’ll drag-drop, build things accordingly.”

And under the hood, it’s doing all the configuration of services and that’s great. But suddenly, you’ve made being a founder of a software company—fundamentally—accessible to people who are not themselves software engineers. And I know that’s anathema to some people, and I don’t even slightly care because I am done with gatekeeping.

Richard: Yeah. No, it’s exciting if that can pull off. I mean, it’s not the years ago where, how much capital was required to find the rack and do all sorts of things with tech, and hire some developers. And it’s an amazing time to be software creators, now. The more we can enable that—yeah, I’m along for that journey, sign me up.

Corey: I’m looking forward to seeing how it winds up shaking out. So, I want to talk a little bit about the paradox of choice problem that I just mentioned. If you take a look at the various compute services that every cloud provider offers, there are an awful lot of different choices as far as what you can run. There’s the VM model, there’s containers—if you’re in AWS, you have 17 ways to run those—and you wind up—any of the serverless function story, and other things here and there, and managed services, I mean and honestly, Google has a lot of them, nowhere near as many as you do failed messaging products, but still, an awful lot of compute options. How do customers decide?

What is the decision criteria that you see? Because the worst answer you can give someone who doesn’t really know what they’re doing is, “It depends,” because people don’t know how to make that decision. It’s, “What factors should I consider then, while making that decision?” And the answer has to be something somewhat authoritative because otherwise, they’re going to go on the internet and get yelled at by everyone because no one is ever going to agree on this, except that everyone else is wrong.

Richard: Mm-hm. Yeah, I mean, on one hand, look, I like that we intentionally have fewer choices than others because I don’t think you need 17 ways to run a container. I think that’s excessive. I think more than five is probably excessive because as a customer, what is the trade-off? Now, I would argue first off, I don’t care if you have a lot of options as a vendor, but boy, the backends of those better be consistent.

Meaning if I have a CI/CD tool in my portfolio and it only writes to two of them, shame on me. Then I should make sure that at least CI/CD, identity management, log management, monitoring, arguably your compute runtime should be a late-binding choice. And maybe that’s blasphemous because somebody says, “I want to start up front knowing it’s a function,” or, “I want to start it’s a VM.” How about, as a developer, I couldn’t care less. How about I just build cool software and maybe even at deploy time, I say, “This better fits in running in Kubernetes.” “This is better in a virtual machine.”

And my cost of changing that later is meaningless because, hey, if it is in the container, I can switch it between three or four different runtimes, the identity management the same, it logs the exact same way, I can deploy CI/CD the same way. So, first off, if those things aren’t the same, then the vendor is messing up. So, the customer shouldn’t have to pay the cost of that. And then there gets to be other actual criteria. Look, I think you are looking at the workload itself, the team who makes it, and the strategy to figure out the runtime.

It’s easy for us. Google Compute Engine for VMs, containers go in GKE, managed services that need some containers, there are some apps around them, are Cloud Functions and Cloud Run. Like, it’s fairly straightforward and it’s going to be an OR situation—or an AND situation not an OR, which is great. But we’re at least saying the premium way to run containers in Google Cloud for systems is GKE. There you go. If you do have a bunch of managed services in your architecture and you’re stitching them together, then you want more serverless things like Cloud Run and Cloud Functions. And if you want to just really move some existing workload, GCE is your best choice. I like that that’s fairly straightforward. There’s still going to be some it depends, but it feels better than nine ways to run Kubernetes engines.

Corey: I’m sure we’ll see them in the fullness of time.

Richard: [laugh].

Corey: So, talk about Anthos a bit. That was a thing that was announced a while back and it was extraordinarily unclear what it was. And then I looked at the pricing and it was $10,000 a month with a one-year minimum commitment, and is like, “Oh, it’s not for me. That’s why I don’t get it.” And I haven’t really looked back at it since. But it is something else now. It almost feels like a wrapper brand, in some respects. How’s it going? [unintelligible 00:29:26]?

Richard: Yeah. Consumption, we’ll talk more upcoming months on some of the adoption, but we’re finally getting the hockey stick, which always comes delayed with platforms because nobody adopts platforms quickly. They buy the platform and a year later they start to actually build new development, migrate the things they have. So, we’re starting to see the sort of growth.

But back to your first point. And I even think I poorly tried to explain it a year ago with you. Basically, look, Anthos is the ability to manage fleets of GKE clusters, wherever they are. I don’t care if they’re on-prem, I don’t care if they’re in Google Cloud, I don’t care if they’re Amazon. We have one customer who only uses Anthos on AWS. Awesome, rock on.

So, how do I put GKE clusters everywhere, but then do fleet management because look, some people are doing an app per cluster. They don’t want to jam 50 apps in the cluster from different teams because they don’t like the idea that this app requires root access; now you can screw around with mine. Or, you didn’t update; that broke the cluster. I don’t want any of that. So, you’re going to see companies more, doing even app per cluster, app per developer per cluster.

So, now I have a fleet problem. How do I keep it in sync? How do I make sure policy is consistent? Those sorts of things. So, Anthos is kind of solving the fleet management challenge and replacing people’s first-gen app platform.

Seeing a lot of those use cases, “Hey, we’re retiring our first version of Docker Enterprise, Mesos, Cloud Foundry, even OpenShift,” saying, “All right, now’s the time for our next version of our app platform. How about GKE, plus Cloud Run on top of it, plus other stuff?” Sounds good. So, going well is a, sort of—as you mentioned, there’s a brand story here, mainly because we’ve also done two things that probably matter to you. A, we changed the price a lot.

No minimum commit, remarkably at 20% of the cost it was when we launched, on purpose because we’ve gotten better at this. So, much cheaper, no minimum commit, pay as you go. Be on-premises, on bare metal with GKE. Pay by the hour, I don’t care; sounds great. So, you can do that sort of stuff.

But then more importantly, if you’re a GKE customer and you just want config management, service mesh, things like that, now you can buy all of those independently as well. And Anthos is really the brand for fleet management of GKE. And if you’re on Google Cloud only, it adds value. If you’re off Google Cloud, if you’re multi-cloud, I don’t care. But I want to manage fleets of compute clusters and create them. We’re going to keep doubling down on that.

Corey: The big problem historically for understanding a lot of the adoption paradigm of Kubernetes has been that it was, to some extent, a reimagining of how Google ran and built software internally. And I thought at the time, the idea was—from a cynical perspective—that, “All right, well, your crappy apps don’t run well on Google-style infrastructure so we’re going to teach the entire world how to write software the way that we do.” And then you end up with people running their blog on top of Kubernetes, where it’s one of those, like, the first blog post is, like, “How I spent the last 18 months building Kubernetes.” And, okay, that is certainly a philosophy and an approach, but it’s almost approaching Windows 95 launch level of hype, where people who didn’t own computers were buying copies of it, on some level. And I see the term come up in conversations in places where it absolutely has no place being brought up. “How do I run a Kubernetes cluster inside of my laptop?” And, “It’s what you got going on in there, buddy?”

Richard: [laugh].

Corey: “What do you think you’re trying to do here because you just said something that means something that I think is radically different to me than it is to you.” And again, I’m not here to judge other people’s workflows; they’re all terrible, except for mine, which is an opinion held by everyone about their own workflow. But understanding where people are, figuring out how to get there, how to meet customers where they are and empower them. And despite how heavily Google has been into the Kubernetes universe since its inception, you’re very welcoming to companies—and loud-mouth individuals on Twitter—who have no use for Kubernetes. And working through various products you offer, I don’t ever feel like a second-class citizen. There’s really something impressive about that, of not letting the hype dictate the product and marketing decisions of it.

Richard: Yeah, look, I think I tweeted it recently, I think the future of software is managed services with containers in the gap, for the most part. Whereas—if you can use managed services, please do. Use them wherever you can. And if you have to sling some code, maybe put it in a really portable thing that’s really easy to run in lots of places. So, I think that’s smart.

But for us, look, I think we have the best container workflow from dev tools, and build tools, and artifact registries, and runtimes, but plenty of people are running containers, and you shouldn’t be running Kubernetes all over the place. That makes sense for the workload, I think it’s better than a VM at the retail edge. Can I run a small cluster, instead of a weird point-of-sale Windows app? Maybe. Maybe it makes sense to have a lightweight Kubernetes cluster there for consistency purposes.

So, for me, I think it’s a great medium for a subset of software. Google Cloud is going to take whatever you got, which is great. I think containers are great, but at the same time, I’m happily going to let you deploy a function that responds to you adding a storage item to a bucket, where at the same time give you a SaaS service that replaces the need for any code. All of those are terrific. So yeah, we love Kubernetes. We think it’s great. We’re going to be the best version to run it. But that’s not going to be your whole universe.

Corey: No, and I would argue it absolutely shouldn’t be.

Richard: [laugh]. Right. Agreed. Now again, for some companies, it’s a great replacement for this giant fleet of VMs that all runs at eight percent utilization. Can I stick this into a bunch of high-density clusters? Absolutely you should. You’re going to save an absolute fortune doing that and probably pick up some resilience and functionality benefits.

But to your point, “Do I want to run a WordPress site in there?” I don’t know, probably not. “Do I need to run my own MySQL?” I’d prefer you not do that. So, in a lot of cases, don’t use it unless you have to. That should go for all compute nowadays. Use managed services.

Corey: I’m a big believer in going down that approach just because it is so much easier than trying to build it yourself from popsicle sticks because you theoretically might have to move it someday in the future, even though you’re not.

Richard: [laugh]. Right.

Corey: And it lets me feel better about a thing that isn’t going to be used by anything that I’m doing in the near future. I just don’t pretend to get it.

Richard: No, I don’t install a general purpose electric charger in my garage for any electric car I may get in the future; I charge for the one I have now. I just want it to work for my car; I don’t want to plan for some mythical future. So yeah, premature optimization over architecture, or death in IT, especially nowadays where speed matters, don’t waste your time building something that can run in nine clouds.

Corey: Richard, I want to thank you for coming on again a year later to suffer my slings, arrows, and other various implements of misfortune. If people want to learn more about what you’re doing, how you’re doing it, possibly to pull a Forrest Brazeal and go work with you, where can they find you?

Richard: Yeah, we’re a fun place to work. So, you can find me on Twitter at @rseroter—R-S-E-R-O-T-E-R—hang out on LinkedIn, annoy me on my blog seroter.com as I try to at least explore our tech from time to time and mess around with it. But this is a fun place to work. There’s a lot of good stuff going on here, and if you work somewhere else, too, we can still be friends.

Corey: Thank you so much for your time today. Richard Seroter, director of outbound product management at Google. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry comment
into which you have somehow managed to shove a running container.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Micheal

Micheal Benedict leads Engineering Productivity at Pinterest. He and his team focus on developer experience, building tools and platforms for over a thousand engineers to effectively code, build, deploy and operate workloads on the cloud. Mr. Benedict has also built Infrastructure and Cloud Governance programs at Pinterest and previously, at Twitter -- focussed on managing cloud vendor relationships, infrastructure budget management, cloud migration, capacity forecasting and planning and cloud cost attribution (chargeback).

Links:

  • Pinterest: https://www.pinterest.com
  • Teletraan: https://github.com/pinterest/teletraan
  • Twitter: https://twitter.com/micheal
  • Pinterestcareers.com: https://pinterestcareers.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: You know how git works right?

Announcer: Sorta, kinda, not really. Please ask someone else!

Corey: Thats all of us. Git is how we build things, and Netlify is one of the best way I’ve found to build those things quickly for the web. Netlify’s git based workflows mean you don't have to play slap and tickle with integrating arcane non-sense and web hooks, which are themselves about as well understood as git. Give them a try and see what folks ranging from my fake Twitter for pets startup, to global fortune 2000 companies are raving about. If you end up talking to them, because you don't have to, they get why self service is important—but if you do, be sure to tell them that I sent you and watch all of the blood drain from their faces instantly. You can find them in the AWS marketplace or at www.netlify.com. N-E-T-L-I-F-Y.com

Corey: This episode is sponsored in part by our friends at Vultr. Spelled V-U-L-T-R because they’re all about helping save money, including on things like, you know, vowels. So, what they do is they are a cloud provider that provides surprisingly high performance cloud compute at a price that—while sure they claim its better than AWS pricing—and when they say that they mean it is less money. Sure, I don’t dispute that but what I find interesting is that it’s predictable. They tell you in advance on a monthly basis what it’s going to going to cost. They have a bunch of advanced networking features. They have nineteen global locations and scale things elastically. Not to be confused with openly, because apparently elastic and open can mean the same thing sometimes. They have had over a million users. Deployments take less that sixty seconds across twelve pre-selected operating systems. Or, if you’re one of those nutters like me, you can bring your own ISO and install basically any operating system you want. Starting with pricing as low as $2.50 a month for Vultr cloud compute they have plans for developers and businesses of all sizes, except maybe Amazon, who stubbornly insists on having something to scale all on their own. Try Vultr today for free by visiting: vultr.com/screaming, and you’ll receive a $100 in credit. Thats v-u-l-t-r.com slash screaming.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Every once in a while, I like to talk to people who work at very large companies that are not in fact themselves a cloud provider. I know it sounds ridiculous. How can you possibly be a big company and not make money by selling managed NAT gateways to an unsuspecting public? But I’m told it can be done here to answer that question. And hopefully at least one other is Pinterest. It’s head of engineering productivity, Micheal Benedict. Micheal, thank you for taking the time to join me today.

Micheal: Hi, Corey, thank you for inviting me today. I’m really excited to talk to you.

Corey: So, exciting times at Pinterest in a bunch of different ways. It was recently reported—which of course, went right to the top of my inbox as 500,000 people on Twitter all said, “Hey, this sounds like a ‘Corey would be interested in it’ thing.” It was announced that you folks had signed a $3.2 billion commitment with AWS stretching until 2028. Now, if this is like any other large-scale AWS contract commitment deal that has been made public, you were probably immediately inundated with a whole bunch of people who are very good at arithmetic and not very good at business context saying, “$3.2 billion? You could build massive data centers for that. Why would anyone do this?” And it’s tiresome, and that’s the world in which we live. But I’m guessing you heard at least a little bit of that from the peanut gallery.

Micheal: I did, and I always find it interesting when direct comparisons are made with the total amount that’s been committed. And like you said, there’s so many nuances that go into how to perceive that amount, and put it in context of, obviously, what Pinterest does. So, I at least want to take this opportunity to share with everyone that Pinterest has been on the cloud since day one. When Ben initially started the company, that product was launched—it was a simple Django app—it was launched on AWS from day one, and since then, it has grown to support 450-plus million MAUs over the course of the decade.

And our infrastructure has grown pretty complex. We started with a bunch of EC2 machines and persisting data in S3, and since then we have explored an array of different products, in fact, sometimes working very closely with AWS, as well and helping them put together a product roadmap for some of the items they’re working on as well. So, we have an amazing partnership with them, and part of the commitment and how we want to see these numbers is how does it unlock value for Pinterest as a business over time in terms of making us much more agile, without thinking about the nuances of the infrastructure itself. And that’s, I think, one of the best ways to really put this into context, that it’s not a single number we pay at the end [laugh] of the month, but rather, we are on track to spending a certain amount over a period of time, so this just keeps accruing or adding to that number. And we basically come out with an amazing partnership in AWS, where we have that commitment and we’re able to leverage their products and full suite of items without any hiccups.

Corey: The most interesting part of what you said is the word partner. And I think that’s the piece that gets lost an awful lot when we talk about large-scale cloud negotiations. It’s not like buying a car, where you can basically beat the crap out of the salesperson, you can act as if $400 price difference on a car is the difference between storm out of the dealership and sign the contract. Great, you don’t really have to deal with that person ever again.

In the context of a cloud provider, they run your production infrastructure, and if they have a bad day, I promise you’re going to have a bad day, too. You want to handle those negotiations in a way that is respectful of that because they are your partner, whether you want them to be or not. Now, I’m not suggesting that any cloud provider is going to hold an awkward negotiation against the customer, but at the same time, there are going to be scenarios in which you’re going to want to have strong relationships, where you’re going to need to cash in political capital to some extent, and personally, I’ve never seen stupendous value in trying to beat the crap out of a company in order to get another tenth of a percent discount on a service you barely use, just because someone decided that well, we didn’t do well in the last negotiation so we’re going to get them back this time.

That's great. What are you actually planning to do as a company? Where are you going? And the fact that you just alluded to, that you’re not just a pile of S3 and EC2 instances speaks, in many ways, to that. By moving into the differentiated service world, suddenly you’re able to do things that don’t look quite as much like building a better database and start looking a lot more like servicing your users more effectively and well.

Micheal: And I think, like you said, I feel like there’s like a general skepticism in viewing that the cloud providers are usually out there to rip you apart. But in reality, that’s not true. To your point, as part of the partnership, especially with AWS and Pinterest, we’ve got an amazing relationship going on, and behind the scenes, there’s a dedicated team at Pinterest, called the Infrastructure Governance Team, a cross-functional team with folks from finance, legal, engineering, product, all sitting together and working with our AWS partners—even the AWS account managers at the times are part of that—to help us make both Pinterest successful, and in turn, AWS gets that amazing customer to work with in helping build some of their newer products as well. And that’s one of the most important things we have learned over time is that there’s two parts to it; when you want to help improve your business agility, you want to focus not just on the bottom line numbers as they are.
It’s okay to pay a premium because it offsets the people capital you would have to invest in getting there.

And that’s a very tricky way to look at math, but that’s what these teams do; they sit down and work through those specifics. And for what it’s worth, in our conversations, the AWS teams always come back with giving us very insightful data on how we’re using their systems to help us better think about how we should be pricing or looking things ahead. And I’m not the expert on this; like I said, there’s a dedicated team sitting behind this and looking through and working through these deals, but that’s one of the important takeaways I hope the users—or the listeners of this podcast then take away that you want to treat your cloud provider as your partner as much as possible. They’re not always there to screw you. That’s not their goal. And I apologize for using that term. It is important that you set that expectations that it’s in their best interest to actually make you successful because that’s how they make money as well.

Corey: It’s a long-term play. I mean, they could gouge you this quarter, and then you’re trying to evacuate as fast as possible. Well, they had a great quarter, but what’s their long-term prospect? There are two competing philosophies in the world of business; you can either make a lot of money quickly, or you can make a little bit of money and build it over time in a sustained way. And it’s clear the cloud providers are playing the long game on this because they basically have to.

Micheal: I mean, it’s inevitable at this point. I mean, look at Pinterest. It is one of those success stories. Starting as a Django app on a bunch of EC2 machines to wherever we are right now with having a three-plus billion dollar commitment over a span of couple of years, and we do spend a pretty significant chunk of that on a yearly basis. So, in this case, I’m sure it was a great successful partnership.

And I’m hoping some of the newer companies who are building the cloud from the get-go are thinking about it from that perspective. And one of the things I do want to call out, Corey, is that we did initially start with using the primitive services in AWS, but it became clear over time—and I’m sure you heard of the term multi-cloud and many of that—you know, when companies start evaluating how to make the most out of the deals they’re negotiating or signing, it is important to acknowledge that the cost of any of those evaluations or even thinking about migrations never tends to get factored in. And we always tend to treat that as being extremely simple or not, but those are engineering resources you want to be spending more building on the product rather than these crazy costly migrations. So, it’s in your best interest probably to start using the most from your cloud provider, and also look for opportunities to use other cloud providers—if they provide more value in certain product offerings—rather than thinking about a complete lift-and-shift, and I’m going to make DR as being the primary case on why I want to be moving to multi-cloud.

Corey: Yeah. There’s a question, too, of the numbers on paper look radically different than the reality of this. You mentioned, Pinterest has been on AWS since the beginning, which means that even if an edict had been passed at the beginning, that, “Thou shalt never build on anything except EC2 and S3. The end. Full stop.”

And let’s say you went down that rabbit hole of, “Oh, we don’t trust their load balancers. We’re going to build our own at home. We have load balancers at home. We’ll use those.” It’s terrible, but even had you done that and restricted yourselves just to those baseline building blocks, and then decide to do a cloud migration, you’re still looking back at over a decade of experience where the app has been built unconsciously reflecting the various failure modes that AWS has, the way that it responds to API calls, the latency in how long it takes to request something versus it being available, et cetera, et cetera.

So, even moving that baseline thing to another cloud provider is not a trivial undertaking by any stretch of the imagination. But that said—because the topic does always come up, and I don’t shy away from it; I think it’s something people should go into with an open mind—how has the multi-cloud conversation progressed at Pinterest? Because there’s always a multi-cloud conversation.

Micheal: We have always approached it with some form of… openness. It’s not like we don’t want to be open to the ideas, but you really want to be thinking hard on the business case and the business value something provides on why you want to be doing x. In this case, when we think about multi-cloud—and again, Pinterest did start with EC2 and S3, and we did keep it that way for a long time. We built a lot of primitives around it, used it—for example, my team actually runs our bread and butter deployment system on EC2. We help facilitate deployments across a 100,000-plus machines today.

And like you said, we have built that system keeping in mind how AWS works, and understanding the nuances of region and AZ failovers and all of that, and help facilitate deployments across 1000-plus microservices in the company. So, thinking about leveraging, say, a Google Cloud instance and how that works, in theory, we can always make a case for engineering to build our deployment system and expand there, but there’s really no value. And one of the biggest cases, usually, when multi-cloud comes in is usually either negotiation for price or actually a DR strategy. Like, what if AWS goes down in and us-east-1? Well, let’s be honest, they’re powering half the internet [laugh] from that one single—

Corey: Right.

Micheal: Yeah. So, if you think your business is okay running when AWS goes down and half the internet is not going to be working, how do you want to be thinking about that? So, DR is probably not the best reason for you to be even exploring multi-cloud. Rather, you should be thinking about what the cloud providers are offering as a very nuanced offering which your current cloud provider is not offering, and really think about just using those specific items.

Corey: So, I agree that multi-cloud for DR purposes is generally not necessarily the best approach with the idea of being able to failover seamlessly, but I like the idea for backups. I mean, Pinterest is a publicly-traded company, which means that among other things, you have to file risk disclosures and be responsive to auditors in a variety of different ways. There are some regulations to start applying to you. And the idea of, well, AWS builds things out in a super effective way, region separation, et cetera, whenever I talk to Amazonians, they are always
surprised that anyone wouldn’t accept that, “Oh, if you want backups use a different region. Problem solved.”

Right, but it is often easier for me to have a rehydrate the business level of backup that would take weeks to redeploy living on another cloud provider than it is for me to explain to all of those auditors and regulators and financial analysts, et cetera why I didn’t go ahead and do that path. So, there’s always some story for okay, what if AWS decides that they hate us and want to kick us off the platform? Well, that’s why legal is involved in those high-level discussions around things like risk, and indemnity, and termination for convenience and for cause clauses, et cetera, et cetera. The idea of making an all-in commitment to a cloud provider goes well beyond things that engineering thinks about. And it’s easy for those of us with engineering backgrounds to be incredibly dismissive of that of, “Oh, indemnity? Like, when does AWS ever lose data?” “Yeah, but let’s say one day they do. What is your story going to be when asked some very uncomfortable questions by people who wanted you to pay attention to this during the negotiation process?” It’s about dotting the i’s and crossing the t’s, especially with that many commas in the contractual commitments.

Micheal: No, it is true. And we did evaluate that as an option, but one of the interesting things about compliance, and especially auditing as well, we generally work with the best in class consultants to help us work through the controls and how we audit, how we look at these controls, how to make sure there’s enough accountability going through. The interesting part was in this case, as well, we were able to work with AWS in crafting a lot of those controls and setting up the right expectations as and when we were putting proposals together as well. Now, again, I’m not an expert on this and I know we have a dedicated team from our technical program management organization focused on this, but early on we realized that, to your point, the cost of any form of backups and then being able to audit what’s going in, look at all those pipelines, how quickly we can get the data in and out it was proving pretty costly for us. So, we were able to work out some of that within the constructs of what we have with our cloud provider today, and still meet our compliance goals.

Corey: That’s, on some level, the higher point, too, where everything is everything comes down to context; everything comes down to what the business demands, what the business requires, what the business will accept. And I’m not suggesting that in any case, they’re wrong. I’m known for beating the ‘Multi-cloud is a bad default decision’ drum, and then people get surprised when they’ll have one-on-one conversations, and they say, “Well, we’re multi-cloud. Do you think we’re foolish?” “No. You’re probably doing the right thing, just because you have context that is specific to your business that I, speaking in a general sense, certainly don’t have.”

People don’t generally wake up in the morning and decide they’re going to do a terrible job or no job at all at work today, unless they’re Facebook’s VP of Integrity. So, it’s not the sort of thing that lends itself to casual tweet size, pithy analysis very often. There’s a strong dive into what is the level of risk a business can accept? And my general belief is that most companies are doing this stuff right. The universal constant in all of my consulting clients that I have spoken to about the in-depth management piece of things is, they’ve always asked the same question of, “So, this is what we’ve done, but can you introduce us to the people who are doing it really right, who have absolutely nailed this and gotten it all down?” “It’s, yeah, absolutely no one believes that that is them, even the folks who are, from my perspective, pretty close to having achieved it.”

But I want to talk a bit more about what you do beyond just the headline-grabbing large dollar figure commitment to a cloud provider story. What does engineering productivity mean at Pinterest? Where do you start? Where do you stop?

Micheal: I want to just quickly touch upon that last point about multi-cloud, and like you said, every company works within the context of what they are given and the constraints of their business. It’s probably a good time to give a plug to my previous employer, Twitter, who are doing multi-cloud in a reasonably effective way. They are on the data centers, they do have presence on Google Cloud, and AWS, and I know probably things have changed since a couple of years now, but they have embraced that environment pretty effectively to cater to their acquisitions who were on the public cloud, help obviously, with their initial set of investments in the data center, and still continue to scale that out, and explore, in this case, Google Cloud for a variety of other use cases, which sounds like it’s been extremely beneficial as well.

So, to your point, there is probably no right way to do this. There’s always that context, and what you’re working with comes into play as part of making these decisions. And it’s important to take a lot of these with a grain of salt because you can never understand the decisions, why they were made the way they were made. And for what it’s worth, it sort of works out in the end. [laugh]. I’ve rarely heard a story where it’s never worked out, and people are just upset with the deals they’ve signed. So, hopefully, that helps close that whole conversation about multi-cloud.

Corey: I hope so. It’s one of those areas where everyone has an opinion and a lot of them do not necessarily apply universally, but it’s always fun to take—in that case, great, I’ll take the lesser trod path of everyone’s saying multi-cloud is great, invariably because they’re trying to sell you something. Yeah, I have nothing particularly to sell, folks. My argument has always been, in the absence of a compelling reason not to, pick a provider and go all in. I don’t care which provider you pick—which people are sometimes surprised to hear.

It’s like, “Well, what if they pick a cloud provider that you don’t do consulting work for?” Yeah, it turns out, I don’t actually need to win every AWS customer over to have a successful working business. Do what makes sense for you, folks. From my perspective, I want this industry to be better. I don’t want to sit here and just drum up business for myself and make self-serving comments to empower that. Which apparently is a rare tactic.

Micheal: No, that’s totally true, Corey. One of the things you do is help people with their bills, so this has come up so many times, and I realize we’re sort of going off track a bit from that engineering productivity discussion—

Corey: Oh, which is fine. That’s this entire show’s theme, if it has one.

Micheal: [laugh]. So, I want to briefly just talk about the whole billing and how cost management works because I know you spend a lot of time on that and you help a lot of these companies be effective in how they manage their bills. These questions have come up multiple times, even at Pinterest. We actually in the past, when I was leading the infrastructure governance organization, we were working with other companies of our similar size to better understand how they are looking into getting visibility into their cost, setting sort of the right controls and expectations within the engineering organization to plan, and capacity plan, and effectively meet those plans in a certain criteria, and then obviously, if there is any risk to that, actively manage risk. That was like the biggest thing those teams used to do.

And we used to talk a lot trade notes, and get a better sense of how a lot of these companies are trying to do—for example, Netflix, or Lyft, or Stripe. I recall Netflix, content was their biggest spender, so cloud spending was like way down in the list of things for them. [laugh]. But regardless, they had an active team looking at this on a day-to-day basis. So, one of the things we learned early on at Pinterest is that start investing in those visibility tools early on.

No one can parse the cloud bills. Let’s be honest. You’re probably the only person who can reverse… [laugh] engineer an architecture diagram from a cloud bill, and I think that’s like—definitely you should take a patent for that or something. But in reality, no one has the time to do that. You want to make sure your business leaders, from your finance teams to engineering teams to head of the executives all have a better understanding of how to parse it.

So, investing engineering resources, take that data, how do you munch it down to the cost, the utilization across the different vectors of offerings, and have a very insightful discussion. Like, what are certain action items we want to be taking? It’s very easy to see, “Oh, we overspent EC2,” and we want to go from there. But in reality, that’s not just that thing; you will start finding out that EC2 is being used by your Hadoop infrastructure, which runs hundreds of thousands of jobs. Okay, now who’s actually responsible for that cost? You might find that one job which is accruing, sort of, a lot of instance hours over a period of time and a shared multi-tenant environment, how do you attribute that cost to that particular cost center?

Corey: And then someone left the company a while back, and that job just kept running in perpetuity. No one’s checked the output for four years, I guess it can’t be that necessarily important. And digging into it requires context. It turns out, there’s no SaaS tool to do this, which is unfortunate for those of us who set out originally to build such a thing. But we discovered pretty early on the context on this stuff is incredibly important.

I love the thing you’re talking about here, where you’re discussing with your peer companies about these things because the advice that I would give to companies with the level of spend that you folks do is worlds apart from what I would advise someone who’s building something new and spending maybe 500 bucks a month on their cloud bill. Those folks do not need to hire a dedicated team of people to solve for these problems. At your scale, yeah, you probably should have had some people in [laugh] here looking at this for a while now. And at some point, the guidance changes based upon scale. And if there’s one thing that we discover from the horrible pages of Hacker News, it’s that people love applying bits of wisdom that they hear in wildly inappropriate situations.

How do you think about these things at that scale? Because, a simple example: right now I spend about 1000 bucks a month at The Duckbill Group, on our AWS bill. I know. We have one, too. Imagine that. And if I wind up just committing admin credentials to GitHub, for example, and someone compromises that and start spinning things up to mine all the Bitcoin, yeah, I’m going to notice that by the impact it has on the bill, which will be noticeable from orbit.

At the level of spend that you folks are at, at company would be hard-pressed to spin up enough Bitcoin miners to materially move the billing needle on a month-to-month basis, just because of the sheer scope and scale. At small bill volumes, yeah, it’s pretty easy to discover the thing that spiking your bill to three times normal. It’s usually a managed NAT gateway. At your scale, tripling the bill begins to look suspiciously like the GDP of a small country, so what actually happened here? Invariably, at that scale, with that level of massive multiplier, it’s usually the simplest solution, an error somewhere in the AWS billing system. Yes, they exist. Imagine that.

Micheal: They do exist, and we’ve encountered that.

Corey: Kind of heartstopping, isn’t it?

Micheal: [laugh]. I don’t know if you remember when we had the big Spectre and the Meltdown, right, and those were interesting scenarios for us because we had identified a lot of those issues early on, given the scale we operate, and we were able to, sort of, obviously it did have an impact on the builds and everything, but that’s it; that’s why you have these dedicated teams to fix that. But I think one of the points you made, these are large bills and you’re never going to have a 3x jump the next day. We’re not going to be seeing that. And if that happens, you know, God save us. [laugh].

But to your point, one of the things we do still want to be doing is look at trends, literally on a week-over-week basis because even a one percentage move is a pretty significant amount, if you think about it, which could be funding some other aspects of the business, which we would prefer to be investing on. So, we do want to have enough rigor and controls in place in our technical stack to identify and alert when something is off track. And it becomes challenging when you start using those higher-order services from your public cloud provider because there’s no clear insights on how do you, kind of, parse that information. One of the biggest challenges we had at Pinterest was tying ownership to all these things.

No, using tags is not going to cut it. It was so difficult for us to get to a point where we could put some sense of ownership in all the things and the resources people are using, and then subsequently have the right conversation with our ads infrastructure teams, or our product teams to help drive the cost improvements we want to be seeing. And I wouldn’t be surprised if that’s not a challenge already, even for the smaller companies who have bills in the tunes of tens and thousands, right?

Corey: It is. It’s predicting the spend and trying to categorize it appropriately; that’s the root of all AWS bill panic on the corporate level. It’s not that the bill is 20% higher, so we’re going to go broke. Most companies spend far more on payroll than they do on infrastructure—as you mentioned with Netflix, content is a significantly larger [laugh] expense than any of those things; real estate, it’s usually right up there too—but instead it’s, when you’re trying to do business forecasting of, okay, if we’re going to have an additional 1000 monthly active users, what will the cost for us be to service those users and, okay, if we’re seeing a sudden 20% variance, if that’s the new normal, then well, that does change our cost projections for a number of years, what happens? When you’re public, there starts to become the question of okay, do we have to restate earnings or what’s the deal here?

And of course, all this sidesteps past the unfortunate reality that, for many companies, the AWS bill is not a function of how many customers you have; it’s how many engineers you hired. And that is always the way it winds up playing out for some reason. “It’s why did we see a 10% increase in the bill? Yeah, we hired another data science team. Oops.” It’s always seems to be the data science folks; I know I’d beat up on those folks a fair bit, and my apologies. And one day, if they analyze enough of the data, they might figure out why.

Micheal: So, this is where I want to give a shout out to our data science team, especially some of the engineers working in the Infrastructure Governance Team putting these charts together, helping us derive insights. So, definitely props to them.

I think there’s a great segue into the point you made. As you add more engineers, what is the impact on the bottom line? And this is one of the things actually as part of engineering productivity, we think about as well on a long-term basis. Pinterest does have over 1000-plus engineers today, and to large degree, many of them actually have their own EC2 instances today. And I wouldn’t say it’s a significant amount of cost, but it is a large enough number, were shutting down a c5.9xl can actually fund a bunch of conference tickets or something else.

And then you can imagine that sort of the scale you start working with at one point. The nuance here is though, you want to make sure there’s enough flexibility for these engineers to do their local development in a sustainable way, but when moving to, say production, we really want to tighten the flexibility a bit so they don’t end up doing what you just said, spin up a bunch of machines talking to the API directly which no one will be aware of.

I want to share a small anecdote because when back in the day, this was probably four years ago, when we were doing some analysis on our bills, we realized that there was a huge jump every—I believe Wednesday—in our EC2 instances by almost a factor of, like, 500 to 600 instances. And we’re like, “Why is this happening? What is going on?” And we found out there was an obscure job written by someone who had left the company, calling an EC2 API to spin up a search cluster of 500 machines on-demand, as part of pulling that ETL data together, and then shutting that cluster down. Which at times didn’t work as expected because, you know, obviously, your Hadoop jobs are very predictable, right?

So, those are the things we were dealing with back in the day, and you want to make sure—since then—this is where engineering productivity as team starts coming in that our job is to enable every engineer to be doing their best work across code building and deploying the services. And we have done this.

Corey: Right. You and I can sit here and have an in-depth conversation about the intricacies of AWS billing in a bunch of different ways because in different ways we both specialize in it, in many respects. But let’s say that Pinterest theoretically was foolish enough to hire me before I got into this space as an engineer, for terrifying reasons. And great. I start day one as a typical software developer if such a thing could be said to exist. How do you effectively build guardrails in so that I don’t inadvertently wind up spinning up all the EC2 instances available to me within an account, which it turns out are more than one might expect sometimes, but still leave me free to do my job without effectively spending a nine-month safari figuring out how AWS bills work?

Micheal: And this is why teams like ours exist, to help provide those tools to help you get started. So today, we actually don’t let anyone directly use AWS APIs, or even use the UI for that matter. And I think you’ll soon realize, the moment you hit, like, probably 30 or 40 people in your organization, you definitely want to lock it down. You don’t want that access to be given to anyone or everyone. And then subsequently start building some higher-order tools or abstraction so people can start using that to control effectively.

In this case, if you’re a new engineer, Corey, which it seems like you were, at some point—

Corey: I still write code like I am, don’t worry.

Micheal: [laugh]. So yes, you would get access to our internal tool to actually help spin up what we call is a dev app, where you get a chance to, obviously, choose the instance size, not the instance type itself, and we have actually constrained the instance types we have approved within Pinterest as well. We don’t give you the entire list you get a chance to choose and deploy to. We actually have constraint to based on the workload types, what are the instance types we want to support because in the future, if we ever want to move from c3 to c5—and I’ve been there, trust me—it is not an easy thing to do, so you want to make sure that you’re not letting people just use random instances, and constrain that by building some of these tools. As a new engineer, you would go in, you’d use the tool, and actually have a dev app provisioned for you with our Pinterest image to get you started.

And then subsequently, we’ll obviously shut it down if we see you not being using it over a certain amount of time, but those are sort of the guardrails we’ve put in over there so you never get a chance to directly ever use the EC2 APIs, or any of those AWS APIs to do certain things. The similar thing applies for S3 or any of the higher-order tools which AWS will provide, too.

Corey: This episode is sponsored by our friends at Oracle Cloud. Counting the pennies, but still dreaming of deploying apps instead of "Hello, World" demos? Allow me to introduce you to Oracle's Always Free tier. It provides over 20 free services and infrastructure, networking databases, observability, management, and security.

And - let me be clear here - it's actually free. There's no surprise billing until you intentionally and proactively upgrade your account. This means you can provision a virtual machine instance or spin up an autonomous database that manages itself all while gaining the networking load, balancing and storage resources that somehow never quite make it into most free tiers needed to support the application that you want to build.

With Always Free you can do things like run small scale applications, or do proof of concept testing without spending a dime. You know that I always like to put asterisks next to the word free. This is actually free. No asterisk. Start now. Visit https://snark.cloud/oci-free that's https://snark.cloud/oci-free.

Corey: How does that interplay with AWS launches yet another way to run containers, for example, and that becomes a valuable potential avenue to get some business value for a developer, but the platform you built doesn’t necessarily embrace that capability? Or they release a feature to an existing tool that you use that could potentially be a just feature capability story, much more so than a cost savings one. How do you keep track of all of that and empower people to use those things so they’re not effectively trying to reimplement DynamoDB on top of
EC2?

Micheal: That’s been a challenge, actually, in the past for us because we’ve always been very flexible where engineers have had an opportunity to write their own solutions many a times rather than leveraging the AWS services, and of late, that’s one of the reasons why we have an infrastructure organization—an extremely lean organization for what it’s worth—but then still able to achieve outsized outputs. Where we evaluate a lot of these use cases, as they come in and open up different aspects of what we want to provide say directly from AWS, or build certain abstractions on top of it. Every time we talk about containers, obviously, we always associate that with something like Kubernetes and offerings from there on; we realized that our engineers directly never ask for those capabilities. They don’t come in and say, “I need a new container orchestration system. Give that to me, and I’m going to be extremely productive.”

What people actually realize is that if you can provide them effective tools and that can help them get their job done, they would be happy with it. For example, like I said, our deployment system, which is actually an open-source system called Teletraan. That is the bread and butter at Pinterest at which my team runs. We operate 100,000-plus machines. We have actually looked into container orchestration where we do have a dedicated Kubernetes team looking at it and helping certain use cases moved there, but we realized that the cost of entire migrations need to be evaluated against certain use cases which can benefit from being on Kubernetes from day one. You don’t want to force anyone to move there, but give them the right incentives to move there. Case in point, let’s upgrade your OS. Because if you’re managing machines,
obviously everyone loves to upgrade their OSes.

Corey: Well, it’s one of the things I love savings plans versus RIs; you talk about the c3 to c5 migration and everyone has a story about one of those, but the most foolish or frustrating reason that I ever saw not to do the upgrade was what we bought a bunch of Reserved Instances on the C3s and those have a year-and-a-half left to run. And it’s foolish not on the part of customers—it’s economically sound—but on the part of AWS where great, you’re now forcing me to take a contractual commitment to something that serves me less effectively, rather than getting out of the way and letting me do my job. That’s why it’s so important to me at least, that savings plans cover Fargate and Lambda, I wish they covered SageMaker instead of SageMaker having its own thing because once again, you’re now architecturally constrained based upon some ridiculous economic model that they have imposed on us. But that’s a separate rant for another time.

Micheal: No, we actually went through that process because we do have a healthy balance of how we do Reserved Instances and how we look at on-demand. We’ve never been big users have spot in the past because just the spot market itself, we realized that putting that pressure on our customers to figure out how to manage that is way more. When I say customers, in this case, engineers within the organization.

Corey: Oh, yes. “I want to post some pictures on Pinterest, so now I have to understand the spot market. What?” Yeah.

Micheal: [laugh]. So, in this case, when we even we’re moving from C3 to C5—and this is where the partnership really plays out effectively, right, because it’s also in the best interest of AWS to deprecate their aging hardware to support some of these new ones where they could also be making good enough premium margins for what it’s worth and give the benefit back to the user. So, in this case, we were able to work out an extremely flexible way of moving to a C5 as soon as possible, get help from them, actually, in helping us do that, too, allocating capacity and working with them on capacity management. I believe at one point, we were actually one of the largest companies with a C3 footprint and it took quite a while for us to move to C5. But rest assured, once we moved, the savings was just immense. We were able to offset any of those RI and we were able to work behind the scenes to get that out. But obviously, not a lot of that is considered in a small-scale company just because of, like you said, those constraints which have been placed in a contractual obligation.

Corey: Well, this is an area in which I will give the same guidance to companies of your scale as well as small-scale companies. And by small-scale, I mean, people on the free tier account, give or take, so I do mean the smallest of the small. Whenever you wind up in a scenario where you find yourself architecturally constrained by an economic barrier like this, reach out to your account manager. I promise you have one. Every account, even the tiny free tier accounts, have an account manager.

I have an account manager, who I have to say has probably one of the most surreal jobs that AWS, just based upon the conversations I throw past him. But it’s reaching out to your provider rather than trying to solve a lot of this stuff yourself by constraining how you’re building things internally is always the right first move because the worst case is you don’t get anywhere in those conversations. Okay, but at least you explored that, as opposed to what often happens is, “Oh, yeah. I have a switch over here I can flip and solve your entire problem. Does that help anything?”

Micheal: Yeah.

Corey: You feel foolish finding that out only after nine months of dedicated work, it turns out.

Micheal: Which makes me wonder, Corey. I mean, do you see a lot of that happening where folks don’t tend to reach out to their account managers, or rather treat them as partners in this case, right? Because it sounds like there is this unhealthy tension, I would say, as to what is the best help you could be getting from your account managers in this case.

Corey: Constantly. And the challenge comes from a few things, in my experience. The first is that the quality of account managers and the technical account managers—the folks who are embedded many cases with your engineering teams in different ways—does vary. AWS is scaling wildly and bursting at the seams, and people are hard to scale.

So, some are fantastic, some are decidedly less so, and most folks fall somewhere in the middle of that bell curve. And it doesn’t take too many poor experiences for the default to be, “Oh, those people are useless. They never do anything we want, so why bother asking them?” And that leads to an unhealthy dynamic where a lot of companies will wind up treating their AWS account manager types as a ticket triage system, or the last resort of places that they’ll turn when they should be involved in earlier conversations.

I mean, take Pinterest as an example of this. I’m not sure how many technical account managers you have assigned to your account, but I’m going to go out on a limb and guess that the ratio of technical account managers to engineers working on the environment is incredibly lopsided. It’s got to be a high ratio just because of the nature of how these things work. So, there are a lot of people who are actively working on things that would almost certainly benefit from a more holistic conversation with your AWS account team, but it doesn’t occur to them to do it just because of either perceived biases around levels of competence, or poor experiences in the past, or simply not knowing the capabilities that are there. If I could tell one story around the AWS account management story, it would be talk to folks sooner about these
things.

And to be clear, Pinterest has this less than other folks, but AWS does themselves no favors by having a product strategy of, “Yes,” because very often in service of those conversations with a number of companies, there is the very real concern of are they doing research so that they can launch a service that competes with us? Amazon as a whole launching a social network is admittedly one of the most hilarious ideas I [laugh] can come up with and I hope they take a whack at it just to watch them learn all these lessons themselves, but that is again, neither here nor there.

Micheal: That story is very interesting, and I think you mentioned one thing; it’s just that lack of trust, or even knowing what the account managers can actually do for you. There seems to be just a lack of education on that. And we also found it the hard way, right? I wouldn’t say that Pinterest figured this out on day one. We evolved sort of a relationship over time. Yes, our time… engagements are, sort of, lopsided, but we were able to negotiate that as part of deals as we learned a bit more on what we can and we cannot do, and how these individuals are beneficial for Pinterest as well. And—

Corey: Well, here’s a question for you, without naming names—and this might illustrate part of the challenge customers have—how long has
your account manager—not the technical account managers, but your account manager—been assigned to your account?

Micheal: I’ve been at Pinterest for five years and I’ve been working with the same person. And he’s amazing.

Corey: Which is incredibly atypical. At a lot of smaller companies, it feels like, “Oh, I’m your account manager being introduced to you.” And, “Are you the third one this year? Great.” What happens is that if the account manager excels, very often they get promoted and work with a smaller number of accounts at larger spend, and whereas if they don’t find that AWS is a great place for them for a variety of reasons, they go somewhere else and need to be backfilled.

So, at the smaller account, it’s, “Great. I’ve had more account managers in a year than you’ve had in five.” And that is often the experience when you start seeing significant levels of rotation, especially on the customer engineering side where you wind up with you have this big kickoff, and everyone’s aware of all the capabilities and you look at it three years later, and not a single person who was in that kickoff is still involved with the account on either side, and it’s just sort of been evolving evolutionarily from there. One thing that we’ve done in some of our larger accounts as part of our negotiation process is when we see that the bridges have been so thoroughly burned, we will effectively request a full account team cycle, just because it’s time to get new faces in where the customer, in many cases unreasonably, is not going to say, “Yeah but a year-and-a-half ago you did this terrible thing and we’re still salty about it.” Fine, whatever. I get it. People relationships are hard. Let’s go ahead and swap some folks out so that there are new faces with new perspectives because that helps.

Micheal: Well, first off, if you had so many switches in account manager, I think that’s something speaks about [laugh] how you’ve been working, too. I’m just kidding. There are a bu—

Corey: Entirely possible. In seriousness, yes. But if you talk to—like, this is not just me because in my case, yeah, I feel like my account
manager is whoever drew the short straw that week because frankly, yeah, that does seem like a great punishment to wind up passing out to someone who is underperforming. But for a lot of folks who are in the mid-tier, like, spending $50 to $100,000 a month, this is a very common story.

Micheal: Yeah. Actually, we’ve heard a bit about this, too. And like you said, I think maintaining context is the most thing. You really want your account manager to vouch for you, really be your champion in those meetings because AWS, like you said is so large, getting those exec time, and reviews, and there’s so many things that happen, your account manager is the champion for you, or right there. And it’s important and in fact in your best interest to have a great relationship with them as well, not treat them as, oh yet another vendor.

And I think that’s where things start to get a bit messy because when you start treating them as yet another vendor, there is no incentive for them to do the best for you, too. You know, people relationships are hard. But that said though, I think given the amount of customers like these cloud companies are accruing, I wouldn’t be surprised; every account manager seems to be extremely burdened. Even in our case, although I’ve been having a chance to work with this one person for a long time, we’ve actually expanded. We have now multiple account managers helping us out as we’ve started scaling to use certain aspects of AWS which we’ve never explored before.

We were a bit constrained and reserved about what service we want to use because there have been instances where we have tried using something and we have hit the wall pretty immediately. API rate limits, or it’s not ready for primetime, and we’re like, “Oh, my God. Now, what do we do?” So, we have a bit more cautious. But that said, over time, having an account manager who understands how you work, what scale you have, they’re able to advocate with the internal engineering teams within the cloud provider to make the best of supporting you as a customer and tell that success story all the way out.

So yeah, I can totally understand how this may be hard, especially for those small companies. For what it’s worth, I think the best way to really think about it is not treat them as your vendor, but really go out on a limb there. Even though you signed a deal with them, you want to make sure that you have the continuing relationship with them to have—represent your voice better within the company. Which is probably hard. [laugh].

Corey: That’s always the hard part. Honestly, if this were the sort of thing that were easy to automate, or you could wind up building out something that winds up helping companies figure out how to solve these things programmatically, talk about interesting business problems that are only going to get larger in the fullness of time. This is not going away, even if AWS stopped signing up new customers entirely right now, they would still have years of growth ahead of them just from organic growth. And take a company with the scale of Pinterest and just think of how many years it would take to do a full-on exodus, even if it became priority number one. It’s not realistic in many cases, which is why I’ve never been a big fan of multi-cloud as an approach for negotiation. Yeah, AWS has more data on those points than any of us do; they’re not worried about it. It just makes you sound like an unsophisticated negotiator. Pick your poison and lean in.

Micheal: That is the truth you just mentioned, and I probably want to give a call out to our head of infrastructure, [Coburn 00:42:13]. He’s also my boss, and he had brought this perspective as well. As part of any negotiation discussions, like you just said, AWS has way more data points on this than what we think we can do in terms of talking about, “Oh, we are exploring this other cloud provider.” And it’s—they would be like, “Yeah. Do tell me more [laugh] how that’s going.”

And it’s probably in the best interest to never use that as a negotiation tactic because they clearly know the investments that’s going to build on what you’ve done, so you might as well be talking more—again, this is where that relationship really plays together because you want both of them to be successful. And it’s in their best interest to still keep you happy because the good thing about at least companies of our size is that we’re probably, like, one phone call away from some of their executive team, where we could always talk about what didn’t work for us. And I know not everyone has that opportunity, but I’m really hoping and I know at least with some of the interactions we’ve had with the AWS teams, they’re actively working and building that relationship more and more, giving access to those customer advisory boards, and all of them to have those direct calls with the executives. I don’t know whether you’ve seen that in your experience in helping some of these companies?

Corey: Have a different approach to it. It turns out when you’re super loud and public and noisy about AWS and spend too much time in Seattle, you start to spend time with those people on a social basis. Because, again, I’m obnoxious and annoying to a lot of AWS folks, but I’m also having an obnoxious habit of being right in most of the things I’m pointing out. And that becomes harder and harder to ignore. I mean, part of the value that I found in being able to do this as a consultant is that I begin to compare and contrast different customer environments on a consistent ongoing basis.

I mean, the reason that negotiation works well from my perspective is that AWS does a bunch of these every week, and customers do these every few years with AWS. And well, we do an awful lot of them, too, and it’s okay, we’ve seen different ways things can get structured and it doesn’t take too long and too many engagements before you start to see the points of commonality in how these things flow together. So, when we wind up seeing things that a customer is planning on architecturally and looking to do in the future, and, “Well, wait a minute. Have you talked to the folks negotiating the contract about this? Because that does potentially have bearing and it provides better data than what AWS is gathering just through looking at overall spend trends. So yeah, bring that up. That is absolutely going to impact the type of offer you get.”

It just comes down to understanding the motivators that drive folks and it comes down to, I think understanding the incentives. I will say that across the board, I have never yet seen a deal from AWS come through where it was, “Okay, at this point you’re just trying to hoodwink the customer and get them to sign on something that doesn’t help them.” I’ve seen mistakes that can definitely lead to that impression, and I’ve seen areas where they’re doing data is incomplete and they’re making assumptions that are not borne out in reality. But it’s not one of those bad faith type—

Micheal: Yeah.

Corey: —of negotiations. If it were, I would be framing a lot of this very differently. It sounds weird to say, “Yeah, your vendor is not trying to screw you over in this sense,” because look at the entire IT industry. How often has that been true about almost any other vendor in the fullness of time? This is something a bit different, and I still think we’re trying to grapple with the repercussions of that, from a negotiation standpoint and from a long-term business continuity standpoint, when your faith is linked—in a shared fate context—with your vendor.

Micheal: It’s in their best interest as well because they’re trying to build a diversified portfolio. Like, if they help 100 companies, even if one of them becomes the next Pinterest, that’s great, right? And that continued relationship is what they’re aiming for. So, assuming any bad faith over there probably is not going to be the best outcome, like you said. And two, it’s not a zero-sum game.

I always get a sense that when you’re doing these negotiations, it’s an all-or-nothing deal. It’s not. You have to think they’re also running a business and it’s important that you as your business, how okay are you with some of those premiums? You cannot get a discount on everything, you cannot get the deal or the numbers you probably want almost everything. And to your point, architecturally, if you’re moving in a certain direction where you think in the next three years, this is what your usage is going to be or it will come down to that, obviously, you should be investing more and negotiating that out front rather than managed NAT [laugh] gateways, I guess. So, I think that’s also an important mindset to take in as part of any of these negotiations. Which I’m assuming—I don’t know how you folks have been working in the past, but at least that’s one of the key items we have taken in as part of any of these discussions.

Corey: I would agree wholeheartedly. I think that it just comes down to understanding where you’re going, what’s important, and again in some cases knowing around what things AWS will never bend contractually. I’ve seen companies spend six weeks or more trying to get to negotiate custom SLAs around services. Let me save everyone a bunch of time and money; they will not grant them to you.

Micheal: Yeah.

Corey: I promise. So, stop asking for them; you’re not going to get them. There are other things they will negotiate on that they’re going to be highly case-dependent. I’m hesitant to mention any of them just because, “Well, wait a minute, we did that once. Why are you talking about that in public?” I don’t want to hear it and confidentiality matters. But yeah, not everything is negotiable, but most things are, so figuring out what levers and knobs and dials you have is important.

Micheal: We also found it that way. AWS does cater to their—they are a platform and they are pretty clear in how much engagement—even if
we are one of their top customers, there’s been many times where I know their product managers have heavily pushed back on some of the requests we have put in. And that makes me wonder, they probably have the same engagement even with the smallest of customers, there’s always an implicit assumption that the big fish is trying to get the most out of your public cloud providers. To your point, I don’t think that’s true. We’re rarely able to negotiate anything exclusive in terms of their product offerings just for us, if that makes sense.

Case in point, tell us your capacity [laugh] for x instances or type of instances, so we as a company would know how to plan out our scale-ups or scale-downs. That’s not going to happen exclusively for you. But those kind of things are just, like, examples we have had a chance to work with their product managers and see if, can we get some flexibility on that? For what it’s worth, though, they are willing to find a middle ground with you to make sure that you get your answers and, obviously, you’re being successful in your plans to use certain technologies they offer or [unintelligible 00:48:31] how you use their services.

Corey: So, I know we’ve gone significantly over time and we are definitely going to do another episode talking about a lot of the other things that you’re involved in because I’m going to assume that your full-time job is not worrying about the AWS bill. In fact, you do a fair number of things beyond that; I just get stuck on that one, given that it is but I eat, sleep, breathe, and dream about.

Micheal: Absolutely. I would love to talk more, especially about how we’re enabling our engineers to be extremely productive in this new world, and how we want to cater to this whole cloud-native environment which is being created, and make sure people are doing their best work. But regardless, Corey, I mean, this has been an amazing, insightful chat, even for me. And I really appreciate you having me on the show.

Corey: No, thank you for joining me. If people want to learn more about what you’re up to, and how you think about things, where can they find you? Because I’m also going to go out on a limb and assume you’re also probably hiring, given that everyone seems to be these days.

Micheal: Well, that is true. And I wasn’t planning to make a hiring pitch but I’m glad that you leaned into that one. Yes, we are hiring and you can find me on Twitter at twitter dot com slash M-I-C-H-E-A-L. I am spelled a bit differently, so make sure you can hit me up, and my DMs are open. And obviously, we have all our open roles listed on pinterestcareers.com as well.

Corey: And we will, of course, put links to that in the [show notes 00:49:45]. Thank you so much for taking the time to speak with me today. I really appreciate it.

Micheal: Thank you, Corey. It was really been great on your show.

Corey: And I’m sure we’ll do it again in the near future. Micheal Benedict, Head of Engineering Productivity at Pinterest. I am Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with a long rambling comment about exactly how many data centers Pinterest could build instead.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Amy

Amy (she/her) has spent the better part of the last 15 years in the tech start-up world, starting off as a front-end software engineer before transitioning into leadership. She has built and led teams across the software and product development spectrum, including web and mobile development, QA, operations and infrastructure, customer support, and IT.

These days, Amy is building the software engineering team at EdTech startup, Unicycle, and challenging the archetype of what a tech leader should be. She strives to be a real-life success story for other leaders who believe that safe, welcoming, and equitable environments can exist in tech.

Links:

  • Unicycle: https://www.unicycle.co
  • AmyChanta: https://twitter.com/AmyChanta

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored by our friends at Oracle Cloud. Counting the pennies, but still dreaming of deploying apps instead of "Hello, World" demos? Allow me to introduce you to Oracle's Always Free tier. It provides over 20 free services and infrastructure, networking databases, observability, management, and security.

And - let me be clear here - it's actually free. There's no surprise billing until you intentionally and proactively upgrade your account. This means you can provision a virtual machine instance or spin up an autonomous database that manages itself all while gaining the networking load, balancing and storage resources that somehow never quite make it into most free tiers needed to support the application that you want to build.

With Always Free you can do things like run small scale applications, or do proof of concept testing without spending a dime. You know that I always like to put asterisks next to the word free. This is actually free. No asterisk. Start now. Visit https://snark.cloud/oci-free that's https://snark.cloud/oci-free.

Corey: Writing ad copy to fit into a 30 second slot is hard, but if anyone can do it the folks at Quali can. Just like their Torque infrastructure automation platform can deliver complex application environments anytime, anywhere, in just seconds instead of hours, days or weeks. Visit Qtorque.io today and learn how you can spin up application environments in about the same amount of time it took you to listen to this ad.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. A famous quote was once uttered by Irena Dunn who said, “A woman without a man is like a fish without a bicycle.” Now, apparently at some point, people just, you know, looked at the fish without a bicycle thing, thought, “That was overwrought. We can do a startup and MVP it. Why do two wheels? We’re going to go with one.”

And I assume that’s the origin story of Unicycle. My guest today is Amy Chantasirivisal who is the Director of Engineering at Unicycle. Amy, thank you for putting up with that incredibly tortured opening. But that’s okay; we torture metaphors to death here.

Amy: [laugh]. Thank you for having me. That was a great intro.

Corey: So, you are, at the time of this recording at least, a relatively new hire to Unicycle, which to my understanding is a relatively new company. What do you folks do over there?

Amy: Yes, so Unicycle is not even a year old, so a company born out of the pandemic. But we are building a product to reimagine what the digital classroom looks like. The product itself was thought up right during a time during the pandemic when it became very clear how much students and teachers are struggling with converting their experience into online platforms. And so we are trying to just bring better workflows, more efficiency into that. And right now we’re starting with email, but we’ll be expanding to other things in the future.

Corey: I am absolutely the wrong person to ask about a lot of this stuff, just because my academic background, tortured doesn’t really begin to cover it. I handle academia about as well as I handled working for other people. My academic and professional careers before I started this place were basically a patchwork of nonsense and trying to pretend I was something other than I was. You, on the other hand, have very much been someone who’s legitimate as far as what you do and how you do it. Before Unicycle, you were the Director of Engineering at Wildbit, which is a name I keep hearing about and a bunch of odd places. What did you do there?

Amy: [laugh]. I will have to follow up and ask what the odd places are but—so I was leading a team there of engineers that were fully distributed across the US and also in Europe. And we were building an email product called Postmark, which some of your listeners might use, and then also a couple of other smaller things like People-First Jobs and Beanstalk—not AWS’s Beanstalk, but a developer repository and workflow tool.

Corey: Forget my listeners for a minute; I use Postmark. That’s where I keep seeing you on the invoices because it’s different branding. As someone who has The Duckbill Group, but also the Last Week in AWS things, it’s the brand confusion problem is very real. That does it. Sorry. Thank you for collapsing the waveform on that one. And of course, before that you were at PagerDuty, which is a company that most folks in the ops space are aware of, founded to combat the engineer’s true enemy: sleep.

Amy: Absolutely. It’s the product that engineers love to hate, but also can’t live without, to some degree. Or maybe they want to live without it, but uh… [laugh] are not able to.

Corey: So, I have a standing policy on this show of not talking to folks who are not wildly over-represented—as I am—and effectively disregarding the awesome stuff that they’ve done professionally in favor of instead talking about, “Wow, what’s it like not to be a white guy in the room? I can’t even imagine such a thing. It sounds hard.” However, in your case, an awful lot of the work you have done and are most proud of centers around DEI, diversity, equity, and inclusion. Tell me about that.

Amy: Absolutely. I would say that it’s the work that I’ve spent my time focusing on in recent years, but also that I’m still learning, right, and as someone who is Asian American, and also from a middle-class socioeconomic background, I have a bunch of privileges that I still have to unpack and that show up in the way that I work every day, as well. And so just acknowledging that, you know, while I spend a lot of time on DEI, still have just barely scratched the surface on it, really, in the grand scheme of things. But what I will say is that, you know, I’ve been really fortunate in my career in that I started in tech 15 or so years ago, and I started at a time when it wasn’t super hard for someone who has no CS degree to actually get into some sort of coding job. And so I fell into my first role; I was building HTML and CSS landing pages for a marketing team, for an ISP that was based in San Francisco.

So, I was cobbling together a bunch of technical skills, and I got better and better. And then I reached this point in my career where I didn’t really have a lot of mentors, and so I was like, “I don’t know what’s next for me.” But then I am also frustrated that it is so hard for our team to get things done. And so I took it upon myself to figure out Scrum and project management type of stuff for my team, and then made the jump into people management from there. So, people management and leadership through project management.

But when I look back on my career, I think about, “Oh, if I had a mentor, would that still have been my fate? Would I have continued down this track of becoming a very senior technical person and just doing that for my whole career?” Because letting go of the code was definitely a hard, hard thing. And I was lucky enough that I really did enjoy the people and the process side of all of this. And so [laugh] this relates to DEI in the fact that there’s research and everything that backs this up, but that women and women of color generally tend to get less mentorship overall and get less actionable feedback about their job performance.

And you think about how that potentially compounds over time, over the course of someone’s career and that may be one of the reasons why women and people of color get pushed out of tech because they’re not getting the support that they need, potentially. They’re not getting feedback, they’re not being advocated for in meetings, and then there’s also all the stuff that you can add on around microaggressions, or just aggressions period, potentially, depending on the culture of the team that you’re working on. And so all of those things compounded are the types of things that I think about now when I reflect on my own career and the types of teams that I want to be building in the future.

Corey: Back when I was stumbling my way through piecing my career together. I mean, as mentioned, I don’t have a degree; I don’t have a high school diploma, as it turns out, and—that was a surprise when I discovered midway through my 20s that the school I had graduated from wasn’t accredited—but I would tell stories, and I found ways to weasel my way through and I gave a talk right around 2015 or 2016, about, “Weasel Your Way to the Top: How to Handle a Job Interview,” and looking back, I would never give that talk again. I canceled it as soon as someone pointed out something that was only obvious in hindsight, that the talk was built out of things that had worked for me. And it’s easy to sit here and say that, well, I had to work for what I have; none of this was handed to me. And there’s an element of truth to that, except for the part where there was nothing fighting against me as I went.

There was not this headwind of a presumed need for me to have to prove myself; I am presumed competent. I sometimes say that as a white guy in tech, my failure mode is a board seat and a book deal, and it’s not that far from wrong. It takes, I guess, a lot of listening and a lot of interaction with folks from wildly different backgrounds before you start to see some of these things. It takes time. So, if you’re listening to this, and you aren’t necessarily convinced that this might be real or whatnot, talk less, listen more. There are a lot of stories out there in the world that I think that it’s not my place to tell but listen. That’s how I approach it.

What’s interesting about your pathway into management is it’s almost the exact opposite of mine, where I was craving novelty, and okay, I wanted to try and managing a team of people. Years later, in hindsight—I’m not a good manager and I know that about myself, and I explicitly go out of my way these days to avoid managing people wherever possible, for a variety of reasons, but at the time, I didn’t know. I didn’t know that. I wanted to see how it went.

First, I had to disabuse myself of this notion that, oh, management is a promotion. It’s not. It’s an orthogonal skill.

Amy: Yes.

Corey: The thing I really learning—management or not—now, is that the higher in the hierarchy you rise, if you want to view it that way, the less hands-on work you do, which means everything that you are responsible for that—and oh, you are responsible—isn’t something you can jump in and do yourself. You can only impact the outcome via influence. And that was a hard lesson to learn.

Amy: Right. And there are some schools of thought, though, where you can affect the outcome by control. And that’s not what I’m about. I think I’m more aligned with what you’re saying in terms of, it’s really the influence and the ability to clear the way for people who are smarter than you to do the things that they need to do. Just get out of their way, and remove the roadblocks, and just help give them what they need. That’s really, sort of like, my overall approach. But I know that there are some folks out there who lead the opposite way of, “It’s my way, and I’m going to dictate how things should be done, and really you’re here to take and follow orders.”

Corey: It’s always fun interviewing people to manage teams. “So, why do you want to be a manager?” It’s, “Oh, I want to tell people what to do.” And I have to say that as an interviewer, there is nothing that takes the pressure off nearly as well as a perfectly wrong answer. And, yes, that at least to my world, is a perfectly wrong answer to this. There aren’t that many pass-fail questions, but you can fail any question if you try hard enough.

Amy: [laugh]. Oh, gosh, yeah, it’s true. But also, at the same time, I would say that there are organizations that are built that way. Because—all it takes is the one person who wants to tell people what to do, and then they start a company, and then they hire other people who want to tell people what to do. And so there are ways where organizations like that exist and come into being even today, I would say.

Corey: The question that I have for you about engineering leadership is, back when I was an engineer, and thinking, all right, it’s time for me to go ahead and try being a manager—let’s be clear, I joke about it, but the actual reason I wanted to try my hand at management was that I found people problems more interesting than computer problems at that point. I still do, but these days, especially when it comes to, you know, cloud services marketing and such, yeah, generally, the technical problems are, in fact, people problems at their core. But talking to my manager friends of how do I go and transition from being an engineer into being a manager, the universal response I got at the time was, “Ehh, I don’t know.” Every person I knew who’d had made that transition was in the right place at the right time, and quote-unquote, “Got lucky.”

Amy: Absolutely.

Corey: And then once they had management on their resume, then they could go and transition back to being an IC and then to management again. But it’s that initial breakthrough that becomes a challenge.

Amy: Absolutely. And I fell into it as well. I mean, I got into it, partially for selfish reasons because I was, an IC, I was doing development work, and I was frustrated, and I had teammates who were coming to me and they were frustrated about how hard it was for us to get our work done, or the friction involved in shipping code. And so I took it upon myself to say, “I think I see a pattern about why this is happening, and so I will try to solve this problem for the team.” And so that’s where the Agile and Scrum thing come in, and the project management side.

And then, when I was at this company—this was One Kings Lane; this was, like, the heyday of flash sales websites and stuff like that, so it was kind of a rocket ship at that time—and because we were also growing so fast and I was interviewing folks as well, I just fell into this management role of, “Well, if I’m interviewing these people, then I guess I should be [laugh] managing them, too.” And that happens for so many people, similar stories of getting into management. And I think that’s where it starts to go wrong for a lot of organizations because, like you said, it’s not an up-leveling; it’s a changing of your role, and it requires training and learning and figuring out how to be effective as a manager. And a lot of people just stumble their way through it and make a lot of mistakes—myself included—through that process.

And that becomes really troubling knowing that you can make these really big mistakes, but these mistakes that you make don’t affect just yourself. It’s the careers of the people that you manage as well and sort of where they’re headed in their lives. And so it’s troubling to think that most leaders that are out there today have not received any sort of training on how to be a good manager and how to be effective as a manager.

Corey: I would agree with that wholeheartedly. It seems that in many cases, companies take the best engineer that they have on their team and promote them to manager. It’s brilliant in some respects in just how short-sighted it is. You are taking a great engineer and trading them for a junior and unproven manager, and hoping for the best. And there is no training on any of these things, at least—

Amy: Right.

Corey: —not the companies that I ever worked at. Of course, there are ways you can learn to be a better manager; there are people who specialize in exactly this. There are companies that do exactly this. But tech has this weird thing where it just tries to solve itself from first principles rather than believing for a minute that someone might possibly have prior experience that could be useful for these things. And—

Amy: Absolutely.

Corey: —that was a challenge. I had a lot of terrible managers before I entered management myself, and I figured, ah, I’ll do the naive thing and I’m just going to manage based upon doing the exact opposite of what those terrible managers all did. And I got surprisingly far with it, on some level. But you don’t see the whole picture when you’re an individual contributor who’s writing code—crappy in my case—most of the time, and then only seeing the aspects of your manager that they allow you to see. They don’t share—if they’re any good—the constraints that they have to deal with, that they’re managing expectations around the team, conflicting priorities, strategic objectives, et cetera because it’s not something that gets shown to folks. So—

Amy: Absolutely.

Corey: —if you bias for that, in my experience you become an empathetic manager to the people on your team, but completely ineffective at managing laterally or upwards.

Amy: Mm-hm, absolutely. And you know, I’m exploring this idea of further. Being at a very small company, I think allows me to do that. And exploring this idea of, does it have to be that way? Can you be transparent about what the constraints are as a leader while still caring for your team and supporting them in the ways that they need and helping them grow their careers and just being open about one of the challenges that you have in building the company?

And I don’t know, I feel like I have some things to prove there, but I think it’s possible to achieve some sort of balance there, something better or more beyond just what exists now of having that entire leadership layer typically be very opaque and just very unclear why certain decisions are made.

Corey: The hard part that extends that these to me beyond that is it’s difficult to get meaningful feedback, on some level, when you’re suddenly thrust into that position. I also, in hindsight, realize that an awful lot of those terrible managers that I had weren’t nearly as terrible as I thought they were. I will say that being on the other side of that divide definitely breeds empathy. Now that I’m the co-owner of The Duckbill Group, and we’re building out a leadership team and the rest, hiring managers of managers is starting to be the sort of thing that I have to think about.

It’s effectively, how do I avoid inadvertently doing end-runs around people? And oh, I’m just going to completely undermine a manager by reaching out to one of their team and retasking them on something because obviously whatever I have in mind is much more important. What could they possibly be working on that’s better than the Twitter shitpost I’m borrowing them to help out with? Yeah, you learn a lot by getting it wrong, and there becomes a power imbalance that even if you try your best to ignore it—which you should not—I assure you, the person who has less power in that relationship cannot set that aside. Even when I have worked with people I consider close friends, that friendship gained some distance during the duration of their employment because there has to be that professional level of separation. It’s a hard thing to learn.

Amy: It’s a very hard line to walk in terms of recognizing the power that you have over someone’s career and the power over, you know, making decisions for them and for the team and for the company, and still being empathetic towards their personal needs. And if they’re going through a tough time, but then you also know from a business perspective that X, Y, or Z needs to happen, and how do you push but not push too hard, and try to balance needs of people who are humans and have things that happen and go on sometimes, and the fact that we work in a capitalist society and we still need to make money to make the business run. And that’s definitely one of the hardest things to learn, and I am still learning. I definitely don’t have that figured out, but I err on the side of, let’s listen to what people are saying because ultimately, I’m not going to be the one to write the code. I haven’t done that in years, and also I would probably suck at it now. And so it behooves leaders to listen to the people who were doing the work and to try, to the best of their abilities in whatever role whether that’s exec-level leadership or mid-level… sort of like, middle management type of stuff to do what is in your power to help set them up to succeed.

Corey: I want to get back a little bit to the idea of building diverse teams. It’s something that you spend an inordinate amount of time and effort on. I do too. It’s one of those areas where it’s almost fraught to talk about it because I don’t want to sound like I’m breaking my arm by patting myself on the back here. I certainly have a hell of a lot to learn, and mostly—and I’m ashamed to admit this—I very often learn only by really putting my foot in it sometimes. And it’s painful, but that is, I think, a necessary prerequisite for growth. From your perspective, what is the most challenging part of building diverse teams?

Amy: I think it’s that piece that you said of making the mistakes or just putting yourself in a position where you are going to be uncomfortable. And I think that a lot of organizations that I’ve been in talk about DEI on a very surface level in terms of, “Oh, well, you know, we want to have more candidates from diverse backgrounds in our pipelines for hiring,” and things like that. But then not really just thinking about, but how do we work as a team in a way that potentially makes retention of those folks a lot harder? And for myself, I would say that when I was earlier on in all of this in my learning, I would say that I was able to kickstart my learning by thinking about my own identity, the fact that I was often the only Asian person on my team, the only woman on my team, and then more recently, the only mom on my team. And that has happened to me so many times in my career. More often than not.

And so being able to draw on those experiences and those feelings of oh, okay, no one wants to hear about my kid because everyone else is, you know, busy going out to drink or something on the weekends. And like that feeling of, you know, that not belonging, and feeling of feeling excluded from things, and then thinking about how then this might manifest for folks with different identities for myself. And then going there and learning about it, listening, doing more listening than talking, and yeah, and that’s, that’s really just been the hardest part of just removing myself from that equation and just listening to the experiences of other people. And it’s uncomfortable. And I think a lot of people are—you have to be in the right mindset, I guess, to be uncomfortable; you have to be willing to accept that you will be uncomfortable. And I think a lot of folks maybe are not ready to do that on a personal level.

Corey: The thing that galls me the most is I do try on these things, and I get it wrong a fair bit. And my mistakes I find personally embarrassing, and I strive not to repeat them. But then I look around the industry—and let’s be clear, a lot of this is filtered through the unhealthy amount of time I spend on Twitter—but it seems that I’m trying and I’m failing and attempting to do better as I go, and then I see people who are just, “Nope. Not at all. In fact, we’re not just going to lean into bias, we’re going to build a startup around it.”

And I look at this and it’s at some level hard to reconcile the fact that… at first, that I’m doing badly at all, which is the easy cop-out of, “Oh, well, if that is considered acceptable on some level, then I certainly don’t even have to try,” which I think is a fallacy. But further it’s—I have to step beyond myself on that and just, I cannot fathom how discouraging that must be, particularly to people who are early in their careers because it looks like it’s just a normal thing that everyone thinks and does that just someone got a little too loud with it. And it’s abhorrent. And if people are listening to this and thinking that is somehow just entrenched, and normalized, and everyone secretly thinks that… no. I assure you it is not something that is acceptable, even in the quote-unquote, “Private white dude who started companies” gathering holes. Yeah, people articulating sentiments like that suddenly find themselves not welcome there anymore, at least in every one of those types of environments I’ve ever found myself in.

Amy: Yeah, the landscape is shifting. It’s slow, but it is shifting. And, myself on Twitter, like, I do a lot of rant-y stuff too sometimes, but despite all of that, I feel like I am ultimately an optimist because I have to be. Otherwise, I would have left tech already because every time I am faced with a job search for myself, I’m like, “Should I—is this it? Am I done in tech? Do I want to go do something else? Am I going to finally go open that bakery that I’ve always wanted to open?” [laugh].

And so… I have to be an optimist. And I see that—even in the most recent job search I’ve done—have seen so many new founders and new CEOs, really, with this mindset of, “We want to build a diverse team, but we’re also doing it—and we’re using diversity as a foundation for what we want to build; it’s part of our decision-making process and this is how we’re going to hold ourselves accountable to it.” And so it is shifting, and while there are those bad actors out there still, I’m seeing a lot of good in the industry now. And so that’s why I stick around; that’s why I’m still here.

Corey: I want to actually call something out as concrete here because it’s easy for me to fall into the trope of just saying vague things. I’ll be specific about something, give us a good example. We’ve done a decent job, I think, of hiring a diverse team, but—and this is a problem that I see spread across an awful lot of companies—as you look at the leadership team, it gets a lot wider and a lot more male. And that is an inherent challenge. In our particular case, my business partner is someone who I’ve been close friends with for a decade.

I would not be able to start a business with someone I didn’t have that kind of relationship with just because your values have to be aligned or there’s trouble down the road. And beyond that, it winds up rapidly, on some level, turning into what appears to be a selection bias. When you’re trying to hire senior leaders, for example, there’s a prerequisite to being a senior leader, which is embodied in the word senior, which implies tenure of having spent a fair bit of time in an industry that is remarkably unfriendly in a lot of different ways to a lot of different people. So, there’s a prerequisite of being willing to tolerate the shit for as long as it takes to get to that level of seniority, rather than realizing at any point as any of us can, there are easier jobs that don’t have this toxicity inherent to them and I’ll go do that instead. So, there’s a tenure question; there’s a survivorship bias question.

And I don’t have the answers to any of this, but it’s something that I’m seeing, and it’s one of those once you see it, you can’t unsee it any more moments. At least for me.

Amy: Yeah, absolutely.

Corey: Please tell me I’m not the only person who see [laugh]—who is encountering these problems. Like, “Wow, you just sound terrible.” Which might very well be a fair rejoinder here. I’m just trying to wrap my head around how to think about this properly.

Amy: Yeah. I mean, this is why I was saying that I am very optimistic about [laugh] new companies that are coming—like, up-and-coming these days, new startups, primarily, because you’re right that a lot of people just end up quitting tech before they get to that point of experience and seniority, to get into leadership. I mean, obviously, there’s a lot of bias and discrimination that happens at those leadership levels, too, but I will say that, you know, it’s both of those things. There are also more things on top of that. But this is why I’m like, so excited to see people from diverse backgrounds as founders of new companies and why I think that being able to be in a position to potentially either help fund, or advocate, or sponsor, or amplify those types of orgs, I think is where the future is that because ultimately, I think a lot of the established companies that are out there these days, it’s going to be really hard for them to walk back on what their leadership team looks like now, especially if it is a sizable leadership team and they’re all white men.

Corey: Yeah. I’m going to choose to believe we say sizable leadership team that it’s also not—we’re talking about the horizontal scaling that happens to some of us, especially during a pandemic as we continue to grow into our seats. You’re right, it’s a problem as well, where you can cut a bit of slack in some cases to small teams. It’s, “Okay, we don’t have any Black employees, but we’re three people,” is a lot more understandable-slash-relatable than, “We haven’t hired any Black people yet and we’re 3000 people.” One of those is acceptable—or at least understandable, if not acceptable—the other is just completely egregious.

Amy: Yes. And I think then the question that you have to ask if you’re looking at, you know, a three-person company, or [laugh] I guess, like in my case, I was looking at the seven-person company, is that, “Okay. There are currently no Black people on your team. And why is that?” And then, “What are you doing to change that? And how are you going to make sure that you’re holding ourselves accountable to it?”

Because I think it’s easy to say, “Oh, you know, the first couple of hires were people we just worked with in the past, and they just happened to, you know, look like us and whatnot.” And then you blink becau—and you do that a handful of times, and you blink, and then suddenly you have a team of 25 and there are no people of color on your team. And maybe you have, like, one woman on the team or something. And you’re like, “Huh. That’s strange. I guess we should think about this and figure out what we can do.”

And then I think what ends up happening at that point is that there are so many already established behaviors, and cultural norms, and things like that, that have organically grown within a team that are potentially not welcoming towards people from different backgrounds who have different backgrounds. So, you go and attempt to hire someone who is different, and they come in, and they’re just sort of like, “This is how you work? I don’t feel like I belong here.” And then they don’t stay, and then they leave. And then people sit there and scratch their heads like, “Oh, what did we do wrong?” And, “I don’t get it.”

And so there’s this conversation, I think, in the industry of like, “Oh, it’s a pipeline problem, and if we were just able to hire a lot of people from diverse backgrounds, the problem is solved.” Which really isn’t the case because once people are there and at your company, are they getting promoted at the same rate as white men? Are they staying with the company for as long? And who’s in leadership? And how are you working to break down the biases that you may have?

All those sorts of things, I think, generally are not considered as part of all of this DEI work. Especially when, in my experience in startups, the operational side of all that is so immature a lot of the times, just not well developed that deeper thought process and reflection doesn’t really happen.

Corey: This episode is sponsored in part by something new. Cloud Academy is a training platform built on two primary goals. Having the highest quality content in tech and cloud skills, and building a good community the is rich and full of IT and engineering professionals. You wouldn’t think those things go together, but sometimes they do. Its both useful for individuals and large enterprises, but here's what makes it new. I don’t use that term lightly. Cloud Academy invites you to showcase just how good your AWS skills are. For the next four weeks you’ll have a chance to prove yourself. Compete in four unique lab challenges, where they’ll be awarding more than $2000 in cash and prizes. I’m not kidding, first place is a thousand bucks. Pre-register for the first challenge now, one that I picked out myself on Amazon SNS image resizing, by visiting cloudacademy.com/corey. C-O-R-E-Y. That’s cloudacademy.com/corey. We’re gonna have some fun with this one!

Corey: I do my best to have these conversations in public as frequently as is practical for me to do, just because I admit, I get things wrong. I say things that are wrong and I’m doing a fair bit of learning in public around an awful lot of that. Because frankly, I can withstand the heat, if it comes down to someone on Twitter gets incredibly incensed by something I’ve said on this podcast, for example. Because it isn’t coming from a place of ill intent when someone accuses me of being ableist or expressing bias. My response is generally to suppress the initial instinctive flash of defensiveness and listen and ask.

And that is, even if I don’t necessarily agree with what they’re saying after reflection, I have to appreciate on some level the risk-taking inherent in calling someone out who is in my position where, if I were a trash fire, I could use the platform to turn it into, “All right. Now, let’s go hound the person that called me out.” No. I don’t do that, full stop. If I’m going to harass people, it’s going to be—not people, despite what the Supreme Court might tell us—but it’s going to be a $2 trillion company—one in particular—because that’s who I am and that’s how I roll.

Whenever I get a DM—which I leave open because I have the privilege to do that—from folks who are early career who are not wildly over-represented, I just have to stop and marvel for a minute at the level of risk-taking inherent to that because there is risk to that. For me, when I DM people, the only risk I feel like I’m running at any given point is, “Are they going to think that I’m bothering them? Oh, the hell with it. I’m adorable. They’ll love me.” And the fact that I’m usually right is completely irrelevant to that. There’s just that sense of I don’t really risk a damn thing in the grand scheme of things compared to the risk that many people are taking just living who they are.

Amy: Yeah. And someone DMs you and you suppress that initial sort of defensiveness: I would say that that is an underrated skill. [laugh].

Corey: Well, a DM is a privilege, too. A call in—

Amy: Yes.

Corey: —is deeply appreciated; no one owes it to me. I often will get people calling me out on Twitter and I generally stop and think about that; I have a very close circle of friends who I trust to be objective on these things, and I’ll ask them, “Did I get this wrong?” And very often the answer is yes. And, “Well, I thought the joke was funny and I spent time building it.” “Yeah, but if people hear a joke I’m making and feel bad about it, then is it really that good of a joke or should I try harder?” It’s a process, and I look back at who I was ten years ago and I feel a sense of shame. And I believe that if anyone these days doesn’t, either they were effectively a saint, or they haven’t grown.

Amy: Yes.

Corey: And that’s my personal philosophy on this stuff, anyway.

Amy: Yeah, absolutely. And that growth is so important. And part of that growth really is being able to suppress your desire to make it about you, [laugh] right? That initial, “Oh, I did something bad,” or, “I’m a horrible person because I said this thing,” right? It’s not about you, there’s, like, the impact that you had on someone else.

And I’ve been giving this some thought recently, and I—you know, I also similarly have a group of trusted friends who I often talk about these things with, and you know, we always kind of check ourselves in terms of, did we mess something up? Did we, you know, put our foot in our mouths? Stuff like that. And think what it really comes down to is being able to say, “Maybe I did something wrong and I need to suppress that desire to become defensive and put up walls and guard and protect myself from feeling vulnerable, in order to actually learn and grow from this experience.”

Corey: It’s hard to do, but it’s required because I—

Amy: Extremely, yes.

Corey: —used to worry about, “Ohh, what if I get quote-unquote, ‘canceled?’” well, I’ve done a little digging into this and every notable instance of this I can find is when someone is called out for something crappy, they get defensive, and they double-down and triple-down and quadruple-down, and they keep digging a hole nice and deep to the point where no one with a soul can really be on their side of this issue, and now they have a problem. I have never gotten to that point because let’s be honest with you, there are remarkably few things I care that passionately about that I’m going to pick those fights publicly. The ones that I do, I am very much on the other side [laugh] of those issues. That has not been a realistic concern.

I used to warn every person here before I hired them—to get this back to engineering management—that there was a risk that I could have a bad tweet and we don’t have a company anymore. I don’t give that warning anymore because I no longer believe that it’s true.

Amy: Mm-hm. Mm-hm. I also wonder about, in general, because of the world that we live in, and our history with white supremacy and oppression and all those things, I also wonder if this skill of being able to self-reflect and be uncomfortable and manage your own reaction and your emotions, I wonder if that’s just a thing that white people generally haven’t had a lot of practice for because of the inherent privileges that are afforded to white people. I wonder if a lot of this just stems from the fact that white people get to navigate this world and not get called out, and thus don’t have this opportunity to exercise this skill of holding on to that and listening more than talking.

Corey: Absolutely agree. And it gets piled on by a lot of folks, for example—I’ll continue to use myself as an example in this case—I live in San Francisco. I would argue that I’m probably not, “In tech,” quote-unquote, the way that I once was, but I’m close enough that there’s no discernible difference. And my social circle is as well. Back before I entered tech, I did a bunch of interesting jobs, telemarketing to pay the bills, I was a recruiter for a while, I worked construction a couple of summers.

These days, everyone that I engage with for meaningful periods of time is more or less fairly tech adjacent. It really turns into a one-sided perspective. And I can sit here and talk about what folks who are not living in the tech bubble should be doing or how they should think about this, but it’s incredibly condescending, it’s incredibly short-sighted, and fails to appreciate a very different lived experience. And I can remind myself of this now, but that lack of diversity and experience is absolutely something where it feels like the tech bubble, especially for those folks in this bubble who look a lot like me, it is easy to fall into a pattern of viewing ourselves as the modern aristocracy where we deserve the nice things that we have, and the rest. And that’s a toxic pattern. It takes vigilance to avoid it. I’m not saying I get it right all the time, by a landslide, but ugh, the perils of not doing that are awful.

Amy: Agreed. And it shows up, you know, getting back to the engineering manager and leadership and org building piece of things, that shows up even in the way that we talk about career development and career ladders, for those of us in tech, and software engineering specifically for me, where we’ve kind of like come up with all these matrices of job levels, and competencies, all that, and humans just are so vastly different. Every person is an individual, and yet we talked about career ladders and how to advance your career in this two-dimensional matrix. And, like, how does that actually work, right?

And I’ve seen some good career ladders that account for a larger variety of competencies than just, “Can you code?” And, “What are your system design skills?” And, “Do you understand distributed systems?” And so on and so forth, but I think a lot gets left behind and gets left on the table when it comes to thinking about the fact that when you get a group of people together working on some sort of common cause or a product, that there’s so much more to the dynamic than just the writing of the code. It’s how do you work with each other? How do you support each other? How do you communicate with each other? And then all my glue work—that is what I call it—like, the glue work that goes into a successful team and building products, a lot of that is just not captured in the way that we talk about career development for folks. And it’s just incredibly two-dimensional, I think.

Corey: One last question that I have for you before we wrap the episode here is, you spend a lot of time focusing on this, and I have some answers, but I’m very interested to hear yours instead because I assure you, the world hears enough from me and people who look like me, what is the biggest mistake that you see companies making in their attempts to build diverse teams?

Amy: I would say that there’s two major things. One is that there have been a lot of orgs in my own past that think about diversity, equity, inclusion as a program and not a mindset that everyone should be embracing. And that manifests itself into, sort of like, this secondary problem of stopping at the D part of D, E, and I. That’s the whole, “We’re going to hire a bunch of people from different backgrounds and then just we’re going to stop with that because we’ve solved the problem.” But by not adopting that mindset of the equity, the inclusion, and also the welcoming and the belonging piece of things internally, then anyone that you hire who comes in from those marginalized or minority backgrounds is not going to want to stay long-term because they don’t feel like they fit in, they don’t feel like they belong.

And so, it becomes this revolving door of you hire in people and then those people leave after some amount of time because they’re not getting what they need out of either the role or for themselves personally in terms of just emotional support, even. And so I would say that’s the problem that I see is not a numbers game—although the metrics and the numbers help hold you accountable—but the metrics and the numbers are not the end goal. The end goal is really around the mindset that you have in building the org and the way that people behave. And the way that you work together is really core to that.

Corey: What I tend to see on the other side is the early intake funnels. People will reach out to me sometimes, “Hey, do you know any diverse speakers we can hire to do a speaking engagement here?” It doesn’t… work that way. There’s a lot more to it than that. It is not about finding people who check boxes, it is not about quote-unquote, “Diversity hires.”

It’s about—at least in my experience—structuring job ads, for example, in ways that are not coded—unconsciously in most cases, but ehh—that are going to resonate towards folks who are in certain cultures and not in others. It’s about being more equitable. It’s about understanding that not everyone is going to come across in a job interview as the most confident person in the room. Part of the talk that I gave on how to handle job interviews, there was a strong section in it on salary negotiation. Well, turns out when I do it, I’m an aggressive hard-charger and they like that, whereas if someone who is not male does that, well, in that case, they look like they’re being difficult and argumentative and pushy and rising above their station. It was awful.

One of the topics I’m most proud of was the redone version of that talk that I gave with a friend, Sonia Gupta, who has since left tech because of how shitty it is, and that was a much better talk. She was a former attorney who had spent time negotiating in much higher-stakes situations.

Amy: Yeah.

Corey: And it was terrific to see during the deconstruction and rebuilding of that talk, just how much of my own unconscious bias had crept in. It’s, again, I look back at the early version of those talks and I’m honestly ashamed. It wasn’t from ill will, but it’s always impact over intent as far as how this has potentially made things worse. It’s, if nothing else, if I don’t say the right things when I should speak up, that’s not great, but I always prefer that to saying things that are actively harmful. So—

Amy: Absolutely.

Corey: —it’s hard. I deserve no sympathy for this, to be clear. It is incumbent upon all of us because again, as mentioned, my failure mode is a non-issue in the world compared to the failure mode for folks for against whom the deck has been stacked unfairly for a very long time. At least, that’s how I see it.

Amy: Right. And that’s why I think that it’s important for folks who are in positions of power to really reflect on—even operationally, right, you were mentioning your job ads, and how to structure that to include more inclusive language, and just doing that for everything, really, in the way that you work. How do decisions get made? And by whom? And why? How do you structure things like compensation? Even, like, how do you do project planning, right?

Even in my own reflections, now when I think back towards Scrum and Agile and all of that, I think that the base foundation of all of that was like was good, but then ultimately the implementation of how that works at most companies is problematic in a lot of ways as well. And then to just be able to reflect and really think about all of your processes or policies—all of that—and bring that lens of equity, really, equity and inclusion to those things, and to really dig deep and think about how those things might manifest and affect people from different backgrounds in different ways.

Corey: So, before we wrap, something that I think you… are something of an empathetic party on is when I see companies in the space who are doing significant DE&I initiatives, it seems like it’s all flash; it feels like it’s all sizzle, no steak to appropriate a phrase from the country of Texas. Is that something that you see, too?

Amy: I do think that it is pretty common, and I think it’s because that’s… that’s the easy route. That’s the easy way to do it because the vanity metrics, and the photo of the team that is so diverse, and all these things that show up on a marketing website. I mean, there—it’s, like, a signal for someone, potentially, who might be considering a job at your company, but ultimately the hard work that I feel like is not happening is really in that whole reflecting on the way you do business, reflecting on the way that you work. That is the hard work and it requires a leadership team to prioritize it, and to make time for it, and to make it really a core principle of the way that you build an org., and it doesn’t happen enough, by far, in my opinion.

Corey: It feels like it’s an old trope of the company that makes a $100,000 donation and then spends $10 million dollars telling the world about it, on some level. It’s about, “Oh, look at us, we’re doing good things,” as opposed to buckling down and doing the work. Then the actual work falls to folks who are themselves not overrepresented as unpaid emotional labor, and then when the company still struggles with diversity issues, those people catch the blame. It’s frustrating.

Amy: Yeah. And as an organization, if you have the money to donate somewhere, that’s great, but it can’t just stop at that. And a lot of companies will just stop at that because it’s the optics of, “Oh, well, we spent x millions of dollars and we’ve helped out this nonprofit or this charity or whatnot.” Which is great that you’re able to do that, but that can’t be it because then ultimately, what you have internally and within your own company doesn’t improve for people from those backgrounds.

Corey: I want to thank you for taking so much time to chat with me about these things. Some of these topics are challenging to talk about and finding the right forum can be difficult, and I’m just deeply appreciative that you were able to clear enough time to have that chat with me today.

Amy: Yeah, thank you for having me. I mean, I think it’s important for us to recognize, even between the two of us that, I mean, obviously, you as a white man have benefited a lot in this space, and then even myself as, you know, that model minority whole thing, but growing up very adjacent to white people and just being ingrained in that culture and raised in that culture, you know, that we have those privileges and there’s still parts of the conversation, I think, that are not captured by [laugh] by the two of us are the nuances as well, and so just recognizing that. And it’s just a learning process. And I think that everyone could benefit from just realizing that you’ll never know everything. And there’s always going to be something to learn in all of this. And yes, it is hard, but it’s something that is worthwhile to strive for.

Corey: Most things worthwhile are. If people want to learn more about who you are, how you think about these things, potentially consider working with you, et cetera. Where can they find you?

Amy: So, I am on Twitter. I am the queen of very, very long threads, I should just start a blog or something, but I have not. But in any case, I’m on Twitter. I am AmyChanta, so @A-M-Y-C-H-A-N-T-A.

Our website is unicycle.co, if you’re thinking about applying for a role, and working with me, that would be awesome. Or just, you know, reach out. I’d also just love to network with anyone, even if there’s not an open position now. I just, you know, build that relationship and maybe there will be in the future. Or if not at Unicycle, then somewhere else.

Corey: And we will, of course, put links to that in the [show notes 00:48:13]. Thank you so much, once again. I appreciate your time.

Amy: Thanks for having me.

Corey: Amy Chantasirivisal, Director of Engineering at Unicycle. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with a comment pointing out that it’s not about making an MVP of a bicycle that turns into a unicycle so much as it is work-life balance.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Dann

Dann Berg is a Senior CloudOps Analyst at Datadog, and has nearly a decade of experience working in the cloud and optimizing multi-million dollar budgets. He is also an active member of the larger technical community, hosting the monthly New York City FinOps Meetup, and has been published multiple times in places such as MSNBC, Fox News, NPR, and others. When he’s not saving companies millions of dollars, he’s writing plays, and has had two full-lengh plays produced in New York City and China.

Links:

  • Datadog: https://www.datadoghq.com
  • Personal Website: https://dannb.org
  • LinkedIn: https://www.linkedin.com/in/dannberg/
  • Twitter: https://twitter.com/dannberg
  • Monthly newsletter: https://dannb.org/newsletter/
  • Previous SITC episode with Dann Berg, Episode 51: https://www.lastweekinaws.com/podcast/screaming-in-the-cloud/episode-51-size-of-cloud-bill-not-about-number-of-customers-but-number-of-engineers-you-ve-hired/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Vultr. Spelled V-U-L-T-R because they’re all about helping save money, including on things like, you know, vowels. So, what they do is they are a cloud provider that provides surprisingly high performance cloud compute at a price that—while sure they claim its better than AWS pricing—and when they say that they mean it is less money. Sure, I don’t dispute that but what I find interesting is that it’s predictable. They tell you in advance on a monthly basis what it’s going to going to cost. They have a bunch of advanced networking features. They have nineteen global locations and scale things elastically. Not to be confused with openly, because apparently elastic and open can mean the same thing sometimes. They have had over a million users. Deployments take less that sixty seconds across twelve pre-selected operating systems. Or, if you’re one of those nutters like me, you can bring your own ISO and install basically any operating system you want. Starting with pricing as low as $2.50 a month for Vultr cloud compute they have plans for developers and businesses of all sizes, except maybe Amazon, who stubbornly insists on having something to scale all on their own. Try Vultr today for free by visiting: vultr.com/screaming, and you’ll receive a $100 in credit. Thats v-u-l-t-r.com slash screaming.

Corey: This episode is sponsored by our friends at Oracle Cloud. Counting the pennies, but still dreaming of deploying apps instead of "Hello, World" demos? Allow me to introduce you to Oracle's Always Free tier. It provides over 20 free services and infrastructure, networking databases, observability, management, and security.

And - let me be clear here - it's actually free. There's no surprise billing until you intentionally and proactively upgrade your account. This means you can provision a virtual machine instance or spin up an autonomous database that manages itself all while gaining the networking load, balancing and storage resources that somehow never quite make it into most free tiers needed to support the application that you want to build.

With Always Free you can do things like run small scale applications, or do proof of concept testing without spending a dime. You know that I always like to put asterisks next to the word free. This is actually free. No asterisk. Start now. Visit https://snark.cloud/oci-free that's https://snark.cloud/oci-free.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. If there’s one thing that I love, it is certainly not AWS billing, but for better or worse, that’s where my career has led me. Way back in Episode 51, I had Dann Berg, the CloudOps analyst at Datadog. And now he’s back for more. Things have changed. He’s now a senior CloudOps analyst, and I’m hoping my jokes have gotten better. Dann, thanks for being bold enough to come out and find out.

Dann: Yeah. I’m excited to see if these jokes have gotten better. That’s the main reason for coming back.

Corey: Exactly. Because it turns out that death, taxes, and AWS bills are the things that are inevitable and never seem to change.

Dann: Yeah. They just keep coming. They never stop, and they’re always slightly different than you expect. I guess, just like death and taxes.

Corey: So, when we spoke back in, I want to say 2019 is when it aired, so probably that—ish—is when we had the conversation, if not a little bit before that, you were effectively a team of one, and as mentioned, had the CloudOps analyst title. Now, you’re a senior CloudOps analyst, which I assume just means you’re older. Is the team larger as well? What does that process look like? How has it evolved in the last couple years?

Dann: Yeah, it’s been interesting, especially being a single organization and that organization being Datadog, that to be able to grow the team a little bit. So, as you said, it was just me. Now, it’s a total of four people, including myself, so three others. And, yeah, it’s been interesting just in terms of my own professional development, being able to identify what needs to be done, how much capacity I have, and being able to grow it over time, especially in this fairly new space of being specifically focused on cloud cost billing. So, kind of that bridge between engineering and finance, which itself is kind of a fairly new space, still.

Corey: It is. And my favorite part of having these conversations with folks who have no idea what this space is, is learning—when I was starting—out how to talk about this in a way that didn’t lead down weird paths. It’s, “Oh, you save money on Amazon bills? Can you help me save money on socks?” It’s like, “No. Well, yes. Get the Prime card, it gives you 5% off. But no.” And yeah, I talk about camelcamelcamel and other ways of working around the retail side, but that’s not really what I do.

It’s similar to back when I was doing SRE-style work. I made it a point never to talk about being someone involved in working in tech, or suddenly you’re the neighborhood printer repair person. Similarly, you have, I guess, gone in a strange direction because you weren’t, to my recollection, someone who had a strong SRE background. That’s not where you came from in the traditional sense, is it?

Dann: No, not an SRE background at all. Yeah, I mean, it’s really interesting. So, talking about this space, I mean, people are calling it a lot of different things, cloud economics, the term FinOps—financial operations—is being used a lot, now—

Corey: Cloud financial management is another popular one. Oh, swing a dead cat, you’ll hit 15 different words, and I give—my advice on that, even though I hate some of the terms is, cool. If people are going to pay you to have a title, even if you think it’s ridiculous, you can take the money or you can die on a petty naming hill and here we are.

Dann: Yeah. And it’s interesting because the role that I was hired for at Datadog was very much this niche, very specific role that I didn’t realize was a niche, very specific role at the time. So previously, I was at a company and I was building out their data centers, so I was working with vendors, buying servers, sometimes going on-site, installing, racking those, dealing with RMAs. And I was getting more involved as their cloud usage was growing and bringing some of those hardware capitalization cost procedures to the cloud. And so I found myself in this kind of niche role in my previous company.

And a Datadog, they basically had the exact same role that was dealing with all of the billing stuff around the cloud—kind of from an engineering perspective because it was on the engineering team—but working closely with finance, and I was like, “Oh, these are the skills that I have.” And it kind of fit perfectly. And it wasn’t until after I got to Datadog and was doing more research about this specific space that I discovered just how wide open it was. And I mean, meeting you was one of the earliest things that I did in the industry. Discovering the FinOps Foundation and a few other things has kind of like opened my eyes to this as an actual career path.

Corey: It’s an expensive problem that isn’t going away anytime soon, and it is foundational and core to the entire rest of how companies are building things these days. My argument has been for a while that when it comes to cloud, cost and architecture are the exact same thing. You don’t have the deep SRE architect background, but you’re also now a member of a four-person team. Does everyone in the team have the same skill set as you, or do you wind up effectively tagging in subject matter experts from different areas? How is the team composed? People love to ask me this question, and I strongly believe there’s no one way to do it. But what’s your answer?

Dann: Yeah, I mean, the team works very much in terms of everybody kind of taking on tasks that they need to do, but we did hire for specific skill sets when we tried to find people. So, the first person that we hired, we wanted them to have more of a developer engineer type background, writing code, stuff like that. The third hire, we were looking for somebody that was more of a generalist. I’ve seen myself more as a generalist in the space; anything that’s going on, I can pick it up and make some progress on it and build something out. And then the fourth person, we were lacking some of the deeper FP&A or FinOps experience, and so we found somebody with more of that kind of background and less of the engineering experience, but they were eager to, kind of, move from finance into more of an engineering role. And I feel like this is the perfect role for that because I feel like there are a lot of non-engineers that want to break into engineering and don’t really know how to do it. And if you are in finance, in FP&A, finding one of these more cloud-cost-optimization-specific roles is the great way to bridge that gap, I
feel.

Corey: The last time we spoke, I was independent, doing this all myself, and it turns out that taking all of the things that make me and trying to find those in other people is a relatively heavy lift, even if you discount the things like ‘obnoxious on Twitter.’ So, how do you start decomposing that? Well, now we’re a dozen people and we’ve found ways to do it. But by and large in our experience, for the way that we interact—and I want to get to that in a second—is that it’s easier for us to teach engineers how finance works than it is the opposite direction. And there are exceptions to that, and as we scale, I can easily see a day in the near future where that is no longer the case.

However, we also have two very specific styles of engagement. We do our cost optimization projects, where we go into an environment and, “Oh, fix this. Turn that thing off. Do you really need eight copies of those four petabytes of data? Oh, you didn’t realize they were there. Great, maybe delete it.” And we look like wizards from the future and things are great.

The other project that we do is contract negotiation with AWS, especially at large scale. It’s never as simple as people would have you believe because, “Oh, you’re doing co-marketing efforts, and you have a very specific use case, and there are business partnerships on 15 different levels, and that all factors into how this works.” It’s nuanced and challenging, and of course, because it’s a series of anecdata, I can’t really tell too many stories in public about that. But those are the two things that we wind up focusing on. You are focusing on a very different problem.

You’re not moving from company to company, basically reimplementing the same global problem, solving it locally for them. You are embedded in an account for the duration, almost four years now by my count. And, “Okay, I guess I could just do a whole bunch of cost optimization projects on a quarterly basis in an environment like that,” doesn’t seem like it solves the problem in any meaningful way. What
does your team do?

Dann: Yeah. Well, I mean, that’s such an interesting question. Just in terms of—yeah, if you’re doing consulting, you’re starting from square one every time you get a new contract, a new engagement, and being at the same company for, like you said, about four years, going on four years now, you really have a chance to dive in and think about, “Okay, what does it mean to work cloud cost optimization into just the regular business cycle of how it works?” Because I mean, you have the triangle that everybody’s familiar with: things can either be cheaper, faster, efficient and at different stages in the product lifecycle, you want to be focusing on these areas, more or less. And so, on our team, the different things that I’m thinking about is, first is visibility, is you want to provide engineers visibility into their cost. And not just numbers, right? Actionable visibility where if something needs to change, they need to do something, they know what that is.

And a lot of the times, that means not just costs, but also efficiency. So, these are the metrics that this particular application should be scaling against. As this application grows, as usage grows, are we remaining as cost-efficient? Then there’s also the piece—as you’re saying—like discovering things within the infrastructure that, “Hey, if we make this change, or if you turn this off, if we do things this way, we’ll save a bunch of money. Let’s do those.”

There’s things like reservations, committed use discounts for GCP, all of those kinds of things we manage. And then dealing closely with verifying our bill, working with finance—FP&A—on cost modeling forecasting, both short-term—like, within a month; like, what are we going to be at the end of this month and it’s the 10th right now?—and also, what does our next quarter look like? What are our next two years look like? And that bleeds into the contract negotiations, those kind of things as well.

So, I mean, it’s setting up the cycles of how do you prioritize this work? What is the company focusing on at the time? And what can you do when the company is not focusing explicitly on deciding to save money?

Corey: One of the more interesting aspects of my work that I didn’t expect is, whenever I wind up starting an engagement, or even in the prospect stage, I love asking the dumbest possible questions I can think of because it turns out they’re not. And the most common one that I always love to start with is, “Oh, okay. Your AWS bill is too high. Why do you care?” And that often takes people aback, but once you dig down underneath the surface just a little bit, it becomes pretty clear that the actual goal is not that it’s too much money—because spoiler, payroll always cost more than infrastructure—instead, it’s, “How do I think about this? How do I rationalize what the additional costs are going to be per thousand monthly active users or whatever metric it is you’re choosing to use?”

And how do you wind up forecasting that because the old days of data centers where you—“Well, we’re going to spend a boatload of money, and then we’ll have capacity for the next, ehh, two years, maybe down to eighteen months, depending on growth,” that’s easier for companies to rationalize around, rather than this idea of incremental cost on a per-unit basis, but not exactly because it also turns out that architecture changes, problems of scale, AWS pricing changes from time to time, all tend to impact that. What I think is not well understood in this space is that yeah, if you have a 20% overage this month, people are going to have some serious questions, but they’re also going to have those same questions if you’re 20% low.

Dann: Yeah. I mean, understanding why people care about the cost is definitely the first step because with a single company, so it’s just
constantly looking at the numbers rather than understanding exactly what motivations a company has to contact somebody like you, like a consultant, right? Because usually, I imagine that it’s going to be a bill, maybe two bills, three bills come in, and they keep going up and up and up, and they need to go down. And they’re going to have an explicit reason why it needs to go down; finance is going to say, “Margins are x, y, and z,” or, “Revenue has done this; our costs can’t do this.” There’s going to be explicit reasons because if there aren’t reasons, then they shouldn’t necessarily be focusing on costs at that moment in time.

What you want to do is have—I mean, this is way more complicated than just saying it out loud, but have a culture of cloud cost mindfulness, where people aren’t just spinning up resources willy nilly. But also, my goal is for people not to have to really think about cost that much other than just in a way that helps them do their work. Because I mean, I want engineers to be able to build stuff and build stuff fast—that’s what the cloud is all about—but I also want to be able to do it in a way that isn’t inappropriately high in cost.

Corey: I have my thoughts on this, and I’ve shared them before and I’ll dive into them again, but how do you approach that? If Datadog makes a grievous error and hires me to write code somewhere as an engineer, what is the, I guess, cost approach training for me as I wind up going through my onboarding as part of an SRE team or an application team?

Dann: I mean, this feels so basic as to not even be the right answer, but honestly, visibility is the easiest and best thing that you can give people, and so we’ve built out some visibility reports that engineers get on a regular basis. We also meet with our top—what is it—ten or fifteen spending internal engineering teams on a monthly basis to go over those costs so that they understand what they’re looking at so that we understand the context behind it, so that we can understand what’s on the roadmap going forward so that when things in the cost happen, we’re aware. And then we’re just staying on top of things. And if we have questions, we have an open dialogue with engineers and things like that.

In an ideal space, it would be great to have cost, I guess, more fit into the product development lifecycle in a more deeply ingrained way, but at the same time, I really don’t want to serve as a gatekeeper. Our goal is not to stop any sort of engineering process. And we haven’t needed to do anything like that although I guess every company is going to be different in terms of what their needs are. But yeah, I’m totally happy to being a little bit more reactionary in terms of looking at the numbers and responding, and then proactive just in terms of the regular communication with people.

Corey: I tend to take the perspective that engineers need to know enough about cost to maybe fill an index card at most because you don’t want them, I guess, over-fixating on it. Left to my own devices in my personal account, I’ll see a $7 a month bill and, “Oh, I’m going to spend two weeks knocking that down to $4.” And of course, I can do it, but is that the best use of my time? Absolutely not.

Very often what is a lot of money to an engineer is absolutely not to the business. And vice versa when you bring in a data science team; it’s, “Oh, yeah, we need at least four more exabytes of data because we never learned to do a join properly.” Yeah, maybe don’t do that. Understanding the difference between those two approaches is key. But I’ve always been of the mindset that I would rather bias for letting developers build and experiment and have things that catch outsized things quickly, then trying to wind up putting a culture of fear around cost because I’d much rather see whether the thing they’re trying to build is possible to build, then go back and optimize it later, once that’s proven out. But again, this is a nuanced thing.

Everyone seems to think I have this back pocket answer that will apply to all companies. And you’ve been doing this at Datadog for almost four years with a team of people. I am an outsider; I see the global trend, I see what works in different ways in different companies, but the idea that I can sit down and say, “Oh. Well, clearly the thing you’re doing is completely wrong because that’s not how I think about it,” is the hallmark of a terrible consultant. There are reasons that things are the way that they are and it’s generally not that people are expecting to do a terrible job today. You know, unless they work in the Facebook ethics department, which is neither here nor there.

Dann: Yeah, I mean, like I said, the product lifecycle, when you’re building something new, you want to go as fast as possible. When you’re launching it, you want it to be as reliable as possible. Once you’re launched, once you’re reliable, then you can start focusing on costs is, kind of like, not the universal rule, but kind of the flow that I tend to see. So, as you’re at a company that is regularly innovating, creating new products, going through that cycle, you’re going to have these kind of periods.

As well as you have the products that have been around. There’s a lot of legacy code, there’s a lot of stuff going on, that maybe isn’t the best, or some efficiency work that has been deprioritized for whatever reason, that maybe it’s time to start considering doing this. So, keeping track of all of that. And like I said, if for whatever reason the business wants to focus on cloud cost efficiency, or a team has decided that in a particular quarter or for a particular reason they want to focus on that, being able to assist as much as you can, being able to save all that work so that there’s kind of like a queue that you can go to when it is time to focus on cost efficiency stuff.

Corey: So, here’s a fun one for you. As of the time of this recording, it’s a couple weeks old, but if you’re anything like what we do here for some of our more sophisticated clients, we do occasionally build out prediction models, models of economics that wind up defining how some architectural patterns should be addressed, et cetera, et cetera. What’s always fun is the large clients who have this significant level of spend on an outlier service. Every once in a while—it was great that we got to do a deep dive into the Washington Post’s use of Lambda because normally, Lambda is a rounding error on the bill; they had a specific challenge and they did a whole blog post on this for the AWS blog. I believe the Monitoring Tools blog, but don’t take that at face value; I never remember which AWS blog is which because AWS doesn’t speak with a single voice on anything.

But yeah, most of the time is block, tackle, baseline stuff that is the big driver of spend, but a few weeks ago, they change the pricing dimensions for S3 intelligent tiering, where there’s no longer a monitoring charge for objects that are smaller than 128 kilobytes, and there’s no 30-day minimum. So, the fact that those two things went away removed almost every caveat that I can picture for using S3 intelligent tiering, which means that for most use cases, that should now be the default. I imagine you caught that change as well, since that’s one of those wake up and take notice, no matter what time of the world [laugh] it is where you are when that gets dropped. How did that change your modeling? Or did that not significantly shift how you view any of this?

Dann: No, I mean, I think part of our role within the organization is to pay attention to stuff like that, and then you just have those conversations with the teams that I know were either exploring intelligent tiering. We do some pricing modeling for different products, S3 storage for different types, so updating those and being like, “Hey, this might be something we want to actually use and explore now.” Similar and I guess, more of something that I actively worked on that I consider in the same category is when Amazon announced savings plans as replacing convertible reservations. Because at first they announced, and being like, “Okay, well, it’s going to automatically rebalance between… different instance families across regions, too”—which convertible RIs could never do it—“And it’s going to be the exact same price for a compute savings plan as a convertible RI.” And we were kind of like, what’s the catch? And we spent a few weeks doing a deep dive working with our data science team, kind of like being, “Where is the catch here?”

Corey: Yeah, the real catch is that you can’t sell it on the secondary market if it—

Dann: Yeah.

Corey: —turns out you bought the wrong thing, which if that’s your Plan A, then good luck.

Dann: Yeah. We definitely don’t use that secondary market. I don’t have as much experience there, although I’m sure some people can use it to their advantage.

Corey: Almost no one does. In fact, the reason that it exists—my pet theory—is that once upon a time, companies would try and classify some of the reserved instance purchases as capital expenditures, which there has since been guidance from regulatory authorities not to do that. But at the time, the fact that you could sell it to a third-party on the secondary market would help shore up that argument. If you’re listening to this, and you’re classifying some of your RIs as CapEx, please don’t do that. Feel free to reach out to me, I can dig out the actual regulation and send it to you. There are two of them. It’s a nuanced topic. If you’re listening to this and have no idea what I’m talking about, God, do I envy you.

Dann: [laugh]. Yeah, definitely don’t do that. [laugh].

Corey: There was a lot that was interesting about savings plans. When I was read in the month or so in advance of them being announced, it was, “Great. I want to see this and this and these other things, too.” And some of those things came to pass. It was extended to work with Lambda.

Now, I don’t believe that is financially useful in almost every case, but it doesn’t need to be because so much of cloud economics from where I sit is psychological in nature, where, “Oh, we have this workload that lives on EC2 instances and we want to move it to Lambda, but we already bought the reserved instances so we’re not going to do it because of sunk cost fallacy.” Which is not much of a fallacy when it’s that kind of money, in some cases. Okay, great. Now, if it can migrate to Lambda and still wind up getting the discounts you’ve paid for it, you have removed an architectural barrier. And that’s significant.

Now, I want to see that same thing apply to oh if you move from EC2 to RDS, or DynamoDB or anything else, that should be helpful, too. But whatever you do, don’t do what SageMaker did and launch their own separate savings plan that is not compatible with the compute savings plans, so effectively, it’s great; you’re locked-in architecturally to one or the other because machine learning is, once again, a marvelously executed scam to sell pickaxes into a digital gold rush.

Dann: I mean, I like savings plans a lot and we’ve been slowly, as convertible RIs have expired, replacing them with savings plans. And I think that it is pushing the other cloud providers forward—because we’re definitely multi-cloud—and so that’s really useful and I hope more people will take on the compute savings plan type model, just because it makes our lives so much easier. Or it makes my life so much easier in terms of planning it, selling the commitment internally, just everything about it has made my life easier. So, I mean, how many years later are we? I definitely haven’t found any big gotchas, I guess, from the secondary market. But that doesn’t really impact me.

Corey: Yeah, I spent a lot of time looking forward, too, doing deep analyses of okay, for which instance classes in which regions is there a price discrepancy? And I finally got someone to go semi on record and say, “Yeah. There should not be any please ping us if you find one.” “Oh, okay, great. That is enough for me to work with.”

Dann: Exactly, we got that, too. I didn’t believe it so we were downloading price sheets and doing comparisons, doing all that stuff.

Corey: Oh, trust but verify. And when we’re talking this kind of money, I don’t trust very far. They make mistakes on billing issues from time to time. And I get it; it’s hard, but there are challenges here and there. I am glad you mentioned a minute ago that you are multi-cloud because my position on that has often been misconstrued.

I think that designing something from day one to work on multiple cloud providers is generally foolish. I think that unless you have a compelling reason not to go all-in on one cloud provider, that’s what you should do. Pick a cloud—I don’t care which—and go all-in. Conversely, you have a product like Datadog where your customers are in multiple clouds, and first, no one wants to pay egress to send all the telemetry from where they are into AWS, and secondly, they’re not going to put up, in many cases, with their data going to a cloud provider they have explicitly chosen not to work with, so you have to meet your customers where they are. In your case, it is absolutely the right thing to do. And Twitter often gets upset and calls me hypocrite on stuff like this because Twitter believes that two things that take opposite visions cannot possibly both be true, but the world is messy.

Dann: Yeah. And I mean, the nice thing about us being in multiple clouds is we are our own biggest user. And that’s actually one of the reasons why I love working at Datadog is because I get to use Datadog all the time. And not only that, Datadog is on everything and we have all of our products. I’m very spoiled [laugh] with all of this. But I mean, we are running in these different cloud providers; we are using Datadog in those different cloud providers, and that is just helping everything overall, too. In addition to supporting customers that are in each cloud because that is a huge reason as well.

Corey: This episode is sponsored in part by something new. Cloud Academy is a training platform built on two primary goals. Having the highest quality content in tech and cloud skills, and building a good community the is rich and full of IT and engineering professionals. You wouldn’t think those things go together, but sometimes they do. Its both useful for individuals and large enterprises, but here's what makes it new. I don’t use that term lightly. Cloud Academy invites you to showcase just how good your AWS skills are. For the next four weeks you’ll have a chance to prove yourself. Compete in four unique lab challenges, where they’ll be awarding more than $2000 in cash and prizes. I’m not kidding, first place is a thousand bucks. Pre-register for the first challenge now, one that I picked out myself on Amazon SNS image resizing, by visiting cloudacademy.com/corey. C-O-R-E-Y. That’s cloudacademy.com/corey. We’re gonna have some fun with this one!

Corey: One of the problems that I keep running into across the board is that with things like Datadog—and again, not to single you out; every monitoring vendor to some extent has aspects of this problem—it’s that when I’m a customer and I’m hooking my accounts up to Datadog, I want you to tell me about things that are going on, but the CloudWatch charges can be so egregious on the customer side, where it is bizarre and, frankly, abhorrent to me when I wind up paying more for the CloudWatch charges than I am for Datadog. And let’s be clear here; I am, in fact, a Datadog customer. I pay you folks money. Not a lot of money, but I pay you money because I have certain things that I need to know are working for a variety of excellent reasons.

And the problem that I keep smacking into on this is—it’s not your fault; there’s not anything you can do. In fact, you are one of the better providers as far as not only not being egregious with the way that you slam the CloudWatch endpoints, but also in giving guidance to customers on how to tune it further. And I really wish that more folks in your space would do things like that. It always bugs me when I wind up using a tool that tries to save money that in turn winds up costing me more than it saves.

Dann: Yeah. Yeah, it’s tricky there. I have less experienced myself setting up Datadog and running it in my own infrastructure as I’m more digging deep into the cost stuff and us using the cloud, so I can’t speak to that specifically. But yeah, you’re not the first person that I’ve heard have that experience. [laugh].

Corey: And again, it’s not your fault at all. I’ve been beating up the CloudWatch team for years on this, and I will continue to do so until I’m safely dead, which—depending on Amazon’s level of patience—might be in mere minutes.

Dann: In the larger-picture-wise, we have to remember that we’re super early in the cloud adoption, even looking at the cloud economics FinOps cloud cost optimization world. I feel like most businesses at this stage in their journey are still in data centers and they’re dealing with the problem of how do we move to the cloud and do it cost-efficiently? How do we set everything up? And that’s where the world is right now.

And I think that dealing with, “Okay, we are one hundred percent running in the cloud. What are the processes that we have in place? How do we think of finance and the finance organization not through the lens of ‘we once had data centers and now we don’t,’ but how do we look through that in the lens of ‘okay, we are cloud-native from day one? What does the finance department look like?’” And dealing with those problems is really interesting because Datadog has never been in a data center. We are cloud-native from the very beginning, and so it was interesting for me to join the company and build up a lot of these processes because it is different than what a lot of other people were dealing with and doing. And it presents some really interesting problems and questions that I think are going to be the foundation for the next decade of building companies and operating in the cloud.

Corey: I always love having conversations with folks who are building out teams to handle these things because usually the folks I keep talking to, or who want to have conversations like this are building tools themselves to solve this problem through the miracle of SaaS, where they will bend over backwards to avoid ever talking to a customer. And we’re all dealing with the same AWS APIs; there’s not that much of a new spin you can put on most of these things. But understanding what customers are actually trying to do instead of falling down the rabbit hole trap of, “Hey, turn off those idle instances that are all labeled ‘drsite’ because you probably don’t need them,” is foolish. And after a few foolish recommendations, tooling doesn’t get there. I am a big believer that tools can assist the process and narrow down what to look at.

I believe they shouldn’t have to exist; I think that the billing dashboard should be a hell of a lot better natively than having to pay a third party to make sense of it for me. But by and large, I do believe this is a problem that is best solved from a consultative approach. When I started this place, I was planning to build out some software, tried doing it—called DuckTools—and wound up mothballing the whole thing because what we were building was not what the industry claimed to want and, frankly, educating people into a position where then they see the value and only then will they buy is never been a game that I wanted to play.

Dann: Yeah, I really liked that article that you guys published about exploring that product and the reason why you decided not to pursue it. But it’s super interesting in terms of where the industry is going and building out those tools because I found that there isn’t really any new thing that you can do with the tools. All the tools that exist for looking at your costs are largely the same. The main differences that I’ve seen is that the UI is slightly different and they have different sales teams. And if the sales teams are better, they’re going to get more of the market share. And if the sales teams are not as good, it’s going to be a smaller market share. And it’s weird, too, be in this industry for as long as we have been, and seeing okay, well, Andreessen Horowitz just funded this new company, and this other company got invited into Y Combinator, or all of these things that are happening, and I’m kind of like, okay, but what is this tool really doing differently? And there are a few of them that are; that are doing something innovative and different, but there’s also a few that are just like, this is a space where people are in, there’s money here, we’re doing the same thing, but we got our sales team, and we’ll carve out our little corner, and then we’ll get acquired, and that’ll be that. Although I guess we’re just at that stage of innovation in this space, I guess.

Corey: Yeah, I have no earthly idea what the story is around how these companies plan to differentiate because it seems to me that they’re directly attempting to compete with Cost Explorer, which—

Dann: Yeah.

Corey: —it’s taken some time for that thing to improve to the point where it is now and it’ll take further time for it to improve beyond it, but long-term, I don’t think you’re going to outrun AWS on a straight line like that.

Dann: Yeah, I mean, when you work for one of these third-party cost tooling things, and you’re working with one of your customers, and they’re like, “How do I view this?” And it’s kind of like, that is the easiest thing to find in Cost Explorer as well, it’s—I can’t imagine being like,
“Well, you should pay me thousands, tens of thousands, hundreds of thousands of dollars a month to view it here,” when Cost Explorer is free. And I think Cost Explorer, it doesn’t do everything, but it’s gotten a lot better at what it does, and it could probably solve 90% of people’s
problems without using a third-party tool.

Corey: You are at significant scale in multiple clouds, so the answer that these companies always give is, “Ah, but we provide a single
dashboard so that you can look at costs across multiple providers in one place.” Is that even slightly useful to you?

Dann: Man, if you need dashboards, get a dashboard tool. Don’t get this crazy cost analysis tool. I mean, there are some great dashboard solutions that you can get where you can connect your detailed billing, cost and usage report—whatever cloud provider is calling it, but, like, that really detailed gigabytes per hour report—and then visualize it, build reports, do all that kind of stuff because that’s not something that the tooling does well right now, in terms of building out cost dashboards and stuff. But that’s also right now. It could in the future.

Corey: Yeah. If you’re a BI tool, wind up passing out templates that normalize these things? I am so tired of building it all from scratch in Tableau myself. If you’re Tableau, sell me a whole bunch of things that I can use to view this stuff through, so I don’t have to wind up continually reinventing that particular wheel.

Dann: Yeah.

Corey: Oh, I like your approach. I didn’t know the answer when I was asking the question. I was about to learn something if you’d gone the other direction, but nope, but it’s good to know that my impressions remain intact.

Dann: Yeah, I mean, I’ve used different tools in the past. Again, I hesitate to name any of them, but there’s a few in this space that I feel like everybody—if they’re in this space, they know which tools I’m talking about—

Corey: Yes, we do.

Dann: —and… yeah, I’ve used them. They’re okay—a few of them are okay, a few of them are better than others, but I mean, I was trying to
evaluate the value-add over me manually setting some things up and having some sort of visualization, and just the value-add in terms of what they were charging, even if it was like a significantly smaller percent of the bill because that alone, like, percent of bill is such a difficult cost model—

Corey: Oh—

Dann: —to do.

Corey: I hate that. Pricing is hard. Let’s start there.

Dann: Yeah. Yeah. Yeah, yeah.

Corey: I hate the percent of bill because then it’s, “Let me get this straight. I’m paying you a percentage of things like data transfer charges that I know are fixed, that I can’t optimize? I’m paying you a percentage of my AWS enterprise support subscription? I’m paying you a percentage of the marketplace?” And so on and so forth. And it doesn’t work. At some point of scale as well it’s, I could hire a team of 20 people and save money versus what you’re charging me. The other side of it though, “Ah, we’ll charge you percentage of savings.” Well, then you wind up with people doing a whole bunch of things like before they bring you in, they’ll make a bunch of ill-advised reserved instance purchases or savings plan purchases you have to then unwind after the fact. When I was setting this place up, I looked long and hard at different billing models and the only thing I found that worked is fixed fee. The end. Because at that point, suddenly everyone’s on board with, “Hey, let’s solve the problem and then get out as soon as possible.” We’re not trying to build ourselves a forever job nestled in the heart of your company. And it’s the only model I found that removes a whole swath of conflicts of interest. And that’s the hard part. We have no partners with anyone in this space—including AWS themselves—just because as soon as we do, it becomes extremely disingenuous when we suggest doing something for your sake that happens to benefit them, such as, “Maybe back that S3 bucket up somewhere.” Well, okay, if we’re partnered with them, does that mean we’re trying to influence spend in the other direction? And it just becomes a morass that I never found it worth the time to deal with.

Dann: Yeah, I—

Corey: But that doesn’t work for SaaS.

Dann: Yeah, that makes a lot of sense. And I haven’t actually thought about pricing model for consulting in this space that closely, but I mean, when you’re charging a percent of bill or percent of savings, you have the opportunity to screw the customer, right, through all the things that you were saying. If you charge a fixed fee, you have the possibility of undervaluing yourself, which the only one that’s screwed in that case is you, potentially, and if you’re okay with that risk and you’re okay with those dollars, that’s great. Because yeah, if you’re able to be like, “Okay, here’s the services that I do, here’s the fixed costs.” “Done.” “Done.” That just sets everybody’s expectations for the relationship in a much better way that you’re not constantly worried about, like, upsells and other things that might happen along the way that screws the customer.

Corey: And that’s the hardest part, I think, is that people lose sight of the entire customer obsession piece of it. That’s one of the things Amazon gets super right. I wish more companies embrace that. Dann, I want to thank you for taking so much time out of your day to suffer my slings, and arrows, and half-formed opinions. If people want to learn more about who you are and what you’re up to, where can they find you?

Dann: Yeah, I have a website you guys can go to that links everywhere else. It is dannb.org. And I spell my name with two ns, so D-A-N-N-B dot org. And I have LinkedIn, I have Twitter, I have a monthly newsletter that is not really about FinOps or anything, but I really enjoy it; I’ve been doing it for a year, now, that you should sign up for.

Corey: And links to that will, of course, be in the [show notes 00:36:26]. Dann, thanks again for your time. I really appreciate it.

Dann: Yeah. Thanks so much for having me again. It’s been a blast.

Corey: It really has. Dann Berg, senior CloudOps analyst at Datadog. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with a comment featuring a picture of several corkboards full of post-it notes and string, and a deranged comment telling me that you have in fact finally found the catch in savings plans.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Betty

Betty Junod is the Senior Director of Multi-Cloud Solutions at VMware helping organizations along their journey to cloud. This is her second time at VMware, having previously led product marketing for end user computing products. Prior to VMware she held marketing leadership roles at Docker and solo.io in following the evolution of technology abstractions from virtualization, containers, to service mesh. She likes to hang out at the intersection of open source, distributed systems, and enterprise infrastructure software. @bettyjunod

Links:

  • Twitter: https://twitter.com/BettyJunod
  • Vmware.com/cloud: https://vmware.com/cloud

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: You know how git works right?

Announcer: Sorta, kinda, not really Please ask someone else!

Corey: Thats all of us. Git is how we build things, and Netlify is one of the best way I’ve found to build those things quickly for the web. Netlify’s git based workflows mean you don't have to play slap and tickle with integrating arcane non-sense and web hooks, which are themselves about as well understood as git. Give them a try and see what folks ranging from my fake Twitter for pets startup, to global fortune 2000 companies are raving about. If you end up talking to them, because you don't have to, they get why self service is important—but if you do, be sure to tell them that I sent you and watch all of the blood drain from their faces instantly. You can find them in the AWS marketplace or at www.netlify.com. N-E-T-L-I-F-Y.com

Corey: This episode is sponsored in part by our friends at Vultr. Spelled V-U-L-T-R because they’re all about helping save money, including on things like, you know, vowels. So, what they do is they are a cloud provider that provides surprisingly high performance cloud compute at a price that—while sure they claim its better than AWS pricing—and when they say that they mean it is less money. Sure, I don’t dispute that but what I find interesting is that it’s predictable. They tell you in advance on a monthly basis what it’s going to going to cost. They have a bunch of advanced networking features. They have nineteen global locations and scale things elastically. Not to be confused with openly, because apparently elastic and open can mean the same thing sometimes. They have had over a million users. Deployments take less that sixty seconds across twelve pre-selected operating systems. Or, if you’re one of those nutters like me, you can bring your own ISO and install basically any operating system you want. Starting with pricing as low as $2.50 a month for Vultr cloud compute they have plans for developers and businesses of all sizes, except maybe Amazon, who stubbornly insists on having something to scale all on their own. Try Vultr today for free by visiting: vultr.com/screaming, and you’ll receive a $100 in credit. Thats v-u-l-t-r.com slash screaming.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Periodically, I like to poke fun at a variety of different things, and that can range from technologies or approaches like multi-cloud, and that includes business functions like marketing, and sometimes it extends even to companies like VMware. My guest today is the Senior Director of Multi-Cloud Solutions at VMware, so I’m basically spoilt for choice. Betty Junod, thank you so much for taking the time to speak with me today and tolerate what is no doubt going to be an interesting episode, one way or the other.

Betty: Hey, Corey, thanks for having me. I’ve been a longtime follower, and I’m so happy to be here. And good to know that I’m kind of like the ultimate cross-section of all the things [laugh] that you can get snarky about.

Corey: The only thing that’s going to make that even better is if you tell me, “Oh, yeah, and I moonlight on a contract gig by naming AWS services.” And then I just won’t even know where to go. But I’ll assume they have to generate those custom names in-house.

Betty: Yes. Yes, I think they do those there. I may comment on it after the fact.

Corey: So, periodically I am, let’s call it miscategorized, in my position on multi-cloud, which is that it’s a worst practice that when you’re designing something from scratch, you should almost certainly not be embracing unless you’re targeting a very specific corner case. And I stand by that, but what that has been interpreted as by the industry, in many cases because people lack nuance when you express your opinions in tweet-sized format—who knew—as me saying, “Multi-cloud bad.” Maybe, maybe not. I’m not interested in assigning value judgment to it, but the reality is that there are an awful lot of multi-cloud deployments out there. And yes, some of them started off as, “We’re going to migrate from one to the other,” and then people gave up and called it multi-cloud, but it is nuanced. VMware is a company that’s been around for a long time. It has reinvented itself in a few different ways at different periods of its evolution, and it’s still highly relevant. What is the Multi-Cloud Solutions group over at VMware? What do you folks do exactly?

Betty: Yeah. And so I will start by multi-cloud; we’re really taking it from a position of meeting the customer where they are. So, we know that if anything, the only thing that’s a given in our industry is that there will be something new in the next six months, next year, and the whole idea of multi-cloud, from our perspective, is giving customers the optionality, so don’t make it so that it’s a closed thing for them. But if they decide—it’s not that they’re going to start, “Hey, I’m going to go to cloud, so day one, I’m going to go all-in on every cloud out there.” That doesn’t make sense, right, as—

Corey: But they all gave me such generous free credit offers when I founded my startup; I feel obligated to at this point.

Betty: I mean, you can definitely create your account, log in, play around, get familiar with the console, but going from zero to being fully operationalized team to run production workloads with the same kind of SLAs you had before, across all three clouds—what—within a week is not feasible for people getting trained up and actually doing that. Our position is that meeting customers where they are and knowing that they may change their mind, or something new will come up—a new service—and they really want to use a new service from let’s say GCP or AWS, they want to bring that with an application they already have or build a new app somewhere, we want to help enable that choice. And whether that choice applies to taking an existing app that’s been running in their data center—probably on vSphere—to a new place, or building new stuff with containers, Kubernetes, serverless, whatever. So, it’s all just about helping them actually take advantage of those technologies.

Corey: So, it’s interesting to me about your multi-cloud group, for lack of a better term, is there a bunch of things fall under its umbrella? I believe Bitnami does—or as I insist on calling it, ‘bitten-A-M-I’—I believe that SaltStack—which I wrote a little bit of once upon a time, which tells me you folks did no due diligence whatsoever because everything I’ve ever written is molten garbage—

Betty: Not [unintelligible 00:04:33].

Corey: And—so to be clear, SaltStack is good; just the parts that I wrote are almost certainly terrible because have you met me?

Betty: I’ll make a note. [laugh].

Corey: You have Wavefront, you have CloudHealth, you have a bunch of other things in the portfolio, and yeah, all those things do work across multiple clouds, but there’s nothing that makes using any of those things a particularly bad idea even if you’re all-in on one cloud provider, too. So, it’s a portfolio that applies to a whole bunch have different places from your perspective, but it can be used regardless of where folks stand ideologically.

Betty: Yes. So, this goes back to the whole idea that we meet the customers where they are and help them do what they want to do. So, with that, making sure these technologies that we have work on all the clouds, whether that be in the data center or the different vendors, so that if a customer wants to just use one, or two, or three, it’s fine. That part’s up to them.

Corey: The challenge I’ve run into is that—and maybe this is a ‘Twitter Bubble’ problem, but unfortunately, having talked to a whole bunch of folks in different contexts, I know it isn’t—there’s almost this idea that you have to be incredibly dogmatic about a particular technology that you’re into. I joke periodically about the Rust Evangelism Strikeforce where their entire job is talking about using Rust; their primary IDE is PowerPoint because they’re giving talks all the time about it rather than writing code. And great, that’s a bit of an exaggeration, but there are the idea of a technology purist who is taking, “Things must be this way,” well past a point of being reasonable, and disregarding the reality that, yeah, the world is messy in a way that architectural diagrams never are.

Betty: Yeah. The architectural diagrams are always 2D, right? Back to that PowerPoint slide: how can I make pretty boxes? And then I just redraw a line because something new came out. But you and I have been in this industry for a long time, there’s always something new.

And I think that’s where the dogmatism gets problematic because if you say we’re only going to do containers this way—you know, I could see Swarm and Kubernetes, or all-in on AWS and we’re going to use all the things from AWS and there’s only this way. Things are generational and so the idea that you want to face the reality and say that there is a little bit of everything. And then it’s kind of like, how do you help them with a part of that? As a vendor, it could be like, “I’m going to help us with a part of it, or I’m going to help address certain eras of it.” That’s where I think it gets really bad to be super dogmatic because it closes you off to possibly something new and amazing, new thinking, different ways to solve the same problem.

Corey: That’s the problem is left to our own devices, most of us who are building things, especially for random ideas, yeah, there’s a whole modern paradigm of how I can build these things, but I’m going to shortcut to the thing I know best, which may very well the architectures that I was using 15 years ago, maybe tools that I was using 15 years ago. There’s a reason that Vim is still as popular as it is. Would I recommend it to someone who’s a new user? Absolutely not; it’s user-hostile, but back in my days of being a grumpy sysadmin, you learned vi because it was on everything you could get into, and you never knew in what environment you were going to be encountering stuff. These days, you aren’t logging in to remote systems to manage them, in most cases, and when it happens, it’s a rarity and a bug.

The world changes; different approaches change, but you have to almost reinvent your entire philosophy on how things work and what your career trajectory looks like. And you have to give up aspects of what you’ve considered to be part of your identity and embrace something new. It was hard for me to accept that, for example, Docker and the wave of containerization that was rolling out was effectively displacing the world that I was deep in of configuration management with Puppet and with Salt. And the world changes; I said, “Okay, now I’ll work on cloud.” And if something else happens, and mainframes are coming back again, instead, well, I’m probably not going to sit here railing against the tide. It would be ridiculous to do that from my perspective. But I definitely understand the temptation to fight against it.

Betty: Mm-hm. You know, we spend so much time learning parts of our craft, so it’s hard to say, “I’m now not going to be an expert in my thing,” and I have to admit that something else might be better and I have to be a newbie again. That can be scary for someone who’s spent a lot of time to be really well-versed in a specific technology. It’s funny that you bring up the whole Docker and Puppet config management; I just had a healthy discussion over Slack with some friends. Some people that we know and comment about some of the newer areas of config management, and the whole idea is like, is it a new category or an evolution of? And I went back to the point that I made earlier is like, it’s generations. We continually find new ways to solve a problem, and one thing now is it [sigh] it just all goes so much faster, now. There’s a new thing every week. [laugh] it seems sometimes.

Corey: It is, and this is the joy of having been in this industry for a while—toxic and broken in many ways though it is—is that you go through enough cycles of seeing today’s shiny, new, amazing thing become tomorrow’s legacy garbage that we’re stuck supporting, which means that—at least from my perspective—I tend to be fairly conservative with adopting new technologies with respect to things that matter. That means that I’m unlikely to wind up looking at the front page of Hacker News to pick a framework to build a banking system in, and I’m unlikely to be the first kid on my block to update to a new file system or database, just because, yeah, if I break a web server, we all laugh, we make fun of the fact that it throws an error for ten minutes, and then things are back up and running. If I break the database, there’s a terrific chance that we don’t have a company anymore. So, it’s the ‘mistakes will show’ area and understanding when to be aggressive and when to hold back as far as jumping into new technologies is always a nuanced decision. And let’s be clear as well, an awful lot of VMware’s customers are large companies that were founded, somehow—this is possible—before 2010. Imagine that. Did people—

Betty: [laugh]. I know, right?

Corey: —even have businesses or lives back then? I thought we all used horse-driven carriages and whatnot. And they did not build on cloud—not because of any perception of distrust; because it functionally did not exist at the time that they were building these things. And, “Oh, come out into the cloud. It’s fine now.” It… yeah, that application is generating hundreds of millions in revenue every quarter. Maybe we treat that with a little bit of respect, rather than YOLO-ing it into some Lambda-driven monster that’s constructed—

Betty: One hundred—

Corey: —out of popsicle sticks and glue.

Betty: —percent. Yes. I think people forget that. And it’s not that these companies don’t want to go to cloud. It’s like, “I can’t break this thing. That could be, like, millions of dollars lost, a second.”

Corey: I write my weekly newsletters in a custom monstrosity of a system that has something like 30-some-odd Lambda functions, a bunch of API gateways that are tied together with things, and periodically there are challenges with it that break as the system continues to evolve. And that’s fine. And I’m okay with using something like that as a part of my workflow because absolute worst case, I can go back to the way that my newsletter was originally written: in Google Docs, and it doesn’t look anywhere near the same way, and it goes back to just a text email that starts off with, “I have messed up.” And that would be a better story than most of the stuff I put out as a common basis. Similarly, yeah, durability is important.

If this were a serious life-critical app, it would not just be hanging out in a single region of a single provider; it would probably be on one provider, as I’ve talked about, but going multi-region and having backups to a different cloud provider. But if AWS takes a significant enough outage to us-west-2 in Oregon, to the point where my ridiculous system cannot function to write the newsletter, that too, is a different handwritten email that goes out that week because there’s no announcement they’ve made that anyone’s going to give the slightest toss about, given the fact that it’s basically Cloud Armageddon. So, we’ll see. It’s about understanding the blast radius and understanding your use case.

Betty: Yep. A hundred percent.

Corey: So, you’ve spent a fair bit of time doing interesting things in your career. This is your second outing at VMware, and in the interim, you were at solo.io for a bit, and before that you were in a marketing leadership role at Docker. Let’s dive in, if you will. Given that you are no longer working at Docker, they recently made an announcement about a pricing model change, whereas it is free to use Docker Desktop for anyone’s personal projects, and for small companies.

But if you’re a large company, which they define is ten million in revenue a year or 250 employees—those two things don’t go alike, but okay—then you have to wind up having a paid plan. And I will say it’s a novel approach, but I’m curious to hear what you have to say about it.

Betty: Well, I’d say that I saw that there was a lot of flutter about that news, and it’s kind of a, it doesn’t matter where you draw the line in the sand for the tier, there’s always going to be some pushback on it. So, you have to draw a line somewhere. I haven’t kept up with the details around the pricing models that they’ve implemented since I left Docker a few years ago, but monetization is a really important part for a startup. You do have to make money because there are people that you have to pay, and eventually, you want to get off of raising money from VCs all the time. Docker Desktop has been something that has been a real gem from a local developer experience, right, giving the—so that has been well-received by the community.

I think there was an enterprise application for it, but when I saw that, I was like, yeah, okay, cool. They need to do something with that. And then it’s always hard to see the blowback. I think sometimes with the years that we’ve had with Docker, it’s kind of like no matter what they do, the Twitterverse and Hacker News is going to just give them a hard time. I mean, that is my honest opinion on that. If they didn’t do it, and then, say, they didn’t make the kind of revenue they needed, people would—that would become another Twitter thread and Hacker News blow up, and if they do it, you’ll still have that same reaction.

Corey: This episode is sponsored by our friends at Oracle Cloud. Counting the pennies, but still dreaming of deploying apps instead of "Hello, World" demos? Allow me to introduce you to Oracle's Always Free tier. It provides over 20 free services and infrastructure, networking databases, observability, management, and security.

And - let me be clear here - it's actually free. There's no surprise billing until you intentionally and proactively upgrade your account. This means you can provision a virtual machine instance or spin up an autonomous database that manages itself all while gaining the networking load, balancing and storage resources that somehow never quite make it into most free tiers needed to support the application that you want to build.

With Always Free you can do things like run small scale applications, or do proof of concept testing without spending a dime. You know that I always like to put asterisks next to the word free. This is actually free. No asterisk. Start now. Visit https://snark.cloud/oci-free that's https://snark.cloud/oci-free.

Corey: It seems to be that Docker has been trying to figure out how to monetize for a very long time because let’s be clear here; I think it is difficult to overstate just how impactful and transformative Docker was to the industry. I gave a talk “Heresy in the Church of Docker” that listed a bunch of things that didn’t get solved with Docker, and I expected to be torn to pieces for it, and instead I was invited to give it at ContainerCon one year. And in time, a lot of those things stopped being issues because the industry found answers to it. Now, unfortunately, some of those answers look like Kubernetes, but that’s neither here nor there. But now it’s, okay, so giving everything that you do that is core and central away for free is absolutely part of what drove the adoption that it saw, but goodwill from developers is not the sort of thing that generally tends to lead to interesting revenue streams.

So, they had to do something. And they’ve tried a few different things that haven’t seemed to really pan out. Then they spun off that pesky part of their business that made money selling support contracts, over to Mirantis, which was apparently looking for something now that OpenStack was no longer going to be a thing, and Kubernetes is okay, “Well, we’ll take Docker enterprise stuff.” Great. What do they do, as far as turning this into a revenue model?

There’s a lot of the, I guess, noise that I tend to ignore when it comes to things like this because angry people on Twitter, or on Hacker News, or other terrible cesspools on the internet, are not where this is going to be decided. What I’m interested in is what the actual large companies are going to say about it. My problem with looking at it from the outside is that it feels as if there’s significant ambiguity across the board. And if there’s one thing that I know about large company procurement departments, it’s that they do not like ambiguity. This change takes effect in three or four months, which is underwear-outside-the-pants-superhero-style speed for a lot of those companies, and suddenly, for a lot of developers, they’re so far removed from the procurement side of the house that they are never going to have a hope of getting that approved on a career-wide timespan.

And suddenly, for a lot of those companies, installing and running Docker Desktop just became a fireable offense because from the company’s perspective, the sheer liability side of it, if they were getting subject to audit, is going to be a problem. I don’t believe that Docker is going to start pulling Oracle-like audit tactics, but no procurement or risk management group in the world is going to take that on faith. So, the problem is not that it’s expensive because that can be worked around; it’s not that there’s anything inherently wrong with their costing model. The problem is the ambiguity of people who just don’t know, “Does this apply to me or doesn’t this apply to me?” And that is the thing that is the difficult, painful part.

And now, as a result, the [unintelligible 00:17:28] groups and their champions of Docker Desktop are having to spend a lot more time, energy, and thought on this than it would simply be for cutting a check because now it’s a risk org-wide, and how do we audit to figure out who’s installed this previously free open-source thing? Now what?

Betty: Yeah, I’ll agree with you on that because once you start making it into corporate-issued software that you have to install on the desktop, that gets a lot harder. And how do you know who’s downloaded it? Like my own experience, right? I have a locked-down laptop; I can’t just install whatever I want. We have a software portal, which lets me download the approved things.

So, it’s that same kind of model. I’d be curious because once you start looking at from a large enterprise perspective, your developers are working on IP, so you don’t want that on something that they’ve downloaded using their personal account because now it sits—that code is sitting with their personal account that’s using this tool that’s super productive for them, and that transition to then go to an enterprise, large enterprise and going through a procurement cycle, getting a master services agreement, that’s no small feat. That’s a whole motion that is different than someone swiping a credit card or just downloading something and logging in. It’s similar to what you see sometimes with the—how many people have signed up for and paid 99 bucks for Dropbox, and then now all of a sudden, it’s like, “Wow, we have all of megacorp [laugh] signed up, and then now someone has to sell them a plan to actually manage it and make sure it’s not just sitting on all these personal drives.”

Corey: Well, that’s what AWS’s original sales motion looked a lot like they would come in and talk to the CTO or whatnot at giant companies. And the CTO would say, “Great, why should we pick AWS for our cloud needs?” And the answer is, “Oh, I’m sorry. You have 87 distinct accounts within your organization that we’ve [unintelligible 00:19:12] up for you. We’re just trying to offer you some management answers and unify the billing and this, and probably give you a discount as well because there is price breaks available at certain sizing.” It was a different conversation. It’s like, “I’m not here to sell you anything. We’re already there. We’re just trying to formalize the relationship.” And that is a challenge.

Again, I’m not trying to cast aspersions on procurement groups. I mean, I do sell enterprise consulting here at The Duckbill Group; we deal with an awful lot of procurement groups who have processes and procedures that don’t often align to the way that we do things as a ten-person, fully remote company. We do not have commercial vehicle insurance, for example, because we do not have a commercial vehicle and that is a prerequisite to getting the insurance, for one. We’re unlikely to buy one to wind up satisfying some contractual requirements, so we have to go back and forth and get things like that removed. And that is the nature of the beast.

And we can say yes, we can say no on a lot of those questionnaires, but, “It depends,” or, “I don’t know,” is the sort of thing that’s going to cause giant red flags and derail everything. But that is exactly what Docker is doing. Now, it’s the well, we have a sort of sloppy, weird set of habits with some of our engineers around the bring your own device to work thing. So, that’s the enterprise thing. Let me be very clear, here at The Duckbill Group, we have a policy of issuing people company machines, we manage them very lightly just to make sure the drives are encrypted, so they—and that the screensaver comes out with a password, so if someone loses a laptop, it’s just, “Replace the hardware,” not, “We have a data breach.”

Let’s be clear here; we are responsible about these things. But beyond that, it’s oh, you want to have some personal thing installed on your machine or do some work on that stuff? Fine. By all means. It’s a situation of we have no policy against it; we understand this is how work happens, and we trust people to effectively be grownups.

There are some things I would strongly suggest that any employee—ours or anyone else—not cross the streams on for obvious IP ownership rights and the rest, we have those conversations with our team for a reason. It’s, understand the nuances of what you’re doing, and we’re always willing to throw hardware at people to solve these problems. Not every company is like that. And ten million in revenue is not necessarily a very large company. I was doing the math out for ten million in revenue or 250 employees; assuming that there’s no outside investment—which with VC is always a weird thing—it’s possible—barely—to have a $10 million in revenue company that has 250 employees, but if they’re full time they are damn close to a $15 an hour minimum wage. So, who does it apply to? More people than you might believe.

Betty: Yeah, I’m really curious to how they’re going to like—like you say, if it takes place in three or four months, roll that out, and how would you actually track it and true that up for people? So.

Corey: Yeah. And there are tools and processes to do this, but it’s also not in anyone’s roadmap because people are not sitting here on their annual planning periods—which is always aspirational—but no one’s planning for, “Oh, yeah, Q3, one of our software suppliers is going to throw a real procurement wrench at us that we have to devote time, energy, resources, and budget to figure out.” And then you have a problem. And by resources, I do mean resources of basically assigning work and tooling and whatnot and energy, not people. People are humans, they are not resources; I will die on that hill.

Betty: Well, you know, actually resource-wise, the thing that’s interesting is when you say supplier, if it's something that people have been able to download for free so far, it’s not considered a supplier. So, it’s—now they’re going to go from just a thing I can use and maybe you’ve let your developers use to now it has to be something that goes through the official internal vetting as being a supplier. So, that’s just—it’s a whole different ball game entirely.

Corey: My last job before I started this place, was a highly regulated financial institution, and even grabbing things were available for free, “Well, hang on a minute because what license is it using and how is it going to potentially be incorporated?” And this stuff makes sense, and it’s important. Now, admittedly, I have the advantage of a number of my engineering peers in that I’ve been married to a corporate attorney for 11 years and have insight into that side of the world, which to be clear, is all about risk mitigation which is helpful. It is a nuanced and difficult field to—as are most things once you get into them—and it’s just the uncertainty that befuddles me a bit. I wish them well with it, truly I do. I think the world is better with an independent Docker in it, but I question whether this is going to find success. That said, it doesn’t matter what I think; what matters is what customers say and do, and I’m really looking forward to seeing how it plays out.

Betty: A hundred percent; same here. As someone who spent a good chunk of my life there, their mark on the industry is not to be ignored, like you said, with what happened with containers. But I do wish them well. There’s lot of good people over there, it’s some really cool tech, and I want to see a future for them.

Corey: One last topic I want to get into before we wind up wrapping this episode is that you are someone who was nominated to come on the show by a couple of folks, which is always great. I’m always looking for recommendations on this. But what’s odd is that you are—if we look at it and dig a little bit beneath the titles and whatnot, you even self-describe as your history is marketing leadership positions. It is uncommon for engineering-types to recommend that I talk to marketing folks.s personally I think that is a mistake; I consider myself more of a marketer than not in some respects, but it is uncommon, which means I have to ask you, what is your philosophy of marketing because it very clearly is differentiated in the public eye.

Betty: I’m flattered. I will say that—and this goes to how I hire people and how I coach teams—it’s you have to be super curious because there’s a ton of bad marketing out there, where it’s just kind of like, “Hey, we do these five things and we always do these five things: blah, blah, blah, blah, blah.” But I think it’s really being curious about what is the thing that you’re marketing? There are people who are just focused on the function of marketing and not the thing. Because you’re doing your marketing job in the service of a thing, this new widget, this new whatever, and you got to be super curious about it.

And I’ll tell you that, for me, it’s really hard for me to market something if I’m not excited about it. I have to personally be super excited about the tech or something happening in the industry, and it’s, kind of like, an all-in thing for me. And so in that sense, I do spend a ton of time with engineers and end-users, and I really try to understand what’s going on. I want to understand how the thing works, and I always ask them, “Well”—so I’ll ask the engineers, like, “So… okay, this sounds really cool. You just described this new feature and you’re super excited about it because you wrote it, but how is your end-user, the person you’re building this for, how did they do this before? Help me understand. How did they do this before and why is this better?”

Just really dig into it because for me, I want to understand it deeply before I talk about it. I think the thing is, it shows a tremendous amount of respect for the builder, and then to try to really be empathetic, to understand what they’re doing and then partner with them—I mean, this sounds so business-y the way I’m talking about this—but really be a partner with them and just help them make their thing really successful. I’m like the other end; you’re going to build this great thing and now I’m going to make it sound like it’s the best thing that’s ever happened. But to do that, I really need to deeply understand what it is, and I have to care about it, too. I have to care about it in the way that you care about it.

Corey: I cannot effectively market or sell something that I don’t believe in, personally. I also, to be clear because you are a marketing professional—or at least far more of one than I ever was—I do not view what I do is marketing; I view it as spectacle. And it’s about telling stories to people, it’s about learning what the market thinks about it, and that informs product design in many respects. It’s about understanding the product itself. It’s about being able to use the product.

And if people are listening to this and think, “Wait a minute, that sounds more like DevRel.” I have news for you. DevRel is marketing, they’re just scared to tell you that. And I know people are going to disagree with me on that. You’re wrong. But that’s okay; reasonable people can disagree.

And that’s how I see it is that, okay, I’ll talk to people building the service, I’ll talk to people using the service, but then I’m going to build something with the service myself because until then, it’s all a game of who sounds the most convincing in the stories that they tell. But okay, you can tell an amazing story about something, but if it falls over when I tried to use it, well, I’m sorry, you’re not being accurate in your descriptions of it.

Betty: A hundred percent. I hate to say, like, you’re storytellers, but that’s a big part of it, but it’s kind of like you want to tell the story, so you do something to that people believe a certain thing. But that’s part of a curated experience because you want them to try this thing in a certain way. Because you’ve designed it for something. “I built a spoon. I want you to use that to eat your soup because you can’t eat soup with a fork.”

So, then you’ll have this amazing soup-eating experience, but if I build you a spoon and then not give you any directions and you start throwing it at cars, you’re going to be like, “This thing sucks.” So, I kind of think of it in that way. To your point of it has to actually work, it’s like, but they also need to know, “What am I supposed to use it for?”

Corey: The problem I’ve always had on some visceral level with formal marketing departments for companies is that they can say that a product that they sell is good, they can say that the product is great, or they can choose to say nothing at all about that product, but when there’s a product in the market that is clearly a turd, a marketing department is never going to be able to say that, which I think erodes its authenticity in many respects. I understand the constraints behind, that truly I do, but it’s the one superpower I think that I bring to the table where even when I do sponsorship stuff it’s, you can buy my attention but not my opinion. Because the authenticity of me being trusted to call them like I see them, for lack of a better term, to my mind at least outweighs any short-term benefit from saying good things about a product that doesn’t deserve them. Now, I’ve been wrong about things, sure. I have also been misinformed in both directions, thinking something is great when it’s not, or terrible when it isn’t or not understanding the use case, and I am thrilled to engage in those debates. “But this is really expensive when you run for this use case,” and the answer can be, “Well, it’s not designed for that use case.” But the answer should not be, “No it’s not.” I promise you, expensive is in the eye of the customer not the person building the thing.

Betty: Yes. This goes back to I have to believe in the thing. And I do agree it’s, like not [sigh]—it’s not a panacea. You’re not going to make Product A and it’s going to solve everything. But being super clear and focused on what it is good for, and then please just try it in this way because that’s what we built it for.

Corey: I want to thank you for taking the time to have a what for some people is no doubt going to be perceived as a surprisingly civil conversation about things that I have loud, heated opinions about. If people want to learn more, where can they find you?

Betty: Well, they can follow me on Twitter. But um, I’d say go to vmware.com/cloud for our work thing.

Corey: Exactly. VM where? That’s right. VM there. And we will, of course, put links to that in the [show notes 00:30:07].

Betty: [laugh].

Corey: Thank you so much for taking the time to speak with me. I appreciate it.

Betty: Thanks, Corey.

Corey: Betty Junod, Senior Director of Multi-Cloud Solutions at VMware. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with a loud, ranting comment at the end. Then, if you work for a company that is larger than 250 people or $10 million in revenue, please also Venmo me $5.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Ed

Ed Boyajian, President and CEO of EDB, drives the development and execution of EDB’s strategic vision and growth strategy in the database industry, steering the company through 47 consecutive quarters of recurring revenue growth. He also led EDB’s acquisition of 2ndQuadrant, a deal that brought together the world’s top PostgreSQL experts and positioned EDB as the largest dedicated provider of PostgreSQL products and solutions worldwide. A 15+ year veteran of the open source software movement, Ed is a seasoned enterprise software executive who emphasizes that EDB must be a technology-first business in order to lead the open source data management ecosystem. Ed joined EDB in 2008 after serving at Red Hat, where he rose to Vice President and General Manager of North America. While there, he played a central leadership role in the development of the modern business model for bringing open source to enterprises.

Links:

  • EDB: https://enterprisedb.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Honeycomb. When production is running slow, it's hard to know where problems originate: is it your application code, users, or the underlying systems? I’ve got five bucks on DNS, personally. Why scroll through endless dashboards, while dealing with alert floods, going from tool to tool to tool that you employ, guessing at which puzzle pieces matter? Context switching and tool sprawl are slowly killing both your team and your business. You should care more about one of those than the other, which one is up to you. Drop the separate pillars and enter a world of getting one unified understanding of the one thing driving your business: production. With Honeycomb, you guess less and know more. Try it for free at Honeycomb.io/screaminginthecloud. Observability, it’s more than just hipster monitoring.

Corey: This episode is sponsored in part by our friends at Jellyfish. So, you’re sitting in front of your office chair, bleary eyed, parked in front of a powerpoint and—oh my sweet feathery Jesus its the night before the board meeting, because of course it is! As you slot that crappy screenshot of traffic light colored excel tables into your deck, or sift through endless spreadsheets looking for just the right data set, have you ever wondered, why is it that sales and marketing get all this shiny, awesome analytics and inside tools? Whereas, engineering basically gets left with the dregs. Well, the founders of Jellyfish certainly did. That’s why they created the Jellyfish Engineering Management Platform, but don’t you dare call it JEMP! Designed to make it simple to analyze your engineering organization, Jellyfish ingests signals from your tech stack. Including JIRA, Git, and collaborative tools. Yes, depressing to think of those things as your tech stack but this is 2021. They use that to create a model that accurately reflects just how the breakdown of engineering work aligns with your wider business objectives. In other words, it translates from code into spreadsheet. When you have to explain what you’re doing from an engineering perspective to people whose primary IDE is Microsoft Powerpoint, consider Jellyfish. Thats Jellyfish.co and tell them Corey sent you! Watch for the wince, thats my favorite part.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Today’s promoted episode is a treasure and a delight. Longtime listeners of this show know that it’s not really a database—unless of course, it’s Route 53—and of course, I don’t solve pronunciation problems with answers that make absolutely everyone hate me. Longtime listeners of the show know that if there’s one thing I adore when it comes to databases—you know, other than Route 53—it is solving pronunciation holy wars in such a way that absolutely everyone is furious with me as a result, and today is no exception. My guest is Ed Boyajian, the CEO of EDB, a company that effectively is the driving force behind the Postgres-squeal database. Ed, thank you for joining me.

Ed: Hey, Corey.

Corey: So, I know that other people pronounce it ‘post-gree,’ ‘Postgresql,’ ‘Postgres-Q-L,’ all kinds of other things. We know it’s decidedly not ‘Postgres-squeal,’ which is how I go for it. How do you pronounce it?

Ed: We say ‘Postgres,’ and this is one of the great branding challenges this fantastic open-source project has endured over many years.

Corey: So, I want to start at the very beginning because when I say that you folks are the driving force behind Postgres—or Postgres-squeal—I mean it. I’ve encountered folks from EDB—formerly EnterpriseDB—in the wild in consulting engagements before, and it’s great because whenever we found an intractable database problem, back at my hands-on keyboard engineering implementation days, very quickly after you folks got involved, it stopped being a problem, which is kind of the entire point. A lot of companies will get up there and say, “Oh, it’s an open-source project,” with an asterisk next to it and 15 other things that follow from it, or, “Now, we’re changing our license so the big companies can’t compete with us.” Your company’s not named after Postgres-squeal and you’re also—when you say you have people working on it, we’re not talking just one or two folks; your fingerprints are all over the codebase. How do you engage with an open-source project in that sense?

Ed: First and foremost, Postgres itself is, as you know, an independent open-source project, a lot like Linux. And that means it’s not controlled by a company. I think that’s inherently one of Postgres’s greatest strengths and assets. With that in mind, it means that a company like EDB—and this started when I came to the company; I came from Red Hat, so I’ve been in open-source for 20 years—when I came to the company back in 2008, it starts with a commitment and investment in bringing technology leaders in and around Postgres into a business like EDB, to help enterprises and customers. And that dynamic intersection between building the core database in the community and addressing customer needs in a business, at that intersection is where the magic happens. And we’ve been doing that since I joined EDB in 2008; it was really an explicit focus for the company.

Corey: I’d like to explore a little bit, well first and foremost, this story of is there a future for running databases in cloud environments yourself? And I have my own angry, loud opinion on this that I’m sure we’ll get to momentarily, but I want to start with yours. Who is writing their own databases in the Year of our Lord 2021, rather than just using whatever managed thing is their cloud provider of choice today is offering for them?

Ed: Well, let me give you context, Corey, because I think it matters. We’ve been bringing enterprise Postgres solutions to companies now, since the inception of the company, which dates back to 2004, and over that trajectory, we’ve been helping companies as they’ve done really two things: migrate away, in particular from Oracle, and land on Postgres, and then write new apps. Probably the first ten of the last 13 years since I’ve been in the company, the focus was in traditional on-prem database transformations that companies were going through. In the last three years, we’ve really seen an acceleration of that intersection of their traditional deployments and their cloud deployments. Our customers now, who are represented mostly in the Fortune 500 and Global 2000, 40% of our customers report they’re deploying EDB’s Postgres in the cloud, not in a managed context, but in a traditional EC2 or GCP self-managed cloud deployment.

Corey: And that aligns with what I’ve seen, a fair bit. Years ago, I wound up getting the AWS Cloud Practitioner Certification—did a whole blog post on it—not because it was opening any doors for me, but because it let me get into the certified lounge at re:Invent, and ideally charge a battery and have some mostly crappy coffee. The one question I got wrong was I was honest when I answered, “How long does it take to restore an RDS database from snapshot backup?” Rather than giving the by-the-book answer, which is way shorter than I found in practice a fair bit of the time. And that’s the problem I always ran into is that when you’re starting out and building something that needs a database, and it needs a relational database that runs in that model so all the no SQL options are not viable for whatever reason, great, RDS is great for getting you started, but there’s only so much that you can tune and tweak before you start to run into issues were, for particular workloads as they scale-out, it’s no longer a fit for a variety of reasons.

And most of the large companies that I work with that are heavily relational-database-driven have either started off or migrated to the idea of, “Oh, we’re going to run our own databases on top of EC2 instances,” for a variety of reasons that, again, the cloud providers will say, “Oh, that’s not accurate, and they’re doing the wrong thing.” But, you know, it takes a certain courage to tell a large-scale customer, “You’re doing it wrong.” “Well, why is that?” “Because I have things to sell you,” is kind of a terrible answer. How do you see it? Let’s not pick on RDS, necessarily, because all of the cloud providers offered managed database offerings. Where do those make sense and where do they fall down?

Ed: Yeah, I think many of our customers who made their first step into cloud picked a single vendor to do it, and we often hear AWS is been that early, early—

Corey: Yeah, a five-year head start makes a pretty compelling story.

Ed: That’s right. And let’s remember what these vendors are mostly. They are mostly infrastructure companies, they build massive data centers and set those up, and they do that beautifully well. And they lean on software, but they’re not software companies themselves. And I think the early implementation of many of our customers in cloud relied on what I’ll call relatively lightweight software offerings from their cloud vendor, including database.

They traded convenience, ease of use, an easy on-ramp, and they traded some capability in some depth for that. And it was a good trade, in fact. And for a large number of workloads it may still be a good trade. But our more sophisticated customers, enterprise customers who are running Postgres or databases at scale in their traditional environments have long depended on a very intimate relationship with their database technology vendor. And that relationship is the intersection of their evolving and emerging needs and the actual development of the database capabilities in support of that.

And that’s the heart of who we are at EDB and what we do with Postgres and the many people we have committed to doing that. And we don’t see our customers changing that appetite. So, I think for those customers, they’ve emerged more aware of the need to have a primary relationship with a database vendor and still be in cloud. And so I think that’s how this evolves to see two different kinds of services side-by-side, what they really want is a Database as a Service from the database vendor, which is what we just announced here at Microsoft Ignite event.

Corey: So, talk to me a little bit more about that, where it’s interesting in 2021 to see a company launching a managed service offering, especially in the database space, when there’s been so much pushback in different ways against the large cloud providers—[cough] Amazon—who tend to effectively lose sleep at night over the haunting fear that someone who isn’t them is making money, somehow. And they will take whatever is available to them and turn it into a managed service offering. That’s always been the fear, so people play games with licenses and the rest. Well, they’ve been running Postgres offerings for a long time. It is an independent open-source project.

I don’t think you can wind up forcing a license change through that says everyone except big companies can run this themselves and don’t do a managed service with it because that cat is very much out of the bag. How is it that you’re taking something to market now and expecting that to fare competitively?

Ed: So, I think there’s a few things that our customers are clearly telling us they want, and I think this is the most important thing: they want control of their data. And if you step back, Corey, look at it historically, they made a huge trade to big proprietary database companies, companies like Oracle, and they made that trade actually for convenience. They traded data to that database vendor. And we all know the
successes Oracle’s had, and the sheer extraordinary expense of those technologies. So, it felt like a walled garden.

And that’s where EDB and Postgres entered to really change that equation. What’s interesting is the re-platforming that happened and the transformation to cloud actually had the same, kind of, binding effect; we now moved all that data over to the public cloud vendors, arguably in an even stickier context, and now I think customers are realizing that’s created a dimension of inflexibility. It’s also created some—as you rightly pointed out—some deficiencies in technical depth, in database, and in software. So, our customers have sorted that out and are kind of coming back to middle. And what they’re saying is, “Well, we want all the advantages of an open-source database like a Postgres, but we want control of the data.”

And so what control looks like is more the ability to take one version of that software—in our case, we’re worrying about Postgres—and deploy the same thing everywhere they go. And that opens the door up for EDB to be their partner as a traditional on-prem partner, in the cloud where they run our Postgres and they manage it themselves, and as their managed service, Postgres Database as a Service Provider, which is what we’re doing.

Corey: I’ve been something of a bear on the idea of, “I’m going to build a workload to run everywhere in every cloud provider,” which I get. I think that’s generally foolish, and people chasing that, with remarkably few exceptions, are often going after the wrong thing. That said, I’m also a fan of having a path to strategic Exodus, where Google’s Cloud Spanner is fascinating, DynamoDB is revelatory, Cosmos DB is a security nightmare, which is neither here nor there, but the idea that I can take a provider’s offering that even if it solves a bunch of problems for me, well, if I ever need to move this somewhere else for any reason, I’m re-architecting, my data model and re-architecting the built-in assumptions around how the database acts and behaves, and that is a very heavy lift. We have proof of that from Amazon, who got up on stage and told a story about how much they hate Oracle, and they’re migrating everything off of Oracle to Aurora, which they had to build in order to get off of Oracle, and it took them three years to migrate things. And Oracle loves telling that story, too.

And it’s, you realize you both sound terrible when you tell that story? It’s, “This is a massive undertaking that even we struggle with, so you should probably not attempt it.” Well, what I hear from that is good God, don’t wind up getting locked into a particular database that is only available from one source. So, if you’re all-in on a cloud provider, which I’m a fan of, personally—I don’t care which one but pick a cloud provider—having a database that is not only going to work in that environment is just a reasonable step as far as how I view things. Trading up that optionality has got to pay serious dividends, and in many database use cases, I’ve just don’t see it.

Ed: Yeah, I think you’re bringing up a really important point. So, let’s unpack it for a minute.

Corey: Please.

Ed: Because I think you brought up some really prominent specialty database technologies, and I’m not sure there’s ever a way out of that intersection and commitment to a single vendor if you pick their specialty database. But underneath this is exactly one of the things that we’ve worried about here at EDB, which is to make Postgres a more capable, robust database in its entirety. A Postgres superpower is its ability to run a vast array of workloads. Guess what, it’s not sexy. It’s not sexy not to be that specialty database, but it’s incredibly powerful in the hands of an enterprise who can do more.

And that really creates an opportunity, so we’re trying to make Postgres apply to a much broader set of workloads, from traditional systems of record, like your ERP systems; systems of analysis, where people are doing lightweight analytic workloads or reporting, you can think in the world of data warehouse; and then systems of engagement, where customers are interacting with a website and have a database on the backend. All areas Postgres has done incredibly well in and we have customer experience with. So, when you separate out that core capability and then you look at it on a broader scale like Postgres, you realize that customers who want to make Postgres strategic, by definition need to be able to deploy it wherever they want to deploy it, and not be gated or bound by one cloud vendor. And all the cloud vendors picked up Postgres offerings, and that’s been great for Postgres and great for enterprises. But that corresponding lock-in is what people want to get away from, at this point.

Corey: There’s something to be said for acknowledging that there is a form of lock-in as far as technology selection goes. If you have a team of folks who are terrific at one database engine and suddenly you’re switching over to an entirely different database, well, folks who spent their entire career working on one particular database that’s still in widespread use are probably not super thrilled to stick around for that. Having something that can migrate from environment to environment is valuable and important. When you say you’re launching this as a database as a service offering, how does that actually work? Is that going to be running in your own cloud environment somewhere and people just make queries across the wire through standard connections to the database like they would something locally? Are you running inside of their account or environment? Is it something else?

Ed: So, this is a fully-managed database as a service, just like you’d get from any cloud vendor or DBAAS vendor that you’ve worked with in the past, just being managed and run by EDB. And with that, you get lot of the goodies that we bring, including our compatibility, and all our deep Postgres expertise, but I think one of the other important attributes is we’re going to run that service in our clients’ account, which gives them a level of isolation and a level of independence that we think is really important. And as different as that is, it’s not heroic; it’s exactly what our customers told us they wanted.

Corey: There’s something to be said for building the thing that your customers have said that they want and make sense for you to build as opposed to, “We’re going to build this ridiculous thing and we’re sure folks are going to love it.” It’s nice to see that shaping up in the proper order. And I’ve fallen victim to that myself; I think most technologists have to some extent. How big is EDB these days?

Ed: So, we have over 650 employees. Now, around the world, we have 6000 customers. And of the 650 employees, about 300 of those are focused on Postgres. A subset of that are 30-odd core team members in the Postgres community, committers in the Postgres community, major contributors, and contributors in the Postgres community. So, we have a density of technical depth that is really unparalleled in Postgres.

Corey: You’re not, for lack of a better term, pulling an Amazon, insofar as you’re, “Well, we have three people working on open-source projects, so we’re going to go ahead and claim we’re an open-source company,” in other words. Conversely, you’re also not going down the path of this is a project that you folks have launched, and it claims to be open-source because we love it when people volunteer for for-profit entities, but we exercise total control over the project. You have a lot of contributors, but you’re also still a minority, I think the largest minority, but still a minority of people contributing to Postgres.

Ed: That’s right. And, look, we’re all-in on Postgres, and it’s been that way since I got here. As I mentioned earlier, I came from Red Hat where I was—I was at Red Hat for a little over six years, so I’ve been an open-source now for 20 years. So, my orientation is towards really powerful, independent open-source projects. And I think we’ll see Postgres really be the most transformative open-source technology since Linux.

I think we’ll see that as we look forward. And you’re right, though, I think what’s powerful about Postgres is it’s an independent project, which means it’s supported by thousands of contributors who aren’t tied to single companies, around the world. And it just makes the software—we develop innovation faster, and I think it makes the software better. Now, EDB plays a big part in there. Roughly, a little less than a third of the last res—actually, the 13 release—were contributions that came from contributors who came from EDB.

So, that’s not a majority, and that’s healthy. But it’s a big part of what helps move Postgres along and there aren’t—you know, the next set of companies are much, much—next set of combined contributors add up to quite small numbers. But the cloud vendors are virtually non-existent in that contribution.

Corey: This episode is sponsored in part by something new. Cloud Academy is a training platform built on two primary goals. Having the highest quality content in tech and cloud skills, and building a good community the is rich and full of IT and engineering professionals. You wouldn’t think those things go together, but sometimes they do. Its both useful for individuals and large enterprises, but here's what makes it new. I don’t use that term lightly. Cloud Academy invites you to showcase just how good your AWS skills are. For the next four weeks you’ll have a chance to prove yourself. Compete in four unique lab challenges, where they’ll be awarding more than $2000 in cash and prizes. I’m not kidding, first place is a thousand bucks. Pre-register for the first challenge now, one that I picked out myself on Amazon SNS image resizing, by visiting cloudacademy.com/corey. C-O-R-E-Y. That’s cloudacademy.com/corey. We’re gonna have some fun with this one!

Corey: Something else that does strike me as, I guess, strange, just because I’ve seen so many companies try to navigate this in different ways with varying levels of success. I always encountered EDB—even back when it was EnterpriseDB, which was, given their love of acronyms, I’m still somewhat partial to. I get it; branding, it’s a thing—but the folks that I engaged with were always there in a consulting service’s capacity, and they were great at this. Is EDB a services company or a product company?

Ed: Yeah, we are unashamedly a product technology company. Our business is over 90% of our revenue is annually recurring subscription revenue that comes from technical products, database server, mostly, but then various adjacent capabilities in replication and other areas that we add around the database server itself. So no, we’re a database technology company selling a subscription. Now, we help our customers, so we do have a really talented team of consultants who help our customers with their business strategy for Postgres, but also with migrations and all the things they need to do to get Postgres up and running.

Corey: And the screaming, “Help, help, help, fix it, fix it, fix it now,” emergencies as well.

Ed: I think we have the best Postgres support operation in the world. It is a global 24/7 organization, and I think a lot of what you likely experienced, Corey, came out of our support organization. So, our support guys, these guys aren’t just handling lightweight issues. I mean, they wade into the gnarly questions and challenges that customers face. But that’s a support business for us. So, that’s part and parcel. You get that, it’s included with the subscription.

Corey: I would not be remembering this for 11 years later, if it hadn’t been an absolutely stellar experience—or a horrible experience, for that matter; one or the other. You remember the superlatives, not the middle of the road ones—and if it hadn’t been important. And it was. It also noteworthy; with many vendors that are product-focused, their services may have an asterisk next to it because it’s either a, “Buy our product and then we’ll support it,” or it’s, “Ohh, we’re going to sell you a whole thing just to get us on the phone.” And as I recall, there wasn’t a single aspect of upsell involved in this.

It was, “Let’s get you back up and running and solve the problem.” Sure, later in time, there were other conversations, as all good businesses will have, but there was no point during those crisis moments where it felt like, “Oh, if you had gone ahead and bought this thing that we sell, this wouldn’t happen,” or, “You need to buy this or we won’t help you.” I guess that’s why I’ve contextualized you folks as a services company, first and foremost.

Ed: Well, I’m glad you have that [laugh] experience because that’s our goal. And I think—look, this is an interesting point where customers want us to bring that capability to their managed DBAAS world. Step back again, go back to what I said about the big cloud vendors; they are, at their core, infrastructure companies. I mean, they’re really good at that. They’re not particularly well-positioned to take your Postgres call, and I don’t think they want that call.

We’re the other guys; we want to help you run your Postgres, at scale, on-prem, in the cloud, fully managed in the cloud, by EDB, and solve those problems at the same time. And I think that’s missing in the market today. And we can step back and look at this overall cloud evolution, and I think some might think, “Gee, we’re into the mature phase of cloud adoption.” I would tell you, since the Red Sox have done well this year, I think in a nine-inning baseball game—for those of your listeners who follow American baseball—we’re in, like, the top of the second inning, maybe. Maybe the bottom of the second inning. So, we’ve been able to listen and learn from the experiences our customers have had. I think that’s an incredible advantage as we now firmly plant ourselves in the cloud DBAAS market alongside our robust Postgres capabilities that you experienced.

Corey: The world isn’t generating less data, and it’s important that we’re able to access that in a bunch of different ways. And the last time I really was playing with relational databases, you can view my understanding of it as Excel with a weirder interface, and you’re mostly there. One thing that really struck me since the last time I went deep into database-land over in the Postgres-squeal world has been just the sheer variety of native data types that it winds up supporting. The idea of, “Here’s some JSON. Take this and store it that way,” or it’s GIS data that it can represent, or the idea of having data types that are beyond just string or var or whatever other somewhat limited boolean values or whatnot. Without having just that traditional list, which is of course all there as well. It also seems to have extensively improved its coverage that just can only hint to my small mind about these things and what sort of use cases people are really putting these things into.

Ed: Yeah, I think this is one of Postgres’ superpowers. And it started with Mike Stonebraker’s original development of Postgres as an object-relational database. Mike is an adviser to EDB, which has been incredibly helpful as we’ve continued to evolve our thinking about what’s possible in Postgres. But I think because of that core technology, or that core—because of that core technical capability within Postgres, we have been able to build a whole host of data types. And so now you see Postgres being used not just as the context of a traditional relational database, but we see it used as a time-series database. You pointed out a geospatial database, more and more is a document-oriented database with JSON and JSONB.

These are all the things that make Postgres have much more universal appeal, universal appeal to developers—which is worth talking about in the recent StackOverflow developer survey, but we can come back to that—and I think universal applicability for new applications. This is what’s bringing Postgres forward faster, unlike many of the specialty database companies that you mentioned earlier.

Corey: Now, this is something that you can use for your traditional CRUD app, the my first hello world app that returns something from a database, yeah, that stuff works. But it also, for example, has [cyter 00:25:09] data types, where you can say, give me the results where the IP range contains this address, and it’ll do that. Before that, you’re trying to solve a whole bunch of very messy things in application logic that’s generally awful. The database now does that for you automatically, and there’s something—well, it would if I were smart and used it instead of storing it as strings because I make terrible life choices, but for sensible people, it solves a lot of those problems super well. And it’s taken the idea of where logic should live in application versus database, and sort of turn a lot of those assumptions I was starting my career with on their head.

Ed: Yeah, I think if you look now at the appeal of Postgres to developers, which we’ve paid a lot of attention to—one of our stated strategies at EDB is to make Postgres easier. That’s been true for many years, so a drive for engineering and development here has been that call to action. And if you measure that, over time, we’ve been contributing—not alone, but contributing to making Postgres more approachable, easier to use, easier to engage with. Some of those things we do just through edb.com, and the way we handle EDB docs is a great example of that, and our developer advocacy and outreach into adjacent communities that care about Postgres. But here’s where that’s landed us. If you looked at the last Stack Overflow developer survey—the 2021 Stack Overflow developer survey, which I love because I think it’s very independent-oriented—and they surveyed, I think this past year was 80,000 developers.

Corey: Oh yeah, if Stack Overflow is captured by any particular constituency, it’s got to be ‘Big Copy and Paste’ that is really behind them. But yeah, other than the cabal of keyboard manufacturers for those copy-and-paste stories, yeah, they’re fairly objective when it comes to stuff like this.

Ed: And if you look at that survey, Corey, if you just took and summed it because it’s helpful to sum it, most used, most loved, and most wanted database: Postgres wins. And I find it fascinating that if you—having been here, in this company for 13 years and watch the evolution from—you know, 13 years ago, Postgres needed help, both in terms of its awareness in the market and some technical capabilities it just lacked, we’ve come so far. For that to be the new standard for developers, I think, is a remarkable achievement. And I think it’s a representation of why Postgres is doing so well in the market that we’ve long served, in the cloud market that we are now serving, and I think it speaks to what’s ahead as a transformational database for the future.

Corey: There really is something to be said for a technology as—please don’t take this term the wrong way—old. As a relational database, Postgres has been around for a very long time, but it’s also not your grandparents’ Postgres. It is continuing to evolve. It continues to be there in a bunch of really interesting ways for developers in a variety of different capacities, and it’s not the sort of thing that you’re only using in, “Legacy environments,” quote-unquote. Instead, it’s something that you’ll see all over the place. It is rare that I see an environment that doesn’t have Postgres in it somewhere these days.

Ed: Yeah, I think quite the contrary to the old-school database, which I love that; I love that shade because when you step away from it, you realize, the Postgres community represents the very best of what’s possible with open-source. And that’s why Postgres continues to accelerate and move forward at the rate that it does. And obviously, we’re proud to be a contributor to that, so we don’t just watch that outcome happen; we’re actually part of creating it. But I also think that when you see all that Postgres has become and where it’s going, you really start to understand why the market is adopting open-source.

Corey: It’s one of those areas where even if some company comes out with something that is amazing and transformatively better, and you should jump into it with both feet and never look back, yeah, it turns out that it takes a long time to move databases, even when they’re terrible. And you can lobby an awful lot of accusations at Postgres—or Postgres-squeal—but you can’t call it terrible. It’s used in enough interesting applications by enough large-scale companies out there—and small as well—that it’s very hard to find a reason not to explore it. It’s my default relational database when Route 53 loses steam. It just makes sense in a bunch of ways that other things really didn’t for me before.

Ed: Yeah, and I think we’ll continue to see that. And we’re just going to keep making Postgres better. And it gets better because of that intersection, as I mentioned, that intimate intersection between enterprise users, and the project, and the community, and the bridge that a company like EDB provides for that. That’s why it’ll get better faster; the breadth of use of Postgres will keep it accelerating. And I think it’s different than many of the specialty databases.

Look, I’ve been in open-source now for 20 years and it’s intriguing to me how many new specialty open-source databases have come to market. We tend to forget the amount of roadkill we’ve had over the course of the past ten years of some of those open-source projects and companies. We certainly are tuned into some of the more prolific ones, even today. And I think again, here again, this is where Postgres shines, and where I think Postgres is a better call for a long-term. Just like Linux was.

Corey: I want to thank you for taking so much time out of your day to talk to me about databases, which given my proclivities, is probably like pulling teeth for you. If people want to learn more, where can they find you?

Ed: So, come to enterprisedb.com. You still get EnterpriseDB, Corey. Just come to enterprise—

Corey: There we go. It’s hidden in the URL, right in plain sight.

Ed: Come to enterprisedb.com. You can learn all the things you need about the technology, and certainly more that we can do to help you.

Corey: And we will, of course, put links to that in the [show notes 00:31:10]. Thank you once again for your time. I really do appreciate it.

Ed: Thanks, Corey. My pleasure.

Corey: Ed Boyajian, CEO of EDB. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with a long angry comment because you are one of the two Amazonian developers working on open-source databases.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Mark

Mark loves to teach and code.

He is an award winning university instructor and engineer. He comes with a passion for creating meaningful learning experiences. With over a decade of developing solutions across the tech stack, speaking at conferences and mentoring developers he is excited to continue to make an impact in tech. Lately, Mark has been spending time as a Developer Relations Engineer on the Angular Team.

Links:

  • Twitter: https://twitter.com/marktechson

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Vultr. Spelled V-U-L-T-R because they’re all about helping save money, including on things like, you know, vowels. So, what they do is they are a cloud provider that provides surprisingly high performance cloud compute at a price that—while sure they claim its better than AWS pricing—and when they say that they mean it is less money. Sure, I don’t dispute that but what I find interesting is that it’s predictable. They tell you in advance on a monthly basis what it’s going to going to cost. They have a bunch of advanced networking features. They have nineteen global locations and scale things elastically. Not to be confused with openly, because apparently elastic and open can mean the same thing sometimes. They have had over a million users. Deployments take less that sixty seconds across twelve pre-selected operating systems. Or, if you’re one of those nutters like me, you can bring your own ISO and install basically any operating system you want. Starting with pricing as low as $2.50 a month for Vultr cloud compute they have plans for developers and businesses of all sizes, except maybe Amazon, who stubbornly insists on having something to scale all on their own. Try Vultr today for free by visiting: vultr.com/screaming, and you’ll receive a $100 in credit. Thats v-u-l-t-r.com slash screaming.

Corey: This episode is sponsored in part by something new. Cloud Academy is a training platform built on two primary goals. Having the highest quality content in tech and cloud skills, and building a good community the is rich and full of IT and engineering professionals. You wouldn’t think those things go together, but sometimes they do. Its both useful for individuals and large enterprises, but here's what makes it new. I don’t use that term lightly. Cloud Academy invites you to showcase just how good your AWS skills are. For the next four weeks you’ll have a chance to prove yourself. Compete in four unique lab challenges, where they’ll be awarding more than $2000 in cash and prizes. I’m not kidding, first place is a thousand bucks. Pre-register for the first challenge now, one that I picked out myself on Amazon SNS image resizing, by visiting cloudacademy.com/corey. C-O-R-E-Y. That’s cloudacademy.com/corey. We’re gonna have some fun with this one!

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Anyone who has the misfortune to follow me on Twitter is fairly well aware that I am many things: I’m loud, obnoxious, but snarky is most commonly the term applied to me. I’ve often wondered, what does the exact opposite of someone who is unrelentingly negative about things in cloud look like? I’m here to answer that question is lightness and happiness and friendliness on Twitter, personified. His Twitter name is @marktechson. My guest today is Mark Thompson, developer relations engineer at Google. Mark, thank you for joining me.

Mark: Oh, I’m so happy to be here. I really appreciate you inviting me. Thanks.

Corey: Oh, by all means. I’m glad we’re doing these recordings remotely because I strongly suspect, just based upon the joy and the happiness and the uplifting aspects of what it is that you espouse online that if we ever shook hands, we’d explode as we mutually annihilate each other like matter and antimatter combining.

Mark: Feels right. [laugh].

Corey: So, let’s start with the day job; seems like the easy direction to go in. You’re a developer relations engineer. Now, I’ve heard of developer advocates, I’ve heard of the DevRel term, a lot of them get very upset when I refer to them as ‘devrelopers’, but that’s the game that we play with language. What is the developer relations engineer?

Mark: So, I describe my job this way: I like to help external communities with our products. I work on the Angular team, so I like to help our external communities but then I also like to work with our internal team to help improve our product. So, I see it as helping as a platform, as a developer relations engineer. But the engineer part is, I think, is important here because, at Google, we still do coding and we still write things; I’m going to contribute to the Angular platform itself versus just only giving talks or only writing blog posts to creating content, they still want us to do things like solve problems with the platform as well.

Corey: So, this is where my complete and abject lack of understanding of the JavaScript ecosystem enters the conversation. Let’s be clear here, first let me check my assumptions. Angular is a JavaScript framework, correct?

Mark: Technically a TypeScript framework, but you could say JavaScript.

Corey: Cool. Okay, again, this is not me setting you up for a joke or anything like that. I try to keep my snark to Twitter, not podcast because that tends to turn an awful lot into me berating people, which I try to reserve for those who really have earned it; they generally have the word chief somewhere in their job title. So, I’m familiar with sort of an evolution of the startups that I worked at where Backbone was all the rage, followed by, “Oh, you should never use Backbone. You should be using Angular instead.”

And then I sort of—like, that was the big argument the last time I worked in an environment like that. And then I see things like View and React and several other things. At some point, it seems like, pick a random name out of the air; if it’s not going to be a framework, it’s going to be a Pokemon. What is the distinguishing characteristic or characteristics of Angular?

Mark: I like to describe Angular to people is that the value-add is going to be some really incredible developer ergonomics. And when I say that I’m thinking about the tooling. So, we put a lot of work into making sure that the tooling is really strong for developers, where you can jump in, you can get started and be productive. Then I think about scale, and how your application runs at scale, and how it works at scale for your teams. So, scale becomes a big part of the story that I tell, as well, for Angular.

Corey: You spend an awful lot of time telling stories about Angular. I’m assuming most of them are true because people don’t usually knowingly last very long in this industry when they just get up on stage and tell lies, other than, “This is how we do it in our company,” which is the aspirational conference-ware that we all wish we ran. You’re also, according to your bio, which of course, is always in the [show notes 00:04:16], you’re an award-winning university instructor. Now, award-winning—great. For someone who struggled mightily in academia, I don’t know much about that world. What is it that you teach? How does being a university instructor work? I imagine it’s not like most other jobs where you wind up showing up, solving algorithms on a whiteboard, and they say, “Great, can you start tomorrow?”

Mark: Sure. So, when I was teaching at university, what I was teaching was mostly coding bootcamps. So, some universities have coding bootcamps that they run themselves. And so I was a part of some instructional teams that work in the university. And that’s how I won the Teaching Excellence Award. So, the award that I won actually was the Distinguished Teaching Excellence Award, based on my performance at work when I was teaching at university.

Corey: I want to be clear here, it’s almost enough to make someone question whether you really were involved there because the first university, according to your background that you worked on was Northwestern, but then it was through the Harvard Extension School, and I was under the impression that doing anything involving Harvard was the exact opposite of an NDA, where you’re contractually bound to mention that, “Oh, I was involved with Harvard in the following way,” at least three times at any given conversation. Can you tell I spent a lot of time dealing with Harvard grads?

Mark: [laugh]. Yeah, Harvard is weird like that, where people who’ve worked there or gone there, it comes up as a first thing. But I’ll tell the story about it if someone asks me, but I just like to talk about univer—that’s why I say ‘university,’ right? I don’t say, “Oh, I won an award at Northwestern.” I just say, “University award-winning instructor.”

The reason I say even the ‘award-winning’, that part is important for credibility, specifically. It’s like, hey, if I said I’m going to teach you something, I want you to know that you’re in really good hands, and that I’m really going to do my best to help you. That’s why I mention that a lot.

Corey: I’ll take that even one step further, and please don’t take this as in any way me casting aspersions on some of your colleagues, but very often working at Google has felt an awful lot like that in some respects. I’ve never seen you do it. You’ve never had to establish your bona fides in a conversation that I’ve seen by saying, “Well, at Google this is how we do it.” Because that’s a logical fallacy of appeal to authority in many respects. Yeah, I’m sure you do a lot of things at Google at a multinational trillion-dollar company that if I’m founding a four-person startup called Twitter for Pets might not necessarily be the same constraints that I’m faced with.

I’m keenly appreciative folks who recognize that distinction and don’t try and turn it into something else. We see it with founders, too, “Oh, we’re a small scrappy startup and our founders used to work at Google.” And it’s, “Hmm, I’m wondering if the corporate culture at a small startup might be slightly different these days.” I get it. It does resonate and it carries weight. I just wonder if that’s one of those unexamined things that maybe it’s time to dive into a bit more.

Mark: Hmm. So, what’s funny about that is—so people will ask me, what do I do? And it really depends on context. And I’ll usually say, “Oh, I work for a company on the West Coast,” or, “For a tech company on the West Coast.” I’ll just say that first.

Because what I really want to do is turn the conversation back to the person I’m talking to, so here’s where that unrelenting positivity kind of comes in because I’m looking at ways, how can I help boost you up? So first, I want to hear more about you. So, I’ll kind of like—I won’t shrink myself, but I’ll just be kind of vague about things so I could hear more about you so we’re not focused on me. In this case, I guess we are because I’m the guest, but in a normal conversation, that’s what I would try to do.

Corey: So, we’ve talked about JavaScript a little bit. We’ve talked about university a smidgen. Now, let me complete the trifecta of things that I know absolutely nothing about, specifically positivity on Twitter. You have been described to me as the mayor of wholesome Twitter. What is that about?

Mark: All right, so let me be really upfront about this. This is not about toxic positivity. We got to get that out in the open first, before I say anything else because I think that people can hear that and start to immediately think, “Oh, this guy is just, you know, toxic positivity where no matter what’s happening, he’s going to be happy.” That is not the same thing. That is not the same thing at all.

So, here’s what I think is really interesting. Online, and as you know, as a person on Twitter, there’s so many people out there doing damage and saying hurtful things. And I’m not talking about responding to someone who’s being hurtful by being hurtful. I mean the people who are constantly harassing women online, or our non-binary friends, people who are constantly calling into question somebody’s credibility because of, oh, they went to a coding bootcamp or they came from self-taught. All these types of ways to be really just harmful on Twitter.

I wanted to start adding some other perspective of the positivity side of just being focused on value-add in our interactions. Can I craft this narrative, this world, where when we meet, we’re both better off because of it, right? You feel good, I feel good, and we had a really good time. If we meet and you’re having a bad time, at least you know that I care about you. I didn’t fix you. I didn’t, like, remove the issue, but you know that somebody cares about you. So, that’s what I think wholesome positivity comes into play is because I want to be that force online. Because we already have plenty of the other side.

Corey: It’s easy for folks who are casual observers of my Twitter nonsense to figure, “Oh, he’s snarky and he’s being clever and witty and making fun of big companies”—which I do–And they tend to shorthand that sometimes to, “Oh, great. He’s going to start dunking on people, too.” And I try mightily to avoid that it’s punch up, never down.

Mark: Mm-hm.

Corey: I understand there’s a school of thought that you should never be punching at all, which I get. I’m broken in many ways that apparently are entertaining, so we’re going to roll with that. But the thing that incenses me the most—on Twitter in my case—is when I’ll have something that I’ll put out there that’s ideally funny or engaging and people like it and it spreads beyond my circle, and then you just have the worst people on the internet see that and figure, “Oh, that’s snarky and incisive. Ah, I’m like that too. This is my people.”

I assure you, I am not your people when that is your approach to life. Get out of here. And curating the people who follow and engage with you on Twitter can be a full-time job. But oh man, if I wind up retweeting someone, and that act brings someone who’s basically a jackwagon into the conversation, it’s no. No-no-no.

I’m not on Twitter to actively make things worse unless you’re in charge of cloud pricing, in which case yes, I am very much there to make your day worse. But it’s, “Be the change you want to see in the world,” and lifting people up is always more interesting to me than tearing people down.

Mark: A thousand percent. So, here’s what I want to say about that is, I think, punching up is fine. I don’t like to moderate other people’s behavior either, though. So, if you’d like punching up, I think it’d be funny. I laugh at jokes that people make.

Now, is it what I’ll do? Probably not because I haven’t figured out a good way for me to do it that still goes along my core values. But I will call out stuff. Like if there’s a big company that’s doing something that’s pretty messed up, I feel comfortable calling things out. Or when drama happens and people are attacking someone, I have no problem with just be like, “Listen, this person is a stand-up person.”

Putting myself kind of like… just kind of on the front line with that other person. Hey, look, this person is being attacked right now. That person is stand-up, so if you got a problem them, you got a problem with me. That’s not the same thing as being negative, though. That’s not the same thing as punching down or harming people.

And I think that’s where—like I say, people kind of get that part confused when they think that being kind to people is a sign of weakness, which is—it takes more strength for me to be kind to people who may or may not deserve it, by societal standards. That I’ll try to understand you, even though you’ve been a jerk right now.

Corey: Twitter excels at fomenting outrage, and it does it by distancing us from being able to easily remember there’s a person on the other side of these things. It is ways you’re going to yell at someone, even my business partner in a text message. Whenever we start having conversations that get a little heated—which it happens; business partnership is like a marriage—it’s oh, I should pick up the phone and call him rather than sending things that stick around forever, that don’t reflect the context of the time, and five years later when I see it, I feel ashamed." I’m not here to advocate for other people doing things on Twitter the way that I do because what I do is clever, but the failure mode of clever in my case is being a complete jerk, and I’ve made that mistake a lot when I was learning to do it when my audience was much
smaller, and I hurt people. And whenever I discovered that that is what happened, I went out of my way, and still do, to apologize profusely.

I’ve gotten relatively good at having to do less of those apologies on an ongoing basis, but very often people see what I’m doing and try to imitate what they’re seeing; it just comes off as mean. And that’s not acceptable. That’s not something that I want to see more of in the world. So, those are my failure modes. I have to imagine the only real failure mode that you would encounter with positivity is inadvertently lifting someone up who turns out to be a trash goblin.

Mark: [laugh]. That and I think coming off as insincere. Because if someone is always positive or a majority of the time, positive, if I say something to you, and you don’t know me that actually mean it, sincerity is incredibly hard to get over text. So, if I congratulate you on your job, you might be like, “Oh, he’s just saying that for attention for himself because now he’s being the nice guy again.” But sincerity is really, really hard to convey, so that’s one of the failure modes is like I said, being sincere.

And then lifting up people who don’t deserve to be lifted up, yeah, that’s happened before where I’ve engaged with people or shared some of their stuff in an effort to boost them, and find out, like you said, legit trash goblin, like, their home address is under a bridge because they’re a troll. Like, real bad stuff. And then you have back off of that endorsement that you didn’t know. And people will DM you, like, “Hey, I see that you follow this person. That person is a really bad person. Look at what they’re saying right now.” I’m like, “Well, damn, I didn’t know it was bad like that.”

Corey: I’ve had that on the podcast, too, where I’ll have a conversation with someone and then a year or so later, they’ll wind up doing something horrifying, or something comes to light and the rest, and occasionally people will ask, “So, why did you have that person on this show?” It’s yeah, it turns out that when we’re having a conversation, that somehow didn’t come up because as I’m getting background on people and understanding who they are and what they’re about in the intake questionnaire, there is not a separate field for, “Are you terrible to women?” Maybe there should be, but that’s something that it’s—you don’t see it. And that makes it easy to think that it’s not there until you start listening more than you speak, and start hearing other people’s stories about it. This is the challenge.

As much as I aspire at times to be more positive and lift folks up, this is the challenge of social media as it stands now. I had a tweet the other day about a service that AWS had released with the comment that this is fantastic and the team that built it should be proud. And yeah, that got a bit of engagement. People liked it. I’m sure it was passed around internally, “Yay, the jerk liked something.” Fine.

A month ago, they launched a different service, and my comment was just distilled down to, “This is molten garbage.” And that went around the tech internet three times. When you’re positive, it’s one of those, “Oh, great. Yeah, that’s awesome.” Whereas when I savage things, it’s, “Hey, he’s doing it again. Come and look at the bodies.” Effectively the rubbernecking thing. “There’s been a terrible accident, let’s go gawk at it.”

Mark: Right.

Corey: And I don’t quite know what to do with that because it leads to the mistaken and lopsided impression that I only ever hate things and I don’t think that a lot of stuff is done well. And that’s very much not the case. It doesn’t restrict itself to AWS either. I’m increasingly impressed by a lot of what I’m seeing out of Google Cloud. You want to talk about objectivity, I feel the same way about Oracle Cloud.

Dunking on Oracle was a sport for me for a long time, but a lot of what they’re doing on a technical and on a customer-approach basis in the cloud group is notable. I like it. I’ve been saying that for a couple of years. And I’m gratified the response from the audience seems to at least be that no one’s calling me a shill. They’re saying, “Oh, if you say it, it’s got to be true.” It’s, “Yes. Finally, I have a reputation for authenticity.” Which is great, but that’s the reason I do a lot of the stuff that I do.

Mark: That is a tough place to be in. So, Twitter itself is an anomaly in terms of what’s going to get engagement and what isn’t. Sometimes I’ll tweet something that at least I think is super clever, and I’m like, “Oh, yeah. This is meaningful, sincere, clever, positive. This is about to go bananas.” And then it’ll go nowhere.

And then I’ll tweet that I was feeling a depression coming on and that’ll get a lot of engagement. Now, I’m not saying that’s a bad thing. It’s just, it’s never what I think. I thought that the depression tweet was not going to go anywhere. I thought that one was going to be like, kind of fade into the ether, and then that is the one that gets all the engagement.

And then the one about something great that I want to share, or lifting somebody else up, or celebrating somebody that doesn’t go anywhere. So, it’s just really hard to predict what people are going to really engage with and what’s going to ring true for them.

Corey: Oh, I never have any idea of how jokes are going to land on Twitter. And in the before times, I had the same type of challenge with jokes in conference talks, where there’s a joke that I’ll put in there that I think is going to go super well, and the audience just sits there and stares. That’s okay. My jokes are for me, but after the third time trying it with different audiences and no one laughs, okay, I should keep it to myself, then. Other times just a random throwaway comment, and I find it quoted in the newspaper almost. And it’s, “Oh, okay.”

Mark: [laugh].

Corey: You can never tell what’s going to hit and what isn’t.

Mark: Can we talk about that though? Like—

Corey: Oh, sure.

Mark: Conference talking?

Corey: Oh, my God, no.

Mark: Conference speaking, and just how, like—I remember one time I was keynoting—well I was emceeing and I had the opening monologue. And so [crosstalk 00:17:45]—

Corey: We call that a keynote. It’s fine. It is—I absolutely upgrade it because people know what you’re talking about when you say, “I keynoted the thing.” Do it. Own it.

Mark: Yeah.

Corey: It’s yours.

Corey: So, I was emcee and then I did the keynote. And so during the keynote rehearsals—and this is for all the academia, right, so all these different university deans, et cetera. So, in the practice, I’m telling this joke, and it is landing, everybody’s laughing, blah, blah, blah. And then I get in there, and it was crickets. And in that moment, you want to panic because you’re like, “Holy crap, what do I do because I was expecting to be able to ride the wave of the laughter into my next segment,” and now it’s dead silent. And then just that ability to have to be quick on your feet and not let it slow you down is just really hard.

Corey: This episode is sponsored by our friends at Oracle HeatWave is a new high-performance accelerator for the Oracle MySQL Database Service. Although I insist on calling it “my squirrel.” While MySQL has long been the worlds most popular open source database, shifting from transacting to analytics required way too much overhead and, ya know, work. With HeatWave you can run your OLTP and OLAP, don’t ask me to ever say those acronyms again, workloads directly from your MySQL database and eliminate the time consuming data movement and integration work, while also performing 1100X faster than Amazon Aurora, and 2.5X faster than Amazon Redshift, at a third of the cost. My thanks again to Oracle Cloud for sponsoring this ridiculous nonsense.

Corey: It’s a challenge. It turns out that there are a number of skills that are aligned but are not the same when it comes to conference talks, and I think that is something that is not super well understood. There’s the idea of, “I can get on stage in front of a bunch of people with a few loose talking points, and just riff,” that sort of an improv approach. There’s the idea of, “Oh, I can get on stage with prepared slides and have presenter notes and have a whole direction and theme of what I’m doing,” that’s something else entirely. But now we’re doing video and the energy is completely different.

I’ve presented live on video, I’ve done pre-recorded video, but in either case, you’re effectively talking to the camera and there is no crowd feedback. So, especially if you’d lean on jokes like I tend to, you can’t do a cheesy laugh track as an insert, other than maybe once as its own joke. You have to make sure that you can resonate and engage with folks, but there are no subtle cues from the audience like half the front row getting up and walking out. You have to figure out what it is that resonates, what it is that doesn’t, why people should care. And of course, distinguishing and differentiating between this video that you’re watching now and the last five Zoom meetings that you’ve been on that look an awful lot the same; why should you care about this talk?

Mark: The hardest thing to do. I think speaking remotely became such a big challenge. So, over time it became a little easier because I found some of the value in it, but it was still much harder because of all the things that you said. What became easier was that I didn’t have to go to a place. That was easier.

So, I could take three different conference talks in a day for three different organizations. So, that was easier. But what was harder, just like you said, not being able to have that energy of the crowd to know when you’re on point because you look for that person in the audience who’s nodding in agreement, or the person who’s shaking their head furiously, like, “Oh, this is all wrong.” So, you might need to clarify or slow down or—you lose all your cues, and that’s just really, really hard. And I really don’t like doing video pre-recorded talks because those take more energy for me than they do the even live virtual because I have to edit it and I have to make sure that take was right because I can’t say, “Oh, excuse me. Well, I meant to say this.”

And I guess I could leave that in there, but I’m too much of a—I love public speaking, so I put so much pressure on myself to be the best version of myself at every opportunity when I’m doing public speaking. And I think that’s what makes it hard.

Corey: Oh, yeah. Then you add podcasts into the mix, like this one, and it changes the entire approach. If I stumble over my words in the middle of a sentence that I’ve done a couple of times already, on this very show, I will stop and repeat myself because it’s easier to just cut that out in post, and it sounds much more natural. They’ll take out ums, ahs, stutters, and the rest. Live, you have to respond to that very differently, but pre-recorded video has something of the same problem because, okay, the audio you can cut super easily.

With video, you have to sort of a smear, and it’s obvious when people know what they’re looking at. And, “Wait, what was that? That was odd. They blew a take.” You can cheat, which is what I tend to do, and oh, I wind up doing a bunch of slides in some of my talks because every slide transition is an excuse to cut because suddenly for a split second I’m not on the camera and we can do all kinds of fun things.

But it’s all these little things, and part of the problem, too, with the pandemic was, we suddenly had to learn how to be A/V folks when previously we had the good fortune slash good sense to work with people who are specialist experts in this space. Now it’s, “Well, I guess I am the best boy grip today,” whate—I’m learning what that means [laugh] as we—

Mark: That’s right.

Corey: —continue onward. Ugh. I never signed up for this, but it’s the thing that happens to you instead of what you plan on. I think that’s called life.

Mark: Feels right. Feels right, yeah. It’s just one of those things. And I’m looking forward to the time after this, when we do get back to in-person talks, and we do get to do some things. So, I have a lot of hot takes around speaking. So, I came up in Toastmasters. Are you familiar
with Toastmasters at all?

Corey: I very much am.

Mark: Oh, yeah. Okay, so I came up in Toastmasters, and for people at home who don’t know, it’s kind of like a meetup where you go and you
actually practice public speaking, based on these props, et cetera. For me, I learned to do things like not say ‘um’ and ‘ah’ on stage because there’s someone in the room counting every time you do it, and then when you get that review at the end when they give you your feedback, they’ll call that out. Or when you say ‘like you know,’ or too many ‘and so’, all these little—I think the word is disfluencies that you use that people say make you sound more natural, those are things that were coached out with me for public speaking. I just don’t do those things anymore, and I feel like there are ways for you not to do it.

And I tweeted that before, that you shouldn’t say ‘um’ and ‘ah’ and have someone tell me, “Oh, no, they're a natural part of language.” And then, “It’s not natural and it could freak people out.” And I was like, “Okay. I mean, you have your opinion about that.” Like, that’s fine, but it’s just a hot take that I had about speaking.

I think that you should do lots of things when you speak. The rate that you walk back and forth, or should you be static? How much should be on your slides? People put a lot of stuff on slides, I’m like, “I don’t want to read your slides. I’d rather listen to you use your slides.” I mean, I can go on and on. We should have another podcast called, “Hey, Mark talks about public speaking,” because that is one of my jams. That and supporting people who come from different paths. Those two things, I can go on for hours about.

Corey: And they’re aligned in a lot of respects. I agree with you on the public speaking. Focusing on the things that make you a better speaker are not that hard in most cases, but it’s being aware of what you’re doing. I thought I was a pretty good speaker when I had a coach for a little while, and she would stand there, “Give just the first minute of your talk.” And she’s there and writing down notes; I get a minute in and it’s like, “Okay, I can’t wait to see what she doesn’t like once I get started.” She’s like, “Nope. I have plenty. That will cover us for the next six weeks.” Like, “O…kay? I guess she doesn’t know what she’s doing.”

Spoiler she did, in fact, know what she was doing and was very good at it and my talks are better for it as a result. But it comes down to practicing. I didn’t have a thing like Toastmasters when I was learning to speak to other folks. I just did it by getting it wrong a lot of times. I would speak to small groups repeatedly, and I’d get better at it in time.

And I would put time-bound on it because people would sit there and listen to me talk and then the elevator would arrive at our floor and they could escape and okay, they don’t listen to me publicly speaking anymore, but you find time to practice in front of other folks. I am kidding, to be clear. Don’t harass strangers with public speaking talks. That was in fact a joke. I know there’s at least one person in the audience who’s going to hear that and take notes and think, “Ah, I’m going to do that because he said it’s a good idea.” This is the challenge with being a quote-unquote, “Role model” sometimes. My role model approach is to give people guidance by providing a horrible warning of what not to do.

Mark: [laugh].

Corey: You’ve gone the other direction and that’s kind of awesome. So, one of the recurring themes of this show has been, where does the next generation come from? Where do we find the next generation of engineer, of person working in cloud in various ways? Because the paths that a lot of us walked who’ve been in this space for a decade or more have been closed. And standing here, it sounds an awful lot like, “Oh, go in and apply for jobs with a firm handshake and a printed copy of your resume and ask to see the manager and you’ll have a job before dark.”

Yeah, what worked for us doesn’t work for people entering the workforce today, and there have to be different paths. Bootcamps are often the subject of, I think, a deserved level of scrutiny because quality differs wildly, and from the outside if you don’t know the space, a well-respected bootcamp that knows exactly what it’s doing and has established long-term relationships with a number of admirable hiring entities in the space and grifter who threw together a website look identical. It’s a hard problem to solve. How do you view teaching the next generation and getting them into this space, assuming that that isn’t something that is morally reprehensible? And some days, I wonder if exposing this industry to folks who are new to it isn’t a problem.

Mark: No, good question. So, I think in general—so I am pro bootcamp. I am pro self-taught. I was not always. And that’s because of personal insecurity. Let’s dive into that a little bit.

So, I’ve been writing code since I was probably around 14 because I was lucky enough to go to a high school to had a computer science program on the south side of Chicago, one school. And then when I say I was lucky, I was really lucky because the school that I went to wasn’t a high resource school; I didn’t go to a private school. I went to a public school that just happened that one of the professors from IIT, also worked on staff a few days a week at my school, and we could take programming classes with this guy. Total luck. And so I get into computer science that way, take AP Computer Science in high school—which is, like, the pre-college level—then I go into undergrad, then I go into grad school for computer science.

So, like, as traditional of a path that you can get. So, in my mind, it was all about my sweat equity that I had put in that disqualified everybody else. So, Corey, if you come from a bootcamp, you haven’t spent the time that I spent learning to code; you haven’t sweat, you haven’t had to bleed, you haven’t tried to write a two’s complement algorithm on top of your other five classes for that semester. You haven’t done it, definitely you don’t deserve to be here. So, that was so much of my attitude, until—until—I got the opportunity to have my mind completely blown when I got asked to teach.

Because when I got to asked to teach, I thought, “Yeah, I’m going to have my way of going in there and I’m going to show them how to do it right. This is my chance to correct these coding bootcampers and show them how it goes.” And then I find these people who were born for this life. So, some of us are natural talents, some of us are people who can just acquire the talent later. And both are totally valid.

But I met this one student. She was a math teacher for years in Chicago Public Schools. She’s like, “I want a career change.” Comes to the program that I taught at Northwestern, does so freaking well that she ends up getting a job at Airbnb. Now, if you have to make her go back four years at university, is that window still open for her? Maybe not.

Then I meet this other woman, she was a paralegal for ten years. Ten years as a paralegal was the best engineer in the program when I taught, she was the best developer we had. Before the bootcamp was over, she had already gotten the job offer. She was meant for this. You see what I’m saying?

So, that’s why I’m so excited because it’s like, I have all these stories of people who are meant for this. I taught, and I met people that changed the way I even saw the rest of the world. I had some non-binary trans students; I didn’t even know what pronouns were. I had no idea that people didn’t go by he/him, she/her. And then I had to learn about they and them and still teach you code without misgendering you at the same time, right because you’re in a classroom and you’re rapid-fire, all right, you—you know, how about this person? How about that person? And so you have to like, it’s hard to take—

Corey: Yeah, I can understand async, await, and JavaScript, but somehow understanding that not everyone has the pronouns that you are accustomed to using for people who look certain ways is a bridge too far for you to wrap your head around. Right. We can always improve, we can always change. It’s just—at least when I screw up async, await, I don’t make people feel less than. I just make—

Mark: Totally.

Corey: —users feel that, “Wow, this guy has no idea how to code.” You’re right, I don’t.

Mark: Yeah, so as I’m on my soapbox, I’ll just say this. I think coding bootcamps and self-taught programs where you can go online, I think this is where the door is the widest open for people to enter the industry because there is no requirement of a degree behind this. I just think that has just really opened the door for a lot of people to do things that is life-changing. So, when you meet somebody who’s only making—because we’re all engineers and we do all this stuff, we make a lot of money. And we’re all comfortable. When you meet somebody where they go from 40,000 to 80,000, that is not the same story for—as it is for us.

Corey: Exactly. And there’s an entire school of thought out there that, “Oh, you should do this for the love because it is who you are, it is who you were meant to be.” And for some people, that’s right, and I celebrate and cherish those folks. And there are other folks for whom, “I got into tech because of the money.” And you know what?

I celebrate and cherish those folks because that is not inherently wrong. It says nothing negative about you whatsoever to want to improve your quality of life and wanting to support your family in varying ways. I have zero shade to throw at either one of those people. And when it comes to which of those two people do I want to hire, I have no preference in either direction because both are valid and both have directions that they can think in that the other one may not necessarily see for a variety of reasons. It’s fine.

Mark: I wanted to be an engineering manager. You know why? Not because I loved leadership; because I wanted more money.

Corey: Yes.

Mark: So, I’ve been in the industry for quite a long time. I’m a little bit on the older side of the story, right? I’m a little bit older. You know, for me, before we got ‘staff’ and ‘principal’ and all this kind of stuff, it was senior software engineer and then you topped out in terms of your earning potential. But if you wanted more, you became a manager, director, et cetera.

So, that’s why I wanted to be a manager for a while; I wanted more money, so why is my choice to be a manager more valuable than those people who want to make more money by coming into engineering or software development? I don’t think it is.

Corey: So, we’ve talked about positivity, we’ve talked about dealing with unpleasant people, we’ve talked about technology, and then, of course, we’ve talked about getting up on soapboxes. Let’s tie all of that together for one last topic. What is your position on open-source in cloud?

Mark: I think open-source software allows us to do a lot of incredible things. And I know that’s a very light, fluffy, politically correct answer, but it is true, right? So, we get to take advantage of the brains of so many different people, all the ideas and contributions of so many different people so that we can do incredible things. And I think cloud really makes the world more accessible in general because—so when I used to do websites, I had to have a physical server that I would have to, like, try to talk to my ISP to be able to host things. And so, there was a lot of barriers to entry to do things that way.

Now, with cloud and open-source, I could literally pick up a tool and deploy some software to the cloud. And the tool could you open-source so I can actually see what’s happening and I could pick up other tools to help build out my vision for whatever I’m creating. So, I think open-source just gives a lot of opportunity.

Corey: Oh, my stars, yes. It’s even far more so than when I entered the field, and even back then there were challenges. One of the most democratizing aspects of cloud is that you can work with the same technologies that giant companies are using. When I entered the workforce, it’s, “Wow, you’re really good with Apache, but it seems like you don’t really know a whole lot about the world of enterprise storage. What’s going on with that?”

And the honest answer was, “Well, it turns out that on my laptop, I can compile Apache super easily, but I’m finding it hard, given that I’m new to the workforce, to afford a $300,000 SAN in my garage, so maybe we can wind up figuring out that there are other ways to do it.” That doesn’t happen today. Now, you can spin something up in the cloud, use it for a little bit. You’re done, turn it off, and then never again have to worry about it except over in AWS land where you get charged 22 cents a month in perpetuity for some godforsaken reason you can’t be bothered to track down and certainly no one can understand because, you know, cloud billing.

Mark: [laugh].

Corey: But if that’s the tax versus the SAN tax, I’ll take it.

Mark: So, what I think is really interesting what cloud does, I like the word democratization because I think about going back to—just as a lateral reference to the bootcamp thing—I couldn’t get my parents to see my software when I was in college when I made stuff because it was on my laptop. But when I was teaching these bootcamp students, they all deployed to Heroku. So, in their first couple of months, the cloud was allowing them to do something super cool that was not possible in the early days when I was coming up, learning how to code. And so they could deploy to Heroku, they could use GitHub Pages, you know like, open-source still coming into play. They can use all these tools and it’s available to them, and I still think to me that is mind-blowing that I would have to bring my physical laptop or desktop home and say, “Mom, look at this terminal window that’s doing this algorithm that I just did,” versus what these new people can do with the cloud. It’s like, “Oh, yeah, I want to build a website. I want to publish it today. Publish right now.” Like, during our conversation, we both could have probably spent up a Hello World in the cloud with very little.

Corey: Well, you could have. I could have done it in some horrifying way by using my favorite database: DNS. But that’s a separate problem.

Mark: [laugh]. Yeah, but I go to Firebase deploy and create a quick app real quick; Firebase deploy. Boom, I’m in the cloud. And I just think that the power behind that is just outstanding.

Corey: If I had to pick a single cloud provider for someone new to the field to work with, it would be Google Cloud, and it’s not particularly close. Just because the developer experience for someone who has not spent ten years marinating in cloud is worlds apart from what you’re going to see in almost every other provider. I take it back, it is close. Neck-and-neck in different ways is also DigitalOcean, just because it explains things; their documentation is amazing and it lets people get started. My challenge with DigitalOcean is that it’s not thought of, commonly, as a tier-one cloud provider in a lot of different directions, so the utility of learning how that platform works for someone who’s planning to be in the industry for a while might potentially not get them as far.

But again, there’s no wrong answer. Whatever interests you, whenever you have to work on, do it. The obvious question of, “What technology should I learn,” it’s, “Well, the ones that the companies you know are working with,” [laugh] so you can, ideally, turn it into something that throws off money, rather than doing it in your spare time for the love of it and not reaping any rewards from it.

Mark: Yeah. If people ask me what should they use it to build something? And I think about what they want to do. And I also will say, “What will get you to ship the fastest? How can you ship?”

Because that’s what’s really important for most people because people don’t finish things. You know, as an engineer, how many side projects you probably have in the closet that never saw the light of day because you never shipped. I always say to people, “Well, what’s going to get you to ship?” If it’s View, use View and pair that with DigitalOcean, if that’s going to get you to ship, right? Or use Angular plus Google Cloud Platform if that’s going to get you to ship.

Use what’s going to get you to ship because—if it’s just your project you’re trying to run on. Now, if it’s a company asking me, that’s a consulting question which is a different answer. We do a much more in-detail analysis.

Corey: I want to thank you so much for taking the time to speak with me about, honestly, a very wide-ranging group of topics. If people want to learn more about who you are, how you think, what you’re up to, where can they find you?

Mark: You can always find me spreading the love, being positive, hanging out. Look, if you want to feel better about yourself, come find me on Twitter at @marktechson—M-A-R-K-T-E-C-H-S-O-N. I’m out there waiting for you, so just come on and have a good time.

Corey: And we will, of course, throw links to that in the [show notes 00:36:45]. Thank you so much for your time today.

Mark: Oh, it’s been a pleasure. Thanks for having me.

Corey: Mark Thompson, developer relations engineer at Google. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry, deranged comment that you spent several weeks rehearsing in the elevator.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Chris

Chris Williams is a Enterprise Architect for World Wide Technology — a technology solution and service provider. There he helps customers design the next generation of public, private, and hybrid cloud solutions, specializing in AWS and VMware. His first computer was a Commodore 64, and he’s been playing video games ever since.

Chris blogs about virtualization, technology, and design at Mistwire. He is an active community leader, co-organizing the AWS Portsmouth User Group, and both hosts and presents on vBrownBag. He is also an active mentor, helping students at the University of New Hampshire through Diversify Thinking—an initiative focused on empowering girls and women to pursue education and careers in STEM.

Chris is a certified AWS Hero as well as a VMware vExpert.

Fun fact that Chris doesn’t want you to know: he has a degree in psychology so you can totally talk to him about your feelings.

Links:

  • WWT: https://www.wwt.com/
  • Twitter: https://twitter.com/mistwire
  • Personal site: https://mistwire.com
  • vBrownBag: https://vbrownbag.com/team/chris-williams/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Honeycomb. When production is running slow, it's hard to know where problems originate: is it your application code, users, or the underlying systems? I’ve got five bucks on DNS, personally. Why scroll through endless dashboards, while dealing with alert floods, going from tool to tool to tool that you employ, guessing at which puzzle pieces matter? Context switching and tool sprawl are slowly killing both your team and your business. You should care more about one of those than the other, which one is up to you. Drop the separate pillars and enter a world of getting one unified understanding of the one thing driving your business: production. With Honeycomb, you guess less and know more. Try it for free at Honeycomb.io/screaminginthecloud. Observability, it’s more than just hipster monitoring.

Corey: This episode is sponsored in part by our friends at Vultr. Spelled V-U-L-T-R because they’re all about helping save money, including on things like, you know, vowels. So, what they do is they are a cloud provider that provides surprisingly high performance cloud compute at a price that—while sure they claim its better than AWS pricing—and when they say that they mean it is less money. Sure, I don’t dispute that but what I find interesting is that it’s predictable. They tell you in advance on a monthly basis what it’s going to going to cost. They have a bunch of advanced networking features. They have nineteen global locations and scale things elastically. Not to be confused with openly, because apparently elastic and open can mean the same thing sometimes. They have had over a million users. Deployments take less that sixty seconds across twelve pre-selected operating systems. Or, if you’re one of those nutters like me, you can bring your own ISO and install basically any operating system you want. Starting with pricing as low as $2.50 a month for Vultr cloud compute they have plans for developers and businesses of all sizes, except maybe Amazon, who stubbornly insists on having something to scale all on their own. Try Vultr today for free by visiting: vultr.com/screaming, and you’ll receive a $100 in credit. Thats v-u-l-t-r.com slash screaming.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. One of the things I miss the most from the pre-pandemic times is meeting people at conferences or at various business meetings, not because I like people—far from it—but because we go through a ritual that I am a huge fan of, which is the exchange of business cards. Now, it’s not because I’m a collector or anything here, but because I like seeing what people’s actual titles are instead of diving into the morass of what we call ourselves on Twitter and whatnot. Today, I have just one of those folks with me. My guest is Chris Williams, who works at WWT, and his business card title is Enterprise Architect, comma AWS Cloud. Chris, welcome.

Chris: Hi. Thanks for having me on the show, Corey.

Corey: No, thank you for taking the time to speak with me. I have to imagine that the next line in your business card is, “No, I don’t work for AWS,” because you know a company has succeeded when they get their name into people’s job titles who don’t work there.

Chris: So, I have a running joke where the next line should actually be cloud therapist. And my degree is actually in psychology, so I was striving to get cloud therapist in there, but they still don’t want to let me have it.

Corey: Former guest Bobby Allen is now a cloud therapist over at Google Cloud, which is just phenomenal. I don’t know what they’re doing in a marketing context over there; I just know that they’re just blasting them out of the park on a consistent, ongoing basis. It’s really nice to see. It’s forcing me to up my game a little bit. So, one of the challenges I’ve always had is, I don’t like putting other companies’ names into the title.

Now, I run the Last Week in AWS newsletter, so yeah, okay, great, there’s a little bit of ‘do as I say, not as I do’ going on here. Because it feels, on some level, like doing unpaid volunteer work for a $2 trillion company. Speaking of, you are an AWS Community Hero, where you do volunteer work for a $2 trillion company. How’d that come about? What did you do that made you rise to their notice?

Chris: That was a brilliant segue. Um—[laugh]—

Corey: I do my best.

Chris: So I, actually prior to becoming an AWS Community Hero, I do a lot of community work. So, I have run and helped to run four different community-led organizations: the Virtualization Technology User Group of New England; the AWS Portsmouth User Group, now the AWS Boston User Group; I’m a co-host and presenter for vBrownBag; I also do the New England AWS Community Day, which is a conglomeration of all the different user groups in one setting; and various and sundry other things, as well, along the way. Having done all of that, and having had a lot of the SAs and team members come and do speaking presentations for these various and sundry things, I was nominated internally by AWS to become one of their Community Heroes. Like you said, it’s basically unpaid volunteer work where I go out and tout the services. I love talking about nerd stuff, so when I started working on AWS technologies, I really enjoyed it, and I just, kind of like, glommed on with other people that did it as well. I’m also a VMware vExpert, which basically use the exact same accolade for VMware. I have not been doing as much VMware stuff in the recent past, but that’s kind of how I got into this gig.

Corey: One of the things that strikes me as being the right move with respect to these, effectively, community voice accolades is Microsoft got something very right—they’ve been doing this a long time—they have their MVP program, but they have to re-invite people who have to requalify for it by whatever criteria they are, every year. AWS does not do this with their Heroes program. If you look at their Heroes page, there’s a number of folks up there who have been doing interesting things in the cloud years ago, but then fell off the radar for a variety of reasons. In fact, the only way that I’m aware that you can lose Hero status is via getting a job at AWS or one of AWS competitors.

Now, the hard part, of course, is well, who is Amazon’s competitors? Basically everyone, but it mostly distills down to Microsoft, Google, and Oracle, as best I can tell, for Hero status. How does VMware fall on that spectrum? To be more specific, how does VMware fall on the spectrum of their community engagement program and having to renew, not, “Are they AWS’s competitor?” To which the answer is, “Of course.”

Chris: So, the renewal process for the VMware vExpert program is an annual re-up process where you fill out the form, list your contribution of the year, what you’ve done over the previous year, and then put it in for submission to the board of VMware vExperts who then give you the thumbs up or thumbs down. Much like Nero, you know, pass or fail, live or die. And I’ve been fortunate enough, so my vBrownBag contributions are every week; we have a show that happens every week. It can be either VMware stuff, or cloud in general stuff, or developer-related stuff. We cover the gamut; you know, people that want to come on and talk about whatever they want to talk about, they come on. And by virtue of that, we’ve had a lot of VMware speakers, we’ve had a lot of AWS speakers, we’ve had a lot of Azure speakers. So, I’ve been fortunate enough to be able to qualify each year with those contributions.

Corey: I think that’s the right way to go, from my perspective at least. But I want to get into this a little bit because you are an enterprise architect, which is always one of those terms that is super easy to make fun of in a variety of different ways. Your IDE is probably a whiteboard, and at some point when you have to write code, I thought you had a team of people who would be able to do that all for you because your job is to cogitate, and your artifacts are documentation, and the entire value of what you do can only be measured in the grand sweep of time, et cetera, et cetera, et cetera.

Chris: [laugh].

Corey: But you don’t generally get to be a Community Hero for stuff like that, and you don’t usually get to be a vExpert on the VMware side, by not having at least technical chops that make people take a second look. What is it you’d say it is you do hear for, lack of a better term?

Chris: “What would you say ya, do you here, Bob?” So, I’m not being facetious when I say cloud therapist. There is a lot of working at the eighth layer of the OSI model, the political layer. There’s a lot of taking the requirements from the customer and sending them to the engineer. I’m a people person.

The easy answer is to say, I do all the things from the TOGAF certification manual: the requirements, risks, assumptions, and constraints; the logical, conceptual, and physical diagrams; the harder answer is the soft skill side of that, is actually being able to communicate with the various levels of the industry, figuring out what the business really wants to do and how to technically solution that and figure out how to talk to the engineers to make that happen. You’re right EAs get made fun of all the time, almost as much as consultants get made fun of. And it’s a very squishy layer that, you know, depending upon your personality and the personality of the customer that you’re dealing with, it can work wonderfully well or it can crash and burn immediately. I know from personal experience that I don’t mesh well with financials, but I’m really, really good with, like, medical industry stuff, just the way that the brain works. But ironically, right now I’m working with a financial and we’re getting along like a house on fire.

Corey: Oh, yeah. I’ve been saying for a while now that when it comes to cloud, cost and architecture are the same things, and I think that ties back to a lot of different areas. But I want to be very clear here that we talk about, I’m not super deep into the financials, that does not mean you’re bad at architecture because working on finance means different things to different folks. I don’t think that it is possibly a good architect in the cloud environment and not have a conception of, “Huh, that thing seems really expensive if I do it that way.” That is very different than having the skill of reading a profit and loss statement or understanding various implications of the time value of money calculation that a company uses, or how things get amortized.

There are nuances piled on top of nuances in finance, and it’s easy to sit here and think that oh, I’m not great at finance means I don’t know how money works. That is very rarely true. If you really don’t know how money works, you’ll go start a cryptocurrency startup.

Chris: [laugh]. So, I plugged back to you; I was listening to one of your old shows and I cribbed one of your ideas and totally went with it. So, I just said that there’s the logical, conceptual, and physical diagrams of an environment; on one of your shows, you had mentioned a financial diagram for an environment, and I was like, “That’s brilliant.” So, now when I go into a customer, I actually do that, too. I take my physical diagram, I strip out all of the IP addresses, and our names, and everything like that, and I plot down how much it’s going to cost, like, “This is the value of the EC2 instance,” or, “This is how much this pipe is going to cost if you run this over it.” And they go bananas over it. So, thanks
for providing that idea that I mercilessly stole.

Corey: Kind of fun on a lot of levels. Part of the challenge is as things get cloudier and it moves away from EC2 instances, ideally the lie we
would like to tell ourselves that everything’s in an auto-scaling group. Great—

Chris: Right.

Corey: —stepping beyond that when you start getting into something that’s even more intricately tied to a specific user, we’re talking about effectively trying to get unit economic measures of every user, every thousand users is going to cost me X dollars to service them on average, on top of a baseline of steady-state spend that is going to increase differently. At that point, talking to finance about predictive models turn into, “Well, this comes down to a question of business modeling.” But conversely, for engineering minds that is exactly what finance is used to figuring out. The problem they have is, “Well, every time we hire a new engineer, we wind up seeing our AWS bill increase.” Funny how that works. Yeah, how do you map that to something that the business understands? That is part of what they do. But it does, I admit, make it much more challenging from a financial map of an environment.

Chris: Yeah, especially when the customer or the company is—you know, they’ve been around for a while, and they’re used to just like that large bolus of money at the very beginning of a data center, and they buy the switches, and they buy the servers, and they virtualize them, and they have that set cost that they knew that they had to plunk down at the beginning. And it’s a mindset shift. And they’re coming around to it, some faster than others. Oddly enough, the startups nowadays are catching on very quickly. I don’t deal with a lot of startups, so it takes some finesse.

Corey: An interesting inflection that I’ve seen is that there’s an awful lot of enterprises out there that say, “Oh, we’re like a startup.” Great. You mean with weird cultural inflections that often distill down to cult of personality, the constant worry about whether you’re going to wind up running out of runway before finding product-market fit? And the rooms filled with—

Chris: The eighty-hour work weeks? The—[laugh]—

Corey: And they’re like, “No, no, no, it’s like the good parts.” “Oh, so you mean out the upside.” But you don’t hear it the other way around where you have a startup that you’re interviewing with, “Ha-ha, we’re like an enterprise. We have a six-month interview process that takes 18 different stages,” and so on and so forth. However, we do see startups having to mature rapidly, and move up the compliance path as they’re dealing with regulated entities and the rest, and wanting to deal with serious customers who have no sense of humor about, “Yeah, we’ll figure that part out later as part of an audit document.”

So, what we also see, though, is that enterprises are doing things that look a lot more startup-y. If I take a look at the common development environments and tools and techniques that big enterprises use, it looks an awful lot like how startups were doing it five or ten years ago. That is the slow and steady evolution of time. And what startups are doing today becomes enterprise tomorrow, and I can’t shake the feeling that there’s a sea of vendors out there who, in the event that winds up happening are eventually going to find themselves without a market at all. My model has been that if I go and found a Twitter for Pets style startup tomorrow and in ten years, it has grown to become an S&P 500 component—which is still easier to take seriously than most of what Tesla says—great.

During that journey, at what point do I become a given company’s customer because if there is no onboarding story for me to become your customer, you’re in a long-tail decline phase. That’s been my philosophy, but you are a—trademarked term—Enterprise Architect, so please feel free to tell me if I’m missing any of the nuances there, which I’m sure I am because let’s face it, nuance is hard; sweeping statements are easy.

Chris: As an architect, [laugh] it would be a disservice to not say my favorite catchphrase, it depends. There are so many dependencies to those kinds of sweeping statements. I mean, there’s a lot of enterprises that have good process; there are a lot of enterprises that have bad process. And going back to your previous statement of the startup inside the enterprise, I’m hearing a lot of companies nowadays saying, “Oh, well, we’ve now got this brand new incubator system that we’re currently running our little startup inside of. It’s got the best of both worlds.”

And I’m not going to go through the litany of bad things that you just said about startups, but they’ll try to encapsulate that shift that you’re talking about where the cheese is moving so quickly now that it’s very hard for these companies to know the customer well enough to continue to stay salient and continue to be able to look into that crystal ball to stay relevant in the future. My job as an EA is to try to capture that point in time where what are the requirements today and what are the known detriments that you’re going to see in your future that you need to protect against? So, that’s kind of my job—other than being a cloud therapist—in a nutshell.

Corey: I love the approach. My line has been that I do a lot of marriage counseling between engineering and finance, which is a fun term that also just so happens to be completely accurate.

Chris: Absolutely. [laugh]. I’m currently being a marriage counselor right now.

Corey: It’s an interesting time. So, you had a viral tweet recently that honestly, I’m a bit jealous about. I have had a lot of tweets that have done reasonably well, but I haven’t ever had anything go super-viral, where it was just a screenshot of a conversation you had with an AWS recruiter. Now, before we go into this, I want to make a couple of disclaimers here. Before I entered tech myself, I was a technical recruiter, and I can say that these people have hard jobs.

There is a constant pressure to perform, it is a sales job that is unlike most others. If you sell someone a pen, great, you can wrap your head around what that’s like. But you don’t have to worry about the pen deciding it doesn’t want to go home with the buyer. So, it becomes a double sale in a lot of weird ways, and there’s a constant race to the bottom and there’s a lot of competition in the space. It’s a numbers game and a lot of folks get in and wash out who have terrible behaviors and terrible patterns, so the whole industry gets tainted—in some respects—like that. A great example of someone who historically has been a terrific example of recruiting done right has been Jill Wohlner. And she’s one of the shining beacons of the industry as far as how to do these things in the right way—

Chris: Yes.

Corey: —but the fact that she is as exceptional as she is is in no small part because there’s a lot of random folks coming by. All which is to say that our conversation going forward is not and should not be aimed at smacking around individual recruiters or recruiting as a whole because that is unfair. Now, that disclaimer has been given. Great, what happened?

Chris: So, first off, shout out to Jill; she actually used to be a host on vBrownBag. So, hey girl. [laugh]. What happened was—and I have the utmost empathy and sympathy for recruiting; I actually used to have a side gig where I would go around to the local recruiting places around my area here and teach them how to read a cloud resume and how to read a req and try to separate the wheat from the chaff, and to actually have good conversations. This was back when cloud wasn’t—this was, like, three or four years ago.

And I would go in there and say, “This is how you recruit a cloud person nowadays.” So, I love good recruiters. This one was a weird experience in that—so when a recruiter reaches out to me, what I do is I take an assessment of my current situation: “Am I happy where I’m at right now?” The answer is, “Yes.” And if they ping me, I’ll say, “Hey, I’m happy right now, but if you have something that is, you know, a million dollars an hour, taste-testing margaritas on St. John island in the sand, I’m all ears. I’m listening. Conversely, I also am a Community Hero, so I know a ton of people out in the industry. Maybe I can help you out with landing that next person.”

Corey: I just want to say for the record, that is absolutely the right answer. And something like that is exactly what I would give, historically. I can’t do it now because let’s be clear here. I have a number of employees and, “Hey, Corey’s out there doing job interviews,” sends a message that isn’t good when it comes to how is that company doing anyway. I miss it because I enjoyed the process and I enjoyed the fun, but even when I was perfectly happy, it’s, “Well, I’m not actively on the market, but I am interested to have a conversation if you’ve got something
interesting.”

Because let’s face it, I want to hear what’s going on in the market, and if I’m starting to hear a lot of questions about a technology I have been dismissive of, okay, maybe it’s time to pay more attention. I have repeatedly been able to hire the people interviewing me in some cases, and sometimes I’ve gone on interviews just to keep my interview skills sharp and then wound up accepting the job because it turned out they did have something interesting that was compelling to me even though I was reasonably happy at the time. I will always take the meeting; I will always at least have a chat about what they’re doing, and I think that doing otherwise is doing yourself a disservice in the long arc of your career.

Chris: Right. And that’s basically the approach that I take, too. I want to hear what’s out there. I am very happy at World Wide right now, so I’m not interested, interested. But again, if they come up with an amazing opportunity, things could happen. So, I implied that in my response to him.

I said, “I’m happy right now, thanks for asking, but let’s set up the meeting and we can have a chat.” The response was unexpected. [laugh]. The response was basically, “If you’re not ready to leave right now, it makes no sense for me to talk to you.” And it was a funny… interaction.

I was like, “Huh. That’s funny.” I’m going to tweet about that because I thought it was funny—I’m not a jerk, so I’m going to block out all of the names and all of the identifying information and everything—and I threw it up. And the commiseration was so impressive. Not impressive in a good way; impressive in a bad way.

Every person that responded was like, “Yes. This has happened to me. Yes, this is”—and honestly, I got a lot of directors from AWS reaching out to me trying to figure out who that person was, apologizing saying that’s not our way. And I responded to each and every single one of them. And I was like, “Somebody has already found that person; somebody has already spoken to that person. That being said, look at all of the responses in the timeline. When you tell me personally, that’s not the way you do things, I believe that you believe that.”

Corey: Yeah, I believe you’re being sincere when you say this, however the reality of what the data shows and people’s lived experience in the form of anecdotes are worlds apart.

Chris: Yeah. And I’m an AWS Hero. [laugh]. That’s how I got treated. Not to blow my own horn or anything like that, but if that’s happening to me, either A, he didn’t look me up and just cold-called me—which is probably the case—and b, if he treats me like that, imagine how he’s treating everybody else?

Corey: This episode is sponsored in part by something new. Cloud Academy is a training platform built on two primary goals. Having the highest quality content in tech and cloud skills, and building a good community the is rich and full of IT and engineering professionals. You wouldn’t think those things go together, but sometimes they do. Its both useful for individuals and large enterprises, but here's what makes it new. I don’t use that term lightly. Cloud Academy invites you to showcase just how good your AWS skills are. For the next four weeks you’ll have a chance to prove yourself. Compete in four unique lab challenges, where they’ll be awarding more than $2000 in cash and prizes. I’m not kidding, first place is a thousand bucks. Pre-register for the first challenge now, one that I picked out myself on Amazon SNS image resizing, by visiting cloudacademy.com/corey. C-O-R-E-Y. That’s cloudacademy.com/corey. We’re gonna have some fun with this one!

Corey: Every once in a while I get some of their sourcers doing outreach to see folks who are somewhat aligned on them via LinkedIn or other things, and, “Oh, okay, yeah; if you look at the things I talked about in various places, I can understand how I might look like a potentially interesting hire.” And they send outreach emails to me, they’re always formulaic, and once in a while, I’ll tweet a screenshot of them where I redact the person’s name, and it was—and there’s a comment, like, “Should I tell them?” Because it’s fun; it’s hilarious. But I want to be clear because that often gets misconstrued; they have done absolutely nothing wrong. You’ve got to cast a wide net to find talent.

I’m surprised I get as few incidents of recruiter outreach as I do. I am not hireable and that’s okay, but I don’t begrudge people reaching out. I either respond with a, “No thanks,” if it’s a particularly good email, or I just hit the archive button and never think about it again. And that’s fine, too. But I don’t make people feel like a jerk for asking, and that is an engineering behavioral pattern that drives me up a wall.

It’s, “So, I’m thinking about a job here and I’m wondering if you might be a fit,” and your response is just to set them on fire? Well, guess what an awful lot of those people sending out those emails in the sourcing phase of recruiting are early career, and guess what, they tend to get promoted in the fullness of time. Sometimes they’re no longer recruiting at all; sometimes they wind up being hiring managers in different ways or trying to figure out what offer they’re going to extend to someone. And if you don’t think that people in those roles remember when they’re treated poorly as a response to their outreach, I have news for you. Don’t do it. Your reputation lingers long after you no longer work there.

Chris: Just exactly so. And I feel really bad for that guy.

Corey: I do hope that he was not reprimanded because he should not be. It is clearly a systemic problem, and the fact that one person happened to do this in a situation where it went viral does not mean that they are any worse than other folks doing it. It is a teachable opportunity. It is, “I know that you have incredible numbers of roles to hire for, all made all the more urgent by the fact that you’re having some significant numbers of departures—clearly—in the industry right now.” So, I get it; you have a hard job. I’m not going to waste your time because I don’t even respond to them just because, at AWS particularly, they have hard work to do, and just jawboning with me is not going to be useful for them.

Chris: [laugh].

Corey: I get it.

Chris: And you’re trying to hire the same talent too. So.

Corey: Exactly. One of the most egregious things I’ve seen in the course of my career was when that whole multiple accounts opened for Wells Fargo’s customers and they wound up firing 3500 people. Yeah, that’s not individual tellers doing something unethical. That is a systemic problem, and you clean house at the top because you’re not going to convince me that you’re hiring that many people who are unethical and setting out to do these things as a matter of course. It means that the incentives are wrong, it means that the way you’re measuring things are wrong, and people tend to do things out of fear or because there’s now a culture of it. And if you fire individuals for that, you’re wrong.

Chris: And that was the message that I conveyed to the people that reached out to me and spoke to me. I was like, there is a misaligned KPI, or OKR, or whatever acronym you want to use, that is forcing them to do this churn-and-burn mentality instead of active, compassionate recruiting. I don’t know what that term is; I’m very far removed from the recruiting world. But that person isn’t doing that because they’re a jerk. They’re doing that because they have numbers to hit and they’ve got to grind out as many as humanly possible. And you’re going to get bad employees when you do that. That’s not a long-term sustainable path. So, that was the conversation that I had with them. Hopefully, it resonated and hits home.

Corey: I still remember from ten years ago—and I don’t always tell the story, but I absolutely will now—I went up to San Francisco when I lived in Los Angeles; I interviewed with Yammer. I went through the entire process—this was not too long before they got acquired by Microsoft so that gives you some time basis—and I got a job offer. And it was a not ridiculous offer. I was going to think about it, and I [unintelligible 00:24:19], “Great. Thank you. Let me sleep on this for a day or two and I’ll get back to you definitely before the end of the week.”

Within an hour, I got a response rescinding the offer claiming it had been sent by mistake. Now, I believe that that is true and that they are being sincere with this. I don’t know that if it was the wrong person; I don’t know if that suddenly they didn’t have the req or they had another candidate that suddenly liked better that said no and then came back and said yes, but it’s been over a decade now and every time I talk to someone who’s considering something in that group, I tell this story. That’s the sort of thing that leaves a mark because I have a certain philosophy of I don’t ever resign from a job before I wind up making sure everything is solid—things are signed, good to go, the background check clears, et cetera—because I don’t want to find myself suddenly without income or employment, especially in that era. And that was fine, but a lot of people don’t do that.

As soon as the offer comes in, they’re like, “I’m going to go take a crap on my boss’s desk,” which, let’s be clear, I don’t recommend. You should write a polite and formulaic resignation letter and then you should email it to your boss, you should not carve it into their door. Do this in a responsible way, and remember that you’re going to encounter these people again throughout your career. But if I had done that, I would have had serious problems. And so that points to something systemically awful at a company.

I have never in my career as a hiring manager extended an offer and then rescinded it for anything other than we can’t come to an agreement on this. To be clear, this is also something I wonder about in the space, when people tell stories about how they get a job offer, they attempt to negotiate the offer, and then it gets withdrawn. There are two ways that goes. One is, “Well if you’re not happy with this offer, get out of here.” Yeah, that is a crappy company, but there’s also the story of people who don’t know how to negotiate effectively, and in turn, they come back with indications that you do not know how to write a business email, you do not know how negotiations work, and suddenly, you’re giving them a last-minute opportunity to get out before they hire someone who is going to be something of a wrecking ball in the company, and, “Whew, dodged a bullet on that.”

I haven’t encountered that scenario myself, but I’ve seen it from other folks and emails that have been passed around in various channels. So, my position on this is everyone should negotiate offers, but visit fearlesssalarynegotiation.com, it’s run by my friend, Josh; he has a whole bunch of free content on his site. Look at it. Read it. It is how to handle this stuff effectively and why things are the way that they are. Follow his advice, and you won’t go too far wrong. Again, I have no financial relationship, I just like what he’s done a lot and I’ve been talking to him for years.

Chris: Nice. I’ll definitely check that out. [laugh].

Corey: Another example is developher—that’s develop H-E-R dot com. Someone else I’ve been speaking to who’s great at this takes a different perspective on it, and that’s fine. There’s a lot of advice out there. Just make sure that whoever it is you’re talking to about this is in a position to know what they’re talking about because there’s crap advice that’s free. Yeah. How do you figure out the good advice and the bad advice? I’m worried someone out there is actually running Route 53 is a database for God’s sake.

Chris: That’s crazy talk. Who would do that? That’s madness.

Corey: I can’t imagine it.

Chris: We’re actually in the process of trying to figure out how to do a panel chat on exactly that, like, do a vBrownBag on salary negotiations, get some really good people in the room that can have a conversation around some of the tough questions that come around salary negotiation, what’s too much to ask for? What kind of attitude should you go into it with? What kind of process should you have mentally? Is it scrawling in crayon, “No. More money,” and then hitting send? Or is it something a little bit more advanced?

Corey: I also want to be clear that as you’re building panels and stuff like that—because I got this wrong early on in my public speaking career, to be clear—I built talks aligned with this based on what worked for me—make sure that there are folks on the panel who are not painfully over-represented as you and I are because what works for us and we’re considered oh, savvy business people who are great negotiators comes across as entitled, or demanding, or ooh, maybe we shouldn’t hire her—and yes, I’m talking about her in a lot of these scenarios—make sure you have a diverse group of folks who can share lived experience and strategies that work because what works for you and me is not universal, I promise.

Chris: So, the only requirement to set this panel is that you have to be a not-white guy; not-old-white guy. That’s literally the one rule. [laugh].

Corey: I like the approach. It’s a good way to do it. I don’t do manels.

Chris: Yes. And it’s tough because I’m not going to get into it, but the mental space that you have to be in to be a woman in tech, it’s a delicate balance because when I’m approaching somebody, I don’t want to slide into their DMs. It’s like this, “Hey, I know this other person and they recommended you and I am not a weirdo.” [laugh]. As an old white guy, I have to be very not a weirdo when I’m talking to folks that I’m desperate to get on the show.

Because I love having that diverse aspect, just different people from different backgrounds. Which is why we did the entire career series on vBrownBag. We did data science with Ayodele; we did how to get into cybersecurity with Christoph. It was a fantastic series of how to get into IT. This was at the beginning of the pandemic.

We wanted to do a series on, okay, there’s a lot of people out there that are furloughed right now. How do we get some people on the show that can talk to how to get into a part of IT that they’re passionate about? We did a triple series on how to get into game development with Dennis Diack, the founder of Apocalypse Studios. We had a bunch of the other AWS Heroes from serverless, and Lambda, and AI on the show to talk, and it was really fantastic and I think it resonated well with the community.

Corey: It takes work to have a group of guests on things like podcasts like this. You’ve been running vBrownBag for longer than I’ve been running this, and—

Chris: 13 years now.

Corey: Yeah. This is I think, coming up on what, four years-ish, maybe three, in that range? The passing of time, especially in a pandemic era, is challenging. And there’s always a difference. If I invite a white dude to come on the podcast, the answer is yes before I get the word podcast fully out of my mouth, whereas folks who are not over-represented, they’re a little more cautious. First, there’s the question of, “Am I a trash bag?” And the answer is, “No.” Well, no, not in the way that you’re concerned about other ways—

Chris: [laugh]. That you’re aware of. [laugh].

Corey: Oh, God, yes, but—yeah. And then—and that’s part of it, and then very often, there’s a second one of, “Well, I don’t think I have
anything, really, to talk about,” is often a common objection here. And it’s, yeah, if I’m inviting you on this show, I promise that’s not true. Don’t worry about that piece of it. And then it’s the standard stuff that just comes with being me, of, “Yeah, I’ve read your Twitter feed; you got to insult me here?” It’s, “No, no, not really the same tone. But great question; throw the”—it goes down to process. But it takes constant work, you can’t just put an open call out for guest nominations, and expect that to wind up being representative of our industry. It is representative of our biases, in many respects.

Chris: It’s a tough needle to thread. Because the show has been around for a long time, it’s easier for me now, because the show has been around for 13 years. We actually just recorded our two thousandth and sixtieth episode the other night. And even with that, getting that kind of outreach, [#techtwitter 00:31:32] is wonderful for making new recommendations of people. So, that’s been really fun. The rest of Twitter is a hot trash fire, but that’s beside the point. So yeah, I don’t have a good solution for it. There’s no easy answer for it other than to just be empathic, and communicative, and reach people on their level, and have a good show.

Corey: And sometimes that’s all it takes. The idea behind doing a podcast—despite my constant jokes—it’s not out of a love affair of the sound of my own voice. It’s about for better or worse, for reasons I don’t fully understand, I have a platform. People listen to the show and they care what people have to say. So, my question is, how can I wind up using that platform to tell stories that lift up narratives that are helpful for folks that they can use as inspiration—in my case, as critical warnings of what to avoid—and effectively showcasing some of the best our industry has to offer, in many respects.

So, if the guest has a good time and the audience can learn something, and I’m not accidentally perpetuating horrifying things, that’s really more than I have any right to ask from a show like this. The fact that it’s succeeded is due in no small part to not just an amazing audience, but also guests like you. So, thank you.

Chris: Oh no, Thank you. And it is. It’s… these kinds of shows are super fun. If it wasn’t fun, I wouldn’t have done it for as long as I have. I still enjoy chatting with folks and getting new voices.

I love that first-time presenter who was, like, super nervous and I spend 15 minutes with them ahead of the show, I say, “Okay, relax. It’s just going to be me and you facing each other. We’re going to have a good time. You’re going to talk about something that you love talking about, and we’re going to be nerds and do nerd stuff. This is me and you in front of a water cooler with a whiteboard just being geeks and talking about cool stuff. We’re also going to record it and some amount of people is going to see it afterwards.” [laugh].

And yeah, that’s the part that I love. And then watching somebody like that turn into the keynote speaker at a conference ten years down the road. And I get to say, “Oh, I knew that person when.”

Corey: I just want to be remembered by folks who look back fondly at some of the things that we talk about here. I don’t even need credit, just yeah. People who see that they’ve learned things and carry them forward and spread to others, there’s so many favors that people have done for us that we can only ever pay forward.

Chris: Yeah, exactly. So—and that’s actually how I got into vBrownBag. I came to them saying, “Hey, I love the things that you guys have done. I actually passed my VCIX because of watching vBrownBags. What can I do to help contribute back to the community?” And Alistair said, “Funny you should mention that.” [laugh]. And here we are seven years later.

Corey: Well, to that end, if people are inspired by what you’re saying and they want to hear more about what you have to say or, heaven forbid, follow in your footsteps, where can they find you?

Chris: So, you can find me on Twitter; I am at mistwire.com—M-I-S-T-W-I-R-E; if you Google ‘mistwire,’ I am the first three pages of hits; so I have a blog; you can find me on vBrownBag. I’m hard to miss on Twitter [laugh] I discourage you from following me there. But yeah, you can hit me up on all of the formats. And if you want to present, I’d love to get you on the show. If you want to learn more about what it takes to become an AWS Hero or if you want to get into that line of work, I highly discourage it. It’s a long slog but it’s a—yeah, I’d love to talk to you.

Corey: And we of course put links to that in the [show notes 00:35:01]. Thank you so much for taking the time to speak with me, Chris. I really appreciate it.

Chris: Thank you, Corey. Thanks for having me on.

Corey: Chris Williams, Enterprise Architect, comma AWS Cloud at WWT. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with a comment telling me that while you didn’t actively enjoy this episode, you are at least open to enjoying future episodes if I have one that might potentially be exciting.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Scott

Scott first typed ‘docker run’ in 2013 and hasn't looked back. He’s been with Docker since 2014 in a variety of leadership roles and currently serves as CEO. His experience previous to Docker includes Sun Microsystems, Puppet, Netscape, Cisco, and Loudcloud (parent of Opsware). When not fussing with computers he spends time with his three kids fussing with computers.

Links:

  • Docker: https://www.docker.com
  • Twitter: https://twitter.com/scottcjohnston

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Liquibase. If you’re anything like me, you’ve screwed up the database part of a deployment so severely that you’ve been banned from touching every anything that remotely sounds like SQL, at at least three different companies. We’ve mostly got code deployments solved for, but when it comes to databases we basically rely on desperate hope, with a roll back plan of keeping our resumes up to date. It doesn’t have to be that way. Meet Liquibase. It is both an open source project and a commercial offering. Liquibase lets you track, modify, and automate database schema changes across almost any database, with guardrails to ensure you’ll still have a company left after you deploy the change. No matter where your database lives, Liquibase can help you solve your database deployment issues. Check them out today at liquibase.com. Offer does not apply to Route 53.

Corey: This episode is sponsored in part by something new. Cloud Academy is a training platform built on two primary goals. Having the highest quality content in tech and cloud skills, and building a good community the is rich and full of IT and engineering professionals. You wouldn’t think those things go together, but sometimes they do. Its both useful for individuals and large enterprises, but here's what makes it new. I don’t use that term lightly. Cloud Academy invites you to showcase just how good your AWS skills are. For the next four weeks you’ll have a chance to prove yourself. Compete in four unique lab challenges, where they’ll be awarding more than $2000 in cash and prizes. I’m not kidding, first place is a thousand bucks. Pre-register for the first challenge now, one that I picked out myself on Amazon SNS image resizing, by visiting cloudacademy.com/corey. C-O-R-E-Y. That’s cloudacademy.com/corey. We’re gonna have some fun with this one!

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Once upon a time, I started my public speaking career as a traveling contract trainer for Puppet; I’ve talked about this before. And during that time, I encountered someone who worked there as an exec, Scott Johnston, who sat down, talked to me about how I viewed things, and then almost immediately went to go work at Docker instead. Today’s promoted episode brings Scott on to the show. Scott, you fled to get away from me, became the CEO of Docker over the past, oh what is it, seven years now. You’re still standing there, and I’m not making fun of Docker quite the way that I used to. First, thanks for joining me.

Scott: Great to be here, Corey. Thanks for the invitation. I’m not sure I was fleeing you, but we can recover that one at another time.

Corey: Oh, absolutely. In that era, one of my first talks that I started giving that anyone really paid any attention to was called, “Heresy in the Church of Docker,” where I listed about 10 to 13 different things that Docker didn’t seem to have answers for, like network separation, security, audit logging, et cetera, et cetera. And it was a fun talk that I used to basically learn how to speak publicly without crying before and after the talk. And in time, it wound up aging out as these problems got addressed, but what surprised me at the time was how receptive the Docker community was to the idea of a talk that wound up effectively criticizing something that for, well, a number of them it felt a lot of the time like it wasn’t that far from a religion; it was very hype-driven: “Docker, Docker, Docker” was a recurring joke. Docker has changed a lot. The burning question that I think I want to start this off with is that it’s 2021; what is Docker? Is it a technology? Is it a company? Is it a religion? Is it a community? What is Docker?

Scott: Yes. I mean that sincerely. Often, the first awareness or the first introduction that newcomers have is in fact the community, before they get their hands on the product, before they learn that there’s a company behind the product is they have a colleague who is, either through a Zoom or sitting next to them in some places, or in a coffee shop, and says, “Hey, you got to try this thing called Docker.” And they lean over—either virtually or physically—and look at the laptop of their friend who’s promoting Docker, and they see a magical experience. And that is the introduction of so many of our community members, having spoken with them and heard their own kind of journeys.

And so that leads to like, “Okay, so why the excitement? Why did the friend lean over to the other friend and introduce?” It’s because the tools that Docker provides just helps devs get their app built and shipping faster, more securely, with choice, without being tied into any particular runtime, any particular infrastructure. And that combination has proven to be a breakthrough dopamine hit to developers since the very beginning, since 2013, when Docker is open-source.

Corey: It feels like originally, the breakthrough of Docker, that people will say, “Oh, containers aren’t new. We’ve had that going back to LPARs on mainframes.” Yes, I’m aware, but suddenly, it became easy to work with and didn’t take tremendous effort to get unified environments. It was cynically observed at the time by lots of folks smarter than I am, that the big breakthrough Docker had was how to make my MacBook look a lot more like a Linux server in production. And we talk about breaking down silos between ops and dev, but in many ways, this just meant that the silo became increasingly irrelevant because, “Works on my machine” was no longer a problem.

“Well, you better back up your email because your laptop’s about to go into production in that case.” Containers made it easier and that was a big deal. It seems, on some level, like there was a foray where Docker the company was moving into the world of, “Okay, now we’re going to run a lot of these containers in production for you, et cetera.” It really feels like recently, the company as a whole and the strategy has turned towards getting back to its roots of solving developer problems and positioning itself as a developer tool. Is that a fair characterization?

Scott: A hundred percent. That’s very intentional, as well. We certainly had good products, and great customers, and we’re solving problems for customers on the ops side, I’ll call it, but when we stood back—this is around 2019—and said, “Where’s the real… joy?” For lack of a better word, “Where’s the real joy from a community standpoint, from a product experience standpoint, from a what do we do different and better and more capable than anyone else in the ecosystem?” It was that developer experience. And so the reset that you’re referring to in November 2019, was to give us the freedom to go back and just focus the entire company’s efforts on the needs of developers without any other distractions from a revenue, customer, channel, so on and so forth.

Corey: So, we knew this was going to come up in the conversation, but as of a couple of weeks ago—as of the time of this recording—you announced a somewhat, well, let’s say controversial change in how the pricing and licensing works. Now, as of—taking effect at the end of this year—the end of January, rather, of next year—Docker Desktop is free for folks to use for individual use, and that’s fine, and for corporate use, Docker Desktop also remains free until you are a large company defined by ten million in revenue a year and/or 250 employees or more. And that was interesting and I don’t think I’d seen that type of requirement placed before on what was largely an open-source project that’s now a developer tool. I believe there are closed-source aspects of it as well for the desktop experience, but please don’t quote me on that; I’m not here to play internet lawyer engineer. But at that point, the internet was predictably upset about this because it is easy to yell about any change that is coming, regardless.

I was less interested in that than I am in what the reception has been from your corporate customers because, let’s be clear, users are important, community is important, but goodwill will not put food on the table past a certain point. There has to be a way to make a company sustainable, there has to be a recurring revenue model. I realize that you know this, but I’m sure there are people listening to this who are working in development somewhere who are, “Wait, you mean I need to add more value than I cost?” It was a hard revelation for [laugh] me back when I had been in the industry a few years—

Scott: [laugh]. Sure.

Corey: —and I’m still struggling with that—

Scott: Sure.

Corey: Some days.

Scott: You and me both. [laugh].

Corey: So, what has the reaction been from folks who have better channels of communicating with you folks than angry Twitter threads?

Scott: Yeah. Create surface area for a discussion, Corey. Let’s back up and talk on a couple points that you hit along the way there. One is, “What is Docker Desktop?” Docker Desktop is not just Docker Engine.

Docker Desktop is a way in which we take Docker Engine, Compose, Kubernetes, all important tools for developers building modern apps—Docker Build, so on and so forth—and we provide an integrated engineered product that is engineered for the native environments of Mac and Windows, and soon Linux. And so we make it super easy to get the container runtime, Kubernetes stack, the networking, the CLI, Compose, we make it super easy just to get that up and running and configured with smart defaults, secured, hardened, and importantly updated. So, any vulnerabilities patched and so on and so forth. The point is, it’s a product that is based on—to your comments—upstream open-source technologies, but it is an engineered commercial product—Docker Desktop is.

Corey: Docker Desktop is a fantastic tool; I use it myself. I could make a bunch of snide comments that on Mac, it’s basically there to make sure the fans are still working on the laptop, but again, computers are hard. I get that. It’s incredibly handy to have a graphical control panel. It turns out that I don’t pretend to understand those people, but some folks apparently believe that there are better user interfaces than text and an 80-character-wide terminal window. I don’t pretend to get those people, but not everyone has the joy of being a Linux admin for far too long. So, I get it, making it more accessible, making it easy, is absolutely worth using.

Scott: That’s right.

Corey: It’s not a hard requirement to run it on a laptop-style environment or developer workstation, but it makes it really convenient.

Scott: Before Docker desktop, one had to install a hypervisor, install a Linux VM, install Docker Engine on that Linux VM, bridge between the VM and the local CLI on the native desktop—like, lots of setup and maintenance and tricky stuff that can go wrong. Trust me how many times I stubbed my own toes on putting that together. And so Docker Desktop is designed to take all of that setup nonsense overhead away and just let the developer focus on the app. That’s what the product is, and just talking about where it came from, and how it uses these other upstream technologies. Yes, and so we made a move on August 31, as you noted, and the motivation was the following: one is, we started seeing large organizations using Docker Desktop at scale.

When I say ‘at scale,’ not one or two or ten developers; like, hundreds and thousands of developers. And they were clamoring for capabilities to help them manage those developer environments at scale. Second is, we saw them getting a lot of benefit in terms of productivity, and choice, and security from using Docker Desktop, and so we stood back and said, “Look, for us to scale our business, we’re at 10-plus million monthly active developers today. We know there’s 45 million developers coming in this decade; how do we keep scaling while giving a free experience, but still making sure we can fund our engineers and deliver features and additional value?” We looked at other projects, Corey.

The first thing we did is we looked outside our four walls, said, “How have other projects with free and open-source components navigated these waters?” And so the thresholds that you just mentioned, the 250 employees and the ten million revenue, were actually thresholds that we saw others put in place to draw lines between what is available completely for free and what is available for those users that now need to purchase subscription if they’re using it to create value for their organizations. And we’re very explicit about that. You could be using Docker for training, you could be using Docker for eval in those large organizations; we’re not going to chase you or be looking to you to step up to a subscription. However, if you’re using Docker Desktop in those environments, to build applications that run your business or that are creating value for your customers, then purchasing a subscription is a way for us to continue to invest in a product that the ecosystem clearly loves and is getting a lot of value out of. And so, that was again, the premise of this change. So, now to the root of your question is, so what’s the reaction? We’re very, very pleased. First off, yes, there were some angry voices out there.

Corey: Yeah. And I want to be clear, I’m not trivializing people who feel upset.

Scott: No.

Corey: When you’re suddenly using a thing that is free and discovering that, well, now you have to pay money for it, people are not generally going to be happy about that.

Scott: No.

Corey: When people are viewed themselves as part of the community, of contributing to what they saw as a technical revolution or a scrappy underdog and suddenly they find themselves not being included in some way, shape or form, it’s natural to be upset, I don’t want to trivialize—

Scott: Not at all.

Corey: People’s warm feelings toward Docker. It was a big part of a lot of folks’ personality, for better or worse, [laugh] for a few years in there. But the company needs to be sustainable, so what I’m really interested in is what has that reaction been from folks who are, for better or worse, “Yes, yes, we love Docker, but I don’t get to sign $100,000 deals because I just really like the company I’m paying the money to. There has to be business value attached to that.”

Scott: That’s right. That’s right. And to your point, we’re not trivializing either the reaction by the community, it was encouraging to see many community members got right away what we’re doing, they saw that still, a majority of them can continue using Docker for free under the Docker Personal subscription, and that was also intentional. And you saw on the internet and on Twitter and other social media, you saw them come and support the company’s moves. And despite some angry voices in there, there was overwhelmingly positive.

So, to your question, though, since August 31, we’ve been overwhelmed, actually, by the positive response from businesses that use Docker Desktop to build applications and run their businesses. And when I say overwhelmed, we were tracking—because Docker Desktop has a phone-home capability—we had a rough idea of what the baseline usage of Docker Desktops were out there. Well, it turns out, in some cases, there are ten times as many Docker Desktops inside organizations. And the average seems to be settling in around three times to four times as many. And we are already closing business, Corey.

In 12 business days, we have companies come through, say, “Yes, our developers use this product. Yes, it’s a valuable product. We’re happy to talk to a salesperson and give you over to procurement, and here we go.” So, you and I both been around long enough to know, like 12 working
days to have a signed agreement with an enterprise agreement is unheard of.

Corey: Yeah, but let’s be very clear here, on The Duckbill Group’s side of things where I do consulting projects, I sell projects to companies that are, “Great, this project will take, I don’t know, four to six weeks, whatever it happens to be, and, yeah, you’re going to turn a profit on this project in about the first four hours of the engagement.” It is basically push button and you will receive more money in your budget than you had when you started, and that is probably the easiest possible enterprise sale, and it still takes 60 to 90 days most of the time to close deals.

Scott: That’s right.

Corey: Trying to get a procurement deal for software through enterprise procurement processes is one of those things when people say, “Okay, we’re going to have a signature in Q3,” you have to clarify what year they’re talking about. So, 12 days is unheard of.

Scott: [laugh]. Yep. So, we’ve been very encouraged by that. And I’ll just give you a rough numbers: the overall response is ten times our baseline expectations, which is why—maybe unanticipated question, or you going to ask it soon—we came back within two weeks—because we could see this curve hit right away on the 31st of August—we came back and said, “Great.” Now, that we have the confidence that the community and businesses are willing to support us and invest in our sustainability, invest in the sustainable, scalable Docker, we came and we accelerated—pulled forward—items in our roadmap for developers using Docker Desktop, both for Docker Personal, for free in the community, as well as the subscribers.

So, things like Docker Desktop for Linux, right? Docker Desktop for Mac, Docker Desktop for Windows has been out there about five years, as I said. We have heard Docker Desktop for Linux rise in demand over those years because if you’re managing a large number of developers, you want a consistent environment across all the developers, whether they’re using Linux, Mac, or Windows desktops. So, Docker Desktop for Linux will give them that consistency across their entire development environment. That was the number two most requested feature on our public roadmap in the last year, and again, with the positive response, we’re now able to confidently invest in that. We’re hiring more engineers than planned, we’re pulling that forward in the roadmap to show that yes, we are about growing and growing sustainably, and now that the environment and businesses are supporting us, we’re happy to double down and create more value.

Corey: My big fear when the change was announced was the uncertainty inherent to it. Because if there’s one thing that big companies don’t like, it’s uncertainty because uncertainty equates to risk in their mind. And a lot of other software out there—and yes, Oracle Databases I am looking at you—have a historical track record of, “Okay, great. We have audit rights to inspect your environment, and then when we wind up coming in, we always find that there have been licensing shortfalls,” because people don’t know how far things spread internally, as well as, honestly, it’s accounting for this stuff in large, complex organizations is a difficult thing. And then there are massive fines at stake, and then there’s this whole debate back and forth.

Companies view contracts as if every company behaves like that when it comes down to per-seat licensing and the rest. My fear was that that risk avoidance in large companies would have potentially made installing Docker Desktop in their environment suddenly a non-starter across the board, almost to the point of being something that you would discipline employees for, which is not great. And it seems from your response, that has not been a widespread reaction. Yes of course, there’s always going to be some weird company somewhere that does draconian things that we don’t see, but the fact that you’re not sitting here, telling me that you’ve been taking a beating from this from your enterprise buyers, tells me you’re onto something.

Scott: I think that’s right, Corey. And as you might expect, the folks that don’t reach out are silent, and so we don’t see folks who don’t reach out to us. But because so many have reached out to us so positively, and basically quickly gone right to a conversation with procurement versus any sort of back-and-forth or questions and such, tells us we are on the right track. The other thing, just to be really clear is, we did work on this before the August 31 announcement as well—this being how do we approach licensing and compliance and such—and we found that 80% of organizations, 80% of businesses want to be in compliance, they have a—not just want to be in compliance, but they have a history of being in compliance, regardless of the enforcement mechanism and whatnot. And so that gave us confidence to say, “Hey, we’re going to trust our users. We’re going to say, ‘grace period ends on January 31.’”

But we’re not shutting down functionality, we’re not sending in legal [laugh] activity, we’re not putting any sort of strictures on the product functionality because we have found most people love the product, love what it does for them, and want to see the company continue to innovate and deliver great features. And so okay, you might say, “Well, doesn’t that 20% represent opportunity?” Yeah. You know, it does, but it’s a big ecosystem. The 80% is giving us a great boost and we’re already starting to plow that into new investment. And let’s just start there; let’s start there and grow from there.

This episode is sponsored by our friends at Oracle Cloud. Counting the pennies, but still dreaming of deploying apps instead of "Hello, World" demos? Allow me to introduce you to Oracle's Always Free tier. It provides over 20 free services and infrastructure, networking databases, observability, management, and security.

And - let me be clear here - it's actually free. There's no surprise billing until you intentionally and proactively upgrade your account. This means you can provision a virtual machine instance or spin up an autonomous database that manages itself all while gaining the networking load, balancing and storage resources that somehow never quite make it into most free tiers needed to support the application that you want to build.

With Always Free you can do things like run small scale applications, or do proof of concept testing without spending a dime. You know that I always like to put asterisks next to the word free. This is actually free. No asterisk. Start now. Visit https://snark.cloud/oci-free that's https://snark.cloud/oci-free.

Corey: I also have a hard time imagining that you and your leadership team would be short-sighted enough to say, “Okay, that”—even 20% of companies that are willing to act dishonestly around stuff like that seems awfully high to me, but assuming it’s accurate, would tracking down that missing 20% be worth setting fire to the tremendous amount of goodwill that Docker still very much enjoys? I have a hard time picturing any analysis where that’s even a question other than something you set up to make fun of.

Scott: [laugh]. No, that’s exactly right Corey, it wouldn’t be worth it which is why again, we came out of the gate with like, we’re going to trust our users. They love the community, they love the product, they want to support us—most of them want to support us—and, you know, when you have most, you’re never going to get a hundred percent. So, we got most and we’re off to a good start, by all accounts. And look, a lot of folks too sometimes will be right in that gray middle where you let them know that they’re getting away with something they’re like, “All right, you caught me.”

We’ve seen that behavior before. And so, we can see all this activity out there and we can see if folks have a license or compliance or not, and sometimes just a little tap on the shoulder said, “Hey, did you know that you might be paying for that?” We’ve seen most folks at the time say, “Ah, okay. You caught me. Happy to talk to procurement.”

So, this does not have to be heavy-handed as you said, it does not have to put at risk the goodwill of the 80%. And we don’t have to get a hundred percent to have a great successful business and continuing successful community.

Corey: Yeah. I’ll also point out that, by my reading of your terms and conditions and how you’ve specified this—I mean, this is not something I’ve asked you about, so this could turn into a really awkward conversation but I’m going to roll with it anyway, it explicitly states that it is and will remain free for personal development.

Scott: That is correct.

Corey: When you’re looking at employees who work at giant companies and have sloppy ‘bring your own device’ controls around these things, all right, they have it installed on their work machine because in their spare time, they’re building an app somewhere, they’re not going to get a nasty gram, and they’re not exposing their company to liability by doing that?

Scott: That is exactly correct. And moreover, just keep looking at those use cases, if the company is using it for internal training or if the company is using it to evaluate someone else’s technology, someone else’s software, all those cases are outside the pay-for subscription. And so we believe it’s quite generous in allowing of trials and tests and use cases that make it accessible and easy to try, easy to use, and it’s just in the case where if you’re a large organization and your developers are using it to build applications for your business and for your customers, thus you’re getting a lot of value using the product, we’re asking you to share that value with us so we can continue to invest in the product.

Corey: And I think that’s a reasonable expectation. The challenge that Docker seems to have had for a while has been that the interesting breakthrough, revelatory stuff that you folks did was all open-source. It was a technology that was incredibly inspired in a bunch of different ways. I am, I guess, mature enough to admit that my take that, “Oh, Docker is terrible”—which was never actually my take—was a little short-sighted. I’m very good at getting things wrong across the board, and that is no exception.

I also said virtualization was a flash in the pan and look how that worked out. I was very anti-cloud, et cetera, et cetera. Times change, people change, and doubling down on being wrong gains you nothing. But the question that was always afterwards what is the monetization strategy? Because it’s not something you can give away for free and make it up in volume?

Even VC money doesn’t quite work like that forever, so there’s a—the question is, what is the monetization strategy that doesn’t leave people either resenting you because, “Remember that thing that used to be free isn’t anymore? Doesn’t it suck to be you?” And is still accessible as broadly as you are, given the sheer breadth and diversity of your community? Like I can make bones about the fact that ten million in revenue and 250 employees are either worlds apart, or the wrong numbers, or whatever it is, but it’s not going to be some student somewhere sitting someplace where their ramen budget is at risk because they have to spend $5 a month or whatever it is to have this thing. It doesn’t apply to them.

And this feels like, unorthodox though it certainly is, it’s not something to be upset about in any meaningful sense. The people that I think would actually be upset and have standing to be upset about this are the enterprise buyers, and you’re hearing from them in what is certainly—because I will hear it if not—that this is something they’re happy about. They are thrilled to work with you going forward. And I think it makes sense. Even when I was doing stuff as an independent consultant, before I formalized the creation of The Duckbill Group and started hiring people, my policy was always to not use the free tier of things, even if I fit into them because I would much rather personally be a paying customer, which elevates the, I guess, how well my complaints are received.

Because I’m a free user, I’m just another voice on Twitter; albeit a loud one and incredibly sarcastic one at times. But if I’m a paying customer, suddenly the entire tenor of that conversation changes, and I think there’s value to that. I’ve always had the philosophy of you pay for the things you use to make money. And that—again, that is something that’s easy for me to say now. Back when I was in crippling debt in my 20s, I assure you, it was not, but I still made the effort for things that I use to make a living.

Scott: Yeah.

Corey: And I think that philosophy is directionally correct.

Scott: No, I appreciate that. There’s a lot of good threads in there. Maybe just going way back, Docker stands on the shoulders of giants. There was a lot of work with container tech in the Linux kernel, and you and I were talking before about it goes back to LPAR on IBMs, and you know, BS—Berkeley’s—

Corey: BSD jails and chroots on Linux. Yeah.

Scott: Chroot, right? I mean, Bill Joy, putting chroot in—

Corey: And Tupperware parties, I’m sure. Yeah.

Scott: Right. And all credit to Solomon Hykes, Docker’s founder, who took a lot of good up and coming tech—largely on the ops side and in Linux kernel—took the primitives from Git and combined that with immutable copy-on-write file system and put those three together into a really magical combination that simplified all this complexity of dependency management and portability of images across different systems. And so in some sense, that was the magic of standing on these giant shoulders but seeing how these three different waves of innovation or three different flows of innovation could come together to a great user experience. So, also then moving forward, I wouldn’t say they’re happy, just to make sure you don’t get inbound, angry emails—the enterprise buyers—but they do recognize the value of the product, they think the economics are fair and straight ahead, and to your point about having a commercial relationship versus free or non-existing relationship, they’re seeing that, “Oh, okay, now I have insight into the roadmap. Now, I can prioritize my requirements that my devs have been asking for. Now, I can double-down on the secure supply chain issues, which I’ve been trying to get in front of for years.”

So, it gives them an avenue that now, much different than a free user as you observed, it’s a commercial relationship where it’s two way street versus, “Okay, we’re just going to use this free stuff and we don’t have much of a say because it’s free, and so on and so forth.” So, I think it’s been an eye-opener for both the company but also for the businesses. There is a lot of value in a commercial relationship beyond just okay, we’re going to invest in new features and new value for developers.

Corey: The challenge has always been how do you turn something that is widely beloved, that is effectively an open-source company, into money? There have been a whole bunch of questions about this, and it seems that the consensus that has emerged is that a number of people for a long time mistook open-source for a business model instead of a strategy, and it’s very much not. And a lot of companies are attempting to rectify that with weird license changes where, “Oh, you’re not allowed to take our code and build a service out of it if you’re a cloud provider.” Amazon’s product strategy is, of course, “Yes,” so of course, there’s always going to be something coming out of AWS that is poorly documented, has a ridiculous name, and purports to do the same thing for way less money, except magically you pay them by the hour. I digress.

Scott: No, it’s a great surface area, and you’re right I completely didn’t answer that question. [laugh]. So—

Corey: No, it’s fair. It’s—

Scott: Glad you brought it back up.

Corey: —a hard problem. It’s easy to sit here and say, “Well, what I think they should do”—but all of those solutions fall apart under ten seconds of scrutiny.

Scott: Super, super hard problem which, to be fair, we as a team and a community wrestled with for years. But here’s where we landed, Corey. The short version is that you can still have lots of great upstream open-source technologies, and you’ll have an early adopter community that loves those, use those, gets a lot of progress running fast and far with those, but we’ve found that the vast majority of the market doesn’t want to spend its time cobbling together bits and bytes of open-source tech, and maintaining it, and patching it, and, and, and. And so what we’re offering is an engineered product that takes the upstream but then adds a lot of value—we would say—to make it an engineered, easy to use, easy to configure, upgraded, secure, so on and so forth. And the convenience of that versus having to cobble together your own environment from upstream has proved to be what folks are willing to pay for. So, it’s the classic kind of paying for time and convenience versus not.

And so that is one dimension. And the other dimension, which you already referenced a little bit with AWS is that we have SaaS; we have a SaaS product in Docker Hub, which is providing a hosted registry with quality content that users know is updated not less than every 30 days, that is patched and maintained by us. And so those are examples of, in some sense, consumption [unintelligible 00:27:53]. So, we’re using open-source to build this SaaS service, but the service that users receive, they’re willing to pay for because they’re not having to patch the Mongo upstream, they’re not having to roll the image themselves, they’re not having to watch the CVEs and scramble when everything comes out. When there’s a CVE out in our upstream, our official images are patched no less than 24 hours later and typically within hours.

That’s an example of a service, but all based on upstream open-source tech that for the vast majority of uses are free. If you’re consuming a lot of that, then there’s a subscription that kicks in there as well. But we’re giving you value in exchange for you having to spend your time, your engineers, managing all that that I just walked through. So, those are the two avenues that we found that are working well, that seem to be a fair trade and fair balance with the community and the rest of the ecosystem.

Corey: I think the hardest part for a lot of folks is embracing change. And I have encountered this my entire career where I started off doing large-scale email systems administration, and hey, turns out that’s not really a thing anymore. And I used to be deep in the bowels of Postfix, for example. I’m referenced in the SVN history of Postfix, once upon a time, just for helping with documentation and finding weird corner cases because I’m really good at breaking things by accident. And I viewed it as part of my identity.

And times have changed and moved on; I don’t run Postfix myself for anything anymore. I haven’t touched it in years. Docker is still there and it’s still something that people are actively using basically everywhere. And there’s a sense of ownership and identity for especially early adopters who glom on to it because it is such a better way of doing some things that it is almost incomprehensible that we used to do it any other way. That’s transformation.

That’s something awesome. But people want to pretend that we’re still living in that era where technology has not advanced. The miraculous breakthrough in 2013 is today’s de rigueur type of environment where this is just, “Oh, yeah. Of course you’re using Docker.” If you’re not, people look at you somewhat strangely.

It’s like, “Oh, I’m using serverless.” “Okay, but you can still build that in Docker containers. Why aren’t you doing that?” It’s like, “Oh, I don’t believe in running anything that doesn’t make me pay AWS by the second.” So okay, great. People are going to have opinions on this stuff. But time marches on and whatever we wish the industry would do, it’s going to make its own decisions and march forward. There’s very little any of us can do to change that.

Scott: That’s right. Look, it was a single container back in 2013, 2014, right? And now what we’re seeing—and you kind of went there—is we’re separating the implementation of service from the service. So, the service could be implemented with a container, could be a serverless function, could be a hosted XYZ as a service on some cloud, but what developers want to do is—what they’re moving towards is, assemble your application based on services regardless of the how. You know, is that how a local container? To your point, you can roll a local serverless function now in an OCI image, and push it to Amazon.

Corey: Oh, yeah. It’s one of that now 34 ways I found to run containers on AWS.

Scott: [laugh]. You can also, in Compose, abstract all that complexity away. Compose could have three services in it. One of those services is a local container, one of those services might be a local serverless function that you’re running to test, and one of those services could be a mock to a Database as a Service on a cloud. And so that’s where we are.

We’ve gone beyond the single-container Docker run, which is still incredibly powerful but now we’re starting to uplevel to applications that consist of multiple services. And where do those services run? Increasingly, developers do not need to care. And we see that as our mission is continue to give that type of power to developers to abstract out the how, extract out the infrastructure so they can just focus on building their app.

Corey: Scott, I want to thank you so much for taking the time to speak with me. If people want to learn more—and that could mean finding out your opinions on things, potentially yelling at you about pricing changes, more interestingly, buying licenses for their large companies to run this stuff, and even theoretically, since you alluded to it a few minutes ago, look into working at Docker—where can they find you?

Scott: No, thanks, Corey. And thank you for the time to discuss and look back over both years, but also zoom in on the present day. So, www.docker.com; you can find any and all what we just walked through. They’re more than happy to yell at me on Twitters at @scottcjohnston, and we have a public roadmap that is in GitHub. I’m not going to put the URL here, but you can find it very easily. So, we love hearing from our community, we love engaging with them, we love going back and forth. And it’s a big community; jump in, the waters warm, very welcoming, love to have you.

Corey: And we’ll of course, but links to that in the [show notes. 00:32:28] Thank you so much for your time. I really do appreciate it.

Scott: Thank you, Corey. Right back at you.

Corey: Scott Johnston, CEO of Docker. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with a comment telling me that Docker isn’t interested in at all because here’s how to do exactly what Docker does in LPARs on your mainframe until the AWS/400 comes to [unintelligible 00:33:02].

Scott: [laugh].

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Chloe

Chloe is a Bay Area based Cloud Advocate for Microsoft. Previously, she worked at Sentry.io where she created the award winning Sentry Scouts program (a camp themed meet-up ft. patches, s’mores, giant squirrel costumes, and hot chocolate), and was featured in the Grace Hopper Conference 2018 gallery featuring 15 influential women in STEM by AnitaB.org. Her projects and work with Azure have ranged from fake boyfriend alerts to Mario Kart 'astrology', and have been featured in VICE, The New York Times, as well as SmashMouth's Twitter account. Chloe holds a BA in Drama from San Francisco State University and is a graduate of Hackbright Academy. She prides herself on being a non-traditional background engineer, and is likely one of the only engineers who has played an ogre, crayon, and the back-end of a cow on a professional stage. She hopes to bring more artists into tech, and more engineers into the arts.

Links:

  • Twitter: https://twitter.com/ChloeCondon
  • Instagram: https://www.instagram.com/gitforked/
  • YouTube: https://www.youtube.com/c/ChloeCondonVideos

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Vultr. Spelled V-U-L-T-R because they’re all about helping save money, including on things like, you know, vowels. So, what they do is they are a cloud provider that provides surprisingly high performance cloud compute at a price that—while sure they claim its better than AWS pricing—and when they say that they mean it is less money. Sure, I don’t dispute that but what I find interesting is that it’s predictable. They tell you in advance on a monthly basis what it’s going to going to cost. They have a bunch of advanced networking features. They have nineteen global locations and scale things elastically. Not to be confused with openly, because apparently elastic and open can mean the same thing sometimes. They have had over a million users. Deployments take less that sixty seconds across twelve pre-selected operating systems. Or, if you’re one of those nutters like me, you can bring your own ISO and install basically any operating system you want. Starting with pricing as low as $2.50 a month for Vultr cloud compute they have plans for developers and businesses of all sizes, except maybe Amazon, who stubbornly insists on having something to scale all on their own. Try Vultr today for free by visiting: vultr.com/screaming, and you’ll receive a $100 in credit. Thats v-u-l-t-r.com slash screaming.

Corey: This episode is sponsored in part by Honeycomb. When production is running slow, it's hard to know where problems originate: is it your application code, users, or the underlying systems? I’ve got five bucks on DNS, personally. Why scroll through endless dashboards, while dealing with alert floods, going from tool to tool to tool that you employ, guessing at which puzzle pieces matter? Context switching and tool sprawl are slowly killing both your team and your business. You should care more about one of those than the other, which one is up to you. Drop the separate pillars and enter a world of getting one unified understanding of the one thing driving your business: production. With Honeycomb, you guess less and know more. Try it for free at Honeycomb.io/screaminginthecloud. Observability, it’s more than just hipster monitoring.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Somehow in the years this show has been running, I’ve only had Chloe Condon on once. In that time, she’s over for dinner at my house way more frequently than that, but somehow the stars never align to get us together in front of microphones and have a conversation. First, welcome back to the show, Chloe. You’re a senior cloud advocate at Microsoft on the Next Generation Experiences Team. It is great to have you here.

Chloe: I’m back, baby. I’m so excited. This is one of my favorite shows to listen to, and it feels great to be a repeat guest, a friend of the pod. [laugh].

Corey: Oh, yes indeed. So, something-something cloud, something-something Microsoft, something-something Azure, I don’t particularly care, in light of what it is you have going on that you have just clued me in on, and we’re going to talk about that to start. You’re launching something new called Master Creep Theatre and I have a whole bunch of questions. First and foremost, is it theater or theatre? How is that spelled? Which—the E and the R, what direction does that go in?

Chloe: Ohh, I feel like it’s going to be the R-E because that makes it very fancy and almost British, you know?

Corey: Oh, yes. And the Harlequin mask direction it goes in, that entire aesthetic, I love it. Please tell me what it is. I want to know the story of how it came to be, the sheer joy I get from playing games with language alone guarantee I’m going to listen to whatever this is, but please tell me more.

Chloe: Oh, my goodness. Okay, so this is one of those creative projects that’s been on my back burner forever where I’m like, someday when I have time, I’m going to put all my time [laugh] and energy into this. So, this originally stemmed from—if you don’t follow me on Twitter, oftentimes when I’m not tweeting about ’90s nostalgia, or Clippy puns, or Microsoft silly throwback things to Windows 95, I get a lot of weird DMs. On every app, not just Twitter. On Instagram, Twitter, LinkedIn, oh my gosh, what else is there?

Corey: And I don’t want to be clear here just to make this absolutely crystal clear, “Hey, Chloe, do you want to come back on Screaming in the Cloud again?” Is not one of those weird DMs to which you’re referring?

Chloe: No, that is a good DM. So, people always ask me, “Why don’t you just close your DMs?” Because a lot of high profile people on the internet just won’t even have their DMs open.

Corey: Oh, I understand that, but I’m the same boat. I would have a lot less nonsense, but at the same time, I want—at least in my case—I want people to be able to reach out to me because the only reason I am what I am is that a bunch of people who had no reason to do it did favors for me—

Chloe: Yes.

Corey: —and I can’t ever repay it, I can only ever pay it forward and that is the cost of doing favors. If I can help someone, I will, and that’s hard to do with, “My DMs are closed so hunt down my email address and send me an email,” and I’m bad at email.

Chloe: Right. I’m terrible at email as well, and I’m also terrible at DMs [laugh]. So, I think a lot of folks don’t understand the volume at which I get messages, which if you’re a good friend of mine, if you’re someone like Corey or a dear friend like Emily, I will tell you, “Hey, if you actually need to get ahold of me, text me.” And text me a couple times because I probably see it and then I have ADHD, so I won’t immediately respond. I think I respond in my head but I don’t.

But I get anywhere from, I would say, ohh, like, 30 on a low day to 100 on a day where I have a viral tweet about getting into tech with a non-traditional background or something like that. And these DMs that I get are really lovely messages like, “Thank you for the work you do,” or, “I decided to do a cute manicure because the [laugh] manicure you posted,” too, “How do I get into tech? How do I get a job at Microsoft?” All kinds of things. It runs the gamut between, “Where’s your shirt from?” Where—[laugh]—“What’s your mother’s maiden name?”

But a lot of the messages that I get—and if you’re a woman on the internet with any sort of presence, you know how there’s that, like—what’s it called in Twitter—the Other Messages feature that’s like, “Here’s the people you know. Here’s the people”—the message requests. For the longest time were just, “Hey,” “Hi,” “Hey dear,” “Hi pretty,” “Hi ma’am,” “Hello,” “Love you,” just really weird stuff. And of course, everyone gets these; these are bots or scammers or whatever they may be—or just creeps, like weird—and always the bio—not always but I [laugh] would say, like, these accounts range from either obviously a bot where it’s a million different numbers, an account that says, “Father, husband, lover of Jesus Christ and God.” Which is so [laugh] ironic… I’m like, “Why are you in my DMs?”

Corey: A man of God, which is why I’m in your DMs being creepy.

Chloe: Exactly. Or—

Corey: Just like Christ might have.

Chloe: And you would be shocked, Corey, at how many. The thing that I love to say is Twitter is not a dating site. Neither is LinkedIn. Neither is Instagram. I post about my boyfriend all the time, who you’ve met, and we adore Ty Smith, but I’ve never received any unsolicited images, knock on wood, but I’m always getting these very bait-y messages like, “Hey, beautiful. I want to take you out.” And you would be shocked at how many of these people are doing it from their professional business account. [laugh]. Like, works at AWS, works at Google; it’s like, oh my God. [laugh].

Corey: You get this under your name, right? It ties back to it. Meanwhile—again, this is one of those invisible areas of privilege that folks who look like me don’t have to deal with. My DM graveyard is usually things like random bot accounts, always starting with, “Hi,” or, “Hey.” If you want to guarantee I never respond to you, that is what you say. I just delete those out of hand because I don’t notice or care. It is either a bot, or a scam, or someone who can’t articulate what they’re actually trying to get from me—

Chloe: Exactly.

Corey: —and I don’t have the time for it. Make your request upfront. Don’t ask to ask; just ask.

Chloe: I think it’s important to note, also, that I get a lot of… different kinds of these messages and they try to respond to everyone. I cannot. If I responded to everybody’s messages that I got, I just wouldn’t have any time to do my job. But the thing that I always say to people—you know, and managers have told me in the past, my boyfriend has encouraged me to do this, is when people say things like, “Close your DMs,” or, “Just ignore them,” I want to have the same experience that everybody else has on the internet. Now, it’s going to be a little different, of course, because I look and act and sound like I do, and of course, podcasts are historically a visual medium, so I’m a five-foot-two, white, bright orange-haired girl; I’m a very quirky individual.

Corey: Yes, if you look up ‘quirky,’ you’re right there under the dictionary definition. And every time—like, when we were first hanging out and you mentioned, “Oh yeah, I used to be in theater.” And it’s like, “You know, you didn’t even have to tell me that, on some level.” Which is not intended to be an insult. It’s just theater folks are a bit of a type, and you are more or less the archetype of what a theatre person is, at least to my frame of reference.

Chloe: And not only that, but I did musicals, so you can’t see the jazz hands now, but–yeah, my degree is in drama. I come from that space and I just, you know, whenever people say, “Just ignore it,” or, “Close your DMs,” I’m like, I want people to be able to reach out to me; I want to be able to message one-on-one with Corey and whoever, when—as needed, and—

Corey: Why should I close my DMs?

Chloe: Yeah.

Corey: They’re the ones who suck. Yeah.

Chloe: [laugh]. But over the years, to give people a little bit of context, I’ve been working in tech a long time—I’ve been working professionally in the DevRel space for about five or six years now—but I’ve worked in tech a long time, I worked as a recruiter, an office admin, executive assistant, like, I did all of the other areas of tech, but it wasn’t until I got a presence on Twitter—which I’ve only been on Twitter for I think five years; I haven’t been on there that long, actively. And to give some context on that, Twitter is not a social media platform used in the theater space. We just use Instagram and Facebook, really, back in the day, I’m not on Facebook at all these days. So, when I discovered Twitter was cool—and I should also mention my boyfriend, Ty, was working at Twitter at the time and I was like, “Twitter’s stupid. Who would go on this—[laugh] who uses this app?”

Fast-forward to now, I’m like—Ty’s like, “Can you please get off Twitter?” But yeah, I think I’ve just been saving these screenshots over the last five or so years from everything from my LinkedIn, from all the crazy stuff that I dealt with when people thought I was a Bitcoin influencer to people being creepy. One of the highlights that I recently found when I was going back and trying to find these for this series that I’m doing is there was a guy from Australia, DMed me something like, “Hey, beautiful,” or, “Hey, sexy,” something like that. And I called him out. And I started doing this thing where I would post it on Twitter.

I would usually hide their image with a clown emoji or something to make it anonymous, or not to call them out, but in this one I didn’t, and this guy was defending himself in the comments, and to me in my DM’s saying, “Oh, actually, this was a social experiment and I have all the screenshots of this,” right? So, imagine if you will—so I have conversations ranging from things like that where it’s like, “Actually I messaged a bunch of people about that because I’m doing a social experiment on how people respond to, ‘Hey beautiful. I’d love to take you out some time in Silicon Valley.’” just the weirdest stuff right? So, me being the professional performer that I am, was like, these are hilarious.

And I kept thinking to myself, anytime I would get these messages, I was like, “Does this work?” If you just go up to someone and say, “Hey”—do people meet this way? And of course, you get people on Twitter who when you tweet something like that, they’re like, “Actually, I met my boyfriend in Twitter DMs,” or like, “I met my boyfriend because he slid into my DMs on Instagram,” or whatever. But that’s not me. I have a boyfriend. I’m not interested. This is not the time or the place.

So, it’s been one of those things on the back burner for three or four years that I’ve just always been saving these images to a folder, thinking, “Okay, when I have the time when I have the space, the creative energy and the bandwidth to do this,” and thankfully for everyone I do now, I’m going to do dramatic readings of these DMs with other people in tech, and show—not even just to make fun of these people, but just to show, like, how would this work? What do you expect the [laugh] outcome to be? So Corey, for example, if you were to come on, like, here’s a great example. A year ago—this is 2018; we’re in 2021 right now—this guy messaged me in December of 2018, and was like, “Hey,” and then was like, “I would love to be your friend.” And I was like, “Nope,” and I responded, “Nope, nope, nope, nope.” There’s a thread of this on Twitter. And then randomly, three weeks ago, just sent me this video to the tune of Enrique Iglesias’ “Rhythm Divine” of just images of himself. [laugh]. So like, this comedy [crosstalk 00:10:45]—

Corey: Was at least wearing pants?

Chloe: He is wearing pants. It’s very confusing. It’s a picture—a lot of group photos, so I didn’t know who he was. But in my mind because, you know, I’m an engineer, I’m trying to think through the end-user experience. I’m like, “What was your plan here?”

With all these people I’m like, “So, your plan is just to slide into my DMs and woo me with ‘Hey’?” [laugh]. So, I think it’ll be really fun to not only just show and call out this behavior but also take submissions from other people in the industry, even beyond tech, really, because I know anytime I tweet an example of this, I get 20 different women going, “Oh, my gosh, you get these weird messages, too?” And I really want to show, like, A, to men how often this happens because like you said, I think a lot of men say, “Just ignore it.” Or, “I don’t get anything like that. You must be asking for it.”

And I’m like, “No. This comes to me. These people find us and me and whoever else out there gets these messages,” and I’m just really ready to have a laugh at their expense because I’ve been laughing for years. [laugh].

Corey: Back when I was a teenager, I was working in some fast food style job, and one of my co-workers saw customer, walked over to her, and said, “You’re beautiful.” And she smiled and blushed. He leaned in and kissed her.

Chloe: Ugh.

Corey: And I’m sitting there going what on earth? And my other co-worker leaned over and is like, “You do know that’s his girlfriend, right?” And I have to feel like, on some level, that is what happened to an awful lot of these broken men out on the internet, only they didn’t have a co-worker to lean over and say, “Yeah, they actually know each other.” Which is why we see all this [unintelligible 00:12:16] behavior of yelling at people on the street as they walk past, or from a passing car. Because they saw someone do a stunt like that once and thought, “If it worked for them, it could work for me. It only has to work once.”

And they’re trying to turn this into a one day telling the grandkids how they met their grandmother. And, “Yeah, I yelled at her from a construction site, and it was love at first ‘Hey, baby.’” That is what I feel is what’s going on. I have never understood it. I look back at my dating history in my early 20s, I look back now I’m like, “Ohh, I was not a great person,” but compared to these stories, I was a goddamn prince.

Chloe: Yeah.

Corey: It’s awful.

Chloe: It’s really wild. And actually, I have a very vivid memory, this was right bef—uh, not right before the pandemic, but probably in 2019. I was speaking on a lot of conferences and events, and I was at this event in San Jose, and there were not a lot of women there. And somehow this other lovely woman—I can’t remember her name right now—found me afterwards, and we were talking and she said, “Oh, my God. I had—this is such a weird event, right?”

And I was like, “Yeah, it is kind of a weird vibe here.” And she said, “Ugh, so the weirdest thing happened to me. This guy”—it was her first tech conference ever, first of all, so you know—or I think it was her first tech conference in the Bay Area—and she was like, “Yeah, this guy came to my booth. I’ve been working this booth over here for this startup that I work at, and he told me he wanted to talk business. And then I ended up meeting him, stupidly, in my hotel lobby bar, and it’s a date. Like, this guy is taking me out on a date all of a sudden,” and she was like, “And it took me about two minutes to just to be like, you know what? This is inappropriate. I thought this is going to be a business meeting. I want to go.”

And then she shows me her hands, Corey, and she has a wedding ring. And she goes, “I’m not married. I have bought five or six different types of rings on Wish App”—or wish.com, which if you’ve never purchased from Wish before, it’s very, kind of, low priced jewelry and toys and stuff of that nature. And she said, “I have a different wedding ring for every occasion. I’ve got my beach fake wedding ring. I’ve got my, we-got-married-with-a-bunch-of-mason-jars-in-the-woods fake wedding ring.”

And she said she started wearing these because when she did, she got less creepy guys coming up to her at these events. And I think it’s important to note, also, I’m not putting it out there at all that I’m interested in men. If anything, you know, I’ve been [laugh] with my boyfriend for six years never putting out these signals, and time and time again, when I would travel, I was very, very careful about sharing my location because oftentimes I would be on stage giving a keynote and getting messages while I delivered a technical keynote saying, “I’d love to take you out to dinner later. How long are you in town?” Just really weird, yucky, nasty stuff that—you know, and everyone’s like, “You should be flattered.”

And I’m like, “No. You don’t have to deal with this. It’s not like a bunch of women are wolf-whistling you during your keynote and asking what your boob size is.” But that’s happening to me, and that’s an extra layer that a lot of folks in this industry don’t talk about but is happening and it adds up. And as my boyfriend loves to remind me, he’s like, “I mean, you could stop tweeting at any time,” which I’m not going to do. But the more followers you get, the more inbound you get. So—

Corey: Right. And the hell of it is, it’s not a great answer because it’s closing off paths of opportunity. Twitter has—

Chloe: Absolutely.

Corey: —introduced me to clients, introduced me to friends, introduced me to certainly an awful lot of podcast guests, and it informs and shapes a lot of the opinions that I hold on these things. And this is an example of what people mean when they talk about privilege. Where, yeah, “Look at Corey”—I’ve heard someone say once, and, “Nothing was handed to him.” And you’re right, to be clear, I did not—like, no one handed me a microphone and said, “We’re going to give you a podcast, now.” I had to build this myself.

But let’s be clear, I had no headwinds of working against me while I did it. There’s the, you still have to do things, but you don’t have an entire cacophony of shit heels telling you that you’re not good enough in a variety of different ways, to subtly reinforcing your only value is the way that you look. There isn’t this whole, whenever you get something wrong and it’s a, “Oh, well, that’s okay. We all get things wrong.” It’s not the, “Girls suck at computers,” trope that we see so often.

There’s a litany of things that are either supportive that work in my favor, or are absent working against me that is privilege that is invisible until you start looking around and seeing it, and then it becomes impossible not to. I know I’ve talked about this before on the show, but no one listens to everything and I just want to subtly reinforce that if you’re one of those folks who will say things like, “Oh, privilege isn’t real,” or, “You can have bigotry against white people, too.” I want to be clear, we are not the same. You are not on my side on any of this, and to be very direct, I don’t really care what you have to say.

Chloe: Yeah. And I mean, this even comes into play in office culture and dynamics as well because I am always the squeaky wheel in the room on these kind of things, but a great example that I’ll give is I know several women in this industry who have had issues when they used to travel for conferences of being stalked, people showing up at their hotel rooms, just really inappropriate stuff, and for that reason, a lot of folks—including myself—wouldn’t pick the conference event—like, typically they’ll be like, “This is the hotel everyone’s staying at.” I would very intentionally stay at a different hotel because I didn’t want people knowing where I was staying. But I started to notice once a friend of mine, who had an issue with this [unintelligible 00:17:26], I really like to be private about where I’m staying, and sometimes if you’re working at a startup or larger company, they’ll say, “Hey, everyone put in this Excel spreadsheet or this Google Doc where everyone’s staying and how to contact them, and all this stuff.” And I think it’s really important to be mindful of these things.

I always say to my friends—I’m not going out too much these days because it’s a pandemic—and I’ve done Twitter threads on this before where I never post my location; you will never see me. I got rid of Swarm a couple [laugh] years ago because people started showing up where I was. I posted photos before, you know, “Hey, at the lake right now.” And people have shown up. Dinners, people have recognized me when I’ve been out.

So, I have an espresso machine right over here that my lovely boyfriend got me for my birthday, and someone commented, “Oh, we’re just going to act like we don’t see someone’s reflection in the”—like, people Zoom in on images. I’ve read stories from cosplayers online who, they look into the reflection of a woman’s glasses and can figure out where they are. So, I think there’s this whole level. I’m constantly on alert, especially as a woman in tech. And I have friends here in the Bay Area, who have tweeted a photo at a barbecue, and then someone was like, “Hey, I live in the neighborhood, and I recognize the tree.”

First of all, don’t do that. Don’t ever do that. Even if you think you’re a nice, unassuming guy or girl or whatever, don’t ever [laugh] do that. But I very intentionally—people get really confused, my friends specifically. They’re like, “Wait a second, you’re in Hawaii right now? I thought you were in Hawaii three weeks ago.” And I’m like, “I was. I don’t want anyone even knowing what island or continent I’m on.”

And that’s something that I think about a lot. When I post photo—I never post any photos from my window. I don’t want people knowing what my view is. People have figured out what neighborhood I live in based on, like, “I know where that graffiti is.” I’m very strategic about all this stuff, and I think there’s a lot of stuff that I want to share that I don’t share because of privacy issues and concerns about my safety. And also want to say and this is in my thread on online safety as well is, don’t call out people’s locations if you do recognize the image because then you’re doxxing them to everyone like, “Oh”—

Corey: I’ve had a few people do that in response to pictures I’ve posted before on a house, like, “Oh, I can look at this and see this other thing and then intuit where you are.” And first, I don’t have that sense of heightened awareness on this because I still have this perception of myself as no one cares enough to bother, and on the other side, by calling that out in public. It’s like, you do not present yourself well at all. In fact, you make yourself look an awful lot like the people that we’re warned about. And I just don’t get that.

I have some of these concerns, especially as my audience has grown, and let’s be very clear here, I antagonize trillion-dollar companies for a living. So, first if someone’s going to have me killed, they can find where I am. That’s pretty easy. It turns out that having me whacked is not even a rounding error on most of these companies' budgets, unfortunately. But also I don’t have that level of, I guess, deranged superfan. Yet.

But it happens in the fullness of time, as people’s audiences continue to grow. It just seems an awful lot like it happens at much lower audience scale for folks who don’t look like me. I want to be clear, this is not a request for anyone listening to this, to try and become that person for me, you will get hosed, at minimum. And yes, we press charges here.

Chloe: AWSfan89, sliding into your DMs right after this. Yeah, it’s also just like—I mean, I don’t want to necessarily call out what company this was at, but personally, I’ve been in situations where I’ve thrown an event, like a meetup, and I’m like, “Hey, everyone. I’m going to be doing ‘Intro to blah, blah, blah’ at this time, at this place.” And three or four guys would show up, none of them with computers. It was a freaking workshop on how to do or deploy something, or work with an API.

And when I said, “Great, so why’d you guys come to this session today?” And maybe two have iPads, one just has a notepad, they’re like, “Oh, I just wanted to meet you from Twitter.” And it’s like, okay, that’s a little disrespectful to me because I am taking time out to do this workshop on a very technical thing that I thought people were coming here to learn. And this isn’t the Q&A. This is not your meet-and-greet opportunity to meet Chloe Condon, and I don’t know why you would, like, I put so much of my life online [laugh] anyway.

But yeah, it’s very unsettling, and it’s happened to me enough. Guys have shown up to my events and given me gifts. I mean, I’m always down for a free shirt or something, but it’s one of those things that I’m constantly aware of and I hate that I have to be constantly aware of, but at the end of the day, my safety is the number one priority, and I don’t want to get murdered. And I’ve tweeted this out before, our friend Emily, who’s similarly a lady on the internet, who works with my boyfriend Ty over at Uber, we have this joke that’s not a joke, where we say, “Hey if I’m murdered, this is who it was.” And we’ll just send each other screenshots of creepy things that people either tag us in, or give us feedback on, or people asking what size shirt we are. Just, wiki feed stuff, just really some of the yucky of the yuck out there.

And I do think that unless you have a partner, or a family member, or someone close enough to you to let you know about these things—because I don’t talk about these things a lot other than my close friends, and maybe calling out a weirdo here and there in public, but I don’t share the really yucky stuff. I don’t share the people who are asking what neighborhood I live in. I’m not sharing the people who are tagging me, like, [unintelligible 00:22:33], really tagging me in some nasty TikToks, along with some other women out there. There are some really bad actors in this community and it is to the point where Emily and I will be like, “Hey, when you inevitably have to solve my murder, here’s the [laugh] five prime suspects.” And that sucks. That’s [unintelligible 00:22:48] joke; that isn’t a joke, right? I suspect I will either die in an elevator accident or one of my stalkers will find me. [laugh].

Corey: It’s easy for folks to think, oh, well, this is a Chloe problem because she’s loud, she’s visible, she’s quirky, she’s different than most folks, and she brings it all on herself, and this is provably not true. Because if you talk to, effectively, any woman in the world in-depth about this, they all have stories that look awfully similar to this. And let me forestall some of the awful responses I know I’m going to get. And, “Well, none of the women I know have had experiences like this,” let me be very clear, they absolutely have, but for one reason or another, they either don’t see the need, or don’t see the value, or don’t feel safe talking to you about it.

Chloe: Yeah, absolutely. And I feel a lot of privilege, I’m very lucky that my boyfriend is a staff engineer at Uber, and I have lots of friends in high places at some of these companies like Reddit that work with safety and security and stuff, but oftentimes, a lot of the stories or insights or even just anecdotes that I will give people on their products are invaluable insights to a lot of these security and safety teams. Like, who amongst us, you know, [laugh] has used a feature and been like, “Wait a second. This is really, really bad, and I don’t want to tweet about this because I don’t want people to know that they can abuse this feature to stalk or harass or whatever that may be,” but I think a lot about the people who don’t have the platform that I have because I have 50k-something followers on Twitter, I have a pretty big online following in general, and I have the platform that I do working at Microsoft, and I can tweet and scream and be loud as I can about this. But I think about the folks who don’t have my audience, the people who are constantly getting harassed and bombarded, and I get these DMs all the time from women who say, “Thank you so much for doing a thread on this,” or, “Thank you for talking about this,” because people don’t believe them.

They’re just like, “Oh, just ignore it,” or just, “Oh, it’s just one weirdo in his basement, like, in his mom’s basement.” And I’m like, “Yeah, but imagine that but times 40 in a week, and think about how that would make you rethink your place and your position in tech and even outside of tech.” Let’s think of the people who don’t know how this technology works. If you’re on Instagram at all, you may notice that literally not only every post, but every Instagram story that has the word COVID in it, has the word vaccine, has anything, and they must be using some sort of cognitive scanning type thing or scanning the images themselves because this is a feature that basically says, hey, this post mentioned COVID in some way. I think if you even use the word mask, it alerts this.

And while this is a great feature because we all want accurate information coming out about the pandemic, I’m like, “Wait a minute. So, you’re telling me this whole time you could have been doing this for all the weird things that I get into my DMs, and people post?” And, like, it just shows you, yes, this is a global pandemic. Yes, this is something that affects everyone. Yes, it’s important we get information out about this, but we can be using these features in much [laugh] more impactful ways that protects people’s safety, that protects people’s ability to feel safe on a platform.

And I think the biggest one for me, and I make a lot of bots; I make a lot of Twitter bots and chatbots, and I’ve done entire series on this about ethical bot creation, but it’s so easy—and I know this firsthand—to make a Twitter account. You can have more than one number, you can do with different emails. And with Instagram, they have this really lovely new feature that if you block someone, it instantly says, “You just blocked so and so. Would you like to block any other future accounts they make?” I mean, seems simple enough, right?

Like, anything related—maybe they’re doing it by email, or phone number, or maybe it’s by IP, but like, that’s not being done on a lot of these platforms, and it should be. I think someone mentioned in one of my threads on safety recently that Peloton doesn’t have a block user feature. [laugh]. They’re probably like, “Well, who’s going to harass someone on Peloton?” It would happen to me. If I had a Peloton, [laugh] I assure you someone would find a way to harass me on there.

So, I always tell people, if you’re working at a company and you’re not thinking about safety and harassment tools, you probably don’t have anybody LGBTQ+ women, non-binary on your team, first of all, and you need to be thinking about these things, and you need to be making them a priority because if users can interact in some way, they will stalk, harass, they will find some way to misuse it. It seems like one of those weird edge cases where it’s like, “Oh, we don’t need to put a test in for that feature because no one’s ever going to submit, like, just 25 emojis.” But it’s the same thing with safety. You’re like, who would harass someone on an app about bubblegum? One of my followers were. [laugh].

Corey: This episode is sponsored by our friends at Oracle HeatWave is a new high-performance accelerator for the Oracle MySQL Database Service. Although I insist on calling it “my squirrel.” While MySQL has long been the worlds most popular open source database, shifting from transacting to analytics required way too much overhead and, ya know, work. With HeatWave you can run your OLTP and OLAP, don’t ask me to ever say those acronyms again, workloads directly from your MySQL database and eliminate the time consuming data movement and integration work, while also performing 1100X faster than Amazon Aurora, and 2.5X faster than Amazon Redshift, at a third of the cost. My thanks again to Oracle Cloud for sponsoring this ridiculous nonsense.

Corey: The biggest question that doesn’t get asked that needs to be in almost every case is, “Okay. We’re building a thing, and it’s awesome. And I know it’s hard to think like this, but pivot around. Theoretically, what could a jerk do with it?”

Chloe: Yes.

Corey: When you’re designing it, it’s all right, how do you account for people that are complete jerks?

Chloe: Absolutely.

Corey: Even the cloud providers, all of them, when the whole Parler thing hit, everyone’s like, “Oh, Amazon is censoring people for freedom of speech.” No, they’re actually not. What they’re doing is enforcing their terms of service, the same terms of service that every provider that is not trash has. It is not a problem that one company decided they didn’t want hate speech on their platform. It was all the companies decided that, except for some very fringe elements. And that’s the sort of thing you have to figure out is, it’s easy in theory to figure out, oh, anything goes; freedom of speech. Great, well, some forms of speech violate federal law.

Chloe: Right.

Corey: So, what do you do then? Where do you draw the line? And it’s always nuanced and it’s always tricky, and the worst people are the folks that love to rules-lawyer around these things. It gets worse than that where these are the same people that will then sit there and make bad faith arguments all the time. And lawyers have a saying that hard cases make bad law.

When you have these very nuanced thing, and, “Well, we can’t just do it off the cuff. We have to build a policy around this.” This is the problem with most corporate policies across the board. It’s like, you don’t need a policy that says you’re not allowed to harass your colleagues with a stick. What you need to do is fire the jackwagon that made you think you might need a policy that said that.

But at scale, that becomes a super-hard thing to do when every enforcement action appears to be bespoke. Because there are elements on the gray areas and the margins where reasonable people can disagree. And that is what sets the policy and that’s where the precedent hits, and then you have these giant loopholes where people can basically be given free rein to be the worst humanity has to offer to some of the
most vulnerable members of our society.

Chloe: And I used to give this talk, I gave it at DockerCon one year and I gave it a couple other places, that was literally called “Diversity is not
Equal to Stock Images of Hands.” And the reason I say this is if you Google image search ‘diversity’ it’s like all of those clip arts of, like, Rainbow hands, things that you would see at Kaiser Permanente where it’s like, “We’re all in this together,” like, the pandemic, it’s all just hands on hands, hands as a Earth, hands as trees, hands as different colors. And people get really annoyed with people like me who are like, “Let’s shut up about diversity. Let’s just hire who’s best for the role.” Here’s the thing.

My favorite example of this—RIP—is Fleets—remember Fleets? [laugh]—on Twitter, so if they had one gay man in the room for that marketing, engineering—anything—decision, one of them I know would have piped up and said, “Hey, did you know ‘fleets’ is a commonly used term for douching enima in the gay community?” Now, I know that because I watch a lot of Ru Paul’s Drag Race, and I have worked with the gay community quite a bit in my time in theater. But this is what I mean about making sure. My friend Becca who works in security at safety and things, as well as Andy Tuba over at Reddit, I have a lot of conversations with my friend Becca Rosenthal about this, and that, not to quote Hamilton, but if I must, “We need people in the room where it happens.”

So, if you don’t have these people in the room if you’re a white man being like, “How will our products be abused?” Your guesses may be a little bit accurate but it was probably best to, at minimum, get some test case people in there from different genders, races, backgrounds, like, oh my goodness, get people in that room because what I tend to see is building safety tools, building even product features, or naming things, or designing things that could either be offensive, misused, whatever. So, when people have these arguments about like, “Diversity doesn’t matter. We’re hiring the best people.” I’m like, “Yeah, but your product’s going to be better, and more inclusive, and represent the people who use it at the end of the day because not everybody is you.”

And great examples of this include so many apps out there that exists that have one work location, one home location. How many people in the world have more than one job? That’s such a privileged view for us, as people in tech, that we can afford to just have one job. Or divorced parents or whatever that may be, for home location, and thinking through these edge cases and thinking through ways that your product can support everyone, if anything, by making your staff or the people that you work with more diverse, you’re going to be opening up your product to a much bigger marketable audience. So, I think people will look at me and be like, “Oh, Chloe’s a social justice warrior, she’s this feminist whatever,” but truly, I’m here saying, “You’re missing out on money, dude.” It would behoove you to do this at the end of the day because your users aren’t just a copy-paste of some dude in a Patagonia jacket with big headphones on. [laugh]. There are people beyond one demographic using your products and applications.

Corey: A consistent drag against Clubhouse since its inception was that it’s not an accessible app for a variety of reasons that were—

Chloe: It’s not an Android. [laugh].

Corey: Well, even ignoring the platform stuff, which I get—technical reasons, et cetera, yadda, yadda, great—there is no captioning option. And a lot of their abuse stuff in the early days was horrific, where you would get notifications that a lot of people had this person blocked, but… that’s not a helpful dynamic. “Did you talk to anyone? No, of course not. You Hacker News’ed it from first principles and thought this might be a good direction to go in.” This stuff is hard.

People specialize in this stuff, and I’ve always been an advocate of when you’re not sure what to do in an area, pay an expert for advice. All these stories about how people reach out to, “Their black friend”—and yes, it’s a singular person in many cases—and their black friend gets very tired of doing all the unpaid emotional labor of all of this stuff. Suddenly, it’s not that at all if you reach out to someone who is an expert in this and pay them for their expertise. I don’t sit here complaining that my clients pay me to solve AWS billing problems. In fact, I actively encourage that behavior. Same model.

There are businesses that specialize in this, they know the area, they know the risks, they know the ins and outs of this, and consults with these folks are not break the bank expensive compared to building the damn thing in the first place.

Chloe: And here’s a great example that literally drove me bananas a couple weeks ago. So, I don’t know if you’ve participated in Twitter Spaces before, but I’ve done a couple of my first ones recently. Have you done one yet—

Corey: Oh yes—

Chloe: —Corey?

Corey: —extensively. I love that. And again, that’s a better answer for me than Clubhouse because I already have the Twitter audience. I don’t have to build one from scratch on another platform.

Chloe: So, I learned something really fascinating through my boyfriend. And remember, I mentioned earlier, my boyfriend is a staff engineer at Uber. He’s been coding since he’s been out of the womb, much more experienced than me. And I like to think a lot about, this is accessible to
me but how is this accessible to a non-technical person? So, Ty finished up the Twitter Space that he did and he wanted to export the file.

Now currently, as the time of this podcast is being recorded, the process to export a Twitter Spaces audio file is a nightmare. And remember, staff engineer at Uber. He had to export his entire Twitter profile, navigate through a file structure that wasn’t clearly marked, find the recording out of the multiple Spaces that he had hosted—and I don’t think you get these for ones that you’ve participated in, only ones that you’ve hosted—download the file, but the file was not a normal WAV file or anything; he had to download an open-source converter to play the file. And in total, it took him about an hour to just get that file for the purposes of having that recording. Now, where my mind goes to is what about some woman who runs a nonprofit in the middle of, you know, Sacramento, and she does a community Twitter Spaces about her flower shop and she wants a recording of that.

What’s she going to do, hire some third-party? And she wouldn’t even know where to go; before I was in tech, I certainly would have just given up and been like, “Well, this is a nightmare. What do I do with this GitHub repo of information?” But these are the kinds of problems that you need to think about. And I think a lot of us and folks who listen to this show probably build APIs or developer tools, but a lot of us do work on products that muggles, non-technical people, work on.

And I see these issues happen constantly. I come from this space of being an admin, being someone who wasn’t quote-unquote, “A techie,” and a lot of products are just not being thought through from the perspective—like, there would be so much value gained if just one person came in and tested your product who wasn’t you. So yeah, there’s all of these things that I think we have a very privileged view of, as technical folks, that we don’t realize are huge. Not even just barrier to entry; you should just be able to download—and maybe this is a feature that’s coming down the pipeline soon, who knows, but the fact that in order for someone to get a recording of their Twitter Spaces is like a multi-hour process for a very, very senior engineer, that’s the problem. I’m not really sure how we solve this.

I think we just call it out when we see it and try to help different companies make change, which of course, myself and my boyfriend did. We reached out to people at Twitter, and we’re like, “This is really difficult and it shouldn’t be.” But I have that privilege. I know people at these companies; most people do not.

Corey: And in some cases, even when you do, it doesn’t move the needle as much as you might wish that it would.

Chloe: If it did, I wouldn’t be getting DMs anymore from creeps right? [laugh].

Corey: Right. Chloe, thank you so much for coming back and talk to me about your latest project. If people want to pay attention to it and see what you’re up to. Where can they go? Where can they find you? Where can they learn more? And where can they pointedly not audition to be
featured on one of the episodes of Master Creep Theatre?

Chloe: [laugh]. So, that’s the one caveat, right? I have to kind of close submissions of my own DMs now because now people are just going to be trolling me and sending me weird stuff. You can find me on Twitter—my name—at @chloecondon, C-H-L-O-E-C-O-N-D-O-N. I am on Instagram as @getforked, G-I-T-F-O-R-K-E-D. That’s a Good Placepun if you’re non-technical; it is an engineering pun if you are. And yeah, I’ve been doing a lot of fun series with Microsoft Reactor, lots of how to get a career in tech stuff for students, building a lot of really fun AI/ML stuff on there. So, come say hi on one of my many platforms. YouTube, too. That’s probably where—Master Creep Theatre is going to be, on YouTube, so definitely follow me on YouTube. And yeah.

Corey: And we will, of course, put links to that in the [show notes 00:37:57]. Chloe, thank you so much for taking the time to speak with me. I really appreciate it, as always.

Chloe: Thank you. I’ll be back for episode three soon, I’m sure. [laugh].

Corey: Let’s not make it another couple of years until then. Chloe Condon, senior cloud advocate at Microsoft on the Next Generation Experiences Team, also chlo-host of the Master Creep Theatre podcast. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with a comment saying simply, “Hey.”

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Nick

Nick Heudecker leads market strategy and competitive intelligence at Cribl, the observability pipeline company. Prior to Cribl, Nick spent eight years as an industry analyst at Gartner, covering data and analytics. Before that, he led engineering and product teams at multiple startups, with a bias towards open source software and adoption, and served as a cryptologist in the US Navy. Join Corey and Nick as they discuss the differences between observability and monitoring, why organizations struggle to get value from observability data, why observability requires new data management approaches, how observability pipelines are creating opportunities for SRE and SecOps teams, the balance between budgets and insight, why goats are the world’s best mammal, and more.

Links:

  • Cribl: https://cribl.io/
  • Cribl Community: https://cribl.io/community
  • Twitter: https://twitter.com/nheudecker
  • Try Cribl hosted solution: https://cribl.cloud

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Thinkst. This is going to take a minute to explain, so bear with me. I linked against an early version of their tool, canarytokens.org in the very early days of my newsletter, and what it does is relatively simple and straightforward. It winds up embedding credentials, files, that sort of thing in various parts of your environment, wherever you want to; it gives you fake AWS API credentials, for example. And the only thing that these things do is alert you whenever someone attempts to use those things. It’s an awesome approach. I’ve used something similar for years. Check them out. But wait, there’s more. They also have an enterprise option that you should be very much aware of canary.tools. You can take a look at this, but what it does is it provides an enterprise approach to drive these things throughout your entire environment. You can get a physical device that hangs out on your network and impersonates whatever you want to. When it gets Nmap scanned, or someone attempts to log into it, or access files on it, you get instant alerts. It’s awesome. If you don’t do something like this, you’re likely to find out that you’ve gotten breached, the hard way. Take a look at this. It’s one of those few things that I look at and say, “Wow, that is an amazing idea. I love it.” That’s canarytokens.org and canary.tools. The first one is free. The second one is enterprise-y. Take a look. I’m a big fan of this. More from them in the coming weeks.

Corey: This episode is sponsored in part by our friends at Jellyfish. So, you’re sitting in front of your office chair, bleary eyed, parked in front of a powerpoint and—oh my sweet feathery Jesus its the night before the board meeting, because of course it is! As you slot that crappy screenshot of traffic light colored excel tables into your deck, or sift through endless spreadsheets looking for just the right data set, have you ever wondered, why is it that sales and marketing get all this shiny, awesome analytics and inside tools? Whereas, engineering basically gets left with the dregs. Well, the founders of Jellyfish certainly did. That’s why they created the Jellyfish Engineering Management Platform, but don’t you dare call it JEMP! Designed to make it simple to analyze your engineering organization, Jellyfish ingests signals from your tech stack. Including JIRA, Git, and collaborative tools. Yes, depressing to think of those things as your tech stack but this is 2021. They use that to create a model that accurately reflects just how the breakdown of engineering work aligns with your wider business objectives. In other words, it translates from code into spreadsheet. When you have to explain what you’re doing from an engineering perspective to people whose primary IDE is Microsoft Powerpoint, consider Jellyfish. Thats Jellyfish.co and tell them Corey sent you! Watch for the wince, thats my favorite part.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. This promoted episode is a bit fun because I’m joined by someone that I have a fair bit in common with. Sure, I moonlight sometimes as an analyst because I don’t really seem to know what that means, and he spent significant amounts of time as a VP analyst at Gartner. But more importantly than that, a lot of the reason that I am the way that I am is that I spent almost a decade growing up in Maine, and in Maine, there’s not a lot to do other than sit inside for the nine months of winter every year and develop personality problems.

You’ve already seen what that looks like with me. Please welcome Nick Heudecker, who presumably will disprove that, but maybe not. He is currently a senior director of market strategy and competitive intelligence at Cribl. Nick, thanks for joining me.

Nick: Thanks for having me. Excited to be here.

Corey: So, let’s start at the very beginning. I like playing with people’s titles, and you certainly have a lofty one. ‘competitive intelligence’ feels an awful lot like jeopardy. What am I missing?

Nick: Well, I’m basically an internal analyst at the company. So, I spend a lot of time looking at the broader market, seeing what trends are happening out there; looking at what kind of thought leadership content that I can create to help people discover Cribl, get interested in the products and services that we offer. So, I’m mostly—you mentioned my time in Maine. I was a cryptologist in the Navy and I spent almost all of my time focused on what the bad guys do. And in this job, I focus on what our potential competitors do in the market. So, I’m very externally focused. Does that help? Does that explain it?

Corey: No, it absolutely does. I mean, you folks have been sponsoring our nonsense for which we thank you, but the biggest problem that I have with telling the story of Cribl was that originally—initially it was, from my perspective, “What is this hokey nonsense?” And then I learned and got an answer and then finish the sentence with, “And where can I buy it?” Because it seems that the big competitive threat that you have is something crappy that some rando sysadmin has cobbled together. And I say that as the rando sysadmin, who has cobbled a lot of things like that together. And it’s awful. I wasn’t aware you folks had direct competitors.

Nick: Today we don’t. There’s a couple that it might be emerging a little bit, but in general, no, it’s mostly us, and that’s what I analyze every day. Are there other emerging companies in the space? Are there open-source projects? But you’re right, most of the things that we compete against are DIY today. Absolutely.

Corey: In your previous role, which you were at for a very long time in tech terms—which in a lot of other cases is, “Okay, that doesn’t seem that long,” but seven and a half years is a respectable stint at a company. And you were at Gartner doing a number of analyst-like activities. Let’s start at the beginning because I assure you, I’m asking this purely for the audience and not because I don’t know the answer myself, but what exactly is the purpose of an analyst firm, of which Gartner is the most broadly known and, follow up, why do companies care what Gartner thinks?

Nick: Yeah. It’s a good question, one that I answer a lot. So, what is the purpose of an analyst firm? The purpose of an analyst firm is to get impartial information about something, whether that is supply chain technology, big data tech, human resource management technologies. And it’s often difficult if you’re an end-user and you’re interested in say, acquiring a new piece of technology, what really works well, what doesn’t.

And so the analyst firm because in the course of a given year, I would talk to nearly a thousand companies and both end-users and vendors as well as investors about what they’re doing, what challenges they’re having, and I would distill that down into 30-minute conversations with everyone else. And so we provided impartial information in aggregate to people who just wanted to help. And that’s the purpose of an analyst firm. Your second question, why do people care? Well, I didn’t get paid by vendors.

I got paid by the company that I worked for, and so I got to be Tron; I fought for the users. And because I talk to so many different companies in different geographies, in different industries, and I share that information with my colleagues, they shared with me, we had a very robust understanding of what’s actually happening in any technology market. And that’s uncommon kind of insight to really have in any kind of industry. So, that’s the purpose and that’s why people care.

Corey: It’s easy from the engineering perspective that I used to inhabit to make fun of it. It’s oh, it’s purely justification when you’re making a big decision, so if it goes sideways—because find me a technology project that doesn’t eventually go sideways—I want to be able to make sure that I’m not the one that catches heat for it because Gartner said it was good. They have an amazing credibility story going on there, and I used to have that very dismissive perspective. But the more I started talking to folks who are Gartner customers themselves and some of the analyst-style things that I do with a variety of different companies, it’s turned into, “No, no. They’re after insight.”

Because it turns out, from my perspective at least, the more that you are focused on building a product that solves a problem, you sort of lose touch with the broader market because the only people you’re really talking to are either in your space or have already acknowledged and been right there and become your customer and have been jaded to see things from your point of view. Getting a more objective viewpoint from an impartial third party does have value.

Nick: Absolutely. And I want you to succeed, I want you to be successful, I want to carry on a relationship with all the clients that I would speak with, and so one of the fun things I would always ask is, “Why are you asking me this question now?” Sometimes it would come in, they’d be very innocuous;, “Compare these databases,” or, “Compare these cloud services.” “Well, why are you asking?” And that’s when you get to, kind of like, the psychology of it.

“Oh, we just hired a new CIO and he or she hates vendor X, so we have to get rid of it.” “Well, all right. Let’s figure out how we solve this problem for you.” And so it wasn’t always just technology comparisons. Technology is easy, you write a check and you hope for the best.

But when you’re dealing with large teams and maybe a globally distributed company, it really comes down to culture, and personality, and all the harder factors. And so it was always—those were always the most fun and certainly the most challenging conversations to have.

Corey: One challenge that I find in this space is—in my narrow niche of the world where I focus on AWS bills, where things are extraordinarily yes or no, black or white, binary choices—that I talked to companies, like during the pandemic, and they were super happy that, “Oh, yeah. Our infrastructure has auto-scaling and it works super well.” And I look at the bill and the spend graph over time is so flat you could basically play a game of pool on top of it. And I don’t believe that I’m talking to people who are lying to me. I truly don’t believe that people make that decision, but what they believe versus what is evidenced in reality are not necessarily congruent. How do you disambiguate from the stories that people want to tell about themselves? And what they’re actually doing?

Nick: You have to unpack it. I think you have to ask a series of questions to figure out what their motivation is. Who else is on the call, as well? I would sometimes drop into a phone call and there would be a dozen people on the line. Those inquiry calls would go the worst because everyone wants to stake a claim, everyone wants to be heard, no one’s going to be honest with you or with anyone else on the call.

So, you typically need to have a pretty personal conversation about what does this person want to accomplish, what does the company want to accomplish, and what are the factors that are pushing against what those things are? It’s like a novel, right? You have a character, the character wants to achieve something, and there are multiple obstacles in that person’s way. And so by act five, ideally everything wraps up and it’s perfect. And so my job is to get the character out of the tree that is on fire and onto the beach where the person can relax.

So, you have to unpack a lot of different questions and answers to figure out, well, are they telling me what their boss wants to hear or are they really looking for help? Sometimes you’re successful, sometimes you’re not. Not everyone does want to be open and honest. In other cases, you would have a team show up to a call with maybe a junior engineer and they really just want you to tell them that the junior engineer’s architecture is not a good idea. And so you do a lot of couples therapy as well. I don’t know if this is really answering the question for you, but there are no easy answers. And people are defensive, they have biases, companies overall are risk-averse. I think you know this.

Corey: Oh, yeah.

Nick: And so it can be difficult to get to the bottom of what their real motivation is.

Corey: My approach has always been that if you want serious data, you go talk to Gartner. If you want [anec-data 00:09:48] and some understanding, well, maybe we can have that conversation, but they’re empowering different decisions at different levels, and that’s fine. To be clear, I do not consider Gartner to be a competitor to what I do in any respect. It turns out that I am not very good at drawing charts in varying shades of blue and positioning things just so with repeatable methodology, and they’re not particularly good at having cartoon animals as their mascot that they put into ridiculous situations. We each have our portion of the universe, and that’s working out reasonably well.

Nick: Well, and there’s also something to unpack there as well because I would say that people look at Gartner and they think they have a lot of data. To a certain degree they do, but a lot of it is not quantifiable data. If you look at a firm like IDC, they specialize in—like, they are a data house; that is what they do. And so their view of the world and how they advise their clients is different. So, even within analyst firms, there is differentiation in what approach they take, how consultative they might be with their clients, one versus another. So, there certainly are differences that you could find the more exposure you get into the industry.

Corey: For a while, I’ve been making a recurring joke that Route 53—Amazon’s managed DNS service—is in fact a database. And then at some point, I saw a post on Reddit where someone said, “Yeah, I see the joke and it’s great, but why should I actually not do this?” At which point I had to jump in and say, “Okay, look. Jokes are all well and good, but as soon as people start taking me seriously, it’s very much time to come clean.” Because I think that’s the only ethical and responsible thing to do in this ecosystem.

Similarly, there was another great joke once upon a time. It was an April Fool’s Day prank, and Google put out a paper about this thing they called MapReduce. Hilarious prank that Yahoo fell for hook, line, and sinker, and wound up building Hadoop out of it and we’re still paying the price for that, years later. You have a bit of a reputation from your time at Gartner as being—and I quote—“The man who killed Hadoop.” What happened there? What’s the story? And I appreciate your finally making clear to the rest of us that it was, in fact, a joke. What happened there?

Nick: Well, one of the pieces of research that Gartner puts out every year is this thing called a Hype Cycle. And we’ve all seen it, it looks like a roller coaster in profile; big mountain goes up really high and then comes down steeply, drops into a valley, and then—

Corey: ‘the trough of disillusionment,’ as I recall.

Nick: Yes, my favorite. And then plateaus out. And one of the profiles on that curve was Hadoop distributions. And after years of taking inquiry calls, and writing documents, and speaking with everybody about what they were doing, we realized that this really isn’t taking off like everyone thinks it is. Cluster sizes weren’t getting bigger, people were having a lot of challenges with the complexity, people couldn’t find skills to run it themselves if they wanted to.

And then the cloud providers came in and said, “Well, we’ll make a lot of this really simple for you, and we’ll get rid of HDFS,” which is—was a good idea, but it didn’t really scale well. I think that the challenge of having to acquire computers with compute storage and memory again, and again, and again, and again, just was not sustainable for the majority of enterprises. And so we flagged it as this will be obsolete before plateau. And at that point, we got a lot of hate mail, but it just seemed like the right decision to make, right? Once again, we’re Tron; we fight for the users.

And that seemed like the right advice and direction to provide to the end-users. And so didn’t make a lot of friends, but I think I was long-term right about what happened in the Hadoop space. Certainly, some fragments of it are left over and we’re still seeing—you know, Spark is going strong, there’s a lot of Hive still around, but Hadoop as this amalgamation of open-source projects, I think is effectively dead.

Corey: I sure hope you’re right. I think it has a long tail like most things that are there. Legacy is the condescending engineering term for ‘it makes money.’ You were at Gartner for almost eight years and then you left to go work at Cribl. What triggered that? What was it that made you decide, “This is great. I’ve been here a long time. I’ve obviously made it work for me. I’m going to go work at a startup that apparently, even though it recently raised a $200 million funding round”—congratulations on that, by the way—“It still apparently can’t afford to buy a vowel in its name.” That’s C-R-I-B-L because, of course, it is. Maybe another consonant, while you’re shopping. But okay, great. It’s oddly spelled, it is hard to explain in some cases, to folks who are not already feeling pain in that space. What was it that made you decide to sit up and, “All right, this is where I want to be?”

Nick: Well, I met the co-founders when I was an analyst. They were working at Splunk and oddly enough—this is going to be an interesting transition compared to the previous thing we talked about—they were working on Hunk, which was, let’s use HDFS to store Splunk data. Made a lot of sense, right? It could be much more cost-effective than high-cost infrastructure for Splunk. And so they told me about this; I was interested.

And so I met the co-founders and then I reconnected with them after they left and formed Cribl. And I thought the story was really cool because where they’re sitting is between sources and destinations of observability data. And they were solving a problem that all of my customers had, but they couldn’t resolve. They would try and build it themselves. They would look at—Kafka was a popular choice, but that had some challenges for observability data—works fantastically well for application data.

And they were just—had a very pragmatic view of the world that they were inhabiting and the problem that they were looking to solve. And it looked kind of like a no-brainer of a problem to solve. But when you double-click on it, when you really look down and say, “All right, what are the challenges with doing this?” They’re really insurmountable for a lot of organizations. So, even though they may try and take a DIY approach, they often run into trouble after just a few weeks because of all the protocols you have to support, all the different data formats, and all the destinations, and role-based access control, and everything else that goes along with it.

And so I really liked the team. I thought the product inhabited a unique space in the market—we’ve already talked about the lack of competitors in the space—and I just felt like the company was on a rocket ship—or is a rocket ship—that basically had unbounded success potential. And so when the opportunity arose to join the team and do a lot of the things I like doing as an analyst—examining the market, talking to people looking at competitive aspects—I jumped at it.

Corey: It’s nice when you see those opportunities that show up in front of you, and the stars sort of align. It’s like, this is not just something that I’m excited about and enthused about, but hey, they can use me. I can add something to where they’re going and help them get there better, faster, sooner, et cetera, et cetera.

Nick: When you’re an analyst, you look at dozens of companies a month and I’d never seen an opportunity that looked like that. Everything kind of looked the same. There’s a bunch of data integration companies, there’s a bunch of companies with Spark and things like that, but this company was unique; the product was unique, and no one was really recognizing the opportunity. So, it was just a great set of things that all happen at the same time.

Corey: It’s always fun to see stars align like that. So—

Nick: Yeah.

Corey: —help me understand in a way that can be articulated to folks who don’t have 15 years of grumpy sysadmin experience under their belts, what does Cribl do?

Nick: So, Cribl does a couple of things. Our flagship product is called LogStream, and the easiest way to describe that is as an abstraction between sources and destinations of data. And that doesn’t sound very interesting, but if you, from your sysadmin background, you’re always dealing with events, logs, now there’s traces, metrics are also hanging around—

Corey: Oh, and of course, the time is never synchronized with anything either, so it’s sort of a giant whodunit, mystery, where half the eyewitnesses lie.

Nick: Well, there’s that. There’s a lot of data silos. If you got an agent deployed on a system, it’s only going to talk to one destination platform. And you repeat this, maybe a dozen times per server, and you might have 100,000 or 200,000 servers, with all of these different agents running on it, each one locked into one destination. So, you might want to be able to mix and match that data; you can’t. You’re locked in.

One of the things LogStream does is it lets you do that exact mixing and matching. Another thing that this product does, that LogStream does, is it gives you ability to manage that data. And then what I mean by that is, you may want to reduce how much stuff you’re sending into a given platform because maybe that platform charges you by your daily ingest rates or some other kind of event-based charges. And so not all that data is valuable, so why pay to store it if it’s not going to be valuable? Just dump it or reduce the amount of volume that you’ve got in that payload, like a Windows XML log.

And so that’s another aspect that it allows you to do, better management of that stuff. You can redact sensitive fields, you can enrich the data with maybe, say, GeoIPs so you know what kind of data privacy laws you fall under and so on. And so, the story has always been, land the data in your destination platform first, then do all those things. Well, of course, because that’s how they charge you; they charge you based on daily ingest. And so now the story is, make those decisions upfront in one place without having to spread this logic all over, and then send the data where you want it to go.

So, that’s really, that’s the core product today, LogStream. We call ourselves an observability pipeline for observability data. The other thing we’ve got going on is this project called AppScope, and I think this is pretty cool. AppScope is a black box instrumentation tool that basically resides between the application runtime and the kernel and any shared libraries. And so it provides—without you having to go back and instrument code—it instruments the application for you based on every call that it makes and then can send that data through something like LogStream or to another destination.

So, you don’t have to go back and say, “Well, I’m going to try and find the source code for this 30-year old c++ application.” I can simply run AppScope against the process, and find out exactly what that application is doing for me, and then relay that information to some other destination.

Corey: This episode is sponsored in part by Liquibase. If you’re anything like me, you’ve screwed up the database part of a deployment so severely that you’ve been banned from touching every anything that remotely sounds like SQL, at at least three different companies. We’ve mostly got code deployments solved for, but when it comes to databases we basically rely on desperate hope, with a roll back plan of keeping our resumes up to date. It doesn’t have to be that way. Meet Liquibase. It is both an open source project and a commercial offering. Liquibase lets you track, modify, and automate database schema changes across almost any database, with guardrails to ensure you’ll still have a company left after you deploy the change. No matter where your database lives, Liquibase can help you solve your database deployment issues. Check them out today at liquibase.com. Offer does not apply to Route 53.

Corey: I have to ask because I love what you’re doing, don’t get me wrong. The counterargument that always comes up in this type of
conversation is, “Who in their right mind looks at the state of the industry today and says, ‘You know what we need? That’s right; another observability tool.’” what differentiates what you folks are building from a lot of the existing names in the space? And to be clear, a lot of the existing names in the space are treating observability simply as hipster monitoring. I’m not entirely sure they’re wrong, but that’s a different fight for a different time.

Nick: Yeah. I’m happy to come back and talk about that aspect of it, too. What’s different about what we’re doing is we don’t care where the data goes. We don’t have a dog in that fight. We want you to have better control over where it goes and what kind of shape it’s in when it gets there.

And so I’ll give an example. One of our customers wanted to deploy a new SIEM—Security Information Event Management—tool. But they didn’t want to have to deploy a couple hundred-thousand new agents to go along with it. They already had the data coming in from another agent, they just couldn’t get the data to it. So, they use LogStream to send that data to their new desired platform.

Worked great. They were able to go from zero to a brand new platform in just a couple days, versus fighting with rolling out agents and having to update them. Did they conflict with existing agents? How much performance did it impact on the servers, and so on? So, we don’t care about the destination. We like everybody. We’re agnostic when it comes to where that data goes. And—

Corey: Oh, it’s not about the destination. It’s about the journey. Everyone’s been saying it, but you’ve turned it into a product.

Nick: It’s very spiritual. So, we [laugh] send, we send your observability data on a spiritual [laugh] journey to its destination, and we can do quite a bit with it on the way.

Corey: So, you said you offered to go back as well and visit the, “Oh, it’s monitoring, but we’re going to call it observability because otherwise we get yelled out on Twitter by Charity Majors.” How do you view that?

Nick: Monitoring is the things you already know. Right? You know what questions you want to ask, you get an alert if something goes out of bounds or something goes from green to red. Think about monitoring as a data warehouse. You shape your data, you get it all in just the right condition so you can ask the same question over and over again, over different time domains.

That’s how I think about monitoring. It’s prepackaged, you know exactly what you want to do with it. Observability is more like a data lake. I have no idea what I’m going to do with this stuff. I think there’s going to be some signals in here that I can use, and I’m going to go explore that data.

So, if monitoring is your known knowns, observability is your unknown unknowns. So, an ideal observability solution gives you an opportunity to discover what those are. Once you discover them. Great. Now, you can talk about how to get them into your monitoring system. So, for me, it’s kind of a process of discovery.

Corey: Which makes an awful lot of sense. The problem I’ve always had with the monitoring approach is it falls into this terrible pattern of enumerate the badness. In other words, “Imagine all the ways that this system can fail,” and then build an alerting that lets you know when any of those things happen. And what happens next is inevitable to anyone who’s ever dealt with the tricksy devils known as computers, and what happens, of course, is that they find new ways to fail and you generally get to add to the list of things to check for, usually at two o’clock in the morning.

Nick: On a Sunday.

Corey: Oh, absolutely. It almost doesn’t matter when. The real problem is when these things happen, it’s, “What day, actually, is it?” And you have to check the calendar to figure out because your third time that week being woken up in the dead of night. It’s like an infant but less than endearing.

So, that has been the old school approach, and there’s unfortunately still an awful lot of, we’ll just call it nonsense, in the industry that still does exactly the same thing, except now they call it observability because—hearkening back to earlier in our conversation—there’s a certain point in the Gartner Hype Cycle that we are all existing within. What’s the deal with that?

Nick: Well, I think that there are a lot of entrenched interests in the monitoring space. And so I think you always see this when a new term comes around. Vendors will say, “All right, well, there’s a lot of confusion about this. Let me back-fit my product into this term so that I can continue to look like I’m on the leading edge and I’m not going to put any of my revenues in jeopardy.” I know, that’s a cynical view, but I’ve seen it over and over again.

And I think that’s unfortunate because there’s a real opportunity to have a better understanding of your systems, to better understand what’s happening in all the containers you’re deploying and not tearing down the way that you should, to better understand what’s happening in distributed systems. And it’s going to be a real missed opportunity if that is what happens. If we just call this ‘Monitoring 2.0’ it’s going to leave a lot of unrealized potential in the market.

Corey: The big problem that I’ve seen in a lot of different areas is—I’ll be direct—consolidation where you have a company that starts to do a thing—and that’s great—and then they start doing other things that are tied to it. And in turn, they start, I guess, gathering everything in the ecosystem. If you break down observability into various constituent parts, I—know, I know, the pillars thing is going to upset people; ignore that for now—and if you have an offering that’s weak in a particular area, okay, instead of building it organically into the product, or saying, “Yeah, that’s not what we do,” there’s an instinct to acquire a company or build that functionality out. And it turns out that we’re building what feels the lot to me like the SaaS equivalent of multifunction printers: they can print, they can scan, they can fax, and none of those three very well, so it winds up with something that dissatisfies everyone, rather than a best-of-breed solution that has a very clear and narrow starting and stopping point. How do you view that?

Nick: Well, what you’ve described is a compromise, right? A compromise is everyone can work and no one’s happy. And I think that’s the advantage of where LogStream comes in. The reality is best-of-breed. Most enterprises today have 30 or more different monitoring tools—call them observability tools if you want to—and you will never pry those tools from the dead hands of those sysadmins, DevOps engineers, SREs, et cetera.

They all integrate those tools into how they work and their processes. So, we’re living in a best-of-breed world. It’s like that in data and analytics—my former beat—and it’s like that in monitoring and observability. People really gravitate towards the tools they like, they gravitate towards the tools their friends are using. And so you need a way to be able to mix and match that stuff.

And just because I want to stay [laugh] on message, that’s really where the LogStream story kind of blends in because we do that; we allow you to mix and match all those different pieces.

Corey: Joke’s on you. I use Nagios and I have no friends. I’m not convinced those two things are entirely unrelated, but here we are. So here’s, I guess, the big burning question that a lot of folks—certainly not me, but other undefined folks, ‘lots of people are saying’—so you built something interesting that actually works. I want to be clear on this.

I have spoken to customers of yours. They swear by it instead of swearing at it, which happens with other companies. Awesome. You have traction, you’re moving forward, things are going great. Here’s $200 million is the next part of that story, and on some level, my immediate reaction—which does need updating, let’s be clear here—is like, all right.

I’m trying to build a product. I can see how I could spend a few million bucks. “Well, what can you do with I don’t know, 100 times that?” My easy answer is, “Something monstrous.” I don’t believe that is the case here. What is the growth plan? What are you doing that makes having that kind of a war chest a useful and valuable thing to have?

Nick: Well, if you speak with the co-founders—and they’ve been open about this—we view ourselves as a generational company. We’re not just building one product. We’ve been thinking about, how do we deliver on observability as this idea of discovery? What does that take? And it doesn’t mean that we’re going to be less agnostic to other destinations, we still think there’s an incredible amount of value there and that’s not going away, but we think there’s maybe an interim step that we build out, potentially this idea of an observability data lake where you can explore these environments.

Certainly, there’s other types of options in the space today. Most of them are SQL-based, which is interesting because the audience that uses monitoring and observability tools couldn’t care less about SQL right? They want search, they want regex, and so you’ve got to have the right tool for that audience. And so we’re thinking about what that looks like going forward. We’re doubling down on people.

Surprisingly, this is a very—like anything else in software, it is people-intensive. And so certainly those are other aspects that we’re exploring with the recent investment, but definitely, multiproduct company is our future and continued expansion.

Corey: Expansion is always a fun one. It’s the idea of, great, are you looking at going deeper into the areas you’re already active within, or is it more of a, “Ah, so we’ve solved the, effectively, log routing problem. That’s great. Let’s solve other problems, too.” Or is it more of a, I guess, a doubling down and focusing on what’s working? And again, that probably sounds judgmental in a way I don’t intend it to at all. I just have a hard time contextualizing that level of scale coming from a small company perspective the way that I do.

Nick: Yeah. Our plan is to focus more intently on the areas that we’re in. We have a huge basis of experience there. We don’t want to be all things to all people; that dilutes the message down to nothing, so we want to be very specific in the audiences we talk to, the problems we’re trying to solve, and how we try to solve them.

Corey: The problem I’ve always found with a lot of the acquisition, growth thrashing of—let me call it what I think it is: companies in decline trying to strain relevancy, it feels almost like a, “We don’t see a growth strategy. So, we’re going to try and acquire everything that hold still long enough, at some level, trying to add more revenue to the pile, but also thrashing in the sense of, okay. They’re going to teach us how to do things in creative, awesome ways,” but it never works out that way. When you have a 50,000 person company acquiring a 200 person company, invariably the bigger culture is going to dominate. And I don’t understand why that mistake seems to continually happen again, and again, and again.

And people think I’m effectively alluding to—or whenever the spoken word version of subtweeting is—a particular company or a particular acquisition. I’m absolutely not, there are probably 50 different companies listening right now who thinks, “Oh, God. He’s talking about us.” It’s the common repeating trend. What is that?

Nick: It’s hard to say. In some cases, these acquisitions might just be talent. “We need to know how to do X. They know how to do X. Let’s do it.” They may have very unique niche technology or software that another company thinks they can more broadly apply.

Also, some of these big companies, these may not be board-level or CEO-level decisions. A business unit might decide, “Oh, I like what that company is doing. I’m going to go acquire it.” And so it looks like MegaCorp bought TinyCorp, but it’s really, this tiny business unit within MegaCorp bought tiny company. The reality is often different from what it looks like on the outside.

So, that’s one way. Another is, you know, if they’re going to teach us to be more effective with tech or something like that, you’re never going to beat culture. You’re never going to be the existing culture. If it’s 50,000, against 200, obviously we know who wins there. And so I don’t
know if that’s realistic.

I don’t know if the big companies are genuine when they say that, but it could just be the messaging that they use to make people happy and hopefully retain as many of those new employees for as long as they can. Does that make sense?

Corey: No, it makes perfect sense. It’s the right answer. It does articulate what is happening there, and I think I keep falling prey to the same failure. And it’s hard. It’s pernicious, but companies are not monolithic entities.

There’s no one person at all of these companies each who is making these giant unilateral decisions. It’s always some product manager or some particular person who has a vision and a strategy in the department. It is not something that the company board is agreeing on every little decision that gets made. They’re distributed entities in many respects.

Nick: Absolutely. And that’s only getting more pervasive as companies get larger [laugh] through acquisition. So, you’re going to see more and more of that, and so it’s going to look like we’re going to put one label on it, one brand. Often, I think internally, that’s the exact opposite of what actually happened, how that decision got made.

Corey: Nick, I want to thank you for taking so much time to speak with me about what you’re up to over there, how your path has shaped, how you view the world, and also what Cribl does these days. If people want to learn more about what you’re up to, how you think about the world,
or even possibly going to work at Cribl which, having spoken to a number of people over there, I would endorse it. How do they find you?

Nick: Best place to find us is by joining our community: cribl.io/community, and Cribl is spelled C-R-I-B-L. You can certainly reach out there, we’ve got about 2300 people in our community Slack, so it’s a great group. You can also reach out to me on Twitter, I’m @nheudecker, N-H-E-U-D-E-C-K-E-R. Tell me what you thought of the episode; love to hear it. And then beyond that, you can also sign up for our free cloud tier at cribl.cloud. It’s a pretty generous one terabyte a day processing, so you can start to send data in and send it wherever you’d like to be.

Corey: To be clear, this free as in beer, not free as an AWS free tier?

Nick: This is free as in beer.

Corey: Excellent. Excellent.

Nick: I think I’m getting that right. I think it’s free as in beer. And the other thing you can try is our hosted solution on AWS, fully managed cloud at cribl.cloud, we offer a free one terabyte per day processing, so you can start to send data into that environment and send it wherever you’d like to go, in whatever shape that data needs to be in when it gets there.

Corey: And we will, of course, put links to that in the [show notes 00:35:21]. Thank you so much for your time today. I really appreciate it.

Nick: No, thank you for having me. This was a lot of fun.

Corey: Nick Heudecker, senior director, market strategy and competitive intelligence at Cribl. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with a comment explaining that the only real reason a startup should raise a $200 million funding round is to pay that month’s AWS bill.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Abby

With over twenty years in the tech world, Abby Kearns is a true veteran of the technology industry. Her lengthy career has spanned product marketing, product management and consulting across Fortune 500 companies and startups alike. At Puppet, she leads the vision and direction of the current and future enterprise product portfolio. Prior to joining Puppet, Abby was the CEO of the Cloud Foundry Foundation where she focused on driving the vision for the Foundation as well as growing the open source project and ecosystem. Her background also includes product management at companies such as Pivotal and Verizon, as well as infrastructure operations spanning companies such as Totality, EDS, and Sabre.

Links:

  • Cloud Foundry Foundation: https://www.cloudfoundry.org
  • Puppet: https://puppet.com
  • Twitter: https://twitter.com/ab415

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Liquibase. If you’re anything like me, you’ve screwed up the database part of a deployment so severely that you’ve been banned from touching every anything that remotely sounds like SQL, at at least three different companies. We’ve mostly got code deployments solved for, but when it comes to databases we basically rely on desperate hope, with a roll back plan of keeping our resumes up to date. It doesn’t have to be that way. Meet Liquibase. It is both an open source project and a commercial offering. Liquibase lets you track, modify, and automate database schema changes across almost any database, with guardrails to ensure you’ll still have a company left after you deploy the change. No matter where your database lives, Liquibase can help you solve your database deployment issues. Check them out today at liquibase.com. Offer does not apply to Route 53.

Corey: This episode is sponsored in part by Honeycomb. When production is running slow, it's hard to know where problems originate: is it your application code, users, or the underlying systems? I’ve got five bucks on DNS, personally. Why scroll through endless dashboards, while dealing with alert floods, going from tool to tool to tool that you employ, guessing at which puzzle pieces matter? Context switching and tool sprawl are slowly killing both your team and your business. You should care more about one of those than the other, which one is up to you. Drop the separate pillars and enter a world of getting one unified understanding of the one thing driving your business: production. With Honeycomb, you guess less and know more. Try it for free at Honeycomb.io/screaminginthecloud. Observability, it’s more than just hipster monitoring.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Once upon a time, I was deep into the weeds of configuration management, which explains a lot, such as why it seems I don’t know happiness in any meaningful sense. Then I wound up progressing into other areas of exploration, like the cloud, and now we know for a fact why happiness isn’t a thing for me. My guest today is the former CEO of the Cloud Foundry Foundation and today is the CTO over at a company called Puppet, which we’ve talked about here from time to time. Abby Kearns,
thank you for joining me. I appreciate your taking the time out of your day to suffer my slings and arrows.

Abby: Thank you for having me. I have been looking forward to this for weeks.

Corey: My stars, it seems like things are slow over there, and I kind of envy you for that. So, help me understand something; you went from this world of cloud-native everything, which is the joy of working with Cloud Foundry, to now working with configuration management. How is that not effectively Benjamin Button-ing your career. It feels like the opposite direction that most quote-unquote, “Digital transformations” like to play with. But I have a sneaking suspicion, there’s more to it than I might guess from just looking at the label on the tin.

Abby: Beyond I just love enterprise infrastructure? I mean, come on, who doesn’t?

Corey: Oh, yeah. Everyone loves to talk about digital transformation, reading about books like a Head in the Cloud to my children used to be a fun nightly activity before it was formally classified as child abuse. So yeah, I hear you, but it turns out the rest of the world doesn’t necessarily agree with us.

Abby: I do not understand it. I have been in enterprise infrastructure my entire career, which has been a really, really long time, back when Unix and Sun machines were still a thing. And I’ll be a little biased here; I think that enterprise infrastructure is actually the most fascinating part of technology right now. And why is that? Well, we’re in the process of actively rewritten everything that got us here.

And we talk about infrastructure and everyone’s like, “Yeah, sure, whatever,” but at the end of the day, it’s the foundation that everything that you think is cool about technology is built on. And for those of us that really enjoy this space, having a front-row seat at that evolution and the innovation that’s happening is really, really exciting and it creates a lot of interesting conversation, debate, evolution of technologies, and innovation. And are they all going to be on the money five, ten years from now? Maybe not, but they’re creating an interesting space and discussion and just the work ahead for all of us across the board. And I’m kind of bucketing this pretty broadly, intentionally so because I think at the end of the day, all of us play a role in a bigger piece of pie, and it’s so interesting to see how these things start to fit together.

Corey: One of the things that I’ve noticed is that the things that get attention on the keynote stage of, “This is this far future, serverless, machine-learning Kubernetes, dingus nonsense,” great is—

Abby: You forgot blockchain. [laugh].

Corey: Oh, yeah. Oh, yeah blockchain as well. Like, what other things can we wind up putting into the buzzword thing to wind up guaranteeing that your seed round is at least $200 million? Great. There’s that.

But when you look at the actual AWS bill—my specialty, of course—and seeing where the money is actually going, it doesn’t really look that different, as far as percentages go—even though the numbers are higher—than it did ten years ago, at least in the enterprise world. You’re still buying a bunch of EC2 instances, you’re still potentially modernizing to some of the managed services like RDS—which is Amazon’s reimagining of what a database could be if you still had to manage the finicky bits, but had no control over when and how they worked—and of course, data transfer and disk. These are the basic building blocks of everything in cloud. And despite how much we talk about the super neat stuff, what we’re doing is not reflected on the conference stage. So, I tend to view the idea of aspirational architecture as its own little
world.

There are still seasoned companies out there that are migrating from where they are today into this idea of, well, virtualization, we’ve just finally got our heads around that. Now, let’s talk about this cloud thing; seems like a fad—in 2021. And people take longer to get to where they think they’re going or where they intend to go than they plan for, and they get stuck somewhere and instead of a cloud migration, they’re now hybrid because they can redefine things and declare victory when they plant that flag, and here we are. I’m not here to make fun of these companies because they’re doing important work and these are super hard problems. But increasingly, it seems that the technology is not the thing that’s holding them back or even responsible for their outcome so much as it is people.

The more I work with tech, the more I realized that everything that’s hard becomes people issues. Curious to get your take on that, given your somewhat privileged perspective as having a foot standing very deeply in each world.

Abby: Yeah, and that’s a super great point. And I also realized I didn’t fully answer the first question either. So, I’ll tie those two things together.

Corey: That’s okay, we’re going to keep circling around until you get there. It’s fine.

Abby: It’s been a long week, and it’s only Wednesday.

Corey: All day long, as it turns out.

Abby: I have a whole soapbox that I drag around behind me about people and process, and how that’s your biggest problem, not technology, and if you don’t solve for the people in the process, I don’t care what technology you choose to use, isn’t going to fix your problem. On the other hand, if you get your people and process right, you can borderline use crayons and paper and get [laugh] really close to what you need to solve for.

Corey: I have it on good authority that’s known as IBM Cloud. Please continue.

Abby: [laugh]. And so I think people and process are at the heart of everything. They’re our biggest accelerators with technology and they’re our biggest limitation. And you can cloud-native serverless your way into it, but if you do not actually do continuous delivery, if you did not actually automate your responses, if you do not actually set up the cross-functional teams—or sometimes fondly referred to as two-pizza teams—if you don’t have those things set up, there isn’t any technology that’s going to make you deliver software better, faster, cheaper. And so I think I care a lot about the focus on that because I do think it is so important, but it’s also—the reason a lot of people don’t like to talk about it and deal with it because it’s also the hardest.

People, culture change, digital transformation, whatever you want to call it, is hard work. There’s a reason so many books are written around DevOps. And you mentioned Gene Kim earlier, there’s a reason he wrote The Phoenix Project; it’s the people-process part is the hardest. And I do think technology should be an enabler and an accelerator, but it really has to pair up nicely with the people part. And you asked your earlier question about my move to Puppet.

One of the things that I’ve learned a lot in running the Cloud Foundry Foundation, running an open-source software foundation, is you could a real good crash course in how teams can collaborate effectively, how teams work together, how decisions get made, the need for that process and that practice. And there was a lot of great context because I had access to so much interesting information. I got to see what all of these large enterprises were doing across the board. And I got to have a literal seat at the table for how a lot of the decisions are getting made around not only the open-source technologies that are going into building the future of our enterprise infrastructure but how a lot of these companies are using and leveraging those technologies. And having that visibility was amazing and transformational for myself.

It gave me so much richness and context, which is why I have firmly believed that the people and process part were so crucial for many years. And I decided to go to a company that sold products. [laugh]. You’re like, “What? What is she talking about now? Where is this going?”

And I say that because running an open-source software foundation is great and it gives you so much information and so much context, but you have no access to customers and no access to products. You have no influence over that. And so when I thought about what I wanted to do next, it’s like, I really want to be close to customers, I really want to be close to product, and I really want to be part of something that’s solving what I look at over the next five to ten years, our biggest problem area, which is that tweener phase that we’re going to be in for many years, which we were just talking about, which is, “I have some stuff on-prem and I have some stuff in a cloud—usually more than one cloud—and I got to figure out how to manage all of that.” And that is a really, really, really hard problem. And so when I looked at what Puppet was trying to do, and the opportunity that existed with a lot of the fantastic work that Puppet has done over the last 12 years around Desired State Configuration management, I’m like, “Okay, there’s something here.”

Because clearly, that problem doesn’t go away because I’m running some stuff in the cloud. So, how do we start to think about this more broadly and expansively across the hybrid estate that is all of these different environments? And who is the most well-positioned to actually drive an innovative product that addresses that? So, that’s my long way of addressing both of those things.

Corey: No, it’s a fair question. Friend of the show, Matt Stratton, is famous for saying that, “You cannot buy DevOps, but I sure would like to sell it to you,” and if you’re looking at it from that perspective, Puppet is not far from what that product store look like in some ways. My first encounter with Puppet was back around 2009, 2010 or so, and I was using it in an environment I was working within and thought, “Okay, this is terrible, and it’s crap, and obviously, I know what I’m doing far better than this, and the problem is the Puppet’s a bad product.” So, I was one of the early developers behind SaltStack, which was a terrific, great way of approaching the problem from a novel perspective, and it wasn’t crap; it was awesome. Right up until I saw the first time a customer deployed it and looked at their environment, and it wasn’t crap, it was worse because it turns out that you can build a super finely crafted precision instrument that makes a fairly bad hammer, but that’s how customers are going to use it anyway.

Abby: Well, I mean, [sigh] look, you actually hit something that I think we don’t actually talk about, which is how hard all of this shit really is. Automation is hard. Automation for distributed systems at scale is super duper hard. There isn’t an easy way to solve that problem. And I feel like I learned a lot working with Cloud Foundry.

Cloud Foundry is a Platform as a Service and it sits a layer up, but it had the same challenges in that solving the ability to run cloud-native applications and cloud-native workloads at scale and have that ephemerality to it and that resilience to it, and the things everyone wants but don’t recognize how difficult it is, actually, to do that well. And I think the same—you know, that really set me up for the way that I think about the problem, even the layer down which is, running and managing desired state, which at the end of the day is a really fancy way of saying, “Does your environment look like the way you think it should? And if it doesn’t, what are you going to do about it?” And it seems like, in this year of—what year are we again? 2021, maybe? I don’t know. It feels like the last two years of, sort of, munged together?

Corey: Yeah, the passing of time is something it’s very hard for me to wrap my head around.

Abby: But it feels like, I know some people, particularly those of us that have been in tech a long time are probably like, “Why are we still talking about that? Why is that a thing?” But that is still an incredibly hard problem for most organizations, large and small. So, I tend to spend a lot of time thinking about large enterprises, but in the day, you’ve got more than 20 servers, you’re probably sitting around thinking, “Does my environment actually look the way I think it does? There’s a new CVE that just came out. Am I able to address that?”

And I think at the end of the day, figuring out how you can solve for that on-prem has been one of the things that Puppet has worked for, and done really, really well the last 12 years. Now, I think the next challenge is okay, how do you extend that out across your now bananas complex estate that is—I got a huge data estate, maybe one or two data centers, I got some stuff in AWS, I got some stuff in GCP, oh yeah, got a little thing over here and Azure, and oh, some guy spun up something on OCI. So, we got a little bit of everything. And oh, my God, the SolarWinds breach happened. Are we impacted? I don’t know. What does that mean? [laugh].

And I think you start to unravel the little pieces of that and it gets more and more complex. And so I think the problems that I was solving in the early aughts with servers seems trite now because you’re like, I can see all of my servers; there’s eight of them. Things seem fine. To now, you’ve got hundreds of thousands of applications and workloads, and some of them are serverless, and they’re all over the place. And who has what, and where does it sit?

And does it look like the way that I think it needs to so that I can run my business effectively? And I think that’s really the power of it, but it’s also one of those things that I don’t feel like a lot of people like to acknowledge the complexity and the hardness of that because it’s not just the technology problem—going back to your other question, how do we work? How do we communicate? What are our processes around dealing with this? And I think there’s so much wrapped up in that it becomes almost like, how do you eat an elephant story, right? Yes, one bite at a time, but when you first look at the elephant, you’re like, “Holy shit. This is big. What do I need to do?” And that I think is not something we all collectively spend enough time talking about is how hard this stuff is.

Corey: One of the biggest challenges I see across the board is this idea of conference-ware style architecture; the greatest lie you ever see is someone talking about their infrastructure in public because peel it back a little bit and everything’s messy, everything’s disastrous, and everything’s a tire fire. And we have this cult in tech—

Abby: [laugh].

Corey: —it’s almost a cult where we have this idea that anything that isn’t rewritten completely within the last six months based upon whatever is the hot framework now that is designed to run only in Google Chrome running on the latest generation MacBook Pro on a gigabit internet connection is somehow less than. It’s like, “So, what does that piece of crap do?” And the answer is, “Well, a few $100 million a quarter in revenue, so how about you watch your mouth?” Moving those things is delicate; moving those things is fraught, and there are a lot of different stakeholders to the point where one of the lessons I keep learning is, people love to ask me, “What is Amazon’s opinion of you?” Turns out that there’s no Ted Amazon who works over there who forms a single entity’s opinion. It’s a bunch of small teams. Some of them like me, some of them can’t stand me, far and away the majority don’t know who I am. And that is okay. In theory; in practice, I find it completely unforgivable because how dare you? But I understand it’s—

Abby: You write a memo, right now. [laugh].

Corey: Exactly. Companies are people and people are messy, and for better or worse, it is impossible to patch them. So, you have to almost route around them. And that was something that I found that Puppet did very well, coming from the olden days of sysadmin work where we spend time doing management [bump 00:15:53] the systems by hand. Like, oh, I’m going to do a for loop. Once I learned how to script. Before that, I use Cluster SSH and inadvertently blew away a University’s entire config file what starts up on boot across their entire FreeBSD server fleet.

Abby: You only did it once, so it’s fine.

Corey: Oh, yeah. I’m never going to screw up again. Well, not like that. In other ways. Absolutely, but at least my errors will be novel.

Abby: Yeah. It’s learning. We all learn. If you haven’t taken something down in production in real-time, you have not lived. And also you [laugh] haven’t done tech. [laugh].

Corey: Oh, yeah, you either haven’t been allowed close enough to anything that’s important enough to be able to take down, you’re lying to me, or thirdly—and this is possible, too—you’re not yet at a point in your career where you’re allowed to have access to the breaky parts. And that’s fine. I mean, my argument has always been about why I’d be a terrible employee at Google, for example, is if I went in maliciously on day one, I would be hard-pressed to take down google.com for one hour. If I can’t have that much impact intentionally going in as a bad actor, it feels like there’d be how much possible upside, positive impact can I have what everyone’s ostensibly aligned around the same thing?

It’s the challenge of big companies. It’s gaining buy-in, it’s gaining investment in the idea and the direction you’re going in. Things always take longer, you have to wind up getting multiple stakeholders on board. My consulting practice is entirely around helping save money on the AWS bill. You’d think it would be the easiest thing in the world to sell, but talking to big companies means a series of different sales conversations with different folks, getting them all on the same page. What we do functionally isn’t so much look at the computer parts as it is marriage counseling between engineering and finance. Different languages, different ways of thinking about things, ostensibly the same goals.

Abby: I mean, I don’t think that’s a big company problem. I think that’s an every company problem if you have more than, like, five people in your company.

Corey: The first few years here, it was just me and I had none of those problems. I had very different problems, but you know—and then we started bringing other people in, it’s like, “Oh, yeah, things were great until we hired people. Ugh, mistake. Never do that.” And yeah, it turns out that’s not particularly sustainable.

Abby: Stakeholder management is hard. And you mentioned something about routing around. Well, you can’t actually route around people, unfortunately. You have to get people to buy in, you have to bring people along on the journey. And not everybody is at the same place in the way they think about the work you’re doing.

And that’s true at any company, big or small. I think it just gets harder and more complex as the company gets bigger because it’s harder to make the changes you need to make fast enough, but I’d say even at a company the size of Puppet, we have the exact same challenges. You know, are the teams aligned? Are we aligned on the right things? Are we focusing on the right things?

Or, do we have the right priorities in our backlog? How are we doing the work that we do? And if you’re trying to drive innovation, how fast are we innovating? Are we innovating fast enough? How tight are our feedback loops?

It’s one of those things where the conversations that you and I have had externally with customers are the same conversations I have internally all the time, too. Let’s talk about innovators’ dilemma. [laugh]. Let’s talk about feedback loop. Let’s talk about what does it mean to get tighter feedback loops from customers and the field?

And how do you align those things to the priorities in your backlog? And it’s one of those never-ending challenges that’s messy and complicated. And technology can enable it, but the technology is also messy and hard. And I do love going to conferences and seeing how pretty and easy things could look, and it’s definitely a great aspiration for us to all shoot for, but at the end of the day, I think we all have to recognize there’s a ton of messiness that goes on behind to make that a reality and to make that really a product and a technology that we can sell and get behind, but also one that we buy in, too, and are able to use. So, I think we as a technology industry, and particularly those of us in the Bay Area, we do a disservice by talking about how easy things are and why—you know, I remember a conversation I had in 2014 where someone asked me if Docker was already passe because everybody was doing containerized applications, and I was like, “Are they? Really? Is that an everyone thing? Or is that just an ‘us’ thing?” [laugh].

Corey: Well, they talk about it on the conference stages an awful lot, but yeah. New problems that continue to arise. I mean, I look back at my early formative years as someone who could theoretically be brought out in public and it was through a consulting project, where I was a traveling trainer for Puppet back in 2014, 2015, and teaching people who hadn’t had exposure before what Puppet was about. And there was a definite experience in some of the people attending class where they were very opposed to the idea. And dig down a little bit, it’s not that they had a problem with the software, it’s not that they had a problem with any of the technical bits.

It’s that they made the mistake that so many technologists made—I know I have, repeatedly—of identifying themselves with the technology that they work on. And well, in some cases, yeah, the answer was that they ran a particular script a bunch of times and if you can automate that through something like Puppet or something else, well, what does that mean for them? We see it much larger-scale now with people who are, okay, I’m in the data center working on the storage arrays. When that becomes just an API call or—let’s be serious, despite what we see in conference stages—when it becomes clicking buttons in the AWS console, then what does that mean for the future of their career? The tide is rising.

And I can’t blame them too much for this; you’ve been doing this for 25 years, you don’t necessarily want to throw all that away and start over with a whole new set of concepts and the rest because unlike what Twitter believes, there are a bunch of legitimate paths in this industry that do treat it as a job rather than an all-consuming passion. And I have no negative judgment toward folks who walk down that direction.

Abby: Most people do. And I think we have to be realistic. It’s not just some. A lot of people do. A lot of people, “This is my nine-to-five job, Monday through Friday, and I’m going to go home and I’m going to spend time with my family.”

Or I’m going to dare I say—quietly—have a life outside of technology. You know, but this is my job. And I think we have done a disservice to a lot of those individuals who for better or for worse, they just want to go in and do a job. They want to get their job done to the best of their abilities, and don’t necessarily have the time—or if you’re a single parent, have the flexibility in your day to go home and spend another five, six hours learning the latest technology, the latest programming language, set up your own demo environment at home, play around with AWS, all of these things that you may not have the opportunity to do. And I think we as an industry have done a disservice to both those individuals, as well in putting up really imaginary gates on who can actually be a technologist, too.

Corey: This episode is sponsored by our friends at Oracle Cloud. Counting the pennies, but still dreaming of deploying apps instead of "Hello, World" demos? Allow me to introduce you to Oracle's Always Free tier. It provides over 20 free services and infrastructure, networking databases, observability, management, and security.

And - let me be clear here - it's actually free. There's no surprise billing until you intentionally and proactively upgrade your account. This means you can provision a virtual machine instance or spin up an autonomous database that manages itself all while gaining the networking load, balancing and storage resources that somehow never quite make it into most free tiers needed to support the application that you want to build.

With Always Free you can do things like run small scale applications, or do proof of concept testing without spending a dime. You know that I always like to put asterisks next to the word free. This is actually free. No asterisk. Start now. Visit https://snark.cloud/oci-free that's https://snark.cloud/oci-free.

Corey: Gatekeeping, on some level, is just—it’s a horrible thing. Something I found relatively early on is that I didn’t enjoy communities where that was a thing in a big way. In minor ways, sure, absolutely. I wound up gravitating toward Ubuntu rather than Debian because it turned out that being actively insulted when I asked how to do something wasn’t exactly the most welcoming, constructive experience, where they, “Read the manual.” “Yeah, I did that and it was incomplete and contradictory, and that’s why I’m here asking you that question, but please continue to be a condescending jackwagon. I appreciate that. It really just reminds me that I’m making good choices with my life.”

Abby: Hashtag-RTFM. [laugh].

Corey: Exactly. In my case, fine, its water off a duck’s back. I can certainly take it given the way that I dish it out, but by the same token, not everyone has a quote-unquote, thick skin, and I further posit that not everyone should have to have one. You should not get used to personal attacks as a prerequisite for working in this space. And I’m very sensitive to the idea that people who are just now exploring the cloud somehow feel that they’ve missed out on their career, and that so there’s somehow not appropriate for this field, or that it’s not for them.

And no, are you kidding me? You know that overwhelming sense of confusion you get when you look at the AWS console and try and understand what all those services do? Yeah, I had the same impression the first time I saw it and there were 12 services; there’s over 200 now. Guess what? I’ve still got it.

And if I am overwhelmed by it, I promise there’s no shame in anyone else being overwhelmed by it, too. We’re long since past the point where I can talk incredibly convincingly about AWS services that don’t exist to AWS employees and not get called out on it because who in the world has that entire Rolodex of services shoved into their heads who isn’t me?

Abby: I’d say you should put out… a call for anyone that does because I certainly do not memorize the services that are available. I don’t know that anyone does. And I think even more broadly, is, remember when the landscape diagram came out from the CNCF a couple of years ago, which it’s now, like… it’s like a NASCAR logo of every logo known to man—

Corey: Oh today, there’s over 400 icons on it the last time I saw—I saw that thing come out and I realized, “Wow, I thought I was going to shit-posting,” but no, this thing is incredible. It’s, “This is great.” My personal favorite was zooming all the way in finding a couple of logos on in the same box three times, which is just… spot on. I was told later, it’s like, “Oh, those represent different projects.” I’m like, “Oh, yeah, must have missed that in the legend somewhere.” [laugh]. It’s this monstrous, overdone thing.

Abby: But the whole point of it was just, if I am running an IT department, and I’m like, “Here you go. Here’s a menu of things to choose,” you’re just like, “What do I do with this information? Do I choose one of each? All the above? Where do I go? And then, frankly, how do I make them all work together in my environment?” Because they all serve very different problems and they’re tackling different aspects of that problem.

And I think I get really annoyed with myself as an industry—like, ourselves as an industry because it’s like, “What are we doing here?” We’re trying to make it harder for people, not only to use the technology, to be part of it. And I think any efforts we can make to make it easier and more simple or clear, we owe it to ourselves to be able to tell that story. Which now the flip side of that is describing cloud-native in the cloud, and infrastructure and automation is really, really hard to do [laugh] in a way that doesn’t use any of those words. And I’m just as guilty of this, of describing things we do and using the same language, and all of a sudden you’re looking at it this says the same thing is 7500 other websites. [laugh]. So.

Corey: Yep. I joke at RSA’s Expo Hall is basically about twelve companies selling different things. Sure, each one has a whole bunch of booths with different logos and different marketing copy, but it’s the same fundamental product. Same challenge here. And this is, to me, the future of cloud, this is where it’s going, where I want something that will—in my case, I built a custom URL shortener out of DynamoDB, API Gateway, Lambda, et cetera, and I built this thing largely as a proof of concept because I wanted to have experience playing with these tools.

And that was great, not but if I’m doing something like that in production, I’m going with Bitly or one of the other services that provide this where someone is going to maintain it full time. Unless it is the core of what I’m doing, I don’t want to build it myself from popsicle sticks. And moving up the stack to a world of folks who are trying to solve a business problem and they don’t want to deal with the ten prerequisite services to understand the cloud, and then a whole bunch of other things tied together, and the billing, and the flow becomes incredibly problematic to understand—not to mention insecure: because we don’t understand it, you don’t know what your risk exposure is—people don’t want that. They—

Abby: Or to manage it.

Corey: Yeah.

Abby: Just the day-to-day management. Care and feeding, beyond security. [laugh].

Corey: People’s time is free. So, yeah. For example, do I write my own payroll system? Absolutely not. I have the good sense to pay a turnkey company to handle that for me because mistakes will show.

I started my career running email systems. I pay for Google workspaces—or GSuite, or Gmail, or whatever the hell they’re calling it this week—because it’s not core and central to my business. I want a thing that winds up solving a business problem, and I will pay commensurately to the value that thing delivers, not the individual constituent costs of the components that build it together. Because until you’re significantly scaled out and it is the core of what you do, you’re spending more on people to run the monstrous thing than you are for the thing itself. That’s always the way it works.

So, put your innovation where it matters for your business. I posit the for an awful lot of the things we’re building, in order to achieve those outcomes, this isn’t it.

Abby: Agreed. And I am a big believer in if I can use off-the-shelf software, I will because I don’t believe in reinventing everything. Now, having said that, and coming off my soapbox for just a hot minute, I will say that a lot of what’s happening, and going back to where I started around the enterprise infrastructure, we’re reinventing so many things that there is a lot of new things coming up. We’ve talked about containers, we’ve talked about Kubernetes, around container scheduling, container orchestration, we haven’t even mentioned service mesh, and sidecars, and all of the new ways we’re approaching solving some of these older problems. So, there is the need for a broad proliferation of technology until the contraction phase, where it all starts to fundamentally clicks together.

And that’s really where the interesting parts happen, but it’s also where the confusion happens because, “Okay, what do I use? How do I use it? How do these pieces fit together? What happens when this changes? What does this mean?”

And by the way, if I’m an enterprise company, I’m a payroll company, what’s the one thing I care about? My payroll software. [laugh]. And that’s the problem I’m solving for. So, I take a little umbrage sometimes with the frame that every company is a software company because every company is not a software company.

Every company can use technology in ways to further their business and more and more frequently, that is delivering their business value through software, but if I’m a payroll company, I care about delivering that payroll capabilities to my customer, and I want to do it as quickly as possible, and I want to leverage technology to help me do that. But my endgame is not that technology; my endgame is delivering value to my customers in real and meaningful ways. And I worry, sometimes, that those two things get conflated together. And one is an enabler of the other; the technology is not the outcome.

Corey: And that is borderline heresy for an awful lot of folks out there in the space, I wish that people would wake up a little bit more and realize that you have to build a thing that solves customer pain, ideally, an expensive customer pain, and then they will basically rush to hurl money at you. Now, there are challenges and inflections as you go, and there’s a whole bunch of nuances that can span entire fields of endeavor that I am hand-waving over here, and that’s fine, but this is the direction I think we’re going and this is the dawning awareness that I hope and trust we’ll see start to take root in this industry.

Abby: I mean, I hope so. I do take comfort in the fact that a lot of the industry leaders I’m starting to see, kind of, equate those two things more closely in the top [track 00:31:20]. Because it’s a good forcing function for those of us that are technologists. At the end of the day, what am I doing? I am a product company, I am selling software to someone.

So clearly, obviously, I have a vested interest in building the best software out there, but at the end of the day, for me, it’s, “Okay, how do I make that truly impactful for customers, and how do I help them solve a problem?” And for me, I’m hyper-focused on automation because I honestly feel like that is the biggest challenge for most companies; it’s the hardest thing to solve. It’s like getting into your auto-driving car for the first time and letting go the steering wheel and praying to the software gods that that software is actually going to work. But it’s the same thing with automation; it’s like, “Okay, I have to trust that this is going to manage my environment and manage my infrastructure in a factual way and not put me on CNN because I just shut down entire customer environment,” or if I’m an airline and I’ve just had a really bad week because I’ve had technology problems. [laugh]. And so I think we have to really take into consideration that there are real customer problems on the other end of that we have to help solve for.

Corey: My biggest problem is the failure mode of this is not when people watch the conference-ware presentations is that they’re not going to sit there and think, “Oh, yeah, they’re just talking about a nuanced thing that doesn’t apply to our constraints, and they’re hand-waving over a lot of stuff,” it’s that, “Wow, we suck.” And that’s not the takeaway anyone should ever have. Even Netflix doesn’t operate the way that Netflix says that they do in their conference talks. It’s always fun sitting next to someone from the company that’s currently presenting and saying something to them, like, “Wow, I wish we did things that way.” And they said, “Yeah, I wish we did, too.”

And it’s always the case because it’s very hard to get on stage and talk for 45 minutes about here’s what we completely screwed up on, especially at the large publicly traded companies where it’s, “Wait, why did our stock price just dive five perce—oh, my God, what did you say on stage?” People care [laugh] about those things, and I get it; there’s a risk factor that I don’t have to deal with here.

Abby: I wish people would though. It would be so refreshing to hear someone like, “You know what? Ohh, we really messed this up, and let me walk you through what we did.” [laugh]. I think that would be nice.

Corey: On some level, giving that talk in enough detail becomes indistinguishable from rage-quitting in public.

Abby: [laugh].

Corey: I mean, I’m there for it. Don’t get me wrong. But I would love to see it.

Abby: I don’t think it has to be rage-quitting. One of the things that I talk to my team a lot about is the safety to fail. You can’t take risk if you’re too afraid to fail, right? And I think you can frame failure in a way of, “Hey, this didn’t work, but let me walk you through all the amazing things we learned from this. And here’s how we used that to take this and make this thing better.”

And I think there’s a positive way to frame it that’s not rage-quitting, but I do think we as an industry gloss over those learnings that you absolutely have to do. You fail; everything does not work the first time perfectly. It is not brilliant out the gate. If you’ve done an MVP and it’s perfect and every customer loves it, well then, you sat on that for way too long. [laugh]. And I think it’s just really getting comfortable with this didn’t work the first time or the fourth, but look, at time seven, this is where we got and this is what we’ve learned.

Corey: I want to thank you for taking so much time out of your day to wind up speaking to me about things that in many cases are challenging to talk about because it’s the things people don’t talk about in the real world. If people want to learn more about what you’re up to, who you are, et cetera, where can they find you?

Abby: They can find me on the Twitters at @ab415. I think that’s the best way to start, although I will say that I am not as prolific as you are on Twitter.

Corey: That’s a good thing.

Abby: I’m a half-assed Tweeter. [laugh]. I will own it.

Corey: Oh, I put my full ass into it every time, in every way.

Abby: [laugh]. I do skim it a lot. I get a lot of my tech news from there. Like, “What are people mad about today?” And—

Corey: The daily outrage. Oh, yeah.

Abby: The daily outrage. “What’s Corey ranting about today? Let’s see.” [laugh].

Corey: We will, of course, put a link to your Twitter profile in the [show notes 00:35:39]. Thank you so much for taking the time to speak with me. I appreciate it.

Abby: Hey, it was my pleasure.

Corey: Abby Kearns, CTO at Puppet. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with a comment telling me about the amazing podcast content you create, start to finish, at Netflix.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Ewere

Cloud, DevOps Engineer, Blogger and Author

Links:

  • Infrastructure Monitoring with Amazon CloudWatch: https://www.amazon.com/Infrastructure-Monitoring-Amazon-CloudWatch-infrastructure-ebook/dp/B08YS2PYKJ
  • LinkedIn: https://www.linkedin.com/in/ewere/
  • Twitter: https://twitter.com/nimboya
  • Medium: https://medium.com/@nimboya
  • My Cloud Series: https://mycloudseries.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Honeycomb. When production is running slow, it's hard to know where problems originate: is it your application code, users, or the underlying systems? I’ve got five bucks on DNS, personally. Why scroll through endless dashboards, while dealing with alert floods, going from tool to tool to tool that you employ, guessing at which puzzle pieces matter? Context switching and tool sprawl are slowly killing both your team and your business. You should care more about one of those than the other, which one is up to you. Drop the separate pillars and enter a world of getting one unified understanding of the one thing driving your business: production. With Honeycomb, you guess less and know more. Try it for free at Honeycomb.io/screaminginthecloud. Observability, it’s more than just hipster monitoring.

Corey: This episode is sponsored in part by Liquibase. If you’re anything like me, you’ve screwed up the database part of a deployment so severely that you’ve been banned from touching every anything that remotely sounds like SQL, at at least three different companies. We’ve mostly got code deployments solved for, but when it comes to databases we basically rely on desperate hope, with a roll back plan of keeping our resumes up to date. It doesn’t have to be that way. Meet Liquibase. It is both an open source project and a commercial offering. Liquibase lets you track, modify, and automate database schema changes across almost any database, with guardrails to ensure you’ll still have a company left after you deploy the change. No matter where your database lives, Liquibase can help you solve your database deployment issues. Check them out today at liquibase.com. Offer does not apply to Route 53.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I periodically make observations that monitoring cloud resources has changed somewhat since I first got started in the world of monitoring. My experience goes back to the original Call of Duty. That’s right: Nagios.

When you set instances up, it would theoretically tell you when they were unreachable or certain thresholds didn’t work. It was janky but it kind of worked, and that was sort of the best we have. The world has progressed as cloud has become more complicated, as technologies have become more sophisticated, and here today to talk about this is the first AWS Hero from Africa and author of a brand new book, Ewere Diagboya. Thank you for joining me.

Ewere: Thanks for the opportunity.

Corey: So, you recently published a book on CloudWatch. To my understanding, it is the first such book that goes in-depth with not just how to wind up using it, but how to contextualize it as well. How did it come to be, I guess is my first question?

Ewere: Yes, thanks a lot, Corey. The name of the book is Infrastructure Monitoring with Amazon CloudWatch, and the book came to be from the concept of looking at the ecosystem of AWS cloud computing and we saw that a lot of the things around cloud—I mostly talked about—most of this is [unintelligible 00:01:49] compute part of AWS, which is EC2, the containers, and all that, you find books on all those topics. They are all proliferated all over the internet, you know, and videos and all that.

But there is a core behind each of these services that no one actually talks about and amplifies, which is the monitoring part, which helps you to understand what is going on with the system. I mean, knowing what is going on with the system helps you to understand failures, helps you to predict issues, helps you to also envisage when a failure is going to happen so that you can remedy it and also [unintelligible 00:02:19], and in some cases, even give you a historical view of the system to help you understand how a system has behaved over a period of time.

Corey: One of the articles that I put out that first really put me on AWS’s radar, for better or worse, was something that I was commissioned to write for Linux Journal, back when that was a print publication. And I accidentally wound up getting the cover of it with my article, “CloudWatch is of the devil, but I must use it.” And it was a painful problem that people generally found resonated with them because no one felt they really understood CloudWatch; it was incredibly expensive; it didn’t really seem like it was at all intuitive, or that there was any good way to opt out of it, it was just simply there, and if you were going to be monitoring your system in a cloud environment—which of course you should be—it was just sort of the cost of doing business that you then have to pay for a third-party tool to wind up using the CloudWatch metrics that it was gathering, and it was just expensive and unpleasant all around. Now, a lot of the criticisms I put about CloudWatch’s limitations in those days, about four years ago, have largely been resolved or at least mitigated in different ways. But is CloudWatch still crappy, I guess, is my question?

Ewere: Um, yeah. So, at the moment, I think, like you said, CloudWatch has really evolved over time. I personally also had that issue with CloudWatch when I started using CloudWatch; I had the challenge of usability, I had the challenge of proper integration, and I will talk about my first experience with CloudWatch here. So, when I started my infrastructure work, one of the things I was doing a lot was EC2, basically. I mean, everyone always starts with EC2 at the first time.

And then we had a downtime. And then my CTO says, “Okay, [Ewere 00:04:00], check what’s going on.” And I’m like, “How do I check?” [laugh]. I mean, I had no idea of what to do.

And he says, “Okay, there’s a tool called CloudWatch. You should be able to monitor.” And I’m like, “Okay.” I dive into CloudWatch, and boom, I’m confused again. And you look at the console, you see, it shows you certain metrics, and yet [people 00:04:18] don’t understand what CPU metric talks about, what does network bandwidth talks about?

And here I am trying to dig, and dig, and dig deeper, and I still don’t get [laugh] a sense of what is actually going on. But what I needed to find out was, I mean, what was wrong with the memory of the system, so I delved into trying to install the CloudWatch agent, get metrics and all that. But the truth of the matter was that I couldn’t really solve my problem very well, but I had [unintelligible 00:04:43] of knowing that I don’t have memory out of the box; it’s something that has to set up differently. And trust me, after then I didn’t touch CloudWatch [laugh] again. Because, like you said, it was a problem, it was a bit difficult to work with.

But fast forward a couple of years later, I could actually see someone use CloudWatch for a lot of beautiful stuff, you know? It creates beautiful dashboards, creates some very well-aggregated metrics. And also with the aggregated alarms that CloudWatch comes with, [unintelligible 00:05:12] easy for you to avoid what to call incident fatigue. And then also, the dashboards. I mean, there are so many dashboards that simplified to work with, and it makes it easy and straightforward to configure.

So, the bootstrapping and the changes and the improvements on CloudWatch over time has made CloudWatch a go-to tool, and most especially the integration with containers and Kubernetes. I mean, CloudWatch is one of the easiest tools to integrate with EKS, Kubernetes, or other container services that run in AWS; it’s just, more or less, one or two lines of setup, and here you go with a lot of beautiful, interesting, and insightful metrics that you will not get out of the box, and if you look at other monitoring tools, it takes a lot of time for you to set up, for you to configure, for you to consistently maintain and to give you those consistent metrics you need to know what’s going on with your system from time to time.

Corey: The problem I always ran into was that the traditional tools that I was used to using in data centers worked pretty well because you didn’t have a whole lot of variability on an hour-to-hour basis. Sure, when you installed new servers or brought up new virtual machines, you had to update the monitoring system. But then you started getting into this world of ephemerality with auto-scaling originally, and later containers, and—God help us all—Lambda now, where it becomes this very strange back-and-forth story of, you need to be able to build something that, I guess, is responsive to that. And there’s no good way to get access to some of the things that CloudWatch provides, just because we didn’t have access into AWS’s systems the way that they do. The inverse, though, is that they don’t have access into things running inside of the hypervisor; a classic example has always been memory: memory usage is an example of something that hasn’t been able to be displayed traditionally without installing some sort of agent inside of it. Is that still the case? Are there better ways of addressing those things now?

Ewere: So, that’s still the case, I mean, for EC2 instances. So before, now, we had an agent called a CloudWatch agent. Now, there’s a new agent called Unified Cloudwatch Agent which is, I mean, a top-notch from CloudWatch agent. So, at the moment, basically, that’s what happens on the EC2 layer. But the good thing is when you’re working with containers, or more or less Kubernetes kind of applications or systems, everything comes out of the box.

So, with containers, we’re talking about a [laugh] lot of moving parts. The container themselves with their own CPU, memory, disk, all the metrics, and then the nodes—or the EC2 instance of the virtual machines running behind them—also having their own unique metrics. So, within the container world, these things are just a click of a button. Everything happens at the same time as a single entity, but within the EC2 instance and ecosystem, you still find this there, although the setup process has been a bit easier and much faster. But in the container world, that problem has totally been eliminated.

Corey: When you take a look at someone who’s just starting to get a glimmer of awareness around what CloudWatch is and how to contextualize it, what are the most common mistakes people make early on?

Ewere: I also talked about this in my book, and one of the mistakes people make in terms of CloudWatch, and monitoring in generalities: “What am I trying to figure out?” [laugh]. If you don’t have that answer clearly stated, you’re going to run into a lot of problems. You need to answer that question of, “What am I trying to figure out?” I mean, monitoring is so broad, monitoring is so large that if you do not have the answer to that question, you’re going to get yourself into a lot of trouble, you’re going to get yourself into a lot of confusion, and like I said, if you don’t understand what you’re trying to figure out in the first place, then you’re going to get a lot of data, you’re going to get a lot of information, and that can get you confused.

And I also talked about what I call alarm fatigues or incident fatigues. This happens when you configure so many alarms, so many metrics, and you’re getting a lot of alarms hitting and notification services—whether it’s Slack, whether it’s an email—and it causes fatigue. What happens here is the person who should know what is going on with the system gets a ton of messages and in that scenario can miss something very important because there’s so many messages coming in, so many integrations coming in. So, you should be able to optimize appropriately, to be able to, like you said, conceptualize what you’re trying to figure out, what problems are you trying to solve? Most times you really don’t figure this out for a start, but there are certain bare minimums you need to know about, and that’s part of what I talked about in the book.

One of the things that I highlighted in the book when I talked about monitoring of different layers is, when you’re talking about monitoring of infrastructure, say compute services, such as virtual machines, or EC2 instances, the certain baseline and metrics you need to take note of that are core to the reliability, the scalability, and the efficiency of your system. And if you focus on these things, you can have a baseline starting point before you start going deeper into things like observability and knowing what’s going on entirely with your system. So, baseline understanding of—baseline metrics, and baseline of what you need to check in terms of different kinds of services you’re trying to monitor is your starting point. And the mistake people make is that they don’t have a baseline. So, we do not have a baseline; they just install a monitoring tool, configure a CloudWatch, and they don’t know the problem they’re trying to solve [laugh] and that can lead to a lot of confusion.

Corey: So, what inspired you from, I guess, kicking the tires on CloudWatch—the way that we all do—and being frustrated and confused by it, all the way to the other side of writing a book on it? What was it that got you to that point? Were you an expert on CloudWatch before you started writing the book, or was it, “Well, by the time this book is done, I will certainly know [laugh] more about the service than I did when I started.”

Ewere: Yeah, I think it’s a double-edged sword. [laugh]. So, it’s a combination of the things you just said. So, first of all, I have experienced with other monitoring tools; I have love for reliability and scalability of a system. I started Kubernetes at some of the early times Kubernetes came out, when it was very difficult to deploy, when it was very difficult to set up.

Because I’m looking at how I can make systems a little bit more efficient, a little bit more reliable than having to handle a lot of things like auto-scaling, having to go through the process of understanding how to scale. I mean, that’s a school of its own that you need to prepare yourself for. So, first of all, I have a love for making sure systems are reliable and efficient, and second of all, I also want to make sure that I know what is going on with my system per time, as much as possible. The level of visibility of a system gives you the level of control and understanding
of what your system is doing per time. So, those two things are very core to me.

And then thirdly, I had a plan of a streak of books I want to write based on AWS, and just like monitoring is something that is just new. I mean, if you go to the package website, this is the first book on infrastructure monitoring AWS with CloudWatch; it’s not a very common topic to talk about. And I have other topics in my head, and I really want to talk about things like networking, and other topics that you really need to go deep inside to be able to appreciate the value of what you see in there with all those scenarios because in this book, every chapter, I created a scenario of what a real-life monitoring system or what you need to do looks like. So, being that I have those premonitions, I know that whenever it came to, you know, to share with the world what I know in monitoring, what I’ve learned in monitoring, I took a [unintelligible 00:12:26]. And then secondly, as this opportunity for me to start telling the world about the things I learned, and then I also learned while writing the book because there are certain topics in the book that I’m not so much of an expert in things, like big data and all that.

I had to also learn; I had to take some time to do more research, to do more understanding. So, I use CloudWatch, okay? I’m kind of good in CloudWatch, and also, I also had to do more learning to be able to disseminate this information. And also, hopefully, X-Ray some parts of monitoring and different services that people do not really pay so much attention into.

Corey: What do you find that is still the most, I guess, confusing to you as you take a look across the ecosystem of the entire CloudWatch space? I mean, every time I play with it, I take a look, and I get lost in, “Oh, they have contributor analyses, and logs, and metrics.” And it’s confusing, and every time I wind up, I guess, spiraling out of control. What do you find that, after all of this, is a lot easier for you, and what do you find that’s a lot more understandable?

Ewere: I’m still going to go back to the containers part. I’m sorry, I’m in love containers. [laugh].

Corey: No, no, it’s fair. Containers are very popular. Everyone loves them. I’m just basically anti-container based upon no better reason than I’m just stubborn and bloody-minded most of the time.

Ewere: [laugh]. So, pretty much like I said, I kind of had experience with other monitoring tools. Trust me, if you want to configure proper container monitoring for other tools, trust me, it’s going to take you at least a week or two to get it properly, from the dashboards, to the login configurations, to the piping of the data to the proper storage engine. These are things I talked about in the book because I took monitoring from the ground up. I mean, if you’ve never done monitoring before, when you take my book, you will understand the basic principles of monitoring.

And [funny 00:14:15], you know, monitoring has some big data process, like an ETL process: extraction, transformation, and writing of data into an analytic system. So, first of all, you have to battle that. You have to talk about the availability of your storage engine. What are you using? An Elasticsearch? Are you using an InfluxDB? Where do you want to store your data? And then you have to answer the question of how do I visualize the data? What method do I realize this data? What kind of dashboards do I want to use? What methods of representation do I need to represent this data so that it makes sense to whoever I’m sharing this data with. Because in monitoring, you definitely have to share data with either yourself or with someone else, so the way you present the data needs to make sense. I’ve seen graphs that do not make sense. So, it requires some level of skill. Like I said, I’ve [unintelligible 00:15:01] where I spent a week or two having to set up dashboards. And then after setting up the dashboard, someone was like, “I don’t understand, and we just need, like, two.” And I’m like, “Really?” [laugh]. You know? Because you spend so much time. And secondly, you discover that repeatability of that process is a problem. Because some of these tools are click and drag; some of them don’t have JSON configuration. Some do, some don’t. So, you discover that scalability of this kind of system becomes a problem. You can’t repeat the dashboards: if you make a change to the system, you need to go back to your dashboard, you need to make some changes, you need to update your login, too, you need to make some changes across the layer. So, all these things is a lot of overhead [laugh] that you can cut off when you use things like Container Insights in CloudWatch—which is a feature of CloudWatch. So, for me, that’s a part that you can really, really suck out so much juice from in a very short time, quickly and very efficiently. On the flip side, when you talk about monitoring for big data services, and monitoring for a little bit of serverless, there might be a little steepness in the flow of the learning curve there because if you do not have a good foundation in serverless, when you get into [laugh] Lambda Insights in CloudWatch, trust me, you’re going to be put off by that; you’re going to get a little bit confused. And then there’s also multifunction insights at the moment. So, you need to have some very good, solid foundation in some of those topics before you can get in there and understand some of the data and the metrics that CloudWatch is presenting to you. And then lastly, things like big data, too, there are things that monitoring is still being properly fleshed out. Which I think that in the coming months and years to come, they will become more proper and they will become more presentable than they are at the moment.

Corey: This episode is sponsored by our friends at Oracle HeatWave is a new high-performance accelerator for the Oracle MySQL Database Service. Although I insist on calling it “my squirrel.” While MySQL has long been the worlds most popular open source database, shifting from transacting to analytics required way too much overhead and, ya know, work. With HeatWave you can run your OLTP and OLAP, don’t ask me to ever say those acronyms again, workloads directly from your MySQL database and eliminate the time consuming data movement and integration work, while also performing 1100X faster than Amazon Aurora, and 2.5X faster than Amazon Redshift, at a third of the cost. My thanks again to Oracle Cloud for sponsoring this ridiculous nonsense.

Corey: The problem I’ve always had with dashboards is it seems like managers always want them—“More dashboards, more dashboards”—then you check the usage statistics of who’s actually been viewing the dashboards and the answer is, no one since you demoed it to the execs eight months ago. But they always claim to want more. How do you square that?I guess, slicing between what people asked for and what they actually use.

Ewere: [laugh]. So yeah, one of the interesting things about dashboards in terms of most especially infrastructure monitoring, is the dashboards people really want is a revenue dashboards. Trust me, that’s what they want to see; they want to see the money going up, up, up, [laugh] you know? So, when it comes to—

Corey: Oh, yes. Up and to the right, then everyone’s happy. But CloudWatch tends to give you just very, very granular, low-level metrics of thing—it’s hard to turn that into something executives care about.

Ewere: Yeah, what people really care about. But my own take on that is, the dashboards are actually for you and your team to watch, to know what’s going on from time to time. But what is key is setting up events across very specific and sensitive data. For example, when any kind of sensitive data is flowing across your system and you need to check that out, then you tie a metric to that, and in turn alarm to it. That is actually the most important thing for anybody.

I mean, for the dashboards, it’s just for you and your team, like I said, for your personal consumption. “Oh, I can see all the RDS connections are getting too high, we need to upgrade.” Oh, we can see that all, the memory, there was a memory spike in the last two hours. I know that’s for you and your team to consume; not for the executive team. But what is really good is being able to do things like aggregate data that you can share.

I think that is what the executive team would love to see. When you go back to the core principles of DevOps in terms of the DevOps Handbook, you see things like a mean time to recover, and change failure rate, and all that. The most interesting thing is that all these metrics can be measured only by monitoring. You cannot change failure rates if you don’t have a monitoring system that tells you when there was a failure. You cannot know your release frequency when you don’t have a metric that measures number of deployments you have and is audited in a particular metric or a particular aggregator system.

So, we discovered that the four major things you measure in DevOps are all tied back to monitoring and metrics, at minimum, to understand your system from time to time. So, what the executive team actually needs is to get a summary of what’s going on. And one of the things I usually do for almost any company I work for is to share some kind of uptime system with them. And that’s where CloudWatch Synthetics Canary come in. So, Synthetic Canary is a service that helps you calculate that helps you check for uptime of the system.

So, it’s a very simple service. It does a ping, but it is so efficient, and it is so powerful. How is it powerful? It does a ping to a system and it gets a feedback. Now, if the status code of your service, it’s not 200 or not 300, it considers it downtime.

Now, when you aggregate this data within a period of time, say a month or two, you can actually use that data to calculate the uptime of your system. And that uptime [unintelligible 00:19:50] is something you can actually share to your customers and say, “Okay, we have an SLA of 99.9%. We have an SLA of 99.8%.” That data should not be doctored data; it should not be a data you just cook out of your head; it should be based on your system that you have used, worked with, monitored over a period of time so that the information you share with your customers are genuine, they are truthful, and they are something that they can also see for themselves.

Hence companies are using [unintelligible 00:20:19] like status page to know what’s going on from time to time whenever there is an incident and report back to their customers. So, these are things that executives will be more interested in than just dashboards, [laugh] dashboards, and more dashboards. So, it’s more or less not about what they really ask for, but what you know and what you believe you are going to draw value from. I mean, an executive in a meeting with a client and says, “Hey, we got a system that has 99.9% uptime.”

He opens the dashboard or he opens the uptime system and say, “You see our uptime? For the past three months, this has been our metric.” Boom. [snaps fingers]. That’s it. That’s value, instantly. I’m not showing [laugh] the clients and point of graphs, you know? “Can you explain the memory metric?” That’s not going to pass the message, send the message forward.

Corey: Since your book came out, I believe, if not, certainly by the time it was finished being written and it was in review phase, they came out with Managed Prometheus and Managed Grafana. It looks almost like they’re almost trying to do a completely separate standalone monitoring stack of AWS tooling. Is that a misunderstanding of what the tools look like, or is there something to that?

Ewere: Yeah. So, I mean by the time those announced at re:Invent, I’m like, “Oh, snap.” I almost told my publisher, “You know what? We need to add three more chapters.” [laugh]. But unfortunately, we’re still in review, in preview.

I mean, as a Hero, I kind of have some privilege to be able to—a request for that, but I’m like, okay, I think it’s going to change the narrative of what the book is talking about. I think I’m going to pause on that and make sure this finishes with the [unintelligible 00:21:52], and then maybe a second edition, I can always attach that. But hey, I think there’s trying to be a galvanization between Prometheus, Grafana, and what CloudWatch stands for. Because at the moment, I think it’s currently on pre-release, it’s not fully GA at the moment, so you can actually use it. So, if you go to Container Insights, you can see that you can still get how Prometheus and Grafana is presenting the data.

So, it’s more or less a different view of what you’re trying to see. It’s trying to give you another perspective of how your data is presented. So, you’re going to have CloudWatch: it’s going to have CloudWatch dashboards, it’s going to have CloudWatch metrics, but hey, this different tools, Prometheus, Grafana, and all that, they all have their unique ways of presenting the data. And part of the reason I believe AWS has Prometheus and Grafana there is, I mean, Prometheus is a huge cloud-native open-source monitoring, presentation, analytics tool; it packs a lot of heat, and a lot of people are so used to it. Everybody like, “Why can’t I have Prometheus in CloudWatch?”

I mean—so instead of CloudWatch just being a simple monitoring tool, [unintelligible 00:22:54] CloudWatch has become an ecosystem of monitoring tool. So, we got—we’re not going to see cloud [unintelligible 00:23:00], or just [unintelligible 00:23:00] log, analytics, metrics, dashboards, no. We’re going to see it as an ecosystem where we can plug in other services, and then integrate and work together to give us better performance options, and also different perspectives to the data that is being collected.

Corey: What do you think is next, as you take a look across the ecosystem, as far as how people are thinking about monitoring and observability in a cloud context? What are they missing? Where’s the next evolution lead?

Ewere: Yeah, I think the biggest problem with monitoring, which is part of the introduction part of the book, where I talked about the basic types of monitoring—which is proactive and reactive monitoring—is how do we make sure we know before things happen? [laugh]. And one of the things that can help with that is machine learning. There is a small ecosystem that is not so popular at the moment, which talks about how we can do a lot of machine learning in DevOps monitoring observability. And that means looking at historic data and being able to predict on the basic level.

Looking at history, [then are 00:24:06] being able to predict. At the moment, there are very few tools that have models running at the back of the data being collected for monitoring and metrics, which could actually revolutionize monitoring and observability as we see it right now. I mean, even the topic of observability is still new at the moment. It’s still very integrated. Observability just came into Cloud, I think, like, two years ago, so it’s still being matured.

But one thing that has been missing is seeing the value AI can bring into monitoring. I mean, this much [unintelligible 00:24:40] practically tell us, “Hey, by 9 p.m. I’m going to go down. I think your CPU or memory is going down. I think I’m line 14 of your code [laugh] is a problem causing the bug. Please, you need to fix it by 2 p.m. so that by 6 p.m., things can run perfectly.” That is going to revolutionize monitoring. That’s going to revolutionize observability and bring a whole new level to how we understand and monitor the systems.

Corey: I hope you’re right. If you take a look right now, I guess, the schism between monitoring and observability—which I consider to be hipster monitoring, but they get mad when I say that—is there a difference? Is it just new phrasing to describe the same concepts, or is there something really new here?

Ewere: In my book, I said, monitoring is looking at it from the outside in, observability is looking at it from the inside out. So, what monitoring does not see under, basically, observability sees. So, they are children of the same mom. That’s how I put it. One actually needs the other and both of them cannot be separated from each other.

What we’ve been working with is just understanding the system from the surface. When there’s an issue, we go to the aggregated results that come out of the issue. Very basic example: you’re in a Java application, and we all know Java is very memory intensive, on the very basic layer. And there’s a memory issue. Most times, infrastructure is the first hit with the resultant of that.

But the problem is not the infrastructure, it’s maybe the code. Maybe garbage collection was not well managed; maybe they have a lot of variables in the code that is not used, and they’re just filling up unnecessary memory locations; maybe there’s a loop that’s not properly managed and properly optimized; maybe there’s a resource on objects that has been initialized that has not been closed, which will cause a heap in the memory. So, those are the things observability can help you track. Those are the things that we can help you see. Because observability runs from within the system and send metrics out, while basic monitoring is about understanding what is going on on the surface of the system: memory, CPU, pushing out logs to know what’s going on and all that.

So, on the basic level, observability helps gives you, kind of, a deeper insight into what monitoring is actually telling you. It’s just like the result of what happened. I mean, we are told that the symptoms of COVID is coughing, sneezing, and all that. That’s monitoring. [laugh].

But before we know that you actually have COVID, we need to go for a test, and that’s observability. Telling us what is causing the sneezing, what is causing the coughing, what is causing the nausea, all the symptoms that come out of what monitoring is saying. Monitoring is saying, “You have a cough, you have a runny nose, you’re sneezing.” That is monitoring. Observability says, “There is a COVID virus in the bloodstream. We need to fix it.” So, that’s how both of them act.

Corey: I think that is probably the most concise and clear definition I’ve ever gotten on the topic. If people want to learn more about what you’re up to, how you view about these things—and of course, if they want to buy your book, we will include a link to that in the [show notes 00:27:40]—where can they find you?

Ewere: I’m on LinkedIn; I’m very active on LinkedIn, and I also shared the LinkedIn link. I’m very active on Twitter, too. I tweet once in a while, but definitely, when you send me a message on Twitter, I’m also going to be very active.

I also write blogs on Medium, I write a couple of blogs on Medium, and that was part of why AWS recognized me as a Hero because I talk a lot about different services, I help with comparing services for you so you can choose better. I also talk about setting basic concepts, too; if you just want to get your foot wet into some stuff and you need something very summarized, not AWS documentation per se, something that you can just look at and know what you need to do with the service, I talk about them also in my blogs. So yeah, those are the two basic places I’m in: LinkedIn and Twitter.

Corey: And we will, of course, put links to that in the [show notes 00:28:27]. Thank you so much for taking the time to speak with me. I appreciate it.

Ewere: Thanks a lot.

Corey: Ewere Diagboya, head of cloud at My Cloud Series. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you hated this podcast, please leave a five-star review on your podcast platform of choice along with a comment telling me how many more dashboards you would like me to build that you will never look at.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Tim
Tim’s tech career spans over 20 years through various sectors. Tim’s initial journey into tech started as a US Marine. Later, he left government contracting for the private sector, working both in large corporate environments and in small startups. While working in the private sector, he honed his skills in systems administration and operations for largeUnix-based datastores.

Today, Tim leverages his years in operations, DevOps, and Site Reliability Engineering to advise and consult with clients in his current role. Tim is also a father of five children, as well as a competitive Brazilian Jiu-Jitsu practitioner. Currently, he is the reigning American National and 3-time Pan American Brazilian Jiu-Jitsu champion in his division.

Links:

  • Twitter: https://twitter.com/elchefe
  • The Duckbill Group: https://duckbillgroup.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Honeycomb. When production is running slow, it's hard to know where problems originate: is it your application code, users, or the underlying systems? I’ve got five bucks on DNS, personally. Why scroll through endless dashboards, while dealing with alert floods, going from tool to tool to tool that you employ, guessing at which puzzle pieces matter? Context switching and tool sprawl are slowly killing both your team and your business. You should care more about one of those than the other, which one is up to you. Drop the separate pillars and enter a world of getting one unified understanding of the one thing driving your business: production. With Honeycomb, you guess less and know more. Try it for free at Honeycomb.io/screaminginthecloud. Observability, it’s more than just hipster monitoring.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Periodically, I have a whole bunch of guests come on up, second time. Now, it’s easy to take the naive approach of assuming that it’s because it’s easier for me to find a guest if I know them and don’t have to reach out to brand new people all the time. This is absolutely correct; I’m exceedingly lazy. But I don’t have too many folks on a third time, but that changes today.

My guest is Tim Banks. I’ve had him on the show twice before, both times it led to really interesting conversations around a wide variety of things. Since those episodes, Tim has taken the job as a principal cloud economist here at The Duckbill Group. Yes, that is probably the strangest interview process you can imagine, but here we are. Tim, thank you so much for joining me both on the show and in the business.

Tim: My pleasure, Corey. It was definitely an interesting interview process, you know, but I was glad to be here. So, I’m happy to be here a third time. I don’t know if you get a jacket like you do in Saturday Night Live, if you host, like, a fifth time, but we’ll see. Maybe it’s a vest. A cool vest would be nice.

Corey: We can come up with something.[ effectively, it can be like reverse hangman where you wind up getting a vest and every time you come on after that you get a sleeve, then you get a second sleeve, and then you get a collar, and we can do all kinds of neat stuff.

Tim: I actually like that idea a lot.

Corey: So, I’m super excited to be able to have this conversation with you because I don’t normally talk a lot on this show about what cloud economics is because my guest usually is not as deep into the space as I am, and that’s fine; people should never be as deep into this space as I am, in the general sense, unless they work here. Awesome. But I do guest on other shows, and people ask me all kinds of questions about AWS billing and cloud economics, and that’s fine, it’s great, but they don’t ask the questions about the space in the same way that I would and the way that I think about it. So, it’s hard for me to interview myself. Now, I’m not saying I won’t try it someday, but it’s challenging. But today, I get to take the easy path out and talk to you about it. So Tim, what the hell is a principal cloud economist?

Tim: So, a principal cloud economist, is a cloud computing expert, both in architecture and practice, who looks at cloud cost in the same way that a lot of folks look at cloud security, or cloud resilience, or cloud performance. So, the same engineering concerns you have about making sure that your API stays up all the time, or to make sure that you don’t have people that are able to escape containers or to make sure that you can have super, super low response times, is the same engineering fundamentals that I look at when I’m trying to find a way to reduce your AWS bill.

Corey: Okay. When we say cloud cost and cloud economics, the natural picture that leads to mind is, “Oh, I get it. You’re an Excel jockey.” And sometimes, yeah, we all kind of play those roles, but what you’re talking about is something else entirely. You’re talking about engineering expertise.

And sure enough, if you look at the job postings we have for roles on the team from time to time, we have not yet hired anyone who does not have an engineering and architecture background. That seems odd to folks who do not spend a lot of time thinking about the AWS bill. I’m told those people are what is known as ‘happy.’ But here we are. Why do we care about the engineering aspect of any of this?

Tim: Well, I think first and foremost because what we’re doing in essence, is still engineering. People aren’t putting construction paper up on [laugh] AWS; sometimes they do put recipes up on there, but it still involves working on a computer, and writing code, and deploying it somewhere. So, to have that basic understanding of what it is that folks are doing on the platform, you have to have some engineering experience, first and foremost. Secondly, the fact of the matter is that most cost optimization, in my opinion, can be done on the whiteboard, before anything else, and really I think should be done on the whiteboard before anything else. And so the Excel aspect of it is always reactive. “We have now spent this much. How much was it? Where did it go?” And now we have to figure out where it went.

I like to figure out and get a ballpark on how much something is going to cost before I write the first line of code. I want to know, hey, we have a tier here, we’re using this kind of storage, it’s going to take this kind of instance types. Okay, well, I’ve got an idea of how much it’s going to cost. And I was like, “You know, that’s going to be expensive. Before we do anything, is there a way that we can reduce costs there?”

And so I’m reverse engineering that on already deployed workloads. Or when customers want to say, “Hey, we were thinking about doing this, and this is our proposed architecture,” I’m going to look at it and say, “Well, if you do this and this and this and this, you can save money.”

Corey: So, it sounds like you and I have a bit of a philosophical disagreement in some ways. One of my recurring talking points has always been that, “Oh, by and large, application developers don’t need to think overly much about cloud cost. What they need to know generally fits on an index card.” It’s, okay, big things cost more than small things; if you turn something on, it will never get turned off and will bill you in perpetuity; data transfer has some weird stuff; and if you store data, you pay for data, like, that level of baseline understanding. When I’m trying to build something out my immediate thought is, great, is this thing possible?

Because A, I don’t always know that it is, and B, I’m super bad at computers so for me, it may absolutely not be, whereas you’re talking about
baking cost assessments into the architecture as a day one type of approach, even when sketching ideas out on the whiteboard. I’m curious as to how we diverge there. Can you talk more about your philosophy?

Tim: Sure. And the reason I do that is because, as most folks that have an engineering background in cloud infrastructure will tell you, you want to build resilience in, on the whiteboard. You certainly want to build performance in, on the whiteboard, right? And security folks will tell you you want to do security on the whiteboard. Because those things are hard to fix after they’re deployed.

As soon as they’re deployed, without that, you now have technical debt. If you don’t consider cost optimization and cost efficiency on the whiteboard, and then you try and do it after it’s deployed, you not only have technical debt, you may have actual real debt.

Corey: One of the comments I tend to give a lot is that architecture and cost are the same thing in the world of cloud. And I think that we might be in violent agreement, as Liz Fong-Jones is fond of framing it, where I am acutely aware of aspects of cost and that does factor into how I build things on the whiteboard—let’s also be very clear, most of the things that I build are very small scale; the largest cost by a landslide is the time I spend building it—in practice, that’s an awful lot of environments; people are always more expensive than the AWS environment they’re working on. But instead, it’s about baking in the assumptions and making sure you’re not coming up with something that is going to just be wasteful and horrible out of the gate, and I guess part of that also is the fact that I am at a level of billing understanding that I sort of absorbed these concepts intrinsically. Because to me, there is no difference between cost and architecture in an environment like this. You’re right, there’s always an inherent trade-off between cost and durability. On the one hand, I don’t like that. On the other, it feels like it’s been true forever and I don’t see a way out of it.

Tim: It is inescapable. And it’s interesting because you talk about the level of an application developer or something like that, like what is your level of concern, but retroactively, we’ll go in for cost optimization houses—and I’ve done this as far back as when I was working at AWS has a TAM—and I’ll ask the question to an application developer or database administrator, and I’m like, “Why do you do this? What do you have a string value for something that could be a Boolean?” And you’ll ask, “Well, what difference does that make?” Well, it makes a big difference when you’re talking about cycles for CPU.

You can reduce your CPU consumption on a database instance by changing a string to a Boolean, you need fewer instances, or you need a less powerful instance, or you need less memory. And now you can run a less expensive instance for your database architecture. Well, maybe for one node it’s not that biggest difference, but if you’re talking about something that’s multi-AZ and multi-node, I mean, that can be a significant amount of savings just by making one simple change.

Corey: And that might be the difference right there. I didn’t realize that, offhand. It makes sense if you think about it, but just realizing that I’ve made that mistake on one of my DynamoDB tables. It costs something like seven cents a month right now, so it’s not something I’m rushing to optimize, but you’re right, expand that out by a factor of a million or so, and we’re talking serious money, and then that sort of optimization makes an awful lot of sense. I think that my position on it is that when you’re building out something small scale as a demo or a proof of concept, spending time on optimizations like this is not the best use of anyone’s time or brain sweat, for lack of a better term. How do you wind up deciding when it’s time to focus on stuff like that?

Tim: Well, first, I will say that—I daresay that somewhere in the 80% of production workloads are just—were the POC, [laugh] right? Because, like, “It worked for this to get funding, let’s run it,” right?

Corey: Let they who does not have a DynamoDB table in production with the word ‘test’ or ‘dev’ in it cast the first stone.

Tim: It’s certainly not me. So, I understand how some of those decisions get made. And that’s why I think it’s better to think about it early. Because as I mentioned before, when you start something and say, “Hey, this works for now,” and you don’t give consideration to that in the future, or consideration for what it’s going to be like in the future, and when you start doing it, you’ll paint yourself into corners. That’s how you get something like static values put in somewhere, or that’s how you get something like, well, “We have to run this instance type because we didn’t build in the ability to be more microservice-based or stateless or anything like that.”

You’ve seen people that say, “Hey, we could save you a lot of money if you can move this thing off to a different tier.” And it’s like, “Well, that would be an extensive rewrite of code; that’d be very expensive.” I daresay that’s the main reason why most AS/400s are still being used right now is because it’s too expensive to rewrite the code.

Corey: Yeah, and there’s no AWS/400 that they can migrate to. Yet. Re:Invent is nigh.

Tim: So, I think that’s why, even at the very beginning, even if you were saying, “Well, this is something we will do later.” Don’t make it impossible for you to do later in your code. Don’t make it impossible for you to do later in your architecture. Make things as modular as possible, so that way you can say, “Hey”—later on down the road—“Oh, we can switch this instance type.” Or, “Here’s a new managed service that we can maybe save money on doing this.”

And you allow yourself to switch things out, or turn different knobs, or change the way you do things, and give yourself more options in the future, whether those options are for resilience, or those options or for security, or those options are for performance, or they’re for cost optimizations. If you make binding decisions earlier on, you’re going to have debt that’s going to build up at some point in the future, and then you’re going to have to pay the piper. Sometimes that piper is going to be AWS.

Corey: One thing that I think gets lost in a lot of conversations about cloud economics—because I know that it happened to me when I first started this place—where I am planning to basically go out and be the world’s leading expert in AWS cost analysis and understanding and optimization. Great. Then I went out into the world and started doing some of my first engagements, and they looked a lot less like far-future cost attribution projections and a lot more like, “What’s a reserved instance?” And, “We haven’t bought any of those in 18 months.” And, “Oh, yeah, we shut down an entire project six months ago. We should probably delete all the resources, huh?”

The stuff that I was preparing for at the high end of the maturity curve are great and useful and terrific to have conversations about in some very nuanced depth, but very often there’s a walk before you can run style of conversation where, okay, let’s do the easy stuff first before we start writing a whole bunch of bespoke internal stuff that maps your business needs to the AWS bill. How do you, I guess, reconcile those things where you’re on the one hand, you see the easy stuff and on the other, you see some of the just the absolutely challenging, very hard,
five-years-of-engineering-effort-style problems on the other?

Tim: Well, it’s interesting because I’ve seen one customer very recently who has brilliant analyses as to their cost; just well-charted, well-tagged, well-documented, well—you know, everything is diagrammed quite nicely and everything like that, and they’re very, very aware of their costs, but they leave test instances running all weekend, you know, and their associated volumes and things like that. And that’s a very easy thing to fix. That is a very, very low-hanging fruit. And so sometimes, you just have to look at where they’re spending their efforts where sometimes they do spend so much time chasing those hard to do things because they are hard to do and they’re exciting in an engineering aspect, and then something as simple as, “Hey, how about we delete these old volumes?” It just isn’t there.

Or, “How about we switch to your S3 bucket storage type?” Those are easy, low-hanging fruits, and you would be surprised how sometimes they just don’t get that. But at the same time, sometimes customers have, like, “Hey, we could knock this thing out, we knock this thing out,” because it’s Trusted Advisor. Every AI cost optimization recommendation you can get will tell you these five things to do, no matter who you are or where you are, but they don’t do the conceptual things like understanding some of the principles behind cost optimization and cost optimization architecture, and proactive cost optimization versus react with cost optimizations. So, you’re doing very conceptual education and conversations with folks rather than the, “Do these five things.” And I’ve not often found a customer that you have to do both on; it’s usually one or the other.

Corey: It’s funny that you made that specific reference to that example. One of my very first projects—not naming names. Generally, when it comes to things like this, you can tell stories or you can name names; I bias for stories—I was talking to a company who was convinced that their developer environments were incredibly overwrought, expensive, et cetera, and burning money. Okay, great. So, I talked about the idea of turning those things off at night or between test runs, deleting volumes to snapshot, and restore them on a schedule when people come in in the morning because all your developers sit in the same building in the same time zones. Great. They were super on board with the idea, and it was going to be a little bit of work, but all right, this was in the days before the EC2 Instance Scheduler, for example.

But first, let’s go ahead and do some analysis. This is one of those early engagements that really reinforced my idea of, yeah, before we start going too far down the rabbit hole, let’s double-check what’s going on in the account. Because periodically you encounter things that surprise people. Like, “What’s up with those Australia instances?” “Oh, we don’t have anything in that region.” “I believe you’re being sincere when you say this, however, the API generally doesn’t tell lies.”

So, that becomes a, oh, security incident time. But looking at this, they were right; they had some fairly sizable developer instances that were running all the time, but doing some analysis, their developer environment was 3% of their bill at the time and they hadn’t bought RIs in a year-and-a-half. And looking at what they were doing, there was so much easier stuff that they could do to generate significant savings without running the potential of turning a developer environment off at night in the middle of an incident or something like that. The risk factor and effort were easier just do the easy stuff, then do another pass and look at the deep stuff. And to be clear, they weren’t lying to me; they weren’t wrong.

Back when they started building this stuff out, their developer environments were significantly large and were a significant portion of their spend. And then they hit product-market fit, and suddenly their production environment had to scale significantly in a short period of time. Which, yay, cloud. It’s good at that. Then it just became such a small portion that developer environments weren’t really a thing. But the narrative internally doesn’t get updated very often because once people learn something, they don’t go back to relearn whether or not it’s still true. It’s a constant mistake; I make it myself frequently.

Tim: I think it’s interesting, there are things that we really need to put into buckets as far as what’s an engineering effort and what’s an administrative effort. And when I say ‘administrative effort,’ I mean if I can save money with a stroke of a pen, well, that’s going to be pretty easy, and that’s usually going to be RIs; that’s going to be EDPs, or PPAs or something like that, that don’t require engineering effort. It just requires administrative effort, I think RIs being the simplest ones. Like, “Oh, all I have to do is go in here and click these things four times and I’m going to save money?” “Well, let’s do that.”

And it’s surprising how often people don’t do that. But you still have to understand that, and whether it’s RIs or whether it’s a savings plan, it’s still a commitment of some kind, but if you are willing to make that commitment, you can save money with no engineering effort whatsoever. That’s almost free money.

Corey: So, much of what we do here comes down to psychology, in many ways, more than it does math. And a lot of times you’re right, everything you say is right, but in a large-scale environment, go ahead and click that button to buy the savings plan or the reserved instance, and that’s a $20 million purchase. And companies will stall for months trying to run a different series of analyses on this and what if this happens, what if that happens, and I get it because, “Yeah, I’m going to click this button that’s going to cost more money than I’ll make in my lifetime,” that’s a scary thing to do; I get it. But you’re going to spend the money, one way or the other, with the provider, and if you believe that number is too high, I get it; I am right there with you. Buy half of them right now and then you can talk about the rest until you get to a point of being comfortable with it.

Do it incrementally; it’s not all or nothing, you have one shot to make the buy. Take pieces out of it that makes sense. You know you’re probably not going to turn off your database cluster that handles all of production in the next year, so go ahead and go for it; it saves some money. Do the thing that makes sense. And that doesn’t require deep-dive analytics that requires, on some level, someone who’s seen a lot of these before who gets what customers are going through. And honestly, it’s empathy in many respects, becomes one of those powerful things that we can apply to our customer accounts.

Tim: Absolutely. I mean, people don’t understand that decision paralysis, about making those commitments costs you money. You can spend months doing analysis, but those months doing analysis, you’re going to spend 30, 40, 50, 60, 70% more on your EC2 instances or other compute than you would otherwise, and that can be quite significant. But it’s one of those cases where we talk about psychology around perfect being the enemy of good. You don’t have to make the perfect purchase of RIs or savings plans and have that so tuned perfectly that you’re going to get one hundred percent utilization and zero—like, you don’t have to do that.

Just do something. Do a little bit. Like you said, buy half; buy anything; just something, and you’re going to save money. And then you can run analysis later on, while you’re saving money [laugh] and get a little better and tune it up a little more and get more analysis on and maybe fine-tune it, but you don’t actually ever need to have it down to the penny. Like, it never has to be that good.

Corey: At some point, one of the value propositions we have for our customers has always been that we tell you when to stop focusing on saving money because there’s a theoretical cap of a hundred percent of the cloud bill that you can save, but you can make so much more than that by launching the right feature to the right market a little sooner; focus on that. Be responsible stewards of the money that’s invested with you, but by and large, as a general piece of guidance, at some point, stop cutting and go back to doing the thing that makes your company work. It’s not all about saving money at all costs for almost all of us. It is for us, but we’re sort of a special case.

Tim: Well, it’s a conversation I often have. It’s like, all right, are you trying to save money on AWS or are you trying to save money overall? So, if you’re going to spend $400,000 worth of engineering effort to save $10,000 on your AWS bill, that doesn’t make no sense. So—[laugh]—

Corey: Right. There has to be a strategic reason to do things like that—

Tim: Exactly.

Corey: —and make sure you understand the value of what you’re getting for this. One reason that we wind up charging the way that we do—and we’ve gotten questions on this for a while—has been that we charge a fixed fee for what we do on engagements. And similarly—people have asked this, but haven’t tied the two things together—you talk about cost optimization, but never cost-cutting. Why is that? Is that just a negative term?

And the answer has been no, they’re aligned. What we do focuses on what is best for the customer. Once that fixed fee is decided upon, every single thing that we say is what we would do if we were in the customer’s position. There are times we’ll look at what they have going on and say, “Ah, you really should spend more money here for resiliency, or durability,” or, “Okay, that is critical data that’s not being backed up. You should consider doing that.”

It’s why we don’t take percentages of things because, at that point, we’re not just going with the useful stuff, it’s, well we’re going to basically throw the entire kitchen sink at you. We had an early customer and I was talking to their AWS account manager about what we were going to be doing and their comment was, “Oh, saving money on AWS bills is great, make sure you check the EBS snapshots.” Yeah, I did that. They were spending 150 bucks a month on EBS snapshots, which is basically nothing. It’s one of those stories where if, in the course of an hour-long meeting, I can pay for that entire service, by putting a quarter on the table, I’m probably not going to talk about it barring [laugh] some extenuating circumstances.

Focus on the big things, not the things that worked in a different environment with a different account and different constraints. It’s hard to context switch like that, but it gets a lot easier when it is basically the entirety of what we do all day.

Tim: The difference I draw between cost optimization and cost-cutting is that cost optimization is ensuring that you’re not spending money unnecessarily, or that you’re maximizing your dollar. And so sometimes we get called in there, and we’re just validation for the measures they’ve already done. Like, “Your team is doing this exactly right. You’re doing the things you should be doing. We can nitpick if you want to; we’re going to save you $7 a year, but who cares about that? But y’all are doing what you should be doing. This is great. Going forward, you want to look for these things and look for these things and look for these things. We’re going to give you some more concepts so that you are cost-optimized in the future.” But it doesn’t necessarily mean that we have to cut your bill. Because if you’re already spending efficiently, you don’t need your bill cut; you’re already cost-optimized.

Corey: Oh, we’re not going to nitpick on that, you’re mostly optimized there. It’s like, “Yeah, that workload’s $140 million a year and rising; please, pick nits.” At which point? “Okay, great.” That’s the strategic reason to focus on something. But by and large, it comes down to understanding what the goals of clients are. I think that is widely misunderstood about what we do and how we do it.

The first question I always ask when someone does outreach of, “Hey, we’d like to talk about coming in here and doing a consulting
engagement with us.” “Great.” I always like to ask the quote-unquote, “Foolish question” of, “Why do you care about the AWS bill?” And occasionally I’ll get people who look at me like I have two heads of, “Why wouldn’t I care about the AWS bill?” Because there are more important things to care about for the business, almost certainly.

Tim: One of the things I try and do, especially when we’re talking about cost optimization, especially trying to do something for the right now so they can do things going forward, it’s like, you know, all right, so if we cut this much from your bill—if you just do nothing else, but do reserved instances or buy a savings plan, right, you’re going to save enough money to hire four engineers. Think about what four engineers would do for your overall business? And that’s how I want you to frame it; I want you to look at what cost optimization is going to allow you to do in the future without costing you any more money. Or maybe you save a little more money and you can shift it; instead of paying for your AWS bill, maybe you can train your developers, maybe you can get more developers, maybe you can get some ProServ, maybe you can do whatever, buy newer computers for your people so they can do—whatever it is, right? We’re not saying that you no longer have to spend this money, but saying, “You can use this money to do something other than give it to Jeff Bezos.”

Corey: This episode is sponsored in part by Liquibase. If you’re anything like me, you’ve screwed up the database part of a deployment so severely that you’ve been banned from touching every anything that remotely sounds like SQL, at at least three different companies. We’ve mostly got code deployments solved for, but when it comes to databases we basically rely on desperate hope, with a roll back plan of keeping our resumes up to date. It doesn’t have to be that way. Meet Liquibase. It is both an open source project and a commercial offering. Liquibase lets you track, modify, and automate database schema changes across almost any database, with guardrails to ensure you’ll still have a company left after you deploy the change. No matter where your database lives, Liquibase can help you solve your database deployment issues. Check them out today at liquibase.com. Offer does not apply to Route 53.

Corey: There was an article recently, as of the time of this recording, where Pinterest discussed what they had disclosed in one of their regulatory filings which was, over the next eight years, they have committed to pay AWS $3.2 billion. And in this article, they have the head of engineering talking to the reporter about how they’re thinking about these things, how they’re looking at things that are relevant to their business, and they’re talking about having a dedicated team that winds up doing a whole bunch of data analysis and running some analytics on all of these things, from piece to piece to piece. And that’s great. And I worry, on some level, that other companies are saying, “Oh, Pinterest is doing that. We should, too.” Yeah, for the course of this commitment, a 1% improvement is $32 million, so yeah, at that scale I’m going to hire a team of data scientists, too, look at these things. Your bill is $50,000 a month. Perhaps that’s not worth the effort you’re going to put into it, barring other things that contribute to it.

Tim: It’s interesting because we will get folks that will approach us that have small accounts—very small, small spend—and like, “Hey, can you come in and talk to us about this whatever.” And we can say very honestly, “Look, we could, but the amount of money we’re going to charge you is going to—it’s not going to be worth your while right now. You could probably get by on the automated recommendations, on the things that already out there on the internet that everybody can do to optimize their bill, and then when you grow to a point where now saving 10% is somebody’s salary, that’s when it, kind of, becomes more critical.” And it’s hard to say what point that is in anyone’s business, but I can say sometimes, “Hey, you know what? That’s not really what you need to focus on.” If you need to save $100 a month on your AWS bill, and that’s critical, you’ve got other concerns that are not your AWS bill.

Corey: So, back when you were interviewing to work here, one of the areas of focus that you kept bringing up was the concept of observability, and my response to this was, “Ah, hell. Another one.” Because let’s be clear, Mike Julian—my business partner and our CEO—has written a book called Practical Monitoring, and apparently what we learned from this is as soon as you finish writing a book on the topic, you never want to talk about that topic ever again, which yeah, in hindsight makes sense. Why do you care about observability when you’re here to look at cloud costs?

Tim: Because cloud costs is another metric, just like you would use for performance, or resilience, or security. You do real-time monitoring to see if somebody has compromised the system, you do real-time monitoring to see if you have bad performance, if response times are too slow. You do real-time monitoring to know if something has gone down and then you need to make adjustments, or that the automated responses you have in response to that downtime are working. But cloud costs, you send somebody a report at the end of the month. Can you imagine, if you will—just for a second—if you got a downtime report at the end of month, and then you can react to something that has gone down?

Or if you get a security report at the end of the month, and then you can react to the fact that somebody has your root keys? Or if you get [laugh] a report at the end of month, this said, “Hey, the CPU on this one was pegged. You should probably scale up.” That’s outrageous to anybody in this industry right now. But why do we accept that for cloud cost?

Corey: It’s worse than that. There are a number of startups that talk about, “Oh, real-time cloud cost monitoring. Okay, the only way you’re going to achieve such a thing is if you build an API shim that interprets everything that you’re telling your cloud control plane to do, taking cost metrics out of it, and then passing it on to the actual cloud control plane.” Otherwise, you’re talking about it showing up in the billing record in—ideally, eight hours; in practice, several days, or you’re talking about the CloudTrail events, which is not holistic but gives you some rough idea, but it’s also in some cases, 5 to 20 minutes delayed. There’s no real-time way to do this without significant disruption to what’s going on in your environment.

So, when I hear about, “Oh, we do real-time bill analysis.” Yeah, it feels—to be very direct—you don’t know enough about the problem space you’re working within to speak intelligently about it because anyone who’s played in this space for a while knows exactly how hard it is to get there. Now, I’ve talked to companies that have built real-time-ish systems that take that shim approach and acts sort of as a metadata sidecar ersatz billing system that tracks all of this so they can wind up intercepting potentially very expensive configuration mistakes. And that’s great. That’s also a bit beyond for a lot of folks today, but it’s where the industry is going. But there is no way to get there today, short of effectively intercepting all of those calls, in a way that is cohesive and makes sense. How do you square that circle given the complete lack of effective tooling?

Tim: Honestly, I’m going to point that right back at the cloud provider because they know how much you’re spending, real-time. They know exactly how much you spend in real-time. They’ve figured it out. They have the buckets, they have APIs for it internally. I’m sure they do; it would make no sense for them not to. Without giving anything anyway, I know that when I was at AWS, I knew how much they were spending, almost real-time.

Corey: That’s impressive. I wish that existed. My never having worked at AWS perspective on it is that they, of course, have the raw data effective immediately, or damn close to it, but the challenge for the billing system is distilling and summarizing and attributing all of that in a reasonable timeframe; it is an exabyte-scale problem. I’ve talked to folks there who have indicated it is comfortably north of a petabyte in raw data per day. And that was a couple of years ago, so one can only imagine as the footprint has increased, so has all of this.

I mean, the billing system is fundamentally magic from the outside. I’m not saying it’s good magic, but it is magic, and it’s something that is unappreciated, that every customer uses, and is one of those areas that doesn’t get the attention it deserves. Because, let’s be clear, here, we talk about observability; the bill is still the only thing that AWS offers that gives you a holistic overview of everything running in your account, in one place.

Tim: What I think is interesting is that you talk about this, the scale of the problem and that it makes it difficult to solve. At the same time, I can have a conversation with my partner about kitty litter, and then all of a sudden, I’m going to start getting ads about kitty litter within minutes. So, I feel like it’s possible to emit cost as a metric like you would CPU or disk. And if I’m going to look at who’s going to do that, I’m going to look right back at AWS. The fun part about that, though, is I know from AWS’s business model, that if that’s something they were to emit, it would also cost you, like, 25 cents per call, and then you would actually, like, triple your cloud costs just trying to figure out how much it costs you.

Corey: Only with 16 other billing dimensions because of course it would. And again, I’m talking about stuff, because of how I operate and how I think about this stuff, that is inherently corner case, or [vertex 00:31:39] case in many cases. But for the vast majority of folks, it’s not the, “Oh, you have this really weird data transfer paradigm between these two resources,” which yeah, that’s a problem that needs to be addressed in an awful lot of cases because data transfer pricing is bonkers, but instead it’s the, “Huh. You just spun up a big cluster that’s going to cost $20,000 a month.” You probably don’t need to wait a full day to flag that.

And you also can’t put this on the customer in the sense of, “Oh, just set some budget alarms, that’s great. That’s the first thing you should do in a new AWS account.” “Well, jackhole, I’ve done an awful lot of first things I’m supposed to do in an AWS account, in my dedicated test account for these sorts of things. It’s been four months, I’m not done yet with all of those first things I’m supposed to do.” It’s incredibly secure, increasingly expensive, and so far all it runs is a single EC2 instance that is mostly there just so that everything else doesn’t error out trying to
divide by zero.

Tim: There are some things that are built-in. If I stand up an EC2 instance and it goes down, I’m going to get an alert that this instance terminated for some reason. It’s just going to show up informationally.

Corey: In the console. You’re not going to get called about it or paged about it, unless—

Tim: Right.

Corey: —you have something else in the business that will, like a boss that screams at you two o’clock in the morning. This is why we have very little that’s production-facing here.

Tim: But if I know that alert exists somewhere in the console, that’s easy for me to write a trap for. That’s easy for me to write, say hey, I’m going to respond to that because this call is going to come out somewhere; it’s going to get emitted somewhere. I can now, as an engineer, write a very easy trap that says, “Hey, pop this in the Slack. Send an alert. Send a page.”

So, if I could emit a cost metric, and I could say, “Wow. Somebody has spun up this thing that’s going to cost X amount of money. Someone should get paged about this.” Because if they don’t page about this and we wait eight hours, that’s my month’s salary. And you would do that if your database server went down; you would do that if someone rooted that database server; you would do that if the database server was [bogging 00:33:48] you to scale up another one. So, why can’t you do that if that database server was all of sudden costing you way more than you had calculated?

Corey: And there’s a lot of nuance here because what you’re talking about makes perfect sense for smaller-scale accounts, but even some of the very large accounts where we’re talking hundreds of millions a year in spend, you can set compromised keys up on GitHub, put them in Payspin, whatever, and then people start spinning up Bitcoin miners everywhere. Great. It takes a long time to materially move the needle on that level of spend; it gets lost in the background noise. I lose my mind when I wind up leaving a managed NAT gateway running and it cost me 70 bucks a month in my $5 a month test account. Yeah, but you realize you could basically buy an island and it gets lost in the AWS bill at some of the high watermarks for some of these larger accounts.

“Oh, someone spun up a cluster that’s going to cost $400,000 a year?” Yeah, do I need to re-explain to you what a data science team does? They light money on fire in return for questionable returns, as a general rule. You knew that when you hired them; leave them alone. Whereas someone in their developer account does this, yeah, you kind of want to flag that immediately.

It always comes down to rules and context. But I’d love to have some templates ready to go of, “I’m a starving student, please alert me anytime it looks like I might possibly exceed the free tier,” or better yet, “Don’t let me, and if I do, it’s on you and you eat the cost.” Conversely, it’s, “Yeah, this is a Netflix sub-account or whatnot. Maybe don’t bother me for anything whatsoever because freedom and responsibility is how we roll.” I imagine that’s what they do internally on a lot of their cloud costing stuff because freedom and responsibility is ingrained in their culture. It’s great. It’s the freedom from having to think about cloud bills and the responsibility for paying it, of the cloud bill.

Tim: Yeah, we will get internally alerted if things are [laugh] up too long, and then we will actually get paged, and then our manager would get paged, [laugh] and it would go up the line. If you leave something that’s running too expensive, too long. So, there is a system there for it.

Corey: Oh, yeah. The internal AWS systems for employees are probably my least favorite AWS service, full stop. And I’ve seen things posted about it; I believe it’s called Isengard, for spinning up internal accounts and the rest—there’s a separate one, I think, called Conduit, but I digress—that you spin something up, and apparently if it doesn’t wind up—I don’t need you to comment on this because you worked there and confidentiality is super important, but to my understanding it’s, great, it has a whole bunch of formalized stuff like that and it solves for a whole lot of nifty features that bias for the way that AWS focuses on accounts and how they’ve view security and the rest. And, “Oh, well, we couldn’t possibly ship this to customers because it’s not how they operate.” And that’s great.

My problem with this internal provisioning system is it isolates and insulates AWS employees from the real pain of working with multiple accounts as a customer. You don’t have to deal with the provisioning process of Control Tower or whatnot; you have your own internal thing. Eat your own dog food, gargle your own champagne, whatever it takes to wind up getting exposure to the pain that hits customers and suddenly you’ll see those things improve. I find that the best way to improve a product is to make the people building it live with the painful parts.

Tim: I think it’s interesting that the stance is, “Well, it’s not how the customers operate, and we wouldn’t want the customers to have to deal with this.” But at the same time, you have to open up, like, 100 accounts if you need more than a certain number of S3 buckets. So, they are very comfortable with burdening the customer with a lot of constraints, and they say, “Well, constraints drive innovation.” Certainly, this is a constraint that you could at least offer and let the customers innovate around that.

Corey: And at least define who the customer is. Because yeah, “I’m a Netflix sub-account is one story,” “I’m a regulated bank,” is another story, and, “I’m a student in my dorm room, trying to learn how this whole cloud thing works,” is another story. From risk tolerance, from a data protection story, from a billing surprise story, from a, “I’m trying to learn what the hell this is, and all these other service offerings you keep talking to me about confuse the hell out of me; please streamline the experience.” There’s a whole universe of options and opportunity that isn’t being addressed here.

Tim: Well, I will say it very simply like this: we’re talking about a multi-trillion dollar company versus someone who, if their AWS bill is too high, they don’t pay rent; maybe they don’t eat; maybe they have other issues, they don’t—medical bill doesn’t get paid; child care doesn’t get paid. And if you’re going to tell me that this multi-trillion dollar company can’t solve for that so that doesn’t happen to that person and tells them, “Well, if you come in afterwards, after your bill gets there, maybe we can do something about it, but in the meantime, suffer through this.” That’s not ethical. Full stop.

Corey: There are a lot of things that AWS gets right, and I want to be clear that I’m not sitting here trying to cast blame and say that everything they’re doing is terrible. I feel like every time I talk about billing in any depth, I have to throw this disclaimer in. Ninety to ninety-five percent of what they do is awesome. It’s just the missing piece that is incredibly painful for customers, and that’s what I spend most of my time focusing on. It should not be interpreted to think that I hate the company.

I just want them to do better than they are, and what they’re doing now is pretty decent in most respects. I just want to fix the painful parts. Tim, thank you for joining me for a third time here. I’m certain I’ll have you back in the somewhat near future to talk about more aspects of this, but until then, where can people find you slash retain your services?

Tim: Well, you can find me on Twitter at @elchefe. If you want to retain my services for which you would be very, very happy to have, you can go to duckbillgroup.com and fill out a little questionnaire, and I will magically appear after an exchange of goods and services.

Corey: Make sure to reference Tim by name just so that we can make our sales team facepalm because they know what’s coming next. Tim, thank you so much for your time; it’s appreciated.

Tim: Thank you so much, Corey. I loved it.

Corey: Principal cloud economist here at The Duckbill Group, Tim Banks. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, wait at least eight hours—possibly as many as 48 to 72—and then leave a comment explaining what you didn’t like.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Emma

Emma Bostian is a Software Engineer at Spotify in Stockholm. She is also a co-host of the Ladybug Podcast, author of Decoding The Technical Interview Process, and an instructor at LinkedIn Learning and Frontend Masters.

Links:

  • Ladybug Podcast: https://www.ladybug.dev
  • LinkedIn Learning: https://www.linkedin.com/learning/instructors/emma-bostian
  • Frontend Masters: https://frontendmasters.com/teachers/emma-bostian/
  • Decoding the Technical Interview Process: https://technicalinterviews.dev
  • Twitter: https://twitter.com/emmabostian

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Jellyfish. So, you’re sitting in front of your office chair, bleary eyed, parked in front of a powerpoint and—oh my sweet feathery Jesus its the night before the board meeting, because of course it is! As you slot that crappy screenshot of traffic light colored excel tables into your deck, or sift through endless spreadsheets looking for just the right data set, have you ever wondered, why is it that sales and marketing get all this shiny, awesome analytics and inside tools? Whereas, engineering basically gets left with the dregs. Well, the founders of Jellyfish certainly did. That’s why they created the Jellyfish Engineering Management Platform, but don’t you dare call it JEMP! Designed to make it simple to analyze your engineering organization, Jellyfish ingests signals from your tech stack. Including JIRA, Git, and collaborative tools. Yes, depressing to think of those things as your tech stack but this is 2021. They use that to create a model that accurately reflects just how the breakdown of engineering work aligns with your wider business objectives. In other words, it translates from code into spreadsheet. When you have to explain what you’re doing from an engineering perspective to people whose primary IDE is Microsoft Powerpoint, consider Jellyfish. Thats Jellyfish.co and tell them Corey sent you! Watch for the wince, thats my favorite part.

Corey: This episode is sponsored in part by Liquibase. If you’re anything like me, you’ve screwed up the database part of a deployment so severely that you’ve been banned from touching every anything that remotely sounds like SQL, at at least three different companies. We’ve mostly got code deployments solved for, but when it comes to databases we basically rely on desperate hope, with a roll back plan of keeping our resumes up to date. It doesn’t have to be that way. Meet Liquibase. It is both an open source project and a commercial offering. Liquibase lets you track, modify, and automate database schema changes across almost any database, with guardrails to ensure you’ll still have a company left after you deploy the change. No matter where your database lives, Liquibase can help you solve your database deployment issues. Check them out today at liquibase.com. Offer does not apply to Route 53.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. One of the weird things that I’ve found in the course of, well, the last five years or so is that I went from absolute obscurity to everyone thinking that I know everyone else because I have thoughts and opinions on Twitter. Today, my guest also has thoughts and opinions on Twitter. The difference is that what she has to say is actually helpful to people. My guest is Emma Bostian, software engineer at Spotify, which is probably, if we can be honest about it, one of the least interesting things about you. Thanks for joining me.

Emma: Thanks for having me. That was quite the intro. I loved it.

Corey: I do my best and I never prepare them, which is a blessing and a curse. When ADHD is how you go through life and you suck at preparation, you’ve got to be good at improv. So, you’re a co-host of the Ladybug Podcast. Let’s start there. What is that podcast? And what’s it about?

Emma: So, that podcast is just my three friends and I chatting about career and technology. We all come from different backgrounds, have different journeys into tech. I went the quote-unquote, “Traditional” computer science degree route, but Ali is self-taught and works for AWS, and Kelly she has, like, a master’s in psychology and human public health and runs her own company. And then Sydney is an awesome developer looking for her next role. So, we all come from different places and we just chat about career in tech.

Corey: You’re also an instructor at LinkedIn Learning and Frontend Masters. I’m going to guess just based upon the name that you are something of a frontend person, which is a skill set that has constantly eluded me for 20 years, as given evidence by every time I’ve tried to build something that even remotely touches frontend or JavaScript in any sense.

Emma: Yeah, to my dad’s disdain, I have stuck with the frontend; he really wanted me to stay backend. I did an internship at IBM in Python, and you know, I learned all about assembly language and database, but frontend is what really captures my heart.

Corey: There’s an entire school of thought out there from a constituency of Twitter that I will generously refer to as shitheads that believe, “Oh, frontend is easy and it’s somehow less than.” And I would challenge anyone who holds that perspective to wind up building an interface that doesn’t look like crap first, then come and talk to me. Spoiler, you will not say that after attempting to go down that rabbit hole. If you disagree with this, you can go ahead and yell at me on Twitter so I know where you’re hiding, so I can block you. Now, that’s all well and good, but one of the most interesting things that you’ve done that aligns with topics near and dear to my heart is you wrote a book.

Now, that’s not what’s near and dear to my heart; I have the attention span to write a tweet most days. But the book was called Decoding the Technical Interview Process. Technical interviewing is one of those weird things that comes up from time to time, here and everywhere else because it’s sort of this stylized ritual where we evaluate people on a number of skills that generally don’t reflect in their day-to-day; it’s really only a series of skills that you get better by practicing, and you only really get to practice them when you’re interviewing for other jobs. That’s been my philosophy, but again, I’ve written a tweet on this; you’ve written a book. What’s the book about and what drove you to write it?

Emma: So, the book covers everything from an overview of the interview process, to how do you negotiate a job offer, to systems design, and talks about load balancing and cache partitioning, it talks about what skills you need from the frontend side of things to do well on your JavaScript interviews. I will say this, I don’t teach HTML, CSS, and JavaScript in-depth in the book because there are plenty of other resources for that. And some guy got mad at me about that the other day and wanted a refund because I didn’t teach the skills, but I don’t need to. [laugh]. And then it covers data structures and algorithms.

They’re all written in JavaScript, they have easy to comprehend diagrams. What drove me to write this is that I had just accepted a job offer in Stockholm for a web developer position at Spotify. I had also just passed my Google technical interviews, and I finally realized, holy crap, maybe I do know what I’m doing in an interview now. And this was at the peak of when people were getting laid off due to COVID and I said, “You know what? I have a lot of knowledge. And if I have a computer science degree and I was able to get through some of the hardest technical interviews, I think I should share that with the community.”

Because some people didn’t go through a CS degree and don’t understand what a linked list is. And that’s not their fault. It’s just unfortunately, there weren’t a lot of great resources—especially for web developers out there—to learn these concepts. Cracking the Coding Interview is a great book, but it’s written in backend language and it’s a little bit hard to digest as a frontend developer. So, I decided to write my own.

Corey: How much of the book is around the technical interview process as far as ask, “Here’s how you wind up reversing linked lists,” or, “Inverting a binary tree,” or whatever it is where you’re tracing things around without using a pointer, how do you wind up detecting a loop in a recursive whatever it is—yeah, as you can tell, I’m not a computer science person at all—versus how much of it is, effectively, interview 101 style skills for folks who are even in non-technical roles could absorb?

Emma: My goal was, I wanted this to be approachable by anyone without extensive technical knowledge. So, it’s very beginner-friendly. That being said, I cover the basic data structures, talking about what traditional methods you would see on them, how do you code that, what does that look like from a visual perspective with fake data? I don’t necessarily talk about how do you reverse a binary tree, but I do talk about how do you balance it if you remove a node? What if it’s not a leaf node? What if it has children? Things like that.

It’s about [sigh] I would say 60/40, where 40% is coding and technical stuff, but maybe—eh, it’s a little bit closer to 50/50; it kind of depends. I do talk about the take-home assessment and tips for that. When I do a take-home assessment, I like to include a readme with things I would
have done if I had more time, or these are performance trade-offs that I made; here’s why. So, there’s a lot of explanation as to how you can improve your chances at moving on to the next round. So yeah, I guess it’s 50/50.

I also include a section on tips for hiring managers, how to create an inclusive and comfortable environment for your candidates. But it’s definitely geared towards candidates, and I would say it’s about 50/50 coding tech and process stuff.

Corey: One of the problems I’ve always had with this entire industry is it feels like we’re one of the only industries that does this, where we bring people in, and oh, you’ve been an engineer for 15 years at a whole bunch of companies I’ve recognized, showing career progression, getting promoted at some of them transitioning from high-level role to high-level role. “Great, we are so glad that you came in to interview. Now, up to the whiteboard, please, and implement FizzBuzz because I have this working theory that you don’t actually know how to code, and despite the fact that you’ve been able to fake your way through it at big companies for 15 years, I’m the one that’s going to catch you out with some sort of weird trivia question.” It’s this adversarial, almost condescending approach and I don’t see it in any other discipline than tech. Is that just because I’m not well-traveled enough? Is that because I’m misunderstanding the purpose of all of these things? Or, what is this?

Emma: I think partially it was a gatekeeping solution for a while, for people who are comfortable in their roles and may be threatened by people who have come through different paths to get to tech. Because software engineer used to be an accredited title that you needed a degree or certification to get. And in some countries it still is, so you’ll see this debate sometimes about calling yourself a software engineer if you don’t have that accreditation. But in this day and age, people go through boot camps, they can come from other industries, they can be self-taught. You don’t need a computer science degree, and I think the interview process has not caught up with that.

I will say [laugh] the worst interview I had was at IBM when I was already working there. I was already a web developer there, full-time. I was interviewing for a role, and I walked into the room and there were five guys sitting at a table and they were like, “Get up to the whiteboard.” It was for a web development job and they quizzed me about Java. And I was like, “Um, sir, I have not done Java since college.” And they were like, “We don’t care.”

Corey: Oh, yeah, coding on a whiteboard in front of five people who already know the answer—

Emma: Horrifying.

Corey: —during a—for them, it’s any given Tuesday, and for you, it is a, this will potentially determine the course that your career takes from this point forward. There’s a level of stress that goes into that never exists in our day-to-day of building things out.

Emma: Well, I also think it’s an artificial environment. And why, though? Like, why is this necessary? One of the best interviews I had was actually with Gatsby. It was for an open-source maintainer role, and they essentially let me try the product before I bought it.

Like, they let me try out doing the job. It was a paid process, they didn’t expect me to do it for free. I got to choose alternatives if I wanted to do one thing or another, answer one question or another, and this was such an exemplary process that I always bring it up because that is a modern interview process, when you are letting people try the position. Now granted, not everyone can do this, right? We’ve got parents, we’ve got people working two jobs, and not everyone can afford to take the time to try out a job.

But who can also afford a five-stage interview process that still warrants taking vacation days? So, I think at least—at the very least—pay your candidates if you can.

Corey: Oh, yeah. One of the best interviews I’ve ever had was at a company called Three Rings Design, which is now defunct, unfortunately, but it was fairly typical ops questions of, “Yeah, here’s an AWS account. Spin up a couple EC2 instances, load balance between them, have another one monitored. You know, standard op stuff. And because we don’t believe in asking people to work for free, we’ll pay you $300 upon completion of the challenge.”

Which, again, it’s not huge money for doing stuff like that, but it’s also, this shows a level of respect for my time. And instead of giving me a hard deadline of when it was due, they asked me, “When can we expect this by?” Which is a great question in its own right because it informs you about a candidate’s ability to set realistic deadlines and then meet them, which is one of those useful work things. And they—unlike most other companies I spoke with in that era—were focused on making it as accommodating for the candidate as possible. They said, “We’re welcome to interview you during the workday; we can also stay after hours and have a chat then, if that’s more convenient for your work schedule.”

Because they knew I was working somewhere else; an awful lot of candidates are. And they just bent over backwards to be as accommodating as possible. I see there’s a lot of debate these days in various places about the proper way to interview candidates. No take-home because biases for people who don’t have family obligations or other commitments outside of work hours. “Okay, great, so I’m going to come in interview during the day?” “No. That biases people who can’t take time off.” And, on some level, it almost seems to distill down to no one likes any way that there is of interviewing candidates, and figuring out a way that accommodates everyone is a sort of a fool’s errand. It seems like there is no way that won’t get you yelled at.

Emma: I think there needs to be almost like a choose your own adventure. What is going to set you up for success and also allow you to see if you want to even work that kind of a job in the first place? Because I thought on paper, open-source maintainer sounds awesome. And upon looking into the challenges, I’m like, “You know what? I think I’d hate this job.”

And I pulled out and I didn’t waste their time and they didn’t waste mine. So, when you get down to it, honestly, I wish I didn’t have to write this book. Did it bring me a lot of benefit? Yeah. Let’s not sugarcoat that. It allowed me to pay off my medical debt and move across a continent, but that being said, I wish that we were at a point in time where that did not need to exist.

Corey: One of the things that absolutely just still gnaws at me even years later, is I interviewed at Google twice, and I didn’t get an offer either time, I didn’t really pass their technical screen either time. The second one that really sticks out in my mind where it was, “Hey, write some code in a Google Doc while we watch remotely,” and don’t give you any context or hints on this. And just it was—the entire process was sitting there listening to them basically, like, “Nope, not what I’m thinking about. Nope, nope, nope.” It was… by the end of that conversation, I realized that if they were going to move forward—which they didn’t—I wasn’t going to because I didn’t want to work with people that were that condescending and rude.

And I’ve held by it; I swore I would never apply there again and I haven’t. And it’s one of those areas where, did I have the ability to do the job? I can say in hindsight, mostly. Were there things I was going to learn as I went? Absolutely, but that’s every job.

And I’m realizing as I see more and more across the ecosystem, that they were an outlier in a potentially good way because in so many other places, there’s no equivalent of the book that you have written that is given to the other side of the table: how to effectively interview candidates. People lose sight of the fact that it’s a sales conversation; it’s a two-way sale, they have to convince you to hire them, but you also have to convince them to work with you. And even in the event that you pass on them, you still want them to say nice things about you because it’s a small industry, all things considered. And instead, it’s just been awful.

Emma: I had a really shitty interview, and let me tell you, they have asked me subsequently if I would re-interview with them. Which sucks; it’s a product that I know and love, and I’ve talked about this, but I had the worst experience. Let me clarify, I had a great first interview with them, and I was like, “I’m just not ready to move to Australia.” Which is where the job was. And then they contacted me again a year later, and it was the worst experience of my life—same recruiter—it was the ego came out.

And I will tell you what, if you treat your candidates like shit, they will remember and they will never recommend people interview for you. [laugh]. I also wanted to mention about accessibility because—so we talked about, oh, give candidates the choice, which I think the whole point of an interview should be setting your candidates up for success to show you what they can do. And I talked with [Stephen 00:14:09]—oh, my gosh, I can’t remember his last name—but he is a quadriplegic and he types with a mouthstick. And he was saying he would go to technical interviews and they would not be prepared to set him up for success.

And they would want to do these pair programming, or, like, writing on a whiteboard. And it’s not that he can’t pair program, it’s that he was not set up for success. He needed a mouthstick to type and they were not prepared to help them with that. So, it’s not just about the commitment that people need. It’s also about making sure that you are giving candidates what they need to give the best interview possible in an artificial environment.

Corey: One approach that people have taken is, “Ah, I’m going to shortcut this and instead of asking people to write code, I’m going to look at their work on GitHub.” Which is, in some cases, a great way to analyze what folks are capable of doing. On the other, well, there’s a lot of things that play into that. What if they’re working in environment where they don’t have the opportunity to open-source their work? What if people consider this a job rather than an all-consuming passion?

I know, perish the thought. We don’t want to hire people like that. Grow up. It’s not useful, and it’s not helpful. It’s not something that applies universally, and there’s an awful lot of reasons why someone’s code on GitHub might be materially better—or worse—than their work product. I think that’s fine. It’s just a different path toward it.

Emma: I don’t use GitHub for largely anything except just keeping repositories that I need. I don’t actively update it. And I have, like, a few thousand followers; I’m like, “Why the hell do you guys follow me? I don’t do anything.” It’s honestly a terrible representation.

That being said, you don’t need to have a GitHub repository—an active one—to showcase your skills. There are many other ways that you can show a potential employer, “Hey, I have a lot of skills that aren’t necessarily showcased on my resume, but I like to write blogs, I like to give tech talks, I like to make YouTube videos,” things of that nature.

Corey: I had a manager once who refused to interview anyone who didn’t have a built-out LinkedIn profile, which is also one of these bizarre things. It’s, yeah, a lot of people don’t feel the need to have a LinkedIn profile, and that’s fine. But the idea that, “Oh, yeah, they have this profile they haven’t updated in a couple years, it’s clearly they’re not interested in looking for work.” It’s, yeah. Maybe—just a thought here—your ability to construct a resume and build it out in the way that you were expecting is completely orthogonal to how effective they might be in the role. The idea that someone not having a LinkedIn profile somehow implies that they’re sketchy is the wrong lesson to take from all of this. That site is terrible.

Emma: Especially when you consider the fact that LinkedIn is primarily used in the United States as a social—not social networking—professional networking tool. In Germany, they use Xing as a platform; it’s very similar to LinkedIn, but my point is, if you’re solely looking at someone’s LinkedIn as a representation of their ability to do a job, you’re missing out on many candidates from all over the world. And also those who, yeah, frankly, just don’t—like, they have more important things to be doing than updating their LinkedIn profile. [laugh].

Corey: On some level, it’s the idea of looking at a consultant, especially independent consultant type, when their website is glorious and up-to-date and everything’s perfect, it’s, oh, you don’t really have any customers, do you? As opposed to the consultants you know who are effectively sitting there with a waiting list, their website looks like crap. It’s like, “Is this Geocities?” No. It’s just that they’re too busy working on the things that bring the money instead of the things that bring in business, in some respects.

Let’s face it, websites don’t. For an awful lot of consulting work, it’s word of mouth. I very rarely get people finding me off of Google, clicking a link, and, “Hey, my AWS bill is terrible. Can you help us with it?” It happens, but it’s not something that happens so frequently that we want to optimize for it because that’s not where the best customers have been coming from. Historically, it’s referrals, it’s word of mouth, it’s people seeing the aggressive shitposting I engage in on Twitter and saying, “Oh, that’s someone that should help me with my Amazon bill.” Which I don’t pretend to understand, but I’m still going to roll with it.

Emma: You had mentioned something about passion earlier, and I just want to say, if you’re a hiring manager or recruiter, you shouldn’t solely be looking at candidates who superficially look like they’re passionate about what they do. Yes, that is—it’s important, but it’s not something that—like, I don’t necessarily choose one candidate over the other because they push commits, and open pull requests on GitHub, and open-source, and stuff. You can be passionate about your job, but at the end of the day, it’s still a job. For me, would I be working if I had to? No. I’d be opening a bookstore because that’s what I would really love to be doing. But that doesn’t mean I’m not passionate about my job. I just show it in different ways. So, just wanted to put that out there.

Corey: Oh, yeah. The idea that you must eat, sleep, live, and breathe is—hell with that. One of the reasons that we get people to work here at The Duckbill Group is, yeah, we care about getting the job done. We don’t care about how long it takes or when you work; it’s oh, you’re not feeling well? Take the day off.

We have very few things that are ‘must be done today’ style of things. Most of those tend to fall on me because it’s giving a talk at a conference; they will not reschedule the conference for you. I’ve checked. So yeah, that’s important, but that’s not most days.

Emma: Yeah. It’s like programming is my job, it’s not my identity. And it’s okay if it is your primary hobby if that is how you identify, but for me, I’m a person with actual hobbies, and, you know, a personality, and programming is just a job for me. I like my job, but it’s just a job.

Corey: And on the side, you do interesting things like wrote a book. You mentioned earlier that it wound up paying off some debt and helping
cover your move across an ocean. Let’s talk a little bit about that because I’m amenable to the idea of side projects that accidentally have a way of making money. That’s what this podcast started out as. If I’m being perfectly honest, and started out as something even more self-serving than that.

It’s, well if I reach out to people in this industry that are doing interesting things and ask them to grab a cup of coffee, they’ll basically block me, whereas if I ask them to, would you like to appear on my podcast, they’ll clear time on their schedule. I almost didn’t care if my microphone was on or not when I was doing these just because it was a chance to talk to really interesting people and borrow their brain, people reached out asking they can sponsor it, along with the newsletter and the rest, and it’s you want to give me money? Of course, you can give me money. How much money? And that sort of turned into a snowball effect over time.

Five years in, it’s turned into something that I would never have predicted or expected. But it’s weird to me still, how effective doing something you’re actually passionate about as a side project can sort of grow wings on its own. Where do you stand on that?

Emma: Yeah, it’s funny because with the exception of the online courses that I’ve worked with—I mentioned LinkedIn Learning and Frontend Masters, which I knew were paid opportunities—none of my side projects started out for financial reasonings. The podcast that we started was purely for fun, and the sponsors came to us. Now, I will say right up front, we all had pretty big social media followings, and my first piece of advice to anyone looking to get into side projects is, don’t focus so much on making money at the get-go. Yes, to your point, Corey, focus on the stuff you’re passionate about. Focus on engaging with people on social media, build up your social media, and at that point, okay, monetization will slowly find its way to you.

But yeah, I say if you can monetize the heck out of your work, go for it. But also, free content is also great. I like to balance my paid content with my free content because I recognize that not everyone can afford to pay for some of this information. So, I generally always have free alternatives. And for this book that we published, one of the things that was really important to me was keeping it affordable.

The first publish I did was $10 for the book. It was like a 250-page book. It was, like, $10 because again, I was not in it for the money. And when I redid the book with the egghead.io team, the same team that did Epic React with Kent C. Dodds, I said, “I want to keep this affordable.” So, we made sure it was still affordable, but also that we had—what’s it called? Parity pricing? Pricing parity, where depending on your geographic location, the price is going to accommodate for how the currency is doing. So, yes, I would agree. Side project income for me allows me to do incredible stuff, but it wasn’t why I got into it in the first place. It was genuinely just a nice-to-have.

Corey: I haven’t really done anything that asks people for money directly. I mean, yeah, I sell t-shirts on the website, and mugs, and drink umbrellas—don’t get me started—but other than that and the charity t-shirt drive I do every year, I tend to not be good at selling things that don’t have a comma in the price tag. For me, it was about absolutely building an audience. I tend to view my Twitter follower count as something of a proxy for it, but the number I actually care about, the audience that I’m focused on cultivating, is newsletter subscribers because no social media platform that we’ve ever seen has lasted forever. And I have to imagine that Twitter will one day wane as well.

But email has been here since longer than we’d been alive, and by having a list of email addresses and ways I can reach out to people on an ongoing basis, I can monetize that audience in a more direct way, at some point should I need them to. And my approach has been, well, one, it’s a valuable audience for some sponsors, so I’ve always taken the asking corporate people for money is easier than asking people for personal money, plus it’s a valuable audience to them, so it tends to blow out a number of the metrics that you would normally expect of, oh, for this audience size, you should generally be charging Y dollars. Great. That makes sense if you’re slinging mattresses or free web hosting, but when it’s instead, huh, these people buy SaaS enterprise software and implement it at their companies, all of economics tend to start blowing apart. Same story with you in many respects.

The audience that you’re building is functionally developers. That is a lucrative market for the types of sponsors that are wise enough to understand that—in a lot of cases these days—which product a company is going to deploy is not dictated by their exec so much as it is the bottom-up adoption path of engineers who like the product.

Emma: Mm-hm. Yeah, and I think once I got to maybe around 10,000 Twitter followers is when I changed my mentality and I stopped caring so much about follower count, and instead I just started caring about the people that I was following. And the number is a nice-to-have but to be honest, I don’t think so much about it. And I do understand, yes, at that point, it is definitely a privilege that I have this quote-unquote, “Platform,” but I never see it as an audience, and I never think about that “Audience,” quote-unquote, as a marketing platform. But it’s funny because there’s no right or wrong. People will always come to you and be like, “You shouldn’t monetize your stuff.” And it’s like—

Corey: “Cool. Who’s going to pay me then? Not you, apparently.”

Emma: Yeah. It’s also funny because when I originally sold the book, it was $10 and I got so many people being like, “This is way too cheap. You should be charging more.” And I’m like, “But I don’t care about the money.” I care about all the people who are unemployed and not able to survive, and they have families, and they need to get a job and they don’t know how.

That’s what I care about. And I ended up giving away a lot of free books. My mantra was like, hey if you’ve been laid off, DM me. No questions asked, I’ll give it to you for free. And it was nice because a lot of people came back, even though I never asked for it, they came back and they wanted to purchase it after the fact, after they’d gotten a job.

And to me that was like… that was the most rewarding piece. Not getting their money; I don’t care about that, but it was like, “Oh, okay. I was actually able to help you.” That is what’s really the most rewarding. But yeah, certainly—and back really quickly to your email point, I highly agree, and one of the first things that I would recommend to anyone looking to start a side product, create free content so that you have a backlog that people can look at to… kind of build trust.

Corey: Give it away for free, but also get emails from people, like a trade for that. So, it’s like, “Hey, here’s a free guide on how to start a podcast from scratch. It’s free, but all I would like is your email.” And then when it comes time to publish a course on picking the best audio and visual equipment for that podcast, you have people who’ve already been interested in this topic that you can now market to.

This episode is sponsored by our friends at Oracle Cloud. Counting the pennies, but still dreaming of deploying apps instead of "Hello, World" demos? Allow me to introduce you to Oracle's Always Free tier. It provides over 20 free services and infrastructure, networking databases, observability, management, and security.

And - let me be clear here - it's actually free. There's no surprise billing until you intentionally and proactively upgrade your account. This means you can provision a virtual machine instance or spin up an autonomous database that manages itself all while gaining the networking load, balancing and storage resources that somehow never quite make it into most free tiers needed to support the application that you want to build.

With Always Free you can do things like run small scale applications, or do proof of concept testing without spending a dime. You know that I always like to put asterisks next to the word free. This is actually free. No asterisk. Start now. Visit https://snark.cloud/oci-free that's https://snark.cloud/oci-free.

Corey: I’m not sitting here trying to judge anyone for the choices that they make at all. There are a lot of different paths to it. I’m right there with you. One of the challenges I had when I was thinking about, do I charge companies or do I charge people was that if I’m viewing it through a lens of audience growth, well, what stuff do I gate behind a paywall? What stuff don’t I? Well, what if I just—

Emma: Mm-hm.

Corey: —gave it all away? And that way I don’t have to worry about the entire class of problems that you just alluded to of, well, how do I make sure this is fair? Because a cup of coffee in San Francisco is, what, $14 in some cases? Whereas that is significant in places that aren’t built on an economy of foolishness. How do you solve for that problem? How do you deal with the customer service slash piracy issues slash all the other nonsense? And it’s just easier.

Emma: Yeah.

Corey: Something I’ve found, too, is that when you’re charging enough money to companies, you don’t have to deal with an entire class of customer service problem. You just alluded to the other day that well, you had someone who bought your book and was displeased that it wasn’t a how to write code from scratch tutorial, despite the fact that he were very clear on what it is and what it isn’t. I don’t pretend to understand that level of entitlement. If I spend 10 or 20 bucks on an ebook, and it’s not very good, let’s see, do I wind up demanding a refund from the author and making them feel bad about it, or do I say, “The hell with it.” And in my case, I—there is privilege baked into this; I get that, but it’s I don’t want to make people feel bad about what they’ve built. If I think there’s enough value to spend money on it I view that as a one-way transaction, rather than chasing someone down for three months, trying to get a $20 refund.

Emma: Yeah, and I think honestly, I don’t care so much about giving refunds at all. We have a 30-day money-back guarantee and we don’t ask any questions. I just asked this person for feedback, like, “Oh, what was not up to par?” And it was just, kind of like, BS response of like, “Oh, I
didn’t read the website and I guess it’s not what I wanted.” But the end of the day, they still keep the product.

The thing is, you can’t police all of the people who are going to try to get your content for free if you’re charging for it; it’s part of it. And I knew that when I got into it, and honestly, my thing is, if you are circulating a book that helps you get a job in tech and you’re sending it to all your friends, I’m not going to ask any questions because it’s very much the sa—and this is just my morals here, but if I saw someone stealing food from a grocery store, I wouldn’t tell on them because at the end of the day, if you’re s—

Corey: Same story. You ever see someone’s stealing baby formula from a store? No, you didn’t.

Emma: Right.

Corey: Keep walking. Mind your business.

Emma: Exactly. Exactly. So, at the end of the day, I didn’t necessarily care that—people are like, “Oh, people are going to share your book around. It’s a PDF.” I’m like, “I don’t care. Let them. It is what it is. And the people who wants to support and can, will.” But I’m not asking.

I still have free blogs on data structures, and algorithms, and the interview stuff. I do still have content for free, but if you want more, if you want my illustrated diagrams that took me forever with my Apple Pencil, fair enough. That would be great if you could support me. If not, I’m still happy to give you the stuff for free. It is what it is.

Corey: One thing that I think is underappreciated is that my resume doesn’t look great. On paper, I have an eighth-grade education, and I don’t have any big tech names on my resume. I have a bunch of relatively short stints; until I started this place, I’ve never lasted more than two years anywhere. If I apply through the front door the way most people do for a job, I will get laughed out of the room by the applicant tracking system, automatically. It’ll never see a human.

And by doing all these side projects, it’s weird, but let’s say that I shut down the company for some reason, and decide, ah, I’m going to go get a job now, my interview process—more or less, and it sounds incredibly arrogant, but roll with it for a minute—is, “Don’t you know who I am? Haven’t you heard of me before?” It’s, “Here’s my website. Here’s all the stuff I’ve been doing. Ask anyone in your engineering group who I am and you’ll see what pops up.”

You’re in that same boat at this point where your resume is the side projects that you’ve done and the audience you’ve built by doing it. That’s something that I think is underappreciated. Even if neither one of us made a dime through direct monetization of things that we did, the reputational boost to who we are and what we do professionally seems to be one of those things that pays dividends far beyond any relatively small monetary gain from it.

Emma: Absolutely, yeah. I actually landed my job interview with Spotify through Twitter. I was contacted by a design systems manager. And I was in the interview process for them, and I ended up saying, “You know, I’m not ready to move to Stockholm. I just moved to Germany.”

And a year later, I circled back and I said, “Hey, are there any openings?” And I ended up re-interviewing, and guess what? Now, I have a beautiful home with my soulmate and we’re having a child. And it’s funny how things work out this way because I had a Twitter account. And so don’t undervalue [laugh] social media as a tool in lieu of a resume because I don’t think anyone at Spotify even saw my resume until it actually accepted the job offer, and it was just a formality.

So yeah, absolutely. You can get a job through social media. It’s one of the easiest ways. And that’s why if I ever see anyone looking for a job on Twitter, I will retweet, and vouch for them if I know their work because I think that’s one of the quickest ways to finding an awesome candidate.

Corey: Back in, I don’t know, 2010, 2011-ish. I was deep in the IRC weed. I was network staff on the old freenode network—not the new terrible one. The old, good one—and I was helping people out with various things. I was hanging out in the Postfix channel and email server software thing that most people have the good sense not to need to know anything about.

And someone showed up and was asking questions about their config, and I was working with them, and teasing them, and help them out with it. And at the end of it, his comment was, “Wow, you’re really good at this. Any chance you’d be interested in looking for jobs?” And the answer was, “Well, sure, but it’s a global network. Where are you?”

Well, he was based in Germany, but he was working remotely for Spotify in Stockholm. A series of conversations later, I flew out to Stockholm and interviewed for a role that they decided I was not a fit for—and again, they’re probably right—and I often wonder how my life would have gone differently if the decision had gone the other way. I mean, no hard feelings, please don’t get me wrong, but absolutely, helping people out, interacting with people over social networks, or their old school geeky analogs are absolutely the sorts of things that change lives. I would never have thought to apply to a role like that if I had been sitting here looking at job ads because who in the world would pick up someone with relatively paltry experience and move them halfway around the world? This was like a fantasy, not a reality.

Emma: [laugh].

Corey: It’s the people you get to know—

Emma: Yeah.

Corey: —through these social interactions on various networks that are worth… they’re worth gold. There’s no way to describe it other than that.

Emma: Yeah, absolutely. And if you’re listening to this, and you’re discouraged because you got turned down for a job, we’ve all been there, first of all, but I remember being disappointed because I didn’t pass my first round of interviews of Google the first time I interviewed with them, and being, like, “Oh, crap, now I can’t move to Munich. What am I going to do with my life?” Well, guess what, look where I am today. If I had gotten that job that I thought was it for me, I wouldn’t be in the happiest phase of my life.

And so if you’re going through it—obviously, in normal circumstances where you’re not frantically searching for a job; if you’re in more of a casual life job search—and you’ve been let go from the process, just realize that there’s probably something bigger and better out there for you, and just focus on your networking online. Yeah, it’s an invaluable tool.

Corey: One time when giving a conference talk, I asked, “All right, raise your hand if you have never gone through a job interview process and then not been offered the job.” And a few people did. “Great. If your hand is up, aim higher. Try harder. Take more risks.”

Because fundamentally, job interviews are two-way streets and if you are only going for the sure thing jobs, great, stretch yourself, see what else is out there. There’s no perfect attendance prize. Even back in school there wasn’t. It’s the idea of, “Well, I’ve only ever taken the easy path because I don’t want to break my streak.” Get over it. Go out and interview more. It’s a skill, unlike most others that you don’t get to get better at unless you practice it.

So, you’ve been in a job for ten years, and then it’s time to move on—I’ve talked to candidates like this—their interview skills are extremely rusty. It takes a little bit of time to get back in the groove. I like to interview every three to six months back when I was on the job market. Now that I, you know, own the company and have employees, it looks super weird if I do it, but I miss it. I miss those conversations. I miss the aspects—

Emma: Yes.

Corey: —of exploring what the industry cares about.

Emma: Absolutely. And don’t underplay the importance of studying the foundational language concepts. I see this a lot in candidates where they’re so focused on the newest and latest technologies and frameworks, that they forgot foundational JavaScript, HTML, and CSS. Many companies are focused primarily on these plain language concepts, so just make sure that when you are ready to get back into interviewing and enhance that skill, that you don’t neglect the foundation languages that the web is built on if you’re a web developer.

Corey: I’d also take one last look around and realize that every person you admire, every person who has an audience, who is a known entity in the space only has that position because someone, somewhere did them a favor. Probably lots of someones with lots of favors. And you can’t ever pay those favors back. All you can do is pay it forward. I repeatedly encourage people to reach out to me if there’s something I can do to help. And the only thing that surprises me is how few people in the audience take me up on that. I’m talking to you, listener. Please, if I can help you with something, please reach out. I get a kick out of doing that sort of thing.

Emma: Absolutely. I agree.

Corey: Emma, thank you so much for taking the time to speak with me today. If people want to learn more, where can they find you?

Emma: Well, you can find me on Twitter. It’s just @EmmaBostian, I’m, you know, shitposting over there on the regular. But sometimes I do tweet out helpful things, so yeah, feel free to engage with me over there. [laugh].

Corey: And we will, of course, put a link to that in the [show notes 00:35:42]. Thank you so much for taking the time to speak with me today. I appreciate it.

Emma: Yeah. Thanks for having me.

Corey: Emma Bostian, software engineer at Spotify and oh, so very much more. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an incoherent ranting comment mentioning that this podcast as well failed to completely teach you JavaScript.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Jason

Jason is now the Managing Director at Redpoint Ventures.

Links:

  • GitHub: https://github.com/
  • @jasoncwarner: https://twitter.com/jasoncwarner
  • GitHub: https://github.com/jasoncwarner
  • Jasoncwarner/ama: https://github.com/jasoncwarner/ama

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Honeycomb. When production is running slow, it's hard to know where problems originate: is it your application code, users, or the underlying systems? I’ve got five bucks on DNS, personally. Why scroll through endless dashboards, while dealing with alert floods, going from tool to tool to tool that you employ, guessing at which puzzle pieces matter? Context switching and tool sprawl are slowly killing both your team and your business. You should care more about one of those than the other, which one is up to you. Drop the separate pillars and enter a world of getting one unified understanding of the one thing driving your business: production. With Honeycomb, you guess less and know more. Try it for free at Honeycomb.io/screaminginthecloud. Observability, it’s more than just hipster monitoring.

Corey: This episode is sponsored in part by Liquibase. If you’re anything like me, you’ve screwed up the database part of a deployment so severely that you’ve been banned from touching every anything that remotely sounds like SQL, at at least three different companies. We’ve mostly got code deployments solved for, but when it comes to databases we basically rely on desperate hope, with a roll back plan of keeping our resumes up to date. It doesn’t have to be that way. Meet Liquibase. It is both an open source project and a commercial offering. Liquibase lets you track, modify, and automate database schema changes across almost any database, with guardrails to ensure you’ll still have a company left after you deploy the change. No matter where your database lives, Liquibase can help you solve your database deployment issues. Check them out today at liquibase.com. Offer does not apply to Route 53.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’m joined this week by Jason Warner, the Chief Technology Officer at GifHub, although he pronounces it differently. Jason, welcome to the show.

Jason: Thanks, Corey. Good to be here.

Corey: So, GitHub—as you insist on pronouncing it—is one of those companies that’s been around for a long time. In fact, I went to a training conducted by one of your early folks, Scott Chacon, who taught how Git works over the course of a couple of days, and honestly, I left more confused than I did when I entered. It’s like, “Oh, this is super awful. Good thing I’ll never need to know this because I’m not really a developer.” And I’m still not really a developer and I still don’t really know how Git works, but here we are.

And it’s now over a decade later; you folks have been acquired by Microsoft, and you are sort of the one-stop-shop, from the de facto perspective of, “I’m going to go share some code with people on the internet. I’ll use GitHub to do it.” Because, you know, copying and pasting and emailing Microsoft Word documents around isn’t ideal.

Jason: That is right. And I think that a bunch of things that you mentioned there, played into, you know, GitHub’s early and sustained success. But my God, do you remember the old days when people had to email tar files around or drop them in weird spots?

Corey: What the hell do you mean, by, “Old days?” It still blows my mind that the Linux kernel is managed by—they use Git, obviously. Linus Torvalds did write Git once upon a time—and it has the user interface you would expect for that. And the way that they collaborate is not through GitHub or anything like that. No, they use Git to generate patches, which they then email to the mailing list. Which sounds like I’m making it up, like, “Oh, well, yeah, tell another one, but maybe involve a fax machine this time.” But no, that is actually what they do.

Jason: It blew my mind when I saw that, too, by the way. And you realize, too, that workflows are workflows, and people will build interesting workflows to solve their use case. Now, obviously, anyone that you would be talking to in 2021, if you walked in and said, “Yeah, install Git. Let’s set up an email server and start mailing patches to each other and we’re going to do it this way.” They would just kind of politely—or maybe impolitely—show you out of the room, and rightfully [laugh] so. But it works for one of the most important software projects in history: Linux.

Corey: Yeah, and it works almost in spite of itself to some extent. You’ve come a long way as a company because initially, it was, “Oh, there’s this amazing, decentralized version control system. How do we make it better? I know, we’re going to take off the decentralized part of it and give it a central point that everything can go through.” And collaboratively, it works well, but I think that viewing GitHub as a system that is used to sell free Git repositories to people is rather dramatically missing the point. It feels like it’s grown significantly beyond just code repository hosting. Tell me more about that.

Jason: Absolutely. I remember talking to a bunch of folks right around when I was joining GitHub, and you know, there was still talk about GitHub as, you know, GitHub for lawyers, or GitHub for doctors, or what could you do in a different way? And you know, social coding as an aspect, and maybe turning into a social network with a resume. And all those things are true to a percentage standpoint. But what GitHub should be in the world is the world’s most important software development platform, end-to-end software development platform.

We obviously have grown a bunch since me joining in that way which we launched dependency management packages, Actions with built-in CI, we’ve got some deployment mechanisms, we got advanced security underneath it, we’ve Codespaces in beta and alpha on top of it now. But if you think about GitHub as, join, share, and see other people’s code, that’s evolution one. If you see it as world’s largest, maybe most developed software development platform, that’s evolution two, and in my mind, its natural place where it should be, given what it has done already in the world, is become the world’s most important software company. I don’t mean the most profitable. I just mean the most
important.

Corey: I would agree. I had a blog post that went up somewhat recently about the future of cloud being Microsoft’s to lose. And it’s not because Azure is the best cloud platform out there, with respect, and I don’t need you to argue the point. It is very clearly not. It is not like other clouds, but I can see a path to where it could become far better than it is.

But if I’m out there and I’m just learning how to write code—because I make terrible life choices—and I go to a boot camp or I follow a tutorial online or I take a course somewhere, I’m going to be writing code probably using VS Code, the open-source editor that you folks launched after the acquisition. And it was pretty clear that Atom wasn’t quite where the world was going. Great. Then I’m going to host it on GitHub, which is a natural evolution. Then you take a look at things like GitHub Actions that build in CI/CD pipelines natively.

All that’s missing is a ‘Deploy to Azure’ button that is the next logical step, and you’re mostly there for an awful lot of use cases. But you can’t add that button until Azure itself gets better. Done right, this has the potential to leave, effectively, every other cloud provider in the dust because no one can touch this.

Jason: One hundred percent. I mean, the obvious thing that any other cloud should be looking at with us—or should have been before the acquisition, looking at us was, “Oh, no, they could jump over us. They could stop our funnel.” And I used internal metrics when I was talking to them about partnership that led to the sale, which was I showed them more about their running business than they knew about themselves. I can tell them where they were stacked-ranked against each other, based on the ingress and egress of all the data on GitHub, you know, and various reactions to that in those meetings was pretty astounding.

And just with that data alone, it should tell you what GitHub would be capable of and what Azure would be capable of in the combination of those two things. I mean, you did mention the ‘Deploy to Azure’ button; this has been a topic, obviously, pre and post-acquisition, which is, “When is that coming?” And it was the one hard rule I set during the acquisition was, there will be no ‘Deploy to Azure’ button. Azure has to earn the right to get things deployed to, in my opinion. And I think that goes to what you’re saying is, if we put a ‘Deploy to Azure’ button on top of this and Azure is not ready for that, or is going to fail, ultimately, that looks bad for all of us. But if it earned the right and it gets better, and it becomes one of those, then, you know, people will choose it, and that is, to me, what we’re after.

Corey: You have to choose the moment because if you do it too soon, you’ll set the entire initiative back five years. Do it too late, and you get leapfrogged. There’s a golden window somewhere and finding it is going to be hard. And I think it’s pretty clear that the other hyperscalers in this space are learning, or have learned, that the next 10 years of cloud or 15 years of cloud or whatever they want to call it, and the new customers that are going to come are not the same as the customers that have built the first half of the business. And they’re trying to wrap their heads around that because a lot of where the growth is going to come from is established blue chips that are used to thinking in very enterprise terms.

And people think I’m making fun of them when I say this, but Microsoft has 40 years’ experience apologizing to enterprises for computer failures. And that is fundamentally what cloud is. It’s about talking computers to business executives because as much as we talk about builders, that is not the person at an established company with an existing IT estate, who gets to determine where $50 million a year in cloud-spend is going to go.

Jason: It’s [laugh] very, [laugh] very true. I mean, we’ve entered a different spot with cloud computing in the bell curve of adoption, and if you think that they will choose the best technology every time, well, history of computing is littered with better technologies that have failed because the distribution was better on one side. As you mentioned, Microsoft has 40 years, and I wager that Microsoft has the best sales organizations and the best enterprise accounts and, you know, all that sort of stuff, blah, blah, blah, on that side of the world than anyone in the industry. They can sell to enterprises better than almost anyone in the industry. And the other hyperscalers—there’s a reason why [TK 00:08:34] is running Google Cloud right now. And Amazon, classically, has been very, very bad assigned to the enterprises. They just happened to be the first mover.

Corey: In the early days, it was easy. You’d have an Amazon salesperson roll up to a company, and the exec would say, “Great, why should we consider running things on AWS?” And the answer was, “Oh, I’m sorry, wrong conversation. Right now you have 80 different accounts scattered throughout your org. I’m just here to help you unify them, get some visibility into it, and possibly give you a discount along the way.” And it was a different conversation. Shadow IT was the sole driver of cloud adoption for a long time. That is no longer true. It has to go in the front door, and that is a fundamental shift in how you go to market.

Jason: One hundred percent true, and it’s why I think that Microsoft has been so successful with Azure, in the last, let’s call it five years in that, is that the early adopters in the second wave are doing that; they’re all enterprise IT, enterprise dev shops who are buying from the top down. Now, there is still the bottoms-up adoption that going to be happening, and obviously, bottom-up adoption will happen still going forward, but we’ve entered the phase where that’s not the primary or sole mechanism I should say. The sole mechanism of buying in. We have tops-down selling still—or now.

Corey: When Microsoft announced it was acquiring GitHub, there was a universal reaction of, “Oh, shit.” Because it’s Microsoft; of course they’re going to ruin GitHub. Is there a second option? No, unless they find a way to ruin it twice. And none of it came to pass.

It is uniformly excellent, and there’s a strong argument that could be made by folks who are unaware of what happened—I’m one of them, so maybe I’m right, maybe I’m wrong—that GitHub had a positive effect on Microsoft more than Microsoft had an effect on GitHub. I don’t know if that’s true or not, but I could believe it based upon what I’ve seen.

Jason: Obviously, the skepticism was well deserved at the time of acquisition, let’s just be honest with it, particularly given what Microsoft’s history had been for about 15—well, 20 years before, previous to Satya joining. And I was one of those people in the late ’90s who would write ‘M$’ in various forums. I was 18 or 19 years old, and just got into—

Corey: Oh, hating Microsoft was my entire personality.

Jason: [laugh]. And it was, honestly, well-deserved, right? Like, they had anti-competitive practices and they did some nefarious things. And you know, I talked about Bill Gates as an example. Bill Gates is, I mean, I don’t actually know how old he is, but I’m going to guess he’s late ’50s, early ’60s, but he’s basically in the redemption phase of his life for his early years.

And Microsoft is making up for Ballmer years, and later Gates years, and things of that nature. So, it was well-deserved skepticism, and particularly for a mid-career to older-career crowd who have really grown to hate Microsoft over that time. But what I would say is, obviously, it’s different under Satya, and Scott, and Amy Hood, and people like that. And all we really telling people is give us a chance on this one. And I mean, all of us. The people who were running GitHub at the time, including myself and, you know, let Scott and Satya prove that they are who they say they are.

Corey: It’s one of those things where there’s nothing you could have said that would have changed the opinion of the world. It was, just wait and see. And I think we have. It’s now, I daresay, gotten to a point where Microsoft announces that they’re acquiring some other beloved company, then people, I think, would extend a lot more credit than they did back then.

Jason: I have to give Microsoft a ton of credit, too, on this one for the way in which they handled acquisitions, like us and others. And the reason why I think it’s been so successful is also the reason why I think so many others die post-acquisition, which is that Microsoft has basically—I’ll say this, and I know I won’t get fired because it feels like it’s true. Microsoft is essentially a PE holding company at this point. It is acquired a whole bunch of companies and lets them run independent. You know, we got LinkedIn, you got Minecraft, Xbox is its own division, but it’s effectively its own company inside of it.

Azure is run that way. GitHub’s got a CEO still. I call it the archipelago model. Microsoft’s the landmass underneath the water that binds them all, and finance, and HR, and a couple of other things, but for the most part, we manifest our own product roadmap still. We’re not told what to
go do. And I think that’s why it’s successful. If we’re going to functionally integrate GitHub into Microsoft, it would have died very quickly.

Corey: You clearly don’t mix the streams. I mean, your gaming division writes a lot of interesting games and a lot of interesting gaming platforms. And, like, one of the most popularly played puzzle games in the world is a Microsoft property, and that is, of course, logging into a Microsoft account correctly. And I keep waiting for that to bleed into GitHub, but it doesn’t. GitHub is a terrific SAML provider, it is stupidly easy to log in, it’s great.

And at some level, I wish that would bleed into other aspects, but you can’t have everything. Tell me what it’s like to go through an acquisition from a C-level position. Because having been through an acquisition before, the process looks a lot like a surprise all-hands meeting one day after the markets close and, “Listen up, idiots.” And [laugh] there we go. I have to imagine with someone in your position, it’s a slightly different
experience.

Jason: It’s definitely very different for all C-levels. And then myself in particular, as the primary driver of the acquisition, obviously, I had very privy inside knowledge. And so, from my position, I knew what was happening the entire time as the primary driver from the inside. But even so, it’s still disconcerting to a degree because, in many ways, you don’t think you’re going to be able to pull it off. Like, you know, I remember the months, and the nights, and the weekends, and the weekend nights, and all the weeks I spent on the road trying to get all the puzzle pieces lined up for the Googles, or the Microsofts, or the eventually AWSs, the VMwares, the IBMs of the world to take seriously, just from a product perspective, which I knew would lead to, obviously, acquisition conversations.

And then, once you get the call from the board that says, “It’s done. We signed the letter of intent,” you basically are like, “Oh. Oh, crap. Okay, hang on a second. I actually didn’t—I don’t actually believe in my heart of hearts that I was going to actually be able to pull that off.” And so now, you probably didn’t plan out—or at least I didn’t. I was like, “Shit if we actually pulled this off what comes next?” And I didn’t have that what comes next, which is odd for me. I usually have some sort of a loose plan in place. I just didn’t. I wasn’t really ready for that.

Corey: It’s got to be a weird discussion, too, when you start looking at shopping a company around to be sold, especially one at the scale of GitHub because you’re at such a high level of visibility in the entire environment, where—it’s the idea of would anyone even want to buy us? And then, duh, of course they would. And you look the hyperscalers, for example. You have, well, you could sell it to Amazon and they could pull another Cloud9, where they shove it behind the IAM login process, fail to update the thing meaningfully over a period of years, to a point where even now, a significant portion of the audience listening to this is going to wonder if it’s a service I just made up; it sounds like something they might have done, but Cloud9 sounds way too inspired for an AWS service name, so maybe not. And—which it is real. You could go sell to Google, which is going to be awesome until some executive changes roles, and then it’s going to be deprecated in short order.

Or then there’s Microsoft, which is the wild card. It’s, well, it’s Microsoft. I mean, people aren’t really excited about it, but okay. And I don’t think that’s true anymore at all. And maybe I’m not being fair to all the hyperscalers there. I mean, I’m basically insulting everyone, which is kind of my shtick, but it really does seem that Microsoft was far and away the best acquirer possible because it has been transformative. My question—if you can answer it—is, how the hell did you see that beforehand? It’s only obvious—even knowing what I know now—in hindsight.

Jason: So, Microsoft was a target for me going into it, and the reason why was I thought that they were in the best overall position. There was enough humility on one side, enough hubris on another, enough market awareness, probably, organizational awareness to, kind of, pull it off. There’s too much hubris on one side of the fence with some of the other acquirers, and they would try to hug us too deeply, or integrate us too quickly, or things of that nature. And I think it just takes a deep understanding of who the players are and who the egos involved are. And I think egos has actually played more into acquisitions than people will ever admit.

What I saw was, based on the initial partnership conversations, we were developing something that we never launched before GitHub Actions called GitHub Launch. The primary reason we were building that was GitHub launches a five, six-year journey, and it’s got many, many different phases, which will keep launching over the next couple of years. The first one we never brought to market was a partnership between all of the clouds. And it served a specific purpose. One, it allowed me to get into the room with the highest level executive at every one of those companies.

Two allow me to have a deep economic conversation with them at a partnership level. And three, it allowed me to show those executives that we knew what GitHub’s value was in the world, and really flip the tables around and say, “We know what we’re worth. We know what our value is in the world. We know where we sit from a product influence perspective. If you want to be part of this, we’ll allow it.” Not, “Please come work with us.” It was more of a, “We’ll allow you to be part of this conversation.”

And I wanted to see how people reacted to that. You know how Amazon reacted that told me a lot about how they view the world, and how Google reacted to that showed me exactly where they viewed it. And I remember walking out of the Google conversation, feeling a very specific way based upon the reaction. And you know, when I talked to Microsoft, got a very different feel and it, kind of, confirmed a couple of things. And then when I had my very first conversation with Nat, who have known for a while before that, I realized, like, yep, okay, this is the one. Drive hard at this.

Corey: If you could do it all again, would you change anything meaningful about how you approached it?

Jason: You know, I think I got very lucky doing a couple of things. I was very intentional aspects of—you know, I tried to serendipitously show up, where Diane Greene was at one point, or a serendipitously show up where Satya or Scott Guthrie was, and obviously, that was all intentional. But I never sold a company like this before. The partnership and the product that we were building was obviously very intentional. I think if I were to go through the sale, again, I would probably have tried to orchestrate at least one more year independent.

And it’s not—for no other reason alone than what we were building was very special. And the world sees it now, but I wish that the people who built it inside GitHub got full credit for it. And I think that part of that credit gets diffused to saying, “Microsoft fixed GitHub,” and I want the people inside GitHub to have gotten a lot more of that credit. Microsoft obviously made us much better, but that was not specific to Microsoft because we’re run independent; it was bringing Nat in and helping us that got a lot of that stuff done. Nat did a great job at those things. But a lot of that was already in play with some incredible engineers, product people, and in particular our sales team and finance team inside of GitHub already.

Corey: When you take a look across the landscape of the fact that GitHub has become for a certain subset of relatively sad types of which I’m definitely one a household name, what do you think the biggest misconception about the company is?

Jason: I still think the biggest misconception of us is that we’re a code host. Every time I talk to the RedMonk folks, they get what we’re building and what we’re trying to be in the world, but people still think of us as SourceForge-plus-plus in many ways. And obviously, that may have been our past, but that’s definitely not where we are now and, for certain, obviously, not our future. So, I think that’s one. I do think that people still, to this day, think of GitLab as one of our main competitors, and I never have ever saw GitLab as a competitor.

I think it just has an unfortunate naming convention, as well as, you know, PRs, and MRs, and Git and all that sort of stuff. But we take very different views of the world in how we’re approaching things. And then maybe the last thing would be that what we’re doing at the scale that we’re doing it as is kind of easy. When I think that—you know, when you’re serving almost every developer in the world at this point at the scale at which we’re doing it, we’ve got some scale issues that people just probably will never thankfully encounter for themselves.

Corey: Well, everyone on Hacker News believes that they will, as soon as they put up their hello world blog, so Kubernetes is the only way to do anything now. So, I’m told.

Jason: It’s quite interesting because I think that everything breaks at scale, as we all know about from the [hyperclouds 00:20:54]. As we’ve learned, things are breaking every day. And I think that when you get advice, either operational, technical, or managerial advice from people who are running 10 person, 50 person companies, or X-size sophisticated systems, it doesn’t apply. But for whatever reason, I don’t know why, but people feel inclined to give that feedback to engineers at GitHub directly, saying, “If you just…” and in many [laugh] ways, you’re just like, “Well, I think that we’ll have that conversation at some point, you know, but we got a 100-plus-million repos and 65 million developers using us on a daily basis.” It’s a very different world.

Corey: This episode is sponsored by our friends at Oracle HeatWave is a new high-performance accelerator for the Oracle MySQL Database Service. Although I insist on calling it “my squirrel.” While MySQL has long been the worlds most popular open source database, shifting from transacting to analytics required way too much overhead and, ya know, work. With HeatWave you can run your OLTP and OLAP, don’t ask me to ever say those acronyms again, workloads directly from your MySQL database and eliminate the time consuming data movement and integration work, while also performing 1100X faster than Amazon Aurora, and 2.5X faster than Amazon Redshift, at a third of the cost. My thanks again to Oracle Cloud for sponsoring this ridiculous nonsense.

Corey: One of the things that I really appreciate personally because, you know, when you see something that company does, it’s nice to just thank people from time to time, so I’m inviting the entire company on the podcast one by one, at some point, to wind up thanking them all individually for it, but Codespaces is one of those things that I think is transformative for me. Back in the before times, and ideally the after times, whenever I travel the only computer I brought with me for a few years now has been an iPad or an iPad Pro. And trying to get an editor on that thing that works reasonably well has been like pulling teeth, my default answer has just been to remote into an EC2 instance and use vim like I have for the last 20 years. But Code is really winning me over. Having to play with code-server and other things like that for a while was obnoxious, fraught, and difficult.

And finally, we got to a point where Codespaces was launched, and oh, it works on an iPad. This is actually really slick. I like this. And it was the thing that I was looking for but was trying to have to monkey patch together myself from components. And that’s transformative.

It feels like we’re going back in many ways—at least in my model—to the days of thin clients where all the heavy lifting was done centrally on big computers, and the things that sat on people’s desks were mostly just, effectively, relatively simple keyboard, mouse, screen. Things go back and forth and I’m sure we’ll have super powerful things in our pockets again soon, but I like the interaction model; it solves for an awful lot of problems and that’s one of the things that, at least from my perspective, that the world may not have fully wrapped it head around yet.

Jason: Great observation. Before the acquisition, we were experimenting with a couple of different editors, that we wanted to do online editors. And same thing; we were experimenting with some Action CI stuff, and it just didn’t make sense for us to build it; it would have been too hard, there have been too many moving parts, and then post-acquisition, we really love what the VS Code team was building over there, and you could see it; it was just going to work. And we had this one person, well, not one person. There was a bunch of people inside of GitHub that do this, but this one person at the highest level who’s just obsessed with make this work on my iPad.

He’s the head of product design, his name’s Max, he’s an ex-Heroku person as well, and he was just obsessed with it. And he said, “If it works on my iPad, it’s got a chance to succeed. If it doesn’t work on my iPad, I’m never going to use this thing.” And the first time we booted up Codespaces—or he booted it up on the weekend, working on it. Came back and just, “Yep. This is going to be the one. Now, we got to work on those, the sanding the stones and those fine edges and stuff.”

But it really does unlock a lot for us because, you know, again, we want to become the software developer platform for everyone in the world, you got to go end-to-end, and you got to have an opinion on certain things, and you got to enable certain functionality. You mentioned Cloud9 before with Amazon. It was one of the most confounding acquisitions I’ve ever seen. When they bought it I was at Heroku and I thought, I thought at that moment that Amazon was going to own the next 50 years of development because I thought they saw the same thing a lot of us at Heroku saw, and with the Cloud9 acquisition, what they were going to do was just going to stomp on all of us in the space. And then when it didn’t happen, we just thought maybe, you know, okay, maybe something else changed. Maybe we were wrong about that assumption, too. But I think that we’re on to it still. I think that it just has to do with the way you approach it and, you know, how you design it.

Corey: Sorry, you just said something that took me aback for a second. Wait, you mean software can be designed? It’s not this emergent property of people building thing on top of thing? There’s actually a grand plan behind all these things? I’ve only half kidding, on some level, where if you take a look at any modern software product that is deployed into the world, it seems impossible for even small aspects of it to have been part of the initial founding design. But as a counterargument, it would almost have to be for a lot of these things. How do you square that circle?

Jason: I think you have to, just like anything on spectrums and timelines, you have to flex at various times for various things. So, if you think about it from a very, very simple construct of time, you just have to think of time horizons. So, I have an opinion about what GitHub should look like in 10 years—vaguely—in five years much more firmly, and then very, very concretely, for the next year, as an example. So, a lot of the features you might see might be more emergent, but a lot of long-term work togetherness has to be loosely tied together with some string. Now, that string will be tightened over time, but it loosely has to see its way through.

And the way I describe this to folks is that you don’t wake up one day and say, “I’m going on vacation,” and literally just throw a finger on the map. You have to have some sort of vague idea, like, “Hey, I want to have a beach vacation,” or, “I want to have an adventure vacation.” And then you can kind of pick a destination and say, “I’m going to Hawaii,” or, “I’m going to San Diego.” And if you’re standing on the East Coast knowing you’re going to San Diego, you basically know that you have to just start marching west, or driving west, or whatever. And now, you don’t have to have the route mapped out just yet, but you know that hey, if I’m going due southeast, I’m off course, so how do I reorient to make sure I’m still going in the right direction?

That’s basically what I think about as high-level, as scale design. And it’s not unfair to say that a lot of the stuff is not designed today. Amazon is very famous for not designing anything; they design a singular service. But there’s no cohesiveness to what Amazon—or AWS specifically, I should say, in this case—has put out there. And maybe that’s not what their strategy is. I don’t know the internal workings of them, but it’s very clear.

Corey: Well, oh, yeah. When I first started working in the AWS space and looking through the console, it like, “What is this? It feels like every service’s interface was designed by a different team, but that would—oh…” and then the light bulb went on. Yeah. You ship your culture.

Jason: It’s exactly it. It works for them, but I think if you’re going to try to do something very, very, very different, you know, it’s going to look a certain way. So, intentional design, I think, is part of what makes GitHub and other products like it special. And if you think about it, you have to have an end-to-end view, and then you can build verticals up and down inside of that. But it has to work on the horizontal, still.

And then if you hire really smart people to build the verticals, you get those done. So, a good example of this is that I have a very strong opinion about the horizontal workflow nature of GitHub should look like in five years. I have a very loose opinion about what the matrix build system of Actions looks like. Because we have very, very smart people who are working on that specific problem, so long as that maps back and snaps into the horizontal workflows. And that’s how it can work together.

Corey: So, when you look at someone who is, I don’t know, the CTO of a wildly renowned company that is basically still catering primarily to developers slash engineers, but let’s be honest, geeks, it’s natural to think that, oh, they must be the alpha geek. That doesn’t really apply to you from everything I’ve been able to uncover. Am I just not digging deeply enough, or are you in fact, a terrible nerd?

Jason: [laugh]. I am. I’m a terrible nerd. I am a very terrible nerd. I feel very lucky, obviously, to be in the position I’m in right now, in many ways,
and people call me up and exactly that.

It’s like, “Hey, you must be king of the geeks.” And I’m like, “[laugh], ah, funny story here.” But um, you know, I joke that I’m not actually supposed to be in tech in first place, the way I grew up, and where I did, and how, I wasn’t supposed to be here. And so, it’s serendipitous that I am in tech. And then turns out I had an aptitude for distributed systems, and complex, you know, human systems as well. But when people dig in and they start talking about topics, I’m confounded. I never liked Star Wars, I never like Star Trek. Never got an anime, board games, I don’t play video games—

Corey: You are going to get letters.

Jason: [laugh]. When I was at Canonical, oh, my goodness, the stuff I tried to hide about myself, and, like, learn, like, so who’s this Boba Fett dude. And, you know, at some point, obviously, you don’t have to pretend anymore, but you know, people still assume a bunch stuff because, quote, “Nerd” quote, “Geek” culture type of stuff. But you know, some interesting facts that people end up being surprised by with me is that, you know, I was very short in high school and I grew in college, so I decided that I wanted to take advantage of my newfound height and athleticism as you grow into your body. So, I started playing basketball, but I obsessed over it.

I love getting good at something. So, I’d wake up at four o’clock in the morning, and go shoot baskets, and do drills for hours. Well, I got really good at it one point, and I end up playing in a Pro-Am basketball game with ex-NBA Harlem Globetrotter legends. And that’s just not something you hear about in most engineering circles. You might expect that out of a salesperson or a marketing person who played pro ball—or amateur ball somewhere, or college ball or something like that. But not someone who ends up running the most important software company—from a technical perspective—in the world.

Corey: It’s weird. People counterintuitively think that, on some level, that code is the answer to all things. And that, oh, all this human interaction stuff, all the discussions, all the systems thinking, you have to fit a certain profile to do that, and anyone outside of that is, eh, they’re not as valuable. They can get ignored. And we see that manifesting itself in different ways.

And even if we take a look at people whose profess otherwise, we take a look at folks who are fresh out of a boot camp and don’t understand much about the business world yet; they have transformed their lives—maybe they’re fresh out of college, maybe didn’t even go to college—and 18 weeks later, they are signing up for six-figure jobs. Meanwhile, you take a look at virtually any other business function, in order to have a relatively comparable degree of earning potential, it takes years of experience and being very focused on a whole bunch of other things. There’s a massive distortion around technical roles, and that’s a strange and difficult thing to wrap my head around. But as you’re talking about it, it goes both ways, too. It’s the idea of, “Oh, I’ll become technical than branch into other things.” It sounded like you started off instead with a non-technical direction and then sort of adopted that from other sides. Is that right, or am I misremembering exactly how the story unfolds?

Jason: No, that’s about right. People say, “Hey, when did I start programming?” And it’s very in vogue, I think, for a lot of people to say, “I started programming at three years old,” or five years old, or whatever, and got my first computer. I literally didn’t get my first computer until I was 18-years-old. And I started programming when I got to a high school co-op with IBM at 17.

It was Lotus Notes programming at the time. Had no exposure to it before. What I did, though, in college was IBM told me at the time, they said, “If you get a computer science degree will guarantee you a job.” Which for a kid who grew up the way I grew up, that is manna from heaven type of deal. Like, “You’ll guarantee me a job inside where don’t have to dig ditches all day or lay asphalt? Oh, my goodness. What’s computer science? I’ll go figure it out.”

And when I got to school, what I realized was I was really far behind. Everyone was that ubergeek type of thing. So, what I did is I tried to hack the system, and what I said was, “What is a topic that nobody else has an advantage on from me?” And so I basically picked the internet because the internet was so new in the mid-’90s that most people were still not fully up to speed on it. And then the underpinnings in the internet, which basically become distributed systems, that’s where I started to focus.

And because no one had a real advantage, I just, you know, could catch up pretty quickly. But once I got into computers, it turned out that I was probably a very average developer, maybe even below average, but it was the system’s thinking that I stood out on. And you know, large-scale distributed systems or architectures were very good for me. And then, you know, that applies not, like, directly, but it applies decently well to human systems. It’s just, you know, different types of inputs and outputs. But if you think about organizations at scale, they’re barely just really, really, really complex and kind of irrational distributed systems.

Corey: Jason, thank you so much for taking the time to speak with me today. If people want to learn more about who you are, what you’re up to, how you think about the world, where can they find you?

Jason: Twitter’s probably the best place at this point. Just @jasoncwarner on Twitter. I’m very unimaginative. My name is my GitHub handle.
It’s my Twitter username. And that’s the best place that I, kind of, interact with folks these days. I do an AMA on GitHub. So, if you ever want to ask me anything, just kind of go to jasoncwarner/ama on GitHub and drop a question in one of the issues and I’ll get to answering that. Yeah, those are the best spots.

Corey: And we will, of course, include links to those things in the [show notes 00:33:52]. Thank you so much for taking the time to speak with me today. I really appreciate it.

Jason: Thanks, Corey. It’s been fun.

Corey: Jason Warner, Chief Technology Officer at GitHub. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review in your podcast platform of choice anyway, along with a comment that includes a patch.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Lauren

Lauren Hasson is the Founder of DevelopHer, an award-winning career development platform that has empowered thousands of women in tech to get ahead, stand out, and earn more in their careers. She also works full-time on the frontlines of tech herself. By day, she is an accomplished software engineer at a leading Silicon Valley payments company where she is the architect of their voice payment system and messaging capabilities and is chiefly responsible for all of application security.

Through DevelopHer, she’s partnered with top tech companies like Google, Dell, Intuit, Armor, and more and has worked with top universities including Indiana and Tufts to bridge the gender gap in leadership, opportunity, and pay in tech for good. Additionally, she was invited to the United Nations to collaborate on the global EQUALS initiative to bridge the global gender divide in technology.

Sought after across the globe for her insight and passionate voice, Lauren has started a movement that inspires women around the world to seek an understanding of their true value and to learn and continually grow.

Her work has been featured by industry-leading publications like IEEE Women in Engineering Magazine and Thrive Global and her ground-breaking platform has been recognized with fourteen prestigious awards for entrepreneurship, product innovation, diversity and leadership including the Women in IT Awards Silicon Valley Diversity Initiative of the Year Award, three Female Executive of the Year Awards, and recognition as a Finalist for the United Nations WSIS Stakeholder Prize.

Links:

  • DevelopHer: https://developher.com
  • The DevelopHer Playbook: https://www.amazon.com/DevelopHer-Playbook-Simple-Advocate-Yourself-ebook/dp/B08SQM4P5J

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: You could build you go ahead and build your own coding and mapping notification system, but it takes time, and it sucks! Alternately, consider Courier, who is sponsoring this episode. They make it easy. You can call a single send API for all of your notifications and channels. You can control the complexity around routing, retries, and deliverability and simplify your notification sequences with automation rules. Visit courier.com today and get started for free. If you wind up talking to them, tell them I sent you and watch them wince—because everyone does when you bring up my name. Thats the glorious part of being me. Once again, you could build your own notification system but why on god’s flat earth would you do that?

Corey: This episode is sponsored in part by Honeycomb. When production is running slow, it's hard to know where problems originate: is it your application code, users, or the underlying systems? I’ve got five bucks on DNS, personally. Why scroll through endless dashboards, while dealing with alert floods, going from tool to tool to tool that you employ, guessing at which puzzle pieces matter? Context switching and tool sprawl are slowly killing both your team and your business. You should care more about one of those than the other, which one is up to you. Drop the separate pillars and enter a world of getting one unified understanding of the one thing driving your business: production. With Honeycomb, you guess less and know more. Try it for free at Honeycomb.io/screaminginthecloud. Observability, it’s more than just hipster monitoring.

Corey: You could build you go ahead and build your own coding and mapping notification system, but it takes time, and it sucks! Alternately, consider Courier, who is sponsoring this episode. They make it easy. You can call a single send API for all of your notifications and channels. You can control the complexity around routing, retries, and deliverability and simplify your notification sequences with automation rules. Visit courier.com today and get started for free. If you wind up talking to them, tell them I sent you and watch them wince—because everyone does when you bring up my name. Thats the glorious part of being me. Once again, you could build your own notification system but why on god’s flat earth would you do that?

Corey: This episode is sponsored in part by our friends at Jellyfish. So, you’re sitting in front of your office chair, bleary eyed, parked in front of a powerpoint and—oh my sweet feathery Jesus its the night before the board meeting, because of course it is! As you slot that crappy screenshot of traffic light colored excel tables into your deck, or sift through endless spreadsheets looking for just the right data set, have you ever wondered, why is it that sales and marketing get all this shiny, awesome analytics and inside tools? Whereas, engineering basically gets left with the dregs. Well, the founders of Jellyfish certainly did. That’s why they created the Jellyfish Engineering Management Platform, but don’t you dare call it JEMP! Designed to make it simple to analyze your engineering organization, Jellyfish ingests signals from your tech stack. Including JIRA, Git, and collaborative tools. Yes, depressing to think of those things as your tech stack but this is 2021. They use that to create a model that accurately reflects just how the breakdown of engineering work aligns with your wider business objectives. In other words, it translates from code into spreadsheet. When you have to explain what you’re doing from an engineering perspective to people whose primary IDE is Microsoft Powerpoint, consider Jellyfish. Thats Jellyfish.co and tell them Corey sent you! Watch for the wince, thats my favorite part.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. A somewhat recurring theme of this show has been the business of cloud, and that touches on a lot of different things. One thing I’ve generally cognizant of not doing is talking to folks who don’t look like me and asking them questions like, “Oh, that’s great, but let’s ignore everything that you’re doing, and instead talk about what it’s like not to be a cis-gendered white dude in tech,” because that’s crappy. Today, we’re sort of deviating from that because my guest is Lauren Hasson, the founder of DevelopHer, which is a career development platform that empowers women in tech to get ahead. Lauren, thanks for joining me.

Lauren: Thanks so much for having me, Corey.

Corey: So, you’re the founder of DevelopHer, and that is ‘develop-her’ as in ‘she’. I’m not going to be as distinct on that pronunciation, so if you think I’m saying ‘developer’ and it doesn’t make intellectual sense, listener, that’s what’s going on. But you’re also a speaker, you’re an author, and you work on the front lines of tech yourself. That’s a lot of stuff. What’s your story?

Lauren: Yeah, I do. So, I’m not only the founder-developer, but I’m just like many of your listeners: I work on the front lines of tech myself. I work remotely from my home in Dallas for a Silicon Valley payments company, where I’m the architect of our voice payment system, and I up until recently was chiefly responsible for all of application security. Yeah, and I do keep busy.

Corey: It certainly seems like it. Let’s go back to, I guess, the headline item here. You are the founder of DevelopHer, and one thing that always drives me a little nutty is when people take a glance at what I do and then try and tell the story, and then effectively mess the whole thing up. What is DevelopHer?

Lauren: So, DevelopHer is what I wish I had ten years ago—or actually nine years ago. It’s an empowerment platform that helps individual women—men, too—get ahead in their careers, earn more, and stand out. And part of my story, you know, I have the degrees from undergrad in electrical engineering and computer science, but I went a completely different direction after graduating. And at the end of the Great Recession, I found myself with no job with no technical skills, and I mean, no job prospects, at all. It was really, really bad, ugly crying on my couch bad, Corey.

And I took a number of steps to get ahead and really relearn my tech skills, and I only got one offer to give myself a chance. It was a 90-days to prove myself, to get ahead, and teach myself iOS. And I remember it was one of the most terrifying things I’ve ever done. And within two years, I not only managed to survive that 90-day period and keep that job, but I had completely managed to thrive. My work had been featured in Apple’s iOS7 keynote, I’d won the company-wide award at a national agency four times, I had won the SXSW international Hackathon, twice in a row.

And then probably the pinnacle of it all is I was one of 100 tech innovators worldwide invited to attend the [UKG 00:03:41] Innovation Conference. And they flew me there on a private 747 jet, and it was just unreal. And so I founded DevelopHer because I needed this ten years ago, when I was at rock bottom, to figure out how to get ahead: how do I get into my career; how do I stand out? And of course, you know there’s more to the story, but I also found out I was underpaid after achieving all of that, that a male peer was paid exactly what I was paid, with no credentials, despite all of the awards that I won. And I went out and learned to negotiate, and tripled my salary in two years, and turned around and said, “I’m going to teach other women—and men, too—how to get real change in their own life.”

Corey: I love hearing stories where people discover that they’re underpaid. I mean, it’s a bittersweet moment because on the one hand it’s, “Wait, you mean they’ve been taking advantage of me?” And you feel bad for people, but at the same time, you’re sort of watching the blindfold fall away from their eyes of, “Yeah, but it’s been this way, and now you know about it. And now you’re in a position to potentially do something about it.” I gave a talk at a tech conference a few years back called “Weasel your Way to the Top: How to Handle a Job Interview” and it was a fun talk.

I really enjoyed it, but what I discovered was after I’d given it I got some very direct feedback of, “That’s a great talk and you give a lot of really useful advice. What if I don’t look like you?” And I realized, “Oh, my God, I built this out of things that worked for me and I unconsciously built all of my own biases and all of my own privilege into that talk.” At which point I immediately stopped giving it until I could relaunch it as a separate talk with a friend of mine, Sonia Gupta, who does not look like me. And between the two of us, it became a much stronger, much better talk.

Lauren: It’s good that you understand what you were bringing to the table and how you can appeal to an even larger audience. And what I’ve done is really said, “Here’s my experience as a woman in tech, and here’s what’s worked for me.” And what’s been surprising is men have said, “Yeah, that’s what I did.” Except for I put a woman in tech spin on it and… I mean, I knew it worked for me; I have more than quintupled my base salary—just my base salary alone—in nine years. And the results that women are getting from my programming—I had one woman who earned $80,000 more in a single negotiation, which tells me, one, she was really underpaid, but she didn’t just get one offer at $80,000 more; she got at least two. I mean, that changed her life.

And I think the lowest I’ve heard is, like, $30,000 difference change. I mean, this is, this is life-changing for a lot of women. And the scary thing is that it’s not just, say it’s $50,000 a year. Well, over ten years, that’s half a million dollars. Over 20 years, that’s a million, and that’s not even interest and inflation and compounding going into that. So, that’s a huge difference.

Corey: It absolutely is. It’s one of those things that continues to set people further and further back. One thing that I think California got very right is they’ve outlawed recently asking what someone’s previous compensation was because, “Oh, we don’t want to give someone too big of a raise,” is a way you perpetuate the systemic inequality. And that’s something that I wish more employers would do.

Lauren: It’s huge. I know the women and proponents who had moved that forward; some of them are personal friends of mine, and it’s huge. And that’s actually something that I trained specifically for is how to handle difficult questions like, “How much are you currently making?” Which you can’t legally get asked in California, although it still happens, so how do you handle it if you still get asked and you don’t want to rule yourself out? Or even worse—which they still can ask—which is, “How much do you want to make?”

And a lot of times, people get asked that before they know anything about the job. And they basically, if you give an answer upfront, you’re negotiating against yourself. And so I tackle tough things like that head-on. And I’m very much an engineer at heart, so for me, it’s very methodical; I prepare scripts in advance to handle the pushback that I’m going to get, to handle the difficult questions. Without a doubt, I know all of my numbers, and that’s where I’m getting real results for women is by taking the methodical approach to it.

Corey: So, I spent my 20s in crippling credit card debt, and I was extremely mercenary, as a result. This wasn’t because of some grand lost vision or something. Nope. I had terrible financial habits. So, every decision I made in that period of my life was extraordinarily mercenary. I would leave jobs I enjoyed for a job I couldn’t stand because it paid $10,000 more.

And the thing that I picked up from all of this, especially now having been on the other side of that running a company myself, is I’m not suggesting at any point that people should make career decisions based upon where they can make the most money, but that should factor in. One thing we do here at The Duckbill Group, in every job posting we put up is we post the salary range for the position. And I want to be clear here, it is less than anyone here could make at one of the big tech unicorns or a very hot startup that’s growing meteorically, and we’re upfront about that. We know that if money is the thing you’re after and that is the driving force behind what you’re going for, great; I don’t fault you for that.

This might not be the best role for you and that’s perfectly okay. I get it. But you absolutely should know what your market worth is so you can make that decision from a place of being informed, rather than being naive and later discovering that you were taken advantage of.

Lauren: So, I want to unpack just a couple things. There’s just so many gold nuggets in that. Number one, for any employer listening out there, that is such a great best practice, to post the range. You’re going to attract the right candidates when you post the right range. The last thing you want is to get to the end of the process to find out that, hey, you guys were totally off, and all the time invested could have been avoided if you’d had some sort of expectation set, upfront.

That said, that’s actually where I start with my negotiation training. A lot of people think I start with the money and that it’s all about the money. That’s not where I start. The very first thing I train women, and the men who’ve taken it, too, on the course is, figure out what success looks like to you. And not just the number success, but what does your life look like? What does your lifestyle look like? What does it feel like? What kinds of things do you do? What kinds of things do you value?

Money is one of those components, but it’s not all. And here’s the reason I did that: because at a certain point in my life, I only got out at—broke even out of debt, you know, within the last five years. That’s how underpaid I was at the time. But then once I started climbing out of debt, I started realizing it’s not all about money. And that’s actually how I ended up in my dream position.

I mean, I’m living out how I define success today. Could I be making a lot more money at a big tech unicorn? Yeah, I could. But I also have this incredible lifestyle; it’s sustainable. I get on apps like Blind and other internet forums, and I hear just horror stories of people burning out and the toxic cultures they work with. I don’t have that at all. I have something that I could easily do for the next 50 years of my life if I live that long.

But it’s not by accident that I’m in the role that I’m in right now. I actually took the time to figure out what success looks like to me, and so when this opportunity came along—and I was looking at it alongside other opportunities that honestly paid more, I recognized this opportunity for what it was because I’d put in the work up front to figure out what success looks like to me. And so that’s why what you guys are saying, “Hey, it’s a lifestyle that you guys are supporting and mission that you’re joining that’s so important.” And you need to know that and do that
work up front.

Corey: That’s I think what it really comes down to is understanding that in many cases… in fact, I’m going to take that back—in all cases, there’s an inherent adversarial nature to the discussions you have about compensation with your employer or your prospective employer. And I say ‘adversarial’ not antagonistic because you are misaligned as far as the ultimate purpose of the conversation. I’m not going to paint myself as some saint here and say that, oh, I’m on the side of every person I’m negotiating against, trying to get them to take a salary that’s less than they deserve. Because, first, although I view myself that I’m not in that position, you have to take that on faith from me, and I think that is too far of a bridge to cross. So, take even what I’m saying now from the position as someone who has a vested interest in the outcomes of that negotiation.

I mean, we’re not one of those unicorn startups; we can’t outbid Netflix and we wouldn’t even try to. We’re one of those old-fashioned businesses that has taken no investment and we fund ourselves through the magic of revenue and profitability, which means we don’t have a SoftBank-sized [laugh] war chest sitting in the bank that we can use to just hurl ridiculous money at people and see who pans out. Hiring has to be intentional and thoughtful because we’re a very small team. And if you’re looking for something that doesn’t align with that, great; I certainly don’t blame you. That isn’t this, and that’s okay, I’m not trying to hire everyone.

And if it’s not going to work out, why wouldn’t we say that upfront to avoid trying to get to all the way at the end of a very expensive interview process—both in terms of time and investment and emotionally—only to figure out that we’re worlds apart on comp, and it’s never going to work.

Lauren: A hundred percent agree. I mean, I’ve been through it on both ends, both as someone who is being hired and also as a hiring manager, and I understand it. And you need to find alignment, and that’s what negotiation is all about is finding an alignment, finding something where everyone feels like they’re winning in the situation. And I’m a big proponent—and this is going to go so counterculture—I think a lot of people overlook a lot of opportunities that are just golden nuggets. I think there’s a lot of idol worship of the big tech companies.

And don’t get me wrong; I’m sure they pay really well, great opportunity for your career, but I think people are overlooking a lot of really great career opportunities to get experience, and responsibility, and have good pay and lifestyle. And I’m a big proponent and looking for those golden nuggets rather than shooting for one of the big tech unicorns.

Corey: And other people are going to have a very different perspective on that, and that is absolutely okay. So, tell me a little bit more about what it is that DevelopHer does and how you go about doing it because it’s one thing to say, “Oh, we help women figure out that they are being underpaid,” but there’s a whole lot of questions that opens up because great. How do you do that?

Lauren: I do a number of things. So, it’s not all about pay either. Part of it’s building your value, building your confidence, standing out, getting ahead. DevelopHer started, actually, as a podcast. Funny story; I wanted to solve the problem of, we need more technical women as visible leaders out there, and I said, “Where are the architects? Where are the CTOs? Where are the CSOs?”

And I didn’t think anyone would care about me. I mean, I’m not Sheryl Sandberg; I’m not [laugh] the CEO of Facebook. Who’s going to listen to me? And then I was actually surprised when people cared about my own story, about coming back from being underpaid and then getting back into tech and figuring out how to stand out in such a short amount of time. And other women were saying, “Well, how did you do it?”

And it wasn’t just women; it was men, too, saying, “Hey, I also don’t know how to effectively advocate for myself.” And then it was companies saying, “Hey, can you come in and help us build our internal bench, recruit more women to come work for us, and build our own women leaders?” And then I’ve started working with universities to help bridge the gap before it even starts. I partnered with major universities to license my program and train them, not only how do you negotiate for what you’re worth, for your first salary, but also how do you come in and immediately make an impact and accelerate your career growth? And then, of course, I work with individual women.

I’ve talked about I have a salary negotiation course that’s won a couple awards for the work, the results that it’s getting, but then I just recently wrote a book because I wanted to reach women and men at scale and help them really get ahead. And this was literally my playbook. It’s called The DevelopHer Playbook. And it’s, how did I break into tech? And then once I was in tech, how did I get ahead so quickly? And it’s not rocket science. And that’s what I’m working on is training other people do it. And look, I’m still learning; I’m still paving my own path forward in tech, myself.

Corey: This episode is sponsored in part by our friends at Jellyfish. So, you’re sitting in front of your office chair, bleary eyed, parked in front of a powerpoint and—oh my sweet feathery Jesus its the night before the board meeting, because of course it is! As you slot that crappy screenshot of traffic light colored excel tables into your deck, or sift through endless spreadsheets looking for just the right data set, have you ever wondered, why is it that sales and marketing get all this shiny, awesome analytics and inside tools? Whereas, engineering basically gets left with the dregs. Well, the founders of Jellyfish certainly did. That’s why they created the Jellyfish Engineering Management Platform, but don’t you dare call it JEMP! Designed to make it simple to analyze your engineering organization, Jellyfish ingests signals from your tech stack. Including JIRA, Git, and collaborative tools. Yes, depressing to think of those things as your tech stack but this is 2021. They use that to create a model that accurately reflects just how the breakdown of engineering work aligns with your wider business objectives. In other words, it translates from code into spreadsheet. When you have to explain what you’re doing from an engineering perspective to people whose primary IDE is Microsoft Powerpoint, consider Jellyfish. Thats Jellyfish.co and tell them Corey sent you! Watch for the wince, thats my favorite part.

Corey: I feel like no one really has a great plan for, “Oh, where are you going next in tech? Do you have this whole thing charted out?” “Of course not. I’m doing this fly by night, seat of my pants, if I’m being perfectly honest with you.” And it’s hard to know where to go next.

What’s interesting to me is that you talk about helping people individually—generally women—through your program, but you also work directly with companies. And when you’re talking about things like salary negotiation, I think a natural question that flows from that is, are there aspects of what you wind up talking to individuals about versus what you do when talking to companies that are in opposition to each other?

Lauren: Yeah, so that’s a great question. So, the answer is there are some progressive companies that have brought me in to do salary negotiation training. Complete candor, most companies aren’t interested. It’s my Zero-To-Hero DevelopHer Playbook program which is, how do you get ahead? How do you build your value, become an asset at the company?

So, it’s less focused on pay, but more how do you become more valuable, and get ahead and add more value to the company? And that’s where I work with the individuals and the companies on that front.

Corey: It does seem like it would be a difficult sell, in most enterprise scenarios, to get a company to pay someone to come in to teach their staff how to more effectively [laugh] negotiate their next raise. I love the vision.

Lauren: It has happened. I also thought it was crazy, but it has happened. But no, most of my corporate clients say, “We not only want to encourage more women into tech, but we already have a lot of women who are already in our ranks, and we want to encourage them to really feel like they’re empowered and to stand out and reach the next levels.” And that’s my sweet spot for corporate.

Corey: Somewhat recently, I was asked on a Twitter Spaces—which is like Clubhouse but somehow different and strange—did I think that the privilege that I brought to what I do had enabled me to do these things, being white, being a man, being cis-gendered—speaking English as my primary language was an interesting one that I hadn’t heard contextualized like that before—and whether that had advantaged me as I went through these things? And I think it’s impossible to say anything other than absolutely because it’s easy to, on some level, take a step back and think, “Well, I’ve built this company, and this media platform, and the rest. And that wasn’t given to me; I had to build it.” And that’s absolutely true. I did have to build it, and it wasn’t given to me.

But as I was building it, the winds were at my back not against me. I was not surrounded by people who are telling me I couldn’t do it. Every misstep I made wasn’t questioned as, well, you sure you should be doing this thing that you’re not really doing? It was very much a fail-forward. And if you think that applies to everyone, then you are grievously mistaken.

Lauren: I think that’s a healthy perspective, which is why I consider you one of developers in my strongest allies, the fact that you’re willing to look at yourself and go, “What advantages did I have? And how might I need to adapt my messaging or my advice so that it’s applicable to even more people?” But it’s also something I’ve experienced myself. I mean, I set out to help women in tech because I’m in women in tech myself. And I was surprised by a couple of things.

Number one, I was surprised that men were [laugh] asking me for advice as well. And individuals and medicine, and finance, and law, in business not even related to tech, but what I’m really proud of that I didn’t set out to build because I didn’t feel qualified, but I’m really glad that I’ve been able to serve is that there were three populations that I’ve been really able to serve, especially at the university level. Number one, international students who, you mentioned yourself, English might not be their first language, but they’re not familiar with the US hiring and advancement and pay process, and I help normalize that. And that’s something that I myself in the benefit of, having been born here in the US. People who, where English isn’t their first language; you think it’s hard enough to answer, “Why do you think you should be promoted?”

Or, “How much do you think you should make for this role? What do you want?” In your first language? Try answering it in your third, right? And then when I’m really proud of is, especially at the university level, I’ve been really able to help students where they’re first-generation college students, where they don’t have a professional mentor within their immediate family.

And providing them a roadmap—or actually, the playbook to how to get ahead and then how to advocate for yourself. And these were things that I didn’t feel qualified to help, but these are the individuals who’ve ended up coming and utilizing my program, and finding a lot of benefit from that. And it made me realize that I’m doing something bigger than I even set out to do, and that is very meaningful to me.

Corey: You mentioned that you give guidance on salary negotiation and career advancement to not just women, but also men, and not just people who are in tech, but people who are in other business areas as well. How does what you’re advising people to do shift—if at all—from folks who are women working in tech?

Lauren: So, that’s the key is it really doesn’t shift. What I’m teaching are fundamentals and, spoiler alert, I teach grounding yourself in data, and knowing your data, and taking the emotion out of the process, whether you’re trying to get ahead, to stand out, to earn more. And I teach fundamentals, which is five-point process.

Number one, you got to figure out what success looks like to you. I talked a little bit about that earlier, but it’s foundational. I mean, I start with that because that alone changed my life. I would still be pursuing success today and not have reached it, but I’m living out how I defined success because I started there.

Then you got to really know your worth. Absolutely without a doubt, know how much you’re worth. And for me, this was transformational. I mean, eye-opening. Like you said earlier, the blindfold coming off. When I saw for a fact how much employers paid other people with my skill sets, it was a game-changer for me. And so I—without a shadow of a doubt, I use four different strategies, multiple resources in each strategy to know comprehensively how much I’m worth.

And then I teach knowing your numbers. It’s not an emotional thing; it’s very much scientific, so I talked about knowing your key numbers, your target, your ask, and your walk away, and those are all very dependent on your employment and financial situation, so it’s different from person to person. And then I talk about—and this is a little different than what other people teach—is I talk about finding leverage, what you uniquely bring to the table, or identifying companies where you uniquely add value, where you can either lock in an offer or negotiate a premium.

And then I prepare. I prepare. Just like you prepare for an interview, I prepare for a negotiation, and if I’m asking for the right amount of money, I am going to be prepared for pushback and I want to be able to handle that, and I don’t want to just know it on the fly; I want to have scripts and questions prepared to handle that pushback. I want to be prepared to answer some of the most difficult questions that you’re going—get asked, like we talked about earlier.

And then the final step is I practice over and over and over again, just like a sporting event. I am ready to go into action and get a great thing. So, those are the fundamentals. I’ve marketed to women in tech because I’m a woman in tech and we don’t have enough women in tech, and women are 82 cents on the dollar in tech, but what I found is that doctors were using the same methodology. I wasn’t marketing it to them. Lawyers, business people, finance people were using it because I was teaching such fundamentals.

Corey: Taking it one step further, if someone is listening to this and starting to get a glimmering of the sense that they’re not where they could be career-wise, either in terms of compensation, advancement, et cetera, what advice would you have for them as far as things to focus on first? Not to effectively extract the entire content of your course into podcast form, but where do they start?

Lauren: Yeah. So, you start by investing in yourself and investing in the change that you want. And that first investment might be figuring out how much you’re worth, you know, doing that research to figure out how much you’re worth. And then going out and learning the skills. And look, I have a course, I have a book that you can use to get ahead; if I’m not the right fit, there are a ton of resources out there. The trick is to find the best fit for you.

And my only regret as I look back over the last 10, 15 years of my career is that I didn’t invest in myself sooner and that I didn’t go out and figure out how much I was worth, and that I—when they said, “Well, you’re just not there yet,” when I asked for more money, that I believed them. And that was on me that I didn’t go out and go, “I wonder how much I’m worth?” And do the research. And then, I regret not hiring a career coach earlier. I wish I’d gotten back into tech sooner.

And I wish that I had learned to negotiate and advocate for myself sooner. But my knack, Corey—and I believe things happen to me for a reason—is my special skills is I take things that were meant not necessarily intentionally to harm me, but things that hurt me, I learned from them, I turn it around in the best way possible, and then I teach and I create programs to help uplift other people. And that’s my special skill set; that’s sort of my mission and purpose in life, and now I’m just trying to really exploit it and make this into a big movement that impacts millions of lives.

Corey: So, what’s next for you? You’ve built this platform, you’ve put yourself out there, you’ve clearly made a dent in the direction that you’re heading in. What’s next?

Lauren: [laugh]. I am looking to scale. I’m just like any company; I’ve really focused on delivering value proof of concept. What a lot of people don’t realize is not only did I build DevelopHer in quote, “my spare time,” but I did this without any outside investors. I funded it at all myself, built it on my own sweat equity—

Corey: [laugh]. That one resonates.

Lauren: Yeah. [laugh]. I know you know what that feels like. And so for me, I’m focused on scale: bringing in more corporate partners; bringing in more university clients, to scale and bridge the gap before it even starts; and scaling and reaching more women and men and anyone who wants to figure out how to get ahead, stand out, and earn more. And so the next year, two years are really focused on scale.

Corey: If people want to learn more about what you do, how you do it, or potentially look at improving their own situations, where can they find you?

Lauren: I am online. Go to developher.com. I have resources for individuals; I have a book, which is a great, cost-effective way to learn a lot.

I have an award-winning negotiation course that helps you go out and earn what you’re truly worth, and I have a membership to connect with me and other like-minded individuals. If you’re a company leader, I work with companies all the time to train their women—and men, too—to get ahead and build their value. And then also, I work with universities as well to help bridge the gender wage gap before it starts, and builds future leaders.

Corey: And we will, of course, include links to that in the [show notes 00:27:55]. Thank you so much for taking the time to speak with me today. I really appreciate it.

Lauren: Corey, thank you so much for having me, and I really mean it. You know, Corey is a strong ally. We connected, and I am glad to count you as not only my own ally but an ally of DevelopHer.

Corey: Well, thank you. That’s incredibly touching to hear. I appreciate it.

Lauren: I mean it.

Corey: Thank you. Sometimes all you can say to a sincere compliment is, “Thank you.” Arguing it is an insult, and I’m not that bold. [laugh].

Lauren: That’s actually really good advice that I give women is, so many times, we cut down our own compliments. And so that’s a great example right there, and it is not just women who sometimes I have a challenge with it; men, too. When someone gives you a compliment, just say, “Thank you.”

Corey: Good advice for any age, in any era. Lauren Hasson, founder of DevelopHer, speaker, author, frontline engineer some days. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice and an insulting comment telling me that my company is never going to succeed if I don’t attempt to outbid Netflix.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Courtney

Courtney Nash is a researcher focused on system safety and failures in complex sociotechnical systems. An erstwhile cognitive neuroscientist, she has always been fascinated by how people learn, and the ways memory influences how they solve problems. Over the past two decades, she’s held a variety of editorial, program management, research, and management roles at Holloway, Fastly, O’Reilly Media, Microsoft, and Amazon. She lives in the mountains where she skis, rides bikes, and herds dogs and kids.

Links:

  • Verica: https://www.verica.io
  • Twitter: https://twitter.com/courtneynash
  • Email: courtney@verica.io

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at the Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Jellyfish. So, you’re sitting in front of your office chair, bleary eyed, parked in front of a powerpoint and—oh my sweet feathery Jesus its the night before the board meeting, because of course it is! As you slot that crappy screenshot of traffic light colored excel tables into your deck, or sift through endless spreadsheets looking for just the right data set, have you ever wondered, why is it that sales and marketing get all this shiny, awesome analytics and inside tools? Whereas, engineering basically gets left with the dregs. Well, the founders of Jellyfish certainly did. That’s why they created the Jellyfish Engineering Management Platform, but don’t you dare call it JEMP! Designed to make it simple to analyze your engineering organization, Jellyfish ingests signals from your tech stack. Including JIRA, Git, and collaborative tools. Yes, depressing to think of those things as your tech stack but this is 2021. They use that to create a model that accurately reflects just how the breakdown of engineering work aligns with your wider business objectives. In other words, it translates from code into spreadsheet. When you have to explain what you’re doing from an engineering perspective to people whose primary IDE is Microsoft Powerpoint, consider Jellyfish. Thats Jellyfish.co and tell them Corey sent you! Watch for the wince, thats my favorite part.

Corey: This episode is sponsored in part by our friends at VMware. Let’s be honest—the past year has been far from easy. Due to, well, everything. It caused us to rush cloud migrations and digital transformation, which of course means long hours refactoring your apps, surprises on your cloud bill, misconfigurations and headache for everyone trying manage disparate and fractured cloud environments. VMware has an answer for this. With VMware multi-cloud solutions, organizations have the choice, speed, and control to migrate and optimize

applications seamlessly without recoding, take the fastest path to modern infrastructure, and operate consistently across the data center, the edge, and any cloud. I urge to take a look at vmware.com/go/multicloud. You know my opinions on multi cloud by now, but there's a lot of stuff in here that works on any cloud. But don’t take it from me thats: VMware.com/go/multicloud and my thanks to them again for sponsoring my ridiculous nonsense.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Periodically, websites like to fall into the sea and explode. And it’s sort of a thing that we’ve accepted happens. Well, most of us have. My guest today is Courtney Nash, Internet Incident Librarian at Verica. Courtney, thank you for joining me.

Courtney: Hi, Corey. Thanks so much for having me.

Corey: So, I’m going to assume that my intro is somewhat accurate, that we’ve sort of accepted that sites will crash into the sea, the internet will break, and then everyone tears their hair out and complains on Twitter, assuming that’s not the thing that fell over this time—

Courtney: [laugh].

Corey: —but what does an Internet Incident Librarian do?

Courtney: Yeah, I’ll come back to the first part about how—some people have accepted it and some people haven’t, I think is the interesting part. So technically, I think my official real title is, like, research analyst or something really boring, but I have a background in the cognitive sciences and also in technology, and I’m really—have always been fascinated by how these socio-technical systems work. And so as an Internet Incident Librarian, I am doing a number of things to try to better understand—both for myself and, obviously, the company I work for, but for the industry as a whole—what do we really know about how incidents happen, why they happen, when they happen, and what do we do when they happen? And how do we learn from that? So, one of the first things that I’m doing along those lines is actually collecting a database of all of the public write-ups of incidents that happened at companies that are software-related.

So, there’s already bodies of work of people who collect airline incidents and other kinds of things. And we don’t have that [laugh] as an industry, which I think is—I want to solve that problem because I think other industries that have spent some time introspecting about why things fall down, or when things fall down and how they fall down. Take the airline industry for example; planes don’t really fall out of the sky very often.

Corey: No. When it does, it makes news and everyone’s scared about flying, but at the same time, it’s yeah, do you have any idea how many people die in car crashes in a given hour?

Courtney: Yeah, yeah. And we’ll come back to how the media covers things in a minute because that is definitely something I have opinions about. But, I’m not trying to say I want to create the NTSB of the internet; I don’t think that’s quite the same thing, and I really want something in the spirit of software, and the internet, and open-source that’s more collaborative and it’s very open to all of us. So, the first step is to just get them in one place. There is no single place where you could go and say, “Oh, where all of the X incident reports? Where all the ones that Microsoft’s written, and also Amazon, or Google, or, you know, whoever.”

Corey: They have them, but they hide them so thoroughly. It turns out that they don’t really put that in big letters on their corporate blog with links to it. And when you look at one incident report, they don’t say, “Here, look at our previous incident reports.” They really—

Courtney: Yeah.

Corey: —should but no one does.

Courtney: And I think that’s fascinating because there’s a precedent. So, there’s two precedents, and I just gave you basically one side of the two, which is, the airline industry has done this and it’s not like people don’t fly, right? So, a lot of internet companies, a lot of software-based companies, seem to be afraid of what their customers, or what the stock market, or what folks will think. Mind you, these are publicly traded [laugh] airline companies. People aren’t going to stop using Amazon just because you give more of this information out.

And so I think that piece is—I would love to see that stop being the case. Because the flip side of the coin is that this is a rising tide lifts all boats kind of thing, which granted, not all companies agree on, especially really big ones because their boats already mowing all the little ones out of the ocean. But that’s another story.

Corey: Sure, but also, it’s easy to hide an outage. “Our site is down for you can say three days. Great, if a customer didn’t try to access the site at all during those three days, was the site really down in the first place?”

Courtney: Oh, the tree in the forest of internet outages. Yes, it’s true, although I think that companies are—they know that people go complain on social media, right? I think there’s more and more of that happening now. It’s not like you can hide it as easily as you could have before Twitter or Instagram or—

Corey: Right. Whereas a plane falls out of the sky, generally it’s one of those things that people notice.

Courtney: Yeah. Even if you weren’t interested in that flight at all.

Corey: Right. When it lands in your garden, you sort of have a comment on this.

Courtney: [laugh]. Yeah. Pieces fall out of the sky. That has happened. But I think the other flip side of the coin I already mentioned is the safety of airline industry has increased so significantly over the past, you know, whatever, 30, 40 years because of this concerted effort.

And the other piece of it, then, as an industry, as technologists, as people who use software to run their businesses, some of those things are now safety-critical. And this comes back to the whole software is running the world now. Planes now actually could fall out of the sky because of software, not just because of hardware failures. And nuclear power plants are [laugh] run by software, and your electronic grid, and your health care systems, heart rate monitors, insulin pumps. There are a lot of really critical things, and now our phone services and our internet stuff is so entwined in our lives, that people can’t be on their Zoom calls, people can’t run their businesses. So, this stuff has a massive impact on people’s lives. It’s no longer just pictures of cats on the internet, which admittedly, we’ve really honed the machine for that.

Corey: No, but now when software goes down, the biggest arguments people make, the stories people tell is, “Oh, well, it meant that the company lost this much money during that timeframe.” And great, maybe. We can argue about is that really true or is it not? It depends entirely on the company’s business model, but I don’t like to tend to accept those things at face value. But yeah, that’s the small-scale thing, especially when you start getting to these massive platform providers. There are a lot of second and third-order effects that are a lot more interesting slash important to people’s lives, than, well, we couldn’t show ads to people for an hour and a half.

Courtney: Right. Yes. Absolutely. So, T-Mobile had this outage, what is it, how is time—time is still not working very well, for me. I’m trying to remember if it was earlier this year, or if it was in—it was last year. I think it was 2020. And you’re like, T-Mobile, oh okay, whatever. You know, like, cell phones, yadda, yadda. 911 stopped working. [laugh].

And it was a fascinating outage because these are now actually regulated industries that are heavily software-backed. There was a government investigation into that the same way we have NTSB investigations into airline accidents, and they looked at all of those, kind of, second or third-order effects of people who—you know, a grandma who was stranded on the road, people who couldn’t call 911, those kinds of things that are really significant impacts on people’s lives. And the second-order effect is, oh, yeah, AWS goes down—like you said—and Amazon or people like to say, Jeff Bezos—I guess, now, are they going to complain about how much money Andy loses? I guess so—but [laugh] what lives on AWS, that’s crazy to think about, right?

Corey: Yeah, the more I learn the answer to that question, the more disturbed I become.

Courtney: Well, you’d probably know a better answer to that question [laugh] than a lot of people.

Corey: They have the big companies they can talk about. What’s really interesting is the companies that they don’t and can’t. An easy example: financial services is an industry that is notorious for never granting logo rights. Like, at some point, they’ll begrudgingly admit, “Yes, our multinational bank does use computers.” But it’s always like pulling teeth, and I get it on some level; the entire philosophy of a lot of these companies is risk-mitigation, rather than growth and advancing the current awareness of knowledge. But it does become a problem.

Courtney: Yeah. It’s interesting, I need more data, which we’ll get to—help me, people—but I am able to start seeing some of those interesting graphs of, kind of these cascading effects of these kinds of outages. And so I strongly believe that we need to talk about them more, that more companies need to write them up, and publish them, and be a lot more transparent about it. And I think there’s a number of companies that are showing the way there that—and it has to do with your first question which is, we’ve all sort of accepted this, right? But I disagree with that.

I think those of us who are super close to these kinds of complex, dynamic distributed systems totally know that they’re going to fail, and that’s not shocking, nor the case of incompetence. We are building systems that are so big and so complex, no one person, no 10X engineer out there could possibly model or hold the whole thing in their head. Especially because it’s not even just your systems… we were just talking about, right? Your stuff’s on GitHub; it’s on AWS; there’s, like, three other upstream providers; there’s this API from over there. These systems are too intricate, too complex; they’re going to fail.

Corey: So, we’re back to why all these things failed simultaneously and it comes out it’s a Northern woods, middle of nowhere backhoe incident. That’s right, if we look at the natural food chain of things, fiber optic cable has a natural predator in the form of a backhoe. To the point where if I’m ever lost in the woods, I will drop a length of fiber, kick some dirt over it, wait a few minutes; a backhoe will be along to sever it. Then I can follow the backhoe back to civilization. They don’t teach that one and the boy scout manual, but they really should.

Courtney: Yeah. Oh, my gosh. There was a beaver outage in Canada, which is the—[laugh] God, that’s the most Canadian thing ever.

Corey: Can you come up with a more Canadian—

Courtney: No.

Corey: —story than that? I would posit you could not, but give it a shot.

Courtney: No, probably not. Anyhoo. So, I think, like I was saying, those of us close to it accept that, understand it, and are trying to now think about, okay, well, how do we change our approach and our philosophy about this, knowing that things will fall down? But I think if you look at a lot of the rest of the world, people are still like, “What are those idiots doing over there? Why did their site fall down?”

Corey: Oh, my God—

Courtney: Right?

Corey: —the general population is the worst on stuff like this. The absolute worst.

Courtney: The media is the worst. [laugh].

Corey: It’s, “How did they wind up to going down?” “Yeah, because this stuff is complicated.” Back when I was getting started in tech, I thought the whole thing worked on magic, so I started figuring out different pieces of it worked. And now I’m convinced; it runs on magic. The most amazing thing is this all works together. Because—

Courtney: Yeah.

Corey: —spit and duct tape and baling wire holding this stuff together would be an upgrade from a lot of the stuff that currently exists in the real world. And it’s amazing.

Courtney: I know the secret, Corey. You know what holds it all together?

Corey: Hit me with it. Hope? Tears?

Courtney: People.

Corey: Mmm.

Courtney: Technology is Soylent Green, Corey. It’s Soylent Green. It’s made of people.

Corey: And that’s the thing that always bugs me on Twitter. The whole HugOps movement has it right. When you see a big provider taking an outage, all their competitors are immediately there with, “Man, hope things get back together soon. Best of luck. Let us know if we can help.” And that’s super reassuring because today is their outage; tomorrow it’s yours.

Courtney: Yep.

Corey: And once in a blue moon, you see someone who’s relatively new to the industry starting trying to market their stuff based on someone else’s outage, and they basically get their butts fed to them, just because it’s this—it’s not what you do, and it’s not how we operate. And it’s one of the few moments where I look at this and realize that maybe people’s inherent nature isn’t all terrible.

Courtney: [laugh]. Oh. Oh, I would hope that would be something that comes out of all of this.

Corey: Yeah.

Courtney: No one goes to work at their day job doing what we do, to suck. [laugh]. Right? To do a bad job.

Corey: Right. Unless you’re in Facebook’s ethics department, I completely agree with you.

Courtney: Okay. Yes. All right. There are a few caveats to that, probably. But you know, we all want to show up and do good stuff. So, nobody’s going in trying to take the site down, barring bad actor stuff that’s not relevant.

Corey: When Azure takes an outage, AWS is not sitting there going, “Ah, we’re going to win more cloud deals because of this,” because they’re smarter than that. It’s, no, people are going to look at this and say, “Ah, see. Told you the cloud was dangerous.” It sets the entire industry back.

Courtney: Yeah. That’s why we need to talk about it more, and we need to just normalize that these things happen and that we can all level up as an industry if we get a lot smarter about how we, A) think about that, and B) how we react to them. And we will develop much more useful models of our safety boundaries, right? That’s really it. You don’t know—no one at any of these companies hardly knows if you’re five steps from the cliff, five feet, driving a Ferrari 90 miles an hour towards the edge of it.

Like, we don’t know, it’s amazing to me just how much in the dark we are as an industry and how much of the world we’re running. So, I think this is one tiny, first little step in what could be sort of a sea change about how all of this works. So, that’s a big part of why I’m doing what I’m doing.

Corey: Well, let’s talk about something else you’re doing. So, tell me a little bit about VOID?

Courtney: Yeah. So, that’s the first iteration of this. So, it’s the [Verica Open Incident Database 00:14:10]. I feel like I have to say this almost every time John Allspaw would like me to say that it’s the Verica Open Incident Report Database, but VOID is way cooler than—

Corey: VOIRD?

Courtney: VOIRD.

Corey: Yeah, that sounds like you’re trying to make fun of someone ineffectively.

Courtney: Yeah. And there’s a reason why he’s not in marketing. But what this is is a collection of all of the publicly available incident reports in one place, easily searchable. You can search by company, you can search by technology, you can filter things by the types of, sort of, kinds of failure modes that we’re seeing. And it’s, I hope, valuable to a wide swath of folks, both technologists and otherwise: researchers, media and press types, analysts, and whatnot.

And my biggest desire is that people will look at it, realize how incomplete it is, and then help me fill it. [laugh]. Help me fill the VOID, people. I think I have right now, at the time we’re talking, about 1700, maybe 1800 of these. And they run the gamut. And I know some people who like to quibble about language—and I am one of those people having been an editor in various flavors of my life—not all of these are what a lot of people directly related to these, sort of, incident management and whatnot would call ‘incident reports.’

I wanted to collect a corpus that reflects all of the public information about software-related incidents. So, it’s anything from tweets—either from a company or just from people—to a status page, to a media article, a news article, an online article, to a full-blown deep-dive retrospective or post-mortem from a company that really does go into detail. It’s the whole gamut. It’s all of those things. I have no opinionated take on that.

I want that all to be available to people. And we’ve collected some metadata on all of the incidents as well. So, we’re collecting the obvious things like when did it happen? What date was it, if we can figure it out, or if it’s explicit—how long was it? And those kinds of things and then we collect some metadata, like I said. We add some tags: was this a complete production outage, was it a partial outage? Those kinds of things.

And this is all directly just taken from the language of the report. And we’re not trying—like I said—we’re trying not to have any sort of really subjective takes on any of that, but a bit of metadata that helps people spelunk some of this stuff. So, if it is the kind of report—these are usually from a status page, or a company post about it—what kinds of things were involved in this outage? So, sometimes you’ll get lucky and the company will tell you, “It was DNS,” because, you know, it’s always DNS.

Corey: On some level, it always is. That’s why—

Courtney: It always is.

Corey: —DNS is my database. It’s a database problem.

Courtney: It’s a database problem. And sometimes you get even more detail. And so we will put as much of that that’s in the report into a set of metadata about these things. So, I think there’s some fascinating, really easy things that I’ve already seen from some of these data, and we kind of hit on one of these, which is the way that companies themselves talk about these outages versus the way that press and media and other types of organizations talk about these things. So, I think there’s a whole bunch of really fascinating analysis that’s going to be available to nerdy research-minded type folks like myself.

I think it’s a place, though, where technologists can also go and spelunk things that they’re interested in, looking for patterns, anything that’s really—there’s an opportunity for experts in the field to add insights to what we can discern from these public incident reports. They are, like, two orders abstracted from what happened internally, but I think there’s still a lot that we can learn from those. So, the first iteration of the VOID will allow people to get a first look at some of the data and to help me, hopefully, add to it, grow that corpus over time, and we’ll see where that goes.

This episode is sponsored by our friends at Oracle Cloud. Counting the pennies, but still dreaming of deploying apps instead of "Hello, World" demos? Allow me to introduce you to Oracle's Always Free tier. It provides over 20 free services and infrastructure, networking databases, observability, management, and security.

And - let me be clear here - it's actually free. There's no surprise billing until you intentionally and proactively upgrade your account. This means you can provision a virtual machine instance or spin up an autonomous database that manages itself all while gaining the networking load, balancing and storage resources that somehow never quite make it into most free tiers needed to support the application that you want to build.

With Always Free you can do things like run small scale applications, or do proof of concept testing without spending a dime. You know that I always like to put asterisks next to the word free. This is actually free. No asterisk. Start now. Visit https://snark.cloud/oci-free that's https://snark.cloud/oci-free.

Corey: I love the idea of having a centralized place where outages, post-mortems, root cause analyses—I’ll let you tear into that in a minute—and other things that are all tied to where can I find a list of outages. Because companies list these on their websites, they put them in blog posts, and it’s always very begrudging; they don’t link them from any other place, you have to know the magic incantation to find the buried link on their site. Having something that is easily searchable for outages is really something that’s kind of valuable.

Courtney: Yeah. And I mean, some of them are like—I’m looking at you, Microsoft—I like you for a lot of reasons, but hey, I have to scroll your status page. I can’t link directly to their write-ups, and—this is Azure—and it [laugh] please stop. Make it easier. [laugh]. You’re driving me crazy; I don’t even have a data model to figure out how to make this work for people, other than, like, taking screenshots of them.

So yeah, so there’s shades of grey and black in how much they’ll share, or how easy it is to find these things. So, it’ll be interesting to see if there’s any less-than-positive [laugh] reactions to all of this being available in one place. I’m anticipating at least a little bit of that.

There is one other type of metadata that we collect for the VOID. And that is the type of analysis that is conducted if it is clear what that type of analysis is. And there, some companies explicitly say, or call it an RCA, “We did a Root Cause Analysis.” There’s a few other types; some people talk about having a Contributing Factors Analysis. Most people don’t consider a formal analysis type, but I am trying to collect and categorize these because I do think there are some fascinating implications buried therein, and I would like to see if I can keep track of whether or not those change over time. And yes, you’ve hit on one of my favorite hot-take soapbox things, which is root cause.

Corey: Please, take it away.

Courtney: Yeah. Well, and anyone who’s close to these systems and has watched these things fall down has the inherent sense that there is no root cause. Like—[laugh]—let’s—great. One of my favorite ones: human error. We don’t have enough hours for this, Corey. I’m sorry. That’s one of my favorite other ones. But let’s say somebody fat-fingers a config change. Which happens—

Corey: That was fundamentally the S3 service disruption back in—

Courtney: Yes.

Corey: —2017 that took down S3 for hours on end.

Courtney: And took down so many other people that relied on S3.

Corey: Everything was tied to that. And that’s an interesting question; when something like that hits, does that mean that everything it takes down get its own entry in VOID?

Courtney: I hope so. If everybody writes them up, then yes. [laugh]. So, if S3 goes down, and you go down, and you write it up, and you put it in the VOID, then we can see those things, which would be so cool. But let’s go back to the fat-fingered config file—which if you haven’t ever done, you’re lying, first of all—

Corey: Or you haven’t been allowed to touch anything large and breakable yet, which, either way, you’re lying on some level. So, please—

Courtney: Yeah. I mean, I took down [Halloway’s 00:20:53] homepage when it was on Hacker News because of YAML. So, anywho. Even if you fat-finger a config change, that’s not the root cause because you have this system wherein a fat-fingered configure change can take down S3. That is a very big, complex, and I might add, socio-technical system.

There are decisions that were made long ago about why it was structured that way, or why this happens that way, or what kinds of checks and balances you have. It’s just, get over it people. There is no root cause. These are complex, highly dynamic systems that when they fail, they fail in unpredictable and weird ways because we’ve built them that way. They’re complex because you’re successful at pushing the envelope and your safety boundaries.

So, if we could get past the root cause thing as an industry, I mean, I could probably just retire happy, honestly. [laugh]. I’m a simple woman; could we just get one thing, people? [laugh]. First of all, then it gives non-technologists, people outside of our bubble, the media, you can’t hang it on these things anymore. We all have to then grapple with the complexity, which admittedly humans, not big fans of, but—

Corey: People want simple stories, simple narratives. When people say, “Oh, remember the S3 outage?” They don’t want to sit there and have to recount 50,000 different details. They want to say, “Oh, yeah. It took down a few big sites like Instagram, United Airlines, and it was a real mess.” The end. They want something that fits in a tweet, not something that fits in a thesis.

Courtney: Well, and if you have a single root cause, then you can fix the root cause and it will never happen again. Right?

Corey: That’s the theory. If we’re just a little bit more careful, we’re never going to have outages anymore.

Courtney: Yeah, if we could just train those humans to not try to make the best possible high-quality decision they could possibly make in that situation given the information they have at the time, then we’ll do better. But I mean, that’s why your system stay up most of the time, if you think about it. It’s shocking how well these things actually work the vast majority of the time. And that’s what we could learn from this, too. We could, you know—oh if we would write near-misses up, please.

I mean, if I could have one more wish, I think one of the coolest things the airline industry and the government side of that did was start writing up near-misses. It’s, wow, what do we learn from when we’re successful, versus trying to, like, spelunk and nitpick the failures.

Corey: Most of us aren’t so good at the whole introspection part. We need failures, we need painful outages to really force us to make difficult, introspective, soul-searching decisions and learn from them.

Courtney: Yeah. And I don’t disagree with that. I just wish one of the things we would learn is that we should study our successes, too. There’s more to be mined from our successes, if we can figure out how to do that, then there is from our failures. So, I have a metadata category in the VOID called ‘near-miss.’

And oh man, I really wish people would write those up more. I mean, I think there’s, like, five things in there that I’ve found so far. Because the humans hold these systems together. We make these things work the vast majority of the time. That’s why there is no root cause, and even when we’re involved in these things, we’re also involved in preventing them, or solving them, or remediating them. So, yeah, there’s no root cause. Humans aren’t the problem. Those are my big hot button ones.

Corey: I really wish more places would embrace that. Even Amazon uses the ‘root cause’ terminology internally, and I’m not going to sit here and tell them how to run large things at scale; that’s what I pay them to figure out for me. But I can’t shake the feeling that by using that somewhat reductive terminology that they’re glossing over an awful lot of things the rest of us could really benefit from.

Courtney: Well, so the question then—one of the other things that I look at is, personally when I read and analyze these incident reports, these public ones a lot, I always ask myself, “Who’s the audience for this?” And there are different audiences for different types of incident reports and different things. The vast majority of them are for customers, partners, investors.

Corey: The stock market. Yes. Yes.

Courtney: They’re not actually for the organization. There’s usually an internal one that we don’t get to see—maybe—that’s for the organization. But a lot of places feel that if you have a process, and a template, and a checklist, and a list of action items at the end, then you’ve done the right thing. You’ve had your incident, you’ve talked about it, you’ve got your action items. Move on.

Corey: Right, and it always seems with companies, that as you get further into the company, the more honest and transparent the actual analysis is. Like, at some point, you wind up with the, like, they’re very public and very cagey, and under NDA, they open up a little bit more, and a little bit more, and finally, when you work there, their executive team, it turns out, the actual thing was, “Well, Dewey was carrying arm full of boxes in the data center, tripped, went cascading face-first into the EPO cutoff switch that cut power to the entire facility.” The cagier they get, the—I guess, not to be unkind here—but the more ridiculous whatever the actual answer is. It’s one of those things where, “Really? Someone tripped and hit a button. You didn’t have a plan for that?” “Well, not really. We sort of assumed that people would”—

Courtney: Why would you have a plan for that, right?

Corey: Right.

Courtney: I mean like—[laugh].

Corey: Why would you have a plan for that, the first time?

Courtney: Yeah. I mean, so imagine this exercise: sitting down in a room with a bunch of people and going, “What are all the things that could go wrong?” I mean, [laugh] ain’t nobody got time for that? That’s not how it works. You all have other jobs to do, too, and systems to build, and pressures, and customers, and partners, and features to build, so admit and acknowledge that you just won’t know all of the antecedents and how do you respond when things happen?

Which is a whole other, you know—I know you told me you recorded an episode with Dr. Christina Maslach on burnout, which I’m so happy you did, and there’s a whole ‘nother piece of incidents and incident response, and burning people out, and blaming people, and all that stuff that’s a whole ‘nother pod—it sounds like you might—you know, probably not incidents with her. But still, these things take a toll on people. And people who, like I said, show up every day really hoping to do their best job, and go up a ladder, and get a promotion, and whatever. So, I think not just treating those things as checklists has broader implications as well, just for the wellbeing of your organization.

Corey: On some level, the biggest problem that I think we’ve run into is that, as you said, it all comes down to people. Unfortunately, legally,
we can’t patch those. Yet.

Courtney: No, [laugh]. No, no. Not most kinds of patches, no. And that’s messy. And I know some people are like, “Everyone should learn to code.” And I’m like, “Actually, everyone should get a liberal arts degree.” Come on, help me out people. Because there’s so much of these socio-technical systems where the socio part of it is more relevant than the actual technical part.

Corey: I believe you’re right, for better or worse; there’s no way around it. Thank you so much for taking the time to speak with me. If people want to learn more about what you’re up to, where can they find you? And we will, of course, throw a link to VOID in the [show notes 00:28:06].

Courtney: Yeah, I also like to talk on Twitter, like you do. I’m not as good at it as you are, but I try. So yeah, I’m @courtneynash on Twitter. And at Verica, you can find me at Verica as well, courtney@verica.io. And those are the best ways to find me, I would say. And yeah, please people, write up your incidents, send them to the VOID and let’s all learn and get better together, please.

Corey: Thank you so much for taking the time to speak with me today. I really do appreciate it.

Courtney: Thank you for having me on. I know—do people say this: I’m like, “Yeah, big fan,” but I am. I’m a [laugh] big fan [laugh] of the podcast.

Corey: Oh, dear Lord, find better things to listen to. My God.

Courtney: [laugh]. But it’s been a treat. Thank you.

Corey: Courtney Nash, Internet Incident Librarian at Verica. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with a comment making it very clear that for whatever reason the website is down, it is most certainly not your fault.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need the Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Jackie

Jackie Singh is an Information Security professional with more than 20 years of hacking experience, beginning in her preteen years. She began her career in the US Army, and deployed to Iraq in 2003. Jackie subsequently spent several years in Iraq and Africa in cleared roles for the Department of Defense.

Since making the shift to the commercial world in 2012, Jackie has held a number of significant roles in operational cybersecurity, including Principal Consultant at Mandiant and FireEye, Global Director of Incident Response at Intel Security and McAfee, and CEO/Cofounder of a boutique consultancy, Spyglass Security.

Jackie is currently Director of Technology and Operations at the Surveillance Technology Oversight Project (S.T.O.P.), a 501(C)(3), non-profit advocacy organization and legal services provider. S.T.O.P. litigates and advocates to abolish local governments' systems of mass surveillance.

Jackie lives in New York City with her partner, their daughters, and their dog Ziggy.

Links:

  • Disclose.io: https://disclose.io
  • Twitter: https://twitter.com/hackingbutlegal

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at VMware. Let’s be honest—the past year has been far from easy. Due to, well, everything. It caused us to rush cloud migrations and digital transformation, which of course means long hours refactoring your apps, surprises on your cloud bill, misconfigurations and headache for everyone trying manage disparate and fractured cloud environments. VMware has an answer for this. With VMware multi-cloud solutions, organizations have the choice, speed, and control to migrate and optimize

applications seamlessly without recoding, take the fastest path to modern infrastructure, and operate consistently across the data center, the edge, and any cloud. I urge to take a look at vmware.com/go/multicloud. You know my opinions on multi cloud by now, but there's a lot of stuff in here that works on any cloud. But don’t take it from me thats: VMware.com/go/multicloud and my thanks to them again for sponsoring my ridiculous nonsense.

Corey: This episode is sponsored in part by “you”—gabyte. Distributed technologies like Kubernetes are great, citation very much needed, because they make it easier to have resilient, scalable, systems. SQL databases haven’t kept pace though, certainly not like no SQL databases have like Route 53, the world’s greatest database. We’re still, other than that, using legacy monolithic databases that require ever growing instances of compute. Sometimes we’ll try and bolt them together to make them more resilient and scalable, but let’s be honest it never works out well. Consider Yugabyte DB, its a distributed SQL database that solves basically all of this. It is 100% open source, and there's not asterisk next to the “open” on that one. And its designed to be resilient and scalable out of the box so you don’t have to charge yourself to death. It's compatible with PostgreSQL, or “postgresqueal” as I insist on pronouncing it, so you can use it right away without having to learn a new language and refactor everything. And you can distribute it wherever your applications take you, from across availability zones to other regions or even other cloud providers should one of those happen to exist. Go to yugabyte.com, thats Y-U-G-A-B-Y-T-E dot com and try their free beta of Yugabyte Cloud, where they host and manage it for you. Or see what the open source project looks like—its effortless distributed SQL for global apps. My thanks to Yu—gabyte for sponsoring this episode.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. The best part about being me—well, there’s a lot of great things about being me, but from my perspective, the absolute best part is that I get to interview people on the show who have done awesome and impressive things. Therefore by osmosis, you tend to assume that I’m smart slash know-what-the-living-hell-I’m-talking-about. This is proveably untrue, but that’s okay.

Even when I say it outright, this will fade into the depths of your mind and not take hold permanently. Today is, of course, no exception. My guest is Jackie Singh, who’s an information security professional, which is probably the least interesting way to describe who she is and what she does. Most recently, she was a senior cybersecurity staffer at the Biden campaign. Thank you so much for joining me. What was that like?

Jackie: Thank you so much for having me. What was that like? The most difficult and high-pressure, high-stress job I’ve ever had in my life. And, you know, I spent most of my early 20s in Iraq and Africa. [laugh].

Corey: It’s interesting, you’re not the first person to make the observation that, “Well, I was in the military, and things are blowing up all around, and what I’m doing next to me is like—‘oh, the site is down and can’t show ads to people?’ Bah, that’s not pressure.” You’re going the other direction. It’s like, yeah, this was higher stress than that. And that right there is not a common sentiment.

Jackie: I couldn’t anticipate, when I was contacted for the role—for which I had applied to through the front door like everyone else, sent in my resume, thought it looked pretty cool—I didn’t expect to be contacted. And when I was interviewed and got through the interviews and accepted the role, I still did not properly anticipate how this would change my life and how it would modify my life in the span of just a few months; I was on the campaign for five to six months.

Corey: Now, there’s a couple of interesting elements to this. The first is it’s rare that people will say, “Oh, I had a job for five to six months,” and, a, put it on their resume because that sounds like, “Ah, are you one of those job-hopper types?” But when you go into a political campaign, it’s very clearly, win or lose, we’re out of jobs in November. Ish. And that is something that is really neat from the perspective of career management and career planning. Usually is, “Hey, do you want a six-month job?” It’s, “Why? Because I’m going to rage quit at the end of it. That seems a little on the weird side.” But with a campaign, it’s a very different story. It seems like a different universe in some respects.

Jackie: Yes, absolutely. It was different than any other role I’d ever had. And being a political dilettante, [laugh] essentially, walking into this, I couldn’t possibly anticipate what that environment would be like. And, frankly, it is a bit gatekept in the sense that if you haven’t participated on a campaign before, you really don’t have any idea what to expect, and they’re all a bit different to, like, their own special snowflake, based on the people who are there, and the moment in time during which you are campaigning, and who you are campaigning for. And it really does change a perspective on civic life and what you can do with your time if you chose to spend it doing something a little bigger than your typical TechOps.

Corey: It also is a great answer, too, when people don’t pay close enough attention. “So, why’d you leave your last job?” “He won.” Seems like a pretty—

Jackie: [laugh].

Corey: —easy answer to give, on some level.

Jackie: Yes, absolutely. But imagine the opposite. Imagine if our candidate had lost, or if we had had data walk out the door like in 2016. The Democratic National Convention was breached in 2016 and some unflattering information was out the door, emails were hacked. And so it was difficult to anticipate… what we had control over and how much control we could actually exert over the process itself, knowing that if we failed, the repercussions would be extremely severe.

Corey: It’s a different story than a lot of InfoSec gigs. Companies love to talk like it is the end of the universe if they wind up having a data breach, in some effect. They talk about that the world ends because for them it kind of does because you have an ablative CSO who tries to also armor themselves with ablative interns that they can blame—if your SolarWinds. But the idea being that, “Oh yeah, if we get breached we are dunzo.”

And it’s, first, not really. Let’s not inflate the risks here. Let’s be honest; we’re talking about something like you’re a retailer; if you get breached, people lose a bunch of credit card numbers, the credit card companies have to reissue it to everyone, you get slapped with a fine, and you get dragged in the press, but statistically, look at your stock price a year later, it will be higher than at the time of the breach in almost every case. This is not the end of the world. You’re talking about something though that has impacts that have impossible-to-calculate repercussions.

We’re talking about an entire administration shift; US foreign policy, domestic policy, how the world works and functions is in no small part tied to data security. That’s a different level of stress than I think most security folks, if you get them honest enough, are going to admit that, yeah, what I do isn’t that important from an InfoSec perspective. What you did is.

Jackie: I appreciate that, especially having worked in the military. Since I left the military, I was always looking for a greater purpose and a larger mission to serve. And in this instance, the scope of work was somewhat limited, but the impact of failing would have been quite wide-ranging, as you’ve correctly identified. And walking into that role, I knew there was a limited time window to get the work done. I knew that as we progressed and got closer and closer to election day, we would have more resources, more money rolls in, more folks feel secure in the campaign and understand what the candidate stands for, and want to pump money into the coffers. And so you’re also in an interesting situation because your resourcing is increasing, proportional to the threat, which is very time-bound.

Corey: An inherent challenge is that unlike in a corporate environment, in many respects, where engineers can guard access to things and give the business clear lines of access to things and handle all of it in the background, one of the challenges with a campaign is that you are responsible for data security in a variety of different ways, and the interfaces to that data explode geometrically and to people with effectively no level whatsoever of technical sophistication. I’m not talking about the candidate necessarily—though that’s of course, a concern—but I’m talking organizers, I’m talking volunteers, I’m talking folks who are lifelong political operatives, but they tend not to think in terms of, “Oh, I should enable multi-factor authentication on everything that I have,” because that is not what they are graded on; it’s pass-fail. So, it’s one of those things where it is not the number one priority for anyone else in your organization, but it is yours and you not only have to get things into fighting shape, you have to furthermore convince people to do the things that get them there. How do you approach that?

Jackie: Security awareness [laugh] in a nutshell. We were lucky to work with Bob Lord, who is former CSO at Yahoo, OAuth, Rapid7, and has held a number of really important roles that were very wide in their scope, and responsible for very massive data sets. And we were lucky enough to, in the democratic ecosystem, have a CSO who really understood the nature of the problem, and the way that you described it just now is incredibly apt. You’re working with folks that have no understanding or very limited understanding of what the threat actors were interested in breaching the campaign, what their capability set is, and how they might attempt to breach an organization. But you also had some positives out of that.

When you’re working with a campaign that is distributed, your workforce is distributed, and your systems are also distributed. And when you lose that centralization that many enterprises rely on to get the job done, you also reduce opportunities for attackers to compromise one system or one user and move laterally. So, that was something that we had working for us. So, security awareness was incredibly important. My boss worked on that quite a bit.

We had an incredible IT help desk who really focused on connecting with users and running them through a checklist so everyone in the campaign had been onboarded with a specific set of capabilities and an understanding of what the security setup was and how to go about their business in a secure way. And luckily, very good decisions had been made on the IT side prior to the security team joining the organization, which set the stage for a strong architecture that was resistant to attack. So, I think a lot of the really solid decisions and security awareness propagation had occurred prior to myself and my boss joining the campaign.

Corey: One of the things that I find interesting is that before you started that role—you mentioned you came in through the front door, which personally I’ve never successfully gotten a job like that; I always have to weasel my way in because I have an eighth-grade education and my resume—

Jackie: [laugh].

Corey: —well, tenure-wise, kind of, looks like a whole bunch of political campaigns. And that’s fine, but before that, you were running your own company that was a focused security consultancy. Before that, your resume is a collection of impressive names. You were a principal consultant at Mandiant, you were at Accenture. You know what you’re talking about.

You were at McAfee slash Intel. You’ve done an awful lot of corporate world stuff. What made you decide to just wake up one day and decide, “You know what sounds awesome? Politics because the level of civil discourse there is awesome, and everyone treats everyone with respect and empathy, and no one gets heated or makes ridiculous arguments and the rest. That’s the area I want to go into.” What flipped that switch for you?

Jackie: If I’m completely honest, it was pure boredom. [laugh]. I started my business, Spyglass Security, with my co-founder, Jason [Shore 00:11:11]. And our purpose was to deliver boutique consulting services in a way that was efficient, in a way that built on prior work, and in a way that helped advance the security maturity of an organization without a lot of complex terminology, 150-page management consulting reports, right? What are the most effective operational changes we can make to an organization in how they work, in order to lead to some measurable improvement?

And we had a good success at the New York City Board of Elections where we were a subcontractor to a large security firm. And we were in there for about a year, building them a vulnerability management program, which was great. But generally speaking, I have found myself bored with having the same conversations about cybersecurity again and again, at the startup level and really even at the enterprise level. And I was looking for something new to do, and the role was posted in a Slack that I co-founded that is full of digital forensics and information security folks, incident responders, those types of people.

And I didn’t hear of anyone else applying for the role. And I just thought, “Wow, maybe this is the kind of opportunity that I won’t see again.” And I honestly sent my resume and didn’t expect to hear anything back, so it was incredible to be contacted by the chief information security officer about a month after he was hired.

Corey: One of the things that made it very clear that you were doing good work was the fact that there was a hit piece taken out on you in one of the absolute worst right-wing rags. I didn’t remember what it was. It’s one of those, oh, I’d been following you on Twitter for a bit before that, but it was one of those okay, but I tend to shortcut to figuring out who I align with based upon who yells at them. It’s one of those—to extend it a bit further—I’m lazy, politically speaking. I wind up looking at two sides yelling at each other, I find out what side the actual literal flag-waving Nazis are on, and then I go to the other side because I don’t ever want someone to mistake me for one of those people. And same story here. It’s okay, you’re clearly doing good work because people have bothered to yell at you in what we will very generously term ‘journalism.’

Jackie: Yeah, I wouldn’t refer to any of those folks—it was actually just one quote-unquote journalist from a Washington tabloid who decided to write a hit piece the week after I announced on Twitter that I’d had this role. And I took two months or so to think about whether I would announce my position at the campaign. I kept it very quiet, told a couple of my friends, but I was really busy and I wasn’t sure if that was something I wanted to do. You know, as an InfoSec professional, that you need to keep your mouth shut about most things that happened in the workplace, period. It’s a sensitive type of role and your discretion is critical.

But Kamala really changed my mind. Kamala became the nominee and, you know, I have a similar background to hers. I’m half Dominican—my mother’s from the Dominican Republic and my father is from India, so I have a similar background where I’m South Asian and Afro-Caribbean—and it just felt like the right time to bolster her profile by sharing that the Biden campaign was really interested in putting diverse candidates in the world of politics, and making sure that people like me have a seat at the table. I have three young daughters. I have a seven-year-old, a two-year-old, and a one-year-old.

And the thing I want for them to know in their heart of hearts is that they can do anything they want. And so it felt really important and powerful for me to make a small public statement on Twitter about the role I had been in for a couple of months. And once I did that, Corey, all hell broke loose. I mean, I was suddenly the target of conspiracy theorists, I had people trying to reach out to me in every possible way. My LinkedIn messages, it just became a morass of—you know, on one hand, I had a lot of folks congratulate me and say nice things and provide support, and on the other, I just had a lot of, you know, kind of nutty folks reach out and have an idea of what I was working to accomplish that maybe was a bit off base.

So yeah, I really wasn’t surprised to find out that a right-wing or alt-right tabloid had attempted to write a hit piece on me. But at the end of the day, I had to keep moving even though it was difficult to be targeted like that. I mean, it’s just not typical. You don’t take a job and tell people you got a job, [laugh] and then get attacked for it on the national stage. It was really unsurprising on one hand, yet really quite shocking on another; something I had to adjust to very quickly. I did cry at work. I did get on the phone with legal and HR and cry like a baby. [laugh].

Corey: Oh, yeah.

Jackie: Yeah. It was scary.

Corey: I guess this is an example of my naivete, but I do not understand people on the other side of the issue of InfoSec for a political campaign—and I want to be clear, I include that to every side of an aisle—I think there are some quote-unquote, “Political positions” that are absolutely abhorrent, but I also in the same breath will tell you that they should have and deserve data security and quality InfoSec representation. In a defensive capacity, to be clear. If you’re—“I’m the offensive InfoSec coordinator for a campaign,” that’s a different story. And we can have a nuanced argument about that.

Jackie: [laugh].

Corey: Also to be very clear, for the longest time—I would say almost all of my career until a few years ago—I was of the impression whatever I do, I keep my politics to myself. I don’t talk about it in public because all I would realistically be doing is alienating potentially half of my audience. And what shifted that is two things. One of them, for me at least, is past a certain point, let’s be very clear here: silence is consent. And I don’t ever want to be even mistaken at a glance for being on the wrong side of some of these issues.

On another, it’s, I don’t accept, frankly, that a lot of the things that are currently considered partisan are in fact, political issues. I can have a nuanced political debate on either side of the aisle on actual political issues—talking about things like tax policy, talking about foreign policy, talking about how we interact with the world, and how we fund things we care about and things that we don’t—I can have those discussions. But I will not engage and I will not accept that, who gets to be people is a political issue. I will not accept that treating people with respect, regardless of how high or low their station, is a political issue. I will not accept that giving voice to our worst darkest impulses is a political position.

I just won’t take it. And maybe that makes me a dreamer. I don’t consider myself a political animal. I really don’t. I am not active in local politics. Or any politics for that matter. It’s just, I will not compromise on treating people as people. And I never thought, until recently, that would be a political position, but apparently, it is.

Jackie: Well, we were all taught the golden rule is children.

Corey: There’s a lot of weird things that were taught as children that it turns out, don’t actually map to the real world. The classic example of that is sharing. It’s so important that we teach the kids to share, and always share your toys and the rest. And now we’re adults, how often do we actually share things with other people that aren’t members of our immediate family? Turns out not that often. It’s one of those lessons that ideally should take root and lead into being decent people and expressing some form of empathy, but the actual execution of it, it’s yeah, sharing is not really a thing that we value in society.

Jackie: Not in American society.

Corey: Well, there is that. And that’s the challenge, is we’re always viewing the world through the lens of our own experiences, both culturally and personally, and it’s easy to fall into the trap that is pernicious and it’s always there, that our view of the world is objective and correct, and everyone else is seeing things from a perspective that is not nearly as rational and logical as our own. It’s a spectrum of experience. No one wakes up in the morning and thinks that they are the villain in the story unless they work for Facebook’s ethics department. It’s one of those areas of just people have a vision of themselves that they generally try to live up to, and let’s be honest people fell in love with one vision of themselves, it’s the cognitive dissonance thing where people will shift their beliefs instead of their behavior because it’s easier to do that, and reframe the narrative.

It’s strange how we got to this conversation from a starting position of, “Let’s talk about InfoSec,” but it does come back around. It comes down to understanding the InfoSec posture of a political campaign. It’s one of those things that until I started tracking who you were and what you were doing, it wasn’t something really crossed my mind. Of course, now you think about, of course there’s a whole InfoSec operation for every campaign, ever. But you don’t think about it; it’s behind the scenes; it’s below the level of awareness that most people have.

Now, what’s really interesting to me, and I’m curious if you can talk about this, is historically the people working on the guts of a campaign—as it were—don’t make public statements, they don’t have public personas, they either don’t use Twitter or turn their accounts private and the rest during the course of the campaign. You were active and engaging with people and identifying as someone who is active in the Biden campaign’s InfoSec group. What made you decide to do that?

Jackie: Well, on one hand, it did not feel useful to cut myself off from the world during the campaign because I have so many relationships in the cybersecurity community. And I was able to leverage those by connecting with folks who had useful information for me; folks outside of your organization often have useful information to bring back, for example, bug bounties and vulnerability disclosure programs that are established by companies in order to give hackers a outlet. If you find something on hardwarestore.com, and you want to share that with the company because you’re a white hat hacker and you think that’s the right thing to do, hopefully, there’s some sort of a structure for you to be able to do that. And so, in the world of campaigning, I think information security is a relatively new development.

It has been, maybe, given more resources in this past year on the presidential level than ever before. I think that we’re going to continue to see an increase in the amount of resources given to the information security department on every campaign. But I’m also a public person. I really do appreciate the opportunity to interact with my community, to share and receive information about what it is that we do and what’s happening in the world and what affects us from tech and information security perspective.

Corey: It’s just astonishing for me to see from the outside because you are working on something that is foundationally critically important. Meanwhile, people working on getting people to click ads or whatnot over at Amazon have to put ‘opinions my own’ in their Twitter profile, whereas you were very outspoken about what you believe and who you are. And that’s a valuable thing.

Jackie: I think it’s important. I think we often allow corporations to dictate our personality, we allow our jobs to dictate our personality, we allow corporate mores to dictate our behavior. And we have to ask ourselves who we want to be at the end of the day and what type of energy we want to put out into the world, and that’s a choice that we make every day. So, what I can say is that it was a conscious decision. I can say that I worked 14 hours a day, or something, for five, six months. There were no weekends; there was no time off; there were a couple of overnights.

Corey: “So, what do you get to sleep?” “November.”

Jackie: Yeah. [laugh]. My partner took care of the kids. He was an absolute beast. I mean, he made sure that the house ran, and I paid no attention to it. I was just not a mom for those several months, in my own home.

Corey: This episode is sponsored by our friends at Oracle HeatWave is a new high-performance accelerator for the Oracle MySQL Database Service. Although I insist on calling it “my squirrel.” While MySQL has long been the worlds most popular open source database, shifting from transacting to analytics required way too much overhead and, ya know, work. With HeatWave you can run your OLTP and OLAP, don’t ask me to ever say those acronyms again, workloads directly from your MySQL database and eliminate the time consuming data movement and integration work, while also performing 1100X faster than Amazon Aurora, and 2.5X faster than Amazon Redshift, at a third of the cost. My thanks again to Oracle Cloud for sponsoring this ridiculous nonsense.

Corey: Back in 2019, I gave a talk at re:Invent—which is always one of those things that’s going to occasion comment—and the topic that we covered was building a vulnerability disclosure program built upon the story of a vulnerability that I reported into AWS. And it was a decent enough experience that I suggested at some point that you should talk about this publicly, and they said, “You should come talk about it with us.” And I did and it was a blast. But it suddenly became very clear, during the research for that talk and talking to people who’ve set those programs up is that look, one way or another, people are going to find vulnerabilities in what you do and how you do them. And if you don’t give them an easy way to report them to you, that’s okay.

You’ll find out about them in other scenarios when they’re on the front page of the New York Times. So, you kind of want to be out there and accessible to people. Now, there’s a whole story we can go into about the pros and cons of things like bug bounties and the rest, and of course, it’s a nuanced issue, but the idea of at least making it easy for people to wind up reporting things from that perspective is one of those key areas of outreach. Back in the early days of InfoSec, people would explore different areas of systems that they had access to, and very often they were charged criminally. Intel wound up having charges against one of their—I believe it was their employee or something, who wound up founding something and reporting it in an ethical way.

The idea of doing something like that is just ludicrous. You’re in that space a lot more than I am. Do you still see that sort of chilling effect slash completely not getting it when someone is trying to, in good faith, report security issues? Or has the world largely moved on from that level of foolishness?

Jackie: Both. The larger organizations that have mature security programs, and frankly, the organizations that have experienced a significant public breach, the organizations that have experienced pain are those that know better at this point and realize they do need to have a program, they do need to have a process and a procedure, and they need to have some kind of framework for folks to share information with them in a way that doesn’t cause them to respond with, “Are you extorting me? Is this blackmail?” As a cybersecurity professional working at my own security firm and also doing security research, I have reported dozens of vulnerabilities that I’ve identified, open buckets, for example. My partner at Spyglass and I built a SaaS application called Data Drifter a few years ago.

We were interviewed by NBC about this and NBC followed up on quite a few of our vulnerability disclosures and published an article. But what the software did was look for open buckets on Azure, AWS, and GCP and provide an analyst interface that allows a human to trawl through very large datasets and understand what they’re looking at. So, for example, one of the finds that we had was that musical.ly—musical-dot-L-Y, which was purchased by TikTok, eventually—had a big, large open bucket with a lot of data, and we couldn’t figure out how to report it properly. And they eventually took it down.

But you really had to try to understand what you were looking at; if you have a big bucket full of different data types, you don’t have a name on the bucket, and you don’t know who it belongs to because you’re not Google, or Amazon, or Microsoft, what do you do with this information? And so we spent a lot of time trying to reconcile open buckets with their owners and then contacting those owners. So, we’ve received a gamut of ranges of responses to vulnerability disclosure. On one hand, there is an established process at an organization that is visible by the way they respond and how they handle your inquiry. Some folks have ticketing systems, some folks respond directly to you from the security team, which is great, and you can really see and get an example of what their routing is inside the company.

And then other organizations really have no point of reference for that kind of thing, and when something comes into either their support channels or even directly into the cybersecurity team, they’re often scrambling for an effective way to respond to this. And it could go either way; it could get pretty messy at times. I’ve been threatened legally and I’ve been accused of extortion, even when we weren’t trying to offer some type of a service. I mean, you really never walk into a vulnerability disclosure scenario and then offer consulting services because they are going to see it as a marketing ploy and you never want to make that a marketing ploy. I mean, it’s just not… it’s not effective and it’s not ethical, it’s not the right thing to do.

So, it’s been interesting. [laugh]. I would recommend, if you are a person listening to this podcast who has some sort of pull in the information security department at your organization, I would recommend that you start with disclose.io, which was put together by Casey John Ellis and some other folks over at Bugcrowd and some other volunteers. It’s a really great starting point for understanding how to implement a vulnerability disclosure program and making sure that you are able to receive the information in a way that prevents a PR disaster.

Corey: My approach is controversial—I know this—but I believe that the way that you’re approaching this was entirely fatally flawed, of trying to report to people that they have an open S3 bucket. The proper way to do it is to upload reams of data to it because my operating theory is that they’re going to ignore a politely worded note from a security researcher, but they’re not going to ignore a $4 million surprise bill at the end of the month from AWS. That’ll get fixed tout suite. To be clear to the audience, I am kidding on this. Don’t do it. There’s a great argument that you can be charged criminally for doing such a thing. I’m kidding. It’s a fun joke. Don’t do it. I cannot stress that enough. We now go to Jackie for her laughter at that comment.

Jackie: [laugh].

Corey: There we go.

Jackie: I’m on cue. Well, a great thing about Data Drifter, that SaaS application that allowed analysts to review the contents of these open buckets, was that it was all JavaScript on the client-side, and so we weren’t actually hosting any of that data ourselves. So, they must have noticed some transfer fees that were excessive, but if you’re not looking at security and you have an infrastructure that isn’t well monitored, you may not be looking at costs either.

Corey: Costs are one of those things that are very aligned spiritually with security. It’s a trailing function that you don’t care about until right after you really should have cared about it. With security, it’s a bit of a disaster when it hits, whereas with those surprise bills, “Oh, okay. We wasted some money.” That’s usually, a, not front-page material and, b, it’s okay, let’s be responsible and fix that up where it makes sense, but it’s something that is never a priority. It’s never a ‘summon the board’ story for anything short of complete and utter disaster. So, I do feel a sense of spiritual alignment here.

Jackie: [laugh]. I can see that. That makes perfect sense.

Corey: Before we call this an episode, one other area that you’ve been active within is something called ‘threat modeling.’ What is it?

Jackie: So, threat modeling is a way to think strategically about cybersecurity. You want to defend, effectively, by understanding your organization as a collection of people, and you want to help non-technical staff support the cybersecurity program. So, the way to do that is potentially to give a human-centric focus to threat modeling activities. Threat modeling is a methodology for linking humans to an effective set of prioritized defenses for the most likely types of adversaries that they might face. And so essentially the process is identifying your subject and defining the scope of what you would like to protect.

Are you looking to protect this person’s personal life? Are you exclusively protecting their professional life or what they’re doing in relation to an organization? And you want to iterate through a few questions and document an attack tree. Then you would research some tactics and vulnerabilities, and implement defensive controls. So, in a nutshell, we want to know what assets does your subject have or have access to, that someone might want to spy, steal, or harm; you want to get an idea of what types of adversaries you can expect based on those assets or accesses that they have, and you then want to understand what tactics those adversaries are likely to use to compromise those assets or accesses, and you then transform that into the most effective defenses against those likely tactics.

So, using that in practice, you would typically build an attack tree that starts with the human at the center and lists out all of their assets and accesses. And then off of those, each of those assets or accesses, you would want to map out their adversary personas. So, for example, if I work at a bank and I work on wire transfers, my likely adversary would be a financially motivated cybercriminal, right? Pretty standard stuff. And we want to understand what are the methods that these actors are going to employ in order to get the job done.

So, in a common case, in a business email compromised context, folks might rely on a signer at a company to sign off on a wire transfer, and if the threat actor has an opportunity to gain access to that person’s email address or the mechanism by which they make that approval, then they may be able to redirect funds to their own wallet that was intended for someone else or a partner of the company. Adversaries tend to employ the least difficult approach; whatever the easiest way in is what they’re going to employ. I mean, we spend a lot of time in the field of information security and researching the latest vulnerabilities and attack paths and what are all the different ways that a system or a person or an application can be compromised, but in reality, the simplest stuff is usually what works, and that’s what they’re looking for. They’re looking for the easiest way in. And you can really observe that with ransomware, where attackers are employing a spray and pray methodology.

They’re looking for whatever they can find in terms of open attack surface on the net, and then they’re targeting organizations based on who they can compromise after the fact. So, they don’t start with an organization in mind, they might start with a type of system that they know they can easily compromise and then they look for those, and then they decide whether they’re going to ransomware that organization or not. So, it’s really a useful way, when you’re thinking about human-centric threat modeling, it’s really a useful way to completely map your valuables and your critical assets to the most effective ways to protect those. I hope that makes sense.

Corey: It very much does. It’s understanding the nature of where you start, where you stop, what is reasonable, what is not reasonable. Because like a lot of different areas—DR, for example—security is one of those areas you could hurl infinite money into and still never be done. It’s where do you consider it reasonable to start? Where do you consider it reasonable to stop? And without having an idea of what the model of threat you’re guarding against is, the answer is, “All the money,” which it turns out, boards are surprisingly reluctant to greenlight.

Jackie: Absolutely. We have a recurring problem and information security where we cannot measure return on investment. And so it becomes really difficult to try to validate a negative. It’s kind of like the TSA; the TSA can say that they’ve spent a lot of money and that nothing has happened or that any incidents have been limited in their scope due to the work that they’ve done, but can we really quantify the amount of money that DHS has absorbed for the TSA’s mission, and turned that into a really wonderful and measurable understanding of how we spent that money, and whether it was worth it? No, we can’t really. And so we’re always struggling with that insecurity, and I don’t think we’ll have an answer for it in the next ten years or so.

Corey: No, I suspect not, on some level. It’s one of those areas where I think the only people who are really going to have a holistic perspective on this are historians.

Jackie: I agree.

Corey: And sadly I’m not a cloud historian; I’m a cloud economist, a completely different thing I made up.

Jackie: [laugh]. Well, from my perspective, I think it’s a great title. And I agree with your thought about historians, and I look forward to finding out how they felt about what we did in the information security space, both political and non-political, 20, 30, and 40 years from now.

Corey: I hope to live long enough to see that. Jackie, thank you so much for taking the time to speak with me today. If people want to learn more about what you’re up to and how you view things, where can they find you?

Jackie: You can find me on Twitter at @hackingbutlegal.

Corey: Great handle. I love it.

Jackie: Thank you so much for having me.

Corey: Oh, of course. It is always great to talk with you. Jackie Singh, principal threat analyst, and incident responder at the Biden campaign. Obviously not there anymore. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast provider of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with a comment expressing an incoherent bigoted tirade that you will, of course, classify as a political opinion, and get you evicted from said podcast provider.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Jordan

Jordan is a self proclaimed “hacker.”

Links:

  • Twitter: https://twitter.com/jordansissel

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by “you”—gabyte. Distributed technologies like Kubernetes are great, citation very much needed, because they make it easier to have resilient, scalable, systems. SQL databases haven’t kept pace though, certainly not like no SQL databases have like Route 53, the world’s greatest database. We’re still, other than that, using legacy monolithic databases that require ever growing instances of compute. Sometimes we’ll try and bolt them together to make them more resilient and scalable, but let’s be honest it never works out well. Consider Yugabyte DB, its a distributed SQL database that solves basically all of this. It is 100% open source, and there's not asterisk next to the “open” on that one. And its designed to be resilient and scalable out of the box so you don’t have to charge yourself to death. It's compatible with PostgreSQL, or “postgresqueal” as I insist on pronouncing it, so you can use it right away without having to learn a new language and refactor everything. And you can distribute it wherever your applications take you, from across availability zones to other regions or even other cloud providers should one of those happen to exist. Go to yugabyte.com, thats Y-U-G-A-B-Y-T-E dot com and try their free beta of Yugabyte Cloud, where they host and manage it for you. Or see what the open source project looks like—its effortless distributed SQL for global apps. My thanks to Yu—gabyte for sponsoring this episode.

Corey: This episode is sponsored in part by our friends at VMware. Let’s be honest—the past year has been far from easy. Due to, well, everything. It caused us to rush cloud migrations and digital transformation, which of course means long hours refactoring your apps, surprises on your cloud bill, misconfigurations and headache for everyone trying manage disparate and fractured cloud environments. VMware has an answer for this. With VMware multi-cloud solutions, organizations have the choice, speed, and control to migrate and optimize applications seamlessly without recoding, take the fastest path to modern infrastructure, and operate consistently across the data center, the edge, and any cloud. I urge to take a look at vmware.com/go/multicloud. You know my opinions on multi cloud by now, but there's a lot of stuff in here that works on any cloud. But don’t take it from me thats: VMware.com/go/multicloud and my thanks to them again for sponsoring my ridiculous nonsense.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’ve been to a lot of conference talks in my life. I’ve seen good ones, I’ve seen terrible ones, and then I’ve seen the ones that are way worse than that. But we don’t tend to think in terms of impact very often, about how conference talks can move the audience.

In fact, that’s the only purpose of giving a talk ever—to my mind—is you’re trying to spark some form of alchemy or shift in the audience and convince them to do something. Maybe in the banal sense, it’s to sign up for something that you’re selling, or to go look at your website, or to contribute to a project, or maybe it’s to change the way they view things. One of the more transformative talks I’ve ever seen that shifted my outlook on a lot of things was at [SCALE 00:01:11] in 2012. Person who gave that talk is my guest today, Jordan Sissel, who, among many other things in his career, was the original creator behind logstash, which is the L in ELK Stack. Jordan, thank you for joining me.

Jordan: Thanks for having me, Corey.

Corey: I don’t know how well you remember those days in 2012. It was the dark times; we thought oh, the world is going to end; that wouldn’t happen until 2020. But it was an interesting conference full of a bunch of open-source folks, it was my local conference because I lived in Los Angeles. And it was the thing I looked forward to every year because I would always go and learn something new. I was in the trenches in those days, and I had a bunch of problems that looked an awful lot like other people’s problems, and having a hallway track where, “Hey, how are you solving this problem?” Was a big deal. I missed those days in some ways.

Jordan: Yeah, SCALE was a particularly good conference. I think I made it twice. Traveling down to LA was infrequent for me, but I always enjoyed how it was a very communal setting. They had dedicated hallway tracks. They had kids tracks, which I thought was great because folks couldn’t usually come to conferences if they couldn’t bring their kids or they had to take care of that stuff. But having a kids track was great, they had kids presenting. It felt more organic than a lot of other conferences did, and that’s kind of what drew me to it initially.

Corey: Yeah, it was my local network. It turns out that the Southern California tech community is relatively small, and we all go different lives. And it’s LA, let’s face it, I lived there for over a decade. Flaking as a way of life. So yeah, well, “Oh, we’ll go out and catch dinner. Ooh, have to flake at the last minute.” If you’re one of the good people, you tell people you’re flaking instead of just no-showing, but it happens.

But this was the thing that we would gather and catch up every year. And, “Oh, what have you been doing?” “Wow, you work in that company now? Congratulations, slash, what’s wrong with you?” It was fun, just sort of a central sync point. It started off as hanging out with friends.

And in those days, I was approaching the idea of, “You know what? I should learn to give a conference talk someday. But let’s be clear. People don’t give conference talks; legends give conference talks. And one day, I’ll be good enough to get on stage and give a talk to my peers at a conference.”

Now, the easy, cynical interpretation would be, “Well, but I saw your talk and I figured, hey, any jackhole can get up there. If he can do it, anyone can.” But that’s not at all how it wound up impacting me. You were talking about logstash, which let’s start there because that’s a good entry point. Logstash was transformative for me.

Before that, I’d spent a lot of time playing around with syslog, usually rsyslog, but there are other stories here of when a system does something and it spits out logs—ideally—how do you make sure you capture those logs in a reliable way so if you restart a computer, you don’t wind up with a gap in your logs? If it’s the right computer, it could be a gap in everything’s logs while that thing is coming back up. And let’s avoid single points of failure and the rest. And I had done all kinds of horrible monstrosities, and someone asked me at one point—

Jordan: [laugh]. Guilty.

Corey: Yeah. Someone said, “Well, there are a couple of options. Why don’t you use Splunk?” And the answer is that I don’t have a spare
princess lying around that I can ransom back to her kingdom, so I can’t afford it. “Okay, what about logstash?” And my answer was, “What’s a logstash?” And thus that sound was Pandora’s Box creaking open.

So, I started playing with it and realized, “Okay, this is interesting.” And I lost track of it because we have demands on our time. Then I was dragged into a session that you gave and you explained what logstash was. I’m not going to do nearly as good of a job as you can on this. What the hell was logstash, for folks who are not screaming at syslog while they first hear of it.

Jordan: All right. So, you mentioned rsyslog, and there’s—old is often a pejorative of more established projects because I don’t think these projects are bad. But rsyslog, syslog-ng, things like that were common to see for me as a sysadmin. But to talk about logstash, we need to go
back a little further than 2012. So, the logstash project started—

Corey: I disagree because I wasn’t aware of it until 2012. Until I become aware of something it doesn’t really exist. That’s right, I have the object permanence of an infant.

Jordan: [laugh].That’s fair. And I’ve always felt like perception is reality, so if someone—this gets into something I like to say, but if someone is having a bad time or someone doesn’t know about something, then it might as well not exist. So, logstash as a project started in 2008, 2009. I don’t remember when the first commits landed, but it was, gosh, it’s more than ten years ago now.

But even before that in college, I was fortunate to, through a network of friends, get a job as a sysadmin. And as a sysadmin, you stare at logs
a lot to figure out what’s going on. And I wanted a more interesting way to process the logs. I had taught myself regular expressions and it wasn’t finding joy in it… at all, like pretty much most people, probably. Either they look at regular expressions and just… evacuate with disgust, which is absolutely an appropriate response, or they dive into it and they have to use it for their job.

But it wasn’t enjoyable, and I found myself repeating stuff a lot. Matching IP addresses, matching strings, URLs, just trying to pull out useful information about what is going on?

Corey: Oh, and the timestamp problem, too. One of the things that I think people don’t understand who have not played in this space, is that all systems do have logs unless you’ve really pooched something somewhere—

Jordan: Yeah.

Corey: —and it shows that at this point in time, this thing happened. As we start talking about multiple computers and distributed systems—but even on the same computer—great, so at this time there was something that showed up in the system log because there was a disk event or something, and at the same time you have application logs that are talking about what the application running is talking about. And that is ideally using a somewhat similar system to do this, but often not. And the way that timestamps are expressed in these are radically different and the way that the log files themselves are structured. One might be timestamp followed by hostname followed by error code.

The other one might be hostname followed by a timestamp—in a different format—followed by a copyright notice because a big company got to it followed by the actual event notice, and trying to disambiguate all of these into a standardized form was first obnoxious, and secondly, very important because you want to see the exact chain of events. This also leads to a separate sidebar on making sure that all the clocks are synchronized, but that’s a separate story for another time. And that’s where you enter the story in many respects.

Jordan: Right. So, my thought around what led to logstash is you can take a sysadmin or software IT developer—whatever—expert, and you can sit them in front of a bunch of logs and they can read them and say, “That’s the time it happened. That’s the user who caused this action. This is the action.” But if you try and abstract and step away, and so you ask how many times did this action happen? When did this user appear? What time did this happen?

You start losing the ability to ask those questions without being an expert yourself, or sitting next to an expert and having them be your keyboard. Kind of a phenomenon I call the human keyboard problem where you’re speaking to a computer, but someone has to translate for you. And so in around 2004, I was super into Perl. No shocker that I enjoyed—ish. I sort of enjoyed regular expressions, but I was super into Perl, and there was a Perl module called Regexp::Common which is a library of regular expressions to match known things: IP addresses, certain kinds of timestamps, quoted strings, and whatnot.

Corey: And this stuff is always challenging because it sounds like oh, an IP address. One of the interview questions I hated the most someone asked me was write a regular expression to detect an IP address. It turns out that to do this correctly, even if you bound it to ipv4 only, the answer takes up multiple lines on a screen.

Jordan: Oh, for sure.

Corey: It’s enormous.

Jordan: It’s like a full page of—

Corey: It is.

Jordan: —of code you can’t read. And that’s one of the things that, it was sort of like standing on the shoulders of the person who came before; it was kind of an epiphany to me.

Corey: Yeah. So, I can copy and paste that into my code, but someone who has to maintain that thing after I get fired is going to be, “What the hell is this and what does it do?” It’s like it’s the blessed artifact that the ancients built it and left it there like it’s a Stargate sitting in your code. And it’s, “We don’t know how it works; we’re scared to break it, so we don’t even look at that thing directly. We just know that we put nonsense in, an IP address comes out, and let’s not touch it, ever again.”

Jordan: Exactly. And even to your example, even before you get fired and someone replaces you and looks at your regular expression, the problem I was having was, I would have this library of copy and pasteable things, and then I would find a bug, and edge case. And I would fix that edge case but the other 15 scripts that were using the same way regular expression, I can’t even read them anymore because I don’t carry that kind of context in my head for all of that syntax. So, you either have to go back and copy and paste and fix all those old regular expressions. Or you just say, “You know what? We’re not going to fix the old code. We have a new version of it that works here, but everywhere else this edge case fails.”

So, that’s one of the things that drew me to the Regexp::Common library in Perl was that it was reusable and things had names. It was, “I want to match an IP address.” You didn’t have to memorize that long piece of text to precisely and accurately accept only regular expressions and rejects things that are not. You just said, “Give me the regular expression that matches an IP.” And from that library gave me the idea to write grok.

Well, if we could name things, then maybe we could turn that into some kind of data structure, sort of the combination of, “I have a piece of log data, and I as an expert, I know that’s an IP address, that’s the username, and that’s the timestamp.” Well, now I can apply this library of regular expressions that I didn’t have to write and hopefully has a unit test suite, and say, now we can pull out instead of that plain piece of text that is hard to read as a non-expert, now I can have a data structure we can format however we want, that non-experts can see. And even experts can just relax and not have to be full experts all the time, using that part of your brain. So, now you can start getting towards answering search-oriented questions. “How many login attempts happened yesterday from this IP address?”

Corey: Right. And back then, the way that people would do these things was Elasticsearch. So, that’s the thing you shove all your data into in a bunch of different ways and you can run full-text queries on it. And that’s great, but now we want to have that stuff actually structured, and that is sort of the magic of logstash—which was used in conjunction with Elasticsearch a lot—and it turns out that typing random SQL queries in the command line is not generally how most business users like to interact with this stuff, seems to be something dashboard-y-like, and the project that folks use for that was Kibana. And ELK Stack became a thing because Elasticsearch in isolation can do a lot but it doesn’t get you all the way there for what people were using to look at logs.

Jordan: You’re right.

Corey: And Kibana is also one of the projects that Elastic owned, and at some point, someone looks around, like, “Oh, logstash. People are using that with us an awful lot. How big is the company that built that? Oh, it’s an open-source project run by some guy? Can we hire that guy?” And the answer is, “Apparently,” because you wound up working as an Elastic employee for a while.

Jordan: Yeah. It was kind of an interesting journey. So, in the beginning of logstash in 2009, I kind of had this picture of how I wanted to solve log processing search challenges. And I broke it down into a couple of parts of visualization—to be clear, I broke it down in my head, not into code, but visualization, kind of exploration, there’s the processing and transmission, and then there’s storage and search. And I only felt confident really attending to a solution for one of those parts. And I picked log processing partly because I already had a jumpstart from a couple of years prior, working on grok and feeling really comfortable with regular expressions. I don’t want to say good because that’s—

Corey: You heard it here first—

Jordan: [laugh].

Corey: —we found the person that knows regular expressions. [laugh].

Jordan: [laugh]. And logstash was being worked on to solve this problem of taking your data, processing it, and getting it somewhere. That’s why logstash has so many outputs, has so many inputs, and lots of filters. And about I think a year into building logstash, I had experimented with storage and search backends, and I never found something that really clicked with me. And I was experimenting with Leucine, and knowing that I could not complete this journey because that the problem space is so large, it would be foolish of me to try to do distributed log stores or anything like that, plus visualization.

I just didn’t have the skills or the time in the day. I ended up writing a frontend for logstash called logstash-web—naming things is hard—and I wasn’t particularly skilled or attentive to that project, and it was more of a very lightweight frontend to solve the visualization, the exploration aspect. And about a year into logstash being alive, I found Elasticsearch. And what clicked with me from being a sysadmin and having worked at large data center companies in the past is I know the logs on a single system are going to quickly outgrow it. So, whatever storage system will accept these logs, it’s got to be easy to add new storage.

And Elasticsearch first-day promise was it’s distributed; you can add more nodes and go about your day. And it fulfilled that promise and I think it still fulfills that promise that if you’re going to be processing terabytes of data, yeah, just keep dumping it in there. That’s one of the reasons I didn’t try and even use MySQL, or Postgres, or other data systems because it didn’t seem obvious how to have multiple storage servers collecting this data with those solutions, for me at the time.

Corey: It turns out that solving problems like this that are global and universal lead to massive adoption very quickly. I want to get this back a bit before you wound up joining Elastic because you get up on stage and you talked through what this is. And I mentioned at the start of this recording, that it was one of those transformative talks. But let’s be clear here, I don’t remember 95% of how logstash works. Like, the technology you talked about ten years ago is largely outmoded slash replaced slash outdated today. I assure you, I did not take anything of note whatsoever from your talk regarding regular expressions, I promise. And—

Jordan: [laugh]. Good.

Corey: But that’s not the stuff that was transformative to me. What was, was the way that you talked about these things. And there was the first time I’d ever heard the phrase that if a new user has a bad time, it’s a bug. This was 2012. The idea of empathy hadn’t really penetrated into the ops and engineering spaces in any meaningful way yet. It was about gatekeeping, it was about, “Read the manual fool”—

Jordan: Yes.

Corey: —if people had questions. And it was actively user-hostile. And it was something that I found transformative of, forget the technology piece for a second; this is a story about how it could be different. Because logstash was the vehicle to deliver a message that transcended far beyond the boundaries of how to structure your logs, or maybe the other boundaries of regular expressions, I’m never quite sure where those things start and stop. But it was something that was actively transformative where you’re on stage as someone who is a recognized authority in the space, and you’re getting up there and you’re sending an implicit message—both explicitly and by example—of be nice to people; demonstrate empathy. And that left a hell of an impact. And—

Jordan: Thank you.

Corey: I wound up doing a spot check just now, and I wound up looking at this and sure enough, early in 2013, I wound up committing—it’s still in the history of the changelog for logstash because it’s open-source—I committed two pull requests and minutes apart, two submissions—I don’t know if pull requests were even a thing back then—but it wound up in the log. Because another project you were renowned for was fpm: Effing Package Manager if I’m—is that what the acronym stands for, or am I misremembering?

Jordan: [laugh]. We’ll go with that. I’m sure, vulgar viewers will know what the F stands for, but you don’t have to say it. It’s just Effing Package Management.

Corey: Yeah.

Jordan: But yeah, I think I really do believe that if a user, especially if a new user has a bad time, it’s a bug, and that came from many years of participating at various levels in open-source, where if you came at it with a tinkerer’s or a hacker’s mindset and you think, “This project is great. I would like it to do one additional thing, and I would like to talk to someone about how to make it do that one additional thing.” And you go find the owners or the maintainers of that project, and you come in with gusto and energy, and you describe what you want to do and, first, they say, “What you want to do is not possible.” They don’t even say they don’t want to do it; they frame the whole universe against you. “It’s not possible. Why would you want to do that? If you want to make that, do it yourself.”

You know, none of these things are an extended hand, a lowered ladder, an open door, none of those. It’s always, “You’re bothering me. Go away. Please read the documentation and see where we clearly”—which they don’t—“Document that this is not a thing we’re interested in.” And I came to the conclusion that any future open-source or collaborative work that I worked on, it’s got to be from a place where, “You’re welcome, and whatever contributions or participation levels you choose, are okay. And if you have an idea, let’s talk about it. If you’re having a bad time, let’s figure out how to solve it.”

Maybe the solution is we point you in the right direction to the documentation, if documentation exists; maybe we find a bug that we need to fix. The idea that the way to build communities is through kindness and collaboration, not through walls or gatekeeping or just being rude. And I really do think that’s one of the reasons logstash became so successful. I mean, any particular technology could have succeeded in the space that logstash did, but I believe that it did so because of that one piece of framework where if a new user has a bad time, it’s a bug. Because to me, that opens the door to say, “Yeah, you know what? Some of the code I write is not going to be good. Or, the thing you want to do is undocumented. Or the documentation is out of date. It told you a lie and you followed the documentation and it misled you because it’s incorrect.”

We can fix that. Maybe we don’t have time to fix it right now. Maybe there’s no one around to fix it, but we can at least say, “You know what? That information is incorrect, and I’m sorry you were misled. Come on into the community and we’ll figure it out.” And one of the patterns I know is, on the IRC channel, which is where the logstash real-time community chat… I don’t know how to describe that.

Corey: No, it was on freenode. That’s part of the reason I felt okay, talking to you. At that point. I was volunteer network staff. This is before freenode turned into basically a haven for Nazis this past year.

Jordan: Yeah. It was still called lilo… lilonet [crosstalk 00:20:20]—

Corey: No, the open freenode network, that predates me. This was—yeah, lilo—

Jordan: Okay.

Corey: —died about six years prior. But—

Jordan: Oh, all right.

Corey: Freenode’s been around a long time. What make this thing work was that I was network staff, and that means that I had a bit of perceived authority—it’s a chat room; not really—but it was one of those things where it was at least, “Okay, this is not just some sketchy drive-by rando,” which I very much was, but I didn’t present that way, so I could strike up conversations. But with you talking about this stuff, I never needed to be that person. It was just if someone wants to pitch in on this, great; more hands make lighter work. Sure.

Jordan: Yeah, for sure.

Corey: And for me, the interesting part is not even around the logstash aspects so much; it’s your other project, fbm. Well, one of your other projects. Back in 2012, that was an interesting year for me. Another area that got very near and dear to my heart in open-source world was the SaltStack project; I was contributor number 15. And I didn’t know how Python worked. Not that I do now, but I can fake it better now.

And Tom Hatch, the guy that ran the project before it was a company was famous for this where I could send in horrifying levels of code, and every time he would merge it in and then ten minutes later, there would be another patch that comes in that fixes all bugs I just introduced and it was just such a warm onboarding. I’m not suggesting that approach and I’m not saying it’s scalable, but I started contributing. And I became the first Debian and Ubuntu packager for SaltStack, which was great. And I did a terrible job at it because—let me explain. I don’t know if it’s any better now, but back in those days, there were multiple documentation sources on the proper way to package software.

They were all contradictory with each other, there was no guidance as to when to follow each one, there was never a, “You know nothing about packaging; here’s what you need to know, step-by-step,” and when you get it wrong, they yell at you. And it turns out that the best practice then to get it formally accepted upstream—which is what I did—is do a crap-ass job, and then you’ll wind up with a grownup coming in, like, “This is awful. Move.” And then they’ll fix it and yell at you, and gatekeep like hell, and then you have a package that works and gets accepted upstream because the magic incantation has been said somewhere. And what I loved about fpm was that I could take any random repo or any source tarball or anything I wanted, run it through with a single command, and it would wind up building out a RPM and a Deb file—and I don’t know what else it’s supported; those are the ones I cared about—that I could then install on a system. I put in a repo and add that to a sources list on systems, and get to automatically install so I could use configuration management—like SaltStack—to wind up installing custom local packages. And oh, my God, did the packaging communities for multiple different distros hate you—

Jordan: Yep.

Corey: —and specifically what you had built because this was not the proper way to package. How dare you solve an actual business problem someone has instead of forcing them to go to packaging school where the address is secret, and you have to learn that. It was awful. It was the clearest example that I can come up with of gatekeeping, and then you’re coming up with fbm which gets rid of user pain, and I realized that in that fight between the church of orthodoxy of, “This is how it should be done,” and the, “You’re having a problem; here’s a tool that makes it simple,” I know exactly what side of that line I wanted to be on. And I hadn’t always been previously, and that is what clarified it for me.

Jordan: Yeah, fbm was a really delightful enjoyment for me to build. The origins of that was I worked at a company and they were all… I think, at that time, we were RPM-based, and then as folks tend to do, I bounced around between jobs almost every year, so I went from one place that—

Corey: Hey, it’s me.

Jordan: [laugh]. Right? And there’s absolutely nothing wrong with leaving every year or staying longer. It’s just whatever progresses your career in the way that you want and keeps you safe and your family safe. But we were using RPM and we were building packages already not following the orthodoxy.

A lot of times if you ask someone how to build a package for Fedora, they’ll point you at the Maximum RPM book, and that’s… a lot of pages, and honestly, I’m not going to sit down and read it. I just want to take a bunch of files, name it, and install it on 30 machines with Puppet. And that’s what we were doing. Cue one year later, I moved to a new company, and we were using Debian packages. And they’re the same thing.

What struck me is they are identical. It’s a bunch of files—and don’t pedant me about this—it’s a bunch of files with a name, with some other sometimes useful metadata, like other names that you might depend on. And I really didn’t find it enjoyable to transfer my knowledge of how to build RPMs, and the tooling and the structures and the syntaxes, to building Debian packages. And this was not for greater publication; this was I have a bunch of internal applications I needed to package and deploy with, at the time it was Puppet. And it wasn’t fun.

So, I did what we did with grok which was codify that knowledge to reduce the burden. And after a few, probably a year or so of that, it really dawned on me that a generality is all packaging formats are largely solving the same problem and I wanted to build something that was solving problems for folks like you and me: sysadmins, who were handed a pile of code and they needed to get it into production. And I wasn’t interested in formalities or appeasing any priesthoods or orthodoxies about what really—you know, “You should really shine your package with this special wax,” kind of thing. Because all of the documentation for Debian packages, Fedora packages are often dedicated to those projects. You’re going to submit a package to Fedora so that the rest of the world can use it on Fedora. That wasn’t my use case.

Corey: Right. I built a thing and a thing that I built is awesome and I want the world to use it, so now I have to go to packaging school? Not just once but twice—

Jordan: Right.

Corey: —and possibly more. That’s awful.

Jordan: Or more. Yeah. And it’s tough.

Corey: This episode is sponsored in part by our friends at Jellyfish. So, you’re sitting in front of your office chair, bleary eyed, parked in front of a powerpoint and—oh my sweet feathery Jesus its the night before the board meeting, because of course it is! As you slot that crappy screenshot of traffic light colored excel tables into your deck, or sift through endless spreadsheets looking for just the right data set, have you ever wondered, why is it that sales and marketing get all this shiny, awesome analytics and inside tools? Whereas, engineering basically gets left with the dregs. Well, the founders of Jellyfish certainly did. That’s why they created the Jellyfish Engineering Management Platform, but don’t you dare call it JEMP! Designed to make it simple to analyze your engineering organization, Jellyfish ingests signals from your tech stack. Including JIRA, Git, and collaborative tools. Yes, depressing to think of those things as your tech stack but this is 2021. They use that to create a model that accurately reflects just how the breakdown of engineering work aligns with your wider business objectives. In other words, it translates from code into spreadsheet. When you have to explain what you’re doing from an engineering perspective to people whose primary IDE is Microsoft Powerpoint, consider Jellyfish. Thats Jellyfish.co and tell them Corey sent you! Watch for the wince, thats my favorite part.

Corey: And this gets back to what I found of—it was rare that I could find a way to contribute to something meaningfully, and I was using logstash after your talk, I’d started using it and rolling it out somewhere, and I discovered that there wasn’t a Debian package for it—the environment I was in at that time—or Ubuntu package, and, “Hey Jordan, are you the guy that wrote fpm and there isn’t a package here?” And the thing is is that you would never frame it this way, but the answer was, of course, “Pull requests welcome,” which is often an invitation to do free volunteer work for companies, but this was an open-source project that was not backed by a publicly-traded company; it was some guy. And of course, I’ll pitch in on that. And I checked the commit log on this for what it is that I see, and sure enough, I have two commits. The first one was on Sunday night in February of 2013, and my commit message was, “Initial packaging work for Deb building.” And sure enough, there’s a bunch of files I put up there and that’s great. And my second and last commit was 12 minutes later saying, “Remove large binary because I’m foolish.” Yeah.

Jordan: Was that you? [laugh].

Corey: Yeah. Oh, yeah, I’m sure—yeah, it was great. I didn’t know how Git worked back then. I’m sure it’s still in the history there. I wonder how big that binary is, and exactly how much I have screwed people over in the last decade since.

Jordan: I’ve noticed this over time. And every now and then you’d be—I would be or someone would be on a slow internet connection—which again, is something that we need to optimize for, or at least be aware of and help where we can—someone would be cloning logstash on an airplane or something like that, or rural setting, and they would say, “It gets stuck at 76% for, like, ten minutes.” And you would go back and dust off your tome of how to use Git because it’s very difficult piece of software to use, and you would find this one blob and I never even looked at it who committed it or whatever, but it was like I think it was 80 Megs of a JAR file or a Debian package that was [unintelligible 00:28:31] logstash release. And… [laugh] it’s such a small world that you’re like, yep, that was me.

Corey: Oh, yeah. Oh, yeah. Let’s check this just for fun here. To be clear, the entire repository right now is 167 Megs, so that file that I had up there for all of 13 minutes lives indelibly in Git history, and it is fully half of the size—

Jordan: Yep.

Corey: —of the entirety of the logstash project. All right, then. I didn’t realize this was one of those confess your sins episodes, but here we are.

Jordan: Look, sometimes we put flags on the moon, sometimes we put big files in git. You could just for posterity, we could go back and edit the history and remove that, but it never became important to do it, it wasn’t loud, people weren’t upset enough by it, or it didn’t come up enough to say, “You know what? This is a big file.” So, it’s there. You left your mark.

Corey: You know, we take what we can get. It’s an odd time. I’ll have to do some digging around; I’m sure I’ll tweet about this as soon as I get a bit more data on it, but I wonder how often people have had frustration caused by that. There’s no ill intent here, to be very clear, but it was instead, I didn’t know how Git worked very well. I didn’t know what I was doing in a lot of respects, and sure enough in the fullness of time, some condescending package people came in and actually made this right.

And there is a reasonable, responsible package now because, surprise, of course there is. But I wonder how much inadvertent pain I caused people by that ridiculous commit. And it’s the idea of impact and how this stuff works. I’m not happy that people are on a plane with a slow connection had a wait an extra minute or two to download that nonsense. It’s one of those things that is, oops. I feel like a bit of a heel for that, not for not knowing something, but for causing harm to folks. Intent doesn’t outweigh impact. There is a lesson in there for it.

Jordan: Agreed. On that example, I think one of the things… code is not the most important thing I can contribute to a project, even though I feel very confident in my skills in programming in a variety of environments. I think the number one thing I can do is listen and look for sources of pain. And people would come in and say, “I can’t get this to work.” And we would work together and figure out how to make it work for their use case, and that could result in a new feature, a bug fix, or some documentation improvements, or a blog post, or something like that.

And I think in this case, I don’t really recall any amount of noise for someone saying, “Cloning the Git repository is just a pain in the butt.” And I think a lot of that is because either the people who would be negatively impacted by that weren’t doing that use case, they were downloading the releases, which were as small as we can possibly get them, or they were editing files using the GitHub online edit the file thing, which is a totally acceptable, it’s perfectly fine way to do things in Git. So, I don’t remember anyone complaining about that particular file size issue. The Elasticsearch repository is massive and I don’t think it even has binaries. It just has so much more—

Corey: Someone accidentally committed their entire production test data set at one point and oops-a-doozy. Yeah, it’s not the most egregious harm I’ve ever caused—

Jordan: Yeah.

Corey: —but it’s there. The thing that, I guess, resonates with me and still does is the lessons I learned from you, I could sum them up as being not just empathy-driven—because that’s the easy answer—but the other layers were that you didn’t need to be the world’s greatest expert in
a thing in order to credibly give a conference talk. To be clear, you were miles ahead of me and still are in a lot of different areas—

Jordan: Thanks.

Corey: —and that’s fine. But you don’t need to be the—like, you are not the world’s greatest expert on empathy, but that’s what I took from the talk and that’s what it was about. It also taught me that things you can pick up from talks—and other means—there are things you can talk about in terms of technology and there are things you can talk about in terms of people, and the things about people do not have expiration dates in the same way that technology does. And if I’m going to be remembered for impact on people versus impact on technology, for me, there’s no contest. And you forced me to really think about a lot of those things that it started my path to, I guess, becoming a public speaker and then later all the rest that followed, like this podcast, the nonsense on Twitter, and all the rest. So, it is, I guess, we can lay the responsibility for all that at your feet. Enjoy the hate mail.

Jordan: Uhh, my email address is now closed. I’m sorry.

Corey: Exactly.

Jordan: Well, I appreciate the kind words.

Corey: We’ll get letters on this one.

Jordan: [laugh].

Corey: It’s the impact that people have, and someti—I don’t think you knew at the time that that’s the impact you were having. It matters.

Jordan: I agree. I think a lot of it came from how do I want to experience this? And it was much later that it became something that was really outside of me, in the sense that it was building communities. One of the things I learned shortly after—or even just before—joining Elastic was how many folks were looking to solve a problem, found logstash, became a participant in the community, and that participation could just be anything, just hanging out on IRC, on the mailing list, whatever, and the next step for them was to get a better paying job in an environment they enjoyed that helped them take the next step in their career. Some of those people came to work with me at Elastic; some of them started to work on the logstash team at some point they decided because a lot of logstash users were sysadmins.

And on the logstash team, we were all developers; we weren’t sysadmins, there was nothing to operate. And a lot of folks would come on board and they were like, “You know what? I’m not enjoying writing Ruby for my job.” And they could take the next step to transition to the support team or the sales engineer team, or cloud operations team at Elastic. So, it was really, like you mentioned, it has nothing to do with the technology of—to me—why these projects are important.

They became an amplifier and a hand to pull people up to go the next step they need to go. And on the way maybe they can make a positive impact in the communities they participate in. If those happen to be fpm or logstash, that’s great, but I think I want folks to see that technology doesn’t have to be a grind of getting through gatekeepers, meeting artificial barriers, and things like that.

Corey: The thing that I took, too, is that I gave a talk in 2015 or’16, which is strangely appropriate now: “Terrible ideas in Git.” And yes, checking large binaries in is one of the terrible ideas I talk about. It’s Git through counter-example. And around that time, I also gave a talk for a while on how to handle a job interview and advance your career. Only one of those talks has resulted in people approaching me even years later saying that what I did had changed aspects of their life. It wasn’t the Git one. And that’s the impact it comes down to. That is the change that I wanted to start having because I saw someone else do it and realized, you know, maybe I could possibly be that good someday. Well, I’d like to think I made it, on some level.

Jordan: [laugh]. I’m proud of the impact you’ve made. And I agree with you, it is about people. Even with fpm where I was very selfishly tickling my own itch, I don’t want to remember all of this stuff and I also enjoy operating outside of the boundaries of a church or whatever the priesthoods that say, “This is how you must do a thing,” I knew there was a lot of folks who worked at jobs and they didn’t have authority, and they had to deploy something, and they knew if they could just package it into a Debian format, or an RPM format, or whatever they needed to do, they could get it deployed and it would make their lives easier. Well, they didn’t have the time or the energy or the support in order to learn how to do that and fpm brought them that success where you can say, “Here’s a bunch of files; here’s a name, poof, you have a package for whatever format you want.”

Where I found fpm really take off is when Gem and Python and Node.js support were added. The sysadmins were kind of sandwiched in between—in two impossible worlds where they are only authorized to deploy a certain package format, but all of their internal application developer teams were using Node.js and newer technologies, and all of those package formats were not permitted by whoever had the authority to permit those things at their job. But now they had a tool that said, “You know what? We can just take that thing, we’ll take Django and Python, and we’ll make it an RPM and we won’t have to think a lot about it.”

And that really, I think—to me, my hope was that it de-stresses that sort of work environment where you’re not having to do three weeks of brand new work every time someone releases something internally in your company; you can just run a script that you wrote a month ago and maintain it as you go.

Corey: Wouldn’t that be something?

Jordan: [laugh]. Ideally, ideally.

Corey: Jordan, I want to thank you for not only the stuff you did ten years ago, but also the stuff you just said now. If people want to learn more about you, how you view the world, see what you’re up to these days, where can they find you?

Jordan: I’m mostly active on Twitter, at @jordansissel, all one word. Mostly these days, I post repair stuff I do on the house. I’m a stay-at-home full0 time dad these days, and… I’m still doing maintenance on the projects that need maintenance, like fpm or xdotool, so if you’re one of those users, I hope you’re happy. If you’re not happy, please reach out and we’ll figure out what the next steps can be. But yeah. If you like bugs, especially spiders—or if you don’t like spiders and you want to like spiders, check me out on Twitter. I’m often posting macro photos, close-up photos of butterflies, bees, spiders, and the like.

Corey: And we will, of course, throw links to that in the [show notes 00:38:10]. Jordan, thank you so much for your time today. It’s appreciated.

Jordan: Thank you, Corey. It’s good talking to you.

Corey: Jordan Sissel, founder of logstash and currently, blissfully, not working on a particular corporate job. I envy him, some days. I’m Cloud
Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an angry comment in which you have also embedded a large binary.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Jesse

Jesse Vincent is the cofounder and CTO of Keyboardio, where he designs and manufactures high-quality ergonomic mechanical keyboards. In previous lives, he served as the COO of VaccinateCA, volunteered as the project lead for the Perl programming language, created both the leading open source issue tracking system RT: Request tracker and K-9 Mail for Android.

Links:

  • Keyboardio: https://keyboard.io
  • Obra: https://twitter.com/obra

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: You could build you go ahead and build your own coding and mapping notification system, but it takes time, and it sucks! Alternately, consider Courier, who is sponsoring this episode. They make it easy. You can call a single send API for all of your notifications and channels. You can control the complexity around routing, retries, and deliverability and simplify your notification sequences with automation rules. Visit courier.com today and get started for free. If you wind up talking to them, tell them I sent you and watch them wince—because everyone does when you bring up my name. Thats the glorious part of being me. Once again, you could build your own notification system but why on god’s flat earth would you do that?

Corey: This episode is sponsored in part by our friends at Jellyfish. So, you’re sitting in front of your office chair, bleary eyed, parked in front of a powerpoint and—oh my sweet feathery Jesus its the night before the board meeting, because of course it is! As you slot that crappy screenshot of traffic light colored excel tables into your deck, or sift through endless spreadsheets looking for just the right data set, have you ever wondered, why is it that sales and marketing get all this shiny, awesome analytics and inside tools? Whereas, engineering basically gets left with the dregs. Well, the founders of Jellyfish certainly did. That’s why they created the Jellyfish Engineering Management Platform, but don’t you dare call it JEMP! Designed to make it simple to analyze your engineering organization, Jellyfish ingests signals from your tech stack. Including JIRA, Git, and collaborative tools. Yes, depressing to think of those things as your tech stack but this is 2021. They use that to create a model that accurately reflects just how the breakdown of engineering work aligns with your wider business objectives. In other words, it translates from code into spreadsheet. When you have to explain what you’re doing from an engineering perspective to people whose primary IDE is Microsoft Powerpoint, consider Jellyfish. Thats Jellyfish.co and tell them Corey sent you! Watch for the wince, thats my favorite part.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. As you folks are well aware by now, this show is at least ostensibly about the business of cloud. And that’s intentionally overbroad. You can fly a boat through it, which means it’s at least wider than the Suez Canal.

And that’s all well and good, but what do all of these cloud services have in common? That’s right, we interact with them via typing on keyboards. My guest today is Jesse Vincent, who is the founder of Keyboardio and creator of the Model 01 heirloom-grade keyboard, which is sitting on my desk that sometimes I use, sometimes it haunts me. Jesse, thank you for joining me.

Jesse: Hey, thanks so much for having me, Corey.

Corey: So, mechanical keyboards are one of those divisive things that, back in the before times when we were all sitting in offices, it was an express form of passive aggression, where, “I don’t like the people around me, and I’m going to show it to them with things that can’t really complain about. So, what is the loudest keyboard I can get?” Style stuff. And some folks love them, some folks can’t stand them. And most folks to be perfectly blunt, do not seem to care.

Jesse: So, it’s not actually about them being loud, or it doesn’t have to be. Mechanical keyboards can be dead silent; they can be as quiet as anything else. There’s absolutely a subculture that is into things that are as loud as they possibly can be; you know, sounds like there’s a cannon going off on somebody’s desk. But you can also get absolutely silent mechanical switches that are more dampened than your average keyboard. For many, many people, it’s about comfort, it is about the key feel.

A keyboard is supposed to have a certain feeling and these flat rectangles that feel like you’re typing on glass, they don’t have that feeling and they’re not good for your fingers. And it’s been fascinating over the past five or six years to watch this explosion in interest in good keyboards again.

Corey: I learned to first use a computer back on an old IBM 286 in the ’80s. And this obviously had a Model M—or damn close to it—style buckling spring keyboard. It was loud and I’m nostalgic about the whole thing. True story I’ve never told on this podcast before; I was a difficult child when I was five years old, and I was annoyed because my parents went out of the house and my brother was getting more attention than I was. I poured a bucket of water into the keyboard.

And to this day, I’m surprised my father didn’t murder me after that. And we wound up after having a completely sealing rubber gasket on top of this thing. Because this was the ’80s; keyboards were not one of those, “Oh, I’m going to run down to the store and pick up another one for $20.” This was at least a $200 whoops-a-doozy. And let’s just say that it didn’t endear me to my parents that week.

Jesse: That’s funny because that keyboard is one that actually probably would have dried out just fine. Not like the Microsoft Naturals that I used to carry in the mid-’90s. Those white slightly curved ones. That was my introduction to ergonomic keyboards and they had a fatal flaw as many mid-’90s Microsoft products did. In this case, they melted in the rain; the circuit traces inside were literally wiped away by water. If a cup of water got in that keyboard, it was gone.

Corey: Everyone has a story involving keyboard and liquids at some point, or they are the most careful people that are absolutely not my people whatsoever because everyone I hang out with is inherently careless. And over time I used other keyboards as I went through my life and never had strong opinions on them, and then I got to play with a mechanical keyboard had brought all that time rushing back to me of, “Oh, yeah.” And my immediate thought is, “Oh, this is great. I wonder if I could pour water into it? No, no.”

And I started getting back into playing with them and got what I thought was the peak model keyboard from Das Keyboards which, there was the black keyboard with no writing on it at all. And I learned I don’t type nearly as well as I thought I did in those days. And okay. That thing sat around gathering dust and I started getting a couple more and a couple more, and it turns out if you keep acquiring mechanical keyboards, you can turn an interest into a problem but you can also power your way through to the other side and become a collector. And I started building my own for a while and I still have at least a dozen of them in various states of assembly here.

It was sort of a fun hobby that I got into, and for me at least it was, why do I want to build a keyboard myself? Is it, do I believe intrinsically that I can build a better keyboard than I can buy? Absolutely not. But everything else I do in my entire career as an engineer until that point had been about making the bytes on the screen go light up in different patterns. That was it.

This was something that I had built that I could touch with my hands and was still related to the thing that I did, and was somewhat more forgiving than other things that I could have gotten into, like you know, woodworking with table saws that don’t realize my arm it just lopped off.

Jesse: Oh, you can burn yourself pretty good with a soldering iron.

Corey: Oh, absolutely I can.

Jesse: But yeah, no, I got into this in a similar-sounding story. I had bad wrists throughout my career. I was a programmer and a programming
manager and CEO. And my wrist hurts all the time, and I’d been through pretty much every ergonomic keyboard out there. If you seen the one where you stick your fingers into little wells, and each finger you can press back forth, left, right, and down, the ones that looked like they were basically a pair of flat capacitive surfaces from a company that later got bought by Apple and turned into the iPads touch technology,
Microsoft keyboards, everything. And nothing quite felt right.

A cloud startup I had been working on cratered one summer. Long story short, the thing went under for kind of sad reasons and I swore I was going to take a year off to screw around and figure out what the next thing was going to be. And at some point, I noticed there were people on the internet building their own keyboards. This was not anything I had ever done before. When I started soldering, I did figure out that I must have soldered before because it smelled familiar, but this was supposed to be a one-month project to build myself a single keyboard.

And I saw that people on the internet were doing it, I figured, eh, how hard could it be? Just one of those things that Perl hackers are apt to say. Little did I know. It’s now, I want to say something like eight years later, and my one-month project to build one keyboard has failed thousands and thousands and thousands of times over as we’ve shipped thousands of keyboards to, oh God, it’s like 75 or 78 countries.

Corey: And it’s great. It’s well made. The Model 01 that I got was part of an early Kickstarter batch. My wife signed me up for it—because she knew I was into this sort of thing—as a birthday gift. And then roughly a year later, if memory serves, it showed up and that was fine.

Again, it’s Kickstarter is one of those, this might just be an aspirational gift. We don’t know. And—because, Kickstarter—but it was fun. And I use it. It’s great.

I like a lot of the programmability aspects of it. There are challenges. I’m not used to using ergonomic keyboards, and the columnar layout is offset to a point where I miss things all the time. And if you’re used to typing rapidly, in things like chats, or Twitter or whatnot, were rapid responses valuable, it’s frustrating trying to learn how a new keyboard layout works.

Jesse: Absolutely. So, we got some advice very early on from one of the research scientists who helped Microsoft with their design for their natural keyboards, and one of the things that he told us was, “You will probably only ever get one chance to make a keyboard; almost every company that makes a keyboard fails, and so you should take one of the sort of accepted designs and make a small improvement to help push the industry forward. You don’t want to go do something radical and have nobody like it.”

Corey: That’s very reasonable advice and also boring. Why bother?

Jesse: Well, we walked away from that with a very different take, which was, if we’re only going to get one chance of this, we’re going to do the thing we want to make.

Corey: Yeah.

Jesse: And so we did a bunch of stuff that we got told might be difficult to do or impossible. We designed our own keycaps from scratch. We milled the enclosure out of hardwood. When we started, we didn’t know where we were manufacturing, but we did specify that the wood was going to be Canadian maple because it grows like a weed, and as you know, not in danger of being made extinct. But when you’re manufacturing in southern China and you’re manufacturing with Canadian maple, that comes on a boat from North America.

Corey: There’s something to be said for the globalization supply chain as we see things shipped back and forth and back and forth, and it seems ridiculous but the economics are there it’s—

Jesse: Oh, my God. Now, this year.

Corey: Yeah [laugh], there’s that.

Jesse: Supply chains are… how obscenity-friendly is this podcast? [laugh].

Corey: Oh, we can censor anything that’s too far out. Knock yourself out.

Jesse: Because what I would ordinarily say is the supply chains are [BLEEP].

Corey: Yep, they are.

Jesse: Yeah. This time around, we gave customers the—for the Model 100, which is our new keyboard that the Kickstarter just finished up for—we gave customers the choice of that nice Canadian maple or walnut. We got our quotes in advance. You know, our supplier confirmed wood was no problem a few months in advance. And then the night before the campaign launched, our wood supplier got in touch and said, “So, there are no walnut planks that are wide enough to be had in all of southern China. There are some supply chain issues due to the global container shortage. We don’t know what we’re going to be able to do. Maybe you could accept it if we did butcher block style walnut and glued planks together.”

They made samples and then a week later, instead of FedExing us the samples, I got a set of photographs with a whole bunch of sad faces and crying face emojis saying, “Well, we tried. We know there’s no way that this would be acceptable to your customers.” We asked, “So, where’s this walnut supposed to be coming from that you can’t get it?” They’re like, “It’s been sitting on the docks at the origin since March. It’s being forested in Kentucky in the United States.”

Corey: The thing that surprised me the most about the original model on Kickstarter campaign was how much went wrong across the board. I kept reading your updates. It was interesting, at some point, it was like, okay, this is clearly a Ponzi scheme. That’s the name of the keyboard: ‘The Ponzi’, where there’s going to be increasingly outlandish excuses.

Jesse: I don’t think a Ponzi scheme would be the right aspersion to be casting.

Corey: There’s that more pedestrian scam-style thing. We could go with that.

Jesse: We have a lot of friends who’ve been in industry longer than us, and every time we brought one of the problems that our factory
seemed to be having to them, they said, “Oh, yeah, that’s the thing that absolutely happens.”

Corey: Yeah, it was just you kept hitting every single one of these, and I was increasingly angry on your behalf, reading these things about, “Oh, yeah. Just one of your factory reps just blatantly ripped you off, and this was expected to be normal in some cases, and it’s like”—and you didn’t even once threatened to burn the factory now, which I thought was impressive.

Jesse: No, nobody threatened to burn the factory down, but one of the factories did have a fire.

Corey: Which we can neither confirm nor deny—I kid, I kid, I kid.

Jesse: Yeah, yeah, yeah. But so what our friends who had been in industry longer that said, it was like, “Jesse, but, you know, nobody has all
the problems.” And eventually, we figured out what was going on, and it was that our factory’s director of overseas sales was a con artist grifter who had been scamming both sides. She’d been lying to us and lying to the factory, and making up stories to make her the only trusted person to each side, and she’d just been embezzling huge sums of money.

Corey: You hear these stories, but you never think it’s going to be something that happens to you. Was this your first outing with manufacturing a physical product?

Jesse: This was our first physical product.

Corey: But I’m curious about it; are you effectively following the trope of a software person who thinks, “Ah, I could do hardware? How hard
could it be? I could ship code around the world seconds, so hardware will be just a little bit slower.” How close to that trope are you?

Jesse: So, when we went into the manufacturing side, we knew that we knew nothing, and we knew that it was fraught with peril. And we gave ourselves an awful lot of padding on timing, which we then blew through for all sorts of reasons. And we ran through a hardware incubator that helped us vet our plans, we were working with companies on the ground that helped startups work with factories. And honestly, if it hadn’t been for this one individual, yes we would have had problems, but it wouldn’t have been anything of the same scale. As far as we can tell, almost everything bad that happened had a grain of truth in it, it’s just that… you know, a competent grifter can spin a tiny thing into a giant thing.

And nobody in China suspected her, and nobody in China believed that this could possibly be happening because the penalties if she got caught were ten years in a Chinese prison for an amount of money that effectively would be a down payment on an apartment instead of the price of a full apartment or fully fleeing the country.

Corey: It seems like that would be enough of a deterrent, but apparently not.

Jesse: Apparently not. So, we ended up retaining counsel and talking to friends who had been working in southern China for 15 years for about who they might recommend for a lawyer. We ended up retaining a Chinese lawyer. Her name’s [Una 00:13:36]; she’s fantastic.

Corey: Referrals available upon request.

Jesse: Oh, yeah. No, absolutely. I’m happy to send her all kinds of business. She looked at the contract we had with the factory, she’s like, “This is a Western contract. This isn’t going to help you in the Chinese courts. What we need to do is we need to walk into the factory and negotiate a new agreement that is in Chinese, written by a Chinese lawyer, and get them to sign it.”

And part of that agreement was getting them to take full joint responsibility for everything. And she walked in with me to the factory. She dressed down: t-shirt and jeans. They initially thought she was my translator, and she made a point of saying, “Look, I’m Jesse’s counsel. I’m not your lawyer. I do not represent your interests.”

And three-party negotiations with the factory: the factory’s then former salesperson, and us. And she negotiated a new agreement. And I had a long list of all the things that we needed to have in our contract, like all the things that we really cared about. Get to the end of the day and she hands it to me and she’s like, “What do you think?” And I read it through and my first thought is that none of the ten points that we need in this agreement are there.

And then I realized that they are there, they’re just very subtle. And everybody signs it. The factory takes full joint responsibility for everything that was done by their now former salesperson. We go outside; we get into the cab, and she turns to me—and she’s not a native speaker of English, but she is fluent—and she’s like, how do you think that went, Jesse? I’m like, I think that went pretty well. And she’s like, “Yes. I get my job satisfaction out of adverse negotiation, and the factory effectively didn’t believe in lawyers.”

Corey: No, no. I’ve seen them. They exist. I married one of them.

Jesse: Oh, yeah. As it turned out, they also didn’t really believe in the court system and they didn’t believe in not pissing off judges. Nothing could help us recover the time we lost; we did end up recovering all of our tooling, we ended up recovering all of our product that they were holding, all with the assistance of the Chinese courts. It was astonishing because we went into this whole thing knowing that there was no chance that a Chinese court would find for a small Western startup with no business presence in China against a local factory, and I think our goal was that they would get a black mark on their corporate social credit report so that nobody else would do business with this factory that won’t give the customer back their tooling. And… it turns out that, no, the courts just helped us.

Corey: It’s nice when things work the way they’re supposed to, on some level.

Jesse: It is.

Corey: And then you solve your production problems, you shipped it out. I use it, I take it out periodically.

Jesse: We’d shipped every customer order well before this.

Corey: Oh, okay. This was after you had already done the initial pre-orders. This was as you were ongoing—

Jesse: Yeah, there were keycaps we owed people, which were—

Corey: Oh, okay.

Jesse: Effectively the free gift we promised aways in for being late on shipping.

Corey: That’s what that was for. It showed up one day and I wondered what the story behind that was. But yeah, it was—

Jesse: Yeah.

Corey: They’re great.

Jesse: Yeah. You know, and then there was a story in The Verge of, this Kickstarter alleges that—da, da, da, da, da. We’re like, “I understand that AOL’s lawyers make you say ‘alleges,’ but no, this really happened, and also, we really had shipped everything that we owed to customers long before all this went down.”

Corey: Yeah. This is something doesn’t happen in the software world, generally speaking. I don’t have to operate under the even remote possibility that my CI/CD system is lying to me about what it’s doing. I can generally believe things that show up in computers—you would think—but there are—

Jesse: You would think. I mean—

Corey: There a lot of [unintelligible 00:17:19] exceptions to that, but generally, you can believe it.

Jesse: In software, you sometimes we’ll work with contractors or contract agencies who will make commitments and then not follow through on those commitments, or not deliver the thing they promised. It does sometimes happen.

Corey: Indeed.

Jesse: Yeah, no, the thing I miss the most from software is that if there is a defect, the cost of shipping an update is nil and the speed at which you can ship an update is instantly.

Corey: You would think it would be nil, but then we look at AWS data transfer pricing and there’s a giant screaming caveat on that. It’s you think that moving bytes would cost nothing. Yeah.

Jesse: [unintelligible 00:17:53] compared to international shipping costs for physical goods, AWS transfer rates are incredibly competitive.

Corey: No, no, to get to that stage, you need to add an [unintelligible 00:18:02] NAT gateway with their data processing fee.

Jesse: [laugh].

Corey: But yeah, it’s a different universe. It’s a different problem, a different scale of speed, a different type of customer, too, on some levels. So, after you’ve gotten the Model 01’s issues sorted out, you launched a second keyboard. The ‘a-TREE-us’, if I’m pronouncing that correctly. Or ‘A-tree-us’.

Jesse: So Phil, who designed it, pronounces is ‘A-tree-us’, so we pronounce it A-tree-us. And so, this is a super minimalist keyboard designed to take with you everywhere, and it was something where Phil Hagelberg, who is a software developer of some repute for a bunch of things, he had designed this sort of initially for his own use and then had started selling kits. So, laser-cut plywood enclosures, hand-built circuit boards, you just stick a little development board in the middle of it, spend some time soldering, and you’re good to go. And he and I were internet buddies; he had apparently gotten his start from some of my early blog posts. And one day, he sent me a note asking if I would review his updated circuit board design because he was doing a revision.

I looked at his updated circuit board design and then offered to just make him a new circuit board design because it was going to be pretty straightforward to do something that’s going to be a little more reliable and a lot more cost-effective. We did that and we talked a little more, and I said, “Would you be interested in having us just make this thing in a factory and sell it with a warranty and send you a royalty?” And he said, but it’s GPL. You don’t have to send me a royalty.

Corey: I appreciate that I am not compelled to do it. However—yeah.

Jesse: Yeah, exactly. It’s like, “No. We would like to support people who create things and work with you on it.”

Corey: That’s important. We periodically have guest authors writing blog posts on Last Week in AWS. Every single one of them is paid for what
they do, sometimes there for various reasons that they can’t or won’t accept it and we donate it to a charity of their choice, but we do not expect people to volunteer for a profit-bearing entity, in some respects.

Jesse: Yeah.

Corey: Now, open-source is a whole separate universe that I still maintain that is rapidly becoming a, “Would you like to volunteer for a trillion-dollar company in your weekend hours?” Usually not, but there’s always an argument.

Jesse: Oh, yeah. We have a bunch of open-source contributors to our open-source firmware and we contribute stuff back upstream to other projects, and it is a related but slightly different thing. So, Phil said yes; we said yes. And then we designed and made this thing. We launched an ultra-portable keyboard designed to take with you everywhere.

It came with a travel case that had a belt loop, and basically a spring-loaded holster for your keyboard if you want to nerd out like that. All of the Kickstarter video and all the photography sort of showed how nice it looked in a cafe. And we launched it, like, the week the first lockdowns hit, in the spring of 2019.

Corey: I have to say I skipped that one entirely. One of the things that I wound up doing—keyboard-wise—when I started this company four years ago and change, now was, I wound up getting a fairly large desk, and it’s 72 inches or something like that. And I want a big keyboard with a numpad—yeah, that’s right, big spender here—because I don’t need a tiny little keyboard. I find that the layer-shifting on anything that’s below a full-size keyboard is a little on the irritating side. And this goes beyond. It is—it requires significant—

Jesse: Oh, yeah. It’s—

Corey: Rewiring of your brain, on some level.

Jesse: And there are ergonomic reasons why some people find it to be better and more comfortable. There’s less reaching and twisting. But it is a very different typing experience and it’s absolutely not for everybody. Nothing we’ve made so far is intended to be a mass-market product. When we launched the Model 01, we were nervous that we would make something that was too popular because we knew that if we had to fulfill 50,000 of them, we’d just be screwed. We knew how little we knew.

But the Atreus, when we launched it on Kickstarter, we didn’t know if we were going to have to cancel the campaign because no one was going to want their travel keyboard at the beginning of a pandemic, but it did real well. I don’t remember the exact timing and numbers, but we hit the campaign goal, I want to say early on the first day, possibly within minutes, possibly within hours—it’s been a while now; I don’t remember exactly—ultimately, we sold, like, 2600 of them on Kickstarter and have done additional production runs. We have a distributor in Japan, and a distributor in the US, and a distributor in the UK, now. And we also sell them ourselves directly online, from keyboard.io.

So, this is one of the other fascinating logistics things, is that we ship globally through Hong Kong. Which, before the pandemic was actually pretty pleasant. Inexpensive shipping globally has gotten kind of nuts because most discount carriers, the way they operated historically is, they would buy cargo space on commercial flights. Commercial international flights don’t happen so much.

Corey: Yes, suddenly, that becomes a harder thing to find.

Jesse: Early on, we had a couple of shipping providers that were in the super-slow, maybe up to two weeks to get your thing somewhere by air taking, I want to say we had things that didn’t get there for three months. They would get from Hong Kong to Singapore in three days; they would enter a warehouse, and then we had to start asking questions about, “Hey, it’s been eight weeks. What’s going on?” And they’re like, “Oh, it’s still in queue for a flight to Europe. There just aren’t any.”

Corey: It seems like that becomes a hard problem.

Jesse: It becomes a hard problem. It started to get a little better, and now it’s starting to get a little worse again. Carriers that used to be ultra-reliable are now sketchy. We have FedEx losing packages, which is just nuts. USPS shipments, we see things that are transiting from Hong
Kong, landing at O’Hare, going through a sorting center in Chicago, and just vanishing for weeks at a time, in Chicago.

Corey: I don’t pretend to understand how this stuff works. It’s magic to me; like, it is magic, on some level, that I can order toilet paper on the internet, it gets delivered to my house for less money than it costs me to go to the store and buy it. It feels like there’s some serious negative externalities in there. But we don’t want to look too closely at those because we might feel bad about things.

Jesse: There’s all kinds of fascinating stuff for us. So, shipping stuff, especially by air, there are two different ways that the shipping weight can get calculated. It can either get calculated based on the weight on a scale, or it can get calculated using a formula based on the dimensions. And so bulky things are treated as weighing an awful lot. I’m told that Amazon’s logistics teams started doing this fascinating thing where ultra-dense, super-heavy shipments they pushed on to FedEx and UPS, whereas the ultra-light stuff that saved on jet fuel, they shoved onto their own planes.

Corey: This episode is sponsored by our friends at Oracle Cloud. Counting the pennies, but still dreaming of deploying apps instead of "Hello,
World" demos? Allow me to introduce you to Oracle's Always Free tier. It provides over 20 free services and infrastructure, networking databases, observability, management, and security.

And - let me be clear here - it's actually free. There's no surprise billing until you intentionally and proactively upgrade your account. This means you can provision a virtual machine instance or spin up an autonomous database that manages itself all while gaining the networking load, balancing and storage resources that somehow never quite make it into most free tiers needed to support the application that you want to build.

With Always Free you can do things like run small scale applications, or do proof of concept testing without spending a dime. You know that I always like to put asterisks next to the word free. This is actually free. No asterisk. Start now. Visit https://snark.cloud/oci-free that's
https://snark.cloud/oci-free.

Corey: I want to follow up because it seems like, okay, pandemic shipping is a challenge; you clearly are doing well. You still have them in stock and are selling them as best I’m aware, correct?

Jesse: Yes.

Corey: Yeah. I may have to pick one up one of these days just so I can put it on the curiosity keyboard shelf and kick it around and see how it works. And then you recently concluded a third keyboard Kickstarter, in this case. And—

Jesse: Yeah.

Corey: —this is not your positioning; this is my positioning of what I’m picking up of, “Hey, remember that Model 01 keyboard we sold you that you love and we talked about and it’s amazing? Yeah, turns out that’s crap. Here’s the better version of it.” Correct that misapprehension, please. [laugh].

Jesse: Sure. So, it absolutely is not crap, but we’ve been out of stock in the Model 01 for a couple of years now. And we see them going used for as much or sometimes more than we used to charge for them new. It went out of stock because of the shenanigans with that first factory. And shortly before we launched the Atreus, we’d been planning to bring back an updated version of the Model 01; we’ve even gotten to the point of, like, designing the circuit boards and starting to update the tooling, the injection molding tooling, and then COVID, Atreus, life, everything.

And so it took us a little longer to get there. But there is a larger total addressable market for a keyboard like the Model 01 than the total number that we ever sold. There are certainly people who had Model 01s who want replacements, want extras, want another one on another desk. There are also plenty of people who wanted a Model 01 and never got one.

Corey: Here’s my question for you, with all three of these keyboards because they’re a different layout, let’s be clear. Some more so than others, but even the columnar layout is strange here. Once upon a time, I had a week in which I wasn’t doing much, and I figured, ah, I’ll Dvorak—which is a different keyboard layout—and it’s not that it’s hard; it’s that it’s rewiring a whole bunch of muscle memory. The problem I ran into was not that it was impossible to do, by any stretch, but because of what I was doing—in those days help desk and IT support—I was having to do things on other people’s computers, so it was a constant context switching back and forth between different layouts.

Jesse: Yeah.

Corey: Do you see that being a challenge with layouts like this, or is it more natural than that?

Jesse: So, what we found is that it is easier to switch between an ergonomic layout and a traditional layout, like a columnar layout, and what’s often called a row-stagger layout—which is what your normal keyboard looks like—than it is to switch between Dvorak and Qwerty on a traditional keyboard. Or the absolute bane of my existence is switching between a ThinkPad and a MacBook. They are super close; they are
not the same.

Corey: Right. You can’t get an ergonomic keyboard layout inside of a laptop. I mean, looking at the four years of being gaslit by Apple, it’s clear you can barely get a keyboard into a MacBook for a while. It’s, “Oh, it’s a piece of crap, but you’re using it wro”—yeah. I’m not a fan of their entire approach to keyboards and care very than what Apple has to say about anything even slightly keyboard-related, but that’s just me being bitter.

Jesse: As far as I can tell, large chunks of Apple’s engineering organization felt the same way that you did. Their new ones are actually decent again.

Corey: Yes, that’s what I’ve heard. And I will get one at some point, but I also have a problem where, “Oh, yeah, you know that $3,000 laptop with a crappy keyboard, you can’t use for anything? Great. The solution is to give us 3000 more dollars, and then we’ll sell you one that’s good.” And it’s, I feel like I don’t want to reward the behavior.

Jesse: I hear you. I ditched Mac OS for a number of years. I live the dream: Linux on the desktop. And it didn’t hurt me a lot—printing worked fine, scanning worked fine, projectors were fine—but when I was reaching for things like Photoshop, and Lightroom, and my mechanical CAD software, it was the bad kind of funny.

Corey: I have to be careful, now for the first time in my life I’m not updating to new operating systems early on, just because of things like the audio stuff I have plugged into my nonsense and the media nonsense that I do. It used to be that great, my computer only really needs to be a web browser and a terminal and I’m good. And worst case, I can make do with just the web browser because there are embedded a terminal into a web page options out there. Yeah, now it turns out that actually have a production workflow. Who knew?

Jesse: Yep. That’s the point where I started thinking about having separate machines for different things. [laugh].

Corey: Yeah, I’m rapidly hitting that point. Yeah, I do want to get into having fun with keyboards, on some level, but it’s the constant changing of what you’re using. And then, of course, there’s the other side of it where, in normal years, I spent an awful lot of time traveling and as much fun as having a holster-mounted belt keyboard would be, in many cases, it does not align with the meetings that I tend to be in.

Jesse: Of course.

Corey: It’s, “Oh, great. You’re the CFO of a Fortune 500. Great, let me pair my mini keyboard that looks like something from the bowels of your engineering department’s reject pile.” Like, what is this? It’s one of those things that doesn’t send the right message in some cases. And let’s be honest; I’m good at losing things.

Jesse: This is a pretty mini keyboard, but I hear you.

Corey: Or I could lose it, along with my keys. It will be great.

Jesse: Yeah. There are a bunch of things I’ve wanted to do around reasonable keyboards for tablets.

Corey: Yes, please do.

Jesse: Yeah. We actually started looking at one point at a fruit company in Cupertino’s requirements around being able to do dock-connector
connected keyboards for their tablets, and… it’s nuts. You can’t actually do ergonomic keyboards that way, it would have to be Bluetooth.

Corey: Yeah. When I travel on the road these days, or at least—well, ‘these days’ being two years ago—the only computer I’d take is an iPad. And that was great; it works super well for a lot of my use cases. There’s still something there, and even going forward, I’m going to be spending a lot more time at home. I have young kids now, and I want to be here to watch them grow up.

And my lifestyle and use cases have changed for the last year and a half. I’ve had an iMac. I’ve never had one of those before. It’s big screen real estate; things are great. And I’m looking to see whether it’s time to make a full-on keyboard evolution if I can just force myself over the learning curve, here. But here’s the question you might not be prepared to answer yet. What’s next? Do you have plans on the backburner for additional keyboards beyond what you’ve done?

Jesse: Oh, yeah. We have, like, three more designs that are effectively in the can. Not quite ready for production, but if this were a video podcast, I’d be pulling out and waving circuit boards at you. One of the things that we’ve been playing with is what is called in the trade a symmetric staggered keyboard where the right half is absolutely bog-standard normal layout like you’d expect, and the left side is a mirror of that. And so it is a much more gentle introduction to an ergonomic-style keyboard.

Corey: Okay, I can almost wrap my head around that.

Jesse: Because if you put your hands on your keyboard and you feel the angles that you have to move on your right side, you’ll see that your fingers move basically straight back and forth. On the left side, it’s very different unless you’re holding your hand at a crazy, crazy angle.

Corey: Yeah.

Jesse: And so it’s basically giving you that same comfort on the right side and also making the left side comfy. It’s not a weird butterfly-shaped keyboard; it is still a rectangle, but it is just that little bit better. We’re not the first people who have done this. Our first prototype of this thing was, like, 2006, something like that. But it was a one-off, like, “I wonder if I would like this.” And we were actually planning to do that one next after the Model 01 when the Atreus popped up, and that was a much faster, simpler, straighter-forward thing to bring to production.

Corey: The one thing I want from a keyboard—and I haven’t found one yet; maybe it exists, maybe I have to build it myself—but I want to do the standard mechanical keyboard—I don’t even particularly care about the layout because it all passes through a microcontroller on the device itself. Great. And those things are programmable as you’ve demonstrated; you’ve already done an awful lot of open-source work that winds up being easily used to control keyboards. And I love it, and it’s great, but I also want to embed a speaker—a small one—into the keyboard so I can configure it that every time I press a key, it doesn’t just make a clack, it also makes a noise. And I want to be able to—ideally—have it be different keys make different noises sometimes. And the reason being is that when we eventually go back to offices, I don’t want there to be any question about who is the most obnoxious typist in the office; I will—

Jesse: [laugh].

Corey: —win that competition. That is what I want from a keyboard. It’s called the I-Don’t-Want-Anyone-Within-Fifty-Feet-Of-Me keyboard. And I don’t quite know how to go about building that yet, but I have some ideas.

Jesse: So, there’s absolutely stuff out there. There is prior art out there.

Corey: Oh, wonderful.

Jesse: One of the other options for you is solenoids.

Corey: Oh, those are fun.

Jesse: So, a solenoid is—there is a steel bar, an electromagnet, and a tube of magnetic material so that you can go kachunk every time you
press a key.

Corey: It feels functionally like a typewriter to my understanding.

Jesse: I mean, it can make it feel like a typewriter. The haptic engine in an iPhone or a Magic Trackpad is not exactly a solenoid but might give
you the vaguest idea of what you’re talking about.

Corey: Yeah, I don’t think I’m going to be able to quite afford 104 iPhones to salvage all of their haptic engines so that I can then wind up hooking each one up to a different key but, you know, I am sure someone enterprising come up with it.

Jesse: Yeah. So, you only need a couple of solenoids and you trigger them slightly differently depending on which key is getting hit, and you’ll get your kachunk-kachunk-kachunk-kachunk-kachunk.

Corey: Yeah, like spacebar for example. Great. Or you can always play a game with it, too, like, the mystery key: whenever someone types in the hits the mystery key, the thing shrieks its head off and scares the heck out of them. Especially if you set it to keys that aren’t commonly used, but ever so frequently, make everyone in the office jumpy and nervous.

Jesse: This will be perfect for Zoom.

Corey: Oh, absolutely, it would. In fact, one thing I want to do soon if this pandemic continues much longer, is then to upgrade my audio setup here so I can have a second microphone pointed directly into my keyboard so that people who are listening at a meeting with me can hear me typing as we go. I might be a terrible colleague. One wonders.

Jesse: You might be a terrible colleague, but you might be a wonderful colleague. Who knows?

Corey: It all depends on the interests we have. I want to thank you for taking the time to walk me through the evolution of Keyboardio. If people want to learn more, or even perhaps buy one of these things, where can they do that?

Jesse: They can do that at keyboard.io.

Corey: And hence the name. Thank you so much for taking the time to speak with me about all this. I really appreciate it.

Jesse: Cool. Thanks so much for having me. I had fun.

Corey: I did, too. Jesse Vincent—obra on Twitter, and of course, the CTO of Keyboardio. I am Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry comment, but before typing it, switch your keyboard to Dvorak.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Nipun

Nipun Agarwal is Vice President, MySQL HeatWave and Advanced Development, Oracle. His interests include distributed data processing, machine learning, cloud technologies and security. Nipun was part of the Oracle Database team where he introduced a number of new features. He has been awarded over 170 patents.

Links:

  • HeatWave: https://oracle.com/heatwave

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: You could build you go ahead and build your own coding and mapping notification system, but it takes time, and it sucks! Alternately, consider Courier, who is sponsoring this episode. They make it easy. You can call a single send API for all of your notifications and channels. You can control the complexity around routing, retries, and deliverability and simplify your notification sequences with automation rules. Visit courier.com today and get started for free. If you wind up talking to them, tell them I sent you and watch them wince—because everyone does when you bring up my name. Thats the glorious part of being me. Once again, you could build your own notification system but why on god’s flat earth would you do that?

Corey: This episode is sponsored in part by our friends at VMware. Let’s be honest—the past year has been far from easy. Due to, well, everything. It caused us to rush cloud migrations and digital transformation, which of course means long hours refactoring your apps, surprises on your cloud bill, misconfigurations and headache for everyone trying manage disparate and fractured cloud environments. VMware has an answer for this. With VMware multi-cloud solutions, organizations have the choice, speed, and control to migrate and optimize

applications seamlessly without recoding, take the fastest path to modern infrastructure, and operate consistently across the data center, the edge, and any cloud. I urge to take a look at vmware.com/go/multicloud. You know my opinions on multi cloud by now, but there's a lot of stuff in here that works on any cloud. But don’t take it from me thats: VMware.com/go/multicloud and my thanks to them again for sponsoring my ridiculous nonsense.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Today’s promoted episode is slightly off the beaten track. Normally in tech, we tend to find folks that have somewhere between an 18 to 36-month average tenure at companies. And that’s great, however, let’s do the exact opposite of that today. My guest is Nipun Agarwal, who’s the VP of MySQL HeatWave and Advanced Development at Oracle, where you’ve been an employee for 27 years, is it?

Nipun: That’s absolutely right. 27 years and that was my first job out of school. So, [laugh] yes.

Corey: First, thank you for joining me. It is always great to talk to people who have focused on an area that I only make fun of from a distance, in this case, databases which, you know, DNS works well enough for most use cases, but occasionally customers have other constraints. You are clearly at or damn near at the top of your field. In my pre-show research, I was able to unearth that you have—what is it now, 170, 180 filed patents that have been issued?

Nipun: That’s right. 180 issued patents. [laugh].

Corey: You clearly know what you’re doing when it comes to databases.

Nipun: Thank you for the opportunity. Yes, thank you.

Corey: So, being a VP at Oracle, but starting off as your first job as almost a mailroom to the executive suite style story, we don’t see those anymore. In most companies, it very much feels like the path to advance is to change jobs to other companies. It’s still interesting seeing that that’s not always the path forward, for some folks. I think that the folks who have been in companies for a long time need more examples and role models to look at in that sense, just because it is such an uncommon narrative these days. You’re not bouncing around between four companies.

Nipun: Yeah. I’ve been lucky enough to have joined Oracle, and although I had been at Oracle, I’ve been on multiple teams at Oracle and there has been a great opportunity of talent, colleagues, and projects, where even to this day, I feel that I have a lot more to learn. And there are opportunities within the company to learn and to grow. So no, I’ve had an awesome ride.

Corey: Let’s dive in a little bit to something that’s been making the rounds recently, specifically you’ve released something called HeatWave, which has been boasting some, frankly, borderline unbelievable performance benchmarks, and of course, everyone loves to take a crack at Oracle for a variety of reasons, so Twitter is very angry. But I’ve learned at some point, through the course of my career, to disambiguate Twitter’s reactions from what’s actually happening out there. So, let’s start at the beginning. What is HeatWave?

Nipun: HeatWave is an in-memory query accelerator for MySQL. It accelerates complex, long-running, analytic queries. The interesting thing about HeatWave is, with HeatWave we now have a single MySQL database which can run all your applications, whether they’re OLTP, whether they’re mixed workloads, or whether they’re analytics, without having to move the data out of MySQL. Because in the past, people would need to move the data from MySQL to some other database running analytics, so people would end up with two different databases. With this single database, no need for moving the data, and all existing tools and applications which worked with MySQL continue to work, except they will be much faster. That’s what HeatWave is.

Corey: The benchmarks that you are publishing are fairly interesting to me, specifically, the ones that I’ve seen are, you’ve classified HeatWave as six-and-a-half times faster than Amazon Redshift, seven times faster than Snowflake, nine times faster than BigQuery, and a number of other things, and fourteen hundred times faster than Amazon Aurora. And what’s interesting to me about the things that you’re naming is they’re not all data-warehouse style stuff. Aurora, for example, is Amazon’s interpretation of an in-house developed managed database service named after a Disney Princess. And it tends to be aimed at things that are not necessarily massive scale. What is the sweet spot, I guess, of HeatWaves data sizes when it comes to really being able to shine?

Nipun: So, there are two aspects where our customers are going to benefit from HeatWave. One characteristics is the data size, but the other characteristics is the complexity of the queries. So, let’s first do the comparison with Aurora—and that’s a very good question—the 1400 times comparison we have shown, yes, if you take the TPC-H queries on a four terabyte workload and if you run them, that’s what you’re going to see. Now, the interesting thing is this: not only is it 1400 times faster it’s also at half the price because for most of these systems, if you throw more gear, if you throw more hardware, the performance would vary. So, it’s very important to go with how much of performance and at what price.

So, for pure analytics—say, for four terabytes—is 1400 times faster at half the price. So, if it provides truly 800 times better price performance compared to Aurora for pure analytics. Now, let’s take the other extreme. 100 gigabytes—which is a much smaller, your bread and butter database—and this is for mixed workloads. So, something like a CH-benCHmark, which has a combination of say, some TPC-C transactions, and then some added IPP-CH queries, which—the CH benCHmark.

Here we have 42 times advantage price performance over Aurora because we are 42% of the cost, less than half the cost of Aurora and for the complex queries, we are about 18 times faster, and for pure OLTP, we are at par. So, the aggregate comes out to be about 42 times better. So, the mileage varies depending upon the data size and depending upon the complexity of the queries. So, in the case of Aurora, it will be anywhere from 42 times better price performance all the way to 2800.

Corey: Does this have an upper bound, for example? Like, if we take a look at something like Redshift or something like Snowflake, where they’re targeting petabyte-scale workloads at some point, that becomes a very different story for a lot of companies out there. Is that something that this can scale to, or is there a general reasonable upper bound of, okay, once you’re above X number of terabytes, it’s probably good to start looking at tiering data out or looking at a different solution?

Nipun: We designed HeatWave primarily for those customers who had to move the data out of MySQL database into some other database for running analytics. The upper bound for the data in the MySQL database is 64 terabytes. Based on the demand and such we are seeing, we support 32 terabytes processing in HeatWave at any given point in time. You can still have 64 terabytes in the MySQL database, but the amount of data you can load into the HeatWave cluster at any given point in time is 32 terabytes.

Corey: Which is completely reasonable. I would agree with you from not having much database exposure myself in the traditional sense, but from a cloud economics standpoint alone, anytime you have to move data to a different database for a different workload, you’re instantly jacking costs through the roof. Even if it’s just the raw data volumes, you now have to store it in two different places instead of one. Plus, in many cases, the vaguearities of data transfer pricing in many places wind up meaning that you’re paying money to move things out, there’s a replication story, there’s a sync factor, and then it just becomes a management overhead problem. If there’s a capacity to start using the data where it is in more intelligent ways, that alone has a massive economic wind, just from a time it takes your team to not have to focus on changing infrastructure and just going ahead to run the queries. If you want to start getting into the weeds of all the different ways something like this is an economic win, there’s a lot of angles to look at it from.

Nipun: That’s an excellent point and I’m very glad you brought it up. So, now let’s take the other set of benchmarks we were talking about: Snowflake. So, HeatWave is seven times faster and one-fifth the cost; it’s about 35 times better price performance. Compared to let’s say Redshift AQUA, six-and-a-half times faster at half the cost, so 13 times better price performance. And it goes on and on.

Now, these numbers I was quoting is for 10 terabytes TPC-H queries. And the point which you said is very, very valid. When we are talking about the cost for these other systems, it’s only the cost for analytics without including the cost of the source database or without including the cost of moving the data or managing to different databases. Whereas when you’re talking about the cost of HeatWave, this is the cost which includes the cost of both transaction processing as well as the analytics. So, it’s a single database; all the cost is included, whereas, for these other vendors, it’s only the cost of the analytic database. So, the actual cost to a user is probably going to be much higher with these other databases. So, the price performance advantage with HeatWave will perhaps be even higher.

Corey: Tell me a little bit about how it works. I mean, it’s easy to sit here and say, “Oh, it’s way faster and it’s better in a bunch of benchmark stuff,” and we will get into that in a little bit, but it’s described primarily as an in-memory query accelerator. Naively, I think, “Oh, it’s just faster because instead of having data that lives on disk, it winds up having some of it live in RAM. Well, that seems simple and straightforward.” Like, oh, yeah, I’m going to go on a limb and assume that there aren’t 160 patents tied to the idea that RAM is faster than disk. There’s clearly a lot more going on. How does this work? What is it foundationally?

Nipun: So, the thing to realize is HeatWave has been built from the ground up for the cloud and it is optimized for the Oracle Cloud. So, let’s take these things one at a time. When I say designed from the ground up for the cloud, we have actually invented and implemented new algorithms for distributed query processing, which is what gives us such a good advantage in terms of operations like joint processing, window functions, aggregations. So, we have come up—invented, implemented new algorithms for distributed query processing. Secondly, we have designed it for the cloud.

And by that what I mean is, A, we have a lot of emphasis on scalability, that it scales to thousands of cores with a very, very good scale factor, which is very important for the cloud. The next angle about the cloud is that not only have we optimized it for the cloud, but we have gone with commodity cloud services, meaning, for instance, when you’re looking at the storage, we are looking at the least expensive price. So, for instance, we use object store; you don’t use, for instance, locally attached SSDs because that will be expensive. Similarly, for compute: instead of using Intel, we use AMD chips because they are less expensive. Similarly, networking: standard networking.

And all of this has been optimized for the specific Oracle Cloud infrastructure shapes we have, for the specific VMs we use, for the specific networking bandwidth we get, for the object store bandwidth and such; so that’s the third piece, optimized for OCI. And the last bit is pervasive use of machine learning in the service. So, a combination of these four things: designed for the cloud, using commodity cloud services, optimized for the quality cloud infrastructure, and finally the pervasive use of machine learning is what gives us very good performance, very good scale, at a very inexpensive price.

Corey: I want to dig into the idea of the pervasive use of machine learning. In many cases, machine learning is the answer to how do I wind up bilking a bunch of VCs out of money? And Oracle is not a venture-backed company at this stage of its existence, it is a very large, publicly-traded entity; you have no need to do that. And I would also further accept that this is one of those bounded problem spaces where something that looks machine-learning-like could do very well. Is that based upon what it observes and learns from data access patterns? Is it something that it learns based from a specific workload in question? What is the gathering, and is it specific to individual workloads that a given customer has, or is it holistically across all of the database workloads that you see in Oracle Cloud?

Nipun: So, there are multiple parts to this question. The first thing is—and I think as you’re noting—that with the cloud, we have a lot more opportunity for automation because we know exactly what is the hardware stack, we know the software stack, we know the configuration parameters.

Corey: Oh yes, hell is other people’s data centers, for sure.

Nipun: [laugh]. And the approach we have taken for automation is machine-learning-based automation because one of the big advantages is that we can have a model which is tailored to a specific instance and as you run more queries, as you run more workloads, the system gets more intelligent. And we can talk about that maybe later about, like, specific things which make it very, very compelling. The third thing, I think, which you were alluding to, is that there are two aspects in machine learning: data, and the models or the algorithms. So, the first thing is, we have made a lot of enhancements, both to the MySQL engine as well as HeatWave, to collect new kinds of data.

And by new kinds of data, I mean, that not only do we collect statistics of data, but we collect statistics of, say, the queries: what was the compilation time? What was the execution time? And then, based on this data which we’re collecting, we have then come up with very advanced algorithms—machine learning algorithms—which are, again, a lot of them, there is, like, you know, patterns or [IP 00:14:13] which we have built on top of the existing state of art. So, for instance, taking these statistics and extrapolating them on larger data sizes. That’s completely an innovation which we did in-house.

How do we sample a very small percentage of the data and still be accurate? And finally, how do we come up with these machine learning models which are accurate without hiring an army of engineers? That’s because we invented our AutoML, which is very efficient. So, that’s basically the ecosystem of the machine learning which we have, which has been used to provide this.

Corey: It’s easy for folks to sit there and have a bunch of problems with Oracle for a variety of reasons, some of which are no longer germane, some of which are, I’m not here to judge. But I think it’s undeniable—though it sometimes gets eclipsed by people’s knee-jerk reactions—the reason that Oracle is in so many companies that it is in is because it works. You folks have been pioneers in the database space for a very long time and that’s undeniable. If it didn’t deliver performance that was untouchable for a long time, it would not have gotten to the point where you now are, where it is the database of record for an awful lot of shops. And I know it’s somehow trendy, sometimes, for the startup set to think, “Oh, big companies are slow and awful. All innovation comes out of small, scrappy startups here.”

But your customers are not fools. They made intelligent decisions based upon constraints that they’re working within and problems that they need to solve. And you still have an awful lot of customers that are not getting off of Oracle anytime soon because it works. It’s one of those things that I think is nuanced and often missed. But I do feel the need to ask about the lock-in story. Today, HeatWave is available only on the managed MySQL service in Oracle Cloud, correct?

Nipun: Correct.

Corey: Is there any licensing story tied to that? In other words, “Well, if I’m going to be using this, I need to wind up making a multi-year commitment. I need to get certain support things, as well,” the traditional on-premises Oracle story. Or is this an actual cloud service, in that you pay for what you use while you use it, and when you turn it off, you’re done? In theory. In practice, we know in cloud economics, no one ever turns anything off until the company goes out of business.

Nipun: So, it’s exactly the letter what you said that this is a managed service. It’s pay as you go, you pay only for what you consume, and if you decide to move on, there’s absolutely no license or anything that is holding you back. The second thing—and I’m glad you brought it up—about the vendor lock-in. One of the very important things to realize about HeatWave is, A, it’s just an accelerator for MySQL, but in the process of doing so, we have not introduced any proprietary syntax. So, if customers have the MySQL application running on some other cloud, they can very easily migrate to OCI and try MySQL HeatWave.

But for whatever reason, if they don’t like it, and they want to move out, there is absolutely nothing which is holding them back. So, the ease of which they can come in with the same ease they can walk out because we don’t have any vendor lock-in. There is absolutely no proprietary extensions to HeatWave.

Corey: There is the counter-argument as far as lock-in goes, and we see this sometimes with companies we talk to that were considering Google Cloud Spanner, as an example. It’s great, and you can use it in a whole bunch of different places and effectively get ACID-compliance-like behavior across multiple regions, and you don’t have to change any of the syntax of what it is you’re using except the lock-in there is one of a strategic or software architecture lock-in because there’s nothing else quite like that in the universe, which means that if you’re going to migrate off of the single cloud where that’s involved, you have to re-architect a lot, and that leads to a story of lock-in. I’m curious as to whether you’re finding that customers are considering that as far as the performance that you’re giving for MySQL querying is apparently unparalleled in the rest of the industry; that leads to a sort of lock-in itself when people get used to that kind of responsiveness and build applications that expect that kind of tolerances. At some point, if there’s nothing else in the industry like it, does that means that they find themselves de-facto locked in?

Nipun: If you were to talk about some functionality which we are offering which no one else is offering, perhaps you could, kind of, make that case. But that’s not the case for performance because when we are so much faster—so suppose I said, okay, we are so much faster; we are six-and-a-half times faster than Redshift at half the cost. Well, if someone wanted the same performance, they can absolutely do it Redshift on a much larger cluster, and pay a lot more. So, if they want the best performance at the best price, they can come to Oracle Cloud; if they want the same performance but they will have to pay more, they can go anywhere else. So, I don’t think that’s a vendor lock-in at all.

That’s a value which we are bringing in that for the same performance, we are much cheaper. Or you can have that kind of a balance that we are faster and cheaper. So, there is no lock-in. So, it’s not to say that, okay, we have made some extensions to MySQL which are only available in our cloud. That is not at all the case.

Now, for some other vendors and for some other applications—you brought up Spanner; that’s one. But we have had multiple customers of MySQL who, when they were trying Google BigQuery, they mentioned this aspect that, okay, Google BigQuery had these proprietary extensions and they feel locked in. That is not the case at all with HeatWave.

Corey: This episode is sponsored by our friends at Oracle HeatWave is a new high-performance accelerator for the Oracle MySQL Database Service. Although I insist on calling it “my squirrel.” While MySQL has long been the worlds most popular open source database, shifting from transacting to analytics required way too much overhead and, ya know, work. With HeatWave you can run your OLTP and OLAP, don’t ask me to ever say those acronyms again, workloads directly from your MySQL database and eliminate the time consuming data movement and integration work, while also performing 1100X faster than Amazon Aurora, and 2.5X faster than Amazon Redshift, at a third of the cost. My thanks again to Oracle Cloud for sponsoring this ridiculous nonsense.

Corey: I do want to call out, just because it seems like there’s a lies, damned lies, and database benchmarks story here where, for example, Azure for a while was doing a campaign where they were five times less expensive for database workloads than AWS until you scratched beneath the surface and realize it’s because they’re playing ridiculous games with licensing, making it very expensive to run a Microsoft SQL Server on anything that wasn’t Azure. Customers are not necessarily as credulous as they once were when it comes to benchmarking. And Oracle for a long time hasn’t really done benchmarking, and in fact, has actively discouraged it. For HeatWave, you’ve not only published benchmarks, which okay, vendors can say anything they want, and I’m going to wait until I see independent returns, but you put not just the benchmarks, but data sets, and your entire methodology onto GitHub as well. What led to that change? That seems like the least Oracle-like
thing I could possibly imagine.

Nipun: I couldn’t take credit for the idea. The idea actually was from our Chief Marketing Officer, that was really his idea. But here is the reason
why it makes a lot more sense for us to do it for MySQL HeatWave. MySQL is pervasive; pretty much any cloud vendor you can think about has a MySQL-based managed service. And obviously, MySQL runs on premise, like a lot of customers and applications do it.

Corey: That’s one of the baseline building blocks of any environment. I don’t even need to be in the cloud; I can get MySQL working somewhere. Everyone has it, and if not, why don’t you? And I can build it in a VM myself in 20 minutes.

Nipun: That’s right.

Corey: It is a de-facto standard.

Nipun: That’s right. So, given that is the case and many other cloud vendors are innovating on top of it—which is great—how do you compare the innovation or the value proposition of Cloud Vendor A with us? So, for that, what we felt was that it is very important and very fair that we publish our scripts so that people can run those same scripts with a HeatWave, as well as with other cloud offerings, and make a determination for themselves. So, given the popularity of MySQL and given that pretty much all cloud vendors provide an offering of MySQL, and many of them have enhanced it, in order for customers to have an apples-to-apples comparison, it is imperative that we do this.

Corey: I haven’t run benchmarks myself just yet, just because it turns out, there’s a lot of demands on my time and also, as mentioned, I’m not a deep database expert, unless it comes to DNS. And we keep waiting for people to come back with, “Aha. Here’s why you’re completely comprised of liars.” And I haven’t heard any of that. I’ve heard edges and things here about, “Well, if you add an index over here, it might speed things up a bit,” but nothing that leads me to believe that it is just a marketing story.

It is a great marketing story, but things like this fall apart super quickly in the event that it doesn’t stand up to engineering scrutiny. And it’s been out long enough that I would have fully expected to have heard about it. Lord knows if anyone is listening and has thoughts on this, I will be getting some letters after this episode, I expect. But I’ve come to expect those; please feel free to reach out. I’m always thrilled to do follow-up episodes and address things like this.

When does it make sense from your perspective for someone to choose HeatWave on top of the Oracle Cloud MySQL service instead of using some of the other things we’ve talked about: Aurora, Redshift, Snowflake, et cetera? When does that become something that a customer should actively consider? Is it for net-new workloads? Should they consider it for migration stories? Should they run their database workloads in Oracle Cloud and keep other stuff elsewhere? What is the adoption path that you see that tends to lead to success?

Nipun: All customers of MySQL, or all customers of any open-source database, those would be absolutely people who should consider MySQL HeatWave. For the very simple reason: first, regardless of the workload, whether it is OLTP only, or mixed workloads, or analytics, the cost is going to be significantly lower. I’ll say at least it’s going to be half the cost. In most of the cases, it’s probably going to be less than half the cost. So, right off the bat, customers save half the cost by moving to MySQL HeatWave.

And then depending upon the workload you have, as you have more complex queries, the performance advantage starts increasing. So, if you were just running only OLTP, if you only had transactions and you didn’t have any complex queries—which is very unlikely for real-world applications, but even if that was the case, you’re going to save 60% by going to MySQL HeatWave. But as you have more complex queries you will start finding that the net advantage you’re going to get with performance is going to keep increasing and will go anywhere from 10 times aggregate to as much as 1400 times. So, all open-source, MySQL-based applications, they should consider moving. Then you mentioned about Snowflake, Redshift, and such; for all of them, it depends on what the source database is and what is it that they’re trying to
do.

If they are moving data from, say, some open-source databases, if they are ETL-ing from MySQL, not only will MySQL HeatWave be much faster and much cheaper, but there’s going to be a tremendous value proposition to the application because they don’t need to have two different applications for two different databases. They can come back to MySQL, they can have a single database on which they can run all their applications. And then again, you have many of these cloud-native applications are born in the cloud where people may be looking for a simple database which does the job, and this is a great story—both in terms of cost as well as in terms of performance—and it’s a single database for all your applications, significantly reduces the complexity for users.

Corey: To turn the question around a little bit, what sort of workloads is MySQL HeatWave not a fit for? What sort of workloads are going to lead to a poor customer experience? Where, yeah, this is not a fit for that workload?

Nipun: None, except in terms of the data size. So, if you have data sizes which are more than 64 terabytes, then yes, MySQL HeatWave is not a good fit. But if your data size is under 64 terabytes, you’re going to win in all the cases by moving to MySQL HeatWave, given the
functionality and capabilities of MySQL.

Corey: I’d also like to point out that recently, HeatWave gained the MySQL Autopilot capability, which I believe is a lot of the machine learning technologies that you were speaking about a few minutes ago. Are there plans to continue to expand what HeatWave does and offer additional functionality? And—if you can talk about any of that. I know that roadmap is always something that is difficult to ask about, but it’s clear that you’re investing in this. Is your area of investment looking more like it’s adding additional features? Is it continuing to improve existing performance? Something else entirely? And of course, we also accept you can’t tell me any of [laugh] that has a valid answer.

Nipun: Well, we just got started, so we just had our first [GF 00:27:03] HeatWave in December, and you saw that earlier this week we had our second major release of HeatWave. We are just getting started, so absolutely we are investing a lot in this area. But we are pretty much going to attempt all the things that you said. We have feedback from existing customers which is very high up on the priority list. And some of these are just one, say, class of enhancements which [unintelligible 00:27:25], can HeatWave handle larger sizes of data? Absolutely, we have done that; we will continue doing that.

Second is, can HeatWave accelerate more constructs or more queries? Absolutely, we will do that. And then you have other kinds of capabilities which customers are asking which you can think of are, like you know, bigger features, which for instance, we announced the support for scale-out data storage which improves recovery time. Well, you’re going to improve the recovery time or you’re going to improve the time it takes to restart the database. And when I say improve, we are talking about not an improvement of 2X or 3X, but it’s 100 times improvement for, let’s say, a 10 terabyte data size.

And then we have a very good roadmap which, I mean, it’s a little far out that I can’t say too much about it, but we will be adding a lot of very good new capabilities which will differentiate HeatWave even more, compared to the competitive services.

Corey: You have very clearly forgotten more about databases than most of us are ever going to know. As you’ve been talking to folks about HeatWave, what do you find is the most common misunderstanding that folks like me tend to come away with when we’re discussing the technology? What is it that is, I guess, a nuance that is often being missed in the industry’s perspective as they evaluate the new technology?

Nipun: One aspect is that many times, people just think about a service to be here some open-source code or some on-premise code which is being hosted as a managed service. Sure, there’s a lot of value to having a managed service, don’t get me wrong, but when you have innovations, particularly when you have spent years in years or decades of innovation for something which is optimized for the cloud, you have an architectural advantage which is going to pay dividends to customers for years and years to come. So, there is no substitute for that; if you have designed something for the cloud, it is going to do much better whether it’s in terms of performance, whether it’s in terms of scalability, whether it’s in terms of cost. So, that’s what people have to realize that it takes time, it takes investment, but when we start getting the payoff, it’s going to be fairly big. And people have to think that okay, how many technologies or services are out there which have made this kind of investment?

So, what I’m really excited about is, MySQL is the most popular database amongst developers in the world; we spend a lot of time, a lot of person-years investing over the last, you know, decade, and now we are starting to see the dividends. And from what we have seen so far, the response has been terrific. I mean, it’s been really, really good response, and we are very excited about it.

Corey: I want to thank you for taking so much time to speak with me today. If people want to learn more, where can they go?

Nipun: Thank you very much for the opportunity. If they would like to know more, they can go to oracle.com/heatwavewhere we have a lot of details, including a technical brief, including all the details of the performance numbers we talked about, including a link to the GitHub where they can download the scripts. And we encourage them to download the scripts, see that they’re able to reproduce the results we said, and then try their workloads. And they can find information as to how they can get free credits to try the service for free on their own and make up their mind themselves.

Corey: [laugh]. Kicking the tires on something is a good way to form an opinion about it, very often. Thank you so much for being so generous with your time. I appreciate it.

Nipun: Thank you.

Corey: Nipun Agarwal, Vice President of MySQL HeatWave and Advanced Development at Oracle. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an insulting comment formatted as a valid SQL query.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Adam

Adam Zimman is a start-up Advisor providing guidance on leadership, platform architecture, product marketing, and GTM strategy. He has over 20 years of experience working in a variety of roles from software engineering to technical sales. He has worked in both enterprise and consumer companies such as VMware, EMC, GitHub, and LaunchDarkly. Adam is driven by a passion for inclusive leadership and solving problems with technology. As an Advisor he works with a number of startups and nonprofits. His perspective on life has been shaped by a background in Physics and Visual Art, an ongoing adventure as a husband and father, and a childhood career as a fire juggler.

Links:

  • Twitter: https://twitter.com/azimman

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

This episode is sponsored in part by our friends at VMware. Let’s be honest—the past year has been far from easy. Due to, well, everything. It caused us to rush cloud migrations and digital transformation, which of course means long hours refactoring your apps, surprises on your cloud bill, misconfigurations and headache for everyone trying manage disparate and fractured cloud environments. VMware has an answer for this. With VMware multi-cloud solutions, organizations have the choice, speed, and control to migrate and optimize

applications seamlessly without recoding, take the fastest path to modern infrastructure, and operate consistently across the data center, the edge, and any cloud. I urge to take a look at vmware.com/go/multicloud. You know my opinions on multi cloud by now, but there's a lot of stuff in here that works on any cloud. But don’t take it from me thats: VMware.com/go/multicloud and my thanks to them again for sponsoring my ridiculous nonsense.

Corey: This episode is sponsored in part by our friends at Jellyfish. So, you’re sitting in front of your office chair, bleary eyed, parked in front of a powerpoint and—oh my sweet feathery Jesus its the night before the board meeting, because of course it is! As you slot that crappy screenshot of traffic light colored excel tables into your deck, or sift through endless spreadsheets looking for just the right data set, have you ever wondered, why is it that sales and marketing get all this shiny, awesome analytics and inside tools? Whereas, engineering basically gets left with the dregs. Well, the founders of Jellyfish certainly did. That’s why they created the Jellyfish Engineering Management Platform, but don’t you dare call it JEMP! Designed to make it simple to analyze your engineering organization, Jellyfish ingests signals from your tech stack. Including JIRA, Git, and collaborative tools. Yes, depressing to think of those things as your tech stack but this is 2021. They use that to create a model that accurately reflects just how the breakdown of engineering work aligns with your wider business objectives. In other words, it translates from code into spreadsheet. When you have to explain what you’re doing from an engineering perspective to people whose primary IDE is Microsoft Powerpoint, consider Jellyfish. Thats Jellyfish.co and tell them Corey sent you! Watch for the wince, thats my favorite part.

Corey: Welcome to Screaming in the Cloud. I’m Cloud Economist Corey Quinn, and periodically I like to talk to people about different aspects of the industry. One that I think is interesting that doesn’t get spoken about a lot directly is the idea of leadership. My guest today is Adam Zimman, who’s a startup advisor providing guidance on—as mentioned—leadership, platform architecture, Product Marketing, and GTM
Strategy—GTM, of course, standing for go-to-market. Who goes to market? That’s right, little piggies. Adam, thank you for joining me.

Adam: Thank you, Corey. It’s a pleasure to be here.

Corey: I imagine that you usually don’t advise your clients to call their GTM execs, little piggies?

Adam: Well, I mean, I guess it depends. You know, if you’re actually a bacon manufacturer then that might be actually a reasonable thing to do.

Corey: Yeah, that’s a level of investment in the product that you usually don’t see in most environments, but we take what we can get. So, snark and cynicism aside, what is it you do?

Adam: Ultimately, I look for ways in which I can add value. And I’ve had the privilege in my career to be exposed to a lot of amazing companies, and I look for ways to be able to take the lessons that I’ve learned, mainly through mistakes and failure, and be able to translate those into success for others.

Corey: Most recently, you were at LaunchDarkly for a while, taking a number of different VP roles. While you were there we spoke, back in 2017, briefly while you were in that environment. And in fact, my first guest on the show was one of the folks on your team, Heidi Waterhouse, who has been back at least once since then, and hopefully more than that. But it’s been an interesting ride there. Before that you were at places like GitHub—or JIF-ub as I insist on pronouncing it—EMC-slash-VMware—where does one start and the other stop? Hard to say, it’s sort of a giant corporate shell game—but you’ve spent a lot of time in large companies and small ones as well, and now you’re effectively hanging out your shingle as a strategic advisor.

Adam: This is true. I mean, I think that one of the things that I’ve found is that doesn’t really matter what size of company you’re at; you’re going to find new and interesting challenges, and you really don’t have to look that hard. And so one of the things that I found consistently, and I would say that this was most pointedly phrased for me by Emily Freeman in the context of, “DevOps is this amazing thing of people, process, and technology. And the reality is, is the only one that’s complicated is the people.” And oddly enough, small companies, you still got people; big companies, you still got people. So, therein lies some of the challenges.

Corey: And people are inherently non-deterministic; you never know what you’re going to get by applying the same input, even to the same person just separated out by time. It’s a challenge, and the problem that I see across the industry is that very often, you’ll have a team of engineers and you’ll pick the best and brightest one of those engineers, and, “Congratulations, you manage the team now.” Now, management’s inherently orthogonal skill, and what you’ve simultaneously done is gotten rid of a great engineer and introduced a terrible manager. And that’s through no fault of this person’s own. But when I started managing teams, I got surprisingly far by just doing the exact opposite of all the stuff that my previous terrible bosses have done.

And that works really well right up until it doesn’t in a variety of probably fairly easily predictable ways. And the challenge that I’m seeing is that there is no book on how to do these things. If you want to climb an engineering ladder, great; there’s a bunch of very qualified people who will tell you how to go from wherever you are technically, to where you want to go, and what you have to demonstrate, and what you have to do. Leadership is squishy, in that sense. At least it always has been to me.

Adam: The interesting part that I would challenge you a little bit on is that there are thousands of interesting books on leadership, even smaller subsection on management specifically. I think one of the challenges there is that they’re not well circulated within tech as an industry. I think that there are a few that people come back to, like Andy Grove’s book on his experience building Intel. There are a lot of books out there that have done a lot for talking about how to manage people and how to think about what are the specific tactical things that you do. It’s having one-on-ones, it’s having meetings with clear agendas, it’s being able to look for ways to set expectations with your organization.

I think one of the challenges that I see pretty consistently, is the fact that that effort to be able to go out and find that information or to learn those skills is something that is put on to, as you said, this individual who is coming to management through punishment. They’ve been extraordinarily successful and now you will punish them by putting them in a role where they can no longer do all the things that they enjoyed, that made them successful. And I think that you see time and time again, where organizations put people in these roles, but they don’t do anything to either prepare them for it or do anything to continue that notion of professional development or training for those individuals once they’re in those roles.

Corey: There are a lot of books out there for any discipline under the sun; some are good, some are terrible, most are somewhere in the middle of the road law of averages winds up working out. I think a key difference, on some level, is I can take to Twitter, or a forum, or something like that, and complain about software; the computer isn’t doing the thing I think the computer should be doing. And that’s great. I can’t very well go and complain about managerial issues while actively having a team and not find myself no longer having managerial issues, if you catch my meaning. It’s hard to find communities around this stuff.

Adam: I think that you’re right. And I think that this is one of those things where not only that, but I think that we also in tech have predominantly taken a very hierarchical structure to the way that we think about management and leadership, to the sense where oftentimes, it is not only discouraged but downright forbidden for an individual contributor to challenge their manager if they want to continue to have gainful employment. And I think that this is a cultural thing that, you know, it’s funny; I know that you recently did an episode with John Allspaw and were talking about incident remediation. And I think that one of the things that I’ve always tried to do as a manager, as a leader, is think about opportunities for being able to do that type of incident response, for people. If you have a person that leaves, whether that is forced attrition, whether that is voluntary attrition, whether that is something that you wanted to happen, something that you didn’t want to happen, what are you doing from a perspective of kind of a post-incident assessment to learn from that? And I think that the next level that is, how do you do it so that you actually, in some way, incorporate that for the individual that’s actually leaving. Because ideally, they’re learning from that experience, as well.

Corey: Back when I was a generally terrible employee, I decided at some point, I was tired of dealing with computer problems and wanted to deal with people problems instead. Now, let’s be clear, I found a path to do that in a very different direction than I expected at the time, but at the time, it was, “Great. I’m going to go ahead and become a manager of a team.” And I talked to a number of folks about all right, what is the path to go from decent technical engineer—I was a senior SRE type at most of these places—into management. And not just talking to people at the companies I was at, but talking to people in the larger community, and every engineering manager who I respected and talked to about,
it always seemed like they got this lucky break at just the right time and that made them a manager for the first time.

And once you have a track record of having managed people, then you’re in. You can go back and forth between IC and management roles. But, “Well, you’ve never managed people before, so we’re not going to take a chance on you to manage people.” The way that I did it, honestly, was I—a few times—I wound up joining startups where I was effectively the only ops person; we suddenly started scaling and having fun problems, and well, I did negotiate for that director title, so all right, I have teams now. I was more of a team lead than most things, in some cases.

But it led to a really pretty interesting evolution in how I approach these things. I find now that the right answer is for me not to manage people at all because what I fundamentally do here at The Duckbill Group is basically become the loud, obnoxious center of attention. And I think that what managers need to do is showcase their people instead. And those two things, at least in my view, are opposed. And it’s very challenging to do both of them, let alone well. For me at least, I tend to back away from the management side of things almost entirely and abdicate the
role. Which is great. People self-manage, right?

Adam: Well, I mean, I think that there are individuals who definitely will take—have the ability to self-organize and self-manage to a degree. I think that the challenge that you run into is, as the organization scales, as the nature of their role tends to change with that scaling organization, it becomes more challenging for them to navigate through those changes. A great example would be, I have had the pleasure and the privilege a number of times in my career of managing extraordinarily senior individuals; these are individuals who, to your point, don’t need a whole lot of care and feeding. But what they do sometimes need is they need someone who is able to be in rooms that they’re not in, whether that’s from a higher-level leadership meeting understanding larger organizational goals, or they need someone that’s going to check them; they need someone that they can trust, someone that they can bounce their ideas off of to know is this something that’s going to be perceived value or something that’s going to actually take me in the wrong direction, or somebody that’s, kind of like, paying attention to the work product that they’re doing and giving them some coaching, whether that’s cheerleading or whether that’s connecting of saying, “Hey, there’s also this other person you should talk to.” Those types of things are really valuable for those individuals who are, to your point, a little bit more self-sufficient.

Corey: On some level, I ran into this trap a lot, and having over drinks conversations with a bunch of people who went on similar paths, it’s blindingly obvious that it’s a dumb move in hindsight, but an awful lot of us did it, where we’re sitting there as engineers with the belief of, “Ah, if I can make my manager—or beyond, several skip-levels up—look incredibly foolish in the middle of a large meeting, they will inherently see the value of what I have to say and will thus elevate me to management.” As it turns out, they elevate you to customer because you’re not working there anymore, in many cases. And when I talk to people about this, it usually has that lightbulb coming on moment of as soon as you hear it, of course, it is blindingly obvious that you aren’t going to sarcastically obnoxious your way into being management. Instead, the path there—in hindsight, also blindly obvious—is act as if: act managerial; help to effectively carry on your manager’s message to the rest of the team, and when you have reservations or whatnot, talk to them in private rather than calling them out. And it’s the obvious stuff of who gets promoted to management? Well, the people that look managerial. And that is what that looks like, in many respects.

Adam: And this is one of the reasons why, when I talk about management I like to separate the notion of management from leadership. Because I think that anyone can be a leader. You don’t actually have to be the administrative manager of an individual to be a leader to them.

Corey: I saw a great poster once when I was younger. “Leaders are like eagles. We don’t have either of them here.”

Adam: [sigh]. Yeah, yeah. Ugh. I do miss good motivational posters.

Corey: Oh, yeah.

Adam: You know, I think that there’s some truth to it. I think that finding people who are genuinely invested in being able to enable the success of others—which is how I define leadership—is challenging. I think that, especially in rather capitalistic-type industry like we’re in, there is a lot of measurement of people’s success by their own personal achievements and by their ability to beat their own drum. And I think that it’s something that is, frankly, a failing of our industry, where we don’t do a better job of encouraging folks, and rewarding folks that actually look out for others and enable the success of others. Because I think that’s something that is—ultimately you think about how you build strong teams, and it’s not about getting a bunch of individuals who can do amazing things individually. It’s about getting individuals who are capable of working together and being able to do more than they would be able to if they were simply working individually.

Corey: Do you ever find that people are chasing management in many respects because they think that it’s something very different than what it is, and then find themselves in situations where well, I’m the dog that caught the car that I was chasing and only now do I realize that I have no idea how to drive the thing?

Adam: Oh, absolutely. So, this is something that has been interesting me a lot recently, in the sense that I think we as an industry also do a very poor job of measuring management, measuring leadership. We give a lot of power to managers through performance reviews to measure their individual contributors, but there are very few companies who actually efficiently do things like 360 reviews, which has always confused me because I think that implies that you’re getting feedback from all around you, as opposed to what you really want is you want feedback pointed back at you, which would be 180. But maybe that’s just—

Corey: Let’s be clear, that was also pioneered by the German [Wehrmacht 00:13:48] in World War II, which is yeah, basically how some people I’ve worked with do tend to manage.

Adam: Yeah. I think that if we can think about how do we measure the success of a manager, is it simply a function of the output of their team, or are there other efficiency metrics that you should be looking at? Very obvious one is how efficient is a manager from a perspective of the utilization of their resources? And when I think about that, I think about are they actually able to effectively hire? Are they able to effectively retain the people that they hire?

What does it look like for the people on their organization from a promotion perspective in terms of skill growth? Do they become more valuable over time? Those are ways in which we can think about how we measure the manager, potentially, directly. And then there’s indirect things like what’s the qualitative aspect of those individuals that work for them? Are they people who are enjoying the work that they’re doing?

Are they motivated to continue to work towards the company’s vision and mission, to be able to actually make their manager look good, but also make the company successful?

Corey: A challenge, too, because I’ve seen this myself is, all right, you’re not elevated to manager. Congratulations. It’s not really a promotion. It’s a lateral move. However, a lot of companies don’t treat it that way.

They don’t compensate it that way, et cetera. And oh, okay, management, it turns out is not for me. There’s no real good way to say, “I’m going back to being an IC,” especially at the same company, without it being perceived by many—rightly or wrongly—as a demotion or a failure.

Adam: This question of, like, motivation to people, why do they want to go into management? I think that oftentimes this is misplaced. A lot of times the number one motivation that I’ve heard has nothing to do with wanting to actually help people or solve people problems, as you said earlier; it has to do with I want a bigger paycheck, I want more seniority, I want more responsibility, and therefore the only path available to me is management. In fact, many career ladders at organizations require an individual contributor to go to a management position before they can become a principal or a staff-level engineer, which is nonsense. First of all, why would you torture the individual to do something that is so completely and utterly outside of where their interests are? Secondly, why would you just decimate your lower-level individual contributors, your newer individual contributors by having someone who is completely non-inclined towards management be responsible for them? Oh.

Corey: Oh, yeah. Used to be your peer; now they manage you, and great. I think people underestimate exactly how broad the blast radius of a manager is.

Adam: Yeah. Talk to anyone, and they’ll be more than happy to tell you the worst manager that they’ve ever had. At the same time, they’ll also probably be able to tell you the best manager they’ve ever had.

Corey: Oh, yeah. I called both of those out—only one the one of those by name, by the way—in conference talks that I’ve had because it’s—yeah, you can probably guess which one I would call out and which one I would not name publicly—yeah—

Adam: It depends on the conference, I guess. But yeah.

Corey: Oh, yeah, absolutely. If it was you-know-what-your-problem-is con, yeah, it went super well.

Adam: [laugh].

Corey: It was fun. And management, especially in the current era is getting interesting, as we’re seeing the heating up of the market in a bunch of different ways. And I understand, to be clear, that Twitter is not a perfect microcosm of the industry, but there’s a recurring theme that I’m seeing among a number of engineering types that seemed to get—and again, I don’t want to get letters for this, so if I misstate it, audience, please go ahead and be kind—but there seems to be a certain thread running through engineering communities that the purpose of a company is to provide a utopian work environment for its staff. Now, as someone who runs a company myself, yeah, I absolutely want to provide the kind of working environment I wish I’d had in a bunch of different environments. And that’s not going to work for everyone, but that’s okay.

But fundamentally we’re here to make money, and ideally, enough monies that we can keep the lights on. And that does mean that, however, we want to treat our staff that has to be subordinate to can we continue as a going concern? So yeah, it turns out, we can’t—sustainably—outbid Netflix on every hire that we make and we aren’t able to wind up having three catered meals a day as a full remote company delivered to everyone’s house. Now, I’d like to, in a world where money flows like water, but it doesn’t. For better or worse, there are constraints, and constraints shape us.

But there’s a thread that I’m starting to see of… I hesitate to call it entitlement, but it trends slightly toward the direction of folks who are in tech, and in some ways seem very far removed from business realities—now, let’s be clear in the FAANG world, yeah, it’s pretty attenuated. And in startup land where well, we’re the VC backed, so we’re losing money by the billion but we’re making it up in volume. Great. That is not necessarily what I’m talking about here. I’m seeing a thread where, oh, engineers are clearly the smartest people in any company, which means that every other department should defer to them. I disagree with that position.

Adam: I want to follow that thread a little bit with regards to engineers. So, I’ve worked as a software developer—

Corey: My condolences.

Adam: Yeah. I’ve worked as a technical salesperson. I’ve had the opportunity to work in pretty much every department with the exceptions of HR and finance. So, that has been part of my career of jack of all trades, master of none, but it has given me some interesting insights in terms of the value that different organizations, different individuals, bring to a company. And I think that—one of the things that I will say is that for the longest time, in large organizations, especially non-tech industry organizations, the engineer or the developer was at the same expectations or the role as someone in the janitorial staff.

It was basically, “You’re part of the plumbing. You just do the things so that the tech just works, and we’re going to have the other business folks that are more responsible for actually making decisions that are going to make our business money.” The quintessential example is someone like Kraft Foods or someone like John Deere, right, where you’re building tractors; for the longest time, the guy who ran the website wasn’t going to be the guy who was going to make or break John Deere’s quarterly earnings. Now, you’ve got tractors that literally are more computers than they are mechanical devices and so you suddenly have this change in dynamic with regards to the importance of that developer. But I think that something that’s interesting, also, is that those other people who worked at the company didn’t go away.

They’re still there; they’re still important. In fact, they’re still oftentimes making the buying decisions on behalf of the developers. The developers aren’t the ones that are making those choices. And so you need to figure out, how do you actually make the technology choices and the technology outcomes accessible to individuals that are in roles that were, historically, had nothing to do with tech.

This episode is sponsored by our friends at Oracle Cloud. Counting the pennies, but still dreaming of deploying apps instead of "Hello, World" demos? Allow me to introduce you to Oracle's Always Free tier. It provides over 20 free services and infrastructure, networking databases, observability, management, and security.

And - let me be clear here - it's actually free. There's no surprise billing until you intentionally and proactively upgrade your account. This means you can provision a virtual machine instance or spin up an autonomous database that manages itself all while gaining the networking load, balancing and storage resources that somehow never quite make it into most free tiers needed to support the application that you want to build.

With Always Free you can do things like run small scale applications, or do proof of concept testing without spending a dime. You know that I always like to put asterisks next to the word free. This is actually free. No asterisk. Start now. Visit https://snark.cloud/oci-free that's https://snark.cloud/oci-free.

Corey: I’ve always been a big believer in the idea that if you’re going to transition into a new field, be it into tech, out of tech, et cetera, great. In almost every case, you should find ways to do that laterally. I think that this idea that, oh, you’re going to go ahead and just start over with an entry-level job after you’ve been in a field for five years—no. Find the position that’s halfway between where you are and where you think you want to go next and start getting exposure there. In time, it’s those niches that add value that distinguish you from other folks.

It turns out that they don’t generally want to hire someone in almost any role that comes from Central Casting, where it’s alright, give me a
standard MBA with the following pedigree and drop them in as my new executive, whatever. No. They want to see things like industry experience; they want to see things that distinguish folks, and having experience in industries that are not traditionally, purely what this role is, is super helpful in a lot of different ways. What I do pretty clearly blends finance and tech; that goes reasonably well. Increasingly it starts to blend media, which is something I don’t pretend to understand. But here we are, he said into the microphone.

Adam: Yeah. Well, as long as you’re not starting the next Fox News, I’m fine with that.

Corey: No, no. Generally not.

Adam: Okay, fair enough. But I think that you’re right. This is one of the things where, trailing back, we’ve throughout this conversation to the notion of leadership, this is something that I found extraordinarily rewarding and empowering that I’ve done with individuals that I’ve brought into new organizations, either through initial conversations during an interview process, or during, as part of their onboarding, is I sit down, and I actually talk to them about what are their plans? What are their expectations? What are their goals, not only for the next 30, 60, 90 days in this role that we’re talking about but what are they thinking about from a perspective of what do they want to do in the next year? In the next three years? Five years? Ten years? What are those checkpoints of what do you want to do in this role? What do you want to do at this company? What do you want to do with your career? Like, where do you see it headed?

And it doesn’t mean that you’re writing this in stone, or that I’m going to hold you to it, but I think that one of those things that’s really empowering for a leader is to be able to help those individuals find those connective threads that tie one position to the next and help them get there. If they’re somebody who is saying, “Hey, look, I’m currently a developer, but I really wish that I could give more talks.” Okay, well, that’s great for me to know. Let’s put you on some projects that maybe actually would result in great content for a talk that you could give at a conference. And then we’ll figure out, how do we work with the marketing department to be able to help you bring that to fruition?

There’s a lot of ways to be able to leverage this experience that you have as a leader, as a manager, to an individual who’s coming up in their career and saying, “Hey, look. This is how some more ancillary things are connected.” And being able to bring those back to them.

Corey: I really wish, on some level, that there was a more defined path toward a lot of these things, where the stuff is explained to folks. So often, I had terrible managers that, in hindsight, weren’t that terrible. Because I didn’t understand where the role started and stopped, I tended to view the role of the manager is there to protect the team. The end. And be our advocate in the organization, and get us the thing that we want, and what do we want? Comfy chairs.

And it turns out that isn’t ever how it really works. If I had to define management, it would basically be, balancing competing priorities more than it is almost anything else. And counterintuitively, the higher you rise in an organization, the more responsibility you have, and the less you can actually directly do. Everything you do drives influence. And that’s it. That’s how it distills down.

Adam: You talk about the engineer that wants to move into management role because that’s how they see their career progressing. This is a close corollary to the engineer that wants to move into a product management role because they want to have greater oversight into the decisions that are being made about what’s getting built. And what you come to realize, for any engineer who successfully made that transition, is it’s really complicated and difficult to be able to have that mental switch take place between this is how I’m going to build it versus this is the priority of what needs to get built next. And all too often you see engineers that land in product management roles that are dictating how something should be built, and suddenly the engineers are just like, “No, I have no respect for you. Because that’s not your job.”

And likewise, in a management role, oftentimes people view that as an opportunity for them to make all the choices, make all the decisions, and suddenly lose sight of the fact that they used to be on the other side of that outcome themselves, and were disappointed when they weren’t included in some way, shape or form, or their priorities weren’t taken into consideration.

Corey: As you look at your own career, what is the worst job experience you’ve ever had? Or the worst job you’ve ever had? Or the worst boss you’ve ever had? That’s always a good one to do.

Adam: [laugh].

Corey: Pick a superlative and not the good kind. Hit me.

Adam: Yeah, no, I mean, look, I think that probably the worst… experience that I ever had with a manager, with a boss, was actually when I was first a software developer. And my manager would occasionally just come up behind me and just stand and watch me code. And we’re not talking about peer programming, where it was just like, we’re working together. No, it was, literally would come up, stand behind me on my shoulder, and just stand there. Not saying anything; just watching me write Java code. And that was probably the most disconcerting
experience that I’ve ever had in a job ever. I lasted about six months and then I was just like, “I need to move on to something else.”

Corey: It turns out one of my failure modes was that I was great for the first three months in new ops roles because things were invariably a fire, and—

Adam: [laugh].

Corey: —I know how to solve those things. And then it becomes a maintenance role, and I’m bad at that. For longest time, I thought I was just a crap employee. And I am, but for different reasons. Instead, though, for me, it turned into a, I need to find the thing that I’m good at and embrace that. And I have to say, it was not being, basically, a cloud comedian on Twitter where my primary means of communication is shitposting. But you know, here we are, and this is how we’ve gotten there.

Adam: I mean, know your strengths, man. Know your strengths.

Corey: Yeah, lean into it. I mean, you went to college in Maine; you know what it’s like there. It’s dark and cold nine months out of the year, so all we do is sit inside and develop personality disorders. And well, here we are.

Adam: Well, hey, I mean, I took a break from tech after that first job in software development and I actually went back and worked for a guy that I met while I was in school, and I worked for him, he was a general contractor. So, I have an appreciation for Maine winters in a way that I never gained as a privileged college student, when I was actually digging snow out of ditches to be able to pour concrete at six in the morning and then later in the day, I got to go up and use 80-pound weight shingles to reshingle the roof in 20-degree weather. So, it was an eye-opening experience. But I’ll tell you, I learned pretty much everything that I know about how to build infrastructure from that eight months that I spent doing everything from framing, ditch-digging, to electrical, and plumbing, and roofing.

Corey: Kind of fun how often is that we wind up trying other things. And this is part of it, too. As much fun as it is to complain about various jobs and whatnot that we have, let’s be very clear here for a minute that I’m not dealing with hot tar, being paid seven bucks an hour. There are advantages to the [unintelligible 00:28:08] jobs I have.

Adam: I mean, that was a number of years ago, but I still got ten bucks an hour.

Corey: My first job at the University of Maine call center working in tech, in those days, I think I was being paid something like $5.35 an hour. To answer phones, which again, not that hard of a job. I made a lot more money a couple years later when I moved to construction. Yeah, I wouldn’t recommend any of those things for me these days, but it was instructive.

Adam: But at the same time, I would argue that you also have benefited from those experiences in the way that you approach the things that you do now. And I think that’s one of the things that I’ve tried to bring forward in my career is look for those opportunities to make those
connections, and understand the value of those experiences, and be able to help to enable other people because I’ve had those experiences.

Corey: To me at least, the answer is to turn whatever you’ve done or whatever happened to you into some form of empathy. The idea of well, I had to struggle coming up, so you should, too. Let’s instead focus on making it better for people who follow us. Send the elevator back down,
as it were.

Adam: I mean, I think that’s great advice, and I think that it’s something that’s done far too infrequently. One of the things that I’ve noticed is that that aspect, unless somebody has actually been through the experience where somebody has done that for them, it is oftentimes something that is a lot harder for people to see. This goes to your earlier statement around the expectations that maybe are changing, and they’re not such great ways with regards to what people are expecting from companies, what people are expecting from managers. I think that there is a distinct lack of expectation setting that takes place at companies in terms of what is the role of the company, what is the role of an employee, and how can those two come together to still have a positive interaction, but aren’t overstepping on either side? Because that’s really where you get into problems. That’s where all of a sudden you have these companies that are looking to fill the role of, I will take care of all aspects of your life, when in reality that’s not a very healthy relationship for an individual to have with a company.

Corey: So, I want to thank you for coming and speak to me. What are you up to these days, and where can people find you? And why should
people find you?

Adam: Well, I don’t know that anybody should find me.

Corey: “I hope this email finds you never. I hope you’re free.”

Adam: Yeah, exactly. No, I mean, I would love to find folks that I can add value to and help out. It’s easy enough to find me on Twitter. It’s just @-A-Z-I-M-M-A-N—azimman. And they’re welcome to reach out to me there. My DMs are open—much to my displeasure sometimes—but happy to help people who are looking for help. I’m particularly interested in spending my time with those individuals who maybe are coming from underrepresented backgrounds in tech and looking for ways to be able to either get into tech or to move up within leadership roles in tech.

But I’m spending a lot of my time doing a lot of coaching, doing a lot of advising for small startups, and then also just as a small side project have been working pretty extensively with James Governor and a woman by the name of Kim Harrison on this little thing called Progressive Delivery, which is, as far as we’re concerned, it is the next iteration of the software development lifecycle that we’ve written about and talked about pretty extensively. James and Kim and I are working on a book together to be able to capture all those ideas and bring them and coalesce them for people, to make more consumable. But ultimately, we’re trying to say, “Hey, look. The way that we’ve done things leading up till now, moving from waterfall to agile to continuous delivery into what’s next?” And look at some of the market conditions that have changed. A lot of stuff that you talk about. I think that you would be the first to point out how things have changed since the launch of AWS.

Corey: Oh, yes. It’s more confusing now.

Adam: Oh, way more confusing. And the ways in which people consume cloud-based services has radically changed. And so I think that the way that we are building software and the way that we’re consuming software is something that we need to put some serious thought into. And the players that are—you know, as I spoke about earlier on this talk with you—are different. It’s no longer just your developers that care about your AWS choices or care about the cloud service choices that you’re making.

You’ve got other individuals, whether it’s the finance side you focus on or thinking about it from the perspective of the marketing team, or the HR team that’s thinking about which cloud service HRIS are they going to use. There’s a lot of people that need to be party to those choices that you’re making and how you build out your company stack, as it were. And the Progressive Delivery model looks to take into consideration that changing and evolving group of people.

Corey: And we will, of course, have links to that in the [show notes 00:32:46]. Thank you so much for taking the time to speak with me. I appreciate it.

Adam: Corey, thank you so much for having me. It was a pleasure.

Corey: Adam Zimman, startup advisor, and oh, so much more. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you hated this podcast, please leave a five-star review on your podcast platform of choice, along with a scathing comment telling me why you as an engineer are best suited to be the manager of everything.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Karthik

Karthik was one of the original database engineers at Facebook responsible for building distributed databases including Cassandra and HBase. He is an Apache HBase committer, and also an early contributor to Cassandra, before it was open-sourced by Facebook. He is currently the co-founder and CTO of the company behind YugabyteDB, a fully open-source distributed SQL database for building cloud-native and geo-distributed applications.

Links:

  • Yugabyte community Slack channel: https://yugabyte-db.slack.com/
  • Distributed SQL Summit: https://distributedsql.org
  • Twitter: https://twitter.com/YugaByte

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: You could build you go ahead and build your own coding and mapping notification system, but it takes time, and it sucks! Alternately, consider Courier, who is sponsoring this episode. They make it easy. You can call a single send API for all of your notifications and channels. You can control the complexity around routing, retries, and deliverability and simplify your notification sequences with automation rules. Visit courier.com today and get started for free. If you wind up talking to them, tell them I sent you and watch them wince—because everyone does when you bring up my name. Thats the glorious part of being me. Once again, you could build your own notification system but why on god’s flat earth would you do that?

Corey: This episode is sponsored in part by “you”—gabyte. Distributed technologies like Kubernetes are great, citation very much needed, because they make it easier to have resilient, scalable, systems. SQL databases haven’t kept pace though, certainly not like no SQL databases have like Route 53, the world’s greatest database. We’re still, other than that, using legacy monolithic databases that require ever growing instances of compute. Sometimes we’ll try and bolt them together to make them more resilient and scalable, but let’s be honest it never works out well. Consider Yugabyte DB, its a distributed SQL database that solves basically all of this. It is 100% open source, and there's not asterisk next to the “open” on that one. And its designed to be resilient and scalable out of the box so you don’t have to charge yourself to death. It's compatible with PostgreSQL, or “postgresqueal” as I insist on pronouncing it, so you can use it right away without having to learn a new language and refactor everything. And you can distribute it wherever your applications take you, from across availability zones to other regions or even other cloud providers should one of those happen to exist. Go to yugabyte.com, thats Y-U-G-A-B-Y-T-E dot com and try their free beta of Yugabyte Cloud, where they host and manage it for you. Or see what the open source project looks like—its effortless distributed SQL for global apps. My thanks to Yu—gabyte for sponsoring this episode.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Today’s promoted episode comes from the place where a lot of my episodes do: I loudly and stridently insist that Route 53—or DNS in general—is the world’s greatest database, and then what happens is a whole bunch of people who work at database companies get upset with what I’ve said. Now, please don’t misunderstand me; they’re wrong, but I’m thrilled to have them come on and demonstrate that, which is what’s happening today. My guest is CTO and co-founder of Yugabyte. Karthik Ranganathan, thank you so much for spending the time to speak with me today. How are you?

Karthik: I’m doing great. Thanks for having me, Corey. We’ll just go for YugabyteDB being the second-best database. Let’s just keep the first [crosstalk 00:01:13]—

Corey: Okay. We’re all fighting for number two, there. And besides, number two tries harder. It’s like that whole branding thing from years past. So, you were one of the original database engineers at Facebook, responsible for building a bunch of nonsense, like Cassandra and HBase. You were an HBase committer, early contributor to Cassandra, even before it was open-sourced.

And then you look around and said, “All right, I’m going to go start a company”—roughly around 2016, if memory serves—“And I’m going to go and build a database and bring it to the world.” Let’s start at the beginning. Why on God’s flat earth do we need another database?

Karthik: Yeah, that’s the question. That’s the million-dollar question isn’t it, Corey? So, this is one, fortunately, that we’ve had to answer so many times from 2016, that I guess we’ve gotten a little good at it. So, here’s the learning that a lot of us had from Facebook: we were the original team, like, all three of us founders, we met at Facebook, and we not only build databases, we also ran them. And let me paint a picture.

Back in 2007, the public cloud really wasn’t very common, and people were just going into multi-region, multi-datacenter deployments, and Facebook was just starting to take off, to really scale. Now, forward to 2013—I was there through the entire journey—a number of things happened in Facebook: we saw the rise of the equivalent of Kubernetes which was internally built; we saw, for example, microservice—

Corey: Yeah, the Tupperware equivalent, there.

Karthik: Tupperware, exactly. You know the name. Yeah, exactly. And we saw how we went from two data centers to multiple data centers, and nearby and faraway data centers—zones and regions, what do you know as today—and a number of such technologies come up. And I was on the database side, and we saw how existing databases wouldn’t work to distribute data across nodes, failover, et cetera, et cetera.

So, we had to build a new class of databases, what we now know is NoSQL. Now, back in Facebook, I mean, the typical difference between Facebook and an enterprise at large is Facebook has a few really massive applications. For example, you do a set of interactions, you view profiles, you add friends, you talk with them, et cetera, right? These are supermassive in their usage, but they were very few in their access patterns. At Facebook, we were mostly interested in dealing with scale and availability.

Existing databases couldn’t do it, so we built NoSQL. Now, forward a number of years, I can’t tell you how many times I’ve had conversations with other people building applications that will say, “Hey, can I get a secondary index on the SQL database?” Or, “How about that transaction? I only need it a couple of times; I don’t need it all the time, but could you, for example, do multi-row transactions?” And the answer was always, “Not,” because it was never built for that.

So today, what we’re seeing is that transactional data and transactional applications are all going cloud-native, and they all need to deal with scale and availability. And so the existing databases don’t quite cut it. So, the simple answer to why we need it is we need a relational database that can run in the cloud to satisfy just three properties: it needs to be highly available, failures or no, upgrades or no, it needs to be available; it needs to scale on demand, so simply add or remove nodes and scale up or down; and it needs to be able to replicate data across zones, across regions, and a variety of different topologies. So availability, scale, and geographic distribution, along with retaining most of the RDBMS features, the SQL features. That’s really what the gap we’re trying to solve.

Corey: I don’t know that I’ve ever told this story on the podcast, but I want to say it was back in 2009. I flew up to Palo Alto and interviewed at Facebook, and it was a different time, a different era; it turns out that I’m not as good on the whiteboard as I am at running my mouth, so all right, I did not receive an offer, but I think everyone can agree at this point that was for the best. But I saw one of the most impressive things I’ve ever seen, during a part of that interview process. My interview is scheduled for a conference room for must have been 11 o’clock or something like that, and at 10:59, they’re looking at their watch, like, “Hang on ten seconds.” And then the person I was with reached out to knock on the door to let the person know that their meeting was over and the door opened.

So, it’s very clear that even in large companies, which Facebook very much was at the time, people had synchronized clocks. This seems to be a thing, as I’ve learned from reading the parts that I could understand of the Google Spanner paper: when you’re doing distributed databases, clocks are super important. At places like Facebook, that is, I’m not going to say it’s easy, let’s be clear here. Nothing is easy, particularly at scale, but Facebook has advantages in that they can mandate how clocks are going to be handled throughout every piece of their infrastructure. You’re building an open-source database and you can’t guarantee in what environment and on what hardware that’s going to run, and, “You must have an atomic clock hooked up,” is not something you’re generally allowed to tell people. How do you get around that?

Karthik: That’s a great question. Very insightful, cutting right to the chase. So, the reality is, we cannot rely on atomic clocks, we cannot mandate our users to use them, or, you know, we’d not be very popularly used in a variety of different deployments. In fact, we also work in on-prem private clouds and hybrid deployments where you really cannot get these atomic clocks. So, the way we do this is we come up with other algorithms to make sure that we’re able to get the clocks as synchronized as we can.

So, think about at a higher level; the reason Google uses atomic clocks is to make sure that they can wait to make sure every other machine is synchronized with them, and the wait time is about seven milliseconds. So, the atomic clock service, or the true time service, says no two machines are farther apart than about seven milliseconds. So, you just wait for seven milliseconds, you know everybody else has caught up with you. And the reason you need this is you don’t want to write on a machine, you don’t want to write some data, and then go to a machine that has a future or an older time and get inconsistent results. So, just by waiting seven milliseconds, they can ensure that no one is going to be older and therefore serve an older version of the data, so every write that was written on the other machine see it.

Now, the way we do this is we only have NTP, the Network Time Protocol, which does synchronization of time across machines, except it takes 150 to 200 milliseconds. Now, we wouldn’t be a very good database, if we said, “Look, every operation is going to take 150 milliseconds.” So, within these 150 milliseconds, we actually do the synchronization in software. So, we replaced the notion of an atomic clock with what is called a hybrid logical clock. So, one part using NTP and physical time, and another part using counters and logical time and keep exchanging RPCs—which are needed in the course of the database functioning anyway—to make sure we start normalizing time very quickly.

This in fact has some advantages—and disadvantages, everything was a trade-offs—but the advantage it has over a true time-style deployment is you don’t even have to wait that seven milliseconds in a number of scenarios, you can just instantly respond. So, that means you get even lower latencies in some cases. Of course, the trade-off is there are other cases where you have to do more work, and therefore more latency.

Corey: The idea absolutely makes sense. You started this as an open-source project, and it’s thriving. Who’s using it and for what purposes?

Karthik: Okay, so one of the fundamental tenets of building this database—I think back to your question of why does the world need another database—is that the hypothesis is not so much the world needs another database API; that’s really what users complain against, right? You create a new API and—even if it’s SQL—and you tell people, “Look. Here’s a new database. It does everything for you,” it’ll take them two years to figure out what the hell it does, and build an app, and then put it in production, and then they’ll build a second and a third, and then by the time they hit the tenth app, they find out, “Okay, this database cannot do the following things.” But you’re five years in; you’re stuck, you can only add another database.

That’s really the story of how NoSQL evolved. And it wasn’t built as a general-purpose database, right? So, in the meanwhile, databases like Postgres, for example, have been around for so long that they absorb and have such a large ecosystem, and usage, and people who know how to use Postgres and so on. So, we made the decision that we’re going to keep the database API compatible with known things, so people really know how to use them from the get-go and enhance it at a lower level to make a cloud-native. So, what is YugabyteDB do for people?

It is the same as Postgres and Postgres features of the upper half—it reuses the code—but it is built on the lower half to be [shared nothing 00:09:10], scalable, resilient, and geographically distributed. So, we’re using the public cloud managed database context, the upper half is built like Amazon Aurora, the lower half is built like Google Spanner. Now, when you think about workloads that can benefit from this, we’re a transactional database that can serve user-facing applications and real-time applications that have lower latency. So, the best way to think about it is, people that are building transactional applications on top of, say, a database like Postgres, but the application itself is cloud-native. You’d have to do a lot of work to make this Postgres piece be highly available, and scalable, and replicate data, and so on in the cloud.

Well, with YugabyteDB, we’ve done all that work for you and it’s as open-source as Postgres, so if you’re building a cloud-native app on Postgres that’s user-facing or transactional, YugabyteDB takes care of making the database layer behave like Postgres but become cloud-native.

Corey: Do you find that your users are using the same database instance, for lack of a better term? I know that instance is sort of a nebulous term; we’re talking about something that’s distributed. But are they having database instances that span multiple cloud providers, or is that
something that is more talk than you’re actually seeing in the wild?

Karthik: So, I’d probably replace the word ‘instance’ with ‘cluster’, just for clarity, right?

Corey: Excellent. Okay.

Karthik: So, a cluster has a bunch—

Corey: I concede the point, absolutely.

Karthik: Okay. [laugh]. Okay. So, we’ll still keep Route 53 on top, though, so it’s good. [laugh].

Corey: At that point, the replication strategy is called a zone transfer, but that’s neither here nor there. Please, by all means, continue.

Karthik: [laugh]. Okay. So, a cluster database like YugabyteDB has a number of instances. Now, I think the question is, is it theoretical or real?
What we’re seeing is, it is real, and it is real perhaps in slightly different ways than people imagine it to be.

So, I’ll explain what I mean by that. Now, there’s one notion of being multi-cloud where you can imagine there’s like, say, the same cluster that spans multiple different clouds, and you have your data being written in one cloud and being read from another. This is not a common pattern, although we have had one or two deployments that are attempting to do this. Now, a second deployment shifted once over from there is where you have your multiple instances in a single public cloud, and a bunch of other instances in a private cloud. So, it stretches the database across public and private—you would call this a hybrid deployment topology—that is more common.

So, one of the unique things about YugabyteDB is we support asynchronous replication of data, just like your RDBMSs do, the traditional RDBMSs. In fact, we’re the only one that straddles both synchronous replication of data as well as asynchronous replication of data. We do both. So, once shifted over would be a cluster that’s deployed in one of the clouds but an asynchronous replica of the data going to another cloud, and so you can keep your reads and writes—even though they’re a little stale, you can serve it from a different cloud. And then once again, you can make it an on-prem private cloud, and another public cloud.

And we see all of those deployments, those are massively common. And then the last one over would be the same instance of an app, or perhaps even different applications, some of them running on one public cloud and some of them running on a different public cloud, and you want the same database underneath to have characteristics of scale and failover. Like for example, if you built an app on Spanner, what would you do if you went to Amazon and wanted to run it for a different set of users?

Corey: That is part of the reason I tend to avoid the idea of picking a database that does not have at least theoretical exit path because
reimagining your entire application’s data model in order to migrate is not going to happen, so—

Karthik: Exactly.

Corey: —come hell or high water, you’re stuck with something like that where it lives. So, even though I’m a big proponent as a best practice—and again, there are exceptions where this does not make sense, but as a general piece of guidance—I always suggest, pick a provider—I don’t care which one—and go all-in. But that also should be shaded with the nuance of, but also, at least have an eye toward theoretically, if you had to leave, consider that if there’s a viable alternative. And in some cases in the early days of Spanner, there really wasn’t. So, if you needed that functionality, okay, go ahead and use it, but understand the trade-off you’re making.

Now, this really comes down to, from my perspective, understand the trade-offs. But the reason I’m interested in your perspective on this is because you are providing an open-source database to people who are actually doing things in the wild. There’s not much agenda there, in the same way, among a user community of people reporting what they’re doing. So, you have in many ways, one of the least biased perspectives on the entire enterprise.

Karthik: Oh, yeah, absolutely. And like I said, I started from the least common to the most common; maybe I should have gone the other way. But we absolutely see people that want to run the same application stack in multiple different clouds for a variety of reasons.

Corey: Oh, if you’re a SaaS vendor, for example, it’s, “Oh, we’re only in this one cloud,” potential customers who in other clouds say, “Well, if that changes, we’ll give you money.” “Oh, money. Did you say ‘other cloud?’ I thought you said something completely different. Here you go.”
Yeah, you’ve got to at some point. But the core of what you do, beyond what it takes to get that application present somewhere else, you usually keep in your primary cloud provider.

Karthik: Exactly. Yep, exactly. Crazy things sometimes dictate or have to dictate architectural decisions. For example, you’re seeing the rise of compliance. Different countries have different regulatory reasons to say, “Keep my data local,” or, “Keep some subset of data are local.”

And you simply may not find the right cloud providers present in those countries; you may be a PaaS or an API provider that’s helping other people build applications, and the applications that the API provider’s customers are running could be across different clouds. And so they would want the data local, otherwise, the transfer costs would be really high. So, a number of reasons dictate—or like a large company may acquire another company that was operating in yet another cloud; everything else is great, but they’re in another cloud; they’re not going to say, “No because you’re operating on another cloud.” It still does what they want, but they still need to be able to have a common base of expertise for their app builders, and so on. So, a number of things dictate why people started looking at cross-cloud databases with common performance and operational characteristics and security characteristics, but don’t compromise on the feature set, right?

That’s starting to become super important, from our perspective. I think what’s most important is the ability to run the database with ease while not compromising on your developer agility or the ability to build your application. That’s the most important thing.

Corey: When you founded the company back in 2016, you are VC-backed, so I imagine your investor pitch meetings must have been something a little bit surreal. They ask hard questions such as, “Why do you think that in 2016, starting a company to go and sell databases to people is a viable business model?” At which point you obviously corrected them and said, “Oh, you misunderstand. We’re building an open-source database. We’re not charging for it; we’re giving it away.”

And they apparently said, “Oh, that’s more like it.” And then invested, as of the time of this recording, over $100 million in your company. Let me to be the first to say there are aspects of money that I don’t fully understand and this is one of those. But what is the plan here? How do you wind up building a business case around effectively giving something away for free?

And I want to be clear here, Yugabyte is open-source, and I don’t have an asterisk next to that. It is not one of those ‘source available’ licenses, or ‘anyone can do anything they want with it except Amazon’ or ‘you’re not allowed to host it and offer it as a paid service to other people.’ So, how do you have a business, I guess is really my question here?

Karthik: You’re right, Corey. We’re 100% open-source under Apache 2.0—I mean the database. So, our theory on day one—I mean, of course, this was a hard question and people did ask us this, and then I’ll take you guys back to 2016. It was unclear, even as of 2016, if open-source companies were going to succeed. It was just unclear.

And people were like, “Hey, look at Snowflake; it’s a completely managed service. They’re not open-source; they’re doing a great job. Do you really need open-source to succeed?” There were a lot of such questions. And every company, every project, every space has to follow its own path, just applying learnings.

Like for example, Red Hat was open-source and that really succeeded, but there’s a number of others that may or may not have succeeded. So, our plan back then was to tread the waters carefully in the sense we really had to make sure open-source was the business model we wanted to go for. So, under the advisement from our VCs, we said we’d take it slowly; we want to open-source on day one. We’ve talked to a number of our users and customers and make sure that is indeed the path we’ve wanted to go. The conversations pretty clearly told us people wanted an open database that was very easy for them to understand because if they are trusting their crown jewels, their most critical data, their systems of record—this is what the business depends on—into a database, they sure as hell want to have some control over it and some transparency as to what goes on, what’s planned, what’s on the roadmap. “Look, if you don’t have time, I will hire my people to go build for it.” They want it to be able to invest in the database.

So, open-source was absolutely non-negotiable for us. We tried the traditional technique for a couple of years of keeping a small portion of the features of the database itself closed, so it’s what you’d call ‘open core.’ But on day one, we were pretty clear that the world was headed towards DBaaS—Database as a Service—and make it really easy to consume.

Corey: At least the bad patterns as well, like, “Oh, if you want security, that’s a paid feature.”

Karthik: Exactly.

Corey: No. That is not optional. And the list then of what you can wind up adding as paid versus not gets murky, and you’re effectively fighting your community when they try and merge some of those features in and it just turns into a mess.

Karthik: Exactly. So, it did for us for a couple of years, and then we said, “Look, we’re not doing this nonsense. We’re just going to make everything open and just make it simple.” Because our promise to the users was, we’re building everything that looks like Postgres, so it’s as valuable as Postgres, and it’ll work in the cloud. And people said, “Look, Postgres is completely open and you guys are keeping a few features not open. What gives?”

And so after that, we had to concede the point and just do that. But one of the other founding pieces of a company, the business side, was that DBaaS and ability to consume the database is actually far more critical than whether the database itself is open-source or not. I would compare this to, for example, MySQL and Postgres being completely open-source, but you know, Amazon’s Aurora being actually a big business, and similarly, it happens all over the place. So, it is really the ability to consume and run business-critical workloads that seem to be more important for our customers and enterprises that paid us. So, the day-one thesis was, look, the world is headed towards DBaaS.

We saw that already happen with inside Facebook; everybody was automated operations, simplified operations, and so on. But the reality is, we’re a startup, we’re a new database, no one’s going to trust everything to us: the database, the operations, the data, “Hey, why don’t we put it on this tiny company. And oh, it’s just my most business-critical data, so what could go wrong?” So, we said we’re going to build a version of our DBaaS that is in software. So, we call this Yugabyte Platform, and it actually understands public clouds: it can spin up machines, it can completely orchestrate software installs, rolling upgrades, turnkey encryption, alerting, the whole nine yards.

That’s a completely different offering from the database. It’s not the database, it’s just on top of the database and helps you run your own private cloud. So, effectively if you install it on your Amazon account or your Google account, it will convert it into what looks like a DynamoDB, or a Spanner, or what have you with you, with Yugabyte as DB as the database inside. So, that is our commercial product; that’s source available and that’s what we charge for. The database itself, completely open.

Again, the other piece of the thinking is, if we ever charge too much, our customers have the option to say, “Look, I don’t want your DBaaS thing; I’m going to the open-source database and we’re fine with that.” So, we really want to charge for value. And obviously, we have a completely managed version of our database as well. So, we reuse this platform for our managed version, so you can kind of think of it as portability, not just of the database but also of the control plane, the DBaaS plane.

They can run it themselves, we can run it for them, they could take it to a different cloud, so on and so forth.

Corey: I like that monetization model a lot better than a couple of others. I mean, let’s be clear here, you’ve spent a lot of time developing some of these concepts for the industry when you were at Facebook. And because at Facebook, the other monetization models are kind of terrifying, like, “Okay. We’re going to just monetize the data you store in the open-source database,” is terrifying. Only slightly less would be the Google approach of, “Ah, every time you wind up running a SQL query, we’re going to insert ads.”

So, I like the model of being able to offer features that only folks who already have expensive problems with money to burn on those problems to solve them will gravitate towards. You’re not disadvantaging the community or the small startup who wants it but can’t afford it. I like that model.

Karthik: Actually, the funny thing is, we are seeing a lot of startups also consume our product a lot. And the reason is because we only charge for the value we bring. Typically the problems that a startup faces are actually much simpler than the complex requirements of an enterprise at scale. They are different. So, the value is also proportional to what they want and how much they want to consume, and that takes care of
itself.

So, for us, we see that startups, equally so as enterprises, have only limited amount of bandwidth. They don’t really want to spend time on operationalizing the database, especially if they have an out to say, “Look, tomorrow, this gets expensive; I can actually put in the time and money to move out and go run this myself. Why don’t I just get started because the budget seems fine, and I couldn’t have done it better myself anyway because I’d have to put people on it and that’s more expensive at this point.” So, it doesn’t change the fundamentals of the model; I just want to point out, both sides are actually gravitating to this model.

Corey: This episode is sponsored in part by our friends at Jellyfish. So, you’re sitting in front of your office chair, bleary eyed, parked in front of a powerpoint and—oh my sweet feathery Jesus its the night before the board meeting, because of course it is! As you slot that crappy screenshot of traffic light colored excel tables into your deck, or sift through endless spreadsheets looking for just the right data set, have you ever wondered, why is it that sales and marketing get all this shiny, awesome analytics and inside tools? Whereas, engineering basically gets left with the dregs. Well, the founders of Jellyfish certainly did. That’s why they created the Jellyfish Engineering Management Platform, but don’t you dare call it JEMP! Designed to make it simple to analyze your engineering organization, Jellyfish ingests signals from your tech stack. Including JIRA, Git, and collaborative tools. Yes, depressing to think of those things as your tech stack but this is 2021. They use that to create a model that accurately reflects just how the breakdown of engineering work aligns with your wider business objectives. In other words, it translates from code into spreadsheet. When you have to explain what you’re doing from an engineering perspective to people whose primary IDE is Microsoft Powerpoint, consider Jellyfish. Thats Jellyfish.co and tell them Corey sent you! Watch for the wince, thats my favorite part.

Corey: A number of different surveys have come out that say overwhelmingly companies prefer open-source databases, and this is waved around as a banner of victory by a lot of—well, let’s be honest—open-source database companies. I posit that is in fact crap and also bad data because what the open-source purists—of which I admit, I used to be one, and now I solve business problems instead—believe that people are talking about freedom, and choice, and the rest. In practice, in my experience, what people are really distilling that down to is they don’t want a commercial database. And it’s not even about they’re not willing to pay money for it, but they don’t want to have a per-core licensing challenge, or even having to track licensing of where it is installed and how, and wind up having to cut checks for folks. For example, I’m going to dunk on someone because why not?

Azure for a while has had this campaign that it is five times cheaper to run some Microsoft SQL workloads in Azure than it is on AWS as if this was some magic engineering feat of strength or something. It’s absolutely not, it’s that it is really expensive licensing-wise to run it on things that aren’t Azure. And that doesn’t make customers feel good. That’s the thing they want to get away from, and what open-source license it is, and in many cases, until the source-available stuff starts trending towards, “Oh, you’re going to pay us or you’re not going to run it at all,” that scares the living hell out of people, then they don’t actually care about it being open. So, at the risk of alienating, I’m sure, some of the more vocal parts of your constituency, where do you fall on that?

Karthik: We are completely open, but for a few reasons right? Like, multiple different reasons. The debate of whether it purely is open or is completely permissible, to me, I tend to think a little more where people care about the openness more so than just the ability to consume at will without worrying about the license, but for a few different reasons, and it depends on which segment of the market you look at. If you’re talking about small and medium businesses and startups, you’re absolutely right; it doesn’t matter. But if you’re looking at larger companies, they actually care that, like for example, if they want a feature, they are able to control their destiny because you don’t want to be half-wedded to a database that cannot solve everything, especially when the time pressure comes or you need to do something.

So, you want to be able to control or to influence the roadmap of the project. You want to know how the product is built—the good and the bad—you want a lot of people testing the product and their feedback to come out in the open, so you at least know what’s wrong. Many times people often feel like, “Hey, my product doesn’t work in these areas,” is actually a bad thing. It’s actually a good thing because at least those people won’t try it and [laugh] they’ll be safe. Customer satisfaction is more important than just the apparent whatever it is that you want to project about the product.

At least that’s what I’ve learned in all these years working with databases. But there’s a number of reasons why open-source is actually good. There’s also a very subtle reason that people may not understand which is that legal teams—engineering teams that want to build products don’t want to get caught up in a legal review that takes many months to really make sure, look, this may be a unique version of a license, but it’s not a license the legal team as seen before, and there’s going to be a back and forth for many months, and it’s just going to derail their product and their timelines, not because the database didn’t do its job or because the team wasn’t ready, but because the company doesn’t know what the risk it’ll face in the future is. There’s a number of these aspects where open-source starts to matter for real. I’m not a purist, I would say.

I’m a pragmatist, and I have always been, but I would say that a number of reasons why–you know, I might be sounding like a purist, but a number of reasons why a true open-source is actually useful, right? And at the end of the day, if we have already established, at least at Yugabyte, we’re pretty clear about that, the value is in the consumption and is not in the tech if we’re pretty clear about that. Because if you want to run a tier-two workload or a hobbyist app at home, would you want to pay for a database? Probably not. I just want to do something for a while and then shut it down and go do my thing. I don’t care if the database is commercial or open-source. In that case, being open-source doesn’t really take away. But if you’re a large company betting, it does take away. So.

Corey: Oh, it goes beyond that because it’s not even, in the large company story, whether it costs money because regardless, I assure you, open-source is not free; the most expensive thing that we see in all of our customer accounts—again, our consultancy fixes AWS bills, an expensive problem that hits everyone—the environment in AWS is always less expensive than the people who are working on the environment. Payroll is an expense that dwarfs the AWS bill for anyone that is not a tiny startup that is still not paying a market-rate salary to its founders. It doesn’t work that way. And the idea, for those folks is, not about the money, it’s about the predictability. And if there’s a 5x price hike from their database manager that suddenly completely disrupts their unit economic model, and they’re in trouble. That’s the value of open-source in that it can go anywhere. It’s a form of not being locked into any vendor where it’s hosted, as well as, now, no one company that has put it out there into the world.

Karthik: Yeah, and the source-available license, we considered that also. The reason to vote against that was you can get into scenarios where the company gets competitive with his open-source site where the open-source wants a couple other features to really make it work for their own use case, like you know, case in point is the startup, but the company wants to hold those features for the commercial side, and now the startup has that 5x price jump anyway. So, at this point, it comes to a head-on where the company—the startup—is being charged not for value, but because of the monetization model or the business model. So, we said, “You know what? The best way to do this is to truly compete against open-source. If someone wants to operationalize the database, great. But we’ve already done it for you.” If you think that you can operationalize it at a lower cost than what we’ve done, great. That’s fine.

Corey: I have to ask, there has to have been a question somewhere along the way, during the investment process of, what if AWS moves into your market? And I can already say part of the problem with that line of reasoning is, okay, let’s assume that AWS turns Yugabyte into a managed database offering. First, they’re not going to be able to articulate for crap why you should use that over anything else because they tend to mumble when it comes time to explain what it is that they do. But it has to be perceived as a competitive threat. How do you think about that?

Karthik: Yeah, this absolutely came up quite a bit. And like I said, in 2016, this wasn’t news back then; this is something that was happening in the world already. So, I’ll give you a couple of different points of view on this. The reason why AWS got so successful in building a cloud is not because they wanted to get into the database space; they simply wanted their cloud to be super successful and required value-added services like these databases. Now, every time a new technology shift happens, it gives some set of people an unfair advantage.

In this case, database vendors probably didn’t recognize how important the cloud was and how important it was to build a first-class experience on the cloud on day one, as the cloud came up because it wasn’t proven, and they had twenty other things to do, and it’s rightfully so. Now, AWS comes up, and they’re trying to prove a point that the cloud is really useful and absolutely valuable for their customers, and so they start putting value-added services, and now suddenly you’re in this open-source battle. At least that’s how I would view that it kind of developed. With Yugabyte, obviously, the cloud’s already here; we know on day one, so we’re kind of putting out our managed service so we’ll be as good as AWS or better. The database has its value, but the managed service has its own value, and so we’d want to make sure we provide at least as much value as AWS, but on any cloud, anywhere.

So, that’s the other part. And we also talked about the mobility of the DBaaS itself, the moving it to your private account and running the same thing, as well as for public. So, these are some of the things that we have built that we believe makes us super valuable.

Corey: It’s a better approach than a lot of your predecessor companies who decided, “Oh, well, we built the thing; obviously, we’re going to be the best at running it. The end.” Because they dramatically sold AWS’s operational excellence short. And it turns out, they’re very good at running things at scale. So, that’s a challenging thing to beat them on.

And even if you’re able to, it’s hard to differentiate among the differences because at that caliber of operational rigor, it’s one of those, you can only tell in the very niche cases; it’s a hard thing to differentiate on. I like your approach a lot better. Before we go, I have one last question for you, and normally, it’s one of those positive uplifting ones of what workloads are best for Yugabyte, but I think that’s boring; let’s be more cynical and negative. What workloads would run like absolute crap on YugabyteDB?

Karthik: [laugh]. Okay, we do have a thing for this because we don’t want to take on workloads and, you know, everybody have a bad experience around. So, we’re a transactional database built for user-facing applications, real-time, and so on, right? We’re not good at warehousing and analytic workloads. So, for example, if you were using a Snowflake or a Redshift, those workloads are not going to work very well on top of Yugabyte.

Now, we do work with other external systems like Spark, and Presto, which are real-time analytic systems, but they translate the queries that the end-user have into a more operational type of query pattern. However, if you’re using it straight-up for analytics, we’re not a good bet. Similarly, there’s cases where people want very high number of IOPS by reusing a cache or even a persistent cache. Amazon just came out with a [number of 00:31:04] persistent cache that does very high throughput and low-latency serving. We’re not good at that.

We can do reasonably low-latency serving and reasonably high IOPS at scale, but we’re not the use case where you want to hit that same lookup over and over and over, millions of times in a second; that’s not the use case for us. The third thing I’d say is, we’re a system of record, so people care about the data they put, and they don’t absolutely don’t want to lose it and they want to show that it’s transactional. So, if there’s a workload where there’s a lot of data and you’re okay if you want to lose, and it’s just some sensor data, and your reasoning is like, “Okay, if I lose a few data points, it’s fine.” I mean, you could still use us, but at that point you’d really have to be a fanboy or something for Yugabyte. I mean, there’s other databases that probably do it better.

Corey: Yeah, that’s the problem is whenever someone says, “Oh, yeah. Database”—or any tool that they’ve built—“Like, this is great.” “What workloads is it not a fit for?” And their answer is, “Oh, nothing. It’s perfect for everything.”

Yeah, I want to believe you, but my inner bullshit sense is tingling on that one because nothing’s fit for all purposes; it doesn’t work that way. Honestly, this is going to be, I guess, heresy in the engineering world, but even computers aren’t always the right answer for things. Who knew?

Karthik: As a founder, I struggled with this answer a lot, initially. I think the problem is, when you’re thinking about a problem space, that’s all you’re thinking about, you don’t know what other problem spaces exist, and when you are asked the question, “What workloads is it a fit for?” At least I used to say, initially, “Everything,” because I’m only thinking about that problem space as the world, and it’s fit for everything in that problem space, except I don’t know how to articulate the problem space—

Corey: Right—

Karthik: —[crosstalk 00:32:33]. [laugh].

Corey: —and at some point, too, you get so locked into one particular way of thinking that the world that people ask about other cases like, “Oh, that wouldn’t count.” And then your follow-up question is, “Wait, what’s a bank?” And it becomes a different story. It’s, how do you wind up reasoning about these things? I want to thank you for taking all the time you have today to speak with me. If people want to learn more about Yugabyte—either the company or the DB—how can they do that?

Karthik: Yeah, thank you as well for having me. I think to learn about Yugabyte, just come join our community Slack channel. There’s a lot of people; there’s, like, over 3000 people. They’re all talking interesting questions. There’s a lot of interesting chatter on there, so that’s one way.

We have an industry-wide event, it’s called the Distributed SQL Summit. It's coming up September 22nd, 23rd, I think a couple of days; it’s a two-day event. That would be a great place to actually learn from practitioners, and people building applications, and people in the general space and its adjacencies. And it’s not necessarily just about Yugabyte; it’s generally about distributed SQL databases, in general, hence it’s called the Distributed SQL Summit. And then you can ask us on Twitter or any of the usual social channels as well. So, we love interaction, so we are pretty open and transparent company. We love to talk to you guys.

Corey: Well, thank you so much for taking the time to speak with me. Well, of course, throw links to that into the [show notes 00:33:43]. Thank you again.

Karthik: Awesome. Thanks a lot for having me. It was really fun. Thank you.

Corey: Likewise. Karthik Ranganathan, CTO, and co-founder of YugabyteDB. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry comment, halfway through realizing that I’m
not charging you anything for this podcast and converting the angry comment into a term sheet for $100 million investment.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Laurie

Laurie is a Senior Software Engineer at Netflix. You can also find her creating content and educating the technology industry as an egghead instructor, member of the TC39 Educators committee, and technical blogger.

Links:

  • Twitter: https://twitter.com/laurieontech
  • Netflix: https://www.netflix.com
  • Egghead: https://egghead.io
  • The Art of the Subtle Subtweet: https://laurieontech.com/book-launch/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: You could build you go ahead and build your own coding and mapping notification system, but it takes time, and it sucks! Alternately, consider Courier, who is sponsoring this episode. They make it easy. You can call a single send API for all of your notifications and channels. You can control the complexity around routing, retries, and deliverability and simplify your notification sequences with automation rules. Visit courier.com today and get started for free. If you wind up talking to them, tell them I sent you and watch them wince—because everyone does when you bring up my name. Thats the glorious part of being me. Once again, you could build your own notification system but why on god’s flat earth would you do that?

Corey: This episode is sponsored in part by our friends at VMware. Let’s be honest—the past year has been far from easy. Due to, well, everything. It caused us to rush cloud migrations and digital transformation, which of course means long hours refactoring your apps, surprises on your cloud bill, misconfigurations and headache for everyone trying manage disparate and fractured cloud environments. VMware has an answer for this. With VMware multi-cloud solutions, organizations have the choice, speed, and control to migrate and optimize

applications seamlessly without recoding, take the fastest path to modern infrastructure, and operate consistently across the data center, the edge, and any cloud. I urge to take a look at vmware.com/go/multicloud. You know my opinions on multi cloud by now, but there's a lot of stuff in here that works on any cloud. But don’t take it from me thats: VMware.com/go/multicloud and my thanks to them again for sponsoring my ridiculous nonsense.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’m joined this week by Laurie Barth, but no one really knows that’s her last name. In fact, @laurieontech is how most people think of her. She’s a senior software engineer at a company called Netflix, which primarily streams movies and gives conference talks—in the before times—about how you’re doing it wrong.

She also creates a lot of content and educates the technology industry as an instructor at Egghead. She’s a member of the TC39 Educator’s Committee, and of course, is a technical blogger. Laurie, thank you for suffering the slings and arrows I’m no doubt going to be hurtling your way.

Laurie: This is the most fun I’ve had all week. [laugh].

Corey: Well, it’s a pandemic on, so presumably that isn’t that high of a bar for the pony to stumble over.

Laurie: Yeah, unfortunately not. I think that’s maybe the problem.

Corey: So, you’re someone that I have been aware of for an awfully long time. You’re always sort of omnipresent in conversations. You are someone who has a lot of great opinions that present well; you talk about an awful lot of things that are germane to my interests, educating the next generation of engineers, for example. And of course, you recently started at Netflix, at which point, well, if you’re not familiar with what Netflix is doing in the cloud, have you ever even talked to an AWS employee for more than 35 seconds because they’ll go reference Netflix for a variety of wonderful reasons, both based on technical excellence, as well as because AWS is so bad at telling the story of what you can build out of their popsicle stick service collection that they just punt to companies like Netflix to demonstrate what you could do. So, you’re sort of this omnipresent force on Twitter, but we’ve never really had a conversation before, so it was long past time to rectify this.

Laurie: I mean, you sent me two cents. So… I think that was pretty—[laugh].

Corey: That’s what the Tip Jar is for. You just wind up hurling very small amounts of money at people along with insulting comments, and it’s a new form of social media. That is the micro-transaction way.

Laurie: I quite enjoyed that. So, for context, I was one of the first people to be part of the A/B testing for Tip Jar on Twitter and Corey was the first person to send me money with, of course, a very on-brand Corey message, which there’s a screenshot of on Twitter somewhere. And a couple of people followed, but it was great fun. And I think that’s the first time we had ever directly interacted in a message or something, other than obviously, in threads and that sort of thing.

Corey: Yeah, that’s an interesting point to lead into here because I’m also in the A/B test for Tip Jar and I’ve largely turned it off, except for when I’m doing something very small and very focused, usually aimed at some sort of charitable benefit or whatnot, and even then, it’s not the right way to do it. And it’s weird, there was a time I absolutely would have turned it on, but it doesn’t seem right for me to do it now and that’s partially due to the fact that—first, I don’t need tips from the audience in order to sustain myself. I’m not that kind of creator. I have a company that solves very expensive problems for large companies and that works out really well for, you know, keeping the lights on here.

I’m not trying to disparage creators in any way, folks who are in a position of needing that to cover their lifestyle a variety of different ways. And even if they’re well beyond that, I don’t begrudge that to them at all. I mean, from a very selfish capitalist perspective, I don’t want you to feel that you’ve paid your debt to me for entertaining you by sending me $5. I want you to repay that debt by signing a five-figure consulting agreement.

Laurie: Yeah, those aren’t really the same thing, are they?

Corey: No, no. Turns out signing authority caps out at different places for different folks.

Laurie: [laugh].

Corey: Who knew? But it was a fun experiment. I’m glad that they’re doing it. I’m glad to see Twitter coming out of its stasis for a long time and trying new things, even if we don’t like some of them.

Laurie: Well, they have this whole Super Follows thing now, and I got waitlisted for it the other day because they said they accepted too many people, whatever that means. I think—

Corey: Same here.

Laurie: Yeah, I think a bunch of us got that. And I’m interested, my sense is it's sort of like a Patreon hosted in Twitter sort of thing. And I’ve never had a Patreon; I have a mailing list that I made based on an April Fool’s joke this past year where I made an entire signup workflow for the pre-order of my new book, The Art of the Subtle Subtweet. I was very pleased with this joke.

This was, like, very elaborate: I had a whole website, I had a signup flow, and I now have a mailing list which I’ve done nothing with. So, I have all of these things, but that’s not really been my—there’s too many things to do as a content creator, and so I’ve sort of not explored most of those other avenues. And so, Super Follows, I was like, “This could be interesting. I could try doing it,” but, you know, alas, they don’t want me to. So, [laugh] I don’t know that it matters.

Corey: It’s an interesting problem, too, because at the start of the pandemic, I had a third of the Twitter followers that I do as of the time of this recording, which is something like 63,000. When I started what I do, five years ago, and I had just left a company which was highly regulated, so, “Don’t tweet,” was basically their social media policy, it was a, okay, I had something like 2000 followers at the time. I was—it had taken me seven years to get there, let’s be very clear here. And since then, my following has exploded, and yours has as well. You have, I think the last time we checked, was it something like 30,000 and change?

Laurie: Yeah, something like that.

Corey: And it changes the way that people interact with you. This is one of those things that there aren’t that many people that we can have this kind of honest conversation with because let’s be very clear here, for folks who have not established an audience like that it sounds absolutely like it’s either a humblebrag—which I’m not intending that to come across that way—or it’s one of those, “Wish I had those problems.” And in some ways, yeah, it’s a weird problem to have, and it’s also not a sympathetic problem to have, but something that has been very clear to me has been that the way that people perceive me and the way that they interact with me has shifted significantly as my Twitter notoriety has increased.

Laurie: Yeah.

Corey: I’m curious about how you have experienced that?

Laurie: Yeah, so I’m half your size and especially in the front-end universe, there’s plenty of people with between 100,000 to, you know, I think Dan Abramov is at, like, 400,000 at this point. Like—

Corey: Oh yeah, my Twitter following would explode if I either knew JavaScript or was funny. Either one would just absolutely kick me into the stratosphere, but we work with what we’ve got.

Laurie: I either don’t know JavaScript or I’m not funny or maybe both because apparently not. But yeah, there’s these huge, huge, huge, huge scales, and I’m sure by many people’s judgment, pretty, pretty large. But comparing to other people in my ecosystem, maybe not so much. And I didn’t understand it until I was living it. I actually had the opportunity to meet Emily Freeman at a conference in DC, probably… three years ago now, when I had less than a thousand followers. And I thought getting my first hundred was a big deal; I thought getting my first 500—and it is. Don’t get me wrong. Those things are very cool milestones. And I [crosstalk 00:07:18]—

Corey: I still celebrate the milestones, but I do it less publicly now.

Laurie: Yeah, exactly. And I had a whole conversation with her and she gave me some really, really helpful advice: sort of, don’t look at your follower count as it goes back and forth, five people, six people you’ll think people are unfollowing you; they’re probably not. It doesn’t matter. And recognize that the larger you get, the more careful you have to be, and try to keep me sane before I was ever there. And it’s all sort of come true.

There’s two things that have stuck out to me, I think, during the pandemic, especially. One is I can write the most nonsensical, silly tweet and people will like it because they think it says something insightful whether it does or it doesn’t. They’re projecting onto the tweet something funnier, or more relevant than the reason I wrote it in the first place. Which, okay, that’s cool. I’m not as smart as you’re giving me credit for, but sure.

The other thing which is the downside to that is, everyone assumes that if they’re having a conversation with me, they’re having a conversation with me. So one-on-one, back and forth. That’s not untrue, but I’m having a similar conversation in parallel with—if it’s a popular tweet—a hundred other people at the same time. And what that means is, if you’re being a little bit of a jerk, and a little bit troll-y, you’re not being a little bit troll-y, you’re being a little bit troll-y times the a hundred other little bit troll-y people. And so my reaction to you is not going to be necessarily equivalent to what you say, and that can get me in trouble. But there’s no mental, emotional spectrum that was designed to work with the scale of social media.

Corey: Oh, absolutely not. In fact, let’s do an experiment now, while we’re having this conversation. I am making a tweet as we speak. “Some mornings, it’s just not worth chewing through the leather straps.” It’s not particularly insightful.

It’s not particularly deep, and before the end of this episode, we will check and see what that does in terms of engagement just because you can say anything, and there’s some folks who will wind up automatically engaging. And again, that’s fine; everyone engages with Twitter in a bunch of different ways. For me, what’s been very odd is I have talked to a couple of very large companies who I talk about on Twitter from time to time, and it turns out that they are reluctant to engage with me directly on Twitter or promote anything that I do or do retweets of me, not because of me, but because of an element of the audience, in some cases, of what people will chime in and say because it doesn’t align with corporate brands and a bunch of different perspectives. Which, again, I have some sympathy for this; it’s hard to deal with folks who are now suddenly given a soapbox and a platform that rewards clever insults better than it does meaningful heartfelt content, and that is something that I think everyone is still struggling with. Let’s also be very clear here. I’m a white dude in tech; my failure mode is a board seat and a book deal.

Laurie: [laugh].

Corey: When I post something about Git, for example—which I did a few days ago—and someone responds explaining the joke back to me, my response to them was, “Thank you for explaining Git to me.” And that was all I said, and it’s led to a mini-pile-on of this person because it’s like

Laurie: Oh, yeah.

Corey: “Don’t you know who Corey is?” Yet I have seen the same dynamic happen with women tweeting about these things and it’s not just one response that explains Git; it’s all of them. And when people say—like, Abby Fuller, for example—

Laurie: Yep.

Corey: —will tweet about password manager challenges and how annoying some of them are, and it leads to a cavalcade of people suggesting password managers to her. That is not why she’s tweeting it, and she explicitly says, “I do not want you to recommend password managers to me.” And people continue to do it. And I don’t for the life of me understand what goes on in some people’s heads.

Laurie: Yeah. I mean, I’ve watched that happen countless times. I think the frustration—there’s a point at which no matter how big of a following you have, you just want to be yourself. I think most people who get to that amount of interaction have been theirself most of the way, along the way. Or they’re just being totally fake for the sense of growth hacking, in which case, okay, you do you.

But most people, I think, are being themselves because it’s exhausting to spend that much time on a platform and pretend to be someone else or be fake the whole time. So, I’m pretty much myself. And that means that sometimes when someone’s being a total jerk, I really want to treat them and be like, “Yeah, you suck.” But the problem is when I say that, I’m siccing 30,000 other people on them to defend me. And I can’t do that.

So instead, I’ve become sort of famous for subtweeting. And I will wait a couple of days to do it, or I will totally change the framing of the situation so I can get out my same sort of frustration, and annoyance, and just needing to blow off steam, or venting, or whatever it is and not point at the person. Because if I point at the person, I discovered very, very quickly that there’s a whole crowd of people willing to take them down. If they’re being blatantly terrible, I will do it. There is a line here.

Someone recommending that I use a different tool because I decided to bitch about TypeScript, for example, or telling me I don’t understand TypeScript, okay, fine. Someone’s saying, “You only have followers because you’re a pretty girl.” Yeah, you’re an asshole. No, I’m not protecting you. Also, by the way, I tweeted two minutes ago, do all tweets deserve a ‘like,’ question mark, and we’ll see how much that—

Corey: Yeah.

Laurie: —interaction gets. [laugh].

Corey: I’m looking forward to seeing how that plays out. It’s a responsibility, which sounds odd, but if I complain about a company, what I’m fundamentally doing is I have the potential to be calling out an airstrike on top of them. And not every customer service failure deserves that. I deleted all of my tweets prior to 2015 a while back. And the reason most people delete tweets, or the reason we hear about most people deleting tweets, there was nothing especially problematic in my tweets other than jokes that were mean in different ways and punching down in ways that I didn’t realize were at the time.

It was not full of slurs; it was just things that weren’t particularly great. But that wasn’t the real reason I did it. The honest reason was is that I looked at my early tweets and they were cringy beyond belief. I was shilling for the company I worked for in many respects, and there were swaths which I didn’t engage with Twitter, and the only time I really did is I was out there complaining about various customer service failures, so it’s just this neverending stream of complaints about different companies that had wronged me in trivial ways.

Laurie: [laugh].

Corey: And, I don’t know at some point if somebody is going to build something where it’s easy to explore early tweets of a particular account. I don’t want them to do that and then figure out that this is how you get started being me. It’s like, I succeeded in spite of that nonsense, not because of it. And it’s not something good that I want to put out into the world.

Laurie: Yeah. So, I have, I think, only once added a company when I was having a customer service issue on a weekend, and we were in really dire straits. And I was just like, “Okay, it’s a weekend. I’m going to at.” And I’ve never gotten a response so fast.

And my husband looked at me and he was like, “Wait, what?” And I’d done this with an ol—I have this really ancient Twitter account that I got rid of because I was mostly just screaming about politics [laugh] and I didn’t want—I think I got @laurieontech in like, 2016, 2017—and I’d done that before. I’d been like, “Hey, you know”—I’m making something up—“At Spirit Airlines”—they seem like an easy one to—I’ve never flown Spirit, so—but I mean, I never got a response. And so there—realizing that you have power from a brand perspective is really weird.

But I almost want to go back to your point when you were talking about when you worked for a company and you had your account and, you know, they don’t want you to tweet, basically. Or companies are not going to tweet at you now, in your current state. I think it’s really hard to be a company on the internet in tech because you’re either going to make a joke that lands well, or everyone’s going to think that you’re shilling for yourself. There’s no in-between and so—this is a hot take and I might get in trouble for that—companies have realized that the best way to get around that is to hire people who have their own personal names and get your company name associated with them. And all of a sudden, it looks less disingenuous.

Corey: And even that’s a problem because I’ve talked to companies who are hiring folks with large followings for DevRel style jobs, and—I’ve interviewed for a few of those, once upon a time, about midway through when I was debating do I shut this consulting thing down and get a real job again because that’s always how I sort of assumed it would be for the first couple years. And then, “No, I’m going to get serious about it.” And I took on a business partner and got very serious, and here we are. But talking to folks, my question was, in the interview process, I would talk to my prospective manager and ask questions of the form, “So, what is your plan for when we eventually part ways? How are you
structuring that?”

And they looked at me like that was a bizarre question. It’s, understand that, done right, my personal brand will, in some areas and some corners, eclipse that of the company, so as soon as I leave for whatever reason, the question is going to be, “Were you mistreated? Did someone wrong you there? We’ll drag them just preemptively on the off chance.” And you need to have a plan in place to mitigate some of that and have a structured exit for what that is going to look like. And they looked at me like I was coming from a different planet. But I still think I’m right.

Laurie: You are right. And, oh goodness, I’ve seen this in a lot of different places. I mean, I have left companies in the past and I have had to decide how I was going to position that publicly. And how much I was going to say or not say, how complimentary I was going to be or not because the thing is, when you leave a place, you’re not just leaving the company, you’re also leaving your colleagues. And what does that mean for their experience?

You’re gone. You don’t want to be saying, “Hey, this place is horrible, while your really close friends you were working with on Friday are still there.” At the same time, companies don’t think about this from the DevRel perspective and, I want to be very clear, I have friends who work in DevRel who are themselves brands. They are all fantastic people; they work incredibly hard; this is not a knock on them in any way—

Corey: It looks easy from the outside. I want to be very clear on that.

Laurie: [laugh]. It’s not easy. All this stuff is great, but part of the reason I decided to go to a place like Netflix is because I knew my brand had no bearing on them and so I could be myself and just do my own thing and they weren’t going to try and leverage me, or there was no hit to them based on who I was. Granted, did I go after someone the other day, sort of, in deep in a thread for being a jerk and did they try and at Netflix engineering and say, “Is this the kind of person you want representing your brand?” And at egghead.io, “Is this the kind of person wanting your brand?” Yeah, they did.

So, that part’s still a problem, but that’s a problem for me rather than being a problem for my company, if I decide that, you know, I don’t always want to—like, no one cares if I talk about the new Marvel show. No one cares. I like Marvel; I’m allowed to like Marvel. I also love the stuff on Netflix, right, but when you’re at a company that isn’t like that, honestly, when I was at Gatsby, I couldn’t be tweeting about Next or Nuxt, or even Vue for that matter, because it just doesn’t look right. Because my brand had more of an impact in that smaller pond than it does now.

Corey: People have said, “Oh, well, what if AWS acquires you so you can work on their behalf?” Or, “What if Google acquires you?” Or something like that, and it’s—what people don’t get is that my persona—again, to be clear, I am genuine on Twitter. I emphasize aspects of my personality, but I don’t get up there and say things I don’t necessarily believe. We’ll get back to that in a minute.

But what I do as a small company, making fun of trillion-dollar publicly traded entities is funny and it works, but if suddenly I work at a different publicly-traded company, it just looks like I work for my employer, bagging on a competitor. And even if I’m speaking in ‘an opinions my own’ sense, which is apparently Amazon’s corporate motto, based on how often I see it in their employee’s Twitter bios—

Laurie: Oh, yeah. [laugh].

Corey: —is going to be perceived as me smacking at a competitor regardless. Further, I will not be the person that craps on my own employer on Twitter because that sends terrible signal in many respects. I won’t even crap on previous employers who frankly kind of deserve it because when you do that, it does not look good to people who are not familiar with the situation, and no one’s as familiar with it as you are. It just looks like sour grapes, regardless of how legitimate your grievance was. To be very clear, I’m not saying don’t call out abuse when you encounter it—

Laurie: Yeah.

Corey: —that’s fine. I’m not going down that path—

Laurie: Yeah, yeah, yeah, yeah.

Corey: —let’s clear here. But, “Yeah, they have a terrible management culture, and they don’t promote internally, and I hate those people,” it just makes you look bad, and it doesn’t help anything.

Laurie: Yeah. I had always made a commitment to never talk about a former employer in any way that was easily identifiable. I’ve changed that policy a little bit. There’s a story I shared a couple of times where my CEO didn’t want to give me a pay raise because he thought it was my parents’ and boyfriend at the time’s job to take care of me financially. Like, that kind of stuff, I will say publicly.

No one’s going to know who it is; you’d have to go back and figure it out and, like, you don’t have enough context so how would you know? But it’s stuff like that, that I’m like, okay. I don’t want to hide stories like that because that’s not protecting anybody.

Corey: No, I’m not talking about covering up for misbehavior. I’m talking run-of-the-mill just bad management, poor company culture, terrible technical decisions, et cetera. Yeah, if it’s like, yeah, they sexually harassed every woman on the team, out. Yeah, tell that story. I—thank you, I should absolutely clarify my stance. Heaven forbid I get letters.

Laurie: But yeah, it’s the problem is that you can’t—and everyone has a slightly different experience with this, but from what I’ve seen, it doesn’t matter if you say their management is shitty and they didn’t promote versus there was a ton of sexual harassment. If you’re one person saying it—if it’s the Blizzard situation where there’s tons of receipts and it’s made it into national media, then that’s a little bit different. But if you’re one person saying it about one company, people are going to think it’s sour grapes. And unfortunately, it doesn’t reflect on the company; it reflects on you. So, unless there’s a sort of like, where there’s smoke, there’s fire situation where a bunch of people are doing it at once, you have to weigh stuff really carefully.

Especially because your next employer doesn’t want you out there talking about your previous employer because then their fear is what are you going to say about them when you leave? There’s lots of nuance and it gets—if you are screaming into the void—we’re screaming into the cloud here—

Corey: Ahhhh. Yes.

Laurie: Ahhhh. [laugh]. If you’re screaming into the void, it doesn’t matter if you’re you. And I mean… [sigh] I hate saying, “If you’re me,” right?
That’s such an obnoxious statement to make, but at 30,000, they probably care.

Corey: There are inflection points. I started seeing—around 40,000 is when I started seeing a couple of brands reaching out to me to, “Hey, you want to promote some nonsense.” And I’ve never sold any social media promotion for anything. I sell sponsorships for newsletters, this podcast, I do webinars stuff, I do paid speaking engagements. My Twitter account is mine.

It is not the company’s and that is by design. It’s me; that’s what it comes down to. That does lead to challenges in some arenas because I talk to companies about their AWS bill and these companies do not have much of a sense of humor about spending tens of millions of dollars, in some cases a month, on a cloud provider. These are serious problems and they’re a little worried, in some cases, the first time we have conversations that they’re dealing with some kind of internet clown.

Laurie: [laugh].

Corey: And often with talking to folks to convince them to come on this podcast, it’s, “Look, this is not me dragging you and making you look awful because if I do that, I’ll never get another guest again.” And if I do it in the context of a consulting project it’s, “That was a hilarious entertaining intro here. Get out and never come back.” It is not useful. People have generally taken a risk personally on bringing the Duckbill Group in.

If we can’t deliver and cannot present professionally, then they have some serious damage control to do, for a variety of excellent reasons. And we’ve never put someone in that position and we won’t. I talked to brands who sponsor all of these things, and the ones that are the best sponsors intrinsically understand it, that [unintelligible 00:23:56] once I start getting after some serious maleficence style stuff—no one is going to not do business with you because I make fun of your company on Twitter—

Laurie: Yeah.

Corey: —but an awful lot of people are going to hear about you for the first time and advertising in the newsletter and having fun with that, or I talk about you in the podcast ads, it winds up being engaging in many cases depending how far I can stretch it. And it works. I did a tour at re:Invent last year—virtual re:Invent—where I led a Twitch tour for an hour around the virtual expo hall into a bunch of different sponsored virtual booths and made fun of them all, and I got thank you notes from the sponsors because that led to a bunch of leads because people cared about the—oh, people paying attention because Amazon did a crap job of advertising the Sponsor Expo. And it was something that people could grasp, and have fun with, and get attention for. It’s top-of-funnel work and that’s fine, but I just don’t do it with the boring stodgy stuff. I like to have fun with it. Bring a personality or don’t bother.

Laurie: Yeah. And you can’t take yourself too seriously. I’m not the stand-up comedian that you are. I like to fashion myself as a little bit funny but not that funny. I’m not a stand-up comedian and I don’t have a consultancy to represent anymore.

There was a time where I did; I was not the owner of it but I worked there. So, now it’s sort of, I represent me, which is good in the way that you say it. Like, it’s clearly you. It’s not Duckbill Group; it’s your account. But at the same time, it freaks me out when in real life people know that it’s me.

So, in my brain, Twitter is the internet and I have my actual real day-to-day life, and never the two shall cross. [laugh]. And my—one of my—I had this popular tweet where I talked about all the companies I’d been rejected from, and it turned into a bit of a retweet situation with everyone sharing all these companies that they’d been rejected from. And the screenshots made it onto LinkedIn and made it into my cousin’s feed, and she sent me a text message with a screenshot. And she’s like, “You’re on my LinkedIn.”

And I was like, “No, no, this is not okay. This is not”—I have my little circle of the world and it should not expand beyond that. I go to a conference, even a tech conference, and someone’s like, “Oh, you’re blue shirt, crossed arms.” I’m like, “No, this is not okay.” Like, [laugh] I only exist on the internet.

Corey: This episode is sponsored by our friends at Oracle HeatWave is a new high-performance accelerator for the Oracle MySQL Database Service. Although I insist on calling it “my squirrel.” While MySQL has long been the worlds most popular open source database, shifting from transacting to analytics required way too much overhead and, ya know, work. With HeatWave you can run your OLTP and OLAP, don’t ask me to ever say those acronyms again, workloads directly from your MySQL database and eliminate the time consuming data movement and integration work, while also performing 1100X faster than Amazon Aurora, and 2.5X faster than Amazon Redshift, at a third of the cost. My thanks again to Oracle Cloud for sponsoring this ridiculous nonsense.

Corey: My business partner was, a week or so ago, at a cafe and someone came by and saw his Last Week in AWS sticker on his laptop. It’s like, “Oh, you read that, too? I love Corey’s work.” Turns out the guy works at IBM Cloud. And yes, you should hear the air quotes around the word, ‘cloud’ in there. But still.

Laurie: [laugh].

Corey: It’s—I haven’t been out in the world since I really started focusing on this, and now it’s—like, I wear a mask so it’s fine, but I’m starting to wonder, am I going to get stopped on the street when I go back into the universe out there? And it’s weird because you can’t really unring that bell?

Laurie: No.

Corey: It’s a weird transition, and on some level, it’s constraining in some ways. Like, at some point of celebrity—I don’t know if I’m there yet or not—there’s going to become a day where I can’t just unload on a waiter for crappy service at a restaurant—not that that’s how I—

Laurie: I mean, you shouldn’t do that anyway. [laugh].

Corey: —operate anyway—without it potentially going viral, and, “Oh, he’s a jerk when you actually get to know him.” And everyone has this
idea of you and this impression of who you are, based upon the curated selection of what it is you put out into the world. I’ve tried to be as true to life as I can on this. In conversations, I generally don’t drop nothing but one-liners, but I think I’m pretty true to life as far as how I present on the internet versus how I present in person.

Laurie: More than I expected, to be honest.

Corey: Yeah. That also does surprise people. Like, they think there’s some sort of writing team behind me. And it’s, if you look at the timing of some of my tweets where I will respond with a witty, snarky thing in less than a minute, it’s, I wish I had a writing team with that kind of latency. I think that’d be terrific.

Laurie: I always assumed it was you, but I figured there was like a persona that you turn on and turn off and I realize now that it’s an always on sort of thing. [laugh].

Corey: One thing I did experiment with for a little bit was having my team write tweets for my approval to promote episodes of this podcast, for example, because I am not the sort of person going to sit there and build the thing out correctly and schedule at the right time. And I have people who can do things like that, but it’s the sort of thing that led to a situation of never getting much engagement and those tweets never did very well, so why even bother? We have a dedicated Twitter feed for that stuff and everyone’s happier. Especially since I don’t have to share access to this thing through anyone. Speaking of, let’s see her tweets did.

Laurie: Oh, yeah. Okay, hold on. How’d we do? All right. So, I have, “Do all tweets deserve a like?” Was posted 19 minutes ago. It has 12 comments, 1 retweet, and 22 likes.

Corey: My, “Some mornings, it’s just not worth chewing through the leather straps.” Was posted at a similar timeframe has 10 likes and 3 replies. Someone said that, “Organic, eh? Probably better than nylon.” Someone said, “Is this an NDA subtweet?” And someone said—with a GIF of Leonardo DiCaprio, saying, “You had my curiosity. Now, you have my attention.” That’s it. So yeah, not exactly a smash-it-out-of-the-park success.

Laurie: Yeah, but I got to say, “Do all tweets deserve a like?” Is pretty mundane. For that amount of response.

Corey: You included a question mark, which is an open invitation—

Laurie: Oh, right.

Corey: —to the internet randos to engage, so there is—

Laurie: Oh, yeah.

Corey: —a potential there.

Laurie: I going to have to retweet this and say that I’m not grifting and it was done for this podcast [laugh] and they should all listen to it. [laugh].

Corey: Oh, of course. By all means. I am thrilled in any point to wind up helping people learn more things about the environment.

Laurie: [laugh].

Corey: I want to thank you so much for taking the time to speak with me. I have to honestly say that I wasn’t quite sure what was coming, but of all the things you could have asked me to predict about this episode, not talking about how Netflix works in cloud was absolutely not one of them. So wow, are you sure you work at Netflix? That’s one of those odd moment things.

Laurie: Yeah, I got to say I’m pretty abstracted from the cloud these days, so that—maybe that means that I don’t know enough to talk about it intelligently.

Corey: I would argue that extends to lots of folks. To be clear, Netflix has a lot of really neat thing.

Laurie: That never stopped anyone before? Bu-dum-shh.

Corey: Oh, yeah. It’s like, I like to get up there, sometimes I’ll talk about how we do things at Netflix, periodically, on conference stages even though I’ve never worked there, but people don’t correct me because why not? I’m a white man in tech. And I say something, of course, it’s right. It’s just—if you don’t want them to get right, you just don’t have enough context. That’s the rule.

Laurie: Corey, I’m going to need you to take the last minute or so of this episode, and please explain your feelings on how to optimize your use of JavaScript on the front-end, please.

Corey: Oh, wonderful; you pay smart people who know what they’re doing to look deep into the JavaScript side of it—

Laurie: [laugh].

Corey: —because honestly, every time I’ve tried to get into JavaScript, I go back at it and I feel even more foolish than when I started. Async stuff just completely blows my mind, especially by default. How in God’s flat earth is that supposed to work? And—

Laurie: You work in cloud. [laugh].

Corey: It doesn’t make sense to me, in a clear sense. At least with Python, which is the—I would say it’s the language I know best, but it’s not. Crappy Python is. And I can at least do things top to bottom and it works about like I would expect unless explicitly instructed otherwise. But the JavaScript world is just a big question mark and doesn’t work the way that I would expect to. To be clear, the failure here is entirely mine.

Laurie: ‘JavaScript is a big question mark and doesn’t work the way I would expect it to’ should be JavaScript’s tagline.

Corey: That’s fair because I have this ridiculous belief from the Dark Ages—because I spent 20 years as a systems admin—that computer behavior should be deterministic and if there’s one thing that we learned about the internet, it’s not.

Laurie: Yeah, no. There’s that whole user thing, and then that whole browser thing, and then that whole device thing. It’s a whole bunch of non-deterministic behaviors. Just stick to the cloud, and there’s one consumer and one producer, and you’re good.

Corey: One thing I will say—in the moment of pure seriousness here—is that if I were looking at getting into tech today, the first language I would learn would be JavaScript. It is clearly the way of the future. It is a first-class citizen on every platform out there. It is the lingua franca of, effectively, everyone coming out of a boot camp. And it is going to be the way that computers are built.

I say this not from a position of being an advocate for JavaScript. I don’t know it; I can’t stand it personally, but it is clear as day to me that is the direction the world is moving in, so if you’re debating what language to pick up, you’d be hard-pressed to convince me not to recommend JavaScript as the first one.

Laurie: And do you want me to be my serious self, and you’re going to laugh at what I’m about to say?

Corey: Hit me with it.

Laurie: If you’re looking to get into technology because of boot camps and some other things, we have an oversaturation of newbie front-end developers and they’re all way more talented than I was at that point in my career, and yet there aren’t nearly the front-door opportunities for being a—I hate the term junior, but newbie. And where there is the opportunity, it’s cloud. And security.

Corey: I will absolutely point out further that I understand this runs the risk of being ‘boomer gives career advice’—

Laurie: Yeah, right? [laugh].

Corey: —but let’s be clear here. I think that if you are going to enter the front-end space—and this does speak to cloud and it speaks to security as well—distinguish slash differentiate yourself by having another discipline or area of intense interest that you can bring into it as well because when you have a company that’s looking to hire from a sea of new boot camp grads that generally tend to look more or less identical from a resume perspective, the one that will stand out is the one that can bring in another discipline and especially if that niche winds up aligning with a company’s business, or at least an intense interest in something that is directly germane to the company, that will distinguish you. And everyone has something like that; no one is one-dimensional. So, find the thing that is the in-between space, and focus on finding jobs in companies that do those things. And if you’re a mid-career switcher, let me be very clear here.

It is not a go back to entry-level roles-style story. I’ve never understood that philosophy. I do have steps from thing I’m doing now toward thing I want to go to. Well, is there a job I can find to do next that blends the two of them together in different ways, and then once I’m there, then make a further transition. And of course, find someone who’s—in any career, in any path you’re on, find someone who is five years ahead of you, and ask them for their advice.

“What would you do in my shoes?” If the answer is, “Go to a boot camp,” okay. Talk to a few people who’ve done this and make sure it validates it. If it’s, “Get a degree,” okay, but make sure you’re not doing it because you think that’s what you’re supposed to do. You’ll very rarely find me recommending six figures of debt in order to advance your career, but there are occasions.

By and large, they’ll find someone who’s been there before who knows what’s going on, you can have a conversation with and give them context appropriate to your situation and then see what’s right. We turned this into last-minute career advice and I’m not even—I don’t even [unintelligible 00:34:45] have a problem with that.

Laurie: Well, I was about to say that it’s 2020. 21 2020—wow, I—you knew what I meant—it’s 2021, and I guess I need to start taking my half-steps towards becoming a Lego master before I retire. [laugh].

Corey: Oh, yes, the Lego world is vast and deep, and they have gotten no worse since I was a child at separating parents from money to buy LEGO sets. My daughter’s four and his way into them already. So, it’s great. It’s something that we can bond over.

Laurie: If I ever have kids, we’re going to need separate sets because they’re not touching mine. [laugh].

Corey: Yeah, I’m looking at stuff like, oh, well, I’d love to buy that awesome big Star Destroyer—wait, it’s how much money? And it turns into this—yeah. It’s wow, on some level, I never ever thought I would find a hobby that was more expensive than my mechanical keyboards hobby, but here we are.

Laurie: Oh, yeah, I blame Cassidy Williams for getting me into that one, too. I have a shiny one beneath me. And that’s my first.

Corey: She is a treasure and a delight.

Laurie: She’s a treasure, a delight, and dangerous if you want to save money because she will draw you into the mechanical keyboards, and there’s just, there’s no resisting. I tried for a very long time. I failed, ultimately.

Corey: One of these days, she and I are going to have a keyboard-off at some point, once it’s no longer a deadly risk to do so. It’ll be fun.

Laurie: Do it.

Corey: I’m looking forward to it. Thank you so much for taking the time to speak with me. I really appreciate it.

Laurie: Absolutely. Thanks for having me.

Corey: Of course. Laurie Barth, senior software engineer at Netflix, also instructor at Egghead, also a member of the TC39 Educator Committee, and prolific blogger. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with a horrifying comment explaining anything we just talked about, back to us.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Ev

Ev Kontsevoy is Co-Founder and CEO of Teleport. An engineer by training, Kontsevoy launched Teleport in 2015 to provide other engineers solutions that allow them to quickly access and run any computing resource anywhere on the planet without having to worry about security and compliance issues. A serial entrepreneur, Ev was CEO and co-founder of Mailgun, which he successfully sold to Rackspace. Prior to Mailgun, Ev has had a variety of engineering roles. He holds a BS degree in Mathematics from Siberian Federal University, and has a passion for trains and vintage-film cameras.

Links:

  • Teleport: https://goteleport.com
  • Teleport GitHub: https://github.com/gravitational/teleport
  • Teleport Slack: https://goteleport.slack.com/join/shared_invite/zt-midnn9bn-AQKcq5NNDs9ojELKlgwJUA
  • Previous episode with Ev Kontsevoy: https://www.lastweekinaws.com/podcast/screaming-in-the-cloud/the-gravitational-pull-of-simplicity-with-ev-kontsevoy/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at VMware. Let’s be honest—the past year has been far from easy. Due to, well, everything. It caused us to rush cloud migrations and digital transformation, which of course means long hours refactoring your apps, surprises on your cloud bill, misconfigurations and headache for everyone trying manage disparate and fractured cloud environments. VMware has an answer for this. With VMware multi-cloud solutions, organizations have the choice, speed, and control to migrate and optimize

applications seamlessly without recoding, take the fastest path to modern infrastructure, and operate consistently across the data center, the edge, and any cloud. I urge to take a look at vmware.com/go/multicloud. You know my opinions on multi cloud by now, but there's a lot of stuff in here that works on any cloud. But don’t take it from me thats: vmware.com/go/multicloud and my thanks to them again for sponsoring my ridiculous nonsense.

Corey: You could build you go ahead and build your own coding and mapping notification system, but it takes time, and it sucks! Alternately, consider Courier, who is sponsoring this episode. They make it easy. You can call a single send API for all of your notifications and channels. You can control the complexity around routing, retries, and deliverability and simplify your notification sequences with automation rules. Visit courier.com today and get started for free. If you wind up talking to them, tell them I sent you and watch them wince—because everyone does when you bring up my name. Thats the glorious part of being me. Once again, you could build your own notification system but why on god’s flat earth would you do that?

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Roughly a year ago, I had a promoted guest episode featuring Ev Kontsevoy, the co-founder and CEO of Teleport.

A year has passed and what a year it’s been. Ev is back to tell us more about what they’ve been up to for the past year and, ideally, how things may have changed over in the security space. Ev, thank you for coming back to suffer the slings and arrows I will no doubt be hurling your way almost immediately.

Ev: Thanks for having me back, Corey.

Corey: So, it’s been a heck of a year. We were basically settling into the pandemic when last we recorded, and people’s security requirements when everyone is remote were dramatically changing. A year later, what’s changed? It seems like the frantic, grab a bucket and start bailing philosophy has largely been accepted with something that feels almost like a new normal, ish. What are you seeing?

Ev: Yes, we’re seeing exact same thing, that it’s really hard to tell what is normal. So, at the beginning of the pandemic, our company, Teleport was, so we were about 25 people. And then once we got the vaccines, and the government restrictions started to, kind of, disappear, people started to ask, “So, when are we going to go back to normal?” But the thing is, we’re 100 employees now, which means that three-quarters of the company, they joined us during the pandemic, so we have no normal to go back to. So, now we have to redefine—not redefined, we just basically need to get comfortable with this new, fully remote culture with fully remote identity that we have, and become comfortable with it. And that’s what we’re doing.

Corey: Beyond what, I guess, you’re seeing, as far as the culture goes, internally as well, it feels like there’s been a distinct shift in the past year or so, the entire security industry. I mean, I can sit here and talk about what I’ve seen, but again, I’m all over the place and I deal with a very select series of conversations. And I try not to confuse anecdotes with data. Anecdata is not the most reliable thing. You’re working in this space. That is the entire industry you’re in. How has the conversation in the industry around security shifted? What’s new? What trends are emerging?

Ev: So, there are several things actually happening. So, first of all, I wouldn’t call ourselves, like, we do all of security. So, we’re experts in access; like, how do you act this everything that you have in your cloud or in your data centers? And that space has been going through one transformation after another. It’s been basically under the same scaling stress as the rest of cloud computing industry.

And we can talk about historical changes that have been happening, and then we can talk a little bit about, kind of, latest and greatest. And in terms of what challenges companies have with secure access, maybe it helps if I just quickly describe what ‘access’ actually means.

Corey: Please, by all means. It’s one of those words that everyone knows, but if you ask three people to define it, you’ll get five definitions—

Ev: [laugh]. Exactly.

Corey: —and they don’t really align. So please, you’re the expert on this; I am here to listen because I guarantee you I am guilty of misusing the term at least once so far, today.

Ev: Can’t blame you. Can’t blame you. We are—I was same way until I got into this space. So, access basically means four things. So, if you want to have access done properly into your cloud resources, you need to think about four things.

First is connectivity. That’s basically a physical ability to deliver an encrypted packet from a client to destination, to a resource whatever that is, could be database, could be, like, SSH machine, or whatever it is you’re connecting to. So, connectivity is number one. So, then you need to authenticate. Authentication, that’s when the resource decides if you should have access or not, based on who you are, hopefully.

So, then authorization, that’s the third component. Authorization, the difference—like, sometimes people confuse the two—the difference between authentication and authorization is that authorization is when you already authenticated, but the resource decides what actions you are allowed to perform. The typical example is, like, is it read-only or read-write access? So, that’s authorization, deciding on which actions you’re allowed to perform. And the final component of having access properly is having audit or visibility which is, again, it could be real-time and historical.

So ideally, you need to have both. So, once you have those two solved, then you solved your access problem. And historically, if you look at how access has been done—so we had these giant machines, then we had microcomputers, then we had PCs, and they all have these things. So, you login into your Mac, and then if you try to delete certain file, you might get access denied. So, you see there is connectivity—in this case, it’s physical, a keyboard is physically connected to the [laugh] actual machine; so then you have authentication that you log in in the beginning; then authorization, if you can or cannot do certain things in your machine; and finally, your Mac keeps an audit log.

But then once the industry, we got the internet, we got all these clouds, so amount of these components that we’re now operating on, we have hundreds of thousands of servers, and load-balancers, and databases, and Kubernetes clusters, and dashboards, all of these things, all of them implement these four things: connectivity, authentication, authorization, audit.

Corey: Let me drive into that for a minute first, to make sure I’m clear on something. Connectivity makes sense. The network is the computer, et cetera. When you don’t have a network to something, it may as well not exist. I get that.

And the last one you mentioned, audit of a trail of who done it and who did what, when, that makes sense to me. But authentication and authorization are the two slippery ones in my mind that tend to converge a fair bit. Can you dive a little bit in delineate what the difference is between those two, please?

Ev: So authentication, if you try to authenticate into a database, database needs to check if you are on the list of people who should be allowed to access. That’s authentication, you need to prove that you are who you claim you are.

Corey: Do you have an account and credentials to get into that account?

Ev: Correct. And they’re good ways to do authentication and bad ways to do authentication. So, bad way to do authentication—and a lot of companies actually guilty of that—if you’re using shared credentials. Let’s say you have a user called ‘admin’ and that user has a password, and those are stored in some kind of stored—in, like 1Password, or something like Vault, some kind of encrypted Vault, and then when someone needs to access a database, they go and borrow this credentials and they go and do that. So, that is an awful way to do authentication.

Corey: Now, another way I’ve seen that’s terrible as been also, “Oh, if you’re connecting from this network, you must be allowed in,” which is just… yeee.

Ev: Oh, yeah. That’s a different sin. And that’s a perimeter security sin. But a much better way to do authentication is what is called identity-based authentication. Identity means that you always use your identity of who you are within the company.

So, you would go in through corporate SSO, something like Okta, or Active Directory, or even Google, or GitHub, and then based on that information, you’re given access. So, the resource in this case database, [unintelligible 00:07:39] say, “Oh, it’s Corey. And Corey is a member of this group, and also a member of that group.” And based on that it allows you to get in, but that’s where authentication ends. And now, if you want to do something, like let’s say you want to delete some data, now a database needs to check, ah, can you actually perform that action? That is the authorization process.

And to do that, usually, we use some mechanism like role-based access control. It will look into which group are you in. Oh, you are an admin, so admins have more privileges than regular people. So, then that’s the process of authorization.

And the importance of separating the two, and important to use identity because remember, audit is another important component of implementing access properly. So, if you’re sharing credentials, for example, you will see in your audit log, “Admin did this. Admin did that.” It’s exact same admin, but you don’t know who actually was behind that action. So, by sharing credentials, you’re also obscuring your own audit which is why it’s not really a good thing.

And going back to this industry trends is that because the amount of these resources, like databases and servers and so on, in the cloud has gotten so huge, so we now have this hardware pain, we just have too many things that need access. And all of these things, the software itself is getting more complicated, so now we have a software pain as well, that you have so many different layers in your stack that they need to access. That’s another dimension for introducing access pain. And also, we just have more developers, and the development teams are getting bigger and bigger, the software is eating the world, so there is a people-ware pain. So, on the one hand, you have these four problems you need to solve—connectivity, authentication, authorization, access—and on the other hand, you have more hardware, more software, more people, these pain points.

And so you need to consolidate, and that’s really what we do is that we allow you to have a single place where you can do connectivity, authentication, authorization, and audit, for everything that you have in the cloud. We basically believe that the future is going to be like metaverse, like in those books. So, all of these cloud resources are slowly converging into this one giant planetary-scale computer.

Corey: Suddenly, “I live on Twitter,” is no longer going to be quite as much of a metaphor as it is today.

Ev: [laugh]. No, no. Yeah, I think we’re getting better. If you look into what is actually happening on our computing devices that we buy, the answer is not the lot, so everything is running in data centers, the paradigm of thin client seems to be winning. Let’s just embrace that.

Corey: Yeah. You’re never going to be able to shove data centers worth compute into a phone. By the time you can get there, data centers will have gotten better. It’s the constant question of where do you want things to live? How do you want that to interact?

I talk periodically about multi-cloud, I talk about lock-in, everyone is concerned about vendor lock-in, but the thing that people tend to mostly ignore is that you’re already locked in throught a variety of different ways. And one way is both the networking side of it as well as the identity management piece because every cloud handles that differently and equating those same things between different providers that work different ways is monstrous. Is that the story of what you’re approaching from a Teleport perspective? Is that the primary use case, is that an ancillary use case, or are we thinking about this in too small a term?

Ev: So, you’re absolutely right, being locked in, in and—like, by itself is not a bad thing. It’s a trade-off. So, if you lack expertise in something and you outsourcing certain capability to a provider, then you’re developing that dependency, you may call it lock-in or not, but that needs to be a conscious decision. Like, well, you didn’t know how to do it, then someone else was doing it for you, so you should be okay with the lock-in. However, there is a danger, that, kind of, industry-wide danger about everyone relying on one single provider.

So, that is really what we all try to avoid. And with identity specifically, I feel like we’re in a really good spot that fairly early, I don’t see a single provider emerging as owning everyone’s identity. You know, some people use Okta; others totally happy tying everything to Google Apps. So, then you have people that rely on Amazon AWS native credentials, then plenty of smaller companies, they totally happy having all of their engineers authenticate through GitHub, so they use GitHub as a source of identity. And the fact that all of these providers are more or less compatible with each other—so we have protocols like OpenID Connect and SAML, so I’m not that concerned that identity itself is getting captured by a single player.

And Teleport is not even playing in that space; we don’t keep your identity. We integrate with everybody because, at the end of the day, we want to be the solution of choice for a company, regardless of which identity platform they’re using. And some of them using several, like all of the developers might be authenticating via GitHub, but everyone else goes through Google Apps, for example.

Corey: And the different product problem. Oh, my stars, I was at a relatively small startup going through an acquisition at one point in my career, and, “All right. Let’s list all of the SaaS vendors that we use.” And the answer was something on an average of five per employee by the time you did the numbers out, and—there were hundreds of them—and most of them because it started off small, and great, everyone has their own individual account, we set it up there. I mean, my identity management system here for what most of what I do is LastPass.

I have individual accounts there, two-factor auth enabled for anything that supports it, and that is it. Some vendors don’t support that: we have to use shared accounts, which is just terrifying. We make sure that we don’t use those for anything that’s important. But it comes down to, from our perspective, that everyone has their own ridiculous series of approaches, and even if we were to, “All right, it’s time to grow up and be a responsible business, and go for a single-sign-on approach.” Which is inevitable as companies scale, and there’s nothing wrong with that—but there’s still so many of these edge cases and corner case stories that don’t integrate.

So, it makes the problem smaller, but it’s still there rather persistently. And that doesn’t even get into the fact that for a lot of these tools, “Oh, you want SAML integration? Smells like enterprise to us.” And suddenly they wind up having an additional surcharge on top of that for accessing it via a federated source of identity, which means there are active incentives early on to not do that. So it’s—

Ev: It’s absolutely insane. Yeah, you’re right. You’re right. It’s almost like you get penalized for being small, like, in the early days. It’s not that easy if you have a small project you’re working on. Say it’s a company of three people and they’re just cranking in the garage, and it’s just so easy to default to using shared credentials and storing them in LastPass or 1Password. And then the interesting way—like, the longer you wait, the harder it is to go back to use a proper SSO for everything. Yeah.

Corey: I do want to call out that Teleport has a free and open-source community edition that supports GitHub SSO, and in order to support enterprise SSO, you have to go to your paid offering. I have no problem with this, to be clear, that you have to at least be our customer before we’ll integrate with your SSO solution makes perfect sense, but you don’t have a tiering system where, “Oh, you want to add that other SSO thing? And well, then it’s going to go from X dollars per employee to Y dollars.” Which is the path that I don’t like. I think it’s very reasonable to say that their features flat-out you don’t get as a free user. And even then you do offer SSO just not the one that some people will want to pick.

Ev: Correct. So, the open-source version of Teleport supports SSO that smaller companies use, versus our enterprise offering, we shaped it to be more appealing for companies at certain scale.

Corey: Yeah. And you’ve absolutely nailed it. There are a number of companies in the security space who enraged people about how they wind up doing their differentiation around things like SSO or, God forbid, two-factor auth, or once upon a time, SSL. This is not that problem. I just want to be explicitly clear on that, that is not what I’m talking about. But please, continue.

Ev: Look, we see it the same way. We sometimes say that we do not charge for security, like, top-level security you get, is available even in the open-source. And look, it’s a common problem for most startups who, when you have an open-source offering, where do you draw the line? And sometimes you can find answers in very unexpected places. For example, let’s look into security space.

One common reason that companies get compromised is, unfortunately, human factor. You could use the best tool in the world, but if you just by mistake, like, just put a comma in the wrong place and one of your config files just suddenly is out of shape, right, so—

Corey: People make mistakes and you can’t say, “Never make a mistake.” If you can get your entire company compromised by someone in your office clicking on the wrong link, the solution is not to teach people not to click on links; it’s to mitigate the damage and blast radius of someone clicking on a link that they shouldn’t. That is resilience that understand their human factors at play.

Ev: Yep, exactly. And here’s an enterprise feature that was basically given to us by customer requests. So, they would say we want to have FedRAMP compliance because we want to work with federal government, or maybe because we want to work with financial institutions who require us to have that level of compliance. And we tell them, “Yeah, sure. You can configure Teleport to be compliant. Look, here’s all the different things that you need to tweak in the config file.”

And the answer is, “Well, what if we make a mistake? It’s just too costly. Can we have Teleport just automatically works in that mode?” In other words, if you feed it the config file with an error, it will just refuse to work. So basically, you take your product, and you chop off things that are not compliant, which means that it’s impossible to feed an incorrect config file into it, and here you got an enterprise edition.

It’s a version that we call its FIPS mode. So, when it runs FIPS mode, it has different runtime inside, it basically doesn’t even have a crypto that
is not approved, which you can turn on by mistake. It will just not work.

Corey: By the time we’re talking about different levels of regulatory compliance, yeah, we are long past the point where I’m going to have any comments in the slightest is about differentiation of pricing tiers and the rest. Yeah, your free tier doesn’t support FedRAMP is one of those ludicrous things that—who would say that [laugh] actually be sincere [insane 00:18:28]?

Ev: [laugh].

Corey: That’s just mind-boggling to me.

Ev: Hold on a second. I don’t want anyone to be misinformed. You can be FedRAMP compliant with the free tier; you just need to configure it properly. Like the enterprise feature, in this case, we give you a thing that only works in this mode; it is impossible to misconfigure it.

Corey: It’s an attestation and it’s a control that you need—

Ev: Yep. Yep.

Corey: —in order to demonstrate compliance because half the joy of regulatory compliance is not doing the thing, it’s proving you do the thing. That is a joy, and those of you who’ve worked in regulated environments know exactly what I’m talking about. And those of you who have not, are happy but please—

Ev: Frankly, I think anyone can do it using some other open-source tools. You can even take, like, OpenSSH, sshd, and then you can probably build a different makefile for just the build pipeline that changes the linking, that it doesn’t even have the crypto that is not on the approved list. So, then if someone feeds a config file into it that has, like, a hashing function that is not approved, it will simply refuse to work. So, maybe you can even turn it into something that you could say here’s a hardened version of sshd, or whatever. So, same thing.

Corey: I see now you’re talking about the four aspects of this, the connectivity, the authentication, the authorization, and the audit components of access. How does that map to a software product, if that makes sense? Because it sounds like a series of principles, great, it’s good to understand and hold those in your head both, separately and distinct, but also combining to mean access both [technical 00:19:51] and the common parlance. How do you express that in Teleport?

Ev: So, Teleport doesn’t really add authorization, for example, to something that doesn’t have it natively. The problem that we have is just the overall increasing complexity of computing environments. So, when you’re deploying something into, let’s say, AWS East region, so what is it that you have there? You have some virtual machines, then you have something like Kubernetes on top, then you have Docker registry, so you have these containers running inside, then you have maybe MongoDB, then you might have some web UI to manage MongoDB and Grafana dashboard. So, all of that is software; we’re only consuming more and more of it so that our own code that we’re deploying, it’s icing on a
really, really tall cake.

And every layer in that layer cake is listening on a socket; it needs encryption; it has a login, so it has authentication; it has its own idea of role-based access control; it has its own config file. So, if you want to do cloud computing properly, so you got to have this expertise on your team, how to configure those four pillars of access for every layer in your stack. That is really the pain. And the Teleport value is that we’re letting you do it in one place. We’re saying, consolidate all of this four-axis pillars in one location.

That’s really what we do. It’s not like we invented a better way to authorize, or authenticate; no, we natively integrate with the cake, with all of these different layers. But consolidation, that is the key value of Teleport because we simply remove so much pain associated with configuring all of these things. Like, think of someone like—I’m trying not to disclose any names or customers, but let’s pick, uh, I don’t know, something like Tesla. So, Tesla has compute all over the world.

So, how can you implement authentication, authorization, audit log, and connectivity, too, for every vehicle that’s on the road? Because all of these things need software updates, they’re all components of a giant machine—

Corey: They’re all intermittent. You can’t say, “Oh, at this time of the day, we should absolutely make sure everything in the world is connected to the internet and ready to grab the update.” It doesn’t work that way; you’ve got to be… understand that connectivity is fickle.

Ev: So, most—and because computers growing generally, you could expect most companies in the future to be more like Tesla, so companies like that will probably want to look into Teleport technology.

Corey: This episode is sponsored in part by “you”—gabyte. Distributed technologies like Kubernetes are great, citation very much needed, because they make it easier to have resilient, scalable, systems. SQL databases haven’t kept pace though, certainly not like no SQL databases have like Route 53, the world’s greatest database. We’re still, other than that, using legacy monolithic databases that require ever growing instances of compute. Sometimes we’ll try and bolt them together to make them more resilient and scalable, but let’s be honest it never works out well. Consider Yugabyte DB, its a distributed SQL database that solves basically all of this. It is 100% open source, and there's not asterisk next to the “open” on that one. And its designed to be resilient and scalable out of the box so you don’t have to charge yourself to death. It's compatible with PostgreSQL, or “postgresqueal” as I insist on pronouncing it, so you can use it right away without having to learn a new language and refactor everything. And you can distribute it wherever your applications take you, from across availability zones to other regions or even other cloud providers should one of those happen to exist. Go to yugabyte.com, thats Y-U-G-A-B-Y-T-E dot com and try their free beta of Yugabyte Cloud, where they host and manage it for you. Or see what the open source project looks like—its effortless distributed SQL for global apps. My thanks to Yu—gabyte for sponsoring this episode.

Corey: If we take a look at the four tenets that you’ve identified—connectivity, authentication, authorization, and audit—it makes perfect sense. It is something that goes back to the days when computers were basically glorified pocket calculators as opposed to my pocket calculator now being basically a supercomputer. Does that change as you hit cloud-scale where we have companies that are doing what seem to be relatively pedestrian things, but also having 100,000 EC2 instances hanging out in AWS? Does this add additional levels of complexity on top of those four things?

Ev: Yes. So, there is one that I should have mentioned earlier. So, in addition to software, hardware, and people-ware—so those are three things that are exploding, more compute, more software, more engineers needing access—there is one more dimension that is kind of unique, now, at the scale that we’re in today, and that’s time. So, let’s just say that you are a member of really privileged group like you’re a DBA, or maybe you are a chief security officer, so you should have access to a certain privileged database. But do you really use that access 24/7, all the time? No, but you have it.

So, your laptop has an ability, if you type certain things into it, to actually receive credentials, like, certificates to go and talk to this database all the time. It’s an anti-pattern that is now getting noticed. So, the new approach to access is to make a tie to an intent. So, by default, no one in an organization has access to anything. So, if you want to access a database, or a server, or Kubernetes cluster, you need to issue what’s called ‘access request.’

It’s similar to pull request if you’re trying to commit code into Git. So, you send an access request—using Teleport for example; you could probably do it some other way—and it will go into something like Slack or PagerDuty, so your team members will see that, “Oh, Corey is trying to access that database, and he listed a ticket number, like, some issue he is trying to troubleshoot with that particular database instance. Yeah, we’ll approve access for 30 minutes.” So, then you go and do that, and the access is revoked automatically after 30 minutes. So, that is this new trend that’s happening in our space, and it makes you feel nice, too, it means that if someone hacks into your laptop at this very second, right after you finished authenticating and authorization, you’re still okay because there is no access; access will be created for you if you request it based on the intent, so it dramatically reduces the attack surface, using time as additional dimension.

Corey: The minimum viable permission to do a thing. In principle, least-access is important in these areas. It’s like, “Oh, yeah, my user account, you mean root?” “Yeah, I guess that works in a developer environment,” looks like a Docker container that will be done as soon as you’re finished, but for most use cases—and probably even that one—that’s not the direction to go in. Having things scoped down and—

Ev: Exactly.

Corey: —not just by what the permission is, but by time.

Ev: Exactly.

Corey: Yeah.

Ev: This system basically allows you to move away from root-type accounts completely, for everything. So, which means that there is no root to attack anymore.

Corey: What really strikes me is how, I guess, different aspects of technology that this winds up getting to. And to illustrate that in the form of question, let me go back to my own history because, you know, let’s make it about me here. I’ve mentioned it before on the show, but I started off my technical career as someone who specialized in large-scale email systems. That was a niche I found really interesting, and I got into it. So did you.

I worked on running email servers, and you were the CEO and co-founder of Mailgun, which later you sold the Rackspace. You’re a slightly bigger scale than I am, but it was clear to me that even then, in the 2006 era when I was doing this, that there was not going to be the same need going forward for an email admin at every company; the cloudification of email had begun, and I realized I could either dig my heels in and fight the tide, or I could find other things to specialize in. And I’ve told that part of the story, but what I haven’t told is that it was challenging at first as I tried to do that because all the jobs I talked to looked at my resume and said, “Ah, you’re the email admin. Great. We don’t need one of those.”

It was a matter of almost being pigeonholed or boxed into the idea of being the email person. I would argue that Teleport is not synonymous with email in any meaningful sense as far as how it is perceived in the industry; you are very clearly no longer the email guy. Does the idea being boxed in, I guess—

Ev: [laugh].

Corey: —[unintelligible 00:27:05] resonate at all with you? And if so, how did you get past it?

Ev: Absolutely. The interesting thing is, before starting the Mailgun, I was not an email person. I would just say that I was just general-purpose technologist, and I always enjoyed building infrastructure frameworks. Basically, I always enjoyed building tools for other engineers. But then gotten into this email space, and even though Mailgun was a software product, which actually had surprisingly huge, kind of, scalability requirements early on because email is much heavier than HTTP traffic; people just send a lot of data via emails.

So, we were solving interesting technical challenges, but when I would meet other engineers, I would experience the exact same thing you did. They would put me into this box of, “That’s an email guy. He knows email technology, but seemingly doesn’t know much about scaling web apps.” Which was totally not true. And it bothered me a little bit.

Frankly, it was one of the reasons we decided to get acquired by Rackspace because they effectively said, “Why don’t you come join us and we’ll continue to operate as independent company, but you can join our cloud team and help us reinvent cloud computing.” It was really appealing. So, I actually moved to Texas after acquisition; I worked on the Rackspace cloud team for a while. So, that’s how my transition from this being in the email box happened. So, I went from an email expert to just generally cloud computing expert. And cloud computing expert sounds awesome, and it allows me to work—

Corey: I promise, it’s not awesome—

Ev: [laugh].

Corey: —for people listening to this. Also, it’s one of those, are you a cloud expert? Everyone says no to that because who in the world would claim that? It’s so broad in so many different expressions of it. Because you know the follow-up question to anyone who says, “Yeah,” is going to be some esoteric thing about a system you’ve never heard of before because there’s so many ridiculous services across totally different providers, of course, it’s probably a thing. Maybe it’s actually a Pokemon, we don’t know. But it’s hard to consider yourself an expert in this. It’s like, “Well, I have some damage from [laugh] getting smacked around by clouds and, yeah, we’ll call that expertise; why not?”

Ev: Exactly. And also how frequently people mispronounce, like, cloud with clown. And it’s like, “Oh, I’m clown computing expert.” [laugh].

Corey: People mostly call me a loud computing expert. But that’s a separate problem.

Ev: But the point is that if you work on a product that’s called cloud, so you definitely get to claim expertise of that. And the interesting thing that Mailgun being, effectively, an infrastructure-level product—so it’s part of the platform—every company builds their own cloud platform and runs it, and so Teleport is part of that. So, that allowed us to get out of the box. So, if you working on, right now we’re in the access space, so we’re working closely with Kubernetes community, with Linux kernel community, with databases, so by extension, we have expertise in all of these different areas, and it actually feels much nicer. So, if you are computing security access company, people tend to look at you, it’s like, “Yeah, you know, a little bit of everything.” So, that feels pretty nice.

Corey: It’s of those cross-functional things—

Ev: Yeah, yeah.

Corey: —whereas on some level, you just assume, well, email isn’t either, but let’s face it: email is the default API that everything, there’s very little that you cannot configure to send email. The hard part is how to get them to stop emailing you. But it started off as far—from my world at least—the idea that all roads lead to email. In fact, we want to talk security, a long time ago the internet collectively decided one day that our email inbox was the entire cornerstone of our online identity. Give me access to your email, I, for all intents and purposes, can become you on the internet without some serious controls around this.

So, those conversations, I feel like they were heading in that direction by the time I left email world, but it’s very clear to me that what you’re doing now at Teleport is a much clearer ability to cross boundaries into other areas where you have to touch an awful lot of different things because security touches everything, and I still maintain it has to be baked-in and an intentional thing, rather than, “Oh yeah, we’re going to bolt security on after the fact.” It’s, yeah, you hear about companies that do that, usually in headlines about data breaches, or worse. It’s a hard problem.

Ev: Actually, it’s an interesting dilemma you’re talking about. Is security built-in into everything or is it an add-on? And logically—talk to anyone, and most people say, “Yeah, it needs to be a core component of whatever it is you’re building; making security as an add-on is not possible.” But then reality hits in, and the reality is that we’re running on—we’re standing on the shoulder of giants.

There is so much legacy technologies that we built this cloud monster on top of… no, nothing was built in, so we actually need to be very crafty at adding security on top of what we already have, if we want to take advantage of all this pre-existing things that we’ve built for decades. So, that’s really what’s happening, I think, with security and access. So, if you ask me if Teleport is a bolt-on security, I say, “Yes, we are, but it works really well.” And it’s extremely pragmatic and reasonable, and it gives you security compliance, but most of all, very, very good user experience out of the box.

Corey: It’s amazing to me how few security products focus on user experience out of the box, but they have to. You cannot launch or maintain a security product successfully—to my mind—without making it non-adversarial to the user. The [days of security is no 00:32:26] are gone.

Ev: Because of that human element insecurity. If you make something complicated, if you make something that’s hard to reason about, then it will never be secure.

Corey: Yeah.

Ev: Don’t copy-paste IP table rules without understanding what they do. [laugh].

Corey: Yeah, I think we all have been around long enough in data center universes remember those middle of the night drives to the data center for exactly that sort of thing. Yeah, it’s one of those hindsight things of, set a cron job to reset the IP table rules for, you know, ten minutes from now in case you get this hilariously wrong. It’s the sort of thing that you learn right after you really could have used that knowledge. Same story. But those are the easy, safe examples of I screwed up on a security thing. The worst ones can be company-ending.

Ev: Exactly, yeah. So, in this sense, when it comes to security, and access specifically, so this old Python rule that there is only one way to do something, it’s the most important thing you can do. So, when it comes to security and access, we basically—it’s one of the things that Teleport is designed around, that for all protocols, for all different resources, from SSH to Kubernetes to web apps to databases; we never support passwords. It’s not even in the codebase. No, you cannot configure Teleport to use passwords.

We never support things like public keys, for example, because it’s just another form of a password. It’s just extremely long password. So, we have this approach that certificates, it’s the best method because it supports both authentication and authorization, and then you have to do it for everything, just one way of doing everything. And then you apply this to connectivity: so there is a single proxy that speaks all protocols and everyone goes to that proxy. Then you apply the same principle to audit: there is one audit where everything goes into.

So, that’s how this consolidation, that’s where the simplicity comes down to. So, one way of doing something; one way of configuring everything. So, that’s where you get both ease of use and security at the same time.

Corey: One last question that I want to ask you before we wind up calling this an episode is that I’ve been using Teleport as a reference for a while when I talk to companies, generally in the security space, as an example of what you can do to tell a story about a product that isn’t built on fear, uncertainty, and doubt. And for those who are listening who don’t know what I’m referring specifically, I’m talking about pick any random security company and pull up their website and see what it is that they talk about and how they talk about themselves. Very often, you’ll see stories where, “Data breaches will cost you extraordinary piles of money,” or they’ll play into the shame of what will happen to your career if you’re named in the New York Times for being the CSO when the data gets breached, and whatnot. But everything that I’ve seen from Teleport to date has instead not even gone slightly in that direction; it talks again and again, in what I see on your site, about how quickly it is to access things, access that doesn’t get in the way, easily implement security and compliance, visibility into access and behavior. It’s all about user experience and smoothing the way and not explaining to people what the dire problems that they’re going to face are if they don’t care about security in general and buy your product specifically. It is such a refreshing way of viewing storytelling around a security product. How did you get there? And how do I make other people do it, too?

Ev: I think it just happened organically. Teleport originally—the interesting story of Teleport, it was not built to be sold. Teleport was built as a side project that we started for another system that we were working on at the time. So, there was a autonomous Kubernetes platform called Grá—it doesn’t really matter in this context, but we had this problem that we had a lot of remote sites with a lot of infrastructure on them, with extremely strict security and compliance requirements, and we needed to access those sites or build tools to access those sites. So, Teleport was built like, okay, it’s way better than just stitching a bunch of open-source components together because it’s faster and easier to use, so we’re optimizing for that.

And as a side effect of that simplification, consolidation, and better user experience is a security compliance. And then the interesting thing that happened is that people who we’re trying to sell the big platform to, they started to notice about, “Oh, this access thing you have is actually pretty awesome. Can we just use that separately?” And that’s how it turned into a product. So, we built an amazing secure access solution almost by accident because there was only one customer in mind, and that was us, in the early days. So yeah, that’s how you do it, [laugh] basically. But it’s surprisingly similar to Slack, right? Why is Slack awesome? Because the team behind it was a gaming company in the beginning.

Corey: They were trying to build a game. Yeah.

Ev: Yeah, they built for themselves. They—[laugh] I guess that’s the trick: make yourself happy.

Corey: I think the team founded Flickr before that, and they were trying to build a game. And like, the joke I heard is, like, “All right, the year is 2040. Stuart and his team have now raised $8 billion trying to build a game, and yet again it fails upward into another productivity tool company, or something else entirely that”—but it’s a recurring pattern. Someday they’ll get their game made; I have faith in them. But yeah, building a tool that scratches your own itch is either a great path or a terrible mistake, depending entirely upon whether you first check and see if there’s an existing solution that solves the problem for you. The failure mode of this is, “Ah, we’re going to build our own database engine,” in almost every case.

Ev: Yeah. So just, kind of like, interesting story about the two, people will [unintelligible 00:38:07] surprised that Teleport is a single binary. It’s basically a drop-in replacement that you put on a box, and it runs instead of sshd. But it wasn’t initially this way. Initially, it was [unintelligible 00:38:16], like, few files in different parts of a file system. But because internally, I really wanted to run it on a bunch of Raspberry Pi’s at home, and it would have been a lot easier if it was just a single file because then I just could quickly update them all. So, it just took a little bit of effort to compress it down to a single binary that can run in different modes depending on the key. And now look at that; it’s a major benefit that a lot of people who deploy Teleport on hundreds of thousands of pieces of infrastructure, they definitely taking advantage of the fact that it’s that simple.

Corey: Simplicity is the only thing that scales. As soon as it gets complex, it’s more things to break. Ev, thank you so much for taking the time to sit with me, yet again, to talk about Teleport and how you’re approaching things. If people want to learn more about you, about the company, about the product in all likelihood, where can they go?

Ev: The easiest place to go would be goteleport.com where you can find everything, but we’re also on GitHub. If you search for Teleport in GitHub, you’ll find this there. So, join our Slack channel, join our community mailing list and most importantly, download Teleport, put it on your Raspberry Pi, play with it and see how awesome it is to have the best industry, best security practice, that don’t get in the way.

Corey: I love the tagline. Thank you so much, once again. Ev Kontsevoy, co-founder and CEO of Teleport. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with a comment that goes into a deranged rant about how I’m completely wrong, and the only way to sell security products—specifically yours—is by threatening me with the New York Times data breach story.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Katie

Katie Sylor-Miller, Frontend Architect at Etsy, has a passion for design systems, web performance, accessibility, and frontend infrastructure. She co-authored the Design Systems Handbook to spread her love of reusable components to engineers and designers. She’s spoken at conferences like Smashing Conf, PerfMatters Conf, JamStack Conf, JSConf US, and FrontendConf.ch (to name a few). Her website ohshitgit.com (and the swear-free version dangitgit.com) has helped millions of people worldwide get out of their Git messes, and has been translated into 23 different languages and counting.

Links:

  • Etsy: https://www.etsy.com/
  • Design Systems Handbook: https://www.designbetter.co/design-systems-handbook
  • Book of staff engineering stories: https://www.amazon.com/dp/B08RMSHYGG
  • staffeng.com: https://staffeng.com
  • ohshitgit.com: https://ohshitgit.com
  • dangitgit.com: https://dangitgit.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Thinkst. This is going to take a minute to explain, so bear with me. I linked against an early version of their tool, canarytokens.org in the very early days of my newsletter, and what it does is relatively simple and straightforward. It winds up embedding credentials, files, that sort of thing in various parts of your environment, wherever you want to; it gives you fake AWS API credentials, for example. And the only thing that these things do is alert you whenever someone attempts to use those things. It’s an awesome approach. I’ve used something similar for years. Check them out. But wait, there’s more. They also have an enterprise option that you should be very much aware of canary.tools. You can take a look at this, but what it does is it provides an enterprise approach to drive these things throughout your entire environment. You can get a physical device that hangs out on your network and impersonates whatever you want to. When it gets Nmap scanned, or someone attempts to log into it, or access files on it, you get instant alerts. It’s awesome. If you don’t do something like this, you’re likely to find out that you’ve gotten breached, the hard way. Take a look at this. It’s one of those few things that I look at and say, “Wow, that is an amazing idea. I love it.” That’s canarytokens.org and canary.tools. The first one is free. The second one is enterprise-y. Take a look. I’m a big fan of this. More from them in the coming weeks.

Corey: This episode is sponsored in part by our friends at Jellyfish. So, you’re sitting in front of your office chair, bleary eyed, parked in front of a powerpoint and—oh my sweet feathery Jesus its the night before the board meeting, because of course it is! As you slot that crappy screenshot of traffic light colored excel tables into your deck, or sift through endless spreadsheets looking for just the right data set, have you ever wondered, why is it that sales and marketing get all this shiny, awesome analytics and inside tools? Whereas, engineering basically gets left with the dregs. Well, the founders of Jellyfish certainly did. That’s why they created the Jellyfish Engineering Management Platform, but don’t you dare call it JEMP! Designed to make it simple to analyze your engineering organization, Jellyfish ingests signals from your tech stack. Including JIRA, Git, and collaborative tools. Yes, depressing to think of those things as your tech stack but this is 2021. They use that to create a model that accurately reflects just how the breakdown of engineering work aligns with your wider business objectives. In other words, it translates from code into spreadsheet. When you have to explain what you’re doing from an engineering perspective to people whose primary IDE is Microsoft Powerpoint, consider Jellyfish. Thats Jellyfish.co and tell them Corey sent you! Watch for the wince, thats my favorite part.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’m joined this week by Katie Sylor-Miller, who is a frontend architect at Etsy. Katie, thank you for joining me.

Katie: Hi, Corey. Thanks for having me.

Corey: So, I met you a long time ago—before anyone had ever heard of me and the world was happier for it—but since then you’ve done a lot of things. You’re obviously a frontend architect at Etsy. You’re a co-author of the Design Systems Handbook, and you were recently interviewed and included Will Larson’s book of staff engineering stories that people are mostly familiar with at staffeng.com.

Katie: Yeah.

Corey: So, you’ve done a lot of writing; you’ve done some talking, but let’s begin with the time that we met. To my understanding, it’s the only
time we’ve ever met in person. And this harkens back to the first half—as I recall—of 2016 at the frontend conference in Zurich.

Katie: Yes, before either of us were known for anything. [laugh].

Corey: Exactly. And it was, oh, great. And I wound up getting invited to speak at a frontend conference. And my response was, “Uh, okay. Zurich sounds lovely. I’m thrilled to do it. Do you understand who you’re asking?”

There are frontend folks—which, according to the worst people on the internet is the easiest form of programming; it isn’t a real engineering job, and if that’s your opinion, please stop listening to anything I do ever again—secondly, then there’s the backend folks who write the API side of things and what the deep [unintelligible 00:02:03] and oh, that’s the way of the future. And people look at me and they think, “Oh, you’re a backend person,” if their frontend. If they’re backend, they look at me and think, “Oh, you’re a DevOps person.” Great. And if you’re on the DevOps space, you look at me and think, “What is wrong with this person?” And that’s mostly it.

But I was actually invited to speak at a frontend conference. And the reason that they invited me at all—turns out wasn’t a mistake—was that I was giving a talk that year called, “Terrible Ideas in Git,” which is the unifying force that ties all of those different specialties together by confusing the living hell out of us.

Katie: Yes. [laugh].

Corey: So, I gave a talk. I thought it was pretty decent. I’ve done some Twitter threads on similar themes. You did something actually useful that helps people and is more lasting—and right at that same conference, I believe, you were building slash kicking it off—ohshitgit.com.

Katie: Yes. Yeah. It was—

Corey: Which is amazing.

Katie: Thank you. Yeah, it was shortly thereafter. I think the ideas were kind of starting to percolate at that conference. Because you know—yeah I was—

Corey: Because someone gave a talk about Git. Oh, I’m absolutely stealing credit for your work.

Katie: No, Corey—

Corey: “Oh, yeah. You know, that was my idea.”

Katie: [laugh].

Corey: Five years from now, I’m going to call myself the founder of it, and you’re just on the implementation details.

Katie: I don’t—nonononono—

Corey: That’s right. I’m going to D.C. Bro my way through all of this.

Katie: [laugh]. No, no, no, no. See, my recollection is that my talk about being a team player and a frontend expert with a T-shape happened at exactly the same time as your talk about Git because I remember I wanted to go watch your talk because at the time, I absolutely hated Git. I was still kind of learning it. So yeah, so I don’t think you really get any credit because I have never actually heard that talk that you gave. [laugh].

Corey: A likely story.

Katie: [laugh]. However, however, I will say—so, before I was up to give my talk, the emcee of the conference was teasing me, you know, in a very good-natured ribbing sort of way, he was teasing me about my blog being totally empty and having absolutely nothing in it. And I got on the plane home from Zurich, and I was starting to think, “Oh, okay. What are some things that I could blog about? What do I have to say that would be at all interesting or new to anyone else?”

And like I think a lot of people do, I had a really hard time figuring out, okay, what can I say that’s, maybe, different? And, I went back home, I went back to work, and at one point, I had this idea, I had this file that I had been keeping ever since I started learning Git and I call it, like, gitshit.txt. And hopefully, your listeners don’t mind lots of swears because I’m probably going to swear quite a bit.

Corey: No, no. I do want to point out, you’re accessible to all folks: dangitgit.com, also works but doesn’t have the internal rhyming mechanism which makes it, obviously, nowhere near where it needs to be.

Katie: [laugh]. Well—

Corey: It’s sort of a Subversion to Git if you will.

Katie: Yes, exactly.

Corey: I—Subversion fans, don’t yell at me.

Katie: [laugh]. Anyways, so I remember I tweeted something like, “Oh, what about if I took this text file that I had,” where every time I got into a Git mess, I would go on to Stack Overflow—as you do—and I would Google and I—it was so hard. I couldn’t find the words to find the answers to what I was trying to fix. Because one of the big problems with Git that we can talk about it a bit more in detail later is that Git doesn’t describe workflows, Git describes internal plumbing commands and everything that it exposes in its API. So, I had a really hard time with it; I had a hard time learning it.

And, you know, what I said, “Okay, well, maybe if I published on my blog about these Git tips that I had saved for myself.” And I remember I tweeted, and I got a handful of likes on the tweet, including from Eric Meyer, who is one of my big idols in the frontend world. He’s one of the godfathers of modern CSS. And he liked my tweet, and I was like, “Oh, okay. Maybe this is a real thing. Maybe people will actually find this interesting.”

And then I had this brilliant idea for this URL, ohshitgit.com, and it was available, and I bought it. And I swear to you, I think I spent two hours writing some HTML around my text file and publishing it up to my server. And I tweeted about it, and then I went to bed.

And I kind of expected maybe half a dozen of my coworkers would get a little sensible chuckle out of it, and like, that would be the end of it. But I woke up the next morning and my Twitter had blown up; I was on the front page of Hacker News. I had coworkers pinging me being like, “Oh, my God, Katie, you’re on Hacker News. This is insane.” And—

Corey: Wait, wait, for a good thing, or the horrifying kind of thing because, Hacker News?

Katie: Well, [laugh] as I have discovered with Hacker News, whenever my site ends up on Hacker News, the response is generally, like, a mix of, “Ha ha ha, this is great. This is funny,” and, “Oh, my God, somebody actually doesn’t understand Git and needs this. Wow, people are really stupid.” Which I fundamentally disagree with and I’m sure that you fundamentally [laugh] disagree with as well.

Corey: Oh, absolutely.

Katie: Yeah. So—

Corey: It’s one of those, “Oh, Git confuses you. You know what that means? It means you’re human.” It confuses everyone. The only question is, at what point does it escape your fragile mortal understanding? And if you are listening to this and you don’t believe me, great. I’m easy to find, I will absolutely have that discussion with you in public because I promise, one of us is going to learn something.

Katie: [laugh]. Awesome. I love—I hope that people take you up on that because—

Corey: Oh, that would be an amazing live stream, wouldn’t it?

Katie: It would. It would because Git is one of those things that I think that people who don’t understand it, look at it and think, “Gosh, you know, I must be stupid,” or, “I must not be cut out to be a developer,” or, “I must not know what I’m doing.” And I know that this is how people feel because that’s exactly how I felt myself, even when I made ohshitgit.com, that became this big reference that everybody looks at to help them with Git, like, I still didn’t understand it. I didn’t get Git at all.

And since then, I’ve kind of been forced because people started asking me all these questions, and, “Well, what about this? What about that?” And I was just like, “Uh… I don’t know. Uh…” and I didn’t like that feeling, so I did what, you know, obviously, anyone would do in my situation and I sent out a proposal to give a talk about Git at a conference. [laugh].

And what that did is when my talk got accepted, I had to then go off and actually learn Git and understand how it works so that I could go and teach it to other people at this conference. But it ended up being great, I think because I found a lot of really awesome books. There’s A Book Apart book called Git for Humans, which is incredibly good. There’s a couple of websites like learngitbranching.com.

There’s a bunch more that I can’t think of off the top of my head. But I went out and I sort of slowly but surely developed this mental model, internally, of how Git works. And I’m a visual thinker and I’m a visual learner, and so it’s a very visual model. And for what it’s worth, I think that was my biggest problem with Git was, like, I came from Microsoft .NET environment before that, and we used a program called TFS, Team Foundation Server, which is basically like a SVN or a CVS type source control system that was completely integrated into Visual Studio.

So, it was completely visual; you could see everything happening in your IDE as you were doing it. And then making this switch to the command line, I just could not figure it out until I had this visual mental model. So yeah, so ever since then I’ve just been going around and trying to teach people about Git and teach people this visual mental model that I’ve developed, and the tips and the tricks that I’ve learned for navigating Git especially on the command line. And I give talks, I do full-day training workshops, I do training workshops at work. And it’s become my thing now, which is flabbergasting [laugh] because I never intended [laugh] for—I didn’t set out to go and be this Git expert or to be, quote-unquote, “Famous” for a given value of famous, for knowing stuff about Git. I’m a frontend engineer. There’s still a piece of me that looks at it, and is like, “How on earth did this even happen to me?” So, yeah, I don’t know. So, that’s my Oh shit, Git!?! story. And now—

Corey: It’s a great one. It’s—

Katie: Thank you.

Corey: Git is one of those weird things where the honest truth of were, “Terrible Ideas in Git”—my talk—came from was that I kept trying and failing to understand Git, and I realized, “How do I fix this? I know. I will give a talk about something.” That is what we know as a forcing function. If I’m not quite ready, they will not move the conference. I know because I checked.

Katie: Yep. [laugh]

Corey: And one in Zurich was not the first time I’d given it, but it was very clearly something that everyone had problems with. The first version of that talk would have absolutely killed it, if I’d been able to give it to the core Git maintainers. And all, you know, seven of those people would have absolutely loved it, and everyone else would have been incredibly confused. So, I took the opposite tack and said, “All right. How do I expand this to as broad an audience as possible?”

And in one of the times I gave it, I said, “Look, I want to make sure it is accessible to everyone, not just people who are super deep into the weeds but also be able to explain Git to my mother.” And unlike virtually every other time where that, “Let me explain something to my mom.” And that is basically coded ageism and sexism built into one. In that case, it was because my mother was sitting in the front row and does not understand what Git is. And she got part of the talk and then did the supportive mother thing of, and as for the rest of it. “Oh, you’re so well-spoken. You’re so funny. And people seem to love it.” Like, “Did you enjoy my discussion of rebases?”

Katie: [laugh].

Corey: She says, “Just so good at talking. So, good.” And it was yeah.

Katie: [laugh]. Oh, yeah. No, I, I—totally—I understand that. There’s this book that I picked up when I was doing all of this research, and I’m looking over at my bookshelf, it’s called Version Control with Git. It’s an O’Reilly book.

And if I remember correctly, it was written by somebody who actually worked at Git. And the way that they started to describe how Git works to people was, they talked about all kinds of deep internals of Unix, and correlated these pieces of the deep internals of Git to these deep Unix internals, which, at the time, makes sense because Git came out of the Unix kernel project as their source control methodology, but, like, really? Like, [laugh] this book, it says at the beginning, that it’s supposed to teach people who are new to Git about how to use it. And it’s like, well, the first assumption that they make is that you understand the 15 years’ worth of history of the Linux kernel project and how Linux works under the hood. And it’s like, you’ve got to be absolutely kidding me that this is how anyone could think, “Oh, this is the right way to teach people Git.”

I mean, it’s great now, going back in and rereading that book more recently, now that I’ve already got that understanding of how it works under the hood. This is giving me all of this detail, but for a new person or beginner, it’s absolutely the wrong way to approach teaching Git.

Corey: When I first sat down to learn Git myself it was in 2008, 2009, Scott Chacon from GitHub at the time wound up doing a multi-day training at the company I worked at the time. And it was very challenging. I’m not saying that he was a bad teacher by any stretch of the imagination, but back in those days, Git was a lot less user-friendly—[laugh] not that it’s tremendously good at it now—and people didn’t understand how to talk about it, how to teach it, et cetera. You go to GitHub or GitLab or any of the other sites that do this stuff, and there’s a 15-step intro that you can learn in 15 minutes and someone who has never used Git before now knows the basics and is not likely to completely shatter things. They’ve gotten the minimum viable knowledge to get started down to a very repeatable, very robust thing. And that is no small feat. Teaching people effectively is super hard.

Katie: It really is. And I totally agree with you that if you go to these providers that they’ve invested in improving the user experience and making things easier to learn. But I think there’s still this problem of what happens when everything goes wrong? What happens if you make a mistake, or what happens if you commit a file on the wrong branch? Or what happens if you make a commit but you forgot to add one of the files you wanted to put in the commit?

Or what happens if you want to undo something that you did in a previous commit? And I think these are things that are still really, for some reason, not well understood. And I think that’s kind of why Oh Shit, Git!?! has fallen into this little niche corner of the Git world is because the focus is really like, “Oh, shit. I just made a mistake and I don’t know what to do, and I don’t know what terminology to even Google for to help me figure out how to fix this problem.” And I’ve come out and put these very simple, like, here: step one, step two, step three.

And people might disagree or argue [laugh] with some of the commands and some of the orders, but really, the focus is, like, people have this idea in their head, I think, particularly at their jobs, that Git is this big, important thing and if you screw up, you can’t fix it. When really a lot of helping people to become more familiar and comfortable with Git is about ensuring them that no, no, no, the whole point of Git is that just about everything can be undone, and just about everything is fixable, and here’s how you do it. So, I still think that we have a long way to go when it comes to teaching Git.

Corey: I would agree wholeheartedly. And I think that most people are not thinking about this from a position of educators, they’re thinking about it from the position of engineering, and it’s a weird combination of the two. You’re not going to generally find someone who has no engineering experience to be able to explain things in a context that resonates with the people who will need to apply it. And on the other side, you’re not going to find that engineers are great at explaining things without having specific experience in that space. There are exceptions, and they are incredibly rare and extremely valuable as a result. The ability to explain complex things simply is a gift.

Katie: It really is.

Corey: It’s also a skill and you can get better at it, but a lot of folks just seem to never put the work in in the first place.

Katie: Well, you know, it’s quote-unquote, “soft skills.” So [laugh].

Corey: Oh, God. They’re hard as hell, so it’s a terrible name.

Katie: [laugh]. Yeah. Though I could not agree more, I think something that I really look at as a trait of a super senior engineer is that they are somebody who has intentionally worked on and practiced developing that skill of taking something that’s a really complex technical concept, and understanding your audience, and having some empathy to put yourself in the shoes of your audience and figure out okay, how do I break this down and explain it to someone who maybe doesn’t have all the context that I do? Because when you think about it, if you’re working at a big company, and you’re an engineer, and you want to, like, do the new hotness, cool thing, and you want to make Kubernetes the thing or whatever other buzzword term you want to use, in order to get that prioritized and on a team’s backlog, you have to turn around and explain to a product person why it’s important for product reasons, or what benefits is this going to bring to the organization as far as scalability, and reliability. And you have to be able to put yourself in the shoes of someone whose goals are totally different than yours.

Like, product people’s goals are all around timelines, they’re around costs, they’re around things short-term versus long-term improvements. And if you can’t put yourself into the shoes of that person, and figure out how to explain your cool hot tech thing to them, then you’re never going to get your project off the ground. No one’s ever going to approve it, nobody’s going to give budget, nobody’s going to put it in a team’s backlog unless you have that skill.

Corey: That’s the hard part is that people tend to view advancement as an individual contributor or engineer purely through a lens of technical ability. And it’s not. The higher you rise, the more your job involves talking to people, and the less it involves writing code in almost every case.

Katie: One hundred percent. That’s absolutely been my experience as an architect is that, gosh, I almost never write code these days. My entire job is basically writing docs, talking to people, meeting with people, trying to figure out, where, what is the left hand doing and what is the right hand doing so I can somehow create a bridge between them. You know, I’m trying to influence teams, and their approach, and the way that they think about writing software. And, yes there is a foundation of technical ability that has to be there.

You have to have that knowledge and that experience, but at this point, it’s like, my God—you know, I write more SQL as a frontend architect that I write HTML, or CSS, or JavaScript because I’m doing data analysis and [laugh] I’m doing—I’m trying to figure out what does the numbers tell us about the right thing to choose or the right way to go, or where are we having issues? And, yeah, I think that people’s perceptions and the reality don’t always match up when it comes to looking at the senior IC technical track.

Corey: This episode is sponsored by our friends at Oracle Cloud. Counting the pennies, but still dreaming of deploying apps instead of "Hello, World" demos? Allow me to introduce you to Oracle's Always Free tier. It provides over 20 free services and infrastructure, networking databases, observability, management, and security.

And - let me be clear here - it's actually free. There's no surprise billing until you intentionally and proactively upgrade your account. This means you can provision a virtual machine instance or spin up an autonomous database that manages itself all while gaining the networking load, balancing and storage resources that somehow never quite make it into most free tiers needed to support the application that you want to build.

With Always Free you can do things like run small scale applications, or do proof of concept testing without spending a dime. You know that I always like to put asterisks next to the word free. This is actually free. No asterisk. Start now. Visit https://snark.cloud/oci-free that's https://snark.cloud/oci-free.

Corey: At some level, you hear people talking about wanting to get promoted, and what they’re really saying—and it doesn’t seem that they realize this—is, “I love what I do, so I’m really trying to get promoted so I can do less of what I love and a lot more of things I hate.”

Katie: [laugh]. Yes. Yeah. Yeah. [laugh]. In some ways, in some ways, I think that you’ve got to kind of learn to accept it. And there are some people, I think that once you get past the senior engineer, or maybe even the staff engineer, maybe they don’t even want to go there because they don’t want to do the kind of sales pitch, people person, data numbers pitching, trying to get people to agree with you on the right way forward is really hard, and I don’t think it’s for everyone. But I love it. [laugh]. I absolutely love it. It’s been great for me. And I feel like it really—it plays to my strengths in a lot of ways.

Corey: What I always found that worked for me, as far as getting folks on board with my vision of the world is, first, I feel like I have to grab their attention, and my way is humor. With the Git talk, I have to say giving that talk a few times made me pretty confident in it. And then I was invited to the frontend conference. And in hindsight, I really, really should have seen this coming, but I’m there, I’m speaking in the afternoon, I’m watching the morning talks, and the slides are all gorgeous.

Katie: Yes. [laugh].

Corey: And then looking at my own, and they are dogshit. Because this was before I had the sense to hire a designer to help with these things. It was effectively black Helvetica text on a white background. And I figured, “All right, this is a problem. I only have a few hours to go, what do I do?”

And my answer was, “Well, I’m not going to suddenly become an amazing designer in the four hours I have.” So, I changed some of the text to Comic Sans because if you’re doing something bad, do it worse, and then make it look intentional. It was a weird experience, and it was a successful talk in that no one knew what the hell to make of what I was doing. And it really got me thinking that this was the first time I’d spoken to an audience who was frontend, and it reminded me that the DevOps problems that I normally talked about, were usually fairly restricted to DevOps. But the things that everyone touches, like Git, for example, start to be things that resonate and break down walls and silos better than a given conference ever can. But talking instead about shared pain and shared frustrations.

Katie: Yes. Yes. Everyone likes to know that they are not alone in the world, particularly folks who are maybe underrepresented minorities in tech and who are afraid to speak up and say, “Oh, I don’t understand.” Or, “That doesn’t make any sense to me,” because they’re worried that they’re already being taken not as seriously as their white, male counterparts. And I feel like something I really try to lean into as a very senior woman in a very male-dominated field is if I don’t understand something, or if I have a question, or something doesn’t make sense is I try to raise my hand and ask those questions and say, out loud, “Okay, I don’t get this.”

Because I can’t even tell you, Corey, the number of times I’ve had somebody reach out to me after a meeting and say, “Thank you. I didn’t understand it either.” Or, “I thought maybe I just didn’t understand the problem space, or maybe I just wasn’t smart enough to understand their explanation.” And having somebody who’s very senior who folks look up to, to be able to say, “Wait a minute, this doesn’t make sense.” Or, you know, I don’t understand that explanation.

Can you explain it a different way? It’s so powerful and it unblocks people and it gives them this confidence that, hey, if that person up on stage, or leading this meeting, or writing this blog post doesn’t get this either, maybe I’m not so stupid, or maybe I do deserve to be in this industry, or maybe it’s not just me. And I really hope that more and more people can feel empowered to do that in their daily lives more. I think that’s been something that has been a tremendous learning through all of this experience with Oh shit, Git!?!

For me is the number of people that come up to me after conference talks, or tweet me, or send me a message, just saying, “Thank you. I thought I was alone. I thought I was the only one that didn’t get this.” And knowing that not just am I not the only one, but that people are universally frustrated, and universally Git makes them want to swear all the time, I mean, that’s the best compliments that I get is when folks come up to me and say, “Thank you, I thought I was alone.”

Corey: That’s one of the things that I find that is simultaneously the most encouraging and also the most galling. Every once in a while I will have some company reach out to me—over a Twitter thread or something—where I’m going through their product from a naive user perspective of, like, I’m not coming at this with 15 years of experience and instinct that feed into how I approach this, but instead the, I actually haven’t used this product before. I’m not going to jump ahead and make assumptions that tend to be right. I’m going to follow the predictable user path flow. And they are very often times where, “Okay. I’m hitting something. I don’t understand this. Why is it like this? This is not good.”

And usually, companies are appreciative when I do stuff like that, but every once in a while, I’ll get some dingus who will come in, and like, “I didn’t appreciate the fact that you end up intentionally misinterpreting what we’re saying.” And that’s basically license for me to take the gloves off and say, “No, this was not me being intentionally dumb. Sure, I didn’t apply a whole bunch of outside resources I could have to this, but it wasn’t me intentionally failing to get the point. I did not understand this, and you’re coming back to me now reinforces that you are too close to the problem. And, on some level, when your actual customers have problems with this, they are hearing an element of contempt from you.”

Katie: Totally.

Corey: “This is an opportunity to fix it and make it more approachable because spoiler, not a lot of people love paying money to something that makes them feel stupid.”

Katie: [laugh]. See, Corey, I don’t know. You say that you’re not really a frontend person, but that is a very strong UX mindset. Like that—

Corey: Oh, my frontend stuff is actually pretty awesome because as soon as I have to do something that even borders on frontend, I have the insight and I guess, willingness to do the smart thing, which is to immediately stop talking and pay someone who knows what they’re doing.

Katie: [laugh]. Thank you. On behalf of all frontend engineers everywhere, I applaud that, and I appreciate it.

Corey: It comes down to specialty. I mean, again, it would also be sort of weird from my perspective, which is my entire corporate position is I fix the horrifying AWS bill. So, if you’re struggling with the bill in various capacities, first, join basically everyone, but two, you’re not alone so maybe hire someone who is an expert in this specific thing to come in and help you with it. And wouldn’t it be a little hypocritical of me to go in and say, “Oh, yeah, but I’m just going to YOLO my way through this nonsense?”

Katie: Mm-hm. [laugh]. Yeah, [laugh] I don’t know we’ll want to include this in the final recording, but I have a really hilarious story, actually, about Amazon. So—

Corey: Oh, please. They listen to this and they love customer feedback.

Katie: [laugh].

Corey: I’m not being sarcastic. I’m very sincere here.

Katie: Well, this is many, many, many years ago. I mean, probably, oh, gosh, this is probably eight years ago at this point. I was interviewing for a job at Amazon. It was a job to be a frontend engineer on the homepage team, which at the time, I was like, “Oh, my God, this is Amazon. This is such an honor. I’m so excited.”

Corey: And you look at amazon.com’s front page, and it’s, “Oh, I can fix this. There’s so much to fix here.”

Katie: Yes.

Corey: And then reality catches up if I might not be the first person in the world to have made that observation.

Katie: [laugh].

Corey: What’s—

Katie: Well—

Corey: Going on in there?

Katie: Yeah. Well, I’ll tell you what’s going on. So, I think I did five different phone interviews. You know, before they invite you out to Seattle, there’s—and again, this was eight years ago, so this was well before everyone was working at home. And in those five hours of phone interviews, I want you to make a guess at how many minutes we spent talking about HTML, CSS, and JavaScript.

Corey: I am so unfamiliar with the frontend world, I don’t know what the right answer is for an interview, but it’s either going to be all the time or none of it, based on the way you’re framing it.

Katie: Yes. [laugh]. It was basically, like, half an hour. So, when you are a frontend engineer, your job is to write HTML, CSS, and JavaScript. And in five hours, I talked about that for probably half an hour.

It was one small question and one small discussion, and all the rest of the time was algorithms, and data structures, and big O notation, and oh, gosh, I think they even did the whole, like, “I typed something into my browser, tell me what happens after I type a URL into my browser.” And I think that just told [laugh] me everything that I needed to know about how Amazon approached the frontend and why their website was such a hot mess was because they weren’t actually hiring anyone with real frontend skills to work on the frontend. They were hiring backend people who probably—not to say that they weren’t capable or didn’t care, but I don’t know. That’s my favorite Amazon story that I have is trying to go work there, and they basically were like, “Yes, we want a frontend engineer.” And then they didn’t actually ask about any frontend engineering skill sets in the job. They didn’t offer me anyth—I don’t think I got invited to go to Seattle, but I probably wouldn’t have anyways.

Corey: No. Having done it a couple of times now, again, I like the people I meet at Amazon very, very much. I want to be very clear on that. But some of their processes on the other hand, oh, my God. It shows that being a big company is clearly not necessarily a signal that you solved all of these problems. In some cases, you’re basically just crashing through the problem space by sheer power of inertia.

Katie: Yeah, definitely. I think you can see that when looking at their frontend. Harkening back a little bit to what we were talking about earlier is you don’t go to Amazon and learn patterns of interaction that are applicable to every single site on the web. Amazon kind of expects that users are going to learn the Amazon way of shopping and that users are going to adjust how they navigate the web in order to accommodate Amazon. You know, people learn, “Oh, this is what I do on Amazon.” And then, you know, they’re—

Corey: Oh, that’s the biggest problem with bad user experience is people feel dumb.

Katie: Mm-hm.

Corey: They don’t think, “This company sucks at this thing.” They think, “I must not get it.” And I know this, and I am subject to it. I run into this problem all the time myself.

Katie: Oh, yes.

Corey: And that is a problem.

Katie: Yeah. It’s why I think, like you said earlier, it’s so important when you work somewhere to figure out how do you get that distance between being a power user enough so that you can understand and appreciate what it’s like for a regular user who’s not a power user of your site. And what do they do? And UX researchers are amazing. A good UX researcher is worth absolutely their weight in gold because, I don’t know if you’ve ever sat in on a UX session where the researcher is walking a user through completing a specific task on a website, but oh my God, it’s painful.

It’s because [laugh] you just want to, you want to push them in the right direction, and you want to be like, “Oh, but what about in the upper right over there, that big orange button,” and you can’t do that. You can’t push people. You have to be very open-ended, you have to ask them questions. And every single time I’ve listened in on a UX research recording, or a call, I want to scream through the computer and be like, “Oh, my gosh. This is how you do it.”

But, you know, you can’t do that. So, [laugh] I think it’s important to try to develop that kind of skill set on your own of, “Okay, if I didn’t stare at this website every day, what would it be like for me to try to navigate? If I was using a keyboard for navigation or a screen reader instead of a mouse, what would my experience be like?” Having that empathy, and that ability to get outside of yourself is just really important to be a successful engineer on the web, I think.

Corey: Yeah. And you really wish, on some level, that they would be able to articulate this as an industry. And I say ‘they,’ I guess I’m speaking of about three companies in particular. I have a lot more sympathy for a small startup that is having problems with UX than I am for enormous companies who can basically hurl all the money at it. And maybe that’s unfair, but I feel like, at some point of market dominance, it is beholden on you to set the shining example for how these things are going to work.

I don’t feel that way, necessarily about architecture on the backend. Sure, it can be a dangerous, scary tire fire, but that’s not something your customers or users need to think about or worry about, as long as it is up from their perspective. UX is very much the opposite of that.

Katie: Totally. And I think, working at a former startup, there’s a tendency to really focus a lot on those backend problems. You know, you really look at, “Okay, we’re going to nitpick every single RPC request. We’re going to have all kinds of logging and monitoring about, okay, this is the time that it takes for a database API request to return.” And just the slightest movement and people freak out.

But it’s been a process that I’ve been working really hard on the last couple of years, to get folks to have that same kind of care and attention to the stuff that they ship to the frontend, especially for a lot of organizations that really focus on, “Well, we’re a tech company,” it’s easy to get into this, oh, engineering is all of these big hard systems problems, when really your customers don’t care about all of that. Yes, ultimately, it does affect them because if your database calls are really, really slow, then it has an effect on how quickly the user gets a response back and we know that slow-performing websites, folks are more likely to abandon them. Not that it doesn’t matter completely, but personally, I would really love it to see more universally around the industry that frontend is seen as this is the entirety of your product and if you get that wrong, then none of the rest of your architecture, or your infrastructure, or how great your DevOps is matters because you need customers to come to your site and buy things.

Corey: It turns out that the relationship between customers coming to your site and buying things and the salaries engineering likes to command is sometimes attenuated in ways that potentially shouldn’t be. These are interesting times, and it does help to remember the larger context of the work we do, but honestly, at some point, you wind up thinking about that all the time, and not the thing that you’re brought in specifically to fix. These are weird times.

Katie: Yes.

Corey: Katie, thank you so much for taking the time to speak with me about several things. Usually—it’s weird. Normally, when someone says thank you for speaking to me about Git, there is no way that isn’t a sarcastic—

Katie: [laugh].

Corey: —statement. But in this case, it is in fact genuine.

Katie: Yes, I will bitch about Git until I am blue on the face, so I appreciate you having me on board to talk about it, Corey. Thank you.

Corey: Of course. If people want to learn more, where can they find you?

Katie: They can find me at ohshitgit.com, or as you pointed out, the dangitgit.com swear-free version. As a little plug for the site, we now have had the site translated by volunteers in the community into 28 different languages. So, if English is not your first language, there’s a really good chance you’ll find a version of OSG—as I like to call it—that is in your language.

Corey: Terrific. And we will, of course, put links to these wonderful things in the [show notes 00:39:16]. Thank you so much for taking the time to speak with me. I really appreciate it.

Katie: Thank you, Corey. It’s been lovely to reconnect, and gosh, look at where we are now compared to where we were almost five years ago.

Corey: I know. It’s amazing how the world works.

Katie: Really.

Corey: Katie Sylor-Miller, frontend architect at Etsy. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star
review on your podcast platform of choice along with a comment written in what is clearly your preferred user interface: raw XML.

Katie: [laugh].

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

This has been a HumblePod production. Stay humble.

View Details

About EricEric Dynowski, Managing Partner and Chief Solutions Officer at Deft, has been developing software, designing global infrastructures, and managing large technology installations for over 20 years. His background in complex infrastructure design and integration has helped him reduce customer budgets by millions.

Links:

  • Deft: https://www.deft.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: You could build you go ahead and build your own coding and mapping notification system, but it takes time, and it sucks! Alternately, consider Courier, who is sponsoring this episode. They make it easy. You can call a single send API for all of your notifications and channels. You can control the complexity around routing, retries, and deliverability and simplify your notification sequences with automation rules. Visit courier.com today and get started for free. If you wind up talking to them, tell them I sent you and watch them wince—because everyone does when you bring up my name. Thats the glorious part of being me. Once again, you could build your own notification system but why on god’s flat earth would you do that?

Corey: This episode is sponsored in part by Thinkst. This is going to take a minute to explain, so bear with me. I linked against an early version of their tool, canarytokens.org in the very early days of my newsletter, and what it does is relatively simple and straightforward. It winds up embedding credentials, files, that sort of thing in various parts of your environment, wherever you want to; it gives you fake AWS API credentials, for example. And the only thing that these things do is alert you whenever someone attempts to use those things. It’s an awesome approach. I’ve used something similar for years. Check them out. But wait, there’s more. They also have an enterprise option that you should be very much aware of canary.tools. You can take a look at this, but what it does is it provides an enterprise approach to drive these things throughout your entire environment. You can get a physical device that hangs out on your network and impersonates whatever you want to. When it gets Nmap scanned, or someone attempts to log into it, or access files on it, you get instant alerts. It’s awesome. If you don’t do something like this, you’re likely to find out that you’ve gotten breached, the hard way. Take a look at this. It’s one of those few things that I look at and say, “Wow, that is an amazing idea. I love it.” That’s canarytokens.org and canary.tools. The first one is free. The second one is enterprise-y. Take a look. I’m a big fan of this. More from them in the coming weeks.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. For a while I’ve been talking about how I started The Duckbill Group as a highly niche, highly focused consultancy aimed at one very expensive problem—the AWS bill—and that’s all we do. We don’t do implementation; we simply do the thing that it says on the tin. And it turns out that you can get fairly good at the problem like that in not a tremendous amount of time.

And today’s promoted episode of Screaming in the Cloud. My guest is Eric Dynowski, Chief Solutions Officer and Partner at Deft, which is a consulting company that is almost inverted, in the sense that they do an awful lot of stuff. And we’re going to argue about it now. Eric, thank you for taking the time to speak with me today.

Eric: Great to be here, Corey. Can’t wait to dig in. [laugh].

Corey: So, when I started this place back almost five years ago now, I was coming out of an engineering career that I found deeply unsatisfying. I had a bunch of working with computers skill sets, and I wanted to apply them to something that I could use as an independent consultant—because who would ever build a team or a company out of this?—that I could wind up doing repeatedly, so I could, you know, make money, not have a boss, and ideally work less than 80 hours a week. And fixing the AWS bill was the first one I tried; it turned out that, yeah, there is, in fact, an awful lot of expensive business problem hidden in there. And that’s what I’ve been focusing on ever since. I get the distinct impression that your story doesn’t quite sound like that one.

Eric: Well, it’s a little bit like that one, in the sense that I was once an engineer and a technician doing things, and had the crazy idea to go off and start a company that consulted and helped people with their technology needs. And maybe we’re a little bit alike in the sense that we also have a very singular focused offering—I’m going to tease you a little bit here, Corey—is that we do one thing: we provide technology solutions for our customers. But probably what you are getting at is, “Well, that’s a really broad statement, Eric, what is a technology solution?”

Corey: It is. And the reason I did this is because my marketing budget was 50 bucks. So, I wanted something that the Rolodex effect is what I was after, where when someone says at a party to someone else, “Well, I have a problem: my AWS bill.” I want the answer to be, “Hey, I have someone for you to talk to.” You don’t generally tend to see that with more broad statements around positioning most of the time.

Eric: Sure. Sure.

Corey: I also felt like if I was going to be, all right, I’m the best cloud architect advisor ever. Great, now I’m competing against folks like you and the giant consultancies that are in every country. And honestly, those folks have better airport ads than I’m going to be able to put up, at least at the time. Now, I have a platypus for a mascot. So, one wonders whether that would still hold true. You took a very different approach and have done fantastically—

Eric: Yeah.

Corey: —well with it.

Eric: Yeah, I mean, maybe a little bit of history here is helpful. When I started my company, which was Turing Group back in 2013, we did actually focus pretty tightly just on AWS, and that’s all we did. We wanted to help customers get in AWS, fix their problems in AWS, scale in AWS, manage their costs, all these sorts of things. And along the way, we had customers coming back to us saying, “We love you guys. You’re doing great in [laugh] AWS, but I have other needs too.”

And I started saying, “Well, I know a guy over there that does that.” And, “I’ve got a good friend over here at this company that does it.” And, you know, we’re referring back and forth. And kind of parallel to that, one of my business partners who is running a company called Server Central was having the exact inverse problem. And, you know, they were providing managed services within the data center, managed networking capability, things of that nature, you know, helping customers build out an infrastructure to operate at scale, where it’s like, “Oh, we need 1000 servers, and we need them really inexpensive, and we have to be able to manage them at scale.”

And they were crushing it, doing a great job except his customer start getting back to him saying, “Love what you guys are doing. We’re also using AWS, and we want some help with that.” And so, he would pick up the phone and call me and say, “Eric, can you guys help here?” And in some sense, we were competitors. I wanted to move everybody out of the data center into the cloud.

And he wanted to move everyone out of the cloud back into the data center or keep them in the data center. And it was like, “Okay, this is weird, but we both have the same problem.” And so we went out to lunch and started this conversation of like, “Well, what if we weren’t competitors?” [laugh].

Corey: “Maybe there’s alignment here.” Yeah. I think Ben Franklin once said that three moves is as good as a fire when it comes to cleaning out old cruft. And migrations are like that. I do want to call out that since I make a frequent practice of saying that multi-cloud as a best practice is foolish, I want to be clear that is in the absence of other constraints.

If you’re building something new, you probably should pick a provider and go all-in. I don’t care which one you do—you might; I don’t—but beyond that, at a significant point of scale, when a company says, “All right, we’re in one provider—or a data center—and we’re going to move to a cloud or other clouds.” Generally, they’re correct. They have context that I don’t when I’m speaking in the general case. I am not anti-multi-cloud; I am anti-multi-cloud when it is foolish and when it’s badly done. Just to set my bias out there and, ideally, avoid getting some letters.

Eric: You know, I tend not to disagree with that in the right context as well. When we were doing just AWS only, I think I would have argued that multi-cloud, yeah, there’s no place for it. If you’re going multi-cloud, you’re giving up on all of the greatness that a single public cloud has to offer. You know, and by the greatness, I mean, the proprietary services they offer, and the APIs, and things like SNS and SQS and Route 53, all of those things that you could build into your application and just start using them without having to build all the infrastructure to run it. And so I would agree with you, I think, in that sense.

But I wish the world were that simple and I wish the companies that we worked with operated in a nice, clean, unambiguous context. But the more you dig in, I think you realize that when you start dealing with a company, maybe, that has 20,000 employees and offices all across the country, and 25 years of legacy applications—or maybe a vision for the future that just is so massive that it requires a different point of view—and this is really where we engage. And it’s interesting that you mentioned about the context, and usually, if a customer approaches us and says, “We want to go multi-cloud,” the first thing we do is put the brakes on and get away from the discussion around multi-cloud and move straight into, “What are you trying to accomplish here?” And navigate that conversation before it turns into—you know, let’s not start with the solution type of discussion.

Corey: Part of the problem, it seems, that when you start talking to folks about these things, especially in some vendor corners, an awful lot of self-interest that winds up informing the answers that immediately come from it, where it’s, “Oh, yeah, you want to go multi-cloud.” And you scratch underneath the surface and the reason is that if you go all-in on one provider, they have nothing left to sell you, in some cases. Other times, it’s by a cloud provider themselves pushing multi-cloud strongly because they know that if you go all-in on one provider, it will absolutely not be them. So, the hard part is finding someone who can serve as a trusted advisor. And I mean that in the actual sense of a person, not the crappy AWS service that tells you everything’s fine when it isn’t. That’s ‘Plausible Advisor’ at best. Let’s be clear here.

Eric: Well, come on. If it wasn’t for Trusted Advisor, we wouldn’t have a market for all the third-party analysis tools [laugh] and people like you to help us manage our costs.

Corey: Believe me, if I thought a tool could solve the problem, I would have built it years ago. Tried, turned out it didn’t, and well, here we are. You position what you do at Deft as starting from an advisory perspective. You’re not pushing a particular product, you’re not pushing any particular vendors that I’m aware of, you have a partner list that is not a single vendor, what a concept. So, it’s clear that you’re doing something that goes one of two ways.

Either it is, yeah, we’ll take money from anyone who will pay us, or alternately you’re approaching it from a thoughtful perspective of trying to figure out what’s going on with the customer. Based upon our conversations, I’m going to go ahead and guess it’s that one.

Eric: It definitely is the latter, Corey. We’ve certainly had many customers approach us over the years and ask for things that are a bad fit. And a bad fit might be, they’re asking for technology experience we don’t have. If someone came to us and said, “We want you guys to be Oracle DBAs because you do technology,” our answer is very likely going to be no. Could we learn to be Oracle DBAs? Yeah, we probably could. Do we want to probably not? Maybe?

Corey: There are some very qualified Oracle DBAs out in the world, and it’s—

Eric: There’s—

Corey: —great.

Eric: —there’s places for people to specialize. And I think one of our virtues is to know and recognize where we belong and where we don’t belong. And the good thing is, not only have we sort of built up our own partner list in terms of technology partners, strategic partners, but we also have our own internal list of referral partners where we know that something’s out of our wheelhouse, and I got another company and another team here that I know can crush it and help them out. There’s other areas of work that we just don’t get into. If you want to outsource your IT and have a company that’s going to help you figure out why your printer is not working, definitely not us.

You’re going to be wasting your money with a firm like us. Or maybe you want to partner in a way that isn’t going to take the best advantage of the capabilities that we have, meaning you just want to take advantage of us in a halfway manner, and you want to keep an internal team and the two teams are up against one another and fighting about stuff constantly. We need to have good strong trust between our two companies and our partnerships. And so if we feel like there isn’t an opportunity for that, we might walk away from it as well.

So, it’s not a case of everything that walks in front of us, here’s a proposal. [laugh]. We definitely do some opportunity vetting and analysis. We ask a lot of questions upfront about how our customers work, what their internal teams are like, what their expectations of us are, what they want in that relationship. Is it transactional or is it strategic? We’re interested in the strategic partnerships with our customers.

Corey: I think this is something that is not well understood by a lot of the fly-by-night folks, for lack of a better term. I don’t mean to sound disparaging, but the folks who don’t seem to understand that long-term reputation is important. I mean, both of our consulting companies, although radically different in focus, have pages on our site where we list reference customers. In fact, there’s some overlap between our customers. And as we look at this, you aren’t allowed to put a customer logo up if they’re going to take umbrage to you doing it, first off. And secondly, you don’t want to put a customer logo up if people are going to ask them about their experience with you, and the response is, “Oh, they were crap.” At some point, no, let’s not do that.

Eric: [laugh]. That’s right.

Corey: The only way to get there is to deliver on an engagement in such a way that the response is, “That was great. Would you do it again?” “Can I?” And the idea of excitement of delivering an outcome where people who you’ve worked with become some of your biggest advocates, that’s how I always viewed the proper way of building a business.

Eric: Yeah. Now, I’d ask you, in terms of when you guys are providing advice to your customers about AWS spend, or someone approaches you, obviously, you’re probably first thinking, “Okay, well, how much AWS spend do you have in a month and is this worth my time?” But there’s probably another element of evaluation that you must do in terms of is this a good customer for us, and can we do the right thing for them? What are pieces that you guys think about?

Corey: Oh, absolutely. As a general rule, we do a lot of AWS contract negotiation. And that is, if you have an AWS offer in front of you for committing to something, come talk to us; it’s fun. That’s half of our engagements today. The other half are cost optimization projects, and generally speaking, we aren’t going to be able to effectively deliver return on investment for much less than about a million bucks a month in spend.

So, I do at some point want to explore how to help people who are not already paying a king’s ransom to AWS every month, but that is down the road. The next step is a conversation. It’s a, “So great, you want to optimize your AWS bill.” And then my favorite is the—I get the quote-unquote, “Dumb question.” “Why? Why do you care? Why is this an actual problem other than it looks like a phone number and your CFO has some questions, what is the actual concern?”

Very often, we’ll find that it’s not that you’re spending too much on cloud, in many cases, it’s that it’s not understood what it’s doing. “Okay, the bill is 20% higher this month. Is that new normal? Is that something that’s going to inform our planning and we do adjust our expectations for what this is going to cost to run this? Is it just a mistake that someone left up?” The same questions would have arisen if the bill were 20% lower, except somehow it never is.

Eric: Right, right. You know what’s interesting about that, Corey is, too, also I think that you tend to tease AWS from time to time. And also I think the work you do would not be in Amazon’s interest, right? They want customers to spend more money on their infrastructure and their services and their capabilities, and you’re helping customers spend less. But what’s interesting about that is that we’re one of the few managed service partners in the country.

And I don’t know how much you know about that program, and what it takes to get into that program and to maintain the certification in that program. It’s like a three-day audit; there’s 500 control items that we have to go through. In fact, it was that program that took our business to the next level. It was that program and its rigor that took what we were doing and actually matured it and turned it into what I would consider a respectable world-class operation. But one of the interesting aspects of that audit is that there’s several control items in there that ask us to show Amazon that we are taking steps to manage our customers’ costs on AWS and reduce spend. And it’s interesting that Amazon is the one pushing that on us and instilling that requirement as we support our customers in Amazon.

Corey: It’s counterintuitive, but this is one of those areas where there’s no one on the other side of this issue. Of course, Amazon wants customers to spend more money with Amazon—I swear the company spends half its time lying awake at night worrying someone who isn’t them is making money somewhere, at least that’s how it feels some days. But they want that spend to be intelligent. They don’t want the narrative to be that the cloud is just as expensive a bunch of nonsense. If there’s a bunch of instances that are sitting there idle, they will advise you—if they’re on top of the game—to turn them off because that is the goal. They want it to be—

Eric: That’s how they deliver on their promise. Right.

Corey: Well, yeah. Pandemic aside, with most of our customers, what we notice a year after an engagement is that they’re in fact spending more than they were when we started. But it’s more efficient; it’s growth that’s tying into this. It becomes a component of cost of goods sold where, “Yeah, we’re doing more business, so it costs us more to fulfill that business; we’re perfectly happy,” is generally the response to that. And I think that everyone with a vision that extends beyond this quarter’s numbers is likely to start to get into that, on some level.

Eric: Right.

Corey: One thing I do want to ask you is—relevant to what you just said—one thing that we do at The Duckbill Group explicitly is we have no partners, full stop. And the reason behind that is because with what we’re doing around billing, and money, and contract negotiation, and the rest, as soon as we have a partner, it suddenly gives rise to a bunch of real or perceived conflicts of interest. And in this particular niche, it makes an awful lot of sense not to do that. Now, if you’re in any other arena, where you’re in—“Oh, you’re in security, for example. Oh, we have no partners with any vendors,” the answer question becomes, “Well, what’s wrong with you? What, there’s no one willing to trust you? Do you think somehow you’re better than all these other people?” It’s the wrong answer. So, my question for you is, how do you evaluate whether you should partner with a particular company or not?

Eric: Sure. Great question. Deft’s reason for existence—when we think about ourselves, we reframe it as our purpose—is to deliver on the promise of technology. And if you unpack that statement a little bit is like, “Well, what the heck is the promise of technology? What does that even mean?”

And what we get down to is that technology itself doesn’t make any kind of promises. That router you just bought, it doesn’t promise really anything; that EC2 instance you just booted in AWS doesn’t really ultimately, at the end of the day, promise anything. It’s incapable of making promises, but people are. And what we promise is that we can wrangle that technology, we can configure it, we can set it up in a particular kind of way, we can bring in the right components into the solution, and deliver on a promise of, “Yes, you can scale to 100 million users,” or, “Yes, you can reduce latency and improve the customer experience for your customers.” It’s all about the people, and that’s what we have the most of.

And that’s the best thing that we have in our house.s we have an inventory of highly qualified, talented, empathetic, compassionate, excited people. So, when we start thinking about our partners and who we want to partner with, what we take into consideration is what technology, tools, and capabilities do our people need to have in their toolbox, such that when we start working with our customers crafting that promise and that solution, we’ve got the right things at hand and at the right time. And then the second piece of it is, does the partner align well with us in terms of our vision? And in some sense, keeping us relatively technology-neutral, in the same sense that you’re trying to stay neutral from that billing perspective and making sure that you’re looking out and advocating for your customers first.

So, when we’re thinking about our partners as well, it’s not that, oh, well, we want to build our whole business on top of AWS, or Azure, or in our data center. And those are the on—you know, we try to remove dogma from the picture in that sense, and try to probably be dogmatic mostly about the customer and what it is that they’re trying to achieve, and being honest with them. So, it’s more of a, “Hey, let’s scan the horizon. Let’s listen to our customers, let’s understand what problems they’re trying to solve, what challenges they have today. Let’s evaluate the technology options on the table across the world.” Our partner might be [unintelligible 00:18:36], it might be VM, or it might be Amazon, it might be a small little company somewhere that does a niche service. But our job is to come together with all of those things and present a cohesive solution.

Corey: And it’s clearly working. You were the Turing Group and you wound up partnering with—you said your business partner—who was over at Server Central, which I’m just guessing from the name and assuming I hadn’t paid attention to the industry for a while, sounds like it might not be fully cloud-focused, on some level, given the name. What did they do? And why was merging the right answer?

Eric: Yeah. I mean, it’s funny that you bring that up. Server Central. Wow, a server; who’s talking about servers anymore, right? It’s containers and virtualization and—

Corey: But they’ve got to run somewhere.

Eric: That’s right. Serverless applications. Like, hmm, is this the right name? And I think it speaks to the 20-year heritage that Server Central has had and how they built their business. And they do and did, and we do have a cloud focus that’s not related to the public cloud.

We have a significant number of customers that operate on private clouds that we’ve built for them and manage for them, for various reasons. Some are legit and some maybe not so legit and mostly about how they feel about something. And some of them are technically driven. And after we brought the companies together, we realized that hey, you know what, we have a lot of brand equity and history and Server Central and we have to respect that. Turing Group had its own set of brand equity in the market that we had established and promoted a certain kind of ideology and thinking.

And so for a short period of time, we were a little bit unsure of how are we going to bring this together in any kind of cohesive fashion? And it actually went out into the world for about a year as Server Central Turing Group. And I think my tongue twisted as I said it, [laugh]. It’s a lot of words and it mostly just confuses people and makes them scratch their head. And so we went off on a journey to figure out what our reimagined new company is going to look like with all the combined services and capabilities that we have.

And that’s how we arrived at Deft and the idea that’s how we want to engage with our customers; that’s what we want our solutions to be like; that’s what we want the experience in working with us to be. And we want to remove the friction and anxiety that technology can bring. It reminds me of when I was starting my first company. I spent a lot of time sort of navel-gazing, saying, “What do I like about this, and why am I doing this?” And went back to my early days as an engineer, and at the core of it was a really simple idea and it was the idea that when I helped somebody with a technology problem, they were elated. They thought it was magic. They thought it was black magic.

They didn’t understand how I took this goofy, strange, cryptic thing and made it do what they wanted, and I did it quickly and I did it deftly. And there was joy and they were happy and I loved that; I loved that response. I loved knowing that I helped fix this mysterious problem for somebody that just didn’t even know where to begin. And I did it time and time again and it helped me grow my career. And when we started the first company, it was sort of like, I want to continue that feeling.

I want to create that feeling for our customers where they feel like maybe they’re stuck with some crazy complex technology problem, and because I happen to have the innate skill for understanding these things and figuring these things out and I have a team that can do it, we can create that same feeling for our customers. And we want to continue doing that today.

Corey: This episode is sponsored by our friends at Oracle HeatWave is a new high-performance accelerator for the Oracle MySQL Database Service. Although I insist on calling it “my squirrel.” While MySQL has long been the worlds most popular open source database, shifting from transacting to analytics required way too much overhead and, ya know, work. With HeatWave you can run your OLTP and OLAP, don’t ask me to ever say those acronyms again, workloads directly from your MySQL database and eliminate the time consuming data movement and integration work, while also performing 1100X faster than Amazon Aurora, and 2.5X faster than Amazon Redshift, at a third of the cost. My thanks again to Oracle Cloud for sponsoring this ridiculous nonsense.

Corey: I do want to point out that, at least in my mind, there’s always a little bit of, I guess, we call it technical elitism, on some level, where, “Oh, someone is working through a partner. They must be a company that’s stuck in the past.” But a glance at companies that you’re working with, make it very clear that’s not the case. I mean, Ars Technica, New Relic, Wildbit. You’ve got some companies that are very forward-looking, and by no definition are these companies that don’t understand cloud or understand how the internet works. It’s something that I think is not fully understood among a subset of the industry that, in many cases, having a third-party partner, in many cases winds up helping you go faster, further.

Eric: Yeah. It might be cliche to say—oftentimes, in cliches, there’s a little bit of truth—which is, focus on what you’re good at, and focus on what you’re best at, and focus on your core products. And with a lot of these technology companies where you might read on the surface that, “Oh, yeah, they’re a smart bunch over there. Why do they need a partner?” Technology is complicated. The stack is deep.

Whether you’re talking about deployment pipelines, or should I use a fiber connection on this or should I use copper? Or should we have jumbo frames enabled? Or should we be using API Gateway and Lambda functions for this? I just listed a broad range of technologies and things that solve different problems. And these customers have their own products that they have to put out into the world; those products need to be meaningful and thoughtful and aligned with their customers.

And because that technology stack is complex and deep, it creates an opportunity for companies like ours—for partners—to step in, and grab a piece of that complexity, and manage it, and handle it, and help a customer with it to create the space for them to create the most excellent product. And so even though they are technology companies because you’re managing this big wrangling layered technologies, abstractions, and—well, even when we talk about containerization, right, and running a small application, there’s seven layers before you get to the CPU. [laugh]. And within that seven layers, there’s I don’t know how many lines of code, and there’s how many hidden assumptions and configuration files, and you name it. And there’s areas of that entire stack that we’re really good at and customers derive value from that.

Corey: One area that you’ve been relatively active within is the hybrid universe. My talking point on that has generally looked a lot like the snarky take of, “Well, you have a company in a data center today, and they’re going to go all-in on the cloud. And it turns out halfway through that it’s hard to move some workloads. There is no AWS/400 and they have a mainframe.” So, what are they going to do? They give up halfway, plant a flag, declare victory, and now we’re hybrid as a best practice. That is not entirely accurate, but there’s an element of accuracy in some cases to it.

Eric: Yeah.

Corey: But I don’t get the sense that’s how you see it. I’m left with a strong impression that it’s a very intentional choice for some of your customers, that in some cases, workloads that are live in data centers, were at one point living in a cloud provider. Talk to me about that.

Eric: Yeah, I like to think of it not as, like, a binary situation. And something that exists on degrees, and often times has a lot to do with the lifecycle of a product or company and the scale of a company. And we touched on this earlier in our conversation, which is that if someone approached us—maybe a startup or a smaller company—trying to migrate off of half-dozen servers and move into the public cloud, and they approached us and said, “Yeah, we need to be hybrid for this thing,” I would probably question that and I would question it really hard, and say, “Really, what are you going after here? You’re going to give up a lot if you choose to go hybrid, and you won’t really take advantage of some of the amazing opportunities that a full-on single cloud solution has to offer.” On the flip side, we’ve seen companies that started like that, were a hundred percent in public cloud on a single provider, everything’s working fine.

There was no issues whatsoever, except the bill, or maybe a fear of what the bill could be. And this is something that happens at scale. There’s just a point where the public cloud just doesn’t make sense anymore, even despite those benefits. And for the bottom line and in terms of the margins and your cost of revenue, giving up some of those additional benefits that allow you to grow and scale is worth it. And if you look at the technology landscape, there’s a reason Facebook’s not in AWS. [laugh].

There’s a reason a lot of these larger technology companies where we have hundreds of millions of users or bazillions of petabytes of data, that move out and get out of there. I mean, look at Dropbox right? [laugh]. Is Dropbox storing all their data in the public cloud? Not really, and they are a public cloud in a sense on their own, right?

Corey: Yeah, they did just launch a 34-petabyte data warehouse for analytics on AWS, and they’ve made a bunch of big—

Eric: Yeah, yeah. I mean—

Corey: —deals out of that, but the core storage workload, yeah, that does not economically make sense, given their access patterns and how they have built that offering. So yeah, that is a very well understood, very specific, very niche workload. Yeah, that does not belong in AWS.

Eric: We’ve even gone as far as launching our own multi-petabyte managed object storage solution, it’s totally S3 compatible. It works identically to the way S3 works, but we have customers that actually can do better on our platform, either because we can provide lower latency, we can provide custom contracts that aren’t just purely pay as you go; there’s all kinds of different options that we can give our customers that are more custom and tailored to their needs that you’re going to get from, “Here’s your API keys have fun.” [laugh]. And so there’s still a market for that stuff and there’s still a need for that stuff.

Corey: There really is.

Eric: So, the answer is it depends, Corey. [laugh].

Corey: It seems to be the answer to any nuanced question. So, if people want to learn more about what you’re doing at Deft, and potentially whether it might help them with some of the challenges they’re facing, where can they find you?

Eric: Easy. deft.com, D-E-F-T dot com. Great, short four-letter domain name that you wouldn’t believe what we had to go through to get. [laugh].

Corey: I can only imagine. Thank you so much for taking the time to speak with me today. I really appreciate it.

Eric: You’re welcome, Corey.

Corey: Eric Dynowski, Chief Solutions Officer and Partner at Deft. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you hated this podcast, please leave a five-star review on your podcast platform of choice, along with a comment telling me that no, customers should in fact go all-in on your third-rate cloud.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Francesca

Francessca is the leader of the AWS Technology Worldwide Commercial Operations organization. She is recognized as a thought leader of business technology cloud transformations and digital innovation, advising thousands of startups, small-midsize businesses, and enterprises. She is also the cofounder of AWS workforce transformation initiatives that inspire inclusion, diversity, and equity to foster more careers in science and technology.

Links:

  • Twitter: https://twitter.com/Francessca_V
  • LinkedIn: https://www.linkedin.com/in/francesscavasquez/

TranscriptAnnouncer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by “you”—gabyte. Distributed technologies like Kubernetes are great, citation very much needed, because they make it easier to have resilient, scalable, systems. SQL databases haven’t kept pace though, certainly not like no SQL databases have like Route 53, the world’s greatest database. We’re still, other than that, using legacy monolithic databases that require ever growing instances of compute. Sometimes we’ll try and bolt them together to make them more resilient and scalable, but let’s be honest it never works out well. Consider Yugabyte DB, its a distributed SQL database that solves basically all of this. It is 100% open source, and there's not asterisk next to the “open” on that one. And its designed to be resilient and scalable out of the box so you don’t have to charge yourself to death. It's compatible with PostgreSQL, or “postgresqueal” as I insist on pronouncing it, so you can use it right away without having to learn a new language and refactor everything. And you can distribute it wherever your applications take you, from across availability zones to other regions or even other cloud providers should one of those happen to exist. Go to yugabyte.com, thats Y-U-G-A-B-Y-T-E dot com and try their free beta of Yugabyte Cloud, where they host and manage it for you. Or see what the open source project looks like—its effortless distributed SQL for global apps. My thanks to Yu—gabyte for sponsoring this episode.

Corey: And now for something completely different!

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. It’s pretty common for me to sit here and make fun of large cloud companies, and there’s no cloud company that I make fun of more than AWS, given that that’s where my business generally revolves around. I’m joined today by VP of Technology, Francessca Vasquez, who is apparently going to sit and take my slings and arrows in person. Francessca, thank you for joining me.

Francessca: Hi, Corey, and thanks for having me. I’m so excited to spend this time with you, snarking away. I’m thrilled.

Corey: So, we’ve met before, and at the time you were the Head of Solutions Architecture and Customer Solutions Management because apparently someone gets paid by every word they wind up shoving into a job title and that’s great. And I vaguely sort of understood what you did. But back in March of this year, you were promoted to Vice President of Technology, which is both impressive, and largely non-descriptive when one works for a technology company. What is it you’d say it is you do now? And congratulations, by the way.

Francessca: Thank you, I appreciate it. By the way, as a part of that, I also relocated to our second headquarters, so I’m broadcasting with you out of HQ2, or Arlington, Virginia. But my team, essentially, we’re a customer-facing organization, Corey. We work with thousands of customers all over the globe, from startups to enterprises, and we ultimately try to ensure that they’re making the right technology architecture decisions on AWS. We help them in driving people and culture transformation when they decide to migrate onto the cloud.

And the last thing that we try to do is ensure that we’re giving them tools so that they can build cultures of innovation within the places that they work. And we do this for customers every day, 365 days a year. And that’s what I do. And I’ve been doing this for over 20 years, so I’m having a blast.

Corey: It’s interesting because when I talk to customers who are looking at what their cloud story is going to be—not just where it is, but where they’re going—there’s a shared delusion that they all participate in—and I’m as guilty as anyone. I have this same, I guess, misapprehension as well—that after this next sprint concludes, I’m going to suddenly start making smart decisions; I’m going to pay off all of my technical debt; I’m going to stop doing this silly thing and start doing the smart thing, and so on and so forth. And of course, it’s a myth. That technical debt is load-bearing; it’s there for a reason. But foundationally, when talking to customers at different points along their paths, I often find that the conversation that I’m having with them is less around what they should be doing differently from a tactical and execution perspective and a lot more about changing the culture.

As a consultant, I’ve never found a way to successfully do that, that sticks. If I could I’d be in a vastly different, vastly more lucrative consulting business. But it seems like culture is one of those things that, in my experience, has to be driven from within. Do you find that there’s a different story when you are speaking as AWS where, “Yeah, we’re outsiders, but at the same time, you’re going to be running production on us, which means you’re our partner whether you want to be or not because you can’t treat someone who owns production as a vendor anymore.” Does that position you better to shift culture?

Francessca: I don’t know if it positions us better. But I do think that many organizations, you know, all of them are looking at different business drivers, whether that be they want to move to more digital, especially since we’re going through COVID-19 and coming out of it. Many of them are looking at things like cost reduction, some organizations are going through mergers and acquisitions. Right now I can tell you new customer experiences driven by digital is pretty big, and I think what a lot of companies do, some of them want to be the north star; some of them aspire to be like other companies that they may see in or outside the industry. And I think that sometimes we often get a brand as having this culture of innovation, and so organizations very much want to understand what does that look like: what are the ingredients on being able to build cultures of innovation?

And sometimes organizations take parts of what we’ve been able to do here at AWS and sometimes they look at pieces from other companies that they view as north star, and I see this across multiple industries. And I think the one that is the toughest when you’re trying to drive big change—even with moving to the cloud—oftentimes it’s not the services or the tech. [smile]. It’s the culture. It’s people. It’s the governance. And how do you get rallied around that? So yeah, we do spend some time just trying to offer our perspective. And it doesn’t always mean it’s the right one, but it certainly has—it’s worked for us.

Corey: On some level, I’ve seen cloud adoptions stall, in some scenarios, by vendors being a little too honest with the customer, if that doesn’t—

Francessca: Mmm. Mm-hm.

Corey: —sound ridiculous, where it’s—so they take the customer will [unintelligible 00:05:24], reasonable request. “Here’s what we built. Here’s how we want to migrate to the cloud. How will this work in your environment?” And the overly honest answer from a certain provider—I don’t feel the need to name at the moment—is, “Well, great. What you’ve written is actually really terrible, and if you were to write it better, with smarter engineers, it would run great in the cloud. So, do that then call us.”

Surprisingly, that didn’t win the deal, though it was, unfortunately, honest. There was a time where AWS offerings were very much aligned with that, and depending on how you wind up viewing what customers should be doing is going to depend on what year it was. In the early days, there was no persistent storage on EC2—

Francessca: Mm-hm.

Corey: So, if you had a use case that required there had to be a local disk that could survive a reboot, well, that wasn’t really the place for you to run. In time, it has changed, and we’re still seeing that evolution to the point where there are a bunch of services that come out on a consistent, ongoing basis that the cloud-native set will look at and say, “Oh, that hasn’t been written in the last 18 months on the latest MacBook and targeting the developer version of Chrome. Then why would I ever care about that?” Yeah, there’s a bigger world than San Francisco. I’m sorry but it’s true.

And there are solutions that are aimed at customer segments that don’t look anything like a San Francisco startup. And it’s easy to look at those and say, “Oh, well, why in the world would I wind up needing something like that?” And people point at the mainframe and say, “Because of that thing.” Which, “Well, what does that ancient piece of crap do?” “Oh, billions a year in revenue, so maybe show some respect.” ‘Legacy,’ the condescending engineering term for ‘it makes money.’

Francessca: [smile]. Yeah, well, first off, I think that our approach today is you have to be able to meet customers where they are. And there are some customers, I think, that are in a position where they’ve been able to build their business in a far more advanced state cloud-natively, whether that be through tools like serverless, or Lambda, et cetera. And then there are other organizations that it will take a little longer, and the reason for that is everyone has a different starting point. Some of their starting points might be multiple years of on-premise technology.

To your point, you talked about tech debt earlier that they’ve got to look at and in hundreds of applications that oftentimes when you’re starting these journeys, you really have to have a good baseline of your application portfolio. One of my favorite stories—hopefully, I can share this customer name, but one of my favorite stories has been our organization working with Nationwide, who sort of started their journey back in 2017 and they had a goal, a pretty aggressive one, but their goal is about 80% of their applications that they wanted to get migrated to the cloud in, like, three to four years. And this was, like, 319 different migrations that we started with them, 80 or so production cut-overs. And to your point, as a result of us doing this application portfolio review, we identified 63 new things that needed to be built. And those new things we were able to develop jointly with them that were more cloud-native. Mainframe is another one that’s still around, and there’s a lot of customers still working on the mainframe. We work with a very—

Corey: There is no AWS/400 yet.

Francessca: [smile]. There is no AWS [smile] AS/400. But we do have mainframe migration competency partners to help customers that do want to move into more–I don’t really prefer the term modernize, but more of a cloud-native approach. And mostly because they want to deliver new capability, depending on what the industry is. And that normally happens through applications.

So yeah, I think we have to meet customers where they are. And that’s why we think about our customers in their stage of cloud adoption. Some that are business-to-consumer, more digital native-based, you know, startups, of course; enterprises that tend to be global in nature, multinational; ISVs, independent software vendors. We just think about our customers differently.

Corey: Nationwide is such a great customer story. There was a whole press release bonanza late last year about how they selected AWS as their preferred cloud provider. Great. And I like seeing stories like that because it’s easy on some level—easy—to wind up having those modernized startups that are pure web properties and nothing more than that—not to besmirch what customers do, but if you’re a social media site, or you’re a streaming video company, et cetera, it feels differently than it does—oh, yeah, you’re a significantly advanced financial services and insurance company where you’re part of the Fortune 100. And yeah, when it turns out that the computers that calculate out your amortization tables don’t do what you think they’re going to do, those are the kinds of mistakes that show. It’s a vote of confidence in being able to have a customer testimonial from a quote-unquote, “More serious company.” I wouldn’t say it’s about modernization; I’d say it’s about evolution more than anything else.

Francessca: Yeah, I think you’re spot on, and I also think we’re starting to see more of this. We’ve done work at places like GE—in Latin America, Itaú is the bank that I was just referring to on their mainframe digital transformation. Capital One, of course, who many of the audience probably knows we’ve worked with for a long time. And, you know, I think we’re going to see more of this it for a variety of reasons, Corey. I think that definitely, the pandemic has played some role in this digital acceleration.

I mean, it just has; there’s nothing I can say about that. And then there are some other things that we’re also starting to see, like sustainability, quite frankly, is becoming of interest for a lot of our customers as well, and as I mentioned earlier, customer experience. So, we often tend to think of these migration cloud journeys as just moving to infrastructure, but in the first part of the pandemic, one of the interesting trends that we also saw was this push around contact centers wanting to differentiate their customer experience, which we saw a huge increase in Amazon Connect adoption as well. So, it’s just another way to think about it.

Corey: What else have you seen shift during the pandemic now that we’re—I guess, you could call it post-pandemic because here in the US, at least at this time of this recording, things are definitely trending in the right direction. And then you take a step back and realize that globally we are nowhere near the end of this thing on a global stage. How have you seen what customers are doing and how customers are thinking about things shift?

Francessca: Yeah, it’s such a great question. And definitely, so much has changed. And it’s bigger than just migrations. The pandemic, as you
rightfully stated, we’re certainly far more advanced in the US in terms of the vaccine rollout, but if you start looking at some of our other emerging markets in Asia Pacific, Japan, or even AMEA, it’s a slower rollout. I’ll tell you what we’ve seen.

We’ve seen that organizations are definitely focused on the shift in their company culture. We’ve also seen that digital will play a permanent fixture; just, that will be what it is. And we definitely saw a lot of growth in education tech, and collaboration companies like Zoom here in the US. They ended up having to scale from 10 million daily users up to, like, 300. In Singapore, there is an all-in company called Grab; they do a lot of different things, but in their top three delivery offerings—what they call Grabfood, Grabmart, and GrabExpress—they saw, like, an increase of 30% user adoption during that time, too.

So, I think we’re going to continue to see that. We’re also going to continue to see non-technical themes come into play like inclusion, diversity, and equity in talent as people are thinking about how to change and evolve their workforce. I love that term you used; it’s about an evolution: workforce and skills is going to be pretty important. And then globally, the need around stronger data privacy and governance, again, is something else that we’ve started to see in a post-COVID kind of era. So, all industries; there’s no one industry doing anything any different than the others, but these are just some observations from the last, you know, 18 months.

Corey: In the early days of the pandemic, there was a great meme that was going around of who was the most responsible for your digital transformation: CIO, CTO, or COVID-19?

Francessca: [smile].

Corey: And, yeah, on some level, it’s one of those ‘necessity breeds innovation’ type of moments. And we’re seeing a bunch of acceleration in the world of digital adoption. And I don’t think you get to put the genie back in that particular bottle in a bunch of different respects. One area that we’re seeing industry-wide is talent discovering that suddenly you can do a whole bunch of things that don’t require you being in the same eight square miles of an earthquake zone in California. And the line that I heard once that really resonated with me was that talent is evenly distributed; opportunity is not. And it seems that when you see a bunch of companies opening up to working in new ways and new places, suddenly it taps a bunch of talent that previously was considered inaccessible.

Francessca: That’s right. And I think it’s one of those things where—[smile] I love the meme—you’ll have to send me that meme by the way—that just by necessity, this has been brought to the forefront. And if you just think about the number of countries that, sort of, account for almost half the global population, there’s only, like, we’ll say eight of them that at least represent close to 60-plus percent. I don’t think that there’s a company out there today that can really build a comprehensive strategy to drive business agility or to look at cost, or any of those things digitally without having an equally determined workforce strategy. And that workforce strategy, how that shows up with us is through having the right skills to be able to operate in the cloud, looking at the diversity of where your customer base is, and making sure that you’re driving a workforce plan that looks at those markets.

And then I think the other great thing—and honestly, Corey, maybe why I even got into this business—is looking at, also, untapped talent. You know, technology’s so pervasive right now. A lot of it’s being designed where it’s prescriptive, easier to use, accessible. And so I also think we’re tapping into a global workforce that we can reskill, retrain, in all sorts of different facets, which just opens up the labor market even more. And I get really excited about that because we can take what is perceived as, sort of, traditional talent, you know, computer science and we can skill a lot of people who have, again, non-traditional tech backgrounds. I think that’s the opportunity.

Corey: Early on in my career, I was very interested in opening the door for people who looked a lot like me, in terms of where their experience level was, what they’d done because I’d come from a quote-unquote, non-traditional background; I don’t even have a high school diploma at this point. And opening doors for folks and teaching them to come up the way that I did made sense for a while. The problem that I ran into pretty quickly is that the world has moved on. It turns out that if you want to start working in cloud in 2021, the path I walked is closed. You don’t get to go be an email systems administrator who’s really good at Unix and later Linux as your starting point because those jobs don’t exist the way that they once did.

Before that, the help desk roles aren’t really there the way that they once were either, and they’ve become much more systematized. You don’t have nearly as much opportunity to break the mold because now there is a mold. It used to be that we were all these artisanally crafted, bespoke technologists. And now there are training curriculums for this. So, it leads to a recurring theme on the show of, where does the next generation really wind up coming from?

Because trying to tell people to come up the way that I did is increasingly reminiscent of advice of our parents’ generation, “Oh, go out and pound the bricks, and have a firm handshake, and hand your resume to the person at the front desk, and you’ll get a job today.” Yeah, sure you will. How do you see it?

Francessca: You know, I see it where we have an opportunity to drive this talent, long-term, in a variety of different places. First off, I think the
personas around IT have shifted quite a bit where, back in the day, you had a storage admin, a sysadmin, maybe you had a Solaris, .NET, Linux developer. But pretty straightforward. I think now we’ve evolved these roles where the starting point can be in data, the starting point can be in architecture.

The personas have shifted from my perspective, and I think you have more starting points. I also think our funnel has also changed. So, for people that are going down the education route—and I’m a big proponent of that—I think we’re trying to introduce more programs like AWS Educate, which allows you to go and start helping students in universities really get a handle on cloud, the curriculum, all the components that make up the technology. That’s one. I think there are a lot of people that have had career pivots, Corey, where maybe they’ve taken time out of the workforce.

We disproportionately, by the way, see this from our female and women who identify, coming back to the workforce, maybe after caring for parents or having children. So, we’ve got—there are different programs that we try to leverage for returners. My family and I, we’ve grown up all around the military veterans as well, and so we also look at when people come out of, perhaps in the US, military status, how do we spend time reskilling those veterans who share some of the same principles around mission, team, the things that are important to us for customers. And then to your point, it’s reskill, just, non-traditional backgrounds. I mean, a lot of these technologies, again, they’re prescriptive; we’re trying to find ways to make them certainly more accessible, right, equitable sort of distribution of how you can get access to them.

But, anyone can start programming in things like Python now. So, reskill non-traditional backgrounds; I don’t think it’s just one funnel, I think you have to tap into all these funnels. And that’s why, in addition to being here in AWS, I also try to spend time on supporting and volunteering at nonprofit companies that really drive a focus on underserved-based communities or non-traditional communities as different pathways to tech. So, I think it’s all of the above. [smile].

Corey: This episode is sponsored in part by CircleCI. CircleCI is the leading platform for software innovation at scale. With intelligent automation and delivery tools, more than 25,000 engineering organizations worldwide—including most of the ones that you’ve heard of—are using CircleCI to radically reduce the time from idea to execution to—if you were Google—deprecating the entire product. Check out CircleCI and stop trying to build these things yourself from scratch, when people are solving this problem better than you are internally. I promise. To learn more, visit circleci.com.

Corey: Yeah, I have no patience left, what little I had at the beginning, for gatekeeping. And so much of technical interviewing seems to be built around that in ways that are the obvious ones that need not even be called out, but then the ones that are a little bit more subtle. For example, the software developer roles that have the algorithm questions on a whiteboard. Well, great. You take a look at the average work of
software development style work, you don’t see those things coming up in day-to-day. Usually.

But, “Implement quicksort.” There’s a library for that. Move on. So, it turns out that biases for folks who’ve recently had either a computer science formal education or computer science formal-like education, and that winds up in many ways, weeding people out have been in the workforce for a while. I take a look at some of the technical interviews I used to pass for grumpy Unix sysadmin jobs; I don’t remember half of the terminology.

I was looking through some my old question lists of what I used to ask candidates, and I don’t remember how 90% of this stuff works. I'd have to sit there and freshen up on it if I were to go and take a job interview. But it doesn’t work in the same way. It’s more pernicious than that, though, because I look at what I do and how I approach it; the skills you use in a job interview are orthogonal, in many cases, to the skills you’ll need in the workforce. How someone performs with their career on the line at a whiteboard in front of a few very judgy, judgy people is not representative of how they’re going to perform in a collaborative technical environment, trying to solve an interesting problem, at least in my experience.

Francessca: Yeah, it’s interesting because in some of our programs, we have this conversation with a lot of the universities, as well, in their curriculums, and I think ultimately, whether you’re a software developer, or you’re an architect, or just in the field of tech and you’re dealing with customers, I think you have to be very good at things like problem-solving, and being able to work in teams. I have a mental model that many of the tech details, you can teach. Those things are teachable.

Corey: “Oh, you don’t know what port some protocol listens on. Oh, it’s a shame you never going to be able to learn that. You didn’t know that in the interview off the top of your head and there’s no possible way you could learn that. It’s an intrinsic piece of knowledge you’re born with.” No, it’s not.

Francessca: [smile]. Yeah, yeah, those are still things every now and then I have to go search for, or I’ve written myself some nice little Textract. Uh… [smile] [unintelligible 00:22:28] to go and search my handwritten notes for things. But yeah, so problem-solving, being able to effectively communicate. In our case, writing has been a muscle that I’ve really had to work at hard since joining here.

I haven’t done that in a while, so that is a skill that’s come back. And I think the one that I see around software development is, really, teams. It’s interesting because when you’re going through some of the curriculums, a lot of the projects that are assigned to you are individual, and what happens when you get into the workplaces, the projects become very team-oriented, and they’re more than one people. We’re all looking at how we publish code together to create a process, and I think that’s one of the biggest surprises making a transition [smile] into the workforce is, you will work in teams. [smile].

Corey: Oh, dear Lord. The group project; the things that they do in schools is one of those, great, there’s one person who’s going to be diligent—which was let’s be clear, never me—they’re going to do 90% of the work on it and everyone shares credit equally. The real world very rarely works that way with that sense of one person carries the team, at least ideally. But on the other side of it, too, you don’t wind up necessarily having to do these things alone, you don’t have to wind up with dealing with those weird personal dynamics in small teams, for the most part, and setting people up with the expectation, as students, that this is how the real world works is radically different. One of the things that always surprised me growing up was hearing teachers in middle school and occasionally beyond, say things like, “When you’re in the real world”—always ‘the real world’ as if education is somehow not the real world—that, “Oh, your boss is never going to be okay with this, or that, or the other thing.”

And in hindsight, looking back at that almost 30 years later, it’s, “Yeah, how would you know? You’ve been in academia your entire life.” I’m sorry, but the workplace environment of a public middle school and the workplace environment of a corporate entity are very culturally different. And I feel confident in saying that because my first Unix admin job was at a university. It is a different universe entirely.

Francessca: Yeah. It’s an area where you have to be able to balance the academia component with practitioner. And by the way, we talk about this in our solutions architecture and our customer solutions team—that’s a mouthful—in our organization, that how we like to differentiate our capabilities with customers is that we are users, we are practitioners of the services, we have gone out and obtained certifications. We don’t always just speak about it, we’d like to say that we’ve been in the empty chair with the customer, and we’ve also done. So yeah, I think it’s a huge balance, by the way, and I just hope that over the next several years, Corey, that again, we start really shifting the landscape by tapping into what I think is an incredible global workforce, and of users that we’ve just not inspired enough to go into these disciplines for STEM, so I hope we do more of that.

And I think our customers will benefit better from it because you’ll get more diversity in thought, you’ll get different types of innovation for your solution set, and you’ll maybe mirror the customer segments that you’re responsible for serving. So, I’m pretty bullish on this topic. [smile].

Corey: I think it’s hard not to be because, sure, things are a lot more complex now, technically. It’s a broader world, and what’s a tech company? Well, every company, unless they are asleep at the wheel, is a tech company. And that that can be awfully discouraging on some level, but the other side of it has really been, as I look at it, is the sheer, I guess, brilliance of the talent that’s coming up.

I’m not talking the legend of industry that’s been in the field for 30 years; I’m talking some of the folks I know who are barely out of high school. I’m talking very early career folks who just have such a drive, and such an appetite for being able to look at how these things can solve problems, the ability to start thinking in innovative ways that I’ve never considered when I was that age, I look at this. And I think that, yeah, we have massive challenges in front of us as people, as a society, et cetera, but the kids are all right, for lack of a better term.

Francessca: [smile].

Corey: And I want to be clear as well; when we talk about new to tech, I’m not just talking new grads; I’m talking about people who are career-changing, where they wound up working in healthcare or some other field for the first 10 years of their career—20 years—and they want to move into tech. Great. How do we throw those doors open, not say, “Well, have you considered going back and getting a degree, and then taking a very entry-level job?” No. A lateral move, find the niches between the skill you have and the skill you want to pick up and move into the field in half steps. It takes a little longer, sure, but it also means you’re not starting over from square one; you’re making a lateral transition which, because it’s tech, generally comes with a sizable pay bump, too.

Francessca: One of the biggest surprises that I’ve had since joining the organization, and—you know, we have a very diverse, large global field organization, and if you look at our architecture teams, our customer solution teams, even our product engineering teams, one of the things that might surprise many people is many of them have come from customers; they’ve not come from what I would consider a traditional, perhaps, sales and marketing background. And that’s by design. They give us different perspective, they help us ensure that, again, what we’re designing and building is applicable from an end-user perspective, or even an industry, to your point. We have lots of different services now, over a hundred and seventy-five plus. I mean, we’ve—close to two hundred, now.

And there are some customers who want the freedom to be able to build in the various domains, and then we have some customers who need more help and want us to put it together as solutions. And so having that diversity in some of the folks that we’ve been able to hire from a customer or developer standpoint—or quite frankly, co-founder standpoint—has really been amazing for us. So.

Corey: It’s always interesting whenever I get the opportunity to talk to folks who don’t look like me—and I mean that across every axis you can imagine: people who didn’t come up, first off, drowning in the privilege that I did; people who wound up coming at this from different industries; coming at this from different points of education; different career trajectories. And when people say, “Oh, yeah. Well, look at our team page. Everyone looks different from one another.” Great. That is not the entirety what diversity is.

Francessca: Right.

Corey: “Yeah, but you all went to Stanford together and so let’s be very realistic here.” This idea that excellence isn’t somehow situational, the story we see about, “Oh, I get this from recruiters constantly,” or people wanting to talk about their companies where, yes, ‘founded by Google graduates’ is one of my personal favorites. Google has 140,000 people and they founded a company that currently has five folks, so you’re telling me that the things that work at Google somehow magically work at that very small scale? I don’t buy that for a second because excellence is always situational. When you have tens of thousands of people building infrastructure for you to work on, back in the early days was always the story that, that empowered folks who worked at places like Google to do amazing things.

What AWS built, fundamentally, was the power to have that infrastructure at the click of a button where the only bound—let’s be realistic here—is your budget. Suddenly, that same global infrastructure and easy provisioning—‘easy,’ quote-unquote—becomes something everyone can appreciate and get access to. But in the early days, that wasn’t the thing at all. Watching our technology has evolved the state of the art and opened doors for folks to be just as awesome where they don’t need to be in a place like Google to access that, that’s the magic of cloud to me.

Francessca: Yeah. Well, I’m a huge, just, technology evangelist. I think I just was born with tech. I like breaking things and putting stuff together. I’ll tell you just maybe two other things because you talked about excellence and equity.

There’s two nonprofits that I participate in. One I got introduced through AWS, our current CEO, Andy Jassy, and our Head of Sales and Marketing, Matt Garman. But it’s called Rainier Scholars, and it’s a 12-year program. They offer a pathway to college graduation for low-income students of color. And really, ultimately, their mission is to answer the question of how do we build a much more equitable society?

And for this particular nonprofit, education is that gateway, and so spent some time volunteering there. But then to your point on the opportunity side, there’s another organization I just recently became a part of called Year Up. I don’t know if you’ve heard of them or worked with them before—

Corey: I was an instructor at Year Up, for their [unintelligible 00:31:19] course.

Francessca: Ahh. [smile].

Corey: Oh, big fan of those folks.

Francessca: So, I just got introduced, and I’m going to be hopefully joining part of their board soon to offer up, again, some guidance and even figuring out how we can help. But so you know, right? They’re then focused on serving a student population and decreasing, shrinking the opportunity divide. Again, focused on equitable access. And that is what tech should be about; democratizing technology such that everyone has access. And by the way, it doesn’t mean that I don’t have favorite services and things like that, but it does mean—[smile] providing [crosstalk 00:31:58]—

Corey: They’re like my children; I can’t stand any of them.

Francessca: [smile]. That’s right. I do have favorite services, by the way.

Corey: Oh, as do we all. It’s just rude to name them because everyone else feels left out.

Francessca: [smile] that’s right. I’ll tell you offline. Providing that equitable access, I just think is so key. And we’ll be able to tap in, again, to more of this talent. For many of these companies who are trying to transform their business model, and some—like last year, we saw companies just surviving, we saw some companies that were thriving, right, with what was going on.

So again, I think you can’t really talk about a comprehensive tech strategy that will empower your business strategy without thinking about your workforce plan in the process. I think it would be very naive for many companies to do that.

Corey: So, one question that I want to get to here has been that if I take a look at the AWS service landscape, it feels like Perl did back when that was the language that I basically knew the best, which is not saying much.

Francessca: You know you’re dating yourself now, Corey.

Corey: Oh, who else would date me these days?

Francessca: [smile].

Corey: My God. But, “There’s more than one way to do it,” was the language’s motto. And I look at AWS environments, and I had a throwaway quip a few weeks back from the time of this recording of, “There are 17 ways to deploy containers on AWS.” And apparently, it turned into an internal meme at AWS, which is just—I love the fact that I can influence company cultures without working there, but I’ll take what I can get. But it is a hard problem of, “Great, I want to wind up doing some of these things. What’s the right path?” And the answer is always, “It depends.” What are you folks doing to simplify the onboarding journey for customers because, frankly, it is overwhelming and confusing to me, so I can only imagine what someone who is new to the space feels. And from customers, that’s no small thing.

Francessca: I am so glad that you asked this question. And I think we hear this question from many of our customers. Again, I’ve mentioned earlier in the show that we have to meet customers where they are, and some customers will be at a stage where they need, maybe, less prescriptive guidance: they just want us to point them to the building blocks, and other customers who need more prescriptive guidance. We have actually taken a combination of our programs and what we call our solutions and we’ve wrapped that into much stronger prescriptive guidance under our migration and again, our modernization initiative; we have a program around this. What we try to help them do first is assess just where they are on the adoption phase.

That tends to drive then how we guide them. And that guidance sometimes could be as simple as a solution deployment where we just kind of give them the scripts, the APIs, a CloudFormation template, and off they go. Sometimes it comes in the form of people and advice, Corey. It really depends on what they want. But we’ve tried to wrap all of this under our migration acceleration program where we can help them do a fast, sort of, assessment on where they are inclusive of driving, you know, a quick business case; most companies aren’t doing anything
without that.

We then put together a fairly fast mobilization plan. So, how do they get started? Does it mean—can they launch a control foundation, control tower solutions to set up things like accounts, identity and access management, governance. Like, how do you get them doing? And then we have some prescriptive guidance in our program that allows them to look at, again, different solution sets to solve, whether that be data, security. [smile].

You mentioned containers. What’s the right path? Do I go containers? Do I go serverless? Depending on where they are. Do I go EKS, ECS Anywhere, or Fargate? Yeah. So, we try to provide them, again, with some prescriptive guidance, again, based on where they are. We do that through our migration acceleration initiative. To simplify. So.

Corey: Oh, yeah. Absolutely. And I give an awful lot of guidance in public about how X is terrible; B is the better path; never do C. And whenever I talk—for example, I’m famous for saying multi-cloud is the wrong direction. Don’t do it.

And then I talk to customers who are doing it and they expect me to harangue them, and my response is, “Yeah, you’re probably right.” And they’re taken aback by this. “Does this mean you’re saying things you don’t believe?” No, not at all. I’m speaking to the general case, where if, in the absence of external guidance, this is how I would approach things.

You are not the general case by definition of having a one-on-one conversation with me. You have almost certainly weighed the trade-offs, looked at the context behind what you’re doing and why, and have come to the right decision. I don’t pretend to know your business, or your constraints, or your capabilities, so me sitting here with no outside expertise, looking at what you’ve done, and saying, “Oh, that’s not the right way to do it,” is ignorant. Why would anyone do that? People are surprised by that because context matters an awful lot.

Francessca: Context does matter, and the reason why we try not to just be overly prescribed, again, is all customers are different. We try to group pattern; so we do see themes with patterns. And then the other thing that we try to do is much of our scale happens through our partner ecosystem, Corey, so we try to make sure that we provide the same frameworks and guidance to our partners with enough flexibility where our partners and their IP can also support that for our customers. We have a pretty robust partner ecosystem and about 150-plus partners that are actually with our migration, you know, modernization competency. So yeah, it’s ongoing, and we’re going to continue to iterate on it based on customer feedback. And also, again, our portfolio of where customers are: a startup is going to look very different than 100-year-old enterprise, or an independent software vendor, who’s moving to SaaS. [smile].

Corey: Exactly. And my ridiculous build-out for my newsletter pipeline system leverages something like a dozen different AWS services. Is this the way that I would recommend it for most folks? No, but for what I do, it works for me; it provides a great technology testbed. And I think that people lose sight pretty quickly of the fact that there is in fact, an awful lot of variance out there between use cases’ constraints. If I break my newsletter, I have to write it by hand one morning. Oh, heavens, not that. As opposed to, you know, if Capital One goes down and suddenly ATMs starts spitting out the wrong balance, well, there’s a slightly different failure domain there.

Francessca: [smile].

Corey: I’m not saying which is worse, mind you, particularly from my perspective, however, I’m just saying it’s different.

Francessca: I was going to tell you, your newsletter is important to us, so we want to make sure there’s reliability and resiliency baked into
that.

Corey: But there isn’t any because of my code. It’s terrible. This—if—like, forget a region outage. It’s far more likely I’m going to make a bad push or discover some weird edge case and have to spend an hour or two late at night fixing something, as might have happened the night before this recording. Ahem.

Francessca: [smile]. Well, by the way, I’m obligated, as your Chief Solution Architect, to have you look at some form of a prototype or proof of concept for Textract if you’re having to handwrite out all the newsletters. You let me know when you’d like me to come in and walk you through how we might be able to streamline that. [smile].

Corey: Oh, I want to talk about what I’ve done. I want to start a new sub-series on your site. You have the This is my Architecture. I want to have something, This is my Nonsense Architecture. In other words, one of these learning by counterexample stories.

Francessca: [smile]. Yeah, Matt Yanchyshyn will love that. [smile].

Corey: I’m sure he will. Francessca, thank you so much for taking the time to speak with me. If people want to learn more about who you are, what you believe, and what you’re up to, where can they find you?

Francessca: Well, they can certainly find me out on Twitter at @FrancesscaV. I’m also on LinkedIn. And I also want to thank you, Corey. It’s been great just spending this time with you. Keep up the snark, keep giving us feedback, and keep doing the great things you’re doing with customers, which is most important.

Corey: Excellent. I look forward to hearing more about what you folks have in store. And we’ll, of course, put links to that in the [show notes 00:40:01]. Thank you so much for taking the time to speak with me.

Francessca: Thank you. Have a good one.

Corey: Francessca Vasquez, VP of Technology at AWS. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with a comment telling me why there is in fact an AWS/400 mainframe; I just haven’t seen it yet.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Angie

Angie Jones is a Java Champion and Senior Director who specializes in test automation strategies and techniques. She shares her wealth of knowledge by speaking and teaching at software conferences all over the world, writing tutorials and technical articles on angiejones.tech, and leading the online learning platform, Test Automation University.

As a Master Inventor, Angie is known for her innovative and out-of-the-box thinking style which has resulted in more than 25 patented inventions in the US and China. In her spare time, Angie volunteers with Black Girls Code to teach coding workshops to young girls in an effort to attract more women and minorities to tech.

Links:

  • Applitools: https://applitools.com
  • Black Girls Code: https://www.blackgirlscode.com
  • Test Automation University: https://testautomationu.applitools.com
  • Personal website: https://angiejones.tech
  • Twitter: https://twitter.com/techgirl1908

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by CircleCI. CircleCI is the leading platform for software innovation at scale. With intelligent automation and delivery tools, more than 25,000 engineering organizations worldwide—including most of the ones that you’ve heard of—are using CircleCI to radically reduce the time from idea to execution to—if you were Google—deprecating the entire product. Check out CircleCI and stop trying to build these things yourself from scratch, when people are solving this problem better than you are internally. I promise. To learn more, visit circleci.com.

Corey: This episode is sponsored in part by Thinkst. This is going to take a minute to explain, so bear with me. I linked against an early version of their tool, canarytokens.org in the very early days of my newsletter, and what it does is relatively simple and straightforward. It winds up embedding credentials, files, that sort of thing in various parts of your environment, wherever you want to; it gives you fake AWS API credentials, for example. And the only thing that these things do is alert you whenever someone attempts to use those things. It’s an awesome approach. I’ve used something similar for years. Check them out. But wait, there’s more. They also have an enterprise option that you should be very much aware of canary.tools. You can take a look at this, but what it does is it provides an enterprise approach to drive these things throughout your entire environment. You can get a physical device that hangs out on your network and impersonates whatever you want to. When it gets Nmap scanned, or someone attempts to log into it, or access files on it, you get instant alerts. It’s awesome. If you don’t do something like this, you’re likely to find out that you’ve gotten breached, the hard way. Take a look at this. It’s one of those few things that I look at and say, “Wow, that is an amazing idea. I love it.” That’s canarytokens.org and canary.tools. The first one is free. The second one is enterprise-y. Take a look. I’m a big fan of this. More from them in the coming weeks.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. If there’s one thing that I have never gotten the hang of, its testing. Normally, I just whack the deploy button, throw it out into the general ecosystem, and my monitoring system is usually called ‘customers.’ And if I don’t want to hear from them, I just stopped answering calls from the support desk. Apparently, that is no longer state of the art because it’s been about 15 years. Here to talk about testing from a more responsible direction is Angie Jones, a senior director and developer at Applitools. Thanks for joining me.

Angie: Hey, Corey. [laugh]. I am cracking up at your confession there and I appreciate it because you’re not unique in that story. I find that a lot of engineers [laugh] follow that same trend.

Corey: There are things we talk about and there are the things that we really do instead. We see it all over the place. We talk about infrastructure as code, but everyone clicks around for a few things in the Cloud Console, for example. And so on, and so forth. We all know we should in theory be doing things, but expediency tends to win the day.

And for better or worse, talking about testing, in many cases, makes some of us feel better about not actually doing testing. And one of these days, it’s one of those, “I really should learn how TDD would work in an approach like this.” But my primary language has always been, well, always been a crappy version of whatever I’m using, but for the last few years, it’s been Python. There are whole testing frameworks around all of these things, but I feel like it requires me to actually have good programming practices to begin with which, let’s be very clear here, I most assuredly doubt.

Angie: [laugh]. That’s a fair assessment, but I would also argue, in cases like those, you need testing even more, right? You need something to cover your butt. So, what are you doing? You’re just, kind of, living on the edge here?

Corey: Sort of. In my case, it’s always been that I’ll bring in an actual developer who knows what they’re doing to—

Angie: Ah.

Corey: —turn some of my early scripts into actual tools. And the first question is, “Okay, can you explain what this is doing for me?” “Great. So, we’re going to throw it away and completely replace it with—so what are the inputs, what are the outputs, and do you want me to preserve the bugs or not?” At which point, it’s great.

It’s more or less like I’m inviting someone to come in and just savage my code, which is apparently also a best practice. But for better or worse, I’ve never really thought of myself as an engineer, so it’s one of those areas where it’s it doesn’t cut to the core of my identity in any particular way. I do know it would be nice that, oh yeah, when I wind up doing an iterative deployment of a Lambda function or something, if it takes five minutes to get updated, and then I forgot to put a comma in or something ridiculous like that. Yeah. Would have been nice to have something—you know, a pre-commit hook—that caught something like that.

Angie: Yeah, yeah. It’s interesting. You said, “Well, maybe one of these days, I’ll learn.” And that’s the issue I find. No matter what route you took to learn how to become—whatever you are, software engineer, whatever—testing likely wasn’t part of that curriculum.

So, we focus—when teaching—very heavily on teaching you how to code and how to build something, but very little, if any, on how to ensure you built the right thing and that it stands the test of time.

Corey: My approach has always been well, time to write some code, and it started off as just, as a grumpy systems administrator, it was always shell scripts, which, okay, great. Instead of doing this thing on 15 machines, run upon a for loop and just iterate through them. And in time, you start inheriting other people’s crappy tooling, and well, I could rewrite the entire thing and a week-and-a-half, or I could figure out just enough Perl to change that one line in there, and that’s how they get you. You sort of stumbled your way into it in that direction. Naive questions I always like to ask around testing that never really get answers for because I don’t think to ask these when other people are in the room and it’s not two o’clock in the morning and the power is gone out.

You have a basic linter test of, do you have basic syntax errors in the code? Will it run? Seems to be a sort of baseline, easy acceptance test. But then you get into higher-level testing of unit tests, integration tests, and a bunch of others I’m sure I’m glossing over because—to be direct—I tend to conflate all these in my head. What is the hierarchy of testing if there is such a thing?

Angie: Yeah, so Mike Cohn actually created a model that is very heavily used within the industry, and it’s called the ‘Test Automation Pyramid.’ And what this model suggests is that you have your unit tests; you have some kind of, like, integration-type tests in the middle, and then you have these end-to-end tests on top. So, think of a pyramid divided into three sections. But that’s not divided equally; the largest part of that pyramid, which is the base, is the unit test. So, this suggests that the bulk of your test suite should comprise of unit tests.

The idea here is that these are very small, they’re very targeted, meaning they’re easier to write, they take less time to run, and if you have an error, it kind of pinpoints exactly what’s wrong in the system. So, these are great. The next level would be your integration. So, now how do two units integrate together? So, you can test this layer multiple different ways: it might be with APIs, it might be the business logic itself, you know, calling into functions or something like that.

And this one is smaller than the unit test but not as large as the final part, which is the end-to-end test. And that one is your smallest piece, and it doesn’t even have to be end-to-end. It could be UI, actually. That’s how it’s labeled by Mike Cohn in his book: UI tests. So, the UI tests, these are going to be your most fragile tests, these are going to take the most time to write as well as the most time to execute.

If something goes wrong, you have to dig down to figure out what exactly broke to make this happen. So, this should be the smallest chunk of your overall testing strategy.

Corey: People far smarter than I have said that in many cases—along with access—testing, and monitoring—or observability, which is apparently a term for hipster monitoring—are lying on the same axis. Where in the olden days of systems administration, you can ping the machine and it responds just fine, but the only thing that’s left on that crashed machine is just enough of the network stack to return a ping, so everything except the thing that tells you it’s fine is in fact broken. So, as you wind up building more and more sophisticated applications, the idea being that the testing and the ‘is everything all right’ monitoring ping tends to, more or less, coalesce into the same thing. Is that accurate from your view of the world? Is that something that is an oversimplification of something much more nuanced? Or did I completely misunderstand what they were saying, which is perfectly possible?

Angie: You kind of lost me somewhere in the middle. So, I’m just going to nod and say yes. [laugh].

Corey: [laugh]. No, no, it—the hard part that I’ve always found is… I lie to myself, when I’m writing code: “Oh, I don’t need to write a unit test for this,” because I’d gotten it working, I tested it with something that I know is good, it returns what I expect; I tested with something bad and well, some undefined behavior happens—because that’s a normal thing to happen with code—and great, I don’t need to have a test for that because I’ve already got it working. Problem solved.

Angie: Right. Right.

Corey: It’s a great lie.

Angie: Yeah.

Corey: And then I make a change later on that, in fact, does break it. It’s the, “But I’m writing this code once and why would I ever go back to this code and write it again? It’s just a quick-and-dirty patch that only needs to exist for a couple of weeks.” Yeah, the todo: remove this later, and that code segment winds up being load-bearing decades into the future. I’m like, “Yeah, one of these days, someone’s going to go back and clean up all of my code for me.” Like, the code fairies are going to come in the middle of the night with the elves, and tidy everything up. I would love to hire those mythical creatures, but can’t find them.

Angie: This mythical sprint, where it’s, “Oh, let’s only clean up this entire sprint.” You know, everybody’s kind of holding out and waiting for that. But no, you hit the nail on the head with the reason why you need to automate your tests, essentially. So, I find a lot of newer folks to the space, they really don’t understand, why on earth would I spend time writing code to represent this test? Just like you said, “I implemented the feature. I tried it out, it worked.” [laugh]. “And hey, I even tried a non-happy path. And when it broke, I had a nice little error message to tell the user what to do.”

And they feel really good about that, so they can’t understand, “Why would I invest the time—which I don’t have—to write some tests?” The reason for that it’s just as you said: this is for regression. Unless that’s the end of this application and you’re not going to touch it ever again for any reason, then you need to write some tests [laugh] because you’re going to constantly change the application, whether that be refactoring, whether that be adding new features to it, it’s going to change in some way and you cannot be sure that the tests of yesterday still work today because whenever you make the change, you’re just going to poke around manually at that little area not realizing there could be some integration things that you totally screwed up here and you miss that until it goes out into prod.

Corey: The worst developer I’ve ever met—hands down—was me, six months before I’m looking at whatever it is that I’ve written. And given that I do a lot of my stuff in a vacuum and I’m the only person to ever touch these repositories, I could run Git blame, but I already know exactly what it’s going to tell me—

Angie: “It’s me.” [laugh].

Corey: —so we’re just going to skip that part. Like it’s a test. And, “Yeah, we’re just going to try and fix that and never speak about it again.” But I can’t count the number of times I have looked at code that I’ve written—and I do mean written; not blindly copy-and-pasted out of Stack Overflow, but actually wrote, and at the time, I understood exactly what it did—and then I look at it, and it is, “What on earth was I thinking? What—what—it technically doesn’t even return anything; it can’t be doing anything. I can just remove that piece entirely.” And the whole thing breaks.

I’ve out-clevered myself in many respects. And I love the idea, the vision, that testing would catch these things as I’m making those changes, but then I never do it. It’s getting started down that path and developing a more nuanced, and dare I say it, formal understanding of the art and science of software development. Always feels like the sort of thing I’ll get to one of these days, but never actually got around to. Nowadays, my testing strategy is to just actually deploy things into someone else’s account and hope for the best.

And, “Oh, good. Well, everyone has a test account; ideally, it’s not their own production account.” And then we start to expand on beyond that. You have come to this from a very different direction in a number of different ways. You are—among other things—a Java Champion, which makes it sound like you fought the final boss at the end of the developer internet. And they sound really hard. What is a Java Champion?

Angie: Yeah. So, a Java Champion is essentially an influencer in the Java ecosystem. You can’t just call yourself this; like you say, you got to fight the guy at the end, you know? But seriously, in order to become one, a current Java Champion has to nominate you, and all of the other Java Champions has to review your package, basically looking at your work. What have you contributed to the developer community, in terms of Java?

So, I’ve done a number of courses that I’ve taught; I’ve taught at the university level, as well; I am always talking about testing and using Java to show how to do that, as well as talks and all of this stuff. So apparently, I had enough [laugh] for folks to vote me in. So, it is an organization that’s kind of ordained by Oracle, the Gods of Java. So, it’s a great accomplishment for me. I’m extremely happy about it. And just so happens to be the first black woman to become a Java Champion. So, the news made a big deal about that. [laugh].

Corey: Congratulations. Anytime you wind up getting that level of recognition in any given ecosystem, it’s something to stop and take note of. But that’s compounded by just the sheer scale and scope of the Java community as a whole. Every big tech company I know has inordinate amounts of Java scattered throughout their infrastructure, a lot of their core services are written in Java, which makes me feel increasingly strange for not really knowing anything about it, other than that, it’s big and that there are—this entire ecosystem of IDs, and frameworks, and ways to approach these things that it feels like those of us playing around in crappy bash-scripting-land have the exact opposite experience of, “Oh, I’m just going to fire up an empty page and fill it with a bunch of weird commands and run it, and it fails, and run it again, and it fails. And it finally succeeds when I fixed all the syntax errors, and that’s great.” It feels like there is a much more structured approach to writing Java compared to other languages, be they scripts or full-on languages.

Angie: Yeah. That’s been a gift and a curse of the language. So, as newer frameworks have come out, or even as JavaScript has made its way to the front of the line, people start looking at Java, it’s kind of bloated, and all of these rules and structures were in place, but that feels like boilerplate stuff and cumbersome in today’s development space. So, fortunately, the powers that be have been doing a lot of changes in Java. We went for quite a while where releases were about, mmm, every three years or so.

And now they’ve committed to releases every six months. So, [laugh] most people are on Java 8 still, but we’re actually at, like, Java 16, now. So, now it’s kind of hard to keep up but that makes it fun as well. There’s all of these newer features and new capabilities, and now you can even do functional programming in Java, so it’s pretty nice.

Corey: Question I have is, does testing lend itself more easily to Java versus other language? And I promise I’m not trying to start a language war here. I just know that, “Well, how do I effectively test my Python code?” Leads to a whole bunch of? “Well, it depends.”

It’s like asking an attorney any question on the planet; same story. Like, “Well, it really depends on a whole bunch of things.” Is it a clearer, more structured path in Java, or is it still the same murky there are 15 different ways to do it and whichever one you pick, there’s a whole cacophony of folks telling you you’ve done it wrong?

Angie: Yeah, that’s a very interesting question. I haven’t dug into that deep, but Java is by far the most popular programming language for UI test automation. And I wonder why that is because you don’t use Java for building front end. You use Java scripts. I don’t know how this ca—I—well, I do know how it came to be.

Like, back in the day, when we first started doing test automation, JavaScript was a joke, right? People would laugh at you if you said that you were going to use JavaScript. It’s, you know, “I’m going to learn JavaScript and try to enter the workforce.” So, you know, that was a big no-no, and kind of a joke back then. So, Java was what a lot of your developers were using even if they were only using it for the backend, maybe.

You didn’t really have a [unintelligible 00:16:32] language on the client-side, back then. You had your PHP on the back end, you just did some HTML and some CSS on the front end. So, there wasn’t a whole lot of scripting going on back then. So, Java was the language that people chose to use. And so there’s a whole community out there for Java and testing.

Like, the libraries are very mature, there’s open-source products and things like this. So, this is by far the most popular language that people use, no matter what their application is built in.

This episode is sponsored by our friends at Oracle Cloud. Counting the pennies, but still dreaming of deploying apps instead of "Hello, World"
demos? Allow me to introduce you to Oracle's Always Free tier. It provides over 20 free services and infrastructure, networking databases, observability, management, and security.

And - let me be clear here - it's actually free. There's no surprise billing until you intentionally and proactively upgrade your account. This means you can provision a virtual machine instance or spin up an autonomous database that manages itself all while gaining the networking load, balancing and storage resources that somehow never quite make it into most free tiers needed to support the application that you want to build.

With Always Free you can do things like run small scale applications, or do proof of concept testing without spending a dime. You know that I always like to put asterisks next to the word free. This is actually free. No asterisk. Start now. Visit https://snark.cloud/oci-free that's https://snark.cloud/oci-free.

Corey: If I were looking to get a job in enterprise these days, it feels like Java is the direction to go in, with the counterpoint that, let’s say that I go the path that I went through: I don’t have a college degree; I don’t have a high school diploma. If I were to start out trying to be a software engineering today, or advising someone to do the same, it feels like the lingua franca of everything today seems to be JavaScript in many different respects. It does front end; it does back end; people love to complain about it, so you know it’s valid. To be clear, I find myself befuddled every time I pick it up. I’m not coming at this from a JavaScript fanboy perspective in any respect.

The asynchronous execution flow always messes with my head and leaves me with more questions than answers. Is that assessment though—of starting languages—accurate? Are there cases where Java is absolutely the right answer, as far as what to learn first?

Angie: Yeah. So, I first started with C++, and then I learned Java. Well, what I find is, Java because it’s so strict—it’s a statically typed language, and there’s lots of rules, and you really need to understand paradigms and stuff like that with this language—it’s harder to learn, but once you learn it, it’s much easier to pick up other languages, even if they’re dynamically typed, you know? So, that’s been my experience with this. As far as jobs, so the last time I looked at this, someone did some research and wrote it up—this was 2019—and they looked at the job openings available at the time, and they divided it by language. And Java was at, like, 65,000 jobs open, Python was a close second was 62,000, and JavaScript was third place with 39,000.

So, quite a big difference. But if you looked at tech Twitter, you’d think, like, JavaScript is all there is. Most of my followers and folks that I follow are JavaScript folks, front-end folks. So, it is a language I think you definitely need to learn; it’s becoming more and more prevalent. If you’re going to do any sort of web app, [laugh] you definitely want to know it.

So, I’m definitely not saying, “Oh, just learn Java and that’s it.” I think there’s definitely a need for adding JavaScript to your repertoire. But Java, there does seem to be more jobs, especially the big enterprise-type jobs, in Java.

Corey: The reason I ask so much about some of the early-stage stuff is that in your spare time—which it sounds like you have so much of these days—you volunteer with Black Girls Code to help teach coding workshops to young girls in an effort to attract more women and minorities to tech. Which is phenomenal. Few years ago, I was a volunteer instructor for Year Up before people really realized, “Oh, maybe having an instructor who teaches by counterexample isn’t necessarily the best approach of teaching folks who are new to the space.”

But the curriculum I was given for teaching people how Linux worked and how to build a web servers and the rest, started off with a three-day module on how to use VI, an arcane text editor that no one understands, and the only reason we use it is because we don’t know how to quit it.

Angie: [laugh].

Corey: And that’s great and all, but I’m looking at this and my immediate impression was, “We’re scrapping that, replacing it with nano,” which is basically what you see is what you get, and something that everyone can understand and appreciate without three days of training. And it felt an awful lot like we’re teaching people VI almost as a form of gatekeeping. I’m curious; when you presumably go down the path of teaching people who are brand-new to the space? How do you wind up presenting testing as something that they should start with? Because it feels like a thing you have to know first before you can start building anything at scale, but it resonates, on some level, with feeling like it’s, ah, you
must be able to learn this religion first; then you’ll be able to go and proceed further. How do you square that circle?

Angie: Yeah. So, I had the privilege of being an adjunct professor at a college, and I taught Java programming to freshmen. This was really interesting because there’s so much to teach, and this is true of all the courses. So, when I say that they don’t include it in the curriculum, that’s not really that much of a slight on them. Like, it’s just so much you have to cover.

So I, me, the testing guru, I still couldn’t find space to devote an entire sitting, a chapter, or whatever on testing. So, I kind of wove it into my teaching style. So, I would just teach the concept, let’s say I’m teaching loops today, and I’ll have a little exercise that you do in class. So, we do things together, and then I say, okay, now you try it by yourself. Here’s a problem; call me over when you’re done.

And as they would call me over when they’re done, I would break it; I would break their code, right? I’d do some input that they weren’t expecting and all of a sudden is broken. And they started expecting me to do this, you know? “She’s going to come and she’s going to break my stuff.” So, they start thinking themselves, “Let me test it before I give it to my user,” who is Professor Angie, or whatever.

So, that’s how I taught them that. Same with homework assignments. So, they would submit it, I would treat it like a code review, go through line by line, I didn’t have any automated systems to test their homework assignments. I did it like a code review, gave them feedback on how to improve their style, but also I would try to break it and give them, “Here’s all the areas that you didn’t think of.” So, that was my way of teaching them that quality matters in how to think about beyond the requirement.

The requirement is going to say, “Someone needs to be able to log in.” It’s not going to give you all of the things that should happen, you know if there’s a wrong password, so these are things, as an engineer, you need to think beyond that one line requirement that you’ve got and realize that this is part of it as well.

Corey: So, it’s almost a matter of giving people context beyond just the writing of the code, which frankly, seems to be something that’s been missing for many aspects of engineering culture for a while, the understanding the people involved, understanding that it is not just you, or your department, or even your company in some cases.

Angie: Exactly. And I tried to stress that very heavily in each lecture: who is your end-user? And your end-user cannot see your code, they cannot see your comments in the code that’s telling them, “Make sure you input it this way,” or whatever. None of that is seen so you have to be very explicit in your messages, and your intent, and behavior with the end-user.

Corey: One last area I wanted to cover with you, when I was doing some research on you before the show, is that you are an IBM Master Inventor, which I had no idea what that was. Is that a term of art? Let me Google it. And it turns out that you have, according to LinkedIn at least, 27 patents in your name. And it’s, “Oh.”

Yeah, it’s one of those areas where you look at something like, what gives someone the hubris to call themselves—or the grounds to call themselves that? And, “Oh, yeah. Oh, they’re super accomplished, and they have a demonstrated track record of inventing things that are substantial and meaningful. I guess that would do it.” I’d never heard the term until now. What is that? And how are you that prolific, for lack of a better term?

Angie: Yeah, so I used to work at IBM and they’re really big on innovation. And I haven’t kept track in a while, but for many, many years, they were the number one producer of patents [laugh] of this year or whatever. So, it was kind of in the culture to innovate. Now, I will say, like, a very small percentage of people—employees—there would take it as far as I did to actually go and patent something—[laugh]—

Corey: Oh, it’s the ‘don’t offer if you’re not serious,’ model.

Angie: Yeah. [laugh]. But I mean, it was there; it was a program there where, hey, you got an idea for a software patent? Write it up, we’ll have our lawyers, our IP lawyers review it, and then they’ll take your little one-page doc and turn it into a twenty-five-page legal document that we submit to the USPTO—United States Patent Trademark Office—who then reviews it and decides if this is novel enough and grants it, or dismisses it. And, “Hey, we’ll pay you for these patents. We’ll pay for the whole process.” And so I thought, “Heck, why not?”

And I kind of got hooked. [laugh]. So, it just so happens that I got a lot of good ideas. And I would collaborate with people from other areas of the business, and it was an excellent way for me to learn about new technologies. If something new was coming out, I would jump on that to explore, play with it, and think about, are there any problems that this technology is not aimed to solve, but if I tweak it in some way, or if I integrate it with some other concept or some other technology, do I get something unique and novel here?

And it got to the point where I just started walking through life and as I’m hit with problems—like, I’ll give you an example. I’m in the grocery store, right, and this inevitably happens to everyone, what, you choose the wrong line in the grocery store. “This one looks like it’s moving, I’m
going to go here.” And then the whole time, you’re looking to your right, and that line is moving. And you’re, like, stuck.

Corey: Every single time.

Angie: Every time. So, it got—[laugh]—

Corey: Toll booths are the same way.

Angie: —it got to the point where I started recognizing when I’m frustrated, and say, “This is a problem. How can I use tech to solve this?” And so I, in that problem, I came up with this solution of how I could be able to tell which one of these is the right line to get into. And that consisted of lots of things like scanning the things in everyone’s cart. On your cart, you have these smart carts that know what’s inside of them, polling the customers’ spending or their behavior; so are they going to come up here and send the clerk back to go get cigarettes, or alcohol, or are they going to pull out 50 coupons? Are they going to write a check, which takes longer?

So, kind of factoring in all of these habitual behaviors and what’s in your cart right now, and determining an overall processing time. And that way, if you display that over each queue, which one would be the fastest to get into. So, things like that is what I started doing and patenting.

Corey: Well, my favorite part of that story is that it is clearly a deeply technical insight into this, but you’ve told the story in a way that someone who is not themselves deeply technical can wrap their heads around. And I just—making sure you’re aware of exactly how rare and valuable that particular skill set is. So, often there are people who are so in love with a technology that they cannot explain to another living soul who is not equally in love with that technology. That alone is one of the biggest reasons I wanted to have you on this show was your repeated, demonstrated ability to explain complex things simply in a way that—I know this is anathema for the tech industry—that is not condescending. I come away feeling I understand what you were talking about, now.

Angie: Thank you so much. That is one of the skills I pride myself on. When I give talks, I want everyone in that room to understand it, even if they’re not technical. And lots of times I’ve had comments from anyone from, like, the janitor to the folks who are working A/V who, they don’t work with computers or anything at all and they’ve come to me after these talks like, “Okay, I heard a lot of talks in here. Everybody is over my head. I understood everything you said. Thank you.” And yet it’s still beneficial to those who are deeply technical as well. Thank you so much for that.

Corey: No, it’s a very valuable thing and it’s what I look for the most. In fact, my last question for you is tying around that exact thing. You have convinced me. I want to learn more about test automation, and learn how this works and with an eye toward possibly one day applying it to some of my crappy nonsense that I’m writing. Other than going on Google and typing in a variety of search terms that will lead me to, probably, a Stack Overflow thread that has been closed as off-topic, but still left up to pollute Google search results, where should I go?

Angie: Yeah. So, I’ve actually started an entire university devoted to testing, and it’s called Test Automation Universityand I got my employer, Applitools, to sponsor this, so all of the courses are free.

And they are taught by myself as well as other leading experts in the test automation space. So, you know that it’s trusted; I vet all of the instructors, I’m very [laugh] involved in going through their material and making sure that it’s correct and accurate so the courses are of top quality. We have about a little over 85,000 students at Test Automation University, so you definitely need to become one if you want to learn more about testing. And we cover all of the languages, so Java, JavaScript, Python, Ruby, we have all of the frameworks, we have things around mobile testing, UI testing, unit testing, API testing. So, whatever it is that you need, we got you covered.

Corey: You also go further than that; you don’t just break it down by language, you break it down by use case. If I—

Angie: Yeah.

Corey: —look at Python, for example, you’ve got a Web UI path, you’ve got an—

Angie: Exactly.

Corey: API path, you’ve got a mobile path. It aligns not just with the language but with the use case, in many respects.

Angie: Mm-hm.

Corey: I’m really glad I asked that question, and we will, of course, include a link to that in the [show notes 00:31:10]. Thank you so much for taking the time to speak with me. If people want to learn more, other than going to Test Automation University, where can they find you?

Angie: Mm-hm. So, my website is angiejones.tech—T-E-C-H—and I blog about test automation strategies and techniques there, so lots of good info there. I also keep my calendar of events there, so if you wanted to hear me speak or one of my talks, you can find that information there. And I live on Twitter, so definitely give me a follow. It’s @techgirl1908.

Corey: And we will, of course, include links to all of that. Thank you so much for being so generous with your time and insight. I really appreciate it.

Angie: Yeah, thank you so much for having me. This was fun.

Corey: Angie Jones, Java Champion and senior director at Applitools. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you hated this podcast, please leave a five-star review on your podcast platform of choice along with a long, ranting, incoherent comment that fails to save because someone on that platform failed to write a test.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Natalie

Natalie is the Director of Marketing at the Duckbill Group. Her background includes marketing roles in the localization and SaaS industries. In her free time, she teaches yoga, creates beadwork, and tries to keep up with her toddler. All of which impacts how she approaches growth and storytelling. Natalie resides in Missoula, Montana with her husband, daughter, and two wild corgis

Links:

  • Twitter: https://twitter.com/natveiswilliams

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at the Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part my Cribl Logstream. Cirbl Logstream is an observability pipeline that lets you collect, reduce, transform, and route machine data from anywhere, to anywhere. Simple right? As a nice bonus it not only helps you improve visibility into what the hell is going on, but also helps you save money almost by accident. Kind of like not putting a whole bunch of vowels and other letters that would be easier to spell in a company name. To learn more visit: cribl.io

Corey: And now for something completely different!

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Periodically, it seems that I’ve misunderstood the fundamental concept of marketing and interpret it through a lens of aggressively shitposting on Twitter and other places. I have since been informed that what I do is less about marketing and more about creative stunts in public, which are apparently close, but not exactly the same thing. This led to a natural evolution of the Duckbill Group’s understanding of what marketing is, and effectively culminated in our hiring, earlier this year, of Natalie Williams, who joins me today to tell me what a Director of Snarketing might actually do that differs from my ridiculous nonsense. Natalie, thanks for joining me.

Natalie: Thanks, Corey. It’s great to be here.

Corey: At a high level, something that is one of the most misunderstood concepts across the board is what marketing even is. So, before we proceed down the barrel of inevitable, ridiculous commentary I’m about to levy at you, what is marketing?

Natalie: Marketing is a way to attract the audience that you’re looking for and to frame your services, your content, whatever you’re selling, to your audience in a way that makes sense to them. And I think that another piece of it that’s important is, everybody has a pain point; you really have to be able to find the emotion behind what you’re selling in order to attract your audiences.

Corey: So, fundamentally an authenticity story more than anything else.

Natalie: Yeah, it’s finding that—yeah, the authenticity in what you’re selling, and meeting your audience with their pain point.

Corey: So, many times, it seems like the next follow-up question for most companies would be, “Great. So, what emotion is it exactly where people reflexively reach for their wallet and hand all the cash in it to our company?” It tries to be combined with aspects of sales; it tries to, on some level, wind up doing the entire job, in some cases of not even having a decent offering in the first place, but if the marketing story is strong enough, it’ll sell. Is that overly cynical?

Natalie: I don’t think so. I think that marketing, it’s challenging in this time, we have so many options, so many choices there, our attention spans are so short. And so I don’t think it’s cynical because there’s a lot of noise to cut through, and I think what’s important is to think of your audience as—or what you’re selling to your audience in a way that, how can I make their lives easier? I think that’s what’s going to, like you said, get people to pull the money out of their pocket and give it to you. What can you do to make their life easier and how can you show them that what you’re selling is going to make their life easier?

Corey: There’s a recurring trope that engineers, developers, whatever we’re calling ourselves this week—anything except DevOps is fine—that we don’t like marketing because no one likes being marketed to et cetera, et cetera. And I feel like that, in many cases, is because most marketing is terrible at its job. It comes across as smarmy: if I were to talk to people in person the way that marketing talks to people, I’d get punched in the mouth a lot. And I feel that is not an accurate view. It’s almost like looking at salespeople through the lens of the worst experience you’ve ever had buying a used car. And I feel like by judging an entire field by some of its worst examples, people are prone to prematurely dismiss a very challenging field.

Natalie: Absolutely, I think that it’s very easy to pick out that dishonesty, like you said, the slimy marketing sales tactic that just feels, it feels like it’s too much. I think people are very, very good at being able to see that. It’s like the spam that you get in your email or your LinkedIn direct messages. It’s just this constant barrage of tactics to try and tell you what you need without actually getting to know what you need, what you’re looking for, and how it can help you. So, I think there are a lot of those tactics that really give marketing and sales I think, too, a bad name.

An easy way to cut through that is just to be human about what you’re doing. Think of your audience as human beings and not just a target or a persona that isn’t a real person. And so I think, like you said earlier, the authenticity piece of it; it’s not difficult to… well, I don’t think it’s difficult to be human. I think it’s easier to take the emotion out of it and just try and put a sales pitch or a marketing pitch together that hits on all the points of the services you provide, and, you know, “I’m so great, and this is what we do,” without being able to talk to the people that you're marketing to as the humans that they are.

Corey: It feels almost like I became a marketing person of sorts without ever intending to. When I started this company as an independent consultant, I was finding my initial source of clients to my personal network, as most small independent consultants do. And one thing that almost came by accidentally was that newsletter, Last Week in AWS, that took on an additional series of—it gets a different level of meaning to the ecosystem, and it turned out, sort of, growing like a weed. At which point, great, now it turns out that is marketing, although I didn’t realize it at the time. And at least from my perspective, I didn’t consider it marketing, which meant that I built it the way that I would want to receive things because, as it turned out, I was a relatively close match to the people that I was going to be emailing every week.

And as a result, I didn’t do a bunch of scammy nonsense that would alienate people; there’s never any tracking put into it in the form of, “Oh, if you click on a link, that implies that you’re going to be interested in what I’m selling. We’re now going to fire off drip campaign number 17.” We stepped away from a lot of that because I’ve been on the receiving end of it, and it always rings hollow and strange. When we were interviewing to fill your role, we made it a point very early in the interview process to make clear that, yeah, this is not going to be one of those scenarios where you walk in and effectively get to instrument everything with all of the latest marketing technology and the rest. Okay, yeah, we have a list now of almost 27,000 subscribers.

We don’t mine that for leads. Now, should we? Maybe in an idealized sense, but it turns out that the reputational management of people trusting us not to do skeezy things far outweighs any temporary benefit we get from doing things like that. Now, that wound up driving an awful lot of people away, which is why we mentioned this in the early stages; it didn’t drive you away. You decided to say yes; you actually showed up on Monday for your first day, and then continued to show up every day since. What was it about, I guess, that statement that would drive some folks that we spoke to away, but it wasn’t a deal-breaker for you?

Natalie: I remember clearly that part of the interview, and I remember it feeling like a bit of a breath of fresh air. I think that when you get into all the tech, all the marketing tools, there is that trade-off of losing a little bit of that human touch with your audience where all of a sudden they’re not humans, they turn into data points, they turn into the demographics about themselves. And all of that stuff is important and has a role, but when you lose that personal connection of the storytelling piece of it and you’re just driven on a lot of these metrics, which I think it’s very easy to almost paralyze yourself by looking at too many metrics and getting all this set of data versus just what are the most important pieces? What does our audience need? And you know that by knowing your audience, and so I think that’s why it was such a good point when you talked about, you know, you didn’t really know you were doing marketing, but you were having success with it because you were doing it in a way that was genuine to your mission and to your audience, and you weren’t doing it with the intent of how do I get X number of subscribers?

You were doing it in a way of how do I provide something that I think is beneficial to people like me. You had that pain point and you filled that pain point, and it resonated with your audience. And so, yeah, in the interview, when we were talking about those things, I just—it excited me to be able to have that connection with my audience and not be so focused on the data and the intel. Which, I’m not the kind of person that wants everybody to know every single thing that I do online anyway. I think that some of those tactics are, quite frankly, a little bit terrifying. So, knowing that wasn’t a push of knowing every single thing that you could possibly know about all of your audience was something that was very refreshing to me.

Corey: One of the things I built fairly early on was my own custom click-tracker because sponsors demand this, on some level, but I also found it useful for my own purposes in the aggregate. Specifically what I built is something that shows me the number of people that click on a given link in a particular issue, but it doesn’t tell me who any of those people were. It dedupes it, so it isn’t one person clicking a link 500 times winds up inflating the count, but that was as far as I went. And I use that data to periodically put together a best-of issue when there’s a slow week, when Amazon, I don’t know, winds up in a super-contentious company-wide argument about what is the absolute worst name to give something, and then there’s not a whole lot of releases that week.

Great. I can go back and highlight things that people found interesting and useful in the past because no one reads everything, and even if I’ve talked about it before, it’s new to someone. And that took a surprising amount of work because every time I would talk about what I was building with folks in the space it was, “But yeah, then don’t you want to track the people who click it as they move through your website, and as they wind up moving across the internet so you can start putting them into cohorts and the rest?” And no. First, I don’t have the patience to wind up doing that sort of tracking; I don’t speak Excel that fluently.

And also, it’s somewhat marginal as far as the ability to improve any meaningful business outcome. We don’t tend to manage by a whole bunch of iron metrics here around things like that. When I wind up doing ridiculous stunts, like a music video making fun of some Amazon exec or whatnot, I do it because I think it’d be funny and it will probably resonate, but there’s no business case behind it. It’s, “Ehh, why not? Worst case, we learn something.” And I get the sense, talking to folks, that is increasingly rare.

Natalie: I think it is, too, but I like the approach. I think the click-through rate, what’s important there is the success of your content. If people are clicking through it, your content is resonating. And I think that people really underestimate people’s ability to decide if they want to move forward and reach out to become a customer or go further down in the funnel. I think people are pretty smart.

Just because they clicked on a link, you don’t need to berate them with 15 more touchpoints to say, “Hey, are you ready to make this purchase? Are you ready for this?” I mean, all of these things, people can get there, and if you frame your funnel in a way that makes it easy for them to get to point A and point B, really shouldn’t have to have that much involvement in their journey. And some marketing people might not agree with me on that, but I think as long as you make it easy for them to get what they need, to find what they need, to learn as much as they can about your service before they purchase, you don’t need to have all of those touchpoints, you don’t need to know that they clicked on all of these things. And I think, too, as a user, if I know that if I click on a link, I’m going to get a barrage of
emails selling a service if I’m not ready for it, it’s going to make me hesitate to click a link.

And so I think that that approach is just a more user-friendly approach, and it really just makes the experience of the user that much more smooth knowing that there’s kind of a mutual trust there. I’ll let you know when I’m ready and you can help me get there, but don’t try and do it for me.

Corey: The piece that I’ve always found surprising was that everyone talks in the ad tech space as if without this data, our businesses will wither and die on the vine. But the ability to track people as they move across the internet and go through their lives is a relatively recent horrible invention. Companies existed and sold goods and services phenomenally well for millennia before the advent of any of these things. And does it improve around margins? Of course, it does; I’m not saying the industry is built on fraud. I’m just saying that it winds up making a trade-off that I’m not comfortable making myself, and therefore I’m unwilling to ask my audience members to make similar trade-offs themselves.

Natalie: Yeah. And it’s like, at what point do you get to be too much, where you know too much about your audience? And like I said before, people are smart; we don’t need constant advertisements across all of these platforms to remind us of something we thought about or we looked about. I mean, it’s good; there’s certainly a time and a place to have a reminder come through or be retargeted in a way that isn’t creepy, but I mean, it’s those things where I think about something and all of a sudden I see an ad, those still freak me out, so I think that people don’t need this over-advertised approach to make a purchase. Like you said, people have been purchasing things for a very long time, and I think that people are capable, people are smart, people know where they’re at in the journey better than, really, anybody else.

And so there is a time and a place for it, but there is also—it seems to be moving in too much information, too much interruption in people’s lives and it is certainly—yeah, if I don’t want it, I’m certainly not going to—or if I don’t want to be targeted that way, I’m not going to do it to anybody else either.

Corey: Turns out there’s an entire seedy underbelly to the web, and as soon as you have a website that has a little bit of traction, you wind up getting exposed to countless piles of spam about it. There’s the standard SEO stuff of how to optimize your results in search engines. I have this ancient approach that seems to serve me pretty well, which is, I write fun, engaging original content and then put it on the website, and then it sort of takes over from there. I don’t write with an eye towards, “Well, use the following phrase more than 20 times but less than 50 in the course of the next month.” It’s, “no.”

But then you wind up with folks guaranteeing, “Oh, use us for search engine optimization,” which I always think is the weirdest pitch because let’s face it, if I want someone to do my search engine optimization work, wouldn’t you think I would just type the word search engine optimization into Google and then click the number one result? Seems like they have the proof and no one else would. But there’s also this sea of folks asking if they can pay me to put a guest blog post on the website. And the answer is no. When we have guest blog posts—usually on Fridays—that comes from people that we pay to write them.

It’s not some random thing that’s only tangentially tied to what we do but includes a suspicious number of links to some other third-party website. It doesn’t make sense for me to just start devaluing the experience of the reader like that. I try to be as respectful of people’s time as I can be when it comes to content creation.

Natalie: Right. And I think I’m becoming a broken record on this point, but I will say it again: people are smart; they can see through those tactics. And yes, there are certain things that will help your content rank better: where you put your keywords, having a title and a meta description, and all those things are important, but at the end of the day you need to write for your audience and not for the search engines. They’re getting smarter, right, the search engines are getting smarter, but they are still not the humans and so if you overstuff your keywords, if you put in a bunch of links, it’s going to look like you’re doing it solely for the search engine and people will see that, they’ll recognize that, and they won’t value your content because they will feel like they’re being sold to. And they won’t get what they need out of it; they won’t see it as a thought leader that’s genuinely producing content in order to help inform them and it’ll be devalued.

Corey: This episode is sponsored by our friends at Oracle HeatWave is a new high-performance accelerator for the Oracle MySQL Database Service. Although I insist on calling it “my squirrel.” While MySQL has long been the worlds most popular open source database, shifting from transacting to analytics required way too much overhead and, ya know, work. With HeatWave you can run your OLTP and OLAP, don’t ask me to ever say those acronyms again, workloads directly from your MySQL database and eliminate the time consuming data movement and integration work, while also performing 1100X faster than Amazon Aurora, and 2.5X faster than Amazon Redshift, at a third of the cost. My thanks again to Oracle Cloud for sponsoring this ridiculous nonsense.

Corey: It’s disturbing in some ways to see the number of companies losing their collective minds [unintelligible 00:16:29] there’s a Google algorithm change, and suddenly, as a web browser myself—as in the person, not the software—when suddenly I’m no longer seeing a bunch of eHow links or whatnot dominating search results, or the godawful Experts Exchange site for a while where they would pretend to have a paywall, but scroll down far enough and they have the full text of the thing to keep within Google’s rules. And Google got better at these things with time. If a single change to a search engine algorithm can destroy your business, I wonder how sustainable business is in the long term.

Natalie: Right. And I think that, like you said, it’s—I mean, it’s a game, right, and the game constantly changes, it constantly involves. The search engines are getting smarter, but if you just try and constantly keep up with the rules of the game in that way, you’re going to lose because it’s just going to be a constant change. I think the most important thing is to just show up consistently and write good content. I mean, when SEO became a thing, if you put in five million keywords, you’re going to rank well, but it’s not any value.

And so if your intent is just to always produce good content, be aware of the rules of the game and apply them when it makes sense to, and figure out what the ones that are most important to you are, you will do well, but it’s just that consistently showing up, producing good content, and having your readers in mind is what’s going to help you go forward. If you’re just focused on the algorithms and making it fit in that, it’s going to look that way and you’re never going to be able to catch up.

Corey: For me, the big indicator that I’m doing something right is periodically, because it is a collection of links, I’ll link to some site that has some sort of attack malware on it, usually from an ad network that was compromised somewhere—surprise; it’s just a thing that happens—and as a result, some companies will start blocking it or will start putting it in the spam folder. And when you wind up with larger outside spam folks reaching out to you because of the number of complaints they’re getting from their customers about it being blocked, that’s my indication that we’re probably doing something kind of right, as opposed to it being incumbent on us to track that down from our side.

Natalie: Yeah. And like I said before, I mean, the users know what—they also know the rules of the game, too, right? Or they’re at least aware of it, and so they can recognize that. But like I said, it’s just having the goal of informing your audience and producing good content that you know is going to resonate with your audience. It’s just a smart way to go. I mean, having the user in mind always, and not it as a way to see them as just a dollar sign or a number, but knowing how can I really help them learn about this topic or share what I know, is going to always benefit.

Corey: The piece that I think resonates in some respects as well is what I’ll periodically do even on this podcast, where it’s, I will ask people about what they’re good at and listen to the answer, and then ask follow-up questions. It’s not exactly a pioneering technique; it’s just storytelling. And for whatever reason, it seems that marketing—the way I view it—is always tied back to storytelling. It’s about understanding who your audience is for anything that you’re working on, understand what action you want them to take. Maybe it is to fix their S3 bucket permissions.

Maybe it’s to reach out and ask me about AWS bill consulting. Whatever it is, understand who the audience is and what the expected outcome is, and then just tell a story around that. It’s not that compellingly difficult, but for some reason, it seems like it’s the most novel thing in the world to some companies. I always equated marketing with storytelling. Is that rare? Am I wrong on something?

Natalie: Not in my book. I mean, that’s how I think about it as well, too. I think that some companies get so focused on talking about themselves, and this is what we do and this is why we’re the best. But they miss that storytelling piece of it of why it matters to the audience. You can talk about your company until you’re blue in the face, but if you don’t connect it to your audience, they aren’t going to know how it helps them.

And maybe you have a service that they don’t know that they need yet, but by producing content and telling them a story of how what you’re doing is going to make their lives easier, you’re going to resonate with them. I mean, that’s why people look for services: because they have a problem. And maybe it’s not identified, or they don’t know how to fix it, but if you frame your content in a way that takes them through a journey of, this is where you’re at, this is how we can help you and then—or this can ask, how can you make your life better, and this is where you’ll be at the end of it, it’s going to resonate with your audience. And so yeah, I think that if you’re focused on shouting the benefits without connecting it to the emotion of why it matters, it’s going to come across that way. And so I think, yeah, storytelling, having the audience, their perspective, who they are, what they need, and telling the story in a way that is a good experience for them, is always the way to go.

Corey: For whatever reason, it seems like the folks that are the worst offenders of that tend to be the big cloud companies themselves, where, “Okay, here’s a service we built.” “Great, how am I going to use that to solve an actual business problem that I have?” “Next up, we have a guest speaker from Netflix to tell us what they did with it.” “Cool. Maybe I have problems that don’t look like Netflix-scale problems, maybe I’m just trying to contextualize this thing that you’ve released with the 200 other things that you have, and figure out what component that will replace, or extend, or what feature it will grant inside of my existing environment. Or even if it’s something new, it’ll give me ideas for different directions to move things in.”

And they never do that. They talk about features, they talk about capabilities, they talk about how innovative they are, and they just completely abdicate the entire role of telling a compelling story around this. I mean, Jeff Barr’s blog posts are a terrific counterexample to this. I think it’s the only time that AWS actually tells a story of, “I’m going to set out to build this thing to do this. Here’s how I do it, and here’s how it works.”

And it’s incredibly engaging, it has a distinct voice—because Jeff has a personality and can express that personality—and I get the sense, at times, he’s the only person at Amazon allowed to do that. Everyone else just winds up falling into these same somewhat tired tropes of, “Here’s a bunch of feature announcements.” The release feed doesn’t even let you include images, so wind up—people trying to describe the layout of a new console page to you. And it just doesn’t resonate or make sense. I mean, half the value of building the audience that I have is, whenever something comes out, people’s immediate question for me is—they’ll tell me that it’s out by, “Can you explain this to me because Amazon, once again, has failed to do so.” I didn’t show up here, due to the express purpose of being the de facto head of AWS marketing; it just kind of happened because functionally, it feels like it’s a pretty empty seat.

Natalie: Yeah, and I think that those are all really important points because they put out all this information about their services and their features, but they don’t complete the cycle to say how it affects the user, how they would use it, and so it’s like you’re not finishing the race there. And so you also have to acknowledge that you might target a similar set of companies, so maybe you’re targeting a bunch of companies in one space, but everybody’s pain points are going to be a little bit different, so if your audience can’t relate to how you’re implement—or how you’re telling them to implement a service or a feature, they’re not going to feel that. And so you have to tell your audience how what you’re doing is going to improve their life, but also doing it in a way that is going to apply to them. And so that’s part of understanding your audience and understanding the different pain points associated with each of your target audiences. And so that’s why it’s important to have some of those intel on your audience, but finding out a way to resonate it across all of the companies that you target and make it relatable is what’s going to be success.

Corey: It’s always strange to me that when we talk to sponsors who are—they want to wind up telling people about, great, I try to have conversations with them; imagine that? And understand, okay, how does this differentiate from, for example, an AWS service that does something very similar? And the answer is often, “Wait, that’s a real thing that exists?” If people in that market aren’t aware that there’s an AWS release that does the thing that they do, what hope do customers have? None.

Natalie: Right.

Corey: It feels like it’s not just a matter of will AWS launch a service that competes with you it’s, will they do an even passable job of telling a story around it so people know it exists? They tend to do everything very frugally, which means that they can release things that are competing with, in some cases, many billion-dollar companies out there. And great, okay, so you’ve got a 60-person team or whatnot, trying to compete with a company with a $30 billion valuation. Yeah, my money is on not Amazon for stuff like this, until they start trying to do things that reek of anti-competitive acts. I’m not super worried about Amazon as a competitor for a lot of these upstart services.

Natalie: Yeah. And I think another thing, too, that you touched on was having conversations with your audience and really understanding how they are perceiving what you’re selling. I think if you get caught up in, “This is what we’re going to push; this is how we’re going to say it; this is what we’re going to do,” but you don’t have that touchpoint of how are they perceiving it, are they aware of this? There’s a big missed opportunity there. Because if something’s not working, there’s a disconnect; there’s a miscommunication. So, being able to have some kind of engagement with your audience, where you’re surveying them, or whatever, to see how their perception of what you’re selling is vital in shaping your marketing story.

Corey: You really would think that this wouldn’t be anywhere near as challenging as it has clearly become.

Natalie: Mm-hm.

Corey: But here we are.

Natalie: And I think people, they try and make things more complicated, I think, from time to time. I think that just going back to the reason why you’re selling things: the reason why you’re providing a service is to help your audience fill a need, hit a pain point. And so I think that it can certainly get over complicated quickly and you can forget, really, the mission behind a lot of what you’re doing.

Corey: I wish more people shared your viewpoint. I guess, so far, all we can do is set a good example and hope the rest of the industry follows us wherever we’re going.

Natalie: Right. And I mean, you know, like I said, there’s always kind of moving in that direction of a lot of tech, a lot of these big things that are supposed to make your life easier in marketing, but I think it’ll always come back to the audience, it’ll always come back to the storytelling, it’ll always come back to the user experience. And so, it’s just riding the waves, and seeing what’s going to happen, and how things evolve. Because marketing and everything, it’s just a constant keeping track of the trends and the evolution of it.

Corey: And it never seems to hold still. If people want to hear more about what you’re up to and how you think about these things, where can they find you?

Natalie: You can find me on Twitter at @natveiswilliams.

Corey: I will, of course, put a link to that in the [show notes 00:27:31]. Thank you so much for taking the time to speak with me.

Natalie: Yep, Thanks, Corey. This was wonderful. I really enjoyed it.

Corey: Natalie Williams, Director of Snarketing here at the Duckbill Group. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an insulting comment telling me exactly which features on the checklist this episode should have had instead, phrased in the worst way possible.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need the Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Forrest

Forrest is a cloud educator, cartoonist, author, and Pwnie Award-winning songwriter. He currently leads the content marketing team at Google Cloud. You can buy his book, The Read Aloud Cloud, from Wiley Publishing or attend his talks at public and private events around the world.

Links:

  • The Cloud Bard Speaks: https://www.lastweekinaws.com/podcast/screaming-in-the-cloud/the-cloud-bard-speaks-with-forrest-brazeal/
  • The Read Aloud Cloud: https://www.amazon.com/Read-Aloud-Cloud-Innocents-Inside/dp/1119677629
  • The Cloud Resume Challenge Book: https://forrestbrazeal.gumroad.com/l/cloud-resume-challenge-book/launch-deal
  • The Cloud Resume Challenge: https://cloudresumechallenge.dev
  • Twitter: https://twitter.com/forrestbrazeal

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part my Cribl Logstream. Cirbl Logstream is an observability pipeline that lets you collect, reduce, transform, and route machine data from anywhere, to anywhere. Simple right? As a nice bonus it not only helps you improve visibility into what the hell is going on, but also helps you save money almost by accident. Kind of like not putting a whole bunch of vowels and other letters that would be easier to spell in a company name. To learn more visit: cribl.io

Corey: This episode is sponsored in part by Thinkst. This is going to take a minute to explain, so bear with me. I linked against an early version of their tool, canarytokens.org in the very early days of my newsletter, and what it does is relatively simple and straightforward. It winds up embedding credentials, files, that sort of thing in various parts of your environment, wherever you want to; it gives you fake AWS API credentials, for example. And the only thing that these things do is alert you whenever someone attempts to use those things. It’s an awesome approach. I’ve used something similar for years. Check them out. But wait, there’s more. They also have an enterprise option that you should be very much aware of canary.tools. You can take a look at this, but what it does is it provides an enterprise approach to drive these things throughout your entire environment. You can get a physical device that hangs out on your network and impersonates whatever you want to. When it gets Nmap scanned, or someone attempts to log into it, or access files on it, you get instant alerts. It’s awesome. If you don’t do something like this, you’re likely to find out that you’ve gotten breached, the hard way. Take a look at this. It’s one of those few things that I look at and say, “Wow, that is an amazing idea. I love it.” That’s canarytokens.org and canary.tools. The first one is free. The second one is enterprise-y. Take a look. I’m a big fan of this. More from them in the coming weeks.

Corey: Welcome to Screaming in the Cloud. I am Cloud Economist Corey Quinn, and as an industry, we stand on the precipice of change. There’s an awful lot of movement lately. It feels like the real triggering event for this was when Andy Jassy ascended from being the CEO of AWS—the cloud computing division of Amazon—to being the CEO of all of Amazon, including things like not just AWS, but also the underpants store. Suddenly, we have people migrating between different cloud providers constantly.

Today’s guest is a change I would not have expected and didn’t see coming. So, last year, on episode 127, called The Cloud Bard Speaks I had Forrest Brazeal from A Cloud Guru joining me. Forrest, welcome back.

Forrest: Hey, thanks, Corey. Big fan of the show; always great to be here.

Corey: At the time that we’re recording this, you are unemployed, which is great because it’s Screaming in the Cloud. Screaming at people on your day off is always fun. But by the time it airs, you’ll have started your new job as the Head of Content for Google Cloud.

Forrest: Yes. And of course, that’s definitely a career change for me coming directly from A Cloud Guru, which was a wonderful place to be and it was exciting to be with them right up through their acquisition earlier this summer, but when it came time to make the next move, I ended up going to Google Cloud. I’ll be starting there on Monday after this recording has been completed, and just really looking forward to helping tell the story of the cloud at a much bigger scale, something that I’ve been doing throughout my career with increasing levels of scale. It’s exciting to do it at the level of an entire cloud provider.

Corey: We’ll get to the future in a minute, but I want to start by looking at the past. From my perspective, you were a consultant for a while at Trek10; we’ve talked about that before. You have an engineering background of building things with computers, at least presumably computers—you’ve been a big serverless advocate and I’m told that runs on computers somewhere, but I don’t want to get into that particular debate—to the point where you were—I assume were, not are anymore—an AWS Serverless Hero?

Forrest: Yes, that’s right, and even going back prior to Trek10, my background is in enterprise software. I helped to migrate some of the world’s largest enterprise applications from data centers to cloud when I was at Infor and continued to work on that kind of thing as a consultant later on. And in that time, I was working a lot with AWS, which was the only game in town for a lot of those years, right? You go back to 2014, 2015, I’m putting an enterprise app in the cloud, what am I going to put it on? Probably AWS if I’m serious about what I’m doing.

But it’s been amazing to see how the industry has grown and changed and the other options that have come along. And one of the cool things about my work in A Cloud Guru is that I really got a chance to branch out and expand, not just to AWS, but also to get a much better feel for the other cloud providers, for Azure and GCP, and even beyond to Oracle and some of the other vendors that are out there. And just to get a better understanding of how these different cloud providers thrive in different niches. So yes, it is absolutely a change for me; I obviously won’t be an AWS Hero anymore, I’m having to close that chapter, sadly; I love those people and that program, but it is going to be a new and interesting change. I’m going to have to be back in learning mode, back in catch-up mode as I get busy on GCP.

Corey: So, one thing that I think gets occluded with you because it definitely does with me is that you and I are both distinguishable personalities in the cloud community—historically AWS, let’s be clear here—and you do your own custom songs; you write a newsletter that instead of snarky is insightful—of which I’m jealous—but it still has a personality that shines through; you wrote a children’s book, The Read Aloud Cloud; you wound up having a new book that just came out last week for folks listening to this the day of release, called The Cloud Resume Challenge Book, if I’m getting the terms all in the right order?

Forrest: Yeah, exactly.

Corey: It’s like naming cloud services only naming books instead? It’s still challenging to keep all the words in the right order?

Forrest: You know, I think it actually transcends industries; naming things is hard whether you’re in computer science or not.

Corey: Whereas making fun of things’ names is a lot easier. It’s something you did not do—to my understanding—as an employee of A Cloud Guru, The Cloud Resume Challenge, but it’s something you did as a side project because it interested you. It’s effectively, you want to get into tech, into cloud.

Great. Here’s a list of things I want you to do. And it ranges the gamut. And we talked about it before, but to my understanding it’s, build a statically hosted website that winds up building your resume, and a blog post, and how to do all these things, CI/CD, frontend, backend, the works. It’s a lot of work, but by the time you’re done, you know a heck of a lot more about the cloud provider you’re working with than you did when you started.

Forrest: Yeah, not only do you know more than you did when you started, but quite frankly, you’re going to know more than a lot of people who’ve even been doing this kind of thing for a couple of years. That’s why we have people that take The Cloud Resume Challenge, who are not only aspiring cloud engineers but who have been doing this for a while, maybe even are hiring people, and they see this project and say, “Wow. That would look good on my resume. I’ve never actually sat down and plugged a frontend and a backend together on AWS,” and, “Maybe I’ve never had to actually sit down and think carefully about how I would build a CI/CD pipeline,” or, “I really want to get my hands dirty with Terraform,” or something like that. So, we see a whole range of people.

I did a survey on this actually, and I found that about 40% of all the people who take The Cloud Resume Challenge have three years or more of professional IT experience. So, that should tell you how impressive it is, if you can figure this out as a brand new person to cloud. That’s why we’ve seen so many of these folks change careers and go from things like plumbing, and working in a bank, and working in HR, and whatever else to starting roles, now, as cloud engineers and DevOps engineers. It’s not entirely due to the challenge; not even mostly due to the challenge. These are folks who are self-motivated, quick learners, and are going to succeed no matter what, but The Cloud Resume Challenge was the thing that came on at the right time for them to build those skills and show what they had.

Corey: And the fact that you put this together is incredibly uplifting for folks new to the field. And that’s amazing, and it’s great, and it’s more content, the kind that I think that we need in this industry. You also launched a newsletter last week: the cloud jobs newsletter, which is fantastic. It’s a pay-to-subscribe newsletter—which I’ve always debated experimenting with but never did—and lists curated jobs in the industry, sorted by level of experience required and things that you find personally interesting. You might have sponsored job listings in the future that you’ve already said would be clearly delineated from the others, which is the ethically right thing to do. You are seemingly everywhere in the cloud space.

Forrest: Well, I mean look, I’m trying to give back. I’ve benefited from folks like yourself and others who have made time to help lift my career over the years, and I really want to be here to help others as well. The newsletter that you mentioned the Best Jobs in Cloud, it does have a small fee associated with it, but that’s really just to help gate my [laugh] referrals so that they don’t end up getting overwhelmed. You actually can get free access to the newsletter with the purchase of The Cloud Resume Challenge Book we talked about before. It’s really intended to be a package deal where you prepare your resume by doing these projects, and there’s a lot of other advice in that book about how to get yourself positioned for a great career in the cloud.

And then you have this newsletter coming into your inbox every couple of weeks that lays out a list of jobs and they’re broken down by, you know, these are jobs that are best for juniors, these are jobs where you’re going to need some senior-level experience. Because what I found—and honestly, I’ve been kind of acting as a talent agent for a lot of engineers over the past several years as my network has grown, and I’ve tried to give back to others and help to connect folks who are eagerly trying to find great engineers for cool projects that are working on with folks who are eagerly looking for those opportunities. And what I’ve realized is whether you’re a junior or whether you’ve been doing this for a long time, let’s face it, most of us are not spending all of our time being those distinguishable personalities that you mentioned a minute ago. I like how you said distinguishable and not distinguished by the way; those are two very different words. But most of us are not spending our time doing that.

You know, we’re working engineers; we’re working, right? We’re not blogging and tweeting all the time and building these gigantic personal networks. So, it helps if you can have a trusted friend standing alongside you so that when you are thinking about maybe making a switch, or maybe you’re not thinking about making a switch but you should be because of where the market is, that friend is coming alongside you and saying, “Hey, this is an awesome opportunity that I think you should consider checking out; why not just do the interview. Even if you’re not really looking to move, it’s always important to keep your skills fresh.” That’s what this newsletter is designed to do. I hope that it’ll be helpful for you, no matter where you are in your cloud career, as long as you’re staying in the cloud space.

Corey: And the fact that’s how you view this is the answer to a question a lot of folks have asked me over drinks with theoretical conversations for years of, “Well, Corey, if you went to go work at one of these big cloud providers, it destroy everything you’ve built because how in the world could you be authentic while working for one of these companies?” And the answer is exactly what you’re doing. It’s, “Yeah, the people who pay you don’t own you.” I cannot imagine that even Google could afford to buy your authenticity from you because once that’s gone, you don’t get it back, and you’re one of those people in this space, that—I’m not entirely sure that you understand where you are in this space, so let me help enlighten you with that for a minute.

Forrest: Oh, great. [laugh].

Corey: Oh, yeah, like, the first thing I was starting to talk about that we have in common is that we do a lot of content, both of us and that sometimes occludes the very real fact that we have a distinct level of technical expertise, historically. You and I can both feel relatively deep technical questions about cloud services, but because our job doesn’t have the word engineer in the title, it doesn’t lead to the same type of recognition of that fact. But I want to be very clear: you are technically excellent at what you'll do. You also have a distinguished personality and brand in the space, and your authenticity is also unparalleled. When you say something is good, it is believed that it is because you say it, and the inverse is also true.

You’re also someone that is very clearly aligned with fighting for the user if you want to quote Tron. It’s the, you’re not here to shill for things that don’t get people ahead in their careers; you’re not here to prop things up just because that’s where the money is blowing. Your position on this is unimpeachable. And I’m going to be clear here: I am more interested in Google Cloud now than I was before you made this announcement. That is the value of having someone like you aboard, and frankly, I’m astonished they managed to grab you. It shows a forward-looking ability that historically I have not associated with cloud marketing groups.

Forrest: Yeah, well I mean, the space changes fast. And I think you’ve said this yourself as well, even with the services; you look away for six months and you look back and it’s not the same industry you remember. And that actually is a challenge when you talk about that technical credibility because that can go away very, very quickly. So, it does require some constant effort to stay fresh on that, especially if you’re not building every single day. But to your point about the forward-looking-ness of Google Cloud, I really am excited about that and that’s honestly the biggest thing that attracted me to what they’re doing.

They clearly understand, I think, their position in the space. We know they’re three out of three and trying to catch up, and because of that, they’re able to [laugh] be really creative. They’re able to make bold choices and try things that you might not try if you were trying to maintain a market-leading position. So, that’s exciting to me. I’m a creative person, I like to do things that are outside the box and I think you can look forward to seeing some more outside-the-box things coming at Google Cloud here over the next couple of years.

Corey: I’d be astounded if it were otherwise. The question I have for you is that ‘Head of Cloud’ is not a junior role. That’s not something entry-level that you’re just going to pick some rando off of LinkedIn to fill. They’re going to pick a different rando: you specifically as one of those randos. And to my understanding, you’ve never really touched Google Cloud in anger from a technical level before. Is that right? Am I dramatically misunderstanding, “Oh yeah, you don’t remember the whole musical, and three-act stage play that you put on, and the music video, and the rock opera all about Google Cloud?” It’s, “No, I must have been sick that week,” because that’s the level of prolific you tend to be?

Forrest: [laugh].

Corey: What is your experience with it?

Forrest: That’s yet to come. So, check back on the Google Cloud rock opera; we’ll see if that takes place. So no, I’m going to be learning about Google Cloud. This will be a chance for me to kind of start over a little bit from first principles. In another sense, I’ve been interacting with Google services for years.

Keep in mind that Google Cloud is not just Google Cloud Platform, but it’s G Suite as well, and there’s a lot going on there. So, I definitely am going to be going back to being a beginner a little bit here. They do say if you can teach something to a beginner, you have to really understand it at an expert level. And I know that whether I’m doing this officially on behalf of Google or otherwise, I’m going to be continuing to try to help and educate folks wherever I can. So, it’s going to be incumbent on me, if I want to keep doing that, to go deep quickly and continue to learn.

I’m excited about that challenge. I’ve been doing a lot with AWS for a long time, I don’t know everything. In fact, I know less every day with the amount that they’re continuing to roll out, but this is a chance for me to expand, become a more well-rounded person to see how the other cloud lives. I’m taking that very seriously; I’m not going to be an expert overnight, but stick around, follow me. I’m going to be learning, I’m going to share what I learned, and maybe we’ll all get a little better Google Cloud together.

Corey: The thing I can’t quite get past is that when you told me that you had resigned from A Cloud Guru, I want to be selfish here and say that there were two things that went through my mind. The first was, “Okay, it’s probably AWS. I hope it’s AWS,” because the alternative is you’re going somewhere potentially independent, and I know you keep arguing with me on this point but you are one of the few people I could point out that could start something on the basis of cloud content with a personal brand that I would view as potentially being an audience split for what I do. And it’s, “Oh, you’re going to go work for a big cloud company. That’s awesome. Is it AW—no, it’s not.” And that one threw me for a different loop where it’s, that is very odd because you have identified, clearly, publicly as the leading voice in AWS in many contexts. It just really surprised me. Did you consider looking at AWS as an alternative?

Forrest: I mean first, I don’t know that it’s fair to say that I was a leading voice for AWS. There’s many wonderful people that [crosstalk 00:14:13]—

Corey: To be clear, Forrest, that was not a question. You are a leading voice in the community for AWS and understanding how it works. That is one of those things that no one knows their own reputation. This is one of those areas. Take it from me—a thought leader—that it’s true. Please continue.

Forrest: You have led my thoughts in that direction, so thanks for that, Corey. But to your question, Corey, regarding how did I decide what career move to make, and definitely was a challenge. And it was a struggle for me to say, well, I’m going to leave behind this warm, friendly AWS community that I know, and try something brand new. But it’s not the first time I’ve done something like that in my career. You mentioned already that I spent a number of years as a very, very technical person and I identified strongly as an engineer.

I had multiple degrees in computer science and I had worked as a frontend/backend software engineer, I’d worked as a database administrator, I’d worked as a cloud engineer, and a manager of cloud engineers, and I’d consulted for companies from startups all the way up to the Fortune 50, always on cloud and always very hands-on and writing code. I’ve never had a job where I didn’t have an IDE open and wasn’t writing code every day. And it was a tremendous shock to my system when I started moving away from that, moving a little bit more into the business side of cloud, learning more about marketing, learning how to impact the bottom line of a company in other ways. That was a real challenge, and I went through months where I kind of felt like I was having an identity crisis because if I’m not writing code if I didn’t create YAML today, who am I? Can I call myself an engineer? What worth do I have?

And I know a lot of folks have struggled with this, and a lot of times, I think that’s what sometimes holds people back in their career, saying, “Well, I can only do what I’ve already done because I’ve identified myself so strongly with it.” So, I’m encouraging anyone who’s listening, if you’re at that point where you feel like, “I don’t know if I can leave behind what I know because will I still be able to succeed?” I would encourage you to go ahead and take that step and commit to it if you really believe that you have an opportunity because growth is ultimately going to be a good thing for you. Getting outside your comfort zone and feeling those unpleasant cracks as you start to grow and change into a different person, that ultimately is a strength-building thing.

If you’re not growing, you’re not struggling, you’re not going to be the person that you want to be. So, tying all that back, I went through one round of that already, Corey, when I moved a little bit away from technical delivery. I’m about to go through a second round of that when I move away a little bit farther from the AWS community. I believe that’s going to be a growth opportunity. But yeah, it’s going to be hard.

Corey: It really is. The idea of walking away from the thing that you’ve immersed yourself in is really an interesting thing to think about. Forgive me in advance for the next question; I have to ask it. As a part of your interview process at Google, do they make you write code in a Google Doc?

Forrest: Not as a part of this interview process. I interviewed at Google years ago for a developer advocate position, actually, and made it all the way through their interview process, writing many lines of code in many Google Docs, but not this time.

Corey: Yeah, I confess, I did the same with an SRE job many years ago at Google, and again, you are better at writing code than I am; I did not progress past this stage. But it was moot, honestly, because the way that the interview was conducted, the person I was talking to was so adversarial at the time and so, I got to be honest, condescending that I swore I would never put myself through that process again. But I was also under the impression that the ritualistic algorithmic hazing via whiteboarding code was sort of a requirement for every role at Google. So, things change, times change, people change. I’m gratified to know that was not a part of your interview process.

Forrest: Well, I mean, I think it was more just about the role. My favorite whiteboard interview—

Corey: Nonsense. Every accountant must be able to solve code on a whiteboard.

Forrest: No, I don’t think that’s true. But my favorite whiteboard interview story and I’m sure you have a few, I remember being in an interview with someone—I won’t say who it was or what company it was, but it wasn’t not Google—it was some sort of problem where I was having to lay out, I don’t know, a path for a robot to take through an environment or something like that. And I wrote the code, and it was fine. It was, like, iterative. It was what you would do if you had ten minutes to write something.

And then the interviewer looked at the code, and he said, “Great, now write it again, but don’t use any variables.” And I remember sitting there
for a minute thinking, “In what professional context [laugh] would someone encourage you to do that in a pair programming situation?”

Corey: Right. The response there is, “What the hell does your codebase in production look like?”

Forrest: [laugh]. And of course, the answer is you’re supposed to be using, like, the stack, and it’s kind of like this thought exercise with the local stack. But even if you were to do that, the performance hit would be tremendous. It would not be a wise or logical way to actually write the code. So, it was a pure trivial, kind of like a just academic exercise that they were recommending. And I remember being really turned off by that. So, I guess if you’re considering putting problems like that in your interview process, don’t. They’re not helpful.

Corey: Yeah, I remember hearing at one point one of the Microsoft brain teasers which they’ve since done away with—credit where due—where someone was asked, “How would you go about finding out the weight of a Boeing 747?” And the person responded with the exact weight of a Boeing 747 because their previous job had been at Boeing for seven years. And that was apparently not what they were expecting to hear. But yeah, it’s sort of an allegory as well for, first, this has no bearing on your ability to do the job, and two, expertise is important. There’s a lot of ways I could try and Hacker News first principles my way through something like that, but the easier answer is for me to call someone at Boeing and ask them, or Google it, depending on exactly how precise I need to be and whether lives hang in the balance of the [laugh] answer to the question. That’s a skill that seems lost somewhere, too.

Forrest: Yeah, and this takes us all the way back to the conversation about The Cloud Resume Challenge, Corey. And why it works is it takes the burden of proof off of you in the interview, or the burden of proof off the interviewer to have to come up with some kind of trivial problem that you’ve done under time pressure, and instead, it lets the conversation flow naturally back to, “Well, what have you done? Tell me about a story about a problem that you have solved, a challenge you ran into, and how you got past it.” That’s all work that has taken place prior to the interview that you’ve reflected on, that’s built you as a person and as an engineer, even if you don’t necessarily have professional experience. That’s how I try to conduct interviews and I think it’s a much healthier and more sustainable way to find people that you’ll like to work with.

Corey: Is this going to be your first outing at a giant multinational tech company?

Forrest: No, although it will be my first time with a public company. When I worked at Infor, Infor was the largest privately owned software company in the world. I don’t know if that’s still technically true or not, but it’ll be my first time with a publicly-traded company.

Corey: Fantastic. The nice thing from my perspective is it gives me a little bit more context into what companies can and can’t do, and how things are structured. It feels like your content—I mean, the music videos and things and whatnot that you do—I mean, you have something that I don’t, which is commonly known as musical talent. And that’s great. I can write funny lyrics, but you are not just able to write lyrics, you’re able to perform, you’re able to sing, the unanswered question for the entire interview right now is whether you can also dance. So, we’re going to find that out at some point.

Forrest: You would think that I could, Corey. I definitely seem like someone who should be able to tap dance. I regret to tell you that I can’t, but I want to learn.

Corey: For a lot of this, it’s clearly you’re doing this in front of your own piano with a microphone in front of you, doing it live, and having a—I don’t know if it is a built-in webcam to a laptop that’s sitting in front of you or something else, but—

Forrest: I’m playing with that.

Corey: Yeah, well don’t take this the wrong way; it’s not a high definition 4k camera, et cetera. It’s the Lightning’s—eh, it’s your home office. You’re comfortable there. It’s not a studio. What I’m most excited about—from my perspective, I know what you’re excited about—but you’re now going to be producing content for Google and I checked the numbers in preparation for this interview.

It’s okay, can Google wind up affording a production house of some sort to work on your videos to upscale the production value of some of what you’re doing? And I have checked; it is not the likeliest scenario—and I have no inside knowledge for those who are trying to trade on
this—but yes, it turns out that Google could, in fact, shore up your content by buying you Disney.

Forrest: I think that’s technically true, and I do expect that to happen in the next three to six months, so that is completely inside information.

Corey: Oh, exactly. Have reasonable expectations, but you could let it go as long as a year because that’s when the first annual review cycle comes in and you want to give people time to let that clear through M&A and make sure that they are living up to their commitments to you, of course.

Forrest: That’s right, yeah. We’re just about to go into the quiet period there. No, but kind of to that point, though, and you bring up the amateurish quality of a lot of these videos that I put together in terms of the lighting and the staging, and everything else. And I am doing a little bit to help with that. Like, it would be great if you could see—

Corey: To be clear, that is not a criticism. I’m in the same boat as you are on this. It’s—[laugh]—

Forrest: So, far from a criticism, it’s actually pretty deliberate. The fact of the matter is, there’s something very raw, very authentic about just seeing someone sitting in their house, at their piano, playing and singing. There’s no tricks, there’s no edits, there’s no glitz, there’s no makeup team behind the scenes, there’s no one who’s involved with this other than just me caring a lot about something and sitting down and singing about it. And I think some of that is what helps come across to people and it helps these things travel. So yeah, I’m looking forward a lot to being able to collaborate with other fantastic people at Google, and I can’t exactly promise what will come out of that, but I’m quite sure there will be more fun content to come.

But I hope never to lose that, kind of, DIY sensibility. Because, again, my background is as an engineer, and the things I create, whether it’s music, whether it’s cartoons, whether it’s books, or other things I write, I never want to lose that sense of just excitement about the technologies I’m working with and the fact that I get to use the tools that are available at my disposal to share them with you as directly and honestly and humanly as possible.

Corey: Up next we’ve got the latest hits from Veem. Its climbing charts everywhere and soon its going to climb right into your heart. Here it is!

Corey: No matter how hard you try, you’re not able to hide the sheer joy you take from even talking about this sort of stuff, and I think that’s a powerful lesson. For folks listening to this who want to expand into their own content story and approach things that they find interesting in a way that they enjoy, don’t try and do what I do; don’t try to do what Forrest does; do the thing that makes you happy. I would love to be able to sing, but I can’t. I can write funny lyrics, but those don’t do well in pure text form. I’m fortunate that I was able to construct a structure on my end where I can pay people who do know how to sing—like Adeem the Artist and many more—to participate in a lot of the things that I get to work on.

But find the way that you want to express things and do you. You’re only ever going to be second best at being Forrest or being Corey, but you’re always going to be number one at being whoever you happen to be. I think that’s a lesson that gets overlooked an awful lot.

Forrest: Yeah, I’ve been playing with this thought for a while that the only real [moat 00:24:24] out there is originality, is your personality. Everything else can be cloned, but you are an individual. And I mean that to us specifically, Corey, and also the general ‘you’ to anybody listening to this. So, find what makes you tick. It sounds like the most cliche device in the world, but another way, it’s also the only useful advice that’s out there.

Corey: I want to be clear, you don’t work there yet and I’m not here to effectively give undue praise to large companies, but I just want to say again how the sheer vision of hiring you is just astounding to me. That it makes perfect sense, don’t get me wrong, but because I know that every large company, somewhere, at some point, internally has had a conversation of, “We really should hire Corey, except…” well, I’ve got to level with you, Corey without the except parts looks an awful lot like you.

Forrest: Yeah, you know, you brought up earlier this idea that well, hopefully, Forrest doesn’t lose his authenticity at Google. And one of the things that I appreciate about the team that I’ve talked to there so far, is that they really do understand the power of individuals and voices. And so that’s not going to happen. You know, my authenticity is not for sale. And frankly, I’m useless without it, so it wouldn’t be in anyone’s best interest to buy it anyway. And that would be true for you as well, Corey. Whatever you end up doing, whether you someday ascend to the head of AWS Marketing, as is apparently your divine destiny, I know that—

Corey: Well, I’m starting to worry that there’s not too many people left in that org, so I’m worried people took me seriously and they think I’ve got this in hand or something.

Forrest: You may be the last man standing for all we know. You may be able to go in and just, kind of, do this non-hostile takeover where there’s just no one there to defend against you, anymore.

Corey: Well, speaking about takeovers and whatnot, we talk about Google acquiring Disney so you now have a production studio on this. But let’s talk about actual hard problems you’re going to be solving there. Do you think you can bring back Google Reader?

Forrest: That would be my dream. I have no inside knowledge of what would even be required to bring that off, but I think it’s obvious that it’s not just about that particular product that people like—because yes, you or I could go make a startup and create something that did what Google Reader did—but it’s about what it represents. It’s about the commitment that it would mean to Google’s customers and to their products. So yeah, something like bring Google Reader back would be a wonderful thing for everyone that subscribes to Google but it would also be a fantastic storytelling element for Google as well. So yes, I’d be entirely in favor of something like that. I hope we can make it happen someday.

Corey: Oh, as would I. YOu’re in Brian Hall’s org, correct?

Forrest: Yes.

Corey: Brian is a man who was the VP of Product Marketing over at AWS, went to Google for the same role, was sued by AWS under the auspices of a non-compete, which is just the most ridiculous thing in the world, and I want to be very clear here, you can say an awful lot about Brian Hall. I say an awful lot about Brian Hall. AWS says a lot about Brian Hall in very poorly conceived depositions and lawsuits that should never have been allowed to continue, and at least have an editor go over them, but that’s a separate problem. But one thing you cannot say about Brian is that he is not incredibly intelligent. And the way that I find that manifesting is, I do not accept that he is someone with such a limited vision that he would be prepared to even entertain the idea of hiring you without giving you what amounts to effectively full creative control of the things you’re going to be working on.

You are not someone it would make any sense to hire and then try and shove into a box. That is my assessment of everything I’ve read on every conversation I’ve had with Googlers in the marketing org; it all speaks to something like this. Was that your impression during the interview? Specifically that you have carte blanche, not that Brian is smart. You’re about to be in his org; you’re obligated to say it. That’s okay. We’ll meet at the bar until the real Brian stories later but I’m talking about their remit here.

Forrest: No, my authenticity is not for sale, but at the same time. I am a big fan of Brian’s and have been since his AWS days, which was honestly one of the big reasons why I ended up joining his org. But yeah, to your question about what is that role going to look like, day to day, of course obviously, that remains to be seen, but it is my understanding that it will have a consultative element and that I will have some opportunity to help to drive some influence across some different teams. Something that I’ve learned as I’ve grown in my career a little bit and I’ve moved into more of management type of roles is that the people that report to you are such a small fraction of the overall influence that you should be having to be really successful in a role like that, any kind of leadership role, so much more of your leadership is going to happen indirectly and by influence, and it’s going to happen slowly over time, as you build support for what you’re doing and you start to show value and encourage other people to come around to your side. That’s just the reality of making change in large organizations.

And of course, this is by far the largest organization I’ve ever worked in, so I know it’s going to take time. But my understanding is I do have a little bit of leeway to bring some of my ideas in, and I’m excited about that, and you can sort of judge for yourself, how successful I am, over time.

Corey: My last question for you is that sort that has the potential to get you in trouble, except I think I’m going to agree with your answer to this. Do you believe that they’re going to Google Reader Google Cloud?

Forrest: If I believed that I wouldn’t be joining? So obviously, no, I don’t believe that.

Corey: I have to confess that for the longest time, I was convinced that this was yet another Google misadventure, where they were going to dabble with it, sort of half-ass it, and then shut it down. Because that seems to be the fate of so many Google products out there. The first AWS service that entered beta was Simple Queuing Service. What is a queue but a messaging system, and we know how Google treats messaging products. Same problem; same story.

I have to say over the last year or so, my perspective has evolved considerably. They are signing ten-year deals with very large banks; they are investing heavily in hiring, in R&D, in marketing clearly, in a bunch of different areas that are doing the right thing for the long-term. The financial analysts like to beat Google Cloud up because I think two quarters ago, they showed a $5 billion loss, either for the year or for the quarter, and, “It’s not making money.” It’s, “No. Given Google’s position in the market, I’d be horrified if it were. The only way it shouldn’t be turning a profit is if there’s nowhere left to invest in the platform.”

They’re making the investments, they’re doing the right things. And I have to say I’ve gone from, “I don’t know if I would trust that without an exodus plan,” to, “Yeah, you should have a theoretical exodus plan the same way you should with any provider, but it’s not the sort of thing that I feel the need to yank away on 30-days’ notice.” I have crossed that bridge myself. In all sincerity, cheap, easy jokes aside, it’s clear to me from what I’ve seen that Google Cloud is going to be around for the long term. Now, we are talking long-term in terms of tech companies, not 150-year-old companies based in Europe, but we can aspire to it. I expect it to outlive me, and not just because I have a big mouth and piss off large companies.

Forrest: Yeah. Some of my closest friends and longest-tenured colleagues, people I’ve worked with for years are GCP engineers, people who are not working for GCP, but they’re building on GCP services at various companies. And they always come to me and I’ve noticed a steady increase in this over the past, I would say 12 to 18 months where they say, “I love working on GCP. I love these services. I love the way the IAM is designed. I love the way the projects are put together. It just feels right. It feels natural to me. It scratches some sort of an itch in my engineering brain.”

And then they pause and they say, “Why don’t more people get this? Why don’t more people understand this story?” That’s a problem that I can help to solve. So, I’m really excited about helping to tell the story of Google Cloud. And yeah, that chapter is just about to be written.

Corey: I can’t wait to see what happens next. If people want to learn more about what you’re up to, and how you’re approaching these things, and sign up for your various newsletters, where’s the entry point? Where can they find you?

Forrest: I would say go to my Twitter. I’m on Twitter @forrestbrazeal and there’ll be a link in my bio that has links to all the things we’ve mentioned: The Cloud Resume Challenge Book, my other extremely bizarre book about cloud which is called The Read Aloud Cloud. And there you can sign up for that Best Jobs in Cloud newsletter and all the other things we talked about. So, I’ll see you there.

Corey: I look forward to including those links in the [show notes 00:32:24]. That’s how I wind up expressing my support for all of my guests’ nonsense, but particularly yours. Forrest, thank you so much for taking the time to speak with me.

Forrest: Much appreciated, Corey. Always a pleasure.

Corey: Forrest Brazeal, currently unemployed, but by the time you listen to this, the Head of Content at Google Cloud. I am Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with a long, obnoxious, insulting comment, and then rewrite the entire insulting comment without using vowels.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Ant

Ant Co-founded A Cloud Guru, ServerlessConf, JeffConf, ServerlessDays and now running Senzo/Homeschool, in between other things. He needs to work on his decision making.

Links:

  • A Cloud Guru: https://acloudguru.com
  • homeschool.dev: https://homeschool.dev
  • aws.training: https://aws.training
  • learn.microsoft.com: https://learn.microsoft.com
  • Twitter: https://twitter.com/iamstan

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Thinkst. This is going to take a minute to explain, so bear with me. I linked against an early version of their tool, canarytokens.org in the very early days of my newsletter, and what it does is relatively simple and straightforward. It winds up embedding credentials, files, that sort of thing in various parts of your environment, wherever you want to; it gives you fake AWS API credentials, for example. And the only thing that these things do is alert you whenever someone attempts to use those things. It’s an awesome approach. I’ve used something similar for years. Check them out. But wait, there’s more. They also have an enterprise option that you should be very much aware of canary.tools. You can take a look at this, but what it does is it provides an enterprise approach to drive these things throughout your entire environment. You can get a physical device that hangs out on your network and impersonates whatever you want to. When it gets Nmap scanned, or someone attempts to log into it, or access files on it, you get instant alerts. It’s awesome. If you don’t do something like this, you’re likely to find out that you’ve gotten breached, the hard way. Take a look at this. It’s one of those few things that I look at and say, “Wow, that is an amazing idea. I love it.” That’s canarytokens.org and canary.tools. The first one is free. The second one is enterprise-y. Take a look. I’m a big fan of this. More from them in the coming weeks.

Corey: This episode is sponsored in part my Cribl Logstream. Cirbl Logstream is an observability pipeline that lets you collect, reduce, transform, and route machine data from anywhere, to anywhere. Simple right? As a nice bonus it not only helps you improve visibility into what the hell is going on, but also helps you save money almost by accident. Kind of like not putting a whole bunch of vowels and other letters that would be easier to spell in a company name. To learn more visit: cribl.io

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Every once in a while I talk to someone about, “Oh, yeah, remember that time that you appeared on Screaming in the Cloud?” And it turns out that they didn’t; it was something of a fever dream. Today is one of those guests that I’m, frankly, astonished I haven’t had on before: Ant Stanley. Ant, thank you so much for indulging me and somehow forgiving me for not having you on previously.

Ant: Hey, Corey, thanks for that. Yeah, I’m not too sure why I haven’t been on previously. You can explain that to me over a beer one day.

Corey: Absolutely, and I’m sure I’ll be the one that buys it because that is just inexcusable. So, who are you? What do you do? I know that you’re a Serverless Hero at AWS, which is probably the most self-aggrandizing thing you can call someone because who in the world in their right mind is going to introduce themselves that way? That’s what you have me for. I’ll introduce you that way. So, you’re an AWS Serverless Hero. What does that mean?

Ant: So, the Serverless Hero, effectively I’ve been recognized for my contribution to the serverless community, what that contribution is potentially dubious. But yeah, I was one of the original co-founders of A Cloud Guru. We were a serverless-first company, way back when. So, from 2015 to 2016, I was with A Cloud Guru with Ryan and Sam, the two other co-founders.

I left in 2016 after we’d run ServerlessConf. So, I led and ran the first ServerlessConf. And then for various reasons, I decided, hey, the pressure was too much; I needed a break, and a few other reasons I decided to leave A Cloud Guru. A very amicable split with my former co-founders. And then yeah, I kind of took a break, took some time off, de-stressed, got the serverless user group in London up and running; ran a small conference in London called JeffConf, which was a take on a blog that Paul Johnson, who was one of the folks who ran JeffConf with me, wrote a while ago saying we could have called it serverless—and we might as well have called it Jeff. Could have called it anything; might as well have called it Jeff. So, we had this joke about JeffConf. Not a reference to Mr. Bazos.

Corey: No, no. Though they do have an awful lot of Jeffs working over there. But that’s neither here nor there. ‘The Land of the Infinite Jeffs’ as it were.

Ant: Yeah, exactly. There are more Jeffs than women in the exec team if I remember correctly.

Corey: I think it’s now it’s a Dave problem instead.

Ant: Yeah, it’s a Dave problem. Yeah. [laugh]. It’s not a problem either way. Yeah. So, JeffConf morphed into SeverlessDays, which is a group of community events around the world. So, I think AWS said, “Hey, this guy likes running serverless events for some silly reason. Let’s make him a Serverless Hero.”

Corey: And here we are. Which is interesting because a few directions you can take this in. One of them, most recently, we were having a conversation, and you were opining on your thoughts of the current state of serverless, which can succinctly be distilled down to ‘serverless sucks,’ which is not something you’d expect to hear from a Serverless Hero—and I hope you can hear the initial caps when I say ‘Serverless Hero’—or the founder of a serverless conference. So, what’s the deal with that? Why does it suck?

Ant: So, whole serverless movement started to gather momentum in 2015. The early adopters were all extremely experienced technologists, folks like Ben Kehoe, the chief robotics scientist at iRobot—he’s incredibly smart—and folks of that caliber. And those were the kinds of people who spoke at the first serverless conference, spoke at all the first serverless events. And, you know, you’d kind of expect that with a new technology where there’s not a lot of body of knowledge, you’d expect these high-level, really advanced folks being the ones putting themselves out there, being the early adopters. The problem is we’re in 2021 and that’s still the profile of the people who are adopting serverless, you know? It’s still not this mass adoption.

And part of the reason for me is because of the complexity around it. The user experience for most serverless tools is not great. It’s not easy to adopt. The patterns aren’t standardized and well known—even though there are a million websites out there saying that there are serverless patterns—and the concepts aren’t well explained. I think there’s still a fair amount of education that needs to happen.

I think folks have focused far too much on the technical aspects of serverless, and what is serverless and not serverless, or how you deploy something, or how you monitor something, observability, instead of going back to basics and first principles of what is this thing? Why should you do it? How do you do it? And how do we make that easy? There’s no real focus on user experience and adoption for inexperienced folks.

The adoption curve, the learning curve for serverless, no matter what platform you do, if you want to do anything that’s beyond a side project it’s really difficult because there’s no easy path. And I know there’s going to be folks that are going to complain about it, but the Serverless Stack just got a million dollars to solve this problem.

Corey: I love the Serverless Stack. They had a great way of building things out.

Ant: Yeah.

Corey: I cribbed a fair bit from what they built when I was building out my own serverless project of the newsletter production pipeline system. And that’s awesome. And I built that, and I run it mostly as a technology testbed. But my website, lastweekinaws.com?

I pay WP Engine to host it on WordPress and the reason behind that is not that I can’t figure out the serverless pieces of it, it’s because when I want to hire someone to do something that’s a bit off the beaten path on WordPress, I don’t have to spend $400 an hour for a consultant to do it because there’s more than 20 people in the world who understand how all this stuff fits together and integrates well. There’s something to be said for going in the direction the rest of the market is when there’s not a lot of reason to differentiate yourselves. Yeah, could I save thousands of dollars a year in infrastructure costs if I’d gone with serverless? Of course, but people’s time is worth more than that. It’s expensive to have people work on these things.

And even on the serverless stuff that I’ve built, if it’s been more than six months since I’ve touched a component, someone else may have written it; I have to rediscover what the hell I was thinking and what the constraints are, what the constraints I thought existed there in the platform. And every time I deal with Lambda or API Gateway, I come away with a spiraling sense of complexity tied to all of it. And the vision of serverless I believe in, truly, but the execution has lagged from all providers.

Ant: Yeah. I agree with that completely. The execution is just not there. I look at the situation—so Datadog had their report, “The State of Serverless Report” that came out about a month or two ago; I think it’s the second year they’ve done it, now, might be the third. And in the report, one of the sections, they talked about tooling.

And they said, “What’s the most adopted tools?” And they had the Serverless Framework in there, they had SAM in there, they had CloudFormation, I think they had Terraform in there. But basically, Serverless Framework had 70% of the respondents. 70% of folks using Datadog and using serverless tools were using Serverless Framework. But SAM, AWS’s preferred solution, was like 12%.

It was really tiny and this is the thing that every single AWS demo example uses, that the serverless developer advocates push heavily. And it’s the official solution, but the Serverless Application Model is just not being adopted and there are reasons for that, and it’s because it’s the way they approach the market because it’s highly opinionated, and they don’t really listen to end-users that much. And their CDK out there. So, that’s the other AWS organizational complexity as well, you’ve got another team within AWS, another product team who’ve developed this different way—CDK—doing things.

Corey: This is all AWS’s fault, by the way. For the longest time, I’ve been complaining about Lambda edge functions because they are not at all transparent; you have to wait for a CloudFront deployment for it to update every time, only to figure out that in my case, I forgot a comma because I’ve never heard of a linter. And it becomes this awful thing. Only recently did I find out they only run at regional edge caches, not just in all of the CloudFront pop, so I said, “The hell with it,” ripped it out of everything I was using it with, and wound up implementing it in bog-standard Lambda because it was easier. But then rather than fixing that, they’ve created their—what was it—their CloudFront Workers. Or is it—is it CloudFront Workers, or is it CloudFront Functions?

Ant: No, CloudFront Functions.

Corey: I don’t even remember it because rather than fixing the thing, you just released a different thing that addresses these problems in very different ways that aren’t directly compatible. And it’s oh, great, awesome. Terrific. As a customer, I want absolutely not this. It’s one of these where, honestly, I’ve left in many cases with the resigned position of, if you’re not going to take this seriously, why am I?

Ant: Yeah, exactly. And it’s bizarre. So, the CloudFront Functions thing, it’s based on Nginx’s [little 00:08:39] JavaScript engine. So, it’s the Nginx team supporting it—the engine—which is really small number of users; it’s tiny, there’s no foundation behind it. So, you’ve got these massive companies reliant on some tiny organization to support the runtime of one of their businesses, one of their services.

And they expect people to adopt it. And on top of that, that engine supports primary language is JavaScript’s ES5 or ES2015, which is the 2015 edition of JavaScript, so it’s a six-year-old version of JavaScript. You cannot use one JavaScript with it which also means you can’t use any other tools in the JavaScript ecosystem for it. So basically, anything you write for that is going to be vanilla, you’re going to write yourself, there’s no tooling, no community to really leverage to use that thing. Again, like, why have you even done that? Why if you now gone off and taken an engine no one uses—they will say someone uses it, but basically no one uses—

Corey: No one willingly uses or knowingly uses it.

Ant: Yeah. No one really uses. And then decided to run that. Why not look at WebAssembly—it’s crazy—which has a foundation behind it and they’re doing great things, and other providers are using WebAssembly on the edge. I just don’t understand the thought process—well, I say I don’t understand, but I do understand the thought processes behind Amazon. Every single GM in Amazon is effectively incentivized to release stuff, and build stuff, and to get stuff out the door. That’s how they make money. You hear the stories—

Corey: Oh, it’s been clear for years. They only recently stopped—in their keynotes every year—talking about the number of feature releases that they’ve had over the past 12 months. And I think they finally had it clued into them by someone snarky on Twitter—ahem—that the only people that feel good about that are people internal to AWS because customers see that and get horrified of, “I haven’t kept up with most of those things. How many of those are important? How many of them are nonsense?”

And I’m sure somewhere you have released a serverless that will solve my business problem perfectly so I don’t have to build it together myself out of Lambda functions, and string, and popsicle sticks, but I’ll never hear about it because you’re too busy talking about nonsense. And that problem still exists and it’s writ large. There’s a philosophy around not breaking existing workloads—which I get; that’s a hard problem to solve for—but their solution is, rather than fixing existing services will launch a new one that doesn’t have those constraints and takes a different approach to it. And it’s horrible.

Ant: Yeah, exactly. If you compare Amazon to Apple, Apple releases a net-new product once a year, once every two years.

Corey: You’re talking about new generations of products, that comes out on an annualized basis, but when you’re talking about actual new
product, not that frequently. The last one—

Ant: Yeah.

Corey: —I can really think of is probably going to be AirPods, at least of any significance.

Ant: AirTags is the new one.

Corey: Oh, AirTags. AirTags is recent, which is a neat—but it’s an accessory to the rest of those things. It is—

Ant: And then there’s AirPods. But yeah, it’s once—because they—everything works. If you’re in that Apple ecosystem, everything works. And everything’s back-ported and supported. My four-year-old phone still works and had a five-year-old MacBook before this current one, still worked, you know, not a problem.

And those two philosophies—and the Amazon folk are heavily incentivized to release products and to grow the usage of those products. And they’re all incentivized within their bubbles. So, that’s why you get competing products. That’s why Proton exists when CodeBuild and
CodePipeline, and all of those things exist, and you have all these competing products. I’m waiting for the container team to fully recreate AWS on top of containers. They’re not far away.

Corey: They’re already in the process of recreating AWS on top of Lightsail. It’s more or less the, “Oh, we’re making this the simpler version.” Which is great. You know who likes simplicity? Freaking everyone.

So, it’s the vision of a cloud, we could have had but didn’t. “Oh, you want a virtual machine. Spin up a Lightsail instance; you’re going to get a fixed amount of compute, disk, RAM, and CPU that you can adjust, and it’s going to cost you a flat fee per month until you exceed some fairly high limits.” Why can’t everything be like that, on some level? Because in many cases, I don’t care about wanting to know exactly to the penny shave things off.

I want to spin up a fleet of 20 virtual machines, and if they cost me 20 bucks a pop each a month, I can forecast that, I can budget for that, I can do a lot and I don’t actually care in any business context about the money there, but dialing it in and having the variable charges and the rest, and, “Oh, you went through a managed NAT gateway. That’s going to double your bandwidth price and it’s going to be expensive. Surprise, you should have looked more closely at it,” is sort of the lesson of the original AWS services. At some level, they’ve deviated away from anything resembling simplicity and increasingly we’re seeing a world where in order to do something effectively with cloud, you have to spend 12 weeks going to cloud school first.

Ant: Oh, yeah. Completely. See, that’s one of the major barriers with serverless. You can’t use serverless for any of the major cloud providers until you understand that cloud provider. So yeah, do your 12 weeks of cloud school. And there’s more than enough providers.

Corey: Whoa, whoa, whoa. Before you spin up a function that runs code, you have to understand the identity and security model, and how the network works, and a bunch of other ancillary nonsense that isn’t directly tied to business value.

Ant: And all these fun things. How are you’re going to test this, and how are you’re going to do all that?

Corey: How do you write the entry point? Where is it going to enter? What is it expecting? What objects are getting passed in, if any? What format is it going to take?

I’ve spent days, previously, trying to figure out the exact invocation for working with a JSON object in Python, what that’s going to show up as, and how specifically to refer to it. And once you’ve done that a couple of times, great, fine, it’s easy. Copy and paste it from the last time you did it. But figuring it out from first principles, particularly in a time when there isn’t a lot of good public demonstrations of this—especially early days—it’s hard to do.

Ant: Yeah. And they just love complexity. Have you looked at the second edition—so the third version of the AWS SDK for JavaScript?

Corey: I don’t touch JavaScript with my hands most days, just because I’m bad at it and I don’t understand the asynchronous model and computers are really not my thing most.

Ant: So, unfortunately for my sins, I do use JavaScript a lot. So, version two of the SDK is effectively the single most popular Cloud SDK of any language, anything out there; 20 million downloads a week. It’s crazy. It’s huge—version two. And JavaScript’s a very fast-evolving language, though.

Basically, it’s a bit like the English language in that it adopts things from other languages through osmosis, and co-opts various other features of other languages. So, JavaScript has—if there’s a feature you love in your language, it’s going to end up in JavaScript at some point. So, it becomes a very broad Swiss Army knife that can do almost anything. And there’s always better ways to do things. So, the problem is, the version two was written in old JavaScript from years twenty fifteen years five years six kind of level.

So, from 2015, 2016, I—you know, 2020, 2021, JavaScript has changed. So, they said, “Oh, we’re going to rewrite this.” Which good; you should do. But they absolutely broke all compatibility with version two. So, there is no path from version two to version three without rewriting what you’ve got.

So, if you want to take anything you’ve written—not even serverless—anything in JavaScript you’ve written and you want to upgrade it to get some of the new features of JavaScript in the SDK, you have to rewrite your code to do that. And some instances, if you’re using hexagonal architecture and you’re doing all the right things, that’s a really small thing to do. But most people aren’t doing that.

Corey: But let’s face it, a lot of things grow organically.

Ant: Yeah.

Corey: And again, I can sit here and tell you how to build things appropriately and then I look at my own environment and… yeah, pay no attention to that burning dumpster fire behind the camera. And it’s awful. You want to make sure that you’re doing things the right way but it’s hard to do and taking on additional toil because the provider decides the time to focus on this is a problem.

Ant: But it’s completely not a user-centric way of thinking. You know, they’ve got all their 14—is it 16 principles now? Did they add two principles, didn’t they?

Corey: They added two to get up to 16; one less than the numbers of ways to run containers in AWS.

Ant: Yeah. They could barely contain themselves. [laugh]. It’s just not customer-centric. They’ve moved themselves away from that customer-centric view of the world because the reality is, they are centered on the goals of the team, the goals of the GM, and the goals of that particular product.

That famous drawing of all the different organizational charts, they got the Facebook chart, and the Google Chart, and the Amazon chart has all these little circles, everyone pointing guns at each other. And the more Amazon grows, the more you feel like that’s reality. And it’s hurting users, it’s massively hurting users. And we feel the pain every day, absolutely every day, which is not great. And it’s going to hurt Amazon in the long run, but short-term, they’re not going to see that pain quarterly, they’re not going to see that pain, probably within 12 months.

But they will see the pain long run. And if they want to fix it, they probably should have started fixing it two years ago. But it’s going to take years to fix because that’s a massive cultural shift to say, “Okay, how do we get back to being more customer-focused? How do we stop that organizational targets and goals from getting in the way of delivering value to the customer?”

Corey: It’s a good question. The hard part is getting customers to understand enough of what you put out there to be able to disambiguate what you’ve built, and what parts to trust, what parts not the trust, what parts are going to be hard, et cetera, et cetera, et cetera, et cetera. The concern that I’ve got across the board here is, how do you learn? How do you get started with this? And the way that I came into this was I started off, in the early days of AWS, there were a dozen services, and okay, I could sort of stumble my way through it.

And the UI was rough, but it got better with time. So, the answer for a lot of folks these days is training, which makes sense. In the beginning, we learned through things like podcasts. Like there was a company called Jupiter Broadcasting which did a bunch of Linux-oriented podcasts and learned how this stuff works. And then they were acquired by Linux Academy which really focused on training.

And then A Cloud Guru acquired Linux Academy. And then Pluralsight acquired A Cloud Guru and is now in the process of itself being acquired by Vista Equity Partners. There’s always a bigger fish eating something somewhere. It feels like a tremendous, tremendous consolidation in the training market. Given that you were one of the founders of A Cloud Guru, where do you stand on that?

Ant: So, in terms of that actual transaction, I don’t know the details because I’m a long time out of A Cloud Guru, but I’ve stayed within the whole training sphere, and so effectively, the bigger fish scenario, it’s making the market smaller in terms of providers are there. You really don’t have many providers doing cloud-specific training anymore. On one level you don’t, but then another level, you’ve got lots of independent folks doing tons of stuff. So, you’ve got this explosion at the bottom end. If you go to Udemy—which is where A Cloud Guru started, on Udemy—you will see tons of folks offering courses at ten bucks a pop.

And then there’s what I’m doing now on homeschool.dev; there’s serverless-focused training on there. But that’s really focused on a really small niche. So, there’s this explosion at the bottom end of lots of small people doing lots of things, and then you’ve got this consolidation at the top end, all the big providers buying each other, which leaves a massive gap in the middle.

And on top of that, you’ve got AWS themselves, and all the other cloud providers, offering a lot of their own free training, whether it’s on their own platforms—there’s aws.training now, and Microsoft have similar as well—I think it’s learn.microsoft.com is theirs. And you’ve got all these
different providers doing their own training, so there’s lots out there.

There’s actually probably more training for lower costs than ever before. The problem is, it’s like the complexity of too many services, it’s the 17 container problem. Which training do you use because the actual cost of the training is your time? It’s not the cost of the course. Your time is always going to be more expensive.

Corey: Yeah, the course is never going to be anywhere comparable to the time you spend on it. And I’ve never understood, frankly, why these large companies charge money for training on their own platform and also charge money for certifications because I don’t care what you’re going to pay for those things, once you know a platform well enough to hit a certification, you’re going to use the thing you know, in most cases; it’s a great bottom-up adoption story.

Ant: Yeah, completely. That was actually one of Amazon’s first early problems with their trainings, why A Cloud Guru even exists, and Linux Academy, and Cloud Academy all actually came into being is because Amazon hired a bunch of folks from VMware to set up their training program. And VMware’s training, back in the day, was a profit center. So, you’d have a one-and-a-half thousand, two thousand dollar training course you’d go on for three to five days, and then you’d have a couple hundred dollars to do the certification. It was a profit center because VMware didn’t really have that much competition. Zen and Microsoft’s Hyper V were so late to the market, they basically own the market at the time. So—

Corey: Oh, yeah. They still do in some corners.

Ant: Yeah. They’re still massively doing in this place as they still exist. And so they Amazon hired a bunch of ex-VMware folk, and they said, “We’re just going to do what we did at VMware and do it at Amazon,” not realizing Amazon didn’t own the market at the time, was still growing, and they tried to make it a profit center, which basically left a huge gap for folks who just did something at a reasonable price, which was basically everyone else. [laugh].

This episode is sponsored by our friends at Oracle Cloud. Counting the pennies, but still dreaming of deploying apps instead of "Hello, World" demos? Allow me to introduce you to Oracle's Always Free tier. It provides over 20 free services and infrastructure, networking databases, observability, management, and security.

And - let me be clear here - it's actually free. There's no surprise billing until you intentionally and proactively upgrade your account. This means you can provision a virtual machine instance or spin up an autonomous database that manages itself all while gaining the networking load, balancing and storage resources that somehow never quite make it into most free tiers needed to support the application that you want to build.

With Always Free you can do things like run small scale applications, or do proof of concept testing without spending a dime. You know that I always like to put asterisks next to the word free. This is actually free. No asterisk. Start now. Visit https://snark.cloud/oci-free that's https://snark.cloud/oci-free.

Corey: The challenge I found with a few of these courses as well, is that they teach you the certification, and the certifications are, in some
ways, crap when it comes to things you actually need to know to intelligently use a platform. So, many of them distill down not to the things you need to know, but to the things that are easy to test in a multiple-choice format. So, it devolves inherently into trivia such as, “Which is the right syntax for this thing?” Or, “Which one of these CloudFormations stanzas or functions isn’t real?” Things like that where it’s, no one in the real world needs to know any of those things.

I don’t know anyone these days—sensible—who can write CloudFormation from scratch without pulling up some reference somewhere because most people don’t have that stuff in their head. And if you do, I’d suggest forgetting it so you can use that space to remember something that’s more valuable. It doesn’t make sense for how people interact with these things. But I do see the value as well in large companies trying to upskill thousands and thousands of people. You have 5000 people that are trying to come up to speed because you’re migrating into cloud. How do you judge people’s progress? Well, certifications are an easy answer.

Ant: Yeah, massively. Probably the most successful blog post ever written—I don’t think it’s up anymore, but it was when I was at A Cloud Gurus—like, what’s the value of a certification? And ultimately, it came down to, it’s a way for companies that are hiring to filter people easily. That’s it. That’s really it. It’s if you’ve got to hire ten people and you get 1000 CVs or resumes for those ten roles, first thing you do is you filter by who’s certified for that role. And then you go through anything else. Does the certification mean you can actually do the job? Not really. There are hundreds of people who are not cer—thousands, millions of people who are not certified to do jobs that they do. But when you’re getting hired and there’s lots of people applying for the same role, it’s literally the first thing they will filter on. And it’s—so you want to get certified, it’s hard to get through that filter. That’s what the certification does, it’s how you get through that first filter of whatever the talent tracking system they’re using is. That’s it. And how to get into the dev lounge at re:Invent.

Corey: Oh yeah, that’s my reason for getting a certification, originally. And again, for folks who learn effectively that way, I have no problem with people getting certifications. If you’re trying to advance in your career, especially early stage, and you need a piece of paper that says you know what you’re talking about, a certification is a decent approach. In time, with seniority, that gets replaced by a piece of paper, it’s called your resume or your CV, but that is a longer-term more senior-focused approach. I don’t begrudge people getting certifications and I don’t think that they’re foolish for doing it.

But in time, it feels like the market for training is simultaneously contracting into only a few players left, and also, I’m curious as to whether or not the large companies out there are increasing their spend with the training providers or not. On the community side, the direct-to-consumer approach, that is exploding, but at the same time, you’re then also dealing—forgive me, listeners—with the general public and there is nothing worse than a customer, from a customer service perspective, who was only paying a little money to you. I used to work in a web hosting company that $3,000 a month customers were great to work with. The $2999 a month customers were hell on earth who expected that they were entitled to 80 hours a month of systems engineering time. And you see something similar in the training space. It’s always the small individual customers who are spending personal money instead of corporate money that are more difficult to serve. You’ve been in the space for a while. What do you see around that?

Ant: Yeah, I definitely see that. So, the smaller customers, there’s a correlation between the amount of money you spend and the amount of hand-holding that someone needs. The more money someone spends, the less hand-holding they need, generally. But the other side of it, what training businesses—particularly for subscription-based business—it’s the same model as most gyms. You pay for it and you never use
it.

And it’s not just subscription; like, Udemy is a perfect example of that, you know, people who have hundreds of Udemy courses they’ve never done, but they spend ten bucks on each. So, there’s a lot of that at the lower end, which is why people offer courses at that level. So, there’s people who actually do the course who are going to give you a lot of a headache, but then you’re going to have a bunch of folk who never do the course and you’re just taking their money. Which is also not great, either, but those folks don’t feel bad because I only spent 10, 20 bucks on it. It’s like, oh, it’s their fault for not doing it, and you’ve made the money.

So, that’s kind of how a lot of the training works. So, the other problem with training as well is you get the quality is so variable at the bottom end. It’s so, so variable. You really struggle to find—there’s a lot of people just copying, like, you see instances where folks upload videos to Udemy that are literally they’ve downloaded someone’s, video resized it, cut out a logo or something like that, and re-uploaded it and it’s taken a few weeks for them to get caught. But they made money in the meantime.

That’s how blatant it does get to some level, but there are levels where people will copy someone else’s content and just basically make it their own slides, own words, that kind of thing; that happens a lot. At the low end, it’s a bit all over the place, but you still have quality, as well, at the low end, where you have these cheapest smaller courses. And how do you find that quality, as well? That’s the other side of it. And also people will just trade in their name.

That’s the other problem you see. Someone has a name for doing X whatever, and they’ll go out and bring a course on whatever that is. Doesn’t mean they’re a good teacher; it means they’re good at building a brand.

Corey: Oh, teaching is very much its own skill set.

Ant: Oh, yeah.

Corey: I learned to speak publicly by being a corporate trainer for Puppet and it teaches you an awful lot. But I had the benefit, in that case, of a team of people who spent their entire careers building curricula, so it wasn’t just me throwing together some slides; I would teach a well-structured curriculum that was built by someone who knew exactly what they’re doing. And yeah, I needed to understand failure modes, and how to get things to work when they weren’t working properly, and how to explain it in different ways for folks who learn in different ways—and that is the skill of teaching right there—but curriculum development is usually not the same thing. And when you’re bootstrapping, learning—I’m going to build my own training course, you have to do all of those things, and more. And it lends itself to, in many cases, what can come across as relatively low-quality offerings.

Ant: Yeah, completely. And it’s hard. But one thing you will often see is sometimes you’ll see a course that’s really high production quality, but actually, the content isn’t great because folks have focused on making it look good. That’s another common, common problem I see. If you’re going to do training out there, just get referrals, get references, find people who’ve done it.

Don’t believe the references you see on a website; there’s a good chance they might be fake or exaggerated. Put something out on Twitter, put out something on Reddit, whatever communities—and Slack or Discord, whatever groups you’re in, ask questions. And folks will recommend. In the world of Google where you could search for anything, [laugh], the only way to really find out if something is any good is to find out if someone else has done it first and get their opinion on it.

Corey: That’s really the right answer. And frankly, I think that is sort of the network effect that makes a lot of software work for folks. Because you don’t want to wind up being the first person on your provider trying to do a certain thing. The right answer is making sure that you are basically 8,000th person to try and do this thing so you can just Google it and there’s a bunch of results and you can borrow code on GitHub—which is how we call ‘thought leadership’ because plagiarism just doesn’t work the same way—and effectively realizing this has been solved before. If you find a brand new cloud that has no customers, you are trailblazing every time you do anything with the platform. And that’s personally never where I wanted to spend my innovation points.

Ant: We did that at Cloud Guru. I think when we were—in 2015 and we had problems with Lambda and you go to Stack Overflow, and there was no Lambda tag on Stack Overflow, no serverless tag on Stack Overflow, but you asked a question and Tim Wagner would probably be the one answering. And he was the former head of product on Lambda. But it was painful, and in general you don’t want to do it. Like [sigh] whenever AWS comes out with a new product, I’ve done it a few times, I’ll go, “I think I might want to use this thing.”

AWS Proton is a really good example. It’s like, “Hey, this looks awesome. It looks better than CodeBuild and CodePipeline,” the headlines or what I thought it would be. I basically went while the keynote was on, I logged in to our console, had a look at it, and realized it was awful. And then I started tweeting about it as well and then got a lot of feedback [laugh] on my tweets on that.

And in general, my attitude from whatever the new shiny thing is if I’m going to try it, it needs to work perfectly and it needs to live up to its billing on day one. Otherwise, I’m not going to touch it. And in general with AWS products now, you announce something, I’m not going to look at it for a year.

Corey: And it’s to their benefit that you don’t look at it for a year because the answer is going to be, ah, if you’re going to see that it’s terrible, that’s going to form your opinion and you won’t go back later when it’s actually decent and reevaluate your opinion because no one ever does. We’re all busy.

Ant: Yeah, exactly.

Corey: And there’s nothing wrong with doing that, but it is obnoxious they’re not doing themselves favors here.

Ant: Yeah, completely. And I think that’s actually a failure of marketing and communication more than anything else. I don’t blame the product
teams too much there. Don’t bill something as a finished glossy product when it’s not. Pitch it at where it is.

Say, “Hey, we are building”—like, I don’t think at the re:Invent stage they should announce anything that’s not GA and anything that it does not live up to the billing, the hype they’re going to give it to. And they’re getting more and more guilty of that the last few re:Invents, of announcing products that do not live up to the hype that they promote it at and that are not GA. Literally, they should just have a straight-up rule, they can announce products, but don’t put it on the keynote stage if it’s not GA. That’s it.

Corey: The whole re:Invent release is a whole separate series of arguments.

Ant: [laugh]. Yeah, yeah.

Corey: There are very few substantial releases throughout the year and then they drop a whole bunch of them at re:Invent, and it doesn’t matter what you’re talking about, whose problem it solves, how great it is, it gets drowned out in the flood. The only thing more foolish that I see than that is companies that are not AWS releasing things during re:Invent that are not on the re:Invent keynote stage, which in turn means that no one pays attention. The only thing you should be releasing is news about your data breach.

Ant: [laugh]. Yeah. That’s exactly it.

Corey: What do I want to bury? Whenever Adam Selipsky gets on stage and starts talking, great, then it’s time to push the button on the, “We regret to inform you,k” dance.

Ant: Yeah, exactly. Microsoft will announce yet another print spooler bug malware.

Corey: Ugh, don’t get me started on that. Thank you so much for taking the time to speak with me today. If people want to hear more about your thoughts and how you view these nonsenses, and of course to send angry emails because they are serverless fans, where can they find you?

Ant: Twitter is probably the easiest place to find me, @iamstan—

Corey: It is a place for outrage. Yes. Your Twitter user account is?

Ant: [laugh], my Twitter user account’s all over the place. It’s probably about 20% serverless. So, yeah @iamstan. Tweet me; I will probably respond to you… unless you’re rude, then I probably won’t. If you’re rude about something else, I probably will. But if you’re rude about me, I won’t.

And I expect a few DMs from Amazon after this. I’m waiting for you, [unintelligible 00:32:02], as I always do. So yeah, that’s probably the
easiest place to get hold of me. I check my email once a month. And I’m actually not joking about that; I really do check my email once a month.

Corey: Yeah, people really need me then they’ll find me. Thank you so much for taking the time to speak with me. I appreciate it.

Ant: Yes, Corey. Thank you.

Corey: Ant Stanley, AWS Serverless Hero, and oh so much more. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an angry comment defending serverless’s good name just as soon as you string together the 85 components necessary to submit that comment.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Craig

Craig McLuckie is a VP of R&D at VMware in the Modern Applications Business Unit. He joined VMware through the Heptio acquisition where he was CEO and co-founder. Heptio was a startup that supported the enterprise adoption of open source technologies like Kubernetes. He previously worked at Google where he co-founded the Kubernetes project, was responsible for the formation of CNCF, and was the original product lead for Google Compute Engine.

Links:

  • VMware: https://www.vmware.com
  • Twitter: https://twitter.com/cmcluck
  • LinkedIn: https://www.linkedin.com/in/craigmcluckie/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at the Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part my Cribl Logstream. Cirbl Logstream is an observability pipeline that lets you collect, reduce, transform, and route machine data from anywhere, to anywhere. Simple right? As a nice bonus it not only helps you improve visibility into what the hell is going on, but also helps you save money almost by accident. Kind of like not putting a whole bunch of vowels and other letters that would be easier to spell in a company name. To learn more visit: cribl.io

Corey: This episode is sponsored in part by Thinkst. This is going to take a minute to explain, so bear with me. I linked against an early version of their tool, canarytokens.org in the very early days of my newsletter, and what it does is relatively simple and straightforward. It winds up embedding credentials, files, that sort of thing in various parts of your environment, wherever you want to; it gives you fake AWS API credentials, for example. And the only thing that these things do is alert you whenever someone attempts to use those things. It’s an awesome approach. I’ve used something similar for years. Check them out. But wait, there’s more. They also have an enterprise option that you should be very much aware of canary.tools. You can take a look at this, but what it does is it provides an enterprise approach to drive these things throughout your entire environment. You can get a physical device that hangs out on your network and impersonates whatever you want to. When it gets Nmap scanned, or someone attempts to log into it, or access files on it, you get instant alerts. It’s awesome. If you don’t do something like this, you’re likely to find out that you’ve gotten breached, the hard way. Take a look at this. It’s one of those few things that I look at and say, “Wow, that is an amazing idea. I love it.” That’s canarytokens.org and canary.tools. The first one is free. The second one is enterprise-y. Take a look. I’m a big fan of this. More from them in the coming weeks.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. My guest today is Craig McLuckie, who’s a VP of R&D at VMware, specifically in their modern applications business unit. Craig, thanks for joining me. VP of R&D sounds almost like it’s what’s sponsoring a Sesame Street episode. What do you do exactly?

Craig: Hey, Corey, it’s great to be on with you. So, I’m obviously working within the VMware company, and my charter is really looking at modern applications. So, the modern application platform business unit is really grounded in the work that we’re doing to make technologies like Kubernetes and containers, and a lot of developer-centric technologies like Spring, more accessible to developers to make sure that as developers are using those technologies, they shine through on the VMware infrastructure technologies that we are working on.

Corey: Before we get into, I guess, the depths of what you’re focusing on these days, let’s look a little bit backwards into the past. Once upon a time, in the dawn of the modern cloud era—I guess we’ll call it—you were the original product lead for Google Compute Engine or GCE. How did you get there? That seems like a very strange thing to be—something that, “Well, what am I going to build? Well, that’s right; basically a VM service for a giant company that is just starting down the cloud path,” back when that was not an obvious thing for a company to do.

Craig: Yeah, I mean, it was as much luck and serendipity as anything else, if I’m going to be completely honest. I spent a lot of time working at Microsoft, building enterprise technology, and one of the things I was extremely excited about was, obviously, the emergence of cloud. I saw this as being a fascinating disrupter. And I was also highly motivated at a personal level to just make IT simpler and more accessible. I spent a fair amount of time building systems within Microsoft, and then even a very small amount of time running systems within a hedge fund.

So, I got, kind of, both of those perspectives. And I just saw this cloud thing as being an extraordinarily exciting way to drive out the cost of operations, to enable organizations to just focus on what really mattered to them which was getting those production systems deployed, getting them updated and maintained, and just having to worry a little bit less about infrastructure. And so when that opportunity arose, I jumped with both feet. Google obviously had a reputation as a company that was born in the cloud, it had a reputation of being extraordinarily strong from a technical perspective, so having a chance to bridge the gap between enterprise technology and that cloud was very exciting to me.

Corey: This was back in an era when, in my own technical evolution, I was basically tired of working with Puppet as much as I had been, and I was one of the very early developers behind SaltStack, once upon a time—which since then you folks have purchased, which shows that someone didn’t do their due diligence because something like 41 lines of code in the current release version is still assigned to me as per git-blame. So, you know, nothing is perfect. And right around then, then I started hearing about this thing that was at one point leveraging SaltStack, kind of, called Kubernetes, which, “I can’t even pronounce that, so I’m just going to ignore it. Surely, this is never going to be something that I’m going to have to hear about once this fad passes.” It turns out that the world moved on a little bit differently.

And you were also one of the co-founders of the Kubernetes project, which means that it seems like we have been passing each other in weird ways for the past decade or so. So, you’re working on GCE, and then one day you want to, what, sitting up and deciding, “I know, we’re going to build a container orchestration system because I want to have something that’s going to take me 20 minutes to explain to someone who’s never heard of these concepts before.” How did this come to be?

Craig: It’s really interesting, and a lot of it was driven by necessity, driven by a view that to make a technology like Google Compute Engine successful, we needed to go a little bit further. When you look at a technology like Google Compute Engine, we’d built something that was fabulous and Google’s infrastructure is world-class, but there’s so much more to building a successful cloud business than just having a great infrastructure technology. There’s obviously everything that goes with that in terms of being able to meet enterprises where they are and all the—

Corey: Oh, yeah. And everything at Google is designed for Google scale. It’s, “We built this thing and we can use it to stand up something that is world-scale and get 10 million customers on the first day that it launches.” And, “That’s great. I’m trying to get a Hello World page up and maybe, if I shoot for the moon, it can also run WordPress.” There’s a very different scale of problem.

Craig: It’s just a very different thing. When you look at what an organization needs to use a technology, it’s nice that you can take that, sort of, science-fiction data center and carve it up into smaller pieces and offer it as a virtual machine to someone. But you also need to look at the ISV ecosystem, the people that are building the software, making sure that it’s qualified. You need to make sure that you have the ability to engage with the enterprise customer and support them through a variety of different functions. And so, as we were looking at what it would take to really succeed, it became clear that we needed a little more; we needed to, kind of, go a little bit further.

And around that time, Docker was really coming into its full. You know, Docker solved some of the problems that organizations had always struggled with. Virtual machine is great, but it’s difficult to think about. And inside Google, containers we’re a thing.

Corey: Oh, containers have a long and storied history in different areas. From my perspective, Docker solves the problem of, “Well, it works on my machine,” because before something like Docker, the only answer was, “Well, backup your email because your laptop’s about to be in production.”

Craig: [laugh]. Yeah, that’s exactly right. You know, I think when I look at what Docker did, and it was this moment of clarity because a lot of us had been talking about this and thinking about it. I remember turning to Joe while we were building Compute Engine and basically said, “Whoever solves the packaging the way that Google did internally, and makes that accessible to the world is ultimately going to walk away with a game.” And I think Docker put lightning in a bottle.

They really just focused on making some of these technologies that underpinned the hyperscalers, that underpinned the way that, like, a Google, or a Facebook, or a Twitter tended to operate, just accessible to developers. And they solved one very specific thing which was that packaging problem. You could take a piece of software and you could now package it up and deploy it as an immutable thing. So, in some ways, back to your own origins with SaltStack and some of the technologies you’ve worked on, it really was an epoch of DevOps; let’s give developers tools so that they can code something up that renders a production system. And now with Docker, you’re able to shift that all left. So, what you produced was the actual deployable artifact, but that obviously wasn’t enough by itself.

Corey: No, there needed to be something else. And according to your biography, not only it says here that, I quote, “You were responsible for the formation of the CNCF, or Cloud Native Computing Foundation,” and I’m trying to understand is that something that you’re taking credit for or being blamed for? It really seems like it could go either way, given the very careful wording there.

Craig: [laugh]. Yeah, it could go either way. It certainly got away from us a little bit in terms of just the scope and scale of what was going on. But the whole thesis behind Kubernetes, if you just step back a little bit, was we didn’t need to own it; Google didn’t need to own it. We just needed to move the innovation boundary forwards into an area that we had some very strong advantages.

And if you look at the way that Google runs, it kind of felt like when people were working with Docker, and you had technologies like Mesos and all these other things, they were trying to put together a puzzle, and we already had the puzzle box in front of us because we saw how that technology worked. So, we didn’t need to control it, we just needed people to embrace it, and we were confident that we could run it better. But for people to embrace it, it couldn’t be seen as just a Google thing. It had to be a Google thing, and a Red Hat thing, and an Amazon thing, and a Microsoft thing, and something that was really owned by the community. So, the inspiration behind CNCF was to really put the technology forwards to build a collaborative community around it and to enable and foster this disruption.

Corey: At some point after Kubernetes was established, and it was no longer an internal Google project but something that was handed over to a foundation, something new started to become fairly clear in the larger ecosystem. And it’s sort of a microcosm of my observation that the things that startups are doing today are what enterprises are going to be doing five years from now. Every enterprise likes to imagine itself a startup; the inverse is not particularly commonly heard. You left Google to go found Heptio, where you were focusing on enterprise adoption of open-source technologies, specifically Kubernetes, but it also felt like it was more of a cultural shift in many respects, which is odd because there aren’t that many startups, at least in that era, that were focused on bringing startup technologies to the enterprise, and sneaking in—or at least that’s how it felt—the idea of culture change as well.

Craig: You know, it’s really interesting. Every enterprise has to innovate, and people tend to look at startups as being a source of innovation or a source of incubation. What we were trying to do with Heptio was to go the other way a little bit, which was, when you look at what West Coast tech companies were doing, and you look at a technology like Kubernetes—or any new technology: Kubernetes, or KNative, or there’s some of these new observability capabilities that are starting to emerge in this ecosystem—there’s this sort of trickle-across effect, where it’s starts with the West Coast tech companies that build something, and then it trickles across to a lot of the progressive forward-leaning enterprise organizations that have the scale to consume those technologies. And then over time, it becomes mainstream. And when I looked at a technology like Kubernetes, and certainly through the lens of a company like Google, there was an opportunity to step back a little bit and think about, well, Google’s really this West Coast tech company, and it’s producing this technology, and it’s working to make that more enterprise-centric, but how about going the other way?

How about meeting enterprise organizations where they are—enterprise organizations that aspire to adopt some of these practices—and build a startup that’s really about just walking the journey with customers, advocating for their needs, through the lens of these open-source communities, making these open-source technologies more accessible. And that was really the thesis around what we were doing with Heptio. And we worked very hard to do exactly as you said which is, it’s not just about the tech, it’s about how you use it, it’s about how you operate it, how you set yourself up to manage it. And that was really the core thesis around what we were pursuing there. And it worked out quite well.

Corey: Sitting here in 2021, if I were going to build something from scratch, I would almost certainly not use Kubernetes to do it. I’d probably pick a bunch of serverless primitives and go from there, but what I respect and admire about the Kubernetes approach is companies can’t generally do that with existing workloads; you have to meet them where they are, as you said. ‘Legacy’ is a condescending engineering phrase for ‘it makes money.’ It’s, “Oh, what does that piece of crap do?” “Oh, about $4 billion a year.” So yeah, we’re going to be a little delicate with what it does.

Craig: I love that observation. I always prefer the word ‘heritage’ over the word legacy. You got to—

Corey: Yeah.

Craig: —have a little respect. This is the stuff that’s running the world. This is the stuff that every transaction is flowing through.

And it’s funny, when you start looking at it, often you follow the train along and eventually you’ll find a mainframe somewhere, right? It is definitely something that we need to be a little bit more thoughtful about.

Corey: Right. And as cloud continues to eat the world well, as of the time of this recording, there is no AWS/400, so there is no direct mainframe option in most cloud providers, so there has to be a migration path; there has to be a path forward, that doesn’t include, “Oh, and by the way, take 18 months to rewrite everything that you’ve built.” And containers, particularly with an orchestration model, solve that problem in a way that serverless primitives, frankly, don’t.

Craig: I agree with you. And it’s really interesting to me as I work with enterprise organizations. I look at that modernization path as a journey. Cloud isn’t just a destination: there’s a lot of different permutations and steps that need to be taken. And every one of those has a return on investment.

If you’re an enterprise organization, you don’t modernize for modernization’s sake, you don’t embrace cloud for cloud’s sake. You have a specific outcome in mind, “Hey, I want to drive down this cost,” or, “Hey, I want to accelerate my innovation here,” “Hey, I want to be able to set my teams up to scale better this way.” And so a lot of these technologies, whether it’s Kubernetes, or even serverless is becoming increasingly important, is a capability that enables a business outcome at the end of the day. And when I think about something like Kubernetes, it really has, in a way, emerged as a Goldilocks abstraction. It’s low enough level that you can run pretty much anything, it’s high enough level that it hides away the specifics of the environment that you want to deploy it into. And ultimately, it renders up what I think is economies of scope for an organization. I don’t know if that makes sense. Like, you have these economies of scale and economies of scope.

Corey: Given how down I am on Kubernetes across the board and—at least, as it’s presented—and don’t take that personally; I’m down on most modern technologies. I’m the person that said the cloud was a passing fad, that virtualization was only going to see limited uptake, that containers were never going to eat the world. And I finally decided to skip ahead of the Kubernetes thing for a minute and now I’m actually going to be positive about serverless. Given how wrong I am on these things, that almost certainly dooms it. But great, I was down on Kubernetes for a long time because I kept seeing these enterprises and other companies talking about their Kubernetes strategy.

It always felt like Kubernetes was a means to an end, not an end in and of itself. And I want to be clear, I’m not talking about vendors here because if you are a software provider to a bunch of companies and providing Kubernetes is part and parcel of what you do, yeah, you need a Kubernetes strategy. But the blue-chip manufacturing company that is modernizing its entire IT estate, doesn’t need a Kubernetes strategy as such. Am I completely off base with that assessment?

Craig: No, I think you’re pointing at something which I feel as well. I mean, I’ll be honest, I’ve been talking about [laugh] Kubernetes since day one, and I’m kind of tired of talking about Kubernetes. It should just be something that’s there; you shouldn’t have to worry about it, you shouldn’t have to worry about operationalizing it. It’s just an infrastructure abstraction. It’s not in and of itself an end, it’s simply a means to an end, which is being able to start looking at the destination you’re deploying your software into as being more favorable for building distributed systems, not having to worry about the mechanics of what happens if a single node fails? What happens if I have to scale this thing? What happens if I have to update this thing?

So, it’s really not intended—and it never was intended—to be an end unto itself. It was really just intended to raise the waterline and provide an environment into which distributed applications can be deployed that felt entirely consistent, whether you’re building those on-premises, in the public cloud, and increasingly out to the edge.

Corey: I wound up making a tweet, couple years back, specifically in 2019, that the nuclear hot take: “Nobody will care about Kubernetes in five years.” And I stand by it, but I also think that’s been wildly misinterpreted because I am not suggesting in any way that it’s going to go away and no one is going to use it anymore. But I think it’s going to matter in the same way as the operating system is starting to, the way that the Linux virtual memory management subsystem does now. Yes, a few people in specific places absolutely care a lot about those things, but most companies don’t because they don’t have to. It’s just the way things are. It’s almost an operating system for the data center, or the cloud environment, for lack of a better term. But is that assessment accurate? And if you don’t wildly disagree with it, what do you think of the timeline?

Craig: I think the assessment is accurate. The way I always think about this is you want to present your engineers, your developers, the people that are actually taking a business problem and solving it with code, you want to deliver to them the highest possible abstraction. The less they have to worry about the infrastructure, the less they have to worry about setting up their environment, the less they have to worry about the DevOps or DevSecOps pipeline, the better off they’re going to be. And so if we as an industry do our job right, Kubernetes is just the water in which IT swims. You know, like the fish doesn’t see the water; it’s just there.

We shouldn’t be pushing the complexity of the system—because it is a fancy and complex system—directly to developers. They shouldn’t necessarily have to think like, “Oh, I need to understand all of the XYZ is about how this thing works to be able to build a system.” There will be some engineers that benefit from it, but there are going to be other engineers that don’t. The one thing that I think is going to—you know, is a potential change on what you said is, we’re going to see people starting to program Kubernetes more directly, whether they know it or not. I don’t know if that makes sense, but things like the ability for Kubernetes to offer up a way for organizations to describe the desired state of something and then using some of the patterns of Kubernetes to make the world into that shape is going to be quite pervasive, and I’m really seeing signs that we’re seeing it.

So yes, most developers are going to be working with higher abstractions. Yes, technologies like Knative and all of the work that we at VMware are doing within the ecosystem will render those higher abstractions to developers. But there’s going to be some really interesting opportunities to take what made Kubernetes great beyond just, “Hey, I can put a Docker container down on a virtual machine,” and start to think about reconciler-driven IT: being able to describe what you want to have happen in the world, and then having a really smart system that just makes the world into that shape.

Corey: This episode is sponsored by our friends at Oracle HeatWave is a new high-performance accelerator for the Oracle MySQL Database Service. Although I insist on calling it “my squirrel.” While MySQL has long been the worlds most popular open source database, shifting from transacting to analytics required way too much overhead and, ya know, work. With HeatWave you can run your OLTP and OLAP, don’t ask me to ever say those acronyms again, workloads directly from your MySQL database and eliminate the time consuming data movement and integration work, while also performing 1100X faster than Amazon Aurora, and 2.5X faster than Amazon Redshift, at a third of the cost. My thanks again to Oracle Cloud for sponsoring this ridiculous nonsense.

Corey: So, you went from driving Kubernetes adoption into the enterprise as the founder and CEO of Heptio, to effectively, acquired by one of the most enterprise-y of enterprise companies, in some respects, VMware, and your world changed. So, I understand what Heptio does because, to my mind, a big company is one that is 200 people. VMware has slightly more than that at last count, and I sort of lose track of all the threads of the different things that VMware does and how it operates. I could understand what Heptio does. What I don’t understand is what, I guess, your corner of VMware does. Modern applications means an awful lot of things to an awful lot of people. I prefer to speak it with a condescending accent when making fun of those legacy things that make money—not a popular take, but it’s there—how do you define what you do now?

Craig: So, for me, when you talk about modern application platform, you can look at it one of two ways. You can say it’s a platform for modern applications, and when people have modern applications, they have a whole variety of different ideas in the head: okay, well, it’s microservices-based, or it’s API-fronted, it’s event-driven, it’s supporting stream-based processing, blah, blah, blah, blah, blah. There’s all kinds of fun, cool, hip new patterns that are happening in the segment. The other way you could look at it is it’s a modern platform for applications of any kind. So, it’s really about how do we make sense of going from where you are today to where you need to be in the future?

How do we position the set of tools that you can use, as they make sense, as your organization evolves, as your organization changes? And so I tend to look at my role as being bringing these capabilities to our existing product line, which is, obviously, the vSphere product line, and it’s almost a hyperscale unto itself, but it’s really about that private cloud experience historically, and making those capabilities accessible in that environment. But there’s another part to this as well, which is, it’s not just about running technologies on vSphere. It’s also about how can we make a lot of different public clouds look and feel consistent without hiding the things that they are particularly great at. So, every public cloud has its own set of capabilities, its own price-performance profile, its own service ecosystem, and richness around that.

So, what can we do to make it so that as you’re thinking about your journey from taking an existing system, one of those heritage systems, and thinking through the evolution of that system to meet your business requirements, to be able to evolve quickly, to be able to go through that digital transformation journey, and package it up and deliver the right tools at the right time in the right environment, so that we can walk the journey with our customers?

Corey: Does this tie into Tanzu, or is that a different VMware initiative slash division? And my apologies on that one, just because it’s difficult for me to wrap my head around where Tanzu starts and stops. If I’m being frank.

Craig: So, [unintelligible 00:21:49] is the heart of Tanzu. So Tanzu, in a way, is a new branch, a new direction for VMware. It’s about bringing this richness of capabilities to developers running in any cloud environment. It’s an amalgamation of a lot of great technologies that people aren’t even aware of that VMware has been building, or that VMware has gained through acquisition, certainly Heptio and the ability to bring Kubernetes to an enterprise organization is part of that. But we’re also responsible for things like Spring.

Spring is a critical anchor for Java developers. If you look at the Spring community, we participate in one and a half million new application starts a month. And you wouldn’t necessarily associate VMware with that, but we’re absolutely driving critical innovation in that space. Things, like full-stack observability, being able to not only deploy these container-packaged applications, but being able to actually deal with the day two operations, and how to deal with the APM considerations, et cetera. So, Tanzu is an all-in push from VMware to bring the technologies like Kubernetes and everything that exists above Kubernetes to our customers, but also to new customers in the public cloud that are really
looking for consistency across those environments.

Corey: When I look at what you’ve been doing for the past decade or so, it really tells a story of transitions, where you went from product lead on GCE, to working on Kubernetes. You took Kubernetes from an internal Google reimagining of Borg into an open-source project that has been given over to the CNCF. You went from running Heptio, which was a startup, to working at one of the least startup-y-like companies, by some measures, in the world.s you seem to have gone from transiting from one thing to almost its exact opposite, repeatedly, throughout your career. What’s up with that theme?

Craig: I think if you look back on the transitions and those key steps, the one thing that I’ve consistently held in my head, and I think my personal motivation was really grounded in this view that IT is too hard, right? IT is just too challenging. So, the transition from Microsoft, where I was responsible building package software, to Google, which was about cloud, was really marking that transition of, “Hey, we just need to do better for the enterprise organization.” The transition from focusing on a virtual machine-based system, which was the state of the art at the time to unlocking these modern orchestrated container-based system was in service of that need, which was, “Hey, you know, if you can start to just treat a number of virtual machines as a destination that has a distributed operating system on top of it, we’re going to be better off.” The need to transition to a community-centric outcome because while Google is amazing in so many ways, being able to benefit from the perspective that traditional enterprise organizations brought to the table was significant to transitioning into a startup where we were really serving enterprise organizations and providing that interface back into the community to ultimately joining VMware because at the end of the day, there’s a lot of work to be done here.

And when you’re selling a startup, it’s—you’re either selling out or you’re buying in, and I’m not big on the idea of selling out. In this case, having access to the breadth of VMware, having access to the place where most of the customers are really cared about were living, and all of those heritage systems that are just running the world’s business. So, for me, it’s really been about walking that journey on behalf of that individual that’s just trying to make ends meet; just trying to make sure that their IT systems stay lit; that are trying to make sure that the debt that they’re creating today in the IT environment isn’t payday loan debt, it’s more like a mortgage. I can get into an environment that’s going to
serve me and my family well. And so, each of those transitions has really just been marked by need.

And I tend to look at the needs of that enterprise organization that’s walking this journey as being an anchor for me. And I’m pleased with every transition I’ve made. Like, at every point we’ve—sort of, Joe and myself, who’s been on this journey for a while, have been able to better serve that individual.

Corey: Now, I know that it’s always challenging to talk about the future, but do you think you’re done with those radical transitions, as you
continue to look forward to what’s coming? I mean, it’s impossible to predict the future, but you’re clearly where you are for a reason, and I’m assuming part of that reason is because you see an opportunity; you see a transformation that is currently unfolding. What does that look like from where you sit?

Craig: Well, I mean, my work in VMware [laugh] is very far from done. There’s just an amazing amount of continued opportunity to deliver value not only to those existing customers where they’re running on-prem but to make the public cloud more intrinsically accessible and to increasingly solve the problems as more computational resources fanning back out to the edge. So, I’m extremely excited about the opportunity ahead of us from the VMware perspective. I think we have some incredible advantages because, at the end of the day, we’re both a neutral party—you know, we’re not a hyperscaler. We’re not here to compete with the hyperscalers on the economies of scale that they render.

But we’re also working to make sure that as the hyperscalers are offering up these new services and everything else, that we can help the enterprise organization make best use of that. We can help them make best use of that infrastructure environment, we can help them navigate the complexities of things like concentration risk, or being able to manage through the luck and potential that some of these things represent. So, I don’t want to see the world collapse back into the mainframe era. I think that’s the thing that really motivates me, I think, the transition from mainframe to client-server, the work that Wintel did—the Windows-Intel consortium—to unlock that ecosystem just created massive efficiencies and massive benefits from everyone. And I do feel like with the combination of technologies like Kubernetes and everything that’s happening on top of that, and the opportunity that an organization like VMware has to be a neutral party, to really bridge the gap between enterprises and those technologies, we’re in a situation where we can create just tremendous value in the world: making it so that modernization is a journey rather than a destination, helping customers modernize at a pace that’s reasonable to them, and ultimately serving both the cloud providers in terms of bringing some critical workloads to the cloud, but also serving customers so that as they live with the harsh realities of a multi-cloud universe where I don’t know one enterprise organization that’s just all-in on one cloud, we can provide some
really useful capabilities and technologies to make them feel more consistent, more familiar, without hiding what’s great about each of them.

Corey: Craig, thank you so much for taking the time to speak with me today about where you sit, how you see the world, where you’ve been, and little bits of where we’re going. If people want to learn more, where can they find you?

Craig: Well, I’m on Twitter, @cmcluck, and obviously, on LinkedIn. And we’ll continue to invite folks to attend a lot of our events, whether that’s the Spring conferences that VMware sponsors, or VMWorld. And I’m really excited to have an opportunity to talk more about what we’re doing
and some of the great things we’re up to.

Corey: I will certainly be following up as the year continues to unfold. Thanks so much for your time. I really appreciate it.

Craig: Thank you so much for your time as well.

Corey: Craig McLuckie, Vice President of R&D at VMware in their modern applications business unit. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with a comment that I won’t bother to read before designating it legacy or heritage.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need the Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Clint

Clint is the CEO and a co-founder at Cribl, a company focused on making observability viable for any organization, giving customers visibility and control over their data while maximizing value from existing tools.

Prior to co-founding Cribl, Clint spent two decades leading product management and IT operations at technology and software companies, including Splunk and Cricket Communications. As a former practitioner, he has deep expertise in network issues, database administration, and security operations.

Links:

  • Cribl: https://cribl.io
  • Cribl sandbox: https://sandbox.cribl.io
  • Cribl.cloud: https://cribl.cloud
  • Jobs: https://cribl.io/jobs

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part my Cribl Logstream. Cirbl Logstream is an observability pipeline that lets you collect, reduce, transform, and route machine data from anywhere, to anywhere. Simple right? As a nice bonus it not only helps you improve visibility into what the hell is going on, but also helps you save money almost by accident. Kind of like not putting a whole bunch of vowels and other letters that would be easier to spell in a company name. To learn more visit: cribl.io

Corey: And now for something completely different!

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. My guest this week for this promoted episode is Clint Sharp, the CEO and co-founder of a company called Cribl. Clint, thank you for joining me, and let’s get the big question out of the way first: what is Cribl?

Clint: Yeah, so Cribl makes a stream processing engine for log and metric data. And that sounds really dry and boring, but what it really means is, we help connect, in the observability and security world, lots of log and metric sources, so you can take stuff from anywhere and put it to anywhere. And you can think of it like ETL or you can think of it like middleware; it sits there in this particular space, and it’s built for SRE and security people.

Corey: Now, I looked into this a little bit previously, and I had a sneaking suspicion when I started kicking a few of the tires on this, that there’s probably going to be an economic story of optimization and saving money because of a couple things. One, that’s what I do; I pay attention to things that save customers money in the end run, and to your company’s called Cribl—that’s C-R-I-B-L. That should probably have another L and certainly, you should buy a vowel to go in there somewhere, but that’s someone optimizing but still keeping things intact enough to be understood slash pronounceable. It really does feel like in this space, saving money on vowels is a notable tenet for companies that focus on saving money.

Clint: Yeah, so what’s interesting about enterprises is they care about money, and then they don’t care about money. And so it’s a really good way to get a meeting. We definitely do help people save a ton of money, but ultimately, I think what the value people get out of the product is helping connect all the things that they have. And so one of the biggest problems that we see in the spaces is, “Hey, I have all these agents deployed.” Maybe it’s Fluentd or Fluent Bit, or Elastic Beats or Splunk’s Forwarder.

And I want to get this data over to my fancy new data lake, or over to my machine learning and AI systems, and maybe I want to put it on a Kafka Topic, but it’s only designed to work with the thing it’s designed to work with. So, if I have Beats deployed, it works with Elastic. Okay, great. How do I also use that same data elsewhere? And really, that’s the big problem that we end up solving for our customers.

Corey: It’s the many-to-many problem. There’s a lot of work that’s implemented multiple times in multiple ways; it feels like it’s effectively you’re logging the same thing 15 different times in 15 different ways.

Clint: Well, then you look at the endpoint, and you find, “Oh, hey, we’ve got, like, eight agents rolled out here,” which is, you know, one from each vendor, they’re all collecting the same thing. And then people are like, “Oh, man, this is chewing up a ton of resources and we’re spending 20 or 30% of every box just, like, collecting security data and IT data. And couldn’t that be better?” And then oh, by the way, each one of those agents has their own security surface area, so you have to make sure that those agents themselves are secure because they’re often making outbound connections; they’re listening for inbound connections. So, we really kind of help at the edge, help people reuse existing resources.

Corey: One thing you said a few sentences ago caught me a little bit off guard and I want to dive into that a little bit. You talked about the observability and security world. Now, every time I talk to folks in one of those two spaces, they’re sort of tangentially aware of the other one exists, on some level, but they’re always framed as two very distinct universes. And you talk about them as if they’re effectively one and the same. Was that intentional?

Clint: Well, the data is the same. And it starts there because we’re collecting log data, and that log data may go into a SIEM tool, and people are using that to try to understand their security posture, and malicious actors, and threats. Oh, and by the way, that same log data is also used for understanding the performance and availability of your systems. The same type of metric data is used in both, the same type of catalogs that say, hey, what is my inventory, and what assets do I have, and where are they deployed? And all of that is relevant for both sides.

And the tooling often ends up being very similar, if not identical. And I used to work in Splunk many years ago; that’s a tool that’s well known for being popular in both camps. And so I developed this decade-long perspective of like, man, I’d show up and actually, they’re sitting right next to each other; there’s DevOps—DevSecOps now, which are now trying to marry those things. And so certainly, there’s just a ton of overlap.

Corey: It’s still all just sparkling systems administration, but people fight me on that one.

Clint: Oh, yeah. Well, yeah, so SRE is sysadmin plus, plus, plus, plus, plus.

Corey: Now, I’ve told it—what is it, it’s SRE if it’s in the Mountain View region of Silicon Valley. Otherwise, it’s just sparkling DevOps? Yep. Same story. It’s from my perspective, we called ourselves sysadmins, and then if we called ourselves DevOps, but, “I know, but DevOps isn’t a job title.”

Great, but it is a 40% raise so I’m going to be quiet about the purity of titles and take the money was my approach back then. And now there are 10 or 15 different ways you can refer to people who are more or less doing the same job and there’s no consistency between company to company in many respects. They almost become buzzwords and trite at some point, but it’s easier than trying to have a 15-minute conversation in response to, “So, what do you do at whatever company you work at?”

Clint: Well, also the grizzled sysadmin persona very much now a security person as well, right? So, you know, coming out of that sysadmin lineage, now I have to learn a whole bunch of new words, and security very much as a discipline, what I would criticize as saying, is very gatekeeper-y in terms of, “Okay, we’re going to come up with their own vernacular so that we know that you’re not one of us.” That’s one of my big criticisms of security. But the skill set, the same people who were sysadmin 20 years ago are definitely becoming security specialists, they’re becoming SREs. And so if you share the same lineage, then you’re really not all that different.

Corey: Well, that’s why I launched Last Week in AWS security newsletter podcast combo that just as just recently started launching as of the time that this airs because, “Security is everyone’s job,” but strangely, they don’t pay everyone like that. And it ties into an entire ecosystem of folks who have to care about security, but the word security doesn’t appear in their job title. And most security products seem to be pitched at the executive level where they use the same tired wording that you’ll see on airport ads everywhere, or they’re talking to InfoSec practitioners—whatever those might look at—and tying into, in some cases, a very hostile community. In other cases, they’re talking extensively about the ins and outs of how to overcome and defeat particular attack styles, or the—worst of all worlds—where it just reduces down into compliance and auditing checkboxes, which no one gets super excited about. I’m not interested in any of that.

I want to tell stories about, okay, as someone who has other work to get done, what’s the security impact of what’s happening lately? How do you round it up and distill it down into something useful, instead of something that winds up just acting as a giant distraction and becoming a budget justifier?

Clint: Well, security detection, I think, is a really fascinating area. You’re seeing a lot of consolidation now between traditional SIEM companies that—Splunk would be in there, but then you’ve got newer players like Exabeam, you got newer players like CrowdStrike who are coming from the EDR space, and they’re coming very strongly and saying, “Hey, look, I own the endpoint but really what I need to be able to do is analyze all this data.” And that’s where really these things are combining because tell me that XDR is not fundamentally the same—like, I keep using the word lineage, but the same type of product that I was building a SIEM from before. And most people I talked to are having a really hard time. Like, “What’s the difference between XDR and SIEM? Aren’t these things largely the same?”

But at the same time, then when you look at observability, it’s the same problem; I need to be able to ask and answer arbitrary questions of data. And security detection is fundamentally the same problem, I have all this data that’s being egressed from my complex systems, all my endpoints, all of my VMs, my containers, all of my infrastructure, all my applications, and I need to be able to detect when someone is doing something wrong, like, some malicious actor is doing something wrong. Tell me that’s not observability.

Corey: Of course it is. And the same problems apply to both where, if I have something happened in my application and my observability tooling doesn’t tell me for 20 minutes, that’s kind of a problem in the same way that you have that in the security space. Yet somehow, AWS’s CloudTrail takes about that, on average, to wind up surfacing various things that are happening in the environment. In many cases, the entire event can be over by the time CloudTrail says, “Hey, there’s a thing going on.” For those who aren’t familiar, CloudTrail effectively captures management events that happen talking to the AWS APIs.

So, someone creates something, someone accesses something, et cetera, et cetera. That’s useful when you need that, but if you’re going to take action based on that, you want to know sooner rather than later. Same story with any sort of monitoring tool that, “Oh, yeah, the site’s taking an outage and our system will let us know in only 20 short minutes.” Oh, I assure you customers will tell us long before then.

Clint: That’s sort of dovetails into some of the things that we see in the marketplace that we help with which are—talk about CloudTrail, people say all data is security relevant but I have to pay for all that data, too, so that data has to go somewhere. Do I care about every cloud—of course, I don’t care about every CloudTrail event; I care about some subset of those.

Corey: And honestly, in the full sweep of time, you really care about that one specific CloudTrail thing, but it’s the needle in the haystack.

Clint: And so AWS, this is a constant conflict between people who have to observe and secure systems need all the data because I may not know in advance what question I want to ask, but at the same time, I do know that not all of that is necessarily interesting right now, and so there’s a fundamental tension between, okay, the developer says, “Well, look. You can’t ask a question of data that’s not there, so I’m going to put everything in the log. Literally every byte of data, everything that I could ever think of, I’m going to put in that log.”

And then the receiver of that says like—I’ll give a good example. We’ve been talking about EDR. CrowdStrike EDR logs, phenomenal data source, have a ton of really interesting information about the security of your endpoints, and they also have an extra 100 fields that nobody gives a crap about. So, what do I do with that data? Do I pay to ingest all that data because all my vendors are charging me based off the bytes of data that are going into their platforms? And so there’s a real optimization potential there to have a really strong opinion on what good data is.

Corey: Part of the problem, too, is that you absolutely want the totality of everything captured around the specific event you care about. But by and large, we’ve all been in environments where we have a low-traffic app, and we see giant piles of web server logs. “Okay, great. Let’s take a look at what those web server logs are.” And by volume, it’s 98% load balancer health checks showing up.

It seems to me there might either be a way to strip them out entirely or alternately express those in a way that is a lot more compact and doesn’t fill things out. I still feel like there’s some terrible company somewhere where their entire way of getting signal from noise is to pay a whole bunch of interns to read the entire log by hand. I like to imagine that is me speaking hyperbolically, but I’m kind of scared it’s not.

Clint: Yeah. And then the question is, well, then how do I achieve a goal of actually getting the right data to the right place? So, that’s something that we help out about. I think that the—I feel a lot for the persona of this kind of sysadmin, this type of security person because they’re caught in this tension: like, do I go write code? My skill set as an SRE or my skill set as a security person is being an expert in the data itself.

I know that event is good, and I know that event is bad. Am I also supposed to be a person who then needs to go write a bunch of pipelines and Lambda functions, and how do I actually achieve the goal because there’s always way more demand than there is capacity to be able to onboard all of this data. So fundamentally, how do we get the right thing to the right place?

Corey: That’s, on some level, a serious problem. I will say that looking at what you do and how you do it, you take a whole bunch of different disparate data sources, and then effectively reduce all of those into passing through the Cribl log stream, and then sending the data out to exactly where it needs to go. And I have to imagine that when you talk about what you’re doing to typical VCs and whatnot, their question is, “Ah, but what if AWS launches a thing to do that?” To which I can only assume that your response must have been, “You’re right, if AWS does learn to speak coherently and effectively across all of their internal service teams, we’re going to have a serious problem.” At which point, I can only imagine that your VCs threw back their heads, you shared a happy laugh, and then they handed you another $200 million, which you have just raised. Congratulations, by the way.

Clint: Thank you so much. It’s, you know, people say a lot of times in startup-land, like, “Oh, we shouldn’t celebrate the fundraising.” I’ll tell you, as a person who’s done it a few times, I celebrate. That’s a shitload of work.

Corey: Oh, absolutely. I looked into it in the very early days of, okay, as I’m building out what would become The Duckbill Group, do I talk to VCs and the rest? And I did a little bit of investigation, and it’s, wow, that it’s so much work to build the pitch deck and have all the meetings and wind up doing all of that. I’d rather just go and sell things to customers and see how that works. And oh, that turned out to raise money that I don’t have to repay.

Okay, that seems like a different path. And there are advantages and disadvantages to every approach you can take on this. I mean, yeah, no shade here on how you decide to build out a technology company using VC-backed up resourcing, which is a sensible way to do it, but it’s a different style. And the sheer amount of work that very clearly goes into raising a fundraising round is just staggering to me. And that’s for seed-level rounds; I can only imagine down the path. This is not your first round.

Clint: Yeah, I mean, it’s a validation, I think, of where we’re going, and really, kind of, our vision because we’ve been talking a lot about how data moves, but I think one of the other key concepts that we’re advocating for that there’s a net-new concept in the industry is this concept of an observability lake. And back to that tension of there’s always way more data, S3 as an example provides excellent economics, but very few people provide a way for you to use just raw data that I end up going and dumping into S3. And that’s really the fallback for it. Like, if I don’t know what to do with this data, I don’t want to delete it because what if it becomes security relevant? Let’s talk about the SUNBURST SolarWinds attack.

Everybody in the industry wishes that they had every flow log, every log from every endpoint dating back two or three years so that they could actually go do a detailed investigation of, “Okay. That SolarWinds box got breached, and what all was it talking to?” And they can actually build a graph from that and go understand that. But most people have deleted all that data. They’ve decided that I can’t afford to have it anymore.

And so really, this concept of a lake is like, well, look, I can finally at least put it somewhere as an insurance policy and make sure that’s actually going to be relevant. And then eventually what’s going to be happening is people are going to go help you make use of that data—and we will as well—be going out there to help you take petabytes and petabytes and petabytes of logs data, metric data, trace data, observability data and give you the ability to analyze that effectively.

Corey: My constant complaint about the term ‘data lake’—because I’ve seen this happen in various client environments, AWS will release something that specifically targets data lakes, and I’ll talk to my client about that service. “This is a data lake solutions, but it would be awesome.” And they look at me like I’m very foolish and say, “Yeah, we don’t have a data lake.” To which my response is, “Great. What’s that eight petabytes of data sitting in S3?” “Oh, it’s mostly logs.”

And I don’t think that they’re foolish, I don’t think I’m foolish, but very often talking to folks who have data lakes do not recognize what they have as being a data lake because that feels almost like it’s a marketing term that has been inflicted on people. Like, they would consider it—because we all consider it this way—as more of a data morass. You’re not really sure what’s in there; you’re told by your data science teams, who are incredibly expensive, that one day we’ll unlock value in all of those web server logs, the load balancer health checks dating back to 2012, but we just don’t know what that is yet. But do you really want to risk deleting it? And it becomes this, effectively, deadstock that sits there.

So, you want to retain it, particularly if you have compliance obligations. There’s—theoretically at least—business value locked up in those things and you need to be able to access that in a reasonable way. And anytime I see tooling that winds up billing based upon amount of data stored in it, so just cut retention significantly. It feels like it cuts against the grain of what they’re trying to do.

Clint: I mean, yeah, retention, I mean, especially for security people—this is the difference between security and operations because operations is like, “Last 24 hours a data, I need. Pretty much after that, give me some aggregated statistics and I’m good.” Security people want full-fidelity data dating back years. But I think one of the other important concepts that we haven’t seen in the industry, and part of what we’re trying to change is, you know, I put data into a tool today. It’s that tool’s data, right?

So—and it doesn’t matter which tool it is that I’m put—they’re all the same. But fundamentally, I put data into a metrics or time-series database and put data into a logging tool, and that data is now owned by that vendor. And the big difference that we see in the concept of a lake is raw data at rest in S3 buckets—or other object storage depending on your cloud provider, depending on who, on-prem, is providing you that interface—in a way in which I can choose in the future, what tool is going to use to analyze that and I’m no longer locked in. And I think that’s really what we’ve been trying to advocate as an industry is that every enterprise I’ve talked to has everything. They’ve got one of every single tool and none of them are going away.

There is no such thing as a single pane of glass; that’s a myth that we’ve been talking about for 30 freaking years and it’s just never actually going to happen. And so really, what you need to be able to do is integrate things better and just make sure that people can actually use the tool that they want to use to analyze the data in the way that they see fit, and not be bound by the decision that was made six months ago as to which tool to put it in.

Corey: This episode is sponsored in part by Thinkst. This is going to take a minute to explain, so bear with me. I linked against an early version of their tool, canarytokens.org in the very early days of my newsletter, and what it does is relatively simple and straightforward. It winds up embedding credentials, files, that sort of thing in various parts of your environment, wherever you want to; it gives you fake AWS API credentials, for example. And the only thing that these things do is alert you whenever someone attempts to use those things. It’s an awesome approach. I’ve used something similar for years. Check them out. But wait, there’s more. They also have an enterprise option that you should be very much aware of canary.tools. You can take a look at this, but what it does is it provides an enterprise approach to drive these things throughout your entire environment. You can get a physical device that hangs out on your network and impersonates whatever you want to. When it gets Nmap scanned, or someone attempts to log into it, or access files on it, you get instant alerts. It’s awesome. If you don’t do something like this, you’re likely to find out that you’ve gotten breached, the hard way. Take a look at this. It’s one of those few things that I look at and say, “Wow, that is an amazing idea. I love it.” That’s canarytokens.org and canary.tools. The first one is free. The second one is enterprise-y. Take a look. I’m a big fan of this. More from them in the coming weeks.

Corey: I can tell this story—why not. I don’t imagine it was your direct fault, but nine years ago, now—so I should disclaim this. I am not even suggesting this is the way it is today. I was at a startup and we reached out to Splunk to look at handling a lot of our log analysis needs because it turned out we had a bunch of things that were spewing out logs. Nothing compared to what most sites look at these days, but back then for us, it felt like a lot of data.

And we got a quote that was more than the valuation of the company at the time. Because it seems like their biggest market headwind at the time was the rise of democracy basically making monarchies go out of fashion, and there were fewer princesses that we could kidnap for ransom in order to pay the Splunk bill. And, to their credit, they reached out every quarter and said, “Oh, have your needs change any?” “No, we have not massively inflated the value of this company so we can afford your bill. Thank you for asking.”

But the problem that I had is when I pushed back on them on this—because it’s not just one of those make fun of it and move on stories because Splunk was at the time very much the best-of-breed answer here—their response was, “Oh, just go ahead and log less and that brings your bill back into something that’s a lot more cohesive and understandable.”

Clint: Which destroys the utility of the whole tool to begin with.

Corey: Exactly. The entire reason to have a tool like that is to go through vast quantities of data and extract meaning from it. And if you’re not able to do that because you have less data, it completely defeats the value proposition of what it is you’re bringing to the table. Because in the security space, in many ways in the observability space, and certainly in my world of the cost optimization space, it’s an optimization story. It does not speed your time to market, it does not increase revenue in almost every case, so it’s always going to be a trailing function behind things that do.

Companies are structured top to bottom in order to increase revenue and enter new markets with the right offerings at the right times and serve customers because that can massively increase the value of the company. Reduction and, I guess, the housekeeping stuff is things people get really excited about for short windows of time and then not again. It’s inconsistent.

Clint: Yeah, about every time the bill comes due is when they get really excited about it.

Corey: Exactly. And I have to assume on some level, this was one of those, “Okay, first start using it. You’ll see how valuable it becomes, and then you’ll start logging more data.” But it didn’t feel right because it’s either being disingenuous, or it’s saying that, “Oh, don’t worry. You’ll find the money somehow.”

Which is not true in that scenario. Now, they’ve redone their pricing multiple times since then. There are other entrants in the market that help us look at data in a bunch of different ways, but across the board, it’s frustrating seeing that there are all these neat tools that I wanted to use and I was perfectly positioned to use back then, and now nine years later, when someone says, “Oh, we use Splunk.” My immediate instinctive reaction is, “Oh, wow. You must have a lot of money to spend on services.” Which is not necessarily even close to reality in some cases, but first impressions like that really stick around a long time.

Clint: Oh, absolutely. They stick around often because they’re reaffirmed multiple times throughout [laugh] people’s continued interactions. And I think there’s just really a fundamental tension in the marketplace where the value proposition is massive amounts of data. And massive is different, depending on the size of your organization: if you’re a big Fortune 100, massive might be, you know, a 100 petabytes at rest and a petabyte a day of data moving; or for you, massive might be a terabyte a day moving, and maybe a 50 terabytes at rest. But—and by the way, that’s not going down.

So, some of the bigger trends that we’re seeing with the advent of zero trust, with the advent of remote work, with just in general growth of cloud containerized workloads, microservices, people are seeing a lot more data today than they were seeing two years ago, three years ago. And by the way, it’s not like IT went from 2% of the budget to 10% of the budget. The budget’s the same, so I got to do more with less. And it’s a tension between data growth and cost and capacity. And so we got to get smarter.

Corey: I like the fact that you’re saying that you have to get smarter as you think about this from a tool perspective of being able to serve your customers, as opposed to a lot of tooling out there seems to inherently and intrinsically take the world view—and I don’t know if this is an actual choice or just an unfortunate side effect—of, “Yeah, we have to educate our customers because right now, our customers are fairly dumb and we’d like it if they were smarter. If you were smart enough to appreciate how we do things, then things will go super well.” And I always found that to be a condescending attitude that doesn’t serve customers super well. And it also leaves a lot of money on the table because for better or worse, you have to meet customers where they are: at their level of understanding, at their expression of the problem. And I’ve talked to a number of folks over at Cribl and, similar to certain large cloud providers, one of the things that you focus on is the customer; it’s clearly a value of the company. How do you think about that?

Clint: I’m a thousand percent agree with you. And for us, what I found after having been a practitioner for a decade and then working my way over to the vendor side, it’s really nothing specific about one particular employer. Being a vendor is so complex. There’s all these things that you’re trying to con—you have investors, and you have the press, and analysts, and you have people who are constantly trying to influence where it is that you’re—“I need to be in the upper right of the Gartner Magic Quadrant, so I have to make sure that those analysts really believe what it is that I’m saying.” And then pretty soon, just nobody even talks about the customer anymore.

It’s like, well, do people actually want to buy it? Is this thing actually solving real problems? And so from the beginning, me and my co-founders, we just wanted to make sure that the concept of the customer was embedded at the core of the company. And every time that an employee at Cribl is interacting and talking about what should we do next, and what features should we build, and how should we market, and how should we sell, let’s make sure the customer is there. Customers first always is the value, including in how we sell.

We actively leave money on the table when it’s not in the customer’s right interest because we know that we want them to come back and buy from us again, later. When we market, we try to make sure that we’re speaking to our customers in a language that is their language. When we’re building a product, we use the product, we try to make sure that this is actually everyday, we don’t look at, hey, it needs to look like this and have these features to meet these criteria and be called this. It’s just like, “Well, does it actually help the customer solve a real problem for them? If so, let's build it. And if not, then who gives a [BLEEP]?”

Corey: Exactly. It’s understanding what your customers’ pain points are. I mean, I ran into some similar problems when I was starting my consultancy where I—it turns out that I knew people who were more or less top of their class when it came to AWS bill understanding, reconciliation, and the rest. And those are the people I reached out to because I assumed that they knew what they’re doing. There must be lots of people like them, everyone must be like these folks.

And I talked to them about how they looked at their AWS bill. And, okay, “They said I would—I’d love to hire you to come in and do this as a consultant, but I would expect this, this, this, this, and this.” And, “Okay, I better come loaded for bear.” And so I did. And it turns out there’s a lot more people out there who have never heard of a savings plan or a reserved instance before or, “Wait. You mean continues to charge me even after I’m still using it if I don’t turn it off?” Yes, that is generally how it works.

There’s nothing wrong with that level of understanding of these things—well, there are several things wrong but that’s beside the point—but understanding where folks are and understanding how you can meet them where they are and get them to a better place is way more important than trying to prove that I’m the smartest kid in town when it comes to a lot of the edge case, and corner cases, and nuanced areas. And so many tools seem to have fallen in love with their own tooling, and in love with how smart they are, and how clear their lines of thought leadership are, that they’ve almost completely forgotten that there are people in the world who do not think like that, who do not have the level of visibility or deep thought into the problem space; they just know that the logs are unmanageable, or the bill for this thing is really expensive, or whatever their expression or experience of that problem is, there are tools out there that can help them, but all of the messaging, all of the marketing distills down to, “Oh, you must be at least this smart to enter,” like it’s an amusement park ride with a weird sign.

Clint: Software is fundamentally a people business and when you end up implementing a tool—what’s become fascinating to me as I’ve become the CEO of this company, rather than just kind of a product guy, so now I’ve had to sell it and I’ve had to market it, and I had to start very much from scratch, is that this stuff doesn’t just get implemented by magic; even if they download the tool and is the easiest to use tool that you’ve ever used, they still don’t have the time to learn all the details and intricacies of your product, and so hey, they actually want some professional services people to come and go install that; they want a salesperson to help them understand the value. I know a lot of people, especially coming from my background in, like, SRE or sysadmin from when I was doing it, kind of, “Oh, salespeople.” But, like, they do a real job; they help you articulate the value of this thing so that your bosses understand what you’re actually buying. The sales engineers help you understand what those features are. And so having a customer-aligned company means that every interaction that they have with you needs to be a really, really great interaction so that they want to interact with you again because fundamentally, even though the bits are really awesome and they solve this really awesome technology challenge, nobody really cares about it.

Ultimately, they’re buying from people, they’re implementing software built by people, and they’re calling for support—which is another important part—from people who fundamentally care about them as well. So, in every interaction, fundamentally software is a people business, and you got to have the best people and the people that care.

Corey: I wish more people took that philosophy because, frankly, it’s missing from an awful lot of different expressions of what companies do. It’s oh, if we can make the code just a little bit smarter, a little bit more predictive, then we never have to talk to the customer at all. It’s, “No. You shouldn’t write a line of anything before doing a whole bunch of customer research to validate that your understanding of the problem space aligns with theirs.”

Clint: A good way to find out that doesn’t work is to fail for a while, too. So, [laugh] so we did our fair share of that, too, and kind of pontificating and trying to figure out what we thought was best at the market, and it turned out that really what you needed to be able to do was to work closely with customers and understand their problems and tightly pair that sales cycle, that marketing messaging, that product all towards customer pain. And if you do that, customers are great because they see the people who care, and they will reward you by becoming your customer and continuing to advocate for you and talk about you. And it’s so rewarding if you can take the right perspective.

Corey: So, we’ve covered a fair number of things: your philosophy on the world of security versus observability; we’ve talked about meeting customers where they are; we’ve talked about AWS being so inept at communicating internally and cross-functionally that you’re able to raise staggeringly large rounds, and we’ve talked about, I guess, how we wind up viewing the world of log collection, for lack of a better term. If people want to learn more about what you’re up to, and how you get there, where can they find you?

Clint: Yeah, go to cribl.io. If you’re a hands-on product person and you just want to see what we do, you can go to sandbox.cribl.io. And there’s an online learning course, takes about an hour, walk you through the product. We'd love for you to try it.

Corey: Oh, I don’t have to speak to a salesperson?

Clint: No, you don’t have to talk to anybody. You can download the bits, you can try our cloud product for free at cribl.cloud. We are all about making sure that engineers can get access to the product before you have to talk to us. And if you think that’s valuable, if this helps you solve a problem, then and only then should you engage with us and we’ll see if we can figure out a way to sell you some software.

Corey: Customer-focused. I’m also going to take a spot check here. I’m going to guess that given your recent funding news, you’re also aggressively hiring.

Clint: We are hiring across every function, and if you are interested in working for our customers-first software company and this sounds refreshing, please check out cribl.io/jobs, and we’ve hiring everywhere.

Corey: I can endorse. We used to hang out, back before you wound up starting this place, and you were kicking around this idea of, “I have an idea for a company,” and my general perception is, “Eh, I don’t know. Doesn’t sound like it has legs to me.” And well, here we are. I sure can pick them. Badly. Clint, thank you so much for taking the time to speak with me.

Clint: Thanks, Corey. It’s been a pleasure.

Corey: Clint Sharp, CEO and co-founder of Cribl. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you hated this podcast, please leave a five-star review on your podcast platform of choice along with an insulting comment telling me exactly why I’m wrong about the phrase ‘data lake’ and tell me how many petabytes of useless material you have sitting in S3.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Rich

Rich Burroughs is a Senior Developer Advocate at Loft Labs where he's focused on improving workflows for developers and platform engineers using Kubernetes. He's the creator and host of the Kube Cuddle podcast where he interviews members of the Kubernetes community. He is one of the founding organizers of DevOpsDays Portland, and he's helped organize other community events. Rich has a strong interest in how working in tech impacts mental health. He has ADHD and has documented his journey on Twitter since being diagnosed.

Links:

  • Loft Labs: https://loft.sh
  • Kube Cuddle Podcast: https://kubecuddle.transistor.fm
  • LinkedIn: https://www.linkedin.com/in/richburroughs/
  • Twitter: https://twitter.com/richburroughs
  • Polywork: https://www.polywork.com/richburroughs

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part my Cribl Logstream. Cirbl Logstream is an observability pipeline that lets you collect, reduce, transform, and route machine data from anywhere, to anywhere. Simple right? As a nice bonus it not only helps you improve visibility into what the hell is going on, but also helps you save money almost by accident. Kind of like not putting a whole bunch of vowels and other letters that would be easier to spell in a company name. To learn more visit: cribl.io

Corey: This episode is sponsored in part by Thinkst. This is going to take a minute to explain, so bear with me. I linked against an early version of their tool, canarytokens.org in the very early days of my newsletter, and what it does is relatively simple and straightforward. It winds up embedding credentials, files, that sort of thing in various parts of your environment, wherever you want to; it gives you fake AWS API credentials, for example. And the only thing that these things do is alert you whenever someone attempts to use those things. It’s an awesome approach. I’ve used something similar for years. Check them out. But wait, there’s more. They also have an enterprise option that you should be very much aware of canary.tools. You can take a look at this, but what it does is it provides an enterprise approach to drive these things throughout your entire environment. You can get a physical device that hangs out on your network and impersonates whatever you want to. When it gets Nmap scanned, or someone attempts to log into it, or access files on it, you get instant alerts. It’s awesome. If you don’t do something like this, you’re likely to find out that you’ve gotten breached, the hard way. Take a look at this. It’s one of those few things that I look at and say, “Wow, that is an amazing idea. I love it.” That’s canarytokens.org and canary.tools. The first one is free. The second one is enterprise-y. Take a look. I’m a big fan of this. More from them in the coming weeks.sca

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Periodically, I like to have, well, let’s call it fun, at the expense of developer advocates; the developer relations folks; DevRelopers as I insist on pronouncing it. But it’s been a while since I’ve had one of those come on the show and talk about things that are happening in that universe. So, today we’re going back to change that a bit. My guest today is Rich Burroughs, who’s a Senior Developer Advocate—read as Senior DevReloper—at Loft Labs. Rich, thanks for joining me.

Rich: Hey, Corey. Thanks for having me on.

Corey: So, you’ve done a lot of interesting things in the space. I think we first met back when you were at Sensu, you did a stint over at Gremlin, and now you’re over at Loft. Sensu was monitoring things, Gremlin was about chaos engineering and breaking things on purpose, and when you’re monitoring things that are breaking that, of course, leads us to Kubernetes, which is what Loft does. I’m assuming. That’s probably not your marketing copy, though, so what is it you folks do?

Rich: I was waiting for your Kubernetes trash talk. I knew that was coming.

Corey: Yeah. Oh, good. I was hoping I could sort of sneak it around in there.

Rich: [laugh].

Corey: But yeah, you know me too well.

Rich: By the way, I’m not dogmatic about tools, right? I think Kubernetes is great for some things and for some use cases, but it’s not the best tool for everything. But what we do is we really focus a lot on the experience of developers who are writing applications that run in Kubernetes cluster, and also on the platform teams that are having to maintain the clusters. So, we really are trying to address the speed bumps, the things that people bang their shins on when they’re trying to get their app running in Kubernetes.

Corey: Part of the problem I’ve always found is that the thing that people bang their shins on is Kubernetes. And it’s one of those, “Well, it’s sort of in the title, so you can’t really avoid it. The only way out is through.” You could also say, “It’s better never begin; once begun, better finish.” The same thing seems to apply to technology in a whole bunch of different ways.

And that’s always been a strange thing for me where I would have bet against Kubernetes. In fact, I did, and—because it was incredibly complicated, and it came out of Google, not that someone needed to tell me. It was very clearly a Google-esque product. And we saw it sort of take the world by storm, and we are all senior YAML engineers now. And here we are.

And now you’re doing developer advocacy, which means you’re at least avoiding the problem of actually working with Kubernetes day-in-day out yourself, presumably. But instead, you’re storytelling about it.

Rich: You know, I spent a good part of my day a couple days ago fighting with my Kubernetes cluster at Docker Desktop. So, I still feel the pain some, but it’s a different kind of pain. I’ve not maintaining it in production. I actually had the total opposite experience to you. So, my introduction to Kubernetes was seeing Kelsey Hightower talk about it in, like, 2015.

And I was just hooked. And the reason that I was hooked is because of what Kubernetes did, and I think especially the service primitive, is that it encoded a lot of these operational patterns that had developed into the actual platform. So, things like how you check if an app is healthy, if it’s ready to start accepting requests. These are things that I was doing in the shops that I was working at already, but we had to roll it ourselves; we had to invent a way to do that. But when Kelsey started talking about Kubernetes, it became apparent to me that the people who designed this thing had a lot of experience running applications in distributed systems, and they understood what you needed to be able to do that competently.

Corey: There’s something to be said for packaging and shipping expertise, and it does feel like we’re on a bit of a cusp, where the complexity has risen and risen and risen, and it’s always a sawtooth graph where things get so complicated that you then are paying people a quarter-million dollars a year to run the thing. And then it collapses in on itself. And the complexity is still there, but it’s submerged to a point where you don’t need to worry about it anymore. And it feels like we’re a couple years away from Kubernetes hitting that, but I do view that as inevitable. Is that, basically, completely out to sea? Is that something that you think is directionally correct, or something else?

Rich: I mean, I think that the thing that’s been there for a long time is, how do we take this platform and make it actually usable for people? And that’s a lot more about the whole CNCF ecosystem than Kubernetes itself. How do we make it so that we can easily monitor this thing, that we can have observability, that we can deploy applications to it? And I think what we’ve seen over the last few years is that, even more than Kubernetes itself, the tools that allow you to do those other things that you need to do to be able to run applications have exploded and gotten a lot better, I think.

Corey: The problem, of course, is the explosion part of it because we look at the other side, at the CNCF landscape diagram, and it is a hilariously overwrought picture of all of the different offerings and products and tools in the space. There are something like 400 blocks on it, the last time I checked. It looks like someone’s idea of a joke. I mean, I come up with various shitposts that I’m sort of embarrassed I didn’t come up with one anywhere near that funny.

Rich: I left SRE a few years ago, and this actually is one of the reasons. So, the explosion in tools gave me a huge amount of imposter syndrome. And I imagine I’m not the only one because you’re on Twitter, you’re hanging around, you’re seeing people talk about all these cool tools that are out there, and you don’t necessarily have a chance to play with them, let alone use them in production. And so what I would find myself doing is I would compare myself to these people who were experts on these tools. Somebody who actually invented the thing, like Joe Beda or something like that, and it’s obviously unfair to do because I’m not that person. But my brain just wants to do that. You see people out
there that know more than you and a lot of times I would feel bad about it. And it’s an issue, I really think it is.

Corey: So, one of the problems that I ran into when I left SRE was that I was solving the same problem again and again, in rapid succession. I was generally one of the first early SRE-type hires, and, “Oh, everything’s on fire, and I know how to fix those things. We’re going to migrate out of EC2 Classic into VPCs; we’re going to set up infrastructure as code so we’re not hand-building these things from scratch every time.” And in time, we wind up getting to a point where it’s, okay, there are backups, and it’s easy to provision stuff, and things mostly work. And then it becomes tedium, where the day starts to look too much alike.

And I start looking for other problems elsewhere in the organization, and it turns out that when you don’t have strategic visibility into what other orgs are doing but tell them what they’re doing wrong, you’re not a popular person; and you’re often wrong. And that was a source of some angst in my case. The reason I started what I do now is because I was looking to do something different where no two days look alike, and I sort of found that. Do you find that with respect to developer advocacy, or does it fall into some repetitive pattern? Not there’s anything wrong with that; I wish I had the capability to do that, personally.

Rich: So, it’s interesting that you mentioned this because I’ve talked pretty publicly about the fact that I’ve been diagnosed with ADHD a few months ago. You talked about the fact that you have it as well. I loved your Twitter thread about it, by the way; I still recommend it to people. But I think the real issue for me was that as I got more advanced in my career, people assumed that because you have ‘senior’ in your title, that you’re a good project manager. It’s just assumed that as you grow technically and move into more senior roles, that you’re going to own projects. And I was just never good at that. I was always very good at reactive things, I think I was good at being on call, I think I was good at responding to incidents.

Corey: Firefighting is great for someone with our particular predilections. It’s, “Oh, great. There’s a puzzle to solve. It’s extremely critical that we solve it.” And it gets the adrenaline moving. It’s great, “Cool, now fill out a bunch of Jira tickets.” And those things will sit there unfulfilled until the day I die.

Rich: Absolutely. And it’s still not a problem that I’ve solved. I’ll preface this with the kids don’t try this at home advice because everybody’s situation is different. I’m a white guy in the industry with a lot of privilege; I’ve developed a really good network over the years; I don’t spend a lot of time worried about what happens if I lose my job, right, or how am I going to get another one. But when I got this latest job that I’m at now, I was pretty open with the CEO who interviewed me—it’s a very small company, I’m like employee number four.

And so when we talked to him ahead of time, I was very clear with him about the fact that bored Rich is bad. If Rich gets bored with what he’s doing, if he’s not engaged, it’s not going to be good for anyone involved. And so—

Corey: He’s going to go find problems to solve, and they very well may not align with the problems that you need solved.

Rich: Yeah, I think my problem is more that I disengage. Like, I lose my passion for what it is that I’m doing. And so I’ve been pretty intentional about trying to kind of change it up, make different kinds of content. I happen to be at this place that has four open-source projects, right, along with our commercial project. And so, so far at least, there’s been plenty for me to talk about. I haven’t had to worry about being bored so
far.

Corey: Small companies are great for that because you’re everyone does everything to some extent; you start spreading out. And the larger a company gets, the smaller your remit is. The argument I always made against working at Google, for example was, let’s say that I went in with evil in mind on day one. I would not be able—regardless of how long I was there, how high in the hierarchy I climbed—to take down google.com for one hour—the search engine piece.

If I can’t have that much impact intentionally, then the question really becomes how much impact can I have in a positive direction with everyone supposedly working in concert with me? And the answer I always came up with was not that much, not in the context of a company like that. It’s hard for me to feel relevant to a large company. For better or worse, that is the thing that keeps me engaged is, “You know, if I get this wrong enough, we don’t have a company anymore,” is sort of the right spot for me.

Rich: [laugh]. Yeah, I mean, it’s interesting because I had been at a number of startups last few years that were fairly early stage, and when I was looking for work this last time, my impulse was to go the opposite direction, was to go to a big company, you know, something that was going to be a little more stable, maybe. But I just was so interested in what these folks were building. And I really clicked with Lukas, the CEO, when we talked, and I ended up deciding to go this route. But there’s a flip side to that.

There’s a lot of responsibility that comes with that, too. Part of me wanting to avoid being in that spotlight, in a way; part of me wanted to back off and be one of the million people building things. But I’m happy that I made this choice, and like I said, it’s been working out really well, so far.

Corey: It seems to be. You seem happy, which is always a nice thing to be able to pick up from someone in how they go about these things. Talk to me a little bit about what Loft does. You’re working on some virtual cluster nonsense that mostly sails past me. Can you explain it using small words?

Rich: [laugh]. Yeah, sure. So, if you talk to people who use Kubernetes, a lot, you are—

Corey: They seem sad all the time. But please continue.

Rich: One of the reasons that they’re sad is because of multi-tenancy in Kubernetes; it just wasn’t designed with that sort of model in mind. And so what you end up with is a couple of different things that happen. Either people build these shared clusters and feel a whole lot of pain trying to share them because people commonly use namespaces to isolate things, and that model doesn’t completely work. Because there are objects like CRDs and things that are global, that don’t live in the namespace, and so that can cause pain. Or the other option that people go with is that they just spin up a whole bunch of clusters.

So, every team or every developer gets their own cluster, and then you’ve got all this cluster sprawl, and you’ve got costs, and it’s not great for the environment. And so what we are really focused a lot on with the virtual cluster stuff is it provides people what looks like a full-blown Kubernetes cluster, but it just lives inside the namespace on your host cluster. So, it actually uses K3s, from the Rancher folks, the SUSE folks. And literally, this K3s API server sits in the namespace. And as a user, it looks to you like a full-blown Kubernetes cluster.

Corey: Got it. So, basically a lightweight [unintelligible 00:13:31] that winds up stripping out some of the overwrought complexity. Do you find that it winds up then becoming a less high-fidelity copy of production?

Rich: Sure. It’s not one-to-one, but nothing ever is, right?

Corey: Right. It’s a question of whether people admit it or not, and where they’re willing to make those trade-offs.

Rich: Right. And it’s a lot closer to production than using Docker Compose or something like that. So yeah, like you said, it’s all about trade-offs, and I think that everything that we do as technical people is about trade-offs. You can give everybody their own Kubernetes cluster, you know, would run it in GK or AWS, and there’s going to be a cost associated with that, not just financially, but in terms of the headaches for the people administering things.

Corey: The hard part from where I’ve always been sitting has just been—because again, I deal with large-scale build-outs; I come in in the aftermath of these things—and people look at the large Kubernetes environments that they’ve built and it’s expensive, and you look at it from the cloud provider perspective, and it’s just a bunch of one big noisy application that doesn’t make much sense from the outside because it’s obviously not a single application. And it’s chatty across availability zone boundaries, so it costs two cents per gigabyte. It has no [affinity 00:14:42] for what’s nearby, so instead of talking to the service that is three racks away, it talks the thing over an expensive link. And that has historically been a problem. And there are some projects being made in that direction, but it’s mostly been a collective hand-waving around it.

And then you start digging into it in other directions from an economics perspective, and they’re at large scale in the extreme corner cases, it always becomes this, “Oh, it’s more trouble than it’s worth.” But that is probably unfair for an awful lot of the real-world use cases that don’t rise to my level of attention.

Rich: Yeah. And I mean, like I said earlier, I think that it’s not the best use case for everything. I’m a big fan of the HashiCorp tools. I think Nomad is awesome. A lot of people use it, they use it for other things.

I think that one of the big use cases for Nomad is, like, running batch jobs that need to be scheduled. And there are people who use Nomad and Kubernetes both. Or you might use something like Cloud Run or AppRun, whatever works for you. But like I said, from someone who spent literally decades figuring out how to operate software and operating it, I feel like the great thing about this platform is the fact that it does sort of encode those practices.

I actually have a podcast of my own. It’s called Kube Cuddle. I talk to people in the Kubernetes community. I had Kelsey Hightower on recently, and the thing that Kelsey will tell you, and I agree with him completely, is that, you know, we talk about the complexity in Kubernetes, but all of that complexity, or a lot of it, was there already.

We just dealt with it in other ways. So, in the old days, I was the Kubernetes scheduler. I was the guy who knew which app ran on which host, and deployed them and did all that stuff. And that’s just not scalable. It just doesn’t work.

Corey: This episode is sponsored by our friends at Oracle Cloud. Counting the pennies, but still dreaming of deploying apps instead of "Hello, World" demos? Allow me to introduce you to Oracle's Always Free tier. It provides over 20 free services and infrastructure, networking databases, observability, management, and security.

And - let me be clear here - it's actually free. There's no surprise billing until you intentionally and proactively upgrade your account. This means you can provision a virtual machine instance or spin up an autonomous database that manages itself all while gaining the networking load, balancing and storage resources that somehow never quite make it into most free tiers needed to support the application that you want to build.

With Always Free you can do things like run small scale applications, or do proof of concept testing without spending a dime. You know that I always like to put asterisks next to the word free. This is actually free. No asterisk. Start now. Visit https://snark.cloud/oci-free that's https://snark.cloud/oci-free.

Corey: The hardest part has always been the people aspect of things, and I think folks have tried to fix this through a lens of, “The technology will solve the problem, and that’s what we’re going to throw at it, and see what happens by just adding a little bit more code.” But increasingly, it doesn’t work. It works for certain problems, but not for others. I mean, take a look at the Amazon approach, where every team communicates via APIs—there’s no shared data stores or anything like that—and their entire problem is a lack of internal communication. That’s why the launch services that do basically the same thing as each other because no one bothers to communicate with one another. And half my job now is introducing Amazonians to one another. It empowers some amazing things, but it has some serious trade-offs. And this goes back to our ADHD aspect of the conversation.

Rich: Yeah.

Corey: The thing that makes you amazing is also the thing that makes you suck. And I think that manifests in a bunch of different ways. I mean, the fact that I can switch between a whole bunch of different topics and keep them all in state in my head is helpful, but it also makes me terrible, as far as an awful lot of different jobs, where don’t come back to finish things like completing the Jira ticket to hit on Jira a second time in the same recording.

Rich: Yeah, I’m the same way, and I think that you’re spot on. I think that we always have to keep the people in mind. You know, when I made this decision to come to Loft Labs, I was looking at the tools and the tools were cool, but it wasn’t just that. It’s that they were addressing problems that people I know have. You hear these stories all the time about people struggling with the multi-tenancy stuff and I could see very quickly that the people building the tools were thinking about the people using them, and I think that’s super important.

Corey: As I check your LinkedIn profile, turns out, no, we met back in your Puppet days, the same era that I was a traveling trainer, teaching people how to Puppet and hoping not to get myself ejected from the premises for using sarcastic jokes about the company that I was conducting the training for. And that was fun. And then I worked at a bunch of places, you worked in a bunch of places, and you mentioned a few minutes ago that we share this privilege where if one of us loses our job, the next one is going to be a difficult thing for us to find, given the skill set that we have, the immense privilege that we enjoy, and the way that this entire industry works. Now, I will say that has changed somewhat since starting my own company. It’s no longer the fear of, “Well, I’m going to land on my feet.” Rich: Right.

Corey: Yeah, but I’ve got a bunch of people who are counting on me not to completely pooch this up. So, that’s the thing that keeps me awake at night, now. But I’m curious, do you feel like that’s given you the flexibility to explore a bunch of different company types and a bunch of different roles and stretch yourself a little with the understanding that, yeah, okay. If you’ve never last five years at the same company, that’s not an inherent problem.

Rich: Yeah, it’s interesting. I’ve had conversations with people about this. If you do look up my LinkedIn, you’re going to see that a lot of the recent jobs have been less than two years: year, year and a half, things like that. And I think that I do have some of that freedom, now. Those exits haven’t always been by choice, right?

And that’s part of what happens in the industry, too. I think I’ve been laid off, like, four or five times now in my career. The worst one by far was when the bubble burst back in 2000. I was working at WebMD, and they ended up closing our office here.

Corey: You were Doctor Google.

Rich: I kind of was. So, I was actually the guy who would deploy the webmd.com site back then. And it was three big Sun servers. And I would manually go in and run shell scripts and take one out of the load balancer and roll the new code on it, and then move on to the next one. And those are early days; I started in the industry in about ’95. Those early days, I just felt bulletproof because everybody needed somebody with my skills. And after that layoff in 2000, it was a lot different. The market just dried up, I went 10 months unemployed. I ended up taking a job where I took a really big pay cut in a role that wasn’t really good for me, career-wise. And I guess it’s been a little bit of a comfort to me, looking back. If I get laid off now, I know it’s not going to be as bad as that was. But I think that’s important, and one of the things that’s helped me a lot and I’m assuming it’s helped you, too, is building up a network, meeting people, making friends. I sort of hate the word networking because it has really negative connotations to it to me. The salespeople slapping each other on the back at the bar and exchanging business cards is the image that comes to my mind when I think of networking. But I don’t think it has to be like that. I think that you can make genuine friendships with people in the industry that share the interests and passions that you have.

Corey: That’s part of it. People, I think, also have the wrong idea about passion and how that interplays with career. “Do a thing that you love, and the money will follow,” is terrific advice in the United States to make about $30,000 a year. I assure you, when I started this place, I was not deeply passionate about AWS billing. I developed a passion for it as I rapidly had to become an expert in this thing.

I knew there was an expensive business problem there that leveraged the skill set that I already had and I could apply it to something that was valuable to more than just engineers because let’s face it, engineers are generally terrible customers for a variety of reasons. And by doing that and becoming the expert in that space, I developed a passion for it. I think it was Scott Galloway who in one of his talks said he had a friend who was a tax attorney. And do you think that he started off passionate about tax law? Of course not.

He was interested in making a giant pile of money. Like, his preferred seat when he flies is ‘private.’ So, he’s obviously very passionate about it now, but he found something that he could enjoy that would pay a bunch of money because it was an in-demand, expensive skill. I often wonder if instead of messing around and computers, my passion had been oil painting, for example. Would I have had anything approaching to the standard of living I have now?

The answer is, “Of course not.” It would have been a very different story. And that says many deeply troubling things about our society across the board. I don’t know how to fix any of them. I’m one of those people that rather than sitting here talking how the world should be; I deal with the world as I encounter it.

And at times, that doesn’t feel great, but it is the way that I’ve learned to cope, I guess, with the existential angst. I’m envious in some ways of the folks who sit here saying, “No, we demand a better world.” I wish I shared their optimism or ability to envision it being different than it is, but I just don’t have it.

Rich: Yeah, I mean, there are oil painters who make a whole lot of money, but it’s not many of them, right?

Corey: Yeah, but you shouldn’t have to be dead first.

Rich: [laugh]. I used to… know a painter who Jim Carrey bought one of his big canvases for quite a lot of money. So, they’re not all dead. But again, your point is very valid. We are in this bubble in the tech industry where people do make on average, I think, a lot more money than people do in many other kinds of jobs.

And I recently started thinking about possibly going into ADHD coaching. So, I have an ADHD coach myself; she has made a very big difference in my life so far. And I actually have started taking classes to prepare for possibly getting certified in that. And I’m not sure that I’m going to do it. I may stay in tech.

I may do some of both. It doesn’t have to be either-or. But it’s been really liberating to just have this vision of myself working outside of tech. That’s something that I didn’t consider was even possible for quite a long time.

Corey: I have to confess I’ve never had an ADHD coach. I was diagnosed when I was five years old and back then—my mother had it as well, and the way that it was dealt with in the ’50s and ’60s when she was growing up was, she had a teacher once physically tie her to a chair. Which—

Rich: Oh, my gosh.

Corey: —is generally frowned upon these days. And coaching was never a thing. They decided, “Oh, we’re going to medicate you to the gills,” in my case. And that was great. I was basically a zombie for a lot of my childhood.

When I was 17, I took myself off of it and figured I’d white-knuckle it for the next 10 years or so. Again, everyone’s experience is different, but for me, didn’t work, and it led to some really interesting tumultuous times in my ’20s. I’ve never explored coaching just because it feels like so much of what I do is the weirdest aspects of different areas of ADHD. I also have constraints upon me that most folks with ADHD wouldn’t have. And conversely, I have a tremendous latitude in other areas.

For example, I keep dropping things periodically from time to time; I have an assistant. Turns out that most people, they bring in an assistant to help them with stuff will find themselves fired because you’re not supposed to share inside company data with someone who is not an employee of that company. But when you own the company, as I do, it’s well, okay, I’m not supposed to share confidential client data or give access to it to someone who’s not an employee here. “Da da da da da. Welcome aboard. Your first day is Monday.”

And now I’ve solved that problem in a way that is not open to most people. That is a tremendous benefit and I’m constantly aware of how much privilege is just baked into that. It’s a hard thing for me to reconcile, so I’ve never explored the coaching angle. I also, on some level—and this is an area that I understand is controversial and I in no way, shape or form, mean any—want anyone to take anything negative away from this. There are a number of people I know where ADHD is a cornerstone of their identity, where that is the thing that they are.

That is the adjective that gets hung on them the most—by choice, in many cases—and I’m always leery about going down that path because I’m super strange ever on a whole bunch of different angles, and even, “Oh, well he has ADHD. Does that explain it?” No, not really. I’m still really, really strange. But I never wanted to go down that path of it being, “Oh, Corey. The guy with ADHD.”

And again, part of this is growing up with that diagnosis. I was always the odd kid, but I didn’t want to be quote-unquote, “The freak” that always had to go to the nurse’s office to wind up getting the second pill later in the day. I swear people must have thought I had irritable bowel syndrome or something. It was never, “Time to go to the nurse, Corey.” It was one of those [unintelligible 00:27:12]. “Wow, 11:30. Wow, he is so regular. He must have all the fiber in his diet.” Yeah, pretty much.

Rich: I think that from reading that Twitter thread of yours, it sounds like you’ve done a great job at mitigating some of the downsides of ADHD. And I think it’s really important when we talk about this that we acknowledge that everybody’s experience is different. So, your experience with ADHD is likely different than mine. But there are some things that a lot of us have in common, and you mentioned some of them, that the idea of creating that Jira ticket and never following through, you put yourself in a situation where you have people around you and structures, external structures, that compensate for the things that you might have trouble with. And that’s kind of how I’m looking at it right now.

My question is, what can I do to be the most successful Rich Burroughs that I can be? And for me right now, having that coach really helps a lot because being diagnosed as an adult, there’s a lot of self-image problems that can come from ADHD. You know that you failed at a lot of things over time; people have often pointed that out to you. I was the kid in high school who the counselors or my teachers were always telling me I needed to apply myself.

Corey: “If you just tried harder and suck a little less, then you’ll be much better off.” Yeah, “Just to apply yourself. You have so much potential, Rich.” Does any of that ring a bell?

Rich: Yeah, for sure. And, you know, something my coach said to me not too long ago, I was talking about something and I said to her, I can’t do X. Like, I’m just not—it’s not possible. And her response was, “Well, what if you could?” And I think that’s been one of the big benefits to me is she helps me think outside of my preconceptions of what I can do.

And then the other part of it, that I’m personally finding really valuable, is having the goal setting and some level of accountability. She helps with those things as well. So, I’m finding it really useful. I’m sure it’s not for everybody. And like we said, everybody’s experience with ADHD isn’t the same, but one of the things that I’ve had happened since I started talking about getting diagnosed, and what I’ve learned since then, is I’ve had a bunch of people come to me.

And it’s usually on Twitter; it’s usually in DMs; you know, they don’t want to talk about it publicly themselves, but they’ll tell me that they saw my tweets and they went out and got diagnosed or their kid got diagnosed. And when I think about the difference that could make in someone’s life, if you’re a kid and you actually get diagnosed and hopefully get better treatment than it sounds like you did, it could make a really big positive impact in someone’s life and that’s the reason that I’m considering putting doing it myself is because I found that so rewarding. Some of these messages I get I’m almost in tears when I read them.

Corey: Yeah. The reason I started talking about it more is that I was hoping that I could serve as something of, if not a beacon of inspiration, at least a cautionary tale of what not to do. But you never know if you ever get there or not. People come up and say that things you’ve said or posted have changed the trajectory of how they view their careers and you’ve had a positive impact on their life. And, I mean, you want to talk about weird Gremlins in our own minds?

I always view that as just the nice things people say because they feel like they should. And that is ridiculous, but that’s the voice in my head that’s like, “You aren’t shit, Corey, you aren’t shit,” that’s constantly whispering in my ear. And it’s, I don’t know if you can ever outrun those demons.

Rich: I don’t think I can outrun them. I don’t think that the self-image issues I have are ever going to just go away. But one thing I would say is that since I’ve been diagnosed, I feel like I’m able to be at least somewhat kinder to myself than I was before because I understand how my brain works a little bit better. I already knew about the things that I wasn’t good at. Like, I knew I wasn’t a good project manager; I knew that already.

What I didn’t understand is some of the reasons why. I’m not saying that it’s all because of ADHD, but it’s definitely a factor. And just knowing that there’s some reason for why I suck, sometimes is helpful. It lets me let myself off the hook, I guess, a little bit.

Corey: Yeah, I don’t have any answers here. I really don’t. I hope that it becomes more clear in the fullness of time. I want to thank you for taking so much time to speak with me about all these things. If people want to learn more, where can they find you?

Rich: I’m @richburroughs on Twitter, and also on Polywork, which I’ve been playing around with and enjoying quite a bit.

Corey: I need to look into that more. I have an account but I haven’t done anything with it, yet.

Rich: It’s a lot of fun and I think that, speaking of ADHD, one of the things that occurred to me is that I’m very bad at remembering the things that I accomplish.

Corey: Oh, my stars, yes. People ask me what I do for a living and I just stammer like a fool.

Rich: Yeah. And it’s literally this map of, like, all the content I’ve been making. And so I’m able to look at that and, I think, appreciate what I’ve done and maybe pat myself on the back a little bit.

Corey: Which is important. Thank you so much again, for your time, Rich. I really appreciate it.

Rich: Thanks for having me on, Corey. This was really fun.

Corey: Rich Burroughs, Senior Developer Advocate at Loft Labs. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with a comment telling me what the demon on your shoulder whispers into your ear and that you can drive them off by using their true name, which is Kubernetes.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and
we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Will

Will is recovering System Administrator with a decade's worth of experience in technology and management. He now embraces the never-ending wild and exciting world of Information Security.

Links:

  • Color Health: https://www.color.com
  • Twitter: https://twitter.com/willgregorian

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at the Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by CircleCI. CircleCI is the leading platform for software innovation at scale. With intelligent automation and delivery tools, more than 25,000 engineering organizations worldwide—including most of the ones that you’ve heard of—are using CircleCI to radically reduce the time from idea to execution to—if you were Google—deprecating the entire product. Check out CircleCI and stop trying to build these things yourself from scratch, when people are solving this problem better than you are internally. I promise. To learn more, visit circleci.com.

Corey: Up next we’ve got the latest hits from Veem. Its climbing charts everywhere and soon its going to climb right into your heart. Here it is!

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Sometimes I like to talk about my previous job being in a large regulated finance company. It’s true. I was employee number 41 at a small startup that got acquired by BlackRock. I was not exactly a culture fit, as you probably can imagine by basically every word that comes out of my mouth and then imagining that juxtaposed but they’re a highly regulated finance company.

Today, my guest is someone who knows me from those days because we worked together back in that era. Will Gregorian is the head of Information Security at Color Health, and is entirely too used to my nonsense, to the point where he becomes sick of it, and somehow came back around. Will, thanks for joining me.

Will: Hello. How are you?

Corey: It’s been a while, and so far, things are better now. It turns out that I don’t have—well, I was going to say I don’t have the same level of scrutiny around my social media usage that you do at large regulated finance companies anymore, but it turns out that when you basically spend your entire day shitposting about a $1.8 trillion company in the form of Amazon, oh, it turns out your tweets get an awful lot of scrutiny. Just, you know, not by the company that pays you.

Will: That’s very true. And you knew how to actually capitalize on that.

Corey: No, I sort of basically figured that one out by getting it wrong as I went from step to step to step. No, it was a wild and whirlwind time because I joined the company as employee 41. I was the first non-developer ops hire, which happens at startups a fair bit, and developers try to interview you and ask you a bunch of algorithm questions you don’t do very well at. And they say, “Well, I have no further questions. Do you?”

And of course, there’s nothing that says bad job interview like short job interview. “Yeah, just one. What are you actually working on in an ops context?” And we talked about, I think, migrating from EC2 Classic to VPC back in those days, and I started sketching on the whiteboard, “Let me guess it breaks here, here, and here.” And suddenly, there are three more people in the room watching me do the thing on the whiteboard.

Long story short, I get hired and things sort of progressed from there. The acquisition comes down and then how, uh, we suddenly, it turns out, had this real pressing need for someone to do InfoSec on a full-time slash rigorous basis. Which is where you came in.

Will: That’s exactly where I came in. I came in a month after the acquisition, if I remember correctly. That was fun. I actually interviewed with you, didn’t I?

Corey: You did. You passed, clearly.

Will: I did pass. That’s pretty hard to pass.

Corey: It was fun, to be perfectly blunt. This is the whole problem with startup FinTech in some ways, where you’re dealing in regulated industries, but at what point do you start bringing security in, as someone—where that becomes its own function? And how do you build that out? You can get surprisingly far without it until right afterwards then you really can’t. But for a startup in the finance space, your first breach can very much be something of a death knell for the company.

Will: That’s very true. And there’s no really good calculation on when you bring those security people in, which is probably the reason why—brace yourself—we’re talking about DevSecOps.

Corey: Oh, good. Let’s put more words into DevOps because goes well.

Will: Yeah. It does. It really does. I love it. You should look at my Twitter feed; I do make fun of it. But the thing is, it’s mostly about risk. And founders ought to know what that risk is, so maybe that’s the reason why they hired me because they felt like there’s existential risk around brand and reputation, which is the reason why I joined. But yeah, [sigh] fundamentally, the problem with that is that if you hire a security practitioner, especially the first one, it’s kind of like dating, in a way—

Corey: Oh, yes.

Will: If you don’t set them up correctly, then they’re doomed to be failed, and there are plenty of complexities as a result. Imagine you’re a scrappy FinTech startup, you have a bunch of developers, they want to start writing code, they want to do big and great things, and all of a sudden security comes in and says, “Thou shalt not do the following things.” That’s where it fails. So, I think it’s part culture, part awareness from a founder perspective, part DevOps because let’s face it, most of the stuff happens in infra side. And that’s not to slam on anybody. And delicious goes on.

Corey: Yeah. Something that I developed a keen appreciation for when I went into business for myself after that and started the Duckbill Group, is that when you talk to attorneys, that was really the best way to I found to frame it because they’ve been doing this for 2000 years. It turns out InfoSec isn’t quite that old, although occasionally it feels like some of the practices are. Like, you know, password rotation every 30 days. I digress.

And lawyers will never tell you what to do, or at least anyone who’s been doing this for more than six months. Instead, the answer to everything is, “It depends. Here are the risk factors to consider; here are the trade-offs.” My wife is a corporate attorney and I learned early on not to let her have any crack at my proposal documents in those days because it’s fundamentally a sales document, but her point was, “Well, this exposes you to this risk, and this risk, and this risk, and this risk.” And it’s, “Yes, I’m aware of all of that. If I don’t know how to do what I do, effectively, I’m not going to be able to fulfill this. It’s not the contract; it is the proposal and worst case I’ll give them their money back with an apology and life goes on.”

Because at that point, I was basically a tiny one-man band, and there was no real downside risk. Worst case, the entity gets sued into oblivion; I have to go get a real job again. Maybe Amazon’s hiring, I don’t know. And it’s sort of progressed from there. Left to their logical conclusion and letting them decide how it’s going to work, it becomes untenable, and it feels like InfoSec is something of the same story where the InfoSec practitioners I’ve known would not be happy and satisfied until every computer was turned off, sunken into concrete, and then dropped into Challenger Deep out in the Pacific.

Will: Yep. And that’s part of the issue is that InfoSec, generally speaking, hasn’t kept up with the modern practices, technologies, and advancements around even methodologies and culture. They’re still very much [unintelligible 00:06:32] approaching the information security conversation, militaristically speaking; everything is very much based on DOD standards. Therein lies the problem. And funny enough, you mentioned password rotation. I vividly remember we had that conversation. Do you remember that?

Corey: It does sound familiar. I’ve picked that fight so many times in so many different places. Yeah. My current thing that drives me up a wall is, in AWS’s IAM console, you get alerts for any IAM credential parents older than 90 days and it’s not configurable. And it’s, yes, if I get a hold of someone’s IAM credentials, I’m going to be exploiting it within seconds.

And there are studies; you can prove this empirically. Turns out it’s super economical to mine Bitcoin in someone else’s Cloud account. But the 90-day idea is just—all that does—the only good part of that to me is it enforces that you don’t have those credentials stashed somewhere that they become load-bearing and you don’t understand what’s going on in your infrastructure. But that’s not really the best-practice hill, I would expect AWS to wind up staking out.

Will: Precisely. And there lies the problem is that you have basically industry standards that really haven’t adopted the cloud mentality and methodologies. The 90-day rotation comes from the world of PCI as well as a few other frameworks out there. Yeah, I agree. It only takes a
few seconds, and if somebody is account—for example, in this case, IAM account—has programmatic access, game over.

Yeah, they’re going to basically spin up a whole bunch of EC2 instances and start mining. And that’s the issue is that you’re basically trying to bolt on a very passe and archaic standard to this fast-moving world of cloud. It just doesn’t work. So, things have gotten considerably better. I feel like our last conversation was, what, circa 2015, ’16?

Corey: Yeah. That was the year I left: 2016. And then it was all right, maybe this cloud thing has legs? Let’s find out.

Will: It does. It does. It actually really does. But it has gotten better and it has matured in dramatic ways, even on the cybersecurity side of the house. So, we’re no longer having to really argue our way through, “Why do we have to rotate passwords every 90 days?”

And I’ve been part of a few of these conversations with maybe the larger institutions to say, look, we have compensating controls—and I speak their language: ‘compensating controls’—you want to basically frame it that way and you want to basically try to rationalize why, technically speaking, that policy doesn’t make sense. And if it does, well, there is a better way to do it.

Corey: I feel very similarly about the idea of data being encrypted at rest in a cloud context. Yeah in an old data center story this has happened, where people will drive a pickup truck through the wall of the data center, grab a rack into the bed and peel out of there, that’s not really a risk factor in a time of cloud, especially with things like S3 where it is pretty clear that your data does not all live in easily accessible format in one facility. You’d have to grab multiple drives from different places and assemble it all together however it is they’re doing it—I presume—and great. I don’t actually need to do any encryption at rest story there. However, every compliance regime out there winds up demanding it and it’s easier for me to just check the box and get the thing encrypted—which is super easy, and no noticeable performance impact these days—than it is for me to sit here and have this argument with the auditor.

It’s one of the things I’ve learned that would arguably make me a way better employee than I was when we worked together is I’ve learned to pick my battles. Which fights do I really need to fight and which are, fine, whatever, click the ridiculous box. Life goes on.

Will: Ah, the love of learning from mistakes. The basic model of learning.

Corey: Someday I aspire to learn from mistakes of others instead of my own. But, you know, baby steps.

Will: Exactly. And you know, what’s funny about it is that I just tweeted about this. EA had a data breach and apparently, their data breach was caused by a Slack conversation. Now, here’s my rebuttal. Why doesn’t the information security community come together and actually talk about those anti-patterns to learn from one another?

We all keep it in a very in a confidential mode. We locked it away, throw the keys away, and we never talk about why this thing happened. That’s one problem. But, yeah, going back to what you were talking about, yeah, it’s interesting. Choose your battles carefully, frankly, speaking.

And I feel like there’s a lesson to be learned there—and I do experience this from time to time—is that, look, our hands are tied. We are basically in the world of relevance and we still have to make money. Some of these things don’t make sense. I wholeheartedly agree with my engineering counterparts where these things don’t make sense. For example, the encryption at rest.

Yeah, if you encrypt the EBS volume, does really get you a whole lot? No. You have to encrypt the payload in order to be able to secure and keep the data that you want confidential and that’s a massive lift. But we don’t ever talk about that. What we talk about and how we basically optimize our conversations, at least in the current form, is let’s harp on that compliance framework that doesn’t make sense.

But that compliance frameworks makes us the money. We have to generate revenue in order to remain employed and we have to make sure that—let’s face it, we work in startups—at least I do—and we have to basically demonstrate at least some form of efficacy. This is the only thing that we have at our disposal right now. I wish that we would get to the world where we can in fact practice the true security practices that make a fundamental difference.

Corey: Absolutely. There’s a bunch of companies that would more or less look all the same on the floor of the RSA Expo—

Will: Yep.

Corey: —and you walk up and down and they’re selling what seems to be the same product, just different logos and different marketing taglines. Okay. And then AWS got into the game where they offered a bunch of native tools that help around these things, like CloudTrail logs, et cetera, and then you had GuardDuty to wind up analyzing this, and Macie to analyze this, but that’s still [unintelligible 00:12:12], and they have Detective on top of that, and Security Hub that ties it all together, and a few more. And then, because I’m a cloud economist, I wind up sitting here and doing the math out on this and yes, it does turn out the data breach would be cheaper. So, at what point do you stop hurling money into the InfoSec basket on some level?

Because it’s similar to DR; it’s a bit of a white elephant you can throw any amount of money at and still get it wrong, as well as at some point you have now gone so far toward the security side of things that you have impaired usability for folks who are building things. Obviously, you need your data to be secure, but you also need that data to be useful.

Will: Yep. The short answer to that is, I would like to find anybody who can give you the straight answer for that one. There is no [unintelligible 00:13:00] to any of this. You cannot basically say, “This is a point of stop.” If you will, from an expenditure perspective.

The fundamental difference right now is we’re trying to basically cross that chasm. Security has traditionally been in a silo. It hasn’t worked out really well. I think that security really needs to buck up and collaborate. It cannot basically remain in a control function, which is where we are
right now.

A lot of security practitioners have the belief that they are the master of everything and no one is right. That fundamentally needs to stop. Then we can have conversations around when we can basically stop spending the expenditure on security. I think that’s where we are right now. Right now, it still feels very much disparate in a not-so-good way.

It has gotten better, I think; the companies in the Valley are really trying to basically figure out how to do this correctly. I would say the larger organizations are still not there. And I want to really, sort of, sit from the sideline and watch the digital transformation thing happen. One of the larger institutions just announced that they’re going to go with AWS Cloud, I think you know who I’m talking about.

Corey: I do indeed.

Will: Yeah. [laugh]. So, I’m waiting to see what’s going to happen out of that. I think that a lot of their security practitioners are up for a moment of wake-up. [laugh].

Corey: They really are. And moving to cloud has been a fascinating case study in this. Back in 2012, when I was working in FinTech, we were doing a fair bit of work on AWS, so we did a deal with a large financial partner. And their response was, “So okay, what data centers are you
using?” “Oh, yeah, we’re hosting in AWS.”

And their response was, “No, you’re not. Where are you hosting?” “Okay, then.” I checked recently and sure enough, that financial partner now is all-in on Cloud. Great. So, I said—when one of these deals was announced—that large finance companies are one of the bellwether institutions, that when they wind up publicly admitting that they can go all-in on cloud or use a cloud provider, that is a signal to a lot of companies that are no longer even finance-adjacent, but folks who look at that and say, “Okay, cloud is probably safe.”

Because when someone says, “Oh, our data is too sensitive to live on the cloud.” “Really? Because your government uses it, your tax authority uses it, your bank uses it, your insurance underwriter uses it, and your auditor uses it. So, what makes your data so much more special than that?” And there aren’t usually a lot of great answers other than just curmudgeonly stubbornness, which, hey, I’m as guilty of as anyone else.

Will: Well, I mean, there’s a bunch of risk people sitting there and trying to quantify what the risk is. That’s part of the issue is that you have
your business people who may actually be embracing it, but then you—and your technologists, frankly speaking. But then you have the entire risk arm, who is potentially reading some white paper that they read, and they’re concluding that the cloud is insecure. I always challenge that.

Corey: Yeah, it’s who funded this paper, what are they trying to sell? Because no one says that without a vested interest.

Will: Well, I mean, there’s a bunch of server manufacturers that are going to be left out of the conversation.

Corey: A recurring pattern is that a big company will acquire a startup of some sort, and say, “Okay, so you’re on the cloud.” And they’ll view that through a lens of, “Well, obviously of course you’re on the cloud. You’re a startup; you can’t afford to do a data center build-out, but don’t worry. We’re here now. We can now finance the CapEx build-out.”

And they’re surprised to see pushback because the thing that they miss is, it was not an economic decision that drove companies to cloud. If it started off that way, it very quickly stopped being that way. It’s a capability story, it’s if I need to suddenly scale up an entire clone of the production environment to run a few tests and then shut it down, it doesn’t take me eight weeks and a whole bunch of arguing with procurement to get that. It takes me changing an argument to, ideally a command line or doing some pull request or something like that does this all programmatically, waiting a few minutes and then testing it there. And—this is the part everyone forgets—McLeod economic side—and then turning it back off again so you don’t pay for it in perpetuity.

It really does offer a tremendous boost in terms of infrastructure, in terms of productivity, in terms of capability stories. So, we’re going to move back to a data center now that you’ve been acquired has never been a really viable strategy in many respects. For starters, a bunch of you engineers are not going to be super happy with that, and are going to take their extremely hard-to-find skill set elsewhere as soon as that becomes a threat to what they’re doing.

Will: Precisely. I have seen that pattern. And the second part to that pattern, [laugh] which is very interesting is trying to figure out the compromise between cloud and on-prem. Meaning that you’re going to try to bolt-on your on-prem solutions into the cloud solution, which equally doesn’t work if not it makes it even worse. So, you end up with this quasi-hybrid model of sorts, and that doesn’t work. So, it’s all-in or nothing. Like I said, we’ve gotten to the point where the realization is cloud is the way to do it.

Corey: This episode is sponsored by our friends at Oracle HeatWave is a new high-performance accelerator for the Oracle MySQL Database Service. Although I insist on calling it “my squirrel.” While MySQL has long been the worlds most popular open source database, shifting from transacting to analytics required way too much overhead and, ya know, work. With HeatWave you can run your OLTP and OLAP, don’t ask me to ever say those acronyms again, workloads directly from your MySQL database and eliminate the time consuming data movement and integration work, while also performing 1100X faster than Amazon Aurora, and 2.5X faster than Amazon Redshift, at a third of the cost. My thanks again to Oracle Cloud for sponsoring this ridiculous nonsense.

Corey: For the most part, yes. There are occasional use cases where not being in cloud or not being in a particular cloud absolutely makes sense. And when companies come to me and talk to me that this is their perspective and that’s why they do it, my default response is, “You’re probably right.” When I talk about these things, I’m speaking about the general case. But companies have put actual strategic thought into
things, usually.

There’s some merit behind that and some contexts and constraints that I’m missing. It’s the old Chesterton’s Fence story, where it’s a logic tool to say, okay, if you come to a fence in the middle of nowhere, the naive person, “Oh, I’m going to remove this fence because it’s useless.” The smarter approach is, “Why is there a fence here? I should probably understand that before I take it down.” It’s one of those trying to make sure that you understand the constraints and the various strategic objectives that lend themselves to doing things in certain ways.

I think that nuance gets lost, particularly in mass media, where people want these nuanced observations somehow distilled down into something that fits in a tweet. And that’s hard to do.

Will: Yep. How many characters are we talking about now? 280.

Corey: 280 now, but you can also say a lot with gifs. So, that helps.

Will: Exactly, yeah. A hundred percent.

Corey: So, in your career, you’ve been in a lot of different places. Before you came over and did a lot of the financial-regulated stuff. You were at Omada Health where you were focusing on healthcare-regulated side of things. These days, you’re in a bit of a different direction, but what have you noticed that, I guess, keeps dragging you into various forms of regulated entities? Are those generally the companies that admit that they, while still in startup stage, actually need someone to focus on security? Or is there more to it that draws you in?

Will: Yeah, I know. There’s probably several different personas to every company that’s out there. You have your engineering-oriented companies who are wildly unregulated, and I’m talking about maybe your autonomous vehicle companies who have no regulations to follow, they have to figure it out on their own. Then you have your companies that are in highly regulated industries like healthcare and financial industry, et cetera. I have found that my particular experience is more applicable to the latter, not the former.

I think when you basically end up in companies that are trying to figure it out, it’s more about engineering, less about regulations or frameworks, et cetera. So, for me, it’s been a blend between compliance and security and engineering. And that’s where I strive. That doesn’t mean that I don’t know what I’m doing, it just means that I’m probably more effective in healthcare and FinTech. But I will say—you know, this is an interesting part—what used to take months to implement now is considerably shorter from an implementation timeline perspective.

And that’s the good news. So, you have more opportunities in healthcare and FinTech. You can do it nimbly, you can do things that you generally had to basically spend massive amounts of money and capital to implement. And it has gotten better. I find myself that, you know, I struggle less now, even in the AWS stack trying to basically implement something that gets us close to what is required, at least from a bare minimum perspective.

And by the way, the bare minimum is compliance.

Corey: Yes.

Will: That’s where it starts, but it doesn’t end there.

Corey: A lot of security folks start off thinking that, “Oh, it’s all about red team and pentesting and the rest, and no, no, an awful lot of InfoSec is in fact compliance.” It’s not just, do the right thing, but how do you demonstrate you’re doing the right thing? And that is not for everyone.

Will: I would caution anybody who wants to get into security to first consider how many different colors there are to the rainbow in the security side of the house, and then figure out what they really want to do. But there is a misconception around when you call security often, to your point, people kind of default to, “Oh, it’s red teaming.” Or, “It’s basically trying to break or zero-days.” Those happens seldom, although seems they’re happening far more often than they should.

Corey: They just have better marketing now.

Will: Yeah. [laugh].

Corey: They get names and websites and a marketing campaign. And who knows, probably a Google Ad buy somewhere.

Will: Yep, exactly. So, you have to start with compliance. I also would caution my DevOps and my engineering counterparts and colleagues to,
maybe, rethink the approach. When you approach a practitioner from a security side, it’s not all about compliance, and if you ask them, “Well,
you only do compliance,” they’re going to may laugh at you. Think of it as it’s all-inclusive.

It is compliance mixed with security, but in order for us to be able to demonstrate success, we have to start somewhere, and that’s where compliance is—that’s the starting point. That becomes sort of your northern light in a referential perspective. Then you figure out, okay, how do we up our game? How do we refine this thing that we just implemented? So, it becomes evolving; it becomes a living entity within the company. That’s how I usually approach it.

Corey: I think that’s the only sensible way to go about these things. Starting from a company of one to, at the time is recording, I believe we’re nine people but don’t quote me on that. I don’t want to count noses. One of the watershed moments for us when we started hiring people who—gasp, shock—did not have backgrounds as engineers themselves—it turns out that you can’t generally run most companies with only people who have been spending the last 15 years staring at computers. Who knew?—and it’s a different mindset; it’s a different approach to
these things.

And because again, it’s that same tension, you don’t want to be the Department of No. You don’t want to make it difficult for people to do their jobs. There’s some low bar stuff such as you don’t want people using a password of ‘kitty’ everywhere and then having it on a post-it note on the back of their laptop in an airport lounge, but you also don’t want them to have to sit there and go through years of InfoSec training to make this stuff makes sense. So, building up processes like we have here, like security awareness training, about half of it is garbage; I got to be perfectly honest. It doesn’t apply to how any of us do business. It has a whole bunch of stuff that presupposes that we have an office. We don’t. We’re full remote with no plans to change that. And it’s a lot of frankly, terrible advice, like, “Never click a link in email.” It’s yeah, in theory, that makes sense from a security perspective, but have you met humans?

Will: Yeah, exactly.

Corey: It’s this understanding of what you want to be doing idealistically versus what you can do with people trying to get jobs done because they are hired to serve a purpose for the company that is not security. “Security is everyone’s job,” is a great slogan and I understand where it’s going, but it’s not realistic.

Will: Nope, it’s not. It’s funny it’s you mentioned that. I’m going through a similar experience from a security awareness training perspective and I have been cycling through several vendors—one prominent one that has a Chief Hacking Officer of sorts—and amazingly enough, their content is so very badly written and so very badly optimized on the fact that we’re still in this world of going to a office or doing things that don’t make sense. “Don’t click the link?” You’re right. Who doesn’t click the link? [laugh].

Corey: Right. Oh, yeah. It’s a constant ongoing thing where you continually keep running into folks who just don’t get it, on some level. We all have that security practitioner friend who only ever sends you email that is GPG encrypted. And what do they say in those emails?

I don’t know. Who has the time to sit there and decrypt it? I’m not running anything that requires disclosure. I just don’t understand the mindset behind some of these things. The folks living off the grid as best they can, they don’t participate in society, they never have a smartphone, et cetera, et cetera. Having seen some things I’ve seen, I get it, but at some point, it’s one of those you… you don’t have to like it,
but accepting that we live in a society sort of becomes non-optional.

Will: Exactly. There lies the issue with security is that you have your wonks who are overly paranoid, they’re effectively like the your talented engineer types: they know what they’re talking about and obviously, they use open-source projects like GPG, et cetera. And that’s all great, but they don’t necessarily fit into the contemporary context of the business world and they’re seen as outliers who are basically relied on to do things that aren’t part of the normal day-to-day business operations. Then you have your folks who are just getting into it and they’re reading your CISSP guides, and they’re saying, “This is the way we do things.” And then you have people who are basically trying to cross that
chasm in between. [laugh].

And that’s where the security is right now. And it’s a cornucopia of different personalities, et cetera. It is getting better, but what we all have to collectively realize is that it is not perfect. To your point, there is no one true way of practicing security. It’s all based on how the business perceived security and what their needs are, first and foremost, and then trying to map the generalities of security into the business context.

Corey: That’s always the hardest part is so many engineering-focused solutions don’t take business context into account. I feel very aligned with this from the cost perspective. The reason I picked cost instead of something like security—because frankly, me doing basically what I’m doing now with a different position of, “Oh, I will come in and absolutely clear up the mistakes you have made in your IAM policies.” And, “Oh, we haven’t made any mistakes in our IAM policies.” You ever met someone for who not only is that true, but also is confident enough to say that? Because, “Great. We’ll do an audit. You want to bet? If we don’t find anything, we’ll give you a refund.” [laugh]. And it’s fun, but are people going to call you with that in the middle of the night and wake you up? The cloud economics thing, it is strictly a business hours problem.

Will: Yeah, yeah. It’s funny that you mention that. So, somebody makes a mistake in that IAM cloud policy. They say, “Everybody gets admin.” Next thing you know, yes, that ends up causing an auth event, you have a bunch of EC2 instances that were basically spun up by some bad actor, and now you have a $1 million bill that you have to pay.

Corey: Right. And you can get adjustments to your bill by talking to AWS support and bending the knee. And you’re going to have to get yelled at, and they will make you clean up your security policies, which you really shut it down anyway, and that’s the end of it. For the most part.

Will: I remember I spun up a Macie when it had just came out.

Corey: Oh, no.

Will: Oh, yeah.

Corey: That was $5 per gigabyte of data ingested, which is right around the breakeven point of hire a bunch of college interns to do it instead, by hand.

Will: Yeah, I remember the experience. It ended up costing $24,000 in a span of 24 hours.

Corey: Yep.

Will: [laugh].

Corey: And it was one of the most blindsidingly obvious things, to the point where they wound up releasing something like a 90% pay cut with the second generation of billing. And the billing’s still not great on something like that. I was working with a client when that came out, and their account manager immediately starts pushing it to them and they turn to me almost in unison, and, “Should we do it?”—good. We have them trained well, and I, “Hang on,”—envelope math—“Great. Running this on the data you have an S3 right now would cost for the first month, $76 million, so I vote we go with Option B, which is literally anything that isn’t that, up to and including we fund our own startup that will do this ourselves, have them go through your data, then declare failure on Medium with a slash success post of our incredible journey has come to an end; here’s what’s next. And then you pocket the difference and use it for something good.”

And then—this is at the table with the AWS account manager. Their response, “So, you’re saying we have a pricing problem with Macie?” It’s like well, “Whether it’s a problem or not really depends on what side of that transaction [laugh] you’re on, but I will say I’ll never use the thing.” And only four short years later, they fixed the pricing model.

Will: Finally. And that was the problem is that you want to do good; you end up doing bad as a result. And that was my learning experience. And then I had to obviously talk to them and beg, borrow, and steal and try to explain to them why I made that mistake. [laugh]. And then finally, you know [crosstalk 00:29:52]—

Corey: Oh, yeah. It’s rare that you can make an honest, well-intentioned mistake and not get that taken care of. But that is not broadly well known. And they of course can’t make guarantees around it because as soon as you do that you’re going to open the door for all kinds of bad actors. But it’s something where, this is the whole problem with their billing model is they have made it feel dangerous to experiment with it. “Oh, you just released a new service. I’m not going to play with that yet.”

Not because you don’t trust the service and not because you don’t trust the results you’re going to get from it, but because there’s this haunting fear of a bill surprise. And after you’ve gone through that once or twice, the scars stick with you.

Will: Yep. PTSD. I actually learned from that mistake, and let’s face it, it was a mistake and you learn from that. And I feel like I sort of honed in on the fact that I need to pay attention to your Twitter feed because you talk about this stuff. And that was really, like, the first and last mistake that I made with a AWS service stack.

Corey: Following on my Twitter feed? Yeah, first and last mistake a lot of people make.

Will: Oh, I mean, it was—that’s too, but you know, that’s a good mistake to make. [laugh]. But yeah, it was really enlightening in a good way. And I actually—you know, what’s funny about it is if you start with a AWS service that has just basically been released, be cautious and be very calculated around what you’re implementing and how you’re implementing it. And I’ll give you one example: AWS Shield, for example.

Corey: Oh, yeah. The free version or the $3,000 per month with a one-year commitment?

Will: [unintelligible 00:31:15] version. Yeah, you start there, and then you quickly realize the web application firewall rules, et cetera, they’re just not there yet. And that needs to be refined. But would I pay $3,000 for AWS Shield Advanced or something else? I probably will go with something else.

There lies the issue is that AWS is very quick to release new features and to corner that market, but they just aren’t fast enough to, like, at least in the current form—you know, from a security perspective, when you look at those services, they’re just not fast enough to refine. And there is, maybe, an issue with that, at least from my experience perspective. I would want them to pay a little bit more attention to, not so much your developers, but your security practitioners because they know what they’re looking for. But AWS is nowhere to be found on that side of the house.

Corey: Yeah. It’s a hard problem. And I’m not entirely sure the best way to solve for it, yet.

Will: Yeah, yeah. And there lies a comment where I said that we’re crossing that chasm right now…. We’re just not there yet.

Corey: Yeah. One of these days. If people want to hear more about what you’re up to and how you view these things, where can they find you?

Will: Twitter.

Corey: Always a good decision. What’s your username? And we will, of course, throw a link to it in the [show notes 00:32:33].

Will: Yeah, @willgregorian. Don’t go to LinkedIn. [laugh].

Corey: No. No one likes—LinkedIn is trying to be a social network, but not anywhere near getting there. Thank you so much for taking the time to basically reminisce with me if nothing else.

Will: This was awesome.

Corey: Really was. Will Gregorian, head of information security at Color Health. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an ignorant comment telling me why I’m wrong about rotating passwords every 60 days.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need the Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Cliff

Cliff is an Agile Consultant and self proclaimed “computer botherer.”

Links:

  • Agile Manifesto: https://agilemanifesto.org
  • Twitter: https://twitter.com/moonpolysoft

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at the Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part my Cribl Logstream. Cirbl Logstream is an observability pipeline that lets you collect, reduce, transform, and route machine data from anywhere, to anywhere. Simple right? As a nice bonus it not only helps you improve visibility into what the hell is going on, but also helps you save money almost by accident. Kind of like not putting a whole bunch of vowels and other letters that would be easier to spell in a company name. To learn more visit: cribl.io

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. My guest today has done an awful lot of things over the course of his career: startup engineering; software work; founded two startups; has been an engineering manager a bunch of places; has been the CTO at UpGuard, for example; and consulted at one point on HBO’s Silicon Valley. Also of note, he is now these days, a renowned Agile consultant. Cliff Moon, welcome to the show.

Cliff: [laugh]. Hi, Corey. Thanks for having me.

Corey: So, you and I have had energetic conversations about Agile, and based upon that context calling you an Agile consultant for enterprises is basically a deadly insult at this point. Let’s get some context on that. For those who have not heard of the term because they live wonderful, blessed lives. What is Agile? A lot of people talk about it, but always the presupposition that people listening know what it is.

Cliff: Yeah, that’s a great place to start. So, let’s go back into, sort of, prehistory. What we call Agile today is, I guess, several generations removed from the original thoughts of Agile. So, in case folks aren’t aware, to kind of lay the background, there was a group of software developers, I think it might have been in 2000, might even have been earlier than that, who came up with what they call the Agile Manifesto. And I’m not going to go through a point by point, one, because I don’t remember it; two, it’s not germane. But—

Corey: But it’s called a manifesto. I mean, if you take a look at things that have been written historically that are called manifestos, very few of them are good. Like, generally ‘manifesto’ sounds like something you wind up writing in a cabin somewhere, right before you wind up doing some sort of horrible crime that winds up living in infamy for 30 years.

Cliff: Yeah, yeah. Manifestos, they get a bad rap for good reason. Anyway, let’s not go down that [laugh] that rabbit hole. But yeah, so Agile Manifesto, right. Basically, it was this group of people, they said, “Hey, we don’t think we’re developing software the right way. This is unnecessarily painful. We’re doing things in kind of a silly fashion. Let’s refocus it around the customer. Let’s do this,” yadda, yadda, yadda.

And again, like you’re saying, it’s a manifesto; it’s not very prescriptive about what to do to solve the problem, it really just points out problems and then gives a bunch of vague statements about, “Here’s the things we should value,” or whatever. And so again, like we’re saying, as a manifesto it kind of mutates from there, and then everyone agrees; they say, “Yep, this is the wrong way. Let’s try a better way.” And down through the line, what ended up happening was a lot of people figured out that they can make money doing this, and make money being an Agile consultant, or an Agilist, if you’re, [laugh] if you like that title. So, it’s people who come in, and I guess, try to teach you how to do Agile the right way.

And the problem is that the right way usually ends up being Scrum. And so Scrum, if we can get into that, is this idea of having a two-week sprint, you plan out the work you’re going to do for that sprint, there’s a bunch of meetings you have to do that are kind of mandatory: you do a stand up every day, you do a retrospective, you do sprint planning, et cetera, et cetera. And so it’s this, like, cake-in-a-box, right? So, it’s like a ready-made thing. And, like, ta-da, now we have Agile, we’ve implemented Agile, we’ve done this Scrum thing.

The problem, of course, is that most people when they implement this—and it’s certainly not the Agile consultant in most cases because they’re basically there to just bake the cake, and then keep soaking hours until they’re forced to leave.

Corey: The problem that I had, whenever I wound up dealing with, I guess, Agile consultants in large enterprises, they always looked a lot like Agile trainers. And I don’t know what they were charging because it doesn’t matter what they were charging; they wound up gathering the entire engineering department in part of the building for two or three days to talk about how tickets worked, and how planning worked, and how to iterate forward, and how to wind up planning for spikes, and all the various terminology, and how to work with different tooling and the rest. And the reason didn’t matter what these people cost was because it was absolutely dwarfed by the sheer cost of every engineer in the company sitting through this nonsense for the better part of a week.

Cliff: Oh, absolutely. And then it’s an ongoing thing, right?

Corey: Well, it’s supposed to be an ongoing thing.

Cliff: Yeah, it’s an ongoing concern. You end up having all of these meetings that you have to do every two weeks or, God forbid, every one week, or whatever the iteration speed that they’ve laid out is. And the thing that gets lost in the sauce here is, why are you doing this? What is the point of all of this? And I think one of my favorite things to do if I’m at a company that they’re implementing Scrum, and they’ve got their sprints, and then we have our retrospective, one of the things I love to do is I love to touch the third wire during retrospective because that’s when you’re supposed to bring up, “Hey, what can we do better? What could have gone well?”

And that’s what I like to say, “Hey, can we just not do this?” [laugh]. And the response that I get is usually an indication of how hard of a job I’m going to have of trying to deprogram people. Because what it ends up being is that—and especially if you have an Agile consultant, and Agile teacher, whatever, they’re not going to like the sound of that at all, right? It’s like, “Why are we doing any of this? What is the point?”

And when you dig into it, especially when we talk about Scrum, it sort of confuses a bunch of different goals that a lot of companies don’t even necessarily have anymore. One of the tenets of Scrum is that every two weeks—or whatever your iteration speed is, whether it’s one week, two weeks, whatever—the whole system has to be shippable. So, that means that everything has to work together correctly, and then you can ship an entire, like, vertical, or monolith, or whatever. The problem with this is that, especially related to if people are deploying to the cloud or if they’re running some sort of SaaS service, this is a meaningless statement; the way that people develop software in that arena today is things get shipped immediately. The system is always shippable because the system is always up because prod is always up.

And so you ship your component, everything is backwards compatible, and then your features are behind a feature flag. So, the idea that, oh, everything has to be set in stone on a two-week cycle or whatever, it doesn’t mean anything anymore, unless of course, you have a physical artifact that you actually send to a customer, like a CD, or an image to download, or something. But if you’re doing cloud-based software—

Corey: Or a giant rocket.

Cliff: Yeah. Or giant rocket, yes. Oh, God. [laugh].

Corey: For some things, you always using waterfall. It’s like, giant rockets going to space or—more realistically for most of us or more prosaic at least—is billing systems; people don’t generally tend to iterate forward on things that charge customer credit cards. It’s a lot of planning, a lot of testing, and they roll it out, and everyone’s sweating bullets for a while.

Cliff: Right. And I would submit that, at least in my experience, most companies which have tried to implement a Scrum-type process, what they’re actually doing is they’re running a two-week waterfall. Because a lot of times they’ve got a lot of technical debt, so the idea that you can ship things immediately might be a little bit shaky, and so what you end up doing is you have this iteration speed of, like, two weeks, and then you have to plan everything out for that, and then you have to go through testing, QA, acceptance. The whole cycle has to run in a two-week sprint. And it truly is a sprint in that case because it’s too much work, it’s too much stuff, and everything just falls apart. And then they wonder why they can’t ship any software anymore. Well, it’s because you adopted this process. [laugh].

Corey: Oh, I’ve been in environments where we’ll sit down and do quarterly or, God forbid, annual planning about what we’re going to build this year, via Agile. It seems a little unlike what Agile professes to be. Now, other than the sheer aspect of hypocrisy surrounding all of this, you
take it a step further and say that it in many ways causes harm.

Cliff: Yeah. It causes harm in a couple of different ways. One of the ways that I think it is most harmful is the effect that it has on junior engineers, so people who are just starting their career, folks who are just coming out of college, and, you know, in most colleges, they don’t really teach you software engineering processes, or software engineering practices beyond the nuts and bolts of the code, or the theory of the code, or whatever, but they don’t teach you how to work in a professional environment. And so then you get a lot of folks just entering their first job, and they learn the way to do things at that job. And then they go on, they move to another job.

And someone might have, you know… they might go through ten different companies in their career, maybe some more, but they learn a certain way of doing things, and then all of a sudden, it’s like, “Yeah. I know how to do Agile. We did it at company X, Y, and Z.” And then they cargo-cult it and take it to the next place if they don’t already have it. And so it’s this sort of inculcation of younger engineers into this way of doing things that is completely harmful because most places, they don’t sit you down and tell you why we’re doing this because they don’t necessarily know why we’re doing it either.

Like I said, a two-week sprint with Scrum, the system is shippable every two weeks, you have to go through testing, and yadda, yadda, yadda, this may actually make sense in some cases. And professionally, in my experience, I’ve designed certain processes that are similar to that. Longer timeframes, but they were designed towards both the product and the team, and, sort of, the interval that they had to ship on. But in most places I’ve been, no one’s thinking about it from that perspective; they’re not thinking, “We have to design our processes around the software, or customers, or whatever.” They just kind of do, either whatever the Agile consultant tells them or whatever they learned at the last place. And so it has this effect of replicating a cargo-cult mentality throughout the entire industry, which is sad.

Corey: I’ve talked to a number of relatively Junior folks who have not heard of Agile or Scrum or any of these higher-level concepts about software development methodologies. They just walked into the workplace one day, and everyone’s doing two-week planning sessions. I’ve had people ask me six months into their career, “Why is it called a sprint?” Or, “What is up with the swim-lane style things? It seems weird, but everyone I talk to is used to it. Is it this company thing, or is this an industry thing?” And, on some level, it’s, “No, it’s just a terrible thing [laugh] that’s sort of like a mind virus that wind up taking root in an awful lot of people’s minds.”

Cliff: Yeah, absolutely. And so when we talk about the damage being done, I think that’s the worst. When you think about new people getting into the industry, having a fresh perspective, and that perspective or having an opportunity to forge a new way to do things, that kind of gets ground out of them immediately and they have to do things this set way. And this especially goes for people who end up at large companies where it’s just like, you’re not going to change anything. You’re going to get in there and you’re going to do it their way and then that’s it.

In the rare case when someone comes out of college or comes out of a training program and then they go to a startup that doesn’t have as much structure, those are really the only sorts of areas where you even have an opportunity to innovate in terms of the process of how we develop software. Because otherwise, it’s just set in stone at this point.

Corey: So, you’re given a blank slate—or a blank whiteboard, as the case may be, or God forbid, a blank Jira board—how would you structure it instead? How would you advise companies to think about software development? Since I think it’s pretty clear that an awful lot of what they’re doing today either isn’t working or is some weird bastardized hybrid of different methodologies that doesn’t really have a name other than something cynical, like ScrumBut.

Cliff: [laugh]. Yeah, so that’s a great question. So, I think where I’ll take this is, I can talk a little bit about—I mean, I’ve done this before, right, so I’ve been hired into several different startups as either, like, an engineering manager, or a director, or basically, like, hired management, and typically when a start-up hires an engineering manager or someone on that management chain, they only really do so when the pain has gotten so bad that they want to throw money at the problem. So, I specialized in that for a little bit; very thankless job, but it was interesting because what happens is that every team fails in its own unique and beautiful sort of way. [laugh].

So, one of the first places where I did this, there was a person running product; he had learned his Agile methodology from being at Booz Allen Hamilton, which, I mean, it is a nightmare factory in every metric you can measure it on. But apparently, they specialize in Agile as well as the military-industrial complex, so great. [laugh]. He was running things on a one-week sprint. And it was a shippable system, so it had a cloud component but it also had a component that was forward-deployed into a customer network.

So, he was running this where basically everyone would work on a one-week sprint; they would then do a bug-bash every Friday, and it was very much a case of, you keep doing the same thing and you keep getting the same results, and you keep doing it to see if you get different results. And they were very much in that kind of mode. So, they would do this every single week. You would have a bug-bash where the same bugs came up every single time; they wouldn’t get fixed, no one would triage them. So, the same bug was in the system in maybe, like, 10 or 15 different tickets.

No one was triaging it. And it was just a mess. And so when I got there, part of my job was to just kind of break apart this crazy structure that was happening, and again, try to design something that would actually work, again, for the product. You have to design something that has to work for how the product gets deployed. So, as I said earlier, if you have cloud services, they can deploy whenever so structuring them around
some sort of timeframe doesn’t really make a ton of sense.

However, when you have something that gets deployed into a customer network, like an agent, or some kind of desktop software, or anything that’s on machines that you do not have direct control over, you have to factor that into the speed at which you ship, you have to factor that into your engineering process. Because if you can only ship out that executable once every quarter, or—it’s like, how fast can your customer actually consume these things, right? Most places, if you give them updates every two weeks, they’re going to say, “What are you doing? Why are you making my life hard?” In a lot of places, the fastest—especially if you’re selling to an enterprise—the fastest they can consume a forward-deployed component is once a month at the very fastest.

Usually, they prefer on a quarterly or even a six-month basis. But if that’s the case, you have to design your engineering process to account for that. Then the other part is that when you land in a place like this, you can’t just pull the rug out from everybody immediately. It’s similar to saying, “Oh, we got to do a rewrite.” It’s like, well, you can’t just do a rewrite of your engineering process either; you have to incrementally make changes to it so that people are not confused about what they’re supposed to be doing, but you’re making changes towards things running in a more smooth fashion.

So, what I typically try to do is I try to design a process, and then get the team bought into it, and then hopefully get them moving faster. And the first time I tried this, it was a disaster. That was the company I was just talking about where they were running one-week sprints. I did not know what I was doing at the time; that was a very difficult situation. Landed at another place after that where this one was a two-week Scrum; similar problems around okay, frequency of testing, you have a component that gets deployed into a customer network, how fast can
we deploy that?

And similar sorts of problems, and so now that I could see what the pattern was, I could now develop a—I had a much better time developing a process that actually worked and helped the team ship with confidence. Which is really what you want the process to do is you want the process to be something that takes burden off of the engineering team, as opposed to something that makes your job as the engineering manager easier, which I think a lot of engineering managers approach it from that perspective of, “Oh, I can get a report at Jira and then I don’t have to talk to everybody every day,” or whatever. If you’re trying to make your job easier through the process, you are necessarily putting more burden onto your team.

Corey: Your company might be stuck in the middle of a DevOps revolution without even realizing it. Lucky you! Does your company culture discourage risk? Are you willing to admit it? Does your team have clear responsibilities? Depends on who you ask. Are you struggling to get buy in on DevOps practices? Well, download the 2021 State of DevOps report brought to you annually by Puppet since 2011 to explore the trends and blockers keeping evolution firms stuck in the middle of their DevOps evolution. Because they fail to evolve or die like dinosaurs.
The significance of organizational buy in, and oh it is significant indeed, and why team identities and interaction models matter. Not to mention weither the use of automation and the cloud translate to DevOps success. All that and more awaits you. Visit: www.puppet.com to download
your copy of the report now!

Corey: It feels like this is almost the early version of a similar political machination playing out where, we see it now with—there are these large companies that, once upon a time, had these big mono repos, and they had 5000 developers, and every one wound up causing problems because a group of developers is collectively referred to as a merge conflict. Then they wound up building out, “Ah, we’re going to break the monolith apart into microservices and it solves that political problem super well.” And then you wind up with a bunch of startups with five engineers working there, and they have 600 microservices running in their environment, and it feels like someone took an idea outside of the context in which it was designed for and applied it to a bunch of inappropriate areas and just bred an awful lot of complexity while actively making everything worse. Please don’t email me if people disagree with that statement. But it feels like an echo of that, doesn’t it?

Cliff: Oh, absolutely, yeah. I mean, I think a lot of this is a reflection of our relative infancy as an industry. When you think about the amount of time software engineering has been around, and has been a going concern of itself, as opposed to other engineering disciplines, I think we are still very much in our infancy. Like I was saying, they don’t really teach this sort of stuff in school, and certainly not the theory behind why you would structure things this way versus that way. In fact, most people who get promoted as a manager, you get promoted from being an engineer—someone who codes all day, or codes and does design work, but basically, someone who works as an individual contributor—then you get promoted to being a manager, and very few places give you any training or education or anything at all about how do you even do this job. And so you either sink or swim.

Corey: It’s an orthogonal skill set that basically bears little relation to what you were doing before?

Cliff: Exactly. The thing that sort of gives you any sort of ability to swim in that type of job is having the clout or the respect of your former peers as you get promoted into that. And the people who do well with that, they basically learn on the job and rely on that inbuilt respect to basically screw up a lot until they can get the hang of it. But yeah, most of the time, you don’t get an education and management or any of the other things that are not just specific to people management, but people management for software engineering, which I do think is its own discipline.

Corey: And some, I guess, almost borderline ridiculous level, it feels like no one really knows what they’re doing when it comes to management, especially in engineering. In other disciplines, it seems that management is treated as a distinct key skill, but very often—the way my management strategy evolved—and those people think I’m kidding whatever I say this, but I assure you I’m not—it came out of looking at what my terrible managers had done in the past and what didn’t work for me, and what made me quit slash become demoralized slash convince others to quit, et cetera, et cetera; or, you know, steal office supplies. Whatever it is that—how it is that you act out, and then I just did the exact opposite of those things. And I’ve been told repeatedly, “Wow, you’re a great manager.” Not really. I just don’t do all the things I hated. It gets you surprisingly far.

Cliff: Well, yeah, absolutely. But that gets you far with your own reports. There’s a whole other side to being a manager, which is dealing with the outside world. And then that’s, especially if you’re in a large organization, even in some smaller ones as well, there’s a whole dimension to the job that you as an individual engineer, you don’t even see.

That’s the politics part of the job about how do you justify what you’re doing? How do you advocate for your team? How do you operate as a quote-unquote, “Shit umbrella” for your team? And all these sorts of other things where you provide a safe harbor within the company for your team to operate, and then try to procure resources and make sure that the decision-makers above you understand the importance of what you’re doing. And no one teaches you how to do that.

Corey: Oh, never. You’re absolutely right on this. I was mostly focused on managing my reports. I completely failed in those roles managing up and, in some cases, managing sideways as well, just because that was never clear to me when I was an independent contributor working on engineering problems. It’s an evolution, on some level, of figuring out what it is that the role really is.

And all this stuff is not that complicated to teach people, but for some reason, culturally, we don’t do it. We take the Hacker News approach to things and try and figure out complex forms of interaction from first principles. And it really feels like there are some giants upon whose shoulders we could stand.

Cliff: Yeah. I mean, I agree with that. I mean, there’s definitely people in the industry who’ve written books and who are starting to try to put down that first layer of institutional knowledge to share with other folks. You got people like Camille Fournier and other folks who’ve written books specifically for engineering managers who work in the software world. Which I think is a really great first step.

But yeah, when it comes down to it, it’s like, “Okay, we’re going to implement this process; we’re going to do these things; I don’t know why.” It’s almost like no one got fired for buying IBM; no one got fired as an engineering manager for implementing Scrum. But if you try to go and do some other weird stuff, you’re running the risk of getting fired, if you fail.

Corey: There is the question of whether someone at IBM will get fired for buying Red Hat, but that’s not the analogy that people always fall back on for the last 25 years. I think that there’s also the idea that people will try and build their own thing where it makes sense for them. In complex engineering areas, that often makes sense, and sometimes it doesn’t, but then they try and approach human interaction like it’s an engineering problem, and that can lead to a lot of, frankly, disastrous outcomes, on some level. I feel like this does tie into the, I guess, almost unthinking adoption of Agile and similar methodologies or perversions of those methodologies in many large enterprises. Do you see a fix for this, or is this something that we all more or less have to live with, and watch people continue to make the same mistakes for another ten years?

Cliff: I think, for the most part, you have to—I guess, change starts at home. [laugh]. What I would advocate for is that if you have problems or qualms with the process that your particular organization is following, and you have ways you think you could fix it or changes that you’d want to make to it, then start advocating for those. And you’d be surprised about how far you can get sometimes with just saying, “Hey, can we just stop doing this, or can we do this a different way?” But I would also say that, like—one of the things you just said sort of knocked something loose from my mind, which is that even when companies share, like, “Oh, we’ve done something amazing here. We’ve designed this amazing new process, it really works well for us.”

And they write a big blog post about it, turns out if anyone ever follows up on that, they either never did it or it was never as described. And they certainly don’t do it today. So, I think a good example of this would be like when Spotify put out there—this was a number of years ago—Spotify put out their big creed about, like, “Here’s how Spotify develops software.” And they had this whole bespoke thing about they’ve got these pods of people, and you’ve got a matrix management, and they reinvented a whole bunch of stuff. And then you talk to anyone who was at Spotify during that time, and they’re like, “Yeah, we tried that; it didn’t work.” But they still put out the blog post. So. [laugh].

Corey: And I think it’s still up and hasn’t been taken down yet. It’s, “Yeah, did this work for other people?” “No, absolutely not. But it might work for us.”

Cliff: [laugh]. Yeah, it’s the same kind of trick that companies do with open-source, which is you open-source something to a bunch of fanfare and try to get people to adopt it when it hasn’t even been adopted internally. And anyone who tries figures out it’s not the right thing, and they don’t even like it. And so, but it’s like, “Oh, yeah, we can open-source it and then it comes with the imprimatur of whatever company it comes from.” I mean, this is a pretty classic joke. It’s like that old movie, The Gods Must Be Crazy, you throw the Coke bottle out of the plane; someone on the ground picks it up, and eventually ruins your life, even though it’s just a Coke bottle. Same thing with open-source; same thing with management processes.

Corey: It seems like it’s going to be one of those areas that continues to evolve whether we want it to or not. Or at least I hope because the failure is, it doesn’t.

Cliff: Yeah, I mean, hopefully it evolves. And like I said, I would say change starts at home. Try to advocate for changes on your own team and think outside of the box; try to figure out what you can get away with and try to figure out, I guess, ways to break down the walls and the rituals that the Agile consultants have set up.

Corey: Ugh. [sigh]. I hope you’re right. If people want to hear more about your thoughts on these and many other matters, where can they find you?

Cliff: Yeah, so typically, I’m just usually tweeting. So, my Twitter account is @moonpolysoft, and that’s usually where I’m doing most of my stuff. Yeah.

Corey: And we will, of course, include a link to that in the [show notes 00:25:15]. Thank you so much for taking the time to chat with me today. I really appreciate it.

Cliff: Yeah, Corey. It was great, and thanks for having me.

Corey: Cliff Moon: absolutely everything except an Agile consultant. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with a comment that you’re going to continue to iterate on and update every two weeks, like clockwork.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need the Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About John

John Allspaw has worked in software systems engineering and operations for over twenty years in many different environments. John’s publications include the books The Art of Capacity Planning (2009) and Web Operations (2010) as well as the forward to “The DevOps Handbook.” His 2009 Velocity talk with Paul Hammond, “10+ Deploys Per Day: Dev and Ops Cooperation” helped start the DevOps movement.

John served as CTO at Etsy, and holds an MSc in Human Factors and Systems Safety from Lund University

Links:

  • The Art of Capacity Planning: https://www.amazon.com/Art-Capacity-Planning-Scaling-Resources/dp/1491939206/
  • Web Operations: https://www.amazon.com/Web-Operations-Keeping-Data-Time/dp/1449377440/
  • The DevOps Handbook: https://www.amazon.com/DevOps-Handbook-World-Class-Reliability-Organizations/dp/1942788002/
  • Adaptive Capacity Labs: https://www.adaptivecapacitylabs.com
  • John Allspaw Twitter: https://twitter.com/allspaw
  • Richard Cook Twitter: https://twitter.com/ri_cook
  • Dave Woods Twitter: https://twitter.com/ddwoods2

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at the Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Corey: This episode is sponsored in part by CircleCI. CircleCI is the leading platform for software innovation at scale. With intelligent automation and delivery tools, more than 25,000 engineering organizations worldwide—including most of the ones that you’ve heard of—are using CircleCI to radically reduce the time from idea to execution to—if you were Google—deprecating the entire product. Check out CircleCI and stop trying to build these things yourself from scratch, when people are solving this problem better than you are internally. I promise. To learn more, visit circleci.com.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’m joined this week by John Allspaw, who’s—well, he’s done a lot of things. He was one of the founders of the DevOps movement—although I’m sure someone’s going to argue with that—he’s also written a couple of books, The Art of Capacity Planning and Web Operations and the foreword of The DevOps Handbook. But he’s also been the CTO at Etsy and has gotten his Master’s in Human Factors and System Safety from Lund University before it was the cool thing to do. And these days, he is the founder and principal at Adaptive Capacity Labs. John, thanks for joining me.

Corey: And now for something completely different!

John: Thanks for having me. I’m excited to talk with you, Corey.

Corey: So, let’s start at the beginning here. So, what is Adaptive Capacity Labs? It sounds like an experiment in auto-scaling, as is every use of auto-scaling, but that’s neither here nor there. I’m guessing it goes deeper.

John: Yeah. So, I managed to trick, or let’s say convince some of my heroes, Dr. Richard Cook and Dr. David Woods, these folks are what you would call heavies in the human factors, system safety, and resilience engineering world, Dave Woods is credited with creating the field of resilience engineering. And so what we’ve been doing for the past—since I left Etsy is bringing perspectives, techniques, approaches to the software world that are, I guess, some of the most progressive practices that saved other safety, critical domains, like aviation, and power plants, and all of the stuff that makes news.

And the way we’ve been doing that is largely through the lens of incidents. And so we do a whole bunch of different things, but that’s the core of what we do is activities and projects for clients that have a concern around incidents; both, are we learning well? Can you tell us that? Or can you tell us how to understand incidents and analyze them in such a way that we can learn from them effectively?

Corey: Generally speaking, my naive guess, based upon the times I spent working in various operations role has been, “Great. So, how do we learn from incidents?” Well, if you’re like most of the industry, you really don’t. You wind up blaming someone in a meeting that’s called blameless, so instead of using the person’s name, you use a team or a role name, and then you wind up effectively doing a whole bunch of reactive process work that in long enough timeline and enough incidents ossifies you into a whole bunch of processes and procedure that is just horrible. And then how do you learn from this?

Well, by the time it actually becomes a problem, you’ve rotated CIOs four times and there’s no real institutional memory here. Great. That’s my cynical approach, and I suspect it’s not entirely yours because if it were, you wouldn’t be doing a business in this because otherwise, it would be this wonderful choreographed song-and-dance number of, “Doesn’t it suck to be you? Da-da.” And that’s it. I suspect you do more as a consultant than that. So, what does my lived experience of terrible companies differing in what respects from the folks you talk to?

John: Oh, well, I mean, just to be blunt, you’re absolutely spot on. [laugh]. The industry is terrible at this.

Corey: Well, crap.

John: I mean, look, the good news is, there are inklings, there are signals for some organizations that have been doing the things that they’ve been told to do by some book or website that they read, and they’re doing all the things and they realize, “All right, well, whatever we’re doing doesn’t seem to be—it doesn’t feel—we’re doing all the things, checking the boxes, but we’re having incidents”—and even more disturbing to them is we’re having incidents that seem as if—it’d be one thing to have incidents that were really difficult, hairy, complicated, and complex, and certainly those happen, but there is a view that they’re just simply not getting as much out of these sometimes pretty traumatic events as they could be. And that’s all that’s needed, yeah.

Corey: In most companies, it seems like, on some level, you’re dealing with every incident that looks a lot like that. Sure, it was a certificate expired, but then you wind up tying into all the relevant things that are touching that. It seems like it’s an easy, logical conclusion. Oh, wow. It turns out in big enterprises, nothing is straightforward or simple.

Everything becomes complicated, and issues like that happen frequently enough that it seems like the entire career can be spent in pure firefighting reactive mode.

John: Yeah, absolutely. And again, I would say that just like these other domains that I mentioned earlier, there’s a lot of, sort of, intuitive perspectives that are, let’s just say, sort of unproductive. And so in software, we write software; it makes sense if all of our discussions after an incident trying to make sense of it, is entirely focused on the software did this, and Postgres has this weird thing, and Kafka has this tricky bit here. But the fact of the matter is, people and—engineers and non-engineers—are struggling when an incident arises, both in terms of what the hell is happening, and generating hypotheses, and working through whether the hypothesis is valid or not, adjusting it if signals show up that it’s not, and what can we do, what are some options? If we do feel like we’re on a good [unintelligible 00:06:09] productive thread about what’s happening, what are some options that we can take?

That opens up a doorway for a whole variation of other questions. But the fact of the matter is, handling incidents, understanding really, effectively, time-pressured problem solving, almost always amongst multiple people with different views, different expertise, and piecing together across that group what’s happening, and what to do about it, and what are the ramifications of doing this thing versus that thing? This is all what we would call above-the-line work. This is expertise. It shows up in how people weigh ambiguities, and things are uncertain.

And that doesn’t get this lived experience that people have, it just we’re not used to talking about—we’re used to talking about networks, and applications, and code, and network. We’re not used to talking about and even have vocabulary for what makes something confusing? What makes something ambiguous? And that is what makes for effective incident analysis.

Corey: Do you find that most of the people who are confused about these things tend to be more aligned with being individual contributor type engineers, who are effectively boots-on-the-ground, for lack of a better term? Is it high-level executives who are trying to understand why it seems like they’re constantly getting paraded in the press? Or is it often folks somewhere between the two?

John: Yes.

Corey: [laugh].

John: Right? Like there is something that you point out, which is this contrast between boots-on-the-ground, hands-on keyboard, folks who are resolving incidents, who are wrestling with these problems, and leadership. And sometimes leadership who remember their glory days of being an individual contributor sometimes are a bit miscalibrated. They still believe they have a sufficient understanding of all the messy details when they don’t. And so, I mean, the fact of the matter is, there’s the age-old story of Timmy stuck in a well, right?

There’s the people trying to get Timmy out of the well, and then there’s what to do about all of the news reporters surrounding the well asking for updates and questions, and how did Timmy get in the well? These are two different activities. And I’ll tell you pretty confidently, if you get Timmy out of the well, pretty fluidly, if you can set situations up where people who ostensibly would get Timmy out of the well are better prepared with anticipating Timmy is going to be in the well, and understanding all the various options and tools to get Timmy out of the well, the more you can set up those and have those conditions be in place, there’s a whole host of other problems that simply don’t go away. And so, these things kind of get a bit muddled. And so when you say ‘learning from incidents,’ I would separate that very much from what you tell the world externally from your company about the incident because they’re not at all the same.

Public write-ups about an incident are not the results of an analysis. It’s not the same as an internal review, were the review to be effective. Why? Well, first thing is you never see apologies on internal post-incident reviews because who are you going to apologize to?

Corey: It’s always fun watching the certain level of escalating transparency as you go up through the spectrum of the public explanation of an outage, to ones you put internal customers, to ones you show under NDA to special customers, to the ones who are basically partners who are going to fire you contractually if you don’t, to the actual internal discussion about it. And watching that play out is really interesting. As you wind up seeing the things that are buried deeper and deeper, yeah, you wind up with this flowery language on the outside, and it gets more and more transparent, and at the end, it’s, “Someone tripped and hit the emergency power switch in a data center.” And it’s this great list of how this stuff works.

John: Yeah. And to be honest, it would be strange and shocking if they weren’t different. Because like I said, the purpose of a public write-up is entirely different than an internal write-up and the audience is entirely different. And so that’s why they’re cherry-picked. There’s a whole bunch of things that aren’t included in public write-up because the purpose is, “I want a customer or potential customer to read this and feel at least a little bit better.”

Or really, I want them to at least get this notion that we’ve got a handle on it. “Wow, that was really bad, but nothing to see here, folks. It’s all been taken care of.” But again, this is very different, the people inside the organization, even if it’s just sort of tacit, they’ve got a knowledge. Tenured people who have been there for some time, see connections, even if they’re not made explicit, between one incident to another incident.

To that one that happened—“Remember that one that happened three years ago, that big one? Oh, sorry, you’re new. Oh, let me tell you the story. Oh, it’s about this and blah, blah, blah. And who knew that Unix pipes only passes 4k across it.” Blah, blah, blah, something—some weird, esoteric thing.

And so our focus, largely, although we have done projects with companies about trying to be better about their external language about it, the vast majority of what we do and where our focuses is, is to capture the richest understanding of an incident for the broadest audience. And like I said at the very beginning, the bar is real low. There’s a lot of, I don’t want to say falsehoods, but certainly a lot of myths that just don’t play out in the data about whether people are learning. Whenever we have a call with a potential client, we always ask the same question. Ask them about what their post-incident activities look like, and they tell us and throw in some cliches, and everyone—never want a crisis go to waste.

And, “Oh, yes. And we always try to capture the learnings and we put them in a document.” And we always ask the same question, which is, “Oh. So, you put these documents, these write-ups in an area?” Oh, yes, we want that to be shared as much as possible.

And then we say, “Who reads them?” And that tends to put a bit of a pause because most people have no idea whether they’re being read or not. And the fact is, when we look, very few of these write-ups are being read. Why? I’ll be blunt: because they’re terrible. [laugh].

There’s not much to learn from there because they’re not written to be read. They’re written to be filed. And so we’re looking to change that. And there’s a whole bunch of other things that are unintuitive, but just like all of the perspective shifts, DevOps, and continuous deployment, they sound obvious, but only in hindsight after you get it. That’s characterization of our work.

Corey: It’s easy to wind up, from the outside, seeing a scenario where things go super well in an environment like that, where, okay, we brought you in as a consultant, suddenly, we have better understanding about our outages. Awesome. But outages still happen. And it’s easy to take a cynical view of, okay, so other than talking to you a lot, we say the right things, but how do we know that companies are actually learning from what happened as opposed to just being able to tell better stories about pretending to learn?

John: Yeah, yeah. And this is, I think, where the world of software has some advantages over other domains. And the fact is, software engineers don’t pay any attention to anything they don’t think the attention is warranted, or they’re not being judged, or scored, or rewarded for. And so there’s no single signal that accompanies learning from incidents. It’s more like a constellation, like, a bunch of smaller signals.

So, for example, if more people are reading the write-ups. If more people are attending group review meetings. In organizations that do this really well, engineers who start attending meetings, we ask them, “Well, why are you going to this meeting?” And they’ll report, “Well, because I can learn stuff here that I can’t learn anywhere else. Can’t read about it in a runbook, can’t read about it on the wiki, can’t read about it in an email, or hear about it in an all-hands.”

And that they can see a connection between, even incidents handled in some distant group, they can see a connection to their own work. And so those are the sort of signals—we’ve written about this on our blog—those are the sort of signals that we know that progress is building momentum. But a big part of that is capturing this, again, this experience. Usually, we’ll see, there’s a timeline, and this is when memcached did X, and this alert happened, and then blah, blah, blah, blah, blah. Right?

But very rarely are captured the things that, when you ask an engineer, “Tell me your favorite incident story.” People who will even describe themselves, “Oh, I’m not really a storyteller, but listen to this.” And they’ll include parts that make for a good story. Social construct is, if you’re going to tell a story, you’ve got the attention of other people, you’re going to include the stuff that was not usually kept or captured in write-ups. For example, like, what was confusing?

A story that tells about what was confusing, well—“And then we looked, and it said, ‘zero tests failed.’”—this is an actual case that we looked at—“It says ‘zero tests failed.’ And so, okay. So, then I deployed. Well, the site went down.” “Okay, well, so what’s the story there?” “Well, listen to this. As it turns out, at a fixed font, zeros, like, in Courier or whatever, have a slash through it and at a small enough font, a zero with a slash through it looks a lot like an eight. There were eight tests failed, not zero.” So, that’s about the display. And so those are the types of things that make a good story. We all know stories like this, right? The Norway problem with YAML. You ever heard of that Norway problem?

Corey: Not exactly. I’m hoping you’ll tell me.

John: Well, so lay [laugh] it’s excellent, and of course it works out that the spec for YAML will evaluate the value no—N-O—to false as if it was a boolean. Yes, for true. Well, but if your YAML contains a list of abbreviations for countries, then you might have Ireland, Great Britain, Spain, US, false instead of Norway. And so that’s just an unintuitive surprise. And so, those are the types of things that don’t typically get captured in incident writeups.

There might be a sentence like, “There was a lack of understanding.” Well, that’s unhelpful. At best. Don’t tell me what wasn’t there. Tell me what was there. “There was confusion.” Great. “What made it confusing?” “Oh, yeah. N-O is both ‘no’ and the abbreviation for Norway.”

Red herrings is another great example. Red herrings happen a lot; they tend to stick in people’s memories; and yet, they never really get captured. But it’s, like, one of the most salient aspects of the case that ought to be captured. People don’t follow red herrings because they know they’re a red herring. They follow red herrings because they think it’s going to be productive.

So therefore, you better describe for all your colleagues what brought you to believe that this was productive. Turns out later—you find out later that it wasn’t productive. Those are some of the examples. And so if you can capture what’s difficult, what’s ambiguous, what’s uncertain, and what made it difficult, ambiguous, or uncertain, that makes for good stories. If you can enrich these documents, it means people who maybe don’t even work there yet, when they start working there, they’ll be interested; they have a set expectation they’ll learn something by reading these things.

Corey: This episode is sponsored by our friends at Oracle Cloud. Counting the pennies, but still dreaming of deploying apps instead of "Hello, World" demos? Allow me to introduce you to Oracle's Always Free tier. It provides over 20 free services and infrastructure, networking databases, observability, management, and security.

And - let me be clear here - it's actually free. There's no surprise billing until you intentionally and proactively upgrade your account. This means you can provision a virtual machine instance or spin up an autonomous database that manages itself all while gaining the networking load, balancing and storage resources that somehow never quite make it into most free tiers needed to support the application that you want to build.

With Always Free you can do things like run small scale applications, or do proof of concept testing without spending a dime. You know that I always like to put asterisks next to the word free. This is actually free. No asterisk. Start now. Visit https://snark.cloud/oci-free that's https://snark.cloud/oci-free.

Corey: There’s an inherent cynicism around… well, from at least from my side of the world, around any third-party that claims to fundamentally shift significant aspects of company culture, and if the counter-argument to that is that you and DORA and a whole bunch of other folks have had significant success with doing it, it’s just very hard to see that from the outside. So, I’m curious as to how you wind up telling stories about that because the problem is inherently whenever you have an outsider coming into an enterprise-style environment, is, “Oh, cool. What are they going to be able to change?” And it’s hard to articulate that value, and not—well, given what you do, to be direct—come across as an engineering apologist, where it’s well, “Engineers are just misunderstood, so they need empathy, and psychological safety, and blameless post-mortems.” And it sounds to crappy executives, if I’m being direct, that, “Oh, in other words, I just can’t ever do anything negative to engineers who, from my perspective, just failed me or are invisible, and there’s nothing else in my relationship with them.” Or am I oversimplifying?

John: No, no. I actually think you’re spot on. I mean, that’s the thing is that if you’re talking with leaders—remember, a.k.a. People who are, even though they’re tasked with providing the resources and setting conditions for practitioners—the hands-on folks who get their work done—they’re quite happy to talk about these sort of abstract concepts, like psychological safety and insert other sorts of hand-wavy stuff.

What is actually pretty magical about incidents is that these are grounded, concrete, messy phenomena that practitioners have, and will remember; they’re sometimes visceral experiences. And so that’s why we don’t do theory at Adaptive Capacity Labs. We understand the theory, happy to talk to you about it, but it doesn’t mean as much without the practicality. And the fact of the matter is that the engineer apologist is, “If you didn’t have the engineers, would you have a business?” That’s at the flip side; this is, like, the core unintuitive part of the field of resilience engineering, which is that Murphy’s Law is wrong.

What could go wrong almost never does, but we don’t pay much attention to that. And the reason why you’re not having nearly as many incidents as you could be is because, despite the fact that you make it hard to learn from incidents, people are actually learning. But they’re just learning out of view from leaders. When we go to an organization and we see that most of the people who are attending post-incident review meetings are managers, that has a very particular signal. That tells me that the real post-incident review is happening outside that meeting, it probably happened before that meeting, and those people are there to make sure that whatever group that they represent in their organization isn’t unnecessarily given the brunt of the bottom of a bus.

And so it’s a political due diligence. But the notion that you shouldn’t punish or be harsh on engineers for making mistakes completely misses the point. The point is to set up the conditions so that engineers can understand the work that they do. And if you can amplify that, as Andrew Schaffer has said, “You’re either building a learning organization, or you’re losing to someone who is.” And a big part of that is you need
people; you have to set up conditions for people to give detailed story about their work, what’s hard.

This part of the codebase is really scary, right? All engineers have these notions: this part is really scary, this part is really not that big of a deal, this part is somewhere in between. But there’s no place for that outside of the informal discussions. But I would assert that if you can capture that, the organization will be better prepared. The thing that I would end on that is that it’s a bit of a rhetorical device to get this across, but one of the questions we’ll ask is, “How can you tell the difference between a difficult case—a difficult incident—handled well, or a straightforward incident handled poorly?”

Corey: And from the outside, it’s very hard to tell the difference.

John: Oh, yeah. Well, certainly if what you’re doing is averaging how long these things take. But the fact of the matter is that all the people who were involved in that, they know the difference between a difficult case handled well, and a straightforward one handled poorly. They know it, but there’s nowhere, there’s no place to give voice to that lived experience.

Corey: So, on the whole, what is the tech industry missing when it comes to learning effectively from the incidents that we all continually experience and what feels to be far too frequently?

John: They’re missing what is captured in that age-old parable of the blind men and the elephant. And I would assert that these blind men that the king sends out—“Go find an elephant and come back and tell me about the elephant”—they come back and they all have—they’re all valid perspectives, and they argue about, “No, an elephant is this big flexible thing,” and other one is, “Oh, no, an elephant is this big wall,” and, “No, an elephant is a big flappy thing.” If you were to make a synthesis of their different perspectives, then you’d have a richer picture and understanding of an elephant. You cannot legislate—and this is where what you brought up—you cannot set ahead, a priori, some amount of time and effort. And quite often what we see are leaders saying, “Okay, we need to have some sort of root cause analysis done within 72 hours of an event.” Well, if your goal is to find gaps, and come up with remediation items, that’s what you’re going to get. Remediation items might actually not be that good because you’ve basically contained the analysis time.

Corey: Which does sort of feel, on some level, like it’s very much aligned as—from a viewpoint of, yeah, remediation items may not be useful as far as driving lasting change, but without remediation items, good luck explaining to your customers that will never ever, ever happen again.

John: Right, yeah. Of course. Well, you’ll notice something about those public write-ups; you’ll notice that they don’t tend to link to previous
incidents that have similarities to them because that would undermine the whole purpose, which is to provide confidence. And a reader might actually follow a hyperlink to say, “Wait a minute. You said this wouldn’t happen again.”

Turns out it would. Of course, that’s horseshit. But you’re right. And there’s nothing wrong with remediation items, but if that’s the goal, then that goal is—you know, what you look for is what you find, and what you find is what you fix. If I said, “Here’s this really complicated problem and I’m only giving you an hour to describe it,” and it took you eight hours to figure out the solution.

Well then, what you come up with in an hour is not actually going to be all that good. So, then the question is, how good are the remediation items? Quite often what we see is—and I’m sure you’ve had this experience—an incident’s been resolved and you and your colleagues are like, “Wow, that was a huge pain in the ass. Oh, dude. I didn’t see that coming. That was weird. Yeah.” And one of you might say, “You know what? I’m just going to make this change because I don’t want to be woken up tonight, or I know that making this change is going to help things. I’m not waiting for the post-mortem. We’re just going to do that.” “Is that good?” “Yep.” “Okay, yeah, please do it.”

Quite frequently, those things, those actions, those aren’t listed as action items, and yet it was a thing so important that it couldn’t wait for the post-mortem—arguably the most important action item—and it doesn’t get captured that way. We’ve seen this take place. And so again, in the end, it’s about those who have the lived experience. The live experience is what fuels how reliable you are today.

You don’t go to your senior technical people and say, “Hey, listen. We got to do this project. We don’t know how. I want you to figure out—we’re going to—let’s say we’re going to move away from this legacy thing, so I want you to get in a room, come up with two or three options. Gather a group of folks who know what they’re talking about. Get some options, and then show me what the options. Oh, and by the way, I’m prohibiting you from taking into account any experience you’ve ever had with incidents.” It sounds ridiculous when you would say that, and yet, that is what [unintelligible 00:27:54].

So, if you can fuel people’s memory, you can’t say you’ve learned something if you can’t remember it. At least that’s what my kids’ teachers tell me. And so yeah, you have to capture the lived experience, and including what was hard for people to understand. And those make for good stories. That makes for people reading them. That makes for people to have better questions about it. That’s what learning looks like.

Corey: If people want to learn more about what you have to say and how you view these things, where can they find you?

John: You can find me and my colleagues at adaptivecapacitylabs.com where we talk all about the stuff on our blog. And myself, and Richard Cook, and Dave Woods are also on Twitter, as well.

Corey: And we’ll, of course, include links to that in the [show notes 00:28:42]. John, thank you so much for taking the time to speak with me today. I really appreciate it.

John: Yeah, thanks. Thanks for having me. I’m honored.

Corey: John Allspaw, co-founder and principal at Adaptive Capacity Labs. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you hated this podcast, please leave a five-star review on your podcast platform of choice, along with a comment giving me a list of suggested remediation actions that I can take to make sure it never happens again.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need the Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Nader

  • Currently working to help build the decentralized future at Edge and Node.
  • Previously led Developer Advocacy for Front End Web and Mobile at Amazon Web Services.
  • Specializing in GraphQL, cross platform, & cloud enabled web & mobile application development
  • Developing applications & reference architectures using a combination of GraphQL & serverless technologies built on AWS
  • 4 years experience training fortune 500 companies on web & mobile application development, with the last two focused on React and React Native Training (clients include Microsoft, Amazon, US Army Corps of Engineers, Visa, ClassPass, American Express, Indeed, & Warner Bros).
  • Mobile consultant specializing in cross platform web & mobile application development
  • Author of React Native in Action (Manning Publications)
  • Author of Full Stack Serverless (O'Reilly Publications)
  • International speaker
  • Creator of React Native Elements
  • Creator of JAMstack CMS & JAMStack ECommerce

Links:

  • Edge & Node: https://edgeandnode.com
  • Js.la: https://js.la
  • Twitter: https://twitter.com/dabit3
  • Youtube: https://www.youtube.com/naderdabit

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at the Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Your company might be stuck in the middle of a DevOps revolution without even realizing it. Lucky you! Does your company culture discourage risk? Are you willing to admit it? Does your team have clear responsibilities? Depends on who you ask. Are you struggling to get buy in on DevOps practices? Well, download the 2021 State of DevOps report brought to you annually by Puppet since 2011 to explore the trends and blockers keeping evolution firms stuck in the middle of their DevOps evolution. Because they fail to evolve or die like dinosaurs. The significance of organizational buy in, and oh it is significant indeed, and why team identities and interaction models matter. Not to mention weither the use of automation and the cloud translate to DevOps success. All that and more awaits you. Visit: www.puppet.com to download your copy of the report now!

Corey: Up next we’ve got the latest hits from Veem. Its climbing charts everywhere and soon its going to climb right into your heart. Here it is!

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’m joined this week by Nader Dabit, who until recently was a senior developer advocate at AWS, and now is heading to a new company that, as of the time of this recording, you haven’t started the job yet, but you’ll be doing Developer Relations at Edge & Node.

Nader: That’s right.

Corey: First, congratulations. Career mobility is great. Secondly, welcome to the show.

Nader: Thank you. Thank you.

Corey: And thirdly, what is Edge & Node?

Nader: So yeah, this is actually the first time I’ve even spoken to anyone about this, so this is very early stuff. So, it’s exciting to even be acknowledging it. But yeah, thank you for having me here. And I’m a big fan of yours, and I know we’ve met a few times in person, so it’s always fun to keep up with the ideas that you throw out there. Yeah, Edge & Node is basically a new company, it just started in February of 2021, and it’s in the decentralized finance, or web 3 world as well.

And the general idea that they’re trying to accomplish is to facilitate companies and people that are looking to build applications in this space. It’s a fairly new space compared to the traditional web space, which would be things like what I’m doing at AWS now, even as we speak. But the general idea is that they want to facilitate companies building out these centralized applications in general, and any types of applications that are, kind of, falling into that same category.

Corey: I’m in a weird spot because I think that a lot of the technology around the distributed financial stuff—by which we, of course, mean blockchain. Cryptocurrency et cetera—is goofy, there are challenges with it across the board, et cetera, et cetera, et cetera. However, I’m deeply envious of the passion that the people who are really into that space exude around it. And maybe I’m just old and bitter and twisted, but man, I wish I could get that excited about anything.

Nader: [laugh]. Well, I mean, I’m leaving probably the best job I’ve ever had in my entire life by far. I’m leaving the most comfortable situation I’ve ever been in, in my entire life as well. I have a great team that I work with, I’ve been here for a little over three years, the stuff that we’re doing is amazing. So, for me to actually leave this position, you know, says a lot to me—or the idea that I would even leave kind of says a lot to me.

And everyone I’ve told so far, which is a few people like my wife, and a few family members, and of course, the people that work, they are also kind of shocked by it. But I think that I’ve really enjoyed what I’ve been doing, but after a little over three years, I’ve gotten a little bit—I’m the type of person that is always looking for the next new thing to do, and something to challenge me, and the more I’ve looked into the space, and also the team that I’m joining, and also the more that I look at maybe some of the challenges that we’re seeing in the world in general, that, kind of—the idea that some of these technologies are aimed to solve, especially if you look at what’s happening in countries in the last few years, like Lebanon, where you cannot take money out of the bank and they’re having massive inflation, Venezuela and other parts of the world like Turkey and Nigeria as well, there is a problem that needs to be solved. And a lot of the solutions, in my opinion, lie in these decentralized financial applications. It’s still in the early days, though, so I think it’s a really cool opportunity to get in on something, and learn it, and learn the ecosystem, and maybe even help some people while I’m doing it. Really, really passionate about the team I’m going to be working with, and I’m really excited about it.

So yeah, it’s an interesting part of my career. If I looked at the job that I have now AWS five years ago, I would have done anything to have it, right? It is such a great role, so for me to actually leave this is, it’s mind-bending even to myself. So, it’s kind of a weird and interesting time in my career, but I’m really excited.

Corey: I look forward to hearing how it treats you. But that’s a forward-looking conversation and almost one for another day. Let’s look back a bit. You’ve been at AWS for almost four years. And you’ve been a senior developer advocate around Amplify, which means that you and I haven’t spent a whole lot of time interacting because everyone says, “Oh, front end, and thus JavaScript by extension, are easy, beginner, that’s baby stuff.”

Yeah, the hell with everything about that. It is dark magic that every time I try to understand I come away more confused than I was when I started. And in anything remotely resembling a professional context, I have the good sense to pay professionals rather than try to unpack it myself. I wind up tying myself in knots. The few times I have dabbled with it, you’ve been around and have reached out to help with things that I get stuck on.

And I’m gratified, in many cases, that you’re looking at this going, “Huh. That’s interesting.” Which is engineer talk for, “What the hell is this?” And it makes me feel, “Oh, good. I’m not alone. This stuff really is confusing.” But you’ve had something of a meteoric rise as far as a known voice in space in the front end world. Tell me about that.

Nader: Yeah, it’s been really exciting to have had all this stuff happen over the last few years, and it means a lot to me, and I’m really always extremely grateful for where I am. I mean, I don’t really know. I mean, I joined AWS in January of 2018. Before that, I was doing consulting in my own company for about a year and a half, in company called React Native Training. And I think that my community involvement started really, maybe in 2016-ish when I started funding my own trips to go to conferences.

So, in Mississippi, the first company that I was working at when I started getting interested in conferences and stuff wasn’t really going to pay for me to do that type of stuff, so I started just finding events outside of Mississippi because there isn’t anything—and that’s where I live in Mississippi—there’s no tech here, right? So, I started going to these events, and I started becoming really inspired by the people that I saw that were working at these companies like Google and Facebook and Amazon, and even startups that were having really successful careers, and they were very, very inspired and excited about what they were doing. And they were speaking and I was like, wow, this seems so, so awesome because I could just see how great all that stuff was. At least it was for me. I started thinking, “Okay, if I want to become involved in these communities, what do I need to do?”

And you know, I started doing the things that I thought would get me to that point. And it was more like instead of—I don’t really want to be known; I just want to become friends with these people, and get to know them, and have these opportunities come up for me, maybe. So, I started doing writing. So, I wrote my first blog post, it was talking about something like webpack configuration or something like that—2016, that was my very first blog post—and put it out there and it actually did really well. It has over a million views at this point and it probably doesn’t even work, but people still do it.

But having that initial blog post and putting it out there having people read it and like it really spurred my continued content creation, you could call that even though I hate the word ‘content creation’—or the phrase. So, I just started doing more of that. So, I started writing, I started contributing to answers on Stack Overflow and stuff when I could. And then I started an actual meetup here in Mississippi and I started speaking at the meetup myself because there was no one here to speak at it at first. And one thing kind of led to another and then I would say, over time I started speaking at conferences.

When I joined AWS people started to take me a little more seriously because when you’re working, I guess, at a company that people don’t know and you have that on your Twitter profile, for some reason people just are less likely to follow you. Or at least I noticed that when I joined AWS, people were more likely to follow me. So, I mean, I don’t know, it just happened over time, and it’s been really exciting. And I think the thing that I like the most about what’s happened is all of the relationships that I’ve built, and all of the people that I’ve become friends with. I’m friends with so many people, and almost all of my friends now, I could say pretty much came from the tech world outside of my hometown of Jackson, Mississippi, you know.

Corey: I grew up in a small town in Maine. I have somewhat similar origin stories, except I never actually amounted to anything useful or good; I just went funny and obnoxious instead. And my road to tech was by starting off as a grumpy Unix admin, and over time evolving into a DevOps person, SRE person, managing the same teams, and then… doing whatever the hell it is that I do now. What I find interesting is that if I were talking to someone who wanted to get into tech today, I would suggest radically different things. The path that I walked as well and truly closed, so instead, what I see is this world of being able to get into tech through boot camps, through having more established paths for beginners, self-study is just as valid, too.

And I would suggest the first language someone pick up, even though I don’t understand it, to be JavaScript because it is very clear that that is the future of technology in many respects. And I’m not saying that because I love JavaScript. I don’t. I don’t like it if I’m being honest. But it is the most accessible language that is broadly being taught and the number of people in a position to learn it absolutely dwarfs the existing tech sector. I have a very hard time seeing a future where JavaScript is not the lingua franca of just about everything that requires coding. Is that overstating it?

Nader: No, I agree. I agree a hundred percent. I mean, and the way that I look at it is unless you kind of have a clear, clear understanding about where you want to go into tech and it does not align with what JavaScript has to offer, then you might as well choose JavaScript to start off with. Like you mentioned, it’s just about the numbers; with JavaScript, it’s kind of like the highest language that people are actually hiring for—if you look at all the different charts, and all the different studies and all that stuff recently—but you can also do the most with it, so if you’ve learned JavaScript and you become—if you want to say, “Oh, I don’t want to do front end, I want to do back end,” you can do that. You can build mobile apps, you can build desktop apps, you can pretty much do anything.

So, I think it’s a really great way to learn programming and also find the path that you want to have at the same time without having to learn an entire language and ecosystem and then throw all that knowledge away to learn something else because that’s not going to align with where you want to go. So yeah, I agree completely.

Corey: And I’ve know that I’m going to get letters for this, but I think that’s a good thing. It would be kind of nice. On some level, if I had at least a working knowledge of a language that could be simply and commonly used for simultaneously front end, back end tooling, et cetera. The Node ecosystem is vast. Instead, bad Python is my lingua franca, and there are ways you can theoretically transpile it into front end, and the working consensus is, “Oh, my God, never do this.”

And that’s okay. I feel like, on some level, I’m from a different era and the direction that I go in is radically different than where I would if I were starting out today. That’s okay. It’s a big industry; there’s room for all kinds. If you don’t agree with what I’m saying, that’s fine, too, because again, there are so many paths and so many ways to get there that there’s plenty of room for everyone. The pie is getting bigger, and that’s what I think it’s important to focus on, rather than trying to maximize the amount of the existing pie you can claim for yourself. And that’s something that you’ve been doing for a long time. You entered tech, what is it, 11, 12 years ago, now?

Nader: Yeah, it was in 2012. So, I guess about nine years ago.

Corey: And since then, you’ve had a number of roles that were effectively pure development. But additionally, there’s a constant and recurring theme that you wound up sort of smacking into going from senior web developer, front end web developer, software developer, web application developer—that’s when it starts to shift—software developer, software developer, then at React Native Training, you were a founder and trainer. And then you went to AWS and did senior developer advocate work. And you have shifted, in many respects, away from the person that writes the production code to the person that teaches other people to write code. And that is a massive and fundamental shift that often goes unnoticed.

Nader: Yeah, I would say that going back a lot, [laugh] I guess, when I started learning how to code, I was 29 or 30, and I tried to find a job shortly after, in Mississippi, and I’m not coming from a traditional background, either. I don’t have a diploma, even from high school; I don’t have a college degree or anything like that. So, I was coming from a very, very non-traditional background. In Mississippi there just wasn’t anything here, so I started applying for roles all over the place, like southeast United States. Didn’t get any opportunities there.

Started hitting the East Coast and the West Coast. Hired someone to spruce up my resume, I did have some stuff to put on there, just from me playing around on GitHub, and just learning, and also I did build out an e-commerce site that was actually working and stuff like that. But the general thing that I was just looking for was my foot in the door. And the first opportunity that I got was in LA. So, the LA opportunity was a contract; it wasn’t even a full time role, but I took it and I just moved to Los Angeles over the weekend because I had to be there on Monday.

And when I got to LA I was, like, put into this real developer role, right, with real developers. And it was a complete shock in every way for me, but it was probably the most exciting thing that ever happened to me in my life. Because I was, I was kind of, like, enjoying coding, and then I was like, now put into this space where I was around a bunch of really good coders, and they were just coming to work in this warehouse wearing, like, flip flops, and there was free snacks and free drinks, and there was a dog and a beanbag chair, stuff that was completely wild and new to me. So, I became really excited. And then when I was in LA, the developers that I was working with started taking me and introducing me to the community of JavaScript, I guess, or really the coding community in LA, but it was really mainly JavaScript is what we were working in.

So, when I first went to my first meetup, I was just completely blown away because not only was it just probably one of the better meetups in the world—because I’ve been to a bunch of them since then—but it was also my very first one. I didn’t even know these things existed. So, I’m going into this really gritty part of downtown LA into this huge warehouse with them, and I don’t know what to expect. And we walk in, and there’s waiters walking around with food, there’s free drinks, there’s alcoholic and non-alcoholic beverages, there’s a huge stage set up with hundreds of chairs. Everyone’s walking around talking about coding and stuff.

People from Google were there trying to hire people. To me, it was just something that I’d never imagined, that I thought was just so cool. And then I sit down and people are just going one by one teaching people things that they’ve learned in their free time. They’re taking time out of their day, this is their free time, to teach other people about how to do things. And I was just completely really in love with that idea from that moment on.

So, I got introduced to these meetups, and ever since that happened—I moved back to Mississippi about a year later, and being here, there just wasn’t anything here. There was no meetups, there was no conferences, of course, and I really missed that. So, I started getting into ways that I could do the same thing, but have it here in Mississippi. So, I first started a meetup here; we ran the meetup for a little over four years, I think, three and a half to four years maybe, where we would have bi-weekly meetups where we would do pretty much the same thing we did in LA—it was js.la was the meetup; it’s still happens—and I actually got to go back there and speak at Google, maybe a year and a half ago, which was really meaningful for me. And then after the meetup, I also started a local coding school, which never really took off, but I did it for about two years and we never really made—in fact we lost money because most of the classes were free.

Corey: It is a school. If it loses money, you’re doing it right. Until you apparently, around the bend somewhere and have an endowment for your school that’s in the billions, or you just pivot to pure for-profit, and suddenly you start making trade-offs on behalf of your students. It’s a mess.

Nader: Yeah. Exactly, yeah, I was just actually really hoping to break even and be able to have this community resource here in Mississippi because there wasn’t any. But there just wasn’t enough people interested in learning how to code. That, or maybe I wasn’t as good of a marketer as I could have been. Maybe not enough people knew about it.

But I was doing everything I could to get the word out. And we would have between one to six people maybe show up to these classes that were between one to eight hours long, and we would teach people how to write code, we would teach them how to build apps with Angular, and towards the end, I was doing React because that was a thing. And from then on—I really enjoyed that. I really enjoyed the experience of being in LA and having someone teach me, and then I learned so much. I kind of like, always was wanting to give that back, and also be involved in those types of events. So yeah, I’ve always enjoyed the teaching aspect and the learning aspect.

Corey: I really love installing, upgrading, and fixing security agents in my cloud estate! Why do I say that? Because I sell things, because I sell things for a company that deploys an agent, there's no other reason. Because let’s face it. Agents can be a real headache. Well, now Orca Security gives you a single tool that detects basically every risk in your cloud environment -- and that’s as easy to install and maintain as a smartphone app. It is agentless, or my intro would’ve gotten me into trouble here, but it can still see deep into your AWS workloads, while guaranteeing 100% coverage. With Orca Security, there are no overlooked assets, no DevOps headaches, and believe me you will hear from those people if you cause them headaches. and no performance hits on live environments. Connect your first cloud account in minutes and see for yourself at orca.security. Thats “Orca” as in whale, “dot” security as in that things you company claims to care about but doesn’t until right after it really should have.

Corey: Oh, I loved my time in LA; it was a fun place to be, there were a lot of interesting thing happens, and there are days I sincerely miss that. But, you know, life happens, things go on, we drift in different directions. One of my personal breakthrough moments was, I was contracting for another company and they sent me to various places to do various things, and they were effectively a DevOps slash sysadmin style bodies for hire, and [unintelligible 00:17:41] skills that you would accept, and one day, Puppet—at the time called Puppet Labs—reached out to them with a very weird ask, which is, “Great. We want someone who obviously checks all the technically capable boxes to come and be a traveling trainer to teach people how to use Puppet,” which these days is borderline considered negligence, but at the time, it was a good approach. And it really taught me how to, first, dive into things and understand them myself and, two, how to explain them in different ways and reach people who all have different learning styles.

And hands down, I thought that becoming a network engineer for a brief time was one of the best things I ever did to advance my career. Yeah, it doesn’t hold a candle to learning to teach other people about complex things. Because if they don’t understand what you’re talking about, it’s your failure, not theirs.

Nader: Right. Yeah, I agree. I agree completely. So, did you find that a lucrative career at the time?

Corey: It was a four-month contract. I found it very lucrative in the sense of—I mean, I was on salary for the consulting company. It’s not something that I would set out to do as an independent consultant without significant forethought. It was rough. I was on the road every week for four straight months, I was on a first-name basis with different aircrews to the different cities, and it was very wearying.

And it was also the problem that I ran into, at least for my proclivities is you’re teaching a different group of folks in a different city every week, but it’s the exact same curriculum, which is set by a different group. And the first one, you’re terrified to give it because, “Oh, my God, I’m going to mess it up.” And you somehow muddle your way through. And on the second one, you’re like, “I’ve got it.” And your third one, “Oh no, I don’t got it,” because something goes wrong.

And by the eighth, it’s repetitive and it’s the same thing, and you can use the same jokes and make people laugh and have all kinds of fun conversations. But it’s a weird problem. I mean, this was exacerbated by the fact that the training at the time was a couple $1,000. It was three days. A lot of people who were taking the class were angry at Puppet because they saw, rightly or wrongly, that this was automation software that was coming to take their job away and they were nervous and scared and didn’t want to deal with that, and I’m the representation of that company in front of them, and oh, by the way, if I get negative enough ratings, I’m fired.

So, good luck, send us a postcard when you get there. And it was how I learned to speak publicly off the cuff because you do a demo, the demo breaks—because it’s a demo—[laugh] and you have to tell a story while fixing the demo. But it became repetitive at some point, to the point where now the class does an exercise, great. And someone has a problem, and without even getting up from the front of the room, it’s, “Yeah, you forgot a comma.” And they look at you like you’re a wizard from the future.

It’s no, just at that point, there’s always one person who forgets a comma. You start seeing these things. I really enjoyed that. And I enjoyed meeting people and telling stories with them and it was a great experience; I miss a lot of it and I try and recapture it with the other stuff that I do now, but it was absolutely a watershed moment in my career.

Nader: Yeah, I would say that I was doing something similar when I was doing my React training for, like, a little over a year, and the thing that I like about AWS is—what I’m doing now—is I’m teaching different things. I’m teaching the same ideas, but I’m teaching them applied in a bunch of different ways, so I don’t really get too bored. But when I was doing React Native Training, I was kind of like you, teaching the exact same thing over and over and over, and it did start getting to the point where I was just unhappy with myself because I was, like, not learning anything myself. And even though I like teaching, I also enjoyed learning, so when I’m not pushing myself, I seem to get really antsy.

Corey: Oh, yeah. It’s one of those evolve-or-die moments where it’s… at least the way I see the world, you have to be able to effectively understand multiple points of view, you have to be able to recognize frustration when you see it and not take it personally. It requires so many things that we call soft skills, but oh my God are they hard as hell.

Nader: Yeah, yeah, absolutely. People skills or whatever you want to call it, I don’t really know what we should call it, but it’s one of those things that, for me, didn’t come naturally and it just came with practice. So, like, speaking at events, like you said, going to meetups and just interacting with people in the developer community, being on podcasts, and hosting my own podcast at one point, yeah, over time, you just get better at it. I mean, there are people born with very talented skills, like they just are really pleasant people, but for some people like myself, and a lot of other people I’m sure, it’s one of those things that you have to practice and become aware, keep looking at yourself and identifying areas where you can improve and be very, like you said, not taking anything personally. Like if someone—if they call you out on something, acknowledge that they’re probably right, and then look back into yourself and see what you can do to improve.

Corey: One of the common myths that we see across the board is that you have a sizable audience on Twitter, and I say that as someo—we’re roughly equivalent. And it’s one of those, I take a look at that and it’s weird, it’s not something that you ever, I think, come to accept. You say, like, “Wow, he has a lot of followers.” And I realize, “Wait, I have roughly the same number. Oh, wow, I have a lot of followers.”

Man, is it weird talking to an audience that is several times the size of the town I grew up in? But it feels almost like you’ve always been there that you have emerged, fully-formed, from the forehead of some god. But you’ve only been in tech for roughly a decade—please don’t take this the wrong way—you’re a smidgen older than that. And your LinkedIn doesn’t talk about what you did before that. It’s very focused on the existing tech narrative, which from a business perspective, makes sense. But let’s unpack that a bit; what were you doing before you learn to first, write code, and later, teach it?

Nader: So, this is a really important part of I would say my story whenever I really get into this because I think a lot of people have a similar path, or they’re maybe wanting to become on a similar path that I’m on now. And just hearing other people’s stories and how they’ve done it is really encouraging to a lot of people. So, I do like to share this and go into the details of before I became a developer. So, like I mentioned, I dropped out of high school because I didn’t ever get a diploma. In Mississippi, you can actually get your GED and then go into community college. So, I did that.

Corey: Hey, you have more educational credentials than I do. I don’t even have a GED.

Nader: Oh, really? Wow. [laugh]. I did not know that. I would have never guessed that.

Corey: Sidebar: Yeah, I was expelled from two boarding schools, wound up getting a diploma from some random homeschooling organization let me test out of it; found out years later they weren’t accredited, and then failed out of college. But on paper, I have an eighth-grade education and no one can ever take that away from me.

Nader: Man, you and I have a lot in common it seems like. [laugh].

Corey: Sure seems like it.

Nader: So yeah, basically, I was kind of a [BLEEP]-up from the age of, I don’t know, 18 until 29. And I had my good and bad moments, without going into all the details. But during that time, I just didn’t know what I wanted to do with my life. My father has a clothing store here in Mississippi, and he sells men’s suits in what you would say is a part of town that’s the lower-income part of town and stuff. So, and I worked there often, whenever I couldn’t find a job anywhere else, my dad would always let me come in and work for him, so I did have that as a fallback mechanism.

So, I would even consider that like a privilege almost, right, because a lot of people, you lose your job, you’re kind of out in the street. But during the time between the age of 18 and 29, I was doing all kinds of stuff. I was trying different things. I worked at restaurants for a few years early on, and then later on in my career, everything from a host to a bartender to a waiter to a manager, even. Looking back on the restaurant business, I actually don’t like it; I have so much respect for anyone that is in it.

I have a lot of respect for waiters and stuff because I understand what they’ve gone through. So, I did that and I tried working in retail for my dad, also tried working in retail for other people—like shoe stores and stuff like that—as a salesperson. I got my real estate license, and I got my real estate license in 2007-ish, like, right before the bust of 2008.

Corey: Oh, yes. And everyone was getting into real estate, and if you weren’t, people looked at you like you were nuts.

Nader: [laugh]. Exactly.

Corey: Oh, I remember those days.

Nader: So, I got my real estate license and six months later, the bust happened, and I was a failure big time in that endeavor. If I had stuck around, who knows, right? But it didn’t work out, for better or for worse. So yeah, I tried all kinds of stuff, honestly. And I was never really, like, good at anything; I never really succeeded in anything.

So, one of the times when I was working with my dad in his store, we basically wanted to put our suits online, e-commerce—actually, when I say ‘we’ it was just me. I was kind of interested in app development and web development, I was really more interested in the idea of having an app in the app store and making money off of it. That was kind of really, if I want to be completely honest, that’s how I really wanted to learn coding to do. But in the meantime, we were trying to hire developers to build us a website to sell suits online, and I was so, so out of touch with what actually needs to be done for that to happen because I had no tech background, that we were continually being disappointed because we were hiring people, you know, local developers and trying to pay them just, like, a few $1,000. Turns out, for $2 or $3,000, you can’t actually build an entire e-commerce platform, even back then. So, we failed a few times with that.

And I was like, “You know what? Let me try to build this myself.” And I took an HTML programming class in community college 10 years prior, so I knew how to write HTML and I figured that would be enough [laugh] to build out an e-commerce platform. Anyway, after doing some digging, I discovered WordPress, and with WordPress, you can basically build an e-commerce site with plugins and stuff, so that actually worked out. We, over the course of, like, nine months I, I would say, built out this e-commerce site, learned a little bit of PHP and CSS, and very, very little JavaScript, and HTML, of course, and was able to get this e-commerce store up and running and also profitable and making money.

And at certain points, we were making more money on the website than we were making in the store. So, it was a success, and it was the first real success that I had in my life, honestly. And it felt really good. And I also knew that this was my thing. I wanted to get into coding; this is the thing that seems to speak to me, that I’m okay at. And the rest is kind of history.

That was when I decided to look for a job. And I found a job, like I said, in LA and moved out over the course of a weekend so I could have my first day. And I don’t want to go into a lot of more details, but I got fired out for that, like two months later. My wife had just come out there, and a few days after she got there, I got fired from my contract. So, that was a really tough time. But got through that.

Corey: Getting fired is one of my core competencies. It’s really an underrated skill.

Nader: [laugh]. I mean, it taught me a lot. I mean, I wouldn’t ever want anyone to go through it in a really critical time like that, but here I am; I’m still around. So.

Corey: [laugh]. It seems like a lot of this built to where you are by giving you exposure to a lot of different areas. There’s a certain, I guess, hustle required in those moments when you’re, “Oh. It turns out that that money I was counting on is no longer going to be coming in because that job I thought I had doesn’t exist and I have a limited runway here.” And you combine that with various roles that have exposed you to the wonder known as the general public, and it forces you to be able to have civil conversations with people you would prefer not to. [laugh].

And I really feel like this is the sort of stuff that, although it sucks at the time when you’re going through a lot of it, it has the opportunity to really help shape a future where you can blend technology with people. And that really seems to be where the most interesting work is being done.

Nader: Yeah, and I really look forward to the future and I see a lot more people coming into tech from non-traditional backgrounds. If I could go back and do it all over, I would have loved to actually gone and gotten my computer science degree, and been a very good student, and all that stuff. Like, if I could go back with the discipline and all this stuff, I guess, maturity that I have now. But at the same time, a lot of people are like me and they go to college and maybe they don’t want to do that thing that they got a degree in, or maybe they just are not going to college at all. And there’s so many resources online now that people can use to get there.

And people start at all ages these days. I mean, there’s people that I see that are online that are starting in their 50s, in their 40s, and even in their 60s, that are just starting. And the awesome thing about technology is that there’s a lot of stuff online that you can just pick up and learn for free. And there’s different barriers to entry, and different levels of privilege, and things like that, of course, that you have to take into consideration, but at the end of the day, I feel like there is more of an opportunity for someone to come into tech and make a name for themselves and become successful than almost any other traditional discipline like engineering, or medicine, or law, where you have to have these accredited things. With tech, you don’t really have to be accredited. You are your accreditation. Like, what you say and the things that you provide online are kind of how people are going to vet you.

Corey: It’s a hard lesson, and if you can learn it, it really acts as something of a superpower. And a sad number of people seem to never quite get around there. We talk about coming from positions of privilege—and we all do to some extent—but, on some level, having to scrape a little bit early in your professional career really feels like it is, in some ways, a benefit. Now, let’s be clear, I don’t wish that on anyone, but if you have to go through it I can’t shake the feeling that it does lend itself to interesting things later career. But of course, you’ve got to get through that first.

Nader: Right. Exactly, yeah. I mean, one thing that always is in the back of my mind is, when you have gone through a lot of tough times, you almost have a post-traumatic stress syndrome that you have, that you always remember those really, really, really tough times that are financial tough times, and they always spur you to continue doing the thing that you’re doing. And I don’t know if it’s that way for people that never had to deal with any of that stuff because I have no idea, but for me for sure that’s part of it. It’s almost hard for me to say no to anything that might advance my career at this point, and it’s not a good thing, honestly.

I would like to be able to say no to more things, but because of having those bad times, you always are like, “Oh, I don’t ever want to go back to that point, so I’m going to do everything I can to continue going forward.” That’s my mental state right now. And like I said, it’s probably not the best place to be all the time because it does stress you out sometimes, but it’s one of the things I’ve never been able to shake. And I don’t know if it’s a good thing or a bad thing. It just is. It is what it is.

Corey: It is. [laugh]. Nader, thank you so much for taking the time to speak with me today. If people want to learn more about what you’re up to these days, where can they find you?

Nader: Thank you for having me, Corey. Yeah, so the number one place will probably be Twitter. So @dabit3, o D-A-B-I-T and the number three on Twitter. I’ve been on Clubhouse a little bit lately at dabit, D-A-B-I-T. And I have a YouTube channel, it’s youtube slash naderdabit. So, those three places are probably where you’ll see me around.

Corey: And we will, of course, include links to them in the [show notes 00:32:55]. Thank you so much for taking the time to speak with me. I
really appreciate it. And best of luck in your new role.

Nader: I really enjoyed the conversation. Thanks for having me.

Corey: Of course. Nader Dabit, currently about to embark on a new journey in developer relations at Edge & Node. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice and a gatekeeping comment telling me that no, JavaScript is not a good language, written entirely in Perl.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need the Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Serena

Serena Tiede is a SRE at Optum, a healthcare technology company that manages everything from the delivery of care to the management of patient data. Prior to becoming an SRE they were a Kafka operator for real time security logging and ingestion. In their off time, they moonlight as the proud admin of an incredibly over engineered Minecraft server.

Links:

  • Optim: https://www.optum.com/
  • Twitter: https://twitter.com/SerenaTiede
  • Personal Blog: https://blog.serenacodes.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Your company might be stuck in the middle of a DevOps revolution without even realizing it. Lucky you! Does your company culture discourage risk? Are you willing to admit it? Does your team have clear responsibilities? Depends on who you ask. Are you struggling to get buy in on DevOps practices? Well, download the 2021 State of DevOps report brought to you annually by Puppet since 2011 to explore the trends and blockers keeping evolution firms stuck in the middle of their DevOps evolution. Because they fail to evolve or die like dinosaurs. The significance of organizational buy in, and oh it is significant indeed, and why team identities and interaction models matter. Not to mention weither the use of automation and the cloud translate to DevOps success. All that and more awaits you. Visit: www.puppet.com to download your copy of the report now!

Corey: This episode is sponsored in part by Thinkst. This is going to take a minute to explain, so bear with me. I linked against an early version of their tool, canarytokens.org in the very early days of my newsletter, and what it does is relatively simple and straightforward. It winds up embedding credentials, files, that sort of thing in various parts of your environment, wherever you want to; it gives you fake AWS API credentials, for example. And the only thing that these things do is alert you whenever someone attempts to use those things. It’s an awesome approach. I’ve used something similar for years. Check them out. But wait, there’s more. They also have an enterprise option that you should be very much aware of canary.tools. You can take a look at this, but what it does is it provides an enterprise approach to drive these things throughout your entire environment. You can get a physical device that hangs out on your network and impersonates whatever you want to. When it gets Nmap scanned, or someone attempts to log into it, or access files on it, you get instant alerts. It’s awesome. If you don’t do something like this, you’re likely to find out that you’ve gotten breached, the hard way. Take a look at this. It’s one of those few things that I look at and say, “Wow, that is an amazing idea. I love it.” That’s canarytokens.org and canary.tools. The first one is free. The second one is enterprise-y. Take a look. I’m a big fan of this. More from them in the coming weeks.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. A recurring theme of this show has been for a while, where does the next generation of cloud engineer come from because the path I walked of being a grumpy Unix admin isn’t really as commonly available as it once was, and honestly, I wouldn’t wish my path on anyone in good conscience. My guest today is Serena Tiede, who’s a site reliability engineer at Optim and didn’t start their career as a grumpy systems administrator. Serena, welcome to the show.

Serena: Hey, thanks for having me. I’m so pumped to be here.

Corey: Don’t worry, that will soon pass. What I’m wondering is, you didn’t come to be an SRE through a giant ops background of clawing your way up by dealing with hardware and data centers and driving at unsafe speeds in the middle of the night because someone tripped over a patch cable in the data center. You have a combination of traditional/non-traditional background. Tell me about that.

Serena: Yeah. So, it’s funny you mentioned hardware. So, I went to school for electrical engineering, went to University of Minnesota because you want to do engineering, you pretty much going to one of the big state schools in the Midwest. So, I grew up and was like, “I want to be a
hardware designer.” I’m terrible at it. So terrible. [laugh].

Corey: Wait, I didn’t realize that you could want to be things you were bad at. If somebody told me that early on my career, it’s, “Huh. This might have taken a very different turn, and far more productive one.” I just assumed if I wasn’t good at something I should give up and never try it again.

Serena: Oh, I took the courses and was like, “Whoa, this is circuit design? Not for me.” Then I ended up just taking a bunch of engineering math courses. So, I took communications, the digital signal processing, controls, and started programming. I was like, all right, let’s do embedded systems. No one was hiring and then come internship time, there’s this little company that I’ve never heard of called Optim. And they’re like, “We want software engineers.” Well, I can write C. Does that count?

Corey: Oh, question, of course, to really ask is, “Oh, can you really write C having gone through it?” The more I talk to people who’ve been writing C for their entire career, and you ask them, “Can you write C?” The answer is, “Not really slash reliably. I can basically type and sometimes it works.” And, “Oh, thank God they’re mortal, too.” Was my response.

Serena: Oh, my opinion: no one should learn C unless there are specific reasons why. And those reasons are: you’re doing embedded systems where I had to learn how to write in assembly, for three weeks, and then my professor at the end said, “Hey, we’re writing C. Be thankful; it’s a high-level language.”

Corey: That is terrifying. But let’s get back to this idea of you going to school for electrical engineering, and you didn’t just dabble in it; you graduated with a degree in electrical engineering, didn’t you?

Serena: Oh, yes, I did. I graduated. It was fun even though, unfortunately, it still had my dead name on the diploma. So, I refer to that as my… matrix, Mr. Smith moment. [laugh].

Corey: They won’t go back and edit and reissue it under your actual name?

Serena: I haven’t bothered to look, but I almost consider it just kind of hilarious and just keeping it that way.

Corey: No. Again, I am not one to ever advise people how to deal with names. When I changed my name back in 2010 or so I wound up getting a whole lot of strange looks over it. And honestly, it is no one’s business, except how you interact with a name. Not the direction that we need to go in on this. I’m more interested in understanding, on some level, how you got a degree as an electrical engineer and then immediately landed a job writing software. That one feels a little strange. Can you talk me through it?

Serena: Oh, yeah. So, pretty much I took a bunch of operating systems classes and was like, “Wow, this computer science thing is cool.” But I was too far in the electrical engineering track to change degrees. So, I got the degree, ended up working at Optim. I originally started off in security, oddly enough, for my internship, then came back, did a—you know, we have a rotational program so I did security for six months and then… I wound up on this team for my second rotation where their literal job description, “Write RESTful APIs in streaming applications.”

Corey: So, it wasn’t even a software job that focused on the close-to-the-hardware stuff where you’re doing embedded systems. Like, that would at least make a bit more intuitive sense to the way I see the world. No, this was full-on up-the-stack REST API stuff.

Serena: Oh, yeah. I tried embedded, but in my market, it was all medical devices, and between all of us listening here, I don’t do well with medicine. Get very squeaked out, very faint. So decided, all right, let’s go up the stack, and turns out, it’s, like, okay, Kafka Streams. And then we were trying to figure out, “Okay, why our services—like, how do we know if it’s saturated?”

I’m like, “Oh, well, we have this Prometheus thing. This sounds cool.” And it was deployed on, like, you know, a rudimentary Kubernetes cluster. “Oh, hey, there’s this cool service discovery thing. Let’s do that.” And then one thing led to another. Thanos was coming out, and before it had a release candidate, I decided my claim to fame at the company was like, “All right, let’s do this Thanos thing because it seems really cool. I read about it on Reddit.” And the distinguished engineer in the room was like, “Oh, yeah, I heard about it on Hacker News. Do it.” I did; it was rough, but it was so cool. And then I come back, like, a year later because I went back to security for a wee bit, and the same monitoring stack is still there. And they were like, “Hey, can you do more monitoring things and pivot to observability?”

Corey: Yeah, let’s skip past the obvious joke that I could make about someone at a healthcare company saying, “Let’s do it because I read about it last night on Hacker News,” because it’s just too easy at some point. It’s odd, though, because I always held the belief, somewhat publicly, that an SRE role was not going to be a junior role. It was something that required quasi-significant experience to wind up moving into it, it’s always felt like a transition from traditional ops roles or folks who are deep in the weeds that have been doing software engineering at scale to a point where they see how these systems fail over time in production scenarios. It doesn’t sound like that was your path at all. Not to delegitimize your path by any stretch of the imagination. This is more to do with me reevaluating how I view SRE, as a field that people get into and how they approach it.

Serena: I just fell into it. And the reason why I bring up my digital signal processing background is a lot of the SRE stuff I look at all of our time-series metrics, and it’s like, “Oh. Well, this is just a real-time stream of data that we scrape periodically.” And it’s like, “Oh, cool. So, we can look at our averages, percentiles, I can eventually do some really cool fancy digital filtering.” And kind of was like, “Oh, wow. I, kind of, know the math behind a lot of this stuff and just have to just brute force apply it in places.”

Corey: Tell me a little bit more about that because with my approach to SRE—which let’s be clear, was fairly crap—the math that I tended to do was mostly addition and subtraction, and for the really deep stuff, I used the most common tool to manage anything at scale, Microsoft Excel, and that mostly handled even the addition and subtraction for me what math?

Serena: So, for me, a lot of it comes down to—I actually have my signals book in the other room—the big concept behind all these systems is the concept of sampling. You’re not going to, real-time, get memory and CPU data every second. Processors are running at gigahertz of speed, you would need double that to recreate your signal with full fidelity. That’s the Nyquist sampling theorem. But you kind of can fudge the numbers a little bit and just say, “Ehh, do we need that granular detail?”

We’re not trying to reproduce what happens in the past, we’re just trying to see what’s going on now. So, I say okay, 15-second scrape interval, things are looking good and then rolling into what I’m doing later of applying, like, “All right, let’s do some fun control loops,” because people wanted service-level objectives. People want service-level objectives; everyone loves them some SLOs and SLAs. No one wants to figure out, by hand, what their baseline is. But again, some fancy—this is more controls math—figure out what your baseline is just automatically and do some little magic in the frequency domain, courtesy of Laplace transforms, and that’s it. I can just automate that for you and remove the human from the equation.

Corey: I’m still somewhat astounded by the fact that people calculate these things out mathematically instead of, you know, dead reckoning and confident-sounding estimation.

Serena: It’s really just bringing that electr—like, controls background to software. Honestly, I’m kind of baffled that no one else is found this hack because I’m just thinking, “Oh, well, I can’t be that unique. Someone else has to have done that.” And then I talk to the people in the room and it’s like, “Oh, wait, no, I am the only person here.” [laugh].

So, that’s my whole thing. Everything is just applied math. And all of our human dead reckoning, it’s great, but it doesn’t scale well. You know, my boss wanted me to figure out how to do our SLOs for the entire team, and turns out realist—and when it came time to hire, realistically, cloning myself was not an option. [laugh]. So—

Corey: For better or worse, it seems like it isn’t. So, what was your first exposure to the SRE-style space? You started off in security, but looking at the timelines on this, it wasn’t that long ago. It feels like you were probably not exposed, in many cases, to physical data centers as much as you would be cloud, or at least not having to image bare-metal systems. Were you up at the AMI level, or was it beyond that in having virtual machines that moved around into full-on containers, or serverless?

Serena: So, I started my internship in 2016, and got my full-time offer in 2017. And we started having our—container platforms started becoming this up-and-coming thing. You know, my lead engineers were like, “All right, you've got to learn this thing called ‘Docker.’” And I have never heard of it, but I was just amazed that, “Wow, I can just run these little, little itty bitty pods anywhere on this hardware.” And later on, I did do some, like, virtual machine stuff, but I’ve had the luxury of all of these years of pain and toil, to be able to say, “Oh, yeah. I can just manage things with Ansible, create my Docker files, and do everything from a code deploy pipeline style. And it was awesome.” And I just can’t fathom what it’s like to work without those tools, but knowing… the past, it’s kind of like, “Wow, we have gotten a lot farther. Things are abstracted. This is actually kind of nice.”

Corey: It kind of is, on some level. I feel like my initial reticence towards containers—I gave a talk: “Heresy in the Church of Docker,” which sort of put me on the speaker map once upon a time—and it was about all the things that Docker as a technology didn’t really have good answers for. Honestly, the reason that I gave the talk was I assumed that it did have answers and I was just unaware of them, and I just gave the talk so I could publicly become the idiot who didn’t know what they were talking about and then get “well actually’d” to death by [ducks 00:12:40] slash Googlers. And it turns out that no, no, at that point in time, these things were not well understood or solved for. The observability stories, the metrics, the logging, the orchestration, the security story, the how you handle things like state, et cetera, et cetera, et cetera.

And Kubernetes these days has largely solved a lot of those problems, but I don’t dabble in those spaces just because of outright ornery. Back then it was a weird problem, but these problems have largely gotten solved in some ways. But I sort of just skipped over the whole Kubernetes slash container renaissance, and personally, I went directly into the serverless world. What’s your take on that?

Serena: Oh, so as someone who loved Kubernetes, I was a serverless skeptic, initially. I was like, “Well, I can just build my Docker file and write the deployment manifest. No big deal.” And then I started working on my side project. For, I think, better purposes, my iCloud account is tied to my credit card and I have to actually be on the hook for cloud bills. And I use GCP for my home lab and lo and behold, 1 million requests a month for free. And I love the sound of free when it’s my money on the line.

Corey: Oh, yeah, company money versus enterprise money, radically different scales. I mean, if you try and sell me personally a $50 hamburger, I’m going to tell you to go to hell. If you try to sell me, as representative of my company, a $50 hamburger, I’m going to need a receipt.

Serena: Exactly. And then also, like I’m just running through, I was redoing one of my serverless functions and watching the deploy steps. And then one of my coworkers introduced me, he’s like, “Hey, Serena, you hear this thing called ‘Buildpack?’” and I’m like, “No. What on earth is that thing?” And he’s like, “Oh, well, you take your code, and then it just magically turns into a container.” I’m like, “Well, crap. Show me.” And lo and behold, code goes in one end, nice little container comes out the other. And that crap was magic.

Corey: It really does change the world if you let it. I think. I know it sounds like a ridiculous, I guess, hype-driven thing to say, but for the right use case, it’s great because it removes the entire administrative burden from running services. Now, critics are going to say that well, that means you’re just putting all of your reliability in the hands of your cloud provider. Yeah, we’re kind of all doing that already; serverless just, sort of, forces us to be a little bit more honest with ourselves about that.

Serena: Oh, yeah. I mean, even if you self-host things, you’re relying on your data center ops people to, like, make sure, oh, I don’t know, your machines don’t literally catch fire. We literally had a bug one time where it’s like, “Why is this one node bad?” “Oh, actually—hey, did you increase the fan speed?” Someone had to literally go increase the fan speed for whatever servers, which, again, in the serverless and cloud provider world, I don’t think about that. The cloud is just infinite to me. It’s just computers and APIs as far as the eye can see. It’s wonderful.

Corey: It really is. It’s amazing, and it’s high level, and on some level, you went from getting a degree that required you to write assembly and super low-level stuff and figure it out hardware works into, let’s be honest, writing in your primary language, which for all of us in SRE-land is, of course, YAML.

Serena: Oh, I am a very spicy YAML engineer. YAML and a little bit of Go for what I need to make things go.

Corey: You ever notice there’s never a language called ‘Stay,’ or ‘Stop,’ or anything like that? It’s always about moving to the next thing. And we in engineering always have sprint after sprint after sprint. Never a, “It’s a marathon, not a sprint. Relax. Walk. Enjoy the journey.” Nope, nope, nope. Faster, further, sooner.

Serena: Yeah, it is honestly weird because my relatively short career span, you know, it’s 2021 and I graduated in 2017. The company is like, “Hey, you’re a senior software engineer now.” Here’s a program, here’s a budget. Go forth.

Corey: Oh, that’s lucky. It must have been amazing to have an actual budget. When I started out, I was in one of those shops where it’s, “Yeah, Palo Alto wants $4,000 for that appliance. That’s okay. We have some crappy instances and pfSense, and you know, we could wind up spending eight weeks of your time to build something not as good. Get on it.”

Serena: While the hilarious part is I’m stressing out about every single dollar I’m spending and then my boss is like, “Oh, you know, your budget is super small potatoes, right, compared to like our other stuff? Don’t sweat it. It’s fine.”

Corey: I keep making this point to the cloud providers where they’re somewhat parsimonious free tiers are damaging longer-term adoption because I look at building something myself, in my spare time in my dorm room or whatnot, and I’m spinning up some instances that talk to each other and I want to use a load balancer and I want to use a managed NAT gateway—God forbid—and at the end of the month, I get a bill for $300. And it’s, what the hell is this? I thought I was on the free tier and it scares the living hell out of us. So, we learn not to use those services that are higher level and differentiated. And then when we start working in environments that have budgeting and are corporate, we still remember that, and, “Oh, don’t use that thing. It’s expensive.” And you’ll inadvertently spend 80 times as much in what your employer is paying for your time, rather than using the high-level thing because they could not care less about a $500 a month charge. And it’s this weird thing that really serves as a drag on adoption.

Serena: It’s super wei—I actually literally had this conversation with one of my engineers who wanted to, “Hey, we’re trying to expose a GRPC thing.” And I had issues getting it to work with an ingress. And he’s like, “Do you want me to take a crack at that?” And I’m like, “Look at the price of the load balancer.” And I’m like, “Unless you can figure it out in half an hour… it is literally more expensive for you to continue tilting at that windmill than for us to just leave it be.” [laugh]. And it’s also weird. I have my personal stuff where I’m trying to keep my cloud bill to, you know, maybe a humble $100 a month max, versus, “Oh, the enterprise? Oh, yeah. That’s just logging that you’re paying for.” Which is baffling to me.

Corey: I feel like as engineers, we always, always, always fall into this trap. And maybe I fall into it worse than others because my entire business is actually lowering the bill. But when I started as an independent consultant, my bill was something like seven bucks a month, which yeah, I’m pretty content with that. And I started looking at ways to golf it lower, which in most cases is never worth the time, but in my case, I should really understand every penny of the AWS bill or I’m going to have a problem someday. And now I look at it recently because we have a number of engineers building things here, and our bill was over $2,000 a month.

And true story, by the way, it turns out that your AWS bill is not so much a function of how many customers you have; it’s how many engineers you have. And I look at this and, “Oh, my God, we need to fix that immediately.” And I spent a little bit of time on it and knocked 500 bucks off, and, “Whew, that’s better.” And it still bugs me to see a $1500 bill; it feels like it’s an awful lot of money. I mean, think of what you can buy for 1500 bucks a month.

And then in the context of the larger business picture, compared to payroll, compared to all the other nonsense we use, like Tableau, for example, it’s nothing. It is a rounding error that gets lost in the weeds. I never understood that before having access to company budgets. When I was an employee, this was never explained to me, so I was always optimizing for absolutely the wrong thing in hindsight. It feels like this is part of the problem that we run into as a culture when we don’t give our staff context to make the right decisions.

Serena: Yeah, I actually do appreciate the way my company does things because I am, like—not personally, my bank account, but I am, like, responsible if someone should ask, “Hey, what’s this charge for?” I have to say, “Oh, well, it’s for all of these things, and we need that.” But for the most part, it’s been really weird to, kind of, learn, like, one of the ways I, kind of, sped up my, like, “Okay, I need to learn how business works. What do I do?” Well, quite honestly, a lot of my cloud cost tips I have learned from your various podcasts. [laugh].

Corey: Uh-oh, that’s a problem.

Serena: No, but like, all of a sudden, all this stuff and just hanging out on tech Twitter and hearing all the advice of people and then… it was, kind of a weird way of, like, yeah, years-wise, yeah, some people might look me askance and be, like, “You’re really a senior engineer?” But then they hear me speak and it’s all about like, “Oh, well, I”—again—“I stand on the shoulders of giants,” which is awesome, and I’m honestly just hoping that one day I will write something that is very cool and then someone will say, “Oh, well, they were right on these things, but not right on this. Let’s edit this to make it a little bit better.” And the standing on the shoulders of giants trend continues.

Corey: This episode is sponsored in part my Cribl Logstream. Cirbl Logstream is an observability pipeline that lets you collect, reduce, transform, and route machine data from anywhere, to anywhere. Simple right? As a nice bonus it not only helps you improve visibility into what the hell is going on, but also helps you save money almost by accident. Kind of like not putting a whole bunch of vowels and other letters that would be easier to spell in a company name. To learn more visit: cribl.io

Corey: I’m a little taken aback by the fact that you’ve learned a lot of this stuff from the podcast because I tend to envision when I’m telling stories about this, companies that show ads, or my mythical Twitter for Pets startup. I have to remember that there are banks, like, is one of the examples of serious businesses that I use all the time. But you’re in healthcare. I’m sorry, that’s more serious than finance, just because—I hate to say this because it sounds incredibly privileged and I don’t even care—it’s only money. What is money compared to the worth of someone’s life?

I don’t think that you can ever draw an equivalent and I feel dirty every time I try. When you’re working with things that impact people’s ability to access healthcare, that is more important than showing banner ads. And a lot of the stories I tell about, “Maybe it’s okay to have downtime.” Because yeah, if AWS takes a region down issue for an afternoon and you can’t show ads to people or your website isn’t working, yeah, that’s kind of sad and it’s obviously not great for your business, but at the same time, the stories in the news are always about Amazon’s issue, not about your specific issue. If you’re in an environment where there’s a possibility that people will die if what you have built is not available, we’re having a radically different conversation.

Serena: Exactly. Fortunately, for me, I personally, not working in the, like, kind of, care delivery space, but the stuff I’m working on right now is supporting, you know, that lovely end-of-the-year where it’s open enrollment, all the employers are saying, “Hey, time to re-up your benefits.” Yeah, it’s kind of a big deal that our site doesn’t go down. Because—

Corey: Yeah. And open enrollment, to my understanding, changes based upon what plan you’re on. I’ve known companies that have open enrollment in the summertime. I believe ours winds up coinciding pretty closely with the calendar year, but I’ve certainly worked in environments where that wasn’t true. So, being able to say, “Oh, it’s fine. It’s April; no one’s doing open enrollment now.” Is it actually true?

Serena: So, it totally depends on which part of your business. If you’re going through the healthcare exchanges, that’s usually more in the fall. I think the Medicare plans, those are a little bit before the individual enrollments. And there’s a ton of these things that even though I just work tangentially, that I’m just not even in the know for. And then, of course, we talk about open enrollment, but the thing that a lot of people don’t really talk about is, so what happens when your plan goes live on January first of the next year? Yep. Our site’s still got to be up. And it’s a responsibility I take really seriously because it impacts so many people.

Corey: It really does. And it shouldn’t, to be clear. I try to avoid getting overly political on this podcast, but the state of healthcare in the United States as of the time of this recording is barbaric. And I really, really, really hope there comes a day where someone’s listening to this and laughing because it’s such an antiquated and outmoded story that isn’t true anymore. But I’m terrified that it won’t be.

And yeah, having access to a website lets you sign up for healthcare during a limited period of availability, if you miss that window, you don’t have healthcare, in many cases, until the following year when open enrollment opens again, or honestly, you wind up changing jobs because that is a qualifying event to change healthcare. “Well, I missed the open enrollment window, so I have to quit and take a job somewhere else,” is a terrifying thing. It’s bad for the business for a variety of reasons, but that pales in comparison to the fact that people have to make life-altering career decisions based upon a benefit that is routed through an employer when it should not be. Okay, I’ll climb off my soapbox.

Serena: Oh, it’s bizarre to me. Honestly, for better or worse—I argue worse—but I’m honestly optimistic. One of the weirdest things I saw that stuck out from the most recent stimulus bill was, “Oh, hey. We’re having a special enrollment period during a pandemic.” And I’m like, “You know, it’s not a hundred percent.

Maybe we should just extend it to the whole year.” But it’s better than what was the previous state, where it’s like I can’t make—I mean, even in my work life, I can’t make everything perfect. I can’t make outages go away, but I can make things just a touch better. And that’s all I can do.

Corey: Sometimes all we can do, and I wish there were better ways to handle that. I don’t know what the future is going to hold, but I also think that there are bright areas. There are aspects that are promising as far as the future being brighter than today. The overall trend—I hope—is for humanity to uplift itself.

Serena: Totally.

Corey: Again, I do want to highlight that you went in a very strange direction where you went from software engineering—a generally pleasant job—to SRE, which is horrible and would not be recommended to anyone. What guidance would you have for people who are, for some godforsaken reason, trying to figure out what their career trajectory is going to be like, and thinking that they might want to become an SRE—even if they’re not in tech yet—because for some reason they hear the stories and think there’s some nobility in suffering or whatnot?

Serena: Well, for starters, for me, it kind of came down to get real good with this great math. It’s boring, but that’s kind of the bread and butter of the concepts I’ve learned. Also for junior people, if you’re also just curious—say you’ve written an app, go over to OpenTelemetry. Go, like, instrument your stuff and see how many requests you get in a day. Start getting your hands dirty with instrumentation.

Look at how cool it is, and then maybe you want to start structuring your logs; maybe you start end up doing tracing. But at the end of the day, it’s all, for me, I think best learning is just experiential, and you know, one of the things where how do you learn from production outages? Go to happy hour with some of the senior people and listen to the stories that they tell. With enough time they become funny, but they’re also valuable learning things.

Corey: The aspect I would push back on is the hard requirement around discrete math. I don’t deny that it has been helpful for what you’ve done and how you do it. I don’t know how any of that stuff works on paper; I have an eighth-grade education. That was never my path and never my strong suit. I would agree that knowing it would have made aspects of what I do easier, but the bulk of it I don’t necessarily know that I would agree. I guess, my counterpoint slash pushback would be that if you thought you’d like this, but you don’t want to deal with the math, it’s not a hard requirement, and I don’t think that I would frame it as one.

Serena: Actually, that is a very good catch. It is not a hard requirement. I am not sitting here in my notebook, scribbling away at equations. But the concepts that I’ve learned from a while back, it’s the concepts are way more important than the actual computation itself. Because
computers do that, and a computer will absolutely run circles around me.

Corey: Most of us do, unless, you know, the computer is an overheating processor from Intel. But that’s a little bit of a low blow. Not that it
stopped me. But it was a low blow.

Serena: Well, I mean, your local science supply shop might have some liquid nitrogen. Maybe.

Corey: So, what’s next for you? You started off in security slash software engineering, transitioned on over to SRE work. What’s the next step?
What’s the beyond for you?

Serena: Ohh, great question. So, I don’t really know. I’m enjoying the SRE thing. At some point, might write a book trying to make all the concepts I have learned from my electrical engineering degree, maybe a bit more accessible, be it a series of blog posts, maybe a book. I would love to get a book published. And honestly, just writing more because knowledge should be shared, and if someone learns something from my nonsense experiments on my home lab, then cool; it’s all worth it.

Corey: I’d agree with that. I’m a big fan of learning in public. One of the, I guess, magical things that I do, for lack of a better term, is that I will stumble my way through learning a new concept that I have no idea what I’m doing, and when I get lost, I call it out because invariably, I’m not the only person who runs into that problem. But for folks who don’t have—I don’t know if it’s the platform, the seniority, the perceived gravitas, the very intentional misdirection where I fooled the entire world into thinking I know what the hell I’m doing, whatever that is, most people have a problem with admitting they don’t know something and learning in public, so anytime I can take up that mantle or that burden, I love doing it, just because I don’t have any technical credibility to lose from my point of view. I wish that were more accepted and more common. That’s why I’m so intentional about being able to talk, on some level, about the things I don’t understand or the things that I don’t get.

Serena: I love that. I used to read a bunch of philosophy books, way back when, and my big thing, this great quote—I always get it confused, Plato or Socrates, but it’s, “I know that I know nothing,” and I just run with that because I mean, even though fortunately, for me, my corner of the internet, as a non-binary person, no one’s really mean to me when I say, “Okay, I broke my DNS,” because, honestly, I knew DNS conceptually when I was setting up my Minecraft server for friends, but I never really got it until I, well, kind of, broke it, [laugh] and eventually fixed it. But I hope that over time, it becomes more acceptable to say, “I don’t know things.” Within my team, I tell anyone that’s working with me when they’re asking me a question, say, “I don’t know, but I have a feeling this rabbit hole, this trail of crumbs might lead us to an answer.” And then it’s a fun little adventure.

Corey: I miss the days when I could describe what I do is a fun little adventure. It’s now, “Oh, dear Lord, it’s this bullshit again.” [sigh]. That was my sign that I was burned out, in time, find other things to do than keeping sites up.

Now, I have no on-call responsibilities because there’s no real site to keep up. Thank you, serverless, I get to sleep at night again. But there are times I miss aspects of working in the trenches, of being able to dive deep into a problem on a very large scale architecture. The grass is always greener, somehow.

Serena: The grass is always greener. In a weird way, I actually, I complain about my on-call weeks, but I actually kind of love them. There’s a weird camaraderie about all of us dealing with a shared thing. And on my team, it’s really cool because we do this whole thing where, you know, I have these junior people asking, “Oh, am I going to go on call?” And we’re like, “Well, unfortunately, you’re not quite fully baked yet. Not quite ready. Once you’re here longer with us, then yeah, we’ll go walk you through a game day and make sure you can do all the things. But being on-call, it should not be a punishment for people.” Honestly, it’s just the greatest feedback mechanism that guides me because I say, “Wow, this stinks. This could be better.” And then try to make it better.

Corey: If people want to learn more about what you’re up to, how you think about these things, or potentially even reach out for advice, where can they find you?

Serena: So, I am on Twitter at @Serena—S-E-R-E-N-A—Tiede—T-I-E-D-E. DMs are open; come bug me. I got my lovely blog. It’s just blog.serenacodes.com. It’s pretty bare-bones, but I’ll have some new content up there hopefully pretty soon, once I get around to writing it. And say hi. I like meeting new people and learning new things. Adventures await.

Corey: And we will, of course, put a link to that in the [show notes 00:34:30]. Thank you so much for taking the time to speak with me. I really appreciate it, Serena.

Serena: Hey, thank you. I am so happy to be here. This was one of my life goals, and now I don’t know what to do now that I’ve gone up here.

Corey: That’s the problem with achieving these bucket list items. It’s, “Oh, well, I wake up the following day. Now, what do I do?” And when life
eventually returns to normal, on some level. [laugh]. Thanks so much for your time. I really appreciate it.

Serena: Thank you. Have a great day.

Corey: Serena Tiede, site reliability engineer at Optim. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you hated this podcast, please leave a five-star review on your podcast platform of choice along with a comment saying that if you think that C is a high-level language, oh, just wait until you explore the beauty and majesty of Rust.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Amy

With over ten years industry experience, Amy Arambulo Negrette has built web applications for a variety of industries including Yahoo! Fantasy Sports and NASA Ames Research Center. One of her projects modernized two legacy systems impacting the entire research center and won her a Certificate of Excellence from the Ames Contractor Council. More recently, she built APIs for enterprise clients for a cloud consulting firms and led a team of Cloud Software Engineers. Amy has survived acquisitions, layoffs, and balancing life with two small children.

Links:

  • The Duckbill Group: http://duckbillgroup.com/
  • @nerdypaws: https://twitter.com/nerdypaws

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Your company might be stuck in the middle of a DevOps revolution without even realizing it. Lucky you! Does your company culture discourage risk? Are you willing to admit it? Does your team have clear responsibilities? Depends on who you ask. Are you struggling to get buy in on DevOps practices? Well, download the 2021 State of DevOps report brought to you annually by Puppet since 2011 to explore the trends and blockers keeping evolution firms stuck in the middle of their DevOps evolution. Because they fail to evolve or die like dinosaurs. The significance of organizational buy in, and oh it is significant indeed, and why team identities and interaction models matter. Not to mention weither the use of automation and the cloud translate to DevOps success. All that and more awaits you. Visit: www.puppet.com to download your copy of the report now!

Corey: And now for something completely different!

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’m joined this week by my colleague, Amy Arumbulo Negrette, who’s a cloud economist here at The Duckbill Group. Amy, thank you for taking the time to, basically, deal with my slings and arrows instead of the ones that clients throw your way.

Amy: It’s perfectly fine. It’s not as if we are… not the kindest people within the Slack channels anyway. So, I am totally good. [laugh].

Corey: [laugh]. So, you’ve been at The Duckbill Group, as of the time of this recording, which when you’re releasing things in the future, it’s always a question of how long will you have been here by then? No. We’re playing it straight here from a perspective of, as of the time of this recording, you’ve been here six months, all of which, of course, have been during the global pandemic. So first, what’s that been like?

Amy: It has been very loud. And that’s to say, I live in a house with five other people in it, so it’s one thing for me to be a remote worker and just being at my desk, working quietly, but also having to manage noise that you can’t really control, it’s been an extra level of stress that I could possibly do without. It’s fine. [laugh].

Corey: One of the whole problems with the pandemic, from our perspective, has been that we’ve run this place as a full remote operation since it was started, and people come at this from a perspective of, “Well, this whole experience we’ve had with working remote is awful. It’s terrible. No one likes it. I’m not productive.” Let’s be very clear here. There’s been a global pandemic; this is not like most years, and there are stressors and things that absolutely suck about this that don’t normally impact the remote work story quite the way that they have.

Amy: I totally agree. At least before, in one of my previous companies had an office in Chicago, so I would be there once a week, but effectively I was remote because that was an all meetings type of day. And the difference between that and now is that you had very explicit work hours; you had client hours; your work sometimes brought you out the house, if you had to go on travel, or on-site. This is just everything is done within the same ten feet of basically where you work, and you eat, and you sleep, if you have really unhealthy living habits like I do. And while I’m trying to get better at it, I’m also not the best at having time-based boundaries. I’m only good with physical boundaries. So, I
have to turn off the work computer, to turn on the fun computer, which are physically next to each other, but I have to look in a different
direction. And that is as close as I get.

Corey: Yeah. It’s the good screen versus the bad screen model of, “Oh, yeah, we’re going to stop doing work now and just move our gaze slightly to one side and look at the fun screen and work on those things instead.” And at first, I was trying to be militant when we started the whole pandemic thing and working full remote of booting people of, “Hey, all right, it’s quitting time. Go home and stop it.” The other side of that, though, is some people are, in some ways, using work to escape.

So, we’ve modified our approach to get the work done. If you’re working consistently more than 40 hours to get it done, let us know; that’s a problem. But let people work when they want to work, how they want to work and be empathetic humans. And that carries surprisingly far. Now, the question, of course, becomes, does this scale to a company that has 50,000 employees? I don’t know. That’s never been a problem that anyone has asked me to solve. But it works for us at small scale.

Amy: I find that the attitude working here has been really understanding as to, we know what our lives are like and we know what kind of work that we actually have to get done within a certain time period. And all of us make those. We don’t feel the need to explain how we were able to get that work done or what time slots happen. You’ll see me and Jesse—one of the other cloud economists—it’s like we’ll be hitting document at the same time at—for my time would be closer to later in the evening, but that’s just because I spent most of the day either taking care of kids or handling house management sort of duties. So, having that flexibility on where my schedule goes without having to answer that question of, “Well, why were you working at that hour?” I feel gives me a lot more control and takes one less thing away that I have to worry about.

Corey: One approach that we’ve always taken here has been that we treat people functionally like adults. And that’s sort of an insulting way to frame it. Like, “What are you saying. That a lot of employers treat their staff like kids?” Well, basically, yes, is the short answer, where it’s, “great, we’re going to trust you with root in production, or a bunch of confidential customer data, but we’re also not going to trust you to make a $50 purchase on the company credit card without a bunch of scrutiny because we don’t trust that you’re not embezzling.”

The cross-incentives of different organizational structures are so twisted at that point that it’s very hard to self-correct. I mean, our approach going into this was always never to go down that dark path, and so far, so good. Will it bite us someday if we continue to grow to that 50,000 person company? Undoubtedly. But I have to believe it can be done.

Amy: I would also think the pressure of managing 50,000 people would break every single person in this company. Just thinking about it gives me panic.

Corey: Oh, yeah. When every person becomes effectively, what, 500 people, that’s divisional stuff. I don’t think that anyone wants to stick around and see a small company go through those kinds of transformations because functionally, it becomes such a radically different place. For better or worse, that’s not something we have to worry about, at least not anytime soon.

Amy: We work really well with our fairly tight teams, I think. And it’s I think it’s one of the virtues of the type of work that we do. You don’t need a team of 15 people looking at one document.

Corey: No. And invariably, it seems like that tends to slow things down. Let’s talk a little bit about something that makes you a bit of an outlier insofar as when you are only dealing with a small number of people everyone’s inherently an outlier. You’re the only cloud economist at The Duckbill Group with a background as a software engineer.

Amy: Yeah. I did not realize that when I was joining. So, this deal with my infrastructure background is that I’ve only ever done enough infrastructure to support my applications. Like, Pete, I come from a startup background initially in my career, where you had to manage your application, you had to manage your own routes, you had to manage your own database connections, your own storage connections, and all of that, which, looking back on it and saying it out loud, sounds like a really bad idea now, but this was life before infrastructure as code and it was the quote-unquote, “Wild West” of San Francisco startups, where you can make a product out of basically anything. And the anything I went into was fantasy football of all things.

So, I always made sure to have enough knowledge so that if I knew something broke, that I could blame it on networking, and then I would be able to show the paper trail to prove it. That was the extent of my knowledge. And then I started getting into the serverless space, where I started building things out of cloud services, instead of just spinning up more EC2 because I was young and impressionable. And that gave me a lot more understanding of what infrastructure engineering was like. But beyond that, I build APIs, I do business logic; that’s really where my comfort levels are.

Corey: For the rest of us, we seem to come from a background of grumpy Unix sysadmin types where we were running infrastructures, but as far as the code that was tied into it, eh, that was always the stuff we would kind of hand-wave over or we’d go diving into it as little as possible. And that does shape how you go through your career. For example, most companies are not going to wind up needing someone in that role until they’ve raised at least a Series B round, whereas in many cases, “Oh, you’re an engineer. Great, why don’t you start the company yourself?” From a software engineering perspective. It’s a different philosophy in many respects, and one that I think is a little bit on the strange side if that makes sense.

Amy: It is. It’s extremely strange how completely dependent the two are on each other, yet the mindset to get into either and of that level of engineering is completely different. My husband and my father are both on the infrastructure side, and they’ve tried to explain networking to me my entire life, and I just—the minute the word subnet comes up, my brain is gone. It’s like I’m replaying Star Trek episodes in my head because I can no longer handle this [laugh] conversation.

Corey: That’s how a lot of us feel about various code constructs at some point. For better or worse, we’ve made our peace with it, and let’s be very direct here for a minute, we’ve learned to talk our way around customer questions that go too deep into the software engineering space, by and large. What’s it like on the other side of that, where there’s an expectation that you have a lot more in-depth infrastructure experience than perhaps you do? Or isn’t there at that expectation?

Amy: There is, I think a lot of that is just because of the type of industry this is. Cloud consulting is always infrastructure first because that is what the cloud is selling; they are selling managed infrastructures. They are giving you data center alternatives, but they’re not giving you are full-blown apps. And whenever they do—let’s say Lightsail—it’s an expensive thing that you, somewhere in your mind go, “I could build that cheaper. Why am I paying for this service?”

So, when I am on the phone with clients and they have a situation that is obviously going to be a software solution, where their infrastructure is growing, but it’s because their software has a specific requirement, either for logging, or for surge, that they’re using either Elasticsearch, or Kubernetes, or CloudWatch Metrics for, and it’s turning into an expensive solution, it gives me an inside, “Well, this is the kind of engineering effort that’s going to need to happen in order for you to write all of these problems and to reduce these costs. And these aren’t as simple as hitting an option within AWS console to bring all of that down.” It’s always going to be seen as more of an effort, but you also get a bit of empathy from the engineers you’re talking to because you now are explaining to them that you understand what they built. You understand why they built it a specific way, and you’re just trying to give them a path out.

Corey: You mentioned the now antiquated idea of going on-site and talking to clients. I mean, before pandemic, Mike and I would head out to a lot of our clients for the final wrap-up meeting, or even in some cases kickoffs, because it made sense for us to do it and get everyone in the same room and on the same page. And over the past year, we’ve found ways to solve for these problems in ways that I don’t necessarily know are going to go away once the pandemic is over. Is it more effective for us to travel somewhere and sit down in the same room with people, who in many cases have to travel in for wherever they live themselves? I don’t know. There is going to be a higher bandwidth story there, of course, and the communication is going to be marginally more effective, but is it going to be so effective that it’s worth more or less throwing a wrench into everyone’s schedule for that meeting? That leaves me somewhat unconvinced.

Amy: One of the strange things is that previously, I would go on-site to clients and fly out to where they are because as many startups as there are within Chicago, and the [unintelligible 00:13:10], and within Illinois, I’m always being sent to New York, or Atlanta, or Denver, for some reason because they’re far and there’re planes there. But we always end up having to talk to some amount of people that don’t even work in that time zone, or maybe even then in this country, so we’re talking to resources in Asia, resources in Europe, which meant we were flying people in to be on somebody else’s phone. And I’m glad to not do that. I’m glad to not have to hang out in an airport, there is a burrito place in O’Hare that I truly enjoy, but that is the one thing I miss about traveling. I almost have my punch card done and it stopped right before then, but I’m kind of okay with that.

Corey: I was chasing the brass ring of airline status and all the rest for a long time. I can’t wait to finally go and hit the next tier and the rest, and where, well, the pandemic through all of that into a jumble and I take a step back and look at it and, you know, I don’t miss it as much. What I do miss is that the opportunity, in my case, to get away from everyone that I spend all of my time with now, just for a day or two, and clear my head and recenter myself. But there are probably ways to do that that doesn’t keep me on the road for 140,000 miles a year.

Amy: I think, or at least I hope, that this will give us a chance to as an industry just reevaluate how we treat travel. A lot of clients treated it as essentially a status level that came with your engagement where we need you on-site so we can show off we have consultants coming in on-site so frequently to give us personal reports, even though we are all in a room on a conference call with other people. So hopefully, even if they’re not forcing that every other week—sometimes weekly, depending on what your engagement is—type of cadence on travel, then maybe it’ll just increase the quality of life for some of us. It would be nice. It would be super nice. I honestly don’t see them forcing that anytime soon, but once everyone gets vaccinated, and there’s a successful pediatric vaccine that comes out, it’s like, I don’t see them, just letting us stay at home and continue doing our job the way we have been for the past year, going on two years.

Corey: So, dialing back into the mists of the distant past, it’s always a question of where do cloud economists—or clouds economist, depending upon how we choose to mis-pluralize things—come from. And everyone here is a different story and there’s not a whole lot of common points between those stories. You, for example, spent some time doing work with NASA. What was that about?

Amy: There’s a lot of misconceptions about working for NASA like you need to be a doctor. [laugh]. And trust me, you don’t. I knew a lot of people who work there that they basically got their degree, and then they just did code work forever and they are lifers there. And it’s such an interesting place to be because, on one hand, you have that mission of space and exploration and trying to do better by the world, but also, it’s still a federal agency and there’s still a lot of problems with federal agencies in that how you get paid is essentially at the beginning of the year, that’s when all the budgets are done.

So, you can’t do the startup thing where you go, I’m going to try a bunch of things, and if one doesn’t work, I’m going to pivot to something else because you’re essentially answering taxpayers and they don’t let you do that. No one wants their taxes going to someone who tried a thing and then found out they messed up. Which is unfortunate, but also a really hard reality of the way these work.

Corey: I really love installing, upgrading, and fixing security agents in my cloud estate! Why do I say that? Because I sell things, because I sell things for a company that deploys an agent, there's no other reason. Because let’s face it. Agents can be a real headache. Well, now Orca Security gives you a single tool that detects basically every risk in your cloud environment -- and that’s as easy to install and maintain as a smartphone app. It is agentless, or my intro would’ve gotten me into trouble here, but it can still see deep into your AWS workloads, while guaranteeing 100% coverage. With Orca Security, there are no overlooked assets, no DevOps headaches, and believe me you will hear from those people if you cause them headaches. and no performance hits on live environments. Connect your first cloud account in minutes and see for yourself at orca.security. Thats “Orca” as in whale, “dot” security as in that things you company claims to care about but doesn’t until right after it really should have.

Corey: I must confess that I’m somewhat disappointed that you opened with, “You don’t really need to be a rocket scientist to work at NASA,” just because, honestly, I was liking the mystique of, “Oh, yeah. You need to be a rocket scientist to understand AWS billing constructs.” But I suppose if I’m being honest, that might be a slight overreach.

Amy: I knew a lot of people there who had multiple PhDs, and they could barely keep their computer on, so really, I’m finding that I respect very smart people, but it also does not imply your world intelligence anywhere else outside of that very specific field. One of the really weird things about having worked there—I worked at NASA Ames as part of their IT department—one of the things we did as outreach was, we did a booth over at SiliCon Valley Comic Con once, and it was great. We had a vintage display of old electronics, like CRT monitors, and full keyboards, and all of this nonsense, and kids would go, “These are so old. Why would anyone use this? It’s so boring.” [laugh].

And the entire IT department showed up to volunteer. We’re like, “No, you don’t understand why it’s interesting. It’s great.” It was so, so hard to watch young people just [laugh] not be interested in what the past of digital devices were. Very sad. And on the other hand, we did get a lot of interest, but it was also having to have that conversation in real life can be a little disheartening.

Corey: It really is.

Amy: But it was fun because it was one of the few things that you don’t really get to see NASA at a comic book convention, so that was actually a really cool thing to do. Also, we got free tickets, so that was great.

Corey: So, what was your background before you got to the point of, “You know what I want to do? Work for a consultancy, whose entire mascot is a platypus, and from there, go ahead and fix AWS bills,” which sounds like, to folks who aren’t steeped in it, the worst thing ever? What series of, I guess, decisions led you here?

Amy: I know you don’t remember this, but we actually met at Serverlessconf, and you opened your talk with, “I am a cloud economist, a title I completely made up.” Your talk was right before mine, so that’s why you didn’t remember I was there because I was actually on the backstage getting prepared for my talk.

Corey: That’s right. I would have been breathing into a paper bag right before or right after my talk, trying not to pass out. People say, “Oh, you won’t be nervous once you give enough talks.” I’m still waiting.

Amy: It never happens. I did finally stop having blackouts, so that’s an improvement. It gets better, but it never goes away. And when you told me that, and I saw the listing, I’m like, “I don’t know what this job is. There’s an easy way to find out what the job is, and that’s to apply.”

And that is when I started going through the process of applying, and then you hired me some months afterwards. And the thing that I found out about looking through AWS billing is that I found out I have a very specific skill set, in trying to find a discount while looking through receipts. This is a thing I thought only applied to my personal life because I don’t really like paying retail prices for anything, so I’m always looking for a way to squeeze out another 20, 30% out of something because 10% is really just taking taxes off of stuff. And the fact that I was able to apply that very specific skill set to an actual technical job is so much fun for me because I like being able to tell stories out of what people spend, just because it is—as we say around here—it’s the sum of all of your engineering decisions. Because everything you do, there’s a price tag on it. And knowing how you got there, and that you can optimize the architecture by looking through the bill is super fascinating to me.

Corey: So, now that it’s been six months, is the job what you expected it to be? Is it something radically different? Is it something else?

Amy: The tone of our engagements have actually changed within the past six months, partially because of the way AWS has made some organizational changes internally as far as we can tell. But really, it’s also what types of companies are finding out that they need a service like this. Before, when I was interviewing, and when I started, I was talking to Pete and Jesse about the types of engagements you do. It was for larger companies; they’re looking for some amount of savings, and then we run some tooling and then we get back to them. And now it’s turning into… some are relatively small companies, companies that wouldn’t get use out of an EDP, for example, because they’re spending so low.

But also other companies, they don’t even really want this saving specifically, they want validation on what the process is like, they want validation on their unit economics and what the cost allocation strategies are. So, it’s fascinating what people actually want now that they understand that that option is out there.

Corey: One thing that you mentioned a minute ago, was the idea of going and giving talks—in the before times, at Serverlessconf and things like that—how do you find that that has changed for you over the past year? And how are you viewing a slow, but effectively guaranteed in the long enough timeframe, return to normalcy?

Amy: So, when everything went virtual, it was a really hard transition for a lot of communities. I’m part of AWS Chicago, I’m an organizer at a meetup group called Write/Speak/Code where our whole deal was to give women and people of underrepresented genders the opportunity to learn how to do open-source, how do you learn how to do technical speaking, we helped them with their CFPs, and all of that required a physical community in order for us to be able to give each other that type of support. Well, we can’t do that anymore. The last big event we did was specifically around getting everyone set up into the organization. So, we do one big one every year where we tell you how to do slides, we tell you how to do everything, and then the rest of the year smaller meetups where we do feedback and prompts and other types of
support events.

Now, that we can’t do that, we found that a lot of the people who ran these events, they’re extroverts and they’re social to begin with, so they
got burnt out very quickly, which meant we not only had to find new ways of supporting them but also reaching out to members and making sure everyone could still do the types of thing we promised we’d help them to do. The other issue is that with things going virtual, there used to be clear lines on the types of events you could apply to, which were events that I could reach, that my company would pay for, what have you. But now everything’s virtual, I accidentally applied to a meetup that I spoke at that was based in Australia, which meant I was talking at 10 o’clock at night, just because I explained to them the mix-up and essentially begged to get an earlier slot. And it’s interesting because it presents both a wide array of opportunities, but it also means that there’s now so much noise, and so much burnout and fatigue from these one direction types of conferences, which are long Zoom meetings, which is basically everyone’s workday now, which is just full day’s worth of Zoom meetings. And it’s hard to get people interested.

And really, what these events did best was give people who generally don’t have that type of visibility—like me—who, I am a female engineer, and I am a person of color, so my opportunities aren’t always going to be as well as they could be, but this visibility also gave me a boost. So, when we lost out on physical events, my organization personally lost out on a lot of things that we could do for those people, which is unfortunate, but it’s also what was happening, really, everywhere. Now, we’re gearing more towards how to do less Zoom meeting type of events, and we’re now using a tool called gather.town, which lets everyone go into a space, you can walk around, you can drop in and out of conversations like you would in a hallway, and it has this cute little eight-bit kind of avatar feel to it so it looks kind of like a game, but it’s also—if you go around a large group of people, it pulls up everyone’s picture and then you’re suddenly speaking to each other over voice without having to physically join in or wait for a breakout room or wait to be let into a different room. So, it’s been difficult to try to find ways of managing it, but it’s also been very interesting seeing the tools that had come out of this.

Now, once we start going back to physical events and on-site events, personally, I have a lot of anxiety about that just because my allergies and everything makes it hard for me to do travel in the first place, but also, I’m not really sure how everyone’s going to react being in the same space anymore, or how full they’re going to be. So, to me, it does cause me a little bit of anxiety and I’m going to wait a little until things are more settled and are a little more stable than they are right now. That said, I believe AWS Midwest is doing something physical at the end of this year. But I’m not entirely positive about that.

Corey: One of the things I want to avoid is going back to an old style of re:Invent, where it’s only open to folks who are able to spend $2,000 on a ticket, travel to Las Vegas, and get the time off and afford the hotel stay and the rest. One thing I loved about the 2020 version of that was that everyone was coming from a baseline of its full remote. There was no VIP ticket option that got you a better experience than anyone else had. And I do worry, on some level, that as soon as they can, they’re going to go running back to a story like that. I hope not.

Amy: I hope not, too, because one of the great things about virtual re:Invent was everyone saw the same things at the same time. And you didn’t have to give up one panel because you’re too busy being in line for another panel. And I’ve never liked that type of convention activity where you know you’re actively giving something up just because the line for the thing you want to get to is so super long. That said, they could have done better a bit on the website. The website was really hard to navigate and confusing and the schedule was weird.

If they get that kind of usability fixed and a little more reasonable so that things are surprisingly searchable and easier to navigate, and they have more social type of events instead of AWS is going to—it’s going to announce another service that they’re calling a product, even though it’s just… it’s a service enhancement, and that’s all it is. I would like for all of that to kind of be streamlined so it’s not just more announcements. I can read announcements. I don’t need to watch two hours’ worth of announcements.

Corey: I really hope that at some point, some of the AWS service teams learn that, “Hey. We could announce services anytime of the year,” rather than in a three-week sprint that leaves no one able to pay attention because there’s just too much.

Amy: And it’s hard because it’s hard when you have to hold a release because re:Invent’s coming up. Or you have to be part of their Developer Relations Group where you have to do all of your training and all of your docs right beforehand, just so that it’s prepared for this launch that people may not hear about because it gets drowned out in all the other noise.

Corey: And that’s sometimes part of the entire problem.

Amy: Yeah.

Corey: Thank you so much for taking the time to speak with me. As always, it’s a pleasure. If people want to learn more about what you’re up to, where can they find you?

Amy: They can hire me for engagement here. [laugh] but also—

Corey: Good answer.

Amy: —I do technical talks. I don’t have anything lined up right now, just because it’s spring in my brain took a break, like everyone else did. I’m on Twitter as @nerdypaws because that was a handle I had since college and have not changed.

Corey: Excellent. And we will, of course, leave a link to that in the [show notes 00:31:10], as we always do. Thank you so much for taking the time to speak with me. It’s always a pleasure, and it’s deeply appreciated.

Amy: It’s always a good time talking to you, Corey.

Corey: Amy Arumbulo Negrette, cloud economist here at The Duckbill Group. I am Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with a comment telling me why you do in fact need to be a rocket surgeon in order to properly work on AWS bills.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

This has been a HumblePod production. Stay humble.

View Details

About Jesse

Jesse is a seasoned operations engineer with a deep passion for understanding complex technical and organizational systems. He's spent his career helping Engineering teams achieve their business goals by improving how they interact with their technical systems, and with each other. He's currently a Cloud Economist with Duckbill Group, guiding organizations along their journey of cloud cost optimization and management.

Links:

  • The Duckbill Group: https://www.duckbillgroup.com/
  • Jesse’s Twitter: https://twitter.com/jesse_derose
  • AWS Morning Brief: https://www.lastweekinaws.com/podcast/aws-morning-brief/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Your company might be stuck in the middle of a DevOps revolution without even realizing it. Lucky you! Does your company culture discourage risk? Are you willing to admit it? Does your team have clear responsibilities? Depends on who you ask. Are you struggling to get buy in on DevOps practices? Well, download the 2021 State of DevOps report brought to you annually by Puppet since 2011 to explore the trends and blockers keeping evolution firms stuck in the middle of their DevOps evolution. Because they fail to evolve or die like dinosaurs. The significance of organizational buy in, and oh it is significant indeed, and why team identities and interaction models matter. Not to mention weither the use of automation and the cloud translate to DevOps success. All that and more awaits you. Visit: www.puppet.com to download your copy of the report now!

Corey: Up next we’ve got the latest hits from Veem. Its climbing charts everywhere and soon its going to climb right into your heart. Here it is!

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’m joined this week by Jesse DeRose, my colleague, and cloud economist at The Duckbill Group. Jesse, thank you for joining me, even though when I asked you, it isn’t exactly like you felt you had much of a choice.

Jesse: [laugh]. I appreciate being on this podcast with you. I think you’ve had an opportunity to talk to a few other folks from our organization so far, so I’m just happy to be included. I’m hoping that I get the Members Only jacket after this recording.

Corey: Oh, absolutely. The swag that goes out to guests is secret and wonderful, all at the same time. So, let’s start at the very beginning. It turns out that despite the prevalent narrative that we put out there, are clouds’ economist did not just spring fully-formed from the forehead of some God, and appear—fully-formed—ready to slice AWS bills to ribbons. There’s a process and there’s always an origin story. Where do you come from? Where were you before this?

Jesse: So, my background is mixed. I started with a management of information systems degree, which is inherently interdisciplinary. You’re linking, sort of, the importance of both technical systems and business goals and outcomes into a single degree. And when I graduated with this degree, I went out into the working world and said, “Okay, I’m here. I’m ready. Here’s my degree. Here’s what I’m interested in. Let’s do this.”

And nobody knew what to do with me, Corey. Nobody knew what management of information systems was. They simultaneously said, “You know, you don’t have a computer science or a computer engineering degree, so we don’t believe that you’ve got programmer chops.” Especially since this is the golden age of boot camps and quickstart programming books, and this movement that anybody can be a developer, which makes it even stranger that I’m not spending my weekends learning every programming language under the sun and automating the little tasks.

Corey: Oh, absolutely. In fact, I got the exact same feedback; I don’t have a degree, and, “Oh, you couldn’t possibly be a good programmer because you don’t have the degree, you didn’t go to the right schools, you didn’t go to a boot camp. And when I asked you to write some sample code to demonstrate how to program, you just started crying instead.” Because yeah, it turns out, I’m actually not a good programmer, I just, sort of, brute force my way through it. And so far, so good.

Unfortunately, people aren’t paying me to program these days, for excellent reason. In fact, I strongly suspect some people are paying me not to. But yeah, now it’s funny to laugh about, but back when you’re getting started, and going out in that space, in the world of operations, which we both came up in, and looking at it through a lens of the SRE movement, suddenly, “Hey, you used to be really, really good at all these Linux things and working on systems and keeping them up. Great. Now, learn to code.” And that was a big lift, at least for me.

Jesse: Absolutely. I struggled with that so much because every company that I interviewed with, every company that I even talked to, just assumed that I had some kind of programming experience and didn’t want to talk to me if I didn’t have programming experience. And to me, looking back, I think I really look at it like I may not have the programming chops to be the software engineer that is going to write the code for you, day in and day out, but I am the person who knows enough that I can have the conversation with the software engineers. I can be the SRE, I can be the ops person that has the conversation with your software engineers and knows the things that they’re talking about, but also knows when to get out of their way and let them be the expert and do what they do best.

Corey: There’s something to be said for valuing expertise in areas that are—how to put this—not the thing that you think you’re looking for. I mean, back when I was getting jobs, before The Duckbill Group and I would be the first ops hire into a team of developers—which happened a few times—the process was always the same, where you’d have a bunch of developers asking what they thought were ops questions, or just giving up on that entirely and trying to figure out how decent have an ops person I would be by how badly I programmed.

Or, “Oh, okay, cool. You’re an ops person. Great. Can you invert a binary tree on a whiteboard?” It’s, “No, but I can invert a rack in your data center. I’ll go rage-flip the rack. Why not?” And it takes time. You have to guide those interviews and those conversations. But it’s always weird.] interviews are always weird because you’re being judged on a skill set that only matters when you’re interviewing for a job.

Jesse: Yes. This is one of the things that I struggled with the most because I knew that most of the people who were interviewing me were either business people themselves—so they assumed that because their engineering team thought a certain way and acted a certain way that I should act that way—specifically, too—the same way as everybody else on the team. Or they were software engineers themselves, so they said, “Okay, I know how to invert a binary tree”—to your point, Corey—“So, do how to invert a binary tree. If you know how to do that, then sure, you can be part of my team because you know how to do these things and think about these things the same way I do.” Whereas because I was coming from this operations space, I knew other things that were equally as important but weren’t part of the conversations
that they were used to having, day-to-day.

So, they didn’t understand that just because I didn’t have the same engineering chops as them, that I didn’t have important information to share and wasn’t able to stand on my own two feet in other ways. And that was one of the things that I really struggled with when I was starting out in the industry because I was thinking to myself, well I have such passion to be part of these conversations, to have that conversation between the business side of the organization and the engineering side of the organization, from an ops perspective, from a business perspective, from a technical perspective, and if I can’t convince these people of my own volition, of my own passion for the good of the company, maybe data will help. Maybe there’s something that I can find from a scientific research perspective. Maybe there’s something that—I’m sure somebody else has already researched this topic or found the same problems that I have in this space, and maybe they’re already talking about it, and maybe I can ride their coattails, so to speak, or follow in their footsteps and use the information that other people are talking about the industry to help me not just land these jobs, but ultimately better sell myself and help these companies move forward.

Corey: That is fundamentally an encapsulation of what I believe the ops role to be of, make things better, and move them forward. But man, do we get stuck in an awful lot of weird and strange places. And interviewing itself is a skill. Giving an interview, very often, it’s a, “I know a bunch of things that are trivia, but I know them. And if I know them, everyone must know them, therefore, if you don’t know them, you must be bad at things.” And it turns out—for better or worse—being able to memorize the documentation and spit out answers is not indicative of whether someone is a good ops person or a terrible one.

Jesse: Absolutely. I think that is one of the biggest problems that I have faced and one of the biggest problems in the interviewing space today because it’s not just about, can you regurgitate this information, but it’s about how do you think? How do you look at problems? How do you communicate to the rest of your team, and within the rest of the organization? Those technical skills are ultimately important because you do need to understand some amount of technical information to have those conversations, but the soft skills are also super, super important to be able to communicate effectively, to be able to think collaboratively, and help everybody, not just yourself, but help the team build that shared purpose and move forward together.

Corey: We’re talking right now, so far, about traditional ops roles. Then we have what we do here, which is beyond the rest of all of that, where all of what we just said is necessary but not sufficient. Then it comes down to great, okay. So, you understand how systems work together; we found that, for what we do and how we do it, you need to be a competent ops person as a fundamental tenet.

Otherwise, learning what all these AWS services do will occupy you for the next three years. So okay, we start off there. Then on top of that is, okay, there are consulting skills that it turns out are possible to teach, but incredibly challenging and time-consuming because a lot of them boil down to, can you be in a meeting with stakeholders of various levels? Can you deliver bad news in a way that they don’t hate you? Because they don’t really want to pay you just say yes to whatever they think.

And can you do that in such a way without, you know, actively insulting them, which sounds like a strange thing until you realize, oh, wait, that’s right. I do that, too. So ooh, yeah. Corey is going to have that problem, isn’t he? Yeah. And that’s part of the beautiful part about this place is that finally, I’m able to hire people like you.

You were the first cloud economist here, which meant suddenly I didn’t have to do it all myself and my mouth slowly stopped getting me in trouble in consulting engagements, so I could spend more time having my mouth getting me in trouble on Twitter.

Jesse: Yeah, I have to tell you, Corey, when I originally spoke with you and Mike about this role, I had just taken another operations position with a tech startup, and I was about two months into the role. And Mike sat down with me for coffee one morning and said, “Hey, we’re thinking about doing this thing. Are you interested?” And I said, “Yes, but I just started this other operations gig. I can’t up and leave them; I really care about the team, I really care about the company. And it would look really bad on me if I just, you know, two months left.”

Because—unless they were a really, really awful employer, which they weren’t. So, I said, “Sure. I’m interested in doing some kind of part-time work.” And that’s ultimately where I started with you and your business partner, Mike. And I have to admit to when Mike originally approached me and said, “Hey, this is what we are thinking about; this is what we’re doing,” I didn’t really think twice about the opportunity because I wanted to work with you and Mike again.

But the way that Mike described the work, just didn’t stick with me. It didn’t resonate with me, it was more about, “Hey, I would love to work with Mike and Corey again,” than, “Oh, my God. This sounds like the dream role that I want to be a part of.” And then, when I came back to Mike, probably, I don’t know, a month or two later, after I had started working part-time with both of you, I said, you know, “Mike, I don’t think I really made myself clear. I want to make sure that I help you understand, ultimately, the things that I want to do are having these conversations, being that bridge with the business side, and being able to talk tech with the tech side, and being able to talk business, and make sure that both sides of the conversation are aligned.”

And he just looked at me and said, “What do you think we’re doing? What do you think we sold you on?” And it was that aha moment where I thought, “Oh, my god, yes.” I had already said yes; I was already working with both of you part-time, but that was the moment that really solidified it for me of, this is what I’ve kind of been moving towards. I’ve been wanting to be that person that can speak both languages and have a conversation with both sides of the table, and speak to multiple different audiences, and now I’m finally getting the chance to do that. I’m getting the chance to grow both skill sets, which I think is extremely rare in a lot of the smaller tech spaces that we see today.

Corey: You’ve hit on one of the secrets of The Duckbill Group if I can be so grandiose as to claim that. And it’s true because we take a look at people that we bring in, and things that they’re good at, and things that we do—the things we do publicly and the things that we do, sort of, behind the scenes and there’s no reason we don’t talk about them publicly, but there’s no real reason for us to do so. Easy example, and what I want to talk to you about next is, you are deep into improving understanding of complex systems, both technical—okay, great people expect that—and organizational, which sometimes throws people for a loop. And it sounds like a weird thing to focus on here because we fix the AWS bill. We do not bill ourselves as management consultants, we do not bill ourselves as coming in and we will restructure your organization because that sounds patently ridiculous, and no one in their right mind is going to buy that thing.

I wouldn’t buy it, at least not for me. My God, there are large consultancies that specialize in these sorts of things. I don’t know how they do it because I certainly don’t. We’re not here to sell that, though. Fixing the AWS bill—I mean actually fixing it. Fixing the business problem tied to it mandates an understanding of those complex systems. And your expertise and interest in that area is incredibly helpful here. Tell me more about it.

Jesse: This is one of the things that I’ve been really fascinated by ever since I joined Duckbill Group. I think everybody in Duckbill Group has a superpower or has a really passionate hobby to some extent, which makes each of us really interesting, unique individuals that can focus in different areas of a client’s bill or a client’s pain points when it comes to cloud cost management and help, and find the parts that are frustrating, find the levers that can be moved, and point them out and say, “Okay, this is ultimately where you want to pull this lever or not pull this lever to make these changes.” And to your point, Corey, the one that is most interesting and passionate for me is that organizational development space. It is really understanding, not just the small things that we can do today to help you save money on your AWS bill, but how we, and collectively how our clients can think about costs long term to save money on their AWS bill. And I know that sounds really, really broad, and that’s part of why I think that there is a lot of nuance in this space, to your point about other organizations or other vendors that are providing these consulting services, and I think is also something that is also difficult to sell, which is why it’s not our expertise in terms of what we are on the cover trying to sell to any of our clients.

But I definitely think that there is opportunity to have some of those conversations within each of our clients’ spaces to talk about some of the pain points that we see that may ultimately lead to better cost management practices long term, things that ultimately might help the engineering teams communicate better with finance on a long term basis, help the finance team and any of the leadership team more collaboratively talk with the engineering teams about understanding how much money is the product costing us? How much can we continue to spend on this product? Or how much can we discount one of our products for our customers before we are losing shares, losing money? What are the fine lines that we understand, based on how much money we’re ultimately spending on these features, on these products, that will help us make better data-driven decisions about other parts of the company?

Corey: I really love installing, upgrading, and fixing security agents in my cloud estate! Why do I say that? Because I sell things, because I sell things for a company that deploys an agent, there's no other reason. Because let’s face it. Agents can be a real headache. Well, now Orca Security gives you a single tool that detects basically every risk in your cloud environment -- and that’s as easy to install and maintain as a smartphone app. It is agentless, or my intro would’ve gotten me into trouble here, but it can still see deep into your AWS workloads, while guaranteeing 100% coverage. With Orca Security, there are no overlooked assets, no DevOps headaches, and believe me you will hear from those people if you cause them headaches. and no performance hits on live environments. Connect your first cloud account in minutes and see for yourself at orca.security. Thats “Orca” as in whale, “dot” security as in that things you company claims to care about but doesn’t until right after it really should have.

Corey: And I want to call out that this is something that we are comparatively enthusiastic amateurs around. It’s valuable; it’s important; it’s an awful lot of deep work, but I’m not sure that we go more than three working days without referencing Dr. Nicole Forsgren’s work, internally, as we think about these things. So, if you’re hearing this, and you think that okay, AWS bill, fine, whatever. We really want to talk about organizational challenge and improvement, oh, my God, talk to Dr. Forsgren. Holy crap. She’s been on this show, at least I think, three times now, and every time I feel like I’m lucky to get her. Most weeks, you know, I’m stuck with people like you. My God, Jesse.

Jesse: [laugh].

Corey: But no, her work is seminal in this space. And in seriousness, every time I start to question the value of expertise, I look at how deep she goes on all of these things and the level both of understanding that’s baked into this, and the amount of sheer work that it takes for her to take all of that very deep, penetrating analysis, and make it accessible and understandable. But every time I look at her work, I come away more impressed than I started, and that wasn’t a low bar, to begin with.

Jesse: Yes. And this gets back to my earlier comment about, I just want to be here to help. And in a lot of cases, when I was starting out, folks didn’t know what to do with me because I didn’t have data, I didn’t have any information. But we have folks in the industry like Dr. Nicole Forsgren, and other folks who are doing the research, who are knowledgeable in this space, who are putting in the effort to run these studies, to analyze this data, to share the results—and to your comment, Corey—to share the results in a way that makes sense to everybody, that’s easy to read, it’s approachable, it’s understandable.

And I am so thankful to have folks in the industry who are doing that work because that is not my expertise. But that means that I get to say to the folks that I’m working with—internally and with our clients—“Hey, don’t take my word for it. There are other folks who have done research, and here’s what the data says, and here’s how we can help you apply this work, or you can apply this work within your own organization.”

Corey: And this is the challenge in some cases, too, where there’s a lot of organizational theory, and that is being advanced heavily, and in ways that makes teams more effective is super helpful. The challenge, of course, is that sitting here and talking about the theoretical layout of teams and how to improve functioning as an organization is all well and good, but we’re brought in by our clients to help them with their AWS billing situation, so at some point the conversation has to evolve beyond, “Okay, so here’s what you could do in theory, in a vacuum, assuming spherical cows, et cetera, et cetera.” And their response is, “That’s great. You actually going to fix the bill or just pontificate for a while here?” So, for better or worse, we don’t really get to sit there and have deep organizational conversations at length with our clients, just because that’s not the problem we’re there to solve. Everyone’s busy and we want to make sure that we’re respectful of their time.

Jesse: Yeah, one of the things that I’ve learned through my time with Duckbill Group and through other similar roles in the past, is that I may have a strong passion, I may have this strong guiding light in my head, but it’s not the same guiding light that our clients or our customers have. And that’s fine because we don’t need to necessarily have the same goal in sight, but that means that I, to best serve our clients or best serve our customers, need to make sure that I am aligning, that Duckbill Group is aligning with the clients that we’re working with, with the organizations that we’re working with. So, I may be thinking to myself, “Oh, my gosh, I would love to come in with this long list of organizational development practices and share a million different things,” but that’s not ultimately what they need. Maybe that’s something that they’re going to want long-term, but it’s not what we’re here for today. And it’s more important to help serve the client that is in front of me that is asking for things now, today, than try to educate them on quote-unquote, “What their problem is,” and then sell them on a solution.

Corey: The thing that I think gets lost as well, whenever I start talking in-depth about what we do on Twitter, for example, is generally from other engineers whose response is, “Okay, yeah, sounds great. You come in and say that it will save a bunch of money before you rearchitect our application. But that’s an awful lot of engineering time, so I bet engineers most hate you.” And the honest answer is, “Yeah. We know that you would save some significant money if you rearchitected your application, but we also pay attention to organizational dynamics, and we know you’re not going to do it because there’s no business justification for doing it. So, we’re not going to bring it up, other than, possibly, in passing, just so people are aware of the relative benefit if they want to bake that in.”

But we don’t go in and suggest nonsense that is abhorrent. We all started as engineers ourselves. We are sensitive to engineering time, both in terms of what engineers enjoy working on—which is more important than people think—as well as the sheer cost of engineering time. People get concerned about the AWS bill, but it’s invariably pale compared to the cost of the people working on the AWS infrastructure. You want to optimize the right things, and then you want to get back to doing what your company does, not continue to iterate forward and spend thousands of dollars to save tens of dollars.

Jesse: Yeah, it’s an extremely difficult balance to find. And it’s really important to think about, is the ROI on this change worth the change? How much money am I ultimately going to save for the amount of engineering effort that I’m going to put in? Because we don’t want to run in and tell your engineering teams, “Rearchitect everything,” if it’s going to maybe save them tens of dollars. We want to make it very clear that here are the different levers that you can pull to affect change within your AWS bill, to optimize your AWS bill.

It’s up to you which ones you want to pull. We are just giving you that guiding path, and then you have the option to say, “Yes, I want to spend the engineering effort to get this kind of ROI.” Or, “You know what? I don’t think that’s the priority for us right now.” One of the things that I’ve noticed with a number of our clients is that balance of, do we focus more time on building new features which brings in new customers, or do we spend time on the existing infrastructure and making changes to the existing workloads that we have?

And it’s this delicate balance of internal work—the things that you have put on the backburner over time that you ultimately say, “Well, I’ll come back to that,” versus the new things that are ultimately going to bring in new customers, bring in new users, potentially bring in more money. Because both are important, but there needs to be a delicate balance of both, and I feel like that’s one of the biggest challenges that we face when managing an AWS bill and trying to optimize, and organize, and better manage cloud costs.

Corey: And that’s fundamentally what it comes down to what is the best outcome for the client? And the answer to ‘what is best?’ Varies based upon their constraints, what they’re focused on, what they’re trying to achieve. You could look at the end result of one of our analyses for a customer, and take issue with it in a vacuum of but they’re spending all this money on things that they shouldn’t be doing, or don’t need to be. Why didn’t you suggest this, or this, or this, or this?

And the answer is because, based upon our conversations, we knew that they weren’t going to do it, and suggesting things that we know they’re not going to do is one of the best ways to erode trust. We’re there to deliver an outcome. We’re not naive software that is just running a pile of tools on an Amazon bill and saying, “Here you go. Have fun.” We’re not the billing equivalent of a Nessus scan that someone slaps their logo on top of, drops off, then it’s 700 pages long. “Have a good one. Check, please.” It doesn’t work. Not well, anyway. It doesn’t drive to lasting change.

Jesse: Yeah. And I think that’s ultimately part of where we come in best because we ultimately want to be that bridge between the engineering teams, and finance, and leadership, and essentially the business side of the business. We want to give both sides the information that they need to be able to speak effectively and collaboratively with the other side. We want to make sure that the finance side understands enough tech that they can work with the engineering teams to guide them in terms of what goals are important for finance, in terms of managing budget, in terms of forecasting spend. And then from the flip side, we want to make sure that engineering understands that the business collaboratively wants to manage these things, and also help build the organization, and engineering has great, great potential to do little things, take little steps to help the organization get the data that they need to make these decisions.

And that scratches that itch for me. That scratches that itch of how can I really be both sides of the conversation? How can I flex the business lingo and also flex the tech lingo, maybe not in the same conversation but with the same client over time, to really help both sides understand that both sides are important, and both sides need to understand each other in order to help the business succeed?

Corey: And that’s ultimately what it comes down to. Jesse, thanks for taking the time to speak with me. If people want to hear more about what you have to say, and how you like to say it, where can they find you, other than go into The Duckbill Group and get me a consulting engagement underway?

Jesse: So, there’s two places that folks can find me. The first is on Twitter at @jesse_derose. And then the second is our other podcast, the AWS Morning Brief podcasts, on the mornings where we don’t get to hear your melodious voice, Corey.

My colleague, Pete Cheslock and I are talking about all of the interesting things that we’ve seen on AWS, from client-specific situations to things that we’ve seen on the job in previous organizations that we use to work at, to new features that AWS is releasing. There’s a whole slew of interesting things that we get to talk about from a more practical perspective. Because there’s so many releases coming out day-to-day that you’re obviously helping us stay on top of, Corey, but there’s so many other things that we want to make sure that we are talking about from the real-world applicable perspective.

Corey: And we will, of course, put links to these things in the [show notes 00:27:48], but you should already be aware of most of them, at
least the ones that are on the company side. Jesse, thank you so much for taking the time to speak with me.

Jesse: Thank you for having me.

Corey: Jesse DeRose, cloud economist here at The Duckbill Group. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an insulting comment telling me that organizational dynamics really aren’t that hard and you could solve it better than Dr. Forsgren does, in a weekend.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

This has been a HumblePod production. Stay humble.

View Details

About James

James is the Redmonk co-founder, sunshine in a bag, industry analyst loves developers, "motivating in a surreal kind of way". Came up with "progressive delivery". He/Him

Links:

  • RedMonk: https://redmonk.com/
  • Twitter: https://twitter.com/MonkChips
  • Monktoberfest: https://monktoberfest.com/
  • Monki Gras: https://monkigras.com/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Your company might be stuck in the middle of a DevOps revolution without even realizing it. Lucky you! Does your company culture discourage risk? Are you willing to admit it? Does your team have clear responsibilities? Depends on who you ask. Are you struggling to get buy in on DevOps practices? Well, download the 2021 State of DevOps report brought to you annually by Puppet since 2011 to explore the trends and blockers keeping evolution firms stuck in the middle of their DevOps evolution. Because they fail to evolve or die like dinosaurs. The significance of organizational buy in, and oh it is significant indeed, and why team identities and interaction models matter. Not to mention weither the use of automation and the cloud translate to DevOps success. All that and more awaits you. Visit: www.puppet.com to download your copy of the report now!

Corey: And now for something completely different!

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’m joined this week by James Governor, analyst and co-founder of a boutique analysis shop called RedMonk. James, thank you for coming on the show.

James: Oh, it’s my pleasure. Corey.

Corey: I’ve more or less had to continue pestering you with invites onto this for years because it’s a high bar, but you are absolutely one of my favorite people in tech for a variety of reasons that I’m sure we’re going to get into. But first, let’s let you tell the story. What is it you’d say it is that you do here?

James: We—industry analysts; we’re a research firm, as you said. I think we do things slightly differently. RedMonk has a very strong opinion about how the industry works. And so whilst there are plenty of research firms that look at the industry, and technology adoption, and process adoption through the lens of the purchaser, RedMonk focuses on it through the lens of the practitioner: the developer, the SRE, the people that are really doing the engineering. And so, historically IT was a top-down function: it required a lot of permission; it was something that was slow, you would make a request, you might get some resources six to nine months later, and they were probably the resources that you didn’t actually want, but something that was purchased from somebody that was particularly good at selling things.

Corey: Yes. And the thing that you were purchasing was aimed at people who are particularly good at buying things, but not using the things.

James: Exactly right. And so I think that RedMonk we look at the world—the new world, which is based on the fact there’s open-source software, there’s cloud-based software, there are platforms like GitHub. So, there’s all of this knowledge out there, and increasingly—it’s not a permission-free world. But technology adoption is more strongly influenced than ever by developers. That’s what RedMonk understands; that’s what makes us tick; that’s what excites us. What are the decisions that developers are making? When and why? And how can we tap into that knowledge to help everyone become more effective?

Corey: RedMonk is one of those companies that is so rare, it may as well not count when you do a survey of a landscape. We’ve touched on that before on the show. In 2019, we had your colleague, Rachel Stevens on the show; in 2020, we had your business partner Stephen O’Grady on, and in 2021 we have you. Apparently, you’re doling out staff at the rate of one a year. That’s okay; I will outlast your expansion plans.

James: Yeah, I think you probably will. One thing that RedMonk is not good at doing is growing, which may go to some of the uniqueness that you’re talking about. We do what we do very well, but we definitely still haven’t worked out what we’re going to be when we grow up.

Corey: I will admit that every time I see a RedMonk blog post that comes across my desk, I don’t even need to click on it anymore; I don’t need to read the thing because I already get that sinking feeling, because I know without even glancing at it, I’m going to read this and it’s going to be depressing because I’m going to wish I had written it instead because the points are always so pitch-perfect. And it feels like the thing that I struggle to articulate on the best of days, you folks—across the board—just wind up putting out almost effortlessly. Or at least that’s how it seems from the outside.

James: I think Stephen does that.

Corey: It’s funny; it’s what he said about you.

James: I like to sell his ideas, sell his work. He’s the brains and the talent of the operation in terms of co-founders. Kelly and Rachel are both incredibly smart people, and yeah, they definitely do a fantastic job of writing with clarity, and getting ideas across by stuff just tends to be sort of jumbled up. I do my best, but certainly, those fully formed, ‘I wish I had written that’ pieces, they come from my colleagues. So, thank you very much for that praise of them.

Corey: One of the central tenets that RedMonk has always believed and espoused is that developers are kingmakers, to use the term—and I steal that term, of course, from your co-founder’s book, The New Kingmakers, which, from my read, was talking about developers. That makes a lot of sense for a lot of tools that see bottom-up adoption, but in a world of cloud, where you’re seeing massive deals get signed, I don’t know too many developers out there who can sign a 50 million dollar cloud services contract more than once because they get fired the first time they outstrip their authority. Do you think that that model is changing?

James: So, ‘new kingmakers’ is quite a gendered term, and I have been asked to reconsider its use because, I mean, I don’t know whether it should be ‘new monarchmakers?’ That aside, developers are a fundamentally influential constituency. It’s important, I think, to say that they themselves are not necessarily the monarchs; they are not the ones sitting in Buckingham Palace [laugh] or whatever, but they are influences. And it’s important to understand the difference between influence and purchase. You’re absolutely right, Corey, the cloud is becoming more, like traditional IT.

Something I noticed with your good friends at GCP, this was shortly after the article came out that they were going to cut bait if they didn’t get to number two after whatever period of time it was, they then went intentionally inside a bunch of 10-year deals with massive enterprises, I guess, to make it clear that they are in it for the long haul. But yeah, were developers making that decision? No. On the other hand, we don’t talk to any organizations that are good at creating digital products and services—and increasingly, that’s something that pretty much everybody needs to do—that do not pay a lot more attention to the needs and desires of their developers. They are reshoring, they are not outsourcing everything, they want developers that are close to the business, that understand the business, and they’re investing heavily in those people.

And rather than seeing them as, sort of, oh, we’re going to get the cheapest possible people we can that have some Java skills and hope that these applications aren’t crap. It may not be Netflix, “Hey, we’re going to pay above market rate,” but it’s certainly what do they want? What tools do they want to use? How can we help them become more effective? And so yeah, you might sign a really big deal, but you still want to be thinking, “Hang on a minute, what are the skills that people have? What is going to make them happy? What do they know? Because if they aren’t productive, if they aren’t happy, we may lose them, and they are very, very important talent.”

So, they may not be the people with 50 million dollars in budget, but their opinion is indeed important. And I think that RedMonk is not saying there is no such thing as top-down purchasing anymore. What we are saying is that you need to be serving the needs of this very important constituency, and they will make you more productive. The happier they are, the more flow they can have, the more creative they can be with the tools at hand, the better the business outcomes are going to be. So, it’s really about having a mindset and an organizational structure that enables you to become more effective by better serving the needs of developers, frankly. It used to just be the only tech companies had to care about that, but now everybody does. I mean, if we look at, whoever it is: Lego, or Capital One, or Branch, the new insurance company—I love Branch, by the way. I mean—

Corey: Yeah. They’re fantastic people, I love working with them. I wish I got to spend more time talking with them. So far, all I can do is drag them on to the podcast and argue on Twitter, but one of these days, one of these days, they’re going to have an AWS bill bigger than 50 cents a month, and then, oh, then I’ve got them.

James: There you go. But I think that the thing of him intentionally saying we’re not going to set up—I mean, are they in Columbus, I think?

Corey: They are. The greater Ohio region, yes.

James: Yes. And Joe is all about, we need tools that juniors can be effective with, and we need to satisfy the needs of those juniors so they can be productive in driving our business forward. Juniors is already—and perhaps as a bad term, but new entrants into the industry, and how can we support them where they are, but also help them gain new skills to become more effective? And I just think it’s about a different posture, and I think they’re a great example because not everybody is south of Market, able to pay 350 grand a year plus stock options. That’s just not realistic for most businesses. So, it is important to think about developers and their needs, the skills they learned, if they’re from a non-traditional background, what are those skills? How can we support them and become more effective?

Corey: That’s really what it comes down to. We’re all trying to do more with less, but rather than trying to work twice as hard, how to become more effective with the time we have and still go home in time for dinner every day?

James: Definitely. I have to say, I mean, 2020 sucked in lots of ways, but not missing a single meal with my family definitely was not one of
them.

Corey: Yeah. There are certain things I’m willing to trade and certain things I’m not. And honestly, family time is one of them. So, I met you—I don’t even recall what year—because what is even time anymore in this pandemic era?—where we sat down and grabbed a drink, I want to say it was at Google Cloud Next—the conference that Google does every year about their cloud—not that Google loses interest in things, but even their conference is called ‘Next’—but I didn’t know what to expect when I sat down and spoke with you, and I got the sense you had no idea what to make of me back then because I was basically what I am now, only less fully formed. I was obnoxious on Twitter, I had barely coherent thoughts that I could periodically hurl into the abyss and see if they resonated, but stands out is one of the seminal grabbing a drink with someone moments in the course of my career.

James: Well, I mean, fledgling Corey was pretty close to where he is now. But yeah, you bring something unique to the table. And I didn’t totally know what to expect; I knew there would be snark. But yeah, it was certainly a pleasure to meet you, and I think that whenever I meet someone, I’m always interested in if there is any way I can help them. And it was nice because you’re clearly a talented fellow and everything else, but it was like, are there some areas where I might be able to help? I mean, I think that’s a good position as a human meeting another human. And yeah, it was a pleasure. I think it was in the Intercontinental, I guess, in [unintelligible 00:11:00].

Corey: Yes, that’s exactly where it was. Good memory. In fact, I can tell you the date: it was April 11 of 2019. And I know that because right after we finished having a drink, you tweeted out a GIF of Snow White carving a pie, saying, “QuinnyPig is an industry analyst.” And the first time I saw that, it was, “I thought he liked me. Why on earth would he insult me that way?”

But it turned into something where when you have loud angry opinions, if you call yourself an analyst, suddenly people know what to do with you. I’m not kidding, I had that tweet laser engraved on a piece of wood through Laser Tweets. It is sitting on my shelf right now, which is how I know the date because it’s the closest thing I have to a credential in almost anything that I do. So, congratulations, you’re the accrediting university. Good job.

James: [laugh]. I credentialed you. How about that?

Corey: It’s true, though. It didn’t occur to me that analysts were a real thing. I didn’t know what it was, and that’s part of what we talked about at lunch, where it seemed that every time I tried to articulate what I do, people got confused. Analyst is not that far removed from an awful lot of what I do. And as I started going to analyst events, and catching up with other analysts—you know, the real kind of analyst, I would say, “I feel like a fake analyst. I have no idea what I’m actually doing.”

And they said, “You are an analyst. Welcome to the club. We meet at the bar.” It turns out, no one really knows what is going on, fully, in this zany industry, and I feel like that the thing that we all bond over on some level is the sense of, we each only see a piece of it, and we try and piece it together with our understanding of the world and ideally try and make some sense out of it. At least, that’s my off-the-cuff definition of an industry analyst. As someone who’s an actual industry analyst, and not just a pretend one on Twitter, what’s your take on the subject?

James: Well, it’s a remarkable privilege, and it’s interesting because it is an uncredentialed job. Anybody can be, theoretically at least, an industry analyst. If people say you are and think you are, then then you are; you walk and quack like a duck. It’s basically about research and trying to understand a problem space and trying to articulate and help people to basically become more effective by understanding that problem space themselves, more. So, it might be about products, as I say, it might be about processes, but for me, I’ve just always enjoyed research.

And I’ve always enjoyed advice. You need a particular mindset to give people advice. That’s one of the key things that, as an industry analyst, you’re sort of expected to do. But yeah, it’s the getting out there and learning from people that is the best part of the job. And I guess that’s why I’ve been doing it for such an ungodly long time; because I love learning, and I love talking to people, and I love trying to help people understand stuff. So, it suits me very well. It’s basically a job, which is about research, analysis, communication.

Corey: The research part is the part that I want to push back on because you say that, and I cringe. On paper, I have an eighth-grade education. And academia was never really something that I was drawn to, excelled at, or frankly, was even halfway competent at for a variety of reasons. So, when you say ‘research,’ I think of something awful and horrible. But then I look at the things I do when I talk to companies that are building something, and then I talked to the customers who are using the thing the company’s building, and, okay, those two things don’t always align as far as conversations go, so let’s take this thing that they built, and I’ll build something myself with it in an afternoon and see what the real story is. And it never occurred to me until we started having conversations to view that through the lens of well, that is actual research. I just consider it messing around with computers until something explodes.

James: Well, I think. I mean, that is research, isn’t it?

Corey: I think so. I’m trying to understand what your vision of research is. Because from where I sit, it’s either something negative and boring or almost subverting the premises you’re starting with to a point where you can twist it back on itself in some sort of ridiculous pretzel and come out with something that if it’s not functional, at least it’s hopefully funny.

James: The funny part I certainly wish that I could get anywhere close to the level of humor that you bring to the table on some of the analysis. But look, I mean, yes, it’s easy to see things as a sort of dry. Look, I mean, a great job I had randomly in my 20s, I sort of lied, fluked, lucked my way into researching Eastern European art and architecture. And a big part of the job was going to all of these amazing museums and libraries in and around London, trying to find catalogs from art exhibitions. And you’re learning about [Anastasi Kremnica 00:15:36], one of the greatest exponents of the illuminated manuscript and just, sort of, finding out about this interesting work, you’re finding out that some of the articles in this dictionary that you’re researching for had been completely made up, and that there wasn’t a bibliography, these were people that were writing for free and they just made shit up, so… but I just found that fascinating, and if you point me at a body of knowledge, I will enjoy learning stuff.

So, I totally know what you mean; one can look at it from a, is this an academic pursuit? But I think, yeah, I’ve just always enjoyed learning stuff. And in terms of what is research, a lot of what RedMonk does is on the qualitative side; we’re trying to understand what people think of things, why they make the choices that they do, you have thousands of conversations, synthesize that into a worldview, you may try and play with those tools, you can’t always do that. I mean, to your point, play with things and break things, but how deep can you go? I’m talking to developers that are writing in Rust; they’re writing in Go, they’re writing in Node, they’re writing in, you know, all of these programming languages under the sun.

I don’t know every programming language, so you have to synthesize. I know a little bit and enough to probably cut off my own thumb, but it’s about trying to understand people’s experience. And then, of course, you have a chance to bring some quantitative things to the table. That was one of the things that RedMonk for a long time, we’d always—we were always very wary of, sort of, quantitative models in research because you see this stuff, it’s all hockey sticks, it’s all up into the right—

Corey: Yeah. You have that ridiculous graph thing, which I’m sorry, I’m sure has an official name. And every analyst firm has its own magic name, whether it’s a ‘Magic Quadrant,’ or the ‘Forrester Wave,’ or, I don’t know, ‘The Crushing Pit Of Despair.’ I don’t know what company is which. But you have the programming language up-and-to-the-right line graph that I’m not sure the exact methodology, but you wind up placing slash ranking all of the programming languages that are whatever body of work you’re consuming—I believe it might be Stack Overflow—

James: Yeah.

Corey: —and people look for that whenever it comes out. And for some reason, no one ever yells at you the way that they would if you were—oh, I don’t know, a woman—or someone who didn’t look like us, with our over-represented faces.

James: Well, yeah. There is some of that. I mean, look, there are two defining forces to the culture. One is outrage, and if you can tap into people’s outrage, then you’re golden—

Corey: Oh, rage-driven development is very much a thing. I guess I shouldn’t be quite as flippant. It’s kind of magic that you can wind up publishing these things as an organization, and people mostly accept it. People pay attention to it; it gets a lot of publicity, but no one argues with you about nonsense, for the most, part that I’ve seen.

James: I mean, so there’s a couple of things. One is outrage; universal human thing, and too much of that in the culture, but it seems to work in terms of driving attention. And the other is confirmation bias. So, I think the beauty of the programming language rankings—which is basically a scatterplot based on looking at conversations in StackOverflow and some behaviors in GitHub, and trying to understand whether they correlate—we’re very open about the methodology. It’s not something where—there are some other companies where you don’t actually know how they’ve reached the conclusions they do.

And we’ve been doing it for a long time; it is somewhat dry. I mean, when you read the post the way Stephen writes it, he really does come across quite academic; 20 paragraphs of explication of the methodology followed by a few paragraphs explaining what we found with the research. Every time we publish it, someone will say, “CSS is not a programming language,” or, “Why is COBOL not on there?” And it’s largely a function of methodology. So, there’s always raged to be had.

Corey: Oh, absolutely. Channeling rage is basically one of my primary core competencies.

James: There you go. So, I think that it’s both. One of the beauties of the thing is that on any given day when we publish it, people either want to pat themselves on the back and say, “Hey, look, I’ve made a really good choice. My programming language is becoming more popular,” or
they are furious and like, “Well, come on, we’re not seeing any slow down. I don’t know why those RedMonk folks are saying that.”

So, in amongst those two things, the programming language rankings was where we began to realize that we could have a footprint that was a bit more quantitative, and trying to understand the breadcrumbs that developers were dropping because the simple fact is, is—look, when we look at the platforms where developers do their work today, they are in effect instrumented. And you can understand things, not with a survey where a lot of good developers—a lot of people in general—are not going to fill in surveys, but you can begin to understand people’s behaviors without talking to them, and so for RedMonk, that’s really thrilling. So, if we’ve got a model where we can understand things by talking to people, and understand things by not talking to people, then we’re cooking with gas.

Corey: I really love installing, upgrading, and fixing security agents in my cloud estate! Why do I say that? Because I sell things, because I sell things for a company that deploys an agent, there's no other reason. Because let’s face it. Agents can be a real headache. Well, now Orca Security gives you a single tool that detects basically every risk in your cloud environment -- and that’s as easy to install and maintain as a smartphone app. It is agentless, or my intro would’ve gotten me into trouble here, but it can still see deep into your AWS workloads, while guaranteeing 100% coverage. With Orca Security, there are no overlooked assets, no DevOps headaches, and believe me you will hear from those people if you cause them headaches. and no performance hits on live environments. Connect your first cloud account in minutes and see for yourself at orca.security. Thats “Orca” as in whale, “dot” security as in that things you company claims to care about but doesn’t until right after it really should have.

Corey: One of the I think most defining characteristics about you is that, first, you tend to undersell the weight your words carry. And I can’t figure out, honestly, whether that is because you’re unaware of them, or you’re naturally a modest person, but I will say you’re absolutely one of my favorite Twitter follows; @monkchips. If you’re not following James, you absolutely should be. Mostly because of what you do whenever someone gives you a modicum of attention, or of credibility, or of power, and that is you immediately—it is reflexive and clearly so, you reach out to find someone you can use that credibility to lift up. It’s really an inspirational thing to see.

It’s one of the things that if I could change anything about myself, it would be to make that less friction-full process, and I think it only comes from practice. You’re the kind of person I think—I guess I’m trying to say that I aspire to be in ways that are beyond where I already am.

James: [laugh]. Well, that’s very charming. Look, we are creatures of extreme privilege. I mean, I say you and I specifically, but people in this industry generally. And maybe not enough people recognize that privilege, but I do, and it’s just become more and more clear to me the longer I’ve been in this industry, that privilege does need to be more evenly distributed. So, if I can help someone, I naturally will. I think it is a muscle that I’ve exercised, don’t get me wrong—

Corey: Oh, it is a muscle and it is a skill that can absolutely be improved. I was nowhere near where I am now, back when I started. I gave talks early on in my speaking career, about how to handle a job interview. What I accidentally built was, “How to handle a job interview if you’re a white guy in tech,” which it turns out is not the inclusive message I wanted to be delivering, so I retired the talk until I could rebuild it with someone who didn’t look like me and give it jointly.

James: And that’s admirable. And that’s—

Corey: I wouldn’t say it’s admirable. I’d say it’s the bare minimum, to be perfectly honest.

James: You’re too kind. I do what I can, it’s a very small amount. I do have a lot of privilege, and I’m aware that not everybody has that privilege. And I’m just a work in progress. I’m doing my best, but I guess what I would say is the people listening is that you do have an opportunity, as Corey said about me just now, maybe I don’t realize the weight of my words, what I would say is that perhaps you have privileges you can share, that you’re not fully aware that you have.

In sharing those privileges, in finding folks that you can help it does make you feel good. And if you would like to feel better, trying to help people in some small way is one of the ways that you can feel better. And I mentioned outrage, and I was sort of joking in terms of the programming language rankings, but clearly, we live in a culture where there is too much outrage. And so to take a step back and help someone, that is a very pure thing and makes you feel good. So, if you want to feel a bit less outraged, feel that you’ve made an impact, you can never finish a day feeling bad about the contribution you’ve made if you’ve helped someone else.

So, we do have a rare privilege, and I get a lot out of it. And so I would just say it works for me, and in an era when there’s a lot of anger around, helping people is usually the time when you’re not angry. And there’s a lot to be said for that.

Corey: I’ll take it beyond that. It’s easy to cast this in a purely feel-good, oh, you’ll give something up in order to lift people up. It never works that way. It always comes back in some weird esoteric way. For example, I go to an awful lot of conferences during, you know, normal years, and I see an awful lot of events and they’re all—hmm—how to put this?—they’re all directionally the same. The RedMonk events are hands down the exception to all of that. I’ve been to Monktoberfest once, and I keep hoping to go to—I’m sorry, was it Monki Gras is the one in the UK?

James: Monki Gras, yeah.

Corey: Yeah. It’s just a different experience across the board where I didn’t even speak and I have a standing policy just due to time commitments not to really attend conferences I’m not speaking at. I made an exception, both due to the fact that it’s RedMonk, so I wanted to see what this event was all about, and also it was in Portland, Maine; my mom lived 15 minutes away, it’s an excuse to go back, but not spend too much time. So, great. It was more or less a lark, and it is hands down the number one event I will make it a point to attend.

And I put that above re:Invent, which is the center of my cloud-y universe every year, just because of the stories that get told, the people that get invited, just the sheer number of good people in one place is incredible. And I don’t want to sound callous, or crass pointing this out, but more business for my company came out of that conference from casual conversations than any other three conferences you can name. It was phenomenal. And it wasn’t because I was there setting up an expo booth—there isn’t an expo hall—and it isn’t because I went around harassing people into signing contracts, which some people seem to think is how it works. It’s because there were good people, and I got to have great conversations. And I kept in touch with a lot of folks, and those relationships over time turned into business because that’s the way it works.

James: Yeah. I mean, we don’t go big, we go small. We focus on creating an intimate environment that’s safe and inclusive and makes people feel good. We strongly curate the events we run. As Stephen explicitly says in terms of the talks that he accepts, these are talks that you won’t hear elsewhere.

And we try and provide a platform for some different kind of thinking, some different voices, and we just had some magical, magical speakers, I think, at both events over the years. So, we keep it down to sort of the size of a village; we don’t want to be too much over the Dunbar number. And that’s where rich interactions between humans emerge. The idea, I think, at our conference is, is that over a couple of days, you will actually get to know some people, and know them well. And we have been lucky enough to attract many kind, and good, and nice people, and that’s what makes the event so great.

It’s not because of Steve, or me, or the others on the team putting it together. It’s about the people that come. And they’re wonderful, and that’s why it’s a good event. The key there is we focus on amazing food and drink experiences, really nice people, and keep it small, and try and be as inclusive as you can. One of the things that we’ve done within the event is we’ve had a diversity and inclusion sponsorship.

And so folks like GitHub, and MongoDB, and Red Hat have been kind enough—I mean, Red Hat—interestingly enough the event as a whole, Red Hat has sponsored Monktoberfest every year it’s been on. But the DNI sponsorship is interesting because what we do with that is we look at that as an opportunity. So, there’s a few things. When you’re running an event, you can solve the speaker problem because there is an amazing pipeline of just fantastic speakers from all different kinds of backgrounds. And I think we do quite well on that, but the DNI sponsorship is really about having a program with resources to make sure that your delegates begin to look a little bit more diverse as well. And that may involve travel stipends, as well as free tickets, accommodation, and so on, which is not an easy one to pull off.

Corey: But it’s necessary. I mean, I will say one of the great things about this past year of remote—there have been a lot of trials and tribulations, don’t get me wrong—but the fact that suddenly all these conferences are available to anyone with an internet connection is a huge accessibility story. When we go back to in-person events, I don’t want to lose that.

James: Yeah, I agree. I mean, I think that’s been one of the really interesting stories of the—and it is in so many dimensions. I bang on about this a lot, but so much talent in tech from Nigeria. Nigeria is just an amazing, amazing geography, huge population, tons of people doing really interesting work, educating themselves, and pushing and driving forward in tech, and then we make it hard for them to get visas to travel to the US or Europe. And I find that to be… disappointing.

So, opening it up to other geographies—which is one of the things that free online events does—is fantastic. You know, perhaps somebody has some accessibility needs, and they just—it’s harder for them to travel. Or perhaps you’re a single parent and you’re unable to travel. Being able to dip into all of these events, I think is potentially a transformative model vis-à-vis inclusion. So, yeah, I hope, A) that you’re right, and, B) that we as an industry are intentional because without being intentional, we’re not going to realize those benefits, without understanding there were benefits, and we can indeed lower some of the barriers to entry participation, and perhaps most importantly, provide the feedback loop.

Because it’s not enough to let people in; you need to welcome them. I talked about the DNI program: we have—we’re never quite sure what to call them. We call them mentors or things like that, but people to welcome people into the community, make introductions, this industry, sometimes it's, “Oh, great. We’ve got new people, but then we don’t support them when they arrive.” And that’s one of the things as an industry we are, frankly, bad at, and we need to get better at it.

Corey: I could not agree with you more strongly. Every time I wind up looking at building an event or whatnot or seeing other people’s events, it’s easy to criticize, but I try to extend grace as much as possible. But whenever I see an event that is very clearly built by people with privilege, for people with privilege, it rubs me the wrong way. And I’m getting worse and worse with time at keeping my mouth shut about that thing. I know, believe it or not, I am capable of keeping my mouth shut from time to time or so I’m told. But it’s irritating, it rankles because it’s people not taking advantage of their privileged position to help others and that, at some point, bugs me.

James: Me too. That’s the bottom line, we can and must do better. And so things that, sort of, make you proud of every year, I change my theme for Monki Gras, and, you know, it’s been about scaling your craft, it’s been about homebrews—so that was sort of about your side gig. It wasn’t about the hustle so much as just things people were interested in. Sometimes a side project turns into something amazing in its own right.

I’ve done Scandinavian craft—the influence of the Nordics on our industry. We talk about privilege: every conference that you go to is basically a conference about what San Francisco thinks. So, it was nice to do something where I looked at the influence of Scandinavian craft and culture. Anyway, to get to my point, I did the conference one year about accessibility. I called it ‘accessible craft.’

And we had some folks from a group called Code Your Future, which is a nonprofit which is basically training refugees to code. And when you’ve got a wheelchair-bound refugee at your conference, then you may be doing something right. I mean, the whole wheelchair thing is really interesting because it’s so easy to just not realize. And I had been doing these conferences in edgy venues. And I remember walking with my sister, Saffron, to check out one of the potential venues.

It was pretty cool, but when we were walking there, there were all these broken cobblestones, and there were quite a lot of heavy vehicles on the road next to it. And it was just very clear that for somebody that had either issues with walking or frankly, with their sight, it just wasn’t going to fly anymore. And I think doing the accessibility conference was a watershed for me because we had to think through so many things that we had not given enough attention vis-à-vis accessibility and inclusion.

Corey: I think it’s also important to remember that if you’re organizing a conference and someone in a wheelchair shows up, you don’t want to ask that person to do extra work to help accommodate that person. You want to reach out to experts on this; take the burden on yourself. Don’t put additional labor on people who are already in a relatively challenging situation. I feel like it’s one of those basic things that people miss.

James: Well, that’s exactly right. I mean, we offered basically, we were like, look, we will pay for your transport. Get a cab that is accessible. But when he was going to come along, we said, “Oh, don’t worry, we’ve made sure that everything is accessible.” We actually had to go further out of London.

We went to the Olympic Park to run it that year because we’re so modern, and the investments they made for the Olympics, the accessibility was good from the tube, to the bus, and everything else. And the first day, he came along and he was like, “Oh, I got the cab because I didn’t really believe that the accessibility would work.” And I think on the second day, he just used the shuttle bus because he saw that the experience was good. So, I think that’s the thing; don’t make people do the work. It’s our job to do the work to make a better environment for as many people as possible.

Corey: James, before we call it a show, I have to ask. Your Twitter name is @monkchips and it is one of the most frustrating things in the world trying to keep up with you because your Twitter username doesn’t change, but the name that goes above it changes on what appears to be a daily basis. I always felt weird asking you this in person, when I was in slapping distance, but now we’re on a podcast where you can’t possibly refuse to answer. What the hell is up with that?

James: Well, I think if something can be changeable, if something can be mutable, then why not? It’s a weird thing with Twitter is that it enables that, and it’s just something fun. I know it can be sort of annoying to people. I used to mess around with my profile picture a lot; that was the thing that I really focused on. But recently, at least, I just—there are things that I find funny, or dumb, or interesting, and I’ll just make that my username.

It’s not hugely intentional, but it is, I guess, a bit of a calling card. I like puns; it’s partly, you know, why you do something. Because you can, so I’ve been more consistent with my profile picture. If you keep changing both of them all the time, that’s probably suboptimal. Sounds good.

Corey: Sounds good. It just makes it hard to track who exactly—“Who is this lunatic, and how did they get into my—oh, it’s James, again.” Ugh, branding is hard. At least you’re not changing your picture at the same time. That would just be unmanageable.

James: Yeah, no, that’s what I’m saying. I think you’ve got to do—you can’t do both at the same time and maintain—

Corey: At that point, you’re basically fleeing creditors.

James: Well, that may have happened. Maybe that’s an issue for me.

Corey: James, I want to thank you for taking as much time as you have to tolerate my slings, and arrows, and other various vocal devices. If people want to learn more about who you are, what you believe, what you’re up to, and how to find you. Where are you hiding?

James: Yeah, I mean, I think you’ve said already, that was very kind: I am at @monkchips. I’m not on topic. I think as this conversation has shown, I [laugh] don’t think we’ve spoken as much about technology as perhaps we should, given the show is normally about the cloud.

Corey: The show is normally about the business of cloud, and people stories are always better than technology stories because technology is always people.

James: And so, yep, I’m all over the map; I can be annoying; I wear my heart on my sleeve. But I try and be kind as much as I can, and yeah, I tweet a lot. That’s the best place to find me. And definitely look at redmonk.com.

But I have smart colleagues doing great work, and if you’re interested in developers and technology infrastructure, we’re a great place to come and learn about those things. And we’re very accessible. We love to talk to people, and if you want to get better at dealing with software developers, yeah, you should talk to us. We’re nice people and we’re ready to chat.

Corey: Excellent. We will, of course, throw links to that in the [show notes 00:37:03]. James, thank you so much for taking the time to speak with me. I really do appreciate it.

James: My pleasure. But you’ve made me feel like a nice person, which is a bit weird.

Corey: I know, right? That’s okay. You can go for a walk. Shake it off.

James: [laugh].

Corey: It’ll be okay. James Governor, analyst and co-founder at RedMonk. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you hated this podcast, please leave a five-star review on your podcast platform of choice along with an insulting comment in which you attempt to gatekeep being an industry analyst.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Sean

Sean Kilgore is an Architect at Twilio, where he draws boxes, lighthouses and soapboxes. In Sean’s spare time, he enjoys reading, walking, gaming, and a well-made drink.

Links:

  • Twilio: https://www.twilio.com/
  • Silvia Botros's Twitter: https://twitter.com/dbsmasher
  • Sean's Twitter: https://twitter.com/log1kal

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at the Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Your company might be stuck in the middle of a DevOps revolution without even realizing it. Lucky you! Does your company culture discourage risk? Are you willing to admit it? Does your team have clear responsibilities? Depends on who you ask. Are you struggling to get buy in on DevOps practices? Well, download the 2021 State of DevOps report brought to you annually by Puppet since 2011 to explore the trends and blockers keeping evolution firms stuck in the middle of their DevOps evolution. Because they fail to evolve or die like dinosaurs. The significance of organizational buy in, and oh it is significant indeed, and why team identities and interaction models matter. Not to mention weither the use of automation and the cloud translate to DevOps success. All that and more awaits you. Visit: www.puppet.com to download your copy of the report now!

Corey: Up next we’ve got the latest hits from Veem. Its climbing charts everywhere and soon its going to climb right into your heart. Here it is!

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’m joined this week by Sean Kilgore who’s an architect at a small company called Twilio. Sean, welcome to the show.

Sean: Corey, it’s a pleasure to be here.

Corey: It really is. You’re one of those fun people that I always mean to catch up with and really do a deep dive in, but we keep passing like ships in the night. And in fact, I want to go back to more or less what is pretty damn close to your first real job in technology. You were a
network administrator at Lutheran High School in Orange, California.

Sean: I was.

Corey: And at that same time, I was a network administrator down the street at Chapman University, also in Orange, California. And despite that, we have traveled in many of the same circles since, but we have never met in person despite copious opportunities to do so.

Sean: That is amazing.

Corey: Talking to you is like looking into a funhouse mirror of what would it have been like if I could, you know, hold down a job, and was actually good at things. It’s really fun. Apparently, I’d be able to grow a better beard.

Sean: I don’t know about that. My beard is pretty patchy.

Corey: Yeah, I look like an angry 14-year-old trying to prove a point to Mommy and Daddy. But that’s not really the direction that we need to take this in today. And you’ve done a lot of stuff that aligns with things that are near and dear to my heart. For the last—what is it now?—six years and change? Seven years and change? You’ve been at SendGrid, then Twilio through acquisition?

Sean: Mm-hm.

Corey: And you have done basically every operations-looking job at that company. You’ve had a bunch of titles. You wound up going from DevOps engineer to a team lead, then to a senior DevOps engineer again, and you call—you voluntarily move back to an individual contributor role. Let’s start there. What was management like?

Sean: The management was interesting. My first go at that, I had no idea what I was doing, and so I didn’t know how to ask for the help that I needed. And so my wife and I refer to that time as the time that I played a lot of video games. Just, I wasn’t prepared for the emotional outlay that managing humans costs. And so I would end up spending my nights just playing video games trying to unwind and unpack from all that.

I’ve managed twice, now. The second time was—it went much better. I knew more of what I was doing, I had more support. The manager of the team that had left, I had worked with that team very closely in the past, I’d been part of it. And so my whole purpose there was to make sure that we didn’t lose anybody after that until we found a new manager.

And that actually worked out pretty well. I had to have some really difficult conversations with some people along the way, but they all stayed, they all told me that they really enjoyed me managing them. And I’ve had people ask me to manage them again after that, which was super bonkers.

Corey: It’s always flattering when you have an impact as a manager and people seek you out to work with you again. I dabbled in management as well; no one ever asked me to do it twice, and I have effectively avoided managing people here at The Duckbill Group, just because my belief of what a good manager is—and I think it aligns with what you’ve already said—requires a certain selflessness and ability to focus on others and grow them, whereas my role here is very much as face of the company, and it’s about me. That’s not a recipe for a successful
outcome for managing people and not having them rage-quit.

Sean: Definitely.

Corey: One thing that I find is interesting about management, the higher you rise in an organization, it’s counterintuitive but the more responsibility you get, but the less you can directly affect yourself. Your entire world becomes about effectively delegating work to others and about influence. In your case now, you are an architect, which means different things at different companies. So, I’m not entirely sure I know what it means at Twilio. So first, what is an architect at Twilio? And where does your responsibility start and stop, I guess is where I want to go, there?

Sean: So, architecture at Twilio is kind of different. Some of the architects that I’ve worked with in the past is the ‘I design everything, and then engineers go and build the thing that I designed.’

Corey: Oh, yes. The enterprise architect approach where I’m going to sit in my ivory tower and dispatch, effectively, winged monkeys to
implement things that I don’t fully understand, but I have the flashy title. You’re saying that’s not what it is?

Sean: There’s a fine distinction here because I think that some of the people I work with would definitely say that I am up in my ivory tower. It is more about—if I’m looking five years out—what capabilities my teams need to be able to provide to execute against a business strategy. That landscape is going to change immensely along the way, and so my job isn’t to say, “We’re going to use Kubernetes because that’s what we need to do five years out.” It happens to be what I’m saying right now, and I’m sure we’re going to go into that, Corey, but it’s less about this is how each of these things should be strung together to achieve that goal and more direction setting. So, I worked on something that I call the ‘Lighthouse,’ and it’s a vision of the future of where we want to go but the caveat is that if you actually go to the Lighthouse, it means you
hit the rocks.

It’s describing what I call the ‘Bay of Appropriate Futures,’ and you want to land somewhere in that Bay, but it’s not going to be the thing I wrote down five years before we get there. And so it’s much more of a technical leadership position, trying to help other technologists make good technology decisions. And so it’s more about making sure that all of the right questions get asked, not having all the right answers. That’s the difference between some architects that I’ve worked with in the past.

Corey: And one of the challenges in that role is that you’re not managing people directly. So, what that means is, you are, on some level, not doing a lot of the hands-on keyboard implementation work yourself and, unless I dramatically misunderstand your corporate culture, you’re not empowered to unilaterally fire people, which means that you can only really lead via example and influence. Tell me about that.

Sean: When people ask me if they want to be architects, I ask them if they can influence without authority, or if that’s even interesting to them. That is definitely the thing. And so when it comes down to, “Man, I really wish I could fire this person.” A, that never happens. But, B, it’s definitely about modeling behaviors. And there’s a bit of management here in that making expectations clear of senior engineers is part of my job, and helping them also be examples for other engineers is definitely a thing I get graded on.

Corey: Influence without authority is sort of the definitional characteristic of being a consultant. It turns out, you can’t even force anything; it has to be the strength of your ideas combined, in some shops, and I admit I have a bit of a somewhat cynical view on this, but also the ability for the client to commit to their sunk-cost fallacy of, “Well, we paid a lot of money to hear this advice, we should probably do something with it.” And there’s always a story of making sure that you’re serving the organization with which you work, well. But when you can only influence rather than direct, it becomes a much more nuanced thing, and I feel like the single differentiator between success and failure in that role is, fundamentally, empathy. Am I wrong on that?

Sean: Not at all. When I’m working with a very in-the-code engineer who comes to me and is trying to convince their team that they should do something, one of the things that’s a stumbling block for them is that they don’t realize that other people need to be influenced in ways that might be different than the way that they’re influenced. So, as an example, I work with a team of very senior people; I know that some people will respond well, if I site Accelerate, for example. And some people want to hear, “Well, Google does this, so it’s obvious that we should do that.” And trying to thread that needle with everyone in a way that gets everyone on board in the best way for all of us, when you can do that, you can be an architect.

Corey: Some weeks, I feel like I’m closer to architect than I do others. It seems that the idea now—solutions architect being a job role that a lot of companies have they hurl out into the universe—is in many ways vastly misunderstood. You want to talk about some kind of architecture story, where I’m going to go ahead and design an architecture that solves some business problem on a whiteboard. That’s not hard to do.

The hard part is then controlling for constraints such as, “Yeah, we already have a thing that doesn’t look anything like that, and we want to get it to that point. And oh, by the way, 18 months of downtime while we do that is not acceptable.” Nothing is ever truly greenfield. And adapting to constraints, and making compromises, and being realistic seems to separate the folks who are good at that from the folks who are
playacting.

Sean: There’s another thing in there where people who’ve worked on the same thing for a long time sometimes have a problem seeing where you could go. And so the constraints there can be really useful in designing things. A lot of people think that greenfield is awesome, but greenfield just means the possible outcomes are the entire universe. I like working in constraints, essentially.

Corey: So, I want to talk to you a little bit about your tenure at Twilio, where you started off at SendGrid and then there was an acquisition, and everyone I know got super-quiet for a while because it turns out when there’s a pending acquisition, talking to people about it is frowned upon, and that goes doubly so in the context of someone who basically shitposts for a living. And I get that; I respect confidentiality, but I also don’t want people to jeopardize their own positions. So, it’s one of those, “Yeah, whatever you’re comfortable telling me, or not, is fine.” So, we didn’t talk for a while. And then the acquisition happened, and now you’re there. And you’ve been there at Twilio for a couple of years or so and haven’t rage-quit, so apparently, it’s worked. What was the transition like?

Sean: The transition was really interesting. A lot of people were telling me that acquisitions were universally horrible. And that’s not how it worked. This is the first acquisition I’ve been through, so I have no context. People I trust told me that this one went well.

So, in my role as architect, this acquisition was kind of interesting because SendGrid had a very robust, and we’ve done architecture for years. And Twilio’s architecture was a little bit different. It was more like, “There are some really, really senior people at Twilio who have seen some things, and you should probably ask them their opinion on stuff.” But there wasn’t really, like, an architecture review process. There was definitely, “A you need to write down a bunch of stuff and get some people to look at it,” but it wasn’t a, you need to get approved by your local architect or a group of architects.

Part of that is to provide visibility across the org so that we’re not duplicating work and stuff. But Twilio basically adopted SendGrid’s architecture process, but it grew 10x. So, at SendGrid, we had, I think, six architects. At Twilio, we have, like, 40 now, so not quite 10x. But trying to copy and paste that process was kind of rough.

We’re still, kind of, making that better. And then there were a couple things—as the acquired company, you kind of expect, I don’t know, maybe some housecleaning to happen. And that isn’t what happened. We saw a lot of like really senior leaders move into positions at Twilio of leadership. So, on day one, I think SendGrid’s sales leader became the sales leader for all of Twilio.

And that sounded—I haven’t done this before, but it sounded like that’s not normal. And that’s happened in a couple different spots. It’s been pretty neat. And when I think about the acquisition, not just of acquiring another channel for Twilio, but kind of doing an acqui-hire of a bunch of key positions, that was a pretty valuable one.

Corey: Let’s talk about one aspect of working at Twilio that I profoundly envy you for, which is working with one of the greatest people in the world: Silvia. Let’s talk about Silvia.

Sean: She interviewed me at SendGrid. She’s been here almost a year longer than I have, and it’s been such a joy to work with her. Not just because everything gets Botros’ed around her, and so we have our own built-in chaos monkeys, but also, there’s no one that cares more about making sure that what teams are building won’t come back and bite the team later. She’s worked with, I think, maybe a 10th of the company now—and at Twilio, that’s a lot of teams—trying to just help them do better and make sure that the stuff that they’re building is not going to page them all the time, is actually going to serve the customer in ways that isn’t surprising to the customer. I can probably talk for half an hour about my appreciation for Silvia.

Corey: Well, she was a great guest in the early days of this podcast. Silvia Botros is phenomenal. She has the Twitter handle of @dbsmasher so she’s my default go-to on misusing things as databases. And she also was just one of the most genuinely kind people I know.

She also has an aura effect, where she is basically a walking EMP, and every time someone tries to show her a piece of technology, it explodes in novel and interesting ways, which, frankly, as an acceptance gate for technology is a terrific skill set to have. Does it cause problems in the office?

Sean: Not normally. It causes more problems in the office when we are actually in an office together because Silvia, maybe predictively, is also a giant klutz. And so the joke is that she also EMPs herself. In the office, she does break things, but it’s never in an intended way. Or it’s just like a fun, “Oh, man, the WiFi’s down. It must be Silvia.”

Corey: Exactly. It’s always nice to have someone you can blame for these things.

Sean: It’s SOP.

Corey: Yeah, oh, absolutely. At some point do you ever wind up missing things such as, “Oh, it’s probably just Silvia. No, it was actually a problem somewhere?”

Sean: So, we actually determine that Silvia’s EMP works at a distance. She flew somewhere close to one of our hardware data centers, and at the time that she passed it, we had an outage. Like, the data center went dark, kind of thing. And so it still happens even if she’s not around. We’re pretty sure it’s not a local phenomenon.

Corey: So, the thing that I know is probably going to sound completely boring and ridiculous to half the audience while the other half the audience sits and listens raptly; before I started this place, I never stayed anywhere for longer than two years, because as previously disclosed in multiple directions, I am a terrible employee. First, why did you stay at the same place for as long as you have, and what’s it like? And I’m really hoping you have an answer that isn’t just, “Oh, I have a complete lack of ambition,” because I won’t believe that for a second. But it is a tempting cop-out so let me just shut that down now.

Sean: No, it’s more I’ve been here for almost eight years now, and I’ve never done the same thing. The fun fact that I tell people when they onboard or I’m interviewing them is, I think I’ve had more titles than anyone at Twilio. I’m up to 11, I think. And so 11 titles in eight years? It hasn’t been the same company.

When I started, I was employee number 150. There were 80 of us when I started at SendGrid. I work at a company with 4500 people now and going through that growth, the company that I work for today, and even pre-acquisition, you would not recognize from the day I started at SendGrid. And so if I had been doing the same thing all the time, I wouldn’t still be here. There was a point before we started architecture at SendGrid, I was definitely in a spot kind of a rut, like, “Cool. I can continue to do the same thing over and over,” but I felt like there wasn’t a lot
of growth to do.

I needed to go see something else, kind of thing. I knew really well how to do our mail stuff, and I felt like I needed to broaden my horizons a
bit, or I needed to level up. And at the time, the only place to go up was to management. And then we brought in a chief architect, J.R. Jasperson, and I I remember very clearly, it was like his third day or something, we had an all-company meeting—like a lunch thing—and I walked up to him and said, “You don’t know this yet, but you’re my new mentor because it seems like what you’re talking about is really interesting to me.” And he didn’t know it but the subtext there was like, “And if you don’t, I’m out.”

And since then, the work that I do day-to-day is completely different. Like, I work for a platform org. This platform org is 130 people right now; it spans everything from building EC2 instances to, recently, it was, like, Twilio API Edge. There’s such a breadth in there that I never do the same thing every day.

Corey: I really love installing, upgrading, and fixing security agents in my cloud estate! Why do I say that? Because I sell things, because I sell things for a company that deploys an agent, there's no other reason. Because let’s face it. Agents can be a real headache. Well, now Orca Security gives you a single tool that detects basically every risk in your cloud environment -- and that’s as easy to install and maintain as a smartphone app. It is agentless, or my intro would’ve gotten me into trouble here, but it can still see deep into your AWS workloads, while guaranteeing 100% coverage. With Orca Security, there are no overlooked assets, no DevOps headaches, and believe me you will hear from those people if you cause them headaches. and no performance hits on live environments. Connect your first cloud account in minutes and see for yourself at orca.security. Thats “Orca” as in whale, “dot” security as in that things you company claims to care about but doesn’t until right after it really should have.

Corey: That’s functionally, I think, the problem that I had in working in environments as a DevOps type because for the first three months in a job where I’m the first ops person, “Great everything’s on fire.”—I’m an adrenaline junkie in that sense—“Cool. Oh wow, all these problems that I know how to fix.” And then it gets to a reasonable level of working and now it’s just care and feeding of same. Okay, now I’m getting slightly bored, so let me look for other problems in other parts of the org.

And that doesn’t go super well when you’re not welcome in those parts of the org which leads to a whole bunch of challenges I’ve had in my career. This is incidentally why being a consultant aligns so well with me and the way I approach things. It’s cool. I’m going to come in; I’m going to fix things, and then I get to leave. On day one, we know this is a time-bound engagement and that’s okay.

Instead of going down the path of the lies everyone tells themselves where average tenure in this space is 18 to 24 months, but magically we’re all going to lie and pretend in the interview that this is their forever job and suddenly you’re going to stay here for 25 years and get a pension and a gold watch when you retire. And it’s, oh wow that’s amazing it sounds like everybody having these conversations wearing old-timey stovepipe hats. There’s just so much that isn’t realistic in those conversations. So, I talk to people who’ve been down those paths who’ve been at the same company for a decade or two, and the common failure mode there is that they have a year or so of experience that they repeat 10 or 20 times. And that’s sad; people get stuck. What you say absolutely resonates with me in that every year is a different thing that you’re working on. You’re not doing the same thing twice. I get antsy when too many days look the same, one to the next.

Sean: I definitely hear that. If my every day was, come in join a stand-up, talk about the problem that I had last week and still have today, it wouldn’t work for me. I feel lucky that I work for an organization where outside input is actually requested and honored, so if I go to a team and just happened to have noticed something and say, “Hey, this right here you might want to take a look at. And I have some opinions here if you’d like to hear them.” I normally get asked for that opinion, and it normally turns out pretty good.

There’s definitely times where it’s been, like, “No, Sean. You don’t know what you’re talking about.” And normally they’re right. It’s definitely not the same. People say you should be at a company for 18 to 24 months. And that’s true if your company is totally shortchanging you. When I ask my peers at other companies about, have I gotten stiffed by staying at the same place for this long? It’s definitely not. And if that wasn’t true, if Twilio was holding back my compensation, maybe this would have gone a different way, but it’s not what’s happened.

Corey: Oh, true and to be clear that is very often the biggest criticism I have of people who stay at one company for a long time. They don’t realize what market rate is anymore and they find themselves in a scenario where, “Wow, I could go somewhere else and triple my salary,” which is not an exaggeration and an unpleasant discovery when people realize that they’ve been taken advantage of. And credit where due, I have had conversations with people at Twilio who have been there a long time. And I have never gotten the impression that that is what’s going on there. Your compensation is fair. I want to be very clear here. This is not one of those, “Oh, yeah, I’m just trying to be polite because someone’s being taken advantage of and doesn’t even know it.” No. They’re doing right by their employees. The fact that I have to call them out explicitly as an example of a rarity of a company doing right by its employees, is monstrous.

Sean: It is. We’ve hired a few people recently where I found out that I think their pay was close to doubled just by coming here, and I just wish that it was more okay to call out their prior employees publicly and be like, “Cool. If you work for this company will probably pay you a ton more.”

Corey: And that’s the other side of it, too, which I did early on in my career. It’s, “Oh, I’m leaving this company and screw you all.” “Well, why are you leaving?” “Oh, because I’m getting a 5% raise to change jobs.” I’m not saying that money should not factor into it, but at some point, when all is said and done at that scale, it works out to be 100 bucks a paycheck, or so, is it really worth changing for that? Maybe.

If there are things you don’t like about the environment, please don’t let me dissuade people from interviewing for jobs. You always should be doing that, on some level, just so what the market looks like. But I’m also a big believer in, you don’t need to be as mercenary as I was early on in my career. A lot of it was shaped by environments—not Chapman, I want to be clear—that were not particularly kind to staff. And that I felt taken advantage of because I was. And as a result, “Oh, screw me? Screw you.” And it became a very mercenary approach that didn’t serve me well. That is now a baked-in aspect of how I view careers in some respects, and that is something of a problem that I wrestle with.

Sean: The mercenary thing?

Corey: Yeah. I wrestle with the mercenary thing just because when I talk to someone who’s having a challenge at work, or something, my default instinctive gut reaction that I’ve learned to suppress is, “Oh, screw ‘em. Quit and find another job.”

Sean: Ah, gotcha.

Corey: That’s not the most constructive way to work in the context of a company where you’re building a career trajectory, and a reputation, and you’ve been there for five years, and maybe rage-quitting because you didn’t wind up getting to pick the title of that presentation isn’t the best answer. I can be remarkably petty, for the reasons I’ll leave a company. But that’s not constructive, and I try very hard to avoid giving that advice unnecessarily to people.

Sean: It’s definitely just, like, incidents, right? It’s never a root cause; it’s a contributing factor, and pay is just one contributing factor. I find that a lot of people, even if they’re being taken advantage of compensation-wise, they won’t leave unless there’s something else wrong.

Corey: Yeah, compensation is absolutely a symptom, and in most cases, that’s not the real reason people are going to leave. I assure you, people who work at The Duckbill Group could make more money, objectively, somewhere else. But there’s a question of what people value. We pay people well, but we don’t offer FAANG money, with the equity upside and the rest. We’re not trying to pull the Netflix and pay absolute top-of-market in all cases to all people.

I would love to be able to do that; our margins don’t yet support that. Thanks to our sponsors, we’re going to continue to ratchet those prices way up. I kid. I kid. But there are business reasons why things are the way they are. What we do offer instead is things that contribute to a workplace we want to work at. More of us are parents than aren’t.

We don’t expect people to work outside of business hours in almost any scenario, short of, you know, re:Invent or something. There’s a very human approach to it. We’re not VC-backed at all, so we don’t ever have to worry about trying to sprint to hit milestones over debt as a company. We have this insane secret approach called ‘revenue’ and ‘profitability’ that means we can continue to iterate month to month, and as long as the trend line continues, we’re happy.

Sean: That kind of sustainability is awesome, and is a really good indicator that a company is going to be successful, to me. Especially smaller companies; the decision to not take VC money. And to chase sustainable revenue growth, I know everyone wants to chase the hockey sticks, but at what cost?

Corey: Yeah. And I think that people put this on job-seekers way too much. I have been confronted, at one point—I will not name the company—when I was interviewing years ago, and I was asked by the hiring manager, “Well, it seems like you’ve done a fair bit of job-hopping in the course of your career. What’s up with that?” And they pulled up my LinkedIn profile and went through it, and I said, “You realize most of those were contract gigs?” “Well, I don’t kno—oh, yeah. I guess it was. Oh, that was—huh. I guess so.”

So, it was a failed attempt to call bullshit on my job history. And because I don’t take things like that well, I turned it around right back on him, and I said, “No, I appreciate that. Thank you for clarifying.” That’s a warning sign is when I thank you for insulting me because what’s coming next is always going to sting. But while we’re on the topic of turnover, “Your team has lost 80% of its members in the last six months. What’s going on with that? Is there a problem here that I should be aware of?” And suddenly, the back-peddling was phenomenal. I did get an offer from that company; I did not accept it.

Sean: Good.

Corey: You can tell a lot about a company by how they buy their people. And if you’re actively being insulted or hazed in the job interview process, no. I want people who I choose not to hire, to come away from the experience feeling respected and that they enjoyed the experience to the point where they would say nice things about us if asked, or even evangelize us without ever even having to be asked. And so far, we’ve done that because we’re very intentional on how we approach things. And man, am I tired of people doing this badly.

Sean: When I interview someone, I want them to leave, and then if they don’t take the job, I want it to be because it wasn’t the right job for them. Or, like, the team wasn’t the right fit, not because anything happened in the interview process that was a red flag. That’s the worst. I want Twilio to be a spot where there are no red flags. That would be ideal.

Corey: Absolutely. I think that so many folks get it wrong, where there’s this idea of, “Oh, I’m going to interview you. And oh, you’re an ops person. Great. I want you to implement Quicksort on the whiteboard.” And it’s, “Question one: do you do that a lot here? And two: no, of course you don’t because I’ve seen your services list. There’s no rhyme or reason to the order it appears in. Maybe someone should implement Quicksort in production.”

And then there’s the other side, too, of, “Oh, great. There’s this broad skill set across the entire space. I’m going to figure out where you’re weak and then needle you on those.” I don’t like hiring for absence of weakness; I like hiring for, you’re really good at things we need here and you’re acceptable at the things that are non-negotiable, and able to improve in areas where it becomes helpful.

Sean: Yeah. The best interview process I ever had, they flew me up to San Francisco, I worked with them for a day on a real problem that they had, like, pair programming. They offered me the job—it didn’t work out because I didn’t want to move to San Francisco, it turned out—but that interview process was super valuable to me as the candidate because I found out exactly how a day at that job would work; what it would
look like.

Corey: I had a very similar experience once and the cherry on the top was they paid me a nominal contracting rate for the day—

Sean: Same.

Corey: —because it was touching things that they were doing. And I think that that’s another anti-pattern of, that was a thing that also just happens to be a thing we’re going to use in production, but we’re not going to tell you that we’re not going to compensate you for it. I’ll work on toy problems; not production in an interview context.

Sean: I wanted to circle back to one thing about leaving a company, like, rage-quitting. It’s essentially—if you rage-quit because of a problem, like, a small thing, you’re missing an opportunity to grow. And especially if I had one superpower, I would say that it’s probably managing up. Part of this is just, I have a lot of privilege that lets me do that, but it is definitely a skill that I wish more people had for their own careers.

Corey: I really do, too. We spent all this time practicing how to be a candidate in a job interview, and almost no time training people how to be a good interviewer, and what you’re looking for. And you wind up with terrible things like, “I had this problem once in production that I thought was super clever, so I’m going to set it up for you and see how you would solve that problem. And if you don’t follow the exact same path that I did, then we’re going to go ahead and just keep shooting down anything else you suggest.” No, stop it.

Sean: I do a lot of interviewing, and so I love when I learn something from a candidate because I can ask them questions that are like, “How did you figure this out? How did you even notice that this was a problem?”s and you get to go really deep in something they know, the way they know it. We used to do the, like, “Build us an LRU cache in the best big-O notation time.” And if you didn’t get it, you didn’t get the job, if you did it in slower than optimum time.

And I remember leaving one of the interviews and doing the recap, and it was like, if anyone came to work and did this, I would be upset at them for wasting time. This is part of the standard library of all the things that we do. Why are we asking this question? I know for a while we stopped asking the question, which is great. I don’t do a lot of code interviews at Twilio, so I don’t know if we do something similar there, yet. I should go find out.

Corey: I do not know either way, to be clear. None of the stories I’m talking about involved Twilio. Though I will say, I went on an interview years ago at SendGrid in Anaheim, and I don’t know if I ever got a formal rejection or not afterwards, but regardless, they did not opt to hire me. In hindsight, good decision.

Sean: I wonder if we were in the office at the same time.

Corey: It would have been 2006, so I think it might have been a bit before your time.

Sean: That was before my time.

Corey: And very much, credit where due, I started my career in large-scale email systems, so SendGrid was one of those. Oh, I could probably apply the skill set there. The problem, of course, was that it became pretty apparent, even in those days that eventually there weren’t going to be that many companies that needed that skillset. The days of an email admin in every company were drawn to a close, and it was time to evolve or die.

Sean: You’re welcome.

Corey: Of course. And again SendGrid today, under the hood—deep under the hood—does still power Last Week in AWS. You folks send emails and get them where they need to go, for which I thank you, and the rest of the world probably does not most weeks.

Sean: [laugh]. Yeah.

Corey: Ugh. So, we’ve covered a lot of wide-ranging topics. If people want to hear more about who you are, and what you have to say, where can they find you?

Sean: I’m on Twitter at @log1kal with a one and a K because I hate people who want to find me, apparently. But that’s @log1kal. Twitter’s probably the only thing.

Corey: Excellent. We’ll, of course, put a link to that in the [show notes 00:30:39]. Sean, thank you so much for speaking with me today. I really appreciate it.

Sean: Thank you so much, Corey. I love having these kinds of conversations. I love that there is no plan; we’re just going to have a conversation and record it. I love listening to these kinds of podcasts.

Corey: Well, I like creating these kinds of podcasts because the other ones take way too much work.

Sean: [laugh].

Corey: Sean Kilgore, architect at Twilio. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with a comment explaining how it is almost certainly the fault of Silvia Botros’s aura.

Sean: [laugh].

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

This has been a HumblePod production. Stay humble.

View Details

About Nigel

Nigel Kersten’s day job is Field CTO at Puppet where he leads a group of engineers who work with Puppet’s largest customers on cultural and organizational changes necessary for large-scale DevOps implementations - among other things. He’s a co-author of the industry-leading State Of DevOps Report and likes to evenly talk about what went right with DevOps and what went wrong based on this research and his experience in the field. He’s held multiple positions at Puppet across product and engineering and came to Puppet from the Google SRE organization, where he was responsible for one of the largest Puppet deployments in the world. Nigel is passionate about behavioral economics, electronic music, synthesizers, and Test cricket. Ask him about late-stage capitalism, and shoes.

Links:

  • Puppet: https://puppet.com
  • 2020 State of DevOps Report: https://puppet.com/resources/report/2020-state-of-devops-report/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by LaunchDarkly. Take a look at what it takes to get your code into production. I’m going to just guess that it’s awful because it’s always awful. No one loves their deployment process. What if launching new features didn’t require you to do a full-on code and possibly infrastructure deploy? What if you could test on a small subset of users and then roll it back immediately if results aren’t what you expect? LaunchDarkly does exactly this. To learn more, visit launchdarkly.com and tell them Corey sent you, and watch for the wince.

Corey: Your company might be stuck in the middle of a DevOps revolution without even realizing it. Lucky you! Does your company culture discourage risk? Are you willing to admit it? Does your team have clear responsibilities? Depends on who you ask. Are you struggling to get buy in on DevOps practices? Well, download the 2021 State of DevOps report brought to you annually by Puppet since 2011 to explore the trends and blockers keeping evolution firms stuck in the middle of their DevOps evolution. Because they fail to evolve or die like dinosaurs. The significance of organizational buy in, and oh it is significant indeed, and why team identities and interaction models matter. Not to mention weither the use of automation and the cloud translate to DevOps success. All that and more awaits you. Visit: www.puppet.com to download your copy of the report now!

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. This promoted episode is sponsored by a long time… I wouldn’t even say friends so much as antagonist slash protagonist slash symbiotic company with things I have done as I have staggered through the ecosystem. There’s a lot of fingers of blame that I can point throughout the course of my career at different instances, different companies, different clients, et cetera, et cetera, that have shaped me into the monstrosity than I am today. But far and away, the company that has the most impact on the way that I speak publicly, is Puppet.

Here to accept the recrimination for what I become and how it’s played out is Nigel Kersten, a field CTO at Puppet—or the field CTO; I don’t know how many of them they have. Nigel, welcome to the show, and how unique are you?

Nigel: Thank you, Corey. Well, I—you know, reasonably unique. I think that you get used to being one of the few Australians living in Portland who’s decided to move away from the sunny beaches and live in the gray wilderness of the Pacific Northwest.

Corey: So, to give a little context into that ridiculous intro, I was a traveling contract trainer for the Puppet fundamentals course for an entire summer back in I want to say 2014, but don’t hold me to it. And it turns out that when you’re teaching a whole bunch of students who have paid in many cases, a couple thousand dollars out of pocket to learn a new software where, in some cases, they feel like it’s taking their job away because they view their job, rightly or wrongly, is writing the same script again and again. And then the demo breaks and people are angry, and if you don’t get a good enough rating, you’re not invited to continue, and then the company you’re contracting through hits you with a stick, it teaches you to improvise super quickly. So, I wasn’t kidding when I said that Puppet was in many ways responsible for the way that I give talks now. So, what do you have to say for yourself?

Nigel: Well, I have to say, congratulations for surviving, opinionated defensive nerds who think not only you but your entire product you’re demoing could be replaced by a shell script. It’s a tough crowd.

Corey: It was an experience. And some of these were community-based, and some of them were internal to a specific company. And if people have heard more than one episode of this show, I’m sure they can imagine how that went. I gave a training at Comcast once and set a personal challenge for myself of how many times could I use the word ‘comcastic’ in a three-day training. And I would work it in and talk about things like the schedule parameter in Puppet where it doesn’t guarantee something’s going to execute in a time window; it’s the only time it may happen.

If it doesn’t fire off, and then it isn’t going to happen. It’s like a Comcast service appointment. And then they just all kind of stared at me for a while and, credit where due, that was the best user rating I ever got from people sitting through one of my training. So, thanks for teaching me how to improve at, basically, could have been a very expensive mistake on Puppet’s part. It accidentally worked out for everyone.

Nigel: Brilliant, brilliant. Yes, you would have survived teaching the spaceship operator to that sort of a crowd.

Corey: Oh, I mostly avoided that thing. That was an advanced Puppet-ism, and this was Puppet fundamentals because I just need to be topically good at things, not deep-dive good at things. But let’s dig into that a little bit. For those who have not had the pleasure of working with Puppet, what is it?

Nigel: Sure? So, Puppet is a pretty simple DSL. You know, DSLs aren’t necessarily in favor these days.

Corey: Domain-specific language, for those who have not—

Nigel: Yep.

Corey: —caught up on that acronym. Yes.

Nigel: So, a programming language designed for a specific task. And, you know, instead, we’ve decided that the world will rest on YAML. And we’ve absorbed a fair bit of YAML into our ecosystem, but there are things that I will still stand by are just better to do in a programming language. ‘if x then y,’ for example, it’s just easier to express when you have actual syntax around you and you’re not, sort of, forcing everything to be in a data specification language. So, Puppet’s pretty simple in that it’s a language that lets you describe the state that infrastructure should be.

And you can do this in a modular and composable way. So, I can build a little chunk of automation code; hand it to Corey; Corey can build something slightly bigger with it; hand it to someone else. And really, this sort of collaboration is one of the reasons why Puppet’s, sort of, being at the center of the DevOps movement, which at its core is not really about tools. It’s about reducing friction between different groups.

Corey: Back when I was doing my traveling training shtick, I found that I had to figure out a way to describe what Puppet did to folks who were not deep in the space, and the analogy that I came up with that I was particularly partial to was, imagine you get a brand new laptop. Well, what do you do with it? You install your user account and go through the setup; you install the programs that you use, some which have licenses on it; you copy your data onto it; you make sure that certain programs always run on startup because that’s the way that you work with these things; you install Firefox because that’s the browser of choice that you go with, et cetera, et cetera. Now, imagine having to do that for, instead of one computer, a thousand of them, and instead of a laptop, they’re servers. And that is directionally what Puppet does.

Nigel: Absolutely. This is the one I use for my mother as well. Like, I was working around Puppet for years before—and the way I explained it was, “You know when you get a new iPad, you’ve got to set up your Facebook account and your email. Imagine you had ten thousand of these.” And she was like—I was like, “You know, companies like Google, company like big banks, they all have lots and lots and lots of computers.” And she was like, “They run all those things on iPads.” And I was like, “This is not really where my analogy was going.” But.

Corey: Right. And increasingly, though, it seems like the world has shifted in some direction where, when you explain that to your mother and she comes back with, “Well, wouldn’t they just put the application into Docker and be done with it?” Oh, dear. But that seems to be in many ways that the direction that the zeitgeist has moved in, whether or not that is the reality in many environments, where when you’re just deploying containers everywhere—through the miracle of Kubernetes—if you’ll pardon the dismissive scorn there, that you just package up your application, shove into a container, and then hurl it from the application team over the operations team, like a dead dog cast into your neighbor’s yard for him to worry about. And then it sort of takes up the space of you don’t have to manage state anymore because everything is mostly stateless in theory. How have you seen it play out in practice in the last five years?

Nigel: I mean, that’s a real trend. And, you know, the size of a container should be [laugh] smaller than an operating system. And the reality is, I’m a sysadmin; I love operating systems, I nerded out on operating systems. They’re a necessary evil, they’re terrible, terrible things: registry keys, config files, they’re a pain in the neck to deal with. And if you look at, I think what a lot of operations folks missed about Docker when it started was that it didn’t make their life better. It was worse.

It was, like, this actual, sort of, terrible toolchain where you sort of tied together all these different things. But really importantly, what it did is it put control into the hands of the developers, and it was the developers who were trying to do stuff who were trying to shift into applications. And I think Docker was a really great technology, in the sense of, you know, developers could ship value on their own. And that was the huge, huge leveling up. It wasn’t the interface, it wasn’t the user experience, it wasn’t all these things, it was just that the control got taken away from the IT trolls in their basement going, “No, don’t touch my servers,” and instead given straight to the developers. And that’s huge because it let us ship things faster. And that’s ultimately the whole goal of things.

Corey: The thing that really struck me the most from conducting the trainings that I did was meeting a whole bunch of people across the country, in different technological areas of specialty, in different states of their evolution as technologists, and something that struck me was just how much people wound up identifying with the technology that they worked on. When someone is the AIX admin, and the AIX machines are getting replaced with Linux boxes, there’s this tendency to fight against that and rebel, rather than learning Linux. And I get it; I’m as subject to this as anyone is. And in many cases, that was the actual pushback that I saw against adopting something like Puppet. If I identify my job as being the person that runs all these carefully curated scripts that I’ve spent five years building, and now that all gets replaced with something that is more of a global solution to my local problem, then it feels like a thing that made me special is eroding.

And we see that with the migration to cloud as well. When you’re the storage admin, and it just becomes an API call to S3, that’s kind of a scary thing. And when you’re one of the server hugger types—and again, as guilty as anyone of this—and you start to see cloud coming in as, like, a rising tide that eats up what it was that you became known for, it’s scary and it becomes a foundational shift in how you view yourself. What I really had a lot of sympathy for was the folks who’ve been doing this for 20 years. They were, in some cases, a few years away from retirement, and they’ve been doing basically the same set of tasks every year for 25 years.

It’s one year of experience repeated 25 times. And they don’t have that much time left in their career, intentionally, so they want to retire, but they also don’t really want to learn a whole bunch of new technologies just to get through those last few years. I feel for them. But at the same time—

Nigel: No, me too, totally. But what are you going to do? But without sounding too dismissive there, I think it’s a natural tendency for us to identify with the technology if that’s what you’re around all the time. You know, mechanics do this, truck drivers with brands of trucks, people, like, to build attachments to the technology they work with because we fit them into this bigger techno-social system. But I have a lot of empathy for the people in enterprise jobs who are being asked to change radically because the cycle of progress is speeding up faster and faster.

And as you say, they might be a few years away from retirement. I think I used to feel more differently about this when I was really hot-headed and much more of a tech enthusiast, and that’s what I identified with. In terms of, it’s okay for a job to just be a job for people. It’s okay for someone to be doing a job because they get good health care and good benefits and it’s feeding their family. That’s an important thing. You can’t expect everyone to always be incredibly passionate about technology choices in the same way that I think many of us who live on Twitter and hanging out in this space are.

Corey: Oh, I have no problem whatsoever with people who want to show up for 40 hours a week-ish, work on their job, and then go home and have lives and not think about computers at all. There’s this dark mass of developers out there that basically never show up on Twitter, they aren’t on IRC, they don’t go to conferences, and that’s fine. I have no problem with that, and I hope I don’t come across as being overly dismissive of those folks. I honestly wish I could be content like that. I just don’t hold still very well.

Nigel: [laugh]. Yeah, so I think you touched on a few interesting things there. And some of those we sort of cover in the State of DevOps Report, which is coming out in the next few weeks.

Corey: Indeed, and the State of DevOps Report started off at Puppet, and they’ve now done it for, what, 10 years?

Nigel: This is the 10th year, which is completely crazy. So, I was looking at the stats as I was writing it, and it’s 10 years of State of DevOps Reports; I think it’s 11 years of DevOps Weekly, Gareth Rushgrove’s newsletter; it’s 12 or 13 years of DevOpsDays that have been going on. This is longer than I spent in primary and high school put together. It’s kind of crazy that the DevOps movement is still, kind of, chugging along, even if it’s not necessarily the coolest kid on the block, now that GitOps, SRE flavor of the month, various kinds of permutations of how we work with technology, have perhaps got a little bit cooler. But it’s still very, very relevant to a lot of enterprises out there.

Corey: Yeah. As I frequently say, legacy is a condescending engineering term for ‘it makes money,’ and there’s an awful lot of that out there. Forget cloud, there are still companies wrestling with do we explore this virtualization thing? And that was something I was very against back in 2006, let’s be very honest. I am very bad at predicting the future of technology.

And, “I can see this for small niche edge workload cases, where you have a bunch of idle servers, but for the most part, who’s really going to use this in production?” Well, basically everyone because that, in turn, is what the cloud runs on. Yeah, I think we can safely say I got that one hilariously wrong. But hey, if you’re aren’t going to make predictions, then what’s it matter?

Nigel: But the industry pushes you in these directions. So, there was this massive bank in Asia who I’ve been working with for a long time and they were always resistant to adopting virtualization. And then it was only four or five years ago that I visited them; they’re like, “Right. Okay. It’s time. We’re rolling out VMware.” And I was like, “So, I’m really curious. What exactly changed in the last year or two in, like, 2014, 2015 that you decided virtualization was the key?” And I’m like—

Corey: Oh, there was this jackwagon who conducted this training? Yeah, no, no, sorry. I can’t take credit for that one.

Nigel: They couldn’t order one rack unit servers with CD drives anymore because their whole process was actually provisioning with CDs before that point.

Corey: Welcome to the brave new world of PXE booting, which is kind of hard, so yeah, virtualization is easier. You know, sometimes people have to be dragged into various ways of technological advancement. Which gets to the real thing I want to cover, since this is a promoted episode, where you’re talking about the State of DevOps Report, I’m almost less interested in what this year’s has to say specifically, than what you’ve seen over the last decade. What’s changed? What was true 10 years ago that is very much not true now? Bonus points if you can answer that without using the word Kubernetes more than twice.

Nigel: So, I think one of the big things was the—we’ve definitely passed peak DevOps team, if you may remember, there was a lot of arguments and there’s still regular, is DevOps a job title? Is it a team title? Is it a [crosstalk 00:14:33]—

Corey: Oh, I was much on the no side until I saw how much more I would get paid as a DevOps engineer instead of a systems administrator for the exact same job. So, you know, I shut up and I took the money. I figured that the semantic arguments are great, but yeah.

Nigel: And that’s exactly what we’ve written in the report. And I think it’s great. The sysadmins, we were unloved. You know, we were in the basement, we weren’t paid as much as programmers. The running joke used to be for developers, DevOps meant, “I don’t need ops anymore.” But for ops people, it was, “I can get paid like a developer.”

Corey: In many cases, “Oh, well, systems administrators don’t want to learn how to code.” It’s, yeah, you’re remembering a relatively narrow slice of time between the modern era, where systems administrator types need to be able to write in the lingua franca of everything—which is, of course, YAML, as far as programming languages go—and before that, to be a competent systems administrator, you needed to have a functional grasp of C. And—

Nigel: Yeah.

Corey: —there is only a limited window in which a bunch of bash scripts and maybe a smidgen of Perl would have carried you through. But the deeper understanding is absolutely necessary, and I would argue, always has been.

Nigel: And this is great because you’ve just linked up with one of the things we found really interesting about the report is that you know when we talk about legacy we don’t actually mean the oldest shit. Because the oldest shit is the mainframes; it’s a lot of bare metal applications. A lot of that in big enterprises—

Corey: We’re still waiting for an AWS/400 to replace some of that.

Nigel: Well, it’s administered by real systems engineers, you know, like, the people who wrote C, who wrote kernel extensions, who could debug things. What we actually mean by legacy is we mean late ’90s to late 2000s, early 2010s. Stuff that was put together by kids who, like me, happened to get a job because you grew up with a computer, and then the dotcom explosion happened. You weren’t necessarily particularly skilled, and a lot of people, they didn’t go through the apprenticeships that mainframe folks and systems engineers actually went through. And everyone just held this stuff together with, you know, duct tape and dental floss. And then now we’re paying the price of it all, like, way back down the track. So, the legacy is really just a certain slice of rapid growth in applications and infrastructure, that’s sort of an unmanageable mess now.

Corey: Oh, here in San Francisco, legacy is anything prior to last night’s nightly build. It’s turned into something a little ridiculous. I feel like the real power move as a developer now is to get a job, go in on day one, rebase everything in the Git repository to a single commit with a message, ‘legacy code’ and then force push it to the main branch. And that’s the power move, and that’s how it works, and that’s also the attitude we wind up encountering in a lot of places. And I don’t think it serves anyone particularly well to tie themselves so tightly to that particular vision.

Nigel: Yep, absolutely. This is a real problem in this space. And one of the things we found in the State of DevOps Report is that—let me back up a little and give a little bit of methodology of what we actually do. We survey people about their performance metrics, you know, like how quickly can you do deploys? What’s your mean time to recovery? Those sorts of things, and what practices do you actually employ?

And we essentially go through and do statistical analysis on this, and everyone tends to end up in three cohorts, they separate pretty easily, of low, medium, and high evolution. And so one of the things we found is that everyone at the low level has all sorts of problems. They have issues with what does my team do? What does the team next to me do? How do I talk to the team next to me?

How do I actually share anything? How do I even know what my goals are? Like, fundamental company problems. But everyone at all levels of evolution is stuck on two big things: not being able to find enough people with the right skills for what they need, and their legacy infrastructure holding them back.

Corey: The thing that I find the most compelling is the idea of not being able to find enough people with the skills that they need. And I’m going to break my own rule and mentioned Kubernetes as a prime example of this. If you are effective at managing Kubernetes in production, you will make a very comfortable living in any geographical location on the planet because it is incredibly complex. And every time we’ve seen this in previous trends, where you need to get more and more complexity, and more and more expertise just to run something, it looks like a sawtooth curve, where at some point that complexity, it gets abstracted away and compressed down into something that is basically a single line somewhere, or it happens below the surface level of awareness. My argument has been that Kubernetes is something no one’s going to care about in roughly three years from now, not because we’re not using it anymore, but because it’s below the level of awareness that we have to think about, in the same way that there aren’t a whole lot of people on the planet these days who have to think about the Linux virtual memory management subsystem. It’s there and a few people really care about it, but for the rest of us, we don’t have to think about that. That is the infrastructure underneath our infrastructure.

Nigel: Absolutely. I used to make a living—and it’s ridiculous looking back at this—for a year or two, doing high-performance custom compiled Apaches for people. Like, I was really really good at this.

Corey: Well yeah, Apache is a great example of this, where back in the ’90s, to get a web server up and running you needed to have three days to spare, an in-depth knowledge of GCC compiler flags, and hope for the best. And then RPM came out and then, okay, then YUM or other things like that—

Nigel: Exactly.

Corey: —on top of it. And then things like Puppet started showing up, and we saw, all right now, [unintelligible 00:20:01] installed. Great. And then we had—it took a step beyond that, and it was, “Oh, now it’s just a Docker-run whatever it is,” and these days, yeah, it’s a checkbox in S3.

Nigel: So, let me get your Kubernetes prediction down, right. So, you’re predicting Kubernetes is going to go away like Apache and highly successful things. It’s not an OpenStack failure state; it’s Apache invisibility state?

Corey: Absolutely. My timeline is a bit questionable, let’s be fair, but—it’s a little on the aggressive side, but yeah, I think that Kubernetes is inherently too complex for most people to have to wind up thinking about it in that way. And we’re not talking small companies; we’re talking big ones where you’re not in a position, if you’re a giant blue-chip Fortune 50, to hire 2000 people who all know Kubernetes super well, and you shouldn’t have to. There needs to be some flattening of all of that high level of complexity. Without the management tools, though, with things like Puppet and the things that came before and a bunch of different ways, we would all not be able to get anything done because we’d be too busy writing in assembly. There’s always going to be those abstractions on top abstractions on top abstractions, and very few people understand how it works all the way down. But that’s, in many cases, okay.

Nigel: That’s civilization, you know? Do you understand what happens when you plug in something to your electricity socket? I don’t want to know; I just want light.

Corey: And more to the point, whenever you flip the switch, you don’t have that doubt in your mind that the light is going to come on. So, if it doesn’t, that’s notable, and your first thought is, “Oh, the light bulb is out,” not, “The utility company is down.” And we talk about the cloud being utility computing.

Nigel: Has someone put a Kubernetes operator in this light switch that may break this process?

Corey: Well, okay, IoT does throw a little bit of a crimp into those works. But yeah. So, let’s talk more about the State of DevOps Report. What notable findings were there this year?

Nigel: So, one of the big things that we’ve seen for the last couple of years has been that most companies are stuck in the middle of the evolutionary progress. And anyone who deals with large enterprises knows this is true. Whatever they’ve adopted in terms of technology, in terms of working methods, you know, agile, various different things, most companies don’t tend to advance to the high levels; most places stay mired in mediocrity. So, we wanted to dive into that and try and work out why most companies actually stuck like this when they hit a certain size. And it turns out, the problems aren’t technology or DevOps, they really fundamental problems like, “We don’t have clear goals. I don’t understand what the teams next to me do.”

We did a bunch of qualitative interviews as well as the quantitative work in the survey with this report, and we talked to one group of folks at a pretty large financial services company who are like, “Our teams have all been renamed so many times, if I need to go and ask someone for something, I literally page up and down through ServiceNow, trying to find out where to put the change request.” And they’re like, “How do I know where to put a network port opening request for this particular service when there are 20 different teams that might be named the right thing, and some are obsolete, and I get no feedback whether I’ve sent it off to the right thing or to a black hole of enterprise despair?”

Corey: I really love installing, upgrading, and fixing security agents in my cloud estate! Why do I say that? Because I sell things, because I sell things for a company that deploys an agent, there's no other reason. Because let’s face it. Agents can be a real headache. Well, now Orca Security gives you a single tool that detects basically every risk in your cloud environment -- and that’s as easy to install and maintain as a smartphone app. It is agentless, or my intro would’ve gotten me into trouble here, but it can still see deep into your AWS workloads, while guaranteeing 100% coverage. With Orca Security, there are no overlooked assets, no DevOps headaches, and believe me you will hear from those people if you cause them headaches. and no performance hits on live environments. Connect your first cloud account in minutes and see for yourself at orca.security. Thats “Orca” as in whale, “dot” security as in that things you company claims to care about but doesn’t until right after it really should have.

Corey: That doesn’t get better with a lot of modernization. I mean, I feel like half of my job—and I’m not exaggerating—is introducing Amazonians to one another. Corporate communication between departments and different groups is very far from a solved problem. I think the tooling can help but I’ve never been a big believer in solving political problems with technology. It doesn’t work. People don’t work that
way.

Nigel: Absolutely. One of my earliest times working at Puppet doing, sort of, higher-level sales and services and support, huge national telco walk in there; we’ve got the development team, the QA team, the infrastructure team. In the course of this conversation, one of them makes a comment about using apt-get, and the others were like, “What do you mean? We’re on RHEL.” And it turned out, production was running on RHEL, the QA team running on CentOS and the developers were all building everything on Ubuntu. And because it was Java wraps, they almost didn’t have to care. But write once, debug everywhere.

Corey: History doesn’t repeat, but it rhymes; before Docker, so much of development in startup-land was how do I make my MacBook Pro look a lot more like an EC2 Linux instance? And it turns out that there’s an awful lot of work that goes into that maybe isn’t the best use of people’s time. And we start to see these breakthroughs and these revelations in a bunch of different ways. I have to ask. This is the tenth year that you’ve done the State of DevOps Report. At this point, why keep doing it? Is it inertia? Are you still discovering new insights every year on top of it? Or is it one of those things where well someone in marketing says we have to do it, so here we are?

Nigel: No, actually, it’s not that at all. So definitely, we’re going to take stock after this year because ten years feels like a really good point to, sort of—it’s a nice round number in certain kind of number system. Mainly the reason is, a lot of my job is going and helping big enterprises just get better at using technology. And it’s funny how often I just get folks going, “Oh, I read this thing,” like people who aren’t on the bleeding edge, constantly discussing these things on Twitter or whatever, but the State of DevOps Report makes its way to them, and they’re like, “Oh, I read a thing there about how much better it is if we standardized on one operating system. And that made a really huge difference to what we were actually doing because you had all this data in there showing that that is better.”

And honestly, that’s the biggest reason why I ended up doing it. It’s the fact that it seems to be a tool that has made its way through to very hard to penetrate enterprise folks. And they’ll read it and managers will read things that are like, “If you set clear goals for your team and get them to focus on optimizing the legacy environment, you will see returns on it.” And I’m being a little bit facetious in the tone that I’m saying because a lot of this stuff does feel obvious if you’re constantly swimming in this stuff day-to-day, but it’s not just the practitioners who it’s just a job for in a lot of big companies. It’s true, a lot of the management chain as well. They’re not necessarily going out and reading up on modern agile IT management practices day-to-day, for fun; they go home and do something else.

Corey: One of my favorite conferences is Gene Kim’s DevOps Enterprise Summit, and the specific reason behind that is, these are very large companies that go beyond companies, in some cases, to institutions, where you have the US Air Force as a presenter one year and very large banks that are 200 years old. And every other conference, it seems, more or less involves people getting on stage, deliver conference-ware and tell stories that make people at those companies feel bad about themselves. Where it’s, “We’re Twitter for Pets, and this is how we deploy software,” or the ever-popular, “This is how Netflix does stuff.” Yeah, Netflix has basically no budget constraints as far as hiring engineering folks go, and lest we forget, their failure mode is someone can’t watch a movie right now. It’s not exactly the same thing as the ATM starts spitting out the wrong balance in the streets.

And I think that there’s an awful lot of discussion where people look at the stories people tell on conference stages and come away feeling bad from it. Very often, I’ll see someone from a notable tech companies talk about how they do things. And, “Wow, I wish my group did things like that.” And the person next to me says, “Yeah, me too.” And I check and they work at the same company.

And the stories we tell are not necessarily the stories that we live. And it’s very easy to come away discouraged from these things. And that goes triply so for large enterprises that are regulated, that have significant downside risk if the technology fails them. And I love watching people getting a chance to tell those stories.

Nigel: Let me jump in on that really quickly because—

Corey: Please, by all means.

Nigel: —one is, you know, having done four years at Google, things are a shitshow internally there, too—

Corey: You’re talking about it like it’s prison. I like it.

Nigel: —you know. [laugh]. People get horrified when they turn up and they’re like, “Oh, what it’s not all gleaming, perfect software artifacts, delivered from the hand of Urs.” But I think what Gene has done with DevOps Enterprise Summit is fantastic in how people share more openly their failure states, but even there—and this is an interesting result we found from a few years ago, State of DevOps Report—even those executives are being more optimistic because it’s so beaten into you as the senior executive; you’re putting on a public face, and even when they’re trying to share the warts-and-all story, they can’t help but put a little bit of a positive spin on it. Because I’ve had exactly the same experience there where someone’s up there telling a war story, and then I look, turn to the person next to me, and they work at that same 300-year-old bank, and they’re like, “Actually, it’s much, much worse than this, and we didn’t fix it quite as well as that.” So, I think the big tech companies have terrible inside unless they’re Netflix, and the big enterprises are also terrible. But they’re also—

Corey: No, no, I’ve talked to Netflix people, too. They do terrible things internally there, too. No one talks about the fact that their internal environments are always tire fires, and there are two stories: the stories we tell publicly, and the reality. And if you don’t believe me on that, look at any company in the world’s billing system. As much as we all talk about agile and various implementations thereof when it comes to
things that charge customers money, we’re all doing waterfall.

Nigel: Absolutely. [laugh].

Corey: Because mistakes show when you triple-charge someone’s credit card for the cost of a small country’s GDP. It’s a problem. I want to normalize those sorts of things more. I’m looking forward to reading this year’s report, just because it’s interesting to see how folks who are in environments that differ from the ones that I get to see experience in this stuff and how they talk about it.

Nigel: Yeah. And so one of the big results I think there for big companies that’s really interesting is that one of the, sort of, anti-patterns is having lots of different types of teams. And I kind of touched on this before about having confusing team titles being a real problem. And not being able to cross organizational boundaries quickly is really, really—you know, it’s a huge inhibitor and cause, source of friction. But turns out the pattern that is actually really great is one that the Team Topologies guys have discovered.

If you’ve been following what Matthew Skelton and Manuel Pais have been doing for a while, they’ve basically been documenting a pattern in software organizations of a small number of team types, of a platform team, value stream teams, complicated subtest system teams, and enabling teams. And so we worked with Manuel and Matt on this year’s report and asked a whole bunch of questions to try and validate the Team Topologies model, and the results came back and they were just incredibly strong. Because I think this speaks to some of the stuff you mentioned before that no one can afford to hire an army of Kubernetes developers, and whatever the hottest technology is in five years, most big companies can’t hire an army of those people either. And so the way you get scale internally before those things become commoditized is you build a small team and create the situation where they can have outsized leverage inside their organization, like get rid of all the blockers to fast flow and make their focus self-service to other people. Because if you’re making all of your developers learn distributed systems operations arcane knowledge, that’s not a good use of their time, either.

Corey: It’s really not. And I think that’s something that gets lost a lot is, I’ve never yet seen a company beyond the very early startup stage, where the AWS bill exceeded the cost of the people working on the AWS bill. Payroll is always a larger expense than infrastructure unless you’re doing something incredibly strange. And, oh, I want to save some money on the cloud bill is very often offset by the sheer amount of time that you’re going to have to pay people to work on that because, contrary to what we believe as engineering hobbyists, people’s time is very far from free. And it’s also the opportunity cost of if you’re going to work on this thing instead of something else, well, is that really the best choice? It comes down to contextualizing what technology is doing as well as with what’s happening over in the world of business strategy. And without having a bridge between those, it doesn’t seem to work very well.

Nigel: Absolutely. It’s insane. It’s literally insane that, as an industry, we will optimize 5%, 3% of our infrastructure bill or application workload and yet not actually reexamine business processes that are causing your people to spend 10% of their time in synchronous meetings. You can save so much more money and achieve so much more by actually optimizing for fast flow, and getting out of the way of the people who cost lots of money.

Corey: So, one last topic that I want to cover before we call it an episode. You talk to an awful lot of folks, and it’s easy to point at the
aspirational stories of folks doing things the right way. But let’s dish for a minute. What are you seeing in terms of people not using the cloud properly? I feel like you might have a story or two on that one.

Nigel: I do have a few stories. So, in this year’s report, one of the things we wanted to find out of, like, are people using the cloud in the way we think of cloud; you know, elastic, consumption-based, all of these sorts of things. We use the NIST metrics, which I recognize can be a little controversial, but I think you’ve got to start somewhere as a certain foundation. It turns out just about everyone is using the public cloud. And when I say cloud, I’m not really talking about people’s internal VMware that they rebadged as cloud; I’m talking about the public cloud providers.

Everyone’s using it, but almost no one is taking advantage of the functionality of the cloud. They’re instead treating it like an on-premise VMware installation from the mid-2000s, they’re taking six weeks to provision instances, they’re importing all of their existing processes, they keep these things running for a long time if they fall over, one person is tasked with, “Hey, do you know how pet number 45 is actually doing here?” They’re not really treating any of these things in the way that they’re actually meant to. And I think we forget about this a lot of the time when we talk about cloud because we jump straight to cloud-native, you know, the sort of bleeding edge of folks in serverless, highly orchestrated containers. I think if you look at the actual numbers, the vast majority of cloud usage, it’s still things like EC2 instances on AWS. And there’s a reason: because it’s a familiar paradigm for people. We’re definitely going to progress past there, but I think it’s easy to leave the people in the middle behind when we’re talking about cloud and how to improve the ecosystem that they all operate in.

Corey: Part of the problem, too, is that whenever we look at how folks are misusing cloud, it’s easy to lose sight of context. People don’t generally wake up and decide I’m going to do a terrible job today unless they work in, you know, Facebook’s ethics department or something. Instead, it’s very much a people are shaped by the constraints they’re laboring under from a bunch of different angles, and they’re trying to do the best with what they have. Very often, the reason that a practice or a policy exists is because, once upon a time, there was a constraint that may or may not still be there, and going forward the way that they have seemed like the best option at the time. I found that the default assumption that people are generally smart and doing the right thing with the information they have carries you a lot further, in many respects than what I did is a terrible junior consultant, which is, “Oh, what moron built this?” Invariably to said moron, and then the rest of the engagement rapidly goes downhill from there. Try and assume good faith, and if you see something that makes no sense, ask, “Why is it like this?” Rather than, “Why is it like this?” Tone counts for a lot.

Nigel: It’s the fundamental attribution bias. It’s why we think all other drivers on the road are terrible, but we actually had a good reason for swerving into that lane.

Corey: “This isn’t how I would have built it. So, it’s awful.”

Nigel: Yeah, exactly.

Corey: Yeah. And in some cases, though, there are choices that are objectively bad, but I tried to understand where they came from there. Company policy, historically, around things like data centers, trying to map one-to-one to cloud often miss some nuances. But hey, there’s a reason it’s called the digital transformation, not a project that we did.

Nigel: [laugh]. And I think you’ve got to always have empathy for the people on the ground. I quite often have talked to folks who’ve got, like, a terrible cloud architecture with the deployment and I’m like, “Well, what happened here?” And they went, “Well, we were prepared to deploy this whole thing on AWS, but then Microsoft’s salespeople got to the CTO and we got told at the last minute we’re redeploying everything on Azure.” And so these people were often—you know, you’re given a week or two to pivot around the decision that doesn’t necessarily make any sense to them.

And there may have been a perfectly good reason for the CTO to do this: they got given really good kickbacks in terms of bonuses for, like, how much they were spending on the infrastructure—I mean, discounts—but people on the ground are generally doing the best with what they can do. If they end up building crap, it’s because our system, society, capitalism, everything else is at fault.

Corey: [laugh]. I have to say, I’m really looking forward to seeing the observations that you wound up putting into this report as soon as it drops. I’m hoping that I get a chance to speak with you again about the findings, and then I can belligerently tell you to justify yourself. Those
are my favorite follow-ups.

Nigel: [unintelligible 00:37:05].

Corey: If people want to get a copy of the report for themselves or learn more about you, where can they find you?

Nigel: Just head straight to puppet.com, and it will be on the banner on the front of the site.

Corey: Excellent. And will, of course, put a link to that in the show notes, if people can’t remember puppet.com. Thank you so much for taking the time to speak with me. I really appreciate it.

Nigel: Awesome. No worries. It was good to catch up.

Corey: Nigel Kersten, field CTO at Puppet. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice as well as an insulting comment telling me that ‘comcastic’ isn’t a funny word, and tell me where you work, though we already know.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

This week Corey is joined by Anurag Gupta, founder and CEO of Shoreline.io. Anurag guides us through the large variety of services he helped launch to include RDS, Aurora, EMR, Redshift and other. The result? Running things almost like a start-up—but with some distinct differences.

Eventually Anurag ended up back in the testy waters of start-ups. He and Corey discuss the nature of that transition to get back to solving holistic problems, tapping into conveying those stories, and what Anurag was able to bring to his team at Shoreline.io where automation is king. Anurag goes into the details of what Shoreline is and what they do. Stay tuned for me.

Links:

  • Shoreline.io: https://shoreline.io
  • LinkedIn: https://www.linkedin.com/in/awgupta/
  • Email: anurag@Shoreline.io

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Your company might be stuck in the middle of a DevOps revolution without even realizing it. Lucky you! Does your company culture discourage risk? Are you willing to admit it? Does your team have clear responsibilities? Depends on who you ask. Are you struggling to get buy in on DevOps practices? Well, download the 2021 State of DevOps report brought to you annually by Puppet since 2011 to explore the trends and blockers keeping evolution firms stuck in the middle of their DevOps evolution. Because they fail to evolve or die like dinosaurs. The significance of organizational buy in, and oh it is significant indeed, and why team identities and interaction models matter. Not to mention weither the use of automation and the cloud translate to DevOps success. All that and more awaits you. Visit: www.puppet.com to download your copy of the report now!

Corey: If your familiar with Cloud Custodian, you’ll love Stacklet. Which is made by the same people who made Cloud Custodian, but put something useful on top of it so you don’t have to be a need to be a YAML expert to work with it. They’re hosting a webinar called “Governance as Code: The Guardrails for Cloud at Scale” because its a new paradigm that enables organizations to use code to manage and automate various aspects of governance. If you’re interested in exploring this you should absolutely make it a point to sign up, because they’re going to have people who know what they’re talking about—just kidding they’re going to have me talking about this. Its doing to be on Thursday, July 22nd at 1pm Eastern. To sign up visit snark.cloud/stackletwebinar and I’ll talk to you on Thursday, July 22nd.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. This promoted episode is brought to you by Shoreline, and I’m certain that we’re going to get there, but first, I’m notorious for telling the story about how Route 53 is in fact a database, and anyone who disagrees with me is wrong. Now, AWS today is extraordinarily tight-lipped about whether that’s accurate or not, so the next best thing, of course, is to talk to the person who used to run all of AWS’s database offerings and start off there and get it from the source. Today, of course, he is not at an Amazon, which means he’s allowed to speak with me. My guest is Anurag Gupta, the founder and CEO of Shoreline.io. Anurag, thank you for joining me.

Anurag: Thanks for having me on the show, Corey. It’s great to be on, and I followed you for a long time. I think of you as AWS marketing, frankly.

Corey: The running gag has been that I am the de facto head of AWS marketing as a part-time gag because I wandered past and saw an empty seat and sat down and then got stuck with the role. I mostly kid, but there does seem to be, at times, a bit of a challenge as far as expressing stories and telling those stories in useful ways. And some mistakes just sort of persist stubbornly forever. One of them is in the list of services, Route 53 shows up as ‘networking and content delivery,’ which I think regardless of the answer, it doesn’t really fit there. I maintain it’s a database, but did you have oversight into that along with Glue, Athena, all the RDS options, managed blockchain—for some reason—as well. Was it considered a database internally, or was that not really how they viewed it?

Anurag: It’s not really how they view it. I mean, certainly there’s a long IP table, right, and routing tables, but I think we characterized it in a whole different org. So, I had responsibility for Analytics, Redshift, Glue, EMR, et cetera, and transactional databases: Aurora, RDS, stuff like that.

Corey: Very often when you have someone who was working at a very large company—and yes, Amazon has a bunch of small teams internally, but let’s face it, they’re creeping up on $2 trillion in valuation at the time of this recording—it’s fairly common to see that startups are, “Oh, this person was at Amazon for ages.” As if it’s some sort of amazing selling point because a company with, what is it, 1.2 million people give or take is absolutely like a relatively small just-founded startup culturally, in terms of resources, all the rest. Conversely, when you’re working at scales like that, where the edge case becomes the common case, and the corner case becomes something that happens 18 times an hour, it informs the way you think about things radically differently. And your reputation does precede you, so I’m going to opt for assuming that this is, rather than being the story about, “Oh, we’re just going to try and turn this company into the second coming of Amazon,”
that there’s something that you saw while you were at AWS that you thought it was an unmet need in the ecosystem, and that’s what
Shoreline is setting out to build. Is that slightly accurate? Or no you’re just basic—there’s a figurehead because the Amazon name is great for getting investors.

Anurag: No, that’s very astute. So, when I joined AWS, they gave me eight people and they asked me to go disrupt data warehousing and transaction processing. So, those turned into Redshift and Aurora, respectively, and gradually I added on more services. But in that sense, Amazon does operate like a startup. They really believe in restricting the number of resources you get so that you have time and you’re forced to think and be creative.

That said, you don’t really wake up at night sweating about whether you’re going to hit payroll. This is, sort of, my fourth startup at this point and there are sleepless nights at a startup and it’s different. I’d go launch a service at AWS and there’ll be 1000 people who are signed up to the beta the next day, and that’s not the way startups work. But there are advantages as well.

Corey: I can definitely empathize with that. My last job before I started this place was at a small scrappy startup which was great for three months and then BlackRock bought us, and then, oh, large regulated finance company combined with my personality ended about the way you think it would. And where, so instead of having the fears and the challenges that I dealt with then, I’m going to go start my own company and have different challenges. And yeah, they are definitely different. I never laid awake at night worrying about how I was going to make payroll, for example.

There’s also the freedom, in some ways, at large companies where whatever function needs to get done, whatever problem you have, there is some department somewhere that handles that almost exclusively, whereas in scrappy startup land, it’s, well, whatever problem needs to get done today, that is your job right now. And your job description can easily fill six pages by the end of month two. It’s a question of trade-offs and the rest. What did you see that gave you the idea to go for startup number four?

Anurag: So, when I joined AWS thinking I was going to build a bunch of database engines—and I’ve done that before—what I learned is that building services is different than building products. And in particular, nobody cares about your performance or features if your service isn’t up. Inside AWS, we used to talk about utility computing, you know, metering and providing compute storage database the way, you know, my local utility provider, PG&E, provides power and gas. And if I call up PG&E and say that the power is out at my house, I don’t really want to hear, “Oh, did you know that we have six nines power availability in the state of California?” I mean, the power is still out; go come over here and fix it. And I don’t really care about fancy new features they’re doing back at the plant. Really, all I care about is cost and availability.

Corey: The idea of utility computing got into that direction, too, in a lot of ways, in some strange nuances, too. The idea that when I flip the light switch, I don’t stop and wonder, is the light going to turn on? You know, until I installed IoT switches and then everything’s a gamble in the wild times again. And if the light doesn’t come on, I assume that the fuse is out, or the light bulb is blown. “Did PG&E wind up dropping service to my neighborhood?” Is sort of the last question that I have done that list. It took a while for cloud to get there, but at this point, if I can’t access something in AWS, my default assumption is that is my local internet, not the cloud provider. That was hard-won.

Anurag: That’s right. And so I think a lot of other SaaS companies—or anybody operating in the cloud—are now working and struggling to get that same degree of availability and confidence to supply to their customers. And so that’s really the reason for Shoreline.

Corey: There’s been a lot of discussion around the idea of availability and what that means for a business outcome where, I still tell the story from time to time that back in 2012 or so, I was going to buy a pair of underpants on amazon.com, where I buy everything, and instead of completing the purchase, it threw one of the great pictures of staff dogs up. Now, if you listen to a lot of reports on availability, then for one day out of the week, I would just not wear underwear. In practice, I waited an hour, tried it again, the purchase went through and it was fine. However, if that happened every third time I tried to make a purchase, I would spend a lot more money at Target.

There has to be a baseline level of availability. That doesn’t mean that your site is never down, period, because that is, in many cases, an unrealistic aspiration and it turns every outage that winds up coming up down the road into an all-hands-on-deck five-alarm fire, which may not be warranted. But you do need to have a certain level of availability that meets or exceeds your customer’s expectations of same. At least that’s the way that I’ve always viewed it.

Anurag: I think that’s exactly right. I also think it’s important to look at it from a customer perspective, not a fleet perspective. So, a lot of people do inward-facing SRE measurements of fleet-wide availability. Now, your customer really cares about the region they’re in, or perhaps even the particular host they’re on. And that’s even more true if they’ve got data. So, for example, an individual database failing, it’ll take a long time for it to come back up elsewhere. That’s different than something more ephemeral, like an instance, which you can move more easily.

Corey: Part of the challenge that I’ve noticed as well when dealing with large cloud providers, a recurring joke has been the AWS status page: it is the purest possible expression of a static site because it never changes. And people get upset when things go down and the status page isn’t updated, but the challenge is when you’re talking about something that is effectively global scale, it stops being a question of is it up or is it down and transitions long before then into how up or how down is it? And things that impact one customer may very well completely miss another. If you’re being an absolutist, it will always be a sea of red, which doesn’t tell people anything useful. Whereas if a customer is down and their site is off, they don’t really care that most other customers aren’t affected.

I mean, on some level, you kind of want everyone to be down because that differs headline risk, as well as if my site is having a problem, it
could be days before someone gets around to fixing a small bug, whereas if everything is down, oh, this will be getting attention very rapidly.

Anurag: That’s exactly right. Sounds like you’ve done ops before.

Corey: Oh, yes. You can tell that because I’m cynical and bitter about everything.

Anurag: [laugh].

Corey: It doesn’t take long working in operationally-focused roles to get there. I appreciate your saying that though. Usually, people say, “Let me guess. You used to be an ops person.” “How can you tell?” “Because your code is garbage,” is the other way that people go down that path.

And yeah, credit where due; they’re not wrong. You mentioned that back when you were in Amazon, you were given a team of eight people and told to disrupt the data warehouse. Yeah, I’ve disrupted the data warehouse as a single person before so it doesn’t seem that hard. But I’m guessing you mean something beyond causing an outage. It’s more about disrupting the space, presumably.

Anurag: [crosstalk 00:10:57].

Corey: And I think, looking back from 2021, it’s hard to argue that Amazon hasn’t disrupted the data warehouse space and fifteen other spaces besides.

Anurag: Yeah, so that’s what we were all about, sort of trying to find areas of non-consumption. So clearly, data was growing; data warehousing was not growing at the same rate. We figured that had to do with either a cost problem, or it had to do with a simplicity problem, or something else. Why aren’t people analyzing the data that they’re collecting? So, that led to Redshift. A similar problem in transaction processing led to Aurora and various other things.

Corey: You also said a couple of minutes ago that Amazon tends to talk more about features than they do about products, and building a product at a startup is a foundationally different experience. I think you’re absolutely on to something there. Historically, Amazon has folks get on stage at re:Invent and talk about this new thing that got released, and it feels an awful lot like a company saying, “Yeah, here’s some great bricks you can use to build a house.” “Well, okay. What kind of house can I build with those bricks?” “Here to talk about the house that they built as our guest customer speaker from Netflix.”

And it seems like they sort of abdicated, in many respects, the storytelling portion to a number of their customers. It is a very rare startup that has the luxury of being able to just punt on building a product and its product story that goes along with it. Have you found that your time at Amazon made storytelling something that you wound up missing a bit more, or retelling stories internally that we just don’t get to see from the outside, or is, “Oh, wow. I never learned to tell a story before because at Amazon, no one does that, and I have to learn how to do that now that I’m at a startup again?”

Anurag: No, I think it really is a storytelling experience. I mean, it’s a narrative-based culture there, which is, in many ways, a storytelling experience. So, we were trying to provide a set of capabilities so that people could build their own things, you know, much as Kindle allows people to self-publish books; we’re not really writing books of our own. And so I think that was the experience there. Outside, you are trying to solve more holistic problems, but you’re still only a puzzle piece in the experience that any given customer has, right? You don’t satisfy all of their needs, you know, soup to nuts.

Corey: And part of the challenge too, is that if I’m a small, scrappy startup, trying to get something out the door for the first time, the problems that I’m experiencing and the challenges that I have are radically different than something that has attained hyperscale and now has whole optimization stories or series of stories going on. It’s, will this thing even work at all is my initial focus. And in some ways, it feels like conference-ware cuts against a lot of that because it’s hard not to look at the aspirational version of events that people tell on stage at every event I’ve ever seen, and not come away with a takeaway of, “Oh. What I’ve built is actually terrible, and depressing, and sad.” One of the things that I find that resonates about what you’re building over at Shoreline is, it’s not just about the build things from scratch and get them provisioned for the first time. It’s about the ongoing operationalization, I think—if that’s a word—about that experience, and how to wind up handling the care and feeding of something that exists and is running, but is also subject to change because all things are continually being iterated on.

Anurag: That’s right. I feel like operation is sort of an increasingly important but underappreciated part of the service delivery experience much as, maybe, QA was a couple of decades ago. And over time we’ve gone and we built pipelines to automate our test infrastructure, we have deployment tools to deploy it, to configure it, but what’s weird is that there are two parts of the puzzle that are still highly manual: developing software and operating that software in production. And the other thing that’s interesting about that is that you can decide when you are working on developing a piece of code, or testing it, or deploying it, or configuring it. You don’t get to decide when the disk goes down or something breaks. That’s why you have 24/7 on-call.

And so the whole point of Shoreline is to break that into two problems: the things that are automatable, and make it easy, as trivial to automate those things away so you don’t wake up to do something for the tenth time; and then for the remaining things that are novel, to make diagnosing and repairing your fleet, as simple and straightforward as diagnosing and repairing a single box. And we do a lot of distributed systems [techs 00:16:01] underneath the covers to make that the case. But those are the two things that we do, and so hopefully that reduces people’s downtime and it also brings back a lot of time for the operators so they can focus on higher-value things, like working with you to reduce their AWS bill.

Corey: Yeah, for better or worse, working on the AWS bill is always sort of a backseat function, or a backburner function, it’s never the burning priority unless things have gone seriously awry. It’s a good governance thing; it’s the idea of where, let’s optimize this fixed unit economics. It is rarely the number one most pressing area of business for a company. Nor should it be; I think people are sometimes surprised to hear me say that. You want to be reasonable stewards of the money entrusted to you and you obviously want to continue to remain in business by not losing money on everything you sell, but trying to make it up in volume. But at some point, it’s time to stop cutting and focus instead on revenue growth. That is usually the path to success for almost every company I’ve ever spoken to, unless they are either very out of kilter, or in a very strange spot in the industry.

Anurag: That’s true, but it does belong, I think, in the ops function to do optimization of your experience, whether—and, you know, improving your resources, improving your security posture, all of those sorts of things fall into production ops landscape, from my perspective. But people just don’t have time for it because their fleets are growing far, far faster than their headcount is. So, the only solution to that is automation.

Corey: And I want to talk to you about that. Historically, the idea has been that you have monitoring—or observability these days, which I consider to be hipster monitoring—figuring out what’s going on in your environment. Then you wind up with incidents being declared when certain things wind up triggering, which presumably are things that actually matter and not, you’re waking someone up for vague reasons like ‘load average is high on these nodes,’ which tells you nothing in isolation whatsoever. So, you have the incident management portion of that [next 00:18:03], and that handles a lot of the waking folks up and getting everyone onto the call. You’re focusing on, I guess, a third tranche here, which is the idea of incident automation. Tell me about that.

Anurag: That’s exactly right. So, having been in the trenches, I never got excited about one more dashboard to look at, or someone routing a ticket to the right person, per se, because it’ll get there, right?

Corey: Oh, yeah. Like, one of the most depressing things you’ll ever see in a company is the utilization numbers from the analytics on the dashboards you build for people. They look at them the day you build them and hand it off, and then the next person visiting it is you while running this report to make sure the dashboard is still there.

Anurag: Yeah. I mean, they are important things. I mean, you get this huge sinking feeling something is wrong and your observability tool is also down like CloudWatch was in some large-scale events. Or if your ticketing system is down and you don’t even notify somebody and you don’t even know to wake up. But what did excite me—so you need those things; they’re necessary, but they’re not sufficient.

What I think is also needed is something that actually reduces the number of tickets, not just lets you observe them or find the right person to act upon it. So, automation is the path to reducing tickets, which is when I got excited because that was one less thing to wake up on that gave me more time back to wo—do things, and most importantly, it improved my customer availability because any individual issue handled manually is going to take an hour or two or three to deal with. The issue being done by a computer is going to take a few seconds or a few minutes. It’s a whole different thing. It’s the difference between a glitch and having to go out on an apology tour to your customers.

Corey: I really love installing, upgrading, and fixing security agents in my cloud estate! Why do I say that? Because I sell things, because I sell things for a company that deploys an agent, there's no other reason. Because let’s face it. Agents can be a real headache. Well, now Orca Security gives you a single tool that detects basically every risk in your cloud environment -- and that’s as easy to install and maintain as a smartphone app. It is agentless, or my intro would’ve gotten me into trouble here, but it can still see deep into your AWS workloads, while guaranteeing 100% coverage. With Orca Security, there are no overlooked assets, no DevOps headaches, and believe me you will hear from those people if you cause them headaches. and no performance hits on live environments. Connect your first cloud account in minutes and see for yourself at orca.security. Thats “Orca” as in whale, “dot” security as in that things you company claims to care about but doesn’t until right after it really should have.

Corey: Oh, yes. I feel like those of us who have been in the ops world for long enough, we always have a horror story or to have automation around incidents run amok. A classic thing that we learned by doing this, for example, is if you have a primary and a secondary, failover should be automated. Failing back should not be, or you wind up in these wonderful states of things thrashing back and forth. And in many cases in data center land, if you have a phantom router ready to step in, if the primary router goes offline, more outages are caused by a heartbeat failure between those two devices, and they both start vying for power.

And that becomes a problem. Same story with a lot of automation approaches. For example, if oh, every time a disc winds up getting full, all right, we’re going to fire off something automatically expand the volume. Well, without something to stop that feedback loop, you’re going to potentially wind up with an unbounded growth problem and then you wind up with having no more discs to expand the volume to, being the way that winds up smacking into things. This is clearly something you’ve thought about, given that you have built a company out of this, and
this is not your first rodeo by a long stretch. How do you think about those things?

Anurag: So, I think you’re exactly right there, again. So, the key here is to have the operator, or the SRE, define what needs to happen on an individual box, but then provide guardrails around them so that you can decide, oh, a lot of these things have happened at the same time; I’m going to put a rate limiter or a circuit breaker on it and then send it off to somebody else to look at manually. As you said, like failover, but don’t flap back and forth, or limit the number of times, but something is allowed to fail before you send it [unintelligible 00:21:44]. Finally, everything grounds that a human being looking at something, but that’s not a reason not to do the simple stuff automatically because wasting human intelligence and time on doing just manual stuff again, and again, and again, is pointless, and also increases the likelihood that they’re going to cause errors because they’re doing something mundane rather than something that requires their intelligence. And so that also is worse than handing it off to be automated.

But there are a lot of guardrails that can be put around this—that we put around it—that is the distributed systems part of it that we provide. In some sense, we’re an orchestration system for automation, production ops, the same way that other people provide an orchestration system for deployments, and automated rollback, and so forth.

Corey: What technical stacks do you wind up supporting for stuff like this? Is it anything you can effectively SSH into? Does it integrate better with certain cloud providers than others? Is it only for cloud and not for folks with data center environments? Where do you start? Where do you stop?

Anurag: So, we have started with AWS, and with VMs and Kubernetes on AWS. We’re going to expand to the other major cloud providers later this year and likely go to VMware on-prem next year. But finally, customers tell us what to do.

Corey: Oh, yeah. Looking for things that have no customer usage is—that’s great and all, but talking to folks who are like, “Yeah, it’d be nice if it had this.” “Will you buy it if it does?” “No.” “Yeah, let’s maybe put that one on the backlog.”

Anurag: And you’ve done startups, too, I see that.

Corey: Oh, once or twice. Talk to customers; I find that’s one of those things that absolutely is the most effective use of your time you can do. Looking at your site—Shoreline.io for those who want to follow along at home—it lists a few different remediations that you give as examples. And one of them is expanding disk volumes as they tend to run out of space. I’m assuming from that perspective alone, that you are almost certainly running some form of Agent.

Anurag: We are running an Agent. So, part of that is because that way, we don’t need credentials so that you can just run inside the customer environment directly and without your having to pass credentials to some third party. Part of it is also so you can do things quickly. So, every second, we’ll scrape thousands of metrics from the Prometheus exporter ecosystem, calculate thousands more, compare them against hundreds of alarms, and then take action when necessary. And so if you run on-box, that can be done far faster than if you go on off-box.

And also, a lot of the problems that happen in the production environment are related to networking, and it’s not like the box isn’t accessible, but it may be that the monitoring path is not accessible. So, you really want to make sure that the box can protect itself even if there’s some issues somewhere in the fleet. And that really becomes an important thing because that’s the only time that you need incident automation: when something’s gone wrong.

Corey: I assume that Agent then has specific commands or tasks it’s able to do, or does it accept arbitrary command execution?

Anurag: Arbitrary command execution. Whatever you can type in at the Linux command prompt, whether it’s a call to the AWS CLI, Kube control, Linux commands like top, or even shell scripts, you can automate using Shoreline.

Corey: Yeah. That was one of the ways that Nagios got it wrong, once upon a time, with their NRP, their Nagios Remote Plugin engine, where you would only be allowed to run explicit things that had been pre-approved and pushed out to things in advance. And it’s one of the reasons, I suspect, why remediation in those days never took off. Now, we’ve learned a lot about observability and monitoring, and keeping an eye on things that have grown well beyond host-based stuff, so it’s nice to see that there is growth in that. I’m much more optimistic about it this time around, based upon what you’re saying.

Anurag: I hope you’re right because I think the key thing also is that I think a lot of these tools vendors think of themselves as the center of the universe, whereas I think Shoreline works the best if it’s entirely invisible. That’s what you want from a feedback control system, from a automation system: that it just give you time back and issues are just getting fixed behind the scenes. That’s actually what a lot of AWS is doing behind the scenes. You’re not seeing something whenever some rack goes down.

Corey: The thing that is always taken me back—and I don’t know how many times I’m going to have to learn this lesson before it sticks—I fall into the common trap of take any one of the big internationally renowned tech companies, and it’s easy to believe that oh, everything inside is far future wizardry of, everything works super well, the automation is flawless, everything is pristine, and your environment compared to that is relative garbage. It turns out that every company I’ve ever spoken with and taken SREs from those companies out to have way too many drinks until they hit honesty levels, they always talk about it being a sad dumpster fire in a bunch of different ways. And we’re talking some of the companies that people laud as the aspirational, your infrastructure should be like these companies. And I find it really important to continue to socialize that point, just because the failure mode otherwise is people think that their company just employs terrible engineers and if people were any good, it would be seamless, just like they say on conference stages. It’s like comparing your dating life to a romantic comedy; it’s not an accurate depiction of how the world works.

Anurag: Yeah, that’s true. That said, I’d say that, like, the average DBA working on-prem may be managing a hundred databases; the average DBA in RDS—or somebody on call—might be managing a hundred thousand.

Corey: At that point, automation is no longer optional.

Anurag: Yeah. And the way you get there is, every week you squash and extinguish one thing forever, and then you start seeing less and less frequent things because one in a million is actually occurring to you. But if it was one in a hundred, that would just crush you. And so you just need to, you know, very diligently every week, every day, remove something. Yeah, Shoreline is in many ways the product I wish I had had at AWS because it makes automating that stuff easy, a matter of minutes, rather than months. And so that gives you the capability to do automation. Everyone wants automation, but the question is, why don’t they do it? And it’s just because it takes so much time and we’re so busy, as operators.

Corey: Absolutely. I don’t mean to say that these large companies working at hyperscale have not solved for these problems and done truly impressive things, but there’s always sharp edges, there’s always things that are challenging and tricky. On this show, we had Dr. Christina Maslach recently as an expert on burnout, given that she spent her entire career studying occupational burnout as an academic. And it turns out that it’s not—to equate this to the operations world—it’s not waking up at two in the morning to have to fix a problem—generally—that burns people out. It’s being woken up to fix a problem at 2 a.m. consistently, and it’s always the same problem and nothing ever seems to change. It’s the worst ops jobs I’ve ever seen are the ones where you have to wake up to fix a thing, but you’re not empowered to actually fix the cause, just the symptom.

Anurag: I couldn’t agree more and that’s the other aspect of Shoreline is to allow the operators or SREs to build the remediations rather than just put a ticket into some queue for some developer to get prioritized alongside everything else. Because you’re on the sharp edge when you’re doing ops, right, to deal with all the consequences of the issues that are raised. And so it’s fine that you say, “Okay, there’s this memory leak. I’ll create a ticket back to dev to go and fix it.” But I need something that helps me actually fix it here and now. Or if there’s a log that’s filling up my disk, it’s fine to tell somebody about it, but you have to grow your disk or move that log off the disk. And you don’t want to have to wake up for those things.

Corey: No. And the idea that everything like this gets fixed is a bit of a misnomer. One of my hobbies is whenever a site goes down and it is uncovered—sometimes very publicly, sometimes in RCEs—that the actual reason everything broke was due to an expired certificate.

Anurag: Yep.

Corey: I like to go and schedule out a couple of calendar reminders on that one for myself, of check it in 90 days, in case they’re using a refresh from Let’s Encrypt, and let’s check it as well in one year and see if there’s another outage just like that. It has a non-zero success rate because as much as we want to convince ourselves that, oh, that bit me once, and I’ll never get bitten like that again, that doesn’t always hold true.

Anurag: Certificates are a very common source of very widespread outages. And it’s actually one of the remediations we provide out of the box. So, alongside making it possible for people to create these things quickly, we also provide what we call Op Packs, which are basically getting started things which have the metrics, alarms, actions, bots, so they can just fix it forever without actually having to do very much other than review what we have done.

Corey: And that’s, on some level, I think, part of the magic is abstracting away the toil so that people are left to solve interesting problems and think about these things, and guiding them down a path where, okay, what should I do on an automatic basis if the disk fills up? Well, I should extend the volume. Yeah. But maybe you should alert after the fifth time in an hour that you have to extend the same volume because—just spitballing here—maybe there’s a different problem here that putting a bandaid on isn’t going to necessarily solve. It forces people to think about what are those triggers that should absolutely result in human intervention because you don’t necessarily want to solve things like memory leaks, for example, oh our application leaks memory so we have to restart it once a day.

Now, in practice, the right way to solve that is to fix the application. In practice, there are so many cron jobs out there that are set to restart things specifically for that reason because cron jobs are quick and easy and application developer time is absolutely not easy to come by in many of these shops. It just comes down to something that helps enforce more of a process, more of a rigor. I like the idea quite a bit; it aligns both with where people are and how a better tomorrow starts to look. I really do think you’re onto something here.

Anurag: I mean, I think it’s one of these things where you just have to understand it’s not either-or, that it’s not a question of operator pain or developer pain. It’s, let’s go and address it in the here and now and also provide the information, also through an automated ticket generation, to where someone can look to fix it forever, at source.

Corey: Oh, yeah. It’s always great of the user experience, too. Having those tickets created automatically is also sometimes handy because the worst way to tell someone you don’t care about their problem when they come to you in a panic is, “Have you opened a ticket?” And yes, of course, you need a ticket to track these things, but maybe when someone is ghost pale and scared to death about what they think just broke the data, maybe have a little more empathy there. And yeah, the process is important, but there should be automatic ways to do that. These things all have APIs. I really like your vision of operational maturity and managing remediation, in many cases, on an automatic basis.

Anurag: I think it’s going to be so much more important in a world where deployments are more frequent. You have microservices, you have multiple clouds, you have containers that give a 10x increase in the number of things you have to manage. There’s a lot for operators to have to keep in their heads. And things are just changing constantly with containers. Every minute, someone comes and one goes. So, you just really need to—even if you’re just doing it for diagnosis, it needs to be collecting it and putting it aside, is really critical.

Corey: If people want to learn more about what you’re building and how you think about these things, where can they find you?

Anurag: They can reach out to me on LinkedIn at awgupta, or of course, they can go to Shoreline.io and reach out there, where I’m also anurag@Shoreline.io if they want to reach out directly. And we’d love to get people demos; we know there’s a lot of pain out there. Our mission is to reduce it.

Corey: Thank you so much for taking the time to speak with me today. I really appreciate it.

Anurag: Yeah. This was a great privilege to talk to you.

Corey: Anurag Gupta, CEO and founder of Shoreline.io. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with a comment telling me that I’m wrong and that Amazonians are the best at being on call
because they carry six pagers.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Michael

Michael Garski is the Director of Platform Engineering at Fender Musical Instruments, where he leads the teams responsible for service development & testing, devops, and data. He’s been with Fender for over 5 years and prior to that worked as a software engineer & architect on back-end systems at Viant, MySpace, Countrywide Home Loans & Fandango. He is passionate about application reliability and observability and their impact on customer satisfaction.

Links:

  • LinkedIn: https://www.linkedin.com/in/mgarski/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Your company might be stuck in the middle of a DevOps revolution without even realizing it. Lucky you! Does your company culture discourage risk? Are you willing to admit it? Does your team have clear responsibilities? Depends on who you ask. Are you struggling to get buy in on DevOps practices? Well, download the 2021 State of DevOps report brought to you annually by Puppet since 2011 to explore the trends and blockers keeping evolution firms stuck in the middle of their DevOps evolution. Because they fail to evolve or die like dinosaurs. The significance of organizational buy in, and oh it is significant indeed, and why team identities and interaction models matter. Not to mention weither the use of automation and the cloud translate to DevOps success. All that and more awaits you. Visit: www.puppet.com to download your copy of the report now!

Corey: If your familiar with Cloud Custodian, you’ll love Stacklet. Which is made by the same people who made Cloud Custodian, but put something useful on top of it so you don’t have to be a need to be a YAML expert to work with it. They’re hosting a webinar called “Governance as Code: The Guardrails for Cloud at Scale” because its a new paradigm that enables organizations to use code to manage and automate various aspects of governance. If you’re interested in exploring this you should absolutely make it a point to sign up, because they’re going to have people who know what they’re talking about—just kidding they’re going to have me talking about this. Its doing to be on Thursday, July 22nd at 1pm Eastern. To sign up visit snark.cloud/stackletwebinar and I’ll talk to you on Thursday, July 22nd.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. We talk to a lot of people here on this show who are deep in the weeds of SaaS companies, or cloud vendors, or cloud vendors cosplaying as SaaS companies. Today, we’re taking a bit of a different direction. My guest is Michael Garski, Director of Platform Engineering at Fender Musical Instruments. They make guitars among many other things. Michael, thank you for joining me.

Michael: Oh, thanks for having me on, Corey.

Corey: So, one of the things that I really appreciate about what you do as a company is I can, at least presumably, explain it to someone who is not super deep in technical weeds without 45 minutes of explainer first. The easy answer is, “Oh, Fender. You folks make guitars.” These days, no one just does one thing, I have to imagine. How do you describe what the company does?

Michael: Oh, well, to quote Leo Fender, his view was that artists are angels and it’s our job to give them wings. So, in addition to actually making and developing guitars and amplifiers, we’ve branched off into consumer-facing products to actually teach people how to play those instruments.

Corey: You folks have been relatively outspoken about the various things you’re doing at different AWS events. I mean, my approach to that tends to be that if AWS is great at making bricks that you can use to build amazing things with, “Well, great, can you draw a picture of the house that you can build with this?” “No, we’re going to have a customer come out and talk about that stuff instead.” You folks have been focusing on a lot of serverless work, and you’ve been very public about the fact that you are almost entirely serverless-driven in terms of architecture if I’m not mistaken.

Michael: That is true.

Corey: Tell me about that. How did you get there and what brought it about?

Michael: So, I work in the digital division in Fender. We started, let’s see, we’re coming up on five years I’ve been there. So, what we did was, initially, we started building services that could run within a container, or on an EC2 instance, but we started looking at Lambda functions. We had need to ingest a product catalog, so the IT team was able to drop us off a product catalog into an S3 bucket, and the easiest thing to do then was just trigger a Lambda function to then process that file. And it just kind of snowballed in from there.

Corey: I think the common problem when people hear ‘serverless’ is they think, “Oh, great. More discussions about Lambda functions.” And Lambda is almost getting something of a tarred reputation in some circles because when we can build amazing things with it ourselves, we love it, but when we ask AWS how to wind up integrating two services, or about a feature gap, their response is, “Oh, use a Lambda function for it,” It starts to feel like they’re using it as spackle and the spackle has become load-bearing. Do you view serverless as being purely function-driven or is it broader than that?

Michael: It’s much broader than that. Serverless is a mindset where you’re looking beyond just Lambda functions to using a lot of third-party services so that you can actually focus on your core business. Like, we use Zuora as a subscription provider for web-based subscriptions; we use Algolia for full-text search; we use a variety of other services so that we can just focus on the core business.

Corey: One thing that’s been on everyone’s mind, somewhat recently, has been the idea of dramatic changes as far as user behavior goes. And in the more traditional environments where you see things like EC2 instances or on-premises data centers, back when the pandemic first hit and companies that were very focused on a model of business that aligned directly with people behaving in certain ways that they suddenly didn’t, would the 80% drop-offs or more in their user traffic, but their infrastructure spend just kept hanging out exactly where it was, in a straight line. So, at some level, it feels like yes, the whole point of cloud is that it can be elastic, except no one builds it that way for a variety of reasons. When COVID hit, what changed for your business?

Michael: Change for our business is we launched a program called Playthrough, okay we did this about a year ago; we started it, we gave away three months of Fender Play for free. It was a single-use code that a user would redeem and no credit card required, and over a period of five days, we saw our traffic increase by more than ten times. And we had very little changes we needed to make. Everything scaled up, we had no issue with—we used a lot of Lambda functions, DynamoDB, everything just scaled up fine. The only point that became a bottleneck was our Elasticsearch cluster. However, beefing up the nodes and adding a few more nodes that resolved that issue immediately.

Corey: So, I’m going to go out on a limb and postulate that you folks increased pickup when the lockdowns hit, if for no other reason then, “Well, I’m trapped at home and I’m tired of staring at the guitar on the wall. I may as well learn to play it.” I would guess. I could be way off base on that.

Michael: No, no, that’s very true. Even since then, even after that program has expired—of course, not everyone then converts and sticks around—but many, many did, many more than we thought would did stick around, and our usage and our goals were exceeded for this last year, and we’re in a healthy place, and looking at continuing to grow and expand in the future.

Corey: So, one of the applications that I think gets a fair bit of attention—rightfully so—lately, is something called Fender Play, and as best I can tell, that is a app that works in web, it works on mobile, and it’s a video-based instruction tool for guitar at least, but some other instruments as well. How did that come to be? Did that exist before COVID hit? Has that been something that’s been in the works for a while? Or was it, “Well, we’re going to do a two-week sprint and build this thing from scratch?”

Michael: No, we launched that—this June we’re coming up on the fourth anniversary since it’s been launched, so we launched this in summer of 2017.

Corey: One of the problems I’ve always found is that it’s challenging to learn to do something that is as, I guess, physical and intricate, et cetera, as playing an instrument without having someone in the room looking at you and smacking you with a stick whenever you do things that are wrong. “Nope, that’s a bad habit. If you keep doing that it’s going to hurt you.” How do you approach that as a company from a non-interactive perspective of someone who’s going to watch a video and do things and maybe it’ll work, maybe it won’t? Particularly in light of things like, well, the competition is YouTube, which, you know, I’m going to roll the dice and sometimes I’ll see a great tutorial, sometimes I’ll see one that I don’t realize teaching me terrible things, and then it’s going to recommend some baseless conspiracy theory because YouTube. How do you differentiate that? What makes Fender Play different?

Michael: So currently, you’re right; it’s just a video-based instruction app. There’s not any way to, like, provide direct feedback to students within the web and mobile applications. However, we do have an online community, and our Fender Play instructors do an office hours feature, is where they’ll actually answer questions live and talk to students. We are investigating and doing some earlier research in some, possibly, being able to provide that type of feedback to users, but it’s very challenging problem, just due to the nature of you’re playing an instrument that has multiple strings, so you’re trying to pick out the chord that they’re playing in, and the timing. But it’s something we definitely need to add.

Corey: There’s something to be said as well for the kind of care and attention that you folks wind up putting into your media where, “This is how you finger a chord,” and someone on the YouTube video will do it for two-tenths of a second, and they’re filming it with a potato that isn’t focused properly and pointing at the wrong part of the guitar. You folks have a high bar for quality on this. Is that done in-house? Do you wind up just going through a bunch of random folks that you just wind up offering a bunch of gift cards to, or free guitars to do this? How does the program work on the back end?

Michael: So, we have an in-house curriculum team that puts together the lesson plans to really help people learn in small bite-sized lessons so that it’s not too overwhelming at once. And that curriculum then is shot and filmed by an in-house video team that put that together; they upload the data into S3 for the final cut, then that gets transcoded via MediaConvert, and we serve it up via CloudFront.

Corey: It’s rare to wind up talking to a company that is something of a household name about something that they’re doing, and hear the AWS services that they’re using not trend toward a baseline mean if I can be so bold. Normally, you’ll see some of the case studies, like, “Oh, this is an online bank. What services are they using?” “Oh, they’re using EC2, and S3, and load balancing because did you miss the part where it’s a bank?” They’re not going to use these far-future services due to regulatory risk, among other things, in many cases.

You’re using Elemental MediaConvert, which is one of those relatively high-up-the-stack offerings that isn’t broadly known. It’s one of those services that is focused on specific use cases and specific industry verticals in a way that a baseline primitive service isn’t. What does MediaConvert do?

Michael: What it does is it takes the final edit of the video, and we have several different presets so that it will put it into an HLS format with different bitrates so that the user is getting the best quality video depending on their bandwidth.

Corey: When I looked into it in the early days when it was first launching, I found that it looked an awful lot like Elastic Transcoder, which is a service that they’ve had for a while, only they changed up some of the capabilities. It’s obviously far more capable as a service, but they also added something that felt like 15 different billing dimensions to it, “So, what is this going to cost me?” “Well, we’re going to run it for a month and find out if we’re still in business.” And it seemed like it was one of those very difficult to get started with and run experiments with service. Now, obviously, services evolve over time. When you started looking into it was that experience roughly akin to what you felt, or am I completely and unfairly slandering in the product?

Michael: We actually started out using Elastic Transcoder and then moved over to MediaConvert, I believe it was last year. We found it to be a little bit easier to use, and the pricing overall in transcoding the videos for us is really a drop in the bucket as compared to actually hosting them and serving them up via CloudFront. And when we switched over to MediaConvert, we adjusted our settings to lower the maximum bitrate for a given video, we found that after a certain point, the quality to the user just doesn’t really improve, and yet we’re paying to serve the larger video.

Corey: One statistic that I found was that in March of 2020—you know which I believe we’re still in at this point; just, it’s the Endless September model, applied to March—you wound up seeing over an order of magnitude in traffic increase within five days, and looking at that through a lens of traditional architecture, that means that nobody sleeps a whole heck of a lot. Given that you’re in on the serverless story, and you have been since before that hit, what was that scaling experience like for you?

Michael: Scaling experience was completely seamless. We use a lot of Lambda, DynamoDB, Kinesis, SNS, to glue things together, and no problems whatsoever. Just had to bump up our Elasticsearch cluster a bit, that was really the only thing because we saw some latency starting to rise on some of our APIs.

Corey: Let me ask the uncomfortable question then because whenever I tried to scale things up quickly in a cloud environment, what was your experience with smacking into various AWS service limits as the traffic grew?

Michael: Initially, we actually requested some service limits increase to make sure we weren’t hitting the concurrent Lambda invocation limit, and same thing with Cognito, making sure that we weren’t going to hit any limits as far as sign-ins and things like that. So, we were able to just put in requests, and they served us around pretty quick turnaround time on that, as well.

Corey: It really does seem like there’s a strong benefit on the serverless space, but I had to double-check before we started recording that you do, in fact, work at Fender because you are a staunch advocate for observability. And usually, when someone is that passionate about observability, you can guess that they work at an observability-slash-monitoring company. It’s akin to the idea of someone selling mattresses telling you that mattresses are great and you should have four of them. You’re on the customer side of that and still very passionate about it. Where’d that come from?

Michael: Came from my time years ago, when I worked at MySpace—if anyone can still remember that—working on the search systems there. And as the company started winding down, to laying people off, and being one of the only people left working on those systems, being able to know and understand them, you just have to, so you have to continue to monitor and find ways to monitor, and that really ingrained how important instrumentation is and being able to really understand the health of your application as it’s running so that you can see, yes, everything is good, and then when something doesn’t look right so that you can know where to start looking, and you can be alerted of a problem.

Corey: So, I tend to view the world in olden terms where monitoring was what we did, and we use something like Nagios, which was the second-worst option out there because everything else felt like it was tied for first. I also take a somewhat regressive view that observability is to monitoring as DevOps is to being a systems administrator. It’s the same thing, but by using the more modern terminology, you can charge more for it. I’m going to go out on a limb and guess that you take a somewhat contrarian [laugh] view to that.

Michael: Yes, yes, I do. It’s about really understanding how your applications is running. It’s not just looking at, oh, how many HTTP 500s am I serving up per hour, if I hit a threshold for the last hour? It’s a lot more than that. It’s really being able to really dig in and see what the issue is or what’s working really well.

And to that end, we rely on two services for this. We use Honeycomb and Epsagon. Honeycomb, kind of, acts as our top layer because it gives us the really good high-cardinality metrics where I can punch in a user ID and I can see all the API traffic that this user has performed. As well as, even just like when we launched the Playthrough when our traffic rose, that the reason we discovered that our latency was dropping was due to a service-level objective being triggered in Honeycomb on latency. And we were able to respond to that using that before customers really noticed anything at all.

Corey: As an Epsagon customer myself, I’m always conflicted when I find myself going into their service and using it to figure out what the heck’s going on with my giant pile of Lambda functions, and API gateways, and whatnot, wired together because the experience is uniformly excellent, but I’m also frustrated in that it needs a third-party to even begin to allude to what’s going on. It feels, on some level, like the vendor that is providing this service to me should be reasonably effective at telling me what it’s doing, and when it’s breaking. I understand that how I wish the world is and how it actually is are two radically different things but does that ever strike you as well?

Michael: Whether or not AWS should be providing that type of level, that seems… that seems like more of a service that you can have competition and other vendors that really specialize and get in the weeds on it. I don’t think AWS needs to provide every service you could possibly use for your application. That’s not something I’m too concerned about. I don’t really even think it’s their place, frankly.

Corey: No, no, I understand. The problem I keep running into, on some level, whenever I try and diagnose it natively is, I look at CloudWatch and it’s difficult to understand that is this—in my case because again, I’m still early days with a lot of these things—is it the API gateway that’s having the problem? Is it the CloudFront distribution that is tied to that? Is it the Lambda function? Where’s the handoff?

Trying to understand where in a complicated application the failure is occurring is a challenge. And let’s be clear, most of that is a problem of my own making because I didn’t have the good sense to instrument this thing in a reliable repeatable way when I built it. It feels like everything is tied together with duct tape, and baling wire, and spit, and a bit of luck. As a counterpoint, the more companies I talk to, the more I realize that no, no, this is actually how most people feel [laugh] when they look at things that are working. It’s, yeah, it’s terrible. It’s a trash fire, but it makes money so we’re going to roll with it.

And there’s always, on some level, a sense of what we’ve built is very far from the platonic ideal of what we should have built. Does that resonate with you, or do you take a step back and look at what you’ve achieved with a perspective of, “This is awesome. More people should do it exactly like this.” And honestly, if it’s that one, I’d love to take a look at what you’ve built.

Michael: I think there’s always room for us to improve on what we’re doing because we’re constantly learning and evolving to improve both, even at such a low level of like, “Okay, how do we lay out the files in our service repository to make the best organization to make sense?” All the way up to, “Okay, how are we going to do tracing? And what kind of information do we need to get from that so that we can find problems when they occur?” We’re always looking to learn what others are doing, and talking to others in this space. No one will ever be a hundred percent right. There’s always room for improvement everywhere.

Corey: This episode is sponsored in part by LaunchDarkly. Take a look at what it takes to get your code into production. I’m going to just guess that it’s awful because it’s always awful. No one loves their deployment process. What if launching new features didn’t require you to do a full-on code and possibly infrastructure deploy? What if you could test on a small subset of users and then roll it back immediately if results aren’t what you expect? LaunchDarkly does exactly this. To learn more, visit launchdarkly.com and tell them Corey sent you, and watch for the wince.

Corey: One thing that you folks have done that I think was really interesting and didn’t get as much play as I think it really deserved, was that, especially in the early days of the pandemic, you wound up seeing that massive increase due to giving out almost a million free three-month subscriptions to Playthrough. Additionally, you also worked closely with LAUSD, the Los Angeles Unified School District, to add Fender Play to their middle school music program’s curriculum to help supplement their remote learning programs. First, was that all in the same timeframe? Or—and, two, what has it been like, I guess, working with a organization that is, I guess, on some level, not particularly cloud-first. I would imagine. When I lived in Los Angeles, I never got the sense that LAUSD was full-on serverless, full on-board with cloud, full on-board with remote learning. And then the pandemic of course exacerbates all of that.

Michael: Yeah, so those were really two different projects. So, that the Playthrough project that started in March, and we started working with Los Angeles Unified School District last year during their summer school program; started out with 1500 students and we put it together very quickly. Essentially, we use the same three-month codes that we used for that Playthrough promotion so that we could set things up very quickly for students and gave out, through our nonprofit arm of Fender, the Fender Play Foundation, gave out 1500 instruments to these students to use during the summer school program. And that program became so successful, we continued on with them in the fall, and now in the current semester, and we will be again this summer. I believe there’s 7000 students in the program now.

And working with their IT team has actually been quite nice. And in dealing with partners, you wouldn’t think much of, “Oh, it’s a school district, what do they have?” But as far as just ease of working with them, we actually hooked into their SAML provider in Cognito so that LAUSD students could authenticate when they come in through the remote learning systems. And they were great to work with and very helpful and cooperative.

Corey: One of the arguments that you’ll see that comes up against serverless, from time to time, is that you are now indelibly linked to your provider, but you can’t take what you’ve built with all of these services and just move it over to Azure or GCP on a moment’s whim. Now, in practice, people who tend to build for that, just build everything on top of EC2 and very little else, and then run it entirely in AWS and never move it to any of those other places. But was there friction with making that, I guess, architectural commitment to a single vendor?

Michael: Oh, you’re bringing up the vendor lock-in Boogeyman.

Corey: Oh, I absolutely am. Most people who bring that—when I bring it up as a straw man so you can attack it, most people who bring up the vendor lock-in Boogeyman, “Oh, you have to go multi-cloud,” are either trying to sell you something that is required if you want to go multi-cloud, or they’re a cloud provider themselves who know that if you go all-in on one provider, it will certainly not be theirs.

Michael: I think if you properly architect your applications with separations of concerns that you could move to, say—okay, say Lambda wasn’t working out for us anymore, and we needed to take our applications and, where, we’re going to put them into a container, but we’re going to stay in AWS. Our applications are set up in such a way that Lambda is basically a deployment pattern. We could easily convert those individual function handlers into route handlers with a minimal effort because the business logic and then the underlying data storage are separated. So, it would be feasible for us if we wanted to, say, move to Azure and use Azure Functions and whatever comparable service they have to DynamoDB. I’m not too familiar with a lot of their offerings.

But that would certainly be possible to do it with, obviously, some effort and really, at the end of the day, the resources you have working on the applications are end up going to costing you much more than any, sort of like, software licensing or specific savings you’re going to get from a cloud vendor, so might as well go ahead and just use those service that they’re providing. So that you can just focus on the business.

Corey: My approach has almost universally been that looking at an awful lot of companies and their AWS bills, it is a challenge to find an environment where the resources in the environment cost more than the people who are operating them. In the context of business, AWS bills seemed giant and enormous, right up until you look at payroll and then it’s, “Oh, okay.” That’s counterintuitive for folks who are learning this, and I fall prey to it myself is, when I’m playing around as a hobbyist trying to build something I value, my time is free because I’m learning as this goes, and then in that context, especially when I was starting out as a student, it was, “Oh, great. So, this winds up costing me $7 a month. Oh, that’s a lot of money. That’s my ramen budget, so I’m instead going to wind up spending eight hours avoiding it charging me anything.” It’s the exact opposite from the direction you want staff that you’re paying to work on these things to go in. How do you approach the idea of increasing the cloud cost if it will save time for your team?

Michael: It’s a balance between, where do we need to build this ourselves? And then not only build it, you have to operate it and maintain it? Or what is the cost of getting this third-party service? And that’s really what it comes down to in all of them. And do we actually want to spend time working on this piece of infrastructure that these other people are specializing in and do so well? I’ve got better things I can have people doing than that.

Corey: Speaking of people, one thing that you talk about, as you self-describe, is that you wind up not writing a whole lot of code anymore, but you’re something of a stickler for observability and enforcing consistency between services, so you’ll periodically do things like submit a PR to tweak a log message to put your mind at ease, was one example that you gave. Given that you’re a director, which is generally manager of managers style approaches, how do you avoid having those PRs come across to your team as either micromanagement or a condemnation of what they’ve built? Because I get it; when I see something that’s easy and small to tweak, I want to go ahead and get it fixed immediately. I don’t want to go back and forth and play those games; I just want it done. But I’m also always weighing that against, I don’t want to have people think that I’m judging them somehow for something I’m very much not.

Michael: That’s a very good point. The larger technical decisions on how things are laid out, I generally just try to—I don’t insert myself into. I let the team go ahead, and make those decisions, and leave that direction, and let them take the charge on that, and I take the approach of looking at it as more of a guiding, and mentoring and teaching to really hone and instill that discipline in really being able to understand what the applications are doing. And as our team is growing, I have less and less time to even do those things, but I can go through the systems and go, “Hey, how come we’re not tracing this call to the reCAPTCHA servers? Let’s add that in there.” And I’ll just at this point now, I mainly just write Jira tickets to have someone else actually do the work.

Corey: The more I do this, the more I realize that as complicated as the technology is, the people are in many ways, far more complicated. And let’s be fair here, non-deterministic things that work super well on one person one month could work entirely differently a following month, or even with the same person, or between teams. It’s a constant balancing act, on some level. And giving people a sense of psychological safety has always been the biggest challenge. The thing that surprised me about management, back when I was running ops teams was the more, I guess, responsibility you accrue as you rise from individual contributor into the management—or ‘rise’ is sort of a wrong term; it’s an orthogonal transition—is that you spend a lot more time on the people problems, and your ability to directly control or affect change diminishes because you have to do everything via influence. You get a lot more responsibility with a lot less direct power [laugh] over the outcome in some ways. Does that align with how you see it, or am I just—do I have very strange approaches on management? Which may be true, and why I got out of it as fast as I could.

Michael: No, that is a good point because you are having to [unintelligible 00:27:05], like, influence, and guide, and more take a higher-level view, as opposed to really getting into the weeds of like, “Okay, what methods are we going to put on this interface? How are we going to, say, architect the internals of an application?” Those are details I just really don’t have time for anymore. But larger things as to making sure that we’re okay, it’s like, “What’s the performance of this?” And, “Overall, is something that can be adapted as the business needs change, and as we change? And as we learn, what can we do to modify it?” And more just things like guiding, and mentoring, and really taking a higher-level view of that.

Corey: I’m going to selfishly ask about something that I struggle with myself. That goes a bit more into the technical area, but you talk about
enforcing consistency across all of your different services. What does that mean? Similar coding style? Similar instrumentation?

Because I look at the things I built and microservices that power my internal nonsense, and each one of those is very different than all the rest. So, whatever your version of consistency is, I know I’m not doing it. But how do you view it?

Michael: So, there’s really two types of consistency. The one I really refer to the most is in observability. So that, if you’ve got a thousand Lambda functions out there, and each one is logging things slightly differently, that’s just a pain to deal with, and realistically, dealing with a thousand unicorns is a real pain. So, through that observability, at least in Lambda, we use an internally developed middleware to make sure that the logging is consistent, and it’s easy enough to use. And then other consistency, like, just within projects of how we lay things out.

That’s something that’s been consistently evolving. What’s the folder structure in how we organize the code? And we’ve kind of been evolving that over the last three years. And within about the last six months, we’ve come up with a really good pattern and a template for the future. And it’s not much different from what we started out with, but it’s a little bit easier, really, to comprehend as a new engineer coming in. It makes more sense.

Corey: I have to ask—and I understand if you don’t want to give a particular endorsement in any direction—but do you go through Serverless Framework, SAM CLI, the CDK, using the console and then lying about it? What is the template that you wind up using for that uniformity? Because even internally, I use three or four of those different things and professional advice: don’t do that.

Michael: Let’s see. So, in our development, QA, production environments, infrastructure is all managed with Terraform. Each engineer has their own personal AWS account so that they can work on things there—

Corey: Oh, that makes billing granularity super easy.

Michael: Oh, yes. You can tell who’s got EC2 instances running up for too long. But for the most part, we’ll use Serverless Framework in that regard to say—for the engineer can just deploy into your local environment. Although we are working on ways to reuse the Terraform infrastructure and deploy that. But we have our own build and deployment pipeline that we built using CircleCI, and all of our Lambda
functions are in Go.

And so having to compile, say, 20 binaries in a service, that gets kind of slow, one of our DevOps engineers actually came up with a way to use Lambda to build the Lambdas, so that we can build them all in a distributed parallel fashion during the build process.

Corey: One thing that I do love about the whole serverless approach—and it is a neat part about Lambda—is no two people ever seem to do it quite the same way. You can tie things together in so many different and exciting ways, and it’s fun. It’s almost like a modern version of playing with Lego. And I know that if Jeff Barr is listening, he just perked up at that. But I love the concept that you can take so many different ways to achieve similar outcomes. And it almost gives a bigger sense of creativity in how you approach problems. Has that been your experience?

Michael: Oh, definitely. It’s not only the creativity; it’s also the flexibility in how you solve it, and the ability to adapt and evolve as services evolve, or change, or there’s new ones are added. And to the point of using AWS, kind of, saying, “Oh, using a Lambda function to do this.” Like, using Lambda functions for customizing behavior of Cognito with the Cognito triggers, is to me, I think, a perfect way to customize the service to do exactly what you need to do.

Corey: I want to thank you so much for taking the time to speak with me today. It’s always appreciated. If people want to hear more about what you have to say and how you view these things or even, possibly, decide to work with you, okay can they find you?

Michael: I’m somewhat active on LinkedIn. LinkedIn is the best place to find me. Please go ahead and connect to me; tell me you heard me on the podcast here.

And yes, we are hiring. We have, all within our technical organization, from client, to web, and mobile engineers, data engineers, DevOps, API, we’re always hiring and if we don’t have something right now that fits your experience, let me know that you’re interested and I’ll put you on the list so that when we do have an opening, we’ll reach out right away.

Corey: And we will, of course, include links to that in the [show notes 00:32:20]. Thank you so much for being so generous with your time. I appreciate it.

Michael: Thanks for having me on, Corey. It was nice talking to you.

Corey: Michael Garski, Director of Platform Engineering at Fender Musical Instruments. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with a comment telling me that I’m almost certainly doing that chord incorrectly.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

Links:

  • CTO Whitepaper: Reinventing Enterprise Networks for the Cloud Era
    www.alkira.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part byLaunchDarkly. Take a look at what it takes to get your code into production. I’m going to just guess that it’s awful because it’s always awful. No one loves their deployment process. What if launching new features didn’t require you to do a full-on code and possibly infrastructure deploy? What if you could test on a small subset of users and then roll it back immediately if results aren’t what you expect? LaunchDarkly does exactly this. To learn more, visitlaunchdarkly.com and tell them Corey sent you, and watch for the wince.

Corey: If your familiar with Cloud Custodian, you’ll love Stacklet. Which is made by the same people who made Cloud Custodian, but put something useful on top of it so you don’t have to be a need to be a YAML expert to work with it. They’re hosting a webinar called “Governance as Code: The Guardrails for Cloud at Scale” because its a new paradigm that enables organizations to use code to manage and automate various aspects of governance. If you’re interested in exploring this you should absolutely make it a point to sign up, because they’re going to have people who know what they’re talking about—just kidding they’re going to have me talking about this. Its doing to be on Thursday, July 22nd at 1pm Eastern. To sign up visit snark.cloud/stackletwebinar and I’ll talk to you on Thursday, July 22nd.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. On this promoted episode, we’re returning to something I did a while back on the AWS Morning Brief. I took us on a twelve-week exploration of networking in the cloud and how that wound up impacting how companies do business. Today, my guest is cloud networking evangelist Rasam Tooloee, and he works over at a company called Alkira. Rasam, thanks for joining me.

Rasam: Thank you for having me, Corey. Pleasure.

Corey: So, let’s start with the obvious. What is a cloud networking evangelist? I’ve heard of people evangelizing all kinds of things, some of which make more sense than others, but this is the first evangelism title that actually made me sit up and say, “Ooh, this is relevant to my interests.”

Rasam: That’s funny. A cloud networking evangelist, to me—you know, what I consider my charter is really helping our customers and prospective customers understand that networking the way they’re used to doing it historically, if you look at legacy networking and the way that networking has evolved, specifically in the cloud, is just not sufficiently agile; it’s not sufficiently natively enterprise-grade from a visibility and control and compliance perspective. And that there’s just a better way of doing networking for this cloud era that we’re in. And I find enterprises that I talk to every day struggle with the complexity associated with how to get their properties into the cloud. There are many of them have become natively multi-cloud just by default, you know, some through business imperatives and priorities, some through acquisitions, but ultimately in the context of networking, that has led to a whole lot of complexities that they grapple with, and the evangelist in me is looking to help them find the better options that are out there for them.

Corey: In the earlier days, before I got into Cloud, I was deep into configuration management. And oh, we manage all of these systems via configuration drift detection, and every time they run, they remediate the drift, and it’s great. Cool, so how do we wind up managing the networking equipment? Well, there’s this thing called RANCID. It’s made out of some horrifying Perl and if you turn on ‘strict,’ the whole
thing breaks.

And it was this awful sort of dark ages technology to approaching networking. It felt like the DevOps movement towards agility really didn’t come to networking in any meaningful sense for a while after that. Is that accurate? Or was I just hanging out in the wrong shops?

Rasam: No, that’s absolutely accurate. For sure.

Corey: Your career has been fascinating. You went from Cisco, where you presumably worked on networking because that’s kind of the thing they’re known for, then Salesforce, which is sort of definitionally SaaS, as says on the tin. You went to cloud with Microsoft for a while, and now you’re at Alkira, where you’re sort of in the perfect center of all three of those things. Tell me a little about how you got to where you are?

Rasam: Yeah. Well, I have my roots in networking. I worked for Cisco—a great company—for a long time, and really got the opportunity to tackle networking for many different facets, both core networking as well as some of the advanced technologies that Cisco forayed into, and absolutely loved the ride and learned so much. And then there came a time where SaaS was clearly the next big wave. And being at Cisco, I watched that wave grow from afar, and at a certain point in my career, I decided to take the leap and go to the SaaS leader at the time, Salesforce, and really learn that business from the ground up, both in terms of the underlying constructs of how SaaS business model works, but also the core business that Salesforce has in CRM.

And then from there, I transitioned to Microsoft because SaaS was the tip of the spear that launched the cloud revolution, but then there’s a lot more to cloud than just SaaS. And going to Microsoft really helped me to understand that business from the ground up. I’ve always had this inclination of really being intrigued and curious about what’s next and trying to take that next leap before I have the opportunity to really enjoy the uptrend of the ride. I’m looking for the next emerging significant innovation, but what’s interesting is, come full circle, I’m back in networking. Which kind of begs the question of, like, what happened?

And to me, what happened was, in Alkira, I find the coming together of everything that is fascinating to me about cloud and SaaS with where I really earned my chops in technology, which is networking, and really solving this problem for this era, this moment in time in a truly innovative way. So, I feel like it’s actually—it may look like full circle, but it’s actually a continuous trend of seeking out the next innovative thing.

Corey: Back when I wound up getting into networking, the reason I did it was because first, it was the 2008 financial crisis and no one was hiring; we just had a salary freeze and I was demoralized at my job. But I realized that as a systems administrator, I was always sort of hand waving over the networking pieces. And all right, let’s figure out how this whole thing works. So, I got my CCNA, my Cisco Certified Network Administrator, cert, back in the days when that didn’t have a whole bunch of different derivative adjectives after, telling you exactly what kind. And what happened next was that, okay, now I understand it a lot better.

Every time I find myself basically scratching my head, trying to figure out exactly what the deal is with something that I’m working on technically, and I don’t understand what’s there, dig deeper into that and you’ll often discover that it makes everything else make a bit more sense. And then came cloud. And now we have cloud networking, and anyone who tells you they understand how cloud networking works is generally lying to you. It feels like it is complexity stacked on top of complexity, and these days, it more or less distills down to you fire up your cloud provider of choice, you click a few buttons in the console, and really hope you did it right. I’m guessing that you have not automated the clicking of the buttons in the console, so how do you folks approach it? What have you done that’s different?

Rasam: So, to build on your point about the complexity of cloud networking, there’re a number of reasons why it is so cumbersome and so complex for enterprises to tackle the challenge of cloud networking. One is, it tends to be rather rudimentary in nature, and there’s a lot of manual effort involved, there’s hop-by-hop configuration, you have to do unnatural things to solve for some basic challenges. For example, we often find enterprises or service chaining firewalls so they can have symmetric traffic routing. And they will do things like have a separate path for ingress and egress traffic; the segmentation is extremely hard. So, one part of it is—probably my personal opinion; I don’t think cloud was built with the idea of let’s solve for the networking from the ground up because that’s so important to how people are going to have to manage their compute and storage.

It was all about compute and storage, and networking was kind of an afterthought. And that shows. That it shows in the way that you actually have to configure your cloud network. Now, the other thing that really complicates it is the fundamentals of networking don’t change, but the way in which the fundamentals are applied and the vernacular that’s used to describe it and the UI that’s used to control it varies from cloud to cloud. So, just because you learn Azure doesn’t mean you know AWS, and doesn’t mean you know GCP.

So, if you are becoming multi-cloud, now you got to go learn all this stuff separately, and then actually as an amplification of the challenge that is associated with the hop-by-hop configuration, when you bring up a region, for example, in cloud provider A, that doesn’t mean that you all of a sudden did most of the heavy lifting required for region B, you got to go do the same thing in region B.

Corey: Oh, I had a client that had a very deeply skilled networking team that spent months without much success trying to get Terraform to set up IPsec between GCP and AWS at one point, and ultimately they gave up in disgust. My argument about, “Oh, we’re going to go multi-cloud to avoid locking.” “Well sorry, you already have lock-in, both in terms of what your staff’s up to speed on, but also the identity model, the security model, and critically, the networking approach.”

Rasam: Yeah, absolutely. And back to your original question of how we do it differently. So, what we have done is really looked at the problem differently through a new way of thinking. Again, this goes back to my prior point about the network isn’t sufficiently agile, and the reason it’s not agile is for all the reasons that I explained. And when our founders who come from decades of experience in networking looked at this problem and they looked at the native value proposition of cloud—which in our mind is agility—is the fact that cloud is a competitive imperative these days, that’s where innovation is happening, and we see enterprises every day increase their investment in cloud because if they don’t, they’re going to get left behind; they’re going to get left behind because the next digital disrupter or their competitor is going to do, in cloud, the things that their customers expect, and the things that are required to truly compete in today’s marketplace.

So, because it’s a competitive imperative, and at the heart of it is agility, our founders really contrasted what networking was like in the cloud and in the cloud era, which is really largely fragmented, and silos, highly complex, slow to deploy to your point, often CapEx heavy because you’re making substantial investments in things like colocation and dedicated bandwidth. There are a lot of delays and limitations like I talked about in terms of the various constructs between the cloud providers.

They contrasted that with DevOps and said, “Okay, look. DevOps is all about automation. It’s about rapid iteration. It’s about abstracting the underlying complexity on an elastic platform that scales with you. You can actually go into cloud with minimum upfront investments, test and iterate, and then scale as you need to with velocity and agility. That’s what cloud is about. That’s the way DevOps has adapted to the constructs of cloud. But network isn’t the case, so how do we rethink networking from the ground up, so that it is more in line with the business imperative of why businesses go in the cloud in the first place?”

And to do that, what they really did was design a unified fabric that’s a multi-cloud unified fabric that delivers a full stack of networking services that meet the vast majority of the use cases that an enterprise would have from a networking perspective, and does so in a way that’s natively multi-cloud, and does so in a way that natively addresses some of the complexities with things like security, compliance, visibility, control, et cetera.

Corey: And I’ve been very vocal about opposing multi-cloud as a best practice, and people sometimes are surprised to discover that as soon as I find a customer who’s doing multi-cloud, I dive right into discussions about that, and, “We thought you were going to yell at us.” Look, do I think it’s a best practice in the general sense? No, but you have specific constraints, and you have an environment that is how it is, and sitting here saying, “Oh, you should have made a different series of decisions six years ago,” it turns out is not the most compelling story. And there are always specifics that override general guidance. So, whether I like multi-cloud or not as a guidance perspective, I don’t think that I can intelligently deny the reality that it very much exists in an awful lot of places.

And sitting here just trying to be a purist by going through one cloud, whatever it happens to be, and nothing else doesn’t really solve any pain that customers have. Hybrid is and will be a big story for a long time. In my more cynical moments, I tend to view hybrid as, “Well, we tried to do an all-in cloud migration and got stuck halfway through because it turns out, it’s hard to move some things, so we gave up and called it hybrid and now we’re calling it good.” That might be overly cynical, but it takes time to move these things. It takes time to wind up wrapping around a bunch of different environments.

So, if you have something that makes it a lot, I guess, more straightforward to rationalize about and around the network layer, that really feels like it’s a great equalizer because that is one of the most differentiated aspects of all the different clouds.

Rasam: Yeah, absolutely. I mean, the proof is in the pudding, right? So, we find the challenge of getting to cloud, getting cloud networking enterprise-ready from a security, governance, compliance perspective, high availability perspective, disaster recovery perspective, to be a monumental challenge. And for an enterprise, it could be an effort of months, or years, sometimes, for a single cloud, much less a multi-cloud. And just because you did it with Cloud A doesn’t make Cloud B all that much easier.

And I agree with you; I think multi-cloud isn’t necessarily an easy and desirable place to find yourself, but that’s besides the point because enterprises are finding themselves there for a myriad of reasons. It could be business imperatives, partnerships, acquisitions, it just happens. And when it happens, you need the best possible strategies and tools to deal with that. And for us the proof is in the pudding because we’ve had customers be able to contract the amount of time that it would have taken them to get from Cloud A to Cloud B from months and months to a matter of weeks. We can provision something that would take multiple weeks of change control and manual effort, and do it in a matter of hours.

So, I don’t want to overstate how much the technology simplifies things, but the technology does literally simplify things that much. There’s still business process involved, there’s still change control involved, there’s still the human element of making sure that the change is well orchestrated, but the actual process of getting your cloud networking and multi-cloud networking up and running is simplified in a way that I think you have to see to believe and, you know, the proof is in the pudding, and when we have a chance to actually demonstrate that to our prospective customers, it truly is game-changing.

Corey: It’s clear that you’ve built something that works. You have a laundry list of customers on your website who are referenced customers, and these are logos and names people recognize. It’s not, “Oh, wow. That sounds like you made half of those up, and weren’t three of those the big evil corporation in some movie somewhere?” No, these are real companies solving real problems.

And digging a bit into what you’ve built before you came on the show, it is clear that you folks offer a TCO story that lowers the total cost of ownership, but lies, damn lies, and TCO analyses tend to be the three forms of lies people tell. I’m much more interested in the story of how you accelerate time-to-market because speaking as someone who focuses on AWS bills and cost reduction, it always takes a backseat to accelerating features being released. So, there’s a capability story that goes along with this, which it sounds like they’re very much is. That’s the real win; the fact that it saves money is almost icing on the cake.

Rasam: Yeah, absolutely. You’re right, the Holy Grail is time-to-market, which really, time-to-market is very much for me, synonymous with this idea of agility and the ability to pivot, and to get to the next iterative desired outcome for your organization, whatever that may be, quickly. That’s consistent with this idea of velocity, and iterative testing, and scale that the cloud provides. For example, recently, I’ve been working with one of our prospective customers who’s, really, underlying challenge is, “Look, I’ve already built this really robust infrastructure from a cloud networking perspective. It is really colo-centric; that’s my model for my cloud interconnects, but I am now in a global expansion phase. I need to go to all these new geographies, and if I were to do what I just did to build out my cloud networking footprint, I’m looking at a substantial CapEx investment and a substantial amount of time and runway to get that operational, and I just don’t have the CapEx or the time for that.” So—

Corey: What, they can’t just copy and paste the config from one to the other again and again and again in the true StackOverflow tradition?

Rasam: Or get the circuits dropped in the colo, or get all that hardware delivered, and deal with all the complexities of international customs control, et cetera, et cetera. So, what we bring to them as a value proposition is the fact that our points of presence are virtual; they’re software-defined constructs that run atop the hyperscale cloud provider. We can spin them up anywhere in the world where the hyperscale cloud provider has a footprint, and we are in many regions across the globe. And if we’re not in one, we can get one up and running in a matter of days. And most of that time is actually just spent testing it to make sure that is operationally viable; the actual provisioning and turn up of it is very, very quick.

So, the ability for us to be a virtual PoP for this particular customer and give them the ability to quickly expand into brand new geos in a way that also concurrently, natively streamlines and simplifies the complexities of cloud networking that we’ve already covered is extremely attractive to them. And from time-to-service perspective, it’s taking their ability to deliver the needed services in the cloud to their business users from something that would have taken months and months to something that can be up and running in a matter of weeks.

Corey: If your mean time to WTF for a security alert is more than a minute, it's time to look at Lacework. Lacework will help you get your security act together for everything from compliance service configurations to container app relationships, all without the need for PhDs in AWS to write the rules. If you're building a secure business on AWS with compliance requirements, you don't really have time to choose between antivirus or firewall companies to help you secure your stack. That's why Lacework is built from the ground up for the Cloud: low effort, high visibility and detection. To learn more, visit lacework.com.

Corey: Can you give me an example of a customer pain point that you’ve resolved? Because, again, you have customers willing to say nice things about you, but one of the challenges I’ve often found with a lot of the, shall we say larger, more enterprise-y offerings is, “Well, what did you actually do for the customer?” And the answer requires two hours and at least 40 PowerPoint slides and at the end, you say you get it just to get the person to stop talking. What is the value, the better outcome that you’ve delivered for a customer?

Rasam: Yeah, sure. So, our customer, Koch Industries—and they’re a public reference for us; you can check out their story more in-depth on our website—but they were your traditional enterprise, originally designed for cloud using a hub-and-spoke architecture, which consisted of using the data centers as the focal point for data center interconnect, cloud interconnect, high-speed bandwidth, private Lan, et cetera, that comprise their overall architecture. And over time, they simplified—and I use the word simplified loosely here—but they simplified with a more cloud-native, cloud-transit type of architecture, where they leveraged more of the default capabilities and networking services on cloud, which helped considerably. There was a ramp involved in learning the native-cloud constructs and associated networking and security aspects of that, but over time, they did simplify. They were able to condense their overall provisioning time of a cloud interconnect from what they originally shared with us was eighteen months down to about six, and consolidated across about a dozen transit hubs from a cloud networking perspective.

But then, as we discussed previously in the podcast, when they took a step back and looked at it, what they still saw was an enormous level of complexity in networking, an enormous level of complexity in operations, and they still were seeking a better way, a way that was operationally viable in the long run with a lower total cost of ownership, and the ability to really consume networking services in a way that moved at the speed of business in a way that was more in line with the way that we’re using cloud computing storage, and in line with the speed and agility with which their business wanted to move. And that’s where Alkira came into the picture, and there was a real alignment of vision between how they saw their networking strategy moving forward and how Alkira delivered services. Long story short, they are now able to take their planning process down from six months to a matter of weeks, and the actual provisioning process of cloud networking to a matter of hours, sometimes less. And that has brought immense value to their business and to their IT organization, again, in terms of agility, in terms of total cost of ownership, in terms of visibility and control, in terms of governance. And another added benefit was historically they were single cloud, AWS, but in the process of their journey, with Alkira, the need came up to go into Azure for some Azure native services in a scenario where the data still resided inside of AWS, and that request historically would have been months and months of due diligence to get the environment up and running, and in their case, they were able to do that all within a day because they were already leveraging the Alkira multi-cloud platform.

So, a tremendous amount of value for them across a myriad of fronts that, again, have been pivotal to their long-term strategy and how they address cloud networking moving forward.

Corey: If we go back to the early days of cloud, we started off with some of the advanced stuff like, you know, virtual machines—some places called them instances—and there was a lot of competitive variation between them. “Well, these instances cost a fifth of what this other cloud providers do.” “Yes, but that other cloud provider [unintelligible 00:19:47] don’t fall over every 20 minutes and have persistent disk.” In the fullness of time, everything’s sort of commoditized to the point where now, in many cases, if you’re just running a bunch of virtual machines on cloud providers, it’s largely a matter of price. The same story has happened in many respects with object store. Do you think that the network will eventually wind up commoditizing as well, or do you think that there’s still going to be significant variances as the rest of the cloud world grows up on top of that bedrock foundation?

Rasam: I think that’s a brilliant question, and I think the answer is yet undetermined. I don’t think it’s clear. I think there are a lot of different approaches to trying to solve for the challenge of networking in a cloud, first world, right? And most of the solutions on the market address some subset of the problem: some gets you to the edge of cloud, some really reside on the edge and try to interconnect you to the various clouds that you want to be in, some are meant to help you orchestrate your cloud footprint once you’re in the cloud. The underlying challenge remains that, at its core, cloud networking itself remains extremely complex and extremely siloed.

If you zoom out and look at your traditional enterprise architecture, it’s a bunch of siloed solutions that have been stitched together to meet the end-to-end workflow. Well, cloud is kind of a microcosm of that. The same thing happens in cloud is, you have a lot of manual intervention of stitching together the various pieces to meet the end-to-end workflow. None of the existing approaches on the market are really operationalizing cloud from a networking perspective the way DevOps and containerization has done with compute and storage and made it really a seamless part of an end-to-end infrastructure as code strategy. So, I think everyone is really trying to tackle that problem in a way that hopefully, the end state will be one that is aligned with the underlying value proposition of what cloud brings to an enterprise.

But how that is going to end up looking and whether or not it ends up being a singular sort of end-to-end infrastructure as code strategy that ties the pieces together elegantly, or ends up being all these various piece-parts that are solving a best of breed problem but still need to get stitched together, I think remains to be seen.

Corey: One of the things that I think networking has had in common or is at least spiritually aligned with the world of security is that when it isn’t working, “Well, we’re going to go ahead and make things broader and broader and broader, and we’re going to go ahead and grant everything access to everything, and once we get it working, then we’re going to go back and dial that back down because we want to be secure.” Yeah, no one ever remembers to go back and dial things back down. Once it’s working, we’re on to the next ticket, in many cases. So, the complexity doesn’t just act as a drag on feature velocity; it also acts as significant security risk in many environments. How do you folks tackle that, or think about that? Or is that one of those, “Oh, that’s the best kind of problem: someone else’s.”

Rasam: I think at the root of that problem is the visibility and control problem because it’s easy to do something, to turn some knobs to get something up and running and then forget about it. And if you don’t ever go and touch that part of your network again, then you can easily end up in a situation like the one that you described. And that’s why we really think of the idea of solving for this problem as needing a new paradigm and a new way of thinking, which is a unified fabric, end-to-end, in a multi-cloud world, with a full stack of network services that addresses the vast majority of the use cases that an enterprise would have. So, we’re literally giving you a single user interface for full visibility and control end-to-end for all of your networking use cases, be they on-prem, for your remote users, for your branches, or any of the clouds that you might be in.

Corey: When you find that you’re talking to your prospective customers that, in the fullness of time, become actual customers, and they wind up going from, “Okay, this might work,” to, “This is awesome,” what do you find that they’re, first, the most surprised about during the adoption? And secondly, what do you think their biggest misunderstanding along the way was?

Rasam: You know, the way that you leverage the Alkira Network Cloud—which is what we call it. We call it Alkira Network Cloud because it is in fact a network cloud that delivers all your full stack of network services in a cloud model. But the way you leverage the Alkira Network Cloud is you go through a multi-step, really simple workflow. So, we have this concept of cloud exchange points, which you can think of as virtual PoPs, and they reside all over the world. So, the first thing you do is you pick your virtual PoP or PoPs—you can have one or multiple of them, as many as you need—and the next thing you do is you attach your sites to this fabric.

And there are multiple ways you can do that. You can do that through high-speed dedicated connectivity like AWS Direct Connect, you can do it by extending your SD-WAN fabric into the Alkira fabric, you can do it through IPsec connections, you could do it through remote access for your users. But that’s the first step. And then the next step is to attach your cloud VPCs or VNets. So, you go through a process of providing your credentials for your cloud properties, and you attach the cloud properties to the Alkira fabric, and in the middle, there’s the step of defining your segments.

So, you define logically what your segments will be, and then you assign your sites, or your users, or your cloud properties to that segment. And literally, I mean, that’s five steps, and at the end of those five steps, you just established end-to-end multi-cloud connectivity from your sites, and branches, and data centers, and users to your cloud properties end-to-end with full visibility and control. And usually, that process can take 30 minutes, if you have all of your credentials and the necessary data lined up for what you’re connecting and the sequence that you want to go through, and at the end of that half-hour, people that are new to the platform will stop and say, “That’s it? We’re done? It can’t be that easy.” And in fact, it was that easy. And that’s really the big aha moment for a lot of our enterprise customers that see the platform for the first time of, like, “Wait. This is way, way different than anything I’ve seen before.”

Corey: Your website has a 30-minute challenge for configuring a network, and I haven’t run myself through it yet with a stopwatch, but the fact that you can even make that claim means that there’s something radically different because frankly, it takes that long to find that the networking section of the console in many of the cloud providers. Something you just said was—talking about your enterprise clients; do you find that you’re generally working in the enterprise space, or do you tend to have offerings that make sense at the SMB scale? In other words, when is it time to start talking to you folks? Invariably, “After someone probably should have,” seems to be a common refrain, but at what scale does Akira begin to make sense?

Rasam: Yeah. I think I use the term ‘enterprise’ sort of, more generically than your large enterprise.

Corey: Oh, to me, a big company is anything with more than 200 people, so I’m the wrong person to ask on that score. But yeah.

Rasam: Yeah, and I would say I agree with you, and that’s kind of the definition of when I say enterprise for me. Because networking is a horizontal problem. Every company needs networking and no matter what the size of your organization, if you’re going into cloud, you’re going to have to deal with the challenges of cloud and operationalizing the challenges of cloud. Now, the larger you are and the more clouds you’re in, the greater the complexity that you have to deal with and the greater the operationalization of that complexity. So, we deal with large enterprises that are deep into their cloud journey and find themselves back-ended into complexity and looking to simplify.

And we also have enterprises that are born in—I’m sorry. When I say enterprise, I’m talking about customers that are born in the cloud, startups that are really looking for a simplified and operationally aligned networking solution with the way that they’re intending to leverage cloud. So really, if you’re getting into cloud, and you’re getting into cloud networking, and you have a cloud-first strategy, regardless of the size of your organization, the chances are pretty good that Alkira is going to be a good fit for you.

Corey: Thank you so much for taking the time to speak with me today. If people want to learn more about what you’re up to, how you view these things or basically take it for a spin themselves where can they find you?

Rasam: On alkira.com. So, www dot alkira—A-L-K-I-R-A dot com, and take a look at our resources page. It’s packed with great content. And like I said earlier, you really have to see this to believe it, so we’re happy to show you; request a demo and we’ll get online for you and take you through the journey.

Corey: Excellent. Well, thank you so much for taking the time to speak with me. I really do appreciate your being so generous with your time.

Rasam: Thank you, Corey. I really appreciate it.

Corey: Rasam Tooloee, cloud networking evangelist. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you hated this podcast, please leave a five-star review on your podcast platform of choice along with a comment containing the proper Terraform configuration to get IPsec working between two different clouds.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Jason

Jason Yee is Director of Advocacy at Gremlin where he helps companies build more resilient systems by learning from how they fail. He also leads the internal Chaos Engineering practices to make Gremlin more reliable. Previously, he worked at Datadog, O’Reilly Media, and MongoDB. His pandemic-coping activities include drinking whiskey, cooking everything in a waffle iron, and making craft chocolate.

Links:

  • Break Things On Purpose podcast: https://www.gremlin.com/podcast/
  • Twitter: https://twitter.com/gitbisect

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the Enterprise (not the starship). On-prem security doesn’t translate well to cloud or multi-cloud environments, and that’s not even counting IoT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IoT devices, detects these threats up to 35 percent faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at extrahop.com/trial.

Corey: This episode is sponsored in part by LaunchDarkly. Take a look at what it takes to get your code into production. I’m going to just guess that it’s awful because it’s always awful. No one loves their deployment process. What if launching new features didn’t require you to do a full-on code and possibly infrastructure deploy? What if you could test on a small subset of users and then roll it back immediately if results aren’t what you expect? LaunchDarkly does exactly this. To learn more, visit launchdarkly.com and tell them Corey sent you, and watch for the wince.

Corey: Jason, thanks for joining me.

Jason: Thanks for having me, Corey.

Corey: So, you’re one of those people that we’ve always passed at conferences and other events, sort of like ships in the night. We hang out in group settings, but strangely, for whatever reason, despite traveling in the same circles for years now, we’ve never really sat down had an in-depth conversation with each other to the point where I feel like both of us are sort of wondering on some level, “Does he just not like me?” It’s been one of those items for me of, I want to catch up with Jason at some point and learn what makes him tick. And then pandemic happened. Well, no more. Thank you for talking to me.

Jason: Yeah. And again, thanks for having me. I’ve always felt the same way. We’re always at these speaker dinners, or just hanging out with friends, and for some reason, I’m, like, at one end of the table, and you’re at the other. And we’ve just never had this opportunity.

Corey: Exactly. Because you actually do a lot of good in the community, and I’m usually at the kids table. Which is, frankly, what happens, and honestly, it’s the right call. But you and I, I guess, are aligned in a few weird and interesting ways. And—well, let’s talk about what you do. You’re the Director of Advocacy at Gremlin. What is Gremlin, first off, and then what is a Director of Advocacy really do?

Jason: So, Gremlin is a chaos engineering platform, or a reliability platform as we’re trying to sell it now. Because we started out doing chaos engineering, so some of the folks that were doing chaos engineering back at Netflix and back at Amazon, decided, most people aren’t Netflix, most people aren’t Amazon; let’s build something that everybody can use. So, Kolton and Forni, our founders, got together, they started this up. And the idea is really, how can we help people make things more reliable? And obviously, chaos engineering is one of those ways, so that’s what they started off with.

And we’ve got a platform that really just makes that easy and safe to do. So, the second question about what is Director of Advocacy? I know you like to make fun of AWS naming, and I feel like it is sort of a weird, nonsense name because it doesn’t actually explain anything. But essentially, it’s developer relations. So, I have the task of talking to all sorts of folks who aren’t customers—really, just anybody in tech—about chaos engineering and why they should be doing it, and how to make applications and systems more reliable.

And then, aside from that, I also get to interact with our customers and help them out. So, I’m a combination of customer success or success engineer slash support slash the advocate side is advocating for their needs within the organization. So, when they make a product request, I pass that on, see what we can do about that. So, it’s sort of a mishmash of all these different roles.

Corey: I want to draw a bit of a parallel that DevRel slash advocacy slash evangelism universe to the sysadmin world where then we started calling ourselves DevOps and that led to an enormous schism around is DevOps a job title or not? “No, but it pays a lot better, so yes.” Then SRE. “Well, you’re not real, SRE,” and the rest. It comes down to quibbling over definition of terms instead of, you know, doing work. And I feel like, on some level, the whole DevRel space has, in some respects, gotten twisted around something that resembles the same axle. Is that unfair?

Jason: No, that’s absolutely correct. There is that question of what is DevRel? How do you define it? And part of that is how do I justify my job? And on top of that, how did—at least pre-pandemic, how do I justify the company spending tens of thousands, if not hundreds of thousands of dollars, not only for my salary but to fly me around the world to get on stage and say things.

Corey: Right. And it looks from a distance, an awful lot like, okay, you cost as much as an engineer, you don’t write any code to make what we do any better. Your expense budget is about the same as your salary in some cases, and then you travel far away to what looks like a giant party to hang out with your friends. And you get on stage and say, “I work at company X. Thanks. They’re great. Now, for the next 45 minutes, let’s talk about the right standing desk for you.” And it becomes a very difficult sell internally. And for a group that prides itself on advocating for its company. They don’t often seem to do as good of a job advocating for themselves, internally.

Jason: Absolutely. There’s always the discussion of KPIs. How do we measure the impact of what developer evangelism, DevRel does? And it’s a hard thing, partly because every company is a little bit different. Because nobody’s really defined this, DevRel often is very fluid and just fills in the cracks of whatever a company needs.

So, for some companies that might be doing support, right? I’ve heard people being called DevRel, and they literally are just on forums all day answering questions, or writing documentation, or speaking. So, it’s really just this nebulous thing of whatever a company needs.

Corey: It becomes almost this weird expression, in some respects, of marketing. Of course, a lot of DevRel folks will scramble at the objection, “Oh, we are not in marketing.” And that’s always said with a very sneering tone towards marketing because those people are terrible. I argue that marketing is, A) wildly misunderstood, B) incredibly valuable, and C) where DevRel in many respects finds its spiritual home because it’s very hard to tie your marketing budget as a company to definable results and do attribution effectively, but there’s clear value to the company in things that can’t necessarily be measured, or at least not without a heck of a lot of work. That is the piece, in many respects, the DevRel is missing. But the first thing that they want to make clear is that we don’t work for marketing. It’s a very weird feeling.

Jason: It’s very weird because as I explain that DevRel often is filling in the cracks and is very fluid, that’s because my personal perspective of DevRel is inclusive. I try to get involved in as many teams as I can, so I’m constantly working with engineering, and with marketing, and with customer success, and really everybody. And then on the flip side, you have people that define it by what it’s not. I’m not marketing, I’m not this. And you end up cutting yourself off.

Corey: And neither are you an accountant, but I didn’t ask if you were, so yeah.

Jason: But at the same time, you’re not an accountant, but you should have some sort of notion of what the finances of the company are because that gives you some sort of indication on whether you’re going to get laid off, for one, but also just for the success of the company. And I think maybe it’s just the engineering mindset that I’ve had from being an engineer of you take everything that and you try to learn everything that you can and put it together. And so, for me, that comes from having experience working in marketing, having experience working in engineering; how can I put these things that I know together to solve a problem? So, rather than saying, “I’m not marketing,” I’m going to ignore that because as you mentioned, marketing’s super valuable, especially the way that they’ve done data-driven marketing now. It used to be like madmen days, you’d throw up a billboard, and who knows if it works, but you paid a bunch of money for it. And now they’re so data-driven, and everything’s tracked. And, yeah, you may not be able to directly connect a few things, but you get a much better sense of where your value is, and where your time should be spent.

Corey: Absolutely. And you can get—I don’t know—the 80% of the way there, and then the last 20% will drive you mad, so at some point, you just shrug, give up, and that’s okay. Similar in many respects to an AWS bill. It just becomes such a weird process to explore. And from a certain lens, when you have those cross-cutting functional types who are doing DevRel, they start to sound almost enthusiastic amateurs in the various disciplines that they bring together.

“Yes, I’m an engineer, but not as deep on the engineering side, as some of my colleagues who do engineering 40 hours a week and then some.” “Oh, we're part of product.” But strangely, to work in product you usually have significant experience and training in how to conduct user experience studies and user interviews, whereas an awful lot of the DevRel input back to product is ‘word on the street style’ stuff.

Jason: Yeah. And both are extremely valuable. It’s obviously very valuable to have that process of doing user studies and actually getting that hard data, but as we all know, that word on the street and what’s the general vibe of folks at a conference or folks at a meetup really informs things that usually doesn’t get asked in those formal user studies.

Corey: Completely. And telling stories from my own world, back when I was, you know, having a real job and able to be fired by a whole bunch of different people—and was—there was the constant justification story of why should you go to that conference and speak? Why would we spend that money? Why shouldn’t it just be a personal thing that you take vacation for? Now that I own the company, it’s a different story because I know that when I go out and participate in the community, good things happen, but I don’t have the need anymore to justify it, other than to myself and possibly to my business partner.

There are very real stories that I’ve looked at here where I go to a conference, I start talking to someone, we keep in touch, they wind up changing companies, we continue to talk, suddenly, they have an AWS bill problem, and now they become a customer. Yeah, it turns out that’s super hard to predict when you’re looking at flight prices to go to that conference in the first place. And there are many other conferences that nothing came out of it, I think, but you never really know.

Jason: Yeah. One of the nice things about my job and one of the reasons that I joined Gremlin was the idea that chaos engineering is still pretty new. And so in my past experience with DevRel, it very much was your exact experience; how has what you said on stage or the introduction of our brand to an audience made an impact? And since chaos engineering has been so new, I’ve gotten to take a little bit of a step back from that. Obviously, I want people to get Gremlin or to try Gremlin, but even if folks just try chaos engineering and have a better understanding of it, that’s a big goal of my job. That means that I win if you try chaos engineering, even if that’s with an open-source tool. So, that’s one of the reasons that I’m super happy about where I’m at right now in terms of DevRel is, I get to be DevRel for an entire practice, rather than just a company.

Corey: And, on some level, you get to define what success and failure looks like among your team. But turn it around for a second; how do you wind up articulating the value and story of what you do to the larger business? Because I’ve seen the approach if you can’t measure DevRel that way—regardless of what that way is—and it’s always this, don’t ask us for metrics. Don’t ask us to really, functionally, be accountable for much. And from a business strategic point of view, where you’re not deeply involved with aspects of what that leads to, “Okay, so it rounds to zero, and wow, I’m spending an awful lot of money on something that doesn’t really add any value. I could spend that money on things that do instead.” And then you see a bunch of negative things happen. Like, as soon as there’s a layoff or a downturn, that entire group winds up getting decimated in some cases, even when, in reality, that’s the thing that should be invested in the most.

Jason: Absolutely, yeah. One of the things that I’ve always loved is people talk about metrics. And yes, we definitely get that from the marketing side. And so I do have metrics on things like how many workshops we run. And those people are obviously, we capture those leads, they go through the marketing funnel, et cetera, et cetera.

But then there’s the idea of how many engineers out there have those same metrics? We always complain about you shouldn’t count the number of lines of code because that’s stupid. You shouldn’t count all these other things. But generally, most engineering teams are working off of quarterly OKRs or some sort of time period, what those goals are and the product that they’re going to ship. And so I’ve tried to adopt the same thing in every DevRel organization that I’ve been in, is what are the high-level goals?

And if you can get leadership to buy off on those, for example, we’re currently working on an online learning platform. We don’t have tight metrics about how many people should be registered and complete the course and be certified yadda, yadda, but we have a good sense that if we build this, it’s going to be very beneficial in a number of ways. And leadership agrees, and they’ve bought off on that, and they’ve signed their names to it. And so for us, what does success look like in terms of this is actually implementing that and shipping it.

Corey: It’s a really strange and really powerful thing, but you take a look at so many different companies who have done well and companies that haven’t done well, and the way that they engage not just with the ecosystem, but with the community specifically, in many cases seems to be the path that it follows. I mean, not to pick on them unnecessarily, but Chef had a wonderful community; they engaged absolutely flawlessly, from what I could tell, even when I didn’t agree with people or particularly like them in some cases, the people who worked at Chef almost demanded respect, and it was pretty clear, even as someone who didn’t use it myself, that they were a force to be reckoned with. And then they wind up effectively losing a lot of the people that made it special, the community moved on, they sold it to a company no one had ever heard of, and now it’s one of those, oof, they deserved a better end. Maybe that’s unfair, but that is the perception.

Jason: Yeah, I would say the same thing sort of happened with Puppet, the idea that they built a nice community, and back to my point of, like, you have a project, you work on shipping that, you don’t really track those numbers. That’s what I saw from both communities Chef and Puppet is they had these strong communities, they were doing things, and the goal was the community. And I don’t know—I haven’t talked to Nathan, I haven’t talked to folks at Puppet, but I suspect that they weren’t simply about how many people—like, what’s the total number of people that we would say are in our community? There was a value on, we want to do this thing and we have a sense of the quality of the community, and how much people just are engaged, and interested, and want to help each other.

Corey: The piece that also gets lost as well is companies are out there to turn a profit. And building a vibrant open-source community who loves your open-source offering but aren’t in a position to either champion or purchase the thing is often viewed as a complete waste of time by the business. So, they in turn, then pivot business models and do things that insult or alienate the community, and suddenly are perplexed by the massive groundswell of negative publicity they get, of people actively advocating that companies not use them. And their position is somewhat understandable in a form of, “What the hell is this? You weren’t spending money on us before. Now, you’re still not spending money on us, but you hate us. What gives?” Community is a weird thing to wrap your arms around.

Jason: Absolutely. I would say it’s hard to wrap your arms around it when you’re not valuing the relationship. It’s like any relationship where you have ulterior motives. If you can’t actually connect with people, it’s never going to go right.

Corey: No. And it also can’t be self-serving, or seem to be self-serving—spoiler, the best way to make sure you’re not perceived a certain way is to not actually be that way—we take a look at Last Week in AWS, my newsletter, it is explicitly aimed at people who want to keep up with what’s going on in the world of AWS, which is fair. It is not aimed at people who have a big AWS bill and don’t know what to do about it. And sure I reference periodically in that newsletter what I do, but it’s not a sales piece. It’s not every week hammering home, buy whatever it is I’m selling because that’s how you alienate and lose the audience.

I’ve always felt that by being top-of-mind for the problem and reminding people I exist every week with something that’s useful and ideally a bit funny, then, when they have that expensive problem, they’ll think of me. That was my theory four years ago, and I’m still here, so apparently, it wasn’t completely off base.

Jason: Yeah, well, that works, right, because nobody wants to subscribe to a newsletter to hear about the service. If they knew they needed your service, they would just buy your service. So, what’s the value of the newsletter? What’s the value that you’re offering to people? And that is, well, the fact that there’s so much freaking news about AWS every week that it does require a newsletter.

Similarly for me, what’s the value? Well, if people knew that they needed Gremlin, they would just come talk to me. But they don’t. They were concerned about the needs that they have, about how do I build a more reliable application, “My stuff’s always breaking. I’m having too many incidents. I’ve done everything that I can think of. What’s next.” So, it’s just offering that.

Corey: If your mean time to WTF for a security alert is more than a minute, it's time to look at Lacework. Lacework will help you get your security act together for everything from compliance service configurations to container app relationships, all without the need for PhDs in AWS to write the rules. If you're building a secure business on AWS with compliance requirements, you don't really have time to choose between antivirus or firewall companies to help you secure your stack. That's why Lacework is built from the ground up for the Cloud: low effort, high visibility and detection. To learn more, visit lacework.com.

Corey: And let’s be very clear here, you have a much harder challenge than I do. Because it turns out that you don’t need to be deep into the weeds of corporate finance, to understand the concept of wasting money on the AWS bill might not be the best thing in the world. Once you get more into the nuances, you start to realize, “Oh, being able to predict the AWS bill sounds super awesome, too.” But none of those are a particularly heavy lift, whereas, “Wow, your site is crappy and falls over a lot. Have you considered breaking it on purpose?” Sounds deranged the first time someone hears it.

Jason: Absolutely, yeah. That’s the number one thing that I hear all the time is—and people joke about it. I don’t need chaos engineering; I do regular deploys.

Corey: That sounds almost like someone was sitting in a blameless post mortem and got carried away trying to keep it blameless because otherwise, it was going to be their fault, and accidentally invented entire field.

Jason: Yeah, yeah. I mean, it’s definitely blameless if everybody is causing things to break; then we all share the blame. It is a funny thing. It’s a tricky thing to sell the people and I think it’s tricky because we have these misconceptions about what that actually means, the idea of breaking things on purpose. And trying to move away from that because the breaking really isn’t the goal.

And oftentimes, they’re not actually even breaking things; you’re stressing them out or you’re simulating things, so nothing’s really broken. But once you start thinking of it as that idea of I’m going to test my assumptions, right? I think that things work this way, but I don’t know, I’m not super confident that it actually will do that. And we do that all the time when we’re developing applications or infrastructure. I set things up, I’m pretty sure that it’s going to work a certain way.

Documentation says that this app works this way. Does it actually do that? Well, I can either find out when it doesn’t do that at some random point, or I can actually try to force it to act in that way, or to encounter that bad environment that I’m a little suspect about. And so we do this all the time with other things. And oftentimes, we’ll do this just mentally as, “What would happen if—” and you kind of play it out in your mind.

And that’s actually a great way to start with chaos engineering, rather than actually doing it, just that mental game. “What do you think would happen if this goes wrong?” Play that out in your head? Cool. Once you’re comfortable with that you’re like, I think this is what my next steps would be. I’m pretty sure there’s documentation here, or I’ve gone and checked and assured that there’s docs, or run books, or whatever, why not give it a try?

Corey: It’s one of those areas where what have you got to lose? I mean, as you just said, your site breaks all the time anyway, before you even touch it’s stability, what happens if the database just suddenly increases latency through the roof? What happens if suddenly all of us-east-1 is hard down? In many cases the answer is, we don’t really care about our website anymore because the world is not going to care about the internet not working that day, in the context of what we do. In other shops, yeah, that matters, and we kind of still need the power grid to work.

So, there’s a definite question of what failure modes are worth planning for and what aren’t, but even going through that exercise is fantastic. I used to do things like that from a sysadmin perspective, asking companies when I was asked to build out a mail server. “Great, how much downtime is acceptable?” And they said, “Absolutely none.” I said, “Great. I’ll need a budget of $20 billion to start, and when that runs out, I’ll come back for more.” And they said, “Wait, what are you talking about?”

And we said, “Oh, now we’re negotiating with the business.” And it turned out what they really meant was, “It would be nice if the mail server worked during business hours most of the time.” And, “Oh, okay. I can do that for slightly less.” And it really just came down to what do you value? What is important to your business?

Jason: Yeah. How much reliability do you need? Although one of the key things that I always point out is, a lot of times people are like, “Oh, you don’t need 99.9% reliability; you could probably get by with less than 90 because people aren’t using your application at night, they’re not using it on the weekends, yadda, yadda.” The other problem with that, though, is you rarely control when those outages happen.

So sure, if it happens in the middle of the night, and nobody’s using it, great. Just keep sleeping. As you start to work on this, though, there is the idea of it could happen at any time, so let’s actually test things to ensure that if it happens at the least opportune time, things actually work the way that we expect.

Corey: And that’s an incredibly valuable thing. See, you’re already convincing me on this. And clearly, you’re very effective at that advocacy role. How do you hire and how do you determine who’s a great fit? Because I’m imagining that bringing someone in, in an advocate role, and their position being, “Oh, at no point, can you ever measure me on any context, and just assume that what I’m doing is amazing and great.”

That becomes a hard thing to do. When I was talking to companies about possibly doing evangelist style roles, years ago, I asked, “How will you know if I’m being successful in this job?” And one of the answers was, “Well, you speak at a certain number of tier-one conferences a year.” “Cool, what are those?” And, they listed off a bunch and cool, there’s only one in that list that I’m not scheduled to speak at this year, so do I get a raise?

People try and aim at the wrong thing in their quest to articulate what they really value, but what they really value is hard to measure. So, how do you evaluate people on a basis of are they doing what they should be doing, or are there ways that they can be coached to improve, or are they just not effective in the role at all?

Jason: Yeah. Well, I think you mentioned two great things, are they doing what they’re supposed to be doing? And it comes back to every quarter, we’re laying out the goals of what do we want to accomplish this quarter? And we make them achievable, so hopefully, by the end of the quarter, you’ve achieved this thing that not only the team, but senior leadership has decided is a good thing for the company. And to that point, if it’s not, if we do that thing and nothing happens, and it’s—or it’s bad for the company, at least we can say, “Hey, senior leadership, you are the people that thought this was a good idea, too.” But that said, we try not to do the blame. We try to iterate on things and experiment a lot. Especially at Gremlin, we’re all about experimentation, so we’re constantly trying things. But ultimately, it’s are you getting this thing done that we’ve agreed that we’re going to get done?

But you also mentioned that second thing about growth. I think that’s something that I always look for with anybody, whether that’s DevRel or engineering. I want people that are interested enough in the job that they want to do it well. There’s something about it that they really love or they’re really into, and they want to master that. And so part of my goal as a leader is trying to help people along that path of what do you find interesting? For example, last year, we were working on those tiers, as we’re trying to figure out what does it actually look like. Because we’re really small team at Gremlin, and so as I’m starting to consider how do I promote people?

What are the various, like, levels or tiers of going from an advocate, to a senior advocate, to whatever is beyond that? So, I asked the team, really, “What do you think that would look like? What do you think the next level for your career is? What is the thing that you want to master?” Because ultimately, people have more investment when they’re choosing their destination and they’re choosing their direction.

And so if I can help people do that, just define what’s the next thing that you want to tackle? What do you think mastery or the next level of your career looks like? How can we help you get there? So, that’s what I am for.

Corey: For better or worse, it seems to be working. I remember back when Gremlin was a rando startup idea a couple people had and now I’m starting to see you folks, basically everywhere.

Jason: Yeah. Again, we’ve got a small team, but it’s a great team. So, Ana Medina has been on the team, actually, before I joined, but she’s been doing a fantastic job and she has been working on a lot of our educational outreach. And then Pat Higgins on the team actually started on the engineering side. So, he was one of our front-end engineers; he’s been working on a lot of really great tools.

He helped me restart the Break Things On Purpose podcast. So, we’re into season two of that now—and by the way, we should have you on that show as well. But yeah, we’re doing a lot of fun stuff, and folks are happy. So, try to keep them challenged, and we’ll see what’s next.

Corey: Yeah, I’m really looking forward to seeing how the story continues to evolve. It’s a fascinating field that went from, “That is ridiculous,” to, “Oh, that’s great but it would not apply to what I do,” to, in my case, it actually would not help me in any way with what I do because it turns out, well, what if an AWS region goes down and you can’t produce your newsletter the usual way? Oh, I’ll write it by hand that way because suddenly I have a much bigger story to talk about that week.

Jason: I am curious, though, speaking of having you on the podcast. Oftentimes, we talk about reliability, and having never had to deal with AWS bills because they always go to somebody else in finance, I am curious how reliability ties into the cost of what you’re paying for AWS? Because I can imagine things like—a common thing that we hear about is, “I’m moving a lot of stuff to Lambdas.” Like, great. Serverless. It’s cool, it’s hot. How is that charged?

Corey: Right.

Jason: Obviously, by time.

Corey: Oh, yeah.

Jason: So, if it’s charged by how long something takes, what if your latency goes up? What if your resources are constrained? How does this actually affect things? And how does that impact how you think about reliability not just from a is it up or down? How’s my customer looking at it? But maybe from what your AWS bill looks like?

Corey: I love where you’re going with that. And it’s the conversations everyone loves to have as about three levels beyond where most companies actually are. Easy example that sounds like something in the distant past, but it’s very real today: I want to store data in multiple availability zones for durability purposes and making sure that we are reliably up. Well, every time a gigabyte crosses an availability zone boundary, that cost two cents. And then you have to pay to store it twice.

So, there’s a question of how much is having multiple sets of that data worth? And the cloud-native answer to that is, “Oh, put it in S3. There’s no cross-charges there. Their durability is ridiculous, and you can access it a whole bunch of different ways, provided your application supports it.” But that’s not a fit for everything.

And you find that saving money, and being reliable, are at some point completely at odds with each other. And this is incidentally, why we don’t do this as a tool, we do it as a consulting engagement. There are times where, for business purposes, you will want to spend more on reliability. Because saving money that accidentally takes your company down for a month is not money you should be saving.

Jason: Yeah.

Corey: Now, the real fun thing I want to see from Gremlin one of these days from a implementation perspective is, just for fun, we’re going to run a chaos injection experiment where we decide to cancel the credit card tied to the account and then also remove the increasingly frantic alerts from your email when that happens, and see how long it takes you to realize the giant single point of failure that no one really thinks about existing, but absolutely does.

Jason: So, I am curious, for folks that are listening who are engaged with the chaos engineering community, or at least follow Corey’s newsletter and have seen updates, AWS has announced their own chaos engineering tool, the Fault Injection Simulator, which to Coreys skill of poorly named things, that actually isn’t a simulator. It does inject real faults, so it may be—S should be service. One of their faults, though, that they can do is API throttling, which essentially could simulate the idea of, you haven’t paid your bill; we’re turning things off. So, Gremlin is working with the AWS folks, we’re trying to figure out great ways that we can work together so that people can use both Gremlin and AWS FIS. So, I’ll let you know if that becomes a thing, and maybe we can get some API access to billing as well.

Corey: I’d love to see it. Please keep me looped in. Thanks so much for taking the time to basically go all over the world of DevRel and probably make some lifelong enemies in the process. If people want to hear more about what you have to say, where can they find you?

Jason: Yeah, I’m on Twitter. My Twitter handle is @gitbisect—and by the way, if anybody tweets about Git bisect, it is a fantastic tool, fantastic utility within Git—oftentimes, I will respond. But that’s where to find me on Twitter. Otherwise, you can find me on [unintelligible 00:31:30] podcast, Break Things On Purpose. It’s available in all the platforms.

Corey: Excellent. We will, of course, put links to that in the [show notes 00:31:37]. Thanks so much for taking the time to speak with me. I really do appreciate it.

Jason: Yeah, thanks, again. It’s been long overdue, and I’m glad we finally made it happen.

Corey: Awesome. Jason Yee, Director of Advocacy at Gremlin, I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you hated this podcast, please leave a five-star review on your podcast platform of choice, along with a comment saying that the best thing to test breaking in production is your DevRel team.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

This has been a HumblePod production. Stay humble.

View Details

About Cassidy

Cassidy is a Principal Developer Experience Engineer at Netlify. She's worked for several other places, including CodePen, Amazon, and Venmo, and she's had the honor of working with various non-profits, including cKeys and Hacker Fund as their Director of Outreach. She's active in the developer community, and one of Glamour Magazine's 35 Women Under 35 Changing the Tech Industry and LinkedIn's Top Professionals 35 & Under. As an avid speaker, Cassidy has participated in several events including the Grace Hopper Celebration for Women in Computing, TEDx, the United Nations, and dozens of other technical events. She wants to inspire generations of STEM students to be the best they can be, and her favorite quote is from Helen Keller: "One can never consent to creep when one feels an impulse to soar." She loves mechanical keyboards and karaoke.

Links:

  • Netlify: https://www.netlify.com/
  • TikTok: https://www.tiktok.com/@cassidoo
  • Newsletter: https://cassidoo.co/newsletter/
  • Scrimba: https://scrimba.com/teachers/cassidoo
  • Udemy: https://www.udemy.com/user/cassidywilliams/
  • Skillshare: https://www.skillshare.com/user/cassidoo
  • O’Reilly: https://www.oreilly.com/pub/au/6339
  • Personal website: https://cassidoo.co
  • Twitter: https://twitter.com/cassidoo
  • GitHub: https://github.com/cassidoo
  • CodePen: https://codepen.io/cassidoo/
  • LinkedIn: https://www.linkedin.com/in/cassidoo

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Thinkst. This is going to take a minute to explain, so bear with me. I linked against an early version of their tool, canarytokens.org in the very early days of my newsletter, and what it does is relatively simple and straightforward. It winds up embedding credentials, files, that sort of thing in various parts of your environment, wherever you want to; it gives you fake AWS API credentials, for example. And the only thing that these things do is alert you whenever someone attempts to use those things. It’s an awesome approach. I’ve used something similar for years. Check them out. But wait, there’s more. They also have an enterprise option that you should be very much aware of canary.tools. You can take a look at this, but what it does is it provides an enterprise approach to drive these things throughout your entire environment. You can get a physical device that hangs out on your network and impersonates whatever you want to. When it gets Nmap scanned, or someone attempts to log into it, or access files on it, you get instant alerts. It’s awesome. If you don’t do something like this, you’re likely to find out that you’ve gotten breached, the hard way. Take a look at this. It’s one of those few things that I look at and say, “Wow, that is an amazing idea. I love it.” That’s canarytokens.org and canary.tools. The first one is free. The second one is enterprise-y. Take a look. I’m a big fan of this. More from them in the coming weeks.

Corey: This episode is sponsored in part by our friends at Lumigo. If you’ve built anything from serverless, you know that if there’s one thing that can be said universally about these applications, it’s that it turns every outage into a murder mystery. Lumigo helps make sense of all of the various functions that wind up tying together to build applications. It offers one-click distributed tracing so you can effortlessly find and fix issues in your serverless and microservices environment. You’ve created more problems for yourself; make one of them go away. To learn more, visit lumigo.io.

Corey: I’m Corey Quinn. I’m joined this week by Cassidy Williams, principal developer experience engineer at Netlify. Cassidy, thanks for joining me.

Cassidy: Thanks for having me.

Corey: So, you’re famous in many circles for things that have nothing to do with your actual job. Or at least that’s the perception. So, let’s at least start there because I’m not sure we’ll get back to it. What is Netlify? And what does a principal developer experience engineer do at such a place?

Cassidy: Yeah, so the shortest answer is, it’s a place where you can host your website. The longer answer is it’s a whole development workflow. You can build whatever types of complex websites that you want, and we make it very easy to get it up and running. And my job there is on the developer experience team. And basically, what we do is we are developer experience engineers. We try to build things and show developers how to make their apps, their websites, their

various products, and projects easier to build on Netlify.

Corey: Sort of the whole idea of what I used to think of, I guess, as static websites and various ways to host it, which I think is now called Jamstack. But that probably also misses a fair bit of nuance because I’m going to be completely transparent here: I am crap at all things frontend.

Cassidy: It takes all kinds to make a project work. Yeah, so it is more than static. I like to think of it more as static first. The way I’ve defined Jamstack, that kind of clicks with most people is, writing Jamstack—and for those who don’t know, it initially was an acronym, where it was:, JavaScript, APIs, and Markup stack. And so, it’s less about technologies and more about the philosophy of building websites.

But the philosophy of it is, it’s kind of like building mobile applications, but in the browser, where you try to build as much as you can upfront, and then pull data in as needed. Because in a mobile application, when you have something native, you don’t, server-side, render the UI every single time. The UI is built pretty—

Corey: Well, not with that attitude anyway.

Cassidy: [laugh]. That’s true. That’s true. But when you’re on a mobile app, you don’t normally pull in the UI every single time. It’s built-in, and then you pull in data as needed; sometimes it’s local, sometimes it’s on a server somewhere. And that’s what Jamstack is all about. It’s building as much as you can upfront and then pulling in data as needed.

Corey: The idea is incredibly compelling, and it gets at a emerging trend that I don’t think that there’s any escaping, and—maybe this is overblown, I’d love to get your feedback on it—I can’t shake the feeling that JavaScript is the future—not necessarily a frontend—in general, when it comes to, effectively, computers. We’re seeing it on the backend, we’re seeing it on the frontend, the major cloud providers are all moving in a direction of approaching folks who have JavaScript experience, and that’s the only certainty in that persona that they wind up identifying. It is very clearly not going away while getting more capable. Is that fair? Is that missing something? What’s the deal there?

Cassidy: I keep hearing there’s, like, a rule that people are saying, like, “If it can be built in JavaScript, it will,” because I think it started as kind of this toy language that people didn’t really take seriously. But it has not only become more powerful, but also browsers have become more powerful too, and you can just build more and more with it. And because it’s kind of a low barrier-to-entry language, it’s relatively simple to at least initially learn JavaScript before you get into all the nuances of everything, that I think, just because there are more people using it and it’s easier and faster to pick up then something like assembly or C++ or something. I hesitate to make generalizations because you never know, but it does feel like that sometimes, that JavaScript is just the way that things are going.

Corey: And I admit, a couple of times I have tried to get into the JavaScript world, and it isn’t clicking for me. My lingua franca is crappy Python. And it’s just crappy enough to run, but it’s neither elegant nor well-designed. It is also barely functional. And every time I have brought in an actual developer to turn some of my scripts into something a bit more robust, they ask me what it does, they smile and nod a lot and never take their eyes off me for a second, and then immediately get rid of everything I might possibly have touched.

This is, of course, a best practice where I’m involved. But it runs. Like, “This is the worst code I’ve ever run.” “Ah, yes, but it does run.” The problem I have with JavaScript is that I do not understand it. The idea of asynchronous calls on a

browser completely melt my brain whenever I look at it.

That’s caused a few of my early naive mistakes where, “Oh, go ahead and set this value and then use it here down below, and—wait. Why is it completing before it has that value and it’s not you—what is going on here?” And now I understand the general principles of it, but I’m still getting lost and confused in the weeds. Now, is this just another expression of being secretly terrible? Or is there a nuance here that I’m not picking up on?

Cassidy: I was smiling the entire time you were saying this because I feel like almost everybody who is new to JavaScript coming from another language has had the exact same issues. So, you’re not alone, and you’re not a total idiot. [laugh].

Corey: So, I decided that it was time to learn it the second time, and I—all right, I’m going to break my own rule, which is the way I normally learn something new is I’ll dive into it and start building something and then we’ll see what happens. Sure, it means I’m a full stack overflow developer, and my primary IDE is copying and pasting, but I can get something sort of functional that works. That approach wasn’t working for me, so what I did on my second attempt was odd. I’m going to go actually do the unthinkable for me, and read some documentation and/or some tutorials.

And I was almost immediately blown off course there because suddenly, I find myself just wandering onto what I can only describe as a battlefield between all of the different frameworks I could have chosen between, and it seemed like the winning move was not to play. What am I missing? Are these frameworks hard requirements for doing anything that even remotely resembles frontend in a responsible way? Are they nice-to-haves? Is it effectively an aside current debate that I got suckered into and lost the forest for the trees?

Cassidy: You probably got sucked into many debates because there are so many in this world, I do not think you need a framework to do complex web apps or any web apps. I mean, my personal website, as much as I love React—and I’m deep in the React world—I did that with vanilla HTML, CSS, and JavaScript, and that’s all it is. And plenty of the projects that I do, I start with vanilla, and then I add React as needed. I think it’s something where these frameworks, you don’t need them, but it’s really nice once you start building large applications where you don’t want to reinvent the wheel. Because there have been plenty of times on my own projects on other projects, where I start to basically start implementing state-driven components, and trying to parse templates and stuff that I end up making for myself. Where if I did React, I probably wouldn’t need to actually implement all of those. And so you don’t need these frameworks. That being said, they can be very helpful as you make more complex projects.

Corey: So, I periodically post an architectural diagram of the pipeline slash workflow thing I use to write my newsletter every week. And I was on the verge of just hiring a frontend developer to build something frontend because it turns out that there’s not a great experience in using a whole bunch of shell scripts that require a CLI to post at random API endpoints. And then a discovered Retool, which is one of those low-code tools that more or less is Visual Basic for frontend. It was transformative because suddenly, it’s, “Oh. When I click this button, make this query that hits some API that I can define,” and oh my stars. It was transformative, and I was actively annoyed I hadn’t discovered it years ago.

Cassidy: [laugh]. Yeah, all of those low-code tools for web devs, they’ve been growing, that is a really interesting realm of the web that I’m curious about. I’ve played around with quite a few of them, and some of them, I kind of end up just wishing that I built it myself in the first place, and then for some of the others, I’m like, “You know, this saved me some time.” And yeah, I think those things are really, really powerful. I don’t know if they’ll ever fully replace having an actual developer, but for a lot of individual smaller tasks, it’s really nice to not have to, again, reinvent the wheel.

Corey: And you’re right. These tools are getting more capable. The problem I have is, whenever I talk to the teams building these things, they’re super excited about them and can’t wait to show them off. And then I say, “Just a quick question. Of all your engineers here, how many of them don’t know JavaScript?”

And the answer is always the same. None of them? Great. Yeah. Now, there’s an opportunity to present this to existing frontend developers so they can get back to what they were doing when they build a quick internal tool for someone else in a business unit, but there’s an entire untapped market of people like me who don’t understand JavaScript. So, when we see these things described in JavaScript context, it looks like it’s not for us, even though it very much is. There’s something to be said for making things accessible to an audience that

would benefit from them.

Cassidy: Yeah. I’ve actually given a few talks where it’s geared towards a backend developer who might want to dip their toe into frontend but have no idea where to start. And that is a whole world of people who are like you who just don’t understand the DOM in the browser, and how the interactions happen, and how the async await stuff works, and how promises work and everything. And they’re very weird concepts that just aren’t in other parts of programming, typically. And I think that’s a marketing problem where a lot of these low-code tools or no-code tools don’t understand the opportunity that’s

available to them.

Corey: I think that there’s a misunderstanding in many respects, where I’ve also seen a fair bit of, I can only call it technical bigotry, I guess, is the best framing here of, “Oh, where frontend is easy, and backend is the hard stuff, and that’s really where it’s at.” And having worked with qualified teams on both sides and looking at all the intricacies on both sides, where the hell does that come from?

Cassidy: You know, I think it just comes from the past.

Corey: So, do I. And I don’t agree with it. It’s just such a misunderstanding and a trivialization of such a valuable area of things. It kills me every time I see it.

Cassidy: Yeah, it’s frustrating, I admit, because I’ve faced that a lot in my career. I actually—I used to do backend. I used to do Python stuff, and I have a computer science degree, but plenty of times, there’s some kind of backend dev who’s just like, “Eh, well, I know HTML and CSS, so I know frontend.” And that’s about it. Or they’ll say, “Well, do you really need to know this kind of algorithm or this way of doing things in an optimized way because you’re just putting a pretty face on the data that we’re producing for you.”

And it’s an annoying sentiment. And I really think that it’s just from a previous time because a long time ago, from five to seven to ten years ago, that might have been more true because we didn’t have some of these frameworks that do a lot on the frontend. And we didn’t have things like GraphQL, and really powerful tools on the frontend. Where back then, it was a lot of the backend doing stuff, and then the frontend making it look good. But now the work is distributed a bit more where our backend teams, I can say, “Build however you want. You can change your language to Rust, to Go, to whatever, do whatever you want; as long as the data is exposed to me, I can use it and run with it.”

And then all the routing ends up happening on the frontend, all of the management of that data happens on the frontend, all of the organization and optimizing for the browser happens on the frontend. And so I think both sides have been empowered in recent years in that regard because, again, with that modularity, you can scale a lot better, but those lingering sentiments are still there. And they’re annoying, but unfortunately, we’ve got to live with them sometimes.

Corey: So, let’s talk as well about, I guess, sort of the elephant in the room. Your Twitter feed is one of the most obnoxious parts of my day, specifically because every time you post something I am incredibly envious about the insight it provides, the humor inherent in it. “I wish I had thought to go in that direction,” is almost always my immediate response. And, ugh, it kills me. Let’s talk a little bit about that. How did it start? And how is it continuing?

Cassidy: That’s a good question. So, I’ve always been a bit of a clown, both on and off the internet, but I was never very, very public about it, for a while there. Either that or just had a small audience and people were just like, “There she goes again. Maybe she’ll shut up someday.” And so I’ve always had those little drops of humor where I can because I think I’m amusing myself at least.

But about a year and a half ago, I discovered TikTok. And with TikTok, basically, it has such a good video editor—that was the only reason why I got the app because it made it so easy to make videos on my phone—where I was able to suddenly not just type my tweet jokes and my snarky humor, I could make a video about it, I could add music to it, I could make a dumb face. And people seem to like it, and it’s worked out.

And I try to approach things rather from a realistic or educational perspective first and then drop in the humor later, I don’t try to lead with the joke, but at the same time, it’s always fun to have a joke in there because people like to say, “Oh, something funny is happening. I’m getting ready for it.” And it’s kind of fun that I’m able to do that a lot more now that people actually expect humor.

[laugh].

Corey: When I was an employee—which I was, let’s be very clear here, terrible at. There is no denying that—it was always a problem for me where the biggest fear that anyone had—start to finish—was that I would open my mouth and say something. And credit where due, my last job was at a large finance company. And at that point, they’re under such scrutiny that anytime someone opens their mouth on anything, it has the potential to trigger an SEC investigation, and no one knows what I’m going to say. Yeah, there’s a lot of validity and being concerned about that.

I felt like I couldn’t ever just shoot my mouth off and be me. And I always had this approach of, no company in the world would ever be willing to tolerate my shenanigans, therefore, I should never look to either do these things in public or later, go to be an employee again. You’re living proof that it is in fact possible to have both.

Cassidy: Yeah. It brings a levity to our very serious industry—I used to be in FinTech; I know how serious that can be—but then just in tech in general, a lot of tech people take themselves way too seriously. And I understand we’re doing awesome work. Some people think they’re gods because they can think something and make it into an app. There’s ego there, but I feel like making fun of the problems, pointing out the problems in the industry and, kind of, just making light of it and making certain tech jokes and making certain concepts humorous as well as educational, I think bringing that approach to things is just

really, really effective.

And I’m really happy to be on my team, honestly, at Netlify because a bunch of them are just dorks [laugh] where pretty much every single meeting, we try to make it a little bit fun. And it makes our meeting so much more enjoyable and productive because we’re not just seriously staring at our screens and saying, “Okay, let’s make this decision for our OKRs,” or anything like that. We have a good time in these meetings while being productive, and it makes for a really nice team dynamic. And I think there should be more of that, in general, in tech.

Corey: One of the things that you have always done with your platform that I am, I guess, slowly warming up to is that you’re never mean, or in the rare occasions where you punch at something, it’s a dynamic; it’s not a company and it’s not a person. I have a strong rule of not punching at people, but large companies have always been fair game from my perspective. And that is a mixed bag. Yours is—how to put this—unrelentingly positive where it’s always about building people up, and shining a light on things that used to be confusing, and reminding people that they’re not alone in being confused by those things. And that’s no small thing.

Cassidy: Yeah. I appreciate you noticing that. I do try to do that, not only, necessarily, to be just like, “I want to be the positive star in tech,” but also because you never know what someone is dealing with, and someone might be pretty mean, and there have been plenty of people who have said some not great things towards me or towards other people and that cuts deep. And so I do try to avoid those kinds of pointed things. Believe me, it’s difficult; sometimes I do just want to call people out and be just like, “I know what you did to this group of people, and I hate it.” But you never know what people are going through, and I’d rather just make sure that the people who are doing well are the ones who are uplifted, and they get the attention that they need, or deserve, rather.

Corey: I did a little research—I know, I know; shock—before I wound up inviting you here, and it’s not just your Twitter account. It’s not just your TikToks, it’s not just your weekly multi-hour livestreaming on Twitch—or ‘Twetch’ or however it’s pronounced. I’m old, and that’s fine—it’s not the platforms; it’s the fact that no matter where you are, you’re constantly teaching people things. And I want to be clear, that doesn’t seem like it’s in your job description, is it?

Cassidy: No, but it’s something that I really care about. I really like teaching in general. A lot of the resources that I provide and the things that I do are me trying to give people things that I didn’t have when I was in the industry, trying to give advice that I wish I had, trying to give resources that I didn’t have. Because a lot of times, people don’t know where to look, and if I can be that person that can help them along, some of the greatest joys I’ve ever felt have been when people say, “This blog post that you wrote helped me get my first job,” or, “This thing that you said, was the kick in the pants that I needed to start my own company.” Little things like that. I love hearing it because I really just love making people successful and helping them get to that next step in their careers. And that’s my passion project, and I tried to do that and all the things that I do.

Corey: There’s really something to be said about being able to reach people who have pain and have needs. I mean, the one crossover talk that I gave that really transformed the way that I saw things was “Terrible Ideas in Git” because if there’s one thing that unites frontend, backend, ops folks, data scientists, et cetera, et cetera, et cetera, it’s Git as being the common thing that no one really understands. And by teaching people how to use Git, first, it was sort of my backdoor, sneaky hack into finally having to teach myself how Git works. But then it was a problem of where, now I need to go ahead and find a way to present this in a way that’s engaging, and fun, and doesn’t require being deep into the weeds. And I was invited to speak at Frontend Conference, Zurich,

which was just a surreal experience.

Incredibly nice people, very gracious community and I’m sitting there for the first half of the day watching the talks, and it’s a frontend conference and everyone’s slides are gorgeous. And this was before I started having a designer help me with my slides. So, it was always a black Helvetica text on a white background. And mine looked like crap, and I only had a few hours until my talk, so what do I do because I’m feeling incredibly out of place? I changed the font on everything to Comic Sans and leaned in on that.

And it definitely got a reaction. The talk was great. It really did work. And it was fun. And in hindsight, I don’t think I’d do it again because I keep hearing rumors that I can’t quite confirm, but it’s significant enough that I want to be clear, that Comic Sans is apparently super accessible when it comes to people with dyslexia, and I don’t want to crap on something like that. It’s not funny when it makes people feel out of place.

Cassidy: Yeah. These kinds of things, it’s delicate to talk about because you have to figure out, okay, how can I make this accessible to as many people as possible? How can I communicate this information? And then, meanwhile, when you are this person, that just means your DMs are very, very full of people who want one-on-one help and you have to figure out how to scale yourself, and how can you make these statements that are helpful for as many people as possible, provide as many resources as you can, and hope that people don’t feel bad when you can’t answer every DM that comes your way. And yeah, there’s a delicacy when it comes to all the different things that you could be poking fun at, or saying you don’t like, and stuff, and my answer to pretty much everything has turned into just, “It depends.”

Whenever people are just like, “What’s the best framework to learn?” I’m kind of like, “Eh, it depends on what you want to build.” Because first of all, that’s true, but second of all, there’s enough opinions out there in the world saying, like, “This is the worst font.” “This is the best font.” “This is the worst way to build web apps.” “This is the only way to build web apps.” I mean, you hear this constantly throughout the tech industry. And I think if more people said, “It depends,” we would be a [laugh] much happier industry in general.

Corey: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the Enterprise (not the starship). On-prem security doesn’t translate well to cloud or multi-cloud environments, and that’s not even counting IoT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IoT devices, detects these threats up to 35 percent faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at extrahop.com/trial.Corey: I really think that you’re right, and I think the hardest part is getting there. You say that the answer to, “What framework should I pick?” Is, “Well, it depends.” And that’s very true. The counterargument is that it’s also supremely unhelpful. It’s—

Cassidy: Right.

Corey: —“I’m looking to build a web page that has a form on it, and when I click a button, it does a thing.” And at that point, it feels like it’s, “Well, there are an entire field of yaks before you, all of them need to be shaved before the form will exist.” And it just becomes this. “Oh, my god, are you just trying to tell me not to bother?” And no, that’s never the response.

But having a blessing, I guess, golden path of where you can focus to get something done, and then where it makes sense to deviate gets signaled, I like that approach. But people are for some reason worried about being overly prescriptive. And I get that too.

Cassidy: Yeah, there’s a balance there. But I should append to my previous answer. I say, “It depends, but here’s how I would do it.” And that gives some direction. Some people might be just like, “Oh, well, I don’t want to use React,” or something like that, and I’m like, “Well, then, unfortunately, I can’t help you. You’re on your own. But I’m sure it’ll work for you.” And just kind of roll with it from there because you never know.

Corey: Yeah, what I’ve never liked the questions that the asker already has an answer they want to hear, and they’re looking for, almost, confirmation bias.

Cassidy: Yeah.

Corey: Yeah.

Cassidy: That’s common.

Corey: At that point, why bother? Just say, “This is what I’m thinking about doing. Please tell me it’s not ridiculous.” And if it is, people will generally try and

be kinder about it. But we’ll see.

Cassidy: Yeah, a lot of times, too—and I hate to say it, but a lot of times, too, people come in with such an arrogant air, and oftentimes, that’s either because they’re insecure about something, or they don’t have a lot of experience in something. But [unfortunately 00:23:27], that’s almost always the case. There have been times on my stream, for example, where someone will say, “If you use this framework, it will solve 99% of your problems.” And I’m kind of like, “Eh, will it though?” And I don’t want to just straight up say you’re wrong, but I kind of have to keep asking questions and try to be one of those teachers where I’m saying, “Okay, I’m going to ask you these questions. Are you sure that this edge case is in that 1%? I think you’re being a little bold here.” And not trying to specifically humble them, and know that they are wrong, but also turn it into a moment where you have to learn that nothing really solves 99% of your

problems. [laugh].

Corey: And whenever someone says something like that, I always assume conflict of interest somewhere. It’s like, “With this framework you’re suggesting, I don’t know, just so happened to integrate super well with the thing your company does? Huh, how about that?” Whenever someone can’t identify an area that they’re offering is crap in, I assume that they’re, effectively, evangelizing something with almost a religious fervor, and aren’t really people to take overly seriously. I have technologies that I adore, but if I can’t articulate use cases in which they would be wildly inappropriate, then I’m not really being fair, either to the person I’m talking to, honestly, the product itself.

Cassidy: Exactly. There’s always cons. Yes, there might be a lot of pros and the pros may outweigh the cons, but you have to be able to speak to those if you’re going to give a credible answer to any sort of recommendation like that.

Corey: So, let’s talk about platforms a little bit. You have a newsletter which I’m a fan of, and will of course link in the [show notes 00:25:05]. You stream on Twitch, which is similar to a podcast, only it’s video and it’s live so, unlike here where we can edit heavily if someone winds up breaking down crying, like I tend to every third episode—

Cassidy: Yeah, we should cut out those farts earlier, by the way.

Corey: Oh, yeah. Oh, we’ve already edited that out.

Cassidy: Okay, great. [laugh].

Corey: We’re already set. We do this in real-time here. But you have to do things like that in real-time on Twitch; as soon as something happens on camera, it’s done, it’s out there, and it’s a very different experience. You do it also on hard mode, where you and I are having a conversation back and forth, whereas when you do Twitch, you’re doing it solo. You are effectively in an empty room—or what appears to be one anyway—and you’re talking to the camera, and there’s

no other audio other than you and a lovely backing track.

There’s no conversation, you are monologuing for the duration of that. People mention things in the chat with a slight delay, and then you can take action based upon that. But that feels like an awful lot of pressure to wind up filling the dead air while you’re waiting for the next question to come in.

Cassidy: Yeah, it’s something that has taken practice. And I think it’s something that because I have done quite a bit of public speaking, I’ve done a bunch of teaching, I am comfortable with the silence. And the music also helps that a lot. Some people when they are about to livestream, or they’re learning how to livestream for the first time they kind of panic at the silence. They’re like, “Oh, my gosh, how am I going to fill it?”

Meanwhile, with me, I’m just like, “Ah, nobody’s asking a question. I can take a drink of water now.” And try to keep it as natural as possible. I try to make this stream—I started doing it more regularly during the pandemic, as something that’s kind of just co-working and kind of having something in the background, because usually when people are in the office or working at a cafe or something, you get to hear interesting conversations, and a voice, and you can chime in on occasion. And I try to make that what the stream is where people don’t have to be paying excessive attention, but I open it up where you can ask me pretty much anything and I will give you an honest answer, and just try to make it a space where people can not worry about asking a stupid question because I think that none of these questions, whether it’s about tech jobs, or certain frameworks, or opinions about things, none of them are dumb.

Sometimes it’s just people who aren’t sure what the answer should be, or they aren’t sure if their biases are correct or anything like that. And I really enjoy the livestream because it gives me a connection with the community that I can help teach further. And then as they ask questions, I can take that and run with it, and build a demo, help them come up with project ideas, show how I would build something, something like that.

Corey: Oh, there’s an incredible authenticity to what you do, and that is, I think, one of the most impressive aspects of what you do. I’ve never yet seen you make someone feel like a jerk for asking a question. I’ve also never once seen you claim you knew how something worked when you didn’t. You point people at resources to find the right answer. You are constantly gracious, you’re always incredibly authentic, and it’s become really easy to consume your materials because I know you’re not going to make it up if you don’t know the answer. And that’s no small thing.

Cassidy: Thank you. [laugh]. I appreciate that. It’s not easy, but it’s very fun. And I do hope that it makes people more comfortable with the concept of streaming, coding, and any of that.

Corey: You also seem to have some of the same problems than I do, specifically—not the jerk problems. That’s unique to me—but the problem in the context of answering a difficult question, namely, “So, what is it you do?” Because as mentioned, you have the newsletter, you have the job, you have the Twitch stream, you have the TikTok, you have the Twitter. You do courses from time to time, if I’m not mistaken, as well?

Cassidy: Yes, I do. I have a few online courses on Scrimba on Udemy on Skillshare on O’Reilly. I like teaching JavaScript and showing people how React works, and stuff, under the hood. And you’re right, it’s hard to explain what I do sometimes. [laugh].

Corey: And that’s the hard part is when someone asks, “So, what do you do?” What’s your default answer?

Cassidy: I have created this tagline that I’m kind of just sticking with, and we’ll see how long it lasts me. But I say, “I make memes, streams, and software.” And I just kind of leave it at that, and people be like, “Okay, Cassidy, shut up.” [laugh] and I leave it at that. But yeah, if someone asks me what I do, I kind of start with, “I code.”

And then if they press further, I’ll be like, “Well, I teach people how to code, and I show people how to code best.” And usually, that’s where my grandpa stops asking. He’s just like, “Okay, it’s that computer stuff.” If it’s a tech person, I start diving more and more into all of the things, and it’s very hard to explain. I wish there were a word for trying to make people laugh, and teach, and build things, and stuff, but I don’t know what that word would be.

Corey: Yeah, it’s a hard problem. My answer has always been to spin it depending upon who I’m talking to.

Cassidy: True.

Corey: If it’s at a neighborhood barbecue and people ask what I do, I try and

make myself sound like some sort of esoteric accountant because if I say even slightly incorrectly what I do, suddenly people are asking me about their printers. And honestly, how do I fix a printer? I throw it away and I buy a new one, but that’s not really helpful to people who are looking for actual help. So, it’s a matter of aligning what I do with people’s expectations. “I make fun of Amazon for a living,” is technically accurate, but boy does that get some strange looks.

Cassidy: [laugh]. Yeah, it definitely, definitely varies on the audience. If I’m, for example, going to some kind of church barbecue, I just say, oh, I’m a software engineer. Questions stop there, and I leave it at that. If I’m at a tech meetup, I’ll be just like, “Oh, well, I specialize on frontend things, but I also do some dev advocacy and stuff.”

And I can generally stop there. But you’re right, depending on the audience, I

have to be careful because I don’t want people to just ask me to fix their WiFi all the time, even though they do anyway. And to them. I usually say oh, I build computer things. I don’t know how to work them, though. And I leave it at that.

Corey: Oh, hey, I’m building a computer, too. Can you recommend some parts? Absolutely not. Is my—

Cassidy: Nope.

Corey: —I don’t know what I’m doing there.

Cassidy: I kind of just Google and accept whatever I’m told. [laugh].

Corey: Yeah. And the other side of it, too, is if you’re not direct enough and say, “Working with technology,” people tend to think that you’re being condescending. It’s like, “Oh, I do some cloud computing finance work.” And they’re like, “Oh, so what, you fix an AWS bill?” Yeah, exactly. “You could just say that, you know?” “Well, yeah. To you, but there’s a whole world of people out there to whom that sounds like I’m blowing them off with geekspeak.”

Cassidy: Yeah. Yeah, exactly. And it’s almost harder if it’s a mixed group of people, too, because sometimes people who are in tech but I don’t know the rest of the people, they might say, “Oh, she makes tech jokes on Twitter.” And they’ll say, “Oh, really? Say something funny.” I’m like, “Uh—I don’t know how.” [laugh]. It’s not that easy. It’s interesting trying to figure out how you define that for other people.

Corey: “Oh, you’re a comedian. Great. Make me laugh.” Like, “Oh, God.”

Cassidy: Just please, no.

Corey: Yeah, that’s the best setup for a good belly laugh is command performance of, be funny when you weren’t expecting it?

Cassidy: Yeah. Ugh, can’t handle it. I just freeze up and give up.

Corey: Ugh. Again, these are not common problems. One thing that I did find incredibly funny was that when we started talking, we talked about things that we had encountered as we wound up going through expanding audiences on Twitter and whatnot. And you sent a screenshot, at one point, of tracking your Twitter follower count over time in a private Slack channel that you had. And you said, “This is ridiculous, and no one ever does it.” And then I responded with a screenshot of me doing the exact same thing, which is—

Cassidy: So funny.

Corey: —first, hilarious because I’ve never seen anyone else do that. And, two, a bit of product feedback, perhaps, for the team at Twitter.

Cassidy: It really is. Yeah, no, when I found out you did that, too, I laughed so hard because so many times people have been just like, “You know there’s tools for this? You don’t have to just write a number in DM to yourself on Slack.” But this is the tool that works for me. It’s quick. It’s done. I can see, generally, how things are going. Someday I should put it in a graph of some type, but eh.

Corey: But it’s always forward-looking, too, because all those tools don’t go back in time to your account’s inception. And, “Oh, you had this person follow you at this time.” There’s no historical record there.

Cassidy: Yeah. It is totally product feedback. I have no idea how I’d be able to say, “Hey, look at this DM, fix this problem,” to a specific Twitter person, but, eh.

Corey: Four years ago, I had 1500 Twitter followers and it had taken me seven years to do it. And people ask, “What were the big inflection points when you wound up getting significant audience boosts?” And if I had dates on that stuff, I could absolutely do some correlation like, “Oh, there’s re:Invent.” “Oh, that’s where I was visibly thrown out of a bar on the news.” Kidding. But being able to tie it to things like that would be helpful, but it’s happened, it’s gone. I just have to basically try and remember, and assume I’m somewhat close to accurate.

Cassidy: Yeah. And I don’t do it consistently, mind you, there’s definitely weeks where I just totally miss it. But sometimes, for example, if I’m about to tweet something funny, I’ll mark it and then make the post and just see where it goes. And it’s more just interesting for me; I will probably never share this with people, besides you when we talk about our [laugh] strategies. But yeah, I mean, I guess that also speaks to building what’s best for you is often the best solution.

Corey: Yeah, and it changes, too. And the part of the reason that these conversations tend to happen behind closed doors because the easy, naive response is, “Oh, that’d be super interesting to watch and see how those problems get addressed.” But so much of what we’re doing and how we approach it is not helpful until you’re at a certain point of scale. If you have 200 Twitter followers, for example, frankly, you’re making better life choices and either one of us are, but the things that we are concerned with and have to pay attention to, just don’t apply in any meaningful way.

Cassidy: Right.

Corey: Conversely, if you have a small following Twitter account, that is a freedom that we don’t really have because past a certain point, as I’m sure you can attest, you can’t say that you like waffles without getting someone asking, “Well, what’s the problem you have with pancakes?” And then insulting you and following you around until you block them.

Cassidy: It’s so true. I was talking with someone about this yesterday because it’s not like I ever say things that are particularly controversial or anything, but word choice matters so much more when there are a lot of eyes on you. And so many times I’ll make a joke, and then I have to do a follow-up tweet saying, “This is a joke. Please don’t tell me how to exit vim.” Or something like that. Because oh, my word. People just will never take things the right way en masse.

Corey: No, I have learned there’s no possible way to say something without it being misinterpreted. And I try and wind up turning it back around, and every time I read something, I do my best to assume good faith. I don’t always succeed, and sometimes I look like a fool for basically taking a troll seriously, but I’d rather that than the alternative of someone asks a naive question, and I assume they’re just being a jerk and block them or I mock them. Because the failure mode of me looking like I got hoodwinked is better than making someone else feel crappy.

Cassidy: Right. Exactly. I remember a while ago, this was, like, a couple years ago, there was someone who was not being nice to me in the mentions, and I was just like, “Why would you respond to me like this? Just leave me alone.” I said something like that.

And it was a lesson for me and for them, where they ended up getting really upset with me and yelling at me in the DMs because they were getting all of this negative commentary on there and for being the mean one, and then I end up looking like a jerk because I ended up spotlighting this person who might have been having a bad day. You never know. And the algorithm works against you when you have a lot of eyes who are looking at what you’re tweeting about. And so, yeah, you have whenever stuff like that happens, you kind of just have to ignore it and learn to pick your battles, I guess.

Corey: Oh, yeah. And I assume that’s going on now. I imagine that one day, the AWS Twitter account is going to finally snap and just quote-tweet me with some incredible roast and there will be no coming back from that for me. I look forward to that day. It would be so nice to see that come out of them. I worry, I

may die before it gets there, but hope springs eternal.

Cassidy: [laugh].

Corey: Cassidy, thank you so much for taking the time to speak with me. If \

people want to hear more about what you have to say—as they damn well should—okay can they find you? Take a deep breath; run through the list.

Cassidy: All right, they can find me on all sorts of platforms. You could look up Cassidy Williams, and you’ll find either me or a Scooby-Doo character, and I’m not the Scooby-Doo character. Or you could look up cassidoo—C-A-S-S-I-D-O-O—cassidoo.co is my website, cassido on Twitter on GitHub on CodePen on LinkedIn all those platforms. That’s where you can find me.

Corey: And we will put links to all of those things in the [show notes 00:38:03] because honestly, that’s someone else’s job, and I am going to hurl that mess to them.

Cassidy: [laugh]. Perfect.

Corey: Thank you so much for taking the time to speak with me. I really appreciate that.

Cassidy: It was really fun. It was good chatting with you, too.

Corey: It really was. Cassidy Williams, principal developer experience engineer at Netlify. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an aggressive comment encouraging me to fight you on Twitch, however that might work.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Nick

Nick Frichette is a Penetration Tester and Team Lead for State Farm. Outside of work he does vulnerability research. His current primary focus is developing techniques for AWS exploitation. Additionally he is the founder of hackingthe.cloud which is an open source encyclopedia of the attacks and techniques you can perform in cloud environments.

Links:

  • Hacking the Cloud: https://hackingthe.cloud/
  • Determine the account ID that owned an S3 bucket vulnerability: https://hackingthe.cloud/aws/enumeration/account_id_from_s3_bucket/
  • Twitter: https://twitter.com/frichette_n
  • Personal website:https://frichetten.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Thinkst. This is going to take a minute to explain, so bear with me. I linked against an early version of their tool, canarytokens.org in the very early days of my newsletter, and what it does is relatively simple and straightforward. It winds up embedding credentials, files, that sort of thing in various parts of your environment, wherever you want to; it gives you fake AWS API credentials, for example. And the only thing that these things do is alert you whenever someone attempts to use those things. It’s an awesome approach. I’ve used something similar for years. Check them out. But wait, there’s more. They also have an enterprise option that you should be very much aware of canary.tools. You can take a look at this, but what it does is it provides an enterprise approach to drive these things throughout your entire environment. You can get a physical device that hangs out on your network and impersonates whatever you want to. When it gets Nmap scanned, or someone attempts to log into it, or access files on it, you get instant alerts. It’s awesome. If you don’t do something like this, you’re likely to find out that you’ve gotten breached, the hard way. Take a look at this. It’s one of those few things that I look at and say, “Wow, that is an amazing idea. I love it.” That’s canarytokens.org and canary.tools. The first one is free. The second one is enterprise-y. Take a look. I’m a big fan of this. More from them in the coming weeks.

Corey: This episode is sponsored in part by our friends at Lumigo. If you’ve built anything from serverless, you know that if there’s one thing that can be said universally about these applications, it’s that it turns every outage into a murder mystery. Lumigo helps make sense of all of the various functions that wind up tying together to build applications. It offers one-click distributed tracing so you can effortlessly find and fix issues in your serverless and microservices environment. You’ve created more problems for yourself; make one of them go away. To learn more, visit lumigo.io.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I spend a lot of time throwing things at AWS in varying capacities. One area I don’t spend a lot of time giving them grief is in the InfoSec world because as it turns out, they—and almost everyone else—doesn’t have much of a sense of humor around things like security. My guest today is Nick Frechette, who’s a penetration tester and team lead for State Farm. Nick, thanks for joining me.

Nick: Hey, thank you for inviting me on.

Corey: So, like most folks in InfoSec, you tend to have a bunch of different, I guess, titles or roles that hang on signs around someone’s neck. And it all sort of distills down, on some level—in your case, at least, and please correct me if I’m wrong—to ‘cloud security researcher.’ Is that roughly correct? Or am I missing something fundamental?

Nick: Yeah. So, for my day job, I do penetration testing, and that kind of puts me up against a variety of things, from web applications, to client-side applications, to sometimes the cloud. In my free time, though, I like to spend a lot of time on security research, and most recently been focusing pretty heavily on AWS.

Corey: So, let’s start at the very beginning. What is a cloud security researcher? “What is it you’d say it is you do here?” For lack of a better phrasing?

Nick: Well, to be honest, the phrase ‘security researcher’ or ‘cloud security researcher’ has been, kind of… I guess watered down in recent years; everybody likes to call themselves a researcher in some way or another. You have some folks who participate in the bug bounty programs. So, for example, GCP, and Azure have their own bug bounties. AWS does not, and too sure why. And so they want to find vulnerabilities with the intention of getting cash compensation for it.

You have other folks who are interested in doing security research to try and better improve defenses and alerting and monitoring so that when the next major breach happens, they’re prepared or they’ll be able to stop it ahead of time. From what I do, I’m very interested in offensive security research. So, how can I as, a penetration tester, or red teamer or, I guess, an actual criminal, [laugh] how can I take advantage of AWS, or try to avoid detection from services like GuardDuty and CloudTrail?

Corey: So, let’s break that down a little bit further. I’ve heard the term of ‘red team versus blue team’ used before. Red team—presumably—is the offensive security folks—and yes, some of those people are, in fact, quite offensive—and blue team is the defense side. In other words, keeping folks out. Is that a reasonable summation of the state of the world?

Nick: It can be, yeah, especially when it comes to security. One of the nice parts about the whole InfoSec field—I know a lot of folks tend to kind of just say, “Oh, they’re there to prevent the next breach,” but in reality, InfoSec has a ton of different niches and different job specialties. “Blue teamers,” quote-unquote, tend to be the defense side working on ensuring that we can alert and monitor potential attacks, whereas red teamers—or penetration testers—tend to be the folks who are trying to do the actual exploitation or develop techniques to do that in the future.

Corey: So, you talk a bit about what you do for work, obviously, but what really drew my notice was stuff you do that isn’t part of your core job, as best I understand it. You’re focused on vulnerability research, specifically with a strong emphasis on cloud exploitation, as you said—AWS in particular—and you’re the founder of Hacking the Cloud, which is an open-source encyclopedia of various attacks and techniques you can perform in cloud environments. Tell me about that.

Nick: Yeah, so Hacking the Cloud came out of a frustration I had when I was first getting into AWS, that there didn’t seem to be a ton of good resources for offensive security professionals to get engaged in the cloud. By comparison, if you wanted to learn about web application hacking, or attacking Active Directory, or reverse engineering, if you have a credit card, I can point you in the right direction. But there just didn’t seem to be a good course or introduction to how you, as a penetration tester, should attack AWS. There’s things like, you know, open S3 buckets are a nightmare, or that server-side request forgery on an EC2 instance can result in your organization being fined very, very heavily. I kind of wanted to go deeper with that.

And with Hacking the Cloud, I’ve tried to gather a bunch of offensive security research from various blog posts and conference talks into a single location, so that both the offense side and the defense side can kind of learn from it and leverage that to either improve defenses or look for things that they can attack.

Corey: It seems to me that doing things like that is not likely to wind up making a whole heck of a lot of friends over on the cloud provider side. Can you talk a little bit about how what you do is perceived by the companies you’re focusing on?

Nick: Yeah. So, in terms of relationship, I don’t really have too much of an idea of what they think. I have done some research and written on my blog, as well as published to Hacking the Cloud, some techniques for doing things like abusing the SSM agent, as well as abusing the AWS API to enumerate permissions without logging into CloudTrail. And ironically, through the power of IP addresses, I can see when folks from the Amazon corporate IP address space look at my blog, and that’s always fun, especially when there’s, like, four in the course of a couple of minutes, or five or six. But I don’t really know too much about what they—or how they view it, or if they think it’s valuable at all. I hope they do, but really not too sure.

Corey: I would imagine that they do, on some level, but I guess the big question is, you know that someone doesn’t like what you’re doing when they send, you know, cease and desist notices, or have the police knock on your door. I feel like at most levels, we’re past that in an InfoSec level, at least I’d like to believe we are. We don’t hear about that happening all too often anymore. But what’s your take on it?

Nick: Yeah, I definitely agree. I definitely think we are beyond that. Most companies these days know that vulnerabilities are going to happen, no matter how hard you try and how much money you spend, and so it’s better to be accepting of that and open to it. And especially because the InfoSec community can be so, say, noisy at times, it’s definitely worth it to pay attention, definitely be appreciative of the information that may come out. AWS is pretty awesome to work with, having disclosed to them a couple times, now.

They have a safe harbor provision, which essentially says that so long as you’re operating in good faith, you are allowed to do security testing. They do have some rules around that, but they are pretty clear in terms of if you were operating in good faith, you wouldn’t be doing anything like that. It tends to be pretty obviously malicious things that they’ll ask you to stop.

Corey: So, talk to me a little bit about what you’ve found lately, and been public about. There have been a number of examples that have come up whenever people start googling your name or looking at things you’ve done. But what’s happening lately? What have you found that’s
interesting?

Nick: Yeah. So, I think most recently, the thing that’s kind of gotten the most attention has been a really interesting bug I found in the AWS API. Essentially, kind of the core of it is that when you are interacting with the API, obviously that gets logged to CloudTrail, so long as it’s compatible. So, if you are successful, say you want to do, like, Secrets Manager, ListSecrets, that shows up in CloudTrail. And similarly, if you do not have that permission on a role or user and you try to do it, that access denied also gets logged to CloudTrail.

Something kind of interesting that I found is that by manually modifying a request, or mal-forming them, what we can do is we can modify the content-type header, and as a result when you do that—and you can provide literally gibberish. I think I have VS Code window here somewhere with a content-type of ‘meow’—when you do that, the AWS API knows the action that you’re trying to call because of that messed up content type, it doesn’t know exactly what you’re trying to do and as a result, it doesn’t get logged to CloudTrail. Now, while that may seem kind of weirdly specific and not really, like, a concern, the nice part of it though is that for some API actions—somewhere in the neighborhood of 600. I say ‘in the neighborhood of’ just because it fluctuates over time—as a result of that, you can tell if you have that permission, or if you don’t without that being logged to CloudTrail. And so we can do this enumeration of permissions without somebody in the defense side seeing us do it. Which is pretty awesome from a offensive security perspective.

Corey: On some level, it would be easy to say, “Well, just not showing up in the logs isn’t really a security problem at all.” I guess that you
disagree?

Nick: I do, yeah. So, let’s sort of look at it from a real-world perspective. Let’s say, Corey, you’re tired of saving people money on their AWS bill, you’d instead maybe want to make a little money on the side and you’re okay with perhaps, you know, committing some crimes to do it. Through some means you get access to a company’s AWS credentials for some particular role, whether that’s through remote code execution on an EC2 instance, or maybe find them in an open location like an S3 bucket or a Git repository, or maybe you phish a developer, through some means, you have an access key and a secret access key. The new problem that you have is that you don’t know what those credentials are associated with, or what permissions they have.

They could be the root account keys, or they could be literally locked down to a single S3 bucket to read from. It all just kind of depends. Now, historically, your options for figuring that out are kind of limited. Your best bet would be to brute-force the AWS API using a tool like Pacu, or my personal favorite, which is enumerate-iam by Andres Riancho. And what that does is it just tries a bunch of API calls and sees which one works and which one doesn’t.

And if it works, you clearly know that you have that permission. Now, the problem with that, though, is that if you were to do that, that’s going to light up CloudTrail like a Christmas tree. It’s going to start showing all these access denieds for these various API calls that you’ve tried. And obviously, any defender who’s paying attention is going to look at that and go, “Okay. That’s, uh, that’s suspicious,” and you’re going to get shut down pretty quickly.

What’s nice about this bug that I found is that instead of having to litter CloudTrail with all these logs, we can just do this enumeration for roughly 600-ish API actions across roughly 40 AWS services, and nobody is the wiser. You can enumerate those permissions, and if they work fantastic, and you can then use them, and if you come to find you don’t have any of those 600 permissions, okay, then you can decide on where to go from there, or maybe try to risk things showing up in CloudTrail.

Corey: CloudTrail is one of those services that I find incredibly useful, or at least I do in theory. In practice, it seems that things don’t show up there, and you don’t realize that those types of activities are not being recorded until one day there’s an announcement of, “Hey, that type of activity is now recorded.” As of the time of this recording, the most recent example that in memory is data plane requests to DynamoDB. It’s, “Wait a minute. You mean that wasn’t being recorded previously? Huh. I guess it makes sense, but oh, dear.”

And that causes a reevaluation of what’s happening in the—from a security policy and posture perspective for some clients. There’s also, of course, the challenge of CloudTrail logs take a significant amount of time to show up. It used to be over 20 minutes, I believe now it’s closer to 15—but don’t quote me on that, obviously. Run your own tests—which seems awfully slow for anything that’s going to be looking at those in an automated fashion and taking a reactive or remediation approach to things that show up there. Am I missing something key?

Nick: No, I think that is pretty spot on. And believe me, [laugh] I am fully aware at how long CloudTrail takes to populate, especially with doing a bunch of research on what is and what is not logged to CloudTrail. I know that there are some operations that can be logged more quickly than the 15-minute average. Off the top of my head, though, I actually don’t quite remember what those are. But you’re right, in general, the majority at least do take quite a while.

And that’s definitely time in which an adversary or someone like me, could maybe take advantage of that 15-minute window to try and brute force those permissions, see what we have access to, and then try to operate and get out with whatever goodies we’ve managed to steal.

Corey: Let’s say that you’re doing the thing that you do, however that comes to be—and I am curious—actually, we’ll start there. I am curious; how do you discover these things? Is it looking at what is presented and then figuring out, “Huh, how can I wind up subverting the system it’s based on?” And, similar to the way that I take a look at any random AWS services and try and figure out how to use it as a database? How do you find these things?

Nick: Yeah, so to be honest, it all kind of depends. Sometimes it’s completely by accident. So, for example, the API bug I described about not logging to CloudTrail, I actually found that due to [laugh] copy and pasting code from AWS’s website, and I didn’t change the content-type header. And as a result, I happened to notice this weird behavior, and kind of took advantage of it. Other times, it’s thinking a little bit about how something is implemented and the security ramifications of it.

So, for example, the SSM agent—which is a phenomenal tool in order to do remote access on your EC2 instances—I was sitting there one day and just kind of thought, “Hey, how does that authenticate exactly? And what can I do with it?” Sure enough, it authenticates the exact same way that the AWS API does, that being the metadata service on the EC2 instance. And so what I figured out pretty quickly is if you can get access to an EC2 instance, even as a low-privilege user or you can do server-side request forgery to get the keys, or if you just have sufficient permissions within the account, you can potentially intercept SSM messages from, like, a session and provide your own results. And so in effect, if you’ve compromised an EC2 instance, and the only way, say, incident response has into that box is SSM, you can effectively lock them out of it and, kind of, do whatever you want in the meantime.

Corey: That seems like it’s something of a problem.

Nick: It definitely can be. But it is a lot of fun to play keep-away with incident response. [laugh].

Corey: I’d like to reiterate that this is all in environments you control and have permissions to be operating within. It is not recommended that people pursue things like this in other people’s cloud environments without permissions. I don’t want to find us sued for giving crap advice, and I don’t want to find listeners getting arrested because they didn’t understand the nuances of what we’re talking about.

Nick: Yes, absolutely. Getting legal approval is really important for any kind of penetration testing or red teaming. I know some folks sometimes might get carried away, but definitely be sure to get approval before you do any kind of testing.

Corey: So, how does someone report a vulnerability to a company like AWS?

Nick: So AWS, at least publicly, doesn’t have any kind of bug bounty program. But what they do have is a vulnerability disclosure program. And that is essentially an email address that you can contact and send information to, and that’ll act as your point of contact with AWS while they investigate the issue. And at the end of their investigation, they can report back with their findings, whether they agree with you and they are working to get that patched or fixed immediately, or if they disagree with you and think that everything is hunky-dory, or if you may be mistaken.

Corey: I saw a tweet the other day that I would love to get your thoughts on, which said effectively, that if you don’t have a public bug bounty program, then any way that a researcher chooses to disclose the vulnerability is definitionally responsible on their part because they don’t owe you any particular duty of care. Responsible disclosure, of course, is also referred to as, “Coordinated vulnerability disclosure” because we’re always trying to reinvent terminology in this space. What do you think about that? Is there a duty of care from security researchers to responsibly disclose the vulnerabilities they find, or coordinate those vulnerabilities with vendors in the absence of a public bounty program on turning those things in?

Nick: Yeah, you know, I think that’s a really difficult question to answer. From my own personal perspective, I always think it’s best to contact the developers, or the company, or whoever maintains whatever you found a vulnerability in, give them the best shot to have it fixed or repaired. Obviously, sometimes that works great, and the company is super receptive, and they’re willing to patch it immediately. And other times, they just don’t respond, or sometimes they respond harshly, and so depending on the situation, it may be better for you to release it publicly with the intention that you’re informing folks that this particular company or this particular project may have an issue. On the flip side, I can kind of understand—although I don’t necessarily condone it—why folks pursue things like exploit brokers, for example.

So, if a company doesn’t have a bug bounty program, and the researcher isn’t expecting any kind of, like, cash compensation, I can understand why they may spend tens of hours, maybe hundreds of hours chasing down a particularly impactful vulnerability, only to maybe write a blog post about it or get a little head pat and say, “Thanks, nice work.” And so I can see why they may pursue things like selling to an exploit broker who may pay them hefty sum, if it is a—

Corey: Orders of magnitude more. It’s, “Oh, good. You found a way to remotely execute code across all of EC2 in every region”—that is a hypothetical; don’t email me—have a t-shirt. It seems like you could basically buy all the t-shirts for [laugh] what that is worth on the export market.

Nick: Yes, absolutely. And I do know from some experience that folks will reach out to you and are interested in, particularly, some cloud exploits. Nothing, like, minor, like some of the things that I’ve found, but more thinking more of, like, accessing resources without anybody knowing or accessing resources cross-account; that could go for quite a hefty sum.

Corey: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the Enterprise (not the starship). On-prem security doesn’t translate well to cloud or multi-cloud environments, and that’s not even counting IoT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IoT devices, detects these threats up to 35 percent faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at extrahop.com/trial.

Corey: It always feels squicky, on some level, to discover something like this that’s kind of neat, and wind up selling it to basically some arguably terrible people. Maybe. We don’t know who’s buying these things from the exploit broker. Counterpoint, having reported a few security problems myself to various providers, you get an autoresponder, then you get a thank you email that goes into a bit more detail—for the well-run programs, at least—and invariably, the company’s position is, is whatever you found is not as big of a deal as you think it is, and therefore they see no reason to publish it or go loud with it. Wouldn’t you agree?

Because, on some level, their entire position is, please don’t talk about any security shortcomings that you may have discovered in our system. And I get why they don’t want that going loud, but by the same token, security researchers need a reputation to continue operating on some level in the market as security researchers, especially independents, especially people who are trying to make names for themselves in the first place.

Nick: Yeah.

Corey: How do you resolve that dichotomy yourself?

Nick: Yeah, so, from my perspective, I totally understand why a company or project wouldn’t want you to publicly disclose an issue. Everybody wants to look good, and nobody wants to be called out for any kind of issue that may have been unintentionally introduced. I think the thing at the end of the day, though, from my perspective, if I, as some random guy in the middle of nowhere Illinois finds a bug, or to be frank, if anybody out there finds a vulnerability in something, then a much more sophisticated adversary is equally capable of finding such a thing. And so it’s better to have these things out in the open and discussed, rather than hidden away, so that we have the best chance of anybody being able to defend against it or develop detections for it, rather than just kind of being like, “Okay, the vendor didn’t like what I had to say, I guess I’ll go back to doing whatever [laugh] things I normally do.”

Corey: You’ve obviously been doing this for a while. And I’m going to guess that your entire security researcher career has not been focused
on cloud environments in general and AWS in particular.

Nick: Yes, I’ve done some other stuff in relation to abusing GitLab Runners. I also happen to find a pretty neat RCE and privilege escalation in the very popular open-source project. Pi-hole. Not sure if you have any experience with that.

Corey: Oh, I run it myself all the time for various DNS blocking purposes and other sundry bits of nonsense. Oh, yes, good. But what I’m trying to establish here is that this is not just one or two companies that you’ve worked with. You’ve done this across the board, which means I can ask a question without naming and shaming anyone, even implicitly. What differentiates good vulnerability disclosure programs from terrible ones?

Nick: Yeah, I think the major differentiator is the reactivity of the project, as in how quickly they respond to you. There are some programs I’ve worked with where you disclose something, maybe even that might be of a high severity, and you might not hear back four weeks at a time, whereas there are other programs, particularly the MSRC—which is a part of Microsoft—or with AWS’s disclosure program, where within the hour, I had a receipt of, “Hey, we received this, we’re looking into it.” And then within a couple hours after that, “Yep, we verified it. We see what you’re seeing, and we’re going to look at it right away.” I think that’s definitely one of the major differentiators for programs.

Corey: Are there any companies you’d like to call out in either direction—and, “No,” is a perfectly valid [laugh] answer to this one—for having excellent disclosure programs versus terrible ones?

Nick: I don’t know if I’d like to call anybody out negatively. But in support, I have definitely appreciated working with both AWS’s and the MSRC—Microsoft’s—I think both of them have done a pretty fantastic job. And they definitely know what they’re doing at this point.

Corey: Yeah, I must say that I primarily focus on AWS and have for a while, which should be blindingly obvious to anyone who’s listened to me talk about computers for more than three and a half minutes. But my experiences with the security folks at AWS have been uniformly positive, even when I find things that they don’t want me talking about, that I will be talking about regardless, they’ve always been extremely respectful, and I have never walked away from the conversation thinking that I was somehow cheated by the experience. In fact, a couple of years ago at the last in-person re:Invent, I got to give a talk around something I reported specifically about how AWS runs its vulnerability disclosure program with one of their security engineers, Zach Glick, and he was phenomenally transparent around how a lot of these things work, and what they care about, and how they view these things, and what their incentives are. And obviously being empathetic to people reporting things in with the understanding that there is no duty of care that when security researchers discover something, they then must immediately go and report it in return for a pat on the head and a thank you. It was really neat being able to see both sides simultaneously around a particular issue. I’d recommend it to other folks, except I don’t know how you make that lightning strike twice.

Nick: It’s very, very wise. Yes.

Corey: Thank you. I do my best. So, what’s next for you? You’ve obviously found a number of interesting vulnerabilities around information disclosure. One of the more recent things that I found that was sort of neat as I trolled the internet—I don’t believe it was yours, but there was a ability to determine the account ID that owned an S3 bucket by enumerating by a binary search. Did you catch that at all?

Nick: I did. That was by Ben Bridts, which is—it’s pretty awesome technique, and that’s been something I’ve been kind of interested in for a while. There is an ability to enumerate users’ roles and service-linked roles inside an account, so long as the account ID. The problem, of course, is getting the account ID. So, when Ben put that out there I was super stoked about being able to leverage that now for enumeration and maybe some fun phishing tricks with that.

Corey: I love the idea. I love seeing that sort of thing being conducted. And AWS’s official policy as best I remember when I looked at this once, account IDs are not considered confidential. Do you agree with that?

Nick: Yep. That is my understanding of how AWS views it. From my perspective, having an account ID can be beneficial. I mentioned that you can enumerate users’ roles and service-linked roles with it, and that can be super useful from a phishing perspective. The average phishing email looks like, “Oh, you won an iPad,” or, “Oh, you’re the 100th visitor of some website,” or something like that.

But imagine getting an email that looks like it’s from something like AWS developer support, or from some research program that they’re doing, and they can say to you, like, “Hey, we see that you have these roles in your account with account ID such-and-such, and we know that you’re using EKS, and you’re using ECS,” that phishing email becomes a lot more believable when suddenly this outside party seemingly knows so much about your account. And that might be something that you would think, “Oh, well only a real AWS employee or AWS would know that.” So, from my perspective, I think it’s best to try and keep your account ID secret. I actually redact it from every screenshot that I publish, or at the very least, I try to. At the same time, though, it’s not the kind of thing that’s going to get somebody in your account in a single step, so I can totally see why some folks aren’t too concerned about it.

Corey: I feel like we also got a bit of a red herring coming from AWS blog posts themselves, where they always will give screenshots explaining what they do, and redact the account ID in every case. And the reason that I was told at one point was, “Oh, we have an internal provisioning system that’s different. It looks different, and I don’t want to confuse people whenever I wind up doing a screenshot.” And that’s great, and I appreciate that. And part of me wonders on one level how accurate is that?

Because sure, I understand that you don’t necessarily want to distract people with something that looks different, but then I found out that the system is called Isengard and, yeah, it’s great. They’ve mentioned it periodically in blog posts, and talks, and the rest. And part of me now wonders, oh, wait a minute. Is it actually because they don’t want to disclose the differences between those systems, or is it because they don’t have license rights publicly to use the word Isengard and don’t want to get sued by whoever owns the rights to the Lord of the Rings trilogy. So, one wonders what the real incentives are in different cases. But I’ve always viewed account IDs as being the sort of thing that eh, you probably want to share them around all the time, but it also doesn’t necessarily hurt.

Nick: Exactly, yeah. It’s not the kind of thing you want to share with the world immediately, but it doesn’t really hurt in the end.

Corey: There was an early time when the partner network was effectively determining tiers of partner by how much spend they influenced, and the way that you’ve demonstrated that was by giving account IDs for your client accounts. The only verification at the time, to my understanding was that, “Yep, that mapped to the client you said it did.” And that was it. So, I can understand back in those days not wanting to muddy those waters. But those days are also long passed.

So, I get it. I’m not going to be the first person to advertise mine, but if you can discover my account ID by looking at a bucket, it doesn’t really keep me up at night.

So, all of those things considered, we’ve had a pretty wide-ranging conversation here about a variety of things. What’s next? What interests you as far as where you’re going to start looking and exploring—and exploiting as the case may be—various cloud services? hackthe.cloud—which there is the dot in there, which also turns it into a domain; excellent choice—is absolutely going to be a great collection for a lot of what you find and for other people to contribute and learn from one another. But where are you aimed at? What’s next?

Nick: Yeah, so one thing I’ve been really interested in has been fuzzing the AWS API. As anyone who’s ever used AWS before knows, there are hundreds of services with thousands of potential API endpoints. And so from a fuzzing perspective, there is a wide variety of things for us to potentially affect or potentially find vulnerabilities in. I’m currently working on a library that will allow me to make that fuzzing a lot easier. You could use things like botocore, Boto3, like, some of the AWS SDKs.

The problem though, is that those are designed for, sort of like, the happy path where you can format your request the way Amazon wants. As a security researcher or as someone doing fuzzing, I kind of want to send random gibberish sometimes, or I want to malform my requests. And so that library is still in production, but it has already resulted in a bug. While I was fuzzing part of the AWS API, I happened to notice that I broke Elastic Beanstalk—quite literally—when [laugh] when I was going through the AWS console, I got the big red error message of, “[unintelligible 00:29:35] that request parameter is null.” And I was like, “Huh. Well, why is it null?”

And come to find out as a result of that, there is a HTML injection vulnerability in the Elastic—well, there was a HTML injection vulnerability in the Elastic Beanstalk, for the AWS console. Pivoting from there, the Elastic Beanstalk uses Angular 1.8.1, or at least it did when I found it. As a result of that, we can modify that HTML injection to do template injection. And for the AngularJS crowd, template injection is basically cross-site scripting [laugh] because there is no sandbox anymore, at least in that version. And so as a result of that, I was able to get cross-site scripting in the AWS console, which is pretty exciting. That doesn’t tend to happen too frequently.

Corey: No that is not a typical issue that winds up getting disclosed very often.

Nick: Definitely, yeah. And so I was excited about it, and considering the fact that my library for fuzzing is literally, like, not even halfway done, or is barely halfway done, I’m looking forward to what other things I can find with it.

Corey: I look forward to reading more. And at the time of this recording, I should point out that this has not been finalized or made public, so I’ll be keeping my eyes open to see what happens with this. And hopefully, this will be old news by the time this episode drops. If not, well, [laugh] this might be an interesting episode once it goes out.

Nick: Yeah. I hope they’d have it fixed by then. They haven’t responded to it yet other than the, “Hi, we’ve received your email. Thanks for checking in.” But we’ll see how that goes.

Corey: Watching news as it breaks is always exciting. If people want to learn more about what you’re up to, and how you go about things, where can they find you?

Nick: Yeah, so you can find me at a couple different places. On Twitter I’m @frichette_n. I also write a blog where I contribute a lot of my research at frechetten.com as well as Hacking the Cloud. I contribute a lot of the AWS stuff that gets thrown on there. And it’s also open-source, so if anyone else would like to contribute or share their knowledge, you’re absolutely welcome to do so. Pull requests are open and excited for anyone to contribute.

Corey: Excellent. And we will of course include links to that in the [show notes 00:31:42]. Thank you so much for taking the time to speak with me. I really appreciate it.

Nick: Yeah, thank you so much for inviting me on. I had a great time.

Corey: Nick Frechette, penetration tester and team lead for State Farm. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with a comment telling me why none of these things are actually vulnerabilities, but simultaneously should not be discussed in public, ever.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Christina

Christina Maslach, PhD, is a Professor of Psychology (Emerita) and a researcher at the Healthy Workplaces Center at the University of California, Berkeley. She received her A.B. from Harvard, and her Ph.D. from Stanford. She is best known as the pioneering researcher on job burnout, producing the standard assessment tool (the Maslach Burnout Inventory, MBI), books, and award-winning articles. The impact of her work is reflected by the official recognition of burnout, as an occupational phenomenon with health consequences, by the World Health Organization in 2019. In 2020, she received the award for Scientific Reviewing, for her writing on burnout, from the National Academy of Sciences. Among her other honors are: Fellow of the American Association for the Advancement of Science (1991 -- "For groundbreaking work on the application of social psychology to contemporary problems"), Professor of the Year (1997), and the 2017 Application of Personality and Social Psychology Award (for her research career on job burnout).

Links:

  • The Truth About Burnout: https://www.amazon.com/Truth-About-Burnout-Organizations-Personal/dp/1118692136

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Thinkst. This is going to take a minute to explain, so bear with me. I linked against an early version of their tool, canarytokens.org in the very early days of my newsletter, and what it does is relatively simple and straightforward. It winds up embedding credentials, files, that sort of thing in various parts of your environment, wherever you want to; it gives you fake AWS API credentials, for example. And the only thing that these things do is alert you whenever someone attempts to use those things. It’s an awesome approach. I’ve used something similar for years. Check them out. But wait, there’s more. They also have an enterprise option that you should be very much aware of canary.tools. You can take a look at this, but what it does is it provides an enterprise approach to drive these things throughout your entire environment. You can get a physical device that hangs out on your network and impersonates whatever you want to. When it gets Nmap scanned, or someone attempts to log into it, or access files on it, you get instant alerts. It’s awesome. If you don’t do something like this, you’re likely to find out that you’ve gotten breached, the hard way. Take a look at this. It’s one of those few things that I look at and say, “Wow, that is an amazing idea. I love it.” That’s canarytokens.org and canary.tools. The first one is free. The second one is enterprise-y. Take a look. I’m a big fan of this. More from them in the coming weeks.

Corey: This episode is sponsored in part by our friends at Lumigo. If you’ve built anything from serverless, you know that if there’s one thing that can be said universally about these applications, it’s that it turns every outage into a murder mystery. Lumigo helps make sense of all of the various functions that wind up tying together to build applications. It offers one-click distributed tracing so you can effortlessly find and fix issues in your serverless and microservices environment. You’ve created more problems for yourself; make one of them go away. To learn more, visit lumigo.io.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. One subject that I haven’t covered in much depth on this show has been a repeated request from the audience, and that is to talk a bit about burnout. So, when I asked the audience who I should talk to about burnout, there were really two categories of responses. The first was, “Pick me. I hate my job, and I’d love to talk about that.” And the other was, “You should speak to Professor Maslach.” Christina Maslach is a Professor of Psychology at Berkeley. She’s a teacher and a researcher, particularly in the area of burnout. Professor, welcome to the show.

Dr. Maslach: Well, thank you for inviting me.

Corey: So, I’m going to assume from the outset that the reason that people suggest that I speak to you about burnout is because you’ve devoted a significant portion of your career to studying the phenomenon, and not just because you hate your job and are ready to go do something else. Is that directionally correct?

Dr. Maslach: That is directionally correct, yes. I first stumbled upon the phenomenon back in the 1970s—which is, you know, 45, almost
50 years ago now—and have been fascinated with trying to understand what is going on.

Corey: So, let’s start at the very beginning because I’m not sure in, I guess, the layperson context that I use the term that I fully understand it. What is burnout?

Dr. Maslach: Well, burnout as we have been studying it over many years, it’s a stress phenomenon, okay, it’s a response to stressors, but it’s not just the exhaustion of stress. That’s one component of it, but it actually has two other components that go along with it. One is this very negative, cynical, hostile attitude toward the job and the other people in it, you know, “Take this job and shove it,” kind of feeling. And usually, people don’t begin their job like that, but that’s where they go if they become more burned out.

Corey: I believe you may have just inadvertently called out a decent proportion of the tech sector.

Dr. Maslach: [laugh].

Corey: Or at least, that might just be my internal cynicism rising to the foreground.

Dr. Maslach: No, it’s not. Actually, I have heard from a number of tech people over the past decades about just this kind of issue. And so I think it’s particularly relevant. The third component that we see going along with this, it usually comes in a little bit later, but I’ve heard a lot about this from tech people as well, and that is that you begin to develop a very negative sense of your own self, and competence, and where you’re going, and what you’re able to do. So, the stress response of exhaustion, the negative cynicism towards the job, the negative evaluation of yourself, that’s the trifecta of burnout.

Corey: You’ve spent a lot of your early research at least focusing on, I guess, occupations that you could almost refer to as industrial, in some respects: working with heavy equipment, working with a variety of different professionals in very stressful situations. It feels weird, on some level, to say, “Oh, yeah, my job is very stressful. In that vein, I have to sit in front of a computer all day, and sometimes I have to hop on a meeting with people.” And it feels, on some level, like that even saying, “I’m experiencing burnout,” in my role is a bit of an overreach.

Dr. Maslach: Yeah, that’s an interesting point because, in fact, yes, when we think about OSHA, you know, and occupational risks and hazards, we do think about the chemicals, and the big equipment, and the hazards, so having more psychological and social risk factors, is something that probably a lot of people don’t resonate to immediately and think, well, if you’re strong, and if you’re resilient, and whatever, you can—anybody can handle that, and that’s really a test almost of your ability to do your work. But what we’re finding is that it has its own hazards, psychological and social as well. And so, burnout is something that we’ve seen in a lot of more people-oriented professions, from the beginning. Healthcare has had this for a long time. Various kinds of social services, teaching, all of these other things. So, it’s actually not a sign of weakness as some people might think.

Corey: Right. And that’s part of the challenge and, honestly, one of the reasons that I’ve stayed away from having in-depth discussions about the topic of burnout on the show previously is it feels that—rightly or wrongly, and I appreciate your feedback on this one either way—it feels like it’s approaching the limits of what could be classified as mental health. And I can give terrible advice on how computers work—in fact, I do on a regular basis; it’s kind of my thing—and that’s usually not going to have any lasting impact on people who don’t see through the humor part of that. But when we start talking about mental health, I’m cautious because it feels like an inadvertent story or advice that works for some but not all, has the potential to do a tremendous bit of damage, and I’m very cautious about that. Is burnout a mental health issue? Is it a medical issue that is recognized? Where does it start, okay does it stop on that spectrum?

Dr. Maslach: It is not a medical issue—and the World Health Organization, which just came out with a statement about this in 2019 on burnout, they’re recognizing it as an occupational risk factor—made it very clear that this is not a medical thing. It is not a medical disease, it doesn’t have a certain set of medical diagnoses, although people tend to sometimes go there. Can it have physical health outcomes? In other words, if you’re burning out and you’re not sleeping well, and you’re not eating well, and not taking care of yourself, do you begin to impair your physical health down the road? Yes.

Could it also have mental health outcomes, that you begin to feel depressed, and anxious, and not knowing what to do, and afraid of the future? Yes, it could have those outcomes as well. So, it certainly is kind of like—I can put it this way, like a stepping stone in a path to potential negative health: physical health, or mental health issues. And I think that’s one of the reasons why it is so important. But unfortunately, a lot of people still view it as somebody who’s burned out isn’t tough enough, strong enough, they’re wimpy, they’re not good enough, they’re not a hundred percent.

And so the stigma that is often attached to burnout, people not only indulge it, but they feel it directed towards them, and often they will try to hide the kinds of experiences they’re having because they worry that they are going to be judged negatively, thrown under the bus, you know, let go from the job, whatever, if they talk about what’s actually happening with them.

Corey: What do you see, as you look around, I guess, the wide varieties of careers that are susceptible to burnout—which I have a sneaking suspicion based upon what you’ve said rounds to all of them—what do you think is the most misunderstood, or misunderstood aspects of burnout?

Dr. Maslach: I think what’s most misunderstood is that people assume that it is a problem of the individual person. And if somebody is burned out, then they’ve got to just take care of themselves, or take a break, or eat better, or get more sleep, all of those kinds of things which cope with stressors. What’s not as well understood or focused on is the fact that this is a response to other stressors, and these stressors are often in the workplace—this is where I’ve been studying it—but in essentially in the larger social, physical environment that people are functioning in. They’re not burning out all by themselves.

There’s a reason why they are feeling the kind of exhaustion, developing that cynicism, beginning to doubt themselves, that we see with burnout. So there, if you ever want to talk about preventing burnout, you really have to be focusing on what are the various kinds of things that seem to be causing the problem, and how do we modify those? Coping with stressors is a good thing, but it doesn’t change the stressors. And so we really have to look at that, as well as what people can bring about, you know, taking care of themselves or trying to do the job better or differently.

Corey: I feel like it’s impossible to have a conversation like this without acknowledging the background of the past year that many of us have spent basically isolated, working from home. And for some folks, okay, they were working from home before, but it feels different now. At least that’s the position I find myself in. Other folks are used to going into an office and now they’re either isolated—and research shows that it has been worse, statistically, for single people versus married people, but married people are also trapped at home with their spouse, which sounds half-joking but it is very real. At some point, distance is useful.

And it feels like everyone is sort of a bit at their wit’s end. It feels like things are closer to being frayed, there’s a constant sense that there’s this, I guess, pervasive dread for the past year. Are you seeing that that has a potential to affect how burnout is being expressed or perceived?

Dr. Maslach: I think it has, and one of the things that we clearly see is that people are using the word burnout, more and more and more and more. It’s almost becoming the word du jour, and using it to describe, things are going wrong and it’s not good. And it may be overstretching the use of burnout, but I think the reason of the popularity of the term is that it has this kind of very vivid imagery of things going up in smoke, and can’t handle it, and flames licking at your heels, and all this sort of stuff so that they can do that. I even got a comment from a colleague in France just a few days ago, where they’re talking about, “Is burnout the malady of the century?” you know, kind of thing. And it’s being used a lot; it’s sometimes maybe overused, but I think it’s also striking a chord with people as a sign that things are going badly, and I don’t know how to deal with it in some way.

Corey: It also feels, on some level, for those of us who are trapped inside, it kind of almost feels like it’s a tremendous expression of privilege because who am I to have a problem with this? Oh, I have to go inside and order a lot of takeout and spend time with my family. And I look at how folks who are nowhere near as privileged have to go and be essential workers and show up in increasingly dangerous positions. And it almost feels like burnout isn’t something that I’m entitled to, if that makes sense.

Dr. Maslach: [laugh]. Yeah. It’s an interesting description of that because I think there are ways in which people are looking at their experience and dealing with it, and like many things in life, I find that all of these things are a bit of a double-edged sword; there’s positive and there’s negative aspects to them. And so when I’ve talked with some people about now having to work from home rather than working in their office, they’re also bringing up, “Well, hey, I’ve noticed that the interviews I’m doing with potential clients are actually going a little better”—you know, this is from a law office—“And trying to figure out how—are we doing it differently so that people can actually relate to each other as human beings instead of the suit and tie in the big office? What’s going on in terms of how we’re doing the work that there may be actually a benefit here?”

For others. It’s been, “Oh, my gosh. I don’t have to commute, but endless meetings and people are thinking I’m not doing my job, and I don’t know how to get in touch, and how do we work together effectively?” And so there’s other things that are much more difficult, in some sense. I think another thing that you have to keep in mind that it’s not just about how you’re doing your work, perhaps differently, or you’re under different circumstances, but people, so many people have lost their jobs, and are worried that they may lose their jobs.

That we’re actually finding that people are going into overdrive and working harder and more hours as a way of trying to protect from being the next one who won’t have any income at all. So, there’s a lot of other dynamics that are going on as a result of the pandemic, I think, that we need to be aware of.

Corey: One thing that I’d like to point out is that you are a Professor Emerita of Psychology at Berkeley, which means you presumably wound up formulating this based upon significant bodies of peer-reviewed research, as opposed to just coming up with a thesis, stating it
as if it were fact, and then writing an entire series of books on it. I mean, that path, I believe, is called being a venture capitalist, but I may be mistaken on that front. How do you effectively study something like burnout? It feels like it is so subjective and situation-specific, but
it has to have a normalization aspect to it.

Dr. Maslach: Uh, yeah, that’s a good point. I think, in fact, the first time I ever wrote about some of the stuff that I was learning about burnout back in the mid ’70s—I think it was ’75, ’76 maybe—and it was in a magazine, it wasn’t in a journal. It wasn’t peer-reviewed because not even peer-reviewed journals would review this; they thought it was pop psychology, and eh. So, I would get, in those days, snail mail by the sackfuls from people saying, “Oh, my God. I didn’t know anybody else felt like this. Let me tell you my story.”

You know, kind of thing. And so that was really, after doing a lot of interviews with people, following them on the job when possible to, sort of, see how things were going, and then writing about the basic themes that were coming out of this, it turned out that there were a lot of people who responded and said, “I know that. I’ve been there. I’m experiencing it.” Even though each of them were sort of thinking, “I’m the only one. What’s wrong with me? Everybody else seems fine.”

And so part of the research in trying to get it out in whatever form you can is trying to share it because that gives you feedback from a wide variety of people, not only the peers reviewing the quality of the research, but the people who are actually trying to figure out how to deal effectively with this problem. So it’s, how do I and my colleagues actually have a bigger, broader conversation with people from which we learn a lot, and then try and say, okay, and here’s everything we’ve heard, and let’s throw it back out and share it and see what people think.

Corey: You have written several books on the topic, if I’m not mistaken. And one thing that surprises me is how much what you talk about in those books seems to almost transcend time. I believe your first was published in 1982—

Dr. Maslach: Right.

Corey: —if I’m not mistaken—

Dr. Maslach: Yes.

Corey: —and it’s an awful lot of what it talks about still feels very much like it could be written today. Is this just part of the quintessential human experience? Or has nothing new changed in the last 200 years since the Industrial Revolution? How is it progressing, if at all, and
what does the future look like?

Dr. Maslach: Great questions and I don’t have a good answer for you. But we have sort of struggled with this because if you look at older literature, if you even go back centuries, if you even go back in parts of the Bible or something, you’re seeing phrases and descriptions sometime that says sounds a lot like burnout, although we’re not using that term. So, it’s not something that I think just somehow got invented; it wasn’t invented in the ’70s or anything like that. But trying to trace back those roots and get a better sense of what are we capturing here is fascinating, and I think we’re still working on it.

People have asked, well, where did the term ‘burnout’ as opposed to other kinds of terms come from? And it’s been around for a while, again, before the ’70s or something. I mean, we have Graham Greene writing the novel A Burnt-Out Case, back in the early ’60s. My dad was an engineer, rarefied gas dynamics, so he was involved with the space program and engineers talk about burnout all the time: ball bearings burn out, rocket boosters burn out. And when they started developing Silicon Valley, all those little startups and enterprises, they advertised as burnout shops. And that was, you know, ’60s, into the ’70s, et cetera, et cetera. So, the more modern roots, I think probably have some ties to that use of the term before I and other researchers even got started with it.

Corey: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the Enterprise (not the starship). On-prem security doesn’t translate well to cloud or multi-cloud environments, and that’s not even counting IoT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IoT devices, detects these threats up to 35 percent faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at extrahop.com/trial.

Corey: This is one of those questions that is incredibly self-serving, and I refuse to apologize for it. How can I tell whether I’m suffering from burnout, versus I’m just a jerk with an absolutely terrible attitude? And that is not as facetious a question as it probably sounds like.

Dr. Maslach: [laugh]. Yeah. Well, part of the problem for me—or the challenge for me—is to understand what it is people need to know about themselves. Can I take a diagnostic test which tells me if I am burned out or if I’m something else?

Sort of the more important question is, what is feeling right and what is not feeling so good—or even wrong—about my experience? And usually, you can’t figure that all out by yourself and you need to get other input from other people. And it could be a counselor or therapist, or it could be friends or colleagues who you have to be able to get to a point where we can talk about it, and hear each other, and get some feedback without putdowns, just sort of say, “Yeah, have you ever thought about the fact that when you get this kind of a task, you usually just go crazy for a while and not really settle down and figure out what you really need to do as opposed to what you think you have to do?” Part of this, are you bringing yourself in terms of the stress response, but what is it that you’re not doing—or that you’re doing not well—to figure out solutions, to get help or advice or better input from others. So, it takes time, but it really does take a lot of that kind of social feedback.

So, when I said—if I can stay with it a little bit more—when I first was writing and publishing about and all these people were writing back saying, “I thought I was the only one,” that phenomenon of putting on a happy face and not letting anybody else see that you’re going through some difficult challenges, or feeling bad, or depressed, or whatever is something we call pluralistic ignorance; means we don’t have good knowledge about what is normal, or what is being shared, or how other people are because we’re all pretending to put on the happy face, to pretend and make sure that everybody thinks we’re okay and is not going to come after us. But if we all do that, then we all, together, are creating a different social reality that people perceive rather than actually what is happening behind that mask.

Corey: It feels, on some level, like this is an aspect of the social media problem, where we’re comparing our actual lives and all the bloopers that we see to other people’s highlight reels because few people wind up talking very publicly about their failures.

Dr. Maslach: Oh, yeah. Yeah. And often for good reason because they know they will be attacked and dumped. And there could be some serious consequences, and you just say, “I’m going to figure out what I’m going to do on my own.”

But one of the things that when I work with people, and I’m asking them, “What do you think would help? What sort of things that don’t happen could happen?” And so forth, one of the things that goes to the top of the list is having somebody else; a safe relationship, a safe place where we can talk, where we can unburden, where you’re not going to spill the beans to everybody else, and you’re getting advice, or you’re getting a pat on the back, or a shoulder to cry on, and that you’re there for them for the same kind of reason. So, it’s a different form of what we think of as social network. It used to be that a network like that meant that you had other people, whether family, friends, neighbors, colleagues, whoever, that you knew, you could go to; a mentor, an advisor, a trusted ally, and that you would perform that role for them and other people, as well.

And what has happened, I think, to add to the emphasis on burnout these days, is that those social connections, those trusts, between people has really been shredding, and, you know—or cut off or broken apart. And so people are feeling isolated, even if they’re surrounded by a lot of other people, don’t want to raise their hand, don’t want to say, “Can we talk over coffee? I’m really having a bad day. I need some help to figure out this problem.” And so one of those most valuable resources that human beings need—which is other people—is, if we’re working in environments where that gets pulled apart, and shredded, and it’s lacking, that’s a real risk factor for burnout.

Corey: What are the things that contribute to burnout? It doesn’t feel, based upon what you’ve said so far, that it’s one particular thing.
There has to be points of commonality between all of this, I have to imagine.

Dr. Maslach: Yeah.

Corey: Is it possible to predict that, oh, this is a scenario in which either I or people who are in this role are likely to become burned out faster?

Dr. Maslach: Mm-hm. Yeah. Good question and I don’t know if we have a final answer, but at this point, in terms of all the research that’s been done, not just on burnout, but on much larger issues of health, and wellbeing, and stress, and coping, and all the rest of it, there are clearly six areas in which the fit between people and their job environment are critical. And if the fit is—or the match, or the balance—is better, they are going to be at less risk for burnout, they’re more likely to be engaged with work.

But if some real bad fits, or mismatches, occur in one or more of these areas, that raises the risk factor for burnout. So, if I can just mention those six quickly. And these are not in any particular order because I find that people assume the first one is the worst or the best, and it’s not. Any rate, one of them has to do with that social environment I was just talking about; think of it as the workplace community. All the people whose paths you cross at various points—you know, coworkers, the people you supervise, your bosses, et cetera—so those social relationships, that culture, do you have a supportive environment which really helps people thrive? Can you trust people, there’s respect, and all that kind of thing going on? Or is it really what people are now describing as a socially toxic work environment?

A second area has to do with reward. And it turns out not so much salary and benefits, it’s more about social recognition and the intrinsic reward you get from doing a good job. So, if you work hard, do some special things, and nothing positive happens—nobody even pats you on the back, nobody says, “Gee, why don’t you try this new project? I think you’re really good at it,” anything that acknowledges what you’ve done—it’s a very difficult environment to work in. People who are more at risk of burnout, when I asked them, “What is a good day for you? A good day. A really good day.” And the answer is often, “Nothing bad happens.” But it’s not the presence of good stuff happening, like people glad that you did such good work or something like that.

Third area has to do with values—and this is one that also often gets ignored, but sometimes this is the critical bottom line—that you’re doing work that you think is meaningful, where you’re working has integrity, and you’re in line with that as opposed to value conflicts and where you’re doing things that you think are wrong: “I want to help people, I want to help cure patients, and here, I’m actually only supposed to be trying to help the hospital get more money.” When they have that kind of value conflict, this is often where they have to say, “I don’t want to sell my soul and I’m leaving.”

The fourth area is one of fairness. And this is really about that whatever the policies, the principles, et cetera, they’re administered fairly. So, when things are going badly here—the mismatch—this is where discrimination lives, this is where glass ceilings are going on, that people are not being treated fairly in terms of the work they do, how they’re promoted, or all of those kinds of things. So, that interpersonal respect, and, sort of, social justice is missing.

The next two areas—the fifth and six—are probably the two that had been the most well-known for a long time. One has to do with workload and how manageable it is. Given the demands that you have, do you have sufficient resources, like time, and tools, and whatever other kind of teams support you need to get the job done. And control is about the amount of autonomy and the opportunities you have to perhaps improvise, or innovate, or correct, or figure out how to do the job better in some way. So, when people are having mismatches in work overload; a lack of control; you cannot improvise; where you have unfairness; where there is values that are just incompatible with what you believe is right, a sort of moral issue; where you’re not getting any kind of positive feedback, even when it’s deserved, for the kind of work you’re doing; and when you’re working in a socially toxic relationship where you can’t trust people, you don’t know who to turn to, people are having unresolved conflicts all the time. Those six areas are, those are the markers really of risk factors for burnout.

Corey: I know that I’m looking back through my own career history listening to you recount those and thinking, “Oh, maybe I wasn’t just a terrible employee in every one of those situations.”

Dr. Maslach: Exactly.

Corey: I’m sure a lot of it did come from me, I want to be very clear here. But there’s also that aspect of this that might not just be a ‘me’
problem.

Dr. Maslach: Yeah. That’s a good way of putting it. It’s really in some sense, it’s more of a ‘we’ problem than a ‘me’ problem. Because again, you’re not working in isolation, and the reciprocal relationship you have with other people, and other policies, and other things that are happening in whatever workplace that is, is creating a kind of larger environment in which you and many others are functioning.

And we’ve seen instances where people begin to make changes in that environment—how do we do this differently? How can we do this better, let’s try it out for a while and see if this can work—and using those six areas, the value is not just, “Oh, it’s really in bad shape. We have huge unfairness issues.” But then it says, “It would be better if we could figure out a way to get rid of that fairness problem, or to make a modification so that we have a more fair process on that.” So, they’re like guideposts as well.

As people start thinking through these six areas, you can sort of say, “What’s working well, in terms of workload, what’s working badly? Where do we run into problems on control? How do we improve the social relationships between colleagues who have to work together on a team?” They’re not just markers of what’s gone wrong, but they can—if you flip it around and look at it, let’s look at the other end—okay is a path that we could get better? Make it right?

Corey: If people want to learn more about burnout in general, and you’re working in it specifically, where can they go to find your work and learn more about what you have to say?

Dr. Maslach: Obviously, there’s been a lot of articles, and now lots of things on the web, and in past books that I’ve written. And as you said, in many ways, they are still pretty relevant. The Truth About Burnout came out, oh gosh, ’97. So, that’s 25 years ago and it’s still work.

But my colleague, Michael Leiter from Canada, and I have just written up a new manuscript for a new book in which we really are trying to focus on sharing everything we have learned about, you know, what burnout has taught us, and put that into a format of a book that will allow people to really take what we’ve learned and figure out how does this apply? How can this be customized to our situation? So, I’m hoping that that will be coming out within the next year.

Corey: And you are, of course, welcome back to discuss your book when it releases.

Dr. Maslach: I would be honored if you would have me back. That would be a wonderful treat.

Corey: Absolutely. But in return, I do expect a pre-release copy of the manuscript, so I have something intelligent to talk about.

Dr. Maslach: [laugh]. Of course, of course.

Corey: Thank you so much for your time. I really appreciate it.

Dr. Maslach: Well, thank you for having me. I appreciate the opportunity to share this, especially during these times.

Corey: Indeed. Professor Christina Maslach, Professor Emeritus of Psychology at Berkeley, I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an insulting comment telling me why you’re burned out on this show.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your
business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Scott

Scott is a web developer who has been blogging at https://hanselman.com for over a decade. He works in Open Source on ASP.NET and the Azure Cloud for Microsoft out of his home office in Portland, Oregon. Scott has three podcasts, http://hanselminutes.com for tech talk, http://thisdeveloperslife.com on developers' lives and loves, and http://ratchetandthegeek.com for pop culture and tech media. He's written a number of books and spoken in person to almost a half million developers worldwide.

Links:

  • Hanselminutes Podcast: https://www.hanselminutes.com/
  • Personal website: https://hanselman.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Thinkst. This is going to take a minute to explain, so bear with me. I linked against an early version of their tool, canarytokens.org in the very early days of my newsletter, and what it does is relatively simple and straightforward. It winds up embedding credentials, files, that sort of thing in various parts of your environment, wherever you want to; it gives you fake AWS API credentials, for example. And the only thing that these things do is alert you whenever someone attempts to use those things. It’s an awesome approach. I’ve used something similar for years. Check them out. But wait, there’s more. They also have an enterprise option that you should be very much aware of canary.tools. You can take a look at this, but what it does is it provides an enterprise approach to drive these things throughout your entire environment. You can get a physical device that hangs out on your network and impersonates whatever you want to. When it gets Nmap scanned, or someone attempts to log into it, or access files on it, you get instant alerts. It’s awesome. If you don’t do something like this, you’re likely to find out that you’ve gotten breached, the hard way. Take a look at this. It’s one of those few things that I look at and say, “Wow, that is an amazing idea. I love it.” That’s canarytokens.org and canary.tools. The first one is free. The second one is enterprise-y. Take a look. I’m a big fan of this. More from them in the coming weeks.

Corey: This episode is sponsored in part by our friends at Lumigo. If you’ve built anything from serverless, you know that if there’s one thing that can be said universally about these applications, it’s that it turns every outage into a murder mystery. Lumigo helps make sense of all of the various functions that wind up tying together to build applications. It offers one-click distributed tracing so you can effortlessly find and fix issues in your serverless and microservices environment. You’ve created more problems for yourself; make one of them go away. To learn more, visit lumigo.io.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’m joined this week by Scott Hanselman of Microsoft. He calls himself a partner program manager—or is called a partner program manager. But that feels like it’s barely scraping the surface of who and what he is. Scott, thank you for joining me.

Scott: [laugh]. Thank you for the introduction. I think my boss calls me that. It’s just one of those HR titles; it doesn’t really mean—you know, ‘program manager,’ what does it even mean?

Corey: I figure it means you do an awful lot of programming. One of the hardest questions is, you start doing different things—and Lord knows you do a lot of them—is that awful question that you wind up getting at cocktail parties of, “So, what is it you do exactly?” How do you answer that?

Scott: Yeah, it’s almost like, if you spent any time on Clubhouse recently, there was a wonderful comedian named Spunky Brewster on Instagram who had a whole thing where she talked about the introductions at the beginning of a Clubhouse thing, where it’s like, you’re a multi-hyphenate sandwich artist slash skydiver slash programmers slash whatever. One doesn’t want to get too full of one’s selves. I would say that I have for the last 30 years been a teacher and a professional enthusiast around computing and getting people excited about computing. And everything that I do, whether it be writing software, shipping software, or building community, hangs off of the fact that I’m an enthusiastic teacher.

Corey: You really are. And you’re also very hard to pin down. I mean, it’s pretty clear to basically the worst half of the internet, that you’re clearly a shill. The problem is defining exactly what you’re a shill for. You’re obviously paid by Microsoft, so clearly you push them well beyond the point when it would make sense to.

You have a podcast that has been on for over 800 episodes—which puts this one to shame—called Hanselminutes, and that is, of course, something where you’re shilling for your own podcast. You’ve recently started on TikTok, which I can only assume is what the kids are into these days. You’re involved in so many different things and taking so many different positions, that it’s very hard to pin down what is the stuff you’re passionate about.

Scott: I’m going to gently push back and say—

Corey: Please do.

Scott: That if one were to care to look at it holistically, I am selling enthusiasm around free and open-source software on primarily the Windows platform that I’m excited about, and I am selling empowerment for the next generation of people who want to do computing. Before I went to Microsoft, my blog and my podcast existed, and I was consistent in my, “Hey, have you heard the news?” Message to anyone who would listen. And I taught at both Portland Community College and Oregon Institute of Technology, teaching web services and history of the web and C# and all that kind of stuff. So, I’m one of those people where if you touch on a topic that I’m interested in, I’ll be like, “Oh, my goodness, let’s”—and I’ll just like, you know, knock everything off the desk and I’m going to be like, “Okay, let’s build a model, a working model of the solar system here, now. The orange is the sun.”

And it’s like, suddenly now we’re talking about science, like Hank Green or whatever. My family will ask me, “Why isn’t the remote control working?” And then I’ve taken it apart and I’m explaining to them how the infrared LED inside works. And, you know, how can you not be excited about all these things? And that’s my whole thing about computing and the power that being able to program computers represents to me.

Corey: I would agree with that. I’d say that one thing that is universal about everything you’re involved in is the expression I heard that I love and am going to recapture has been, “Sending the elevator back down.”

Scott: Oh, yeah. Throwing ladders, ropes, elevators. I am very blessed to have made it out of my neighborhood, and I am very hopeful that anyone who is in a situation that they do not want to be in could potentially use coding, programming, IT, computing as the great equalizer and that I can I could somehow lend my privilege to them to get the things done and solve the problems that they want to solve with computers.

Corey: I’m sure that you’ve been asked ad nauseum about—you work in free and open-source software. You’ve been an advocate for this, effectively, for your entire career; did no one tell you you work at Microsoft? But that’s old Microsoft in many respects. That’s something that we’ve covered with a bunch of different guests previously from Microsoft, and it’s honestly a little—it’s becoming a bit of a tired trope. It was a really interesting conversation a few years back that, oh, it’s clearly all just for show.

Well, that is less and less obvious, and more tired and frankly bad take as time progresses. So, I want to go back a bit further into my own personal journey because it turns out that the number one reason to reach out to you for anything is tech support on various things. I don’t talk about this often, but I started my career moonlighting as a Windows admin, back in the Windows 2003 server days; and it was an experience, and licensing was a colossal pain, and I finally had enough of it one day, in 2006, switched over to Unix administration on BSD, and got a Mac laptop, and that was really the last time that I used Windows in anger. Now, it’s been 15 years since that happened, and I haven’t really been tracking the Windows ecosystem. What have I missed?

Scott: [laugh]. There’s a lot there that you just said. So first, different people have their religions and they’re excited about them, and I encourage everyone to be excited about the religion that they’re excited about. It’s great to be excited about your thing, but it’s also really not cool to be a zealot about your thing. So hey, be excited about Windows, be excited about Linux, be excited about Mac.

Just don’t tell me that I’m going to heck because I didn’t share your enthusiasm. Let’s just be excited together and we can be friends together. I’ve worked on Linux at Nike, I’ve worked on Mac, I’ve worked on Windows, you know, I’ve been there before these things existed and I’ll be there afterwards.

Corey: Exactly. At some point being a zealot for a technology just sort of means you haven’t been around the block enough to understand how it’s going to break, how it’s going to fail, how it’s going to evolve, and it doesn’t lead to a positive outcome for anyone. It fundamentally becomes a form of gatekeeping more than anything else, and I just don’t have the stomach for it.

Scott: Yeah. And ultimately, we’re just looking for—you know, we got these smart rocks that we taught how to think with lightning, and they’re running for loops for us. And maybe they’re running them in the cloud, maybe they’re running locally. So, I’m not really too worried about it. Windows is my thing of choice, but just, you know, one person’s Honda is another person’s Toyota; you get excited about the brand that you start out with.

So, that’s that. Currently, though, Windows has gone, at least in the last maybe 20 years, from one of those things where there’s generational pain, and, like, “Microsoft killed my Pappy, and I’ll never forgive you.” And it’s like, yeah, there was some dumb stuff in the ’90s with Internet Explorer, but as a somewhat highly placed middle manager at Microsoft, I’ve never been in an active mustache-twirling situation where I was behind closed doors and anyone thought anything nefarious. There’s only a true, “What’s the right thing for the customer? What is the right thing for the people?”

My whole thing is to make it so developers can develop more easily on Windows, so I’m very fortunate to be helping some folks in a partnership between the Windows division and the developer division that I work in to make Windows kick butt when it comes to dev. Historically, the Windows terminal, or what’s called cmd.exe which is run by a thing called the console host has sucked; it has lagged behind. So, if you drop out to the command line, you’ve got the, you know, the old, kind of, quote-unquote, “DOS shell” with a cmd processor—it’s not really DOS—running in an old console host. And it’s been there for gosh, probably early ’90s. That sucks.

But then you got PowerShell. And again, I want to juxtapose the difference between a console—or a terminal—and a shell. They’re different things. There’s lots of great third-party terminals in the ecosystem. There’s lots of shells to choose from, whether it be PowerShell, PowerShell Core—now PowerShell 7.0—or the cmd, as well as bash, and Cygwin, and zsh, and fish.

But the actual thing that paints the text on Windows has historically not been awesome. So, the new open-source Windows terminal has been the big thing. If you’re a Machead and you use iTerm2, or Hyper, or things like that, you’ll find it very comfortable. It’s a tabbed terminal, split-screen, ripping fast, written in, you know, DirectX, C++ et cetera, et cetera, all open-source, and then it lets you do transparency, and background colors, and ligature fonts, and all the things that a great modern terminal would want to do. That is kind of the linchpin of making Windows awesome for developers, then gets even awesomer when you add in the ability that we’re now shipping an actual Linux kernel, and I can run N number of Linuxes side-by-side, in multiple panes, all within the terminal.

This getting to the point about juxtaposing the difference between a terminal and a console and a shell. So, I’ve got, on the machine, I’m talking to you on right now, on my third monitor, I’ve got Windows terminal open with PowerShell on Windows on the left, Ubuntu 18.04 LTS on the right, with the fish shell. And then I’ve got another Ubuntu 20.04 with bash, a standard bash shell.

And I’m going and testing stuff in Docker, and running .NET in Docker, and getting ready to deploy my own podcast website up into Azure. And I’m doing it in a totally organic way. It’s not like, “Oh, I’m just running a virtual machine.” No, it’s integrated. That’s what I think you’d be impressed with.

Corey: That right there is the reason that I generally tended to shy away from getting back into the Windows ecosystem for the longest time—and this is not a slam on Windows, by any stretch of the—

Scott: No of course. Sure, sure, sure.

Corey: —imagination—my belief has always been that you operate within the environment as it’s intended to be operated within, and it felt at the time, “Oh, install Cygwin, and get all this other stuff going, and run a VM to do it.” It felt like I was fighting upstream in some respects.

Scott: Oh, yeah, that’s a great point. Let’s talk about that for a second. So—

Corey: Let’s do it.

Scott: So, Cygwin is the GNU utilities that are written in a very nice portable C, but they are written against the Windows kernel. So, the example I like to use is ls, you type ls, you list out your directory, right? So, ls and dir are the same thing for this conversation. Which means that someone has to then call a system call—syscall in Linux, Windows kernel call in Windows—and say, “Hey, would you please enumerate these files, and then give me information about them, and check the metadata?” And that has to call the file system and then it’s turtles all the way down.

Cygwin isn’t Linux. It’s the bash and GNU utilities recompiled and compiled against the Windows stuff. So, it’s basically putting a bash skin on Windows, but it’s not Linux; it’s bash. Okay? But WSL is actually Linux, and rather than firing up a big 30 gig Hyper-V, or VirtualBox, or Parallels virtual machine, which is, like, a moment—“I’m firing up the VM; call me in an hour when it comes back up.”—and when the VM comes up, it’s, like, a square on your screen and now you’re dealing with another thing to manage.

The WSL stuff is actually a utility virtual machine built on a lower subsystem, the virtualization platform, and it starts in less than a second. You can start it faster than you can say, one one-thousand. And it goes instantly up, it automatically allocates and deallocates memory so that it’s smart about memory, and it’s running the actual Linux kernel, so it’s not pretending to be Linux. So, if your goal is a Linux environment and you’re a Linux developer, the time of Linux on the desktop is happening, in this case, on the Windows desktop. Where you get interesting stuff, and where I think your brain might explode is, imagine you’re in the terminal, you’re at the Linux file system at the bash prompt, and you type ‘notepad.exe.’ What would you expect to happen? You’d expect it to try to find it in a Linux path and fail.

Corey: Right. And then you’re trying to figure out, am I in this environm—because you generally tend to run these things in the same-looking terminal, but then all the syntax changes as soon as you go back into the Windows native environment, you’re having to deal with line-ending issues on a constant basis, and you just—

Scott: Oh, yeah. All that stuff, where.

Corey: And as soon as you ask for help because back in those days, I was looking primarily into using freenode as my primary source of support because I network staff on the network for the better part of a decade, and the answer is, “I’m having some trouble with Linux,” and the response is, “Oh, you’re doing this within a Windows environment? Get a real computer, kid.” Because it’s still IRC, and being condescending and rude to anyone who makes different choices than you do is apparently the way that was done back then.

Scott: Well, today in 2020 because we don’t want to just have light integration with Windows—and by light integration, like, I don’t know if you remember firing up a virtual machine on Windows and then, like, copy-pasting a file, and we were all going like, “Oh, my God, that’s amazing.” I drug the file in and then it did a little bit of magic and then moved the file from Windows into Linux. What we want is to blur the lines between the two so you can move comfortably. When you type explorer.exe or notepad.txt in Linux on Windows, Linux says no, and then Windows gets the chance, fires it up, and can access the Linux file system.

And since Notepad now understands line endings, just happily, you can open up your .profile, your bash_profile, your csh file in Notepad, or—here’s where it gets interesting—Visual Studio Code, and comfortably run your Windows apps, talking to your Linux file system, or in the—coming soon, and we’ve blogged about this and announced it at Build last year, run Linux GUI apps seamlessly so that I could have two browsers up, two Chromes, one Windows and one Linux, side-by-side, which is going to make web testing even that much easier. And I’m moving seamlessly between the two. Even cooler, I can type explorer.exe and then pass in dot, which represents the current folder, and if the current folder is the Linux file system, we seamlessly have a Plan 9 server—basically a file server that lets you access your Linux file system—from—

Corey: Is it actually running Plan 9?

Scott: It is a Plan 9 server.

Corey: That is amazing. I’m sorry, that is a blast from the past.

Scott: I’m glad. And we can run N number of Linuxes; this isn’t just one Linux. I’ve got Kali Linux, two different Ubuntus, and I could tar up the user mode files on mine, zip them up, give them to you, and you could go and type ‘wsl–import,’ and then have my Linux file system. Which means that we could make a custom Screaming in the Cloud distro, put it in the Windows Store, put it up on GitHub, build our own, and then the company could standardize on our Linux distro and run it on Windows.

Corey: That is almost as terrible an idea as using a DNS service as a database.

Scott: [laugh].

Corey: I love it. I’m totally there for it.

Scott: It’s really nice because it’s extremely—the point is, it has to have no friction, right? So, if you think about it this way, I just moved—I blogged about this; if people want to go and learn about it—I just moved my blog of 20 years off of a Windows Server 2008 server running under someone’s desk at a host, into Azure. This is a multi-month-long migration. My blog, my main site, kind of the whole Hanselman ecosystem moved up in Azure. So, I had a couple things to deal with.

Am I going to go from Windows to Linux? Am I going to go from a physical machine to a virtual machine? Am I going to go from a physical machine to a virtual machine to a Platform as a Service? And when I do that, well, how is that going to change the way that I write software? I was opening it in Visual Studio, pressing F5, and running it in IIS—the Internet Information Server for Windows—for the last 15, 20 years.

How do I change that experience? Well, I like Visual Studio; I like pressing F5; I like interactive debugging sessions. But I also like saving money running Linux in the cloud, so how can I have the best of all those worlds? Because I wrote the thing in .NET, I moved into .NET 5, which runs everywhere, put together a Docker file, got full support for that in Visual Studio, moved it over into WSL so I can test it on both Windows and Linux.

I can go into my folder on my WSL, my Windows subsystem for Linux, type code dot, open up Visual Studio Code. Visual Studio Code splits in half. The Windows client of Visual Studio Code runs on Windows; the server, the Visual Studio Code server, runs in WSL providing the bridge between the two worlds, and I can press F5 and have interactive debugging and now I’m a Linux developer even though I’ve never left Windows. Then I can right-click publish in Visual Studio to GitHub Actions, which will then throw it into the cloud, and I moved everything over into Azure, saved 30%, and everything’s awesome. I’m still a Windows developer using Visual Studio. So, it’s pretty much I don’t know, non-denominational; kind of mixing the streams here.

Corey: It is. And let me take it a step further. When I’m on the road, the only computer I bring with me these days—well, in the before times, let’s be very realistic. Now, when ‘I’m on the road,’ that means going to the kitchen for a snack—the only computer I bring with me is my iPad Pro, which means that everything I do has a distinct application. For when I want to get into my development environment, historically it was, use some terminal app—I’m a fan of Blink, but everyone has their own; don’t email me.

And everything else I tended to use looked an awful lot like a web app. If there wasn’t a dedicated iOS app, it was certainly available via a web browser. Which leads me to the suspicion that we’re almost approaching a post-operating-system world where the future development operating system begins to look an awful lot—and people are going to yell at me for this—Visual Studio Code.

Scott: Mmm.

Corey: It supports a bunch of remote activities now that GitHub Codespaces is available—at least to my account; I don’t know if it’s generally available yet—but I’ve been using it; I love it; everything it winds up doing is hosted remotely in Azure; I don’t have to think about managing the infrastructure; it’s just another tab within GitHub, and it works. My big problem is that I’m trying to shake, effectively, 20 years of muscle memory of wrestling with Vim, and it takes a little bit of a leap in order to become comfortable with something that’s a more visually-oriented IDE.

Scott: Why don’t you use the VsVim, Jared Parsons Vim plugin for Visual Studio?

Corey: I’ve never yet found a plugin that I like for something else to make it behave like Vim. Vimperator is a browser extension, all of it just tends to be unfortunate and annoying in different ways. For whatever reason, the way that I’m configured or built, it doesn’t work for me in the same way. And it goes back to our previous conversation about using the native offering as it comes, rather than trying to make it look like something else.

Scott: Okay. I would just offer to you and for other Vim people who might be listening, that VS Code Vim does have 2.5 million installs, over 2 million people happily using that. And they are—

Corey: Come to find it only has 200,000 actual users; there was an installation bug and one person just kept trying over and over and over. I kid, I kid.

Scott: No, seriously though, these are actual Vim-heads and Jared Parsons is a developer at Microsoft who is like, out of his cold dead hands you’ll pull his Vim. So, there’s solutions; whether you’re Vim or Emacs, you know, we welcome all comers. But to your point, the Visual Studio, once it got split in half, where the language services, those services that provide context to Python, Ruby, C# C++ et cetera, once those extensions can be remoted, they can run on Windows, they can run on Linux, they can run on the cloud. So, VS Code being split in half as a client-server application has really made it shine. And for me, that means that I don’t notice a difference, whether I’m running VS Code on Windows or running VS Code to a remote Linux install, or even using SSH and coding on Windows remotely to a Raspberry Pi.

Corey: I love the idea. I’ve seen people do this, in some respects, back in the days of Code Server being a project on GitHub, and it took a fair bit of wrangling to get that to work in a way that wasn’t scarily insecure and reliable. But once it was up and running, you could effectively plug a Raspberry Pi in underneath your iPad and effectively have a portable computer on the go that did local development. I’m looking at this and realizing the future doesn’t look at all like what I thought it was going to, and it’s really still kind of neat.

Scott: Mm-hm.

Corey: There’s a lot of value in being able to make things like this more accessible, and the reason I’m excited about a lot of this, too, is that aligned with a generous free tier opportunity, which I don’t know final pricing for things like GitHub Codespaces, suddenly the only real requirement is something that can render a browser and connect to the internet for an awful lot of folks to get started. It doesn’t require a fancy local overpowered development machine the way a lot of things used to. And yes, I know; there are certain kinds of development that are changing in that respect, but it still feels to me like it has never been easier to get started with all of this technology than ever before, with a counterargument that there’s so many different directions to go in. “Oh, I want to get started using Visual Studio Code or learning to write JavaScript. Great. How do I do this? Let me find a tutorial.” And you find 20 million tutorials, and then you’re frozen with indecision. How do you get past that?

Scott: Yeah, there is and always will be, unfortunately, a certain amount of analysis paralysis that occurs. I started a TikTok recently to try to help people to get involved in coding, and the number one question I get—and I mean, thousands and thousands of them—are like, “Where do I start?” Because everyone seems to think that if they pick the wrong language, that will be a huge mistake. And I can’t think of a wrong language, you know? Like, what human language should I learn?

You know, English, Chinese, Arabic, Japanese. Pick one and then learn another one if you can. Learn a couple. But I don’t think there’s a wrong language to learn because the basics of computer science are the basics of computer science. I think what we need to do is remind people that computers are computers no matter whether they’re an Android phone or a Windows laptop, and that any forward motion at all is a good thing. I think a lot of people have analysis paralysis, and they’re just afraid to pick stuff.

Corey: I agree with what you’re saying, but I’m also going to push back gently on what you’re saying, as well. If someone who is new to the field was asking me what language to learn, I would be hard-pressed to recommend a language that was not JavaScript. I want to be clear, I do not understand or know JavaScript at all, but it’s clear from what I’m seeing, that is, in many ways, the language of the future. It is how frontend is being interacted with; there are projects from every cloud provider that wind up managing infrastructure via JavaScript primitives. There are so many on-ramps for this, and the user experience for new folks is phenomenal compared to any language that I’ve worked with in my career. Would you agree with that or disagree with that assessment?

Scott: So, I’ve written blog posts on this topic, and my answer is a little more ‘it depends.’ I say that people should always learn JavaScript and one other language, preferably a systems language, which also may be JavaScript. But rather than thinking about things language-first, we think about things solutions-first. If someone says, “I want to do a lot of data science,” you don’t learn JavaScript. If someone says, “I want to go and write an Android app,” yeah, you could do that in JavaScript, but JavaScript is not the answer to all questions.

Just as the English language, while it may be the lingua franca, no pun intended, it is not the only language one should pick. I usually say, “Well, what do you want to do?” “Well, I want to write a video game for the Xbox.” Okay, well, you’re probably not going to do that in JavaScript. “Oh, I want to do data science. I want to write an iPhone app.” JavaScript is the language you should learn if you’re going to be doing things on the web, yes, but if you’re going to be writing the backend for WhatsApp, then you’re not going to do that JavaScript.

Corey: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the Enterprise (not the starship). On-prem security doesn’t translate well to cloud or multi-cloud environments, and that’s not even counting IoT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IoT devices, detects these threats up to 35 percent faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at extrahop.com/trial.

Corey: Yeah, I think you’re right. It comes down to what is the problem you’re trying to solve for? Taking the analogy back to human languages, well, what is your goal? Is it just to say that you’ve learned a language and to understand, get a glimpse at another culture through its language? Yeah, there is no wrong answer. If it’s that you want to go live in France one day and participate in French business discussions, I have a recommendation for you, and it’s probably not Sanskrit.

At some point, you have to align with what people want to do and the direction they’re going in with the language selection. What I like about JavaScript is, frankly, it’s incredible versatility as far as problems to which it can be applied. And without it, I think you’re going to struggle as you enter the space. My first language was crappy Perl—slash bash because everyone does bash when you’re a systems administrator—and then it has later evolved now to crappy Python as my language of choice. But I’m not going to be able to effectively do any frontend work in Python, nor would I attempt to do so.

My way of handling frontend work now is to have the good sense to pay a professional. But if you’re getting started today and you’re not sure what you want to do in your career, my opinion has always been that if you think you know what you want to do in your career, there’s a great chance you’re going to be wrong, but pursuing the thing that you think you want to do will open other opportunities and doors, and present things to you that will catch your interest in a way you might not be able to anticipate. So, especially early on in careers, I like biasing for things that give increased options, that boost my optionality as far as what I’m going to be able to do.

Scott: Okay. I think that’s fair. I think that no one ever got fired for picking IBM; [laugh] no one ever jeopardized their career by choosing JavaScript. I do think it’s a little more nuanced, as I mentioned.

Corey: It absolutely is. I am absolutely willing to have a disagreement with you on that front. I think the thing that we’re aligned on is that whatever you pick, make sure it’s something you’re interested in. Don’t do it just for—like, “Well, I’m told I can make a lot of money doing X.” That feels like it’s the worst reason to do things, in isolation.

Scott: That’s a tough one. I used to think that, too, but I am thinking that it’s important to note and recognize that it is a valid reason to get into tech, not for the passion because for no other reason that I want to make a lot of money.

Corey: Absolutely. I could not agree with you more, and that is… something I’ve gotten wrong in the past.

Scott: Yeah. And I have been a fan of saying, you know, “Be passionate and work on these things on the side,” and all that kind of stuff. But all of those things involve a lot of assumptions and a lot of privileges that, you know, people have: that you have spare time and that you have a place to work on these things. I work on stuff on the side because it feeds my spirit. If you work on woodworking, or drones, or gardening on the side, you know, not everything you work on the side has to be steeped in hustle culture and having a startup, or something that you’re doing on the side.

Corey: Absolutely. If you’re looking at a position of wanting to get into technology because it leads to a better financial outcome for you and that is what motivates you, you’re not wrong.

Scott: Exactly.

Corey: The idea that, “Oh, you have to love it or you’ll never succeed.” I think that some of the worst advice we ever wind up giving folks early in their career—particularly young people—is, ‘follow your passion.’ That can be incredibly destructive advice in some contexts, depending upon what it is you want to do and what you want your life to look like.

Scott: Yeah, exactly.

Corey: One of the things that I’ve always been appreciative of from afar with Microsoft has been there’s an entire developer ecosystem, and historically, it’s focused on languages I can barely understand: ASP.NET, the C# is deep in that space, F#, I think, is now a thing as well. There’s an entire ecosystem around this with Visual Studio the original, not Visual Studio Code—turns out naming is one of those things that no tech companies seems to get right—but it feels almost like there’s an entire ecosystem there for those of us who spent significant time—and I’m speaking for myself here, not you—in the open-source community talking about things like Perl and whatnot, I never got much exposure to stuff like that. I would also classify Enterprise Java as being in that direction as well. Is there a bifurcation there that I’m not seeing, or was I just never talking to the right people? All the above? Maybe I was just—maybe I had blinders on; didn’t realize it.

Scott: There was a time when the Microsoft developer ecosystem meant write things for Windows, do things on Windows, use languages that Microsoft made and created. And now, with the rise of the cloud and with the rise of Software as a Service, Microsoft is a much simpler company, which is a funny thing to say for such a complicated company. Microsoft would love to run your for loop in the cloud for money. We don’t care what language you use; we want you to use the language that makes you happy. Somewhere around five to seven years ago, in the developer division, we started optimizing for developer happiness.

And that’s why you can write Ruby, and Perl, and Python, and C, and C++ and C# and all those different things. Even C# now, and .NET, is owned by the .NET Foundation and not by Microsoft. Microsoft, of course, is one of the primary users, but we’ve got a lot of—Samsung is a huge contributor, Google is a huge contributor, Amazon Web Services is a big contributor to .NET.

So, Microsoft’s own zealotry towards—and bias towards our own languages has, kind of, gone away because Office is on iPhone, right? Like, anywhere that you are, we’ll go there. So, we’re really going where the customer is rather than trying to funnel the customer into where we want them to be, which is a really an inverted way of doing things over the way it was done 20, 30 years ago. In my opinion.

Corey: This gets back to the idea of the Microsoft cultural transformation. It hasn’t just been an internal transform; it’s been something that is involved with how it’s engaging with its customers, how it’s engaging with the community, how it’s becoming available in different ways to different folks. It’s hard to tell where a lot of these things start and where a lot of these things stop. I don’t pretend to be a Microsoft “fanboy,” quote-unquote, but I believe it is impossible to look at what has happened, especially in the world of cloud, and not at the very least respect what Microsoft has been able to achieve.

Scott: Well, I came here to open source stuff. I’m surely not responsible for the transformation, I’m just a cog in the machine, but I can speak for the things that I own, like .NET and Visual Studio Community, and I think one of the things that we have gotten right is we are trying to create zero-distance products. You could be using Visual Studio Code, find a bug, suggest a feature, have a conversation in public with the PMs and devs that own the thing, get an insider’s build a few days later, and see that promoted to production within a week or two. There is zero distance between you the consumer and the creator of the thing.

And if you wanted to even fix the bug yourself, submit a pull request, and see that go into production, you could do that as well. You know, some of our best C# compiler folks are not working for Microsoft and they are giving improvements, they are making the product better. So, zero-distance in many ways, if you look at the other products at Microsoft, like PowerToys is a great thing, which is [unintelligible 00:32:06] an incubator for Windows features. We’re adding stuff to the PowerToys open-source project like launchers, and a thing called FancyZones that is a window tiling manager, you know, features that prosumers and enthusiasts always wished Windows could have, they can now participate in, thereby creating a zero-distance product in Windows itself.

Corey: And I want to point out as well that you are still Microsoft. You, the collective you. I suppose you personally; that is where your email address ends. But you’re still Microsoft. This is still languages, and tools, and SDKs, and frameworks used by the largest companies in the world. This zero-distance approach is being done on things that service banks, who are famously not the earliest adopters of some code that I wrote last night; it’s probably fine.

Scott: Do you know what my job was before I came here?

Corey: Tell me.

Scott: I was the chief architect at a finance company that created software for banks. I was responsible for a quarter of the retail online banking systems in North America, built on .NET and open-source software. [laugh].

Corey: So, you’ve lived that world. You’ve been that customer.

Scott: Trying to convince a bank that open-source was a good idea in the early 2000s was non-trivial. You know, sitting around in 2003, 2004, talking about Agile, and you know, continuous integration, and build servers, and then going and saying, “Hey, you should use the software,” trying to deal with lawyers and explain to them the difference between the MIT, Apache, and GPL licenses and what it means to their bank was definitely a challenge. And working through those issues, it has been challenging. But open-source software now pervades. Just go and look at the license.txt in the Visual Studio Program Files folder to see all of the open-source software that is consumed by Visual Studio.

Corey: One last topic that I want to get to before we call it a show is that you’ve spent a significant portion of your career, at least recently, focusing on, more or less, where the next generation of engineers, developers, et cetera, come from. And to that end, you’ve also started recently with TikTok, the social media platform. Are those two things related, first off, or am I making a giant pile of unwarranted assumption?

Scott: [laugh]. I think that is a fair assumption. So, what’s going on is I want to make sure that as I fade away and I leave the software industry in the next, you know, N number of years, that I’m setting up as many people as possible for success. That’s where my career started when I was a professor, and that’s hopefully where my career will end when I am a professor again. Hopefully, my retirement gig will have me teaching at some university somewhere.

And in doing that, I want to find the next million developers, right? Where are they, the next 10 million developers? They’re probably not on Twitter. They might be a lot of different places: they might be on Discord, they might be on Reddit, they might be on forums that I haven’t found yet. But I have found, on TikTok, a very creative and for the most part kind and inclusive community.

And both myself and also recently, the Visual Studio Code team have been hanging out there, and sharing our creativity, and having really interesting conversations about how you the listener can if not be a programmer, be a person that knows better the tools that are available to you to solve problems.

Corey: So, I absolutely appreciate and enjoy the direction that you’re going in, but again, people invite you to things and then spring technical support questions on you. Can you explain what TikTok is? I’m still trying to wrap my head around it because I turned around and discovered I was middle-aged one day.

Scott: Sure. Well, I mean, I am an old man on TikTok, to be clear. TikTok, like Twitter, revels in its constraints. If you recall, there was a big controversy when Twitter went from 140 characters to 280 because people thought it was just letting the constraint that we were so excited about—which was artificial because it was the length of a standard message service text—

Corey: I’m one of those people who bitterly protested it. I was completely wrong.

Scott: Right? But the idea that something is constrained, that TikTok is either 15 seconds, or less than 60, it’s similar to Vine in that it is a tiny video; what can I do in one minute? Additionally, before they allowed uploading of videos, everything was constrained within the TikTok editor, so people would do amazing and intricate 30 and 40 shot transitions within a 60 second period of time. But one of the things I find most unique about TikTok is you can reply to a text comment with a video. So, I make a video—maybe I do 60 seconds on how to be a software engineer—somebody replies in text, I can then reply to that text with a video, and then a TikTok creator can do what’s called a stitch and reply to my video with a video.

So, I could take 15 seconds of yours, a comment that you made, and say, “Oh, this is a great comment. Here’s my thoughts on that comment.” Or we could even do a duet where you record a video and then I record one, side-by-side. And we either simulate that we’re actually having a conversation, or I react to your video as well. Once you start teaching TikTok about yourself by liking things, you curate a very positive place for yourself.

You might get on TikTok, not logged in, and it’s dancing, and you might find some inappropriate things that you don’t necessarily want to see, or you’re not interested in, but one of the things that I’ve noticed as I talk about my home network and coding is people will say, “Oh, I finally found adjective TikTok; I finally found coding TikTok I finally found IT TikTok. Oh, I’m going to comment on your post because I want to stay on networking TikTok.” And then your feed isn’t just a feed of the people that you follow, but it’s a feed of all the things that TikTok thinks you’re excited about. So, I am on this wonderful TikTok of linguistics and languages, and I’m learning about cultures, and I’m on indigenous TikTok, and I’m on networking TikTok. And the mix of creativity and the constraint of just 60 seconds has been, really, a joy. And I’ve only been there for about a month and I’ve blessed to have 80,000 people hanging out with me there.

Corey: It sounds like you’re quite the fan of the platform, which alone in isolation, is enough to get me to look at it in more depth.

Scott: I am a fan of creativity. I would also say though, it’s very addictive once you find your people. I’ve had to put screen time limits on my own phone to keep me from burning time there.

Corey: That is all of tempting, provocative, and disturbing. I—

Scott: You should hang out with me on YouTube, then. I just got my 100,000 YouTube Silver Play Button in the mail. That’s where I spend my time doing my long-form. I just did, actually, 17 minutes on WSL and how to use Linux. That might be a good starter for you.

Corey: It very well might. So, if people want to learn more about what you’re up to, and how you think about the wide variety of things you’re interested in, where can they find you?

Scott: They should start at my last name dot com: Hanselman.com. They used to be able to Google for Scott, and I was in an epic battle with Scott brand toilet paper tissue, and then they trademarked the name Scott and now I’m somewhere in the distant second or third page. It was a tragedy. But as an early comer—

Corey: Oh, my condolences.

Scott: Yeah, oh my God. As an early comer to the internet, it was me and Scott Fly Rods on the first page, for many, many years. And then—

Corey: If it helps, you and Scott Fly Rods are both on page two.

Scott: Oh. Well, the tyranny of the Scott toilet paper conspiracy against me has been problematic.

Corey: Exactly.

Scott: [laugh].

Corey: Thank you so much for taking the time to speak with me today. I really do appreciate it.

Scott: It’s my pleasure.

Corey: Scott Hanselman, partner program manager at Microsoft and so much more. I’m Cloud Economist Corey Quinn. This is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with a crappy comment that starts with a comment that gatekeeps a programming language so we know to ignore it.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Martin

Martin Mao is the co-founder and CEO of Chronosphere. He was previously at Uber, where he led the development and SRE teams that created and operated M3. Prior to that, he was a technical lead on the EC2 team at AWS and has also worked for Microsoft and Google. He and his family are based in our Seattle hub and he enjoys playing soccer and eating meat pies in his spare time.

Links:

  • Chronosphere: https://chronosphere.io/
  • Email: contact@chronosphere.io

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Thinkst. This is going to take a minute to explain, so bear with me. I linked against an early version of their tool, canarytokens.org in the very early days of my newsletter, and what it does is relatively simple and straightforward. It winds up embedding credentials, files, that sort of thing in various parts of your environment, wherever you want to; it gives you fake AWS API credentials, for example. And the only thing that these things do is alert you whenever someone attempts to use those things. It’s an awesome approach. I’ve used something similar for years. Check them out. But wait, there’s more. They also have an enterprise option that you should be very much aware of canary.tools. You can take a look at this, but what it does is it provides an enterprise approach to drive these things throughout your entire environment. You can get a physical device that hangs out on your network and impersonates whatever you want to. When it gets Nmap scanned, or someone attempts to log into it, or access files on it, you get instant alerts. It’s awesome. If you don’t do something like this, you’re likely to find out that you’ve gotten breached, the hard way. Take a look at this. It’s one of those few things that I look at and say, “Wow, that is an amazing idea. I love it.” That’s canarytokens.org and canary.tools. The first one is free. The second one is enterprise-y. Take a look. I’m a big fan of this. More from them in the coming weeks.

Corey: If your mean time to WTF for a security alert is more than a minute, it's time to look at Lacework. Lacework will help you get your security act together for everything from compliance service configurations to container app relationships, all without the need for PhDs in AWS to write the rules. If you're building a secure business on AWS with compliance requirements, you don't really have time to choose between antivirus or firewall companies to help you secure your stack. That's why Lacework is built from the ground up for the Cloud: low effort, high visibility and detection. To learn more, visit lacework.com.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’ve often talked about observability, or as I tend to think of it when people aren’t listening, hipster monitoring. Today, we have a promoted episode from a company called Chronosphere, and I’m joined today by Martin Mao, their CEO and co-founder. Martin, thank you for coming on the show and suffering my slings and arrows.

Martin: Thanks for having me on the show, Corey, and looking forward to our conversation today.

Corey: So, before we dive into what you’re doing now, I’m always a big sucker for origin stories. Historically, you worked at Microsoft and Google, but then you really sort of entered my sphere of things that I find myself having to care about when I’m lying awake at night and the power goes out by working on the EC2 team over at AWS. Tell me a little bit about that. You’ve hit the big three cloud providers at this point. What was that like?

Martin: Yeah, it was an amazing experience, I was a technical lead on one of the EC2 teams, and I think when an opportunity like that comes up on such a core foundational project for the cloud, you take it. So, it was an amazing opportunity to be a part of leading that team at a fairly early stage of AWS and also helping them create a brand new service from scratch, which was AWS Systems Manager, which was targeted at fleet-wide management of EC2 instances, so—

Corey: I’m a tremendous fan of Systems Manager, but I’m still looking for the person who named Systems Manager Session Manager because, at this point, I’m about to put a bounty out on them. Wonderful service; terrible name.

Martin: That was not me. So, yes. But yeah, no, it was a great experience, for sure, and I think just seeing how AWS operated from the inside was an amazing learning experience for me. And being able to create foundational pieces for the cloud was also an amazing experience. So, only good things to say about my time at AWS.

Corey: And then after that, you left and you went to Uber where you led development and SRE teams that created and operated something called M3. Alternately, I’m misreading your bio, and you bought an M3 from BMW and went to drive for Uber. Which is it?

Martin: I wish it was the second one, but unfortunately, it is the first one. So yes, I did leave AWS and joined Uber in 2015 to lead a core part of their monitoring and eventually larger observability team. And that team did go on to build open-source projects such as M3—which perhaps we should have thought about the name and the conflict with the car when we named it at the time—and other projects such as Jaeger for distributed tracing as well, and a logging backend system, too. So, yeah, definitely spent many years there building out their observability stack.

Corey: We’re going to tie a theme together here. You were at Microsoft, you were at Google, you were at AWS, you were at Uber, and you look at all of this and decide, “All right. My entire career has been spent in large companies doing massive globally scaled things. I’m going to go build a small startup.” What made you decide that, all right, this is something I’m going to pursue?

Martin: So, definitely never part of the plan. As you mentioned, a lot of big tech companies, and I think I always got a lot of joy building large distributed systems, handling lots of load, and solving problems at a really grand scale. And I think the reason for doing a startup was really the situation that we were in. So, at Uber as I mentioned, myself and my co-founder led the core part of the observability team there, and we were lucky to happen to solve the problem, not just for Uber but for the broader community, especially the community adopting cloud-native architecture. And it just so happened that we were solving the problem of Uber in 2015, but the rest of the industry has similar problems today.

So, it was almost the perfect opportunity to solve this now for a broader range of companies out there. And we already had a lot of the core technology built-in open-source as well. So, it was more of an opportunity rather than a long-term plan or anything of that sort, Corey.

Corey: So, before we dive into the intricacies of what you’ve built, I always like to ask people this question because it turns out that the only thing that everyone agrees on is that everyone else is wrong. What is the dividing line, if any, between monitoring and observability?

Martin: That’s a great question, and I don’t know if there’s an easy answer.

Corey: I mean, my cynical approach is that, “Well, if you call it monitoring, you don’t get to bring in SRE-style salaries. Call it observability and no one knows what the hell we’re talking about, so sure, it’s a blank check at that point.” It’s cynical, and probably not entirely correct. So, I’m curious to get your take on it.

Martin: Yeah, for sure. So, you know, there’s definitely a lot of overlap there, and there’s not really two separate things. In my mind at least, monitoring, which has been around for a very long time, has always been around notification and having visibility into your systems. And then as the system’s got more complex over time, being able to understand that and not just have visibility into it but understand it a little bit more required, perhaps, additional new data types to go and solve those problems. And that’s how, in my mind, monitoring sort of morphed into observability. So, perhaps one is a subset of the other, and they’re not competing concepts there. But at least that’s my opinion. I’m sure there are plenty out there that would, perhaps, disagree with that.

Corey: On some level, it almost hits to the adage of, past a certain point of scale with distributed systems, it’s never a question of is the app up or down, it’s more a question of how down is it? At least that’s how it was explained to me at one point, and it was someone who was incredibly convincing, so I smiled and nodded and never really thought to question it any deeper than that. But I look back at the large-scale environments I’ve been in, and yeah, things are always on fire, on some level, and ideally, there are ways to handle and mitigate that. Past a certain point, the approach of small-scale systems stops working at large scale. I mean, I see that over in the costing world where people will put tools up on GitHub of, “Hey, I ran this script, and it works super well on my 10 instances.”

And then you try and run the thing on 10,000 instances, and the thing melts into the floor, hits rate limits left and right because people don’t think in terms of those scales. So, it seems like you’re sort of going from the opposite end. Well, this is how we know things work at large scale; let’s go ahead and build that out as an initially smaller team. Because I’m going to assume, not knowing much about Chronosphere yet, that it’s the sort of thing that will help a company before they get to the hyperscaler stage.

Martin: A hundred percent, and you’re spot on there, Corey. And it’s not even just a company going from small-stage, small-scale simple systems to more complicated ones, actually, if you think about this shift in the cloud right now, it’s really going from cloud to cloud-native. So, going from VMs to container on the infrastructure tier, and going from monoliths to microservices. So, it’s not even the growth of the company, necessarily, or the growth of the load that the system has to handle, but this shift to containers and microservices heavily accelerates the growth of the amount of data that gets produced, and that is causing a lot of these problems.

Corey: So, Uber was famous for disrupting, effectively, the taxi market. What made you folks decide, “I know. We’re going to reinvent observability slash monitoring while we’re at it, too.” What was it about existing approaches that fell down and, I guess, necessitated you folks to build your own?

Martin: Yeah, great question, Corey. And actually, it goes to the first part; we were disrupting the taxi industry, and I think the ability for Uber to iterate extremely fast and respond as a business to changing market conditions was key to that disruption. So, monitoring and observability was a key part of that because you can imagine it was providing all of the real-time visibility to not only what was happening in our infrastructure and applications, but the business as well. So, it really came out of a necessity more than anything else. We found that in order to be more competitive, we had to adopt what is probably today known as cloud-native architecture, adopt running on containers and microservices so that we can move faster, and along with that, we found that all of the existing monitoring tools we were using, weren’t really built for this type of environment. And it was that that was the forcing function for us to create our own technologies that were really purpose-built for this modern type of environment that gave us the visibility we needed to, to be competitive as a company and a business.

Corey: So, talk to me a little bit more about what observability is. I hear people talking about it in terms of having three pillars; I hear people talking about it, to be frank, in a bunch of ways so that they’re trying to, I guess, appropriate the term to cover what they already
are doing or selling because changing vocabulary is easier than changing an entire product philosophy. What is it?

Martin: Yeah, we actually had a very similar view on observability, and originally we thought that it is a combination of metrics, logs, and traces, and that’s a very common view. You have the three pillars, it’s almost like three checkboxes; you tick them off, and you have, quote-unquote, “Observability.” And that’s actually how we looked at the problem at Uber, and we built solutions for each one of those and we checked all three boxes. What we’ve come to realize since then is perhaps that was not the best way to look at it because we had all three, but what we realized is that actually just having all three doesn’t really help you with the ultimate goal of what you want from this platform, and having more of each of the types of data didn’t really help us with that, either. So, taking a step back from there and when we really looked at it, the lesson that we learned in our view on observability is really more from an end-user perspective, rather than a data type or data input perspective.

And really, from an end-user perspective, if you think about why you want to use your monitoring tool or your observability tool, you really
want to be notified of issues and remediate them as quickly as possible. And to do that, it really just comes down to answering three questions. “Can I get notified when something is wrong? Yes or no? Do I even know something is wrong?”

The second question is, “Can I triage it quickly to know what the impact is? Do I know if it’s impacting all of my customers or just a subset of them, and how bad is the issue? Can I go back to sleep if I’m being paged at two o’clock in the morning?”

And the third one is, “Can I figure out the underlying root cause to the problem and go and actually fix it?” So, this is how we think about the problem now, is from the end-user perspective. And it’s not that you don’t need metrics, logs, or distributed traces to solve the problem, but we are now orienting our solution around solving the problem for the end-user, as opposed to just orienting our solution around the three data types, per se.

Corey: I’m going to self-admit to a fun billing experience I had once with a different monitoring vendor whom I will not name because it turns out, you can tell stories, you can name names, but doing both gets you in trouble. It was a more traditional approach in a simpler time, and they wound up sending me a message saying, “Oh, we’re hitting rate limits on CloudWatch. Go ahead and open a ticket asking for them to raise it.” And in a rare display of foresight, AWS respond to my ticket with a, “We can do this, but understand at this level of concurrency, it will cost something like $90,000 a month on increased charges, with that frequency, for that many metrics.” And that was roughly twice what our AWS bill was in those days, and, “Oh.” So, I’m curious as to how you can offer predictable pricing when you can have things that emit so much data so quickly. I believe you when you say you can do it; I’m just trying to understand the philosophy of how that works.

Martin: As I said earlier, we started to approach this by trying to solve it in a very engineering fashion where we just wanted to create more efficient backend technology so that it would be cheaper for the increased amount of data. What we realized over time is that no matter how much cheaper we make it, the amount of data being produced, especially from monitoring and observability, kept increasing, and not even in a linear fashion but in an exponential fashion. And because of that, it really switched the problem not to how efficiently can we store this, it really changed our focus of the problem to how our users using this data, and do they even understand the data that’s being produced? So, in addition to the couple of properties I mentioned earlier, around cost accounting and rate-limiting—those are definitely required—the other things we try to make available for our end-users is introspection tools such that they understand the type of data that’s being produced. It’s actually very easy in the monitoring and observability world to write a single line of code that actually produces a lot of data, and most developers don’t understand that that single line of code produces so much data.

So, our approach to this is to provide a tool so that developers can introspect and understand what is produced on the backend side, not what is being inputted from their code, and then not only have an understanding of that but also dynamic ways to deal with it. So that again, when they hit the rate limit, they don’t just have to monitor it less, they understand that, “Oh, I inserted this particular label and now I have 20 times the amount of data that I needed before. Do I really need that particular label in there> and if not, perhaps dropping it dynamically on the server-side is a much better way of dealing with that problem than having to roll back your code and change your metric instrumentation.” So, for us, the way to deal with it is not to just make the backend even more efficient, but really to have end-users understand the data that they’re producing, and make decisions on which parts of it is really useful and which parts of it do they, perhaps not want or perhaps want to retain for shorter periods of time, for example, and then allow them to actually implement those changes on that data on the backend. And that is really how the end-users control the bills and the cost themselves.

Corey: So, there are a number of different companies in the observability space that have different approaches to what they solve for. In some cases, to be very honest, it seems like, well, I have 15 different observability and monitoring tools. Which ones do you replace? And the answer is, “Oh, we’re number 16.” And it’s easy to be cynical and down on that entire approach, but then you start digging into it and they’re actually right.

I didn’t expect that to be the case. What was your perspective that made you look around the, let’s be honest, fairly crowded landscape of observability companys’ tools that gave insight into the health status and well being of various applications in different ways, and say, “You know, no one’s quite gotten this right, yet. I have a better idea.”

Martin: Yeah, you’re completely correct, and perhaps the previous environments that everybody was operating in, there were a lot of different tools for different purposes. A company would purchase an infrastructure monitoring tool, or perhaps even a network monitoring tool, and then they would have, perhaps, an APM solution for the applications, and then perhaps BI tools for the business. So, there was always historically a collection of different tools to go and solve this problem. And I think, again, what has really happened recently with this shift to cloud-native recently is that the need for a lot of this data to be in a single tool has become more important than ever. So, you think about your microservices running on a single container today, if a single container dies in isolation without knowing, perhaps, which microservice was running on it doesn’t mean very much, and just having that visibility is not going to be enough, just like if you don’t know which business use case that microservice was serving, that’s not going to be very useful for you, either.

So, with cloud-native architecture, there is more of a need to have all of this data and visibility in a single tool, which hasn’t historically happened. And also, none of the existing tools today—so if you think about both the existing APM solutions out there and the existing hosted solutions that exist in the world today, none of them were really built for a cloud-native environment because you can think about even the timing that these companies were created at, you know, back in early 2010s, Kubernetes and containers weren’t really a thing. So, a lot of these tools weren’t really built for the modern architecture that we see most companies shifting towards. So, the opportunity was really to build something for where we think the industry and everyone’s technology stack was going to be as opposed to where the technology stack has been in the past before. And that was really the opportunity there, and it just so happened that we had built a lot of these solutions for a similar type environment for Uber many years before. So, leveraging a lot of our lessons learned there put us in a good spot to build a new solution that we believe is fairly different from everything else that exists today in the market, and it’s going to be a good fit for companies moving forward.

Corey: So, on your website, one of the things that you, I assume, put up there just to pick a fight—because if there’s one thing these people love, it’s fighting—is a use case is outgrowing Prometheus. The entire story behind Prometheus is, “Oh, it scales forever. It’s what the hyperscalers would use. This came out of the way that Google does things.” And everyone talks about Google as if it’s this mythical Valhalla place where everything is amazing and nothing ever goes wrong. I’ve seen the conference talks. And that’s great. What does outgrowing Prometheus look like?

Martin: Yeah, that’s a great question, Corey. So, if you look at Prometheus—and it is the graduated and the recommended monitoring tool for cloud-native environments—if you look at it and the way it scales, actually, it’s a single binary solution, which is great because it’s really easy to get started. You deploy a single instance, and you have ingestion, storage, and visibility, and dashboarding, and alerting, all packaged together into one solution, and that’s definitely great. And it can scale by itself to a certain point and is definitely the recommended starting point, but as you really start to grow your business, increase your cluster sizes, increase the number of applications you have, actually isn’t a great fit for horizontal scale. So, by default, there isn’t really a high availability and horizontal scale built into Prometheus by default, and that’s why other projects in the CNCF, such as Cortex and Thanos were created to solve some of these problems.

So, we looked at the problem in a similar fashion, and when we created M3, the open-source metrics platform that came out of Uber, it was also approaching it from this different perspective where we built it to be horizontally scalable, and highly reliable from the beginning, but yet, we don’t really want it to be a, let’s say, competing project with Prometheus. So, it is actually something that works in tandem with Prometheus, in the sense that it can ingest Prometheus metrics and you can issue Prometheus query language queries against it, and it will fulfill those. But it is really built for a more scalable environment. And I would say that once a company starts to grow and they run into some of these pain points and these pain points are surrounding how reliable a Prometheus instance is, how you can scale it up beyond just giving it more resources on the VM that it runs on, vertical scale runs out at a certain point. Those are some of the pain points that a lot of companies do run into and need to solve eventually. And there are various solutions out there, both in open-source and in the commercial world, that are designed to solve those pain points. M3 being one of the open-source ones and, of course, Chronosphere being one of the commercial ones.

Corey: This episode is sponsored in part by Salesforce. Salesforce invites you to “Salesforce and AWS: Whats Ahead for Architects, Admins and Developers” on June 24th at 10AM, Pacific Time. Its a virtual event where you’ll get a first look at the latest innovations of the Salesforce and AWS partnership, and have an opportunity to have your questions answered. Plus you’ll get to enjoy an exclusive performance from Grammy Award winning artist The Roots! I think they’re talking about a band, not people with super user access to a system. Registration is free at salesforce.com/whatsahead.

Corey: Now, you’ve also gone ahead and more or less dangled raw meat in front of a tiger in some respects here because one of the things that you wind up saying on your site of why people would go with Chronosphere is, “Ah, this doesn’t allow for bill spike overages as far as what the Chronosphere bill is.” And that’s awesome. I love predictable pricing. It’s sort of the antithesis of cloud bills. But there is the counterargument, too, which is with many approaches to monitoring, I don’t actually care what my monitoring vendor is going to charge me because they wind up costing me five times more, just in terms of CloudWatch charges. How does your billing work? And how do you avoid causing problems for me on the AWS side, or other cloud provider? I mean, again, GCP and Azure are not immune from this.

Martin: So, if you look at the built-in solutions by the cloud providers, a lot of those metrics and monitoring you get from those like
CloudWatch or Stackdriver, a lot of it you get included for free with your AWS bill already. It’s only if you want additional data and additional retention, do you choose to pay more there. So, I think a lot of companies do use those solutions for the default set of monitoring that they want, especially for the AWS services, but generally, a lot of companies have custom monitoring requirements outside of that in the application tier, or even more detailed monitoring in the infrastructure that is required, especially if you think about Kubernetes.

Corey: Oh, yeah. And then I see people using CloudWatch as basically a monitoring, or metric, or log router, which at its price point, don’t
do that. [laugh]. It doesn’t end well for anyone involved.

Martin: A hundred percent. So, our solution and our approach is a little bit different. So, it doesn’t actually go through CloudWatch or any of these other inbuilt cloud-hosted solutions as a router because, to your point, there’s a lot of cost there as well. It actually goes and collects the data from the infrastructure tier or the applications. And what we have found is that not only does the bill for monitoring climb exponentially—and not just as you grow; especially as you shift towards cloud-native architecture—our very first take of solving that problem is to make the backend a lot more efficient than before so it just is cheaper overall.

And we approached it that way at Uber, and we had great results there. So, when we created an—originally before M3, 8% of Uber’s infrastructure bill was spent on monitoring all the infrastructure and the application. And by the time we were done with M3, the cost was a little over 1%. So, the very first solution was just make it more efficient. And that worked for a while, but what we saw is that over time, this grew again.

And there wasn’t any more efficiency, we could crank out of the backend storage system. There’s only so much optimization you can do to the compression algorithms in the backend and how much you can get there. So, what we realized the problem shifted towards was not, can we store this data more efficiently because we’re already reaching limitations there, and what we noticed was more towards getting the users of this data—so individual developers themselves—to start to understand what data is being produced, how they’re using it, whether it’s even useful, and then taking control from that perspective. And this is not a problem isolated to the SRE team or the observability team anymore; if you think about modern DevOps practices, every developer needs to take control of monitoring their own applications. So, this responsibility is really in the hands of the developers.

And the way we approached this from a Chronosphere perspective is really in four steps. The first one is that we have cost accounting so that every developer, and every team, and the central observability team know how much data is being produced. Because it’s actually a hard thing to measure, especially in the monitoring world. It’s—

Corey: Oh, yeah. Even AWS bills get this wrong. Like if you’re sending data between one availability zone to another in the same region, it charges a penny to leave an AZ and a penny to enter an AZ in that scenario. And the way that they reflect this on the bill is they double it. So, if you’re sending one gigabyte across AZ link in a month, you’ll see two gigabytes on the bill and that’s how it’s reflected. And that is just a glimpse of the monstrosity that is the AWS billing system. But yeah, exposing that to folks so they can understand how much data their application is spitting off? Forget it. That never happens.

Martin: Right. Right. And it’s not even exposing it to the company as a whole, it’s to each use case, to each developer so they know how much data they are producing themselves. They know how much of the bill is being consumed. And then the second step in that is to put up bumper lanes to that so that once you hit the limit, you don’t just get a surprise bill at the end of the month.

When each developer hits that limit, they rate-limit themselves and they only impact their own data; there is no impact to the other developers or to the other teams, or to the rest of the company. So, we found that those two were necessary initial steps, and then there were additional steps beyond that, to help deal with this problem.

Corey: So, in order for this to work within a multi-day lag, in some cases, it’s a near certainty that you’re looking at what is happening and the expense that is being incurred in real-time, not waiting for it to pass its way through the AWS billing system and then do some tag attribution back.

Martin: A hundred percent. It’s in real-time for the stream of data. And as I mentioned earlier, for the monitoring data we are collecting, it goes straight from the customer environment to our backend so we’re not waiting for it to be routed through the cloud providers because, rightly so, there is a multi-day or multi-hour delay there. So, as the data is coming straight to our backend, we are actively in real-time measuring that and cost accounting it to each individual team. And in real-time, if the usage goes above what is allocated, will actually limit that particular team or that particular developer, and prevent them by default from using more. And with that mechanism, you can imagine that’s how the bill is controlled and controlled in real-time.

Corey: So, help me understand, on some level; is your architecture then agent-based? Is it a library that gets included in the application code itself? All of the above and more? Something else entirely? Or is this just such a ridiculous question that you can’t believe that no one has ever asked it before?

Martin: No, it’s a great question, Corey, and would love to give some more insight there. So, it is an agent that runs in the customer environment because it does need to be something there that goes and collects all the data we’re interested in to send it to the backend. This agent is unlike a lot of APM agents out there where it does, sort of, introspection, things like that. We really believe in the power of the open-source community, and in particular, open-source standards like the Prometheus format for metrics. So, what this agent does is it actually goes and discovers Prometheus endpoints exposed by the infrastructure and applications, and scrapes those endpoints to collect the monitoring data to send to the backend.

And that is the only piece of software that runs in our customer environments. And then from that point on, all of the data is in our backend, and that’s where we go and process it and get visibility into the end-users as well as store it and make it available for alerting and dashboarding purposes as well.

Corey: So, when did you found Chronosphere? I know that you folks recently raised a Series B—congratulations on that, by the way; that generally means, at least if I understand the VC world correctly, that you’ve established product-market fit and now we’re talking about let’s scale this thing. My experience in startup land was, “Oh, we’ve raised a Series B, that means it’s probably time to bring in the first DevOps hire.” And that was invariably me, and I wound up screaming and freaking out for three months, and then things were better. So, that was my exposure to Series B.

But it seems like, given what you do, you probably had a few SRE folks kicking around, even on the product team because everything you’re saying so far absolutely resonates with the experiences someone who has run these large-scale things in production. No big surprise there. Is that where you are? I mean, how long have you been around?

Martin: Yeah, so we’ve been around for a couple of years thus far—so still a relatively new company, for sure. A lot of the core team were the team that both built the underlying technology and also ran it in production the many years at Uber, and that team is now here at Chronosphere. So, you can imagine from the very beginning, we had DevOps and SREs running this hosted platform for us. And it’s the folks that actually built the technology and ran it for years running it again, outside of Uber now. And then to your first question, yes, we did establish fairly early on, and I think that is also because we could leverage a lot of the technology that we had built at Uber, and it sort of gave us a boost to have a product ready for the market much faster.

And what we’re seeing in the industry right now is the adoption of cloud-native is so fast that it’s sort of accelerating a need of a new monitoring solution that historical solutions, perhaps, cannot handle a lot of the use cases there. It’s a new architecture, it’s a new technology stack, and we have the solution purpose-built for that particular stack. So, we are seeing fairly fast acceleration and adoption of our product right now.

Corey: One problem that an awful lot of monitoring slash observability companies have gotten into in the last few years—at least it feels this way, and maybe I’m wildly incorrect—is that it seems that the target market is the Ubers of the world, the hyperscalers where once you’re at that scale, then you need a tool like this, but if you’re just building a standard three-tier web app, oh, you’re nowhere near that level of scale. And the problem with go-to-market in those stories inherently seems that by the time you are a hyperscalers, you have already built a somewhat significant observability apparatus, otherwise you would not have survived or stayed up long enough to become a hyperscalers. How do you find that the on-ramp looks? I mean, your website does talk about, “When you outgrow Prometheus.” Is there a certain point of scale that customers should be at before they start looking at things like Chronosphere?

Martin: I think if you think about the companies that are born in the cloud today and how quickly they are running and they are iterating their technology stack, monitoring is so critical to that. It’s the real-time visibility of these changes that are going out multiple times a day is critical to the success and growth of a lot of new companies. And because of how critical that piece is, we’re finding that you don’t have to be a giant hyperscalers like Uber to need technology like this. And as you rightly pointed out, you need technology like this as you scale up. And what we’re finding is that while a lot of large tech companies can invest a lot of resources into hiring these teams and building out custom software themselves, generally, it’s not a great investment on their behalf because those are not companies that are selling monitoring technology as their core business.

So generally, what we find is that it is better for companies to perhaps outsource or purchase, or at least use open-source solutions to solve some of these problems rather than custom-build in-house. And we’re finding that earlier and earlier on in a company’s lifecycle, they’re needing technology like this.

Corey: Part of the problem I always ran into was—again, I come from the old world of grumpy Unix sysadmins—for me, using Nagios was my approach to monitoring. And that’s great when you have a persistent stateful, single node or a couple of single nodes. And then you outgrow it because well, now everything’s ephemeral and by the time you realize that there’s an outage or an issue with a container, the container hasn’t existed for 20 minutes. And you better have good telemetry into what’s going on and how your application behaves, especially at scale because at that point, edge cases, one-in-a-million events happen multiple times a second, depending upon scale, and that’s a different way of thinking. I’ve been somewhat fortunate in that, in my experience at least, I’ve not usually had to go through
those transformative leaps.

I’ve worked with Prometheus, I’ve worked with Nagios, but never in the same shop. That’s the joy of being a consultant. You go into one environment, you see what they’re doing and you take notes on what works and what doesn’t, you move on to the next one. And it’s clear that there’s a definite defined benefit to approaching observability in a more modern way. But I despair the idea of trying to go from one to the other. And maybe that just speaks to a lack of vision for me.

Martin: No, I don’t think that’s the case at all, Corey. I think we are seeing a lot of companies do this transition. I don’t think a lot of companies go and ditch everything that they’ve done. And things that they put years of investment into, there’s definitely a gradual migration process here. And what we’re seeing is that a lot of the newer projects, newer environments, newer efforts that have been kicked off are being monitored and observed using modern technology like Prometheus.

And then there’s also a lot of legacy systems which are still going to be around and legacy processes which are still going to be around for a very long time. It’s actually something we had to deal with that at Uber as well; we were actually using Nagios and a StatsD Graphite stack for a very long time before switching over to a more modern tag-like system like Prometheus. So—

Corey: Oh, modern Nagios. What was it, uh… that’s right, Icinga. That’s what it was.

Martin: Yes, yes. It was actually the system that we were using Uber. And I think for us, it’s not just about ditching all of that investment; it’s really about supporting this migration as well. And this is why both in the open-source technology M3, we actually support both the more legacy data types, like StatsD and the Graphite query language, as well as the more modern types like Prometheus and PromQL. And having support for both allows for a migration and a transition.

And not even a complete transition; I’m sure there will always be StatsD, Graphite data in a lot of these companies because they’re just legacy applications that nobody owns or touches anymore, and they’re just going to be lying around for a long time. So, it’s actually something that we proactively get ahead of and ensure that we can support both use cases even though we see a lot of companies and trending towards the modern technology solutions, for sure.

Corey: The last point I want to raise has always been a personal, I guess, area of focus for me. I allude to it, sometimes; I’ve done a Twitter thread or two on it, but on your website, you say something that completely resonates with my entire philosophy, and to be blunt is why in many cases, I’m down on an awful lot of vendor tooling across a wide variety of disciplines. On the open-source page on your site, near the bottom, you say, and I quote, “We want our end-users to build transferable skills that are not vendor or product-specific.”
And I don’t think I’ve ever seen a vendor come out and say something like that. Where did that come from?

Martin: Yeah. If you look at the core of the company, it is built on top of open-source technology. So, it is a very open core company here at Chronosphere, and we really believe in the power of the open-source community and in particular, perhaps not even individual projects, but industry standards and open standards. So, this is why we don’t have a proprietary protocol, or proprietary agent, or proprietary query language in our product because we truly believe in allowing our end-users to build these transferable skills and industry-standard skills. And right now that is using Prometheus as the client library for monitoring and PromQL as the query language.

And I think it’s not just a transferable skill that you can bring with you across multiple companies, it is also the power of that broader community. So, you can imagine now that there is a lot more sharing of, “Hey, I am monitoring, for example, MongoDB. How should I best do that?” Those skills can be shared because the common language that they’re all speaking, the queries that everybody is sharing with each other, the dashboards everybody is sharing with each other, are all, sort of, open-source standards now. And we really believe in the power that and we really do everything we can to promote that. And that is why in our product, there isn’t any proprietary query language, or definitions of dashboarding, or [learning 00:35:39] or anything like that. So yeah, it is definitely just a core tenant of the company, I would say.

Corey: It’s really something that I think is admirable, I’ve known too many people who wind up, I guess, stuck in various environments where the thing that they work on is an internal application to the company, and nothing else like it exists anywhere else, so if they ever want to change jobs, they effectively have a black hole on their resume for a number of years. This speaks directly to the opposite. It seems like it’s not built on a lock-in story; it’s built around actually solving problems. And I’m a little ashamed to say how refreshing that is [laugh] just based upon what that says about our industry.

Martin: Yeah, Corey. And I think what we’re seeing is actually the power of these open-source standards, let’s say. Prometheus is actually having effects on the broader industry, which I think is great for everybody. So, while a company like Chronosphere is supporting these from day one, you see how pervasive the Prometheus protocol and the query language are that actually all of these probably more traditional vendors providing proprietary protocols and proprietary query languages all actually have to have Prometheus—or not ‘have to have,’ but we’re seeing that more and more of them are having Prometheus compatibility as well. And I think that just speaks to the power of the industry, and it really benefits all of the end-users and the industry as a whole, as opposed to the vendors, which we are really happy to be supporters of.

Corey: Thank you so much for taking the time to speak with me today. If people want to learn more about what you’re up to, how you’re thinking about these things, where can they find you? And I’m going to go out on a limb and assume you’re also hiring.

Martin: We’re definitely hiring right now. And you can find us on our website at chronosphere.io or feel free to shoot me an email directly. My email is martin@chronosphere.io. Definitely massively hiring right now, and also, if you do have problems trying to monitor your cloud-native environment, please come check out our website and our product.

Corey: And we will, of course, include links to that in the [show notes 00:37:41]. Thank you so much for taking the time to speak with me
today. I really appreciate it.

Martin: Thanks a lot for having me, Corey. I really enjoyed this.

Corey: Martin Mao, CEO and co-founder of Chronosphere. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an insulting comment speculating about how long it took to convince Martin not to name the company ‘Observability Manager Chronosphere Manager.’

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About AJ

AJ Yawn is a seasoned cloud security professional that possesses over a decade of senior information security experience with extensive experience managing a wide range of cybersecurity compliance assessments (SOC 2, ISO 27001, HIPAA, etc.) for a variety of SaaS, IaaS, and PaaS providers.

AJ advises startups on cloud security and serves on the Board of Directors of the ISC2 Miami chapter as the Education Chair, he is also a Founding Board member of the National Association of Black Compliance and Risk Management professions, regularly speaks on information security podcasts, events, and he contributes blogs and articles to the information security community including publications such as CISOMag, InfosecMag, HackerNoon, and ISC2.

Before Bytechek, AJ served as a senior member of national cybersecurity professional services firm SOC-ISO-Healthcare compliance practice. AJ helped grow the practice from a 9 person team to over 100 team members serving clients all over the world. AJ also spent over five years on active duty in the United States Army, earning the rank of Captain.

AJ is relentlessly committed to learning and encouraging others around him to improve themselves. He leads by example and has earned several industry-recognized certifications, including the AWS Certified Solutions Architect-Professional, CISSP, AWS Certified Security Specialty, AWS Certified Solutions Architect-Associate, and PMP. AJ is also involved with the AWS training and certification department, volunteering with the AWS Certification Examination Subject Matter Expert program.

AJ graduated from Georgetown University with a Master of Science in Technology Management and from Florida State University with a Bachelor of Science in Social Science. While at Florida State, AJ played on the Florida State University Men's basketball team participating in back to back trips to the NCAA tournament playing under Coach Leonard Hamilton.

Links:

  • ByteChek: https://www.bytechek.com/
  • Blog post, Everything You Need to Know About SOC 2 Trust Service Criteria CC6.0 (Logical and Physical Access Controls): https://help.bytechek.com/en/articles/4567289-everything-you-need-to-know-about-soc-2-trust-service-criteria-cc6-0-logical-and-physical-access-controls
  • LinkedIn: https://www.linkedin.com/in/ajyawn/
  • Twitter: https://twitter.com/AjYawn

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Thinkst. This is going to take a minute to explain, so bear with me. I linked against an early version of their tool, canarytokens.org in the very early days of my newsletter, and what it does is relatively simple and straightforward. It winds up embedding credentials, files, that sort of thing in various parts of your environment, wherever you want to; it gives you fake AWS API credentials, for example. And the only thing that these things do is alert you whenever someone attempts to use those things. It’s an awesome approach. I’ve used something similar for years. Check them out. But wait, there’s more. They also have an enterprise option that you should be very much aware of canary.tools. You can take a look at this, but what it does is it provides an enterprise approach to drive these things throughout your entire environment. You can get a physical device that hangs out on your network and impersonates whatever you want to. When it gets Nmap scanned, or someone attempts to log into it, or access files on it, you get instant alerts. It’s awesome. If you don’t do something like this, you’re likely to find out that you’ve gotten breached, the hard way. Take a look at this. It’s one of those few things that I look at and say, “Wow, that is an amazing idea. I love it.” That’s canarytokens.org and canary.tools. The first one is free. The second one is enterprise-y. Take a look. I’m a big fan of this. More from them in the coming weeks.

Corey: This episode is sponsored in part by our friends at Lumigo. If you’ve built anything from serverless, you know that if there’s one thing that can be said universally about these applications, it’s that it turns every outage into a murder mystery. Lumigo helps make sense of all of the various functions that wind up tying together to build applications. It offers one-click distributed tracing so you can effortlessly find and fix issues in your serverless and microservices environment. You’ve created more problems for yourself; make one of them go away. To learn more, visit lumigo.io.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’m joined this week by AJ Yawn, co-founder, and CEO of ByteChek. AJ, thanks for joining me.

AJ: Thanks for having me on, Corey. Really excited about the conversation.

Corey: So, what is ByteChek? It sounds like it’s one of those things—‘byte’ spelled as in computer term, not teeth, and ‘chek’ without a second C in it because frugality looms everywhere, and we save money where we can by sometimes not buying the extra letter or vowel. So, what is ByteChek?

AJ: Exactly. You get it. ByteChek is a cybersecurity compliance software company, built with one goal in mind: make compliance suck less. And the way that we do that is by automating the worst part of compliance, which is evidence collection and taking out a lot of the subjective nature of dealing with an audit by connecting directly where the evidence lives and focusing on security.

Corey: That sound you hear is Pandora’s Box creaking open because back before I started focusing on AWS bills, I spent a few months doing a deep dive PCI project for workloads going into AWS because previously I’ve worked in regulated industries a fair bit. I’ve been a SOC 2 control owner, I’ve gone through the PCI process multiple times, I’ve dabbled with HIPAA as a consultant. And I thought, “Huh, there might be a business need here.” And it turns out, yeah, there really is.

The problem for me is that the work made me want to die. I found it depressing; it was dull; it was a whole lot of hurry up and wait. And that didn’t align with how I approach the world, so I immediately got the hell out of there. You apparently have a better perspective on, you know, delivering things companies need and don’t need to have constant novel entertainment every 30 seconds. So, how did you start down this path, and what set you on this road?

AJ: Yeah, great question. I started in the army as a information security officer, worked in a variety of different capacities. And when I left the military—mainly because I didn’t like sleeping outside anymore—I got into cybersecurity compliance consulting. And that’s where I got first into compliance and seeing the backwards way that we would do things with old document requests and screenshots. And I enjoyed the process because there was a reason for it, like you said.

There’s a business value to this, going through this compliance assessments. So, I knew they were important, but I hated the way we were doing it. And while there, I just got exposed to so many companies that had to go through this, and I just thought there was a better way. Like, typical entrepreneur story, right? You see a problem and you’re like, “There has to be a better way than grabbing screenshots of the EC2 console.” And set out to build a product to do that, to just solve that problem that I saw on a regular basis. And I tell people all the time, I was complicit in making compliance stuff before. I was in that role and doing the things that I think sucked and not focused on security. And that’s what we’re solving here at ByteChek.

Corey: So, I’ve dabbled in it and sort of recoiled in horror. You’ve gone into this to the point where you are not only handling it for customers but in order to build software that goes in a positive direction, you have to be deeply steeped in this yourself. As you’re going down this process, what was your build process like? Were you talking to auditors? Were you talking to companies who had to deal with auditors? What aspects of the problem did you approach this from?

AJ: It’s really both aspects. And that’s where I think it’s just a really unique perspective I have because I’ve talked with a lot of auditors; I was an auditor and worked with auditors’ hand-in-hand and I understood the challenges of being an auditor, and the speed that you have to move when you’re in the consulting industry. But I also talked to a lot of customers because those were the people I dealt with on a regular basis, both from a sales perspective and from, you know, sitting there with the CTOs trying to figure out how to design a secure solution in AWS. So, I took it from the approach of you can’t automate compliance; you can’t fix the audit problem by only focusing on one side of the table, which is what currently happens where one side of the table is the client, then you get to automate evidence collection. But if the auditors can’t use that information that you’ve automated, then it’s still a bad process for both people. So, I took the approach of thinking about this from both, “How do I make this easier for auditors but also make it easier for the clients that are forced to undergo these audits?”

Corey: From a lot of perspectives, having compliance achieved, regardless of whether it’s PCI, whether it’s HIPAA, whether it’s SOC 2, et cetera, et cetera, et cetera, the reason that a companies go through it is that it’s an attestation that they are, for better or worse, doing the right things. In some cases, it’s a requirement to operate in a regulated industry. In other cases, it’s required to process credit card transactions, which is kind of every industry, and in still others, it’s an easy shorthand way of saying that we’re not complete rank amateurs at these things, so as a result, we’re going to just pass over the result of our most recent SOC 2 audit to our prospective client, and suddenly, their security folks can relax and not send over weeks of questionnaires on the security front. That means that, for some folks, this is more or less a box-checking exercise rather than an actual good-faith effort to improve processes and posture.

AJ: Correct. And I think that’s actually the problem with compliance is it’s looked at as a check-the-box exercise, and that’s why there’s no security value out of it. That’s why you can pick up a SOC 2 report for someone that’s hosted on AWS, and you don’t see any mention of S3 buckets. You can do a ctrl+F, and you literally don’t see anything in a security evaluation about S3 buckets, which is just insane if you know anything about security on AWS. And I think it’s because of what you just described, Corey; they’re often asked to do this by a regulator, or by a customer, or by a vendor, and the result is, “Hurry up and get this report so that we can close this deal,”—or we can get to the next level with this customer, or with this investor, whatever it may be—instead of, let’s go through this, let’s have an auditor come in and look at our environment to improve it, to improve this security, which is where I hope the industry can get to because audits aren’t going anywhere; people are going to continue to do them and spend thousands of dollars on them, so there should be some security value out of them, in my opinion.

Corey: I love using encrypting data at rest as an example of things that make varying amounts of sense because, sure, on your company laptops, if someone steals an employee’s laptop from a coffee shop, or from the back of their car one night, yeah, you kind of want the exposure to the company to be limited to replacing the hardware. I mean, even here at The Duckbill Group, where we are not regulated, we’ve gone through no formal audits, we do have controls in place to ensure that all company laptops have disk encryption turned on. It makes sense from that perspective. And in the data center, it was also important because there were a few notable heists where someone either improperly disposed drives and corporate data wound up on eBay or someone in one notable instance drove a truck through the side of the data center wall, pulled a rack into the bed of the truck and took off, which is kind of impressive [laugh] no matter how you slice it. But in the context of a hyperscale cloud provider like AWS, you’re not going to be able to break into their data centers, steal a drive—and of course, it has to be the right collection of drives and the right machines—and then find out how to wind up reassembling that data later.

It’s just not a viable attack strategy. Now, you can spend days arguing with auditors around something like that, or you can check the box ‘encrypt at rest’ and move on. And very often, that is the better path. I’m not going to argue with auditors about that. I’m going to bend the knee, check the box, and get back to doing the business thing that I care about. That is a reasonable approach, is it not?

AJ: It is, but I think that’s the fault of the auditor because good security requires context. You can’t just apply a standard set of controls to every organization, as you’re describing, where I would much rather the auditor care about, “Are there any public S3 buckets? What are the security group situation like on that account? How are they managing their users? How are they storing credentials there in the cloud environment as well?

Are they using multiple accounts?” So, many other things to care about other than protecting whether or not someone will be able to pull off the heist of the [laugh] 21st century. So, I think from a customer perspective, it’s the right model: don’t waste time arguing points with your auditors, but on the flip side, find an auditor that has more technical knowledge that can understand context, because security work requires good context and audits require context. And that’s the problem with audits now; we’re using one framework or several frameworks to apply to every organization. And I’ve been in the consulting space, like you, Corey, for a while. I have not seen the same environment in any customers. Every customer is different. Every customer has a different setup, so it doesn’t make sense to say every control should apply to every company.

Corey: And it feels on some level like you wind up getting staff accustomed to treating it as a box-checking exercise. “Right, it’s dumb that we wind up having to encrypt S3 buckets, but it’s for the audit to just check the box and move on.” So, people do it, then they move on to the next item, which is, “Okay, great. Are there any public S3 buckets?” And they treat it with the same, “Yeah, whatever. It’s for the audit,” box-checking approach? No, no, that one’s actually serious. You should invest significant effort and time into making sure that it’s right.

AJ: Exactly. Exactly. And that’s where the value of a true compliance assessment that is focused on security comes into play because it’s no longer about checking the box, it’s like, “Hey, there’s a weakness here. A weakness that you probably should have identified. So, let’s go fix the weakness, but let’s talk about your process to find those weaknesses and then hopefully use some automation to remediate them.”

Because a lot of the issues in the cloud you can trace back to why was there not a control in place to prevent this or detect this? And it’s sad that compliance assessments are not the thing that can catch those, that are not the other safeguard in place to identify those. And it’s because we are treating the entire thing like a check-the-box exercise and not pulling out those items that really matter, and that’s just focusing on security. Which is ultimately what these compliance reports are proving: customers are asking for these reports because they want to know if their data is going to be secure. And that’s what the report is supposed to do, but on the flip side, everyone knows the organization may not be taking it that serious, and they may be treating it like a check-the-box exercise.

Corey: So, while I have you here, we’ll divert for a minute because I’m legitimately curious about this one. At a scale of legitimate security concern to, “This is a check-the-box exercise,” where do things like rotating passwords every 60 days or rotating IAM credentials every 90 days fall?

AJ: I think it again depends on the organization. I don’t think that you need to rotate passwords regularly, personally. I don’t know how strong of a control that is if people are doing that, because they’re just going to start to make things up that are easy—

Corey: Put the number at the end and increment by one every time. Great. Good work.

AJ: Yep. So, I think again, it just depends on your organization and what the organization is doing. If you’re talking about managing IAM access keys and rotating those, are your engineers even using the CLI? Are they using their access keys? Because if they’re not, what are you rotating?

You’re just rotating [laugh] stale keys that have never been used. Or if you don’t even have any IAM users, maybe you’re using SSO and they’re all using Okta or something else and they’re using an IAM role to come in there. So, it’s just—again, it’s context. And I think the problem is, a lot of folks don’t understand AWS or they don’t understand the cloud. And when I say, folks, I mean auditors.

They don’t understand that, so they’re just going to ask for everything. “Did you rotate your passwords? Did you do this? Did you do that?” And it may not even make sense for you based off of your environment, but again, is it worth the fight with the auditor, or do you just give them whatever they want and so you can go about your way, whether or not it’s a legit security concern?

Corey: Yeah. At some point, it’s not worth fighting with auditors, but if you find yourself wanting to fight the auditor all the time, at some level, you start to really resent the auditor that you have. To put that slightly more succinctly, how do you deal with non-technical auditors who don’t understand your environment—what they’re looking at—without strangling them?

AJ: Great question. I think it goes back to before you hire your auditor. Oftentimes, in the sales process, there’s questions around, “Who’s come from the Big Four on your staff?” Or, “What control frameworks do you all specialize in?” Or, “How long will this take? How much will it cost?” But there’s very rarely any questions of, “Who on your staff knows AWS?”

And it’s similar to going to the doctor: you wouldn't go to an eye doctor to get foot surgery. So, you shouldn’t go to an auditor who has never seen AWS, that doesn’t know what EC2 is, to evaluate your AWS environment. So, I think organizations have to start asking the right questions during the sales process. And it’s not about price or time or anything like that when you’re assessing who you’re going to work with from an auditing firm. It’s, are they qualified to actually evaluate the threats facing your organization so that you don’t get asked the stupid question.

If you’re hosted on AWS, you shouldn’t be getting asked where are your firewall configurations. They should understand what security groups are and how they work. So, there’s just a level of knowledge that should be expected from the organization side. And I would say, if you’re working with a current auditor that you’re having those issues with, continue to ask the hard questions. Auditors that are not technical—I have a blog post on our website, and it says this is the section your auditors are the most scared of, and it’s the logical access section of your SOC 2 report.

And auditors that are not technical run away from that section. So, just keep asking the hard questions, and they’ll either have to get the knowledge or they realize they’re not qualified to do the assessment and the marriage will split up kind of naturally from there. But I think it goes back to the initial process of getting your auditor. Don’t worry about cost or time, worry about their technical skills and if they’re qualified to assess your environment.

Corey: And in 2021, that’s a very different story than it was the first few times I encountered auditors discovering the new era. At a startup, the auditor shows up. “Great, how do we get access to your Active Directory?” “Yeah, we don’t have one of those.” “Okay, how do we get on the internet here?” “Oh, here’s the wireless password.” “Wait, there’s not a separate guest network?” “That’s right.” “Well, now I have privileged access because I’m on your network.”

It’s like, “Technically, that’s true because if you weren’t on this network, you wouldn’t be able to print to that printer over there in the corner. But that’s the only thing that it lets you do.” Everything else is identity-based, not IP address allow listing, so instead, it’s purely just convenience to get the internet; you’re about as privileged on this network as you would be at a Starbucks half a world away. And they look at you like you’re an idiot. And that should have been the early warning sign that this was not going to be a typical audit conversation. Now, though in 2021, it feels like it’s time to find a new auditor.

AJ: Exactly. Yeah. Especially because organizations—unfortunately, last year security budgets were some of the things that were first cut when budgets were cut due to the global pandemic, S0—

Corey: Well, I’m sure that’ll have no lasting repercussions.

AJ: Right. [laugh]. That’s always a great decision. So compliance, that means compliance budgets have been significantly slashed because that’s the first thing that gets cut is spending money on compliance activities. So, the cheaper option, oftentimes, is going to mean even less technical resources.

Which is why I don’t think manual audits, human audits are going to be a thing moving forward. I think companies are realizing that it doesn’t make sense to go through a process, hire an auditor who’s selling you on all this technical expertise, and then the staff that’s showing up and assigned to your project has never seen inside the AWS console and truly doesn’t even know what the cloud is. They think that iCloud on their phone is the only cloud that they’re familiar with. And that’s what happens; organizations are sold that they’re going to get cybersecurity technical experts from these human auditors and then somebody shows up without that experience or expertise. So, you have to start to rely on tools, rely on technologies, and that can be native technologies in the cloud or third-party tools.

But I don’t think you can actually do a good audit in the cloud manually anyways, no matter how technical you are. I know a lot about AWS but I still couldn’t do a great audit by myself in the cloud because auditing is time-based, you bill by the hour and it doesn’t make sense for me to do all of those manual things that tools and technologies out there exist to do for us.

Corey: So, you started a software company aimed at this problem, not a auditing firm and not a consulting company. How are you solving this via the magic of writing code?

AJ: It’s just connecting directly where the evidence lives. So, for AWS, I actually tried to do this in a non-software way prior, when I was just a typical auditor, and I was just asking our clients to provision us cross-account access to go in their environment with some security permissions to get evidence directly. And that didn’t pass the sniff test at my consulting firm, even though some of the clients were open to it. But we built software to go out to the tools where the evidence directly lives and continuously assess the environment. So, that’s AWS, that’s GitHub, that Jira, that’s all of the different tools where you normally collect this evidence, and instead of having to prove to auditors in a very manual fashion, by grabbing screenshots, you just simply connect using APIs to get the evidence directly from the source, which is more technically accurate.

The way that auditing has been done in the past is using sampling methodologies and all these other outdated things, but that doesn’t really assess if all of your data stores are configured in the right way; if you’re actually backing up your data. It’s me randomly picking one and saying, “Yes, you’re good to go.” So, we connect directly where the evidence lives and hopefully get to a point where when you get a SOC 2 report, you know that a tool checked it. So, you know that the tool went out and looked at every single data store, or they went out and looked at every single EC2 instance, or security group, whatever it may be, and it wasn’t dependent on how the auditor felt that day.

Corey: This episode is sponsored in part by ChaosSearch. As basically everyone knows, trying to do log analytics at scale with an ELK stack is expensive, unstable, time-sucking, demeaning, and just basically all-around horrible. So why are you still doing it—or even thinking about it—when there’s ChaosSearch? ChaosSearch is a fully managed scalable log analysis service that lets you add new workloads in minutes, and easily retain weeks, months, or years of data. With ChaosSearch you store, connect, and analyze and you’re done. The data lives and stays within your S3 buckets, which means no managing servers, no data movement, and you can save up to 80 percent versus running an ELK stack the old-fashioned way. It’s why companies like Equifax, HubSpot, Klarna, Alert Logic, and many more have all turned to ChaosSearch. So if you’re tired of your ELK stacks falling over before it suffers, or of having your log analytics data retention squeezed by the cost, then try ChaosSearch today and tell them I sent you. To learn more, visit chaossearch.io.

Corey: That sounds like it is almost too good to be true. And at first, my immediate response is, “This is amazing,” followed immediately by that’s transitioning into anger, that, “Why isn’t this a native thing that everyone offers?” I mean, to that end, AWS announced ‘Audit Manager’ recently, which I haven’t had the opportunity to dive into in any deep sense yet, because it’s still brand new, and they decided to release it alongside 15,000 other things, but does that start getting a little bit closer to something companies need? Or is it a typical day-one first release of an Amazon service where, “Well, at least we know the direction you’re heading in. We’ll check back in two years.”

AJ: Exactly. It’s the day-one Amazon service release where, “Okay. AWS is getting into the audit space. That’s good to know.” But right now, at its core, that AWS service, it’s just not usable for audits, for several reasons.

One, auditors cannot read the outputs of the information from Audit Manager. And it goes back to the earlier point where you can’t automate compliance, you can’t fix compliance if the auditors can’t use the information because then they’re going to go back to asking dumb questions and dumb evidence requests if they don’t understand the information coming out of it. And it’s just because of the output right now is a dump of JSON, essentially, in a Word document, for some strange reason.

Corey: Okay, that is the perfect example right there of two worlds colliding. It’s like, “Well, we’re going to put JSON out of it because that’s the language developers speak. Well, what do auditors prefer?” “I don’t know, Microsoft Word?” “Okay, sounds good.” Even Microsoft Excel is a better answer than [laugh] that. And that is just… okay, that is just Looney Tunes awful.

AJ: Yep. Yeah, exactly. And that’s one problem. The other problem is, Audit Manager requires a compliance manager. If we think about that tool, a developer is not going to use Audit Manager; it’s going to be somebody responsible for compliance.

It requires them to go manually select every service that their company is using. A compliance manager, one, doesn’t even know what the services are; they have no clue what some of these services are, two, how are they going to know if you’re using Lambda randomly somewhere or, or a Systems Manager randomly somewhere, or Elastic Beanstalk’s in one account or one region. Config here, config—they have to just go through and manually—and I’m like, “Well, that doesn’t make any sense because AWS knows what services you’re using. Why not just already have those selected and you pull those in scope?” So, the chances of something being excluded are extremely high because it’s a really manual process for users to decide what are they actually assessing.

And then lastly, the frameworks need a lot of work. Auditing is complex because their standards or regulations and all of that, and there’s just a gap between what AWS has listed as a service that addresses a particular control that—there was a few times where I looked at Audit Manager and I had no clue what they were mapping to and why they’re mapping. So, it’s a typical day-one service; it has some gaps, but I like the direction it’s going. I like the idea that an organization can go into their AWS console, hit to a dashboard, and say, “Am I meeting SOC 2?” Or“ am I meeting PCI?” I feel like this is a long time coming. I think you probably could have done it with Security Hub with less automation; you have to do some manual uploads there, but the long answer to say it has a long way to go there, Corey.

Corey: I heard a couple of horror stories of, “Oh, my god, it’s charging me $300 a day and I can’t turn it off,” when it first launched. I assume that’s been fixed by now because the screaming has stopped. I have to assume it was. But it was gnarly and surprising people with bills. And surprising people with things labeled ‘audit’ is never a great plan.

AJ: Right. Yeah, the pricing was a little ridiculous as well. And I didn’t really understand the pricing model. But that’s typical of a new AWS service, I never really understand. That’s why I’m glad that you exist because I’m always confused at first about why things cost so much, but then if you give it some time, it starts to make a little bit more sense.

Corey: Exactly. The first time you see a new pricing dimension, it’s novel and exciting and more than a little scary, and you dive into it. But then it’s just pattern recognition. It’s, “Oh, it’s one of these things again. Great.” It’s why it lends itself to a consulting story.

So, you were in the army for a while. And as you mentioned, you got tired of sleeping on the ground, so you went into corporate life. And you were at a national cybersecurity professional services firm for a while. What was it that finally made you, I guess, snap for lack of a better term and, “I’m going to start my own thing?” Because in my case, it was, “Well, okay. I get fired an awful lot. Maybe I should try setting out my own shingle because I really don’t have another great option.” I don’t get the sense, given your resume and pedigree, that that was your situation?

AJ: Not quite. I surprisingly, don’t do well with authority. So, a little bit I like to challenge things and question the norm often, which got me in trouble in the military, definitely got me in trouble in corporate life. But for me it was, I wanted to change; I wanted to innovate. I just kept seeing that there was a problem with what we were doing and how we were doing it, and I didn’t feel like I had the ability to innovate.

Innovating in a professional services firm is updating a Google Sheet, or adding a new Google Form and sending that off to a client. That’s not really the innovation that I was looking to do. And I realized that if I wanted to create something that was going to solve this problem, I could go join one of the many startups out there that are out there trying to solve this problem, or I could just try to go do it myself and leverage my experience. And two worlds collided as far as timing and opportunity where I financially was in a position to take a chance like this, and I had the knowledge that I finally think I needed to feel comfortable going out on my own and just made the decision. I’m a pretty decisive person, and I decided that I was going to do it and just went with it.

And despite going about this during the global pandemic, which presented its own challenges last year, getting this off the ground. But it was really—I collected a bunch of knowledge. I realized, maybe, two and a half years ago, actually, that I wanted to start my own business in this space, but I didn’t know what I wanted to do just yet. I knew I wanted to do software, I didn’t know how I wanted to do it, I didn’t know how I was going to make it work. But I just decided to take my time and learn as much as I can.

And once I felt like I acquired enough knowledge and there was really nothing else I could gain from not doing this on my own, and I knew I wasn’t going to go join a startup to join them on this journey, it was a no-brainer just to pull the trigger.

Corey: It seems to have worked out for you. I’m starting to see you folks crop up from time-to-time, things seem to be going well. How big are you?

AJ: Yeah, we’re doing well. We have a team of seven of us now, which is crazy to think about because I remember when it was just me and my co-founder staring at each other on Zoom every day and wondering if they’re ever going to be anybody else on these [laugh] calls and talking to us. But it’s going really well. We have early customers that are happy and that’s all that I can ask for and they’re not just happy silently; they’re being really public about being happy about the platform, and about the process. And just working with people that get it and we’re building a lot of momentum.

I’m having a lot of fun on LinkedIn and doing a lot of marketing efforts there as well. So, it’s been going well; it’s been actually going better than expected, surprisingly, which I don’t know, I’m a pretty optimistic entrepreneur and I thought things will go well, but it’s much better than expected, which means I’m sleeping a lot less than I expected, as well.

Corey: Yeah, at some point, when you find yourself on the startup train, it’s one of those, “Oh, yeah. That’s right. My health is in the gutter, my relationships are starting to implode around me.” Balance is key. And I think that that is something that we don’t talk about enough in this world.

There are periodically horrible tweets about how you should wind up focusing on your company, it should be the all-consuming thing that drives you at all hours of the day. And you check and, “Oh, who made that observation on Twitter? Oh, it’s a VC.” And then you investigate the VC and huh, “You should only have one serious bet, it should be your all-consuming passion” says someone who’s invested in a wide variety of different companies all at the same time, in the hopes that one of them succeeds. Huh.

Almost like this person isn’t taking the advice they’re giving themselves and is incentivized to give that advice to others. Huh, how about that? And I know that’s a cynical take, but it continues to annoy me when I see it. Where do you stand on the balance side of the equation?

AJ: Yeah, I think balance is key. I work a lot, but I rest a lot too. And I spend—I really hold my mornings as my kind of sacred place, and I spend my mornings meditating, doing yoga, working out, and really just giving back to myself. And I encourage my team to do the same. And we don’t just encourage it from just a, “Hey, you guys should do this,” but I talk to my team a lot about not taking ourselves too seriously.

It’s our number one core value. It’s why our slogan is ‘make compliance suck less’ because it’s really my military background. We’re not being shot at; we’re sleeping at home every night. And while compliance and cybersecurity, it’s really important, and we’re protecting really important things, it’s not that serious to go all-in and to not have balance, and not to take time off not to relax. I mean, a part of what we do at ByteChek is we have a 10% rule, which means 10% of the week, I encourage my team to spend it on themselves, whether that’s doing meditation, going to take a nap.

And these are work hours; you know, go out, play golf. I spent my 10% this morning playing golf during work hours. And I encourage all my team, every single week, spend four hours dedicated to yourself because there’s nothing that we will be able to do as a company without the people here being correct and being mentally okay. And that’s something that I learned a long time ago in the military. You spend a year away from home and you start to really realize what’s important.

And it’s not your job. And that’s the thing. We hire a lot of veterans here because of my veteran background, and I tell all the vets that come here when you’re in the military, your job, your rank, and your day-to-day work is your identity. It’s who you are. You’re a Marine or you’re a Soldier, or you’re a Sailor; you’re an Airman if that’s a bad choice that you made. Sorry for my Air Force guys.

Corey: Well, now there’s a Spaceman story as well, I’m told. But I don’t know if they call them spacemen or not, but remember, there’s a new branch to consider. And we can’t forget the Coast Guard either.

AJ: If they don’t call themselves Spacemen, that is their name from now on. We just made it, today. If I ever meet somebody in the Space Force, [laugh] I’m calling them the Spacemen. That is amazing. But I tell our interns that we bring from the military, you have to strip that away.

You have to become an individual because ByteChek is not your identity. And it won’t be your identity. And ByteChek’s not my identity. It’s something that I’m doing, and I am optimistic that it’s going to work out and I really hope that it does. But if it doesn’t, I’m going to be all right; my team is going to be all right and we’re going to all continue to go on.

And we just try to live that out every day because there’s so many more important things going on in this world other than cybersecurity compliance, so we really shouldn’t take ourselves too seriously. And that advice of just grinding it out, and that should be your only focus, that’s only a recipe for disaster, in my opinion.

Corey: AJ, thank you so much for taking the time to speak with me. If people want to hear more about what you have to say, where can they find you?

AJ: They can find me on LinkedIn. That’s my one spot that I’m currently on. I am going to pop on Twitter here pretty soon. I don’t know when, but probably in the next few weeks or so. I’ve been encouraged by a lot of folks to join the tech community on Twitter, so I’ll be there soon.

But right now they can find me on LinkedIn. I give four hours back a week to mentoring, so if you hear this and you want to reach out, you want to chat with me, send me a message and I will send you a link to find time on my calendar to meet. I spend four hours every Friday mentoring, so I’m open to chat and help anyone. And when you see me on LinkedIn, you’ll see me talking about diversity in cybersecurity because I think really the only way you can solve a cybersecurity skills shortage is by hiring more diverse individuals. So, come find me there, engage with me, talk to me; I’m a very open person and I like to meet new people. And that’s where you can find me.

Corey: Excellent. And we’ll of course throw a link to your LinkedIn profile in the [show notes 00:29:44]. Thank you so much for taking the time to speak with me. It's really appreciated.

AJ: Yeah, definitely. Thank you, Corey. This is kind of like a dream come true to be on this podcast that I’ve listened to a lot and talk about something that I’m passionate about. So, thanks for the opportunity.

Corey: AJ Yawn, CEO and co-founder of ByteChek. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you hated this podcast, please leave a five-
star review on your podcast platform of choice along with a comment that’s embedded inside of a Word document.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Mike

Beside his duties as The Duckbill Group’s CEO, Mike is the author of O’Reilly’s Practical Monitoring, and previously wrote the Monitoring Weekly newsletter and hosted the Real World DevOps podcast. He was previously a DevOps Engineer for companies such as Taos Consulting, Peak Hosting, Oak Ridge National Laboratory, and many more. Mike is originally from Knoxville, TN (Go Vols!) and currently resides in Portland, OR.Links:

  • Software Engineering Daily podcast: https://softwareengineeringdaily.com/category/all-episodes/exclusive-content/Podcast/
  • Duckbillgroup.com: https://duckbillgroup.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Thinkst. This is going to take a minute to explain, so bear with me. I linked against an early version of their tool, canarytokens.org in the very early days of my newsletter, and what it does is relatively simple and straightforward. It winds up embedding credentials, files, that sort of thing in various parts of your environment, wherever you want to; it gives you fake AWS API credentials, for example. And the only thing that these things do is alert you whenever someone attempts to use those things. It’s an awesome approach. I’ve used something similar for years. Check them out. But wait, there’s more. They also have an enterprise option that you should be very much aware of canary.tools. You can take a look at this, but what it does is it provides an enterprise approach to drive these things throughout your entire environment. You can get a physical device that hangs out on your network and impersonates whatever you want to. When it gets Nmap scanned, or someone attempts to log into it, or access files on it, you get instant alerts. It’s awesome. If you don’t do something like this, you’re likely to find out that you’ve gotten breached, the hard way. Take a look at this. It’s one of those few things that I look at and say, “Wow, that is an amazing idea. I love it.” That’s canarytokens.org and canary.tools. The first one is free. The second one is enterprise-y. Take a look. I’m a big fan of this. More from them in the coming weeks.

Corey: This episode is sponsored in part by our friends at Lumigo. If you’ve built anything from serverless, you know that if there’s one thing that can be said universally about these applications, it’s that it turns every outage into a murder mystery. Lumigo helps make sense of all of the various functions that wind up tying together to build applications. It offers one-click distributed tracing so you can effortlessly find and fix issues in your serverless and microservices environment. You’ve created more problems for yourself; make one of
them go away. To learn more, visit lumigo.io.

Corey: This episode is sponsored in part by ChaosSearch. As basically everyone knows, trying to do log analytics at scale with an ELK stack is expensive, unstable, time-sucking, demeaning, and just basically all-around horrible. So why are you still doing it—or even thinking about it—when there’s ChaosSearch? ChaosSearch is a fully managed scalable log analysis service that lets you add new workloads in minutes, and easily retain weeks, months, or years of data. With ChaosSearch you store, connect, and analyze and you’re done. The data lives and stays within your S3 buckets, which means no managing servers, no data movement, and you can save up to 80 percent versus running an ELK stack the old-fashioned way. It’s why companies like Equifax, HubSpot, Klarna, Alert Logic, and many more have all turned to ChaosSearch. So if you’re tired of your ELK stacks falling over before it suffers, or of having your log analytics data retention squeezed by the cost, then try ChaosSearch today and tell them I sent you. To learn more, visit chaossearch.io.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I spent the past week guest hosting the Software Engineering Daily podcast, taking listeners over there on a tour of the clouds. Each day, I picked a different cloud and had a guest talk to me about their experiences with that cloud.

Now, there was one that we didn’t talk about, and we’re finishing up that tour here today on Screaming in the Cloud. That cloud is the obvious one, and that is your own crappy data center. And my guest is Duckbill Group’s CEO and my business partner, Mike Julian. Mike, thanks for joining me.

Mike: Hi, Corey. Thanks for having me back.

Corey: So, I frequently say that I started my career as a grumpy Unix sysadmin. Because it isn’t like there’s a second kind of Unix sysadmin you’re going to see. And you were in that same boat. You and I both have extensive experience working in data centers. And it’s easy sitting here on the tech coast of the United States—we’re each in tech hubs cities—and we look around and yeah, the customers we talked to have massive cloud presences; everything we do is in cloud, it’s easy to fall into the trap of believing that data centers are a thing of yesteryear. Are they?

Mike: [laugh]. Absolutely not. I mean, our own customers have tons of stuff in data centers. There are still companies out there like Equinix, and CoreSite, and DRC—is that them? I forget the name of them.

Corey: DRT. Digital Realty [unintelligible 00:01:54].

Mike: Digital Realty. Yeah. These are companies still making money hand over fist. People are still putting new workloads into data centers, so yeah, we’re kind of stuck with him for a while.

Corey: What’s fun is when I talked to my friends over in the data center sales part of the world, I have to admit, I went into those conversations early on with more than my own fair share of arrogance. And it was, “[laugh]. So, who are you selling to these days?” And the answer was, “Everyone, fool.” Because they are.

People at large companies with existing data center footprints are not generally doing fire sales of their data centers, and one thing that we learned about cloud bills here at The Duckbill Group is that they only ever tend to go up with time. That’s going to be the case when we start talking about data centers as well. The difference there is that it’s not just an API call away to lease more space, put in some racks, buy some servers, get them racked. So, my question for you is, if we sit here and do the Hacker News—also known as the worst website on the internet—and take their first principles approach to everything, does that mean the people who are building out data centers are somehow doing it wrong? Did they miss a transformation somewhere?

Mike: No, I don’t think they’re doing it wrong. I think there’s still a lot of value in having data centers and having that sort of skill set. I do think the future is in cloud infrastructure, though. And whether that’s a public cloud, or private cloud, or something like that, I think we’re getting increasingly away from building on top of bare metal, just because it’s so inefficient to do. So yeah, I think at some point—and I feel like we’ve been saying this for years that, “Oh, no, everyone’s missed the boat,” and here we are saying it yet again, like, “Oh, no. Everyone’s missing the boat.” You know, at some point, the boat’s going to frickin’ leave.

Corey: From my perspective, there are advantages to data centers. And we can go through those to some degree, but let’s start at the beginning. Origin stories are always useful. What’s your experience working in data centers?

Mike: [laugh]. Oh, boy. Most of my career has been in data centers. And in fact, one interesting tidbit is that, despite running a company that is built on AWS consulting, I didn’t start using AWS myself until 2015. So, as of this recording, it’s 2021 now, so that means six years ago is when I first started AWS.

And before that, it was all in data centers. So, some of my most interesting stuff in the data center world was from Oak Ridge National Lab where we had hundreds of thousands of square feet of data center floor space across, like, three floors. And it was insane, just the amount of data center stuff going on there. A whole bunch of HPC, a whole bunch of just random racks of bullshit. So, it’s pretty interesting stuff.

I think probably the most really interesting bit I’ve worked on was when I was at a now-defunct company, Peak Hosting, where we had to figure out how to spin up a data center without having anyone at the data center, as in, there was no one there to do the spin up. And that led into interesting problems, like you have multiple racks of equipment, like, thousands of servers just showed up on the loading dock. Someone’s got to rack them, but from that point, it all has to be automatic. So, how do you bootstrap entire racks of systems from nothing with no one physically there to start a bootstrap process? And that led us to build some just truly horrific stuff. And thank God that’s someone else’s problem, now. [laugh].

Corey: It makes you wonder if under the hood at all these cloud providers if they have something that’s a lot cleaner, and more efficient, and perfect, or if it’s a whole bunch of Perl tied together with bash and hope, like we always built.

Mike: You know what? I have to imagine that even at AWS at a—I know if this is true at Facebook, where they have a massive data center footprint as well—there is a lot of work that goes into the bootstrap process, and a lot of these companies are building their own hardware to facilitate making that bootstrap process easier. When you’re trying to bootstrap, say, like, Dell or HP servers, the management cards only take you so far. And a lot of the stuff that we had to do was working around bugs in the HP management cards, or the Dell DRACs.

Corey: Or you can wind up going with some budget whitebox service. I mean, Supermicro is popular, not that they’re ultra-low budget. But yeah, you can effectively build your own. And that leads down interesting paths, too. I feel like there’s a sweet spot where working on a data center and doing a build-out makes sense for certain companies.

If you’re trying to build out some proof of concept, yeah, do it in the cloud; you don’t have to wait eight weeks and spend thousands of dollars; you can prove it out right now and spend a total of something like 17 cents to figure out if it’s going to work or not. And if it does, then proceed from there, if not shut it down, and here’s a quarter; keep the change. With data centers, a lot more planning winds up being involved. And is there a cutover at which point it makes sense to evacuate from a public cloud into a physical data center?

Mike: You know, I don’t really think so. This came up on a recent Twitter Spaces that you and I did around, at what point does it really make sense to be hybrid, or to be all-in on data center? I made the argument that a large-scale HPC does not fit cloud workloads, and someone made a comment that, like, “What is large-scale?” And to me, large-scale was always, like—so Oak Ridge was—or is famous—for having supercomputing, and they have largely been in the top five supercomputers in the world for quite some time. A supercomputer of that size is tens of thousands of cores. And they’re running pretty much constant because of how expensive that stuff is to get time on. And that sort of thing would be just astronomically expensive in a cloud. But how many of those are there really?

Corey: Yeah, if you’re an AWS account manager listening to this and reaching out with, “No, that’s not true. After committed spend, we’ll wind up giving you significant discounts, and a whole bunch of credits, and jump through all these hoops.” And, yeah, I know, you’ll give me a bunch of short-term contractual stuff that’s bounded for a number of years, but there’s no guarantee that stuff gets renewed at that rate. And let’s face it. If you’re running those kinds of workloads today, and already have the staff and tooling and processes that embrace that, maybe ripping all that out in a cloud migration where there’s no clear business value derived isn’t the best plan.

Mike: Right. So, while there is a lot of large-scale HPC infrastructure that I don’t think particularly fits well on the cloud, there’s not a lot of that. There’s just not that many massive HPC deployments out there. Which means that pretty much everything below that threshold could be a candidate for cloud workloads, and probably would be much better. One of the things that I noticed at Oak Ridge was that we had a whole bunch of SGI HPC systems laying around, and 90% of the time they were idle.

And those things were not cheap when they were bought, and at the time, they’re basically worth nothing. But they were idle most of the time, but when they were needed, they’re there, and they do a great job of it. With AWS and GCP and Azure HPC offerings, that’s a pretty good fit. Just migrate that whole thing over because it’ll cost you less than buying a new one. But if I’m going to migrate Titan or Gaia from Oak Ridge over to there, yeah, some AWS rep is about to have a very nice field day. That’d just be too much money.

Corey: Well, I’d be remiss as a cloud economist if I didn’t point out that you can do this stuff super efficiently in someone else’s AWS
account.

Mike: [laugh]. Yes.

Corey: There’s also the staffing question where if you’re a large blue-chip company, you’ve been around for enough decades that you tend to have some revenue to risk, where you have existing processes and everything is existing in an on-prem environment, as much as we love to tell stories about the cloud being awesome, and the capability increase and the rest, yadda, yadda, yadda, there has to be a business case behind moving to the cloud, and it will knock some nebulous percentage off of your TCO—because lies, damned lies, and TCO analyses are sort of the way of the world—great. That’s not exciting to most strategic-level execs. At least as I see the world. Given you are one of those strategic level execs, do you agree? Am I lacking nuance here?

Mike: No, I pretty much agree. Doing a data center migration, you got to have a reason to do it. We have a lot of clients that are still running in data centers as well, and they don’t move because the math doesn’t make sense. And even when you start factoring in all the gains from productivity that they might get—and I stress the word might here—even when you factor those in, even when you factor in all the support and credits that Amazon might give them, it still doesn’t make enough sense. So, they’re still in data centers because that’s where they should be for the time because that’s what the finances say. And I’m kind of hard-pressed to disagree with them.

Corey: While we’re here playing ‘ask an exec,’ I’m going to go for another one here. It’s my belief that any cloud provider that charges a penny for professional services, or managed services, or any form of migration tooling or offering at all to their customers is missing the plot. Clearly, since they all tend to do this, I’m wrong somewhere. But I don’t see how am I wrong or are they?

Mike: Yeah, I don’t know. I’d have to think about that one some more.

Corey: It’s an interesting point because it’s—

Mike: It is.

Corey: —it’s easy to think of this as, “Oh, yeah. You should absolutely pay people to migrate in because the whole point of cloud is that it’s kind of sticky.” The biggest indicator of a big cloud bill this month is a slightly smaller one last month. And once people wind up migrating into a cloud, they tend not to leave despite all of their protestations to the contrary about multi-cloud, hybrid, et cetera, et cetera. And that becomes an interesting problem.

It becomes an area—there’s a whole bunch of vendors that are very deeply niched into that. It’s clear that the industry as a whole thinks that migrating from data centers to cloud is going to be a boom industry for the next three decades. I don’t think they’re wrong.

Mike: Yeah, I don’t think they’re wrong either. I think there’s a very long tail of companies with massive footprint staying in a data center that at some point is going to get out of a data center.

Corey: For those listeners who are fortunate enough not to have to come up the way that we did. Can you describe what a data center is like inside?

Mike: Oh, God.

Corey: What is a data center? People have these mythic ideas from television and movies, and I don’t know, maybe some Backstreet Boys music video; I don’t know where it all comes from. What is a data center like? What does it do?

Mike: I’ve been in many of these over my life, and I think they really fall into two groups. One is the one managed by a professional data center manager. And those tend to be sterile environments. Like, that’s the best way to describe it. They are white, filled with black racks. Everything is absolutely immaculate. There is no trash or other debris on the floor. Everything is just perfect. And it is freezingly cold.

Corey: Oh, yeah. So, you’re in a data center for any length of time, bring a jacket. And the soulless part of it, too, is that it’s well-lit with
fluorescent lights everywhere—

Mike: Oh yeah.

Corey: —and it’s never blinking, never changing. There are no windows. Time loses all meaning. And it’s strange to think about this because you don’t walk in and think, “What is that racket?” But there’s 10,000, 100,000 however many fans spinning all the time. It is super loud. It can clear 120 decibels in there, but it’s a white noise so you don’t necessarily hear it. Hearing protection is important there.

Mike: When I was at Oak Ridge, we had—all of our data centers, we had a professional data center manager, so everything was absolutely pristine. And to get into any of the data centers, you had to go through a training; it was very simple training, but just, like, “These are things you do and don’t do in the data center.” And when you walked in, you had to put in earplugs immediately before you walked in the door. And it’s so loud just because of that, and you don’t really notice it because you can walk in without earplugs and, like, “Oh, it’s loud, but it’s fine.” And then you leave a couple hours later and your ears are ringing. So, it’s a weird experience.

Corey: It’s awful. I started wearing earplugs every time I went in, just because it’s not just the pain because hearing loss doesn’t always manifest that way. It’s, I would get tired much more quickly.

Mike: Oh, yeah.

Corey: I would not be as sharp. It was, “What is this? Why am I so fatigued?” It’s noise.

Mike: Yeah. And having to remember to grab your jacket when you head down to the data center, even though it’s 95 degrees outside.

Corey: At some point, if you’re there enough—which you probably shouldn’t be—you start looking at ways to wind up storing one locally. I feel like there could be some company that makes an absolute killing by renting out parkas at data centers.

Mike: Yeah, totally. The other group of data center stuff that I generally run into is the exact opposite of that. And it’s basically someone has shoved a couple racks in somewhere and they just kind of hope for the best.

Corey: The basement. The closet. The hold of a boat, with one particular client we work with.

Mike: Yeah. That was an interesting one. So, we had a—Corey and I had a client where they had all their infrastructure in the basement of a boat. And we’re [laugh] not even kidding. It’s literally in the basement of a boat.

Corey: Below the waterline.

Mike: Yeah below the waterline. So, there was a lot of planning around, like, what if the hold gets breached? And like, who has to plan for that sort of thing? [laugh]. It was a weird experience.

Corey: It turns out that was—was hilarious about that was while they were doing their cloud migration into AWS, their account manager wasn’t the most senior account manager because, at that point, it was a small account, but they still stuck to their standard talking points about TCO, and better durability, and the rest, and it didn’t really occur to them to come back with a, what if the boat sinks? Which is the obvious reason to move out of that quote-unquote, “data center?”

Mike: Yeah. It was a wild experience. So, that latter group of just everything’s an absolute wreck, like, everything—it’s just so much of a pain to work with, and you find yourself wanting to clean it up. Like, install new racks, do new cabling, put in a totally new floor so you’re not standing on concrete. You want to do all this work to it, and then you realize that you’re just putting lipstick on a pig; it’s still going to be a dirty old data center at the end of the day, no matter how much work you do to it. And you’re still running on the same crappy hardware you had, you’re still running on the same frustrating deployment process you’ve been working on, and everything still sucks, despite it looking good.

Corey: This episode is sponsored in part by ChaosSearch. As basically everyone knows, trying to do log analytics at scale with an ELK stack is expensive, unstable, time-sucking, demeaning, and just basically all-around horrible. So why are you still doing it—or even thinking about it—when there’s ChaosSearch? ChaosSearch is a fully managed scalable log analysis service that lets you add new workloads in minutes, and easily retain weeks, months, or years of data. With ChaosSearch you store, connect, and analyze and you’re done. The data lives and stays within your S3 buckets, which means no managing servers, no data movement, and you can save up to 80 percent versus running an ELK stack the old-fashioned way. It’s why companies like Equifax, HubSpot, Klarna, Alert Logic, and many more have all turned to ChaosSearch. So if you’re tired of your ELK stacks falling over before it suffers, or of having your log analytics data retention squeezed by the cost, then try ChaosSearch today and tell them I sent you. To learn more, visit chaossearch.io.

Corey: The worst part is playing the ‘what is different here?’ Game. You rack twelve servers: eleven come up fine and the twelfth doesn’t.

Mike: [laugh].

Corey: It sounds like, okay, how hard could it be? Days. It can take days. In a cloud environment, you have one weird instance. Cool, you terminate it and start a new one and life goes on whereas, in a data center, you generally can’t send back a $5,000 piece of hardware willy nilly, and you certainly can’t do it same-day, so let’s figure out what the problem is.

Is that some sub-component in the system? Is it a dodgy cable? Is it, potentially, a dodgy switch port? Is there something going on with that node? Was there something weird about the way the install was done if you reimage the thing? Et cetera, et cetera. And it leads
down rabbit holes super quickly.

Mike: People that grew up in the era of computing that Corey and I did, you start learning tips and tricks, and they sound kind of silly these days, but things like, you never create your own cables. Even though both of us still remember how to wire a Cat 5 cable, we don’t.

Corey: My fingers started throbbing when you said that because some memories never fade.

Mike: Right. You don’t. Like, if you’re working in a data center, you’re buying premade cables because they’ve been tested professionally by high-end machines.

Corey: And you still don’t trust it. You have a relatively inexpensive cable tester in the data center, and when—I learned this when I was racking stuff the second time, it adds a bit of time, but every cable that we took out of the packaging before we plugged it in, and we tested on the cable tester just to remove that problem. And it still doesn’t catch everything because, welcome to the world of intermittent cables that are marginal that, when you bend a certain way, stop working, and then when you look at them, start working again properly. Yes, it’s as maddening as it sounds.

Mike: Yeah. And then things like rack nuts. My fingers hurt just thinking about it.

Corey: Think of them as nuts that bolts wind up screwing into but they’re square and they have clips on them so they clip into the standard rack cabinets, so you can screw equipment into them. There are different sizes of them, and of course, they’re not compatible with one another. And you have—they always pinch your finger and make you bleed because they’re incredibly annoying to put in and out. Some vendors have quick rails, which are way nicer, but networking equipment is still stuck in the ‘90s in that context, and there’s always something that winds up causing problems.

Mike: If you were particularly lucky, the rack nuts that you had were pliable enough that you could pinch them and pull them out with your fingers, and hopefully didn’t do too much damage. If you were particularly unlucky, you had to reach for a screwdriver to try to pry it out, and inevitably stab yourself.

Corey: Or sometimes pulling it out with your fingers, it’ll—like, those edges are sharp. It’s not the most high-quality steel in some cases, and it’s just you wind up having these problems. Oh, one other thing you learn super quickly, is first, always have a set of tools there
because the one you need is the one you don’t have, and the most valuable tool you’ll have is a pair of wire cutters. And what you do when you find a bad cable is you cut it before throwing it away.

Mike: Yep.

Corey: Because otherwise someone who is very well-meaning but you will think of them as the freaking devil, will, “Oh, there’s a perfectly good cable sitting here in the trash. I’ll put it back with the spares.” So you think you have a failed cable you grab another one from the pile of spares—remember, this is two in the morning, invariably, and you’re not thinking on all cylinders—and the problem is still there. Cut the cable when you throw it away.

Mike: So, there are entire books that were written about these sorts of tips and tricks that everyone working [with 00:19:34] data center just remembers. They learned it all. And most of the stuff is completely moot now. Like, no one really thinks about it anymore. Some people are brought up in computing in such a way that they never even learned these things, which I think it’s fantastic.

Corey: Oh, I don’t wish this on anyone. This used to be a prerequisite skill for anyone who called themselves a systems administrator, but I am astonished when I talk to my AWS friends, the remarkably senior engineers I talk to who have never been inside of an AWS data center.

Mike: Yeah, absolutely.

Corey: That’s really cool. It also means you’re completely divorced from the thing you’re doing with code and the rest, and the thing that winds up keeping the hardware going. It also leads to a bit of a dichotomy where the people racking the hardware, in many cases, don’t understand the workloads that are on there because if you have the programming insight, and ability, and can make those applications work effectively, you’re probably going to go find a role that compensates far better than working in the data center.

Mike: I [laugh] want to talk about supply chains. So, when you build a data center, you start planning about—let’s say, I’m not Amazon. I’m just, like, any random company—and I want to put my stuff into a data center. If I’m going to lease someone else’s data center—which you absolutely should—we’re looking at about a 180-day lead time. And it’s like, why? Like, that’s a long time. What’s—

Corey: It takes that long to sign a real estate lease?

Mike: Yeah.

Corey: No. It takes that long to sign a real estate lease, wind up talking to your upstream provider, getting them to go ahead and run the thing—effectively—getting the hardware ordered and shipped in the right time window, doing the actual build-out once everything is in place, and I’m sure a few other things I’m missing.

Mike: Yeah, absolutely. So yeah, you have all these things that have to happen, and all of them pay for-freaking-ever. Getting Windstream on the phone to begin with, to even take your call, can often take weeks at a time. And then to get them to actually put an order for you, and then do the turnup. The turnup alone might be 90 days, where I’m just, “Hey, I’ve bought bandwidth from you, and I just need you to come out and connect the [BLEEP] cables,” might be 90 days for them to do it.

And that’s ridiculous. But then you also have the hardware vendors. If you’re ordering hardware from Dell, and you’re like, “Hey, I need a couple servers.” Like, “Great. They’ll be there next week.” Instead, if you’re saying, “Hey, I need 500 servers,” they’re like, “Ooh, uh, next year, maybe.” And this is even pre-pandemic sort of thing because they don’t have all these sitting around.

So, for you to get a large number of servers quickly, it’s just not a thing that’s possible. So, a lot of companies would have to buy well ahead of what they thought their needs would be, so they’d have massive amounts of unused capacity. Just racks upon racks of systems sitting there turned off, waiting for when they’re needed, just because of the ordering lead time.

Corey: That’s what auto-scaling looks like in those environments because you need to have that stuff ready to go. If you have a sudden inrush of demand, you have to be able to scale up with things that are already racked, provisioned, and good to go. Sometimes you can have them halfway provisioned because you don’t know what kind of system they’re going to need to be in many cases, but that’s some up-the-stack level thinking. And again, finding failed hard drives and swapping those out, make sure you pull the right or you just destroyed an array. And all these things that I just make Amazon’s problem.

It’s kind of fun to look back at this and realize that we would get annoyed then with support tickets that took three weeks to get resolved in hardware, whereas now three hours in you and I are complaining about the slow responsiveness of the cloud vendor.

Mike: Yeah, the amount of quick turnaround that we can have these days on cloud infrastructure that was just unthinkable, running in data centers. We don’t run out of bandwidth now. Like, that’s just not a concern that anyone has. But when you’re running in a data center, and, “Oh, yeah. I’ve got an OC-3 line connected here. That’s only going to get me”—

Corey: Which is something like—what is an OC-3? That’s something like, what, 20 gigabit, or—

Mike: Yeah, something like that. It’s—

Corey: Don’t quote me on that.

Mike: Yeah. So, we’re going to have to look that up. So, it’s equivalent to a T-3, so I think that’s a 45 megabit?

Corey: Yeah, that sounds about reasonable, yeah.

Mike: So, you’ve got a T-3 line sitting here in your data center. Like that’s not terrible. And if you start maxing that out, well, you’re maxed out. You need more? Again, we’re back to the 90 to 180 day lead time to get new bandwidth.

So, sucks to be you, which means you’d have to start planning your bandwidth ahead of time. And this is why we had issues like companies getting Slashdotted back in the day because when you capped the bandwidth out, well, you’re capped out. That’s it. That’s the game.

Corey: Now, you’ve made the front page of Slashdot, a bunch of people visited your site, and the site fell over. That was sort of the way of the world. CDNs weren’t really a thing. Cloud wasn’t a thing. And that was just, okay, you’d bookmark the thing and try and remember to check it later.

We talked about bandwidth constraints. One thing that I think the cloud providers do—at least the tier ones—that are just basically magic is full line rate between any two instances almost always. Well, remember, you have a bunch of different racks, and at the top of every rack, there’s usually a switch called—because we’re bad at naming things—top-of-rack switches. And just because everything that you have plugged in can get one gigabit to that switch—or 10 gigabit or whatever it happens to be—there is a constraint in that top-of-rack switch. So yeah, one server can talk to another one in a different rack at one gigabit, but then you have 20 different servers in each rack all trying to do something like that and you start hitting constraints.

You do not see that in the public cloud environments; it is subsumed away, you don’t have to think about that level of nonsense. You just complain about what feels like the egregious data transfer charge.

Mike: Right. Yeah. It was always frustrating when you had to order nice high-end switching gear from Cisco, or Arista, or take your pick of provider, and you got 48 ports in the top-of-rack, you got 48 servers all wired up to them—or 24 because we want redundancy on that—and that should be a gigabit for each connection, except when you start maxing it out, no, it’s nowhere even near that because the switch can’t handle it. And it’s absolutely magical, that the cloud provider’s like, “Oh, yeah. Of course, we handle that.”

Corey: And you don’t have to think about it at all. One other use case that I did want to hit because I know we’ll get letters if we don’t, where it does make sense to build out a data center, even today, is if you have regulatory requirements around data residency. And there’s no cloud vendor in an area that suits. This generally does not apply to the United States, but there are a lot of countries that have data residency laws that do not yet have a cloud provider of their choice region, located in-country.

Mike: Yeah, I’ll agree with that, but I think that’s a short-lived problem.

Corey: In the fullness of time, there’ll be regions everywhere. Every build—a chicken in every pot and an AWS availability zone on every corner.

Mike: [laugh]. Yeah, I think it’s going to be a fairly short-lived problem, which actually reminds me of even our clients that have data centers are often treating the data center as a cloud. So, a lot of them are using your favorite technology, Corey, Kubernetes, and they’re treating Kubernetes as a cloud, running Kube in AWS, as well, and moving workloads between the two Kube clusters. And to them, a data center is actually not really data center; it’s just a private cloud. I think that pattern works really well if you have a need to have a physical data center.

Corey: And then they start doing a hybrid environment where they start expanding to a public cloud, but then they treat that cloud like just a place to run a bunch of VMs, which is expensive, and it solves a whole host of problems that we’ve already talked about. Like, we’re bad at replacing hard drives, or our data center is located on a corner where people love to get drunk on the weekends and smash into the power pole and take out half of the racks here. Things like that great, yeah, cloud can solve that, but cloud could do a lot more. You’re effectively worsening your cloud experience to improve your data center experience.

Mike: Right. So, even when you have that approach, the piece of feedback that we give the client was, you have built such a thing where you have to cater to the lowest common denominator, which is the constraints that you have in the data center, which means you’re not able to use AWS the way that you should be able to use it so it’s just as expensive to run as a data center was. If they were to get rid of the data center, then the cloud would actually become cheaper for them and they would get more benefits from using it. So, that’s kind of a business decision for how they’ve structured it, and I can’t really fault them for it, but there are definitely some downsides to the approach.

Corey: Mike, thank you so much for joining me here. If people want to learn more about what you’re up to, where can they find you?

Mike: You know, you can find me at duckbillgroup.com, and actually, you can also find Corey at duckbillgroup.com. We help companies lower their AWS bills. So, if you have a horrifying bill, you should chat.

Corey: Mike, thank you so much for taking the time to join me here.

Mike: Thanks for having me.

Corey: Mike Julian, CEO of The Duckbill Group and my business partner. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice and then challenge me to a cable-making competition.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Guy

Guy Raz is a Sr. Systems Engineer at ExtraHop with previous experience as a Network Engineer and Solution Architect. Guy is one of the SMEs leading the unique ExtraHop approach to cloud-native NDR for the hybrid multi-cloud enterprise. Before joining the Sales Engineer team, Guy was one of the ExtraHop Solution Architects, responsible for conducting deep technical and business discovery sessions, assisting in troubleshooting and problem resolution during war-room and security/network investigations, and developing strategies for acquiring high-value data from the wire; requiring in-depth technical understanding of L2-L7 networking principles.

Links:

  • https://www.extrahop.com/
  • https://extrahop.com/demo

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Thinkst. This is going to take a minute to explain, so bear with me. I linked against an early version of their tool, canarytokens.org in the very early days of my newsletter, and what it does is relatively simple and straightforward. It winds up embedding credentials, files, that sort of thing in various parts of your environment, wherever you want to; it gives you fake AWS API credentials, for example. And the only thing that these things do is alert you whenever someone attempts to use those things. It’s an awesome approach. I’ve used something similar for years. Check them out. But wait, there’s more. They also have an enterprise option that you should be very much aware of canary.tools. You can take a look at this, but what it does is it provides an enterprise approach to drive these things throughout your entire environment. You can get a physical device that hangs out on your network and impersonates whatever you want to. When it gets Nmap scanned, or someone attempts to log into it, or access files on it, you get instant alerts. It’s awesome. If you don’t do something like this, you’re likely to find out that you’ve gotten breached, the hard way. Take a look at this. It’s one of those few things that I look at and say, “Wow, that is an amazing idea. I love it.” That’s canarytokens.org and canary.tools. The first one is free. The second one is enterprise-y. Take a look. I’m a big fan of this. More from them in the coming weeks.

Corey: This episode is sponsored in part by our friends at Lumigo. If you’ve built anything from serverless, you know that if there’s one thing that can be said universally about these applications, it’s that it turns every outage into a murder mystery. Lumigo helps make sense of all of the various functions that wind up tying together to build applications. It offers one-click distributed tracing so you can effortlessly find and fix issues in your serverless and microservices environment. You’ve created more problems for yourself; make one of them go away. To learn more, visit lumigo.io.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Once a year in San Francisco, if I find myself being overly cheerful, all I have to do is walk up and down the RSA Expo Hall and look at a bunch of vendors talking about how their on-premises product kind of sort of works in the cloud, and then I’m not overly cheerful anymore. One notable exception to this is a company called Extrahop. I’ve spoken about them before and on this promoted episode, we’re going to dive a little bit deeper. Today, my guest is Senior Systems Engineer Guy Raz. Guy, thanks for taking the time to speak with me.

Guy: Thanks, Corey, happy to be here.

Corey: So, for those who have not caught previous episodes, or heard me ranting from the rooftop about it, at a very basic level for folks who have not even, I guess, dip their toes in the RSA space because they, you know, want to be happy with their lives, what is ExtraHop?

Guy: ExtraHop is a cloud-native approach for analyzing wire data. Historically, customers have, kind of, looked at TAP SPANs, but with cloud, there’s a ton of ways of getting this natively. You know, AWS, GCP, Azure give us ways of collecting this data. ExtraHop is a platform for analyzing that network traffic and, in real-time, providing context to application and
security teams.

Corey: So, when you take a look at that from, I guess, the perspective of security, it’s easy to sit here and say, “Oh, so how do you wind up thinking about security in a place or time of cloud?” Because there’s an awful lot of ways to view it: you can go down the path of, “Ah, I’m going to just use all the first-party tooling from my provider, and that’s it,” which, that could be fair. Alternatively, you could go down a different path of, “I’m going to just go ahead and buy whatever they’ll sell me at RSA,” which is great because the hardest part there is the booth attendees not making actual cash register sounds with their mouths when you walk past with an open checkbook. But security always feels like a thing that’s kind of an afterthought. It’s something that is tied too closely, on some level, to this idea that you’re never going to be secure, so you may as well just give up. It’s also something people only care about after it’s been a little too late, where they really should have been caring about it. How do you see that?

Guy: It’s a really unfortunate space, but you’re absolutely right, Corey, there. What we end up seeing as a lot of customers, and just the industry as a whole tends to be an afterthought when it comes to cloud. They assume cloud-native solutions or built-in free solutions have their best foot forward, have their best instance in mind. And that’s not always the case. There’s a lot of, like you mentioned, built-in solutions that these cloud providers can give us.

And while a lot of them are kind of scratching the surface of what security in the cloud can provide, there’s a lot that it kind of leaves unanswered. And the unfortunate thing is, the cloud journey isn’t always the easiest. There’s a lot of lift-and-shift, there’s a lot of refactor, and sometimes the security portion of that gets put on the side street until it becomes a priority or an event happens.

Corey: So, given that you can effectively not even swing a dead cat anymore without hitting 15 different security vendors all claiming to do everything you’d want, start to finish, what makes ExtraHop different? How do you approach security that’s differentiated from the rest of the, I guess, entire security industry?

Guy: Yeah, that’s a really good question. I think my favorite part, and one of the reasons I love our product is the data stream that we collect. Network data is a huge source of information that’s just sitting there silently, kind of, waiting to be consumed and analyzed. In the old on-premise environment, there were legacy packet capture solutions, or ways of grabbing this information from a SPAN or a TAP. But it’s still the same data stream as we go to the cloud, it’s just a slightly different way of collecting it.

So, the biggest thing that I would encourage people is, use the data that’s there. The network traffic is passing your infrastructure: it’s EC2s hitting your S3 buckets, it’s RDS instances going through a load balancer to a Lambda function. It’s all just traversing through infrastructure that you just don’t own anymore, but getting that information is a huge differentiator. You’re talking about every packet of every transaction being analyzed in real-time at a cloud-scale, which, you know, you need a smaller instance today—it’s smaller today—you need a bigger instance tomorrow, it just auto-scales up.

Corey: Now, back in the world of data centers, I agreed an awful lot with what you’re saying, as far as looking at the network as the first point of, I guess, the arbiter of truth, for lack of a better term. And, on some level in cloud, I feel like I’ve drifted away from that. Now, back in our days at data centers, you don’t know what’s running on these systems; you don’t know what various engineers are shoved onto them, but generally speaking, you can mostly trust the network. Please don’t email me. So, once you move into a cloud world, everything sort of changes a bit.

You don’t really have to think about any of the layer 2 networking, and most of the layer 3 networking sort of goes away, too. Plus, let’s be very realistic; from the perspective of the virtual machines you’re running in a cloud environment, everything beyond that is kind of a lie. There’s a bunch of encapsulation, you’re higher up the stack, you’re not on hardware anymore so, on some level, it always felt that, eh, networking is not really the same thing in the cloud environment. I can ignore it. And I have to admit, back when I first started talking to you folks, I was something of a skeptic.

And then you, more or less, made me change my perspective through a very sneaky approach of spinning up a test account for me with ExtraHop, and now I get it in a way I never did before. Is that aha moment common to the, I guess, the cloud-native set, or do most people come into this with a much more rational and reasoned approach to networking in the cloud?

Guy: I would say it’s both. We have customers who are familiar with the type of information we can provide going through their cloud journey, or are starting their cloud journey and they want the same type of visibility. But for our net-new customers, when we hit that whitespace, that aha moment comes, and it’s so much fun to see. Someone who had no idea what this type of data can provide; they’re used to legacy telemetry or log information. So, that aha moment is something that, as someone who gets to interact with customers, is one of my favorite parts of the job. And I would say it’s fun to play with and show that.

Corey: Now, I want to be clear that, again, in the interest of full disclosure, now, since I’ve put this in my test account, ExtraHop is now the second most expensive consumer of AWS services. But it’s not as bad as folks might think. It’s using a VPC mirror in order to look at traffic, and that costs me the princely sum of somewhere between $10 and $11 a month. And that doesn’t really vary, regardless of how much traffic I shove through this thing. It’s not doing a whole lot in the AWS
account; if I didn’t know that was there and that’s what it was doing, I would ignore the spend line entirely. How does this
work? What are you doing in order to get access to seeing what is happening, “On the wire,” quote-unquote, in a cloud
environment?

Guy: Just focusing on AWS for a second, since that’s what you called out. It’s using a native built-in functionality that Amazon provides. It’s called VPC packet mirroring. It’s super simple: you deploy an ExtraHop collector into your VPC, you set that up as a destination of your traffic, and then you configure what’s called a monitoring session in VPC. You can say I want it to do based on these tags, I want it to send traffic based on this subnet—or any there combination of—and it just kind of works. You know, it’s beautiful.

And where we’re kind of taking this to the next step is using some intelligent Lambda automation to ensure that anytime a new instance gets spun up, whether it’s tagged, untagged, deployed into a different VPC, or is a different instance size, it gets automatically added into this data feed. So, you know, you talk about the ephemerality of the cloud and how instances can spin up and spin down almost instantaneously, as soon as an instance is up, before it even gets any traffic sent to it, traffic is [laugh] coming to the ExtraHop, right? We’ll see IMDS traffic, we’ll see instance metadata, we’ll get the ENI information, all just by sitting there, passively listening.

Corey: One of the things that I found particularly, I guess—appreciated about your entire approach is I didn’t have to change anything about what was actually running in this account. I didn’t have to teach the EC2 instances that something else was going on. I didn’t have to reconfigure anything on an application basis. This was purely done in the underlying VPC configuration. It was done without any downtime whatsoever.

And I feel like that is an understated benefit for an awful lot of tooling. “Oh, just go ahead and roll this thing out to all of your environment.” Like, yeah, there are tens of thousands of instances and VMs scattered throughout our entire estate. Exactly how long do you think we’re going to spend on this? You don’t have that problem here, and it’s kind of nice.

Guy: It is really nice. And not to take anything away from some agent solutions because they do have their [crosstalk 00:09:46]—

Corey: Oh, I will, but please go on.

Guy: [laugh]. But this approach to security and monitoring in the cloud, to your point, Corey, is seamless. Application owners don’t know it’s there. It doesn’t add any added load. I’m a former network engineer. Troubleshooting different instances or different virtual machines, the first thing I used to do is turn off those agents, right? Is this consuming CPU resources? Is this slowing down my agent? That’s no longer the case in cloud. That’s no longer the case with this network-based approach.

Corey: I’ll also point out that it always feels like there’s a false dichotomy when we’re talking about security vendors. And it either feels like, oh, you’re in a bunch of data-center style environments, you’re migrating into the cloud, but basically today, your environment is a bunch of VMs, and maybe a load balancer or an object store. And a lot of tooling speaks super well to that use case. But then if you take a step back and look at well, the lie that all these companies love to tell themselves, and I’m no more immune to this than they are, to be very clear here, but we all tell ourselves this beautiful lie which is after this next sprint ends, then, then we’re going to go ahead and pay off all of our technical debt and things are going to be done properly with a capital P. And it never happens, but it’s the lie we tell ourselves.

And we make financial decisions, in some cases, tied to that false vision of, “Well, why would I wind up embracing something that is aimed at that particular use case because once we wind up going full-on cloud-native and embracing our provider of choice, all of this stuff is going to change?” What I like about ExtraHop is, all right, assume you’re in that mythical born-in-the-cloud world where you have a significant estate that everything runs on top of these higher-level services. ExtraHop is still there, still working, and still doing exactly the sorts of things we’re talking about here. No matter where you are on that transformational journey, it feels like there’s an answer here. Is that accurate? Have I been gargling the marketing tea too heavily? What’s the story here?

Guy: No, that’s pretty accurate. And it doesn’t really matter where you are on your cloud journey; security can’t be foregone for the sake of this cloud instance. We see this day in, day out. You know, if you subscribe to as many news alerts as I do, it’s a scary world. Just even recently this past weekend, we had a—not our customer, but there was an attack against an oil pipeline.

That came through a cloud vulnerability. IAM account leakage, and service accounts, and open S3 buckets. It’s a scary part of this cloud journey. We want to make sure that we’re scaling, we want to reduce our physical footprint, but we can’t forgo the security and the trust that our customers have in our applications. And that means that having an approach to security in the cloud needs to be top of mind, regardless of where you are in that cloud journey.

Corey: I think one of the, I guess, biggest concerns in the security space is very similar to what I deal with in the cost optimization space, which is people care about it only after they really, really, really should have cared about it, on some level. Now, over in the billing world that I live in, people generally have a failure mode of, “Well, we spent a little too much money,” and that is generally a very survivable thing. I used to say—tongue-in-cheek, only I was being completely serious—one of the reasons I went with AWS billing as my direction of choice was that no one is going to come and call me at two o’clock in the morning with a billing emergency; it is strictly a business hours problem. Security is a very different world. But if you screw up the bill, you spent too much money.

If you screw up security, well, your company’s name is mud, you could try and pull a SolarWinds with a ring of ablative interns to wind up trying to pass the buck off onto, but in practice, you’re probably losing a CSO and a few other high-level execs as a sort of token offering to the market gods. And it’s painful, and I’m hard-pressed to name a company these days that has not suffered at least some form of data breach somewhere. It almost feels like it’s a losing game.

Guy: It’s not a losing game, but it is a post-breach world, right? It’s not a question of, if you get breached. It’s more a question of what security holes have been left open, and what can they collect from these holes? And minimizing that attack surface is obviously critical, but understanding the damage and reacting to it as fast as possible is just as important. And honestly, that’s, kind of, my favorite parts about the cloud.

You know, I can see something like a suspicious transaction, or a large increase in web traffic, and then fire off an API to Lambda that says, “Deploy the security group onto this instance.” That whole process takes milliseconds. So, the reaction time that we have with the cloud vastly surpasses what we ever had in the data center. And yeah, you’re right, maybe that adds up costing a little bit more, or creates a slightly higher bill because we called a couple Lambda functions, but no exfiltration of data; no loss of customer information. You can’t trade that off, at the end of the day.

Corey: The thing that always, I guess, sort of bothered me about various breaches or various security reports is whenever companies will say definitively, “We have never suffered a security breach,” that might mean that they are absolutely on point—though, you always have this probabilities question—but it could also mean that they have no effective visibility or effective logging, and that is the dangerous part. It’s similar to this idea of back once upon a time in the early days of unbreakable Linux, when Oracle was pushing that and they said, “It is unhackable.” The entire internet proved them wrong within hours because everything can be broken into at some point. It’s just a question of how high do you raise that bar? Ideally, a little bit above random people just scanning S3 buckets.

Guy: Yeah, and you know, that’s really scary, kind of, the data that we get to see when—you know, you called this earlier that aha moment. Because we’re an always-on solution, we get to see the hygiene of the network, too. I can tell you when someone hit an insecure S3 bucket, or an IAM role logged in at two in the morning that it never has before, or someone sent an API command to Lambda to spin up another instance at two in the morning, using a service account that has admin permissions. It’s a scary world in the cloud, and making sure you have that surface covered gets you to those aha moments quicker.

Corey: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the Enterprise (not the starship). On-prem security doesn’t translate well to cloud or multi-cloud environments, and that’s not even counting IoT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IoT devices, detects these threats up to 35 percent faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at extrahop.com/trial.

Corey: One thing that I do want to draw a little bit of attention to as well, having kicked the tires on ExtraHop for a few months now, I keep forgetting that I have it in place. And the only time I really get reminded is that $10 a month for that attachment to the VPC that I see on my bill when I go over that thing with a fine-tooth comb because of who I am and what I do. My point being is that I have instances in that account that are doing a bunch of relatively strange things from time to time. And the behavior is not consistent from day to day. One of them has an IRC bouncer hanging out on it because I used to spend a disproportionate amount of my time on freenode, and it does a whole bunch of different things that looks super weird.

And every time I wind up pointing a typical security product at it, it starts shrinking its head off—if it can even get that far into it—of, “This thing is clearly exploited. Shut it down, shut it down, shut it down.” And none of that happens. I mean, this thing looks very weird on the network, I’m not going to deny otherwise. This is my development box.

When I’m on the road—remember back when we used to travel places?—and I would just be connecting from an iPad and remoting into this thing, and then I would have it do all of the things I would normally do on a desktop computer. But it doesn’t make noise. Now, to be clear, I also have a somewhat decent security posture on this thing so it’s not a story of it getting actively exploited and it should be making noise. But it just doesn’t say anything. It just sort of sits there quietly in the background. And it works. Whenever I log in, I have to click around to make sure it actually is still working because there’s nothing on the dashboard where it’s just giving you noise to talk about noise. Why is this such a rarity?

Guy: [laugh]. So, your environment is probably pretty secure. I imagine you’re not deploying hundreds and thousands of containers and EC2s and spinning up all this type of data, but—

Corey: No. It’s tiny, I spend 50 bucks a month on this account.

Guy: So, it’s not atypical, the behavior you see. You know, I’ve been in POCs and proof of values where we deployed the ExtraHop, and it doesn’t see too much. And so one thing I’ve started doing for a lot of my customers is deploying a lab for them. Do you trust that something like an ExtraHop will see ransomware? Do you trust that ExtraHop will see credential harvesting, and lateral movement, and exfiltration?

Or are you using your ExtraHop to troubleshoot your web applications? Let me spin up a lab for you, throw some workloads in there. We’ll drop a Kali instance or a Kubernetes cluster and show you what an attack surface can look like. Not to scare or, kind of, build on what customers are experiencing, but knock on wood, I don’t want any of my customers to be attacked, but I also have to build that confidence that if or when something happens, they’re covered.

Corey: Back when I first had ExtraHop demoed for me, I was convinced it was going to be garbage, let me be very honest with you. And the reason was that the dashboard looked like it was demoware. It was well-designed, well-executed, it had a very colorful interface. It felt like bossware if I’m being perfectly honest. My belief has always been, you either get a good interface that works and is easy to use and navigate within, or you get something that looks super flashy when you do a demo on stage somewhere, but it is almost impossible to wind up effectively nailing both of those use cases. And then I started using this and I am having to eat those words because you actually did it. You wound up building something that looks great and is easy to navigate. How much work did that actually take? I mean, is that where all the engineering
on this product has gone?

Guy: We really appreciate it. Our UX team and our engineering group work very, very hard. We spend more on R&D and research than we do on a lot of our marketing and front-end sectors and it shows. The product kind of speaks for itself. And the experience that you’re describing with the easy-to-consume UI, with the data to support that experience behind it is our goal. And I’m happy to hear that you’re enjoying it in your lab.

Corey: I just did a little poking around while I have you on the phone, and if I dig deep enough, it does tell me that there’s some weak ciphers in use. And every single one of these things is talking to an AWS-owned endpoint, which is, first, a little bit on the hilarious side, since I keep this thing current. Awesome. Secondly, the fact that I had to dig for that and it wasn’t freaking out about it. There are no alerts; it doesn’t show up on the dashboard.

I had to really start diving into this. Because, yeah, it’s good to know if I’m doing some sort of audit activity, it’s good to know if I need to dive in and look at these things, but it doesn’t need to wake me up at two in the morning because, “Holy crap. The Boto3 library isn’t quite using the latest cipher suite.” How much tuning did this take?

Guy: Not much. So, there is a learning period, as with any application that has a backend on behavioral analytics. But most of my customers, usually two to three weeks after we start seeing a data feed, are in a state of excellent tuning. Very little manual tuning required, the system will learn normalities, it’ll learn behaviors, and it’ll flag anomalies, kind of, on its own. So, the same experience that you’re having where you’re running a compliance scan, or you’re running an audit, or you’re trying to look for, in this world where—I’m going to make a joke here—we all have free time, and you have the time to go look at, you know, “How do I clean up some of these hygienic issues that are not currently causing me heartache?” The data is there. That’s the beauty of the network is some of your users may be familiar with Wireshark, or something like a tcpdump. There’s boatloads of data in. There are thousands and thousands of data points you can analyze though. If you want the data, it’s there, but like you said, no reason to wake you up at two in the morning unless we see things that are super critical.

Corey: Encrypt everything sort of becomes the theme, especially when Amazon’s CTO slaps it on a t-shirt, and then in some cases charges extra for it; but that’s a diversion. What is the story as you start seeing more and more traffic wind up being encrypted at a bunch of different levels? In fact, I’ll take it a step further. With the rise of customer-managed keys and things like KMS in the AWS world, does that mean that ExtraHop is effectively losing visibility beyond just the typical TCP flow?

Guy: So, ExtraHop is unique in the space that we have the ability to decrypt TLS 1.3 data. It came out a couple years ago and it’s a way of encrypting traffic between servers and clients in a manner that isn’t as breakable as historic encryption mechanisms were. We can parse that data, we can ingest those decryption mechanisms, we can—in real-time, without being a man-in-the-middle so we’re not breaking any of this trust chain that you have to explicitly build to the internet in a lot of cases, or you don’t have to upload any of your private keys to the ExtraHop. So, it’s a super unique approach for how we can unpack that envelope.

This goes back to when we were kids, and we all got those Christmas presents and you check the box and you try to guess what’s inside. And maybe you’re right, maybe you’re not, but until you open that wrapper, you can’t really know what’s being said. So, something like a hidden database transaction underneath a web call just shows up as a web call when you’re not unpacking the envelopes. Decryption is an underrated feature, in my opinion, and I would—you know, true security posture team should probably have something where they can look inside those payloads.

Corey: This is where it starts to get a little weird, too, because, on some level, great, the whole premise of TLS is that my application talks to something far away—or nearby. It doesn’t really matter—but there’s a bit of a guarantee that from the point it leaves that application and hits the encryption side on the instance to the other end, there should be no decryption there. The only way I’ve ever seen that get around that is effectively man-in-the-middling these things, which in some level, “Oh, decrypt all of your secure traffic in the name of security,” always felt a little on the silly side.

Guy: Not only is it silly, it’s a little harder to manage when we talk about cloud because those man-in-the-middle decryption mechanisms typically involve building explicit trust so that they can decrypt the traffic, and then the client and the server both agree that, “Yeah, sure. You can read my information. You use your own certificate. I don’t care.” That gets harder to do as you start talking about containers, as you start talking about ephemeral instances.

Sure, you can build a golden image of a container and make it trust your IPS—which most people should have—but you still have to have the ability to see this traffic when you’re bypassing certain metrics. If you’re bypassing traffic back to your data center so you can [unintelligible 00:24:45] your point of sale application, or if, maybe, you’re a multi-cloud environment where you have to pass from cloud to consume all of your data space. You still have to be able to see that data to understand what’s really being said during the conversation without always being able to break that trust chain.

Corey: One thing that I want to make very clear I call out because otherwise, I am going to get letters on this. This is a promoted episode. You folks have paid to sponsor. Thank you. It is appreciated. But I want to be very clear you buy my attention, not my opinion. I know I’ve been, sort of, gushing about what ExtraHop does, and how it works, and how I view these things, but that’s not because you’re paying me to do that. I am legitimately excited about the product itself.

This is one of those things where it finally is giving me visibility into something that I understand from my olden sysadmin network admin days combined with how I know the cloud works today, and I’m looking at this and the strange spots that I see of, “Ohh, I would improve that a bit,” there aren’t that many and they’re not that big. This is something that is legitimately awesome, and I would encourage people to kick the tires and see what they think.

Guy: Yeah, we appreciate that feedback, Corey. A lot of us are previous users. I myself, you know, before coming to ExtraHop, used ExtraHop at a previous job and that was one of the big reasons I came to work for the company is I believe in the software. A lot of our people here and we have long-time-term employees believe in what we do. And our goal is to build this partnership and trust with our customers, too, so that they have the same experience that you do. It’s a fun product to play with, and kicking around and tires is fun and we’d love to show you.

Corey: When you start talking to folks who are going through their, I guess, ExtraHop journey of discovery—don’t ever use that term. It sounds awful—what do you find that they are getting the most confused about? What do they misunderstand that would be helpful for them to have more clarity around?

Guy: There’s a lot of what ExtraHop can provide when it comes to data ingestion, and data collection, and even data aggregation, but where a lot of my customers fall in the confusion space tends to be in, “Do I care about this data? Should I care about this information?” And that really falls down to the individual user’s responsibility. A security team cares about all of it, whereas an application team may only care about the website’s performance, or the network latency, or the error rates. And it spans the gambit.

So, one thing that I do with a lot of my customers is weekly training sessions, or give them access to videos that we’ve recorded in advance so they can self-teach. As an engineer myself, I hate when people talk me into things: I like to play, and I like to see. So, let me give you a guide, you want to play with it, kind of poke the toes, kick the tires, have fun, that seems to get customers excited, and again, back to that aha moment a lot quicker. There’s so much data that gets exposed, and sometimes it can be overwhelming. But when it comes to visibility, it’s all stuff that’s useful at the end of the day.

Corey: If people want to learn more, where can they go next? How do they begin this journey? And of course, mention me just because every time someone talks to a sponsor and brings my name up, the reflexive wince is just my favorite look in the world.

Guy: Yeah, so definitely mentioned Corey’s name. [laugh]. We have online demos where people can play with the lab, you go to extrahop.com/demo. We also offer AWS trials if you want to actually deploy one and see what it looks like in your environment for a period of time. And we have teams all over the world, from the United States, EMEA, APACs, that are happy to help answer questions, help deploy, and help automate a lot of this, whether it be through something like a CloudFormation template, or Terraform scripts, whatever infrastructure as code language you choose to use.

Corey: Excellent. Thank you so much for taking the time to speak with me today. I really do appreciate it.

Guy: Yeah, Corey, it’s been a pleasure talking to you. And I’m looking forward to maybe having another one with you in the future.

Corey: Oh, I would expect so. I’m curious to see what happens next. Guy Raz, senior systems engineer at ExtraHop. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice and an insulting comment that will no doubt get flagged by ExtraHop as being something that shouldn’t be on the network.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Austin

Austin makes problems with computers, and sometimes solves them. He’s an open source maintainer, observability nerd, devops junkie, and poster. You can find him ignoring HN threads and making dumb jokes on Twitter. He wrote a book about distributed tracing, taught some college courses, streams on Twitch, and also ran a DevOps conference in Animal Crossing.

Links:

  • Lightstep: https://lightstep.com/
  • Lightstep Sandbox: https://lightstep.com/sandbox
  • Desert Island DevOps: https://desertedislanddevops.com
  • lastweekinAWS.com Resources: https://lastweekinAWS.com/resources
  • Distributed Tracing in Practice: https://www.amazon.com/Distributed-Tracing-Practice-Instrumenting-Microservices/dp/1492056634
  • Twitter: https://twitter.com/austinlparker
  • Personal Blog: https://aparker.io

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at the Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Thinkst. This is going to take a minute to explain, so bear with me. I linked against an early version of their tool, canarytokens.org in the very early days of my newsletter, and what it does is relatively simple and straightforward. It winds up embedding credentials, files, that sort of thing in various parts of your environment, wherever you want to; it gives you fake AWS API credentials, for example. And the only thing that these things do is alert you whenever someone attempts to use those things. It’s an awesome approach. I’ve used something similar for years. Check them out. But wait, there’s more. They also have an enterprise option that you should be very much aware of canary.tools. You can take a look at this, but what it does is it provides an enterprise approach to drive these things throughout your entire environment. You can get a physical device that hangs out on your network and impersonates whatever you want to. When it gets Nmap scanned, or someone attempts to log into it, or access files on it, you get instant alerts. It’s awesome. If you don’t do something like this, you’re likely to find out that you’ve gotten breached, the hard way. Take a look at this. It’s one of those few things that I look at and say, “Wow, that is an amazing idea. I love it.” That’s canarytokens.org and canary.tools. The first one is free. The second one is enterprise-y. Take a look.
I’m a big fan of this. More from them in the coming weeks.

Corey: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the Enterprise (not the starship). On-prem security doesn’t translate well to cloud or multi-cloud environments, and that’s not even counting IoT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IoT devices, detects these threats up to 35 percent faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at extrahop.com/trial.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’m joined this week by Austin Parker, who’s a principal developer advocate at Lightstep. Austin, welcome to the show.

Austin: Hey, it’s great to be here.

Corey: It really is. I love coming here. It’s one of my favorite places to go. So, let’s get the obvious stuff out of the way. You’re a principal developer advocate at Lightstep. I know this because I said it a whole sentence ago, which is about the limit of my attention span. What is Lightstep? And what does your job mean?

Austin: So, Lightstep is an observability platform. We take traces, and metrics, and logs, and all that good stuff, throw them together in a big old swamp of data, and then, kind of, give you some really cool workflows to help you make sense of it, figure out, hey, where is the slow SQL query? Where is the performance bad?

Corey: The way to figure out, in most of my environments, where’s the performance bad is git blame, figure out what part I wrote.

Austin: But imagine there were, like, 1000, or 100,000 of you all working on this massive distributed system, and you didn’t know half—

Corey: It would snark itself to death before it ever got off the ground.

Austin: Yeah. I mean, I think that’s actually most large companies, right? We deliver shippable software only through inertia.

Corey: Yeah. Just because at some point, it bounces off all the walls, there’s nowhere else for it to go but to production.

Austin: Yep. But yeah, you have thousands of people, hundreds of people, however many people, right? I think the whole distributed workforce thing that most people are dealing with now has really made observability rise to the top of your concern list because you don’t have the luxury of just going and poking your head around the corner and saying, “Hey, Joanne. What the heck? Why did things break?” You can’t just poke someone anymore. Or you can, but you never know what you’re going to have to deal with.

Corey: It feels weird to call them at home or bug their family members to poke them or whatnot. It just seems weird.

Austin: It does. And until Amazon comes out with a minder drone that just, kind of like, hovers over your shoulder at all times, and pokes you, when someone is like, “Hey, you broke the build.” Then I think we’re going to need observability so that people can sort of self-serve, figure out what’s going on with their systems.

Corey: Cool. One of the things I’m going to point out is that I’ve had a bunch of people attempt to explain what distributed tracing is and how observability works, and it never really stuck. And one of the things that I found that did help explain it—and we didn’t even talk about this in the pre-show, while we figure out how to pronounce each other’s names—but one of the things that has always stuck with me is the interactive sandbox on Lightstep, which used to be prominently featured on your page; now it’s buried in the menu somewhere. But it’s an interactive sandbox that sets up a scenario, problem you’re trying to solve, gives you data—so it gets away from the problem of, “Step one, have a distributed application where it’s all instrumented and reporting things in.” Because in a lot of shops, that’s not exactly a small lift that you can do in an afternoon to start testing things like this out. It’s genius. It shows what the product does, how it works, mapped to the type of problems people will generally encounter. And after I played with this, “Oh, my stars, I get it.”

Austin: We actually just recently updated that to add some new stuff to it because we shipped a feature called ‘Change Intelligence’ where you can take actual time-series metrics, and then overlay those on traces and say, “Hey, I saw a weird spike,” and highlight that, and then we go through, look at all the traces for that service and its related services during that time, and tell you, “Hey, we think it might be this. Here’s things that are highly correlated in those time windows.” So, if you haven’t checked it out recently, go back and check it out. It’s—yeah, a little more hidden than it used to be, but I believe you can find it at lightstep.com/sandbox.

Corey: Yeah. And there’s no sign up to do this. It’s free access. It asked for an email address, but that’s okay, I just use yours. No, I’m not kidding. I actually did. And, yeah, it works; it shows exactly what it is. It even has, instead of ‘start’ it says ‘play’ because that’s fundamentally what it is. If you’re trying to wrap your head around distributed tracing, take a look at this.

Austin: Yes, definitely. I have a long-standing Jira ticket to add achievements to that.

Corey: Oh, that could be fun. You could bury some, too, like misusing services as databases—

Austin: Ooh.

Corey: —or most expensive query to get the right answer.

Austin: Yeah. And then maybe, like, there’s just one span, kind of, hidden there where it’s ‘using Route 53 as a database.’

Corey: I keep seeing that cropping up more and more places. That’s something I get to own and that’s an awful lot of fun. Speaking of gamification and playing in strange ways, one of the things you did last year that I wasn’t paying attention to—because, you know, there was a pandemic on—was you were one of the organizers behind Desert Island DevOps which is a strange thing that I’ve only recently delved into—delven into—gone spelunking inside of. There we go.

It wasn’t instrumented for observability—buh-dum-tss. But it’s fundamentally a DevOpsDays that takes place inside the animated world of Animal Crossing’s New Horizon, which is apparently a Nintendo game, which is apparently a game company.

Austin: Yeah.

Corey: It is not really my space. I don’t want to misspeak.

Austin: No, you hit it. ‘Deserted.’ Deserted Island [crosstalk 00:05:43].

Corey: Oh, ‘Deserted.’ Ah, got it. And don’t spell it as ‘dessert’ either, as in this would be a delicious game to play.

Austin: I mean, it is a delicious and comforting sort of experience. If you aren’t familiar with Animal Crossing, the short 30-second explanation is it is a life simulator, building game where, you as your character, you are on an island, and there are relatively adorable animal NPCs that are your villagers, and you can talk to them, and they will say funny things to you. You can go around and do chores like picking up fruit or fishing. And the purpose is, kind of, do these chores, get some in-game currency, and then go spend that in-game currency on furniture so that you can make a pretty house, or buy pretty clothing. And it came out at a perfect time last year because everyone was about to bundle inside for the—well, we’re still inside—but everyone had to go inside. And suddenly, here’s this like, “Oh, it’s just this cute, sort of like, putz around and do whatever.”

Corey: It was community-oriented. It was more of a building-oriented game than a destruction game.

Austin: Yeah.

Corey: It’s the sort of thing that is a great way of taking your mind off your troubles. It is accessible to a bunch of people that aren’t generally perceived as gamers when you think of that subculture. It really is an encompassing, warm, wonderful thing—by all accounts—and you looked at it and figured, “All right, how can we ruin something?” And the correct answer you got to is, “Let’s pour DevOps on it.”

Austin: Yeah. Let’s use this as an event platform, and let’s really just tech-bro this shit up.

Corey: And it seems to work super well. At the time of this recording, I have submitted a talk that I live-streamed my submission around, and I have not heard in either direction. To be perfectly frank, I forget what I wound up submitting, which is always a bit of a challenge, just because I make so many throwaway random jokes that, cool. Well, we’ll see how it plays out. I think you were even in the audience for that on the Twitch stream.

Austin: Yeah. You found some bugs on the CFP form [laugh] that I had to fix.

Corey: To be clear, the reason I do those things is not because it’s a look how clever I am, but rather to instead talk about how it’s not scary to submit a talk proposal. Everyone has a story that they can tell. And you don’t need a big platform or decades of experience in this space to tell a story. And that was my goal, and I think I succeeded. You would have the numbers more than I do; I hope people wound up submitting based upon seeing that. I want to hear voices that, frankly, aren’t ours all the time.

Austin: I think in, like, a week, we basically got more submissions than we did for the entire CFP last year. One thing that I kind of think is interesting to bring up because you bring up, oh, we don’t hear a variety of voices, right? One thing I tell people, and I know that it’s not universally applicable advice, but I got into DevRel as a—not quite luck, but, like, everything in my life is luck, on some level. It always plays some level of importance. But I didn’t go to school to get into DevRel, I didn’t do a lot of things.

I have actually been in tech, maybe—depending on how you want to count it—in terms of actually being in a software development job or primarily software development job, maybe, like, five or six years, give or take. And before that, I did a lot of stuff. I was a short-order cook; I worked at gas stations; I did tech support for Blackberry, and I did a lot of community organization. I was a union organizer for a little while. I like DevRel because it’s like, oh, this kind of integrates a lot of things I’m interested in, right?

I enjoy teaching, helping people, and helping people learn, but I also like talking; I like to go and be a public figure, and I like to build a platform and use that to get a message out. And I think what I did with Deserted Island, or what the impetus there was, we suddenly were in a situation where it’s like, “Hey, there’s a bunch of people that normally get together and they fly around the globe in decent airplane seats, and people come and see us talk.” Because why? Because they think we know what we’re talking about, or because we have something that shows we know what we’re talking about, or however you want to say it. But in a lot of cases, I think people are coming for that sort of community, they’re coming because, “Hey, I can go to a room and I can sit in some weird little hotel, or conference center, or whatnot, and everyone I look at, everyone I see is someone that is doing what I’m doing, on some level. These are all people that are working in technology, they’re building things, they’re solving problems.”

And that goes away really quickly when you get into this remote-first world, and when we can’t travel, and we don’t have that visual aspect. So, what I wanted to do with Deserted Island, what I thought what was important about it is, I was already sick of Zoom by the time, everyone went to Zoom; I was already sick of the idea of, oh my god, a year or two years of these sort of events and these community things just being, like, everyone’s staring at a bunch of slides and a talking head. Didn’t sound very appealing, so what if we try something different? What if we do something where it’s like, look, we’re going to take people out of their day; we’re going to put them in somewhere else. And maybe that’s somewhere else is just, hey, you’re watching people run around on an Animal Crossing Island on a Twitch stream.

But that sort of moment of just, like, this isn’t what you would normally be doing, I think takes people’s heads out of their normal routine and puts them in a place where they can learn, and they can feel community, and they can feel, like, a kinship. I also think it’s really important because it’s that whole stupid New Yorker joke of, “On the internet, nobody knows you’re a dog.” We have this really cool opportunity to craft who we are as people, and how we present that to the world. And for a lot of people, you’re stuck inside; you don’t get that self-expression, so here’s a way to be expressive, right? Here’s a way to communicate who you are on a level that isn’t just a profile picture or something, or things that don’t work as well over Zoom.

It’s a way to help project your identity. And that, I think, gives more weight to what you’re saying because when you feel like, “Hey, this is more of who I am,” or, “This is a representation of me. I can show something about who I am.” And that helps you speak. And that helps you deliver, I think, an effective talk. And that, again, builds community and builds these bonds.

Corey: I want to talk to you about that, specifically because you are one of those people that aligns very much with my view of the world on developer marketing. But I don’t want to lead you too much on this, so why don’t you start? Take it away. Where do you stand on developer marketing? And what do people get wrong?

Austin: I think the thing that a lot of people get wrong is that they try to monetize the idea of community. If you go and you search, insert major company name here; you search “Amazon community,” or you search “Microsoft community,” or you search “Google community,”—well, if you do that, you’ll get no results, but whatever, right? You get the picture that marketers in a way have turned the idea of developer community into something that you can just throw a KPI or throw an OKR on and squeeze it for money. And I don’t like that. I’m not very comfortable with that idea of community—because I think community in a lot of ways, it’s like family. And the families that you like the best are the ones you choose. I think this is—

Corey: The family you choose is an important concept.

Austin: Right. And for the most part… so much of human experience activity is built around finding those people you choose, and those communities develop out of that. I use AWS sometimes, I don’t necessarily know if I would put myself in a community with every other AWS user. I—

Corey: Oh, I certainly wouldn’t. This is the problem. Everyone thinks when you talk about community or a group of people
doing something, they’re ‘other people’ that are in some level of otherness. And that’s—like there are entire communities around AWS that I do not talk to, I do not see, I do not pretend to understand.

Austin: Yeah, even at Lightstep. We’re not a massive, massive company by any means, but we have a bunch of different users that are using our tool in different ways. And they all have different needs, and they all have different wants. So, I could say, “Oh, here’s the Lightstep community.” But it’s not a useful abstraction.

It’s not a useful way to abstract all of our users because any tool that’s worth using is going to be this collection of other abstractions and building blocks. Like, you… I don’t know, look at something like Notion, or look at something like Airtable, or the popularity of low or no-code stuff, where someone built a platform and then other people are building stuff on top of that platform, if you go to those user groups or you go to those forums, and it’s just like, there’s a million, million different varied use cases, and people are doing it in different ways, and some people are building this kind of application, or that kind of application, or whatever. So, the idea of, oh, there’s a community and we can monetize that community somehow, I’m uncomfortable with that from, sort of, a base level. And I’m uncomfortable with the idea of the DevRel industry—or the developer marketing industry—kind of moving towards this idea of, like, we’re going to become community marketers or whatever. I think you have to approach people as individuals.

And individuals are motivated by a lot of things. They’re motivated by, can you solve this problem? Do I like you? Are you funny? Whatever. And I believe that if you’re a developer tool, and you are trying to attract developers, then [sigh] it works a lot better, I think, to have just individuals, to have people that can help influence the much broader—the superset of all developers that might have an interest in what you’re doing by being different, I guess.

Being something that’s like, hey, this is entertaining, or this is informative, or this is interesting. The world is not a meritocracy. The world is governed by many, many different things. You’re not going to win over the developer industry simply by going out and having the best white papers, or having one more ad read than your competitor. You need to do something to get people interested and excited in [sigh] a way that they can see themselves using it.

It’s like, why did Apple go and do ‘Think Different’ ads? Because it’s like, you using a Mac, that’s kind of like being Einstein, or that’s kind of like being Picasso. This is basic marketing stuff that I feel like a lot of technical marketers or developer marketers sort of leave at the door because they think the audience is too sophisticated for it, or their—

Corey: I’ll even soft-launch it here because I haven’t at this point in time, talked about it in public, but if you go to lastweekinAWS.com/resources we wrote our own developer marketing guide because I got tired of explaining the same type of thing again, and again, and again. It asks for an email address and it sends it to you—I know, I’m as guilty as any. And I, of course, called it ‘Devreloper,’ which is absolutely a problem with me and I talk about things. But I’m right.

And it goes to an awful lot of what you’re saying. An example that you just talked about of giving people something rather than trying to treat them as metrics, one of the best marketing things I’ve seen you do, for example, is you wrote O’Reilly’s Distributed Tracing in Practice which means if someone has a question about distributed tracing and how it’s supposed to work, well, that’s not a half-bad resource. And okay, I’ve read it and I have some further questions. Let me track down the author and ask them. Oh, you work at a company that is in this space? Huh. Maybe I’ll look into this. And it’s a very long-tail story. And how do you attribute that as far as, did this lead come from someone who read your book or not will drive
marketers crazy.

Austin: Oh, it’s super hard. And it does drive them crazy. [laugh].

Corey: Yeah, my answer is, I don’t know and I don’t care. One of the early sponsors of this podcast sponsored for a month and then didn’t continue because they saw no value. A month goes by, they bought out everything that held still long enough, and, “Thank you for your business.” “Can you explain to me what changed?” “Oh, we talked to some of our big customers and it turned out the two of them had heard about us for the first time on your show.”

And that inspired them to start digging into it and reaching things out, but big companies, corporate games of telephone, there was no way to attribute that. My firm belief is, on some level, that if you get in front of an audience with a message that resonates and—and this is the part some people miss—is something that solves an actual problem that they have. It works. It’s not necessarily predictable and it’s hard to say that this thing is going to go big and this thing isn’t. So, the solution, on some level is just keep publishing things that speak to your audience. But it works, long term. I’m living proof of
this.

Austin: Yeah. I think that it makes a lot more sense to… rather than to do, sort of, I don’t want to say vanity metrics, but kind of vanity metrics around, like, oh, this many stars, or this many forks, or whatever. There’s a lot of people, especially in this OSS proximate world. Where you have a lot of businesses that are implicitly or explicitly built on top of an open-source project, not everyone that is using your open-source project is going to, one, be capable of converting into a paid user, or two, be super interested in it. And I would rather spend time thinking about, well, what is the value someone gets out of this product?

And even if that only thing is, is that, hey, we know what we’re talking about because we’ve got a bunch of really smart people that are building this product that would solve their problem. If you want to go out and build your own internal observability solution using completely open-source tools Grafanas and Prometheuses of the world, great. Go for it. I’m not going to hold you back. And for a lot of people, if they come to me and say, “Well, this is what we got, and this what we’re
thinking about.”

I’ll say, “Yeah. Go for it. You don’t need what we’re offering.” But I can guarantee you that as it scales and as it grows, then you’re going to have a moment where you have to ask yourself the question of, “Do I want to keep spending a bunch of time stitching together all these different data sources, and care and feeding of these databases, and this long term storage, and dealing with requests from end-users, or I just want to pay someone else to solve that problem for me? And if I’m going to pay someone else, shouldn’t I pay the people who literally spend all day every day thinking about these problems and have had decades of experience solving these problems at really big companies that have a lot of time and effort to invest in this?”

Corey: This episode is sponsored in part by our friends at Lumigo. If you’ve built anything from serverless, you know that if there’s one thing that can be said universally about these applications, it’s that it turns every outage into a murder mystery. Lumigo helps make sense of all of the various functions that wind up tying together to build applications. It offers one-click distributed tracing so you can effortlessly find and fix issues in your serverless and microservices environment. You’ve created more problems for yourself; make one of them go away. To learn more, visit lumigo.io.

Corey: Oh, yeah. We’re doing some new content experiments on our site, and what we’re doing is we’re having some folks write content for us. Now, when people hear that, what a lot of marketers will immediately do is dive down the path of, “Ah. I’m going to go ahead and hire some content farm.” Well, that doesn’t work, I found that we wound up working with individual people that work super well.

And these are people who are able to talk about these things because their day job is managing a team of 30 SREs or something like that, where they are very clearly experts in the space. And I want to be very clear, I’m not claiming credit for our content writers; they get their own bylines on these things.

Austin: Yeah.

Corey: And it turns out that that, over time, leads to good outcomes because it helps people what they need. There’s the mystical SEO Juju that I don’t pretend to understand, but okay, I’m told it’s important, so fine, whatever. And it makes for an easier onboarding story, where there are now resources that I can trust and edit if I need to, as things change, that I can point people to, that isn’t a rotating selection of sketchy sites.

Austin: Mm-hm. I think that’s one thing that I would love to see more of, just not in any one particular part of the tech industry, but overall, the one thing I’ve noticed, at least in the pandemic, during this whole work-from-home, whatever, whatever, we don’t talk enough. And it sounds maybe weird, but I think this actually goes back to what you’re saying earlier, about everyone having a story to tell. People don’t feel comfortable, I think, putting their opinion out there or saying, “Hey, this is what worked. This is what didn’t work.”

And so if you want to go find that out—like, if I wanted to go write something about, hey, these are the five things you should do to ensure you have great observability, then that’s going to involve a lot of me going around and sort of Sherlocking my way through StackOverflow posts, and forums, and reaching out to people individually for stories and comments and whatever. And I would love to see us get to a point where we’re just like, “Actually, no. This isn’t—we should just be sharing this. Let’s write blogs about it.” If you’re sitting there thinking no one’s going to find this useful, right—like, you solve a problem, or you see something that could have worked better, and you’re like, “Eh, no one else is going to find that valuable.”

I can almost guarantee you that someone is going to find that valuable. Maybe not today, maybe not tomorrow, but go ahead and write about your experiences, write about the problems you’ve solved, write about the things that have vexed you, and put that on the internet because it’s really easy to publish stuff on the internet.

Corey: Yes. Which is a blessing and a curse. That is very much a double-edged sword.

Austin: That very much is a double-edged sword. But I think that by biasing towards being more open, by biasing towards
transparency and sharing what works, what doesn’t work, and having that just kind of be the default state, I’m a big proponent of things like radical transparency in terms of incident reports, or outages, or hiring, or anything. The more information that you can put in the world is going to—it might not make it better, but it at least helps change the conversation, gives more data points. There was a whole blow-up on Twitter this week, where someone posted like, “Hey, this is a salary I’m looking for.” I think you—

Corey: Oh, yeah. She’s great.

Austin: Yeah, she’s worth it, right? And the thing that got everyone’s bee in a bonnet was, like, she’s saying, “Oh, I want $185k.” And it’s like, “Well, why don’t we just publish that information?” Why isn’t everyone just very open and honest about their salary expectations? And I know why: because the paucity of information is a benefit to employers and it works against employees.

There was a lady that left—gosh, where was it? [sigh] I forget the company, but she left because she found out she was systemically underpaid compared to their male peers. Having these sort of information imbalances don’t really help the people at the bottom of the pyramid. They don’t help the little guys. They really only help the people that are in the very large companies with a lot of clout and ability to control narratives.

And they want it to stay that way; they don’t necessarily want you to know what everyone’s salary is because then it gives you, as someone trying to get a job, a better negotiating position because you know what someone with your level of experience is worth to them.

Corey: It’s important to understand the context behind these salary negotiations and how to go about getting interviews and the rest. The entire job-hunting process is heavily biased in favor of employers because, especially at large employers, they go through this multiple times a week, whereas we go through this, as employees basically, every time we change jobs. Which for most people is every couple of years and for me, because of my mouth, it’s every three weeks.

Austin: Yeah. I’m not saying it’s a simple solution. I am advocating for, sort of, societal, or just cultural shifts, but I think that it all comes full circle in the sense that, hey, a big part of observability is the idea that you need to be able to ask arbitrary questions. You want to know about unknown unknowns. And maybe that’s why I like it so much as a field, why I like tracing, why I like this idea.

Because, yeah, a lot of things in the world would be interesting, and different, and maybe more equitable if we did have more observability about not just, hey, I use Kafka, I use these parameters on it, and that gives me better throughput, but what if you had observability for how HR runs? What if you had observability for how hiring is done? And that was something that you could see outside of the organization as well. What if we shared all this stuff more, and more, and more, and we treated a few less things as trade secrets? I don’t know if that’s ever going to happen in my lifetime, but it’s
my default position. Let’s share more rather than less.

Corey: Yes, absolutely. Especially those of us with inordinate amounts of privilege. And that privilege takes different forms; there’s the usual stuff people are talking about in terms of the fact that we are over-represented in tech in many respects, but there are other forms of privilege, too. There’s a privilege that comes with seniority in the space, there is a privilege with being a published author, in your case, there is privilege in having a broad audience, like I do. And it just becomes this incredibly nuanced story.

The easiest part of it to lose sight of—at least for me—is I tell stories about what has worked for me and how I’ve done what I do, and I have to be constantly conscious of the fact that there is that privilege baked in and call it out where I can. I’ve gotten much better at that, but it’s an ongoing process. Because what works for me does not work for other people
across a wide variety of different axes. And I don’t want people to feel bad based upon what I say.

Austin: Oh, yeah, absolutely. I mean, I’m in the same boat. Like, I tend to be very irreverent and/or shitpost-y and I don’t have much of an explanation other than, I learned at some point in my life, that it’s just… [sigh] I would rather go through life shitposting on Twitter, rather than be employable. It’s just who I am. There’s—I’m sure some people think I come off as rude. I don’t know. I also agree, you’d never punch down. You only punch up. But you never know how other people are going to take that, and I don’t think that it always gets interpreted in the spirit it was meant. And I can always do better, right?

Corey: As can we all. The hard part for, I think, a lot of us is to suppress that initial flash of defensiveness when someone says you didn’t quite get there, and learn from the experience. One of the ways I do that, personally, is I walk away before responding, sometimes. I want to be a better version of myself, but when I get called out of—like, this tweet thread is the whitest thing I’ve seen since I redid my bathroom walls, and I get a flash of defensiveness, “Excuse me. That’s not accurate.”

And… and then I stop and I think, and then sanity prevails, where it’s, yeah. There’s a lot of privilege baked into my existence, and if I don’t see it, that doesn’t mean it’s not there. I have made it a firm rule of not responding defensively to things like that, ever. And there are times when I get called out for aspects of how I present that I don’t believe are justified, to be very honest. But that is a me thing; that is not them, and I welcome the feedback, regardless. If you make people feel like a jerk for giving you feedback, they stop giving you feedback. And then where are you?

Austin: Yeah. Funny anecdote. I wrote a blog for my personal blog a little while ago about, oh, togetherness, community, something like that. But I wrote—the intro was something like, talking about why people love Sweet Caroline, right? Favorite song in the world.

Corey: [sings].

Austin: [joins in]. Yeah.

Corey: Yeah. I’m not allowed to play with that song here at The Duckbill Group because one of our employees is named Caroline and, firm rule: don’t make fun of people’s names. They’re sensitive about it, and let’s not kid ourselves here, I own the company. Even if she says, “It’s fine, I love it.” That doesn’t help because I own the company. There is a power imbalance here.

Austin: Yeah.

Corey: I don’t know that she would feel that she had the psychological safety to say, “That’s not funny.” I absolutely hope she would because that’s the culture that I spend significant effort on building, but I can’t depend on that. So, I don’t go down the path of making those jokes. But I—yes, I love the intro to the song. Please continue.

Austin: It’s great. Everyone loves it. So, the intro of my initial paragraph was ruminating on that. And this post went around enough that it got submitted to Hacker News a few times, and the only comment it got was some mendacious busybody Hacker News type going on about why I would be so racist against white people. [laugh]. And I was just like, “And this is why I don’t come to this website at all.”

Corey: Yeah. There are so many things on Twitter that are challenging and difficult and obnoxious, and it’s still the best thing we have for a sense of community. This has replaced IRC for me, to be perfectly honest.

Austin: Yeah. No, I used to be big on IRC, and then I left because [sigh], well, a couple reasons. One, I really liked being able to post gifs.

Corey: Yeah, that is something where the IRC experience is substandard. I was Freenode network staff for years—

Austin: Oh wow.

Corey: —and that was the thing to do. Now, turns out that the open-source dialogue and the community dialogue have shifted form. And I still hang out there periodically for specific things, but by and large, it’s not where the discourse is.

Austin: Yeah, it is interesting. It’s something that concerns me, kind of, in a long term sense about not only our identity but also, sort of, the actual organic communities we formed, we’ve put on to these extremely unaccountable privately held platforms whose goal is monetization and growth so that they can continue to make money. And for as much as anyone can rightfully say, “Hey, Twitter’s missed the mark,” a lot of times, it is a hard balance to strike. They don’t have simple questions to answer, and I don’t necessarily know if the nuance of their solutions has really risen to the challenge of answering those well, but it’s a hard thing for them to do. That said, I think we’re in a really awkward position where suddenly you’ve got the world’s collection of open-source software is being hosted on a platform that is run by Microsoft, and I am old enough to remember. “Embrace, extend, extinguish.”

Corey: Oh, yeah. I made an entire personality out of hating Microsoft.

Austin: Yeah. And I mean, a lot of people still do. I read MacRumors sometimes, and they’re all posting there still. Or Slashdot.

Corey: I wondered where they’d gone. I didn’t think everyone had changed their mind.

Austin: I had just a very out-of-body moment yesterday because someone replied to a comment on mine about Slashdot on it, and then the Slashdot Twitter account liked it. And there exists a photo of me from when I was a teenager, where I owned a Slashdot ballcap. And that picture is somewhere in the world. Probably not on the internet, though, for very good reason.

Corey: I’m mostly just still reeling at the discovery that there’s a Slashdot Twitter account. But I guess time does evolve.

Austin: It does. It makes fools of us all.

Corey: It really does. Well, Austin, thank you so much for taking the time to speak with me. If people want to learn more about what you’re up to, how you view the world, et cetera, et cetera, et cetera. Where can they find you?

Austin: So, you can find me on Twitter, mostly, at @austinlparker. You can find my blog with various musings that is updated frequently at aparker.io and you can learn more about Deserted Island DevOps 2021, coming on April 30th this year, at desertedislanddevops.com.

Corey: Excellent. And we will put links to all of that in the [show notes 00:34:01]. Thank you so much for taking the time to speak with me. I appreciate it.

Austin: Thank you for having me. This was a lot of fun.

Corey: It really was. Austin Parker, principal developer advocate at Lightstep. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice and then a giant series of comments that all reference one another and then completely lose track of how they all interrelate and be unable to diagnose performance issues.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

This has been a HumblePod production. Stay humble.

View Details

About Donnie

Donnie is VP of Products at Docker and leads product vision and strategy. He manages a holistic products team including product management, product design, documentation & analytics. Before joining Docker, Donnie was an executive in residence at Scale Venture Partners and VP of IT Service Delivery at CWT leading the DevOps transformation. Prior to those roles, he led a global team at 451 Research (acquired by S&P Global Market Intelligence), advised startups and Global 2000 enterprises at RedMonk and led more than 250 open-source contributors at Gentoo Linux. Donnie holds a Ph.D. in biochemistry and biophysics from Oregon State University, where he specialized in computational structural biology, and dual B.S. and B.A. degrees in biochemistry and chemistry from the University of Richmond.

Links:

  • Docker: https://www.docker.com/
  • Twitter: https://twitter.com/dberkholz

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Thinkst. This is going to take a minute to explain, so bear with me. I linked against an early version of their tool, canarytokens.org in the very early days of my newsletter, and what it does is relatively simple and straightforward. It winds up embedding credentials, files, that sort of thing in various parts of your environment, wherever you want to; it gives you fake AWS API credentials, for example. And the only thing that these things do is alert you whenever someone attempts to use those things. It’s an awesome approach. I’ve used something similar for years. Check them out. But wait, there’s more. They also have an enterprise option that you should be very much aware of canary.tools. You can take a look at this, but what it does is it provides an enterprise approach to drive these things throughout your entire environment. You can get a physical device that hangs out on your network and impersonates whatever you want to. When it gets Nmap scanned, or someone attempts to log into it, or access files on it, you get instant alerts. It’s awesome. If you don’t do something like this, you’re likely to find out that you’ve gotten breached, the hard way. Take a look at this. It’s one of those few things that I look at and say, “Wow, that is an amazing idea. I love it.” That’s canarytokens.org and canary.tools. The first one is free. The second one is enterprise-y. Take a look. I’m a big fan of this. More from them in the coming weeks.

Corey: This episode is sponsored in part by our friends at Lumigo. If you’ve built anything from serverless, you know that if there’s one thing that can be said universally about these applications, it’s that it turns every outage into a murder mystery. Lumigo helps make sense of all of the various functions that wind up tying together to build applications. It offers one-click distributed tracing so you can effortlessly find and fix issues in your serverless and microservices environment. You’ve created more problems for yourself; make one of them go away. To learn more, visit lumigo.io.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Today I’m joined by Donnie Berkholz, who’s here to talk about his role as the VP of Products at Docker, whether he knows it or not. Donnie, welcome to the show.

Donnie: Thanks. I’m excited to be here.

Corey: So, the burning question that I have that inspired me to reach out to you is fundamentally, and very bluntly and directly, Docker was a thing in, I want to say the 2015-ish era, where there was someone who gave a parody talk for five minutes where they got up and said nothing but the word Docker over and over again, in a bunch of different tones, and everyone laughed because it seemed like, for a while, that was what a lot of tech conference talks were about 50% of the way. It’s years later, now, and it’s 2021 as of the time of this recording. How is Docker relevant today?

Donnie: Great question. And I think one that a lot of people are wondering about. The way that I think about it, and the reason that I joined Docker, about six months back now, was, I saw the same thing you did in the early 2010s, 2013 to 2016 or so. Docker was a brand new tool, beloved of developers and DevOps engineers everywhere. And they took that, gained the traction of millions of people, and tried to pivot really hard into taking that bottom-up open-source traction and turning it into a top-down, kind of, sell to the CIO and the VP operations, orchestration management, kind of classic big company approach. And that approach never really took off to the extent that would let Docker become an explosive success commercially in the same way that it did across the open-source community and building out the usability of containers as a concept.

Now, new Docker, as of November 2019, divested all of the top-down operations production environment stuff to Mirantis and took a look at what else there was. And the executive staff at the time, the investors thought there might be something in there, it’s worth making a bet on the developer-facing parts of Docker to see if the things that built the developer love in the first place were commercially viable as well. And so looking through that we had things left like Docker Hub, Docker Engine, things like Notary, and Docker Desktop. So, a lot of the direct tools that developers use on a daily basis to get their jobs done when they’re working on modern applications, whether that’s twelve-factor, whether that’s something they’re trying to lift and shift into a container, whatever it might look like, it’s still used every day. And so the thought was, there might be something in here.

Let’s invest some money, let’s invest some time and see what we can make of it because it feels promising. And fast-forward a couple of years—we’re in early 2021—we just announced our Series B investment because the past year has shown that there’s something real there. People are using Docker heavily; people are willing to pay for it, and where we’re going with it is much higher level than just containers or just a registry. I think there’s a lot more opportunity
there. When I was watching the market as a whole drifting toward Kubernetes, what you can see is, to me, it’s a lot like a repeat of the old OpenStack days where you’ve got tons of vendors in the space, it’s extremely crowded, everybody’s trying to sell the same thing to the same small set of early adopters who are ready for it.

Whereas if you look at the developer side of containers, it’s very sparsely populated. Nobody’s gone hard after developers in a bottom-up self-service kind of way and helped them adopt containers and helped them be more productive doing so. So, I saw that as a really compelling opportunity and one where I feel like we’ve got a lot of runway ahead of us.

Corey: Back in the early days—this is a bit of a history lesson that I’m sure you’re aware of, but I want to make sure that my understanding winds up aligning with yours is, Docker was transformative when it was announced—I want to say 2012, in Santa Clara, but don’t quote me on that one—and, effectively, what it promised to solve was—I mean, containerization was not a new idea. We had that with LPARs on mainframes way before my time. And it’s sort of iterated forward ever since. What it fundamentally solved was the tooling around those things where suddenly it got rid of the problem of, “Well, it worked on my machine.” And the rejoinder from the grumpy ops person—which I very much was—was, “Great. Then backup your email because your laptop’s about to go into production.”

By having containers, suddenly you have an environment or an application that was packaged inside of a mini-environment that was able to be run basically anywhere. And it was, write once, deploy basically as many times as you want. And over time, that became incredibly interesting, not just for developers, but also for folks who were trying to migrate applications. You can stuff basically anything into a container. Whether you should or not is a completely separate conversation that I am going to avoid by a wide margin. Am I right so far in everything that I have said there?

Donnie: Yep. Absolutely.

Corey: Awesome. So, then we have this container runtime that handles the packaging piece. And then people replaced Docker in that cherished position in their hearts—which is the thing that they talk about, even when you beg them to stop—with Kubernetes, which is effectively an orchestration system for containers, invariably Docker. And now people
are talking about that constantly and consistently. If we go back to looking at similar things in the ecosystem, people used to care tremendously about what distribution of Linux they ran.

And then—well, okay. If not the distro, definitely the OS wars of, is this Windows or is this a Linux workload? And as time has gone on, people care about that less and less where they just want the application to work; they don’t care what it’s running in under the hood. And it feels that the container runtime has gotten to that point as well. And soon, my belief is that we’re going to see the orchestrator slip below that surface level of awareness of things people have to care about, if for no other reason than if you look at Kubernetes today, it is fiendishly complicated, and that doesn’t usually last very long in this space before there’s an abstraction layer built that compresses all of that into something you don’t really have to think about, except for a small number of people at very specific companies. Does that in any way change, I guess, the relevance of Docker to developers today? Or am I thinking about this the wrong way with viewing Docker as a pure technology, instead of an ecosystem?

Donnie: I think it changes the relevance of Docker much more to platform teams and DevOps teams—as much as I wish that wasn’t a word or a term—operations groups that are running the Kubernetes environments, or that are running applications at scale in production, where maybe in the early days, they would run Docker directly in prod, then they moved to running Docker as a container runtime within Kubernetes, and more recently, the core of Docker—which was containerd—as a replacement for that overall Docker, which used dockershim. So, I think the change here is really around, what does that production environment look like? And where we’re really focusing our effort is much more on the developer experience. I think that’s where Docker found its magic in the first place was in taking incredibly complicated technologies and making them really easy in a way that developers love to use. So, we continue to invest much more on the developer tools part of it, rather than what does the shape of the production environment look like?

And how do we horizontally scale this to hundreds or thousands of containers? Not interesting problems for us right now. We’re much more looking at things like how do we keep it simple for developers so they can focus on a simple application. But it is an application and not just a container, so we’re still thinking of moving to things that developers care about. They don’t necessarily care about containers; they care about their app.

So, what’s the shape of that app, and how does it fit into the structure of containers? In some cases, it’s a single container, in some cases, it’s multiple containers. And that’s where we’ve seen Docker Compose pick up as a hugely popular technology. When we look at our own surveys, when we look at external surveys, we see on the order of two-thirds of people who use Docker using Compose to do it, either for ease of automation and reproducibility or for ease of managing an application that spans across multiple containers as a logical service, rather than try and shove it all in
one and hope it sticks.

Corey: I used to be relatively, I guess, cynical about Docker. In fact, one of my first breakout talks started life as a lightning talk called “Heresy in the Church of Docker,” where I just came up with a list of a few things that were challenging and didn’t fully understand. It was mostly jokes, and the first half of it was set to the backstory of an embarrassing chocolate coffee explosion that a boss of mine once had. And that was great. Like, what’s the story here? What’s the relevance? Just a story of someone who didn’t understand their failure modes of containers in
production. Cue laugh.

And that was great. And someone came up to me and said, “Hey, can you give the full version of that talk at ContainerCon?” To which my response was, “There’s a full version?” Followed immediately by, “Absolutely.” And it sort of took life from there.

Now, I want to say that talk hasn’t aged super well because everything that I highlighted in that talk has since been fixed. I was just early and being snarky, and I genuinely, when I gave that first version, didn’t understand the answers. And I was expecting to be corrected vociferously by an awful lot of folks. Instead, it was, “Yeah, these are challenges.” At which point I realized, “Holy crap, maybe everyone isn’t 80 years ahead of me in technical understanding.” And for
better or worse, it’s set an interesting tone.

Donnie: Absolutely. So, what do you think people really took out of that talk that surprised you?

Corey: The first thing that I think, from my perspective, that caught me by surprise was that people are looking at me as some sort of thought leader—their term, not mine—and my response was, “Holy crap. I’m not a thought leader. I’m just a loud, white guy in tech.” And yep, those are pretty much the same thing in some circles, which is its own series of problems. But further, people were looking at this and taking it seriously, as in, “Well, we do need to have some plans to mitigate this.”

And there are different discussions that went back and forth with folks coming up with various solutions to these things. And my first awareness, at least, that pointing out problems where you don’t know the answer is not always a terrible thing; it can be a useful thing as well. And it also—let me put a bit of a flag there as far as a point in time because looking back at that talk, it’s naive. I’ve done a bunch of things since then with Docker. I mean, today, I run Docker on my overpowered Mac to have a container that’s listening with our syslog.

And I have a bunch of devices around the house that are spitting out their logs there, so when things explode I have a rough idea of what happened. It solves weird problems. I wind up doing a number of deployment processes here for serverless nonsense via Docker. It has become this pervasive technology that if I were to take an absolutist stance that, “Oh, Docker is terrible. I’m never going to use Docker.”

It’s still here for me, and it’s still available and working. But I want to get back to something you said a minute ago because my use of Docker is very much the operations sysadmin-with-title-inflation whatever we’re calling them this week; that use case and that model. Who is Docker viewing as its customer today? Who as a company are you identifying as the people with the painful problem that you can solve?

Donnie: For us, it’s really about the developer, rather than the ops team. And specifically it’s about the development team. And this to me is a really important distinction because developers don’t work in isolation; developers collaborate together on a daily basis, and a lot of that collaboration is very poorly solved. You jump very quickly from, “I’m doing remote pairing in my code editor,” to, “It’s pushed to GitHub, and it’s now instantly rolling into my CI pipeline on its way to production.” There’s not a lot of intermediate ground there.

So, when we think about how developers are trying to build, share, and run modern applications, I think there’s a ton of whitespace in there. We’ve been sharing a bunch of experiments, for anybody who’s interested. We do community all-hands every couple of months where we share, here’s some of the things we’re working on. And importantly, to me, it’s focused on problems. Everything you were describing in that heresy talk was about problems that exist, and pointing out problems.

And those problems, for us, when we talk to developers using Docker, those problems form the core of our roadmap. The problems we hear the most often as the most frustrating and the most painful, guess what? Those are the things we’re going to focus on as great opportunities for us. And so we hear people talking about things like they’re using Docker, or they’re using containers, but they have a really hard time finding the good ones. And they can’t create good ones, they are just looking for more guidance, more prescription, more curation, to be able to figure out where’s this good stuff amidst the millions of containers out there? How do I find the ones that are worth using, for me as an individual, for me as a team, and for me as a company. I mean, all of those have different levels of requirements and
expectations associated with them.

Corey: One of the perceptions I’ve had of the DevOps movement—as someone who started off as a grumpy Linux systems administrator—is the sense that they’re trying to converge application developers with infrastructure engineers at some point. And I started off taking a very, “Oh, I’m not a developer. I don’t write code.” And then it was, “Huh. You know, I am writing an awful lot of configuration, often in something like Ruby or Python.” And of course, now it seems like everyone has converged as developers with the lingua franca of all development everywhere, which is, of course, YAML. Do you think there’s a divide between the ops folks and the application developers in 2021?

Donnie: You know, I think it’s a long journey. Back when I was at RedMonk, I wrote up a post talking about the way those roles were changing, the responsibilities were shifting over time. And you step back in time, and it was very much, you know, the developer owns the dev stack, the local stack, or if there’s a remote developer environment, they’re 100% responsible for it. And the ops team owned production, 100% responsible for everything in that stack. And over the past decade, that’s clearly been evolving.

They could still own their code in production and get the value out of understanding how that was used, the value of fast iteration cycles, without having to own it all, everywhere, all of the time, and have to focus their time on things that they had really no time or interest to spend it on. So, those things have both been happening to me, not in parallel, quite; I think DevOps in terms of ops learning development skillsets and applying those has been faster than development teams who were taking ownership for that full lifecycle and that iteration all the way to production, and then back around. Part of that is cultural in terms of what developer teams have been willing to do. Part of it is cultural in terms of what the old operations teams—now becoming platform engineering teams—have been willing to give up, and their willingness to sacrifice control. There’s always good times like PCI compliance, and how do you fight those sorts of battles.

And when I think about it, it’s been rotating. And first, we saw infrastructure teams, ops teams, take more ownership for being a platform, in a lot of cases, either guided by the emerging infrastructure automation config management tools like CFEngine back in the early 90s, which turned into Puppet and Chef, which turned into Ansible and Salt, which now continue to evolve beyond those. A lot of those enabled that rotation of responsibilities where infrastructure could be a platform rather than an ops team that had to take ownership of overall production. And that was really, to me, it was ops moving into a development mindset, and development capabilities, and development skillsets. Now, at the same time, development teams were starting to have the ability to take over ownership for their code running into production without having to take ownership over the full production stack and all the complexities involved in the hardware, and the data centers, and the colos, or the public cloud production environments, whatever they may be.

So, there’s a lot of barriers in the way, but to me, those have been all happening alongside, time-shifted a little bit. And then really, the core of it was as those two groups become increasingly similar in how they think and how they work, breaking down more of the silos in terms of how they collaborate effectively, and how they can help solve each other’s problems, instead of really being separate worlds.

Corey: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the Enterprise (not the starship). On-prem security doesn’t translate well to cloud or multi-cloud environments, and that’s not even counting IoT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IoT devices, detects these threats up to 35 percent faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at extrahop.com/trial.

Corey: Docker was always described as a DevOps tool. And well, “What is DevOps?” “Oh, it’s about breaking down the silos between developers and the operations folks.” Cool, great. Well, let’s try this. And I used to run DevOps teams. I know, I know, don’t email me. When you’re picking your battles, team naming is one of the last ones I try to get to.

But then we would, okay, I’m going to get this application that is in a container from development. Cool. It’s—don’t look inside of it, it’s just going to make you sad, but take these containers and put them into production and you can manage them regardless of what that application is actually doing. It felt like it wasn’t so much breaking down a wall, as it was giving a mechanism to hurl things over that wall. Is that just because I worked in terrible places with bad culture? If so, I don’t know that I’m very alone in that, but that’s what it felt like.

Donnie: It’s a good question. And I think there’s multiple pieces to that. It is important. I just was rereading the Team Topologies book the other day, which talks about the idea of a team API, and how do you interface with other teams as people as well as the products or platforms they’re supporting? And I think there’s a lot of value in having the ability to throw things over a wall—or down a pipeline; however you think about it—in a very automated way, rather than going off and filing a ticket with your friendly ITSM instance, and waiting for somebody else to take action based on that.

So, there’s a ton of value there. The other side of it, I think, is more of the consultative role, rather than the take work from another team and then go do another thing with it, and then pass it to the next team down and then so on, unto eternity. Which is really, how do you take the expertise across all those teams and bring it together to solve the problems when they affect a broader radius of groups. And so, that might be when you’re thinking about designing the next iteration of your application, you might want to have somebody with more infrastructure expertise in the room, depending on the problems you’re solving. You might want to have somebody who has a really deep understanding of your security requirements or compliance requirements if you’re redesigning an application that’s dealing with credit card data.

But all those are problems that you can’t solve in isolation; you have to solve them by breaking down the barriers. Because the alternative is you build it, and then you try and release it, and then you have a gatekeeper that holds up a big red flag, delays your release by six months so you can go back and fix all the crap you forgot to do in the first place.

Corey: While on the topic of being able to, I guess, use containers as sort of as these agnostic components, I suppose, and the effects that that has, I’d love to get your take on this idea that I see that’s relatively pervasive, which is, “I can build an application inside of containers”—and that is, let’s be clear, that is the way an awful lot of containers are being built today. If people are telling you otherwise, they’re wrong—“And then just run it in any environment. You’ve built an application that is completely cloud agnostic.” And what cloud you’re going to run it in today—or even your own data center—is purely a question of either, “What’s the cheapest one I can use today?” Or, “What is my mood this morning?” And you press a button and the application lives in that environment flawlessly, regardless of what that provider is.
Where do you stand on that, I guess, utopian vision?

Donnie: Yeah, I think it’s almost a dystopian vision, the way I think about it—which is the least common denominator approach to portability—limits your ability to focus on innovation rather than focusing on managing that portability layer. There are cases where it’s worth doing because you’re at significant risk, for some reason, of focusing on a specific portability platform versus another one, but the bulk of the time, to me, it’s about how do you focus your time
and effort where you can create value for your company? Your company doesn’t care about containers; your company doesn’t care about Kubernetes; your company cares about getting value to their customers more quickly. So, whatever it takes to do that, that’s where you should be focusing as much time and energy as possible. So, the container interface is one API of an application, one thing that enables you to take it to different places, but there’s lots of other ones as well.

I mean, no container runs in isolation. I think there’s some quote, I forget the author, but, “No human is an island” at this point. No container runs in isolation by itself. No group of containers do, either. They have dependencies, they have interactions, there’s always going to be a lot more to it, of how do you interact with other services?

How do you do so in a way that lets you get the most bang for your buck and focus on differentiation? And none of that is going to be from only using the barest possible infrastructure components and limiting yourself to something that feels like shared functionality across multiple cloud providers or multiple other platforms.

Corey: This gets into the sort of the battle of multi-cloud. My position has been that, first, there are a lot of vendors that try and push back against the idea of going all-in on one provider for a variety of reasons that aren’t necessarily ideal. But the transparent thing that I tend to see—or at least I believe that I see—is that well, if fundamentally, you wind up going all-in on a provider, an awful lot of third-party vendors will have nothing left to sell you. Whereas as long as you’re trying to split the difference and ride multiple horses at once, well, there’s a whole lot of painful problems in there that you can sell solutions to. That might be overly cynical, but it’s hard to see some stories like that.

Now, that’s often been misinterpreted as that I believe that you should always have every workload on a single provider of choice and that’s it. I don’t think that makes sense, either. I mean, I have my email system run in GSuite, which is part of Google Cloud, for whatever reason, and I don’t use Amazon’s offering for the same because I’m not nuts. Whereas my infrastructure does indeed live in AWS, but I also pay for GitHub as an example—which is also in the Azure business unit because of course it is—and different workloads live in different places. That’s a naive oversimplification, but in large companies, different workloads do live in different places.

Then you get into stories such as acquisitions of different divisions that are running in completely different providers. I don’t see any real reason to migrate those things, but I also don’t see a reason why you have to have single points of control that reach into all of those different application workloads at the same time. Maybe I’m oversimplifying, and I’m not seeing a whole subset of the world. Curious to hear where you stand on that one?

Donnie: Yeah, it’s an interesting one. I definitely see a lot of the same things that you do, which is lots of different applications, each running in their own place. A former colleague of mine used to call it ‘best execution venue’ over at 451. And what I don’t see, or almost never see, is that unicorn of the single application that seamlessly migrates across multiple different cloud providers, or does the whole cloud-bursting thing where you’ve got your on-prem or colo workload, and it seamlessly pops over into AWS, or Azure, or GCP, or wherever else, during peak capacity season, like tax season if you’re at a tax company, or something along those lines. You almost never see anything that realistically does that because it’s so hard to do and the payoff is so low compared to putting it in one place where it’s the best suited for it and focusing your time and effort on the business value part of it rather than on the cost minimization part and the risk mitigation part of, if you have to move from one cloud provider to another, what is it going to take to do that? Well, it’s not going to be that easy. You’ll get it done, but it’ll be a year and a half later, by the time you get there and your customers might not be too happy at that point.

Corey: One area I want to get at is, you talk about, now, addressing developers where they are and solving problems that they have. What are those problems? What painful problem does a developer have today as they’re building an application that Docker is aimed at solving?

Donnie: When we put the problems that we’re hearing from our customers into three big buckets, we think about that as building, sharing, and running a modern application. There’s lots of applications out there; not all of them are modern, so we’re already trying to focus ourselves into a segment of those groups where Docker is really well-suited and containers are really well suited to solve those problems, rather than something where you’re kind of forklift-ing it in and trying to make it work to the best of your ability. So, when we think about that, what we hear a lot of is three common themes. Around building applications, we hear a lot about developer velocity, about time being wasted, both sitting at gatekeepers, but also searching for good reusable components. So, we hear a lot of that around building applications, which is, give me a developer velocity, give me good high-trust content, help me create the good stuff so that when I’m publishing the app, I can easily share it, and I can easily feel confident that it’s good.

And on the sharing note, people consistently say that it’s very hard for them to stay in sync with their teams if there’s multiple people working on the same application or the same part of the codebase. It’s really challenging to do that in anything resembling a real-time basis. You’ve got the repository, which people tend to think of—whether that’s a container repository, or whether that’s a code repository—they tend to think of that as, “I’m publishing this.” But where do you share? What do you collaborate on things that aren’t ready to publish yet?

And we hear a lot of people who are looking for that sort of middle ground of how do I keep in sync with my colleagues on things that aren’t ready to put that stamp on where I feel like it’s done enough to share with the world? And then the third theme that we hear a lot about is around running applications. And when I distinguish this against old Docker, the big difference here is we don’t want to be the runtime platform in production. What we want to do is provide developers with a high-fidelity, consistent kind of experience, no matter which environment they’re working with. So, if they’re in their desktop, if they’re in their CI pipeline, or if they’re working with a cloud-hosted developer environment, or even production, we want to provide them with that same kind of feeling experience.

And so an example of this was last year, we built these Compose plugins that we call code-to-cloud plugins, where you could deploy to ECS, or you could deploy to ACI cloud container instances, in addition to being able to do a local Compose up. And all of that gives you the same kind of experience because you can flip between one Docker context and the other and run, essentially, the same set of commands. So, we hear people trying to deal with productivity, trying to deal with collaboration, trying to deal with complex experiences, and trying to simplify all of those. So, those are really the big areas we’re looking at is that build, share, run themes.

Corey: What does that mean for the future of Docker? What is the vision that you folks are aiming at that goes beyond just, I guess—I’m not trying to be insulting when I say this, but the pedestrian concerns of today? Because viewed through the lens of the future, looking back at these days, every technical problem we have is going to seem, on some level, like it’s, “Oh, it’s easy. There’s a better solution.” What does Docker become in 15 years?

Donnie: Yeah, I think there’s a big gap between where people edit their code, where people save their source code, and
that path to production. And so, we see ourselves as providing a really valuable development tools that—we’re not going to be the IDE and we’re not going to be the pipeline, but we’re going to be a lot of that glue that ties everything together. One thing that has only gotten worse over the years is the amount of fragmentation that’s out there in developer toolchains, developer pipelines, similar with the rise of microservices over the past decade, it’s only gotten more complicated, more languages, more tools, more things to support and an exponentially increasing number of interconnections where things need to integrate well together. And so that’s the problem that, really, we’re solving is all those things are super-complicated, huge pain to make everything work consistently, and we think there’s a huge amount of value there and tying that together for the individual, for the team.

Corey: Donnie, thank you so much for taking the time to speak with me today. If people want to learn more about what you’re up to, where can they find you?

Donnie: I am extremely easy to find on the internet. If you Google my name, you will track down, probably, ten different ways of getting in touch. Twitter is the one where I tend to be the most responsive, so please feel free to reach out
there. My username is @dberkholz.

Corey: And we will, of course, put a link to that in the [show notes 00:29:58]. Thanks so much for your time. I really appreciate the opportunity to explore your perspective on these things.

Donnie: Thanks for having me on the show. And thanks everybody for listening.

Corey: Donnie Berkholz, VP of products at Docker. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an insulting comment that explains exactly why you should be packaging up that comment and running it in any cloud provider just as soon as you get Docker’s command-line arguments squared away in your own head.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

This has been a HumblePod production. Stay humble.

View Details

About Owen

Owen Rogers wears many hats at 451 Research; he’s research director of cloud transformation and digital economics and head of the quantum computing centre of excellence. Prior to these positions, Owen was a doctoral researcher in cloud computing at the University of Bristol, completing his PhD thesis in 2013; a product portfolio manager at Claranet; and an independent product management and cloud computing consultant, among other positions.

Join Corey and Owen as they talk about what it’s like when two cloud economists meet at an event but only one has a PhD, what exactly an industry analyst does, how 451 Research found that 53% of companies increased cloud spend during the pandemic and what resources they’re investing in, the Law of Cloud Entropy and why Owen believes the cloud will only get more disordered in the future, why it’s easy for cloud costs to spiral out of control, how organizations are trying to rein in cloud spend despite using more cloud services, why Owen doesn’t believe we’ll reach cloud commoditization anytime soon, and more.

Links:

  • 451 Research: https://451research.com/
  • Cloud Price Index: https://451research.com/services/price-indexing-benchmarking/cloud-price-index
  • Twitter: https://twitter.com/owenrog

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at the Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Thinkst. This is going to take a minute to explain, so bear with me. I linked against an early version of their tool, canarytokens.org in the very early days of my newsletter, and what it does is relatively simple and straightforward. It winds up embedding credentials, files, that sort of thing in various parts of your environment, wherever you want to; it gives you fake AWS API credentials, for example. And the only thing that these things do is alert you whenever someone attempts to use those things. It’s an awesome approach. I’ve used something similar for years. Check them out. But wait, there’s more. They also have an enterprise option that you should be very much aware of canary.tools. You can take a look at this, but what it does is it provides an enterprise approach to drive these things throughout your entire environment. You can get a physical device that hangs out on your network and impersonates whatever you want to. When it gets Nmap scanned, or someone attempts to log into it, or access files on it, you get instant alerts. It’s awesome. If you don’t do something like this, you’re likely to find out that you’ve gotten breached, the hard way. Take a look at this. It’s one of those few things that I look at and say, “Wow, that is an amazing idea. I love it.” That’s canarytokens.org and canary.tools. The first one is free. The second one is enterprise-y. Take a look. I’m a big fan of this. More from them in the coming weeks.

Corey: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the Enterprise (not the starship). On-prem security doesn’t translate well to cloud or multi-cloud environments, and that’s not even counting IoT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IoT devices, detects these threats up to 35 percent faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at extrahop.com/trial.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’m joined this week by Owen Rogers, who’s a research director of cloud and managed services transformation, at 451 Research, a part of S&P Global Market Intelligence. Owen, thank you for tolerating my slings and arrows.

Owen: Lovely to be here, and I look forward to them.

Corey: So, you got your PhD back in 2013, in cloud economics. And I know this because when we first met at a cloud economics event, I called myself a cloud economist and you lit up like a Christmas tree. “Oh, my God. Someone else does what I do.” And I just gave myself the title because I thought it was something I’d invented, whereas you actually got a PhD in it, and you made the understandable assumption that I knew what I was talking about.

And I knew I had two directions I could go in. The first was, to be honest and come clean, and the other was to basically string you along until we co-publish a book. It comes out in two weeks, and its title is—I’m kidding. I’m kidding. But thank you for being as gracious about me stomping on your credentials back then, as you were.

Owen: No, I was relieved, actually, that someone else was looking into this because I thought I was just on my own and I’d made this big gamble by coming up with this PhD and taking a bit of a risk. And when I saw you at that event, I was like, “Thank God it’s not just me. Thank God, this actually might be a thing to pay attention to.” So, thanks for being there.

Corey: No, by all means, it turns out that there’s a very narrow subset of people who care about these things, and that tends to be a somewhat insular circle. So, it was nice to finally meet someone who was a bit outside of the nuts and bolts of lowering bills and looking at the broader implications across the market. But we’ll get there. You’re an industry analyst. What does that mean?

Owen: So essentially, we are the intermediary party between buyers and sellers. So, we help sellers of services and goods work out who to sell to, how to sell, and we help buyers work out what product is best for them. And we do this by conducting market research across buyers and sellers, pretty much. So, my specialism is cloud economics, so it’s my job to help solve a lot of this complexity that’s going on in cloud pricing for enterprises, and to help service providers and tech vendors sell and price appropriately.

Corey: Well, let’s define terms as well. One of the hardest ones is ‘cloud.’ In fact, the reason I called myself a cloud economist is because it was two words that no one could accurately define in almost any context. Cloud generally meant a bunch of people’s computers that weren’t yours, and economists generally meant someone who claimed to know everything about money but dressed like a disaster victim. And put the two together and no one had any idea what the hell I did, and in many cases, they still don’t, which is kind of ideal from my perspective.

But when you say cloud, what does that mean? Are we talking Infrastructure as a Service? Are we moving up the stack into SaaS? Are we taking a Microsoft-ian definition and including LinkedIn revenue as part of their cloud unit for some godforsaken reason? What is it? Where do you start? Where do you stop?

Owen: I mean, when we both started looking at cloud economics, it was all about the infrastructure because people wanted virtual machines and storage, and there was a huge amount of complexity there. But I think as the market has become more mature, increasingly we’re seeing people want to use all these other services, up to the platform level and the software level. So, I’m interested primarily in infrastructure and platform, but the thing about cloud providers nowadays is, there doesn’t seem to be any barriers to where they want to go. And I think you and I and others who are in this field are going to have to broaden our horizons and start thinking about everything because cloud is becoming the center of all IT, really.

Corey: What I find strange is that as the further I go afield from my core competency in this space, which is looking at the AWS side, Infrastructure as a Service spend—which is not small in most companies and not getting any smaller—as soon as you start diverging from there, the requests that I start seeing from customers are all over the map. Some of them are trying to work on their Microsoft licensing. Others are trying to optimize some random SaaS tool’s billing because that’s top of mind at some point. It immediately shatters into 1000 different niches. But the common thread that I’ve always found was the AWS bill. And let me be very clear: that’s a function of who I talk to in my market in which I live, here in San Francisco. That does not mean that AWS is in every type of company of every profile; just the ones I started talking to and figured out that I could help.

Owen: Yeah. That would make sense to me. I think COVID particularly, has made companies realize that cloud is an option. So, even though not everybody was using the cloud hyperscalers one or two years ago, I think over the past year, even if you weren’t dabbling in the cloud, perhaps you’ve started to play around. And even though optimization might be the first thing you think of when you start using the cloud, as time goes on and things start getting carried away, then that’s when this optimization is going to become more important.

So, for example, we found 53% of enterprises we surveyed are using more cloud services as a result of the pandemic because they’ve had to change things. And even though they’ve probably spun it up really quickly because they’ve needed to grab hold of it quickly and take hold of the opportunity as soon as they could, in years to come there’s going to be so much complexity, and they’re all going to have different requirements—as you said—because they’re all using things in different ways to address their different needs as the pandemic has gone on.

Corey: Talk to me a little bit about the COVID spike that you’re seeing. Is that people spinning up a bunch of new VMs? Is it people leveraging different video conferencing services, like Zoom or—God forbid—WebEx? Is it something else entirely?

Owen: I think there’s two areas, and pretty much what you said. So, there’s some companies who are scaling their existing cloud resources. So, those who have built scalable applications, they’re just adding more and more virtual machines or auto-scaling so that they can keep up with demand. And as you said, that could be something like video conferencing, authentication, VDI, anything like that. But then there’s also companies who have had to really rapidly change business model because a year ago, they weren’t selling things online, they weren’t able to deliver, everything was bought in a shop.

And now they’ve had to rapidly get access to cloud resources and rebuild their businesses, almost. And I think there’s lots of new cloud users as a result of the pandemic who thought, “How are we going to suddenly get this infrastructure? Or we’re going to have to go to whenever hyperscalers to get used to it, get hold of it right now.”

Corey: One of the things that caught me by surprise, in the early days of the pandemic was a number of different companies whose business models weren’t really extending to online things in the same way. They were all based around real-world, physical commerce, for lack of a better term, saw their user traffic and their e-commerce traffic fall off a cliff, but their infrastructure spend remained relatively stable. At which point we realize, “Oh, interesting, when everyone talks about being able to auto-scale, they just mean up.” And that makes sense on some visceral level because if you don’t scale up, you drop customers on the floor. If you don’t scale down, well, it just costs a little bit of extra money and that’s not the end of the world, comparatively. And suddenly seeing people in somewhat dire straits, in a company context, and having to renegotiate their commits with different providers was something of an eye-opener.

Owen: Yeah, yeah. And I identified this term, which is ‘cloud entropy.’ And I came up with a term called the ‘Law of Cloud Entropy.’ And what that essentially means is, cloud is only likely to get more disordered over time because most enterprises, as you said, would rather just leave things running, would rather scale up than risk scaling down because if they scale up, it means their applications can still continue to run, they’re not going to be shot in the foot by a server going down, but if they’re scaling down too quickly, then they’ve got a lot to lose that only a tiny cost saving to make. So, it’s almost like the cloud model is inherently risky, in terms of costs running away with it because things can happen automatically and there’s so much to lose by getting it wrong.

Corey: Absolutely. After your cost-saving exercise winds up causing an outage, you’re generally not allowed to save money anymore.

Owen: Yeah. Yeah, totally. Like, why risk saving a few cents, when it could bring your business down? But that’s the thing. I mean, it's not just a few cents anymore, is it? Because over time, people consume more and more resources, things aren’t being managed correctly; it’s really easy for those costs to spiral out of control.

And that is not just a few cents. It’s thousands, tens of thousands of dollars. And there is a point where you think, “Well, actually, I am going to have to do this now because costs are spiraling, and it is time to take that step into optimizing and cutting my costs.”

Corey: I’ve got a level with you, it does not stop at tens of thousands of dollars. Many of my clients wish it did. “Sure, we can eat that no problem.” It becomes something so much deeper, and it grows without any bounds on it. If you spin up an instance with the idea that you’ll just experiment on something and then turn it off in a couple of days. If you don’t proactively turn it off yourself, you’re going to retire before that instance does. It’ll sit there costing you, every hour of the day.

Owen: Well, I’ve done that, as I’m sure you have. I’ve sped up a virtual machine, played around, and then six months later, I’ve been like, “Oh, this is surprisingly expensive.” And then I’m like, “Oh, well. I’m not going to be able to expense this, am I?” It’s only going to get worse.

So, 49% of enterprises we survey say that cost savings are going to be a greater priority since COVID. So ironically, even though people are using more cloud, and perhaps these costs are spiraling out of control a bit, the fact of the matter is, they’ve never been under more pressure to try and quell it.

Corey: Oh, yeah. And it doesn’t get any easier when people look at these things in their own right. And it doesn’t lend itself to easy analysis, especially as you start getting into large swings, you have seasonal cycles, you have people buying reserved instances, or savings plans, or whatever the other provider equivalent is in bulk at certain times of the year. And it’s very difficult to do accurate projections, especially when you don’t know the answer to a number of very pressing business questions. It almost becomes, in my case, marriage counseling between Finance and engineering.

Owen: That’s such a shrewd observation because I think there is this huge disjoint between IT and Finance. And I can’t really see that being solved anytime soon because they’re both—it’s Finance’s job is to save money, but IT is to keep things ticking over and to innovate. And unfortunately, it’s a compromise. But to get a compromise when neither party really understands each other’s field is really tricky.

Corey: Absolutely. If AWS were to somehow wave a magic wand and fix their billing—and, my God, I wish they would—I still have a business here. I still have credibility when talking to a customer about, “Is this the right level of spend? The right level of commit?” That you’re never going to have when your email address ends in the same domain as the vendor’s. And the ability to help them negotiate what those commits look like with that vendor is one of those business models that never goes out of style.

Owen: Yeah, there’s always going to be that negotiation. Although I think it’s not as big as it used to be. So, when it was, like, server hardware 20 years ago, the list price was nothing like the price that you’d actually pay. Whereas in cloud, I think, I don’t know if you agree, but the variation seems far smaller to me. It almost seems 10 to 15% rather than the 50 to 60% it might have been 20 years ago.

Corey: At certain points of scale, that no longer holds true.

Owen: Interesting. Interesting.

Corey: Not to name names, or specify numbers. Again, confidentiality matters. But at some point, when you wind up being a significant portion of a given service’s revenue, again, no one is paying retail, or anything even close to retail, at a certain point.

Owen: And things are only going to get more complicated. So, we track all the things for sale from AWS, Google, Microsoft, and every week, we scan, now, 2 million individual line items for sale from those cloud providers. So, even if there was some kind of standardization list price with everything, that’s not going to apply to all of those different line items. So, I think for people like Duckbill, a lot of the need is to look at this whole bill, look at everything that’s being used for opportunities to optimize and negotiate, not just on the handful of services, which most enterprises are using.

Corey: When you say that you can consume all of those pieces of information in a single week, that tells me you’re doing some definite data crunching and big number processing, largely because it’s impossible to get that much clarity within a week. Do you find that the cloud providers themselves change pricing—other than on things like preemptible instances or the spot market—without an announcement?

Owen: Interesting. So, the Cloud Price Index, which I manage, essentially, every week we look at the websites of all these cloud providers, we go through their APIs, and we look at every price item they have and we compare it to the week before. And sometimes prices just go up and down just like a blip. It’s almost something’s gone wrong on the website or the API. But in 2020, we saw 4000 significant price cuts.

So, a significant price cut is one that is greater than 10%. So, sometimes prices go down over time and the cloud providers don’t make a big song-and-dance about it anymore. But other times prices do go up, and those prices seem to go up, in particular, when a product goes from almost a beta into general availability. And different cloud providers do it in different ways. But yeah, I think prices are almost continually changing, and it’s almost like a sea of prices rather than thinking, “Oh, well, everything’s going down, or some things are going up.”

Things going up and down all the time and it’s tricky to really know what’s going on. I think this is why cost optimization is going to be needed on an ongoing basis. Because it’s not just a one-off thing anymore, or where you go and buy a bunch of reserved instances. You need to be constantly reassessing this all the time. And, like, we were talking about the synergy between IT and Finance, you need to work out what the company is going to be doing in a year’s time so if it’s worthwhile investing in something to make those savings.

Corey: When you say that they’re thousands of price changes, generally decreases, are they often correlated—in other words, if, “We’re going to be reducing the cost of the X instance family. The end.” But then there’s thousands of SKUs on some cases because they’re in all of the different regions, they have all the different pricing for the committed price, the reserve price, et cetera, et cetera, or are they making large-scale cuts and just not mentioning it? Because there was a time on the AWS side, which is where I live, where they would trumpet every minor cost reduction in some far-flung region for some service that basically no one used.

Owen: You’re right. A lot of those cuts are because of a family, or a particular region, or a generation. And obviously, one cut translates to thousands of individual line items, which again, shows the complexity for companies to deal with because they’ve got to understand that one change can affect a whole range of different things. It’s not just one change anymore; it’s tens of thousands.

Corey: What I hope is that, at some point, we’re going to start seeing something approaching commoditization in the space, but the price that has never materially changed—well, that’s unfair. The price that has generally never materially
changed has been the egress fee for data transfer.

Owen: See, I don’t think we’re going to reach commoditization for a long time yet. And I think of it as a gas station analogy. So, if there’s a bunch of gas stations all on the same road, we all know that the cheapest gas station will be the one that probably gets most of the business because people are only buying gas. But the reality is, people go to gas stations for loads of different things. They go because one might have a nice restaurant; one might sell different chocolates and candy.

So, it’s not really about the commoditized offering of the gas. There’s loads of other things that would drive why you might choose one gas station over another. And I think that is the same with cloud providers. That yeah, they all might get similar prices for virtual machines at some point. But still, there’s going to be a reason why you might choose AWS, or Google, or Microsoft, or Oracle, or IBM, or Alibaba, or any of these folks. It’s going to be because of their whole portfolios and everything else they offer in trust and reliability, and regional access, and not just that single commodity price point which is their core business.

Corey: Part of the problem is, at least in my experience, when I look at the customer profile that I tend to engage with, they have the bulk of their expenses, across a very small number of services, almost always EC2, RDS, S3, Elastic Block Store, and data transfer. And everything else is, sort of, a bit of a rounding error. There are always going to be exceptions on this, but what that tells me is that despite all the high-level services that get trumpeted, and despite the flashy abilities, and capabilities, and savings opportunities, et cetera, et cetera, that get trotted out, during all the provider keynotes,
people are still largely using this to run virtual machines and store data. Is that a fair assessment from what you’re seeing?

Owen: I would strongly agree with you. And it’s because people know how to build applications on servers. There’s different skills, but people have got the skills already to some degree. Whereas if you want to use serverless, or these new analytics tools, or IoT, or machine learning, that’s a whole new skill set. I’m with you; I think the bulk of it is still the basic infrastructure items.

Corey: It really seems to be. And I can’t shake the feeling that as much as they want to give attention to the new stuff, it’s not a massive driver of people who are debating adopting the cloud. I really don’t think that it’s going to change anytime soon. If we take a look at AWS that has an annualized $51 billion run rate, and revenue at this point, which is just astonishing, it’s pretty clear that the next $51 billion is not going to come from the same customer profile. If anything, it’s going to come from what looks an awful lot like blue-chip companies, some of whom are in manufacturing, some of whom are in logistics, et cetera, et cetera.

They’re not web properties; they’re not Netflix-style companies. And to meet those people where they are, they have to embrace the edge a lot more closely, they have to tell a story where you can manage the data center and the cloud environment similarly. And if anything, it’s going to increase that trend, not decrease it.

Owen: How do you feel about multi-cloud, mate?

Corey: I was hoping we would get there.

Owen: [laugh].

Corey: I have thoughts on the matter, but I will do you the service of letting you start.

Owen: So, [laugh]. So, 59% of enterprises we talk to are pursuing a hybrid approach to IT. So, what that means in our language is, essentially enterprises want to make sure they can use different cloud providers. And the top reason they want to do that is because they want to choose which is the best expertise from each individual cloud provider. So yeah, they might want a cloud provider A because they’ve got really cheap infrastructure, yadda, yadda, yadda.

But they still want to have the freedom to use cloud provider B because they’ve got these cool, sexy new analytics and stuff. And for me, I think the hyperscalers almost have to have these newer sexier services, not necessarily because lots of companies are going to use them and it’s going to erode all the commodity business, but more because if they don’t have them, it almost appears like a bit of a weakness because their competitors all have the same thing. And considering enterprises are so willing to consider multiple cloud environments, I think that more appropriately shows that you have to have these things because companies will look elsewhere if you don’t.

Corey: This episode is sponsored in part by our friends at Lumigo. If you’ve built anything from serverless, you know that if there’s one thing that can be said universally about these applications, it’s that it turns every outage into a murder mystery. Lumigo helps make sense of all of the various functions that wind up tying together to build applications. It offers one-click distributed tracing so you can effortlessly find and fix issues in your serverless and microservices environment. You’ve created more problems for yourself; make one of them go away. To learn more, visit lumigo.io.

Corey: I’m not going to disagree with what you’re saying because I’m not sure of the direction it’s going to go in, yet. I’ve often been mischaracterized when I rant about multi-cloud being a worst practice that I’m saying that you should absolutely pick one provider and go all in. Full stop. And that’s never been what I intended to say. For example, personally, all my infrastructure lives on AWS, give or take a few things that are hosted WordPress, for example.

But my Git repositories live on GitHub because code commit is a funny joke that people just haven’t realized as a joke yet. And I use G Suite for email and the rest because work mail and work docs are services that even now, you’re not sure I’m not making up. And that’s the way that I tend to view the multi-cloud story: different workloads in different places. And that tends to be fine because in this case, there’s not a whole lot of interaction between those things. The dumb version of multi-cloud, to my world—and I think you called it hybrid in some respects—is the idea of, “I’m going to take a workload that can seamlessly go to any different cloud provider at any time.” And in practice, it never does that, and it also winds up trading off a lot of the benefit of going to public cloud in the first place.

Owen: Yeah. That makes sense to me. And actually, in our data, that’s exactly what we found. So, it’s something like the average number of clouds used by the enterprises we survey is 2.2 on average, but 80% of their workloads are deployed in one cloud.

So, I think you’re right; it is almost an aspiration. It’s just keeping your option open. Having the ability to move workloads between clouds constantly, we don’t see it either. It’s more about just having the ability to if you really needed to. Do you think some of that is because people are just scared of that lock-in, in your experience? Is it more of a psychological worry, than actual—a worry in reality?

Corey: Partially that. It’s also that there’s a vendor ecosystem where if you’re selling a shared control plane that can speak to all three of the primary tier-one cloud providers, and people aren’t using multiple cloud providers, you suddenly have nothing left to sell them. It’s also being sold, in many respects, by cloud providers who are painfully aware that if you go all-in on one cloud, it will not be theirs.

Owen: Mmm. Yeah, makes sense. But it’s been interesting over the past year or so how the hyperscalers have started talking mature about hybrid and multi-cloud. So, five years ago, I didn’t think AWS would ever have something like Outposts. And also their competitors. So Google, have Anthos where you can move workloads, Microsoft came out with Arc. So, it’s surprising to me that they’re all embracing this concept so readily.

Corey: Well, I do want to call out that there is a distinct difference in my mind, between using multiple cloud providers and having a hybrid structure where you have a data center and a cloud provider because everyone goes through a migration process there. In fact, a failed cloud migration is called, “We’re hybrid now,” because it turns out midway through, it’s super hard to move something so you give up and declare victory. No one generally sets out to live permanently with a foot in each world. What invariably happens then is they improve their data center at the expense of their cloud environment. And they really tend to treat the cloud more or less as an incredibly expensive place just to run a bunch of virtual machines, compared to what they would get economically on-prem. Now, that said, the raw infrastructure cost is only a small part of the story.

Owen: Yeah because you’ve got the labor cost of running it yourself as well, right?

Corey: Which is always more expensive, than the infrastructure. It’s an incredible rarity when we see the AWS bill costing
more than payroll.

Owen: Then again, I think, you know, it’s not just the cost savings of having some of your cloud stuff on-prem, though, is it? I mean, the world is a complex place at the moment, pandemic, politics. I think some buyers like to have their data somewhere where they know, in a country they understand the compliance and the sovereignty. I mean, even though cloud is an easily accessible place, you still don’t have ownership of it all. There’s some things you just want to keep close to home. And that is a lot of the driver we see for the hybrid model. The public cloud gives the flexibility but the on-prem cloud lets you still have some flexibility, but keep it all controlled and in your own arms.

Corey: To be very clear, I’m also speaking in the very general case. When I talk to individual clients who have made a different decision, my default assumption there is that they’ve thought through these things and have a reason for things being the way that they are. My problem—and why I started making noise about this topic—is that, in my experience, no one else was saying it, which means that if you don’t really know what’s going on and you listen to just the vendor hype, then you would think that, “Oh, I absolutely must build everything that I’m doing in the cloud to work on multiple providers on day one.” And that’s just not the case.

Owen: I see what you mean. Yes, that’s not the case. Just because people are using multiple venues doesn’t mean they’re all necessarily working together in any sensible way.

Corey: Exactly. This is part of the reason I have no partnerships with any vendor in this space. It’s the reason I don’t charge percentages of things. It never goes well. I wind up charging fixed-fee to my clients and then I tell them to do what I would
do in their position. And I’ll explain my logic as I go, and everyone’s generally pretty happy with that.

Owen: Mmm. Makes sense.

Corey: For better or worse, it seems to solve the problem that folks have. But it’s a growing market; I’m never going to be able to talk to more than a very small percentage of it, and this problem has to be solved, on some level, systemically. Because if we look at cloud spend as an unbounded growth problem, well, first, it means that in the cloud business is a great place to be if you’re one of the ones that’s making money at it. But it also means that at some point, there’s going to be some kind of a reckoning where people need to go back and play cloud environment archaeologist. And this isn’t just a big company problem.

I’m the only person that was in most of our early accounts here at The Duckbill Group and I have to figure out what that idiot moron known as my past self was thinking when he tried to build some of this nonsense. And the short answer is, he had no idea, but it seemed like a good idea at the time.

Owen: I totally love that: ‘cloud bill archaeologist’ and I will be stealing that for a future reference. And I think you’re right, even during the pandemic, I bet loads of people have scaled up straightaway, thinking, “Well, we’ve got to capture the opportunity now, or we’re going to risk losing business.” And no one’s really planned it or looked at it for a year because, quite understandably, they’ve had bigger things to worry about. And in five years’ time, no one’s going to know what’s going on, what workload is tied to what specific application, who owns it. And the thought of even understanding it is challenging, let alone trying to optimize it.

And I was having a debate with my colleague, Jean Atelsek, today, and she was asking me if I thought one day, this could all be automated away. And I don’t think it can be automated away because there still needs to be someone who understands the business, to understand scaling, to understand if something’s worth an investment, to understand if you should scale up or down in response to a specific demand or project. So, I always think there’s going to be some kind of human intervention, just because humans will understand the needs of the business and relate them to how the cloud has to change.

Corey: The only other approach as I see it, than my own are, “Oh, we’re going to build some tools that will solve all of this for you.” And they just don’t work. That’s terrific to wind up finding specific things, absolutely, but there’s no context to them. There’s no idea of, “Should I optimize for this cluster for the long haul, or should I instead wind up focusing on it as this thing that I should immediately ignore?” As soon as you start getting three or four terrible recommendations in a row, you wind up in a space of not trusting the tool at all. Bad recommendations are worse than no recommendations.

Owen: So, why do you think that is? Why do you think the tools—I suppose the tools can predict the future so they don’t know the context of what needs to be done. Are there any other reasons to see those tools has not been able to adapt? Why do you not think tools will have a longer-term impact.

Corey: Because in many cases, there’s no way to tell from a programmatic perspective. “Those idle instances that are sitting there? I’m going to recommend that they get turned off.” Well, a little more digging shows that they’re the DR site and you need three seconds of warning or so before they’re going to be under load. You can’t turn them off.

Whereas, buy a bunch of reserved instances on that particular cluster that someone just spun up for a one-week experiment and then they’re turning it off, doesn’t make a lot of sense, either. And as you step down this path, it becomes nuanced. There are times where that is this tiny little test environment, so no one is going to look at it or care. Except that that tiny little test environment is about to go hyperscale once the business deal gets signed, so now is absolutely the time to optimize stuff like that. There’s the idea of well, this data could be migrated to infrequent access one zone, and it winds up costing less money. Cool. That’s true, but if that data goes away, it winds up effectively destroying aspects of the business.

So, in that case, you should spend more money on backing that data up securely, in many cases, to another cloud provider. That’s the level of nuance. There’s a whole bunch of different things that a naive approach would suggest would be a good idea. But a deeper dive into what the business is actually doing and the model that they’re working under, make it the wrong direction to go.

Owen: I strongly agree.

Corey: And it gets worse than that because there’s this false narrative that companies care tremendously about saving money on the bill; that’s the thing that drives them. And it is just not true. Because it’s an inversion of monetary philosophy that people take on a personal level. If I offer you the opportunity, you can either make another $1,000, or save $1,000, you’re typically going to say you’d rather save the money because, well, you can cut Netflix out, you can stop eating out, and that works out well, whereas having to go ahead and make more money, that means you have to ask your boss for a raise and start doing odd jobs and update your resume, and it’s just a pain. Companies, on the other hand, are structured to drive revenue. There’s a theoretical cap of whatever they’re spending on cloud in total, that they’ll ever be able to cut off, but they can make multiples of that by launching the right feature to the right market at the right time.

Owen: I wonder if cost optimization is perhaps the wrong word and it should be something along the lines of value optimization because obviously, I don’t want any company reducing their virtual machines to save money because as you said, it’s going to reduce their opportunities to gain new revenue if it’s their web applications. Really, it’s about, “Well, this is where you should be putting your money, and this is where you’re wasting it.” It’s about optimizing their value, not saving them their costs.

Corey: Precisely. It comes down to what’s right for them, given the constraints that they’re working under. And again, it’s easy to go ahead and play, more or less, Wild West architecture, where you look at what they’re doing and say, “Oh, yeah. This is all wrong, you should be doing it this way.” And you sketch out a beautiful architecture on a whiteboard—also known as a lie—and, yeah, in theory, it’s great.

In practice, they have existing business, it’s driving revenue, and you’re not going to be allowed to turn everything off for 18 months while you rebuild it. The money that you save doesn’t matter if you’re not in business by the time you’re in a
position to realize those savings.

Owen: And perhaps after COVID, there will be loads of these servers and virtual machines and objects on object storage which are left there, just because it’s not really worth removing them. Because nobody knows what they are, it might bring down the whole business. Better just leave them there for the time being.

Corey: Well, that does bring up the last topic I wanted to bounce off of you. What is the outcome of all of this COVID stuff, once it is all past? What is the lasting after effect, if any, of COVID on cloud?

Owen: I think COVID will be a catalyst for cloud adoption. Some companies have changed their business models; they’ve aged collaboration; they’ve been able to change their businesses in a matter of weeks. And that’s been enabled because of the rapid scalability of cloud because they’ve been able to get a third-party to do physical server management and because they’ve been able to concentrate on changing and evolving their businesses instead of worrying about infrastructure. I think those who have succeeded by doing that are likely to keep doing that because it puts them in good stead during the past challenging eight months. And those who hadn’t done that will now think, “Well, perhaps we should have done that.” And again, they’ll look at the cloud as a way of moving forward. So COVID, although horrifically terrible for so many people, will probably be a catalyst for cloud adoption, and has demonstrated to the industry that cloud is a suitable venue for many, many workloads.

Corey: Owen, thank you so much for taking the time to speak with me and suffer my, I guess, less educated slash informed opinions on cloud economics. If people want to hear more about what you have to say, where can they find you?

Owen: So, you can find me on 451Research.com, or on Twitter; I’m @owenrog.

Corey: And we will of course, put links to that in the [show notes 00:33:43]. Thank you so much for taking the time to speak with me. I really appreciate it.

Owen: No, thank you very much.

Corey: Owen Rogers, research director, and cloud economist at 451, Division of S&P Global. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with a comment listing all 4000 prices that changed.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

This has been a HumblePod production. Stay humble.

View Details

About Chadd

Chadd Kenney is the Vice President of Product at Clumio. Chadd has 20 years of experience in technology leadership roles, most recently as Vice President of Products and Solutions for Pure Storage. Prior to that role, he was the Vice President and Chief Technology Officer for the Americas helping to grow the business from zero in revenue to over a billion. Chadd also spent 8 years at EMC in various roles from Field CTO to Principal Engineer. Chadd is a technologist at heart, who loves helping customers understand the true elegance of products through simple analogies, solutions use cases, and a view into the minds of the engineers that created the solution.

Links:

  • Clumio: https://clumio.com/
  • Clumio AWS Marketplace: https://aws.amazon.com/marketplace/pp/prodview-ifixh6lnreang

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by ChaosSearch. As basically everyone knows, trying to do log analytics at scale with an ELK stack is expensive, unstable, time-sucking, demeaning, and just basically all-around horrible. So why are you still doing it—or even thinking about it—when there’s ChaosSearch? ChaosSearch is a fully managed scalable log analysis service that lets you add new workloads in minutes, and easily retain weeks, months, or years of data. With ChaosSearch you store, connect, and analyze and you’re done. The data lives and stays within your S3 buckets, which means no managing servers, no data movement, and you can save up to 80 percent versus running an ELK stack the old-fashioned way. It’s why companies like Equifax, HubSpot, Klarna, Alert Logic, and many more have all turned to ChaosSearch. So if you’re tired of your ELK stacks falling over before it suffers, or of having your log analytics data retention squeezed by the cost, then try ChaosSearch today and tell them I sent you. To learn more, visit chaossearch.io.

Corey: This episode is sponsored in part by our friends at Lumigo. If you’ve built anything from serverless, you know that if there’s one thing that can be said universally about these applications, it’s that it turns every outage into a murder mystery. Lumigo helps make sense of all of the various functions that wind up tying together to build applications.

It offers one-click distributed tracing so you can effortlessly find and fix issues in your serverless and microservices environment. You’ve created more problems for yourself; make one of them go away. To learn more, visit lumigo.io.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Periodically, I talk an awful lot about backups and that no one actually cares about backups, just restores; usually, they care about restores right after they discover they didn’t have backups of the thing that they really, really, really wish that they did. Today’s promoted guest episode is sponsored by Clumio. And I’m speaking to their VP of product, Chadd Kenney. Chadd, thanks for joining me.

Chadd: Thanks for having me. Super excited to be here.

Corey: So, let’s start at the very beginning. What is a Clumio? Possibly a product, possibly a service, probably not a breakfast cereal, but again, we try not to judge.

Chadd: [laugh]. Awesome. Well, Clumio is a Backup as a Service offering for the enterprise, focused in on the public cloud. And so our mission is, effectively, to help simplify data protection and make it a much, much better experience to the end-user, and provide a bunch of values that they just can’t get today in the public cloud, whether it’s in visibility, or better protection, or better granularity. And we’ve been around for a bit of time, really focused in on helping customers along their journey to the cloud.

Corey: Backups are one of those things where people don’t spend a lot of time and energy thinking about them until they are, I guess, befallen by tragedy in some form. Ideally, it’s something minor, but occasionally it’s, “Oh, yeah. I used to work at that company that went under because there was a horrible incident and we didn’t have backups.” And then people go from not caring to being overzealous converts. Based upon my focus on this, you can probably pretty safely guess which side of that [chasm 00:02:04] I fall into. But let’s start, I guess, with positioning; you said that you are backup for the enterprise. What does that mean exactly? Who are your customers?

Chadd: We’ve been trying to help customers into their cloud journey. So, if you think about many of our customers are coming from the on-prem data center, they have moved some of their applications, whether they’re lift-and-shift applications, or whether they’ve, kind of, stalled doing net-new development on-prem and doing all net new development in the public cloud. And we’ve been helping them along the way and solving one fundamental challenge, which is, “How do I make sure my data is protected? How do I make sure I have good compliance and visibility to understand, you know, is it working? And how do I be able to restore as fast as possible in the event that I need it?”

And you mentioned at the beginning backup is all about restore and we a hundred percent agree. I feel like today, you get this [unintelligible 00:02:51] together a series of solutions, whether it’s a script, or it’s a backup solution that’s moved from on-prem, or it’s a snapshot orchestrator, but no one’s really been able to tackle the problem of, help me provide data protection across all of my accounts, all of my regions, all of my services that I’m using within the cloud. And if you look at it, the enterprise has transitioned dramatically to the cloud and don’t have great solutions to latch on to solve this fundamental problem. And our mission has been exactly that: bring a whole bunch of cool innovation. We’re built natively in the public cloud; we started off on a platform that wasn’t built on a whole bunch of EC2 instances that look like a box that was built on-prem, we built the thing mostly on Lambda functions, very event-driven. All AWS native services. We didn’t build anything proprietary data structure for our environment. And it’s really been able to build a better user experience for our end customers.

Corey: I guess there’s an easy question to start with, of why would someone consider working with Clumio instead of AWS Backup, which came out a few months after re:Invent, I want to say 2018, but don’t quote me on that; may have been 2019. But it has the AWS label on the tin, which is always a mark of quality.

Chadd: [laugh]. Well, there’s definitely a fair bit to be desired on the AWS Backup front. And if you look at it, what we did is we spent, really, before going into development here, a lot of time with customers to just understand where those pains are. And I’ve nailed it, kind of, to four or five different things that we hear consistently. One is that there’s near zero insights; “I don’t know what’s going on with it. I can’t tell whether I’m compliant or not compliant, or protecting not enough
or too much.”

They haven’t really provided sufficient security on being able to airgap my data to a point where I feel comfortable that even one of my admins can’t accidentally fat-finger a script and delete, you know, whether the primary copy or secondary copy. Restore times have a lot to be desired. I mean, you’re using snapshots. You can imagine that doesn’t really give you a whole bunch of fine-grained granularity, and the timeframe it takes to get to it—even to find it—is kind of a time-consuming game. And they’re not cheap.

The snapshots are five cents per gig per month. And I will say they leave a lot to be desired for that cost basis. And so all of this complexity kind of built-in as a whole has given us an opportunity to provide a very different customer experience. And what the difference between these two solutions are is we’ve been providing a much better visibility just in the core solution. And we’ll be announcing here, on May 27, Clumio Discover which gives customers so much better visibility than what AWS Backup has been able to deliver.

And instead of them having to create dashboards and other solutions as a whole, we’re able to give them unique visibility into their environment, whether it’s global visibility, ensuring data is protected, doing cost comparisons, and a whole bunch of others. We allow customers to be able to restore data incredibly faster, at fine-grained granularities, whether it’s at a file level, directory level, instance level, even in RDS we go down to the record level of a particular database with direct query access. And so the experience just as a whole has been so much simpler and easier for the end consumer, that we’ve been able to add a lot of value well beyond what AWS Backup uses. Now, that being said, we still use snapshots for operational recovery at some level, where customers can still use what they do today but what Clumio brings is an enhanced version of that by actually using airgap protection inside of our service for those datasets as well. And so it allows you to almost enhance AWS Backup at some level if you think about it. Because AWS Backups really are just orchestrating the snapshots; we can do that exact same thing, too, but really bring the airgap protection solution on top of that as well.

Corey: I’ve talked about this periodically on the show. But one of the last, I guess, big migration projects I did when I was back in my employee days—before starting this place—was a project I’d done a few times, which was migrating an environment from EC2-Classic into a VPC world. Back in the dark times, before VPCs were a thing, EC2-Classic is what people used. And they were not just using EC2 in those environments, they were using RDS in this case. And the way to move an RDS database is to stop everything, take a final snapshot, then restore that snapshot—which is the equivalent of backup—to the new environment.

How long does that take? It is non-deterministic. In the fullness of time, it will be complete. That wasn’t necessarily a disaster restoration scenario, it was just a migration, and there were other approaches we theoretically could have taken, but this was the one that we decided to go with based upon a variety of business constraints. And it’s awkward when you’re sitting there, just waiting indefinitely for, it turns out, about 45 minutes in this case, and you think everything’s going well, but there’s really nothing else to do during those moments.

And that was, again, a planned maintenance, so it was less nerve-wracking then the site is down and people are screaming. But it’s good to have that expectation brought into it. But it was completely non-transparent; there was no idea what was going on, and in actual disasters, things are never that well planned or clear-cut. And at some level, the idea of using backup restoration as a migration strategy is kind of a strange one, but it’s a good way of testing backups. If you don’t test your backups, you don’t really have them in the first place. At least, that’s always been my philosophy. I’m going to theorize, unless this is your first day in business, that you sort of feel the same way, given your industry.

Chadd: Definitely. And I think the interesting parts of this is that you have the validation that backups occurring, which is—you need visibility on that functioning, at some level; like, did it actually happen? And then you need the validation that the data is actually in a state that I can recover—

Corey: Task failed successfully.

Chadd: [laugh]. Exactly. And then you need validation that you can actually get to the data. So, there’s snapshots which give you this full entire thing, and then you got to go find the thing that you’re looking for within it. I think one of the values that we’ve really taken advantage of here is we use a lot of the APIs within AWS first to get optimization in the way that we access the data.

So, as an example—on your EC2 example—we use EBS direct APIs, and we do change block tracking off of that, and we send the data from the customers tenancy into our service directly. And so there’s no egress charges, there’s no additional cost associated to it; it just goes into our service. And the customer pays for what they consume only. But in doing that, they get a whole bunch of new values. Now, you can actually get file-level indexing, I can search globally for files in an instance without having to restore the entire thing, which seems like that would be a relatively obvious thing to get to.

But we don’t stop there. You could restore a file, you could go browse the file system, you could restore to an AMI, you could restore to another EC2 instance, you could move it to another account. In RDS, not an easy service to protect, I will say. You know, you get this game of, “I’ve got to restore the entire instance and then go find something to query the thing.” And our solution allows you direct query access, so we can see a schema browser, you can go see all of your databases that are in it, you can see all the tables, the rows in the table, you can do advanced queries to join across tables to go [unintelligible 00:10:00] results.

And that experience, I think, is what customers are truly looking forward to be able to provide additional values beyond just the restoration of data. I’ll give you a fun example that a SaaS customer was using. They have a centralized customer database that keeps all of the config information across all of the tenants.

Corey: I used to do something very similar with Route 53, and everyone looks at me strangely when I say it, but it worked at the time. There are better approaches now. But yeah, very common pattern.

Chadd: And so you get into a world where it’s like, I don’t want to restore this entire thing at that point in time to another instance, and then just pull the three records for that one customer that they screwed up. Instead, it would be great if I could just take those three records from a solution and then just imported into the database. And the funny part of this is that the time it takes to do all these things is one component, the accidentally forgetting to delete all the stuff that I left over from trying to restore the data for weeks at a time that now I pay for in AWS is just this other thing that you don’t ever think about. It’s like, inefficiencies built in with the manual operations that you build into this model to actually get to the datasets. And so we just think there’s a better way to be able to see and understand datasets in AWS.

Corey: One of my favorite genres of fiction is reading companies’ DR plans for how they imagine a disaster is going to go down. And it’s always an exercise in hilarity. I was not invited to those meetings anymore after I had the temerity to suggest that maybe if the city is completely uninhabitable and we have to retreat to a DR site, no one cares about this job that much. Or if us-east-one has burned to the ground over in AWS land, that maybe your entire staff is going to go quit to become consultants for 100 times more money by companies that have way bigger problems than you do. And then you’re not invited back.

But there’s usually a certain presumed scale of a disaster, where you’re going to swing into action and exercise your DR plan. Okay, great. Maybe the data center is not a smoking crater in the ground; maybe even the database is largely where; what if you lost a particular record or a particular file somewhere? And that’s where it gets sticky, in a lot of cases because people start wondering, “Do I just spend the time and rebuild that file from scratch, kind of? Do I do a full restore of the”—all I have is either nothing or the entire environment. You’re talking about row-level restores, effectively, for RDS, which is kind of awesome and incredible. I don’t think I’ve ever seen someone talking about that before. How does that map as far as, effectively, a choose-your-own-disaster dial?

Chadd: [laugh]. There’s a bunch of cool use cases to this. You’ve definitely got disaster recovery; so you’ve got the instance where somebody blew something away and you only need a series of records associated to it; maybe the SQL query was off. You’ve got compliance stuff. Think about this for a quick sec: you’ve got an RDS instance that you’ve been backing up, let’s say you keep it for just even a year.

How many versions of that RDS database has AWS gone through in that period of time so that when you go restore that actual snapshot, you’ve got to rev the thing to the current version, which would take you some time [laugh] to get up and running, before you can even query the thing. And imagine if you do that, like, years down the road, if you’re keeping databases out there, and your legal team’s asking for a particular thing for discovery, let’s say. And you’ve got to now go through all of these iterations to try to get it back. The thing we decided to do that was genius on the [unintelligible 00:13:19] team was, we wanted to decouple the infrastructure from the data. So, what we actually do is we don’t have a database engine that’s sitting behind this.

We’re exporting the RDS snapshot into a Parquet file, and the Parquet file then gets queried directly from Athena. And that allows us to allow customers to go to any timeframe to be able to pull not-specific database engine data into—whether it’s a restore function, or whether I want to migrate to a new database engine, I can pull that data out and re-import it into some other engine without having to have that infrastructure be coupled so closely to the dataset. And this was, really, kind of a way for customers to be able to leverage those datasets in all sorts of different ways in the future, with being able to query the data directly from our platform.

Corey: It’s always fun talking to customers and asking them questions that they look at me as if I’ve grown a second head, such as, “Okay. So, in what disaster scenario are you going to need to restore your production database to a state that was in nine months ago?” And they look at me like I’ve just asked a ridiculous question because, of course, they’re never going to do that. If the database is restored to a copy that backed up more than 15 minutes or so in the past, there are serious problems. That’s why the recovery point objective—or RPO—of what is your data window of loss when you do a restore is so important for these plannings.

And that’s great. “Okay then, why do you have six years of snapshots of your database taken on an interval going back all that time, if you’re never going to ever restore to any of them?” “Well, something compliance.” Yeah. There are better stories for that. But people start keeping these things around almost as digital packrats, and then they wind up surprised that their backup bill has skyrocketed. I’m going to go out on a limb presume—because if not, this is going to be a pretty
awkward question—that you do not just backup management but also backup aging as far as life cycles go.

Chadd: Yeah. So, there’s a couple different ways that are fun for us is we see multiple different tiers within backup. So, you’ve got the operational recovery methodology, which is what people usually use snapshots for. And unfortunately, you pay that at a pretty high premium because it’s high value. You’re going to restore a database that maybe went corrupt, or got somehow updated incorrectly or whatever else, and so you pay a high number for that for, let’s say, a couple days; or maybe it’s just even a couple hours.

The unfortunate part is, that’s all you’ve got, really, in AWS to play with. And so, if I need to keep long-term retention, I’m keeping this high-value item now for a long duration. And so what we’ve done is we’ve tried to optimize the datasets as much as possible. So, on EC2 and EBS, we’ll dedupe and compress the datasets, and then store them in S3 on our tenancy. And then there’s a lower cost basis for the customer.

They can still use operational recovery, we’ll manage that as part of the policy, but they can also store it in an airgap protected solution so that no one has access to it, and they can restore it to any of the accounts that they have out there.

Corey: Oh, separating access is one of those incredibly important things to do, just because, first, if someone has completely fat-fingered something, you want to absolutely constrain the blast radius. But two, there is the theoretical problem of someone doing this maliciously, either through ransomware or through a bad actor—external or internal—or someone who has compromised one of your staff’s credentials. The idea being that people with access to production should never be the people who have access to, in some cases, the audit logs, or the backups themselves in some cases. So, having that gap—an airgap as you call it—is critical.

Chadd: Mm-hm. The only way to do this, really, in AWS—and a lot of customers are doing this and then they move to us—is they replicate their snapshots to another account and vault them somewhere else. And while that works, the downside—and it’s not a true airgap, in a sense; it’s just effectively moving the data out of the account that it was created in. But you double the cost, so that sucks because you’re keeping your local copy, and then the secondary copy that sits on the other account. The admins still have access to it, so it’s not like it’s just completely disconnected from the environment. It’s still in the security sphere, so if you’re looking at a ransomware attack, trust me, they’ll find ways to get access to that thing and compromise it. And so you have vulnerabilities that are kind of built into this altogether.

Corey: “So-what’s-your-security-approach-to-keeping-those-two-accounts-separated?” “The sheer complexity that it takes to wind up assuming a role in that other account that no one’s going to be able to figure it out because we’ve tried for years and can’t get it to work properly.” Yeah, maybe that’s not plan A.

Chadd: Exactly. And I feel like while you can [unintelligible 00:17:33] these things together in various scripts, and solutions, and things, people are looking for solutions, not more complexity to manage. I mean, if you think about this, backup is not usually the thing that is strategic to that company’s mission. It’s something that protects their mission, but not drives their mission. It is our mission and so we help customers with that, but it should be something we can take off their hands and provide as a service versus them trying to build their own backup solution as a whole.

Corey: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the Enterprise (not the starship). On-prem security doesn’t translate well to cloud or multi-cloud environments, and that’s not even counting IoT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IoT devices, detects these threats up to 35 percent faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at extrahop.com/trial.

Corey: Back when I was an employee if I was being honest, people said, “So, what is the number one thing you’re always sure to do on a disaster recovery plan?” My answer is, “I keep my resume updated.” Because, on some level, you can always quit and go work somewhere else. That is honest, but it’s also not the right answer in many cases. You need to start planning for these things before you need them.

No one cares about backups until right after you really needed backups. And keeping that managed is important. There are reasons why architectures around this stuff are the way that they are, but there are significant problems around how a lot of AWS implements these things. I wound up having to use a backup about a month or so ago when some of my crappy code deleted data—imagine that—from a DynamoDB table, and I have point-in-time restores turned on. Cool. So, I just roll it back half an hour and that was great. The problem is, there was about four megabytes of data in that table, and it took an hour to do the restore into a new table and then migrate everything back over, which was a different colossal pain. And I’m sure there are complicated architectural reasons under the hood, but it’s like, that is almost as slow as someone who’s retyped it all by hand, and it’s an incredibly frustrating experience. You also see it with EBS snapshots: you backup an EBS volume with a snapshot—it just copies the data that’s there. Great—every time there’s another snapshot taken, it just changes the delta. And that’s the storage it gets built to. So, what does that actually cost? No one really knows. They recently launched direct APIs for EBS snapshots; you can start at least getting some of that data out of it if you just write a whole bunch of code—preferably in a Lambda function because that’s AWS’s solution for everything—but it’s all plumbing solution where you’re spending all your time building the scaffolding and this tooling. Backups are right up there with monitoring and alerting for the first thing I will absolutely hurl to a third party.

Chadd: I a hundred percent agree. It’s—

Corey: I know you’re a third-party. You’re, uh, you’re hardly objective on this.

Chadd: [laugh].

Corey: But again, I don’t partner with anyone. I’m not here to shill for people. You can’t buy my opinion on these things. I’ve been paying third parties to back things up for a very long time because that’s what makes sense.

Chadd: The one thing that I think, you know, we hit on at the beginning a little bit was this visibility challenge—and this was one of the big launch around Clumio Discover that’s coming out on May 27th there—is we found out that there was near-zero visibility, right? And so you’re talking about the restore times, which is one key component, but [laugh]—

Corey: Yeah, then you restore after four hours and discover you don’t have what you thought you did.

Chadd: [laugh]. And so, I would love to see, like, am I backing things up? How much am I paying for all of these things? Can I get to them fast? I mean, the funny thing about the restore that I don’t think people ever talk about—and this is one of the things that I think customers love the most about Clumio—is, when you go to restore something, even that DynamoDB database you talked about earlier, you have to go actually find the snapshot in a long scroll.

So first, you had to go to the service, to the account, and scroll through all of the snapshots to find the one that you actually want to restore with—and by the way, maybe that’s not a monster amount for you, but in a lot of companies that could be thousands, tens of thousands of snapshots they’re scrolling through—and they’ve got a guy yelling at them to go restore this as soon as possible, and they’re trying to figure out which one it is; they hunt-and-peck to find it. Wouldn’t it be nice if you just had a nice calendar that showed you, “Here’s where it is, and here’s all the different backups that you have on that point in time.” And then just go ahead and restore it then?

Corey: Save me from the world of crappy scripts for things like this that you find on GitHub. And again, no disrespect to the people writing these things, but it’s clear that people are scratching their own itch. That’s the joy of open-source. Yeah, this is the backup script—or whatever it is—that works on the ten instances I have in my environment. That’s great.

You roll that out to 600 instances and everything breaks. It winds up hitting rate limits as it tries to iterate through a bunch of things rather than building a queue and working through the rest of it. It’s very clearly aimed at small-scale stuff and built by people who aren’t working in large-scale environments. Conversely, you wind up with sort of the Google problem when you look at solving it for just the giant environments. Great, that you wind up with this overengineered, enormously unwieldy solution. Like, “Oh yeah, the continental saw. We use it to wind up cutting continents in half.” It’s, “I’m trying to cut a sandwich in half here. What’s the problem here?”

It becomes a hard problem. The idea of having something that scales and has decent user ergonomics is critically important, especially when you get to something as critical as backups and restores. Because you’re not generally allowed to spend months on end building a backup solution at a company, and when you’re doing a restore, it’s often 3 a.m. and you’re bleary-eyed and panicked, which is not the time to be making difficult decisions; it’s the time to be clicking the button.

Chadd: A hundred percent agree. I think the lack of visibility, this being a solution, less a problem I’m trying to solve [laugh] on my own is, I think, one area no one’s really tackled in the industry, especially around data protection. I will say people have done this on-prem at a decent level, but it just doesn’t exist inside the public cloud today. Clumio Discover, as an example, is one thing that we just heard constantly. It was like, “Give me global visibility to see everything in one single pane of glass across all my accounts, ensure all of my data is protected, optimized the way that I’m spending in data protection, identify if I’ve got massive outliers or huge consumers, and then help me restore faster.”

And the cool part with Discover is that we’re actually giving this away to customers for free. They can go use this whether they’re using AWS Backup or us, and they can now see all of their environment. And at the same time, they get to experience Clumio as a solution in a way that is vastly different than what they’re experiencing today, and hopefully, they’ll continue to expand with us as we continue to innovate inside of AWS. But it’s a cool value for them to be able to finally get that visibility that they’ve never had before.

Corey: Did, you know, that AWS users can have multiple accounts and have resources in those accounts in multiple regions?

Chadd: Oh, yeah. Lots of them.

Corey: Yeah. Because—the reason that you know that, apparently, is that you don’t work for AWS Backup where, last time I checked, there are still something like eight or nine regions that they are not present in. And you have to wind up configuring this, in many cases, separately, and of course, across multiple accounts, which is a modern best practice: separate things out by account. There we go. But it is absolutely painful to wind up working with.

Sure, it’s great for small-scale test accounts where I have everything in a single account and I want to make sure that data doesn’t necessarily go on walkabout. Great. But I can’t scale that in the same way without creating a significant management problem for myself.

Chadd: Yeah, just the amount of accounts that we see in enterprises is nuts. And with people managing this at an account level, it’s unbearable. And with no visibility, you’re doing this without really an understanding of whether you’re successfully executing this across all of those accounts at any point in time. And so this is one of the areas that we really want to help enterprises with. It’s, not only make the protection simple but also validate that it’s actually occurring. Because I think the one thing that no one likes to talk about in this is the whole compliance game, right? Like—

Corey: Yeah, doing something is next to useless; you got to prove that you’re doing the thing.

Chadd: Yeah. I got an auditor who shows up once a quarter and says, “Show me this backup.” And then I got to go fumble to try to figure out where that is. And, “Oh, my God. It’s not there. What do I tell the guy?” Well, wouldn’t it be nice if you had this global compliance report that showed you whether you were compliant, or if it wasn’t—which, you know, maybe it wasn’t for a snapshot that you created—at least would tell you why. [laugh]. Like, an RPO was exceeded on the amount of time it took to take the snapshot. Okay, well, that’s good to know. Now, I can tell the guy something other than just make
something up because I have no information.

Corey: So, you’d have multiple snapshots in flight simultaneously; always a great plan. Talk to me a little bit about Discover, your new offering? What is it? What led to it?

Chadd: I love talking to customers, for one, and we spend a lot of time understanding where the gaps exist in the public cloud. And our job is to help fill those gaps with really cool innovation. And so the first one we heard was, “I cannot see multiple services, regions, accounts in one view. I had to go to each one of the services to understand what’s going on in it versus being able to see all of my assets in one view. I’ve got a lot of fragment reporting. I’ve got no compliance view whatsoever. I can’t tell if I’m over-protecting or under-protecting.”

Orphan snapshots are the bane of many people’s existence, where they’ve taken snapshots at some point, deleted an EC2 instance, and they pay monthly for these things. We’ve got an orphan snapshot report. It will show you all of the snapshots that exist out there with no EC2 instance associated to it, and you can go take action on it. And so, what Discover came from is customers saying, “I need help.” And we built a solution to help them.

And it gives them actionable insights, globally, across their entire set of accounts, across various different services, and allows them to do a whole bunch of fun stuff, whether it’s actionable and, “Help me delete all my orphan snapshots,” to, “I’ve got a 30-day retention period. Show me every snapshot that’s over 30 days. I’d like to get rid of that one, too.” Or, “How much are my backups costing me in snapshots today?”

Corey: Yeah, today, the answer is, “[mumble].”

Chadd: [laugh]. And imagine being able to see that with, effectively, a free tool that gives you actionable insights. That’s what Discovery is. And so you pair that with Clumio Protect, which is our backup solution, and you’ve got a really awesome solution to be able to see everything, validate it’s working, and actually go protect it, whether it’s operational recovery, or a true airgap solution, of which it’s really hard to pull off in AWS today.

Corey: What problem that’s endemic to the backup space is that from a customer perspective, you are either invisible, or you have failed them. There are remarkably few happy customers talking about their experience with their backup vendor.
So, as a counterpoint to that, what do the customers love about you, folks?

Chadd: So, first and foremost, customers love the support experience. We are a SaaS offering, and we manage the backups completely for the end-user; there’s no cloud infrastructure the customer has to manage. You know, there’s a lot of these fake SaaS offerings out there where I better deploy a thing and manage it in my tenancy. We’ve created an experience that allows our support organization to help customers proactively support it, and we become an extension to those infrastructure teams, and really help customers to make sure they have great visibility and understanding what’s going on in their environment. The second part is just a completely new customer experience.

You’ve got simplicity around the way that I add accounts, I create a policy, I assign a tag, and I’m off and running. There’s no management or hand-holding that you need to do within the system. The system scales to any size environment, and you know, you’re off and running. And if you want to validate anything, you can validate it via compliance reports, audit reports, activity reports. And you can see all of your accounts, data assets, in one single pane of glass, and now with Clumio Discover, you get the ability to be able to see it in one single view and see history, footprint, and all sorts of other fun stuff on top of it. And so it’s a very different user experience than what you see in any other solution that’s out there for data protection today.

Corey: Thank you so much for taking the time to speak with me today. If people want to learn more about Clumio and kick the tires for themselves, what should they do?

Chadd: So, we are on AWS Marketplace, so you can get us up and running there and test us out. We give you $200 of free credits, so you can not only use our operational recovery, which is, kind of, snapshot management, similar database backup, which is free. You can check out Clumio Discover, which is also free, and see all of your accounts and environments in one single pane of glass with some awesome actionable insights, as we mentioned. And then you can reach out to us directly on clumio.com, where you can see a whole bunch of great content, blog posts, and the like, around our solution and service. And we’re looking forward to hearing from you.

Corey: Excellent. And we will, of course, throw links to that in the [show notes 00:29:57]. Thank you so much for taking the time to speak with me today. I appreciate it.

Chadd: Well, thank you so much for having me. I had an awesome time. Thank you.

Corey: Chadd Kenney, VP of product at Clumio. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with a very long-winded comment that you accidentally lose because the page refreshes, and you didn’t have a backup.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Aviad

Aviad Mor is the Co-Founder & CTO at Lumigo. Lumigo’s SaaS platform helps companies monitor and troubleshoot serverless applications while providing actionable insights that prevent business disruptions. Aviad has over a decade of experience in technology leadership, heading the development of core products in Check Point from inception to wide adoption.

Links:

  • Lumigo: https://lumigo.io/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the Enterprise (not the starship). On-prem security doesn’t translate well to cloud or multi-cloud environments, and that’s not even counting IoT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IoT devices, detects these threats up to 35 percent faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at extrahop.com/trial.

Corey: This episode is sponsored in part by our friends at Lumigo. If you’ve built anything from serverless, you know that if there’s one thing that can be said universally about these applications, it’s that it turns every outage into a murder mystery. Lumigo helps make sense of all of the various functions that wind up tying together to build applications.

It offers one-click distributed tracing so you can effortlessly find and fix issues in your serverless and microservices environment. You’ve created more problems for yourself; make one of them go away. To learn more, visit lumigo.io.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I periodically talk about how I bolt together a whole bunch of different serverless tools in horrifying ways to write my newsletter every week. At last count, I was up to something like four API Gateways, twenty-nine Lambda functions, and counting. How do I figure out if something’s broken in there? Well, honestly, I just keep clicking the button until it works, which is a depressingly honest story.

Now, that doesn’t work for everyone. Today’s promoted episode is brought to us by Lumigo. And my guest today is Aviad Mor, their CTO, and co-founder. Aviad, thanks for taking the time to suffer my slings and arrows.

Aviad: Thank you, Corey. I’m very glad to be here today.

Corey: So, let’s begin at, I guess, the very easy, high-level question: what is Lumigo and is ‘loom-ago’ an accepted alternate pronunciation?

Aviad: [laugh]. So, Lumigo is a monitoring and debugging platform for serverless environments. And yes, you can call it whatever you want as long as it’s Lu-mi-go. What we do is we integrate with the customer’s AWS account, we do a very quick connection to its Lambdas, and then we’re able to show him exactly what’s going on in his system: what’s going well, what’s going wrong, and how we can fix it.

Corey: So, let’s make sure that we hit a few points here at the beginning. It is AWS specific at this time?

Aviad: Yes, it is. We’re not officially exclusive with AWS, but right now we see the most interesting serverless environments in AWS, so it’s a pretty easy call. But we are keeping our eye open to, you know, Google, Microsoft, even Oracle.

Corey: Oh, Oracle Cloud has some phenomenally interesting serverless stories that I don’t think the world is paying enough attention to yet. But one of these days, I’m hoping that that’s going to change just because they have so much savvy locked up in that platform.

Aviad: Right. They do have serverless functions. Yeah, so.

Corey: They acquired the iron.io folks a while back, and those people were way ahead of Lambda at the time.

Aviad: Right, right. So, we’re waiting for the big breakout of serverless in Oracle, and then we’ll build the best monitoring solution for them.

Corey: So, something else, I think, that you have successfully navigated as far as, I guess, the traps that various observability tooling falls into, you also talk on your site about monitoring AWS Lambda as the center around which everything winds up being captured. You also, of course, integrate with the things that tied directly into it, such as API Gateway—or ‘opi-gateway,’ as I’m sure they mispronounce it at AWS—but that’s sort of where you stop. You don’t also show all of the container workloads that folks are running, and, “Oh, hey. While we have access to your API, here’s a whole story about ECS, and RDS, and all the rest.” And eventually, it feels like everything, in the fullness of time, tries to become Datadog version two.

And that always drove me nuts because I want something that’s going to talk to me specifically about what it is that I’m looking at in a serverless application context, not trying to be all things to all people at once. Is that a fair assessment of
the product strategy you folks have pursued?

Aviad: Right. So, we’re very focused on serverless. We think there’s a lot of interesting things that we can do there, and we’re actually seeing more and more use cases of serverless. And it is important to say that when we say serverless, it’s very clear what is serverless. So Lambda, of course, and API Gateway, DynamoDB, S3, and so on.

There’s a lot of services in data ecosystem, and seeing them all being tied together in a serverless cloud application, we’re able to do all of that to monitor it; not only monitor it at the high level, but also get into the details and show you things which are very specific because this is what we do all day, and sometimes all night. And then there are those boundaries of where do we go beyond serverless. So, there are some hybrid environments out there. And when I say ‘hybrid,’ the easy hybrid, which is you have two different applications which just happen to be on the same AWS account; one of them is completely serverless, and then the other one is EC2. So, that’s kind of hybrid.

But the more interesting hybrid is those applications which start with an API Gateway in the Lambda, and then are directly connected to something else, which is maybe Fargate, ECS, EKS, and so on. So, we are very much focused on serverless, but we are getting also a lot of requests from our customers, “So, show us, also, the other parts.” We’re starting to look at that, but we’re not losing our focus. Our focus is still very much on the serverless while allowing you to tie together if you do have some other aspects in your environment to see them all together.

Corey: So, you’ve done a number of things that I would consider best in class as you’ve gone through this. First and foremost, let’s begin with the easy stuff. It doesn’t appear that your dashboard, your tooling itself, is purely serverless itself. I can tell this because when I click around in your site, the site loads super quickly. It’s not waiting for cold starts or waiting for the latency inherent to the Lambda.

It’s clear that you have not gone so far down the path of being, I guess, religiously correct around everything must be serverless all the times in favor of improving customer experience. That’s something that I’ve seen a number of different vendors fall into the trap of, of, “Why is the dashboard so slow to load?” “Ah, because everything is itself a Lambda function.” Is that accurate, or if you just found a way to improve Lambda [laugh] function performance in an ungodly way?

Aviad: [laugh]. We are serverless—we call ourselves serverless first, but the customer is always—he’s really the first. So, if there’s a place where serverless is not the best solution, we’re going to use whatever is the best solution. But the truth is, we’re, I’d say, something like 99% serverless. And specifically, anything which is dashboard-facing customer-facing, that’s actually completely serverless.

So, we did have to put in a lot of work, but also, I have to say that AWS went a very long way, like, in the last two years, allowing us to give much better latencies in different parts of the dashboard. So, all of that is serverless, and it goes together with the new features of Lambdas, and API Gateways, and a lot of small things we had to do in order to provide the best experience to the customer.

Corey: The next thing that I think was interesting, as far as, I guess, capturing the way in which people use these things. One of the earliest problems I had, in the early days of these, I guess, new breed of serverless tools was getting everything instrumented correctly. It felt like it took in some cases more time to get the observability pieces working than it did to write the thing in the first place. So, you’re integrating out of the gate with a lot of the right things as best I can tell. Your website advertises that you integrate with the Serverless Framework, you integrate with a bunch of other [processes 00:07:52] as well. Chalice, which I haven’t seen used in anger too much, but okay; Terraform, which everyone’s using; Stackery, et cetera. Is AWS’s SAM on the list as well?

Aviad: Yes, it actually is. And once we started seeing more and more users using SAM, we had to provide a way to allow them to easily do the integration. Because one of the things that we learned is, no, our users are developers and, just like you said, they don’t want to spend time on anything, which is not like doing the thing that they want to do. Especially in serverless, because the whole serverless premise is work on what you do best, and don’t spend time on everything else. So, we actually spend a lot of time ourselves in order to make the integrations as easy as possible, as quickly as possible, and that also means that working with a lot of different tools to fit all the different environments our users are using out there.

Corey: It looks like you’re doing this through the application—judiciously—of a bunch of custom layers. In other words, whatever you wind up using winds up being built as an underpinning of the existing Lambda functions, so it’s super easy to wind up grabbing the necessary dependencies for anything that you folks support without having to go back and refactor existing applications. Is that directionally correct?

Aviad: Right. That’s correct. We’re using layers in order to, on one hand, do this deep integration we do with the Lambda, allowing us to do different instrumentations, collecting the data that’s being passed into the Lambda, being passed out of the Lambda, on one hand. But on the other hand, so the developer doesn’t have to make any code changes, and he can do whatever changes he wants to do. Doesn’t have to think about Lumigo at any point, and serverless layer does everything for him automatically.

Corey: How do you handle support of the Lambda@Edge functions, which seem an awful lot like regular Lambda functions, except they’re worse in nearly every single way, every time I’ve tried to use them? In fact, in my experience, the best practice has been to immediately rip out Lambda@Edge and replace it with something else. Recently, it was formally disclosed that they only ran in a subset of 13 regional cache locations, and they still took a full CloudFront distribution update cycle every time you did a deployment, which dramatically slowed everything down for deploying it; they were massively expensive to run at significant scale, and they would log to whatever region was closest so it was a constant game of whack a mole to figure out what was going on. But, you know, other than that, they were great. How do you approach those?

Aviad: Lambda@Edge are not very easy to use, and they’re, like, let’s say they’re full of surprises [laugh] because not everything they do is exactly what you find in the documentation. But again, since our users are using them, we had to make sure that we give them proper support. And giving them the proper support—other than running and collecting the data—is things that you mentioned, like the fact that it will log to the specific region it’s running in, so you have to go and collect all this data from different places, and you don’t really know exactly where it’s going to run. So, the main thing here is just to make things easy. It’s a bit of a mess when you’re looking at it directly, and taking all the information, putting it in one place so you as a user can just go ahead and read it and you don’t care where it’s running and what it’s doing, that was the main challenge which we worked on and added to the product.

Corey: So, across the board, it seems like you folks have been evolving in lockstep with the underlying platform itself. Have you had time to evaluate their new CloudFront Functions, I believe is what they’re calling it. Or is it CloudFront Workers? I can never quite keep it straight; between all the different providers, all the words start to sound alike. But the thing that only runs for a millisecond or two, only in JavaScript, only in the actual all the CloudFront edge locations, et cetera, et cetera. Rather than fixing Lambda@Edge, they decided to roll something completely different out, and I haven’t looked at anything approaching the observability story yet because I’m still too angry about it.

Aviad: [laugh]. Right. So, there’s a lot of things coming out, and we’re also very close partners with AWS, so in many cases, we’re actually beta users of new services or new functionality in Lambda. And one of the hardest parts is—and then we cannot spend all our time checking everything new. So, this is one of the things which is still in the to-do list; we’re going to check it very close, in a very close time.

I think it’s interesting to see how we can actually use it and is it as quickly as they say. What they say usually works; we’ll see if it works already today, or do we have to wait a little bit until it works exactly like they said. But no, that’s one of the things that are on my to-do list. I’m really looking forward to checking it out.

Corey: So, it looks like once I set this up and it starts monitoring my account—or observing my account. I know I know, observability is just hipster monitoring, but no one agrees with me on that, so I’m going to keep rolling with it anyway just to irritate people—it looks like I can effectively more-or-less click a button, and suddenly, you will roll out an underlying Lambda layer to all of my existing Lambda functions. How does that get maintained whenever I wind up, for example, doing a new deployment with the serverless framework or something like it that isn’t aware of that underlying layer, so it—presumably—would revert that layer itself in the definition? Or am I misunderstanding how that works?

Aviad: No, no. You’re actually getting it right. So, unless you, for example, are using a serverless plugin, so this is an integral part of your deployment, one of the things that we need to do is to automatically identify that a deployment is happening so we can automatically update the Lambda layer to be the right one, so you won’t miss anything. And this deep integration, which is happening without the user having to know anything about it, this is, I think, one of the most important parts because in serverless, as you know, you have so many components, and very easily you can reach, you know, hundreds of Lambdas, which are things that we’re seeing. So, if a user has to take care and maintain something across a hundred Lambdas or more, you can be sure that it won’t be maintained because he has, like, something much more important to do.

So, behind the scenes, immediately as the deployment is happening, we can recognize that it’s happening, and then update the layer that’s required. And by the way, now the layers have a new part called extensions, which allow us and everybody else to do a lot more with those layers, basically allowing the code to run in parallel to the Lambda. So, this is a new thing that AWS has started to roll out, and we think will allow us to give even better experience to our users.

Corey: Let’s have a look across the, I guess, the ecosystem of have different approaches to this stuff. One thing that has always annoyed me about a whole raft of observability and monitoring tools is they wind up charging me whatever it is they charge me; it’s generally fine—and I don’t really have a problem with that. You know, in advance going in what things are going to cost you. Incidentally, what is your pricing model?

Aviad: So, our pricing model is according to the number of invocations you have. So, we have basically two models which we’re using right now, and each one can decide what he wants better. So, if you want to know in advance exactly how much you’re going to pay, you can go with the tiered model meaning, I want to pay for, let’s say, a million invocations each month, and then you’re sure that you’re paying exactly for what you have a budget for. And it’s always related to how much your AWS account is working, so similar to how much you’re paying for your Lambdas. And then there’s another way, which is dynamic pricing, which is very similar to serverless payment.

So, it’s really according to the number of invocations you have; you don’t need to decide in advance, and at the end of each month, according to the number of invocations you have, you get the bill. And that way it’s not based on the invocation in general; it’s exactly according to the number of invocations.

Corey: And let’s be clear, if I wind up exceeding the number of invocations under my plan, it just stops tracing and observing these things, it doesn’t break my app.

Aviad: Yeah, right. [laugh].

Corey: Always good to triple-check those things. It seems like that might hurt.

Aviad: That’s very important. You’re totally correct. And, yeah, we never do anything bad to your Lambdas. That’s written on the top of our door: “Never hurt a Lambda.” And we make sure that nothing bad happens, we just stopped collecting data.

And by the way, even as you pass your limit, we still collect the basic metrics so you can see what’s going on in your system. But you won’t be able to see the rich information, all the information that allowing you to do the debugging, or seeing the full traceability end-to-end of all the invocations and see how they’re connected to each other.

Corey: This episode is sponsored in part by Thinkst. This is going to take a minute to explain, so bear with me. I linked against an early version of their tool, canarytokens.org in the very early days of my newsletter, and what it does is relatively simple and straightforward. It winds up embedding credentials, files, that sort of thing in various parts of your environment, wherever you want to; it gives you fake AWS API credentials, for example. And the only thing that these things do is alert you whenever someone attempts to use those things. It’s an awesome approach. I’ve used something similar for years. Check them out. But wait, there’s more. They also have an enterprise option that you should be very much aware of canary.tools. You can take a look at this, but what it does is it provides an enterprise approach to drive these things throughout your entire environment. You can get a physical device that hangs out on your network and impersonates whatever you want to. When it gets Nmap scanned, or someone attempts to log into it, or access files on it, you get instant alerts. It’s awesome. If you don’t do something like this, you’re likely to find out that you’ve gotten breached, the hard way. Take a look at this. It’s one of those few things that I look at and say, “Wow, that is an amazing idea. I love it.” That’s canarytokens.org and canary.tools. The first one is free. The second one is enterprise-y. Take a look. I’m a big fan of this. More from them in the coming weeks.

Corey: So, the pricing makes perfect sense, and that is in line with what I would expect, but the thing that irritates me then is, “Great. I know what I’m going to be paying you folks on a monthly basis, and that’s fine.” And then I use the monitoring tool and it cost me over three times as much in AWS charges, both direct and indirect, where it’s, “Oh, now CloudWatch is going to suddenly be the largest component of my bill and data transfer for sending everything externally winds up spiking into the stratosphere.” What’s your experience been around that?

Aviad: So, since we are collecting data and we are doing API calls, it will affect your AWS bill. But because we don’t want to irritate you, or anybody else, we are putting a lot of focus to see that we’re doing the absolute minimal possible effect on your system. So, for example, as we’re collecting data from your Lambda, we’re doing our best to add only milliseconds to the running time of your Lambda so you don’t end up paying a lot more for the runtime. Or for the API calls or data transfer, we have a lot of optimizations that we did, so the billing on your AWS account is really very, very small; it’s not something that you will notice. And sometimes when people do ask us, we go together with them into their account and show them exactly how their billing was affected by Lumigo so they’ll have assurance that nothing crazy is going on there.

Corey: Which is I guess one of the fundamental problems of the underlying platform itself. I have a hard time blaming you for any of this stuff. This is the perpetual joyless story of getting to work with a variety of AWS services. It’s not something that I see that you folks have a good way around just on basis of how the underlying platform works.

Aviad: Yeah. And then there are a lot of different prices for a lot of small things that you do, and you need to be able to collect it all in order to have the big picture of the effect. And yeah, we don’t have a silver bullet for it, but we can show exactly where we’re going, what we’re adding, to show how low it is.

Corey: One of the things that I think is not well understood for folks who are not into the serverless ecosystem is just how these applications tend to look. In most organic environments, you’ll see a whole bunch of Lambda functions that are all
tied together with basically spit and baling wire. They talk to each other, either directly on the back end—which is an anti-pattern in many respects, let’s not kid ourselves—or alternately, they’re approaching through a lens of, we’re going to now talk to each other through hardened REST APIs, which is generally preferred, but also a little finicky to get up and running. So, at some point, you have a request come in, and it winds up bouncing around through a whole bunch of different subsystems. Tracing, and a lot of the observability story around serverless is figuring out, all right, somewhere in that rat’s nest, it winds up failing.

Where did it break? What was it that actually threw the exception? What was it that prevented something from working? Or alternately, adding latency: where is the bulk of the time serving that request being spent? And you would think that this is the sort of thing that AWS could handle itself.

And they’ve tried with their X-Ray distributed tracing option, which more or less feels like a proof of concept demonstrating what not to do. And if you take a look from their application view, and all the rest, it is the best sales pitch I can possibly imagine for any of the serverless monitoring tools that I’ve seen because it is so badly articulated. You have to instrument all of your stuff by hand. There’s none of this, oh, I’ll look at it and figure out what it’s talking to and build an automated trace approach, the way that Lumigo does. And that’s always frustrated me because I shouldn’t have to turn every weird analysis into a murder mystery. Am I missing something obvious in how this could be done using native tools directly, or is it really as bad as I believe it is?

Aviad: [laugh]. I won’t say it as bad as you’re saying it is. I think X-Ray is a great place to start with. So, if you have, like, just a few Lambdas; you’re starting to check out the serverless world, X-Ray can be good enough if you don’t want to start with a third-party tool right at the beginning. But then as it gets a little bit complex, it’s going to get hard, especially if you’re trying to do it yourself.

That’s usually the wow part when people start using Lumigo when we show them a demo, is seeing how everything is tied together. So, once you see how everything is tied together: the whole system, which components are talking to each other, and how they’re affecting each other. And for example, if one of them goes down, does it mean that the whole system now is not working, or maybe, eh, wasn’t that important, and everything is working. I’ll fix it next week. But I think the most important part is actually what we call the transactions.

So, as you said, there’s an API call at the very beginning with an API Gateway or AppSync, and then it can go through dozens of components. Some of them are not directly related, so it’s like, Lambda calling, putting something into a DynamoDB, which triggers a DynamoDB stream. And then another Lambda is being called, and so on, and so on. It’s crucial to be able to see how everything is connected, both very visually, so you can understand it. There’s only so much you can understand when looking at a list as a human being, right?

You need to see it visually how everything is connected. But then after you understand how everything is connected in this specific transaction, if, for example, you have an issue in a specific invocation, you need to understand the story of that invocation. And maybe you’re looking at a Lambda which starts to throw an exception, and you didn’t change anything in its code today, yesterday, or the day before that, so take care of that exception, but the root cause is probably not in that Lambda, it’s probably upstream. So, you need to be able to understand exactly what was the chain of events, all the calls being made until that specific Lambda was called to see the data being passed, including the data that Lambda maybe passed to a third-party API—like Stripe or PayPal—and what it got in return. And only when you’re able to see all of that you’re able to solve an issue quickly, not a murder mystery like it might be. Time over time without having to think about how will I make sure that I make all the code changes in order to keep getting these transactions?

Corey: So, taking a look at the somewhat crowded space—if I’m being perfectly honest with you—of the entirety of, let’s call it the serverless observability space—or ‘observerless,’ as I’m a big fan of calling it—what is it the differentiates Lumigo from a number of other offerings that people could wind up pulling out of the hat?

Aviad: Right. So, that’s a great question. And every time somebody asks me, the first thing I can say is, the more I see people getting into this space, I think that that’s a great sign. Because that means there’s more serverless activity, there’s more companies doing serverless and it means that our serverless space is interesting. People see an opportunity there, and they want to try and solve the issues that we’re seeing there.

And I think that there’s a few things: one of them is the serverless expertise. So, if you look at a lot of the big companies—like I’ll mention Datadog and New Relic—they’re doing a lot of great things, but in the end, in the serverless environment, there are very specific things which you need to know, have to do in order to be able to do that distributed tracing, the distributed tracing which allows you to correlate specific transactions together and then bring in those metrics which are relevant and bring in the logs which are relevant for a specific transaction. That’s a lot of hard work which we put in in order to be able to do the transactions with a distributed tracing in the best way possible, and then showing it to you in the simplest way possible. And today, I think that Lumigo does that in a very good way. And also, if we’re looking around at other players, which are not only the big ones, also players, which are doing more specifically serverless, I think that if we’re looking at companies which are very focused on serverless, and serverless is the thing that they do, you’ll still see that Lumigo is the one which is doing serverless the most, let’s call it.

So, as serverless is expanding, we’re still not becoming generic—something that we mentioned before—and this allows us not only to do the best distributed tracing but also allow us to show you, out of the box, a lot of issues which might be hiding in your environment. So, it’s not only, “Okay, you have an exception here,” it’s also more specific things to serverless. Like for example, because it’s event-driven, so sometimes you’ll get duplicate events that Kinesis or SQL might send you over and over the same event. The fact that we can show you it automatically and put a spotlight on it can save you a lot of time in trying to understand why things are not working the way you think they should be working. And allowing us to scan your environment and show you misconfigurations which are specific to serverless, this is the kind of things that once you use Lumigo, you get automatically without having to do anything special and that can save you a lot of time.

Corey: I think that’s a relatively astute position to take. I’m a big believer in getting the right tool for the right job. I don’t necessarily want the one single pane of glass to look at everything. I think that is something that is very often misunderstood. Yeah, I might be using three or four different clouds in different ways.

I don’t need to see a roundup of all of them; I don’t necessarily care what the billing looks like on all of them; I don’t necessarily want to spend my time thinking about different aspects of these things juxtaposed to one another, and it’s a pain in the butt to have to sort through to find the thing I actually care about. So yeah, on some level, I definitely want there to be a specific tool. And let’s be clear, you have a terrific stack of things that you integrate with for alerting, for opening tickets, for remediation—or issues, as the case may be. Nomenclature is always a distraction. Don’t at me—but yeah, across the board, I see that you’re doing a lot of things right that if I were going to be entering the space, I would make a lot of those decisions very similarly. And then expect to hear it from the internet. You’ve been around for years now and are continuing to grow. What’s next for you, folks?

Aviad: So, that’s a great question, which I asked myself every morning. I’ll actually take together the two things that you mentioned. One is how we’re focused on serverless, and the second is where do we want to grow from there? And when you do this great focus, you have to make sure that what you’re focusing on is big enough. So, as we’re growing, we’re very happy to see that the serverless is growing with us.

We’re seeing more and more places using serverless. We see a lot more users, companies, developers going into serverless. And we see new types of users. So, it’s not only those bleeding-edge technologies that people want to use, and they are really trying to find out how they can use it. We’re seeing more and more places, for example, enterprises that had maybe one architect in the beginning that said, “Okay, I’m going to use serverless.”

And now a year or two afterwards, they see that it’s working, and it saving them money. They’re able to build faster, and now it’s both spreading virally to other teams which are starting to use that, and also the initial project, which was started two years ago, is now growing and becoming bigger and more complex. And also, that team which was just starting with serverless two years ago now has maybe a second and third product. So, what we’re doing is we’re looking how we can give serverless better and better monitoring for the new services that are entering that field. And also, we’re very strong believers that developers today are doing much of that monitoring—or observability, you can choose whatever you want—and that means that it goes all the way into debugging.

So, we think that doing those two together, bringing together the monitoring and debugging is a great opportunity for our users just to save them more time because it’s the same person who’s going to do both those things, and trying to keep being best of breed in serverless, and doing those two together, I think that’s going to be hard. And that’s exactly the challenge that we’re taking, and we want to see how we’re doing it to the best.

Corey: And I think that that is probably the best way to approach it. If people want to learn more about what you’re up to, how you view these things, and ideally, kick the tires on Lumigo and see for themselves, where can they find you?

Aviad: So, easiest thing you can do, just search for Lumigo in Google, you’ll get to lumigo.io. And from there, it’s very easy to try us out.

Corey: And we will, of course, put links to that in the [show notes 00:31:24]. Thank you so much for taking the time to speak with me today. I really appreciate it.

Aviad: Thank you, Corey. It was great fun and looking forward for the next time.

Corey: Absolutely. Aviad Mor, co-founder and CTO at Lumigo. I’m Chief Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you hated this podcast, please leave a five-star review on your podcast platform of choice along with a long rambling comment telling me how very wrong I am on the wonder that is Lambda@Edge.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Nathen
Nathen Harvey, Cloud Developer Advocate at Google, helps the community understand and apply DevOps and SRE practices in the cloud.

Nathen formerly led the Chef community, co-hosted the Food Fight Show, and managed operations and infrastructure for a diverse range of web applications.

Links:

  • cloud.google.com/devops: https://cloud.google.com/devops
  • 97 Things every Cloud Engineer Should Know: https://shop.aer.io/oreilly/p/97-things-every/9781492076735-9149
  • Twitter: https://twitter.com/nathenharvey

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Thinkst. This is going to take a minute to explain, so bear with me. I linked against an early version of their tool, canarytokens.org in the very early days of my newsletter, and what it does is relatively simple and straightforward. It winds up embedding credentials, files, that sort of thing in various parts of your environment, wherever you want to; it gives you fake AWS API credentials, for example. And the only thing that these things do is alert you whenever someone attempts to use those things. It’s an awesome approach. I’ve used something similar for years. Check them out. But wait, there’s more. They also have an enterprise option that you should be very much aware of canary.tools. You can take a look at this, but what it does is it provides an enterprise approach to drive these things throughout your entire environment. You can get a physical device that hangs out on your network and impersonates whatever you want to. When it gets Nmap scanned, or someone attempts to log into it, or access files on it, you get instant alerts. It’s awesome. If you don’t do something like this, you’re likely to find out that you’ve gotten breached, the hard way. Take a look at this. It’s one of those few things that I look at and say, “Wow, that is an amazing idea. I love it.” That’s canarytokens.org and canary.tools. The first one is free. The second one is enterprise-y. Take a look. I’m a big fan of this. More from them in the coming weeks.

Corey: This episode is sponsored in part by our friends at Lumigo. If you've built anything from serverless, you know that if there's one thing that can be said universally about these applications, it's that it turns every outage into a murder mystery. Lumigo helps make sense of all of the various functions that wind up tying together to build applications.

It offers one-click distributed tracing so you can effortlessly find and fix issues in your serverless and microservices environment. You've created more problems for yourself. Make one of them go away. To learn more, visit lumigo.io.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’m joined this week by Nathen Harvey, a cloud developer advocate at a small startup called Google. Nathen, thank you for joining me.

Nathen: Hey, Corey. It’s really great to be here.

Corey: We’ll get to the Google bits in a little, but first, I want to start back in the beginning with your origin story. It turns
out, for example, that you were at a lot of places, and the first thing going through your history that I really recognized was way back at the end of 2009, where you were the web operations manager at Custom Ink. They’re a t-shirt company—and other apparel—that I’ve been using for three years now for the charity t-shirt drive here, as well as other sundry things. Longtime listeners of the show might remember we had Ken Collins on to talk about Ruby in Lambda and other horrifying things, before it was cool.

Nathen: Yes, indeed, I was at Custom Ink. And, you know, you talk about them being a t-shirt company, and I don’t know… maybe I’m still a shill for Custom Ink, but I really look at them as an experience company. And you’ve recognized that yourself. They produce and help people, really encourage that group and experiences, and really drive what does it mean to connect with other humans, and how can you do that through custom apparel? To me, that’s what Custom Ink has always been about. They’re not selling t-shirts; they are selling an experience.

Corey: In my case, I view them as a t-shirt company because, let’s be fair here, I wind up doing charity t-shirt drives, and they’ve always been extremely supportive of—well, there’s really no other way to put this—my ridiculous nonsense. The year I had linked campaigns of the ‘AMI has three syllables’ shirt that was on sale, and then for the Amazonians, ‘ah-mi’ is how it’s pronounced instead and that one was $10 more because there’s a price to being wrong. And all proceeds, of course, went to benefit the charity of the year. And that was a fun thing. And I talked to a number of other folks on this, and they look at me very strangely, and Custom Ink didn’t even blink.

Nathen: Right, right. Absolutely. Absolutely.

Corey: And yes, they said lots of other apparel, but for whatever reason, it seems that sending out complicated multiple options of things that need each hit minimum order quantities to print during a fundraiser, and the fact that I don’t have to deal with the money because they just wind up sending it over directly. It’s just easier. It’s one of those things where back when I was a single person who was doing this stuff, I didn’t have to worry about it. Now that I’ve grown and my needs have multiplied, I still like doing business with them. Great folks.

Nathen: Absolutely. And that’s exactly what I mean by—like, they’ve sold you on that experience. That’s why you continue to do business with them. It’s not just because of the t-shirts. It’s the whole package that goes along with it.

Corey: And then in 2012, the world didn’t end. But yours kind of did because you stopped working at Custom Ink and went to another company called Chef. You were there for a little over six years. You started off as a community director and then became the VP of Community Development. And I think you did an amazing job, but first tell me about that, then I will give my hot take.

Nathen: All right, great. I’m always up for the hot takes. So listen, Chef was an amazing community of people. Oh, it was also a company. And so I really fell in love with—while I was at Custom Ink, actually, we were using Chef, and I fell in love with the community.

And I was doing a lot of community support, running my own podcast, or participating with some co-hosts on a podcast called the Food Fight Show back in the day—it was all about Chef—running meetups and so forth. And at one point I decided, you know, what I should do maybe is stop being on call and start supporting this community full time. And that’s exactly what I did. I went to Chef and yes, as you mentioned, spent just over six years there, or just about six years there, and it was really, really an incredible time. Lots of hugs to be given, and just a great community in the DevOps space.

Corey: I took a somewhat, I guess, agreeing or disagreeing position. I was on the Puppet side of the configuration management debate, and it was challenging. And then, ah, I was one of the very early developers behind SaltStack because clearly, the problem with all of these things was that no one had written it correctly, and we were going to fix that. And it turns out no, no, the problem was customers the whole time. But that’s a separate debate.

So, I was never in the Chef ecosystem. That was the one system I never really touched in anger. And it’s easy to turn this into a, “Oh, you folks were the competition,” despite the fact I’ve never actually worked directly for either of those companies. But it was never like that because our real enemies were people configuring things by hand, for one because that’s unnecessary toil; don’t do it, and it was also just such an uplifting sense of community. Some of the best people I knew were in the Chef ecosystem, in the Chef orbit.

For a while, they’re, on some level—and this is something I’d love to get your thoughts on—it seems that a failure mode that Chef exhibited was hiring directly from its community, where if someone was a big fan of Chef, start a stopwatch, they’re going to be working there before the month is out.

Nathen: I think that Chef, the company definitely pulled a lot of community members into the organization. And frankly, when the company started, that was really, really great because it was an early startup. And as the company grew, it was still wonderful, of course, to pull in people from the community to really help drive the future direction, how our customers are using it. But like you said, there is a little bit of a challenge or concern when you start pulling too many of your most vocal supporters out of the community and putting them into the company, sometimes in places or roles where they didn’t have the opportunity to be as vocal, as big a champions for the product, for the services.

Corey: I think at some level, it was—again, it helps to have people who are passionate about the product working there, but on the other, it felt like over time, it wound up weakening the community in some respects, just because everyone who worked there eventually found themselves in a scenario of well, I work here, it’s what we do, and now I have to say nice things. It winds up weakening the grassroot story.

Nathen: Mmm. There’s definitely some truth to that, but I think there’s also some truths to just the evolution of community as you went from a community in the early days where there were a lot of contributors to over time—gratefully so—the community that—or sort of the proportion of the community that were consumers of Chef versus contributors to Chef, that balance changed. And so you had a lot more customers using the product. So, I don’t disagree with you, but I do think that it’s part of the natural evolution of community as well.

Corey: And all things must end. And of course, Chef got acquired, I believe, after you left. So, I mean, at that point, you left, they were rudderless and what else were they to do? And you went to Google. And that is always an interesting story because Google’s community interaction before the time you wound up there, and after—I don’t know that you were necessarily the proximate cause, but I’m going to hang that around your neck because it’s all been a positive change since then—look radically different.

Nathen: Yeah. Well, thank you. It is definitely not something that I should wear or carry alone, but going to Google was an interesting choice for me and I recognize that. And, you know, honestly, Corey, one of the things that drove me to Google was a good friend of mine, Seth Vargo. And just to kind of tie the complete throughline here, Seth and I worked together at Custom Ink, we’ve worked together at Chef, he left Chef and went to Hashi, and then went to Google. And the day after I knew that he was going to Google, I called him up and I said, “Seth, come on. Google’s so big. Why? Why? And how? I don’t understand. I don’t understand the move.”

Corey: I asked him many of the same questions back in episode three of this show. He was a very early guest when I was scared speechless having conversations. It’s improved since then, a couple hundred in. But yeah, very friendly; very open; very warm.

Nathen: Yeah. And, you know—

Corey: “Why are you at Google?” was sort of the next follow-on question there in that era.

Nathen: [laugh]. Yes, indeed. And I do think that Google, and specifically Google Cloud, has really taken to heart this idea that there’s a lot that we can learn from each other. And I don’t mean from each other within Google. Although, of course, we can learn a lot from each other.

But we can also learn a lot from our community, from our customer base. How are they using Google Cloud? How are they using technology to drive their business forward? These are all things that we can learn. It turns out, not every company has Google, and that’s a good thing.

Not every company should be Google or Google-sized, and certainly don’t have Google customers. And I think that it’s really important that we recognize that when we work with a customer, they’re the experts in their customers, and in their systems, and so forth.

Corey: A lot has changed with Google’s approach to, well, basically everything. It turns out that when you’re a company that is, what, 26 years old now—27, something like that—starting with humble beginnings and then becoming a trillion-dollar entity, things change. Culture change, your community changes, what you do changes, and that becomes something that I think is not necessarily fully appreciated or fully understood in some corners. But then 2018 hit. You went to Google; what did you do then?

Because it is such a large company that it is very difficult to know what any individual is up to there, and the primary means that I engage in the DevRel community space—specifically via aggressively shitposting on Twitter—isn’t really your means of interacting with the community. So, from that particular point of view, it’s, “Oh, yeah, he went to Google, and no one ever heard from him again.” What is it you say it is you do there?

Nathen: Yeah. So, for sure. What I do here as a cloud advocate, is I really focus in on kind of two areas, I would say: DevOps—and I recognize that is a terrible, terrible word because when I say it, we all think of different things, but I definitely focus on the DevOps—and then SRE practices as well, or Site Reliability Engineering. And specifically what I work on is how do we bring the principles and practices of DevOps, of SRE, into our customer base and into the community at large? How do we drive what is the state of the art?

How do we approach these particular topics? And so that’s really what I’ve been focused on since joining Google. Well, frankly, I was focused on that while at Chef, as well, maybe without the SRE bend so much, but certainly at Google SRE comes in, but it’s always—for the past decade for me—been about DevOps and how do we use technology to align the humans and work towards the business outcomes that we’re driving for?

Corey: And business outcomes become an interesting story in the world of cloud because it distills down, for a cloud service provider is, we would like people to use our cloud, more of it, in perpetuity. It is not a complicated business model—if I can be direct—because business models inherently are not. “Whatever it is your company does, we would like you to do it here.” And that turns into a bunch of differentiated services across the spectrum, in some cases hilariously so, when it turns into basically pick an English word, and there’s a 50/50 shot that’s part of a service name somewhere. But a lot of it distills down to baseline distinct primitives.

You’re talking about the DevOps aspect of it, which is—we talk about, is it culture? Is it tools? No, it’s a means to sell conferences, and books, and things like that. But what is it in the context of a cloud service provider? Specifically, Google because let’s be clear here, DevOps apparently for other providers is Azure DevOps. That’s right. It’s a service name, and DevOps Guru on the AWS side because everything is terrible.

Nathen: Absolutely. Look, I think that I used to snark that the only DevOps tool was the manager of DevOps. But the truth is that DevOps is… it is tooling, and it is culture, and to separate the two is really a fool’s errand. I think that your tooling amplifies your culture, your culture amplifies your tooling. Together, this is how we make progress.

Now, when it comes to Google, what do we mean when we say DevOps? Well, one of the good things is, shortly after I joined Google Cloud, Google Cloud acquired DORA, the DevOps Research and Assessment Organization.

Corey: Jez Humble, and Dr. Nicole Forsgren. And then, for all intents and purposes, they googled it. Relatively shortly thereafter, by which I mean, we never really heard from DORA again. In 2020, the “State of DevOps Report” didn’t exist, which was what they were famous for doing. And it was, “Oh, yep. That’s a Google acquisition all right.” Is that what
happened? Did I miss some nuance there?

Nathen: Yeah. Let’s talk about that. So first, you’re right, it was Dr. Nicole Forsgren, who founded DORA. So, when the acquisition happened, she came along to Google Cloud, Jez Humble came along through that acquisition as well. And frankly, what happened in 2020? Well, Corey, I don’t know if you noticed, but there was a lot happening in 2020, much of it not very good. I think when we look at the global scale, like, 2020 was not a great year for us—

Corey: It was a rebuilding year.

Nathen: Oh, all right, fair enough. Fair enough. [laugh]. A rebuilding year. But so here’s what happened with DORA, quite frankly. We—Google Cloud—continue to invest in that research program. And really, in a sense, 2020 was a rebuilding year, in that our focus was really about how do we help our customers and our community apply the lessons of DORA?

And so one of the things that we’ve done is we’ve released much of the research under cloud.google.com/devops, including right there, a DevOps quick check where, as a team you can go in and, using the metrics and the research program from DORA, you can assess, are you a low, medium, high, or elite performer?

And then beyond that assessment, actually use the research to help you identify which capabilities should my team invest in improving. So, those capabilities might be technical capabilities, things like continuous integration; it might be process or measurement capabilities, or in fact, cultural capabilities. So, all of these capabilities come together to help you improve your overall software delivery and operations performance. And so in 2020, the big thing that we did was release and continue to update this Quick Check, release the research, make it fully available. We’ve also spent some time internally on the program that, you know, is not super interesting to talk about on the podcast.

But the other thing that we did in 2020 with the DORA research program was update the ROI research, the return on investment research. This is something that maybe your listeners don’t care about, but their managers might care about, their CIOs, CTOs, CFOs might care about. How do we get money back on this transformation thing? And the research paper really digs into exactly that. How do we measure that? What returns can we expect? And so forth. So, that was
released in 2020.

Corey: I have a whole bunch of angry thoughts about a lot of takes in that space, but this is neither the time nor the place for me to begin ranting incoherently for an hour and a half. But yeah, I get that it was a year that was off, and now you’re doing it again, apparently, in 2021. And the one thing I never really saw historically because I don’t know if I’m playing in the wrong environments, or I’m certainly not the target [laugh] audience now, if I ever really was, but most years, I missed the release of the survey of where people can go to fill in these questions. I would be interested to know where that is now. And then I would be interested to know, how have you been socializing in that in the past? In other words, where are you finding these people?

Nathen: Yeah, for sure. So, the place you go to find the survey right now is cloud.google.com/devops, you’ll find a button on the page that says something like, “Take the survey,” or, “Take the 2021 survey.”

And what we’ve done in the past, and really what DORA has done in the past is use a number of different ways to get out information about the survey, when the survey is open, and so forth. Primarily Twitter, but also we have partners, and DORA historically has used partners as well to help share that the survey itself is open. So, I would absolutely recommend that you go and check out the survey because I’ll tell you what, one of the things that’s really interesting, Corey, over the years, I’ve talked to a bunch of people that have taken the survey, and that have read the State of DevOps Report that comes out each year, and some of the consistent feedback I’ve heard from folks is that simply taking the survey and considering the questions that are asked as part of the survey gives great insight immediately into how their team can improve. What things, what capabilities are they lacking? Or what capabilities are they doing really well with and they don’t need to make investments on? They can immediately see that just by answering and carefully considering the questions that are part of the survey.

Corey: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the Enterprise (not the starship). On-prem security doesn’t translate well to cloud or multi-cloud environments, and that’s not even counting IoT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IoT devices, detects these threats up to 35 percent faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at extrahop.com/trial.

Corey: Very often, in some cases looking at things like maturity models and the like, the actual report is less valuable than the exercise of filling it out and going through the process. I mean, compliance reports, audit framework, et cetera, often lead to the same outcomes. The question is, are you taking it seriously, or are you one of those folks who is filling out a survey because do this and you’ll be entered to win a $25 gift card somewhere? Probably Blockbuster because it no longer exists. I get those in my email constantly of, “Yeah, give half an hour of your time in return for some paltry chance to win something.”

No, I have a job to do. And I worry if at that level of that approach, who are you actually getting that’s going to sit down and fill this thing out? That said, the State of DevOps Reports have been for a long time, sort of the gold standard in this
space and I would encourage people listening to this to absolutely take the time to fill that out. cloud.google.com/devops.

I’m looking forward to seeing what comes out of it. And I love it because of the casual shade you can use to throw at other companies, too. Like, “Are you an elite team?” With the implicit baked-in sentiment being, no, you’re not, but I want to hear you say it.

Nathen: Yeah, one of the things that really sets DORA apart, also, I think, is just the—well, two of the things I guess I would say. One is the length of time that the research program has been running. It’s going on seven years now that this research program has been running, and so given that, you have tens of thousands of IT professionals that have taken the survey and provided insights into sort of what’s the state of our industry today, and where are we heading, but it’s also an academically rigorous survey. The survey and the research itself has always and continues to be completely platform and program agnostic. This is not a survey about Google Cloud.

This is not a survey where we’re trying to help understand exactly what products on Google Cloud should you use in order to be an elite performer. No. That’s not what this is about. It is about, truly, capabilities that your team needs in order to
improve their software delivery and operations performance. And I think that’s really, really important.

Dr. Nicole Forsgren who founded DORA, she didn’t come up with all of these ideas: “Hey, I think that you get better by doing this.” No. Instead, she researched all of these ideas. She got this input from across organizations of all sizes, organizations in every industry, and that, I think, really sets it apart.

And our ability to really stay committed to that academic rigor, and the platform-agnostic approach to capturing and investigating these capabilities, I think is so important to this research. And again, this is why you should participate in the survey because you truly are going to help us move the state of the art of our industry.

Corey: No, historically, there’s been a challenge where the mantle of thought leadership in conjunction with Google have intersected because there’s a common trope—historical—and I think that it is no longer accurately true. It’s an easy cheap shot, but I don’t think it holds water like it once did. Where, “Oh, Googler. It’s another word for condescending.” And there is an element of “Oh, this is how DevOps should be; this is how we’re moving things forward.” How do you distance it from
being Google says you should do it like this?

Nathen: Yeah. This comes up a lot. And frankly, I get in conversations with customers asking, “How does Google do this? How does Google do that?” And my answer always is, “You know, I can tell you how Google does something, and that might be interesting, but the fact is, it’s not much more than that, much more than interesting. Because what really matters is how are you going to do this? How are you going to improve your outcomes, whether that’s you’re delivering faster, you’re delivering more reliable, you’re running more reliable services? You’re the experts. As I mentioned earlier, you’re the experts in your teams, in your technology, and your customers. So, I’m here to learn right along with you. How are you going to do this? How are you going to improve?” Knowing how Google does it, eh, it’s interesting, but it’s not the path that you will follow.

Corey: I think that’s one of those statements that can’t ever be outright stated on a marketing website, somewhere; it’s one of those shifts that you have to live. And I think that Google’s done a pretty decent job of that. The condescending Googler jokes are dated at this point, and it’s not because there was ever an ad campaign about, we’re not condescending anymore. It was a very subtle shift in the way that Google spoke to its customers, spoke about themselves. I no longer feel the need to stand up in a blinding white rage in the Q&A portion of conference talks given by Google employees.

A lot has changed, and it’s not one thing that I can point to, it’s a bunch of different things that all add up to dramatically shifted credibility models. Realistically, I feel like that is a microcosm of a DevOps transformation. It’s not a tool; it’s not a single person being hired; it’s not, we’re taking an agile class for three days for all of engineering, and now things will be better. It’s a whole bunch of sustained work with a whole bunch of thought, and effort put into making it an actual transformation, which is such a toxically overloaded term, I dare not use it.

Nathen: Indeed. And there’s no maturity model that shows, are you there yet? And it is something that you don’t flip on or flip off like a switch, right? It doesn’t happen overnight. It takes iteration and iterative change across the entire organization.

And just like every change that you have across an organization, there are places where it’s going better than other places. And how do you learn from that? I think that’s really, really important. And to recognize and to bring some of that humility to the table is so important.

Corey: So, what’s interesting about folks that I talk to on this show—well, there are many interesting things, but one of the interesting things is, is that they have a higher rate than the general population of having at one point in their careers, written a book of some form, and you are, of course, no exception. You and Emily Freeman co-authored recently, a book entitled 97 Things every Cloud Engineer Should Know. And it’s interesting because it only has one nine in the title. Okay, that is at least an attempt at being available. I know it’s available wherever most books are sold. Tell me more.

Nathen: Yeah, so first, let’s start with the 97. Why 97? Corey, I don’t know if this or not, but 100% is the wrong reliability target for just about everything. So, 97. That feels achievable.

Corey: It also feels like three people said they would do it and then backed out at the last minute, but that’s my cynicism
speaking.

Nathen: Well, for better or worse, O’Reilly. Has a whole 97 Things series and this is part of it. So, it is, in fact, 97 things. The other thing that I think is really important about the book: you mentioned that Emily and I wrote it, and the beauty is, for a long time, I’ve wanted to have written a book, and I have never wanted to be writing a book.

Corey: That is what every author has ever said. It’s, no one wants to write a book; they want to have written one. And then you get a couple of beers into people and ask them, “So, I’m debating writing a book. Should I?” The response is, “No. Absolutely not. No.” And at some point, when you calmed them down again, and they stop screaming, they tell you the horrifying stories, and you realize, “Oh, wow, I really never want to write a book.”

Nathen: [laugh]. Yes. Well, the beauty of 97 Things and this book in particular, or the whole series, really, is its subtitle is Collective Wisdom from the Experts. So, in fact, we had over 80 different contributors sharing things that other cloud engineers should know. And I think this is also really, really important because having 80-plus contributors to this book gave us, not 97 things that Emily and Nathen think every cloud engineer should know, but instead, a wide variety of experience levels, a wide variety of perspectives, and so I think that is the thing that makes the book really powerful.

It also means that those 80-some folks that contributed to the book, had to write a very short article. So, of course, with 80 authors and 97 Things, the book is not—it doesn’t weigh 27 pounds, right? It’s less than 300 pages long, where you get these 97 tidbits. But really, the hope and the intent behind the book is to give you an idea about what should you explore deeper and, just as importantly, who are some people that you can, maybe, reach out to and talk to about a particular topic, a particular thing that a cloud engineer should know. Here are 80 people that are here, helping you and really cheering you on as you take this journey into cloud engineering.

Corey: I think there’s something to be said from having the stories for this is what we do, this is how we do it. But the lessons learned stories, those are the great ones, and it’s harder to get people on stage to talk about that without turning into, “And that’s how we snagged victory from the jaws of defeat.” No one ever gets on stage and says and that’s why the entire project was a boondoggle and four years later, we’re still struggling to recover. Especially, you know, publicly traded companies tend not to say those things. But it’s true.

You wind up with people getting on stage and talking instead about these high-level amazing things that they’ve done in the project went flawlessly, and you turn to the person next to you and say, “Yeah, I wish I could work in a place like that.” And they say, “Yeah, me too.” And you check, and they work at the same place as the presenter. Because it’s conference-ware; it’s never a real story. I’m hoping that these stories go a bit more in-depth into the nitty-gritty of what worked, what
didn’t work, and it’s not always ‘author as hero protagonist.’

Nathen: Oh, you will definitely find that in this book. These are true stories. These are stories of pain, of heartache, of victory and success, and learnings along the way. Absolutely. And frankly, in the DevOps space, we do an okay job of talking openly about our failures.

We often talk about things that we tried that went wrong, or epic failures in our systems, and then how we recovered from them. And yes, oftentimes, those stories have a great sort of storybook ending to them, but there’s a lot of truth in a lot of those stories as well because we all know that no organization is uniformly good at everything. That may be the stories that they want to share most, but, you know, there’s some truth in those stories that hopefully we can find. And certainly, in this book, you will find the good, the bad, the ugly, the learnings, and all of the lessons there.

Corey: Where can people find it if they want to buy it?

Nathen: Oh, you know, you can find it wherever you buy books. There are of course, ebooks, O’Reilly’s website, you know with the—

Corey: Wherever fine books are pirated. Yes, yes.

Nathen: That’s a good place to go for books, yeah. For sure.

Corey: And we will, of course, throw a link to the book in the [show notes 00:29:12]. Thank you so much for taking the time to speak with me. If people want to learn more about the rest of what you’re up to, how you’re thinking about it, what wise wisdom you have for the rest of us, okay can they find you, other than the book?

Nathen: Yeah, a great place to reach out to me is on Twitter. I am at @nathenharvey. But I should warn you, my father
misspelled my name. So, it’s N-A-T-H-E-N-H-A-R-V-E-Y. So, you can find me on Twitter; reach out to me there.

Corey: And we will of course include links to all of that in the [show notes 00:29:43] as well. Thank you so much for speaking to me today. I really appreciate it.

Nathen: Thank you, Corey. It’s been a pleasure.

Corey: Nathen Harvey, cloud developer advocate at Google. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with an angry comment telling me that I’m completely wrong. You can instantly get DevOps in your environment if I only purchase whatever crap it is your company sells.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Sanjay

Sanjay Poonen is the former COO of VMware, where he was responsible for worldwide sales, services, support, marketing and alliances. He was also responsible for the Security strategy and business at VMware.

Prior to SAP, Poonen held executive roles at SAP, Symantec, VERITAS and Informatica, and he began his career as a software engineer at Microsoft, followed by Apple.

Poonen holds two patents as well as an MBA from Harvard Business School, where he graduated a Baker Scholar; a master's degree in management science and engineering from Stanford University; and a bachelor's degree in computer science, math and engineering from Dartmouth College, where he graduated summa cum laude and Phi Beta Kappa.

Links:

  • VMware: https://www.vmware.com/
  • leadership values: https://www.youtube.com/watch?v=lxkysDMBM0Q
  • Twitter: https://twitter.com/spoonen
  • LinkedIn: https://www.linkedin.com/in/sanjaypoonen/
  • spoonen@vmware.com: mailto:spoonen@vmware.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Thinkst. This is going to take a minute to explain, so bear with me. I linked against an early version of their tool, canarytokens.org in the very early days of my newsletter, and what it does is relatively simple and straightforward. It winds up embedding credentials, files, that sort of thing in various parts of your environment, wherever you want to; it gives you fake AWS API credentials, for example. And the only thing that these things do is alert you whenever someone attempts to use those things. It’s an awesome approach. I’ve used something similar for years. Check them out. But wait, there’s more. They also have an enterprise option that you should be very much aware of canary.tools. You can take a look at this, but what it does is it provides an enterprise approach to drive these things throughout your entire environment. You can get a physical device that hangs out on your network and impersonates whatever you want to. When it gets Nmap scanned, or someone attempts to log into it, or access files on it, you get instant alerts. It’s awesome. If you don’t do something like this, you’re likely to find out that you’ve gotten breached, the hard way. Take a look at this. It’s one of those few things that I look at and say, “Wow, that is an amazing idea. I love it.” That’s canarytokens.org and canary.tools. The first one is free. The second one is enterprise-y. Take a look. I’m a big fan of this. More from them in the coming weeks.

Corey: Let’s be honest—the past year has been a nightmare for cloud financial management. The pandemic forced us to move workloads to the cloud sooner than anticipated, and we all know what that means—surprises on the cloud bill and headaches for anyone trying to figure out what caused them. The CloudLIVE 2021 virtual conference is your chance to connect with FinOps and cloud financial management practitioners and get a behind-the-scenes look into proven strategies that have helped organizations like yours adapt to the realities of the past year. Hosted by CloudHealth by VMware on May 20th, the CloudLIVE 2021 conference will be 100% virtual and 100% free to attend, so you have no excuses for missing out on this opportunity to connect with the cloud management community. Visit cloudlive.com/coreyto learn more and save your virtual seat today. That’s cloud-l-i-v-e.com/corey to register.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I talk a lot about cloud in a variety of different contexts; this show is about the business of cloud. But, fundamentally, where cloud comes from was this novel concept, once upon a time, of virtualization. And that gave rise to a whole bunch of other things that later became, then containers, now it becomes Kubernetes, and if you want to go down the serverless path, you can.

But it’s hard to think of a company that has had more impact on virtualization and that narrative than VMware. My guest today is Sanjay Poonen, Chief Operating Officer of VMware. Thank you for joining me.

Sanjay: Thanks, Corey Quinn, it’s great to be with you and with your audience on this show.

Corey: So, let’s start with the fun slash difficult questions. It’s easy to look at VMware as a way of virtualizing existing bare-metal workloads and moving those VMs around, but in many respects, that is perceived by some—ehem, ehem—to be something of a legacy model of cloud interaction where it solves the problem of on-premises, which is I’m really bad at running data centers so I’m just going to treat the cloud like a data center. And for some companies and some workloads, where, great, that’s fine. But isn’t that, I guess, a V1 vision of cloud, and if it is, why is VMware relevant to that?

Sanjay: Great question, Corey. And I think it’s great to be straight up on a topic [unintelligible 00:02:01]. Yeah, I think you’re right. Listen, the ‘V’ in VMware is virtualization. The ‘VM’ is virtual machines.

A lot of what is the underpinning of what made the private cloud, as we call it today, but the data center of the past successful was this virtualization technology. In the old days, people would send us electricity bills, before and after VMware, and how much they’re saving. So, this energy-saving concept of virtualization has been profound in the modernization of the data center and the advent of what’s called the private cloud. But as you looked at the public cloud innovate, whether it was AWS or even the SaaS applications—I mean, listen, the most popular capability initially on AWS was EC2 and S3, and the core of EC2 is virtualization. I think what we had to do, as this happened, was the foundation was certainly those services like EC2 and S3, but very quickly, the building phenomenon that attracted hundreds of thousands and I think now probably a few million customers to AWS was the large number of services, probably now 150, 200-odd services, that were built on top of that for everything from data, to AI, to a variety of other things that every year Andy Jassy and the team would build up.

So, we had to make sure that over the course of the last, I’d say, certainly the last five to maybe eight years, we were becoming relevant to our customers that were a mix. There were customers who were large—I mean, we have about half a million customers—and in many cases, they have about 80, 90% of their workloads running on-prem and they want to move those workloads to the cloud, but they can’t just refactor and re-platform all of those apps that are running in the on-premise world. When they will try to do it by the end of the year—they may have 1000 applications—they got 10 done.

Corey: Oh, and it’s not realistic and it’s unfair. I mean, there’s the idea of, “Oh, that’s legacy,” which is condescending engineering speak for it actually makes money because it’s been around for longer than six months. And sure you can have Twitter For Pets roll stuff out every day that you want; when you’re a bank, you have different constraints forced upon you. And I’m very sympathetic to folks who are in scenarios where they aren’t, for whatever reason, able to technically, culturally, or for regulatory reasons, be able to do continuous deployment of everything. I want to be very clear that I’ve in no way passing judgment on an entire sector of enterprise.

Sanjay: But while that sector is important, there was also another sector starting to emerge: the Airbnbs, the Pinterests, the modern companies who may not need VMware at all as they’re building native, but may need some of our container in a new open-source capabilities. SaltStack was one of them; we will talk about that, I’m sure. So, we needed to be relevant to both customer communities because the Airbnbs of today, will be the Marriotts of tomorrow. So, we had to really rethink what is the future of VMware, what’s our existence in a public cloud phenomenon? That’s really what led to a complete watershed moment.

I called publicly in the past sort of a Berlin Wall moment where Amazon and VMware were positioned pretty much as competitors for a long period of time when AWS was first started. Not that Andy was going around talking negatively about VMware, but I think people view these as two separate doors, and never the twain would meet. But when we decided to partner with them—I then quite frankly, the precursor to that was us divesting our public cloud strategy. We’d tried to build a competitive public cloud called vCloud Air between the period of 2012 and 2015, 2016—we had to reach an end of that movement, and catharsis of that, divest that asset, and it opened the door for a strategic partnership. But now we can go back to those customers and help them move their applications in a way that’s highly efficient, almost like a house on wheels, and then once it’s in that location in AWS—or one of the other public clouds—you can modernize it, too.

So, then you get to both get the best of both worlds: get it into the public cloud, maybe retire some of your data centers if that’s what you want to do, and then modernize it with all the beautiful services. And that’s the best of both worlds. Now, if you have 1000 applications, you’re moving hundreds of them into the public cloud, and then using all of the powerful developer services on that VMware stack that’s built on the bare metal of AWS. So, we started out with AWS, but very quickly then, all the other public clouds, maybe the five or six that are named in the Gartner Magic Quadrant, came to us and said, “Well, if you’re doing that with AWS, would you consider doing that with us, too?”

Corey: There’s definitely been an evolution of VMware. I mean, it’s in the name; you have the term VM sitting there. It’s easy to, at least from where I sit, think of, “Oh, VMware, back when running virtual machines was novel.” And there was a lot of skepticism around the idea. I’m going to level with you; I was a skeptic around virtualization. Then around cloud. Then around containers.

And now I’m trying—all right I’m going to be in favor of serverless, which is almost certain to doom it because everything else that I’ve been skeptical of in this sense beyond any reasonable measure. So, there is this idea that VMs are this sort of old-school thinking. And that’s great if you have an existing workload that needs to be migrated, but there are a finite number of those in the world. As we turn towards net-new and greenfield build-outs, a lot of things are a lot more cloud-native than just hosting a bunch of—if you take the AWS example—EC2 instances hanging out in the network talking to other EC2 instances. Taking advantage of native offerings definitely seems to be on the rise. And there have been acquisitions that VMware has made. You talk about SaltStack, which was a great example, given that I wrote part of that very early on, and I don’t think the internet’s ever forgiven me for it. But also Bitnami—or BittenAMI, as I insist on pronouncing it—and you also acquired Wavefront. There’s a lot of interesting stuff that feels almost like a setting up a dichotomy of new VMware versus old VMware. What are the points of commonality there? What is the vision for the next 15 years of the company?

Sanjay: Yeah, I think when we think about it, it’s very important that, first off, we acknowledge that our roots are what gives us sustenance because we have a large customer base that uses us. We have 80 million workloads running on that VMware infrastructure, formerly ESX, now vSphere. And that’s our heritage, and those customers are happy. In fact, they’re not, like, fleeing like birds into there, so we want to care for those customers.

But we have to have a north star, like a magnet that pulls us into the modern world. And that’s been—you know, I talked about phase one was this really charting of the future of VMware for the cloud. Just as important has been focused on cloud-native and containers the last three, four years. So, we acquired Heptio. As you know, Heptio was founded by some of the inventors of Kubernetes who left Google, Joe Beda, and Craig McLuckie.

And with that came a strong I would say relevancy, and trust to the Kubernetes, we’ve become one of the leading contributors to open-source Kubernetes. And that brain trust now, some of whom are at VMWare and many are in the community think of us very differently. And then we’ve supplemented that with many other moves that are much more cloud-native. You mentioned two or three of them: Bitnami, for that sort of marketplace; and then SaltStack for what we have been able to do in configuration management and infrastructure automation; Wavefront for container-based workloads. And we’re not done, and we think, listen, there will be many, many more things that the first 10, 15 years of VMware was very much about optimizing the private cloud, the next 10, 15 years could be optimizing for that app modernization cloud-native world.

And we think that customers will want something that can work in a multi-cloud fashion. Now, multi-cloud for us is certainly private cloud and edge cloud, which may have very little to do with hardware that’s in the public cloud, but also AWS, Azure, and two or three other clouds. And if you think of each of these public clouds as mini skyscrapers—so AWS has 50 billion in revenue; I’m going to guess Azure is, like, 30, and then Google is I don’t know 12, 13; and then everyone else, and they’re all skyscrapers are different—it’s like, if we can be that company that fills the crevices between them with cement that’s valuable so that people can then build their houses on top of that, you’re probably not going to be best served with a container Stack that’s trapped to just one cloud. And then over time, you don’t have reasonable amount of flexibility if you choose to change that direction. Now, some people might say, “Listen, multi-cloud is—who cares about that?”

But I think increasingly, we’re hearing from customers a desire to have more than just one cloud for a variety of reasons. They want to have options, portability, flexibility, negotiating price, in addition to their private cloud. So, it’s a two plus one, sometimes it might be a two plus two, meaning it’s a private cloud and the edge cloud. And I think VMware is a tremendous proposition to be that Switzerland-type company that’s relevant in a private cloud, one or two public clouds, and an edge cloud environment, Corey.

Corey: Are you seeing folks having individual workloads that they want to flow from one cloud to another in a seamless way, or is it more aligned along an approach of having workload A lives in this cloud and workload B lives in this cloud? And you’re in a terrific position to opine on that more than most, given who you are.

Sanjay: Yeah. We’re not yet as yet seeing these floating workloads that start here and move around, that’s—usually you build an application with purpose. Like, it sits here in this cloud and of course. But we’re seeing, increasingly, interest at customers’ not tethering it to proprietary services only. I mean, certainly, if you’re going to optimize it for AWS, you’re going to take advantage of EC2, S3, and then many of the, kind of, very capable [unintelligible 00:11:24], Aurora, there are others that might be there.

But over time, especially the open-source movement that brings out open-source data services, open-source tooling, containers, all of that stuff, give ultimately customers the hope that certainly they should add economic value and developer productivity value, but they should also create some potential portability so that if in the future you wanted to make a change, you’re not bound to that cloud platform. And a particular cloud may not like us saying this, but that’s just the fact of how CIOs today are starting to think much more so as they build these up and as many of the other public clouds start to climb in functionality. Now, there are other use cases where particular SaaS applications of SaaS services are optimized for a particular [unintelligible 00:12:07], for example, Office 365, someone’s using a collaboration app, typically, there’s choices of one or two, you’re either using a G Suite and then it’s tied to Google, or it’s Office 365. But even there, we’re starting to see some nibbling around the edges. Just the phenomenon of Zoom; that wasn’t a capability that Microsoft brought very—and the services from Google, or Amazon, or Microsoft was just not as good as Zoom.

And Zoom just took off and has become the leading video collaboration platform because they’re just simple, easy to use, and delightful. It doesn’t matter what infrastructure they run on, whether it’s AWS, I mean, now they’re running some of their workloads on Oracle. Who cares? It’s a SaaS service. So, I think increasingly, I think there will be a propensity towards SaaS applications over custom building. If I can buy it why would I want to build a video collaboration app myself internally, if I can buy it as a SaaS service from Zoom, or whoever have you?

Corey: Oh, building it yourself would be ludicrous unless that was one of your core competencies.

Sanjay: Exactly.

Corey: And Zoom seems to have that on lock.

Sanjay: Right. And so similarly, to the extent that I think IT folks can buy applications that are more SaaS than custom-built, or even on-prem, I mean, Salesforce—the success of Salesforce, and Workday, and Adobe, and then, of course, the smaller ones like Zoom, and Slack, and so on. So, it’s clear evidence that the world is going to move towards SaaS applications. But where you have to custom build an application because it’s very unique to your business or to something you need to very snap quickly together, I think there’s going to be increasingly a propensity towards using open-source types of tooling, or open-source platforms—Kubernetes being the best example of that—that then have some multi-cloud characteristics.

Corey: In a similar note, I know that the term is apparently, at least this week on Twitter, being argued against, but what about cloud repatriation? A lot of noise has been made about people moving workloads from public cloud back to private cloud. And the example they always give is Dropbox moving its centralized storage service into an on-prem environment, and the second example is basically a pile of tumbleweeds because people don’t really have anything concrete to point at. Does that align with your experience? Is there a, I guess, a hidden wave of people doing a reverse cloud migration that just doesn’t get discussed?

Sanjay: I think there’s a couple of phenomenons, Corey, that we watch here. Now, clearly a company of the scale of Dropbox has economics on data and storage, and I’ve talked to Drew and a variety of the folks there, as well as Box, on how they think about this because at that scale, they probably could get some advantages that I’m sure they’ve thought through in both the engineering and the cost. I mean, there’s both engineering optimization and costs that I’m sure Drew and the folks there are thinking through. But there’s a couple of phenomena that we do—I mean, if you go back to, I think, maybe three or four quarters ago, Brian Moynihan, the CEO of Bank of America, I think in 2019, mid to late 2019 made a statement in his earnings call, he was asked, “How do you think about cloud?” And he said, “Listen, I can run a private cloud cheaper and better than any of the public clouds, and I save 240%,” if I remember the data right.

Now, his private cloud and Bank of America is a key customer [unintelligible 00:15:04] of us, we find that some of the bigger companies at scale are able to either get hardware at really good pricing, are able to engineer—because they have hundreds of thousands—they’re almost mini VMware, right, [unintelligible 00:15:18] themselves because they’ve got so many engineers. They can do certain things that a company that doesn’t want to hire those many—companies, Pinterest, Airbnb may not do. So, there are customers who are going to basically say, even prior to repatriation, that the best opportunity is a private cloud. And in that place, we have to work with our private cloud partners, whether it’s Dell or others, to make sure that stack of hardware from them plus the software VMware in the containers on top of that is as competitive and is best cost of ownership, best ROI. Now, when you get to your second—your question around repatriation, what we have found in certain regions outside the US because of sovereign data, sovereign clouds, sometimes some distrust of some of those countries of the US public cloud, are they worried about them getting too big, fear by monopoly, all those types of things, lead certain countries outside the US to think about something that they would need that’s sovereign to their country.

And the idea of sovereign data and sovereign clouds does lead those to then investing in local cloud providers. I mean, for example in France, there is a provider called OVH that’s kind of trying to do some of that. In China, there’s a whole bunch of them, obviously, Alibaba being the biggest. And I think that’s going to continue to be a phenomenon where there’s a [federated said 00:16:32], we have a cloud provider program with this 4000 cloud providers, Corey, who built their stack on VMware; we’ve got to feed them. Now, while they are an individual revenue way smaller than the public clouds were, but collectively, they represent a significant mass of where those countries want to run in a local cloud provider.

And from our perspective, we spent years and years enabling that group to be successful. We don’t see any decline. In fact, that business for us has been growing. I would have thought that business would just completely decline with the hyperscalers. If anything, they’ve grown.

So, there’s a little bit of the rising tide is helping all boats rise, so to speak. And the hyperscaler’s growth has also relied on many of these, sort of, sovereign clouds. So, there’s repatriation happening; I think those sovereign clouds will benefit some, and it could also be in some cases where customers will invest appropriately in private cloud. But I don’t see that—I think if anything, it’s going to be the public cloud growing, the private cloud, and edge cloud growing. And then some of these, sort of, country-specific sovereign clouds also growing. I don’t see this being in a huge threat to the public cloud phenomena that we’re in.

Corey: This episode is sponsored in part by our friends at Lumigo. If you've built anything from serverless, you know that if there's one thing that can be said universally about these applications, it's that it turns every outage into a murder mystery. Lumigo helps make sense of all of the various functions that wind up tying together to build applications. It offers one-click distributed tracing so you can effortlessly find and fix issues in your serverless and microservices environment. You've created more problems for yourself. Make one of them go away. To learn more, visit lumigo.io.

Corey: I want to very clear, I think that there’s a common misconception that there’s this, somehow, ongoing fight between all the cloud providers, and all this cloud growth, and all this revenue is coming at the expense of other cloud providers. I think that it is simultaneously workloads that are being migrated from on-premises environments—yes—but a lot of it also feels like it’s net-new. It’s not just about increasingly capturing ever larger portions of the market but rather about the market itself expanding geometrically. For a long time, it felt like that was what tech was doing. Looking at the global IT spend numbers coming out of Gartner and other places, it seems like it’s certainly not slowing down. Does that align with your perception of it? Or are there clear winners and losers that are I guess, differentiating out?

Sanjay: I think, Corey, you’re right. I think if you just use some of the data, the entire IT market, let’s just say it’s about $1 trillion, some estimates have it higher than that. Let’s break it down a little bit. Inside that 1 trillion market it is growing—I mean, obviously COVID, and GDP declined last year in calendar 2020 did affect overall IT, but I think let’s assume that we have some kind of U-shape or other kind of recovery, going into the second half of certainly into next year; technology should lead GDP in terms of its incline. But inside that trillion-dollar market, if you add up the SaaS market, it’s about $115 billion market.

And these are companies like Salesforce, and Adobe, and Workday, and ServiceNow. You add them all up, and those are growing, I think the numbers were in the order of 15 or 20% in aggregate. But that SaaS market is [unintelligible 00:19:08]. And that’s growing, certainly faster than the on-prem applications market, just evidenced by the growth of those companies relative to on-premise investments in SAP or Oracle. And then if you look at the infrastructure market, it’s slightly bigger, it’s about $125 billion, growing slightly faster—20, 25%—and there you have the companies like AWS, Azure, and Google, and Alibaba, and whoever have you. And certainly, that growth is faster than some of the on-premise growth, but it’s not like the on-premise folks are declining. They’re growing at slower paces.

Corey: It is harder to leave an on-premise environment running and rack up charges and blow out the bill that way, but it—not impossible, I suppose, but it’s harder to do than it is in public cloud. But I definitely agree that the growth rate surpasses what you would see if it were just people turning things on and forgetting to turn them off all the time.

Sanjay: Yeah, and I think that phenomenon is a shift in spending where certainly last year we saw more spending in the cloud than on-premise. I think the on-premise vendors have a tremendous opportunity in front of them, which is to optimize every last dollar that is going to be spent in the data centers, private cloud. And between us and our partners like Dell and others, we’ve got to make sure we do that for our customer base that we’ve accumulated over last 10, 15 years. But there’s also a significant investment now moving to the edge. When I look at retailers, CPG companies—consumer packaged good companies—manufacturers, the conversation that I’m having with their C-level tech or business executives is all about putting compute in the stores.

I mean, listen, what is the retailer concerned about? Fraud, and some of those other things, and empowering a quick self-service experience for a consumer who comes in and wants to check out of a Safeway or Walmart really quickly. These are just simple applications with local compute in the store, and the more that we can make that possible on top of almost like a nano data center or micro data center, running in the store with those applications resident there, talking—you know, you can’t just take all of that data, go back and forth to the cloud, but with resident services and capability right there, that’s a beautiful opportunity for the VMware and the Dells of the world. And that’s going to be a significant place where I think you’re going to see expansion of their focus. The Edge market today is I think, projected to be about $6 or $8 billion this year, and growing to $25 billion the next four or five years.

So, much smaller than the previous numbers I shared—you know, $125, $115 billion for SaaS and IaaS—but I think the opportunity there, especially these industries that are federated: CPG, consumer packaged goods, manufacturing, retail, and logistics, too—you know, FedEx made a big announcement with VMware and Dell a few months ago about how they’re thinking about putting compute and local infrastructure at their distribution sites. I think this phenomenon, Corey, is going to happen in a number of different [unintelligible 00:21:48], and is a tremendous opportunity. Certainly, the public cloud vendors are trying to do that with Outposts and Azure Stack, but I think it does favor the on-premise vendors also having a very strong proposition for the edge cloud.

Corey: I assumed that the whole discussion with FedEx started by someone dramatically misunderstanding what it meant to ship code to production.

Sanjay: [laugh]. I mean, listen, at the end of the day, all of these folks who are in traditional industries are trying to hire world-class developers—like software companies—because all of them are becoming software companies. And I think the open-source movement, and all of these ways in which you have a software supply chain that’s more modernized, it’s affecting every company. So, I think if you went into the engineering product teams of Rob Carter, who runs technology for FedEx, you’ll find them and they may not have all of the sophistication as a world-class software company, but they’re getting increasingly very much digital in their focus of next generation. And same thing with UPS.

I was talking to the CEO of UPS, we had her come and speak at our kickoff. It’s amazing how much her lingo—she was the former CFO of Home Depot—I felt like I was talking to a software executive, and this is the CEO of UPS, a logistics company. So, I think increasingly, every company is becoming a software company at their core. And you don’t need to necessarily know all the details of containers and virtualization, but you need to understand how software and digital transformation, how technology can power your digital transformation.

Corey: One thing that I’ve noticed the more I get to talk to people doing different things in different roles was, at first I was excited because I get to talk to the people where they’re really doing it right and everything’s awesome. And I’ve increasingly of the opinion that those sites don’t actually exist. Everyone talks about the great thing is that they’re doing and aspirationally in certain areas in the terms of conference-ware, but you get down into the weeds, and everyone views their environment as being a burning tire fire of sadness and regret. Everyone thinks other people are doing it way better than they are. And in some cases they’re embarrassed about it, in some cases they’re open about it, but I feel like we’re still in the early days where no one is doing things in the quote-unquote, “Right ways,” but everyone thinks everyone else is.

Sanjay: Yeah, I think, Corey, that’s absolutely right. We are very much early days in all of this phenomenon. I mean, listen, even the public cloud, Andy himself would say it’s [laugh]—he wouldn’t say it’s quite day one, but he would say it’s very early [unintelligible 00:24:03], even though they’ve had 15 years of incredible success and a $50 billion business. I would agree. And when you look at the customers and their persona—when I ask a CIO what percentage of—of an established company, not one of the modern ones who are built all cloud-native—but what percentage of your workloads are in a public cloud versus private cloud, the vast majority is still in a data center or private cloud.

But with the intent—if it’s 90/10, let’s say 90 private 10—for that to become 70/30, 50/50. But very rarely do I hear a one of these large companies say it’s going to be 10/90 the opposite way in three, five years. Now, listen, I think every company as it grows that is more modern. I mean the Zooms of the world, the Modernas, the Airbnbs, as they get bigger and bigger, they represent a completely new phenomenon of how they are building applications that are all cloud-native. And the beautiful thing for me is just as a former engineering and developer, I mean, I grew up writing code in C, and C++ and then came BEA WebLogic, and IBM WebSphere, and [JGUI 00:25:04].

And I was so excited for these frameworks. I’m not writing code, thankfully, anymore because it would create lots of problems if I did. But when I watched the phenomena, I think to myself, “Man, if I was a 22 year old entering the workforce now, it’s one of the most exciting times to write code and be a developer because what’s available to you, both in the combination of these cloud frameworks and open-source frameworks, is immense.” To be able to innovate much, much faster than we did 25, 30 years ago when I was a developer.

Corey: It’s amazing there’s the pace of innovation, if cloud has changed nothing else, from my perspective, it’s been the idea that you can provision things without these hefty waiting periods. But I want to shift gears slightly because we’ve been talking about cloud for a bit in the context of infrastructure, and containers, and the rest, but if we start moving up the stack a little bit, that’s also considered cloud, which just seems to have that naming problem of namespace collision, just to confuse folks. But VMware is also active in this space, too. You’ve got things like Workspace ONE, you’ve got a bunch of other endpoint options as well that are focused on the security space. Is that aligned?

Is that just sort of a different business unit? How does that, I guess, resonate between the various things that you folks do? Because it turns out, you’re kind of a big company, and it’s difficult to keep it all straight from an external perspective.

Sanjay: Well, I think—listen, we’re roughly a little less than $12 billion in revenue last year. You can think of us in two buckets: everything in the first bucket is all that we talked about. Think of that as modernization of applications and cloud infrastructure, or what people might think about PaaS and IaaS without the underlying hardware; we’re not trying to build servers and storage and networking at the hardware level, you know, and so and so. But the software layer is about, that’s the first conversation we had for the last 15, 20 minutes. The second part of our business is where we’re touching end-users and infrastructure, and securing it.

And we think that’s an important part because that also is something through software, and the cloud could be optimized. And we’ve had a long-standing digital workspace. In fact, when I came to VMware, it was the first business I was running in terms of all the products and end-user computing. And our thesis was many of the current tools, whether it’s the virtual desktop technology that people have from existing vendors, or even today, the security tools that they use is just too cumbersome. It’s too heavy.

In many cases, people complain about the number of agents they have on their laptops, or the way in which they secure firewalls is too expensive and too many. We felt we could radically—VMware gets involved in problems where we can radically simplify thing with some disruptive innovation. And the idea was, first in the digital workspace was to radically reduce cost with software that was built for the cloud. And Workspace ONE and all of those things radically reduce the need for disparate technologies for virtual desktops, identity management, and endpoint management. We’ve done very well in that.

We’re a leader in that segment, if you look at any of the analysts ratings, whether it’s Gardner or others. But security has been a more recent phenomenon where we felt like it leads us very quickly into securing those laptops because on those same laptops, you have antivirus, you have a variety of tools, and on the average, the CSOs, the Chief Security Officers tell me they have way too many agents, way too many consoles, way too many alerts, and if we could reduce that and have a single agent on a laptop, or maybe even agentless technology that secure this, that’s the Nirvana. And if you look at some of the recent things that have happened with SolarWinds, or Petya, WannaCry in the past, security’s of top concern, Corey, to boards. And the more that we could do to clean that up, I think we can emerge—which we’re already starting to—as a cybersecurity layer. So, that’s a smaller part of our business, but, I mean, it’s multi-billion now, and we think it’s a tremendous opportunity for us to take what we’re doing in workspace and security and make that a growth vector.

So, I think both of these core areas, the cloud infrastructure, and modern applications—topic number one—workspace and security—topic number two—I’m both tremendous opportunities for VMware in our journey to grow from a $12 billion company to one day, hopefully, a $20 billion company.

Corey: Would that we all had such problems, on some level. It’s really interesting seeing the evolution of companies going from relatively small companies and humble beginnings to these giant—I guess, I want to use the term Colossus, but I’m not sure if that’s insulting or [laugh] not—it’s phenomenal just to see the different areas of business that VMware has expanded into. I mean, I’ve had other folks from your org talking about what a Tanzu is or might be, so we aren’t even going to go down that rabbit hole due to time constraints at this point, but one thing that I do want to get into, slightly, has been a recurring theme in the show, which is where does the next generation of leaders come from? Where do the next generation engineers come from? And you’ve been devoting a bit of time to this. I think I saw one of your YouTube videos somewhat recently about your leadership values. Talk to me a little bit about that.

Sanjay: Yeah. Corey, listen, I’m glad that we’re closing out this on some of the soft topics because I love talking to you, or other talented analysts and thought leaders around technology. It’s my roots; I’m a technical person at heart. I love technology. But I think the soft stuff is often the hard stuff.

And the hard stuff is often the soft stuff. And what I mean by that is, when all this peels away, what your lasting legacy to the company are the people you invest in, the character you build. And, I mean, as an immigrant who came to this country, when I was 18 years old, $50 in my pocket, I was very fortunate to have a scholarship to go to a really nice University, Dartmouth College, to study computer science. I mean, I grew up in India and if it wasn’t for the opportunity to come here on a scholarship, I wouldn’t have [been here 00:30:32]. So, everything I consider a blessing and a learning opportunity where I’m looking at the advent of life as a growth mindset: what can I learn? And we all need to cultivate more and more aspects of that growth mindset where we move from being know-it-alls to learn-it-alls.

And one of the key things that I talk about—and all of your listeners on this, listening to this, I welcome to go to YouTube and search Sanjay Poonen and leadership, it’s a 10-minute video—I’ll pick one of them. Most often as we get higher and higher in an organization, leaders tend to view things as a pyramid, and they’re kind of like this chief bird sitting at the top of the pyramid, and all these birds that are looking—below them on branches are looking up and all they see is crap falling down. Literally. That’s what happens when you look at the bird up. And our job as leaders is to invert that pyramid.

And to actually think about the person who is on the front lines. In a software company, it’s an engineer and a sales rep. They are the folks on the frontline: they’re writing code or selling code. They are the true people who are making things happen. And when we as leaders look at ourselves as the bottom of the pyramid—some people call that, “Servant leadership.”

Whatever way you call it, the phrase isn’t the point—the point is, invert that pyramid and to take obstacles out of people from the frontline. You really become not interested as much around what your own personal wellbeing, it’s about ensuring that those people in the middle layers and certainly at the leaf levels of the organization are enormously successful. Their success becomes your joy, and it becomes almost like a parent, right? I mean, Corey, you have kids; I’ve got kids. Imagine if you were a parent and you were jealous of your kid’s success.

I mean, I want my three children, my daughter, my two children to do better than me, running races or whatever it is that they do. And I think as a leader, the more that we celebrate the successes of our teams and people, and our lasting legacy is not our own success; it’s what we have left behind, other people. I’ve say often there’s no success without successors. So, that mindset takes a lot of work because the natural tendency of the human mind and the human behavior is to be selfish and think about ourselves. But yeah, it’s a natural phenomenon.

We’re born that way, we live in act that way, but the more that we start to create that, then taking that not just to our team, but also to the community allows us to build a better society. And that’s something I’m deeply passionate about, try to do my small piece for it, and in fact, I’m sometimes more excited about these topics of leadership than even technology.

Corey: It feels like it’s the stuff that lasts; it has staying power. I could record a video now about technology choices and how to work with those technologies and unless it’s about Git, it’s probably not going to be too relevant in 10 years. But leadership is one of those eternal things where it’s, once you’ve experienced a certain level of success, you can really see what people do with that the people that I like to surround myself with, generally make it a point to send the elevator back down, so to speak.

Sanjay: I agree, Corey, it’s—glad that you do it. I’m always looking for people that I can learn from, and it doesn’t matter where they are in society. I mean, I think you often—I mean, this is classic Dale Carnegie; one of the books that my dad gave to me at a young age that I encourage everyone to read, How to Win Friends and Influence People, talked about how you can detect a person’s character based on the way they treat the receptionist, or their assistants, the people who might be lower down the totem pole from them. And most often you have people who kiss up and kick down. And I think when you build an organization that’s that typical.

A lot of companies are built that way where they kiss up and kick down, you actually have an inverted sense of values. And I think you have to go back to some of those old-school ways that Dale Carnegie or Steven Covey talked about because you don’t have to build a culture that’s obnoxious; you can build a company that’s both nice and competitive. It doesn’t mean that anything we’ve talked about for the last few minutes means that I’m any less competitive and I don’t want to beat the competition and win a deal. What you can do it nicely. And even that’s something that I’ve had to grow in.

So, I think when we all look at ourselves as sculptures, work in progress, and we’re perfecting our craft, so to speak, both on the technical front, and the product front and customer relationship, but then also on the leadership and the personal growth front, we actually become both better people and then we also build better companies.

Corey: And sometimes that’s really all that we can ask for. If people want to learn more about what you have to say and get your opinion on these things, okay can they find you?

Sanjay: Listen, I’m very approachable. You can follow me on Twitter, I’m on LinkedIn [unintelligible 00:34:54], or my email spoonen@vmware.com. I’m out there.

I read voraciously, and probably not as responsive, sometimes, but I try—certainly, customers will hear from me within 24 hours because I try to be very responsive to our customers. But you can connect with me on social media. And I’m honored to be on your show, Corey. I’ve been reading your stuff since it first came out, and then, obviously, a fan of the way you’re thinking about things. Sometimes I feel I need to correct your opinion, and some of that we did today. [laugh]. But you’ve been very—

Corey: Oh, I would agree. I come out of this conversation with a different view of VMware than I went into it with. I’m being fully transparent on that.

Sanjay: And you’ve helped us. I mean, quite frankly, your blogs and your focus on this and, like, is the V in VMware, like, a bad word? Is it legacy? It’s forced us to think, so I think it’s iron sharpens iron. I’m very delighted that we connected, I don’t know if it was a year or two years ago.

And I’ve been a fan; I watch the stuff that you do at re:Invent, so keep going with what you’re doing. I think all of what you write and what you talk about is hopefully making an impact on people who read and listen. And look forward to continuing this dialogue, not just with me, but I think you’re talking to other people in VMware in the future. I’m not the smartest person at VMware, but I’m very fortunate to be [laugh] surrounded by many of them. So hopefully, you get to talk to them, also, in the near future.

Corey: [laugh]. I will, of course, will put links to all that in the [show notes 00:36:11]. Thank you so much for taking the time to speak with me today. I really appreciate it.

Sanjay: Thanks, Corey, and all the best of you and your organization.

Corey: Sanjay Poonen, Chief Operating Officer of VMware, I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with a condescending comment telling me that in fact, it is a best practice to ship your code to production via FedEx.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

This has been a HumblePod production. Stay humble.

View Details

Transcript

Corey: This episode is sponsored in part byLaunchDarkly. Take a look at what it takes to get your code into production. I’m going to just guess that it’s awful because it’s always awful. No one loves their deployment process. What if launching new features didn’t require you to do a full-on code and possibly infrastructure deploy? What if you could test on a small subset of users and then roll it back immediately if results aren’t what you expect? LaunchDarkly does exactly this. To learn more, visitlaunchdarkly.com and tell them Corey sent you, and watch for the wince.

Jesse: Welcome to AWS Morning Brief: Fridays From the Field. I’m Jesse DeRose.

Amy: I’m Amy Negrette.

Jesse: This is the podcast within a podcast where we talk about all the ways we’ve seen AWS used and abused in the wild, with a healthy dose of complaining about AWS for good measure because I mean, who doesn’t love to complain about AWS? I feel like that’s always a good thing that we can talk about, no matter the topic. Today, we’re going to be talking about the ‘cloud cost management starter kit.’ So, the starter kit seems to be a big fad that’s going around. If you’re listening to this episode, you’re probably thinking, “It’s already done. It’s over.”

But I still want to talk about it. I think that this is a really relevant topic because I think a lot of companies are trying to get started, get their hands started in cloud cost management. So, I think this would be a great thing for us to talk about: what’s in our cloud cost management starter kit?

Amy: And it really will help answer that question that I get asked a lot on: what is even a cloud economist, and what do you do?

Jesse: Yeah, I mean, given the current timeframe, I haven’t gone to any parties recently to talk about what I do, but I do feel like anytime I try to explain to somebody what I do, there’s always that moment of, “Okay. Yes, I work with computers, and we’ll just leave it at that.”

Amy: It’s easier to just think about it as we look at receipts, and we kind of figure things out. But when you try to get into the nuts and bolts of it, it’s a very esoteric idea that we’re trying to explain. And no, I don’t know why this is a real job. And yet it is.

Jesse: This is one of the things that always fascinates me. I absolutely love the work that I do, and I definitely think that it is important work that needs to be done for any organization, to work on their cloud cost management best practices, but it also boggles my mind that AWS, Azure, GCP, haven’t figured out how to bake this in more clearly and easily to all of their workflows and all their services. It still boggles my mind that this is something that exists as—

Amy: As a thing we have to do.

Jesse: As a thing we have to do. Yeah, absolutely.

Amy: Well, the good news is, they’re going to change their practices once every six weeks, and we’ll have a new thing to figure it out. [laugh].

Jesse: [laugh]. So, let’s get started with the first item on our cloud cost management starter kit. This one is something that Amy is definitely passionate about; I am definitely passionate about, as well. Amy, what is it?

Amy: Turn on your CUR. Turn on your CUr. If you don’t know what it is, just Google AWS CUR. Turn it on. It will save you a headache, and it will save anyone you bring in to help you [laugh] [unintelligible 00:02:59] a huge headache. And it keeps us from having to yell at people, even though that’s the thing that if you pay us to do it, we will totally do it for you.

Jesse: If you take nothing away from this episode, go check out the AWS Cost and Usage Report—otherwise known as CUR—turn it on for your accounts, ideally enable it in Parquet format because that’s going to allow you to get all that sweet, sweet data in an optimized manner, living in your S3 bucket. It is a godsend. It gives you all the data from Cost Explorer, and then some. It allows you to do all sorts of really interesting business intelligence analytics on your billing data. It’s absolutely fantastic.

Amy: It’s like getting all of those juicy infrastructure metrics, except getting that with a dollar sign attached to it so you know what you actually doing with that money.

Jesse: Yeah, this definitely is, like, the first step towards doing any kind of showback models, or chargeback models, or even unit economics to figuring out where your spend is going. The Cost and Usage Report is going to be a huge first step in that direction.

Amy: Now, the reason why we yell at people about this—or at least I do—is because AWS will only show you the data from the time that it is turned on. They do have it for historical periods, but if you enable it at a specific point, all of your reports are going to start there. So, if you’re looking to do forecasting, or you want to be able to know what your usage is going to be looking like from this point on, turn it on as early as possible.

Jesse: Absolutely. If you are listening to this now and you don’t have the CUR enabled, definitely go pause this episode, enable it now, and come back and listen to the rest of the episode because the sooner you have the CUR enabled, the sooner you’ll be able to get those sweet, sweet metrics for all of your—

Amy: And it’s free.

Jesse: [laugh]. Yeah, that’s even the more important part. It’s free. There’s going to be a little bit of data storage costs if you send this data to S3, but overall, the amount of money that you spend on that storage is going to be optimized because you’re saving that CUR data in Parquet format. It’s absolutely worthwhile.

All right, so number two; the second item on our cloud cost management starter kit, is getting to know your AWS account manager and account team. This one, I feel like a lot of people don’t actually know that they have an AWS account manager. But let me tell you now: if you have an AWS account, you have an AWS account manager. Even if they haven’t reached out to you before they do exist, you have access to them, and you should absolutely start building a rapport with them.

Amy: Anytime you are paying for a support plan, you also have an account manager. This isn’t just true for AWS; I would be very surprised for any service that charged you for support but did not give you an account manager.

Jesse: So, for those of you who aren’t familiar with your account manager, they are generally somebody who will be able to help you navigate some of the more complex parts of AWS, especially when you have any kind of questions about your bill or about technical things using AWS. They will help you navigate those resources and make sure that your questions are getting to the teams that can actually answer them, and then make sure that those questions are actually getting answered. They are the best champion for you within AWS.

If you have more than a certain threshold of spend on AWS, if you’re paying for enterprise support, you likely also have a dedicated technical account manager as well, who will be basically your point person for any technical questions. They are a great resource for any technical questions, making sure that your technical questions are answered, making sure that any concerns that you have are addressed, and that they get to the right teams. They can give you some guidance on possibly how to set up new features, new architecture within AWS. They can give you some great, great guidance about the best ways to use AWS to accomplish whatever your use case is. So, in the cases where you’ve got a dedicated technical account manager as well, get to know them because again, they are going to be your champion. They are here to help you. Both your account manager and your technical account manager want to make sure that you are happy with AWS and continue to use AWS.

Amy: And the thing to know about the account manager is, like, if you ever run into that situation where, oh, something was left on erroneously and we ended up with a spike, or this is how I was understanding the service to work and it didn’t work that way, and now I have some weird spend, but I turned it off immediately, if you ever want to get a refund or a credit or anything, these are the people to talk to; they’re the ones who are going to help you out.

Jesse: Yeah, that’s a great point. It’s like, whenever you call into any kind of customer support center, if you treat the person who answers the phone with kindness, they are generally more likely to help you solve your problem, or generally more likely to go out of their way to help you solve your problem. Whereas if you just call in and yell at them, they have no interest in helping you. So—

Amy: You’ll never see that refund.

Jesse: Exactly. So, the more that you can create that rapport with your account manager—and your technical account manager if you have one—the better chances that they will fight for you internally to go above and beyond to make sure that you can get a refund if you accidentally left something running, or make sure that any billing issues are taken care of extremely fast because they ultimately have already built that rapport with you. They care about you and the way that you care about them and the way that you care about continuing to use AWS.

Amy: There’s another note about the technical managers where if you are very open with them on what your architecture plans are—“We’re going to move into this type of EKS deployment. This is the kind of traffic we think we’re going to run, and we think it’s going to be shaped this way”—they’ll help you out and build that in most efficient way possible because they also don’t want the resources out there either being overutilized or just being run poorly. They’ll help you out in trying to figure out the best way of building that. They’ll also—if AWS launches a new program and you spent a lot of money on AWS, maybe there’s a preview program that they think will help you solve a very edge case kind of issue that you didn’t think you had before.

Jesse: Absolutely.

Amy: Yeah. So, it’s a great way to get these paths and get these relationships because it helps both parties out.

Corey: This episode is sponsored in part by VM Ware. Because lets face it, the past year hasn’t been kind to our AWS bills or, honestly, any cloud bills. The pandemic had a bunch of impacts. It forced us to move workloads to the cloud sooner than we would otherwise. We saw strange patterns such as user traffic drops off but infrastructure spend doesn't. What do you do about it? Well, the CloudLive 2021 Virtual Conference is your chance to connect with people wrestling with the same type of thing. Be they practitioners, vendors in the space, leaders of thought—ahem, ahem. And get some behind the scenes look into the various ways different companies are handling this. Hosted by Cloudhealth by VM Ware on May 20th the CloudLive 2021 Conference will be 100% virtual and 100% free to attend. So you really have no excuses for missing out on this opportunity to deal with people who care about cloud bills. Vist cloudlive.com/corey to learn more and save your virtual seat today. Thats cloud l-i-v-e.com/corey c-o-r-e-y. Drop the “e,” we’re all in trouble. My thanks for VM Ware for sponsoring this ridiculous episode.

Jesse: So, the third item on our cloud cost management starter kit is identifying all of your contracts. Now, I know you’re probably thinking, “Well, wait. I’ve just got my AWS bill, what else should I be thinking about?” There’s other contracts that you might have with AWS. Now, you as the engineer may not know this, but there may be other agreements that your company has entered into with AWS: you might have an enterprise discount program agreement, you might have a private pricing addendum agreement, you might have an acceleration program—migration program—agreement. There’s multiple different contracts that your company might have with AWS, and you definitely want to make sure that you know about all of them.

Amy: If you’re ever in charge of an architecture, you’re going to want to know not just what your costs are at the end of the day, but also what they are before all your discounts because those discounts can maybe camouflage a heavy usage if you’re also getting that usage covered by refunds and discounts.

Jesse: Absolutely, totally agreed. Yeah, it’s really, really important to understand, not just your net spend at the end of the day, but your actual usage spend. And that’s a big one that I think a lot of people don’t think about regularly and is definitely important to think about when you’re looking at cloud cost management best practices and understanding how much your architecture is actually costing you on a team-by-team or product-by-product basis.

Amy: Also, make sure if you’re doing reservations that you know when those reservations and savings plans ent—

Jesse: Yes.

Amy: —because you don’t want to have to answer the question, “Why did all of your costs go up when you actually have made no changes in your infrastructure?”

Jesse: Yeah. Half the battle here is knowing that these contracts and reservations exist; the other half of the battle is knowing when they expire so that you can start having proactive conversations with teams about their usage patterns to make sure that they’re actually fully utilizing the reservations, and fully utilizing these discounts, and that they’re going to continue utilizing those discounts, continue utilizing those reservations so that you could ultimately end up purchasing the right reservations going forward, or ultimately end up renegotiating at the correct discount amount or commitment amount so that you are getting the best discount for how much money you’re actually spending.

So, the last item on our cloud cost management starter kit is thinking about the non-technical parts of projects. Amy, when you think about the non-technical parts of projects, what do you think about?

Amy: Non-technical always makes you think of people and process. So, this would be the leadership making the decisions on what those cost initiatives are. Maybe they want to push this down to the team lead level: it would include that. Or maybe they want to push it down to the engineering level, or the individual contributor level. There are some companies that are small enough that an engineer can be completely cognizant and responsible for the spend that they make.

Jesse: Yeah. I think that this is a really, really critical item to include in our starter kit because leadership needs to be bought into and back whatever work is being done, whatever cloud cost management work is being done. But also teams need to be empowered to make the changes that they want to make, make the changes that will ultimately provide those cloud cost management optimization opportunities and better cost visibility across teams. So, does everybody know what their teams are empowered to do, what their teams are capable of? Does everybody know what their teams are responsible for on the flip side? Do they ultimately know that they are responsible for managing their own spend, or do they think that the spend belongs to somebody else? Also, do they understand which resources are part of their budget or part of their spend?

Amy: It’s the idea that ownership of—whether it’s a bill, whether it’s a resource—comes down to communication, and level setting. Do we know who owns this? Do we know who’s paying for it? Do they know the information in the same way? Is there someone who’s outside who can figure out this information for themselves? Just making sure that it’s done in a clear enough way that everyone knows what’s going on.

Jesse: Absolutely. Well, that will do it for us this week. Those are our four main items for our cloud cost management starter kits. If you’ve got questions you’d like us to answer, please go to lastweekinaws.com/QA, fill out the fields and submit your questions.

If you’ve enjoyed this podcast, please go to lastweekinaws.com/review and give it a five-star review on your podcast platform of choice, whereas if you hated this podcast, please go to lastweekinaws.com/review, give it a five-star rating on your podcast platform of choice and tell us, what would you put in your ideal starter kit?

Announcer: This has been a HumblePod production. Stay humble.

View Details

Links:

  • Github main site: https://github.com/
  • The Open Guide to Amazon Web Services:https://github.com/QuinnyPig/og-aws
  • Slackhatesthe.cloud: slackhatesthe.cloud
  • Global Diversity CFP Day: https://www.globaldiversitycfpday.com/
  • Twitter: https://twitter.com/deniseyu21
  • Personal site: deniseyu.io

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Thinkst. This is going to take a minute to explain, so bear with me. I linked against an early version of their tool, canarytokens.org in the very early days of my newsletter, and what it does is relatively simple and straightforward. It winds up embedding credentials, files, that sort of thing in various parts of your environment, wherever you want to; it gives you fake AWS API credentials, for example. And the only thing that these things do is alert you whenever someone attempts to use those things. It’s an awesome approach. I’ve used something similar for years. Check them out. But wait, there’s more. They also have an enterprise option that you should be very much aware of canary.tools. You can take a look at this, but what it does is it provides an enterprise approach to drive these things throughout your entire environment. You can get a physical device that hangs out on your network and impersonates whatever you want to. When it gets Nmap scanned, or someone attempts to log into it, or access files on it, you get instant alerts. It’s awesome. If you don’t do something like this, you’re likely to find out that you’ve gotten breached, the hard way. Take a look at this. It’s one of those few things that I look at and say, “Wow, that is an amazing idea. I love it.” That’s canarytokens.org and canary.tools. The first one is free. The second one is enterprise-y. Take a look. I’m a big fan of this. More from them in the coming weeks.

Corey: This episode is sponsored in part by VM Ware. Because lets face it, the past year hasn’t been kind to our AWS bills or, honestly, any cloud bills. The pandemic had a bunch of impacts. It forced us to move workloads to the cloud sooner than we would otherwise. We saw strange patterns such as user traffic drops off but infrastructure spend doesn't. What do you do about it? Well, the CloudLive 2021 Virtual Conference is your chance to connect with people wrestling with the same type of thing. Be they practitioners, vendors in the space, leaders of thought—ahem, ahem. And get some behind the scenes look into the various ways different companies are handling this. Hosted by Cloudhealth by VM Ware on May 20th the CloudLive 2021 Conference will be 100% virtual and 100% free to attend. So you really have no excuses for missing out on this opportunity to deal with people who care about cloud bills. Vist cloudlive.com/corey to learn more and save your virtual seat today. Thats cloud l-i-v-e.com/corey c-o-r-e-y. Drop the “e,” we’re all in trouble. My thanks for VM Ware for sponsoring this ridiculous episode.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Denise Yu, who's a senior software engineer at a company that mispronounces itself as GitHub, but it's actually pronounced Jif-ub. Denise, welcome to the show.

Denise: Thanks so much for having me, Corey. I am excited to scream in some clouds with you.

Corey: Yes, and I do want to point out that Jif-ub is not accepted by the Jif-ub marketing team, and they're upset. But that's the nice thing about this being my show; they don't get to dictate the nonsense that I say. But this is not you endorsing my correct pronunciation, as an employee, in any way. Because that's not your role. You're a senior software engineer. What do you do?

Denise: So, I work on a bunch of different things. I actually originally joined GitHub back in March at the beginning of this global pandemic that we're still living through, somehow, in the year 2021. But I joined a year ago to work on a team called ‘Community and Safety.’ That team has been rolled into other teams by this point, partly because the original team did too good of a job at creating abuse prevention tools for the platform. But these days, I still work on community-facing tools, which is very exciting.

I'm actually in a department called Communities. And our whole charter is to build tools that make life better for open-source maintainers. So that's a mission that I'm actually really excited about. I think [laugh] of all the products that I've worked on, this is one of the few that I can say I actually really want to make life better for the end-users. [laugh].

Corey: Oh, absolutely. For better or worse, GitHub—or Jif-up—has taken a central role in the open-source community. And on some level, I find it—how do I frame this—kind of weird in that we've taken this protocol that is widely decentralized, and that's its entire point. Well, what do we do about this? Oh, immediately, we're going to re-centralize it. That's the right answer.

And it just never ceases to amuse me that that's what we have taken from it as a society. But it works. I can't argue otherwise. And I maintain—and this is a bet that I am definitely going to the wall with—that in five years or so we are going to look back at Microsoft acquiring GitHub as a transitional moment.

Denise: Yeah, I think so. I think even though the Git technology is built for resilience and distributed work, I think what GitHub as a platform has shown that the process of building software is more than just pushing lines of code. The code is the artifact of collaboration, but we still need a place to do that collaboration, which is why—I’ve been using GitHub for five or six years now, and I only realized a few years ago that things like pull requests and issues don't exist outside of GitHub. That's an abstraction that's so, so deeply baked into my idea of what it means to work collaboratively. That didn't exist until GitHub invented it, which is pretty wild.

Corey: Yeah, it's bizarre in a number of different levels. But so much of what we think is part of Git is not. It is the GitHub abstraction, but it's also something that is widely copied by all of GitHub’s competitors, in many respects. So, the line has gotten very, very blurry. And how people come to Git is also fascinating.

I used to go to a Git training, I think in 2009, conducted by a GitHub employee—I may be misremembering the year by a few years in either direction—and it was a multi-day process, and it was complex, and I left it feeling in many ways that I had more questions than I answered. Now, I point people to a GitHub page that talks about how to use Git, and they're mostly there. So, it isn't that Git necessarily has improved as a product, it's that GitHub has made it far more accessible. And let's face it, after a few million times practicing, you get very good at explaining complicated things in simple ways.

Denise: Oh, yeah. For sure. It's not a huge surprise that internally, we use GitHub for everything. So, if you want to, I don't know, collaborate on writing a—even a new blog post, new marketing copy, or new documentation, all of that happens through GitHub. And I think that people with different levels of proficiency with command line—actually these days, you don't even need command line. You do everything through the UI, which I think is pretty neat.

Corey: Oh, it's phenomenal. And I do want to call out that I am the maintainer of an open-source project on GitHub. We know it's good because it has [28.3000 00:06:18] stars; at the time of this recording, almost 3000 forks. And what it is, is the “Open Guide to AWS” it's a big markdown document that more or less explains what all of the AWS services do.

There are over 200 contributors, we have an online Slack community at slackhatesthe.cloud because they don't like open-source community stuff very well, and we make fun of them for it. But my point being is that even without knowing how to write a single line of code, this is more or less just a big readme that explains different aspects of what AWS is because it's incredibly hard to adjust to that. And that is every bit as much an open-source project as anything that's included in NPM, or any command line tool that you'll use in any aspect of your job.

Denise: Yeah, exactly. And I think the nice thing about editing a document like an AWS Guide, or anything else, really. A couple years ago, I would have said, “Okay, that sounds like it should be a wiki page.” Wiki pages support revisions, collaboration—or maybe a Google Doc or something. But the nice thing about putting it somewhere like GitHub, is like, well, you're already—probably all your other tools already integrate into GitHub. Like, you maybe get Slack notifications, when you have a PR requires review, or whatever it is. It's just so much easier to have everything in one place. And you also get the cool green squares for contributing. [laugh].

Corey: Oh, of course, the most important thing is fill that thing up. It’s—talking about gamification causing weird behavior patterns. Yeah, we have a GitHub workflow on this, that fires off on every pull request that looks at the entire site, and validates the thousand and some odd links are still resolving and valid because when you link to this many different sites across the internet, link rot’s a real thing, and it's depressing and sad to be able to look at this and, “Oh, that didn't work.” Now, we do have to fix some of this because it winds up, in some cases, flagging people's submitted pull requests as not being valid, but it has nothing to do with what they're submitting because I'm bad at this. But the fact that that's built in and is available for use is incredible. Every time I look at GitHub’s site, it's gotten better than the last time I looked at it.

Denise: Oh, that's so nice.

Corey: Well, can't shake the feeling that at some point, there's going to be at the proper moment, a deploy to Azure button that as soon as someone clicks, whatever they've written is suddenly running in some environment in Azure. Now, if it comes too soon, it's going to be terribly implemented, and no one's going to trust it; if it comes too late, then people will already have solved this problem in different ways, so timing is going to be critical. But that is going to be—well-executed at least—a potential sea change in how people approach Cloud and what the default environment to run something on it.

Denise: Wait, have you looked at Codespaces? Or seen the demo around that?

Corey: I was hoping you would bring those up. Back in the before times, I would travel kind of a lot. And the only computer I brought with me was my iPad.

Denise: Mm-hm. Yeah, a lot of people do this.

Corey: Oh, yeah. Trying to get a environment on an iPad, it turns out, it's not terrific. The solution that I fell back to was, well, I've been using VI for 20 years, so I'll just SSH into an instance and call it good. But there were interesting projects of varying degrees of success that would run VS Code in a instance somewhere and then just present its interface as a web page. GitHub has actually integrated this as a core offering: all right, you click a button, it spins it up, and it's there.

Denise: Yep.

Corey: And oh, my God, is that better.

Denise: Yeah. I think it's still in limited beta. So, I think you still have to request access to it—I don't know if this will still be the case in March, but yeah, I've tried it out a couple of times, and it works really, really well. My only gripe is that I don't like to use VS Code, so I'm going to have to learn VS Code in order to use it.

Corey: I've got to say that is my single biggest barrier to adoption. I already know how to use VI, but a lot of the VS Code stuff, of you having a full featured IDE that does all of these things, is extraordinarily heavy of a lift for me. It more or less is breaking 20 years of muscle memory.

Denise: Yeah. Tons of my co-workers love VS Code and use it every day. I only know how to use RubyMine for, [laugh] like, large monolithic Ruby development. And if I'm not doing work on the monolith, I don't want to use an IDE otherwise; I want to just use VI. I'm working in a small enough codebase where I know the name of every file, then I don't want to use the IDE for it.

The VS Code to me sits in this middle ground that I know I need to learn how to exist in that middle ground, but it requires me going off to the internet and hunting down every single plugin that I need to put onto VS Code to make it feel like RubyMine again. And that's the thing that I don't want to do.

Corey: Well, I'll take it a step further, and most of the tutorials I see about how to get up with VS Code as an IDE are JavaScript based—or YavaScript. [unintelligible 00:11:01] pronunciation once again—and my lingua franca is crappy Python. So, whenever I look for, oh, I want a VS Code tutorial for Python, the first eight pages of it are here the different ways you can do virtual environments and dependency management because that is Python. It focuses on that, and it winds up getting hung up on those implementation details before ever getting to, “Here's a reasonable Python project that someone who's not very good at Python can understand, and here's how VS Code can add value to it.” I mean, those posts exist, but they're hard to find. And it honestly makes me feel like an imposter for not knowing JavaScript.

Denise: Yeah, the feature that I use the most in a full IDE like RubyMine is command-click to go to definition. And with Ruby, you get the correct definition about 70 percent of the time, but it's still better than nothing. Because if I'm not doing that, I'm doing a full text search for the method definition, and that's like—my brain can only work with one of those two things. With VS Code, after 15 minutes of trying to find the right plug-in that does click to go to definition and failing, I was just like, this is just not worth it. Nobody should have to live through this much pain. [laugh]. I don't want to do this anymore. [laugh].

Corey: Exactly. It becomes this situation where at least when you're starting to learn something, it's breaking everything you're used to, and it feels like instead of helping you, you're fighting with it.

Denise: Yeah, this reminds me of how people talk about switching their keyboard layout to Dvorak, or whatever—

Corey: Oh, my God. I want to do that on one hand because I'm deep into the mechanical keyboard space. But on the other is, I don't really have six months to teach myself to type all over again.

Denise: Yeah, that's the thing. It's like, maybe it'll speed up my typing by, like, two percent. But I already type really fast. I think I would have extremely diminishing returns on that. But also, I don't want my brain to just not work for weeks or months.

Corey: Oh, it is worse than that because as soon as you leave that and go on to someone else's computer—which admittedly is not nearly as frequent as it used to be when I was in help desk and desktop support—“Oh, can you type in your password on this thing?” “Absolutely not. Thank you for asking.” It's effectively having to go back and restart it over. I understand that it's more efficient, but you need a societal shift to it, not individuals doing it piecemeal.

Denise: Yeah, that's true.

Corey: So you're relatively recent, as far as hires over a Jif-ub. You joined mid-pandemic, or early pandemic, depending on how you want to slice that, and you've never been to their headquarters. What is that like?

Denise: Everyone keeps telling me how awesome HQ is like, and I’m like, “Cool.” I actually have a brick wall in my office that looks like some of the spaces in HQ, I guess, so people get super confused when I tell them that I'm in Toronto.

Corey: Oh, yeah, I've been to HQ a bunch of times for various things. It's great. You’ve really—

Denise: Yeah.

Corey: —got to see it. I'm not helping, am I?

Denise: [laugh]. You're really not helping. I really want to go to HQ. I was meant to go to HQ in March. So, my start date was March 13th, or 14th—I can't remember—in 2020. And I was part of the first onboarding class to not get to go out to HQ because that was when the travel restrictions around COVID were finally being implemented. So I'm very sad about that.

Corey: I will say that GitHub has the advantage of a lot of other companies in that—to my understanding—they had a remarkably resilient remote culture before this all hit. Is that correct, or am I misremembering my history?

Denise: No, you're right. I think the figure was 75 percent of employees don't live within commuting distance to HQ when I first joined.

Corey: Well, not with that attitude.

Denise: [laugh]. Yeah, that's true. We just got to try harder. But I think that number is a lot higher now because we have been hiring internationally. And we've grown a lot in the last year or so that I've been here.

Yeah, so I think we were better equipped than a lot of other companies to go full remote because it basically just was telling that 25 percent of employees, “Don't come into HQ anymore.” And giving people a little bit more office budget, I guess, to get set up fully for full time remote work. There are a lot of challenges though, especially when you join as a remote employee and you've never worked in a fully remote environment before. It took me a really, really long time to get used to the fact that Slack was my portal to everything.

Corey: Oh, yeah. Here at the Duckbill Group, we were full remote from the beginning, just because I decided to merge my consultancy with my business partner’s ten days before he moved from across the street to Portland, Oregon. So, we at that point decided, all right, we're full remote. And we've never had a critical mass of staff in one city to the point where they're became a de facto headquarters because as soon as that happens, it changes your culture. Trying to retrofit something like full remote to a company that historically has always been face-to-face, it shatters a lot of the dynamics. So, on some level it was easy for us. On the other, it feels like we're at a bit of a disadvantage in some respects, once things return to normal. But that's a future problem, unfortunately.

Denise: Yeah, I think the biggest thing is learning to work asynchronously, which GitHub was already pretty good at. So I'm pretty used to opening a PR on Wednesday evening, for example, and not really expecting to get feedback on and until the next day, Thursday, or even Friday, especially if I needed a feedback from a different team. Actually, that's something that I've learned since joining GitHub; that wasn't something I was used to before because I've only ever worked in co-located environments before. Someone didn't review my PR in time, I would take a Nerf gun and go over to them [laugh] and gently ask them to [laugh] review my PR. But it's been very interesting.

So, I also used to work at Pivotal, which I'm sure you've heard rumors about our kind of cultish pairing [laugh] culture. So, my default state of working was to constantly be talking to another human about what I was doing all the time. And so, suddenly joining a fully remote company in a time when we have to work a little bit more asynchronously than before because people's lives are just completely upended by this pandemic, that was a big adjustment.

Corey: So, in the before time, something else that you were effectively renowned for was not only giving conference talks yourself, but helping other people give conference talks, too. Let's talk a little bit about that because speaking at conferences was always one of my passion projects because I love the sound of my own voice, simply love it. And helping other people do that, too, was also something I was very interested in doing, but it was always hard for me to get traction around getting people to let me help them for reasons that are probably blindingly obvious to everyone except me. So, tell me about that. What's the story?

Denise: Yeah, so I have always been pretty comfortable speaking in public. I think it's something that I almost took for granted. And the reason for that is because I spent upwards of 10 years of my teenage and adult years doing competitive debate. It was traveling to debate tournaments pretty much every single weekend in college and spent a lot of time doing both structured and off-the-cuff speaking in front of large groups of people. And when I got into tech, it took me a long time to realize that not everyone is like me. I kind of had this special strand, I guess of imposter syndrome, where I think that if I know how to do a thing, then everybody else knows how to do that thing also. Which is aggressively not a correct assumption to make, ever.

Corey: Oh, same here across the board. It's one of those I thought for a while I fit in tech because the stereotypes seem to fit me super well. It's like, yeah, “We're bad with people.” “Absolutely.” “We prefer dealing with computers to humans.” “Absolutely.” “I hate my boss.” “Absolutely.” “I'm good at programming…” “Oh… okay, never mind.” But yeah, there's a lot that I think people miss about what it takes to give a good talk. It's not technical mastery. It is an ability to tell a story.

Denise: Exactly. I think you hit the nail on the head there. And so after I realized that, okay, I am, at least at the median. I actually think I'm quite good at public speaking and telling stories about technical subject matter, and breaking things down in a way that someone without as much technical context as me can understand it. The thing that got me into doing more and more public speaking, specifically at tech conferences was, after I gave my first talk, people just kept coming up to me, basically, and being like, “Hey, I saw you talk to this conference. Do you want to come talk at this other conference that I'm running?” or whatever. I think like, once you've got a couple of larger stage conference talks under your belt, you almost don't really have to look for speaking opportunities anymore.

Corey: No, they fall into your lap. And then you have different problems such as, oh, I've been invited to speak at this conference. Let me do some checking to figure out are they a trash fire, or are they a reasonable place?

Denise: Yeah. Am I going to be the only not white guy speaking at this conference?

Corey: Am I going to be on a manel? I won't do it. Is there not a code of conduct that has actual teeth and doesn't read as a joke? I won't do it. And that stuff is way more important to me than the technical content of the talk.

Denise: For sure. For sure.

Corey: But tying it back to even the stuff that GitHub does, one of the talks that was transformative for how I approach this was a Git talk, “Terrible Ideas in Git.” And I like to teach people things by counterexample, and the first iteration of that talk would have been great if the Git maintainers themselves had been the target audience for this, but two people would have loved that and the rest of the audience would have been looking at this with what the hell are you talking about? So, I had to refactor it, and it made it a far stronger talk all the way to the point where there's jokes for everyone in there. But my favorite time giving the talk, I said that, “I now need to be able to explain Git in such a way that my mother would understand it. And that is a problematic way of framing on ageism and sexism basis in most cases, but not this time because she's sitting in the front row. Hi, mom.” That was one of my favorite speaker moments that was just, it's hard to get that one pulled off, mostly because getting my mother to fly to various parts of the world and watch me speak is a bit of a heavier lift than one expected. But here we are.

Denise: Yeah, I use a similar litmus test for my stepdad. My stepdad is not a technologist. He is a biologist by training, actually. But he asks me a lot of questions about what I do. And back when I worked at Pivotal, I had to explain cloud virtualization to him every time I went home. [laugh].

I don't know if it's gotten easier. With GitHub, I feel like GitHub is a little bit easier to explain to the average person than Cloud Foundry. But my bar kind of is, every time I put together a talk, do I think that my stepfather could understand the content.

Corey: This episode is sponsored in part by our friends at New Relic. If you’re like most environments, you probably have an incredibly complicated architecture, which means that monitoring it is going to take a dozen different tools. And then we get into the advanced stuff. We all have been there and know that pain, or will learn it shortly, and New Relic wants to change that. They’ve designed everything you need in one platform with pricing that’s simple and straightforward, and that means no more counting hosts. You also can get one user and a hundred gigabytes a month, totally free. To learn more, visit newrelic.com. Observability made simple.

Corey: Increasingly, I found the best way to explain GitHub to folks who aren't familiar with it and are non-technical but work in the business space has been, “You know how Microsoft Word has track changes? Yeah, imagine that only without the copy-of-copy-of-file-final-use-this-one-date.bak.PDF.doc? Yeah, imagine that. Instead, this makes this streamlined and easy for multiple people to work on at the same time,” at which point they stare at you hungrily and say, “Oh, my God, how can I use this for my work?” And I feel like there's a breakout moment waiting for GitHub if they ever decide to focus on that part of the market because my God, is there an appetite for it?

Denise: Yeah, I think like that kind of revision driven editing, I think that's applicable to so many different domains. I've heard people experimenting with the idea for—I think, Figma does something similar for designs for visuals, but I've also heard the idea floated of this being done for, like, sound mixing, sound engineering, which I think would be super, super cool. I don't know how you could visualize that, but someone who has more knowledge of that space, I'm sure could come up with a way.

Corey: Oh, yeah, the challenge I find with stuff like that, for podcasts, for example, is GitHub gets very, very angry because Git gets very, very angry when you start doing things like checking in large uncompressed media files, and then making small changes to them, but you can't diff that so every revision for video work is, “Oh, here's another 150 gig file to put up,” and that becomes a problem.

Denise: Yeah.

Corey: Yeah, there's a lot that's neat out there. I also want to credit GitHub with fixing the state of working with Git because let's face it, Git makes everyone feel stupid at some point. The question is, is just how far along are you before that hits? And invariably, the approach was always, when you get into that territory, you wind up asking the question on some forum somewhere, and then you wind up getting a few minutes later, this giant essay where it starts out, “Well, Git is not really a version tracker, it's instead of distributed graph that”—and then you just skip down eight paragraph to the bottom where they give you the one-liner that fixes the problem.

GitHub has fixed the result of queries for, “How do I X in Git?” And just gives you the answer outright. It's transformative, and it's one of the things that everyone takes for granted because no one really stops to think this documentation is awesome, and accessible, and answers my problem, but the result of it is, is that you just don't see people complaining about how hard Git is anymore.

Denise: Yeah, yeah. Oh, my God, Git is so hard. I think the first-time I correctly did an interactive rebase—like, I typed all the arguments correctly. I think I just got up and did a lap around the office. [laugh]. There's just all this random stuff that you have to remember when you're using—especially some people use the desktop GUIs or whatever. I always just use command line. But it took forever and I still make a ton of mistakes with it.

Corey: Oh, we all do. That's the beautiful part of it. The nice thing is, “Oh, well the nice thing about Git is that nothing ever really gets truly lost.” And you say, “Well, how do you figure?” “Oh, I copy the entire directory to a backup before I do anything.” Which is on the one hand, awful version control and, two, has saved my bacon multiple times. So, if it's stupid and it works, it's not really stupid, now is it?

Denise: It's not stupid if it works. [laugh].

Corey: So, a passion project of yours that I want to talk about has been teaching people to submit to conferences, to give talks at conferences, to convince them that yes, you can do this and you're good at it, you have a story that people want to tell, the things you find easy are not easy for everyone. ‘Give that story voice’ has been a recurring theme that you do. How has that manifested during this year of isolation?

Denise: Ohh. This year’s definitely been tough. So, usually when I do that pitch, I run workshops specifically around a certain CFP. Like I help out conference organizers in the Toronto area. Like, I did a workshop for devopsdays last year, and another one for GoCon, also meant to be last year.

And usually it's like, okay, here's how you write a strong proposal—and I'll have the committee in the room—and be like, here's what the committee is looking for; ask these people questions about what they would choose to program. And this is how you get your talk submitted. This is how you get accepted; you need to understand who your audience is. It's a product question. It's exactly the same mindset that you need to get into when you're playing product manager, or if you're actually a product manager.

You need to understand who your target customer or audience is, what they're looking for, and how do you give that to them on a plate. This year, it's slowed down a lot, obviously because there have been fewer conferences, so I've just had fewer opportunities to do this kind of pitch without—it feels weird telling people to write talks without there being a place to submit a talk to. But I think—I found that there still is room for content creation, even if that isn't a conference talk. So, within the company, within my networks, I do try to encourage people to write blog posts. Err on the side of writing what you've learned; if you spent a couple days doing a technical investigation and you found it interesting, write it down so other people can learn about it.

On the conference talk front, I think the unique benefit of doing conference talks and getting up on stage is you do build a personal brand. After a little while that kind of transcends the company that you're at, granted that you're not talking about your company's products in every single talk.

Corey: Absolutely. And a lot of companies get profoundly insecure about that. And I'm sorry; they're wrong. That is not a bad thing. Anyone who gives a talk for your company will eventually become better known than your brand for some aspects and in some areas—harder at places like GitHub than it is Twitter for Pets, but that's okay—and you have to be okay with that and have a plan for how you're going to transition them out when they outgrow the role and move somewhere else. But that's okay. And companies that are insecure about that drive me up a wall.

Denise: Yeah, there have been a few times when I went to very cloud specific conferences and talked about things that my company is working for, but other times, I've gone to just very general conferences where it was just like, let's talk about Agile, or let's talk about software, like super, super generally. And the ones who have gone to the really general conferences and talk about things that matter to me—like cross-functional collaboration, or test driven development, or whatever it is—those are the ones that I've actually managed to get people to apply to my company without even saying that we're hiring. It's a little mind blowing.

Corey: It is. People want to work with interesting people, and when you find someone who's able to tell a story that touches you on some level, which is the point of telling stories, then at some point, the idea of I'd like to work with that person on ongoing basis becomes actively compelling. People ask so well, it's not like you could hire a whole lot of people who do what you do, so hiring and recruiting is going to be a problem. It's like, well, it is but it's because it's finding the right people with the right alignment, but we have an awful lot of candidates to choose from for basically any role, just because for better or worse, I make a splash.

Denise: Yeah, exactly.

Corey: So, as people are looking to get into the space and they're new to speaking, how do you get them to a point where there's at least a certain level of confidence? I would say, “How do you get them to not be terrified?” But I’ve been speaking almost full time for years, and I'm still scared to death every time I give a talk when I start out.

Denise: Me too. I'm still terrified. I almost switch into a different persona when I go on stage, and I don't remember anything that happens. I don't know what—maybe there's psychology [laugh] around this, but I literally blank out the 40 minutes that I'm on stage. So, the most common objection that I hear about why someone doesn't want to give a talk is they'll say, “But everybody already knows about this. The thing that I want to talk about has already been talked about. There's already too much content out there. I wouldn't add anything new.”

And for a long time, my response to that was just, I don't care. I still want to hear what you have to say, you're going to offer a fresh perspective, you're going to offer a new experience, you're going to explain something in a way that clicks for someone. Your explanation might be the one that works for them that they haven't been able to figure out by listening to anybody else. So, I historically said those things, and I still one hundred percent those things, and I still say those things to people but I actually came across something interesting on Twitter today. Someone was actually talking about fanfiction. I don't know if you’ve written or read fanfiction in the past. When I was a kid I used to—

Corey: I can neither confirm nor deny…

Denise: [laugh]. There’s varying levels of fanfiction out there. But fanfiction communities are actually incredibly collaborative places. I remember being, I don't know, like a 10 or 11-year-old and writing a lot of anime fanfiction, basically just writing myself into the Pokemon universe and things like that. And we would post these on forums.

It'd be forums of strangers working together, and we’d post drafts. And then we would give each other feedback on those drafts. And in the end, there'd be 60 different pieces of fanfiction, roughly about the same universe, every story is a little bit different, and nobody was ever self-conscious that this had already been done. Because the whole point of fanfiction is that there's a lot of it, the value of the community is in the fact that the content is abundant. Not that every single piece of content is unique.

So, I think a similar thing is true for content that's technical storytelling. Just because someone else has said this thing before, it doesn't matter. The fact that there's more content out there about this thing is good for the community, building a community around this thing. So, I don't know, that was something that just resonated with me a lot, and I just really liked that framing.

Corey: We have this weird perspective that the things that we know are easy, the things we don't know are hard, so once you've learned how to do something, everyone else knows it. I mean, one of my most popular Twitter threads that exploded that I didn't see coming was just me talking about random things that live in my shell config file, and what they do. It's, “Well, this is stuff everyone knows, and everyone has something like this, right?” And it got into a dissection of what these things do, and how it works, and of course, it's snarky because that's my excuse for a personality, but it also shined a light on something that I forgot, once again: not everyone knows this. A lot of the viral tweets you see going around are people noticing something that we've all seen and been ignoring because that's just the way it is, and someone makes an observation about it and, oh, yeah, that's right. That is something that everyone can relate to.

Denise: Yeah.

Corey: People want stories.

Denise: Yeah, exactly. I think about every six months, I try to retweet—every review cycle around mid-year, end-of-year. It’s like, “Hey, just a reminder, if someone on your team has done something for you write this down. Catalog what they did, write it down, send it to them, send it to their manager, put it into their file so they can use it at review time.” And every time I talk about this, I always get a ton of engagement.

And I have to remind myself that this is not obvious. Even if someone was told to do this at some point in the past, reminding people to be deliberate and be active about giving each other peer feedback, writing stuff down. You just can't remind people to do that enough, I think.

Corey: Yeah, it's one of those things where, “Well, I've already given that talk, so I have to give a different one.” No, you don't, I promise you, you have not given your talk to everyone in the industry, and the reason I know that is because by that point, you would see the view count on YouTube or whatnot—or the video winds up—and be flabbergasted because the industry is bigger than anyone thinks. I feel like I tell the same joke from time to time, but they always get laughs from the audience because it's new to them. My business partner and my romantic partner find it incredibly challenging, both—those are two people, incidentally—they find it incredibly irritating because, “Yes, yes, we've heard that one.” And every once a while they sit up, “Oh, Corey got a new joke.” But the rest of the world doesn't respond that way.

Denise: Yeah, exactly. So, I don't know if we've talked about this already, but that closes the loop on the whole imposter syndrome piece where—just because I think one strain of imposter syndrome is telling yourself, just because I know this thing, means everybody else must also know it. I must be the last person in the world to learn this piece of information.

Corey: There is no piece of technology, or anything even tangentially related to technology, that you couldn't do a tweet thread on and people would learn things about it. How an if-then-else loop works, for example. That is—sorry—ahh—do you see what I'm talking about? That's exactly what I mean. See, I meant to say ‘Conditional,’ and instead I said ‘Loop’ and we all trip over these things; they're all challenging, and the fact that we can still learn new things about that of, how does a flow control conditional actually work?

What does that mean? Sure, anyone who spends their days writing code all the time is going to have some topical understanding of it, but not everyone's going to, first, be that person, or secondly, understand why that's valuable and relevant. You could do an entire conference talk on nothing more than that concept alone and it would be a dynamite talk.

Denise: Yeah, exactly. I did go through a wave of imposter syndrome around last year, or maybe—uh, I don't even know what year is it. [laugh]. A little while back, I was feeling uncertain about this talk about distributed systems that I've done a few times. And then I got a random message on Twitter from a student who told me that she had just gotten her first engineering role out of university—ever, ever—and she said, I watched your talk on distributed systems; you gave me the vocabulary to talk about this stuff in an intelligent way, and I've landed my first role as an entry level—I forget exactly what kind of engineer but—I mean software engineering—but I forget if it was like infrastructure, or app dev, I don't know. But it was like, that kind of feedback, just—that's really cool to hear.

Corey: I'm going to take it a step further and request anyone listening to this who’s seen the talk that has potentially changed the way they view things or helped them advance in their career, track down whoever gave that talk and send them a message. One of the hardest parts about being a speaker is that you very rarely get feedback. I mean, even on this podcast, I almost never get direct feedback online. People talk to me about it in person, when they run into me at a conference, in the before times. But it's not a thing where people wind up drafting an email, “Dear sir, I found your podcast to be compelling, and also you are a jackass.” I would love letters like that.

Denise: Well, I don't love letters like that, so please don't write me letters like that. [laugh].

Corey: Well, in my case, they're well-deserved. In your case, it would just be offensive.

Denise: [laugh]. But yeah, reach out to those strangers whose content you've learned from, I think it can actually make someone's day.

Corey: Absolutely. So, on the other side of that, if someone's listening to this, and they want to start giving a talk, but they don't know where to start, how to build a CFP, where to begin, where should they start? What is the next step for those people?

Denise: So there's an initiative that I, in the before times, mentored and helped organize local chapters for. It's called Global Diversity CFP—it's either Global CFP Diversity or Global Diversity CFP, I always forget—day. And even if you're not a person who has a marginalized identity, they've done a really, really great job of consolidating a bunch of resources to get first-time speakers off the ground.

Corey: And we will put a link to that in the [show notes 00:36:34], of course.

Denise: Cool. So, I recommend reading through what they have. So it's actually an event that take place on a full day; it's a little bit different this year, I think they might be running it remotely this year. But all of their content is geared for first-time speakers, and they have a lot of material that's pepping people up to take the first step. But I would say that beyond that, the best way to get started, honestly, is to just pick a topic that you care about, and just try to write a talk.

And if you don't have a place to deliver that talk—if there's no meetup, or conference or anything coming up, you should just record yourself giving the talk to your computer, actually. This is the thing that my debate coach used to have us do in college: the best way to improve as a speaker is to record yourself speaking and watch it. It's going to be terrible the first few times because you're going to be like, “Oh, God. I look so awkward.”

Corey: Power through it. I still can't watch myself speak.

Denise: [laugh]. Like, what are my hands doing? But you can only catch those things and improve at them once you're aware that you're doing them. So, that is the fastest way to improve your presence and your speaking ability.

Corey: Yeah. I want to thank you for taking the time to speak with me and go through all these various topics today. If people want to learn more about you, where can they find you?

Denise: So, you can find me on the internet on the cursed blue website at @deniseyu21 is my handle on Twitter. My personal website is deniseyu.io because—I used to have .com but some kids stole it and squatted on it after I forgot to renew it. So I'm not deniseyu.com anymore.

Corey: Yeah, the joy of changing usernames on different platforms and the rest. I hear you.

Denise: Yeah. Yeah, probably those two places. If you want to talk about anything that I talked about on the podcast, feel free to DM me on Twitter. Please don't send me mean feedback that starts with “Dear sir,” in my Twitter DMs, but other than that, I—

Corey: Generally never a good plan for anyone.

Denise: [laugh]. No.

Corey: Thanks so much for taking the time to speak with me. I really appreciate it.

Denise: Yeah. Thanks so much, Corey for having me on. It was a lot of fun. I feel like we covered a lot of ground. Probably too much. [laugh].

Corey: It really was. And we did. Denise Yu, senior software engineer at Jif-ub. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you've hated this podcast, please leave a five-star review on your podcast platform of choice along with a comment containing a correctly formatted Git rebase command.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey atscreaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Brad

Author and Senior Executive Editor, Bloomberg Technology

Brad Stone is the author of four books, including Amazon Unbound: Jeff Bezos and the Invention of a Global Empire,published by Simon & Schuster in May 2021. It traces the transformation of Amazon into one of the largest and most feared companies of the world and the accompanying emergence of Jeff Bezos as the richest man alive. Brad is also the author of The Everything Store: Jeff Bezos and the Age of Amazon, which chronicled the foundational early years of the company and its founder. The book, a New York Times and Wall Street Journal bestseller, was translated into more than 35 languages and won the 2013 Financial Times/Goldman Sachs Business Book of the Year Award. In 2017, he also published The Upstarts: Uber, Airbnb, and the Battle for the New Silicon Valley.

Brad is Senior Executive Editor for Global Technology at Bloomberg News

where he oversees a team of 65 reporters and editors that covers high-tech companies, startups, cyber security and internet trends around the world. Over the last ten years, as a writer for Bloomberg Businessweek, he’s authored over two dozen cover stories on companies such as Apple, Google, Amazon, Softbank, Twitter, Facebook and the Chinese internet juggernauts Didi, Tencent and Baidu. He’s a regular contributor to Bloomberg’s technology newsletter Fully Charged, and to the daily Bloomberg TV news program, Bloomberg Technology. He was previously a San Francisco-based correspondent for The New York Times and Newsweek. A graduate of Columbia University, he is originally

from Cleveland, Ohio and lives in the San Francisco Bay Area with his wife

and three daughters

Links:

  • The Everything Store: https://www.amazon.com/Everything-Store-Jeff-Bezos-Amazon/dp/0316219282/
  • Amazon Unbound: https://www.amazon.com/Amazon-Unbound-Invention-Global-Empire/dp/1982132612/
  • Andy Jassy book review: https://www.amazon.com/gp/customer-reviews/R1Q4CQQV1ALSN0/ref=cm_cr_getr_d_rvw_ttl?ie=UTF8&ASIN=B00FJFJOLC

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Thinkst. This is going to take a minute to explain, so bear with me. I linked against an early version of their tool, canarytokens.org in the very early days of my newsletter, and what it does is relatively simple and straightforward. It winds up embedding credentials, files, that sort of thing in various parts of your environment, wherever you want to; it gives you fake AWS API credentials, for example. And the only thing that these things do is alert you whenever someone attempts to use those things. It’s an awesome approach. I’ve used something similar for years. Check them out. But wait, there’s more. They also have an enterprise option that you should be very much aware of canary.tools. You can take a look at this, but what it does is it provides an enterprise approach to drive these things throughout your entire environment. You can get a physical device that hangs out on your network and impersonates whatever you want to. When it gets Nmap scanned, or someone attempts to log into it, or access files on it, you get instant alerts. It’s awesome. If you don’t do something like this, you’re likely to find out that you’ve gotten breached, the hard way. Take a look at this. It’s one of those few things that I look at and say, “Wow, that is an amazing idea. I love it.” That’s canarytokens.org and canary.tools. The first one is free. The second one is enterprise-y. Take a look. I’m a big fan of this. More from them in the coming weeks.

Corey: This episode is sponsored in part by VMware. Because let’s face it, the past year hasn’t been kind to our AWS bills, or honestly any cloud bills. The pandemic had a bunch of impacts: it forced us to move workloads to the cloud sooner than we would have otherwise, we saw strange patterns such as user traffic drops off but infrastructure spend doesn’t. What do you do about it? Well, the CloudLIVE 2021 virtual conference is your chance to connect with people wrestling with the same type of thing, be they practitioners, vendors in the space, leaders of thought—ahem, ahem—and get some behind the scenes look into various ways different companies are handling this. Hosted by CloudHealth by VMware on May 20, the CloudLIVE 2021 conference will be 100% virtual and 100% free to attend, so you really have no excuses for missing out on this opportunity to deal with people who care about cloud bills. Visit cloudlive.com/coreyto learn more and save your virtual seat today. That’s cloud L-I-V-E slash Corey. C-O-R-E-Y. Drop the E, we’re all in trouble. My thanks to VMware for sponsoring this ridiculous episode.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. Sometimes people tell me that I should write a book about Amazon. And that sounds awful. But to be sure, today, my guest is Brad Stone, someone who has written not one, but two books about Amazon, one of which coming out on May 11th, or as most of you will know while listening to this, today. Brad, thanks for joining me.

Brad: Corey, thanks for having me.

Corey: So, what on earth would inspire you to not just write a book about one of what is in many ways an incredibly secretive company, but then to go back and do it again?

Brad: Yeah. I’m a glutton for punishment. And Corey, my hair right now is completely white way before it should be, and I think that Amazon might be responsible for some of that. So, as you contemplate your own project, consider that this company—will you already know: it can age you. They are sometimes resistant to scrutiny.

So, to answer your question, I set out to write The Everything Store back in 2011, and this was a much smaller company. It was a cute little tiny internet company of about $100 billion in market value. And poor, impoverished Jeff Bezos maybe had, I’d be guessing maybe $50 billion.

So anyway, it was a much different time. And that was a great experience. The company was kind of flowering as the book came out. And to my surprise, it was embraced not by Bezos or the management team, who maybe we’ll talk about didn’t love it, but by Amazon employees, and customers, and competitors, and prospective employees. And I was really proud of it that this had become a kind of definitive account of the early years of the company.

And then a funny thing happened. The little cute little internet company became a juggernaut, a $1.5 trillion market cap. Bezos is the wealthiest guy in the world now with a $200 billion fortune, and Alexa, and the rise of AWS, and the Go store, and incursions into India and Mexico and other countries, I mean, so much had changed, and my definitive history felt a little out of date. And so back in 2017—also a different world, Bezos is a happily married man; he’s the CEO of Amazon, Amazon’s headquarters are in Seattle only—I set out to research and write Amazon Unbound. And as I was writing the story, yeah, just, like, the ground kept shifting under my feet.

Corey: Not a lot changes in the big sphere. I mean, one of the things that Bezos said is, “Oh, what’s going to be different in 10 years? I think the better question is, ‘what’s going to not be different in 10 years?’” but watching the company shift, watching it grow, just from the outside has been a real wild ride, I’ve got to say. And I restrict myself primarily to the AWS parts because well, there’s too much to cover if you go far beyond that, and two, it’s a very different place with very different challenges around it.

I viewed The Everything Store when it came out and I read that, almost like it was a biography of Jeff Bezos himself. And in some respects, Amazon Unbound feels like it hews in that direction as well, but it also goes beyond that. How do you approach separating out the story of Amazon from the story of Jeff Bezos?

Brad: Yeah, you’re putting your finger on almost the core challenge, and the adjoining challenge, which is how do you create a narrative, a linear story? Often readers want a chronological story out of a miasma of overlapping events, and initiatives, and challenges. Amazon’s really decentralized; everything is happening at once. Bezos is close to some things, he was very close to Alexa. He is really distant from other things.

Andy Jassy for years had a lot of independence to run AWS. So, how do you tell that story, and then keep Bezos in the center? I mean, Andy Jassy and Jeff Wilke and everyone, I mean, those are great business people. Not necessarily dynamic personalities as, Corey, you know well, but people want to read about Jeff Bezos. He is a larger-than-life figure.

He’s a pioneer. He’s an innovator. He’s controversial. And so the challenge all along is to, kind of, keep him in the center. And so that’s just, like, a writing challenge. It’s a narrative challenge.

And the lucky thing is that Amazon does tend to orbit around Jeff Bezos’s brain. And so in all the storytelling, even the AWS bits of the book, which we can talk about, as an author, you can always bring Bezos back just by following the facts. You’ll eventually get, in the evolution of any story, to an S Team meeting, or to an acquisition discussion where Jeff had an impact, said something insightful, walked out of a meeting, raise the bar, had impossibly high standards. So, the last thing I’ll say is, because Amazon’s so decentralized, when you write these books you have to talk to a lot of people. And then you get all the pieces of the puzzle, and you start to assemble them, and the challenge as a writer is to, kind of, keep Bezos, your main character in the lens at all times, never let him drift too far out.

Corey: One of the things that I learned from it was just the way that Bezos apparently talks to his senior executives, as far as, “I will invest in this project, more than you might think I would.” I guess I’ve never really heard of a budget meeting talking about, “I”—in the first person—“Will invest.” Like, that is what happens, but for some reason the business books never put it quite that starkly or frame it quite that way. But in hindsight, it made a lot of things of my own understanding of Amazon fall into place. That makes sense.

Brad: He’s got a lot of levers, ways in which he’ll back a new initiative or express his support. And one of them is simply how he spends his time. So, with Alexa in the early years, he would meet once or twice a week with that team. But another lever is just the amount of investment. And oftentimes teams will come to him—the India team is a great example—they’ll come to the S Team with a budget, and they’ll list out their priorities and their goals for the coming year, and he’ll say, “You know, you’re thinking about this all wrong. Don’t constrain yourself. Tell us what the goals are, tell us what the opportunity is, then we’ll figure out how much it costs.”

And his mindset is like you can kind of break up opportunity into two categories: one are the land grabs, the big immediate opportunities where he will go all out, and India was a great example of that, I think the failed fire phone was another example, Prime Video, he doesn’t cap the investment, he wants to win. And then there are the more greenfield opportunities that he thinks he can go slower on and groceries for a long time was in that category. And there the budgets might be more constrained. The other example is the much older businesses, just like the retail business. That’s 20 years old—I have a chapter about that—and the advertising business, and he recognized that the retail business wasn’t profitable and it was depending on advertising as a crutch, and he blew it up because he thinks that those older divisions shouldn’t require investment; they should be able to stand on their own.

Corey: One quote you had as well, that just really resonated with me, as far as basically my entire ethos of how I make fun of Amazon is—and I’m going to read the excerpt here. My apologies. You have to listen to your own words being read back toward you—

Brad: [laugh].

Corey: These were typically Amazonian names: geeky, obscure, and endlessly debated inside AWS since—according to an early AWS exec—Bezos had once mused, “You know, the name is about 3% of what matters, but sometimes 3% is the difference between winning and losing.” And I just want to call that out because I don’t think I’ve ever seen an AWS exec ever admit that names might be even 3% worth of important. Looking at how terrible some of their service names are, I would say that 3% might be an aspirational target for their worldview.

Brad: [laugh]. Let me throw this back at you, Corey. Have you ever figured out why certain AWS services are Amazon and why others are AWS?

Corey: I did. I got to sit down—in the before times—with then the VP of Global Marketing, Ariel Kelman—who’s now Oracle’s Chief Marketing Officer—and Jeff Barr. And the direction that they took that in was that if you could use an AWS service without getting into the AWS weeds of a bunch of other services, then it was called Amazon whatever. Amazon S3, for example, as a primitive service doesn’t need a bunch of other AWS services hooked into it, so that gets the Amazon moniker. Whereas if you’re dealing with a service that requires the integration of a whole bunch of AWS in the weeds stuff—

Brad: Mmm, right.

Corey: —then it’s AWS. For example, AWS Systems Manager is useless without a whole bunch of other Amazon services. And they say they don’t get it perfectly right all the time, but that is the direction that it’s gone in. And for better or worse, I still have to look a lot of them up myself because I don’t care nearly as much as their branding people do.

Brad: Right. Well, I’ll tell you in the chapter about AWS, that quote comes up when the team is contemplating the names of the databases. And they do go into long debates, and I remember talking to Charlie Bell about the search for Redshift, and they go back and forth on it, and the funny thing about that one was, of course, Oracle interpreted it as a competitive slight. Its corporate color, I guess, being red, which he intended it more as a physics term. But yeah, when they were launching Aurora and Redshift, they contemplated those names quite a bit. And I don’t know if it’s 3%. I don’t know if it does matter, but certainly, those services have become really important to a lot of businesses.

Corey: Oh, yeah. And once you name something, it’s really hard to rename it. And AWS does view for—better or worse—APIs as a promise, so when you build something and presented a certain way, they’re never going to turn it off. Our grandkids are going to have to deal with some of these decisions once they get into computers. That’s a problem.

And I understand the ethos behind it, but again, it’s easy to make fun of names; it’s an accessible thing because let’s be very real here, a lot of what AWS does is incredibly inaccessible to people who don’t live in this space. But naming is something that everyone can get behind making fun of.

Brad: Absolutely. Yep. And [laugh] it’s perhaps why they spend a lot of time on it because they know that this is going to be the shingle that they hang out to the world. I don’t know that they’re anticipating your ridicule, but it’s obviously key to the marketing process for them.

Corey: Some of the more aware ones do. But that’s a different topic for a different time. One question I have for you that I wrestle with myself is I’ve been spending the last four years or so basically studying AWS all the time. And there’s a lot of things they get right; there’s a lot of things that they get wrong. But for better or worse, it’s very difficult not to come away from an in-depth study with an appreciation for an awful lot of the things that they do. At least for me.

I’m not saying that I fall in love with the company and will excuse them their wrongs; I absolutely do not do that. But it is hard, bordering on impossible for me, to not come away with a deep respect for a lot of the things that they do and clearly believe. How do you feel about that? Looking at Amazon, do you come away with this with, “Ooh. Remind me to never to become a Prime member and get rid of everything with an Amazon logo in my house,” versus the you’re about to wind up wondering if they can hire you for some esoteric role? Where do you fall on that spectrum?

Brad: I think I’m probably with you. I come away with an admiration. And look, I mean, let me say upfront, I am a Prime member. I have a Alexas in my home, probably more than my wife and kids are comfortable with. We watch Prime Video, we have Prime Video.

We order from Amazon all the time, we ordered from Whole Foods. I’m an Amazon customer, and so part of my appreciation comes from, like all other customers, the fact that Amazon uniquely restores time to our lives rather than extracts it. I wouldn’t say that about the social networks, right? You know, those can be time-wasters. Amazon’s a great efficiency machine.

But in terms of my journalism, you know, now two books and this big in-depth study in Amazon Unbound, and you have to admire what they have built. I mean, a historic American institution that has not only changed our economic reality, in ways good and bad but over the last year and a half, in the pandemic was among the few institutions that functioned properly and served as a kind of lifeline. And there is a critique in Amazon Unbound and we can talk about it, but it’s hard to come away—I think you said it well—it’s hard to come away after studying this company and studying the top executives, and how Jeff Bezos, thinks and how he has conceived products without real admiration for what they have built over the last 25 years.

Corey: Well, let’s get into your critique of Amazon. What do you think is, from what you’ve seen with all of the years of research you put into this company, what’s the worst thing about them?

Brad: Well, that’s a good way to put it, Corey. [laugh]. Let me—

Corey: [laugh]. It’s like, talk about a target-rich opportunity. Like, “Oh, wow. It’s like my children. I can’t stand any of them. How in the world could I pick just one?” But give it a shot.

Brad: Right. Well, let me start this way, which is I often will listen to their critiques from Amazon critics—and I’m sure you might feel this way as well—and just think, like, “Do they get it?” They’ll argue that Amazon exercised its size and might to buy the companies that led to Alexa. As I write in the Alexa chapter, that’s not true at all. They bought a couple of small companies, and those executives were all horrified at what Amazon was trying to do, and then they made it work.

Or the critics will say, “Fifty percent or more of internet users start their product searches on Amazon. Amazon has lock-in.” That’s not true either. Lock-in on the internet is only as strong as a browser window that remains open. And you could always go find a competitor or search on a search engine.

So, I find at least some of the public criticism to be a little specious. And often, these are people that complained about Walmart for ten years. And now Amazon’s the big, bad boogeyman.

Corey: Oh, I still know people who refuse to do business with Walmart but buy a bunch of stuff from Amazon, and I’m looking at these things going, any complaints you have about Walmart are very difficult to avoid mapping to Amazon.

Brad: Here’s maybe the distillation of the critique that’s an Amazon Unbound. We make fun of Facebook for, “Move fast and break things.” And they broke things, including, potentially, our democracy. When you look at the creation of the Amazon Marketplace, Jeff wanted a leader who can answer the question, “How would you bring a million sellers into the Amazon Marketplace?” And what that tells you is he wanted to create a system, a self-service system, where you could funnel sellers the world over into the system and sell immediately.

And that happened, and a lot of those sellers, there was no friction, and many of them came from the Wild West of Chinese eCommerce. And you had—inevitably because there were no guardrails—you had fraud and counterfeit, and all sorts of lawsuits and damage. Amazon moved fast and broke things. And then subsequently tried to clean it up. And if you look at the emergence of the Amazon supply chain and the logistics division, the vans that now crawl our streets, or the semi-trailers on our highways, or the planes.

Amazon moved fast there, too. And the first innings of that game were all about hiring contractors, not employees, getting them on the road with a minimum of guidance. And people died. There were accidents. You know, there weren’t just drivers flinging packages into our front yards, or going to the bathroom on somebody’s porch.

That happened, but there were also accidents and costs. And so I think some of the critique is that Amazon, despite its profession that it focuses only on customers, is also very competitor-aware and competitor-driven, and they move fast, often to kind of get ahead of competitors, and they build the systems and they’re often self-service systems, and they avoid employment where it’s possible, and the result have been costs to society, the cost of moving quickly. And then on the back-end when there are lawsuits, Amazon attempts to either evade responsibility or settle cases, and then hide those from the public. And I think that is at the heart of what I show in a couple of ways in Amazon Unbound. And it’s not just Amazon; it’s very typical right now of corporate America and particularly tech companies.

And part of it is the state of the laws and regulations that allow the companies to get away with it, and really restrict the rights of plaintiffs, of people who are wronged from extracting significant penalties from these companies and really changing their behavior.

Corey: Which makes perfect sense. I have the luxury of not having to think about that by having a mental division and hopefully one day a real division between AWS and Amazon’s retail arm. For me at least, the thing I always had an issue with was their treatment of staff in many respects. It is well known that in the FAANG constellation of tech companies, Facebook, Amazon, Apple, Netflix, and Google, apparently, it’s an acronym and it’s cutesy. People in tech think they’re funny.

But the problem is that Amazon’s compensation is significantly below that. One thing I loved in your book was that you do a breakdown of how those base salaries work, how most of it is stock-based and with a back-loaded vesting and the rest, and looking through the somewhat lengthy excerpt—but I will not read your own words to you this time—it more or less completely confirms what I said in my exposé of this, which means if we’re wrong, we’re both wrong. And we’ve—and people have been very convincing and very unified across the board. We’re clearly not wrong. It’s nice to at least get external confirmation of some of the things that I stumble over.

Brad: But I think this is all part of the same thing. What I described as the move fast and break things mentality, often in a race with competition, and your issues about the quality, the tenor of work, and the compensation schemes, I think maybe and this might have been a more elegant answer to your question, we can wrap it all up under the mantle of empathy. And I think it probably starts with the founder and soon-to-be-former CEO. And look, I mean, an epic business figure, a builder, an inventor, but when you lay out the hierarchy of qualities, and attributes, and strengths, maybe empathy with the plight of others wasn’t near the top. And when it comes to the treatment of the workforce, and the white-collar employees, and the compensation schemes, and how they’re very specifically designed to make people uncomfortable, to keep them running fast, to churn them out if they don’t cut it, and the same thing in the workforce, and then the big-scale systems and marketplace and logistics—look, maybe empathy is a drag, and not having it can be a business accelerant, and I think that’s what we’re talking about, right?

That some of these systems seem a little inhumane, and maybe to their credit, when Amazon recognizes that—or when Jeff has recognized it00, he’s course-corrected a little bit. But I think it’s all part of that same bundle. And maybe perversely, it’s one of the reasons why Amazon has succeeded so much.

Corey: I think that it’s hard to argue against the idea of culture flowing from the top. And every anecdote I’ve ever heard about Jeff Bezos, never having met the man myself, is always filtered through someone else; in many cases, you. But there are a lot of anecdotes from folks inside Amazon, folks outside Amazon, et cetera, and I think that no one could make a serious argument that he is not fearsomely intelligent, penetratingly insightful, and profoundly gifted in a whole bunch of different ways. People like to say, “Well, he started Amazon with several $100,000 and loan from his parents, so he’s not really in any ways a self-made anything.” Well, no one is self-made. Let’s be very clear on that.

But getting a few $100,000 to invest in a business, especially these days, is not that high of a stumbling block for an awful lot of folks similarly situated. He has had outsized success based upon where he started and where he wound up ending now. But not a single story that I’ve ever heard about him makes me think, yeah, that’s the kind of guy I want to be friends with. That’s the kind of guy I want to invite to a backyard barbecue and hang out with, and trade stories about our respective kids, and just basically have a social conversation with. Even a business conversation doesn’t feel like it would be particularly warm or compelling.

It would be educational, don’t get me wrong, but he doesn’t strike me as someone who really understands empathy in any meaningful sense. I’m sure he has those aspects to him. I’m sure he has a warm, wonderful relationship with his kids, presumably because they still speak to him, but none of that ever leaks through into his public or corporate persona.

Brad: Mmm, partially agree, partially disagree. I mean, certainly maybe the warmth you’re right on, but this is someone who’s incredibly charismatic, who is incredibly smart, who thinks really deeply about the future, and has intense personal opinions about current events. And getting a beer with him—which I have not done—with sound fantastic. Kicking back at the fireplace at his ranch in Texas, [laugh] to me, I’m sure it’s tremendously entertaining to talk to him. But when it comes to folks like us, Corey, I have a feeling it’s not going to happen, whether you want to or not.

He’s also incredibly guarded around the jackals of the media, so perhaps it doesn’t make a difference one way or another. But, yeah, you’re right. I mean, he’s all business at work. And it is interesting that the turnover in the executive ranks, even among the veterans right now, is pretty high. And I don’t know, I mean, I think Amazon goes through people in a way, maybe a little less on the AWS side. You would know that better than me. But—

Corey: Yes and no. There’s been some turnover there that you can also pretty easily write down to internal political drama—for lack of a better term—palace intrigue. For lack of a better term. When, for example, Adam Selipsky is going to be the new CEO of AWS as Andy Jesse ascends to be the CEO of all Amazon—the everything CEO as it were. And that has absolutely got to have rubbed some people in unpleasant ways.

Let’s be realistic here about what this shows: he quit AWS to go be the CEO of Tableau, and now he’s coming back to run AWS. Clearly, the way to get ahead there is to quit. And that might not be the message they’re intending to send, but that’s something that people can look at and take away, that leaving a company doesn’t mean you can’t boomerang and go back there at a higher level in the future.

Brad: Right.

Corey: And that might be what people are waking up to because it used to be a culture of once you’re out, you’re out. Clearly not the case anymore. They were passed over for a promotion they wanted, “Well, okay, I’m going to go talk to another company. Oh, my God, they’re paying people in yachts.” And it becomes, at some level, time for something new.

I don’t begrudge people who decide to stay; I don’t begrudge people who decide to leave, but one of my big thrusts for a long time has been understand the trade-offs of either one of those decisions and what the other side looks like so you go into it with your eyes open. And I feel like, on some level, a lot of folks there didn’t necessarily feel that they could have their eyes open in the way that they can now.

Brad: Mm-hm. Interesting. Yeah. Selipsky coming back, I never thought about that, sends a strong message. And Amazon wants builders, and operators, and entrepreneurial thinking at the top and in the S Team. And the fact that Andy had a experienced leadership team at AWS and then went outside it for the CEO could be interpreted as pretty demotivating for that team. Now, they’ve all worked with Adam before, and I’ve met him and he seems like a great guy so maybe there are no hard feelings, but—

Corey: I never have. He left a few months before I started this place. So, it—I get the sense that he knew I was coming and said, “Well, better get out of here. This isn’t going to go well at all.”

Brad: [laugh]. I actually went to interview him for this book, and I sat in his office at Tableau thinking, “Okay, here’s a former AWS guy,” and I got to tell you, he was really on script and didn’t say anything bad, and I thought, “Okay, well, that wasn’t the best use of my time.” He was great to meet, and it was an interesting conversation, but the goss he did not deliver. And so when I saw that he got this job, I thought, well, he’s smart. He smartly didn’t burn any bridges, at least with me.

Corey: This episode is sponsored in part by our friends at ChaosSearch., you could run Elasticsearch or Elastic Cloud—or OpenSearch as they’re calling it now—or a self-hosted ELK stack. But why? ChaosSearch gives you the same API you’ve come to know and tolerate, along with unlimited data retention and no data movement. Just throw your data into S3 and proceed from there as you would expect. This is great for IT operations folks, for app performance monitoring, cybersecurity. If you’re using Elasticsearch, consider not running Elasticsearch. They’re also available now in the AWS marketplace if you’d prefer not to go direct and have half of whatever you pay them count towards your EDB commitment. Discover what companies like HubSpot, Klarna, Equifax, Armor Security, and Blackboard already have. To learn more, visit chaossearch.io and tell them I sent you just so you can see them facepalm, yet again.

Corey: No. And it’s pretty clear that you don’t get to rise to those levels without being incredibly disciplined with respect to message. I don’t pity Andy Jesse’s new job wherein a key portion of the job description is going to be testifying before Congress. Without going into details, I’ve been in situations where I’ve gotten to ask him questions before in a real-time Q&A environment, and my real question hidden behind the question was, “How long can I knock him off of his prepared talking points?” Because I—

Brad: Good luck. [laugh].

Corey: Yeah. I got the answer: about two and a half seconds, which honestly was a little bit longer than I thought I would get. But yeah, incredibly disciplined and incredibly insightful, penetrating answers, but they always go right back to talking points. And that’s what you have to do at that level. I’ve heard stories—it may have been from your book—that Andy and Adam were both still friendly after Adam’s departure, they would still hang out socially and clearly, relationships are still being maintained, if oh, by the way, you’re going to be my successor. It’s kind of neat. I’m curious to see how this plays out once that transition goes into effect.

Brad: Yeah, it’ll be interesting. And then also, Andy’s grand homecoming to the other parts of the business. He started in the retail organization. He was Jeff’s shadow. He ran the marketing department at very early Amazon.

He’s been in all those meetings over the years, but he’s also been very focused on AWS. So, I would imagine there’s a learning curve as he gets back into the details of the other 75% of Amazon.

Corey: It turns out that part of the business has likely changed in the last 15 years, just a smidgen when every person you knew over there is now 10,000 people. There was an anecdote in your book that early on in those days, Andy Jesse was almost let go as part of a layoff or a restructuring, and Jeff Bezos personally saved his job. How solid is that?

Brad: Oh, that is solid. An S Team member told me that, who was Andy’s boss at the time. And the story was, in the late 90s—I hope I remember this right—there was a purge of the marketing department. Jeff always thought that marketing—in the early days marketing was purely satisfying customers, so why do we need all these people? And there was a purge of the marketing department back when Amazon was trying to right-size the ship and get profitable and survive the dotcom bust.

And Jeff intervened in the layoffs and said, “Not Andy. He’s one of the most—yeah, highest ceiling folks we have.” And he made him his first full-time shadow. Oh, and that comes right from an S Team member. I won’t say the name because I can’t remember if that was on or off the record.

But yeah, it was super interesting. You know what? I’ve always wondered how good of a identifier of talent and character is Bezos. And he has some weaknesses there. I mean, obviously, in his personal life, he certainly didn’t identify Lauren Sánchez’s brother as the threat that he became.

You know, I tell the story in the book of the horrific story of the CEO of Amazon Mexico, who Jeff interviewed, and they hired and then later ended up what appears to be hiring an assassin to kill his wife. I tell the story in the book. It’s a horrible story. So, not to lay that at the feet of Jeff Bezos, of course, but he often I think, moves quickly. And I actually have a quote from a friend of his in the book saying, “It’s better to not be kind of paranoid, and the”—sort of—I can’t remember what the quote is.

It’s to trust people rather than be paranoid about everyone. And if you trust someone wrongly, then you of course-correct. With Andy, though, he somehow had an intuitive sense that this guy was very high potential, and that’s pretty impressive.

Corey: You’re never going to bet a thousand. There’s always going to be people that slip through the cracks. But learning who these people are and getting different angles on them is always interesting. Every once in a while—and maybe I’m completely wrong on this, but never having spent time one on one with Andy Jassy, I have to rely on other folks and different anecdotes, most of them, I can’t disclose the source of, but every time that I wind up hearing about these stories, and maybe I’m projecting here, but there are aspects of him where it seems like there is a genuinely nice person in there who is worried, on some level, that people are going to find out that he’s a nice person.

Brad: [laugh]. I think he is. He’s extraordinarily nice. He seems like a regular guy, and what’s sort of impressive is that obviously he’s extraordinarily wealthy now, and unlike, let’s say Bezos, who’s obviously much more wealthy, but who, who really has leaned into that lifestyle, my sense is Andy does not. He’s still—I don’t know if he’s on the corporate jet yet, but at least until recently he wasn’t, and he presents humbly. I don’t know if he’s still getting as close from wherever, [unintelligible 00:32:50] or Nordstroms.

Corey: He might be, but it is clear that he’s having them tailored because fit is something—I spent a lot of time in better years focusing on sartorial attention, and wherever he’s sourcing them from aside, they fit well.

Brad: Okay, well, they didn’t always. Right?

Corey: No. He’s, he’s… there’s been a lot of changes over the past decade. He is either discovered a hidden wellspring of being one of the best naturally talented speakers on the planet, or he’s gone through some coaching to improve in those areas. Not that he was bad at the start, but now he’s compelling.

Brad: Okay. Well, now we’re talking about his clothes and his speaking style. But—

Corey: Let’s be very honest here. If he were a woman, we would have been talking about that as the beginning topic of this. It’s on some level—

Brad: Or we wouldn’t have because we’d know it’s improper these days.

Corey: We would like to hope. But I am absolutely willing to turn it back around.

Brad: [laugh]. Anyway.

Corey: So, I’m curious, going back a little bit to criticisms here, Amazon has been criticized roundly by regulators and Congress and the rest—folks on both sides of the aisle—for a variety of things. What do you see is being the fair criticisms versus the unfair criticisms?

Brad: Well, I mean, I think we covered some of the unfair ones. But there’s one criticism that Amazon uses AWS to subsidize other parts of the business. I don’t know how you feel about that, but until recently at least, my reading of the balance sheet was that the enormous profits of AWS were primarily going to buy more AWS. They were investing in capital assets and building more data centers.

Corey: Via a series of capital leases because cash flow is king in how they drive those things there. Oh, yeah.

Brad: Right. Yeah. You know, and I illustrate in the book how when it did become apparent that retail was leaning on advertising, Jeff didn’t accept that. He wanted retail to stand on its own, and it led to some layoffs and fiercer negotiations with brands, higher fees for sellers. Advertising is the free cash flow that goes in Prime movies, and TV shows, and Alexa, and stuff we probably don’t know about.

So, this idea that Amazon is sort of improperly funneling money between the divisions to undercut competitors on price, I think we could put that in the unfair bucket. In the fair bucket, those are the things that we can all look at and just go, “Okay, that feels a little wrong.” So, for an example, the private brand strategy. Now, of course, every supermarket and drugstore is going to line their shelves with store brands. But when you go to an Amazon search results page these days, and they are pockmarked with Amazon brands, and Whole Foods brands, and then sponsored listings, the pay-to-play highest bidder wins.

And then we now know that, at least for a couple of years, Amazon managers, private label managers were kind of peeking at the third-party data to figure out what was selling and what they should introduce is a private Amazon brand. It just feels a little creepy that Amazon as the everything store is so different than your normal Costco or your drugstore. The shelves are endless; Amazon has the data, access to the data, and the way that they’re parlaying their valuable real estate and the data at their disposal to figure out what to launch, it just feels a little wrong. And it’s a small part of their business, but I think it’s one where they’re vulnerable. The other thing is, in the book, I tried to figure out how can I take the gauge of third-party sellers?

There’s so many disgruntled voices, but do they really speak for everyone? And so instead of going to the enemies, I went to every third-party seller that had been mentioned in Jeff Bezos’s shareholder letters over the past decade. And these were the allies. These were the success stories that Bezos was touting in his sacrosanct investor letter, and almost to a one, they had all become disgruntled. And so the way in which the rules of the marketplace change, the way that the fees go up, and the difficulty that sellers often have in getting a person or a guiding hand at Amazon to help them with those changes, that kind of feels wrong.

And I think that maybe that’s not a source of regulation, but it could be a source of disruptive competition. If somebody can figure out how to create a marketplace that caters to sellers a little better with lower fees, then they could do to Amazon with Amazon years ago did to eBay. And considering that Marketplace is now a preponderance of sales more than even retail on amazon.com, that can end up hurting the company.

Corey: Yeah, at some point, you need to continue growing things, and you’ve run out of genuinely helpful ways, and in turn in start to have to modify customer behavior in order to continue doing things, or expand into brand new markets. We saw the AWS bleeding over into Alexa as an example of that. And I think there’s a lot of interesting things still to come in spaces like that. It’s interesting watching how the Alexa ecosystem has evolved. There’s still some very basic usability bugs that drive me nuts, but at the same token, it’s not something that I think we’re going to see radically changing the world the next five years. It feels like a hobby, but also a lucrative one, and keeps people continuing to feed into the Amazon ecosystem. Do you see that playing out differently?

Brad: Wait, with Alexa?

Brad: Absolutely.

Brad: Yeah. I agree with you. I mean, it feels like there was more promise in the early years, and that maybe they’ve hit a little bit of a wall in terms of the AI and the natural language understanding. It feels like the ecosystem that they tried to build, the app store-like ecosystem of third-party skills makers, that hasn’t crystallized in the way we hoped—in the way they hoped. And then some of these new devices like the glasses or the wristband that have Alexa feel, just, strange, right?

Like, I’m not putting Alexa on my face. And those haven’t done as well. And so yeah, I think they pioneered a category: Alexa plays music and answers basic queries really well, and yet it hasn’t quite been conversational in the way that I think Jeff Bezos had hoped in the early days. I don’t know if it’s a profitable business now. I mean, they make a lot of money on the hardware, but the team is huge.

I think it was, like, 10,000 people the last I checked. And the R&D costs are quite large. And they’re continuing to try to improve the AI, so I think Jeff Bezos talks about the seeds, and then the main businesses, and I don’t think Alexa has graduated yet. I think there’s still a little bit of a question mark.

Corey: It’s one of those things that we remain to see. One last thing that I wanted to highlight and thank you for, was that when you wrote the original book, The Everything Store, Andy Jassy wrote a one-star review. It went into some depth about all the things that, from his perspective, you got wrong, were unfair about, et cetera.

And that can be played off as a lot of different things, but you can almost set that aside for a minute and look at it as the really only time in recent memory that Andy Jassy has sat down and written something, clearly himself, and then posted it publicly. He writes a lot—Amazon has a writing culture—but they don’t sign their six-pagers. It’s very difficult to figure out where one person starts and one person stops. This shows that he is a gifted writer in many respects, and I don’t think we have another writing sample from him to compare it to.

Brad: So, Corey, you’re saying I should be honored by his one-star review of The Everything Store?

Corey: Oh, absolutely.

Brad: [laugh].

Corey: He, he just ignores me. You actually got a response.

Brad: I got a response. Well.

Corey: And we’ll put a link to that review in the [show notes 00:40:10] because of course we will.

Brad: Yes, thank you. Do you—remember, other Amazon executives also left one-star reviews. And Jeff’s wife, and now ex-wife Mackenzie left a one-star review. And it was a part of a, I think a little bit of a reflexive reaction and campaign that Jeff himself orchestrated at my—this was understanding now, in retrospect. After the book came out, he didn’t like it.

He didn’t like aspects of his family life that were represented in the book, and he asked members of the S Team to leave bad reviews. And not all of them did, and Andy did. So, you wonder why he’s CEO now. No, I’m kidding about that. But you know what?

It ended up, kind of perversely, even though that was uncomfortable in the moment, ended up being good for the first book. And I’ve seen Andy subsequently, and no hard feelings. I don’t quite remember what his review said. Didn’t it, strangely, like, quote a movie or something like that?

Corey: I recall that it did. It went in a bunch of different directions, and at the end—it ended with, “Well, maybe someday he’ll write the actual story. And I’m not trying to bait anyone into doing it, but this book isn’t it.” Well, in the absence of factual corrections, that’s what we go with. That is also a very Amazonian thing. They don’t tell their own story, but they’re super quick to correct the record—

Brad: Yeah.

Corey: —after someone says a thing.

Brad: But I don’t recall him making many specific claims of anything I got wrong. But why don’t we hope that there’s a sequel review for Amazon Unbound? I will look forward to that from Andy.

Corey: I absolutely hope so. It’s one of those things that we just really, I guess, hope goes in a positive direction. Now, I will say I don’t try to do any reviews that are all positive. And that’s true. There’s one thing that you wrote that I vehemently disagree with.

Brad: Okay, let’s hear it.

Corey: Former Distinguished Engineer and VP at AWS, Tim Bray, who resigned on conscientious objector grounds, more or less, has been a guest on the show, and I have to say, you did him dirty. You described him—

Brad: How did I—what did I do? Mm-hm.

Corey: Oh, I quote, “Bray, a fedora-wearing software developer”—which is true, but still is evocative in an unpleasant way—“And one of the creators of the influential web programming language, XML”—which is true, but talk about bringing up someone’s demons to haunt them. Oh, my starts.

Brad: [laugh]. But wait. How is the fedora-wearing pejorative?

Corey: Oh, it has a whole implication series of, and entire subculture of the gamer types, people who are misogynist, et cetera. It winds up being an unfair characterization—

Brad: But he does wear a fedora.

Corey: He does. And he can pull it off. He has also mentioned that he is well into retirement age, and it was a different era when he wore one. But that’s not something that people often will associate with him. It’s—

Brad: I’m so naive. You’re referring to things that I do not understand what the implication was that I made. But—

Corey: Oh, spend more time with the children of Reddit. You’ll catch on quickly.

Brad: [laugh]. I try, I try not to do that. But thank you, Corey.

Corey: Of course. So, thank you so much for taking the time to go through what you’ve written. I’m looking forward to seeing the reaction once the book is published widely. Where can people buy it? There’s an easy answer, of course, of Amazon itself, but is there somewhere you would prefer them to shop?

Brad: Well, everyone can make their own decisions. I flattered if anyone decides to pick up the book. But of course, there is always their independent bookstore. On sale now.

Corey: Excellent. And we will, of course, throw a link to the book in the [show notes 00:43:31]. Brad, thank you so much for taking the time to speak with me. I really appreciate it.

Brad: Corey, it’s been a pleasure. Thank you.

Corey: Brad Stone, author of Amazon Unbound: Jeff Bezos and the Invention of a Global Empire, on sale now wherever fine books are sold—and crappy ones, too. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice and then a multi-paragraph, very long screed telling me exactly what I got wrong.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Wesley

Wesley Faulkner is a first-generation American. He is a founding member of the government transparency group Open Austin and ran for Austin City Council in 2016. His professional experience also includes work as a social media and community manager for the software company Atlassian, and various roles for the computer processor company AMD, Dell, and IBM. Wesley Faulkner serves as a board member for South by Southwest Interactive (SXSWi) and is a Developer Advocate for Daily.

Links:

  • Daily website: https://www.daily.co/
  • Twitter: https://twitter.com/wesley83

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: The Apps ON Cloud Summit, hosted by Turbonomic, is a new action-packed not-a-conference happening online May 11th through 13th. It’s for everyone who makes applications in the cloud run, from IT leaders to DevOps pros to you folks. Take a break from screaming into the cloudy void to learn from some of the best, like Kelsey Hightower, AWS Blogger Jon Myer, and yours truly.

Register now at turbonomic.com/screaming. There’s a swag box ready to ship for the first two thousand registrants – don’t miss it!

Corey: Let’s be honest—the past year has been a nightmare for cloud financial management.

The pandemic forced us to move workloads to the cloud sooner than anticipated, and we all know what that means—surprises on the cloud bill and headaches for anyone trying to figure out what caused them.

The CloudLIVE 2021 virtual conference is your chance to connect with FinOps and cloud financial management practitioners and get a behind-the-scenes look into proven strategies that have helped organizations like yours adapt to the realities of the past year.

Hosted by CloudHealth by VMware on May 20th, the CloudLIVE 2021 conference will be 100% virtual and 100% free to attend, so you have no excuses for missing out on this opportunity to connect with the cloud management community. Visit cloudlive.com/corey to learn more and save your virtual seat today. That’s cloud-l-i-v-e.com/corey to register.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Wesley Faulkner, who's a developer advocate at a company called Daily. Either the company's named Daily or there's a misconfigured cron job somewhere. Wesley, welcome to the show. Which is it?

Wesley: It is Daily and you can find them at daily.co. Because of the name, it's kind of hard to Google.

Corey: Yeah. It seems like the more interesting a company is, the harder it is to Google. When it's some common word where they had to cheap out due to a shortfall in investment round and they can't afford to, you know, buy a vowel, it becomes super EC2 Google once Google understands there's not an autocorrect story in there. But it seems like the companies that are really poised for success in some markets are incredibly difficult to Google for Puppet and Chef were some of the previous generation stories. What does Daily do?

Wesley: Daily makes a video API built off WebRTC. And today, as we record this, WebRTC went 1.0 officially. And what we do is allow applications to break out video and audio so they can use it and configure it for their applications any way they want. So, you can do like a Zoom competitor, or you can just take the audio feed and make a Clubhouse competitor.

Or you can make a new way of collaborating with those spatial-audio-virtual work environments. There's a lot of different ways that you can use video to make you more productive, and we allow you to do that with our API.

Corey: It seems like there's a lot of emphasis on video these days, given that as we record this, we're still in a pandemic. I keep hoping that one day someone's going to listen to this and “Oh, yeah. The pandemic. I remember those days,” but so far, no luck. Last year, it always felt like “Well, oh, wow, this must have been recorded a while ago. They weren't even talking about the meteor yet.”

So, it just becomes an escalating series of awful, in a bunch of different ways. But video has been something that's been on a lot of people's mind, usually when it breaks. Is this designed to be embedded in other applications? Is it designed to be a standalone application? Is someone going to have a Daily desktop app at some point?

Wesley: So, we empower companies to build on top of our platform to make their own products. There is a possibility that we might make our own branded plugins or proof of concepts that are out there for people to use, but what we really want to focus on is enabling the next video experience and the configuration is where the innovation happens in terms of how big a bubble is or when it shows up. We just want to make sure that we take care of the table stakes of video and make everything, in terms of the stack, easier for people to implement.

Corey: So, one thing I noticed on your website that's always of interest to me is the magic word, ‘HIPAA,’ which I used to think was the female version of an animal. Turns out no, English doesn't work quite the same way as Spanish does. It's for those who are unaware, or perhaps not based in the US, it's the Health Information Portability and Privacy Act, which means you are certified to carry video for medical uses.

Wesley: Yes. And that we don't track PII or making sure that all the information is secure. And a little secret is that HIPAA, sometimes, is easier because you don't need to get hold logs or store that information. So, HIPAA is the Snapchat of database information, which is really good in some cases.

Corey: So, one of the things you're passionate and talking about is neurodiversity. So, why not? Let's have a conversation about that with a segue that even works, which is whenever I wind up having a video call with someone, I use Zoom or something like it, but when I call my psychiatrist for an appointment, we're going back in time to some ancient thing that they've embedded that performs the impossible and makes WebEx look good. But it feels awful, and whenever I've dug into why is this so bad, HIPPA’s always the answer. It feels like it's stuck ten years in the past, and it's annoying and challenging to work with.

You're a relatively new company from what I can tell. You’re version 1.0.6 on the website, which unless you have a glacially slow release process, means you haven't been around for decades, yet you've got something that works performantly, and seems to be aligned with modern technology. How’d that happen?

Wesley: Actually, in the world of WebRTC, Daily is kind of ahead of the curve. I mentioned that it hit 1.0 today, and so we've been doing WebRTC since the early teens, and even before then the video technology has been something our founders have dabbled in since their graduate years of college. And there's a lot of flow and churn in terms of what to use and how to use it. Back in the flash days, of course, you had to make sure that you both had the same version, and that everyone
had it installed, and that there were no incompatibilities with the network you're on.

And then there is a lot of ebb and flow in terms of, is it going to be good for this application or that application? Since we've been building on top of WebRTC in the early days, we've been able to stay basically cutting edge. And so the problems we're tackling is kind of the NASA of video. I think that if you try to use some of the older technologies, you're kind of stuck into being comfortable and just making it work, and we're just trying to enable all of these different scenarios that you couldn't before because we are trying to solve the problems no one else is looking at.

Corey: That's the problem. It always feels like on some level that—at least in the United States—that technology is going to be dictated by the insurance company and in the medical group, and by the time it gets to the doctor, I mean, ideally, they're doctors and focusing on the whole health aspect, not spending their nights and weekends messing around with various video codecs or whatnot. So, it always felt like an insurmountable problem to me. But honestly, this becomes a differentiator.

Wesley: Yeah. I mean, there's a reason why we exist is because there's a need that's not being met. So, if you need one of those off the shelves, just do a quick video chat, there are plenty of solutions out there. And even with WebRTC, you can probably build your own. So, if you were looking for anything that is somewhat complicated, needs a scale, you need to worry about bandwidth; if you need to worry about having more than 200 300 people on a call, that's something that it's really hard to do on your own.

Corey: Again, talking from a speaking to my psychiatrist perspective, that sounds like hell on earth. It's “Great. So, what am I doing now? I'm going to livestream my mental health sessions. That's going to go super well for everyone.” There's vulnerable and then there's just plain dumb.

Wesley: There's reality TV, which is also, if you need the—

Corey: Oh, my God, you're right. That's exactly what that is.

Wesley: [laugh]. I would pay to watch it.

Corey: So, you've had a fascinating career. Like I normally don't read through people's bios on the air, but I’ll make an exception for you. You're a first-generation American, which is awesome; you're a founding member of the government transparency group Open Austin; you ran for the Austin City Council in 2016—that phrasing tells me you did not win?

Wesley: Yes, that's correct.

Corey: You also worked as a social media and community manager for Atlassian which means at one point you were the poor shmoo who was on the other side of my Jira barbs. Sorry about that. I always hate my past comes back to haunt me.

Wesley: I love hearing the pain of others. I feed off of it.

Corey: Yes. And then you worked at a number of companies: AMD, Dell, and IBM, which were awesome, fascinating places in, you know, 1998, but I don't know that that's when you were there. But that's my prejudices, not yours.

Wesley: Yes. Every company seems to have their little nook of being big corporate company because they’re a big corporate company. I try to make sure that I carve a little space for myself. Sometimes I'm successful, sometimes I’m not.

Corey: Yeah, just as a passion project or something, you’re also apparently a board member for SXSW Interactive?

Wesley: Yeah, been a board member, I think, for over ten years now. Yeah.

Corey: And I'm guessing you are not yet bored with it?

Wesley: It's a great experience. One, you get a free badge every year. And so it’s… didn't do me any good in 2020, but hopefully, when we start getting past this pandemic, we start meeting people face-to-face. It's a sea of people and I love surfing that whenever I can.

Corey: I spend a lot of time at various conferences—in the before times—and one thing I've never been quite clear on is
what the hell SXSW is. Is there basically a quick summary you can give us?

Wesley: SXSW, it started off as three main festivals: music, film, and then interactive. And interactive is the part that has changed throughout the years. Interactive used to be an encyclopedia on a CD ROM. [laugh]. It used to be video game technology.

And then “New media” started—quote-unquote—with social media, Facebook, Twitter. And those companies decided to launch some of their products and send some of their people to South By and it started attracting some of their same companies. And then, when the bubble started with the internet, with companies forming that are internet first, they also came to South By and they brought the VCs with them. And so it didn't just become a meet and greet, it was becoming a place to start and get funding for your company. And that whole crowd just basically made a beachhead with South By, and it grew because of that.

Corey: And it's turned into something that still becomes impossible to define. And I've basically been boycotting it because no one has ever invited me to speak there. And everyone looks at me super strangely when I say that and it's not that kind of thing. But I don't know, invitations work the same way, universally. So, I am insisting that my misunderstanding is something I'm imposing on the world. I'm really tech-bro-ing my way through it.

Wesley: Yeah. That’s the same reason why I don't go to TED or cons.

Corey: Exactly. If they want me there, they will have me on the floor giving a talk. It'll be amazing, and the 360 stage and all the amazing things that scare the heck out of people. But speaking of before times, you also describe yourself as a public speaker. What do you like talking about?

Wesley: I love talking about neurodiversity. I love also talking about diversity and inclusion, which is part of neurodiversity, but since it's one of those things that doesn't really get highlighted, I want to make sure that I break it out to give some focus. And I love talking about networking.

Corey: Well, tell me more about the neurodiversity piece. What exactly does that mean? Where does it start? Where does it stop? It's a term that not everyone is particularly familiar with. Until very recently, I was in that group.

Wesley: So, neurodiversity is the acknowledgment that brains work differently. And that a variety of functioning is considered normal. There have been names associated with different understandings and different processing like Asperger's or autism. I'm dyslexic, ADHD. And there are many different kinds that are there at birth, and then there are some that are inherited, like PTSD, or depression, or anxiety. Those are all different ways of processing and seeing the world or different situations.

Corey: I'm going to talk about something I haven't ever mentioned on this show because why the hell not? If I can't talk about it, who the hell can? When I was five years old, I was diagnosed with ADD which later became ADHD, and the more I talked to people who've been down this path—I mean, I've been on medication for three decades now, aside from a decade where I just white-knuckled it the whole way, and [laugh] that didn't work so well. But it seems that ADHD is very much a spectrum disorder, and everyone who has it experiences it very differently. And the universal constant though, whenever you talk to someone who has it is, they feel like they're a shitty person where they keep dropping the ball, they're a bad friend, they're bad at showing up on time. It's basically the single unifying theme, start to finish, always feels like it's beating yourself up for things that are fundamentally not really a choice.

Wesley: Yes. I would say from all the people that I know have ADHD, me included, that's definitely part of it, of feeling like you can't focus, which means that you're not paying attention to some of the details, but also ADHD, flip-flops between being hyper-focused on one thing or being focused on a lot of things, and it feels like the way that it's seen in the media, it's described differently. And I think that's part of the reason why I like talking about neurodiversity is to bring more awareness around this subject and for people to be able to understand that it's not necessarily someone's fault, but also to highlight some of the advantages of neurodiversity.

Corey: That's a good way of framing that. One of the things I did when I started this company was I built it around things that I was good at, and avoided things I was not. In my initial reports on AWS bills were two pages. And “Oh, yeah, do these five things. It will save you 20% off your bill. Have a nice day,” is what it distilled down to. And it was right, and it worked, but it turns out when you're charging people bespoke consulting project money, they kind of want something that doesn't fit in a tweet. Who knew.

So, as we expanded, I brought in a business partner who's my exact opposite, whose primary language is Excel. And slowly, we wound up building systems around me to at least mitigate an awful lot of my shortcomings in that respect. But I still can't shake the feeling that it's a lot of work people have to do that they wouldn't if I weren't this way. And that tends to disregard the fact that if I work this way, I would not be able to do the things that I do. It's a double-edged sword, and I think it's something I need to be a lot more public about than I have been. So, 2021, here we go.

Wesley: Yeah. I think if you look at entrepreneurs, people who are neurodiverse over-index because they don't necessarily fit in the systems that are considered standard. They problem-solve in different ways, which don't necessarily conform to the standard operating procedure of most established companies. So, if you—thinking that, like, “I had to break out and I had to do this,” what you're doing is you're doing it your way, and you're building a team around you, hiring people, give them a purpose, given them a function in their duty because of yourself and the gifts that you were given.

Corey: And that's sort of part of it. The fact that I can context switch very rapidly is helpful, the fact that I can't pay attention for super long means that I'm better at reading things extremely quickly. When you're trying to sort through everything coming out of AWS, that becomes an asset more than it is a hindrance. And in some ways, it's hard for me to remember and realize—I’m still learning almost every day—that there are different aspects and different manifestations of ADHD, and that in many cases, I've always just written them off for 40 years as “The way that I am.”

Wesley: Also, I mean, you probably can judge things fairly quickly. Like, “This is crap,” “This is put together poorly,” or “They didn't spend enough time actually getting to the root of the matter.” Because you get bored. You're like, “This is a waste of my time.” And that type of feedback comes quicker to you because you know what you're looking for, and you know it's not cutting it. And so when you pass on and you make people, quote-unquote “Adjust” to you, what you're doing is you're making their material more accessible for everyone, which is also really great.

Corey: This episode is sponsored in part by our friends at New Relic. If you’re like most environments, you probably have an incredibly complicated architecture, which means that monitoring it is going to take a dozen different tools. And then we get into the advanced stuff. We all have been there and know that pain, or will learn it shortly, and New Relic wants to change that. They’ve designed everything you need in one platform with pricing that’s simple and straightforward, and that means no more counting hosts. You also can get one user and a hundred gigabytes a month, totally free. To learn more, visit newrelic.com. Observability made simple.

Corey: I'm trying. Every episode of Screaming in the Cloud has a transcript that goes along with it, and that's partially because not everyone wants to listen and other people want to read. I don't have the attention span to sit and listen to these things, even at 2X. I want the transcript so I can read it. And if I have that desire, it's a near certainty that I'm not the only person doing it.

Incidentally, if you're listening to this and didn't realize that, yeah, if you prefer to read, go to screaminginthecloud.comand take a look at all of the episodes: there's a full transcript for everyone because accessibility is important.

Wesley: Not only just important, it helps everyone. Like for you, it probably helps with your SEO, it helps with when someone's looking for a certain subject or a certain person, that you could be a hit on Google. So, it has some knock-on effects that people don't realize, too.

Corey: One of the things that I find continually surprising is just how many people that I have a deep and abiding respect for who when something comes up that alludes to being neurodivergent, which is the opposite of neurotypical for those who are not familiar with the phrase, it seems oh that everyone I talked to and work with extensively has something going on that is a deviation from the norm. And for someone who grew up in my position viewing it as just a set of weaknesses and a shameful thing you never ever talk about. It's kind of liberating in a way to realize, oh my god, I'm not alone. There's a community of us. And in fact, we formed a community and didn't even realize this was a common, shared thing.

Wesley: Oh, absolutely. And I think you mentioned the shame. That shame I have, and I'm still battling. And speaking on neurodiversity is something I'm doing to try to get past that. But as a parent, I want to make the world a better place for everyone, including my kids, so that they can not only accept who they are, but the world can also accept them.

Corey: It's a fascinating thing. One other thing we sort of have in common that I'm very curious about that you alluded to a few minutes ago is, you talk about enjoying networking. And sure enough, I pull up a couple of social media sites, and we have a crapload of people in common on all of them, but this is the first time we've ever spoken other than through a few messages back and forth and writing. How is it that you have traveled so broadly in the same circles that I am in--first off—and secondly, how have we never met before now?

Wesley: Well, I mean, that's one thing about networking is if you find a passion and you do it, you are able to broaden your network fairly quickly, especially using the internet. But for me, in terms of the circles, I'm in a lot of different circles. You mentioned my background, I'm a first-generation American. I'm also a Black American. When that kind of means that I've never really, in some ways, found a place where I felt entirely comfortable, which means that I am generally uncomfortable at a certain level, no matter where I am, which makes the barrier of entry for these different networks, all even. [laugh].

So, what I try to do is just find good people and talk to people. And when you do that, you meet people who do different things. So, in social media space, I've met a lot of people, and in marketing, I've met a lot of people, and in developers and developer relations, I've met a lot of people. How we have never met is one of those things that is just I feel is inevitable. I hopefully will meet everyone that I can, who feels the need to talk to people who care about other humans.

Corey: One of the things that I found that was very aligned with my way of thinking about the world was a book called Never Eat Alone by Keith Ferrazzi, a consultant. And he talked about how it was important to meet people as much as you can, to introduce people every chance you get, and do favors for people without ever really expecting anything in return. And it sounds hokey, it sounds like something you want the people you're about to take advantage of to believe, except for the part where I've lived the last ten years like this, and it absolutely is true. It comes back around in super weird ways. If I help someone without thinking about how this is going to benefit me, I just do it.

And invariably, people reach out randomly at appropriate times, and, “Hey, I've got this thing going on. Can you help us with this?” And it becomes a business opportunity, or it's a chance to wind up meeting someone who's aligned with some research I'm doing. It's one of those things where you put something out, and it sort of comes back in some weird ways. And in fact, there were studies that mentioned in this book that the wealthy, it was not about giving money to each other that led to a lot of their success so much as it was they did favors for each other.

One of them was on the board of a private school, so they put their thumb on the scale and got someone else's kid in. Admittedly not a sympathetic example. But the point stands, when you help people out it comes back in super weird ways. What's your take on that?

Wesley: Well, there's two kinds of economies. There's the money economy, and there's, like, the friend economy, when you have a roommate, and they eat some of your food, you don't necessarily keep track of how much they've eaten and make sure they pay you back. If you both live together, and it kind of just works out. Eventually, if you're really good friends, I am of the philosophy that if you are doing transactional networking, meaning that you're doing something to get something back, you're in essence, making a short term decision, not a long term engagement with that person because you're usually trying to get something because of their position, power, or influence. So, if that person's position, power, or influence changes, or yourself, and what you do and what you need changes, you kind of lost that relationship and it's inherently short term.

But if you are really trying to connect with people you care about, and people you want to be what I call fully present and fully seen, then you won't hold things back like say, “Hey, that's kind of racist,” because you want to cultivate that relationship for that transaction. But if you're truly trying to be yourself and you're calling that person out and that doesn't end the relationship, it'll in fact strengthen that relationship, and it'll be a longer commitment than just that one transaction that you're hoping for. And it'll come back because once you're in someone's life, you become top-of-mind. I give a talk on how to get over the awkwardness of networking. And when you look at your recent called list, and then when you look at who you're connected to on LinkedIn, if you don't have the same feeling of fondness, when you look at those people, then you can realize that those are just transactions and you're not building a network of people you care about, just a network of people who can take advantage of.

Corey: The idea of viewing relationships as transactional leads to an awful lot of negative things, the ability to be able to have a conversation with folks, regardless of who they are, it's its own skill in some ways. I've got to admit, when I first started doing this podcast, I was incredibly intimidated by some of the guests that I've had. And I no longer have that reaction, for whatever reason. And I think a large part of it comes from, I had an intern from Facebook, a while back, and I had the EVP that runs all of Azure, I've had the CEO of Stack Overflow, I've had personally what I find to be the most impressive guest ever, Mai-Lan Tomsen Bukovec, who runs S3 and storage at AWS, there are fantastic people and many more, and if you treat all of these people the same way, as people, it really goes a long way. I don't have a database of these people appointed to reach out for when it's time for a favor.

It doesn't work that way. And I feel like that would be a way of cheapening it in some ways, and in many ways working directly against me. When someone reaches out with a nomination for someone who would be a guest on this show. My response is always “Great, what story will they tell?” Not “Well, what's their job title? How long have they been doing it? Are they an investor? What do the VCs think of them?” It's the wrong angle. And every time I see folks who go down that path, I am discouraged by how apparently this is not what everyone [unintelligible 00:27:02].

Wesley: Yeah. And then let's say you do have the person who's in this high position, and then they are just a total crap person, but you're like, “I'm going to make sure I get them on the show because they have influence.” And then they get arrested. Do you still have them on the show? Probably not.

Because they've lost that position and they've tarnished their reputation publicly, even though that they might have a tarnished reputation privately. But if you're going by your own gut, and your own honor system of the kind of people that you want to honor with their time and your time and give them a platform, you won't fall into that loop or that challenge because you'll be governed by a sense of elevating people who care about other people.

Corey: I've talked about this on Twitter from time to time, but I don't believe I've ever said it explicitly on the show, so I'm going to hijack this to say it right now. If you're listening to this, I have two requests for you: the first is, if I can help you with anything, please ping me. Worst case, I will tell you, “I don't know. But I probably know someone who can.” Secondly, it turns out when you run a podcast, everyone is super nice to you, which means that very often trash fire people don't present that way.

So, if you ever listening to this show, and you hear that I have a guest that you know is garbage, please tell me that. Confidentiality is guaranteed, but I don't want to wind up platforming folks who are frankly, terrible. I'm losing my connection to the back channel network in some ways, as a result. It's a weird thing to ask for, but I'm quite sincere with it.
Please, as a personal favor. And as always, if I can help you, please let me know.

Wesley: You already are.

Corey: I do my best wherever I can. Well, I want to thank you for taking the time to speak with me today. If people want to
learn more, where can they find you?

Wesley: Well, you can find me on Twitter. I'm @wesley83, on Twitter, and I am at Daily. So, go to daily.co, and if you hit the help, it'll get to me, if you don't want to go that route. Those are the two main ways. If you send me a message on LinkedIn, most likely I will not reply. I do treat LinkedIn as my own personal network of people that I care for and about, so if you're sending me a request there, at least to chat with me first on Twitter so we can get to know each other.

Corey: Absolutely. And we will of course put links to all of that in the [show notes 00:29:25]. Thanks so much for taking the time to speak with me. I really appreciate it.

Wesley: Thank you, Corey. I really appreciate being on your show.

Corey: Wesley Faulkner, developer advocate at Daily. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you've hated this entire podcast, please leave a five-star review on your podcast platform of choice, along with a comment telling me why you're the garbage person.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

Screaming in the Cloud with Corey Quinn features conversations with domain experts in the world of Cloud Computing. Topics discussed include AWS, GCP, Azure, Oracle Cloud, and the "why" behind how businesses are coming to think about the Cloud.

View Details

About Bobby

Bobby Allen serves as VP of Strategic Alliances for Turbonomic. Bobby is a veteran of Intel, Bank of America, TIAA and multiple startups including one that was successfully acquired by the former CSC (now DXC). He went into corporate America after being an Intel fellow at the University of Michigan (MS in Computer Science and Engineering) and a Meyerhoff scholar at UMBC (BS in Computer Science). Bobby has been involved in cloud computing startups since 2012.

He frequently advises CXO’s on cloud strategy and logical equivalents in cloud technology. His goal is to provide data-driven output to move decision-makers from information to clarity to insight. Bobby has been a featured speaker in various events and digital formats including VMworld, AWS re:Invent, theCube, crowdchat, The CTOAdvisor and Gigaom’s Voices in the cloud. He’s equally skilled talking to analysts or technical teams but most enjoys helping customers separate fact from fiction.

Bobby also serves as Stewardship Pastor of Wellspring Church – a Gospel centered, multi-ethnic community in Charlotte, NC. Bobby is a member of the preaching team at Wellspring and is responsible for technology, finances and facilities. He’s grateful to be part of the team that helped complete a multi-year building purchase and remodeling project. Wellspring moved into their new home in December 2019.

Links:

  • turbonomic.com: https://turbonomic.com
  • bobbyjallen.me: https://bobbyjallen.me
  • @ballen_clt: https://twitter.com/ballen_clt

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at The Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part byLaunchDarkly. Take a look at what it takes to get your code into production. I’m going to just guess that it’s awful because it’s always awful. No one loves their deployment process. What if launching new features didn’t require you to do a full-on code and possibly infrastructure deploy? What if you could test on a small subset of users and then roll it back immediately if results aren’t what you expect? LaunchDarkly does exactly this. To learn more, visitlaunchdarkly.com and tell them Corey sent you, and watch for the wince.

Corey: The Apps ON Cloud Summit, hosted by Turbonomic, is a new action-packed not-a-conference happening online May 11th through 13th. It’s for everyone who makes applications in the cloud run, from IT leaders to DevOps pros to you folks. Take a break from screaming into the cloudy void to learn from some of the best, like Kelsey Hightower, AWS Blogger Jon Myer, and yours truly. Register now at turbonomic.com/screaming. There’s a swag box ready to ship for the first two thousand registrants – don’t miss it!

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. My promoted guest today works for Turbonomic. Now, Bobby Allen is the VP of Strategic Alliances, but on LinkedIn he’s something else entirely: A cloud therapist. As someone who called himself a cloud economist once upon a time because I figured no one would know what that means, great. That appealed to me. And I like the idea of someone calling themselves what appears to basically be unique in the universe, a cloud therapist. Bobby, welcome to the show, and what is a cloud therapist?

Bobby: Yeah, thank you, Corey. Thank you for having me on the program, first of all. Love the format and all the great guests who’ve had. And I kind of took a page from you, Corey, I kind of made up my own title and some of my own job description. So, I call myself a cloud therapist on LinkedIn.

Corey: Hang on. While I’m sitting here, I’m going to do a search. And as it turns out, if LinkedIn cooperates, yeah, there are—ooh, there is a second person calling themselves a cloud therapist.

Bobby: Interesting.

Corey: Well, someone in AWS, which is good. I had the same problem as well because when I was a cloud economist, someone was calling themselves that as well, at AWS. And well, I can let that skate because AWS inherently is terrible at naming things, and job titles, presumably are going to fall into the same bucket. It’ll be great. But I’m curious, what is a cloud therapist? How does that work?

Bobby: So, a cloud therapist to me, Corey, is really about listening because the reality is, there are a lot of things that happened before I got there. A lot of them honestly, they went very badly. And I’ve noticed that a lot of people have almost a level of PTSD, especially from big transformation projects. So, cloud therapy to me is really two things. One, I’m focused more on how I can help you than what I can sell you, and number two, I’m there to tell you what you need to know now what you want to hear.

Corey: A somewhat similar line that I’ve been using for the economic side of the world has been that I largely view myself as a marriage counselor between engineering and finance. Because that does seem to be two folks who struggle to articulate a common love language if you’ll forgive the allusion. And I feel like the more I look at this, it’s less about math, and it’s less about being able to prove things with technical correctness, it’s not about the technology, it’s not about the tools, it always goes back to the people. And if I were to start my business over again, four years ago, I would almost certainly align it more directly with that ethos in mind.

Bobby: Corey, I couldn’t agree more. I think you and I are so aligned on that. Part of my personal mantra for last year and this year has been—the short version—is that tech is the easy part. The longer version is that tech is the easy part, people are the best part, behavior is the hard part, and humility is the worst part. We don’t like raising our hand to admit that we need help.

Corey: No. And part of the problem too is that, on some level, we live in a society that winds up penalizing people when they raise their hand and say, “I don’t know something.” For better or worse, I figured, well, I have no technical credibility to speak of, so why not admit when I don’t understand something? It sort of snowballed from there where other people started speaking up too;, “Yeah, I don’t get it either.” And, “Oh, good. I see you had to wait for someone to speak up, but all right.” It became an interesting story. And I’m starting to realize now that there’s more psychology that goes into so much of this than I ever would have believed previously.

Bobby: Again, I couldn’t agree more, Corey. I think the thing is the soul of technology goes back to people. It’s so easy to forget that why are we tinkering with things? Why are we playing around with new stuff? It’s got to come back to are we solving a problem for a person or an organization to make life better for someone?

And I think, when I kind of step back at some of the conversations I have with executives, this is one thing I’ll throw out for the audience. I’m a big believer that everything new isn’t good, and everything old isn’t bad. Wisdom is about knowing which new things to embrace and which old things to retain. I call that mastering the remix. And if you’re not willing to ask for help and you’re overwhelmed, that mix of old and new is probably crushing you right now.

Corey: That’s a really astute way of framing it. I want to come back to that, but before we do, this is a promoted guest episode by your employer, Turbonomic.

Bobby: Yes.

Corey: My question for you is why Turbonomic? Now, this isn’t just a ‘what does Turbonomic do?’ We’ll get there in a moment. But you have an interesting and storied history as far as things you’ve done. You were at CloudGenera for a long time, and you wound up doing a bit of a tease as far as, ‘oh, where am I going to go next?’ at the beginning of this year, and the answer was Turbonomic. Where were you, and what made you decide that this was the next thing for you to do?

Bobby: Yeah, it’s a great, great question, Corey. We’ll hopefully get into automation a little bit more in the topic. I’ve got a different take on automation, maybe the [many. 00:05:15]. But I used to be a person, Corey, that talked about—folks would say, “I’m drowning in information,” and I would say, “No, no. What you really need is to move from information to clarity to insight.”

And I think the revelation I had within the past year or so is the gap, Corey, is not from information to insight, it’s really from intent to execution. Even if I tell you what you need to do, do you have the time and the attention to do it, or do you really need me to do it for you? So, the automation that Turbo does kind of helps free you up to go focus on the next thing.

Corey: I have a somewhat conflicted relationship with an awful lot of the cloud cost optimization tooling. And generally speaking, we don’t have a whole lot of them on the show, we don’t have a whole lot of them sponsoring our stuff because of a really strange divide that we’ve found over the years. Either I wind up actively insulting the tool or product—which is not a great look when people sponsor things so that’s a problem—or the other is I wind up saying great things about it; it’s perceived as a full-throated endorsement. And that’s impossible for me to do in this space just because, as we’ve discussed already, I don’t think that it is inherently a tooling problem nearly as much as it is a people problem. That said, one of the things I do appreciate about Turbonomic when I took a whirlwind tour through it was that in many cases, it is hands-off—you configure it to do certain things, and then you don’t have to mess with it anymore. It just works, and it’s a disappears-into-the-background offering. And that is, in many respects exactly what people need around a lot of things that Turbonomic does.

Bobby: Mm-hm. So, I agree, Corey. I think the other part is Turbonomic can kind of meet you halfway because sometimes the things you want to automate still need to be routed through something like a ServiceNow, to get people to bless it, so I think that part is really cool. But I’ll give this analogy to the audience, Corey. So, Turbonomic is not a cost optimization company.

We’re really a performance company focused on applications. And so here’s the difference; I’ll use a gym analogy. There’s a healthy way to lose weight, and there’s an unhealthy way to lose weight. So, for a layman, the way I would sum up Turbonomic is we’re optimization without unintended consequences. We’re not the diet pill you take to drop 50 pounds and your kidneys fail.

Think of us as a combination of the nutritionist and the trainer so that you lose weight and add muscle in a healthy way so that your body has the nutrition and the physical makeup that it needs for you to be sustainable and successful.

Corey: It’s effectively teaching people the long-lasting approaches rather than the quick fix of come in and, “All right, flip that button there. Hit that switch there. No, you flipped that switch incorrectly. Do it more like this. Thank you. Pay me.”

And then you vanish and you haven’t really fixed anything. You’ve just caused a minor inflection and made people feel good, but it isn’t building the lasting muscle that you need in these spaces to ultimately fix the root of the problem, which comes down to having communications, effectively, cross-functionally.

Bobby: Corey, you’ve said it well in other formats. Cost is a proxy for value, but cost is a symptom of, typically, something that’s a lot deeper. A bad cloud bill is poor communication, is overprovisioning because you don’t know what the applications need, there are other symptomatic things. And I think, again, back to being a cloud therapist, if you listen to people, they’ll tell you what’s wrong, they won’t necessarily be able to tell you why, and you’ve got to listen to the bigger point of what’s happening. And so sometimes, I’m screaming about the bill but there’s really a bigger issue. I don’t know what resources my apps need. I don’t know where to draw the line. I don’t know what I should be doing. I need some coaching, and I need some automation so that package can help me operate my environment better.

Corey: I think your framing of the question started out on the right path. “Oh, you get a bad cloud bill. Well, hang on a second. Why is it bad?” I mean, theoretically, you can drop the cloud bill to zero by turning everything off, but that is surprisingly unappealing for most companies.

It comes down to is it too high? Is it really that it’s high? Because people accept when they run businesses, that there’s a cost to providing their services or goods, and that’s a natural order of things, but what they don’t understand in many cases is that when finance is complaining about the bill, it’s not that it’s too much, it’s that it doesn’t align with projections. So, what is it that’s really driving that question? Was it that it wasn’t predictable, that it wasn’t planned for, budgeted for or, in many cases, is it the fact that you just hired a data science team and now the bills in the stratosphere while it’s very hard to articulate the value they’re providing for you so far?

Bobby: I think that’s the key, Corey, is that—you know, I’ll use another analogy. When grandma’s transmission goes out in her car and I’m weighing whether I put a $10,000 repair around that vehicle, I’m not going to do that if she’s going to stop driving next year. So, like I said, cost is a proxy for value. In the end, the real issue that we’re struggling with is what is the value of that application? Because applications are the bridge between technology and people, right?

That’s how we’re delivering something to make someone else’s life better. And we’re struggling because the bill may be going up, but the value to our customers and our users isn’t necessarily going up and we’re not aligned. If a Netflix cost goes up, or Disney+ goes up, that cost going up means they’re adding more value and they’ve got more customers. They’re not complaining about that as much.

Corey: Right. It’s you look at these things, and like, oh, wow, Netflix, or Lyft, or whoever it is, that discloses cloud spend in a variety of different ways through their public filings, it’s easy to look at that and, “Oh, what are they spending all that money on?” It’s the rest of the S-1. The business fundamentals that mean that they have successfully gone from harebrained idea to something that is viable in the public markets. That’s also invariably second-place—at best case—compared to the cost of payroll. In some cases, you’re also going to see office real estate trumps that as well, but let’s be realistic in a post-COVID time, that one’s a bit of a question mark in a lot of places.

Bobby: Exactly. Corey, just to kind of segue, communication and people is kind of a common theme I hope the audience is hearing. And I tell customers all the time, missed expectations sink more projects than bad code or broken APIs. We have all this tech, but we’ve gotten lazy in terms of having the right conversations. We went to the cloud, or we did this app this way to make things better.

What does better actually mean to us? Is it cheaper? Is it faster? Is it bigger? I’ll be specific about that for a second, Corey. When we talk about making something better, do I need to do a good thing faster, or make a mediocre thing better? Because a bad recipe at scale is still nasty.

Corey: I really, really wish that you could have said that story to some of the institutional kitchens I was at at a bunch of my early educational processes. Oh, my stars. “Yeah, could you fix the recipe first before you scale it?” Ugh. But there’s also this misguided belief that every company holds, every engineering department has, specifically with the idea that after this next sprint completes, then, then, Bobby, we’re going to start making good decisions, and pay off all of our technical debt, and start doing the right thing all the time. And we, of course, will agree on what the right thing is. It’s a ludicrous fantasy that everyone holds, on some level.

Bobby: And you’re right again, Corey. I feel like I’m saying that a lot today because we’re probably agreeing more than we usually do, but that’s cool. In grad school, I had a professor who talked a lot about verification and validation. The first question is, do we do it right, but the more important question is, did we do the right thing? We are not asking that question enough, Corey.

We’re building two-story houses for people that are in wheelchairs. And then we look back and we wonder why people are upset. It goes back to expectations, communication. The technology, all the stuff that we can do, Corey, is making us fall in love with a science fair project and not tying it back to is this what this person or this firm needed me to actually do for them? And then we get upset because we spend a lot of time and effort on things that really weren’t relevant to meet the need.

Corey: So, I want to take a bit of a detour as well where I normally would call this a side project, but that in this case, would be a horrific insult. And that is not at all my intention. You’re very upfront about not just being a cloud therapist, you are also a pastor. And I want to be clear, not a cloud pastor, an actual legitimate pastor. Tell me, first, about, I guess, trying to balance those two worlds, and secondly, how they inform each other.

Bobby: Thank you for the question, Corey. One, I know that’s maybe different territory for your audience, so let me try to sum it up.

Corey: Well, let’s be clear here. It’s also different territory for me. I’m a somewhat secular Jew, and the whole pastor, religious services, faith thing is an area that I always felt like a bit of an outsider in American Christian culture. And if I do, I know other people feel much more so. So, it’s time to start normalizing some of these things and say, “Yeah, I don’t know what I don’t know.” Please, continue.

Bobby: Oh, thank you. So, Corey, I serve as one of the pastors at a multicultural church in the south. I live in Charlotte, which I like to call Silicon South to let people know we’re not just country bumpkins sitting on tractors. And so my church is about half-and-a-half black and white. We have three black pastors and two white pastors.

And I honestly, Corey, feel like my life is better because I get to do life with such a diverse group of people. We have white families, for example, that have adopted black children. We get to process things together, we get to talk about is this racist, or is this ignorant? I’m struggling to find the words that have this conversation. We get to do a lot of those things and I think that’s why for me the pastor part of me wants to make sure that I listen for how people are hurting, I listen for, again, what happened before I got there, and I think about how to apply.

Here’s the key, I think, for the audience, too. A lot of times your passion projects can teach you skills that you can transfer to the enterprise and vice versa. Let me tell you what I mean about that. One of the biggest projects I’ve done—probably hardest thing in my life, short of somebody dying was managing a building renovation project. About the only thing harder than doing building stuff for a church is getting a loan for church, no bank will foreclose on the church so they don’t want to lend any money to you.

Managing that project definitely tapped into skills that I acquired doing things like the Bank of America-Merrill Lynch data conversion, managing schedule, resources, budgets. So, when that came down the line on the pastor side, I said, “I’m built for this.” Bank of America trained me for this CloudGenera and ServiceMesh trained me for this. And it also goes the other way. When I’m managing volunteers at church who don’t owe me their allegiance, I’ve got to lead them and motivate them; that applies to influencing people in the enterprise. So, I think they’re very complementary to each other is the way I’d sum it up.

Corey: It’s fascinating watching fo—at least from my perspective, who have this multidimensional aspect to them. I mean, we all have it to some extent. I spend a lot of time in my off hours—such as they are—being a parent, or indulging myself in cooking and whatnot. But they’re very different than the activities I pursue in my professional slash public life. And it always is, frankly, more than aspirational to find someone who’s I guess, non-public, non-professional aspect of what they do is also devoted towards, I guess, either transformation or in this case, helping people.

Let’s call it explicitly what it is; it’s helping uplift people, which is something I’ve threaded through what I do professionally, at least I try to. But in your case, it’s an explicit calling to my understanding. It is a, in many ways, relatively thankless task that is never done. Is that an unfair characterization?

Bobby: No, that’s fair, Corey, you’re not a pastor for the compensation package. You’re in it because you care about people.

Corey: Careful how loudly you say that. I’m sure some VC is going to hear that and their ears are going to go on point.

Bobby: [laugh]. You’re in it because you want to help people, and I would sum it up this way: being a pastor shows me how I’m a beautiful mess. I am flawed, I am mistaken, I’ve also learned a lot, parenthetically—thing that goes right along with being a pastor is being a husband of twenty-one-and-a-half years now. I’ve learned a lot of things from my wife, and humility has taught me one of the biggest things that we as men do often too, Corey, is we don’t listen to our wives enough. I’ve said on social media before, listening to your wife doesn’t make you less of a man, but not listening to your wife may make you a less successful man.

And being a pastor and seeing how many times I’ve been wrong about things, how many times my wife has been right about things has kind of humbled me to understand that. You know, my wife has been my test audience. I’ve also said that before, many guys say something deep, we did something dumb. So, we need to thank the spouses, partners, and mentors who gave us grace, while we figured stuff out, my wife definitely falls into that category.

Corey: I’m in the same boat. I always feel a little, I guess, ashamed, let’s be honest. Ashamed that so much of what I do and who I am is only possible because my wife has a job as a corporate attorney. And yeah, when I was starting this place out and figuring out how I was going to work, I didn’t have to support the family as I was going through that; that gave me certain amounts of latitude. There’s also the aspect of the constant emotional support, the ability to help pitch in and put the girls to bed when I have a late-night event that goes late.

There’s a lot in there and it’s one of those areas where if I did it, as the husband, I would be lauded for these things to help support my wife. But when my wife does that, for me, culturally, that’s very much a well, that’s what wives do. It’s a double standard and it’s terrible, and I am ashamed for the fact that I don’t do a good enough job of calling that out frequently enough.

Bobby: It’s balanced, Corey, I think the other thing that we’ve got to look at is even when we deal with challenges on the corporate side, it’s still informed by people. Sometimes people are frustrated with their jobs because they’re frustrated at home. It is really tough to like your job when your spouse can stand your job. And sometimes that happens because you’ve been giving your best to your job and they’re getting leftover energy at home. My wife and I are very direct with each other, and she’ll tell me—you know, sometimes, Corey, people will say, “You’re being a rock star at work.”

And my wife has said to me before, “I feel like your company is getting a rock star, but I’m not getting rock star at home.” And before I get offended, I remember what another one of my pastor friends said, you need to sit down and ask your wife, “How often do you have my undivided attention?” And be quiet and listen to whatever she has to say. Don’t be defensive. Suck it up and realize there’s some truth there in that feedback that you probably need to process.

Corey: Let’s be honest—the past year has been a nightmare for cloud financial management.

The pandemic forced us to move workloads to the cloud sooner than anticipated, and we all know what that means—surprises on the cloud bill and headaches for anyone trying to figure out what caused them.

The CloudLIVE 2021 virtual conference is your chance to connect with FinOps and cloud financial management practitioners and get a behind-the-scenes look into proven strategies that have helped organizations like yours adapt to the realities of the past year.

Hosted by CloudHealth by VMware on May 20th, the CloudLIVE 2021 conference will be 100% virtual and 100% free to attend, so you have no excuses for missing out on this opportunity to connect with the cloud management community. Visit cloudlive.com/corey to learn more and save your virtual seat today. That’s cloud-l-i-v-e.com/corey to register.

Corey: You have to sit with that for a while. It’s easy to wind up pointing out specific times and specific weeks for example. Like, during re:Invent as far as my family is concerned, I am basically calling in dead. And that at least becomes something that is time-bound. The more pernicious, the more dangerous aspect is when it slowly starts to become everything.

Because, “Oh, it’s just we’re on a tough sprint right now. I’m going to go ahead and focus on that instead.” Or, “Well, right now is a big deal I’m working on, so once this is done, then I’ll go back and have time to make it up to the family.” And then you turn around and you’re retired and elderly, on some level, assuming we’re all fortunate enough to live that long, and you realize you never made time for the things that matter. You missed watching your children grow up, and you missed being there to support the people who matter.

And, frankly, one of the reasons I started this place that I have to continually remind myself to keep in mind is that I did it because I didn’t want to continue working startup hours in the hopes that someday things pay off. And in fact, most of the people who work here work what amounts to a 40-hour a week-ish—schedule. And when I say ‘ish,’ that’s not at a floor; it’s at a ceiling in many cases, and that’s fine. The reason I do it is because I don’t want to do that grind, I don’t want to play those games, and I don’t want to wind up having these awful scenarios of having to figure out what it is that we’re doing, and promising people, and stringing them along, and burning them out. I want this to be sustainable. I want to build a place where I can enjoy my life and my job doesn’t consume me—unless I want it to—but also for other people as well.

And I’ll admit there are times that becomes very challenging, but that’s the beauty of doing things in a bootstrapped way where we’re supported by the magic of revenue and profitability. We don’t have to sprint to hit runway targets before we’re out of business, in the same way.

Bobby: Yeah. That’s why I say we’re beautiful messes, Corey. Especially as husbands and fathers, we’re working on a plane that we’re flying at the same time. And my kids are a little bit older than yours, I’ve got a 12 and a 14-year-old; I’ll give this piece of advice to your audience that may have younger children. I used to think that the kids were going to need me the most when they were younger, right, because babies and toddlers want to talk but don’t have a lot to say.

Teenagers, on the other hand, have a lot to say but don’t want to talk. They need you more as they’re older, where you have to watch body language and what they’re not saying. And again, the way this all comes together, Corey, is being a pastor, being a husband, being a father makes me better at my job, not the other way around. We think that they’re detriments, but learning how to read people, how to connect with them, and how to watch for what they really need, not just what they’re saying, will serve you in any aspect of your career.

Corey: I think that there’s an awful lot that needs to be aligned and built in a way that we start looking at people through the context of whole career. We don’t though. Everything short term; everything is, “Oh, just at this company, I’m going to work like nuts for a few years, and then it’s going to pay off and we’ll hit that equity point.” It’s a lie. It’s always a lie.

Look at how engineering works. We talk about sprints, a two-week sprint. And then what do you do at the end of the sprint? That’s right, you do another sprint. It’s not a sprint, it’s a marathon. If you’re running all the time, you burn out. It always bugged me, back when I first learned how agile is supposed to work, and I don’t think I’ve ever quite gotten over it.

Bobby: So, Corey, if I kind of build on that, but connect, kind of, life and technology because I feel like a lot of technology lessons are really life lessons or vice versa. One thing that I’ve been wrestling with, I feel like in tech and in life, Corey, we’re struggling with how to evaluate better versus different. Do I need a faster horse or do I need a car? And the challenge is, in a world of overwhelming options—watch this—how you choose is more important than what you choose. And most of us are overwhelmed I find because we’re focusing on a choice, not a plan.

Because without a plan, you’re one more option away from being overwhelmed or starting over again prematurely. We need things, Corey, that are going to free up our minds to go focus on the next problem. And so, tying this back to my firm for a second, Turbonomic is not a choice; Turbonomic is the way you choose. We need to pursue things in life that free us up to focus on family, to focus on spending time with people, to focus on solving the next challenge. Because other than that, our minds are still occupied spinning on things that we can’t resolve.

Corey: We spend so much time doing those things. You’re right. One of the things I like about what Turbonomic does, as well as a number of other tools, is it almost takes the big reserved instance or savings plan purchase out of the equation in some respects. And again, you can configure it otherwise, but something I’ve seen in my clients is we talk about the psychology and we talk about the math. Well, yeah, it’s pretty mathematically straightforward—especially in the world of savings plans—to go ahead and say, yeah, you should spend about $20 million on that savings plan. That makes sense.

Everyone can agree that, “Okay, well, what if X, X, and X?” And, “All right, fine. We’ll make it 18 or 17, or whatever it is.” Fine. That gets agreed upon. Okay, just click the button in the console, add it to your cart, now click purchase.

And no one does because it turns out that clicking the buy button on a cart item that is more than you’re likely to make in your career is a gut-check moment. And people sit around for nine months because they’re scared to click the button and it's great. Okay, having tooling do that is helpful. The way we approach it is we’re consultants and we’re sitting there next to them and hold their hands, in some ways, through it and where, if you’re still on the fence, cut it in half, buy that now and then we’ll reevaluate in a month and go from there. But you’re going to spend the money anyway, may as well do it at a discount. And that’s the psychology piece of it. I like your approach to taking that out of the critical decision path where people in some cases need to take a strong drink before clicking that button.

Bobby: Right.

Corey: It’s surreal. I’ve clicked that button before and every time I quadruple check it. I used to take the weasel’s way out back when I was running ops teams; for those big purchases, I would go ahead and make my Amazon account rep do it on the back end. Because if they mess it up—and spoiler, sometimes they did—it’s on them, and they can unwind it with an apology rather than me begging on my knees to please, please, please don’t get me fired and sued here. It’s a different dynamic there, and I really wish that there were more effective ways to do it.

One thing that GCP gets right, in this case, is the idea of sustained use discounts. Use something for at least x hours a month, and you automatically qualify for a discounted rate on that thing. Awesome. All I have to do is not do anything. Well, I’m great at that.

Bobby: That’s a very under-marketed part of their offering. I agree 100%. I think a lot of people don’t know enough about that. I do want to go back to something you talked about, Corey, with when we look at automation in general—and this may be a little controversial, right in ter—freeing up your mind, freeing up your time is a big thing for me. Automation doesn’t mean that you can stop thinking; it means that you can start thinking about the next thing.

And so I believe, personally, automation was really meant to enable thoughtfulness, not push us towards laziness. And that’s how some of us are using it, right? We talked about better versus different. The key is that you automate better so you can think about different. So, you can’t automate everything in the world.

That’s the other place some people get stuck is hitting the buy button, but then is thinking, “Okay, if we’re going to automate, let’s automate everything.” You can’t automate everything yet. Automate the stuff that’s better; focus on the stuff that’s different.

Corey: And that’s part of the challenge, too. I see this industry-wide, where people are approaching everything through a lens of SaaS. And well, great, can we turn this into something that can only be solved with software where maybe we have a pro services engagement from time to time, but it’s going to be software for that recurring revenue approach. When I talked about starting The Duckbill Group with a bunch of people who were kind enough to sit down and give me their unvarnished opinion. Most people were positive on that.

But one person who wasn’t was a successfully exited founder who had done very well, and their response was, “Yeah, this doesn’t seem that great a business to me because yeah, I can see you making 10 million bucks a year out of this place, but I don’t see a path to a half-billion dollar exit.” And I’m sitting there going, “Well, I appreciate your feedback, first, because feedback is a gift. Thank you. Secondly, exactly how much money do you think I need?” But the thing is, they were right, on some level.

If I want to become a celebrity force, who’s basically famous for being able to buy that fame and have an insulting number of commas in my net worth, I’m in the wrong business and I’m approaching this very differently than I should. That’s an intentional choice. It’s not that, oh, well, Corey, couldn’t find a way to succeed at being a billionaire. Well, first, probably right. But secondly, I was never interested in trying just because it’s—I want prosperity insofar as it is a successful outcome for my family where they never wonder if the roof over their head is going to be there tomorrow, whether food on the table is going to be assured, and have a nice life. And beyond that, I’m not sitting here thinking what I really want? Nesting doll yachts.

Bobby: Right, [laugh]. Right, right. Well, Corey, you hit a good point, though, because the reality is we need advice but we don’t like paying for advice. And I think the dilemma is when we tie that back to cloud, and applications, and the way that we look at technology today, cloud, in my opinion, is at best a teenager that just learned how to drive. It still needs adult supervision, it still needs boundaries.

And that’s what you do via consulting; that’s what my company does via software, but cloud is at a dangerous point, Corey, because the capabilities go beyond the ability to comprehend consequences. Help is needed, and I think that there are some people—I’ve heard this from executives—some people only want to pay for advice. Other people don’t want to pay for advice and only want to pay for products. I think there’s a market for both.

Corey: There absolutely is. And I’m not sitting here suggesting there’s only one right answer here. I want to be very clear on that. There are multiple paths to success, and I’ve always said that there is room for many more Duckbills Group in the world, and that is fine. I have no argument there.

I’m just a believer of things that make it easier for people to improve their situation. And one depressing but fortunate for my business aspect of cloud finance is that the AWS bill never gets smaller on its own. When I reach out to someone like, “Oh, no. We’re set with our bill.” Our response is always, “Great. We’ll talk to you in a few months,” because it gets bigger inherently on its own. And I don’t like that aspect of it because it seems to not make people super happy. But there is the counterargument as well of it does make for a somewhat sustainable and ever-growing total addressable market.

Bobby: It definitely does.

Corey: I would give it up if I had the option to and could wave a wand. I absolutely would. There are other problems I would like to tackle. But now I’m worried that I will retire without this problem ever being solved in a meaningful, global way.

Bobby: Well, I mean, that’s part of the total addressable market, as you talked about. I think Duckbill can have plenty of other little ducks and spin-off companies and ducklings or whatever because they’re going to be other adjacent problems that I think are going to crop up, too. People are going to need help, they’re going to start to embrace humility more, in my opinion, that I can’t do this on my own. I’ve talked about this in other forums, Corey. There’s a difference between self-help and self-service.

One is, “Can I do it on my own?” The other is, “Can I be effective on my own?” And again, if cloud is at best, a teenager that just learned how to drive, I probably can’t be effective on my own. I need to raise my hand and bring in people like The Duckbill Group or Turbonomic, to help me figure this thing out.

Corey: I also want to be clear that, from many people’s perspective of how this stuff works, it’s natural to conclude, oh, you’re a consultant, but Turbonomic is a product, so obviously, your competitors. No. Of course, that’s not true. It turns out that virtually every customer we talk to has some software thing somewhere that is, in many cases vendor-provided that handles aspects of this. We are not ever going to be an automated tool that solves these problems.

We explored that with our DuckTools experiment, and it turns out that what we think the industry needs, and what the industry wants to buy are two different things, so we shuttered the thing because frankly, I don’t have the stomach to sit here first educating people about the problem before then selling them a thing that fixes the problem I just taught them about. It doesn’t work. And that is too much swimming uphill.

Bobby: It is very hard. I mean, again, you’ve got a very successful [pocket 00:31:36], Corey. I think people trust that you have their best interests at heart and that you’re using real-world experience and customer exposure to help them not hit all the potholes. The thing that I think you do a great job of—and I talk to executives about this a lot—you want me to tell you what you didn’t ask that you should have. You don’t want to hit every landmine that the people did before you; they bring you in partially, Corey, because they don’t want to hit every landmine on their own. They want to hear your story so you can coach them on how to fix it ahead of time.

Corey: If we can delve into tech for a minute, there are more services than anyone can shake a stick at when it comes to cloud provider offerings. How do you help companies figure out which ones they should use? Which ones they should avoid?

Bobby: Yeah, great question, Corey. Without going into the weeds, I’ll give you this analogy. I call this ‘chicken in the cloud.’ And it’s all about what do you want to be known for? So, think about there being, kind of, five ways that you can engage with chickens.

So, you can grow it from a baby chicken, you can pluck a dead chicken, you can cut up a chicken that you bought at the store, you can cook a chicken that you bought in a pack, or you can use a rotisserie chicken that is already done. And the reality, Corey, is I believe that a lot of people are not going to pridefully say, “I groomed that little VM from when it was a baby chicken to learn how to walk and chirp.” The reality is cooking and creating are not the same thing, and the question we’ve got to ask ourselves is are we putting time and effort into things that really don’t matter? At the end of the day when that plate goes on the table, the person doesn’t care if I grew that chicken or if I started with a rotisserie chicken; they want to know the recipe and the dish is solid. They don’t care where I started from. So, think about what you want to be known for and let that guide your decision around what you choose to engage with.

Corey: Bobby, thank you so much for taking the time to speak with me if people want to learn more about you, about Turbonomic, about anything we’ve talked about, where can they find you?

Bobby: So, my company is at turbonomic.com, you can check us out there; there’s all sorts of information. You can find me at bobbyjallen.me. I’m also on Twitter; think ballen and Charlotte: @ballen_clt to represent Charlotte. And, Corey, I want to thank you. I want to leave your audience with this. This is, kind of, one of my personal mantras. “Greatness is about what everyone sees. Excellence is about what anyone sees. Faithfulness is about what no one sees.” Seek being faithful over being famous and everything is going to work out.

Corey: And that’s a great point to leave in on. Thank you so much for your time and of course your insight, as always.

Bobby: Thank you, Corey.

Corey: Bobby Allen, VP of Strategic Alliances slash cloud therapist at Turbonomic. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with a comment including your pitch deck for Pastor as a Service that I’m certain is moments away from being funded, as I speak.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

This has been a HumblePod production. Stay humble.

View Details

About Talia

Talia Nassi is an international keynote speaker who delivers content on all things testing and quality. She is a developer advocate at Split.io where she works closely with engineering teams globally to ship software more efficiently. She is passionate about feature flagging, canary launches, CI/CD, testing in production, and A/B testing. She has spoken at countless conferences internationally, ranging from audiences of 100 to 4000!

Links:

  • Split Software: https://www.split.io/
  • Flagship Conference: https://flagshipconf.com/
  • Twitter: https://twitter.com/talia_nassi

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: If your mean time to WTF for a security alert is more than a minute, it's time to look at Lacework. Lacework will help you get your security act together for everything from compliance service configurations to container app relationships, all without the need for PhDs in AWS to write the rules. If you're building a secure business on AWS with compliance requirements, you don't really have time to choose between antivirus or firewall companies to help you secure your stack. That's why Lacework is built from the ground up for the Cloud: low effort, high visibility and detection. To learn more, visit lacework.com.

Corey: The apps on cloud summit is a new action packed, not a conference, happening May 11th through 13th online. Its for everyone who makes applications in the cloud run screaming. From IT leaders to DevOps pros to you folks, whoever you might be. Take a break from screaming into the cloudy void with me to learn from some of the best of people who actually know what they’re doing. Like Kelsey Hightower, AWS blogger John Meyer, and also me, because apparently they didn’t listen to me saying I had no idea what I was doing. Register now at turbonomic.com/screaming. Theres a “swag box” ready to ship for the first two thousand registrants, so you don’t want to miss this. Thanks for Turbonomic for sponsoring this ridiculous podcast.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Talia Nassi, who's a developer advocate at Split Software. Talia, thanks for joining me.

Talia: Thanks for having me. I'm excited to be here.

Corey: So, let's start at the beginning. What is a Split Software? And please don't tell me the answer is that there's two of them.

Talia: Split Software is a feature flagging and experimentation platform that allows you to create better software for your teams so you can gain insight into what your customers are doing with the use of feature flags. And Split provides all of that inside of their platform.

Corey: So, in other words, it's more or less the ability to enable certain things on the fly, disable them on the fly without having to, for example, do a whole new code deployment?

Talia: Right, exactly. And it makes the release process just a lot faster and more seamless because you don't have to have a big release, now, you can turn on a button or flip a switch and that feature will be on.

Corey: One of the earliest tech presentations I saw—and it feels like it was probably almost a decade ago—-, was talking about feature flags. It was from one of the big well-known tech companies. I don't even recall which one, which tells you exactly how well that branding thing worked. But it felt like, “Oh, this feels like something from the future that the big tech companies will do, and maybe someday a few other folks will be using it.” And that's kind of the really the last time I went deep dive, checking into that sort of thing. But looking at Split's website, for example, I wind up seeing a whole bunch of logos that I recognize, which tells me it has sort of come down from the ivory tower of the big well-known tech places and into something that, you know, human beings can use.

Talia: Yeah. Yeah, absolutely. So, we do have a ton of integrations with Google Analytics, and mParticle, and Amplitude. And then we have a bunch of supported SDKs, like, all of the main ones. I do a lot of tutorials in Python, and React, and JavaScript, but we have so many different places to start if you're interested in using our tools.

Corey: So, developer advocacy is one of those areas that means a lot of different things to a lot of different people. I mean the running joke is, what, get three DevRel folks in the room and ask what DevRel is, and you'll wind up getting eight answers at least. And again, my expression of it is usually aggressively shitposting on Twitter aimed at trillion-dollar companies. But everyone finds that they tend to embody different aspects of it. So, it sounds like a weird college entrance essay question, but what does developer advocacy mean to you?

Talia: It means someone who's an—I don't want to say an advocate for the developers, but it's someone who the developers can trust. And there's a few parts to that. So, the first part is just having some sort of engineering credibility. So, I started as an engineer, and I was a software engineer in test for five years before becoming a dev advocate. So, I do have a foundation of software engineering.

And so the second thing, I think, is dev advocates are developers at heart, so they're not going to try to sell you on a tool. Like, I don't sell Split; we have a whole sales team that deals with that. But my role is to make it easier for the developers to use Split. So, I work on things like documentation, and I work on video tutorials, and code tutorials, and create sample apps, and just make it so that if anyone wants to use Split, it's easy for them and they have a really great experience.

Corey: At some level, you would think from the naive objective perspective that making sure developers have a great experience is the entire role of product and then, in theory, the engineers that build it. One thing that I've learned through fronting a small company of my own, is that every aspect of business is way more complicated than it looks from the outside, even at tiny scale, and you folks are larger than that. Where does the breakdown fall? Because it's easy to say that, “Oh, product should wind up making sure that the developer experience is a delight.” But if that were true and completely inalienable, then developer advocacy wouldn't exist because product would already handle it? Am I right on that? Am I missing something? Or is there some nuance that's, that’s—

Talia: You’re right. So, basically, what dev advocacy does is it finishes the feedback loop. So, you have product and engineers who create the product, and then you have developers who use the product. But then if there's issues, or problems, or things that can be improved, that feedback needs to somehow get back to the product team. So, I think what dev advocacy does is it fills in this cycle of when something goes wrong, or if something should be improved, or if there is a typo in the documentation or something isn't clear, that feedback all needs to get back to product to continuously improve, and that’s, I think, where dev advocacy comes into that cycle.

Corey: It's one of the I guess, big debates of DevRel across the industry has been, well, what organization does it belong in? At some level, it seems that some folks spend more time talking about what DevRel isn't than what it is: “Oh, we're absolutely not sales.” I agree with that. “We're absolutely not marketing.” Well, I have challenges with that approach. “We're not product.” Well, okay. Yes, but no. “We're not engineering.” Well, really? You write an awful lot of code for someone who's not engineering. And so on and so forth. It feels almost like it's an intersection, and it feels like it's a very nebulous thing to define.

Talia: Yeah, it is. And I think with every company, it's different. I think, in an ideal world, DevRel would just be its own division within a company, not with marketing, not with engineering, it would just be freestanding on its own. And it does combine a lot of engineering and a lot of marketing, but what you’re marketing is not the product, you’re marketing code, and tutorials, and things like that. You're not marketing, like, the actual product.

So, I think there is an intersection, obviously, between product and engineering and DevRel, but in terms of where it should belong in a company, ideally, I feel like it should be its own division.

Corey: The challenge, of course, is from a business strategy perspective. When it's nebulous, it's difficult to assign value and determine appropriate level of investment in it. Back in the before times, for example, it was always tricky to articulate the value. “So, let me get this straight. You cost as much as an engineer, and you spend about that much again in travel to go to conferences that look like giant parties from where we sit. And I talked to the VP of DevRel, whatever that position happens to be, and they say that no, you can't tie a sales quota to you folks. And okay. And I can't tie you to all the metrics I suggest, like, inbound leads are bad. So, without anything I can use to evaluate performance of whether someone is outperforming or underperforming in the group, what happens if I just let the entire group go?”

And you see, in some cases when the pandemic hit, that's exactly what some companies did. I want to disclaim my own biases here. I believe that is a mistake, but I'm curious to get your perspective.

Talia: Yeah, there's companies that really benefit from having dev advocates, and I think those are the startups as well as the big companies. So, Split, for example, I think our team benefits from me because I'm so wonderful. I think we benefit because when developers have questions and when they need tutorials, that's where I come in. And they know that, “Okay, she's not trying to sell us anything, she really just wants us to have such a great experience.” And then there's cloud advocates like in AWS and in Microsoft, and they provide kind of the same tooling and support that we do at smaller companies. So, I think it is beneficial, but you just have to know how to do it the right way.

Corey: Back when I was on the speaking circuit—again in the before times—again, I was an independent consultant for a few years and then I wound up taking on a business partner and expanding, and one of the things I went through with taking on a business partner, who himself is engaged publicly. He's an O'Reilly author. Good for him. But there was a bit of where, how do we quantify the value of me going around and speaking at these events? And the answer that we ultimately came to—and I'm not suggesting this is perfect, ideal, or even good, if I'm being honest, was that we don’t. I know that on some level, if I go out, and I talk about the things that I do in front of an audience, good things happen. But, first, it's impossible to quantify and, two, attribution will drive you out of your mind if you let it.

Talia: Yeah, I agree with you. And I don't think that there's any way to quantify this role because you're not measuring the amount of leads because you're not doing sales, and you can't really quantify the traffic that comes to your site because it might not lead to a quantifiable number of sales. So, I really don't think there's a way to quantify success in the role, it's more of just the quality of what you're doing and how you're engaging with developers, and yeah, things like that.

Corey: And it's challenging in the context of larger organizations because as you start expanding your developer advocacy groups and your developer relations functions, they get—I don't mean to be unkind—pretty expensive at some point. And when you're investing vast seas of money and figuring out where to allocate that, and it's, “Well, this group makes good things happen, but we can't really define it,” isn't good enough from a executive level, the only reason I'm able to get away with it here is because I own half the company and cool, I can basically say, “We're going to do this,” and there isn't a whole lot of recourse. It's the first time in my career where I don't actually have to worry about getting fired. Most people don't have that particular luxury.

Talia: [laugh]. Yeah, the way that would be ideal is if the numbers didn't matter. But if for example, I put up a tutorial on setting up feature flags with Node and React, and then the next day, we see there's a lot of activity on the Node and React documentation pages, and we know that people are building these apps, I mean, it's something that you could use to provide value, but in terms of the quantifiable number, there's not a lot that you can correlate.

Corey: I don't have any answers for this. It's one of those areas that's just difficult to look at because a lot of my friends and associates are in the DevRel space and this is a common discussion. And I'm increasingly of the opinion that no one has really solved this, but I know that taking a step back, companies that have invested heavily in developer advocacy do well in this market when it's executed appropriately, and others who have cut back find themselves floundering, and there's enough data points to make it pretty clear what the right direction is. So, let's talk less than the abstract and more about you personally. There are an awful lot of things that DevRel can encompass. What parts of it do you do? What parts of it do you not do? Where do you start? Where do you stop?

Talia: One of the main things I do is I speak at conferences and meetups, and right now everything's virtual, but I speak at events, and I write blog posts and code tutorials, and I create code samples and sample apps. I also host a monthly roundtable with our developers, so anyone can join and talk about any roadblocks they had or implementation things that went wrong that we can improve. And then we also have community Slack channel where developers can talk and I can also answer questions there. And then—I think at the core of all of this, is I teach. And I think developer advocates, they should be, I think, really great teachers.

So, you're teaching different things to do with the tools that you're advocating for and the products that you're advocating for. So, that's kind of what I do. And I got started with it at my previous company. I was a software engineer at WeWork. And I was doing this really cool thing, and someone suggested, “Hey, you should do a tech talk because this is something really cool.”

And so I did a tech talk internally for the company. And then someone suggested, “Hey this is really great. You should submit it to a conference.” So, I submitted it to speak at a conference, and from there, I just kept getting invitations to more conferences and meetups, and it kind of just skyrocketed from there. And I think I realized that the part of my role that I was missing from being an engineer was this external PR teaching side of dev advocacy. So, now I'm a dev advocate, and I love what I do.

Corey: I adore the way that you started that with the idea of being a teacher. And you're right. One of the most transformative contracting projects I ever did was—must have been six, seven years ago now—where I spent a summer as a traveling trainer teaching people how to use Puppet. And that was fascinating to me from a perspective of… first you have a roomful of twenty people who are within punching-you distance, and you have three days to teach them on this thing. Secondly, they spent not small money to be there, so they're expecting an outcome.

And in many cases, some of them are kind of angry, where there's a perception that, “Oh, this software is coming to take my job away.” So, they're already coming at this from a weird place, then, of course, it's software. And computers are notorious for doing what they do best, and that's breaking. So, at some point, you'll do a demo and it fails, and good luck. Have fun.

Oh, by the way, if they wind up yelling at the company and demand their money back, you don't have a job anymore. So, okay, no pressure. I'm not saying that's the way to learn to speak publicly, but it's one of those drown or swim moments. And that really informed a lot. It was a rapid evolution.

The first class I gave, I was terrified; the second one is I got this; the third one is I don't got this; and by seven, it started to wind up getting rote, and I started experimenting more and being more genuine, and it worked super well as I went down that process.

Talia: I think it takes a specific skill to be a great teacher. And it's one that is really important in being a dev advocate is that you need to be able to teach without belittling and be able to teach people with different backgrounds. Some people who go through our tutorials have twenty years of coding experience and some people have two weeks of experience. So, I think creating content and connecting with people with all different backgrounds is really important.

Corey: There's another thing I learned by doing this was I gave a talk once the early version of it was a talk on Git. And I wanted to go super deep—honestly, I wouldn't recommend this strategy to anyone—but I wanted to learn Git better, so it's all right, how do I force myself to learn Git? I know, I'll sign up to give a talk about it in three months. [laugh]. That's what we call a forcing function because they won't reschedule the conference if you run out of time. I know this; I checked.

And where, time to learn Git. And the first iteration of that talk would have been hilarious to the five people who maintain Git, but it was super deep in the weeds. And I wound up taking a look at this and figured, you know, let's make this more broadly accessible. And it started off by even explaining what Git was. And yeah, there were jokes for all experience levels scattered throughout it, but the lesson I took from that is, it has never been a bad idea to make content more accessible to people.

At some point, you do have to draw a line somewhere. I mean, when I wind up doing content about AWS, I don't start by explaining what AWS is every time; that would get very tiring. But the idea of reducing the prerequisites that are required in order to understand the context of what's happening is very important. And there's no perfect answer. Sometimes there’s—if people don't understand an environment, you aren’t going to be able to teach them effectively because they don't have the fundamentals. So, dialing that in is very important, but given the choice, I will always trend towards being more accessible to more people.

Talia: Yeah, yeah. And I think it also depends on your audience. If you're speaking at a meetup that is college students, and people who don't have twenty years of coding experience, you might want to add those introductory steps. But if you're speaking to all the architects at Google, like, maybe don't tell them how to create a React app.

Corey: No, no, when you architects at Google, they like to get on stage and tell everyone else how they're doing it wrong. There’s a little thing called context, it's, “Yeah, your company has invested tens of billions of dollars in your developer infrastructure. We have about, ehh, eight bucks.” So, yeah, Google scale solutions do not necessarily map to others. And in fact, that's one of the things that started to irritate me and I started getting louder and louder about is when you give a talk about how you do things in the right way to proceed with things and people leave the room feeling bad because what they've built looks nowhere near that good, you've kind of failed.

The theme that I always like to go with is what you're doing is great. Now, here's how you make it even better. And that's an uplifting next step, brighter path, brighter future story. I used to think that there was something wrong with me when I would leave a talk feeling bad. And having done this enough, I'm firmly of the opinion that no, no. That was a bad talk.

Talia: Yeah. And I think you're absolutely right. I think when you leave a talk, you should feel empowered. Like, “Wow, I want to go do this thing that I just learned about and I didn't know that I could do it, and that it was so easy, and this person just enlightened me, and I just want to go do it now.” I think that's how people should be leaving talks that they hear.

Corey: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the Enterprise (not the starship). On-prem security doesn’t translate well to cloud or multi-cloud environments, and that’s not even counting IoT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IoT devices, detects these threats up to 35 percent faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at extrahop.com/trial.

Corey: The whole point is to be uplifting and inspirational and fun. And not everyone is going to have the same approach. Other people's material that's about as well as other people's shoes. But what I find is that my way of grabbing people's attention is humor. If I can make people laugh, I have their attention and then I can teach them something.

Now, some people don't have that particular approach. Their humor doesn't work in that way, or that's not how they contextualize things, they don't know how to tell a joke, or—worst case failure mode—they confuse being funny with being actively insulting. And that doesn't fly. But that's always been my approach. I've turned them into sort of borderline stand up comedy, and everyone thinks, “Oh, that's what you have to do to give a talk?” No, that's my personality defect that I just managed to turn into something of a strength.

Talia: Yeah, I totally agree. Being humorous onstage. And honestly, mostly, what I do that works for me is I just make fun of myself. And that works really well. The other thing is, I'm like a very sarcastic person, and it comes out in my talks.

And so yeah, I think humor is a great place to start. It's been rough during COVID because when I give talks, everyone's muted and I can't tell if they're laughing at my jokes. So, sometimes I'll tell people, if you laugh at my jokes, just unmute yourself, [laugh] so I can hear you.

Corey: Oh, and the delay, too. It turns out that there are a couple of things that are baked into human behavior. I discovered this once giving a talk to the silent disco type of talk. For those who are not familiar because it’s a weird term, there are—usually a big open room, and there are three stages next to each other, and all of the audience for each talk is wearing a different colored set of headphones. So, you're basically whispering into people's ears, and you speak in a normal voice onstage in front of everyone.

But when you're sitting in a room wearing headphones, there's a societal conditioning thing of you don't laugh when someone says something funny. I mean, otherwise you turn into the person who just starts bursting out laughing at nothing while riding the train, and suddenly, oh, you're that person. So, it's challenging when you tell a joke that you know is good and you can see people smiling and laughing, but there's no visible feedback. And video—or recorded video—makes that even worse.

Talia: Yeah. I mean, it's hard for a first talk that you don't know what the audience reaction is going to be. But there's a couple of talks that I've done so many times, and I know that I'm being funny, in certain points.

Corey: So, how have you found that your approach to DevRel has shifted since the before times where now it's, “Hey, do you want to come and hang out in a big room with other people?” “Hell, no.” It's definitely forcing rethinking of a lot of this. What have you been doing?

Talia: Yeah, I feel like right now, there's just more of an emphasis on digital experiences because everything is online from COVID. And I feel like the tech world is just booming right now, so I think, yeah, there's just an emphasis on everything being digital. And I don't know that it necessarily has an effect on my role, specifically. And the only thing that I can think of is that travel is obviously limited right now, and all the events are virtual.

Corey: There's an awful lot of things that a virtual event lets you do, and just as many things it doesn't let you do. And in some ways, it feels like people are still stumbling through figuring this out. It turns out that when people are in Zoom meetings all day long, they don't want to open up Zoom to attend the conference. It's, “Oh, cool. It's just another meeting.”

Anything that doesn't have a strong social channel where you can talk to peers—which frankly, from my perspective has always been the best part of any conference I've gone to—is a problem. There are a lot of weird and broken approaches, and the technology isn't quite there in some respects, and everyone's trying to do the same thing, and, “Oh, we can do digital events. Okay, we'll make it three weeks long.” Why? How much attention do you think that people have?

People don't want to attend online conferences three days a week, as it turns out. So, it's a matter of almost separating signal from noise. I don't have any answers here. It's just a painful problem I keep seeing.

Talia: Yeah, I think you just have to choose your conferences wisely and choose what you go to wisely because there are a lot of virtual events right now, and it can be a little bit draining if you decide to go to as many as you want. I feel like you should just choose wisely, go to the ones that you really feel like will be beneficial for you, and don't go just so that you can say, “Oh, I'm not working right now.” Split has a user conference coming up March 16th and 17th. It's called Flagship, and it's just a two-day interactive virtual conference where we'll have success stories about progressive delivery, and best practices of experimentation, and just, like industry trends. And so we have speakers from Google, and LinkedIn, and ServiceNow. So, that's happening in March.

Corey: Excellent. We will of course put a link to that in the [show notes 00:24:27] because honestly, attending conferences we get actual value out of them is super important. The painful part I've always found is how do you figure out which one it is? And I hate to be judgy but I'm going to do it: the more enterprise-y it looks and the more it looks like you're reading an airport ad billboard, invariably the crappier the conference is going to turn out to be. I'm sorry, but that's how I always feel when I look at this stuff. And credit where due, I don't see that Split is falling into that trap. Hazzah.

Talia: [laugh]. Yeah, it should be a good conference. So, we'll see what happens. I'm doing a workshop at the conference on setting up feature flags. So, yeah, we'll see how it goes.

Corey: It's one of those things I keep wanting to get into. Let's actually talk a little bit about that. Feature flags are something I keep meaning to look into, which makes me feel better about not having really looked into them in any meaningful way because it feels like it needs things that I don't actually do. For example, you know, testing my code. What are the prerequisites for making feature flags something someone should care about?

Talia: There aren't any prerequisites for using feature flags. You just have to have a stable app. But once you have the app, there's so many things that feature flags allow you to do, so things like testing in production, and A/B testing, and canary releases, and percentage rollouts, and things that without feature flags would be so much harder to implement. So, I think it just makes your development much easier, and Split takes care of all of that for you.

Corey: So, testing in production is one of those things that is very frequently talked about, but it feels almost like it's chaos engineering-focused where, it sounds, for the first time you hear about it, like, an absolutely terrible idea. And then it once you look into it, no, that actually makes an awful lot of sense, but it feels like you have to educate the customer before they see the value of it. And that always felt scary to me.

Talia: Yeah. So, one of the things that I talk about—and this was actually the first talk I ever did—was testing in production because it was something that I implemented end-to-end at WeWork. And so testing in production can be really scary because you're running tests in production, you could affect real end-users, you can break things in production. But the great thing is that when you use feature flags for testing in production, you reduce the risk of anything going wrong. So, what I mean by that is you target your internal teammates inside of the feature flag.

And what that means is that you basically create a list of people, and you say only these people can see this new feature. And then you use that list of people and say, “I'm going to test my feature with just these people.” So, now your developers, and your QA, and your product people can go in and see the feature in production, test it out manually, and then turn the feature flag on once it's ready to be turned on. So, basically, you're not using a staging environment or a dummy environment; you're using production because that's where your feature is going to live. Your users aren't going to log into staging to see your feature, they're going to log into production. So, I think people do get scared, but using feature flags is a great way to mitigate that risk because if something does go wrong, your end users won't be affected. It'll just be your internal users, which are your devs, and your product, and your QA.

Corey: Yeah, it's something you really need to get the entire team on board with in some respects. The other piece of it though, is by testing in production, on some level, you at least have the potential to inherently degrade the experience for your customers or your users. At least I have a hard time seeing how that wouldn't be the case.

Talia: So, what you're actually doing is you're making the customer's experience much better. You're making sure that the features work before your customers can see it. So, if I'm releasing a feature, I want to make sure that it works in production, not in staging. So, you just want to make sure that you're doing it safely with feature flags and you have it implemented correctly, but at the end of the day, you're providing a really great user experience.

Corey: Yeah, there is something to be said for accepting a bit of short term pain in favor of making it better for everyone longer term. And I suppose that’s part of the reason why chaos engineering came out of Netflix as opposed to other places, where the failure mode of degrading a particular user's Netflix stream and having them have to restart it or whatnot, isn't super high. As long as you're not continually testing on that same one user. And if you are, honestly, the idea of having a canary list populated with people you don't like is amazing, but at that point, the stakes aren't super high. It works when you're talking about streaming movies. It feels a lot more challenging when you're doing feature enhancements on someone's bank account.

Talia: Yeah. So, we definitely don't recommend using real users to test in production, what I recommend is using test users that are in production. So, the same way that you would log in to Netflix and create an account, so you kind of do the same thing for test account. So, you just create an account in production that's used for testing that acts as a real user, but is not a real user. We had, like, automation bot after automation bot in production, and we would know that whenever any data comes in from these bots that they're actually testing in production and that they're not real users.

And so we don't use actual users when we're doing testing. We're using test users that act like real users and look like real users, but they have this identification of being a test user. We use things like a back end flagging system to identify all test users and test entities in production. So, just some sort of Boolean that's ‘is test equals true?’ Or ‘is test user equals true?’ And that's how we would differentiate test data and real data in production.

Corey: One line that I liked about this was that everyone has a test environment; some people also have an environment where production lives separately. And on one level or another, you're always going to be testing your code, just can you do it in a way that doesn't cause serious damage or harm to users? And getting people onto that path is challenging.

Talia: It is. And I think that everyone has experiences of testing in a staging environment, or a QA environment where they tested their features and it worked so great in staging and they signed off, and then as soon as they pushed to production, there's an issue or something goes wrong. And I think using those experiences to get people on board is a really great approach.

Corey: So, thank you so much for taking the time to speak with me today. If people want to learn more about who you are, or what you do, what you're up to, where can they find you?

Talia: Oh, you can find me on Twitter. I'm @talia_nassi. And yeah, and I hope that you guys come to Flagship in March.

Corey: Excellent. We'll of course put links to those things in the [show notes 00:24:27]. Thanks so much for joining me. I really appreciate it.

Talia: Thanks for having me. This was great.

Corey: It really was. Talia Nassi, developer advocate at Split Software. I'm Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you've hated this podcast, please leave a five-star review on your podcast platform of choice along with an insulting comment telling me that you hope someone tests in my production.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Alan

Developer and DevOps-er; interested in all kinds of cloudy tech, especially deployment pipelines and infrastructure as code. Also building the DevOps capabilities at Hitachi Capital.

Links:

  • Hitachi Capital UK: https://www.hitachicapital.co.uk/
  • Accelerate book: https://www.amazon.com/Accelerate-Software-Performing-Technology-Organizations/dp/1942788339
  • Twitter: https://twitter.com/alanraison
  • GitHub: https://github.com/alanraison

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: If your mean time to WTF for a security alert is more than a minute, it's time to look at Lacework. Lacework will help you get your security act together for everything from compliance service configurations to container app relationships, all without the need for PhDs in AWS to write the rules. If you're building a secure business on AWS with compliance requirements, you don't really have time to choose between antivirus or firewall companies to help you secure your stack. That's why Lacework is built from the ground up for the Cloud: low effort, high visibility and detection. To learn more, visit lacework.com.

Corey: This episode is sponsored in part byLaunchDarkly. Take a look at what it takes to get your code into production. I’m going to just guess that it’s awful because it’s always awful. No one loves their deployment process. What if launching new features didn’t require you to do a full-on code and possibly infrastructure deploy? What if you could test on a small subset of users and then roll it back immediately if results aren’t what you expect? LaunchDarkly does exactly this. To learn more, visitlaunchdarkly.com and tell them Corey sent you, and watch for the wince.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’m joined this week by Alan Raison, who’s currently a developer meets DevOps lead, slash whatever you want to call yourself, really, over at Hitachi Capital in the UK. Alan, thanks for joining me.

Alan: Pleasure to be here.

Corey: So, what is it you do, exactly? It’s a truism that as long as we continue to improve our titles, it obscures meaning to the point where it, on a long enough timeline, is completely meaningless but has 500 words tied to it. Where do you start? Where do you stop?

Alan: Okay, so I’ve got a background in being a developer. So, I started at Hitachi nearly five years ago as a lead developer. But last year, I really sort of—well, I’ve always been interested in the deployment side, infrastructure as code, things like that. And so the opportunity came about a year ago to try and start up a DevOps team inside it, actually. So, I am now DevOps lead, one of two.

We’ve got a small DevOps team. As such, we work alongside the development teams, the testers, trying to set up their build pipelines, their cloud infrastructure, and any, sort of, bits and pieces along the way. So, we’ve got on-prem infrastructure, as well, that we do a bad job of looking after.

Corey: So, I think it’s fair to say on some level, that Hitachi Capital could be considered a ‘legacy’ company. And I use the term legacy in the way that I always hear it when other people use it, which means, it makes money, which is apparently falling out of fashion. It has one of those old-fashioned business models that our grandparents would have understood of, “Oh, you provide a service or good and you [unintelligible 00:02:12] for more money than it takes to wind up providing that service or good, and that leftover is called a ‘profit.’” and that seems to have not really been absorbed as a lesson by the venture capital set, et cetera. Is that a fair way of casting it?

Alan: Well, yeah, it certainly seems to work as a business model for us. [laugh].

Corey: Yeah, “Well, we’re losing money, but making it up in volume,” never really seems to catch on. But you were a listener of this show and reached out with some fascinating feedback. And you said that, “Well, you’re in a company that’s the antithesis of most of the things that I believe.” And you went through a laundry list of those, and we’ll get to some of them, but what I really appreciated about it was that… yeah, I agree with you; I absolutely have a certain perspective that aligns with a certain type of company, just because that’s what I tend to be surrounded with. It’s where my own career has taken me, it’s a large part of where significant and overly vocal—shall we say—portions of my audience tend to reside.

So, it’s very easy to lose sight of the fact that there’s a big world out there, and it doesn’t always align to a SaaS company in the Bay Area building something that is patently ridiculous. So, a lot of what I said doesn’t necessarily apply to what you’re doing and how you’re doing it. And I think that’s great feedback. More of that, please. But of all the weird positions I’ve taken, let’s start at the top. What offends you, slash you disagree with the most?

Alan: [laugh]. Well, I wouldn’t say anything that you say, offends me. I think, in fact, you’re probably on the right lines objecting to some of these things. But we are where we are, I guess. So, for example, multi-cloud.

As developers, we tend to like AWS. It’s in many ways built for developers. It’s the first one developers tend to go to. But hey, we’ve got a sizable Windows estate of servers and whatever else we do in Windows, including a number of COTS systems. And so it makes sense for us to not only use AWS, but we use Azure as well, and we’ve got some Windows engineers who are putting things into Azure, as well as the developers putting most of their stuff into AWS.

That doesn’t sound too crazy for starters; most people, I would imagine if they’re looking to go multi-cloud, would consider two of the big three—who I consider to be, like, AWS, Azure, and GCP—but we also have data centers. We don’t want data centers, but we do have a large Oracle estate, Oracle databases, Oracle WebLogic application servers. And when it came to putting that into AWS, it looked like it was going to be very expensive. Combine this with those Windows servers that these applications seems to talk to, and it was suggested, by I think it was an architect or maybe from some salesperson who spoke to an architect, that we use Oracle Cloud. And believe you me, me and my boss, we fought this for quite some time.

We thought, you know, [laugh] this can’t be a good idea. Three clouds is quite clearly crazy, and we don’t want to go there. But we did look into it, we looked in the numbers and the Oracle salespeople did their technical demonstrations and things, and it does seem that if you want to run the Oracle Database in the cloud, then Oracle Cloud isn’t actually too bad a platform to do it on.

Corey: Okay. A lot to unpack there. Let’s start at the top. Even though I talk about it otherwise, I fall prey to it the same way anyone else does, which is, it’s easy to talk about an ideal world and how you would do things in that mythical environment. It’s the whiteboard fallacy, almost, where design an architecture that solves for X, Y, and Z.

Well, sure, in a vacuum, it’s easy to do. In practice, there are questions of, “Well, what about the real world being messy?” Plus the whole aspect of we’re very rarely working with anything that’s purely greenfield. There’s got to be support for existing workloads that are not built in a way that align with cloud. So, I think that I need to be more cautious about framing things the way that I do.

So first, thanks for the feedback. I really do appreciate being able to catch these things when they slip out and inadvertently start tainting the ecosystem with, I guess, a too forward-looking SaaS-y startup-style company perspective. I mean that sincerely. It’s great to get feedback in a constructive way.

Alan: Well, I still think there is room for that criticism, and it’s not entirely unfounded. But that’s what works for us, and I’m not offended by your snark at the multi-cloud strategy at all.

Corey: [laugh]. I would also say that multi-cloud does make sense for different workloads that have very few points of interaction, especially if they’re already built to embrace the benefits that one cloud provider offers over another. I don’t really have a problem with that. My anti-multi-cloud stance—which I feel like it gets misinterpreted a fair bit—is, “We’re going to build one workload that we could seamlessly move between cloud providers, even though we either never do it or we’re trying to do it for the wrong reasons.” That’s the multi-cloud that I’ve seen being pushed by a variety of vendors and that’s the thing that upsets me.

But personally, I use GitHub, which is an Azure product or rapidly becoming something like that. I use things that are well under the umbrella of GCP. And I run a lot of infrastructure services on top of AWS. So, from that perspective, we’re all multi-cloud unless we’re doing something hilariously wrong.

Alan: [laugh]. Yep.

Corey: Now, as far as Oracle Cloud goes, I’ve played with them before, and I think I’ve been fairly public about it, in that, from a technical perspective, I like an awful lot about what Oracle Cloud is doing, particularly if you’re already an Oracle customer. From migrating a database that is already an Oracle Database onto a cloud, it’s hard to necessarily push back against some of the value propositions that Oracle comes out with. Even on a technical basis from a pure serverless perspective, Oracle Cloud is still pretty decent, based upon what I’ve seen. My problem has always been, honestly, their business practices, their salespeople, their approach to a lot of things that makes it unpalatable for certain customers to enthusiastically dive in.

Alan: Yeah, and I think that’s what made us nervous to begin with.

Corey: So, something else that you mentioned in your email to me was specifically around the idea of serverless development, and I find that fascinating on a couple of levels. The first is that in your parenthetical, you said—right after serverless development—“Not very well,” which is universally true every time I see someone start to work with serverless. It’s, “Are you using serverless?” “Oh, yeah. But we’re really bad at it.

We’re not using it properly.” And I’ve come to the conclusion that if that’s what people think about their own use of serverless, they’re probably doing it right because everyone feels like it’s unfinished. It’s weird. You’re misusing it in a bunch of different ways. But that’s exactly how everyone uses it. Tell me more about that.

Alan: Yeah, I guess so. I mean, there’s always more services to use in AWS. But yeah, we started as, I guess, most people start in AWS: writing applications that run on EC2, and then decided to tear that up and try it out with containers, and then we decided that that wasn’t good enough, so let’s go and try out Lambda.

Corey: And how did that experience unfold for you?

Alan: Yeah, so it’s a massive learning curve for anyone. So, I think I can pick stuff up quite quickly, but when you’re bringing a whole development team along with that, then there’s always going to be gaps in people’s knowledge and things people know better than other people. And so it’s just really hard to get everyone on the same level. And especially because we did, basically, a change of language. We’re typically Java developers, and now who wants to run Java in a Lambda? I don’t think anyone does. So, we now almost exclusively do TypeScript. So yeah, there’s, as I say, a learning on lots of different levels there.

Corey: One thing that seems to be true across the board with serverless, is that it may be one of the first, shall we say, bleeding-edge futuristic-looking technologies that's seeing greater adoption in the enterprise than it is at smaller scale. Is that something that you’re starting to feel yourself? What drove you folks to start looking at serverless, instead of a number of other approaches, be it containers, or something else?

Alan: I suppose it’s the promise of not having to look after infrastructure. We’ve got a number of infrastructure teams at Hitachi Capital, and it always seems to take a long time, or there’s a lot of roadblocks in server provisioning, and security hardening. And we just wanted to sidestep that, really, and go for the managed offering.

Corey: There really is something to be said, for letting the provider handle these things for you. There’s pushback from more traditional ops side of the world where it’s, “Well, at that point, you’re just letting your cloud provider dictate your availability.” Well, yeah, no kidding. You always are. With serverless, you’re just being a little more upfront about acknowledging that.

Alan: Absolutely. Yeah.

Corey: You said that the learning curve is steep. And you’re right, I keep saying that things like Lambda functions and their equivalent on other providers are inherently platforms that are defined by their limitations rather than by their capabilities. And honestly, I’m kind of on board with that, even though we start to see those limitations relaxing on a year-over-year basis, where it forces you to rethink about how it is you’re approaching these types of things. I find that in my experience dealing with enterprises, one of the first things to get Lambda-fied, for lack of a better term, is a cron job somewhere living on a job server where sometimes that job server fails—did it run? Didn’t it run?—being able to effectively shove that into a Lambda function almost feels like step one, where people start to get their feet wet. Does that align with how you approached it? Or did you come at it from a different angle?

Alan: I guess we came at it from a slightly different angle because we were looking to build applications on top of—so we have APIs that have a Lambda back-end; they’re usually, sort of, fairly dumb Lambdas. They’ve passed through to some third-party back-end somewhere, you know, doing a bit of mapping or transformation, a little bit of business logic, perhaps. But the bulk of it is just listening to a web frontend, calls a back-end API that has some Lambda components.

Corey: That’s functionally how I got into it at first, as well. I started off with a shell script that was great and all, to generate my weekly newsletter, and it was, yeah, what happens if I’m not in front of a box that has a console? I wanted to turn into something that vaguely looked like a web app, ideally that I could use from an iPad out on the road. And in time, it was, “Huh, maybe a Lambda function becomes the right answer here.” And it almost certainly wasn’t, but the way I misused it, it became the right answer and it sort of snowballed away from there.

But what you say about feeling like you’re doing it not very well absolutely resonates. Because the entire time I was doing this, I’m sitting here with the, “I am really not using this in the way it was imagined being used.” But I’ve never met anyone who feels differently.

Alan: Yeah, I think the range of different frameworks, as well, is quite daunting. So, you can either use AWS’ SAM framework, or you can just use the APIs directly, or there’s now the serverless framework, or there’s Architects, or a number of different other—

Corey: The CDK, Chalice, Apex, you can keep going on, and on, and on—

Alan: Exactly.

Corey: —and the problem is, if you start following random blog posts, suddenly it’s, “Oh. They’re using a different one. I should start over.” Or you wind up with this ridiculous combination of 15 different things that you’ve gathered from various places, and, “What on holy horror have I built? Well, [laugh] it’s something that works, so here we are.”

Alan: Absolutely. I think we sometimes miss the opportunity to really use some of the power of AWS, so it kind of frustrates me that—because we’ve always done relational databases, we don’t really have the confidence to use DynamoDB, for example, or if we do try and use it, then we kind of hurry back into our comfort zone when things get hard. So, that kind of frustrates me a little bit with serverless applications.

Corey: I can’t shake the feeling that everyone gets confused sooner or later. Speaking of things that are confusing and everyone tends to go in strange directions with, let’s talk about everyone’s favorite tech buzzword in the enterprise world ‘Kubernetes.’ You playing with that at all? Have any strong opinions one way or the other on it?

Alan: Funny you should ask. Yes so, I’ve been actually trying to promote the idea of using Kubernetes in Hitachi. And this may sound odd as well, but as I said before, we’ve got a number of commercial off-the-shelf systems that we use. They’re way too old to actually have heard about Kubernetes. But sometimes you just need somewhere to deploy something.

You know, not everything will run in a serverless platform. Sometimes I just want to deploy a dashboard somewhere or make this application that you can easily download run somewhere in our infrastructure or that includes cloud as well. And I think in that sense, Kubernetes would be a really good fit because we could have a common platform. As developers, they don’t need to care where it’s running, they just point to an API, and then voila, they can see an endpoint and access it without needing to provision hardware, get it certified by InfoSec, and all of that fun stuff.

Corey: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the Enterprise (not the starship). On-prem security doesn’t translate well to cloud or multi-cloud environments, and that’s not even counting IoT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IoT devices, detects these threats up to 35 percent faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at extrahop.com/trial.

Corey: I tend to shriek and something approaching horror whenever I see a Kubernetes deployment beginning because invariably it’s aimed at something that is a ‘Hello World’ style of deployment. And it’s, “Oh, dear lord, this is overwrought for what you’re trying to deploy.” The other side of it, though, is that the applications that people are looking to migrate into a Kubernetes environment are also a lot bigger than the Hello World examples that folks use. Are you already running these apps in containers and looking for a different orchestrator, or is containerization part of the Kubernetes experience as you’re approaching it?

Alan: So, we’ve wanted to run a few things in containers for a while, but they’ve never really sat anywhere, so we’ve just got a few services running just in Docker, and if they fail, then well, who knows? Someone will have to go and restart it when they find out. Whereas the behemoth that is Kubernetes which obviously sought that out for us, and hopefully, let the support engineers sleep better at night, as long as they don’t have to manage too much of the actual hardware parts of it. So, as I indicated, that’s not where my skills lie; I just let someone else deal with that. And I believe that’s probably where most of the complexity of Kubernetes is. So, maybe that explains my viewpoints.

Corey: I think in the aggregate, my take on it is that if you’re building something greenfield or vaguely close to greenfield, serverless feels like the way to go. But if you’re looking at migrating something pre-existing, then Kubernetes slash containers seems like the path that is the least disruptive to the organization. First, do you agree with that? And secondly, do you think that optimizing for not disrupting the organization is the right path?

Alan: I mostly agree with that. So, I would just add to that, that some things you can’t run in a serverless environment. I know that AWS have managed Grafana and managed Prometheus, but before then, if you wanted to deploy those sorts of services, you had to, sort of, roll it yourself or run some crazy infrastructure to do that. Or just deploy it with a Helm Chart into Kubernetes. So, I think there is going to be a place to package applications that don’t necessarily run in a service or as a service. And sometimes you don’t want to buy things in the cloud as a service. Sometimes it’s more convenient to run them yourselves in your own environment. And I think that’s probably where Kubernetes fits in.

Corey: I think that the cloud providers would shriek and scream at the idea of, “Oh, there’s a workload you think isn’t appropriate for the cloud? Well, you’re wrong, based upon nothing other than what you just said.” And that’s not a [laugh] helpful sentiment. “Cool. I’m not just going to grab petabytes of data and shove it on up there.” Or, “I’m not going to just migrate an entire massive fleet overnight.”

It doesn’t work that way. So yeah, I think there are a number of workloads that, based upon business constraints—internal and external—that you approach from that perspective. I think that there’s tremendous validity to having workloads that are at least capable of running in two places at once, at least during the transition, if not beyond. Now, every time you add a provider to that, it feels like you’re inherently limiting the capability story, whereas if you have something on-premises, great. You probably don’t have an object store on-premises that reacts in quite the same way that S3 does, to give an easy example. Does that align with your understanding, or do you feel differently about it?

Alan: Yeah, sure. So, as you say, we’re not going to have unlimited petabytes of storage in our data center, ever—

Corey: Well, you will, but your budget owner is going to have a minor heart attack [laugh] as soon as you try it the first time.

Alan: [laugh]. Yeah, exactly. So, I think our strategy, Hitachi is cloud-first, but some things we know aren’t suitable for that, so take a look at where it’s most suited to run.

Corey: So, looking at how you’ve approached it, is there messaging or storytelling that would have made it easier to wind up heading in this direction sooner? Was it a matter of waiting for certain levels of maturity from the provider perspective? Or was it entirely from an internal point of view where there are enough stakeholders that need to get on board with alternate ways of doing things before you can see any real traction?

Alan: So, I think it’s mainly a case of education. So yeah, definitely getting the stakeholders in the company bought into this crazy new way of working, so having this massive AWS bill land every month and not really being sure who’s paying for it, but it gets paid anyway. And also in terms of the developers, obviously, writing their code in a way that’s appropriate to deploy to the cloud, and also the support staff who monitor and look after these applications.

Corey: So, it’s easy for me to sit here and say, “Oh. You should just go ahead and do X, Y, and Z if you want to wind up exploring down this particular path.” But I don’t work at a large enterprise or even a medium enterprise as you describe it, which is still orders of magnitude bigger than my version of big company. So, what are the biggest stumbling blocks that you’ve encountered going down this path, and how would you address them differently if you had to do all over again?

Alan: Okay. So, one of the biggest issues we’ve had for years is the lack of cloud connectivity. That’s now being addressed, but we’ve had to work around that limitation in quite a few ways. And there’s been various reasons for that. Network infrastructure has changed a bit over the years.

But the lack of cloud connectivity was really a problem in the early days when we were trying to make useful systems run outside of our network, but still, be secure and communicate with the back-end services that we needed in our network.

Corey: So, if you could talk to other folks who are where you were when you started this journey, are there any resources you’d point them to, as far as something that they should understand, something that they should absorb, or just something to help point the way?

Alan: Yeah, okay. So, we’re really trying to change into more of a DevOps culture at Hitachi, and I’m really passionate about how we go about this. And as such, I really love the book Accelerate—which have you had the author on the show? I forget.

Corey: Oh, yes. Dr. Forsgren has been a recurring guest on the show.

Alan: Yeah, I thought so.

Corey: And she’s just spectacular at answering the questions that I wish I had been smart enough to ask in the first place.

Alan: Excellent. Yeah. So, they point out the four key metrics for a high-performing culture, and so I’m always trying to bring those up in team meetings and to the leadership of Hitachi. So: deployment frequency, lead time, change fail percentage, and meantime to restore. So definitely, that book, I think, is really key to understanding why organizations need to look at those metrics, as opposed to measuring lines of code, or hours worked, or anything like that, in terms of getting a high-performance culture.

Corey: And we’ll, of course, throw a link to the Accelerate book in the [show notes 00:22:29] because we are always fans of what Dr. Forsgren is up to, and any chance we get to wind up putting her in front of more people, we’ll take it.

Alan: [laugh]. Great.

Corey: So, I want to thank you for taking the time to speak with me in as much depth as you have. If people want to wind up learning more about what you’re up to, challenges you’re facing, how you’re overcoming them, okay can they find you?

Alan: So, I occasionally Tweet. I’m @alanraison. I’m on the GitHub, also alanraison there. That’s probably about it.

Corey: Excellent. Thank you so much for taking the time to speak with me today. I appreciate it.

Alan: No problem. It’s been a pleasure.

Corey: Alan Raison, DevOps lead at Hitachi Capital UK. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you hated this podcast, please leave a five-star review on your podcast platform of choice, along with a comment giving me feedback the insulting way.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Pete

Pete is a recovering system administrator who got his start with AWS services back in 2009 while at Sonian, the first cloud-based email archiving platform. As one of the earliest and largest users of AWS, Pete ran technical operations and brought DevOps theory into action. Pete has worked for other companies such as Dyn, Threat Stack, and CHAOSSEARCH, managing large scale AWS deployments. A frequent speaker at DevOps and Observability events, Pete brings a product mindset to SaaS operations. Outside of work he spends his free time smoking meats and tweeting about the results.

Links:

  • Twitter: https://twitter.com/petecheslock

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Join me on April 22nd at 1 PM ET for a webcast on Cloud & Kubernetes Failures & Successes in a Multi-everything World. I'll be joined by Fairwinds President Kendall Miller and their Solution Architect, Ivan Fetch. We’ll discuss the importance of gaining visibility into this multi-everything cloud native world. For more info and to register visit www.fairwinds.com/corey.

Corey: The apps on cloud summit is a new action packed, not a conference, happening May 11th through 13th online. Its for everyone who makes applications in the cloud run screaming. From IT leaders to DevOps pros to you folks, whoever you might be. Take a break from screaming into the cloudy void with me to learn from some of the best of people who actually know what they’re doing. Like Kelsey Hightower, AWS blogger John Meyer, and also me, because apparently they didn’t listen to me saying I had no idea what I was doing. Register now at turbonomic.com/screaming. Theres a “swag box” ready to ship for the first two thousand registrants, so you don’t want to miss this. Thanks for Turbonomic for sponsoring this ridiculous podcast.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. For the third year in a row, I am joined by—who is now my colleague, Pete Cheslock. Pete, thanks for coming back.

Pete: It’s great to be here yet again, although under different circumstances than our normal post re:Invent extravaganza.

Corey: Oh, yes. So, every year, for those who have not been following this show since its inception, Pete and I get together to more or less kibitz around what happened at re:Invent. We would have done this in December, but then they put up on their website, three more days happening in January, and where, we’ll wait until after that happens. And guess what happened? Nothing. They did some additional breakout sessions and that’s it. So honestly, it was a giant waste of everyone’s time, kind of like the, you know, sponsor expo hall at a digital event.

Pete: Yeah. Can we start on that topic?

Corey: Do you want to start on the digital event aspect or the crappy expo hall?

Pete: I want to start with the expo hall because I am a former sponsor of re:Invent, many, many times over. Again, I’ve been very lucky. I’ve worked in the cloud world, so I’ve been to pretty much all the re:Invents except for one—I’m not sure we’re going to count this year—but for almost all of those re:Invents, I was in the expo hall as part of a sponsor, whether it was the company I was working at, or whatever, but we were either part of that process of setting up a booth and shilling our wares to all these tech folks that are there, and the experience has been different: everything from, like, you build your own booth; here’s a square, just put whatever you want there, which I think was hilarious in the early days. To the nope, just give us a picture of what you want behind your booth. This year, though, it’s a weird digital thing. I guess they sent some VR things around there. I’m not sure if you heard of that. Some of the—

Corey: Oh, they did. Only for the Heroes, not for anything sponsor expo hall stuff. Now, let’s begin and say that I have a fair bit of sympathy—kind of—because 2020, weird year, pandemic is not something you generally plan for when scheduling events out years in advance. And there we have it, where it’s, suddenly there’s really no other option. Now, AWS absolutely dragged its feet embarrassingly long before announcing it would be digital-only. I think it was October, when they finally said, “All right,”—or damn near it—“All right, it’s going to be a virtual event.” To which the rest of the industry said, “No kidding.” But until they actually come out and say that you can’t bank on it.

Pete: How do you plan for that? I mean, I remember just a few years ago going through the process while I was still at ChaosSearch to set up a sponsorship for re:Invent. People start that process in March, April timeframe; they start thinking about, you know, the strategizing it; they start locking in their deposits because the sooner you get a deposit in the earlier, you can pick your booth location and that’s a big part. You want it to be not in the way far back that no one can find you. You want it to be somewhere near a good walkway. And so there’s a lot of planning that goes in, so pushing it back so late. I mean, I don’t know, do you think that they had this belief that they were going to do an in-person event?

Corey: My honest belief, to be very frank—and I say this in my capacity as gadfly, not in my capacity as self-appointed head of marketing for AWS—is that I think that they just were facing a whole bunch of cancellation fees, because when you book something like that, it’s expensive. They’re frugal, and I feel like for once, they were on the side of a contract that there was no winning move to get out of. And I feel like there was just some bitterness around that, they were hoping for a miracle and finally had to face reality. That’s my gut feeling. I have no inside track on that because although I call myself the head of AWS marketing, they don’t agree. Which is fine, though because, given their messaging or lack of same, I don’t need their agreement to be effective in the role.

Pete: Exactly. I mean, I think what’s most impressive to see is that they still provided an event that people attended; people watched the videos, they took part in, they did a bunch of different, kind of, vehicles and using things like Twitch and this VR thing for the expo hall, as a little silly as it is, at least they tried, right? They tried something new in this new world that we have been living in.

Corey: It was definitely an experience. Because we’re starting with the expo hall, though, I did a virtual walkthrough of the expo hall, and I made fun of things as I typically do in the real one, but it was hard to find, a bunch of people didn’t realize it existed, three weeks in, and what shocked me was the sheer level of enthusiastic outreach I got from some of the sponsor booths I visited, I got phone calls from vendors ask if they can help me with various solutions. “No, no. I was just there to make fun of you.” At which point it’s a very surreal conversation for an account rep to have when they don’t know who I am and what my nonsense looks like. But it was, they had so few leads coming in that they were just really focusing on every one that showed up. And I feel for them.

Pete: Yeah.

Corey: Problem is that they paid top dollar for these things, and got—I got to be honest with you—remarkably little. The minimum buy was something like 35 grand for a tiny little booth, and it went up to 125 plus a whole bunch of extras. And I’m looking at this in my own re:Quinnvent sponsorship nonsense that worked super well for sponsors, and I’m sitting here going, “I really need to start charging more. My God.”

Pete: Yeah. We just think about that for a second is that in normal re:Invent, your booth size is priced differently. If you want a small booth, like, 10 by 10 foot, you’ll pay a certain amount of money, a 20 by 20 booth is a lot more money. It makes sense. It’s a square footage situation here, but we’re talking about, like, computer bits, right? They’re like, yeah, well, you can get the small digital booth for 30,000 or the large, double-decker digital booth for 150.

Corey: Oh, yeah. One of my favorite personal experiences, remember, I talked to a company about sponsoring. Their initial position is usually always the same. “You’re a jerk. You made fun of us on Twitter, you roasted us, why on earth would we ever pay you for sponsorships?”

And my response was, “Look, let me level it with you here. I get up there and make fun of you, and no one in the world is going to stop doing business with you because I said something snarky and sarcastic, but they absolutely will hear of you for the first time.” And then the penny drops. And it goes one of two ways. It’s either, “You’re an ass. No.” Or it’s, “Oh, my god, you’re right. Have some money.”

And it’s really an interesting experience watching that transformation take place. I used to think that I was, I don’t know, somehow fooling people with this. I’m not. It has the benefit of being completely true.

Pete: Yeah, I mean, sometimes all news is good news. Anything that you say, you’re just going to make a joke about someone’s product and—or their marketing strategy because every marketing strategy is a little stupid from time-to-time—but you look at and you go, “Yeah, this is pretty stupid.” And then you say that to other people, and they’re like, “Oh, I’ve never heard that company. What do they do?” And at the very least, it’s that opportunity for them to go to the site and be like, “I don’t know what that company does. Let me go look at it.”

And sure, I may spend about five seconds on your site, but you’ve got that free impression; you’ve got that opportunity. Like, I hope your site is good enough to have a clear statement of what you are because I’m going to give you about five seconds, but still, just mentioning it, I’m going to be like, “Oh, who are they? Let me go check it out.”

Corey: Well, the expo hall was on the other side of that. They had these mini-sites built up. My personal favorite was in week three, I went to a vendor and clicked on their mini-site, and it 404’ed because they were hosting, like, some virtual drink-up the week before, and when it was done, they just packed up their booth and left. You spent an awful lot of money not to drive me to your website. I just don’t pretend to understand marketing, at least that’s what I thought. It turns out that no, just some companies are really bad at it.

Pete: Yeah. And—

Corey: What I do is not, oh, for everyone. I get that. But it makes an otherwise dry subject area kind of fun. At least to me.

Pete: Yeah, I’ve always—I’m a weird person. So, I enjoy re:Invent most mostly because—re:Invent is what you make of it. It’s always been that case, even the very first re:Invent, which feels quaint by comparison to some of the more recent ones, where I think there was maybe 4000 people, 6000 people the first one? I’d love to find the answer to that one. It was small, though. It was very small.

And moving on to the more recent years of re:Invent how, just, big they’ve gotten. I’ve always liked the expo hall. I’ve liked walking around. I mean, granted, being in this industry, oftentimes many of my friends are working the booths as well. Like, I will be working the booth for a while, people will come by, and my friends will be working in a booth and I’ll stop by and say hello, but mostly it’s a really great opportunity to just see what’s out there.

There is so much stuff. It’s hard to follow everything around. I mean just following—right, Corey?—following just the Amazon ecosystem as a full-time job. Following the surround sound, it’s nearly impossible.

So, I’ve always liked the expo hall. I like walking around. I like to see what people are saying. I like to see things that are—what are they doing? What problems are they solving? And it’s a great way to just get that, kind of, streamlined process, you get to see so much in such a short amount of time.

Corey: It’s absolutely one of my favorite parts of the show is walk around the expo hall. It’s a natural meeting place for people. I’ve got to be honest, I don’t go to too many sessions just due to the fact that there’s better uses of my time than standing in line for two hours to make sure I get a seat. So, it’s a natural gathering point; it’s a great way to catch up with people you only get to see once a year. And you get to see what the zeitgeist is, what people are talking about.

You can see whose booth is slammed and who’s not. And you can fool yourself into thinking it’s about the quality of their product rather than the quality of their swag. But it’s an experience. And on the one hand, I’m sad to miss it. On the other, I had so many more productive conversations this year.

There were so many aspects of it that I don’t miss, like the conference crud where you get the flu every time, because you have a lot of people in a small space. And there’s the sense of having to fit everything into one week. Instead, they’ve now expanded it to three weeks, which on the one hand, okay, I actually like the fact that can be a more measured pace. On the other, exactly who do they believe can get three full weeks off from work to sit around and watch videos in a browser?

Pete: Yeah. That was a big concern when I started looking at some of the re:Invent stuff. And even for us, we follow a lot of the Amazon ecosystem, I didn’t even get to take part of a lot of this Amazon re:Invent activities as much as I wanted to. Because we have clients that we want to service and make sure that their needs are met. So, just for us who spend so much time in this ecosystem, we really can only dedicate so many people to understanding what all these changes are and staying up on it.

Corey: I will say that the one thing that tips my entire assessment event over into the positive, and I want to see aspects of this going forward, is how accessible the whole thing became where we suddenly have a scenario where it’s not just restricted to people who, one, can drop two grand—or damn near it—on a ticket; two, can afford to travel to and stay within Las Vegas for a week, and three, can get the time off from work to do it. Suddenly, the only prerequisite was ‘has an internet connection.’

Pete: [laugh]. It’s so true. I mean, let’s not kid anyone out there. It’s a boondoggle. It is one hundred percent a boondoggle. It is a week in Vegas. And I know there’s a bunch of people out there that are like, “Aww, I hate going to re:Invent. I hate Vegas.” And yeah, I can understand that some people just don’t enjoy it. I’m a weird person. I actually like Vegas. I’m weird.

But people go and they enjoy it because even if you hate all of the noise, and the smoke, and the gambling, and the whatever, and you have to walk an hour to get any place, the people there, though, and the connections that you can make—I mean, I meet new people every year at re:Invent which, at an event that is so large, kind of feels counterintuitive. Like it’s so large, you almost feel lost in a sea of people, but yet somehow it’s like, I still run into people. I meet up with someone for coffee, and they may be back to back with meetings, “Oh, do you know this person?” I get to meet new Amazon folks. And I get to actually see a lot of friends, and again, maybe we’re all just missing that personal connections of conferences. But on the flip side, I really enjoyed not having to go to Vegas this year. Even without a pandemic, it was really nice to not have to spend a week and come home sick, and tired, and exhausted.

Corey: There’s something amazing about being able to do it at your own pace. Now again, because Amazon is willing to be misunderstood for long periods of time, which is used as an excuse to completely abdicate any actual marketing work, for the most part, it means that I instead had to guess what was going to happen going in. So, all right, I committed to doing a daily email roundup four days of the week; I committed to doing a bunch of live streams and such. And what I didn’t realize was it was going to be a stop-start thing where there would be a whole bunch of releases one day, and then there’d be nothing the following day. So, I was sort of left on some of those empty days of kind of holding the bag of, “Here’s something you might not have caught yesterday”—and I’m digging deep into the barrels—“There was a minor change to an SDK in a language no one has ever heard of.”

And the fun, nice thing is about AWS is that people just assume someone’s going to care about that. They won’t, but, “Oh, that one just must not be for me.” So, it was a little challenging from a content management perspective. I would absolutely want to see that done differently or at least telegraphed in advance next year.

Pete: So, this is an interesting topic where, obviously, the last 12 months has been interesting for the conference world. I know a lot of more local conferences are, kind of, already writing off the year. Maybe some are holding off to see if, maybe, the end of the year, we reach enough vaccines and things get better. But I mean, Amazon had re:Invent. It was small, then it expanded. Then they broke out these summits, right? They had these city regional summits, then they did like—

Corey: And having gone to a bunch of those summits, it turns out, they’re all basically kind of the same thing. And I went to my third one back in 2019 or so, that year, and it was, “Hey, that’s the same joke in all the keynotes. What’s the”—and then it’s like, “Oh, right. Most people are not, you know, nuts, and don’t travel around the world, like some sort of ridiculous groupie for a rock band going to all of the AWS summits.” I have problems. But it’s fun.

Pete: Yeah, again, as someone who has worked in a lot of these booths and had to go to them, it’s always a little painful when you’re at the Santa Clara Convention Center—which, you know, is in the middle of nowhere, there’s nothing around, there’s nothing to do—and there’s like a Bennigan’s, I think you can go get some dinner at when you’re done for the day. And you finish up and you’re chatting with folks, and then you’re like, “Oh, so you get to be at the Toronto one?” Be like, “Oh, yeah, I’ll be there. I’ll see you in a few weeks.” Like, [laugh] it’s just, it’s the saddest thing.

Corey: Yeah, it really is. Then they also expanded beyond summits, though into re:Inforce, which is security. And that was a summer event in 2019 and they were hoping to do in 2020 and canceled it. And supposedly, it’s going to be coming back; we don’t know when. I like the idea of being able to break out the security-focused stuff into its own event because the biggest problem I’ve always had with re:Invent is that it doesn’t know what it wants to be.

Is it a thing that’s for new product releases? Is it a chance to have a bunch of executive briefings between big customer execs and Amazon folks? It’s a vendor expo hall where people get to shill their wares? Is it a partner gathering so the partners can learn how things work? What’s going on there? And what is it about? And is it a big party? Is it just, effectively, a chance for a bunch of people to get on stage and talk about what they’re working on? And the answer to all that is, “Yes. And more.”

Pete: Yeah, exactly. There is an identity problem for re:Invent. I think they have been doing a decent enough job of trying to break things out, to make, kind of, that re:Invent knowledge transfer, a little bit more accessible with—they’re all free, too, these regional summits, these are free events people can just go to and consume a lot of this content and workshops and things like that. The breaking out of the security stuff, I think was great. Hilariously, I had not been working at a security company when that happened, but I had heard rumors that the sponsorships were all invite-only. There are so many security companies that Amazon was like, “We can’t take all of your money for this, so we’re going to invite you in to sponsor it.” Because that market is so thirsty for sales.

Corey: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the Enterprise (not the starship). On-prem security doesn’t translate well to cloud or multi-cloud environments, and that’s not even counting IoT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IoT devices, detects these threats up to 35 percent faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at extrahop.com/trial.

Corey: What was also bizarre was that at re:Inforce’s expo hall, it was only security companies. And that was something I didn’t fully understand until you just said that. Because I was annoyed that, well, there’s no monitoring companies here or anything else. Does it not occur to folks that may be people who care about AWS security might also have other needs in the cloud computing space? There aren’t too many people who have blinders on that restrict them specifically to security and only security. “Oh, it does monitoring, too? I’m not interested.”

Pete: [laugh]. I think that’s got to be a hard challenge as a security vendor if you’re at an event—and I’ve never been to RSA, and RSA must be just the exact same way—but how many different companies there do, like, threat detection or, like, sims—

Corey: Most of them. It’s the same product with different logos on, as you walk up and down the halls. There are occasionally unique and interesting things, but they’re few and far between.

Pete: Yeah, exactly. So, I’m most curious to see where Amazon goes with re:Invent because I think what this year is giving them is an opportunity to maybe find a better identity for what do they want re:Invent to be. And look, if the answer is we want it to be a big celebration of everyone who uses this, whose job is impacted by it, who makes money off of it, who et cetera, et cetera, then great. Then that’s the event. But there are going to be people, like there might be partner network people, well, we don’t really want to go to the party, but we want to get all the updates because our businesses are fully dependent on this. So, then maybe, do they break that up? So, I don’t know, do they start breaking re:Invent out into more focused things like they did with security? Would it take away from the overall feeling? I don’t know.

Corey: It’s a weird problem, and I don’t know how to solve it because without understanding what the event is intended to be, is really hard to guide it. They get on stage and they say, “Re:Invent is really—it’s not a sales conference. It’s not a release confer—it’s an education conference.” Which is shorthand for we have absolutely no idea what this is.

Pete: It’s so true. I mean, the early re:Invents, we’re, “All right. Let’s hear the latest price cuts for Amazon. Tell us how much cheaper S3 is and tell us how much cheaper EC2 is.” It was like clockwork.

That was what it was about. And it was about new software releases. Weirdly enough, though, the software announcements they did back then were all things that you could get then. They were like, “Hey, we’re announcing these new instances, available today.” They almost didn’t have to say available today because like, of course. Why would you announce it if you can’t get it?

Corey: Right, it seems like half the releases that there are big headlines, was, “Available in preview.” And you know what that means. They’re setting it up for them to just get dragged whenever they pull a Timestream. And, “Yes, we’re announcing it in private preview. It will be available soon.”

Then two years go by. And it’s pretty clear when you see that kind of delay that something happened. And credit we’re due, because they’re Amazon, I prefer that they get it right before launching it rather than launching something that isn’t great, and then we’re stuck with it forever. But if you’re still that early on, don’t announce it. Announce things that aren’t vapor.

Pete: Yeah. And I’d like to think there’s a strategy to it, but there probably isn’t, it’s probably just more of a, “Well, this is a great way for us to identify other customers who might also want to use this and maybe they want to be part of the preview.” And I don’t know, it’s a little frustrating, I will say. But things come out, and you’re like, “I really want to use that.” And it’s like, “Yeah, no. That’s not for you.”

Corey: Yeah. And that’s the challenge, too I think, from a marketing perspective. They do so many different releases of what’s coming out, how they’re going to be talking to satellites in orbit, or talking to manufacturing floors, or whatever it is that they’re talking about, that every company, no matter who they are or what they do, look at that and think, “Huh. That’s not what we do, therefore, AWS is not for me.” And AWS, as a company, is basically an alien organism, compared to going down the path of any other company that can’t really walk and chew gum at the same time.

I don’t know, if Apple, for example, starts doing a big push into filling potholes as a primary function, I’d look at that and think, “Oh, okay. They’re pretty clearly not focused on the Mac, on some respect.” I mean, look, what they did to their laptops for years with the Thunderbolt, the touch bar, the crappy keyboards, et cetera. It’s, yeah, it’s clear that they can’t focus on the iPhone and the Mac at the same time effectively. Or they just hate their customers.

Don’t email me. But what I don’t see is any ability of companies other than Amazon, to be able to focus and execute across this many different things. It’s hard to contextualize. So, it’s very easy for the messaging takeaway to be, “It’s not for me.”

Pete: Yeah. And maybe that ties into the re:Invent identity problem because you’ve got Andy Jassy on stage talking about, look out for whatever this new service that’s going to ruin your Amazon bill. It’s like warehouse logistics stuff, which, yeah, cool. I’m sure that that solves a big problem in the industry, but I’m a DevOps engineer, and I want to hear more about EKS, right? And I have to sit through learning about this predictive, whatever for my warehouse that I don’t have. Does that just become too off-putting, and do I just then zone out and, kind of, ignore all these other interesting things that could be happening?

Corey: It’s unclear. And that’s the biggest problem, I think, that they’re failing to educate people on. Specifically, every service is for someone. No service is for everyone. And that is a difficult thing to hold on to.

We’ve long since passed the point where anyone can hold all the services in their head, we’ve gotten to a point where even I don’t always pick up a fake service someone slips in to see if I know if it exists or not. It’s expanded too far too quickly and where, that’s fine. But the messaging strategy has to change, the marketing strategy has to change, your entire go-to-market has to change.

Pete: Yeah. Amazon is really good at running things. I mean, that’s what they’re good at. It’s operationalizing software, and they continue to find things that people don’t want to run anymore. I don’t blame them.

I don’t want to run things. I’m a cloud economist. I look at bills. I don’t want to run Elasticsearch anymore. I don’t want to deal with Cassandra and get paged at 2 a.m. like, I really want someone else to deal with that stuff. And just think about all of the other verticals, all the other businesses that exist out there with the same people who are having the same complaints, just insert different words. Like, “Oh, I really hate my business intelligence solution. I really hate these Excel spreadsheets that are always locked. There must be a better way.” Right? And it turns out Amazon’s like, “Yeah I got you.”

Corey: Yeah, the idea that Amazon is equally good across all of these different offerings is a bit of a red herring. There are things that they excel at, and they’re things that they struggle at. And I often shorthand that to the infrastructure pieces, the plumbing, they’re phenomenal at. Anything that requires a user interface, or is SaaS, they are hilariously bad at—Honeycode—and most things are somewhere on the spectrum between those two points. And there are exceptions in both directions, but by and large, the more it looks like a big computer rented by the hour, the better the offering is. Would you agree or disagree with that?

Pete: Yeah, I definitely agree. I think the thing that was most surprising in the recent Kinesis outage was just how intertwined Amazon services are internally—AWS services—and how internally, the engineers at AWS are building on top of AWS. It’s a weird Russian nesting doll issue, where it’s just turtles on turtles. And it’s fascinating, and I wonder if the services which are most used internally, as well become those services that are the most stable, the most well-supported, most features coming in? Does Amazon build for Amazon first, and therefore, put a lot of effort into those things that further supports their business and maybe grows revenue? Maybe.

If they’re building for the customer, and they consider themselves a big customer, then theoretically, they’re building as well for some of those features. But of course, they build for everyone; they build for the startup that has two EC2 servers, and then they build for the federal government who wants to beam some bits using Ground Station. Like, who else is going to use that service?

Corey: Yeah. It feels like there’s like five companies out there that might need it, and the rest of us are, “Yeah, I don’t currently have any satellites in orbit this quarter that I need to speak to, so it’s probably not for me.” I will say that every time I meet someone who’s about to go to AWS as an employee, they’re super excited because they’re going to see how it works internally and come out understanding of this Google-like system that is decades ahead of anything else on how they run their stuff operationally. And then a few months go by, and I catch up with them again, and they look haunted. There is no enthusiasm for it at all. Their voice shakes; they tremble a bit; frequently, they’ve developed the drinking problem. And they don’t ever talk about it, but what I’ve managed to piece together is there’s no magic secret sauce. It’s the same nonsense that you would see anywhere else, but they excel at the operational aspects of all of it. And that’s what makes it work.

Pete: I think what actually happens is they find the truth of the m1.medium and the first EC2 instance, and they’re so horrified that those are still running, that they can just not come back from the brink.

Corey: I didn’t know it was a Raspberry Pi.

Pete: [laugh]. I think you are totally right. I mean, every place, largely, is the same. It’s just, it’s got its own history that has framed how everything is. And if you’re on the inside, and especially if you’ve been at Amazon for many years, you’re financially incentivized to love that place.

I mean, if I had a lot of stock grants that were granted many years ago, and the stock continues to climb, like yeah, this is the best place I’ve ever been like, “What are you talking about?” As they look around and everything is on fire, or who knows what.

Corey: As they walk past the conference room filled with people crying? Yeah.

Pete: You know, it’s… I don’t know, what is it? Stockholm Syndrome? Is that the term? You just accept it, and you get used to it, and you get comfortable with it. And yeah, in rare scenarios, there’s folks that I know that have been there for a long time, and they’re there—I hate to say they’re there for the mission. They’re not. They’re there because of the challenge; because technically what they get to work on is cutting edge.

I don’t know if that’s the case for every new service and feature. If you were someone who was working on making QuickSite graphs look better versus, like, the Nitro Hypervisor, maybe depending on who you are, one of those is more thrilling than the others. I don’t really know. But obviously, there seems to be two types of Amazon employee: one who sticks around for their year, gets that bonus, or at least doesn’t get the clawback they need; and the others who stick around for many, many, many years, right? It’s, I think, a different company, depending on who you are.

Corey: That is increasingly the vibe I’m getting from feedback I’ve gotten to blog posts, people yelling at me, people saying that, oh, my assessment of how compensation works at Amazon is either spot on or completely inaccurate. And both of those groups are being fully sincere when they say it. But that’s a conversation for another time.

Pete, thank you for joining me. We will do a second episode in the very near future talking about the actual releases of re:Invent 2020, but thank you for joining me. For those who are unfamiliar with your amazing work. Where can they find you?

Pete: You can find me at @petecheslock on Twitter; it’s probably the best place. It is a mixture of smoked meats and technology hot takes.

Corey: Thank you. As always, it’s a pleasure. Pete Cheslock, cloud economist at the Duckbill Group.

Pete: That’s me.

Corey: I’m Cloud Economist Corey Quinn, also at the Duckbill Group, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you hated this podcast, please leave a five-star review on your podcast platform of choice, along with a comment telling me that I really don’t understand Amazon’s marketing approach, and that no, I don’t really run AWS marketing.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Kevin

Kevin Miller is currently the global General Manager for Amazon Simple Storage Service (S3), an object storage service that offers industry-leading scalability, data availability, security, and performance. Prior to this role, Kevin has had multiple leadership roles within AWS, including as the General Manager for Amazon S3 Glacier, Director of Engineering for AWS Virtual Private Cloud, and engineering leader for AWS Virtual Private Network and AWS Direct Connect. Kevin was also Technical Advisor to Charlie Bell, Senior Vice President for AWS Utility Computing. Kevin is a graduate of Carnegie Mellon University with a Bachelor of Science in Computer Science.

Links:

  • AWS S3: https://aws.amazon.com/S3
  • AWS Twitch: https://www.twitch.tv/aws
  • AWS YouTube: https://www.youtube.com/user/AmazonWebServices
  • AWS Pi Week: https://pages.awscloud.com/pi-week-2021.html

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Join me on April 22nd at 1 PM ET for a webcast on Cloud & Kubernetes Failures & Successes in a Multi-everything World. I'll be joined by Fairwinds President Kendall Miller and their Solution Architect, Ivan Fetch. We’ll discuss the importance of gaining visibility into this multi-everything cloud native world. For more info and to register visit www.fairwinds.com/corey.

Corey: If your mean time to WTF for a security alert is more than a minute, it's time to look at Lacework. Lacework will help you get your security act together for everything from compliance service configurations to container app relationships, all without the need for PhDs in AWS to write the rules. If you're building a secure business on AWS with compliance requirements, you don't really have time to choose between antivirus or firewall companies to help you secure your stack. That's why Lacework is built from the ground up for the Cloud: low effort, high visibility and detection. To learn more, visit www.lacework.com.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’m joined this week by Kevin Miller, who’s currently the general manager for Amazon S3 which presumably needs no introduction itself, but there’s always someone. Kevin, welcome to the show. Thanks for joining us, and what is S3?

Kevin: Well, Corey, thanks for having me. Yes, Amazon S3 was actually the first generally available AWS service. We actually just celebrated our 15-year anniversary here on Pi Day, 3/14. And S3 is an object storage service that makes it easy for customers to put and store any amount of data that they want. We operate in all AWS regions worldwide, and we have a number of features to help customers manage their storage at scale because scalability is really one of the core building blocks, tenets for S3, where we provide the ability for customers to scale up and scale down the amount of storage they use, they don’t have to pre-provision storage, and when they delete objects that they don’t need, they stopped paying for them immediately.

So, we just make it easy for customers to store whenever they need, access it from applications, whether those are applications running in AWS or somewhere else on the internet, and really just want to make it super easy for customers to build storage, use storage with their applications.

Corey: So, a previous guest in, I say the first quarter of the show’s life—as of this time—was Mai-Lan Tomsen Bukovec, who at the time was also the general manager of S3, and she has since ascended to, perhaps, S4 or complex storage service. And you have transitioned from a role where you were the general manager of Glacier—

Kevin: Correct.

Corey: —or Amazon S3 Glacier, and that’s the point of the question. Is Glacier part of S3? Is it something distinct? I know they’re tightly related, but it always seems that it’s almost like the particle-wave experiment in physics where, “is it part of S3 or is it a distinct service?” depends entirely on the angle you’re looking at it through?

Kevin: Right. Well, that’s—Amazon S3 Glacier is a business that we run as a separate business, with a general manager. Joe Fitzgerald looks after that business today. Certainly, most of our customers use Glacier through S3, so they can put data into S3 and they actually can put it directly into the Glacier storage class—or the Glacier Deep Archive storage class—or customers can configure lifecycle policies to move data into Glacier at a certain point. So, the primary interface customers use is through S3, but it is run as a standalone business because there’s just a set of technology and human decisions that need to be made, specific to that type of storage, that archive storage. So, I work very closely with Joe, he and I are peers, but they are run as separate businesses.

Corey: So you, of course, transitioned. You’ve I guess, we’ll say that you’ve thawed. You are no longer the GM of Glacier, you’re now the GM of S3. And you just had a somewhat big announcement to celebrate that 15-year anniversary of S3 Object Lambda.

Kevin: Yes. We’re very excited about S3 Object Lambda. And we’ve spoken to a number of customers who were looking for features with S3, and the way that they described it was that they liked the S3 API, they want to access their data through that standard API, there’s lots of software that knows how to use that including, obviously, the AWS SDK. And so they liked that GET interface to get data out and to put data in, but they wanted a way to change the data a little bit as it was being retrieved. And there’s a lot of use cases for why they wanted to do it.

Everything from redacting certain data to maybe changing the size of an image for particular workloads, or maybe they have a large amount of XML data and for certain applications, they want a JSON formatted input. And so rather than have a lot of complicated business logic to do that, they said, well, why can’t I just put something in the path so that as the data is being retrieved through the GET API, I can make that change, the data can be reformatted.

Corey: It’s similar to the Lambda@Edge approach where instead of having to change or modify the source constantly and have every possible permutation, just operate on the request.

Kevin: Yeah, that’s right. So, I want one copy of my data; I don’t want to have to create lots of derivative copies of it. But I want to be able to make changes to it as it’s going through the APIs. So, that’s what we built is Lambda, it’s integrated with Lambda, it’s full Lambda. So really, it’s pretty powerful.

Customers can do anything you can do in a Lambda function you can do in these functions that are then run. So, an application makes a GET request, that invokes the Lambda function, the function can process the data, and then whatever is returned out is then sent and streamed back to the application. So, customers can build some transformation logic that runs in line with that request, but then transforms that data that goes to applications.

Corey: So, at the time that we’re recording this, the announcement is hours old. This is not something that has had time yet to permeate the ecosystem; people are still working through the various implications of it, so it may very well be that this winds up aging before we can even turn the episode around. But what is the most horrifying use case of this that you’ve seen so far? Because I’m looking at this and I’m thinking, “Oh, you know what I can use this for?” People are thinking, “Oh, a database?” “No, that’s what Route 53 is. Now, I can use S3 as a messaging queue.”

Kevin: Well, possibly. I keep saying that I’m going to use it as a random number generator. But that was—yeah—

Corey: I thought that was the bill.

Kevin: [laugh]. Not quite. We have a lot of use cases that we’re hearing and seeing already in just the first few hours for it. I don’t know that I would call any super-horrifying. But we have everything from what I was saying in terms of redaction and image transformation to one of the things that I think a lot of—will be great will be using it to prepare files for ML training.

I’ve actually done some work with training machine learning models, and oftentimes, there’s just little things you have to tweak in the data. Sometimes you get a row that has an extra piece of data in it that you didn’t expect or it’s missing a field, and that causes the training job to fail. So, just being able to kind of cleanse data and get it ready to feed into an ML training model, that seems like a really interesting use case as well.

Corey: Increasingly, it’s starting to seem like S3’s the biggest challenge over the past 15 years of evolution has been that it was poorly named because it’s EC2 look at this now and come away with the idea that it’s not simple. And if you take a look at what it does, it’s very clearly not. I mean, the idea of having storage that increases linearly, as far as cost goes—you’re billed for what you use, without having to pre-provision a storage appliance at a petabyte at a time and buy a number of shells. “Ooh, if I add one more the vendor discount kicks in, so I may as well over-provision there.” “Oh, we’re running low. Now, we have to panic order and get some more in.” I’ve always said that S3 has infinite storage because it does. It turns out, you folks can provision storage added to S3 faster than I can fill it, I suspect because you just get the drives on Amazon.

Kevin: Well, it’s a little bit more complicated than that. I mean, I think, Corey, that’s a place that you rightly call out. When we say ‘simple storage service,’ although there’s so much functionality in S3 today, I think we go back to some of the core tenets of S3 around the simplicity, and scalability, and resiliency, and those are not easy. There’s a lot of time spent within the team just making sure that we have the capacity, managing the supply chain to a deep level; it’s a little bit harder than just clicking ‘buy now.’ But we have teams that focus on that and do a great job, and also just around looking around corners and identifying how we continue to raise the bar for resiliency, and security, and durability of the service.

So, there’s just, yeah, there’s a lot of work that goes into that. But I do think it goes back to that simplicity of being able to scale up and scale down makes it just really nice to build applications. And now with the ability to build serverless applications where you have you have the ability to put a little code there in the request path so that you don’t have to have complicated business logic in an application. We think that that is still, it’s a simple capability. It goes back to how do we make it EC2 build applications that are integrated with storage?

Corey: Does S3 Object Lambda integrate with all different storage tiers? Is it something that only works on standard? Does it work with infrequent access? Does it work with, for example, the one that still exists, but no one ever talks about: Reduced Redundancy Storage? Does it work with Glacier? Just, it sits there and that thing spins for an awfully long time.

Kevin: It will work with all storage classes, yes. With Glacier you would have to restore an object first and then it would. So, you’d issue the restore initially, although the Lambda function itself could also issue the restore. Then you would most likely then come back for a second request later to retrieve the data from Glacier once it’s been restored. But it does work with S3 Standard, S3 Intelligent Tiering, SIA, and any other storage classes.

Corey: I think my favorite part of all of this is that the interaction model for any code that’s accessing stuff in S3 doesn’t change. It is strictly a talk to the endpoint, make a typical S3 GET and everything that happens on the backend of that is transparent to your application.

Kevin: Exactly. And that was, again, if you go back to the simplicity, how do we make this simple, we said, “Customers love just that simple API.” It’s a GET API, and how do we make it so that that API continues to work, and applications that know how to use a GET, they can continue to use a GET and retrieve the data. But the data will be transformed for them before it comes back.

Corey: Are there any boundaries around what else that Object Lambda is going to be able to talk to? Is it only able to do internal massaging of the data that it sees? Is it going to be able to call out to other services? How extensible is this?

Kevin: The Lambda can do, essentially, whatever a Lambda function can do, including all the different languages. And then also, yeah, it can call out to DynamoDB, for example, if you want to, for example, let’s say you have a CSV file and you want to augment that CSV with an extra piece of data, where you’re looking it up in a DynamoDB table, you can do that. So, you can merge multiple data streams together, you can dip out to an external database to add to that data. It’s pretty flexible there.

Corey: So, at some level, what you’re realistically saying here is that until now, S3 has been able to be configured as a static website hosting facility; now it can also host dynamic websites.

Kevin: Well, S3 Object Lambda today will work with applications that are running within the customer’s account or where they’ve granted access through another account. We don’t support S3 Object Lambda directly as a public website endpoint at this point, so that’s something that we’re definitely listening to feedback from customers on.

Corey: Can I put CloudFront in front of it, and then that can invoke the GET endpoint?

Kevin: Today, you can’t, but that is also something that we’re—we’ve heard from a few use cases. But primarily, the use cases that we’re focused on right now are ones where it’s applications running within the account or within a peer account.

Corey: I was hoping to effectively re-implement WordPress on top of S3. Now, again, not all use cases are valid, or good, or something anyone should do, but that’s most of the ways I tend to approach architecture. I tend to live my life as a warning to others, whenever I get the opportunity.

Kevin: Yeah. [laugh]. I don’t respond to that, Corey. [laugh].

Corey: That’s fine, you don’t need to. So, one thing that was also discussed is that this is the 15-year anniversary, and the service has changed an awful lot during that time. In fact, I will call our, for really no other reason than to be a small petty man, that the very first AWS service in beta was SQS. Someone’s going to win a bar trivia night on that, someday.

Kevin: That’s right.

Corey: But S3 was the first to general availability because obviously, a message queue was needed before storage. And let’s face it, as well, that most people even if they’re not in the space can instinctively wrap their heads around what storage is; a message queue requires a little bit more explanation. But that’s okay, we will do the revisionist history thing, and that’s fine. But it’s evolved beyond that. It had some features that again, are still supported but not advertised.

The Reduced Redundancy Storage is still available, but not talked about. And there’s no economic incentive for doing it, so people should not be using it, I will make that declaration on my part, so you don’t have to. But you can still talk to it using SOAP calls, in the regions where that existed, via XML, which is the One True Data Interchange Format, because I want everyone mad at me. You can still use the, we’ll call it legacy because I don’t believe it’s supporting new regions, the BitTorrent interface for S3 data. A lot of these were really neat when it came out and far future, and they didn’t pan out for one reason or another, but they’re still there. There’s been no change since launch that I’m aware of that suddenly breaks if you’re using S3 and have just gone on walkabout for the last 15 years. Is that correct?

Kevin: You’re right. There’s functionality that we had from early on in S3 that’s still supported. And I think that speaks to the way we think about the service, which is that when a customer starts adopting it, even for features like BitTorrent, which certainly that’s not a feature that is as widely adopted as most of them. But there are customers that use it and so our philosophy is that we continue fully supporting it and helping those customers with that protocol. And if they are looking to do something different, then will help them find a different alternative to it.

But, yeah, the only other thing that I would highlight is just that there have been some changes to the TLS protocols we’ve supported over time, and that’s been something we’ve closely worked with customers to manage those transitions to make sure that we’re hitting the right security benchmarks in terms of the TLS protocol support.

Corey: It’s hard on some level also to talk about S3 without someone going, “Oh, what about that time in 2017 when S3 went down?” Now, I’m going to caveat that before we begin in that, one, it went down in a single region, not globally. To my understanding, the ability to provision new buckets was impacted during the outage, but things hosted elsewhere would have been fine. Everything depends, inherently, on S3 on some level, and that sort of leads to a cascade effect where other things were super wonky for a while. But since then, AWS has been remarkably public about what changed and how things have changed.

I think you mentioned during the keynote at re:Invent, or re:Invent two years ago, that there’s now something like 235 microservices at the time, that power S3 under the hood, which of course, every startup in the world looked at that and said. “Oh, a challenge. We can beat that.” Like they’re somehow Pokemon, and you’ve got to implement at least that many to be a real service. I digress. A lot changed under the hood, to my understanding, almost a complete rewrite, but the customer experience didn’t.

Kevin: Yeah, I think that’s right, Corey. And we are constantly evolving the services that underlie S3. And over the 15 years, that’s been, maybe, the only constant has been the change in the services. And those services change and improve based on the lessons we’ve learned and new bars that we want to hit. And I think one really good example of that it is the launch of S3 Strong Consistency in December of last year. And Strong Consistency, for folks who have used S3 for a long time, that was a very significant change.

Corey: Oh, it was a bi-modal distribution, as far as the response to that. The response was either, “What does that even mean, and why would I care?”

Kevin: Right.

Corey: And the other type of response was people dropping their coffee cup in shock when they heard it.

Kevin: It’s a very significant change. And obviously, we delivered that to all requests, to all buckets was no change to performance and no additional costs. So, it was just something that everyone who uses S3 and—today or in the future—got for free, essentially no additional charge.

Corey: What does Strong Consistency mean, and why is that important, other than as an impressive feat of technical engineering?

Kevin: Right. So, in the original implementation of S3, you could overwrite one object but still receive the initial version of an object in response to a GET request. So, that’s what we call eventual consistency where there can be, generally a short period of time, but some period of time where a subsequent write would not be reflected in a GET request. And so with Strong Consistency, now, the guarantee we provide is that as soon as you receive a 200 response on a PUT request, then all subsequent GET requests and all subsequent LIST requests will include that most recent object version, the most recent version of the data that you’ve provided for that object.

And that’s just an important change because there’s plenty of applications that rely on that idea of I’ve PUT the data and now I’m guaranteed to get the exact data that I’ve PUT in response, versus getting an older version of that data.

Corey: There’s a lot that goes into that, and it’s deceptively complicated because someone thinks about that in the context of a single computer writing to disk—“Well, why is that hard? I edit a file. Then I talk to that file, and my edits are in that file.” Yeah. Distributed systems don’t quite work that way.

And now imagine this at the scale of S3. It was announced in a blog post at the start of this week that 100 trillion objects are stored in S3. That’s something like 16,000 per person alive today. And that is massive. And part of me does wonder how many of those are people doing absolutely horrifying things, but it’s a—customer use cases are weird. There’s no way around that.

Kevin: That’s right. North of 100 trillion objects. I think, actually, 99 trillion are cat pictures that you’ve uploaded, Corey, but—

Corey: Oh, almost certainly. Then I use them as a database. The mood of the cat is how we wind up doing this. It’s not just for sentiment analysis; it’s sentiment-driven.

Kevin: Yeah, that’s right. That’s right. But yes, S3 is a very large distributed system, and so maintaining consistent state across a large distributed system requires very careful protocols. There’s actually, one of the things we talked about this week, that I think it’s pretty interesting about the way that internal engineering in S3 has changed over the last few years, is that we’ve actually been using formal logic and mathematical proofs to actually prove the correctness of our consistency algorithms. So, the team spent a lot of time engineering the consistency services, and all the services that had to change to make consistency work.

Now, there’s a lot of testing that went into it, kind of traditional engineering testing, but then on top of that, we brought in mathematicians, basically, to do formal proofs of the protocols. And they found edge cases. I mean, some of the most esoteric edge cases you can imagine, but—

Corey: But it’s not just startups that are using this stuff, it’s hospitals. Those edge cases need to not exist if you’re going to make guarantees around things like this.

Kevin: That’s right. And you just have to make sure. And it’s hard; they did painstaking work to test, but with our formal logic, we’re able to just to simulate billions of combinations of messages and updates that we’re able to then validate that the correct things are happening relative to consistency. So, there’s a very significant engineering work, it was a multi-year effort, really, to get Strong Consistency to the point it was. But just to go back to your earlier point, that’s just an example of how S3 really has changed under the hood, but the external API, it’s still the external API. So, that’s our north star on all of this work.

Corey: Incidents happen fast, but they don’t come out of nowhere. If they’re watching, your team can catch the sudden shifts in performance, but who has time to constantly check thousands of hosts, services, and containers?

That’s where New Relic Lookout comes in. Part of Full-Stack Observability, it compares current performance to past performance, then displays it in an estate-wide view of your whole system.

Sign up for free at NewRelic.com and start moving faster than ever

Corey: So, you’ve effectively rebuilt the entire car while hurtling down the freeway at 60—or if you’re like me, 85—but it still works the same way. There are some things as a result that you’re not able to change. So, if you woke up, alternate timeline, you knew then what you know now, how would you change the interface? Or what one-way doors did you go through when building S3 early on in its history that in hindsight you would have treated differently?

Kevin: Well, I think that for the customers who used S3 in the very early days, there was an originally this idea that S3 buckets would be global, actually, global in scope. And we realized pretty early on that what we really wanted was regional isolation. And so today, when you create a bucket, you create a bucket in a specific region and that’s the only place that that data is stored. It’s stored in that region. Of course, it’s stored across three physically diverse data centers within that region to provide durability and availability, but it’s stored entirely within that region.

And I think in hindsight, I think if we had known, initially, that we would have moved into that regional model, we may have thought a little bit differently about how buckets are named, for example. But where we are now, we definitely like the regional resiliency, I think that’s a model that has proven itself time and time again, that having that regional resiliency is critical. And customers really appreciate that.

Corey: Something I want to talk about speaks directly to the heart of that resiliency, and the, frankly, ridiculous level of durability and availability the service offers, we’ve had you get on stage talking about these things, we’ve had Mai-Lan several times on stage talking about these things, and Jeff Barr writes blog posts on all of these things. I’m going to go out in the limb and guess that there’s more than just the three of you building this.

Kevin: Oh, yeah.

Corey: What’s involved keeping this site up and running? Who are the people that we don’t get to see? What are they doing?

Kevin: Well, there’s large engineering teams responsible for S3, of course, and they, I would say, in many ways are the unsung heroes of delivering the services that we do. Of course, you know, we get to be on stage and talking about these cool new features, but it’s only with a ton of hard work about the engineering teams day in and day out. And a lot of it is having the right instrumentation, and monitoring the health of the service to an incredibly deep level. It’s down very deep into hardware, of course, very deep into software, and getting all those signals and then making sure that every day, we’re doing the right set of things, both in terms of work that has to be done today, and project work that will help us deliver step-functions improvements, whether it’s adding another degree of availability, or looking at just certain types of data and certain edge cases that we want to strengthen our posture around, there’s constant work to look around corners, and then really just to continuously raise the bar for availability, and resiliency, and durability within the service.

Corey: It almost feels, on some level, like the most interesting changes and the enhancements that come out, almost always without comment, come from the strangest moments. I mean, I remember having a meeting with a couple of folks a year or two ago, when I was—I kept smacking into a particular challenge; I didn’t understand that there was an owner ACL at the time, and it turned out that there were two challenges there. One was that I didn’t fully understand what I was looking at, so people took my bug report more seriously than it probably deserved. And to be clear, no one was ever anything professional on this. And we had a conversation, my understanding dramatically improved, but the second part was a while later, “Oh, yeah. Now, with S3, you can also set an ACL that determines that any object placed into the bucket now has an ownership ID of the bucket owner.”

And I care about that primarily because that directly impacts the cost and usage reports that are what my company spends most of our life staring into. But it made for such an easier time as far as what we have to deploy to customer accounts and how we went up thinking about these things. And it was just a quiet release that was like many others with the same lack of fanfare that, “Oh, the service you don’t use is now available in a region you’ve never heard of. Have fun.” And there are, I think, almost 3000 of various releases last year; this was one of them that move the needle.

It’s little things like that, but it’s not so little because doing anything like this at the scale of something like S3 is massive. People who have worked in very small environments don’t really appreciate it. People who have worked in much larger environments—like, the larger the environment you get to work within the more magical something like this seems.

Kevin: Yeah, I think it’s a good example, you point to the S3 object ownership example, I think that’s a great example of the kind of feature that took us actually quite a bit of work to figure out how we would deliver that in as simple a fashion as possible. That was actually a feature that, at one point, I think there was a 2 or 3D matrix being developed of different ways that we might have to have flags on objects. And we just kept pushing and pushing to say, “It has to be simpler. We have to make this easier to use.” And I think we ended up in a really good spot. And it certainly, for customers that have lots of accounts, which I would say almost all of our large customers end up with many, many accounts—

Corey: Well, we’d like to hope so anyway. There was a time where, “Oh, just one per customer is fine.” And then you got to redefine what ‘large account’ looked like a few times that it was, “Okay, let’s see how this evolves.” Again, the things you learn from customers as you go.

Kevin: Yeah, exactly. And then there’s lots of reasons for different teams, different projects, and so forth, where you have lots of accounts. But for any of those, kind of, large accounts scenarios, or large organization scenarios, there’s almost always cases where you’re writing data across accounts in different buckets. So certainly, that’s a feature that, for folks who use S3, they knew exactly how they were going to use it, turned it on right away.

Corey: It’s the constant quiet source of improvement that is just phenomenal. The argument I always made that I think is one of the most magical parts of cloud that isn’t really talked about is that if I go ahead and I build an environment and I put it in AWS, it’s going to be more durable, arguably more secure, and better run and maintained five years later, if I never touch it again, whereas if I try that in a data center, the raccoons will carry the equipment off into the wilderness right around year three. And that’s something that is generally not widely understood until people have worked extensively with it.

S3 is also one of those things that I find is a very early and very defining moment, when companies look at going through either a cloud migration or a digital transformation, if people will pardon me using the term that I love making fun of, it’s a good metric for how cloud-y for lack of a better term, is your application and your environment. If everything lives on disks attached to instances, well, not very; you’ve just more or less replicated your data center environment into a cloud, which is fine as a step one. It’s not the most efficient, it makes the cloud look a lot more like your data center, and you’re not leveraging a lot of the capability there. Object storage is one of the first things that seems to shift, and one of the big accelerators or drags on adoption always seems like it comes down to how the staff think about those things. What do you see around that?

Kevin: Yeah. I think that’s right, Corey. I think that it’s super exciting to me working with customers that are looking to transform their business because oftentimes it goes right down to the data in terms of, what data am I collecting? What can I do with that data to make better decisions and make more real-time decisions that actually have meaningful impact on my business? And you talk about modern applications, some of it is about developing new modern applications and maybe even applications that open up new lines of business for a customer.

But then we have other customers who also use data and analytics to reduce costs and to better manage their manufacturing or other facilities. We have one customer who runs paper mills, and they were able to use data in S3 and analytics on top of it, to optimize how fast the paper mills run to eliminate the machines or reduce the amount of time that machines are down because they get jammed. And so it’s examples like that, where customers are able to first off, using S3 and using AWS able to just store a lot more than they’ve ever thought they could in a traditional on-premises installation, and then on top of that really make better use of that data to drive their business. And I mean, that’s super exciting to me, but I think you’re right as well about the people side of it. I mean, that is a, I think an area that is really underappreciated in terms of the amount of change and the amount of growth that is possible and yet really untapped at this point.

Corey: On some level, it almost shifts into—and again, this is understandable. I’m not criticizing anyone, I want to be clear here. Lord knows I’ve been there myself—where people start to identify the technology that they work with, as a part of their identity of who they are, professionally or in some cases personally.

Kevin: Yep.

Corey: And it’s an easy misstep to make. If there were suddenly a giant pile of reasons that everyone should migrate back to data centers, my first instinct would be to resist that, regardless of the merits of that argument because well, I’ve spent the last four years getting super deep into the world of AWS. Well, isn’t that my identity now on some level, so I should absolutely advocate for everything to be in AWS at all times. And that’s just not true; it’s never true, but every time it’s a hard step to make, psychologically.

Kevin: Oh, I agree. I think it is, psychologically, a hard step to make. And I think people get used to working with the technology that they do. And change can always be scary. I mean, certainly for myself as well, just in circumstances, where you say, “Well, I don’t know. It’s uncertain; I don’t know if I’m going to be successful at it.”

But I firmly believe that everyone at their core is interested in growth and developing, and doing more tomorrow than they did yesterday. And sometimes it’s not obvious. Sometimes it can be frightening, as I said, but I do think that fundamentally people like to grow. And so I think with the transformation that’s ongoing in terms of moving towards more cloud environments, and then, again, transforming the business on top of that, to really think about IT differently, think about technology differently. I just think there’s tremendous opportunity for folks to grow; people who are maintaining current systems to grow and develop new skills to maintain cloud systems or to build cloud applications even. So, I just think that’s an incredibly untapped portion of the market in terms of providing the training, and the skills and support to transform the culture and the people to have the skills for tomorrow’s environments.

Corey: Thank you so much for taking the time to speak with me about this dizzying array of things that S3 has been doing. What you’ve been up to for the last 15 years, which is always a weird question. “What have you been up to for the last 15 years, anyway?” But usually in a much more accusatory tone. If people want to learn more about what you’re up to, how you’re thinking about these things, okay can they find you?

Kevin: Well, I mean, obviously, they can find the S3 website at aws.amazon.com/S3. But there’s a number of videos on Twitch and YouTube both of myself and many of the folks within the team. Really, we’re excited to share a lot of new material. This week, with our Pi Week we decided Pi Day was not enough; we would extend it to be a four-day event. So, all week we’ve been sharing a ton of information, including some deep dives with some of the principal engineers that really help build S3 and deliver on that higher bar for availability, and durability, and security. And so, they’ve been sharing a little bit of behind-the-scenes, as well as just a number of videos on S3 and the innards there. So, really invite folks to check that out. And otherwise, my [inbox 00:32:54] is always open as well.

Corey: And of course, I would be remiss if I didn’t point out that I just did a quick check, and you have what can only be described as a sarcastic number of job openings within the S3 organization of all kinds of different roles.

Kevin: That’s right. I mean, we’re always hiring software engineers, and then systems development engineers in particular, as well as product management—

Corey: And TPMS, and, you know, of course, I’m assuming naming analysts. Like, “How do we keep it ‘S3,’ but not call it ‘simple’ anymore?” Let me spoil that one for someone: serverless. You call it serverless storage service, and you’re there. Everyone wins. You ride the hype train, everyone’s happy.

Kevin: I’m going to write that up right now, Corey. It’s a good idea.

Corey: Exactly. Well, we find a way to turn that story into six pages, but that’s a separate problem.

Kevin: That’s right.

Corey: Thank you so much for taking the time to speak with me. I really appreciate it.

Kevin: Likewise. It’s been great to chat. Thanks, Corey.

Corey: Kevin Miller, General Manager of Amazon Simple Storage Service, better known as S3. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an insulting comment that whenever someone tries to retrieve it, we’ll have an Object Lambda rewrite it as something uplifting and positive.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Laurie

Laurie has been a web developer for 25 years and cares deeply about making the web bigger and better for everyone. He previously co-founded awe.sm and npm, and is currently a Senior Data Analyst at Netlify.

Links:

  • Netlify: https://www.netlify.com/
  • Twitter: https://twitter.com/seldo
  • Personal website: https://seldo.com/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Fairwinds. Whether you’re new to Kubernetes or have some experience under your belt, and then definitely don’t want to deal with Kubernetes, there are some things you should simply never, ever do in Kubernetes. I would say, “run it at all.” They would argue with me, and that’s okay because we’re going to argue about that. Kendall Miller, president of Fairwinds, was one of the first hires at the company and has spent the last six years the dream of disrupting infrastructure a reality while keeping his finger on the pulse of changing demands in the market, and valuable partnership opportunities. He joins senior site reliability engineer Stevie Caldwell, who supports a growing platform of microservices running on Kubernetes in AWS. I’m joining them as we all discuss what Dev and Ops teams should not do in Kubernetes if they want to get the most out of the leading container orchestrator by volume and complexity. We’re going to speak anecdotally of some Kubernetes failures and how to avoid them, and they’re going to verbally punch me in the face. Sign up now at fairwinds.com/never. That’s fairwinds.com/never.

Corey: The apps on cloud summit is a new action packed, not a conference, happening May 11th through 13th online. Its for everyone who makes applications in the cloud run screaming. From IT leaders to DevOps pros to you folks, whoever you might be. Take a break from screaming into the cloudy void with me to learn from some of the best of people who actually know what they’re doing. Like Kelsey Hightower, AWS blogger John Meyer, and also me, because apparently they didn’t listen to me saying I had no idea what I was doing. Register now at turbonomic.com/screaming. Theres a “swag box” ready to ship for the first two thousand registrants, so you don’t want to miss this. Thanks for Turbonomic for sponsoring this ridiculous podcast.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’m joined this week by Laurie Voss, who is currently a senior data analyst at a company called Netlify. Laurie, thank you for joining me.

Laurie: Thanks for inviting me.

Corey: So, let’s start at the very beginning. What is Netlify?

Laurie: Netlify is a single cohesive build chain for websites. A lot of people don’t think of it that way. I think a lot of people think of Netlify as a web host, but really where people are getting value from Netlify is you build your website, you upload your website, you deploy your website, you host your website, you test your website, you monitor your website, and that can be five or six different services, like a CI service, and a hosting service, and a Git service and all of those things. And Netlify just joins that entire build chain into a single tool where you just hook up a Git repo, hit commit, and it goes out into the world, and it’s incredibly fast and convenient. And that’s really where people get value out of it.

Corey: Perhaps somewhat uncharitably, I would almost think of that as Heroku for this decade.

Laurie: I mean, I would consider that pretty charitable to us and somewhat uncharitable to Heroku, who are still around and chugging.

Corey: Oh, absolutely. I’m a big fan of things like that, where it’s take this code—whatever it looks like, maybe it’s a repository, maybe it’s some, I don’t know, some files I email over, God forbid—and then go ahead and deploy it into something that at least pretends to be able to scale. I often hear Netlify brought up in the context of Jamstack, which seems to be this whole area of cloud computing that I don’t tend to spend a whole lot of time in, at least not knowingly. What is it?

Laurie: So, Jamstack originally stood for JavaScript, APIs, and markup sometimes also referred to—

Corey: But I hate all of those things. Please continue.

Laurie: [laugh] it’s sometimes also referred to as static websites, which is a term I tend to avoid simply because it’s not really very accurate. A static website is one of the things that you can deploy on the Jamstack, certainly, but it’s certainly not the only thing you can deploy. I would say that it is an architecture that lends itself to pre-rendering as much content as is possible, and then caching all of that stuff at the edge, and then pulling in only the bare minimum of dynamic content to improve both scalability and performance. Those are the things that people like about Jamstack websites, is that they tend to be extremely fast.

Corey: So, that makes intuitive sense to me. And you, of course, became fairly broadly known as one of the people behind npm. But now you’re a senior data analyst, which feels like it’s a departure from the things you were doing to the things you’re doing now. Help me either validate that, or tell me what obvious thing I’m missing, or highlight something clever for me because right now, I feel like there’s a missing link in my chain of events here.

Laurie: No, that’s a totally fair question. So, I started npm as the CTO and hired an excellent engineering team underneath me. In fact, one of our very first hires was a lady called C J Silverio, who is just a staggeringly good engineer. And it became very obvious very early on in the life of the company that we really had two people of CTO caliber, and that we didn’t need to have them, but what we did need was somebody to run the operational side of the business. So, relatively early on in the life of the company, we promoted C J to CTO, and I moved my title to COO, you know, obviously, still with a technical bent, but my job as a COO is to do operational things.

So, I was in charge of running the financials and making sure that marketing and sales weren’t going massively over budget or under quota, those sorts of things. And that’s fundamentally a keep-the-lights-on data analysis job. So, while I was CTO, I was sharing fun stats about npm’s internals; while I was COO, I was doing a lot of analysis of our financials. But the common factor was analysis, and I was doing more and more of it. So, towards the end of my time at npm, I became the Chief Data Officer, where I basically specialized down into doing just data things—some financial, some technical—and doing a lot of outward-facing presentations about that kind of thing.

So, that was where my job ended up being. And literally how I pitched my way into Netlify was like, “What if I did that thing that I was doing for npm for you,” and they were like, “Great. You can’t be a C though because you just got here.” [laugh]. I was like, “Fine.”

Corey: Well, of course. We all have to start somewhere. Humility. And it took me a couple of years to unofficially run AWS marketing. My God. Yeah, have some humility as you step through this process. Was it a big barrier to you once you arrived at Netlify, convincing them to buy you the Excel license you obviously need to do all this data analysis, or alternately, are there better tools for it, then the one that we’ve all been using anyway?

Laurie: Honestly, I’ve always been a Google Sheets partisan. I know that the really hardcore financial types will complain about the functions that are missing from Google Sheets versus Excel—

Corey: Oh, will they ever.

Laurie: —but I’m not that person. But we have a pretty great stack that I like quite a lot at Netlify these days. We have a variety of older tools laying around, not all of which we’ve migrated away from, but the core of the new class is this company called Databricks, who are basically Spark clusters as a service. So, you can just throw, essentially, arbitrarily large amounts of log data on to S3 buckets on AWS, and it can query them as if they were databases, which is truly beautiful. And on top of them, we have a system called Mode Analytics, which is a general platform for data analysis, and presentation; draws graphs, that kind of thing; has an SQL interface.

And between those two we’ve got a new open-source project, or relatively new to me anyway, called dbt, which is this very organized, clever way of codifying your best practices around data. So, you’ve probably heard of extract, transform, load jobs; it’s basically a way of quantifying chains of extract, transform, and load jobs such that they’re always tested, and always running, and you know what the dependencies are between them and everything is documented.

Corey: Okay. While I’m in the process of getting everyone in trouble on things, what is your take on machine learning for things like this? Because it seems that whenever you talk about data, it’s inevitable that someone, usually with a crap ton of VC backing, will immediately jump in because they’re clearly getting bonused every time they managed to fit the phrase ‘machine learning’ into basically anything.

Laurie: So, I would step back a bit and say that, before I joined Netlify, I interviewed at a couple other companies just to see what the space was like, for basically the same job at other companies. And there was a really interesting pattern that I noticed, which is that it is quite a common pattern for an early-stage startup, to say, “Oh, we have a data problem. We must hire a data scientist.” And they go and find somebody staggeringly qualified, with a PhD in data science, and they hire that person. And that person immediately runs into trouble because that is not actually the problem that they have.

They don’t have a data science problem; they have a data engineering problem. They have, like, mounds of data lying everywhere, and it’s not organized, nobody knows where it is, nobody can query it efficiently. A data scientist is, at earliest, your fifth hire in your data team. The first five people are people who have to do an enormous amount of plumbing and engineering to be able to just get the data from all of the places that it’s lying around, all of the piles that it’s accumulating in, into any kind of a reasonable format that you can query it and figure out what it does.

Corey: You have to forgive my cynicism on some level because I’ve been in the ops space for, I guess, entirely too long where I’ve been dealing—particularly in the context of AWS bills, with making arguments against data science teams who are insisting that the Apache logs from 2012 that are taking petabytes of space are the key to unlocking the mysteries of the business. They’re not sure how yet, but one day they’re going to become super valuable, so I’m never allowed to delete anything. And on some level, it just almost seems like it’s a big make-work conspiracy for data scientists amongst each other which, hey, respect. Counter-argument; what sorts of insights can you glean from these vast quantities of data because everyone else I’ve talked to about this generally works for a big-data-oriented company. I got to be honest with you, it feels like they’re selling pickaxes into a gold rush because, “Oh, it’s very important to keep all your data so that we can sell you things to go through it.” You’re on the other side of that your buy-side. So, what is the value that this giant data hoard winds up providing?

Laurie: Well, I will say that my initial inclination is to agree with you. There’s definitely a lot of pickaxes being sold to miners who have no idea what they’re doing. I think about ten years ago, there was a huge industry-wide pile-into big data people were like, you need Hadoop, and you need gigantic data processing clusters, and huge data, and massive amounts of processing, and, like, buy this enterprise contract for $100,000 a year. And then everybody did those things and was like, “And now what?” And they were like, “Oh, well, we don’t know. Maybe you can count it up. How many hits did you get?”

That’s not useful analysis. Having all of your data queryable is not, per se, a useful thing to be able to do. And I think in the 10 years since then, people have got smarter about that. They realized medium and small data are actually [laugh] often quite useful. It’s more about how you analyze it, and can you present it to people, and can you make sense of it?

But there was a second gold rush into the ML space. There are certainly use cases where you have enough data and a problem that is amenable to being solved by applying ML to it in some way. Those are a minority of cases; they’re maybe five percent of all data problems are big enough that you can use ML in the first place, and also get an answer that ML can help you with, would be helpful. And the other ninety-five percent, it’s just plumbing and engineering.

Corey: Once upon a time, it felt like the way to address all this data was the… honestly, the result of a prank perpetuated many moons ago by what felt like Google in a white paper, that Yahoo went for hook, line, and sinker for MapReduce, which then led to Hadoop and a bunch of other stuff. I maintain this was a Google April Fool’s prank that everyone took way too seriously and went way too far. These days, it feels like stream processing as that data comes in is sort of the preferred approach. Yes, no, or am I completely misunderstanding most of the point? Or all the above?

Laurie: I would say definitely, the industry has moved away from the batch processing that Hadoop did. I actually worked at Yahoo at the time when they were inventing Hadoop. [laugh].

Corey: Oh, you fell for it, too. Great.

Laurie: [laugh]. I was—we were selling the Kool Aid as opposed to drinking it.

Corey: Oh, if you’re going to be involved in a Kool Aid transaction, that is absolutely the side of it you want to be on. Let’s be very clear here.

Laurie: So yeah, streaming processing, but like semi-real-time processing of things, as opposed to giant batch jobs is certainly where stuff has mostly gone. Although people who are end consumers of data, as an analyst, if I asked you how fresh does this data need to be, they will always say realtime. Like, [laugh] that will be their first answer. And then I’ll be like, “What if it was 24 hours delayed?” And they’re like, “Oh, yeah. Well, obviously, yesterday’s data is fine. I’m not going to care about what happened at noon today when it’s 2 p.m.” And then you’re like, “Well, yes. Well, then it’s a batch job, and it’s, like, an order of magnitude cheaper to provide to you, so let’s do that.” Batch jobs are still very cost efficient and so we do a lot of batch processing, it’s just we don’t make a big song and dance about it anymore because it’s no longer the new shiny thing.

Corey: On some level, it feels like that is the nature of things where something gets announced, and it’s super complicated and hard, and people skill to the peaks of complexity, and they make good money doing it. I mean, in the original dotcom boom, ‘firewall engineer’ was a quarter million dollars a year if you could swing it. Now, it’s just assumed that basically, anyone who touches the network should be able to configure firewall rules; things get simpler with time. It feels, on some level, like an awful lot of the data world is undergoing some of that consolidation as well, where we’re starting to find tools and methods and ways to extract meaning from giant piles of data without the part where, you know, you go and drop $5 million here on a data science team.

Laurie: Well, you’ve sort of arrived at my favorite pet topic, which is the stack. The stack is this abstraction that I wrote about at the beginning of last year. It’s the idea that the ever increasing complexity of technical fields means that we are constantly inventing, adopting, and then forgetting about abstractions. As you said, we’re constantly chasing after the new shiny thing; we make a big song and dance about it; it’s very complicated. People make enormous amounts of money doing it in the early days, and then somebody eventually invents some kind of tool or open-source framework, or possibly, like, a SaaS that makes it one-click to do.

And it’s not any less complicated or any less magical than it was before, it’s just you think about it much less, right? Like I mentioned, Databricks. Every time I run a query Databricks is taking my SQL, converting my SQL into giant MapReduces, running it on a huge cluster of machines of arbitrary size alu—I don’t know what size it is because I don’t need to care anymore—and then pointing it at AWS, where it’s pulling in every single piece of data in every bucket that I put in there. And all of that, ten years ago would have been of a complexity that only Google or Yahoo could do it. And now it’s literally we spin them up by clicking a button and we don’t even remember that it’s happening.

Like, all of that complexity is still happening, all of that magic is still happening, but now it’s just a commodity. And we’re doing that across the tech space. So, we’ve certainly done it in data; a bunch of stuff that used to be very complicated, used to be the thing that you would hire me to do is now just the tool that I use and the thing that I do is the analysis, which is a more useful use of someone’s time, really.

Corey: One of like to hope so. But I do feel like there’s a story—and we see it across the board; this is one of the things I really enjoy about Netlify—once upon a time to put a website on the internet, you had to know a whole bunch of different things all at the same time. It was, how to build a web server, how to maintain and patch that web server so it didn’t become an attack spam cannon, how to get files into a format the web server could understand, how to put that out there, how to get DNS to work, how to handle SSL—if that was even a glimmer in your eye at that point—and so on and so forth. Now, it really requires, click a button. And Netlify is made this way easier because I tend to look at this from the exact opposite side in the industry where I come from an ops background; building all the infrastructure to handle these things is relatively straightforward to me, but then I get to the other side.

Cool, now all that’s done, “Build the web app.” And my response, “Ehhh, what?” Yeah, I can write bad HTML by hand, sort of, and that’s as far as I generally tend to go, whereas it feels like the Jamstack story in general, and Netlify in particular, are aimed at folks in many ways, coming from the other side of the world where it’s, “I picked up JavaScript. I picked up a framework or two. I understand frontend, I understand how web applications get built. What’s the deal with this whole infrastructure piece?” And thanks to the miracle of stacks collapsing in upon themselves in many respects, you don’t have to know about that or care, and you live in this blissful world where the term Kubernetes never crosses your desk. Is that a fair summation of the state of the industry? Am I dramatically misunderstanding what Netlify does and for whom?

Laurie: No, I think that’s pretty much how it goes. One of the reasons that I wrote this blog post about the stack—it was almost exactly a year ago—is because about a year ago is when I joined Netlify and I was suddenly immersed in the things that Netlify does. It became more clear to me that I was seeing a fundamental shift happening.

I was like, “Oh. We are obeying some kind of natural law here, right? We are taking things that used to be people’s whole jobs and turning them into things that are so simple that you don’t even think about them happening anymore.” I’ve definitely met and worked with people in my life whose whole job was managing SSL certificates. And now, it’s literally a checkbox. And it’s on by default. It’s like, “Would you like your site to be secured by SSL?” Yes, obviously. I don’t know why I would turn that off.

And it just comes as part of deploying your website. Way in the background, let’s encrypt is doing it, and there’s a whole bunch of song and dance about refreshing certs every 90 days, and it all just happens completely automatically without you caring even a little bit. And that’s what Netlify is doing. It’s taking things that used to be five or six companies and squishing them down into a single layer that you call your deploy service. And you’re like, “Great. My deploy service does all of those things and I don’t need those other five companies anymore.”

Corey: Now, if you’re one of those five companies, that becomes something of a problem. But again, that’s the pace of innovation. That is the world continuing to evolve.

Laurie: Nobody wants to be commoditized, but on the other hand, the company that gets to do the commoditizing tends to run away with it, right? Like that’s kind of the AWS story. It’s like, there used to be lots and lots of companies that would sell you a server in a rack and then take 24 hours to set it up and you’d pay with a credit card. And AWS was like, “What if that was one button?” And everyone was like, “Yes, I would love that to be one button. I never want to care about what rack it’s in anymore, or whether or not it has enough power, or whether or not the cable in the back has got jiggly. Just virtualize it all the way for me, thank you.” And then AWS completely ran away with it.

Corey: Oh, yes. And it’s AWS, so it was, “What if that button was hidden in a console that doesn’t work super well, and then we give that button a terrible name?” People are like, “Ehh, I’ll risk it.”

Laurie: I mean, the observed behavior of the industry is that we love the terrible console.

Corey: Oh, absolutely. Everyone talks about infrastructure as code, which is basically a polite way of saying I use the console, and then lie about it on conference talks.

Corey: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the Enterprise (not the starship). On-prem security doesn’t translate well to cloud or multi-cloud environments, and that’s not even counting IoT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IoT devices, detects these threats up to 35 percent faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at extrahop.com/trial.

Laurie: [laugh]. Indeed.

Corey: So, since you brought up AWS, terrific, it’s time for me to do my whole conspiracy theory approach here and accuse you of basically war crimes. So, you were big into the npm space for a long time, which is great. I accept the fact that that is a thing that happens—package.json and package-lock.json are basically artifacts of you folks.

Now, AWS has launched their Amazon CodeGuru machine learning—wink, wink, nudge, nudge—powered code review. And of course because it’s AWS, they charge based upon lines of code in a pull request, which tells me that you’re a deep plant for many years now, planning for the day where this one day supports JavaScript—which it doesn’t today—and all someone has to do is check in the package-lock and the package.json files once, and suddenly the entire scheme pays off handsomely. True, false, or I’m not supposed to talk about that in public?

Laurie: It’s true. I’m part of a global cabal whose purpose is to make Node modules infinitely deep until the gravity well sucks in all of programming and we don’t have computers anymore.

Corey: On a slightly more serious note, I do want to talk a little bit about package management—in the context of programming languages as opposed to package management in the context of Linux distributions because, oh, do I have thoughts on that—there are a few different competing tools out there to handle dependencies across different programming languages, in the JavaScript world, in the Python world. And I’m not a JavaScript programmer, except when forced to be, and it’s usually editing something as small-scale as humanly possible and backing away slowly. But my general consensus, looking at it across the board, is that there is no consensus, that there is no clear one right way to do things. Invariably, dependencies always become a challenge. Getting something to a reproducible build while also being secure is a problem.

And no matter what stack you pick, what language you pick, there’s always a—for ‘Hello World’—there’s a step one of setting up your local environment to resemble what the person writing the document’s environment looks like. Is that accurate? Is there some magic tool out there that somehow I’m just unaware of that solves all of this for me?

Laurie: Well, there’s definitely not a single tool that gets it completely right, but I would say that there is a commonality between the things that work that I don’t know that everyone appreciates. So, I’m going to draw a parallel between package.json and Kubernetes right now, so bear with me. Basically the thing that people often don’t like about npm and the thing that people don’t like about package.json is that it says, “All of your dependencies must live here, in your tree. I don’t care how many JavaScript projects are on your computer; I am going to have one copy of every module right here where I can see it, and I’m going to use those and only those.”

It tends to make JavaScript programs a little bit easier to debug because you know that the code that is at fault can’t possibly be anywhere else. It can’t be sitting in userlib unexpectedly, or in some additional libraries folder, or it can’t have been, like, blown away by somebody installing something else. It has to be the one that’s sitting in your tree, and that’s one of the things that made Node so popular in the beginning, and npm so popular at the same time, was that it was very easy to deal with, and in particular, it made it work on Windows, which didn’t have any of those things anyway. And Node's popularity as a development environment, where you could write code on Windows and it would work perfectly in a Linux environment because all of the dependencies were JavaScript and that ran the same on both of those computers is understated. And that’s essentially the Kubernetes story.

Kubernetes is saying, “This thing where we have libraries all over the place, where we have dependencies all over the place, like, they lie all over the operating system. It’s too late to fix that. What if we packaged up the entire operating system and said that that’s the package?” And that’s what Kubernetes is. It’s creating a package.json of your entire computer, and then you run that.

Corey: It sure beats the old approach of, “Oh, it works on your machine. Great. Well, backup your email, Slappy, because your laptop going to production.”

Laurie: Exactly. Right. It’s basically, you’ve packaged up the entire world. And people are like, “Well, this is very wasteful.” And we’re like, “Yes, it’s very wasteful. But it works.” And like the other approach—

Corey: You know what’s less wasteful, then? That’s right, a whole bunch of engineering time spent fixing things. “Well, that’s not the most optimal way of doing it,” say people who seem to consistently mistake their time for being free.

Laurie: Exactly.

Corey: No, and it makes perfect sense. I love the fact that I can use at least some semblance of what other people are using and get it to work. The counter-argument to it is that it’s very—how do I put this—disconcerting when I’m working in a Python project, but I’m using a framework or so that generally installs via npm, and now my Python project has a package.json in there, and I get very confused at first. And, all right, then I run npm install in there and then I’m way more confused. And I mostly just look at this, and I struggled to make sense of it before the penny drops. “Oh, that’s right. It’s because I’m bad at computers.” I wish people would not keep letting me forget that part.

Laurie: [laugh]. Is your objection that you can’t launch a website these days without JavaScript anymore because a lot of people are angry about that, and they send me email more often than you would imagine.

Corey: Well, I assume it’s your personal fault, right?

Laurie: I mean, absolutely. Like, again, the secret cabal; we’re trying to inflate all of your applications with as much extraneous code, with as many security vulnerabilities as we can possibly manage because I work for the people who sell storage and virus scanning, obviously.

Corey: Emailing you about the world requiring JavaScript is evocative of an old story where some town manager angrily emailed the CentOS project maintainers because someone installed a web server in his environment and he pulled it up, and this isn’t our town’s website; it’s the default, “Welcome to CentOS. If you’re seeing this page, you’ve successfully installed Apache. Read these docs to configure it…” and accused them of hacking his website. It seems roughly the same level of technical nuance, blaming you for the proliferation of something in society.

Laurie: I don’t know. I mean, I certainly spent five years cheerleading it, so I feel like people who are, like, “You helped make this popular.” I’m like, “Oh, why thank you. I’m so glad you think I made a difference.” But really, it probably would have happened on its own. Like, I was running after a snowball that was already running very quickly downhill and engulfing villages as it went.

Corey: Absolutely. And I do want to talk to you about that in particular because as people on this podcast often hear, I talk about this podcast, I talk about the AWS Morning Brief my other podcast, and I talk about lastweekinaws.com where my newsletter lives; I don’t urge people to follow me on Twitter, I don’t talk about the Facebook page I don’t have. And the reason behind all of those things, is that I have built an audience on open standards and open platforms so that no one company can change business models and suddenly I have a serious problem.

It’s why I blog on my own website, not on Medium. Their business model changes aren’t going to directly impact what I do and how I do it. Do you think this is naive? Do you think that the open web was a nice idea and now we’re just going to see increasingly walled gardens as time goes on?

Laurie: I think the openness of your website is—or your web app, or your, sort of, technical strategy in general—is always going to be a hybrid; like AWS is… it’s not rolling your own, you’re using a service. If AWS decides that they don’t support your service anymore—which they never do as far as I can tell, but theoretically, they could—you would have to stop doing that; you are to some extent locked into AWS. But I don’t think that a website hosted on AWS is, like, not part of the open web.

Corey: I would agree wholeheartedly on that point, absolutely.

Laurie: Right. I think at that point, you’ve adopted a tool that works for you, and you can move elsewhere. So, there are people who say using JavaScript frameworks, that’s not the open web, you should have been writing your own; you’re dependent on Facebook continuing to maintain React. And I’m like, “Well, kind of, but not really. You don’t have to be. You could write your own website, if you wanted to. This way, it’s just faster, in the same way that hosting it on AWS is faster than spinning up your own machines.”

Corey: Oh, I take it a step further beyond that, I paid WP engine which, they manage WordPress for me, so I don’t have to, and the reason for that is I’ve managed WordPress in the past, and I will not go down that path again for love or money.

Laurie: [laugh]. Right.

Corey: But then, as a fun artifact of that, lastweekinaws.com does in fact live on GCP.

Laurie: [laugh]. Nice.

Corey: But it’s WordPress. Worst case, WP Engine shuts down, or charges me at times more, or decides that now, nope, everything has to move to a new framework, I can migrate it elsewhere. And the fact that I have that strategic exodus means that I don’t need to sit here on everything I build and agonize over, do I go all-in on my current hosting provider or not? It’s something that I can migrate with me. And I try and maintain at least that theoretical exodus path.

I can repoint domains to other places; I own the domains myself, and that has been enough for the way that I view the world. But increasingly, I’m starting to feel like a relic. Oh, follow me on Instagram; follow me on TikTok and it’s if these platforms pull a MySpace and vanish, then you’ve got to rebuild your audience from scratch, whereas email’s been with us longer than I’ve been alive, and it’ll be here long after I’m dead. I can carry that audience with me regardless of what any particular provider has. I just wish I didn’t feel like such a Captain Edgecase, or someone stuck in the past whenever I articulate that to some folks.

Laurie: Well, I’ve been in the industry a long time, so I think if you’re going to, sort of, say, “I’ve got this old opinion,” I’m going to be like, “Me, too. I’m also extremely old.”

Corey: And then we’ll talk about the Great War. “Wasn’t it amazing?” And, yeah, there we are.

Laurie: The browser wars of ’97, and I was ‘16.

Corey: Yes, we’ll make Eternal September references all week.

Laurie: Oh, my God. See, we’re literally doing that thing that I was just joking we were going to do.

Corey: We absolutely are.

Laurie: Yeah, I think you have to pick your battles. I think the one that I personally struggle most with is databases. I spent a good chunk of my career as a DBA; I definitely know how to install and configure databases. I don’t want to. [laugh]. You know, like, using one of the fancy databases as services, where you’re just like, it has an SQL interface and it’s got apparently infinite storage and infinite processor, and I don’t need to worry about it anymore.

Corey: Exactly, and it has those things because what it also has is someone else’s credit card. Done.

Laurie: Right. It’s great. But to some extent, I’m definitely locking myself into that database service, right? To some extent, I have to find an equally capable service if I ever wanted to migrate away. So, am I still open, or am I locked in then?

I don’t think anybody can call themselves truly independent, anybody can call themselves truly open. So, from your perspective of, like, what platform am I on, as long as you’re not only on that platform, as long as it’s not your only bet, I think—sure, pile into the Facebook page. Why not?

Corey: Yeah. I have separate problems with that that we need not get into here.

Laurie: [laugh].

Corey: That’ll be a whole separate episode there. So, as to look across the past—I don’t know, let’s call it eight decades that you and I have been in tech together, what are the themes you’ve seen continue to emerge that people should be paying attention to moving forward?

Laurie: I think one of the most common mistakes that I see in technologists who’ve been in the industry a long time, is—I can tell that they’re doing it because they start ranting about ‘the fundamentals.’ And it is my firmly held conviction—and no one will sway me from it—there is no such thing as the fundamentals. Everybody comes into the industry at a certain time, when a certain set of tools were considered commodities that you don’t need to think about, a certain set of tools were considered, like, the complicated thing that you need to learn, and a certain set of tools were considered, like, fluff on top that are bonus, but those things are always drifting downwards, right? Yesterday’s fluff is today’s bedrock, and the new fluff is stuff that wasn’t invented before. And then they start going, “Well, you should be able to understand HTTP, and roll your own JavaScript framework because those are the fundamentals.”

And I’m like, “Only to you because you came into the industry when that was the complicated thing. The fundamentals to somebody who started 20 years before you did are like, ‘you need to know about power management and how to configure a firewall,’” like you were saying, in the beginning of this thing. Everybody’s fundamentals are somebody else’s fluff.

Corey: Oh, you want to learn how Linux works? Step one—I see this in classes all the time—learn how Vim works.

Laurie: Right, exactly.

Corey: How about not doing that?

Laurie: Oh, my God—

Corey: —and focusing on the differentiated part. My God.

Laurie: The bizarre cargo-culting of Vim. I’m like, “You know why the people who are good at Vim are good at Vim? It’s because they’ve been doing Vim for 30 years. If you do any tool for 30 years, you’re going to be really good at it.”

Corey: So, you say that, but then you look at me with databases, and I don’t know, I might be able to fool you on that one.

Laurie: [laugh]. Use any tool for 30 years, and you’ll be so good at it that the switching cost is too high to go to anything else. But if you’re just starting in the industry, you could start with any editor that you wanted and it would be fine. And by the time you’ve been using it for 30 years, you’ll be like a goddamn wizard at it.

Corey: Mm-hm. Absolutely.

Laurie: So, that’s what I tell people is, like, the things that you learn now, you’re going to have to expect that they get commoditized. The stack that you live on today will get crushed down to nothing and you have to be constantly climbing the stack to what the new thing is.

Corey: [laugh]. I want to thank you for taking so much time to speak with me today. If people want to hear more about what you have to say and how you wish to say it, okay can they find you?

Laurie: I’m most active and responsive on Twitter. My username is @seldo and I also own seldo.com where I blog much less frequently than I would like to.

Corey: And we will, of course, put links to both of those into the [show notes 00:32:33]. Thank you so much for taking the time to speak with me. I really appreciate it.

Laurie: Thanks for the invitation. It’s been a lot of fun.

Corey: Really has. Laurie Voss, senior data analyst at Netlify. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice and an entirely insulting, rambling comment complaining about how I talked about all these different package management systems for different languages and never once mentioned Rust.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Mark

Mark Curphey is the co-founder at Open Raven, a cloud native data security company. Mark’s fingerprints can be found all over the security industry, but perhaps most visibly from his role as the founder of OWASP. His contributions range from his time as a hands-on application security director at Charles Schwab, Product Unit Manager of Microsoft’s MSDN program and his more recent role as founder and CEO of SourceClear. Mark’s obsessed with building elegant products that solve hard problems for discerning customers.

Links:

  • Open Raven: https://www.openraven.com/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at the Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by LaunchDarkly. Take a look at what it takes to get your code into production. I’m going to just guess that it’s awful because it’s always awful. No one loves their deployment process. What if launching new features didn’t require you to do a full-on code and possibly infrastructure deploy? What if you could test on a small subset of users and then roll it back immediately if results aren’t what you expect? LaunchDarkly does exactly this. To learn more, visit https://launchdarkly.com and tell them Corey sent you, and watch for the wince.

Corey: If your mean time to WTF for a security alert is more than a minute, it's time to look at Lacework. Lacework will help you get your security act together for everything from compliance service configurations to container app relationships, all without the need for PhDs in AWS to write the rules. If you're building a secure business on AWS with compliance requirements, you don't really have time to choose between antivirus or firewall companies to help you secure your stack. That's why Lacework is built from the ground up for the Cloud: low effort, high visibility and detection. To learn more, visit https://www.lacework.com.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. A recurring theme of a lot of my nonsense has been finding hapless companies who have not been adequate stewards of the data with which they have been entrusted and giving them the ignominious S3 Bucket Negligence Award. That seems to be something that isn’t well-appreciated in some areas, so I figured, let’s have a conversation about that in a bit more depth. Today’s episode is sponsored slash promoted by Open Raven and I’m joined by Mark Curphey, their co-founder and chief product officer. Mark, thanks for joining me.

Mark: Thanks for having me.

Corey: So, let’s start at the very beginning. As a co-founder and chief product officer, that means that you’re one of those folks who very early on presumably had part of the idea, if not the entire idea for what the company does. What is Open Raven, and where did you folks come from? What problem are you aimed at solving?

Mark: Sure. So actually, it’s an interesting story. I had previously done an application security company called SourceClear that I sold to CA. My co-founder Dave Cole was the early product guy at a company called CrowdStrike, which recently IPO'd. And David and I had always wanted to work together; really didn’t know what to go do.

And the honest truth is we decided to go be good capitalists and went out and asked our chief security officer friends, “What’s the biggest problem that you’ve got?” And resoundingly, it came back that, “I don’t know where my data is. I don’t know what type of data I have. I don’t know how it’s being protected. And data breaches are happening all the time, and it’s probably the big thing that I’m going to get fired for.” So frankly, Dave and I rubbed our hands together and said, “I think we can make money off of that.” And solve a meaningful problem. And hence, the Open Raven company as it is now.

Corey: Which is absolutely something that is increasingly in the public eye. Well, we’d like to hope. At some point, people just shrug, give up, assume that everything about them is public, and that’s the end of privacy to some extent, and get on with their lives. At least, that’s the negative story. I like to believe that on some level, getting better than we are today is possible.

And what infuriates me, and why I started giving out S3 Bucket Negligence Awards personally, isn’t because you wound up getting breached. I view, on some level, that is being akin to taking an outage: it happens to everyone on some level, and you have to prepare for it as best you can. All right, I get that. One of the problems that we tend to see from all corners is that companies that wind up getting breached are, in many cases, exposing data that isn’t theirs, that no one consented to have handled by these folks. We see it, in some cases, with some of the credit reporting agencies and some of the data brokers. And it’s not always S3 buckets, but it is the consistent drumbeat of companies not being adequate stewards of the data that has been entrusted to their care.

Mark: Yeah. I mean, look, it’s certainly true that a lot of people have breached fatigue; this stuff’s been going on for years, and years, and years. I think that the S3 Negligence Awards, or the Bucket Wall of Shame kind of go back down to DEFCON hacker conferences. It’s called the Wall of Shame from passwords. It’s not necessarily a new phenomenon.

And I would also say that whilst S3, you know, you open The Register and every day, there’s an S3 bucket thing, it’s certainly not only S3. We know that; we’ve been doing some profiling of things, and Elastic, and MongoDB, and everything else is hanging out there. But I guess buckets, sort of, tend to be so easy just to make them open and host data on them in the first place. But I think you’re right: companies that have data, whether it’s knowingly capturing it or processing it, you have a duty of care, at managing and holding someone else’s data. And it just feels like people don’t take that duty of care seriously enough.

Corey: And what’s more, is that you’ll often see a company get breached, and, “Oh, your data has been subject to a data breach.” And ideally, you wind up getting that notification before you read about it in the papers. And a lot of the companies that you do business with, that contact you are very quick to point the finger of blame at a third-party contractor. Well, I didn’t hire the third-party contractor. You did, and if you’re not willing to wind up owning up to that, well, you’re effectively trying to outsource the work—which is fair—and the blame, which is not. How do you stand on that?

Mark: Yeah. Well, it’s a system that we happen to be using, but it was someone else’s problem, that the default configuration was their problem. I mean also, Corey, I can tell you I have a lot of friends in the forensics industry who deal with incident response, and still to this day, the vast majority of data breaches and never reported; I know of breaches that have happened in major public companies where all the breach laws are such that they should have notified their customers and they should have notified the authorities, and it just doesn’t happen. So, it’s one of those problems that I think it’s like the iceberg problem, right? And to a large extent, it’s kind of an interesting one in that when someone notifies their customers, they’re doing it from transparency.

And whilst I think you and I will both appreciate that and place more trust in those companies, the reality is a lot of the public wouldn’t. And so the incentives aren’t necessarily aligned up there around why, and why they should do it.

Corey: I would take it even a step further than that. I would argue that I don’t know if it’s a majority, but a significant number of breaches are almost certainly never detected in the first place. On some level the, “Oh, we’ll detect data breaches,” as a pitch that a vendor makes to a company is going to be met on some level was, “Good Lord, no. Why would we want to do that? We are happier not knowing.” And that depresses me.

Mark: Absolutely true. I’ve been building security tools for 20 years, and you’ll be surprised the amount of people that, if I deploy your tool, even as a trial and we find that we have problems, then I’m legally responsible for going and dealing with it, and I won’t touch it. The other thing that’s kind of related to that is that the security guys are incredibly busy as well, and the security tools, historically, generate lots and lots of noise; very low signal, lots and lots of noise. And so the security teams look at it and they go, “Oh, my gosh, I’m going to get a whole bunch more noise that I have to go deal with, and a bunch of more work that I have to go do. Can I bury my head in the sand?” Like, “Sure.” And it happens. That’s just the reality of the world we’re living in, unfortunately.

Corey: It is. And for better or worse, I think that it’s a world that we’re sort of stuck in to some extent. Do you think that the drumbeat of open S3 buckets that have been misconfigured containing sensitive data, it feels like we aren’t seeing as many of those as we once did, but is that just because people aren’t reporting them? Is it something that is going away slowly but surely? Or is it just as bad as it’s ever been, but it’s not making headlines anymore?

Mark: So, I’m actually building a tool to profile the AWS IP space for all the open buckets, and all of the open Elasticsearch and MongoDB thing. There’s a few of those that are out there, like Greyhat Warfare, which you can go search for S3 only, but not all the other stuff. I think when I looked up there the other day, I want to say there were about 750,000 open buckets. Now, of course, not all of those open buckets are—a number of them should be open. That’s kind of why they’re there and et cetera. So it’s—

Corey: Oh, I keep getting alerts constantly about open buckets that I have that are intentionally open, and I get alerts in the console, and I get emails at least quarterly, and this bucket that it starts with the word assets dot and then some domain, yeah. How about that? That is, in fact, designed to be open. On some level, I wind up—a few of them—just slapping a CloudFront distribution in front of it, not because I need it. Just because I want it to stop nagging me.

Mark: Of course. That’s the signal-to-noise problem again. And honestly, that’s part of the reason why Open Raven’s doing very well is that all these companies have been hiring people to go around, chase down open buckets, which are designed by functionality to be open in their organizations but don’t contain any sensitive data. So, I don’t know. Look, to answer your question around the S3 bucket thing, I think again, like breaches, people have got a bit of fatigue.

And there’s only so many articles The Register can do with, “Hey, S3 bucket open to the world,” and—what is it—27 terabits of data or 900 terabits of data. It’s not necessarily one-upmanship on the headlines anymore. But the gut tells me the faster we deploy things, right, the whole kind of DevOps movement, in many ways is moving against the security grain. And rightly so. So, to think that the problem is getting better, I don’t think is reasonable.

And then I think that Amazon, in some ways, have done good things with the security policy and all of those types of things, but they’re largely designed for greenfield environments. And when you step off the reservation, I mean, gosh, what, do you think the average developer is going to be out of figuring out how to configure that XML policy in an average XML bucket? Of course not. They’re just going to make the damn thing open, and they’re just going to move on and do their job. So, we’ve got to figure out how to get better tooling, easier, secure by default. And then, like you said, we’ve got to figure out ways to reduce the noise so that people can act on signal, and not get bombarded by noise, and just shut it down.

Corey: My approach to cutting through that noise on open S3 buckets, as I tweeted out a couple of times now, is to just copy a few petabytes of data into the open buckets. My operating theory is that while you’re going to ignore a politely worded email from a security researcher, you’re probably not going to ignore a bill that is 80 times larger than it’s supposed to be at the end of the month. That seems like it might be—among other things—legally fraught. What’s your approach at Open Raven to solving this particular problem?

Mark: Well, so for us, it’s all about improving the signal-to-noise. So, it’s about setting up a policy that you can say, “On this bucket, this is the type of data that I’m expecting; these are the security controls I’m expecting; if it deviates from that in any way, go, let me know.” And then we use OPA, Open Policy Agent, to go check that and go send alerts or pump stuff out to a firehose, pump it to a security event system, or whatever. So, in general, that’s it. It’s like, define what goods meant to be, and then let me know when good is not occurring so I can go figure out how to deal with it.

And of course, you know, you create generic policies like, “Look, I never want to see financial data on a bucket that’s open to the Internet and that’s unencrypted,” or, “I never want to see something with a CIDR range of 0000 and through some security group,” or something. So, that way, essentially, what companies get to do is, sort of, encode their intended use policy and alert when that’s not there. So, for us, it’s all about that. Part of what we see, and I’ve seen this in the security industry for the last 15 or 20 years, it’s there’s textbook security and then there’s the real-world security. You go look at a lot of, kind of, textbook security solutions, and they’re fine.

They work absolutely fine. I worked at Microsoft for a long time, and I was always amazed at how everything worked perfectly at Microsoft, and then when you stepped off the reservation, nothing worked properly. But everyone in Microsoft would be scratching their heads going, “Well, it all works fine here.” It’s like a developer saying, “Hey, it works on my laptop. It was only when I committed it to CI the problem occurred.” So, it’s that same thing, and you got to design stuff for the real world.

Corey: You also say it goes beyond just S3 buckets, which I believe. For a while there, I think—was it Elasticsearch, or was it Mongo that had a default password of ‘changeme’ or something horrifying like that?

Mark: Yeah. One of those. I forget which one? It was? Elastic’s a big offender, for sure. Mongo is a big offender, for sure. But I mean, you also see, like, Jenkins servers that are sat out there. I mean, that camera thing—what was it—the Verkada thing recently? That was a CI server that happened to have a script that had access to loads of things. But the amount of Jenkins servers that are accessible through the internet is shocking. It’s not just buckets for sure.

Corey: It definitely becomes a weird thing. I don’t know if there’s a fix here—I really don’t—longer term. But instead of looking forward for a minute, let’s go back and visit the past for a bit. You were the founder of the OWASP reporting list. What is OWASP? Is it a list? I’m most familiar with the OWASP Ten. But I’m certain you’ll have a better story on that than I will.

Mark: Yeah, no, no, no. Top Ten was this whole thing. So, I was running software security at Charles Schwab, early 2000s, 2001. Before, it was kind of a really big thing. And we used to get vendors coming in trying to sell me products.

And honestly, it was kind of a joke. My market open would have 8 million accounts, like, a trillion dollars under asset, and people would come in and try and sell me a web application firewall, which, maximum throughput was like 0.01% of my market open traffic, and things. But there was nothing out there on the internet to go point to and to say, “Well, this is good.” It was basically me versus a vendor coming in.

And so I said, “Right. This is kind of crap, right?” And I got together with a bunch of other people that were also doing similar things, some other people at some other banks, some other people in other companies. And I said, “Right. I’m going to go publish something.” And I wrote it over a weekend, literally wrote this guide, called the “OWASP Guide.” And it was basically a set of principles around software security, like lease privilege—you know, nothing sophisticated, but it was those types of things—and published it.

And then OWASP was basically born. So, it’s the Open Web Application Security Project. Then it got a lot of traction because a lot of people had signed up, and a lot of people were then starting referencing this to build their own application security programs. Over time, of course, OWASP got very successful. I think it’s, like, 40,000 people or something like that turn up at those conferences all around the world and chapters all around the world.

And there’s lots and lots of projects that have taken place, one of which is the Top Ten that you referenced, that a lot of people know of. And the Top Ten has been around I want to say since, like, 2004, or something like that. I don’t know, I’d have to go back and check with history. It’s hardly changed since 2004. And you can have a good conversation around why that is. But yes, that’s the history of OWASP.

Corey: And now, of course, you have this list that doesn’t seem to have changed significantly in a while. I mean, back when I was starting up the Meanwhile in Security podcast and newsletter with Jesse Trucks, we talked about that being one of the key problems is everyone wants to know how to handle security in cloud, but if we take a look at how a lot of application vulnerabilities exist, that list hasn’t materially changed. If anything, the advent of cloud has fixed some security issues, in that you’re not allowed to muck with them anymore. Datacenter physical security is no longer a vector for most folks who are all-in on a public cloud provider. But you’re also dealing with this other problem of, where, now it’s a list of enumerated S3 buckets, for example, and if you misconfigure that, it’s something that’s globally known, and I guess it removes the security-through-obscurity argument, insofar as ever was one. Has things changed in a time of cloud or is it just the same thing with new labels on it?

Mark: Well, I mean, there’s a couple of things. So, you’ve got to ask yourself, what is the OWASP Top Ten of, right? Is it the top ten most popular issues? The top ten most severe issues? The top ten voted by security people?

Like, no one's ever really been able to get to that, apart from an arbitrary top ten. And I don’t want to take anything away from it because the Top Ten has been incredibly useful in getting to developers, giving them a tangible, like, ten things; go focus on these ten things and you’ll raise the bar. So, that’s kind of piece number one, but has it changed? Well, should you have expected it to change depends on what you believe it’s based on. If you go look at them, though, like, no.

Things like injection, and broken authentication, and sensitive data exposure, those things haven’t changed because they’re just general things and they’re going to be around forever. You think sensitive data exposure is going to go? Doesn’t matter what technology we change, it’s always going to be there. For me, though, what’s kind of interesting about it, and why maybe I’m a bit of a skeptic about it is that you can eradicate total classes of problems—I believe—by changing patterns. So, a good example is, look, if you go use one of these modern development frameworks, application frameworks.

It’s built-in inherently. And the same with a lot of SQL injection problems that you used to see all over the place. You’d have to intentionally go create those problems, for the most part, now. And I think the cloud is done the same. It’s taken a lot of problems away, it’s extrapolated them into a service, it’s extrapolated them into a pattern, and the pattern can then go away.

So, back to the S3 thing, I think there’s hope [laugh] because if you can make a change upstream, I mean, you’ve probably seen recently all these damn, you know, supply chain attacks. And the bad guys are going further and further upstream where they can affect things downstream. And the good news about all of that is if you can figure out upstream, the way to go secure it, everything downstream gets secured as well. So, I think with a lot of these things, if we can, instead of trying to play whack-a-mole or put the finger in the dike, if we can start thinking about patterns and ways to go solve them at a class or problem level, then we stand a chance of fixing them.

Corey: This episode is sponsored in part by ChaosSearch. As basically everyone knows, trying to do log analytics at scale with an ELK stack is expensive, unstable, time-sucking, demeaning, and just basically all-around horrible. So why are you still doing it—or even thinking about it—when there’s ChaosSearch? ChaosSearch is a fully managed scalable log analysis service that lets you add new workloads in minutes, and easily retain weeks, months, or years of data. With ChaosSearch you store, connect, and analyze and you’re done. The data lives and stays within your S3 buckets, which means no managing servers, no data movement, and you can save up to 80 percent versus running an ELK stack the old-fashioned way. It’s why companies like Equifax, HubSpot, Klarna, Alert Logic, and many more have all turned to ChaosSearch. So if you’re tired of your ELK stacks falling over before it suffers, or of having your log analytics data retention squeezed by the cost, then try ChaosSearch today and tell them I sent you. To learn more, visit chaossearch.io.

Corey: I sure hope you’re right. I mean, in an ideal world, you will be. But it’s, ugh, I have so much trepidation [laugh] around all this. And I don’t know how it’s going to wind up playing out. And I hope that it’s going to go well. But it just feels like you’re constantly railing against the tide. And I don’t know how to wind up addressing that. I really don’t. I wish I did.

Mark: Mm-hm.

Corey: Is there anything you can say that helps them be more optimistic about this, at least?

Mark: [laugh]. Well, I mean, you’re right. Look, I’m no longer in the application security business after spending 15 or 20 years in there because I just gave up trying to convince developers to care about security. I just—and I don’t blame the developers. They’ve got another job to go do and security’s too hard.

So, for me, it was just pushing molasses uphill. And, I think, to your point, yeah, why would you expect anything different if we carry on doing the same thing? And the reality is, we’re moving faster and faster, we’re making it easier and easier to deploy things, we’re getting more and more complex systems. Why would you expect anything different? So, yeah, I don’t think you’re skeptical for a bad reason.

Corey: No. For better or worse, we still wind up having these problems. I don’t know how to solve it. I really don’t.

Mark: I mean, look, for the reality, if you go back to the old days, like, the old school—obviously I’m a bit of an old person, right—but you go back to some of the military things used to be, like, “Trust, but verify.” That motto works incredibly well. You trust people are going to do the right thing, you verify they’ve done the right thing. That means you don’t hinder the speed, but you go back and check and if anything happens, you come back. And it’s like, accepting things.

One of the other ones around that was, like, it’s people, process, and tools. People, process, and technology. And again, technology is never going to solve the problem of security. It’s a people problem. “You can’t patch stupidity,” and all of those phrases.

But if someone gives someone access to a local root account, or whatever the thing is, doesn’t matter how many other security controls you’ve got. I mean, I’ve seen it in cloud environments, as I’m sure you have. Someone goes and creates a security group, 0000 so they hop in the thing from home and don’t have to come in and go through all of the other control points. And it’s just the way stuff works. So, if you have that—if you take that mentality of, “People, process, and technology,” and, “Trust, but verify,” I think, use the right technologies and build the right process around it, then you can at least manage the risk. The risk is never going to be zero, but you can at least manage the risk to an acceptable level.

Corey: Let’s pivot a little bit and talk about the flip side of data security. And that comes down to privacy. There’s been a bunch of regulatory efforts around that. GDPR, for example, California has its own version of that that’s going out, and there’s also a growing school of thought that thinks, on some level, we’re post-privacy. Where do you stand with that?

Mark: Yeah. I mean, look, the privacy regulations are raging right now. You got GDPR; you’ve got CPRA, the California one; you got HIPAA, the Health Information Privacy Protection Act. And they’re all over the world. Japan has them, Australia has them.

They’re all over the place. And I think the US now is talking about having a central breach law around privacy data. The great challenge is that we’re all becoming a data economy, and companies are all becoming data companies, and so they want to gather more and more data. And the reality, I think, is that this whole stuff around cookie consent, I just think it’s just nonsense. When was the last time you said, hey, I’m not going to consent to you using my cookies?

It’s kind of like back in the old days, when you said, “Hey, I’m not going to allow JavaScript to run in my browser.” Like, all of a sudden, nothing works. And you’re like, “Oh. I’ll succumb.” But then before you know it, data it’s been over-reached, right?

You probably saw the Alexa the other day that has the radar so it can watch you sleep in your bed. Sure, of course, they’re not going to use that data for anything bad. But next time a breach happens, or some clever data science person decides to correlate something—I don’t know what it might be—in the middle of the night, it happens. So, I think what you’re starting to see is that you’re starting to see regulators and legal people who don’t really understand technology, regulating to prevent those bad things happening. And then technology trying to figure out how to go and meet those regulations, but meeting it with the absolute minimum bar versus trying to figure out what the actual intention is.

And I think you’re going to see a bigger and bigger gap. I mean, look at what happened with third-party cookies as an example. The whole third-party cookie thing we saw, what was that the CORS headers, we saw anti-cross-site scripting headers because all of those things started happening. And then what does everyone do? They just go call a tracking pixel.

And then all the marketing automation tools carry on working as possible. So, I mean, I think you’ve got a balance between technology working as intended in certain good use cases, and there are people using that for their own use cases, which either break or push over the line of privacy. I don’t know. How do you see it?

Corey: I think on some level, it’s not necessarily that people care necessarily that some company in the aggregate knows what they’re doing. There are some that do, and I’m not disputing that. But for most of us, I don’t necessarily care if Google, for example, knows what I browse on the internet. I care much more if you—personally—know what I—personally—am browsing on the internet. So, there’s a question of, once they have that data, do I really care that much about what they do with an aggregate? Not really? Do I care what they do about it on individualized basis? Kind of, yeah. And do I care if they’re making, then, that individualized data available to third parties? Absolutely.

Mark: Yeah.

Corey: It comes down to what the use of that thing is. Now, I know that I am not going to win friends with that particular argument myself. And I get it. In an ideal world, I think that advertising should be something radically different than it is. There are advertisements in this podcast, for example, and they’re catering to an audience that cares about the topics we talk about on this podcast.

But I have no tracking data of who listens to this, other than raw download numbers and rough GeoIP by continent. It’s not something that is ever going to be attributed—at least from where I sit—to individual listeners, nor would I want it to be.

Mark: Yeah. But, look, here’s where I might be able to convince you otherwise of that. In China, there is a well-known place called the Beijing Genomics Institute, and the Beijing Genomics Institute do genetic engineering, and not necessarily for good. So, it’s not necessarily to find cures for things, it’s also for other nefarious purposes. And the Beijing Genomics Institute acquire DNA data from US hospitals, US healthcare systems when you get your blood checked.

Now, that data is supposedly aggregated, but once you can start pulling apart DNA strands. You can start identifying people at different levels. And I think that’s the danger. There’s been a lot of cases where de-anonymizing information is possible. And so you’re making the assumption that that data is generally de-anonymized and use for the right reasons, but there’s been case after case where that’s not the case. So, maybe you’ll change your mind on that, Corey. I don’t know.

Corey: Maybe. I also, on some level, feel like I’m fighting a losing battle against the tide.

Mark: Yeah, yeah. My wife says, “Aren’t you worried about your credit card going missing?” And I’m like, “I’m sure it’s in many, many databases at this point.” I rely on Visa, at that point.

Corey: Well, that’s also a separate problem, too. I mean, this idea of, “Oh, your identity was stolen because someone else has opened a credit card in your name or stolen your credit card.” My very honest response to that is, “Oh. So, you weren’t cautious about who you decided to lend money to and validate they were the person you thought. And you’re trying to make this my problem because why, exactly?”

Mark: Yeah. I mean, look, in those cases, and that’s why it’s the corporate’s responsibility to deal with those issues. I guess it’s the same with social security numbers, in that they’re out there in so many places on the internet, and they’re pushed around in so many different ways, aren’t they? I think we’ve got to start moving into some of these zero-trust kind of protocols, and zero-knowledge ways, and all of that type of thing and the future.

Corey: Indeed. And I think that there’s one thing that every corporate entity listening to this—or representative of same—can agree on, and that is they prefer this conversation to remain hypothetical and aimed at the abstract not at them right after they’ve had a data breach, which of course brings us back to Open Raven and how it aims at these things. You do have a—at the time of this recording, it is still upcoming—a paper coming out contrasting what you have built with I believe it’s Amazon Macie?

Mark: Mm-hm. That’s right. Yep. Yep. So, when Dave and I founded the company, we went out, like I said, and we asked everyone, what’s the biggest problem, and it was data security. And then when you broke that down, it broke down into, “Let me know where my data stores are.” So, do I have buckets? Do I have stuff in RDS? Do I have stuff on file systems, et cetera? “What type of data do I have there?” “How is that data being protected?” You know, access control, and encryption, and all that things, and who has access to it?

So, it basically broke down to those things. Those things haven’t changed at all. So, think of that piece number two—what type of data do I have—as being data classification. Amazon have a service called Macie, which works on S3. So, we’ve built that feature.

Now, lucky for us, as it turned out—you get few really good breaks in the startup world—is that Amazon Macie it turns out it’s not very good, and incredibly expensive, and very, very slow. So frankly, the way we market it is, “Cheaper, faster and better than Macie.” And we believe in transparency of that. Every vendor will say we’re way better than everything, right? So, we’ve kind of done what you would do with a clinical trial in that we have basically built a—you know, here’s the test.

Here’s exactly what we’re going to test for, kind of like, laying it out in an academic paper. Here is the data, so you can go rerun the test yourself. And here are the results. And we know that we are way, way, way more accurate than Macie. We’re deployed as Lambda functions so we can scale up and run much, much faster than Macie. And then, certainly way, way cheaper than Macie, but that wouldn’t surprise you at all in that case would it?

Corey: No. Even after their massive recent price reduction, it was still, okay. That is in fact, still incredibly expensive, across the board. I mean, my argument with the original Macie and its pricing was I had a customer at that point, eyeing it and doing some math and, yeah, okay, first month would have been $76 million to run it in their existing stuff, which was significantly more than at that point, their annual AWS bill. So, it was, “Okay, let’s go with Option B,” which is literally anything except that and you’ll save money.

Even a data breach wouldn’t have been that disastrous compared to the pricing story. And now they’ve cut it to 20% of that, but that’s still an eight-figure bill to run these analytics on their data set. And that is… that’s not tenable. And on some level, it becomes the differentiated value of doing that isn’t there for customers. If I wound up running all of the various security services that AWS offers on an environment, it’s pretty clear that it would cost more than the data breach would.

Mark: Well, it doesn’t even work. Even if the cost thing was put aside, one of our customers tried it. I think they spent a million and a half on a trial, in a month, and it found 30 first names in a credit card database. I mean, it’s kind of crazy. And when you pick it apart underneath the hood, it’s a giant regex, essentially, and just doesn’t really work.

I mean, the reality is that that thing was built—it was actually a—it was originally an In-Q-Tel project, which is the funding arm of the US intelligence agencies. It was called [Harvest IO 00:28:04]. It was an acquisition that they bought in. And it was built a long, long time ago. If you want to do data classification today, you have to be able to not only identify structured, unstructured, and semi-structured data, and it comes in all places, and it goes into all file formats, in S3 buckets, it’s Parquet files—which are the backend of LakeFormation and Lakehouses and things like that.

But when you find a piece of data, you’ve got to be able to go and validate, is that data real? I mean, take an AWS API key as an example. It’s very easy to go figure out how to push that thing into that format, but is it a real key? Whereas if you use validators, go login to an AWS API and you’ll get a return that will say, “Is this a valid key, and which account is it associated with?” And so we’ve done, both in terms of the accuracy of identifying the information stores, the tests that we’ve got show, in general, we are twice or three times more accurate than Macie on finding the initial piece of data.

But then we have these validators. So, you know, you get a credit card, go call a credit card API. Is it a real credit card or is it just a 16 digit int? And you can go check that stuff. Data classification has moved on since that stuff was there.

So, even if the pricing thing was fixed—and as you point out, it certainly isn’t—it’s just not a good option for people. And then the kind of second piece to that is that the majority of customers that we see, and people, are looking at things like Snowflake. I mean, if you look at these data platforms, Databricks, Cloudera, Snowflake in particular, you know, they’re built on top of AWS services. But people are moving data to those places, so it’s not just an S3 problem, as I said. It’s about people putting data in Elasticsearch, in RDS, in file systems.

The data is everywhere—and backups. Like, all of this stuff gets pushed up into backups and stuff as well. And so you’ve got to have a service which goes and checks it. We decided to go compete with S3 and beat Macie first, but that’s certainly not where the tool and technology is going.

Corey: No, for better or worse, it would seem not. Thank you so much for taking the time to speak with me. If people want to learn more about Open Raven, what you’re doing and how you’re doing it, where can they find you?

Mark: Yeah. openraven.com is the best place to go. We’ve also got a pretty exciting open-source tool coming up soon, which is called Magpie. Magpie is a cloud security posture manager.

So, think of it as we’ll go out and check all of the security settings on your AWS environment. And so we’re releasing that open-source around the end of April as well. So, keep an eye out for Magpie. We’re taking the core out of Open Raven that does all the discovery across the orgs, pulls back all the attributes, or the IAM, or the security groups, and then allows you to go write security rules on top of that. Not data rules, which is what the Open Raven platform does, but security rules. So, also go check that out, but all linked off of openraven.com.

Corey: And we’ll of course put links to that in the [show notes 00:30:43].

Mark: Wonderful.

Corey: Thank you so much for taking the time to speak with me today. I really appreciate it.

Mark: No, thank you very much, Corey. Much appreciated.

Corey: Mark Curphey, co-founder and chief product officer at Open Raven. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you hated this podcast, please leave a five-star review on your podcast platform of choice along with a comment enumerating all of the S3 buckets you have inadvertently left open.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Heidi

Heidi is a transformation advocate with LaunchDarkly. She delights in working at the intersection of usability, risk reduction, and cutting-edge technology. One of her favorite hobbies is talking to developers about things they already knew but had never thought of that way before. She sews all her presentation shirts so they match the pajama pants.

Links:

  • LaunchDarkly: https://launchdarkly.com/
  • Personal website: https://heidiwaterhouse.com
  • Blog: https://medium.com/@wiredferret
  • Twitter: https://twitter.com/wiredferret

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at the Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part byLaunchDarkly. Take a look at what it takes to get your code into production. I’m going to just guess that it’s awful because it’s always awful. No one loves their deployment process. What if launching new features didn’t require you to do a full-on code and possibly infrastructure deploy? What if you could test on a small subset of users and then roll it back immediately if results aren’t what you expect? LaunchDarkly does exactly this. To learn more, visitlaunchdarkly.com and tell them Corey sent you, and watch for the wince.

Corey: If your mean time to WTF for a security alert is more than a minute, it's time to look at Lacework. Lacework will help you get your security act together for everything from compliance service configurations to container app relationships, all without the need for PhDs in AWS to write the rules. If you're building a secure business on AWS with compliance requirements, you don't really have time to choose between antivirus or firewall companies to help you secure your stack. That's why Lacework is built from the ground up for the Cloud: low effort, high visibility and detection. To learn more, visit lacework.com.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. So, this promoted episode is honestly one I’ve been looking forward to for a while. Three years ago, almost exactly from the time of this recording, I started this ridiculous podcast, and we’re now a couple 100 episodes in or so. But my first guest was Heidi Waterhouse, who is now back for more. Heidi, thank you so much for joining me. You are still at LaunchDarkly; you’re a transformation advocate. First, thank you for getting this whole ridiculous thing started. The world may never forgive you.

Heidi: It’s always good to be blamed accurately.

Corey: Exactly. It’s as long as you spell my name right, there’s really no such thing as terrible publicity, now is there.

Heidi: [laugh]. I feel really sorry for the other Corey Quinn on Twitter.

Corey: Me too because he worked in marketing; he had the name longer than I did. And he gets tagged every so often in ridiculous nonsense that I’m in, and at some point, I feel like he’s going to pivot to marketing in the cloud space just because of inertia. You can’t fight the tide forever.

Heidi: Go with what’s working.

Corey: So, LaunchDarkly. There are a few things interesting about this to me. The first is that—thanks, you folks are sponsoring stuff now, like this podcast. So, thank you for that. Secondly, you’re still there, which is not a ding on the company itself, I want to be very clear, but it feels like the average shelf life of an employee in tech these days is somewhere between 12 to 18 months. You’re over twice that.

Heidi: I am, and I am employee 20 or something. It’s kind of amazing, but it turns out that you can retain high-level employees if you treat them well and allow them to keep growing. So, I think that’s the thing people might consider taking under advisement.

Corey: That feels almost like it’s one of those ancient business practices that everyone likes to frown upon, like it’s something from our grandparents’ era, like make more money than it costs you to deliver your goods and services.

Heidi: That seems fake.

Corey: Exactly. One of those old-timey business models? All right, so we talked three years ago about feature flags. For those who have not taken the time to go through every single episode we’ve ever done, what are feature flags?

Heidi: Feature flags are a way to control your code, after it’s live in the world. At its most basic level, actually, that’s what LaunchDarkly does. Feature flags don’t actually have to work live in order to work effectively. A lot of people create feature flags as database queries or environment variables or runtime changes or settings. It’s a pattern that most teams end up needing, and they either recreate it or they buy it.

Corey: A line that I heard recently that reframed the entire topic for me—and perhaps you covered this on the first episode; I don’t actually know, I was way too nervous to pay attention to what you were actually saying back then—is it separates out feature toggles and rolling out features from your code deployment. And that one resonated because there are two kinds of shops out there, really: those who have terrible disturbingly bad code deployment processes, and those who lie about having the exact same thing. No one has a terrific code deployment process to go safely and repeatedly from developer laptops into production. And the only people are going to argue with that are the people who sell something that claims to do that. But it’s always scary. It’s always frightening. And having to do that to change what your application is doing feels like it’s a little bit—how do we put this—overwrought. Is that a fair characterization?

Heidi: Yeah. And the way I think of it is in sort of a Vegas thing. So, I come from an era where we had something called a gold master. We burned things on CD, and that was what you got, so we were really cautious about what we put into a release. And now we live in an era where you can YOLO stuff into production 100 times a day and not have a problem with it.

But that doesn’t mean that people want their website that they’re using to change 100 times a day. That’s bad for the business. We’re not just moving their cheese, it’s just popping around like a video game character. But on the other hand, if we do a release every hour, it gets a lot less scary because we have to have streamlined it. Like, if you can’t run your test suite in enough time to do a code push in under an hour, then you optimize your test suite so it gets better.

So, once we separate this idea of like, “How do we get things on to the server?” From, “How do we deliver the things to people?” We realize those are actually two different roles, and why did we conflate those?

Corey: I first learned about feature flags a while back at a presentation at a meet-up from some big tech company. And it felt, “Oh, that sounds like an awesome idea that you would run at a big tech company, but here in the real world, no one’s actually going to do it until you’re at that scale.” And talking to you the first time, it was clear that that is not strictly accurate. That is the whole point of LaunchDarkly. Three years later, has the industry changed? Is adoption of this pattern becoming more widespread? Are we learning new things as we go? And if the answer to that question is no, I’ve really just painted myself into one heck of a corner, but we’re going to try it anyway.

Heidi: [laugh]. This is honestly something I’m very proud of professionally. We have done, as a company and as a movement, so much to get the idea of feature flags in front of people and teach them that it is not just about how your frontend behaves and it’s not just about A/B testing, which is what most people think it’s for, but it’s actually about being able to mitigate your risk to be more cloud-native, to think in a way that says, “I don’t actually have a deterministic way to say everybody is getting the same thing at the same time unless I’m using a tool to make sure.”

Corey: So, is the idea of feature flags strictly a frontend concept? Is it something that winds up mapping to backend as well? And, God forbid, is there a story about using feature flags for infrastructure?

Heidi: Oh, absolutely. So, I think the most compelling story that I keep running across when I talk to people in Ops, is the idea of a permanent kill switch. So, most people, when they talk about feature flags, they’re talking about deployment, and rollout, and testing. Those are what we call temporary feature flags; you’re going to pull them out after they’ve done their job. But there are also long-lived or permanent feature flags that you put on bits of your infrastructure so that you can control things when shit goes sideways.

So, for instance, imagine you have an inbound API that’s writing to your database. This is a normal thing that happens: you sanitize the data, it’s a known sender, you’re accepting all this data, and you have a monitor on it. And all of a sudden—Datadog or whatever—your monitor goes off and says, “I’m getting 100x traffic. I don’t know what’s happening. Beep, beep, beep, I’m not happy.”

And the first thing that you want is to start shunting that traffic off because it’s probably a DDoS, or some other kind of bad data. So, rather than have to wake somebody up, figure out what the problem is, figure out where to shunt it, you can set up a permanent feature flag that says, hey, if this alarm goes off, I want you to shut all that data to this overflow database. And I will wake up and look at it, but first, maybe I will ingest some coffee or at least, you know, wash my face before I try and stare at a screen. That gives people so much more time to react in a smart way and uses our automation to sort of delay the need for the human in the loop. The human has to be there, but they don’t have to be there instantly.

Corey: The traditional idea of feature flags seemed like it was something that you would use to roll out experiments, on some level, to 1 out of 100 users of your site, and then you could start validating: Is this feature working? Is it breaking? Et cetera. Facebook, I seem to recall, had something vaguely similar. This is misremembering many moons ago, back when I gave the slightest crap what anyone from Facebook had to say to me about anything, just based upon ethical reasons. But they started off with, I think, seven concentric circles, or six concentric circles that spanned from a single developer’s account all the way out to the entire world. That feels like the feature flag story to some extent, isn’t it?

Heidi: It is. And Microsoft uses it, and Lyft uses it, and they all have teams that are doing that. And the story is, put it into production, but nobody can see it. And this works especially well at Facebook and Google because they’re using trunk-based development. So, everything is always live in their codebase, it’s just hidden behind different feature flags.

So, they put it out into production—they’ve deployed it—they turn it on for themselves, they see if it works; they turn it on for their team, they see if it works; they start turning it on for beta users. Microsoft calls this ‘ring deployment.’ And that really works for getting stuff out and making sure that it’s not going to be overwhelming in a weird way. And also, it turns out that even though most of our test engineers are amazing geniuses, and we should take them more seriously, you can’t really test how a distributed cloud environment is going to respond, except in that cloud environment.

Corey: So, the question I have, at the idea of expanding out from effectively just the developer’s test account, all the way out to the entire world regardless of who you are, is there an ethical concern here? I’m not trying to wind up putting you on the spot, but the idea of, I want to roll out a test to my paying customers, in many cases. The idea of chaos engineering and running experiments, and breaking production intentionally came out of Netflix, among other places, but at Netflix, the failure mode was on some level, “Okay, someone has to restart their stream in the event that something goes sideways.” That isn’t really the end of the world. But there is still a question, okay, you’re actually slightly degrading a paying customer’s experience. Given that feature flags are seeing adoption significantly outside of the entertainment space where the stakes are almost invariably going to be higher, okay do you stand on the ethics side?

Heidi: So, I feel like most of the life-critical and financial clients that we have, do manage to work with feature flags in a regulated environment because they already have rules about how to do that. It’s not like I’m saying because you can test on anybody, you can test on everybody. So, if you can test on yourself, that’s fine, but you still need an approval to test on any customers. There’s still an approval process. And in fact, we built in a new set of features that allow you to do approvals and say, “Okay, developers can try this out for themselves, and internal people, but it has to go past the approvals board if it’s going to hit any customer effect.” And actually, the thing that I say about this is that we are all testing in production, it’s just that some of us admit it.

Corey: That’s fair. Do you think that there needs to be additional scaffolding put in place before you can do this in an effective way?

Heidi: So, LaunchDarkly provides you a lot of that scaffolding, if you already have some concept of having to be regulated. I think that if you are a new financial startup, I hope that you are consulting best practices on how to set that up. I think that the thing about feature flags is that they are not a pattern, but a tool that can have many patterns. And that gives you the ability to say okay, the way we’re going to implement feature flags is with this extreme change management control, staging environment, soak time, like, we’re going to have all of these safeguards, or you can be somebody who doesn’t have to be that careful and you can move fast and YOLO things into production and see how it works.

Corey: So, on some level, what you’re saying is that the folks that are going to need to build additional scaffolding to do this responsibly, more or less already need to have built that scaffolding already and they’re already behind the curve to some extent.

Heidi: Right, exactly. Because how else are you going to say, “I can auditably and verifiably say that this person got this exact variant of the deployment, of the release?”

Corey: One of the problems I have with the idea of feature flags—that’s right, I brought you on the show on a promoted episode to basically tell you, “You know what your problem is”—and basically berate you for the way your entire product works because that’s how we roll here. But I have to ask, I wonder how I even begin getting started with something like this. I want to go ahead and test it, sure. It feels like I may have to wind up doing two, three sprints worth of work just to get into a position where I can even test something out. And at that point, it almost doesn’t matter what you’re going to charge me for a product or service. The sheer engineering time investment makes it a relative non-starter.

Heidi: Right. So, one of the things that we found is we are so frequently replacing some homegrown solution. So, people have the concept of being able to control how their software operates, they’ve just set it up in some way, and that some way is not scaling. One of the patterns that I’ve seen when people get started is they get started with something that’s not mission-critical because it’s easier to learn that. So, we have a customer called Xero who does a ton of payroll stuff in Australia and New Zealand.

And the way that we got in was, they wanted to use it for their financial transaction stuff but they didn’t want to do a ton of risky messing around with that. So, what they did was they’re like, “Okay, we’re going to buy a small license, and we’re going to use this to control our website. And once we’ve internalized how it works, then we’re going to use it for financial stuff.” And so these small non-mission-critical projects are learning labs for the team and then they can go on and share that knowledge with a more core business value team.

Corey: When we first spoke a few years back, my initial takeaway was, it sounds awesome. I can see the value. But it also felt like you were more or less running into the wind in that first you had to teach the market about the thing that you solved then immediately afterwards had to go ahead and sell them something. That always felt like a very heavy lift. But looking around now, I can see broad consensus in the customers I talked to about the understanding of the value of feature flags, and, “Oh yeah, it’s a reasonable thing that we should be doing.” It seems like in that respect, the heavy lifting has already been done, and on some level, it, “Oh, it just sort of happened organically. That’s just a natural evolution of the market.”

Heidi: [laugh].

Corey: It feels like that may have partially been what you're doing.

Heidi: Thank you. I have worked really hard. It turns out that category creation is enormously fun. I mean, you know, founder of Duckbill, a consulting group that comes in and says, “You’re doing it wrong and here’s how to fix it. And we’re not going to charge you more for being dumb.”

It was actually kind of a risky stance to take. And in the same way, LaunchDarkly came in and said, “Okay. We see this need and we’re going to explain to you why your life will be better after this.” And fortunately for me, CI/CD really took off, I think there’s been a ton of great work from the IT Revolutions people, putting out Accelerate putting out Project to Product putting out Sooner Faster Happier. This ethos, this zeitgeist of being able to move faster, and safely, is exactly the group movement that I needed to catch on to.

Corey: It’s a really neat thing to see the natural evolution of a product in a space, going from, “What the hell is this thing?” To, “Oh yeah, it’s a best practice and if you don’t do it, you’re probably doing something that is at least marginally dangerous.” It’s really something to behold and I have a hard time identifying other major players in the space that aren’t you folks. Not that I’m asking you to, because, “Yes, now let’s talk about your competitors,” is never a great look on an episode. But as you mentioned before, I strongly suspect your strongest competitor, the one that we all fight against commonly is, “I’m just going to build this internally myself.”

Talk to me about that. What does that usually look like because very often when I see people building things internally themselves, they don’t quite contextualize it in the context of the thing they can go out and buy; they view it as something different. As I look around the shattered remnants of my build system for the crappy software I build and deploy myself internally, what parts of that are going to look like, hey, that’s a feature flag option but I don’t think of it that way.

Heidi: Right. I think that this is actually a super exciting place for tools vendors to look at because it turns out that not only do people build it themselves in large organizations, they continue to build new things themselves. By my count, Google has at least ten different feature flagging systems.

Corey: Wow that’s almost half as many as they have messaging options that they release and then deprecate.

Heidi: Right?

Corey: I don’t know if we’ve seen the google messaging application for 2021 yet, but I’m sure it’s coming.

Heidi: I’m sure it’s coming. And then we will get attached to it, and then they will kill it off. I’m still salty about Reader so, you know.

Corey: As am I. They are never going to live that one down.

Heidi: Never.

Corey: Nor should they.

Heidi: It was rude. It was very rude.

Corey: It absolutely was. It showed a flagrant disregard for an entire ecosystem.

Heidi: The thing that I find when we’re competing with homegrown is that people are solving the problem in front of them, and it is absolutely true that it will take your engineers less than a week to code up some kind of feature flagging system, but it is also true that we have invested a ton of money in all this infrastructure so that we can serve flags at the edge in under 200 milliseconds around the world; that we have done a bunch of integrations, we have like 23 SDKs now; we have all of these abilities to hook into Salesforce and ServiceNow, so that you can have this seamless throughput and so that people don’t have to leave their native tools in order to use feature flags. And your developers aren’t going to replicate that and they’re not dedicated to researching where we need to go. So, when I’m competing with a homegrown solution I’m always like, “Yes you can build it, but it’s free as in puppies. You’re going to have to maintain it.” And that’s the expensive part of any software.

Corey: Any engineers who build this kind of thing just please skip ahead 15 seconds on your podcast players. Go ahead; do that now. Great. Managers, yeah, do you really want this sort of thing being built by the exact same people who built whatever horrifying monstrosity you’re using to deploy your existing software into production? Really stop and think about that for a minute.

Okay, now let’s talk a little bit about the future. Now, I’m not asking for roadmap information because that’s always in flux and no one likes to pre-announce things, but what do you think the future of feature flags is? Now, that it’s broadly accepted, okay’s it going from here?

Heidi: Individualization. The future is personal. And I think that the thing that we want to be able to do is let people set their own experience of their phone, and their web, and their smart home devices, and say, in all of the ways that we sort of have control now, we get more control. So, I have all of these things that I’m like, “I can change the settings on it, but not as much as I want.” And also, the settings are pre-assuming a bunch of things about me.

So, if I had the ability to do some of the things that we can do in browsers—like you can set your browser text to be something more dyslexia-friendly, okay—But not all web pages respect that. If I could force that, it would be awesome. I’m not dyslexic, but I want that ability, I want the ability to say, this is exactly the experience that I have chosen for myself.

And I’d love it to be portable, I’d love it to be a markup language because I think it’s a real accessibility statement to be able to say, this is what my web experience is like. And I have separated the content from the container. It’s a really old tech-writing concept that the words and how they are presented are almost entirely separated from each other. And in the same way, I want somebody to say, “I’m giving you this web content or this application, and how you present it is up to you.” And breaking that linkage is going to be so empowering for so many people.

Corey: Tell me a little bit more about this. I was worried when you said personalization because, “Oh, good. More creepy tracking of people.” But that’s not at all what you’re talking about. You’re talking about something that winds up transcending devices and sticking with a person, but for, I guess, the power of good rather than for the purposes of, basically, spying on people.

Heidi: Right. So, this is my futurist hat and not necessarily LaunchDarkly, like, end goal, but what I want—I’m not allowed to call it ‘Flag Markup Language’—

Corey: For the obvious acronym purpose—

Heidi: Yeah.

Corey: —of course.

Heidi: But Flag Markup Language follows you around and says, “I never want to see day mode. I’m only a night mode person, and if something appears in day mode, I want you to override it.” That’s like the simplest explanation of it. But it would also follow you around and say, “I’ve reduced the screen width.” Or—here’s an important one—I’ve taken out everything that makes this page very heavy because you are accessing it on a very narrow pipe.

Like, my parents don’t have cell phone service, and their WiFi is, well their rural internet is not great. And so, every time I visit a page that’s not a problem when I’m in the city, it takes a minute to download because there’s all this stuff. And I’m like, what if people could still get their ads through, but they were simple text instead of, like, video animation, based on the size of the pipe that is trying to go through.

Corey: I love the idea, but it feels, on some level, like that also requires broad-based acceptance across the board from every site that they visit, wouldn’t it?

Heidi: It would, or you’d have blockers. What it actually requires is broad-based acceptance by the browsers.

Corey: Got it. That feels like it is simultaneously easier and far more difficult all at once.

Heidi: Well, I don’t think it could be a solo project, but I think it would be a fascinating step forward in accessibility to have Lighthouse run and say, “Okay, but your page is not only inaccessible, it’s also too heavy for people who are on this bandwidth. Do you want a reduced fidelity version?”

Corey: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the Enterprise (not the starship). On-prem security doesn’t translate well to cloud or multi-cloud environments, and that’s not even counting IoT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IoT devices, detects these threats up to 35 percent faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at extrahop.com/trial.

Corey: So, one more question before we wind up calling it a show. You were one of the people that basically got me on board the train of giving conference talks from an iPad. It was transformative. It worked super well. I was really in the swing of it, and then the pandemic hit and I’m not traveling at all anymore and my iPad is mostly gathering dust here.

What’s your experience been on that? Are you still using iPads for any of your digital presentation works, or are you basically putting that on hiatus until it’s no longer taking your life into your hands more than it normally is to go on stage in front of a roomful of people?

Heidi: I think the earliest I will be at a physical gathering is possibly November. And honestly, conference organizers, you’re going to have a hell of a time doing anything this fall. And it had better be single-nation. We are not going to be able to cross borders.

Corey: Oh, absolutely. And beyond that, the first few conferences that rush to come back in person, you’ll be able to take a look at who attends those things and realize whose company considers them expendable.

Heidi: Right?

Corey: It’s a harsh thing to say; it is also accurate.

Heidi: Yeah. It’s really interesting to see how this is going to work out. But as far as my iPad, what I’m actually using it for now is I watch talks, I don’t live-tweet them as much because I’m watching on the iPad, and its multi-screen capability is okay-ish, but not great.

Corey: Adorable. I would classify it as adorable.

Heidi: Yeah, like, points for effort.

Corey: There was a solid attempt.

Heidi: But I will watch talks because it turns out that a lot of what makes me work is getting to listen to engineers, and developers, and ops people. And if I don’t have that input, I don’t have any grist in my mill. I figured out this is why I was having such trouble writing blog posts is because I wasn’t talking to anybody out in the world who was having problems. And I guess that whole developer advocate title wasn’t just hot air because what I really care about is making sure that those voices are getting represented.

Corey: There really is something to be said for what has happened to conferences in the past year. Suddenly, attending a conference no longer requires the ability to travel places, take time off, pay for your accommodations while you’re there, and in many cases, pay the not small fee for the conference. Suddenly, there’s a wealth of content that is available online, universally. And sure, the experience is relatively crappy, but it’s universally crappy. There’s not a better experience for some subset of people who are able to spend more, and the folks who can’t cover that expense are sort of forced into a substandard, degraded mode. This is what it’s like for everyone. And I have a hard time seeing how that is going to continue once the world opens back up. It’s one of the vanishingly few items on the list of things I’m going to miss, post-pandemic.

Heidi: Yeah. I was speaking at a conference based out of Russia. And there was a guy logged in from Kinshasa, D.R. Congo. Now, the thing that nobody really knows about me is I grew up, until I was five, in D.R. Congo. And it had never occurred to me that I was going to get to talk to a technologist from Central Africa because they can’t get visas; money is an entirely different thing.

And now in this new time, all you need is an internet connection to be able to attend. And honestly, I kind of choked up because there are all of these technologists all over the world who have been shut out for various reasons: because they can’t get childcare, because their company won’t pay for it, because they’re not like working for either a very large company or a cool Silicon Valley company. It makes me painfully aware of what we haven’t been providing all along, and I hope that conference speakers and me—this is a thing that I’m working on—will start doing more fixed time, like, here’s a webcast, here’s a podcast, here’s a conversation that enables everybody to access it.

Corey: On the other side of it—again, not from a engineering insight and knowledge perspective—but what I think is a slow-dawning awareness among a few people that I’m super excited about, is they’re watching me, effectively, livetweet slash aggressively shitpost various big company keynotes. And they’re finally realizing something that, wait a minute. Corey is not doing this because he had early access to what’s coming out and is ready to go with it. He doesn’t have special front row seats; he hasn’t been behind the stage talking to people before it goes out. He’s just watching this thing in one window, screen-capturing it as he goes, pasting it into his Twitter client, which is just the Twitter web app, adding some stupid commentary, and whacking send.

That’s the entirety of what I do, start to finish. And I apologize for just using the word stupid; let’s say ‘nonsensical’ instead. Let’s be a little less ablest. That is all it is. And I think that there’s now a slow creeping awareness that, wait a minute, it doesn’t require stupendous amounts of money, or access, or privilege beyond the normal level of privilege that I have in this space.

It’s just the ability to do it. And of course, the tremendous privilege I have of not being able to be fired for the things I say on Twitter, which is a separate problem entirely. But it isn’t because I have this magic ability to reach behind the scenes and grab things. It’s just, I’m doing it because no one’s stopping me from doing it. And I glad to see that other people are starting to take notice of that.

Heidi: Yeah. I don’t think you need to be an insider to be a good reporter. And I think that one of the things I’d love to see is for companies to hire some people that they haven’t been thinking about lately. My current campaign at LaunchDarkly is I really want us to hire a librarian. Honeycomb did it and I think it’s frickin genius.

Because we have all of this content, and we don’t have a good information architecture system to find it. And Confluence’s search is… not great. [laugh]. But I want us to hire librarians and I want us to hire investigative journalists, and say, “Look, we are developing so much stuff, but we need somebody to make the connections to make it accessible and usable, to make it something that you can use.” Like, you can absolutely teach yourself programming from YouTube, but where do you start?

Corey: It’s a terrific question. I think it starts with doing it. I wish I had a better answer, just because it’s, “Oh, I just go out and I do the thing that I do.” One thing I admire about you and I’ve always admired about you is you view a primary component of your job as being teaching people to do things. The problem I have is that so much of what I do is an outgrowth of things that have worked for me that it’s not easy for me to teach it because it’s just oh, just be yourself.

Well, most people aren’t themselves, they’re, you know]actually pleasant people. And for me, it’s always been hard to get to that point of being able to articulate what I do in a useful, constructive way. And I am extremely conscious of the danger of, “Oh, just do this thing because it’s what works for me.” That is a perfect recipe to unintentionally create a masterclass in how to be a white guy in tech who is swimming in privilege. Because I am, whether I want to admit that or not. And the things that work for me will not work for someone who is not themselves overrepresented. So, I am very torn on, how do I teach things? What do I teach? What can I effectively teach? And how do I not turn into a monster while doing it?

Heidi: Right. I think actually, the live-tweeting big keynotes is an interesting take because you can be an asshole and nobody is going to fire you. And also, the internet is unlikely to fall on your head because you’re a dude.

Corey: Exactly. My failure mode is a board seat and a book deal. I’ve been using that joke for a long time because it’s not really a joke.

Heidi: It’s not wrong. And I think it’s important to note that every once in a while I pull my punches because I don’t want the internet to fall on my head. And I’m a nice white lady in tech and have axis of privilege along that. So, I think it’s interesting when we’re talking to people who are under-indexed, and we need to be really careful when we listen to that feels unsafe. And one of the things I like about when you’re talking about how you’re funny online is, you talk a lot about punching up and not punching down.

And by sheer odds and percentages every once in a while you get it wrong, and then you apologize and try not to do it again. And I think that’s certainly a model I would like to see a lot more people adopt, where I don’t want people to necessarily be permanently canceled, but I do want them to say, “I caused harm and that was wrong, and here’s what I’m doing to fix it.”

Corey: And sometimes I feel like all I can do is try and lead by good example. And on some level, that’s really what the story about feature flags—and now, that’s a hell of a reach—is. It’s a—[laugh] it’s about setting examples and giving good demos. And that’s what you did. That’s why it became normalized.

Let’s stop beating around that particular bush. The reason that it became normalized is because people like you got up on stage, from an iPad, threw up some random website that you’d just thrown up, and said, “Here’s how feature flags are going to work. And here’s what it took to instrument it.”

And it turns out, not that much. You can do it in a live demo.

And then there was an interactive approach. And seeing that again, and again, went, “Wow, if you can use this for upscale shitposting mid-conference talk, done live, then what’s my excuse for not being able to do this thing in production?” And the answer is, “Oh, right. I’m bad at things.” But for most people, that’s not the answer.

It worked for them. And that, I want to be very clear, is no small thing. I’m not blowing sunshine up your butt when I say that you have fundamentally changed the way the entire industry thinks about feature flags. It’s true. It really is. One of the unique things about—you mentioned at the beginning—is that you have been at the same place for as long as you have, consistently on message but also dragging the rest of the industry forward, politely. You aren’t doing it with my brash way of getting up there and picking fights. You’re doing it by making people feel good by listening to you.

Heidi: Yeah.

Corey: If people take nothing else from this entire episode—forget the feature flags; forget the how to make your software better—if the only thing they take away is just watching you and using that as a life lesson in how to become a better person than they were, then this podcast, every episode has achieved more than I dream to when it’s set up.

Heidi: Oh. All right, but it’s a sponsored podcast. So, I’m going to say my thing.

Corey: Excellent. Please, by all means, take it away.

Heidi: I want you to feel safe about your software. And I want you to be able to do that while your software is operating in production. And the way you can do that is by being able to control it live in production. And if you can’t control it live in production, and if you can’t commit broken code, you’re not doing CI/CD. You’re doing some kind of hideous mini-waterfall. And so when you’re thinking about all of the things that keep you up at night about your deployment and your production, remember that there’s a better way to do it and it involves finer-grained control.

Corey: And again the whole point of this podcast is, I have a bunch of sponsors who say different things at different times. There is a bar: we don’t have sponsors on this show if I am not convinced that there are people for whom their product or service is the right thing. We have rejected sponsors on those grounds before. But to be clear we also don’t ringingly endorse the companies either. With what you do, and what I’ve seen, I do endorse it. I want to be very clear; and that is not something a company can buy, though some have tried.

Heidi: Well, you know, the sacks of cash that VC gives us are ours to spend either responsibly or irresponsibly, and since John, our CTO, is not getting his kombucha tap anytime soon, I think that what we’re going to do with it is do as much as we can to help people sleep better at night. And also—oh, this is a cool thing that just happened at our annual meeting. I just found out that LaunchDarkly is completely carbon neutral. We’re offsetting our AWS, we’re offsetting our travel, we’re offsetting our office space.

Corey: That is no small thing.

Heidi: And we commit to continue doing it. That’s just part of our budget now.

Corey: Well, I guess that really does sort of throw a wrench into the pivot option of starting to do one of those new NFT cryptocurrency things, now doesn't it?

Heidi: Man, I am so angry about that. I am so angry.

Corey: I want to own a representation of a jpeg, but I also want to burn down a forest, boil the oceans, and wind up effectively using more energy to do it than my house uses in 18 months. What have you got for me?

Heidi: I just don’t understand why this—I don’t. I just don’t understand anything that has to do with, I ran my car on idle for two years so that I could do a sudoku that is somehow fungible. The economics of it don’t make sense to me.

Corey: That is a whole separate podcast that I will get to one of these days. Heidi, thank you so much for taking the time to tolerate my slings and arrows and occasionally ridiculous compliments. If people want to learn more, where can they find you?

Heidi: So, find us at launchdarkly.com. We would love to give you a trial or a demo. And you can find me at heidiwaterhouse.com. And every once in a while, I update my blog but not on any regular cadence.

Corey: Excellent. And we will of course throw links to those into the [show notes 00:37:54].

Heidi: Oh, and the place that I really am is Twitter. So that’s—

Corey: As are we all.

Heidi: Right. That’s @wiredferret.

Corey: All one word, of course. Heidi, thank you so much once again. It is always a pleasure to talk with you.

Heidi: I had a great time. Thanks, Corey.

Corey: As did I. Heidi Waterhouse, transformation advocate at LaunchDarkly. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud if you’ve enjoyed this podcast please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast please leave a five-star review on your podcast platform of choice, along with an insulting comment with an embedded feature toggle so that you can wind up changing it to a glowing comment after it’s already been published.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

This has been a HumblePod production. Stay humble.

View Details

About Kendall

Kendall was the first hire at Fairwinds and has been in almost every role in the company. Today he works to establish Fairwinds as a essential name in kubernetes—offering software, services, and open source. Kendall has four kids, a dog, and three weasels. He also co-hosts a podcast on leadership with his friend Rachel at https://authorityissu.es.

Links:

  • Fairwinds: https://www.fairwinds.com/
  • kubernetestheeasyway.com: https://kubernetestheeasyway.com
  • Fairwinds Elements: https://www.fairwinds.com/elements
  • lastweekinaws.com: https://lastweekinaws.com
  • lastweekinazure.com: https://lastweekinazure.com
  • Fairwinds Insights: Fairwinds Insights
  • blatanterror: https://twitter.com/blatanterror
  • Authority Issues: https://authorityissu.es/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Chief Cloud Economist at the Duckbill Group, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by LaunchDarkly. Take a look at what it takes to get your code into production. I’m going to just guess that it’s awful because it’s always awful. No one loves their deployment process. What if launching new features didn’t require you to do a full-on code and possibly infrastructure deploy? What if you could test on a small subset of users and then roll it back immediately if results aren’t what you expect? LaunchDarkly does exactly this. To learn more, visit launchdarkly.com and tell them Corey sent you, and watch for the wince.

Corey: If your mean time to WTF for a security alert is more than a minute, it's time to look at Lacework. Lacework will help you get your security act together for everything from compliance service configurations to container app relationships, all without the need for PhDs in AWS to write the rules. If you're building a secure business on AWS with compliance requirements, you don't really have time to choose between antivirus or firewall companies to help you secure your stack. That's why Lacework is built from the ground up for the Cloud: low effort, high visibility and detection. To learn more, visit lacework.com.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’m joined this week by Kendall Miller, president of Fairwinds and, due to a lapse in judgment and both of our parts, one of my longtime friends. Kendall, welcome to the show.

Kendall: Thank you, Corey. I’m pleased to be here and continue that lack of judgment.

Corey: Excellent. So, we go back, and we will get into that story in a bit. But I’ve known you longer than I’ve been an independent consultant. You were there in my early formative years as a new manager. And I manage people in interesting ways.

There was a lot of empathy to it, but there was a lot of, shall we say, personality, and you had the good graces not to call me a jerk to my face, in so many words. Thanks. I wanted to make sure we got that in before I proceed to destroy what you’re currently doing professionally.

Kendall: Well, first of all, I appreciate that I also—even right there, I want to jump in with a story that, one time, in San Francisco, at a brewery or a bar or something with like, 25 friends, I had a friend who doesn’t work in tech show up and I was walking around the table introducing who everyone was, and this person works there, this person works there, this is what their title is, this is what they do. And I got to you. And I said, “This is Corey Quinn. He’s… a personality.” And I think that’s what you just described yourself as, and I do think that’s still maybe should be your title instead of cloud economist, just ‘personality.’

Corey: Yes, the problem is that ‘personality’ has a lot of implications to it, most of which are absolutely correct, but I prefer to let people discover that on their own. It was not in a bar or a brewpub. What it really was, was at a Chinese restaurant. And I remember this very firmly because we’re sitting around the typical white tech bros, as we are, surrounded by our friends who are fortunately not all looking like us. And the waiter comes by and you turn to the waiter mid-sentence, and completely switch languages and order in—I believe it was Mandarin, but it may have been Cantonese.

Kendall: Mandarin. Yes, probably.

Corey: Yes. And for the longest time, I had to do a fair bit of research to figure out whether or not that was actual legitimate Mandarin or an elaborate prank that you had staged just to make me fall for this and tell the story someday, I actually went in the back room had them send a different waitperson out who you would not have had time to bribe and yeah, sure enough, you can speak Mandarin. So, that was one in a long series of ways in which you surprise me. Every time I think I’ve got you dialed in, you go in a new direction, and I am forced to expand the ever-increasing multi-dimensional representation I have of Kendall Miller.

Kendall: I like to think that it’s all respect-based. But normally, I go into a restaurant like that beforehand and ask a few of them to speak Chinese to me and to listen to me speak and then tell everyone at the table that I actually speak Chinese when I really don’t.

Corey: Exactly. You know, there are dumber things people could do.

Kendall: [laugh]. If impressing friends, it takes a little bit of preparation, I’m in for it.

Corey: Exactly. So, all of that said, while we’re on the topic of dumb things, you are at Fairwinds, which is awesome. And you’re—the tagline for the company is, “Kubernetes done right,” but I checked the page very thoroughly. And you do in fact do Kubernetes which, from my perspective, is not doing it right. The only winning move is of course not to play. What’s the deal with that?

Kendall: Well, for a guy who spends his life criticizing AWS, if you believe that Heroku is the answer and solution to all things, wonderful. I mean, it does run on AWS, and you—marry those two things for me, what is the perfect solution, Corey? Is it all serverless all the time? Is it Heroku all the time? Because if it’s not Kubernetes, if Kubernetes is not your savior, then what are you betting on?

Corey: Operational excellence is sort of the short answer there. But Fairwinds is an interesting company—

Kendall: Bullshit. Whoa, whoa, wait, wait. Can I cuss on this?

Corey: By all means. By all means.

Kendall: [laugh] so, so bullshit. I mean, ‘operational excellence?’ You don’t need operational excellence if you have a simple enough infrastructure.

Corey: Yes] because if there’s one thing for which Kubernetes is renowned, it’s simplicity.

Kendall: Oh, that’s what I’m saying. You need operational excellence if you’re going to run Kubernetes. But if you’re not running Kubernetes, you don’t need operational excellence. If you’re on Heroku you just need a really good understanding of UI.

Corey: And you also need to have outsourced that operational excellence to someone else, which is not an invalid strategy.

Kendall: Oh, no, no. It’s actually an excellent strategy. If Heroku works for you, never ever, ever, ever leave.

Corey: No, Heroku works super well, credit where due—I’m not being snarky—right up until the point where it doesn’t. That point is further out than a lot of people think it is, but I’m—I have no problem with Heroku. If you’re on Heroku and think I’m bagging on you, I assure you, I’m not. We have built some stuff at The Duckbill Group on Heroku for very good reason.

Kendall: I’m with you. Yes. Well, and I regularly tell companies if it doesn’t cost you too much, and it’s not overly simplistic for your needs, never ever leave. It is a great place to be. But, [laugh] you touch on Kubernetes, and how can that even be done right?

And is that the non-starter for even doing things? It may be for certain people. And for a lot of companies, there is a serverless, or a Heroku, or something that is the right solution because be as no-ops as you can, by all means. Like, design things as simple, on great automated self-managed hosted systems wherever possible. I mean, the beauty of the cloud is you don’t have to turn the machine on and off yourself; Amazon will do that for you—or Google, or Azure, or whoever your third-tier cloud may be, as the case may be—but why not carry that all the way through to, they will make sure the service is up and running for you, they will make sure the connections are working for you?

By all means leverage all the things that you can. It’s just that at some point, companies reach a point where they require the ability to dig into that complexity and be hands-on themselves. And when that happens, where would you send them if not Kubernetes, Corey?

Corey: Well, I do want to call out, first and foremost, that there is a potential perceived conflict of interest that I want to be very clear that we express. I was an advisor to ReactiveOps—once upon a time—for almost a year when you were the president of that company. Then I stopped advising you folks, and it became pretty clear because you looked around and said, “Huh. What’s the biggest problem with the name of this place being ReactiveOps? That’s right, people have heard of it, so we’re going to call ourselves ‘Fairwinds’ without even throwing in a following seas joke to go with it.” And it is, in fact, the same company, correct?

Kendall: It is, in fact, the same company. Yes.

Corey: Because there was some great branding there, the pajama pants at KubeCon that were labeled ‘ReactiveOps.’ Genius. I love that idea. I wish I could steal credit for it, but I can’t.

Kendall: You have said to me—and I’ve thought about this a lot—you said ‘ReactiveOps’ wasn’t a great name, but Fairwinds is also not a great name and you forget about it by the time you get to the end of your sentence. And I don’t think you’re wrong, but I also think it was the right decision to change names. Now, we could have done it better. There’s a number of things we could have handled better on the SEO side. Like, don’t get me wrong; changing a name is complicated, messy, we learned some lessons, the hard way that I wish we hadn’t.

But at the end of the day, internal marketing matters a lot. Like, I think a lot about GitLab. GitLab is a really impressive product. In fact, it’s a really impressive suite of products, but most people don’t know that because the name is GitLab. And they’re the—

Corey: Oh, the fact that it has ‘lab’ in the name sounds like a science project or an experiment that’s still waiting to see its viability. It’s like, “Oh, GitLab. Is that GitHub’s development group?” Like, “No, it’s not. Well, well, kind of, but no.” And yeah, the fact that they’re toying with IPO, apparently, and are a multi-billion dollar company, they still have the word lab in the name.

Kendall: Well, and it’s still heavily focused on Git even though everybody knows it’s all SVN under the hood.

Corey: I want to be fair, Fairwinds—to be fair—Fairwinds is not a terrible name in the universe that contains things like AWS Trainium, or Systems Manager Sessions Manager. There’s always going to be a bad name that’s worse.

Kendall: Someone who names things worse than you? Yes.

Corey: Exactly.

Kendall: No, I appreciate that.

Corey: It’s just, I got to be direct, uninspiring.

Kendall: It is uninspiring until it’s been around enough and it has enough market traction, that it doesn’t matter. I mean, I am a big believer that this was a good decision. And I like working for a company that’s not called ReactiveOps. I was the first hire at ReactiveOps, and I asked in my interview, “Why is it called ReactiveOps, not ProactiveOps?”

Corey: Because ProactiveOps was taken.

Kendall: [laugh] well, and then throughout many, many years—I mean, there’s a reason it was called ReactiveOps. And it was a great name, it was the right thing to be called for a while. We were able to get the domain, we were able to grow to a certain number. Fairwinds, we had to buy the domain—it wasn’t free, right—and—like, the same way ReactiveOps was because it was catchier, even if it’s—has problems. But the beauty is, we can grow into it, and it can be anything.

And internally, we are allowed to think of ourselves as anything, inclusive of being an ops company, but not exclusive of everything else. Does that make sense? That’s why I really think it matters is the internal naming really matters a lot because people are affected by internal marketing. It’s hard to think outside the box.

Corey: They absolutely are. It seems like a weird juxtaposition because, credit where due, while the name is uninspiring, the company, in fact, is. The people I have met who work there have been nothing short of stellar in every case. It really set the model—to be direct—with how I wound up staffing The Duckbill Group. Like Fairwinds, we are full remote and we’re built that way from the beginning, not having it bolted on after the fact so you have basically two tiers of employees, or remote in the way that, surprise, everything’s now remote because of the deadly pandemic. No, no. We were full remote in the before times, as were you.

Kendall: Yes.

Corey: And that really led to some interesting conversations and some amazing hires you, quite frankly, otherwise would never have been able to get.

Kendall: Yep, agreed. Well, so I’m a leader in this organization; I know all of our warts inside and out. And there’s no such thing as a company that’s firing on all cylinders and perfect in every way. Although I’m pretty damn proud of where we are, and what we’re doing, and who we’re doing it with. And I tell people in interviews, we don’t hire everyone.

In fact, it’s a difficult job to get, but if you make it through the process, you’re going to like the people you work with. I can almost guarantee that. And to a person, we have a great team of people that are compassionate, that take care of one another; it’s an inclusive environment; there are no stupid questions. There was, one time about two years ago, where somebody said something passive-aggressive in Slack, and the company response was actually just laughter, throughout. I mean, just people DM-ing this around, just, just howling with laughter because nobody ever says something passive-aggressive in Slack. We have a culture of respect and I’m proud of that. So, we do a lot of things, right. The people that work here are one of those things. Very much.

Corey: And as I always said, the best way to run Kubernetes is not to, but if someone forces me to deploy Kubernetes, there are really two options. The one that I would prefer would be to go with you folks. I’ve seen how you run this stuff; it just makes sense. And it covers some of the reasons that people run Kubernetes, but not all of them.

Kendall: Well, so Kubernetes is hard, but part of the reason Kubernetes is hard is because it’s still new to most people the same way that moving from a Windows machine to a Linux machine is hard because you’re not familiar with Linux. And in the early days of Linux, you spent all of your time just trying to get it to work, right? Trying to make sure your screen actually had the right driver installed and had all the right settings. And I mean, it was a huge pain in the ass, I spent a lot of my childhood just trying to get different Linux distributions to work on my old Tiger Machines computer because that was entertaining, just trying to get it right. But you don’t want to have to do that with a production environment for a product that you’re running.

And so if it’s new, and the new is complicated, yeah, just look for help; we offer that help. But now, I mean, we’ve really changed a lot, Corey, even since you worked with us where we were heavily focused on services, and now we have a software product that gives people confidence they’re using it right. So, rather than go hire the experts to make the problem go away—please go hire the experts. Make the problem go away—but if that’s not going to be what you’re going to do, install a piece of software—I mean, we have an open-source solution out there called Polaris. It is widely adopted in the Kubernetes ecosystem, especially the open-source ecosystem. You run this on your cluster and it tells you things you’re doing well, and things you’re doing wrong, and it gives you a score.

And then we’ve built on top of that, including a bunch of other open-source tools heavily focused on security and policy enforcement, et cetera, so that large-scale enterprises can actually roll out Kubernetes with confidence because their engineers don’t know what they’re doing. They are giving people the ability to deploy things into Kubernetes that are horribly, horribly configured unless they have good policy in place and a software tool that enables that enforcement. Because at the end of the day, the reason this exists is we can build great infrastructure for people, but if what people are deploying into that infrastructure is terrible, it only gets you so far. And Kubernetes is difficult like you’re saying, but it doesn’t have to be if you got the right team behind you or the right software to help you. And that’s the end of my plug. No, it’s not. I’m probably going to say all the things [crosstalk 00:13:19].

Corey: Oh, of course—oh, you’re going to be self-promotional the whole way. If not, frankly, you’re not doing your job. But let’s be serious here. I don’t disagree with anything you just said. In fact, I endorse it. The problem I have is with the fundamental conceit of the entire argument, which is that people are attempting to use Kubernetes to get actual work done instead of dicking around. It seems to me that the reason that a lot of folks are going with Kubernetes is because they can’t pass Google’s interview but still want to cause play as a Google SRE.

Kendall: So, it’s resume-driven development? RDD?

Corey: Exactly. There are three or four great reasons to run Kubernetes and five thousand terrible ones. And it’s very often it feels that it is incredibly hype-driven in many respects because every time I tend to see it—that’s not fair. Most times that I see it in the wild, and I start talking to the people who have rolled it out on why you’re running Kubernetes. It goes back to talking points that do not ever tie back to an actual business constraint or problem that they were faced with. I mean yes, if I’m trying to run something hyperscale and I need to make sure that no individual system or rack or even data center could take down that service, yeah, something like Kubernetes makes a hell of a lot of sense.

But I’m trying to run a WordPress blog here and baby seals get more hits than this thing does, some weeks. So, for me, it is stupendous, stupendous overkill. But I see things that are about my level of complexity running in Kubernetes all the time, or let’s be fair, they’re not running in Kubernetes; they’re attempting to run in Kubernetes. Change my mind.

Kendall: So, well, there’s a couple things there. Is it hype-driven? Absolutely. But a lot of the hype is deserved. I mean, when our company was founded, when ReactiveOps started, we set out to build a framework for Infrastructure as Code, and we wrote a shit ton of Ansible and a little bit of Terraform, to go solve the problem of having automated deploys, blue-green deploys.

You know, everybody wants logging, monitoring, alerting, a system for their cloud. Everyone’s needs in the DevOps space are all the same. How they accomplish them is a little bit different. So, we wrote a framework. Again, tons of Ansible.

Kubernetes comes along, and we took a look at it, and it was a lot better. There’s a lot of things that does it just make sense? Is the API different? Is it complicated? Yes, especially if you’re new to it. But honestly, the same way, Corey, that you might spin up a simple Linux instance on Linode, or an AWS to go kick the tires on something or spin up a simple server, that’s easy for you because you’ve lived in the Linux water for a long time. And once you get familiar with it, it doesn’t take a long time. Same thing with Kubernetes. Should most people be deploying WordPress onto a Kubernetes cluster? No. [laugh].

Corey: Absolutely not. I’m hard-pressed offhand to come up with a worse idea.

Kendall: No, it is a terrible idea for so many reasons. But if you live in Kubernetes world, or you’re very familiar with it, or you want something to fiddle with, which is a legitimate reason to kick the tires with Linux is because you want something to fiddle with or Kubernetes, it’s a thing that you can go fiddle with; it’s a thing that you can go learn. The paradigms are new, they’re exciting, it’s fun. This is the way that the world’s going. In the future, all the Herokus of the world, every PaaS is going to be underlied by Kubernetes, every service you’re using is going to be Kubernetes almost everywhere, except for the few places where it really doesn’t make sense.

And I don’t think we’re that far away from that. Should you use the PaaS? Yes. But if you need a PaaS that you’ve built yourself, use Kubernetes. It’s the closest thing we have to a foundation or a framework for cloud infrastructure.

Now, that said, it’s really not a foundation. It’s somebody giving you rebar and cement and saying, “Good luck, buddy.” Right? But if what you’re doing with that rebar and with that cement, you can build a really impressive foundation that’s going to meet your needs for your very, very, very custom-built house. If you have a small house, a small family, no big needs, don’t buy a custom house.

If you just need something simple to live in, don’t buy a custom house. But if you’re a large enterprise, and you need to have dramatic control over all the different things and you want it to be a little bit flexible, Kubernetes does a pretty darn good solution, Corey. Change my mind.

Corey: You’re right. The fundamentally—

Kendall: No, what? No. Stop. We can just end the recording right there.

Corey: Oh, where. We’re just—cut it there. Good. We’re done.

Kendall: [laugh].

Corey: You’re not wrong on a lot of that. And the argument that I see is that you wind up with two sides girding themselves for war, you have the containerized side—which we can distill down to Kubernetes because regardless of what many of us wish happened, it is basically winning in the space—and the other side is, ah, serverless.

Kendall: You—wait, wait. You want a Docker swarm to win?

Corey: No, no. I personally ECS, if you—I still maintain kubernetestheeasyway.com and I have re-pointed it to the ECS product homepage, I will re-point that to the highest bidder.

Kendall: ECS is going to run on Kubernetes. More and more. It’s all—

Corey: Oh, yes. We’ll have that argument some other [crosstalk 00:17:56]. But there’s serverless on the other side—

Kendall: Yep.

Corey: Which is, you just wind up using a bunch of high-level managed services, pay for consumption. And the old-school admins are all very angsty about this. At that point, you’re just handing your availability over to your cloud provider. Well—

Kendall: Sure.

Corey: —no, you’re just being honest about it because you’ve been doing that for 15 years.

Kendall: Absolutely. And yeah, I mean, serverless is the absolute—okay, not the abs—I’m sure there’s going to be things that iterate on serverless but in the old days of, I have a computer running my server in my data center, or honestly, not even my data center. I mean, the startup I worked for in 2004, we had a back room, like, literally a closet with a server rack in it. I’ve taken this server with this install of this operating system and all of the things it takes to run my app, and I’ve given it to the cloud on an instance that now I have to manage in the cloud. And they just continue to abstract those pieces away to literally, here’s the workload; make it happen, Amazon; make my problem go away. Brilliant. Way to go cloud. Way to go serverless people. I give credit all the way back to the Fission.io folks, which I think were Platform9. I don’t think Platform9 talks about that much more, anymore.

Corey: I keep mistaking them with Plan 9. Talk about derivative names. But please, continue.

Kendall: [laugh]. There you go. Well, but—so, I mean, it makes sense. It’s brilliant. The reason to use Kubernetes isn’t because you have a workload you don’t want to worry about. The reason to use Kubernetes is because you have to have fine-grained control over some of the internal networking, some of all the different—you know, I need this to scale up this way, and that to scale up that way, and I need them to talk to each other in this way, and I need to have this control over that thing.

And should you use serverless? Yes. If you can make the whole thing work in serverless, yes, just do it. But in a few years, all the serverless everything is going to be running Kubernetes underneath, and that’s what I’m betting on. So, I don’t care if you run in serverless. Somebody is running that serverless system and it’s probably running on Kubernetes and they’re going to want help.

Corey: The problem that I see with a lot of this, too, is that okay, fine. You’ve convinced me. I’m going to run Kubernetes. Now, okay, and how did you say finding each other? Oh, they need to add something Istio or Envoy or—don’t correct me on that—and something else in front of it.

And then I pull up the Cloud Native Computing Foundation’s landscape. And some wit on Twitter just took a screenshot of that once, and tweeted it with a caption of, “Jesus Christ.” And it got something like 20,000 retweets because it’s hilariously overwrought. I look at this, and it makes the AWS service listing look reasonable. It’s that complex, and vast, and broad.

And there’s an entire universe contained within the things you need to responsibly run Kubernetes. And I look at it, and my entire position on it is, the hell with this. I can go back to running VMs on top of a cloud provider—or instances or whatever you want to call them—in a standard three-tier architecture, and that worked pretty well back in 2012. The world hasn’t changed that much.

Kendall: Well, so this is—you can blame the CNCF for some of this. Why did they create a landscape that literally includes everything? You want to submit something to the CNCF, you basically can; you have to sign a couple of agreements. But then it makes it look like all those things are the things you need. I mean, this goes to your tweet, just, like, yesterday, or the day before where you complained there is no enterprise Kubernetes distribution that excites you.

OpenShift is overfraught. Tanzu is complicated and it’s hard to understand. And Anthos is just a SKU of a whole bunch of Google products. I get it. I mean, we have something similar. So, we run Kubernetes at scale for lots and lots of companies, mostly leveraging open-source things. There is a finite number of things you need to go from Kubernetes to production-grade Kubernetes, and we have those packaged in a thing, on our website, in GitHub. It’s called Fairwinds Elements. It’s all open-source. Just go use those things. You don’t need more than that. If you need more than that, go get help. But there is a finite list of all the things you need to go from click a button, get Kubernetes to, click a button, get production-grade Kubernetes. And it should be easy, and nobody’s defining it easily.

Corey: It just feels, on some level, like Kubernetes is really aimed at people who want to cosplay as cloud providers themselves.

Kendall: That’s like saying Linux is disguised as cosplaying people who want to… I don’t know, run servers. I can’t, I can’t finish that. [laugh].

Corey: That is exactly what it’s for. It’s for people who want to run servers. That’s the problem with Linux as a culture.

Kendall: Yeah, well, so I’m just saying like, yes, it’s fixing the need. Now, here’s the question that I have, though, Corey. Talk to me about this. Google bets on Kubernetes—and there’s some debate about whether Google bet on that or the people who founded Kubernetes bet on that. But Google internally is still using Borg.

Talk to me about that. Why have they not bet on Kubernetes? Is it because of all the things you’re saying, that Kubernetes is overcomplicated and Borg is actually the solution, and we should be open-sourcing Borg as-is?

Corey: Borg, to my understanding, is so deeply baked into how Google does things internally, there’s no way it could ever see the light of day. And I also have it on good faith that Kubernetes being open-sourced is perceived as a strategic blunder internally at Google because once it’s an open-source project, they are discovering to their detriment that they can’t deprecate it.

Kendall: But why have they not then bet on it, or at least dogfooded some way significantly, internally? When I talked to a Google engineer, and I ask them about Kubernetes and they say, “I don’t know Kubernetes. I don’t know anything about it because I use Borg.” How’s that not a problem?

Corey: It’s a massive problem. It’s Google had such an advantage with being the home of Kubernetes that they are excitedly squandering as fast as humanly possible, from my perception.

Kendall: I mean, it’s amazing seeing the other cloud providers catch up to GKE because it wasn’t that long ago that we told every client GKE does it better. And there are—

Corey: Oh, my god. EKS was a punchline.

Kendall: [laugh]. I mean, we handle a lot of workloads on EKS now, and it has come a long ways, and it is a completely fine solution for the vast majority of people. And yes, for a long time, it was really, really, really painful. But it’s not anymore. They’ve caught u—I mean, not caught up, but they’re pretty darn close and honestly, sufficiently.

Corey: Incidents happen fast, but they don’t come out of nowhere. If they’re watching, your team can catch the sudden shifts in performance, but who has time to constantly check thousands of hosts, services, and containers? That’s where New Relic Lookout comes in. Part of Full-Stack Observability, it compares current performance to past performance, then displays it in an estate-wide view of your whole system. Sign up for free at NewRelic.com and start moving faster than ever

Corey: They’re not bad, I will say. At this point, there is no way in the world I would want to run Kubernetes myself on top of bare metal. That sounds like pain. I’d want to get some form of distro around it that doesn’t come with a team of seven people wearing suits trying to sell it to me. That’s the wrong kind of distro.

Kendall: But that’s all the fun of Kubernetes. You’re taking away all the fun of Kubernetes. Sorry, keep going.

Corey: I really am. But I want someone to run it for me. I don’t want to think about it. I get some crap for this sometimes. Someone thought that they were pulling a big aha moment that lastweekinaws.com runs on top of—duh-duh-DUH—GCP because they looked at what was spitting out. And my response was a polite form of, “Yeah, no shit. I pay WP Engine to run WordPress for me because I’m not irresponsible, and I honestly, past that, I don’t care where they put it.” I have so many other things in my life that I care about more than I do that. So, what’s it matter?

Kendall: If there’s anything that shouldn’t run on AWS, it’s Last Week in AWS Corey. I mean, the managed service is great, but that’s the thing is it doesn’t matter how great EKS is if everybody’s deploying terrible things into it, that are horribly insecure, that are set to use terribly way too many—you know, are requesting way too many resources and therefore costing you a fortune. Have I come full circle to, “Buy Fairwinds Insights?” Am I allowed to do that on this podcast? Because I feel like just plugging—

Corey: It’s all about the guest here. By all means, knock yourself out. I’ll talk smack about you on a separate podcast like—

Kendall: Deal.

Corey: —at some point I’m going to go through all the previous episodes, get them all lined up and do a mega episode for an hour and a half, “And now I contradict all the crazy horseshit that my previous guests have said, in one conversation.”

Kendall: Yes, well, you’ve been on my podcast and I just want to say that if you do that, I will go back and do the same thing to you. And I have way fewer listeners than you so it’ll work out great for both of us.

Corey: That works out well because—they say, what is the collective noun for white guys is a ‘podcast?’

Kendall: That’s, that’s, yeah—

Corey: Yeah, the collective noun for developers is a ‘merge conflict.’ But, you know, we all take what we can get.

Kendall: I think my favorite comment like that was, “Where do podcasts come from?” And it was saying, “Well, when two white guys like their ideas very much, dot, dot, dot…” and that’s really stuck with me. Well, so anyways, Corey, we’re coming up on time, I think, from your side. What not Kubernetes should we be talking about?

Corey: It’s adorable you think I’m not going to cut the hell out of this. We’re at minute three, Kendall.

Kendall: Oh, you’re totally going to. But I want to talk about something not Kubernetes-related. What are you working on at Duckbill Group that’s driving you crazy right now that you can share, or is really exciting to you that you could share?

Corey: Oh, the things driving me crazy? Talking to people like you. My God. I mean, I thought that would have been obvious.

Kendall: [laugh]. I’m the most delightful thing in your day-to-day.

Corey: It’s a growth year. We’re looking at expanding the audience; we have some things we’ll be launching in the near future. Nothing to disclose on that right now. We’re toying with expanding in different directions. One of the things that I’m setting for myself is that if we do any more newsletters or things of that nature, I’m not writing them. I don’t want to put more weekly toil on my plate. I can write well, or I can write a lot, but it’s hard for me to do both. Consistently.

Kendall: You sit and read through the AWS blog for a living, which sounds like literal torture. Well, so let me ask you this. You’re a personality, going back to my first story, right?

Corey: Jeez, you come on my show and insult me. I don’t get that very often.

Kendall: I—hey, [laugh] if I don’t insult you on your own podcast, am I actually your friend? I feel like you would think, no. [laugh].

Corey: No, no, it’s fine. Beating the crap out of me is kind of my thing. I’m like, basically the personification, you know, of AWS marketing.

Kendall: That’s right. I mean, I want to ask about this. How has being a personality paid off for you because it’s led to you being able to start a business. If Corey Quinn was a nobody when you start Duckbill Group, it would have been a lot harder to get your wheels off the ground, it would have been a lot harder to hire people. You have a brand that’s allowed you to build a company and in a lot of ways that not having a brand wouldn’t do. I mean, can you talk to me just for a second about how beneficial it is to have the brand that you have?

Corey: Uh, it’s a double-edged sword like most things. It’s nice to be able to go out there and tell a story and people are like, “Oh, you’re the guy from whatever.” It does get super hard when no one has heard of me, and it’s, “So, what do you do exactly?” And it’s, take a deep breath, and rattle off the newsletter, the podcast, the consulting, the Twitter shitposting, et cetera, et cetera.

Kendall: That’s why you just tell people you’re a personality. Keep going.

Corey: Yeah, that happens, but—and it is helpful, but it also means that on some level, it’s—this is going to sound weird—it’s very lonely. Everyone’s sort of engaging with a persona, where it’s—and they have this idea of me rather than me as a person. Like, everyone knows me, I have remarkably few friends. It’s a very strange mixed bag, there.

Kendall: I mean, it’s something that I have spent time thinking about, that the complexity of being known is that people come up to you at an event and they want to be in proximity to you, to say that they were rather than to say, “Hi” because they know you know them back. And the larger that percentage is of people who know you that you don’t know—or that ratio is—the more complicated that gets, I can see that as being lonely. I’ll make sure that next time I see you in person, I give you a big hug.

Corey: Oh, good. But as long as the pandemic is over, it’s fine. The other side of it, too, is that you get used to scrutiny a lot. Everything I say is controversial to someone, and it’s differentiating, someone getting upset because I did or did not use an Oxford comma in a tweet—which, frankly, is not an important battle worth fighting. Don’t email me—and the other side of it, which is someone gets upset because I refer to a group of people collectively as “Guys,” which is valid because that’s something that is exclusionary to folks who do not see themselves encapsulated in the term guys. I get it. I eradicated that word from my vocabulary and replaced it with folks and people can deal with it.

To all the way on the other end of the spectrum, which I’ve never actually had to deal with of, “Wow, your views on race are incredibly problematic.” So, regardless of what you say, or what you do, you’re going to get scrutiny, you’re going to get feedback and disambiguating into where on that spectrum any bit of that feedback falls into of can I safely ignore it because it’s irrelevant, or am I just thinking that because growth is painful, I don’t want to go through that? And are some of the ways that I perceive things actually regressive? It takes time and a commitment to improving, but it’s not easy because you get a lot of feedback. And if you’re not careful in moderating that and taking it to heart and evaluating it on its own merits, it can destroy you.

Kendall: Well, what’s interesting about that is it almost sounds like you had to reach a certain level of fame to have the normal level of scrutiny imposed upon, say, your average woman on Twitter.

Corey: Absolutely. Absolutely. And even now, let’s be fair here, I don’t have anywhere near that level of scrutiny directed at me even now.

Kendall: Sure. Yeah, that’s interesting. And does it give you more empathy, though, for people who make their living in the Twittersphere, that don’t look like you?

Corey: I don’t think I ever was missing that to begin with because I’ve have conversations with a lot of folks who have far more valuable things to say than I ever will and who are, frankly, better people across the board. So, I’ve always been very aware of that. And again, it’s uncomfortable becoming aware of the privileged one carries and that was something that was a definite—it takes an adjustment like anything else. I used to be very different when it comes to my views on these things than I am today. And it just, it takes empathy, it takes walking a mile in someone else’s shoes, and it’s transformative because once you see it, you can’t ever unsee it.

Kendall: Yeah.

Corey: And frankly, at this point, I wouldn’t want to.

Kendall: Yeah. Well, and it’s interesting because now we’re both in positions of power in our organizations, like, actual titles of authority—

Corey: Oh, yeah. I have an authoritative position in the industry, and you have an authoritative position because you’re one of the only people who have gotten Kubernetes to boot up and get the errors to stop scrolling.

Kendall: [laugh]. But it’s the authority in the industry that sets you apart there, too, and it comes with a weight that I know you’re aware of, and I’ve seen you—I mean, one of the things that I like about you, Corey, is I’ve seen a friend call you out for something, you asked a bunch of clarifying questions to understand what it was about what you had said that was wrong, and then you went and removed it because you humbly understood that. And I mean, frankly, that’s a big deal, Corey, not everybody does that. So, if you’re going to be a celebrity, at least carry that weight with a little bit of humility, which now I’m on your podcast brown-nosing. Which, if we can just wrap up, maybe—[laugh].

Corey: No, no. That’s much more expected and normal. We’re used to that. I can handle that.

Kendall: [laugh]. If we can just scroll back now and insert that, you saying, “You’re right. You’re right.” And then just end right there. That would be ideal, probably. Is there anything else you wanted to talk about, Corey?

Corey: No, it’s it’s—you’re the guest, I should be asking you that. Anything else you want to make sure we cover?

Kendall: [laugh]. Um, gosh, what else is going on in the world? I mean, I think it’s really fascinating watching the speed at which Azure is advancing. I think it’s increasingly proof that… I think there’s a lot of ways you can argue Google has some of the best engineering solutions in some of their cloud products. They’re the best—

Corey: Oh, yeah. Just ask them.

Kendall: Well, they’re the best solutions for some of the wrong problems. AWS is willing to build anything, even if it’s the wrong solution, as long as there’s a market for it. And Microsoft can just sell. In fact, it was a Microsoft person who asked me about my different opinions on the clouds, and I was telling them where I thought AWS and Google sat in the market, and they said, “Our only differentiator is that we can sell. We’ve been selling to everyone for forever, and we’re going to continue to be able to sell to everyone for forever.” And it is fascinating to me watching a cloud grow with the speed that Azure is because they have the Rolodex that they do. Nobody has that Rolodex. And that’s fascinating to me. I mean, how long until you launch Last Week in Azure?

Corey: Oh, it exists. When it hits enough subscribers and people care, I’m going to find someone to run it.

Kendall: Oh, wow. Okay.

Corey: I don’t want to keep it. My god. I’m just building the list because enough people will care. lastweekinazure.com. Sign up.

Kendall: Oh, wow, interesting. Okay, there you go. I didn’t know. I didn’t know. But I’m not surprised. You should, you should be there because, at some point, there’s going to be meaningful competition to AWS. And it looks like it’s coming from Azure, not DigitalOcean.

Corey: I would agree. But I don’t think that that market needs to be served by me. I think it needs to be someone like me in that space. I am not going to become that person. And that’s okay.

Kendall: It’s a different kind of snark to attach to Microsoft than it is to attach to Amazon, given the—

Corey: It’s a different audience.

Kendall: Yes.

Corey: It’s a different language in many respects, and there are people who could be much more authoritative in those customer relationships than I can.

Kendall: Yeah, I believe that. Interesting. And do you see any third-party or second-tier or third-tier cloud catching up, ever? Is somebody going to enter the space and make waves? It seems like it’s a little bit too late. It doesn’t seem like Oracle is going to catch up, or DigitalOcean is going to take over.

Corey: Well, yes and no. DigitalOcean and Linode are both doing interesting things. I mean, take a look at them. They’re not shrinking.
Everyone likes to say, “Oh, they’re just withering on the vine.” No, they’re not. They’re everywhere.

Kendall: But they’re not going to catch up either. They’re never going to be number two to Amazon, are they? Or—s I mean, that’s what I’m asking. Will they be?

Corey: Yeah, and isn’t that a sad fate that will only make hundreds of millions instead of many billions in a given quarter. I mean, that’s not a terrible life, from my perspective.

Kendall: It’s true. It is interesting how we measure those things where Google will kill off a product that has more revenue than the vast majority of startups do in their first ten years of business, but it’s such a small number compared to them, they’ll just shut it down. Not to pick on Google, who is infamously shutting things down, but lots of business units that do that in the Apples, in the Googles, in the Amazons. But that’s interesting, the way we measure that.

Corey: There are many paths to success. And I don’t think that it needs to be measured in the context of the GDP of a midsize country.

Kendall: Yeah, yeah. I agree.

Corey: Duckbill won’t get to that kind of revenue for another ten years. That’s okay.

Kendall: Yeah, well, and you’re going to experience an interesting thing, being a bootstrap company who’s trying to make money. And everyone who has venture money around you is going to look down their nose at you, which is a weird thing that—

Corey: And that’s a serious problem if VCs don’t like me. I mean, that—I don’t know what I’m going to do if I wind up in that position. I mean, I need the wisdom that only comes from winning a lottery once and then being able to tell me how I can win a lottery, too, someday.

Kendall: I mean, there’s some nice things about being able to leverage VC money and grow really fast. I get it. I think what’s amusing to me is when a founder backed by VC is looking at a person like you who’s growing a company profitably and thinks to themselves, “Wow, I’m way better at burning money than this guy is at earning money.” And that that somehow gives them an air of superiority. That’s, that’s the thing that amuses me. But our industry is a weird industry and everybody’s all the time trying to size themselves up compared to the next guy. And—

Corey: Oh, I’m an old-fashioned crotchety old man here because I have the kind of business model our grandparents would have understood.

Kendall: [laugh]. It’s true.

Corey: It’s like, “So, you haven’t—where’s your investment all come from?” It’s, yeah, it’s this magical thing called revenue and profitability.

Kendall: Yep, yep, yep.

Corey: Because honestly, I’ve got to be direct here. If I am solving people’s AWS bills and losing money in the process, I don’t think that I would be qualified to do the thing that I do. It’s similar—no joke—back in two years of re:Invent being an in-person thing in Las Vegas, I never would gamble when I was there because I didn’t want the optics of, “Isn’t that the guy that’s supposed to be really good at saving mon—
understanding large, complicated money things sitting at a slot machine?” It’s just the optics aren’t terrific.

Kendall: That’s hilarious. I’ve never thought about that. I’ve been at a re:Invent with you, and I don’t play slot machines because they bore me, as does most gambling, but it never occurred to me that you had the—

Corey: Yeah, if I want to look at flashing lights and get endorphin hits by pushing buttons, that’s what I have Twitter for.

Kendall: [laugh]. That’s right. When somebody hits ‘like.’ The thing is that you have to reach a certain amount of inertia before you get the endorphin hit that you need from Twitter. That’s why so many people fizzle out before they get a reasonable following.

Corey: Credit okay due, it took me seven years to get my first 1500 followers, which is what I was when I launched this place.

Kendall: Yeah, that’s impressive.

Corey: I finally cracked the secret of Twitter. And guess what? Ready? Here it is: be funny. That’s all it is. The end.

Kendall: I mean, is it even that? Doesn’t it show up all the time, and being funny is like a nice to have?

Corey: Okay, be funny frequently. There we go.

Kendall: [laugh]. Be funny, frequently. Yeah. I buy that. That works.

Corey: So, if people want to learn more about what you’re up to, and actually maybe see if your company can solve a real business problem they have, where can they find you?

Kendall: So, the company is Fairwinds. That’s Fairwinds.com as in, “The winds are fair,” because this is Kubernetes, and everything is nautically themed. See, Corey, there’s more to the name than you thought.

Corey: There is. And people want to keep up with you personally because they make the same terrible series of choices I do, okay can they find you?

Kendall: My Twitter handle is @blatanterror as in a mistake that was very obvious. And I also host a podcast on leadership, primarily highlighting people who come from underrepresented backgrounds in tech. And the podcast is Authority Issues. That’s authorityissu.es if you want to check that out.

Corey: Upon which I have guested, and vastly enjoyed the experience. The host, not so much, but I did.

Kendall: Well. That’s why I have a co-host is so I don’t have to be in your shoes in this situation and come up with all the clever things. I mostly just ask questions, and then when I’m having an off day, she carries the load for me which is delightful.

Corey: Excellent. Well, thank you once again for joining me. I appreciate it, despite what you may think.

Kendall: Thanks for having me, Corey, and I’m a little disappointed because if you didn’t appreciate it, I think I would enjoy the spiting you a little bit more. Spiting the professional spiter.

Corey: Kendall Miller, president of Fairwinds. I’m Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you didn’t enjoy this podcast, please leave a five-star review on your podcast platform of choice along with a comment explaining how that despite cutting this episode down to five and a half minutes, somehow Kendall still managed to irritate the living piss out of you.

Corey: If your AWS bill keeps rising and your blood pressure is doing the same, then you need The Duckbill Group. We help companies fix their AWS bill by making it smaller and less horrifying. The Duckbill Group works for you, not AWS. We tailor recommendations to your business and we get to the point. Visit duckbillgroup.com to get started.

This has been a HumblePod production. Stay humble.

View Details

About Laura

Laura Thomson is Vice President of Platform Engineering at Fastly. She is also a member of the Board of Trustees of the Internet Society. Previously, she spent more than a decade at Mozilla, leading engineering and operations teams, and was on the board of Let's Encrypt. Laura has spoken at many conferences worldwide over the last 20 years and is the author of best-selling software development books.

Links:

  • Fastly: https://www.fastly.com/
  • Twitter: https://twitter.com/lxt

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part byLaunchDarkly. Take a look at what it takes to get your code into production. I’m going to just guess that it’s awful because it’s always awful. No one loves their deployment process. What if launching new features didn’t require you to do a full-on code and possibly infrastructure deploy? What if you could test on a small subset of users and then roll it back immediately if results aren’t what you expect? LaunchDarkly does exactly this. To learn more, visitlaunchdarkly.com and tell them Corey sent you, and watch for the wince.

Corey: If your mean time to WTF for a security alert is more than a minute, it's time to look at Lacework. Lacework will help you get your security act together for everything from compliance service configurations to container app relationships, all without the need for PhDs in AWS to write the rules. If you're building a secure business on AWS with compliance requirements, you don't really have time to choose between antivirus or firewall companies to help you secure your stack. That's why Lacework is built from the ground up for the Cloud: low effort, high visibility and detection. To learn more, visit lacework.com.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Laura Thompson, VP of platform engineering at Fastly. Laura, thank you for joining me.

Laura: Thank you for having me on.

Corey: So, Fastly is generally considered to be one of the definitive names in the CDN world. Is that an accurate description of what Fastly is, or is it perceived internally as, “Well, we are a CDN, but there's also much more to it than that.” I always want to make sure that my understanding of the company isn't based upon things that are no longer completely true.

Laura: So, we like to describe ourselves as an edge cloud platform because we've gone beyond the CDN. We are now shipping these products like Compute@Edge—which is essentially serverless only really, really fast because it runs at the edge—and a suite of security products. So it's more than just a CDN. And I think we're becoming more and more general as cloud platforms go.

Corey: Credit where due, I had a customer a while back who was using Fastly as a CDN and had been using a bunch of these edge compute things, and they said, “Oh, yeah, we can't move that off on to a competitor at this point because of all that logic that's there.” And I said, “Oh, so it's a locked-in story?” And they looked at me as if I were simple and said, “No. Because it's awesome, and nothing else is quite like it.” “Well, yeah, I guess there is a certain lock-in story around building something awesome that other people haven’t.”

Laura: That's actually really interesting. We've been talking about—this the product folks that work because the perception that I have is that people choose a particular cloud provider—whether it's general compute, or whether it's CDN, or whatever it is, edge cloud stuff—for the unique features. For the most part, people aren’t really interested in the lowest common denominator stuff, unless they have a very straightforward use case, and people have less and less straightforward use cases over time. So, they want to use the unique features; they want to know what's special, what can I only do here, and that's the basis of their purchasing decision a lot of the time.

Corey: Absolutely. Any sort of cloud comparison analysis between which vendor we're going to pick that breaks down to, “Well, what size of instance is going to cost how much?” And then trying to equate that, like, that is so far from the relevant part of the story when you're doing vendor selection in a modern era.

Laura: Right, exactly. It's interesting for me because I've been on the other side of the desk a lot. My last job at Mozilla, I was partially responsible for making CDN buying choices, so I have pretty good insight into what goes into that. It's interesting to think about.

Corey: It really is. In fact, to that end, you were—in fact, you are currently a member of the board of trustees of the Internet Society. What is the Internet Society, first off?

Laura: So, the Internet Society is about making sure that people have access, equality of opportunity, and that standards are kept open. So, the Internet Society essentially funds the IETF, which is responsible for most of the standards that run the internet. And that is sort of the primary mission.

Corey: I feel like on some level, the default response to that is, “Wait a second, the internet has standards? Since when?”

Laura: [laugh]. Right. Well, it does. I mean, TLS and HTTP would be the two things that people would think of, I would expect. And if those were not standardized, we wouldn't be here having this conversation.

Corey: Exactly. For those who didn't grow up in the ’80s ’90s, et cetera, watching these things evolve, the things that we take for granted today, like different networks can talk to one another all comes down to interoperability around shared standards. But I remember the dark ages, where if you were on CompuServe, you couldn't talk to the people who were on AOL directly. And over time, those walls started breaking down, and a lot of work of the standards bodies like the Internet Society and the IETF are the reason why. It's one of those boring governance things that happens underneath the hood so no one has to think about it. But the reason no one has to think about it is because people like you are doing the hard work of making it that way.

Laura: That’s right. And standards is a thing I've been passionate about for a long time. It is one of the reasons that I was at Mozilla for so long. It's one of the reasons I'm at Fastly, which has, I think, a similar role on the other side. You know, if Mozilla is thinking about how do we make clients use open standards, Fastly is thinking about how do we keep the internet open, and standard, and useful, and safe for everybody?

And [ISoc 00:06:00] and IETF are obviously, in the position of figuring out how things talk to each other. And it seems like maybe this is a solved problem, except it's not. One of the things that's happened on the internet—which as you can see over the last few years is that we're actually getting less and less open. So, we have more walled gardens, we have the internet being dominated by a few really large vendors, and they have a lot of power and standards. If you have the vast majority of the market share in a particular vertical, then you get to dictate what the standards are, and you can do that in a fairly anti-competitive way. I'm deliberately not naming any names here.

Corey: Oh, absolutely. I wouldn't expect you to, but I sure can. There's a reason that I have a podcast here where I control the RSS feed that I can point wherever I need it to live; there's a reason I have an email newsletter; there's a reason I don't have a Facebook page, to be very direct, where it's not about de-platforming in the censorship sense, but more along the lines of if I have built an audience on a platform, and then that platform decides that its business strategy is going to shift—something like Medium, for example—then suddenly, I am beholden to the whims of that provider unless I want to start over from scratch again. That's why it's always been so important for me to build my audience on something that was much more agnostic, not because I'm posting garbage that should be taken off the internet—well, most weeks—but rather because I don't want other people making my business decisions for me.

Laura: Yeah, so my interest in the internet is not about de-platforming, it's about platforming. It's making sure that people have access and the ability to build awesome things as much as they can.

Corey: Absolutely. But you're also a VP of engineering, which I feel like if I had asked you a year and a half ago, “What does it mean to be a VP of engineering?” You would have given me an answer. And now I feel like if I asked that same question in these uncertain times, or these unprecedented times, depending upon framing, you might have a different answer from that. Is that accurate?

Laura: Yeah, I think it's definitely been—I’m really tired of the word unprecedented. It's been a year. [laugh]. It's been a heck of a year.

Corey: It really ha—it feels like it's been longer than that.

Laura: Yeah. So, I started at Fastly on the ninth of March, 2020, which was a week after Fastly closed the offices for COVID. So, I have never set foot in an office, other than in the interview process. And I haven't met almost any of my coworkers. It's super strange, and it's really hard to build those trust relationships without having met anybody.

So that's kind of the first thing, and obviously, you would have a different experience if you'd been there and seeing the change, but this is all I have known in this role. So it's actually pretty interesting.

Corey: I know that there's a lot of zeitgeist awareness around what it's like to be an employee in a pandemic now and being full remote and the rest, but something that doesn't get discussed very much is, how does leadership change when suddenly you aren't able to gather people in a room and have a conversation and hash something out when everyone's remote?

Laura: Right. So, the funny thing is, that part hasn't changed very much. And I will say—for me, I've been a remote employee for about 15 years. And this is not what remote is normally like. I think it's part of the, sort of, important takeaway for people who might be doing it for the first time.

There's a big difference between choosing it and having it thrust upon you, for one thing. When I give people advice on how to work remotely successfully, it's always about, have good boundaries, have a dedicated space, and make sure that you start work at a particular time, finish work at a particular time so that you can walk away and de-stress. And because a lot of people are kind of trapped in the houses, they can’t have as good boundaries. They might have a one-bedroom apartment in San Francisco, and it's all in one room, and their spouse is there trying to work as well, and maybe they have a dog and a baby. It is not a typical remote working experience.

Similarly, it's not the typical remote leadership experience because remote is great. It also is important to meet with people once in a while, and not for meetings—not all sitting around a whiteboard talking through a bunch of dot points--but for the part that is really hard to do remote, which is the relationship-building. To me, you are trying to get to know people so that when things go wrong, they say, “Oh, I know Laura. I know what she's like. I trust her to do this right.”

And when I’ve never met you, it’s a little harder for people to do that; everybody has a little work a little harder. And when I say a little harder that's in the face of all of the cognitive load that we're already under because it's been, as you say, a heck of a year.

Corey: Right. And there are a lot of negative examples about how all this stuff looks terrible when people get it wrong. For example, “Oh, you're not spending an hour commuting every day; now you can spend that time working.” Or, “We have a policy of turning your webcam on, which is a fancy way of saying I'm inviting myself into your home so I can critique it. Excuse me, it's not acceptable that your infant is crying.” At some point, you hear these stories, and it's the biggest gap we have is the ability to strangle people over the internet now because that's horrifying.

Laura: I know. We have done a lot, I think, to make that comfortable for people. I think part of it is, Fastly has been at least sort of, half remote for a long time. [unintelligible 00:10:50], like remote-first companies are good to work for I will say, in general. But making sure that people know that it's okay if you have to have your camera off, or if you have to have a baby on your lap.

In fact, the thing I found is there is nothing quite like a meeting that has a cat, or a baby, or a puppy. And it lightens the mood, it helps people talk to each other. So we're all surrounded by our emotional service babies, and cats, and dogs. It's quite funny, really. I'm going to tell a story, which is—people talk about CEOs, and we have a CEO named Joshua Bixby, who has a really lovely human being. And when the pandemic started, he started a weekly meeting where he would read picture books to people's kids.

Corey: That's an amazing idea and in fact, I’m debating stealing it and claiming I came up with it myself, except that we're doing this on a recording, and suddenly my team will know when this comes out.

Laura: Yeah. I just thought it was an incredibly nice thing to do. And making sure that people knew that we're all in it with our families, we're all stuck at home with kids and whatever, and let's try and make the best of it and get to know each other a little bit better.

Corey: On the one hand, I absolutely agree with the sentiment and the place it comes from. On the other, I've got to say I have an anti-authority streak in a big way. My single biggest stumbling block, along with my personality, back when I was an employee. I mean, my last job was at a regulated finance company. You can imagine how well that worked out for me.

But authority is a problem for me, so whenever I hear, “Oh, go ahead and bring your family into this social gathering we're doing,” my immediate knee-jerk response is, “Is this required?” And that's not helpful as far as the sentiment goes, but it was there. That was my initial flare-up reaction. And I'm always hyper-aware of—now that I managed before myself of being sure to never present as anything other than if you want to.

Laura: Yes. Yes, it's very much been that way here. I too have an anti-authority streak, which is a funny thing to say when you end up in leadership. But that's why, by the way. And I think sort of compulsive joining-ness annoys a lot of people. I'm not one of those people that it annoys. I like hanging out with other people. I'm really extroverted, which may surprise you in an engineer and someone who chooses to work remote, but I like people, right?

I like hanging out with people; happy to hang out with everybody, but I know that not everybody wants to and that is 100 percent okay and you have to be so clear that that's okay. Everybody is kind of walking their own road through this thing. And for some people it’s… [sigh] some people are actually pretty happy being on their own, doing their own thing, so that's pretty important to notice as well and not try and invade people's privacy, and keep it strictly work if that's what you need. And if you need to, sort of, get to know people, then that's okay, too. Part of, I think, being able to lead well is to code switch to what people need. So, one style of management doesn't work for everybody. You have to figure out what works for each individual person and work with them in that way.

Corey: Impedance matching, almost. I periodically reference on this show a boss I once had, who, as I described him, spoke only in metaphor. Where—

Laura: Oh my.

Corey: —that's great. I don't understand what the hell you're talking about. Should I be doing more of this thing or more of that thing? “As the boulder crashes down the mountain through the stream”—it’s like, “Okay. I'm sorry, go write haiku in your own time. I'm trying to figure out exactly what needs to happen. Am I doing well? Am I about to be fired? It'd be really nice to know where I stand with you.” And I never got a clear answer.

Laura: Wow, that's rough. It's funny how those things work out. I once had a boss who had very little in the way of facial expressions. Just the way that their psychology worked; I would venture to guess that they're probably not neurotypical. And at the beginning, I found this, like, super intimidating.

And then I figured out that it didn't matter if you couldn't read his face because he would tell you exactly what he was thinking. And without emotion; it would just be all very factual. But there was no sketching around the issue, there was no trying to figure out what he meant. He would just tell you. And that was actually very relaxing once I figured that out. It's strange because if someone had said, “Would you like to work for someone like this?” I would not have said, “Yes.” But it was actually kind of refreshing.

Corey: There's something to be said about having a sense of—I know people talk about this a fair bit—of psychological safety in employment. And you need that sense of psychological safety, but you also need a sense of job safety. I know that even now, four years into running my own company with my business partner, whenever one of us sends the other a message of, “Can we talk?” And then goes quiet, it's, “Oh, am I about to be fired?” is the instinctive, immediate reaction every time. Even though neither one of us can be fired, it's still—that's a trauma that leaves scars.

Laura: It’s actually really terrifying, I think. I totally agree with you. And having had to do that, reasonably often. When you just need to ask somebody a question. And I think, I end up putting something in Slack that'll be like, “Hey, I need to ask you something quickly, but this is nothing bad. Don't worry, this is not a scary VP coming to be scary.” I find myself qualifying it like that because it's not my goal to make anybody's heart rate spike. [laugh]. They can watch a horror movie if they want that, but they probably don't. Nobody needs any extra stress right now.

Corey: Exactly.

Laura: Making sure I only inflict stress intentionally, and that's very rare.

Corey: So, one of the things that you spoke on, back when conference speaking was a thing—was a periodic focus on ‘Minimum Viable Bureaucracy,’ and as I mentioned, for someone who has a problem with authority, just the very phrase is appealing. Tell me more.

Laura: The reason I started with Minimum Viable Bureaucracy is that I, too, have a problem with authority. I have a problem with process. I'm really goal-oriented, and I don't really mind how we get there. And it's been hard for me to learn as an adult that you actually do need some process. So, figuring out what the balance is of what is the level of process that is the smallest amount required for things to run smoothly, to be predictable, without driving all the people that you work with bonkers. Because there's a lot of engineers who don't like authority, but there are people who also need process, so what's the balance? And there's a few basic principles to that.

One of them is, first of all, to push decision-making down to the edges so that people who know the most about something can make a decision about that thing. That's really important to me. A second principle is that you should always iterate on your process like you would on your code, or in your infrastructure, or in your products. You don't expect to ship v1 and walk away and never improve it; you don't expect to ship with version 1 of your config for your infrastructure and walk away and never improve it, but we tend to get stuck on process. If you are doing something and it seems terrible, then stop and don't do it; do something else instead. I think the ability to change something on a day-to-day basis, to iterate quickly, essentially, to continuously deploy the way that you work is super important to happiness.

Corey: There's also the risk aversion approach where whenever something breaks into a particular way, “Ah. We're going to add a process in to make sure that never happens that way again.” And keep iterating forward, and eventually, you're at a point where there's six thousand things that need to be verified at every point and it becomes unwieldy. It becomes almost ossified and mired in that process.

Laura: Yeah, that's absolutely true. And knee-jerk process is the worst. I think we actually have a really good team here that does incident management, and part of their job is to figure out, well, what parts of this actually require a change? Other changes we could make, other changes we shouldn't make, and to be really strategic about that. And it's super exciting to work with them on that.

Having said that, I think you can target things that need process. And the way I always say this is you should make the boring things boring, and that covers everything from the promotions process: everybody should know how it works. “How do I get promoted?” Everyone knows. It's straightforward.

It's the same every time. Maybe there's some iteration, but in general, it's well understood. It's true for deploying code as well. It should be boring, it shouldn't be exciting. I want the exciting things to be exciting, like the new thing that we're shipping, or the new tech that we're working on, or hiring somebody awesome. And I want the things that should be the same every time to be almost invisible.

Corey: You're the engineering leader I'm not, so you will almost certainly have a way better take on his than I will, but it feels that some organizations approach process and procedure from a position of if we just put enough process around this, we can finally not have the creative expensive types do these jobs, and instead just wind up turning it down to someone who follows a script and that's it. Is that the actual intention? Is that how it just manifests, or am I completely missing something fundamental?

Laura: The thing you've described is not valueless; there's one thing I think it's important to note, so I'll come back to that in a second. But in a place where you don't have any process—everything is terribly artisanal and bespoke and so on, you end up in a situation where you can only have really, really senior people working there, people who have the experience, and judgment, and know how to improvise. And what you have done then is you've made it impossible to have junior engineers, or interns, or people who do process. So, you can go too far the other way. And, to me, it's super important to have up-and-coming people, like people who are relatively new to the industry because they bring new ideas and also, who's going to run it when we all retire?

It's super important to have junior people coming up. And if you have, sort of, absolutely no standard ways of doing anything, it's going to be really, really hard for them to be successful. So that's actually perhaps a really strange reason for having processes, but that's one of the reasons. It's certainly not unreasonable to say, “Let's have the expensive people do the hard things.” Like that's certainly one way of thinking about it. But it's also so that not everybody has to be stressed out by not knowing how things work all the time.

Corey: And that's very fair. I want to be very clear: here at The Duckbill Group, we fixed the horrifying AWS bill, and originally everything was bespoke because it was just me as an independent consultant. And, yeah, I can keep it all in my head. Why not? As we started hiring people, we built out processes and procedures around how these engagements go, how the analysis looks, but it's still nowhere near the point where someone who's not conversant with the relevant technologies, and the relevant terms of art, and the relevant financial strategic requirements of businesses would be able to perform effectively in that.

So it's the, let's get some standards around here and let's make sure that we're not missing things, but it's also never been aimed at driving down what it takes in order to deliver an engagement successfully to a point where we can start having fundamentally unqualified people in those roles.

Laura: Or people that you're trying to train. You want to be able to have an apprentice, or a Padawan learner, or whatever you want to call it.

Corey: Oh, absolutely. Every person we've hired into this role has gone through a training, and onboarding, and upskill approach because no one else does this quite this way. But there's foundational prerequisite knowledge that works. I mean, it’s—the idea of what it takes to operate in large-scale environments is a key example here. You can't teach that without giving someone a high-scale environment to work within. And it turns out, we don't have a lot of those right now.

Laura: Yeah, that's really true. There are some things you can try and there are some things you can't try, and some things are harder to learn than others, I think, too. And not just scale. Because scale, as long as you can get somewhere that has it, you'll pick it up. You'll have to.

There are some things that are much, much harder to learn. And one of those is, I think, really good troubleshooting. A great way to learn that is by shadowing someone working or working with someone who's really good at it, as long as they can talk through what they're doing. Someone who's really good at troubleshooting and can communicate. Another one is, sort of, responding to ops situations.

And I think—what I mean there is incidents, and outages, and firefighting. That's actually kind of an interesting thing. I remember I once toured a fire station in Atlanta. It was actually a heavy rescue station. If you don’t know the difference between firefighting and heavy rescue, firefighting is putting out fires, heavy rescue is running into burning buildings to pull people out.

And so I had asked this fire chief, “What makes someone really good at heavy rescue?” And he said, “Someone who thinks the day we get to run into a burning building is the best day of their life.” And some people in ops are like that. Some people love it. Like, it’s—

Corey: Oh, the adrenaline hit of the firefighting? Oh, yeah. I'm right there with you.

Laura: Yeah, I love that stuff, and some people find that incredibly stressful. That's one of the things I'm not even sure that you can learn it, by the way. I mean, I like to think you can learn anything if you set out to, but it might not be good for you to do it [laugh] honestly. If you find it incredibly stressful, maybe you should take something more—a little further from the fire. Dispatch or something, or R&D. But, you know, some things are hard to learn. And that's one of them.

Corey: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the Enterprise (not the starship). On-prem security doesn’t translate well to cloud or multi-cloud environments, and that’s not even counting IoT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IoT devices, detects these threats up to 35 percent faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at extrahop.com/trial.

Corey: What do you think is changing? Is it even still possible to have that shadow approach in a virtualized environment like we're all working within now? I mean, having someone shadow in the olden days was a tried and true method of getting someone up to speed on various engagements. It's changing now. And I don't want to ever be the kind of company that can't manage to hire junior people: “Oh, everyone here must be super senior.” Great, then where does the next generation come from if that's your approach?

Laura: Right, exactly. I think it is possible. One of the things that makes this interesting is—you know, I have an 11-year-old daughter, and she's doing remote schooling right now. And it is interesting to me how much better at learning without being in person, people who have grown up doing that than people who have grown up doing it in person. One of our friends make jokes about, “Haha, digital natives.”

And I'm like, “Yeah, that's it.” So, I see this kid sitting on a call explaining to an adult, walking them through how to set up that Discord server. Which I think is awesome, by the way. I'm super happy with that. But they don't see anything weird about it.

So, I think that is a thing that, the longer you spend doing it, the easier it becomes. It is hard to visualize how that works if you've done it in person your whole life, but the longer you spend doing it, the longer you start figuring out the tricks. To me, it is harder because you can't necessarily tell when somebody is struggling. So, there has to be a certain level of trust, and building that is probably the hardest thing. Comes back to trust, again.

Corey: Every company claims that they've nailed this. “Our employees love us. We have high trust among our staff.” And it holds still and it makes sense right up until the point where you talk to their staff directly.

Laura: Yeah. I think that's really true. It's really a hard change for people who haven't done it before and who wouldn't choose it. There's a lot of people who have had it thrust upon them this year.

Corey: So, a recurring theme of this show has been where does the next generation come from? And it's pretty clear that Fastly has done a phenomenal job of finding and recruiting extremely capable senior talent. But what are you doing to bring up the next generation? Because it's clear you don't just hire senior people, which is good. I'm just curious how you wind up developing those folks into the same level of amazing that some of your senior folks are.

Laura: So, we actually have some programs here I'm really excited about. To begin with, we hire people from all different backgrounds: we hire them from boot camps, we have people that are self-taught, we have people who have PhDs in computer science. But internally, we have a couple different programs that I am excited about. One is that we have a system where folks who work in our customer support teams can actually do a rotation through engineering where they can work in engineering one or two days a week for a few months. And if they like it, then they can transfer over and work full time in engineering as a junior engineer.

And that's been a really successful program, we've had a number of graduates from that, and some of the most awesome people on our teams. I’m really excited about some of those folks. The second thing, which is a new thing that we've just started trying, we have a kind of a weird team here called resilience engineering. That's something that I started that I'm super, super happy with. I’ll talk about that briefly, and then I'll talk about the apprenticeship program.

So, as you pointed out, we have a lot of senior folks. One of the things that's often the case with senior folks when you hire them because you—you know, they work on a standard, or they're famous for a thing, is that they tend to be specialists. People who know the most in the world about some technology X, whether it's some kind of network thing, or TLS, or whatever it is. And we didn't have a lot of folks looking at the whole system end to end: what are the weakest parts of the system? How can we make it better?

What can we do to make it more resilient? So, we started this team called resilience engineering. And it’s an interesting team. It has a bunch of really senior folks on it; some of them are pretty well known. But one of the things we thought about was that we weren't really training any new systems engineers that way.

So we've actually just rotated in our first junior engineer—well, they're not junior. They're an engineer, engineer—and I think this will be a good pathway for them to work their way up to being a principal engineer by looking at every system at Fastly, and understanding how things work, and understanding that the complexities that you get from having complex systems with huge scale. They don't necessarily behave in predictable, deterministic ways, it has emergent behavior. And there's that old story about the engineer who—knowing where to tap: this is all about knowing where to tap. So I'm pretty excited about that program.

Yeah, we're trying new things. We do not currently have an internship program, but that's something that we would like to do in the next couple of years. They’re all approaches we're getting to get junior people in.

Corey: Internships are always hard because, on the one hand, it apparently has to be about education. Two, if you're bringing interns in and not paying them, don't do that.

Laura: Oh no. Don’t do that.

Corey: That’s garbage. [laugh].

Laura: Don’t do that.

Corey: Yeah.

Laura: Yeah. Don’t.

Corey: Oh, that's not for you, that’s for a couple of companies who are feeling shame when they hear this, and they should because that's monstrous. But a lot of companies view intern programs as a backdoor recruiting funnel—which is fine, thrilled to do it. Until I watched them try to talk people into dropping out of school their final year and come into work full time instead, which feels a little weird. I've got to admit.

Laura: Yeah. No, absolutely not. Try really hard not to do that. So, we talked about starting one this year, but we didn't because of COVID. I was pretty involved at the internship program at Mozilla, so let me talk about that and some of the things we did there because I'm pretty proud of the work that we did there.

Corey: Wonderful.

Laura: So, obviously we had college interns, and we made some changes to that program because for a while we had gone after, sort of, Stanford people and MIT people, and you know exactly what I'm talking about, right? And the thing that we noticed was that there's a certain lack of diversity in folks like that, which is probably completely unsurprising to you. So, we started looking for interns from a different set of colleges and universities, and that was incredibly helpful. So, we did some deliberate recruiting to historically black colleges and universities, to colleges that had lots of professors that were interesting open-source and things like that. And those things were actually just a way to get some different types of folks.

We did another thing, too, which was a non-traditional internship program, and this was for people from any background. One of the people we had that came into it was a chef, previously. And the criteria for application were separate for each particular internship. So, for some of it might be: to apply for this internship, write us a piece of documentation; or, to apply for this internship, write a test for this. So, it was sort of a very open-source approach to things.

And some of the people we got through that program were incredibly brilliant folks from really unusual backgrounds. There was one particular person that I worked with who is one of the smartest engineers I've ever worked with. Generally, you hire somebody and you are constantly blown away by their insights and how hard they work.

Corey: I'm very fortunate to be able to say, “Yes.”

Laura: Yes, exactly. And this person had no background in computer anything. They had a background in mechanical engineering, believe it or not. They were also from a non-traditional background; I think the first person from their entire family to go to college. And they had actually been building museum exhibits for science museums, which is a really, really cool thing to do, by the way, but obviously nothing to do with programming.

The problem with those kinds of jobs is a lot of them are seasonal or contract. And they were looking around saying, “Well, I’d really like something a little bit longer term, and it seems like jobs in tech seem to have those criteria, so I will go and apply for this internship and along the way, hopefully, I will learn how to code,” and ended up being an incredibly brilliant engineer who built some really amazing things. So, yeah, I'm really open to bringing people in. My goal in working with people on the internet is for everyone to have the opportunity to build things they’re passionate about. And not everybody starts from the same point of opportunity, so finding different paths for people to get in is really important to me. I have kind of a non-traditional path at some point in my career, so I’m very empathetic to that.

Corey: I think that most people who have thrived in their career can look back at various points in their career and specific people that they reported to and say, “That person had a profound impact on my career.” Ideally, they'll even say that in a positive way, rather than negative, but I guess my goal has always tried to be one of those people that they say that about, someday. And it's easier said than done because the payoff is far in the future, you'll never know if you succeed, in some cases ever, and in other cases, not for another 15 years. But it's a good aspirational way to aim for, at least in my fumbling attempts at management. What tips would you have for folks who are aspiring to either become managers at all, or—in other words—become better managers than the ones that they had inflicted upon them?

Laura: So, I think the hardest part of any kind of leadership is managing yourself. Before you attempt to manage other people, you have to try and get a handle on yourself. And by that I mean let's imagine that you're at work, and you are really stressed. If I am an individual engineer, let's say I'm really stressed about something going on in my personal life: maybe one of my relatives has COVID, and I'm freaking out about that. Oh, I don't know how I'm going to manage childcare this year, there's so much going on, I can come up with a million examples.

And maybe my work suffers. And maybe my manager comes to talk to me about it. Now, if I'm a manager, or a director, or a VP, I come to work, and I have that leaking out all over the place, it is likely that that gets taken out on the people who work for me. Or at least, if they see that you are freaked out about something or if you're angry about something, people always jump to the worst conclusion. It’s the thing you mentioned earlier, “Can we have a quick chat about something?”

It's that if I see that my manager is super, super grumpy about something, and maybe writing a grumpy email, or the tone is not there, it has a profound negative trickle-down effect. It's not to say that you can't be human and can't have feelings, but you have to have an emotionally mature way of dealing with things. And that's really hard, by the way. Like that's non-trivial, obviously, but I encourage people, if you want to be a good manager, make sure that you have somebody to talk to about it. And that can be your boss; it can be trusted peers outside work; maybe you maybe hang out a social Slack. In a normal year, maybe you would go to a conference and have a cup of coffee with somebody. But I think it's really important to have ways to let off steam that do not involve your employees because it's just not fair to them.

Corey: Yeah. One of the hardest lessons for me to learn was that you can never complain to your directs, or in some cases, your peers. And that becomes a very difficult thing, especially when you're working in companies that aren't aligned in all the ways you wish they were, where there are things you aren't allowed to tell your staff, there are things that you think that your staff is right on, but you're not allowed to communicate in various directions because politics always strike you down. It was never a game I was particularly good at, and my approach was ultimately not to play, which. Really it's not winning; that's abdicating. What's the old line about office politics, where, you're not opting out, you're forfeiting?

Laura: Oh, ow. Ow. That's really painful, but yes, you're right. I think the second piece of advice is related to this, which is that I think you should try to be as transparent with people as you can. And when you can't be, it's okay to say, like, “I can't talk to you about this,” but being transparent by admitting that you can't talk about something.

Like, “There's something going on here; I can't talk about it now.” That's one set of things. And the other thing is to be really direct, if you can. I tell folks that I work with that my goal for communication is to be kind, direct, and prompt. So, be direct, say a thing that needs to be said, even if it's not really a positive thing, but be kind about it; there's no need to be a jerk about it, especially if it's negative feedback.

And to be prompt. So, if there's something that I need to tell you, like, “You just did a great job with this podcast, Corey,” I should tell you today. On the other hand, if it was, like, “Corey, you were a complete jerk. Why did you speak to me like that?” I should tell you today. I shouldn't wait. Three months later, until you've done the thing 20 times.

Corey: Oh, my God. The annual review or whatnot. It’s, “Well, eight months ago, you said something dumb, and we're going to ding you for it.” It’s, “What? What is this?”

Laura: Yeah, exactly. And that's all about managing your own psychology, too because it's really easy for people to want to avoid conflict. But learning to have constructive conflict is an incredibly important skill here, too.

Corey: Oh, it absolutely is. Thank you so much for taking the time to speak with me today about a wide variety of different topics all tied back to engineering leadership in some form or another. If people want to learn more about what you're up to, where can they find you?

Laura: Probably the best way to find me is on Twitter, which sounds terrible, doesn't it? But blogs, and websites, and things all have fallen by the wayside because I have had less and less time to think about anything longer than 140 characters. So, you can find me on Twitter @lxt.

Corey: And I presume you're hiring as well?

Laura: So we've had a very busy year, and as a result, we're hiring a lot of folks. So, please reach out if you're interested in working here.

Corey: Excellent. And we'll of course put links to that in the [show notes 00:35:40]. Thank you so much for taking the time to speak with me today. I appreciate it.

Laura: Of course. Anytime. It was lovely talking to you, as well.

Corey: Laura Thompson, VP of platform engineering at Fastly. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you've hated this podcast, please leave a five-star review on your podcast platform of choice, along with a comment saying why a CDN isn't necessary and even if it were, you could build your own in the course of a weekend.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Matt

Matt Cauthorn oversees the ExtraHop Security Sales Engineering, and enjoys studying the intersection of business and technology. Prior to ExtraHop, Matt was a Sales Engineering Manager at F5. He’s a passionate technologist and evangelist. He holds an MBA from Georgia State University and a Bachelor of Science degree from the University of Florida. Matt speaks at industry events, has been featured on podcasts, and quoted in industry coverage.

Links:

  • ExtraHop cloud solutions
  • WEBINAR with ExtraHop and Corey: "Secure Your Cloud Against Advanced Attacks with Network Detection and Response"

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the Enterprise (not the starship). On-prem security doesn’t translate well to cloud or multi-cloud environments, and that’s not even counting IoT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IoT devices, detects these threats up to 35 percent faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at extrahop.com/trial.

Corey: If your mean time to WTF for a security alert is more than a minute, it's time to look at Lacework. Lacework will help you get your security act together for everything from compliance service configurations to container app relationships, all without the need for PhDs in AWS to write the rules. If you're building a secure business on AWS with compliance requirements, you don't really have time to choose between antivirus or firewall companies to help you secure your stack. That's why Lacework is built from the ground up for the Cloud: low effort, high visibility and detection. To learn more, visit lacework.com.

Corey: Welcome to Screaming in the Cloud I’m Corey Quinn. One of the problems with being me is that it gets kind of lonely because I stand sort of squarely between the worlds of business and technology. You’d think they might be the same world; they’re kind of not. And one way that I tend to make that isolation a little bit more bearable, is to talk to other people who are in similar positions. This episode is promoted by ExtraHop which is a network security vendor that we’re going to dive into because my guest today is Matt Cauthorn, who’s the VP of Security and Cloud at ExtraHop. Matt, thank you for joining me.

Matt: Yeah, thanks for having me, Corey. Good to be here.

Corey: So, ExtraHop was one of those companies that I became aware of as something to pay attention to. And it’s going to sound weird and obnoxious that I don’t even care, but the reason that I started paying attention was because there was an event in the before times here in San Francisco, and I started seeing your name on the side of city buses. The company, not yours personally; when you see a person’s name on a bus, that usually is a different implication.

Matt: Yeah, I have a feeling it was one of several events that we were involved in. Yeah, it’s great. It’s great that you discovered it that way.

Corey: Say what you will about advertising like that: it works. And the problem you run into, in some cases, is that you aren’t able to really convey the depth and intricacy of what a company does. Now, you folks have been a sponsor for a while of my nonsense. And thank you for that; that shows that someone is making excellent decisions on your side. They should be promoted and make more decisions just like that one.

But for those who haven’t been paying attention to the world of security, and all the various nonsense that I do, what is ExtraHop? What do you folks do over there, other than buy advertising on buses?

Matt: So, the technical category that we fall into is network detection and response, which effectively means sophisticated network security analytics for the enterprise in the cloud. And if there’s a network where we can see the packets and process them, we are able to give very, very sophisticated security analytics on that, as well as support for the incident response workflows, and APIs, and much more.

Corey: I’m going to put the shoe on the other foot for a minute here. Whatever I start doing significant sponsorship work with a company, I like to vet them and figure out that okay, is this something that isn’t, you know, complete crap because I’m not necessarily endorsing every sponsor that comes through, but at some point, when you wind up having a sponsor who is next to things that you’re doing, over a long enough timeline, you start to become associated with them.

And the problem with security vendors in many respects is they almost invariably start speaking to security folks who are steeped in that world where a CISSP almost feels like it’s a prerequisite to understand what’s going on. That’s one of the reasons we launched the Meanwhile in Security companion podcast, to specifically cut through that mess. But for me, the way I learned is by rolling something out and using it myself. And I did deploy ExtraHop to my test environment. And I was pleasantly surprised by what you folks have built.

Matt: Oh, thank you. I’m glad to hear it. And you’re exactly right. So—and I assume your test environment or your environment was up in one of the cloud providers, is that correct?

Corey: Yes. AWS because that’s the one that I have the most experience with; our day job is helping companies fix the horrifying AWS bill, and we sometimes discover how that breaks by incurring one ourselves from time-to-time because it’s a problem that stalks all of us throughout the course of life.

Matt: Yeah. I’ve been on the dubious receiving end of that billing as well in another life. So, you’re doing a service; thank you for that. So yeah, one of the big things that not everyone is familiar with, and I think many, many more if not everyone who’s delivering apps and services in the cloud, is that now, in particular, AWS and the other cloud service providers can send network traffic to target interfaces.

And that means vendors like us can process that invoke behavioral analysis on these byte streams and give you transaction analysis and security, forensics investigation, and detection. It’s a very, very powerful—and it’s purely out-of-band, native to cloud. So, the way you’ve deployed is using the native facilities up there, and it works really, really well and it’s a wonderful adjunct to a nascent security strategy or a very mature security practice.

Corey: The way that I wound up contextualizing it is… I started off as a grumpy Unix systems administrator, hands-on hardware in the worst ways possible, and I spent a fair bit of time dabbling as a network engineer as a part of that. In fact, during the financial crisis, back in 2008, I was stuck in a job because no one was hiring. I’d been there a year; there was no real advancement opportunity; there was a salary freeze, so as I hit my one year, I wasn’t able to get a raise. And that led me to be even more disgruntled than I normally am. So, my approach to becoming a better systems administrator was to get a CCNA during that timeframe.

And it sounds counterintuitive, but the more I understood what was going on in the network, the more the rest of the system made sense to me, to the point where now when I start trying to diagnose weird issues, I start from a network-based perspective. The problem is so much of that in a cloud environment is obscured away and not easily discoverable.

Matt: Yeah, the beauty and danger of the cloud, simultaneously, are the layers of abstraction that you just described. That’s is exactly right. On the winning end of it, you get this radical acceleration of traditional infrastructure, deployment, workflows, deploy, destroy, all of this stuff, but the price that you pay is that these levels of abstraction take you, sort of, further and further away from having your finger on the pulse of the environment.

And the ultimate—I’ll just wear out the metaphor here, Corey, but the ultimate connective tissue is the network itself. And in fact, that’s where the preponderance, at least, of the actual behavioral intelligence lies, it’s on that connective tissue. And so without having real awareness of what’s happening on the network itself from a behavioral analysis perspective, you really are kind of flying blind.

Corey: What I want to talk about, too, is that to just give folks an example of what’s happened. In fact, while I have you on the recording, I just pulled up a view into what’s going on in my environment, and it tells me all kinds of interesting views. And honestly, this is one of those visualizations that I wish more companies would discover because, let’s be very clear here, what you’ve built is actually beautiful and a pleasure to use. It almost feels like it’s conference-ware where it’s designed to look good in demos, rather than actually be usable, except that having played with it a bit, it is in fact usable. And it distills down to the EC2 instances that are in the environment, it tells me what’s talking to what, on what port, any sudden spikes, any anomalies, and then it highlights a bunch of different rules here.

And I’m seeing all this from a purely network perspective. Now, that’s great. You can talk to folks about all kinds of tools that do this stuff. All right, so effectively, you’re implementing Wireshark as a service. Okay, that is certainly a way to think about it, except it’s being captured by a VPC mirroring; there was no configuration required on the instance itself; it’s something that can be done account-wide.

It’s something that can be enforced via SCPs within AWS organizations; it’s something that is not, no matter how thoroughly I subvert the EC2 instance that this thing is running on, even if I subvert the entire AWS account itself, as long as I haven’t been able to lateral into the management account for the AWS organization itself, you can’t turn this off and it shows up the truth that lives on the wire.

Matt: Yeah. I love the way you said it. And so I’ll add to the Wireshark metaphor here in a moment, but you’re exactly right, Corey. One of the strengths—and I would encourage like all the listeners—and you’ve got a very broad listener base here, so there’s a veritable mix of different skill sets and folks at different parts of the organization, this is all fine. But I would encourage everyone listening to think about the role of network visibility as it relates to your application and service delivery. The network has a couple of unique—several unique properties. One of them is what you just described: it’s very, very difficult to evade; and it’s very difficult to turn off, and it’s very difficult to manipulate.

Corey: And if the network isn’t working, effectively no cloud service is either. “Oh, it’s doing an awful lot of calculation. Good for it. If I can’t talk to it, what’s the point?”

Matt: Exactly right. So, what we’re doing here with the modern era of analytics, and the state-of-the-art changing so rapidly in the last 10 years or so for network analytics, think of millions of concurrent Wireshark sessions happening with the subsequent expert analysis and behavioral intelligence, with behavioral security detections layered on top. And then if you need to investigate one of those detections that you’re seeing right now, Corey, you click through, you see the asset involved, you see the transactions themselves, that surface to the conclusion that the system came to. And so it’s a very, very powerful thing for just the detection and investigative workflows. But there are far broader use cases as well.

Corey: The real value as well—I want to be very clear to help paint the picture here—you have a web server, or an application server, or database server, if you’re still running those yourself—given some of the database services that are offered, I can’t say I fault you for that particular choice, but I digress—if suddenly those things start talking externally to random botnet command-and-control servers, for example, that’s atypical behavior. And it’s the kind of thing that you sort of would like to know, approximately, immediately, it’s the sort of thing that emerges of, “This is an emergent aberrant behavior and it should be investigated.” Now, the other side of that is, I set this up back at the beginning of the year—thank you for the account, it’s appreciated—and I wound up getting it dialed in on my environment, and I haven’t logged into it in a few months. So, now I’ve logged back into it for this discussion, there are zero alerts waiting for me.

And that’s no small thing because what I do on this development EC2 instance in this account is monstrous. There’s no way around it. I install random stuff from Docker Hub, occasionally, due to poor life choices, effectively the entire software security supply chain—oh, [laugh] that’s a funny joke. I don’t know anyone who—involved in any aspect of it runs in my stack. I may as well just open it to the world.

I have my IRC connection living persistently on this box through Irssi. It does a whole bunch of things and talks to other stuff because that’s the way the world works. It’s messy. When I set this up, it flagged those things immediately and I said, “Okay, don’t alarm on the fact that it’s connecting to Freenode with IRC.” Great. It hasn’t bothered me since as I continue to do monstrous things. There were no alerts waiting for me because the problem of not getting any alerts when things are going wrong is super bad, but getting alerts constantly when things are normal, is in many ways worse because when something happens, it gets masked.

Matt: A hundred percent. Yeah, so what you experienced is the power of the state-of-the-art of network analysis. And behind your instance is machine learning that runs in the cloud at scale. And what that means is, is that the system that you’re running in your environment, right now, Corey, is able to extract observed transactional features that feed the machine learning. And so initially, the IRC, we’re like, “Wow, we don’t normally see this, dude.”

And you’re like, “No, don’t worry about it, ExtraHops.” So, what we learned is, that is normal behavior in your environment. And there’s just a plethora of different use cases and different machine learning models and implementations. That stuff doesn’t really matter for the purposes of this conversation. Suffice it to say, when you think about the network, just if you’re looking at it through the pure lens of as a data source itself, well, what kind of data, what sort of information could I mined from that data source? Then the answer is it’s staggering.

So, then the question becomes, how do I present it—which you’ve mentioned earlier—with our UI? There’s been a ton of R&D, that we’ve got this wonderful R&D team. And the UX team has done a great job at distilling the information down that we surface because we’re just analyzing just insane amounts of raw network data in a given environment and every single day. So then, when you overlay machine learning, it really helps to sort of—you know, there are certain things that machines are really, really good at doing, and extracting features and analyzing those features for real behavioral analysis is one of them.

Corey: I also want to point out as well—because again, I approach the entire world through a lens of AWS billing, and there’s an awful lot of solutions out there that give horrifying impact to the AWS bill by deploying them, to the point where you start doing a cost-benefit analysis and realize, “Huh. I’m reasonably certain an actual data breach would be less expensive.” And you wouldn’t be far from wrong. I just pulled up last month’s bill in the account this is running in, and sure enough, the traffic mirroring, that is what powers your solution is a third of my bill. But I want to say that that third of the bill is $10.08.

And that does not have traffic volumes attached to it; it is strictly a per hour—one and a half cents per hour—that it’s attached. The end. And I’ve got a level with you, if $10 is meaningful to monitor what’s going on on the network in an account, I don’t know what to tell you, other than perhaps you are not the target customer. And I want to get into that a bit with you because I’ve long held the opinion that there are different on-roads for different companies at different times throughout their growth to start working with vendors. Who should be reaching out to you folks, and more importantly, at what stage of the development process does starting to engage a solution that looks at the network traffic and cares about network visibility makes sense in the modern era?

Matt: Very high-level guidances is this, is that if you have any Infrastructure as a Service running in your environment of consequence with risk associated critical assets, with critical services. Generally speaking, Corey, it’s worth reaching out to us about—whether it’s cloud, or enterprise, or hybrid combinations therein, if there’s a network to monitor, we will do that. And we don’t discriminate in that way. So, it’s very, very useful also, for the enterprise cloud journey folks out there, and there’s a lot of them [laugh] at various different stages at this. If it’s early stage, there’s the sort of assessment, the security controls that need to be sort of moved up into cloud.

And a lot of the executives that I talked to, I’ve got—I’m fortunate, I get to talk to CEOs and VPs about this exact scope of concerns, and many of them, their feet really aren’t firmly under them when it comes to cloud. They’ve got their enterprise environment locked in, and they’ve got their security controls well defined, but DevOps is moving and the agility that they’re gaining from the cloud, it’s moving so so fast that the CSOs are kind of caught flat-footed and they’re not exactly sure what this thing should look like in the cloud. And so, for the enterprise folks on the journey into cloud—digital transformation, whatever buzzword you want to throw at it—that’s another wonderful target account for us.

Corey: An observation slash analogy I’ve been making for a little while has been that, imagine tomorrow I go and I file the paperwork to start Twitter for Pets. I already own the dot com, but now it’s a real business. And in the next 10 years, it’s going to become an S&P 500 component where, great, it has gone from ridiculous social network for pets to consequential social network for pets. And as it grows from ridiculous startup to large enterprise, there has to be a reasonable onramp for folks, given the sensibilities of how companies work today.

It can’t be an enterprise transformation story because anything I start tomorrow is going to be born in the cloud anyway. And it’s no guarantee or honestly, not even that likely for a lot of these use cases, there will ever be a physical data center component. There has to be a point during that company’s growth where there’s a natural on-ramp to use a vendor’s product or service because if there isn’t one, they are fundamentally serving what is, in the very long term, a market that is in decline. And that’s always the sort of thing I look for and am cautious about. Oh, we wouldn’t be having this conversation if I thought you didn’t have an option for folks who are in precisely that position. How do you think about that?

Matt: Well, no, it's a really interesting point, you’ve got a very unique voice in the space. Before I continue, I really like the particular angle you’re approaching these problems from because these are conversations that have to take place. So, the operational concern itself bears a certain cost, and a certain level of risk, and a certain level of opportunity cost. And you’re exactly right, at some point in the story arc of a cloud—or business’s experience as they grow into this, there’s a point of diminishing returns with native tooling or hand-rolled tooling. And beyond a certain point of scale, you need to actually fall back on more broad-based utility, broader coverage of the security requirements, the coverage of your security policy and your controls, and just better alignment. And in many, many cases that will be vendor-led. And that’s okay. But you’re exactly right, there is a point beyond which you’re really going to want to engage with experts in that particular domain because it’s not cost-effective to do so yourself.

Corey: One of the most blatantly wrong things that I hear from the world of cloud marketing comes from AWS itself, which is, “There’s no compression algorithm for experience.” There absolutely is. You don’t have to build all of this stuff yourself from scratch. You can compress that experience into hiring experts who are good at that sort of thing, either as employees or consultants. That’s why advisory consultancy is a thing.

You can buy products and services that compress all of that hard-won, hard-fought experience into something that you can buy off the shelf and it solves the problem far more effectively than you’re ever going to be able to build in-house. And that’s a valuable and powerful thing. The hard part, of course, is in the security space, you can effectively spend infinite money on security, and even then there are no guarantees. So, it’s challenging as companies grow—especially in the early days—to make security a priority because it’s always something we’ll focus on later until suddenly, you really should have been paying attention, and now it’s too late.

Matt: Yeah, this is a big one. And I understand how that comes to pass, Corey, as do you and everyone who’s listening. Like, it’s very easy to rationalize yourself into that place, and it’s very understandable. And in fact, I myself have done it in my past in—as my prior life in operations. And there is a certain point beyond which the risk calculus alone and the impact of that, it just reverses the polarity of that whole discussion.

And then the worst case is something bad happens to you when you’ve been in limbo before you’ve implemented your security. Unfortunately, we’ve seen this happen with several organizations where they’ve decided to just freeze budgets on security, whatever, and then bang, there’s a compromise and they end up on the news. I’ve seen this several different times in the last year alone, as a matter of fact. And so this isn’t fear-mongering, and I want to—Corey, part of your brand is calling out things as you see them, and so I think that one of the unfortunate things about the security industry at large is there’s lots and lots of fear-mongering. And I’m not doing that.

Instead, I’m saying understand your risk and understand that calculus and your appetite for impact. Let that be your north star as to when to really get serious about your security controls. And that might be from inception, by the way. And that’s a great answer. To an earlier point, it might be a risk that you’re willing to make up until some sort of financial threshold, beyond which you’re not willing to appetite—it’s a unappetizing risk beyond that.

Corey: Forget dozens of visualization tools and view your entire system in one place with New Relic Explorer, the latest addition to New Relic One. See your system-wide health at a glance with a dense hex view that has your hosts, services, containers, and everything else. And get an estate-wide view of sudden changes, so you can catch issues before they impact customers. So go to https://newrelic.com, sign up for free, and start exploring your system today.

Corey: It really comes down to risk management. I mean, one of the reasons that I focus on the AWS bill is that that is almost ever a company-ending event, it’s, “Oh, I spent too much money,” is the cost of not focusing on it sooner. And that’s almost always both okay and survivable. In the absolute worst case of, “Wow, we normally have $1,000 a month bill and we just got charged $800,000,” AWS is a company that understands the longer-term view, you can reach out to them and get it fixed in almost every case. Security does not work that way.

And it’s much less tangible, as far as being able to sell something effectively into that market. In fact, one of the problems I have is walking around the RSA expo hall—whenever I was able to do that in the before times; last conference I went to before this whole thing started—and you see what feels—past a certain point—the same product being offered again, and again, and again, with different logos and different company names, but the messaging is the same, and it’s incomprehensible, and it just looks like there is no winning here. I found that ExtraHop was a breath of fresh air comparatively. But I’m not going to lead you that far down the road. Tell me what separates you folks out from the industry at large—not specific vendors because no one’s going to look great smacking in the competition, but there’s something refreshing about your approach and how you talk about your approach. Where did that come from?

Matt: It comes from our pedigree of being network-deployed, but application-fluent. So, here’s a fun fact. So, our co-founders, years ago, invented the modern-day application delivery controller, specifically at F5 networks. And this was a long time ago. And in so doing, that device is a very, very, it’s a network-deployed device that’s deeply application-fluent, and all of that domain experience and all of that sensibility towards scale, the ability to see inside decrypted packet streams and do analysis, all of that made its way into our product and then fed the beast of network analytics.

And our worldview really is steeped in this idea of just network analysis and the various outcomes that you can glean from said analysis, like behavioral detections for security, like asset inventory, your security controls, this the visibility that you cited earlier, Corey. It’s like many environments, they don’t know what’s running. And the network will tell you what’s running in a way that’s deeper than just, like, the management console listing the assets and services you’ve got. And so now, down to even the transactions, what types of services? What’s the consumption model of this?

Who’s consuming it? Where’s the traffic going? And is this normal? Yes or no? So, that’s really what makes us different. Most of the folks in our space focus solely on detections, and we believe that the network as a data source can give you much, much more value. And so we strive to deliver that.

Corey: There’s an awful lot of value in being able to deliver value upfront, and getting customers who have worked with you before to say, “Yes, this thing is amazing.” And I have problems with that in the space that I’m in because it turns out that there is a perception—that I disagree with—that fixing bills or talking to someone about a cloud bill that was high is somehow a ding on the company. And it’s not even about being high; it’s about having a lack of visibility or understanding in many cases, but people don’t want to talk about it. It’s hard enough to get testimonials and logo rights in that context. In a security space, it feels like we are thrilled to wind up buying your product now that we see the value of it. If you ever mention our name in any context ever again, we’re going to drive a wrecking ball through your corporate headquarters, legally speaking. How do you get past that?

Matt: It’s understandable, first of all, and you’re right, Corey. In large part, folks are not super eager to talk about security in a very public way. And that’s okay. I wish that there was more, though, not as a vendor representative where we would be the beneficiaries of it, but just more sharing in general really, really needs to happen. And what we’re seeing instead is the big disclosure and the big tech ta—like last year with SUNBURST.

It’s a monster and it’s catastrophic affliction leveled on the industry, and there was a single point of disclosure, which was wonderful, and then the sharing started. And I feel like there’s a lot more opportunity for information sharing, even with the current frameworks that are out there; there are vehicles to do this in a formal way for a given industry. But we need more. And you’re exactly right. It’s discussing the state-of-the-art and threats, and God forbid, attempts at compromise or full-fledged compromises, there needs to be more of that so we can collectively level up.

Corey: I’ll even name names on this because I’m not a security vendor. The Capital One breach a few years back was fascinating for me because it wasn’t just that they had done things badly or irresponsibly, didn’t read the instructions on the tin, it was a series of chained together exploits. There was a exploit in the web application firewall, I believe—according to court filings—that allowed someone to get a foothold. From there, there was an overbroad instance role that allowed them to get access to an S3 bucket that they should not have had access to from that account. It was tying together different things in different ways.

And that, in turn, is the sort of attack that is not easy to see coming, and there’s a lot of things you can learn from that; I’m sympathetic to it. The problem, of course, is that first, they’re are a bank and the lawsuits and the rest means that Capital One at that point, whenever the word ‘cloud’ comes up, felt like for a while they just put their heads down, and there was six more weeks of no talking about cloud whatsoever because they didn’t want to talk about it at all. But that’s the sort of thing where we can all learn so much from what happened. But the instinct is to button up and never say a word about it. Which means that the only people who are able to really go in-depth on this is, in fact, security vendors with the counter-argument that as soon as you start talking about that in your marketing, you get accused of effectively ambulance chasing or that you’re using fear, uncertainty, and doubt to wind up selling your products. And yeah, a lot of vendors do exactly that and it’s awful. But there are valuable learnings here, and it’s not just a sales opportunity for a product but rather an opportunity to uplift the entire ecosystem.

Matt: Yeah. And to the extent that the security market, in general, is a very vendor-wary market as an audience, and I understand why. I was on the receiving end of vendors as well, back in my prior life, as I mentioned. And I understand that, and to that, I would say, is make us prove it. If there’s a decision to be made and you’ve deemed it necessary to engage with us then, as a good security buyer, make us prove it.

And there’s many, many—especially in the cloud—there’s many vehicles at your disposal to test the claims of any given vendor with any given approach, whether it’s a SIM with log analysis, or endpoint, or network, or beyond. So, make us prove it, and then you’ll get a line of sight to whatever claims are being made around catching breaches, or understanding behaviors, or beyond.

Corey: So, with all that in mind, and obviously the way that things used to be and how all of this stuff would tie it together, it feels like the old answers aren’t right for the new era. So, from that perspective in a more forward-looking sense, what does strategic security tooling look like in this cloud era that we all find ourselves, willingly or not, enmeshed within?

Matt: Okay. That’s a super important—in fact, that’s probably like—you’ve asked a bunch of good questions; this one’s at the top of the list as far as I’m concerned. So—

Corey: When you don’t know a lot, you get very good at asking good questions because that’s how you fix that problem.

Matt: [laugh]. Hey, man, I ask a lot of questions myself, so you’re in good company. So, one of the problems in the traditional terrestrial enterprise is that their tooling strategy looks like a shotgun blast. And that shotgun blast is comprised of point solutions that are loosely federated at best, at best. And the only point of integration is the swivel chair that an analyst would sit in, or the Site Reliability Engineer, or DevOps person.

Corey: Don’t forget the screens upon screens upon screens that show amazing things when someone walks by, but if you think about this for more than half a second, you realize people are going to wind up with repetitive strain injuries from trying to pivot to look at all those things on the screen, and wow, maybe that much thing to look at all at the same time, but be incredibly stressful that unpleasant when you’re getting a suntan from the monitors. That’s a problem.

Matt: No, that’s exactly right. The big board of the past, in the terrestrial Data Center—the Security Operation Center or the Ops IT center, whatever, the ‘fishbowl’ we used to call it back in my old place—that really does point to the legacy era. Now, if you hoist that exact same model up into the cloud, or especially in hybrid environments because most—or many. I don’t know about ‘most,’ but many are in this sort of transitionary state. They’re multi-cloud, A, or they’re at some stage of cloud adoption with traditional enterprise workloads.

Well, now what does tooling look like because we have a management plane that can do really, really intelligent stuff, and the APIs are very, very consistent, they’re very actionable, and they happen pretty quickly. Not as quickly as I would like sometimes, but these events are easy to trap, and they’re easy to act on. And so the modern era of security tooling is comprised of, think about your data along the boundaries of its data source. So, for example, I care about my containers and so I want some sort of runtime container visibility. Or if I’m running EC2 instances, I want endpoint visibility because I want to know what’s running and resident in memory, or if it’s whatever; malware or whatever.

Then I want—I’m going to log because you log a lot in the cloud, it turns out, and so I’m going to need some way to make sense of those logs and wrap that into part of my practice. And then lastly, I want to have visibility into the network because of the three things that I just described, endpoint, say, or agent-based approaches, log-based approaches, those things can be evaded, they can be disabled, they can be turned off—and in fact we saw evidence of that, very active evidence, last year with SUNBURST—and the network is the only one that’s truly covert and difficult to evade, manipulate, or disable. And so as part of this collective strategy, now you’ve got—and we’re very complementary to one another: logs are complementary to us; where we leave off, as well as endpoint, and vice versa. And so we call this the ‘Cyber Triad.’ And this is not just our terminology; it’s analysts and others that are out there.

Corey: Always good when you hear the buzzwords, and they didn’t come directly from the vendor.

Matt: In this case, it’s not a buzzword; it’s actually a genuine strategy because we tended—in the past, we haven’t thought about our security tooling from a strategic, sort of, data source perspective. And in the context of cloud, especially, you can wield these data sources in some really, really powerful ways and do, in this sort of DevOps or SRE sense, you can do this event-driven security model. Now, the tooling itself can emit events into the management plane of the cloud, and the cloud, in turn, can take intelligent action. It’s a beautiful and devastatingly powerful new era for real-time security response. So, now in the past, Corey, I would quarantine a process on a system, or maybe if something was really, really bad in a terrestrial, I would just, like, disable that, block it.

Maybe I would do virtual patching on the firewall where I would disable a given service on the firewall. Well, now in the cloud era—and your audience understands this super well, I just call the management plane and redeploy the container. Done. Golden image; it’s fresh, it’s clean, it’s got attribution and I know that if that other one was compromised, I’m just going to get rid of it, because cloud, and redeploy this thing right in its place. It’s beautiful.

And so in the modern era, the cloud itself unlocks a set of operational models for security that are really difficult to achieve otherwise. It’s not impossible; there’s a whole industry dedicated to it, but in the cloud era, it’s much, much, much easier, and it’s easier to wrangle, and you can hoist it higher up into the dev lifecycle, the CI/CD lifecycle itself. So, it’s a really nice time for security ops.

Corey: It really seems to be. Matt, thank you for taking the time to go through, sometimes, the befuddling world of InfoSec, especially from a vendor perspective. If people want to learn more about you, what you’re doing, what you’re up to, where can they find you?

Matt: Well, they can find us at extrahop.com. And there we’ve got cloud case studies, use cases. In fact, we’ve even got an eval that’s out there. We’ve got a live—it’s running in the cloud, actually, a live demo where you can sign up and experience the system running in the cloud, before your very eyes and see the type of visibility gains you can get, and network analysis manifest, really. It’s a real live system up there. So, I would strongly recommend that if anyone’s interested, to have a look at that because it’s quite a powerful model, in my opinion.

Corey: And if folks have questions, do feel free to direct them my way because, remember, the one thing that is never for sale here is my authenticity, for better or worse, which often gets me into serious trouble. Matt, thanks for taking the time to chat with us. I really appreciate it.

Matt: Yeah, likewise, it’s been a pleasure, Corey. Thanks so much.

Corey: Matt Cauthorn, VP of Security and Cloud at ExtraHop. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice along with an insulting comment that you will later be able to disavow because no one was tracking what was happening on the network, so it must just be an application bug.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Katrina

Katrina Bakas is a Senior Technical Product Manager at Amazon, working on CloudFront within AWS. She brings a lifetime of relentless curiosity to her role and a desire to simplify complex technologies to make them accessible for more folks. Previously, she brought the same inquisitiveness and desire to simplify to observability at Pivotal and VMware (upon acquisition), having spent time at start ups and in megacorporate Financial Services before that. She strongly believes the best tech is found at the intersection of psychological safety, Design, Engineering, and Product Management.

Links:

  • “What is a Cloud Platform? What is a Platform as a Service?”: https://medium.com/@kvbakas/what-is-a-cloud-platform-what-is-a-platform-as-a-service-eb2c33cfa38e
  • Twitter: https://twitter.com/katrinabakas

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Fairwinds. Whether you’re new to Kubernetes or have some experience under your belt, and then definitely don’t want to deal with Kubernetes, there are some things you should simply never, ever do in Kubernetes. I would say, “run it at all.” They would argue with me, and that’s okay because we’re going to argue about that. Kendall Miller, president of Fairwinds, was one of the first hires at the company and has spent the last six years the dream of disrupting infrastructure a reality while keeping his finger on the pulse of changing demands in the market, and valuable partnership opportunities. He joins senior site reliability engineer Stevie Caldwell, who supports a growing platform of microservices running on Kubernetes in AWS. I’m joining them as we all discuss what Dev and Ops teams should not do in Kubernetes if they want to get the most out of the leading container orchestrator by volume and complexity. We’re going to speak anecdotally of some Kubernetes failures and how to avoid them, and they’re going to verbally punch me in the face. Sign up now at fairwinds.com/never. That’s fairwinds.com/never.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Katrina Bakas, who is a senior-product-manager-dash-technical-dash-that's-a-corporate-title-if-ever-I-heard-one at Amazon, specifically AWS’s CloudFront group. Katrina, welcome to the show.

Katrina: Thanks for having me,

Corey. Glad to be here.

Corey: So, first, I understand that you're soon to be transitioning to a different team within Amazon, but fortunately, we're able to do this recording before that happens. What is a product manager dash technical? And of course, we will not forget the word ‘senior’ that goes in front of that either.

Katrina: Yeah, a product manager dash technical at Amazon is pretty much the same thing as a technical product manager anywhere else, which is to say, no one really knows what a product manager is and it varies wildly depending on your role, sometimes even within the same company.

Corey: Well, I have a friend who went down that path for a while, and one of the things that he kept encountering was at some companies that I won't name because you can tell a Google story from a mile away. It was great. “Oh, you're a TPM. Wonderful. Please solve this algorithm on the whiteboard first.”

And it's a little ridiculous if you look at the role there where it's, you have to be able to prove that you can code so you can get a job wherein you'll never have to code again. But okay, we'll roll with that. Other companies that he spoke to, were able to have the conversation without ever getting into anything that could even remotely be considered engineering stuff. So, between those two endpoints across a broad spectrum, where does that land for AWS, first, and secondly, for you, personally?

Katrina: Amazon and within AWS, I get the sense that it varies a bit per team. You can sort of choose your own adventure at Amazon with as involved as you want to be. It really is the kind of place where you can find a team that really needs your skillset, and you can dive in really deeply there. I personally do not have a software engineering background, so I stay away from needing to make those implementation decisions. On my team in particular, I'm very much responsible for representing the customer in the room, and balancing what engineering tells me is possible, and know enough technical knowledge to understand if I'm getting sandbagged or not, but not prescribing what the implementation details should be.

Corey: I always try and steer away from saying, “Oh, well, do you write code, or aren't you technical?” I think that's one of the biggest tropes out there that is provably false. Anyone who thinks that there's not technical detail and minutiae in that whole universe around any aspect of business is, frankly, naive. There is so much depth and insight that is valued across the entire board, that, oh, you don't know how to write a particular programming language, therefore, you're not technical is just a terrible take.

Katrina: I totally agree with you. I think one of the things that's even more frustrating in this conversation is folks who say, “Well, great. Now more people can be technical because Jamstack exists, so you can make an app.” I would go one step further to say that to be technical, you don't have to be able to develop an app, whether it's code, no-code, or anything else.

It's about understanding what the thing is that you're working on and knowing the trade-offs of any particular feature, or chore, or technical debt that you're introducing. To me, it's just about understanding systems at least at a level where you're able to make trade-off decisions.

Corey: Yeah, the idea of Jamstack is just fascinating to me and something I plan on getting into a bit down the road if this year lets me. I love the concept; I love how approachable it's becoming. My static website template generator thing that I launched with a few years ago has been replaced by things that actually, you know, work. And it's nice to see that unfold. So, getting, I guess, away from the high level and into a bit more of the weeds of what you do, CloudFront is odd to me in that it feels on some level like it was, give or take, feature-complete—I know, I know.

Don't email me—a while back because it's a CDN that has various points of presence across the world. Above my desk, I have a map of the AWS global infrastructure because I am really sad. And every time I wind up hearing about a new release, I put a new pin in there. And I look at this thing, and it's a sea of CloudFront locations around most of the world at this point. So, turning up new edge locations is valuable, obviously, and necessary. But I don't see a rapid pace of innovation around the product capabilities. And to be very clear, I'm not necessarily sure I want to either. Am I wrong? What am I missing?

Katrina: I think you hit the nail on the head here. A CDN exists as, ideally, something in as much of the background as it can. A CDN will accelerate your website and protect it from DDoS attacks from around the world or various other types of attacks. People come to CloudFront for security, and for reliability, and for acceleration. We do that remarkably well.

Anything on top of that can be nice to have, but we don't focus so much on what are the additional bells and whistles that we can tack onto the core product that does things, as I mentioned, that we do really well today.

Corey: So, I feel like I should also clarify slightly here where, if I look over the past year, there was no big announcement about it, but CloudFront distribution updates used to take—how do I put this directly—an age. And it has consistently dropped to sub-five minutes, which is awesome as a customer. But there was no big announcement about that. It was a quality of life improvement that I thought was incredibly appreciated from my side, but there are also aspects for certain use cases where that feels like an eternity as well. There are, of course, enhancements and stories about making it better to work with across the board.

But it's also something that is relatively arcane. If you talk to the typical person just getting in technology, a CDN is not necessarily intuitive, and it has a bunch of sharp edges, if you're not careful, that can cut folks to ribbons that has nothing to do with CloudFront specifically, and everything to do with the fact that networking on a global stage is super hard.

Katrina: Networking on a global stage is super hard, and as someone that works in it every single day, I still find nuances just about every week that I didn't realize that I needed to know. I mean, even just taking a big step back and thinking about how the internet works at large and then replicating that in a private system so that your data can be more protected, and managing that without necessarily all of the same… let's say, benefits and headaches of liaising with all sorts of government entities as the public internet does. It's quite, quite complex. And the thing that I most look for as a PM for CloudFront is how can I pass the least amount of burden of managing all of that to my users while still delivering the value of that acceleration and protection at the same time.

Corey: It's also a bit of a challenge because I would argue that people's first experience as they learn how all of the various AWS services work with CloudFront isn't particularly friendly. Because if I'm looking at this from the Jamstack perspective—and one of my first outings with CloudFront was exactly this way: “Oh, I'm going to build a static website, and I'm going to have it live in an S3 bucket. Awesome. That's great. Now, I want it to, A) have its own domain name, which is not unreasonable, and I want it to work over SSL/TLS because this is no longer the 1990s.”

Okay, to do that, now I have to bring CloudFront into play and it has an interface with ACM— the certificate management service that AWS offers—it has to integrate with S3, it has to integrate with Route 53—ostensibly a DNS provider, but here in reality, the world's greatest database—you have to get it to hook up to all of these different things, and as an initial starting out process, it's kind of formidable, at least from where I sit. And part of the problem is well, on the CloudFront specific piece is that the things that make it awesome—which is its flexibility— also make it terrifying of, “Wow. Those certainly are an awful lot of options of which I understand approximately none of them. Which ones can I safely ignore/take the default on/which one is going to turn my secure database backups into a new trending topic on Twitter?”

Katrina: Gosh, I'm so glad you brought up this topic, Corey, I came to Amazon, to CloudFront specifically, about a year ago, in particular, to focus on our developer experience. And while we've had some really big wins this year, like you mentioned, bringing our distribution update times down to under five minutes consistently, one of the things that is absolutely on our mind as a team, and on my mind in particular as a PM, is how to soften those sharp edges and just in general, reduce the burden of knowledge from the majority of the users who are coming to CloudFront for CDS—or CDN usage, rather—to something that's much more approachable. What I've found is that most of the users coming to CloudFront—and by most of I mean, on the order of between 80 and 90 percent of the users that come to CloudFront—do not ever change their default values, while still getting all of the perks of using a CDN, that acceleration and protection that we keep talking about. And if that's the case, how can I deliver a user experience to developers so that they know that they can safely ignore all of these other configurations, and when might I bring their attention back to the service when I think there's an optimization that they could make instead, intelligently. And more of those changes will be rolling out over the course of the next year.

Corey: One of the challenges, too, is even right now, the values that you articulate of having a CDN— the security aspect and the performance aspect— when someone is setting out down this path at the first time, they don't actually care about either of those things, to be direct. And there's a strong argument that maybe they shouldn't because I'm building a ‘Hello World’ website, and I want to put that in a S3 equivalent, and then I want to put that out a website to the world. Yeah, if that's not secure, I couldn't possibly care less as long as it's not exposing, you know, access keys for my AWS account because there's nothing in a security sense to worry about other than, “Please don't make the browser pop up a giant warning instead of the website.” And from a performance perspective, when you're building out, oh, this is a simple HTML page that says, “Hello World.” “How impatient are you?” is sort of the default response here of, I want to make sure that this is an easy to hit website, but if you're on the other side of the world, it might take half-a-second round trip to pull that up. But a CDN will make that faster. Well, who can't wait half a second after hitting a link on a website for it to load? The real-world answer is that down the road, things get much more complicated and there are security concerns up the wazoo. And, well turns out most of us write websites like crap.

And there are 30 sequential requests for the same thing, and that half-second latency for each one of those winds up multiplying and now your website takes two and a half minutes to load, so at that point, you may as well just store it in Glacier and wait for the retrieval times.

Katrina: I'm less concerned about that particular user of the Hello World app builder who's then adding on new and new tiers and then making their way to production by adding all of these different things onto their website. Maybe you start with Hello World and end up with a really cool site to sell Billy the Platypus merchandise on, where people can buy it from.

Corey: We are working on that.

Katrina: [laugh]. And that's a great use case. But at the same time, typically, as you're building up those building blocks, something tells you that you need this protection. Maybe when you're implementing your checkout process, you realize that that encrypted field could be better powered by using signed cookies or something like that. And then you get down this rabbit hole and discover that's a feature available of a CDN, or that you don't have to write it yourself, or something like that.

What I'm possibly more interested in or maybe just more top of mind, for me, is the opposite use case where someone comes to AWS and says, “I have the world's coolest merch store. I need to put it up online. I know either not very much, or very little, or I just can't be bothered with knowing the intricacies of what a routing system is, but in figuring out how to set up my first eCommerce site, it tells me that I need something called an SSL/TLS cert, something called domain name registration, something called a CDN because that DDoS thing sounds scary.” How do we make that experience as easy as possible from that entry point, which is that you already have a predefined goal in mind. You already know that you need to serve this content that is really hot for your users in a way that's not going to negatively impact neither them nor you.

Corey: Yeah. There's a lot of opportunity there. I feel like those are the cloud stories of the future. I mean, this idea right now feels almost like a temporary aberration where you have a developer who wants to get into doing these things. It seems like the much more common path, down the road if not today, is going to look a lot more like, I'm just new to the workforce, or I'm changing careers, or I just graduated school, and I have an idea for a business.

That business involves computers in some form, as most things do. Today, the answer is, you have to go to basically software engineering school, or at the very least cloud school to start understanding how a lot of these things fit together. Down the road, I'm more and more convinced all the time that that is not going to be a permanent state of affairs.

Katrina: Oh, I totally agree. And I would challenge you on that even one step further, which is that I'm more curious about the folks who set out to solve any particular problem, and that solution tends to have some sort of technical impact to it usually involving a website in some way, but more so than saying, I'm going to set out to create a business that's going to do this particular thing, sometimes it's just, “I’m really tired of having to keep track of what's in my fridge, so I'm going to start in an Excel spreadsheet of all of the ingredients, and then maybe I want some sort of logic on top of it that removes them as I use them. And oh, that gets really cumbersome in Excel, so I'm actually going to throw that into this simple webpage. And I know how to do that because there are a billion tutorials and blog posts out there.” And now you've just backed into creating your first app when you never set out to create an app. [laugh].

Corey: Yeah, it's almost on some level, too, where we start to see the parallels between that and career aspects and career paths, where I will never be a software engineer professionally, according to some people because I didn't go to Stanford and get a degree. Awesome. That is certainly a take people can have. But you're deeply involved in this and have a level of insight into the running of a CDN that I'm never going to have because I'm not getting the same hands-on experience that you are. But your degree is in I believe international studies and biology.

Katrina: That's right. My undergrad degrees are in international studies and biology. I was pretty convinced I was going to be a neurosurgeon for quite a long time.

Corey: And now you wound up going through a variety of different career moves. You were at Pivotal and then VMware when they did that whole shell game, acquisition, divesting. I don't even know who owns what anymore nor do I particularly care. You were in a bunch of other areas: loyalty marketing, you were in large financial services, and now you're a senior TPM at Amazon. What about your background makes you—I don't know how to put this politely, so I'm just not even going to try—relevant to that?

Katrina: Yeah. The thing that makes me relevant to being a really good product manager is caring, at the core, about what users’ experiences are, whether it was in loyalty marketing, or working in startups on real estate CRMs, or any other job that I've had so far. At the end of the day, I just want to solve interesting problems with simple solutions that actually impact the lives of users in some way. Sometimes, that's a really clever problem, and sometimes it's a problem of how can I make you care less day-to-day about the thing that I work on by the end of the time that I'm working on it. So, for me, that variation in my background, having a lot of different roles that have exposed me to all sorts of different industries and therefore different users, has allowed me to think about or to be better at understanding and seeking out what my users’ needs are, and then bringing those into a business and, frankly, just advocating really well internally, by being able to describe what impacts they have on users and therefore what impact that has on our business, and making good cases for developing better experiences.

Corey: One of the most valuable things that I see time and again, is folks who have come to this industry—by which I will round to Cloud as a whole— through, I guess we can call it non-traditional paths, but that always feels weird because I know more people who have come to this industry through those paths than folks who went and got the degree at Stanford. So, I really do object to that framing on some level. But I find that the folks who have that wild, varied background lead to teams that generally tend to be stronger, which is something that the studies have borne out as well. Now, that is the abstract. In a much more practical sense on some level, it's a CDN.

This is not the sort of thing where I would—going to put up a bit of a straw man for you to knock down, so please don't write me emails, listeners—it seems like, on some level, the idea of having diverse and inclusive teams, including a pretty wide diversity of background isn't the most important thing when you are running a CDN. It's not something that is going to, for example, suffer from algorithmic bias issues where we're seeing a lot about that in the news these days. It's not going to, for the most part, wind up being something that is a differentiation along axes that are problematic to various constituencies. It's basically a, we make the network faster. The end. What is the value of having a more humanistic approach to these things?

Katrina: I think the value of having a more humanistic approach to providing a really robust and usable CDN is understanding that if CloudFront, as an org, only hired die-hard DevOps folks who have only ever worked on DevOps, only have these extremely coding-rich technical backgrounds who just love building extremely beefy apps, then the way that the product is going to be made is going to reflect that. The problem is that users of CDNs vary wildly away from that. There are people, like we mentioned, who want to tinker around with every configuration or who have to because they run really heavy and robust workloads, like maybe video streaming platforms, or other applications with highly concurrent needs, like big banking applications, and things like that, all the way down to the Billy the Platypus commerce site that you're building right now, and not having that same level of expertise. And what I think is the value of having, and what I've seen is the value of having folks come in with a variety of different backgrounds is just a greater ability to empathize because they can relate to different people's experiences. And when you have a thousand of all of the same exact person working on a thing, then you're going to end up with a thing that reflects all of those thousand people.

But if you have a thousand different perspectives working on a thing, you're going to end up with something that works well for, well, more than just that one persona type.

Corey: No, you raise an excellent point. There are ways now to get things up and running effectively with CloudFront for those folks among us whose first language is not JSON. So, you're right in what you're saying. The user story is becoming much more simplified; the ability to have things start being removed from the console. I don't know if RTMP— if I'm getting an acronym right or not— is still an option that is actively promoted on the CloudFront side, the fact that I could do things with WebSockets now if I know what those are, but it doesn't nag the heck out of me about them and make me feel like I'm a fool for not knowing.

It's becoming much more approachable to the point where, do I have to use this, or is there a click by numbers option I can use instead? There is something to be said for folks who are representing different customers getting a voice.

Katrina: Absolutely. And I would also challenge the notion that there is not algorithmic bias necessarily in CDNs or anything like that. I mean, there may not be yet, but as we continue to invest in different compute at the edge and logic at the edge functions where you can write all sorts of really interesting applications that instigate really interesting network impacts, like logic that may, let's say, have downstream impacts on the rest of your infrastructure, but only doing so when a customer requests it, the way that we continue to invest in compute at the edge and as we continue to push different logic to the edge, what we have to keep in mind is that the edge is simply just a proxy to the rest of your system if not configured right, or sometimes if configured exactly right. So, whatever logic is being built and introduced to the system can be really harmful to end-users, and we take that responsibility very seriously that we're giving people, as we continue to build CloudFront, continue to build different at-the-edge functionality, that we're giving people more and more power over these systems that are, I mean, just deeply ingrained in all sorts of networking infrastructure. And the more that you're able to fine-tune and customize there, the more possibilities of [laugh] different attack vectors there are at the same time. That sounds really scary, but all of that to say that the more flexibility that we offer, and the more that that investment in the edge logic continues, I see more and more possibility for just more creative usage.

Corey: I will say that I had the privilege of talking to a sarcastically senior engineer on the CloudFront team a year or two ago after I was doing my typical way of filing bugs against AWS services, which is complaining loudly on Twitter. And they had the grace to always respond with that typical Amazonian response of, “Interesting. Can you tell me more about your use case?” As opposed to the much more direct slash honest, “What the hell's your problem?” Which is basically what I deserve. [laugh].

And the answer that I had at the time was—I was complaining about the 45-minute distribution updates. And the problem that they said—and they're right, by the way— “Do you want us to just lie to you and come back when we wind up having the update not done, and you'll be serving different data than you think to customers? Or do you want it to be technically correct? Because we strive for correctness here.” Because that's a great question.

It takes time to update stragglers and the rest. And my answer was also a very legitimate, “I could not possibly care less. The reason I care about this at all is that in order to update a Lambda-at-edge function, I have to wait for a distribution update to finish.” There's no option for me to just deploy it to a single POP, basically, instantaneously in the region I'm developing in; I have to wait, at the time, 45 minutes to discover, yep, I forgot to put a comma in there and there are syntax errors because see previous notes about terrible coding practices that I do.

And that was a use case that, it turned out, was echoed by more than a few people. We’re all secretly terrible developers, as it turns out. And that understanding was taken and obviously acted on because it's now way faster. But it's still challenging, in some ways to have a four-minute wait when I want to see if I forgot a comma again. But it's more tenable.

It means that writing a Lambda-at-edge function isn't going to be something that's going to take two days for me to spit out before I finally get it right. Things improve all the time, and it's easy to complain about the things we don't have, rather than the things that we've been given.

Katrina: It is often easy to complain about the things that we don't have versus the things that we’re given. I, as a PM, absolutely understand, though. Sort of like, as users, we’re conditioned to always want more to always want better, we're conditioned to want the perfect thing that works for our exact use case. And to understand [laugh]—and, as you said, complain loudly about that when it doesn't work out. I take that responsibility quite seriously.

If there are ways that we can improve our users' lives, like continuing to work on the change propagation time, so that it continues to be lower, lower and you have more of that instantaneous feedback. There are certainly other ways of solving that problem that we've heard lots of requests for, like different sandbox environments, for example, or a blue-green deploy system or, like you said, the ability to specify a specific edge locations to where you can deploy code. For some people, that's a persistent request. They want to always be able to deploy certain logic by edge location at a very granular level because that suits their business needs. Taking all of those and weighing them in not only for what will work for a subset of our user population, but what will work for all of our consumers at large as part of a major in-demand CDN with customers of all different sizes, all over the world, from single dev shops to massive multinational corporations.

We have to consider—what I try to consider as a product manager is what are the features and the, sort of, bells and whistles and the customizations that I can offer that impact the most of our users in the most applicable way, rather than how do I design for just one customer's need?

Corey: Some people need dozens of tools to visualize their stack. But not you. You’ve got New Relic Navigator. All your hosts, services, containers, and anything else you can monitor are in one dense hex view, so you can check system-wide health at a glance.And you can sort by platform, service, app, or any other tag, which makes it easy to navigate yourentire system and see what needs your attention. Go to https://newrelic.com and get started for free.

Corey: And this is the thing that I've had to learn again, and again, and again. Every time I think I've seen it all in the context of a given AWS service, or even AWS as a whole. All I have to do is talk to one more customer and I'll discover a use case that I hadn't considered or hadn't envisioned as being something that real people did. No judgment intended on that one, by the way. And you really do go about learning from customers as you go.

But it's also hard, particularly on the CloudFront level, where of all the environments I've been in, there have been remarkably few requests I've had of a CDN that were not a modified form of either, make it easier, or make it faster, or get out of my way more. Those are the only three requests that everything winds up in a CDN piece that I ever saw going down that path. And I thought that that was it. Okay, CDNs are boring, why does everyone care? And then I started talking to customers. So, from that perspective, what do you think, in what you've seen, is the most misunderstood aspect of CDNs and/or the most curious use cases you've seen?

Katrina: The most misunderstood aspects of CDN that I have experienced in my work has been this notion of—back to the beginning of our conversation in this podcast— the notion of my CDN should be everything to me. This kind of conflation of CDN and application, or Jamstack, or application builder tooling that people tend to sometimes come to CloudFront with a very particular need, or they come to CloudFront thinking, “This is where I'm going to build my entire website, and by the time that I'm done with the console, I'm going to start with nothing and have a working website that will allow my users to purchase the thing and check out.” And that's been a really interesting space to, kind of, step into and figure out what that balance is, for me as a PM, in terms of where I want to continue to advocate for taking the product. Do we want to be everything to everyone, or do we want to do our core competencies really, really well?

In terms of the most interesting use cases that I've seen in the CDN space, gosh, it’s sort of limitless. Some of the more fascinating ones to me are understanding how really large players out there you see CNS in incredibly complex ways. Talking customers who have hundreds, if not thousands of CloudFront distributions, all of that do something a little bit differently and serve a unique purpose for their tech stack, typically always going towards one final application. So, the interconnectedness of it all completely blows my mind, especially when I use the end-user-facing product in realizing that when I navigate, say, to a website, and let's say, watch a movie or something like that— say that I'm watching something on Prime Video, let's say, it blows my mind that complexity that can be underlying for this in a way that customers somehow can keep track of all at the same time and continue to iterate and improve on, and it just kind of magically works for me as an end-user.

Corey: That's the biggest thing that continues to amaze me almost in equal measure. First, the fact that I can wind up having some crappy line of code and push a button and it gets deployed at world scale for fractions of a penny, followed secondly by the fact that I am pissed off about minutia of that experience and want to take to Twitter and drag AWS for it. Those two senses of wonder are always competing in my mind. And on Twitter, I do bias for the ladder, I admit, but it's magic; you have built what is fundamentally magic.

Katrina: We have. We built fundamentally what is magic through a lot of conversations with users and less back-and-forth debates about what is doing the right thing and then, of course, fast iteration and learning from it. There's so much that goes on, as with any technology, but behind the scenes, even in what we'll call a slow feature time, if you'd like to call it that, just all of these underlying improvements and underpinnings that continue to make the experience better and easier to use and at the same time, enable people who didn't know anything about a CDN before to deploy their code to over 220 edge locations around the world. The idea that someone who's never written a line of code before can, within a day for sure, learn all there is about it in tutorials, and online, and whatnot, and then come to a service like CloudFront and deploy that brand new code, the first line they've ever written, absolutely mind-boggling to me.

Corey: One thing that never occurred to me to ask before, but since I've got you on a recording and it's impossible to say no when I ask these things, I'm going to ask my question now about this. AWS is somewhat famous for not having a global control plane; it's avoided things like global outages of every region at the same time. Those boundaries don't really exist with CloudFront, I have never yet seen an announcement of this new feature has come to CloudFront—and there have been a lot of those— but only to the following three global regions with more to come in the future. Features are rolled out either globally, or not at all, and I can't remember, although I'm sure someone will have some document somewhere that’ll yell at me for this, a single global CloudFront outage. How does that happen without the blast radius being either monstrous, which it clearly is or the pace of innovation slowing to a crawl which, snark aside, it isn't?

Katrina: Through a lot of diligence, Corey. That's the [laugh] that's the best answer I can possibly give, having this extremely globalized service, that is still though, by nature, are at the core of being a network of interconnected physical locations deployed all over the world. Those physical locations in each particular place are separated, so we are, to some extent, naturally protected from large natural disasters by having edge locations in each place, much like different data centers or availability zones. There does exist that aspect of that fragmentation. The flip side of that is that we also are extremely conscious of the fact that CloudFront operates at a global scale— with a small caveat of, except for in China because of the Great Firewall of China and having to operate as a separate legal and network entity within China— that globally we have this responsibility of making sure that the software that sits on top of all of those physical edge locations can reroute and does reroute traffic consistently based on what's available, and not only what's available, but what's most performant.

So, we continue to keep it at absolute top of mind to be able to redundantly support workloads through any sort of localized or even regional outages that may happen so that we can continue to serve the information that we need to.

Corey: It's never ceased to amaze me just how much of these various concerns AWS has to hold in tension with each other. Because customers want new features, and they want them quickly. They also don't want you to drop their entire banking presence onto the floor while doing it. And every time I had a feature request, the answer has been, there are extremely complicated reasons why it doesn't do X. I have never yet come to an AWS team with a feature request and had them drop a cup of coffee and, “Holy crap, we never thought of that before.”

It's always the, “We'd like to, but it's really hard.” And on the one hand, I'm very sympathetic about that, and these things do take time. On the other, I'm not paying my fractions of a penny so that AWS can only solve the easy problems. So, again, it's a study of contrast, a study of tension. And I feel like just through this conversation, I now have a better idea of what product management is than I did when I started half an hour ago.

Katrina: That's lovely. That's a great meeting of my goal then, of spending this time with you. [laugh].

Corey: Exactly. So, one question I have before we call this a show is, what is the best and worst you've ever felt about your job as a product manager?

Katrina: Oh, gosh, that's a great question. The best—or the worst, I'll start with, which is more fun.

Corey: The worst is always the good stories. It’s the, “What do you regret the most?” It is always a great question to ask someone. It's super weird when you ask it in a job interview, but that's a separate story.

Katrina: [laugh]. I would love to hear Screaming in the Cloud of just, Corey’s regrets. That would be lovely.

Corey: Oh, my stars. I don't know if anyone else would. That's the problem. Once I start complaining, I don't stop. But you're not getting out of the question that easily.

Katrina: The worst I've ever felt my job as a product manager was sitting down, I had worked on this very large feature release with my team for about six months at another company, had had tons of user interviews felt really good about the research that we did at the design phases; had these iterations with real live customers in hand, testing things out, did all of these recordings, like, couldn't have done anything better by the book. Sat down when it was in beta with the customer, had them test it out. And the customer just struggled. They could not get through the very first thing they needed to do to even log into the application, much less use the new feature, which was this really tricky authentication maneuver. So, they sat down, they struggled, they struggled, they struggled, they couldn't get any further.

And finally, this customer, who I knew really well, had built a good relationship with, turned around to me and said, “I’ve been doing my job for 20 years and I've never felt worse at it than I do today. And it's because this experience has been so rough.” And as a product manager, that's about the worst thing in the world that you can hear from a user is that you're making them feel bad at their job. And the lesson that that very brutal experience taught me was that while being so focused on the new and exciting thing, I was completely ignoring these really fundamental steps that it takes to actually access that new and exciting thing, which in this case was something as simple as a back end authentication.

Corey: That one absolutely resonates across the board. I believe wholeheartedly, that if a user is working with an AWS service and feels dumb, that is a failing of the AWS service team. And that's a very hard line that I draw, in that none of the services that AWS offers—mostly. I'm going to carve out exceptions for some of the advanced ML stuff, and certainly the quantum computing folks. Don't get me started on that.

But for most of these things, the fundamental concept of what the various services do is not that complicated. Someone who has gotten far enough to have spun up an AWS account specifically should be able to grasp on some level what these things are. And every time they're struggling trying to get something done, in an idealized version of the world, that's a service team failure and is an area in which a customer can reasonably expect the cloud provider service team to do better than they have. It's the divinely dissatisfied in perpetuity approach.

Katrina: Absolutely. Anytime that I can free up my users’ busy time to get to the core of whatever it is that they have to do. With the CDN, for example, I want to give my end-users time back to do whatever they want to do. That is not spending hours figuring out which CloudFront configurations apply to them. You would much rather be doing something else, be it coding an application, be it spending time with your kids, be it literally anything else in the world than trying to figure out what this very, very particular setting might be.

And for me, that's a big part of that, is making sure that things that users have to do routinely are easy and fast and get them to the core of the thing that they have to do or that they want to do. So, basically, removing the obstacles from the things that you have to do to get you to the things that you want to do, faster.

Corey: Yes. I do want to caveat my statement as well because I know that there is some junior product manager out there somewhere at some startup, who's going to hear that and think, oh, well, our predictive algorithms are really good, so we're just going to go ahead and send things to our customers that we know that they would like and charge them for it. No. Stop. That is called credit card fraud. No. Stop it, you still need to get consent from users to sell them things. The fact that I'm not entirely sure whether that disclaimer was a joke or not, should tell you a lot about the state of our industry.

Katrina: [laugh]. We want things to be easier to use. I think that's a great fine line to call out, that I don't want to be making predictions about what I think you might want and get that wrong. That feels worse than just giving you the option to do what you want super easily.

Corey: Right. There's a reason that none of the cloud providers have implemented the web console version of Clippy of, “It looks like you're building a CloudFront distribution. Would you like help?” No. Because the first time it gets it wrong, it annoys the hell out of everyone.

Katrina: Damn. Cat’s out of the bag. Roll back the feature release, everyone.

Corey: Exactly. Yeah, but I assume it's also going to be Amazon-centric, so it wouldn't be a paperclip because that's too derivative. Instead of, I don’t know, a binder clip.

Katrina: [laugh].

Corey: Or a nail gun. There we go. The jokes would write themselves at some point. So, that was the worst experience that you had as a product manager. But I don't like ending on down notes. What's the best?

Katrina: The best experience that I've had working as a product manager was—different company again, a couple of years back. I took the time to sit down and, in laypeople's terms write out what a platform as a service meant. It ended up, I guess, setting me up well, for my career at Amazon because it was about six pages of a narrative of just what a platform as a service was, what the Cloud was, how it integrates, how servers play together. Wrote an analogy about how these things interconnect and why it matters. And just sort of understanding that there were probably a lot of people in my same shoes, maybe even working at my same company who, despite working on a platform as a service, didn't quite understand what it was that it did and why that layer of abstraction was important, much less the underlying infrastructure as a service that powered that thing.

And taking the time to write that out in really simplistic terms then created the experience where when I shared it out, I had several folks come up to me and just say, “Hey, I've never really understood that thing. Thanks for making it approachable, and thanks for making me understand this piece of technology in a way that I otherwise wouldn't have.” And lowering that burden of information that you've had to go seek out yourself for other people that I went through the experience of finding myself in, and trying to understand and reading through incredibly dense technical blogs to understand something that could have been simplified much better has probably been the pinnacle of my product management experience so far. Just simply being able to take something incredibly complex, describe it very simply to people, and removing the burden from them having to go self-serve that information through the same pain that I had to understand it was a really great experience for me.

Corey: There's a lot of power in being able to, I guess, have the position in the market—which I do, and I basically took it because I figured, what's the worst-case someone's going to do? Take away my birthday—and say, “I don't get it.” And that tends to be valuable because every single time I've done that, people have come forward and said, “Oh, thank God, it wasn't just me. I just didn't want to say it out loud.”

But there's much more power in being able to take that feedback and fix it; make it understandable in a way that is extraordinarily approachable. There's a reason that people don't just ask for a list of API options as documentation. They want to understand what business problem do I have that this solves for me. They want to see code samples, but they also want to see something in between the very high-level and the very low-level that explains considerations of using this, the trade-offs, the design reason why you would use this, and very often, why you wouldn't use a particular feature, product, or service. Anytime someone says, oh my product’s awesome and you should use it for everything, they're doing a sales job and not a particularly good one at that.

Katrina: Absolutely. I think the really apt analogy here is, you should disseminate whatever information you have and whatever knowledge you have, in a way that is consumable to your users. And the analogy to me that comes up is a water fountain. If someone is thirsty and—thirsty for knowledge, in this case— and you are trying to supply them with water, you might think that the very best thing that you can do is open up a fire hose on them because it's more water than they could ever use, and it's fantastic, and you're giving them all of these things. But you're also going to drown them. If you give them something that's more approachable and let them opt into consuming more as they're ready, it’s a much better experience for users, and you're enabling them to have their needs met, to opt into more information, to self-serve that at the same time, and then continued to grow with them.

Corey: I think that's a terrific sentiment to end the episode on. If people want to hear more about what you have to say, where can they find you?

Katrina: Folks can find me @katrinabakas on Twitter. Feel free to reach out anytime.

Corey: Excellent. Thank you so much for taking the time to speak with me. I appreciate it.

Katrina: Thanks, Corey. It was my pleasure.

Corey: Katrina Bakas, senior technical product manager at AWS’s CloudFront. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you've hated this podcast, please leave a five-star review on your podcast platform of choice, and in the comments. Tell me what year you graduated from Stanford.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey atscreaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Randall

Randall is a Software Engineer and Open Source Developer Advocate at Facebook. Previously of AWS, SpaceX, MongoDB, and NASA.

Links:

  • Totes Not Amazon: totes-not-amazon.com
  • Twitter: https://twitter.com/jrhunt

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the Enterprise (not the starship). On-prem security doesn’t translate well to cloud or multi-cloud environments, and that’s not even counting IoT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IoT devices, detects these threats up to 35 percent faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at extrahop.com/trial.

Corey: If your mean time to WTF for a security alert is more than a minute, it's time to look at Lacework. Lacework will help you get your security act together for everything from compliance service configurations to container app relationships, all without the need for PhDs in AWS to write the rules. If you're building a secure business on AWS with compliance requirements, you don't really have time to choose between antivirus or firewall companies to help you secure your stack. That's why Lacework is built from the ground up for the Cloud: low effort, high visibility and detection. To learn more, visit lacework.com.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I've been waiting for this one for a while. I'm joined this week by Randall Hunt, who is currently a developer advocate at Facebook, but for years, was also a Developer Advocate slash evangelist slash gadfly slash I can't believe he hasn't been thrown from the building yet at AWS. Randall, welcome to the show.

Randall: Hi, Corey, thanks for having me.

Corey: So I've been trying to get you onto the podcast for a long time, and to their credit, AWS PR did everything in their power to never respond when I explicitly requested you by name, which makes perfect sense because you put the two of us together, we feed off of one another. And suddenly, now it's an incident.

Randall: It's a parasitic relationship. Symbiotic? I don't know, one of those two.

Corey: It's a negative feedback loop, I suspect. But now you don't work there anymore and I never have, so the gloves get to come off. But first, before we dive into what was, let's talk a little bit about what you're doing now. You're a developer advocate at Facebook, as mentioned, with an emphasis on AI and more specifically, something called PyTorch. To my understanding, PyTorch is like TensorFlow except Google didn't create it. I'm betting there's more nuance there. Tell me more.

Randall: Yeah. Well, interestingly, I would say TensorFlow is more like PyTorch now. So, PyTorch is the culmination of years and years of open source projects across a couple of different platforms and ideas. And it's really—if you've ever heard of a framework called NumPy that does a lot of kind of statistical analysis and other kinds of data manipulation techniques, and what PyTorch does is it's an open-source framework that does everything NumPy does and more, but with GPU based acceleration, and a lot of kinds of runtime improvements that you couldn't really think about doing. And the really amazing part is TensorFlow, kind of, realized the flexibility of PyTorch back, I would say, 2018 or even earlier, and TensorFlow 2, it uses the same kind of API that PyTorch does.

And I think almost all of these frameworks have started to move towards the execution methodology that PyTorch enables. And it's a giant open-source project, really huge community going around it, and I'm pretty proud of all the work that everyone has done on it, from professors at Cornell to kids sitting in their garage, like, blasting techno music. It's just a pretty huge breadth of people.

Corey: One of the things that I always appreciated about your talks that you gave when you were at AWS—and also back in that era when people went places and sat face-to-face in rooms with no masks on. What a concept.

Randall: I don't remember this at all. [laugh].

Corey: Right. You would get on stage and you would give talks and you were basically everything I swore I would never be on stage. You would write code live and see how it went. Yeah, you were someone who went deep into technical weeds and did live coding rather than a contrived demo is the way that I prefer to do it because the demo gods are indeed spiteful. And sometimes it worked.

Sometimes it didn’t. You always turn it into a story. But more to the point, you were one of the shining lights at AWS when it came to not just telling me how to use a given service, but why. The value it got from this because you would start out by writing some silly bot to do something, and your stated purpose at the beginning was, “Let's build a bot to do X.” And now suddenly, I am coming on board with you. I'm going down that road. Even if it's patently ridiculous, you painted the picture of what was possible. And that's something that I don't think that a lot of AWS storytelling focuses on to its own detriment.

Randall: Well, thank you for saying that. Yeah, I always really enjoyed doing live coding. To be honest, I got my start—we won this TechCrunch Disrupt nonsense with a team back in… I think it was 2011 or something. Maybe even 2012. And I had never live coded on stage before, but our demo broke while we were speaking.

So, I started live coding and the audience just went crazy and laughed, and I was kind of joking as I did it. And I got addicted to the stand up comedy-esque thrill of trying to code live on stage. And I've done it ever since I guess.

Corey: I feel very similar to you do, but the thing that really was revelatory for me and it sounds like it is for you, too, it's the wrong framing to view it as stand up comedy. Because I thought that's what I did. Turns out, absolutely not. Stand up comics, rehearse, and repeat, and prepare, and write constantly, whereas what I do is show up unprepared, which is properly known as improv. So it's stand-up style, but it's a lot less work slash planning slash caring about the outcome, apparently.

Randall: Well, I think there's a rehearsal component to it. So, I found the more successful talks—

Corey: Oh, here's where we diverge. You actually prepare for things. Oops.

Randall: Yes. I had to learn that the hard way. [laugh]. So, I didn't start off doing that. In my early 20s, that was not the case. But I found that the best talks are the ones that are a mix of extemporaneous improvisation, even audience prompted so the audience feels like they have a stake in the game.

So, if you take a poll of the audience, and you say, “Hey, what should we build today?” And you can kind of take suggestions and craft them into something you already have mentally prepared in your head so it seems more extemporaneous than it really is. So I'm kind of revealing my trick here. If you ever see me go do a live coding demo, almost all of them involve some form of audience participation. Voting: so going and saying, oh, let's vote on Vim, or Emacs or something like that, these are things that the audience can get very involved in. And I definitely have every possible way of building a voting app memorized at this point, so it's cheating a little bit. I'm not coding completely from scratch; it's in my head, I've done it before. Most of the time.

Corey: So, help me recapture that old magic, I guess. My position on AI slash machine learning is it's a cleverly executed scam across the board because all the cloud providers are saying up and down that you need to do machine learning to remain competitive. And okay, great. I would expect that when you scratch beneath the surface and you look at what machine learning requires, which is basically a whole bunch of compute and a whole bunch of storage. So yeah, I can get why the big cloud providers who sell both of those things would be interested in advocating for this.

Randall: Dolla dolla bills.

Corey: Exactly. But when you look at the stories of people using machine learning or AI in the real world, they always have the ring of either something that is extraordinarily use-case-specific to one company or is patently ridiculous. I mean, the classic example was WeWork used advanced machine learning algorithms to learn that there was a bottleneck in several of their facilities in certain times in the morning, and they alleviate that by hiring a second barista. Not kidding. The response is, “Wait, you spent how much on data science and machine learning to figure out that people like to drink coffee in the morning?” It just seemed insane from that perspective. So, change my mind. You've been a big advocate of it for a long time; what is the business value other than it makes you sound more expensive and thus can command a higher salary? Which, respect.

Randall: Okay, you mentioned WeWork and I just want to say if anybody from work is listening to this, the WeWork right across the street from my house has had a light shining into my face for literally a year. Please turn it off. Okay. Now let's focus on machine learning. So, when I think about machine learning, I think about it as on the x-axis, you will have say compute, and on the y-axis, you will have the amount of data that's available to you.

And there's a great graphic—[maybe I can send it to you later 00:09:52]—that the [sci-fi 00:09:54] team produced which says which machine learning techniques or deep learning techniques should be used for which sets of data and things like that. Now, traditional machine learning, it's everywhere in the world around us. So, the voice activation filter on the microphone that you're using, that is likely using machine learning now, instead of using some hardware pop filter or something. The noise subtraction that you hear in audio engineering will be using deep learning now, instead of trying to do it with frequency analysis tools. The ad bidding all of this other stuff, the typical specific use cases, those are all still using machine learning as well.

But it's a little bit of a misconception that you need to build these bigger and bigger and grander and grander models. There's definitely a part of the industry that's moving that way. So there's a part of the industry that says, “Hey, let's have these terabyte models that have billions and billions of parameters.” There's a great talk by Geoffrey Hinton back in, I don't know, I can't remember the date. But there's a great talk by Geoffrey Hinton that goes over the number of fixations that a human brain makes.

So, a human brain is what a lot of artificial intelligence machine learning is based on. Like, my entire Twitter account, for instance. And if you take the number of fixations that a human brain makes, it’s, like, 10 billion fixations over the course of its lifetime, where we can outperform human brains with way fewer parameters than that. So, the ways machine learning are being applied and the things that are being done these days are ubiquitous. I mean, it really is everywhere, but the things that are working and the things that are easy are the ones you don't see.

So, machine learning, when it's done correctly, you don't even realize it's happening in the background; you don't see what algorithms are going on and doing these predictions. And so you might not see it, but it's pretty much everywhere you're looking these days.

Corey: Yeah, my experience with it is a bit more toward the ridiculous. One of my employees, a while back, when we were just friends and we didn't work together, we ended up writing totes-not-amazon.com—that's totes dash not dash amazon dot com—and it uses a Markov generator. Every time someone visits the page that is trained on the corpus of all—I think what is it—12,000 AWS service announcements dating back to 2004. And some are hilarious, some make no sense, and some are hilarious and make no sense. And the funniest ones, of course, are where you don't actually know if it was lifted wholesale from an actual release announcement or not.

Randall: Yeah, I think that's a funny thing. These models now, like GPT-3, and Bert, and other things where you're starting to wonder if a human wrote this as satire or [laugh] if it's actually AI.

Corey: Well, to be fair, that’s some of the release announcements that humans apparently do write. Or I assume with humans; it could just be machine learning experiments. My personal favorite is whenever they come out with an announcement that I wind up reading, and at the end of it, it was, “Well, those are certainly a bunch of words, none of which I understood in the context they were used.”

Randall: [laugh]. I just went to the site and the headline that I got was, “Amazon poly ads, bilingual: Indian English, Hindi language support for AWS Code Pipeline, AWS Config, adds support for non-RFC 1918 address ranges.”

Corey: That's amazing. That's the sort of thing we're talking about. It's a blast. I did some experimentation with GPT-2, and that turned into some interesting stuff. But for that I had a bit more fun and turned it loose on, effectively, not the release announcements, but all of the blog posts that Jeff Barr wrote, dating back to the launch of AWS.

And it's fun because Jeff has a personality, and it would capture entire sentences that I thought were hilarious, and that's ridiculous. What a human-sounding bot this is. And I would check the corpus and that exact entire sentence was used. And it was a lot more sensible in the context where it was originally found, it just becomes ridiculous when you surround it with other unrelated things. So honestly, having talked to people who are good at this, it turns out, I don't fully understand how training works and how to tune it to get the best results possible. I'm mostly punting until GPT-3 becomes more available.

Randall: Yeah. Well, there are even things on GPT-3 already as well. You know, I don't know if you remember back in 2018, I did something similar is, I scraped all of the blog posts—I used to write for the AWS blog. I think I wrote 50 or so posts over the years—and I scraped not only my posts but all the posts from all the authors on all AWS blogs. And I created a model that would generate—sort of like this Totes Not Amazon site, it would generate fake blog posts.

And I did this all live on Twitch. So, that kind of goes back to the live coding is I think there's this general idea in the industry that machine learning is some secret. And I think the machine learning engineers are kind of incentivized to keep it that way because they can earn the big bucks as long as it's a specialty. But the reality is, most of the machine learning techniques and training techniques are all, sort of, automated. So you're saying do a softmax here.

Okay, do a activation function, do this layer, do this fully-connected layer, whatever. All of that stuff is completely automated at this point. You don't really need to think about it. You can tune it, you can play with it if you want, but the gross majority of machine learning these days is really just data preparation. If you look at the Twitch video that I did a couple years ago, we spent more time writing a scraper to download all of the blog posts than we did training and deploying the model onto SageMaker.

So it's interesting to see all this focus from all the different industries and blog post talking about how ml is this huge, big thing. But it's data science, again. That's where everything's at is kind of doing feature engineering and data science. And I wish I would see more focus in the industry on that side of things.

Corey: That's the challenge, though is the entire industry is expanding, and it's getting bigger all the time. Which reminds me a lot of things from yesteryear, and really, the reasons years ago, I wanted to have you on to have some of these discussions. So let's do it now. Pivoting away from the machine learning morass for the moment. Let's talk about big cloud providers and their growth and what we're seeing in the market. So let's begin with a provocative thing that will start a fight. If you could change one thing in AWS, what would it be?

Randall: I tweeted about this. I would love for there to be within AWS an S-team level goal, an Andy Jassy level goal that says, all of our APIs that we release and all of our developer experience that we release will be gatekept until this number of internal developers agree that it's a pleasant experience. I can rant more about this, but if I had my druthers, if I could change anything, I think the single most important change that could happen would be to have developer experience as a core goal for every single service team. Except WorkDocs. [laugh].

Corey: Well, I love the idea in principle, but it sounds like it could turn into a few things that are terrible. One being that then you wind up with the perfect becoming the enemy of the good and never shipping anything. And the other is that it's become very apparent to me over the years, watching AWS releases and having a launch-day product that is—how to put it—basically crap. And it learns from its customers, and two years later, you look at that product again and it is world's better. I don't know that you'd be able to get it better if you had, A) launched later in the process without that customer perspective, or if you had not built things that were reflected by the use cases customers put them to.

Randall: I think that's a fair call-out. And I'll just kind of say, yeah, but AWS is pretty well established at this point. They can afford to take an extra seconds to get things done correctly or to do a finesse and polish check. I don't know if you remember the Elastic Kubernetes service launch back in say… I can’t remember. Was it 2017?

That was essentially at the time, not a real service. It was, “Here. Run this AMI.” And that was the entirety of it. That's not a pleasant developer experience. So, everyone who had access to the service in preview was sitting there complaining about the service and the complete, really lack of a coherent presentation layer. They weren't asking for new features because they were too busy complaining about how frustrating the API and that side of things were.

So, I think there's a balance, obviously, and businesses want to move fast. But here's the thing, the kind of thinking that got AWS to where they are right now is not the kind of thinking that's going to take them into the future because the problems that they've created with this rush of onslaught of new services and new information is impossible for an individual developer overcome with a current developer experience. And I can go on at length about this, but—you know, the particular services that I would call out. But if I harken back to the early days of AWS, when EC2 came out, for instance, one of the things that made me go to work for AWS in the first place, was they never had a run instance API. So, you think about launching an EC2 instance, right?

Corey: Mm-hm.

Randall: It's like, “Okay, cool. Run instance.” It wasn't ‘run instance.’ It was always ‘run instances.’ It was plural from the beginning. And that seems obvious to us in 2021. In 2011 or 2010, people weren't thinking about launching multiple instances simultaneously. I mean, some people were obviously but it wasn't a common theme back then.

So, they were thinking ahead about the developer experience and they weren't painting themselves into a corner that's impossible to escape from. This is why you have different versions of different APIs that are expressed across different services now, and you'll see different things, sort of reinventing themselves and relaunching as new services. Because that lets them internally get more buy-in and get more effort kind of devoted to this new shiny thing. And the real value, the things that would drive tremendous value for customers, even if they're not necessarily asking for it, would be to focus on the developer experience, the hands-on-keyboard experience of using AWS.

Corey: The challenge that I think we're seeing, too, is if you go back to its inception, AWS was very clearly aimed at engineers slash builders slash surly sysadmins, folks who for one reason or another, were working with the computers an awful lot. And it’s—at the time of this recording—turned them into what is basically a $46 billion a year business. And it's gotten them super far. And it's carried them super well. The challenge, though, is that their next $46 billion a year is not going to come from the same places that the first one did.

I mean, we talked about a $2 trillion IT industry and growing. A lot of the folks that are still running on-premises are now not building net new on top of Cloud, but they're migrating in. They have a different philosophy, they have a different approach. And when you say ‘developer experience,’ their response, quite reasonably is, yeah, we don't hire developers were just trying to run our corporate IT somewhere. And there's a lot wrong with that, but that is their position.

Randall: And I understand that, but I think it's a little myopic to have that view. I want to look at Microsoft for a second. So, Microsoft has made a pretty ingenious acquisition of GitHub. Microsoft also runs VS Code. The gross majority of new developers are learning and getting started with VS Code as the platform.

That is their IDE, that is what they spend all their time in. From high school and on, that is what they know. So, if you hire someone who's 21 today, it's very likely the only IDE they've ever used is VS Code. So, Microsoft now owns the code—where it lives—they own, with GitHub actions, the CI and CD of the code, they own the IDE, which is the developer experience. And admittedly, thankfully, this is all open source.

But where is AWS in that developer experience story? Because the acquisition of developers is the funnel that drives your long term business success. I agree that you have these great efforts that can go on in parallel for enterprise sales, and inside sales, and all this other go-to-market activities with these much larger customers. However, think about Twilio. Twilio, started by an AWS—or an Amazon product manager, Twilio became what they are today by focusing and obsessing over the developer experience. They have by far one of the best developer experiences. They set the industry standard there.

Corey: As long as we're not talking about their SendGrid division, I would wholeheartedly agree with you.

Randall: Yes. No comment on SendGrid. I haven't used it in probably a decade.

Corey: I used to. I love what it does, as far as an email [campaign 00:23:15] goes, but the API is kind of sad. If you're listening to this at SendGrid, please reach out.

Randall: [laugh]. So, I don't see these two efforts is mutually exclusive. Improving developer experience will improve your sales across the board because companies have to hire new people. Even if they're running on-prem, they're still hiring new people if they're a growing business. And if they're not a growing business, does it really matter if you're getting their business?

I mean, the economy is going to move on. You don't necessarily need to capture these dying companies; you want to capture the company that are going grow along with your business. And the companies that are growing are hiring new developers. Those new developers have a very Microsoft oriented worldview with VS Code, and GitHub, and all these other things. Where is AWS in that story? And that's what I would love to see change at AWS is this huge focus on developer experience. I think it would transform the company.

Corey: Incidents happen fast but they don't come out of nowhere like AWS bills do. If they're watching, your team can catch sudden shifts in performance, but who has time to check thousands of hosts, services and containers. Thats where New Relic Lookout comes in. Part of full stack observability, it compares current performance to past performance just like your not supposed in the stock market, then displays in an estate wide view of your entire system. Sign up for free at NewRelic.com and start moving faster than ever.

Corey: I would agree with you, and part of the challenge is that AWS is fantastic at building the plumbing, and they have trouble with a porcelain. Sure you can view that through whatever toilet analogy lens you want, but they're great at building the blocks you can use to construct something. But I think that they abdicate almost entirely—and we talked about this with the approach you take to telling stories—is that they don't tell the story about what you can do with those things that they're giving for you. Sure you give me a bunch of bricks and talk to me about building a house, but you haven't demonstrated to me what that house might look like. Help me out. And sure your customers can get on stage, but I kind of want something that stands somewhere in between the two continuum endpoints of ‘Hello World’ and Netflix.

Randall: Yes. I agree. And it's not like all of AWS suffers from this. It's just a majority, I would say there are two projects, maybe three projects that I'm a huge fan of right now. One is AWS Amplify. So, Amplify is an open source project, it's a service, and it's an console—

Corey: And a breakfast cereal, too. I kid.

Randall: [laugh]. Oh, I could use a breakfast cereal. So Amplify, has really just been completely focused on the developer experience and you can tell. You can see it in their documentation, you can see it in the way that their advocates go out and are speaking to developers and telling stories, and they're hitting a whole new generation of developers. So, a lot of developers these days, I don't know if you see all this Twitter controversy that goes on. Somebody tweeted about how they would never hire a front end developer, or front end developers are only junior. Did you see that?

Corey: Yes, I did. And in fact that that Twitter rando was the CEO of Shopify, which is—

Randall: Oh, my God.

Corey: —a bit of a challenge. I understand aspects of the sentiment, which is, for example, you will become a better engineer by understanding other aspects than just the area that you're focusing on. But, “Front end is just for junior crappy engineers,” is absolutely not helpful sentiment. I'm also not entirely convinced that that was the point he was trying to make. But there is the nuance, of course that I'm not here to do communications, PR spin, marketing messaging, et cetera for the CEO of Shopify. He can clarify his own statement. The way he put it was tone deaf and bad.

Randall: Right. And I see where that stigma or whatever is coming from, but the reality of the situation is front engineers today are, by necessity, are some of the best engineers out there. And it requires a complete understanding of a complex set of services. If I go right now, and I want to build a back-end that can scale to millions and millions of users—billions of messages per second. Whatever—with AWS, as long as I'm blindly swiping my credit card, that is not a problem.

I can make that work, right? On the front end, you have so much more nuance, you have so much more complications. And one of the things that AWS Amplify is doing is they're making that front end more accessible to more developers, and they're taking that front end skillset and gently introducing the cloud skill set in addition to it. And I don't think people fully grok how good of an onboarding tool that is because while developers might start with AWS Amplify, a couple of years later, they're not just using Amplify, they're using a slew of AWS services. And when they go out and work in other places, and they want to rapidly prototype something, they're not just going to pick Amplify, they're going to pick all of AWS.

And Elastic Beanstalk did something similar to this back in I would say 2014-ish. It never really lived up to the vision, I would say. I think Elastic Beanstalk is awesome. It's my baby; I love it. But if I'm being perfectly blunt, of course, no, it didn't do what we wanted. And there was this other service, I don't know if you remember, do you remember CodeStar?

Corey: Of course I do. You have to understand, I'm a walking encyclopedia of everything that AWS has ever done. It was sort of their unifying approach to take an opinionated build of, I want to set up a new project. Well, it's going to spin up a code commit repository, it's going to set up a code deploy pipeline, it's going to use code build to wind up building these things. Effectively, it was a… I almost want to use the term Potemkin village of all the build tools that AWS offers, which you only use if you don't have the option of using things that are way better at each of the individual things that they do.

Randall: Yeah, pretty much. And CodeStar was supposed to be this onboarding tool, but what CodeStar focused on was kind of a console experience. They didn't focus on either the command line tool, or the developer experience, or the APIs. So, the only real way to access CodeStar in the beginning, was through the console. So, what Amplify has done is they have you using the AWS Amplify SDK before you ever even create an AWS account.

That's a huge, huge difference in developer experience, especially in onboarding. But even going beyond the beginner side of things, Amplify is able to help experienced AWS developers take pain points with things like Cognito make them much simpler to deal with. And there's also this concept of evolving. So there's this thing called Lightsail that AWS launched a couple years ago to make it easier for developers to go and launch stuff without worrying about specific costs and overruns because they would just charge you a flat monthly fee. Lightsail is great, but again, it's not Heroku.

It's not a command line that you're doing to deploy your app. It was more complicated than that. So, Amplify is one project. I'm a huge fan of it, I can talk at length about it and what I think they're doing right, but they're really focused on developer experience. The other thing that I bet this is one, you'll disagree with me on is CDK.

Corey: Oh, I'm thrilled to have that particular debate with you. In fact, as of the time of this recording, roughly a week ago, I did an article on building a toy app to build out a bot that counts my Twitter followers and writes it to Dynamo as an excuse to play with the CDK. My experience is after an hour, I gave up and went back to SAM CLI.

Randall: [laugh]. So, first of all, I love SAM. Chris Munns and I have spent years working together and presenting together at various conferences. Chris Munns is an amazing speaker and the whole serverless AWS Developer Advocate team is a group of really amazing, and talented, and very, very eloquent storytellers. I think that SAM is really playing catch up to a lot of other frameworks.

So, even [BLANK 00:31:23] is doing more faster than SAM is. And part of that is because SAM is based on CloudFormation. And CloudFormation transforms and all of these other different CloudFormation techniques are being developed in parallel to the things that SAM should be doing. For a long time just doing something as simple as S3 event notifications in SAM was phenomenally difficult. And it only really got addressed, I would say, in 2019. So, it took three or four years for it to really become a solved problem. What I like about CDK is—well, let's take a step back. When you define infrastructure, what do you think about?

Corey: Usually I think of—first I have a workload, usually, that I want to put on an infrastructure that's already built or defined somehow, or some form of code I want out there. I think of infrastructure deployment usually being more of a one and then done approach, as opposed to continuing to iterate on the application that lives in that infrastructure. Now, that's a bit of an outmoded way of thinking in some respects, but every time I push code, I don't necessarily want to reprovision the database that holds the data that code talks to, for instance.

Randall: Gotcha. And if you were to walk back that thought to say, the mid-2000s, how do you think about provisioning hardware?

Corey: In the mid-2000s, my entire approach—because I was provisioning hardware then—was you had to do a lot more capacity planning, there was a six-week lead time, instead of a six-second lead time, there was a lot of building excess capacity in, and you were just getting around to this idea of virtualizing things because it was way faster to spin a new VM—even at slow speeds—than it was to provision, wipe, and reinstall something on bare metal.

Randall: Right.

Corey: So, even that it was starting to get in that direction. Now, I love that everything's an API call away, and you can have more compute power than, like, the entire 20th century humanity.

Randall: Yes. It's amazing to me that a t2 nano or a T3 nano has something like one and a half billion times the compute capacity we took to go to the moon. That's pretty insane. So, thinking about CDK and defining your infrastructure as code: I think defining your infrastructure in markup language, like YAML or JSON, that was a little bit of a foreign concept to a lot of people. But I think code is, one, more expressive, and, two, better suited to cloud deployments.

Now, I have lots of reasons for this, but I do want to issue one point of caution because a lot of times—this is something I've even seen very recently in some folks that I've been working with—there are different kinds of engineers who approach problems in different ways. So, a person who is primarily a software engineer, they will follow—most likely—a principle of DRY. So, don't repeat yourself. And that becomes something that they want to take into their infrastructure deployment and into their continuous integration deployments and things like that as well. My… strong suggestion—and of course, this is not always true, is that ‘don't repeat yourself’ is somewhat the enemy of a lot of infrastructure deployment because when you try to get too clever with CDK, when you try and make everything—oh, let me add one aspect to this one stack and have everything deploy magically with one line of code, you've probably over engineered the problem.

It is perfectly okay to repeat yourself a few times when defining CDK code. I have a cool story about a CDK project, which is what brought me onto it. But originally, I was kind of like you, I was like, super, super skeptical, I didn't really buy into the value of it and then I used it in a real world production project.

Corey: I love the concept, at a high level, of what the CDK offers. And it's better than it was when I first played with it a while back. But there are problems with it. There is on some level, in many cases, a distance between the developer and the infrastructure. That's a different philosophy, and it takes time for companies to get used to that and cultures to shift.

Let's skip past that because eventually, you're right; it's going to be unified. The idea of having all of my infrastructure defined in my codebase for the application means that suddenly I have to integrate CI/CD in a much more meaningful, thoughtful way for all of the application workloads, it means that my code—and the structure of my code projects in this repository—is dictated by how the infrastructure looks. It opens up a number of cans of worms. It does absolutely speed iteration and make this more accessible to developers. That's no small thing. But there is a challenge to that: it requires a different way of thinking, and it's very challenging in my experience to wind up mapping that to something that isn't Greenfield.

Randall: I would agree with that, actually. So, I had tried porting existing pieces of infrastructure using CDK. So, I had a production service that I ran at AWS, that I wanted to move into CDK, and I found it challenging just because I had too much random things going on. And that side of the infrastructure, surprisingly, was the most reliable, but the least iterable, if that makes any sense. You know, it was very reliable as long as you don't touch it. [laugh].

Whereas with the CDK project that I built, so this is—I don't know if you saw Amazon Connect chat. They released a couple of things around re:Invent timeframe, and one of the things that they released is built on top of CDK. So, that project was kind of my baby. It was something that I worked on pretty aggressively and it was a Greenfield project. And I started with CDK because there was a little bit of a push internally to explore it for new use cases and stuff. And I was like you: I was kind of hesitant, I was like, “I don't like this project structure. I don't like any of this.” But over time—that's the same way I felt about DynamoDB in the beginning, to be honest. Do you remember single table design?

Corey: I do, indeed. I remember a lot of use cases where it solves problems, and twice as many use cases where it gets even worse.

Randall: Yeah. And it's a different way of thinking about problems that once you grok it, once you kind of have that mental model in place, you can see the value of it, you can see where it can be applied. So I'm not saying CDK is perfect for everything. What I'm saying is, if you do have this Greenfield project, and you can work around some of the new folder structures and where you're keeping things and how you're thinking about the construction of your stack and your CI/CD, I was able to take a project from literally nothing to deployed in multiple regions, running production workloads over the course of about three hours. Or so in three hours, that one stack was working.

And then as I wanted to implement more things, let's say I wanted to put IAM permission boundaries in there. Traditionally, with CloudFormation, or Terraform, or something like that, I'd have to write a whole section that would go and apply either a transform or something else that would go and apply to all of these different resources that were being provisioned. On the CDK side, I applied one aspect of the stack and those IAM permission boundaries were applied across the board. Let's say I wanted to tag things. This is easier now in CloudFormation than it was, but again, I can have programmatically generated tags as the stack is being deployed.

It can look up what region it's in, it can look up what account it’s in, it can roll up, it can do all kinds of good things with organizations and cost reporting. I've found the speed of iteration with CDK and the reliability of that iteration to be drastically improved over say, traditional CloudFormation, or Terraform, or anything like that. And that was the huge one. And that's why I think CDK is such a valuable project, for Greenfield especially.

Corey: I will meet you in the middle and agree to suspend judgment pending further explorations with it, then.

Randall: Yeah, we should chat about this sometime. I can give you a guided introduction; a no nonsense tour. I think you'd like it a lot.

Corey: You'd be the third person to have done so if we were to go down that path. But I will keep my mind open, and maybe we'll even livestream it for fun.

Randall: That would be fun.

Corey: I want to thank you for taking the time to speak with me now that AWS couldn't keep the fire and the gunpowder keg separate anymore.

Randall: Yeah. And hey, I'll say this. I feel like there's a little bit of a sentiment that I dislike AWS or that I have bad feelings towards them, and that's not the case at all. I'm actually a pretty huge AWS fan. I joined AWS because I was a customer of AWS first.

And I was obsessed with the product. I loved it. I met Jeff Barr in my interview, and he's like, “Okay, well, let's fix this blog post. Here's how you log in to Emacs and stuff.” And that was probably my first month of work there.

And I got to spend the next several years working on a series of really, really exciting projects. But the thing that I loved most about my time in AWS, if I can just take everything in its entirety, the customers that I met and the amount of—just the community. And just going to re:Invent every year. Holy smokes, I actually left AWS for a year and came back, and I made a decision to come back my first day at re:Invent as a customer again. [laugh]. It's a palpable energy, and I was sad to miss that this year.

Corey: I really hope that there's a better story at re:Invent this year than there was last year in terms of getting people together. But I also want to keep that strong online component because it's so much more accessible to folks who can't take a week to fly to Las Vegas, and not do work for that entire time frame, and pay $2,000 for a ticket.

Randall: Yep. [laugh].

Corey: So, combine the best of both, I think.

Randall: I desperately need to get out of my home office and meet people again. And there's another aspect to that is that every demo and every talk I've ever made was built on a train, or a plane, or in the back of a van driving through the floods of Bangkok on my way to my next meetup or something. I need that kind of forcing function of the travel to help me think creatively and build new compelling stories to tell developers.

Corey: Yeah. I think we could absolutely come up with something terrifying in the somewhat near future. More to come on that as it unfolds. Randall, thank you so much for taking the time to speak with me today. If people want to hear more about what you're up to, what ridiculous ideas you have, or mostly just want to see you kick people in the shins, where can they find you?

Randall: I am pretty much just on Twitter. So it's twitter.com/jrhunt is my handle. And I post a lot about AWS, and a lot about machine learning, and occasionally about sci-fi books.

Corey: Excellent. We will of course put links to that in the [show notes 00:42:32]. Thanks so much. I appreciate your time.

Randall: Thanks for having me.

Corey: Randall Hunt, developer advocate at Facebook. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple podcasts, whereas if you hated this podcast, please leave a five-star review on Apple Podcasts or your podcast platform of choice along with a comment that is entirely generated by machine learning.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About David

https://dhh.dk

Links:

  • Basecamp: https://basecamp.com/
  • Hey: https://hey.com/
  • Rework: https://www.amazon.com/Rework-Jason-Fried/dp/0307463745/
  • Remote: https://www.amazon.com/Remote-Office-Required-Jason-Fried/dp/0804137501/
  • It Doesn’t Have to Be Crazy at Work: https://www.amazon.com/Doesnt-Have-Be-Crazy-Work/dp/0062874780/
  • Less is More by Jason Hickel: https://www.amazon.com/Less-More-Degrowth-Will-World-ebook/dp/B085L9XSM1

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the Enterprise (not the starship). On-prem security doesn’t translate well to cloud or multi-cloud environments, and that’s not even counting IoT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IoT devices, detects these threats up to 35 percent faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at extrahop.com/trial.

Corey: If your mean time to WTF for a security alert is more than a minute, it's time to look at Lacework. Lacework will help you get your security act together for everything from compliance service configurations to container app relationships, all without the need for PhDs in AWS to write the rules. If you're building a secure business on AWS with compliance requirements, you don't really have time to choose between antivirus or firewall companies to help you secure your stack. That's why Lacework is built from the ground up for the Cloud: low effort, high visibility and detection. To learn more, visit lacework.com.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’m joined today by a guest who probably needs no introduction, but is going to get one anyway. He’s the CTO of a company called Basecamp. Recently launched Hey, a new way of thinking about email.

He’s the creator of Ruby on Rails, is a Le Mans class-winning race car driver, and of course enjoys one of my personal favorite hobbies, kicking trillion-dollar companies when they deserve it. David, thank you for taking the time to speak with me.

David: Thank you for having me on the show.

Corey: So, you’ve been up to a lot lately. Basecamp is one of those things that has become a staple of online work; your books, Rework and Remote, have been seminal as far as shaping the way people think about remote work. And your most recent one is doing very well as well, which is called It Doesn’t Have to Be Crazy at Work. But you’ve also made the headlines for the past year more or less picking a fight with Apple about their 30% charge on everything that gets sold through their app store, which, for anyone who’s building anything that they expect people to use, isn’t really an optional thing anymore.

David: It’s not, which is actually why I’d flip that statement and say, I didn’t pick a fight; Apple picked a fight. Particularly with us; particularly with the launch of Hey, our new email service. Because I’ve been interested in this as sort of a distant observer, annoyed participant in the software community just, like, “Hey, this doesn’t seem quite right.” But it never felt personal because it wasn’t personal. We got on with our business more or less under the unwritten rules that Apple and Google had set out and did not enter into any direct confrontations on this issue.

All that changed last summer when we launched Hey and Apple showed up with the baseball bats and said, “Ayy, what a nice business you got. Give us 30% or we’re going to smash your kneecaps and burn down your store.” And our response was, “No. You’re not going to do any of those things, and we are going to kick and scream until you back down.” Which was then two weeks of terrible kicking and screaming from my side—from our side, and finally Apple backing down at least on that part of the question, whether we were going to hand over 30% of our revenues.

And after that, I got, um, what’s the word? Radicalized. [laugh]. Or at least I got my attention focused on this, not just being, like, one of many, many issues I care about to one of the top five issues I care about, that I follow intensely and that I’m advocating on all the ways I possibly can to have something happen about. And as it were—today, there is a bill in the Arizona House of Representatives that actually goes up for a vote in a few hours, I think, on the question of whether Apple should be able to force app developers to use their exorbitantly priced payment processing system and whether they can retaliate against developers if they don’t.

So, it’s very much on my mind, it’s very much I think, also in the time right now. This topic has just exploded in the last year. We’ve gone from essentially antitrust being something that weird, wonky policy people were talking about it in other rooms than the ones that I were in to, this is an absolutely mainstream story that’s being followed by outlets left and right, and that consumers are starting to care about because they’re starting to learn what’s actually going on and realizing, “Wait a minute. That doesn’t sound right.” So, that’s one of my—I was going to say big beefs.

It is one of the big issues that I care about intensely, and one of the things that we care about intensely as a company, particularly now that we’re in the market of selling email services to consumers. The funny thing about Hey was that when we launched in the summer of last year, we legitimately thought that the main big tech company we were going to go up against was Google, that Gmail is this gorilla in email. They have more than 55% of all emails in the US, bigger than anyone else, and we thought, like, that’s the battle, right? We’re going to go up against Gmail, we’re going to give an alternative to Gmail, and then we realized, oh, before we can fight that boss, we’ve got to fight this other boss—Apple—to even be allowed in the marketplace. That’s kind of bananas.

Corey: You have to fight the person manning the toll booth in order to get to the castle. Yeah.

David: Yes.

Corey: You’re also somewhat uniquely positioned in that you’re one of the only people who can pick this fight effectively, or at least respond to the fight being picked, in that, sure, you wind up upsetting some small indie app developer, well, they’re going to rant like a loon to 500 followers, and it goes nowhere. You go on the other end of the spectrum, and a lot of companies that are VC-backed are going to be basically muzzled by their investors or by their customers of, “Whoa, whoa, whoa, whoa, whoa. We don’t want to rock the boat. We might want to sell Apple things, so let’s not do that.” And for all that you’ve done and all that you have built a reputation for, you are known as someone who it is not wise to cross, I think is probably the best way to frame that, because you have opinions and you’re not shy about sharing them. And frankly, from where I sit, it’s a wonderful thing.

David: It is indeed also a rare thing that most companies don’t end up in this weird Goldilocks part of the spectrum where, A) we’re not big enough that we have quarterly earnings we have to defend to Wall Street or other investors, and we’re also not small enough that Apple pulling this shit last summer would have, like, wiped us out overnight. So, we’ve been in business for 20 years; we’ve been profitable for 20 years, and if you’re profitable for 20 years, like, that accumulates to give you a buffer, to give you options, and to give you courage. And all of those things really do combine to just leading to this notion of, “What’s the good of having [BLEEP] you money if you never say [BLEEP] you to things that deserve a [BLEEP] you?” And Apple’s monopoly abuses absolutely deserve a [BLEEP] you, so I almost feel like it’s my moral duty and obligation, given the fact that we have the privilege to stand up and speak in a way that most other developers—on both ends of the spectrum: the very rich and the not so rich—can’t do because of who we are. So, if we’re not going to speak up, no one is.

Which is, of course, exactly what’s been happening on the monopoly issue, in particular with Apple and their app store, is that tons of developers have been abused in these ways and they’ve just suffered in silence because no one was interested in getting Apple on their bad side because do you know what? I mean, if you get an Apple’s shitlist, you may not even be able to publish any software at all going forward. Like when Epic, which was perhaps the only company that—in the US, at least—has dared to go head-on with Apple. They tried to destroy their 3D engine.

Corey: Yeah, just pull Unity from the App Store, or threaten to do that to everyone who’s using it.

David: Yeah.

Corey: And if that’s not an antitrust issue, I really struggle to identify what it is.

David: And that’s exactly why people don’t do it. I mean, I think Epic is showing exactly why even huge companies would go, “Whoa, whoa, whoa. We can’t speak up against Apple because they have so many leverage points. They’re a conglomerate and they’re in so many different areas. And they probably have a thumbscrew somewhere they can apply to our business where it’s going to really hurt.”

And that was the thing that was just phenomenal about Epic was that Tim Sweeney and Mark Rein just went, “You know what? We’re going to do it anyway.” This doesn’t actually make business sense. It doesn’t make business sense to go up against that, but let’s just put that out there. It didn’t even make business sense for us.

The odds were so long, and the catastrophe if you get it wrong, would normally do—going up against Apple are so severe that I would absolutely not recommend that anyone else try to do what we’ve done. It’s one of those big stickers, “Don’t try this at home, folks.” And that comes apart, again, from this sense of do you know what? If we’re going to get wiped out, Basecamp is going to get destroyed because we stood up for our principles and what we believe it was right, I can be okay with that because I’ve ended up and we’ve ended up in this position of having had such a wonderful run. And I literally thought about that going into the fight with Apple, just thought, “Do you know what? This might be the end.”

I mean, not perhaps in the literal sense that the company is going to shut down tomorrow, but one of the things we looked into was, we’re going to sue Apple. And then we talked to a bunch of lawyers, and they told us, “Yeah, it’s going to be, like, $25 to $40 million. And it’s going to be about ten years for you to get any sort of permanent relief because it has to go through two rounds of appeals. And so do you want to spend”—

Corey: Bring money and patience.

David: [laugh]. Exactly.

Corey: And a lot more of either than you think you’re going to need.

David: And I mean, we stood on the edge of that, talked to several law firms, almost pulled the trigger on that. And thankfully, we didn’t do it. I’m happy that someone like Epic, who actually employs an entire legal department at their company, could go forth with the fight, but that just shows how lopsided the fight is and why Apple acts the way that they do. Because they know they have this kind of power. They know they have this kind of leverage, even against governments.

I talked about this Arizona bill; before that, there was a North Dakota bill that was very similar, which was essentially, again, let developers use their own payment processing and don’t retaliate against them if they do—

Corey: I was sad to see that get shut down.

David: It did. And what was the justification? Why did the senators not want to vote for this? One senator was quoted from the floor saying not, “I don’t believe this is a good bill,” or, “I believe Apple is amazing.” No, no, no. “We should not”—essentially, I’m paraphrasing here—“Upset big companies because they’re very big and we could get sued.” I mean, holy shit.

When you have [laugh] entire states that are afraid of pissing off Apple because what they might be able to do with them in litigation, you know you’re really in trouble on the antitrust side. So, that just gives me even more fuel for my personal fire that carries this forward and for me to care about it and continue to kick and scream about it because it just seems like this is a really dark turn. Not just for the technology industry, right? Because technology industry is not, like, its own little thing anymore. Technology industries is society. The biggest companies in the world—

Corey: Every company is a tech company, whether they know it or not.

David: Exactly. And this is the slogan, the positive slogan for a while in tech was, “Software’s eating the world.” Right? People were cheering that on. “Woo-hoo. Yeah. Software’s eating the world,” but yeah, who’s eating software? Monopolists. So, once monopolists have eaten the software, they’ve eaten the world. Do we want the world eaten by, like, five companies? Uh, I don’t.

Corey: No. Absolutely not. From my perspective, if I wanted to have to bend the knee to giant companies and do business with people I can’t stand, and keep my mouth shut, and keep on the straight and narrow, there’s a word for that: it’s called having a job. If I’m going to deal with all of that, I may as well go get a job at one of those giant companies where I’ll be unhappy and have to keep my mouth shut all the time, but at least then I don’t have to worry about all the other things that come with running a business. I can let that stuff go.

So, if I’m going to be in this position, where I’m not necessarily accountable in the same way that I am to one person being in a good mood for me to continue to put food on the table, I feel like it’s almost incumbent upon me to speak up and basically speak truth to power. Although someday I hope to do it nearly halfway as effectively as you have.

David: That’s exactly it. Actually, that thought crossed both Jason and my mind independently. We talked about it later, during the heyday of the Hey launch, with Apple breathing down our neck. It’s like, “Do I even then want to make software anymore?” If I have to ask Apple for permission to exist if I have to hand over 30% of the business to Apple, do I need this? Like, I’m 41. Could I retire? Could I do something else? Could I become a farmer? Could I just move somewhere? Both of us thought—seriously—about you know what, maybe it’s over.

Maybe we lived through the glory days of software distribution where all you needed was the internet, and you didn’t have to ask anyone for permission—more or less—and you could just do your own thing. If anything, this whole experience just reminds me of just what a unique and simultaneously amazing place the internet at large is. It is such an aberration in the history of software in terms of who controls the platforms and who dictates what you can and can’t do. And we lived through the best parts of that—perhaps. I mean, I’m not willing to give up hope.

I hope that we’re going to see a resurgence of that, that big tech is actually countered—both legislatively and in the marketplace—and we will see sort of a new dawn of this same kind of freedom that we enjoyed with the internet where we could throw up a website basecamphq.com; 2004; here’s the credit card field if you want to use our software. It’s between you and I. That’s it. No middlemen, no middle tiers, no one telling us whether we can be on the internet or not be on the internet.

Corey: No 600 companies that are included on your website, watching everything I do on there under a magnifying glass, then tracking me across half the internet. You know, all these old-fashioned, legacy ideas.

David: Exactly. And that’s—I’ve become ever more militant about bringing back. I lived through web 1.0. I started with the internet back in 95, and I’ve seen that whole cycle. And it wasn’t all great. Let’s first be fair about that. Nostalgia has a way of erasing everything that was—

Corey: Oh, yes. Those rose-colored glasses get ever rosier with time.

David: Exactly. So, it was not great, but there was a lot of it that really was great. And the ethics and the possibilities really were different, and in many cases, superior. Better. Could we bring some of that back? Why do we have to leave that to the past? I don’t believe that at all.

So, this is one of the things by I care so much about privacy, why we’ve focused so much about privacy in Hey.com, why we focused so much about weird things like speed, or not using too much JavaScript, or just having really light pages, why we are oppose spy pixels, where we are oppose tracking of all kind is, I remember what it used to be, and there’s no reason it can’t be more like that today. Web 1.0 had so many things right that we’ve lost over the years. The sort of containment of the internet that has happened with big platforms has caused us amnesia; we’ve forgotten how it could be. And we’re here to bring that back.

Corey: When you launched Hey, you’ve mentioned previously that this was done on top of cloud. Basecamp is, somewhat famously, primarily hosted in an on-premises environment. What triggered that decision point?

David: Some historical factors. One historical factor was that about, let’s see, four or five years ago, I caught the cloud bug. I caught the argument, which is how so many of these things spread—or the meme, if you will. No one is producing their own power anymore. It’s not like, to run a business you have to construct your own power company and run that in your backyard.

Why would you want to run your own infrastructure? It is like power, it should just be something you buy from the power company. And I thought, “Yeah. Yeah, yeah, yeah. Totally. Makes sense. Let’s not do physical hardware anymore. Let’s just go cloud. That’s great.” And then, I mean, things happen, right? The awareness of big tech. The awareness that cloud isn’t some magic thing. It’s just other people’s computers that you rent. And a lot of the promises of cloud, they don’t actually pan out, or didn’t pan out for a lot of companies, including ours. Is cloud simpler? No, not really.

Corey: I don’t think anyone has ever accused it of that.

David: It was part of the pitch in the beginning. And I understand there's a certain context in which it is, but when you run at our scale in something like, Hey or whatever, it absolutely is not. We run both things, we run the bulk of Basecamp on-prem and we run Hey on-cloud. And on-cloud is not simpler in a total perspective. There are aspects of it that are simpler, but by and large, not simple.

It’s just other people’s computers, and you’re renting those computers. Then that comes to other things such as economics. Do you know what? If you’re renting other people’s computers by the hour, you’re going to pay a per-hour rate to rent those computers. Is that rate—

Corey: For the rest of your life because they never automatically turn off. Or as in the case of a data center, get carried it off at some point by the raccoons because you failed to maintain the data center in any reasonable way.

David: Exactly. So, that makes sense for certain things, especially speculative ventures where you’re like, I don’t want to plop down, I don’t know, hundreds of thousands of dollars buying all sorts of equipment. What if my idea doesn’t work? Then I end up with hundreds of thousands of dollars in computers that I don’t know what to do with. That is a waste of money.

And I totally agree. The cloud does that well. It does well, the, “Let’s just see if there’s anything here.” But then the moment that there is something there, the economics totally change. And they’re totally shit, particularly for companies like us where we have a long horizon. We’ve been in business for 20 years. For me to depreciate a piece of hardware over 3, 5, 7, 10 years—I think we actually do have some hardware that’s 10 years old. Not a lot, but some—just completely changes the economics.

Of course it’s going to be monumentally cheaper, in a lot of instances, to buy your own stuff and keep it for five years. Particularly when you come out to the edges of what cloud is good at or designed for. For us, that’s big databases. You can get some really big databases now, on AWS, I’m sure you’re aware. But they are phenomenally expensive.

Some of the calculations we have are, like—for some of the largest databases that we’ve been looking at, the cost to rent them on a per-month basis is something like, in three months we would have bought it. Sometimes even less than that. And you just go, “Wait a minute. That really doesn’t add up.” And then that doesn’t even include the fact that you don’t actually have access to, sort of, the tall end of the spectrum.

And for us, we’re trying to keep our operations as simple as possible, and that includes distributing your computing the least amount possible. And for databases, it means let’s not take on the burden of premature sharding, premature federation, this that and the other thing. Which means we run some pretty big, beefy databases. We’re trying really hard to pretend we just have one database and we can keep using that forever, and that has kind of been the truth for almost 20 years with Basecamp. As early back as 2007, I thought, “All right, we’re going to hit the wall here. We’re going to run up again—we’ve got a shard.” Right?

I wrote it up at the time, I think there’s a blog post on this being from 2008, about how we starved off the sharding thing because we bought a slightly bigger machine. And then things just kept bailing us out, so to speak. Then SSDs came, then all of a sudden you could have a terabyte of RAM, then all these other things changed to the point that for Basecamp, it’s a completely non-issue. Like, technology got bigger, faster, better, quicker than our growth rate. So, that’s not a problem. On Hey, it's a little different. [laugh]. Email is surprisingly resource-intensive. And—

Corey: Oh, yes. And mentally overhead-intensive, too. Is sort of the entire reason you built the product in the first place.

David: Exactly. It’s just a massive amount of data, in a way that the data we accumulate and accrue for Basecamp just isn’t. So, we’re facing some of these issues again and we’re facing those cost constraints again, including the physical constraints of what kind of hardware is available, and now we’re also additionally facing the constraint—as we started to talk about—is big tech. Who do you rent these computers from? Are these companies—

Corey: Which of your favorite three big tech companies are you going to go with? Or are you going to pick something sketchy and dangerous on some level?

David: Exactly.

Corey: And there are things in between. For example, DigitalOcean and Linode are both not unreasonable options, and they, in many cases, are a lot more sensible. But at some point, you’re picking which devil you’re crawling into bed with. And I don’t know that there’s necessarily a way around that in many contexts.

David: And that was my conclusion for the longest time. Do you know what? Just in part, hold your nose, and that’s what it is, right? We’re going to use—and we’ve used both Google’s and Amazon’s cloud offerings. We actually used to use Google a whole lot, and then we had some traumatic, terrible, no good, very bad experiences with that, that really just burned us severely, and vowed we’ll never run anything on Google again in that capacity.

And we moved all our stuff off, and we moved it to Amazon. But it’s not exactly like Amazon is wonderful in the scheme of big tech. So, the conclusion I’ve reached is that, why do we even have to have a devil in our bed? I understand there’s some advantages, and maybe we could take those advantages with us. Kubernetes, which is taking a hell of a lot of a beating as a super complex, convoluted beast, and not unfairly so, but at a certain scale, that complexity is not a showstopper.

We have an ops department today of—what—nine people? Eight people. Nine people, full-time people who work on this stuff. That’s a thing we could do. That’s a thing we could do: we could run our own Kubernetes setup. And we can also run our own servers; we’ve done that for 20 years.

Can we compare those two things or merge those two things such that we get a cloud-like experience, but on our own hardware, without being in the bed with some devil? Yeah. I actually think we could. And that’s the trajectory that we’re currently on. I would like to see less devil-based cloud computing and more learning and running our own clouds.

And I think that other companies that are in this mid-size range, may very well come to the same conclusions. It’s really interesting. It’s almost like the who can afford to speak up thing. Again, you got to be in the Goldilocks phase. When you’re first starting out, you don’t have full-time ops people, you don’t know if this idea is even viable; cloud is a really interesting, cheap way to validate your idea.

And then perhaps when you get to, sort of, all the way to the other end of the spectrum and you’re a billion-dollar company, cost just doesn’t matter. Like, it doesn’t matter whether it cost you $1 billion or $2 billion to run this stuff. There’s just other concerns that matter to you. Well, we’re in the middle phase where cost absolutely matters. We’re spending my money; Jason’s money.

And I care about how we spend it, and could we spend it on other things? And then we also have the capacities in-house to do some of these things. So, it’s interesting. We are in that middle phase where we’re trying to figure out what is the path forward? For example, do we want to host our own files again? I was just talking with Troy, our head of ops, today about this question. Do you know what? S3 really makes it easy to—

Corey: It’s dirt cheap, at least compared to a lot of the other services you can pick.

David: Yes.

Corey: The durability guarantees are phenomenal. It’s getting increasingly intelligent, so you can start staging things from an economic point of view. Heck, look at Glacier Deep Archive: it winds up costing roughly $1,000 a month to store a petabyte of data, which gets us into the realm of why delete things ever again because you might theoretically need it, and in the context of business, $1,000 a month is peanuts. And that leads to some weird behaviors.

David: It does. And first, I’d say S3 is the one service from the whole cloud thing where, like, yeah, actually we can’t do it cheaper. Even if I’m willing to go on a ten-year depreciation thing, the setups that we’ve at least had so far and the benchmarking we’ve done, we can’t beat the pricing. Which is completely different than Hey, the beefiest database server, what does that cost? It’s not even remotely competitive on a pricing basis if you’re willing to buy hardware and depreciate it over the years of service, but storage absolutely is.

But that’s run into this other thing where the storage is no longer physical. It’s not even an inconvenience. It’s a line item on an invoice. That’s it. But you know what? That’s, again, a mirage. This storage isn’t in the cloud; it’s just on other people’s computers. And those computers need power, they need operations, they need cooling, they need all the things that go into having computers that are turned on all the time. And that creates waste. And do you know what—particularly with email where we’ve just seen our data usages growing phenomenally fast. I mean, orders of magnitude faster than anything we’ve ever done in the past because email just takes up a lot and people get a lot of email.

Corey: Yeah. Please go ahead and send me two megabytes of images embedded in the email for a marketing email that I, in the best case, will glance at once. And then where’s it going to go? Oh, it’s going to sit in my inbox for the rest of my life. In fact, I’m sure that my great-great-grandchildren will pass down one of their ancestor’s old email inboxes and go through it. Who are we kidding?

David: Exactly. This was exactly the experience I first got when I exported all my email from Gmail. So, I’ve been using Gmail for like, 12, 14 years? Whatever. I’ve used Gmail since it started until, what is that, like, a year and a half ago when we switched over to Hey.

And I exported all that and I got that whole archive, and I think it was like 40 gigabytes. And then I loaded it up in my local mail app—it’s just an inbox file. They’re quite easy actually to export and then use locally—and I started looking through what was in it. And I was like, “What am I, [BLEEP] packrat horder? Why am I saving old newsletters from 2009?” If any of this stuff was physical, I’d see the insanity of just saving this stuff forever.

Corey: Oh, my God. I’m the digital equivalent of a hoarder? What’s going on?

David: Exactly. And that was exactly what I was thinking. I was just like, “Oh, man, this doesn’t make any se—like, what is going on here?” Those 40 gigabytes combined with the other 2 billion users who each have their own, if not 40 gigabyte, but then many gigabytes, like, what is it cost the world in terms of energy, in terms of waste and pollution, to keep all that storage live, searchable, whatever? That is not insignificant.

I saw one estimate saying that it costs the equivalent of 1.9 billion trees a year in terms of just spam. And this is not spam. The 40 gigabytes I had saved, there’s no spam in that. That was just email that was sent to me, and at the time I accepted in my inbox, and then I just didn’t delete it, and I kept it. This didn’t make sense.

So, this brings us back to the whole Hey thing and us essentially having solved the storage problem to a place where it’s invisible, and I’m not at all sure that’s good. In some regards, I kind of wished that storage was painful again. It’s just that then I would at least think about, like, does this even make sense? The same report that quoted the 1.5—or 9 billion trees, said something about, like, “On a bunch of studies-blah, blah, blah—90% of all data that we keep is never accessed again, after”—I think 30 days or something like that. That just seems like a tremendous, horrendous, terrible, no good use of space.

Corey: Some people need dozens of tools to visualize their stack. But not you. You’ve got New Relic Navigator.All your hosts, services, containers, and anything else you can monitor are in one dense hex view, so you can check system-wide health at a glance.

And you can sort by platform, service, app, or any other tag, which makes it easy to navigate yourentire system and see what needs your attention.Go to https://newrelic.com and get started for free

Corey: And there are occasions where you need that sort of thing. For example, one of the things I did about 10 years ago is I was the first ops person that Expensify and the pattern that we identified pretty quickly is that people would upload receipts and very often never look at the receipt again. But God help them if someday they needed one of those receipts and didn’t have it. That is not this. This is, “Oh, wow. If I never get that important marketing message”—that, again, is festooned with trackers and the rest—“I will be a happier person.”

But the world has gone a different direction. I quick check online, because I haven’t looked into this space in a while, I didn’t realize this. They sell 16 terabyte hard drives now for 300 bucks. And it’s, “Oh, well, yeah, my email archive is growing, but that’s okay because drives are getting bigger faster than I’m collecting digital crap.” But it’s an arms race. And it becomes this enormous untrackable mess at some point where I can’t even begin to contextualize how much data I might have hanging around on my drive. And it’s all for no purpose.

David: Totally. And then if you look at it at an individual scale, like, “Oh, how much data do I use? Is that actually a big deal?” Maybe not. But we see it at a much larger scale. We have tens of thousands of paying customers for Hey, and already at that level, it’s real. And it’s not just another hard drive. It’s multiple computers. We just added three new nodes to our Elasticsearch setup.

Because that’s the other thing; an email that comes in, it’s not stored once. It’s stored, like, 500 different times. It’s got to be stored once for backup, once for distribution, once for resilience, once for search, ba-da-da-da-da-da-da-da, right? This data just multiplies in a way that when you then add it all up, it’s actually real machines. And for what?

If it’s serving a purpose, if it’s, like, oh, it’s giving you an archive for these invoices, and you get audited in, like, five years, and you really need to prove it, where, fine. Totally. Good. Very nice. In fact, just having all that stuff around is really helpful for that. But the myth that Gmail, in particular, is responsible for sort of putting out there’s it’s like, “No, no, no. There’s no cost. Store whatever, for however long, and everything is fine.” And then you go back and think, “Wow, man. Wow, that’s such a generous offer. Why would they do that?”

And you think, “What is Gmail in the business of doing?” Of training ever more sophisticated AI models to figuring out how they can organize and strip you of all your personal data, build ad profiles on users so that they can more efficiently sell you a pair of [BLEEP] sneakers. Okay. Well, I mean, maybe in that perspective it’s not that great that I’m giving Gmail in particular, like, 12 years of history of who I am, what I buy—

Corey: I really want to show Google of all companies, exactly how I evolve as a person—

David: [laugh]. Exactly.

Corey: —over time. Sure that won’t come back and bite me in any meaningful way.

David: Right. And I mean, that’s the thing. When you think about it an abstract terms, you go, “Yeah, weeell, I mean, ehh… what are they going to do, right?” And then you learn that Gmail literally does what Expensify did—except just, like, under the cover of darkness, almost—that they go in, they look through all your purchasing receipts, they mine that for data, and then they use it to build a profile of you, of what you bought, when, and so on and so forth. And you go, like, “Okay, that’s pretty [BLEEP] creepy.”

Corey: Oh, they had a page, ‘slash purchases’ I think it was, that they unveiled—like, they were super proud of this—that showed that of every receipt that had ever interpreted through your Gmail, and they talk about this like, “Look at this. We can show you all the stuff that you bought.”

David: Right.

Corey: And the response was, “You’re [BLEEP]-ing what?” And that became a big kerfuffle. And now they don’t talk about it very much, but it’s still there because of course, it’s still there. Why would they deprecate something horrible when they can deprecate something useful and beloved? But that’s the Google ethos in some respects, as the big Google company that it is.

David: Yes, once, quote-unquote, ‘normal people’ who are not, sort of, tech folks or marketing people follow all this stuff in detail, they learn about some of these things, they go, like, “They’re doing what?” This was what we just encountered with spy pixels. So, Hey tracks, shames, and names when someone sends you an email that has spy pixels in it, those little one-by-one pixels that you can’t see that sends data back about whether you opened the mail, how often you opened the email, where you were when you opened the email, what time, what device you used, all the personal information that’s sent back that techies in large regard—like, they’ve known this, right? They’ve know—this thing that happens. You explain this to normal people, and they go, “Wait, what? I opened the email and they can see that?”

Corey: “You know what about me?” Yeah. Oh, yeah.

David: Exactly. They’re like, “This is n—no. No, no, no, no, no, no, no, no, no. This has got to stop.”

Corey: And as someone who sends a newsletter, it turns out it is incredibly challenging to find an ESP—or email service provider—that even thought about letting you disable this. It’s creepy. And it’s just become so accepted that when you ask, “Oh, can I disable that?” The person you’re talking to at these companies looks at you like you’re simple. It’s, “Why on earth would you want to do that?” “Oh, I don’t know, I have this ancient belief, like you know, respecting my readers.” What a concept. It’s awful.

David: It really is. And I think this is where some of the practices of the industry, it just seeps into the assumptions, it seeps into the ideology that, like, the more data you can extract, the better. Full stop. The more data you can store on someone, the better. Full stop.

Not thinking about, do you—does that person want their data stored in this way? It’s not even part of the conversation. And I think that’s one of the things I’m happy for us to do Hey, not just because I wanted to, quote-unquote, ‘solve email forever.’ Since I’ve been an email user forever, and I’m sort of selfish in that way, and I build software for myself, [laugh] and the majority of features I’ve worked on in Hey are things I wanted email to do forever. But also because it allows us to have this conversation and start these conversations.

And spy pixels is a great example because I’ve been talking about this for a while now, and people just go, “Well, I mean, it kind of doesn’t matter. We’ve been doing this forever.” Yeah, but if the people you’re doing it to are just learning about it now, and they’re realizing they don’t want this done, perhaps it’s time for you to, like, stop doing that. And I hope that we can grow Hey to an extent that it becomes sort of meaningful enough that it can help shift the conversation. Which is, of course, already happening.

Corey: There has to be a shift.

David: Yes.

Corey: If I’m the only newsletter out there that tells sponsors, “Oh, we have no idea about any metrics of this,” they’re going to go elsewhere. There has to be something of a groundswell approach. The philosophy I’ve always held that seems to serve me reasonably well is, anything that I do, I should be able to take any random representative user and take them through everything that I’m doing, and they should come away and look at it and say, “Yeah, that makes perfect sense.” Take this podcast, for example. I have numbers of people who download this—because I have server logs for how many people have downloaded the file—and I can roughly geo-correlate that by continent. And that’s good enough.

Every once in a blue moon, I will have a sponsor or prospective sponsor ask if, “Oh, can we embed a tracking pixel in the RSS feed?” And our answer is a very polite form of, “No, you the [BLEEP] may not.” If you need that in order to sell your product and figure out how to move your wares and grow your company, maybe you aren’t that good of a company, to be very direct.

David: Yes. This is one of the arguments I keep hearing for targeted advertising in general. There’s like, “Oh, there’s all these businesses that couldn’t exist unless they could totally pinpoint target”—do you know what, maybe they just shouldn’t then. Right? First of all, this is completely ahistorical.

Targeted advertising is, in the grand, long scheme of the history of commerce, a blink of an eye. It’s existed for what, 15 years at the very, very most. It’s been super prevalent for maybe 10-plus years, a little more than that. Businesses sold their stuff before there was targeted advertisement. In fact, many businesses still sell stuff.

Corey: A billboard that you look at doesn’t look at you. I hope. Although I’m sure that there’s some VC-backed startup trying to change that.

David: Yeah. I mean, all sorts of Bluetooth tracking and all sorts of things going on. But yes, exactly. There are all these other forms of advertisement. And the reason that people feel like they kind of have to do this is in part because so much of the conversation and attention has been captured by these platforms—whether that’s Facebook, or Instagram, or Twitter—or anything else, that they have such a huge slice of the attention now, that take all that attention, be able to slice it down into, “Oh, this person is between 31 and 32 years old and just got engaged to someone who’s pregnant.” We should show them ads for prams, right?

You go, “Okay. Like, that’s a relevant ad, but actually, I didn’t even tell my parents that we’re pregnant yet. Why do you know this already?” And then you realize, oh, actually, Google knows a whole lot of things based on your search history, and they use that data, and they’re paired with the receipts that you have in your Gmail account. And all of a sudden, yeah, they know that the couple is pregnant before they’ve told anyone else.

And you go, like, “Actually, no. That’s not a world that I’m all that interested in living in. I’d like to have some privacy. I’d like to sort of control who knows that.” And then the argument often goes, “No, no, no. It’s just targeted advertisement. I mean, we’re not selling the information. What the hell are you talking about?”

Corey: Well, how else are we supposed to get in front of people—

David: Right. [laugh].

Corey: —who are interested in cloud computing? Oh, I don’t know maybe find publications that cater directly to that audience and put two and two together. Doesn’t require a whole lot of tracking to realize that if that’s the subject of a media publication, the people who are in the audience for that probably care about the thing.

David: Bingo.

Corey: It’s such a strange thing.

David: Well, it’s such a profitable thing, I think. That’s the problem.

Corey: Ugh.

David: The problem is that, like, exactly the scenario you described, where sort of niche—or not even that niche topic areas whether it’s a newspaper, or it’s a magazine, or it’s a blog, or it’s a podcast, or whatever, they would have a real advantage. They’d be like, “Okay, I mean. Talk about cloud on this podcast.” That’d be a good place to advertise your cloud-monitoring wares, or bug trackers, or whatever, whatever, right? You just, you don’t need that anymore now that you have targeted advertising.

You can just as well show that ad against the most vile, disgusting content on the internet because the content no longer matters, only the attention does because we can pinpoint who’s looking at that. And then all of a sudden, all quality content loses most of its appeal and becomes way harder to make a business producing that kind of content. Because all content is now equal. All eyeball time spent is equal when you can slice and dice with targeted advertisement, exactly who should see what. And that’s just led to a collapse of local newspapers, all sorts of media institutions are completely scrambling. While we have a handful of organizations—primarily for this case, Google and Facebook—just surging into trillion-dollar market caps, posting ever higher results capturing ever more of online advertising. And you go, like—

Corey: And you need double-digit growth—

David: Exactly.

Corey: —year-over-year at all times. And at some point, you’re looking at this and they’re smacking into population limits, and the only way to continue that growth trajectory is to leverage what you’ve already built and expand into horrifying new markets that no one wants someone with your level of access in. Or else you’re looking at doing horrifying things that fundamentally betray user and customer trust.

David: Yes. This is the trap. Sounds so benevolent, but I’ll use it because I think more stringent terms for it might turn off your listeners. But this is the trap of capitalism. The trap of having to post growth year over year over a year.

And once you get to a certain size, particularly when you’re a trillion-dollar company, if you just have to post 5, 10 percent growth every year, you have to find $100 billion worth of new business or value in new business every year. And then the year after that, you have to find out another a hundred—and now ten—billion dollars. This becomes very hard to keep up on the ethical side of the road. Sooner or later you’re going to start realizing, “Shit we’ve run out of all the ethical, good for society ways of making money. Now, we got to just start what’s beneath that.”

And that goes for, I think, both Apple as we’ve seen with their pivot to services once they ran out of people on earth to sell a new $1,000 iPhone to every year, they started realizing, “Oh, shit. I mean, how are we going to get this up? Oh, services. If we just take all the people who bought an iPhone and we make a bunch more money on them.” And then, of course, to spin it all the way back to the start, “Hey, if we could just take 30% of everyone else’s business, oh, that’d be great. We don’t even have to do anything. We don’t even have to come up with any new”—

Corey: Amazon has expanded to advertising to the point now where I go to amazon.com and search for anything, the results are dogshit. They are unusably bad, where it’s not at all clear which ones are sponsored, which ones aren’t. Oh, ‘Amazon’s Choice?’ Turns out that algorithm-powered. Reviews are getting gamed constantly. I no longer have trust that whatever it is I’m going to buy is actually going to be any good.

I’ve had a few bad counterfeit experiences by getting things, and sure they’re quick to do refunds and all, but you feel taken advantage of. You feel like they have abused your trust. And that’s a very hard thing to get back. I am increasingly buying a lot of things not on Amazon. I was an Amazon household for years because it was so convenient and so easy. And now I’m looking at this going, “I don’t want to give these bastards my money.”

David: Right there with you. So, we used to be, like, “Oh, man. Prime is the best thing and you get things so quickly.” In the past, I think it’s been almost two years now, we essentially just stopped using Amazon. And part of that is powered by where I don’t like what Amazon is doing. I think the product, the service—

Corey: Oh, it’ll be there in two hours or—basically if it’s not there with an hour or two, they’re going to be extremely apologetic. Yeah. Okay, cool. I love convenience. Who doesn’t? But at the cost of what?

David: Right.

Corey: Horse-whipping someone who is a contractor, not an employee?

David: Exactly. Yes. Yes.

Corey: You can’t treat employees that way.

David: Right.

Corey: It’s awful. This is not good for society.

David: It really isn’t. It’s not only not good for society, but even the product experience of using the service is getting worse. As you say, you can get counterfeit wares, you can’t trust reviews anymore, you can’t even trust the listings because they’ve been bought. The same thing with Google. Google is realizing that to keep juicing their returns on search, they have to make the search worse, they have to cram the search full of confusing elements of what’s an ad and what’s a not ad, that they take over entire areas of search, “Oh, travel, we’re not going to bother with this linking to the internet business anymore. We’re just going to show our own shit. And that’s how we capture an entire market.”

All the products are just getting worse because there are these imperatives to keep posting growth. And you think, like, “You know what? We’re literally making things worse. Isn’t the future supposed to be better?” And I’m currently looking at the extrapolation of these companies keeping up the growth rates they currently have, and then think about where are they in 10 years? Things are pretty [BLEEP] bad now. Where are they in 10 years?

Tha—tha—no. The compound growth requirements needed to keep up, they’re going to take them down a very dark path unless we start putting up some legislative guard rails and saying, like, do you know what? You just can’t do that. You can’t just buy all your competitors and then kill them because then they might be a threat. You can’t just bully and abuse all your smaller vendors, or developers, or whatever business partners you have on your platform.

That’s just no good. You can’t just both own all the platforms and operate in all the platforms, self-prefacing all your services at the same time. And this is how so many of these conversations lead back to, we have a monopoly problem. We have a big tech company problem, and we should do something about that. And if you had asked me about a year ago, two years ago, “Is that going to happen?” I’d say, “Phuh, I haven’t seen any containment at all on big tech companies since Microsoft was sued by the DOJ in, what, 2000.” It’s literally been 20 years since there’s been any antitrust activity.

And then boom. Right? Boom. In the past year, there’s just been a huge amount of investigations, and inquiries, and hearings on multiple continents, and multiple states, and from multiple parties. And you just go, “Oh. Actually, okay, yeah. Maybe things could change.” I’m a little less gloomy than I was, like, two years ago.

And I think perhaps it is possible for democracy and societies to wrestle back control. Because we’ve faced these issues before. Big tech is not exactly the first set of monopolies to come along in the history of, let’s just say, the U.S. Battled with telecoms, battled with railroads, battled with Standard Oil, battled with the banking trusts. There’s been a bunch of big fights on this. So, it’s in the DNA of the U.S. In particular, to eventually go, all right, we’re going to do something about it. It’s way too late; we should [laugh] have handled this 10 years ago, it would have been way easier if we’d done it then. But okay, now, enough is enough. You’re not going to push around anymore. Here we come.

Corey: Ugh. As much as I want to keep going. I feel like, at some point, I’m already going to get letters from friends and listeners who are at Amazon, and frankly, please send those letters. I love debating these things with folks. But I feel like at this point, our time is really drawing to a close. Any parting thoughts, concepts, things you wish the world saw differently than it does?

David: On particularly this notion of capitalism’s issues with compounding and never-ending growth, I just read a wonderful book by Jason Hickel called Less is More that I would seriously recommend to anyone where this is sort of uncomfortable territory, or maybe you work at a big tech company like this, and you’re like, “Well, yeah, maybe there’s some sort of abstract ideas, or… I don’t know.” Read the book. It’s great, and it discusses both what we talked about with the ecological challenges we’re facing with this never-ending growth, with never-ending, all the storage that we talked about. And then the issues with capitalism, and particularly as it relates to big tech companies. He doesn’t really go into monopolies, but many of the lessons and critiques apply directly.

And I think it would really help someone open up their eyes or blinders of ideology just as I’ve opened up my eyes and blinders of ideology. I didn’t exactly come from, I don’t know, a Red Camp somewhere. I’m a graduate of the Copenhagen Business School. I got indoctrinated very well in [laugh] traditional tools, tactics, and techniques for how to juice growth, and got just a fundamental worldview that was very in line with The Economist and never—unending growth and how the GDP, as long as it keeps going up, the world keeps getting better. And it’s taken me a long time to deprogram from all that, but I still have all it with me.

So, I think if I can go through that journey, and I can open my eyes a bit and realize in which ways things aren’t always wonderful, other people certainly can, too. So, that book is just a great starting point. I’ve just been reading it now. Huge fan of Jason Hickel’s work in general. So, there goes my recommendation. You could even buy it on Amazon.

Corey: We will, of course, include a link to this in the show notes. It is not an affiliate link because, frankly, I don’t feel the need to start doing those games. But it’s—yeah, there’s a lot to the idea of thinking critically about the things we use and the places we work, and if that makes us uncomfortable, well, that’s where growth comes from. You’re not exactly a barefoot hippie in the park saying these things. You understand what it takes to run a business because you’ve done it multiple times now.

You have a perspective that I think the world desperately needs to hear. And of course, we will include your full-on bio and links to your books and whatnot in our show notes as well. Thank you so much for taking the time to speak with me today. It is a rare privilege to get to pick your brain directly like this.

David: It was my pleasure. All my favorite topics, stirred up good. It’s a—that’s a good time. Thanks for having me on.

Corey: Thank you, David Heinemeier Hansson—or DHH—CTO and founder of Basecamp and Hey, creator of Ruby on Rails, et cetera, et cetera, et cetera. Read the show notes.

I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice and an insulting comment that is packed jam full of every spyware tracker you can shove into the thing.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Eric

https://aws.amazon.com/blogs/security/aws-security-profiles-eric-brandwine-vp-and-distinguished-engineer/

Links:

  • Twitter: https://twitter.com/ebrandwine
  • AWS Security Blog: https://aws.amazon.com/blogs/security/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is brought to you in part by our friends at FireHydrant where they want to help you master the mayhem. What does that mean? Well, they’re an incident management platform founded by SREs who couldn’t find the tools they wanted, so they built one. Sounds easy enough. No one’s ever tried that before. Except they’re good at it. Their platform allows teams to create consistency for the entire incident response lifecycle so that your team can focus on fighting fires faster. From alert handoff to retrospectives and everything in between, things like, you know, tracking, communicating, reporting: all the stuff no one cares about. FireHydrant will automate processes for you, so you can focus on resolution. Visit firehydrant.io to get your team started today, and tell them I sent you because I love watching people wince in pain.

Corey: This episode is sponsored in part by ChaosSearch. As basically everyone knows, trying to do log analytics at scale with an ELK stack is expensive, unstable, time-sucking, demeaning, and just basically all-around horrible. So why are you still doing it—or even thinking about it—when there’s ChaosSearch? ChaosSearch is a fully managed scalable log analysis service that lets you add new workloads in minutes, and easily retain weeks, months, or years of data. With ChaosSearch you store, connect, and analyze and you’re done. The data lives and stays within your S3 buckets, which means no managing servers, no data movement, and you can save up to 80 percent versus running an ELK stack the old-fashioned way. It’s why companies like Equifax, HubSpot, Klarna, Alert Logic, and many more have all turned to ChaosSearch. So if you’re tired of your ELK stacks falling over before it suffers, or of having your log analytics data retention squeezed by the cost, then try ChaosSearch today and tell them I sent you. To learn more, visit chaossearch.io.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Eric Brandwine, who's a Distinguished Engineer and VP at AWS. Eric, welcome to the show.

Eric: Hi, Corey. Thanks for having me.

Corey: So, what is it you actually do at AWS? Every time I've mentioned your name to folks in passing, they get this sort of stricken look, and all I can assume is that you're basically Darth Vader.

Eric: Darth Vader, I think is a slightly unfair characterization; perhaps not wholly unfair—

Corey: Because he had a redemption arc?

Eric: Well, there were only three movies. And—

Corey: Exactly. If they'd made a prequel or sequel, it would have probably been a really good movie. Shame they never did.

Eric: There were three Indiana Jones movies, there were three Star Wars movies. And yeah, at the end of the third movie, it kind of sort of worked its way out. Every year at re:Invent, Andy Jassy gets up on stage and he says, “Security is job zero.” And I love it when he says this because one, he's counting from zero, which is how all good computer scientists count, but two, he's very publicly saying how important security is to what we do. And this isn't just Andy getting up on stage at re:Invent; when he comes back to Seattle, this is the behavior that he models for his leaders.

And so, unfortunately, my first interaction with a lot of our employees is during a security event. And so I have met a good number of my coworkers via SEV 2 tickets—SEV 2s or pager tickets. And it's not the best way to make friends and influence people. And so, you've got to put a lot of work into building those relationships and making sure you reach out and contact people after the dust has settled. But I would say that my primary job is making sure that we hold the security bar high, and relentlessly so.

Corey: The idea of security being job zero, when I first heard of that—my instinctive reaction was, “Okay, we made a list of all the things we have to do, oh, crap, we forgot security, so we'll do what everyone does with security bolted it on at the end and put at the top so we don't have to renumber anything.” And it's a funny joke. And it's great to make the cheap shots and whatnot, but let's be very clear here; it is blindingly apparent to everyone who has used AWS in depth that security is baked in. You cannot bolt it on after the fact and expect to see the level of success that AWS has in a security perspective. I want to be explicitly clear on this. I have a laundry list of grievances around AWS—mostly around service naming—but I have never had a problem with how seriously you folks take security.

Eric: That is excellent to hear. I'm glad that that's coming through.

Corey: It's a no-win game when it comes to security because either it's in the way or it's invisible—if it's done well—but no one ever stops—“I really like the security there.” In fact, the only time people really seem to talk about it is after they've had a data breach in public, and they're saying, “We take security seriously,” right after it was exquisitely clear that they did not take security seriously.

Eric: Well, I've heard that perspective many times, and I disagree with it. Because—you're unsure of your security; when you're surrounded by ambiguity, it's deeply unsettling, and the effect of that, the materialization of that in the business is friction and low velocity. And when you work with someone, whether it's one of our service teams or one of our customers, and you give them the data that they need, and the tools to manage that data so that they can understand the security risks that they're facing, and so that they can make data-driven, informed decisions about how quickly they want to move, about which risks they need to mitigate, about which risks they can accept, they become way more comfortable, and that leads to greater velocity for the business, it leads to greater confidence for the leadership, and it leads to greater delivery to customers, which is the reason that the business is there. And so, over my time at AWS, I've seen security go from something that nobody talked about, to something that could only be a deficit, to something that is actually an enabler for us and for our customers.

Corey: I think you were the first people I've spoken to in my life who ever pushed back on the idea of, “Well, obviously, security wasn't the right answer here,” yadda, yadda, yadda. No, I think you're onto something, though. It's always a spectrum between usability and security, and there are trade-offs that have to get made. And on some level, given what AWS is and who your customers are, you can't ever get it wrong in a serious, big way. Because you don't get a second bite at that apple.

Every Cloud-doubter on the planet is going to come back with saying, “See, see. I told you, I told you.” And that's kind of weird. It's a very high-risk story, and so far you've delivered. There haven't been these horrifying nightmare things that on-prem sysadmin, grumpy type—of which I used to be one—was long predicting. It's a track record for what I can only imagine to be a thorough defense, in-depth position. What am I missing?

Eric: Absolutely. It's something we spend a tremendous amount of time and energy on; it does not happen by accident, and we're constantly looking for the maximum leverage we can get out of any defensive mechanism. But I don't think that the world is as boolean as you're describing here. When security events happen, they're absolutely serious. We take them very seriously. We respond immediately.

But if you look at the things that have happened in the world at large, even some of the large newsmaking security issues that we've had recently, it's never a complete company extinction. It's never the end of a line of business. It's definitely disruptive to roadmaps, it is damaging to customer trust, it's not something to be taken casually but it's not like you make a single misstep and it's all over. And I think it's really important to reinforce that. I'm regularly humbled by the amount of customer trust that we've earned.

And I don't take that lightly, I'm not going to play casually with it, but if you're caught up in the belief that a single misstep is going to lead to business extinction, then you're going to be paralyzed, you're going to be unable to move forward, you're going to be unable to objectively consider the risks. And I see security as highly parallel to availability. Availability is something that every service provider thinks about and has deep experience in, and we all think about the risks that we face. It is possible to build a system that is incredibly resilient: you run it in multiple availability zones, you run it in multiple regions, you build two completely separate implementations of it using two different languages and two different runtimes, you completely don't share a [fate 00:09:08] across anything, and you can build an incredibly robust system. Almost no one does that because it's expensive.

And the business either implicitly or explicitly is making decisions about how much they're willing to spend on availability, and which availability risks they're willing to take. And all of these services have some availability risks and sometimes they have availability events. And security is exactly the same. We are surrounded by security risks, no businesses without security risks. And so the way to succeed here is to, as objectively as possible, think about those risks, mitigate the ones that aren't acceptable, prepare to mitigate the ones that are acceptable if it turns out that your analysis was wrong, and to move the business forward.

Corey: The problem that I see is that what you've just said is, first, accurate; I disagree with absolutely none of it. But it's also nuanced. It doesn't fit in easy sound bites; it doesn't fit in tweets—which is my primary form of shitposting—it requires a level of maturity on the part of the listener to understand the nuances. Why is it such a hard concept to convey repeatedly, and well?

Eric: I think there are two things that make security difficult. If you look at availability, we have models for availability. What are the odds that a backhoe is going to cut this fiber? What are the odds that a bit is going to flip in this again? What are the odds that a human is going to violate operational procedures and push code that wasn't completely tested?

And we have models for that and you build this chain of events, and you've got some idea of the likelihood of each of these events, and you multiply through and you come up with some level of assurance that this is an acceptable risk to take, or this is a risk that we need to mitigate this much, we need to drive the likelihood of this event down below this threshold. And when you're dealing with security, you're dealing with a motivated human adversary. There's some reason. And it may be just kids out for the lolz, they're rattling all the doors down the hallway, and if they happen to find yours, you might have an issue. But in general, you're dealing with a motivated human adversary and at that point, probabilities go out the window.

It's wildly unlikely that this event, followed by this event, followed by this event are going to happen unless there's a human at the keyboard making them happen. And I found a lot of engineers shy away from that kind of thinking. And I don't have an explanation for that. I don't know why, but the idea that you're basically playing a blind game of chess with an unknown adversary is unsettling to them.

Corey: And of course, there's more than one adversary and only one of them. And despite what you say, there is the perception that security is a—you only get to fail once, and then it's all over.

Eric: So, you have to be confident in what you're doing. And this is one of the things I love about the culture at AWS. It is an incredibly important thing. If you ask anyone in AWS security what my favorite word is, they will immediately respond ‘Escalate.’ We have a culture of aggressive escalation in Amazon in general, but definitely within AWS. And an escalation is not a vote of no confidence. It's not me saying that you're bad at your job, and I don't trust you, and I'm going to go get a second opinion.

Corey: “I'm going to grab your boss because clearly, you're incompetent.” Yeah, that is in some cultures, how it's perceived. Not, when AWS does it; when people in those environments do it.

Eric: That is correct. That is not what we're doing. We're saying, “I don't think we have the right decision-makers in the room, and rather than getting caught around the axle and having a repetitive conversation that doesn't converge, or having a groundhog day meeting where we have the same document and the same argument again and again and again, we're going to get the right decision-makers in the room and we're going to make high-quality high-velocity decisions.” And so, if I'm uncomfortable with something, I know that I can pick up the phone and I can get a hold of literally any leader in the company. And they trust me: I'm not going to call Andy Jassy because something went bump in the night and I'm scared.

But if I need to get his attention, I know that I can get his attention. And I know that he will listen to what I have to say. And so given that, I know that if there's a decision that I'm uncomfortable with, if there's a path forward that's unclear, I can go get high judgment people that I trust to help me with that decision. And then when we make that decision, it's made with much higher confidence, and that enables me to continue to do my job, to continue to stare into the ambiguity of security. The other thing, I think, that makes security different from other disciplines is availability events happen much more frequently than security events.

And so we just have a larger data set, a larger training set. And so the humans that have to deal with availability issues, have dealt with them way more often than the humans that need to deal with large-scale security issues. And it's a much harder problem to quantify. And I think that's one of the things that the Cloud makes uniquely possible. I've spent a lot of time in security in multiple positions, and I have never had access to the data and the tools that I have access to here at AWS.

Between DNS logging, and Flow Logging, and CloudTrail, and all of the other data sources that we have the amount of visibility that I have, the ability to reconstruct the past, to set up alarming, and then the tools to deal with this data—not just S3 to host it and all of the machine learning and analytics tools, but things like Lambda, where setting up alarming on a new condition is the job for an engineer for an hour, not a major system design, it has completely changed the way that I think about security and the way that team thinks about security.

Corey: Something to emphasize is you're able to do all of that and have that visibility from the hypervisor and network perspective, but not from within the customer environment. And the fact that you could achieve all of this without effectively forcing your customers to make a privacy or data security trade-off is sort of its own minor miracle from where I sit.

Eric: So, I am very happy with how far we've gotten with the data sources that we have. I'm very impressed with the team and what they've managed to accomplish. One of the things that we think about: we're surrounded by constraints. I mean, that's the nature of all human endeavors. And so we have limited time, we have limited money, we have limited human resources, and the human resources are the biggest constraint.

[unintelligible 00:16:08] engineers are a hot commodity, and so every engineer-hour is precious, and making sure that we allocate those—and not optimally because then you wind up spending a lot of time optimizing and not actually delivering, but acceptably optimally is really important. And you look at the leverage of the coverage you're going to get for an invested engineer-hour. And something like Flow Logs was expensive to build, and analyzing Flow Logs is expensive to build as well, but every single thing in AWS talks IP. You can't get into or out of an EC2 instance without talking IP, and so Flow Logs gives us ubiquitous coverage, literally one hundred percent coverage; every packet is accounted for.

And that's huge. It doesn't matter what version of the kernel you're running, it doesn't matter what operating system you're running, it doesn't matter if you're playing with the latest container micro-operating system that we don't have support for, eventually, it's going to turn into IP packets and it's going to wind up in the Flow Logs. And so that's one of the things that we consider when we decide where to invest. And that sort of ubiquitous coverage is incredibly valuable.

Corey: For those of us who are doing things that are—how do I put it—not particularly serious in an AWS environment, where for example, I'm building a Lambda function to wind up taking the status page and make it sarcastic, and worse. And I'm having trouble with it. It's irritating on some level where I'm not able to push a button and grant support access into the environment to look at these things because it's a toy app, and I don't care. And it's easy to lose sight of the fact that, yeah, it doesn't matter. It's a toy app that's doing some nonsense like that, or a bank that is doing something that is incredibly sensitive and valuable and regulated, I get the same level of protection as those workloads. And that's a powerful thing, though it's, I admit, easy to lose sight of that when it's two o'clock in the morning, and I just want the funny joke to work.

Eric: I hear you. And for me, this is one of the most enticing challenges of working at AWS. We don't have grades of service; we don't have different levels of complexity. We have a single suite of services that we offer to our customers. And a novice customer that reads a blog post and wants to try something out, is going to use the same EC2, the same Lambda, the same S3, the same IAM as our most sophisticated government or financial services customers.

And in fact, that novice customer may themselves work for one of these very demanding large customers. And this may be their first foray into AWS, and so today, they're a one instance, one Lambda, one bucket kind of customer, but they're going to evolve over time into one of these very sophisticated, very demanding customers. And so there's this continuum here. And you can't tell the customer, “I’m sorry, that was great. I'm so happy that you liked that. In order to move to the next level, you need to shut everything down, pack it up, and move it over here to the much more rich-featured, complex cloud.”

You have to be able to accommodate the getting started use case, and the mildly more complicated use case, and the early production use case, and all the way on through full corporate governance, multiple accounts, organizations, security audits, compliance audits, et cetera, in a single suite of services. And I don't think we've got it perfect; I don't think we'll ever get it perfect, but figuring out how to accommodate that entire spectrum of use cases in a single service and to grow with your customers and to enable them to tackle complexity incrementally as it becomes meaningful to them, is honestly my favorite part of designing a service.

Corey: The thing that continually eludes me is I accept as fact—because you've clearly demonstrated it—that you can handle, for example, the security in all its sharp and difficult edges around things like an EC2 instance talking to RDS and then storing something in an S3 bucket. That makes sense to me. I don't know how you did it, but you clearly have done it. But then you wind up with the almost Cambrian explosion of higher-level AWS services that are in machine learning, and, “Hey, we have this thing that talks to satellites in orbit,” and oh, “There's this other thing that's Lookout for Equipment,” which is apparently named after a sign on the factory floor somewhere. And all of those things in all those different directions have the same level of security guaranteed, despite what is in many cases, a newly completely alien workflow compared to what the historical expertise has been aimed at. At least that's what it seems like from the outside. Is that accurate? Is there something fundamental that I'm missing, or is this just another demonstration of Amazon doing its operational excellence thing?

Eric: This is my favorite thing about security—as opposed to designing AWS service—is you have someone come to you, and for example, they say, “We would like to have a farm of iOS and Android devices that mobile developers can use to test their applications, and they're going to be awesome because they're going to be located right in our data centers, right next to the EC2 instances that they're using for their development work.” And you go to the bookshelf, and you pull down the big binder of policy—ask anyone in AWS security what my least favorite word is, and they'll tell you ‘Policy’—and the policy is you're not allowed to have mobile devices in the data center, you're not allowed to have cameras, you’re not allowed to have Bluetooth, you're not allowed to have WiFi. And so you run the flow chart that's in the policy, and the answer is clearly no. That is obviously the wrong answer. The right answer is, “Wow, that sounds cool. I bet our customers would love that. Let's figure out how to do it.”

Which leads to the next question, which is, “How?” I have no idea. I have never built a device farm before, but we're going to figure it out. And so we go and we find people that have the specific expertise that's necessary, but there are patterns that crop up over, and over, and over again, multi-tenancy is really challenging, but it's an acquirable skill. Capacity management is really hard, but it's something that you can build expertise in.

And so we have a whole bunch of the fundamental building blocks lying around in different parts of the organization, it's just a matter of getting the specific knowledge necessary to apply to that domain, whether it's the device farm or ground station, or whatever absolutely insane idea our service teams are going to come up with next that's going to delight customers. And it's these crazy ideas, the ones that, prima facie, seem absolutely ludicrous that wind up being really, really valuable to our customers and totally feasible.

Corey: I would be remiss if I didn't make a feature request while I have you in a circumstance in which you can't possibly say no. Now, let me preface this with, I have never yet come to AWS with a feature request and gotten a response of, “Holy crap. We never thought of that.” The answer is always, “The reason we can't do it, quite like your thinking, is nuanced and complicated.” And a couple of times I've been taken down that path, and yeah, there are dragons everywhere and computers are awful is what I take away from it.

But IAM is one of those really—how do I put this—esoteric things for an awful lot of people. It's easier to just grant access to everything, and then in turn—like, later, we'll go back and fix that. Yeah, ten years later, it doesn't work that way. We all write terrible things and we lie to ourselves and others about what we're going to be able to come back and do. It feels like there's an opportunity to build almost a warn-if-reject style IAM approach wherein a test environment—and please only use this in test environments—you could [have 00:24:19] run a Lambda function, for example, through its paces, and it looks at what function it was able to use, it’s allowed to do basically everything, and then it spits out a narrowly scoped down approach.

This is a sort of thing that people have been asking for for a long time, but to my understanding, the closest we've gotten is the IAM Access Analyzer. Is that a reasonable customer request? Is there something that winds up getting missed somewhere when people are asking for this? Or is this one of the ridiculously rare, “Wow, no one ever mentioned that to us. We'll get right on it,” moments?

Eric: I hate to disappoint you, Corey, but this is not the first time we've had this conversation with a customer.

Corey: Well, I am reassured by that, if that helps.

Eric: So, I think that things like IAM Access Analyzer are our preferred path here. And I think that over time IAM Access Analyzer will evolve to be more closely that kind of shrinkwrap that you describe, but what we've often found is that in order to get the right shrinkwrap policy, you have to exercise all of the functionality of that Lambda function, or whatever resource it is that you're attempting to shrinkwrap, and if you miss any branches, and in particular, you often miss the error branches and there are actions that your code takes when things aren't working well that are incredibly important to the survivability of your application. And so it turns out that everything's running fine for a long time, then there's some sort of failure. It's a failure that didn't occur while you were running in test mode to generate the shrinkwrapped policy.

And your code following exactly what you wrote says, “Oh, no. I have to post to this SNS topic in order to let them know that I've had a failure.” And it can't because that wasn't included in the policy. And that kind of latent failure is in some ways worse than an over-scoped policy. And so there's a balancing act here and it winds up, as you said, being nuanced and complicated in practice.

And this is one of the philosophies that we try and help our customers and our service teams understand is that you want to do successive refinement here. The tighter you make the policy, the closer to least privilege you get, the more work you're going to have to do with that policy. You're going to have to spend more time, and in the fully realized corporate governance version of this, there's going to be some other team that has to review your policy changes and approve them. And if you've got a really, really tight policy that allows exactly and only the things that you need, and then you add a feature and that feature happens to use a new SQS Queue, or takes some new feature of S3 and requires yet another API call that's not currently allowed, then you have to go through this whole process of getting this approval, and doing the review, and making sure that it's acceptable. And so as your applications mature, you want the policies to get tighter and tighter, you want the restrictions on changes to have a higher and higher bar, not just for security reasons, but for availability reasons.

The thing that you're playing around with on your own personal time, if it has a complete outage, no one's even going to notice; you might not notice. That production app that your customers are depending on, if it has an outage everyone's going to notice. And so you want to perform successive refinement here, where you keep making the policies tighter, you keep making the operations tighter until you get to a level that's appropriate for your current level of maturity, your current scale of operations, the criticality of the data you're currently dealing with. And so I'm not a huge fan of going all the way to least privilege right off the bat.

Corey: Forget dozens of visualization tools and view your entire system in one place with New Relic Explorer, the latest addition to New Relic One. See your system-wide health at a glance with a dense hex view that has your hosts, services, containers, and everything else.

And get an estate-wide view of sudden changes, so you can catch issues before they impact customers.

So go to NewRelic.com, sign up for free, and start exploring your system today.

Corey: Like everything, security feels like more of a journey, that is a destination. But that does change, for example, when you find yourself on the expo floor of RSA, at which point security is then transformed into something people are attempting to sell you. And my question across the board around that, I think is, do you see that there's a place in the security space for third party offerings to thrive in the context of a—assume a pure AWS environment along with the spherical cow. That's great. Is there a place for partners in that space?

Eric: Absolutely. One of the things that we say all the time is that we're not as smart as the aggregate of our customers. If you're building an AWS service, one of the ways that you know that you got it right is when you learn of some customer that's doing something with your service that you never anticipated, that's absolutely glorious and clever, and enabling for their business, and you got out of their way. And you never even thought about this use case, and they managed to do something that stuns you, even though you helped build this service. And so we're also not as smart as the aggregate of our partners are as the aggregate of the internet as a whole, and we want to make sure that all of these people that have something to offer, that have these differentiating ideas that can make our customers’ experience in the Cloud better have an opportunity to do so.

There's a set of fundamental building blocks that we have to own, things like EC2 itself, or S3, or IAM, or CloudTrail. There's a set of things that customers expect us to offer. For example, GuardDuty: the feedback from our customers is overwhelmingly clear that, as the owners of AWS and as the owners of CloudTrail, they expected us to have a service that would perform security analysis over those logs. And one of the data sources used by GuardDuty, and one of our external security services is CloudTrail. And so that was in response to direct customer feedback.

But we have a very rich ecosystem of partners that help customers out at all sorts of places, and some of these are born in the cloud partners. Some of these are partners that have been working with our customers for years and have made the journey with them from their on-premises data centers into the Cloud. And there is a long and bright future there.

Corey: It seems on some level like there's a bit of a series of terms of art, or its own unique dialect in the security space, where compared to almost every other line of Cloud offerings, or SaaS offerings, or developer tool offerings, that it feels like it speaks in a much more enterprise-style focus way, even when marketing to startups. Is that just because it's so difficult to message that everyone is going from the same playbook, or is there a cultural aspect of infosec done properly at a lot of these companies, that means that I'm just not in that target market, so it's a language that isn't speaking to me?

Eric: You asked me a marketing question.

Corey: Oh, yeah, I'm trying to understand. You started off, once upon a time as an engineer focused type. I mean, you don't generally become a Distinguished Engineer without writing least a couple lines of code. And you used to be hands-on-keyboard, and now you're talking to exactly those folks, and every time I talk to someone in the security space who does speak that dialect, they come away impressed at having spoken with you. So, that tells me that whether it or not, you do speak it, I'm just hoping you can sort of act as my security translator.

Eric: I do think that we've been very clear in our messaging, however. My boss, Steve Schmidt, who is the Chief Information Security Officer of AWS, has talked a lot very publicly about how we think about security and how we treat security is something that's baked in from the beginning; how our messaging with our customers is around helping them move forward with confidence, not about sowing fear, uncertainty, and doubt; it's about making the pie larger and enabling more people to succeed, not in scaring people off from doing things. And so I think, to a large extent, our security marketing, if not our security product marketing is very much in our own voice. And I think it does a good job of conveying the message that we want to convey.

Corey: So, a challenge that I have to imagine is frustrating, if nothing else, is that the reality of AWS and the perception of AWS have some significant gaps, where on the one hand, it's the idea of two-pizza teams, and people iterating rapidly, and bunch of small service teams, each building something as part of a collective whole and on the other, you take a step back, and it's you’re Amazon; your market cap is measured in the trillions. Why is ‘Insert whatever thing annoys you today,’ such a bad experience or whatever it is, how does that tension wind up manifesting in your world.

Eric: So it's true, that the security team has gotten to be reasonably large. And you look across all of AWS—and I've been with the company now for 13 years, and it is dramatically larger than it was when I started. But the job that we're taking on is also dramatically larger than it was when I started. It does not feel like our budgets have gotten any richer. It's just customer expectations have gone up, the expectations in terms of compliance, in terms of security, in terms of availability, in terms of operational excellence, have all gone up at the same time that we've been launching more and more services and features.

And so we're still incredibly parsimonious with our engineer time. And a lot of our best security tools are things that an engineer was tired of dealing with and they went off; in the space of a couple of days, they made an absolutely horrendous prototype. Like, this code should never even have been typed into the computer in the first place, but it made their lives better; it made their job easier, and so another engineer contributed some code and it became a little bit less eye-searing. And over the course of a couple of months, we wind up with a system that's actually really useful. And at some point, you have a discussion, you're like, “Wow, this thing is no longer really useful. This thing is essential to our operations. We've hit another level of scale, and if we didn't have this automation, we wouldn't be able to keep up anymore.”

And so then you build a team around it. And when we say build a team, there's the whole two-pizza team thing, and we don't really talk about buying pizzas and thinking in that term, but these tend to be very, very small teams—you know, a handful of engineers, a software development manager—and now that team owns that thing, and they evolve that thing. And all of the security tools that I see that I really like are things that started off as small tactical answers to an actual problem that we had that accreted functionality over years. And it usually means that they're not beautiful, that there isn't some grand design that some architect sat down and sketched out and thought about all of the future scaling concerns, it means that they tend to be kind of patched together and evolved and as-built. But the reality is that the grant designs that the architect sits down to sketch out usually don't take into account the future that actually happens, and so you wind up with patches and changes, and emergent feature requests anyway. And it is incredible how quickly that value accretes.

Corey: Oh, absolutely. People are familiar on some level with the idea of the mythical man-month, it feels like this is almost a parallel of that the mythical, “Just throw $5 billion at it and wait,” where it's throwing additional resources doesn't lead to better outcomes and in many cases can lead to materially worse ones.

Eric: Absolutely. And so when I'm talking to customers and they want the tools that we have, one of the reasons that our tools are valuable, is that they're tightly integrated with the way we do things. At Amazon, we have a ticketing system, and everything is a ticket: if your laptop needs more memory, it's a ticket; if you want to bring your dog to work, it's a ticket; if the website is down, it's a ticket; if your parking token doesn't work, it's a ticket. Everything is a ticket. And so all of our security tooling is integrated with a ticketing system, we even have security tooling that monitors the ticketing system to make sure that the tickets we've already cut are in a healthy state, and to take metrics on that, so we can report on it, so we can understand if we're spending our time in the right places.

And none of that integration translates. And so what I tell customers that are looking to get started on this journey, customers that want the kinds of tooling that we have, I tell them to just get started. Rather than writing a catalog of all the things you'd like to check and all the Lambdas you'd like to write, just write one. Just pick one.

Corey: Check a single thing, write a quick three-liner, that'll do it, and see how it goes.

Eric: Yes.

Corey: Yeah.

Eric: And the most important thing is not that you have that check, it's that you have the feedback loop. It's that the next time something goes wrong, you think, “Why did this go wrong? What can I check that would prevent this from going wrong?” And then you add that check. And so over time, you're going to accrete this library of validations.

And the way we think about this is in terms of invariants. We call them security invariants. These are statements that should always be true. And they can be incredibly simple, like, “This IAM policy matches this text document exactly.” Or they can be incredibly nuanced, like, “There is no path from the internet through any combination of nodes to any host that's tagged blue.”

And so the validators can be very simple; they can be very complicated. You build this library of invariants. And every time something happens that you don't like, or during the application security process ahead of time, you come up with invariants. And you just keep building this library of invariants. And every single time we've done this, the library of invariants that we've wound up with is very different from the library of invariants we thought we needed, and because it's driven by things that have actually happened or things that we specifically identified in our threat models, they're the things we actually need. And that value accretes incredibly quickly.

Corey: It's a matter of taking a bunch of little things and composing them into something fantastic at the end. It's almost like the microservices story, or some of the architectural diagrams that list a borderline sarcastic number of services, but the outcome is really neat.

Eric: Absolutely. And over time, you will learn that past-you was not as smart as current-you, but that's fine. The principal engineer community has a set of tenets, and one of the tenets is ‘Respect what came before,’ and it's incredibly important to me as an engineer. I've been around long enough that I've seen things where I've said, “Oh, my gosh, what idiot did that?” And you look in the source repo—it's git blame these days, but it was CVS blame back in the day, and your name is next to that line.

Corey: And then you immediately fire up git blame someone else.

Eric: No, no you own it. Like, “I made this decision.” And this—

Corey: Yes. That’s why you use the tool that rewrites history. So it's someone else's fault and not your own. Oh, yeah, I'm right there with you.

Eric: So, the idiots that built the systems of the past weren't idiots. In fact, they're the ones that got us to where we are today. Those systems are what enabled our current business, our current success. Now, every single thing I've ever worked on we've outgrown. You got a couple orders of magnitude scaling out of your design, and then you've got to go back to the drawing board, but you do so making sure that you respect what came before; that you value the systems that got you to where you are, even though they've scaled beyond their utility, even though you think they're old and broken, they embody lessons; they're wise; they're battle-tested.

And you make sure that you take as many of the lessons as you can from the systems that got you to where you are, and you treat them with respect, even as you turn them off in favor of the new shiny thing. And after you've been through that cycle a couple of times the new shiny thing is going to be one of those legacy systems someday soon.

Corey: I tend to view legacy through a lens of being a disparaging engineering term for ‘It makes money.’ It turns out that unlike what we learn in conference talks, you can't generally throw the entire banking system away and replace it with something you built in a weekend off of Hacker News. So, I have an awful lot of sympathy for not just the greenfield stuff, but how you get what exists today into an environment that is better tomorrow? And there's no easy answer.

So, I want to thank you for taking so much time to speak with me about what you're up to, and how you folks view these things. If people want to learn more about what you're up to, where can they find you?

Eric: So, I am on Twitter, @ebrandwine. I'm not very good at the whole social media thing, so caveat emptor. We also have a wealth of material on the AWS Security Blog. A lot of the stuff that I've talked about here about how we think about security and about making incremental progress is well covered there.

Corey: Excellent. We will, of course, throw links to that in the [show notes 00:41:45]. Thanks so much for taking the time, I really appreciate it. One never knows what one's reputation is with different groups at Amazon; there's no unified single opinion Amazon has, so it's nice to know that least some people will still take my calls, and it's very much appreciated.

Eric: I think that your taste is terrible, and the fact that you had me on just confirms that. And the fact that anyone wants to listen to this is mind-boggling to me.

Corey: One person's trash is another person's treasure, and I'm generally the trash. Thanks so much. I appreciate it.

Eric: Thank you, Corey. It's a pleasure. Eric Brandwine, Distinguished Engineer and VP at AWS, I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you've hated this podcast, please leave a five-star review on your podcast platform of choice along with a comment saying that actually there's a job negative one, and tell me what it is.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Dennis

Dennis Gada is SVP and Head for Financial Services, North America at Infosys, where he has executive responsibility for all client relationships and new client acquisitions in the Financial Services sector. Dennis has significant Business Transformation, Innovation, and Financial Services Consulting experience. He is an industry leader in Financial Services with experience in partnering with clients to shape strategies and execute digital transformation programs leveraging business and technology services. Dennis is a frequent speaker at various industry events and is a member of the Institute of Chartered Accountants.

Links:

  • Infosys: https://www.infosys.com/
  • Email: dennis_gada@infosys.com
  • Twitter: https://twitter.com/dennisgada

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the Enterprise (not the starship). On-prem security doesn’t translate well to cloud or multi-cloud environments, and that’s not even counting IoT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IoT devices, detects these threats up to 35 percent faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at extrahop.com/trial.

Corey: If your mean time to WTF for a security alert is more than a minute, it's time to look at Lacework. Lacework will help you get your security act together for everything from compliance service configurations to container app relationships, all without the need for PhDs in AWS to write the rules. If you're building a secure business on AWS with compliance requirements, you don't really have time to choose between antivirus or firewall companies to help you secure your stack. That's why Lacework is built from the ground up for the Cloud: low effort, high visibility and detection. To learn more, visit lacework.com.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’m joined this week by Dennis Gada, head of financial services at Infosys. Dennis, thanks for joining me.

Dennis: Thanks, Corey. Good to be here.

Corey: So, this is the first time I’ve had someone from Infosys on in the entire run that this show has happened. And the reason behind that is—in the interest of full disclosure—my wife is Senior Counsel at Infosys and General Counsel of the Infosys Foundation USA. I mostly try to stay out of her way, given my mouth and attitude problems. I mean, there’s a terrific chance at some point if I start mouthing off, she would have to sue me on behalf of her employer. So, we’re going to try very hard not to cause that to happen this episode. So, let’s see what we can do. Dennis, thank you for joining me. What do you do?

Dennis: Corey, so I head the financial services business for Infosys in North America. I’ve been with the company for almost 16 years. In fact, I complete 16 years this week. And through my journey here, I’ve been working across Asia, Europe for ten years, and last six years in North America. And really focused on delivering end-to-end technology operations, digital services, for our financial services clients. Very exciting times to be at the intersection of technology, and business, and financial services at this point.

Corey: Oh, it absolutely is. I don’t talk about this too often on this show, but my last real job before starting my own nonsense was at a startup that got acquired by BlackRock. So, going down the whole financial services path was something that I got sort of a crash course into, and it definitely leaves an impression. It’s not like a lot of other sectors for a, frankly, very good reason.

Tell me a little bit more about what Infosys does because my understanding personally comes from my wife, who’s an attorney. Love my wife and love attorneys though I do, they always tend to bring something of a unique perspective to things that I’m not sure adequately conveys the fullness of the experience. Helped me figure out, please, so I can give a better answer than I’ve been giving for the last five years of where my wife works, please?

Dennis: Sure, I’ll break it down for you, and you can go and validate it with your wife as well. But I think from an Infosys perspective—so at the foundation, we provide the technology services to kind of run applications, support infrastructure, test applications, and so on. If you look at the next layer, it’s really about helping in the change and transformation of businesses. And much more in the last few years, our business has pivoted towards that. So, whether it’s launching a new digital bank—if I take the financial services example—or launching new payments on products, new digital wallets, and so on, that’s the kind of really transformation-oriented services we provide.

We just announced our results, and almost 50% of our business now comes from digital services—which is kind of the cherry on the cake, so to speak—which is around helping clients on experience, on data analytics, on Cloud and IoT, on modernization of their platforms, and around cybersecurity. That’s where we are headed towards; most of our new business, our future business, is around digital services, Corey.

Corey: It’s an interesting transformation story on the part of a company that effectively helps its own clients with digital transformation. It’s easy to forget from the tech bubble that, well, I find myself ensconced in, let alone what other people do, that companies have been doing business since before the .com bubble. I know it’s hard for folks to believe sometimes, but it is, in fact true. And as much as we want to imagine that all of the innovation and all of the code worth looking at et cetera, et cetera, comes out of startups founded within the last two years in San Francisco, or out of Google, the reality fundamentally is that companies are doing business, they’re not stopping the business that they’re doing, but taking an existing large scale, often regulated workload and moving that into a cloud-native style environment, or modernizing beyond the mainframe era, it takes a lot of work because it’s hard to do, but it’s also because it takes so much effort to ensure that during that migration, everything still works. As much as you might like to, you can’t take the ATM networks down for six weeks while you go ahead and migrate everything. At least not more than once, if you don’t want society to be standing by the end of it.

Dennis: Absolutely. And let me explain that also with a very recent example. We did a partnership recently with the Old National Bank. They’re headquartered in Evansville, Indiana, and one of the regional banks, but largest are in the state of Indiana and also do business in the neighboring states. And this was six months before the pandemic; they had a new CEO, Jim Ryan, who had taken over.

And dropped him an email and went there to meet him, I thought it would be an executive lunch in a boardroom. Jim and I sat in the cafeteria and had a discussion over salad in terms of how the world of banking is changing, and how the innovation that’s happening all around the world, given my global perspective, whether it’s in Silicon Valley, whether it’s in Asia, whether it’s in Europe, et cetera, how all of that can be brought into regional bank which is aiming to transform itself in the state of Indiana. And then we build the relationship over the next six months, and just before the pandemic, we entered into a partnership with Old National Bank. And the entire focus has been around, on one side, how do you drive efficiencies into the run side of the business, the existing infrastructure, the existing services, moving to the Cloud, and so on, and then plow back or invest it into the transform side of the business, how to really build new capabilities in commercial banking or wealth management in this case. And that’s a great example I wanted to mention because innovation is not just happening in pockets in certain parts of the world, it is all over.

And at Infosys and in the work we do in financial services, our aim is to really bring innovation closer to our clients, whether it’s in Evansville, Indiana, whether it’s in Charlotte, North Carolina, whether it’s in Amsterdam in Netherlands, or any part of the world. And that’s what we’re focused on.

Corey: I think the common misconceptions among folks who start off in small businesses—as I tend to. I mean, I admit my own biases here, I tend to think of a big company as someone that employs 200 people. That’s not really how it works, but from that perspective, it doesn’t seem like doing a migration or doing a digital transformation should take that long. So, when you hear about a company taking five years to do a cloud migration, for example, the knee-jerk reaction in some folks—and admit this was me before I started having deeper conversations—is, oh, they must be super bad at it. What are the things that those folks generally tend to overlook?

Dennis: Yeah, I think there are sometimes internal organizational bottlenecks, whether there are governance problems, legacy problems, and also the whole mindset or culture of the way things are done. Those are the issues that large enterprises typically struggle with. The issue is not with the technology, or the issue is not with what capabilities are available in the market, but many times internal culture and governance issues. But one thing I have to say, Corey, is what has changed that dramatically is the pandemic in some ways. When we look at certain organizations, we almost feel that what would have taken ten years to do has been accomplished in the last ten months or so.

In fact, one of my clients said that they feel that they are in 2030. That’s the level of work that has been done in the months during the pandemic. Some of it was because of certain government interventions such as the Paycheck Protection Program—or PPP—here in the US. In some cases, it was more driven by how the expectations of the end clients of many of these banks have completely evolved during the pandemic. And a lot of it was happening before the pandemic as well. But it has really, in my view, got a super acceleration. And then I would love to talk more about that.

Corey: Oh, absolutely. The pandemic has absolutely upended industries from one side to another. I mean, the running joke is that COVID-19 has done more to foster your company’s digital transformation than your last ten CIOs and there’s a strong element of truth to that. And on the one hand, you’d look at that and say, “Oh, well, banking was all mostly virtualized or digital online anyway. That hasn’t really had that much of an impact.”

But you bring up the PPP program, which was an unprecedented effort to get giant piles of money out to businesses in a responsible way. And that was passed on to the banks themselves. What did that look like from an IT infrastructure perspective?

Dennis: So, very interesting point. I think it put a tremendous pressure on the banks. The volume of loans that they had to process increased 20 to 25 times what they have typically done. And that, too, in a very crunch timeline. And different banks adopted different kinds of approaches to handle that massive spike in volumes.

But I must say that the banks that were better prepared and used technology to really address this challenge came out very, very successful. In case of one of the regional banks that we helped in, again, in this case, I’ll just quickly walk through the process. The whole application to accept loans was set up in a couple of days, making it very easy for these small businesses to apply for these loans. The entire underwriting process, they used AI technology from Infosys to extract information from the documents, payslips, et cetera, and for low-value loans, do an automated underwriting. Then they used the automation, the RPA technology to submit those applications to the SBA E-Tran’s website.

APIs were used to really also push applications in big batches to SBA. So, a combination of RPA, AI, APIs, automated underwriting, and a transformed experience of accepting applications was used to really drive volumes. We were working around the clock. Everybody was hands-on, including the technology and business leaders, making sure that this gets done. And I think in my mind, that was an absolutely amazing experience of seeing, one, how technology can really help the business, and how, if everybody comes together on a particular challenge, things can be done so much faster.

Corey: It’s sort of jarring when someone who isn’t really conversant with how the financial system works, starts to peek under the covers of it. I mean, functionally, the way that money moves with checks, as an easy example here, is that effectively it’s a promise to pay that gets passed around, and it’s super easy to effectively cheat. The way that society has fixed this is the audit trails around this are nothing short of astonishing. So, sure, you could get money out pretty quickly and cause problems, but they will absolutely have this on record, and oh by the way, that’s a felony. So, it’s a real strange type of combination between legacy systems that still need to be supported, and technical shortcomings that, in turn, were stop-gapped by regulation, policy, and societal norms.

Again, I speak primarily to the United States. I can’t speak much to international banking state these days, but that was the thing that blew my mind, not the fact that all of that was—how it grew organically. I mean, that makes sense, but the fact that it works at all. And as best I can tell from the outside, works effectively. We didn’t see systemic collapse during the PPP program, we didn’t see banks having their IT systems explode.

I’m looking forward to the future case studies about how this stuff was implemented and achieved. But from the outside perspective, for most people who are not deep into the technical weeds, it just seems like another day of doing business. It very much wasn’t.

Dennis: No, it absolutely wasn’t, right? But I think what it did was that after banks went through that process, and especially the ones that came out very successful, they realized that this could be the new way of working. If this is the way things can be delivered. Literally, in four to six weeks thousands of loan applications with new technology was successfully delivered with all the checks and balances from a fraud and [evolution 00:13:26] perspective, how is it that you can use that model in the future? And in a way, it gave the industry a glimpse of the future.

And I think banks are now looking at how they can sustain that kind of a model. I’ll use an analogy, and I actually borrowed from a chair of one of the largest digital-only banks in the US. In one of the other events, he said that this was the difference between an opera and a flashmob. Normally, in financial services, everything is well-regulated, there is a conductor, you’re told exactly what to do, everybody follows, does it in a structured way. But I think what the whole digital transformation agenda last year converted [teams 00:14:06] into more of flash mobs that were more energetic with new ideas coming together to really drive major transformation, major shifts in very short timeframes. And I think that model has really given new life to how changes can be driven in these large organizations. And we see that that continues to sustain, even in 2021 and beyond, and almost becoming the new ways of working.

Corey: Forget dozens of visualization tools and view your entire system in one place with New Relic Explorer, the latest addition to New Relic One. See your system-wide health at a glance with a dense hex view that has your hosts, services, containers, and everything else.

And get an estate-wide view of sudden changes, so you can catch issues before they impact customers.

So go to NewRelic.com, sign up for free, and start exploring your system today.

Corey: I’m very curious to get your take on this, where having to be able to expand so rapidly to service the PPP program across the board, it seems like it’s kind of unlikely—to understate it massively—that now that, as things begin to return to normal—ideally—that, okay, time to dismantle all of those systems and go back to the old, slow, manual way of doing things. These systems are in place now and they’re not going anywhere. What does this change once society returns to some semblance of normalcy?

Dennis: Banks were already working on digital transformation, even before the pandemic, before PPP. But I think a lot of that was focused maybe on building the right websites, portals, mobile apps, and so on. I think what some of this transformation has shown is that you have to really transform end-to-end, not just your front-end but also your middle and back-office systems and processes. And hence, I think this has really generated a new wave of initiatives within many organizations to really look at a full end-to-end transformation, keeping the customer at the center. If this forced you to process loans within not days or even hours, but within minutes, how is it that you can make the other processes also similar.

Account opening is a common challenge. Every time you try to open an account with a new bank, it involves a lot of documentation, manual submissions, et cetera, and sometimes could take days. How do you make that a completely digital process that can be completed within minutes? How do you bring more real-time analytics into the spend and the saving patterns of your customers? How do you use data and analytics more efficiently to upsell products and services?

So, those kind of new paradigms, much more in the front now. And the underlying problem still is the legacy infrastructure. So, a lot of the banks are now really focused on transforming not just the front-end capabilities, but also the back-end systems and the legacy infrastructure that they have inherited over many years.

Corey: One of the things that also I think gets lost as we take a look across the board—I think finance is one of the easier places you can see it, but it does take place across the board—where there’s this idea that developers are the ones that are always going to be making the technical decisions, the vendor selection process, et cetera. And there was a book ten years ago by the founder of RedMonk called The New Kingmaker that speaks specifically to this. And yeah, the idea of bottom-up adoption is incredibly valuable. The counterargument, of course, is that developers and engineers generally are not empowered to sign $50 million cloud contracts. So, at some point, it’s less about technologists talking to technologists, and it shifts instead into speaking to businesses.

Is that something you’ve seen in your particular practice whereas you’re seeing a shift towards more business stakeholders are involved in discussions than in years past? Or has it always been this way and I’m just finally waking up to a longstanding truth.

Dennis: No, I think there is a significant shift that has happened in the last few years. I think the intermingling of business and technology, both from an organization structure perspective, decision-making perspective, is much more now than it was ever before because businesses are really hungry for innovation. They want to launch new products to the market as fast as they can and improve the experience that they can provide to the customers, and hence the main enabler to do all of that for the survival of the businesses is technology. And hence, they are much more involved in making decisions related to technology. Talking about modernization, in July 2020 we entered into partnership with Vanguard to really advance the digital transformation of their record-keeping business. And Vanguard, as you know, is one of the largest DC asset managers in the US. And this partnership is all about driving transformation in the record-keeping platform, building a new cloud-native platform for record-keeping—the first in the industry—and managing also their operations and technology. And I think this is a partnership that will set the benchmark for how digital transformation should happen in the record-keeping industry.

So we do see a major shift in overall decision making and a much larger interest from the business in making significant decisions that do involve our technology services. I talked about Old National Bank, the partnership we did with them. Again, it was driven by the CEO of the bank. So, it’s not just business, but very senior levels of the organizations are getting involved in making the right decisions, from the technology transformation perspective.

Corey: Historically, there was always a bit of a negative connotation to that. The running joke among technologists—because it wasn’t really a joke—was that the Oracle sales team would completely bypass the technologists and go and talk to the C-Suite and other executives. And the perception was then that, “Oh, they’re just taking them out to play golf, and that’s what wins the deals.” Personally, I find that a little bit cynical. For a long period of time, Oracle Technology offered something that nothing else really did, their business practices notwithstanding.

It was important of course, for the amount of money they were charging, they needed to get executive stakeholder buy-in and it made sense, but I’m curious as to how much the reality mirrors or doesn’t the popular perception. Not necessarily speaking to Oracle, of course, but to large enterprise software vendors and providers in general.

Dennis: Yeah, I think businesses are now much more informed about technology. Many business leaders actually come with a technology background, they have been studying this space quite minutely. And so when they are involved in the decision making, they actually are quite well-informed of what are the options, what’s the impact that a particular platform or technology solution would have on the business. So, I feel it is actually bringing the right focus in the decision-making process, with business stakeholders looking at what is really important for them and how a decision or a choice they would make will directly impact certain outcomes.

I’ll give you a very quick example. We’ve worked with the head of operations for a bank on a challenge related to collections. And the discussion there was around how can you use AI automation—in our platform in that case—to improve collections? And we ran challenger model to the existing manual process of doing the segmentation of the customers, identifying the high-risk customers. And they realized that the challenger model using the AI platform from Infosys was actually giving better results, and then that model was used to enhance the downstream collection capabilities across various products.

Now, the business in this case, the business stakeholder, the head of operations, did not make a decision just because they were very familiar with Infosys, but really using the platform, doing proof of concept, seeing the results, which actually impacted the business outcomes, improved the collections by 25 basis points, and then making a choice that this is the right platform to go for. So, I think businesses have become much smarter in understanding how technology works and using that to make decisions that would be right for them and clients and for the right business outcomes.

Corey: And that’s really, on some level, what I think—to put the shoe on the other foot for a minute—that technologists have really missed out on over the historical trend, where they tend to miss the business forest for the trees and they focus on pure technical merit of, “Well, why would I spend this pile of money on a vendor solution when I can put together a few open-source projects and have it achieve mostly the same thing?” In some cases, they’re right. In other cases, there are concerns about liability, there are concerns about, okay, it’s broken and I don’t really know why. There’s something nice about having a phone number you can call and get things addressed. I don’t think that it’s an easy answer; I don’t think that one side is always going to be right in those, so please, if you listen to this and disagree, don’t email me. But there is an increasing awareness that I’m seeing as well on the technology side, where they are starting to wake up to the realities of business. Is that something that you’re seeing from your perspective as well? Or am I just hanging out with some better people than I used to?

Dennis: I think that is changing, but I wish that happens even faster. And especially it’s very important, Corey because a lot of the new technologies coming in, technologists love to experiment with new ideas and bring forward the new technologies, but they should not do it for the sake of just using the new technology. Again, it needs to have a business impact. So, wherever the technology teams that we work with are able to marry the changes in technology with what it can mean for the business, and then put together those solutions, they’re definitely seeing much more of an impact. And I’ll give you another example—and this was actually for a European bank—where they set up a new retail store and they built a lot of fancy cafe-like experience and so on, in that retail store.

But I went into that retail store of that bank, and they greeted me with a tablet and welcomed me, and then I asked them a specific question about a very complex problem I had with a particular mortgage in that case, and then it all fell apart because they could not figure out who could help, I had to wait for 30 minutes, then somebody came and said, “You know, you have to make changes in two or three different applications, and because of that, it’s going to take ten more days.” So, while the facade, or the face, was very good in terms of a very innovative branch, but when you actually try to solve the real problem, it was still all legacy and no changes, no modernization was done in that case. And I think that’s the problem. So, the technology teams have to really look at it end-to-end from a business perspective and not implement technologies or new technologies in certain silos, but really see how the journey or the experience can be improved from an end-to-end perspective.

Corey: As you look across the landscape right now, what is it you think that financial services as a whole is doing right when it comes to their digital transformation stories, and what do you think that they could do to improve where they are that they’re not doing yet?

Dennis: There is certainly a significant focus on customer experience and going beyond just the mobile app or a great website, really becoming part of lifestyles of customers. So, if you take the example of mortgage, it is not just about the transaction of getting a mortgage, but being part of the home-buying process and the home servicing process. We actually call it ‘living dot com,’ which is about going beyond just a mortgage, and helping clients have a great home and a lifestyle. Similarly, looking at data and analytics from the point of view of how do you do more predictive analytics and personalize the offerings? How do you provide more tailored recommendations?

So, I think a lot of effort is going into customer experience around data and analytics, keeping really the customer at the center, and that’s really the right thing to do. There is also a significant shift in the last 12 to 18 months on the journey to Cloud—and I’m sure you’re quite familiar with that as well, Corey—okay it has become real in the banking world, the movement to Cloud, whether it’s private cloud, public cloud, hybrid cloud, it is very, very real, and banks are looking at, really, how to leverage the inherent capabilities of the Cloud, whether it’s scalability, resilience, security, et cetera, not just from a cost perspective, but really from an innovation and a time to market perspective. So, I think all of those are the right areas to focus on from our transformation. Where I think banks still need to do more is what I talked about earlier, looking at their monolithic legacy platforms. Those are complex challenges to solve, to modernize those over a period of time.

Because otherwise, no matter how much you try to improve the customer experience, use Cloud, use data analytics capabilities, but if the back-end remains legacy, then there will be certain limitations. And I think really bringing in that transformation is where more needs to be done. And it’s a more complex problem to solve, as well.

Corey: It is. I think the thing that everyone hopes for and can’t get to is the idea that somehow there’s an easy trick that if only they just look in the right place, they’ll find this magic thing that gets them where they want to go with none of the downsides. Yeah, me too. If anyone can ever sell that, I would be your first customer, but I don’t [laugh] believe it actually exists.

Dennis: [laugh]. You know, there is no magic trick. This is all a lot of hard work, but as long as you have the customer in the center and then you try and build everything around it, at least it starts making sense. But many times, if you just have technology in the center and you’re doing things for the sake of technology, for the sake of implementing new technology, you really don’t get to the desired objectives. So, I think that’s where we really see firms that have done significantly better.

And everybody is focused and investing in the whole digital transformation journey, but there are clearly firms that are far ahead in the journey, and they have really identified what makes sense from a customer perspective, from a business perspective, and that’s all that they focus on: using technology as more of an enabler to get to that North star faster.

Corey: I think that’s one of the most valuable takeaways we can really have from this conversation. Thank you so much for taking the time to speak with me. If people want to hear more about what you’re up to and how you view these things. Where can they find you?

Dennis: They can reach me, dennis_gada@infosys.com on my Twitter handle, @dennisgada. It was a pleasure, Corey, to speak to you as well, and I wish you all the best.

Corey: Likewise. Thanks so much for taking the time to speak with me today. I really enjoyed it.

Dennis: Thank you.

Corey: Dennis Gada, head of financial services, Infosys. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve hated this podcast, please leave a five-star review on your podcast platform of choice, along with a comment telling me why you wrote some code at 3 a.m. this morning, and now it’s perfectly good to push it out immediately to the nation’s ATM networks.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Jess

Jess Schalz (she/they) is a computer gremlin multiclassing in software development and infosec. She’s a queer disability advocate, and this informs her empathy-based approach to technology. Her hobbies include watercolors and collecting human remains. Talk to her about weird medical history and cats.

Links:

  • Transposit: https://www.transposit.com/
  • Twitter: https://twitter.com/jessica_schalz

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the Enterprise (not the starship). On-prem security doesn’t translate well to cloud or multi-cloud environments, and that’s not even counting IoT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IoT devices, detects these threats up to 35 percent faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at extrahop.com/trial.

Corey: If your mean time to WTF for a security alert is more than a minute, it's time to look at Lacework. Lacework will help you get your security act together for everything from compliance service configurations to container app relationships, all without the need for PhDs in AWS to write the rules. If you're building a secure business on AWS with compliance requirements, you don't really have time to choose between antivirus or firewall companies to help you secure your stack. That's why Lacework is built from the ground up for the Cloud: low effort, high visibility and detection. To learn more, visit lacework.com.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Jess Schalz, who's a workflow engineer at Transposit. Jess, welcome to the show.

Jess: Hi, thanks for having me.

Corey: So I'm going to dive right in and ask what the heck is a workflow engineer? It sounds either amazing, as far as something a company should have to improve development workflows, or alternately it sounds like a coordinator type of role that has been up-leveled to have the word engineer shoved into it in some sort of weird descriptor for a common role. I have the sneaking suspicion it's the former, but I don't want to assume. Tell me more.

Jess: Yeah, a workflow engineer, in my capacity, is someone who works with customers or clients to build automated processes for what they need in their DevOps-y workflows. So, I tend to work with things like APIs to build custom calls or workflows for people that they can then call through, like, Slack. That's kind of my function.

Corey: So, this strikes near and dear to my heart because one of the things that you self-identify as, and is extraordinarily relevant to the way I'm starting to think about the world is that you are a advocate for accessibility and neurodiversity. Is that an accurate description?

Jess: Yeah. I would say I'm an aggressive advocate for both of those things. Yeah.

Corey: The reason I bring that up is, the more that I do things like, oh, I don't know, record podcasts and build workflows around it that people can work with, something that I'm noticing is that what I've accidentally been building here is things that accommodate my own expression of ADHD, where I find it challenging to do certain things on a schedule, so everything I do, start to finish, on this podcast is either automating handoff to other people, getting me the things I need when I start these things and making sure that I'm not scrambling last minute to go ahead and implement stuff. And on the one hand, it's made this feasible; I do a lot of these, and if I had to do all of these things bespoke and by hand every day, it would be an unmanageable burden. But on the other hand of it, I've always felt bad about this kind of thing. If only I wasn't quite so shitty, I'd be able to do this like a functional person, and I wouldn't be burdening everyone I work with. Is this aligned with what you're talking about?

Jess: I think so. I think it is a way for people to approach a problem in a customized way that they are comfortable with and that is approachable for them, so it's not as overwhelming to go into a whole dev platform or a whole codebase that you don't know, to do something. So, for people who aren't familiar with code in general, being able to have it in an accessible and approachable way is very useful. And we see that a lot with, like, automated processes for daily life to like Google homes. So, it just depends on what that accessibility and approachability looks like for you.

Corey: From where I sit, it felt to me like—well, let's back up a bit and remove from the podcast space because admittedly for the audience, that tends to be a little bit niche. It seems like we have more folks listening who are familiar with DevOps workflows, and once upon a time I called myself DevOps because calling yourself a sysadmin means you're doing the same job, but you make a third less. So yeah, absolutely, I'll jump along with the hype, if it means people are going to pay me. Yeah, you can call me anything you want if you're paying me.

Jess: [laugh].

Corey: And the things I saw as a result of that were, workflows were incredibly important. Everyone talks about CI/CD, for example, which made us feel way better about not actually doing CI/CD because there was a lot of scaffolding that was tied into it, there was a lot of boilerplate. If you want to build ‘Hello World’ in a way that you would actually build an application, it's not three lines of code, it's three hundred and most of a morning in some shops to get that stuff set up. So, every chance you get, if you don't have something that is enforcing this, or much more ideally doing it for you, you're stumbling through the, “Well, I'm just going to work around that and take a shortcut here. And I know you should never hard code credentials, but I'm going to take a shortcut and do it anyway.” For one of my production systems—I kid you not—the DynamoDB table has the word test in its name because it was an experiment to see if it works. And then it worked, and well, that sucker is load-bearing now.

Jess: Yep.

Corey: And it feels like that's the sort of thing that you allude to with what you do as a workflow engineer.

Jess: Correct. Yeah. That's very accurate, I think. It's what we try to go for, and what really drew me to the role, in general, is that we wanted to create something so that engineers and non-engineers could use this product in a way that is very user friendly. So, I don't mean to have this be a sales pitch. But—

Corey: Oh, please, sell. Sell.

Jess: [laugh].

Corey: Pretend for a minute that you're a mediocre white guy who’s pushing something. Go for it. The only promotion that is acceptable to talk about is self-promotion, by all means. Shine on you, ridiculous diamond.

Jess: Thank you. Yeah, I think what really drew me to this is that the reason I wanted to be a part of this team was because to me, something that's incredibly important is making sure everyone feels like they have power over their own processes. So, for me with ADHD, it's a big problem in my home life. So, for you, it's organizing the podcast. For me, it's cleaning my apartment.

So I've had several people suggest getting a cleaning service, which isn't responsible right now, but maybe in the future that's something I'll do. And I think being a workflow engineer also ties into that in a lot of ways because it lets us build processes that make parts of our lives easier and less overwhelming.

Corey: The problem that I always ran into, and I'm curious as to how typical this is across the industry, is that when I have problems with workflows, or problems with these things, I look at this, and I put on my consultant gaze at it from 30,000 feet view, and it's like, “Oh, the reason that I'm having trouble with all of this is because I'm shitty. If I were actually good, I would have done all these things properly.” So, every time I see a bad workflow, I take a mental shortcut to, “I'm just super bad at things.” It's my secret shame, so don't let people see.

And by now I know enough to know that this is wrong, and everyone who feels that way is wrong. There are procedural and process boundaries there, but it feels shameful and you feel alone, which means that I don't think people advocate too often for improving workflows because it sounds too much to them like there'd be advocating for replacing themselves somehow with someone who is better at this.

Jess: Right. That's a huge anxiety with especially older sysadmins that are worried about code automating their jobs out of existence. I've heard that from a ton of people. What I drive for in workflow and general life process automation, is that it isn't that you’re shitty, or it isn't that you are bad at your job, or bad at things, it's that your current conditions aren't correct for you, or they aren't correct for the current period of your life. So, something I've recently changed in my life is that now I have timers in my life to tell me when to take breaks. Otherwise, I will continue on an activity for way too long. I think that tied into me for a long time where I just thought, like, “Oh god, I'm just a bad engineer and I can't time manage.” That's not the case. I just needed something to remind me every so often because my brain is funky.

Corey: Yeah, it feels like, I look at how it expresses, for me at least, in terms of not being like other people. And if you ask me to list off how my ADHD manifests and what that means, I will come up with a laundry list of ways that I am difficult to live with both personally as well as professionally, and it's a giant list of my shortcomings. I'll just completely gloss over the other side of it where it empowers me to do a number of things that would not be tenable otherwise. If I take a look at what I actually functionally do, in terms of a marketing context, for the podcasts and the newsletters and the rest, in most companies, that would be the role of three to five people, in most cases.

Jess: Yeah, I think ADHD in a lot of ways is a superpower because we are so able to jump from topic to topic. So, I have found personally, agile development works really well for me because it's so, like, short story; it's so quick and rapid development that it's way less overwhelming to me than things like waterflow style, or like huge project all at once overwhelming mountain of work style of project. And I see that also translating into my daily life, where as an engineer, I find that very helpful for ADHD to let me jump through different hoops. I also find it very helpful for me to notice all of the things around me happening rather than only being able to address one thing at a time.

Corey: I'm always a little reluctant to make sweeping generalizations because there's one thing I've learned, it's that ADHD is very much a spectrum disorder because I—

Jess: Oh, for sure.

Corey: —talk to other folks who have it, and their experiences sound absolutely nothing whatsoever like mine. It may as well be someone from another planet. And then I talk to other folks and some of it sounds alike, and I talk to other folks and it sounds like my mirror image. So I'm very reluctant to make broad sweeping generalizations about it. It feels on some level uncomfortable to talk about even if I get past the stigma aspect because it feels like at that point, I'm just talking about me.

And honestly, I'm not great at talking about how awesome I am because I don't want that to ever be the brand. It's super easy for that to come off as incredibly full of myself and sound like a blowhard. That's why my favorite joke whenever I have to mock someone is myself because, first, I’m not going to piss anyone else off, and secondly, I'm much more comfortable being the butt of a joke than I am getting up and sincerely telling people, “This is why I'm awesome.” Because I don't see myself as awesome. I really don't. And I have enough people that tell me otherwise, who I trust, that I believe, that I'm something different, but I leave the value judgments completely aside.

Jess: Yeah, it's definitely a fair statement to say that ADHD is an incredible spectrum of experience. For me, I think of it as a superpower and I try to think of it as a superpower for everyone, but that might not be the case. Because how it affects me is very individual. So, for me to say it's a superpower, I think it means that our brains just operate different from other people. And that’s great. I think that diversity is beautiful.

Corey: It really is. And we talk about, “Well, what about diversity of thought?” This is the kind of thing I equate that to in the non-horrible context. Yeah, it's the, yeah, different ways of thinking, different lived experiences, different perspectives on stuff; that's super valuable. One of the hardest parts I found in the two years that I was doing this independently, before I took on a business partner and a team was, I was stuck in my own mindset all the time, where I didn't have anyone looking at my code, for example, so I spent weeks beating my head against something relatively recently, and someone was looking at what I was doing and I told him what I was up to, and they asked one question in response and I got viscerally angry at it because it was—that just saved me an entire complicated thing with basically three lines of code.

Jess: Oh, yeah.

Corey: It's so good, it's awful. I'm angry I didn't see it. It's getting outside of your head and rubber ducking?

Jess: Yes, absolutely. My manager is my rubber duck for me, where I'll come with her to problems with, like, “I don't know what's going on. Why is this happening?” And she'll look at me, and she'll look at my process, and she'll be like, “Oh, yeah, give me a second. I think you just need to reframe something.” And then the problem is solved. And that bouncing off is so valuable for different types of brains.

Corey: So, before you were focused on workflows, you were in the world of infosec, which is fascinating to me, not least of all because it seems that the primary tenet that I've ever seen in infosec is be as difficult for people to work with as possible. What's the story?

Jess: Ah, that's been a long history in infosec, hasn't it?

Corey: It really has.

Jess: Yeah.

Corey: It's one of those things where there are two worlds of security one is what people say they do--like, “Security is job zero”—and then there's the reality of it, where, “Well, we built the slide, forgot security, and we're going to slap security at the very top and call it job zero like it wasn't an afterthought.” “Your security is important to us,” is what companies say right after they have publicly demonstrated that it very much was not. And so on. And then you wind up with a world full of vendors who are offering what is basically the same thing with different logos on it. And it's easy to become jaded and annoyed. Then you have the entire community which is, frankly, toxic in many respects and why I got out of that space—

Jess: Absolutely.

Corey: —and, again, if you're [unintelligible 00:14:44] disagree with it, and you're in the infosec community, feel free to email me if you can manage to put together an email that doesn't involve a bunch of profanity.

Jess: [laugh]. Oh, man—

Corey: I’m sorry, I'm being a little unfair, but not by much. I had some bad experiences in the infosec space, across the board.

Jess: No, it's fair.

Corey: So I'm trying to understand how you got to where you are.

Jess: I actually had very similar experiences with infosec, where it was so not beginner-friendly. So, much of it was, if you weren't immediately excellent, then it was incredibly difficult to break in. But the thing that really struck me about infosec was that it was so difficult to measure wins. Like it was so difficult to feel like I had ever accomplished anything because everything is a constant struggle. So, I loved working with users, and I loved working with every aspect of a business and helping them find their solutions in a way that worked for all of us, but my goodness, it's hard to feel like you get a win. So, I ended up switching back over to development, which is what I'm originally trained in because I figured at least I have incremental wins there. And that is easier for me to understand.

Corey: But you got out of infosec. And I'm not trying to cast aspersions; I’m genuinely looking to know, what drove your migration from infosec focused to workflow focused? Because yeah, sure, security is important—at least that's what we always tell ourselves—but as far as developer workflows go, I don't necessarily know that there's a clear path there that I'm seeing.

Jess: Yeah, my path was—so I originally got my degree in software engineering, so that might be part of it. I think what I like about workflow engineering, compared to infosec, why I would make that shift is that with workflow engineering, I still get the aspects of infosec, where I still have to safely handle information, I'm still touching information security and in that realm, but in a way in which there are deliverables, like, immediately. It's not a question of a constant struggle, a constant improvement; you never know when something is going to be finished. A client can say I want X, Y, and Z. And you can deliver that in a way that brings them joy. And I think that is the difference for me of why I made the switch. Because there's still empathy, there's still data handling, and safety, and working with people, and making something for them, but I also can see my accomplishments in a measurable way.

Corey: Forget dozens of visualization tools and view your entire system in one place with New Relic Explorer, the latest addition to New Relic One.

See your system-wide health at a glance with a dense hex view that has your hosts, services, containers, and everything else.

And get an estate-wide view of sudden changes, so you can catch issues before they impact customers.

So go to NewRelic.com, sign up for free, and start exploring your system today.

Corey: One question I have for you is if we take a look across the board of DevOps space, and the only way to get security built-in is to actually build it from the beginning. You don't get to bolt it on after the fact, and that's clear, so the workflow has to embrace that.

Jess: Right.

Corey: But darting back to our earlier point where we're focusing on the idea of empowering people who think in different ways, I look back at how software used to be built in the dark times of waterfall and whatnot. And I would have been a complete non-starter there in the way I'm only mostly a non-starter these days. But I have to wonder, is there some alignment in the way you see things between agile development or the way the architecture has evolved that embraces different patterns of thinking that don't involve, “We're going to make a plan and then for two years, we're going to follow that plan?”

Jess: Yeah, absolutely. The ability to break things into pieces and see things in different perspectives is very, very unique to current development practices. And it shares a lot of thought patterns with people with a variety of different brains. So, when you put a whole bunch of people on a team together with different kinds of brains, you all see things in different ways and that makes the process more efficient. And it lets us build things in ways that we wouldn't be able to see two years into the future. And we get to do that so quickly now. And I think that's amazing and very cool to watch.

Corey: Do you see that this maps differently into the microservices approach versus monoliths? I mean, I know it's not quite the same thing as going from waterfall to agile, but there are echoes of it, the way that I see the services architecture evolving.

Jess: Yes. The building of those two things, I think, takes very different forms, and how I can map my ADHD onto a map of microservices makes perfect sense to me. So, when I'm trying to build a monolith application, it's much more difficult because I have to fit all of these pieces in together that I can't quite visualize, whereas microservices very clearly parallels my brain. So, my thought processes are all over the place; so our microservices. And it works out really well for me; I can visualize that so clearly.

And that to me, makes development much easier. And I think it is easier to digest for a lot of people as well. Like we talked about the curb cut effect with disability and with neurodivergence where it's anything that makes life easier for neurodivergent or disabled people makes things easier for everyone, regardless of disability or neurodivergence. And I think we see that in development, too, where we can move things faster and people can understand systems better when they are broken up more clearly and it's less tangled.

Corey: I wonder if this ties into, I guess, a common complaint I have about microservices, which is people will ask, “Should we use microservices? I guess so because I saw some thought leader talking about it on a blog somewhere.” I'm kidding. It's not on the blog, it's always on a conference stage.

Jess: Oh, yeah.

Corey: Point being, if I want to deploy microservices in my environment. From my perspective, the reason to do that, traditionally, was, I have 5000 engineers all working on a monolith, and it turns out that the collective noun for developers is in fact, ‘Merge conflict’ so maybe we could find a better way of doing this and microservices absolutely solves that people problem super well. And then it looks like it's almost been twisted into parody, where you have five engineers working on 300 microservices, and I look at this and think that it feels like it’s sort of missing the point here. Now, understand that my own architectures are—I want to say intentionally bad, but no, they're hilariously bad. It isn't always intentional. I'm curious as to what your thoughts are on that particular alignment.

Jess: That microservices are just inherently better?

Corey: No. The idea of having microservices be an answer to the people problem, aligning with how you see workflows evolving.

Jess: I don't know if I would call workflows the solution to a people problem, partially because I don't think technology can solve people problems, necessarily. I think people problems are people problems and we need to solve them with people solutions. But the aspect of workflows is that we make those people solutions easier. So, if your people problem is too many engineers on a team, that's not something technology can necessarily solve. That might be an organizational shift, or what have you. But workflows can maybe make the engineers lives easier while they are still working on too many services. So, in that respect, I think technology can only aid people solutions to people's problems.

Corey: On some level, it almost feels like aspects of technology have lost sight of the fact that the actual purpose of this is to solve problems for people, and instead, it seems like some sectors are barreling onto how do we create new problems?

Jess: [laugh]. Have you seen the—it's a product, I can't remember who has done it, but it's a watch that tells your employer your mood throughout the day?

Corey: I did see that. I forgot who was making that, but my comment was that I was introducing the ‘Amazon Fire Me.’

Jess: [laugh]. Yeah, as someone with depression, I was immediately thinking, my manager is just going to think I'm sad all the time and I don't think I want that. [laugh]. But yeah, that's what it feels like to me. It feels like we are disrupting a lot of things with our technology—and I say that so sardonically—that, really, it takes a human connection sometimes. And that can be very difficult. But I think it's a skill that everyone should develop, and it's a skill that everyone should learn over time, having empathy for your fellow engineers, or anyone you work with.

Corey: I really think that there's an empathy story, and very often when people start talking about empathy, it is perceived by some of the worst people in the world as a form of weakness. Where it's, “Yes, yes. Be nice to people. [and 00:23:20] great.” Like, an example of this is if we go back to the Glengarry Glen Ross movie. There's this great monologue about coffee is for closers, and there's this whole rant.

And most people look at this with, “That's an awful place, I would get out of there, and I would storm out.” And they're right. But there's a certain type of person who watches that scene, gets fired up and says, “I’m going to go sell some real estate.” There are certain folks who want to basically sacrifice everything in pursuit of certain goals. And on the one hand, I do admire the idea of that single focus that occludes everything else.

On the other, it feels like it's a hell of a way to live, especially when it's for something as relatively shallow as getting a big paycheck. Now again, I understand I'm incredibly privileged when I say that. I don't worry about where my next meal is coming from and a lot of folks do. If you're in that position, for God's sake, yeah, get money, put food on the table, feed your family first, then eventually you get to the point of thinking aspirationally what dent you want to leave in the universe? I don't want to undersell that.

Jess: Oh, yeah. That's a very good point that financial and food security, and housing security, all those things kind of take precedent over psychological safety in some ways. We see tons of marginalized people, myself included in the past, that will have to compromise their own moral ground and empathy, to ensure their safety. That's for sure. But I think there is an aspect of empathy that still needs to be fostered in the realm of connection among other humans when we're talking about those kinds of work, too.

Rather than test driven development, I tend to do empathy driven development because that actually guides the design of my product. So, I tend to ask people what problem they're trying to solve, and how that problem makes them feel, and then I can actually address the root of the problem because nine times out of ten in my experience, whatever somebody is actually frustrated with is not necessarily the problem they said they're trying to solve. So, it lets me view their problems in a very human way that I think is very valuable.

Corey: The biggest problem that I see in the entire industry is that we always have this short term thinking, regardless of anything else. And we see it everywhere: in the, “To do, I will fix this later,” and you check and it's ten years old—when you use git blame. [unintelligible 00:25:38] ‘git shame’ because that's how we use it, but that's neither here nor there. A challenging piece to that as well is whenever I talk to companies about technical debt, there's always this persistent delusion that never gets corrected, that as soon as we finish the next sprint, then we're going to go back and do everything right; we're going to fix the technical debt; we're going to start making technical and architectural decisions that are not compromises; we're going to do it from a purist, absolute right perspective. And it never happens.

The few companies that actually intend to strive for architectural purity very often run out of money. Their codebase was beautiful, but they never found a business model. And it always feels like it's a short term thinking problem, rather than looking at the bigger picture. And I'm not trying to sound preachy, I'm not trying to get [unintelligible 00:26:21] from on high, but the only way, I think, personally for this to get solved, in some cases, is, one, for people to wake up, which, good luck, and, two, to have the workflows, and the tooling and the way that you approach things to align with doing things right becomes easier than not doing it.

Jess: Yes. I think perfection is generally antithetical to the human experience. Like we see so many people trying to be perfect and build these perfect workflows and tools when maybe that's not what we actually need. Maybe what we need is to feel good about what we do, and to feel more comfortable, and to feel like it's easy. And if that's the case, and if that's what my workflows can provide people, then I'm more than happy to build that work. Because I know that I'm making lives easier doing it.

Corey: When people say they want to start a company to change the world, like, that's really what I'd love to see the actual answer be. Like, “How are you going to change the world?” “I’m going to make people's lives easier.”

Jess: Yeah.

Corey: And a lot of companies will claim, “Oh, that's me.” And sure the intentions are there, but look at the externalities that generate from these hyper unicorns?

Jess: Yeah.

Corey: It's a problem.

Jess: Yeah. I remember in my interviews, actually, for this company, I remember, at one point talking with our leader of product, and we had a discussion about how documentation is in itself a form of empathy because you are writing for the future; you're not writing for yourself currently, so whether it's being kind to yourself or being kind to others, it is still an act of kindness and human connection to write documentation for other people. And that to me is, I think, what workflows are about, what people are about, what engineers are about, and I'm excited to see people really embracing that right now.

Corey: Yeah. It's optimistic for the future and I think that is probably the best note to leave this on. If people want to learn more about what you're up to, where can they find you?

Jess: I am on Twitter at @jessica_schalz. Warning, I am vaguely feral on that account. But if you're into hot takes and bad takes, I'm right there.

Corey: Excellent. And we will of course throw links to that in the [show notes 00:28:28]. Thank you so much for joining me today. I appreciate it.

Jess: Thank you. Thank you for letting me share some optimism and some joy with you.

Corey: It is refreshing to have, and it also just highlights how infrequent it is. Jess Schalz, workflow engineer at Transposit. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you hated this podcast, please leave a five-star review on your podcast platform of choice along with an inarticulate comment saying that no, that's not actually what diversity of thought means.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

Links:

  • Splunk: https://www.splunk.com/
  • Meanwhile in Security: https://meanwhileinsecurity.com/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is brought to you in part by our friends at FireHydrant where they want to help you master the mayhem. What does that mean? Well, they’re an incident management platform founded by SREs who couldn’t find the tools they wanted, so they built one. Sounds easy enough. No one’s ever tried that before. Except they’re good at it. Their platform allows teams to create consistency for the entire incident response lifecycle so that your team can focus on fighting fires faster. From alert handoff to retrospectives and everything in between, things like, you know, tracking, communicating, reporting: all the stuff no one cares about. FireHydrant will automate processes for you, so you can focus on resolution. Visit firehydrant.io to get your team started today, and tell them I sent you because I love watching people wince in pain.

Corey: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the Enterprise (not the starship). On-prem security doesn’t translate well to cloud or multi-cloud environments, and that’s not even counting IoT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IoT devices, detects these threats up to 35 percent faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at extrahop.com/trial.

Corey: Ever notice how security tends to be one of those things that isn’t particularly welcoming to folks who don’t already have the word ‘security’ somewhere in their job title? Introducing our fix to that, Meanwhile in Security. To sign up for the newsletter or to find the podcast, visit meanwhileinsecurity.com. Coming soon from The Duckbill Group.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Jesse Trucks who does, well, two things that we really care about. Let's start with the one that's easier to describe. Jesse, you are the Minister of Magic at Splunk. Thanks for joining me. What the hell do you do?

Jesse: Well, I'm a security specialist. But also people often ask me, “Well, what do you do?” So, I'm essentially an advisor or consultant to Splunk customers, and I help them get better return on their investment for their spend on Splunk products and services and so that they have better security.

Corey: Awesome. Better security; the sort of thing everyone claims is super important to them, usually right after it's been in the papers that security was not important to them, at least, not worth doing anything about. Security has always been a weird space, and you have a strong background in security. You were at Oak Ridge National Lab doing cybersecurity—not the physical security part around the reactors, but doing supercomputer-looking stuff, right?

Jesse: Yes. For a while, I was doing the HPC security for the National Center for Computational Sciences. Before that, I was a sysadmin for Anton supercomputers for D. E. Shaw Research. That was a lot of fun. And then I was a cybersecurity operations team lead for the entire lab’s cybersecurity team.

Corey: Again, but keep going back in time, but one other interesting thing that I'll hit on, and then I promise I'll get off this track, is that you were a D. E. Shaw Research. D. E. Shaw famously being the hedge fund that Jeff Bezos worked at before going to start Amazon.

And you were there 10 years, 20 years later, something in that range, in the D. E. Shaw Research Group, handling provisioning of their supercomputers. And that's great and all. It's one of those things where you look at a place like D. E. Shaw and you can say an awful lot, but one thing you can't say is that they hire dumb.

Jesse: That's true. I often say that—and this is fairly true—that one of the hardest days of my career was the one-day interview back-to-back with different people at D.E. Shaw, where you knew every 55 minutes, they decide whether you go home. And they asked some of the most interesting, realistic, real-world questions. None of us, like, “How many ping-pongs in a jet” crap, but instead it was, “How would you figure out the bandwidth available on a fiber link, and where a problem might be?” Here's a whiteboard.

Corey: Oh, I like that. It's one of those things that sounds deceptively simple but is so open-ended and you're looking for the trap because, an interview, and you want to perform well and… man… that's a fun one. I will not rat-hole on that question today but, man, is that a fun one to go into.

Jesse: It was good.

Corey: Now, you do all this stuff in security and have for a long time, but the other piece of this—and why we're talking about this today here—is, you are launching a combination newsletter and podcast here at the Duckbill Group called Meanwhile in Security, and people are sometimes thinking that I'm going to be the person that writes this. And sure, I can get up there and pretend to know things that I don’t. I'm a white guy in tech; confidently making statements I don't have facts on is basically what's expected of me at that point, right? But you actually know what you're talking about. And there's a reason that we have you telling those stories instead of me, but start at the beginning: what is Meanwhile in Security and why should anyone care?

Jesse: So, as we all know, one of the problems that we see is security is important. And there's all the obvious reasons: don't want to get hacked, don't get exposed, don't want your stuff shut down, blah, blah, blah. But the real reason the security is really important, and to be integrated into how operations are built from the ground up from—not afterwards, “Scan my new application,” or things like that, but built-in from the ground up, especially in cloud-native apps these days because everything is a surface, and therefore everything is so surfaced. So, how do you get more ROI out of your spend on your cloud infrastructure or your on-premise infrastructure? Do better security and do it up front.

And also, one of the things is, nobody knows how to talk about security except for security geeks. There's very, very few people who understand the nuances of security and the actual business impacts of what those things do, and what they are, and how they interact. There's very few people who are willing to go public and talk about having one foot in business and one foot in tech. I mean, Corey, that's how you’ve become successful: you can speak both.

Corey: Yeah, I can blend multiple areas of expertise into something forming a niche, which is great. I mean, the idea of finance and cloud computing is a small but important niche. Being able to talk to computers about security, and being able to talk to people about security is more important I would argue, and also not so much of a niche as it is a common business problem that every company needs to be able to traverse somehow.

Jesse: Yes, absolutely. And the message out there to a lot of people in their various organizations, especially in new DevOps-y kinds of operations is, “You must do better security,” but they're not security people. And security people generally aren't the DevOps-y IT people and so how do you get your SRE types to know better security? Meanwhile in Security.

Corey: I love the buried pun in there, personally, but that's neither here nor there. There are a lot of podcasts about cybersecurity, information security, what is the appropriate term these days? I don't want to deal with the community yelling at me because, honestly, I have no patience left for the InfoSec community as a whole.

Jesse: So, that really depends on who you're talking to. If you're talking to anyone vaguely associated with the federal government or works with the federal government at all, we've all given in, thrown in the towel, and we call it cyber, or cybersecurity. Cyber meant something totally different when I was coming up into computers, back when IRC was the only chat option. But these days old school InfoSec people call it InfoSec, or information security or just security but generally most people just call it cybersecurity because that's what the general layperson or the management people who don't know technology call it. So, really you can call it anything, and I'm going to say cybersecurity and cringe every time I say it.

Corey: Everyone refers to your industry a certain way. You smile and you shrug, and you put up with it. I mean, I was on the side of DevOps shouldn't be in a job title, but then I saw a lot of people paying an awful lot of money for jobs that had the word DevOps in them. And what, smile, nod, take the money. The problem I saw across the board was that there are two categories—very broadly—of folks who work with cybersecurity or InfoSecurity, we're going to call it—is InfoSec okay to call it? Or is that going to annoy people?

Jesse: InfoSec is fine. In fact, it's actually my favorite term. And I like CamelCase. So, it's capital I capital S, but that partially because it kind of irritates some people.

Corey: Oh, absolutely. Irritating people is one of my favorite things to do, especially when it's intentional. The promise—there are two worlds in InfoSec. One is people who are deep into InfoSec talking to other people who are deep into InfoSec. And then you have InfoSec-adjacent folks, where you're a DevOps person, or you're an SRE type, or you're managing a team, and you are very much on the hook for security being implemented properly—and yes, it’s a spectrum. I get it.

And it properly can meet a lot of different things to a lot of different folks—but security is one of those things you have to do, but it's not your core focus. And all the stuff I see in the InfoSec world, there are a bunch of great newsletters and podcasts in the space that speak security to people who are already deep in the weeds on security, and it's not accessible. I don't find that it resonates with me for whatever reason. The more I looked for something that didn't have that approach and instead talked about cloud security—which let's be clear, is security these days—in an accessible way, the more I realized I can't find it. Well, if I'm looking for this, maybe other people are, too. And thus began this our Meanwhile in Security experiment.

Jesse: Yeah, and one of the reasons for the title is because what happens is, is most of us have some job, and then there’s security. So, it's like, “Oh, yeah. Meanwhile, over here in the security land…” so—and again, the pun in case anybody didn't catch it, is Meanwhile in Security. And we call things ‘insecure,’ which always cracks me up because you think of a therapist for a computer. Although the origins are actually in grammar.

And the problem is, is that people have to figure out all the security things that are irrelevant to their job, but their job is not security. And on top of that, they're expected to wear five or six different hats, and the hardest hat is the one that they're focusing on right now. And then in 10 minutes, a different hat, it's a different hat. They're all incredibly hard jobs, and most of them have DevOps-y titles as well, even though really the biggest thing these days is SecDevOps. Have you seen that one yet?

Corey: I've seen DevSecOps, and I'm sure the ordering is important. And at some point, it's one of those how many words you're going to shove into a portmanteau?

Jesse: [laugh]. Exactly.

Corey: What about ‘QA’ as well? And oh, we got to get the word Kubernetes in there somehow because the CNCF demands it.

Jesse: [laugh]. Yes. And also, if people are searching for job titles, then they'll find what you're trying to hire for. And most people I've talked to actually have DevSecOp jobs or SecDevOps, or whatever way you're doing that. Generally what they are is they're people who have a strong programming infrastructure background who have learned some security.

And, you know, that's not always true. Sometimes the security people have learned IT stuff, or like me, you know, I came up through the ranks doing both IT Ops and SecOps together. And the problem is, is that they can't keep up with all of their industries that they have to; there's already too much. There's already so much bifurcation. For instance, you don't have somebody who is an amazing Linux engineer and an amazing Windows engineer, and an amazing Cisco engineer because that's too much to keep track of. They can be good at all of them, but they're only going to be amazing at one of them at best, and—with rare exceptions; there are prodigies—and so in security, same thing happens.

Corey: There's also a question of how do you position something? Because on the one end of the extreme spectrum, there's room for a show that explains InfoSec concepts to the layperson: use a password manager, here's what multi-factor authentication looks like and what it means. And then on the other side of it, you have folks talking about the actual seed algorithms you use behind those MFA devices. And there's an unmet need in the middle, which is, “Okay, I want to know how to approach that AWS account MFA story. Yeah, I get why MFA is important. You don't need to belabor that to me, but I don't care about the underlying algorithm. What I want to know is how do I handle the root account’s MFA device without having to put it in a safe somewhere that burns down or we can't get to in a pandemic, versus making it available to every person who works at my company? What are strategies real companies use around stuff like that?” Because everyone wonders. That's the sort of question I want to see explored. While simultaneously rounding up, what are the interesting stories that happened last week in the wide world of InfoSec that people should pay attention to?

Jesse: Yeah, I think that there's a combination of things that you just hit on. One is just understanding fundamental concepts. And these have to be repeated quite often because even people who are in InfoSec, we often get rusty in things that we don't work on on a frequent basis. And then you have people who of course, are not full-time security people, and you have new people coming along all the time. Like, “Okay, so what is ‘buffer flow’ exactly?” “Okay, well, it's this fundamental concept.” “What's a man-in-the-middle attack? Do people even do these anymore?”

And also, where are some resources if I wanted to deep dive and geek out on it? Where do I go? Which ones are reliable? What's the things that are most digestible? And the last thing is, how do I actually do this? Just functional things.

I guess the other really last thing is what has happened recently? What's in the news? And Meanwhile in Security, we'll cover these concepts, as well as big things that happen in the news and links to things that are current events that are relevant to how we do security today.

Corey: I've also said it before on the show, I'm going to be very transparent about this, I find the InfoSec community is largely a trash fire. And they haven't quite had the same experiences that a lot of the sysadmin community did as the DevOps movement was emerging, built around empathy. I found it so abhorrent in many cases that I got largely out of those spaces. There are good people in that community working for change. I get that, but it still has a weird lingering vibe of that, for me anyway.

And I can expect that some people are going to read Meanwhile in Security and say, “Oh, this is way too basic for my big brain exploration of how security works.” Great. That's not for you. It's also not for someone who is sitting at home and trying to make sure that someone doesn't break into their Facebook account, necessarily. It's not really for them either.

It's for folks who are operating in the world of Cloud slash modern computing—hybrid, of course, is a very real deal, too—and even in data centers, a lot of this stuff all is cohesive and needs a reasonable approach. But everything that even begins to speak to those people, it seems, is vendor-captured. It's by folks who have something to sell them, and magically, everything that they talk about, all roads lead to, “Buy my enterprise software.”

Jesse: Yeah. And that's one of the problems, too, is that, which vendors do you trust? Which vendors have a good account team for you to talk to? Who’s going to be frank with you today? And that changes and evolves.

And most of your main vendors out there, they're going to tell you for perfectly true things, but we're all biased in our experiences, and where we work, and what we're trying to sell. And what I want to do with Meanwhile in Security is to show people what happens when there is no bias. Because this is about what is good security, what is good methodology, and most importantly, one of the things that I've always done is teach somebody how to understand the basics because you know if the basics, you can derive an answer. You can re-derive if you can go back; if the how and the why, you understand where it comes from, then you can make your own decisions, your own informed decisions. Even if you're not an expert in something you know what kinds of questions to ask, you know what kinds of things to look for, whether you're asking that question of your systems, or of your vendors, or your coworkers or yourself.

Corey: I do want to talk a little bit about our sponsorship policy on Meanwhile in Security. We have something very similar here on this show, but with you, it's been much more explicit and drawn out. You don't know who's sponsoring any given episode or issue until you see it come out. We are building a hard firewall between you and anything sponsorship-oriented, which means you have no idea what vendor is going to appear above or below whatever it is you have to say that week, which absolutely could lead to really interesting stories. Theoretically, if I don't know, ‘vulnerability scanners are bad’ is your entire thesis one week, and ‘this episode is sponsored by a vulnerability scanner,’ that's going to be an interesting juxtaposition and we may have to make some apologies, but it's the way to do it because you aren't shifting your coverage based upon who's paying us. I never have either, but in your case, we're making it far more explicit.

Jesse: Yes, absolutely. And one of the things that it's very important to me, which is one of the reasons why I like working with you and your CEO, Mike Julian, is because it's a strong code of ethics, a personal code of ethics. And part of that is that I also come from sort of a tangential to a journalistic background, I worked at a newspaper; it's where I met my wife when she was a reporter and an editor there. And it's important to have the firewall between the journalistic content and the advertising content. I don't want to know who's advertising; it's not relevant to me.

What's relevant to me is me producing the best possible content I can to deliver the message to educate, train. I had so much incredible amounts of help in my career so far to give me the success that I have found, and it's my turn to pay it forward just like I pay it forward in other community things that I do with nonprofits. And that separation is extremely, extremely important to me, and I thank you for that.

Corey: Of course. I want to just be very clear, as well, that there is a vetting process that we put sponsors through. We don't usually talk about it directly with them because we politely decline to pursue anything further if they don't pass through it. But fundamentally, I take a look at their site in the abstract. And, okay, let me understand what this product does, and I dig into it a bit.

And it's a relatively low bar, but there has to be a problem that I can imagine or conceive of where, yeah, this is the right answer for that problem. If I think about that problem, and oh, this would be a terrible solution for that and I don't like anything about what they're doing, I mean, I can always make more money by talking to different sponsors. I can't recover credibility that I've lost by recommending garbage. And while a sponsorship is not an explicit recommendation, it appears next to my name. People are going to associate us on some level. So, I make it a point to not do business with companies I feel the need to apologize for on a consistent ongoing basis. With the possible exception of whoever names AWS services.

Jesse: Yeah, that was a good one. Yeah, one of the things I find really important is to always remember—and I'm glad you mentioned indirectly—at the end of every day, if your entire house burned down, your company went bankrupt, and there's nothing left whatsoever, all you're left with is yourself, your family, your friends, your reputation. That's it. All this other stuff can go. It can be replaced. And so standing on a code of ethics and putting your best face out there, and standing behind everything that you say and do is extremely important.

Corey: Some people need dozens of tools to visualize their stack. But not you. You’ve got New Relic Navigator.

All your hosts, services, containers, and anything else you can monitor are in one dense hex view, so you can check system-wide health at a glance.And you can sort by platform, service, app, or any other tag, which makes it easy to navigate your

entire system and see what needs your attention.

Go to NewRelic.com and get started for free

Corey: It's one of those areas where you can't regain trust after it's been lost. It's one of the things that's true in business; it's true in security. And people will forgive implementation challenges, but they won't forgive you attempting to rip them off, or you making trade-offs about their security so that you can do something you need to do faster. It comes down to messaging, and I think that that's something that gets very much lost. I mean, take two of the security sponsors that I've had on my existing shows, namely ExtraHop, and Lacework.

They're both solid companies that do really interesting things. In both cases, I've used demos of their products in my own AWS accounts, and I see the value of it. Oh, this is actually extraordinarily helpful. But I wanted to look at those things before I accepted it. Just because—especially in the security space, it feels like the RSA expo hall, is a never-ending sea of primarily the same company with different logos and different marketing phrases, selling what for all the world appears to be the exact same thing.

Jesse: Oh, and the security in any industry is bad, but—in terms of trying to figure out what's what, and who's the right vendor, and what things are available, but cloud security, it's a murky morass of unknowns because it's something that's not actually been explored enough, and there's a lot of great vendors doing really great work in SaaS offerings, as well as observability and security related to that, but you got to figure that out. Going into RSA or black hat showrooms, you just look at a sea of what do they all do, and why are they different? And that's the problem. You walk in there, if you don't know who they are, you don't know those vendors—it’s hard enough to figure out when you're somebody who does this all day long. So, I'm hoping to demystify some of that in terms of understanding the technology.

Then you can go ask all the vendors, all the questions you need to, to be able to find out what they're doing and why they're doing it. And especially moving into the cloud infrastructures, you know, companies who are doing the lift and shift, and going hybrid, and before they go cloud-native, it's really important.

Corey: We talked about people being on one side or the other of the security divide, but one thing I see, what makes that the most clear for me, is I look at various security vendors websites, and some of them explain in reasonable terms, what it is they do. And that's great. No argument here. I might, I don't know, argue about a word choice or something, or maybe there's a better story. Fine.

The others are incomprehensible to me. It is very clear, they're written with the CSO in mind who speaks in specific phrases around very specific compliance requirements, or very specific terms of art in the InfoSec field. It feels almost like a CISSP is a prerequisite for an awful lot of the marketing material to even begin to make sense. When I look at something and I don't understand it, I feel like I am not the target market, which on some level, I cheat and just take a shortcut to, “And therefore it's probably crap,” which is unfair. But it's very hard for me to see how those offerings are being articulated to the people who could actually benefit from them, assuming they're good.

Jesse: Well, in a sense it’s the same problem that you solve is, I look at an AWS bill and it just baffles me. It's like, “Well… do I start drinking now or later?” [laugh]. And instead, there's somebody who's an expert in that, that can help somebody get through that path, and I hope to be a little bit of that guiding light. And one of the things I think is important for people to—and you mentioned it earlier: people need to remember security and cloud security are one and the same, but there's an added layer.

In a data center, you can lock the door, and you can put really, really, really fancy good locks on it. It will really raise the bar up high to barrier to entry. So, somebody has to be able to understand physical security, get through, get into your data center, or they can go through the network. So, now you can secure the network; all of your normal security things apply. In the infrastructure you have on the cloud, there's all the cloud infrastructure that is remotely accessible, so your entire data center, effectively for you as the cloud consumer, you have all that extra layers of security. So, security all is the same, plus you have to deal with the cloud security. It's really, really important, and in fact, it's more important to have cloud security infrastructures have better traditional security mechanisms than it is for traditional on-prem solutions.

Corey: I argue on some level, there's a bit of alignment there where you can save an awful lot of money by turning off things you're not using with the simultaneous argument that it's really hard to exploit something that's been turned off. Not impossible I’m told, but still harder to do than most. And there is a spiritual alignment of worshipping at the altar of the Church of Turn that Shit Off that I feel security and cost [laugh] folks can align on.

Jesse: [laugh]. Absolutely, and especially when you're starting looking at containerizing things and cloud-native apps, you don't have to secure something if you never write it.

Corey: Oh, yeah. And another approach that I find that is very similar from the security world and the billing and governance world is you talk to companies across a wide variety of industries, levels of sophistication, et cetera, and no one believes that they've gotten their cost attribution right, or they're billing exactly where it should be, or their security posture where it needs to be. Everyone believes that they really have a lot of work to do to catch up. No one looks at this and says, “Oh, I've nailed this. I'm doing great,” except for people who are painfully naive.

Jesse: Yes, absolutely. And one of the things that it comes up often, is compliance. Compliance, compliance, compliance. Is it PCI? Is it ISO this? ISO that? Is it NIST this? NIST that? Is it FERPA? FISMA? HIPAA? Or things like in Europe: GDPR, which is actually affecting all of us—in the US and everywhere else in the world because we all have European citizens looking at our things or using our products and services, or we have people who are working and living there.

Corey: Yeah, that is a weird law that is, frankly, a little, I want to say overbroad in some respects. I care deeply about privacy and security. This stuff is important. I built a custom link tracker for my newsletter, for example, that will tell me the number of unique newsletters that were sent out, that have been opened, and which links within them have been clicked without ever telling me what you have opened, or you have clicked on because I've got to be honest, I do not care and I struggle to find a way to care less. Because I want to know of all the links I sent out last quarter, which were the ones that were the most appealing to the point where people clicked on them and dove into them.

That helps me build things like ‘best of’ issues. It lets me look in the aggregate and say, “Okay. I don't care about IoT, but the rest of the world seems to so I'm going to have to continue to cover it.” Versus, “I don't care about Windows stuff, and it looks like the readership doesn't either. Thank God. I don't have to talk about it much.”

It helps me inform opinions around what I should write about because I do have an audience; I want to appeal to them. But being able to run a little bit of analysis on that in a privacy-preserving way is super important to me. But all of the tools that you can look at for sending out newsletters stuff is deep in the weeds of oh, and then we can track them across the internet if you want. No. I want you to not do explicitly that.

A lot of stuff that Basecamp with DHH and their hey.com stuff appeals to me at a deep level. I feel like every time we have any form of tracking, it should have above it in bright letters, “We are going to be watching what happens with these links,” because that is fundamentally what it means. And most of the time, we don't need that, so why do we do it?

Jesse: Oh, absolutely. And privacy is one of the things that is near and dear to me, too. I'm also the principal of a small nonprofit school that I founded with my wife and another woman, and I want to make sure that the things that we use and the services we use at that school have no school information. I'm okay with somebody knowing that this customer did XYZ as long as it doesn't tie back to individual information. Otherwise, we have FERPA violations, et cetera.

And so, compliance has always been important to me. I've worked in support of just about any kind of compliance you can imagine—including in classified spaces—before. And I think that that's one of the morasses that's really, really confusing is, how do you do really good data-driven business, but yet maintain privacy of both your employees and your customers?

Corey: That's really what it comes down to is you need to be able to have a functional business. We do a best effort job on our side to comply with it, but to dot the i's and cross the T's on this would effectively require a full-time person whose job is nothing other than maintaining a lot of these compliance systems. And I mean, again, we have a list of what information we have about individuals who subscribe to the newsletter, for example, what IP address did they sign up from? We capture that because I believe that's a requirement. And, “What PII do we have on them?” Well, in our case, we have their email address. “Well, why do we need the email address?” Needs to be a valid reason. “Because it's hard to send you an email if I don't have your email address.”

Jesse: Yes.

Corey: Okay, and did they click the ‘Confirm’ option on the email that we send out to validate that this really is their email address? Yeah, without it, they never show up at all. And great, then we have a record of what was sent to them at various times, and when they unsubscribe that shows up, too. There, we have now disclosed everything we know about what you do, tied to you.

Jesse: And I think it's important to point out that that is actually harder to do than you just listed out because—

Corey: Well, what I just talked about is the ideal. What we actually have—because ESPs don't really offer a way to disable this—I can look at what newsletters you've opened, I can look in their systems at what links you have clicked on. And that is a problem, and there's no good way to disable that with our current ESP. But the saving grace is that their interface is so terrible for looking at these things that no one ever does. And we restrict access to this, obviously. It's not useful for us, and it definitely goes to show that okay, well, what's the value of all this surveillance? The answer here is, not much.

Jesse: Also, I think there's a different aspect of it, too, is what is actually legislated compared to what has been tried in court and proven realistic. I remember when I worked at AT&T—it was SBC at the time, pre-merger—and my team was part of the first SOX audit—Sarbanes Oxley act. We had no idea how bad things would be, we didn't know if we're going to get in trouble, so we did all this work internally. And then when we finally got to the official external audit, turns out, things were just fine because we were paranoid about it for nine months.

Corey: Yeah, it turns out that it's super easy to wind up inadvertently setting a foot wrong. And I looked at all of this when it was first coming out and on—again, we did a best-effort attempt to do all of these things, and we didn't do the dumb things, like assuming that an IP address is indicative of whether someone is considered an EU citizen. Great. Yeah, that doesn't actually work. Who knew?

And again, we weren't tracking all this stuff to begin with, but it's just one of those areas where, yeah, I want Facebook and Google and large companies that are doing tens of millions of dollars a year in ad business—or more—to absolutely have to do these things. But this stuff also applies to rando hobbyist newsletters with a Patreon donation link in them. And where do you start, and where do you stop?

Jesse: Yeah, and a good example, there are, too, is school regulations. I have an actual school in the state of Tennessee. And it's a legal school. We have a full-time teacher and we're subject to all the same regulations, and laws, and inspections, and data privacy for the FERPA act, and things like that. And it doesn't matter that we have eight students in one classroom.

Doesn't matter. We have the exact same burden as a large-scale district with 10,000 students. And so, how do you implement that in a way that's reasonable? Same thing with HIPAA. If you're a single doctor, and you do all your own bookings and you have one room you use, you know, like a therapist, for instance, you're still subject to the exact same requirements, as Kaiser Health.

So, how do you navigate that? I'm hoping that we can help solve some of those things in—if you think about it in Meanwhile in Security, we’ll take you to the side and have a little fireside chat about how you can figure that out by knowing the requirements and how they are in both at a legal and spirit of that law. And that's how you learn.

Corey: That's really what it all comes down to is learning; learning as we go. And that's the goal of this podcast. Some weeks, we do a better job of teaching folks things than others. At least if the takeaway here this week is sign up for Meanwhile in Security at meanwhileinsecurity.com, but also the newsletter.

It's about educating people about what's going on in the world. AWS says that re:Invent is a conference primarily about education, and I get that. I can see an angle from which that is true. The problem with re:Invent is it has no idea what it wants to be, and thus no sense of itself, so it tries to be all things to all people at all times. And it turns out, that's a really hard list of boxes to check.

Jesse: Absolutely. That's why you should pick a narrow focus, and go for it, and do it right.

Corey: Exactly. And in our case, we're doing that with explaining security to practitioners who have other work to do.

Jesse: And we're going to have a little fun.

Corey: We certainly are. Where can people sign up, once again?

Jesse: meanwhileinsecurity.com.

Corey: And of course, the Meanwhile in Security podcast is also available wherever you get your podcasts from. If it's not there quite yet, let us know. By the time this airs, it should be everywhere, but it seems like there's a new podcast platform every week and it is increasingly difficult to catch them all. Let us know. Jesse, thank you for taking the time to speak with me. As always, it is appreciated.

Jesse: Thank you very much, Corey. I love talking to you. It's always a good time and I look forward to Meanwhile in Security helping others learn better about security.

Corey: Jesse Trucks the new author slash editor of Meanwhile in Security. I'm Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you've hated this podcast, please leave a five-star review on your podcast platform of choice, along with an angry and insulting comment about why I truly don't understand the true meaning of InfoSec, if you can manage to get the comment posted through the filters against racial slurs.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Kevin

Kevin's the Lead Product Owner and Engineer on Ducktools, a recently-released set of AWS Cost Management power tools.

Links:

  • Stop Lying Cloud: https://stop.lying.cloud/
  • TabDB.io: https://tabdb.io/
  • Personal website: https://kevinkuchta.com/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by LaunchDarkly. Take a look at what it takes to get your code into production. I’m going to just guess that it’s awful because it’s always awful. No one loves their deployment process. What if launching new features didn’t require you to do a full-on code and possibly infrastructure deploy? What if you could test on a small subset of users and then roll it back immediately if results aren’t what you expect? LaunchDarkly does exactly this. To learn more, visit launchdarkly.com and tell them Corey sent you, and watch for the wince.

Corey: If your mean time to WTF for a security alert is more than a minute, it’s time to look at Lacework. Lacework will help you get your security act together for everything from compliance service configurations to container app relationships, all without the need for PhDs in AWS to write the rules. If you’re building a secure business on AWS with compliance requirements, you don’t really have time to choose between antivirus or firewall companies to help you secure your stack. That’s why Lacework is built from the ground up for the Cloud: low effort, high visibility and detection. To learn more, visit lacework.com.

Corey: Ever notice how security tends to be one of those things that isn’t particularly welcoming to folks who don’t already have the word ‘security’ somewhere in their job title? Introducing our fix to that, Meanwhile in Security. To sign up for the newsletter or to find the podcast, visit meanwhileinsecurity.com. coming soon from The Duckbill Group.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by my colleague, Kevin Kuchta. Kevin, welcome to the show.

Kevin: Hey, Corey. Great to be here.

Corey: So, you were the lead product owner and only engineer on our internal DuckTools suite of projects, which we have recently announced that we are going to be sunsetting. So let's start at the beginning; what's the deal with that?

Kevin: Yeah, that was a fun ride. DuckTools was a suite of tools that we were trying to build around AWS cost management. The plan was to sort of have them as tools that would focus very specifically on individual problems people might have in the Cloud. So, you need to figure out what your AWS savings plan commit level should be. We have a really cool tool that tells you exactly what you want to do for that. We tried to make the set of tools work, and we built some really cool stuff. But also we didn't find product-market fit, and so now we are sunsetting them. That's the sort of super-high-level view of that.

Corey: Oh, yeah, in many respects, I basically take ownership of that because this whole suite of tools idea started years ago when it was just me. And ‘tools’ is kind of a lofty term as you can well attest. It was a bunch of crappy shell scripts that I tied together that took every good programming practice and tossed it out and then took what was left and made a muddling mess of it. But on a good day, I could run these things on an account and get some semblance of an answer, which at least gave you a direction to go in. It was an initial, “Oh, cool. So, what is this thing supposed to be doing? Great, okay, I can build that thing, but I'm starting from scratch, and let us never speak of this again.” That's the direct version of our conversations because again, I'm not good at computering things.

Kevin: Yeah, I remember you reaching out to me much earlier, in 2020, as sort of a Slack message in some joint Slack we were in, saying, “Hey, we've got these internal tools. You're looking for work, want to come work for Duckbill?” And that was a lot of fun. That's how it all started.

Corey: In fact, we first met 10 years ago now when we work together at the same startup, and we've kept in touch ever since. And you've always been top of my list for engineers to work with just because, on the one hand, you've got the technical chops; that is undeniable for anyone who has even seen some of your ridiculous projects, which we'll get into in a little bit. But you also in a relatively uncommon way, are able to grasp business nuance, understand the context behind the larger picture of why something is being done in a way that very often engineers like to, frankly, ignore. And finding both of those skill sets that are incredibly valuable in the same person is one of those, let's grab on to that whenever we can find it, and hold on with two hands in a death grip.

Kevin: Well, thank you. I've been equally looking forward to working with you and Mike again, at some point, for as long as I've known you. Mike's brought me on for a few random contracting gigs over the years, and of course, I've worked with you both on some silly side projects and Expensify, way back in the day.

Corey: Oh, yeah. One of my favorite side projects of yours still remains the Lambda URL shortener that you built, which, okay fine, doesn't sound super interesting until you realize that there is no data store. It is just Lambda, and it works, and it's a monstrosity.

Kevin: You know, actually, you can take some personal credit for inspiring that. I was reading your newsletter, and I saw an article on creating a URL shortener with Lambda. And I thought, “Oh, wait, how'd you do that? That sounds amazing.” And I looked into the article that you'd liked and it turned out they were using Lambda and DynamoDB.

In the instant between seeing the link and clicking on it, I had some ideas as to how it could be done with self-modifying Lambda functions. And that inspired that whole mess. It was one of my more popular projects. It's horrible; you should never do it in real life.

Corey: Oh, yeah, the outreach from the Lambda team was great. First, they’re like, “We want to talk to the person that did this.” And followed immediately by, “We never want to see that person again. We'd like to salt the earth now that we've seen the code.” I'm exaggerating only slightly, but it was definitely a stop-and-take-stock moment of how to approach those things.

And you've got a bunch of other stuff too. You were instrumental in the first version of implementing the stop.lying.cloud status page cleanup project. You took a Chrome extension and then wound up porting that into Lambda@Edge, and it wound up working dynamically on the fly.

Somewhat recently, I wound up modifying that to be much more of a blunt instrument approach. There was some significant latency in it gathering the original seven and a half megabytes of HTML that is the AWS status page, and then dissecting the whole thing and making the modifications I wanted to make and then returning it. And the blunt instrument approach was, this is a status page; we just update it once a minute and render that HTML to a static site, we don't even need to use Lambda@Edge. And the best feature of Lambda@Edge is you are not required to use it. And then it became something we could just hurl up, and it was a very simple serve this static website, disable all caching on it, and it's good.

Kevin: Nice.

Corey: Yeah, it was fun. But back to the whole DuckTools part of the story. The thing that I find so interesting is that I still maintain that there is a use case for every tool that you've built and every tool that we were planning on building. The problem that you alluded to and I want to dive into more, is that when we talked to a bunch of customers and prospective customers, the things that they wanted were not the things that we were building, and I'm never a big fan of having to educate your market and then trying to sell them something.

Kevin: Yeah, I think that's exactly right. It was an interesting experience because I had never done serious customer research interviews before like I was doing here. As the product owner, I was on calls, dozens of calls with people who might potentially want to buy DuckTools and just sort of probing everything I can think of to hear about the problems that they have, the challenges they've got. And I ended up putting together this big doc a few weeks before we shut down, and it was just a list of, here's all the problems people have, and none of them quite lined up with the things we were trying to build. And all of them were problems that, yeah, we could solve them if we were willing to hire two to five engineers and spend four to six years working on it. And we just decided that it wasn't really a good fit for our business. I think we were joking internally, that a lot of people, what they really wanted was Cloudability, or CloudHealth or CloudForecast, but cheaper.

Corey: Right. The entire problem is we're going to pay a percentage of our bill, it doesn't work in some respects because it means that companies grow out of being able to justify the tool that you're providing them. And any other approach that we've looked into is sort of a mess, too because you're leaving an awful lot of money on the table if you just charge a fixed fee per customer. And again, we're not a VC-backed, so that's okay for us. We don't have to capture all of the value, but we do want to wind up charging reasonably for this.

Kevin: Yeah. Finding the right price point was definitely tricky. We've talked to people who would say that, “Oh, yeah, 50 bucks a month, is about the right price point for an incredibly valuable tool that solves all my problems.” We've talked to other people who would say, “Yeah, this tool is going to save me tens of thousands of dollars per month. I’m happy to pay a couple thousand a month for it.”

Corey: yeah, some of the stuff we built too is, “Oh, what does this do?” “Well, it looks at all your DynamoDB tables, looks at the historical usage over a configurable period of time, and it spits out whether they're currently on a on-demand basis or provision capacity basis, and is there an optimization to be made for toggling some of those in some cases, with certain constraints built in, and certain guidance that is assumed or can be configured?” And the typical answer when I described that sort of tool to someone who's in this space, or more commonly not in this space is, “Well, that doesn't sound hard. I can build a script to do that over the weekend.” And there are a bunch of tools that do similar things on GitHub.

And they're all limited and don't do it correctly on a wide variety of different bases here. It's nuanced and challenging. What we fundamentally wound up building internally is a set of power tools for folks who are steeped in this space to use to become more effective at analyzing what's going on in the AWS bill. And that is a difficult thing to wind up selling. I would prefer to it internally from time to time as ‘Photoshop for the AWS bill.’

I mean sure anyone can buy Photoshop—me for example—and does that mean that I can build professional-looking graphics? No. It means I can make graphics that looks like I used a tool a professional might have used had I had the good sense to pay them. I'm not an artist and no tool is going to change that. And that was part of the problem that I saw us running into as well.

Kevin: Yeah, that's definitely true. And our original vision for this product was to be Photoshop for professionals trying to solve these problems, whereas a lot of competing cloud tools that charge an arm and a leg, they instead try to sell you a painting, and try to say, “We will solve your cloud cost problems. You don't have-to-have in-depth knowledge of this sort of stuff.” And I think we were trying to differentiate ourselves from those products by saying, “No, no. We're going to offer you a scalpel; we're going to offer you a really sharp tool that will do exactly what you want.” And unfortunately, the people who want really sharp tools are also the people who think that they can solve these problems with a shell script that they wrote last weekend, or with a free GitHub product. And so selling to those people was kind of tricky.

Corey: And traditional SaaS models of subscriptions and whatnot are also challenging in this space. For example, one of the tools that we built—you built and is working—it analyzes which reserved instances a customer has that are expiring, and then calculates out what the savings plan commitment should be so you don't have to wind up waiting for a week of on-demand so that AWS’s native analyzer can look at it and finally make a recommendation that may or may not fit. Instead, it winds up letting you figure this out in advance so there isn't a gap where you're paying through the nose. And that's great; it's helpful, but it's generally used only when there's a batch of reserved instances expiring. In some companies, that happens once a year. People are going to need that tool once, they're not going to need it on an ongoing, sustained basis.

Kevin: Yeah. We had that problem with both of the tools that we put out for our beta: the tool that you just described for migrating reserved instances and savings plans and our tool for just picking savings plans commit levels. People would say, “Oh, that product looks great. Can I use it for a month to pick a savings plan commit level, then not worry about it again for three years because we just bought a three-year savings plan?”

Corey: Right. The savings plan calculator is something that we've been using for a long time because there's nothing out there that does this. If you ask the AWS system what savings plan commit should I buy? It's going to come back with a recommendation that is, okay, it's directionally correct. But here's the thing that I think they lost sight of: for large accounts, that's going to come back and say, “Great. For the best savings, go ahead and click here and commit to spending $24 million.”

And then—I'm not kidding—there's an add to cart button. No one on this planet is going to say, “Okay,” click the button and hit buy without first asking a few other questions. Namely, they all boil down to scenario modeling. “Well, this past month has been weird because, I don't know, it was a holiday season, and now we're dropping back to baseline load.” Everyone's favorite lie they tell each other.

So, then it was, “Okay. What do we go back three months and look at that timeframe? Great. Oh, okay. That's a much lower number. Cool. What if we only bought half of that? How much would we save?” And the answer is, of course, some money. How much money and the answer of, [growling], “I don’t know,” doesn't work when you're asking someone to spend eight figures, as it turns out. So there's a lot of nuance and scenario modeling and helping to justify the recommendation because no one wants to be left holding the bag on that kind of mistake.

Kevin: Yeah, exactly. I think you described the use case perfectly. We had one internal customer we were trying it out with, and their spend was incredibly seasonal. By seasonal, I mean daily. Their spend would increase about 8X during peak hours during the day versus its lowest hours in the night.

And so they were looking at their spending in savings plan calculator, and they said, “Well, you know, we want to take these big peaks every day, and we're going to put that all into spot instances.” And so they wanted to be able to figure out, okay, what's our savings plan if we only look at the baseline below these peaks. Of course, Amazon's built-in tools will not tell you anything about that and our savings planning calculator was really great for that. I still believe that there's a pretty good value that this tool provides, if only we could have found a way to attach it to a useful business, or a viable business.

Corey: Right. And the challenge, too, is the person who really cares about getting that number dialed in specifically, is very often not the people that we're speaking to in most other contexts. It's very often someone in FP&A, Financial Planning and Analysis. It's someone who's being told to model additional scenarios that their business just wants to make sure they check off the list. And that's all fine, there's a need for that and it's great, but it doesn't lend itself to the audience that we're talking to.

And further, these things also don't really lend themselves to our current sales model of having one of our account folks reach out and talk to people as they move through the process. It turns out, you can't have a high-touch enterprise-style sales conversation when you're trying to sell something that is a reasonably affordable SaaS.

Kevin: Yep. And that was actually a sort of a misstep I made early on when I was trying to sell DuckTools. When I first joined, I sort of assumed, “Well, they have this great sales team. We'll just have them sell DuckTools along with consulting engagements.” And as it turned out, that didn't work for the reasons you described.

And it took me, I don’t know, three months to realize that and at that point, we sort of switched our model. We decided I'll be doing the sales directly, I'll start talking with customers more directly, and that sort of got us on this whole trip to realize some of the flaws with our business plan. Because until that point, I hadn't been talking with our customers nearly enough. That's a piece of advice that every product development book will tell you: talk to customers more. And I wish I had taken that advice a little earlier.

Corey: This is one of the weird things about starting a business that I've learned the fun way is, you hear all the tropes and all the advice that people give, and you think, “Ah, I'm not going to make that mistake.” There's a reason everyone winds up making that mistake. It's like that old joke you see going around with a comment of, “Below this line is madness; no one understands it, but you're going to tilt at it anyway. When you're done, please increment the next line so that it remains accurate.” And the next line is some high number next to hours of people's lives wasted on this. It's the same model where everyone is going to make the same type of mistakes sooner or later, and sometimes the only way you learn is by the hard-won experience of getting it really wrong.

Kevin: Yep. In this case, I think that the best-case scenario, if I’d talk to customers sooner than I did—talked to them more and sooner—the best thing that would have happened was we realize that this product isn't viable a little bit earlier. I don't think we could have turned the product around or just suddenly figured out that, “Oh, yes. Here's just one strategy that will make DuckTools perfect.” So, at worst, we lost six months of time. It's a bummer, but I certainly learned a lot doing it.

Corey: We learned an awful lot. There are things that we'll carry with us, absolutely, as far as just some of the IP we generated along the way, but there's also the experience we get that will inform future product decisions as we look at various aspects of it. This is a constant struggle on some level here. People come up to us and ask us to do all kinds of things, and if you're in a mindset of trying to say yes to everyone, you'll wind up going in circles. “Can you expand to doing Azure instead?”

You want to say yes to that, but we're not really equipped for it. “Hey, how about you go ahead and do an implementation project for us, just, it'll be quick. While we have the time and you've got the expertise.” Down that path lies ridiculous problems. And it's difficult to say no to that, especially if you're in the early stages where you want to say yes to every customer because there's a mindset you have to grow out of that every lead might possibly be your last one ever. And it takes a certain level of repetition—at least for me—to get past that.

Kevin: Yeah, that's definitely true. One of the other factors for why we decided to shut down DuckTools was that it was sort of a bit of a distraction, in that Duckbill has three distinct business lines: it's got the consulting, it's got the media arm, and it has this SaaS tool that we're building. And trying to build three different lines of business at the same time, it reduces our limited pool of focus, especially for the CEO, Mike, who has to deal with all of these things at once. And so yeah, when you say yes to everything in an enterprise business, you end up scatterbrained, jumping at a whole bunch of different opportunities to the detriment of any individual one. And so I think by shutting down tools, we're going to be able to focus a lot more on the media wing and the consulting wing.

Corey: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the Enterprise (not the starship). On-prem security doesn’t translate well to cloud or multi-cloud environments, and that’s not even counting IoT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IoT devices, detects these threats up to 35 percent faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at extrahop.com/trial.

Corey: Yeah, the ability to turn things off was sort of the driving issue in some respects because as my business partner Mike and I continued to look at the business and reason about it for the last six months or so, it seemed more and more that we were really running three very different businesses instead of running only two very different businesses. And it seemed like it was very much the, here are three things: good, fast, cheap; pick two, and you can never have all three. It felt that way with our lines of business. We can come up with a bunch of things that benefit two areas, but not the third. So, it was always a bit of an exercise in frustration, given how we reason about our own business, and not having to account for a SaaS product as a part of all of this does free us up somewhat significantly. It also means that I can be much more intentional with how I wind up directing the famed Duckbill Group’s spite budget.

Kevin: True that. I remember specifically, not that long ago, we were looking to hire a marketing director or something along those lines, and I was talking with Mike and he would say, “Yeah, we can find someone who is good at selling SaaS products and someone who's good at selling consulting engagements, but we can't find someone who's also good at those two things and being able to sell sponsorships for the media wing.” Finding someone who could do all three of those things very well at a high level is basically trying to find a unicorn. But finding two, maybe we could do that.

Corey: Oh, absolutely. And remember, I started off doing all these things myself. And that doesn't mean that I'm this magical, incredible unicorn myself, it means that I am passively mediocre at a bunch of different things. And it does lead to the humbling experience where every single person that I hire is, by definition, better at the thing that I'm hiring them to do than I am.

Kevin: You know, I once spoke to a CEO a couple startups ago that I was working at, and he said that the key jobs of being a CEO, or any founder, is constantly firing yourself. Because yeah, you end up doing a little bit of everything, and you find yourself in a lot of interesting jobs over time. [laugh].

Corey: It's made hiring a fascinating challenge. I mean, when it comes to hiring our cloud economists, yeah, I absolutely know what profile I'm looking for, I absolutely know what skills they can learn, versus which they need to come in with, and I'm very competent at sussing those things out in an interview, but I have no clue how to hire someone who's a marketer, for example. And same story with hiring a salesperson. And a common trap that people fall into when they're in this position, that I'm thrilled to pass on—please don't make this mistake yourselves—is you just wind up trying to wing it, and as a result, you invariably wind up hiring the person that sounds the most confident. And that is a dangerous thing. My way out of that trap is for those roles I don't understand super well, is finding a high-level person in that space whom I can trust, and then paying them as a consultant to help me find the right people.

Kevin: I think that's exactly the right approach to take. A couple startups ago that I was working at a company, and that's pretty much the approach they took. I joined as the fourth employee, and so we spent a lot of time hiring our first sales hire, our first marketing hire. That guy was the second engineer, so at the time, they were also hiring their first engineers. And that was the approach they took: they had hired these outside advisors, ideally, people in their network or just people who are very well-positioned, to try to come in and interview people and help make that decision for them. For that matter, I've always thought that if I ever found myself in the position of hiring a first ops person, you and/or some of the other Duckbills will be first on my list of people to call to help me interview people.

Corey: Oh, yeah. On one hand, I probably should have had a specific focused senior developer come in and have them interview you to suss out your technical skill sets. Now, there are a few problems with this: one, you would be the person I would reach out to you to fill that role, so kind of a problem. Two, we've been working together on and off for a decade, now. I'm not going to suddenly magically figure out in a half-hour interview that you've been faking it all along and don't actually know how to code.

That is ridiculous. And thirdly, this role was so weird and strange and required a combination of things that you are great at all at once, it would have felt incredibly weird to bring someone in that we didn't have the longstanding working relationship with because of the positioning would be this is a complete gamble; there's a terrific chance that this is going to fail, and take that on faith. And you did because we have that shared context, it was an easier conversation, whereas someone who doesn't know me, it sounds honestly like a terrible job.

Kevin: Honestly, this job is pretty much exactly what I was looking for. My curse as an engineer is that I love being a generalist. I love working at every part of a stack. I love being able to build entire features, and entire projects, and entire products from scratch. And it's kind of hard to find that when you're just out interviewing at random dev jobs.

Everyone wants you to be a front-end engineer, a back-end engineer, and an infra-engineer. And, no, I want to build whole features, I want to build whole things. And so when you and Mike came to me and said, “Hey, build this entire product from scratch,”—not just engineering, but also sales, and product, and strategic planning, I thought, “That sounds exactly up my alley.” And it was a lot of fun.

Corey: Yeah, it's one of those weird edge case stories, and for very specific roles like that you almost have a specific person in mind when you're reasoning about them. Now, the problem, of course, is that you carry this to its logical extent, that's how you wind up starting a company and the first thing you do is you hire all of your friends. And invariably, there wind up then being significant diversity issues down the road. And then people step back and think, “Oh, okay, now that we've hired 20 white dudes all named Brad, maybe it's time we look at hiring something different like a, I don’t know, white guy named Steve.” And at that point, it's almost too late because that sends a signal—a strong one, and not a good one—to people who are considering working there.

But by the same token, it's also a very hard type of thing to open up to the outside world. “Hey, do you want to take a massive gamble on a thing that probably won't work out and will leave you looking for a new job in six months?” And that's a hard thing to sell to someone if they're not already conversant with who you are, and how you operate.

Kevin: Yeah, it's a tricky problem to solve. If you're starting a new company, you want to hire your network because those are the people you already have some amount of confidence in. Networks tend to look like the person who they are centered around. And so, at some point, you need to get away from hiring just from your network. How do you go about that? I think probably the best approach if you can is try to build a more diverse network; try to have more friends who don't look like you.

Corey: It's hard to do, and I don't know that there's a terrific answer here. So what's next for you, as of the time that we record this but the understanding that there is always a production delay, and some of this may be out of date.

Kevin: That's true. Looking for another super-early-stage startup. Before Duckbill, I had joined a 400-person startup and they paid me a lot of money, but it wasn't quite what I was looking for, and so when I joined Duckbill, I was really excited to get back to a 10-person company where I could wear a lot of hats, and have a lot of influence, have a lot of impact, be able to spend a lot of my time just building. And so I think that's probably what I'm going to look for next. Maybe it's a VC-backed startup that's five people, or maybe it's a bootstrapped business that's going to be a bit more sustainable. But either way, I think it's going to be pretty small, and probably involve working with as much of the stack as possible.

Corey: I'm curious that you're looking for specifically for an early-stage startup. Tell me more about that because that can mean some good things and some very weird things. Because if I look at this, the most cynical possible lens, what I hear you say is you're looking for a very small team that doesn't really know what it's doing yet, where the founder’s personality quirks are now business problems, and the runway is very uncertain and shaky and keeping your resume updated on a weekly basis is top of mind. Change my mind.

Kevin: I think you've illustrated all the problems with going to an early-stage startup. So, let me talk about the other side of that balance. You've got a startup that's growing fast; there's a lot of opportunity; new problems arise every month or two, new technical problems, new organizational problems, new challenges and opportunities to learn and grow. New organizational roles open up to the lead and to be manager, and to try to jump into different roles within the organization. You're very close to your other team members; you are having lunch with the sales team, which is often one person, you're sitting one desk across the CEO hearing about all the business problems that happen.

You can pick up a lot of context, just via osmosis, about what's going on in the rest of the business because you're just talking to everyone constantly, whereas it a much larger or more stable organization rules are a little bit more ossified, you're a little more stuck in your lane. And, yeah, I feel like the excitement of a small company like that is a lot of fun. Now, as you point out, there are problems; if the CEO is terrible, you're stuck; there's no team to transfer to. You have to transfer out of the company. If the company fails, well, you might get a month of notice at best, to start looking for your next job. These are trade-offs you have to deal with.

Corey: Oh, yeah, when we realized this wasn't going to be a success, we were very clear, we made commitments to ourselves and nothing else, and we wound up giving you, what, a little over two months notice that at this point, we're going to be winding it down; let us know how we can help. And that's true. If for some reason you're listening to this, and Kevin has not yet announced he's going somewhere, I really can't recommend Kevin enough. He is one of the most insightful engineers I've ever worked with and is also incredibly human as he does it. And I think that that is a very rare thing to find, as I mean that. I work with an awful lot of engineers; I don't think of any of them the way that I do, Kevin.

Kevin: Well, thanks. I really appreciate that. And likewise, if you're looking for a job and Duckbill happens to be hiring, you should absolutely go there, spend some of the most fun six months—actually it’s been well more than six months at this point. It's one of the better one of the best tenures at a company I've been at. I’ve really enjoyed everyone I've worked with here from the sponsorship team to the consultant team to the leadership.

Corey: Except for my business partner, Mike because he's not on the podcast to defend himself.

Kevin: Absolutely. We can definitely talk crap about Mike now.

Corey: So, one last thing I want to cover before we call it an episode. Tell me a little bit more about the other shenanigans you have built for fun.

Kevin: Sure. So, probably the first one I did. I was stuck on a train—actually with Mike—in Belgium, and the train had no WiFi, so I was just on my laptop messing around. We were traveling from I want to say Brussels to Ghent, and I came up with this horrible idea of trying to make Ruby look like JavaScript. Ruby has incredibly powerful meta-programming syntax; you can do some truly terrible things with Ruby.

And so I managed to make some pretty complex code that was both valid Ruby and valid JavaScript just by massively abusing Ruby syntax. And that did actually pretty well. I've turned it into a few talks I've given at RubyConf.

Corey: I do things that are super similar but on the other end of it. Yeah, it doesn't matter if you're using this in Python, Ruby, or JavaScript, it won't work in any of those.

Kevin: [laugh].

Corey: Yours worked, and that's the scary part.

Kevin: Oh, yeah. Well, I feel like I've gotten to several of these projects now, and I'm calling them the sort of, ‘cursed code projects.’ They're terrible, horrible things you should never do, but they're also technically interesting. Ideally, someone should look at them and say, “Oh, my God, that's awful, but also, how did you do that? I want to see.”

Corey: That was a blast. I thought that was just fun watching people's brains explode when you showed them this, and like, “All right. So, here I have some Ruby”—or JavaScript [depending 00:29:36] on how you described it—and he said, “Oh. And it is also the other one.” And then just watching people sit there and nod, and like—oh—they just—it's so absurd it sort of sails past them for a second, and then you can see it hit them and they do a mental spit-take.

Kevin: [laugh]. Exactly. Probably the other one that is probably my favorite and was the most popular one I've ever built, went a little bit viral, was a CSS-only chat. It's a web-based chat. It's asynchronous, doesn't require any page reloading, you can chat between two browsers, and there's no client-side JavaScript at all.

It's entirely abusing properties of the HTTP protocol and CSS. And that's definitely one of those, “How did you do it?” Things that people have been constantly asking me and I've got a whole write-up on it if anyone's curious. Just Google CSS-only chat. I think it probably comes up at this point.

Corey: And I got to say, there's one more that was my favorite, too, given my propensity to misuse things as databases; you wound up building a Chrome extension to use Chrome tabs as a database.

Kevin: That's not even a Chrome extension. It's just a—it's actually just a webpage. Is it still up?

Corey: Oh, my mistake. I didn't even realize that. Oh great, you go to a website, and it now is your database. Just make sure that you don't wind up closing that particular window.

Kevin: Exactly.

Corey: And of course, it's hilarious and nothing you should ever do, so I'm certainly at least one bank somewhere is kind of running their transaction database on top of it.

Kevin: Oh, god, I hope not. There we go. tabdb.io. I had to actually look up the URL for that. But yes, it actually just uses the tab titles to store DB content.

Corey: And we will, of course, put a link to that as well as these other nonsense things into the [show notes 00:31:02]. Kevin, thank you for taking the time to unpack what we've been up to here. If people care more about what you're up to or want to see what nonsense you're doing next, where can they find you?

Kevin: Sure. kevinkuchta.com is probably the best URL. It's got my resume, it's got my email, it's got all the useful contact information you might want.

Corey: Excellent. Kevin, thank you so much for, I guess, tolerating the slings and arrows and on some level, I think, disappointment of this not being the thing we'd hoped it would be.

Kevin: No worries. Thanks for having me both at the Duckbill Group and on this podcast.

Corey: Kevin Kuchta, product owner and lead engineer for DuckTools. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you hated this podcast, please leave a five-star review on your podcast platform of choice along with a comment insisting that my thing that doesn't work in any of the languages I named also doesn't work in Rust.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Jonan

Jonan Scheffler is the Director of Developer Relations at New Relic. He has a long history of breaking things in public and occasionally putting them back together again. His interest in physical computing often leads him to experiment with robotics and microelectronics, though his professional experience is more closely tied to cloud services and modern application development. In order to break things more effectively he is particularly excited about observability lately, and he’s committed to helping developers around the world live happier lives by showing them how to keep their apps and their dreams alive through the night.

Links:

  • New Relic: https://newrelic.com/
  • The Relicans: https://www.therelicans.com/
  • New Relic Twitch: https://www.twitch.tv/new_relic
  • Twitter: https://twitter.com/thejonanshow

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is brought to you in part by our friends at FireHydrant where they want to help you master the mayhem. What does that mean? Well, they’re an incident management platform founded by SREs who couldn’t find the tools they wanted, so they built one. Sounds easy enough. No one’s ever tried that before. Except they’re good at it. Their platform allows teams to create consistency for the entire incident response lifecycle so that your team can focus on fighting fires faster. From alert handoff to retrospectives and everything in between, things like, you know, tracking, communicating, reporting: all the stuff no one cares about. FireHydrant will automate processes for you, so you can focus on resolution. Visit firehydrant.io to get your team started today, and tell them I sent you because I love watching people wince in pain.

Corey: If your mean time to WTF for a security alert is more than a minute, it’s time to look at Lacework. Lacework will help you get your security act together for everything from compliance service configurations to container app relationships, all without the need for PhDs in AWS to write the rules. If you’re building a secure business on AWS with compliance requirements, you don’t really have time to choose between antivirus or firewall companies to help you secure your stack. That’s why Lacework is built from the ground up for the Cloud: low effort, high visibility and detection. To learn more, visit lacework.com.

Corey: Ever notice how security tends to be one of those things that isn’t particularly welcoming to folks who don’t already have the word ‘security’ somewhere in their job title? Introducing our fix to that, Meanwhile in Security. To sign up for the newsletter or to find the podcast, visit meanwhileinsecurity.com. coming soon from The Duckbill Group.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Jonan Scheffler, the director of developer relations at New Relic. Jonan, thanks for joining me.

Jonan: Thank you for having me, Corey, this is awesome.

Corey: So let's go back before we even dive into something because I don't know how to actually file grievances or trouble tickets. So, I was a New Relic customer, I want to say circa 2012, and then again, in 2014 or so. And the product was great. I want to be very upfront: it solved a problem and we were thrilled to use it. The sales tactics were—how do we put this—unpleasant where, “Oh, you're using more than we thought you were, so you need to re-up mid-contract cycle after we had things negotiated,” and the rest. And it left a bitter taste in my mouth for years.

Now, in preparation for this, I've reached out to a bunch of current New Relic customers, and I've heard, A) the product has continued to improve, so good work; you've nailed the core competency. But relevant to where I'm going with this, it was also stated that your sales team isn't doing that anymore, which means, oh, thank God, I can do business with them again.

Jonan: Yeah.

Corey: So first, thanks. Secondly, I'm assuming that was entirely you're doing, correct?

Jonan: I did that, yeah. I actually went to them. We've got a new strategy now where we knock on your door in the morning with breakfast to have a sales meeting. It's kind of a cold call strategy: we find you, and then we talk to you over breakfa—that's not actually our sales strategy right now. I'm not in charge of those things.

I'm glad to hear it's better. We have, actually, an entire new model for paying for New Relic than we've had before where you're paying by usage rather than subscription model that fed us so far. So, a lot of those things have changed, recently.

Corey: And to be clear, something that always appeals to me, you have a perpetual free tier. New Relic and free were two words that went together basically like pudding and cheese. So that's also a bit of an eyebrow-raiser as well. Tell me about it.

Jonan: Yeah, so that was actually one of my frustrations. I worked here at New Relic—I'm a boomerang. I was here about five years ago. And I left just shortly after they went public, and went gallivanting around the industry and then came back around to run DevRel for them. But when I left, it was a very different free tier.

A lot of the features that we had at the time, which were fantastic that it started out as enterprise-focused features, you know, they expand and users want access to those. And nowadays, every part of the platform actually is available on our free tier. And what you get when you sign up for New Relic is the entire New Relic platform. Every feature you could want that we have to offer is available and then you pay as you exceed your 100 gigabyte per month free ingest allowance, which I think, personally, is a much better model for this kind of thing. Because it's just mean to get someone into the house and then show them all the fancy, awesome things that they can't play with. And now we have that.

Corey: Oh, and it solves the problem, too, historically, where you have the pricing—where you're incentivizing behaviors that don't benefit the customer, where it's oh, we're going to charge per node, or we're going to wind up going at it through a very weird lens that causes particular aspects of someone's application to spike the price. And the answer to how to control for those costs winds up being, “Oh. Just use our product less.” That seems like it's aimed in the wrong direction. And one counterpoint, your sales team did bring up with me was, “Well, sampling is not a great idea. You kind of want to capture everything.” Yeah, I absolutely do, but I have a budget, so it's a constant tension. And anything that gets away from having to actively predict those things in advance is a clear win. Let's also be clear that when we're talking about an environment that does, you know, auto-scaling, having to do a per-node licensing is sort of like hitting moving targets at any given moment.

Jonan: And that's the thing. I mean, when New Relic became a company, it was a very different landscape in the Cloud. We had a much more likely customer scenario where people were going to scale vertically. And now as is obviously the case, people are scaling horizontally: into Kubernetes, you have a massive number of nodes, or whatever you end up with in your particular infrastructure, but the per-host licensing just doesn't work in this world anymore.

Corey: Oh, back when Lambda first came out, one of the first things I did was I wound up embedding the New Relic library into a test function and then wound up invoking a couple thousand of them at once, and then just waited for the phone to ring. And sure enough, there was a salesperson who called and I don't know how they managed to make a cash register noise with their mouth, but they did. And just to taunt people because I am, first and foremost, basically an internet troll back in that era. So yeah, credit where due, you’ve figured out how Cloud works since then, and I've become a little bit less grating and obnoxious. But I have to say, the overall trend line I'm seeing from New Relic is positive. Good job.

Jonan: I wouldn't be back here if I hadn't been impressed with their progress. And I wish I could talk about more of what's going on inside because it's exciting.

Corey: I want to be also very clear here that this is a promoted episode. You are sponsoring this episode—thanks, we appreciate your business—but there is a barrier here as far as who we let sponsor things that have our name on it. Because if there's a company that wants to sponsor what we're doing, and what they're doing is abjectly awful and no one in their right mind should use it, I don't want to be associated with that. That's awful. There's always a difficult challenge of trying to square those two things of one, we like money; two, we don't like being actively harmful. And I just want to be very clear here, it was not a hard decision when we were talking to New Relic.

Jonan: Yeah. I've always gotten a lot out of New Relic. I've used it since I started in software. I mean, I came into this industry ten years ago through a code school, and that was day two of setting up an application. It's like, “All right, well, here's how you get set up with your staging development production pipeline on Heroku, and then you install New Relic.” It was just considered a best practice and I've used it ever since. I've always loved the product, and it's only getting better.

Corey: Oh, absolutely. It's one of those few things that has a different perspective on monitoring. But I want to get away from the past for the moment and talk instead about, well, I guess the present. Tell me about this whole New Relic Explorer thing.

Jonan: Yeah. So, speaking of getting better, this is the UI I’ve always wanted from New Relic. The screen that I want to see when I log in, the first screen I want to see, is an overview of my entire system. That's what I want from an observability platform like New Relic.

Corey: The mythical single pane of glass.

Jonan: That's exactly it. Right. When I used to deploy things earlier on in my career, we would open up the New Relic dashboard, and we would stare at it and wait for that red line. Nowadays, I think our practices have evolved a little bit, and even then we were a little bit behind the times in that particular role, but we were waiting for that line to go down to figure out if we were going to roll back. And that's what I want to know.

If I'm logging into New Relic, while we would like to, as a company, believe that this is all you look at all day, I imagine you have other work to do. And when you're logging in to diagnose an issue, you want to see it in your face immediately and identify the spread across potentially 2 million different systems, 2 million different entities in your platform that are interacting at a glance. That's a difficult thing to design a UI around, but it's really important. I think it comes to the fundamental flaw that humans have that we don't actually do the glyph thing where we read letters, like, staring at a whole page full of logs, and metrics as numbers and letters, our brains are not designed for that. We want pictures; we want to be able to visualize the thing and hold it in our head so that we can figure out where the pieces are falling off.

Corey: The problem I always had historically with dashboards for monitoring systems is, one, I'll go fairly long periods of time without ever logging into them. That's not necessarily a sign that this is adding no value, it's a sign of, yay, things work and I'm focusing on other stuff. And then there's the scenario that I think almost every vendor falls into, which is, “Huh. It's now three in the morning. Several graphs that are central to things have either spiked to infinity or dropped to zero, and everything else is super wonky.”

So, someone who normally is in Pacific Time is logging in 3 a.m. Pacific Time and they're greeted by a whole host of, “Have you seen what's new in our console?” “Our conference is coming up soon. Click here to buy tickets,” “We have a new user experience.” Then you have Intercom popping up. “Are you happy? Can we talk to you?” And et cetera, et cetera, et cetera. And there's no awareness built into it that the panic moment someone has when they're looking at a monitoring dashboard is, “Oh, God, get the site back up.” And it feels like there's a competition between vendors, as far as who can antagonize their customer base the most in the middle of the night.

Jonan: Yeah.

Corey: Is there anything like that baked into New Relic Explorer?

Jonan: There is no specific feature for that. I'm certain that there are a fair number of popover kind of notifications within the system, but the New Relic One platform lets you design your own dashboard. So now, when you log in there, you are—

Corey: Does it help me or does it give me the worst possible thing in the world, which is an empty screen: “Go ahead and put what you want here.” It doesn't matter if it's an IDE, a content management system when it's time for me to write a post, or a dashboard, design your own from scratch without templates is awful.

Jonan: It is awful. And that's I think one of the things that I appreciate about New Relic is that we offer opinionated solutions, we have pre-built dashboards and pre-built views, and you can also just make your own wherever they are insufficient for your needs, or you would like to see your data differently with React, which is pretty quick to learn. And you can build your own dashboards whenever you want them. But I actually—now that Explorer is here, I don't imagine I'm going to be using that feature as much.

Corey: Consider this something of a feature request. And again, you're in an uncomfortable position, specifically because every time someone is logging into their monitoring system, I can already answer the survey question you send out: “Are you happy as a New Relic customer?” “No, I’m absolutely not happy because, in an ideal world, I wouldn't remember that I was a New Relic customer because my stuff would be working.” It's actively broken right now, and I am annoyed and pissed off at everything that appears in front of me, kick-the-dog-on-the-way-over-to-the-computer-style. So it's a hard problem because on some level, you have people whose primary interaction with what you're doing is during incidents, yet, you are the director of developer relations, which means that not only do you have a whole bunch of pissed off developers because the stuff that they wrote broke, but you have to talk to these people. Tell me about that.

Jonan: Yeah. I mean, it's my job to be there and just to listen to people. I've got kind of a weird background. Before I came into tech, I worked in hospitality for a long time. One of the things they teach you when you're working the front desk at a hotel is that it is always your fault and it's always your problem, and you are prepared to give away the house if necessary, to get this person to someday stay at your hotel's brand.

Again, I worked at the Hilton back in the day. If someone came down to the front desk, and they wanted me to comp their two weeks’ stay in its entirety, and we’re flushing $5,000, we were trained to think about the $10 million lifetime spend of that Hilton customer. And I consider DevRel to be a step in that direction for software. This DevRel team that exists here at New Relic is here because New Relic cares about those interactions. So, it is in fact my job to get out there on Twitter and take the heat.

I would like it—community, if you're listening, if you just sent me some nice tweets, sometimes. I know you all have good experiences with New Relic, too. Just send me the good news. But you don't hear that, and I get that. That's part of what being in a customer service style role is.

I'm okay with that. I want to hear that, and I want to curate it and take it back to the product team so we can improve and iterate. But maybe a smiley face, too, on the tweet when you send it.

Corey: You talk about tweeting and you talk about interacting with folks in different areas. And that's great, and I'm thrilled to have all kinds of different means of interaction, but you're also getting big into Twitch. And what I'm unclear on is that just because really into Fortnight or something, lately, or is there more to it than that?

Jonan: Yeah, I would say there's a little more to it than that. I have not been playing very much Fortnight. The shooter-style games are not my thing. I like Minecraft, World of Warcraft, neither of which I try to play very much anymore because I have a tendency to get sucked into games. But, yeah, Twitch is primarily known for being a gaming platform, but it more and more is being seen as an opportunity to live code with community members and to level up when you're learning some new technology with someone in real-time in a much more engaging format.

It's a choose your own adventure television show where you're in the chat guiding the conversation and participating. I think our generation got a lot of grief over the years for having short attention spans. Y'all got ADHD because Nintendo and whatever. And that may or may not be true, but what I think has actually happened is that we've shifted away from traditional mediums like newspapers, and broadcast news shows because they're one-way communication. Online platforms like Twitch, allow you an opportunity to be a part of the media that you're consuming.

So, I can pair program, I can mob program with a hundred people from the New Relic community who want to come and learn about some new feature, they want to come and write a Nerdlet—one of our programmability dashboards—with me. They can decide the direction that goes, with me, live. I don't think that there has been another such opportunity even in a live conference format, which was the traditional realm of DevRel before the pandemic hit us. We were all out there going to these conferences. But you're standing on a stage—

Corey: Oh, yeah. Most often in the before times, your two most common places to spot DevRel were on stage at a conference, or in the frequent flyer lounge at the airport.

Jonan: Exactly. In both cases, it was a one-way conversation. I mean, DevRel in the lounge because we're really loud, but DevRel on the stage because it’s a stage, and we have a slideshow. And you get the Q&A and the hallway track, but all that's gone now. And now we have the alternative of going to a virtual conference—which is the least engaging parts of that experience—or going to Twitch where we get to have a conversation. And that's why our new team, The Relicans, you can go to the therelicans.com, audience, and check us out.

Corey: Please tell me there's a bird logo.

Jonan: Yeah. It's a pelican and their name is Ellie.

Corey: Excellent, good, good, good. I’m just checking. Sometimes it seems that companies don't connect the dots that seem obvious to me.

Jonan: Yeah, this was actually a little bit of a joke early on in New Relic. We were specifically instructed to call ourselves ‘Relics’ because relican rhymed with pelican. So, when I came back, of course, I had to name the team, ‘The Radicans.’ We chose it together, but I may have influenced the direction there. Yeah, we have a pelican. Their name is Ellie, and Ellie is featured all over therelicans.com.

Anyway, we're all on about Twitch; you can find us on twitch.tv/new_relic. There's ten of us right now streaming more than 100 hours a week on various topics. And maybe only 10 percent of that time is explicitly about New Relic. We're teaching people how to program in a variety of languages. It's just fun. Is a fantastic platform. I love playing on there.

Corey: It really is one of those just foundational type of experiences. I don't know how to describe Twitch effectively, but you're right, absolutely every other online conference—with a bit of rounding error—as aspects of, “Oh. It's just like another Zoom meeting only in some ways worse.”

Jonan: Exactly.

Corey: In my case, it’s because I'm not allowed to take myself off of mute and insult people like I do in my Zoom meetings. But for other folks, it's almost lecture-like. And this is my big problem fundamentally, with the way that DevRel has handled these unprecedented times. And it's that they're trying to take the same thing that they would have done on stage and now they're just doing it in front of a camera. It doesn't work.

Jonan: It doesn't. And it's not going to change. I think once the live events come back, we're going to be back out there because the real value for me was standing in the hallways and having the conversations with people, getting that feedback real-time. And now I have another way to do that is only complimentary. But I don't mean to say that virtual events are not a thing we're participating in; we're still out there and doing those, but they are a select few.

And they tend to be the ones that have found other ways to drive that engagement. Because we're there for the community as much as we are there for our companies. To the community, I represent the company; to the company, I represent the community. I need to get that feedback from people so I can bring it inside the house and help improve New Relic as a platform. That is my job. And if I'm just talking, I'm not getting it.

Corey: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the Enterprise (not the starship). On-prem security doesn’t translate well to cloud or multi-cloud environments, and that’s not even counting IoT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IoT devices, detects these threats up to 35 percent faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at extrahop.com/trial.

Corey: There has to be an engagement model. And I think that we've talked about on the show repeatedly, through various episodes, a recurring theme has been that DevRel is widely misunderstood, including in many cases, by practitioners of same.

Jonan: Yeah.

Corey: It's a hard problem to solve for. People think, “Oh, it's marketing.” But it's not, really. “It’s product.” But it's not, really. It almost lends itself to a definition of being categorized based upon what it isn't. And that's always a hard thing to do. We saw massive layoffs around the DevRel space when the pandemic hit. You've grown the team during that same timeframe.

Jonan: Yeah.

Corey: What is it that you see differently?

Jonan: New Relic right now is focused on developers. And I don't think that I would surprise anyone at New Relic or outside to say that maybe we weren't doing a great job of that for a while. We were focused on growing the company and talking to larger companies. And in that process, the focus naturally trends away. I mean, look, this is a standard thing that happens.

A company with a developer product goes public; they get there by swiping their first Amex from someone, and they're like, “Oh, hey. That just supplanted all the income we ever earned from the $7 a month people that we're spending all of our energy on.” And what they don't realize is the Amex swiped because that CTO sent a message to their whole team—first thing—when they got a contact from a salesperson at New Relic and said, “Hey, what do you think of New Relic?” And eight out of ten people went nuts. They were so excited that they would get to use New Relic instead of whatever home-rolled scripts or whatever they had at the time, that it was an obvious choice.

And the CTO gets to look like a hero in choosing this platform. And that starts to evaporate out from under those leaders as you move your focus to larger enterprise business. And I'm not saying you don't have to focus on both sides, but New Relic now more than ever understands how much developers mean to their success. And I think that is the real heart of New Relic, we have always been a developer-first company. But I think in the last six months since I've been here, I've seen evidence of that in ways that I would never have expected from outside. And this DevRel team is an example of that. We're here for the people. I'm really excited to be doing it.

Corey: A periodic example I keep coming back to is let's say that I start some nonsense Twitter for Pets style company tomorrow. And in the next ten years, it grows to be a component of the Fortune 500. At some point, between it's just me with a dumb idea and household name, there has to be an on-ramp for a vendor to become my vendor. And an awful lot of—if you'll forgive the term—legacy companies, but not in the flattering definition of legacy I usually use of ‘It makes money,’ but companies that were poised at the old model of selling to executives and then having it inflicted on their staff, are instead dealing with, “Oh. You mean, we need to get buy-in from folks?” Sean O'Grady’s book, The New Kingmakers—about how developers are, in many cases, the influencers that drive a lot of these purchases—seems to be a lesson that companies lose sight of as they hit certain milestones such as exceeding certain sizes, going public, having certain percentages of the S&P, or having certain percentages of the Fortune 500 as customers.

Jonan: Yeah. And I mean, you have to focus on different things. There are a certain number of features that an enterprise client demands of your business that you need to really focus on to implement. Like, PCI compliance is not a small thing, HIPAA compliance is not a small thing for an entity to build into their software. So—

Corey: It's not negotiable.

Jonan: Yeah, it's non-negotiable if you want to play in that space. So, if you do you want to serve the full breadth of potential software customers, then yeah, you need that. And it takes a little bit of focus time. I think that a lot of companies spend the appropriate amount of time doing that. What's important to me is that they remember why they're here in the first place. And New Relic very much remembers that. And I appreciate that about this company.

Corey: It's reassuring to hear that a lot of these lessons don't necessarily get lost, they just get obscured from time to time. And companies that are targeting big enterprises, of course, have to hit those compliance goals, but that doesn't mean that that's useless to startups. A further point beyond that is that every big company that I talk to right now tends to perceive itself as, “Weird like a startup,” or, “We're going to be like a startup,” or, “We're going to use some tools a startup might use.” And those tools still have to check those boxes. But you also never hear the opposite where you have a startup, and, “Oh, yes. We have a procurement process that's just like a big enterprise is.” Because the answer to that is, “Holy God, why would you do that?” There's a narrative that people want to fall in line with of how they are perceived. They want to perceive themselves a certain way that tells a certain story. Whether it's accurate or not, is almost beside the point.

Jonan: Yeah, no, I absolutely agree with you. And I think that specifically in the observability space, that move fast and break things mentality that sometimes exists in the startup community is not an ideal place to be. I think New Relic does a very good job of taking their customers seriously and their customers’ data seriously, and their business seriously, and maybe some of the time, not ourselves so seriously. I mean, I'm on a team that has a pelican for a logo. I certainly am in this space in a spirit of play. But there's a line there that is difficult for companies to walk. I agree with you. It's often the biggest mistake that a company at scale will make, is getting that wrong.

Corey: Do you think that it's required to get a lot of stuff wrong before you get it right or is it possible to learn from the mistakes of others and just do the right thing the first time?

Jonan: I've never been a big believer in vicarious learning, right? Ask every kid whose parents told them they didn't ever want to get drunk, “Never drink more than three,” has never worked. I think that we keep learning the same lessons—as evidenced by a lot of our languages and frameworks and tooling—and we eventually iterate towards better as a natural circumstance of that. I mean, there's not as much efficiency as there could be in going from here to five percent better, but I mean, you and I both worked in this industry ten years ago; it was not awesome. And today, I think it is better. But along the way, there was a lot of missteps. So yes, I think that you can certainly learn lessons from those around you, but it's easy also to feel like you're cutting ground where you're not, if that makes sense. And I think a lot of companies fall into that.

Corey: It's also easy to sit here, from my perspective—remember, I don't run any production software here. I'm mostly casting stones. The reason I'm not currently a New Relic customer is primarily because I use hosted WordPress and, frankly, if that site goes down, it's not the end of the world—first off. Secondly, I already have a monitoring system in place that costs less: it's called Twitter because people will absolutely let me know whenever my site has the slightest of hiccups. And application performance on something like that, or even moving to the other aspects of what New Relic does, isn't—to be blunt—that interesting to me.

It's a content website. There's no web app built onto it, so at least through my current lens, I don't view myself as a prospective New Relic customer. So, on some level, I'm sort of LARPing here as, well, what I would do if I had something to monitor today and how I would approach it. But I'm confident in what I'm saying for the most part just because I talk to people who are in that position, who do care about these things, who actually care about their customers. And New Relic is not a bad option these days, I've got to say, by every person I'm able to talk to. Unless I'm missing something. Are you secretly terrible and I just don't know?

Jonan: No. We are not secretly terrible. I think that I'm going to need to check my press brief here. No, it says right here, I can say that officially, “We are not secretly terrible.” It's a fantastic product. It always has been. The New Relic Explorer interface is a huge level-up for us, see all of it at a glance.

I would suggest that even if you don't care about the people reading your blog, you might be interested to know how long or how broken their page-loading experience is. They may not tweet at you just because it takes them 60 seconds to load the page because the blog post eventually shows up. That's the kind of thing we can help with. There's a free tier now, Corey, get on there. It's 100 gigs. You're not going to use that on your WordPress blog.

Corey: Oh, well, you'd be surprised. My logging is super verbose, and I tend to have an awful lot of unoptimized nonsense. But yeah, it's one of those things to where, to be direct, I don't see that there's any meaningful business benefit from improving a lot of the performance on this. Yeah, I know that it was Walmart or Amazon or someone did a study many years ago that's quoted the death about, every 100 milliseconds of latency reduces revenue by some percentage. I feel like that might make sense in the context of whatever they did. But as a global truism, I can't shake the feeling that it's significantly overplayed.

Jonan: Yeah, I think so. I think that there are a lot of opportunities to use a platform like New Relic to get observability for other reasons as well. I mean, ultimately, it is, I think, a focus on the customer and it always should be, but you're not necessarily looking for places where you're losing money; you might be looking for places where your customers are having a bad experience and you're not ever going to hear about it, they're not ever going to click your thumbs up/thumbs down form or your intercom pop up, you're never going to know. They're just going to walk away, and they're going to talk about you to their friends in the hallway, at the conference. That's, I think, where I watch that knowledge.

But again, I'm a very community-focused person, I always have been in my career, but that's what I care about is the end users’ experience using the product because I want to make the world better. I think most people do; call me naive, but I think most of us are in it to improve things a little bit while we're on this planet, and I would like to see developers' lives be better. Because as you and I both know, they're pretty bad a lot of time. Computers are terrible. It was a mistake.

Corey: Computers are terrible, and then we have to add people to that mess, too.

Jonan: Yeah. Of the two, I choose humans.

Corey: Jonan, thank you so much for taking the time to speak with me today. If people care more about what you have to say and how you care to say it, where can they find you?

Jonan: They can find me on the internet on Twitter. I'm @thejonanshow on Twitter. Better place may even be therelicans.com, our new developer relations blog that we launched. And of course, we're all on Twitch all the time; you can find us through the twitch.tv/new_relic channel, which is linked everywhere, and all ten of our individual channels are on there. I do a lot of playing with Kubernetes and Raspberry Pis and other nonsense on there. So, come on by and say hello; I would love to hear from you.

Corey: Excellent. We will of course throw links to that in the [show notes 00:30:08]. Thank you so much for speaking to me. I really do appreciate it.

Jonan: Thanks for having me, Corey.

Corey: Jonan Scheffler, director of developer relations at New Relic. I'm Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you hated this podcast, please leave a five-star review on your podcast platform of choice and an insulting comment talking about how this podcast free trial isn't really free at all.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Tom

Tom is the co-founder and CEO of Command E, an app that provides blazing fast search across all your docs and records in G Suite, Salesforce, LinkedIn, Dropbox, and 20+ more tools via one easy keyboard shortcut.

Links:

  • Command E: https://getcommande.com/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Fairwinds. Whether you’re new to Kubernetes or have some experience under your belt, and then definitely don’t want to deal with Kubernetes, there are some things you should simply never, ever do in Kubernetes. I would say, “run it at all.” They would argue with me, and that’s okay because we’re going to argue about that. Kendall Miller, president of Fairwinds, was one of the first hires at the company and has spent the last six years the dream of disrupting infrastructure a reality while keeping his finger on the pulse of changing demands in the market, and valuable partnership opportunities. He joins senior site reliability engineer Stevie Caldwell, who supports a growing platform of microservices running on Kubernetes in AWS. I’m joining them as we all discuss what Dev and Ops teams should not do in Kubernetes if they want to get the most out of the leading container orchestrator by volume and complexity. We’re going to speak anecdotally of some Kubernetes failures and how to avoid them, and they’re going to verbally punch me in the face. Sign up now at fairwinds.com/never. That’s fairwinds.com/never.

Corey: If your mean time to WTF for a security alert is more than a minute, it's time to look at Lacework. Lacework will help you get your security act together for everything from compliance service configurations to container app relationships, all without the need for PhDs in AWS to write the rules. If you're building a secure business on AWS with compliance requirements, you don't really have time to choose between antivirus or firewall companies to help you secure your stack. That's why Lacework is built from the ground up for the Cloud: low effort, high visibility and detection. To learn more, visit lacework.com.

Corey: Ever notice how security tends to be one of those things that isn’t particularly welcoming to folks who don’t already have the word ‘security’ somewhere in their job title? Introducing our fix to that, Meanwhile in Security. To sign up for the newsletter or to find the podcast, visit meanwhileinsecurity.com. coming soon from The Duckbill Group.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Tom Uebel, who's the CEO and co-founder of a company called Command E. Tom, thanks for joining me.

Tom: Thanks for having me, Corey. It's great to be with you.

Corey: It really is because I'm delightful. Just ask anyone who has been on the show and I basically hold at gunpoint, to admit that I'm delightful. So, for a long time, I have been whining for something that doesn't really exist. And when I discovered a post that you had made about a month before this recording or so, I realized, “Oh, my God, someone built it. I got to get him on the show and yell at him about it.” And what you have built is also the name of the company, Command E. What is Command E because there's a near certainty you'll do a better job of telling that story than I will.

Tom: [laugh]. Yeah, so Command E just gives you one keyboard shortcut to get to any of your docs and records in the Cloud. You can think about any tool that you use at work, Command E is blazing fast search across all of that. So, you can be going about your day, and no matter what you need to get to next, you just hit Command E, quickly type in the search box that pops up, maybe it's like, “Corey Quinn, LinkedIn,” and hit enter and we'll get you there as fast as humanly possible.

Corey: I feel like on some level that you're being completely accurate, while also dramatically underselling how magic this thing is. I want to be clear, you are not sponsoring any of my nonsense, yet. This is me more or less fanboying over what you have built. Fundamentally, from where I sit, it’s… I hit Command E, which is the shortcut, and I type in anything I want. And it does the stuff you'd expect, like files on your computer, but then it goes beyond that.

There's a crap ton of services—I don't know if it's metric crap tons or Imperial crap tons, but whatever it is, there's a lot of them—where you integrate with various third-party APIs. I mean, as I'm pulling it up now, I notice Clubhouse on the list, of all things. You integrate with Dropbox, with Gmail, with GitHub—or Jif-up, depending upon pronunciation—JIRA, Slack, Spotify, and the list goes on. I'm not going to read this off; I'm not part of your marketing department.

But it's impressive in that it ties in and solves the problem of you and I were talking about something two weeks ago; where was that conversation? Was it an email? Was it in Twitter DMs? Was it in LinkedIn conversations? Was it in Slack somewhere? Was it in post-it notes or something? And it doesn't matter. I can search and as long as there's an integration, it just shows up. And that's part of the magic. Have I nailed the salient points, or am I overstating the case so far?

Tom: Yeah, that's it. It's really there's been this proliferation of tools as the cost of building great software has come down. It's any category you think of, there's a few great tools in it. And that's awesome. It's great that you can find best in breed tools for anything you want to get done today, but it involves all these switching costs that maybe didn't exist 10, 20 years ago.

And Command E is that layer of glue that pulls all these things together really nicely, and works the way your brain does where the second you think, “I want to go to this,” we just get you there; you don't have to think about, “was it in Gmail? What was the name of that Google Doc that Corey shared with me last week?” Just the second you know what Google Doc you want to go to, Salesforce record, JIRA ticket, whatever it is, you just hit Command E, and we'll get you there incredibly quickly.

Corey: And you're not exaggerating when you say incredibly quickly because I figured, “All right, great. Typical story here. There's a search, everything has a search and it usually sucks.” This doesn't suck. What is that using under the hood?

Tom: Yeah, so there's a few pieces that roll up to something that I think is really compelling. So, one, it's a desktop app, so you can be anywhere on your computer. It's not just browser-based. Anywhere on your computer, hit Command E—Control E, if you're on Windows—and then the other thing is that we are tapping into all those services and we are doing a pretty good job, I think, of knowing what you're probably going to want to get to so that we kind of have it built up where every search should return in 10 milliseconds or less. A lot of people talk about 100 milliseconds as being the speed at which humans can perceive, “Hey, I'm waiting on something or not.” So, we try to keep all our searches to 10 milliseconds or less so it really does feel like your tools are finally moving as fast as your brain can, and it's an incredibly powerful feeling that I think a lot of us just haven't experienced in a while.

Corey: No, it's very clearly not reaching out to anything across the network, which gets to my next part that to be very clear in disclosure here, you and I did have a conversation about this when I first discovered it because honestly, I was convinced you were doing something horrifying and that I do have some level of authenticity in what I tell people on these shows. And I don't want to be advocating for something that is not working in people's interest. And there's a couple things that scared the hell out of me. The first is, you link with a whole bunch of different APIs, Like, I have you linked right now to multiple Gmail accounts; I have you talking to my Evernote stuff that has a bunch of history from back when I used to use Evernote because they used to care about their customers; I have it tied into Spotify, not that I care about that; a whole bunch of local stuff on my own disk, and so on and so forth. There's some sensitive data in there, is the first part.

So, what is the whole privacy policy? And is my next big announcement going to be a data breach? And followed very quickly afterwards by, what does it cost? Oh, there is no direct way to give you money. Combine the two of these things, and you can sort of see why I jumped to a terrifying conclusion there.

Tom: Definitely, yeah. So, this is something that we thought long and hard about from day one and made sure we did this right. Candidly, I sleep a lot better at night with the security model we went with. It's been kind of interesting seeing this debate play out with Signal versus WhatsApp and various other tools more recently. But from day one, our security model has essentially been, you control your own data, we don't actually have it on our servers or anywhere else.

Your data is synced from these third-party services and stored encrypted on your own device as well as the connections that you make to services to enable us to do our thing, that's stored securely on your own computer. None of that is going to us, so even if we wanted to we don't have access to your service, our terms stipulate that. So there's no surprises here.

Corey: So, my data never leaves my machine?

Tom: Exactly.

Corey: Awesome. I know you’ve said that before, but it's important to get that on the record on these things because it's one of those like, “Well, remember that time we destroyed the company by trusting something?”

Tom: Right. [laugh].

Corey: Which brings me to the next. Like, there's a trope going around that if you're not paying for something, you are the product. On some level that even applies to this podcast. I mean, there are sponsor ads that are put into this, and on some level, the audience listening to it is the product. Now, we take a very restricted view of this; we wind up saying, well, we get this many downloads, and we think GeoIP-wise, this is the generalized distribution globally, and that's all we say because that's all we know.

It's more or less screaming into the void. At the end of that though, it does wind up working because people listen to the ads; if they're compelling and relevant, people will visit them, and it becomes a very valuable thing for the sponsors. So it's not necessarily an inherently awful thing, but it's also the audience here is paying with their attention, for lack of a better term. And with this, I don't see you're doing sponsored ads next to my content of, “Wow, you've just pulled up your W-2, it looks like you should have a new job. Have you considered working at Google?” There could be some horrifying monetization plays here. What is the plan for that?

Tom: Definitely. Yeah. So we're in a fortunate position, we’re early days backed by top venture firms in the Valley. And so right now, the product is free for any individual to try. We'll eventually have a pro tier and team plans in time, so this very much will be traditional SaaS monthly subscription payment at the team level. It's definitely not a data play where we're, even if we had your data, we would be trying to do anything there. This will be a subscription SaaS with team plans, in time.

Corey: One of the things that I noticed about Command E is you could probably view it as something of a weakness—I view it personally as a strength because it reinforces what you've already said—which is, when I installed it on a second computer, it had no earthly idea who I was. It had none of my synced accounts and I had to go through and log back into all of them. Now, on the one hand, is that annoying? Mmm, slightly, but it does validate that if you're stealing my data, you're really bad at it and not using it in any way to benefit the user experience. And credit, where due I not completely as, shall we say, naive as I might appear. I did some packet captures and kept a close eye on this thing and saw what it was talking to, and if you are stealing my data, you have also built one of the best compression algorithms in the universe.

Tom: [laugh]. Yeah, now it's always fun talking to people that say, “Hey, that sounds great, but let me do my own homework and see if that's actually the case.”

Corey: Yeah. Again, I have sensitive data with my clients on the consulting side. I don't want to ever turn this into a story of, “Well, I took it on blind faith, and that seemed like the way to go.” Believe me, if there had been something nefarious, this would not be the conversation we were having right now.

Tom: Definitely, definitely.

Corey: But it's phenomenal, and the problems that I have with it now have extended beyond the, “Oh, the getting up and onboarding, and getting the muscle memory trained,” which in some ways is kind of the hardest part. And now it's going more in the direction of integrating with additional services, limitations of the integrations made available by the various providers themselves. For example, I don't see a Twitter option, particularly to look at DMs because those things are impossible to search natively. I can see that there's a bunch of, in some ways, they look like duplicate services like Gmail and Superhuman, for example; it goes to the same data store. And looking through this, it's interesting, and I'm very curious to know how a lot of these decisions got made. Tell me more.

Tom: Yeah. And I mean, honestly, it's one of the most exciting parts of running a startup, I think, is when you get your product out there and what you keep hearing from people that try it is, “Oh, this is great, just can you please carve off more of my world and support more services?” So that's definitely something we'll continue to work on. We want this to be sort of the command center that you can run your day off. So, continuing to work on supporting more services.

Corey: So I've been whining about this for years, in the fact that this didn't exist. And it turns out that when you have a problem, you can either whine about it, or alternately, you can apparently raise some VC money and go fix it. And you took the path less traveled. Why? I mean, complaining is fun.

Tom: [laugh]. Yeah, so where Command E really came from is my co-founder and I worked together for a few years before this at an early stage VC. And while we were there, we saw a few things. One was, we were on a team of 12, and we counted them up one day, and we were using 26 different cloud systems. So, I think it's become this problem that people are really aware of at this point, but there's just been this proliferation of SaaS tools and great software that I alluded to earlier, and it creates some pain in finding your information.

And then the other piece was that some of our closest friends there, some of the smartest people we've worked with were not in engineering roles; they were living their days between Salesforce and Gmail—multiple Gmail inboxes—LinkedIn, AngelList, Crunchbase, and just, sort of, tearing their hair out at the friction of moving between these systems, quickly getting information in and out. And meanwhile, my co-founder and I were in our code editors. And you're familiar with, you know, anytime you need to move to the next file in your day, it's just this really nice pattern of you hit a keyboard shortcut, a search box pops up, focused, ready for you to type, you type a fragment of the next file name, you hit enter, and you're there, and you move on with your day in less than a second without ever really—

Corey: Well, someone hasn’t configured VIM the same way that I have.

Tom: [laugh]. But it's just really quick; just we didn't experience this problem that they were experiencing because we had great tooling and we realized that some of the underlying tech had shifted in a way that it was possible to build something in the mold of Command E that would just make their lives so much easier and let them be as productive as they're capable of being, and it was just their tools getting in the way.

Corey: One thing that I find a little challenging about the onboarding is not the tool itself and it's not setting up the accounts—I whine about it, but it took me three minutes; not the end of the world. The problem I have is now that it exists, I have to remember that it exists. It's overcoming the learning curve of train yourself to automatically whack Command E. And that took me a little bit of doing. Is that a common thing? Am I just basically a slow learner? What's the deal there?

Tom: Yeah, it's interesting. I think we have more work to do there, but you'll have some people that will just immediately map it to one use case and then expand it to do more things over time. There's some people that it just kind of works for them. Like engineers are really interesting to us, just because it's a pattern that they've experienced before and they don't have to remember this new behavior, it's just mapping it to a few services. But yeah, it's definitely a different way of working, that once you're bought in on and you're ramped, it's impossible to go back. But it does take some doing. You are undoing a decade or so point-and-click muscle memory.

Corey: It's also easy to look at this because I found it a month or so ago, and, “Okay, so you're a month or so old.” It turns out that as I learn things, I'm figuring out—among others—that software does not spring fully formed from the forehead of some God. It takes an iterative design process, and in fact, I believe you folks were founded, when, in 2018?

Tom: Yep.

Corey: So, let me ask you the awkward, difficult question, then. Looking at it, it's a very simple app, as far as the way that is designed. And honestly, that is a testament to its design. It's not at all difficult, from my perspective, to work with it in an intelligent way, to understand what it's doing, but it doesn't feel like there's a lot of there, there, if that makes sense. There's not a whole lot of configuration options, there aren't a whole lot of messing around with other nonsense. Instead, it looks like it's very stripped down. And it feels like if you talked about this on Hacker News, the response is, “Oh, it must have taken you a weekend to build it, right?”

Tom: [laugh]. Yeah.

Corey: Explain to me how this evolved because I'm pretty sure that I'm looking at the end road of an awful lot of customer stories, and iteration, and, honestly, attention to detail.

Tom: Yeah, and we knew early on that this was going to be the kind of thing that just had to work really simply, so what I like to tell people now is you can go to our website, hit the download button, and have the download done, all your accounts connected data synced and have completed your first search in less than five minutes. So I'm pleased to hear that it only took you three. It takes a lot of work. There's definitely some real engineering under the hood. One of the things that my co-founder Ben really kind of nailed early on was he thought a lot about how Slack took IRC, which I think a lot of us that have been on engineering teams had communicated with other people on the team through IRC, but you never really saw sales, marketing, all these other roles in the company in IRC, but then Slack built this amazing desktop app experience around it, and made it really easy to get up and running, delightful to use, and simple, and all of a sudden, you had wall-to-wall deployments; everybody was in it, everybody liked it.

Corey: I still maintain, and this is not necessarily the universal story. But I used to set up IRC servers and ejabberd servers to talk to folks, and Slack started eroding all of that, and people got very angry. Like, “What is the deal here?” And the honest answer was, “Look, I can go and talk to someone in marketing or in accounting whose primary language is not writing code, and give them just this webpage, and suddenly they're up, and they're in the chat, and persistence works.” It fixed the user experience, and people love to overlook the sheer value of that.

Tom: A hundred percent. Yeah, that's really what we're trying to do. I think I really like products that are just very simple and elegant on the surface but have a lot of depth to them. So, I think that's really the needle that we're trying to thread here where anybody can get up and running and it just works the way you'd expect it to, but there's also a lot of power under the hood for people that are power users and want a bunch of advanced features.

Corey: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the Enterprise (not the starship). On-prem security doesn’t translate well to cloud or multi-cloud environments, and that’s not even counting IoT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IoT devices, detects these threats up to 35 percent faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at extrahop.com/trial.

Corey: So, tell me a little bit more about the evolution of this because it's easy to go from the initial problem statement of, “Huh, I have an awful lot of different things I need to search independently and in obnoxious ways,” to, “Okay, now let's start unifying that. And what that looks like.” Is that an iterative process? Is it originally imagined as a web app instead of a desktop app? Was it originally viewed as targeting a single operating system? What was the product evolution here before it sprang out fully-formed in the light of day?

Tom: Definitely. So, my co-founder, Ben, had been talking to me about this while we were working together previously, and then we both left the firm that we were at. And I went and traveled for a bit, and I came back and he had put together a V1. And we went to dinner and he pulled it up for me. And at that point, it was desktop from day one with the keyboard shortcut that pulls it up.

So, this isn't a story of a lot of pivots along the way. It's a story of, I think, just ruthlessly continuing to evolve the initial idea, which I think was really spot-on from the beginning. The first version I saw was a desktop app that pulled up with Command E but searched across all of your Chrome history. And we've since then realized that, okay, we don't want Instagram and Facebook in there, we want this to be a work tool that's very focused on helping you be productive in your workday, and continuing to add integrations, continuing to add features. So, we were very focused on docs and records early on.

And as we move to this command center that you can run your day off, it means adding stuff like calendar so that if I hit Command E five minutes before this call, I shouldn't have to do anything, the first result should just be the event we're about to be on, and I can just hit enter and the Zoom link, I’m just taken straight into the Zoom instead of having to dig into my calendar that went from being purely browser and cloud-focused to now letting you search across your local machine because a lot of people told us, “Hey, I love the fact that you've added cloud search; that wasn't really a thing before, but I just want one pattern that I can build and use across all of my searches.” And so there's more to do, but it's just that continual evolution of trying to carve off everything so that you really do just have one place that, as quickly as possible, gets you to the next thing in your day.

Corey: One thing that you talk about is the [vision 00:21:35], this is cloud search, and I always dropping this Easter egg in for folks who listen to this and pay attention. It's always been irritating to me that Google—the search engine company—will have Google Docs. And you can use the search for Google Docs, and it's rubbish because why would a search engine company build a good search engine for its own work suite?

But if you go to cloudsearch.google.com on a paid account, it's awesome. It does more or less what Command E does, albeit only within the confines of your Google account. And on some level, that becomes almost a great definitive answer to, “Well, what's to stop some other company from coming at you by behind?”

Oh, don't worry, if a Google or an Amazon were to come out and do this, they would really screw it up. See slide B. Because they've done things like this and completely failed to tell that story compellingly. Every time I bring up cloudsearch.google.com, people are astounded and they think I'm messing with them, then they pull it up in the browser so they can yell at me, and then they get angry because why did they not know this existed years ago?

Tom: Yeah, it's really fascinating seeing [laugh] the different cuts of that. I’ll bite my tongue on that one. But, yeah.

Corey: Of course. You have to be nice to people. I don't, which is sort of my entire shtick, and it works for me. So, what can you tell me, if anything, about future roadmap? I mean, it's easy to look at this from a perspective of integrating with different services, I think—I saw you recently started supporting Coda, which I'm using in some weird edge cases, and oh, that becomes super handy. But what's on the roadmap coming up? Is there an idea at some point that I'll be able to invoke local bash scripts, for example, on my system?

Tom: Yeah. So we're trying to let people do more and more. So, I mentioned we just recently added calendar; if you think about the ways that you start units of work in your day, it's often hey, let me pull up that Google Doc; let me get into this meeting—

Corey: Oh, that's right. This is what people with real jobs do. I wind up starting my day by pulling up Twitter mentions and starting with, “Oh, Christ, what now?”

Tom: Yep. [laugh]. They’re just, sort of, continuing to nail away search use cases. So, one thing that we want to get better at is hey, what was that doc that Corey shared with me? And we made it really easy to get into your Zoom, but can we make sure that we surface the relevant information there?

One of the things we're really excited about right now is supporting more collaborative use cases and letting teams work with Command E, I think there's definitely some benefits to solving the problems that we've solved at the single-player level for teams as well. So, just continuing to expand on that.

Corey: What about potential dangers? One of the things that I see when looking at this is the ghost of feature creep, on some level where, hey, we can use this to also—I don't know—do keyboard macro expansions. Or we can use this to automatically start killing processes that are running amok with the right invocation. And at some point, it almost tries to become an electron-based version of the terminal where at some point, you're more or less building an entire operating system into this application. Is that something that you're concerned about? Is that so far down the road that I'm hilarious for even mentioning or thinking of it now? What are the dangers that you see if this doesn't go right?

Tom: Yeah. I think there's a few things you could definitely see tools that have tried to build for all these bespoke use cases and lost the focus. I think simplicity is super important to us. If you think about our guiding principles, it's often focused on performance, so from day one, we've said it's going to be an extremely strong engineering-lead team that solves this problem. And so we focused on just building the best engineering team and then bringing on a great designer to make sure that we're always focused on keeping things really performant and simple.

I do think that we've also very much focused on work use cases. So, we get a lot of requests for different things. And one of the things I like about building B2B software is often your customers can kind of pull out use cases from you. And it's just a matter of applying some discernment and just really building for them. So, we try to stay really focused on things that we think will be used very broadly by the personas that Command E works really well for which is across sales teams, across product teams, engineers, it's just really focused on staying close to the voice of the customer.

Corey: That's part of the challenge that scares me on some level is, I always feel like my use cases are bizarre and more than a little bit frightening, but I knew that Command E was something to, honestly, be super excited about which I hope my enthusiasm is conveying accurately. I'm not shilling; no one's paying me to talk about this. Unfortunately because as it turns out, I'm not as good at business as I thought I was. But there's something deeply compelling about when I saw this, and it just grabbed me. But I wanted to validate that because again, I'm super weird.

So, I showed it to other folks, and every person I've shown this to has lit up as soon as they see what I'm talking about. And I'm not just talking engineering users, I'm talking business folks. And there's really something that you have tapped into here. It seems like it's the sort of thing that has a viral potential tied to it. In the positive way; not the infosec sense.

Tom: Yeah, it's been awesome. Like, I love conversations with people like you that are just like, “Oh, yes, I've been waiting for this thing. Thank you so much for building it. It's great, but also, can you do these three things?” It's just—

Corey: “This is amazing. Now, let me tell you what your problem is.” Yeah.

Tom: [laugh]. But yeah, I just think my co-founder, Ben, and the team have really nailed it. This was this natural evolution, I think when you think about how things have matured on the web and at work, over time. So, I remember finding this clip where Marc Andreessen was talking about the Netscape browser and how previously everything was done through what he called cryptic commands, talking about the command line, which just wasn't really accessible to most people. And then all of a sudden, you had a browser, and people could point-and-click their way around, which makes total sense when the corpus you're talking about is sort of the entirety of the internet.

But so many of us go to work and there are so many of these things where it's like, “I know exactly where I want to go, just get me there as quickly as possible. I don't need this point-and-click thing.” Actually, this command line terminal experience, where performance is the key bit, something like Command E just makes total sense.

Corey: One thing I've been always trying to understand about Marc Andreessen—and apparently, I'm nowhere near alone in this, but he went from following me on Twitter one day to blocking me, and to this day I have no idea why that is. So, if you're listening to this somewhere, Mark, I invite you at any time to come on this show and explain yourself. I would love to hear it, we can talk about whatever you'd like. I somehow don't think he listens, but we always learn new things as we go. Now, that said, tell me about your team.

It's easy to fall into the myth of the single heroic founder working a day job and then not sleeping and writing code all night. And that really is only accurate for terrible code. That's not the way to write anything well. None of us are an island, and there is a team. It's not just you sitting there by yourself building this. Who else is involved?

Tom: Yeah. So we're kind of an early small team. It's seven of us now. It's myself, five senior engineers, and really great designer who have experienced building some of the best Silicon Valley companies that we all admire, previously worked at companies like WhatsApp, Stripe, and Oracle, and Evernote. So, just people that have been thinking about this problem a lot and seen it at past companies they've been at, and just know how to really build great products, and have a lot of experience with the problem that we're going after.

So that's I think one of the best things about being a founder is you really get to build a team that you're excited to get out of bed every morning and work with towards some goal that you're all excited about. So that's been, I think, the best part of building Command E is just having a really special early team and knowing the possibilities of what the next wave of folks we bring on board will look like.

Corey: I know it's dangerous to ask this given that you're still seed round territory, but if you were to go back to 2018 and do something differently, what would it be?

Tom: You know, that's an interesting question. I always kind of struggle with this one because I'm always going to want to move faster. I think a lot of my job is just looking at this whole engine and saying what can we do to move 10x faster on something. But I've also seen that a lot of what we're doing, it's easy to sort of take all the learnings that we've had to date and say, “Oh, yeah, if we could have just done that six months ago, that would have been awesome.” But I think a lot of this is the compounding benefits of spending more time than anyone in the world thinking about this particular problem, and the pieces come together in a way that you'd love to just remove all the work that got you to where you are today, but I'm not sure that [laugh] that's realistic. So.

Corey: Yeah, make all the right decisions upfront. Easy to say that, but honestly, if we could do that and predict the future, why are we talking about this instead of, “Well, first, I would begin by winning the Powerball six weeks in a row.”

Tom: Yeah.

Corey: Everyone talks about, “Oh, I’d buy this stock or that stock.” Oh, no, no, shoot for the moon.

Tom: Yeah, if you could have told me like, “Hey, the team that you have today, you can have them on day one,” sign me up for that. If you could fast forward all of it, I would love to do that, of course. But.

Corey: Those 40 VCs I met with that didn't invest? Yeah, just skip all those meetings; just go to the one that did is way easier, as it turns out.

Tom: Yeah, that's part of it is just getting to a point where you enjoy the day-to-day and take it as it comes and have faith that you're making progress and just keep chipping away at it. It's a decade to build this to what it should be, I think. So.

Corey: And I'm looking forward to seeing how it gets built and following along from the cheap seats, throwing sarcasm and unsolicited product feedback your way.

Tom: Well, like I said, these are some of the most fun conversations we have is just people that have been looking for this for a while and they find it and share their excitement with us. So yeah, I really appreciate you sharing that with us.

Corey: No, and thank you. I appreciate it. If people want to learn more about you and what you're up to, and of course, download it themselves, where can they find you?

Tom: Yep. Our website is just getcommande.com, just the word ‘Get’ and then ‘Command’ and the letter ‘E’, the way you'd expect it and go there. And it's more about the product and download button right at the top.

Corey: And we will, of course, put a link to that as well into the [show notes 00:32:30]. Thanks so much for taking the time to speak with me. I really appreciate it.

Tom: Yeah, likewise, thanks so much, Corey.

Corey: Tom Uebel, CEO and co-founder of Command E. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you've hated this podcast, please leave a five-star review on your podcast platform of choice and that insulting comment just as soon as you use your crappy search option instead to figure out where those reviews go.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

Links:

  • Linux Foundation: https://www.linuxfoundation.org/
  • OpenJS Foundation: https://openjsf.org/
  • Relying on plain-text email is a 'barrier to entry' for kernel development, says Linux Foundation board member:https://www.theregister.com/2020/08/25/linux_kernel_email/
  • Twitter: https://twitter.com/sarahnovotny
  • Website: https://sarahnovotny.com/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is brought to you in part by our friends at FireHydrant where they want to help you master the mayhem. What does that mean? Well, they’re an incident management platform founded by SREs who couldn’t find the tools they wanted, so they built one. Sounds easy enough. No one’s ever tried that before. Except they’re good at it. Their platform allows teams to create consistency for the entire incident response lifecycle so that your team can focus on fighting fires faster. From alert handoff to retrospectives and everything in between, things like, you know, tracking, communicating, reporting: all the stuff no one cares about. FireHydrant will automate processes for you, so you can focus on resolution. Visit firehydrant.io to get your team started today, and tell them I sent you because I love watching people wince in pain.

Corey: This episode is sponsored in part by LaunchDarkly. Take a look at what it takes to get your code into production. I’m going to just guess that it’s awful because it’s always awful. No one loves their deployment process. What if launching new features didn’t require you to do a full-on code and possibly infrastructure deploy? What if you could test on a small subset of users and then roll it back immediately if results aren’t what you expect? LaunchDarkly does exactly this. To learn more, visit launchdarkly.com and tell them Corey sent you, and watch for the wince.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by—oh boy. Well, I'm joined by Sarah Novotny is probably the best way to frame it, but describing what you do takes a bit of work. Sarah, welcome to the show.

Sarah: Thank you, Corey. It's great to be on the show. And yes, not everyone quite gets the threads that run through my career.

Corey: Right. Even if I want to completely ignore everything that you are, and do, et cetera historically, and only focus on the current stuff, I have options to choose from. You are the open-source wonk in Azure’s Office of the CTO. So, you are a Microsoft employee. You are also on the boards of directors for both the Linux Foundation and Node.js.

Sarah: Actually, the Node.js Foundation has now been folded into the OpenJS foundation. So, that was a big merger between Node.js and the JSF—the JS Foundation.

Corey: Given that, honestly, my biggest challenge with JavaScript is pronouncing it—I keep biasing for ‘YavaScript,’ I am not going to dive into most of that today. So, unfortunately, we only have 20 years of your expertise in the open-source world and now, effectively, setting free and open-source software strategy at Microsoft to delve into. So, you know, only that.

Sarah: Just a bit.

Corey: [laugh]. My word. So, at a very high level, the open-source strategy is now being driven from Microsoft's Office of the CTO apparently. Which is a nice departure from 20 years ago when it was clearly driven by their legal department. So, I have had a bunch of guests on previously, from Microsoft who've talked about ‘the new Microsoft’ and the way that they're pivoting to the point where that is now kind of old news. It's exciting. It's great. It's seeing a level of participation that I am honestly shocked to have seen over the past decade or so. My question for you is—and this is possibly dumb, possibly not—why does Microsoft care about open-source at all?

Sarah: There's actually a really simple answer to it. And it's because customers care. In the same way that AWS tries to put the customer perspective first, Microsoft is always focused on the customer and how the customer is interacting. And if customers want open-source to be able to run in Azure, or inside, or in formation with other Microsoft tools, then you need to be able to work in that composable manner. The industry has changed in the last 20 years.

You've mentioned the historical aspect of this in that 20 years ago, we didn't have nearly the broad industry expertise that we do today, and we were primarily sold vertically integrated stacks by some particular vendor that knew way more about their thing than we ever could in an organization. And so that is changing. There are more and more opinions from the technical staff and teams that have grown up in infrastructure and grown up in technology. And so it's much harder to sell a vertically integrated stack. And so if that's the case, then your customers want Microsoft software to work with all sorts of other things. And you meet your customers where they are, or you risk losing them.

Corey: So I'm not always one to drag the compare and contrast to competition stories here, but something that I think has been pretty clear in the open-source space is there's so much whining about AWS offering a competing service to open-source projects—which just so happened to have a VC backed company behind them, but it's unfair to point that out—whereas Azure does not, to my understanding, get virtually any of that pushback. And it's not because you folks aren't offering things that are open-source aligned. Why is that?

Sarah: We have been spending a lot of effort making sure that we are working in the communities and then also with different companies in an ecosystem to showcase their work as well. So, we have several partner-run services on top of Azure that are run by the open-source company that is best known for the open-source project. So, for example, we have an Azure Databricks service, we have Azure Arrow, Azure Red Hat OpenShift, which is run by Red Hat engineers. And, I mean, we have several more of these because there’s—I think there's an Elastic service, and there's a bunch of services from HashiCorp that are happening. So, it is a matter of trying to work with the community, as opposed to seeing yourself as merely a beneficiary of the community. You have to be willing to put in the work in the communities and in the projects to be able to garner that ability to work with them to build a service.

Corey: Well, to be clear, no one out there has ever accused any of the large cloud providers of being good at naming services, so we're going to just skip past a lot of that snark. And instead, I do—

Sarah: Wait, but why? You don't understand; half of my internal open-source escalations around the names of things.

Corey: Oh, it's also, truth be told—I don't want to necessarily shatter my own myth—it is way harder to name a service well than it is to make fun of a crappy name. So, I do have some sympathy for folks, but in the interest of full disclosure, I don't want to crap on the merit of a service very often because then you wind up with the engineers who built it feeling really bad. Whereas no one spent months, and months, and months naming Azure Data Box, and if they did, maybe they should feel bad about it. So there’s—I understand that the direction that this stuff goes in, but at some point, you just need to call things out as being badly named.

Sarah: [laugh]. Trust me, we do a fair number of that internally, even before some of them get out.

Corey: Oh, if anyone ever wants to leak me stuff—everyone’s like, “What can I leak you if I really hate my employer?” The honest answer: the proposed names for services that never saw the light of day. That's the thing I want to see. I don't want to see big customer lists or your internal P&L. No, no, no, I want to see the names that were considered too bad to put into public because I would have a field day with it.

Sarah: [laugh]. I bet you would. Oh, I bet you would. And that's actually a really interesting differentiation, for you to not want to mock or derive a service because of the engineers, but the name, the naming is still hard, and there's still lots of people who work on it and think about it, but not in the way one might if, say, it was a firstborn child or something.

Corey: Right. It's understood, I think, in marketing organizations, that naming is a series of trade-offs and compromises. Engineering is too, but that's not broadly understood in the general market. And it's easy to wind up taking umbrage from an engineering perspective, in a way that doesn't seem to happen when we're talking about naming, by and large. There are always going to be exceptions.

And there are times the service and or the name both absolutely suck to the point where saying otherwise is borderline malpractice, but it's not necessarily something that is going to rub people the wrong way in quite the same direction. Or at least I don't get the letters about naming that I do about insulting engineering teams.

Sarah: [laugh]. Okay.s, good to know. Perhaps that's an audience thing for you, too. Maybe there are fewer marketing people who are listening to your podcast.

Corey: Yeah, the sad thing is, is I looked up one day and discovered I'd become a marketing person myself, and now I don't recognize myself. But speaking of not recognizing things, what is the story with Microsoft—of all companies—having an employee who is also a board member for the Linux Foundation? That’s one of those oil and water stories if you've been asleep for the last decade. But first, is that on some level considered to be a conflict of interest?

Sarah: It is absolutely not a conflict of interest. And in fact, I represent Microsoft on the Linux Foundation Board. So, I am in Microsoft’s seat on that. We've had a seat on the Linux Foundation Board—we Microsoft—for almost four years now. And when John Gossman actually finally said, “No, we need to do this,”s and a lot the Platinum membership, and participated, and was the first board member for Microsoft there, when he arrived at the Linux Foundation Board meeting the first time, he actually had people clapping and super excited that Microsoft had joined. So it's a complete shift from 20 years ago when I was living here in Seattle, and would not even have considered if a recruiter from Microsoft called me. It's completely different now.

Corey: It really is. And again, I used to be a Windows admin very briefly and got out of it, just because I couldn't stand aspects of what I had to do to run Microsoft stack; went to Linux around 2006 and never went back. And looking back, everything I predicted about hating Microsoft—oh, they're going to be awful in Cloud; they're going to wind up crushing Linux; like the GitHub acquisition, everyone was extremely skeptical about it. Looking at it now, and I think that if you're still beating that war drum, you're kind of a relic, I've got to say. It's very clear that the things that the general open-source community feared did not come to pass full stop. And to be very honest with you, I really question whether any company including Microsoft has the kind of long game to wind up submerging the evil for 30 years.

Sarah: Very few people are that good at strategy.

Corey: Or are, frankly, that driven by spite. I'm one of them, but not too many others.

Sarah: Okay, note to self: don't get Corey mad at me.

Corey: Exactly. So, as I look now, across the open-source landscape, everyone who is doing open-source projects, by and large, has them living on GitHub, it's the de facto place for things to be. And an awful lot of folks who, unlike me, don't have 20 years of muscle memory built up in VI are reaching for an editor, it's going to be Microsoft Visual Studio Code. And that's the early developer ecosystem play that really seems to be paying dividends for Microsoft across the board. I maintain on some level when there winds up being a click the button to seamlessly deploy this thing to Azure, like Heroku used to be but then got stalled somewhere along the way, if done right, that becomes something transformative across the entire cloud landscape. And I like it because if nothing else, it makes the developer experience better. It should not be hard to get started in this world, the way that it once was.

Sarah: Or in many cases still is.

Corey: Oh, when I look at how to write code, a lot of the intro level tutorials in a lot of places it starts off with, first here's a chapter on learning VI. Why? Why on earth would you subject people to that? I get it 20 years ago, but now that's not really necessary in any meaningful way. It's the wrong direction; it feels like gatekeeping, and—

Sarah: It is gatekeeping.

Corey: —it's awful. It's making it harder for people to get into this rather than easier. First, you have to learn some weird-ass ASCII text editor is not a viable strategy for making what you're doing inclusive and welcoming.

Sarah: Yeah, it's been really interesting to watch some of the longer-standing open-source projects, understand and engage with newer developers and learn from them. And then also, some different projects struggle to bring in new recruits because they don't have the empathy for that world being changed.

Corey: Oh, yeah, “I had to suffer, so other people going through it should suffer as well.” This idea runs counter to the philosophy of ‘send the elevator back down.’ You had to struggle and fight to get to where you are? Great. Shouldn't you want the next person coming up to not have to do all of that? We all stand on the shoulders of giants. None of us know how to work punch cards these days. Let's extend that philosophy.

Sarah: Yeah, it's interesting because you brought up gatekeeping. In the same way, you really want to have a positive-sum game out of all of this work because if we do work well across the clouds, and if we were ever magically to have good interoperability between them, then you end up seeing a way that the customers are so much more successful, and the pie is big enough for all of us. This is the thing: the total addressable market of Cloud is not small, and we do not have to look at this in terms of if I gain one customer, you lose one customer. We can entirely look at it as building the best infrastructure for the customers who need them and then setting up the support that they need to make all of this work well together. It's the same way that—well, I guess I can't say utilities at this point because I was going to say it’s the same way we eventually moved toward electricity as a utility. But at some point, we see—at least in California—the electric utility not taking care have their systems here as well as they should have. So, maybe that's a bad example.

Corey: No, but there is something to be said for the idea of utilities, which is, whenever I turn the faucet on or flip a light switch, I don't sit here questioning whether it's going to work or not. Now, I have some IoT light switches, so maybe I shouldn't question that a bit more than I do. But by and large, there's a level of dependability that is required for folks to consider something a utility. By and large, cloud providers are getting a lot closer to that than anything else once were.

Sarah: Yeah. But with utility comes regulation and that has always been a thing that the cloud providers have been more concerned about, broadly speaking. And regulation and oversight does make it harder to cut corners and go fast, but it also may make for a much more stable infrastructure and a much more interoperable and less monopolistic looking space.

Corey: This episode is sponsored in part by our friends at New Relic. If you’re like most environments, you probably have an incredibly complicated architecture, which means that monitoring it is going to take a dozen different tools. And then we get into the advanced stuff. We all have been there and know that pain, or will learn it shortly, and New Relic wants to change that. They’ve designed everything you need in one platform with pricing that’s simple and straightforward, and that means no more counting hosts. You also can get one user and a hundred gigabytes a month, totally free. To learn more, visit newrelic.com. Observability made simple.

Corey: There's so much that could be done in order to make things even easier, but you need to get the fundamentals right first. You have to not wonder if the virtual machines are going to boot when you tell them to; you have to not wonder whether the storage is going to lose your data due to a weird disk issue. And in a lot of ways, those problems from the early days of Cloud don't really exist in the same way. Now, people are wondering what's possible when you don't have to think about that general baseline level of work that anything needs, and instead you can have your staff start to focus on things that are much more aligned with the actual business problem they're trying to solve.

Sarah: Yeah. Take away the effort that is commoditized. I will always need a server-ish thing that can do compute. Maybe we could be serverless; it could be a VM; it could be a container, but I will need some compute, and I will need some storage probably to put data somewhere. And then how do you interop all of these different things across the different clouds, and how do you do this across the different needs, and use this cloud for the workloads that they are best optimized for and that cloud for the workloads that they're best optimized for.

And just basically hand empowerment back to customers to choose what's right for them as opposed to trying to tell them that they need to buy something to fix a problem from my company or someone else's company instead of saying, “Where are you and what are you doing right now? We need to make sure this works on our systems as well as it can.”

Corey: So, if you want to wind up getting effectively dragged for positions that you didn't intend to make, there are a few better ways to do that than to sit down for an interview with The Register. They're snarky, they're sarcastic. I'm a big fan of them and have been for many years, but you sort of know what you're getting into when that happens. In August of last year, you sat down and had an interview with them, and the headline, which I will quote verbatim, “Relying on plain-text email is a ‘barrier to entry’ for kernel development, says Linux Foundation board member.” Now, tell me more about that, please. Because it’s certainly provocative.

Sarah: It is provocative. There is work happening inside the Linux community, like, looking at ways that our workflow may have new entry points that are more modern, as opposed to the strict model that they use currently, which is their entire workflow is patches through diffs in email, which is super important for the high velocity that the Linux kernel changes have and the size and breadth of that community. And it's limiting to say this is how you need to engage with this project because not as many people today read their mail through a client that can do just straight plain text for real; it's tough actually. And there's lots of opportunity for these larger, longer-standing communities to reach out and look to growing new contributors and new leaders within that community. And that's the only way that these projects will live on another 10 or 20 years.

Or they will have to be sort of wound down and we will have to look toward the new versions, the new operating system that might happen, or the new web server, or container orchestration system, or whatever. If we're not managing and maintaining these open-source projects and working very hard to bring in new views, bring in new people, and to set them up for succession when succession needs to happen, is incredibly important. We saw a challenge with that with the Python community when Guido Rossum ended up saying, “I’m going to step down,” and there had not been a good setup, particularly, for succession there. And so that community struggled for a while. They seem to be doing better again.

Corey: Yeah, there's an awful lot that companies and organizations can do to make things easier. And once upon a time, plaintext email was the way that these things—

Sarah: Was the easiest way. Yeah, absolutely.

Corey: Yeah. And I was one of those militant folks when I ran large-scale email systems that, “Oh, it's plaintext email, or it's trash.” HTML? No, I want plaintext email. And guess what? We lost that fight across the board.

Sarah: Yep. I gave up Mutt, finally.

Corey: Yeah, it's hard to get Gmail or Outlook to send plaintext email only out of the box. So it's gone from—to make sure that this is something everyone can understand. Now, it's become a gatekeeping Shibboleth that stands in the way of a lot of people who otherwise would love to contribute. Because let's face it, anyone who's reading email in something that only speaks ASCII text, I'm sorry, you are so far of a corner case that it's almost a vanishing point here. I pay attention to this because I write a large email newsletter and the number of people who complain about the plaintext version, I think I had to the first year and nothing since. It just isn't a thing that realistic people are using these days. Now, yes, you can make an argument that the deep geeks who are Linux kernel developers should know enough to do that. Sure—

Sarah: Absolutely.

Corey: I can accept that, but why? What is the actual upside value here?

Sarah: Yeah. There is no big benefit in making them spend the first hour or three hours of whatever they're trying to do in an open-source project where most of the time people are offering their time or offering their time as a tiny slice of what they do for their day job, and they then spend a bunch of time just faffing about and trying to figure out how to get something submitted to a mailing list without looking like they don't know what they're doing. And I guarantee that anyone who has already had any little drag on their career—so pretty much every underrepresented minority out there—is not going to be as comfortable trying to go to something like a mailing list, where they know that it's a hard audience, and try to submit something on their first try and get it right and/or know that if they don't get it right that their first interaction will be a, “Resend your patch. Can't read this,” kind of thing or, “Patch doesn't apply smoothly,” or something.

Corey: Yeah. That was the obnoxious part about a lot of open-source communities back when I helped to run the freenode IRC network. It was incredibly off-putting to have someone go to all the effort of building a patch, making sure that it works on their end, to fix a pain that they had—all on a volunteer basis, may I point out—and then being told, “You didn't format your patch correctly.” Or not even being told how, just, “This isn't in the appropriate format. Read the docs.” And it was, “Why am I helping you people? You're thoroughly unpleasant and I'd rather go somewhere else where I get more of an emotional high from helping.”

Sarah: Yeah. It's actually one of the things we have found most in the Kubernetes community that's most responded to, and most reported, and most talked about is that the community is really open and friendly, and tries to encourage people, and takes on a, “Wow, you look like you're struggling. Let me help,” kind of approach as opposed to an RTFM.

Corey: Yeah. It's incredibly off-putting, too. On some level that's why people started paying more attention—to be direct—to Ubuntu than they did Debian because the response they got from the community when they got stuck—and let's face it, everyone gets stuck somewhere, especially in the early days—was such worlds apart. “Go read the manual,” without any indication of what manual they should read—turns out there's a lot of documentation—what exactly are they're getting stuck on? And people spend more time castigating folks for not asking the question properly than they did actually trying to help people.

Sure, these are all volunteers. No one owed anyone any particular level of support, but Ubuntu’s entire community from day one was aimed at making this accessible to people who were not experts in it, and that alone changed the entire way that it was perceived. And it's why an awful lot of listeners know what Ubuntu is but will have to do a bit of quick googling to understand what Debian is.

Sarah: Mm-hm. And the fact that Debian is the upstream from Ubuntu is lost on every—many. Not everyone.

Corey: Oh, sure. But this also is turning into what the world is doing where originally, what distro I use, what operating system I used, were big, contentious issues. Today, I don't have to care about that. I mean, a few people do in very specific roles and that's their entire world. But I don't care necessarily, as things move up the stack, what operating system it's using, whether it's in a container or a serverless function, I just care that there's an API I can bounce something off of, and I get what I want in return from that API.

And it's going to continue to go beyond that down the road, I'm sure. But this stuff is all slipped below the surface of awareness. And the folks who haven't evolved with it feel like relics, to be very honest with you. Not in their focus—higher up the stack. I mean, people working on these things is important, but folks who are convinced that this is all that they'll need to know, where they aren't evolving their skill set, it's sad to see.

Sarah: Yeah. We touched a little bit on gatekeeping, and I think we are in a very much a generational shift, and it might even be the second generational shift within tech in my career time. I'm sure there have been many more since computers began. But there's a point in every career where you kind of get to the top end of an intermediate-level tech person, and you're the one person who knows how to fix this thing, and you're the go to person. And everybody sort of loves that feeling of being the hero, and always being pulled in to fix something for someone, and being the only person who knows or understands it.

But in a world where we're trying to empower people to use good technology to make good decisions and improve the world, we need to make sure that we are opening technology up for everyone who is coming behind us with more documentation, more philosophical writing, more entry points for people to start, more ways that have guided the directional paths so that people can learn more, and not do the othering that used to happen in the network engineers not talking to the systems engineers because network engineering is so much more critical than systems engineering, and all of the weird silos we had made. So, I actually find this fascinating to watch as these generations change that we know we have this full-stack developer concept now. And it's full-stack in today's world because full-stack no longer means having to rack the hardware and actually talk to the person who was giving you your ISDN circuit or whatever technology you were trying to connect at the time. But from ethernet all the way up to somebody interacting with it used to be a lot more difficult because you'd have to have expertise as opposed to an API, as you mentioned.

Corey: So, the last thing I want to really talk to you about, aligned with this is, if you look across the landscape, there's an awful lot of people who are defining themselves by the technology, or the stack, or the project that they work with. And I get it. I started my career as a large-scale email systems administrator. And it became pretty clear after a year of that, that there was a consolidation happening in the email space and that this was not going to be a long term niche that would have rendered me infinitely employable. And my choice was either learn something else to focus on, or double down and try and deny the reality of what's happening in front of me.

And it's a hard choice, I'm not denying that at all. But open-source seems particularly tied to this. When I went onto freenode somewhat recently and showed up at a few of the old channels I used to haunt a decade ago, a lot of the same people are still there giving the same advice with the same tone. The only difference is, a lot fewer people are asking the questions now.

Sarah: Mm-hm. Yeah.

Corey: So, from that perspective, what's the future of open-source?

Sarah: Oh, the future of open-source still is hearty. I actually think it's a grand space ahead of us. I do think that it is going to look very different than open-source of 20 years ago. We've had companies decide to make concerted, organized strategies. This is the work I do and making sure that they are developing in a way and working with the communities and customers that keep them on track to not capture our customers and keep them through lock-in.

We really want to see—and open-source is a great way to allow this—is to open up the cloud communities and the cloud infrastructures to be able to be what the customers need, not what the companies need. And I think that—focusing on the customer—is going to be, and has been proven to be the growth win. If you can work with your customers, get them what they need, and then also occasionally show them a thing that they didn't know they needed, and that they can jump forward technologically. I go back a lot to the Ford quote of, “If I asked my customers what they wanted, they would have wanted faster horses.” And so I do think that there is a lot of need for big technology companies to be looking at what the sea changes are, what the paradigm shifts are.

But I think people need to be answering and offering what their customers need, while then also teasing out the future in there. But open-source specifically, lots of great opportunity, lots of great collaboration across the industry, and lots of practice in all of the very large companies that are trying to work in open-source now. Practice to get it right, practice to be helpful and hopeful and harmless in the communities, and help grow them to be the successes that they need to be, independent of any product may need to be built.

Corey: That's, I think, a terrific place to leave it. If people want to hear more about what you have to say, where can they find you?

Sarah: They can find me on Twitter, and I am just @sarahnovotny. I have a webpage but it mostly is for people who want bios and headshots. Mostly Twitter; Twitter's probably the best place to find me.

Corey: Excellent. We will of course, as always, leave links in the [show notes 00:31:45]. Thank you so much for taking the time to speak with me today. I appreciate it.

Sarah: Yeah, you're most welcome, Corey. It's been fun and it's been fun to meet you for the first time. Clearly, we have to have more nerdy conversations.

Corey: Oh, yes, I'm a treasure and a joy, simultaneously.

Sarah: It's true. It's true, it's true.

Corey: Sarah Novotny, open-source wonk. Azure’s Office of the CTO as well as the member of the boards of directors at the Linux Foundation, Node.js-turned-OpenJS, et cetera, et cetera. And I'm Cloud Economist Corey Quinn. This is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you've hated this podcast, please leave a five-star review on your podcast platform of choice and a comment telling me what I got wrong, written only in ASCII text.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About James

James Urquhart is a Strategic Executive Advisor for VMware Tanzu customers. Mr. Urquhart brings almost 30 years of experience in distributed applications development, deployment, and operations, focusing on software as a complex adaptive system, cloud native applications and platforms, and automation. Prior to joining VMware, via Pivotal, Mr. Urquhart ran product and engineering teams for AWS, SOASTA, and Dell (via Enstratius). Mr. Urquhart has also written and spoken extensively about cloud computing, software agility and the business opportunities they afford.

Mr. Urquhart was named one of the ten most influential people in cloud computing by both the MIT Technology Review and the Huffington Post, and is a former contributing author to GigaOm and CNET. He recently completed a book on event-driven integration for O'Reilly Publishing titled "Flow Architectures: The Future of Event-Driven Integration".

Mr. Urquhart graduated from Macalester College with a Bachelor of Arts in Mathematics and Physics.

Links:

  • VMWare: https://www.vmware.com/
  • Book: Flow Architectures: The Future of Streaming and Event-Driven integration: https://www.amazon.com/Flow-Architectures-Streaming-Event-Driven-Integration/dp/1492075892
  • Twitter: https://twitter.com/jamesurquhart

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by James Urquhart, who is currently a strategic executive advisor at VMware Tanzu. James, welcome to the show.

James: Well, thanks, Corey. I’ve been looking forward to this for a really long time. So.

Corey: So there's a few things I want to talk to you about. But first, I'm going to indulge myself and have a little bit of a diversion. I was at VMworld, I want to say 2019, but don't quote me on that, where there was a keynote that I was live-tweeting, and they talked about VMware Tanzu, in the context of a made-up company called Tanzu Tees, a tee-shirt company. Now, I'm a jerk on Twitter, in case you didn't realize that, and I immediately went after making fun of the overwrought insanity of building a company to do tee-shirts on all of these different technology stacks, and as a result of that, I sort of lost the plot of what Tanzu actually is. So, as a personal favor, can you set me straight on that, please?

James: Yeah, happy to do so. It is, in fact, not a tee-shirt vertical offering.

Corey: And you don't even sell the shirts is the worst part.

James: No. Not even the decals that you iron on like we used to do back in the 70s. No, none of that stuff. Tanzu is a brand, but it's a brand that is focused on essentially how people run modern applications anywhere, pretty much: in public cloud and in their data center portfolio. A number of technologies that are required in that way, I think probably the demo you saw back then, if they had a do-over, they would do sort of the way we talk about things now.

So there's sort of three main buckets to it, there's build, run, and manage. And so the build components are things like Tanzu Application Service, which is the old Pivotal application service. Tanzu Build Service, which is cloud-native build packs. And then a number of technologies related to everything from CI/CD kind of capabilities, the Harbor Repository, you know, all that kind of stuff that's about how do you build and then deploy software in a way that it can then be automated by operations downstream. And then the operations downstream piece is the run piece, and that'll be things like our Kubernetes offerings, Tanzu Kubernetes Grid, the Tanzu Mission Control, which is a multi-cloud multi-cluster management and operations environment for Kubernetes.

And then Service Mesh, Tanzu Service Mesh, and other products that allow you then to basically get the capacity and the connectivity that you need, simply and easily, from both your existing on-premises VMware estate—we actually have TKG embedded in vSphere—and also from your public clouds as well. And then the final piece with the managed piece is—the probably the highlight in that bucket is Tanzu Observability by Wavefront, the old Wavefront application monitoring, and observability environment. Wavefront’s amazing, probably one of my biggest surprises when I joined the company and got to see the demos the first time. It's a really solid product. To be able to see full-stack, if you're an SRE, if you're in security, a number of different ways that you want to be able to kind of see how all the components are working together to deliver the service that you need to deliver and to be able to debug this thing. So that's kind of the heart and soul of what it is, is sort of that kind of level of environment, sitting on top of things like VMware Cloud Foundation, and vSphere, and things like that.

Corey: Which does a lot to explain the challenge that I had because I'm trying to imagine, wow, this product sounds kind of like it's trying to do all things at once. What the hell kind of multifunction printer is this? Understanding that it's a brand takes me one step further and, “Oh, it's a label over a whole bunch of different things.” That suddenly the wool falls away from my eyes. Thank you. That is helpful.

James: Well, I appreciate that.

Corey: So, I didn't bring you here to talk about VMware, surprisingly. Instead, I wanted to talk about a few other things. First and most notably, you have a new book out from O'Reilly, called Flow Architectures with a subtitle that involves a whole bunch of words that I'll mess up, so I'll let you state it.

James: Yeah, the title is Flow Architectures: The Future of Streaming and Event-Driven integration. What all that boils down to is, it's a book that makes a prediction; it's a book that makes a prediction that we are moving, over a period of maybe five to ten years, but we're moving towards a world in which connecting to a stream of events will be done through very standard interfaces much like we connect information with HTTP today in very, very standard ways. And then that movement towards that standardization and the arrival of those standard interfaces and protocols that enable that are going to reduce the cost of integration of real-time data to such an extent that cross-organizational integration will be significantly easier and cheaper to do. And that will lead to an explosion of new applications that are able to respond in near real-time—close to real-time—to the things you do in your everyday life. Or to weather changes, or to a number of other things.

So, you can imagine having a trucking company that might connect to the National Weather Service to get weather data and might connect to a broker for loads for partial truckloads or whatever it might be, it might connect to traffic systems, but be able to do that without having to custom write an API call for several different API's to do that. Just a simple—

Corey: So, this is more of an integration approach to architecture. I’m trying to distill this down to something I guess, more fundamental, where it’s—like, is there a simple, I guess, skeleton case you can wind up giving as an example of this? I mean, sure, it can be ridiculous, obviously goes far beyond this, but give me the tee-shirt company style example.

James: Yeah. Well, I think the big thing that you look for is where are the places today where either it would make sense for two organizations to share data about what's happening but they don't because the use case doesn't quite generate the value to justify the work to get there. So, this might be—for instance, Walmart has a phenomenal real-time inventory program with their suppliers, but a lot of other online shops don't because to do the work to make that really easy to do is a fair investment. The other area is that you can imagine that there are a lot of combinations of data that, if it was cheap enough to experiment, you might find a really powerful way of correlating pieces of data that weren't easy to do before. So there's efforts underway in the scientific community to look at having, sort of, sensors all over the place, and that those sensors collect everything from weather information, and temperature, and the movements of currents, and the vibrations that are being detected in the ground, and so on.

And that you might be able to find really interesting ways of doing things like everything from earthquake prediction to understanding where economies might be the strongest at any given time. Being able to get really easy and cheap access to data that can help you do these amazing things and also the ability then to not have to write a ton of code to take advantage of that, to be able to just say, hey, I want my library or my application to just point this URL right to the stream and I'm now connected, and I'm now receiving data that quickly.

Corey: Gotcha. So it's on some level, I've had customers who have talked to me about replacing a data lake model with a stream processing model. Is that directionally aligned with the idea of flow architectures, or is that a dramatic misunderstanding slash—

James: No, no, no, it’s—

Corey: —oversimplification? Because you say ‘flow’ I hear ‘stream’ and I think, “Ah, those two words mean the same thing in laypeople language.” So, I tend to take shortcuts and be extremely lazy, and people call me a thought leader for it.

James: Yeah. Well, there you go. And that's probably to a certain extent, how I started the research on this as well as just kind of going, “Aha, well, this is all”—like, data moving through our economy is a lot like water running downhill through a sand dune. Right? It picks its paths but as it reinforces certain paths, and as it determines the equivalent of values that it determines that there's a way downhill towards sea level, then it takes that and it reinforces that further.

So yes, you're right, Corey, it's very much about as you move towards stream processing, the question becomes, what if I need streams from external organizations? What if I actually want to find streams where I don't even want a deep contractual or personal relationship with the other party, I just want to be able to go on the internet and say, oh, here's a stream and I want to connect to it? What would it take to get to a world where that is just everyday and commonplace?

Corey: “A lot,” is the short answer. It's pretty clear, that requires a basic—a complete reimagining of how systems interact with one another.

James: Absolutely. And I go through where the current technology status—and I also talk a little bit about how it might move in the future: that gets very speculative, and I'm very clear about that. But with a Wardley Mapping and Promise Theory that I do as a part of this, they are modeling techniques that let me be a little bit more clear about where the likely, kind of, large brain movements are going to be. And I use that then to, sort of, make arguments about what are kind of the key areas that we're going to have to solve in order to get to that state? And to your point a lot of the integration technologies you hear about today are really sort of advanced versions of enterprise service buses, and those kinds of old core queue, ways of doing things.

And those are useful, except they're very focused on sort of having adapters to all the different things, and then having adapters in the middle to translate data from one format to another. Flow gets to the point where we've agreed upon the protocols and the interfaces to allow us to already understand how metadata is going to be communicated, no matter where it's coming from, where it's going to, and to already understand what are the interfaces to subscribe to a stream to signal, “Hey, I'm getting overwhelmed. Hold off for a second for me,” or whatever it might be. And that eliminates a lot of sort of that having to custom code for all the different players’ pieces. And there's more to it than that, but the fundamental change I see is moving from a highly contextual environment where you have to predefine all these things and you have predefined places you can plug things into, to a composable environment much more like the Linux command line and the pipe command in Linux command lines, where you can just string a bunch of commands together and the data passes between those—in this case would be services and you string a bunch of services together, and the data just passes between services, without a lot of additional work on your part to make it work.

Corey: It sounds like it's one of those things where it's easy to wind up waxing poetic—or snarky, depending on how it works. Sometimes I rhyme and make snarky poetry—in a tweet. But it sounds like this is a deep field that as soon as you scratch beneath the surface, it becomes geometrically more complex. Help me envision the book since I don't have a copy of it yet. I accidentally knock it off the table: how dangerous is it to a dog if it hits the dog? If the dog is small, medium or large?

James: No, no. It's a book, but it's not a thick book. It's definitely an animal cover O'Reilly book, but it's towards the thin side of one of those. Not as thin as the I Love Logging book, but not as thick as most of the programming manuals you're going to come across. It’s six chapters, and six chapters that are all within sort of a reasonable number of, sort of, 25 to 40 pages or something along those lines.

I hope I wrote it as something that's a pretty easy read, that it steps you through from base principles and core principles through the modeling techniques, and then what fits into the models, and then how you use the model to kind of predict the future and what you can do today. And do that in a way that I take you from point to point to point where I don't lose you by sort of jumping to a new term or a new technology or a new something without having given some context for a first.

Corey: A skill I would definitely benefit from with most of my tweeting and almost psychotic shifting gears without a clutch. It's always fun trying to see how you take something that would fit in either a tweet or a tweet thread or a blog post, but then taking that idea of ‘flow’—if you pardon me overloading the term and turning that into a narrative that makes sense with a start, middle, and an end for something as long-form as a book. I'm always in awe of people who have the, basically, attention span to wind up writing books. I do not have that skill.

James: [laugh]. Dude, that's a great term for it, too because it took a little over a year to really write this and there are parts of this—there's in a whole big, large appendix that's sort of an inventory of the technologies that I came across that helped inform all of this and some explanations about their import. And man, it is an attention span thing. It is definitely something where you're constantly going back and rereading the previous section to make sure you're flowing into the current suction. So, yeah, the writing of it was a labor of love and I'm really glad I got it out there because it certainly has been a topic that, I think, it's under-appreciated for the impact it's going to have on the way we economically connect with each other maybe 10, 15, 20 years in the future.

Corey: Is this your first book?

James: It is my first book.

Corey: Oh. So, every time I wind up talking to someone who has been down this path, and I suggest, “Hey, should I write a book?” They get this haunted thousand-yard stare in their eyes before they begin screaming, “No!” My working theory is that no one actually wants to write a book; they want to have written the book.

James: But at the risk of maybe offending some of your listeners, it was explained to me in a really good way, which was like, “Look, have you ever been curious what it's like to give birth?” [laugh].

Corey: Yeah, it feels like that's the equivalent. It's one of those, “Oh, kids are great. You should have kids.” You know who says that? That's right. Parents because we want non-parents to be just as miserable as we are from time to time.

James: Yeah. No, and it's a hard, laborious process to go through, but if you're like me and I always say my motto for the last 15, 20 years has been I write to learn. I write to learn things. And it absolutely did that. So, to go through the process and come out the other side with something I'm holding in my hands, the satisfaction from that is immense.

And just the feedback I will get, whether it's positive or negative, the feedback I get from this point on, it will only help inform and enrich my understanding of the technology world we live in today. And I'm grateful for that. I'm grateful for every reader that provides any form of feedback, regardless of what it is.

Corey: Yeah, it's one of those massive undertakings that someday I'd like to do it, I just—again, I lie; I just contradicted myself. Someday I would like to have written a book. After some point, it's like, just give me a shot in the arm. It's a year later, and I've written it and, “Wow, that's great. And how come I'm malnourished and I have these scars all over myself?” And that—and we'll know why.

James: Well, Corey, let me real quick, I just want to say, man, I blogged for a long time. And at the beginning of the cloud era, I wrote a blog called, The Wisdom Of Clouds that was considered one of sort of a very influenced set of blogs that were out there. And for all that writing and all that attention that I got, and all of the key points I made that entrepreneurs later told me, like, “That was the key that got me on the right track with my product.” And okay, that's great except it does not have the retention. It's not something you can point to as easily out there and say, “Look at this body of work I did.”

At this point in time, it's been probably approaching eight years, nine years since I last wrote that blog. And people who knew it, well know that I wrote it, but people who don't know me very well, I say, I wrote The Wisdom Of Clouds. And they're like, “What—okay. I have no idea what that is.” And so a book is a little bit more something that you can go back and say, “Well, it may be out of date now, might be something you’re using to prop up a monitor somewhere, but I wrote it and it's there.”

Corey: I remember this. But I'm going to go ahead and admit to my own sacred shame because I remember citing parts of what you wrote many years ago. So, I have a almost perfect track record of being completely wrong on foundational shifts. I thought virtualization—early on—was going to be a niche thing. Sure, for edge stuff, but most workloads are not going to wind up working there.

Yeah, we can safely say I was wrong. Cloud: I was very against it early on, in part because my identity at the time was around running systems manually in data centers and it was hard for me to accept that it might have to change that. But there were arguments against it, and some of the economic stuff that you wound up talking about in the early example, was a great way that this was demonstrated. And then I was down in containers for a while, I made fun of Kubernetes a fair bit because it's easy to do. And then I figured just for fun right now, I’m taking the opposite approach of I'm going to be a big fan and champion of the idea of serverless computing, which based upon my track record all but guarantees it will be a flop.

James: [laugh]. Yeah, I got to say, I mean, I have a great history and of making these predictions about—well, it's really obvious to me that this is going to work or this isn’t. I think, the one I always laugh at as I go back, I have this blog post before I even ended up at CNET on the very earliest version of the blog, where I was absolutely convinced since AWS already had SimpleDB, it was already out there, they had no need for DynamoDB, and it was going to be just this tiny thing that really wasn't going to make a difference. Now, of course, I didn't know anything about the relative architectures of the two, and so that was purely crazy. There are a number of other things about how big and competitive the other two clouds were going to be.

I mean, I think we all correctly predicted that Microsoft and Google were going to be the other players in the space, but I don't think we realized how hard it was going to be for them to catch up to AWS in any way, shape or form. And I think this serverless thing has legs if that helps, Corey because it fits into the flow model really well, as well. Serverless is generally about consuming events and creating events, and building applications by linking these tiny things together with events. And so I do believe it has—I would love to see AWS more involved in some of the standardization efforts. They were involved for a while with the cloud event stuff; I'm sure there are still people that are involved with it in some way.

I would love to see the other cloud providers also being aware of the fact that while a nice profitable system may be running in their environment, there may be people in other environments that need to consume events from that system, and facilitating that rather than making it hard to do. It’s data egress, so maybe they can even make money off of it, right? They're so good at that. But that's my thoughts on it. I think serverless is—to me, Functions as a Service is the lowest granularity you can get to—

Corey: The least interesting part of all of it, too. It's the event model that is truly transformational, and it only really, I guess, now occurs to me that this feels like it is directly in line with a—you guessed it—flow architecture.

James: Yeah, it absolutely is. And it's also, when you think about what developers deliver, we used to make them build servers, right, and they didn't care about servers. So, now we got down to the point with containers that we make them build processes and package those processes. That's way closer to what they actually do, they build an executable, they put that in a container, that makes a ton of sense. But when you really look at the unit of work for a developer on a day-by-day basis, it’s likely measured in functions, not executables.

And so if you can successfully break the environment down so that a developer can work in that unit of work and deliver—every time they modify a function or group of functions and solve a problem, they can deploy that really quickly, and the system is instantly updated with its new powers and capabilities, that's awesome. There are challenges here. And that's one of the reasons why I think the major cloud providers in this space are going to end up looking a lot like operating system, or compiler companies. Or they're going to have to figure out ways to have their infrastructure optimize the placement of functions and optimize the technologies that connect and store data and everything else so that they work really efficiently and fast and don't have network latency in every single case.

But I believe that it will only get better over time and it’ll only be more reasonable for you to say, “Hey, look we don't need to build a full big package service in a single executable and deploy that executable in this case. What we really need to do is just deploy a set of functions that are then optimally working with each other within the cloud because the cloud provider makes it so.” And all of that is, again, consuming events and sending events calling API's—by the way, APIs don't go away; you absolutely need the call-response piece as well where it makes sense to do so. And so it all kind of fits together to me really nicely. And it is a simpler, lower toil way for developers to work.

Corey: This episode is sponsored in part by our friends at New Relic. If you’re like most environments, you probably have an incredibly complicated architecture, which means that monitoring it is going to take a dozen different tools. And then we get into the advanced stuff. We all have been there and know that pain, or will learn it shortly, and New Relic wants to change that. They’ve designed everything you need in one platform with pricing that’s simple and straightforward, and that means no more counting hosts. You also can get one user and a hundred gigabytes a month, totally free. To learn more, visit newrelic.com. Observability made simple.

Corey: It's extremely gracious of you to agree to talk with me on this show. Now, let's see if I can make you regret it.

James: [laugh]. Right.

Corey: You did a lot of predictions on this. How would you say that cloud computing evolved differently than you expected it would a decade ago?

James: Yeah, a myriad of answers to that. I think when you looked at where we were seeing the first work with EC2, and then even the turn of the last decade was just an emergence of other patterns besides just compute and the ability to get services that did more higher-level things for you, things like step functions and other things coming out at that time. I think we were thinking it was going to remain very sort of infrastructure, and maybe data. So, the things that were infrastructure components for an application where you might buy it from a vendor, it's now a utility. But I think what's happened is in fact, the most surprising to me is that the cloud is able to invent new forms of working much more easily and much more cost effectively than a lot of on-premises or licensed lenders can do.

And so I think you're seeing a big explosion in the variety of things that can happen, and a slow move towards maybe a little bit more, not so much vertical market focus, but a little bit more kind of niche need focused set of services that are out there. And I don't think we thought the niche needs services would come from the major cloud providers back then. The other thing is, the growth was unbelievable. The growth is consistently been so damn big, that you get a sense that they created a Jevons Paradox situation where they made computing so accessible and so cost effective for inventing that it didn't mean we spent less money on inventing, it meant there were just a massive amount more invention going on out there. And that, to me, signals really, really well for the opportunities that remain in the future, the opportunities that will continue to be created, and that this is really sort of a model that there may be business models that change around Cloud, but this model of running compute as utility is here to stay.

Corey: I think the entire idea of not having to focus on the baseline stuff just to get a rack up and running in order to start experimenting is huge. Ten, fifteen years ago—I keep forgetting time is advancing. I really should have framed this as fifteen instead of ten years—but originally, it was, “Well, you'd better have a very specific niche use case to put something in the cloud.” Now, it's, “You probably should have a very specific nice use case to justify not putting something in the cloud.” I'm not saying those use cases don't exist, but it’s—

James: Yeah. Yeah, I always tell people, there is a financial argument to say that if you have giant scale, then maybe you need to build your own cloud data center. But the scale that we're talking about is approaching Facebook. It's not even target retail scale.

Corey: Yeah, if Google were an AWS customer, I would have some thoughts that maybe they might want to at least run a cost benefit analysis on running their own gear. So yeah, there are an awful lot of shops out there, but it is not the 100 million dollar a year range, I promise.

James: Yeah. I do believe though, that that's going to shift and change. It’s never going to be, okay, on-prem is dead. For certain companies that very well may be true, but I do believe that there are—

Corey: Oh, that's not true at all. I mean, I've worked a lot of places where I pushed the wrong button or trip over the wrong cable, and yep, on-prem is now dead.

James: [laugh]. Well, I've had a couple of situations where I hit the wrong button in the AWS console, and the cloud’s dead, too. But yeah, it's an interesting world, it is a shifting world, it is a world where we have—the public cloud is fulfilling its promise in a huge way. And the biggest thing that I look for is when does the international competition start picking up. When do the Alibabas and the companies that may come out of even Russia or Europe, when did those companies start finding market needs and market niches in the US where they actually begin to get a footprint? Because I do believe that the big three as we know them today, they're not safe if you look at a 20-year timeframe. I think there's a lot of opportunity for disruption, but it has to be another major large player that knows how to run many, many large data centers in order to do that.

Corey: Oh, yeah, the fact that that used to be a skill set that every company needed to have, and now does not, that's transformative. I mean, I know that there's a lot of talk about the big companies out there, but I like to focus the other end of the spectrum: the random, ridiculous small business idea today, that maybe, just maybe, not saying this is guaranteed or even likely, but will one day become a component of a large index, and they're publicly traded in the rest, the barriers to entry are lower and lower because you don't need to raise a bunch of money just to get a computer to run this stuff on. It's now free trials or credit cards, I can spend six hours and evening putting together the bones of an idea before discovering that it's invariably a terrible idea, turn the whole thing off. And my bill is, what, $7 for that, plus the 22 cents in perpetuity for some weird resource and an AWS thing that I will be paying until I die.

James: Yeah. Well, that's the story, right? You know, there's a number of us that wrote early on that really the thing that cloud change was—it used to say the cost of experimenting, and then experimenting again went way, way, way down. And that is ultimately what you want to see as somebody who wants to create a new technology on somebody else's technology; you want to see that your cost of taking a shot at something and it not working out is significantly low. So, I see a myriad of businesses out there that are super cool and even disrupting businesses that were already built on AWS to disrupt their previous market. You're seeing the second generation come along already. And that's because people can keep trying and keep trying until they find a product market fit that gives them a chance.

Corey: Yeah, and there's value to that and validity to it. And it really drives home the unit economics of doing this at tremendous scale. And not only that, if there's an outage because—surprise—they’re computers; they break. When you have a large scale cloud provider—I care not which one—there is a team of some of the best people in the world at that particular subject, rushing to fix it. You won't have that, to some extent. There's that aspect. There’s the business headline story as well, where it's not your site is down today, it's your cloud provider took an outage, and here's a giant list of businesses that are impacted. And if you're lucky, you might make that list.

James: Yeah. Well, on the operation side of things and the personnel there, I worked at AWS for 18 months. I can tell you, I sat in on those operations calls; there is no enterprise that I am aware of anywhere that runs anything like what the professionalism and the science that they apply to operating their environment. And the incentives they have for people to be on their game are phenomenal and so-and I know all the cloud providers have their version of that. And that's the one thing that I absolutely agree with is that you're never going to have a organization that’s as focused at optimizing exactly what they need to do to have a phenomenally high-quality product at a reasonable price with the flexibility they need. You'll never have that in your own environment because there's just too many variables and that you need too many people that are focused on that day by day by day to really hit a home run with it.

Corey: Yeah, it's one of those areas where subject expertise really matters and counts for a lot. I want to thank you for taking the time to speak with me today. If people want to learn more about what you're up to, and of course, buy a copy of your book, where can they find you?

James: Yeah, so I'm by far most active on Twitter where it's just my first and last name, so J-A-M-E-S-U-R-Q-U-H-A-R-T. The book is available on Amazon. Again, it's Flow Architectures: The Future of Streaming and Event-Driven integration. It's available both in Kindle and paperback form. I'm told that it's on Amazon Books already and on a number of other e-book sites, so check your environment. And welcome honest reviews and honest feedback from anybody who reads the book and I would look forward to that. And certainly, reach out to me, anybody who'd like to follow up and have any deeper conversation either on the event-driven stuff or anything else we talked about today.

Corey: Excellent. We'll of course throw links to that in the show notes. Thank you so much for taking the time and I really appreciate it.

James: Thank you so much, Corey. Have a good day.

Corey: James Urquhart, strategic executive advisor at VMware Tanzu. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you hated this podcast, please leave a five-star review on your podcast platform of choice, along with a comment telling me that why data lakes are the future and that streaming is a red herring.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Nora

Nora is the founder and CEO of Jeli. She is a dedicated and driven technology leader and software engineer with a passion for the intersection between how people and software work in practice in distributed systems. In November 2017 she keynoted at AWS re:Invent to share her experiences helping organizations large and small reach crucial availability with an audience of ~40,000 people, helping kick off the Chaos Engineering movement we see today. She created and founded the www.learningfromincidents.io movement to develop and open-source cross-organization learnings and analysis from reliability incidents across various organizations, and the business impacts of doing so.

Links:

  • Jeli main webpage: https://www.jeli.io/
  • Chaos Engineering Book: https://www.amazon.com/Chaos-Engineering-System-Resiliency-Practice/dp/1492043869
  • Learning From Incidents: https://www.learningfromincidents.io/
  • Jeli contact us form: https://www.jeli.io/contact-us/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by LaunchDarkly. Take a look at what it takes to get your code into production. I’m going to just guess that it’s awful because it’s always awful. No one loves their deployment process. What if launching new features didn’t require you to do a full-on code and possibly infrastructure deploy? What if you could test on a small subset of users and then roll it back immediately if results aren’t what you expect? LaunchDarkly does exactly this. To learn more, visit launchdarkly.com and tell them Corey sent you, and watch for the wince.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Norah Jones who, despite having a storied history, is probably best known these days for being the founder and CEO of Jeli. Nora, welcome to the show.

Nora: Thank you, Corey.

Corey: So that's Jeli—J-E-L-I. I’ll avoid the various jam puns we could wind up going with. Let's start at the very beginning. What the heck is Jeli?

Nora: So, first of all, please never stop using the puns. [laugh]. We all use puns internally; we call ourselves ‘Jeli Beans,’ so it's complete pun-dom in our Slack. But Jeli is an incident analysis platform. So we've built the first incident analysis platform that allows companies to not only learn from their incidents, but address everything they can from them so that they're actually understanding what's contributing to some of their major failure modes, what they're doing well at versus what they think they're doing well at, and exposing the delta between those two worlds.

And honestly, using incidents as a catalyst for helping orgs understand themselves better so that they can make better decisions. This can lead to things like helping them with their OKRs, helping them with their headcount on teams, it can lead to a number of things. Really, Jeli is using the incident as a catalyst for helping you understand how you think your org works versus how it actually works.

Corey: Okay, let's back up a little bit. You have a history as a software engineer. You were at Jet, and then you wound up at Netflix, you quite literally wrote the book on chaos engineering, then you went to Slack, and now you found a company of your own that's aimed at this. What's the common thread there?

Nora: So, I was actually in hardware prior to Jet and I was working on reliability there. And I've been focused on developer productivity, reliability, honestly, my entire career. And I was seeing some of almost the exact same flavor of incidents happening at some of the companies I was working at.

Corey: It's always [DNS 00:04:03], a disk fills up that sort of thing, or are you talking about something beyond that?

Nora: Honestly, even with certain tools. Like I love console, I think it's a really great mechanism, but I was seeing the same types of console incidents at certain companies, like, five or six years apart from working at them. And I thought that was kind of incredible how folks were using it in the same unintended ways that were leading to certain failure modes. And I just started thinking, “Wow, it would be really helpful if we, as an industry even started sharing this stuff with each other a little bit more.” And also even giving these companies tools to understand it more.

When I was at Jet, we were having incidents pretty regularly. Our marketing team was crushing it and we were just growing really fast. And it was a trade-off at the time and we had amazing engineers. It was just, it was a lot of speed. And so we were trying to figure out what to do with incidents at that time.

Corey: Well, when you say, a lot of incidents, I mean, I've worked in shops where that means—so how many—s what is a lot? And everyone's going to have questions about that. From my perspective, it's, “Oh, we've had eight incidents.” “Oh, well, of course. Of your company? That’s not that—” “No, no. This morning.”

And it comes down to at that point when everything's an incident, nothing is, and everyone feels sort of trapped into wherever they are because architectural decisions, because of business requirements, et cetera. And it's easy to sit here without running an infrastructure myself and have these conversations. But when you're in the middle of it, it feels exhausting, never-ending, and the rest.

Nora: And that's such a good point, Corey, is that what is an incident at our company? And I've challenged a few orgs that I've worked with in the past to ask—I’ll say, “If I asked five different people at this company, ‘what is an incident here?’ How many different answers am I going to get?” And I'll usually get the answer, “Five.” Some folks will be cheeky, and say, “You'll get 10 answers.”

But that's the problem at some of these organizations is that when everything's an incident, nothing is. And so there needs to be some key business metrics, and that has to be something that people can consistently answer from legal all the way to engineering, to marketing. Everyone should be able to have a consistent answer on what an incident means. And I don't mean, like, a document that you can pull up that's five pages, that I have to figure out if this is an incident, what level incident and I'm trying to find the document, I've wasted ten minutes at this point, and it's two in the morning, and I'm really tired. None of that should happen.

This should just be kind of a consistent KPI-level metric that you can grok and people can pool on. And I realize that's different with different products, but at Jet, we were having a lot of incidents. And when I say incidents, I mean we were having incidents every day. And I worked with amazing folks there that were working around the clock to pull things back together and some of the best and brightest engineers I've ever seen. But when I went to Netflix, I noticed that everyone kind of knew what an incident was at Netflix.

Everyone knew what the key business metrics were, and how they impacted things, and it was just, there was this alignment. And so it was never a question on whether something was an incident, on whether something was worth waking up at two in the morning for, it was just understood and baked into the fabric of the culture pretty early on.

Corey: Unfortunately, this kind of doesn't help in some respects because it feels like it's just another example of, “Oh, well, Netflix is of course, otherworldly, and far beyond what any other mortal company could wind up doing.” And I don't know that that's necessarily true. But it also feels reminiscent of chaos engineering, insofar as getting buy-in to fix things by breaking them on purpose is often a very heavy lift for folks who can't get to a point of stability. Similarly, it feels like learning from incidents is going to be very hard with respect to finding the time to do it when you're buried in them. It almost feels like you have to educate your customer before you can help them. Is that at all accurate or am I misunderstanding something dramatic?

Nora: No, it's a really interesting point, Corey. And when I was at Jet when we were having all of those incidents, we were kind of reaching a point where we were like, “Let's just try a few different things.” And I think a lot of companies reach that point where they're willing to try something, where they're not wanting to wake their engineers up anymore. And so that's when we started trying chaos engineering. And it was helpful from the perspective of helping us understand our culture a little bit more in who we need to rely on.

The hard part was figuring out what to fail, where to inject chaos, and even what to do with the results afterwards. As a lot of software engineers do, it was kind of like thought of it almost as a, let’s automate it away situation: we can just have a tool running in the background, but that kind of defeats the purpose of chaos engineering. And so when I went to Netflix, I was really excited to join the chaos engineering team there because Netflix had made this percolate and work throughout the company, and I was really excited to be in an organization where it was so widely understood. But as I worked on the team a bit more, I mean, I was programming probably, like, 80% of my time at Netflix, and as was the rest of the chaos engineering team, but when I went to check who was actually using the chaos engineering tools on our team, it was mostly the four of us building these tools. Which was fine from a certain perspective, but we weren't getting enough ROI out of it.

The whole purpose of chaos engineering is to actually learn where your weak spots are so that you can be a bit more proactive to them. And we were focusing very much on the injecting failure, and we were really focused on mitigating the blast radius so we could do it safely. We were doing some very fascinating things technically, and we were doing some really great stuff with distributed systems in general and working with other teams, but what we weren't super focused on was the creating the experiment, and what to do with the results. And so it was usually us nudging people to create experiments, some folks would put them in their continuous deployment pipelines and stuff, which can add a little bit more of the benefit. But actually sitting and taking the time to think about where you want to experiment and what you want to do with the results, causes, I think, a bit more ROI from doing chaos engineering.

And so I realized those were problem areas at Netflix. And so I started analyzing incidents to try to make a catalyst for like, “Okay, here's where we should create chaos experiments, and here are the areas where we need to do a lot of stuff with the results.” Basically, I started looking at incidents to try to help my chaos tools a little bit. And then I realized there was so much more to learning from incidents than that. I was finding things like, wow, we bring in this particular engineer all the time.

They are a knowledge island in this organization. Or this team is severely underwater right now. Maybe we should staff them up. And so, yes, it was helpful in informing where we should chaos experiment and what we should do with those results, like if we should prioritize them in our action items, but it was also helpful in a number of things in the business. And we started writing incident reports that were getting read by folks all over the company.

And people were learning more about the system because we had taken a deeper look at analyzing these incidents. And when I say analyzing these incidents, I mean looking through the chat transcripts, understanding who got paged, figuring out what team they were on, what tenure they have, things like that. And so that was clearly beneficial, and it was stuff I started doing at Slack, too, but it was a lot of manual work. And I can't imagine that most companies would invest that time doing that manual work. And so we wanted to help do some of that for them.

And so, basically, Jeli gives you shoulders to stand on with your incidents so that you're not coming in at zero. We're directing your attention towards places that could use your attention organizationally. And so this post mortem that you're doing is not a chore or a checklist item just because it's part of this process you engrained five years ago; it’s actually something that's useful for you. So we're showing you places that deserve more attention, maybe an engineer that got brought in that we hadn't planned on being there, and understanding what specific knowledge that they had; or understanding that with Kubernetes incidents, we don't do a great job as an organization figuring out the right folks to get in the room; or we throw out a lot of theories before we actually figure out what's going on. These are the things that we're showing you. We're helping you get to those places faster so that you can do your post mortems faster. But we're also enhancing the quality of the output at the same time.

Corey: So, when I look back to my operations days, the dealing with incidents was always obnoxious. Let me walk you through a minor example of one and then you can figure out I guess, well, the audience can figure out more easily what is wrong with the places that I've worked. So, things are breaking, getting the right people on the call is important and almost impossible, so you wind up with your great, great grand-boss on, and you have the CEO breathing down your neck—“Is it fixed? Is it fixed? Is it fixed?”—and then it finally comes up and cool, now it's time to do a post mortem, but we don't call them that, so it's going to be an incident retrospective.

And you're sitting in the room and it's a blameless post mortem. Cool. And you say great. “So, that engineer over there screwed this up.” It's like, whoa, whoa, whoa: blameless. Okay. So, an unnamed engineer screwed this up. And it becomes an iterative process, and invariably, it almost feels like it's a, justify while you're still good at your job, exercise, and a lot of these places. Help.

Nora: Yeah it is.

Corey: How do I fix that?

Nora: It's what people know. If your incident is hitting Twitter, or you have customers calling, that's when someone from your C suite is probably going to jump in and they're probably doing more harm than good. We're actually giving you tools to show you where some of that is hurting. If certain folks jumping in or hurting or helping the situation, we're allowing you to analyze that a little bit better so that you can reduce your costs of coordination during these incidents. And I think costs of coordination are not something that folks tend to look at.

I think a lot of companies look at what quote-unquote, “Caused the incident,” and they look at, quote-unquote, “How to prevent it from ever happening again,” when really, they should also be looking at how they worked together in that moment. Did this involve teams that had never spoken to each other prior to this event? Had this—

Corey: Well, not politely anyway.

Nora: Yeah. Had they been in an incident before? Did the CTO actually hurt jumping in or were they helpful? Who knew the right people to bring in the room? How many hops did it take to get to those right people?

And even imagine a world where you can understand who exactly has the information that you're looking for in a particular moment, and allowing the incident to go much more smoothly. And also just showing you how it didn't go smoothly. I think post mortems and incident reviews—and I don't like the word post mortem either; I tend to use incident review, but if that's a word that you want to hold on to, you can hold on to it, but I recommend making little changes at a time. It's already a super-charged event; removing the word post mortem can make it a little less charged, and I think having a tool to help you during that event to point to areas can also make it a little less awkward of a situation where it doesn't feel as blame-y and finger-pointing, even if you are using the term ‘blameless post mortem’ to describe what that meeting is, it can still sometimes feel like that.

Corey: Whenever you talk about something like a tool in this space, I start getting flashbacks to a number of—I don't want to say failed attempts, but that's really what they were—looking at previous patterns and then making AI or machine learning-driven suggestions about what the outage is likely to be. Which generally means you're trying to swindle someone; if you're not sure who it is, it's probably you. And it looks at previous things and it pops up—if you ever get it to this level of development—with exceedingly unhelpful things. Like, “Last week, a disk filled up. Maybe this time, it's a disk that filled up,” when it is very clearly not that.

And with anything that's driven with suggestion oriented or machine learning-based, it feels like two or three bad suggestions in a row means that no one will trust anything it has to say forever, even if it improves. I mean, take a look at the various digital assistants we have floating around. When you ask Siri to do something and it doesn't work the way you expect it to you feel a bit dumb for having asked in the first place. Never mind the fact that a week later it does that thing; you won't go back and try it again.

Nora: Yeah, I completely agree. And I am so skeptical of AI ops and something that's automating everything for you, I think where we need to go as a software engineering industry—and some insights are helpful, and bubbling those things up for you is super helpful. I think about different things that GSuite does sometimes; some of the automated responses are useful, some of them are not. Setting up the Zoom meetings accordingly. But I think the tools that work the best are the ones that you treat like a member of your team, almost.

It's not something that's doing something for you, it's something that you're working with to achieve the best outcome. And they're still putting something on the table. And that's the mindset that we're building Jeli with. We are showing you some insights, but you still have to do some work on your own. What we're really giving you is a playground to play with your incident a little bit more.

Something that's dedicated and built to help you understand this incident a little bit more. To the point where if you signed on to the incident, we've given you some areas to direct your attention, but you still need to put in some time to understand those areas as you would any post mortem. We want to help you facilitate that discussion so that it's productive, and people leave the discussion feeling like, “Wow, this was really a good discussion for us.” It's not like we are telling you which questions to ask or which things to fix, like the disk uses stuff. We are just giving you focus areas so that you can see themes over time.

And I think how we're really different, too, is we are really focusing on the people and how to best enable and help them. I've done this sort of pattern analysis and incident analysis at a number of organizations, and it is useful, and it can provide a lot of recommendations. And I don't want to just automate exactly what I was doing at these organizations, but I do want to automate the beginning stages of that. And that's what we're doing with Jeli is letting you not start from scratch with a post mortem. Because let's face it, when you get assigned a post mortem, you're like, “I have to remember how to do this. Okay, let me open up a Google doc. Okay, let me pull up 15 Chrome tabs. Okay, I needed to DM so-and-so—”

Corey: Or it’s the other side where you're so used to it, it's habit, and it winds up auto completing automatically. And that doesn't feel great either.

Nora: No, it doesn't. And so, we're helping that be productive. We're helping you look better. We're helping your organization be a bit more collaborative in these events, and just feel more confident about your incidents.

Corey: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the Enterprise (not the starship). On-prem security doesn’t translate well to cloud or multi-cloud environments, and that’s not even counting IoT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IoT devices, detects these threats up to 35 percent faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at extrahop.com/trial.

Corey: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the Enterprise (not the starship). On-prem security doesn’t translate well to cloud or multi-cloud environments, and that’s not even counting IoT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IoT devices, detects these threats up to 35 percent faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at extrahop.com/trial.

Corey: So, help me understand, who's your target customer for something like this? Is it going to be the hyperscale companies who already have attained a certain level of operational maturity? Is it a brand new startup that has just committed their first line of code yesterday and won't figure out until the end of the month that it has a disastrous effect on their AWS bill? Or is it folks in between? I mean, who is your ideal customer these days?

Nora: We're working with a number of different companies right now that are getting value out of Jeli in very different ways. We're working with a Series B company that has about 100 people; we are working with a company that has 10,000 people that's been around for a while; we're working with a company that can measure the exact cost of their incidents at this point. The primary criteria is that it's companies that have incidents right now and they want something a little bit better. I think everyone in the industry doesn't feel great about their post-incident process—I think that's a pretty common thread—and wants it to be a little bit better. And so we want to help that be a more delightful and less, kind of, awkward experience for folks, that they're actually getting value out of, and that it's helping with their internal relationships, it's helping with our customer relationships, it's helping with a number of things.

Corey: So, framing it a little bit differently. If I'm an engineer, and I'm constantly frustrated, by the way that incidents seem to always happen, turn into blame festivals, et cetera, is that something Jeli can help with? In other words, what is the pain that I have that is going to transform into you jumping up and saying, “Yes, yes, that's what we fix.” What is the symptom that lets me know as I walk through the world, that I'm a prospective fit for what Jeli is doing?

Nora: So, the pain that engineers are experiencing right now, or anyone that has to write this post-incident document is a pain around creating a timeline, and copying and pasting items, and figuring out what to focus on in their meeting, and the time it takes to do a good job doing this. So, we reduce that time for you, we make that faster for you, and we enhance the quality of the output. And so the target audience, the target pain point is folks that have pain putting this together today, and feel like it's a chore, and feel like it's seldom a great experience. We want to help make that faster for you. And we want you to have a better and different experience. And so that's really the pain, which is why we're working with an array of companies because no matter where you are, you are having incidents if you're having customers. And so it's it's a matter of how you're addressing those—

Corey: Well, not if you ignore them sufficiently.

Nora: That's true. [laugh].

Corey: But if you do that they become not customers anymore and that sort of solves the problem, but not in the way that anyone really wanted it to.

Nora: Yeah.

Corey: So you've been a software engineer, you've been a senior technical leader at a number of different companies, and now you're a founder. What has changed for you or surprised you the most as you went through that path?

Nora: [laugh]. Yeah, it's an interesting question. A lot of things. I mean, I'm definitely a software engineer at heart. So, I still love architecting and writing code. And it's definitely been a big shift to enabling the folks around me to build in this vision, too, and add to it.

I think that's been, it's not really a surprise because I think we have a really great team, but it's been amazing. It's exactly what I want to be doing right now. I can't imagine doing anything else right now. It's just, I kept having this itch at every company I was at that there has to be something better around incidents. And after seeing these patterns at a number of places, I got the urge to go build it. And there's a lot of folks in the industry that are feeling this pain too, and are dedicated towards making a better solution. And I think that's been a lot of fun.

Corey: It's hard to go back on some level. Once you started a company and had the autonomy, it's scary, and it's hard, and it's one of those I don't ever see a future where I go back to what I used to be. It's sort of a one way door that you never really realized that when you're going through it.

Nora: Yeah, absolutely.

Corey: So, something I've noticed about every company, no matter what it does, I mean for my own where I fix Amazon bills, people have hilarious misunderstandings about it. In my case, it's, “Oh, great. How can I save money on socks?” And the answer is I don't really have a good answer, except I actually kind of do. If you get their Prime credit card, it knocks five percent off, but don't quote me on that. And even if it's something relatively straightforward, people don't always get it. What are the most hilarious misapprehensions you've seen so far about what Jeli does?

Nora: I think some of what you alluded to earlier. It's the AI ops kind of thing. We're certainly providing insights for folks, but you're also participating in the insights. And it's not this AI-focused engine. I think folks are not used to understanding the value that they can get from looking at the chat transcript, but there is so much in there that is just kind of waiting to be analyzed. And I get it, I don't want to go read a Slack conversation after it occurs. And so we're making it easier for you to do that, and glean those insights so that you can get the most value out of them. But.

Corey: Yeah. Assuming you can, that alone is valuable. People have always been saying, “Oh, the chat logs become super valuable just as soon as we learn how to work with them.” And they’ve been saying that for 15 years—

Nora: Right.

Corey: —but I've yet to see it really become valuable. I don't find myself scrolling back to look at how conversations unfolded. I search for specific terms: “Oh, there's the URL I was looking for. Oh there's the image.” Getting more signal than that seems inevitable, but I don't see people doing much with it yet.

Nora: No. And what I was doing at various companies I was at is, I was reading the chat transcript. I would sometimes print them out, go at my desk, highlight them, write notes on them, figure out who the people were, figure out what teams they were on, figure out who was getting paged, and I just, you know, I ended up having [crosstalk 00:26:20]—

Corey: “D minus. You can do better than this, please see me after class.” And then mail it to someone.

Nora: [laugh]. I ended up having a desk that looked like a crime scene where I had sticky notes everywhere, and I had yarn attached to different sticky notes, and just trying to connect all the pieces. Like, an investigation was unfolding because that's exactly what it was. And what we're doing is—

Corey: Oh, if someone didn’t know any better, they’d think you were trying to put together Google's messaging strategy.

Nora: [laugh]. But at Jeli, we're aggregating all of that for folks. So, you can have a more comprehensive picture about how people were coordinating in that moment so that you can reduce those coordination costs in the future. And no one's really looking at that today because it's not easy to do. But there's so much data in there that could actually help you really, really improve and make incidents not such a stressful, time-consuming experience.

Corey: So, one other thing that you've been involved with that I wanted to make sure that we got to talk about was you are also the founder of the learningfromincidents.io. Is it a community? Is it a movement? I'm not entirely clear, but it sounds directly aligned with what you're doing now, what you have been doing, and what Jeli is setting out to solve for. What is the relationship between Learning From Incidents as an entity and Jeli as an entity?

Nora: Yeah. While I was at Netflix and I was getting more deeply into incident analysis, I kind of had this thought, “Wow I really want to talk to folks from other organizations that are also looking at incidents under a deep lens. Surely there are more folks.” And I posted something on Twitter, and I think I got, like, hundreds and hundreds of DMs that night. And I kind of wanted to get like-minded folks together so that we could share our experiences and learn from each other in, kind of like, a safe space.

And so I started a Slack community around that, and I got some great people in it. And as we talked for over a year together, we kind of wanted to open-source some of these learnings. As I was mentioning earlier, it would be so helpful if companies talked a little bit more openly about their incidents and understood that they can do so without revealing proprietary business information. And so we started open-sourcing some of our learnings on the learningfromincidents.iowebsite.

And so it's a community of folks that want to use incidents as a catalyst for helping their organization and helping their businesses, but it's also a place to open source some of those learnings and get stories from folks that are doing it as well. And so, yeah, it's a movement; it's definitely a community; it's both of those things. And I think it's a new take on how the software industry is progressing in the reliability world is kind of taking a more human-centered and learning-focused approach because ultimately it can be really good for your business.

Corey: So we've covered an awful lot of ground over the course of this episode. What are the next steps? What should people who are interested in what you're up to do next if they want to learn more, or figure out whether they're potentially a fit for some of the stuff that you're talking about, and offering solutions to very real painful problems?

Nora: Yeah. So, for learningfromincidents.io, definitely go to the webpage and read some of the posts. There's posts from brilliant folks in that community that are actually doing real things, and chopping the wood, and carrying the water. And my focus with that website was I don't want to talk about the theory, I actually want to do this stuff at companies and have folks talk about how that worked out.

And that's what that website is. And I think if your organization is not feeling like you're getting a lot out of your incidents right now and wants a boost, and you're interested in Jeli, you can use the Contact Us form on our webpage right now and we'll reach out to you to set something up.

Corey: Excellent. We will, of course, include links to that in the [show notes 00:30:12]. Nora, thank you so much for taking the time to speak with me today. I really appreciate it.

Nora: Thanks, Corey.

Corey: Norah Jones, founder, and CEO of Jeli. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you hated this podcast, please leave a five-star review on your podcast platform of choice along with a lengthy comment arguing about exactly whose fault it is.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Ana Visneski

Ana Visneski is a Grandmaster of Disaster (responding to them more so than causing them). She has 15+ years of experience in communications and disaster response. Ana was an officer in the U.S. Coast Guard for 12 years responding to disasters such as Hurricane Katrina and the BP Oilspill. After leaving the service she was the Head of Launch, Blog, and Podcast Operations, before becoming the Head of Global Disaster Response. She is now the Sr. Director of Communications and Community for H2O.ai. She has a Master of Communication in Digital Media and a Master of Communication in Networks. She pronounces AMI the right way...as an acronym.

Links Referenced:

  • H2O.ai: https://www.h2o.ai/
  • Twitter: https://twitter.com/acvisneski

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by longtime, I guess, friend/nemesis/I-don't-even-know-anymore, Ana Visneski, who at least in title these days, if we go by business cards, is the senior director of communications and community at H2O.ai. Ana, welcome to the show.

Ana: Thank you, Corey, it's nice to finally get to join you.

Corey: So, at the beginning, let's start with where you are now: H2O.ai. What is that? By just reading off of the tin, the .ai domain tells me that it's machine-learning powered, and the H2O portion tells me that it's machine learning only watered down. Close to accurate? Not so much?

Ana: Yes, the AI part is accurate, but the reason for the H2O part is actually more of the idea that our founder and CEO, Sri, really wants AI to be accessible to everyone, and as transparent as possible. So, a lot of times when you're working with AI, a lot of what goes into an AI application is in this big black box that no one knows what's inside of it. And the whole idea behind H2O.ai is democratize that, to make it transparent, to make it easy to access, and to make it available to anyone and everyone who needs it or wants to use it.

Corey: People who need to use it or want to use it, but do you take a position on whether people should use it? Because lord knows I do, usually cynical.

Ana: I know. We've had many conversations on Twitter about that. I think that it has applications across the board, but it needs to be used carefully. It needs to be used with the mindset of looking for bias, looking for problematic use. One of the things we do is we are very careful to try and hunt down bias within AI and make sure that if it's being used, it's being used the right way. I think AI could genuinely help the world. As you know, my background is a lot of disaster response. I think AI could genuinely help the world. But as with any good tool, it's double-edged, and it needs to be used carefully.

Corey: Sidestepping the issue of bias in AI entirely, let's instead talk about your background because I've been trying to get you on this show for ages. Your background is fascinating. You spent an inordinate amount of time in the United States Coast Guard. Then you went to AWS where first you ran launch operations for a while and then transitioned to a different team running disaster response. The fact that those two are not the same department gives me some pause and I have a lot of crappy comments around that, but tell me a little bit about, first, why leave the Coast Guard after you were in as long as you were? And then, why AWS?

Ana: So, the Coast Guard, it was just time to go. I love the Coast Guard. I'm third generation. I specialized in crisis communications, disaster response, and search and rescue. A little known fact is I was actually the founder of the Coast Guard social media program, the first official blogger for any armed service, and generally a digital-native pain in their ass.

So, there came a point with the Coast Guard where, as the chief of digital media at headquarters, I’d kind of run out of things to do besides going back to search and rescue, and it was time to go. And then why AWS? Well, I went to grad school with Jeff Barr, actually. And so when Jeff posted about needing someone to help him with the blog, I reached out to him and applied for the job, and went through Amazon's very intense interview process, and ended up coming aboard to help Jeff run the blog. That was my first job was the senior director of digital media for the blog.

And why AWS? It looked super exciting. Obviously, getting the chance to work with Jeff was just amazing, and it kind of tied together a lot of the things I had loved building for the Coast Guard into something I could do as a civilian.

Corey: You were instrumental in my sarcastic birthday video for Jeff, where we redid the tune, “Piano Man” to, “Blogging Man.” You were credited at the end of that as Barr Raiser—with two Rs—because of course you were. And thank you for that.

Ana: You're very welcome. That was fun.

Corey: Everyone has asked, what did Ana do with that? The honest answer is that you knew Jeff super well, and I wanted to surprise him with the video, but I didn't want to go too far in a direction that he would inadvertently find insulting or highlighted something he didn't like. It's all fun and games until you realize that you've basically just inadvertently insulted one of the nicest people in tech. So, that was really the core of what I needed your help with—

Ana: Yep.

Corey: —and you delivered admirably.

Ana: Yeah. Jeff, hands down, is—I would go beyond one of the nicest people in tech. He's one of the most genuinely kind, and thoughtful human beings on the planet. I love that guy.

Corey: I also want to point something out now that it is years later, and neither of us has anything to fear or anything to gain. I got to know you when you were hoping to do a lot of the blog stuff and then transitioned into launch operations, helping handle the orchestration of all the various AWS releases. And although we always maintained a good relationship, never once did you tell me anything that was not public knowledge, that was relevant or germane to a release. You never even pointed at—“Go check this thing over there and tell me what you see.”

Some folks have. You never have. You always were the consummate professional on this, which is fine. I mean, my goal was not and has never been, to beat AWS to the punch at announcing or releasing something. I’d like to make sure it is very clearly in public before I talk about it, otherwise, no one's going to believe that I didn't break NDAs. Every once in a while, I wind up periodically getting angry emails, “You just tweeted about this thing that's under NDA. How could you?” It was, “Well, I didn't know about it until I just read it on the official blog.”

Ana: [laugh].

Corey: “Oh, I guess we released it.” Yeah. No kidding. I'm not here telling other people's secrets. Why would I? There's no benefit to that.

Ana: Yeah. I used to joke that Ariel Kelman, who was the VP at the time, knew my background in disaster response and in the Coast Guard, that I had a top-secret clearance. And so he's like, “She'd be great to run launch.” So, I actually went from just running the one blog to being the head of launch blog and podcast, which is, again, kind of similar to my current title: taking a whole bunch of things that have nothing to do with each other and cramming them into one role.

But yeah, I did. I ended up running launch operations, and that's when you—I believe you tweeted something about waterboarding a VP, and one of my jobs was literally to make sure that things didn't leak, and to make sure that the launch process went well. And for anyone who's seen an Andy Jassy keynote, everything he's talking about on stage, there's a second room where we are making things go live as he speaks, which is intense, to say the least.

Corey: That's right. Yeah, I did a tweet thread on you one day, and my exact tweet quote was—and, like, I used to see early leaks, like documentation push too soon, pricing API updates, new endpoints coming online before they were announced. And one day that all stopped. And I first started to see why while waterboarding an AWS SVP trying to get roadmap information out of them. “I can’t.” He sobbed. “Ana will kill me.”

That was mostly sarcastic, but there was an element of truth to that. It was very clear that what you ended up doing was fixing a number of internal processes. You brought discipline to a previously organic process that, honestly, the organization had very clearly from the outside, outgrown. That was my perception of it, and I'm sure curious as to how accurate that assessment really is.

Ana: So, what's interesting is how fast a company like AWS grows. And like the company I'm at now, H2O, they're growing. And at some point, you have to put in a process to scale. Because you can't keep doing everything manually; you can't keep doing everything seat of your pants. But you also can't put a process in place that is so rigid that when things are naturally chaotic, you can't flex.

So, what I mean by that is that, yeah, I put a process in place around launch, and that's kind of my thing, is looking at chaos, untangling it, and figuring out how to build a process around it. But you can't build something so rigid, that whatever was working before goes away. Does that make sense?

Corey: It absolutely does. And it's easy to forget from the outside because we see our AWS bills getting bigger, we see them releasing a bunch of nonsense that isn't ever going to help us directly or personally, but what we don't see is the fact that every person we used to know at AWS is now at least 10 people as the organization has grown. And the processes that make sense, fail to scale. I mean, enterprise is itself a skill. When you're interviewing someone, “So, what is your nature of your relationship, historically, with the sales division?” And if the response is, “Oh, you mean, Steve?” That's not really the right answer. It's understanding the interplay, organizationally, is super challenging. At AWS more so than most.

Ana: Yes.

Corey: I feel like half my job is still introducing Amazonians to one another.

Ana: [laugh]. That's true. I actually had someone figure out who I was because you had tweeted about me, leaving launch. It was kind of funny— and someone at AWS. So, without talking about AWS too much because it kind of feels like talking about an ex-boyfriend.

You know, I don't want to sound obsessive. But the thing is, is that I think there's an aspect to it, to what I do, that every company can benefit from, and it actually comes from my background as a veteran. And I think this is something a lot of companies miss out on or don't think about when they're looking at veterans as potential employees. A lot of times they look at us for what did we specifically do. And I have friends who are aviators and friends who are boat drivers.

Well, who's going to need a helicopter pilot in tech, right? At the same time, that helicopter pilot understands how to triage operations quickly, understands how to make processional decisions. Like there's all of these things we do that are built into how we're trained and how we learn to process information that is fairly different from the civilian sector. And so honestly, a lot of what I did for launch operations, a lot of what I'm doing now, and I'll tell you, a lot of what I did during 2020 and with disaster response for AWS, was literally taking the same processes I used in a command center to make sure that the operations of that command center were flowing smoothly, and I translated them into corporate and used the same methodology. And it worked really well.

Corey: It's curious to me that skills that can be sharpened in one discipline apply so well to other problem domains while there's still so much of an initial resistance to bringing people in who do not have deep expertise in the problem domain they're focusing on now. It continually baffles me. It's, “Oh, we're looking for someone who's been experienced to”—I don't know, in your case—“watering down AI”—which I know is not what you do. Don't @ me—and has been doing it for at least five years.

It's well, great. You haven't been but the thing that you have been doing very clearly lends itself directly to your current role. And it requires a little bit of vision to find people who have skills that may directly translate, and I guess on some of the willing to, quote-unquote, “Take a chance,” not that it's a big gamble. And oh, wow, it turns out that you don't need to know everything about our business when we hire you.

Ana: It was interesting. When I first got to AWS, there was someone where when I was originally putting in a process around Jeff Barr, I actually won the Marketing Newcomer of the Year award in my first year at AWS, and the big joke was because I played guard dog and protected Jeff Barr. And one of the things I did was I built a process around getting in touch with him and an actual ticketing process for blog posts. And someone hadn't put in any tickets and I went into their office and I said, “Hey, we're not going to be able to do these. They need tickets.”

And I don't remember how the whole conversation went, but the thing that stands out to me is I made a joke going, “Well, maybe it's my military background that I like things to be organized.” And this person said back to me, “Well, then you're not going to fit in very well around here.” Now, fast forward to re:Invent, and said person needed me to figure out how to create two blog posts—This was my very first re:Invent, so what 2016— they needed to figure out how to create two blog posts by the end of a keynote.

And I did it. And you know what I used? That military organization skill. And if you think I was too above it to look at the person and say, “So, how do you like those military organization skills now?” You think I'm a better person than I am because I, of course, did the I-told-you-so dance. But it is interesting that, as a veteran, you run into some really interesting, preconceived notions about what your skills are and aren’t, and what you can and can't do. And that has been an eye-opener for me, joining the civilian sector.

Corey: Now, never having served myself, for a variety of excellent reasons, including they didn't want me when I attempted to enlist when I was 18 years old. There's a lot of, I think, preconceptions people have about those who have served and what it means to be in the military. My dad was a Naval Academy grad, so it was pretty clear from the age of five onward that I was never going to live up to the lofty expectations he held for me. So, when I look at it, it's through the—I have a bunch of nostalgic stories that he would tell, and a whole bunch of myths and falsehoods from those who didn't serve, or the stories of those who did, filtered through same. It's surprising to me that, first, there's this expectation that, “Oh, if you were in the military, clearly, you're only ever good at following orders and not solving anything independently.” It's anyone who's ever worked with a large organization knows that that is provably untrue. It's an easy, funny trope to make, but let's not kid ourselves. If that's all it was, then we would have a very different society than we do now.

Ana: Well, yeah. And so it's always going to depend also— there is a bit of the personality and what did you do when you were in. I can speak for the Coast Guard in that the Coast Guard as a service is smaller than the New York Police Department. There's 39,000—give or take—total in the whole country—or in the whole service, at any given point in time on active duty. That's not a lot of people.

And so that means that we were all expected to kind of be a jack of all trades in some way. You know, as an officer, yeah, my operational specialty was search and rescue but my staff specialty was crisis comms and, basically, public relations. And then because we didn't have anyone to do it and I was the one who had been doing it the longest in my own life, I also ended up as the social media guru who launched all of that. So, there's a lot of stuff that people don't necessarily understand. And you're right.

We do get kind of lumped in is we're not going to be good decision-makers, or we're going to be rigid decision-makers, that we're all very rigid, that we're all very hard, that we all are gun-toting. That was one of my favorites is, I must be into guns if I was in the military. Yes, I know how to shoot them. No, I don't own any of them. I was in the Coast Guard. I liked saving people.

Corey: Oh, the one that always resonated with me, it was, “Oh, your ex-military. You must manage by yelling at people.” It's a common misconception that that's not how military people manage. That's how Israelis manage.

Ana: Well, you did get your yellers. But you also have to figure what most people see of quote-unquote, “Military” is what they see on TV or in the movies. And they're not going to make it look as civilized as it is because it's not as entertaining. When you get a chance, watch the movie The Guardian, have a few beers, call me, and I will explain that entire movie to you.

But we don't yell. And if we yell, it's because someone is literally dying. And that was one of the things that I know, Ariel and my managers at AWS appreciated, and Read, my manager now, appreciates is, I don't flip out. Like… I mean, honestly, once you've been in charge of trying to find someone before they drown, or there's literal fires around you, or I was a responder in Katrina, once you've been a responder to a lot of that stuff, it's pretty hard for you to get spun up about anything else.

Corey: Yeah, it's hard to look at this through a lens of—how do I put this politely—giving a crap when historically, the risk was, oh, people will very possibly die. Whereas then you take it to civilian life, and it's, “Oh, heavens. If this doesn't go correctly, then fewer people might be able to view ads for the next 15 minutes.” And I'm not trying to crap on ad tech’s business models—

Ana: [laugh].

Corey: —but it's also not exactly life or death. One thing I've always appreciated about ex-military is they've never seemed to get too worked up about various outages. I understand I'm stereotyping wildly here, but it's, “Oh, the site is down. Let's go through the runbook. Let's be calm and collected.” And that's really what you need when everything's on fire.

Ana: Well. I could tell you, anyone who's worked with me knows I get fired up.

Corey: Anyone who’s ever worked with me knows that I get fired.

Ana: [laugh]. That's true. I've heard the stories. I will say one interesting thing on the benefit side of being a veteran, is that the fact that veterans quote-unquote, “Yell” or are intense, has helped me a lot in some ways because then when I'm being intense, or I'm being really driven, or I'm fired up about something, it's because I'm a veteran, not because I'm a moody woman.

I've basically been able to use one stereotype to fend off another. Which is a horrible thing to think about when you think about the way it is to be a woman in tech. But I know for a fact, I've probably been able to accomplish things because people assume my personality is based on my training, not just because I've been this way since I was a kid.

Corey: Yeah, it's fun. I find that for better or worse, an awful lot of the preconceptions that people have are, “Oh, it must be this experience you went through that makes you this way.” And not the inverse of, “Oh, being this way, naturally caused you to gravitate towards that experience.” It's a chicken versus egg question.

Ana: It’s true. It is a chicken versus egg question. And for me, I had only been in the Coast Guard, gosh, not even a full year when Hurricane Katrina happened. And I was deployed to help there, and I ended up as the federal on-scene coordinator’s public information officer. And that right there changed the path of my life.

But I wouldn't have pushed to be deployed if I didn't have the personality I have. But honestly, Katrina changed the path, in the sense that in my first year in the Coast Guard, I got to see what a difference we could make if we could communicate to the public where they could find water, where safe haven was. If we could communicate to everyone the best way to help. And I realized that I had a unique skill set because I was digital. I had been blogging for years, I was one of the LiveJournal early adopters.

Boy, that ages me. You know, all of this. And it changed how I viewed what I wanted to do within the Coast Guard. But that actually ends up impacting, too, why I left AWS. This need to be a part of something that helps and a part of something bigger than me.

And I used to joke all the time that leaving the Coast Guard was kind of leaving the Justice League. I didn't have my cape anymore. Who was I? So, that part was pretty interesting. But yeah, it does. It's a chicken or the egg. It's Katrina changed my life, but I wouldn't have been there if it wasn't for my personality and my drive.

Corey: One of the saddest things on some level is when we have something that is transformative like that, in the context of a disaster or other form of crisis, how infrequently it seems that that changes anything. But there are so many lessons there if you care to go after them, and it bugs me every time I see an organization, at least visibly, failing to learn from those things.

Ana: You might have seen on Twitter, I just talked about this, that there's a list of things I really hope people learn from 2020. And when I say people, I specifically mean managers and companies. I'm not even going to get into the government side of it; I could rant about that for three years. But at the end of the day, there's a thing called a black swan event, and you have to understand what a black swan event is. It's basically this idea that no matter how much you plan, you can't actually plan for the exact next event.

Yeah, we can look at what happened in COVID-19, we can look at what happened in 2020, we can try to plan to deal with 2020. But in 2021, something else is going to happen, and it's going to come from a different direction. There will never be another Hurricane Katrina, a Deepwater Horizon, a Fukushima event. Every event is at its core different. So, you have to understand that.

But you should also take a look at what we did that was successful in 2020. What did we do that helped? What did we do that made our businesses successful, that helped our people not burn out? Find those things because those, you can still apply. Even though the disaster might be different, the methodologies can be the same.

Corey: And I think that's sort of the overarching question that I have for you, which is, you've gone from handling blog posts to handling all of launch operations to handling disaster response and multiple companies. What's the common thread?

Ana: The common thread is that at the end of the day, they're all common, if that makes sense. So, what I mean by that is, when I'm running launch operations, each launch is an event. Each launch has a beginning, a middle, and an end. There are things that go wrong. There are things that don't work. Well, when you have a disaster, there's a beginning, a middle, and an end.

Corey: One of the hard parts, of course, is here recording this mid-pandemic. I wish I had your optimism that this disaster will end. And I'm sure that's a common thing for folks in the middle of… trauma, for lack of a better term. And it always does end, but it's hard to wrap my head around that right now and feel that emotionally, which is part of the reason we have experts in these things to guide us through it. Ideally.

Ana: At the end of the day, COVID has been a global disaster. If you look at the number of people who have passed away, who didn't need to if we had locked down the right way, and all of that, I could go on for hours, about the way this crisis has been handled or mishandled. But at the end of the day, I look for the hope in it. There are some great things that have happened, too. And on the days when it just seems like there is no light at the end of the tunnel, and when I'm looking at, I have friends who have gotten it, and I've lost friends, and it's just—it feels to impossible to carry for another day, I look and think, “Okay, well, we were able to find this vaccine faster than ever.” And that to me is amazing. Last year was literally the hardest year of my life. That's including running a re:Invent where I launched 106 net new services and/or features. Okay.

Corey: We've changed yours, incidentally. So, it was two years ago, if we're not—

Ana: [laugh].

Corey: —if we’re being very hon—I assume. You're talking about re:Invent 2019

Ana: ’18 was my last one.

Corey: Oh, okay.

Ana: ’19, remember, I actually got to see you for coffee because I was there and not—

Corey: I just figured you were playing hooky.

Ana: [laugh]. Are you kidding? I think it was. I should have worn an ankle bracelet; they would have electrocuted me. But at the end of the day, now what we need to look at is, okay, what worked last year? We can do a lot more online digitally than we thought.

One of the last things I did at AWS was helped our technical teams and helped AWS build out a good disaster response plan moving forward, including command center style watch rotations for teams so that no one burned out, workloads were more evenly spread, so that if someone does get sick, they don't feel guilty—because, of course, in our culture, you never let anyone else take your work because then you're redundant, right? Well, redundancy is important to survive in a pandemic or in a crisis. So, we need to start looking at those things. And I think one of the things that can give us hope is, look how far we've already survived this? I'm not going to say we've succeeded, or we're survivors yet, but look how far we've already come. And—

Corey: Well, those of us who have survived.

Ana: That is true. Those of us who have survived. And, again, don't get me started on the numbers, Corey because I will go off. But we do have a vaccine, and totally barring the lack of response here in the US by the administration, let's just talk about the vaccine from a technical perspective. That vaccine as you know, AWS, one of the things my team did last year was we had established the Diagnostics Development Initiative, with $20 million going to universities, and hospitals, and scientists to try and help speed up finding a vaccine.

Well, if you look, the vaccine is here, largely because of AI. And largely because of the cloud computing power that was given out and that everyone came together to build. It's not dissimilar to efforts in previous pandemics or previous diseases where everyone put differences aside and came together to try and find the best possible vaccine. And that was a really cool thing to see. So, I think as a responder, the one thing I've learned throughout my entire career is, look for the helpers.

Look for the people who want to get into it with you and just dive in, and dig in, and find ways to help. And so for you, Corey, maybe one way to help with seeing the light at the end of the tunnel is find somewhere to volunteer your expertise. I know you do a lot of volunteering with nonprofits to help them with their AWS bills. Well, maybe volunteer and target nonprofits that deal with suicide prevention and domestic violence response because right now we're seeing spikes in both of those due to the lockdown.

Corey: For better or worse. Most of the nonprofits I speak to in those spaces don't have appreciable cloud bills on some level, which makes sense. I would argue that in many cases, a company does not need a massive cloud bill. And if they have one that's indicative that something's gone wrong somewhere. But your point is well taken.

Ana: Well, and even if they don't need the billing help, as some of them are starting to get into working with AWS, they don't even know how to figure out what their bill says. So, there's ways. And to anyone listening, there's ways for them to help too. There's great hackathons going on, there's ways you can help them home, there are—in some countries they need help—there's mapathons because as they're trying to get the vaccine disseminated, there are certain areas of the world where maps haven't been updated enough to figure out what's a house versus a warehouse versus a school. So, help with the mapathons. That can help people in those areas get the vaccines to the homes they need to be at. So, there's a plethora of ways to help. You just have to look for him.

Corey: So, last question, before I wind up calling this an episode because I figure after this one, you're going to want to get the head start, running. Why did you leave AWS?

Ana: A number of reasons. I kind of alluded to a little bit of it earlier in that 2020 was a very hard year for me. And I don't think enough of us talk about burnout. And I have been through 12 years of the Coast Guard, including night watches, and basically every major disaster response from two thousand—what—four to sixteen. And 2020 was a brutal year.

And I was on 18-hour days. It was starting to physically take a toll on me. So, I decided to take a step back and I took a leave of absence. And during that leave of absence, I really thought about it. I had done all the big stuff I wanted to do with AWS: I had fixed the blog platform, I had built a disaster response program. What did I want to do next?

And I was tired, and I was not giving AWS or my manager—who I adore, I still love her— or Andy, the best of me because I was so tired. And then at the end of the day, the other part of it for me was, remember how I said being in the Coast Guard was kind of like being in the Justice League. At my core, I love helping people I love building, and building the new thing, and making technology accessible and usable to those who might not have it already or not know how to use it. That’s what I did the whole time I was in the Coast Guard was taking these new—at the time, which again makes me sound old—technologies like Twitter, and getting them into the hands where we could actually use it appropriately. And H2O offered me that opportunity.

It's a chance to help a company that genuinely wants to use AI to better and to use AI as a tool in a way that will genuinely help people. Our founder actually started it largely because his mother had cancer and he was having trouble getting data to work with to help with her situation. She is in remission and is doing well, but the whole reason the company was founded was about helping people. And so for me, leaving AWS and coming to H2O, was a little bit more of feeling like I was back to where I was happiest: using a new technology to help in new and in engaging, and amazing ways. Plus getting to build things because it is a younger company; it is a startup.

Being able to dive in and build things from the ground up because you asked earlier what the commonality was between everything I've done. Every time I've taken a new role and every job I've done, there has been a problem that I needed to fix and a system I needed to build. Every time. And it's exciting, so the reason I left was I needed that excitement and I needed that mission. And I was just damn tired. [laugh].

Corey: I can definitely understand that. Ana, thank you for taking the time to tolerate my slings, arrows, and difficult questions. If people want to learn more about what you're up to and follow along on your grand adventure, where can they find you?

Ana: The easiest place is probably on Twitter at @acvisneski. Hard to spell but, you know, I'm sure they'll find me. But that's probably the best place to do it. My Instagram is just my dogs.

Corey: Excellent. I should follow you on Instagram is what I'm taking away from this.

Ana: [laugh]. Yeah, I have very cute dogs. [laugh].

Corey: Ana Visneski, senior director of communications and community at H2O.ai. I'm Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you’ve despised this podcast, please leave a five-star review on your podcast platform of choice and an insulting comment about the military.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Sarah Kaiser

I use lasers to melt acrylic and the cisheteropatriarchy alike. Quantum Computing technologist/consultant by day, author and dog mom the rest of the time.

Links:

  • Unitary Fund: https://unitary.fund/
  • Sarah’s Twitch: https://www.twitch.tv/crazy4pi314/
  • Learning Quantum Computing with Python and Q#: https://www.amazon.com/Learn-Quantum-Computing-Python-hands/dp/1617296139
  • Twitter: https://twitter.com/crazy4pi314
  • GitHub: https://github.com/crazy4pi314
  • Personal Website: https://www.sckaiser.com/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined today by Dr. Sarah Kaiser, who is very recently a technical staff member and quantum community lead at Unitary Fund. Sarah, welcome to the show.

Sarah: Hi, Corey. Thanks for having me.

Corey: So there's a lot to unpack. Let's start with the easy stuff. What is Unitary Fund?

Sarah: Yeah, so we're actually a 501(c)(3) nonprofit that is invested in trying to grow the quantum open-source software community. So, we kind of do two main things: we give out micro-grants to help support maintainers on projects, or other sorts of community groups or educational projects that kind of help grow the community of quantum developers; and we also do—given that we kind of spent a lot of time in the quantum open-source software space—we also have our own kind of internal development team and we work on building up software projects in gaps that we see that maybe are not as interesting or kind of boring, but actually are needed to overall help the ecosystem grow.

Corey: So let's start I guess with a big, meaty topic: quantum computing. It feels like it's something that we've been hearing about for 20 years, usually in almost the same sense as we have cold fusion which is, a practical X is always 20 years away. And over time, recently, we've started seeing actual products and services come out from companies that are talking about exciting breakthroughs from a quantum perspective. But the challenge I've always had is, first, what is quantum computing? Every time I've tried to delve into it, it seems that, “Oh, just go through the ‘Hello World’ example.” And the challenge, of course, is that ‘Hello World’ in quantum computing is basically a PhD. Given that you already have one of those, help. What is it?

Sarah: [laugh]. That's a great question. Honestly, the way I like to think about quantum computing, generally as a field, and in this case, in a literal product sense is, it's a hardware accelerator for computation. So, in the same way that we have, like, GPUs, and FPGAs that are bespoke, custom-designed hardware that can accelerate different—whether it's machine learning or graphics processing, quantum computers are—I kind of loathe that they are called computers; they really should just be called quantum hardware accelerators. [laugh].

Corey: Right. You start calling them computers, I start wondering, okay, and I scroll them and huh. They're not for sale at Best Buy. What's the deal here?

Sarah: Exactly. You're not going to check your email; importantly, they're not going to replace all computers.

Corey: Well, not with that attitude. Okay, not to sound cynical here, but I do look at some of these electronic stores and they're selling $200 audio cables because it's better than the $50 one, and it's passing nothing but digital signal anyway. So, if people will buy anything that's hyped well enough, I mean, there's certainly been enough hype poured into quantum computing to the point where, at least from where I said, it's occluding what it actually is and what it's capable of.

Sarah: Yeah. Think I’d entirely agree with that. I think what is important, as someone who has spent over 10 years of their life researching and working in this field, I still do think there still is really interesting and cool stuff here, it's just maybe not exactly the features or things that are hyped for venture capital funding. [laugh].

Corey: Yeah, part of the challenge with raising VC for something like quantum is that the immediate short term returns are hard to demonstrate. And I'm being very charitable with that. Not from a success or a breakthrough story, but from a position of economic viability. It seems that very often, one of the challenges with quantum computing and explaining what it is, is even articulating the type of problem that a quantum computer is built to solve. Is that a fair assessment? Or am I radically misunderstanding something? Probably both?

Sarah: [laugh]. No, I think you're actually pretty spot-on there. One of the ways I like to think about it is, we have a pretty clear description or box that we can say, “These are the types of problems that say a GPU is good at. We know that any problem that's highly parallelizable, if we can make the solution to our problem fit in those sort of constraints, yep, throw it at a GPU. We're good.”

For quantum computing, where we're at is we basically have examples of things that might be in that box, like we know we can speed up this specific problem, we can speed up this specific problem, but we haven't really worked out what the generalization of—you know, the class of problems or the types of problems that we might be able to solve here. So, it makes it hard to say, “Well, yes. Here's your arbitrary problem.” I actually have to sit down and work through and try to figure out a specific solution to that problem, as opposed to being able to have a general framework, like parallelization or something like that, to break your problem down into something that I can use on the device.

Corey: Please don't take this as the deadly insult that it probably comes across as, but it feels similar in some respects to machine learning, where there was a lot of excitement about it, but every time that someone tried to articulate the real-world business value, it was either aimed at an incredibly specific business use case that is only going to work within one or maybe two companies or alternately, it came across as completely ridiculous. The example that springs to mind for that is when WeWork, back when that was a thing, talked about using their machine learning algorithms to determine that there was congestion in their lobbies at certain times, so they wound up bringing in a second barista during that time window. And it's, “Let me get this straight. You spent how much on data science to figure out that people like to drink coffee in the morning? Yeah, have you ever talked to a barista and figured out maybe there's a pattern here that a human could discern way faster?” It feels like it either lends itself to mockery, or to extreme niche use cases. And I know that's wrong, but that is the impression a lot of folks are left with. Is quantum in the same boat?

Sarah: So, I think actually, in some senses, we currently are, in that, kind of like I was saying, we don't really have a good way of generalizing what types of problems are suitable to speed-ups on our quantum devices. Because having a GPU, that doesn't mean it universally speeds up everything on my computer. If I'm I/O bound, or something like that, it doesn't help. Same with a quantum computer, I can have one—like, even if you literally gave me one today, descended from the heavens, it was perfect, error corrected. I honestly wouldn't know what to do with it [laugh] because we really haven't had a chance to try out larger applications.

But I don't—from my perspective, and maybe it's [laugh] just in that I've been doing research on this for a long time, but I think that doesn't necessarily preclude that there won't be. There are no no-go proofs that we won't be able to find other interesting things. And what I think, to me at, like, an almost romantic level, what is really beautiful about quantum computing is that it's an entirely different physical resource that we're actually using to do the computation. FPGAs, GPUs, they're all at some level the same sort of transistors on silicon [laugh] that at some level function the same-ish way. Maybe we put them together differently, but they're the same Legos.

Here, this is, now we got K’Nex. [laugh]. So, we haven't been thinking with K'Nex brains for a long time because we've been building with Legos, so it is kind of hard to actually find what maybe we can build with K'Nex. And that's what I'm really personally excited about exploring and using, honestly, the—we need to build up the hardware, of course, but that's where I actually see quantum software as being a really exciting kind of emerging discipline here, where we can actually start exploring K’Nex type solutions, [laugh] but all in software.

Corey: That, on some level almost seems to lend itself to another comparison, which is something you happen to be renowned for. Specifically, whenever we look at the industry and what they're doing with AI and machine learning, people are starting to find actual use cases for it. And that's exciting, and that's great. And that use case is invariably some form of bias laundering, where I put my biases in, the algorithm does this thing—as if that somehow absolves me of all responsibility—and then it spits my own biases back to me. But now I'm considered to be somewhat blameless. The ethics of quantum computing feel like they're still far away as the actual underlying technology gets built out. But you've been talking about it a fair bit. Tell me more.

Sarah: Yeah, machine learning is a really apt comparison, I think here because it is exactly a form of bias laundering. And as someone who's excited about technology, and always is excited about building it, I always try to keep in the forefront of my mind and in forefront of our discussions, there’s: can we do it? And then the other question is, should we do it? [laugh]. And so I think quantum computing is a technology that does have the possibility of drastically changing our society.

The computational power for at least the problems that we've seen speed-ups on is incredible. And so I have to really sit with myself and think about okay, this is great if I have it, or essentially, good people have it, but what could an adversary or what could a malicious agent do with this technology? And that's why I think it's really important to make sure, as we're building out this technology, this community, this field as a whole, that we really try to involve as many people as possible and get as many people as developers—literally full-stack quantum computing is kind of a thing now, [laugh] so we really need everybody at the table when we're making these decisions, so it doesn't just kind of turn into a bunch of white guys at a table making a choice for everyone.

Corey: Sorry, the other thing you just said about full-stack quantum computing is objectively terrifying on some level. It's one of those “Oh, great. You think boot camps are hard now because JavaScript doesn't make sense”—because it doesn't to me—“Great. Wait until we introduce the rest of this.” And it becomes almost this, I guess, surreal vision of a possible future. Do you believe that quantum computing is a technical inevitability?

Sarah: I do. I think we will, eventually, it's going to be a long pull. Like I know even when I started grad school over 10 years ago, they told me it was 20 years away at that point. I think they still say it's 20 years away, as you said. I have no good idea about hardware timelines, but what I think is here and present, and actually we in the year of our Lord 2021, have a chance to actually influence and change how we're developing the stack around the hardware.

Like I said before, if someone showed up and gave us a working hardware device right now, we wouldn't have the networking, the classical, kind of, dispatch. And that's basically where we're trying to build up: how does quantum computing integrate with the rest of the stack? And honestly, really, the best model for it is [laugh] as a cloud-computing resource. So it's not going to be a device that you have in your house or you put into your gaming PC build, but it'll be a thing that is offered—and is currently offered, actually, from a lot of the major cloud vendors like AWS, and Azure, and whatnot. So, I think trying to figure out what that looks like from a consumer standpoint is a really exciting and really cool place to actually make a difference.

Corey: Whenever I start looking into quantum computing and understanding the various approaches to it—I know AWS launched their Braket service last year, which was interesting in that it oh, it's finally contextualizing this through the lens of something that I spent a lot of time working with. And I pulled it up and, honestly, I don't know if you're familiar with a subreddit VXJunkies or not, but—and I'm telling a bit of a secret here, so I will deny this—fortunately it’s just you and me, and no one will ever listen to this—but the entire subreddit is built upon technobabble of explaining things back and forth that aren't actually real, and people making up technical words. And it's incredibly convincing; no one is entirely sure when they first discover this, whether it's real or not. And it was a remarkable parallel for looking at what these things were. The terminology behind quantum computing is unreal; the entire methodology by which these things get addressed, the concept of qubits, and different types of quantum computers that require different aspects, and some of them apparently are, I don't know, liquid-fueled or something like that.

At this point, it's one of these, is this just a giant attempt to have fun at my expense? Because, honestly, if so, yeah, one, good work. Why is it so radical a departure from the world that most of us are used to?

Sarah: You highlight something that is kind of uniquely challenging about quantum computing, and I think it really comes from the fact that it is a really interdisciplinary field. At a minimum—you know, if you were to sit down and hire a bunch of people to, in a closed room, build a quantum computer, you'd probably need a chemist, you’d need electrical engineers, you’d need mechanical engineers, you’d need physicists, you’d need computer theorists, you’d need mathematicians. And something that I really struggled with through grad school is, like, almost every textbook or resource you look at is a view of quantum computing from that field.

Corey: All you're missing at that point is a bartender for the punchline.

Sarah: Basically. [laugh]. But yeah, what you're seeing there is basically this amalgamation of jargon from four or five different distinct research areas and fields. And frankly, I feel like even the researchers in the fields—we name things ‘magic states,’ a lot of our analogies and papers are about King Arthur; it's like the Quantum Merlin Arthur problem. There's at least some fun had with the terminology, but also, yeah, it's kind of a mess. [laugh]. And so I've written a textbook, on teaching kind of quantum computing with Python and Q#, and one of the things I've tried really hard to do there is strip as much of that away as possible and use common language terms from programming to refer to what are effectively just names for particular types of matrices and stuff like that.

Corey: The challenge, too, at least to my mind, is, every time I step through this, it talks about things like running a quantum shop, for example—which again, does not detract from the idea of having a bartender involved somewhere—but even the idea of doing something like that is bizarre to me because my problem is, I cannot, in layperson's terms, come up with a reasonable explanation for what kind of problem would I have that this would solve? And I'm not even asking for a real business problem. That still feels like it's years away. I'm talking about things like, “Well, all right. You're going to learn to write code. So, all right, we're going to make the program spit out ‘Hello, world.’” “Cool, I can do that.” “Now we're going to make the thing count from one to 10.” “Awesome. Yay. I'm programming.” And honestly, you are at that point. Sure it's at an elementary level, but that is fundamentally what it's about. A lot of things I've looked at, even their basic ‘Hello World’ equivalents are tremendously confusing. Is that just something I'm missing?

Sarah: I don't think there's something exactly that you're missing there. What I think has been happening is a lot of the software and tools that we're developing for quantum computing right now is pretty heavily focused on bootstrapping up those initial quantum devices that folks are building right now. So, we have some number of qubits that you can access from IBM, or IonQ, or Honeywell, or wherever, and most of the software that's getting written is really geared towards that. Which, it's kind of like trying to think about writing programs on your computer in machine code because that's literally—you're thinking about gates. The programs are often called circuits.

Corey: Yeah, assembly is a good analogy here. It feels a lot like the assembly class I took half of once, and then immediately stopped attending because, “Wow, all right, my brain is full time for me to excuse myself.” You’re right, that feels very similar.

Sarah: Yeah, I'm not a firmware person. I don't really like thinking at that level. I'm much more comfortable in Python, where I can just say, “Please give me a variable.” And I don't have to think about pointers, or memory management, or anything.

I'm really excited about building and thinking about what the tools for quantum computing look like at that level of abstraction. And so kind of the closest we have—there are some different programming languages that are kind of being targeted more at the algorithmic level like that, like Q# and stuff like that, some of the Python tools that are out there. And that's where I think is a much better place for people to start with quantum computing because then, I feel like that's more commensurate with what you would see with a Python or Rust or something like that ‘Hello World’ program, as opposed to, here’s all of the assembly instructions to say ‘Hello World’ on the screen.

Corey: This episode is sponsored in part byChaosSearch. As basically everyone knows, trying to do log analytics at scale with an ELK stack is expensive, unstable, time-sucking, demeaning, and just basically all-around horrible. So why are you still doing it—or even thinking about it—when there’s ChaosSearch? ChaosSearch is a fully managed scalable log analysis service that lets you add new workloads in minutes, and easily retain weeks, months, or years of data. With ChaosSearch you store, connect, and analyze and you’re done. The data lives and stays within your S3 buckets, which means no managing servers, no data movement, and you can save up to 80 percent versus running an ELK stack the old fashioned way. It’s why companies like Equifax, HubSpot, Klarna, Alert Logic, and many more have all turned to ChaosSearch. So if you’re tired of your ELK stacks falling over before it suffers, or of having your log analytics data retention squeezed by the cost, then try ChaosSearch today and tell them I sent you. To learn more, visitchaossearch.io.

Corey: That's part of the challenge is that there needs to be at least some grounding at that level of technical competence. I would still argue as a result—and please feel free to contradict me on this one—that if people wind up coming at this without having some grasp of lower-level computing concepts and principles, they're likely to struggle. Is that a fair assessment?

Sarah: Yeah, just like in classical computing, we have developers at every level in the stack: we have people who are working on literal CPU instruction optimization sort of things, all the way to doing web dev, and database, and cloud stuff. Right now, most of our tools in quantum computing, and honestly, most of the educational focus is all at that CPU assembly sort of level. And there yes, it is more necessary to kind of know more about the hardware, to know more about what the qubits are actually doing because you're literally interfacing with it. I think it's a lot easier to kind of start actually at a higher level. And it's kind of like if you want to learn a new Python package or something like that.

I don't usually sit down and read through the API docs. I will start with their just very high-level example, and then as I need to understand things as I use it, I will go and dive in. I think you can take a similar approach to quantum computing, where you say, “All right. I want to actually start at this really high level; let's talk about algorithms, let's talk about built-in functions sorts of things, and then drill down and understand, kind of, once you have that broader picture of what's going on.” I think it's, kind of like, rather than starting zoomed in on a map, on Google Maps at Street View to understand where you are in a city, it's much easier to maybe start zoomed out a lot farther.

Corey: So, I opened this episode by joking about the tutorial being a PhD. That is clearly a bit above and beyond where we actually are in this day and age. But what are the realistic prerequisites? I've never been a fan of gatekeeping and I refuse to accept the answers, “You must have this degree from this university.” “Cool. Then you need to get out of my office,” because there's never just one path to anything.

And I understand that there are absolutely prerequisites that in many cases are hard to find without very specific academic achievements, credentials, and prerequisite, but I don't know that a PhD is one that I would even accept. So, what is the real-world limit of what you should know before diving into this space?

Sarah: First of all, I really do think anyone can actually be a quantum technologist or be a quantum software dev, is really the most critical skill for working on this stuff is basically linear algebra. So, if you can multiply matrices on a computer with whatever programming language, you can already start building quantum software, basically. I actually went into quantum computing—so I was in a regular physics track in undergrad, and then I saw these triple integral crap, and then I was like, “Oh God, this is really hard. I don't want to memorize any of this.”

And then I saw quantum was like, it's just matrices. And I knew how to make my computer—I knew how to make Mathematica multiply those, so I didn't have to do it by hand. So, I really think there's a lot of misconceptions about, exactly as you say, gatekeeping, or you must be this smart to participate. I regularly now in the open-source community work with a ton of folks that have no background in quantum, they were actually a web front-end dev, and they're helping to make contributions to these quantum open-source projects and tools. So, my personal belief is anyone can be a quantum developer, and I would hope people take me up on that and take a look at some of these higher-level approaches. That really—linear algebra. A little bit of statistics is nice, but honestly, that's where having software is helpful because it'll just do that for you. You don't have to think about the details of exactly what's going on there.

Corey: As someone who basically capped out at precalculus, to me, it sounds, oh, okay, this is not going to be accessible to me without a whole lot of study and planning. But the reason I bring that up is not for pity, or for you to, “Oh, no, it'll be fine,” then I go in, and it is very much not fine. But to point out that I believe this is like almost anything else in technology across the board which is today, it might be beyond my capability of easily getting into and assimilating, but the bar always gets lower, never higher. Things simplify over time.

It used to take three weeks to effectively get a web server up and running. Now, it requires basically a passing thought or a checkbox on a website. It gets easier with time. So, my question for you is, do you have a ballpark and very general ‘predict the future’ census of when this starts becoming more accessible to more people without the either math background or math focus? And I understand that's an incredibly loaded question.

Sarah: Yeah, I straight up generally refused to answer the question of, “When are we going to have a quantum computer?” But I think about now how easy it is for me to use PiTorch or something like that to do machine learning sort of things. I can in one or two lines with a folder full of pictures of my dog, [laugh] get it to train on my dog. That's the kind of accessibility that I really hope—kind of as you were describing—that we can get to with quantum computing.

And I really do think that in probably the next five to ten years, we can get the software there. Whether we have hardware necessarily to back it or not, that is sufficiently large for what people want to do, I honestly have never worried about [laugh] and, frankly, don't care. [laugh]. They're working on it; they're engineering problems. They'll get there when they get there.

It took how long to get transistors from the giant triangles of lead down to what's sitting here in my [laugh] PC next to me. But the software and kind of like that user experience, or what does it mean to actually use this technology is somewhere that we can make huge strides in the next five to ten years to have an experience kind of analogous to checking a box or whatnot to add whatever it is, quantum machine learning or whatever, to your projects or whatnot.

Corey: So, you send it you don't ever accept or answer the question of when are we going to have a quantum computer? And that's fair. But let me see if I can sort of do an end-run around that. What do we actually need in order to make quantum computers practical? And you can, of course, solve for ‘practical’ however you'd like.

Sarah: [laugh]. Sure. So, where we're at right now is basically we have a bunch of different kind of competing types of technology. Like you were even talking about the [laugh] liquid-run ones that possibly you were meaning the ones that you have in a bath of liquid helium or nitrogen. But we have superconducting qubits, we have ion trapped qubits, we have optical qubits.

There's lots of different options. And basically, there's five criteria that we need to have for it to be a good scalable type of device. Each of those technologies usually meets three, no problem. Then there's one that's an engineering stretch, but we mostly got in hand. And then there's one that’s, like, kind of an open question.

And so basically, where it seems like we've landed is superconducting is probably, at least at the moment—superconducting and ion trap technologies are kind of the leading candidates. But mostly, we just need time. The nice thing that these devices that they're currently pursuing can leverage, is all of our experience building all of the silicon manufacturing infrastructure. Obviously, we're pretty good at that; that's all of classical computing. And so we can leverage that for miniaturization, and really what they're kind of iterating on right now is reducing noise. So, quantum devices, in general, to stay quantum have to be isolated from the environment, and so it's just working on progressively better isolation.

Corey: Which sounds increasingly hard to do, given that we can't effectively handle isolation, even in a conceptual sense. When we look at things like oh, data security of, oops, did I accidentally turn the database backups into a web server with a wrong mouse click somewhere? If that sounds like getting stuff like that separated out, but still usable, is that as heavy a lift as it sounds like?

Sarah: Yep, pretty much. [laugh]. And honestly, that is kind of one of the most interesting questions to me. Having done some research on it, quantum machine learning is the thing. We've found algorithms that can help us speed up certain machine learning tasks, but the problem is, any advantages we find at an algorithmic level there are entirely blown away once we use—basically load the data [laugh] load the data—like, transferring classical data into the quantum computer for it to operate on it. So, those sorts of protocols that people usually just, let's assume we have that. We need to fill in the homework, and we need to fill in those answers before we can figure out some more applications.

Corey: At some level, it sounds unsatisfying, but the answer is, it's still a work in progress. I will say that it's interesting to see that even as early days as it is, you're still focused squarely on the ethics piece of it. Out of curiosity, is this something different than what we saw with the rise of things like machine learning, or were there folks early on in that process as well, talking about the ethics and thinking about the bigger philosophical questions? In other words, do we have a better chance now of avoiding some of the pitfalls that we keep smacking into as a society because we didn't pay enough attention the first time?

Sarah: I would like to hope, but I honestly don't think we are learning fast enough. I mean, it is early days, but what I've, even in the course of my career, seen it shift strongly from being an only academic pursuit to now a very largely industrial, most everyone I know now works for companies [laugh] working on this stuff—they're not postdocs, they're not professors—and that gives me some hope because weirdly, in general, I think companies are better at being ethical than academic institutions, just because they have lawyers. [laugh]. But I want to hope that we can do better. But right now I don't think we're on a better track, honestly.

Corey: Well, I have a serious problem ending an episode on that much of a downer, so let's ask one more question that expands on something a bit more hopeful. If folks have listened to this episode, and don't have the shrieking aversion, going back in time 20 years to struggling with math class in high school, or whatnot, and think I actually would like to get started with some of this, where would you recommend that they start?

Sarah: Self-promotion-y, come chat with me on office hours. So, I stream a lot on Twitch, both just kind of working on quantum open-source projects, and I also do office hours where people can come just ask me whatever questions you have about quantum computing, or just, kind of, tech stuff in general, or crazy stories about what we blew up in the lab. [laugh]. There's lots of good resources, like myself, on Twitter, and to just, kind of like, actually interact with the people who are currently building this stuff. Because I think, at a personal level, we're probably different than who you might expect is actually working on the technology.

As I mentioned earlier, I also have a textbook—or it's not a textbook, really. It's more of, like a, kind of, PowerShell in a Month of Lunches sort of format. [laugh]. But it's called Learning Quantum Computing with Python and Q#. It is basically geared for your average sort of dev; helps if you know Python, but you don't strictly have to know Python, just any sort of programming experience is good.

There are tons of good open-source sorts of resources out there, there's good awesome lists, stuff like that. The main thing I will caution is, guard yourself against the hype. [laugh]. Kind of like how we opened the episode with. There is a lot of hype out there that we've solved time travel with quantum healing crystals, whatever. Bring your best rational sort of logical skills when, kind of, exploring some of that stuff, and you will gain the most.

Corey: I think that is a much more uplifting vision for the future than a dark cloud over the future of humanity. But that seems to be the season for either one of those, now. People get to choose their own adventure on this one. If people want to learn more about what you're up to specifically, where can they find you? You've already mentioned your Twitch stream, but where else?

Sarah: Yep, I'm on Twitter a lot. [laugh]. My handle there is crazy--the number 4-pi314 I did pi memorization contests in high school, so kind of was my first internet handle, and that's pretty much what I am everywhere on GitHub, on Twitter. You can also find more about what I'm doing on my website, sckaiser.com.

Corey: And we will, of course, put links to that in the [00:32:43 show notes]. Thank you so much for taking the time to speak with me. I appreciate it.

Sarah: Yeah, this has been really fun. And I hope folks are interested to check out some quantum computing stuff, and come make fun of it on Twitter, too. [laugh].

Corey: Sounds good. Thank you once again for your time. Dr. Sarah Kaiser, technical staff member, quantum community lead at Unitary Fund. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you didn't like this podcast, please leave a five-star review on your podcast platform of choice along with an angry incoherent, misspelled comment telling me why all of this stuff is wrong, and you need to have a PhD in order to approach any of this.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey atscreaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Caroline CarterCaroline is our sponsorships manager at The Duckbill Group for our three media publications: Screaming in the Cloud, AWS Morning Brief, and Last Week in AWS. She also helped us create our first-ever re:Quinnvent digital conference in December of 2020. Before joining the Duckbill Group, Caroline sold market insights software to Fortune 500 companies at CB Insights and payment software to businesses at Square. Prior to her sales career, she worked in client operations at FutureAdvisor helping clients invest their money digitally. She lived in Paris for 3 years, which is where she caught the tech bug and did a coding boot camp.

Join Corey and Caroline as they discuss their mutual love of fintech, how learning to code in a foreign language can be tough, why people are reluctant to make changes in their careers, how to find better mentors, and more.

Links:

  • Caroline’s email: caroline@theduckbillgroup.com

Transcript
Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

This episode is brought to you in part by our friends at FireHydrant where they want to help you master the mayhem. What does that mean? Well, they’re an incident management platform founded by SREs who couldn’t find the tools they wanted, so they built one. Sounds easy enough. No one’s ever tried that before. Except they’re good at it. Their platform allows teams to create consistency for the entire incident response lifecycle so that your team can focus on fighting fires faster. From alert handoff to retrospectives and everything in between, things like, you know, tracking, communicating, reporting: all the stuff no one cares about. FireHydrant will automate processes for you, so you can focus on resolution. Visit firehydrant.io to get your team started today, and tell them I sent you because I love watching people wince in pain.

Corey: This episode is sponsored in part by LaunchDarkly. Take a look at what it takes to get your code into production. I’m going to just guess that it’s awful because it’s always awful. No one loves their deployment process. What if launching new features didn’t require you to do a full-on code and possibly infrastructure deploy? What if you could test on a small subset of users and then roll it back immediately if results aren’t what you expect? LaunchDarkly does exactly this. To learn more, visit launchdarkly.com and tell them Corey sent you, and watch for the wince.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by my colleague, Caroline Carter. Now, we've worked together now, but we've also worked together in the past and we've known each other and been friends for years. Caroline, welcome to the show.

Caroline: Hi, Corey, thanks. Great to be on with you today. This is different than my usual working with you.

Corey: Right. Normally, you're an account executive here at The Duckbill Group with an emphasis on media sales. So, whenever there are ads in this podcast, for example, which I'm sure we've already played at least one of by now, that's a company has chosen to sponsor my glorious love affair with the sound of my own voice, and they always go through you to do it. So first, thanks. You're helping put my kids through college someday. It's appreciated.

Caroline: You're welcome. It's always a fun experience to work with you.

Corey: I want to get into that in a little bit, but first I want to talk a little bit about career progression. Because your background is fascinating. You got a liberal arts degree, you went and lived abroad in France for a while, you went to a boot camp, if I'm not mistaken, in France which means that you learn to code in French.

Caroline: That is correct, yes. So, I was learning Ruby on Rails for a summer in French, which was definitely a challenge learning it in another language because obviously, Ruby is its own language in and of itself. But I did learn how to code.

Corey: That's on some level, it seems insane to me of, “Okay, I'm going to take a coding lesson in a language that is not my primary language.” But that's an incredibly, I guess, ethnocentric view of the entire world where, surprise, there are more people in this world who have to learn how to code in English or some form, and that is not their native language, then folks who didn’t if we look at this in terms of aggregate populations. So, on some level, if English is your first language, and that's the language you're learning to code with it, you're sort of playing on easy mode.

Caroline: Yeah, in some ways that you are, so I envy the people who are able to learn to code in English. But I think that's probably one of my strong suits, languages in general; I learned French, and Spanish, my mom speaks Polish, so I learned a little bit of that at home. So, in some ways, coding is just another language.

Corey: So, after you did the boot camp thing and decided that living in France, was for one reason or another, not what you were going to continue doing, you somehow decided to go to California and found yourself working at a fintech startup, which is where we met.

Caroline: That's correct, yes. So, I was in France for about two and a half years after college. I was teaching English and working at a couple of American companies part-time and decided that tech would be this new, amazing field I could join, but I didn't have any technical skills. So that's why I wound up doing the boot camp, and then decided all these companies were out in San Francisco, and so that's really where I should go to start my journey in tech.

Corey: So, what was that like? When I talked to most folks who do the boot camp thing and then go work at a company in San Francisco, the role is generally focused around writing code full-time. Your role wasn’t, so I have to put myself in a different position, in some respects, and try and imagine the insanity of startup culture through the lens of someone whose role does not involve writing code as a part of that culture. Was that as bizarre as it sounds like it was?

Caroline: It was definitely different. But after the coding boot camp, I figured out I really liked working on the client-side, talking to customers every day, so I actually started out in an operations role initially, but I was able to bring in some of the coding background by just spinning up email templates so that our team could automate certain things faster. So, it was definitely interesting joining tech, but I'd say because I was on the client-facing side, I actually wasn't as exposed to engineering teams, as most folks are.

Corey: And what did your role involve?

Caroline: So, I think our first interaction was at that startup I was working at. It was a robo-advisor helping people manage their assets digitally. And this was before people knew about services, kind of—now, I think they're more mainstream, like a Wealthfront, Betterment, Robinhood, so this was one of those earlier iterations of a robo-advisor.

Corey: It was a lot of fun. There were a lot of good stories that came out of that place, and it was a blast to meet some of the people I got to work with there. And you were one of them. It was very interesting because we started talking, and that turned into, “Let's go get a cup of coffee.” And that turned into, “Let's get coffee,” a few times a week.

And over time, we just started having longer, deeper conversations about careers, about job satisfaction, and the rest. And over time, you became, I guess, annoyed with the lack of advancement opportunity at the company. And what I found remarkable was, you complained about it, like we all do. But you didn't stop there; you did something about it.

Caroline: Yes, well, I think it's very easy for people to get complacent. And I see people who complain about their jobs, go in from nine to six every day, stay in the same role, maybe in the same company for three or four years, and you just have to hear this person complain. So, I never wanted to be one of those people. I wanted to always find a way to move on to an opportunity with more responsibility, something that was going to be more challenging for me. So, I know that that's always been a driver when I think about making career moves.

Corey: There's an entire society of people who don't like their jobs. It’s called everyone, and they meet at the bar to quote some comedian here or there. The difference that I found was that there's a certain, I guess, subset of people who will, at some point, decide enough's enough and start looking for something else to do. And as I recall—please correct me if I'm wrong on this—my thought at the time was, figure out what growth looks like here and what success is, set metrics around it, make it clear that you're looking for those opportunities, and time-bound it. So don't let companies lead you on for extended periods of time.

And at some point, like, all right, things have been promised or assured me that are coming, but they have not happened in the timeline I gave them. Let's see what else is out there. Which, frankly, I advise anyone to do, regardless. And again, I know I run a company, I still maintain that everyone—including people who work here, including you—should always keep your eyes open for other options. Just because, you should be where you are because you like it not because you believe it's the best you can do, or there's not some other place where you might fit in better. Validate that assumption, continue to talk to places.

Caroline: Definitely. And I think career stuff is not unlike relationships or other areas of life where you have to look at the words versus actions. So, a company maybe tells you, “Oh, we'll consider you the next go-around for a promotion.” And then it gets delayed, or, “Oh, we’ll think about giving you that raise.” And then it doesn't happen, or the title bump.

So, I think at a certain point, companies do the same thing where they promise these things. And then, like you said, there's a time limit, and you have to decide, okay, am I willing to wait that out for a year or two years? What does that horizon look like for me personally? And something you encouraged me to do that I find to be very helpful, and I tell everyone I know this as far as career advice goes, just take interviews. Take some calls once in a while because honestly, the best time to make a career move can be when you're doing well somewhere because then you can really evaluate opportunities more objectively than if you're in a bad place or not performing well at work. And I think that's always been helpful to me, too, just to get a better sense of what opportunities truly exist out there.

Corey: Let's also not lose sight of the fact that when we're having the job interview conversation and evaluating people to hire, we're fundamentally judging people on the basis of how well they perform in a job interview, which is its own separate set of skills in a lot of respects. And a lot of those skills are only ever relevant in the context of a job interview. So, one of the more hilarious series of interviews you'll find is when you find someone who's been working somewhere at the same place for 15 years and hasn't been interviewing, and now they're starting again, their reflexes are all wrong; the muscle memory for how to feel the question just isn't there. And it takes time to wind up getting back into the swing of things. So, why not do that when you're happy and enjoying what you do, just to see what else is out there and keep those skills sharp?

Caroline: I agree. And I think a lot of people when they realize maybe this opportunity at their current company isn't working out, or they want to go somewhere else, they do it from a place of desperation, or, “Oh, I got to hurry and find something new.” So, when you're under that pressure, you can't make a good logical decision for your next move. Whereas if you're happy in your role, or you're feeling so-so about it, but you look at some other companies, talk to folks maybe in other roles that you're interested in moving towards, I think that can inform you better. And then you're making the decision from a place of confidence and certainty, after doing some research.

Corey: I will also point out, just as a footnote for history, that I was told one of the most abhorrent things I've ever heard in corporate America back when we were having those almost daily coffee chats. Specifically, that the two of us going out and getting coffee once a day had questionable optics, and as a manager at this company, I should definitely make it a point to not spend a lot of time one-on-one with a woman. I really regret in hindsight, not telling that person to go to hell just on the spot. I mean, I'm sorry, I'm not Mike Pence. Stories like this are how folks who don't look like me wind up getting shut out of an awful lot of opportunities that otherwise come about through building those relationships in the workplace.

And yeah, spoiler, that doesn't always happen on company property. And it doesn't always look like people sitting down in the formal confines of a one-on-one in the conference room. I just found that to be one of the most disturbing things I was ever told. And every year that goes by, I find it even more questionable.

Caroline: Yes, I mean, I definitely think a lot of times we talk about mentoring people in a professional setting and how certain groups—women, in particular—that they don't always have role models, or people to look up to, or mentors, and from my perspective, in a way, many of the best mentors, I've had yourself included, it's come about, sort of organically, like you mentioned. We just went to coffee, talked a little bit about tech, traveling, things like that. So, I think it should be more encouraged, and there's so many things I learned from some of the mentors that I might not otherwise have that have allowed me to get to the place I am in my career, and I hope to have those people for years in the future.

Corey: Well, first, thank you, it's very kind of you to say that. Something I've learned is that mentorship is dependent much more upon the protegé than it is the mentor. It has to be driven by the person who is basically trying to improve their own situation. Otherwise, I'm just sitting out here shouting advice, and no one wants to hear it. That's not particularly compelling. That's called being a white dude on tech and Twitter.

Caroline: Definitely. And I remember at that time, I was pretty junior in my career. It was the early days. And so you're right, I think we talk about mentors, but they can really only do so much. It's up to the mentee to go out of their way to ask questions, ask if the mentor will be willing to look over their resume, their work, how they can improve.

And I think the other thing is usually a mentor is someone who, maybe not always age-wise but experience-wise, has a bit more experience under their belt, and so if you can talk about where you're looking to go, or you have some examples of people you look up to, the onus really is on you to ask questions and see how they can help or strategize a plan for what your future will look like. And I think a lot of people think about the next role, but not the longer-term where, “Hey, where do I want to be in 10 or 20 years career-wise?”

Corey: Yeah, forget the next job. Tell me about the job after that and how we help you get there. And I did get some feedback on our spending as much time talking as we did that was less horrifying, and more to do with the form, “Look, Corey, you're in the DevOps group, and she's basically in the customer service org”—which is an understatement of what you did, but okay, fine, it's directionally correct so we'll roll with it—“Why are you spending all this time working with someone who's not going to be in a position to help your team?” Well, turns out, I occasionally take the longer view on things. So let's move on. You left that company, what happened next?

Caroline: After that company, I actually decided—so I was working in a client-facing role; I was onboarding new customers, and I decided, to your point because I really was driven by metrics and always going above and beyond just getting through a to-do list, I decided sales might be a good path forward. So, I actually went on to work at Square. I worked in their New York office’, they were just building out a sales team there. This was back in 2016. And so then I was there for about three years working on their SMB and mid-market sales, selling software. Which I realize sounds crazy because most people when they think of Square, they think of the point-of-sale system, which it is, but they have a lot of software that businesses can use, too.

Corey: I was about to say that. [laugh]. They've really moved up-market in a large degree. They're now directly competitive with Stripe in some respects, for example. And, yeah, they were there for three years, which is obviously not a short term role for basically anyone. I've never stayed at a company other than this one for that long. And so you went there; it was a better situation. You were doing sales, for the first time in your career. Was that a hard transition?

Caroline: It was in the sense of, okay, I have to call a lot of people, run the full sales cycle, prospect to people, but it's really not that different than if you do work in some client-facing role, it's just that there's some more metrics around it. And I have a lot of friends and people, I know who, when they hear sales, they're like, “Oh, I'm just too risk-averse for that. I could never do that. That's terrifying.” But I think if you're friendly, and you like talking to people, it's actually a really great career path for a lot of folks.

Corey: I would also take it a step further and argue that absolutely everyone is doing sales; not everyone knows it. Whether you're selling a product, or a service, or an idea to your colleagues and suddenly the light goes on for some people, you're trying to persuade people in one direction or another. And I think that there's a strong negative reaction to the idea of sales that is almost entirely based upon people's experience with terrible salespeople and terrible sales processes.

Caroline: For sure, and I have actually the same feeling when I see someone that says sales, I immediately cringe because I have had those interactions with bad salespeople in the past, too. But I think now we're moving away from this era of the really pushy car salesman, and with a lot of these technology salespeople, they have to be more customer-oriented, finding custom solutions for folks. And so I think we're hopefully coming into an era of more, I guess, humanizing salespeople.

This episode is sponsored in part by our friends at New Relic. If you’re like most environments, you probably have an incredibly complicated architecture, which means that monitoring it is going to take a dozen different tools. And then we get into the advanced stuff. We all have been there and know that pain, or will learn it shortly, and New Relic wants to change that. They’ve designed everything you need in one platform with pricing that’s simple and straightforward, and that means no more counting hosts. You also can get one user and a hundred gigabytes a month, totally free. To learn more, visit newrelic.com. Observability made simple.

Corey: The idea that you're immune to sales pressure is a common misconception. We see it with advertising in the same boat. And I feel like one of the painful parts about that is no one wants to be persuaded to do something. They want to believe the rational creatures who would absolutely do the right thing. And I get it.

Counterpoint: for anything that is more complex than punch in a form and give a credit card. Very often you need to have conversations with folks, especially at enterprise scale. It turns out that you don't get to go to any website in the world for the big management consultancies and purchase, “One consulting, please,” and put in your credit card number, and then they'll send a team out. These are complex processes; there's a nuanced discussion.

And I think that anyone who winds up saying that, “Oh, sales is useless,” or, “Sales is full of bottom feeders,” or the rest doesn't truly understand what sales is. I think that when you run a company, as I've learned, you've got to be in sales mode to some extent, all the time. But that doesn't mean what most people think it is when you say that.

Caroline: I agree. And I think a lot of sales really is—it boils down to relationship-building. And I was in an enterprise sales role after Square, and in that role, it was very much talking to a lot of teams at Fortune 500 companies, and sometimes even if you were speaking to a team who wasn't going to be the ultimate buyer if you had an in with them, they liked you, they would be a great person to introduce you to the person who was actually going to buy the software. So, I think you are always doing sales in some regard, and if you can build those relationships, it makes the whole sales process a lot easier.

Corey: So, what happened next? You were at Square for three years. And then you left and went somewhere else that was not here. What happened? Where do you go? And why did you leave?

Caroline: [laugh]. You know, three years is a decent amount of time to be at a company, and I decided Square at the time, it was really focused on small to medium businesses, or kind of, the mid-market segment. They have since moved up-market, and they do have a small enterprise team, but I really wanted to get into enterprise sales. I always enjoyed fintech, so I wanted to stay within that area, and had an opportunity to go work for a company called CB Insights where they do market intelligence. So they're looking at all these different data points that allow these huge companies, like a Microsoft or Amazon, to make corporate strategy or corporate innovation decisions based on some of that data.

Corey: That's a much more complicated sale. At that point, no one is going to sign up on a website for something like that, I would imagine. It becomes a much more nuanced story of what value it provides to a company, who when the company is going to care about those things, and it becomes an enterprise sales conversation.

Caroline: Exactly. So, that was within the enterprise sales framework. And it was very interesting because I got to work in a lot of different segments, whether it was digital health, autonomous vehicles, fintech, there was a lot happening there. And so, what a lot of people don't realize is, a corporate innovation team at say, LGs, they are trying to think five or ten years down the road. And so they're trying to analyze technology, but because they themselves are not technical they don't always know what's possible, or where they should be investing their resources, whether they should build technology in-house or outsource it, or potentially considering acquiring a company. So, it was a very interesting sales product as well as sales role.

Corey: Something I've learned is that as you start talking to bigger and bigger organizations, the organizational distance between the person who is paying for the thing you're selling them, and the person who's benefiting from the thing that you're selling them, and the person who procures the thing that you're selling them, and the strategic drive behind the benefit from the thing that you're selling them is all separated out; the organizational distance between those functions dramatically increases, and that makes a much longer drawn-out sales process that requires an awful lot of, I guess, project management style approaches to those sales conversations. And no one is on the other side of that issue in those conversations; it's just hard as you negotiate through a large company process.

Caroline: Exactly. That's a really good way of describing the enterprise sales cycle. And what I found is, yes, the deals take a lot longer, just because you are trying to coordinate between the decision-maker themselves, then you're doing demos with people who are actually going to be using the product, but they obviously aren't the ones with, necessarily, the budget to purchase the software. And then there's the procurement or legal teams you're going to have to go through, so it is an involved process. And they often say that enterprise salespeople are quarterbacks where they're managing all these different relationships, keeping everybody on a timeline, making sure the right stakeholders are involved.

And I think that's part of what makes it so interesting, but I also think if you want to do that, you have to have an internal drive because no one is going to check in on you and say, “Hey, are you getting the deal done?” It’s sort of on you to go to all those different teams and make sure people are staying on track and heading in the direction you want them to, which is usually a signed contract or a closed deal.

Corey: So, how long were you at CB Insights?

Caroline: I was at CB Insights for about a year and a half.

Corey: And then as all things do, it came to an end, and what happened next?

Caroline: So, I decided to actually make a move because I think, you think about that if there's a better opportunity presented. And so I was chatting with you and Mike over at the Duckbill Group about building out a new segment of your business, and I wanted to go somewhere where I could really take ownership of a sales cycle and potentially build out a sales team. So, that was the motivation for coming over to The Duckbill Group and joining you guys.

Corey: And may I just say, that that solves so many problems at once. First, there are never enough hours in the day on my side of things to focus on everything. And the weird thing about going from doing it all yourself to hiring a team is that every person you hire, without exception, is better at the thing you're hiring them to do than you are. It's an incredibly humbling experience.

Caroline: Yes. And the interesting thing is because I've worked at companies that are different sizes, everything from the small startup, to a company that was acquired, to a larger public company, and then now back to a smaller company, I think it depends on what skill sets you have, or which ones you want to cultivate. And I think anytime you're working in a smaller company, you're obviously going to be wearing multiple hats, and a lot of the responsibility falls upon you because you don't have maybe a sales engineer who can do the demos for you or a sales development representative who can do the prospecting, so you have to be comfortable taking on a lot of different roles at once.

Corey: It goes beyond that, too, at least for me, one of the problems I had is that when people started asking if they could sponsor my nonsense, my response was, “Of course you can give me money to mention your name. How much money?” And in time as these just became a larger and larger, I guess, phenomenon, I started to feel really weird about it because, “Ah, my voice talking about you is so valuable that you should pay me an insulting pile of money for it.”

It's weird. I still don't understand why anyone cares what I have to say, but they do. When you find you and the market disagree, assume you're wrong. But it was, seriously, a difficult thing for me. The fact that I don't have to deal with that at all, when people say, “Oh, what does it cost to sponsor your stuff?” I can give the honest answer, “I don't know, talk to Caroline,” and the problem goes away. It is incredible just as a stress relief from that perspective. Which is not a common problem, but it's definitely one that resonated and kept me up at night.

Caroline: Yes, and I think a big advantage you have with that is it allows you to stay independent in the content that you produce. So, you can go about your normal Corey ways, talk about companies in a snarky format, say your honest feelings about them without even knowing who a sponsor will be that week, so there's that nice delineation between the content side of things and the creativity that you bring to the table, versus businesses who want to basically be a part of your audience and get in front of those folks.

Corey: Oh, for those who aren't aware, I know there are ads in this episode, as there are in all of them, but I record the ads in a batch, and I don't ever know who is sponsoring a given episode, which means I don't have to shape what I'm talking about to not inadvertently offend a sponsor. And it's never been a problem, but it's an editorial firewall that I find incredibly stress-relieving as well. There's also—and this is something I want to get into as well with you—there have been times where someone has reached out wanting to sponsor, and we as a company will sit there, think about it and the answer becomes no, for a variety of reasons. Either it's something that we don't generally believe is the right answer for a customer to use, possibly there are ethics violations. There was one security product a while back that I looked at and found absolutely horrifying and didn't want to be associated with. That's got to be challenging as a salesperson to be told, “Yeah, you went to all the trouble of finding and getting someone ready to buy, and now we're not going to be able to proceed with them.” But you haven't rage-quit yet, so apparently, it's manageable. But we'll talk about that. What's that like?

Caroline: You know, every company—and I think we're seeing this, obviously, in the news, when you're a private company, you have the ability to make decisions as to who you want to work with or who you don't want to work with, and something I've always admired about The Duckbill Group specifically is that you guys stay true to your word with that. So that's what makes going to work really rewarding is being able to work with customers that we like and we think are a good fit, and then if they're not, telling that honestly and keeping that credibility.

Corey: My argument has always been that if I wind up turning down money in favor of authenticity and building relationships, either with people individually or with the audience, in time, that relationship can turn into money, but I can't sell that out and then wind up biasing for money, and then turn that money into relationships. I think it's short-sighted to go down that path. I'm just disappointed by how, I guess, infrequently it seems that this mindset takes root in companies.

Caroline: I agree. And maybe because we're a smaller company, it's easier to do that. I think when you have a huge market share and a lot more shareholders, it becomes harder, maybe, to turn down certain business opportunities. But for now, I think The Duckbill Group, we do everything with a sense of integrity, and yeah, it's what makes it worthwhile for sure because I think I would maybe rage-quit if I had to work with companies I didn't believe in or stand for.

Corey: At some point, you have to wind up being true to your own values. Or alternately, if you're just chasing money above all else—which again, that is a choice people make—why are you working in tech as opposed to investment banking? Or going down and making land mines, or whatever it is that is the number one amoral dollar you could possibly make for your time? Good for folks who are into that. I don't ever find that philosophy to be compelling. But again, I'm not trying to judge unduly. Everyone has their own situations and I'm not going to blame people for chasing money.

Caroline: And you talk to a lot of people who maybe aren't happy in their careers or their jobs that they're in right now, and a lot of times it's because they don't always think about, hey, what gives me a sense of purpose when I wake up every day? What is the specific industry I like? And within tech, I think what's so exciting is you can work in digital health; you can work in financial tech stuff; you can work in so many of these different areas, so for me, I always wanted to sell something I was actually interested in because I knew that I’d then be able to do a lot of research, stay up to date, like, that would motivate me and excite me, as opposed to just taking any random job for sales. And I think a lot of people would be better served if they've thought about, “Hey, what do I get excited about?” As opposed to the money because I think they always come together when you can find an area of interest, and then a practical application of it.

Corey: Caroline, thank you so much for taking time away from actually chasing deals down to come on the show and have a conversation with me. If people want to reach out to you to either talk about career trajectory, or buy sponsorships on various properties of The Duckbill Group, or attempt to hire you away and thus lead me into a very expensive bidding war to retain you, where can they find you?

Caroline: Well, Corey, thank you so much for having me on the podcast. It's been a ton of fun. I'm not used to actually being on shows, so thanks for again, having me and hosting. If you'd like to get in touch, feel free to reach out to me. My email is caroline@theduckbillgroup.com and we'd love to chat with you.

Corey: Caroline Carter, account executive at The Duckbill Group. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you've hated this podcast, please leave a five-star review on your podcast platform of choice, along with an incoherent comment explaining why I'm completely wrong to wind up spending time one-on-one with a colleague who happens to be a woman.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Regis WilsonRegis Wilson is the founding engineer at Release, an environments as a service provider. Regis brings more than 25 years of tech experience to this position, having worked as an infrastructure architect and SRE at TrueCar, Inc. and a cloud systems architect at Live Nation, among several other positions.

Links Referenced:

  • Connect with Regis: https://www.linkedin.com/in/regis-wilson-a713609/
  • Personal Website: http://www.zennet.com/
  • Release: https://releaseapp.io/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

This episode is sponsored by our friends at New Relic. If you’re like most environments, you probably have an incredibly complicated architecture, which means that monitoring it is going to take a dozen different tools. And then we get into the advanced stuff. We all have been there and know that pain, or will learn it shortly, and New Relic wants to change that. They’ve designed everything you need in one platform with pricing that’s simple and straightforward, and that means no more counting hosts. You also can get one user and a hundred gigabytes a month, totally free. To learn more, visit newrelic.com. Observability made simple.

Corey: This episode is sponsored in part by LaunchDarkly. Take a look at what it takes to get your code into production. I’m going to just guess that it’s awful because it’s always awful. No one loves their deployment process. What if wanting new features didn’t require you to do a full-on code and possibly infrastructure deploy? What if you could test on a small subset of users and then roll it back immediately if results aren’t what you expect? LaunchDarkly does exactly this. To learn more, visit launchdarkly.com and tell them Corey sent you, and watch for the wince.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. A recurring theme throughout a number of these episodes over the years has been me complaining about various jackwagon bosses that I've had over the course of my career. Today, as a guest I have one of those jackwagon bosses here to suffer my slings and arrows directly. Regis Wilson, welcome to the show.

Regis: Hey, Corey, how's it going? Glad to talk to you.

Corey: It has been an interesting year. I don't think anyone planned it to go quite this way. But yeah, we've worked together repeatedly over the course of well, many years, just because first, for a long time I lived in Los Angeles where you are and LA tech is not as large as San Francisco tech, where at least here I have the ability to not want to work with people, and I never see them again once we change companies. You keep running into the same faces, the same names in the LA market. And honestly, I kind of miss aspects of that.

Regis: Yeah. I remember the first time that I interviewed you. You asked me a question, like a interview question, and I answered it. And then you were like, “Oh, okay.” It was like, “Why was someone that I was interviewing asking me an interview question.” It's pretty funny.

Corey: Yeah, I always have maintained that people who are going down the path of interviewing for a role forget that it's a two way street at their own peril. And that was sort of weird back then when I didn't have much of a history at all, or track record, of being able to deliver on the technical, but I had a certain way of asking weird questions. Now I just do it on podcasts, and I get my kicks that way. But it was a lot of fun working with you. Again, I joke at the intro here, but you were a great person to work with.

It was fun; I never felt like you were asking me to do something that you wouldn’t have done yourself. None of it was unreasonable. And it was just a awful lot of fun. But for me, the real validation was a couple years later, after the last time we worked together, I happened to wander into a different company, and the next interviewer was you. And you, “Oh, Corey. Great to see you.” And you turned to the previous interviewer and said, “Yeah, I don't need to speak to him. He's great. Hire him.”

Regis: [laugh]. Yeah.

Corey: And it was awesome. That was the first time getting some aspect of my own reputation. And I could have assumed that oh, as soon as I leave, you said, “Just kidding. For God's sake, no.” But they extended an offer. So, if you did do the whole, “For God's sake, no.” You were nowhere near persuasive enough.

Regis: Yeah. No, I take that seriously. When I meet people that I like to work with, that I enjoy working with, or working for, or you're working for me, I always want to be able to work with you again. Or not you, but the person that I'm talking to? So, yeah, that was great. It’s a great history. And I think it goes back to, like, 2008, or—well no, 2006. [00:04:43 crosstalk]—

Corey: It’s somewhere in that range, yeah.

Regis: It sounds recent, but it's 14 years ago now.

Corey: Time is speeding up and we're all getting older.

Regis: Yep.

Corey: It's fun, though. We agreed that we could tell stories, but never name names of companies to avoid lawsuits and whatnot. But it was fascinating seeing some of the anti-patterns that emerged in running engineering companies. To tie this back to the idea of Cloud, that was one of my first AWS environments and I want to say this was, ehh, I don't know, circa 2010, give or take a couple years in either direction. Well, put some fudge factor in there to avoid lawsuits as well.

And back in those days, network latency was crap. They didn't have VPCs, or anything like it, or not widely deployed anyway. And I think it was the VP or someone like that had this love affair with this ancient monitoring system.

Regis: Nagios, yeah.

Corey: I think it was Big Brother or something like that.

Regis: Big Brother. Yeah. Oh, yeah.

Corey: One of those. And because it was so flaky from network perspective, but AWS was clearly the future, instead of having one of these things, we had three of them. And whenever someone was on-call, you could just ignore the pages unless you got three of them because then if all of those listening posts found the problem, then you had to swing into action to fix it. And I think it was between nine o'clock at night and six in the morning, there was no alerting whatsoever in the entire estate, just because otherwise no one would get a wink asleep when they were on-call. And looking at this, it was, okay, this is certainly a way to wind up architecting these things. And the lesson I took from that was, I don't ever want to be in an on-call rotation where the person who dictates how on-call stuff works isn't themselves on-call because it was terrible.

Regis: Yeah, exactly. The VP that you're talking about, he designed a monitoring system, but he's not really an engineer capable of being able to design this monitoring solution. And it was terrible. And I mean, the good news is that I was able to deploy this thing into three regions in AWS, back in the day, with EC2-classic instances, just like you said. No VPCs, everything's public. But it was a great experience for that Cloud VM. So.

Corey: Oh, yeah. It was my first outing with AWS in any realistic sense. And it was okay, this is kind of awesome. But I was also short-sighted back then, and it was, “Wow, that seems super crappy. So, obviously, I'll never run anything here that's even slightly concerned with latency or jitters because it's clearly not a fit for that.” It's kind of funny of the limitations of a platform like that tend to erode over time.

Regis: Absolutely. And if you fast forward to, like, 2016, and ’17, when they started releasing all these great features in serverless, and the list goes on and on every year; it's more and more, it really shows the difference between then and now and how much it's taken over.

Corey: It really does. The world has evolved in a bunch of different ways. Well, at least I assume it has because part of it is also, I guess, my own tolerance for things. That place that we worked at briefly the first time was so odd to me across so many different axes.

For example, they were doing this really weird implementation of Scrum meets Agile meets someone's half attempt at an MBA. And throughout the course of the week, there was never more than 45 minutes that any engineer would be able to sit down and work on something for an extended period of time. Which doesn't lead to the best outcomes.

Regis: Right. Yeah, we were planning for meetings, that would be meetings later, and then have meetings to plan for the meetings that we were going to plan things in. It was really bad. Really unfortunate.

Corey: So, time marches on. We're now in the year of our lord 2021. And you're still in Los Angeles. I'm not. But you are the founding engineer at releaseapp.io. Am I pronouncing that correctly?

Regis: Yeah, that's correct.

Corey: Excellent. So, what do you folks do?

Regis: Yeah, so it's really exciting. We've taken this idea of—basically, for my whole career, I've been building, like we've discussed, nothing but application deployment environments. Usually, it's a production environment and a pre-prod environment. In this case, in this founding company we're saying, “What would happen if that was free?” Because you could just spin up environments today, at a whim.

What if you could just do that so much and so fast that it was essentially unlimited? So, you had unlimited number of environments. And that's what we're doing.

Corey: So, tell me a little bit more about that. Because, again, the last time we really worked together in any depth, it was, “Oh, you want to provision something new?” Even in the world of AWS, that was, you're going to want to go and pack a lunch for this; it's going to be a little hairy; it's going to take some time. What's changed since we were sitting there beating on the custom bespoke Nagios box by hand?

Regis: Oh, yeah. Well—

Corey: Kind of a lot. But that's a loaded question.

Regis: Yeah. No, absolutely. So, if you remember, we were racking and stacking physical hardware. And I always was ahead of my time, I think. We always tried to use virtualization on that hardware, but even if he didn't, you would use the hardware, it would take weeks to rack.

Let's say you had a box that was already racked, you needed to provision it. If it was a VM it was a little bit easier, but it was still—as you said, it would take hours— or maybe even days sometimes— to get a system running. Nowadays, you can spin up an EC2 instance—even better, you can spin up a container on Docker, in ECS or EKS, and you can have a running OS with ready to deploy applications in seconds. Minutes. And so what we do now is even built on top of that you have Kubernetes, let's say, EKS clusters, which is what we use under the covers.

You can define your application topology and then you can just deploy it with the template in Kubernetes. We can spin these up in minutes. And so if it used to take you months to build the QA environment and keep it up to date, originally would take months, now it takes hours. We can spin you up these environments in minutes.

Corey: It seems like there's a lot of companies that are tilting at the windmill of attempting to solve for development workflow. And every time that I've tried to come in as the first DevOps hire, a quote, “Non-developer” with a wink and a nudge because I write an awful lot of YAML for someone who's not a developer.

Regis: Right.

Corey: And they wind up saying, “Great, now go ahead and fix our deployment and development process.” Oh, my God, do you run into the hell that is other people's workflows. Some folks are dyed-in-the-wool fans of local development and you need to make their laptop look exactly like production because they were on a plane once that didn't have wi-fi and ever since then, for the last six years, they insist on being able to do everything locally. And then you have other folks who are basically view, “Oh, my feature branch is called master in Git.” And it's everyone's different approach of, “Ooh, this workflow doesn't work for me at all because my special custom IDE doesn't work well in this type of environment because it's super prescriptive,” and what is that IDE you might ask? I couldn't tell you because I've never seen it before or since. And this was all at the same company. So, it becomes a bit of a problem as far as getting that hell organized into one particular direction. What's your take on it?

Regis: Yeah, I totally agree. And we find that, too. Obviously, like, in my history, working in companies just like you’ve described, what we've taken is sort of the approach that these are all environments. So, excluding the laptop—which we could actually do. You could run a Kubernetes cluster on your laptop.

I don't recommend it—but excluding the laptop environment, all of the other environments, typically, they're hard to maintain, they're hard to set up, they're hard to configure, they're hard to make the same. And as you move through QAs, typically you move the QA to staging, to production. Those environments could be different, even in the same company, even with different platforms and trying to make everything the same. What we've done is we've said they're all the same environment.

And it doesn't matter if it's ephemeral, like you use an environment for a PR branch, just like you've said, “Hey, I want to spin up this feature and see what it looks like.” Or if you have a QA place with people running tests against it, or if you have a staging environment where you need to test before you go to production. You have sales demo environments that need to be isolated and pristine so a salesperson can run a demo. We treat them all the same. They're all identical. And, excluding laptop, as I said, you check your code in, if you've got that application topology defined, we spin them up within minutes, and you can start using them. So, it really helps, and it makes it all uniform.

Corey: I mean, yeah, the idea of having a Kubernetes cluster on your laptop, sure, why not? I've heard dumber things—not lately— and everyone seems to want to do it. Good for them. Back in the day, when I could travel, my laptop that I took with me was increasingly just an iPad, which meant everything I was doing was remote development. And as someone who lived on airlines a couple years ago, back when that was a thing, I never found the wi-fi so bad, I couldn't get work done or, worst case, find something else to do for two hours.

So, optimizing for that use case is something I was fairly dismissive of. But it forced me to look at everything through a lens of at any moment, you could steal my local terminal, which is really what I was treating this thing as. But my data and the stuff that I care about should always remain safe in the Cloud. That led to a number of excellent habits and a few really strange ones for me. Do you find that what you're doing is taking a driving stance towards improving remote development experiences? Do you think that you're effectively taking an opinionated stance against local dev? Or is it more nuanced than that?

Regis: I think it's much more nuanced than that. I think what you're saying is that people would use our environments like developers would use our environments. What we're more headed toward is adding value to the business by displaying the code in a running condition that theoretically or even—possibly, depending on a workflow— that you could push a button and say, “This is deployed to production now.” So, that's more where we're headed. We generally don't care if you use an iPad or laptop, Macintosh, or Windows.

Corey: Plan 9, it is.

Regis: Right, right. Plan 9 it is. And exactly, the point there is that once that code hits the repository, that's sort of one we take over. We take that repository, we build it up, we show it to you, you share it, you use it, you test it, you burn it to the ground, you do it again. And that's where we make really easy, really fun.

And then when you're done, we have, even, customers now that we're running and maintaining their production environment. So, once they're done testing, and they've looked at the feature, they can click that button to merge and it goes into production.

Corey: The problem that I've had whenever I've started out with the best of intentions down this path—for example, my ridiculous newsletter has a overly engineered serverless workflow that builds the whole thing—and there are two mistakes I make. One, I tend to hard code things endpoints, which make it super challenging to wind up having this portable between environments, and of course because people ask what my school of thought is, when it comes to development, my answer is, “I'm a pragmatist.” And that means that one of my production tables has the word ‘test’ at the end of it. Another one has ‘dev’ at it because once it's working, well surprise, its production now.

It seems like in order to get that done correctly, there's a significant level of unwinding that needs to get done of refactoring out hard-coded endpoints and credentials, in some cases. How do you approach that? Or is it something that, I guess, enforces those good behaviors upfront, which is really what I need. Basically, if there's a SaaS that offers adult supervision for developers, I'm extremely interested because Lord knows I need it.

Regis: Right. No, you're absolutely correct. In my experience, the development environment, or the POC is actually what becomes production. And what we do, I think we do a combination of both things of what you mentioned. First of all, we sort of do enforce this idea that your end goal is this production environment.

It doesn't have to be, but that's sort of what we do from the beginning. This environment that we create based on your code it's deployed into, it's isolated, it's designed to look like production. Because let's be honest, it probably doesn't function well if it isn't very similar, or even exactly the same as production. And the second part that we do is, we templatize all of those variables out for you. So, if you think about it, if you spin up your environment and it has a database, and a back end, and a front end, you need to wire those all together.

So, all of those variables have to be templatized, and filled in, and formatted. And we do all that for you. So, really, when you start up with this environment and it's the first time you've ever used it, you can play with it until you break it. And then you rebuild it, and you do it again, and you break it. And you mess it and you name the tables ‘test,’ and everything you want to do.

Corey: And then things aren't working and you keep fiddling with something, and one of your fiddling random winds up working suddenly, and, “Oh, my god. Touch nothing. That comment is now load-bearing.”

Regis: Yeah, well, that could happen too. But in any case, the point is you do reach that point where it's stable. And then you save that template—

Corey: Optimist.

Regis: [laugh]. Maybe. And then you're ready to go from there. So, you can make a new version— we version all of your templates; we version all of your application code. So, it's like, you can always go back, you can always move forward, you can feel confident that the one that was working previously, you can always go back to that. And then when you get a new one, you can work from there. And so it really does help and it covers all those bases that you talked about.

Corey: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the Enterprise (not the starship). On-prem security doesn’t translate well to cloud or multi-cloud environments, and that’s not even counting IoT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IoT devices, detects these threats up to 35 percent faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at extrahop.com/trial.

Corey: So, help me understand the career progression here. You went from a sizable environment that was cloud-native that was built out very well because you and the team got to determine that that was the way it was going to be, to go be the founding engineer at a four-person startup. And you did this towards the end of the summer, back when the pandemic was in full swing. At some point, it feels like… what is it? You missed re:Invent and decided, “Well, you know what, I want to go put it all on black somehow.”

Regis: Yeah, exactly.

Corey: What happened?

Regis: It's a huge gamble in the middle of a pandemic to leave a stable, comfortable job that I basically built with my own hands, and as part of a wonderful team of people. And the thing that really drew me to this was, it's sort of a purer software play on what I've already been doing my whole life. So, thinking of going into a data center and building equipment, and racks, and wiring networks, and putting it all together, versus building VMs versus running serverless, versus doing all of these new features with Kubernetes and containers. It's just the natural way of things that are going and being able to write the actual code that makes the code, if you will, or to build the factory that builds the factory, that's really appealing to me. And I just, I couldn't pass it up; it was just the right opportunity at the right time. And in the middle of all this thing, I just— I feel so lucky to have the option to do it at least. And I would have regretted it if I hadn't tried.

Corey: I feel like there's a definite story to be told there, where it's this idea of at some point, people decide to take the gamble, make the leap, and do the startup thing. And I feel like people look at me with starting my own consulting firm aimed AWS bills in that same light. Except it wasn’t. It was, honestly, I am a terrible employee. I mean, all the negative aspects of working with me that you remember or have the good grace to pretend you don't, didn't get better with time in many respects.

At some point, my core competency really became getting fired. And from there, it was— honestly, I started this place because I didn't feel I had another option left that I could stomach. And from my perspective, it was always this entire world of making sure that well, my answer is, “Or starve,” so I guess I'm going to do this, you were very much not in that position. So, I guess I'm just trying to vicariously see how the other half lives.

Regis: Yeah, I think of this is kind of like opposites but the same, if you will. I'm a team player; I like showing up at nine to five; I like saying, “Yes, sir. No, sir.” And you're totally the opposite. I think that's what makes it great. At the core of it, everybody's different, and everybody has different needs and different wants and all that.

Corey: Yeah, it's a hard problem. But if you take a look across the entire ecosystem of all of the various startup options, I mean, I worked with you before: you are no slouch, even in a pure engineering sense, let alone all the other stuff you bring to the table. There was a lot of options as far as directions to go in. You could attempt to swindle people with cryptocurrency, you could attempt to swindle VCs by just pouring AI and ML on something or, God willing, both. But instead, you decided to focus on the developer tooling aspect. What was it that drew you in?

Regis: Well, like I said, it's a pure software playing, what I've done my whole life. So, I feel like if we can do this right, we can stop building DevOps deployment toolsets. It's like, can we solve that problem and get it over with? If you think of, in terms of thermodynamics, most of the work that we do in DevOps are in the back end. It's sort of waste if you think—not negatively, but it's just leftover heat.

So, if you spend 99 watts of energy building all this stuff in the background that doesn't really do anything, that 1 watt of useful stuff that comes out in production, it's kind of like, can you just buy that? Can you just pay, Release to do your production environments or your staging environments, any kind of environment you need, can you just pay us to do it, and we'll do it for you and turn it out, and nobody ever has to do this ever again. That's really what I want to do.

Corey: Okay. I am going to do my best to keep a straight face for this question because I'm sure that your company has been asked this by investors. Are you worried [laugh] about Amazon moving into your space? I almost made it.

Regis: [laugh]. Look, you would be foolish to ignore them, right? In point of fact, most of their products are complete garbage, as you know. So, I am not too worried about it at this point.

Corey: Yeah, that's honestly, the fact of it is. And I know it's not a popular opinion to take, but AWS does awesome work at building infrastructure that scales globally, it becomes this amazing world-spanning thing. That's great. They build the bricks that you use to build, effectively, the digital equivalent of houses. They never pay the picture of the houses and every time they attempt to move up the stack into a SaaS offering… maybe I'm just lacking vision, or I'm forgetting something key, but are there any AWS SaaS products that we can point out and say, “That nailed it?”

Regis: No, course not. And like you said, they make the bricks. They have this concept of undifferentiated lifting. They say, like, “Well, everybody builds a data center, and everybody has to rack and stack servers, so let's take that away. Okay, that's undifferentiated.”

They said, everybody has to write scripts now to make an EC2 instance and a VPC, and, you know, how many CloudFormation templates, and how many Terraform scripts, and all of that stuff that has to be done? So, they've just made it differentiated lifting, I guess, at this point, where everybody does the same but different unicorn lifting. So, we're trying to remove all that. We're trying to actually make it so that you can point-and-click and actually realize that goal of having any kind of environment that you need spun up, and it's isolated and it runs. And it really is that simple. Not like, you start running some scripts and then you copy some things from Git and do all of that.

Corey: It just seems like there's such a lost opportunity here. I worry, honestly, that Azure is going to run away with the cloud world, just because GitHub as an acquisition, in hindsight, was brilliant. And it feels like they are one step away from, “Oh, deploy this thing on to some Azure service with a single click in the GitHub interface.” And at that point, well, wow, that becomes a story that I don't think any other cloud provider can touch.

Regis: Yeah. And what you've just described is what we're trying to do, and we're running like mad to get there.

Corey: Yeah, I really wish you well on it. As you're looking across the entire landscape now, what's the hard part? What are you focusing on? What's your area of, now it's time to do this piece of it? I mean, there's always the chop wood, carry water aspect of it, but there's also the question of in a world of infinite work—welcome to startup life—what do you focus on?

Regis: The main focus for me has always been the same during my entire career is to add business value. But for whom? The business value that I'm now trying to offer is for our customers. So, typically, try to go to a company and say, “Hey, you need to deploy this software, I can make a deploy.”

Well, now I'm saying to all my other customers, you know, all the customers who come to us for this product, I'm saying, “You want to deploy this, I can make a deploy.” The hard part is onboarding. How do we take every unicorn and put them into a box? It's really difficult to say, like—

Corey: Believe they're called ‘paddocks’ for one.

Regis: [laugh]. Yeah. Well, it's like, how do you corral them? How do you box them? How do you ship them?

Corey: Flattery works well.

Regis: Yeah. We really have to get focused on getting these customers onboarded. And once they use the product, it’s wonderful, but I think we just need to get that little ramp-up going nicely.

Corey: So, I have to ask, given my position in the cloud industry, what's your monetization strategy?

Regis: It's really simple actually. We charge basically, based on environments. We're working on the model, so it's, do we charge per our users? Do we charge our users who use the environments? Right now we just charge by the environment. So, if we manage a pack of ten environments for you, we charge monthly for that. And then you scale up, we'll deal with it at an enterprise level at that point.

Corey: Gotcha. It's always challenging to talk to the developer tools folks, just because on some level— I don't know how to put this politely so I'm not going to even bother. Some of the worst customers in the world are engineers—

Regis: Yeah.

Corey: —because you talked about a tool that makes their workflow better, and they look at that and they say something moronic, like, “$2999 a month? Please. That's serious money. I'm going to spend eight weeks building my own crappy version instead.”

And sure, okay, great, I get it. But if that's not the core competency of what their company does, there needs to be a correction story there. On some level, they're challenging to sell to on a variety of different levels. When I was first starting out with my consulting work, talking to engineers, it was, “Well, okay, so you do this, this, this, this, this, this, this, this, and this. What? That doesn't sound hard; I could write a script to do that in a weekend.”

Spoiler, if you can write a script that does everything that I do in a weekend, please call me. I have a job waiting for you. But it wasn't realistic, so at some point, it was, “Okay, I get it. No, you want to build your own stuff. That's cool. Thanks for meeting with me. Oh, by the way, can I chat with your boss real quick?”

It's one of those things of just talking to the folks who are empowered to see the higher-level strategy was sort of the right answer. But I was never selling developer tools. Do you think that there's a go-to-market strategy that works with folks that are, A) skeptical because developers love to build, and, 2) in many cases, where their signing authority caps out at 50 bucks.

Regis: No, you’re absolutely right. We're not targeting any kind of hobbyist. We're definitely targeting startups like ourselves who maybe don't have a lot of DevOps engineers or maybe no DevOps engineers, all the way up to enterprises where we would basically be competing with a team of DevOps engineers. But the way that we've approached it—you're absolutely right—is, first of all, we have to convince the potential person or the people there that these environments are important. So, what would happen if you had only one QA environment and you could only test one feature at a time?

How long would it take you to deploy new features in production? You know, all the things you need to do? What if you had 100 of those? What if your salespeople could go out and they could get a stable version of your product that you sell, but you don't have to worry about deployments and you don't have to worry about another salesperson sharing that environment. Ten minutes before the meeting, the salesperson, say, spins up a brand new environment, displays it to the customer, customer can maybe even—leave it with the customer so the customer can log in and use your product for a week, and then it disappears at the end of that or something.

What are all these values that can be unlocked by the environments? And then you think about even scaling that up to the holy grail of, like, what if you had more than one production environment? We talk about blue-green, for example. What if you had rainbow production environments? What would that even be? What would it look like? We don't even know.

The other thing is, in terms of value for the customers, what we look at is, how long would it take your DevOps team— or you yourself if you're the only DevOps engineer— how long would it take you to spin up a Kubernetes cluster that handles multiple environments, all the templating, with all of the configuration YAMLs, the 30 or 40 configuration YAML files that you need to manage, and all of those deployment pieces that you need glue together, all the integrations with GitHub, all the integrations with Slack, all that. It would not be a weekend exercise, right? We know, anecdotally, it takes about six months to implement sort of a new pipeline or new features that are in the production path, at a minimum. And if you're a startup, you don't have six months. You have a weekend if you're lucky. So, how quickly can you turn this thing on and get it running? That's what we focus on.

Corey: It seems that getting something out the door that developers like, but also solve serious enterprise problems is sort of the sweet spot right now.

Regis: Yeah, exactly.

Corey: I'm looking forward to seeing what happens next. If people want to follow along with your journey, where can they find you?

Regis: They can visit our website, it's just https://releaseapp.io. And they can sign up for free with GitHub]. They get two environments to start with. And for Christmas, we ran a promotion which is still valid, I believe.

You can get a free Minecraft server, so you can read our blog to figure out how to start a Minecraft server, it’s just something fun you can do and shows off what our environments do. But yeah, it's pretty fun.

Corey: It seems it. I will be following along and we'll of course put links to that in the [00:32:13 show notes]. Thank you once again for taking the time to speak with me. I really appreciate it. It's great to catch up.

Regis: Oh, yeah, thank you, too.

Corey: Regis Wilson, founding engineer at releaseapp.io. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you've hated this podcast, please leave a five-star review on your podcast platform of choice, along with a comment that includes an entire performance review over my nonsense over the past year.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About the Tim Banks
Tim Banks is currently with Packet, an Equinix Company, where he is a Principal Solutions Architect. His tech career spans over 20 years through various sectors. Tim’s initial journey into tech started as a US Marine, having originally joined the Marine Corps to be a musician. He was later reassigned into an avionics specialty based on the results of standardized testing. Upon leaving the Marine Corps, he went on to work for hardware manufacturers and defense contractors as a civilian. Later, he left government contracting for the private sector, working both in large corporate environments and in small startups. While working in the private sector, he honed his skills in systems administration and operations for large Unix-based datastores.

Today, Tim leverages his years in operations, DevOps, and Site Reliability Engineering to advise and consult with engineering groups in his current role. Tim is also a husband and a father of five children, as well as a competitive Brazilian Jiu-Jitsu practitioner. Currently, he is the reigning American National and Pan American Brazilian Jiu-Jitsu champion in his division.

Links Referenced:

  • Equinix Metal: https://metal.equinix.com/
  • Twitter: https://twitter.com/elchefe

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I am joined once again by Tim Banks, by popular demand. We spoke last year and it was fantastically well-received to the point where everyone demanded more of him. And you're back, Tim, in a new year with a new job. You were at Mission Cloud, and now you've joined Equinix Metal as a principal solutions architect. First, congrats. Secondly, what gives?

Tim: Well, thank you, Corey. I'm glad to be here, and I appreciate it. So, I spent some time in the TAM world, both at AWS and at Mission, and I made the jump into solution architect—actually principle solution architect at Equinix Metal. And I think there's a couple reasons why I made that jump.

First and foremost is that I still wanted to be in a place where I'm talking to the customer because that's where I think I can excel in both helping the customer and helping my company because I have an engineering background, I have engineering chops, but I also know how to talk to people and more importantly, listen to people and figure out what they need and figure out what they want, and then try and find a way to get those two things to happen. So, sticking within that role versus going to, like, a people management role, or versus going to some other kind of role, I think it was very important to me. The other thing that was very important to me were dollars.

Corey: Oh, yes. Hopes, dreams, and mission are all very important, but I can pay exactly none of my rent with those things.

Tim: No, I wish I could, with best intentions, but they don't translate well to keeping the lights on. You know titles matter, so to get the principal title was very important to me because I've done the work, so to be able to have the title and then the financial backing that goes with that title was extremely important to me. People have asked, “Well, you were only there for a year,” or a little more for a year, and, “What gives?” And the fact of the matter is that for many roles, for many people, and in many companies, you're going to get a far better raise by changing companies and changing jobs than you will by just trying to go through the normal internal promotion track. There's a lot of reasons for that, they vary from org to org, but the numbers don't lie.

There are a lot of people that have a career that looks like that, where you are staying at a place one to two years, and then you change jobs or change companies and there was a significant pay raise that went along with that. And if you look at what the pay raises that come with the internal promotions and review periods and stuff like that, they never match what you'll get by going to another company. I wish to say they would. I wish that companies would do more investment in keeping their people around, but that's just not the case.

Corey: Oh, I'm right there with you. It's one of those I'm lobbying for a three percent versus four percent raise after I've been somewhere for a year versus, well I can get twenty percent if I change jobs. And partially that's because my skills have changed and improved, presumably. But at some point, this obviously does cap out, but if I'm talking to someone else, and I'm able to get a few tens of thousands of dollars in raise, then why not? It feels like it's absolutely one of those things that lends itself to just pure self-interest if nothing else.

Tim: Absolutely. And so I have a strategy behind it. And so every year, on my year anniversary date or so, being with a company, I will look at other jobs and I will take some interviews. And I'll do that for two reasons: the first reason is because searching for a job and interviewing is a skill. It is absolutely a skill, and if you don’t—

Corey: Oh God, yes. And it evaluates you on a list of things that you really only bring into practice when you're looking for a job. So, if you've been somewhere for eight years, that skill has atrophied massively, and you're being judged based on that ability.

Tim: And it's atrophied massively, and the problem is too many people are exercising that skill when they absolutely need to instead of exercising it when they don't have to. It's like do you want to wait to run a mile when you're being chased by wolves, or do you maybe want to run a mile every so often to make sure you can, and then if you do get chased by wolves, you're very comfortable running that mile. If you are not interviewing; if you don't know how to interview; if you don't know what interviews look like now; if you don't know what questions they’re asking; if you don't know what skills they’re looking for; if you don't know how your resume supposed to be [00:06:09 unintelligible]; if any of these things, and then heaven forbid you get laid off, or for some reason like that, and then now you have to interview under duress to make your mortgage payment and to feed your children, that's a lot more stress, combined with being unfamiliar with what you've got to do. So, just as a means of practice and as honing my professional skills, I will go out, and look for, and interview for jobs every year.

Corey: I strongly endorse the entire approach.

Tim: And now the second reason I do that is because when I do go in for my annual review, for my performance review, or whatever, I want to know what the value of my skill set is to outside companies. And that way, when it comes time for review, and the company I'm working for now says, “Hey, well, this is what you are worth to us.” And that's what they're doing when they give you your raise or evaluate for whatever annual raise, or performance rate you’re due for, they're saying, “This is what your job is worth to us. This is what the value of your work to this company.” And if my work is more valuable to another company, I will bring that up to my current employer, and if they don't see that the value is the same, then I think we've reached that point, an amicable parting of the ways.

Corey: I agree wholeheartedly with everything you just said. I've gone on quote-unquote, “Practice interviews” past in my career, and either wound up hiring the interviewer, or that practice interview suddenly became very real, and I took a job there for a variety of different reasons. I don't have many regrets about my current situation in life, but one of them is that given that I own the company, it's very challenging for me to reliably and in good faith do that because it first sends a message about the company that isn't great, and two I can't imagine what something that would lure me out of here—that wasn't an acquisition—would look like. And I don't want to waste people's time and cast those doubts into the ecosystem. But I really enjoyed aspects of the job interview process.

Tim: Yeah, I think the one thing that I take away the most from the job interview process is figuring out how to present my skill set to different groups of people. So, I can say that, “Yes, I know containers, I know best practices for engineering, I know DevOps, I know SRE, I know all these languages, I know all these automation methods.” But how do I convey that to someone in the manner that they're asking the questions? If I interview for a startup folks from Silicon Valley, those questions are going to look a lot different than if I'm interviewing with someone like Amazon, or Google, or Facebook, or something like that. They ask questions differently, they look for different answers to find the same underlying set of skills. And so you have to learn how you can convey your skillset and how you can convey what other qualities have to different people who are asking the same question different ways.

Corey: Absolutely. It's also strongly valuable to understand this is a form of corporate hazing, for lack of a better term. In practice, you are never going to need to invert a binary tree or implement quicksort yourself. That's what lunatics would do.

But people love asking those questions, and it's more or less a, do the CS students’ secret handshake to pass this and show that you belong here. Which is deeply problematic on a number of different levels. But that's where it seems like it goes, and I've never been a fan of the trivia questions. It still feels on some level like these companies are interviewing for a lack of weakness rather than hiring for strengths.

Tim: Yeah. It is, in its very core, just gatekeeping. And not gatekeeping to say, “Oh, we have a standard and this is what's done,” because as you mentioned, most of these have nothing to do with the actual job you’re going to perform.

Corey: Yeah, as opposed to, what, all those companies that have no standards? Please.

Tim: [laugh].

Corey: Except maybe Facebook, but that's a moral issue.

Tim: Oh. No lies detected, Corey. But I think what it ends up doing is it ends up just satisfying the interviewer, which is a whole other problem in and of itself. I will say, though, that I will ask very upfront now, what does the interview process look like? How long is it?

Are there going to be, you know, a coding test, whiteboarding, like that? Now, and at this point in my career, if someone's going to have me up there whiteboarding for an hour, I'm going to pass because I don't have to at this point. And I've worked very hard. I've done a lot of whiteboarding to get to this point, but I don't have to anymore. It is an antiquated practice. People that suffer from various forms of anxieties, ADHD, or other kinds of neurodivergentcy at all, whiteboarding can be crippling and is not a good test of their ability to do a job, unless your job is to actually code on a whiteboard in front of strangers.

Corey: And it's awful. The last time I was subjected to that I successfully pivoted the interview with a, “Look. What is it you're actually trying to uncover?” “Well, whether you have technical competence.” It was for a DevRel style role. “Cool, why are you talking to me?” “Oh, everyone loves your work, and we're highly aware of your technical competence and your authenticity with developers.” “Cool. So, what is the purpose of this again?”

And it turned into a let me show you some code I wrote, and then I can make fun of it for you and see if you find that more engaging than me implementing some API I've never seen before. And sure enough, it worked out super well. But it's this entire idea of what are you actually trying to uncover. Becomes incredibly annoying.

Tim: Exactly. One of the interview styles that I like the most is probably the narrative style interviews where you want me to tell you a story, and from that story, you're going to glean certain things about my personality, certain things about my ability to troubleshoot, certain things about my ability to own a task, certain things about my ability to break something down or explain something, but you're going to want me to tell you a story, versus answering questions. And when I find, if I let people talk and tell a story in their own voice, I can figure out what it is they're trying to communicate to me fairly well. And maybe I'll ask some questions to dig in a little bit if I want to get a little more meat on something.

But I still essentially want them to do most of the talking: I want them to tell me about themselves; I want them to tell me about their experience in whatever way makes them feel the most comfortable. Because in the end, that's what you're going to be doing. If someone feels comfortable when they're working, they're going to be able to get across what they're thinking, what they want to do, how they can help. And that's what people should be looking for interview, versus Jeopardy.

Corey: I wholeheartedly agree with your position on this. But let's move on a little bit. I'm curious as to why you would go to work for Equinix after spending as much time as you have in, shall we say, the hyperscaler cloud world. It feels on some level like that's a very strange direction. Which based upon my understanding of you, is a near certainty that there's something I’m very clearly missing.

Tim: So, I will say that I first started off—because I've been in it for so long—in the days before public clouds, in data centers, and I've still, for the first few years of using public clouds had to break myself of the kind of mentality that most people have when they're working in data centers. I will say though that Equinix Metal—and Equinix as a whole—are moving in a direction that you would not typically expect from a data center provider. They are moving in a direction that is much more what you'd expect from a new up-and-coming cloud provider. Talking about trying to be more elastic and more agile, quicker deployments, using API's versus sending in a ticket to a human to do things, but also moving in a direction where they understand that you're going to have presence in a public cloud, or several public clouds, and how can I tie all these assets together? And that's what's fascinating about the whole thing because I know, having worked in public held long enough, that there are a lot of workloads and a lot of use cases that maybe you wouldn't want to run in a public cloud.

Corey: Well, tell me a little bit more about that because I would agree with you, but the answer I would give to that question in 2015 versus the answer I would give to that question in 2021 are radically different, and honestly, much smaller now than they used to be. It turns out, the Cloud got better. Anything that required deterministic performance was a no-go back then, well, that kind of changed. A lot of HPC workloads weren't effective. Well, now they kind of are. What am I missing?

Tim: So, I think the biggest thing that I see and something that will probably appeal to you, knowing you to be the person that you are, is that when you are running in a data center or when you own your equipment, your costs are very predictable. And that's something that you will typically don't see in a public cloud. If I need to spin up something real quick and in a hurry—I mean, I can cost-cut with Spot; I can cost-cut with Reserved Instances if I've made reservations, or I have an idea of what it's going to be in on-demand. But if everything I'm going to spend is unknown until the end of the month, that can be a little harrowing.

Corey: And let's not kid ourselves either. The things we spin up in a hurry, tend to be load-bearing forever.

Tim: Yeah. [laugh].

Corey: One of my production tables’s in DynamoDB, and my serverless application is called ‘clicktracker-test.’ Yeah, funny how that works out. One of the other tables has the word ‘dev’ buried in it because I make terrible decisions, and is it working? Great. Am I going to deploy it clean? Of course not. Should I? Yes. Will I? No because I'm bad at things.

Tim: Well, there's so many long-running production workloads that are still running off of some developer’s credit card right now. So, there are decisions that we make on the spur of the moment, without a lot of, necessarily, planning, that we end up sticking with—like you said—for that very reason. How many POCs or MVPs are still basically what people are running now? But what you end up with as an unpredictable cost. You can forecast it pretty well at this stage, but it's not predictable.

It's going to change month-by-month, year-by-year in ways that are not definitive. Versus making CapEx, where you buy an allocation of hardware and that's what it's going to cost you. The other thing I think, too, that really gets people on using public cloud is your bandwidth transfer. I know you've seen that before. It’s like—

Corey: Oh my god, yes. It's the Achilles heel across the board. There are still some workloads that are heavily based on that, that are not an appropriate fit for the Cloud. Look at Netflix, for example. They're very public about the fact that they don't stream anything from AWS.

They have their Open Connect custom CDN that they built out themselves. Given who they are, that makes perfect sense. But even imagining the heavily discounted pricing—and yes, anyone who thinks that something the scale of Netflix is paying retail pricing for cloud for anything is missing something very key here—it winds up being a complete difference. It's just a non-starter. Even extortionate levels of discount still make it untenable.

Tim: And it's so wildly unpredictable. Because you can have a workload that, “Well, we know what our compute’s going to look like, and we can predict what infrastructure we're going to need for traffic,” but if your traffic spikes and your bandwidth usage goes way up, that's going to be a bill you weren't expecting. Every time. And I've seen that sticker shock from people; like, “I don't understand how our AWS bill went up $45,000 this month.” And it was all in data transfer.

So, there's a certain amount of predictability around being in a data center on those things, especially within Equinix, that appeals to a lot of people. I do think, too, if you need special custom hardware, you can get similar configurations in public cloud providers, but it's still going to be an approximation. If you really want to be able to turn all the little dials that you need for your specific workload, that's still going to be something that you're going to need to have something on-metal to do.

Corey: True. Now, let's be very clear here; I do want to call things out. I am objective, and that is perceived as authenticity when in reality, I'm just a jerk. But if I pull up and click around in the Equinix website, the bandwidth pricing there—at default egress—starts at five cents a gigabyte, which, sure, it's better than the nine cents that the big providers are charging, but that still is non-trivial. Oracle, for example, starts at below a penny a gigabyte, which is, okay, that feels on some baseline level to be directionally correct. All the big hyperscalers all feel like they're still charging 1998 prices for bandwidth, and it drives me nuts.

Tim: Mm-hm. Well, because they can. There's no incentive for them to charge lower, even though AWS will say, “Well, we lower our prices on this all the time.” And yeah, they do; they're very good at lowering prices on some things, but they're still going to keep that bandwidth pricing, like you said, at that 1998 price.

Corey: “We have lowered the pricing on this type of instance that you don't use in a region you didn't know existed. We're counting that as another price cut.” At some point, it’s come on. It's this one of those things that actually matters to people or not?

Tim: No. No, it doesn't. But it looks good. It makes for good marketing material, so that's how it happens. When folks say, “Oh, we're going to go on AWS and we're going to save money in the long run because we're going to get a EDP and we're going to get a private pricing agreement, and they cut prices on their stuff all the time.”

But you signed an EDP and like—or the private pricing agreement—and you're like, “Oh, we still have to grow twenty percent year-over-year.” Well, can you imagine signing an EDP in 2019, and you're a business that makes its money off of a service industry or people going out places?

Corey: I don't have to imagine it. This is actually what I do as a consultant. I have customers who were hurting and how do they wind up restructuring those EDPs and PPAs? Contract negotiation with AWS is one of the things that we find ourselves doing a lot more than I thought we would when we started doing it. And it leads to great outcomes on both sides of it, but yeah, it turns out that there are hiccups in the road, that the growth is not always hockey stick upward and to the right. The real world is messy, and it happens. And we talk about architecture aspirationally, and then we talk about the architecture that we have here in reality, and one of those things is not nearly so shiny as the others.

Tim: It's funny, when I would have architecture reviews with customers, the first thing I would say is, “Tell me what your architecture looks like now. Let's talk about what you could make your perfect architecture look like and let's find out where in the middle we can land, actually.”

Corey: Oh, absolutely. I've been saying for a long time that you can design a working architecture to solve a customer's problem on a whiteboard pretty easily. That's something that anyone with a baseline level of certification or experience can wind up doing pretty effectively, and it'll work. But the challenge is, all right, where are you now? How do you get to that future state?

And what is the benefit of doing it? Because “It'll save us money in the longer term,” is surprisingly, uncompelling as a reason to undertake a project like that. There has to be a value proposition that goes beyond pure cost savings or it's unlikely to get done.

Tim: It's true. You're going to run up against either, the bean counter—there's always going to be a bean counter of some kind, whether it's engineering hours, or whether it's going to be actual money—as to say whether we can do this or not. And that's where you have to make those kind of proposals lucrative. You have to find out what motivates them to say, okay, well, this is now worth it, addressing whatever concerns I had. And that's hard.

But that's almost never a technical problem. That's almost never a technical problem. It's almost always resolving somebody's fears or resolving someone's insecurity, or sometimes stroking somebody's ego; more times the ego-stroking than I care to admit. But once you get there, then you can finally say, okay, now, we have enough buy-in to be able to make this digital transformation, or whatever it is they want to call it.

Corey: The big challenge that I'm seeing across the board is, we just had re:Invent happen, and I want to talk to you about that in a minute or two, but the challenge that you see is these companies get on stage, and they talk about how amazing they are and how their transformation has worked. And you look at that on stage, and then you look at your own environment, and you feel bad as a direct result. And it's almost like it's conference-ware on some level. In practice, no one is ever completely honest when they're on stage at a conference talking about their architecture. They're talking about a particular workgroup, or they're talking about a proof of concept. It is very rare that you'll see someone get on stage and say, “This is how it works across our organization,” and be accurate about it. Because most things are inherently brownfield or there's legacy that has to be accounted for. Legacy, of course, being engineering-speak for, “It makes money.”

Tim: I think it's interesting when we talk about what it takes to get—especially these large enterprises—to make these kinds of changes, and the thing that they don't tell you on the stage is, “Also, AWS threw us a million dollars worth of credits.”

Corey: Oh, I love it when people wind up making economic decision based upon credits. “Oh, they'll pay for the migration with MAP credits.” Or, “They'll give us a whole bunch of things through Activate.” And that's great and all, I appreciate the value, but things get bigger. I mean, I was astounded, relatively recently, to discover that with the DuckTools launch of our SaaS offering here, as well as the experiments we're running and the stuff we're doing for our customers, we're now a $25,000 a year AWS bill.

And that took me aback on some level. Now in practice, it doesn't really matter. That's two grand a month and change. And compared to the cost of running this place, no one cares. But it keeps going up. It only ever goes in one direction.

Eventually, we're going to care, and we're very well suited for when that day comes, but by the same token, it's also the cobbler’s kids have no shoes style of story. Yeah, optimizing our bill is always going to be a distant second place to optimizing our clients’ bills.

Tim: Yeah. And as business goes, almost everyone does it the same way. What I do think is important to note is that when you talk about being on AWS and that bill is always going to go up, you don’t, sometimes, even have to do anything. You just have to exist. Your bill will go up over time just because you're storing more objects in S3, right?

Because you're storing more EBS volumes. And if you don't do those optimization tips that we both talked about the past, then you will never save money. You can save money even if you are growing, but you have to be very purposeful about it. You have to be very intentional about it. You have to do—you actually have to build cost savings into your architecture; it can't be an afterthought because you're never really going to make anything happen.

If you have an architecture—and then I've worked with folks before where—one of my old customers runs advertisements and monetization for game ads. And a year ago, a year and a half ago, they pushed to change their workload into a very stateless workload that can be run on Kubernetes and in Spotinst. And so this allowed them to have very, very well-orchestrated Spot instances that was easily 95 percent of their infrastructure. So, that way, when the pandemic hit and a bunch of more people were home, and they were playing games—like, kids were playing games, adults were playing games like that—and their usage tripled. Their AWS bill only took a little bit of a hit up because they had built cost optimization into their architecture before they needed to.

Corey: Absolutely. Strongly agree. I mean, the sad part for us was during the pandemic when we saw folks have their user activity drop off a cliff, and their infrastructure spend remains stable. People talk about elasticity, but whether they're living it or not is really a subject of some debate. Many applications still require, if you add or remove a node from a cluster, all the other cluster members need to be made aware of that. That's not a great candidate for auto-scaling as it turns out.

Tim: It is very painful to watch people try to turn things down and realize that they cannot.

Corey: Oh, yes. Because historically—and I understand why. It made sense. If you scale up rapidly, great, you're going to be able to solve for that problem because if you don't, you're dropping customers on the floor. If you fail to scale down, well, you're just spending a little extra money, and that is never as big of a deal as it could be. So, it's always a trailing function; it's always something people think about after the fact. And the consistent lie we always tell ourselves is after this sprint ends, we're going to start doing everything the right way and clear on technical debt.

Tim: [laugh].

Corey: Yeah. Sure, you will.

Tim: Oh. It's like when people clean up the garage, like, “Oh, yeah, we're going to get in there, we're going to throw all this stuff out, and then the garage will be clean and we’ll be able to put both cars in here.” And either you've always been that way, or it's never going to happen. I've never seen anybody who had a dirty garage that's cleaned it out and then from that point forward can always put both cars in the garage. It just doesn't happen.

Corey: The only way it happens is if you wind up, effectively paying someone to do it for you—

Tim: Exactly.

Corey: —which wasn't the metaphor I was intending to go with, but that seems to be the way that I've experienced it. My home office was always a mess until I started paying someone to clean it up behind me. And that's helpful.

Tim: Do you mean paying someone, like, maybe The Duckbill Group to come in straighten out your—

Corey: It wasn't the direction I was going in, but… and again, it's one of those stories of, what is the actual impact a customer is looking to have, it's never the, “Oh, just pay us and all your problems go away.” It very rarely is that simple. But it's a start, and it highlights what you could do, what you should do, what we would do in your shoes. But it's varied; there's a reason that we do this as a consulting service and we're not just selling software in isolation. We're selling some power tools that we do use to solve very specific problems, but we're not doing this as a ‘sign up for SaaS and never worry about your bill again,’ because there's no API for business insight.

Tim: That’s true.

Corey: There's no context that you can glean programmatically to avoid recommending terrible things. And after a few terrible recommendations, people stop caring what you have to say.

Tim: I think that’s something that I talked about when I was working at Mission and something that I like to carry forward in any role: when you talk to customers—what I like to tell people who are new to customer-facing roles is that people can get AI recommendations for basically anything now. People can get machine learning and some kind of algorithmic recommendation for basically anything at this point. But what a machine cannot really effectively give you, right now at least, is insight based on your experience. And if you have a wealth of experience, and you've been able to go back and examine it, especially when the experience is making your own mistakes, right, of which I am very, very versed in, if you don't have that experience, you cannot say to others, “I know what the numbers say, but there's more to the story and here's why.” When you can do that, people will pay you for that because you are going to keep them from making mistakes.

You are going to keep them either from spending more money because they haven't optimized or from spent more money because something bad happened and now they have to pay out a bunch of customers. So, when you talk about a consultancy, you're talking about giving customers—you know, people paying you for you telling them things to do, right? That's important. The ability to say that, “Hey, I have been there, and this is what I have seen, and this is what I see for you coming forward if you don't handle this.” I mean, that's a big thing.

I think the garage cleaning analogy works, especially for us because oftentimes it's way easier for us to throw out somebody else's stuff than it is for us throw out our own. If I'm going into the house and I'm going to clean up the garage, if you don't tell me something absolutely has to stay, I'm going to throw it away. But if I'm going there, you're going into my house—or if I'm going to my house, I will find it, “Oh, I can't throw this away. My girlfriend from my sophomore year of high school, you know, I was wearing this when she looked at me the first time, so I got to keep it.” You know, like, all these little reasons why that you cannot really downsize things when it's your own.

Corey: Right. This is the big problem I see, [00:29:08 unintelligible] conference-ware where it's this idea of, here's what you could do if you wound up approaching it—the same sort with reference architectures, same story with the ‘hello world’ style stuff. The real world is messy, and well, there's context that means I can't safely do that for a variety of reasons. Sure, you can sit there and say that's because my process and culture are shitty, but that process or/and our culture also generates a sizable pile of revenue. So, how do I fix this? And meeting people where they are, rather than shaming them into meeting you where you are—as a vendor—seems to be something that all of the major cloud providers have been focusing on for a while now.

Tim: Yeah, I think it's interesting. We talked about it in the context of—you mentioned re:Invent before, we're talking about cloud providers. One of the things that cloud providers spent a long time doing was telling you that you shouldn't be on-premise and you should go to public clouds and here's various thousands of reasons why. They would do everything they could to get you off of being in a data center, being on metal, into public cloud. And if you notice at this year's re:Invent, they talked about putting a bigger emphasis on AWS Outposts, they talked about EKS Anywhere, ECS anywhere, they talked about the open-source distro for EKS you can run anywhere.

And they mentioned a couple things about migrating databases back and forth, and things like that. And what I think it was is an acknowledgment that people are still in the data center, will continue to be in the data center or will be expanding into the data centers for various reasons. And so instead of trying to fight against that, they are now trying to make sure they can still make their money while people are doing that. Which is brilliant, number one, but number two, it is actual customer obsession because the customers are telling them, “Hey, this is what we want to do.” And AWS is saying, “Okay, that's fine, but we still want your money.”

Corey: Yeah, again, there's a cynical perspective, and there are uplifting perspectives to take on it. But I don't work for AWS. I’ll let them articulate what exactly this [laugh] they mean on this. I am curious as to what you saw coming out of re:Invent. What did you take away from the conference? And by extension, what did we learn about ourselves from all of last year, both as individuals and as a society?

Tim: The thing I took away from re:Invent was they literally put a big emphasis on large customers—enterprises—reinventing themselves technologically, [00:31:25 crosstalk].

Corey: Because that's where the money is.

Tim: Oh, sure. Sure. And they had a couple of startups, but the biggest thing was, they were talking about how large companies changed to meet the demands of 2020. And there were a lot of ways that large companies had to do some innovation. They talked about, was it—I think the bank—I think was a Capital One, they cut that four-hour wait time for folks to talk to an agent to, like, ten minutes based on doing some serverless functions into Lambda stuff, which I thought was brilliant.

And that's something that everyone can understand. If you got laid off from your job, you got a pay cut, and you need to talk to the bank about something, can you imagine having to wait four hours on a weekday to talk to somebody about something super important, like, you know, mortgage payment, or bills, or something like that? That would be nightmarish. That is a real effect. That is a real thing that affects real people.

And I think is important. A lot of things that we do in tech, people don't really know, people don't really feel. You know, like, if you say, “Hey, I reduce response time on this API.” Well, that's all well and good, but most people are never going to know that. But if you said, “Hey, we did a lot of work so that you don't have to wait four hours on the phone to talk to the bank.”

That's something people were going to notice. And I think that's important. I do think that the overall theme of people changing and doing things differently through AWS technologies to meet the demands of 2020 was timely. I think it was inspiring in a lot of ways, but I think a lot of it, they're also kind of telling on themselves a little bit that maybe—the unwritten thing is, maybe the things that they were doing in the past aren't going to serve them going from 2020 on. The answer, the other questions like what do we learn as a society, and as individuals in this year? I think is—

Corey: Yeah, versus what should we have learned as a society, which we're going to stay away from because that is way too depressing.

Tim: Yeah. Here's what I think we learned: I think we learned, number one, that our cultures, our work cultures, and work organizations are nearly not as smoothly designed as maybe we thought they were and that they rely too much on people being in proximity, one with another, to exchange information, or to find information more importantly. I think that we realized that we don't document enough things. I think that we realized that things that we do you have documented are too hard to find. I think that we realized that we don't really know how to talk to, or check in on, or check in with people if we are not physically in front of them.

And I think that a lot of companies that were remote-first companies had a huge advantage over companies that were trying to pivot to being remote-based companies. And what you saw were some companies did it really well, and some companies kind of didn’t. And so those companies, you'll see where they're bringing people back as soon as they can. You know, so whether they should or shouldn’t, that's a different discussion.

But I think that's one thing we discovered about ourselves; we really learned, how is our culture? How is our company culture? How is our organizational culture? How are our practices in dealing with people who have real concerns? Without getting on the soapbox too much, but something I mentioned in the past is like, how do we address our employees’ real-life needs as a company?

Whether it's our female engineers, are we making sure that they're supported, if their mothers, so they don't have to leave the workforce? Are we making sure that people have time for mental health breaks instead of just scheduling them from meeting to back-to-back meeting for 10 hours every day for four weeks straight? Hey, are we checking in on our people just in general? Are we making sure our people—hey, do you have the power on? Do you have high-speed internet at home?

Corey: Even if you are full remote, as we were when we started this whole thing, it's not the same. And if wow, it turns out our company can't ever do remote work. Look how terrible that year was. Yeah, not really a fair test, if I'm being very direct. This year has been remarkably different.

Tim: It has been extremely different. And I think even for us, people who were full remote beforehand, we were able to still have the advantage of, every couple of weeks, couple months, being able to go and fly somewhere and see all the people that we talk to on Twitter, and have fun, and exchange pleasantries, and actually have real physical contact with these people. And you don't have that this year. And that's changed a lot of jobs. Like people who do DevRel, their jobs changed, I think, dramatically in 2020 because they're not doing conferences like they were.

Conferences have changed. The value-add for conferences has changed quite a bit where people are realizing that, do I really want to pay $1500 to fly to a hotel to hear people just give a talk, or would I rather hear people discuss panels or talk to people in the hallway tracks? A lot of people have missed the hallway tracks. A lot of people have missed interacting one with another, exchanging ideas completely off the cuff about something, and just watching where it goes. Just sitting there and having a person just talk on the stage doesn't seem as valuable, I think, anymore because people are sitting there watching videos of people doing the same thing. And it's maybe not as fun. And I'm not saying that's universal across the board. Some people are very compelling speakers—

Corey: The entire world is defined by exceptions. That's fine. No argument here.

Tim: Yeah. But I do think, by and large, people now are realizing the value of sitting around, in a group of people, in a room talking about this thing that comes up and just letting that idea take root, and electrify people, and write stuff down. And I've been at conferences that have started a conversation with someone in the hallway that ended up being a product in two months. And that's the kind of stuff that people pay for. That's the kind of enrichment that you have.

And yeah, the vendor booths and swag are one thing, having been both a participant and sponsor at a conference, the vendor experience is quite different, and I think that that needs to be rethought how that gets done going forward. And I do think another thing that needs to be rethought is, who are we bringing to the conferences? If you noticed, a lot of conferences were free, and people that could otherwise not attend because they couldn't afford, either personally, or for the companies to fly them to someplace and put them in a hotel—you know, it's always going to be fly into some very expensive place and put them in a very expensive hotel, to listen to—

Corey: And it’s people who can basically justify being not at work for a week for the duration of that conference.

Tim: Oh, absolutely. So, all these combinations mean that only a very few people in an interest group would be able to actually participate. But now when you say, “Well, you can do it from your workday, you can pop in the session, then work some. It doesn't cost you anything, and you can literally get to it from anywhere.” Now I saw more people—different types of people—actually being able to participate in some of the activities.

Mind you, they were different than a conference was before, but I do think what you're seeing is that maybe we need to find a different way of doing this so that we can get more people into these things; you get more people into these things, you get more thoughts, you get more products, you get more interactions. That's going to be a net benefit for tech as a whole, for whatever user group, or special interest group, or specialty that you're talking about. The more people you can get in there discussing more ideas, the better that's going to be.

Corey: I absolutely agree with basically everything you've just said. And for better or worse, there's an awful lot of things we need to learn as we move forward as a society. And hopefully, we'd look back on this with some silver lining we can take out of it, otherwise, it just becomes this entire year of nonsense.

Tim: Yeah.

Corey: And that's too depressing to really contemplate.

Tim: Yeah. If we go into 2021, when it becomes quote-unquote, “Normal” again and we're all going back to work, whatever that looks like. And we looked exactly the same that we did in 2019, then all the shit that happened this year was for nothing. And that would be the greatest tragedy of it all, if we didn't learn anything, we take anything away. I would really like to see the overall health of people—physical health, mental health, you know, what's your burnout level, what's your home stress level—become a more important focus for these businesses.

Because if you have a thriving individual there, that person is thriving, and that person is healthy, and that person is happy or as happy as they can be on whatever their circumstances are, they're going to be a better and more productive employee than they would be if they were [BLEEP] miserable.

Corey: Yep. Wholeheartedly agree.

Tim: So, what does this have to do with tech, Tim? Because I've had that question, too. It's like, well, when you're in your stand-up meeting, and you're talking to somebody, and you notice that hey, maybe they sound down, maybe they're a little off, maybe something's not right with them. Take a second and like, you know, “Hey, person, let me get five minutes after work.” And just check in on them.

See how they’re doing, and see if there's anything you can do. Because the real benefit of having people managers—versus tech and engineering leads—is that they're going to address things that aren't necessarily technical but that are good for org, good for the business, good for the employees. Like, if you handle somebody’s pay problems, if you handle somebody’s insurance problems, and certainly you can maybe handle—it’s like, “Hey, you need a day off and I noticed you haven't taken any PTO in three months. Why don't you take a couple days. Something. Get away from the thing. Let me switch around some on-call so you can have a break.” These are things that you should be doing as people managers. People management is a skill. It may be it's not a technical skill, but it is a skill that is very important to technical work.

Corey: It absolutely is. And I think that people lose sight of that at their own peril. And once again, we've had a terrific conversation on this show. If people want to learn more about you and who you are, what you do, what you believe, what you're working on, et cetera, et cetera, where can they find you?

Tim: They can find me on Twitter @elchefe: E-L-C-H-E-F-E.

Corey: Fantastic, and we will of course throw links to that in the [00:41:05 show notes]. Thank you so much for taking the time to speak with me. Every time we have one of these conversations. I always wish we could go longer.

Tim: I appreciate it. Corey, I'm always glad to come back anytime.

Corey: Don't worry, I'll take you up on that.

Tim: Awesome.

Corey: Tim Banks, currently a principal Solutions Architect at Equinix. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts, whereas if you hated this podcast, please leave a five-star review on Apple Podcasts along with an ignorant comment telling me why I should get over the high cost of data transfer egress pricing from cloud providers.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

Links Referenced:

  • Company website: http://summitroute.com
  • flaws.cloud: http://flaws.cloud
  • fwd:cloudsec: https://fwdcloudsec.org/
  • Twitter: https://twitter.com/0xdabbad00

Transcript

This episode is sponsored by our friends at New Relic. If you’re like most environments, you probably have an incredibly complicated architecture, which means that monitoring it is going to take a dozen different tools. And then we get into the advanced stuff. We all have been there and know that pain, or will learn it shortly, and New Relic wants to change that. They’ve designed everything you need in one platform with pricing that’s simple and straightforward, and that means no more counting hosts. You also can get one user and a hundred gigabytes a month, totally free. To learn more, visit newrelic.com. Observability made simple.

Corey: This episode is sponsored in part by LaunchDarkly. Take a look at what it takes to get your code into production. I’m going to just guess that it’s awful because it’s always awful. No one loves their deployment process. What if wanting new features didn’t require you to do a full-on code and possibly infrastructure deploy? What if you could test on a small subset of users and then roll it back immediately if results aren’t what you expect? LaunchDarkly does exactly this. To learn more, visit launchdarkly.com and tell them Corey sent you, and watch for the wince.

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by returning guest, Scott Piper, who is an independent consultant focusing on helping companies secure their AWS environments. Or as I like to think of it: railing against the tide. Scott, welcome back.

Scott: Thank you. Thanks for having me again.

Corey: So, you do an awful lot of stuff in the AWS security space. summitroute.com has become a mainstay of folks who want to have conversations about security. You run flaws.cloud, you're the organizer behind fwd:cloudsec. It's named after an email subject line, as all cloud conferences should be. And that's great.

But today, what I want to talk about is something near and dear to my heart, where you decided to set out and carve yourself a niche as a one-person band in the AWS space. For those who aren't familiar with my backstory, I did something very similar around AWS bills, then I took on Mike Julian, my business partner, two years ago and now we're 10 people. So, I sort of let the thing run away from me, you have kept it a single person operation, and I must confess there are times I'm deeply envious.

Scott: And I am envious of your position as well. The grass is always greener somewhere else.

Corey: It really is, usually because it's fertilized with crap. So, what were you doing before you decided, “You know what, I should be an AWS security person because that seems like the direction to go in.”

Scott: Yeah, so I talked a little bit about it in the last episode that I did with you where I was running security for a company, and I really didn't know a lot about AWS at the time, and specifically AWS security. I was supposed to be in charge of that for our company; I didn't know much about it. I tried looking around to try and identify who was a consultant that I could reach out to, to try and get them to assess our environment or to help me out in some way. And I wasn't able to find anybody. So, I knew there was demand because I wanted it, and I knew that there was no supply because I couldn't find anybody to do it.

So, I knew that that was an opportunity in and of itself. The other thing was, while I was at that job, I released flaws.cloud, which we talked about in the last episode, which is this online CTF. And as a result of that, I started to get a lot of emails from people saying, “Hey, you're the AWS expert. Can you help me out with a thing?” And I was like, “I’m the AWS expert—you know, AWS security expert?” Like, “Uh, okay. Yes. Yes, I am the expert, I mean.”

And so as a result of that I knew that other people were interested in this as well. And so again, I knew that there was an opportunity for this. And so the other big thing was that a meetup group had asked me to talk at one of their upcoming get-togethers, and they were based in another city and I emailed them, said, “Hey, that would be great. However, I'm in—” at the time I was in Denver, and they were in Chicago, and I said, “I would love to, but unfortunately it just doesn't really make sense for me to fly out there, buy a plane ticket, get a hotel room and everything, in order to talk at your Meetup group,” and they're like, “Well, we'll pay you for this.” You know, I wasn't going to profit off of this, but they were willing to pay for the flight and hotel.

And so someone that has never met me before was willing to write me a check for just a few hundred dollars; like, it was probably going to be $300 or less, but they were willing to write me that check. And so I knew there's an opportunity for this. And so as a result of that, I really just, kind of, dove in; one day just kind of announced it to people and said, “Hey, I'm an AWS security consultant. I am ready for business.” And unfortunately, emailed all the people that had emailed me previously when I was working for the startup.

And at the time, I had emailed all of them and said, “Hey, I can't help you. I'm super busy with a startup.” But then I emailed them now after I'd quit, and I said, “Hey, I can help you out.” And they're like, “Oh, we didn't know you want money for this, I thought that maybe you could just help us for free.”

Corey: “I thought you were volunteering for us?” Why not? It's one of those things like at some point if enough people ask you to do the thing you do for free, it's like, do I just look like a sucker is what you start wondering.

Scott: Exactly. So, I immediately started freaking out. Oh, I just quit my job and now I do not have any opportunities for income. What on earth am I doing? This is horrible.

But I ended up being really lucky because I had previously, years prior to this, I had interviewed at a company Duo Security and had basically interviewed for a position with them, and everything was going great, but they wanted me to move out to Ann Arbor, Michigan. And at the time, I was not interested in that. I looked at a map of America and where there's mountains, and there's not a lot of mountains around Detroit, and around Ann Arbor, Michigan. And so I was not interested in that. I want to be by the mountains where I can go hiking, and do those types of activities.

And so I said, “Hey, let's keep in touch. I'd love to work with you in the future at some point, but this is not going to work right now.” And so then when I decided to announce that I was a consultant, the opportunity arose where they said, “Hey, you can do some remote work for us, now.” And now they allow remote work and everything. And so, that was my first client was Duo Security, and giving me that opportunity to try and basically start doing things. So, yes, I mean, that was really how I got my head start on things.

Corey: There's something to be said, for having set yourself up as a consultant in that—when I've been trying to grab people in to do things, like content writing and whatnot, it turns out it's a colossal pain in the butt to pay someone who's not situated as something of a business entity. Whereas when it’s, “Oh, yeah. Here's a thing you can pay with a credit card,” or, “Here's my W-9 for whatever you need,” and just, “Here's how you wire money,” it becomes easier to throw money at folks on a certain level. And of course, depending on the customer, the requirements on that change significantly, where it's going through vendor assessments, which for what I do with billing is always a treasure and a joy. It's folks who don't realize that, yeah, maybe your one size fits all form doesn't fit me super well because I will not be getting $10 million of commercial vehicle insurance because we have no commercial vehicles; we need to buy one first.

And it goes through this one size fits, basically, nobody process. It's annoying, but it works. And we've gotten those down to a science. But as an independent practitioner, that was the biggest pain. I'm mostly biased for not doing extensive work with companies that had that level of requirement. Now, of course, we have 10 people, we can dispatch people to fill out those forms. But at the time, it was aligned with me not wanting to do all that extra work.

Scott: And I think that type of thing scares a lot of people away from trying to become a consultant in some way, and following this, kind of, independent consultant path. But what I have found is they're making a request to you, and like if you were to try and scope out the engagement that you had with them, you're going to have some back and forth about the engagement, where they're going to want a certain pricing, maybe, and you're going to haggle over that, and various things like that. And similarly, when it comes to some of those requirements that they may have, in terms of how much insurance you should have, or various policies that you should have in place, you can just tell them, “Hey, I only have this much business insurance, so if you want to do this engagement or not, you're going to have to change your contract with me to basically specify that I only need this amount of insurance, which happens to be the number that I'm telling you is how much I have.” So, you can do those types of things. Furthermore, a lot of clients that I found is they potentially are frustrated with some of their other vendors, especially some of the larger vendors in different ways and they know that I can be flexible in different ways.

And so as a result, they tend to help coach me to some degree or help me in different ways. They want this done and they recognize that my skills and ability are in this technical area, are in AWS security; my skills and ability are not in how to invoice them, or how to write a statement of work or other type of contract. And so I've been very lucky that a number of these clients are helping me to accomplish those things, to write up those documents and do things like that. So, that I think is important that people recognize that your good clients are going to try and work with you to make these successes happen. Furthermore, I mean, they'll find their own loopholes in their own situations, and so I have—for various clients for various reasons—I have subcontracted under existing vendors that they have where that vendor just takes a cut of however much I'm charging to the client. I have in one situation worked as a proper W-2 employee for a client with the understanding that I would quit after two weeks or something like that. Just ridiculous things where there's just loopholes in their own systems that—

Corey: We do the things we have to do to meet our customers where they are. Are you writing any code for your clients?

Scott: So, yeah, so I have done that. That was part of what I did with Duo Security. And so with the initial engagement, I was doing various other things for them, but then as that engagement was winding down, I told them that I was looking to make sure that—you know, I worked for other clients, they were super happy with me, they would have continued on with that contract, however, I personally wanted to end it so that I could make sure that I saw other environments. I wanted to be a consultant for a number of clients. And so I told them that I was going to be ending things.

They said, “Do you have plans lined up? Do you have a new customer or anything?” And I told them, “Look it, I don't really have plans for that yet, but I know that I want to build some tools. I want to make a tool that helps me visualize AWS environments. I want to make a tool that can help me better improve lease privileges and environments. Those tools didn't exist, I wanted to create them.” And they said, “Those sounds awesome. How about if we pay you to build those tools?”

And they will have them associated with their name and brand, and they will also help build me up by allowing me to you have a blog post on their company site that mentions me as an independent AWS security consultant. And so I'm able to take advantage of their audience that they have. And yeah, at the end of the day, I get paid for building a tool that I was about to build for free. And so as a result, CloudMapper and CloudTracker were the first two tools that I built for Duo Security.

Corey: And I've used them both, and love them.

Scott: Yeah. And it's been interesting, too because—so CloudMapper does network visualizations. And unfortunately, that no longer is something that I really maintain in that project because I built it, and that was the first time that I’d tried to do anything like that, that I'd seen anything like that. And unfortunately, visualizing the proof of concept and demo environments that I was running it in when I was building it, they looked great. If I have five nodes in an environment, it looks fantastic.

When you run it in a real environment with 10,000 EC2s or some other ridiculous number, it just becomes a massive hairball. There's just this absolute insanity, you're not able to make any sense of it. But CloudMapper does have a lot of other functionality in it now. And so it's able to audit people's environments; it's able to generate this web of trust to show you the trust relationships between different AWS accounts. And so it has all this other functionality, and that other functionality really came about as a result of working with follow-on clients who they needed certain types of tasks done in their environment; the easiest way for me to do those tasks was to build a tool to automate that process for myself, and those clients, for whatever reason, they don't want to be publicly associated with it, going through trying to open-source a tool is a difficult process for them, and so I ask them, “Hey, is it cool if I just put this into the existing CloudMapper project and just add that as just additional functionality into that tool,” and they've been fine with it.

So, as a result, CloudMapper has expanded over time, largely as a result of follow-on contracts that I've had with other clients that have been willing to pay me to do certain types of activities for them, which then benefited those open-source tools.

Corey: I think there's a huge value to doing stuff like that. I got out of anything that resembled implementation pretty early on because I found that when I was doing the advisory thing, it was extremely repeatable. Sure, you learn something new from every account, but it's something of a bounded problem space, at least on the costing side. And I can scale to give advice a lot more than I can scale to do actual code. And plus, when I find that I'm writing code for a customer and doing any implementation stuff, I'm suddenly beholden to their release timelines. And that is a great way to basically lose margin on the deal if you do what I do, and don't bill by the hour or by the day. So, it comes down to just having a good answer there.

Scott: So, I think it's important that when you are defining what you're doing for consulting, is that you make the decision up front: do you want to have long term engagements with clients where you have these repeated follow-on contracts, or they essentially have you on retainer—

Corey: Recurring revenue has things to recommend it.

Scott: Yeah. Or do you want to have engagements, which is what I do, where I love my clients that I've had, but I do not want to be in a position where they have to call me in the middle of the night; I don't want to be in a position where—like, some consulting businesses, their model is basically what's sometimes referred to as ‘land and expand.’ You get involved with the client, you do some type of work for them, and you have to constantly maintain it. And so as a result, you've basically infected that client. And you are constantly having to do repeat business with that client.

And so I say, this as a very negative thing, but it can be a legitimate business strategy and done in a good and ethical way. However, that's not the business that I want to do. I want to do short term engagements for clients, where I am doing an assessment for them; I am training their team; I'm doing something along the lines where there is a fixed end date for this contract, and I have delivered value and can walk away from it entirely.

Corey: Oh, yeah, I assume people are going to hate me by the end of the project, once they've seen my code and realize that I'm a giant fraud. So, yeah, I want to make sure that check is cleared at that point, by the time I hand anything over and oh, they're never going to let me back in the building.

Scott: Yep. And so yeah, so I mean, like, the code that I've written has been an additional thing that they can use. CloudMapper is this open-source tool that exists in isolation, that they can run on their own whenever they want. It is not incorporated into their existing applications and things like that. So, I don't have to worry about, am I writing it in the correct programming language? And does it work with the correct operating system or anything else in their environment? It is able to work in isolation. So, again, I think that that is important, as you're defining the engagements that you do, is you define how you want to do those engagements. Is there going to be a set end date? I know you've talked about in the past of having fixed-fee contracts, where you're specifying, I am going to deliver this type of value for this money, and we're not really going to talk about how long that's going to take, or—

Corey: That is the only way we operate, from a consulting basis. I take that back. We also have retainer agreements, where we will manage company’s cloud spend for them on an ongoing basis. But our optimization projects are—yeah, absolutely. We are pretty good at estimating how long it will take based upon a variety of different factors, and we are highly capable of turning that into a repeatable, scalable engagement.

And again, people are like, “Oh, so when are you going to just turn this into software?” Well, first, we turned some power tools into it and called it DuckTools, but the honest answer here is that I don't believe that software does this all on its own. At a certain point, there needs to be business insight. Otherwise, you're just doing things that don't make sense for the company.

Scott: Exactly. And I have talked with companies previously, where they've been interested in doing an assessment. I told them, “Look, I'm super busy with other clients. I can't really help you right now.” They're like, “But we really want an assessment. What is the bare minimum you could do for us?” And I’d say, “Look, I could run, like, CloudMappers auditing in your environment, but you could do it as well.”

But they will pay you for those types of things because they know that when I run it in their environment, I'm not just going to basically copy and paste those results into a report and throw it over the wall, I'm going to explain to them why these different issues are problematic. What is the different severity of those issues based on some things I know about their environment? You bring in some of that human capability to it. And that is the value that you're delivering is being able to tell them why and how, and what are the potential limitations or gotchas of trying to fix these different problems that you're pointing out to them. Because a lot of my clients, they are running vendor solutions in their environment that potentially have told them about different issues; they have generated alerts and alarms and things like that, but the client may just not know how to respond to some of those issues.

And so, that is where you're bringing the human element to coach them on how to actually fix those issues, or how severe is that issue? If the thing is telling you that you do not have something that's encrypted, a lot of the encryption in the cloud may not make that much sense.

Corey: Encrypted at rest on disk, why? That makes sense in a data center or your laptop, for example. Someone's going to take that away from you. But what is the threat model you're really guarding against? That they don't dispose of drives properly? Yeah, good luck.

Scott: Exactly. And on top of that, you will have those types of issues along with you have a public S3 bucket with all of your customer credit card information in it. That is a much higher severity issue to fix than whether or not some EBS volume is encrypted or something like that.

Corey: Oh, absolutely. But at the same time, it's also way better to wind up just checking the box if it's easy to do and doesn't really come at a cost or performance penalty than arguing with your auditor all afternoon.

Scott: Exactly.

Corey: So, it's smile, nod, pick your battles.

Scott: Yep.

Corey: I dabbled briefly when I was consulting with an infosec project or two, or I did assessments—you know, picture the really crappy version of what you do and you're sort of there. And it turns out that the talking points you develop for specific offerings, with specific tools, and specific parts problems sound like crap if you try to pivot it. For example, on the cost side, I am not partnered with any vendor in this space, even AWS, so there's no perceived conflict of interest. When I say, “I’m not partnered with any vendor in this space, even the cloud provider, and now I'm going to look at your security.” It's, “Are you simple? Why wouldn't you partner with people? You’re claiming you can do it all yourself?” The thing that serves as a selling point serves as a giant screaming red flag in a different niche.

Scott: Yeah, it can. I mean, I'm similar. I don't partner with any vendors in any way, and so as a result, there are some customers that don't make sense for me to do engagements for because they are very aligned with various partners. And they have a certain philosophy on how to do security, that is going to leverage different vendor solutions. And as a result, yeah, those customers are not going to make sense.

So, again, I think that that is important is trying to define that niche. And I don't know if we made this point specifically, but I want to make sure people are aware of it is that you do have to define what it is that you do, and to narrow your scope, and to not broaden that, even if a customer is potentially going to write you a check that maybe sounds enticing in some way, if you are broadening the scope of the things that you can do and it's going to weaken your ability to be an expert in a specific area because as you try to broaden that—for me, for example, I only do AWS security. If I was to try and cover Azure, and GCP, and potentially some of the other cloud environments, it's way too much. I'm not capable of doing that. If there's somebody else that is able to follow all those things, good for them, but I'm not that person.

I have to focus solely on AWS. And by doing that, by defining that niche, it's not only in terms of technically being able to follow these things, but it's also the relationships that you have. You're able to reach out to, potentially, people at AWS or potentially people at other companies that work heavily with AWS in order to get questions answered. You're able to sell yourself to customers because if you are just a generalist if you are able to say, “I am a programmer,” that's not helpful. If you're a technical person that's not helpful.

You need to narrowly define that niche and that is going to motivate people to reach out to you. Because all of my business is all inbound: it is all people reaching out to me saying, “Hey, can you help me with a thing?” And I say yes or no to all that. I'm not cold-calling people. I don't want to do that. And it's because I've narrowly defined that niche. When someone wants someone that can help with AWS security, I'm going to be probably on that list of people that are recommended to them.

Corey: Absolutely. It's about, basically, rising to the top of the mental SEO for whatever expensive problem it is that you solve. And I backed into my current positioning, where I had engineering skills, I’d done a lot of systems engineering, systems architect, solutions architect, style work, and it was okay, well, that's a skill set I can bring, what expensive problem can I align that with? Because the broader you go, the harder it is to wind up differentiating yourself. “Oh, I do AWS cloud architecture.”

Yeah, now I'm competing against Deloitte and Accenture. And yeah, they're Deloitte, but I'm Deliotte-ful. And [00:23:20 unintelligible] management consultancy side of it. Now I'm calling myself McQuinnzie and getting sued for it. But it becomes harder to differentiate.

Narrowing it down is the better approach. If I weren't as noisy as I am, I could have kept going. All right, I fix the horrifying AWS bill—which is what I do now—for SaaS companies in the Bay Area. Like at some point, it narrows it down so much that someone hears what you do and pops up with, “That's me.” What I try and do is look for the Rolodex moment where, when I describe what I do, I get basically three answers. One is, “What's that mean?” Cool. The other is, “Oh okay, that that sounds super useful.” But the one that really makes this work is, “Oh my God, I know someone you need to speak to.”

Scott: Yep, exactly. Yeah. And I think also, a lot of people don't recognize that they think that starting your own consulting business like this is a very scary thing that's very high risk. And to some degree, it's much less risk than being an employee for a company. And part of that is because if you were able to get this diversity of clients in different industries, or different geolocations or something like that, that it becomes a lot less risky because, for example, when this pandemic hit, that was a big hit on my business, I was very scared about it.

However, I had some clients in certain industries that suddenly were flush with cash, more or less, and were able to say, “Hey, we've always wanted to have some of this stuff done. We now have the money to do it, can you help us out?” And so I'm able to survive throughout the pandemic because of this. And I recognize that I'm very lucky as are potentially many of your listeners that you work in the Cloud industry in some way, and this is a job that you can do remotely, that you can do working from home, which is the best possible situation under a pandemic situation like this. But I do think it's important that people realize that it really isn't quite as risky going out on your own and starting a consulting business like this as they may think otherwise.

Corey: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the Enterprise (not the starship). On-prem security doesn’t translate well to cloud or multi-cloud environments, and that’s not even counting IoT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IoT devices, detects these threats up to 35 percent faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at extrahop.com/trial.

Corey: Absolutely. I think that's also something people lose sight of massively, to their own detriment, where when I talk to people about, “Oh, become an independent consultant—” again, my business partner back before we were partners, we were friends, and we were each other's first clients. And when he set out on his own to originally focus on application performance monitoring, it was, “Well, Mike, that sounds super risky. What are you doing?” And he made the various stupid observation that, yeah, I was an employee at the time and my core competency was pissing people off.

How many people would I have to piss off to wind up having a surprise meeting with HR in which they don't offer you coffee—that's the tell by the way—and not having any income anymore, versus a diversified business where I have multiple consulting clients. We've long since passed a point where no one company is more than 20 percent of revenue is what we have to attest to on some of the forms. In practice, I think our largest customer is far less than that because we are diversified. It's like, at this point, who would I have to upset in order to not have a viable business anymore? Pretty much everyone. And I'm not saying I can't do that, but I have to work for it now. It is counter-intuitively less risky than traditional employment.

Scott: Exactly. Yep.

Corey: Now, there are things that I absolutely found changed once I started taking on, first a partner, and then staff, where I maintain the hardest part of all of it, bar none—at the beginning, it was hard, now it's worse—is managing my own psychology.

Scott: Yes, [laugh].

Corey: It's lonely, especially when you're dealing with NDA’d stuff, and you can't talk with other people about it. When you have staff, you also can't complain down to your staff, you can only really complain to your peers, and when you own the company, it's hard to find. But the thing that really sticks with me is it’s sort of a blow to your ego when you start hiring people like we did, to do sales and to do cloud economics, and every person I hired was better at the thing I hired them to do than I was. And it was incredibly humbling, and it's awesome. I mean, then there's also the getting past the old protestant work ethic approach of, “Oh, if I'm not busy and doing these things and suddenly I'm handing all of this to someone else, oh, no, that means I might get fired.” Yeah, that doesn't apply anymore. And in fact, in most reasonable companies, it doesn't apply. It's, “Oh, you've automated or delegated something else. Great. Now you can get promoted, or go work on something that adds more value.”

Scott: Yeah, and I have not expanded beyond myself yet, but the psychological aspect of having your own business is a very scary thing at times. And you constantly run into kind of an imposter syndrome as you're trying to define yourself as this expert in this niche. Because no matter what, there's going to be someone that has narrowly defined a niche that is smaller than yours that is within your niche, and they are much more of an expert in that narrowly defined niche than you are.

Corey: Oh yeah, I found experts who can go up one side and down the other on spot instance pricing way better than I can. Good for them. I mean, that is super valuable, and I tag them in from time to time. I talked to Alex DeBrie from time to time on DynamoDB specific stuff because he's amazing. And I have no problem reaching out to folks who are domain-specific experts. It turns out that there's plenty of pie to go around.

Scott: And I mean, you really realize, knowledge is fractal in the sense that if you imagine those fractal images, where you zoom in and becomes more and more complex, and you keep zooming in and just keeps becoming more complex. And you have that same situation, with these niches that we've defined. Even though I focus on AWS security, there are all these other branches and smaller niches within that, whether it is technical niches, for example, if you're focused on IAM policies, or you're focused on networking related things. But then there's also niches that end up being defined based on your use cases, and the industries people work in.

And so there's all these different ways in which that can play out such that you do get very scared sometimes, “Oh, I really don't know anything about this thing, and I'm talking to this person who's asking me these questions, and I can clearly tell by how they're asking them, they know way more than I do.” But at the end of the day, you hopefully have some expertise across the board in that niche such that even though they're an expert in an area of that niche, that you still have some value you can bring to the table in other areas of it.

Corey: Oh, absolutely. For me, another hard part of bringing staff on was the—again, if I wind up completely crashing and burning when I was starting out—and that was a very real possibility. I didn't know what I was doing; I still don’t, but I've convinced enough people that I do, so great. But my entire marketing stick was more or less aggressively shitposting about Cloud on Twitter. And that's fine when it's me, but if I get a wrong take and become, effectively, unemployable industry-wide, then suddenly it's not just me anymore.

I've impacted the livelihoods of people who were depending on me. And that's a very scary thing. And I do want to point out as well that you and I are talking about this, we are both the whitest of white guys, steeped in enormous piles of privilege. The things we're talking about, “Oh, claim, you're an expert and everyone will take you at face value.” Turns out, it doesn't work for people who don't look like our overrepresented selves. And that's a problem.

Scott: I will admit, there's a lot of luck involved in it as well. I feel very lucky, very privileged, that I have been able to be at the right place at the right time, release the right thing, and have the right conversations with the right people, and have been able to compound that interest into something that has become bigger and better. But yeah, it is a situation where potentially not everybody has these opportunities. However, there are some things that I do think people can do to try and set themselves up for success. One of them as we mentioned, is focusing on a narrowly defined niche not to broaden yourself too far.

There's a lot of benefit to being a general practitioner in different ways, but I think focusing on that niche. And I think what you'll find is, as you focus in on that niche, that one, even though you've narrowly defined and focused on one niche, you start touching into other areas; you start to understand and get better at some areas that are somewhat adjacent to what you're doing. But also, one thing that people don't recognize is that it becomes a flywheel of sorts. Is that as you start becoming better known for this thing, that people start reaching out to you and start contacting you about certain things. They start sending you private messages, “Hey, did you know about this certain thing?” “Hey, there's this security issue that I'm looking at on AWS.” I get those types of private messages from people.

So, I become aware of different, potentially, news that's going to come out before the general public does, and so I'm able to digest and better research that before it happens. And so that's an important flywheel that ends up being generated there. But I think also, some of the more well-defined things that people can do is, one, as you're starting out, I think it's very beneficial as you're trying to get your name out there, as you're trying to build your brand, one of the things you can do is to do what's called a survey article. So, not where you're surveying people asking them questions, but a survey in kind of the academic sense, in which you basically write an article that is going through kind of the history of your niche or going through what are the existing tools? What are the pros and cons of them?

Kind of doing a review of what are the current problems in your niche, and what are the best tools for accomplishing that. And so you can create, kind of, some blog post that I think becomes very powerful. But a lot of people start searching for that and it's going to feed into you. If you put that on your own blog, and you include some type of pitch for yourself on there, eventually, people are going to find it. And so doing something like that, and doing things that to you, in your mind, might be somewhat boring, or it's just, like, a normal day for you to do this type of thing.

Or maybe it's some advice you gave to a client and that advice is not restricted by some type of NDA, that you can generally give that advice to other people. To start including that in blog posts as well. Start doing those types of things and that starts leading more and more people to you when you're putting that public effort out there. And the way in which myself and you have done it is to put that information out there for free because there's a finite number of customers that I'm going to be able to work with. And there are only a limited subset of people that have interest in AWS security that are actually going to want to be one of my clients.

So, as a result, if I put this information out there for free, that is going to benefit people that would not be my clients anyways. But some of those people are going to have the ear of, potentially, some of my clients. They're going to be able to recommend me to them because they saw this blog post because I am putting out what I hope is some valuable work for other people to see.

Corey: It absolutely is, and you're right; it builds the reputation. Other things, too, are the unspoken dependencies on this. For example, for a consulting company, something that I took for granted and my business partner had to learn early in his career because we had different upbringing stories, and he's very open about this, is learning to speak to different levels of the organization at the same time without alienating any of those constituencies as you do it. It's a baseline consulting slash professionalism skill. And if you grew up surrounded by it like I tend to—my dad taught me to handle job interviews when I was 12 on a lark, and that's one of the best things he ever did for me.

But if you haven't had that exposure, it's like learning a new language. And it takes time to start sounding like your client in many respects. And there are ways to short circuit that, but we have it all the time. For example, but all the blog posts I do—I don't think I've ever talked about this publicly, but it shouldn't come as any surprise—every blog post that goes out with remarkably few very rapid response exceptions go through a technical editor. I mean, I'm not terrible at writing, but I will absolutely throw it all through an editor who makes it better just by virtue of being able to take that different perspective of being able to look at it differently.

Scott: Yeah, I think that that type of concept of having different types of conversations with people is also kind of eye-opening, is that previously, when I was an employee of a company and I talked with the head of security to another company, ultimately, they were trying to hire me. At the end of the day, if it was a tech company in the Bay Area, they're constantly hiring everybody all the time trying to do whatever they can. So, that conversation becomes very much them talking about their company, and how awesome the snacks are at their company and things like that. Whereas now, when I have a conversation with someone who's the head of security, or otherwise, at a company, it's a very different conversation, where it tends to be focused more on the problems that they're having, as opposed to the benefits of their company. Because the value that they're going to be getting out of the conversation with me is whether or not I can identify areas that I could potentially help them in some way, or give them some free advice on how to do things in a different way. So, you end up having these very different conversations with people as well.

Corey: Absolutely. And one thing that I, every once in a while, I like to do on this show, and I haven't done it for a while, so you get to be the person who sits around while I do it this time. If you're listening to this, and you want advice on something like this, be it setting out on your own, be it career advice, please reach out to me. The entire reason that I have a career at all is that people who didn't owe me a thing did me favors. And you can never repay that you can only ever pay it forward. So, if you're listening to this, and that resonates with you, please don't hesitate to reach out.

Scott: Yep. I do try to give people advice. I will say though, oftentimes, someone will ask me, wanting some mentoring advice and instead they pitch me on a startup idea that they have which, please don't do that. I am in no way able to fund your startup idea or give you, really, any advice on that. If you have a tech startup, I cannot help you. If you are considering doing consulting, I am in a much better position to try and give you some advice there.

Corey: Yeah, I forgot to clarify. Yeah, I'm not looking to cut checks for folks, as much fun as that is. I'm much more interested in being able to provide insight, things I learned. If I were to start this over again knowing what I know now, I could probably get to where I am now in roughly half the time.

Scott: Exactly, yeah.

Corey: There's just something to be said for having someone to talk to about these things. Every once in a while I say something like that, and someone will corner me. It happened at re:Invent a couple of years ago, and he asked for advice, said, “I heard you say ‘reach out,’ so I did.” He's now an AWS community builder, and he is just transforming as far as modernizing his entire career and skill set. It's really something to see when people take you seriously on stuff like that.

Scott: And I will say also, it's important that as people start finding some success in whatever they're doing, that you do try to raise up other people around you that are entering it as well. Even though I've defined this niche of AWS security, there are a number of other people that are doing amazing, fantastic things. And I try as hard as possible to raise them up, to point out who they are, to credit them. You know, if I ever mention any type of tool, or blog post, or anything like that, I try to make sure I also identify who the person is that wrote that tool or blog post. I think that that is really important as well, is that you don't step on other people in order to get to where you're trying to go.

If you raise other people up, for all sorts of reasons it becomes a better situation. One of them being that you create a better community within your niche. This is ultimately what I've decided to spend a lot of my life doing; a lot of my waking hours are focused on AWS security. I don't want the environment of AWS security to have a bunch of jerks within it. And within the security community, specifically, there, unfortunately, have been a number of situations where there have been jerks within the security community.

However, I tried very hard to make sure that that does not happen within the AWS security community. I have a limited ability, obviously, to impact that. But I do where I can, to try and make sure that it is a friendly community, it is a welcoming community. I try and do that through fwd:cloudsec, the conference, that is a nonprofit, and try to make that an opportunity for people to attend, to speak at, to be a part of. There's also the Cloud Security Forum Slack, which is invite-only, however, if you ask me or anybody else that's part of it, we'll invite you.

The only reason we've chosen not to have just an invite link for it is because we have found in some of the other Slacks that we've been a part of that if a Slack is completely open, it ends up getting a bunch of spammers in it, unfortunately. So, that Slack though is something very open and inclusive for people to become part of. And what’s been really great is fwd:cloudsec, the conference because we ended up being virtual this past year as opposed to being an in-person conference, it ended up being a lot less expensive. And we were lucky enough that almost all of our sponsors stuck with us. And they capped at the same price that they were paying us, and our costs were much reduced.

And so as a result, we were able to take the funds of fwd:cloudsec, the conference, and use that to pay for that Slack because once you reach a certain number of people in a Slack, it ends up costing money, but also because it's a nonprofit, we're able to have Slack’s kind of reduced pricing as a nonprofit. So, all these things in a blending in well together.

Corey: Oh, yeah. I run—well, I—one of the admins of the Open Guide to AWS Slack channel, and that has something like 13, 14,000 people in it. You can join it yourself if you’re listening to this at slackhatesthe.cloud. So, that is the URL: slackhatesthe.cloud. And you can wind up popping in there. It's just the direct invite link. Yeah, occasionally we have spammers, but you're right in that it's hard to run community well, and having vibrant communities out there for stuff like this is super important. So, if people want to learn more about who you are, what you do, or just want to pick your brain, where can they find you?

Scott: Going to my webpage, summitroute.com; that's going to be the easiest way to find me. You can try searching for me on Twitter. My handle unfortunately is written in hex speak, it's @0xdabbad00, but just try searching for Scott Piper on Twitter, and hopefully, I'll come up.

Corey: Excellent. I will make it a point to do that. And of course, it'll be in the [00:41:36 show notes]. Scott, thank you so much for joining me once again. It is always a pleasure to talk with you.

Scott: Thank you.

Corey: Scott Piper, AWS security consultant. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you've hated this podcast, please leave a five-star review on your podcast platform of choice, and tell me which vendor I should not partner with next.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About the Guest

Tabitha Sable has been a hacker and sysadmin since the turn of the century. She serves Kubernetes as co-chair of SIG Security and an associate member of the Product Security Committee. She loves to build tools and make friends, and puts those skills to work coordinating the efforts of the infrastructure, security, and product teams at Datadog. Outside of work, she can often be found organizing or participating in Capture the Flag contests and loves "pretty much anything with wheels."

Links Referenced:

  • Datadog: https://www.datadoghq.com/
  • “The Ironies of Automation”: https://www.sciencedirect.com/science/article/abs/pii/0005109883900468
  • International Journal of Proof of Concept, or Get Out The [BLEEP] out: https://pocorgtfo.hacke.rs/
  • “Reliable Code Execution on a Tamagotchi”: https://doc8643.com/pocorgtfo/pocorgtfo02.pdf
  • “What happens when you type google.com into your browser's address box and press enter?": https://github.com/alex/what-happens-when
  • Twitter: https://twitter.com/tabbysable

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored by a personal favorite: Retool. Retool allows you to build fully functional tools for your business in hours, not days or weeks. No front end frameworks to figure out or access controls to manage, just ship the tools that will move your business forward fast. Okay, let’s talk about what this really is. It’s Visual Basic for interfaces. Say I needed a tool to, I don’t know, assemble a whole bunch of links into a weekly sarcastic newsletter that I send to everyone. I can drag various components onto a canvas: buttons, checkboxes, tables, et cetera. Then I can wire all of those things up to queries with all kinds of different parameters: post, get, put, delete, et cetera. It all connects to virtually every database natively, or you can do what I did, and build a whole crap ton of Lambda functions, shove them behind some APIs gateway and use that instead. It speaks MySQL, Postgres, Dynamo—not Route 53 in a notable oversight, but nothing’s perfect. Any given component then lets me tell it which query to run when I invoke it. Then it lets me wire up all of those disparate APIs into sensible interfaces. And I don’t know front end. That’s the most important part here: Retool is transformational for those of us who aren’t front end types. It unlocks a capability I didn’t have until I found this product. I honestly haven’t been this enthusiastic about a tool for a long time. Sure they’re sponsoring this, but I’m also a customer, and a super happy one at that. Learn more and try it for free at retool.com/lastweekinaws. That’s retool.com/lastweekinaws, and tell them Corey sent you because they are about to be hearing way more from me.

Corey: When you think about feature flags—and you should—you should also be thinking of LaunchDarkly. LaunchDarkly is a feature management platform that lets all your teams safely deliver and control software through feature flags. By separating code deployments from feature releases at massive scale—and small scale, too—LaunchDarkly enables you to innovate faster, increase developer happiness—which is more important than you’d think—and drive transformation throughout your organization. LaunchDarkly enables teams to modernize faster. Awesome companies have used them, large, small, and everything in between. Take a look at launchdarkly.com, and tell them that I sent you. My thanks again for their sponsorship of this episode.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Tabitha Sable who, among many other things, is a system security engineer at Datadog. Tabitha, thank you for joining me on the show today.

Tabitha: Thank you very much, Corey. I'm delighted to be here.

Corey: So, you're relatively new at Datadog, which for those who are relatively new to the sector is apparently a monitoring company, not as it turns out when you mispronounce it, Tinder for Pets. I learned that the fun, difficult way when I got yelled at on Twitter for that a year or so ago.

Tabitha: I'm sure that was hilarious.

Corey: I said, “Who wants to date a dog?” Well, ideally, compared to the trash fire that is some people, lots of folks. But I digress. So, you've been there since April, sort of the very beginning of the pandemic era, for lack of a better term. What was that like joining a company where, “You’re here just in time to never meet your coworkers in person.”

Tabitha: It's been really wild. I work remotely from the Midwest. The two biggest Datadog offices are in New York and Paris, and the teams that I'm associated with are, for the most part, split pretty evenly between New York and Paris. But also within each country, a fair number of people are remote, and I'm one of those remote people. So, I had been planning all along to be working remotely, but not in this way.

So, it's interesting that I've been in the Datadog office, but never as an employee. For people like me, there's a few unexpected advantages. Like, for example, when you're meeting someone over Zoom, it's much more difficult to not be able to remember their name at a time of high pressure where you would cause yourself embarrassment because you can just look at the little words that are underneath their picture.

Corey: Well, I've seen people who are still capable of messing that up perfectly effectively. Messing up Zoom etiquette is always a terrific way to do things. Especially if you just kill video, and then we can go back to the old call center days of putting inordinate amounts of faith into the effectiveness and accuracy of the mute button.

Tabitha: [laugh]. Oh my goodness, yeah.

Corey: It’s like, “Well, you've worked with electronics. How far do you trust them?” “Well, not super far. Certainly not bet my entire career on it.” Yeah, there's a reason that more people should exercise caution with those things.

Tabitha: You know, I’ve been accused of being wholesome, and honestly, I think I'm just a little too simple to get into those kinds of situations. I know that I would absolutely be the worst at remembering whether the mute button is on or not. And so I just try to keep it on the up and up, and for the most part that saves me from myself.

Corey: You want to talk about the most destructive virus in the world: something that has Zoom tell you you're muted when you're not. I feel like that would be one of those old-school prank viruses that—

Tabitha: Oh my gosh.

Corey: —you can talk about things that would transform society? That's right there with it.

Tabitha: Yeah, yeah. You would learn so many things that a lot of people would be unhappy with, but in a lot of cases, when somebody shows you their ass, and then that you need to treat that person with care. Like, that hurts in the short term, but it saves you problems in the long term.

Corey: Yeah, it's painful to discover that people that you look up to and admire are trash goblins. But I prefer that to not knowing, and unknowingly introducing them as someone, “Oh, they're great. You should talk to them.” People remember when you introduce them to horrifying nightmare people.

Tabitha: And bad news is unfortunate to get, but it generally beats not knowing. Yeah. A hundred percent.

Corey: Yeah. So, I know that we're talking about things in the past rather than forward-looking. And I see no reason whatsoever to change that. I don't believe this is a core function of your day job, or at least I hope it's not but you moonlight apparently, as a Unix historian.

Tabitha: [laugh]. Yeah, it's fun. Unsurprisingly, there's history there. So, when I was just a tiny little person—and by that, I mean a teenager—I had the experience of being involved in the computing club at school. And this was around the turn of the century.

And so we had what were at that time a trash heap of outdated Unix workstations. We had an R 3000 indigo, like Iris Indigo, like the adorable purple dorm fridge, little SGI box; there was a Sun-3 machine with a VT100, plugged into a 25-pin serial port on the back of it. And there's a funny story there, which I'll share with you later if you're interested in. But those sorts of things; we had these Unix workstations that were 15 years old at the time. And so a lot of my early exposure to Unix and to system administration, was from, kind of, herding these wheezy old pieces of crap into providing some useful services for students. The first server that I ever ran myself was a 386 that I put OpenBSD on so that people can play Rogue.

Corey: Oh, that is a blast from the past. I mean, I started out in the university setting as well; my first professional job in the space was 2006. So, I was a bit after this to some extent, but I was a grumpy Unix admin because, let's be honest, have you ever met a different kind of Unix admin? I haven't. And BSD was the way that we went—full stop—for a year because I had opinions, and if no one really has any better idea what direction to go in, it's not the worst direction you could pick.

Tabitha: Yeah. So, yeah, we had all of these things. And that was where I got a lot of my really early experience with the concepts of system administration, with providing a service that people would actually use, and that would help them achieve some goal or help them have a good time. And so I got interested in those sorts of machines and in the software from that era. And then that intersects really well with this sort of general thing about the way that I understand the world, where I don't really feel comfortable with my understanding of something as it is unless I can connect it to a little bit of history, a little bit of how did it get to be this way.

And so sometimes I'll ask, “What does that word mean?” But I don't mean, “What is the definition of that word in the dictionary?” What I really mean is, “What is the etymology of that word?” So, this kind of point of view sort of permeates my experience of life, and that in common with my early history with Unix, has made me prone to researching what was going on in Murray Hill, New Jersey in the early ’70s.

Corey: Tell me more. So, in the ’80s, there was something worse going on in New Jersey, namely, I lived there. And my favorite part of New Jersey is not living there anymore. But the ’70s are before my time; tell me more.

Tabitha: So, this story starts a little bit earlier than that, in the ’60s. So, computers were expensive, but they were starting to get to be really good. And so—

Corey: Well, Apple is working on reversing that trend, too. They're trying to make computers more expensive, bless their hearts.

Tabitha: So, yeah. In the ’60s, there was a group project, sort of the group project made in heaven I guess, between MIT, Bell Labs, and Honeywell—if I remember correctly—to produce a computer that could be used as a service, like the computer utility. And so they were making Multics, which is this huge operating system that was very highly specified, and some people still believe was the best system that never finished getting built. And so that project was kind of slowly grinding to a halt under the weight of its own design specifications. For example, one of the design specifications of Multics was it had to be written in PL/I, but it turned out that at the time, there was no PL/I compiler.

And so then they had to define the implementable subset of PL/I and then write the operating system in that because even the language was too big to compile at the time. But the project was kind of grinding to a halt, and eventually, Bell Labs pulled out of it. And some of the folks who were working on that were disappointed because they had a huge Honeywell mainframe all to themselves, and it was running great with three or four users, and it was a really collaborative experience that was unlike what you got on other systems at that time. So, they wanted to start doing that same kind of computing, but also management at Bell Labs was very serious, they were not going to get into another operating system project.

Corey: So, many good things came out of Bell Labs and its history could have been so different.

Tabitha: Right?

Corey: We take a look at the modern technologies we take for granted, a majority, it feels like, and trace their lineage back to Bell.

Tabitha: Oh, my goodness, yes. Learning about what it was like to be at Bell Labs in that time is so interesting because it seems like very much a product of its time, but also one of the best products of its time. So, yeah, a bunch of the people that had worked on Multics wished that they had a computer. And they couldn't buy one, so they looked around, and they found a PDP-7 that was, like, sitting in a corner of a lab room because it had been bought for a certain project and the project was finished. And it wasn't big enough to run anything on. So, they were like, “Sweet, can we just use this?” And the answer was, “Yes.” And so that was how Ken Thompson, and Dennis Ritchie, and friends started writing Unix.

Corey: It's amazing just to think back to those days when getting even time on a system was this incredibly valuable thing. How do you get access to this? I mean, the reason that I went in a software direction instead of hardware, personally was, let's not kid ourselves; people who have seen me on Twitter have a pretty good idea of exactly how accident-prone and clumsy I am in social interactions, let alone in the physical world. It maps. When you screw up hardware, that's super expensive, especially when you're me as a kid and have basically no money to use on these things.

And whereas with software, “Well, I can just make a backup copy.” And you always remember to make a backup right after you really could have benefited from having made a backup. But you can unwind software and start over in different ways, whereas with the physical world, generally speaking—though there are exceptions—computers don't ever work the same way after you let the magic smoke out the back.

Tabitha: Yeah, yeah. There's the whole range of horrible things that you can do to your hardware, some of which are recoverable and some of which are not. And where that border is depends very much on what skill level you have, how steady your hands are, and what kind of tools you have in your house.

Corey: So, that does, of course, lead to the question, as someone from the perspective of a Unix historian, here in this year of our lord 2020—as of the time of this recording, but it’ll be 2021 by the time it hits the world—is the operating system relevant anymore? Does it matter, other than to a few folks who are going super deep into the stack for tuning or performance reasons?

Tabitha: I really love that question. That's a hard one. I feel like the answer is yes and no. It kind of doesn't, to the extent that you don't use it. And so a lot of the things that made Unix feel comfortable—for a certain sort of person anyway—are getting to the point where they don't matter anymore because they’re UX concerns, and the UX is replaced.

So, if you're talking about deploying containers onto Kubernetes, or on to some kind of cloud-hosted container service, there's Unix in there, but you don't ever touch it. You write the Docker file and you build the thing, and the rest of it just goes. And so to that extent, it kind of doesn't. But there's a different way of looking at it, I think, which is, what does the developer experience look like? And that gets more and more relevant as the user experience gets less relevant because of the fact that you're shifting your relationship to the operating system.

The operating system never goes away, but you're not banging at a keyboard typing in commands anymore. And so the sorts of concerns—like, how memorable are the command words? How lengthy are they? Are they annoying to type—become much less of a thing, but is the API of high quality? Is it easy to write to the API? Does the API help you to not write bugs? All of those sorts of things become more and more important because that becomes your dominant way of interacting with the operating system. And so, yes and no.

Corey: It feels on some level like it's akin to the network where, in a modern cloud era, you don't have to think about the network because it's all abstracted away. And I've checked with AWS extensively, they will not let you go hands-on hardware with their switching equipment because they don't believe in having fun. But I found that when I got deeper into learning networking—during the Great Recession when suddenly there's a salary freeze, so I want to leave my job, but no one's hiring, so I can't leave my job. What am I going to do to pass the time? I'll get my CCNA.

And I learned a lot about how networks work. And not that I wanted to become a network engineer, but suddenly, I understood a lot more about how the systems that wrote on top of the networks were built and why they did the things that they did, instead of hand waving my way past subnet masking, for example. I understood what the hell it meant. It wasn't the magic spell that you would just repeat because it's what you heard the elder say; now it's something you understand. And that opened up a lot of doors.

Tabitha: Absolutely.

Corey: I can't shake the feeling that understanding operating system fundamentals is heading in that same direction. Sure, most of the time, I'm just writing a Lambda function; I just want it to invoke the code, and I want it to do the thing until suddenly I see an emergent behavior that I don't fully understand. What's it doing under the hood? I don’t know. I used to work in shops where there was no one to really escalate to beyond me, which is, A) a giant problem for those companies, but, B) it teaches a certain level of self-reliance where, how do I start finding out what this looks like?

Why do I build a reproducible skeleton case for a mailing list? How do I frame a question? I mean, far and away, the best way I ever found to solve a problem myself was to start writing an email, “When I wind up doing this, just like it says in the documentation, except the documentation has a comma there, and I don’t—oh.” And then we close that draft and never speak of it again. There's something to be said by framing the question in such a way—like, rubber duck style—that opens up a bunch of understanding. But figuring out what the underlying tools and services are that power this, seems like it's a critical step. It's not super appreciated these days.

Tabitha: I think that this is a really interesting tension that we're having to deal with right now. I feel like I would get voted off the internet if I didn't call out the paper “The Ironies of Automation” right now. The author escapes me right now, but it's easy to search for on the internet. And it's great. And I feel like that applies here, too, where the whole advantage of making these abstractions is that it lets the people who are riding on the upper layers be able to achieve their goals without having to actually understand everything that's going on in the lower layers.

But there has never been an abstraction that didn't leak. There's never been an abstraction that was actually perfect. And so there are always emergent behaviors that come out because of the way that the abstraction is implemented. And the real substrate that's underneath of it affects the layers above in ways that the abstraction says that it's not supposed to. And I think it's beautiful, and also frustrating.

So, that balance there of, I'm so glad that we have all these layers of abstraction because it demonstrably lets people get so much more done, but it also makes room in the world for people like me, who just cannot stop themselves from opening something up to see what's inside of it. Especially if you tell me that I don't need to worry about what's inside of it, or I don't need to think about how that works. Nothing makes me want to dig into it more than that.

Corey: If people are asking for, “Well, I have a couple of weeks off. What should I be doing with that time in my spare time? Because I enjoy this stuff as a hobby, where do I go next?” As if the people have a curriculum here. But my answer is yeah.

Tabitha: We're actually super, super lucky. We kind of do have a curriculum, if your flavor matches my flavor anyway, I cannot recommend anything more highly than the International Journal of Proof of Concept, or Get Out The [BLEEP] Out

Corey: Yes.

Tabitha: —which is a hackerzine. It's been running for several years now, and they do short, conversational articles on an astounding variety of really deep, delightful, and fun, reverse engineering projects.

Corey: And we're throwing a link to that puppy in the [00:21:26 show notes].

Tabitha: Yeah, yeah. One of my favorite past articles there was achieving “Reliable Code Execution on a Tamagotchi”. That's the kind of stuff that you get in POC or GTFO.

Corey: This is amazing. I like that better than my curriculum which was, always look at the thing in your stack of whatever it is you're running, and then magic happens and then that kicks off this other piece. It winds up on covering areas in which you're generally weak. One of my favorite interview questions just from a showing how candidates think perspective, but terrible from a pass or fail perspective is, “I type www.google.com into a browser and I press enter. Assuming that they have not yet deprecated google.com, what happens next to make that work?”

And you can go anywhere with that question: the network, the browser, the DOM discussion, the debouncers in the keyboard, electricity. You can talk about TCP handshakes, you can talk about Google's deprecation policy; it's hard to get me not to. And there's so many different ways to take that, that you can spend an entire lifetime on that. There's a GitHub repository—or GifHub repository—where someone documented everything that happens to make that work. And it's a collaborative effort, and it is enormous at this point. I’ll have to see if I can dig that out and put that in the [00:22:43 show notes] as well.

Tabitha: Yeah, that's one of those kinds of questions that you can use for good or evil. And as someone who would like to use it for good, it breaks my heart that it is so frequently used for evil because I would absolutely never ask someone that question in an interview, not because I am afraid of what I would do to them during that conversation, but because I am afraid of how it would make them feel and what the baggage a question like that will pull in from all of their previous bad experiences being abused by interviewers who just wanted to show off that they were smarter than somebody.

Corey: Yeah, I've just pulled it up, and I'll throw it into the [00:23:28 show notes]. Yeah, it has 291 commits, 69 contributors, 48 pending pull requests, 81 issues. Yeah, it's incredibly deep. This is why I don't like this in the form of a pass or fail question. Job interviews are awful; I will die on that particular hill.

Tabitha: Oh, man, I love them so much, and I absolutely agree with you. They're awful.

Corey: Yes, I love going through them—on either side of the table—because you have conversations; you get to see what other people are into, and that's great. But you're also evaluating people on a skill set that for almost every job, they're not going to need again until they interview for a different job somewhere else. I don't know about you, but I don't tend to write a whole hell of a lot of code on the whiteboard on purpose.

Tabitha: [laugh].

Corey: I don't need to implement quicksort. I don't need to invert binary trees or avoid link lists to wind up finding a cycle loop, or whatever it is that people are asking about these days.

Tabitha: Yeah.

Corey: It’s not interesting to me. It's fun to play the games of who knows more, but they’re games you play with coworkers and friends, not, “I am going to dictate the course of your career by how well you answer this question. But no pressure.”

Tabitha: Yeah. Yeah, yeah. And like, that's why I have taken so strongly to interviewing and why I enjoy the opportunity of interviewing so much because I've had the good fortune in my career, not to have very many of those awful sort of, “If you can jump through these 17 hoops in exactly the way that I prescribed, then you pass,” sort of interviews. I have had a far lower than usual number of them. And I'm grateful for that.

And so I'm also grateful then, for the opportunity to pay that forward by trying to have a good and interesting conversation with every single candidate that I interview. Sometimes we have a great conversation because, actually, they are an astoundingly good fit for the role and, hooray; other times, I might know two minutes in that this is very likely not going to be a yes from me, but for the next 58 minutes, or whatever, my time is that person’s, and I want to make sure that they have a chance to talk about the way they think about technology, the way they approach problem-solving, in whatever way is going to show what they're good at—as well as it can—and hopefully give them things to think about to improve themselves in the future, to improve their interviewing performance in the future, but also just to improve the way that they approach things in the future.

This episode is sponsored by our friends at New Relic. If you’re like most environments, you probably have an incredibly complicated architecture, which means that monitoring it is going to take a dozen different tools. And then we get into the advanced stuff. We all have been there and know that pain, or will learn it shortly, and New Relic wants to change that. They’ve designed everything you need in one platform with pricing that’s simple and straightforward, and that means no more counting hosts. You also can get one user and a hundred gigabytes a month, totally free. To learn more, visit newrelic.com. Observability made simple.

Corey: I want to jump in. I mean, before you get letters, and oh, will you get letters if we're not careful to clarify here. Like, “Well, if you know you’re not going to hire someone, why would you waste their time?” Well, first, if I could reliably say this person is ‘no hire’ in the first two minutes of meeting them, that would be incredible in so many different aspects of life. But I have my initial gut impression, and then I have the actual studied impression so I can back up the decision either way.

Tabitha: Mm-hm. Got to have reasons. Like, just what you feel isn't enough.

Corey: Yeah. And as other benefits. One, it gives this person exposure to the interview process, if that's one of the areas they're weak on; it also lets you see how this person thinks. And again, remember, it's not a, “No,” it's a, “Not right now, for this role.” And it's also a marketing opportunity.

The goal of every interview, in my experience has been that even if you turn down the candidate, or you make an offer and they reject the offer, they should walk away from the experience thinking, “That was such a great experience. I would love to try again for a better fitting role and/or recommend it to other people I know.” Because people remember this stuff. People talk about these things. I still have nightmare stories about awful interview processes. I went through the Google interview process twice, and after the second time, I swore that it was such a degrading experience I would not go through it a third time. And here we are.

Tabitha: [laugh]. Yeah, yeah, yeah. Or to put it the other way, there are organizations that I've had really great interviewing experiences with, and for whatever reason—my reasons or their reasons—did not end up working there, but I met people through the process of interviewing there, and I learned about the organization and I developed a lot of respect for both the people and the organization. And I keep in touch with some of those people, and I send them referrals. I try to get people to work at Datadog because I'm happy here and I want to have more people around that are the sort that I want to work with. But it's not just purely mercenary. I don't actually refer everyone that I know to Datadog, but I do refer people to the places where I think that they'll be happy.

Corey: I will refer people like crazy. It doesn't take much to do it. And introductions always help. But I also am always clear to draw a line between referring someone and recommending someone. That's a much shorter list. And there's a difference between, “You should talk to such and such.” Versus, “If you don't hire this person, I would very much like to know why because one of us is very wrong about something and I don't think it's me.”

Tabitha: [laugh]. That's a great point. And like—

Corey: Yeah, ‘the shut up and hire them’ list versus the ‘some yahoo found their way into my inbox and asked if I could introduce them to someone at your company.’ Yeah, I'm thrilled to do that. And again, if you're listening to this show and there's someone that I know at a company that you would love to work at and want an introduction, please reach out; that's what I do. I enjoy doing that and it helps short circuit a lot of the HR application process that screens on things that, frankly, are not germane to your ability to do the job. I am thrilled to introduce people to one another. But there is a difference between that and, “I have worked with this person in the past; I cannot say enough good things about their work product, and they're a joy and treasure to work with.” Those are two different things. And I have no problem giving one of those recommendations to anyone, and I love when I find someone I can give the other one to.

Tabitha: I am super grateful for you bringing that up. Because, A) I think it's a good thing to be saying here on air, but also, B) that's a concept that I really needed in my life because I did not have that concept before. And so when I have done referrals, I have only done them in the, “I highly recommend this individual based on my personal experience and opinion,” sense. But you make a really good point about how introducing people to each other is not expensive, and can help people in ways that you don't understand.

Corey: And that's part of it, too is, again, don't view life in a transactional manner. My question I always like to ask people after I have a cup of coffee or lunch with him is what are you up to and how can I help? Who can I introduce you to? It takes basically no effort from me, but it has the potential to change the trajectory of someone's life. Our lives are built on relationships.

And that's why I love sitting down and doing the interview conversations. I mean, here at The Duckbill Group, we go to extensive lengths to make sure that our interview process is not, “All right, we're going to find the thing you're weakest at and beat you up on that.” Like, what the hell is that? Are you trying to hire for strengths, or absence of any discernible weakness? The latter leads to a really weird culture.

Tabitha: Yeah, a culture where everybody's afraid they're going to get stabbed in the back by somebody else. Like, that's not actually a high-performing culture.

Corey: Oh, yeah. I mean, if you're talking to me for a technical interview process, and you ask how good I am with database. I'm like, “Oh, I'm great with DNS, thank you for asking.” That's probably a warning sign that I'm not a great fit to bring in when your DBA has a question. However, there are a lot of other things that I am particularly skilled at. It comes down to finding where someone is strong, where someone is weak, and looking holistically at the team and start to build out a highly functional team in the areas you need to have this. We could do a whole show on this sort of thing.

Tabitha: Oh my god, yes.

Corey: I'm just annoyed to the point of ranting now, the culture of hazing that is technical interviewing in an awful lot of shops. Because here's the hell of it: they've done studies on this; they’ve found what works and what doesn't at large scale for hiring effectively, and there's this entire Silicon Valley culture of, “We're good at programming, which probably means we're good at everything else, too, so we're going to discount what the quote-unquote ‘experts’ say and reinvent the job interview from first principles.” Do you hear yourselves? Stop it.

Tabitha: Oh, my goodness, a hundred percent. I always have to be on the lookout for that kind of stuff because my educational background is math and physics, and so I have to learn things from first principles. And if I don't, it never really sits well with me; I never really feel like I get it. But also, that comes with it a need to be quite clear on the distinction between, “I have experience in this area,” versus, “I am speculating from first principles. You should take this for what it's worth,” which might be something but it might be nothing at all.

Corey: Yeah. It's just this entire, weird, horrifying culture that we live in. The thing is, is I find you can tell an awful lot about companies by how they buy their people. And that is something that I think companies lose sight of is: candidates talk, and if you think otherwise, you're about to be sadly disabused of that notion.

Tabitha: Oh my gosh, especially as you get into more specialized areas. Sometimes it feels like there are ten Kubernetes security people and we all know each other. It's not really like that, but it can feel that way sometimes.

Corey: Yeah. I feel like we could do a whole ‘nother episode one of these days, about the joys, trials, and tribulations of technical interviewing. The problem is, I'm not the best person to have that conversation anymore because, at some point, I went through a transition from, “I have to get the job to put food on the table.” Which, respect. I get it. I have been there. I have nothing but respect for people in that position.

And it switched to, “Okay, now that I have options in my career, what's the right company and what's the right fit?” And when I stopped pretending to be something I wasn't and being much more myself, I found that many job interviews were far shorter as a result, but because I was sussing out places I didn't want to work. It's one of those things where the old joke of, “What's your biggest weakness?” “Honesty.” “Well, I don't think honesty is a weakness.” “Well, I don't give a [BLEEP] what you think.”

It comes down to not being a great thing for getting yourself hired, but it's filtering out the places you don't want to work. So, you've just completed the interview process yourself; you're working at Datadog. The fact that you will appear on this show and admit to working at Datadog implies that it's not a terrible place to be. Having gone through that process more recently than I have, any tips you have for anyone who's listening as far as what things to look for in an employer, red flags, things that really make you take a step back on the other hand, and say, “This is a place I want to be.” And, again, if you're listening to this and hiring, pay attention, this is also for you.

Tabitha: Oh, that's also a good one. For me, I was getting a lot from reading between the lines of the interview process. We were just talking about interview process, and trying to have a good and interesting conversation for both parties even if it didn't seem that the interview was necessarily leading towards a hire, and the more that I felt like my time was being well-used in the interview process, the more that I felt like I was enjoying meeting someone, maybe learning something, having an interesting conversation that seemed also to correlate with places where people liked their jobs, where people had positive relationships with their coworkers, where people helped each other out instead of backstabbing each other. And to me, that was a really important selector for where I wanted to work. I want to solve fun tech problems, I want to solve fun people problems with others who have that same goal. I don't want to waste my time on watching my back so it doesn't get stabbed.

Corey: Exactly. People have so much energy, and they can either focus on moving things forward, or they can focus on covering their own ass and doing everything the most defensible way possible. That's a gross oversimplification. I mean, depending on the company, sometimes downside risk management is more important than speeding time to market or getting code out. It turns out most banks don't have a culture of, “I wrote this code last night. Let's YOLO it into production, it's probably fine.” But in most cases, people will optimize for what you reward, culturally.

Tabitha: One hundred percent. For better and for worse.

Corey: I have been in multiple job interviews in my life as a candidate where I wound up successfully recruiting the interviewer because you start talking about work-life balance, and how the rest works, and at some point, they become too honest. And… “Yeah, this isn't the kind of place where you'd be happy,” “Yeah. That's kind of why I am where I am now.” “Are you folks hiring?” “Well, we could be.” And suddenly, we're having a different conversation.

Tabitha: You know, that is an absolute delight. And it reminds me of an interviewing story that was one of the best interviewing experiences that I never had. A friend of mine was a manager, managing a pen testing team, and at that time I was looking for a job, and I was potentially interested in it. So, I didn’t, like, apply, but I just asked her, “Hey, tell me a little bit about this.” And we had the conversation. And eventually, she told me, “I think that you could do this; I think you could be really successful at this, but honestly, I am afraid that you would get bored. You probably don't want to do this.” And that was the truest expression of friendship. And I'm still super grateful to her.

Corey: There's something to be said for if you're going through the process of interviewing and you don't know how to handle things. For God's sake, reach out to people. You're not alone in this, I promise. And find people who've done well in their career and ask them what their tricks are, ask what their secrets are, ask what they wish they'd known. People like the ambition in most cases, and people love giving advice. Just remember, everyone has their own opinion, and not all of them are great.

Tabitha: Yeah, yeah. You got to try to understand what is the context that has led to this opinion because sometimes you can learn from good advice, and in good circumstances, you can also learn from bad advice, as long as you recognize that it's bad advice and read between the lines.

Corey: Absolutely. So, thank you so much for taking the time to speak with me today. If people want to learn more, where can they find you?

Tabitha: Easiest place to find me is on Twitter. My Twitter handle is @tabbysable. Otherwise, you know, I'm around on the internet. I'm on a lot of the same big industry Slacks that you are, and those sorts of things. But I love to get messages on Twitter.

Corey: Oh, it's my favorite thing because then I feel so good about reading them and then forgetting to respond.

Tabitha: Oh—

Corey: Don’t take it personally.

Tabitha: —my gosh, the management of the DM inbox is such a trash fire, and I'm just lucky that I get a manageable number of them.

Corey: Exactly. Thank you once again. I appreciate your taking as much time with me as you have.

Tabitha: It's been great. It's been so much fun. Thank you so much for having me.

Corey: Tabitha Sable, systems security engineer at Datadog. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you've hated this podcast, please leave a five-star review on your podcast platform of choice, along with a comment telling me what happens when I type google.com into my browser and press enter.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Scott Piper
Scott is an independent consultant helping companies secure their AWS environments through private trainings. He created the free training sites flaws.cloud and flaws2.cloud, along with the open-source projects CloudMapper, Parliament, and more.

Links Referenced:

  • Connect with Scott Piper on...
    • LinkedIn
    • Twitter: @0xdabbad00
  • Company website: Summit Route
  • flaws.cloud
  • flaws2.cloud
  • fwd:cloudsec

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

This episode is sponsored by our friends at New Relic. If you’re like most environments, you probably have an incredibly complicated architecture, which means that monitoring it is going to take a dozen different tools. And then we get into the advanced stuff. We all have been there and know that pain, or will learn it shortly, and New Relic wants to change that. They’ve designed everything you need in one platform with pricing that’s simple and straightforward, and that means no more counting hosts. You also can get one user and a hundred gigabytes a month, totally free. To learn more, visit newrelic.com. Observability made simple.

Corey: When you think about feature flags—and you should—you should also be thinking of LaunchDarkly. LaunchDarkly is a feature management platform that lets all your teams safely deliver and control software through feature flags. By separating code deployments from feature releases at massive scale—and small scale, too—LaunchDarkly enables you to innovate faster, increase developer happiness—which is more important than you’d think—and drive transformation throughout your organization. LaunchDarkly enables teams to modernize faster. Awesome companies have used them, large, small, and everything in between. Take a look at launchdarkly.com, and tell them that I sent you. My thanks again for their sponsorship of this episode.

Corey: Welcome to Screaming in the Cloud, I'm Corey Quinn. I'm generally the person that people think of when there's an AWS billing problem, but when there's an AWS security problem, the one person I think of before anyone else is AWS Security Consultant Scott Piper. Scott, welcome to the show.

Scott: Thanks for having me, Corey.

Corey: It's been fascinating, just sort of, I guess, passing like ships in the night for the last, well, three and a half, four years or so. You're an independent consultant, a one-man band—like I was for the first two years before I had the good sense to hire someone whose primary language is spreadsheets—but it's been really interesting seeing you grow and evolve. And, honestly, you have actual expertise in the whole security space, whereas with billing, I mostly faked it for a while.

Scott: Yeah. And I faked it myself for a while because I did not come in with strong AWS experience at all. I basically was at a previous job trying to wear a lot of hats. I was the sole security person at a startup, and as a result, I was doing not only our CloudSec but also our AppSec—or CorpSec—our physical security, badge readers, surveillance cameras, just every aspect of our security in different ways. And doing all of it poorly, and especially on the cloud security side; that was an area that I felt I was very weak and I didn't have a lot of experience in, I didn't really know what my concern should be in there.

And so I started to really just try to understand what are the past incidents that have happened in AWS? And what are the things that I want to make sure that our DevOps guy is aware of when he's trying to build out our AWS infrastructure? And as, kind of, a challenge to myself, I figured, “Hey, I'll turn this into kind of a training program, kind of a game, that I can make online and available to everybody all the time.” And ended up releasing that as flaws.cloud. And so that is available even today. And we're not—

Corey: Yeah, I’m going to just put a plug in right there for that. flaws.cloud is one of the foundational learn-as-you-go exploration stories for learning by doing. It's basically an adventure game is probably the best way I can think of that, where there's an escalating series of open S3 buckets and things like that, where you go from level to level. It's what seven levels or something like that?

Scott: I think five levels for that one. And then I ended up creating flAWS2.cloud afterwards. And flaws2.cloud, I created a number of years later, but the initial flaws.cloud, I released it, and I figured if I release it on the right day on Twitter, maybe a dozen people will come and check it out.

And instead in the first month, 30,000 people visited the site. And so I was just floored. “Like, oh my gosh, there's something here.” Like I didn't know very much about AWS security, and apparently, a lot of other people out there are interested but don't know very much either, and are trying to learn about it. And so—

Corey: Oh, I’ll take a step further than that. The reason I think that what you've built is so compelling is it's real world. It shows what is going on and how this is supposed to work. But if you look across the entire landscape, every security story out there, it’s first pushed by a vendor trying to sell a product of some sort. And two, it's boring as hell. It’s, “Sit down and learn how this whole nonsense works for this fixed period of time.” I have to ask though, given the flaws.cloud and flaws2.cloud are both operating intentionally vulnerable environments, how many freaked out phone calls or emails have you gotten from the AWS security folks over the years for this?

Scott: Yeah, so I have received a number of emails, especially when, as part of flaws.cloud, I give people access to an AWS access key in my environment. And so, that access key has found its way onto a number of GitHub repos over the years, of people creating their own little test utilities, and they needed an access key, and they didn't want to use one in their own environment, so they just grabbed mine and put it in a GitHub repo somewhere. So, as a result, I get a number of emails along those lines. Eventually, AWS told me that I had to change the access key. And I told them, “No, I'm not going to do it. I've changed it a bunch of times already. This is just annoying. I know that it's not a security issue.” And so eventually, they somehow put some type of flag on my account to say, like, this guy is just a hassle to deal with. Just stop reaching out to him all the time.

Corey: I suspect that I have that flag, too. I'm sure it's something obscene as far as the naming convention, there goes. Yeah, that's my approach, too, very often. When they'll release something, like the original version of API gateway that is so convoluted to configure. I'll just bound an access key to that service, throw it in the internet, and then see what attackers do with it. “Oh, that's how it gets configured. Awesome.” I'm mostly kidding, but also not entirely.

Scott: Yep. [laugh].

Corey: Sometimes you learn by watching people break or misuse something in fascinating ways. I want to also highlight that it seems like you have aspects of the same problem that I do, and before you take that as a deadly insult, let me be more specific. It feels like if I ask you for the elevator pitch of, “What is it, you do exactly?” You've got to sabotage the elevator because it's not just the independent security stuff; it's not just the tool stuff; you're also the creator behind fwd:cloudsec. What is that?

Scott: So, fwd:cloudsec came about as a result of basically a number of security researchers and just other security folks had attended AWS’s re:Inforce, which was their security conference that they started in 2019. And we attended that, and there were some, kind of, frustrations that we had with it. And so specifically, we recognize that a lot of people are running on multiple clouds, whether they want to or not, whether they know it or not, they are running on many cloud environments. And obviously, AWS’s security conference is only going to be about AWS. So, that was one aspect.

Another aspect is that AWS is going to add their conferences because it's their conference, they're going to control the message and make sure that AWS and the features, the services that they release, the ones that exist, are always viewed in the best light; they don't want to talk about the limitations so much. And that really, as practitioners, is something that you're most interested in: does this work in all regions? Does it have these various integrations? Does it have CloudFormation support? All those different aspects of it.

And so we wanted to make sure that we had a conference or a platform where we could talk about a lot of those things. So, there's that. We wanted to be able to talk about attacks because we didn't want to just talk about, “Hey, here are some features on AWS that you can use to prevent security misconfigurations.” We wanted to dive into, what are those misconfigurations? What are attackers doing?

What beyond just a technical fix is something that people could try to use to mitigate these issues? So, we wanted to dive into all of that. And then it finally was AWS conferences are very large and trying to meet people at these conferences can take a long time to try and walk across from one end of a conference center to another. And so we wanted a smaller place where we could have a lot of the hallway conversations to talk about things. And so as a result of all of that, we ended up creating fwd:cloudsec to basically become a cloud security conference for practitioners that focused on all clouds and was able to dive into the limitations, the attack research, the other types of defenses that can be used, all those different types of things.

And then on top of that, a number of the organizers, they really wanted to make this have benefits for the greater community as well. And so it's a nonprofit. And so as a result of that, when—in 2020, we originally planned on having an in-person conference in June. Obviously, that did not work out due to this pandemic that has happened. But when we were going to have that in-person conference, we had planned on having basically scholarships for college students that couldn't otherwise afford to attend some of these conferences.

Corey: If there's one thing I'm taking away from the pandemic, it’s absolutely that it is such a better experience when you are not limiting these conferences to the folks who can get a week off of work, and travel to Las Vegas, and put themselves up there, and pay the $2,000 ticket fee. There's just so much else that it could be. And I am a huge fan of just that entire model. I'm also a huge fan of, by the way, of you following AWS’s wonderful footsteps when it comes to naming things, and naming fwd:cloudsec after an email subject line.

Scott: Yep. That was what we decided to do with the name is we were throwing around a whole bunch of different names. And we're like, “You know what? Let's just make fun of AWS with our name.” And so, with AWS using ‘re’ for everything—so re:Inforce, re:Invent, re:MARS, we decided to use the other subject header of forward—F-W-D colon.

Corey: Yeah, I am really looking forward to next year’s. The challenge was I believe the first year of this conference was co-located with re:Inforce, wasn't it?

Scott: Yeah. So, re:Inforce was supposed to have been in Houston this past year. And, you know, it was canceled. And so we decided, though, to continue to move forward with things. And so, we had it as a virtual conference.

And we originally had it planned to be in Houston; it was going to be the day before re:Inforce. And so we're currently trying to figure out what to do for next year. Because currently the next re:Inforce has not been announced yet. We don't know what city it will be in, what date or anything like that, but we have made a couple of decisions upfront, one of them being is that we do want it to be streamed live the day of the conference because we recognize that, again, going back to your earlier point, that one of the benefits of the virtual conferences is that anybody could have access to that conference in some way. And we don't know how vaccines are going to play out, we don't know whether or not we'll be able to have an entirely in-person conference. Our hope is that we will, but again, there's a lot of unknowns there.

Corey: It’s also the last sort of thing that's going to come back. It's all right, I'm going to take a risk now and go out to a restaurant or get my haircut, but travel to a different city for, basically, to sit in a vendor expo for two days and wind up effectively sharing air with 20,000 people, it feels like on some level, it's like, so which of our staff are we sending there? Oh, the expendable ones, of course.

Scott: [laugh]. Yeah. And so even if best-case scenario we're able to have an in-person conference, we still recognize that the people that are able to attend that conference are going to be people from countries that have access to the vaccine—not all countries do—and people that probably can make a plane purchase within a short timeframe, based on that decision making, and so potentially be spending more money on that flight. And so as a result of all that, we recognize that we want to make sure that fwd:cloudsec is accessible to people all around the world, no matter what their current economic or—situation is with whether or not they have access to the vaccine and everything. So, that is one of the decisions we have made for it. But a lot of the things are still up in the air for it.

Corey: One thing I've really come around to is the idea that with online conferences, I love the idea of live streaming the talk, but I feel like those talks should generally be pre-recorded. Whenever you do them live, it feels like you're, one, taunting the demo gods, which never goes well. But what I've also really enjoyed is participating in the live chat Q&A as a part of whatever conference program you're using, and answering questions on the fly as you go. Or if you're me, slash psychotic, live-tweeting your own talk.

Scott: Yeah, and I've seen some amazing conference talks this year, that had been pre-recorded, that had been professionally edited. They had cutscenes back and forth between demos of things and actual physical demonstrations of things as well. So, yeah, that all are things that we're still trying to figure out how we're going to make this make sense in some way. Especially given that there's so many unknowns as to how this next year is going to play out.

Corey: Oh, absolutely. And again, I think that people are extraordinarily patient when it comes to these sorts of things. Do you have any idea when the call for papers is going to open?

Scott: So, we still haven't settled on a date for the conference, specifically, but we expect probably in the next maybe two months.

Corey: For those listening, we are recording in the very last days of the wonderful year 2020. So, yeah, it's always interesting when people listen to these things, and it's a point in time, and sometimes I embarrass myself. Wow, you recorded that episode six months ago. What's the deal with that? And the answer is, generally, legal review. But I digress.

Scott: So, probably February of 2020. But again, we still may change that, and we're still trying to make some decisions on things.

Corey: Absolutely. I do want to get, as a security expert in all things AWS—largely self-appointed, but again, it’s not like there’s a certification board for these things and I've seen enough of your work to say that I unreservedly trust you. When you tell me something is true in the world of security. I take that at face value because I've yet to see you proven wrong.

Scott: Thank you. [laugh].

Corey: My position on this—talk about saying controversial things, get people in trouble down the road when this is played back, probably by a client for you down the road—but my take on the shared responsibility model, which is AWS's overly complicated way of saying, “Here's what the cloud provider worries about versus here's what the customer is responsible for,” is basically an overwrought song-and-dance because the answer actually fits in a tweet, but if you sit there to someone who's just suffered a breach, and tell them the truth of it, which is, “If you get breached in the Cloud, it is almost certainly your fault.”

Scott: Yep.

Corey: “You messed up.” That's all it says because the breaches are not people driving trucks into data centers, grabbing racks into the back, and peeling out. Its misconfigurations of S3 buckets. It's oh, it turns out ‘kitty’ was a terrible password. It turns out that the OSP10 haven't really changed the list of top ten security vulnerabilities in web apps in the past decade because people still aren't sanitizing their inputs, or cross-site scripting; what does that mean exactly?

It's always the same old stuff and there's nothing new under the sun. But it's not oh, the cloud provider forgot to wipe the disk volume after you were done with it and present it to someone else. They have those operational aspects down to a science.

Scott: Yeah, so the shared responsibility model really comes down to anything that you can secure is your responsibility to secure. Any type of configuration change that you can make is your responsibility to make that. And so with the shared responsibility model, though, the confusion, the frustration comes down to there are some things that AWS does very, very well: they are able to operate this amazing cloud infrastructure that very rarely goes down. It has had some definite hiccups, but it tends to stay up, it tends to be able to scale fairly well, they have backwards compatibility, there's a number of things that they do well. But then there's some things that they don't do well, such as having good user interfaces, for example, being able to better understand your cloud environment in different ways.

And as a result of those limitations, that, I think, is where a lot of these security issues come into play is that people don't understand their cloud environments as well as they would like to, and AWS is not really helping people to understand their environments that well. And so as a result of that, that is where I think a lot of the misconfigurations come into play, there are some other aspects of this issue; for example, a number of their security services don't have as broad a coverage as we would like. So, for example—

Corey: Oh, I’ll take it a step further; a lot of the security services suck.

Scott: Ye—[laugh].

Corey: Okay, we're going to put all this stuff into CloudTrail logs, but no one ever reads them, so we're going to consume them with GuardDuty. Oh, and that's super noisy too, so we're going to build Detective on top of that. Now, they're charging you all the way to go up this ladder, and at some level, you look at all the security offerings that they have—and I've looked at some of the big consultancy security architectures for all this stuff—and I'm looking at this because I focus in the world of billing. And I am almost certain that a data breach would be less expensive than running these services.

Scott: [laugh]. Yeah, it can get pretty crazy. And there are these various aspects of their services that really are AWS’s own responsibility to improve them. So, you brought up CloudTrail, for example; there are a number of API calls on AWS that are not recorded in CloudTrail anywhere. So, a number of these are going to be your data-related calls.

And so there are configuration changes that you can make to CloudTrail to be able to see S3 object axises, but for example, CloudWatch put metric data: that call is not recorded anywhere, you have no ability to see that call anywhere. So, as a result of that, AWS's guidance on implementing a least privilege strategy becomes difficult to implement because one way of accomplishing that is to look at your historical access and basically remove privileges that have not been used. But because a number of those actions are not recorded anywhere, you do not have the ability to know whether or not those privileges have actually been used. Furthermore, you can start using some kind of more advanced concepts, like client-side monitoring, which is where you can flip some environmental variables in your application and it will result in you getting a local log of some of these actions that are made. However, within that recording of events, it does not record what resources were accessed, it only records what API calls were made.

And so as a result, if you were to leverage client-side monitoring to basically try and implement a least privileged strategy, you would not be able to restrict down to specific resources or apply certain conditions because you could only restrict down to the specific actions that are made. So, yeah, there's a number of limitations that I do think AWS still needs to improve on things themselves.

Corey: I would absolutely agree with that. The problem is, I feel like security and cost are spiritually aligned, insofar as people really care about them only after they didn't care about them and now they have egg on their face. I'm fortunate on my side of the world where it's just, “Well, we spent a little bit too much money that particular month and now we feel bad.” As opposed to security where it's, “So, what is your primary means of breach detection?” And the answer honestly is, “The front page of the New York Times.”

Scott: Yep, that or their AWS bill because there are a number of attacks on AWS that result in massive bills. The most common one is just going to be cryptocurrency mining. If you put an access key up on GitHub, the first thing that's going to happen—well, there's a race that happens. Basically, can AWS alert you about this problem, and can you take action on it before something bad happens, or the bad thing that'll happen is there's a number of bots that are continuously monitoring GitHub, and they will find that access key and they're going to spin up EC2 instances in your environment to do cryptocurrency mining. But beyond that, there's also the concept of denials—

Corey: [00:20:42 crosstalk] the only thing you have for a complete inventory. Periodically, I run into scenarios with smaller companies where, “Okay, so tell me about those instances in Australia?” “Oh, we don't have anything in that region.” I believe you are being sincere when you say this. However, somewhat paradoxically, above a certain point of scale, you can't really notice those breaches anymore via the bill. I mean, if you're spending, I don't know, $18 million a month on AWS, that's an awful lot of Bitcoin you have to mine.

Scott: Yep. [laugh].

Corey: It just disappears into the rest of the noise.

Scott: And so yeah, so trying to create billing alarms for things, that works when you have a free tier account, to some degree; it is not going to work when you are actively spending a lot of money otherwise. So, doing that type of monitoring is difficult. But I want to touch on, though, a little bit is the concept of, like, denial of wallet attacks are a concept that really didn't occur in the data center world, but now is suddenly an opportunity for attackers in the Cloud world. And what that really means is that if you were someone that didn't like a company for some reason, previously, you could have DDoS’ed them; you could have basically tried to send a lot of bandwidth over to that company in some way to shut down their servers because they're not able to keep up with the amount of traffic that you're sending to them. But in AWS, and in the Cloud world, and across the cloud providers, you now have the ability to basically increase that customer’s—or that company's—amount of spend on AWS.

And so if you are able to get access to an access key to spin up some of these resources or to start making some of these very expensive AWS calls. So, start reserving instances for the next three years that are SQL Server licensed or some other type of licensing option on them, and suddenly, yeah, you can actually burn—I think like, there's a single AWS call you can make that will cost a company $64 million, just because you're spinning up a whole bunch of resources all at once, they're all licensed resources.

Corey: Single API call? That much? I think that's right around the cap of the default console limits for maxing out savings plans.

Scott: [laugh]. Yeah, there's a lot of these different opportunities. Or there's the possibility of using the AWS marketplace in order to purchase something from one of the vendors there, or there's the opportunities to commit various types of white-collar crimes if you were to create your own ABS marketplace offering, and then from your company that you work for, during the day to make a purchase of that, and now suddenly you're able to make this side cash that the company probably isn't going to be very aware of, just because you are purchasing something from a vendor who happens to be yourself that you're moonlighting as. So, there's all those types of things that can potentially happen on AWS. And all the cloud providers.

Corey: Oh, absolutely. I don't think that AWS is particularly vulnerable. If anything, I would say their security posture is I would argue the best of all the major cloud providers, I know that GCP would argue that point strenuously. Where do you stand?

Scott: So, I am a big advocate at AWS. On Twitter, I will make fun of them all day; I will call them out on every single minor mistake that they make—

Corey: That’s why we get along so well.

Scott: [laugh]. But at the end of the day, I advocate to everybody to use AWS. I use it for all of my personal things, everything in my life is backed up on AWS in some way. I trust AWS. And if you look at it, there are a number of government agencies and different ways running on AWS.

Some of them have that special, fancy isolated partition that is not connected to the internet and is used for classified information, but, like, AWS and Amazon, they are able to secure things well. And part of that is just because of the economies of scale. They are making tons and tons of money, or receiving that money from customers, and as such, they can now afford to have their own DDoS response teams; they can have secure enclaves—you know, their Nitro Enclaves and all sorts of different features. Their automated reasoning, for example, like, these are things that you cannot do. And so I do recommend people use AWS, even though I do give them a hard time.

Corey: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the Enterprise (not the starship). On-prem security doesn’t translate well to cloud or multi-cloud environments, and that’s not even counting IoT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IoT devices, detects these threats up to 35 percent faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at extrahop.com/trial.

Corey: I agree with you wholeheartedly. I think that you might have a more fraught relationship dynamic, just because the best approach from the cloud provider perspective when someone discovers a security issue is, “Cool. Could you tell us very quietly? We will never confirm or deny anything there. We will quietly fix it, and you will go away forever.”

Which is kind of not how you build a reputation for excellence in this space. So, there's that constant tension. I gave a talk at re:Invent last year—two years ago now, I guess, since this is going to be 2021 when people are listening—about the vulnerability disclosure program that AWS runs and how they do it. It was a fantastic story based upon some stuff that we collaboratively found. And the thing that surprised me was, every time I find other stuff in the billing area or, this doesn't make sense, they're super friendly and thrilled for me to bring that to their attention. They try not to talk much about anything that's even vaguely security-related, if they can help it.

Scott: Yeah. And for those types of issues, I mean, it makes sense then, but at the same time, like, for me, it can be frustrating at times. And I see that from the community, as well. I mean, we are each in a position where I think a lot of people send us DMs about a lot of different things. We get a lot of emails from people privately asking us, telling us different things.

And one of those things that sometimes I get messages from people about is about security issues with AWS that they will report to me prior to reporting it to AWS, in the expectation that I have some ability to fix it or something like that. Or they just want my advice; they're concerned because there are a number of vendors out there that take legal action against security researchers. And so they are concerned about how that's going to play out if they communicate these issues to AWS. And so I will say that AWS does not—in my experience, and I've never heard of them doing this, of taking legal action against people. Obviously, that's going to depend on how they’re—

Corey: No, they tend to reserve that for their employees as best I can tell.

Scott: Yeah, [laugh] yeah. I mean, it's going to depend on how you're finding these issues, and where you're finding them, and what you're doing. There is an extent to the leniency of what they can do. But yeah, I mean, I think that AWS is very good about interacting with researchers, and that's especially been true because of Zack Glick, specifically, over there at AWS.

Corey: Oh, he's amazing. He was my co-presenter.

Scott: Yeah. And he is kind of the liaison between security researchers and AWS. And so he, a couple years ago, ended up moving into that position of becoming that liaison, specifically because AWS was falling flat on responding to security issues. And so specifically, I had found, for example, some issues with AWS managed policies that were being—they were documented by AWS and advocated for customers to use, and there were a number of flaws in them that were not making those policies work in the way that was expected. And I tried reporting that to AWS repeatedly, trying to get their attention; couldn't get their attention, finally was lucky enough to be working with a client at the time that was spending enough money with AWS that they basically told AWS, “Hey, you need to respond to these issues. These are actual problems, and we're concerned about your security posture if you're not resolving these issues.” And so as a result of that, AWS now does respond much better to security researchers.

Corey: They do. For a while, they had specific pen testing requirements or disclosures for port scanning and the rest, and they've really loosened that, which is nice. There's something to be said for not getting in the way.

Scott: Yep. And those—oh, those pen testing requirements were—they were impossible to try and follow in different ways, just because you were supposed to tell them, which EC2 instances were going to be pen tested, which isn't going to work if you have auto-scaling groups that are spinning up and spinning down instances all the time.

Corey: I'm going to be doing—oh, nevermind, they're gone. Just kidding, they're back.

Scott: Yeah, it was a mess. So, it is good that they removed some of those requirements. They still do have some pen testing requirements, and still, some services are supposed to be off-limits. However, I've never heard of anybody getting in trouble for doing pen tests in different ways against these services. I don’t—I’m not going to—

Corey: No, I imagine if anything, they wind up IP blocked or, at worst case, they get a phone call or an email of, “Would you mind knocking that off?”

Scott: I mean, I haven't heard of them doing any of those things. Just because again, like a pen test that you're performing on AWS is not going to be any worse than the just day-to-day operations of a lot of companies out there, and also just the aspect of a lot of companies—or AWS services that are publicly accessible and are constantly getting free pen tests that aren't being reported in a new way, just because there's attackers out there that are constantly port-scanning, constantly trying to brute force hack, or break into different services.

Corey: One question I have for you, I've been seeing all the announcements about it, and I feel like I’ve almost been gaslit into believing that I'm the one that's looking at this the wrong way, but they've made a strong push towards attribute-based access control, primarily tags, and that terrifies me because historically, for the last 15 years, everyone gets access to tag everything because that was the way—the only way you could reasonably do cost allocation. Suddenly, if you go down that path, everything with tagging permission across your entire estate now becomes a massive security vector. Am I wrong on that?

Scott: So, historically, the best, only security boundary on AWS was having separate AWS accounts. And people realize that people have these large monolithic accounts, and they want to try and segment access in some way. And so the concept of attribute-based access control came about to allow people to tag resources, and only allow certain services or certain developers to work on those resources that are tagged in a certain way. There's a number of limitations with this, though.

The biggest one that people ran into is just that a number of resources didn't support tagging, at all, anywhere. And so AWS has gotten better about that, however, there's still a number of resources that don't support tag on create. And so even though they may support tags, you cannot restrict what tags someone can use when they're creating that resource.

Corey: I think Elastic IP addresses was done during re:Invent this year, or last year, which is just what… what is that?

Scott: There's still a number of them that—it's just kind of mind-boggling that they are trying to push this concept, and yet they don't provide you with the ability to really utilize that concept, except for a number of limited use cases. So, as a result of that, the best security boundary still remains having separate AWS accounts on AWS. However, attribute-based access control, the other big limitation of it is the lack of tooling around it. So, if you want to understand who has access to a certain S3 bucket or other resource on AWS, it's difficult to try and figure that out. I mean, you can use tools like Access Analyzer to try and figure out, generally, which AWS accounts have access, is it public or not, but if you want to identify the specific IAM role and when those roles have different restrictions based on tags or other conditions and things like that, it becomes really complicated to try and figure that out.

So, I think that is one of the big issues is the lack of tooling, that it's hard to try and understand this, it's hard to try and audit those different policies that you may have in some way. So, that all becomes just a big frustration. So, I currently still do not advocate people attempt, really, to use attribute-based access control except for maybe a few limited situations, or unless that's all they can do because they have—

Corey: Or it’s greenfield, potentially. But what's truly greenfield in the world of identity and access control? It's a company that just started.

Scott: But even then, the best practice still is to have multi-account, to try and have different accounts for your different applications, and so trying to do things that way. But then again, you run into another problem is that as you end up with these large number of AWS accounts, people start connecting those AWS accounts in different ways. They create these trust relationships via IAM roles, via S3 bucket policies between your AWS accounts. And so your security boundaries start becoming blurred, or start basically erasing those security boundaries because now you can have two AWS accounts which really become equivalent if someone was to compromise one of those accounts because they can then assume into an admin role inside the other account if they were to compromise an account and obtain admin privileges in that account. So, again, that is another problem that we're starting to run into is how do you understand the relationships between your different accounts?

What are the trust connections that exist between them? And I think that is another area that, one, I would really like to see AWS do more in that area, but, two, I think that that's an opportunity for people to create different types of tooling around that.

Corey: One of these days, I want to just give them a giant wish list of things in security, just from a usability perspective. The idea of being able to throw IAM policies into warn-if-reject mode. In other words, the test account, let me give something basically admin rights, I have the Lambda function step through, it's code paths—which should not be that many—and at the end of it, everything it just did, that's the only thing I want you to ever be allowed to do, and it blocks everything else out. That would be amazing.

Scott: Yep, that type of thing has been on people's wish lists, and people have tried to use that concept of client-side monitoring to try to accomplish that. But again, there's those limitations of client-side monitoring that just don't allow it to work as effectively as you'd like.

Corey: I was playing around with the SAM CLI for a while, and by default, it will build Lambda functions, when you're talking to DynamoDB, with the ability to—you have full access to DynamoDB, every table in the account, every permission. And getting that narrowed down to something that is much more bounded to a particular table requires an awful lot of messing around with it, which tells me that most people aren't doing it.

Scott: Correct. So, what I do for my business is doing assessments for companies. And so as a result, I get to see all sorts of IAM policies, and I will one hundred percent agree with you that people have very open IAM policies that are not as restricted as this ideal utopian world that AWS tries to tell people could potentially exist.

Corey: So, looking at the sheer complexity around security, it feels like the easiest solution is to give up on some level, rather than attempting and obviously failing to get it right. Because let's not kid ourselves: if Capital One, which has its faults but they don't hire dumb, if they, with all the assets that they have to protect can fall victim to this, what chance do the rest of us have? And the sheer complexity of service offerings that are ignoring first-party, the third-party stuff across the board, with everyone trying to sell me something, how is there any hope?

Scott: Yeah. And I mean, that was like, one of the big concerns that came out of the Capital One breaches. They are known as being one of the best for AWS security. And so they have some open source tools; they have a number of people there that are highly respected. And yeah, that breach, unfortunately, happened to them.

So, the big thing, though, that you can do, I think, is to try and do that account separation because that does allow you to make some of the mistakes—or specifically with regards to least privilege—that if one of your accounts gets compromised completely, you know, someone has admin access in it, it still allows your other accounts do not catch on fire when that one account does. It separates that blast radius there. The other thing that I think is still really in its infancy is using SCPs—or Service Control Policies—to start better protecting things. And so for example, if you have an incident response role that you want to allow your security team to be able to assume into so they can remediate issues, or just investigate potential issues, you want to make sure that you have an SCP that protects that IAM role so that if your AWS account gets compromised—one of your accounts—and that attacker starts deleting the incident response role, it won't be able to do that because it would be protected by the SCP. The SCP cannot be bypassed, even by the root user of an AWS account.

And so that allows you, basically, to put in those guardrails, to put in those restrictions, so that not only can you stop someone from turning off CloudTrail, or GuardDuty, or some of those other security features, but you can also better protect, for example, if you are using a vendor to do auto-remediation or to do monitoring, to again, use an SCP to protect that vendor’s IAM role or whatever access that they have into that account so that an attacker cannot disrupt that monitoring from happening. So, I think that that's another powerful thing that people need to do. But along with that—this is another area where I think AWS is still weak—is that if you, as a legitimate user in an account, try to turn off CloudTrail or try to do something that is somehow protected by one of these guardrails, you do not have the ability to know whether or not that was stopped by your IAM policy, or via an SCP, or via IAM boundary, or a session policy. Or if you're trying to mess with an S3 bucket in some way, is that stopped by one of the resource policies that's on the S3 bucket? There's all these different places in which you can define privileges on AWS.

And you do not know what is stopping you. And furthermore, even if you do have access to all that IAM information, you're able to see all those things, as just a developer that is not an expert in IAM, you can have very convoluted policies put on yourself that are difficult to understand. And so having some mechanism, I think, for AWS to be able to tell you, “Hey, you are not able to create that EC2 because you did not tag it with the required tag,” or, “You did not specify that it should have this EC2 instance type,” or something like that. AWS is just going to tell you, “Access denied,” and they're not going to give you any further information. And so I think that is an area that I would really like to see AWS be able to, in some way, provide you with more information, whether that's in the CloudTrail event to be able to dig into that, or if maybe it's some additional privilege that a user has if you have the debug privilege or something like that, that it will tell you why you were denied some action that you were trying to take.

Corey: The why of these failures is the hardest part to work around. And then people go ahead and work around it temporarily, and over-provision things, and they'll fix it later. Yeah. One of the biggest lies we tell ourselves.

Scott: Which again is the reason why I think that having those separate AWS accounts is so important, to not be working inside a monolithic account, to not have your dev-test sandbox environments within your production account, to have those as separate AWS accounts. So, again having those different AWS accounts, and to start focusing on building up those guardrails via SCPs, or auto-remediations where SCPs are not possible, I think that those are two things that people can really start focusing on.

Corey: Yeah, that's a reasonable starting point. I mean, security is a destination. It's a spectrum. It's not a journey. Nothing is completely secure until it's been blasted to pieces. And it's just a question of where your risk tolerance lies.

I strongly believe you will be more secure in a public cloud provider that is not IBM than you will be in your on-premises data centers in almost every case. So, thank you so much for taking the time to go through all this with me. If people want to learn more about who you are and what you do, where can they find you?

Scott: So, summitroute.com is probably the main entrance point to try and figuring out who I am and how to contact me. I am active on Twitter with almost entirely just AWS security-related things. Unfortunately, my Twitter handle, I created it back in the day when I was interested in reverse engineering, so it is @0xdabbad00 written all in hex letters. So, I would recommend just going to summitroute.com and figuring things out from there. Or just searching for me on Twitter as Scott Piper.

Corey: Excellent. We will, of course, throw links to all of that in the [00:40:41 show notes]. Thanks so much for taking the time to speak with me today. I really appreciate it, as always.

Scott: Yeah, thank you.

Corey: Scott Piper, AWS security consultant, I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you've hated this podcast, please leave a five-star review on your podcast platform of choice along with a comment allowing access to a single DynamoDB table written without consulting the documentation.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

Links Referenced:

  • Gaia Platform
  • Hal’s Blog
  • Follow Hal on Twitter

TranscriptAnnouncer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode has been sponsored in part by our friends at Veeam. Are you tired of juggling the cost of AWS backups and recovery with your SLAs? Quit the circus act and check out Veeam. Their AWS backup and recovery solution is made to save you money—not that that’s the primary goal, mind you—while also protecting your data properly. They’re letting you protect 10 instances for free with no time limits, so test it out now. You can even find them on the AWS Marketplace at snark.cloud/backitup. Wait? Did I just endorse something on the AWS Marketplace? Wonder of wonders, I did. Look, you don’t care about backups, you care about restores, and despite the fact that multi-cloud is a dumb strategy, it’s also a realistic reality, so make sure that you’re backing up data from everywhere with a single unified point of view. Check them out at snark.cloud/backitup.

Corey: When you think about feature flags—and you should—you should also be thinking of LaunchDarkly. LaunchDarkly is a feature management platform that lets all your teams safely deliver and control software through feature flags. By separating code deployments from feature releases at massive scale—and small scale, too—LaunchDarkly enables you to innovate faster, increase developer happiness—which is more important than you’d think—and drive transformation throughout your organization. LaunchDarkly enables teams to modernize faster. Awesome companies have used them, large, small, and everything in between. Take a look at launchdarkly.com, and tell them that I sent you. My thanks again for their sponsorship of this episode.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Hal Berenson, who at the moment is the founder of Gaia Platform, which is an industrial-strength, low-code platform for developers. But you've done a lot more than that. Hal, welcome to the show.

Hal: Thank you, Corey, it's great to be here.

Corey: So, it's rare that I see a line that jumps out in a biography like the one that you sent over. “In your 45 years in the industry,” is how it starts, at which point that sort of boggles my mind. In the interest of full disclosure, I am not yet 40 years old, at least at the time of this recording. And it's phenomenal to me to imagine the idea of, first, staying in a particular industry that long, but also having a career that transcends my own lifetime. What's that like?

Hal: You know, it's one of those things that's gone really quickly. And I kind of get shocked when I think about it as being 45 years. And it's really kind of longer than that. I grew up in a computer family, my father was in IT from the late 40s. And so I've always been around computers. And it goes fast. And that's largely because I have so much fun at it.

Corey: And it's fun looking through, I guess, obviously, you're cherry-picking different bits and pieces of your career when you talk about these things because the full-on accounted list would almost undoubtedly take up most of the show. But things that you've done, for example, moving backwards before Gaia Platform where you founded it, you were the Vice President of RDS at AWS—which one of those feels like it should be expanded—you were a Distinguished Engineer and general manager at Microsoft, you were a senior consulting engineer at DEC, or D-E-C, depending on pronunciation. We're not starting that fight again. So, you've done an awful lot of stuff. You do technology and management consulting. You're a song-and-dance man who’s the—what is the acting term? Triple threat?

Hal: [laugh]. Actually, when I started, computers were a hobby, which of course, today is a pretty common thing. But I started out, really, in the 70s, and access to computing was relatively difficult. So, it was fun for me. And when I went to college, I wasn't thinking computers.

My first year I was thinking, microbiology or something, but computing just drew me in. And a lot of the things I'm best known for, like databases, are things I never intended to do, I've tried to get away from multiple times. It really wasn't until I went to AWS where I said, “Okay, I'm giving in on my desire to get away from database work,” and, you know, let it happen.

Corey: To be clear, you were at Amazon doing RDS… VP’ing—I guess we're going to call it that—from the end of 2014 through the end of 2017, which saw a lot of changes. That's when Aurora started coming out. That's when DMS became a thing, and from a personal note, that's when using RDS somehow transitioned so subtly, I didn't even know that it was happening at the time. It went from, “[laugh], that thing's for idiots. If you care about your databases don't use it,” to becoming an extremely viable answer for running large scale production database workloads.

And it was just the reliability story, the maintenance story on it, an awful lot of stuff changed. And it wasn't one feature or one announcement. The backups got better, replication stories got better, the encryption stories became radically more mature. And it was just at some point, you'd look up and realize, huh, this service has gone from, “I don't know,” to something super legitimate. What drove that? Am I imagining this?

Hal: Yeah, I don't think you're imagining it. And while I’d love to take credit for some of that, I can't take credit for a lot of it, either. Before I got there, it had gone through a lot of the maturing that everything at AWS did, or anywhere in the Cloud did, and some hard lessons were learned in terms of things like just how reliable you needed to be. Whereas people knew it theoretically before, by the time I got there, that was what was in the back of the mind: we can't have a big outage, we can't have a backup not work, et cetera, et cetera. So, it was already underway, and then my enterprise background led me to support that and push that direction, instead of trying to change it to something else.

Corey: That's sort of a microcosm of the entire Cloud story. It's the idea that if a cloud provider—any of them; I'm not singling out AWS on this—releases a feature today, it's probably not fully baked for all workloads. Surprise, surprise, every customer is different in some ways. But they also don't get worse over time; they tend to improve, sometimes in leaps, sometimes just by chop wood, carry water.

But over time, they wind up becoming much more acceptable for a wider variety of workloads. And it still sort of boggles my mind that the same underlying service is equally suitable for fast-moving startups who are just trying to get something out the door before they run out of runway next week. And these large enterprises whose primary goal is risk mitigation.

Hal: There's some amount of dichotomy there in terms of how you can do more at the low-end and do more at the high-end at the same time. But you don't have to degrade the low-end experience as you do the high-end, and a lot of people the low-end can benefit as you do this. For example, having more high availability options, well, if you're a startup that needs super-high availability, the fact that what drove somebody to implement all these availability features was a Fortune 500 company screaming for them before they'll adopt your service, doesn't matter, you did this great feature, and now the low-end could use it. One of the things I didn't like about the package software product world is you would take features, and you’d reserve them for a high-end, super-expensive edition, that only your enterprise customers could pay for.

Corey: “Oh, if you wanted it encrypted, that'll cost extra. If you want SAML Federation, that'll cost you. Oh, you want audit logs, [laugh] get out the checkbook.”

Hal: Exactly. And those things harm startups. And it goes to one of my frustrations with SQL Server after I left which was, I had envisioned that what we would do is every new version, we would take a few features that we had put in the enterprise version, and push them down to lower-end versions, to push it to Standard, for example. But people got so excited about the money you got out of enterprise that for a long time, they let Standard stagnate. It's only three or four years ago that they finally moved a lot of those features.

And it really harmed SQL Server growth. I mean, I know that's hard to see, in terms of how big SQL Server got, but things like MySQL would never have gotten traction if Microsoft had done just a few small things. Multi-platform certainly could have been one of them, but the other was pushing more functionality further down on the curve. And I'm not the only former SQL Server leader who sees it that way. I've talked to a number of them who are kind of like, yeah, this was a big miss on our part. But the package product world makes that very common. In the Cloud, it's a lot easier to find differentiation and ways of justifying doing work without making it inaccessible to even an individual starting a business.

Corey: Well, you can even see it in the way that companies talk about themselves, and the things that a startup needs and uses today are what enterprise is going to be using down the road. There are virtually no startups that describe themselves as, “We're like an enterprise.” But basically, every enterprise likes to do the recruiting pitch of, “We're like a startup. Just don't examine that statement too quickly.” And if you can get the things that enterprises value, which let's be very clear here, are important.

I'm not here to talk smack about large enterprises generally. But things like durability, things like being able to trace transactions, and who did what, and being able to trace accountability for every step of the, I guess, custodial supply chain of a piece of software is important. And enterprises require this from a risk mitigation perspective, whereas startups don't need it until suddenly, they very much do.

Hal: Yeah, and we saw that a lot. I mean, this is one of the great things of having all these different jobs is, of course, you learn very different things. So, in the case of going to AWS, I got to experience a lot of these startups who had become daily things. I mean, one day, I did look at my phone and all the apps I was using and realize, “Oh, my God, they all run on AWS and they're household names.” And as you just indicated, they had the same availability requirements.

Look, when you can't get your shared car ride, it's as big a problem as when you can't get $10 out of an ATM. And this goes through everything. People notice it, it makes the front page of the Wall Street Journal when the services are down. So, the startup world and seeing these people become big and go from—well, there's one that didn't have a DBA until they were way into being a household name. And then, of course, they start having the exact same requirements, as the Fortune 500 companies around their mission-critical apps. So, it's fun doing it in that world, and seeing startups grow to need the same things as enterprises. And it's fun in the Cloud, that you can deliver that.

Corey: One of the things that I find gets missed so much is this idea that companies are able to speak to each other and they're using the same language. I mean, one of the biggest pains that I have to deal with, as a small business dealing with AWS bills on a consultancy basis, is the onboarding vendor procurement forms, where we wind up having to go through the enterprise-grade process for what is our data protection strategy? That's super important. We have to get that right and our answers there are pretty solid.

Then they ask things like, “Do you have at least $10 million in commercial vehicle insurance?” For what? [laugh]. We don't have cars, or forklifts, or anything handled by the company on that, so it's just a big, “No, not so much.” It's this idea that the same intake process that enterprises use, that they would use for their caterer or the bus company they're having do a company-wide event don't necessarily apply to pure services-based, high-level-expertise-oriented consultant work.

So, it's this going back and forth and having those conversations. In my previous job, I wound up feeling this very acutely. I was employee number 41, at FutureAdvisor, a FinTech startup that handled—basically a robo-advisor. And three months in, we got acquired by BlackRock, a company that at the time managed something like $4 trillion in assets. There was a bit of a culture gap between those two experiences.

And watching the transformation that hit us at basically warp speed as a direct result of this was eye-opening for me, and forever changed how I view large-scale enterprises and the things that they care about.

Hal: Yeah, I have similar experiences from the independent consulting standpoint, which is every time I deal with a large company, I go through those same things. Somebody will want me to have, like, $25 million in IP liability protection, in case I somehow transfer somebody else's IP to them. And hey, I'm a single proprietor thing and that's a—first we’d have to find that insurance, and then it's really expensive, and I'm probably not going to do enough work for you to justify it. Whereas with a startup, it's like, you talk to the CEO, and the CEO goes, “Yep, can you start yesterday?” And don't go to all this vendor craziness and everything.

And so I have a view of different large companies, based on exactly that problem. And when I first did some consulting back to Microsoft, I was talking to my friend, and now co-founder, David Vaskevitch, who was Microsoft’s CTO at the time, and I said, “Oh, this is really a pain, becoming a vendor and whatever,” and he goes, “Yeah, you're in this interesting position: you're not quite at the point where Steve Ballmer—who was CEO at the time—will just tell them, ‘Make it happen.’ And kind of ignore all the policies. You're just below that level. But of course, you know everybody, and you know everything, and everybody's trying to make it happen fast, but they're not the person who can say blow out the policies.”

But doing the work for Microsoft, dealing with their vendor policies was eye-opening, having seen it from both sides. And I've done that with others. There's another large computer company, I did some work for where their vendor policy was insane, and it needed approval from people in multiple countries as to where the chain went. And I was supposed to fly to China to do the consulting for them, and I didn't have the approvals. And I just decided to go and hope everything got approved because if not, you know, it was going to be, like, tens of thousands of dollars out of pocket when you threw in airfare and hotels for weeks. Fortunately, the day I got there, I walked into the customer, and the approval just came through. But I mean, it took weeks to get through this vendor process. So, it does tend to cloud your view of big companies versus small ones.

Corey: Oh, it consistently surprises me that a lot of these process-driven companies that employees who, I need to write a justification document to spend $50 on a book, would be able to either spin up hundreds of thousands of dollars worth of infrastructure or, more commonly, call massive meetings where the actual cost of salaries for the duration of that meeting are astronomical. And that's never questioned. It's a really weird perspective, and I used to be very cynical toward that. Then I start looking at how these things play out and what drives these, and the more I look into things like this, the more sympathetic I become.

Hal: Yeah, the problem is, they never die, the policies never die or get revised away. That's mostly the case. And so, if you look at Microsoft's continued thing with the vendors, that all goes back to this 1990s lawsuit by contract test people, in which Microsoft lost the suit that had them declared to be employees. And so ever since then, it adds layer and layer on making sure that anybody there on a contract basis will never be considered an employee. And so it gets pretty messy for people they have been working there for—I forget what it is now, but, like, nine months, and then they have to not work there again for another six months, or nine months or something.

Corey: Oh, and I understand that. It's no joke when you misclassify employees. I hear you on that.

Hal: And even if they had work elsewhere, which would kind of automatically say you're a contractor because you're doing work for more than one person, you still can run into those things, but the proxies that get thrown on there and the proxies that get thrown on, yeah, for acquiring things like books because you go and discover what were the past abuses in that, for example. But those policies never kind of go away. At Amazon, you could make the policy go away, there actually was—you could go and write a narrative, you could go justify it, you could say, look, this is violating a leadership principle because you're actually wasting more money than this process is trying to fix. But still, that's a lot of work to go and do it. And so only occasionally with somebody actually tries that route.

This episode is sponsored by our friends at New Relic. If you’re like most environments, you probably have an incredibly complicated architecture, which means that monitoring it is going to take a dozen different tools. And then we get into the advanced stuff. We all have been there and know that pain, or will learn it shortly, and New Relic wants to change that. They’ve designed everything you need in one platform with pricing that’s simple and straightforward, and that means no more counting hosts. You also can get one user and a hundred gigabytes a month, totally free. To learn more, visit newrelic.com. Observability made simple.

Corey: Something I want to talk to you about is a form of diversity and inclusion because it's fun: two white guys talking about this is all well and good, but what I do want to talk about here is the idea that first, we talk about women, we talk about people of color, we talk about the very real struggles that they wind up facing, but an area that we don't talk about is age which, barring tragedy, is going to affect each and every one of us as we move through our career. Very often, I talk to people on this show who are at the beginning or toward the middle of their career, it's uncommon that I speak to someone who's been doing this for 45 years in the industry. What have you seen from that perspective as you've gotten inevitably older as time goes on? But do you find the way that people respond to you or the way that you're treated has shifted?

Hal: Interestingly, I haven't seen it as much as typically gets reported. And I experienced it more on the other end of the scale, actually. I quit college after one year and went into the industry, so I didn't have a degree and I was super young. And I very quickly was ending up in positions where people would have typically been as much as ten years older than me. And so very early on, I experienced a little bit of it.

It was countered by the fact that there was such a demand for expertise that wasn't there, and I had the expertise so people would, kind of, give me a chance. And as time went on, I never really noticed it. When I went to talk to Amazon, and I was already in my late 50s when I went to do that, I didn't think about, “Oh, I'm in my late 50s.” I thought about, you know, “I’ve mostly been retired for four years. What's Amazon going to think about that?” But I didn't really think too much about my age.

And it was kind of funny because I’d learned during the interviews that I would end up being, I think, the oldest VP in AWS when I joined. If not the oldest, pretty much tied for it. But I never really thought about it. And on the other side, I don't have a problem with millennials. I think millennials are great. I don't get this generational thing. In fact, I never have.

From when I was young and others were old, I’ve never gotten it. So, millennials, actually, that worked for me were shocked when I told them my age. I've never found it to be a real issue in my career. I have found other things more important to focus on: experience, expertise, willingness to listen, the ability to deal with people in various ways and at various levels. Those to me all were more important in every aspect of my career than this calendar age thing.

And the thing that hits me now is, yeah, I keep retiring; I've retired three times, I don't know if I'll ever claim to retire again. I may just claim to take a break because I come back into it. And it occurred to me relatively late in that, something my friend, [00:20:53 Jim Gray], had said years ago on my first retirement. He was like, “You're crazy. How could you ever retire? You should work basically, until the day you die,” which sadly, Jim did, having passed prematurely. And I realized, if you're enjoying it, it's true. Why would you ever retire? And people don't really want you to most of the time.

Corey: Well, something I do want to highlight is you started your career as an engineer. And now that you're founding a company, everything gets weird and screwy, but again, your last quote-unquote, “Real job,” or capital J job, or corporate entity job was as a VP. I've talked to people who've been in the industry for 30 years who are still engineers, and oh, my God are they great engineers. But they seem that they have to deal with a headwind of if you didn't transition to management—and executive management at that—then somewhere along the way, you must have gotten stuck. Is that something that you've seen? And is it something you agree with?

Hal: I've seen people get stuck, but not because they didn't transition to management. And I've seen people not get stuck at all. There's an awful lot of people my age who are Fellows and Distinguished Engineers and the like, who have maintained largely individual contributor jobs. I started out, actually, my first jobs were actually more technical support jobs, which I actually highly recommend because it gives you a lot of customer focus, but people get a little scared of it because sometimes you'll find some resistance when you try to transition into engineering. Depends on what you've shown in terms of skills.

But I kind of went through these phases. I really wasn't interested in management. I did switch back and forth between what I'll consider operational line roles and staff consulting kind of roles. I liked having the influence over the direction of things and how organizations were built, and so on, but I didn't really want to manage people. And I did a short stint at management at DEC, which was positive, but not a plan; when my boss had a heart attack, and so I filled in for him for a few months.

And then I went back to individual contributor. And when I got to Microsoft, Microsoft was not as good a culture of having influencers. You needed to own things. DEC had been—the ladders were really peered. And I have to tell you the truth, a consulting engineer was considered a more influential—in fact, it’s a funny thing; I discovered that in the consulting engineer job description that it basically said, I could go stick my nose into business anywhere in the company.

Corey: I've never had a job description like that, and I’d do it anyway, I have problems. I'm a terrible employee.

Hal: There was one time I really exercised it specifically, as a consulting engineer, versus just me being nosy or whatever. I actually did go to somebody at one point and say, “I'm sticking my nose into your business because I think you're screwing up.” In a organization way far away. And they were like, “Fine, come on in.” And so, there were these peer-level kind of roles.

And then Microsoft, Peter Spiro and I, Peter had been my partner-in-crime in the database world at DEC, and we get to Microsoft, and we realized one day that we were having trouble influencing things, which is that the development managers were themselves ignoring us, even when our senior management was supporting us. And so I was standing in Peter's office one day, and he looks at me, goes, “I think we're going to have to become managers to make this happen.” I was like, “Yeah, I think you're right.” And so we took over as development managers, initially, and product unit matters pretty quickly.

Corey: You hoisted the veritable black flag as it were.

Hal: Yeah. And it turned out for both of us, we liked it. It was a weird thing; neither of us had it as a career goal. Both of us loved mentoring people. We loved helping with the direction of things, and both of us liked still being engineers.

I mean, this is the thing: we had the opportunity—this is why Peter never was listed as a—he never became a Corporate Vice President at Microsoft, he became a Technical Fellow; he was one of that first wave of Technical Fellows. And the thing is that we could keep being technical and become managers. Now, over time, that gets harder and harder because of time, but we were able to make that transition. And even with that, I went back to IC roles. This is why I list the Microsoft time as Distinguished Engineer and general manager because sometimes I played General Manager, and sometimes I really played Distinguished Engineer.

At my last management role at Microsoft was, we did a re-org in the security Products Division and we ended up about a half General Manager short. That is, we couldn't find anybody gives some stuff to without overloading them. And my boss, knowing I had the management experience, looked at me and said, “Would you mind being the GM and mentoring the guy who was the product unit manager for this big piece of it until he was ready to be promoted?” And I went, “No, I’d be happy doing that.” And suddenly, I'm back into management. But I go back and forth.

To your specific point, I think it's fine. And I know a lot of senior people never make that management leap. But they either continue to grow in technical ability and, most importantly, judgment, so that they have a broader span, or they get very happy with the niche they're in. There's not a lot of people that like what they're doing. They actually don't want to be a DEC, they didn't want to be a consulting engineer.

At Microsoft, there were people who didn't want to be a Partner or didn't want to be a Distinguished Engineer because it meant that they would have to spend a very large percentage of their time not working on their own technical problem, but helping others, and they were very heads-down kind of people.

Corey: Oh, I hate helping others. I can’t. I can't. But I'm with you on that. It feels like in many companies, industry-wide, there's this perception that if you're a great engineer, cool, then you should get promoted to management, which is not a promotion. It's an orthogonal skill, but it's not compensated that way, and it's not respected that way in most companies, hierarchies, so it's terrific.

And very often what you'll see is, “Wow, we're losing a great engineer and gaining an absolutely terrible manager. Yay, us.” And the way that it's done is you can't ever go back and have someone go back to being an IT role, in most cultures because it would widely be perceived as a demotion, and people won't rightly stand for that. So, it's a really hard problem that I think very few companies seem to get right. Now, I think that from the outside Amazon, and Microsoft, and Google have done a very good job of this. But as far as smaller companies go, I question it. It's hard to say it's small scale.

Hal: Yeah, I think you're right, it varies a lot in company culture. DEC, one of the weaknesses of its culture was it would push people into management who would not be good managers. It didn't focus on their management skills. And so that that frequently happened. Amazon is kind of at the extreme other end of it which is that they're kind of skeptics about people wanting to move into management.

They wait until they’re higher level than most companies would before they'll let them move over. In AWS terms, they mostly wanted managers to be level six or higher, and occasionally, you'd let somebody an L5, be a manager because you were about ready to promote them to L6, you just wanted to get a taste of management. But it was pretty rare. You want it to focus on their management skills, Microsoft’s somewhere kind of in between those two scales. But yeah, there are a lot of companies which, like, “Okay, we need a manager. You're the best we got. You're the manager.” And that can be a disservice, a big disservice to the employees. That's the manager there to serve employees basically and help make them more productive.

Corey: The idea of servant leadership is great. The problem is, look at managers you've had over the course of your career. How many lived it?

Hal: Yeah, I think I've been pretty lucky, but I've been, most of the time, able to select my manager. [laugh]. So, I avoid the ones I don't think would be good, that I won't learn from, or won't be supportive. I had one at DEC who I thought was—he was not only a terrible people manager, he was a terrible large organization manager. And I remember he came to me when I was at Microsoft, looking for a job, and I was very polite, but I was thinking in the back of my head, “No way will ever let this person get hired by Microsoft.” And he's like the only case that really falls into that.

I’ve been pretty good at finding managers who've been very supportive of my career. Even some managers who others didn't care for. And so I had one manager, who was actually very well known, for another company becoming one of the longest-serving top leaders at another company, who I was like, “You don't understand this—” When people complain to me about him. I'd say, “He's fantastic with managing up, and so you just have to think about using him the right way. He's not going to be there for you every day, but if you understand that he's making sure nobody interferes with what you're doing, and what you want to do, and that's where he's putting his energy, and use him the right way, then he's great.”

And that worked for me because I'm pretty self-sufficient. And for others, though, they needed more direction from him and more time with him, and it didn't work for them. So, even some people who get perceived one way are actually pretty good, as long as you can figure out how to interact with them the right way, how to take advantage of the things they're good at. And I just wanted to say, I think overall, there's only been the one person who I really thought was damaging both to myself and to other employees, and to the company. Everyone else has been good to very good, or even excellent.

Corey: Yeah, there's a very different way of framing these things, I suspect. When I hear managing up, I think of something radically different than I think most people intended. What I hear is, telling your boss to go to hell in such a way that they look forward to the trip and begin eagerly packing. I don't think that that's exactly how people mean that phrase, but it's what I hear, and in my experience, it seems to be just about right.

Hal: Well, I'm sure in some cases that that is true, but a lot of cases, who says that you're doing work that three or four or five levels up the management structure, understands, believes in, is willing to fund appropriately, when times get tough they're not going to disproportionately attack it from a budgetary or layoff standpoint. How do you keep something strategic? Again, this is even more important in a big company, how do you keep something strategic to the company, when it has 100 things and it's going, “I can't do a hundred things—or a thousand things—I need to whittle this down.” How do you stay up there? How do you find opportunities for your technology or products to be used more broadly in the organization?

And what happens when you have senior leaders who are enigmatic and strongly opinionated? And how do you keep them supportive of the team overall? And there are people who are very good at that. And you can look at some big company’s executives disappearing all the time because basically, they finally had it out with the CEO or something. And so somebody who can manage one of these very opinionated CEOs is a very important skill. And organizational structures and all that.

But they may not be that great at one-on-one management of people below them, especially as you get more senior if you're expected to be able to operate independently. And you think your manager is going to be there constantly to support you, well, that could be a problem. Now, if you're junior you expect it. I mean, if you're—you expect first and second-level managers to be able to be very supportive of the people below them because they're just not that senior themselves. And they are going to need the guidance.

Corey: I think that's very fair. And also probably a good place to leave this episode. If people want to hear more about what you have to say, where can they find you?

Hal: So, I blog at hal2020.com—H-A-L2020.com. I have not been super active this last year. I go through periods where I kind of get a writer's block thing going. And other times, I’ve been super active. But there's a lot of material there.

There's a lot of historically interesting material there, I think. I've had ex-Microsoft people or current Microsoft people—there were people trying to understand what's going on in Microsoft told me they read my blog to find that out even though I’d been out of Microsoft for years. That's the best place to go. I'm also on Twitter. I don't even remember my Twitter handle. I think it's @halberenson. And I'm on there a fair amount. Again, not as active lately, but I go in and out of being very active.

Corey: Excellent. We'll of course include links to those in the [00:34:18 show notes]. Hal, thank you so much for taking the time to speak with me today. I really appreciate it.

Hal: You're welcome. It's been fun.

Corey: It really has. I'm delightful. And so are you. Hal Berenson, founder at Gaia Platform, LLC, as well as many other things, as we've just discussed. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you've hated this podcast, please leave a five-star review on a podcast platform of choice—once again—and a comment telling me that I don't understand it yet, but I will when I'm older and been in this industry just a little bit longer.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

Links Referenced:

  • JFrog’s Website
  • Follow Kat on Twitter
  • Connect with Kat on LinkedIn
  • Email Kat at katc@jfrog.com

TranscriptAnnouncer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

This episode is sponsored by our friends at New Relic. If you’re like most environments, you probably have an incredibly complicated architecture, which means that monitoring it is going to take a dozen different tools. And then we get into the advanced stuff. We all have been there and know that pain, or will learn it shortly, and New Relic wants to change that. They’ve designed everything you need in one platform with pricing that’s simple and straightforward, and that means no more counting hosts. You also can get one user and a hundred gigabytes a month, totally free. To learn more, visit newrelic.com. Observability made simple.

Corey: This episode has been sponsored in part by our friends at Veeam. Are you tired of juggling the cost of AWS backups and recovery with your SLAs? Quit the circus act and check out Veeam. Their AWS backup and recovery solution is made to save you money—not that that’s the primary goal, mind you—while also protecting your data properly. They’re letting you protect 10 instances for free with no time limits, so test it out now. You can even find them on the AWS Marketplace at snark.cloud/backitup. Wait? Did I just endorse something on the AWS Marketplace? Wonder of wonders, I did. Look, you don’t care about backups, you care about restores, and despite the fact that multi-cloud is a dumb strategy, it’s also a realistic reality, so make sure that you’re backing up data from everywhere with a single unified point of view. Check them out at snark.cloud/backitup.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Kat Cosgrove, who's a developer advocate at J. Frog, and an actual cyborg. Kat, welcome to the show.

Kat: Thank you for having me. I appreciate the mention of the actual cyborg. I feel like that gets glossed over a lot, and I'm not a fan of it.

Corey: It does. Can you tell us more about it? It sounds like if you're going to put that as the opening line to your bio, there's probably a story there.

Kat: Yeah, so I have an NFC chip implanted in my right hand. I got it at Defcon a couple years ago with my best friend; he also got an RFID chip. And I mostly use it when I'm at conferences and—you know, when we're allowed to be at conferences—to give people my contact info. So, if you tap your phone up against my hand, it will load either a link to my LinkedIn page, or to my Twitter page, or a vCard for my contacts.

Corey: It becomes pretty obvious that if Cyberpunk 2077 have been delayed even further, it would have just been a documentary. My God.

Kat: True. Yep. It's pretty rad. They have ones with LED lights now, so that’ll probably be my next one.

Corey: Would you consider it a form of minor surgery at some level? Because what you're saying and—or at least what I’m hearing is, “Well, I had some minor surgery done in the back alley at Defcon,” which, all right, you know, I'm not one to judge out loud, but I look at that and I do have some questions that would be natural follow-ups to that.

Kat: Well. It goes in with a needle, actually. They come prepackaged in an injector that just goes a little bit under the skin between your thumb and your forefinger on the back of your hand. And it’s—doesn't require a stitch, doesn't require any glue, just a band-aid. It's really quick.

Before any of the Bill Gates vaccine microchip conspiracy theorists jump up with their hands in the air, it is a very large needle. There is no universe in which this could be hidden in a vaccine, though I wish we did have chips small enough to do that because it would be rad.

Corey: Yeah. At this point, we pay way too much for our iPhones for people to have to chip us involuntarily.

Kat: True.

Corey: But I digress. So, we'll talk a little bit—in a minute or two—about what you're doing now, but first, let's talk about how you got to where you are. This is one of the problems when you look at folks who are well-known in the space as you are, it's natural to assume that you sprung fully-formed from the forehead of some God. That is, I'm told, not strictly true. So, from whence did you come?

Kat: Yeah, not strictly true, while my dad is very cool. I didn't go to college for computer science. I barely graduated from high school. I went to college for biochemical engineering, and I also dropped out pretty quickly.

Actually, I then went to work as a bartender at a strip club. I don't talk about that a whole lot publicly because I was worried at first that people would look down on me. But now I've decided that I just don't care and it's not something you should be looked down on for. From there, I went to work as a clerk at an independent video store, a pretty big one. Like, an average Blockbuster had about 600 titles in store, we had, like, 41,000. And this was not decades ago; this was like ten years ago.

Kind of like, weaseled into tech from there because they needed their computer replaced, the one that was handling the rental database and it was running Windows 98 SE, in [laugh] 2008, 2009. And I replaced it, and from there taught myself SQL, became their database administrator, which was pretty rad, started teaching myself some other programming languages. Ended up doing some freelance web dev, mostly WordPress. And then I moved to Seattle and went to a coding boot camp.

Corey: And now here you are as a developer advocate at JFrog.

Kat: Yeah. I went from zero to relatively popular on Twitter.

Corey: So, let's tie those two together: things from the very beginning and things at the end. So, yeah at first, I think that the, I guess, social shaming of folks who work either in sex work or adjacent to sex work is bullshit and I have no tolerance for it, but let's talk about what did you learn as a bartender that serves you well in DevRel?

Kat: It can be a difficult task to manage people who are needy and demanding, but you can't meet their needs, but you also have to keep them happy. It teaches you to deal with unruly people in a very specific way that keeps everybody happy. It respects my boundaries and the boundaries of my employer, or society without forcing me into a difficult situation with somebody that I don't want to piss off. Of course, frequently I do just piss people off if I've determined that they're not somebody I need to keep happy.

And that's also something you learn from bartending: where the line is; when to fire a customer, so to speak; when it's not worth engaging with somebody, and it's better off to just ignore them, or get management, or start a fight. It's a difficult thing to learn. It does also, unfortunately, teach you how to deal with tolerating a shitty situation, tolerating difficult people, tolerating somebody who is not going to respect you no matter what you do. I don't think that's something we should have to learn, but it is a lesson I learned and it has made doing my job easier and being a minor Twitter celebrity easier. [laugh].

Corey: What's fascinating is that so much of what you just said, is almost a foreign concept to my lived experience, almost as if not everyone is treated the same way. Namely, by which I of course, mean a cis-hetero white man. And that is a tremendous problem.

Kat: It is. It is a tremendous problem. It's wildly inappropriate. I hate that it's still an issue. But we are collectively being louder, and louder, and louder, and less tolerant of that kind of behavior if recent blow-ups on Twitter over the last few months have been any indication.

Corey: I can't tell if we're actually changing people's minds, or getting them to keep their stupid opinions to themselves. And if I'm being perfectly direct, I'm not sure I care. I like the outcome which is, if I have this misguided belief, that based upon what someone's background is, or what their appearance is, or how they express themselves, [00:07:00 of a] gender, or otherwise, winds up somehow invalidating or validating the legitimacy of their opinion, I kind of want you to shut up and keep that thought to yourself because it's not the people I argue with on Twitter that I care about so much as it is the people who watch that argument unfold, and what they take away from it.

Kat: Right. Every time I pick a fight with somebody on Twitter over some kind of gatekeeping, or judgment of another person's worth in tech-based on whatever about themselves, that argument is never to change the mind of the person I'm fighting with. Because you're not going to, probably. They're set in their ways. This guy's just an asshole, and there's nothing you're going to do to change that. But what you can do is broadcast to everyone who follows you, and everyone who follows them, and everybody who sees it, that this shit is not going to be tolerated anymore. And that's a little bit more valuable for me.

Corey: As someone with relatively recent, newfound Twitter celebrity, as you frame it, how much of a distinction do you draw between Twitter and the real world? That's what I wrestle with a lot.

Kat: I, in a lot of ways, am just me on Twitter. There are some aspects of my personality that are considerably louder on Twitter, and some that are considerably quieter. In real life and on my private social media accounts, I don't post that many selfies, and I don't dump my every single, half-awake thought on to my private social media accounts like I do on Twitter. That is absolutely an engagement thing and making sure people still see my stuff thing. But the one thing that is consistent between internet Kat and real-life Kat is being aggressively intolerant of intolerance. I just do not put up with that online or in real life.

Corey: One of the biggest problems I see across the entire industry is either explicit and direct forms of gatekeeping or backhanded, slowly subtle ways of gatekeeping. Either way, I can't stand it. This industry is difficult enough to master, and the technology is expanding, the surface area is geometrically exploding, and we need more people involved, not less. “Well, you didn't have the following educational credential,” or, “You didn't go to the proper school for the right kind of thing,” is just—this is awful.

I'm sorry, I look at the things I work with on a day-to-day basis, there is no academic program that tackles this sort of thing: the technology is moving too fast. And it comes down to learn how you learn best, but I don't think that there's any particular credential that is required to excel in this space. But let me also call out that I am talking about software, and technology, and evangelization of same. I am not talking about becoming an anesthesiologist.

Kat: [laugh]. Yeah, I think we got into some weird position where people in tech have started to think of themselves as super elite geniuses, and nobody could possibly do this, and we're so valuable to humanity. And yeah, we do a useful thing. We build useful tools. Most of our lives are touched by technology in some way, but we're not God, dude.

It's so, I don't know, gross, I guess, to think of ourselves as better than everybody else because we beep boop, make computer do thing. And it doesn't require a special education for most things, it really just super does not.

Corey: One of the problems that I perpetually run into is—how do I put this—folks who try to imitate aspects of what I do without the nuance, where it's, “Hey, you're out there insulting companies. I'm going to do the same thing.” But instead of, I don't know, a trillion-dollar company, they wind up going after a five-person startup or whatnot, it's no, no, you're just being a dick. But thank you for the attempt.

Kat: Yeah, that's reading the room wrong. There's a difference between being an asshole in a funny way, being mean to AWS, and being mean to some dude trying to start an app company out of his garage. That's not fair.

Corey: Or making fun of an aspect of AWS, versus, “You work at Amazon, therefore, I'm going to corner you, as some rando developer, and accuse you of war crimes.” It doesn't work that way.

Kat: It really, really doesn't. And it's something that, I think, I didn't expect to be shouldered with the responsibility of determining who I mean to on Twitter as early as I did because I went from 4000 followers to 11,000 followers in about 36 hours. And that's the difference between having a moderately popular niche Twitter account and being a minor internet celebrity in a particular field. And whoo boy, that has definitely changed the way I interact with people because I'm terrified of accidentally saying something mean to somebody with 200 followers who didn't realize they were stepping in shit, you know?

Corey: Whenever that happens, and I find I've done it—which I'm getting better at steering away from where that goes—but I'm also looking at this through a lens of there are certain aspects of Twitter that will—nothing you do once you set a foot wrong is ever going to be enough. But I look at the folks who have become—how do we call it—the main character today on Twitter, and it's not just having a bad take, it's about doubling or tripling down on that bad take once people come back with, “Ehh,” moment. And I've gotten things wrong in the past and it turns out that a sincere apology is generally enough for most reasonable people. Now, there's always going to be someone for whom it's never enough. But those aren't the people you necessarily want to engage with and have these perpetual conversations with. And at some point, liberal use of the block button is the right answer.

Kat: It is. I was not aggressive with it when I had a small account, but now I get so many, like, crappy replies and weird DMs that I'm pretty aggressive with it. If you've got no profile picture, and you're following 350 accounts, and only have one follower and you say something weird in my replies, I am not going to gamble. I am just going to block you.

Corey: No, at some point, it doesn't make sense. And it's not just about Twitter followers and the rest, this also has to do with the real-world implications of when someone has positioned themselves as, I don't know, a senior engineer at Google, let's say—

Kat: Yeah.

Corey: —your words carry weight you may not realize that they carry. Right now, my shtick works because I am fundamentally, in the industry, a nobody. I run my own company, and that's great, but it doesn't have the gravitas of someone who's a distinguished engineer at one of the big tech companies, or someone who's been in the space and built a thing. When you're that kind of person, even if you don't intend it, there's no way, no matter what you say about, “Oh, these are my opinions, not my employer’s,” that's not how it's going to get cited in some BuzzFeed article or whatnot. It's always going to come back to, you have caused a problem for your employer. I don't have that burden. And you sort of do, in a different way.

Kat: I do.

Corey: Not that I think that folks necessarily take what you're saying on the internet as the gospel truth, according to JFrog, but there is the problem of having an actual company with a lot of people behind it that, in some ways, are going to see indirect consequences in the things you say, despite your best efforts to the contrary. Do you find that that winds up shaping what you say, and how?

Kat: Yes. Two weeks ago, I would have said no, but now I'm saying yes because it happened to me, and it could happen to you too. So, Kubernetes decided to deprecate docker-shim, and this is something that had been discussed for literally years within the community. This was open knowledge that it was going to happen, but I guess the users didn't know that. So, when the changelog released, a screenshot of it was tweeted, highlighting the fact that docker-shim was being deprecated, meaning that you could no longer use Docker as your container runtime. I tried to be helpful. And instead—

Corey: Oh, that’s where it all starts: “I tried to help,” and then look what happens.

Kat: Yeah, so the moral is don't try to help. That's not true. Do try to help but be aware that there still might be repercussions. And I tweeted a very helpful thread explaining what was actually happening to calm people down because people were rightfully freaking out. And at one point, I made a comment that was like, “Docker isn't dead, yet.”

And I had 4000 followers when I tweeted that, and I thought it was just a cute off the cuff thing, but then the thread took off. And two days and 8000 more followers later, that has weight. I don't know how to predict that. How do you predict that kind of thing?

Corey: Oh, you absolutely don’t. I find that when I do big threads that blow up from time to time, there's usually some offhand comment that I've made somewhere like around tweet 80 or so, but that attracts the sketchy rando vibe, where it's, “Wait, are you saying—” and they go in some completely different direction. And I'll respond to one or two with, “No,” but beyond that, like, I can always tell when a tweet goes beyond my typical readership because if I go out and I tweet right now that multi-cloud is a stupid idea, as it is implemented almost everywhere, most of my followers at this point are used to me taking that opinion. If that blows up beyond a certain sphere, then I start getting a lot of the, shall we say, “Well, actually,” in the response, and that sucks, where it's, “Oh, great. Now I get to clarify exactly what I meant.” And no sooner have I done that to one person asking me, then I have five more with variants of the same thing. And then it takes off, and at some point, you're playing whack a mole, and I just give up and start tweeting something different.

Kat: Yeah. And that's what happened to me in that case. It was a bunch of people rapid-fire asking the same question once the thread took off about, “Well, what do you mean?” People assumed I worked for Google. People assumed that I worked for the CNCF, or was actually a Kubernetes maintainer, people assumed that I worked for Docker.

And my employer is right there in my bio, so they were not even bothering to check. And I was fielding the same angry question a dozen times over the course of a couple of hours. And I just muted the thread because I answered it three times, and then, come on. Just read. Just read the thread, please.

Corey: Yep. People never read as much as you think they will. I've had people who are amazed that AWS hasn't fired me yet. I'm not sure you understand exactly the nature of my employment situation, but that's all right. I'm still waiting for someone to reach out to my business partner, who is the CEO of The Duckbill Group—Mike Julian—and attempt to have me fired over something I have tweeted. That's going to be a glorious day. A spoiler: he can’t.

Kat: You know what, that has happened to me once. Somebody tried to get me fired because of something I said on Twitter. It did not work. But it did happen. And it was funny. So, it's a weird thing to do.

Corey: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the Enterprise (not the starship). On-prem security doesn’t translate well to cloud or multi-cloud environments, and that’s not even counting IoT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IoT devices, detects these threats up to 35 percent faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at extrahop.com/trial.

Corey: The thing that bothers me is when I wind up doing a tweet that inadvertently winds up smacking at people I didn't intend for it to smack at, that's the stuff that keeps me up at night. Not that I'm going to upset some alt-right dingbat, but rather the fact that I'm talking about a product decision being crappy, but I don't want the people who built the product to think that I'm crapping all over their hard work. So, it's a weird problem where things like that tend to focus on one or two specific people, but I want to be nuanced and careful in how I criticize the product—“Because that's how they improve—without ruining someone's week where, wow, that loud jackwagon on the internet thinks that this is awful.” “Well, yes, but I didn't mean for you to take it that way. I mean that there's an opportunity to improve here.” But you never get to add that nuance in 280 characters.

Kat: Yeah, you don't. And it's one of the reasons why I try to avoid aggressively talking shit about other products unless it's something I genuinely, really, really do not like, which I think has only happened once within the last year. I wrote a really nasty blog article about an AWS product, actually, that I now do not remember the name of because they announced it, and I genuinely was like, this is so incredibly stupid.

Corey: Oh, I say that, like, three times a week. But normally, I've started bounding it to making fun of the service name because no one spent 18 months naming Lookout for Metrics like it's a warning sign. And if they did, they should feel bad.

Kat: Yeah. They obviously do not spend time on naming things. And I get it; naming things is hard. That's a long-standing joke in tech that will probably never go away, that the hardest thing in tech is naming things, and off-by-one errors. But it feels like they are not trying, Corey [laugh].

Corey: They're not. At some point, naming things is actually super hard because making fun of the name: relatively easy, but what are you going to call this thing? And, that exists, that's a trademark problem, et cetera, et cetera. I mean, I just launched the DuckTools SaaS offering that we're doing and I'm waiting for someone to draw the Ducktales comparison at Disney and start yelling at us for it. It's hard to name things.

And… we'll see. I'm not terribly concerned. But there's also this problem of how do you wind up having something still tied to a unifying brand, but it's clearly its own thing. And when you start launching something like Systems Manager, great, you don't necessarily know in advance that it's going to have 20 different sub-services and that calling one of them ‘Session Manager’ becomes a valid thing. But at the end of it, of course, the name is dumb in hindsight, but how do you go back and fix it because naming is hard; renaming is worse.

Kat: I don't think you can. I think you just have to own it at that point. But AWS has had a long time to figure this out. And a lot of product families where they have made the same mistake over, and over, and over again, nothing stresses me out like looking at my AWS dashboard. I just don't do it anymore.

Sorry, AWS. I'm actually not even vaguely sorry. But yeah. They’ve made the same mistake repeatedly with naming things. They released a thing called AWS CodeStar that lets you spin up a skeleton of some project in whatever framework and language you—

Corey: And tie together a bunch of CI/CD services on the AWS side that don't have any actual customers here in the real world.

Kat: That's a separate product family. They did CodeStar years ago, and then they just released CodeArtifact and CodePipeline—or whatever it is—within the last year. So, they already had a thing that was called ‘code something,’ and then they went and released a whole family of CI tools called ‘code something,’ and they're not related. It drives me absolutely bonkers.

Corey: And one of the things that drives me nuts is, take a look at Cloud9, I really was interested in that when it launched, it was oh, cool, almost like—like, they focused on the idea of having it as a hosted platform—as opposed to Visual Studio Code, which is focused on building an IDE—and Visual Studio Code evolved to the point where you can now run it online with GitHub Codespaces, and that is freaking incredible.

Kat: Yeah, it's rad.

Corey: And Cloud9 basically then kind of sat on it’s ass for two years, and everyone forgot it was there. And it wound up launching a new API at re:Invent this year that I'm still not sure what it does because I don't care because neither I nor anyone else I know uses Cloud9. And it's not because I don't want to; it's because it's bad. Whereas every problem I have with it is solved, by and large, with GitHub Codespaces. So, rather than screaming about wanting a first-party thing from AWS, I'll use the thing that's actually good.

And I don't understand why AWS just likes to let things sit and marinate for years. I keep expecting there to be some giant release that modernizes all of this, and they don't. Instead, they just do an iterative improvement a few years later. It's like, I don't know if they disbanded and reconstituted a team for it or what, but God, it drives me crazy.

Kat: I would love to know why they do that, actually. I would really, really love to know what the business logic is behind that. Is it somehow beneficial to just keep cranking out a ton of services and just sit on them so nobody else can have it. Or when somebody else releases a better version of that thing that actually takes off, they can take the moral high ground and be like, “Well, we did it first, just worse.” What is it?

Corey: Well, then you have the other side of it: CloudShell, which just came out, which people have been wanting for years. They were late to the market in a rather severe way. Google came out with this five years and two months beforehand, and it was transformative. And since then, Microsoft came out with it, in 2017, for Azure. And then at the beginning of 2020, we saw that Oracle Cloud came out with it, and then, holy crap, IBM Cloud came out with it in June.

So, when you're trailing IBM to market on something, really stop and evaluate what's going on. I love the service, but at this point, I'd mostly mitigated all the things I would have wanted it for. The idea that you log in as an IAM user on an account and have the access to run various things in that account is huge. You don't have to teach beginners to configure their local environment for half a day first. I've been stuck in that particular trap.

And it's still obnoxious because to have access to this, you need the power user admin-level tier in your account. So, okay, that means it's going to be challenging to wind upscaling that down at the moment. But ugh, I want to see a better answer, and I worry this is going to be one of the things we never see an update for, ever again.

Kat: I didn't know that they had finally released that, and the lack of a service like that is a huge reason why I just hard committed to GCP because I don't want to spend half of my day messing with my environment to make sure that I can actually connect to my AWS environments. I can't believe they finally released one. I assumed they were just going to die on that hill.

Corey: Yeah. I don't for the life of me understand it, but here we are.

Kat: Well, good for them. Welcome to 2017, I guess. But if IBM and Oracle beat you to something, you need to really, really look at yourself because nothing makes me think old, slow, and stodgy, like Oracle's overall brand. That's a product of growing up the way I grew up and having a dad who's an engineer and just deeply does not like Oracle either. So, maybe there's some personal bias involved, but oof; big oof on that.

Corey: Yeah, it just becomes a difficult thing to do because, again, as soon as you get into the cloud space, everyone wants to start smacking the competition and the rest, and it's harder to do than you think because there aren't too many people who are going deep into multiple cloud providers because why would they?

Kat: You don't need to. You don't need to. And honestly, for the overwhelming majority of the stuff I do in my personal side projects, I just chuck it on Heroku.

Corey: Yeah, absolutely. No problem of that. I use Heroku myself for a few things.

Kat: Yeah, it's rad. It's easy to use, it's CLI is sane.

Corey: It hasn't really changed much in ten years, and that's kind of a benefit.

Kat: Yeah, that's kind of the benefit. I don't have to learn something new every six months, when I get a harebrained idea and decide, I will start another side project that I will definitely finish this time. I know it's going to be the same deployment process than it was the last time I did this six months ago. With AWS, I have zero guarantee of that.

Corey: Yeah, trying to predict the future in the time of Cloud is always weird. Like, you have these companies that are just smacking at each other, they're trying to find their next big customer as a customer of a different cloud provider. But there's so many folks still moving to Cloud and so many workloads that are still on-prem that, fight for those instead. That's the opportunity. Alternately, focus on getting a bunch of folks to start building their ideas on your cloud provider.

And that's where the big companies of ten years from now are going to come from. It just—it's a longer term game. I think that Microsoft Azure is going to be a behemoth in the space, not because Azure itself is awesome, but because if they play this right, their integration story for puttering around with Codespaces on GitHub and GitHub Actions during the CI/CD testing pieces, and if you can just have a single-click deploy to Azure at that point, well, why wouldn't you? And they stand to more or less capture an entire segment. But Azure has to get better than it is right now, too. So, I'm optimistic about the future there, and I think that something like that would be fantastic for customers, I don't want there to be a one right answer to what cloud provider someone should use. I want people to have different conversations.

Kat: Yeah, Microsoft has done a really, really good job of turning it around lately. Like, obviously, Microsoft has problems politically and with respect to some contracts they have. But they aren't the evil company, I remember growing up with, anymore. Intel still gives me the same vibe that they did when I was a teenager and hated Intel, but Microsoft has really—I don't want to say that they're the good guys now, or whatever, but they don't give me big evil corporation vibes like they used to.

Corey: No, they've done a fantastic job of turning it around. It is the story that I'm seeing right now of corporate turnaround, where there's this incredible value of what they've been able to build and how they were able to achieve it. The other side of it, though is that, cool, that means that the things that they do that upset me and don't align with that, are ever more annoying. And okay, that's great. But there's also the painful part of what are they trying to build toward and what things, corporately, are in their way of really taking it all?

Because they've done such a good job at brand rehabilitation in some areas, but they haven't touched it at all on any of their licensing story. So, great, how do you unify all of the company behind this particular, I guess, addressing of that transformation story? It can't just be piecemeal anymore. And the idea of being able to do this, from a perspective of something that is across the board rather than on a business unit by business unit process is sort of their next obstacle. And I haven't got a clue on how to solve for that thing, but I am curious to see how it works out.

Kat: Yeah, I don't know how they're going to handle that either. And they're not the only company with that problem—of that size. Amazon overall definitely has a reputation for being a place to burn out as a software engineer. You go get hired at Amazon on some AWS team, fresh out of college. They pay you gobs of money, mostly in stock.

You work your ass off for a year, maybe two, and then you burn out and leave and you go to Microsoft. But that's not the reputation Twitch has, and Twitch is owned by Amazon. So, why do they have this cultural divide between teams? Microsoft has the same problem. I assume a similar issue exists at Red Hat, and IBM, and Oracle between some teams being, “Cool—” You can't see me because this is a podcast, but I'm doing air quotes—and being miserable to be on. I have no idea how that happened at any of these companies, and how they correct it, but it's a problem that needs to be addressed, I think.

Corey: Yeah. And I wish them well. And I think that there's going to be an awful lot of stories where this becomes better for everyone across the board. Was it a Bill Gates quote that said something along the lines of, “Do you underestimate how much you can get done in a year, and underestimate how long you get done in a decade.” Or something very close to that. Don't @ me. But the sentiment is there. Like, we take a look at where all these companies were ten years ago, and where the state of technology was ten years ago, and we're living in a completely different world. But on a day-to-day basis, we don't see those sweeping changes.

Kat: Oh, for sure. I found my hard drive from high school. And it's a Buffalo brand TeraStation. This was the first one-terabyte hard drive I bought, and it is enormous. It's, I don't know, around the size of a loaf of bread. And also, I remember paying, like, $500 for that.

And now I've got six terabytes of solid-state storage in my desktop, and I think I spent 400 bucks total across that. But it happens so slowly, I didn't notice. And all of technology is like that, I think. Every once in a while there's an outrageous breakthrough, but for the most part, it's a slow, slow revolution until we wake up one day and realize, “Oh, shit, this changed a lot in a decade, didn't it?”

Corey: Yeah. I take a look at my typical workflow on a day-to-day basis, and, “GitHub? Why would I be going to GitHub? I'm not much of a developer.” And weird, that’s kind of a lie we tell ourselves. I spend more time writing code these days than I would have expected, but I still don't think of myself as a developer. “Amazon? You mean the bookstore? Oh, that Cloud thing took off.” And look at how the world changes in a relatively short period of time. But there's no one day I woke up and thought, “Oh, my God, Cloud. This changes everything.” It's a creeping realization.

Kat: It is. And it has taken over a huge portion of our lives. I remember it being a big deal, and everybody was very confused and very concerned the first time an AWS region went down and it took down half of the internet. And now that happens on a semi-regular basis. And every time we're like, “Oh, what is it this time? Is it Cloudflare? Or is it us-east-1—” or-2? Which one was it this time, that went down?

Corey: 1 is generally the one that experiences issues, and it's so—the problem is, that's such a big region that a blast radius is enormous.

Kat: Yeah.

Corey: And people say, “Oh, us-east-1 is just a constant tire fire of downtime.” But looking at it for the past few years, it’s really not. The days when it has issues, which there have been a couple of them, but it causes massive disruption. But then you look at the data center that you're running for your company, and the reason that it doesn't have that reputation is that your customers don't care because you're not that big.

Kat: Yeah. They run a huge portion of the internet. And it feels like that happened overnight, but it totally didn't. And I never would have thought that we would be in a point where [laugh] an AWS region going down would make it literally impossible for me to do my job for a day. But here we are.

And also, I'm not complaining because I do need the time off every once in a while. Burnout’s real. So, anybody at AWS wants to pull a plug occasionally, just, like, strategically, that'd be cool.

Corey: Yeah, I think that would be worth doing in some respects. Ugh, there’d be days. So, thank you so much for taking the time to speak with me. If people want to find out more about your hot takes, where can they find you?

Kat: You can find me on Twitter at @dixie3flatline. If you don't know what that's a reference to, and thinks I just crammed a bunch of words together from a Google search, it's not nothing. It's not nothing. It's just a much, much deeper cut than I expected. It's the name of a character from a book, Neuromancer by William Gibson. So, if you don't know what that is, you should also read that book because it's really good. If you don't have Twitter, first of all, why?

Corey: That's an excellent life choice, is my counterpoint, but please, continue.

Kat: Yeah, you're right. Sometimes I regret it. You can email me at katc@jfrog.com, but I will warn you I am not great about responding to emails. So, best of luck to you if that's the route you choose.

Corey: We will of course include links to that in the [00:34:33 show notes]. Thank you so much for taking the time to tolerate me, and suffer my slings and arrows, and ridiculous questions. It's appreciated.

Kat: Yeah. Thanks for having me. It's been lovely.

Corey: It really has. Kat Cosgrove, developer advocate at JFrog. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you've hated this topic, please leave a five-star review on your podcast platform of choice and a comment including a link to your 200 follower Twitter account where you tell me I'm an idiot.

Kat: [laugh].

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Paul OsmanPaul Osman is a Software Engineer with 20 years of experience in the industry. He's the Lead Instrumentation Engineer at Honeycomb.io and is passionate about making production a less scary word. Having spent most of his career in the ill-defined space between software development and operations, Paul spends a lot of time thinking about making on-call experiences better, responding to and learning from incidents, and improving ways for software engineers to share knowledge. Before joining Honeycomb.io, Paul worked in Platform and SRE teams at Under Armour, PagerDuty, and SoundCloud.

Links Referenced:

  • Honeycomb.io
  • Follow Paul on Twitter
  • Paul’s Blog

TranscriptAnnouncer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

This episode is sponsored by our friends at New Relic. If you’re like most environments, you probably have an incredibly complicated architecture, which means that monitoring it is going to take a dozen different tools. And then we get into the advanced stuff. We all have been there and know that pain, or will learn it shortly, and New Relic wants to change that. They’ve designed everything you need in one platform with pricing that’s simple and straightforward, and that means no more counting hosts. You also can get one user and a hundred gigabytes a month, totally free. To learn more, visit newrelic.com. Observability made simple.

Corey: This episode has been sponsored in part by our friends at Veeam. Are you tired of juggling the cost of AWS backups and recovery with your SLAs? Quit the circus act and check out Veeam. Their AWS backup and recovery solution is made to save you money—not that that’s the primary goal, mind you—while also protecting your data properly. They’re letting you protect 10 instances for free with no time limits, so test it out now. You can even find them on the AWS Marketplace at snark.cloud/backitup. Wait? Did I just endorse something on the AWS Marketplace? Wonder of wonders, I did. Look, you don’t care about backups, you care about restores, and despite the fact that multi-cloud is a dumb strategy, it’s also a realistic reality, so make sure that you’re backing up data from everywhere with a single unified point of view. Check them out at snark.cloud/backitup.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Paul Osman, who is either a lead engineer or the lead engineer at Honeycomb, overseeing instrumentation. Paul, welcome to the show, and which is it?

Paul: Thanks so much for having me, Corey. Happy to be here. Lead instrumentation engineer? I don't want to say I'm the lead instrumentation engineer; that seems to put too much weight of responsibility on my shoulders.

Corey: Well, that's the whole point. It's all about weight of responsibility. That's why I'm mispronouncing it; it's actually ‘lead’ engineer, and it's all based upon density and the fact that you are never ever going to float.

Paul: Absolutely. [laugh]. Especially when you're putting software libraries in your systems. That's what you want to think about.

Corey: Absolutely. Everyone talks about these lightweight instrumentation frameworks. No, no. You go the opposite. You are the heavyweight instrumentation framework.

Paul: [laugh]. We will become the center of gravity in your system.

Corey: Exactly. Not quite the direction Honeycomb has chosen to go for a variety of reasons, not least among them being that it's a terrible idea. So, what do you do? What does instrumentation engineering look like at a company that is fundamentally—well, I'll get in trouble if I call them anything other than an observability company, but instrumentation is kind of what they do.

Paul: Exactly. Yeah. The most succinct way I can think about it is, my team works on the tools that help you get data into Honeycomb. So, if you think about a system like Honeycomb, you've got the platform, you've got the web UI, and then you've got everything that runs on a user's or customer’s system. And that's my team.

Corey: So, fundamentally, you're in charge of the agents, the embedded SDKs, the libraries that people shove into their systems, the—depending on how you're orchestrating it these days—the 800 Lambda functions, CloudWatch integrations, and whatnot, and run this magic CloudFormation template that instruments all of my AWS accounts to hurl information into your system. That sort of thing?

Paul: Absolutely. And the list is long, you're right. [laugh].

Corey: It turns out that the more that you put into your system, architecturally, the more things there are to monitor.

Paul: Exactly, and to pull data out of. You mentioned Lambda, and there's a whole bunch of interesting ways that you can get data out of Lambda functions. Who knew? It's not just—

Corey: Especially their new extensions API—

Paul: Yeah.

Corey: —which is super interesting. I haven't gone diving into it in any depth yet, but I like the idea.

Paul: I really like the idea. I'm a big fan of serverless in general. You could call me a convert because I was honestly skeptical at first, but the idea of creating a platform where you just ship freaking code and you don't worry about anything else. Now, having the ability to run processes in parallel, run sidecars in a serverless environment, I think is really, really cool.

Corey: There's so much capability that's, I guess, fantastic to see. It's amazing to, I guess, look at the complexity of even toy applications. And, on the one hand, it's, “Wow, what an amazing system I've built with all of these different services, and everything tied together, and the way that it interacts with another, and even if it's well-instrumented, that's great.” And then the other side of it is, “So, what does this application do?” “It shows people pictures of cats.” And that's really it.

And at some point, it feels like this is painfully overwrought. Now, this is not a new problem; it feels like that's a bit of a cyclical thing. Things get so complex it no longer fits in anyone's head anymore, and then there's a collapsing function of an abstraction layer that winds up becoming broadly adopted, and then the cycle repeats anew. At least, that's my impression on this having been spending the better part of the last two decades in the ops engineering space, but you have spent two decades in the ops and engineering space. What's your take on it?

Paul: Yeah. This is something I wrestle with a lot. The idea of complexity, right? You can look at a lot of these sort of architectural guides and just go, “Holy crap, there's a lot there.” And sometimes that's what you need.

So, I think struggling, or balancing, or figuring out the balance between needed complexity and kind of accidental or complexity debt is key there. For a simple thing, you want the simplest thing that could possibly work.

Corey: Yeah, and there's never any real tacit acknowledgment of that. It always seems that these frameworks and tools and the rest have, “Example one is ‘Hello, world.’ Example two is ‘Hello to the entire world.’” and it—great, not all of that stuff is needed for every environment, but you also probably don't want to build something hyperscale on the first example. There has to be some point of complexity where, okay, at this scale, the complexity trade-off is well worth doing and in fact, it's dangerous to not have it. That's not everything. Not every system needs to scale globally at all times. Now, that enrages some people when I point it out, but it's true.

Paul: Oh, yeah. And the type of scaling that you need is also highly dependent on the workloads that you're managing. You mentioned I come from an ops background. I was working as an SRE before I joined Honeycomb, and one of the things I've always tried to stick to, not always successfully, is the least amount of technology possible. If you're dealing with something that just has to horizontally scale out and you've got a pretty consistent workload, maybe you don't need it to be running a container orchestrator. Maybe you just need an ALB that can do horizontal scaling on CPU usage or something.

Corey: Yeah, it winds up being a problem when I talk about my philosophy on things as a best practice because I say other things that tend to fly directly against that. Somewhat recently, I got in trouble—again—on Twitter—again—for bringing up the idea that setting your database to your local timezone is a terrible idea. Put it in UTC, and then let the presentation layer figure it out from there, and the answer—legitimately—was, “Look, it's a local payroll app that's only for a one branch company in a single timezone. Why would you ever need to worry about that?” Well, for that kind of story, my position is if you're building this small thing, great, leave the door open for it to potentially become a big thing. 95 percent of apps will never hit a point of success where they need to go hyperscale, but for those 5 percent that do, don't bury landmines they are going to trip over down the road when that time comes.

Paul: Exactly. One of the things that can be challenging with examples like that is there are defaults. And we're not always aware of the consequences of accepting some of the defaults. And it can be really hard as engineers to think through, “What is the reversibility of this setting that I'm accepting, or this state that I'm accepting?” And if the answer is that it's going to be really hard to reverse, then maybe you want to think twice before doing that.

Corey: What are the problems that I keep seeing is that there's a lack of awareness of how to build hyperscale applications, and it occurred to me that part of the reason is, is that no one knows how to build a web property with hundreds of millions of users. I think that's true. Every company that has done that has had to figure it out as they go, for their particular workload, for their particular constraints. And this is proven out by the fact that if you talk to any hyperscale company about their application architecture, how things are built, ignore what they say at conferences on stage, pay attention to what they say at conferences in the bar after you pour six beers into them, and they all admit that it's crap. “Everything we've done is garbage we're doing as best we can, but there's a lot of rough edges. It feels like we're always a hair's breadth from disaster.” I can't shake the feeling that we're all just making it up as we go along.

Paul: I totally think we are. You mentioned earlier, best practices, and what the hell are best practices when they're so highly dependent on the specific architectural decisions made, on the traffic patterns, on the social aspects of how an organization works? I've had the good fortune of being part of a few teams that have had to scale up to hundreds of millions of users, and no story has been correct. This is one of the things that always used to annoy me about—I'm glad to say it doesn't seem to happen as much anymore, but when people would point at specific technologies, like, “Ruby doesn't scale,” or something like that.

That's a meaningless statement. What does that mean? It certainly has scaled for some people in some environments; it just depends on what you're actually doing. And like you said, there's no blanket advice that seems to work for everybody. There are principles, I think. And if we worked really hard, we could probably dig out some of those principles. But the idea that there's a one size fits all pattern, that seems to come from people who are trying to sell you something.

Corey: Oh, yeah. At the time that we're doing this recording, there was recently a great tweet by GitHub—or GifHub depending upon pronunciations—’s CTO.

Paul: Well, I'm Canadian. So, you know.

Corey: Oh yeah. The best part of this show is mispronouncing things. It's not Postgres, it's Postgres-squeal. I digress, the question that he was asking was, “If you're going to start a new company today, what technical stack do you pick? What cloud provider? What language?” Et cetera, et cetera. And my response to it is, “Oh, that's easy. It's the one that the engineers I'm hiring are conversant with and want to work in.”

Paul: Yeah.

Corey: Because I could look around the landscape and see an awful lot of business failures for a variety of reasons. I'm really hard-pressed to identify any of them as, “Ah, they pick the wrong technical stack.”

Paul: Yeah. How many companies have actually been sunk by a decision like that? It literally never happens. And for what it's worth, I completely agree. The right tech stack is the tech stack that you have experience with, the tech stack that you're comfortable with. Way more important.

And it's funny because people—I don't know, sometimes I feel like we talk about this less, but it’s, how comfortable are you with everything else? Who cares what programming language your code is written in if you're not confident in the way that you actually deploy changes. Or if you're not confident in the way that you configure how traffic is routed to it. That stuff, all—I would say—arguably matters a lot more than the actual expression of business logic that gets converted into machine code.

Corey: It really is. And that's what I want to ask you about, too, is that you have exposure to a bunch of different stacks, presumably because you are the instrumentation engineer who's made of lead, and you wind up building these integrations into every godforsaken stack that all of your customers are going to be using, or any of your customers are going to be using, which means that you get to touch a lot of different languages, you get to touch a lot of different platforms, presumably. Is that correct? Or am I—

Paul: Oh, yeah.

Corey: —dramatically overestimating Honeycomb’s compatibility with different systems?

Paul: Oh, no, no. You are absolutely on the nose there. When I was being interviewed by Honeycomb, we have a coding exercise that we send to a lot of candidates, and the only difference with me from an average product or platform engineer at the company was they had me do it in a number of languages just to see how comfortable I was moving from one platform to another because being on the instrumentation team, that is definitely part of the job.

Corey: So, at this point, it's one of those questions that I always used to ask my parents, “Am I the favorite, or is my brother?” And the answer that they gave was, “You're my children. I can't stand either one of you.” So, to that end, what is your favorite stack to integrate with, and your least favorite stack? Because, you know, it's not really a podcast unless you enrage people.

Paul: I'm pausing intentionally because we've been interrupted by an adorable three-year-old.

Corey: Aw. Yeah, I have one of those, too, lurking around here somewhere.

Paul: [laugh].

Corey: And an infant, but that's a separate problem.

Paul: Oh, yeah. [aside] Hey, can you go play with mama? [pause] She literally just came in, stole my phone, and now ran away.

Corey: Oh, yep. Sounds like a very similar story here. Also, thank you for not apologizing. It drives me nuts when people apologize for having the temerity to have a family.

Paul: Oh, right. Especially now, right? When we're all in our home.

Corey: Like, when the kid wanders in when you’re on a video call, “Excuse me.” She lives here; you don’t.

Paul: Yeah. Especially nowadays, when we're all literally working in our homes, right?

Corey: Oh my God, yes.

Paul: It's like, you're in her home, not the other way around.

Corey: I've also never asked an employee or colleague to turn on their camera.

Paul: Mmm. Oh, very good point. Yeah. Especially right now. That's a great [00:13:54 crosstalk].

Corey: Excuse me, invite me into your home like I'm some sort of godforsaken corporate vampire? No, thank you. We hit a perfect stopping point. What is your favorite stack to integrate with and your least favorite stack? Go.

Paul: Right, yeah, so 100 percent based on what we were saying earlier, the ones that I prefer, I'm going to surprise you: they're the ones that I have the most experience working in. [laugh]. And so I've trained my brain to think in a number of different ways, I think fairly well. I'm a really big fan of functional programming—a little. So, I like languages that tend to support a little bit of functional programming.

I come from a background—accidentally, I ended up doing a lot of Scala at a lot of different companies. And so I'm very happy working there. But conversely, I also really like working in Go, one of the languages that is often kind of made fun of—lovingly—for being a very basic language, and it's not too fancy in terms of features.

Corey: I want to be very clear here that my position is that language bigotry is awful.

Paul: Oh, yeah.

Corey: It's one of those ways of gatekeeping and it drives me nuts. It doesn't matter what language you pick, I can write shitty code and all of them.

Paul: Absolutely. And I have and I will.

Corey: It didn’t even compile, it's so bad. Personally, I don't get JavaScript to save my life. It does not match my understanding of the world. Python, conversely, is something that aligns much better with how I see things, and Ruby was also a great [00:15:12 unintelligible] for me for a while. I was also heavily into Perl for a long time. But again, as an old ops person, my favorite language is and always will be, bash scripting.

Paul: Oh, beautiful. Yes. It's funny, I have a very similar experience. Maybe it's something about us ops people, but JavaScript, I have not trained my brain to work that way. I completely agree with you about language bigotry being awful and a form of gatekeeping, and so my approach is when I see somebody who's proficient in JavaScript and can write wonderful applications in Node, or browser applications in React, I'm in [BLEEP] awe. It's just a way that I haven't managed to make my brain as compatible.

Corey: The challenge, of course, is that it's your responsibility to fundamentally support all stacks. So, how do you approach doing an integration in a language or stack with which you're not familiar?

Paul: Yeah. That's a great question. So, part of it is, you just kind of dive in and kind of work through it, which I think if you've worked in enough companies that have different languages and different stacks, you might have some experience doing. I've worked in companies where [laugh]—I worked in one company once where we started the whole microservices journey, and we regretted this decision—spoiler—but we said everybody can choose whatever language they want to use because it doesn't matter at the end of the day; we're all talking to each other over HTTP and JSON APIs. So, that resulted in this Cambrian explosion and, surprise, if you wanted to go and work on something on a different team, or that a different team had created, it's going to be in a language you may have never even seen before.

And so part of it is, you just got to kind of dive in and be willing to learn. Where there are real gaps or weaknesses, that's where hiring becomes important. It's funny, I've been a hiring manager in the past in previous lives, and I've been involved in hiring processes at a bunch of different companies, and I'm very opposed to just hiring based on specific technology or language experience. But sometimes you'd have to say, “Oh, it's a real bonus if this person fills a gap that we don't have [laugh] on the team at the moment.”

Corey: Oh, absolutely. I think that hiring is one of those hard parts where it's easy to fall into the very common trap of never ever wanting to hire someone who's weak in something, as opposed to, “Okay, great. Maybe your Python is crappy, but we have three engineers already who are great with it. But if you know Ruby and we don’t, cool.” That's a strength, not a weakness.

Hire for strengths. Forget the, “I can come up with some puzzler problem to put on a whiteboard that'll stump you.” Hell with that. Show me what you're best at. I want to see you shine. I don't want to see what it looks like when you're sitting there flailing because you haven't brushed up on your CompSci curriculum in 20 years.

Paul: Oh, God. Absolutely. I was very pleasantly surprised—as an aside when I was interviewing this last round, and I joined Honeycomb about a year ago—I did a pretty extensive job hunt, and I ended up doing a fair number of on-sites, I think it was like six in total, which seems exhausting now just thinking about it. But I was so relieved that no one had asked [00:18:07 crosstalk]—

Corey: Six conversations or six different trips to San Francisco to visit them on-site?

Paul: Three of them were trips, two of them were remote, and one of them was local.

Corey: Okay, those are actual separate interviews.

Paul: That's right.

Corey: With different folks at different times. Okay. Yeah, that's a lot of back and forth.

Paul: It's a decent amount. But I was so pleasantly surprised to see that nobody asked me one of those whiteboard questions. Not a single thing that would show up [BLEEP] Cracking the Coding Interview, or LeetCode, or whatever other tool you want.

Corey: Yeah, part of it is also just this—it's almost corporate hazing sense. It sounds weird, especially given that, let's be honest here, most of the audience of this show has an engineering background, but I personally find hiring folks who are either engineers or engineering adjacent to be way easier than a lot of other hires. For example, if I'm hiring another cloud economist who needs to be able to delve into AWS and have some SRE experience, and be able to look at this from a financial analysis perspective, great. I've done a lot of that myself. I know exactly what to look for, what to ask what to uncover.

Whereas if I'm hiring for, I don't know, a product marketer, or an accountant, or a graphic designer, I have no earthly idea how to even frame the question. Part of the challenge, then, is that in many cases, if you're not reaching out to experts who are great at this stuff to help with the winnowing and interviewing process, you're probably going to wind up hiring the person who sounds the most confidant, which is kind of awful.

Paul: Right. Exactly. I think the only thing I've ever found that can even begin to crack that for me, is ask people what they've done and then delve into really, really specific follow up. If somebody comes in and says, “I'm great at X,” great. Tell me about a time when you use x to a good result.

And obviously, you're going to run into people who are just really good at self-selling, but I think if you ask enough follow-ups and if you look for things like communication skills, their ability to connect their effort with outcomes and things like that, you can still get pretty good results.

Corey: This episode is sponsored in part by ChaosSearch. Now their name isn’t in all caps, so they’re definitely worth talking to. What is ChaosSearch? A scalable log analysis service that lets you add new workloads in minutes, not days or weeks. Click. Boom. Done. ChaosSearch is for you if you’re trying to get a handle on processing multiple terabytes, or more, of log and event data per day, at a disruptive price. One more thing, for those of you that have been down this path of disappointment before, ChaosSearch is a fully managed solution that isn’t playing marketing games when they say “fully managed.” The data lives within your S3 buckets, and that’s really all you have to care about. No managing of servers, but also no data movement. Check them out at chaossearch.io and tell them Corey sent you. Watch for the wince when you say my name. That’s chaossearch.io.

Corey: I think you're probably right. I think that there's a lot to be said for digging into things. What I love is asking open-ended questions in interviews. And at some point, one of us is going to get to a point of, “I don't know.” I'm either learning something, or I'm seeing how people think and what they do when they hit a wall which, especially for senior roles, is incredibly important. You don't want folks who are going to sit and not go anywhere, it's, “Great. I'm blocked. How do I resolve this? What do I do?”

Paul: Yeah.

Corey: And in my case, it's reach out to people, look on the internet, do some searching, but don't sit there and stand at the whiteboard and tear up. It's one of those, yeah, we don't know these things off the top of our heads. No one does. So, ask. That's the point. I want to see people saying that they don't know how to do something.

Paul: Yeah. And this is one of the hardest things to do, but when you do manage it—and I don't have the perfect answer, but I've seen it—when you get some kind of collaboration happening in the actual interview, and you get a sense of, “Oh, my gosh, this is what it would be like working with this person because we're actively collaborating on a problem that none of us know the actual answer to.” In other words, what we're paid to do, day-to-day. [laugh].

Corey: So, to that end, I have to ask you, given that you see a lot of this, what makes writing slash shipping slash producing software harder than it needs to be?

Paul: You know, I think there's a few different things there. Writing and communicating, I mean, that's hard because you're dealing with human beings. And to our previous discussion about software stacks, and tools, and tech, and processes, there's no perfect answer. And so the hard thing is figuring out, what do you actually need to communicate? What do you actually need to do?

In terms of shipping software? I think that that comes from making it harder than it needs to be by creating situations where you're scared to touch anything. My background as an SRE, the thing that always terrifies me the most is the service or the software that people don't touch very often. It's the stuff that, maybe it's harder to find out how it works, or how it breaks or whatnot because, frankly, you just never have a need to interact with it. That's the stuff that really scares the crap out of me.

Corey: From your perspective, I guess, what's the interesting part of software versus what's the part of it that no engineer should ever have to touch? Or do again? What is the valuable part? What should engineers of the future be building, focusing on, working on, and what should folks never think about again? I like the fact that you're coming from an engineering perspective because normally if I ask questions like that, it turns into a sales pitch answer.

Paul: Yeah, [laugh] exactly. It's funny because I find myself kind of conflicted here, between what I like to do and what I believe to be actually correct. And what I mean, there is, I like thinking about all the plumbing that makes software go; I like thinking about infrastructure, and I like thinking about writing tools and helping create things that make it easier for other software developers to push code to production, and help users, and delight users, and all that sort of thing. And that's exactly what most businesses shouldn't have to worry about. They shouldn't have to employ people like you, or I—from ops backgrounds—who just know how to make the stuff go because that should just be a given.

I was talking earlier about serverless and some stuff that I think is hopeful there. The average software developer, I think, who wants to delight users, who wants to create things that create value for a business and for customers, they don't want to care if it's running on Kubernetes or if it's running on Spot instances or things like that. They just want to push it, and they want to go. The tricky part comes in when it breaks. And when it breaks, we want something that we have that sort of ability to introspect and debug, even if it's hidden behind some kind of abstraction. And that's a balance that I don't think we've seen yet in the industry. But I think we're getting closer.

Corey: Well see, when I have conversations with folks like you, and we discuss these types of things, and the answers always seem so eminently reasonable, and then I leave the ivory tower of my podcasting studio and go back into the world, and then I see the nonsense everyone's building instead. It feels like on some level, there's two worlds: the aspirational way that we all want to be doing things, and then the messy way that we really are doing things. And I’m starting to despair of ever being able to fully bridge that gap.

Paul: Oh, interesting. By the ivory tower, what would be an example of an ivory tower perspective or point of view?

Corey: Oh, sure. Any conference talk you've ever seen on any technology under the sun, where they talk about how they wind up seamlessly deploying software into production. CI/CD stories, for example, are notorious for this. It's the—you watch these amazing presentations like, “Wow, I'd love to work in a place that did things like that.” And the person next to you says, “Yeah, me too.” And you look at their badge, and they work at the company the presenter works at.

Paul: [laugh].

Corey: It’s the myths, we tell ourselves. Sometimes individual groups wind up solving these problems within larger companies. Sometimes it's a new thing that they're running in test but haven't rolled out everywhere and, let's not kid ourselves, if it touches the payment system, everyone's doing waterfall development whether they admit it or not. But there's a broader world out there of folks who want to be doing things the right way, they want to be getting rid of the boilerplate and stop reinventing the wheel and re-implementing the wheel and get on to doing the truly interesting and innovative stuff. And those people right now are also listening to this while going back to code a login page. You never get past it on some level. That's what bugs me.

Paul: And you know what's super interesting about that? In my experience, which may not be representative, but the places that I've seen that have accomplished the closest to that kind of story, have done it in really simple, almost kludgy ways. And what I mean by that is, like, I've never personally worked somewhere where we had this great system that tracks state of all of these different services and made sure that there is, like, traffic going from here to there in a way that was canary testing and everything. You know, that all sounds like a lot of moving parts; the best places I've worked have a freaking cron script that just pushes out changes or has a webhook that kicks off something that pulls down a tarball from an S3 bucket and then ships it to a machine. Oftentimes this stuff, I think it doesn't make for sexy conference talks, but it's just roll up your sleeves kind of work to get it happening, and then move on to something else. I think sometimes we maybe trip ourselves up wanting it to be more interesting than it actually is.

Corey: That's part of it. If we were completely honest with people at what we were actually building or working on at any given point in time, the answer would be incredibly depressing and we would just be sadder after explaining our jobs to people. I try not to give talks to classrooms full of schoolchildren anymore on what I do for a living for that specific reason.

Paul: And yet—I agree, but isn't it great sometimes that this shit works? If the point is to deliver value to customers quickly and efficiently, maybe investing just enough to make that work repeatedly and in a way that people trust, and frankly, is simple enough that you can also debug when it doesn't do the thing that it's supposed to do, maybe that's actually enough. Maybe we're sometimes overinvesting in complicated solutions that might fall into that accidental complexity scenario we were talking about earlier.

Corey: Well, all right. Let's take that to its logical extent here. Here's something I know for a fact you have an opinion on. Now, I have opinions on things, too, which would surprise no one who listens to this show, but what do you think stops engineers from wanting to be on-call for the service that they work on? Now, there are a couple of answers I have to that: one is the polite public answer, and one's the real answer. But I'd like to hear your answer.

Paul: Yeah, sure. I want to come back to the difference between the public and the private answer, too. My answer is, there's a whole bunch of stuff. One of them is social—and I think this is more common—is that engineers are on call for things, and they don't feel like they have necessarily the autonomy to actually react to things the way that they need to.

And what I mean by that is, like—I've certainly seen this, and I've done my part to try to fix it, or to encourage others to—power people to fix it, or whatever the hell you need to do but, people get paged and they're like, “Oh, that alert means nothing.” “Okay, so get rid of that alert.” “I can't do that.” “Why not? [laugh]. Just do it.”

If you are getting paged for something, you have the right to change the system that is alerting you to something. And what you're seeing on the ground level, as the on-call engineer, should be gospel; it should be the thing that dictates how the future person experiences that role. And if you don't have that, it's a really shitty experience.

I think the other thing—that's technical—is this notion of accidental complexity, is when you have a system that you're responsible for, that you're on call for, and it's just, whether it's because of over-engineering, or it's just out of necessity complex, you don't know how to insert yourself when it [BLEEP] up, right? Like you can start to look at it and say, “Okay, we've got a drop in traffic,” or, “We've got a spike in error rate,” or something like that, but if you're nervous about the actual mechanics that get your changes from your laptop to the production environment, then it can be a really terrifying experience to make changes. And I've been in environments where people just freeze, and it sucks. So, that's why I always think of, like, if you can make it easier to get the changes from your laptop to production, that is the best investment that you can possibly make, technically.

Corey: I would agree with that sentiment. It feels like when you talk to software developers who are building these systems and then complaining about a problem in production. “Here, log into the prod server and see.” Well, this looks nothing like their IDE; it looks nothing whatsoever like their development environment. And people feel awkward and out-of-sorts there.

I mean, I intentionally in years past when I was working in ops roles made production uncomfortable to work in intentionally so, because that's not your default place to operate in. But if people are used to using Visual Studio Code, for example, then, “Okay, now the only editor we have installed here is VI, so you're going to have to spend some time learning, even to look at what's going on here.” That's an awful experience, not to mention that people are never doing these things during the workday, invariably. It's always two in the morning when you're bleary-eyed and have no idea what you're doing. And, congratulations, you're being confronted by the Puzzle Master. It doesn't go well.

Paul: No. And that's actually a great point that I think is within our control as engineering teams to change. Yeah, it'll happen at two in the morning, that's for sure. Any 24/7 service that you're on call for, it's going to break at an uncomfortable time, and you're going to have to debug it, but that doesn't have to be the only time you do this stuff. And in fact, when it is the only time you do this stuff, that's terrible.

And that's why I'm a big proponent of have fire drills, have game days, break the shit that breaks often so that it breaks when everybody's around. Resolve those kinds of uncertainties as much as you can—because, obviously, some things are just unknowable—but practice those muscles as often as possible. There's a funny thing that we talk about sometimes at Honeycomb, and this sounds like a humblebrag, but it's not—it's just that there are periods of time where we don't have as many incidents, and that makes it actually really hard to make sure that people are primed to be on-call. And so we're thinking through what can we do to just make it more comfortable? Like if someone comes on board, and their first on-call rotation is quiet, that doesn't really help them. So, what can we do to, kind of, force interaction with production as often as possible to make it almost routine and muscle memory?

Corey: I talk with companies back when I was looking at various roles where, “Oh, everyone is on call.” And you hear that during an interview, and having been through many on-call rotations myself, it's, “Yeah, that's not a strong point, to be perfectly honest with you.” That sounds like, if you're not very careful how you position this, that everyone is woken up for every incident, and I won't get a whole lot of sleep working here, and not to be unkind, you're not paying significantly more than other folks who don't subject me to that.

Paul: Yeah, that's terrible. Everybody is on call. That reminds me of the companies that I worked for… I don't know, before a certain time when, I don't know, maybe it was the PagerDuty became a ubiquitous tool that was used in companies of a certain size. But it was that old time when the first person to respond is really the person who's on call. And that's a terrible environment, and that's a recipe for burnout.

You should have a clear escalation path, and you should have clear responsibilities, and every engineer should have a huge chunk of time when they're not on call, and they know that they're not on call so they can delete Slack from their phone; they can turn off all of their alerts. And when they're done at the end of the day, they're just done. So, yeah, I would also run screaming from a company who said that, these days. Anecdotally, there's two questions that I always like to ask companies when I'm looking for jobs, and talking to companies.

One is, “How do you get your code to production? Walk me through as many steps as you're comfortable disclosing in an interview,” which hopefully is a lot. And two is, “How are people put on call? And what's the last major incident you had, and what does it look like? Who was involved? What happened? How did the people get the support they needed?” All those questions are really, really interesting ones to dive into. I wish you had more opportunities than interviewing to ask other companies this stuff. [laugh].

Corey: Oh, yeah. My personal favorite way of responding to that, which is why I generally don't get offered jobs, a whole lot is, “So, you have an on-call rotation here?” “Oh, yes. It's absolutely critical that our site is up all the time.” “Cool. So, why don't you staff multiple shifts of people who are responsible for keeping the site up during those times so that you're not making people wake up in the middle of the night to break things?” And suddenly we're in one of those, what I say versus what I do are different territories. And that becomes a problem.

Paul: Oh, are you talking about, like, follow the sun rotations?

Corey: Either follow the sun or having folks who are either night owls who enjoy night shifts or something for—and I'm not talking small startups here. I'm talking companies that have 1500 engineers working there. It's at some point you have multiple offices in various places. Why are you still waking people up in the primary timezone every week?

Sorry. When I say primary timezone, I should be very explicit on this: there's always a timezone hierarchy in every company. It's the headquarters time, and that is how it's going to be regardless of what companies claim otherwise.

Paul: Of course. It's the center of their universe.

Corey: And invariably, it seems to be Pacific West Coast.

Paul: Yes, exactly. Yeah, I agree completely. At a certain size, and there are plenty of companies that I think are doing this, but you have the opportunity to let folks in North America time zones, just stop working. And then folks in European time zones will take over. And then folks in certain Asian time zones will take over. And yeah, that is a great way to do things, I think, if you can manage it.

Corey: So, I guess my last question for you, since I've been peppering you with these—is if people want to learn more, where can they find you?

Paul: I am on Twitter. I'm not sure how much value you'll get from my tweets, but every once in a while, maybe I'll tweet something that at least provokes some discussion: @paulosman. And I very occasionally blog at paulosman.me. And I think that's it.

Corey: And we'll put links to those in the [00:35:44 show notes].

Paul: Excellent.

Corey: Paul, thank you so much for taking the time to speak with me today. I really appreciate it.

Paul: No problem. Corey. I really enjoyed the discussion. Thanks a lot.

Corey: As did I. Paul Osman, lead—or lead—engineer of instrumentation at Honeycomb. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you've hated this podcast, please leave a five-star review on your podcast platform of choice and a comment telling me of why you're on-call rotation is different and unique.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Brandon ShawBrandon Shaw is a senior program manager in security operations at Discovery, Inc., an entertainment company that owns several premium cable brands, including Discovery Channel, HGTV, Food Network, and TLC. Previously, he worked as an applications manager at CRISP and a senior software applications engineer at CompuGroup Medical, among other positions. Brandon has a slew of certifications, including CISSP, CISM, CDPSE, CCSK, PMP, ITIL, and three from AWS.

Links Referenced:

  • Discovery, Inc. Website
  • Connect with Brandon on LinkedIn

TranscriptAnnouncer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Linode. You might be familiar with Linode; they’ve been around for almost 20 years. They offer Cloud in a way that makes sense rather than a way that is actively ridiculous by trying to throw everything at a wall and see what sticks. Their pricing winds up being a lot more transparent—not to mention lower—their performance kicks the crap out of most other things in this space, and—my personal favorite—whenever you call them for support, you’ll get a human who’s empowered to fix whatever it is that’s giving you trouble. Visit linode.com/screaminginthecloud to learn more, and get $100 in credit to kick the tires. That’s linode.com/screaminginthecloud.

Corey: This episode has been sponsored in part by our friends at Veeam. Are you tired of juggling the cost of AWS backups and recovery with your SLAs? Quit the circus act and check out Veeam. Their AWS backup and recovery solution is made to save you money—not that that’s the primary goal, mind you—while also protecting your data properly. They’re letting you protect 10 instances for free with no time limits, so test it out now. You can even find them on the AWS Marketplace at snark.cloud/backitup. Wait? Did I just endorse something on the AWS Marketplace? Wonder of wonders, I did. Look, you don’t care about backups, you care about restores, and despite the fact that multi-cloud is a dumb strategy, it’s also a realistic reality, so make sure that you’re backing up data from everywhere with a single unified point of view. Check them out at snark.cloud/backitup.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Brandon Shaw, Senior Program Manager of information security at Discovery, Inc. Brandon, welcome to the show.

Brandon: Thanks for having me.

Corey: So, you’re at Discovery, not Discover Financial. I keep running into those companies in various cloud stories, and I always get the two of them confused. So, you're the one where I talk to about Shark Week, not credit cards, right?

Brandon: That's us. Correct.

Corey: Fantastic. And we'll get into that, I'm sure, but first, a little backstory here. I think at this point, we've known each other for, what, 26 years.

Brandon: Yeah, something like that.

Corey: I mean, at this point, our relationship has grown up, gone away to college, come back, moved into our basement, gotten through the surly stages, and is now trying to venture back out into the world, but there's a pandemic on, so it's having some trouble.

Brandon: Yeah, you forgot the part about both of us getting married both of us having kids, both of us—

Corey: That's right. We're both freshly back from paternity leave—

Brandon: That's right.

Corey: —as of the time of this recording.

Brandon: [laugh]. Yeah. And congratulations again.

Corey: And to you as well. We both had our second kid. At some point, it's weird, our lives sort of diverged for a while, and then wound up going back into something that resembles, “Oh, I could actually justify having you on the show now. Great.” But it's amazing how much our lives really have, I guess, split apart and then reconverged.

It feels almost like it's a story about technology. I mean, as soon as we use the word converged, I'm sure Nutanix somewhere is perking up, “Oh, hyper-converged? That's our word you owe us $1 for using it.” But I don't think they're sponsoring this episode. If they are, oops.

Brandon: Yeah, so I wanted to jump in, just for all of your friends and fans—

Corey: Both of my friends, let's not over-exaggerate the case.

Brandon: Excuse me. So, I want to jump in and explain to both of your fans this week what Corey was like, as a kid.

Corey: Oh, dear God, here we go. I feel like I've been bamboozled.

Brandon: You have been. What I want to say is, the reason why we've been friends for 26 years is that you have definitely been the guy—and always have been—who is not afraid to be your most authentic self.

Corey: Oh, that's a very kind way of putting it. Thank you. In practice, it's that I have no filter and no social skills. The only challenge was living long enough to evolve it into a way where it was at least halfway socially acceptable. Let's not overstate that.

Brandon: Well, I mean, that was really the thing. I mean, I always appreciated you and your sense of humor, but—

Corey: I think you were the only one.

Brandon: [laugh]. Still—maybe. But a lot of people when we were growing up, they were just like, “Oh, Corey. That's the same tired joke for the 15th time today. Just give it a rest.” But for you, it's still funny. And what, you never give up on those jokes.

Corey: I never do. My jokes are for me. And if other people like them, that's great. And if they don't, well, get your own podcast.

Brandon: So, I will say that I'm really glad that you've been able to find this forum for yourself as a place to have people listen to you, whether they want to or not.

Corey: Exactly. The nice part about podcasts as a medium is that people tend to listen to these things with headphones on when they're washing the dishes or mowing the lawn, which means that they generally find it very inconvenient to stop the podcast, so they're really forced to listen to me. It's a captive audience model more than anything else.

Brandon: Well, absolutely. And so it's different than being in the middle of French class when you're trying your jokes out. But certainly to, again, both of your fans, whether you find him funny, charming, abrasive, whatever, you are the same guy that I met, like 26 years ago. So—

Corey: Uh-oh. Well, let's talk a little more color on that. Why not? It's my show; I can talk about whatever the hell I want.

Brandon: Sure.

Corey: The problem that I had was that my dad was one of those folks who always wanted to be somewhere else, doing something other than what he was doing. So, I come by it honestly. In his case, that manifested as being a bit of a nomad. So, midway through the seventh grade, I wound up moving to Maine. State motto:, “Not a lot of people come here on purpose.”

And because I was a lucky child, I was midway through a series of various lung surgeries, so I missed a hell of a lot of school—I wouldn't say I was missing it, Bob—and I was socially stunted. I was always having to make new friends, which I was terrible at—still am. And it was this weird dynamic, where you were one of the only people at the time who said, “Hey, look at you, you're a sad, lonely loser. Me too. Want to hang out?”

And it was really something redeeming about that. I mean, despite the fact that our lives have taken us in very different directions, we live on opposite coasts, et cetera, it's always been like coming home whenever you and I have a conversation. “Oh, it's been a year since we last spoke. Great, let's play catch up for 20 minutes.” And then it's as if no time had passed.

The other side of it, and it's a bit of a challenge on a show like this, is whenever I talk to you, suddenly I revert back to being that angsty 12-year-old who needed to go and prove himself, and is terrified that no one really likes him. It's the reason I've never had my mother on the show.

Brandon: I thought that we already covered that before. You still are the same angsty 12-year-old.

Corey: Absolutely. Except I'm larger now and don't have the metabolism that I once did. So, you and I met where we spent our formative years respectively, in a small town in Maine—because it's not like there are large towns in Maine—and some of the best things that we ever achieved on a personal triumph basis was leaving Maine. Now, I know this annoys RedMonk’s Stephen O'Grady, who lives in Portland and loves Maine, but I spent enough years there to say that growing up there, it sucked. I don't miss it in any sense. What's your take on it?

Brandon: Well, let's first look at the numbers, right? And that's the biggest part is Maine is number 37 out of 50 states for economical opportunity.

Corey: It was also the second least diverse state in the country as well.

Brandon: And least diverse state in terms of religious diversity. As well as most importantly, Maine is number 42 out of 50 states in infrastructure and number 45 out of 50 for internet access. People like us today, now, cannot live in Maine. If you enjoy technology, and you want to work from home, chances are you can't if you're living in Maine. We have so many friends who are waiting for 5G to come to their forest. It's not happening. It's not happening. Besides that, there is no opportunity. It's cold. It's awful. The only thing that you can do in Maine is sit around in the cold and develop your personality disorders.

Corey: Yeah, exactly. I dated a girl who was valedictorian of her chemical engineering class, and she was on the front page of the Maine Sunday Telegram—which you can probably already tell is the largest and only newspaper in Maine—and the whole story was about how she had to leave the state to get a job. Every year, the governor goes to the University of Maine and does a whole commencement speech about how—stay in Maine. With what jobs? It's impossible—at least the time that I lived there—to effectively grow yourself professionally living there compared to the opportunity available in other places. And maybe that's changed. It's been 20 years. But I kind of question that.

Brandon: Well, University of Maine's biggest son is Stephen King, who is probably the only person that I know who's made a name for staying in Maine. And anybody on Twitter can go ahead and tell me wrong, right. But look at what he's doing.

Corey: Yeah, Anna Kendrick grew up in Cape Elizabeth, the next town over from where we were, and how did she succeed? She got the hell out of Maine.

Brandon: Yeah, but let's go back to Stephen King for a second. What is he known for? Sitting in the dark, writing things that scared the shit out of you.

Corey: Yep. And all he’s really doing was telling the stories about his typical week.

Brandon: Yeah, this is what it's like living in Maine. Get out. In all of his stories, there's something about trying to get out and you can't do it.

Corey: That is the best synopsis of living in Maine that I can possibly imagine. One of my favorite Twitter accounts—I know you don't do it—is @maine_gov, where it's a parody account where it winds up crapping all over Maine, to the point where the governor's administration had to respond with, “This account and its posts are not affiliated with the state government.” And everyone knew that because it had a personality.

Brandon: All right, that may actually get me on Twitter. You might have sold me there.

Corey: Exactly. That's all it takes. So, the audience generally knows what I've been up to, but let's talk about you. You've been at the Discovery Channel, slash Discovery, Inc. Slash Shark Bait for a while now.

Brandon: Yeah.

Corey: I mean, you want to talk about ways that we're not the same, you've had a gift of being able to be at the same place for multiple years without either getting bored and leaving, or—in my case, more often—getting yourself fired. So, it's a different world, one I'm deeply envious of, but you've been at Discovery for a while, what do you do there?

Brandon: So, my job here is really as the senior program manager for the operations arm of our cybersecurity team. So, it is my role to coordinate incidence response activities and remediation efforts across the entire global enterprise. It is also my job to stay engaged with internal and external teams and make sure that they can understand and can attest to all the appropriate security measures and best practices that we provide to it.

Corey: Now, I'm going to say the quiet part out loud because that's what I specialize at. You’re information security; convince me that you're not effectively ablative armor for the company where your job is to sit around and wait for the data breach, and then be ceremoniously fired to protect other executives with I would say, other loftier positions, or—let's be honest—more political savvy, change my mind.

Brandon: So, that's absolutely fair. I'm not going to say that I am not on the chopping block next time something happens. But really, I will say that Discovery has put itself in the line of fire with a lot of our acquisitions, with our footprint, with everything that we do. We have a massive, massive company.

And the truth is, is we don't want to be that same media company that, back in 2014, hackers accessed and wiped personal data from 10s of thousands of employees and users. And that's just who we don't want to be. So, it's my job to actually make friends with everybody across the environment and to bring them on instead of pull them along, kicking and screaming, to understand why we need to have our security best practices.

Corey: One of the challenges I've always found with infosec is ignoring the trash fire of a community that most of it seems to turn into, invariably, it's been this idea that I deal with on the cost side as well, though at somewhat lower stakes, where people only really seem to care about either cost or security, right after it really would have been beneficial for them to care about those things. So, right now, I'm in this weird space where, at least on the cost side, if people don't optimize their cloud bill, and then they have to bring me in, well, the cost there as they spent a little bit extra than they should have until I'm engaged. On the security side, it feels like there's always an announcement about, “Your data is extraordinarily important to us,” is what companies say, always right before they announce the data breach that showed that security was clearly nowhere near as important to them as it should have been. It always feels like it's a back foot thing that can always be punted until suddenly it bites you in the face. How do you get away from that reactive mindset?

Brandon: That is a great question. And I will say first that, as a program manager, that I work for a great team of people who are futurists and enjoy learning about new technologies, and especially not just use but abuse cases for all of the opportunities that we have. And so, staying ahead of security is really the ideal goal. But really it is about, in my opinion, controlling surface area, and knowing whether you're leaving the back door open, whether the door is open at all, and making sure that you close it. But the truth is, is that information security, it's always moving faster than we think it is.

And especially if you have somebody who is very highly skilled and motivated, they're going to find a way. And I mean, that's the ugly truth about security. So, instead of us saying that we're doing this for lip service, and I mean that pejora—excuse me, instead of just saying that we're doing this for lip service, we mean it. We're not going to be a company without our customers, we're not going to be a company without our product. And it is our job to serve our customers and our content in the best way possible.

Corey: And credit we're due, there's a lot of companies out there where they get it wrong, and people love the twist the knife in them. I don't know too many people who have negative associations with Discovery as a whole. Maybe the Bloodhound Gang song from years past, but that's about as far as it goes. There's something to be said for being a household name, but not pissing people off as you do it. Meanwhile, I'm not a household name, and I still piss everyone off as I do it. So, back when we were going to school together, you were focused on computer engineering. It seems like there's a few steps between that and getting into information security, how did you get to where you are?

Brandon: So, I think that you'll recall that when I got into college, I actually hated computers. And getting out of college, I didn't want to do anything with computers at all. I think I've read some statistic out there that 10 percent of people who graduate with a degree do something tangentially related to their major. And so, my goal was to be part of the 90 percent because I was just burnt out. But after I took a little bit of a detour and taught English in Japan for a couple of years, I came back and realized that I actually missed it.

And so my first real-person job, if you will, was as a software engineer, and it took me about two weeks before they said, “So, what do you think about project management?” Because I am the worst programmer. But just being close, just being in the Washington DC area, I thought that there would be so many opportunities to get into cybersecurity. So, that's what became my graduate degree, in information assurance. And so I just kind of rolled with that. It was really kind of organic; just looking at the opportunities that we had in the area—or that I had in the area, excuse me—and taking advantage of everything that I could.

This episode is sponsored by our friends at New Relic. If you’re like most environments, you probably have an incredibly complicated architecture, which means that monitoring it is going to take a dozen different tools. And then we get into the advanced stuff. We all have been there and know that pain, or will learn it shortly, and New Relic wants to change that. They’ve designed everything you need in one platform with pricing that’s simple and straightforward, and that means no more counting hosts. You also can get one user and a hundred gigabytes a month, totally free. To learn more, visit newrelic.com. Observability made simple.

Corey: How do you work in infosec, without becoming profoundly paranoid? And that's not entirely a tongue in cheek question. Something I've noticed about my friends who've gone down that path is increasingly they start to view everything as this hard divide of everything is a potential vector for being exploited, and it starts to color their personal relationships as well. Like, we all know the types, the folks who email you, but you don’t know what it says because they GPG-encrypt the thing, and who can be bothered to decrypt it in the modern era? You've never gone down that path. Is there a philosophical reason behind that? Are you just a terrible infosec person and no one ever pointed it out before? What's the answer?

Brandon: Well, I will say that my role as a project manager, as a program manager, as the leader of the team, really kind of brings a different perspective than the regular infosec person. And yeah, it's really easy to go to the place where the sky is falling, and nothing's right, and the whole world is burning, but that's a really hard message to sell. Nobody wants to hear that. Nobody wants to be presented with a problem without understanding what you can do to fix it. And so that's my job.

In part, it really is a sales position. I am not—nobody's buying anything here, right? Every contact is an opportunity for me to educate people. And that's really where my passion lies, is learning new technologies and understanding new use cases, and being able to say, “Hey, wait, this is really cool, but maybe we want to pump the brakes.” Or, “This is really cool, and let's go for it. I don't see any problems with this.”

And so, I see infosec as enabling technology. It's really easy for an infosec to say, “No/ whatever it is, no, you can't do that. We're not going to let you.” Please submit a form and we're going to [00:19:52 unintelligible] again in 30 days.

Corey: Oh, yeah, you have infosec, sysadmin types, and engineers. And those are the three real points on the spectrum of, “No.” Where engineers are, “Sure, we can do that.” A sysadmin is, “Neeaah,” and infosec is, “No,” because it's easier. At some point that doesn't work anymore.

You have to come up with something different. The amazing part of all of it to me is just that there's so much that can be done if you work collaboratively with the business. And historically, a lot of folks who got into infosec didn't get to do that. And when I dabbled in it, the thing that really disillusioned me with it—besides the crappy people—was that it was less about breaking into systems or defending systems against being broken into than it was about time to fill out some forms for an auditor. And policy and governance are less exciting than—you know, as all the movies show, wearing hoodies, and gloves, and a mask to type into a laptop at a dark street corner somewhere. Which, of course, so we all write code, right?

Brandon: Yeah. I mean, who doesn't want to drop from the ceiling and break into a server that is guided by laser sensors that you can’t stand around? That sounds really cool.

Corey: That feels like a problem that someone at AWS winds up worrying about every day of their lives.

Brandon: And why shouldn't they?

Corey: But, I pay them to worry about that so I don't have to.

Brandon: You're right; you're absolutely right. It is paperwork, and it is a lot of reading, and it is communicating and collaborating. And I think that people who get into infosec just because they love the power, and they love to say no, are really probably the wrong people to be in such a friendly and well-liked company, as Discovery.

Corey: So, recently, you wound up adding some letters to your name. Now, that doesn't mean that you went out and got a doctorate, it doesn't mean that you've decided to retcon yourself into being the fourth person in your family line. But rather, you got the CISSP, which sounds an awful lot like it's a word I don't want my daughter to hear. So, I'm spelling it to you. What is it really?

Brandon: CISSP, it stands for the ‘Certified Information Systems Security Professional.’

Corey: I looked into it once upon a time because I have my flaws, but one of my strengths is that I test well. So, sure I can sit down and take a test. Turns out in order to even qualify for the exam, you have to have done a fair number of things that I'd have to really stretch the truth to qualify for.

Brandon: Yes, and I will actually say that I have heard stories where people have used upgrading their own personal laptop, as five years of experience for the CISSP.

Corey: Well, that really depends on their laptop, now doesn’t it?

Brandon: I'm just saying, right. But the experience that you need is really, really kind of vaguely worded, and you are looking at eight domains where you need to be able to attest to just a couple of them. So, yeah, you can be a mail admin, or you can be a network engineer, or you can just do a lot of paperwork for an auditor, and you can still attest to five years’ experience towards the CISSP.

Corey: So, I used to have a very negative approach to certifications in general. I thought that they were a waste of time, that I would be much better off having built something myself and putting that on the resume instead. And it took me a long time to realize that that was a very naive perspective on a couple of levels. One is that everyone learns differently. And there's an awful lot of privilege that gets baked into the well just do it for 10 years, and then it's on your resume, it's easy to say from the other side of that.

When you're starting out, it becomes valuable. And from the hiring side as well, it's convenient because I know that if someone has a certification in technology, that does there's at least a good chance that we can have a conversation around the topic in question using localized terms of art without having to stop and make sure that we're on the same page. Now, the other side of it that I haven't experienced a lot of is you're a giant company and you're trying to hire 5000 people with cloud skills, or security skills, or whatever it happens to be. You need a formalized training and testing program if you're uplifting your entire existing staff skill set, who is where on their path. So, certification programs in that sense, also tend to add significant value. So, my old take of certifications are garbage and it's a ding against you was immature and frankly, wrong.

Brandon: Well, it's an easy joke to make to say, “I have 20 or so letters after my name, and so I am just making it really easy for HR professionals to find me in a crowded pool of applicants.” But from my perspective, I mean, I live in the DC area. There's consultants everywhere, and where I live, and it's very much about accumulating letters and displaying all these certifications on a LinkedIn resume. But really where my perspective is, is I'm a project manager. I've been doing this for 13 years now.

But really, what does this mean? Project managers, they serve as a jack of all trades, but a master of none and so I will be called on to fit a variety of technical leadership roles based on the nature of the work, or whatever I want to do. But as soon as I say something dumb, like, “Let's go install the agent on the lambda function,” or, “Why don't you just reboot S3?” Is going to lose all credibility and the project is going to fail. So, it is really, for me, a measure of my success and my capacity as a project manager, not just to speak to these pieces and to speak to all of the letters behind my name, but t0 also not be the dumbest person in the room.

Corey: That's my job.

Brandon: And you do a great job at it. [laugh]. But truly, the best bit of career advice I can say, is, whatever your opinion is on these certifications, I feel that they are a good measure to understand what it means to have a general level of understanding on a topic. And it's also good to have external validation, to prove that I am knowledgeable in a particular subject area. And so it is really working to my advantage not to be the dumbest person in the room.

Corey: I was doing a webinar thing recently, where there was a Q&A section and someone asked, “If I know nothing about security, what do I need to know to get started so that I can wind up building something in the cloud without blowing my own foot off?” And the answer was, “Oh, God, I don't know.” What's your answer to that one?

Brandon: That's actually a really tough question. And I'm going to punt on that one.

Corey: Cool. No, that's fair because I don't have a good answer, either. It's, uh, “Yeah. Oh, God, if you don't know where you're going, where does it start?”

Brandon: “I really just love auditing. How do I get into auditing?” I just don’t know that.

Corey: Yeah. It’s like, “I like computers. How do I get into those?” “Ahh.” It's such a big topic. But it's also a problem because you don't want people to have to go through eight years of security school to be able to write a hello world.

Brandon: No. And that's really kind of the point that I will say is kind of a knock against AWS, is that you really do need to spend, like, hundreds of hours, just figuring out everything that's out there. There is no taking a sip from the firehose with AWS; there is no dipping your toes into security configurations. You really need to dive right in and take the time to ingest all of it. And that's the biggest problem.

My biggest beef with AWS is that there is no guidance that they really do provide, especially for casual users or small startup companies that are not going to pay for a business tier support package.

Corey: That's part of it, too, but there's also this idea that what AWS says and what AWS does are two very different things. You have this idea of a shared responsibility model, which winds up requiring 45 minutes to explain and overly complicated graphs, the honest fact of it is, they handle the physical stuff; you handle the configuration stuff. If you get breached, it's probably your fault. The problem is, you can't tell the customer who just got breached that it was, “Because, your fault. The end.”

Wow. What do we do with the next 44 minutes and 30 seconds of, we sit here awkwardly? So, you need to have a more convoluted story there. And also, a lot of the decisions they made—which I don't blame them for—but the way that the world has evolved has dramatically set people up for failure. Here's S3.

It is simultaneously this public-facing static website service that works at a global scale super easily. It's also a place where you can store your deep dark, encrypted backup company secrets. And the fact that the same thing does both of those means that it's relatively easy to trip and make it do both of those things simultaneously, and that's where people get into trouble.

Brandon: I mean, absolutely. We have a culture where it's so easy to just point fingers. And I can't just write my password down, stick it to my monitor, and then get mad at Post-It for my account getting compromised. That's just not how it works. And when you partner with a cloud provider like AWS, you do get the ability to offload risks, like costs, and managing your own platforms and services, but you don't get to transfer liability or responsibility for this.

Corey: You can transfer work, but not the responsibility. That's the thing that always bothered me about breach disclosures where, “A third-party contractor did—” I'm going to stop you there. Look, I don't have a relationship with your third party contractor; I have a relationship with you. You've been a custodian for my sensitive data that I have chosen—ideally chosen. If you're Equifax, maybe not—but I theoretically have chosen to entrust you with, and your lack of vetting of your contractors is a problem.

To date, one of the best things I've ever seen was the Pokémon Company. There was a Wall Street Journal article on it where they declined to do business with a third party vendor based, in part, upon that vendor’s lack of security controls around how they interacted with S3 buckets. I thought that was phenomenal. If you don't validate that the people you're entrusting with customer data are doing things right, then that's not the vendor’s fault. That's your fault.

Brandon: Oh, absolutely.

Corey: Lord knows I filled out enough vendor security forms from the consulting side that it makes sense. I know what these things look like, and most of them are fairly reasonable. There are times it turns into maddening stuff like, “What kind of antivirus do you run on everyone's laptops?” It’s… “We are not that kind of company. I'm sorry.”

Brandon: No, and that's absolutely the point, though, right? I'm going to push back and I'm going to say that I want to have all that information. If I'm going to partner with you, if I'm going to give you my business, then I need to know that you are a great custodian. And if you can’t answer that—what antivirus you have on your laptops—because you don't know, or you don't care, or you don't do it, maybe you should.

Corey: So, what's next for you? You've been at Discovery for a while, you were an infosec project manager; now you're an infosec program manager, which I assume based upon no idea how real media companies work that you're in charge of programming. So, what can you tell us about the next Shark Week?

Brandon: Awesome question. No, program manager is just really kind of the man behind the curtain for this. It's similar to a project manager, but different letters that you can use for Scrabble. That said, what my next step is? Just hanging out. I'm happy to be at Discovery. We're doing great. I’m happy to be a part of such a technologically progressive company. It is so easy to knock Discovery for just being Shark Week, but we're massive. We're Shark Week, we’re the Food Network, we're HGTV.

Our entire portfolio is over 150 channels broadcast across 220 countries and territories, in over—I believe—50 languages, simultaneously. I have coworkers who have gotten Emmys. I actually sat across the desk from them in the cube farm. That is really cool.

Corey: I’m having a really hard time here just contextualizing this with the fact that we used to rotate around to each other's basements to play D&D 20 years back, and it’s… yeah, the world has changed. I guess we're still playing make-believe, but now the stakes have gotten a little higher. Which, you know, I'll take it. I look back at the confused kids that we were, I'm pretty happy with where we wound up all things considered.

Brandon: Yeah, I think that we're doing all right for ourselves.

Corey: We really are. So, if people want to learn more about who you are, what you're doing, what you're up to, and I guess, honestly reach out to you for inside baseball stories they can use to eviscerate me when I have them on the podcast in the future, where can they find you?

Brandon: I am only on LinkedIn right now. I do not do the social media thing, and I do not do Twitter.

Corey: I'm still learning how to shitpost effectively on LinkedIn. It's a process.

Brandon: Yeah, but no Twitter. I have kind of drawn my line in the sand that is the line that I will not cross.

Corey: I don't understand some of your life choices. But again, different strokes for different folks. Brandon, thank you so much, once again, for taking the time to catch up with me. It's always a pleasure.

Brandon: No, it's mine.

Corey: Brandon Shaw, program manager for information security at Discovery, Inc. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you've hated this podcast, please leave a five-star review on your podcast platform of choice, and tell me why you didn't like our episode here about Snark Week.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Kylie RobisonKylie Robison is a California State University, Sacramento student studying business information systems, technology reporter for the State Hornet, and proud president of her school’s Ski & Snowboard Club. She’s hoping to break into the technology industry when she graduates in May of 2021.

Links Referenced:

  • The State Hornet
  • Covered California
  • Follow Kylie on Twitter
  • Kylie’s Website

TranscriptAnnouncer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

This episode is sponsored by our friends at New Relic. If you’re like most environments, you probably have an incredibly complicated architecture, which means that monitoring it is going to take a dozen different tools. And then we get into the advanced stuff. We all have been there and know that pain, or will learn it shortly, and New Relic wants to change that. They’ve designed everything you need in one platform with pricing that’s simple and straightforward, and that means no more counting hosts. You also can get one user and a hundred gigabytes a month, totally free. To learn more, visit newrelic.com. Observability made simple.

Corey: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the Enterprise (not the starship). On-prem security doesn’t translate well to cloud or multi-cloud environments, and that’s not even counting IoT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IoT devices, detects these threats up to 35 percent faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at extrahop.com/trial.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Kylie Robison, who is currently a Business Information Systems student at Sacramento State, targeting graduation in May of 2021. Kylie, welcome to the show.

Kylie: Thank you for having me.

Corey: So, you do more than that, too. Student is sort of the identifier people slap on themselves. I get that, but you also are currently a technology reporter for the State Hornet, which is either a newspaper of some kind, or possibly California now as an official insect.

Kylie: [laugh]. Yeah, no, my school mascot is a hornet. So, I joined the State Hornet. They're a student-run news organization. But we don't have a newspaper anymore; we went online fully, I think a year or two ago.

Corey: Fantastic. It's definitely more eco-friendly, and the counterpoint is, is that when you publish something that people don't want to read, it's harder to steal all the issues when it's online.

Kylie: Exactly.

Corey: You also work in IT services for Covered California.

Kylie: Yes, I do. It's my first IT job. Before that, I was making sandwiches at Beach Hut Deli.

Corey: Wonderful. Do you ever miss the sandwiches?

Kylie: Oh, definitely. I was just there the other day. You know, service in IT or service for food is about the same anyways.

Corey: I look at what you're doing now, and I've been getting echoes back to when I started my career. I was doing a bunch of different things at once, in similar ways. And I still have this problem now, when people ask—like, we're stuck in an elevator. It’s, “So, what do you do?” And to give an honest answer, I have to pull the emergency stop and talk for five minutes.

It’s, what's the context—oh, like at neighborhood block parties, I tell people I'm some weird kind of accountant because I don't want to become the neighborhood computer repair person. That isn't going to go well. But it's hard because when you have so many different hats that you wear, it's difficult to nail down the, “What do you do?” Story.

Kylie: Yeah, definitely.

Corey: So, for those of us who didn't graduate from college, and even when I tried to and didn't go so well, that was still 20, 25 years ago. What is Business Information Systems? That was never a term that I ever saw back when I was failing out of academia?

Kylie: Yeah. I wanted to be a comp sci major, but I was afraid of how long I would be in college paying for my own tuition; I didn't think it was a viable route. So, I just became a business major, and I looked into the concentrations at Sacramento State and they had Information Systems, which I thought was a great route because business and comp sci together just sounded like a great opportunity. So, information systems in business. Right now, for example, my classes are Information Security, databases—so we learned SQL, stuff like that—what else, project management. So, it's just the intersection of both subjects I would say.

Corey: Oh, my god, that sounds wonderful. I was attending the University of Maine back in 2000, 2001. And the comp sci folks there were just—no disrespect. I get it. A lot of the universities struggle with this, but these were folks who hadn't updated the curriculum in 10 years, and they're talking about what the digital world looks like as the dotcom bubble is exploding around, and on some level in the back of your mind, you're stuck with, “If you know how this stuff works, why are you here, and not, you know, going to found the next MySpace,” because that was the thing people cared about back then. It became a very strange, like, a sense of dissonance. The idea that there's now a curriculum that blends that stuff rather than forcing people to forge it on their own, just sounds incredibly valuable.

Kylie: Yeah, definitely. I would say that getting a college education is a lot more valuable now in terms of comp sci, perhaps than it used to be. So, I was on Twitter, and I talked about how I had this Python course. And someone was like, “I would have never imagined learning Python when I was in college.” Which was fascinating to me.

Corey: Oh, that was absolutely never on the roadmap. When I was doing my comp site courses, they were focusing on assembly, which told me that, okay, I'm not a computer. Good to know. And they were focusing on Pascal because this newfangled thing called Java, which just didn't seem to really be where they wanted to go yet. Yeah, spoiler: I don't care how old you are, Java has never been a newfangled thing. It was born old.

But it was a challenge. I mean, it sounds here, like I'm sitting here complaining about back in the era of the invention of the wheel. But it was at the time, it was challenging; people had to stumble their own way through it. Now, it's not just that there are more programs out there, but those programs themselves have evolved to at least somehow embrace the modern reality that there are so many different paths that are either directly into tech, or heavily influenced by tech, or tangential to tech. I don't think there are too many paths that never touch tech at all, but I learn something new every day.

There are multiple paths to walk, and not having a degree myself for my 20s, I was incredibly defensive, and I was very dismissive about the idea of higher education. It took me a long time to realize I'm a hell of an outlier. If you have a degree, life is easier—in a career sense—than if you don’t. There's really no arguing that. I had to talk my way around it more times than I care to remember.

Kylie: Yeah. I think there's a lot of different paths you can take. And it just depends what works best for you. I know people who don't do well in school but are incredibly smart individuals. One of my smartest friends, who was a computer engineer, he worked at Intel during college, and he had, like, a 2.5 GPA. It's just, school isn't for everyone. It just depends what path you want to take.

Corey: And I think that there's also a bit of a misunderstanding these days, especially when we see economic downturns, where people have a degree, they have trouble finding a job because, spoiler, it's hard to find a job, full stop. I don't care who you are or what you do, it’s always a challenge. And the response that we've basically beaten into people is you should go back to school and get another degree. Yeah, great plan. I always hated that model.

Talk to someone who's doing the thing you think you want to do next. That's a good idea. Ask them, what should I do from where I am to get to where you are? And if the answer is, “You need to have a different degree,” okay, then it's something to consider. You're not going to be able to maneuver around your lack of degree if you want to be—oh, I don't know—a surgeon.

But if you want to go work in machine learning, for example, a PhD might not be a requirement. Maybe it'll help, maybe it won't. I don't know, I'm certainly no expert in that field. But I would talk to people who are and see what they have to say. There's a lot to be said for getting advice from people who've been down that path.

Kylie: Exactly. Yeah.

Corey: At least, that's how I see the world.

Kylie: That's the advice I've been given, too because at this age, at the end of my bachelor's, I'm thinking, “Should I get a master's? Will I be more employable? I'm too scared to leave college because what if I fail?” And if I ask people in the industry, “Should I get a master's?” And they're like, “No. I don't have a master’s, none of my co-workers have master's, you don't need a master's to be successful.” So, yeah.

Corey: Yeah. I've toyed with the idea of getting an Executive MBA or something like that, just because right now the highest official credential I have is an eighth-grade education, which, now it's a funny talking story, but it was challenging to get through my 20s. It was, great, how do you do this? It’s like, tell a very different story. And I change the subject, and I’m great at dissembling.

But what I really appreciate is that what you're doing, and your whole bio right now is touching on a whole bunch of things that resonate. My first job, when I was in the process of basically bombing my way out of school, was IT service desk. And a recurring theme on the show has been where does the next generation of engineer come from in the world of Cloud? I became whatever the hell you'd call me these days in the ops world by being an old grumpy Unix systems administrator because there's no other kind. Those jobs aren't really there anymore. I learned so much of how I view technology, troubleshooting from my helpdesk days. It teaches patterns that are incredibly valuable. But those jobs don't seem to exist in nearly the numbers that they once did.

Kylie: Yeah, I would agree. And in Sacramento, it's impossible to get an IT job here that's not with the state. And that's just something people in my major have accepted that you just need a state job to work in IT.

Corey: I… well, why not? I'll tell stories that irritate people, why not? My first breakthrough role was at a university—private—in California, where I basically bluffed my way through the technical interviewer, the screener who was supposed to verify my answers was out sick that day, they really liked my personality [snort], and they wound up offering me the job on the spot. And I was excited because finally, I was working in education because the benefit then—given my personality that should come as no surprise for anyone—I couldn't get fired. The problem was that I learned very quickly, I was working with a bunch of people who couldn't get fired.

And it became frustrating, and the politics of it drove me nuts. And then I idly talked to some companies in the commercial space, and it was much more compelling, or at least it sounded that way at the time. And I really haven't been back ever since. But it's a strange feeling, going from IT and the Service Desk, and then transitioning that into more either engineering-focused roles, or product-or project-based roles. It was a lot of fun, and I know a lot of people who did the exact same thing.

But I know other people who have been working in the IT service desk for 20 years. And it's always been strange to me seeing where people decide to remain, happily so, or where they decide I want something different. For me, the problem I always had with the Service Desk was, it's hard for me to troubleshoot computers misbehaving without a blistering stream of profanity. And you can't really do that on the phone with a paying customer. More than once anyway.

Kylie: Oh, you know, the paying customers are the ones who have profanity most of the time, for me.

Corey: It's two very hard skills: it’s, one, fix misbehaving technology, and two, the polite to someone who's currently being a rude bastard.

Kylie: Yeah.

Corey: It's very hard for me to walk and chew gum at the same time, with that.

Kylie: It's so true, yeah. I was just talking to someone on Twitter about this the other day about how at IT service desk, the most mean users are the ones that work in legal. I don't know what the correlation is there, but I've been cussed out way too many times by lawyers.

Corey: Oh, me, too. I married one, and the last argument I ever won was, will you marry me? It was all downhill from there. I found the same thing with professors. And the attitude that I got was that, “I’m—have a PhD. I am a world leading expert in this very narrow thing that is incredibly hard. They don't even offer PhDs in fixing my computer, so how hard could it really be? Therefore you're an idiot and I'm smart,” and 45 minutes into the conversation, we realize that they're suffering a power outage. So, it's a difficult problem to wind up overcoming because you can't put people in their place. At least I couldn't. I'm hopeful—maybe—that that has changed since the last time I worked in a service desk. Are you allowed to swear at the customers?

Kylie: You know, I’ve considered it, asked my boss. But it's a firm no, at this time.

Corey: Yeah, that's okay. There's always a next year plan. So, one of the reasons I wanted to have you on this show, which is ostensibly about the business of cloud computing, is a single line in your bio that I bet you may have almost thrown in as filler, but I strongly disagree if that's the case, where you are the president of your school's ski and snowboard club. And I think there is no better metaphor for cloud computing than hurtling down the side of the mountain into a bunch of trees. Tell me more about that.

Kylie: Yeah, that is one of my proudest accomplishments right now is being president of my ski and snowboard club. Just a wonderful group of people that I met in college, and I thought I can make this better. So, I became president last fall.

Corey: So, you can't even talk about it without using the word ‘fall.’

Kylie: Yeah. I'm still learning actually, even as the president, which I think is a really good way to get people to join because every time I tell someone that I'm the president of ski and snowboard club, they're like, “I can't stand on board without falling. It's impossible.” And I'm like, “Yeah, me too. Exactly.” It's really hard but fulfilling to say the least.

Corey: So, ski or snowboard?

Kylie: Snowboard. We do not like skiers in this household.

Corey: Gotcha. That's probably the things we're going to have a bit of agreement on. I can't stand up on a snowboard, but back in one of the boarding schools I was thrown out of, I was briefly competitive on the skiing side. It was fun. I was basically skiing because I was old enough to walk, turned 18, and then hadn't ever been back to a mountain until 2016.

It comes back again. A lot has changed, though. It turns out that everyone wears a helmet now when you're skiing, which is a new thing for me. That was never anything that was done. And at first, it's like, “Heh, look at all these little lilies who are terrified of hitting their head.”

And then I did some math and it's, “Wait a minute, I'm hurtling down the mountain about, eh, 30 miles an hour. Those trees and rocks are awfully hard,” and, “Oh my god, how have I never worn a helmet doing this?” It was very clearly one of those things where, huh, with a little bit of critical analysis, maybe realize it's not the things are safer and everything's sugar-coated these days. It's that we were doing things that were freaking dangerous, once upon a time.

Kylie: Yeah, I'm not going to lie. Probably a lot of people listening to this who are in the tech field are skiers. But yeah, as a snowboarder. We don't wear helmets. It's unfortunate, but I'm not going very fast. I'm still on bunnies. But another fun fact about me and my family is that my uncle on my mom's side, he was a pretty famous skateboarder and snowboarder in the 80s and 90s.

Corey: Wow. So, something else that you mentioned, almost as a throwaway, which I never had the patience to do much work with is gardening, which is odd. Everything else in your bio talks about technology, and leadership, and hurtling down the side of a mountain, and information security, which we'll touch on in a minute. And then gardening throws in there. I have a brown thumb and everything I touch dies, which means it's super problematic now that I have two children. But there's something to be said especially during this interminable year of quarantine. I mean, I have plants in my office now, and other people in my house have strict instructions not to let me kill them. But tell me more about gardening.

Kylie: Yeah. I know the tech field has a lot of inside plant, office plant gardens going on. I have not gotten any inside plants because I live with four college students. I don't want them to kill my plants. So, we have a pretty big backyard and a lot of pots that are cemented into the ground.

So, I thought, why don't I just start growing vegetables? And it has been incredibly relaxing. Tomatoes, kale, broccoli, spinach, just anything and everything, I try to grow in the backyard. Snapdragons; that was really fun. But yeah, that's just been my adventure. It's a really great hobby to have.

Corey: This episode is sponsored in part by our friends at Linode. You might be familiar with Linode; they’ve been around for almost 20 years. They offer Cloud in a way that makes sense rather than a way that is actively ridiculous by trying to throw everything at a wall and see what sticks. Their pricing winds up being a lot more transparent—not to mention lower—their performance kicks the crap out of most other things in this space, and—my personal favorite—whenever you call them for support, you’ll get a human who’s empowered to fix whatever it is that’s giving you trouble. Visit linode.com/screaminginthecloud to learn more, and get $100 in credit to kick the tires. That’s linode.com/screaminginthecloud.

Corey: I found that I'm happier now that I have growing living things in my office. I look around—I think I have three plants. I think a fourth is on order. And it just, it feels nicer. I would say it ties the room together, but I think that's a rug. There's just something nice about having them here, and I was never a plant person.

Kylie: Really?

Corey: For better or worse. Yeah. Again, when you're very good at killing things, it's hard to become attached to something you're bad at, at least in my case. You know, other than technology. But that's a separate argument entirely.

Kylie: Yeah. No, I agree. I have fake plants inside my room just because I love the way plants bring life into a room, just make you feel more positive. But… someday. I'll get real plants someday, but not anytime soon because I don't know if you saw, but one of my roommate’s dogs ate my dinner.

Corey: I did not see that. But first, was it something that was safe for dogs? Important things first.

Kylie: Oh, very safe. They had a delicious—

Corey: Oh, good. And good for that pupper.

Kylie: Oh, yeah. And it was two pounds of raw meat that I was going to use for my dinner that they ate clean off the tray. But that is the exact reason I don't have plants in here because they would all die. [laugh].

Corey: Yeah, we have special problems here where I used to be a dog person before I had kids. I was basically nesting before I was ready to have children, so I wound up fostering dogs for a while. And there was this belligerent little Chihuahua mix at one of the rescue events that just barked every week at people, angrily because that's how she decided to get adopted, and that's how I wound up with Ethel my little teacup Chihuahua. She's 15 years old now, but she's not really a dog; she's a malevolent weasel. And her entire shtick is stealing food.

It's really a challenge because she's this little 15-pound dog, 12-pound dog, something like that, who can just coincidentally—despite being 15 years old—jump onto the counter. It's like a cat and a dog had an unholy offspring thing that can fly. Just the worst little dog in the world. I love her, but she's terrible. And at some point, I've had to learn how to keep food away from her.

That was fine. We established an uneasy truce. Then the kids came along. And it's one of those oh, great. So, if I have to wind up worrying about everything, it's worst case, “Dog. Eat what you want. Worst case, you'll die a little sooner.” I’m not going for the world's oldest dog here. But at some point, you have to let go and prioritize. Pick the battles, as it were.

Kylie: Oh, yeah, definitely.

Corey: Speaking of picking battles, you have a strong interest in infosec, which is pretty clear to anyone who spends more than 10 minutes looking at your Twitter profile. Tell me what about that appeals to you?

Kylie: I think it appeals to the empathetic characteristics in me. When I saw how users and technology interacted, I just saw all of the things that could go wrong. And then I just started writing about them in the State Hornet. My first article was about the Iowa caucuses, and that technology being so faulty. Recently, I wrote about voting technology, and the security there is so poor.

The way people and technology intersect, I just have way more empathy for people than computers and corporations. So, I want people to be protected. And I noticed college students also do not understand the technology they're using half the time, don't really care to. So, I thought it was important to write about it, tell them like, “Hey, this is what TicTok is; this is what Respondus Monitor is; this is what you're agreeing to.” So, security, in that sense has just been really fascinating to learn about for me.

Corey: I guess the strangest thing I've heard on this show, in the however many episodes we've done so far, is one of the things you just said, which is, “Oh, I’m really into information security because of empathy.” I'm sorry, I know I'm going to irritate people who are listening to this, but I started off dabbling rather heavily in infosec for a while, and I got out of it because the community was a freaking trash fire. It was all about who the smartest person in the room is, and glorified punching down. When someone made a mistake, it was all about, “Aha, I got the bastard,” rather than, “Help me understand where you're seeing this so we can all learn together.” It was absolutely a trash fire of a community.

I stopped going to infosec conferences. I really stopped paying attention to security in an industry sense. Not stop paying attention to security engineering sense, which it seems every company with a budget loves to do. But it just became such a draining, awful experience that I found there were other places that were more aligned with what I wanted to be doing. Is that no longer the case? Are there pockets within infosec that are better than that now that I'm not seeing, or is there something else.

Kylie: That's funny because I don't even know what communities I'm in, generally. I just follow who I find interesting, who talks about important things. I have seen that punching down, that toxic community, and I've seen that community attacking people I follow who are prominent security—often women or non-binary people who have a large platform in infosec, and just constantly attacked. So, I have chosen not to delve deeper in the infosec community for that reason. It just doesn't feel like a welcoming environment, and I feel like I would survive without opening myself up to that type of toxicity.

Corey: And to be fair, I went to a few of the conferences, DEF CON, which is probably not a great place to get started. The industry events were always much more focused around checkbox compliance in my experience, and it feels like CSOs’—chief information security officers’—primary role is sort of to be an ablative armor for a company so that they can get fired when there's a breach. It's a cynical perspective, I know. And there's a lot of nuance to it, and I'm not deeply steeped in that space anymore, and that's okay. But as of the time is recording, I recently did a Twitter thread about infosec myths: “Rotate your access credentials every 30 to 60 days.”

Yeah, if your cloud credentials get compromised, you'll find out in 20 seconds and it's going to be expensive. It just feels like busy work for a lot of those things. And there's a lot of myths and, effectively, security theater that go around it. And I figured I was going to get torn to pieces, but you're right, the responses were mostly positive. And even the stuff I intentionally put in there to annoy people didn't really seem to trip anything. So, it may be that I'm bringing old prejudices into a new renaissance of infosec culture. I hope so.

Kylie: That's so funny. Yeah, I would hope so too. I have a lot of hope for the community. I actually have, I think, a bit of naive hopefulness for the industry itself. I constantly think that people in this field are trying to build technology for a better world, non-biased technology. That's my hope, and actually, what I see is my truth. So, sometimes Twitter can damper that for me.

Corey: Talk to me a little bit about Twitter. Back when I was starting out in my career, Twitter didn't exist, which was fortunate. I did not have anything approaching the social wherewithal that I do today. I made a lot of shitty jokes that were punching down, I was not as kind or empathetic as I want to be as a person. And that's a problem.

And, for better or worse in that era, there's some IRC chats, there's a few Usenet posts, but there's not a sort of system of record where one day people decide to go data mining and, “Haha, remember these embarrassing things you said? Oh, here's some career-ruining things you've done in the past.” It's nice to have that freedom to do it. But when I started this whole thing, about four years ago now, I had I think, less than 2000 Twitter followers. You're over 5000 at this point and just getting started. You have a bigger audience. More eyes are upon you. What's it like?

Kylie: Some days good; some days bad, definitely. I use the block button liberally all the time.

Corey: I've started doing it a lot just because life is too short to be arguing with jerks. Also, punching down is never a good thing, so when I see something incendiary responded to me by someone who has six followers, and their name looks like a bot account, I'm not going to bother. Why give an asshole a platform? Block. Move on.

Kylie: Yeah, whenever I see people arguing with prominent people in tech, actually, I'll go ahead and block them, everyone who follows them, if I have the time, everyone who'd like their tweet. I just don't have time. And that has saved me a lot of time because not as many people argue with me and quality filters have really helped, honestly.

Corey: The thing that really throws me, at least, really it was eye-opening for me, is that I used to believe that it's okay to insult people, yell at them, et cetera, as long as you're not punching down. The problem is, is in some context, no matter who you're talking to, you can punch down and not realize it. I've inadvertently punched down at multinational companies before, and that doesn't feel great. When people try and emulate some of the whole snarky sarcastic thing that does well on Twitter, very often, they just end up being mean.

Kylie: That's true, yeah. Sometimes my fault with Twitter is I don't take it so seriously because I'm 22. I see it as a social media platform, as it is. But something I've been reminded of before is that it's more than a social media platform, at least for me, right now in my life; I have to take a little bit more seriously. Someone recently told me I have to treat it like LinkedIn, which I thought was funny.

And I still don't think I can successfully take it seriously, and I don't really plan to because that's not my personality, but I will never, hopefully, punch down. I try not to argue with people at all, actually because it's just a waste of my energy. I used to spend almost all of my time in high school, arguing with people I went to high school with about feminism, meninism at the time. And Black Lives Matter was just starting when I was in high school, so I spent a lot of time defending these things I thought I cared so much about. But at the end of the day, I just realized there's no reason to be arguing the way that I did on Twitter. I just need to pick my battles, move on, use the block button.

Corey: Let me Nostradamus this for a second. At some point, you're going to punch down, you're going to say something that hurt someone's feelings, or is taken away you didn't intend. And I used to live in fear of that. But what I've learned is that when you get it wrong, that's an opportunity to demonstrate who you are. And that in many ways is a better glimpse into humanity than if you hadn't ever screwed up at all if that makes any sense.

Like, I wound up doing a Hitler reacts video—a while back—to his AWS bill. And it was fine. It was funny. I got a kick out of it. Some people were, “Oh, you shouldn't make light of Hitler.” Yeah, I'm the Jewish grandson of a survivor. Shut up. We all have our own relationships with it. That's mine. Don't like it? Move on.

But what I screwed up on was there's one line where someone who is a woman speaks to another woman and the dialogue line I used was, “Oh, that's okay. I get gigabytes and terabytes confused all the time, too.” And it worked out well. It fit the flow. It was great. It never occurred to me that that is how it might come across. But it did.

And as soon as I found out about this, I had two paths. I could either double down and, “No, it's funny. You can deal with it,” or, “Holy shit, people feel bad.” So, I took the video private and did a whole thread on, I screwed up. I'm sorry. I'll do better.

And I haven't made that particular mistake since. And no one's right all the time. For trying to pretend otherwise is just lunacy. It gives you a chance to demonstrate character, I guess that's how I see it. Maybe that's stupid and hopelessly naive and probably overprivileged. But that's how I see it.

Kylie: I like to live my life hopelessly naive. It's brought me far, so far. But yeah, I actually had an incident like this recently, and my Twitter was pure mayhem. When I made a joke about teaching computer nerds how to communicate properly. I think that's actually how you ended up following me.

Corey: Entirely possible. My decision matrix for do I follow someone or not is generally whimsical. It's one of those, “Oh, that's a fun tweet. I'll follow.” Other times it's, “Huh, this person's been making great tweets, and I always like and retweet them, and they showed up on my feed for three months now and I've never followed them.” And there's no rhyme or reason.

If you're listening to this and wish I would follow you, well, we all have dreams. It’s probably not anything personal. But maybe it is. So, social media is a weird, strange, different thing. To wrap this up a little bit. You are doing an awful lot of things, but you're also coming up on a transformation. In May. You're graduating. What's next for you?

Kylie: Oh, so many things hopefully. Right now I've just been applying to jobs, interviewing. We'll see what I catch at the end of the day. I'm hoping to land a job in information security, maybe network engineering would be wonderful. I have so many ideas. I'm not sure which way it'll go. Maybe even tech journalism, if I can land a job at a publication. So, many options; still deciding.

Corey: It definitely seems like there's a wide world ahead as far as different opportunities for what you can focus on. I'm curious to hear what goes for you. So, if people want to follow along on your journey, see what you're up to, and hear your scintillating wit, where can they find you?

Kylie: My Twitter is at @kylie_robison. My website, kylierobison.com. It's nothing fancy. I just needed to put it together so I could answer questions like this. But yeah, Twitter is the main way to reach me.

Corey: Fantastic. And we'll of course throw a link in the [00:30:50 show notes]. Thank you so much for taking the time to speak with me today. I really appreciate it.

Kylie: Yeah. Thank you for having me. Corey.

Corey: Kylie Robison, current student who does all the things. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you've hated this podcast, please leave a five-star review on your podcast platform of choice and tell me why I'm wrong to prefer skiing over snowboarding.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Alex HidalgoAlex Hidalgo is a Site Reliability Engineer and author of the upcoming Implementing Service Level Objectives (O'Reilly Media, September 2020). During his career he has developed a deep love for sustainable operations, proper observability, and using SLO data to drive discussions and make decisions. Alex's previous jobs have included IT support, network security, restaurant work, t-shirt design, and hosting game shows at bars. When not sharing his passion for technology with others, you can find him scuba diving or watching college basketball. He lives in Brooklyn with his partner Jen and a rescue dog named Taco. Alex has a BA in philosophy from Virginia Commonwealth University.

Links Referenced:

  • Buy Implementing Service Level Objectives on bookshop.org
  • Buy Implementing Service Level Objectives on Amazon
  • Follow Alex on Twitter
  • Alex’s personal site
  • Corey’s landing page for Implementing Service Level Objectives

TranscriptAnnouncer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Linode. You might be familiar with Linode; they’ve been around for almost 20 years. They offer Cloud in a way that makes sense rather than a way that is actively ridiculous by trying to throw everything at a wall and see what sticks. Their pricing winds up being a lot more transparent—not to mention lower—their performance kicks the crap out of most other things in this space, and—my personal favorite—whenever you call them for support, you’ll get a human who’s empowered to fix whatever it is that’s giving you trouble. Visit linode.com/screaminginthecloud to learn more, and get $100 in credit to kick the tires. That’s linode.com/screaminginthecloud.

Corey: This episode is sponsored by a personal favorite: Retool. Retool allows you to build fully functional tools for your business in hours, not days or weeks. No front end frameworks to figure out or access controls to manage, just ship the tools that will move your business forward fast. Okay, let’s talk about what this really is. It’s Visual Basic for interfaces. Say I needed a tool to, I don’t know, assemble a whole bunch of links into a weekly sarcastic newsletter that I send to everyone. I can drag various components onto a canvas: buttons, checkboxes, tables, etc. Then I can wire all of those things up to queries with all kinds of different parameters: post, get, put, delete, et cetera. It all connects to virtually every database natively, or you can do what I did, and build a whole crap ton of Lambda functions, shove them behind some APIs gateway and use that instead. It speaks MySQL, Postgres, Dynamo—not Route 53 in a notable oversight, but nothing’s perfect. Any given component then lets me tell it which query to run when I invoke it. Then it lets me wire up all of those disparate APIs into sensible interfaces. And I don’t know front end. That’s the most important part here: Retool is transformational for those of us who aren’t front end types. It unlocks a capability I didn’t have until I found this product. I honestly haven’t been this enthusiastic about a tool for a long time. Sure they’re sponsoring this, but I’m also a customer, and a super happy one at that. Learn more and try it for free at retool.com/lastweekinaws. That’s retool.com/lastweekinaws, and tell them Corey sent you because they are about to be hearing way more from me.

Corey: This episode has been sponsored in part by our friends at Veeam. Are you tired of juggling the cost of AWS backups and recovery with your SLAs? Quit the circus act and check out Veeam. Their AWS backup and recovery solution is made to save you money—not that that’s the primary goal, mind you—while also protecting your data properly. They’re letting you protect 10 instances for free with no time limits, so test it out now. You can even find them on the AWS Marketplace at snark.cloud/backitup. Wait? Did I just endorse something on the AWS Marketplace? Wonder of wonders, I did. Look, you don’t care about backups, you care about restores, and despite the fact that multi-cloud is a dumb strategy, it’s also a realistic reality, so make sure that you’re backing up data from everywhere with a single unified point of view. Check them out as snark.cloud/backitup.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Alex Hidalgo, who's a site reliability engineer and, due to an escalatingly poor series of life decisions, a recently published author, specifically, the book Implementing Service Level Objectives. Alex, welcome to the show.

Alex: Thanks, Corey.

Corey: So, every person I've talked to who's written a book has given me the thousand-yard stare when I ask if I should write one, and then immediately begins screaming, “No, never do this.” I've come to the conclusion that nobody actually wants to write a book; they want to have written a book. How accurate is that?

Alex: I think there's absolutely some truth to that. It is difficult, it is tiring, it is emotionally draining, and having to have written one in the middle of a global pandemic didn't make anything any easier. That being said, it's also been an incredibly rewarding experience, especially when you end up producing something that you're truly proud of. And my book is about people. It sounds like it's about service level objectives, and I guess it kind of is, but it's mostly about how to use those to make people's lives better. And I've seen how this process can do that. And so having something out there that I hope will help people's lives is ultimately rewarding. And there were more tears of frustration than there were tears of joy, but there were both.

Corey: So, let's start at the very beginning here, I know the book is—at the time of this recording, it is in print; they are starting to ship out. I have not yet received my copy, but of course, I have ordered one, I keep an eye on whatever O'Reilly releases, and especially when it's people I know, it's a no brainer. It will, in all honesty, sit on the shelf and never get opened because I will do my actual reading on the Kindle. But having something on the shelf is, for something like this—[00:03:54 unintelligible] you know, you’re the person that wrote it—is just the right thing to do. But start at the beginning for me here because it turns out that I am a white guy in tech, which means my failure mode is a board seat and a book deal somewhere; everyone assumes that I know everything about everything, and I tend to not shatter that illusion very often, but I have no earthly idea what the hell a service level objective is. Since it's just you, me, and the thousands of people listening to this, what is an SLO?

Alex: So, an SLO, well, it's an objective for your service; you can't be perfect. The story I like to tell is, imagine you're using a streaming media service of some sort: a Netflix, a Hulu, a Disney Plus, whatever, and when you're using this, normally we start a new video, it buffers for a few seconds, and you're fine with that. And that's technically not perfect, but it turns out these services don't have to be perfect for you because you're fine if it buffers a bit. But on the same token, if it buffers for, like, 20 seconds, you don't love that, but you're not going to abandon the service unless it buffers for 20 seconds every single time. Then you may say, “Screw this. I'm moving to a competitor.”

So, the idea is, find out what your users can tolerate, and make sure you're only failing that often. If you're not losing users, if people are still happy, in general, with your service, if you only take 20 seconds to buffer 1 in 50 times, then aim for that. Because you're going to spend too many resources, both financially and via your developers and your support engineers if you try to make everything 100 percent all the time.

Corey: It sounds, on some level, like it's a derivative of SLA, service level agreement. What's the difference?

Alex: So, the difference is, a service level agreement—and they've definitely been around much longer. Actually, in some of my research, I’ve found—

Corey: I've seen SLAs in contracts all the time when negotiating those, I have never seen the phrase ‘service level objective’ in a contract, which means that lawyers will not know what I'm talking about if I use SLO, I suspect.

Alex: Yep, exactly. An SLA is something you put into a contract, and it generally implies that you owe somebody something, whether it's credit or actual money, if you violate that. SLOs are an approach to thinking about the reliability of your service. They are promises in some sense, but definitely not contractual ones. They’re tools, they're a bit of data that helps you make decisions. Are we buffering too long, too often? Is this page not loading correctly, too often? You know these things are going to happen, and just make sure it's not too much of the time. It kind of accepts the same thing as an SLA does. SLAs are generally not 100 percent because, again, people realize something will break at some point in time. SLOs think about things in the same way, in that sense, but they're used to help you make decisions. Do we need to focus on this part of our product? Do we need to focus on this?

Corey: So, help me understand this in the context of a story that I've related from time to time on various forms of podcasts and whatnot. Years ago, I was trying to buy a pair of socks on amazon.com. And I clicked the buy button and I got one of their error pages, which of course features dogs. In all honesty, the dog page is more satisfying than any other page on amazon.com.

If I listened to the common wisdom, that would mean that during that outage that lasted about an hour or so, I would have therefore gone to a competitor to buy the pair of socks, or alternately, one day out of the week on the day that that pair of socks should have been there, I would just go without socks whatsoever. In practice, “Oh, that's weird, I ever see that. Haha.” I come back an hour later, I buy the socks, and life goes on. There was no loss of revenue, in my case, for amazon.com during that outage. However, if every third time I tried to buy something at Amazon, I got the dog page instead, I'd probably spend a lot more money at Target. So, is that a naive storytelling, I guess, understanding of a much more complex concept of SLOs? Are they related, or is this completely out in the weeds and it's a boring story we should make sure we drop on the floor in post-production?

Alex: No, that's exactly it. If you're down for an hour, chances are people are going to be like, “Huh, this happens.” Stuff breaks; people are used to it; they're actually mostly okay with it. And you're probably just going to come back and check in an hour. That's exactly correct. That's the whole point. You can be down for an hour, you just can't be down for an hour too often.

So, find out what that is. Find that percentage. Can you be down for an hour, once a month? Twice a month? It's going to be different depending on your service. If you're a specialty retailer—no one else sells your stuff—then you can probably actually be a little bit more lax. And if you're someone like Amazon, who's also expected to be constantly up because they're the largest company in the world, in some sense.

So, no, I think you got it right, exactly. It's just that implementing SLOS, there's a lot of math that goes into it. There's a lot of discussions you have to have, it's not easy to just pick a number. So, they're simple to talk about, not always easy to implement. That's why there's a whole book about it. But your story, that's exactly it. That's exactly the whole point.

Corey: B&H, the photo company, closes for 24 hours for Shabbat every week. And they wind up having their website up, but they say yeah, you can't actually make a purchase until Shabbat ends. On some level, it's kind of their own brand now at this point, and it seems to have worked out reasonably well. But as you say, it comes down to what your story is, as far as approaching the market. I'll go back an hour later to buy socks because I need socks. I'm not going to go back to your website an hour later to click on an ad that wasn't displaying.

Alex: Exactly. So, a meaningful SLI is a measurement of how your service is operating from your users’ perspective. And when people ask me to explain it a bit more, I'm always like, “Well, it's the same thing as a KPI, or key performance indicator, for the business side.” Or if you were to talk about this to a product manager, they would say, “Oh, it's a user journey.”

You’ve got to take everything into account. People are going to have different expectations for how different parts of the service work. As you said, there's a difference in between, perhaps, something like a button not registering a click on the first try but registering a click on the second try, that’s—I think people aren't going to be too upset about that can probably fail more often than just not be able to check out entirely.

Corey: When you're looking at SLOs through a lens of things we should strive to do, how does that keep from becoming a, “We'll try our best.” which sounds great, makes everyone feel good, but isn't something that is easy to represent as either having value or matters at all, to the larger business.

Alex: So, I think you really do just try your best, but I understand, I agree that that doesn't sound like a great sentence, even though it generally is true. You know, you'll try your best, and you’ll try to make sure your service is good. But they can't always be perfect, and I get that that kind of language isn't great, but that's actually exactly why I think the whole SLI, SLO, error budgets—which we haven't even talked about. It’s measurement of your SLO over time, as opposed to, kind of, right now—this is actually an example where the numbers can help, can point to things and say, “Yeah, buts, we were only unreliable for 4 minutes and 32 seconds last month,” or something along those lines.

And that's how you, kind of, help explain to people that yes, in a sense, this is we're just going to try our best because we cannot try our perfect; that's not a thing. People get to understand what you actually mean with that. When you're saying I'm going to try my best, you're actually saying, “Well, we're aiming to be 99.95 percent reliable, and that translates to X number of minutes per month that we may not be unreliable.” And that can often help people understand, “Oh huh. Maybe trying your best is actually good enough.”

Corey: I really wish that more people were explicit about saying trying your best is good enough because I can't shake the feeling that that is not a well-circulated belief in far too many places.

Alex: Totally agreed. But it's the truth. It's how things actually work. Things fail, people fail, and it turns out people actually know that. When you’re running a business, the end goal, of course, is to make money and make as much money as possible—or at least for most businesses. There are absolutely outliers there—and that means your executives or whoever owns the most shares, or your shareholders if you've gone public, that's their goal, right?

At some level, they want to make money and want to make as much money as possible. And therefore, they think the way to do that is to aim for perfection, to aim for 100 percent. But you're always going to falter if you do that, and those are the people who can be most difficult to convince that it's just not prudent to aim for 100 percent, but those people can get there as well. It can take time, it can take a lot of examples, but don't let great be the enemy of the good. It's a concept that, as humans, we know so well that we have an idiom for it.

And you just got to figure out how to translate that to a business where, again, they think they want to be 100 percent because they think they need to do that to make as much money as possible, but you can actually often save in resources by not trying to be perfect because the amount of money you're going to spend, it ends up almost becoming a limit approaching infinity, that curve is going to shoot straight up in how much money it's going to cost for you to try to be as reliable as possible. You're going to have to run multi-region, and you're going to have to pay for quicker replication, and you're going to have to hire more engineers because, like, if you want to be up almost all the time, you're going to have to have people who are on call who can respond within, like, 30 seconds.

The only way you can do that is if you have, like, a follow the sun rotation where you have offices all over the world because otherwise—not everyone's going to wake up at 3 a.m. so you got to make sure it's someone's 3 p.m. instead. And you can see how quickly this escalates. It can actually cost you more money; you can actually make more money by having a more reasonable target. Think about what your users actually need from you, and often—well, again, humans expect failure. They're cool if stuff doesn't work every once in a while, as long as it doesn't work too often.

Corey: One of the only other places I've seen SLOs discussed in any serious capacity was the SRE book that came out of Google. And there was a lot of good stuff in that book, but I had a bit of a negative reaction to it just because that came out right when I was in the middle of getting an awful lot of, “We’re Google, we’re smarter than you,” flak on other fronts. The people who wrote that book, to be very clear, are great. That is not the impression I have of those people.

But it's, “Oh good. How to be more like Google. Just what I don't want to listen to.” So, I largely ignored it for a while. But Liz Fong-Jones, now at Honeycomb, is a big advocate of SLOs. They're one of my consulting reference clients, and we've had—most of our conversations around cost optimization have centered around SLOs. So, what the expectation for their customers is and the commitments that they have made. And it was a really interesting philosophy that I haven’t seen replicated elsewhere, yet.

Alex: Yeah, it's something that's gaining a lot of traction. And actually, so I was at Google. I was on the Customer Reliability Engineering Team with Liz, and one of the things we did was we went out and we taught some of Google's largest cloud customers how to SRE. That was kind of our goal. And the beginning of every journey was, “You need to have SLOs first. This is our common language. This is how we talk about reliability. Reliability to an SRE is defined by service level objectives.”

And so while it's still, outside of Google, still kind of a growing discipline, lots of people are doing it. In some ways, this book is about the two years I spent at Squarespace. It's essentially my story there. I joined and people said, “We want to do SLOs, and we know that you know SLOs, and let's do it.” And I was like, “Okay, sure.”

Except I didn't realize what it takes to build this kind of thing from the ground up. Because suddenly, I wasn't at Google anymore. I didn't have all the tooling that Google has, I didn't have the cultural buy-in, not just by engineers but across various different organizations, and I had to do everything from the bottom up: new tooling had to be created, new software written, I had to drastically change how people measure things, how people think about things, and that's kind of how the book came to be. I was running a lot of SLO workshops just internally where teams could come and I could spend three, four hours with them, and maybe even hopefully, end up with a single defined SLI/SLO pair before they left the room. And it was just getting to be a lot because I was seeing the same thing over and over again.

And I was talking to my colleague, Gabe, and I explained this to him, you know? It's just getting tiring. And then I said, “I wish there was a whole book about this so I just point people to the book.” And Gabe said, “You should write it.” And I said, “No, no, no, no. We need, like, the expert to write it.” And he said, “You are the expert.” And so I'm pretty sure my response was, “[BLEEP],” because I knew I was now going to write a book.

Corey: It sounds on some level, like writing a book is something that is—it used to be this thing that people would do, this aspirational task. “I’m going to write a book.” Increasingly, it's starting to sound, for an awful lot of author folks, that it's more like a dead dog that has been cast into your yard by one of your neighbors, and now it's time for you to worry about it.

Alex: [laughs]. I mean, I don't know if I go quite that far with the metaphor, but yeah, absolutely. I mean, I won't lie. It's awesome that I wrote a book. That's neat. As a kid, my dream was to be an author. I thought it'd be like fantasy books and not a technical manual, but still, it's amazing. I got the first physical copy yesterday, And I bawled. I cried.

It's really, really neat having done that, and I won't pretend that some of the status that you get with that, I won't pretend I’m totally ignoring that. But absolutely, at the root of it, it's more like, yeah, someone needs to do this. I am the right person for it. I write well; I know a lot of people who can help me with this; I helped with the second SRE book, so I already understood the process just a little bit, and it was like this needs to be out there, and so, it may well be me.

This episode is sponsored by our friends at New Relic. If you’re like most environments, you probably have an incredibly complicated architecture, which means that monitoring it is going to take a dozen different tools. And then we get into the advanced stuff. We all have been there and know that pain, or will learn it shortly, and New Relic wants to change that. They’ve designed everything you need in one platform with pricing that’s simple and straightforward, and that means no more counting hosts. You also can get one user and a hundred gigabytes a month, totally free. To learn more, visit newrelic.com. Observability made simple.

Corey: I'm glad it's you because, first, I get to wind up reading a book that I don't have to write, which is great because I'm never going to write one. I don't have the attention span to write a tweet most days. But being able to read an in-depth thought in book form, where you are able to opine on the various angles of this from a whole bunch of different perspectives in a much longer form than a blog post, or this is a tweet thread. Tweet one of 487,000, and so on and so forth. It's great to have a single place for it to go. Also, you tipped me off fairly early on that—talk about bucket list items—that I am referenced in an easter egg within the book.

Alex: Yeah. So, actually got to give props to some co-authors here. I actually ended up writing only about 60 percent of the book. I was always planning on bringing in two or three people who are, like, experts at very niche parts of this, and suddenly I had people volunteering all over the place.

And so, the other 40 percent have a bunch of amazing people from across the industry. And Polina Giralt and Blake Bisset wrote a chapter about data reliability. Data reliability is—you've got to approach it in a very different way than latency, or availability, or error rates. So, we have a whole chapter about data reliability, and there's a part where they discuss what even is a database? And they ask the question, is Route 53 a database? And we were able to get a footnote in that says, “Hi, Corey.” At first, the editor wanted to take it out, and I was pretty adamant that we should leave it in.

Corey: Oh, absolutely. The fact that I am referenced in this now means that I'm getting it framed and hung on the wall. “Oh, did you write that book?” “Absolutely not. One of my stupid jokes made it into that book.”

And that's—yeah, oh, I'm absolutely going to steal credit for you in this sense. I can finally have that as my counterpoint to my business partner’s story. He wrote Practical Monitoring for O'Reilly and has it on display on his bookshelf behind him when he's on Zoom calls. And it's a fun problem because, as it turns out when you've written a book, it's very hard to bring up the fact that you have written a book—because you're proud of it, you spent a disturbing portion of your life for a while on writing that book, but you can't open the sentence with that, or people find it pretentious and ridiculous. So, my position has always been that if I know someone's written a book, I will drop that into virtually every conversation when someone who’s talking to them doesn't know that fact. I try and be the one-person promotional band for stuff like that. So, do you know that you've written a book?

Alex: It still doesn't feel real. Even though I got that physical copy and have paged through it one by one, I spent almost, like, an hour, not really reading but kind of reading, and it still doesn't feel real, to be honest. But I totally hear you. So, I've given myself the opportunity to brag occasionally, to gush, because I'm incredibly proud of this and this was literally the hardest thing I've ever done in my entire life. So, I tried to give myself the opportunity to occasionally talk about it.

But outside of Twitter, where, honestly, I don't care that much; I’m constantly self-promotional there. In the real world, you're absolutely right, it's difficult to bring up. I posted a picture of the book on an internal Slack channel at Squarespace because I let myself say, okay, cool. I haven’t talked about the book at work for several months, here's my once a quarter permitted single message about it, and there were engineers at the company that still didn't know I was writing the book because you're right, it's difficult to bring up. You don't want to sound like a bragger. But I do think you have to try to give yourself permission every once in a while. There's nothing wrong with promoting the things you've done, especially the ones that you're very proud of.

Corey: Looking back, based upon what now, first, would you write the book again, and secondly, what do you wish you'd known before you started?

Alex: Yes, I’d write the book. Again, at the end of the day, the positives outweigh how difficult it was. What would I do differently? I would take more time off. I'm not writing a book again and also working full time.

I did take a few weeks off in December but having to do essentially all of this work on weekends and evenings, that was draining; it was a lot. And I should have known that I needed to give myself more time there. Don't do it when a global pandemic is about to happen. That part was [laugh] terrible. Especially if you're someone who—I can't write at home, I need to be at a coffee shop, at a bar. It's strangely not very unheard of. Lots of writers are this way.

And when I didn't have anywhere to go anymore, that was tough. That was not easy. Of course, you can't really predict that, but I had to throw that out there. [laugh]. And the final thing I’d do is I had a bunch of co-authors and they're all wonderful and every chapter turned out great, but trying to navigate that many people's schedules, and their own commitments, and I think I just want to be more sure, ahead of time, to let these chapter authors know, this is going to be difficult; this is going to take you more time than you think.

Can you take a week or two off? Because otherwise, you're really going to be struggling with this. So, those are the things: just give yourself more time; if you're working with other people, ensure that they're giving themselves enough time. And just try to make sure you're in a good situation in general. That might be a better way to sum up the whole pandemic thing because, of course, you can't ever control that. But make sure you have time, make sure you're comfortable. Make sure you're not going through other life events. Don't do this in the middle of some other crisis, or health problems, or something like that. You need to be in the best possible mental state that you can be in before you embark on something like this.

Corey: In what you might be forgiven for mistaking for a blast from the past, today I want to talk about New Relic. They seem to be a relatively legacy monitoring company, and I would have agreed with that assessment up until relatively recently. But they did something a little out there: they reworked everything. They went open source, they made it so you can monitor your whole stack in one place and, most notably from my perspective, they simplified their pricing into something that is much more affordable for almost everyone. There's even a free tier with one user and 100 gigs per month, totally free. Check it out at newrelic.com.

Corey: That's a hard and heavy lift in 2020. And again, there is a production delay between the time that we record this and the time it goes out. Sorry listeners, when you download something from your podcast, I don’t, quick, get someone on the phone, and have that talk live. I know. Spoiling the production magic for you.

But it seems like this is getting to be such a weird year from week to week. It's, “Wow, they didn't even mention the giant meteor.” And well, here we are. It's a hard problem to solve for as far as finding time to write. I've written a few basic outlines of books I've toyed with writing, and invariably in almost every case, I find another book that's already been written that aligns closely enough that, “Oh, I'll just talk about that thing instead. It's easier.”

Alex: Yeah, that’s… [laugh] I remember, when I was first really getting into service level objectives at Google, I was on a team that was responsible for the monitoring and alerting for everything else across Google. So, we wanted to have well-defined SLOs so other internal engineers could look to our definitions and understand how reliable we were aiming to be so they could ensure that their systems were, you know, handling things correctly and knew how to retry when they had to, and things like that. And suddenly, I had this great idea of building it an SLO repository, a centralized place where everyone can define their SLOs and they get some tooling or dashboarding for free, and it was a centralized place for you discover what your dependencies were aiming for so you could set your targets correctly. And I told one of the staff engineers on my team, and he was ecstatic about it.

And he was like, “Oh, my God. Alex, this is a great idea. This is going to get you your next promotion.” And I spent a few hours starting to outline what it would look like, and then someone else on my team came to me, he’s like, “I just discovered there is an entire team staffed of ten people working on this product.” [laugh]. So, my great idea, immediately up in smoke because someone else already had that great idea first.

But in this case, I looked, I really did. When you fill out a proposal for O'Reilly, they ask you, “What are competing books to yours? We need to know, what are you comparing yourself against?” And I listed the SRE books including Seeking SRE, David Blank-Edelman’s book because they at least talk about the SLOs, but really, it was tangential.

I was like, what I'm reading is strangely new. It's not just much more expansive, but it's actually a pretty different take than how they're described in either of the Google SRE books. So, that was one of the reasons I really felt like I had to do it. Because I looked and I couldn't find what I thought needed to be out there.

Corey: That's probably the greatest sign it's time to write a book, I would imagine. When no one else is talking about the thing that you want to talk about, or they are, they're getting it all wrong across the board. My position has always been to do a snarky take on Twitter or a sarcastic blog post, but there are times you need to go deeper than that. And to be honest, I'm very glad that people like you have attention spans.

It's easy to fall into a trap, in my experience of, in the world of Twitter and things like it, it's easy to attain relative mastery—or absolutely not, but the appearance of relative mastery in the confines of 280 characters. But then you see people who are legitimate experts in things and oh, it turns out that maybe I shouldn't be reinventing all the stuff from first principles, as if I were, suddenly Hacker News come to life. There's a definite value in seeing deep exhaustive research. One of the things I find most worrying about my increased attention to short-form social media is that I'm not reading the long-form stuff that really lets you dive into a topic with anything approaching the frequency that I used to.

Alex: Yeah, and I think it's actually—it's an interesting time in tech because I think a lot of people are pivoting towards understanding that we need to have a more in-depth understanding of how everything works. People have stopped being experts or deep experts at individual things. People are always being asked to be full-stack engineers, and you've got to understand everything. If you try to do that, then you will only ever have a shallow understanding of anything.

And I think it's a really interesting time because people are starting to realize that, and one of the ways I see it manifesting right now is in the fact that people are starting to look to outside industries. We've tried to call ourselves engineers for a long time. There's a lot of debate about whether or not that's applicable or not. It's just semantics, I don't think it's actually important outside of the fact that it is, in fact, a fun debate to have. But what I am seeing is people realizing, “Oh, there are other engineering disciplines, and they've learned all these lessons already.”

Especially in my world, people talk about reliability, and unfortunately, from an industry standpoint, that's mostly come to mean availability. Those things are actually very different. Reliability means so much more than that, and reliability engineering has been around since, I think, was the late 1940s is when the term was first phrased by the US military, in terms of whether some—I can't remember the exact object it was now, but, you know, like some armament, would it function the way it was intended to? The term reliability was coined to mean, “Is this doing what it was designed to do?” And that's a heck of a lot more than just being available.

And we're seeing more people think about that, though. You're seeing more people getting academic about things. And I remember a few years ago, at work, someone was trying to solve a problem with the fact that there were only very low-resolution metrics coming in, only a few API calls per hour. They’re like, “How do we alert off of this?” Because a single error—which might be totally fine—could represent 30 percent of all traffic over this hour window, but do we want to measure over 24 hours because then we wouldn't alert until 24 hours have passed.

And a colleague came over and said, “You know, you can just use a binomial distribution to solve that, right?” And everyone was like, “Huh? What are you talking about?” And he just broke out Wikipedia and showed us and suddenly, the entire team pivoted to understanding like, holy crap, there are statistical models, some of which were developed centuries ago, that we can use to help with so much stuff. And again, I think it's neat because after years and years and years of tech being pretty egocentric, and thinking we must solve it, or thinking we are the smartest in the room, while it’s definitely not gone entirely, I feel like—and I've been doing this for a long time. I've been in the industry in some way or another for almost 20 years—and personally, I feel like I'm seeing more and more people looking outside saying, “How have other people already solved this problem before?”

Corey: One of the problems that I've always seen is that there's this tendency to not look for prior art; instead, sit down and dive right into attempting to solve it internally with the resources you have first. Well, one of those resources is Google. Take a look and see, maybe other people have solved this. One problem that I have around this space is that the term ‘[00:31:17 serverless] level objectives’ is not discoverable if you don't know that the term exists. How do you get this in front of people who are absolutely positioned to benefit from this, but don't know what they don't know, so they don't have the term to look for?

Alex: I don't know if I have all the answers there beyond the fact that, in my opinion, to be a good engineer, you need to understand how to market. If you have a new idea or a new service—

Corey: “Well slow down, there, hasty pudding,” says AWS. But please continue.

Alex: [laugh]. Exactly how to approach this, how to get this in front of everyone, I don't think I know, and I don't know if I'm the exact right person for that. But one thing I have learned is you just repeat yourself a lot. If you have a good idea that you think can help other people just tell them over and over again, maybe not to the point that you're actually annoying them because you do want them to listen to you in the future. But if you think you have a solution, let people know.

And they may ignore you at first, or they may think that they still have a better way to do it. But the one thing I will say, and I come back to this over and over and over again, people listen to stories. Yeah, sometimes data helps, sometimes having some numbers to put in front of someone helps, but overall, people like stories. Tell them a story about how this worked for you. Tell them a story about how it made your life easier. Tell them a story about how it saved a company a ton of money or helped them discover some very obscure bug.

And that's what people really connect to. We've been storytellers for a millennia, and that's the best way to get these kinds of things across, I think. And a lot of the book is written that way. There are some chapters that are very heavy in math, and there's an entire chapter about statistics, but even that one has some great stories about dumplings. And we tried to frame the whole book that way. You need a narrative, that's how you engage people, that's how you keep people listening.

Corey: And of course, there's always the option of telling the story on wonderful podcasts like this one.

Alex: 100 percent. I don't think the medium necessarily matters. You can podcast it, you can write a book, you can write a blog post, you can go out to conferences and tell people these stories verbally, or whatever it is. Yeah, I agree. I don't think that part necessarily matters because different people consume information and consume narratives in different ways.

Corey: Well, this has been an absolutely fantastic experience and incredibly educational, at least for me. If people want to hear more about what you're up to, or learn more about SLOs, where can they find you, slash buy your book?

Alex: You can buy the book wherever you want. It's an O'Reilly published book, it's available widely. Go to your local bookstore if you can. If you're not comfortable currently leaving the house, go to bookshop.org that helps support local bookstores. But if you want to order off Amazon because you got Prime, feel free to do that. Just think about, perhaps, supporting small and local businesses. You can find me on Twitter at @ahidalgosre—that’s A-H-I-D-A-L-G-O-S-R-E—where I often pontificate about these kinds of things. And I have a website at www.alex-hidalgo.com.

Corey: Given that you do obviously care about various ways to purchase the book, let's make it easy on people. If you visit snark.cloud/slobook—that's S-L-O-B-O-O-K—we'll drop you onto a site that shows you how to go about purchasing this in a variety of different ways. That's snark.cloud/slobook. Alex, thank you so much for taking the time to speak with me. I really do appreciate it.

Alex: No, thank you, Corey. This has been a great conversation. I've had a lot of fun.

Corey: Likewise. Alex Hidalgo, site reliability engineer, and author. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts, whereas if you hated this podcast please leave a five-star review on Apple Podcasts and a published statement about exactly how many nines we should have had instead of an SLO, in the comments.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Danyel FisherDanyel Fisher is a Principal Design Researcher for Honeycomb.io. He focuses his passion for data visualization on helping SREs understand their complex systems quickly and clearly. Before he started at Honeycomb, he spent thirteen years at Microsoft Research, studying ways to help people gain insights faster from big data analytics.

Links Referenced:

  • Danyel’s section on Honeycomb’s website
  • Danyel’s Personal Site
  • Follow Danyel on Twitter

Transcript
Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

This episode is sponsored by our friends at New Relic. If you’re like most environments, you probably have an incredibly complicated architecture, which means that monitoring it is going to take a dozen different tools. And then we get into the advanced stuff. We all have been there and know that pain, or will learn it shortly, and New Relic wants to change that. They’ve designed everything you need in one platform with pricing that’s simple and straightforward, and that means no more counting hosts. You also can get one user and a hundred gigabytes a month, totally free. To learn more, visit newrelic.com. Observability made simple.

Corey: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the Enterprise (not the starship). On-prem security doesn’t translate well to cloud or multi-cloud environments, and that’s not even counting IoT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IoT devices, detects these threats up to 35 percent faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at extrahop.com/trial.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Danyel Fisher, a principal design researcher for Honeycomb.io. Danyel, welcome to the show.

Danyel: Thanks so much, Corey. I'm delighted to be here. Thanks for having me.

Corey: As you should be. It's always a pleasure to talk to me and indulge my ongoing love affair with the sound of my own voice. So, you have a fascinating and somewhat storied past. You got a PhD years ago, and then you spent 13 years at Microsoft Research. And now you're doing design research at Honeycomb, which is basically, if someone were to paint a picture of Microsoft as a company, knowing nothing internal, only external, it is almost the exact opposite of the picture people would paint of Honeycomb, I suspect. Tell me about that.

Danyel: I can see where you're coming from. I actually sort of would draw a continuity between these pieces. So, I left academic research and went to corporate research, where my goal was to publish articles and to create ideas for Microsoft, but now sort of on a theme of the questions that Microsoft was interested in. Over that time, I got really interested in how people interact and work with data, and it became more and more practical. And really, where's a better place to do it than the one place where we're building a tool for people who wake up at three in the morning with their pager going, saying, “Oh, my God, I'm going to analyze some data right now.”

Corey: There's something to be said for being able to, I guess, remove yourself from the immediacy of the on-call response and think about things in a more academic sense, but it sounds like you've sort of brought it back around, then. Tell me more. What aligns, I guess, between the giant enterprise story of research and the relatively small, scrappy startup story for research?

Danyel: Well, what they have in common is that in the end, it's humans using the system. And whether it's someone working in the systems that I built as an academic, or working on Excel, or Power BI, or working inside Honeycomb, they're individual humans, who in the end are sitting there at a screen, trying to understand something about a complex piece of data. And I want to help them know that.

Corey: So, talk to me a little bit about why a company that focuses on observability—well and before they focused on that, they focused on teaching the world what the hell observability was. But beyond that, then it became something else where it's, “Okay. We're a company that focuses on surfacing noise from data, effectively.” We can talk about that a bunch of different ways, cardinality, et cetera, but how does that lend itself to, “You know what we need to hire? That's right: a researcher.” It seems like a strange direction to go in. Help me understand that.

Danyel: So, the background on this actually is remarkable. Charity Majors, our CTO, was off at a conference and apparently was hanging out with a bunch of friends, and said, “You know, we're going to have to solve this data visualization thing someday. Any of you know someone who's on the market for data visualization?” And networks wiggled around, and a friend of a friend of a friend overheard that and referred her to me. So, yeah, I do come from a research background.

Fundamentally, what Honeycomb wanted to accomplish was they realized they had a bunch of interfaces that were wiggly lines. Wiggly lines are great. All of our competitors use wiggly lines, right? When you think APM, you think wiggly lines displays. But we wanted to see if there was a way that we could help users better understand what their systems were doing. And maybe wiggly lines were enough, but maybe they weren't. And you hire a data visualization researcher to sit down with you and talk about what users actually want to understand.

Corey: So, what is it that people misunderstand the most about what you do, whether it's your role, or what you focus on in the world because when you start talking about doing deep research into various aspects of user behavior, data visualization, deep analysis, there's an entire group of people—of which I confess, I'm one of them—who, more or less, our eyes glaze over and we think, “Oh, it's academic time. I'm going to go check Twitter or something.” I suffer from attention span issues from time to time. What do people misunderstand about that?

Danyel: The word ‘researcher’ is heavily overloaded. I mean, overloaded in the object-oriented programming sense of the word. There's a lot of different meanings that it takes on. The academic sense of researcher means somebody who goes off and finds an abstract problem that may or may not have any connection to the real world. The form of researcher that many people are going to be more familiar with an industry is often meaning the user researcher: the person who goes out and talks to people so you don't have to.

I happen to be a researcher in user research. I go off and find abstract questions based on working with people. So, the biggest misunderstanding, if you will, is trying to figure out where I do fit into that product cycle, I'm often off searching for problems that nobody realized we had.

Corey: It's interesting because companies often spend most of their time trying to solve established and known problems. Your entire role—it sounds like—is, “Let's go find new problems that we didn't realize existed,” which is, first, it tells a great story about a willingness to go on a journey of self-discovery from a corporate sense, but also a, “What, we don't have enough problems, you're going to go borrow some more?”

Danyel: [laugh]. So, let me give you a concrete example if you don't mind.

Corey: Please, by all means.

Danyel: When I first came to Honeycomb, we were a wiggly lines company. We just had the most incredibly powerful wiggly lines ever. So, premise: you've got your system, you've successfully instrumented with the wonderful wide events, and you've heard other guests before talking about observability and the importance of high cardinality-wide events. So, you've now got hundreds of dimensions. And the question was rapidly becoming, how do you figure out which dimension you care about?

And so I went out, I watched some of our user interviews, I watched how people were interacting with the system, and heck, I even watched our sales calls. And I saw over and over again, there was this weird guessing-game process. “We've got this weird bump in our data. Where'd that come from? Well, maybe it's the load balancer. I'll go split out by load balancer. No, that doesn't seem to be the factor. Ah, maybe it's the user ID let's split out by user ID. No, that one doesn't seem to be relevant either. Maybe it's the status code. Aha, the status code explains why that spike’s coming up. Great, now we know where to go next.”

Corey: Oh, every outage inherently becomes a murder mystery at some point. At some point of scale, you have to grow beyond that into the observability space, but these, from my perspective, it always seemed that well, we're small-scale enough in most cases that I can just play detective in the logs when I start seeing something weird. It's not the most optimal use of anyone's time, but for a little while, it works; just over time, it kind of winds up being less and less valuable or useful as we go.

Danyel: I'm not convinced it’s a scale problem. I'm convinced that it's a dimensionality problem. If we only give you four dimensions to look at, then you can check all four of them and call it a day. But Honeycomb really values those wide, rich events with hundreds, or—we have several production systems with thousands of columns of data. You're not going to be able to flip through it all.

And we can't parallelize murder mystery solving. But what we can do is use the power of the computer to go out and computationally find out why that spike is different. So, we built a tool that allows you to grab a chunk of data, and the system goes in and figures out what dimensions are most different between the stuff you selected and everything else. And suddenly what was a murder mystery becomes a click, and while that ruins Agatha Christie's job, it makes ours a whole lot better.

Corey: There is something to be said for making outages less exciting. It's always nice to not have to deal with guessing and checking, and, “Oh, my ridiculous theoretical approach was completely wrong, so instead, we're going in a different direction entirely.” It's really nice to not have to play those ridiculous, sarcastic games. The other side of it, though, is how do you get there from here? It feels like it's an impossible journey. The folks who are success stories have been doing this for many years getting there, and for those of us in roles where we haven't been, it's, “Oh, great. Step one: build an entire nuclear reactor.” It feels like it is this impossible journey. What am I missing?

Danyel: You know, when I started off at Honeycomb, this was one of my fears, too, that we'd be asking people to boil the ocean. You have to go in, tell your microservices and instrument them all, and pull in data from everything. And suddenly, it's this huge, intimidating task. Over and over, what I've seen is one of our customers says, “You know what? I'm just going to instrument one system. Maybe I'll just start feeding in my load balancer logs. Maybe I'll just instrument the user-facing piece that I'm most worried about.”

And we've hit the point where I will regularly get users reporting that they instrumented some subsystem, started running it, and instantly figured out this thing that had been costing them a tremendous amount on their Amazon bill, or that had been holding up their ability to respond to questions, or had been causing tremendous latency. We all have so many monsters hidden under the rug, that as soon as you start looking at it, things start popping out really quickly and you find out where the weak parts of your system are. The challenge isn’t boiling the ocean or doing that huge observability project, it's starting to look.

Corey: It's one of those stories where it sounds like it's not a destination, it's a journey. This is usually something said by people who have a journey to sell you. It's tricky to get that into something that is palatable for a business, where it's, “Oh. Instead of getting to an outcome, you're just going to continue to have this new process forever.” It feels—to be direct—very similar to digital transformation, which no one actually knows what that means, but we go for it anyway.

Danyel: Wow, harsh but true.

Corey: It's hard to get people to sign up for things like this, believe me.; I tried to sell digital transformation. It's harder than it looks.

Danyel: And now we're trying to sell DevOps transformation and SRE transformation. A lot of these things do turn into something of a discipline. And this is kind of true—unfortunately, this is a little bit true here, too. It's not so much that there's a free lunch, it's that we're giving you a tool to better express the questions that you already have. So, I'd prefer to think of it less as selling transformation, and more is an opportunity to start getting to express the things you want to.

I mean, I'd say it sits under the same category, in my mind, is doing things like switching to a typed language or using test-driven development. Allowing you to express the things that you want to be able to express to your system means that later, you’re able to find out if you're not doing it, and you're not seeing the thing that you thought you were. But yeah, that's a transformation, and I don't love that we have to ask for that.

Corey: One challenge I've seen across the board with all of the—and I know you're going to argue with me on this, but the analytics companies. I would consider observability on some level, to look a lot like analytics—where it all comes down to the value you get from this feels directly correlated with the amount of data you shove into the system. Is that an actual relationship? Is that just my weird perception given that I tend to look at expensive things all the time, and that's data in almost every case?

Danyel: No, I actually wouldn't quibble with that at all. Our strength is allowing you to express the most interesting parts of your system. If you want to only send in two dimensions of data—“This was a request and it's succeeded.” That's fantastic, but you can't ask very interesting questions of that. If you tell me that it's a request, and it did a database call, and the database call succeeded, but only after 15 retries, and it was using this particular query and the index of that was hit this way, the more that you can tell us, the more that we can help you find what factors were in common. This is the curse of analytics; we like data. On the other hand, I think that the positive part is that our ability to help find needles in haystacks actually means that we want you to shovel on more hay as the old joke goes, that's a good thing. I agree that that cost center though.

Corey: Yeah, the counter-argument is, is whenever I go into environments where they're trying to optimize their AWS bill, we take a look at things and, “Well, you're storing Apache weblogs from 2008. Is that still something you need?” And the answer is always a, “Oh, absolutely.” Coming from the data science team, where it's almost this religious belief that with just the right perspective, those 12-year-old logs are going to magically solve a modern business problem, so you have to pay to keep them around forever. It feels like on some level, data scientists are constantly in competition to see whether they can cost more than the data that they're storing, and it's always neck and neck at some level.

Danyel: I will admit that for my personal life, that's true. I've got that email archive from, you know, 1994, that I'm still trying to convince myself I'm someday going to dig through and go chase, like, I don't know, the evolution of what the history of spam is. And so it's really vital that I keep every last message. But in reality, for a system like Honeycomb, we want that rich information, but we also know that that failure from six months ago is much less interesting than what just happened right now. And we're optimizing our system around helping you get to things from the last few weeks.

We recently bumped up our default storage for users from a few days to two months. And that's really because we realized that people do want to kind of know what goes back in time, but we want to be able to tell the story of what your last few weeks, what your last few releases looked like, what your last couple tens of thousands of user hits on your site are. You don't need every byte of your log from 2008. And in fact—this one's controversial, but there's a lot of debates about the importance of sampling. Do you sample your data?

Corey: Oh, you never sample your data, as reported by the companies that charge you for ingest. It’s like, “Hm, there seems to be a bit of a conflict of interest there.” You also take a look at some of your lesser-trafficked environments, and it seems that there's a very common story where, “Oh, for the small services that were shoving the stuff into, 98 percent of those logs are load balancer health checks. Maybe we don't need to send those in.”

Danyel: So, I'm going to break your little rule here and say Honeycomb’s sole source of revenue at this point is ingest price. We charge you by the event that you send us. We don't want to worry about how many fields you're sending us, in fact, because we want to encourage you to send us more, send us richer, send us higher-dimensional. So, we don't charge by the field, we do charge by the number of events that you send us. But that said, we also want to encourage you to sample on your side because we don't need to know that 99 percent of what your site sends us is, “We served a 200 in point one milliseconds.”

Of course you did; that's what your cache is for, that's what the system does, it's fine. You know what? If you send us one in a thousand of those, it will be okay, we'll still have enough to be able to tell that your system is working well. And in fact, if you put a little tag on it that says this is sampled at a rate of one in a thousand, we'll keep that one in thousand, and when you go to look at your account metrics, and your P95s of duration, and that kind of thing, we’ll expand out that, and multiply it correctly so that we still show you the data reflected right. On the other hand, the interesting events: that status 400, the 500 internal error, the thing that took more than a quarter second to send back, send us every one of those so we can give you as rich information back about what's going wrong. If you have a sense of what's interesting to you, we want to know it, so that we can help you find that interestingness again.

Corey: This episode is sponsored in part by our friends at Linode. You might be familiar with Linode; I mean, they’ve been around for almost 20 years. They offer Cloud in a way that makes sense rather than a way that is actively ridiculous by trying to throw everything at a wall and see what sticks. Their pricing winds up being a lot more transparent—not to mention lower—their performance kicks the crap out of most other things in this space, and—my personal favorite—whenever you call them for support, you’ll get a human who’s empowered to fix whatever it is that’s giving you trouble. Visit linode.com/screaminginthecloud to learn more, and get $100 in credit to kick the tires. That’s linode.com/screaminginthecloud.

Corey: No, and as always, it's going to come down to folks having a realistic approach to what they're trying to get out of something. It's hard to wind up saying, “I’m going to go ahead and build out this observability system and instrument everything for it, but it's not going to pay dividends in any real way for six, eight, twelve months.” That takes a bit of a longer-term view. So, I guess part of me is wondering how you demonstrate immediate value to folks who are sufficiently far removed from the day-to-day operational slash practitioner side of things to see that value?

Danyel: Mmm.

Corey: It's the, “How do you sell things to the corner offices?” Is really what I'm talking about here.

Danyel: Mmm. Well, as I was saying before, we've been finding these quick wins in almost every case. We've had times when our sales team is sitting down with someone running an early instrumentation, and suddenly someone pops up and goes, “Oh, crap. That's why that weird thing’s been happening.” And while that may not be the easy big dollars that you can show off to the corner office, at least that does show that you're beginning to understand more about what your system does.

And very quickly, those missed caches, and misconfigured SQLs, and weird recursive calls, and timeouts that were failing for one percent of customers do begin to add up pretty quickly to understand unlocking customer value. When you can go a step further and say, “Oh, now that this sporadic error that we don't understand doesn't wake up my people at 3 a.m., people are more willing to take the overnight shift. Morale is increasing, we've been able to control which of our alerts we actually care about.” It feels like it pays off a lot more quickly than the six to nine months range.

Corey: Yeah. That's always the question where, after we spend all this time and energy getting this thing implemented—and frankly, now that I run a company myself, the actual cost of the software is not the expensive part. It's the engineering effort it takes because the time people are spending on getting something like this rolled out is time they're not spending on other things, so there's always going to be an opportunity cost tied to it. It's a form of that terrible total cost of ownership approach where what is the actual cost going to be for this? And you can sit there and wind up in analysis-paralysis territory pretty easily.

But for me, at least, the reason I know that Honeycomb is actually valuable is I've talked to a number of your customers who can very clearly articulate that value of what it is they got out of it, what they hope to get out of it and where the alignments or misalignments are. You have a bunch of happy customers, and, frankly, given how much mud everyone in this industry loves to throw, we'd have heard about it if that weren't true. So, there is clearly something there. So, any of these misapprehensions that I'm talking about here that do not align with reality are purely my own. This is not the industry perspective; it's mine. I should probably point out that while you folks are customers of The Duckbill Group, we are not Honeycomb customers at the moment because it feels like we really don't have a problem of a scale where it makes sense to start instrumenting.

Danyel: This is where my sales team would say we should get you on our free tier. But that's a different conversation.

Corey: Oh, I’m sure it is. There's not nearly as much software as you might think for some of this. It’s, “Well, okay. Can you start having people push a button every time they change tasks?” And yeah, down that path lies madness, time tracking, and the rest? Yee.

Danyel: Oh, God. Yeah, no, I don't think that's what we do, and I don't think that's what you want us to do. But I want to come back to something you were just talking about: enthusiastic customers. As a user researcher, one of my challenges—and this was certainly the case at Microsoft—was, where do I find the people to talk to? And so, we had entire, like, user research departments who had Rolodexes full of people who they'd call up and go ask to please come in to sit in a mirrored room for us.

It absolutely blew my mind to be working for a company where I could throw out a request on Twitter, or pop up in our internal customer Slack and say, “Hey, folks. We're beta’ing a new piece,” or, “I want to talk to someone about stuff,” and get informed, interested users who desperately wanted to talk. And I think that's actually really helped me accelerate what I do for the product because we've got these weirdly passionate users who really want to talk about data analysis more than I think any healthy human should.

Corey: That's part of the challenge, too, on some level is that—how to frame this—there are two approaches you can take towards selling a service or a product to a company. One is the side that I think you and I are both on: cost optimization, reducing time to failure, et cetera. The other side of that coin is improving time-to-market, speeding velocity of feature releases. That has the capability of bringing in orders of magnitude more revenue and visibility on the company than the cost savings approach. To put it more directly, if I can speed up your ability to release something to the market faster than your competition does, you can make way more money than I will ever be able to save you on your AWS bill.

And it feels like there's that trailing function versus the forward-looking function. In hindsight, that's one of the few things I regret about my business model is that it's always an after-the-fact story. It's not a proactive, get people excited about the upside. Insurance folks have the same problem too, by the way. No one's excited to wind up buying a bunch of business insurance, or car insurance, or fire insurance.

Danyel: Right. You alleviate pain rather than bring forward opportunity.

Corey: Exactly.

Danyel: I've been watching our own marketing team… I don’t want to say struggle; I'll say pivot around to that. I think when I first came in—and that was about two years ago—the Honeycomb story very much was, we're going to let your ops team sleep through the night, and everyone's going to be less miserable. And that part's great, but the other story that we're beginning to tell much more is that when you have observability into your system when what the pieces are doing, it's much less scary to deploy. You can start dropping out new things—and our CTO, Charity, loves to talk about testing in production—what that really means is that you never completely know what's going to happen until you press the button and fire it off. And when you've got a system that's letting you know what things are doing—when you were able to write out your business hypotheses in code, and go look at the monitor, and go see whether your system’s actually doing the thing that you claimed it was, you feel very free to go iterate rapidly and go deploy a bunch of new versions. So, that does mean faster time-to-market, and that does mean better iteration.

Corey: You're right. There's definitely a story about, what outcome are companies seriously after? What is it that they care about at this point in time? There's also a cultural shift, at some point, I think. When a company goes from being small and scrappy, and, “We're going to bend the rules and it's fine, move fast break things, et cetera,” to, “We're a large enterprise, and we have significant downside risks, so we have to manage that risk accordingly.”

Left unchecked, at some point, companies become so big that they're unwilling to change anything because that proves too much of a risk to their existing lines of revenue, and long term they wither into irrelevance. I would have put Microsoft in that category once upon a time until the last 10 years have clearly disproven that.

Danyel: I spent a decade at Microsoft, wondering when they were going to accidentally slip into irrelevance and being completely wrong. It was baffling.

Corey: I'd written them off. I mean, honestly, the reason I got into Linux and Unix admin work was because their licensing was such a bear when I was dabbling as a Windows admin that I didn't want to deal with it. And I wrote them off and figured I'd never hear from them again after Linux ate the world. I was wrong on that one, and honestly, they've turned into an admirable company. It's really a strange, strange path.

Danyel: It is. And I definitely—I should be clear, did not leave Microsoft because it wasn't an exciting place or wasn't doing amazing things. But coming back to the iteration speed, the major reason why I did leave Microsoft is because I found that the time lag between great idea and the ability to put it to software was measured in years there. Microsoft Research was not directly connected to product. We'd come up with a cool idea and then we'd go find someone in the product team who could potentially be an advocate for it.

And if we could then we go through this entire process, and a year, or two, or five later, maybe we'll have been able to shift one of those very big, slow-moving battleships. And this is the biggest difference for me between big corporate life and little tiny startup life was, in contrast, I came to Honeycomb, and I remember one day having a conversation with someone about an idea I had, and the next day we had a storyboard on the wall, and about two weeks later, our users were giving us feedback on it and it was running in production and the world was touching this new thing. It's like, “Wow, you can just… do it.”

Corey: The world changes, and it's really odd just seeing how all of that plays out, how that manifests. And the things we talk about now versus the things we talked about five or ten years ago, are just worlds apart.

Danyel: Yeah.

Corey: Sometimes. Other times, it's still the exact same thing because there's no newness under the sun. So, before we wind up calling it an episode, what takeaway would you want the audience to have? What lessons have you learned that would have been incredibly valuable if you'd learn them far sooner? Help others learn from your mistakes? What advice do you have for, I guess, the next generation of folks getting into either data analytics, research as a whole, making signal appear from noise in the observability context? Anything, really.

Danyel: That is a fantastic question. While I'd love for the answer to be something about data analytics, and something about understanding your data—and I believe, of course, that that's incredibly important. I'm not going to surprise you at all by saying that, in the end, the story has always been about humans. And in the last two years, I've had exposure to different ways of human-ing than I had before. I'm sure you saw some of this in your interview with Charity about management.

I've been learning a lot about how to persuade people about ideas and how to present evidence of what makes a strong, and valuable, and doable thing. And those have been career-changing for me. I had a very challenging couple of months at Honeycomb before I had learned these lessons and then started going, “Oh, that's how you make an idea persuasive.” And the question that I've been asking myself ever since is, “How do I best make an idea persuasive?” And that's actually my takeaway because once what a persuasive idea is, no matter what your domain is going to be, it's what allows you to get things into other people's hands.

Corey: That's, I think, probably a great place to leave it. If people want to hear more about who you are, what you're up to, what you believe, et cetera. Where can they find you?

Danyel: You can find me, of course, at honeycomb.io/danyel, or my personal site is danyelfisher.info. Or because I was feeling perverse, my Twitter is @fisherdanyel.

Corey: Fantastic. And we'll put links to all of that in the [00:31:07 show notes], of course. Well, thank you so much for taking the time to speak with me today. I really appreciate it.

Danyel: Thank you, Corey. That was great.

Corey: Danyel Fisher, principal design researcher for Honeycomb. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on your podcast program of choice, whereas if you've hated this podcast, please leave a five-star review on your podcast program of choice, and tell me why you don't need to get rid of your load balancer logs and should instead save them forever in the most expensive storage possible.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Chris PorterChris Porter is the Director of Solutions Engineering at StackRox, the leader in Kubernetes-native container security. Porter has more than 20 years of experience in pre-sales engineering roles, serving and advising customers on security for email, web, cloud, and now Kubernetes and containers. Porter is a certified AWS Solutions Architect and AWS Security Specialist, is the author of a Cisco Press book on Email Security, and holds a Master’s degree from Stevens Institute of Technology.

Linked Referenced:

  • StackRox
  • StackRox Blog
  • Email Chris directly at chris@stackrox.com
  • Follow Chris on Twitter
  • Connect with Chris on LinkedIn

TranscriptAnnouncer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Linode. You might be familiar with Linode; they’ve been around for almost 20 years. They offer Cloud in a way that makes sense rather than a way that is actively ridiculous by trying to throw everything at a wall and see what sticks. Their pricing winds up being a lot more transparent—not to mention lower—their performance kicks the crap out of most other things in this space, and—my personal favorite—whenever you call them for support, you’ll get a human who’s empowered to fix whatever it is that’s giving you trouble. Visit linode.com/screaminginthecloud to learn more, and get $100 in credit to kick the tires. That’s linode.com/screaminginthecloud.

Corey: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the Enterprise (not the starship). On-prem security doesn’t translate well to cloud or multi-cloud environments, and that’s not even counting IoT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IoT devices, detects these threats up to 35 percent faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at extrahop.com/trial.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’m joined on this promoted episode by Chris Porter, director of solutions engineering over at StackRox. Chris, welcome to the show.

Chris: Thank you.

Corey: So, Kubernetes security is the start and the stop of what StackRox does. Now, my impression of Kubernetes security is coming purely from Ian Coldwater’s Twitter feed, so when we talk about the real world of security that isn’t, you know, 280 characters or less per comment, it changes a little bit. Tell me a little bit about I guess, first what you folks do, and also why it’s a hard problem.

Chris: Yeah, so the challenge for Kubernetes is that, of course, it’s a different platform, but you still have the same security concerns. Every platform that comes along has its own nuances in how the applications are running on that platform. In this case, we’re handing a lot of that control over to developers. And what people forget about, as a platform it’s complicated.

And teams get started, and you think about productivity, and you think about how do I get my application running, and how do I have services that my clients can reach? And they’re forgetting that a lot of the power that’s given to the developers actually has security implications. So, we’re trying to raise the attention there, both by talking about it and letting people know, opening some eyes, and providing a software product that—a little bit of a nagging reminder about some of these security settings that they’ve either forgotten about, they’re not paying attention to, or just flat out are unaware of.

Corey: So, I look at this from a somewhat naive perspective, where, okay, I know how to manage application security—well, I don’t; no one does; we pretend we do—and I know how to handle Linux environment security—again, same caveats apply—what needs to change when I start looking down the, I guess, interminable cliff, that is the Kubernetes slide into microservices complexity? What makes it different?

Chris: Well, the tendency is to treat it like every other platform, and then the containers are just little VMs, and I have VMs in the Cloud, and I have VMs on-premise, or maybe I even have physical machines still, a data center. And I want to measure certain signals and understand what activities are associated with an attack. That typically involves things like hunting for vulnerabilities on running machines, and hunting for configuration errors on those machines, and changing it, or patching and keeping those up to date. They’re long-lived.

As they like to say, they’re pets. We care for them, we keep them alive. Containers, the phrase is, “Treat them like cattle, and not pets.” They’re supposed to be disposable, and containers in Kubernetes are no different. They’re quick to start up and quick to throw away.

And one of the, sort of, characteristics is that you don’t try to fix a container, just like, you know, [laugh] treat cattle as a number. And if they get sick, well there’s always another one in line. So, it’s about dispose of and replace these. And then their definition ultimately comes from some code that was used to construct them. So, when you can’t patch a running container—or at least you’re not supposed to—because you lose that patch the next time things get restarted, doing security in the traditional way becomes really hard. You have to look at changing it in the code.

And that’s kind of the mantra behind microservices, that if you’re fixing a bug, if you’re updating a package, if you’re going to incorporate into service, you change the code, you rewrite, and then you roll out that change on a regular basis. And security needs to be thinking in the same way. Like, what happens if you can’t change the running application? How do you do security in that model? And our answer is typically just, you got to go change this in the code somewhere, some setting that you don’t like, or some vulnerability was introduced early in that process, and you have to go to the beginning of that process, fix that, roll out the change.

Corey: So, one of the problems, I guess, at that point is that you’re trying to track down security, both from a proactive perspective as well as a forensics and diagnostic story where the thing that got exploited no longer exists by the time he wound up discovering that it had, in fact, had been exploited. It seems like it does change the story significantly around how this winds up working. Is that accurate? Is that inaccurate?

Chris: Yeah, you’re absolutely right. In fact, we see the power of forensics, of tracking that event, not just for the purpose of reacting to it and sending a line to a very expensive security event monitoring tool to keep around forever, but to teach us a lesson that the event that occurred in our production application that was related to a security incident—somebody installed and ran a crypto miner, somebody put a backdoor in the environment, somebody exploited the application through a known vulnerability, those root causes are lessons that we can learn to apply to prevent it from happening again: to go back to the source code and say, “Well, this attack involves these three links in a chain, and we can interrupt these attacks by eliminating any one of those things.” So, of course, fixing the known vulnerability—if you’re running Apache, you’re responsible for keeping Apache up-to-date, if you’re running a full operating system in your container, there’s a lot of utilities in there, and you’re either responsible for maintaining them or eliminating them. And so picking out each of those individual items is an opportunity; the attacks here, anything that goes a little in disarray in your environment is an opportunity, actually, to go back and correct the problem in the first place by changing the configuration and hardening against those attacks in the future.

We also have this notion that containers are VMs. They’re really not. And if you’re trying to go in and modify them and install user accounts, and modify packages, and do things on them that you would do in a virtual machine, then you’re making a mistake. You’re not using the model correctly, I suppose. But we could take advantage of that.

Signals that would be a normal day-in and day-out on a virtual machine, like a package being installed, actually end up being a tripwire here for what looks like an attack. So, somebody modifying the configuration of that container is either just making a really weird mistake, like trying to go in to add a user account or install some software on there, or of course, it’s a malicious actor who’s just using the tools that are available to go pull down a compiler, compile a crypto miner and start mining Ethereum on your dime.

Corey: Yeah, this does align, I guess, our respected industries here where I talk about billing, you talk about security, functionally one tends to lead to the other, if you often—you’ll find that the way to discover a bunch of crypto mining stuff that’s spun up in your AWS account is via the bill. Now, that only works up to a certain point of scale because you’re spending, eh, when you’re spending a couple tens of millions a month, it’s hard to pick out that giant spike when compared to the giant normal that is the rest of your bill. But I feel like you’re sort of tied, in that billing and security are aligned spiritually in that people care a lot about both of them only after they have failed to adequately care about both of them. It always feels like a trailing action, not like it’s something that people wind up focusing on, front and center. Is that aligned with what you see in the market? Are you hoping to drive an educational story that changes that, or something else?

Chris: Yeah. So, I mean, you hit upon a bunch of hot topics for me personally. One is that I’ve been doing this for a while as a career, where I’ve been working for security vendors, companies who have a story to tell, and a product that addresses the problem, but we always see security is kind of an afterthought. It’s always something that gets brought in a little bit later. And you’re right, it’s not important until there’s a breach, or until your PCI auditor says you’re not even close to being prepared for PCI audit.

So, it becomes a hot topic all of a sudden. And it’s hard to add security in later, especially in a world like this where your teams are prototyping, and developing, and they spend time and effort on that, and then all of a sudden, you bring in a security tool like ours that says, “Hey, these things are all wide open. This is all far too much privilege to grant these applications, and now you’ve got to go back and revise it so that your application works with those settings down.” And so the correct settings are better—it’s easier to deal with that at the start. The other aspect that you mentioned, like the billing thing here, where people don’t see a way to tackle this, it’s a big nebulous problem, and they don’t see any way—like unless they cut their bill by 50 percent, or they’re in perfect security, it’s hard to make the leap to how do I get there, but we’re making the point that you can make a few little changes here and there.

You won’t be PCI compliant in five days, but you can improve your security incrementally, just a little bit, just like you can reduce your costs a little bit. There’s some low hanging fruit, there’s some easy things that you can do, and then there’s topics that I call ‘Advanced Security for Kubernetes,’ which is really nothing more than just, hey, some of the features in Kubernetes can eliminate entire classes of security risk. And so they’re something you have to think about and make sure your applications can work with it, but if you can get there, you’ll make those dramatic improvements. But it’s nice to have stepping stones to get there. It’s not something that you have to do all at once. And risk is not a thing that has to be zero in all cases. I mean, there’s no such thing as security risk of zero, but better security tomorrow than today is a goal in and of itself.

Corey: Yeah, and let’s not get ourselves. Amazon frequently says that security is job zero. Every time I hear that what I really interpret it as is, they came up with a whole list of jobs, realized, “Oh, crap, we forgot security,” and then put it on as job zero so they can pretend they baked it in from the beginning. In practice, I find that almost no one cares about security until they really had to care about security. Does that match what you see? I mean, I’m not trying to make your customers or folks who suddenly find themselves caring a lot about security feel bad for it. I’m just trying to understand how you see the world?

Chris: Well personally, I feel like the world would be a better place if people naturally thought about the advantages of better security. I mean, I for one would not be chasing down credit agencies and card companies because my identity was stolen a few years ago. And believe me, you’d really don’t want to deal with that process.

Corey: Oh, it offends me to have to deal with that process. It’s from a perspective of, “Let me get this straight. You didn’t exercise diligence over who you gave money to, and now you’re trying to make this my problem?” It bugs me on a visceral level.

Chris: That’s right. And a retailer that can’t detect that an out of state license or other ID was used to open up a credit card at a store, and then immediately going to use it to buy gift cards. Those are an attack chain, which we would call it. These are patterns of attack that we know about, and anything that resembles that attack pattern should at least raise some red flags there. So, yes, I mean, I’d love for folks to treat this not as a problem to be solved, but also as an opportunity.

But yes, you’re right about generally, people don’t care about this until either there’s been a breach, or there is some sort of incident that occurs, or realization, or maybe I’ve got to meet some external auditor’s requirements. It could be something industry like PCI, it could be regulatory compliance like GDPR in Europe, but generally, there’s some driving force other than just the desire to have a more secure platform. Now, many organizations have mature security organizations, and they have requirements and goals, and they think about the kinds of data that they want to have, and how they want to respond to that. And I think Kubernetes security can actually fit into that pretty well. It still has very much a big surface area of stuff that you can measure, signals that are interesting to look at, a set of reactions that you can take when you go in there.

It’s just that the nature of it is a little bit different. It takes, often, some learning, I think that changing the way an organization does security to this code-driven model, to this preventive approach is a great advantage, too, but nobody shifts—and especially big organizations—don’t turn on a dime to take advantage of those kinds of things.

Corey: Now, I’ve been trying for a long time to identify the business value of Kubernetes. And it’s pretty clear from what you’re saying that given the fact that you represent a company whose entire value proposition is security for Kubernetes, that improved security posture is not the slam dunk, hit it out of the park narrative of ‘why Kubernetes?’ That I naively had hoped it was. Is this making it worse? I mean, is it still worth going in the Kubernetes direction if… from your perspective, from a business value standpoint, what is the business value of Kubernetes?

Chris: That’s a great question, and there’s probably 100 different answers. We usually tie it into the business goals. When you look at a business, an organization, and if it’s not a tiny little social media startup with eight world-class engineers, most organizations struggle with how do I deliver more value to my customers? How do I get that banking app to support new payment systems? How do I move at the speed that consumers want with mobile devices and websites?

And the model of software development that traditionally started in a dev team isolated somewhere pushing to an operations team to manage it typically ran into a series of challenges in that. And in my background as a software engineer, I never knew who was going to be running it, or what kind of hardware it was going to be on. There were a whole bunch of challenges about that. So, the promise of container and Kubernetes is that I take my environment with me. And I think that putting that control in the hands of the developers to specify everything avoids a lot of that pain that goes from switching environments and probably get into trouble for mentioning this, but a lot of organizations see Kubernetes as a way to provide a multi-cloud strategy. I don’t know if I agree with that, but Kubernetes being a generic, open-source way to specify a platform for running these applications on.

Corey: Oh, yeah. I mean, we saw with the recent re:Invent announcement that, now that EKS anywhere is available, and they’re open-sourcing it, finally, we’re freed from the old-school problem of only being able to run Kubernetes on top of AWS. Wait a minute, no one ever claimed that was the case. What is the value here? It gets more and more confusing, the more you look at a lot of these things. And every deep discussion of why Kubernetes seems to turn into something resembling circular reasoning.

Chris: Yep, absolutely. I think one aspect of it kind of goes unnoticed is something that some much, much smarter people than me came up with a few years ago that I was working with. And the realization that applications are generally subject to the environment in which they’re running in. Now, with Cloud, you kind of get to this state where the application specifies what it needs. I can specify that I need a certain amount of redundancy, geo-redundancy, I need availability zone redundancy, I need a certain amount of storage performance, and network performance.

So, the application really defines what the infrastructure needs to do, and I think Kubernetes does that; we call it this declarative approach: you declare what you want the running environment to be. Now, it’s not quite as easy as you’d like it to be. I’d like to be able to specify things like, what is my mean time between failures, and my recovery time objectives, and other things, other software-defined service-level objectives a little bit more clearly. But we’re climbing up that ladder, and Kubernetes does that, where I can specify some of these things for my application. How many replicas do I want of this?

So, you’re nearing the point where the software platform can actually do whatever it needs to do to configure whatever resources to meet those higher-level goals. So, that declarative nature, I think, is very powerful. And just because Kubernetes has one API for that doesn’t mean it’s the end all be all, but that idea of the software defining what it needs to run is a really powerful one. And then of course that declarative nature is also really important for security. If we know that certain activity is not required for this application, why not declare it to be impossible to do that? If you don’t need to write files—and you shouldn’t be writing files to your container file system—then make it impossible to write files. And then that way we exclude a whole class of security exploits that require writing some sort of payload to a disk and then running.

Corey: Oh, yeah. I agree wholeheartedly with that. One of the value propositions of AWS Lambda for me was, sure it was a bit of a learning curve, but what do you mean, I can’t write local files to the file system within this function? Oh, I can only write to /tmp, and only for this long, and it’s ephemeral. And it’s a platform that’s defined by its constraints.

And on the one hand, while it’s great to be able to say, “Oh, okay, greenfield. I can build around those constraints, and it’s fine.” It feels like the Kubernetes story has always been focused much more around migrating existing applications that never conceived of a stateless model into a cloud-native style world.

Chris: Yep, you’re right. And I’m nodding my head vigorously at everything you’re saying. Like, using the platform to constrain the applications is really powerful. But you’re right, we see a lot of lift and shift, as they call it, or application modernization. Like, oh, I’ll just put this into a container and it’ll just run.

And it’s interesting because Kubernetes originally really didn’t account for stateful applications. It had kind of assumed that you were going to have some nebulous data store—maybe an RDS, or a DynamoDB, or something—it was going to be outside of Kubernetes, and they didn’t really account for any other type of stateful applications in here. It’s been a few years since they introduced it, but it still feels like it was thought of afterwards. So, the idea of lifting and shifting is a great one, but now there’s a lot of teams on the security side of things, a lot of teams think that just putting something into a container naturally makes it more secure.

And I’m, kind of, still trying to shuffle that in my head. I don’t really think so. I don’t think that containers—just like virtualization—were never really designed as security barriers. This certainly wasn’t. And I think we’re just waiting for some of the possible exploits that might happen between containers.

You’ve already seen good examples of container escape attacks that have been shown to be practical. So, use it for what it’s like, but lift and shift is going to be hard. Again, I think you can get better without being perfect. And so if I have an application that was built starting a few years ago with Java on Linux you’ve run on virtual machines, there’s still probably a lot of things that would need to make it perfectly containerizable. But teams can move on that, I think, incrementally; make little changes here and there to improve the capability of resisting an attack when someone exploits a known vulnerability or an unknown vulnerability.

This episode is sponsored by our friends at New Relic. If you’re like most environments, you probably have an incredibly complicated architecture, which means that monitoring it is going to take a dozen different tools. And then we get into the advanced stuff. We all have been there and know that pain, or will learn it shortly, and New Relic wants to change that. They’ve designed everything you need in one platform with pricing that’s simple and straightforward, and that means no more counting hosts. You also can get one user and a hundred gigabytes a month, totally free. To learn more, visit newrelic.com. Observability made simple.

Corey: That’s part of the problem is, it feels on some level like you’re never going to win with security. It’s always something you can continue to improve at and lead to a better place. The problem is, is the journey, not a destination, and a clear lesson we’ve taken from all of this stuff is that you can’t buy security. Counterpoint: there sure are an awful lot of vendors willing to sell it to me. So—

Chris: [laugh]. That’s right.

Corey: —from that perspective, StackRox is not a services company. You’re a software company.

Chris: That’s right.

Corey: You have a security-oriented platform around Kubernetes. What makes you folks different? What makes you not security-in-a-box, or the checklist compliance audit game that doesn’t materially change your security posture? Why you?

Chris: I think we’re pretty realistic about security. And you won’t get—at least for me—a line about StackRox will solve your security problems. We’re there to—

Corey: Sure. You work in solutions engineering, not that baseless marketing. Please, continue.

Chris: [laugh]. That’s right. So, when I’m talking to my potential customers, I have to deliver. So, the next day—we’re a small company—I will show you how the product works, and then the next day I’ll help you get it running in your environment. And so you mentioned a journey and not a destination.

It’s never going to be done. You’ll never have fixed all of those vulnerabilities; there’s always more to do. So, our software really is about designing and enforcing a process, again, using Kubernetes. So, one of things I think that’s different about us is that we’re never going to show you a solution that says, “Run this StackRox library in your application,” or, “Change out this Kubernetes component for this StackRox component.” What we’re there to do is to show you that, hey, your application is running with a very high privilege level, and that means a Unix process privilege level or it could mean a privileged container, but there’s a setting change that you can make.

And even if you can’t change it today, the security team is aware of that’s running because it increases the likelihood of an exploit being serious. And so we’re there to nag you, you as the developer. You are the one who configured this, or you failed to change the default and the default is a bad one for security, so go change it in the source code. It means that after some time following our instructions, hopefully, our product has made your application better. And I’m not sure if I would say this, but you wouldn’t need our product anymore if that was all there was to it.

Changing those settings will help you, and of course, teams will change and they’ll bring on new applications and they’ll have a whole new set of the same problems again and again, so our solution is about nagging you and reminding you that, hey, this is something you haven’t maintained, this is an image you haven’t updated in a while. This is a network setting that is wide open to the entire internet. Now, there’s a call for that sometimes, but you should be aware of it, and not set that up inadvertently that you’re publishing this service publicly without knowing about it because that’s what bites people.

Corey: How do you handle the noise problem, where you wind up with so many different stories about, “Oh, this problem is going to be massive and the rest,” and, “Holy crap, you haven’t rotated your IAM credentials in 60 days,” and versus the, “Oh, by the way, you have no credential set for root and anyone who hits this endpoint can access stuff.” You wind up with the truly valuable important things getting lost to noise. How do you tackle that?

Chris: It’s a big problem. And especially as you’re typically multiplying that problem as you go from an application that might have been on a few VMs, now, to dozens of containers, you’ve potentially got that over and over again. So, the way that we do it, again, is trying to treat this realistically. You’re not going to fix every vulnerability, and the perfect is the enemy of the good, right? That’s the phrase out there.

So, we try to use a measure of the total risk to help you prioritize. Now, prioritizing things may not seem like the smartest thing, like, leaving something for later, but we know that organizations do this anyway. There’s always going to be things that are left until later, like you said—

Corey: “You need to fix these one to three things right now,” gets better results then, “Here’s the 500 list of things you need to do.” This is the problem with Nessus reports is that, historically, they did a terrible job of highlighting what’s the checkbox compliance stuff versus the actual high level of risk.

Chris: That’s right. That’s right. And a real simple example would be that in a Kubernetes cluster, you’ve got all these pods running, which would have containers in them, they’re all potentially on the network, but because you have to kind of declare that in a configuration as to what level of exposure it has, we know what’s sitting behind that ELB; we know the services that are exposed, we know the ones that have other types of ingress, and so there’s an example of prioritization: you got the same vulnerability in 22 different pods, but it’s the front door that’s going to get that probe; that’s the place to fix it first. So, we can use some of the attributes of the environment to help us prioritize that. I mentioned things like privileges.

Sometimes, you’re just going to have to live with a high level of privilege, something needs to run as root all the time. Well, got to be careful with that one. And maybe we think about other defensive tactics around that. But we also want to make sure that it’s not exposed publicly. This is the SSH-open-to-the-world problem in a traditional security group.

Sometimes it is going to be necessary, but you want to keep the awareness high, you want to keep the number of cases where you have to do that low, and if there’s a vulnerability in SSH, you better make sure that that thing is patched on all the EC2 instances that are exposed in that way. So, it’s about prioritization, it’s about being realistic that you can’t fix everything all at once, and a little bit of improvement is better than doing nothing.

Corey: So, when does it make sense for companies to consider bringing you on board? I mean, the easy answer is, “Oh, at the very beginning, when you’re just sketching out ideas on a whiteboard,” yet, in practice, there are so many other competing priorities, that seems unlikely. When does it make actual business sense to bring you folks in?

Chris: Well, I like to come at this from the developers’ perspective. And let’s face it, developers hate security tools because all they do is nag you, they’re noisy, they constrain what you can do. Generally, they have some sort of dashboard or logging system somewhere else that I have to go and look, and I don’t really want to deal with them at all. So, in general, though, if I’ve got a security problem that is hard to unwind, I almost always wish I knew about this earlier. Like, before I start using that Java library, or that version of Ubuntu Linux, or before I start using a configuration from an image that I got from Docker Hub, it’d be nice to know upfront that, hey, that thing hasn’t been maintained in seven months, or that setting is going to cause a privilege escalation problem.

The earlier the better. Your right, teams aren’t going to go out and think about security before they even have a single cluster, but the Kubernetes clusters themselves, there’s some configuration options in Amazon EKS, that you might want to avoid, again, those wide-open settings. So, as early as possible. From a pure software vendor perspective, we’d love to have everybody thinking about this problem from day one, but from the developer perspective, it is nice to have some insight to be able to assess what you’re using, not just for its usefulness, but for how much trouble am I going to get in with the security teams once this thing is running? We talk to a lot of organizations that are basically at the point where this app is ready to go to production and that’s the first time that security ever even heard about this effort. And [laugh] and it’s a little bit hard to retrofit the security at that point. If it requires fundamentally re-architecting the services and finding alternate sources for base images, those are big fundamental changes that can throw off your production delivery date plans.

Corey: So, if you had one takeaway that people could, I guess, carry forward with them from your approach to what you’ve seen, and how all of this works, what would it be? It’s easy to say, “Oh, just buy a product that solves the problem.” But that’s not enough in its own right; there needs to be something that aligns with a fundamental shift in strategy. What takeaway would you give people so that they can start with today?

Chris: So, we like to point out, again, that if you’re choosing Kubernetes, for whatever reason it might be: to make more dynamic delivery of services better, to meet some internal goals, or just because it’s loads of fun, I’d like people to understand that this is a complicated platform, and there is both negatives and positives to that. The negatives are that it’s a complex platform, it has surface area that can be attacked itself, you’re handing over a lot of infrastructure decisions to developers who don’t always make the best decisions, or aren’t always aware, sometimes, of the security implications of those things. But on the positive, there’s so much you can do with it: that you can actually get better security than you could in traditional environments by using, we talked about earlier, that constraining of our applications. Actually, the folks at Google wrote a really nice document explaining some of this, how do you get to better security with something like Kubernetes? So, I want teams to be aware.

And it’s not just us saying this. The platform has a lot of features in it that interact in ways that are not always obvious, and that some of those decisions, some of the defaults, have those security implications. In fact, the folks behind and supporting Kubernetes, the Cloud Native Computing Foundation, actually paid to have a code audit and a security review, a pen test of Kubernetes itself. They did this last year and presented the results at KubeCon. And that’s awesome because there’s a group of people who are really interested in making sure this is a secure platform.

But some of the lines that we like to use about how complex it is, and how hard it is for teams to figure out exactly what’s going on because of multiple layers of abstractions, really point to that security message that we talk about. So, as far as buying a tool, well if you have enough diligence and you have understanding of all of the topics in Kubernetes, you can do a lot of this yourself, but the tools make that easier. Nobody cares about a Kubernetes setting, except maybe with the developer and the one DevOps guy that’s been saddled with the security stuff. The CSO doesn’t care if that Kubernetes setting is being used. But your CSO does care about whether you’re meeting PCI compliance or whether we’re impacted by that new vulnerability we saw on the news.

So, tying those goals down to the individual settings is the job of a product like StackRox. So, we’re there to help you achieve those higher-level security goals. But it comes down to what’s available in the platform. Use the platform for everything it’s worth. That’s the message. Use it for its security value as well as its productivity value.

Corey: Thank you so much for taking the time to speak with me about, well, a variety of things that I don’t tend to spend enough time talking about, according to everyone trying to sell Kubernetes. But that’s a separate problem. If people want to hear more, where can they find out about you and the company?

Chris: Well, we’re easily found on the web at stackrox.com. You can also reach out to me, I’m just chris@stackrox.com. We’re happy to answer any questions. The way that we sell our product is through customers evaluating it in their own environment, so come and kick the tires with us.

If not, a lot of the information about how to use Kubernetes securely in your environment is available on our blog. So, we publish articles about these features and how to make best use of them. And so, that knowledge and our experience is shared for up on the website.

Corey: Thank you. And we’ll of course throw links to where you can be found into the [00:32:28 show notes]. Chris, thank you so much for taking the time to speak with me today. I appreciate it.

Chris: Thank you, Corey, for having me.

Corey: Chris Porter, director of solutions engineering at StackRox. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you hated this podcast, please leave a five-star review anyway on your podcast platform of choice, and tell me why security is not a problem in serverless, and I should use that instead.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

Links Referenced:

  • Follow @SimpsonsOps on Twitter

TranscriptAnnouncer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored by our friends at New Relic. If you’re like most environments, you probably have an incredibly complicated architecture, which means that monitoring it is going to take a dozen different tools. And then we get into the advanced stuff. We all have been there and know that pain, or will learn it shortly, and New Relic wants to change that. They’ve designed everything you need in one platform with pricing that’s simple and straightforward, and that means no more counting hosts. You also can get one user and a hundred gigabytes a month, totally free. To learn more, visit newrelic.com.

Corey: This episode is sponsored by a personal favorite: Retool. Retool allows you to build fully functional tools for your business in hours, not days or weeks. No front end frameworks to figure out or access controls to manage, just ship the tools that will move your business forward fast. Okay, let’s talk about what this really is. It’s Visual Basic for interfaces. Say I needed a tool to, I don’t know, assemble a whole bunch of links into a weekly sarcastic newsletter that I send to everyone. I can drag various components onto a canvas: buttons, checkboxes, tables, etc. Then I can wire all of those things up to queries with all kinds of different parameters: post, get, put, delete, et cetera. It all connects to virtually every database natively, or you can do what I did, and build a whole crap ton of Lambda functions, shove them behind some APIs gateway and use that instead. It speaks MySQL, Postgres, Dynamo—not Route 53 in a notable oversight, but nothing’s perfect. Any given component then lets me tell it which query to run when I invoke it. Then it lets me wire up all of those disparate APIs into sensible interfaces. And I don’t know front end. That’s the most important part here: Retool is transformational for those of us who aren’t front end types. It unlocks a capability I didn’t have until I found this product. I honestly haven’t been this enthusiastic about a tool for a long time. Sure they’re sponsoring this, but I’m also a customer, and a super happy one at that. Learn more and try it for free at retool.com/lastweekinaws. That’s retool.com/lastweekinaws, and tell them Corey sent you because they are about to be hearing way more from me.

Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Jordan and Richard who have no last names. Instead, they are the brains behind the wildly funny @SimpsonsOps Twitter account. Jordan, Richard, welcome to the show.

Richard: It's great to be here.

Jordan: Thanks for having us.

Corey: So, I am astounded every once in a while when I talk to people who don't spend their entire life on Twitter. I don't understand those people, I'm not entirely convinced they're people at all, but that's going to be choices that some people choose to make. Let's start at the very beginning. Well, not the very beginning with, “What is Twitter?” We're going to sort of tiptoe past that. What is the SimpsonsOps Twitter account for those who have not had the wonderful pleasure of experiencing it themselves?

Jordan: I guess the SimpsonsOps account has become, I think more of a way for us, or hopefully other people in the industry to kind of vent about the current state of DevOps and Cloud, and just to really put a bit of snark behind it, kind of the same way you tend to do with your own Twitter account, sometimes. And I think it's really just a way that we can talk about topics that don't usually get talked about, whether it's in the workplace or just generally on Twitter, and having, I think, that pseudo-anonymity sometimes can help give us maybe a bit more of a platform than we otherwise would. Thankfully, Richard and myself are quite good at memeing things, particularly with problems that we have either in the workplace, or in the industry, or just historical things that we've seen.

Richard: Yeah, well, I mean, I think the Twitter account just started as two dudes that just got bored, and we shitpost a lot. We've worked together quite a lot, and, A) we've always been quick to make jokes about things, and B) a little bit self-deprecating, all that kind of stuff. So, it started off as something really, just like… out of the spur, and yeah, now there's so many followers that I feel like we can have a little bit more of a commentary and a platform to vent and make fun of the industry, which is nice. But yeah, it started off as just something very simple and random.

Corey: For those who've never done it, being snarky about industry events is deceptively complex. People tend to view it as, “Oh, it's easy. I just get out there and I make jokes.” It doesn't work that way. It's incredibly easy, especially as you grow to an audience that's more than a couple people around a diner table somewhere, having jokes that don't land right, things that offend people, and wind up punching down inadvertently. And that’s not funny anymore, that's mean.

You folks have not crossed that boundary ever, that I've seen. Everything you do is very good-natured, it's very aligned with what, I guess—the quiet part of what people don't say because they're busy trying to sell you things, and instead, the folks who experience what's going on in the industry see things through a somewhat more cynical—and I would argue realistic—perspective. And you just perfectly capture the sort of zeitgeist of what's happening in the industry at any given point in time. It's incredibly well done.

Jordan: Thank you. [laugh]. I think we've been very conscious to make sure that we keep the account good-natured, and really just have fun with it. It's never really been there, like you said, to punch down on anyone, even though we have had comments in the past that we are probably AWS shills of some kind, or this is some kind of elaborate Amazon marketing campaign.

Corey: Oh, geez. Get in line. I'm always accused of being an AWS shill or having an axe to grind against AWS, usually in response to the exact same thing. You smile, you nod, you move on.

Jordan: Yeah. But we'll send each other what we're thinking of posting on the day, and we'll be like, “Does this work?” And either I or Richard will kind of tell me and say, “Oh, maybe don't say it this way, maybe say it this way.” And I think just having someone to kind of throw ideas at helps a fair bit. And I think that we've kind of explicitly said publicly, I guess, on the Twitter page, that we do post once a day, also means that we do kind of have enough shit to throw at the wall to actually figure out what sticks. And so if jokes do fall flat, which they do, it's fine. We'll just post something else the next day. But it's really just about keeping things consistent.

Corey: One parody account that I've been accused of—multiple times—of being behind is the—I don't know if it's fake Andy Jassy or something else that I want to not get an explicit tag for this episode for, but it's fk_andyjassy and, “Is that me?” No, it absolutely is not; never has been. It's too, I guess, insulting to individuals and groups. It's never been my brand. I like punching up, not down, and a lot of what that particular account does is punching down.

I don't play those games. I don't make fun of people because there's always going to be someone who is offended on their behalf, with the exception of Larry Ellison because, A) he's not really people, and B) no one loves or cares about Larry Ellison in any positive way, so there's no one to take offense on his behalf. But he's the exception. I find that going out and dragging Andy Jassy, for example, is unfair and is not going to end super well. That's never been my brand of comedy. So, a lot of the parody accounts make this mistake and just wind up being mean, instead of funny, I really can't express my admiration enough that you folks have not fallen into that trap.

Richard: Yeah, like Jordan was saying, we're pretty careful about trying not to just be straight up mean. And we kind of try to bounce those ideas off each other. So, we've been pretty good so far, but there's always some people that will get offended or will find your stuff not funny and make comments at you. But we just try to learn from all the stuff we post out. We post quite a lot, so there's a lot of data to work with to figure out, where the line is if we're getting close to it. So, yeah, I mean, there's no real point bringing toxicity to Twitter. I think there's already enough of that as there is.

Jordan: I think the funny part is that being a meme account, you don't necessarily get seen by people as a person—or people—running something, and so while, like Richard said, we may not be very toxic as an account, we tend to experience a lot of that as the account gets bigger. And people tend to say some really unhinged things, particularly because it's not a person, they're not speaking to a person, they're speaking to a meme account, and you end up getting just some very different behavior than you otherwise would from people. And I can only say this as a white male in tech that has experienced things one way my entire life. So, it’s very different.

Corey: one of the things that I admire about you folks, as well, is that I've set up a few meme accounts over the last couple of years because I have great ideas and plans for what to do with them, and I'm right. They're hilarious for about 20 tweets, and then I'm out of content and I don't revisit them anymore; it just doesn't work out. I was waiting for you folks to peter out, but instead, you’ve gotten better with time. What's the trick? What am I missing?

Jordan: Just, consistency. I might have mentioned earlier, but we do very explicitly state on the account, we post once a day, and we set that as above for ourselves to kind of say, “Yes, we are posting something new today.” And that will actually sometimes involve—if I haven't been watching The Simpsons recently, it will sometimes involve looking things up by their own on Disney+ or YouTube and kind of going through things, and being like, “Ah, yes, this might work for me.” And then doing something. So, I think there is a bit of effort involved. But setting that precedent for ourselves, making that promise, almost, to our followers means that we have to kind of look into this and posting. It feels like a job sometimes, I'm not going to lie. I don't know if you feel like that, Richard.

Richard: Yeah, sometimes it does. But I think the other thing is, The Simpsons as a repertoire for making memes is just so rich, there's so much content out there. So, you can almost just look up a random clip of The Simpsons and be like, “Oh, yeah, I can make that DevOps or I can make that Cloud.” So, that's one of the big factors.

Jordan and I have also tried to make another parody account before—which I won't name—but if the concept that you come up with for the parody account doesn't have a lot of content for you to work off of, it's difficult to stay motivated, or it becomes really hard to keep posting stuff consistently. But I think The Simpsons is just a huge repertoire, and it's universal. People from all over the world know about The Simpsons, and like The Simpsons, so we’ve just hit this really nice niche where we have a lot of content that we can work off of and people relate to both the memes as well as the format of The Simpsons.

Corey: There have been something 32 seasons, now, for The Simpsons, what is it—something like that. That's an awful lot of content. How do you categorize all of it? How do you wind up—because you've had stuff that's from recent seasons, you've had stuff that’s from the early days, and most things in between. What does your curation process look like? How do you find this stuff?

Jordan: Man, I don't think we haven't process [laugh]. I think you'll find that most good Simpsons content does come from the golden years and things that I'm probably most likely to recall things that I probably watched during my childhood, which would be the golden years. Being Australians, we did have Simpsons on free-to-air, so we didn't need cable or anything to watch it, and so the nice thing is that everyone kind of got that experience of being able to watch The Simpsons every day at 6 p.m. but I can't say we have a curation process. Unless you do something, Richard.

Richard: [laugh], yeah, I mean, we follow a lot of Simpsons quote accounts and that kind of stuff, and whenever we see anything that we think could be good material, we just kind of bank that and either favorite it, or bookmark it, or something. But yeah, there's no hard process on how we categorize the different content and that kind of stuff. Jordan is a lot more organized than me. He kind of has a lot of content lined up for the future, whereas I tend to do mine a little bit more last minute. So, I just start watching clips, and I find something that I think will work on the day, and then I just post it. Whereas Jordan prepares stuff a little bit more in advance.

Jordan: Depends on the mood, I’m in.

Richard: Yeah.

Jordan: [laugh]. But I am not a smart man; I’m not a funny man; we just have, like you said, Corey, a lot of content to work with. And that makes the creative process a lot easier as opposed to, let's say, like yourself on your own Twitter account where you're constantly posting things, and they usually hit, and they're very good. I don't think I'm capable of doing the same thing.

Corey: Again, not everything's a hit. And sometimes it annoys people. And sometimes I get it wrong, and I inadvertently upset people, and I didn't intend to do it. When I'm called out on that stuff, my general response is to apologize. It's not ideal, but once you hit a certain point of an audience size, you can't say anything without someone getting annoyed by it on some level, and you sort of have to grow a thicker skin, sometimes, if you want to continue to play those games on Twitter.

I'm a white guy and tech people don't come at me in the same way that they would if I weren't a member of an incredibly overrepresented demographic. And that tends to absolutely alter my experience of it compared to an awful lot of folks, but when I look at the responses I get—some of the same people start cropping up, some of the same folks making excellent points show up, and it almost becomes a self-selecting group of people who are at least aligned somewhat in perspective and what my jokes resonate with them, which is super reassuring. I mean, I don't know about you, folks, but my jokes are for me; I find things funny, so I tell the joke. If other people like them, that's just a bonus.

Richard: Yeah so, a lot of the stuff that we post like we mentioned before, is from either the personal experience or seeing it happen in the industry. So, when we send our ideas to each other, a lot of the time, we'll just be laughing at each other's memes and stuff. So, I mean, it really is just, like, two friends sending things laugh at it about each other. And then we ended up posting them on the Twitter account. And I was actually never really big into Twitter, so as we go along, I'm learning the subtle, unspoken rules about Twitter, like what you should do and what you shouldn't do and, like, etiquette, that kind of stuff. So, it's been a bit of a journey, learning what people will be offended over and how best to deal with that. And like you said, I think part of it is just you need to have a bit of a thick skin.

Corey: Yeah, I'm always cautious to give advice like that because, “Oh, people are treating you like crap on the internet? You should just grow a thicker skin.” That is a way of basically sweeping abuse under the table in many cases, and I don't like that. But that's not what I'm talking about at all, just for people listening and wondering what I'm going after here.

It's important to me that people are not going to find every joke funny, and they're not going to like it. People don't come at me questioning whether I'm a person or not. There's none of that. It's always the, “Yeah, I don't like that joke. AWS bills aren't funny.” “Oh, no. You either laugh or cry and things like that, but okay, you don't have to like it.” And I can't afford emotionally to sit awake worrying about stuff like that. It just doesn't go well.

Jordan: I agree that it's very hard to emotionally invest yourself in a lot of the things that other people are saying, and for the most part, we've started avoiding, just engaging in relatively toxic comments from people. What we have done in the past, though, is, if we do find something that really is a bad take, I will retweet it, and I will call it out publicly and say, do this or don't be this guy.

Corey: Sometimes when I have a bad take, I find that one of the best responses to be public and apologize about it. That's happened a couple of times, I did a video a while back of, Hitler reacts to the AWS bill or something like that, and it was fine. A few people were annoyed, like, “You shouldn't be glorifying Hitler,” it's, “Please, I'm the Jewish grandson of a survivor. I get it.” I disagree, fine, whatever.

But what I inadvertently did when I made that video was that there's only one bit of dialogue from women in this entire thing, one woman turns to another and what I had set up was the, it's okay, I get gigabytes and terabytes confused, too. It never occurred to me to frame it that way. It seemed to fit the moment, what I was inadvertently doing, that I didn't realize at the time was, oh, I have effectively set up a dynamic where women aren't good at computers. And when someone pointed that out to me the next day, I was freaking horrified at it, and I took it down and apologized about it, and explained why in a thread. Because when you get it wrong, apologize.

I don't pretend to know all these things in advance, and after a sincere apology, and then let it go. I didn't continue to engage, or defend it, or keep dragging on that entire saga for weeks. I feel like that's the right answer or something directionally close to the right answer. I mean, the best answer, of course, is don't get it wrong like that in the first place, but I'm human. And it comes down to being empathetic, being aware of who your audience is, what message people might be taking away from this, and when you get it wrong, apologize. I don't know why that's so hard for some folks. You don't lose points when you have to apologize.

Jordan: I've kind of messed up that way in the past. I think it was with Deserted Island DevOps thing that I did a while back. But I had that list of people who I am not, or who this meme account is not, and I had it called out to me after I actually gave the talk that I hadn't included any women, which to me was a bit of a wake-up call to kind of say, “Hey, Jordan, get your shit together. You're not really thinking about the things you should be.” So, and I did apologize. I think I apologized as myself for that one, not as the account.

Corey: Yeah, I vaguely remember this. I'm sure it was retweeted by the account, too.

Jordan: Yep.

Corey: Yeah, it absolutely is an easy mistake to make. I mean, I was very conscious when I started this podcast that I didn't want to have it be a whole bunch of white dudes. It takes work to be inclusive. I saw a great tweet the other day of, “Look around and see who's not in the room and get them into the room.” And it's been interesting, going further and further afield trying to find folks who don't have big platforms of their own. It's work; you have to put the work in. But it's worth doing, otherwise, it becomes this horrible echo chamber, everything looks like me. No one wants that. I'm not that good looking.

Richard: [laugh]. Everyone makes mistakes, and by apologizing for it and admitting that you're wrong, it's kind of hard for people to attack you more for it because they realize, “Oh, it's an actual person that’s saying these things.” So, apologizing for stuff publicly is—it kind of like makes you look a little bit vulnerable in a way, and it helps people empathize with you because everyone makes mistakes. So, when someone does something that I don't like, and they publicly apologize for it, I find it very hard to keep hating on them, or it just makes it so much easier to understand what happened and forgive them.

Jordan: Yeah. And we try our best, and taking that kind of feedback on from people is—I would like people to tell me where I'm [BLEEP] up, especially with this kind of stuff. And we have been called out as well, I think, on the account around accessibility is another issue. Just the fact that we're putting text on an image doesn't necessarily mean that everyone can understand or read what's going on in that image.

Corey: This episode is sponsored in part by our friends at Linode. You might be familiar with Linode; I mean, they've been around for almost 20 years. They offer Cloud in a way that makes sense rather than a way that is actively ridiculous by trying to throw everything at a wall and see what sticks. Their pricing winds up being a lot more transparent—not to mention lower—their performance kicks the crap out of most other things in this space, and—my personal favorite—whenever you call them for support, you'll get a human who's empowered to fix whatever it is that's giving you trouble. Visit linode.com/screaminginthecloud to learn more, and get $100 in credit to kick the tires. That's linode.com/screaminginthecloud.

Corey: Yeah, I struggle with accessibility concerns a lot. It is not straightforward or intuitive through alt text, and most of the major Twitter clients and the things that I'm using, that is an ongoing area of concern for me, and I haven't come up with a great answer. Plus, so many of these jokes, when you're doing image-based nonsense is contextual, and there's no good way to explain the joke in writing. I haven't solved that problem, but that doesn't mean that I think I don't need to focus on that and find a better answer. I just don't have one yet.

Jordan: Yeah. And I think that kind of affects us, as well. Like with alt text, we'll try our best to give people, I think, enough context to understand what's going on in the image. We do video format memes as well, and that makes it even harder because there is all this additional work that kind of—not that I mind the additional work, but there is a bit of additional work, particularly, I think there's like that Twitter Media Studio or something now, that you can kind of upload your subtitles, which a screen reader will actually read for you now, which is great. But reminding ourselves that, yes, we need to include these additional people into our jokes is important.

Corey: One other question I had for you, and this is more of a, I guess, shop perspective, but I've finally wound up doing a few ridiculous image memes and whatnot myself, and the way I did that was, for other purposes, I had to wind up finally biting the bullet and getting an Adobe Cloud Subscription. I have thoughts on that, if you work at Adobe reach out, please. We should talk about that. It's not all negative, but I have thoughts on that.

And I have these other things—“Oh, Premiere Rush. What does this do?” And I've started playing around with it, and I am terrible at it. We're talking complete dog-ass. But I have fun doing it. What's your workflow for captioning, turning movies into GIFs, capturing them accordingly, slapping captions on top of still images, what's your toolkit look like? What's your workflow?

Richard: there's a site called Frinkiac, which is basically a site that has a database of a huge amount of Simpsons episodes, and it lets you create image memes, and also video format memes. So, we can create each panel for an image meme or create separate video files that we can then stitch together to make a video meme. So, I think that's where we both start our workflow, and then we use whatever editors we'd like to use to do the stitch together, or add any extra images and stuff. I'm by no means a good image editor either. So, like, I just use Paint.NET on my Windows machine. So, it's very basic, just, like, stitching together images and putting transparent images on top of memes and stuff, to pick things.

Jordan: We quite obviously don't have Adobe After Effects. And it’s been called out on Twitter in the past, with people like, “This video edited is shit. Why do things keep blinking here?” Or, “Why isn't it properly kind of following this animation?” And I'm like I… I don’t know.

Corey: “I am sorry my free content annoyed you.”

Jordan: [laugh]. Yeah. It—

Corey: “It did not live up to your high standard.” Yeah, I hear you on that.

Jordan: And so, I'm like, well, if someone does want to buy me, Adobe After Effects—I'm not saying someone should, but if you're interested, the SimpsonsOps DMs are open. I'm joking, by the way.

Corey: if you work for Adobe and are listening to this show, I think you might want to look up that opportunity, as well. Speaking of big companies here, I was a little bit worried when I first started the Last Week in AWS newsletter, back before anyone had really heard of me because it's like, “So, what's your plan, Corey?” “Well, I'm going to go find a trillion-dollar company, and then I'm going to metaphorically kick them in the junk.” Because that sounds like a smart thing to do that an intelligent person might try.

I was a little concerned for the first year about ceases and desists, a letter from a lawyer telling me to knock that crap off, however you want to pronounce it, and it never happened because it turns out that AWS is way more nuanced than I would have guessed from the big company descriptor. And they have great people that work there. But that was always in the back of my mind of, is this going to cause problems where I’ll have to pivot to Last Week in the Cloud or something down the road? And it turned out that no, but that did hang over me just in the back of my mind for a while. Have you had any concerns or thoughts about that, from the fact that you're doing this on, basically, Simpsons IP? Is Fox going to potentially get annoyed by this and come after you someday?

Jordan: That thought has crossed my mind so many times. And it's also the main reason that I personally will never try and merchandise The SimpsonsOps account. I don't know if Richard wants to merchandise it or monetize it at all someday—but I'm not saying you do. Richard. I'm sorry—but that was really never the intention, I guess, with the account, as well.

I do see a few accounts—which I won't call out—that are trying to monetize shitposting which… I don't know, I feel like I disagree with ideologically speaking. But yeah, oh, Fox coming in one day and saying, “No. You are not allowed to make fun of our friends on this platform. Larry Ellison is a good man.”

Corey: Yeah. I do have the other side of it, where I do monetize what I do. A meme account seems like a weird thing to monetize and almost impossible to monetize well. At some point, people are not, “Well, that funny thing that makes fun of all the Clouds recommends some monitoring project, or Casper mattresses or something.” It feels like on some level, that winds up not working out well.

I've seen a few mainstream meme accounts that start merchandising or going down the path of doing promoted tweets to some content blog or whatnot, and it ruins the feed to the point where I just can't stand it anymore and I unfollow. It's such a jarring disconnect of, I love your original content, but then you retweet garbage clickbait articles in all the time, and that doesn't work. You have to be respectful of your audience, which is something I think that you folks have really stayed true to. Now, if you decide to make money on it someday, I'm also not going to sit here yelling at you for selling out. I get it. I absolutely do. It's just really interesting to be watching the genesis of this account, and watching it gain traction and speed. Big fan of it.

Richard: I 100 percent agree with that, and the purpose of this account is, kind of—I don't think we've ever really properly discussed what the purpose of the account is, but I think we both share the view that it should just be for fun. And by trying to promote things, or sell merchandise, it ruins the feel of the account, and it's probably not something that we’d ever be interested in. Because, yeah, I've seen that happen to heaps of accounts, not just in tech. On like other meme accounts, which are not related to tech, when they start promoting certain products and that kind of stuff, people get really, really upset, and I do the same. I just [00:28:15 unintelligible].

Jordan: I feel like the bar is set at, like, [BLEEP] Jerry, right? Which is probably the worst meme account ever, to date. If you are following [BLEEP] Jerry, please unfollow because they are making money off stolen content. I don't know if our content is necessarily stolen. In saying that, I feel like I'm kind of treading on eggshells here—borrowed.

Richard: [laugh].

Corey: No, I hear you. It's fun, though. It's sort of a cultural touchpoint, and credit where due, I haven't noticed—at least in the time I've been around—that the IP holders of The Simpsons property is overly litigious around stuff like that. I promise you would not be able to do a Disney parody account like this because yes, there is the argument—no, before I get letters, let me call this out: yes, parody is fair use.

Now, do you have the funds to wind up fighting Disney and their attorneys to prove that in court? No. You're going to stop doing it because only a lunatic would actually pick that fight. So, yes, you can do a thing; there's a chilling effect. But by and large, it seems like it's gotten big enough now that it's almost become a common understanding among ops folks, for better or worse. So, good work. Keep it up, please.

Jordan: Thank you. I never expected this. I never expected to be speaking to you, either. So… it's been a weird year.

Corey: We live in interesting times. Thank you both so much for taking the time to speak with me.

Jordan: Thank you. Thank you so much.

Richard: It's been a pleasure.

Corey: So, if you're not on Twitter: you made good decisions. Keep it that way. If you are on Twitter, you can follow them at @SimpsonsOps. It's well worth the trip.

I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you've hated this podcast, please leave a five-star review on your podcast platform of choice along with an image meme on top of Mickey Mouse telling me why this is terrible.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Matt StrattonMatt Stratton is a Transformation Specialist at Red Hat and a long-time member of the global DevOps community. Back in the day, his license plate actually said “DevOps”. Matt has over 20 years of experience in IT operations, ranging from large financial institutions such as JPMorganChase to internet firms including Apartments.com. He is a sought-after speaker internationally, presenting at Agile, DevOps, and ITSM focused events, including DevOps Enterprise Summit, DevOpsDays, Interop, PINK, and others worldwide. Matt is the founder and co-host of the popular Arrested DevOps podcast, as well as the global chair of the DevOpsDays set of conferences.
He lives in Chicago and has three awesome kids, whom he loves just a little bit more than he loves Doctor Who. He is currently on a mission to discover the best phở in the world. You can find him on Twitter at @mattstratton.

Links Referenced:

  • Red Hat
  • Arrested DevOps
  • Arrested DevOps podcast about DevOpsDays Chicago
  • Follow Matt on Twitter
  • Connect with Matt on LinkedIn
  • Speaking Events
  • DevOpsDay Chicago 2020

TranscriptAnnouncer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Gravitational is now Teleport because when way more people have heard of your product than your company, maybe that’s a sign it’s time to change your branding. Teleport enables engineers to quickly access any computing resource, anywhere on the planet. You know, like VPNs were supposed to do before we all started working from home, and the VPNs melted like glaciers. Teleport provides a unified access plane for developers and security professionals seeking to simplify secure access to servers, applications, and data across all of your environments without the bottleneck and management overhead of traditional VPNs. This feels to me like it’s a lot like the early days of HashiCorp’s Terraform. My gut tells me this is the sort of thing that’s going to transform how people access their cloud services and environments. To learn more, visit goteleport.com.

Corey: This episode is sponsored by a personal favorite: Retool. Retool allows you to build fully functional tools for your business in hours, not days or weeks. No front end frameworks to figure out or access controls to manage, just ship the tools that will move your business forward fast. Okay, let’s talk about what this really is. It’s Visual Basic for interfaces. Say I needed a tool to, I don’t know, assemble a whole bunch of links into a weekly sarcastic newsletter that I send to everyone. I can drag various components onto a canvas: buttons, checkboxes, tables, etc. Then I can wire all of those things up to queries with all kinds of different parameters: post, get, put, delete, et cetera. It all connects to virtually every database natively, or you can do what I did, and build a whole crap ton of Lambda functions, shove them behind some API’s gateway and use that instead. It speaks MySQL, Postgres, Dynamo—not Route 53 in a notable oversight, but nothing’s perfect. Any given component then lets me tell it which query to run when I invoke it. Then it lets me wire up all of those disparate APIs into sensible interfaces. And I don’t know front end. That’s the most important part here: Retool is transformational for those of us who aren’t front end types. It unlocks a capability I didn’t have until I found this product. I honestly haven’t been this enthusiastic about a tool for a long time. Sure they’re sponsoring this, but I’m also a customer and a super happy one at that. Learn more and try it for free at retool.com/lastweekinaws. That’s retool.com/lastweekinaws, and tell them Corey sent you because they are about to be hearing way more from me.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by a returning guest, Matt Stratton, who's a transformation specialist at Red Hat. Matt, welcome back to the show.

Matt: Hey, it's really good to be back here. And if getting on a podcast is how you and I talk to each other these days, then I'm all for it.

Corey: Sounds good to me. So, before we dive into this, are you going by Matt? Are you going by Matty? What do people call you these days?

Matt: So, it's kind of funny. Actually, maybe it's not. Audience, you judge if this is funny or not. My friends call me Matty and I decided I was going to for lack of a better word rebrand myself that way, publicly. I was going to start referring to myself as Matty.

So, that's how I had all my profiles, that's how I would write abstracts, and everywhere you find me on the internet, it was as Matty. Well, after about a year of that, I came to a couple conclusions. So, one is after 40 some years of having the name Matt, I can't refer to myself in the third person as Matty; it just doesn't work. Secondly, my family absolutely refuses to call me Matty, potentially because my niece's name is Madison. So, she's Maddie, but also because of probably the same 40 plus years of one thing. But what I also discovered is I really like it when my friends call me Matty. So, my position now is publicly and in writing, I refer to myself as Matt, but Corey, you're my friend; I like it when you call me Matty.

Corey: And Matty, it shall be. What you're saying resonates quite a bit with me, just because I went through a period where I went by first my middle name, and then the shortened version of my middle name, and so at my wedding, I wound up with people who knew me at various points as Corey, Justin or Jay. And it was always fun having to listen for all of those at different times. Then at the end of it, I changed my last name, just because it sounds like I'm trying to flee a dark and twisted past, which, eh, fair enough. But names are important. What people want to be called matters. I'm glad to hear it's not the equivalent of dead-naming if I call you Matty or something—

Matt: It absolutely is not.

Corey: —because you asked me to call you Matty a while back. I went through it, then you're back in public is Matt, and it's one of those, like your two name changes away from, “Oh, that guy.”

Matt: Well, and it also—so can really confuse things where if I know you from my swing dancing days, you call me Mugsy. So, every now and again, I have friends who will always refer to me as Mugsy and then I have to explain that one to the third party as well. So… there's that.

Corey: So, it's been fun. The last time I had you on the show, it was a combination joint episode of Arrested DevOps and Screaming in the Cloud. And that was challenging on a few levels because it had to align with both directions the show went in, and you were sitting across the table from me so if I got too insulting, you were going to punch me in the face, and let's not kid ourselves here, the closest I get to fitness is fitness entire burrito in my mouth, so there's not really a great outcome there. Now, we're separated by at least two time zones so, yay, I can be much mouthier.

Matt: For listeners, it'll come as no surprise that sitting across the dining room table did not color how Corey talked to me in any way, shape or form. [laugh].

Corey: Well, what's fun is that—I don't know if we told this story last time; we're definitely telling it this time—you used to come over here for dinner periodically, back when inviting people to dinner wasn't a deadly risk, and you'd mentioned living in Chicago at the same time that my wife was going to law school in Chicago. And you mentioned that you went to this whole law school musical that my wife was in. She goes into the other room and comes back with the DVD of that show that you had made. You did the filming, you did the producing of it, it was, “Oh, wow, great. So, what I'm really hearing here is you could have introduced me to my wife six or seven years before I actually met her.” And I'm a little resentful of the fact that you didn’t. Never mind the fact that I didn't meet you until years after I met my wife, but that's no excuse, Matty.

Matt: I mean, we find our way where we find, but that was really funny sitting there and Bethany was just talking about this law school musical at, you know, where she went, and I was like, “Oh, I produced the video of that, one year.” And then I started going back through, like, blog archives to find the year when I did it, and that's when she went back and said, “Oh, and here it is, with your name on it.”

Corey: And you had a mutual friend of hers, that was great.

Matt: Yes.

Corey: And it was like, “Oh, yeah. Have you ever met [00:06:07 Darlin]?” Like, “Yeah. We went to her wedding.”

Matt: Yeah. [laugh].

Corey: It is small freaking world. So, let's talk about something that no one is tired of yet, namely COVID and its impact on DevRel. Now, to my understanding, you're not DevRel anymore?

Matt: That is correct. Yeah, I am technically not in—

Corey: You have transformed beyond that?

Matt: I have transformed beyond that. I've rolled out with the Autobots and become a transformation person. But there's still a lot of what I do that feels very DevRel adjacent. And I'm also still heavily involved in the communities I was involved in before, especially around conferences and events, and I still do speaking and all that. So, when we're—we’re thinking about how DevRel has changed with COVID and stuff, as regards to those things, it's the same for me as it is for someone who actually has that on their, quote-unquote, “business card,” if that's the thing people have anymore.

Corey: The problem that I have with online events is that people are sucking at them. They seem to be more or less the same thing. Now, let me be clear: at the time of this recording, re:Invent has not yet started. This is AWS’s own version of Cloud Next, wherein it's going to take three weeks of sessions and whatnot, of a bunch of video being dumped online as best I can tell, and that has the potential to be—how do I frame it—an unmitigated shit-show because, yeah, re:Invent, for those of us who have been there in person, is a lot. It's a solid week of a bunch of content, but the entire world sort of revolves around that in the Cloud-y space.

Now that you're stretching out over three weeks, approximately, no one is going to be allowed, by their employer, to just take those three weeks off and watch a crap-ton of video. Now, I couldn't be surprised, but based upon AWS’s video and online event stuff in years past, I'm a bit of a skeptic. You said you had some ideas on how to make online events more compelling. What have you got?

Matt: Well, yeah. So, there's a lot to unpack there. When we kind of look at the history of this year with events, if you look around that April timeframe, there are a bunch of virtual events that everybody loved, and they were great. I mean, so what I'm about to say is going to sound like damning with faint praise, but you could kind of have done almost anything, and everyone was so starved for doing something that felt at all like a conference that you were fine. And again, if you organized an event that I was a part of in April of 2020, I'm not saying you did a bad job and you only had attendees because of the time, but what happened then is everybody said, “Oh, well, we can just keep doing that.”

And over time, you're finding that again, that appetite that consumers are being a little more discerning maybe is a nice way to put that. And what it all comes down to is this tendency to say what we want to do when we do a virtual event is, “Well, what do we do with our other event, and how do we just throw that online?” And just to be fair, quote, “Just throwing it online” is a lot of work. I'm not trying to say, like, it’s easy—

Corey: Oh yeah, camerawork, and video, and all the rest. Even doing something crappy like that takes a crap-ton of effort, work, time, and money.

Matt: So, I think the thing that we need to think about when you're changing an event to virtual is that you are changing an event to virtual. “It’s like we like to say, getting promoted to manager is not a promotion, it's a career change.” If you're an individual contributor, and now you're a manager, you don't just do what you were doing before the same way but do it at a different scale; you're fundamentally changing your responsibilities and the outcomes you're trying to achieve. And that's the key. So, I’m on the—the founder of DevOpsDay Chicago, this was our seventh installment—if you will—of that event, and the first time we were ever doing it virtually.

And I'm fairly proud of how we did it. I've said before, the general feedback we got from the event—which was in September—was, “I hate virtual events. I still hate virtual events. But if we have to have them, I would like them to be like this.” And that's about the best I could ever possibly hope for.

And the reason I use this as an example—because a lot of folks are asking me, “Well, how did you do it?” And they want to know how did you set up Discord and set up the bots. That's all fine, but I think it's important to look back to how we approached it. And when we were deciding whether or not—do we cancel the event, or do we make a virtual event? And this was way back in, basically, the month of April is when we were making this decision for a September event.

We said we would only do this if we felt like we could provide an experience that was in the same spirit. And our event, and DevOpsDays in general, are very participant-forward. They're very much about interaction, and you'll notice this is the only time throughout the rest of the day that I'm going to use the term ‘hallway track,’ and I'm only going to use it to say, like, I wish we'd stop saying that. And there's a reason.

Corey: Oh, I’m right there with you. It feels absolutely like it minimizes the value of talking to humans in a bi-directional way. Or many to many, to be honest.

Matt: And it's an implementation, it's not an outcome. The outcome isn't to have a hallway track, the outcome is to have genuine interactions with participants. So, when we were deciding this—and we spent about a month—I said, “Look, while we're making this conversation, I don't want to hear about a single specific platform or technology or anything. Let's take some time and think about what the outcome is that we're trying to accomplish.” Because when you start a problem—and this is true with everything we do as engineers—you lead with the tech, you're losing the art of the possible.

And we never would have landed with the implementations we did if we started thinking about it that way, and the experience would have been different. So, for us, it was a lot about how do we create a space where people can have this interaction? And I realize that community events are a little different than maybe the more marquee or the larger events that definitely are much more about people talking to you, or at you. For us, so we run a single track event, talks in the morning, open spaces in the afternoon, and when I build a program for our traditional event, I always consider that the talks are simply kicking off points for these open space conversations. It's so that people have a common thing to be talking about later when they're interacting with each other because that's the real value.

I mean, Andrew Clay Shafer said with the first DevOpsDays, he wanted them to just be all open spaces but he knew nobody would get their company to let them go if there wasn't at least some people standing on a stage. And so that's the thing; what you have to do when you're talking about your pivot, is again, you're changing the event. It's not the online version of that event, it's a new event that happens to be online, and we find ourselves falling into this trap quite a bit.

There's a lot of things that we do, I hate to say, regular, in a in-person event that are simply because of the laws of physics. It's not feasible to move 1000 people to different rooms every 20 minutes in person. It takes 20 minutes just to move them, so you do things like chunk up your talks, then have your open spaces or things like that. But then in a virtual setting, you can do that. You can move things around.

And I think to give a little credit to the extended idea that maybe re:Invent is doing—that's a little bit of that idea is, like, “Oh, conferences are day-long events because you can't have people flying in and out for an hour every Tuesday for a month.” So, that's a little bit about, I guess, removing that. But when all we're doing is saying, “Okay, we always have talks, and they work like this, so let's just stream them instead,” you're not providing anything new. And you're not taking advantage of the ability to get these folks working together or talking to each other. And again, the reason I said the H-T word is because what people are trying to do is they're trying to re-implement the mechanism that that happened.

Well, let's make it so there's a virtual lunch line, or we'll randomly pair people up or something like that. But you're like, well, those experiences also aren't as random as we like to think they are. It's not about just Corey and I get paired up together randomly, and we'll have a great experience. We probably won't, even if it wasn't me and Corey. But you need to have some way that you're facilitating this stuff.

Corey: Right. When it's the same people who show up to all these online events who generally tend to know each other, which yeah, that's kind of the circles we run in, those of us who've done DevRel-ish things for a while. Great. What about someone who's this is their first event? How do they get dragged into the social story? And that's something that a lot of these events are really missing out on.

Matt: Well, and I think that's the thing because they either do it in a way that's not very interactive—I mean, I always say a webinar with a Slack channel, that’s not a virtual event. That's a—okay, again, a very one-way thing, and then you just sort of throw this quote, “community” out there and hope that it happens. You have to have a lot of intentionality. And then you go on the other side, where there's platforms that just sort of will randomly pair up four people and put them in a virtual table so they can just chat. Well, how do those are people that will have something to talk about?

And the thing about all of this is it's hard, which is why people don't do it [laugh]. It requires intentionality. And it's a lot easier to just sort of say, “We do the things the way that we always did them because we have a process around that we have an understanding. And I don't have to educate people.” That was a big thing with us, and I think is going to continue to be a thing.

When I look back at our event, I say, I wish we had done more sponsor enablement for the event, and I don't mean teaching them how to use our platform. Yes, we did that, but there had to be enablement about—you have to change what it means to be a sponsor at virtual event because at a regular event, you can do the, “I’ll throw up a booth with some swag, and if I build it, they will come,” and people will just show up. You have to work harder to get that virtual engagement, and I think that's something that is definitely the responsibility of someone who wants to sponsor an event to think about what they want to do, but as an event organizer, I think it's always really good to kind of help and say, “Here's some examples. Maybe you could do this. Maybe you could do that.”

And also set them up for success—or at least for less failure—by saying, “If you're going to just hang out in the video channel we give you and wait for people to come to you, that's not going to work. You're going to be disappointed with that result. So, here are some things you could do instead.” And over time everybody will learn how to do this and that's great, but it's going to require us taking that extra effort, I think, for a while.

Corey: This episode is sponsored by our friends at New Relic. If you’re like most environments, you probably have an incredibly complicated architecture, which means that monitoring it is going to take a dozen different tools. And then we get into the advanced stuff. We all have been there and know that pain, or will learn it shortly, and New Relic wants to change that. They’ve designed everything you need in one platform with pricing that’s simple and straightforward, and that means no more counting hosts. You also can get one user and a hundred gigabytes a month, totally free. To learn more, visit newrelic.com.

Corey: And that's part of the problem is that the folks who are organizing these events—at least the corporate ones, in many respects—have competing priorities here. Take re:Invent as an example, where the problem that re:Invent has had for years has been that it tries to be too many things to too many people. Is it a giant partner summit? Yes. Is it the expo hall where they drive a bunch of business to their partners? Yes. Is it a bunch of service announcements, sarcastically so? Yes. Is it the biggest community event for the entire global community of AWS users? Yes.

And trying to be so many different things to so many different people is incredibly challenging at the best of times, and now that they're trying to take that online, which parts do you keep? Which parts do you jettison? If you're going to have a whole partner event, why does that need to be co-located online? I mean, there's no reason to get people in the same room at the same time. The track selection is just a nightmare right now. There's a whole bunch of weird problems that companies are running into trying to figure out how this is going to work. And no one knows. So, we're waiting to see.

Matt: Well, and that goes back to the outcomes, right? Because in a lot of ways, big giant in-person conferences are kind of becoming an anachronism because, for the longest time, that was the best way to reach a lot of people simultaneously because you didn't have streaming. And so, if you think back—I'm thinking back to, like, COMDEX days because that's really where this all comes from. Okay, if I wanted to get a bunch of interested people to hear the same message at the same time, okay, you know what? In the ’90s we didn't have streaming; that's what you did.

We said, “Get yourself on a plane and come to Vegas, and spend a bunch of days.” And then we've continued that model, even though it's maybe not the most effective way to do that. And it can actually be very exclusionary. Those keynotes are also streamed, so if you can't get to the event, why did anybody have to be there in the first place other than that's always how we've done it?

And there's a lot of personal things connected to that, which is, it gives them an excuse to go on a trip that work will pay for, so maybe they'll be more inclined to come and listen to your BS, so that could be a thing. But I think we need to really look at all of it and say that the function of several days all in a row that are all together in that way are all based upon physical requirements. And maybe it's not the most effective way.

Corey: I would argue it's almost certainly not. It's hard to do this, and the people organizing these events on the corporate side, too, are also looking at this through a lens of, they have things to sell, they have a narrative to pitch. They're not in this from the attendee perspective, in many respects. When I worked at larger companies that would participate in events like this, the question was very rarely—from the event folks—around, “Well, how can we make this valuable for the engineers attending this for engineering-focused conferences?” And it just felt like it was missing the point. And that's becoming exaggerated and exacerbated by this move to, surprise, everything's online.

Matt: And I think the threat to those big events is community events because they tend to be more participant-focused because the people who are organizing them are probably actually participants of conferences and that's their focus. And so if I can say, “Well, wait a minute, if I can get a better experience out of this, for lack of a better word, community event or something like that—” I know we aren't having velocity, but why do we have to have velocity? Some of the vendor oriented ones because they're controlling the message, so that's not as much a competition, but I think it's kind of a chicken and the egg thing. When I think about, quote-unquote, “corporate events,” the main reason that you're having that is because you want to get a big captive audience to hear your stuff. So, if all you have is, “Come to this event and hear us talk about our products over and over and over again, and announcements,” that's not enough.

So, that's why you have all this extra content. [00:20:22 unintelligible] event that's practitioners talking, and customers talking, and you can learn, and you can do all this because that's all the loss leader that gets them in so they’ll listen to your product sell. Well, if I don't have to come to see you to do that, and if I can get that kind of content somewhere else, why do I go through all the hassle of this long—whether it's re:Invent or something like that—if it's less appealing. And I think there is a huge amount of attendees at a conference that you will lose when it's not a good excuse to go to Vegas. Then they're like, “Well then, really, why do I have to do your thing?”

Corey: Oh, yeah. I don't miss Las Vegas, my God.

Matt: Yeah. [laugh].

Corey: It's a town that’s built on exploitation. Let's not delude ourselves here. So, it always felt super weird to be going there as if it were a tacit endorsement of that. The fact that I don't have to this year is kind of amazing.

Matt: But let's not delude ourselves that there is a large number of people that would feel exactly the opposite. And what you would see as a bug, they see as a feature.

Corey: Yeah. And that is the nature of the world, for better or worse. So, here's a question for you, you were, for a long time, someone who identified as DevRel, now you're not. Talk to me about that.

Matt: So, the funny thing is, in a lot of ways, I feel like I'm just being a little more honest about what I'm doing right now. So, I've kind of been a little bit on the outside of a lot of Devrelians, with being really honest about what we're here for. So, there's a lot of noise made inside DevRel and dev advocacy and stuff, about, “Well, we're not sales. We're not marketing.” As if—almost it feels a little bit like a better than.

And, first of all, I think as an individual, if you want to get yourself connected, either you're building the thing that's being sold, or you're selling the thing that's being sold, and if you're not connected somehow in a measurable way to either of those things, you are the easiest thing in the world to get rid of. So, that's a little bit. But I spent a couple years in DevRel at PagerDuty, and I did so much work with our field, with our sales folks helping with customers and stuff, and I never felt like that was bad. And I think the reason that people feel this way is it feels like they want the community to think of them as very impartial, right? “You can trust me because I'm not trying to sell you something.”

And the reality is, that happens inherently just by virtue of your title. You don't have to actually not do the things. It's kind of a running joke that a customer will tell a sales engineer something they would never tell their account rep because even though you might have sales engineer in your title, you don't feel like a salesperson, so there's more trust, for better or for worse. And I always used to say that the best salesperson at Chef Software was Nathan Harvey, the VP of community because nobody saw him coming. But did Nathan care about Chef making money? Absolutely because Nathan would like to continue working there, right? And it's actually a partnering thing with your field.

So, anyway, the point is, I did a lot of work with prospects and with customers, and it was never connected to the product. So, for example, if I was going to be—again, this is dating us into the past days when we could be traveling places, but let's say I was going to be in Australia for a conference, and I was going to be there for a week because, generally, you don’t fly in and out. I would have 10 to 11 meetings while I was there with potential PagerDuty customers with our sales team, and at none of those meetings was I talking about our product. They would be meetings to talk about, “Hey, how are you doing, incident response?” “What's your digital transformation look like?” “How are you learning from incidents?”

It was all very high-level culture stuff, but it's incredibly powerful because it does a couple different interesting things. It makes the customer say, “Okay, one of the values of if I use your company, and now I have a relationship with you, and I have access to stuff like this that isn't just coming in and doing a sales pitch.” And it also kind of adds to that, like, “Oh, y'all aren’t just trying to sell me something.” On the side of the account team, this gives them a reason to talk to the customer. I can't tell you how many meetings I would walk out of, and the sales rep would say, “I've been trying to meet that CTO for six months. But they took a meeting with you.”

Corey: Oh my God, yes. Done right, DevRel can open doors. I mean, let's not kid ourselves, why do you think I have an interview show podcast? Honestly, it's to get me in front of people that I have no business speaking to, and also you. Because most folks are going to be, “Oh, you want to just talk to me at random? Well, that's weird. No, I got things to do.” “Do you want to be a guest on a podcast?” And people will clear their freaking calendars and be excited to see me rather than their usual reaction of vaguely annoyed.

Matt: Bryan Berry, who was the guy started the Food Fight show with Nathan Harvey many years ago, had a blog post and he said, “The dirty little secret of tech podcasting is this is how you get people to spend an hour talking to you that you could never get that time from them at a conference.” And it's not because they're rude or anything, it's just like you said: it's sort of like, “Hey, Corey, you want to sit down and just talk to me for an hour while you're trying—” No. Of cour—yeah, again, you would love to because we're friends, but even then, you're like, “I got stuff to do.” “Want to be on my show?” “Absolutely, no problem.” And it's very powerful, and it's why I keep doing it: because I get to have fun conversations with interesting people.

Corey: Yeah, that is really what I view DevRel as being. It's, “Yeah, I'm just going to chat with people I like and I'm interested in and oh, by the way, there's an audience.” But it was weird when I started podcasting—still, I don't get a whole lot of feedback in response to these shows because—oh, I get letters when I send out newsletters, but no one calls in to yell at me about something I say on this show. It feels almost like calling into a radio show—something only dangerous lunatics might do—whereas then I go to a in-person conference—in the before times—and I would get swarmed by people who love the show. It's, “Holy crap. You mean the microphone was working?”

Matt: Mm-hm. A funny thing that happens is you feel like you get to know the host of the shows that you love to listen to even if you've never met them. I mean, my relationship with Paul began with him as a voice in my headphones for years. And then I mean, that dude was at my wedding. And I'm not saying just because you listen to podcasts, you'll become BFFs with every host, but you get to know them that way, and then it can be kind of a little weird to—like, when you meet them for real, and you're like, “Oh, you are an actual three-dimensional person, not just the voice of this thing.”

And I want to go back, just real quick, to something you said about what DevRel should be. So, one of the things too, I think that's hard, is kind of defining what DevRel should do is sort of like defining what engineering means. So, what I was explaining about what I did as an advocate doesn't mean that every single developer advocate should do that, nor would they be necessarily good at it, and they're going to be good at things I'm not. So, there's a lot of components to a good DevRel team. Again, it's about being T-shaped and stuff.

But I also do feel like… I’ve become a real big believer in—I don't think he came up with the term but the person who introduced it that I first heard about was John Allspaw about the difference of work-as-imagined versus work-as-done. So, you can talk to an organization about, for example, “How do you do incident response? How do you learn from things?” And they'll tell you all these things, but then you actually talk to the folks who do the work, and they're like, “No, that's not what we do.”

And it's come up two times recently for me, and this is why it’s so interesting how it applies. So, one is about developer advocacy. So, if you talk to a lot of folks in DevRel, they will tell you that, like, “My job is I advocate for the community, and I'm a voice into the product,” and all this stuff. And normally—if that's happening, that's great. You're actually doing your program really well, but—

Corey: Then you say, “Oh, so it's marketing.” “No, it's not marketing.” Says the angry DevRel person who doesn't understand what marketing does.

Matt: Some of it is marketing. That's the thing.

Corey: Yeah.

Matt: there's a lot of stuff. So, but it's sort of like we have this identity that, in my imagined world, I am abstracted away from that filthy lucre of money, and I'm not connected to revenue, and that doesn't matter. But actually work-as-done, it's super is. And it's funny because it also connects to this idea of virtual events.

So, one of the things that's very polarizing when you're talking about doing a virtual event is, should the talks be live or should they be pre-recorded? And I've given this a lot of thought because I did an event and we fought about it for a long time. Actually, we didn't fight about it very long, but we wanted to. And so when you kind of think about it, the advantage of a live talk is certainly much easier for the speaker in terms of your time commitment is when you give the talk. You don't have to pre-record, you have to all this sort of stuff. And it can feel a lot more… well, it can feel more live, because it is.

And you can be like, hey, and if you're the kind of person that writes your slides, the day before the talk, it works out super good for you, so your content can be super-duper fresh. And then an advantage in the pre-recorded is you minimize connection problems, your schedule can be correct because at least if a speaker's long, you knew they were long, and—if you're intentional about this—can interact with the other participants during their talk. And that's actually really cool. I've done that, and I've seen it happen. It's really kind of fun. It's a little weird at first, to be sitting there and watching yourself while you're talking to other people. But you can answer questions in real-time, you kind of chat. It's fun. But here's the sticker. People will tell you, “Well, the reason that I want to do a live talk—”

Corey: Well, hold on a second, let's not kid ourselves. There's a lot of stickers in DevRel.

Matt: There are so many stickers in DevRel. And they're actually hard to get to people now, so it’s—yeah.

Corey: Oh, and the whole conference thing? Yeah, that's a scheme put on by Big Sticker.

Matt: It really is. So, some proponents of live talks will say, “Well, the reason that live talks are better is because the speakers can riff off each other.” Corey can be like, “Hey, earlier this morning when Matty was talking about this thing, that connects right back to what I'm talking about now, and whatever.” And you know what? That is a great example of work-as-imagined versus work-as-done.

So in the physical space, I will not argue this happened. I can tell you some examples of when it happens super awesome. There was one time I was at a conference, I was the opening keynote, Heidi Waterhouse was the second-day keynote, and Ken Mugrage closed it out. And our talks—not intentionally—all connected to each other. And then Heidi and Ken were able to build on that. But you know why that happened? Because we're in a physical event and we had to sit in a room together for two days. Virtual events, you have even more likelihood of a speaker that's going to come give their talk and peace out, and their interaction will be just for their talk. And by the way, if that's what you do, that's okay; I'm not judging you unless there was a different expectation. So, and then virtual events, it's with one very marked exception: every virtual event I've been a part of—either as a speaker or a participant—this has never happened, this reconnecting of talks. And the only one where it happened was Austin's Deserted Island, DevOps talks with the Animal Crossing one, and I always point people back to that event as a virtual event to learn from and to not get distracted by the gimmick that it was connected to a game. So, the reason that all of us, as speakers, talked about each other's talks was we were all in a Zoom together all day long. Not talking, but the presenter’s Zoom, everybody else who was a speaker was welcome to be in there on mute and just listening. And it was kind of a little bit like a speaker lounge. Like, it gave a connection to the speakers. And we were conscious—but that requires intentionality.

Corey: Ah, but counterpoint, too, for folks who aren't speaking, doesn't that feel isolating?

Matt: To be in that Zoom you mean, or to not?

Corey: To not be in that Zoom.

Matt: Yeah.

Corey: when it feels like all the speakers know each other, it feels like it widens the divide between people who are on the inside track and people who are not, and I've always had issues with that with—I try not to spend too much time in the speaker room for that reason, except when I'm building my talk. Once it's done, other than if I'm in there helping someone get ready for theirs, I try and go and socialize with the attendees because not for nothing Matty, I see you an awful lot more than I do the folks I haven't met yet, so I'd rather get the chance to forge new relationships and bring people in. “Oh, I love your talk. I wish I could give a talk.” “Well, guess what, buddy? You can. Let me help.”

Matt: That's absolutely true.

Corey: It's great to go out and meet with folks, and I find a lot of the exclusionary stuff is a little on the strange side.

Matt: So, I'm going to take a little spin on that and say some of the quote-unquote, “exclusionary stuff” actually helps build up new speakers. Not to say that you shouldn't do exactly what you just said because you're exactly right, but you don't hang out in that speaker lounge, but not just to talk to your buddies. I always try to do that. If I'm in the speaker lounge and there's a speaker that I don't know, who's in there working on—I try not to interrupt them, but I will talk to them. It's a chance to bring them into that fold.

And the same thing is true with speaker dinners. And in my couple years when I lived on the conference circuit almost exclusively, I would run into a situation I'm like, “Oh, the last thing I want to do tonight is go to the speaker dinner.” Because speaker dinners are fun, except when you do them every week. But I continually said, “No. You know what? You're going to do this because yeah, for you, ‘Mr. DevRel,’ this is an annoyance and you're bored with it, but most of these speakers, this might be the only talk they're giving this year.”

This is an exciting and special night for them. And not that it's a special night because I get to spend time with Matty Stratton. I'm not trying to put it that way. But the overall thing, when we are, quote, “professional speakers,” we get very jaded about a lot of stuff. Swag is a good example, too.

As an organizer, I might be like, “Oh, the last thing I need is another hoodie that's special for speakers because I get, like, 15 of these a month.” But if your program is good, it's not going to all be people that this is the 15th hoodie they're getting this month. And so for the speakers where this is a special event for them, that stuff matters, and that including them at that level, even though it feels a little exclusionary for me to say including them into that excluding circle, but it also makes it feel special to do that.

So, I think trying to recreate that in some way because you don't have the simplicity of a speaker dinner. Even that's why I think it's nice to have a private channel in the event chat for the speakers because it's just a place to sort of connect about a shared experience, which is speaking at this event.

Corey: Yeah. It's about trying to forge a sense of community when it's difficult to, I guess, reinforce that there is in fact, the community there when there are other communities, just a browser tab away. It's a hard problem, and I have a lot of empathy for people who are going through it. But I think the cultural tolerance we had at the start of this whole pandemic is wearing thin because, sure, you have two weeks to plan a virtual event versus, “You've had nine months. What's the plan here?”

Matt: And we've also as attendees and a community, we have that content. Like, again, I think at the beginning of this, everyone was just excited to have anything because we just had to deal with having a whole bunch of stuff canceled. We didn't even know if we'd ever see each other on the circuit again this year or anything. We didn't even know what it was going to look like, so we're really excited about it and that forgives a lot of rough edges. Which is great, but we don't have the excuse of rough edges anymore, folks. This has been going on. We've seen what works and we've seen what hasn't. And the hard news is, what works is doing a lot more work than you want to do.

Corey: Absolutely. It's a hard problem to solve for, and I don't have any easy answers. If people want to hear more about you, what you're up to, what you're not up to, and what's upcoming in the wild world of Matty Stratton, where can they find you?

Matt: So, I have a blog post I've been working on for about a month, basically ever since DevOpsDay Chicago ended that is all about this. And it's specifically about how we did our event, and by the time this gets published, it will absolutely be done. So, I presume that Corey will include a link to that in the [00:35:44 show notes]. So, go look for that.

Corey: Oh, absolutely I will.

Matt: We also have an upcoming episode of Arrested DevOps in the next week or so, which will be in the past for you listening to this now, where we talk about the event. So, if you want to hear a little more nerdery about that, there's that. But that being said, if you'd like to catch up with the wide world of Matty Stratton—and in pandemic time of being stuck home eating all the time, that world is definitely getting wider—you can find me on Twitter at @mattstratton. That's probably the absolute best place to connect with me and find me. You can find me on LinkedIn, I tend to be pretty polite and reply to messages and stuff, but that's not the best place to find me, but I'm fairly easy to find over there.

And if you want to see where I'll be upcoming speaking as much as I do, which I'm doing a lot less than I was in the past for lots of reasons, if you go to speaking.mattstratton.com you can find all my past talks, upcoming things like that. And if you'd like to hear me talk on a microphone without Corey most of the time, our podcast, Arrested DevOps, one of the longest still-running DevOps podcasts, arresteddevops.com. We have episodes every couple weeks.

Corey: most podcasts are measured in minutes. That one's measured in years.

Matt: It is. It's really weird.

Corey: [laugh].

Matt: Like, I've actually thought about, maybe we're time to be done, and I just—I can’t. It's too much of a thing, you know? So, we're keeping on going, and we've got a bunch of great content coming for you. But yeah, come find me on Twitter, and only about half of what I post is taking shots at Corey.

Corey: The other half is responding to the shots I've taken at you.

Matt: exactly. Because it's a bi-directional medium. It’s a conversation.

Corey: Exactly. [laugh]. Matty Stratton, transformation specialist at Red Hat. Thank you so much for taking the time to speak with me today.

Matt: Thanks for having me on. This was fun.

Corey: I am Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on your podcast app of choice, whereas if you hated this podcast, leave a five-star review on your Apple Podcasts or any other podcast app of choice you use, along with a comment telling me what your middle name is.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Shelby SpeesShelby Spees has been developing software professionally since 2015 in a range of domains, which has made her appreciate the importance of learning how to learn and creating support systems for lifelong skill development. When she’s not helping teams level up their observability practice, you can find her at home playing on her Switch or singing karaoke with her rescue pitbull Nova.

Links Referenced

  • Follow Shelby on Twitter
  • Connect with Shelby on LinkedIn
  • Shelby’s Personal Site
  • Email Shelby directly at shelby@hey.com

TranscriptAnnouncer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored by our friends at New Relic. If you’re like most environments, you probably have an incredibly complicated architecture, which means that monitoring it is going to take a dozen different tools. And then we get into the advanced stuff. We all have been there and know that pain, or will learn it shortly, and New Relic wants to change that. They’ve designed everything you need in one platform with pricing that’s simple and straightforward, and that means no more counting hosts. You also can get one user and a hundred gigabytes a month, totally free. To learn more, visit newrelic.com.

Corey: This episode is sponsored by a personal favorite: Retool. Retool allows you to build fully functional tools for your business in hours, not days or weeks. No front end frameworks to figure out or access controls to manage, just ship the tools that will move your business forward fast. Okay, let’s talk about what this really is. It’s Visual Basic for interfaces. Say I needed a tool to, I don’t know, assemble a whole bunch of links into a weekly sarcastic newsletter that I send to everyone. I can drag various components onto a canvas: buttons, checkboxes, tables, etc. Then I can wire all of those things up to queries with all kinds of different parameters: post, get, put, delete, et cetera. It all connects to virtually every database natively, or you can do what I did, and build a whole crap ton of Lambda functions, shove them behind some API’s gateway and use that instead. It speaks MySQL, Postgres, Dynamo—not Route 53 in a notable oversight, but nothing’s perfect. Any given component then lets me tell it which query to run when I invoke it. Then it lets me wire up all of those disparate APIs into sensible interfaces. And I don’t know front end. That’s the most important part here: Retool is transformational for those of us who aren’t front end types. It unlocks a capability I didn’t have until I found this product. I honestly haven’t been this enthusiastic about a tool for a long time. Sure they’re sponsoring this, but I’m also a customer and a super happy one at that. Learn more and try it for free at retool.com/lastweekinaws. That’s retool.com/lastweekinaws, and tell them Corey sent you because they are about to be hearing way more from me.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Shelby Spees, a developer advocate at a company called Honeycomb. Shelby, welcome to the show.

Shelby: Thank you so much. It's great to be here.

Corey: It really is, isn't it? It's fun. Folks who were doing developer advocacy in previous years would always have a hard time arranging time to talk on a podcast because, “Well, I'd like to, but I'm about to hit 300,000 flown miles this year,” and going from speaking engagement to speaking engagements. Now that we're all trapped at home, it turns out, it's a little easier to get on people's calendars.

Shelby: Yeah. It's been really interesting because I did a lot of internal speaking and lunch and learns, and workshops, and things like that at previous jobs, but this is my first time being a developer advocate. And I started this job in March, right as we were starting to do lockdown, so I never even got to be really part of the conference circuit. And so, there's a lot of things I'm learning now that I do speaking at events that some people take for granted because they've been doing events for years, and other things are brand new to them because everything's virtual now. And so, I think that the biggest challenge has been realizing if you give a 30-minute talk in the morning, and then someone signs up for a podcast or a webinar or something in the afternoon, that ends up being really draining.

And there's been a couple days over the last couple of months where we realized I have four speaking engagements on my calendar today. Like, that's going to be a lot. But for the most part, it's been really fun, and figuring out that transition between being public-facing and speaking. And then going back to my content work, or interacting with my teammates and things like that. Just having that all in one place is really convenient, at least. And I have a really nice desk setup now. So, [laugh].

Corey: Yeah, I think that's something that we're never going to be able to go back to in corporate offices. I mean, I discovered this a few years early when I started my company and realized that, “Wow, okay, I can bootstrap this and run it for a while, but I'm certainly not going to rent office space in San Francisco because, money.” So, instead, I'm going to go all out and I put some sarcastic budget toward decorating out my home office, and I love what I've built; it's awesome. And it was—to be very direct—dirt cheap, particularly in light of what it costs to hire people.

So, I look at this now and then I look at all of the crappy startup offices I've worked in in the course of my career, and it's, “Wait a minute. You were fighting against getting me a nice standing desk or a good chair? Instead, you got the other one because you wanted to save $300. Are you nuts? Look at how much time I spend between sitting in the chair, and working at my desk, and yelling at my computer screen.” Why cut money in the dumbest possible things? I don't want to go back to it.

Shelby: Yeah, I know. I realized that working remote isn't for everybody, but I do feel that for a lot of us, it's been a big improvement. I specifically wanted a remote role, and this is my first time working remote, so it just happened that I started when everyone else was going remote. And so it was really interesting because it was my first time working remote I didn't own a desk. I barely used my computer at home, I just did everything at work. Occasionally I would have to go online for on-call or something, but otherwise, it was just me at my kitchen island or me on the couch.

And so, when I started Honeycomb, right before I started, I talked to our HR manager, and I'm like, hey, is there, like, a budget or something so I can expense a desk? And this was the first time they encountered that. So, I mean, now we've onboarded a whole bunch of remote people in the last few months. But I'm responsible for us having even a protocol around that. And so that kind of felt cool, and I was happy to work with them.

You know, what's a reasonable amount of money to spend on a standing desk, or how many monitors does it make sense to have? And so Liz has been a big supporter there, where she was like, “You should make sure you get a really good mic and a really good camera because you're going to be doing a lot of events and stuff.” And so it's been great. And I think remote work really does suit me because I'm very sensitive to background noise, I’m very sensitive to other people having conversations around me, and when I first started working in software, I was supposed to share an office, but my officemate worked in a lab, and so she was just gone all the time. So, I basically had this gorgeous office all to myself with a window that overlooked a courtyard, and it was the best possible scenario. But I could still hear my office neighbors talking on the phone during the day. And so I started wearing noise-canceling headphones and stuff.

And you think like that shouldn't be necessary if you have an office with a door that closes, right? And so I found myself not really being able to focus until, like, 3, 4 p.m. and I would just work from four to seven because I would have so much trouble focusing during the normal workday. So, yeah, this really suits me now because I can just be flexible about what works with my brain.

Corey: That's something I'm learning is that it's difficult for people to say that, “Oh, I love remote work,” or, “I hate remote work,” because the thing that really surprised me when I started doing it full time has been that I don't actually want to be remote all the time. And I don't want to be sitting in my coworkers’ laps all the time. There's some areas where I want to be alone, and can work independently, and get things done super well. I'm emphasizing that type of work during the pandemic. There are other times where it's sitting in the same room as someone, having a conversation collaboratively is incredibly valuable. Both schools of thought are completely valid. The problem is when people try and build it into the same one, into this unified way of approaching all things must be like this all the time. That's where you run into trouble.

Shelby: Absolutely. And what's exciting to me, or the opportunity that we've had during the pandemic is figuring out how to do that in the same room collaboration stuff when we can't be in the same room. And my teammates have shown me so many cool tools, and strategies, and with pairing on code stuff, or doing product design or prioritization, meeting using sticky notes and a mural board and things like that, where it ends up being almost more productive than it probably would have been if we were in the same room, just because the tools are so well-designed for that use. But there's also just the camaraderie stuff. I'm figuring out right now, is it safe for me to just go walk to, like, a cafe and sit outside a cafe, one day a week? Because sometimes I just need that change of scenery, I just need it to be on a different screen size for my brain to process it differently. And so I miss that; I miss just the variety.

Corey: And that's part of the thing that I noticed when I started getting deeper into, I guess, what you'd call DevRel these days, has been that the role is different every time. And some days I'm on stage speaking to a bunch of people-less so these days—other times I'm depressingly speaking to a very silent, and more than a little judgy camera that doesn't give me any feedback or energy whatsoever, and other times I'm writing, and other times I'm building something myself to see how it works. And no two days look alike, and for some things—I don't understand, now that I'm doing a fair bit of coding, how the living hell people can do this effectively or well in large open-plan offices. And I look back at the times that I was coding professionally, and I have distinct memories of taking an almost sarcastic number of sick days or working-from-home days, not because I was sick: because it was the only way I could get an isolated area where I wasn't getting distracted constantly.

Shelby: Yeah. That was actually something that came up at my last job when they wanted to reorganize. We had two floors, we had the sales and marketing and media floor, and then upstairs was the tech and product floor. And so, everyone talked about how the third floor was just silent all the time, and we loved it. And then at one point someone moved and came upstairs and started putting on satellite radio things and I looked at my teammates, and it's like, “Okay, no. Mutiny. This isn't going to happen.”

And then you go downstairs, and people are talking, and there's background music, and it's just a very different kind of work. And I think that's okay. And some people can code with background music. I can code as long as there's no lyrics to what I'm listening to, and as long as it's not catchy. But everyone has different ways that they can work, and I think the most important thing—and the reason why open offices are just going, the way of the dodo—is it's so hard to accommodate everyone's different work styles. You can't please everybody, so just give them the opportunity to make the office space that works for them.

Corey: I think that's actually one of the more empathetic and humane approaches to things. I've just never understood why, when you're hiring people who are incredibly talented and not to mention, incredibly expensive, you don't have a conversation with them as a part of that of, “Hi, now that we've made the offer and rest, great, how do you work most effectively?” And then do whatever you can in order to empower that. At some level with junior folks, the answer may very well be, I haven't worked in enough workspaces to really have a good idea. In other cases, it’s great. When I'm doing this working in a conference room is great. During incidents, I like having a room with all the relevant people in, et cetera, et cetera.

Shelby: Yeah.

Corey: But I want to bring up something else that is fun. You and I both spend a—how do we say this—unfortunate amount of time on the Twitters. And you, at the time of this recording, recently had a tweet that went out that I replied to, and oh, boy did I get letters. I don't often get it wrong on Twitter, but I am convinced by the aspects of who has responded to my take on this, that I am in the wrong on this. Specifically—and what you said was completely innocuous—of a website where you list your speaking engagements historically.

Shelby: Mm-hm.

Corey: So, let's start at the beginning there. What inspired you to go ahead and list these things?

Shelby: So, Ian Coldwater, who’s also awesome, was—

Corey: —Ian is amazing.

Shelby: Yes. I looked up to them so much, and they're a big speaker, and—

Corey: —they are technically surprisingly short.

Shelby: Oh, [laugh] good to know. And so, they were tweeting about, it’s both of you I assume that remember all the talks I've ever done. And you quote to me that, and said something about, “Do you remember this tweet that you once made?” And I have written, like, 50,000 tweets or something ridiculous, and—maybe not on this account, but on my previous account—and so it's just like, “Yeah. No, I don't remember that.”

And it's actually been a pain point for me, where I remember being remarkably articulate and smart in some tweet, or some thread at some point in the past, and I can never find it again. And so that's happened enough times that I actually went and created a threads page on my personal site, where I just embed tweet threads because it's too hard to actually keep track of them anywhere else. And it's really simple, it’s just a Hugo embed code and stuff. And at the same time that all this was happening, I was talking to my manager about my speaking strategy and how I'm still learning how to propose talks and apply to CFPs, and come up with good topics and things like that, and he recommended that I try out Notist, which is a service where you can host your slides and link to events and things like that. And so I saw this, and I saw that Matt Stratton uses it and has a sub-domain for his speaking page.

So, I was like, “Well, I'm still a baby DevRel, but nothing wrong on having a speaking page and just keeping track of stuff there.” And I’m, like, a little bit vain. And I just really like having a pretty personal site and things well-organized. So, I was like, “Oh, I’ll get ahead on this, and I can just have everything organized and I'll never run into that problem where I'm trying to find stuff, and it's buried or lost until years or whatever.” So, I don't expect people to find my speaking page and be like, “I need to have her come to my event,” but it's nice to have it for me so that if I want to reference something like, “Oh, here's a talk I gave that I talked about that in more detail.”

And that's usually the way it comes up is I'll be having a conversation with someone on Twitter or someone who's a Honeycomb user, and it's like, “Oh, here's a handy resource that I can point you to.” We also have them on the Honeycomb site, too. There's just another place to store things. So, tell me more about what—like, there was drama?

Corey: Sort of. Indirectly. Because it turned into a bit of a downstream discussion. Because I love the idea, in the abstract, of publishing all of the talks I have given. I would want to remove a few of them just because, honestly, I'm ashamed of some of them, but that's a different story.

So, the problem that I ran into inadvertently was in the statement that I don't generally make my slides available to people after the talk. And the reason behind that, from my perspective, was relatively well-intentioned. It was that, first, my slides are not a report; there are remarkably few words on them and without the context of the presentation, it's not going to make much sense to people. Further, a lot of my slides rely heavily on accompanying sarcasm. And if you just get a picture of the slide, it might look like I'm making a point that I'm not intending to. In some cases, those points might be construed as actively problematic if you remove the context from them, which, from my perspective, makes perfect sense. It turns out, I'm wrong.

Shelby: So, why are you wrong [laugh]?

Corey: Because other people have reminded me, don't always think the same way that other people think. And they say, that's great and all. I might remember vaguely you gave a talk, and seeing the picture will jog my memory. Occasionally, there is useful stuff in one of your slides that would be helpful. And I'm starting to really change my perspective based upon the feedback that I've gotten. It turns out—and this might come as a surprise to some of the worst people on the internet—when presented with new information and given time to reflect, it's okay to change your position on something.

Shelby: Mm-hm. Absolutely. And it makes a lot of sense. And I've had that happen before where I've seen an event and I've listened to someone talk, but there's a point that they made that's like, the slide will ring a bell for me. I mean, that's kind of how I got through college is I took notes and I never really went back and referred to my notes.

But even just visualizing that page of my notes to that page of the textbook, I could answer the question… because I never studied, of course. But yeah, and so for people who—the slides can be a really good resource, and I think maybe I'm a bit more conservative than you about this, but I hesitate to write things in slides that could be misconstrued later on. And this is also, like, I’m new to public speaking a little bit, and I mostly present in a professional—representing my company and things like that. So, my approach would be cover your ass because someone's going to get access to your slides anyway. We also upload my slides to the website and make them available at Honeycomb, and make them available to anyone who wants to download them, and sign up for emails and stuff, and so that's another reason for me to just cover my ass.

But if you're out there and you're an independent speaker, and you're comfortable writing things on slides knowing that it's accompanied with the context, that's a perfectly understandable position. But I think the artifact of the slide deck can be really helpful for a lot of people. And especially, I think there's been so many times when I've wanted to go back and just see a graph, or a chart, or just a list of bullets or something—I'm very bullet oriented—to help jog my memory on a point someone made.

Corey: I guess I've been avoiding looking at this whole issue because I have a perspective on this that I don't know how to reconcile. I firmly believe that accessibility is incredibly important. We are all only temporarily abled, for lack of a better term. Our eyes, our ears will fail if we're fortunate enough to live long enough. And I personally do not have the kind of attention span that enables me to sit down and listen to a full 45-minute-long video; I'd rather read it.

I can consume the information faster. I don't provide transcripts of my talks. I don't do any sort of transcription work. And on the one hand, there is the argument of, well, this stuff is super hard, it's challenging, it becomes expensive, and it's a lot of extra work for, to be frank, most of my talk videos get maybe five views. It's not the primary vector for these things.

I mean, accessibility is important. Every episode of this podcast, dating back to the beginning has been fully transcribed in the show notes, by a human. Especially for the half of them or so, not some random service: we have a named person on contract who does all of these, and they're amazing.

Shelby: Yeah, I've gotten to the point—and I think it was a year or so ago—that I decided I'm not going to post any pictures on Twitter that don't have an alt text.

Corey: I’m so bad at doing that.

Shelby: It's really hard, and it's the standard I hold for myself. And what ends up happening is I almost never post screenshots anymore, and I rarely post pictures now because it's just, I know that I'm going to write alt texts, and if I'm busy, or it's a big long screenshot or something with a lot of text, it's a lot of work to transcribe. But I've also done it when there's something that I think is really, really important, I will sit down on my phone and thumb type all of the text in the screenshot, and switch back and forth between the original context and the Twitter app and stuff. And that's just a standard that I hold for myself. I don't think that's the perfect or ideal way of thinking of accessibility, but I think it can help a lot of people to think of it that way. And I've been streaming less—I have a Twitch channel that nobody watches and—because I feel bad streaming if I don't have captions, and it's sort of, right now, for me to stream independently, it's prohibitively expensive to hire a professional captioner to do live captions, and so—

Corey: Oh, it absolutely is. The reason we have sponsors on this show, in part, is to defray those expenses.

Shelby: Yeah. And so it's sort of this conflict for me between getting content out there that might help more people, but wanting to make it accessible to anybody who might benefit from it. And so the way I see it is if I can't make this accessible to someone who's using a screen reader, or someone who depends on captions, or whatever, then I'm not going to make it available to anybody, which is—some would argue that’s throwing the baby out with the bathwater or something like that, but it's just… this stuff matters.

And I totally rely on subtitles, I rely on transcripts for a lot of things. If you ever watch me live tweeting-I’ll sometimes stream myself live-tweeting, where I have the event open and the live captions open, and my Twitter window open, and I'll go back and reread the live captions to make sure I catch everything. And so, it's just—is a challenge, and I think it's something that we have solutions for, and I think it's less on individuals to make sure it happens. And that's where I'm trying to use my influence at Honeycomb a little bit, is just we have the bandwidth to make sure every single image on the Honeycomb website has alt text. So, I am now enforcing that.

My corner of the website is the blog, and so every image that goes on the blog should have meaningful alt text. I don't get to influence accessibility decisions in the product, but I can do it in my little corner of the world. And just help people start thinking about this stuff. Because that's usually the biggest issue. It’s less that it’s like, “Oh, I don't want people to be able to use screen readers.” It's like, “Oh, well, I never even thought about that before; we just were ignorant of these alternative ways of interacting with our tools.” And so I've gotten to the point where I've enabled the Android screen reader feature to just make sure that alt texts on my tweets actually shows up. So, alt text isn't the be all end all of accessibility, but it's the one that I understand a little bit and I have some influence there. But it's definitely a challenge.

Corey: This episode is sponsored in part by our friends at Linode. You might be familiar with Linode; they’ve been around for almost 20 years. They offer Cloud in a way that makes sense rather than a way that is actively ridiculous by trying to throw everything at a wall and see what sticks. Their pricing winds up being a lot more transparent—not to mention lower—their performance kicks the crap out of most other things in this space, and—my personal favorite—whenever you call them for support, you’ll get a human who’s empowered to fix whatever it is that’s giving you trouble. Visit linode.com/screaminginthecloud to learn more, and get $100 in credit to kick the tires. That’s linode.com/screaminginthecloud.

Corey: It really tends to be. It's hard to make things accessible. It takes work, it takes investment, it takes a focus on doing it in a way that is genuine, empathetic, and meaningful. And right now, I feel that’s something the world needs a lot more of. And I genuinely hadn't considered that there are folks who would benefit from me putting my talks up there.

My secret shame—that also bugs me, and one of the real reasons I never did before—is I reuse talks a lot. If I give a talk at a meetup somewhere—back in-person—and 200 people show up, great. That should not be the only time I give that talk. It takes a lot of work to build a talk that's good. Even years later, if I'll give it a big conference, I'll still give it again because it is new to other folks, assuming it is still reasonable.

As a result, I will also reuse slides, or portions of slides, or even entire deck sometimes for different talks, and I feel like putting all my slides up there is, “Oh, he's not doing as much work as he thinks. He's phoning it in.” Which I understand is a ridiculous fear, but it's still there, the back of my head that voice whispering, “You're a fraud. You've secretly suck at this. They're going to find out.”

Shelby: Yeah. No, I feel like I've been so lucky to work with Charity and Liz and Christine—and now George joined us a little bit after I joined at Honeycomb—where I'm surrounded by people who are very experienced in doing DevRel work. And so I don't even get to question that stuff because they're just like, “Oh, these are the four talks that I'm giving this year,” and I'm doing four talks this year for, like, 30 events, or 400 events, or—Liz has probably done 400 events this year already. And so there's no question about whether you're going to repeat a talk because it's probably a week of work that goes into writing a good talk and making the slides, and stuff, sometimes longer. And so it doesn't make sense for that to be a one-off outcome.

Being able to reuse it and especially—like you said, sharing to a new audience, or even the same audience seeing a talk for the second time, things will sink in better. So, that's something that I haven't even had the chance to question because it was just like, of course, you're going to repeat talks. That's been really helpful. And I've talked to Charity about this, where when she started her speaking career, and she was just figuring everything out by the seat of her pants and stuff, and she said, “You're so lucky you don't have to do all that.” And so I try to acknowledge I really am benefiting from the experience of all my teammates and everyone around me who can give me advice that it's no big deal. So, there's still the things—and I think the hard part for me has been what talks to give and write, and things like that. But, of course, once I write it, it's not set in stone, but it's definitely going to get re-used.

Corey: I think that that's probably the right view to take. And, again, as mentioned earlier, I’m wrong on this.

Shelby: Yeah, it's hard, though. And I think it's harder when you don't h—I don’t know, I just feel so lucky to have the network I have, and I've been able to grow since starting at Honeycomb, where people really believe in me. And it's kind of scary, but it helps keep me a little bit motivated. I don't have to even believe in myself. Because people have these expectations of me, I guess I’d better work on this. [laugh]. I certainly wouldn't be able to do all this as an independent speaker. And so, I don't really understand how people do that. That's super scary to me.

Corey: Speaking takes a lot of work. Something I've noticed is that it's harder for me to do it remotely than it is in person. There’s something about the energy that changes things dramatically. I also don't prep super well for things, so most of my talks are improv. I have bullet points in the speaker notes but that's really about it. And it means that I'm going to give a different talk every time I give it, and I'm going to feedback from the audience. You get no feedback from a teleprompter. I've tried.

Shelby: Yeah. I think of it more like other kinds of performance. And maybe this is the difference is, I grew up doing orchestra, and I never learned how to improvise on violin. I had friends who were like, “Oh, you could totally do violin in jazz band.” I was just like, “Okay, no, that's madness.”

And so it was just like, well, of course I'm going to play something more interesting than Eine kleine Nachtmusik, but it's like you hear Eine kleine Nachtmusik everywhere you go. You hear, “Spring,” or Für Elise, or all these super classic songs everywhere you go, and nobody questions it. And that's okay. It's okay to reuse material because a different performer, or even the same performer at a different event, it'll come out differently. And I think that's helped me with the public speaking side, as well.

It’s just one, like, of the talks I give is a talk that Liz and Danyel wrote. I'm qualified to give it, and I understand it, and it's stuff that I've worked on before, but I didn't write it. And it was only my third speaking event that I actually gave the first talk that I wrote myself. And so it was actually easier to write the talk myself than it was to learn somebody else's talk. But it's the sort of thing where it's like, you put sheet music in front of me and I'll learn how to play it. And so I guess that's helped a lot with my mindset.

Corey: Yeah. That comes down to learning what works for you and what doesn't. And we're all learning as we go with this whole glorious pandemic thing. DevRel is sort of reinventing itself and I kind of like it. It also seems that companies are, “Oh, we're going to do a big online conference for three weeks, and it's a whole bunch of videos that we just throw out there.”

The dynamic is changing; you miss the hallway track for one, which is the most valuable aspect of conferences, at least to me, and two, if I go back in time to my days as an employee of other companies, I can get my boss to sign off on a day or two for a conference, maybe even a week. But AWS doing re:Invent over a three-week span, for example; I don't know too many staff are going to be able to say, “Hey, I'm just going to spend three full weeks watching videos all day and that's going to be okay, right?” You lose something. They might have on in the background and that's great, but that means no one's paying full attention to these things. I don't know how you get around that.

Shelby: Yeah. And honestly, I'm even encouraged to attend events and I get to join on behalf of Honeycomb, or even just as myself, I join the conference Slack and stuff, and I'll live tweet people's talks and stuff, but it's gotten to the point where it's very hard for me to justify it in terms of, okay, this thing is in the middle of my workday, and I have all of these other priorities that I have tighter deadlines, or just are higher priority or whatever, it's been very hard for me to justify attending people's talks. And I want to attend people's talks so bad; people are putting out such good content. And so I can justify some of it because it's synchronous, and I'm there at the same time: I'm in the Slack, I'm talking to people. I really enjoy conference Slacks.

I know they're not for everybody, but the hallway track, I was never really part of it and it sounds super scary to me. I really hate mingling. I hate talking to strangers. I'm sure I would get over it for work stuff, but just the idea of approaching a speaker after their talk is just like, I could never do that. I never talked to my professors. I never went to their office hours.

Corey: I've got to be honest, I have the same exact problem as you do, and that is the reason—I’m not kidding—that I started speaking in the first place. I am terrible at approaching people and striking up the conversation. Give me a 10-second icebreaker and we're fine. It turns out when you're the speaker, people will solve that problem for you and that's what got me started.

Shelby: I really did not grok networking or just even talking to strangers my entire career. And so I feel like I'm in networking cheat mode now where people are just like, “I saw your talk. Let me add you on LinkedIn,” or, “I’ll start following you on Twitter. Let's have this great conversation over Slack,” or whatever. And there's things about virtual conferences that are really exciting for me.

I think being able to bring in attendees who could never afford to fly out to an event and stay at a hotel, encouraging people to sign up for free events, and having company sponsorships, just cover the base-level logistics stuff. I mean, the one big conference I ever attended was re:Invent, which is, like, six conferences in one and so—

Corey: Oh, my God, that is like starting your drug experimentation by effectively, “Just give me everything injected directly into my veins, the entire pharmacy; go with it.”

Shelby: Yeah.

Corey: It’s awful. That is not, [laugh] that is not a conference. That's a monstrosity.

Shelby: It was a lot. I think I only left the hotel, like, twice.

Corey: That's the problem. There are six hotels.

Shelby: Yeah. It was a lot. My idea of a conference is okay, maybe one of those hotels. I'm not going to bother with all of them. And I'm lucky that also my coworkers had been to re:Invent every year, and so they were like, “Oh, okay. Don’t try and get to all the different places.”

They gave me this whole strategy, so I was really lucky there. But yeah, I went to that and I went to a DevOps day in LA, once. And so my understanding of conferences is so warped. But it's so comfortable for me to interact with people online and over text, especially. It gives me a little bit more time to think about what I'm going to say, and especially in groups of more than two or three people, I really struggle to keep up with conversation and be an equal conversation partner.

I get really overwhelmed and intimidated and stuff. And so the mingling and the group conversation circles and stuff, the idea that just scares me. And so I think it’s—not everyone, and especially not people in the DevRel space are going to agree with me, but I think it's encouraging to have online discussions and just bring in more people who are me and who aren't ever going to approach a speaker after their talk.

Corey: Thank you so much for taking the time to speak with me today, talk about how you view, well, several aspects of several different things. If people want to learn more about you, where can they find you?

Shelby: I'm always on Twitter, @shelbyspees is my username. I have a website, shelbyspees.com that I don't really keep updated. Well, I've been playing around with styling stuff, but I'm working on blog posts and things like that. And those are the main two places. Twitter, especially, is great. You can also reach out to me at shelby@hay.com is my personal email, and I check that, like, once every couple of weeks. I'm a very bad pen pal, but I love hearing from people and I'd love to talk more about any of this stuff.

Corey: Excellent. Well, thank you once again for taking the time to speak with me. I appreciate it.

Shelby: It's been wonderful. Thanks, Corey.

Corey: Shelby Spees, developer advocate at Honeycomb. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts, whereas if you hated this podcast, please leave a five-star review on Apple Podcasts along with a comment telling me exactly why I'm wrong and that open-plan offices are the best, while shouting over the loud salesperson sitting three inches away from you.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Alfonso Cabrera
Alfonso Cabrera is the Director of Platform Engineering at Red Ventures, where he helps manage and optimize the extensive AWS footprint. He also spent time at AWS as a Solutions Architect and worked as a DevOps Engineer at a few startups. Alfonso enjoys fostering community and has organized DevOpsDays Charlotte for the past 5 years. Outside of work, he tries to stay in shape by playing sports of all kinds, and gets his adrenaline fix by riding motorcycles.

Links Referenced

  • Red Ventures
  • Follow Alfonso on Twitter
  • Connect with Alfonso on LinkedIn

Transcript
Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Gravitational is now Teleport because when way more people have heard of your product than your company, maybe that’s a sign it’s a time to change your branding. Teleport enables engineers to quickly access any computing resource, anywhere on the planet. You know, like VPNs were supposed to do before we all started working from home, and the VPNs melted like glaciers. Teleport provides a unified access plane for developers and security professionals seeking to simplify secure access to servers, applications, and data across all of your environments without the bottleneck and management overhead of traditional VPNs. This feels to me like it’s a lot like the early days of HashiCorp’s Terraform. My gut tells me this is the sort of thing that’s going to transform how people access their cloud services and environments. To learn more, visit goteleport.com.

Corey: This episode is sponsored by a personal favorite: Retool. Retool allows you to build fully functional tools for your business in hours, not days or weeks. No front end frameworks to figure out or access controls to manage, just ship the tools that will move your business forward fast. Okay, let's talk about what this really is. It's Visual Basic for interfaces. Say I needed a tool to, I don't know, assemble a whole bunch of links into a weekly sarcastic newsletter that I send to everyone. I can drag various components onto a canvas: buttons, checkboxes, tables, etc. Then I can wire all of those things up to queries with all kinds of different parameters: post, get, put, delete, et cetera. It all connects to virtually every database natively, or you can do what I did, and build a whole crap ton of Lambda functions, shove them behind some API’s gateway and use that instead. It speaks MySQL, Postgres, Dynamo—not Route 53 in a notable oversight, but nothing's perfect. Any given component then lets me tell it which query to run when I invoke it. Then it lets me wire up all of those disparate APIs into sensible interfaces. And I don't know front end. That's the most important part here: Retool is transformational for those of us who aren't front end types. It unlocks a capability I didn't have until I found this product. I honestly haven't been this enthusiastic about a tool for a long time. Sure they're sponsoring this, but I'm also a customer and a super happy one at that. Learn more and try it for free at retool.com/lastweekinaws. That's retool.com/lastweekinaws, and tell them Corey sent you because they are about to be hearing way more from me.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Alfonso Cabrera, who is currently the director of platform engineering at Red Ventures. Alfonso, welcome to the show.

Alfonso: Thanks, Corey. I'm excited to be on. I've listened to the show for quite a while now.

Corey: I would suggest it might be time to find a show with, you know, intelligent folks, but then again, that's why I have conversations with other people. Basically, we have this idea in storytelling where you have the hero's journey. I do something similar called the moron’s journey, and I'm the moron. So, it works out super well: every time I talk to someone on the show, I learn something. So, you run platform engineering at Red Ventures. Everyone has ventures; what is a Red Venture? And how did it come to become color-coded?

Alfonso: [laugh]. Great question. So, Red Ventures is a portfolio of influential brands and digital platforms and some partnerships as well. But essentially, we connect millions of people, through our audience with expert advice and information.

We’re in a lot of different industries, so we have some businesses in home industry, in health, in travel, in finance. Some of the brands you may have heard of, things like Healthline, or The Points Guy if you travel a lot, but also things like Bankrate and MYMOVE and creditcards.com.

Corey: Yeah, I get nostalgic and sad about that now because again, I used to live on airline miles, more or less, and now it turns out that no one travels anymore. And I am… I guess I never thought I would actually miss traveling, but here we are.

Alfonso: I know I'm in the same boat. I miss it and ready to get back out there and go travel again.

Corey: [laugh]. So, what makes this fun is that I've met you a few years back when you were not a director at that point; you were in a different role at Red Ventures. I forget offhand what that was.

Alfonso: Yeah, I started pretty early on the cloud ops engineering team over at Ventures about four and a half years ago now. I was one of the first two folks that was helping the company move to the Cloud.

Corey: And then you quit. Because the best thing to do is go work somewhere else, given the opportunity. At least that's always been my philosophy, which is why my resume is a patchwork of tattered broken dreams. But you went to a small company called AWS and acted as a solutions architect for a while. And then you boomeranged back to Red Ventures. So, what was that like?

Alfonso: Yeah, for sure. It was a great opportunity to go to AWS and try out a different career path. I view solutions architecture there as customer-facing, almost similar to a sales engineer type of role. And I had a lot of fun. I learned a ton of stuff at AWS.

Some of the things I really really enjoyed and took away from there that probably will stay with me forever is just the writing culture is crazy to see there. I was really fond of that. I think writing brings a lot of clarity of thought, instead of compared to a PowerPoint deck, for example, where you may have a few words in the slide, but generally, you're kind of just talking through things and when you write things down, whether it's a one-pager, a six-pager or the PR FAQs that they're well known for, it just brings a lot more clarity and gets everybody on the same page. Something that’s definitely stuck with me and I've continued to do as a habit over at Ventures.

Corey: I love your framing of that. Usually when you ask people, “Oh, so at this job, what did you take away from it?” The honest answer is generally, “A bunch of office supplies,” and that's about it. But every time you talk to someone who's left AWS, they talk about, on some level, the lessons that they learned by working there. How much of I guess AWS culture do you find is applicable or portable to other companies?

Alfonso: Yeah, actually, it's funny because one of the other things I really enjoyed there is the principles-based decision-making they do. Obviously, they're very vocal about their leadership principles and reference them in day-to-day decisions, or in meetings and in different projects they’re working on, and that's actually something that I've also seen at Red Ventures. One of the things I love about the culture at Red Ventures is it's very similar. We have belief statements that we discuss often.

Those things are brought up when we're making decisions, or when we're starting new projects, we frame them against our belief statements. And it makes a big difference because it gives some consistency or table stakes to things, to decision-making, the projects you work on, and into discussions you have, and into just the company culture overall. I think it makes a big difference.

Corey: The challenge I always see is that you have these folks who go and look at, “Well, what did Amazon do that made them successful?” And their response is, “Oh, okay.” So, they read the leadership principles, pick their three favorites, and then try to, effectively, slap them into their own company culture without anything approaching context. It’s, “Okay. Frugality. Great. We're going to become super cheap. Two pizza teams, I bought a whole bunch of pizzas for everyone. Why is our code still broken?”

And they don't realize that there's some secret sauce that only really works at Amazon, such as giving things ridiculous names, but that's a bit of a low blow. It's strange seeing how culture is shaped by a company and leadership principles—or their equivalent at other companies—seem like outgrowths of what those companies have become. But in your case, I think that, talking about what your takeaway is, I guess, more nuanced than that. It's not just about, well, we did X thing and Y result happened. Instead, you're talking about things like the writing culture. For folks who haven't been obsessively stalking a $2 trillion company for the past four years—ehem—explain a little bit more about what the writing culture is like, please.

Alfonso: Yeah, so for example, as a solutions architect, you typically are partnered with an account executive, depending on, kind of, what group you're in. I was part of a greenfield team, and there's a lot of writing. There are quarterly, essentially, business reviews; they call them something different internally. But you typically partner with your account exec, and write a six-pager or account territory plan and spend a lot of time and effort getting that right and being really thoughtful about it, and that gets reviewed with the entire area team—including leadership—every quarter. And that is the foundation for how you're going to go talk to your customers, how you're going to get different projects, how you're going to help them succeed.

Obviously, it's a very customer-obsessed company, so it's definitely skewed towards, like, how do we help our customers succeed? What can we offer? What can we help with? What programs do we have? And all that is laid out, every quarter, in a really well-written document.

They're really notorious for bringing red pens to a lot of these business reviews. And the way it works is we get into a room and everyone will read through the person's document with a red pen or some other way to mark up questions or thoughts they have about what's written there—

Corey: D minus you can do better than this. See me after class?

Alfonso: [laugh]. Yep, something like that. But really, it's to help that person think through what they're writing and what their plan is. We all want to see each other succeed, and that person puts a lot of thought into what they wrote, so we want to help them shape that from different people's experiences and diverse points of view. So, it's really awesome to see and something that I really, really enjoyed and will stay with me for the rest of my career, for sure.

Corey: So, what I always find interesting as well, and I'm not necessarily speaking to your specific situation, but the idea of a person leaving a company to go somewhere else, and then boomeranging, or coming back to the company that they left. Now, maybe that's because I can't envision doing that sort of thing myself, mostly because, again, I'm not a very good employee. My resignation letter was at one point carved into my boss's office door, and I'm generally not eligible for rehire. But what inspired you to leave, and what inspired you to return?

Alfonso: Yeah, I think honestly, in hindsight I enjoyed my experience at AWS, but in my mind, I was kind of optimizing for the wrong things. When I reflect back on it, I was optimizing for obviously just the prestige of working at AWS and what that entails. And really, as I reflected on when I was thinking about going back to Red Ventures, I'd rather just optimize for happiness. I think, ultimately, that's what's going to make me successful in my careers is if I do what I enjoy and really just focus on what will make me happy versus other ancillary things like money or prestige, or working at the cool startups in Silicon Valley or things like that. Everybody's going to have a different thing or a different view on what makes them happy and what they want to optimize for, but that's where I ended up.

Corey: There's something to be said as well for watching companies change. I mean, Red Ventures is, to my understanding, not exactly an ancient company. I mean, they've been around for a while, I've been to your office for a DevOpsDays that you folks hosted a couple years back, and, well ‘office’ doesn't really do justice. It's a campus. Especially for those of us who live in San Francisco. It's one of those, “Oh, wow, this would cost only several trillion dollars in real estate to make happen.” It's phenomenally huge and overwhelming, but to my understanding, it's not a company that's been around since the 1800s.

Alfonso: No. And I'll say that is definitely one of the things that also was exciting about coming back to our Ventures is our campus is beautiful, and it's kind of tucked in the suburbs of Charlotte. But yeah, we haven't been around that long. So, Red Ventures was founded in 2000 and we've grown significantly since then. Now we're up to, like, 3000 employees and in 10 cities across the US, and also we have offices in the UK and Brazil.

But our campus, it’s—for me, it'd be hard to see another campus compete with what we have. It's pretty amazing. We have tennis courts, we have pickleball courts, bowling, basketball courts, obviously, we've hosted you for DevOps Day Charlotte in the arena auditorium that we have on the campus as well, so it's really awesome.

Corey: What I find interesting is despite the fact that they've only been around 20 years—my God, was 2000 really 20 years ago—you've become a large company, and obviously, Cloud wasn't really a thing when you were getting started. So, you've migrated into the Cloud—or are migrating into the Cloud. I'm sure you can give much more clarity on that than I can—but it also means that as a big company, it's not quite like a Twitter for Pets style migration, where, okay, I'm going to redeploy my single-tenant application—and even that's never simple—but okay, it took a few months, and then it was done. That's a lot of inertia, a lot of stuff to move, a lot of culture change that has to happen. What is that like? And you were gone for that long, did you see significant progression during your absence between the time you left and the time you returned?

Alfonso: Yeah. Honestly, we've had explosive growth in our usage of Cloud just in the past four and a half years since I first started there. But at a certain point, we took the stance of anything that we're going to develop that's a new project or a new application, we're just going to go straight to the Cloud and build it cloud-native and take advantage of the capabilities in AWS. We're not going to deal with hybrid-cloud. We don't want to spend our energy and focus on that; let's just start building the cloud-native muscles and learn more about the platform.

And then over time, we'll transition to things that are still running in the data center and get those moved to the Cloud, which we're actively in the process of migrating our final apps that are running on-prem. So, it's been really exciting to see the growth, but what comes with that is a lot of waste and a lot of learning. Like, there are things that we didn't know about AWS, and how to optimize costs in the Cloud, and what the features and capabilities are. And even the things that we did learn, those things change.

Now, serverless is a lot more mainstream than when we first started and changes how we approach building and designing applications. We want to take advantage of the ability to scale to zero and save costs, and shut things off when we don't need them, and all those things where we've been learning through over the past four years, and I think we're in a really good spot now.

Corey: People mistakenly believe the Cloud just means it runs on someone else's computer. In practice, it means it runs on money. And the capability story is fantastic, but there's no upper bound the way that there is at a data center where it's, “Well, we're completely out of capacity. We need to break ground on another one,” sort of limits what the upside cost it can potentially be in a short period of time, there is no upper bound with Cloud in the same way, it is not possible via conventional means to fill S3 in a given AWS region.

Found out accidentally that it is not possible, and it reflected that in the bill. But there's definitely a story of discovering then that if everyone has the option to spin something up, that's great. There needs to be something that winds up as almost a following function of, “That's great. That experiment was good to know. We got results out of it. Now what? How do we enforce turning things off when people are done with them?” Because in practice, you spin something up? Great. You're going to retire before that instance does, left to its own devices.

Alfonso: Yeah, for sure. That's something my team has—has been a focus for my group over the past few months is, how do we tackle being more efficient in AWS and reducing our waste? We are a pretty large company with a lot of AWS accounts, and there are things that may not look like they're expensive in a single account, but when you scale out across 50, 100, 200, 300 accounts, that really adds up. It's reflected in the bill. So, for example, some of those smaller things we've been looking at is, do we need four NAT gateways in non-production accounts? Right? Like, obviously, in production—

Corey: No.

Alfonso: [laugh].

Corey: Sorry, did I say that out loud?

Alfonso: Right? So, we have account automation that will create accounts, but as part of that, we want to revisit that and say maybe we can just get by with a single NAT gateway and non-production accounts. We don't need the same level of redundancy there. What about CloudWatch log groups?

Do we have retention policies by default with either our Terraform modules or just the ones that already exist? Are we just wasting money there? Duplicate CloudTrail trails is kind of a sneaky one. If you have an org-wide trail, that's free, but if you create a second trail, you're going to pay for all those events that you're already getting for free. And that adds up when you have hundreds of accounts.

So, it's a lot of these really small things that to a single team that's operating in one or two accounts doesn't look crazy, but when you look at the big picture across the organization, you're spending a lot of money on those things.

Corey: The first, and challenging part, of course, is ensuring you have decent governance, ensuring that there's a relevant way of enforcing good practices, building guardrails in. But then the next question very quickly follows: how do you do that, organization-wide, in a way that doesn't absolutely suck and/or wreck the culture?

Alfonso: Yeah, I think I've definitely battled with that. We've gone through different iterations to figure out what works best, and I think we're still attacking that problem. One of the easy things that we've done is really leveraging what AWS already provides. So, in this case, like Trusted Advisor Cost Optimization Checks. It's a really easy place to start where you can go and get those across however many accounts you have, and just service that information to engineers.

Typically, engineers want to do the right thing and optimize their costs, but they don't know where to start, or they're tied up building features, or providing business value. So, if we can reduce some of that friction there and just go through and aggregate all the findings that AWS is recommending, we don't have to go do that work to get the recommendations ourselves, we can leverage what they've already built, and we just service that to them. Maybe we create JIRA tickets for them; we track those month over month; we show what the total savings opportunity is since AWS already provides that information. So, we're just trying to find where we can plug in and glue things together and reduce some of the friction so we can service the savings opportunities to the right people.

Corey: The challenge, of course, is you want to be responsible, and, “Hey, who copied that extra petabyte of data that no one seems to need somewhere else?” And that's a painful experience that no one really wants to deal with unless they have no choice. But you also want to get out of the way and let people innovate because left to their own devices, first, you're going to have those same guardrails firing off, “Well hang on, you just spin up another instance, that's going to cost $7 a month. You better get a director-level approval to do that.” Which is ludicrous, and it also reinforces some of the bad behaviors that you'll see that make a lot of sense in a physical data center environment that makes zero sense at Cloud.

It shouldn't take you six weeks to provision an instance the way it would a physical server. And the heavier the process gets, the less likely people are to turn things off once they're done using them because it's so painful to get it back if they need to test it out again. I've talked to companies where when they're not running workloads on their test cluster, they'll have it do something like Folding@home, just so it doesn't show as idle, and then accounting yells at them. Which I think is just monstrous as far as having to subvert corporate process to do the right thing goes.

Alfonso: Yeah that's, I think, what we want to avoid. If you have too much process and really make it cumbersome for your engineers to do what they need to do for the projects and businesses they're working on, you're kind of just going to lose some of the trust there and honestly, probably have some more attrition because engineers want to be able to do what they need to do to move the business forward, and if there are roadblocks that don't make sense, or are kind of just arbitrary processes because that's the way we said it was best when we first started using the Cloud, you have to evolve those things. And I'm much more of a fan of the trust but verify model when it comes to both cloud security, but also just cost efficiency in the Cloud, and not necessarily blocking engineers upfront.

Corey: This episode is sponsored by our friends at New Relic. If you’re like most environments, you probably have an incredibly complicated architecture, which means that monitoring it is going to take a dozen different tools. And then we get into the advanced stuff. We all have been there and know that pain, or will learn it shortly, and New Relic wants to change that. They’ve designed everything you need in one platform with pricing that’s simple and straightforward, and that means no more counting hosts. You also can get one user and a hundred gigabytes a month, totally free. To learn more, visit newrelic.com.

Corey: One thing that falls directly under your wheelhouse and of course is near and dear to my heart is managing the AWS costs and managing the cost aspects of the AWS environment. And that's great. The problem that you have with that, start to finish, is that every time you talk to a company about, “How are you managing AWS costs?” They always give the same answer, specifically, “Terribly. We're sort of embarrassed about it, and therefore we're not going to talk about it.”

I mean, I fix AWS bills for a living and I've never yet found any company that will stand up and say, “Oh, we've solved this problem. We're doing it's super well, and it's awesome.” The only folks who do are, “Great. What are you spending every month?” “Twenty whole dollars.” “Great.” That doesn't actually speak too much.

And it gets worse when it becomes a purely engineering problem, specifically because you wind up having engineers who lack context. And I don't mean this as an insult to engineers, but you have this problem where an engineer, left to their own devices will gladly spend eight weeks trying to golf $200 a month off of their AWS developer environment. They cost more than that when they go for their first coffee break. There's no business value to optimizing at that scale. Then of course you have the $2 million a month thing running in AWS. Yeah, maybe spending a couple of weeks knocking a bill in half makes sense.

Alfonso: Yeah, absolutely right. You definitely don't want to optimize for things that, in the grand scheme of things, are not that important, are going to take up too much engineering time for little value in return. There's not a good ROI story there for the business, so you really want to just attack what are the biggest problems? What are the problems we have across all the accounts across the organization? What are the problems we have structurally? How can we improve the culture and have this flywheel effect of just building a culture of cost efficiency? Because I think that is what pays off and has a much better ROI. What are the inputs we can put in initially and just get this flywheel effect?

Corey: Part of the problem, I think, is building that flywheel and not getting lost. The challenge that I've seen about costing, by and large, is that it requires persistent sustained effort, but it's never a lasting priority for most companies. Usually, there's a freak-out whenever the AWS bill shows up that lasts a few days, and then something else takes precedence. But like death and taxes, the AWS bill shows up again at the following month, and restarts that process, and build more urgency. But it takes a lot of getting it wrong before companies start to see the value in getting it right. The painful part for me is that if you start off doing a few things that aren't that much additional work, you save yourself months of effort down the road.

Alfonso: Yeah. And the other thing is once you get to a certain size with your account footprint and your bill, it really, really takes dedicated people that are staying on top of the bill, reviewing the bill, and trying to optimize, managing RIs, which is just a nightmare. Thankfully, they have Compute Savings Plan now, but obviously, it only covers compute. I think my number one AWS wishlist item is for sure a data equivalent of the compute savings plan because that just makes things so much easier. Trying to optimize RI coverage, and then utilization, it’s just a moving target. When you're an organization of a certain size, you have so many engineers—

Corey: Oh, the chargeback or showback model becomes painful then, too, because you have an account that predicts their usage super well and says to buy exactly what they need, but those wind up applying to some other workload who decided to YOLO something into production. How do you attribute that back? Not for blame purposes, but for helping to optimize a footprint or understand a predictive model?

Alfonso: Exactly. And the entire process of how do you even validate that they need an RI for that workload? Is it something they're going to use for a year if you're looking at one-year RIs? There's so much legwork that's manual that has to go in there, it can become a nightmare when you get to a certain size, and staying on top of it's not really a game that you're going to win at with how they're structured today. And that's definitely one of our biggest pain points that we're still trying to get better at.

Corey: The hardest part for me is figuring out how do you fit whatever optimization approach you're taking into the company culture without breaking it. Because you don't want a bunch of engineers to generally have to think about this stuff, for the most part. I've been an advocate for a long time, of the idea that what most engineers need to know about AWS billing, fits on an index card. Sure you need in-house expertise past a certain scale; I'm not denying that at all.

But that doesn't need to percolate out to every engineer working in the environment. There could be a centralized, almost SRE style model, where the cost optimization experts go around on a workload by workload basis and are available in a consult level. That seems to work better than teach every engineer to care about the bill. There's got to be a happy medium somewhere in there, but the hardest part is always culture.

Alfonso: Absolutely. You don't want your business engineers spending their time worrying about RIs or optimizing savings plan, or like, what do we need to buy for next month or some of those other minutiae details. That should be offloaded to a central team, or made easier through things like savings plan. You want them more focused on application optimization, in terms of costs. What are some of the things that we can tweak? Are we building things the right way from the get-go? Are we looking at surplus architectures, things that can scale to zero, in some of those optimizations, and not focused on some of the RIs and other workloads?

Corey: Part of the biggest thing that I think teams need to, I guess, wrap their heads around is it's important to focus on AWS cost control at some point. If it's not a problem for you right now, that's cool. We'll check back on you and a quarter or so because it's never a problem that goes away on its own. But at some point, conversely, it's time to stop focusing on it and stop caring because above some certain level of optimization or focus on attribution, learning to let go and have some wiggle room is always going to be the right answer.

And you were mentioning earlier that serverless is an area of interest for you folks. This is also a different problem, particularly when communicating effectively with accounting. Because the idea behind serverless is that your usage is much more optimized, but it is also way more variable because it's completely beholden to customer workload.

Alfonso: Mm-hm. Yeah, absolutely. Although I still think in the end, you're definitely better off, if your use case fits it, designing almost an event-driven architecture or just working with serverless technologies because you still have the ability to scale to zero. You don't have to think about shutting things off when you're not using them, or when usage is slow or managing infrastructure in your non-production accounts.

And I think we're getting better at that; we're moving more towards that direction, but we built a Terraform module to pause resources and non-prod and sandbox accounts because there was a need there, right? Because we still have a lot of applications that are running on EC2, or even just using things like RDS and Redshift that now you can pause; before you couldn't. So, we built a Terraform module that deploys a Lambda function that pauses those things on a schedule. So, if—obviously, on the weekend, if your team is not working, or overnight, you can pause your infrastructure running in your non-production accounts. And we've also added [00:25:40 unintelligible], like, send those events EventBridge so we can track the usage and the stop and start times. And in the past few months since we've deployed that module, we've saved 290,000 hours, resource hours, which is essentially bending time by saving 33 years in a three-month span.

Corey: There's something incredible to be said about that. It's fantastic to hear stories like that. The challenge, of course, is explaining to accounting that, “Oh, so what is your 18-month prediction look like as far as spend goes?” The honest answer that doesn't lead to nonsense is usually going to look an awful lot like, “Well, it's going to depend as a function of user growth and traffic.” This doesn't actually answer the question, but it does kick the can over to the BI folks rather than engineering in many cases. Again, culture is always the hardest part.

Alfonso: Yeah, I do not [laugh] envy the folks that have to go and figure out the forecasting long term with cloud spend because it's such a hard problem to solve. It's essentially impossible to forecast when you have variable growth and customer usage and you’re growing as a company. It's really hard to track.

Corey: Yeah, we've started doing that as a consulting option on our side, but invariably, it's always a second level of offering after first let's figure out what's in the account, and do an optimization pass, and understand what's going on in a meaningful, intelligent, intuitive way. That's never easy, and figuring out how to get there is always—I guess it's a journey more than it is a destination, which is usually the sort of thing you say about digital transformations when you're discussing that about a cloud bill. That's what drives accountants to drink.

Alfonso: Yep, absolutely. [laugh].

Corey: So, changing gears slightly, you and I first met a few years back. You're an organizer for DevOpsDays Charlotte. And I have a whole story about the time I attended there—in fact I’ve been there a couple of times now. But first, tell me about how you got, more or less, dragged into helping to organize a DevOpsDays, which if you look it up in the dictionary is the poster child for ‘thankless job.’

Alfonso: I don't know about that. It’s definitely rewarding. But yeah, it was a really interesting story how it came about. I remember attending DevOpsDays DC in 2015. I believe it was the first one; it may have been the second one, and Nathan Harvey is one of the organizers there.

And they proposed an open space or breakout topic on organizing DevOpsDays. And I went to that, and Nathan Harvey showed up, and the only other person who showed up was Jason Hand, who was one of the organizers for DevOpsDays Denver at the time. So, I had them all to myself for 30 minutes, so I kind of just asked them what the process was, what do they learn from it? What were some of the pain points?

And I went home from that I was like, man, that sounds really, really cool. It's awesome to get a community of folks together that are all interested in this DevOps thing. And a few months later—that was in June—in November we started the first ever DevOpsDay Charlotte. In a really crappy hotel by the airport. [laugh]. Not knowing what we were doing, but it was the first time and we learned a lot. And this past February, we ended up celebrating the fifth anniversary of DevOpsDay Charlotte and you also were at.

Corey: That still remains one of my favorite talks I've ever given. And it's probably time I told that backstory in public a bit. So, a few years ago, I gave a conference talk on things I've learned interviewing for jobs: how to handle job interviews, how to handle salary negotiation, the usual stuff you'd expect. The problem is, is all of this came from my own experience. So, what I accidentally built was how to handle job interviews as a white guy in tech, which is not the most accessible or inclusive talk, so I stopped giving it.

And what I really liked about the keynote was it let me resurrect a highly, highly modified form of that talk that I co-presented with a friend of mine, Sonia Gupta, who before entering—and later leaving—tech, was an attorney, both as a district attorney as well as a defense attorney. So, she had a lot of thoughts on how negotiation worked, and she made the talk 1000 times better. It probably remains one of the most impactful talks I've ever given because usually, let's not kid ourselves, it's dumb jokes about cloud technology that's not going to stand the test of time. But interviewing skills are the sort of things that absolutely will pay dividends for someone's entire career. And I still get outreach over that talk. People discover the video every month and people ask questions.

Alfonso: That was a fantastic talk. We had a lot of feedback after that talk: the people loved that talk, and it was really, really impactful for them in, kind of, how they think about salary negotiation and job interviewing as a whole. That was an awesome talk.

Corey: Yeah, I hope to give another talk with her again, at some point. Just absolutely a fantastic experience. I found that on some level—and this is going to sound ridiculous, and I don't even care, I'm going for it anyway—I’ve gotten bored with some of the aspects of giving a talk where I get on stage, I know exactly how the entire talk is going to progress start to finish. Having someone else on stage makes it a lot more entertaining for me: you can bounce things off of someone, you're never going to give anything approaching the same talk twice, just because it goes back and forth so well, and it's really neat seeing that interaction. I've learned a lot because since then, I've started this podcast, among other things, and I've gotten better at having the ad-lib recording style of conversation where you don't really know where it's going to end up. I love that style of talk, and I'd love to see more people doing that.

Alfonso: Yeah, I think those are some of my favorite talks. You can tell it's not necessarily scripted, there's a little bit of improv mixed in, especially whenever you're on stage. [laugh]. But those are usually pretty rewarding and enjoyable.

Corey: I don't usually show people behind the scenes, but generally speaking, the way this podcast starts is we talked for five or ten minutes beforehand, just to get a broad outline of, make sure that the bio I have is up to date, and figure out, “Okay, what do you want to make sure we do or don't talk about?” “Cool, let's go.” But I don't have a giant list of topics I'm staring at of going point to point to point. I do it entirely in an improv basis.

That's not because I'm awesome at this, it's because I'm terrible at planning ahead, so you've got to be good at improv when that stuff happens. And more often than not, it leads somewhere pretty uplifting and positive. But it's neat seeing that sort of thing unfold on talks because no two people have the same talk preparation approach, so it really forces you to learn to collaborate with people in a different way.

Alfonso: Yeah, and tying into just different talk formats. One of the other things I really love about DevOpsDays is the open space format, which, you know, you've attended plenty of them and have experience with, but it's just more of an audience-led style of interaction. People walk away from those initially, kind of like, “Eh, I don't know, if I'm going to like that,” or, “I don't know if I want to go to those.” But typically, once they go to a couple, they really enjoy that and it ends up being a favorite part of the conference for them. Because they get to talk to their peers and understand actual war stories, right? Like, things that are happening in companies, and challenges people have solved, and challenges they're still facing and just talking to those things. It's a lot of fun.

Corey: Yeah. I think there's a lot of value to it. And frankly, I've stopped speaking most of the DevOpsDays for the last year or so, primarily because let's face it, I'm trapped at home along with everyone else, and there's only so many times I can talk into a microphone. But also, in part, it's because I think the DevOpsDays are a terrific way to get started talking, and I am thrilled to see voices we have not heard before from folks who don't look like me, who are able to get up and start telling their stories.

My offer that I make periodically, and it stands; I'll renew it now, is if anyone wants to give a talk but doesn't know how to get started or doesn't think they have anything to say, please reach out to me. That's one of the things I'm incredibly passionate about is helping people tell stories. It's not as hard as it looks. There's a lot of things you can do to make it easier on yourself, a whole bunch of shortcuts and cheats as it were. I've done a bunch of tweet threads on this, but those are always ephemeral. It's probably time for another one. And I'm very interested in helping people get on stage and learn to tell the story because I love it. And I like watching other people catch the gift as well.

Alfonso: Yeah, that is awesome. And honestly, the DevOps and DevOpsDays communities, one of the most inclusive and welcoming communities that I've been a part of. It's really awesome to see people share knowledge and are super supportive, especially for first-time speakers. I've noticed a lot of the DevOpsDays conference intentionally asked for first-time speakers because they're going to provide support for them and welcome them. And that's one of the other aspects I really love about being part of DevOpsDays.

Corey: Yeah. I think that's probably a good place to leave it. If people want to learn more about who you are, what you're doing, possibly come and work with you, where can they find you?

Alfonso: Yep, I'm on Twitter at @alfonzo__c. And we are definitely hiring. If any of these problems sounds interesting to you, we're looking for senior platform engineers, and also software engineers across the board. So, you can find us at redventures.com.

Corey: And I have to say, based upon conversations I had when I was out there with you folks, I did not get a big company feel, you know, except for the fact that we were on this enormous, gorgeous campus. But I didn't get a cultural sense of what most people think when they hear enterprise company at a large scale. It's not one of those cultures where nothing ever changes and everything is sort of frozen in amber. I know that that tends to be a common perception, but I don't believe that that's true.

Alfonso: Yeah, no, you're right—

Corey: Not in your case.

Alfonso: —we definitely take pride in acting like a startup. That's one of our core principles.

Corey: I'll take it a step further. You act like a startup without the incredibly toxic, damaging, and trash fire aspects that people often associate with startups.

Alfonso: [laugh]. Fantastic point. Yes, we have some really compelling belief statements that really give me pride about working at Red Ventures. One of those is we believe we can be the change we want to see in the world. And we're a company that believes we have a responsibility to interrupt social injustice and take concrete action to drive change, and you'll see that as well. And happy to chat with anybody that's interested in finding out more.

Corey: Excellent. I will, of course, include links to you in the [00:35:28 show notes] as well. Thank you so much for taking the time to speak with me today. I really appreciate it.

Alfonso: No problem. Thanks, Corey, for having me on.

Corey: Alfonso Cabrera, director of platform engineering at Red Ventures. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review at Apple Podcasts or whatever platform you'd like, whereas if you've hated this podcast, leave a five-star review anyway, but also be sure to include a comment about how my talks aren't nearly as good as I think they are.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

Corey: Alfonso Cabrera, director of platform engineering at Red Ventures. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review at Apple Podcasts or whatever platform you'd like, whereas if you've hated this podcast, leave a five-star review anyway, but also be sure to include a comment about how my talks aren't nearly as good as I think they are.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Richard SeroterRichard Seroter is Director of Outbound Product Management at Google Cloud, with a master’s degree in Engineering from the University of Colorado. He’s also an instructor at Pluralsight, the lead InfoQ.com editor for cloud computing, a frequent public speaker, the author of multiple books on software design and development, and a former 12-time Microsoft MVP for cloud. Richard maintains a regularly updated blog on topics of architecture and solution design and can be found on Twitter as @rseroter.

Links Referenced

  • Connect with Richard on LinkedIn
  • Follow Richard on Twitter
  • Richard’s Personal Blog

TranscriptAnnouncer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: You’ve got an incredibly complex architecture, which means monitoring it takes a dozen different tools. We all know the pain and New Relic wants to change that. They’ve designed everything you need in one platform, with pricing that’s simple and straightforward; no more counting hosts. You can get one user and 100 gigabytes per month, totally free. Check it out at newrelic.com. Observability made simple.

Corey: It’s—at least of this recording—morning on the west coast, which means there’s no better time than to inflict a homework assignment upon you in the form of a 42-page ebook from StackRox. Learn about the dancing flames of EKS cluster security, evade the toxic dumpster of the standard controls, and tame the wild beast of best practices for minimizing the risk around cluster workloads. Become renowned for your feats of daring as you learn the specific requirements for securing and EKS cluster and its associated infrastructure. To learn more, visit snark.cloud/stackrox. That’s snark dot cloud slash stack-R-O-X.

Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Richard Seroter, director of outbound product management at Google Cloud. Richard, welcome to the show.

Richard: Thank you for having me, Corey, this is going to be fun.

Corey: I hope so. So, outbound product management; is that the list of products that Google has deprecated and you're in the process of transitioning into a legacy state and writing the exciting blog post, or is it something more nuanced slash less insulting?

Richard: I wish that was it. Boy, that'd be fun. That would be great if they told me that. No, it's more or less, how do you act as the shim between the day-to-day product work, ship product, plan things, as well as some of the go-to-market, the messaging, the talking to analysts, the customer presentations, and then pulling that feedback back, and then helping overall with strategy for the given portfolio.

Corey: So, it sounds like an almost impossible job because if I take a look at GCP, you don't really have quite as many services as your largest competitor does. And to be fair, most of your products seem like things that—well, how do I put this politely—human beings might use from time to time rather than very weird corner cases, but there's still a lot. More than any one person can reasonably fit in their head. Do you have a subset of those products? Are you sort of the overarching view that winds up then delegating and fanning out? How does that work?

Richard: Yeah, that's a good question. So, there are outbound product management leads like myself for seven or eight areas, things like compute, and storage, and serverless. I covered GKE and Anthos, kind of the hybrid story. So, each one has, again, small teams, senior leaders, people meant to kind of translate the message out and also pull it back in.

Corey: It seems like on some level, it's very hard to find people who are both broad and deep across anything in the world of Cloud. There's just no way around that.

Richard: It is. That's the mysterious 10 years in Kubernetes experience you somehow see in job postings, but that’s—no—

Corey: Well, a good 10X engineer should be able to get at least that near.

Richard: You’d think. Yeah, so it is tough, and the Google hiring process is daunting in its own way. So, hiring people who are g—this is still a product management role. I'm in the product management family at Google. But people who also have great communication skills, people who have been in industry for a while, maybe even worked in the enterprise. So, it's a different skill set, but it is one that is tricky to find. But we've actually picked up a few good folks so far.

Corey: I went through the Google interview process a couple of times early in my career for SRE roles, and the correct answer that Google gave was, “Ha ha ha. No.” Fair enough. I've gotten a little better since then, but in all the wrong ways. Is the product management interview process similar to what you'd see in the engineering side of, “Oh great, you're here to wind up figuring out how to take this to market. Can you invert a binary tree on a whiteboard?”

Richard: Yeah. I mean, now you have to do it with sock puppets versus the whiteboard, but it's still close. But no, on the product management side, we're definitely looking more for how do you think about product? What do you like about it? Can you think about the market size?

So, I might ask you something, Corey, like, “Hey, tell me an application on your phone.” You like, “Terrific.” You tell me what that is. Then I might ask you, “That's great. What would you like to see improved in that?” “Cool. How come they didn't do that?” “Okay, now you run the company. How would you get them to do it?” “Okay, now tell me when you have done it before.” And so we daisy-chain questions, so I think for interview people that we're not always used to that, you're used to being asked about your worst boss, or your favorite job or something, but we actually do tr—

Corey: “If you could be any fruit, what kind would it be?”

Richard: Right. Yeah, “How many oranges fit in the state of Utah?” Questions like that. And, I guess you can get those, but for a lot of time, especially for product management, we're trying to think about how do you decide to get to an endpoint? I don't care about your answer as much as how you got there.

Corey: Yeah. And that tends to be a great way of exposing people's strengths and weaknesses both. I think that there's an overemphasis—at least in engineering—of find the thing that people are worst at, and then beat the crap out of them on that. Maybe I'm weird in that I like to find instead, let's find the thing you're best at because honestly, that's the thing I'm going to have you focusing on in most cases, whereas, “Oh, you're really bad at—I don't know, some sort of deep embedded programming. I guess I'm going to put you on a project where you've got to do a lot of that.” It’s, what are you doing? Why is that important?

Richard: So, are you a ‘spend more time on your strengths than bolstering your weaknesses’ person?

Corey: In my case, I'm a little weird, I tend to have a sort of weird skillset where I can go from knowing nothing about something to, maybe, 50th percentile super quickly. I'm not usually very interested in going from even 50 to 51 at that point. It’s, I know enough to be dangerous and to wield it, and that's great. I’m sort of a generalist in that sense, there are a few areas where I take exception to that, but mostly I tend to focus on those sorts of things.

And then I can find subject matter experts to dive into things and hand it off to folks who are great at it. The nice thing about running my own company is that inherently everyone I hire is better at the thing I'm hiring them for than I am. So, it’s, “Great. Welcome aboard. It's your problem, now. Tell me what I'm doing wrong.” “Oh, I'm sure you're doing fine. Oh, my God, what have you done?” “Yeah, that's sort of what I expected.”

Richard: No, that's a certain confidence that comes with age, I think. As most of us coming out professionally, well you're threatened by someone who's better at your job than you are until you realize that really, you're not promotable until you're replaceable—

Corey: Absolutely.

Richard: —so the whole point should be to find people better than you.

Corey: That's the hard part, I've always found is then trying to replace myself on any of the media stuff that I do because I have a very distinctive personality. I can't very well tag someone else in to write a sarcastic newsletter, for example, and expect them to write it sarcastically and have it come out the same way.

Richard: Yeah, not yet. We're working on that with AI, though. So, we'll try to get the Corey Quinn filter.

Corey: Oh, excellent. Yes, yes, yes. AI is going to solve everything, I'm sure of it. So, at the time that we're recording this episode, you're relatively new to Google. What were you doing before? And how did Google tempt you to come aboard?

Richard: Yeah, three months or so in as we record this, but overall, yeah, relatively new here. I was at VMware prior to this. I was actually acquired through the Pivotal acquisition, so, kind of, at those two companies for about four years. Was at CenturyLink doing cloud stuff, running product management for cloud there for a number of years. Startups, enterprise for five years, which is probably the most important part of my experience because I actually got real empathy for normal humans, not vendors, and then consulting and Microsoft, and things like that for 10 years prior.

So, when Google came calling and said, “Hey, it's an interesting executive kind of position, and we're doing something brand new. We've never done this before, you'll be the first outbound product management person we've ever hired. What do you think?” It’s like, “That seems somewhat terrifying and new, and who knows if I'll succeed? Yes, that sounds like a great job.” So, a few months later, and no regrets.

Corey: I absolutely have to ask this question. I'm sorry. But I wound up going to an AWS Summit once and identifying myself as being with the Last Week in AWS newsletter, which is what most people at the time knew me for. And so whoever is at the front desk, saw Last Week in AWS, shrugged and said, “Cool,” gave me an employee lanyard like I was the director of, “take this job and shove it, but I'm serving out my notice.” When you say that you are the outbound product management director, do people think that you've given your notice, or has that not happened to you yet?

Richard: [laugh]. That has not happened yet. Most people ask me if I'm just marketing in disguise or a number of other things. But so far, they have not looked at as, “Wow, you're actually outbound and leaving the company.” That'll be a good one. I'm looking forward to that.

Corey: I have to assume that you are marketing, but only in disguise insofar as people don't understand marketing. I'd argue that marketing is very much a two-way street: you have conversations with the people using your product, figuring out what they're doing, how they're doing it, and proceeding from there. It's not just a bullhorn, it's also listening keenly.

Richard: Yeah, I mean, besides the enterprise experience, probably the most begrudgingly valuable experience I got was running product marketing stuff at Pivotal and VMware for those years because I've kind of learned everybody is in marketing, whether you admit it or not. You're either selling value—you might not be as explicitly creating an ad but you're trying to demonstrate value so someone gets in your open source project, your product, your whatever. So, I think understanding yeah, market is not actually a bad word if you do it, right. It's about actually trying to translate complex stuff to simple things to drive behavior. That's something we're arguably all trying to do.

Corey: Something that interests me about Google—specifically, Google Cloud—has been, their engineering historically has been awesome. A lot of customers believe—and I endorse this of you—that in some respects, they’re three to five years ahead of AWS. But the marketing has always seemed off pitch, the messaging has been disjointed and strange, and it just hasn't resonated with folks in the same way. In previous years, Google Cloud Next had a whole bunch of people getting on stage and talking about things, but there wasn't a lot of focus on customer stories. There wasn't a lot of getting up and talking about how GCP is transforming what the company is doing.

Now, I'm not going to get into specifics because NDAs are real things, but you invited me to an analyst event with GCP, a couple of weeks before this recording, and one of the things that really struck me was that you had not just a customer telling a story, but then opening it up to Q&A. And yeah, there were warts on the customer story, and it was very clear that it was not a perfect idealized situation, city-on-a-hill style. And that had an authenticity that resonated stupendously. It was incredibly relatable, and even for someone who makes fun of clouds for a living, I found that it's very difficult to mock a customer moving forward and saying things like, “Wow, it's helping us advance our business.” It's a very authentic, very relatable story. I had to triple check, this is a Google event? Did I sign up for something else? So, something is clearly changing, just from that perspective, at GCP. Are you able to tell me what's driving that? Am I just imagining this? Has it been like this for a while, and I'm only now seeing it?

Richard: I don't think you're imagining it. I mean, I think Thomas Kurian coming aboard has driven, not just certain business discipline that we're focused on, is making sure that, look, as much as we're an amazing research and engineering organization, we also want to make sure that we're expressing that in the form of purchasable products, and customer value, and things like that. So, I think that's been a great experience so far, but it is this newer push of saying, look, this isn't just about the technology. So, my dumb mental heuristic to me is Google creates the technology, Microsoft creates markets, Amazon creates products. How do you look at that and say, Microsoft always does a great job kind of figuring out how to create markets out of things.

And Amazon productizes anything anybody else creates, and some of their own stuff, too, which is great. Google sometimes is the source of this, but sometimes you also have to connect this back to the actual value of that tech for the customer, not just putting it out into the ecosystem and saying, “Celebrate everyone,” but back to who got benefit? What was the point of that? And so there was more of a push internally to say, “Hey, it's not just about bursting technology into the universe and celebrating that, but actually making sure people are getting real use from it, there's benefits there.” And that is a bit of a shift versus just build the tech and they will come. No, let’s actually partner with folks that make sure they're getting value and then telling their story.

Corey: It seems like Google takes a very different approach to product management than a lot of other companies. I mean, my running comment has been that your peer at AWS is a post-it note that says, “Hey.” It’s, they're going to build anything, and they clearly have a model down to a science that works. Whereas one of my concerns about Google has been where it feels like GCP came from. And I understand that you're new, but none of this is something that is A) your fault—if I'm right—or it's not something that is even necessarily true.

But my perception—and I'm thrilled to be challenged on this is that Google saw that they were terrific at running large-scale world-spanning infrastructure and that this was something other companies could benefit from. And it feels like they ran into trouble in the early days was they forgot the part that the reason they're so good at this is because everyone writing Google software at Google is a Google employee, and you can enforce certain coding standards, how workloads work in a variety of different ways without having to cajole people. It's one of those, “Great if you want other people to support this, it's going to be in one of these three languages, it's going to have the following supportability characteristics, or it's all your own problem.” Customers don't work that way. And I'm wondering how much of that is accurate, misperception, an origin myth that people filled in from details? Do you have a take on that?

Richard: How do you contrast that with another cloud? So, I think your origin story from the beginning is, look, a lot of the technology that's come out in the form of Google Cloud is things that originated as part of what it takes to run, you know, nine services with a billion users, so there is a certain scale and breadth from underlying messaging systems to storage, to networking, load balancing, that absolutely have now matter manifested themselves as do-it-yourself, kind of self-service cloud services. But do you see an extra opinionation, or an extra kind of YOLO just, hey, figure out our cloud with Google Cloud versus the other ones?

Corey: Historically, in the early days, I would have said, yes. Now, I would say that again—this is possibly an overly cynical take, but my impression of Kubernetes and its rise to prominence really came out of Google, on some level, trying to make the way that the rest of the world built and deployed software a lot more Google-y. Like, this is sort of the ultimate expression of ‘ship your culture.’

Richard: [laugh]. Yeah, I mean, look, Kubernetes, TensorFlow, a number of things that we can look at—heck, Angular, there's a lot of opinionated and foundational systems that have come out of Google open source, frankly, most of the ones that seem to be getting copied, whether it's a Spanner sort of database model, or BigQuery, and other things. So, there are a lot of, sort of, opinions that come out of Google that either go into customers or go to other vendors who also productize them. I'm not sure if I've seen that change over time, or maybe I've observed what you've observed. But I think the focus now is we take it back to kind of today's world, I think we are thinking much more versus just here's a service… kind of, good luck with this, but we think a lot about usability. Frankly, as someone who was also an MVP for Azure for a decade, as well as uses Amazon, and all these other things, I do think our portal experience is really good. I think that our usability of our CLI is really good.

Corey: I'm not going to argue with that. I did a Twitter thread a while back, hoping to dunk on you because I was in a foul mood of, “I’m going to spin up a Linux VM and show you how crappy this is.” And it was distressingly good. There was remarkably little for me to pick on as I went through that process. And it was almost in spite of myself. It was, “Huh. Yeah, I guess this is actually good.”

And I was, again, I mostly exaggerate when I say that this is a bad thing. My ultimate goal—everyone accuses me of being a shill from time to time, although who I'm being a shill for constantly changes, so that tells me something interesting. But my goal is I want to see customers have a choice. I want infrastructure to be better now than it was 10 years ago—and it is—but I want it to be better 10 years from now.

And I don't want it to be, for example, an Amazonian monoculture where you're either running on AWS or you're doing something dumb. I don't want it to be the de facto answer. I want picking a cloud provider to have to be a serious challenge that I put thought and work into. And I guess I'm concerned, on some level, that as I look across the ecosystem of what companies are doing, whenever I see people that are, I guess from my point of view, messing it up, I see that future in jeopardy.

Richard: Interesting.

Corey: That is my bias. I want everything to be better than it is. I don't want to wind up seeing a monoculture. And again, I can make fun of absolutely anything, so there's a lot of good stuff coming out of all of the major cloud providers—

Richard: Absolutely.

Corey: —and rarely, IBM, but there's a definite progression here, where I'm starting to see things work differently than they used to. There's a more serious tone being taken by a lot of the providers. And I'm not sure where it's going.

Richard: Yeah. I mean, look, as you say, we want people making these choices. And there are serious criteria going in. I mean, again, I'm fascinated as to your take, as well as I talk to a number of customers every week, and I'm part of watching their cloud journey and making cloud choices, and we can talk about multi-cloud and those sort of considerations, but I wonder sometimes even is there a ton of crazy thoughtfulness going into it, or is it more organic where, this is we're going to start putting a workload? And then you just kind of back up and go, “Okay. Now we have to stop and actually make thoughtful choices about Cloud.” Do you get an impression, though, that it's much more upfront thoughtfulness from a customer perspective?

Corey: I think there has to be. I’m seeing a few signs from GCP that I'm not asking you to confirm or deny, but it's definitely striking me as the right tone. I know, it's easy to say that, oh, Google, turning things off is only a concern for customer products; real enterprise buyers don't actually have that particular perspective. But I've been in the room when customers are having these discussions, and it has been brought up every single time. But Google going out and saying, “Yeah, that's not a thing anymore and no one believes that,” is patently untrue and transparent.

What I'm seeing instead, as an example of that, is it was announced that with—I want to say Deutsche Bank—there was a ten-year deal with GCP that had been inked. And it was not touted by GCP as, “See?” Which is fine. I think it's great. It speaks for itself. When a bank is saying that we're signing a ten-year contract, you're not going to be able to turn that off anytime soon, regardless of the rest of the business concerns because you would get so horribly dragged in court over something like that, that it would be unthinkable. So, stories like that give me hope and optimism that Google really is in this for the long haul. You can understand why people might question that.

Richard: Yeah. I think that's fair, but at the same time, as you mentioned, I think you’re seeing positive signs. Look, I mean, part of the things that outbound product management teams need to do is both craft and share multi-year roadmaps, which in the Cloud, that's not always something that's pushed forward. It's almost quarter-to-quarter agile planning based on feedback. But if you're going to talk to an enterprise buyer, I can't say, “Hey, this is what we're thinking about doing next month after that—” shoulder shrug. Like that doesn't work in large enterprises.

Corey: They want roadmaps. They want plans on these things. One thing that I will give AWS credit for, even to its own detriment at times, I would argue on some level, they don't deprecate enough, which means that when they launch something terrible, with bad pricing and a dumb name, we're stuck with it. It's not going to go away anytime soon. Which is I know, it sounds like a weird thing to complain about, but it's strange.

I will say credit where due, GCPs deprecations have been relatively minimal and understandable as far as service goes, there was a whole Steve Yegge post about question around the API's and the rest, and how that is communicated. I haven't used it deeply enough to have an informed opinion on that yet.

Richard: Yeah. And to be honest with you, that feedback doesn't go unnoticed. I send out an internal newsletter every Friday, just kind of, “This week in the product” and—

Corey: Oh, I have some thoughts about a weekly newsletter.

Richard: I know you do. So, I referenced that post; we talked about it. So, I mean, we are a listening organization. I think maybe potentially this sort of idea of ivory tower Google stuff is really not the case anymore. It's a definitely more—it’s a smart bunch.

It's one of the—probably the preeminent research-oriented technology companies on the planet. But at the same time, a lot of people recognize what they don't know and that their impressions of what is perfect isn't what the customer thinks. And so I just got done—before we jumped on this podcast—another product review, and how many times in that call it was, we got to meet the customer, this is what they're—we think they need. We've talked to these ten customers before we crafted what we think we want to do. I think that's a new Google, I think that's pretty good.

Corey: This episode is sponsored in part by our friends at Linode. You might be familiar with Linode; I mean, they've been around for almost 20 years. They offer Cloud in a way that makes sense rather than a way that is actively ridiculous by trying to throw everything at a wall and see what sticks. Their pricing winds up being a lot more transparent—not to mention lower—their performance kicks the crap out of most other things in this space, and—my personal favorite—whenever you call them for support, you'll get a human who's empowered to fix whatever it is that's giving you trouble. Visit linode.com/screaminginthecloud to learn more, and get $100 in credit to kick the tires. That's linode.com/screaminginthecloud.

Corey: There's a strong value story here around what this is becoming. It's clear that this is undergoing some maturity. The olden days of software development at Google—and most other places—was, to be direct, typical of many Bay Area companies where if you're not running something that is plugged into the server itself, and you're on the latest model MacBook Pro on the absolute bleeding edge version of Chrome, then you should maybe upgrade your computer. It's much more enterprise-oriented in some respects, and that is just a fascinating shift to watch. I can only imagine the cultural debate happening internally.

Richard: Yeah some of that’s, obviously, positively received. People like seeing the stories you talked about, though, where you saw, during our analyst event where it's a German company that creates air compressors. That is not Snapchat; that's not these Bay Area companies that tell these cool stories all the time that are awesome. It's not just LinkedIn on stage, and Pinterest, and whomever else, and Etsy. This is talking about, okay, let's talk about companies that are doing some things that kind of run the world quietly, behind the scenes. And so I think our people like hearing those.

Corey: Well, I wouldn't call air compressors quiet, but I hear what you're saying there.

Richard: [laugh]. Very true.

Corey: It’s, yeah, these are changes to their IT posture take years, at a minimum, to wind up percolating out, so this is not a short term bet on their part.

Richard: No, no. And it's not short term bets, and there's retraining involved, there's big bets they're making, and to your point, they're locking in for a long period of time because that's what's going to make sense for them. They're not going to switch clouds every week; they're not going to try new services every week. There's a long tail, I'm sure, in all the clouds of which services people actually pay for and which ones they talk about. I think, for the most part, the ones most people still pay for are—what—VMs and storage?

Corey: Oh my god, yes. The majority of AWS global spend is EC2 regardless of all the stories they talk about. It gets five minutes in their keynotes.

Richard: Absolutely.

Corey: Running VMs effectively and well is not exciting, but it's what runs the world.

Richard: No, it's the bread and butter. And so if you can show you can solve that—I think increasingly, we're also seeing more focus on the migration story. It's not just net-new workloads because that's a small fraction of the budget. It’s, how am I getting new value out of the existing stuff? How am I migrating that lower in cost, maybe, or at least getting new agility stuff? The usual cloud value prop.

Then there's definitely that push for, again, I don't just want an arm's length vendor, I am kind of looking for a partner, which is the other move we've been making as well, which is probably new for Google, helping people actually get better at software. So, I like that we're not just a post-deploy platform. We're not just sitting here waiting for the workloads. Can we actually go help people get the right things there and create the right things?

Corey: So, one area I want to, I guess, have a debate with you on is this idea of multi-cloud.

Richard: Yes.

Corey: And I have made loud opinionated perspectives on this, and you have a spouse at times, the exact opposite perspective, I think. Let's figure that out. Where do you stand on it?

Richard: Yes, good. So, first, let's step back. So, what are we defining as multi-cloud? I would argue this is not franken-apps made up of a database in one cloud, a network tier in another, and in VMs in another. I don't think anyone really proposes seriously, that's a good idea. So, let's ignore the franken-app.

Yeah, I think we talk about, okay, is multi-cloud putting the same app in different clouds, for DR purposes, for even load balancing, whatever? Kind of. I think, for the most part, most people think of multi-cloud as an organization is using different clouds, and they're putting different apps in different clouds. So, my approach to this, and reason that I've kind of pushed on this as well is, I think that's 100 percent reality if you're in a decent-sized company. And let's even ignore the public clouds only, although every survey I've seen, every customer I talked to, is not saying literally all our workloads are going to go to the one cloud. I've just—I don't see it. But let’s, even, ignore that.

I think we all know you're going to use more tech than one IaaS cloud. Whether I'm using Twilio, or Salesforce, or Workday, or [00:26:24 Versel], or SnapLogic. Like, there's other clouds, if you will, that will make up my portfolio. So to me, multi-cloud is, I'm not going to single-source, my compute, my data, all those sorts of things. Multi-cloud is about probably using different sort of services for different reasons, but still coupling things close together for performance and manageability. I don't think one ops team, again, is striping things across every cloud, ideally. It's about fit for purpose, and I don't see that stopping. So, instead of arguing that it's a bad idea, let's try to do it the right way.

Corey: One of the biggest concerns and considerations that I see around multi-cloud, it always comes down to definition of terms, where I agree that multi-cloud is a reality and not going anywhere is the idea of using G-Suite for my email system—which now counts as part of GCP, so you can't even argue that one—maybe AWS, my infrastructure, GitHub for my Git stuff. And having different workloads in different spaces. We can take that a step further and call out to specific GCP calls for a machine learning or artificial intelligence API that you've implemented super well, where I need to get data returned from that. And that's fine. I'm not arguing against anything like that.

Where I tend to say this is dumb, is where people are building a greenfield service today that they want to be able to press a button and deploy it, to GCP, to AWS, to DigitalOcean, to Oracle Cloud, to Azure, to my basement Raspberry Pi cluster, et cetera, et cetera. And it feels like that is the wrong direction to go in no matter how you slice it. Am I wrong?

Richard: I think you are right on the whole. Though I think there's something philosophically that separates, sometimes, people think about multi-cloud. Are you thinking about the Cloud is going to kind of adapt to you? Are you going to adapt to the Cloud? And based on how you answer that question, if you say, “I’m going to adapt to the Cloud,” then I'm going to use Dynamo, I'm going to use Spanner, I'm going to use whatever I can, I'm going to use Lambda, the messaging services, the networking tools, IAM, all those fun things.

And that is going to be how I build my app. If it's, “The Cloud is going to adapt to me,” which is I think, frankly, most of the enterprise play, it's, “No, I'm actually going to bring my Redis installation. I'm going to bring Confluent because I like Kafka more. I'm using New Relic for monitoring. I don't want to use what you're using.” And all of my logic is, frankly, in my code.

It's not going to be in your Cloud, it's in my application code. So, I can build—and look, at Pivotal, we had great demos on stage of Dick's Sporting Goods doing an actual live cutover between clouds because they were running Cloud Foundry. Again, I don't think that's the common use case, but I think if you're thinking of, “The Cloud is going to work for us,” versus, “I'm kind of going to work for the Cloud,” your answers are different. And if you think that I'm going to leverage the Cloud versus be wrapped up by it, I'm probably bringing my code, plus some of my other software, not just first-party cloud-native services. And if I do that, then that sort of portability story is possible.

Doesn't mean it's a good idea, in every case. I don't think you should be click button, moving workloads around. That's pretend, sort of, vendor fighting against each other, let me get a cheaper bill, sort of silliness. But I think there is an idea that how much of my app is actually comprised of totally cloud-native services, I think there's a whole industry based on, whether it's HashiCorp, or Elastic, or Cockroach, or F5, that are depending on the fact that you will bring best-of-breed things to your cloud of choice, and I think that's okay. So, I think it's going to depend on how deep you go into a cloud, versus how much you see the Cloud as simply a tool in your tool belt.

Corey: One of the previous episodes of this show was a conversation with HashiCorp’s Mitchell Hashimoto, where his argument was that Terraform was not built for multi-cloud workload portability, it was built for workflow portability, which makes a big difference. And I like that take quite a bit. Where I think it gets lost and becomes sort of strange is that, on some level, people start twisting this in a bunch of different ways. And even at the baseline level, where I only use the common primitives that exist between all providers, that breaks down way earlier than people think.

Richard: Very true.

Corey: I had a customer once spend four months of not cheap engineering time, trying to get Terraform to [00:30:27 pier] up a GCP and AWS pair of VPCs. Well, with all those acronyms and C's in there, it's going to be, okay, how do I say this without making that confusing, but that's neither here nor there. And at the end of it, there was still no great answer because the ways that these things behave is radically different; at an implementation level, things start to break down. What I've seen in almost every case for companies that have tried that is that there's always a primary cloud and a secondary cloud use for something else, maybe for backups or for a particular workload, and invariably—even if they don't intend for it to be that way—the company that becomes their primary cloud provider is the one where their data lives because data gravity tends to be the determining factor on these things. Does that align with your understanding?

Richard: Your primary secondary point seems true. That's what I've seen as well. I think the data gravity point is also clearly fair. Shout out to Dave McCrory coining that years ago. I think that—

Corey: Is that where it came from?

Richard: Yeah, yeah. So, but if you think about the consideration for that, so when I think about even Google's interesting place in the industry is I think—and please tell me, I'm making stuff up—Google Cloud’s the only cloud provider where the enterprise buyer, the person who might care about it, might actually care about the business relationship with Google because their business depends on ads, map, Play Store, apps, whatever else. And so they might say, “Look, we're generating tons of ad data, I want to use BigQuery.” There's a gravity there versus who looks at even any of the other clouds for the most part and says, yes, my business depends on my relationship with that cloud provider. I don't know, is that really the case? So, I think there's a gravity that's going to also happen—

Corey: Only the smart ones, if I'm being perfectly honest with you. It’s, at some point, when a cloud provider is running all of your production infrastructure, they're not your vendor anymore, they're your partner. So, watching companies periodically try to beat the crap out of the cloud provider in, frankly, unfortunate ways during contract negotiation. It's, you're not buying a car here, you have to have a conversation and relationship with these people after the negotiation’s done, and insulting people's intelligence is not really the best way to get there.

Richard: No, you're right.

Corey: That's a frustrating point.

Richard: No, that's a bonkers point. I just think that sort of the business of Google, and therefore, Google Cloud has a different relationship with a lot of large businesses and small businesses because of the nature of what Google provides. And Amazon may be the case as well, you're selling things through them, or whatever have you. But there's an interesting data gravity of things generated by Google proper that you can leverage with Google Cloud.

And I say that because look, my data is also going to sit in Salesforce; my data is going to sit in other systems, as you mentioned, so maybe that gravity will continue to pull into a Redshift data warehouse or data lake in S3; totally possible. I'm just seeing, especially as you do more edge stuff, as you do some things at the retail edge, telco stuff, that that data is going to sit in a lot of places, and the compute is going to follow that. So, as you look at platforms, compute services—that's where Anthos is an interesting play—when you move the compute to the data versus necessarily thinking you can keep consolidating all the data to a single gravity point, I’m not sure if that's what the future is going to play out.

Corey: I'm curious to know how you see it with Anthos and, I guess, it’s current strategy. When Anthos was first announced, it was extremely unclear to me a number of different aspects of it. Who is the product for? What actually is it? The price point, I think, started off at a minimum $120,000 a year, so for my ridiculous serverless nonsense, it was clear the target market was not me. And that's okay, I'm not suggesting it should be. But it does become interesting.

Richard: It does. So, it is not, also, a floor wax, dessert topping, kind of, anything you want it to be. There's a specific thing—

Corey: And the breakfast cereal, honestly, needs some improvement.

Richard: Yeah. No, I know, too much sugar. But if you look at what it is, in reality; look, it's an application platform. It’s Kubernetes, it's Knative, it's Istio, it's a number of components, but it's also kind of a service platform which says, “How do I put new and existing services onto a single platform, that yes, can run on different sorts of infrastructure as a single control plane?” To me, this starts to separate something you and I were kind of spitballing around, but sort of the first generation multi-cloud stuff, and second generation.

And so, first-generation multi-cloud; what was this? So, this was VMs, this was RightScale, this was vRealize, this was, if you remember that period, where all these crazy acquisitions were happening, like Clicker, ElasticBox, all these companies, couple hundred million dollar acquisitions because everyone thought that the secret was going to be, “Hey, we're going to help companies build VMs on different clouds.” I don't know if that really landed.

The second-generation stuff I'm seeing is about apps. It's not about the VM. So, you have more—it's more about control planes of logical compute pools, not just, I have compute sitting in AWS and in my vSphere environment in GCP. It's, I've kind of built a general mesh, which I deployed to. It has an operational sort of setup cost.

And then once I have it, yes, I have a consistent infrastructure layer, and connectivity, and security story—which I think is interesting—that hosts both new and existing serverless-y stuff, thanks to Knative and Cloud Run, or maybe a stateful database, or a Kafka instance. So, I think there's something about these more open control plane application platforms that you're seeing that say, don't lowest common denominator things. I get that that's not a fun part of multi-cloud. Instead, can I normalize at least the lowest level of, hey, here's my container runtime; here's some of the network connectivity; now go use an IoT framework. Go ahead and use this great database. Go ahead and use this messaging layer.

Absolutely, you're not locked out of that, if anything, you're simplifying the original stuff so you can actually use the other stuff. So, I think there's these second-generation platforms are more about orchestrating things for the app versus just standalone VMs. And Anthos is trying to do some interesting stuff here. We'll see where it all goes, but I like this push forward to say, again, how do I do things more open and recognize that compute probably won't consolidate, I am going to be doing more edge, I am going to be using different clouds, I don't want on-premises to suck because even someone who says I'm all in on X cloud is probably on a ten-year journey and their choice is either public cloud is awesome, and on-prem is terrible until that happens, or you start to at least create some sort of behavior and cloud-native patterns and tech to get you some sort of speed on both.

Corey: I think that that's probably a fair way to frame that. I think we get lost in nuance. Where I tend to take objection to this is when companies are just starting out on their cloud journey, and the advice that they are given from a number of places, “Oh, you've got to build with multi-cloud in mind.” Which in turn, means that it's forcing them away from choosing higher-level managed service offerings that are differentiated, or they're not talking about the things that they otherwise would be focusing on the things that add value to their business. Instead, they're trying to figure out how to get a relatively toy web app running inside of Kubernetes, which is a pretty heavy lift in many cases. So, it comes down to when people don't know what to do, and they're asking for advice. I don't love the pattern of, “Well, here's some crap advice we've built out for you.” I like to see an outcome of better storytelling. And the multi-cloud stuff in my perspective tends to detract from that.

Richard: You're not wrong. I think it's a later optimization. And frankly, look, a lot of this is polluted by, obviously vendor intent, and this and that, and look, heck, I'm part of it; I work for a vendor. But I've also—why I do like to believe that being the enterprise buyer, being the service provider, being the consultant, being the developer, at least tries to help me think about this the right way that no, you shouldn't be doing multi-cloud on day one, that's probably a bad idea. You should be using a particular cloud getting value from it, and then optimizing over time, especially as more people use it, and figure out what you want to do. So, there's absolutely a progression.

Corey: Yeah. I am talking best practices, too, to be very clear here. People are like, “Ah, but what if I want to potentially sell to retail, then, and I don't want to have to wind up moving my stuff later? So, I should build with cloud agnosticism in mind.” “No. I suggest you go all in. I would also suggest whatever cloud you pick is not AWS, upon which to go all in. GCP: perfectly valid option. Have you talked to those folks?”

And I don't have a preference for what provider to pick. And I assume by default, whenever I talk to a company has done something in a direction toward a provider or toward a multi-cloud strategy, that they've done the analysis and that my best practices thoughts are no longer directly germane because they have context that I don't as general guidance. So, please don't ever interpret what I'm saying in the general sense to apply to specifically calling one company out for doing terrible things.

Richard: Sure. No, that's fair. I mean, the interesting question is that we've been in all-in, can you still put an eye towards what's the smart thing to do? My old colleague, Josh [00:39:03 McEntee]—smart fellow—used to say, “Look, when you're innovat—”

Corey: Oh yes.

Richard: “—when you're innovating and learning, that's when you definitely should be using whatever proprietary crazy thing you can use in a given cloud.” Look, you should be—you know, if I'm in Google Cloud, I should be using Functions and I should be using Spanner, and BigQuery, and Bigtable, and whatever, right? Because you're just trying to prove something. Who cares if you're trying to genericize this to a relational database, or whatever else. Just frickin’ prove your point.

Now, over the lifecycle of that app, you might over time go, “Well, maybe I didn't really need Mongo in this case. I'm going to just go back to MySQL.” Or maybe it shouldn't be a bunch of Lambda functions. Let's actually just make this into a smaller containerized app. The apps have a journey. And so when you're starting and you're experimenting, you're probably going to be more proprietary as it goes on its lifespan; you're purposely looking for more commodity things. I think that's fine, too. I think that's an interesting way to look at these things, that all of these choices probably go on a spectrum and a spectrum over time. I think as most people stop thinking about software as static, we'll be better off.

Corey: Richard, thanks for taking the time to speak with me today. If people want to learn more about what you're up to, and/or buy your company's products, services, or breakfast cereal, where can they find you?

Richard: Yeah, you can always find me on Twitter—along with Corey—at @rseroter and seroter.com is where I try to blog a couple times a month on my idiot explorations of technology that some people like to follow along with.

Corey: Excellent. We will, of course, code links to that in the [00:40:21 show notes]. Thank you so much for taking the time to speak with me today. I really appreciate it. It's nice to get folks from GCP to talk to me from time to time about something other than what they perceive—maybe rightly, if accidentally—as personal attacks. So, thank you for this.

Richard: My pleasure. Appreciate you having me.

Corey: Richard Seroter, outbound product management at Google Cloud. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts, whereas if you hated this podcast, please leave a five-star review on Apple Podcasts and tell me why this podcast should be deprecated.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Peter Cooper
Founder and editor-in-chief of Cooper Press. Programmer, indexer of all the programming links, former O'Reilly conference chair, and general nerd.

Links Referenced

  • Cooper Press
  • JavaScript Weekly
  • Ruby Weekly
  • Article, “How to use AWS SimpleDB from Ruby”
  • Follow Peter on Twitter
  • Connect with Peter on LinkedIn

TranscriptAnnouncer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Gravitational is now Teleport because when way more people have heard of your product than your company, maybe that’s a sign it’s a time to change your branding. Teleport enables engineers to quickly access any computing resource, anywhere on the planet. You know, like VPNs were supposed to do before we all started working from home, and the VPNs melted like glaciers. Teleport provides a unified access plane for developers and security professionals seeking to simplify secure access to servers, applications, and data across all of your environments without the bottleneck and management overhead of traditional VPNs. This feels to me like it’s a lot like the early days of HashiCorp’s Terraform. My gut tells me this is the sort of thing that’s going to transform how people access their cloud services and environments. To learn more, visit goteleport.com.

Corey: This episode is sponsored by a personal favorite: Retool. Retool allows you to build fully functional tools for your business in hours, not days or weeks. No front end frameworks to figure out or access controls to manage, just ship the tools that will move your business forward fast. Okay, let's talk about what this really is. It's Visual Basic for interfaces. Say I needed a tool to, I don't know, assemble a whole bunch of links into a weekly sarcastic newsletter that I send to everyone. I can drag various components onto a canvas: buttons, checkboxes, tables, etc. Then I can wire all of those things up to queries with all kinds of different parameters: post, get, put, delete, et cetera. It all connects to virtually every database natively, or you can do what I did, and build a whole crap ton of Lambda functions, shove them behind some API’s gateway and use that instead. It speaks MySQL, Postgres, Dynamo—not Route 53 in a notable oversight, but nothing's perfect. Any given component then lets me tell it which query to run when I invoke it. Then it lets me wire up all of those disparate APIs into sensible interfaces. And I don't know front end. That's the most important part here: Retool is transformational for those of us who aren't front end types. It unlocks a capability I didn't have until I found this product. I honestly haven't been this enthusiastic about a tool for a long time. Sure they're sponsoring this, but I'm also a customer and a super happy one at that. Learn more and try it for free at retool.com/lastweekinaws. That's retool.com/lastweekinaws, and tell them Corey sent you because they are about to be hearing way more from me.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by someone slightly offbeat from our normal guest list: Peter Cooper, editor in chief of Cooperpress. Peter, welcome to the show.

Peter: How's it going? ‘offbeat’ is probably one of the best things I've been called. So, thanks for that.

Corey: Exactly. We do our best here. So, normally, I talk to folks who are building things out of technology, for lack of a better term. Maybe it's JavaScript, maybe its cloud provider stuff. Regardless, it's terrible because it all involves computers. You have built, I guess for lack of a better term, something of a media empire, a subject near and dear to my heart. What do you do?

Peter: I guess I’ve kind of just become a meta-programmer in a way. And I don't mean that in the technical form of ‘metaprogramming,’ but I develop things to share them with other developers, essentially. So, I run a media company that publishes primarily email newsletters, so things like JavaScript Weekly, and Ruby Weekly. And you just take a technology and you put ‘Weekly’ on the end—other than AWS, obviously because that's your gig—you put ‘Weekly’ on the end, and I tend to come up at some point or another. So, we publish these for software developers. But yeah. It's kind of a meta-process. I've been the developer, I've done all of that stuff, and I continue to do that stuff for my company. But yeah, I've gone very meta.

Corey: You've been doing this longer than I have—I hope—because realistically, I'm looking at your subscriber counts versus mine, like JavaScript Weekly, according to your website, has 140,000 subscribers. At the time I'm recording this, I’ve got roughly 22,000. So, yeah, you definitely have been building a larger audience than I have, but I suspect you've been doing it longer. I got my start on this in mid-2017.

Peter: Yeah, and actually, your numbers are out of date, probably because you looked on the wrong place because we just don't maintain our website whatsoever. So—

Corey: Yes.

Peter: —on the actual javascriptweekly.com website, it will say the actual live number, which I think is something like 172, somewhere in that, kind of, ballpark now. So, yeah.

Corey: And because of production delays, I'm sure that number will be out of date as well.

Peter: Yeah, things aren't growing as much as they used to. I can tell you that much. But yeah. [laugh].

Corey: You hit a certain point where, one, population limits become a concern. And, two, it's always been challenging for me to figure out what to do about growth on the newsletter. When I wind up talking to people about where they've come from when they sign up for this, it’s, “Oh, my friend told me about it.” That's great. I've mentioned this on stage in front of thousands of people during talks I've given and gotten a couple hundred signups as a boost, but it's very hard to get people in large numbers to do these things that suddenly cause massive inflections. I look at the growth over the past three years, and everything's been a pretty steady curve.

Peter: Yeah, it's definitely changed for us over the years. When we began, we could look at it and say, all right, we're growing X percent per quarter or whatever, but after a certain period of time, after a few years, it began growing in a much more linear way. And so, when you're growing linearly, so let's say you're adding 1000 subscribers a month, just to pick an arbitrary number, but you're losing, you're getting the churn of say, I don't know, 1 percent per month, I don't even—I—that's a good number or not.

But let's say you're losing 1 percent per month, well, eventually your list will get to the size where 1 percent per month is equal to that 1000 people that you've added. So you eventually, with any list, if you're growing in a linear way, you eventually reach this plateau and you don't see a lot of growth at that point. And we've actually reached that point on one of our publications. So, it's something that we knew was going to happen, but yeah, as I say we did start ten years ago so it's going to happen at some point or another with certain technologies.

Corey: On some level, that also becomes a non-issue. It feels like, sure, I sell sponsorships, obviously, in the things that I do, but it's not directly tied to number of people reading it. That serves as a baseline for what you wind up fixing it to, but if we were to magically I don't know, multiply the number of readers or listeners that I have by 100, I'm not going to be able to multiply what I'm charging for ads by 100 because there's maybe three companies in this entire space that could conceivably pay that. And I generally spend most of my time making fun of two of them. So, at some point, you wind up with a growth stops mattering to some extent.

Peter: To some extent, yeah. I think, if you had an email newsletter, for example, that was for billionaires only, and you had, say, 50 people on that list, you're going to be getting some serious sponsorship opportunities on that list just by virtue of the fact of the people that are on there. And you can scale that down to other things. So, if you had a list that was all CTOs, you had, like, the top, I don't know, let's say 200 CTOs in the world all on this one list, that would be super valuable, and that’d be making tons more money than either of us are, or would ever dream to do. So, it's not actually about the size of a list, it's about the quality to a certain extent.

So, that's where I see—your list—so let's say I've got JavaScript Weekly with 172,000, or whatever, that's a very—there’s lots of weird disparate groups within that. You can't just say that all people that are working with Node or they're all working with jQuery, or they're all doing enterprise-level development because they're not, whereas at least with your list about the whole AWS thing is that you pretty much know they're probably going to be using AWS or they used AWS in some way or another, and that is a very commercial thing to use, and so they're not afraid of spending money, which obviously, is your whole shtick, you want to spend less money, but then, of course, that means they're spending money, so they're not scared to do that. Whereas my audience, you don't know if they want to spend money or not. So—

Corey: Oh, yeah.

Peter: —there is a difference in value between the groups.

Corey: You look at the typical sponsor for a lot of what I do, they're monitoring, observability companies. These are folks that are trying to sell things to an audience where the long-term value of a given customer is astronomical, so if one lead converts and becomes a customer, then it pays for an entire year's worth of sponsoring my nonsense, and then some. So, the ROI is ridiculously high. But if I were to sit here and look at it through the optics of just viewing it as raw numbers of who listens, or who subscribes, the sponsorship prices are obscene. And it's easy to look at it from a place of who in the world would ever pay this? Well, people who have that problem: people who are looking to get in front of a technically sophisticated audience who generally block ads and are somewhat skeptical. It works. It's the strangest thing. I mean, I thought when I first started that, oh, no one would ever click on an ad in something like this, and go ahead and then buy something from it. I am provably wrong on that.

Peter: Yeah, I mean, we walk a real tightrope with this, actually, in that, one of my things that I had always from the beginning of running the company is that I wasn't actually setting out to sell advertising; I was actually setting out to sell my own courses. So, that's why I started the email newsletter because I was running a Ruby course each month that I would sell tickets to and do, like, a virtual online thing—which now everyone's doing, but ten years ago, this was reasonably new—and I was selling that each month making about out, I don't know, $10,000 or whatever for the course, and then I would just keep promoting it on the newsletter and fill it up. But eventually, I got bored of doing that and I began to accept some of these inquiries for sponsorship and put them on. And I always wanted to make sure that I didn't price it in such a way that it was prohibitive for people that sell things that aren’t, they can sell one unit and that's tens of thousands of dollars of worth to that company. I also want to be able to sell sponsorships to companies that might sell things that are much lower, kind of, dollar value.

And so, one of our first sponsors actually was a site that sold a template for Rails apps. I can't remember exactly what it was called, but it was, like, a Rails template thing, anyway, and they sold tons of copies of this after one sponsorship of our newsletter and they didn't pay a lot. At the time, it was a few hundred dollars or something, but they sold thousands of dollars of this thing. And I thought, I always want to have that experience where just a single one- or two-person band can come to me and say, “We want to run something; we don’t want to spend more than $1,000, $2,000 maybe.” Depends where you going out their risk profile and everything, but we still want to always be able to cater to that type of customer.

And we've actually managed to keep there, and that means we've left a lot of money on the table. And yeah, so that's always a balance I'm trying to find I do think of it as a bit like being on a tightrope. I could fall one way, and just say, let's take the big, let’s say, IBM Cloud money because they've come up a few times in conversations I've had about being very prolific with their spending, let's say. But then, the same time, I don't want to just hand it all over to a company that can just buy all our inventory. I want to have those one-man-band, one-woman-band type things in the mix as well. I just think it makes it nicer for the readers, and they actually enjoy reading that type of stuff.

Corey: I would agree with you. I mean, I came at this from a very different perspective. I had to be talked into starting a newsletter at all. But when I first started my consultancy, it was I had to keep up with everything Amazon announced. That got me 80 percent of the way towards having the stuff I needed.

And it was, “All right. I'll send this out and see if anyone else finds it interesting.” It turns out that that's the thing people tend to know me for the best. And that became sort of the stepping stone that led to this ridiculous thing. I couldn't do it again if my life depended on it.

What I do see is that a lot of newsletters have started not as a labor of love, but as a, “Oh, I'm going to go ahead and make money out of this thing,” and people are coming at it from day one. First, they're probably making better choices about a monetized newsletter than I am. I mean, the way that I've done this, there's no way I could ever have someone else step in and write it for a protracted period of time just because it's so tied to my personality. Whereas other folks, if you're doing just the facts style, well, that's super easy. Just find someone who can opine intelligently about a topic, which is way easier.

Peter: Yeah. I mean, it probably sounds cliché. But for me, publishing has always been in my blood to some extent or another. I began the school newspaper when I was at my early schooling. And I published a fanzine on Usenet in the ’90s, about programming and doing stuff like that, so doing that as a teenager.

And I didn't actually see it as being unusual to build up an audience and have something to say to that audience, so I was very much into blogging and really heavily into that in its earliest days, and tumblelogging, and LiveJournal, and all that type of thing. And, yeah, I kind of would do this anyway, so that's actually what's happened with all the different things I've done. I've been doing the publishing anyway, and then some sort of money-making activity has come along on the back end and appeared in front of me. And it's like, “Oh, actually, I can turn my little fun hobby into something that makes money.”

And that's the thing that actually, kind of, keeps me on the straight and narrow because I guess if that didn't turn up after a few years, I'd be on to the next thing I'd be doing TikTok or something by now. And yeah, that type of thing. If I was a billionaire or something, I'd just be on a different social network each week just playing around with it, see how it works. But yeah, the money kind of keeps me doing the email. Like I might have not still been doing it now if I didn't need to, but it keeps me honest. I like to say that about money is it's not always a bad thing, it's something that can actually keep you on the straight and narrow and going in a certain direction, which I think is good.

Corey: I find that doing it for money definitely helps me power through slumps, where it's, I don't want to really write a newsletter this weekend, I'd like to let it slip but—

Peter: Exactly.

Corey: —on the other side of it, it's well what I'd also like to do? That's right, continue to wind up having food come in. So, I'm going to go ahead and actually power through writing it, and get it out there, and not have to issue a whole bunch of refunds to people. And that, in turn, is great. Most weeks, it's not like that. It's much better from my perspective, to be able to, I guess, have that forcing function that allows me to get it out when I need to. But most of the time, it is a labor of love. I still enjoy it because I get to see how far I can go with these things.

Peter: Exactly. And I guess you're also consulting at the same time, so—

Corey: Oh, very much so.

Peter: —and you’re doing the podcast as well, so you've actually got your finger in several different pies, I guess.

Corey: Well, having staff absolutely helps. That's a lesson I learned from you about a year or two ago when we last spoke in person. Yeah, back when that was a thing people did without it being a deadly risk. And yeah, it turns out that now I have two podcasts going on, the AWS Morning Brief. And, of course, the Screaming in the Cloud that we are currently recording, but the Last Week in AWS newsletter is also going out multiple times a week now.

It's a question of all right, how do I expand this into something even more than it already is? And it's always been a bit of a tension because we are a consulting company first and foremost. We're not a media company, but the media is our marketing, which in turn means that because it is profitable on media, that means that our cost of customer acquisition on the consulting side becomes a negative number. It's marketing we don't pay for; in fact, we get paid for doing it. It's a weird world.

Peter: Yeah. I mean, you're doing DevRel, almost, for what it is that you normally provide in your day-to-day work.

Corey: I would consider it DevRes done right.

Peter: Yeah. I've seen, actually, a few people saying recently that this type of DevRel stuff, this type of content marketing, whatever, it's basically table stakes now. If you're a technology company, this is going to be happening at some point or another. Like, maybe 20 years ago, you would have been doing SEO all the time, and focusing on that, and getting your keywords right and all that type of thing. Now, it’s, you know, this whole new world is what people are producing media and talking a lot, and talking to each other, and sharing stuff and this is, kind of, just table stakes, now.

Corey: Yeah. But what's weird is the number of people I talked to who are convinced that they have nothing to say, they could not start a newsletter themselves and get any traction, maybe you're right, but I would put myself right in that category, too, when I got started. Try it and see. Now, there are things I would have done radically differently for going down this path myself, but now that I know what I know, I would do it very differently, but I don't know that there's so much that I would have changed that it would have resulted in a materially different outcome.

Peter: Yeah, it's a different time now, though I would say, is one thing compared with ten years ago is that ten years ago, I could just turn around and say, “Oh, I'm doing a weekly newsletter about JavaScript.” And I suddenly got a lot of interest, a lot of people on Twitter were—very prolific people in the JavaScript world at that time were tweeting and saying, “Oh, it's great. Someone's finally doing this, blah, blah, blah.” Well, now if someone turns around and says, “Oh, I'm launching a newsletter about something,” like, their friends, and maybe some of their industry contacts will be like, “Oh, that's cool you're doing that,” and sign up.

But I don't think you get quite the same reaction that we did, just because it was novel at the time, and no one was putting their efforts into that direction. So, I would kind of struggle, I think, to do what we're doing now if I was launching it right now, just because there's so much—not necessarily competition for the exact thing that we do—because there's a lot more competition for your attention nowadays. I think everyone's figured out that that is part of the game now. It's about keeping people's attention, and fostering attention, and coming up with ways to get people's attention, and people were a little bit more naive about that type of thing ten years ago, especially in the developer space.

Corey: I absolutely agree with you. Part of the thing that made the whole thing work for me has been that it's easy enough to—for me, at least—to get out there and grab people's attention because I say the quiet part out loud. The fact that I structured my consultancy so I have no vendor relationships with any vendor in the space—including AWS—means that I don't have to worry about censoring what I say out of fear of offending anyone in that sense—in a corporate sense. Obviously, punch up, never down. Offending people is a whole separate argument.

And because I have staff that handles the sales for sponsorships and the editorial pipeline stuff, I don't see who's sponsoring something until after I have already written it so I don't have to worry about it flavoring the content, which means that I get to keep my voice. No one has ever complained, as a sponsor, about why did you say something that wasn't very nice about us when we're sponsoring you? “Because I've been very upfront about this, the entire time,” was always my planned response, but I never had to give it. No one cares. It becomes a very, I guess, straightforward and easy way of maintaining my authenticity, to the point where I could theoretically shut down consulting, decide this whole boondoggle was a mistake, and just do media going forward. What scares me going down that path, though, is how do you avoid losing the technical experience that gives you authenticity and just becoming a talking head?

Peter: Yes, exactly. I've seen a lot of people have this problem, actually, especially in the screencasting space in, particularly, the Ruby world. So, I used to know most of the people in the Ruby world, it's definitely not true now. I've fallen out—not fallen out. That makes it sound bad, but, kind of, I’m just not in that game all the time anymore.

But some of the people I did know that were doing weekly screencasts, and they turned that into their business where it's like, “You're going to pay me $9 a month, and you'll get these screencasts or whatever,” a lot of them just burned out or got to a point where it was like, they were saying, “I’m not doing this on a day-to-day basis for customers and for people that perhaps I don't necessarily like and have to come up with solutions that are imperfect, and I'm not learning these hard lessons anymore because I'm working this idealized, kind of, ‘here's the perfect way to do X, Y, and Z.’” and they weren't getting that real-life experience that they could put into the video. So, once they've done a certain number, they just kind of burned out. Like, “I’ve told you everything that I need to know; I've told you all my wisdom; this is the end.”

And yeah, that is something that could happen, and that's something that I'm really quite aware of in my own work is that I'm constantly researching things and trying things out for what I'm doing, but I don't necessarily have the day-to-day experience of running a giant Postgres cluster in production, let's say, or how to migrate stuff from AWS to Azure or vice versa. That’s stuff that I know of, and that I talk to people about, but it's not something I've actually done for myself so this goes back to my point about being very meta in what I do, is… I’m meta in that sense, as well as that I pick up stories and I learn stuff almost like a reporter essentially, but I've not had that lived experience.

Corey: Yeah. I wouldn't expect that it's something that's a particularly common occurrence, but the fear of that's always been there. But what sort of reassures me on some level is that although I have no plans to stop consulting anytime soon, for obvious reasons, no one is able to use everything in AWS to its fullest potential, full stop. We've launched this past the point where I can talk convincingly about AWS services that don't exist and not get called out on it by AWS employees, because no one has it all in their head. They never do. No one does. So, the idea of explaining these things to people is absolutely valuable to a whole bunch of folks who—that is, I think, the reason that people keep coming back for my nonsense even on weeks where I'm not particularly brilliant or scintillating.

Peter: Yeah, I think there's a lot of parallels, actually, with this type of work, with things like anthropology and archaeology. And I know it sounds a little bit sort of highfalutin, as it were, but there is a lot involved in analyzing a space and being aware of the different things that are happening with it, but also the history that led up to where things are now. So, one area that perhaps you have a positive point on all of what we've just said is that you will have seen the growth of some of these services and why certain services just kind of fell by the wayside. So, like SimpleDB, for example, you got that story in your head, which allows you to make those jokes about a service like SimpleDB, that people coming into AWS now, well they go and look at the list of services, they see SimpleDB and think, oh, this looks cool. It's got a simple API.

And I've done this as well because I wrote an article about three months ago, “Using SimpleDB with Ruby,” even though you probably shouldn't be using SimpleDB. It kind of works, and it does its job, and I can see some use cases for it. But you have that story, and I guess that's also what I have is when I'm looking back at JavaScript, I was posting on the JavaScript Usenet group in 1996, like back when it was beginning, and so even though I've not been programming in JavaScript every single week since then—I kind of do it on and off—I still got that story, and I know where JavaScript was at that point, and how things have built up, and I can tell that story in an authentic way. And I must admit, a lot of software developers don't necessarily do that. I really respect what full-time software developers do, but they don't always have that background story, unless they're, like, 50, 60 years old, and they've really just been doing it the whole time.

Corey: You’ve got an incredibly complex architecture, which means monitoring it takes a dozen different tools. We all know the pain and New Relic wants to change that. They’ve designed everything you need in one platform, with pricing that’s simple and straightforward; no more counting hosts. You can get one user and 100 gigabytes per month, totally free. Check it out at newrelic.com. Observability made simple.

Corey: Right. And I guess that's part of the authenticity, is people who have experience and know the space, who can speak authoritatively about some areas, but then it turns into—like, as technology evolves, as it always does, those stories start to lose some element of relevance. And it'll happen to me someday, the same way it happens to other folks. Right now I'm able to keep up, but at some point, if JavaScript continues to eat the world, well, I've never been great with it, and I don't anticipate that changing anytime soon. At some point, the industry is going to pivot in a way that I can't follow directly. And then a question of “what do I do?” becomes an open question. That said, we don't see a whole lot of contractions in our space.

Peter: No.

Corey: It is continually expanding, and there's always going to be niches, finding those edge cases between other things are really where I've always been able to make success happen, like blending tech and finance on the consulting side, and blending my ongoing love affair with the sound of my own voice and making fun of things means that this podcast and the newsletter tend to work out. But it's definitely an experience. When you started, were you planning on doing this as a business, or was this something that you just did on a lark to see what would happen?

Peter: I meandered through pretty much from the very start. So, I mean, if you just rewind all the way back to when I was 16, I actually finished school when I was 16 because that's how it worked here at the time: you went to, kind of like, an intermediate college at 16 until you were 18, and if you wanted to go to university and whatever, but I didn't want to do that. I wanted to get out and work straight away. So, I left at 16 and got a job with what was then called a new media company. So, they were building websites and so on, back in the ’90s and I just went from there.

So, I have no serious educational qualifications to be fair. And yeah, everything I've done since then has been almost like a side project that's turned into something that's successful. So, I went into doing blogging, and I went into doing freelance writing about stuff and all this type of thing. But I basically just done something for fun, and then someone's like, “Oh, actually we want this. Would you like to do it?”

And this happened when I was blogging. I had Apress, the publishing company, came to me and said, “Oh, do you want to write a Ruby book?” And I wrote a Ruby book. So, I launched a blog to promote the Ruby book, and then people wanted to sponsor the blog because it did really well, and then the blog turned into an email newsletter, and then people wanting to sponsor the email newsletter. So, then that's turned into this company. And—

Corey: Yeah. “Can we give you money to talk about it?”

Peter: Yes.

Corey: “Can you give me money? Of course, you can give me money. How much money are we talking about?” And that's how it started.

Peter: Yeah, it's just like messing around, really. And that comes from a place of privilege, to be fair. And I've had the opportunities to do some of this stuff, and I've had the time on my hands to just play around with things and people approach me in that way. But yeah, at the same time, there's not a lot of design to it.

Corey: Yeah. And at some point, I keep iterating forward, and learning things that I would do differently, and talking to other folks. What always surprises me is when I talk to other DevRel types—if you'll forgive the term—in the space, there's always a bit of a hand-waving dismissiveness about what I've been doing about how, “Oh, yeah, your results aren't typical, and it won't work for other people.” Well, why the hell not? I have magic on this side. I just started off and stuck with it long enough that it turned into something. But I was running this thing at a loss for six months.

Peter: Yeah, I would actually love for you to speak to him, and I don't know if you've done this yet, not in any of the episodes that I've listened to, at least, but actually speak to what you might call a proper analyst, like one of these very, very high paid people that works at, like, Gartner or whatever. And I don't really understand their market, but they seem to have, kind of, taken some of the elements of the types of things that we do—which is building stories up about technologies, and how technology is joined together, and which one you should use and which ones you shouldn’t—and they've turned it into a billion-dollar business. Now, I look at them perhaps the way that some developers might look at us about like, “Oh, what do they know?” And, you know, how do they make this into a business? Well, I look at what I call profit analysts in that way. I don't understand how their business works whatsoever, and I've asked some, and they’ll, “Oh yeah, like, big companies give us loads of money to tell us—we can tell them what we think.” And I'm like, “Well, that sounds like a really cool job. I want in on that.” So, [laugh] yeah, I'd love to know more about that stuff.

Corey: Yeah, I think that there's a lot of secret sauce that goes into that, and part of the challenge at the big analyst firms have is, on some level, what they do is great. It provides validation around directions big companies are going to go, where mistakes to back out of cost billions. The other side of that, though, is I've never yet met an analyst firm that approaches things the way that I do, which is, “Okay, I listen to the vendors who are building things. Cool. I talk to the customers who are using it. Great.”

And that's what analyst firms all do, but then I take it one step further and I build something with it myself, once it hits a certain point of interest for me, and everyone sort of stares at me like I'm a lunatic whenever I say that. But yeah, someone actually said to me once at an analyst event was, “If you can write code, why don't you go do that, instead of this whole media analyst thing? It pays better.” First, are you sure about that? Secondly, how can you effectively work in advising people technology, if you don't know the reality of how it works? I've never fully grasped that.

Peter: You know what? You've just hit on a really big point there that, actually, I think affects so many things that happen in the broader developer space, but, like, where you're not just a software developer, but all the things outside of that—like being a DevRel, and being in marketing, and analysts, and so on—which is that being a software developer right now does tend to pay, or more obviously has a bigger reward, if you're willing to put in the effort, than doing all of these ancillary things, and that actually takes away a lot of opportunity for those other areas to get some really top talent. So, in all the areas that I'm in, it's actually really hard for me to find curators, or anyone that can, kind of like, replace me within the business because anyone who is, let's say, a particularly good JavaScript developer, they're going to be going to Amazon or wherever it is, and earning $150,000 a year plus being a JavaScript developer. Or they're going to go and create a startup of their own because they want that autonomy. They've got all of these different options; they're not going to go and work for a publishing company that maybe had to pay them towards that amount of money, but not quite get there, if you see what I mean.

And that seems to be an issue in so many areas of the industry. Like, why it's not always the best developers, or the best people writing the books, or publishing the books, or writing for the magazines, or all that type of thing, or even producing the videos. I know so many great YouTube-based developers who produce really good videos, but there's probably other developers that are earning two, three hundred thousand dollars a year that could probably do it a little bit better, and probably have more war stories, and types of things like that. So, yeah, I think that's really touched on a point that a lot of companies run into is that there is a lot of talent out there, but it does kind of gravitate towards the purely software development roles, or the management roles, or the FANG roles, as you might call them.

Corey: Oh, yeah. Part of the trick, too, that I think people lose sight of is that because it would take them forever to sit and write the newsletter. Because I know; I've been there. When I started this, it took me a full day's worth of work to sit there, and read everything, and understand it, and copy and paste into Google Docs, and write the snark and the rest. Now it takes me 20 to 30 minutes every week because I built a bunch of automated systems around this.

And sure, what I built is horrible, but it both gives me exposure to the technologies I'm talking about, and makes it harder for me to screw things up, like forgetting to put the sponsor link in, as I did a few times in the early days, or not having a linter so the actual link didn't work, and validating that the actual destination was sending correct responses. And there was a lot of painful stuff, but now it's really getting there to a point where I'm pretty satisfied with how the automation works. Now, the next trick is, of course, getting it to be something that someone else could manipulate without me.

Peter: Oh, absolutely. And I guess you've also touched on the idea of programming as literacy, which I think is another important thing. It's probably a little bit off-topic for this conversation, but I think is very, very important, in that if everyone learns to program to some extent to improve their jobs, then you're going to see massive change in the world. And we're doing on a much smaller scale. We’re making tools because we know how to build software that actually increase the efficiency of our businesses, but just imagine if almost anyone could do that, imagine if the person down at the local hardware store could—oh, they've not got an app to track certain things are in their business, they can just put something together, and they can do that, then yeah, I think the world would be a very different place.

Corey: I think we're going to get there. I think that is the inevitable direction that we're heading in.

Peter: Yeah, the whole no-code thing that's going on at the moment. That’s—

Corey: Yeah, if I want to build something today, and I have a business idea, and I'm fresh out of school, or coming in from working at a hardware store, for example. Great. Today, I have to go to a boot camp and learn how a bunch of this stuff works. What if I didn't? What if it was, I basically put my idea together in some relatively accessible way, and that's enough to get off the ground and get started and start serving customers?

Sure, you're right. I'm never going to be able to scale that thing to a hyperscale, works at massive web-scale properties, but sometime between my ridiculous idea and becoming a Fortune 500 component, then there's going to be some evolution in there. And most things never need the scale like that.

Peter: No, absolutely. Totally agree.

Corey: So, I can't shake the feeling that there's a lot of opportunity that is being missed right now.

Peter: Yeah. And I think once we figure out a way forward as a society, culture, or even just as an industry to do that, we're going to see some massive productivity gains. But it's always hard to see; you might be able to see this point somewhere off in the future, it's really hard to figure out how can we reach that point. Totally obviously once you do that type of thing, you obviously end up reaching a different point than you expected to reach. That just seems to be how our industry goes. [laugh].

Corey: Yeah. That's the hardest part is getting started. When I started the newsletter, for example, I didn't know any of this stuff worked; I was mostly making it up as I went along. And, all right, I'm reading a lot of AWS releases.

A lot of these are just nonsense. No one actually cares that a service that no one ever heard of is in a region that you're not sure it exists or not. So, how do you skim out the stuff that's actually worth talking about, and ignore the rest? And again, that's opinionated and biased, and it's my own position and no one else's, and that's okay.

But there's an element of, just start writing and get started. You learn from your audience, you get less feedback on any of these things than I would have expected. So, you can still to this day, hit reply to my newsletter, and it winds up in my inbox, but almost no one does, which is why that works.

Peter: Yeah, that's true for us as well, actually. You can do that on any of ours. And I'm sending almost 500,000 emails a week, and the amount of replies we get is actually quite minimal. I do get a fair few that I have to work through sort of each weekend because I haven’t got the time during the week. But it's mostly people submitting stuff and saying, oh check this out, blah, blah, blah, and actually come in they're very valuable, and I always reply to people, and let people know what we're doing and everything. But yeah, if you think like, 500,000 people, it's not like I'm getting 1000 replies each week. I wouldn't be able to cope with that. [laugh].

Corey: Yeah, that becomes a problem.

Peter: Yeah.

Corey: So, as far as what you would do differently, if you were starting over, what tips do you have for someone who would be starting out today. You have an idea to start a newsletter. Where should they begin?

Peter: I think there's some different ways that I would go about it now than the way I went about it. So, the way I went about it was very matter of fact, factual. And I guess that's what you did to an extent in the early days as well, but I think something that you've done that we've not managed to do is you've managed to put some character into it over time. So, you've got your mascot, you've got the way that you speak to people, and the way that you make jokes, and you have this level of irreverence. I think that is actually much more important if you're launching something now than if you were doing it ten years ago, where I could get away with that dry, kind of like, “Here's the news. Bye.” Type thing. Nowadays, you do need to put a little bit of yourself in, and it doesn't have to necessarily be a character that everyone likes. It actually helps you out if there's a certain portion of an audience that is going to be like—

Corey: Oh, I get hate mail from time to time—

Peter: Yeah, exactly.

Corey: —and I’m perfectly okay with that. Great, it's not for everyone and that's fine.

Peter: Yeah, I think that's actually beneficial nowadays because generally if you have people that really dislike your shtick, let's say, then you're generally going to have some hardcore fans as well. And building up that level of hardcore fan is actually really important now compared to how it was ten years ago where I could just build a generic audience up. And that's where I've actually got some problems is that I've got such a weird, wide range of people—which is good because it's kind of diverse, in a way—but they don't all necessarily share my personal sense of humor or sensibility, or even accept some of my views about things. So, for example, when we did some issues where we mentioned the Black Lives Matter thing, for example, and we went into depth about why that was important and stuff like that, we got some really nasty emails coming back saying, “Oh, this is a political thing, blah, blah, blah.” And, okay, people can have their opinions about things.

But I would have appreciated the fact that if you've been subscribed to something that I've been writing for several years, that you could at least accept the fact that I might have an opinion that you disagree with, and that's kind of, okay. I can deal with my readers having opinions that I don't like as long as they're not ramming them down my throat. And it's a shame that we couldn't get to that point. But the thing about being really upfront with your character and saying, “Look, this is who I am, this is how I'm going to talk. This is what I believe in. Off we go. Let's publish something.”

It’s that the audience kind of self-selects to a certain extent. You're not going to get some really dry business type who can't handle your sense of humor subscribing, once they’ve seen what you have to say, and they've seen the mascot and all that type of thing. So, get that character in, and tell stories. You know, that's something that, again, we have really failed at over the years. We've gone very dry, we don't necessarily tell the story. And one of the things that you've done is you've brought on that extra podcast—I must admit, I can't keep up the names of all these different shows—but you've got the podcast—

Corey: [laugh]. On some level, I've made a mistake with that. I really should have everything unified around a central brand, and I didn't because I wanted to go in different directions. If I call this the Last Week in AWS podcast, I never would be able to talk to people who were not themselves focused on AWS. I've had a bunch of GCP employees and Azure folks on the show. That would never have happened if this were AWS branded.

Peter: Yes. So, stories is, you know, just really important nowadays. And I know a lot of people that read different newsletters just to see the opening paragraph, almost, each week from the person that writes the newsletter, even if they're not interested in the topic. It's like, “Oh, I just want to know what so-and-so is going to just start out with and say this week,” or this month, or whatever, even if I don't read the rest of the thing. Whereas we go very dry. We're, like, “This is the biggest headline. Bam, you should know this.”

But I think having some of those stories like you tell, you tell stories about your own business, and you do this on Twitter, you do it in email, you do it in podcast, and doing that is very important. So, yeah, I'm just going to pick two things I would do now very differently, it's, get that character in and be yourself to a certain extent, or… maybe in the TV and media sense of just, like, be a magnified, caricatured version of yourself. Like, they put makeup on people just to go on TV to make them look more like themselves, kind of just do that with how you write, and how you speak, and so on. And yeah, do the same with storytelling. Even if you don't have the most exciting stories, people just find stories very compelling nonetheless. They don't have to be super exciting, they just have to be something that's relevant to the person that's enjoying the story. So, yeah, stories and character: just follow those old, kind of, trusted things. Don't necessarily just go in completely dry, factual, like we've tended to do.

Corey: Yeah. I think there's an absolute definite problem here that people lose sight of the fact that it's all storytelling. I think cloud marketing across the board misses this, where you have to paint a picture for people, you have to be engaging, you have to be fun because let's not kid ourselves, this stuff is pretty freakin’ dry otherwise.

Peter: I guess you would probably say the same thing about if we were advising people on what their Twitter strategy is because I imagine most people listening to this podcast have a Twitter account, or have a LinkedIn account, or have something of that nature. And I have had a Twitter account, pretty much since Twitter began. And so, I've actually kept mine reasonably dry over time—again, this is a problem that I seem to have—but yours has more character, and a lot of the people that I follow on Twitter, nowadays, they have strong characters and they're not afraid to let it loose on various topics of the day with hot takes, and so on. And that seems to be where a lot of the growth is coming now.

It's people that can tell a story and relate their lives and their experiences to other people in a—it doesn't have to be entertaining way, but it has to be a way that makes you turn around and take notice. And you've done that really well, and I'm seeing a lot of people doing that very well. It's something that I think needs to be done. So, it's not just on a medium like an email newsletter, or a blog, or whatever, or even a podcast, it's also on your social as well, if you got that character in, win-win.

Corey: One last topic that I want to make sure that we both talk about here is I think that there's a tremendous value that a lot of folks miss, specifically around owning the platform. I think that, “Oh, I'm going to do this on Medium instead,” or, “I'm going to build a big Twitter following instead of focusing on an email newsletter or writing my own blog,” is a mistake. I feel like being portable and able to take your audience with you as you migrate from thing to thing is important. Podcasts can do it to a point; email seems to be the universal API that everyone tends to understand, but I don't see a lot of people talking about that nearly as much.

Peter: I'm actually a little bit more bearish on that than you might think. So, even though I have about 500,000 subscribers on email, I’m actually reasonably bearish on email, which is funny considering it seems to be a little bit of a bubble moment right now when I've built my whole business around it. But I think things are harder than they look. Deliverability is actually kind of an issue nowadays, I'm starting to experience. Which is stuff that you can work around and you can work in, but there is still a little bit of gatekeeping going on, even with email.

But I think it's just by looking at the people I've seen be really successful, like, let's pick a typical name like Gary Vaynerchuk, for example, or top YouTubers, or people that are really successful on Twitch, like your Dr Disrespect and people of that nature. Like, they’ve built names up on different platforms, and they’ve then taken people to other platforms, especially Gary Vaynerchuk is a good example of that. And he's owned his own media before, but then he seems to have more success on big platforms. Now, you see a lot of people doing this in, kind of like, what I would call the mass-market kind of areas like gaming and, marketing, and things like that. I think we might start to see a little bit more of that in the developer space as well, where you see developers have very big popular YouTube channels then, kind of, parlay that into other things; they might move that group of people somewhere else.

And I don't necessarily think they're all doing it via email. And I'm just trying to think of the people that I’ve followed over the years who I consider to be minor celebrities online, and I've not ever moved between the platforms that they're big on because they sent me an email. Like, that just doesn't happen. And the number of actual personalities that I follow in email is very, very small. And I follow a lot of personalities on YouTube and Instagram and different places, Facebook even.

So, like Robert Scoble, for example, I only follow him on Facebook. And that's not where he began; he began by blogging, and he went on to Instagram, and he went around different places in Twitter, and I think he stopped doing Twitter and moved on to Facebook. And yeah, I don't know. I’m, kind of, bearish on the idea that you need to own your own media now, and your own platforms because I've been there and I've done it. And I don't necessarily think that's where all the value is.

The value is in actually getting a reputation, a personal reputation, and this personal sense of having an audience. But then, as soon as you do a YouTube video on Instagram and say, “Oh, I've decided I'm going to try TikTok, now.” They are all straight onto TikTok, and they’re all, like, “Follow. Bam.” You've now moved part of that audience onto TikTok.

That's the real power, is actually being a compelling enough character and personality that people will, almost at a directive, will move here or there. I don't necessarily think that if you were, I don’t know, kind of a more of a milquetoast kind of person without a serious amount of personality that you could send an email to someone and say, “Oh, go and move over to this other thing that I've done.” I think the important thing now is actually being a character yourself, getting that character out there. And I say that as someone who regrets not doing that more, so that's why I think—

Corey: There’s no time like the present.

Peter: Yeah, yeah, yeah. Exactly. So, this comes from a place of regret that I'm saying this. I don't necessarily think it's all about owning the email, and that's something I've done very well at. So, [laugh] I don't know. We'll see how it plays out, but I'm going to try and do more of that type of thing, and who knows, it might work.

Corey: Yeah. I think there's a lot in there. And you can go in a bunch of different directions. I maintain that I always started out with the idea of being able to take my audience with me. So, for better or worse, I've been able to do that. It's why I care a lot more about growing the newsletter list than I do about my Twitter following. But again, meeting people where they are is important.

Peter: Yeah. But then how many people have your email list come because they've learned about you on Twitter?

Corey: Yeah. That's a really good point. I find that every time I wind up mentioning on Twitter these days that I have a newsletter, I see a sudden rush of subscriptions.

Peter: Exactly. So, if you launched a new one—I guess the good thing about email is that you can usually partially move an audience from one thing to another. So, actually, this is something that we've done really well is that—I call it the domino approach—where we started, let's say, a Ruby newsletter, and we knew that a lot of them did JavaScript, so we launched a JavaScript one and say, “Hey, Ruby people, we’re doing a JavaScript one.” And then they—the domino falls over.

And this is exactly how Facebook grew big as well, and a lot of media companies grew big is that they built off a small audience in one area and then they used this domino effect of, “What part of that audience can I take to this new thing I'm building?” And then that builds that new thing up quicker. And you could do that with email, you could do it with Twitter. But yeah. It's all about the platforms in that. I don't necessarily think it's just all going to be from your personal email that you do that.

Corey: Yeah, I think you're right. So, if people want to hear more about what you're up to, look at the various media properties that you curate, where can they find you?

Peter: Well, technically, we have a website at cooperpress.com, but as I say, it's very poorly maintained. It's the old thing of, the builder’s house is falling down.

Corey: Yeah, when you see a consultancy, for example, the great beautiful website, it's one of those how much work are you folks actually doing these days, you have time to keep your website updated?

Peter: Exactly. So, yeah, @peterc on Twitter is probably the best way to go—so that’s just P-E-T-E-R-C—and I'm always happy to answer any questions that you tweet at me. So, yeah, I'm good in that regard. I might not say I have the strongest character on Twitter, but I'm always listening, always reading everything. So, just reach out to me. I'm happy to answer any questions that you might have.

Corey: Thank you so much for taking the time to speak with me today.

Peter: Thank you. It's been great.

Corey: Peter Cooper, editor-in-chief at Cooper press. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star rating on your podcast platform of choice, whereas if you've hated it, please leave a five-star rating on your podcast platform of choice, along with an insulting comment explaining exactly why an email newsletter has no future.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Divanny Lamas
Divanny Lamas is the CEO of Transposit, the DevOps automation company. Divanny and the Transposit team are creating a world where humans interact with machines successfully to manage today’s complex technology stacks. Divanny is also a managing director at leading venture capital firm Sutter Hill Ventures. She is passionate about working with entrepreneurs to tackle ambitious technical challenges. Prior to Transposit, she began her career at Google and spent seven years at Splunk, where she saw the rise of big data and was one of the early product managers working on building out visualizations and analytics. She was responsible for product strategy, roadmap, and execution for Splunk's marquee product, Splunk Enterprise. She also served as a senior director of customer success and the head of new product introduction at Splunk. Divanny obtained a bachelor’s degree in government and computer science at Harvard University.

Links Referenced

  • Transposit
  • Sutter Hill Ventures
  • Follow Divanny on Twitter
  • Connect with Divanny on LinkedIn

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: You’ve got an incredibly complex architecture, which means monitoring it takes a dozen different tools. We all know the pain, and New Relic wants to change that. They’ve designed everything you need in one platform, with pricing that’s simple and straightforward; no more counting hosts. You can get one user and 100 gigabytes per month, totally free. Check it out at newrelic.com. Observability made simple.

Corey: This episode has been sponsored in part by our friends at Veeam. Are you tired of juggling the cost of AWS backups and recovery with your SLAs? Quit the circus act and check out Veeam. Their AWS backup and recovery solution is made to save you money—not that that’s the primary goal, mind you—while also protecting your data properly. They’re letting you protect 10 instances for free with no time limits, so test it out now. You can even find them on the AWS Marketplace at snark.cloud/backitup. Wait? Did I just endorse something on the AWS Marketplace? Wonder of wonders, I did. Look, you don’t care about backups, you care about restores, and despite the fact that multi-cloud is a dumb strategy, it’s also a realistic reality, so make sure that you’re backing up data from everywhere with a single unified point of view. Check them out at snark.cloud/backitup.

Corey: Welcome to Screaming in the Cloud, I'm Corey Quinn. I'm joined this week for a sponsored episode by Divanny Lamas, CEO of Transposit and a managing director at Sutter Hill Ventures because, you know, both of those sound like part-time jobs Divanny, welcome to the show.

Divanny: Thank you for having me, Corey.

Corey: So, I'm going to go in reverse order there and start with the Sutter Hill Ventures story. So, what is Sutter Hill Ventures, and what does a managing director do?

Divanny: So, Sutter Hill Ventures is a VC firm. We tend to focus on enterprise technology as well as hard tech. So, a lot of B2B, a lot of software, a lot of cybersecurity, IT infrastructure, all those types of things. It's an old firm, been around since 1962. And we take a slightly different approach to investing. A lot of what we do is on the incubation side. So, we work very closely with entrepreneurs to build really fabulous technology companies. And we've had some recent successes. One of the most popular ones this year is Snowflake.

Corey: Oh, yes. So, forgive my naivety on these things; I tend to come from the bootstrap world where my naive interpretation of what a VC does is, “Well, I made a lot of money by effectively winning the lottery, and now I advise other people on how they too can win the lottery.” That's probably overly cynical, but where am I wrong?

Divanny: Well, I don't think you're wrong for the majority of VC. But building companies takes a lot of different skill sets, and it takes a lot of knowledge, and not surprisingly, you get better the more you do it. So, we think that the role of a VC is to be helpful to founders to help them overcome a lot of the challenges and to simplify their lives, to try to give them good advice, and good guidance, help them hire, help them sell, help them do all the things that are important in building a startup.

Corey: Speaking of building startups, you're the CEO of Transposit. So, how did that come about? I mean, before this, you were in management roles for a while at a small company called Splunk. And—

Divanny: Tiny little startup, yeah.

Corey: Yeah, exactly. And before that, you were all kinds of other interesting places, too. You went to Harvard, for example, and you were at Google for a while, and—I don't know if you actually deprecated anything or not, which is the way you can really tell someone was at Google, but I digress. How did you wind up where you are?

Divanny: After spending some time at that startup that you mentioned, I was looking for my next role, and I wanted to get back to building and get back to the early stages. And so I met the guys at Sutter Hill, we connected, we decided that we liked each other and that we wanted to go into business together. And Transposit was one of our portfolio companies that I just got really excited by: a company that was in an area that is near and dear to me, which is APIs and kind of dealing with a lot of the challenges around API's. And when I met Transposit’s founder and CTO Tina, it was just a match made in heaven. So, we teamed up about a year ago and have been building out this company ever since.

Corey: So, what does Transposit do? It’s one of those fun names, in that I don't have to spell it for anyone. But conversely, people hear, “Oh, Transposit. I can tell by the names that they—” and then they start checking bus schedules or something. What's the Transposit and what does it do?

Divanny: So, Transposit is an automation platform for DevOps and SRE teams. So, we help modern operations teams take a lot of the things that you would traditionally call toil and automate it, which kind of on the surface sounds like a very, very broad directive and it is; it's a very ambitious goal. But in practice, we help people streamline their incident communications, we help them kind of deal with a lot of the repetitive tasks that happen as part of managing a modern operations team.

Corey: So, this sounds similar in some respects, the idea of AIOps, you're a VC. So, imagine that the direction you're going in, right?

Divanny: [laugh]. Yeah, I mean, I think I have a slightly negative opinion of the term AIOps. When I think of AIOps, I hear a lot of people who start asking questions about machine learning, and algorithms, and data science, and AI. And I’ve build a machine learning team or two in my day, and I do not consider us an AIOps company. I think that AIOps in general is a little bit of a bullshit buzzword right now, no offense to any AIOps companies out there.

But we think of it more as a data problem. So, we think of ourselves as a data company that really makes sure that all of the things that are happening on our platform are recorded, that we have full insight into what people are doing in the system, and ultimately can use that data to help them improve later on. But that's very different than building a self-learning system that's going to take on Skynet in the future.

Corey: One of the most interesting parts of the entire AI revolution, for lack of a better term, seems to be that, whether it's AI, or machine learning, or math, it's sort of a spectrum, but how you talk about it depends entirely on what you're trying to achieve. If you're an engineer, it basically talks about the polite part that we all tend to ignore, which is it's a bunch of if statements and magic and hope, and if you're trying to launch a company, it's oh, it's this massive AI system that is just a hair's breadth away from a general-purpose AI that will transform the world as it usher is in the singularity. And most conversations wind up somewhere in between those two extremes. What about the idea of building interactive runbooks and handling incident management lends itself to being a big data problem?

Divanny: Well, I think of different types of data, different categories of data. And you've got machine data; I spent a long time working on that, and that's logs and machine exhaust. And then you've got completely structured customer data; that's my records of the people that are buying things on my website. And then you've got this thing that's, kind of, unstructured data, and that unstructured data is the conversation we had on Slack, it's the set of things that humans did as part of a process. And at Transposit, we believe that that set of data has historically been undervalued, not really looked at, and certainly not integrated with those other two categories of data.

So, we kind of think of a lot of the things that happen in the process of an incident, for example, as being a really, really great place to start building systems that can structure that data to drive better insights, to help with automation. And I liken it a little bit to what's happened in terms of the structuring of marketing data. So, think of something like AdWords and the way that we take clicks and things that people do on a website, and are now able to turn that into analytics, and turn that into insights on what people are interested in and what they're buying. And so we're taking a very similar approach, but applying it to an enterprise problem, which is how do you keep the website up and running? How do you make sure that your service is meeting the needs of your customer base?

Corey: This sounds very similar, in some respects, to some of the talk that PagerDuty has been putting out for a long time. They're sort of the name that everyone thinks of in terms of incident response. In my experience, they tend to be the component of the pipeline that calls you at three in the morning with a very polite, “Wake up asshole” ringtone because production is broken and you need to fix it. People love it for that, but the story that they're telling about moving up the stack, about incident management, incident response, event intelligence, et cetera, et cetera, it sounds like they're trying to be something that the industry and its customers don't necessarily want them to be. It feels like they are trying to effectively become you folks, in some respects, or as you started off—as best I’m aware—not ever offering as your core value proposition to wake me up at 3 a.m.

Divanny: Yeah, I think PagerDuty is really good at waking people up. And it's a great company. I think everyone has been woken up by PagerDuty. I've been woken up by PagerDuty. They're really persistent and really, really good at it. But the incident doesn't end when you alert someone. That's just the first step.

Corey: I thought the incident finally ended when you conducted a blameless postmortem and blamelessly concluded that it was the person who's not in the room’s fault.

Divanny: [laugh]. Yeah, it's the pointing fingers. It's, “I’m blamelessly going to point my finger at this person.” No, look; I mean, I think PagerDuty is a great company, but we believe that there is a really big untapped area that happens between the first alert and the resolution of that incident. And today, that process is highly manual, involves a lot of tribal knowledge, and a lot of people running around with their heads cut off trying to figure out what to do.

That process is inconsistent, getting learnings from that is difficult to do, and it's very rarely automated. So, our goal is to take that very scary time, that's very painful time, that very expensive time, and simplify it: really give people guidance to help them be more effective and improve that customer experience. So, I think that we have different goals at the end of the day from companies like PagerDuty.

Corey: So, I guess the eternal question then becomes, where does automation start and stop? Where do humans start and stop? What's the handoff look like? Because every time I've seen a company—and frankly, there have been a lot of them—that try to automate the incident management story, it's always one of those things where they have a few golden examples of, it's this issue, and you just so happened to have hooked this up to all the services and tools you need to find this particular incident. And then okay, great.

What if it had been my previous incident in the real world that I've encountered? And it was, “Oh, yeah, it wouldn't have caught that yet. But great idea, we’ll catch that one too.” And it feels like it's a perpetual game of whack-a-mole as they write ever more if statements. How does that handoff work? What is the reality of that?

Divanny: Yeah I think of automation on a spectrum. And I think that most of the time, when people think of automation, they think of the automation that's been popularized by tools like RPA vendors, like the UIPaths of the world. And I call that deterministic automation. So, that means that every single time, the steps are going to be the same, and what's going to happen. And that works for a lot of back-office tasks; that doesn't really work for incidents because, to your point, incidents are unique; every single time we get an incident, it's different.

But automation can run the gamut from a checklist—a checklist is ultimately automation. It's a set of things that you need to do as a human, and you're the automaton in this scenario, and when are we not automatons—all the way to ML, AI, fully cognitive [laugh] systems. And we like to think of humans-in-the-loop problems, and human-in-the-loop problems are things that you can automate parts of it, and then you can enable the humans to do their jobs faster. So, let the humans be good at what humans are good at. Humans are good at intuition, they're good at context, they're good at understanding pieces of data that might not have gone into the original analysis by the system. But they're really bad at certain types of things. Like, they're really bad at passwords, they're really bad at repetitive tasks, they're really bad at copying and pasting scripts. So, we try to let humans be good at human stuff and let the machines be good at machine stuff.

Corey: And that assumes that you can find the humans and the machines that are respectively good at things, which is never a guarantee, but we're getting better—sometimes—at solving for those particular problems. The nice thing that I've discovered in the, I guess, couple years now that I've been loosely tracking what Transposit’s up to, is that you're not telling the same stories as all of the other, “Yes, I do that, too,” type of company. It seems that you are approaching this from a fundamental point of philosophical difference, specifically around the idea that the current way that people handle incidents is broken. How did you get to that point, and what is it you think the rest of the industry is missing?

Divanny: Well, I think that, again, so much of the world is focused on that initial page. They're focused on how do you get this thing rolling? How do you get things started? And I think that honestly if I'm really frank, I think a lot of people don't think about humans. They don't think about the human problem, you just mentioned it before.

The types of skill sets that are required to run modern operations and things that are built on a modern cloud stack, it's a very rare skill set; you find very few people that are good at that. And so I think if you kind of imagine that every single person who works at a company is a kick-ass SRE DevOps person with 30 years of experience in Kubernetes and you build systems for that person, those systems are going to be fundamentally different than the way that we see the world, which is, there are a lot of companies that are trying to modernize. There are a lot of companies that are trying to be like Google, be like Amazon, be like Facebook, and who don't have an army of Kubernetes experts and who are instead trying to deal with that gap in knowledge. And if you approach it from that perspective, where not every single person is an expert on day one, then the problems that you focus on are different. It's really oriented around knowledge management.

You start thinking a little bit about how do you make someone who's on call for the first time productive at 3 a.m. and how do you make their experience a little bit less traumatic? So, I think ultimately, it's probably based off of my experiences and Tina's experiences in the real world of being people who were on call for these types of services.

Corey: Part of the problem that we see across the industry, whenever we talk about incidents is—you're right. Everyone sort of envisions this as starting off when the pager goes off. But on some level, that's sort of a point of failure because it means that something out of the ordinary has happened to the point where you've got to wake someone up, and you've got to understand what goes on. Very often the postmortems, if you want to call them that, tend to focus on the things leading directly on up to the page of well, the disk filled up. And at some point, it was one of those, “Well, why wasn't there an alert on the disk filling up?”

Without ever getting the actual human factors that factor into all these things, such as why, in the Year of our Lord 2020, we have to care about individual disks on individual systems. That seems like something a computer should worry about more than a person, it. Never seems to take that step beyond. That's always what annoyed me about this stuff anyway. Does that align with your collective view of the world, or am I still not seeing it the right way?

Divanny: No, I think you're getting 100 percent right. I was talking to someone recently about MongoDB. And I hate pointing fingers at the database, but I will for a second.

Corey: That's okay, if they don't like it, they're going to lose that data, too.

Divanny: [laugh]. Yeah, I know, they'll make sure no one hears this. So, I was talking to this company, and the company has a bunch of really, really technical, great people in their SRE team and their DevOps engineering team. And they had one guy in the company that understood how MongoDB worked. And they kept running into issues between MongoDB queries that were causing all sorts of failures at the application tier.

But there was one guy that knew how MongoDB worked, so every single time there was an incident that, kind of, affected this particular part of the infrastructure—the problem wasn't that there was a challenge with that MongoDB query because if the rest of the company had known how to identify the query that was causing the problems, and then kill the query, that's a pretty easy thing to fix. The problem was that there was like one guy that knew how the UI worked, one guy that knew how to kill the queries. And so we helped them build out an automation that will go and create a little bit more self-service of a process around that where, when the incident happens, you can see the listing, you can kill the query, and you're guaranteed that you're not going to bring down the entirety of the system. And that's had a really big impact in how they think about incidents. But that's not a technical problem.

That's not MongoDB failing. Like, it's kind of a configuration problem, maybe. It's a human problem. It's that they don't have enough people that know how that part of the system works. And it's not surprising that with modern systems, they're very complex, and there are lots of pieces to think about, and there's lots of pieces to know. And expecting every single person who might be on call to be an expert in every single part of the system is no longer sustainable.

Corey: This episode is sponsored in part by our friends at Linode. You might be familiar with Linode; they've been around for almost 20 years. They offer Cloud in a way that makes sense rather than a way that is actively ridiculous by trying to throw everything at a wall and see what sticks. Their pricing winds up being a lot more transparent—not to mention lower—their performance kicks the crap out of most other things in this space, and—my personal favorite—whenever you call them for support, you'll get a human who's empowered to fix whatever it is that's giving you trouble. Visit linode.com/screaminginthecloud to learn more, and get $100 in credit to kick the tires. That's linode.com/screaminginthecloud.

Corey: So, I'm going to take that a step further and indulge one of my own personal favorite hobbies of dunking on large enterprises. It feels like at some point when—let's take the example of a disk filling up because it's an easy one for most folks to wrap their head around. If not, just write this down and keep writing it down until your computer stop working. The problem is that a disk fills up and that causes an outage. So, then the company goes into full-on reactive mode, where they're going to now have an alert every time a disk climbs above 80 percent.

And it's going to wake someone up. And then it wakes everyone up and they start dialing it in. And eventually, it's set in stone that all systems have a disk getting full monitor on it. The end. And this is the case for everything that changes where whenever there's a issue in production, we're going to make sure that there's now a check to make sure that specific issue never happens again, and it's like whack a mole.

And you wind up with these change advisory boards that have to go through all of these checks just to get the smallest change out into production. And it feels like they become so ossified by their own process, that the value of startups within the company is you're untethered from all of the process and box-checking that you have to do in most of the company. And that just feels like it's not the right direction at all, but it's also very hard to stem that particular tide of overactive changings.

Divanny: Yeah. No, I mean, I think that Agile, we're still having a hard time—especially in large organizations—figuring out how you implement that, and DevOps more broadly, which is really just Agile in a sexy sort of label. You look at companies and the reasons why they're sensitive to changes and to breaks are logical. It's very logical that you wouldn't want the website to go down if it's going to cost you a million dollars every 10 minutes. But the way that they handle it is so often to overlay very, very heavyweight process.

And frankly, the systems—and this is a little bit self—serving—we don't believe that the system's work for it: they're not able to expose the right level of risk. Understanding what's a high-risk change versus a low-risk change, and letting people take more control over their destinies, we're not really there yet from a tooling perspective. So, I understand why you have a bone to pick on it, but be a little bit empathetic for the big companies that are trying to implement these things, and they're stuck between a rock and a hard place.

Corey: And that's part of the challenge. It's easy to sit here and cast stones from where I sit at a ten-person company. I don't tend to work with a whole lot of regulated data. I don't have a giant pile of process and control, and my upside risk is far greater than my downside risk. At some point, as company grows, that changes significantly.

You don't usually want your bank to YOLO things into production. That's how Wells Fargo got into trouble, in some respects. Well, that was more of an ethical YOLO, but that's beside the point. The problem that I see is, at some point, you have to be much more cautious and cognizant of all the moving parts that touch things, and it feels like it's not a particularly well-bounded problem, which I guess, naively, leads me to believe that it's not really a problem begging for an AI or ML solution. What am I missing?

Divanny: Well, I guess I'd kind of take a slightly different lens on it. So, think about DevOps. So, there's two really big pieces to the phrase DevOps: one of them is Dev, and we've spent a lot of energy as an industry on the dev part. Like, on how do you build these things? How do you make sure that developers are working on the right things at the right time?

But I think the part of the process that's a little bit less paid attention to, from a process perspective, is the ops side. And effectively what we've done is we've taken a set of developers, and we've told them, “Congratulations, you are now a sysadmin.” And we've given them raises, and we've given them nice titles, and we've given people a lot of additional responsibility and much scarier things when they break, but I don't think it's that far off to say that most organizations when they think about SRE, they're basically treating their SREs like they are sysadmins that now have engineer in their title. And to me, it comes back to how do you approach modern operations in a way that is cognizant of the importance of some of these things—like don't bring down the website, don't lose data, be aware of the business needs—but make sure that the system itself is helping manage that, helping manage the interplay between the development side and the operational side, and that you're getting the right insights to the right people, that you're building continuous improvement into the process, that you learn from your incidents, and that you learn from those fancy blameless postmortems, and bring them into your overall operations. And I think that that area is very early on, largely because until now, we've been papering over this problem with people and bodies.

So, you know, how does Google deal with this? Well, I was at Google. They dealt with it by having so many people focused on this problem, and letting those people just build tons of tooling internally to make it happen. But if you're a midsize company, or you're a larger company, not every single company out there is going to build their own custom platform. It's expensive. It's hard.

And so I think that as an industry, we need to be aware of those things and give people the right set of tools that they need to be able to manage that. And that's not entirely an ML or AI problem, to your point. Like, that's not really the solution here. I think you can start off with much simpler concepts.

Corey: So, you just touched on something that I wound up going into some—shall we put this—significant depth on the Puppet, “State of DevOps Report” where they talk about internal platforms and why everyone should build one—at least that was my takeaway—and my response was incoherent screaming. And it turns out that that's fairly controversial. And I, at the time of this recording, need to sit down and really write my thoughts out in some more depth. But I'm curious as to where you stand on that particular position.

Divanny: I think that the guys at Puppet, they really hit on something really interesting because we've been seeing this trend with a lot of our customers, and it's certainly something that has grown over the last year. And I think it's the rise of the platform team. And so you've got to ask, why is the platform team coming into existence? Don't we already have DevOps teams? Don't we already have SRE teams?

What does a platform team that's different? And I think what you're starting to see is this recognition that infrastructure automation alone isn't enough. There's more that's needed, in order to be able to operate—not just develop, not just launch things out into the ether, but to make sure that these applications and these really complex systems are running well, you need a group of people whose entire focus is on helping the rest of the organization do their jobs: to make sure that things are running, that the right tooling is in place, that the right processes are being followed—I know process is a dirty word, but the right processes are being followed, that there's the right level of visibility, and that you're not forcing every single person in your organization to become an expert in every single thing. And so I actually think that the platform trend is a really exciting thing that's happening, where we're finally recognizing that this is a problem that necessitates dedicated resources and thinking and that we're not going to ask developers to solve every single problem, and instead, going to give them a support structure. Now, does that mean that every single company should be building their own custom platform? Absolutely not.

But if you look at the industry today, where is the ServiceNow for a modern operations team? It doesn't really exist. There isn't something I can buy off the shelf that helps me operate, that helps me make sure that things are running, and that has all of the different componentry. We haven't developed an ITIL for modern DevOps. And I think that that's what people are hungry for right now. They're hungry for help in knowing, like, what are the necessary pieces? What are the components of that? How do I implement it? I read your tweetstorm. It was a pretty funny breakdown, and I—

Corey: My condolences on that list.

Divanny: [laugh]. I think you said, like, “Where do I buy a DevOp?” And [laugh]—

Corey: You joke, but there's an awful lot of companies trying to sell me one.

Divanny: There are a lot of companies, but no one's really solving it comprehensively. A lot of people are solving little pieces of it, little bits of it. And so I empathize with the companies that are building up those platform teams. I think that you're going to see a lot of evolution in that space over the next couple of years, but I think it's a positive direction because instead of saying, “My SREs are going to fix everything,” or “My DevOps engineers are fixing—” we're finally recognizing that this is a real problem that deserves a real set of solutions and thinking, and it's not enough to just buy the DevOp, you've also got to build the right team around it.

Corey: It feels like there's entirely too much confusion around terms, and failing to disambiguate what people mean when they say certain things means that—because we live in a world of outrage and immediate reactionary responses to everything—means that when you say something like, “It's time to build an internal platform.” I'm going to take the least charitable interpretation of what I think a internal platform might be and then proceeded to dump all over it as, “Well, that's a big distraction from the thing your company actually does. Trying to wrap every platform-as-a-service or infrastructure-as-a-service tool that you use in some crappy web UI internally, is going to be a disaster.” That is not—I don't think—what a lot of people are talking about. But I've seen it done poorly enough times that I sit here and begin angrily shaking my fist at the concept. It feels like I'm really good at attacking strawmen.

Divanny: [laugh]. Here's the way that I like to think about it. We know that digital transformation is happening. And I know that that's a buzzword, but here's a great example. So, at the beginning of the pandemic—I live in San Francisco, a lot of businesses started closing down.

And as those businesses closed down—specifically the restaurant industry—there was this entire other set of companies that supplied them that suddenly had no customer base. And I really like cooking at home. So, there's one company, in particular, that was a wholesale fish distributor. So, all they did was go and work directly with fishermen, and that morning, the fishermen would drop off the fish, and then by that evening, those fish were being cooked at restaurants.

And so when the restaurant industry stopped, effectively shut down overnight in San Francisco, those fish had nowhere to go. And I really wanted to eat those fish. So, a lot of people like me want to get their hands on that. And that wholesale distributor pivoted to selling to home chefs. So, it started off as a listing, you know, page that they would put on their Instagram page, and then you would call someone and tell them what you wanted.

And then they stood up a website. And as they stood up the website, you still had to send an email with your request. And now they have a full checkout and I think they're launching a mobile app soon. So, I get my delivery, the name of the company is Water2Table if anyone is in San Francisco and wants to get some delicious, delicious freshly caught seafood. But that shift, what's happening there, it's not uncommon.

It's happening to every business, whether it's a small business or a medium-sized business, or a large business. And if you are a midsize business in the middle of the country, and you are trying to set up your digital presence and to support your customers in a way that they want to be supported, and you want to build in a modern way, and you want to be responsible, what do you do? Do you go hire engineers from Google? Do you hire engineers from Facebook? Like they haven't minted enough people.

So, at the end of the day, these processes, these kinds of systems, these platforms, they're an attempt for all of these companies to do what Google did, but in their own scale. And I can't blame them for it, even if I, like you sometimes, look at that, and say, “Do you really need to build your own platform or is this a vanity project?” But I think that, ultimately, you're going to see more and more of these things come more productized and make their way out into the market in a way that really solves problems for these types of companies.

Corey: And that's part of the challenge, I think, is that there's an awful lot of folks building out solutions in this space that target specific company profiles that they have themselves experienced. And you see this, I think, more with first-time founders—at least, I would guess—where they've worked for a couple startups themselves and figure ah, companies look like startups, or alternately, they wind up working at the big E enterprise companies, and then come back with a, “Oh, everything must look like that.” And the thing that surprises me the most the longer I'm in this cloud world is every time I think I've seen it all, all I have to do is talk to one more customer, and I'm going to see something that I'd never imagined would be possible. And sure it's easy to sit there and make fun of it as a gut check first reaction, but here in reality, it's much more nuanced than that. And generally, unless you're working in the Facebook ethics department, you don't show up in the morning hoping to do a crappy job at work today.

Divanny: Yeah I think that a lot of us are—probably a lot of the people listening to your podcast, honestly, live in a bubble. And that bubble that we're in is colored by the experiences that we've been through. And if your experiences and your understanding of the needs of the industry are colored by the FANG companies and hot startups that were building and never had legacy technology to deal with, never had to deal with cloud transformation, you might think, “Oh, this cloud thing seems to be kind of over. What's next? Haven't we already figured out the cloud? I think it's done.” [laugh], “I think we've got it all.” I can name at least 10 CI/CD systems off the top of my head, and you could probably put a question up and get another 50 startups listed. Isn't this a solved problem? But I think what it comes back to is everyone's solving tiny little fractions of it and assembling those is not as intuitive as it seems.

Corey: No, sadly, it never seems to be. So, thank you so much for taking the time to speak with me today. If people want to learn more about what you're up to, where can they find you?

Divanny: They can find us at transposit.com, and I am happy to give them a demo and show them a little bit of what we're building. I think it's a super exciting area, and we think ultimately automation is critical to scaling all of these types of things. The only way we're going to get to the point where we really are able to scale out these skill sets to the number of companies that need them is to embrace automation, and automation, and all angles of it. So, yeah, if anyone wants to learn more, you know where to find me.

Corey: Excellent. Thank you so much for taking the time to speak with me today. I really appreciate it. Divanny Lamas, CEO at Transposit and managing director at Sutter Hill Ventures. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you hated it, please leave a five-star review on your podcast platform of choice along with an illiterate comment explaining exactly what I got wrong about machine data and big learning.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Brooke Mitchell
Brooke is an analytical IT professional skilled at team building, data collection and evaluation, cloud computing as well as cross-functional collaboration. She is proficient at scripting and creating CI/CD pipelines to automate workflows to increase efficiency while reducing the chance of error.

Links Referenced

  • Connect with Brooke on LinkedIn
  • Follow Brooke on Twitter
  • A Cloud Guru Blog post, “Automating CI/CD With AWS CodePipeline”

Transcript
Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Gravitational is now Teleport because when way more people have heard of your product than your company, maybe that’s a sign it’s a time to change your branding. Teleport enables engineers to quickly access any computing resource, anywhere on the planet. You know, like VPNs were supposed to do before we all started working from home, and the VPNs melted like glaciers. Teleport provides a unified access plane for developers and security professionals seeking to simplify secure access to servers, applications, and data across all of your environments without the bottleneck and management overhead of traditional VPNs. This feels to me like it’s a lot like the early days of HashiCorp’s Terraform. My gut tells me this is the sort of thing that’s going to transform how people access their cloud services and environments. To learn more, visit goteleport.com.

Corey: This episode has been sponsored in part by our friends at Veeam. Are you tired of juggling the cost of AWS backups and recovery with your SLAs? Quit the circus act and check out Veeam. Their AWS backup and recovery solution is made to save you money—not that that’s the primary goal, mind you—while also protecting your data properly. They’re letting you protect 10 instances for free with no time limits, so test it out now. You can even find them on the AWS Marketplace at snark.cloud/backitup. Wait? Did I just endorse something on the AWS Marketplace? Wonder of wonders, I did. Look, you don’t care about backups, you care about restores, and despite the fact that multi-cloud is a dumb strategy, it’s also a realistic reality, so make sure that you’re backing up data from everywhere with a single unified point of view. Check them out as snark.cloud/backitup.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by my guest author of the week, Brooke Mitchell. Brooke, welcome to the show.

Brooke: Thank you so much. Hello, everyone.

Corey: So, this is always this weird moment of recording something in advance of when it goes live. To be clear, the day that I'm recording this, I'm in San Francisco and everything is dark, as in, looks middle of the night, ash is strewn through the sky. And, even last week, when I was recording things, that would have sounded ridiculous, and now it's just the state of 2020. So, I'm assuming that we're actually going to still have a newsletter going out, and an internet that works, and a functioning society—insofar as it’s functioning—next month, when this goes up. So, from that lens, it's weird now talking in advance about a thing that has happened in the past, and people are listening to it.

Brooke: Right. I completely get that. 2020 has thrown everyone for a loop, especially me. I had no idea that it would be like this. I went from—

Corey: [laughs]. If anyone had said it would be, we’d have thought they were nuts.

Brooke: Right. Nobody predicted this. [laughs].

Corey: So, I just want to start by thanking you for giving me the chance to actually not be writing a newsletter for, honestly, one of the first times in the four years or so that I've been writing this thing, and spending time with welcoming a new infant into the world. So, thanks. It's incredibly helpful to be able to step back from this and still, ideally, have a newsletter to come back to when I'm ready to return to the world.

Brooke: Oh, yes. I also want to thank you for giving me this opportunity. I'm very excited. And I was in your shoes, actually, earlier this year. So, I know from experience firsthand, it's exciting. Even the second time around, you learn so much all over.

Corey: Oh, yeah. Every time I talk to someone about having kids, like, “Oh, your first?” “No, my second.” And everyone who has a second kid starts laughing at me. Like, “[laughs].” Like, “Oh, why are you laughing?” “Oh, no reason. You'll learn.” So, I'm sure it's all positive surprises and only good things.

Brooke: Right. Right. Right. Nothing less. You'll definitely enjoy every second of it.

Corey: The cloud marketing people have gotten to having kids apparently, and it's all up and to the right, and everyone's happy all the time.

Brooke: Yes, nothing short of that.

Corey: So, what do you do these days, professionally speaking?

Brooke: So, professionally, I work as an analyst at T-Mobile. I work with our business and government accounts with reporting and automating bulk changes that they have.

Corey: So, you wound up coming to the world of Cloud somewhat recently with Forrest Brazeal’s resume challenge. First, what is that? And secondly, how’d you hear about it?

Brooke: Right. So, Forrest created the cloud resume challenge where in short, you had to upload your resume that was hosted by S3. And there was also a back end where you held the visitor count in DynamoDB, and you also had to have an AWS certification to complete it. So, you had to use different—other AWS services such as Route 53, the API Gateway, and Lambda. So, that was pretty fun. I got much more hands-on experience with that I got to dive more into JavaScript and Python, which I’ve also continued to learn in my own time.

Corey: It's fun. When I was looking into it, it was, this is a really great way of getting people to learn how this stuff all works together. And I read on, oh, use S3, and DynamoDB, and API Gateway. Huh, okay. And some of this, I would have to dig into the reference material to do myself, and I've been using AWS for a length of time that can only be described as depressing.

There's a lot that goes into it, and it's fascinating to me, seeing things like this, just because, from my side of the world, it really does open up a door for the next generation that really aligns with the way that I learned personally. If you give me a bunch of books, I'm not going to learn from that. If you send me to a class, I will zone out and not pay attention, then miss things; videos don't really work for me either. But give me a problem to solve, and suddenly I learn the things that I need to learn. I always thought that I was basically a unique unicorn like this, but it turns out, there are lots of us like this.

Brooke: Right. And I also found that I'm like that as well. And it took awhile for me to learn that that was my specific learning style, I had some struggles growing up. But now that I understand this, I'm able to learn things quicker because I know my preferred method of learning. And so with this, I learned so much more about the Cloud.

And something that has been the most rewarding from completing this challenge is the people that I've met, and also having the opportunity to help other people, and seeing them have the aha moment and figuring something out. Maybe just not giving them the answer, but helping point them in the right direction, as people did for me. It's all a rewarding experience.

Corey: One of the areas that I learned the most when I was getting started was on IRC. And that was, on the one hand, terrific because I could speak to people who were building these things that were really, I guess, accessible and willing to share what they learned and help with things. On the other side, though, there's this horrible culture of, “Surprise. Oh, you don't know how to do something. You must be stupid. Go read the manual and then bother us.”

There's this condescending attitude of, “Obviously, we're smarter than you are,” so the presumption that everyone starts from isn't anyone asking for help is a moron. And I hated that. It's so nice to not have to deal with that attitude in a lot of places, these days. And I will say that a lot of what Forrest has done, for example, has been centered around building terrific, welcoming communities where no one is mean, and everyone's encouraging. It's nice to see that the, I guess, next generation of folks learning this stuff, is going to have a way easier time than a lot of us did.

Brooke: Right. I completely agree with that. And I'm very happy to say I haven't encountered anyone who has, maybe, made me feel like I was annoying them, or I was a moron for asking questions. Everyone I've connected with has been 100 percent willing to help, or if they couldn't help me, pointing me in the right direction of someone who could. So, I would 100 percent say that that's true.

Corey: When I got started, in how, I guess, my journey to Cloud—and I've talked about in a few episodes, but apparently there are people who haven't listened to the entire back catalog. For shame.

Brooke: Oh, what’s wrong with them?

Corey: I know, right? It's almost like they have hobbies, or lives, or other things to do than listen to my ongoing love affair with the sound of my own voice.

Brooke: Right? It’s 2020. They're in quarantine. Come on.

Corey: Yeah, what's up with that? But I was a grumpy Unix systems administrator—because it's not like there's a second kind. We're all grumpy, and we're all angry, and it was coming from this world of building things out in data centers where one of the biggest server problems we had was, is it going to fall on me? And there was this evolution to Cloud, and a lot of folks resisted it and a lot of people embraced it. I was one of the folks that resisted it. I would have gotten far further, far faster if I'd been more open-minded.

But going through that transition was really interesting. And I looked around and, “Well, how in the world are we ever going to find people who can do all of this stuff because those Linux admin jobs aren't really there in the same way anymore.” But now we're seeing this resurgence in people coming to use Cloud from a completely different set of backgrounds and skill sets. And look at you. For example, you are already established in your career. You're working as an analyst, and now deciding to look into what this whole Cloud thing is and make a lateral move, rather than, “Well, I guess it's time for me to go back to school, and get a brand new degree, and start over at square one.” Which is where a lot of people erroneously seem to think that this is the path forward.

Brooke: Right. I would also say that's true a little bit about maybe my background, myself. So, maybe a unique background for people out there who are feeling stuck where they are. I had my first daughter at 19. So, I finished my Associate's Degree and after that, I pretty much said I was done with college.

I wanted to get in the working force, get a head start on things. So, I landed a job as an analyst after learning all that I could. After that, I was continuously trying to improve, so I decided again, what could I do for the Cloud. Got the AWS Solutions Architect-Associate Certification, as I mentioned earlier, got hands-on experience with that, and now I'm looking to move into a specific Cloud role.

And I have made a lot of headway with the interviewing process. I am really hopeful to start a Cloud role soon. So, there is a lot of opportunity out there. You don't always just have to go back to school. It’s very important to meet people because there’s—I have learned from firsthand experience that people are willing to help you if you're willing to do the work. And there's so many great people, so many amazing people who are willing to help, again.

Corey: From my side of the world, on paper, I have an eighth-grade education. I've never stayed at a job longer than two years, and I don't have any big tech names on my resume. Every job I've ever had has been through networking, not through applying online through a random form that’s going to automatically kick me out because it has the wrong keyword on it. One of the key lessons I learned while doing this is that you're never going to have every item that any job description wants. They’re aspirational shopping lists. Further, if you could do everything that the job requires on the first day, doesn't it sound like it’d be kind of boring?

Brooke: That makes complete sense. Before I was applying to jobs recently, I was feeling super, maybe, down, in the beginning of the process because I was thinking, “Oh, I won't be able to do this because I don't have all this experience.” But like you said, if you could do everything initially, then it would not be fun. You wouldn't be pushed to learn anything new. And I feel like that's what life is about: continuously learning, continuously improving. If you already know it all, it's kind of like where do you go from there?

Corey: One of the things that always annoyed me the most is these job descriptions that require X years of experience with some rando technology. The problem that I have with that is, there is such a wide difference between folks who have been curious and exploring these things, versus folks who are not. And you can't tell that just by number of years, you've sat in a room staring at a particular tech. I'm not even talking about the 15 years of Kubernetes experience because Kubernetes has only been around for six years.

But instead, it's this idea of, “Oh, you need to have at least X years working with something.” I've known too many people who've been in this industry for 20 years, and they don't have 20 years of experience. They have one year of experience that they repeated 20 times. I find that when I end up being the smartest person in the room about something, historically, that was time for me to find a new room.

Brooke: You know, and I feel 100 percent the same. So, I actually started reading a book today that said, “When you are the smartest or the most wealthy person in a room, it’s time to have new friends,” because I've heard several times before that you are an average of the five people closest to you, and if you are the best out of those five people you're going to be brought down a little bit. So, I think it is important to challenge yourself to go out, to make new friends, to do what you need to do to surround yourself with people who are doing better. And of course, that doesn't mean completely [00:14:13 annex], you know, the people around you once you start to do better than them, but also, seeking out new relationships. Something that I hear a lot of people say is, “No new friends,” but why not? Why not find people who are doing better who can teach you new things?

Corey: One of the things that surprises me the most is how frequently mentors of mine transition somehow into friends, where there have been times where I absolutely looked up to them and more or less took anything they told me in the way of advice as gospel. And—this is what I must do because they're great at this and I'm mediocre—and they've transitioned somehow along the way to people I view as peers. And it's not that they've gotten worse; far from it. They're better than they've ever been. But the things that I needed mentorship in, I got better at it. To a point where I no longer have anything left to learn from a lot of these folks in those specific areas. They're still friends, but it's no longer a mentorship-based relationship.

Brooke: And I think that's a beautiful way to grow. And that shows that you chose the correct people to be mentors. If they were not people that you will eventually want to call a friend, is it somebody that you trust to guide your life? So, I think that's also really important when seeking out people to, you know, to take advice from, do you actually want to mimic this person's life? Okay, they might have more money, or they might have more experience in a certain field, but are they a well-rounded individual? Is there someone that you can actually trust? I think those are other important aspects to look into.

Corey: One of the most humbling things I've done has been starting a company and growing it from just me to the 10 people we have right now at The Duckbill Group. And the reason that that's a humbling experience is that, fundamentally, you're doing everything when it’s just you, and every person you hire—unless you're doing something terribly wrong, is better at the thing that you're hiring them to do than you are. So, almost a part of the onboarding now is, “And now we break for half an hour for you to make fun of the pig’s breakfast that I have turned this thing that you're good at into before you got here.”

And we've seen a lot of—the enterprise salespeople where it's, “Wait. That's your sales process?” And it's obvious that they want to start screaming at us, but they have enough decorum not to. And it's, “Oh, you mean it helps if I email people back when they ask me for a proposal? Huh. Today, I learned.” And it's this ongoing, ever humbling experience of realizing that I'm an enthusiastic amateur and now I'm dealing with experts. And sometimes the best thing that I can do is shut up and get out of the way.

Brooke: Aha. That sounds awesome. I would like to just tell you how inspiring it is to hear that you have built your own business from the ground up. 10 employees, that's huge for something that you feel yourself or just period.

Corey: The weird part for me is that most of the core function is fixing AWS bills, but all the marketing is making fun of what is creeping up on a $2 trillion company. It’s, “Oh, good. So, what do you do for fun?” “Well, I find the most wealthy person in the world, pick this passion project that he built into the most powerful company on the planet, near enough, and then I make fun of them.” “Oh, so you're dumb and have no sense of self-preservation?” “Exactly.” And for some reason, every week that goes by, I have not been sued into oblivion. I'm kind of surprised at that.

Brooke: I think he must love you. [laughs].

Corey: I'm really hoping he has no idea who I am.

Brooke: [laughs].

Corey: I don't know.

Brooke: So, what has been maybe the most challenging part of building this business?

Corey: Hey, who's interviewing who here? If I'm being perfectly honest, it's probably managing my own psychology because when you're dealing with client work, you don't really have anyone you can talk to about this stuff because everything's heavily NDA’d, and no one wants to hear you complain. There's also this attitude that once you're running a company with staff, it's one of those, “Oh, you're making way more money than I am, so you aren't allowed to have problems anymore.” “Yeah, I've also got payroll to meet, rain or shine.” So, there's a lot of other problems that happen here.

And it's one of those things that’s very hard to describe to folks who haven't taken a stab at it themselves because from the outside, it looks easy. I will say that I have so much more respect for virtually every manager I've ever had now that I've done a lot of managing people with the singular exception of one particular boss that I had, who spoke only in metaphor and was a terrible manager. I will never work with that person or let friends work with that person. Sorry, I'm still petty. I've got it.

Brooke: That sounds definitely like an interesting experience. I could see maybe challenges with that. [laughs].

Corey: So, getting back to Cloud a little bit, when you went through the cloud resume challenge, what parts were easy for you, what parts were daunting, which took a lot of time and effort to get over? I mean, my experience of cloud services is based upon starting when there were a dozen of them, and sort of keeping up, kind of, ever since. Now, coming in on day one, it has to be profoundly overwhelming.

Brooke: Absolutely overwhelming, I would say. I did mention a little earlier that I spent six months studying for the Solutions Architect Associate exam. I didn't have any prior cloud experience, and I also wanted to be sure that I couldn't only answer the questions. So, I went through three different courses. And I took practice exams from two different people. I, like, wore those things out.

And I also tried to get hands-on experience in the Cloud. So, I followed along with all of their examples. So, that being said, it was very easy for me to host the website from S3, to set up the CloudFront distribution, and redirect HTTP traffic to HTTPS, but I didn't really have too much JavaScript experience, so I did have a slight challenge figuring out how to get the API to work on the page load. I didn't have too much experience with DynamoDB, but that was pretty simple for me to figure out how to create the table. And I did that through CloudFormation. And the hardest part for me was setting up the CI/CD pipeline, and getting my Lambda function to work. That probably took about two to two and a half weeks for me to figure out. So, I would definitely say that was a challenge. But it was definitely worth it in the end. It was so much fun.

Corey: You’ve got an incredibly complex architecture, which means monitoring it takes a dozen different tools. We all know the pain and New Relic wants to change that. They’ve designed everything you need in one platform, with pricing that’s simple and straightforward; no more counting hosts. You can get one user and 100 gigabytes per month, totally free. Check it out at newrelic.com. Observability made simple.

Corey: Oh, Lambda was a heck of a learning curve. It feels like it's an environment defined by its constraints. My first Lambda function took me two weeks to build because I'm a terrible programmer. And my most recent Lambda function took me four minutes to build because I'm a terrible programmer and don't know what tests are. But it really is awesome and an accelerator, once you get over the learning cliff. It's not even a curve; it's a cliff.

Brooke: I would 100 percent agree with you. But again, I feel like the best way to learn is to just get in there and do it. You're definitely going to bump your head a few times, you may have to reach out for help. Stack Overflow was 100 percent, my friend, okay? Other people in there are great, and I've never—also, I've asked a few questions on there and nobody has ever been rude to me. And I feel like that's maybe, also, people who had to learn and maybe had to deal with people who were not so helpful giving it back. And that's awesome.

Corey: There's really something to be said, for folks who have come to work on these things through non-traditional means. Where—by traditional, I suppose—I don't even know what I would call it these days, I guess it would be going to Stanford, getting a computer science degree, getting injected into this space with a bunch of classmates working at a startup. Of course, you're a white guy in this scenario, because that's where for some reason—all the best people who start these companies are always white guys, and they're always in the Bay Area because that's where all the good programmers live. Just ask them. And it is such a toxic, shitty environment in so many ways. Dealing with startup culture was so unpleasant for me that making fun of it is the only way I can avoid screaming sometimes.

Brooke: [laughs]. Oh, geez. So, I [laughs] don't have that experience working in startup culture, so I guess maybe I'm glad, from the way you described it.

Corey: Pro tip. It's best avoided.

Brooke: But being black in tech, I have already just maybe witnessed a few roadblocks. So, something that I am excited about in the future is breaking barriers for the next generation, for women and women of color, black people specifically, I would like for them to get more into tech, and for it to be more diverse so that everyone can be represented. That's also something that I think is equally as important, especially with artificial intelligence and machine learning.

Corey: I want to ask you about that specifically. For me, artificial intelligence slash machine learning has been something that I have viewed with a lot of scorn and disdain because it seems that every example that is trotted out publicly is either awful in the form of empowering fascism, to be very direct, or ridiculous where WeWork did some analysis and discovered at certain times there were bottlenecks at the coffee station, so they hired a second barista during those times. And it's, “Wait a minute, you spent how much in machine learning and data science to figure out that people like to drink coffee in the morning? You didn't go and talk to someone at Starbucks, downstairs, and get that answer super quickly and save a bunch of money?” There are enough people that I have tremendous respect for who are deeply involved in the space that I know that my take is not accurate. It's cynical; it's funny; not completely accurate. What about the machine learning space appeals to you?

Brooke: I personally like it because I see it as an opportunity, eventually, to become more proactive instead of reactive so companies can anticipate their needs better. Because something that I do notice, just across the board, is people—or companies reacting to situations instead of anticipating them when I feel like customers could be better served if we focused on anticipating our customer's needs.

Corey: The challenge that I found when I started down that path in the world of AWS billing, which, naively, I assumed going into it would be a terrific source for machine learning experience, just because it's a bounded problem space. There are only so many ways you can screw up an AWS bill. I claim. I still keep discovering new ones with every client I have, but I digress. It seems like it's a perfect big data problem that you can throw analytics at.

The problem that I’ve found, it goes back to people's psychology, where, “Hey, turn off those idle, EC2 instances that aren't being used for anything.” “Oh, you mean the DR site that we're going to need with maybe three seconds of warning? How about no?” And it doesn't take too long before you have enough of those recommendations that don't make any business sense before humans stop paying attention to those recommendations and they consider it to be noisy and broken. So, it is such a razor's edge that anything like that that has recommendations around things has to walk before it burns people out and everyone mutes it. It's hard to wind up installing APIs in people, legally, and people are always, for whatever it's worth the hardest problem.

Brooke: Oh, yeah, humans are ever-changing. So, something that I really want to get in personally, though, aside from machine learning—I don't know if I want to work in that specifically because I feel like a lot goes into it. But I personally like Lambda, even though it was very intense for me to learn. And also I like dealing with the API Gateway. I like serverless, pretty much. [laughs].

Corey: I have to ask when you work with the API Gateway, did you use the version 2, their HTTP API—which is the worst name ever because it is impossible to Google—or the original API Gateway that effectively is like a networking Swiss Army knife. It does everything. Not all at once. It is incredibly complicated. And, like a Swiss Army knife, the instruction manual is apparently written in Swiss German.

Brooke: Right. So, I did use HTTP method with the API Gateway. I had a little bit of a learning curve for that, too. But once I finally figured it out, I was like, “Okay, this is pretty great. This is good.”

And since then, I built another project. I built a serverless two-way SMS application, and I used the API Gateway for that. And it was much easier the second time around, of course, because I had that background knowledge from completing the cloud resume challenge. I liked it more as I continue to get my hands dirty. So—oh, also, for anyone who is just starting out, just keep trying because once you get it, it’s much more fun.

Corey: That's a great question. What would you recommend for someone who is starting out where you were a few months ago, that would have made your journey easier?

Brooke: Right. So, even though I'm super excited that I found the cloud resume challenge, I would have started to create my own projects quicker because it is also fun to solve your own problems, so to say, and you will have a little bit more passion about it when you are doing something that you created yourself. I love also getting into challenges with other people, but I feel like the learning just feels a little bit different when I solve a problem that I have for myself, or I create my own challenge, so to say.

Corey: One of the biggest problems that I see in tech, and it doesn't matter whether your front end, back end, mobile, what your stack is, is that when you look for help on the internet—let's not kid ourselves, we're all full Stack Overflow developers, if they ever disable copy and paste in that website, modern technical civilization will grind to a halt. But what I can't quite get past—and this has always been the case for my entire career, and I don't think it's going to get fixed anytime soon—is you start researching for help about what you're trying to do, and with everything we're doing, there's more than one way to do it. And it's, “Oh, why are you using that stack? Use this completely separate thing instead.” And on some level, you can keep sitting, and spinning, and never getting anywhere because you keep going back to the beginning and starting over with a different set of technologies, or a different approach. And that's something that I think is easy to get mired in. I still have to avoid that one myself.

Brooke: Oh, yeah, definitely. Me too. I kind of have faced that with my own personal projects that I have done. It's like, okay, I'm doing something one way and then I read something, it's like, “Oh, you should be doing it this way.” And now I've thrown away, now, everything that I've worked with. So, something that I'm still learning to do is work with what you already have because the work that you have already done and the time that you have spent is valuable, if it can be salvaged. Think about it before you just completely jump on to something new. What can you keep?

Corey: One of the hardest parts for me has always been, “Am I doing it, right?” I mean, I instinctively believe that whatever I'm doing, I'm wrong. And this is something that smart people would wind up getting right. Now, I'm clearly not a smart person. So, how do I wind up getting past that? I wish I had an answer, but I'm still stuck there every time.

Brooke: Wow. And I completely get that. So, something that I am personally looking into, well, one, I'm like, does it work? And then, two, can I create testing for it? And after that, maybe I can read over certain documentation, and then if I want to even take it a step further, I would maybe seek out people that I believe are smarter than me, who have more experience, if they could look at it and review it for me. But my number one thing is, does it work, and is it secure?

Corey: Yeah, I wish. Like, “Oh, cool. I'm going to go wind up finding people who are smarter than I am at this stuff—” which is my perspective, basically everyone, and invariably, the answer I always get is, “Huh. That's interesting. Wow, this is really broken. What did you do?” “I don't know, I am incredibly good at breaking things in creative ways by accident. On some level, I feel like it's going to come back around again and become a superpower. But right now, it's still annoying and terrible.

Brooke: Wow, no. I completely get that. I have broken plenty of things. Even as an analyst, there have been several things that I was like, “Okay, I'm doing this, right. This is going good.” I get to the end. “Okay, I didn't do this right at all.” [laughs]. So, I mean, just stick with it. I mean, nobody is ever going to do something 100 percent right. I've talked to seasoned engineers, they're like, “Hey, you know, I still break things.”

So, it's a never-ending cycle of learning; I would say, don't let that get you down because nobody's perfect, so don't try to hold on to a certain predefined image that you have in your head because it might not work out that way. And if it doesn't, that's okay. It could be leading you to something greater that you couldn't even have imagined before. So, just keep on trying. Nobody gets to a certain milestone without failing in some ways.

Corey: I think that's probably one of the best takeaways we have here. If people want to learn more about what you have to say, where can they find you?

Brooke: So, you can find me on LinkedIn, Brooke Mitchell, I'm also on Twitter. My handle is bdmitchell_. So, B-D-M-I-T-C-H-E-L-L underscore.

Corey: And we will put links to those in the [00:33:08 show notes].

Brooke: Awesome. I would love to connect with everyone. I'm planning on doing additional tutorials soon. I was so excited, Forrest reached out to me about creating a blog post for CI/CD pipelines, so that was published on A Cloud Guru blog. And I'm working on some other things now in the Cloud that I would like to create tutorials for. And if there's a need in the community, someone reach out to me, and if I can figure it out, I'm 100 percent will, and post tutorials for it.

Corey: Excellent. Thank you once again, both for taking the time to speak with me today and taking over the newsletter this week. I really appreciate it.

Brooke: Oh, it is my pleasure too, Corey. I also want to thank you again for having me and giving me the opportunity to take over the newsletter this week. And wishing you and your wife the best with the new baby.

Corey: Well, thank you. At the point in time recording this, it's all optimistic and upside, and we don't know. By the time this airs, we'll probably know. I'm just hoping they're a sleeper.

Brooke: [laughs]. Yeah, sleepers are 100 percent the best. I think you'll get a sleeper.

Corey: [laughs]. We'll find out. Brooke Mitchell, analyst at T-Mobile and rising star in the world of Cloud. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts, whereas if you've hated this podcast, please leave a five-star review on Apple Podcasts and a comment telling me why my entire take on machine learning is ridiculous.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Liz ZalmanLiz Zalman is the Co-Founder & CEO of strongDM. Previously she was Co-Founder and CEO of the cross-device profile company Media Armor. After its acquisition, she served as VP of Analytics at the acquirer, Nomi. With over 15 years of experience leading data-driven organizations, she is an expert in analytics, data privacy, and security.

Links Referenced

  • strongDM
  • Connect with Liz on LinkedIn

TranscriptAnnouncer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: You’ve got an incredibly complex architecture, which means monitoring it takes a dozen different tools. We all know the pain and New Relic wants to change that. They’ve designed everything you need in one platform, with pricing that’s simple and straightforward; no more counting hosts. You can get one user and 100 gigabytes per month, totally free. Check it out at newrelic.com. Observability made simple.

Corey: This episode has been sponsored in part by our friends at Veeam. Are you tired of juggling the cost of AWS backups and recovery with your SLAs? Quit the circus act and check out Veeam. Their AWS backup and recovery solution is made to save you money—not that that’s the primary goal, mind you—while also protecting your data properly. They’re letting you protect 10 instances for free with no time limits, so test it out now. You can even find them on the AWS Marketplace at snark.cloud/backitup. Wait? Did I just endorse something on the AWS Marketplace? Wonder of wonders, I did. Look, you don’t care about backups, you care about restores, and despite the fact that multi-cloud is a dumb strategy, it’s also a realistic reality, so make sure that you’re backing up data from everywhere with a single unified point of view. Check them out as snark.cloud/backitup.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Liz Zalman, the co-founder, and CEO of a company called strongDM. Liz, welcome to the show.

Liz: Thanks, Corey.

Corey: So, let's start at the beginning. What is a strongDM? It sounds like something that involves a little bit too much salty language on Twitter when I want to speak to someone privately. But I'm guessing that's not it.

Liz: That's not it. And in fact, there's an ongoing lottery as to what the DM actually stands for. The current winner is something called ‘Dragon Matrix.’

Corey: Oh, I like that. Or you could go with the test environment called saltyDM or something like that. It'll just go, it'll go super well because no one will ever get the wrong idea. So, what is it really?

Liz: StrongDM is a proxy, and it enables DevOps teams to both manage and audit access to infrastructure. And infrastructure could be servers, or databases, or Kubernetes clusters, or web apps or the cloud drivers themselves.

Corey: So, back in the dark ages, when I was a grumpy Unix systems administrator because it's not like there's another kind of Unix systems administrator out there, I found that all of my access control for infrastructure was gated by VPNs in a somewhat traditional environment. You'd have these fleets of data centers, and then there would be a VPN server, and if you're really clever about it, a backup VPN server and people would connect to that when they needed access for things. And I haven't gone deep into that world since I moved to Cloud, making all of security someone else's problem to worry about, or at least that’s what the brand marketing tells me. I haven't had to think about it since. What's changed?

Liz: So, I think a lot has changed, I think a lot has still stayed the same, simply because traditional security hasn't necessarily been reimagined. So, many of our customers today still have VPNs in place. But even if you have it, it's not enough. The analogy that I like to give there is a banking one. A bank doesn't just lock its front door. That would be the equivalent of a VPN getting you onto some sort of a network. There's also a bank vault, and then there's also security cameras to see what somebody is doing. So, VPNs are good, but they don't tell the entire story.

Corey: It also didn't work well with Cloud in any reasonable way because, back when I was doing this, I would go on the AWS marketplace and spin up the OpenVPN appliance that was managed by hand, it didn't cost a whole heck of a lot. And I just kept that thing as its own beautiful hand configured unicorn there. And after I got out of managing infrastructure, AWS finally looked at this and saw the real problem with it that they then solved, namely, no one was paying AWS by the hour for this thing. So, they launched the AWS Client VPN, which sounds like, oh, good, a managed service that's slightly less hand configured. More expensive, but it still feels like the exact same approach and exact same technology. You're talking about something different. Where's the delineation?

Liz: Yeah, totally. And you can do that in AWS, but who were the people who were actually sitting there in the console spinning it up? But if you zoom out, infrastructure access used to be managed by somebody walking into an office and sitting on the corporate network, and there were only a few people that had access to that SQL Server, Oracle monolith, sitting next to you, or in the colo down the street. And those boxes had Active Directory sitting on top of it. And I think to your point, there were sysadmins managing that.

And then Cloud came, and you all of a sudden had a proliferation of systems, of database management systems, of server operating systems. Kubernetes came into being in production workloads, what, 18 months ago now, and then you have CLI interfaces for the clouds themselves. And if you zoom out, you have, now, an entirely distributed workforce, you have lots of people who need access to lots of things sitting in lots of different places. And you actually need to sit there and say, “Okay, how do I actually manage access to this?”

Well, I can tell you, AD doesn't speak Druid, right? And Octave certainly doesn't directly speak to Sybase. And so what do you do? You have to think about how you get people access to the systems that they need, in the fashion in which they need it, and it needs to be done in an easy, seamless way because security is known for spending a ridiculous amount of money on shelfware. Why does it become shelfware? Because it's so freakin’ hard to deploy.

Corey: It feels to me like there's a delineation somewhere, and you see this at, I guess, most startups, even some of the unicorns that have come out of the world, where there was a certain point in time—and this sort of—it feels like company archaeology, where you look at their office networks, and before a certain point, they had all their on-premises data closets running all the local stuff. And then at some point, there's a shift, where, “So, what do you have at your local office network?” “Uh, WiFi access points, and maybe a printer,” and that's usually it give or take. There's no privileged position for being on that network. It almost forces a rethinking of what it means to get access into an environment. You're no longer granted that access based upon what network you happen to be on, but rather who you are. As they say, “Identity is the new perimeter.” Is that aligned with how you see the world, or is there still a giant missing piece?

Liz: No, I couldn't agree with you more. There's getting access to the network itself, and then there's getting access to the things that you actually need access to, and to your point at the specific role or permission level that you need it. And then all of that has to be done with an ability to make sure nothing bad is going on; having that audit log if you need it. So, that's exactly right. We had a customer sign up with us once, and they're in France, actually, and I believe they launched and when they started with us, they had 17 different VPNs. That's how they were controlling network ingress.

Corey: Oh, and I've walked around the RSA Expo hall floor, and it's clear that the answer is you need a 19th VPN, and I would love to sell it to you.

Liz: [laugh]. Yes. People want to throw that VPN out. They want to be able to essentially dispense with that layer and just create that essentially a point-to-point network, right, if it's based upon who you are. Identity is, “Liz needs access to this particular server with non-sudo privileges.” How do we create that connection in a secure way?

Corey: So, a dedicated circuit vendor just heard part of what you said and made cash register noises with their mouth.

Liz: [laugh].

Corey: But that's not feasible for how people work with internet technologies these days.

Liz: Well, I think it is feasible. That's why we started the company.

Corey: Running dedicated circuits everywhere just isn't feasible to the way that most people use the internet these days, unless you have, you know, Google level of money.

Liz: That is true. And I think even Google has had challenges with taking—I mean, what is the Encore? It's a series of white papers and research papers into how you might create zero trust at scale, and yet how many people have actually implemented that commercially? Google certainly hasn't.

Corey: Oh, what's amazing is you see this at conferences all the time, where people get up and talk about what they're doing at their companies. “And you're like, wow, my environment's crap. I wish I could build something like that.” And the person next to you says, “Yeah, me too.” When you look at their badge and they work at the same company as the speaker. It's conferenceware: it isn't real.

Every time I see something like, this is how Netflix does the DevOps—or whatever it is that they're trying to do this week—yeah, in some groups and some teams, absolutely, but every environment is a burning tire fire of sadness and regret. The only question is how honest people are going to be about that.

Liz: It's true. We've had—I mean, especially get into larger enterprises, the number of times I've heard, “We're using X,” and then you get into the conversation, and there's a team to get X implemented, whatever X might be, but, like, they're three years away from instantiating that. It's nuts. When people say they have a ten-year plan to migrate to the Cloud, I mean, they mean it.

Corey: Oh, absolutely. One of my first jobs was working in a university. And this was, oh I want to say back in 2006, one of the projects I worked on the year I was there was a wireless deployment with an eye toward getting it finally completed sometime between 2010 and 2012. And I'm staring at this going, that is so far future out there, we may not even have universities by then. That's 2010; that's years away.

I developed something of a sense of patience since then, but these multi-year rollouts that seem so set in stone with a vision that is clearly built on technologies that are still shifting, and evolving, and improving, it just doesn't seem like the right path. But there's almost a level of ossification beyond a certain point of scale, where it's hard to do anything else.

Liz: It's true. I wonder if you put yourselves in the shoes of the—what are they—they’re an innovation departments. I mean, you have very, very, very smart people sitting in these departments, and you wonder how much they're actually able to get pushed through. I started off as an analyst.

It was my very first job at a school, and I was talking to my boss, who's the head of analytics. And he said, “Think about what you want your next step to be because analytics has a tendency to be all of the power and none of the responsibility. You can just put a report in front of people and say, ‘here's the answer, and you decide what to do,’” or you can really have a hand in the outcome, which, right, is the same thing as consulting. And I see that tendency in innovation arms as well, where are these people are so smart, and so forward-thinking and yet, to your point about ossification, can they actually go and get these changes recommended and done in a timeframe that makes sense? Technology is always changing, and so what happens in one year? Is it already now out of style? Or is this something better?

Corey: Let's get a little bit more into specifics. Let's say I go to strongDM.com, I go to your buy page, I order one, and put in my credit card information. “Would I like it gift wrapped?” “Absolutely.” Yay, free two-day shipping. It shows up, I take it out of the box. What do I do with it from there? That may not be how SaaS products get sold, but I'm still fairly old-fashioned. What happens next? What is my user experience going through it and what problem is it solving that, as you’ve said—and we've all said that VPNs don't solve problems. What problem does this solve? How is life better now that I’ve bought strongDM?

Liz: You as a buyer, your problem is that you are managing access to infrastructure by hand. Your default state is to essentially—like, Terraform, right, infrastructure as code. And you want to be able to encode access as well and you can't do it today. So, that is the gap that we're trying to close for administrators. You should be able to hire somebody, and they should immediately get access to the things that they need to do their job from a least privilege perspective.

As a user, you should be able to connect to the things that you need and to not have to change your workflow and to not have to have friction, and to not have to call your IT tests and say, “WTF. This is breaking?” Or, “Why can't I get onto the VPN?” Or, “Why is this timing out?” You should just be able to connect to things.

And then the people who are administering it, you should be able to get this up really quickly. Our metric of success internally is, how long does it take you to get greenlit? And ‘greenlit’ means you got a gateway up, it's the entry point to your network, you register a database, or server, or whatever you want to test connecting to, and then you make that connection happen. I think the record is something like just under five minutes. We want you to feel the possibility of what it could be without the VPN, and without the traditional methods or scripts that you've held together today. What if it could be something different? I want you to feel that answer as quickly as possible.

Corey: So, it effectively unifies SSO, secure network access, and a single point of provisioning.

Liz: Yep.

Corey: Is that roughly equivalent?

Liz: That's exactly right. Delegate authentication to your SSO or identity provider, create a secure tunnel between you and the thing you're trying to get access to, never have credentials on the end-users workstation, and get them connected. You have it exactly right, Corey.

Corey: Excellent. So, you wind up integrating with—according to your site—a whole bunch of different technologies, a bunch of different AWS services, you call out databases explicitly, but sadly, not my favorite database: Route 53.

Liz: [laugh].

Corey: But you do wind up talking about a bunch of different, I guess, systems of record. Splunk is on the list, Scalar’s on the list, a bunch of different AWS services, and the rest. Is this really done as a item by item process as far as getting it hooked up to each one of the data stores and the systems that the company needs people to be able to access, or does it wind up instead being something that lends itself to a more unified rollout process?

Because, “Oh, we'll give you a free proof of concept rollout.” “Great. That's going to take me eight months of engineering time, and once I've done that, I may as well buy it because whatever you're charging me is tiny compared to what it cost me in engineering time.” What is the rollout process, and how does that look for existing environments that aren’t greenfield?

Liz: Yep, it is easy. You either buy the product or you don't, I'm not incentivized for you to do more or less.

Corey: Oh, it's a per user per month chart. That's kind of amazing. It's the exact opposite of an AWS billing model.

Liz: That's correct. It is literally one price fits all. You buy it or you don't buy it; everything is included. I just went through a procurement process with an unnamed CRM, and I have to tell you, I don't even know how these companies sell money. I don't want to sit through a PowerPoint. I want to try it on, I want to make sure it fits my use cases, and I want to click on the button to buy.

Corey: What I think that a lot of SaaS companies completely miss is how their products or services get tested out. If I'm trying something new, and you've made a product that solves my problem, very often because I am who I am, I'm having a sleepless night. It's two in the morning; I'm trying to build something and see if it works as a proof of concept, and if I need to talk to a sales team before I can ever get involved in my environment at all, well, you don't generally have salespeople around at 2 a.m. I'm certainly not going to want to have a conversation then. I'm probably going to look elsewhere.

But wow, that was easy; the onboarding story where I can set this up in my ridiculous twitterforpets.com proof-of-concept environment—okay, this works. Now I can start expanding beyond that, and okay, yeah, we're a giant company. Now it's time to do the enterprise sales dance. But there's something about self-service that I think a lot of companies miss out on.

Liz: And it's particularly important for this buyer infrastructure or DevOps. I mean, this is a closed-door community, they have been playing Minecraft and beer pong with their buddies, the same buddies they've had for 30 years. They're sitting in closed-door Slack communities. They trade tips, and tools, and secrets, and they only listen to their peer group. They don't want any sales bullshit, to exactly your point. They want to click on the button, they want to try it on, and they are going to decide what makes sense for them. That's it. And we should honor that, and that's what we've tried to honor.

Corey: Out of curiosity, one of the things that I noticed that you emphasize specifically, on what it is that you integrate with, you're all over the map, you have things like different versions of Linux, different AWS services, but you have a particular affinity for calling out specific databases. Why is that?

Liz: I don't want to replace or change anything that you're already buying today. So, going back to the SSO or identity provider, you have already decided that Active Directory is your IDP of choice; great, we're going to integrate with that. You have already purchased Sumo Logic or Splunk. Wonderful, here's a button you can click to send your logs to them. You already use Duo for MFA. Great, here's how you connect Duo.

So, at the end of the day, strong is designed, I think of zoom out, to be an infrastructure API. I want to control network access and I want to be able to introspect as to what's happening to provide you with an audit trail: layer three and layer seven. We need to be able to natively support every single thing that you're using today. So, strong is the only company that supports databases natively; there's no other one. You can go to other places, perhaps, for SSH or RDP, but then you're just buying a point solution.

So, the vast majority of companies, they have old legacy stuff, Oracle, Sybase, we just built Taradata, Db2. And then they also have memcached and they're also using Kubernetes in production. And so you have to honor where people are at, which is they've got a hodgepodge of stuff and they want something which manages access to all of that stuff seamlessly. That's what we've tried to honor as we've built this system.

Corey: This episode is sponsored in part by our friends at Linode. You might be familiar with Linode; I mean, they've been around for almost 20 years. They offer Cloud in a way that makes sense rather than a way that is actively ridiculous by trying to throw everything at a wall and see what sticks. Their pricing winds up being a lot more transparent—not to mention lower—their performance kicks the crap out of most other things in this space, and—my personal favorite—whenever you call them for support, you'll get a human who's empowered to fix whatever it is that's giving you trouble. Visit linode.com/screaminginthecloud to learn more and get $100 in credit to kick the tires. That's linode.com/screaminginthecloud.

Corey: There's really something to be said for meeting customers where they are. That's something a lot of products seem to miss, especially around the identity story, where it's, “Oh, you've been using your AWS account for ages. Great, take out all of those IAM users and use this thing instead that is going to force a complete rethink of how every single system that does anything approaching provisioning interacts with the environment. Step two…” it's one of those incredible ‘boil the ocean’ style of stories, and God, I hate that phrase, but I'm still using it anyway. It feels like it is almost insurmountable.

I look at how I set up my own AWS environment four years ago, and I'm looking at it and I put my hands on my hips, and I survey the landscape, and I say, “Yep, absolutely. This is trash. What was I thinking?” But I'm stuck with it because migrating to a new form of management across multiple AWS accounts is incredibly daunting. How do you get around that problem? Because if—you don't have it, I know that because when I was doing homework for the show. No one I spoke to wound up complaining about the onboarding, so you've clearly solved for this problem somehow.

Liz: So, I think the thing is, how much of a pain in the ass is it today for the person who's managing that access? I remember, I got off the plane once in SFO, and there was a beautiful, there was a new ad campaign that Redis had launched, and the ad campaign said, “‘I love my slow database,’ said nobody, ever.”

Corey: It's like they’ve never heard of blockchain people.

Liz: [laugh].

Corey: Please continue.

Liz: And with infrastructure, I have never heard anybody say, “I love AWS’s IAM.” AWS is not incentivized to make IAM easy. And even if you're using it, you're using other things that don't speak IAM, or you're on GCP, or you have owned and operated data centers. And so the person is managing access, their fingers are bleeding. They selfishly want to get out of this space. And so if they see a way to do that, to ease the pain, the switching cost almost becomes irrelevant because they get connected and they're like, “Oh, my God, there is a better way, and I can see the other side. I see the light at the end of the tunnel.”

Corey: One question I do have for you. This is more of a business ownership perspective. On strongDM.com, toward the bottom of the page, there's basically a whole section that has—says, “And yes, you can throw out the VPN.” And then there's a button with a trash can on it, and that is labeled ‘Trash VPN.’ Now, to my understanding, Trash VPN is a Cisco product. How did you get permission to use their branding on your website?

Liz: [laugh]. Did you click the button?

Corey: I did, and it is amazing. And I encourage everyone listening to this to click that button. It is absolutely phenomenal. Cisco need not listen to this episode.

Liz: Right. [laugh]. We were designing it, and we were spitballing, and I don't remember who said what [00:21:46 crosstalk], “Wait a second, can we actually put a trashcan there?” And then it animated, and it was glorious.

Corey: Oh, and then there's an emoji if you click it, it has a hidden easter egg that—yep. You need to look at this website, folks, this is worth looking into.

Liz: When you deploy a strongDM gateway, the gateway is the entry point to your network, and you can give it a DNS entry—make it publicly available—or you can put it on the corporate network, or you can put it behind a VPN; it deploys however you want. In roughly half of the cases, people end up asking, “Well, can I just get rid of the VPN thing because the only reason why I had the VPN was to get into the network access itself for infrastructure, and this kind of does the same job.” So, they just end up connecting the dots. There's no line item for a VPN. I don't have a line item in my budget: it's free, it's shitty, and they're free, so we don't sell like that, but it's a conclusion that people come to on their own, which is why we put that in.

Corey: It's phenomenal. So, other things that have come up when I was doing homework for this show, which is kind of awesome. I don't think I've ever found anything quite this incriminating. So, we're going to go with it anyway, and see if it survives editing.

When you were 28 years old, you tried to become a tennis pro, and you lost your first and only match in 20 minutes, which is a great story in its own right, but more to the point, I did things when I was 28 where when I got married, I took a new last name in order to bury it. You are actively open about this. So, first, tell me about being a tennis pro, and secondly, why would you, I guess, not do everything in your considerable power to make sure that that story never saw the light of day?

Liz: Because I'm not afraid to embarrass myself. I'll answer the second question first. I think life is nothing but relationships and experiences, and [BLEEP] it. If I'm going to fail, it doesn't matter. I want it known that I had that experience. I'm proud of it.

Anyways, I was in Chicago, I didn't have any friends—to give background on the story to listeners—and so I decided to join tennis, I played in eighth grade and never picked up a racket again. And I was like, man, I love this. And then somehow I was spitballing one day, and I was like oh my god, you can actually go and register for a qualifying tournament. The prize was $10,000 if you won, and it was essentially free to enter, and there was one in Atlanta.

So, I flew down to Atlanta—a friend of mine lived there—and I entered the tournament and I showed up having never done this before. And I was wearing what I would normally wear when I played tennis, which was a Red Sox hat. I grew up in Boston, and I get on the court for this match, and the referee comes up to me—or the umpire I don’t—see I don't even know what they're called. And she says, “Hey, you can't wear that hat.” And I said, “Why not?” And she said, “Because it's branded with a logo. You can only wear the logo of a company or entity that sponsors you.”

Corey: I'm already getting angry just listening to that story.

Liz: I know, and I didn't have another hat with me, and so the salesperson to me retorts, “Well, how do the Red Sox aren't my sponsor?” [laugh].

Corey: Amazing. Oh, that's fine. They can be but they haven't sponsored us. You've got to pay our fee. It's like the re:Invent expo hall.

Liz: That's right, right. Paid a brand at eye level.

Corey: [laugh]. There's something to be said though, in seriousness, about embracing failure, especially once you for better or worse become something of a role model. And like it or not, you started and founded multiple successful companies. Sorry: you're a role model. Some of us look up to you.

There's something to be said about being very open and transparent about failures because everyone fails, and people don't talk about it nearly enough. So, when people fail, as they invariably do, they start to worry that it's just them. And it's not.

Liz: I agree. I think, on interviews, I'm going to not answer the—are you a role model thing, but [00:25:32 crosstalk]—

Corey: Oh, that wasn't a question. That was a statement. Like it or not, you've got to deal with it.

Liz: [laugh]. Oh, man. On interviews, I get asked, “What's your vision for the company? What's your exit strategy? Tell me what we're going to be like in five years.” And the answer that I give is, “I have no idea.”

Companies get bought, they don't get sold. We had two customers who were supposed to IPO this year. One couldn't do it because they were in residential construction and that got killed, and the other is effectively out of business because they were in retail. And so I can't control anything. The things that I can control are building a company, building a product that people want to pay at least $1 for, and doing so at scale.

And then based on that, you get the ability to have outcomes. And so what I end up saying is, “I will speak openly and honestly about successes and failures.” Everybody at the company knows when we gain a customer, when we lose a customer, when we have a huge win, when we have a major loss, how much revenue we're making. Having been a part of a team in a startup before, I want to actually see the fruits of my labor, and I want transparency in what's happening. Don't sugarcoat it.

Corey: “Don't ever bring me bad news.” “Cool, I'll hide the thing you really need to know.” It's moronic.

Liz: Corey, it's funny, in my last company that I ended up selling, actually, we were running out of money and we weren't sure if we're going to pull an acquisition over the finish line. And I remember all the advice that I got from investors was sit down and tell the truth. And I was terrified; my hands were shaking. And I sat the team down and I said, we have two weeks of cash. I want you to stop working and I want you to start looking for another job, and I'm going to bend over backwards to get you another job. And the team sat there. And then one of them picked up their heads and says, “Are you done speaking? Can we get back to work now?”

Corey: That says a lot. Either they really believe in you or they are freaking terrified, one way or the other. We’ll take the charitable [laugh] interpretation instead. But no, in seriousness, that is…

Liz: Yeah.

Corey: It took me a long time to be convinced to go beyond just being an independent consultant because it terrified me to have other people's welfare resting on me. And I still don't know if I'll ever get used to it or not.

Liz: It's scary. Yeah. I mean, you put food on people's table. And the same thing in interviewing: so many people this year have lost their jobs, and you're on the phone with them, and I—[sigh]… it's just people at the end of the day, and relationships, and empathy. Everybody is fighting a battle and you may or may not know what battle they're fighting. It's a scary position to be in, and it's one that's very humbling.

Corey: Yeah, they say it's lonely at the top. It's not because so few people make it, it's because it's a different class of problem. You can't talk too openly in too many fora about these sorts of things with just anyone because until people have walked this path to some level themselves, they don't experience it or live it nearly as viscerally. It's a lonely road in some ways.

Liz: Thank goodness for co-founders. I am in awe of people who start companies as solo founders. I—

Corey: God, yes. I took on a business partner after two years of doing this, and that was the tipping point. When Mike came aboard, suddenly, I wasn't alone anymore. I could talk to people about NDA’d things without, you know, breaking contracts. So, it was nice to actually be able to vent to someone, and talk through things, and have the support. Hardest part of business running, in my experience, has been managing my own psychology.

Liz: I agree with that.

Corey: I want to thank you for taking the time to speak with me. If people want to learn more about you, and what you're up to, and what your company does, and exactly how strong your DM is, where can they find you?

Liz: [laugh]. It's www dot strongDM—that's D as in David, M as in Mary dot com, not BM like the toilet.

Corey: Excellent. [laugh].

Liz: [laugh].

Corey: Thanks once again for taking the time to speak with me. We'll throw links to that in the [00:29:22 show notes] while being careful of spelling and pronunciation.

Liz: Thank you, Corey.

Corey: Of course. Liz Zalman, co-founder, and CEO of strongDM. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on your podcast app of choice, whereas if you've hated this podcast, please leave a five-star review on your podcast app of choice along with a story about how you could have been a tennis pro but chose not to do.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Thomas DullienThomas Dullien / Halvar Flake is a security researcher / entrepreneur known for his contributions to the theory and practice of vulnerability development and software reverse engineering. He built and ran a company for reverse engineering tools that got acquired by Google; he also worked on a wide range of topics - like turning security patches into attacks turning physics-induced DRAM bitflips into useful attacks. After a few years of Google Project Zero, he is now co-founder of a startup called http://optimyze.cloud that focuses on efficient computation -- helping companies save money by wasting fewer cycles, and helping reduce energy waste in the process.

Links Referenced

  • optimyze.cloud
  • Quoted Tweet
  • Follow Thomas on Twitter
  • Connect with Thomas on LinkedIn
  • Thomas’ personal site

TranscriptAnnouncer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by our friends at Linode. You might be familiar with Linode; I mean, they've been around for almost 20 years. They offer Cloud in a way that makes sense rather than a way that is actively ridiculous by trying to throw everything at a wall and see what sticks. Their pricing winds up being a lot more transparent—not to mention lower—their performance kicks the crap out of most other things in this space, and—my personal favorite—whenever you call them for support, you'll get a human who's empowered to fix whatever it is that's giving you trouble. Visit linode.com/screaminginthecloud to learn more and get $100 in credit to kick the tires. That's linode.com/screaminginthecloud.

Corey: This episode has been sponsored in part by our friends at Veeam. Are you tired of juggling the cost of AWS backups and recovery with your SLAs? Quit the circus act and check out Veeam. Their AWS backup and recovery solution is made to save you money—not that that’s the primary goal, mind you—while also protecting your data properly. They’re letting you protect 10 instances for free with no time limits, so test it out now. You can even find them on the AWS Marketplace at snark.cloud/backitup. Wait? Did I just endorse something on the AWS Marketplace? Wonder of wonders, I did. Look, you don’t care about backups, you care about restores, and despite the fact that multi-cloud is a dumb strategy, it’s also a realistic reality, so make sure that you’re backing up data from everywhere with a single unified point of view. Check them out as snark.cloud/backitup.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Thomas Dullien, the CEO of optimyze.cloud. That's optimyze, O-P-T-I-M-Y-Z-E, so we know it's a startup. Thomas, welcome to the show.

Thomas: Hey, nice to be here. Thank you.

Corey: Of course. So, let's start with my, I guess, snarky comment at the beginning here, which is yeah, you misspell a word so clearly you're a startup. You have a trendy domain, in this case, dot cloud, which is great. What do you folks do as a company?

Thomas: Well, originally, we set out to try to help people reduce their cloud bill by taking a bit of an unorthodox approach.

Corey: That sounds like a familiar story.

Thomas: Well, familiar perhaps, but our approach was that I had seen just a tremendous amount of wasteful computation everywhere. And the hypothesis behind our company was, “Hey, with Moore's Law ending, software efficiency will become important again.” Meaning people will actually care about software being efficient. And the reason for this is A) Moore's law is ending, and the second reason is, now that everybody is a SaaS vendor, instead of a software vendor, all of a sudden the software vendor pays for the inefficiency. Like, in the past, if you bought a copy of Photoshop, you had to buy a new Macintosh along with it. Nowadays, it's the vendor for the software that actually pays for it. So, our entire hypothesis was that there's got to be a way to optimize code and then make things run faster and, yeah, make people happy doing this.

Corey: One of the strange things that I find is that every time I talked to a company who's involved in the cloud cost optimization space—and again, full disclosure, I work at The Duckbill Group. That's what we do on the AWS side. And we take a sort of marriage-counseling-based approach for finance and engineering so they stop talking past each other. And that's all well and good, but that's a relatively standard services story: part tools, part services-based consultancy, and it has its own appeal and drawbacks, of course.

What I find interesting is that most tooling companies always do what comes down to be more or less the same ridiculous thing, which is, “Ah, the dashboards are crappy so we're going to build a better dashboard.” “Great, awesome. What problem are you actually solving?” And it becomes, “Well, we don't like Cost Explorer, so here's something else that we did in Kibana instead,” or whatnot. Great. I don't necessarily see that that solves the customer pain points. You have done something very odd in that view, which is that you're not building a restated dashboard of what people will already find from native tools, you're looking at one very specific area. What is it?

Thomas: Yeah, so in our case, we're really looking at where are people spending their computational cycles. I mean, everybody knows that you can profile a single application on their laptop or on their computer, but then once you get to a certain scale, that gets really, really weird. And I like to think about software in the sense that we're building essentially an operating system for an entire data center these days. Like that's really what's happening inside Google; that's what's happening inside AWS. And you need to measure where your time is spent if you want to make things faster.

Everybody who's profiled their own software usually find some really low hanging fruit, and then goes and fixes them. And to some extent, we have a huge disconnect these days between what the developer is writing and what is actually the feedback loop to tell the developer, “Hey, this change here just caused X dollars of extra cost.” So, in some sense, what we want to build something that tells the developer, this line of code here is generating this amount of cost. And we kind of think that developers will make better decisions once they know what the cost is. Full disclosure, I used to work at Google for a fairly long time, and Google wrote a paper in 2010 about a system they have built that's called the Google-wide profiler, and the results of that system inside Google were quite hilarious. Like, they figured out that they're spending, what, 15 percent of all cycles inside Google on gzip when they first started measuring. So, we thought that’s got to be a useful thing for other people to have, too.

Corey: And I'm a mixed opinion on that one because, first, congratulations on working for Google, all things must end, and you've since left Google, and they’re Google, so they're really good at turning things off. But that's beside the point. This, on some level, might feel like it's a problem that a company like Google will experience, where optimizing their code to make it more cost-effective makes an awful lot of sense, given how they operate and what they do. But then I look at customers I work with where they have massive cloud bills, but fleet-wide, their utilization averages to something like 7 percent.

So, these instances are often sitting there, very bored. Optimizing that codebase to run more efficiently saves them approximately zero dollars, in most cases. That is, in many ways, the reality of large scale companies that are not, you know, the hyperscaling Googles of the world. That said, I can see a use case for this sort of thing when you have specific scaled out workloads that are highly optimized, and being able to tweak them for better performance stories starts to be something that adds serious value.

Thomas: Yeah, I mean, I won't even dispute that assessment. There's so much waste in the Cloud these days. And it starts from underutilized machines, it starts from data that's not compressed, and that just sits there; nobody's ever going to touch it again. And I won't claim that everybody will have the problem of needing to optimize their code. That's quite plainly not the case.

Our calculus is pretty much Google needed to build a system like this internally. Facebook needed to build a system like this internally. At some point, there is a SaaS business of a certain size, where you actually want to know, hey, where is all my money going, and you want to enable your engineers to make those decisions in a smarter way. So, I guess the other side is, I looked at my own skill set and try to figure out, what can I do? Where is my skill set a good match for helping people save money in the Cloud? It turns out, it's not in writing a better dashboard, and it's not necessarily in right-sizing, either. So, I guess I took what I knew how to do and figured out [laughs] how to apply it if that makes sense.

Corey: That's, I think, a great approach. The challenge I always see is in translating things that work for Google—for example—into, I guess, the public world, where I would argue that in the early days, this was one of the big stumbling blocks that GCP ran into—and I don't know how well this was ever communicated, or if it's just my own weird perception—but it feels like in a lot of ways, Google took what it learned from running its own massive global infrastructure, and then turned it into a Cloud. The problem is that when you're internal at Google and running that infrastructure, you can dictate how software has to behave in order to be supported in that environment. With customer workloads, you absolutely cannot do that in any meaningful way. So, it feels like the more Google-y your software was, the better it would run on GCP. But the more—well, let's be very direct—the more it looked like something you'd find in a bank, the less likely to find success it was. And that's changed to a large degree as the product itself has evolved, but is that completely out to sea? Is that an assessment that you would agree with? Or am I completely missing something?

Thomas: I think you're kind of right, not entirely right. I think you're completely correct in that Google internally, you need to rewrite things in a very specific manner in order for them to work. And my personal experience there is, I ran a small reverse engineering company from 2004 to 2011 and we got acquired by Google in 2011. And then we had to take an existing infrastructure that we had and port it to Google's infrastructure—to the internal infrastructure, not to the Google Cloud infrastructure because back then Google Cloud was just App Engine, which was not—I mean, it looks like ahead of its time now that everybody's talking about Function as a Service, but back then App Engine just looked weird. So, we had to rewrite pretty much our entire stack to conform to Google's internal requirements.

And it's a super weird environment because once you do rewrite it in the way that Google wants, everything scales to the sky, immediately. Like, the number of cores essentially becomes a command-line parameter. It’s like make -j16, but you replace 16 with 22,000 and you have 22,000 cores doing your work. Now, that said, GCP never externalized these internal systems, so what you get on GCP, to my great frustration, as an ex-Googler is often a not so great approximation of the internal infrastructure. [laughs].

So, it's neither here nor there, but first of all, I completely agree that Google, historically, in the Cloud has had the problem that they had learned a lot of lessons internally from scaling, and were terrible at communicating these lessons properly to the outside world, and then gave people something that wasn't well explained. App Engine is a great example for this, right? Because if you encounter App Engine for the first time in 2011, and you look at this and you’re like, “Why, how, and what is this?”

Corey: It was brilliant in some ways, don't get me wrong. But privately, I always sort of viewed Google App Engine as, “Cool, we're going to put this thing up and see who develops things super well with it, then we're going to make job offers to those people.” It felt like it was more or less a project provided by Google recruiting.

Thomas: I wouldn't know about that. And having been involved in interviewing, I think you're overestimating Google's recruiting progress. But that said, I mean, App Engine is a really interesting concept, but the average developer does not—at least, the average developer in 2011—does not appreciate what it does and why it does these things. And you can talk about Borg and Kubernetes as well, where Google just didn't explain very well why they decided to build things the way they did when they externalized similar services on GCP. And that's certainly hurt them.

If you look at the history of AWS versus Google, Google gave people something like App Engine, which was weird and strange, and for a particular use case, and not terribly well explained, and AWS gave people VMs, which people understood. And by and large, if I am going to choose a product, and one looks really strange and one looks familiar, I'm going to take the familiar product. And I think the strategic failing that Google did historically in Cloud was when they did have a technically superior solution, they did a very poor job at explaining why this is a technically superior solution. So, in some sense, Google never had a—historically Google didn't have a culture of customer interaction and, to some extent, what you need to do in Cloud is you need to reach out to people, take them by the hand, calm the nerves, and then help them walk to the Cloud, and Google just didn't do that. They gave people strange-looking things and told them, “Hey, this is the better way but we don't tell you why.”

Corey: This, of course, feels like a microcosm for Kubernetes. If I'm going to continue my bad take—and I absolutely will: when you're wrong, by all means, double down on it—it feels like Kubernetes being rolled out was an effort to get the larger ecosystem to write code and deploy it in ways that were slightly more aligned with Google's view of the world. And credit where due it worked. The entire world is going head over heels for Kubernetes.

Thomas: Yeah. Well, Kubernetes is an interesting thing to watch because I'm a huge fan of Google's internal system called Borg. And—

Corey: Oh, yes.

Thomas: —there's a philosophical view at play, right. And the philosophical view is that you really shouldn't treat a data center as a group of computers, you should treat a data center as a computer that happens to be the size of a warehouse. Urs Hölzle wrote a book called Warehouse Sized Computing and a lot of Google's internal engineering philosophy is centered around this thing that we really should be building something that treats an entire data center as if it was one computer. And that's actually a very compelling viewpoint and I think it's a very good viewpoint to take. And then they built Borg and Borg worked brilliantly internally for that purpose.

Corey: Well, there was a great tweet that you wrote back in April, “The trouble with Google's infra is sometimes you just want a slice of bread, but the only thing available is the continental saw, normally used for cutting continents and the 2000 page manual.”

Thomas: Yeah.

Corey: That's what Borg is for. Not everyone needs that level of complexity and scale to deploy, you know, a blog.

Thomas: No, that's entirely true and entirely fair. I guess if you're running a SaaS business of some size—and with some size, where, let's say, speaking about when you need 20, 30, 50 servers—you probably want to have some way of administering these and we've all been in the trenches enough to know that administering 50 machines becomes a bit of a nightmare very quickly. So, I think the view of, we really should be treating those 50 machines as if they were one big machine, that's still a good view, even if you're not Google. To be fair, if you're running a blog or—like, 99 percent of all workflows really don't need to be distributed systems. That's the reality of it.

We've had 40 years of Moore's Law. A single cell phone can do fairly amazing things if you think about the sheer computing power we have there. So, I fully agree that in the majority of cases you don't need the complexity, and good engineering usually means keeping things simple. And you do pay a price in complexity for insane scalability. And I stand by the tweet about the intercontinental saw because what happened so often when I was at Google, that they had these fantastically scalable systems, but it took a long while to wrap your head around how to even use them, and you really just wanted to cut a slice of bread.

Corey: That's one of the real problems is that scale means different things to different people, all the time. And, from my perspective, “Oh, wow, I have a blog post that got an awful lot of hits,” that might mean, I don't know, in the first 24 hours it goes up, it gets 80,000 clicks. That's great and all. Then you look at something that Google launches. It doesn't even matter what it is, but because they're Google, and they have a brand, and people want to get a look at whatever it is they launched before it gets deprecated 20 minutes later, it'll wind up getting 20 million hits in the first hour that it's up. It's a radically different sense of scale, and there's a very different model that ties into it. And understand that Google's built some amazing stuff, but none of the stuff that they've built that powers their own stuff is really designed for small scale because they don't do anything small scale.

Thomas: Oh, that's entirely true. Now, to counter that point a little bit, though, I would argue that… I mean, if there's one thing to be learned from Gangnam Style is that the strangest things can go viral these days, and you may find yourself in a position where you need to scale rapidly within 24 hours or even shorter timeframes. If whatever you're offering gets to be insanely popular and spreads through social media. Because the reality is, like, in the year 2000, if your software got popular, you noticed that it was sold out in stores, right? You could produce more, and that was fine.

But the reality of today is, you may be in a situation where you need to scale really rapidly, and then it may be good if you've built things in a way that they can scale. Now, I'm not advocating that you should always pay the complexity price of making things scalable. I'm just saying that scalability may be more important today from a business perspective than it was a couple of years ago, just because, especially with Software as a Service, people switch things around a lot, people try things out, and it's quite possible that just randomly you get 10 million hits on your service the first day, and then you probably don't want to show the famous fail whale that Twitter was so famous for.

This episode is sponsored by our friends at New Relic. Look, you’ve got a complex architecture because they’re all complicated. Monitoring it takes a dozen different tools. Troubleshooting means jumping between all those dashboards and various silos. New Relic wants to change that, and they’re doing the right things. They’re giving you one user and a hundred gigabytes a month, completely free. Take the time to check them out at newrelic.com, where they’ve done away with almost everything that we used to hate about New Relic. Once again, that’s newrelic.com.

Corey: Oh, yeah. And remember that there's an argument to be made for reliability and when it begins to make serious business sense. You can wind up refactoring your existing code that has no customers until you run out of money, but even bad code that doesn't scale super well—like the Twitter fail whale—can get you to a point where you can afford a team of incredibly gifted people that come in and fix your problems for you. There's validity there, but early optimization becomes a problem. The things that I would write if I'm trying to target 20 million active users versus half an active user at any given point in time—because who really pays attention to me?—would be a very different architectural pattern for the most part.

Thomas: Yeah. I don't disagree. Then again, I think the entire selling point of App Engine back in the days was, you just write this thing in the way that App Engine tells you to do and then, whatever happens, you're insured, right? But yeah, I fully agree. I mean, there's the argument to be made that once you have traction, you also have money to fix the scaling issues, but then the question becomes, can you fix those quickly enough, so people don't get turned off by the unreliability? And that's not a question I can—anybody can answer in any good way upfront because you have to try things, and they'll fail and so forth.

Corey: So, what's interesting to me is that you don't come from a cost optimization background, historically. In fact, you come from one of the more interesting things on the internet—which is fascinating to me, at least—which is Google's Project Zero. And for those who haven't heard of it, what is Project Zero?

Thomas: So, Project Zero is a Google internal team that tries to emulate government attackers, essentially, trying to find vulnerabilities in critical software and by emulating the thought process, tries to nudge the industry in the direction of… well, making better security decisions and fixing the glaring issues. And it arose from Google's experience in 2009 where the Chinese government attacked them and used a bunch of vulnerabilities. And then Google at some point, a couple years later decided, “Hey, we've got all these people on the offensive side being paid to find vulnerabilities and then sell them to governments to hack everybody. Why don't we start a team internally that tries to do the same thing, but then publishes all the techniques, and publishes all the learnings, and so forth, so that the industry can be better informed?” In some sense, the observation was that the defensive side often made poor decisions by being not well-informed about how attacks actually work. And if you don't really understand how a modern attack works, you may misapply your resources. So, the thought process was, let's shine a light on how these things work so the defensive side can make better decisions.

Corey: One of the things I find neat about it—this is of course where Tavis Ormandy works—

Thomas: Yes.

Corey: —and it's fun talking to him on Twitter watching him do these various things. Every time he's like, “Hey, can someone at some random company reach out to me on the security contact side?” It's, “Ooh, this is going to be good.” And everyone likes to gather around because it's one of those rare moments where you get to rubberneck at a car wreck before the accident.

Thomas: Yeah, Tavis is a—I have an extremely high opinion of Tavis. He's a person with great personal integrity, and he's a lot of fun to discuss with, and he’s got a really good intuition where things break. So, in general, the entire experience of having worked at Project Zero was pretty great. I spent a grand total of eight years at Google, five of which in a team that did some malware related stuff, and then two years in Project Zero. And the two years in Project Zero were certainly a fantastic experience.

Corey: The thing that I find most interesting is that you have these almost celebrity bug hunters, for lack of a better term, and what amazes me is how many people freaking seem to hate him. And you do a little digging and, oh, you work at a company that had a massive vulnerability that was disclosed and, hm, one wonders why you have this axe to grind. It's, again… in some levels, it’s people doing you a favor. I've never fully understood aspects of blaming people who point out your vulnerabilities to you in a responsible way? Sure, I know you would prefer that they tell you and never tell anyone else, and you owe them maybe a T-shirt at most. Some of us aren't quite that, I guess, willing to accept that price point for our skill sets.

Thomas: Yeah, so the entire vulnerability disclosure debate is a very complicated and deep one, and it also goes in circles over decades. And it's actually quite tiring after 20 years of going through the same cycle; it feels like Groundhog Day. But my personal view is that to some extent, the software industry incurs risks on behalf of their users in order to make a profit, meaning you gather user data, you store it somewhere, and you can, well move fast and break things, and nothing much will happen if that user data gets leaked, and so forth. So, the incentives in the software industry are usually towards more complexity, more features, and bad security architecture.

And because there's no—I mean, there's no software liability, there's no recourse for wider society against the risks that the software industry takes on behalf of the users, the only thing that may happen is that you get an egg on your face because somebody finds a really embarrassing vulnerability and then writes a blog about it. So, in some sense, Project Zero and the people that work at Project Zero, they wouldn't be doing their job if everybody loved them because to some extent, their job is to be an incentive to actually care. If people say, “Oh, let's do a proper security architecture, otherwise Tavis would tweet at us.” That's at least some incentive [laughs] to have security. It sounds a bit sad, but this is the only thing that—not the only thing, but this is a thing that is necessary. But part of the job of being a Project Zero researcher is not to be everybody's best friend if that makes some sense.

Corey: Yeah. Security is always a weird argument. I started my career dabbling in it and got out of it because frankly, the InfoSec community is a toxic shithole. Yes, I did say that; you did not mishear if you're listening to this and take exception to what I just said. I said what I said.

It's such an off-putting community where it was very clear that the folks who are new to the field were not welcome, so I found places to go where learning how this stuff works was met with encouragement rather than derision. That may have changed since I was in the space. It's been, what, nearly 15 years, but I'm not so sure about that.

Thomas: So, I wouldn't know, right? Because I grew into that community in Europe 20 years ago. And the community I grew into 20 years ago in Europe was a very different community from the community I encountered when I first came to the US and interacted with the US InfoSec community. And also, you tolerate a lot of behavior when you're 16 and you want to be part of a community that you wouldn't tolerate as an adult. So, I'm not sure whether I would have, like, a very clear view on these topics, right?

Because the other thing is, once you reach some level of status, everybody's incredibly nice to you all the time. And at least my experience in security after I turned 18 or 19, was that people were by and large, more friendly than justified to me. Now, that doesn't mean they weren't shitty to everybody else at the same time, right? So, the reality is that I have a skewed view of the security community because I got really lucky if that makes any sense. And then also, I guess I'm kind of picky about who I surround myself with, so the two dozen or so people that are really like, out of the greater security community may just not have that same culture if that makes any sense.

Corey: So, help me understand your personal journey on this. You went from focusing on InfoSec to cloud cost optimization. I have my own thoughts on that, and personally, I think that they're definitely aligned from the right point of view. But I'm curious to hear your story. How did you go from where you were to where you are?

Thomas: Yeah. So, we can call it, perhaps, a midlife crisis of sorts, but the background is, after 20 years of security, you realize security is always, at some level, about the human conflict. It’s always—you do security for somebody against somebody, in some sense; you're securing something against somebody else. And it's very—well, I wouldn't say repetitive, but it's certainly a very difficult job, and at some point, I asked myself, “Why am I doing this? And for what am I doing this?”

And I realized, hey, perhaps I want to do something that has a positive externality. Like, I don't necessarily want to participate in human to human conflict all the time. And I realized that my only credible chance of dying with a negative CO2 balance [laughs], or budget would be to help people compute more efficiently, right? Like, there's no amount of no meat-eating, no car driving that I can do that will erase all the CO2 I've emitted so far, but if I can help people compute more efficiently, then I can actually have a positive impact. There's a triple win to be had: if I do my work well, the customer saves money, I earn some money, and in the meantime, I reduced human wastefulness. So, that had a great appeal on a philosophical basis.

And then, the other thing I realized is that when you do security work, a lot of your work is reading existing legacy code and finding problems, and then have people mad at you because you found the problems in the legacy code, and now they need to be fixed. And it turns out that when it comes to optimization, the workflow is surprisingly similar. The skills you need in terms of lower-level machine stuff, and so forth, it's also surprisingly similar. And if you find the thing to optimize, anything to fix, then people are actually thankful because you are saving money and make things faster. So, it turned out that this was pretty much a match that worked out surprisingly well. And yeah, then that's how I made the jump.

And I have to admit, so far, I really haven't regretted it. The technical problems are super fascinating. To some extent, there's less politics, even, in the cost optimization area because one of the issues with security is, on the defensive side, a lot of good security work is about convincing an organization to change the way they're doing things. So, a lot of good defensive work is actually political in nature. And the purely technical geeks in security are often in some form of offensive role.

For example, the Project Zero stuff, that's offensive for the defensive side, but still, it's a sort of offensive role. And then the majority of these jobs are just in companies that sell exploits to governments. And given that my forte happens to be more on the technical side than on the influencing an entire organization side, I decided that—and given that I didn't want to do offensive work in that sense anymore—I decided that this entire cost optimization thing has the beautiful property of aligning good technical work with an actual business case.

Corey: There's an awful lot of value there to aligning whatever you're doing with business case, I would argue that security and cost optimization are absolutely aligned from a basis of cloud governance. Of course, now, here in reality, don't call it that because no one wants to deal with governance, and it always means something awful, just from a different axis, depending upon who you talk to. But that's the painful part is that there's no great answer around how to solve for these problems. What always confuses me—and annoys me on some level—is when I have a cost project that accidentally turns into a security project, where it’s, “So, tell me about those instances running in that region on the other side of the world.” “Oh, we don't have anything there.” “I believe you are being sincere when you say that. However, the bill doesn't lie.” And suddenly we're in the middle of an incident.

Thomas: It's funny that you mentioned this because the number of security incidents that have been uncovered by billing discrepancies is large. If you go back to Cuckoo's Egg, like Clifford Stoll’s story about finding a bunch of KGB finance german hackers in the DOD networks, that was initially triggered by an accounting discrepancy of I think, 25 cents or something like this. So, yeah, the interesting thing about IT security is, for banking, for example, is if you steal data, nobody normally knows because normally data isn't accounted for properly, except when you cause large data transfer fees because you're exfiltrating too much data out of AWS.

Corey: Yeah. That's always fun when that happens. What’s surprising to me, and that makes perfect sense in hindsight, if you have a $75 AWS account every month and suddenly you get a $60,000 bill, you sort of notice that. But if you wind up getting compromised when you're spending, let's say, $10 million a month, it takes an awful lot of bitcoin mining before that even begins to make a dent in the bill. At some point, it just disappears into the background noise.

Thomas: Oh, yeah, definitely. But I guess that's always the case. If you look at a supermarket, they don't notice half of the shoplifting, right?

Corey: Yeah. Supposedly, anyway. I don't know. I tend to not spend most of my time shoplifting. I usually set my eyes on bigger game, you know, by exfiltrating [00:30:37 unintelligible] data from people's open S3 buckets.

Thomas: Isn't even exfiltrating of S3 buckets old? [laughs].

Corey: No, I've decided that people don't respond to polite notes about those things. Instead, I just copy a whole bunch of data into those open buckets on the theory that while they might ignore my polite note, they probably won't ignore a $4 million dollar bill surprise.

Thomas: That's actually a fairly effective-sounding strategy.

Corey: It's funny—let's be very clear here. I'm almost certain that that could be construed by an aggressive attorney as a felony. And let's not kid ourselves; if you cost a company $4 million their attorneys will always be aggressive. This is not legal advice. Don't touch things you don't own. Please consult someone who knows what they're doing. It's not me. Have I successfully disclaimed enough responsibility? Probably not, but we're going to roll with it.

Thomas: All right.

Corey: So, if people want to hear more about what you're up to, how your journey is progressing, or hear your wry but incredibly astute observations on this ridiculous industry in which we find ourselves, where can they find you?

Thomas: Well, one option is clearly on Twitter. I run a Twitter account under twitter.com/halvarflake. H-A-L-V-A-R-F-L-A-K-E. And that is not only about my professional work, I do have a fairly unfiltered Twitter account. Like, there's nobody ghostwriting my tweets, and there's oftentimes things that I tweet that I regret a day later. But that's the nature of Twitter, I guess.

Corey: All of my tweets are ghostwritten for me. Well, not all of them. Which ones specifically? The ones that you don't like. That's right.

Thomas: [laughs].

Corey: That's called plausible deniability.

Thomas: [laughs]. So, yeah, and if you care about questions, like, “I’ve got 50,000 machines, and I would like to know which lines of code are eating how many of my cores?” Then it's probably a good idea to head over to optimyze.cloud. Remember, optimyze, M-Y-Z-E at the end to make spelling more fun. And sign up for our newsletter.

Corey: You're disrupting the spelling of common words.

Thomas: I'm sorry for that, but the regular domains were too expensive, and trademarks are really hard to get.

Corey: They really are. Well, thank you so much for taking the time to speak with me today. I really do appreciate it.

Thomas: Thank you very much for having me.

Corey: Thomas Dullien, CEO of optimyze.cloud. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple podcasts, whereas if you hated this podcast, please leave a five-star review on Apple podcasts anyway, and tell me why this should be deprecated as a show, along with what division of Google you work in.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Jam LeomiJam Leomi is a penmaker who just so happens to computer. When not found ranting on equality and equity in #infosec and beyond on twitter, they're found doing their day job as Lead Security Engineer at Honeycomb.

Links Referenced

  • Honeycomb
  • Jam's Personal Blog
  • Follow Jam on Twitter
  • Connect with Jam on LinkedIn

TranscriptAnnouncer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

This episode is sponsored by our friends at New Relic. Look, you’ve got a complex architecture because they’re all complicated. Monitoring it takes a dozen different tools. Troubleshooting means jumping between all those dashboards and various silos. New Relic wants to change that, and they’re doing the right things. They’re giving you one user and a hundred gigabytes a month, completely free. Take the time to check them out at newrelic.com, where they’ve done away with almost everything that we used to hate about New Relic. Once again, that’s newrelic.com.

Corey: This episode has been sponsored in part by our friends at Veeam. Are you tired of juggling the cost of AWS backups and recovery with your SLAs? Quit the circus act and check out Veeam. Their AWS backup and recovery solution is made to save you money—not that that’s the primary goal, mind you—while also protecting your data properly. They’re letting you protect 10 instances for free with no time limits, so test it out now. You can even find them on the AWS Marketplace at snark.cloud/backitup. Wait? Did I just endorse something on the AWS Marketplace? Wonder of wonders, I did. Look, you don’t care about backups, you care about restores, and despite the fact that multi-cloud is a dumb strategy, it’s also a realistic reality, so make sure that you’re backing up data from everywhere with a single unified point of view. Check them out as snark.cloud/backitup.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Jam Leoni, lead security engineer at Honeycomb. Jam, welcome to the show.

Jam: Thank you so much, Corey.

Corey: So, I want to start by thanking you for taking over the guest authorship of the newsletter, well, this week when this gets aired. But at the time we're recording this, you haven't actually done it yet. So, it's great to sit here and thank you now for a thing that you haven't done yet, but everyone listening to will have already been aware of because yeah, time is weird and is no longer linear, like anything else in 2020.

Jam: I mean, what is time?

Corey: Exactly. It's such a good year so far, and it's just getting better all the time. Ah. So, let's begin at a high level, I suppose. Who are you? What do you do? What's your story?

Jam: So, what am I? I am a Black, genderqueer security, turned ops, turned security engineer. [laugh].

Corey: Interesting as far as the ops turned a security engineer. Very often it feels like the common path in tech is the opposite direction, where it's, oh, I'm going to do security. And then you know what, it turns out that I don't like aspects of the job, and people want to often broaden out into other arenas. So, it feels, to me at least historically, that most folks go security to ops. Counterpoint, that's also the path I took, so there's a heck of a selection bias here.

Jam: Yeah, but, well, also for me, like, you have to realize that most security people don't actually start off in security. I think I'm the exception because my degree was in security, but I couldn't find a job in security. So, I had to do something else.

Corey: There's an almost unfortunate tendency that I see a lot, which is that people who get a degree in something believe—because they are told. Let's not mince words about that—that, oh, what your degree in is now going to define what you do in terms of your career; it's going to set your trajectory. So, then we are basically asking a bunch of 18-year-olds, in the common case to, yeah, figure out what you want to do with your life. I'm damn near 40 and I don't know what I want to do with my life, yet. So, it feels like it tricks people into believing they have to go a certain way.

Jam: Oh, but I was smart about it, though. So, the reason I went into security was specifically because the security degree program that I was a part of taught me so many things. So, I was like, well, if security doesn't work out, I can always jump into something else in technology like I went in as an 18-year-old thinking about this.

Corey: That must have been nice. I was never an academic. I am, effectively, someone who graduated from high school through an unaccredited organization, so on paper, I have an eighth-grade education. And this gets back, on some level, to what you said in the beginning with your introduction, that you are a Black genderqueer ops turned security engineer, I am a cishet, white guy, and society is built in such a way that it takes people like me and always picks us up and dust us off whenever we stumble as if I'm somehow entitled to take up as much space as I possibly want. And that's a travesty because on some level, don't you kind of want everyone to be able to have that freedom to experiment, freedom to fail alternate paths to success rather than prescribed ones? It just—I'm sorry, there's so much about society and the way it is structured that I do not pretend to understand.

Jam: Yeah, I'm in the same boat with you, and I wish I had answers to that. There are definite answers, you know, when we talk about white male privilege, and the patriarchy, and things like that, but I'd rather focus on—while there definitely is pain in the industry from being a Black genderqueer female-presenting person—especially female-presenting—I think it's also prepared me to always try to seek options. And it's a blessing and a curse for me to be always kind of thinking ahead and risk-managing.

Corey: And this is almost certainly made no easier whatsoever now by the fact that it is 2020 and so much has changed in a relatively short period of time. I feel like if there were a time warp, and one of these podcast episodes slipped through to just a year ago, “What the hell are you talking about?” And here we are now, working in a time of COVID, where everyone is more or less trapped at home, except to be honest, some of the worst people in the world. And everything has changed. It's sort of a weird segue here, but it's a weird year, during the course of this entire event, you've started working at a new job, you have effectively—as had the rest of us—tried to find a new normal. What has your experience of working in tech been like during these changes? How has COVID life changed our industry?

Jam: I think COVID life has changed… it's changed our industry, but I don't know how it's going to suss out. Right now, there's the obvious thing of pretty much we have everybody working from home right now. So, that can be kind of an isolating change. Some people are working with children, which, that can be also a change and a change of dynamic of managing energy.

But for me personally, it's been, it's been slightly isolating. Usually, I'm used to working remotely, but I'm used to being able to go to the cafe and work there or go to a co-working space. And so now I'm having to think up ways to manage the normal ways I get connection through different avenues, whether it's making phone calls, making more use of Zoom calls, or just trying to—in the most safest way possible because, you know, your sanity is needed—try to find ways to socially distance and socialize as well.

Corey: It really has turned an awful lot on its head. What I'm trying to figure out is, at some point, a new normal is going to hit, regardless of what that looks like, and we're not going to be in pandemic stages forever. What's going to go back to the way it was, and what's going to be, I guess, forever changed.

Jam: Well, my hope is that the pool for employment in tech and the equity in tech changes to not just be on the coasts, and be more across the boards available to everyone. Now, that could have negative consequences. That could mean less people in cities, in more rural areas. Though, I think rural areas should have a shot, too. Like, I come from Kentucky, which, like many states, I came from a city but I was also surrounded by rural communities and I grew up around—have friends from rural areas.

So, I want them to have a shot just as much as I do. And many of the times they don't, or you have to tell these people, “Hey, we want you to move halfway across the country away from your support systems, family, friends, and just get adjusted and come work for us.” And I feel like people shouldn't have to make that decision in order to make a living or in order to follow their dreams or their heart.

Corey: That is one of the most aspirational answers to that question because usually, people tend to go in a direction of, “Well, I don't know if tech conferences are going to have quite as much swag anymore.” And you're talking about making tech entirely more accessible to folks who, for example, might not live in eight square miles of an earthquake zone and calling it disruption. There's something to be said that is incredibly valuable for finding folks who have gone through alternate paths to get to where they are now. It winds up providing a diversity of experience. And that is incredibly valuable. As it turns out, maybe the sum totality of human existence isn't embodied by a bunch of people who went to Stanford together. Just as a random shot in the dark.

Jam: Yeah. [laugh]. Again, this is like in terms of there's many conversations in making tech more equitable, and making businesses more equitable, and it just can't be in specific places.

Corey: Yeah, I'd say one of the saddest days for San Francisco was when you moved away. To give a little insight back into, I guess, how long we've been talking to each other, I remember back in 2016, when I was running a DevOps team at my last, quote-unquote, “Real job,” and you interviewed for a role. And you were a terrific candidate; we extended an offer, and you were on the fence between this role and another job. And you and I went out and sat down and talked for a while, and at the end of it, my recommendation—that I stand by—was that you take the other offer. And you did. And when I tell people versions of that story, there's always two different responses. One is, “Oh, yeah. Of course, that's what you do.” And the other response is, “Wait, you did what?” And I've never understood people who take the second perspective.

Jam: Yeah, simply because, especially in this day and age, we're no longer in an era where somebody stays at their job until they retire at the age of 50 or 60. You are constantly—especially in the startup world—you're constantly moving jobs because, for some people, especially who look like me, you have to do that in order to get ahead. So, on that same vein, tech is small. People know each other. So, it behooves you, especially as a manager and a leader, to have good relations, even if the choice is not going to be beneficial to you.

Corey: It's funny that you mentioned this idea of being more mobile in our careers and life, extending beyond the next job, but it's amazing how everyone loves to pretend that in job interview stories, that oh, yeah, now the average tenure in tech is, of course, 18 to 24 months, but once you start here, you're going to work here for 25 years, you're going to wind up leaving with a pocket watch and a pension, and it's this ludicrous fantasy. We even take it a step further, where the stories we tell about someone leaving unexpectedly are always, they get hit by a bus, the bus factor. Great. How about, someone else offers them a 30 percent raise somewhere else that's aligned more directly with what they want to be doing? Because, spoiler, I've had a lot more colleagues leave because of a better offer than ever got hit by a bus.

Jam: Yeah. I'm also trying to wonder who are you talking to that are offering pocket watches, and, like, lifetime—because I've never heard of such a company.

Corey: Oh, yeah. They were all over the place in the 1960s, which is where a lot of that interview advice seems to have come from. It's, “Oh, you want to go ahead and get a job? It's easy. Just walk in the front door, have a firm handshake, and ask for a job. Be sure to call the boss, ‘sir.’” “What if the boss isn't a man?” “I have no idea what you're talking about.” It's this old-timey advice where you expect every video of it to be in crackly audio with black and white, where it's this ancient 1940s approach? Ugh, no, thank you. Yeah, pensions are also hilarious fantasies that are gone.

Jam: Yeah. And just, I feel like a lot of the humanity—not to say that there was any humanity back in the 1940s; there was probably some for some people more than others—but I feel like even more so now there's, kind of, less of it. Because you're talking to a person who's been through two recessions already, and also been through the Enron scandals, as well. As well as many other scandals related to misuse of funds.

Corey: Oh, yeah, I'd love the idea of, “Oh, just put your entire retirement in your company stock, it'll be fine.” What I always love is finding people who give that type of advice and talk about how, “Oh, I work at Google,” or, “I work at Amazon,” and, “Oh, all of my retirement is invested in my company's stock.” And well, that seems to be centralizing an awful lot of risk on that company doing well. And they'll come back with a whole suite of answers about this.

And I see where they're coming from. They're arguing good faith. The counterpoint is that everything that they're saying, without exception, could have been said by an Enron employee right before the collapse. Now, for legal and moral reasons, I do want to point out that I'm not insinuating that Google, and Amazon, and the rest are fraudulently lying to everyone, that there falsifying audit information, et cetera. My point is not that they're engaged in malfeasance, but rather that you never know what the future is going to hold for any given company, and nothing lasts forever so decentralized risk.

Jam: Yeah.

Corey: But oh, does that rub some people the wrong way.

Jam: Yeah. And it's also just like, just like you want to make sure you're very diverse about your company, you kind of want to be diverse about your stocks. Don't have all your things in one pot. Or your investments in one pot. That's something I've even got from my financial advisor.

Corey: Oh, yeah. I should probably disclose this, I don't do it quite often enough that everything I own in equities is part of a broad-based index fund. The single exception to that, I own six shares of Amazon stock that I've held for years and will continue to hold indefinitely, not because I view this as a long term financial play, but rather one day I will manage to shitpost via shareholder resolution. Wait for it, it's going to be amazing. I just need to find the joke worth doing it for.

Jam: That is so great. [laugh].

Corey: We all need stretch goals, and that's one of mine because I make terrible life choices.

Jam: Oh, no, you don't.

Corey: So, tell me a little bit more about your path. You're one of those folks that I get to catch up with from time to time, and I love every chance that we get to sync, but it always seems like there's a lot to catch up on. Where did you first enter tech? And where did you go from there?

Jam: So, I first entered tech—it's funny I entered—if we want to say when I entered tech, it was when I was probably about 12 years old. I joined a computer club at school. It was a program that I was a part of until I graduated from high school called the Student Technology Leadership Program. And from there, I learned a whole lot about computers. This is back in the day when computers were starting to become a thing.

And I really started the journey of doing more technical work when I was in high school, and they were taking computers apart at the high school that I wanted to go to, and I was like, “I want to do that.” And so that kind of started me on my journey of doing more deep dives into technical things.

Corey: One of the, I guess, strange things that I found is that when I talk to folks who've been in the space for a while, they always come from something into a new area, and then we have conversations around these things. But an awful lot of us were old school Unix types or Linux folks very early on with the sysadmin ops story. And those jobs, for better or worse seem to be drying up as more and more things move in a cloud-y type direction, so I find myself spending an awful lot of time wondering and having conversations about the topic. Where does the next generation come from?

Where does the next series of cloud folks wind up originating from? Because the terrible answer to this is, “We're just going to wait until the cloud providers start sponsoring public school curricula. And then they're going to start teaching eighth-graders how to wind up spinning up Elastic Beanstalk,” or God knows what. And I don't think that's the answer anyone wants and I'm hoping people are better ones.

Jam: I don't think it's going to come from the children’s. I think it's coming from the people who are entering the industry from other places. Like, one thing I kind of have an issue with is so many of these big companies are like, “Well, we don't have a pipeline, so we're just going to push it to the children, push it to the children, push it to the children.” Meanwhile, I'm seeing so many people being like, “Man, tech is paying some money so I'm going to transition into that.”

And you have so many of these people either transitioning into support roles or transitioning from support roles and trying to get higher up from different industries in the past five years. And so I think those people are going to tell us what is next. I don't think we have to wait for the kiddos to get 10 years in and be the next generation. I think we already have some of those people here, and I think they're going to push the needle on, tell us what's next.

Corey: I sure hope that you're right. There's a definite hope that I have that this is going to turn into something that's, I guess, lasting and transformational. And I don't like the idea that oh, so the only way to now get into this space is to stop doing whatever you were doing before, whatever it might be, and then go to a boot camp, possibly a boot camp then winds up doing an income repayment and they’ll send you straight to collections if you're unable to pay. And almost these predatory for-profit institutions that tend to not, I guess, really be focused on outcomes other than making money for investors.

And I worry that there's going to become this, I guess, artificial gatekeeping story where you need to either have a degree or go to a boot camp. For someone who was able to talk their way past not having either of those things because well, honestly, look at me, I'm incredibly over-represented in this space, that path is not available to everyone. And it makes the existing biases that we have in this space worse, not better.

Jam: Yeah. And I feel like the boot camp thing is kind of changing because what I am seeing, and this is something I saw a few years ago, you're starting to see people who have degrees going into boot camps. And I think universities are starting to notice because these universities are now trying to create boot camps, as well. I don't know whether it is to get in on that money, or whether it's trying to do some more career extensions to their already vast portfolio, but I think that's something that's kind of helpful, too, especially as the traditional idea of degrees, especially in the land of COVID is going to go in a completely different direction.

Corey: This episode is sponsored in part by our friends at Linode. You might be familiar with Linode; I mean, they've been around for almost 20 years. They offer Cloud in a way that makes sense rather than a way that is actively ridiculous by trying to throw everything at a wall and see what sticks. Their pricing winds up being a lot more transparent—not to mention lower—their performance kicks the crap out of most other things in this space, and—my personal favorite—whenever you call them for support, you'll get a human who's empowered to fix whatever it is that's giving you trouble. Visit linode.com/screaminginthecloud to learn more. That's linode.com/screaminginthecloud.

Corey: I want to be very clear because I've been unclear on this in the past, I am not in any way, shape or form saying that a degree does not hold value, that if you have a degree you've made a poor decision or even that degrees are not absolutely necessary for some roles. What my position is—and remains—is that it's not going to work for everyone, and having a prescribed path for many roles that artificially requires a degree is not doing anyone any particular service. Now, if I'm going to hire an attorney, or I need an anesthesiologist, yeah, I have some degree requirements for those people. That is not really the type of role that lends itself to, I'll figure it out as I go. How hard could it be?

Jam: Yeah, I think the only reason that I myself have a degree is that, as a Black person, that is the only way that I can get my way into the door. Or at least that was the only way I could get my way into the door 10 years ago. I think boot camps are slowly changing that to give people the experience and the street cred to do that. Do I want it to go away?

I hope so someday, and I hope that we can get back to a way of having people do more apprenticeships, kind of do the old school, old school way of having people try out jobs and learn skills. But until we get to that point—because again, we're still trying to think about more equitable ways, and unfortunately, the people making the decisions, the gatekeepers, do not look like me. [laugh]. Until that changes, we're working with what we have.

Corey: One of the best descriptions that I've ever heard for helping break down those gates and making things more accessible comes from Stephen O’Grady over at RedMonk, and it's, “Send the elevator back down.” That mindset is how I try to live my life. I mean, the reason that I have a career at all is that people who had no requirement to do so did favors for me when they didn't have to. And you can't ever repay that, you can only ever pay it forward. And I try—mightily sometimes with mixed success—to wind up doing that. And I hope I get it right more than I get it wrong. But what I don't understand is people with the attitude of well, “Screw you, I got mine.”

Jam: Yeah, I don't get those people either. But our industry is kind of saturated with that. But at the same time, it's slowly changing from the past 10 years when I felt like I saw that. There's still like beacons of people, like Jennifer Davis, who helped mentor me and was a sponsor for me, as well as other people in the industry who I feel like have kept me on a good path, especially in security, like [00:24:28 Kirstin Breaker]. I absolutely love her and she's one of my favorite Black female security people and I admire her so much. But just to be able to talk with those people and really get their wisdom and stuff is super helpful. So, I'm glad for her. As well as you, Corey.

Corey: Oh, please, I did a remarkably small amount of work until somewhat recently at any of this. I'm learning as I go, like anyone else. It's one of those looking back moments where it's, “Huh, I could have done a lot more than I did and I feel bad about it.” All you can really do, unless you have somehow the ability to change the past, is do better moving forward. I think that's something that people often give up on where it's, “Oh, I didn't do such a good job in the past. Well, too late to fix it now. Oh, well.” And nothing ever changes.

Jam: Yeah. And I think that's the thing about time—which has no meaning in 2020—there's always hope for moving forward and always changing things. And I think people don't give that opportunity right now so much. I've seen a lot of intense stuff on InfoSec Twitter and I wish there was more kindness in the accountability. I can understand why some reasons why you can't have the kindness because there's so many people who are hurt, and there's so much trauma everywhere. I'm a person with PTSD, so I understand how triggering it can be. And I have hope that it can change to a place where you can have more empathy.

Corey: I sure hope so. One of my greatest fears is that when we look back at this recording, in a few years, we don't look at this through a lens of, “Oh yeah. That was a dark time in our history.” But instead, “Oh yeah, those were the good old days.” That's what scares the hell out of me.

Jam: Oh…

Corey: “Oh, look how naive we were. We didn't even know about the comet yet.”

Jam: Corey, don’t, like [knocks]—don’t do that.

Corey: Don't put that juju [00:26:28 crosstalk] [laugh].

Jam: I’m like knocking on wood. Like, come on, man [laugh].

Corey: So, back when you were applying for your current job, what were you looking for? What was it that mattered to you from a, I want to work with these people, or that company or that technology perspective.

Jam: So, for me, it was actually funny because when I first decided to take a break after my last job, my plan was okay, I'm going to take a break, and then I'm going to come back out and I'm either going to see if I can work for a VC firm, [laugh] and see if I can do, like, security advising for them, because I wanted to do more of a leadership role in that, or I wanted to do some consulting. And at first, I did actually look into that. And one of the final companies that I worked for was a consulting firm. It was between a consulting firm and the place that I work now, but the reason why my current job worked out is for two reasons.

One, I always love working at places with cool products. And Honeycomb had a really, really cool product and idea that I wanted to dive more into. And the second thing is that I love working with cool people, and Honeycomb had all the cool people. I really admire all the people who work there. I admire all of my coworkers; they're awesome people, and I think that is what attracted me there. On top of the fact that there was the third thing, which is there was the opportunity for leadership experience in security, and growth there, which I don't think I would have gotten in consulting.

Corey: Wholeheartedly agree. Honeycomb is a fantastic company, let's not kid ourselves here. And, of course in the interest of full disclosure, they've been a good recurring sponsor for a lot of my nonsense, but they're also a reference client for my consulting business. So, even if I didn't like all of you, folks, I think at this point, I'm contractually obligated to lie about it. I kid. I love what you folks do. I think there's a tremendous value to the industry across the board in about four different axes, and it's hard for me to think offhand of a company with a better internal culture.

Jam: Yeah. That is super, super true. I wish I could digress into it, but I can’t.

Corey: No, I completely understand that.

Jam: But yeah. Since I've joined, I feel like I've really been able to make a mark, and some impact, and really just challenged myself in new and different ways. So, I'm excited to see where my career goes from here.

Corey: I am, too. I look forward to our next recording where we wind up catching up on, “Oh, and here's the changes since the last time.” So, talk to me a little bit about why a company like Honeycomb, who does observability and/or yelling at people for saying, “Don't deploy on Fridays,” depending on your taste, hires a lead security engineer. Judging by everything else I see in the industry, security is this thing you bolt on after the fact and apologize for while saying how much of a priority it was, even though it clearly wasn’t. How does an observability company need a security engineer?

Jam: Well, here's the thing about technology right now. In the past 10 years, it has changed in that most startups need to—side note. This is my personal opinion and not the opinion of my company. End side note—but for a whole lot of startups that I've seen, a lot of them are selling to enterprise customers. And enterprise customers have that requirement called compliance, and they have certain compliance standards that they need in order to have you as a vendor.

And so we're starting to see more companies who are trying to market to these big-money enterprise customers, and they are needing security people to get the work done because it is becoming a thing where at some point, you just can't bolt it on. Like, you have to have a security person in the room doing the work and telling you, “Okay, maybe you should do this differently so we can stay secure instead of doing the very, very security risky thing,” for lack of a better term.

Corey: Increasingly, it seems like the security risky thing is not hiring security folks. And I guess my problem with cloud security, and I can very rarely bring this up on the podcast when I'm talking to folks who work for one of the cloud vendors, is that they take a simple concept, such as the idea of the services themselves are basically secure 99 times out of 100—or more--any mistake is going to be something you have misconfigured. But rather than saying that sentiment that fits in a tweet, instead, they call it the shared responsibility model and then they turn it into this 500-word article at an absolute minimum, and an incredibly complicated slide, and it makes people miss the point. Is that just me having no attention span whatsoever, or does it feel like they're overly complicating a relatively basic concept?

Jam: I think they're overcomplicating a very, very simple concept, and at the same time, they don't want to be held liable, which I can understand that.

Corey: Yeah, good point. I mean, at some point, you have deniability, and you want to be able to point at something larger and complex when you're getting yelled at by one of your customers for their own misconfiguration that goes beyond, “It's your fault.” That is not a helpful sign to point at when someone is screaming at you, as it turns out.

Jam: Yeah. But at the same time, I do wish that the industry would make things more usable. [laugh].

Corey: Oh, my god, yes.

Jam: Even for beyond—like, one of the things I like about having a more holistic security practice is that I do want to try to be more DevOps-y with it; I do want to try to be more collaborative and not just let security people in, but for other people, for developers and other stakeholders of the business to understand. And sometimes it's super hard to make people understand if they can't see what's in front of them, and so much of the tooling that we've had thus far, have tried to inch closer and closer to it, and I'm starting to see some new players in the game to make that more usable, but for some of the bigger providers, it is still like, “Man, what are you doing?” Like, this is such a big space, and you have so much money. You could do some acquisition that's super cool. Like I've seen and been a part of companies with some products of being like, for lack of a better word, Amazon could buy you and stretch their security game so well. But instead, I have to do stuff where security is just basically unusable by even security people. Like one example is, and I hope it's changed in the future, is Cognito. I've had so many tussles with Cognito.

Corey: Oh, don't get me started on that. I really, really hope that by the time this episode gets published, Cognito is better than it is right now, but, mm, today, it seems almost like it is an incredibly well-executed advertisement for Auth0.

Jam: Yeah, Auth0, or just some of the other ones available, too. But it is… you just want it to be that because it is an integrated service. But it doesn't do some of the things that you imagine it would do. But it's also like, it's Amazon. So, it's Amazon, it's free, and all the security shouldn't be up to us, which is a great thing. And at the same time, yeah, I just wish it were better.

Corey: One other thing that I've never fully understood, the most depressing InfoSec experiences I’ve had was wandering around the RSA expo floor. And, first, I don't think you're allowed to sell anything, legally, if you don't have the word ‘firewall’ somewhere in it, and two, I understand that security is not something you can buy, but holy crap, do a lot of companies want to sell it to me. What's the deal there?

Jam: I think it's that people know that security is always going to be a need and it's ever-encompassing and ever-growing. The thing is, just as people have many different ways of engineering, there are also many different ways that people do security because everybody needs it, but nobody knows the specific security that they want or need. So, I think the issue that we run into right now, is that because we don't have somebody telling us what the best—or they're telling us what the best should be, people tend to get stuck in their tooling, and don't realize until it's too late that it doesn't work for them.

Corey: Yeah, on some level, you sort of have this dream that if you buy the right tool or hire the right person, suddenly all of these issues go away. But it doesn't. I wish it did because if there were a product that solved this, I would love to sell it. But it doesn't work that way. And I don't think it ever will because it's people. It's not always about the tools and it's not about the technology.

Jam: Yeah, I feel like the view that people should have on tools, whether they're security or operations, is that it's an extension, and support for people to do their best work. And so, especially when evaluating vendor tooling right now—because it comes up in my current job, and I also have to keep track of it for trends, to see where the industry is going both technology-wise and security-wise. But when looking at this, you always have to keep the business operations in mind, and I think some security people forget that and jump on, “Oh, we need the shiny new toy for compliance.” Instead of thinking of, “Hey, does the shiny new toy match up to our operational goals?” Do—

Corey: Oh, it used to be DR if it wasn't compliance, or it used to be, “Ah. Redundancy.” Or, there was always a reason to just hurl money at some project or whatnot, where you're never done, but depending on the story you tell, you can unlock massive budget?

Jam: Yeah. So, I’d like to think of it in a new way of being like, okay, does this align with our business values, and is this going to help further our business? And I think security people should keep that more in mind where they're evaluating tools.

Corey: If people want to know more about you, where can they find you?

Jam: So, if you want to find me, I can be found on Twitter at @jamfish728.

Corey: That's right. We are in fact, birthday twins.

Jam: Yes, we are birthday twins. Way to tell my secret about my handle. Um—[laugh]. And yeah, that's pretty much the only place that I have right now. I'm thinking about maybe restarting my blog up. I have a blog at blog.jam.fish that I haven't updated, but I might update it more because I'm starting to get antsy about doing stuff beyond Twitter.

Corey: Yeah, that's sort of what pushed me to doing a whole newsletter, and blog post, and podcast series, and breakfast cereal—next, for all I know. There's always the idea of creating more content. But, eh, it's a burden, you know, because now oh, great. Now you have to update it. But regardless, we'll put links to those in the [00:37:53 show notes]. Thank you so much for taking the time to speak with me today. I really appreciate it.

Jam: Of course, Corey, anytime. Let's do this again soon.

Corey: Deal. And thanks again for covering me for the newsletter so I can enjoy some time with the newborn.

Jam: Yay. I want baby pictures.

Corey: [laugh]. Absolutely. Jam Leoni, lead security engineer at Honeycomb. I'm Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts or your platform of choice, whereas if you've hated this podcast, please leave a five-star review in the same place along with an angry ranting comment about how your degree makes you a better person than me.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Christie BrandaoChristie Brandao is a software engineer at Branch Insurance, a company utilizing a fully serverless infrastructure to sell home, auto, renters, and bundled insurance with just a name and address.

Links Referenced

  • Branch
  • Twitter
  • Christie's portfolio
  • Connect with Christie on LinkedIn

TranscriptAnnouncer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Catchpoint. Look, 80 percent of performance and availability issues don’t occur within your application code in your data center itself. It occurs well outside those boundaries, so it’s difficult to understand what’s actually happening. What Catchpoint does is makes it easier for enterprises to detect, identify, and of course, validate how reachable their application is, and of course, how happy their users are. It helps you get visibility into reachability, availability, performance, reliability, and of course absorbency, because we’ll throw that one in, too. And it’s used by a bunch of interesting companies you may have heard of, like, you know, Google, Verizon, Oracle—but don’t hold that against them—and many more. To learn more, visit www.catchpoint.com, and tell them Corey sent you; wait for the wince.

Corey: This episode has been sponsored in part by our friends at Veeam. Are you tired of juggling the cost of AWS backups and recovery with your SLAs? Quit the circus act and check out Veeam. Their AWS backup and recovery solution is made to save you money—not that that’s the primary goal, mind you—while also protecting your data properly. They’re letting you protect 10 instances for free with no time limits, so test it out now. You can even find them on the AWS Marketplace at snark.cloud/backitup. Wait? Did I just endorse something on the AWS Marketplace? Wonder of wonders, I did. Look, you don’t care about backups, you care about restores, and despite the fact that multi-cloud is a dumb strategy, it’s also a realistic reality, so make sure that you’re backing up data from everywhere with a single unified point of view. Check them out as snark.cloud/backitup.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Christie Brandao, a software engineer at Branch Insurance. Christie, welcome to the show.

Christie: Thanks for having me, Corey.

Corey: So, I know there's a whole spiel about what Branch Insurance does, and I ignore it completely because I refuse to accept it as anything other than what you really want to have when you completely screw up Git.

Christie: [laughs]. Yeah… yep, that could be one perspective of what Branch is, for sure. But at the core of it, I do work for an insurance startup, and our name also just so happens to be Branch.

Corey: Gotcha. Those are two words that almost never go together: insurance, startup. It feels-at least for whatever reason—like, oh, an insurance company, you must obviously be an 80-year-old company with trillions of dollars, and it turns out that's not actually true.

Christie: Right. Yeah, that's not actually true. So, Branch was actually founded in 2018 by our CEO, Steve, and he has 20 years of insurance experience. So, it is quite rare that you have an insurance startup. Normally you hear these big giant company names, but with that giant company name what you get is really legacy technology. So, the fact that we are a startup really helps us bring a super bespoke experience to our customers.

Corey: One of the talking points for Branch on Twitter and other places is that you folks are built entirely on top of serverless infrastructure. So, you're an insurance company that focus on home, auto, renters, the actual usual, quote-unquote “boring” kinds of insurance, not exciting type of insurance, like Git. But you're doing all of those things on top of serverless, which sounds like something that AWS made up to put into a keynote so they could say, “See, there's this company that might be doing a thing like this.” However, after a series of conversations with you, folks, it turns out, it's real.

Christie: Yes. Yeah, it's really real. We have real customers and I'm really working on it on a day-to-day job. So, it's really exciting. I actually knew absolutely nothing about insurance before joining Branch, so it's been a journey learning about insurance and also just serverless in general.

We really do use technology to create the most efficient way to bundle and save with your insurance. So, with just the name and address, we can give you a price on your home, auto, renters, and bundled policies, and from there you can just go ahead and purchase, and—I mean, under 30 seconds. Of course, people take more time to customize and see if all the numbers are correct, but—

Corey: Oh, they don’t tend to do an insurance policy speed run.?

Christie: [laughs]. Yeah, speed run check, 30 seconds in and out. Definitely possible, but probably not recommended. But yeah, it's really exciting to be harnessing the power of serverless in order to do this.

Corey: So, there have been other companies that have come up that have prided themselves on being full serverless one of them was a serverless monitoring company that has since been acquired, and they were big on eating their own dog food slash drinking their own champagne slash whatever ridiculous analogy you want to apply, and using that particular product, it felt like there were always cold starts, latencies in the environment as things spun up, whereas one of its competitors that is still in business as an independent concern—Epsagon at the time of this recording at least because who knows what acquisitions look like in this time—had servers on the front end so that there was always a responsive dashboard and you didn't have cold starts done in quite the same way. How did you folks approach that?

Christie: Yeah, that's really interesting. So, we do have a fully serverless architecture built on AWS, cold starts are still very much a thing for us, specifically for our Lambdas, and so we're probably still struggling with the latency and the speed issues; that’s still a pain point that we're trying to solve.

Corey: So, one of the recurring themes that we talk about on the show from time to time is how people got to where they are in their careers and how we approach, as an industry, finding the next generation. And what makes your story fascinating is that seven months ago, you graduated from a boot camp and took a job at Branch as a software engineer, correct?

Christie: Yep. Yeah, that's right. It's hard to believe that it's been already seven months of—but yeah, I am here. [laughs].

Corey: One of the fascinating parts about that story is that it's not a common one, at least that I've seen yet in the market. You have the whole DevOps engineers where it’s, “What's the job requirement?” “Well, step one: spend 10 years as either a sysadmin or an SRE-type, getting increasingly angry at beating on servers and deep-level Linux, and finally emerge through to the other side with a little bit less anger and bitterness, ideally.” But with serverless, one of the interesting stories here is you got to bypass that entire world of pain in some ways, and although you are a quote-unquote “software engineer,” which in most jobs means infrastructure is someone else's group entirely, with serverless it doesn't generally work out that way. Is that a fair assessment?

Christie: Yeah. Yes, that is an accurate assessment. So, just to backstep a little bit, you're totally right, I did just graduate from a boot camp in 2019, and kind of surprisingly, I also have a bit of formal education in Computer Science as well. I have a minor in Computer Science and I majored in something Computer Sciences adjacent called Geographical Information Science where I ended up learning SQL and Python, so it ended up being great to learn different technologies.

But when I graduated from college, I realized I didn't major in web development. All these job postings have technologies that I don't know. I mean, I learned Java and C++. I didn't learn React and JavaScript. So, I ended up taking a job that had the developer title, but we were using a kind of a no-code authoring tool for eLearning, and probably only eLearning people know this tool. It's called Lectora, and it's basically you just generate a static site by drag and dropping some images and doing some simple boolean logic actions on those images. And you just press publish, and boom, you have your HTML CSS and JavaScript, and you didn't write any of that code.

So, during that position, I really was starting to yearn for more of the computer science and coding that I had studied in college but, of course, everyone knows you don't major in web development and that's the kind of stuff you learn over time. And so I decided to join a boot camp to kind of supplement that formal knowledge of computer science with more vocational skills and skills I can put on my resume, that I can build a full-stack web application with these technologies: React, GraphQL, and Node, and basically enter the job force for web development. So, the boot camp was really helpful in that, super helpful. I mean, it was from March through August, and the majority of the boot camp was focused on pair programming, and then there was a portion of it where you would build your full-stack projects to put on your resume and show that you can actually build these things, then finally there was the job search portion of it where they teach you negotiation for salaries, what to expect, how to interview, and all those things, and it wasn’t—

Corey: So, it’s not just about coding or the algorithm piece, it's effectively a full-service on how to navigate the technical workforce?

Christie: Exactly, exactly. And so as you can only imagine, to have that kind of accelerated curriculum in such a short amount of time, systems design and architecture is not super important to be taught. It's more important to teach the basics, and what an array is and what a string is then to teach about what a server is. Of course, they're equally as important to do your job. But that's not something they teach you to be STE1, let's say, to get your foot in the door. And so I only learned about systems design and architecture and all the different ways you can do those things—scaling, availability, and what all that means—in terms of interviewing for a company so that I can look like I know what I'm talking about on a whiteboard.

Corey: It's an interesting story where you see people going through Computer Science degrees and realizing that there was no minor in clicking the retry button on the failed Jenkins build. And the things that they teach in a formal computer science education don't always have a one-to-one mapping with what's going on in the workforce. So, knowing what now, would you do anything differently? Would you avoid the CompSci degree? Would you avoid the boot camp? Would you do exactly both? Would you do neither and decide, “Wow, I'd be much happier raising goats somewhere instead. Computers are terrible.”

Christie: I do love animals. Maybe I'd go and raise some parrots, or birds, or something. But in terms of tech, that's a really hard question to answer. I always struggle with these types of questions because you can always—so go around back to the fact that well, I wouldn't be here exactly where I am today, and I'm happy where I am today if I didn't do all those things. So, it's hard to say that I wouldn’t do my four-year bachelor's degree in something technical. I think it helped me build the foundation of computer science, but I really do think I could have gotten here with just the boot camp experience because that's where I learned my day-to-day skills that I use on the job at Branch.

Corey: One of the neat things from your perspective that I really want to make sure we get to talk about is, if you talk to people who have been in the space for, let's say, the last five to ten years and then you talk about serverless, everything that they're dealing with and interacting with is filtered through the lens of having to worry about stateful, large server-style systems for that class of folk-and I'm certainly in that class—it's a story of everything we're learning is filtered through a lens, and now we're defining serverless in terms of its constraints, and that can lead to a very frustrating getting started story. It sounds—and please correct me if I'm wrong—like you aren't coming from that background. You are effectively having infrastructure handled for you out of the gate.

Christie: Yeah, exactly. So, when I left the boot camp, and I learned about architecture, and scaling, availability in order to interview properly, I left with the impression that I had to learn those things to be a quote-unquote “good developer.” And everybody defines what a good developer is, and I feel like the goalposts are constantly shifting and that kind of contributes to imposter syndrome a little bit, but these are things that I'm going to have to learn, or at least are handled for me internally in my company that I will be joining. But I was so wrong. I was so wrong because to me when I left and I saw the term serverless on the job posting for Branch, honestly, I thought it was, like, a buzzword at that point. I didn't know much. I thought it was a buzzword, kind of like how people use ‘machine learning’ and ‘artificial intelligence’ is marketing terms.

I'm sure that people still are using serverless is a marketing term, but it meant so much more than that and it played so much more of a bigger role in my day-to-day development job than I thought it would. So, I have been able to contribute—within six months of joining Branch—tickets that complete features and bugs across the entire stack, and I've only been able to appreciate the things that serverless has brought for me, which is the empowerment that I don't have to worry about things like availability and scaling, and I can focus on the problems that directly contribute to the business and features that directly contribute to the business. And that's what brings me joy and satisfaction in my job is to see the difference that I'm making for the business, and not necessarily, “Oh, what's going on with the server?” Because let's let the pros handle that and the vendor that is really good at doing those things.

Corey: I am going to attempt to head off the swarm of replies to what you just said that I'm certain I will get where, “Oh, if your provider is handling your availability, you're going to have a problem someday.” I would call the attention back to what you said at the beginning of this show, that you are an insurance company. If all of AWS Lambda, for example, is down for a couple of hours, everything is more or less going to be broken, and it's not as if people are going to change their mind suddenly and not buy an insurance policy in that sense. That is a global cataclysmic event for what your business does, it sounds like that is not a showstopper for what you do. Is that accurate?

Christie: Yeah, that's super accurate. I remember during my first month at Branch when I was talking to Joe, who is the CTO and my direct boss, I was grilling him with these questions, “What is the purpose of serverless for us?” And, “I don't quite understand.” But it ended up just being that if AWS goes down, we have bigger problems than that. Our problems are not contained to, “Oh, we cannot have this person buy a policy right now.” It's going to be a bigger problem that is affecting way more people than just us. And that's okay.

Corey: It comes down to, I think, understanding blast radius and what's going to break in that eventuality. I've always looked with scorn and derision at disaster recovery and business continuity plans that assume that the city is lying in ruin, but all your employees are going to show up on time to head to the backup site, and… yeah, people are going to care about their families a heck of a lot more than they are about anything that's going on for work. It's the human element, like, view this in the larger context of what's going on.

Christie: Yeah, exactly. You said it better than I could.

Corey: So, when we take a look at what the onboarding experience was like, how was it in the early days? What was the experience of encountering the modern state of serverless? Not serviceless as it was four years ago when the console broke half the time, but today in this era, what was that like not coming to it with a bunch of preconceptions that were earned through ten years hard labor in the data center salt mines?

Christie: So, I mean, it's kind of difficult to answer this question because I haven't been burned by these types of issues that you would have if you weren't serverless, but I do get a little bit of insight from my partner. He's a software engineer as well, with the traditional background, and a five-year experience in the field, and so I kind of talk to him about these things after work. And so I do kind of understand what the pain points would be, but for me it kind of really clicked in my first month on the job, where I was completing a ticket where I needed to create an SSM parameter or something, and I thought, “Oh, man, I'm going to have to contact someone to figure out how to create this,” and then I realized, “Oh, I can just do it in the AWS console.” I go in I type ‘SSM.’ It's not the term that popped up, it was ‘parameter store.’

I'm like, “Okay. It's the third option; that's weird, but I guess this is what I need.” And I go and click manually creating the SSM parameter, and then I have a moment of clarity. I'm like, “Wait, how is this going to get to every other developer’s environment?” And just a side note that every developer on our team has their own AWS account with their own environment that's pretty much a copy of production so that we can deploy our code that we're working on to our environment and not step on anybody's toes if we break something, which is really great.

So, I went ahead and did that, and I realized, “Wait, how is this going to get to everyone else's environments?” I call Joe over and then he just shows me the magic of infrastructure-as-code, and I can't even explain to you how empowering it felt to be able to write some code and create the SSM parameter, and know that this is kind of automatically going to happen for everyone else when you deploy. And so that was kind of when it really clicked for me that this stuff is really powerful, and it's really empowering to be able to work on any ticket across the full stack without worrying about things like availability, and scaling.

This episode is sponsored by our friends at New Relic. Look, you’ve got a complex architecture because they’re all complicated. Monitoring it takes a dozen different tools. Troubleshooting means jumping between all those dashboards and various silos. New Relic wants to change that, and they’re doing the right things. They’re giving you one user and a hundred gigabytes a month, completely free. Take the time to check them out at newrelic.com, where they’ve done away with almost everything that we used to hate about New Relic. Once again, that’s newrelic.com.

Corey: One of the fascinating parts about all of this is—and I think the reason that I had a lot of trouble with serverless at first—was I came to it with this idea of, “Oh, it's like a Linux box only weird and strange,” And I kept—at least [00:19:07 unintelligible]—smacking into and finding the problems with my face because reading the documentation is clearly A) for losers, and B) a lot harder back when the documentation was in a prototype state compared to what it is today. It's, “Oh, what do you mean I can only have so much code in there; I can't package up an entire environment? Huh. What the hell do you mean the file system is read-only? Oh, slash temp is there. What do you mean that file could still be there if it re-invokes or not?”

It took me two weeks to build my first Lambda function because I kept smacking into those things, and I'm a terrible programmer. And by the time I got to the end, the most recent one that I did took less than 20 minutes because I am a terrible programmer, and don't know what tests are. So, it comes down to an understanding, at least, of the constraints, and once you wrap your head around that model, in my experience, it is way faster, more durable, fewer things to break and, let's not kid ourselves, a lot less money.

Christie: Yeah, exactly. So, when I talk to Joe about these things, one of the things he brings up a lot of the times is that we were running for two years on our AWS credits, which is just, kind of, mind-boggling that you can just create a company, a startup, and be able to survive on those credits for serverless, and have it up and running and have actual users, which is kind of mind-boggling for me when I found out about that. With your experiences with Lambda, and taking two weeks to build it up, I guess I'm really lucky that we already had most of that built out for me when I joined, and so all those pain points of starting a serverless architecture was kind of already weeded out for me once I joined. But I can only imagine the type of knowledge that you have to have of how you've been burned before in order to figure out how to create this new architecture and have it resilient, and it kind of naturally works out.

Corey: It also—almost definitely from what you're saying—helps that you have coworkers that you can ask these questions to who have already figured out a lot of these answers. In the early days, when AWS releases something on stage, and everyone's super excited about it, it’s available today, well, just because AWS releases something to production doesn't mean that it's a good idea for you to release it to production. Now there's a community out there that has found the sharp edges, knows how to explain things—at least ideally—and you can reach out to folks to get guidance on this, it feels like just from a community perspective alone, it is so much nicer having that path. But not to put words in your mouth, how did you wind up getting help when you started figuring out how this worked? What resources did you fall back on?

Christie: At first, I went to the docs to read all these things, and I figured that I wasn't really grasping it as well as I wanted to, so like you said, I did fall on the community. I ended up asking Joe for a lot of his recommendations of people to follow on Twitter, did my own research online for people to follow, and resources to learn these things, a lot of YouTube videos, a lot of Googling, but honestly, it was pretty overwhelming. The amount of services that AWS has, just that list when you first look at it—I mean, all I knew about when I joined Branch was Lambda, and so it was news to me that it's actually pretty common to not know all of these services that AWS is offering, and that's okay because you don't have to use all of them. It's just about figuring out which one works for your need, and you don't have to know about it before you need it. So, when I was trying to figure out something that I could grasp into for Branch and, kind of, take ownership over, one of those things was observability. And so, it was a lot of researching on, “Well, okay. How can I use what AWS is offering in order to send these events to whatever vendor that we're going to be using for observability?” And it was just overwhelming, but I think the community—and you’re super right about this—the community does a really great job at breaking it down for those people who don't want to rely on those docs for that information.

Corey: So, what would you say has changed in your understanding or perception of this, in the seven months you've been there, that would alter how you would have approached this on day one? In other words, what feedback would you give to you seven months ago who's entering the workforce today, to make their path easier?

Christie: So, I was just a part of an alumni panel for my boot camp last week, and I think a lot of the questions that people were asking were either directly about imposter syndrome or indirectly about it. So, if I were to give myself advice seven months ago when I was applying for all these jobs and reading about all these technologies that people are using, is to not be afraid to challenge yourself. So, I didn't know anything about serverless. Like I said, I thought it was a marketing term. I thought it was just the newest buzzword in tech, the new sexy thing—

Corey: The first time I heard it, I heard it as ‘Cerberus,’ and it's, “Oh, great. Do I have to fight the Minotaur when I get there?” And it turns out that if you view the API gateway through a certain lens, yes, yes, you do.

Christie: Yeah. It's like, “Serverless? What does that even mean? No servers?” Well, I mean, it just turns out, it's just someone else's computer, right? So, I mean, those are the types of experiences that you really don't fully grasp unless you ask someone who's been in it before. So, I try to assure people that, just because you see these terms and they might be new—or newer—than that doesn't mean that it's going to be some type of otherworldly experience that you have to learn.

It turns out for me, at first it's just JavaScript in a Lambda and that Lambda runs on someone else's computer. There's nothing really to be worried about, or frightened, or scared that you can't learn it. It's all learnable; people have done it before, people will do it again. And that's the kind of advice that I would give to people starting out. You don't have to learn about these things in a boot camp in order to be able to use them in your everyday dev career.

Corey: Do you do local development on your serverless stuff and mock services locally, or do you do remote development in the Cloud where every iteration you push, and then run it in the actual AWS environment?

Christie: I'm writing my code and then our Lambdas are actually, most of them are a bit too large to be just writing them inline in the AWS console—or if I need to do a console log or something like that, a lot of the times I'll just locally invoke it, but for the larger stuff, and most of the time, I do just deploy it to my environment so that I can ensure that it's not going to be breaking something I wasn't expecting it to break because that's, kind of, how I gain assurance that my code is working is to see happen all together in my environment, which is the copy of production.

Corey: I've always found that mocking cloud services—except for the way that I mock them, which generally includes making fun of them—is sort of a losing proposition at some point because you're always going to have an imperfect copy, and developing to that rather than the actual thing never seemed like it was quite the right direction to go in for me. People say, “Oh, well, what about when you're on a plane and need to develop?” “Yeah, I spent entirely too much time in planes most years, and I still don't see that I spend that much time in that situation to the point where it is worth making that the common use case for my development work.”

Christie: Yeah, yeah, that makes a lot of sense. I mean, something that I have found very valuable is I was reading a lot about Charity Majors because I'm working on a lot about observability and how to implement that into our codebase, so she's a huge proponent of testing in production which is super scary, and a lot of people immediately probably wince when they hear that, but I think it’s really important—

Corey: I’m a big fan of testing in other people's production. It's easier. Less risk for me.

Christie: Yeah, exactly. Yeah, sorry, Joe, just messed everything up there. But I think it's really important because you can't ever mock exactly what is going to be happening when a real user is interacting with your app.

Corey: That is, I think one of the biggest problems with technology is it's almost impossible to effectively predict what a user is going to do and how they're going to break things.

Christie: I can't even tell you the amount of times I've audibly laughed out loud at watching people interact with our app through those screen recordings. I'm just like, “Why are you clicking that button? I can't believe you even found it that way.” The types of things that people will go through on your app is unimaginable.

Corey: It's still one of the biggest challenges I see is I've been doing this for a decade and a half now and I don't know where I would tell someone to start if they were starting today because there's so much that brought me to where I am. It's nice to see that there are paths forward that don't require going through the same levels of pain in the same areas.

Christie: Yeah, definitely. It's really freeing to know that the boot camp experience and the non-traditional background is being more accepted now in technology. So, it's really great that alongside me, I have my colleagues that have all different types of backgrounds and come from all different paths of life, and it's really exciting to be able to see that our future generations are going to be able to come in technology and have that experience as well.

Corey: Christie, thank you so much for taking the time to speak with me today. If people want to hear more about what you're up to and how you view these things, where can they find you?

Christie: Sure. Yeah. I'm on Twitter at @ChristieBrandao. And then I have my portfolio website, which I believe you linked below, but it's just cbrandao.dev.

Corey: Excellent, we will indeed put those in the [00:29:02 show notes]. Thanks again for your time. I really appreciate it.

Christie: Thanks, Corey.

Corey: Christie Brandao, software engineer at Branch Insurance. I am Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts, whereas if you've hated it, please leave a five-star review on Apple Podcasts and tell me exactly what kind of servers serverless runs on.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Courtney WilburnCourtney's journey began in her hometown of Memphis, Tennessee, by hardware hacking on personal computers and lurking on Prince message boards. She loves finding unusual and efficient ways to solve problems (both human and technical alike), building tools, developing workflows, and building infrastructure almost as much as she enjoys finding ways to keep activists safe organizing online. When she’s away from her desk, she can be found running, hiking, or biking around Philadelphia, cooking, brewing beer, knitting, building keyboards, or singing karaoke duets with her wife.

Links Referenced

  • Elastic
  • Twitter
  • Connect with Courtney on LinkedIn
  • Courtney's personal site

TranscriptAnnouncer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: In what you might be forgiven for mistaking for a blast from the past, today I want to talk about New Relic. They seem to be a relatively legacy monitoring company, and I would have agreed with that assessment up until relatively recently. But they did something a little out there: they reworked everything. They went open source, they made it so you can monitor your whole stack in one place and, most notably from my perspective, they simplified their pricing into something that is much more affordable for almost everyone. There's even a free tier with one user and 100 gigs per month, totally free. Check it out at newrelic.com.

Corey: nOps will help you reduce AWS costs 15 to 50 percent if you do what tells you. But some people do. For example, watch their webcast, how Uber reduced AWS costs 15 percent in 30 days; that is six figures in 30 days. Rather than a thing you might do, this is something that they actually did. Take a look at it. It's designed for DevOps teams. nOps helps quickly discover the root causes of cost and correlate that with infrastructure changes. Try it free for 30 days, go to nops.io/snark. That's N-O-P-S dot I-O, slash snark.

Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Courtney Winburn, engineering manager of Cloud SRE tooling at Elastic. Courtney, welcome to the show.

Courtney: Thanks for having me. It's really good to be here.

Corey: So, you're a relatively recent hire at Elastic for Cloud SRE tooling, and I swear to you, when I first saw that title, I misread it as cloud SRE trolling, at which point it's, “Oh, what I'm doing actually had turned into a valid career path.” No, that's not what it says, and no, there is not an entire industry of people making fun of things, that I'm aware of yet.

Courtney: Yeah, no, there certainly isn’t, but that's not to say that there won't be in the future. I mean, I feel like there's certainly an avenue for trolling for sure. And certainly in my first few weeks, the amount of questions I had, I certainly felt like I was trolling. So. [laughs].

Corey: I hear you. So, we'll get to what you're currently doing in a bit, but first, what I find fascinating is that you worked at one of my favorite places in the world for a couple of years before you wound up going to Elastic, and that is the Wirecutter, a New York Times company.

Courtney: Yes.

Corey: Because, I don't know about you, but every time I want to buy something, and I don't necessarily want to, I don't know, buy 20 different spatulas and do a bake-off for six weeks, I go ahead to the Wirecutter and I pull up, what is the thing you recommend? Great, I'll go ahead and buy that. It's been my default thing to go to that has replaced Consumer Reports for this generation.

Courtney: Oh yeah. No, I mean, it totally cut out the fat for a lot of the things that I was purchasing. When I first started there, I was like, “Oh, great. I can use this and it feels like I'm supporting my own workplace.” It was great.

I went there again, yesterday, just looking for a new foam roller because I've been working out a lot lately. It’s one of those pandemics sort of shifts in trying to change your headspace because you're in the house a lot more. And I looked for a foam roller, and then I was like, “I know exactly where to go. I don't have to go anywhere else.”

Corey: Yeah. The problem I found is that, especially for things like home decor, whenever the Wirecutter recommends something, there's an entire swath of folks, at least in the circles that I hang out with, that will go directly to the Wirecutter and buy the thing that they recommend. So, there's now a disturbing, almost, monoculture of everyone I know has the same standing desk, they have the same monitor arm they have the same chair that is always way too expensive—but if it's between you in the ground, spend the money. Life tip there—and everyone has more or less the same equipment. And that's great for equipment, but then it’s, oh, everyone has the same couch. Everyone has the same wall hanging. Oh okay, this is getting a little weird.

Courtney: Right, right. I mean, I think the point is to pepper your life with some of these things and replace them, but more power to the people who walk into an empty home and just fill it with Wirecutter recommendations. I certainly would if I could, but the office chair pick, that was certainly a lifesaver, especially working from home. [laughs].

Corey: So, you went from a company that effectively is a household name, at least in the world that I live in. And then you've gone to Elastic which, love them or hate them, I don't think I know too many companies that aren't using Elasticsearch for something.

Courtney: That's true. That's true. I kind of like the fact that Elastic has that level of ubiquity in sites. I mean, even the Wirecutter used Elasticsearch. So, it's knowing that, especially as the amount of things that people are looking for online continues to increase, it doesn't seem like there's going to be any end to the use of Elasticsearch and people's applications of it.

So, it's been an interesting ride, especially as I continue to learn the depths of the domain. I think, being on the other side of being a consumer of Elasticsearch while working at the Wirecutter, you interfaced with it, you don't really understand the—or you don't get a sense of the power of the product. And now, wow, as I continue to dig in and learn more, definitely impressed.

Corey: One of the things that I think is so impressive is that everyone in the world is either chasing Elastic, or building their own Elastic offering, or competing with Elastic, or using Elastic. Some people love the company, some people hate the company, some people like, “Oh, I love the product, I use it constantly.” “Oh, I hate the product, but I'm going to have full API compatibility with it.” And so on and so forth. And it's really, sort of, become the default/only answer for arbitrary searching of data.

I was, sort of, in denial about that for a long time, until then I started seeing Elasticsearch showing up in what's running in our AWS environment here at the consultancy, and, “Yep, okay. If we're using it ourselves, then there's really no argument anymore.” It is incredibly pervasive because we make terrible technology selections and try and go off the beaten path wherever possible because I'm really bad at this. And all roads do, in fact, lead to Elastic. Because it makes sense. “Great. You're going to use DNS is a database. Great move there, but what are you going to search through the logs with?” “Oop, there we go—”

Courtney: There we go.

Corey: —“Elastic it is.”

Courtney: Yep. Oh, yeah. No, absolutely, absolutely. And the applications for what they're doing are endless. And it's good to be a part of this. It's really—specifically working in the cloud division, being a part of the cutting edge of the cutting edge. This is something that is a technology that is ubiquitous, but we're still finding novel ways to make it work for a variety of different use cases. It's been pretty awesome. I can't say more about how much I have enjoyed my time there so far.

Corey: Yeah. Before we go too far into this, I also want to thank you for—at the time we're recording this, I'm about to step out for parental leave, and you are a guest author for the newsletter one of the weeks that I'm out. First, thank you, I really appreciate—

Courtney: Oh, you’re welcome.

Corey: —being able to take time off.

Courtney: No problem. Really, congratulations on your upcoming leave. I'm sure your employees are going to be happy about that. [laughs].

Corey: Oh, getting rid of me. Oh, yeah. They've been trying to plan how to get rid of me for years. The ringleader of that is, of course, Mike, my business partner who just absolutely—like, he wants to absolutely seize power and write the newsletter himself, I'm sure. I'm kidding. He's written his own before, and it was not directly aligned with what he loved doing. I'm weird like that.

But in seriousness, there's a lot of really talented people that I get to work with, but this is the first time I've ever stepped away for more than a week or so during the offseason of, “All right, it's between Christmas and New Year's. No one is going to be doing any announcements here. Why run a newsletter issue at all?” I think I've taken three weeks off until this point over the last four years.

Courtney: Oh my god.

Corey: So, this is a new experience. I am taking multiple weeks in a row off. I'm turning off my access at the end of the week. And I will read the newsletter when it comes out and it hits my email between waking up or feeding. So, you know, when I wake up at two in the afternoon.

Courtney: You’re just being a consumer of it, just like everybody else, right?

Corey: Exactly. And we'll see what happens. And I look at this and, “Okay, great. Well—” again, the trick is, as I'm envisioning it now, it seemed impossible when I was first considering this idea because no one is going to be able to write sarcastically and snarky the way that I do, but they don't have to. Other people's material fits about as well as other people's shoes.

I don't need to play those games. People can tell the stories of what they find interesting, and it's all subjective. I get angry emails sometimes from Amazon folks. “Why didn't you talk about my release?” Well, the answer is, “When I saw it, it didn't really seem that interesting to me. What did I miss?” No one is going to ever agree with everything you say, so I smile, I nod, and I made my peace with that.

Courtney: Yeah. I mean, I think I've definitely had to make my peace with people not agreeing with what I say [laughs] just generally. I feel like in order to blaze your own trail in this industry, especially as a Black woman, you kind of have to be used to not being a part of the in-crowd, sort of being a bit of an iconoclast. You can sit back and get a 10,000-foot view. I doubt my issue will be as snarky as the ones that I've read of yours in the past, but I hope that people will enjoy it nonetheless.

Corey: [laughs]. That's a piece I always wondered where I've given talks at conferences before on how to handle job interviews or salary negotiation, or how to be sarcastic and funny, and the problem with those talks and the reason I don't give them any more independently, without someone else on stage with me to share their experiences, is it's extremely unclear for me how much of what has worked for me has the unwritten prerequisite of, be a loudmouth white guy in tech. I mean my failure mode is apparently a board seat and a book deal somewhere. When I say, “I don't understand how this works?” The answer is, “Oh, the marketing is not great, and they should shore up the documentation.” Where no one ever questions, “Should I be here or not?” That's an incredible baked-in level of presumed competence that, frankly, I certainly don't deserve, but also that is not afforded to everyone.

Courtney: Oh, absolutely. It's most certainly not. I mean, I think I've at least walked into—you know, even before I was working in the industry, everywhere I was, there wasn't necessarily something that was tailored for me. I started my career path, generally, not doing tech stuff. So, you get a better understanding growing up how much of the world is tailored to you, and how much isn’t.

And, you know, for me a lot isn’t, and so you just adjust. You make your own way, you figure it out, it can be easier growing up and having your lived experience knowing the world is not tailored for you. It can be easier in some ways, or at least for me, going into another job, an industry change because you know that that's not necessarily going to be tailored to you either. So, you can say, “Okay. I know who the intended audience is for this. How do I make them listen? How do I make them see me for me, even though this isn't built for me?”

Corey: Now, as of the time of this recording, you’ve spent a month or two as an engineering manager at Elastic. Is there an existing culture of hiring Black people in the abstract, and Black women specifically, into leadership roles? Is this a relatively recent change? Is this something that has been baked into the culture for a long time that they're getting right? Is this something new?

Courtney: I would say it's a bit of both, right. I think, generally, they've been able to, as a fully distributed company, been able to hire people from around the world. That being said, I think because they're in the tech industry, the tech industry is over-represented by white, straight dudes. And for that reason, I was the first Black person hired into engineering leadership in the company's history. So, I didn't know that until I started there, but I don't feel unwelcome.

I wasn't made to feel like I was an artifact in a museum, or everyone's looking at me or, “Oh, crap, I have to get this right or there'll never be another Black person in engineering leadership again.” I don't feel that way. I think they certainly have talked the talk. I think, when it comes to making sure that this is in a warm, inclusive environment for a wide variety of people, I think there is a good level of representation, just, sort of, in individual contributors. I think this is a step in the right direction in terms of them continuing to walk the walk when it comes to making sure that the representation in leadership is as reflective of the representation for individual contributors there.

Corey: There's an awful lot to admire about Elastic as a company. It's easy to drag any company when they take a misstep or offend someone's perceived sense of right and wrong, but from what you're saying, it sounds like they have an awful lot, even internally, that recommends them as a place to work.

Courtney: Right, right. They're fully distributed. I think so many people—I've really been welcomed with open arms by other Black folks that work at Elastic. And it's been an amazing experience. Just, so many people in so many different disciplines doing things that, I think to me, being someone who does engineering work or continually had done engineering work in the past, seeing Black people in marketing blows my mind, you know what I mean? [laughs]. Seeing Black people in sales, that blows my mind. But people from all areas of the company, just being in a community with each other, being encouraging to each other is really nice. And the fact that folks feel so welcoming, and really, truly also believe that what they're doing at the same time, and don't seem to be disheartened or down, it's been great.

Corey: You obviously were not brought in for optical purposes. You have an incredibly strong engineering and engineering leadership background. To that end, you are the engineering manager for Cloud SRE tooling. What does that team do? Where does it start? Where does it stop?

Courtney: The team that helps make tools for cloud engineers at Elastic. So, we're very much working on improving the suite of internally facing tools for Elastic engineers, to help them do their jobs better, specifically for SREs. So, we want to make sure that deployments run smoothly, and that they're able to do that, that people have tools at their disposal to make Cloud products even better, to ship better Elastic Cloud products.

Corey: So, someone could argue on some level then, that your team builds the tools that the cloud providers should, but haven't.

Courtney: Um—

Corey: Or is that too dark and cynical?

Courtney: That's a little too dark and cynical. I guess. I don't know. I mean, I think the use cases of our products are so specific. This is more of… how do we launch the Elastic Cloud product, and what do people need to launch the Elastic Cloud product?

So, it's more to help the people that do that, do that better. We're not necessarily looking to each of those cloud providers, even though—because we launch in multiple clouds, we're not looking to them to provide those tools. This is so that folks at Elastic can launch Elastic into those clouds.

Corey: So, in previous jobs, were you working extensively with Elastic, or was that something that you had avoided, and then have to come up to speed on the technology stack itself in your new role?

Courtney: I was working with Elastic, but not very deeply. It was more of, like, at least for most of my career, Elasticsearch was the magical component that just kind of worked. And I think that's what makes it such a draw for so many other companies that use it, is that it just works. You can launch it locally, and it just kind of works. So, that was how I was using it in the past in my career: if we were embedding Elasticsearch into other sites, as long as you knew how to tune your indexes, it just kind of worked.

It didn't make a difference if you couldn't leverage every single bit of Elasticsearch that you wanted to. It just—and so now I'm in the position of really having to learn the nuts and bolts of the architecture so that I can better serve the folks that I manage a bit better because I want to know what they're facing on the day-to-day. So, I did have to get up to speed. I wasn’t, like, totally behind the eight ball with this one, but I think just generally having experience as an engineer certainly helps. I definitely had to get a little bit past the magic and into what's in the hat.

Corey: Yeah. I think that everyone believes that, “Oh, you must be an expert at a company's product to wind up working there.” Not at all. When I started my own company, I vaguely knew some things about AWS, but I was certainly not deeply steeped in it.

Courtney: Have you been able to figure out who's responsible for all of the drawings for the products in AWS, though?

Corey: I have the distinct impression that, at least from what I've seen, that it feels like it's done by committee, in some respects. Because it is so flat, and I guess, devoid of personality. I’m sorry, are you referring to the diagrams that they use, or are you talking about their icons, or stencils, or—

Courtney: Their icons. The icons for the products. I—really indistinguishable. If I was going solely off of icons for some of their products, I would have no idea what each of them meant.

Corey: I'm having one of my designers build out an entire series of custom icons for architectural diagrams, just because I'm so sick of the unimaginative things. Without a label, most of the ones that AWS provides are useless to me. I do not know what they are or what they look like. And if you just need the label, then I could put any arbitrary image there that I want, so why don't have something that at least is slightly more evocative or, failing that, has a personality. But again, this is the problem of big companies, where they're trying to serve as all customers.

You can't really have a sense of humor about your offering as an enterprise-scale company because your customers most assuredly do not have a sense of humor about their own business. So, it's a difficult messaging line to walk, and I really do empathize with them. I'm very fortunate in that I'm a small enough company that I can still pick and choose who I work with, so if folks don't find my sense of humor appealing, great, we're probably not going to have a great engagement anyway. It's sort of—if having a snarky platypus as a mascot is a deal-breaker, I probably was not going to find success there anyway.

Courtney: Right, right. I mean, so to a certain extent, you're the product, right? Your personality is your company's product. Does that make sense?

Corey: Yeah, on some level. I'm trying to break that out a little bit from the consulting versus the media stuff that I do because I do live in fear of a bad take has the potential to wind up impacting ten people's livelihoods. And that's a heck of a responsibility. It's the reason I was so reluctant to begin hiring people, where, sure if I screw up and it's me, I can go find a job somewhere and I'll survive. But now, it's other people are depending on me to get this stuff right. And that becomes a heavy weight to carry, from time to time.

Courtney: Yeah, I mean, I certainly look at being an engineering manager in that way. I'm responsible for people's lives. I feel like generally in the field, though, of engineering management—I'm fortunate in that I'm getting—this is my second go at being a manager. I was a manager in my twenties, and because I was in my twenties, and my main priorities in my twenties were planning happy hours, I think I have a little bit more perspective about how to be responsible to people. But on the flip side of that, the bar is incredibly low. It really is. For how people feel responsible to other people, and how they live that, how they behave, is it whether or not they're responsible to other people, the bar is, like, underground. It's, [laughs] maybe the bar does not exist.

Corey: This episode is sponsored in part by our friends at Linode. You might be familiar with Linode; they've been around for almost 20 years. They offer Cloud in a way that makes sense rather than a way that is actively ridiculous by trying to throw everything at a wall and see what sticks. Their pricing winds up being a lot more transparent—not to mention lower—their performance kicks the crap out of most other things in this space, and—my personal favorite—whenever you call them for support, you'll get a human who's empowered to fix whatever it is that's giving you trouble. Visit linode.com/screaminginthecloud to learn more. That's linode.com/screaminginthecloud.

Corey: Yeah, on some level, it seems that, “Oh, I said something dumb, and half the company got fired? Well, those people were dumb enough to trust me. I guess they'll know better in the future.” It’s like, no, no, no. That is not they're failing, let's be very clear on this.

Courtney: Yeah. You're responsible for other people's lives… if you are, in any capacity for their lives or livelihood.

Corey: Management's hard. I mean, the reason I generally don't have direct reports, even now, and the reporting structure of rolls through my business partner is that my belief has always been that to manage people effectively, you've got to be extraordinarily promotional of them and what they do and help build them up. Whereas on the media side of what I do with the podcast and the newsletter, I have to be incredibly self-promotional. It's a weird expression of marketing. So, that means DevRel meets a bunch of other things. And I don't know that those two are necessarily congruent, at least as the way that I believe management should be done.

Courtney: Yeah, that makes a lot of sense. I think a lot of being a good manager is turning around and highlighting what other people are doing, showcasing those people, lifting them up, building them up. I've been thinking about what my philosophy as a manager is if that makes a lot of sense. And I think that the closest thing that I've come to is having a trauma-informed approach [laughs] because I feel like working in this industry is very, very—especially for a large company, or companies at the scale of which I've been working up to this point, for lack of a better term, like, burnout or approaching burnout is a traumatic experience for people.

So, how do you head that off? What signs do you look for? If someone's approaching burnout how do you help someone recover from burnout? Because not everyone can afford to take time off and not work for six months in order to recover. So, how do you do that? How do you keep those people happy, and alive, and aware, and still working toward whatever goals that you set?

So, thinking about that in terms of being trauma-aware or, and thinking about the things that people have encountered, either projects not going the way that they wanted to, or getting passed up for opportunities that they felt belonged to them, and what that can do to someone, and how I can be a person to help them heal, but still keeping them happy and productive at the same time.

Corey: That's the hard part from my perspective, is figuring out how to balance all these competing objectives and things that need to get done in certain orders and certain priorities: saying yes to one thing means saying no to something else. How do you prioritize? I didn't realize how much of management was juggling.

Courtney: Oh, yeah. Oh, yeah. It's juggling on a unicycle while music is playing right. It's a lot, but I like it. I certainly was not equipped for it in my twenties. I'm glad that I had a chance to take another stab at it this late in the game.

Corey: Yeah. One other thing that I noticed in your biography that just absolutely resonates with me is, oh, great. You have the same horrible, obnoxious, expensive hobby that I do. Namely, building mechanical keyboards.

Courtney: Yes.

Corey: Tell me a little bit about that.

Courtney: I like throwing money out the window, setting it on fire,k I don't know. I think it actually—seriously though, it actually has been a fun way for me to pick up new skills, or at least it started that way. Now it's just sort of me dumping money into a dumpster and collecting switches and all these sorts of things—

Corey: Yep.

Courtney: —and an outlet for me to have a series of projects in varying degrees of completion.

Corey: Mm-hm. Yeah, it's like, “How many ErgoDoxes have you built?” “Two and a half.” “Well, tell me more about that half.” “I’d really rather not.”

Courtney: [laughs], yeah.

Corey: Yeah, it’s—like the pile of switches, like, “Oh, are those the loud ones or the quiet ones?” “Well, technically, those are the weighted Zealios that have a slightly different model—” Yeah, going down that entire ridiculous rabbit hole. And the worst part of all of that is that now that we're all working remotely, you don't get the payoff from all of this, which is annoying the living hell out of people in the open-plan office next to you.

Courtney: Oh, I've made up for that by annoying the [00:26:26 liver] out of my wife who works from home right now because of the pandemic. She's working a floor below me and every time things get a little too loud with the typing, she shouts, “Dear diary,” [laughs] up toward me because she can hear everything clicking. And it’s like, “Oh, I must be on a roll here.” So, yeah. I mean, it's absolutely an annoying habit. But it really is a way for me to just have fun. And I get to play with things, and keycaps with different colors, and I learned a lot about keycap profiles. That was something that was completely foreign to me. I was just like, “A keyboard is a keyboard.”

Corey: Oh yeah. The SAs quite a bit, but my hands do get tired.

Courtney: Oh, yeah. No, the SAs are nice. It certainly reminds me of, I remember, my mom worked in the la—she was a quality control chemist for most of her career and had this really old calculator that weighed maybe, I don't know, 55, 60 pounds—because, you know, it was a calculator from the 60s—and the keycaps, they definitely, I'm guessing they're probably SA profile, but just to be able to do simple math on it, how hard you had to press the keys, at least for my six- or seven-year-old fingers, it was like the equivalent of running a hammer on these just to—you know, eight plus eight. Okay, let's wait 30 minutes to get the answer. And it's funny how much of that stuff has come full circle. I would probably kill to have the—

Corey: Oh it has. It’s a hipster keyboard, let's not kid ourselves.

Courtney: Oh, yeah.

Corey: But it's such a nice departure from our day-to-day work lives of making the lights on the screen form different patterns which, from a very literalist perspective, is what our jobs entirely are. And that's, “Oh yeah, this is something real in the world I can point at and say I built that.” Or in more common cases, “There's that thing I tried to build, failed, and it's sitting there just taunting me with disappointment in the corner for months on end until I clean it up.”

Courtney: Yeah, no, absolutely. I’m going to be like, “Oh, look at those cold solder joints. Oh, I messed this up. [sighs] oh, I'm going to have to buy another PCB. How much does that cost again? [sighs]. Oh, right.”

So, yeah, it's been so much fun. And I actually have another insanely expensive PCB coming in the mail in November that I'm getting—I'm like, I can't dress my dog up for Halloween and take him trick-or-treating this year, so at least I have this keyboard coming, and that can be the equivalent of a Halloween costume in my mind.

Corey: So, here's the $64,000 question. What is your current keyboard?

Courtney: My current keyboard is a—I use the Laser SA keys from I think MiTo, and it's on the GH60 right now. And those switches are MX CHERRY Browns. So.

Corey: Yeah, those aren't super noisy, so—

Courtney: No. Not super noise.

Corey: —I don’t what your wife is complaining about, now? If they want noise, yeah, here we go. My next project, when I decided to get really fun about this is—I don't even care what the key switches are going to be, but I want to wire in an Arduino or microcontroller of some sort, and a mini-speaker so it plays a sound on every keypress because I don't care what the click is I want it to make a beep, or a boop, or something incredibly annoying, or spark off, “Surfin’ Birds” so it’s, “Bird’s the word” for 30 seconds with every keypress and see how long it takes someone to come in and use that keyboard to beat me to death in my chair.

Courtney: I mean, my keyboard isn't loud, but it certainly makes—I think the edges of the keycaps make enough of noise to be annoying when I'm really on a roll. But I definitely think if I did that, I would absolutely not make it through the week, or maybe through the day if it was making any noise. I would love to make a soundboard though; get a small PCB and make a soundboard that does the, like, air horn, the, “Bier bier bier bieer,” and a couple other fun things, there's definitely—it would be a nice little stress release at the end of the day, just, “Bier bier bier bieer.” Or, you know, make a short loop, some beats, but have some fun keys and fun switches on it instead.

Corey: Yeah. That's a fun hobby, I enjoy playing around with it. And for better or worse, it's not the end of days type of hobby, where it's, “Oh, great. This new thing came out and it cost me $20,000.” You can get started in this space for 50 bucks. It's not something that has an incredibly high barrier to entry. It lets me play with my hands and fool myself for a little while into believing that I'm making a physical change in the world.

Courtney: Yeah. No, I definitely feel like I've done something. I mean, last year was really strange in that we had to have several appliances replaced in our house at the same time, including the hot water heater. And this was just as I was really starting to get into some advanced—more advanced soldering types of things, and the plumber who was putting in the water heater was like, “I have to solder this new pipe this lead pipe into the water heater that goes into your basement.” And I was like, “Did you say soldering?” And that's when I learned how soldering really works. [laughs]. When you're having to do that to solder copper pipes into someone's home for a hot water heater. And I was like, “Ah. What I'm doing is child's play compared to the actual real world application of soldering.” But he was kind enough to let me run the torch a little bit. So, that was fun.

Corey: [laughs]. It's nice to have something that's a little bit less staring at a screen, especially in this era of lockdown. But it can be more than a hobby and into the territory of problem if you're not careful. So, I've sort of gotten out of it for a little bit. So, I think that's probably a good point to wind up leaving it. If people want to hear more about what you have to say, where can they find you?

Courtney: I get saucy on Twitter and other forms of social media @cjwilburn. I'm active on [00:31:59 unintelligible]. I'm not as snarky and fun as you are, but I certainly will give an opinion or two about anything, or if you want to hear a lot about anything related to Prince and his music, another place to look at that as well.

Corey: You are my new go-to for that.

Courtney: Oh, awesome. Well, I may end up getting on your nerves. I talk about Prince, maybe as much as you talk about the Cloud. [laughs].

Corey: Excellent. You know, we all need things to focus on and I think that's as valid of a topic as any other. Thanks so much for taking the time to speak with me. I appreciate it.

Courtney: It was my pleasure.

Corey: Courtney Wilburn, engineering manager at Elastic for Cloud SRE tooling. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on your podcast platform of choice, whereas if you hated this podcast, please leave a five-star review on your podcast platform of choice along with a comment that you angrily type out, and then tell me what kind of keyboard you used to type it.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Jez Humble
Jez Humble is co-author of several books on software including Shingo Publication Award winner Accelerate, The DevOps Handbook, Lean Enterprise, and Jolt Award winner Continuous Delivery. He has spent his 20 year career in software tinkering with code, infrastructure, and product development in companies of varying sizes across three continents, including working for the US Federal Government’s 18F team as part of the Obama Tech Surge, and co-founding startup DevOps Research and Assessment LLC, which was acquired by Google in December 2018. He works for Google Cloud as a technology advocate, and teaches classes on agile software engineering and product management at UC Berkeley’s School of Information.

Links Referenced

  • DORA: https://cloud.google.com/devops/
  • Cloud.gov: https://cloud.gov/
  • NIST Special Publication 800-145: https://nvlpubs.nist.gov/nistpubs/Legacy/SP/nistspecialpublication800-145.pdf
  • The Phoenix Project: https://www.amazon.com/Phoenix-Project-DevOps-Helping-Business/dp/1942788290/
  • The Unicorn Project: https://www.amazon.com/Unicorn-Project-Developers-Disruption-Thriving/dp/1942788762/
  • Continuous Delivery: https://www.amazon.com/Continuous-Delivery-Deployment-Automation-Addison-Wesley/dp/0321601912/
  • The DevOps Handbook: https://www.amazon.com/DevOps-Handbook-World-Class-Reliability-Organizations/dp/1942788002/
  • Accelerate: https://www.amazon.com/Accelerate-Software-Performing-Technology-Organizations/dp/1942788339/
  • Twitter: https://twitter.com/jezhumble
  • LinkedIn: https://www.linkedin.com/in/jez-humble/

Transcript
Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Catchpoint. Look, 80 percent of performance and availability issues don’t occur within your application code in your data center itself. It occurs well outside those boundaries, so it’s difficult to understand what’s actually happening. What Catchpoint does is makes it easier for enterprises to detect, identify, and of course, validate how reachable their application is, and of course, how happy their users are. It helps you get visibility into reachability, availability, performance, reliability, and of course, absorbency, because we’ll throw that one in, too. And it’s used by a bunch of interesting companies you may have heard of, like, you know, Google, Verizon, Oracle—but don’t hold that against them—and many more. To learn more, visit www.catchpoint.com, and tell them Corey sent you; wait for the wince.

Corey: nOps will help you reduce AWS costs 15 to 50 percent if you do what tells you. But some people do. For example, watch their webcast, how Uber reduced AWS costs 15 percent in 30 days; that is six figures in 30 days. Rather than a thing you might do, this is something that they actually did. Take a look at it. It's designed for DevOps teams. nOps helps quickly discover the root causes of cost and correlate that with infrastructure changes. Try it free for 30 days, go to nops.io/snark. That's N-O-P-S dot I-O, slash snark.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Jez Humble who's currently a developer advocate at Google, but is already very well known in the continuous delivery DevOps spaces. And, I guess—well, what would you say you're most known for? Co-founding DORA's got to be up there.

Jez: Thanks very much for having me. Pleasure to be on the show. Yeah, co-founding DORA, the State of DevOps Research Program, which [00:01:57 unintelligible] Nicole Forsgren and I worked on for many years. Yeah, I've written some books. That's probably the biggest stuff.

Corey: Yes. And then you've got Googled, by which I mean, acquired—not the common use case these days of being turned off.

Jez: [laughs].

Corey: In fact, at the time of this recording you’re in the process of coming out with more research. Which was a question. Whenever a large company buys a smaller one—I don't mean to pick on Google unnecessarily on this—it's, “Oh, great. Are we not going to see any of the things that we'd loved out of those folks now, ever again?” And it turns out that oh, in fact, we are.

Jez: Yeah, that's right. So, one of the things that I've been working on recently, and by the time you—the audience—hear this, this will be fully available, is making all the things we learned in DORA publicly available, to the greatest extent possible. So, there's going to be a bunch of more write-ups of the capabilities that we discovered as a part of the DORA research program. We have, for example, write-ups on code maintainability, which was new in 2019; database change management, monitoring, and observability; cloud infrastructure, like what it means to actually, really, use the Cloud in a, kind of, modern, appropriate way, and some other bits and pieces; and a quick check, which lets you see how you're doing.

So, everyone who knows our research knows that we have four key metrics that we found to be valid, and reliable, and predictive of organizational performance. You can find out how you're doing in those, and it will recommend to you some capabilities that we think you might find helpful to work on. And you can take another quiz and find out which of those we think should be a priority for you, some suggestions, and also, crucially, that exposes the questions that we ask. So, taking the quiz will give you a good idea of what we found that actually works in terms of implementing these capabilities as well, which I think is valuable information to have.

So, that's all out there, on page of our research program where you can explore the research and all the articles that we've written about how to implement these capabilities, and what the common obstacles are, and what it means to be good at them. So, we could not have done this without Google's resources. You know, it was a three-person company; we were strong enough focused on serving our customers, there's a ton of stuff that we learned, and my job for the last year and a half has basically been putting it out there in a way that is easy to consume and because it's on the Google Cloud website, it's Creative Commons. So, I'm really pleased that we were able to do that.

Corey: As someone who's in a small company, myself—we’re 10 people at the moment. At least at time of this recording. One never knows. Hiring is a thing—one of the hardest things for me to wrap my head around, having come from small companies most of my career, is the resources available at a big company. I never quite understood it until I was talking to someone once, probably later in my career than I really should have figured this out.

It's, “Yeah, what's the point of big companies? There's process and procedures, and you always have to ask permission, and do the other stuff.” And the person I was talking to sat there and looked at me, and she said, “Think about this for a second, Corey. What's the thing that annoys you the most?” And I mentioned some business function or another. And the answer is, “Yeah. Here we have 200 people that do different aspects of just that function.” “Oh, okay.”

And it turns out that when you can sit there and straight-faced, theoretically, write a check with three commas in it for something—which again, most companies don't do particularly frequently, but the fact that you can do that means that the limit of what is possible is radically different. Here, if I want that kind of outcome, I'm pretty much limited to either doing something monstrous to raise money from SoftBank or else holding people's kids hostage. There's really not a good third way to get there.

Jez: [laughs]. Yes, that's definitely a thing. And there's advantages to both. My career has been kind of a whipsaw between huge organizations and tiny ones. I started my career in startups, and then I was a contractor, and then I joined ThoughtWorks which was a medium-sized company, and then I went to work for another startup, and then I went to work for the US federal government, the world's largest bureaucracy, and then a two-person startup, and now Google. So, I feel like I've [laughs] really experienced the whole gamut. And, you know, there's definitely advantages to both. The great thing about being in a small company is that you can do whatever you want. And that's fabulous.

Corey: The weird combination there was your time with 18F, your time in the federal government. As you say, the world's largest bureaucracy. The 18F program, to my understanding, is a two-year term renewable for up to another two, but then you have to leave, which to, I guess, most lifelong government workers, that sounds terrifying, but for tech folks—at least in my case, until I started this company, I've never stayed anywhere past the two-year mark, it's oh, “That sounds wonderful. There's no illusions going into this.” What was that like?

Jez: It was amazing. It was certainly one of the highlights of my professional career. I got to work with some really incredible people who were very mission-focused and very good at what they did. And we got stuff done. It was an amazing experience.

Helping to build cloud.gov was just wonderful. We built a Platform as a Service that allowed other federal agencies to deploy stuff to the Clouds in a seamless, straightforward, easy way. And we demonstrated that you can do continuous delivery with cloud platforms in, possibly, the world's most heavily regulated organization. I still have my NIST binder, where I have a significant number of NIST documents printed off, double-sided, two pages to a piece of paper, and it's still enormous. And, you know, we—

Corey: That sounds like a lot of companies’ DR plans. Hasn't been touched in years, enormous, sits on a shelf somewhere in a binder.

Jez: Yeah, I mean, although it is a living document, extraordinarily, Special Publication 853, which for the NIST junkies out there will know is the list of information security controls you have to implement when you're deploying a government information system. That's on revision four. When I left, revision five was in process. I don't think it's been officially released yet, but it's there. And then not a lot of people also know that there's also a Special Publication 853a, which is how you test those information security controls.

Corey: NIST has a lot of great stuff, but I keep looking for the 800-page NIST document on why brevity is important.

Jez: [laughs]. I mean, interestingly, they do have a very short document which defines what is a clouds, and that's one of the things: that was a document that we used to build our capability, or test our capability of cloud infrastructure. And that's really good. I actually have a huge amount of respect for the NIST stuff. And that one is very good because it defines these five characteristics of what it means to have a cloud.

And I think well under 50 percent of people who say they have a cloud meet those five characteristics, but the ones who do are 24 times more likely to be elite performers, which is a very large number. Even if it's only a correlation, it's still a big number and demonstrates a big impact. So, for all NIST’s issues, that is a particularly fine and striking example of something concise and powerful.

Corey: So, one of the most surprising acquisitions in the last decade was the DevOps Research and Assessment—or DORA—which you co-founded, by Google because historically, Google was always very in SRE, and for better or worse, the Industry seems to view there being a schism, real or imagined, between SRE and DevOps. And the cynical among us—ahem, ahem—assumed, “Oh, okay, Google really doesn't understand DevOps, so they think they're buying it out, then they can kill it.” It turns out, nope, that was not the strategy and not the plan. But why was DORA something that it made sense for Google to acquire?

Jez: So, I can give you my perspective. My perspective is that Nicole and I had this research program, which was the most academically rigorous program into what it means to be good at software delivery. And I should mention though, it's not just Nicole and me. Obviously, Gene Kim played a big part in it. We work with Puppet Labs, originally, on four years of this before it was acquired by Google.

So, there's a whole bunch of people who worked on this, I want to make sure credit where it's due. But it's the most academically rigorous program investigating what it means to be a high performer in software delivery, and people trust it because it's independent, and it's academically rigorous, and we have an extremely large n, which is a big deal in its own right, literally. And so I think that's why it's valuable is because we have good research that tells people how you should actually do this stuff.

Corey: It definitely aligns because as I look at tenets of DevOps—and again, we could have an entire thesis arguing what DevOps is or is not if we wanted to do—but the interesting piece I take away is the more agile, the easier it is to iterate forward, it almost seems like the ‘cloudier’ an environment becomes. It seems like there are a lot of points of alignment between what it takes to succeed in business today, what it takes to develop and deliver software safely, securely, and fastly, and being able to leverage a hyperscale cloud provider. Is that something I'm imagining? Is that something you're seeing too? Is it something much more nuanced than that? Almost certainly.

Jez: So, my two cents on this is, I mean, we find that if you're doing Cloud, right: if you meet those five characteristics that NIST defines, which I'm going to [laughs] fail to remember off the top of my head.

Corey: I've mentioned them in one of my talks. I put them on a slide, and the reason I put them on slides is so I don't have to remember them in conversation. We go to, “Cut to the slide deck, Ted,” and then, problem solved.

Jez: Yeah, exactly. I do actually have them here, just because we have, in fact, released the article on it. So: on-demand self-service, the ability to provision computing resources as needed, and this one, I mean, in terms of pushing my buttons, this is a big one. When organizations buy a clouds, but then developers still have to raise a ticket or send an email and wait weeks for their environment to be provisioned, or to raise a ticket to get some networking setting changed so they can actually have a route to the thing that they want, it's not a cloud.

So, [laughs] a lot of people fail in step one. It's like, “Well, we spent all this money on the Cloud, but outcomes for the people who use that cloud—developers, and people trying to get their software built—are completely unchanged, so what was the point in that?” And then, broad network access, resource pooling, rapid elasticity, measured service. You can go to cloud.google.com/DevOps and go to the research and go to the article, or read NIST Special Publication—see, I can't remember this number. I'm not that good—800-145 for yourself.

But that has an impact on software delivery performance. And what I will say is, I think a lot of the things that let you take advantage of that are things we talked about in the research program: things like being able to version control the configuration of your system, being able to do test automation, being able to build security. And you have all these different capabilities, that's going to definitely give you the ability to take advantage of the things that Cloud offers, which is the ability to stand up an environment based on a configuration you specified, and do testing in that environment, and be able to rapidly deploy software. If you can do all these things that we talked about in the program, that's going to give you an incredible ability to actually use cloud infrastructure in a way that's a force multiplier. So, that's kind of how I feel about that.

Corey: One of the biggest challenges for organizations of almost any size but particularly large ones is, we can read books like Gene Kim's The Phoenix Project, and then The Unicorn Project from other perspectives several years later, and it's easy to see, oh, we should be doing this, we should be doing that, we should align better with this vision of the future. And the problem, of course, is that it's easy to say that, how do you get there? What is the journey look like to go from where you are to where you're going? And, of course, how do you avoid the trap that we all fall into at every company, which is, “Just after this next sprint finishes, we're going to stop making poor decisions and make good ones instead and all the technical debt gets paid off.” Everyone says it, but it's the biggest lie this industry tells ourselves constantly. And we tell it to ourselves in every company, in every environment I've ever seen.

Jez: Yeah. I mean, in the future, things will be better. And there's never been a worse time to think that. And I think the important thing to realize first is that all these books, certainly all the books I've written, they're composites of a bunch of different experiences.

So, there's no one project on which all the things in the Continuous Delivery book were true, or The DevOps Handbook], or Accelerate. None of these books—and I can only speak on my own behalf, but I would suspect the same is true of other books—they’re composites of a whole bunch of different experience, plus some extrapolations, plus some hindsight bias, plus some narrative fallacy. And so what you’ve got to realize is all of us, when we're writing these things, we've been inspired by things that have happened that, we've experienced, and things that have happened that other people have experienced, and you synthesize and you extrapolate, and you're like, “Oh, there's a thing here.” There's a set of principles; we can test those; we can validate them.

But in real life, I mean, the dirty secret, which we wrote at the end of the DevOps Handbook, is that a lot of those case studies where, la, the amazing thing happens, well some senior manager left and then someone new came in, and then it went back to [BLEEP] again. And so I think, firstly, you've got to lower your expectations and realize that even the people who you look up to who are doing great things, they’re living in the real world, too. Google is often pointed to as some kind of amazing, perfect unicorn. Google is a massive bureaucracy like any other massive bureaucracy, and we can't change the laws of physics. And I think it's important to be humble about that.

And the important thing is, it's really boring. It's just like anything else, you just got to do a little bit every day. It's like trying to diet, or trying to get better at anything, or trying to write—the advice from Ernest Hemingway about writing, about making sure you do a little bit every day, which is probably the most boring thing that Hemingway said, but also one of the most important things. It's just that. It's just that discipline of trying to get a little bit better every day. And that's how you do it. And really good organizations give you the time and capacity and resources to do that. But it's hard, and it takes a long time, and it's relentless, and it can be a bit tedious.

Corey: In what you might be forgiven for mistaking for a blast from the past, today I want to talk about New Relic. They seem to be a relatively legacy monitoring company, and I would have agreed with that assessment up until relatively recently. But they did something a little out there: they reworked everything. They went open source, they made it so you can monitor your whole stack in one place and, most notably from my perspective, they simplified their pricing into something that is much more affordable for almost everyone. There's even a free tier with one user and 100 gigs per month, totally free. Check it out at newrelic.com.

Corey: That's the hard part, is people also look at this through a lens of, “Well, okay, I read the book, and I saw all the things we have to do, and I look at what we're doing across the board, and we're doing approximately none of those things. So, okay, what do we do? Well, we have to fire everyone and burn everything down to the ground.” And it's not really—there's a reason that we call this a journey, not a destination. And not everything has to be, I guess, the top of the scale in every category. It just doesn't work that way.

I mean, we take a look at the same idea of how to go multi-region, and being able to handle extreme high availability at global scale but, I mean, take the ridiculous system that builds my newsletter, it's single region in a single AWS region. There's no redundancy there, and there doesn't need to be because if you think about it through a lens of a business context, for example: cool. The entire AWS region is hard down. I'm going to be sending a very different newsletter out that week if that happens, and I think the world will be quite okay with that. It doesn't have a business need for that, so why invest in the additional expense, not just of infrastructure, but of engineering time and effort of getting that monstrosity to work in a more durable way when there's no business requirement for it? Just get better than it is today, and focus on the thing that's material to your business, not the thing that—with all love and respect—some book say is important for you to focus on instead. Because yeah, there's commonality there, but a book that is written for the masses is never going to have particular context into your situation. That's what judgment is for. And it's why robots don't do our jobs… yet.

Jez: [laughs]. Yes, hopefully. That is a worry, but I've consistently been relieved by the fact that it appears that I have a job for life because I'm still talking about things that we were talking about in the 1960s, and still haven't been able to effectively implement. And those are all human things. I mean, I've got McGregor's Human Side of Enterprise on my bookshelf from 1960 where he talks about theory X and theory Y, and that's still a real problem that we face. We just published an article on transformational leadership which is closely related to those concepts, and effective leadership is still really hard.

All the big obstacles to improvement are people-based, human-based, and those are hard problems to solve because a lot of it depends on building an organization where you can effectively delegate to the people doing the work to make decisions. And the larger the organization gets, the harder that is to do because people don't have visibility into what's going on and how to get better, and they are unable to make good prioritization decisions, and they need to work together with other people who may have different priorities, and trying to get people aligned, and trying to give them the resources and the ability to make change and accept that when they do that, they're going to mess it up a whole bunch of times. And that's okay because that's how you learn. Those are all incredibly difficult things to achieve, and no machine is going to ever solve that problem. And we'll be lucky if any human solves that problem, frankly.

Corey: One of the things that I'm a big fan of is the Dunning–Kruger effect. And I think that that is entirely malign when we look at how it is commonly applied. Because people generally understand it: it's this idea of, “Oh, people who are ignorant of an area don't even know how much they don't know, so they think it's easy.” What I think people miss is that this applies to everyone in various fields of endeavor. Just because you know this thing exists does not mean that you're not susceptible to it.

I mean, I was as guilty as anyone, but when I first started my company aimed at fixing AWS bills, I was worried that AWS was going to fix all the problems with the bill in the next six months, and then I wouldn't have a business anymore. Now that I know what I know, my big concern is that they're never going to be able to fix this because it is such a complex and deep problem that I might not ever get to go focus on other problems I find more interesting as I walk through the world. There is such hidden depths of complexity behind almost every area that you have to constantly remind yourself—at least I do in my case—is whenever I look at something I don't understand well and think, “That doesn't seem hard,” that is almost certainly the whispers of Dunning–Kruger in my ear and it's, you're basically hand-waving over entire fields of expertise that is not being accomplished right now by people who are bad at things. Maybe dig deeper.

Jez: [laughs]. Yes. And I think Silicon Valley is one of the worst when it comes to the Dunning–Kruger effect. You'll never find people more willing to wade in on topics that they're supremely unqualified to pronounce on then a bunch of Silicon Valley tech bros who have solved one problem and think they thus have the ability to solve every problem. And that's something I constantly work on myself because, you know, I'm as susceptible to that as everyone else.

So, fortunately, I have people in my life who tell me to shut the [BLEEP] up, and that's a very valuable service. And I think you're in a very lucky position with The Duckbill Group. I love this quote from Upton Sinclair. “It's difficult to get someone to understand something when their salary depends on them not understanding it.” That's where you are with The Duckbill Group. I don't think Amazon will ever solve that problem, and you're in a great position to solve it. And it's the same with getting people to improve the way they develop and deliver software.

Corey: It's gone through weird iterations because at first, again, naively I thought, “Oh, I'll come in and I'll fix the billing problem they have and save a bunch of money, and there we go, I'm done. And it's a shame I'll never have a second engagement to sell those people.” Yeah. Turns out it doesn't work that way. A) find me a single cloud environment that is static from month to month, where the entire team hasn't disbanded and someone forgot to cancel a credit card. It just doesn't happen.

Jez: No.

Corey: Environments continue to evolve. And it's also not just about lowering the bill—people believe it is—it's about understanding and predicting it. It's about moving up a maturity model of going from—on one end you have finance getting a giant bill from Amazon, and their first question is, “Wow, how many books is engineering buying?” And there needs to be an education there.

And on the other end of the spectrum, it's, “Cool. As we wind up defining our bill every month by a function of users that we have, how do we wind up accurately predicting that over an 18- to 36-month span?” And again, this is not specific to Amazon, these problems. It's just where I focus because, surprise, specificity in marketing really the right answer. If I'm all things, all clouds, to all people, great. Swing a dead cat, you look like a generalist. And that's not an expensive problem that people bring consultants in for. Be specific, solve a problem.

Jez: This reminds me of my funniest experience working at 18F was when I was sat next to a contracting officer talking about an AWS buy that we were trying to get through. And I was basically explaining the problem with the Cloud, there's two axes of variability. There's the resources you decide to buy, which is somewhat under your control, and then there's how much end users decide to use those resources, which is really not under your control as much as you would like. And the contracting officer was, kind of, looking at me as I was having this discussion, and he said, “Let me explain how contracting works in the US government. We put an order in for 100 sandbags, and then we get 100 sandbags.” And he just sat there in, kind of, silence, paused for a little while. And I was like, “Okay. We're going to have to go back to the basics here.”

And, you know, we had a discussion about mobile phones and how mobile phone bills are variable. And he was like, “Yes. This is why we have fixed-price contracts for mobile phones, which is all you can eat.” And you just fundamentally can't do that with the Cloud because of that second-order effect. So, I ended up putting together this crazy spreadsheet, where I basically listed all the things we had consumed and then did some kind of moderately complex statistics with two orders of standard deviation to try and put a buffer in.

But, I knew it was all crap because consumption of resources doesn't follow a normal distribution, and so it was all just padding, and you can't make predictions based on power law, which is Nicholas Taleb’s entire career in a nutshell. And also because what that forces you to do is just spend a large fixed amount of money on something which is probably going to be ending up much cheaper, but that won't end up much cheaper because you've got to buffer it because you want to fixed price. And that's fundamentally what you're dealing with, and it's incredibly problematic.

Corey: It turns out that one of the hardest parts in all of business and all of computing is not—contrary to popular opinion. For example, “Money is always a limiting constraint.” Not really. Look at companies like SoftBank walking the world. You can always find money from someone who has no clue what's going on.

But in fact, that is the actual constraint. It's a lack of understanding. And we all suffer from it in different arenas in different areas, and the thing I never really appreciated until I started working with large enterprises is in a small company, the people I talk to as I go through a consulting engagement are effectively the same person at a small enough company, the person who signs the check, the person who champions bringing me in, the person who benefits from what I wind up doing, great. At an enterprise, the organizational distance between those various people is sufficiently great that no one person anymore can hold the entire problem space in their head, and that's where running and managing teams is incredibly important. Good departmental communication becomes critical.

And my old-school startup approach is very cynical around that. It's, oh, just hire smarter people who can hold it all in their head. Yes, how scalable. Brilliant. Big companies are their own scale, and I never appreciated that. Like, when I started making fun of Amazon in the newsletter, I thought I was going to upset Amazon as if there's someone over there named Steve Amazon that is going to get upset by the things that I say that informs an entire company's opinion of me. Yeah, turns out companies don't have single people who make opinions for the entire environment. It's different groups, different people. Some people love me, some people hate me, but most of them—far and away—do not know who I am.

Jez: Yeah. That's definitely humbling. And the good news about that is that's why it's okay to give the same talk multiple times in different places. So, there is a benefit from that.

Corey: Oh, absolutely. I mix it up a little bit. I try and change the titles so people may not necessarily realize it's a version of the same talk, but also when you suck at preparation—hi—everything that you do, in turn, becomes an exercise in improv. So, yeah, I have the same bullet points, but I'll give the same slide deck twice with radically different results, radically different storytelling approaches, just because, well, I didn't write it down the first time. I guess I'll be coming up with this from whole cloth again. It's fun. Not that I recommend this as a great speaking technique for most people. It's sort of me working around my own shortcomings.

Jez: Well, I think that's an excellent point. And that probably is the major driver of my career as well is working for my own shortcomings, and finding good strategies to deal with my personality defects.

Corey: Well, thank you so much for taking the time to speak with me today. If people want to hear more about what you have to say what you're up to, find you in order to cast various slings and arrows in your direction, where can they track you down?

Jez: So, I tweet at @jezhumble. I am Jez Humble pretty much everywhere on the internet, so have at it. Definitely don't use LinkedIn because nothing is more annoying than unsolicited LinkedIn messages.

Corey: Oh my stars, yes. Like, it's always weird to me—I do want to ask you that one: that's a fun sidebar for there. My approach to LinkedIn has always been that if I've met with someone, and I've had a conversation with them, and I could pick them out of a police lineup—but hopefully won't have to—I will connect to them on LinkedIn. But, “Hey, you gave a talk to a room of 300 people, I'd like to add you on LinkedIn now.” For me, at least, and maybe I'm conceptualizing this in the wrong way, but if I am connected to someone on LinkedIn, I would need to feel comfortable sending them a message asking them for an introduction to someone they know. And if, “Hi, I gave a talk eight years ago, and you were in the room when I gave that talk. Can you introduce me to your boss?” Some people are comfortable asking stuff like that. I’m really not. So, for me, at least, I tend to curate my LinkedIn connections relatively carefully. It seems that most people don't take that approach. So, I have to suspect I'm the one who's wrong.

Jez: I mean, the only thing keeping me from deleting my LinkedIn account right now is my own ego, which unfortunately is a substantial obstacle.

Corey: [laughs]. Okay, I find it reassuring that people who I respect in this space are not going to look at my approach and think it's completely out to lunch.

Jez: Well, honestly, here as in many other ways, Kelsey Hightower has led the way by, in fact, deleting his LinkedIn account. So, just one of the many ways in which we can aspire to be more like Kelsey.

Corey: Oh, that sounds like the promised land where I get to wander the desert for forty years, but I am not allowed to go in.

Jez: [laughs].

Corey: Thanks so much for taking the time to speak with me today. I really do appreciate it.

Jez: It's an absolute pleasure. Thank you so much for having me on.

Corey: Jez Humble, developer advocate at Google. I am Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts, whereas if you've hated this podcast, please leave a five-star review on Apple Podcasts and a comment including the entirety of at least one NIST standard.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Ceora Ford

Ceora Ford is a digital marketer turned software engineer based in Philadelphia. She is really into Python, AWS, education and diversifying tech. She's had the pleasure of teaching with Kode With Klossy, BSD Education, and more recently, egghead.io. When she is not coding, she's usually watching movies and pretending to be a film critic.

Links

  • egghead.io: https://egghead.io/
  • Eight Resources for Learning Python blog post: https://www.ceoraford.com/posts/8-resources-you-can-use-to-learn-python/
  • Twitter: https://twitter.com/ceeoreo_
  • Personal Blog: https://www.ceoraford.com/
  • LinkedIn URL: https://www.linkedin.com/in/ceora-ford/
  • Ceora’s Website: https://ceoraford.com/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Catchpoint. Look, 80 percent of performance and availability issues don’t occur within your application code in your data center itself. It occurs well outside those boundaries, so it’s difficult to understand what’s actually happening. What Catchpoint does is makes it easier for enterprises to detect, identify, and of course, validate how reachable their application is, and of course, how happy their users are. It helps you get visibility into reachability, availability, performance, reliability, and of course, absorbency, because we’ll throw that one in, too. And it’s used by a bunch of interesting companies you may have heard of, like, you know, Google, Verizon, Oracle—but don’t hold that against them—and many more. To learn more, visit www.catchpoint.com, and tell them Corey sent you; wait for the wince.

Corey: nOps will help you reduce AWS costs 15 to 50 percent if you do what tells you. But some people do. For example, watch their webcast, how Uber reduced AWS costs 15 percent in 30 days; that is six figures in 30 days. Rather than a thing you might do, this is something that they actually did. Take a look at it. It's designed for DevOps teams. nOps helps quickly discover the root causes of cost and correlate that with infrastructure changes. Try it free for 30 days, go to nops.io/snark. That's N-O-P-S dot I-O, slash snark.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Ceora Ford, who is among other things, a software engineer, a digital marketer—historically—and an educator at egghead.io. Ceora, welcome to the show.

Ceora: Thanks for having me. I'm so excited.

Corey: So, as of the time that we're recording this show, it is before I go out on parental leave, but when people are listening to this, you will have put out this week's issue of Last Week in AWS, so people should at least have some passing familiarity with who you are, at least among this audience. So, it's interesting having conversations in the past from when people listening to this about something that is yet to happen. So, we're assuming the newsletter goes out on time, that 2020 hasn't gotten worse because oh, wow, they haven't even mentioned the giant meteor, yet. It’s… it turns into interesting times. But first, I want to start by thanking you for letting me take some time off and actually spend time with my family rather than writing these things part-time in between bottle changes.

Ceora: Of course. I'm so excited that I get to have this opportunity.

Corey: The way that I see it, though, is it's always a problem where I started this whole ridiculous newsletter once upon a time where it was this labor of love, it was built on my crappy personality defects—which people misinterpret as a sense of humor—and then turned into something that sort of caught on. But I never built it with the idea of other people being able to take it over which, pro tip, if you're building something and trying to scale it, maybe have a succession plan in place. That becomes a bit of a challenge.

And it was, “Well great. What am I going to do?” Well, now I'm experimenting and finding out. If people aren't listening to the show by this point, great. Good to know: now at least I've learned that, nope, it dies with me; good to know. I don't think that's going to be the case, but we'll find out. So, enough about that nonsense. Let's talk about you. What are you doing these days?

Ceora: Okay. Ah, I feel like I have a million and one things going on, always. I'm one of those people that as soon as I feel like I have a chance to relax. It's like, oh, let me find something else to do. So, right now, I am working on a couple different projects.

So, one of the things I'm doing is, I'm looking for a lot of speaking opportunities. People have been passing on a lot of information about different conferences and things like that, and now since they're remote, I have a lot more availability for different things. So, expect more of that for me as well, more conference talks and things like that. I'm also working on completing my Udacity course and officially becoming a Cloud DevOps engineer. So, probably by the time this is published, I'll be finished with that—hopefully.

And then I have several different projects on the side that I'm working on. One that I just started recently is called “100 Days of Projects.” So, it's a community-based movement, I guess you could say, that just encourages people to work on their projects consistently. It doesn't matter how much you do or how little you do, as long as you're taking the time to put forth some effort on the projects that you build.

I feel like, myself included, a lot of us developers have tons of project ideas and not enough time or, you know, sometimes we just forget about them. And I am in that category so much, so I decided to start this project that will help me and others get more involved in building your projects consistently, and learning in public as you do it. So, yeah, those are some of the things that I have going on, I probably missed a few things because like I said, I just have so much happening right now.

Corey: Oh, believe me, I understand the feeling. I love the idea of learning in public. That's something that, for whatever reason, when I was starting out in my career I never felt like I could do that. It was always a learn, study in your own time, cram, so whatever happens, your employer never finds out that you’re a giant fraud. And I figured that I would wind up being able to stop doing that when I stopped feeling like a giant fraud.

Fifteen years later, still waiting for that moment to hit. So, I've had to, I guess, readjust that approach. I mean, something I've actually learned is that when you're interviewing people for various roles, a mark of seniority is when people admit they don't know something. It's very hard for people to do, for whatever reason.

Ceora: Oh yeah.

Corey: So, whenever I see people doing the learning in public thing, like you do, I'm generally incredibly impressed. It's something that I wish I had had the courage to do back when I was starting out.

Ceora: Yeah. Well, I have, fortunately—especially through my work with egghead—I have been introduced this concept a lot lately. People at egghead are really big on pushing you to learn in public and, even if you're a beginner, create content for others. So, I totally understand your hesitance because I felt that way for the longest. I was always scared that I'm going to be exposed as a fraud or, you know, someone's going to pick out one of my mistakes and hate me or something. Like, coming up with all these different scenarios.

And sometimes, yeah, you do get reply guys who were like, “Oh, well, actually, it's not like that.” But for the most part, I feel like learning in public is great for so many reasons. Especially—I have the personality type where I need external sources of motivation. So, external sources of motivation for me are other people watching me, knowing that I’m—like, expecting something from me. So, if I'm public about, “Oh, I'm learning, I'm trying to work on getting this AWS certification, or I'm working on this project,” and I tweet about it, I share on my blog, or whatever the case may be, I know that people are watching, and they expect something; they expect the finished product.

So, it pushes me to actually do things and finish things. And then also, when you are learning in public, sometimes some forms of that are basically you're teaching other people what you're doing. So, if you're writing an article, if you're streaming on Twitch, or whatever the case may be, you're teaching information that you just recently learned. And I feel that's a way to spur your own understanding exponentially. Because if you have to teach other people, it forces you to really make the ideas and concepts your own, and it forces you to see, like, “Okay, if I'm a beginner and I'm completely new to this concept—” whatever the case may be—you're going to be looking at everything in a much more in-depth point of view, which I think really helps you to learn and solidify your knowledge much faster.

So, I'm a huge fan of learning in public, I wish I would have done it earlier, just like you were saying, but I'm trying to get over to a fear of failure, or imposter syndrome, or whatever the case may be, and just like put myself out there a little bit more because in the end, it ends up being not only beneficial for me but other people who are beginners or other people who are looking to learn something new that I'm also learning. So, yeah.

Corey: One thing that continues to take me aback is people in your position who are looking to, “All right, I'm going to start learning tech,” and one of the technologies and areas of focus that you've talked about is going into learning about AWS, which is, “Okay, that's daunting. So, what's your goal?” “Oh, I'm going to learn all about AWS.” “Yeah, I'd like to do that someday, too.” There's no outcome there. It's one of those looking at how stupendously complicated, and overwrought, and vast just that one company's product offering is. I despair of ever being able to attain mastery over even a small part of it.

Ceora: Oh, yeah. I totally agree. AWS is so… it can seem so daunting, and honestly, I'm not even sure what even convinced me that, “Oh, yeah, this is something I want to tackle,” because now looking back, I'm like, “What was I thinking?” Because it's so—okay, so for me, I don't have an IT background or, like, a networking background, or even a back end engineering background. So, a lot of those concepts, I've become introduced to through AWS, which is great in a lot of ways, but also because AWS is such a huge thing within itself, and sometimes—I've talked about this on Twitter before—but sometimes AWS docs are not very helpful. So—[laughs].

Corey: No.

Ceora: [laughs]—so sometimes I've been so confused about certain concepts that I've been introduced to through AWS. But the thing I am thankful for is, even though AWS is not always the most, you know, welcoming way to be introduced to some of these things—especially, like, networking and stuff like that—I don't think I would have learned about those concepts otherwise. I probably would have just stuck with the basic—not the basic. I don’t want to say basic—but I probably would have just stuck with, you know, front end web development, those kind of things.

But AWS has helped me to broaden my horizons and learn a lot more, which I'm happy about. But yeah, it's so much packed into one product, I guess, that it can seem very daunting. It's something that, it's a valuable skill to know but, like you said, there's no way—if you're one of those people who, “I need to master this, and I need to know everything about it,” you'll be reaching for that for the rest of your life with AWS.

Corey: Oh, yeah. But what amazed me when I started getting into deep into the space is I’ll talk to someone at AWS who is an absolute wizard, when it comes to networking, for example, and how AWS networking works because they work there, and they help build part of it, and it's, “Yeah, I am a giant fraud. I am never ever going to attain anything approaching this person's level of expertise.” And then I'll ask a question about Elastic Beanstalk, and their response is, “Elastic what now? Is that a real thing? Are you making that up to have fun at my expense?” And it's? “Oh.”

I keep forgetting that my failure mode is that I come from a world of small companies where a big company has 200 people. The idea of a company that is so vast that people don't know what other parts of it are even supposed to be building or working on is ridiculous to me, but that's the world we live in. I’ve long since passed the point where I can talk convincingly about services that don't really exist and not get called out on it when I'm speaking to large groups of Amazonians. It’s… no one has this all in their head.

Ceora: Yeah. Yeah, I totally agree. I totally agree with you. So, it's one of those things. I guess this kind of ties back to learning in public, too. When you're learning AWS, you cannot feel bad about not understanding certain things because no one understands everything about AWS. Nobody. So, it's kind of the perfect thing to just share that you're learning because all of us are just figuring it out, honestly. No one's an expert.

That's not something—I learn new things about AWS every single day. And admittedly, I haven't been in the space for long. But still, that’s still major. There are some people who kind of reach a point with certain frameworks or languages where they, kind of, reach expert level, but I haven't met anybody yet. Who's like, “Oh, yeah. I'm definitely considered an AWS expert,” because nobody feels that way. We all know there are things that we just don't know, and probably will never figure out.

Corey: One of the transformative moments I've had is I was getting coffee back in the before times when that was a thing we could do without taking a deadly risk. And I was talking with Jeff Barr, the chief evangelist at AWS, who writes, I think at this point, he is approaching 3500 blog posts published.

Ceora: Wow.

Corey: And I asked him about this because I've started to experience an aspect of this. I don't write nearly as prolifically as he does, but there are times where I will Google something and find the answer. And, “Oh, great. This is an awesome article. Who wrote it?” It was me, and I have no recollection of writing it.

And it turns out that his answer was, oh, as soon as he writes something, he tends to mostly put it out of his head. So, if I come and I asked him a technical question about a blog post he wrote, he will—“Sorry, what service is that? Is that a real service? I guess it would be.” It's one of those moments where you can't retain all this in your head. That is why we write things down.

And one of the things that you mentioned a few minutes ago that I really liked, was the idea of teaching something as a way to learn it. And that really resonates. For me it was I built a conference talk, “Terrible Ideas in Git” because when you get a conference talk accepted and you have to give it in four months about a technology you know almost nothing about, that is what we know in computer science is a forcing function. You're going to learn it, or you're going to give a really crappy talk. And, do I know Git now? Of course not. Git makes everyone feel dumb. The only question is how far along the path you get before the floor drops out from under you. But by building that talk and making it accessible, I understood a lot more about it than I did when I started. There's really something to be said for teaching others as a learning style.

Ceora: Oh, yes. Oh, for sure. Actually, what you're saying right now about giving a conference talk about something you're not very familiar with, that's actually something I've been doing recently. And again, like I said, I need external sources of motivation to do anything, pretty much, at this point because I'm not a very self-motivated kind of person. Like I would lie to myself and say, “Yeah, I'm going to do this thing, and I'm going to do it so well.”

But if I don't tell anyone else about it, it's basically not going to get done. So, that applies to almost every part of my life, especially the tech side of my life. And it seems so bizarre. When you see someone giving a conference talk, you automatically assume that this person must have years and years of experience, and they're an expert. And they know everything, and they're just like, “Oh, they're so confident and that's why they're giving a conference talk.”

I was asked recently to be on a panel for Jamstack which is something not super related to AWS—kind of with the serverless side of things—but at that point, I knew about it, but I didn't really know enough about it to, really, you know, be on a panel. But, you know, I was given the opportunity, so I was like, “Okay, I'll accept it.” And that was the first time I ever confronted this idea of these things that we assume only experts do, making videos, or writing articles, or even giving a conference talk, you do not have to be an expert. I repeat: you do not have to be an expert to do those things. In fact, it's a great way to learn enough to talk about something in a knowledgeable way.

I feel like that's my new learning tactic, is, like, if I want to learn something, the next thing I automatically plan is, okay, I'm going to write this article, or I'm going to sign up to give this talk about this subject so that it's going to force me to learn enough about it to be able to talk to people and know what I'm talking about to a certain degree. So, yeah, I love this idea of you don't have to be an expert. In fact, not being an expert gives you a different perspective that probably is really useful to a lot of people. So, yeah, I love the idea of, yeah, give that talk. You don't know anything about this subject? You don't know anything about AWS? Well, use it as a way to learn about it. So, that's something I've definitely been doing recently, especially, I mentioned earlier, I'm trying to do more speaking opportunities and things like that. So, that's what I've been using as a way to learn and also get more speaking opportunities as well.

Corey: If you're listening to this and looking for a speaker at your events, you should be paying attention to this.

Ceora: [laughs].

Corey: This episode is sponsored in part by our friends at Linode. You might be familiar with Linode; they've been around for almost 20 years. They offer Cloud in a way that makes sense rather than a way that is actively ridiculous by trying to throw everything at a wall and see what sticks. Their pricing winds up being a lot more transparent—not to mention lower—their performance kicks the crap out of most other things in this space, and—my personal favorite—whenever you call them for support, you'll get a human who's empowered to fix whatever it is that's giving you trouble. Visit linode.com/screaminginthecloud to learn more. That's linode.com/screaminginthecloud.

Corey: You had a post on your blog back in May, “Eight Resources for Learning Python.”

Ceora: Yes.

Corey: And reading through it. I've tried most of these resources myself, and I have come to the conclusion that they are all crap, for my way of learning. And that's what it comes down to. Giving a talk for a conference is a great way that works, for me, though, the best way to learn something like a programming language, I've tried classroom courses, I've tried videos, I've tried books, the only thing that works/sticks is me building something with it. And it is not a particularly efficient means of learning things. I

If I was capable of paying attention and absorbing something in a structured sense, I wouldn't spend three hours Googling the difference between a string and an int when I'm trying to diagnose some arcane error. It's one of those effectively—“Oh, cool. So, what's the secret to wind up building a lambda function? Oh, you just go ahead and iterate forward and every time you do a new deployment, it winds up incrementing the version number.” Why are there 5000 versions of this lambda function? Iterative development.

It's one of those, “Oh great. Typos. Did you know that there are editors for writing code that will do syntax checking? Today I learned.” It comes down to getting it hilariously wrong and iterating forward, for me is the right answer. And people look at this, like, oh, some people be called a self-taught learner but it appears you are never taught, period. But it's like, “This is the worst run code I've ever seen.” “Ah, but it does run.” Yeah, that's the—it's awful and I don't recommend that, but it's the only thing that I found that works. And I am incredibly envious of people who are able to learn without breaking things in hilarious fashion.

Ceora: Well, I have an interesting perspective on this because I'm going to get really nerdy right now. But I am a language—like spoken language. I'm not talking about coding languages. Spoken languages are very, very interesting to me. And when I was 15, I made this life goal that I'm going to become a polyglot and I'm going to learn six different languages.

That never happened, but one of the interesting concepts in the language learning space is this idea that you learn enough of the language to hit a certain goal. So, for instance, if your goal is to go to France to go visit Paris for two weeks, you're going to learn enough French to get you through those two weeks. Or if your goal is to like, oh, I'm going to teach at a university in Spain, say you’re a biology professor, you're going to learn all the Spanish biology terminology so that, in that space, you'll be able to have full-fledged conversations, but maybe outside of that, not so much. So, I kind of apply that to—that's how project-based learning is to me. You learn enough to get the job done, and you may not be super knowledgeable about other things involving the language.

You know, if you build something in Python, you might not need to know about tuples. Or you might not need to know about lambda functions in Python and things like that, but you will be able to get something done that works. So, in some ways, it can be very useful which, to me, I love learning in public, and I love project-based learning because that, to me, is super important. And it works for me as well, but it also, in some ways you do miss out in some context sometimes. But I don't think it's that big of a deal. If you're trying to build something, Google is going to be your best friend. If there's something you want to do, Google it. And that's what I do all the time.

Corey: To be clear, you say Google it, you're talking about looking something up in an online search engine—

Ceora: Oh, yes.

Corey: —not turning off something beloved that people have come to depend upon?

Ceora: [laughs]. Yes.

Corey: Excellent.

Ceora: Yeah, that's what I mean. So, in the context of Python, that's probably the language I probably learned in, you know, air quotes, “The traditional way,” taking courses and reading articles, but I've also done a lot of building with it, and breaking stuff and figuring out things in a certain context. So, yeah, probably a combination of those things. But like you said, project building is the best for you. And I think it's important for people to know that everyone learns differently. There is no linear path to anything in tech, really. Everyone is different. Some people do best with having a CS degree. Good for them. Some people do best going to a boot camp, some people—

Corey: I have an eighth-grade education. I was never the book learning type, it doesn't work for me. Also turns out that if you learn only by screwing it up seriously, maybe neurosurgery is not your field.

Ceora: [laughs]. But, like, with tech—

Corey: What do I do next, be a pilot? Yeah, [laughs] doesn't go well?

Ceora: No, no, no. But that's [laughs] not what we're saying here. But with tech, learn, however you want. There is no direct formula for how to become a software engineer or a cloud engineer, or whatever your goal is. So, yeah, I hope that people become more open to this idea that—you know, there are some thought leaders who are like, “You have to do this, and you have to learn this first, and then do that thing, and then read up on that.” Like, no.

Do whatever works for you, honestly. Do whatever gets you to your end goal. Just like with learning a language, sometimes the way that people do it is they absorb a whole lot of words, and some people just have an end goal; they're taking a trip, they want to be able to hold basic conversations, so they learn all the vocabulary they need to know for that. So, whatever your end goal is, do what you need to do to get there. This is how I feel.

Corey: So, something you mentioned a few minutes ago was that you were doing front end work for a while, and then moved into Python. There is a perception on Twitter that is crappy and wrong. That oh, front end is easy, back end is the hard stuff and it's where all the good engineers go. Well, I don't pretend to be a good engineer, but I can beat things together in Python for a back end moderately well, I can effectively take this beautifully crafted precision screwdriver set that is Python and use it as a tremendously crappy hammer. But what I'm not able to do is understand the first thing about front end, or JavaScript, or whatever it is that makes that whole stuff work. I have tried repeatedly, and I end up more confused than I did when I started. A week in, like, “What the hell is ‘asynchronous’ mean, and why is it doing this?” Yeah, doesn't go well. So, my answer for how I wind up handling front end has always been I pay a professional because I find it completely lost. It is, from my perspective, way harder than back end.

Ceora: Oh, yeah.

Corey: What is wrong with that entire misperception?

Ceora: Okay. Actually, I kind of have a few different philosophies on that. Because everyone thinks that front end well—not everyone. But if you search, oh, “Learning to code,” a lot of the boot camps or articles from these tech thought leaders will encourage you to start with front end, HTML, CSS, JavaScript because there's this perception that is easier like we were just saying. So, when I first decided to learn to code, that's the route I was going to take because that's what everyone was telling me to do.

You know, learn HTML, CSS, JavaScript, it's easier. It's a great way to start things out. It was not easy for me. I struggled. It took me so long to get the basics of CSS down, and then when I finally got to JavaScript, I became even more confused and it made me stop. I got to DOM manipulation and was so freaked out that I just was like, “You know what? I'm going to take a break,” and that break ended up lasting for six months.

So, what I decided to do was I'm going to try this back end thing. I had just gotten a scholarship to the Cloud DevOps Udacity program I was just talking about, and one of the things they mentioned was Python for some reason. So, I was like, “Okay, I have this scholarship. I'm going to get into the cloud thing, and then I'm going to learn Python. Maybe this is going to match my brain better.”

So, I decided to learn Python. And I loved it. It made me fall in love with coding again. And I had this idea that the reason why people view front end as easier is because it's very visual. It's a lot of design aspects to it, especially with the HTML, CSS, it's a lot of, you know, colors and all that kind of stuff.

And I feel like a lot of people have this perception that anything design-wise is easier, and that anything remotely artistic is kind of feminine, as well. So, automatically, anything—sometimes for certain people—that's perceived as feminine is, oh, well, that must be easy. Drawing and making animations with CSS and JavaScript is easy. Or creating a front end for e-commerce store is easy. And it's not. It’s not easy. It's the farthest from easy. Like, it’s no.

Corey: It’s easier to learn in public, even, as a white guy for God's sake. If I wind up getting something wrong, my failure mode is a board seat and a book deal somewhere. It's absolutely, effectively, I am playing life on the easiest of easy modes.

Ceora: [laughs]. So—

Corey: Ah.

Ceora: Yeah. One of my first introductions to coding was through an organization called Kode With Klossy, and one of our instructors was saying how, I believe in one of his CS courses, they pushed all the women to do web design because web design is seen as, like, artistic and, “Oh, yeah, you design some things.” And I guess that has a feminine flair to it, and so it's automatically perceived as easier. And I'm like, “No. It's so not easy.”

Like, I wish that we would get rid of this rhetoric. Just demolish it for good because it's really not easy to me at all. And not that anything in tech is easy, but we need to erase this idea that, “Oh, yeah. You know, back end engineering is for the real software engineers. You're not a real engineer unless you do back end stuff.” No. I don't believe that at all. Give all the glory and the praise to front end engineers because they're doing the thing I could never—I only touch front end when I absolutely have to at this point. So, yeah.

Corey: One of the strangest things that I see is that I look at the body of things that I understand and have learned how to do, and my instinctive response that is, oh, that's easy. Whereas all the things I don't know, well, those things are hard. And this is absolutely not correct in any meaningful sense. However, this is a very common thing where people believe that things they don't know are harder than the things that they do no, which means it directly leads to people discounting the things that they have already learned. The alternative is some of the PhDs I used to work with at a university, where they believe, “Wow, I am the world's leading expert in this one incredibly narrow field, therefore, I'm also very good at everything else.”

And that's a bit of a negative expression of that, but there's really something to be said for don't discount the value of what you already know. I'm a huge believer in the idea that everyone could give a credibly compelling conference talk about something that they know that they think everyone else knows, but the response that they're going to get from the audience is, “Holy crap. I just learned something awesome.”

Ceora: Oh, yeah, I agree with this 100 percent. And admittedly, I'm not very good at applying this to myself because I always have this little voice in my head that like, “Oh, everybody knows that already. Nobody wants to hear you talk about it.” But I've told people before, your experience and perspective is uniquely yours. So, everyone has something valuable to bring to the table.

I know that there are some ways that I understand concepts—even with AWS because a lot of coding concepts are very abstract.—so I have a special way of thinking about things. Even though I'm not a huge fan of web design and front end, I am a very visual person, I'm a very visual learner. So, it means that when I tackle something in AWS, for instance, I have to see things, I have to be able to really visualize it to understand it. And that is a perspective that a lot of people don't have, and a lot of people haven't heard of.

And they may be visual learners too, but they never think about things that way. So, me sharing my perspective could be very helpful to someone. Someone tweeted, “Oh, I'm trying to explain some JavaScript concepts, but I don't want to drone on and on.” And I shared with them how I understand classes and object-oriented programming, and it's a very corny—but it's the way I understand it. And I think it helped the person to be able to see different ways that you can teach a concept.

So, you know, usually, when you think of a class, I've heard the car metaphor or whatever, like, “Oh, the make and the model,” or whatever. The way I look at it, I look at classes as little build a bear workshops in your code, so everyone who goes to a Build-A-Bear workshop comes out with a bear. Every bear has eyes, every bear has arms, every bear has feet, but the way they look is different, right? So—

Corey: Well, someone's better at Build-A-Bear than I am.

Ceora: [laughs]. So, some bears will have different outfits, different colors, different eye colors, and stuff like that, but they're all bears. And that's what classes are: they're all going to be the same object, they just might come out differently depending on how you define them. So, each kid is going to go in wanting something different and come out with something different, but as a bear. So, that's how I understand classes and object-oriented programming.

But that's my unique perspective, and everyone has a unique perspective when they approach anything that they're learning or that they're working with. Share that perspective. Share it in a conference talk, and I guarantee you it will blow people's minds. There have been times where I went to a virtual conference and hearing people give talks even about concepts that I don't know anything about, like, React, I don't know anything about React—

Corey: Does anyone?

Ceora: [laughs]. I guess probably one of those AWS type things that you'll never become a master at. But I heard these conference speakers give talks, and they blew my mind. To this day, there are some things I think about from those talks, and I'm sure they thought it was just something normal.

It wasn't a huge revelation to them, but to other people, it can be incredibly valuable, So I think it's super important to share your perspective, and share how you view things, how you view the world. It could be so meaningful to other people, and you might not even realize it, which I'm really bad at doing. Like I'm saying all this stuff, but I don't even take my own advice, but I'm trying to do better. [laughs].

Corey: I feel like that's that is probably one of the best life advice tips that anyone can get. Like, “Yeah, I’m trying to do better.” That's the…this stuff is hard, and you're never able to master it all. But it's definitely of interest watching people evolve and how their learning path goes. If people want to learn more about what you're up to, who you are, and follow your learning in public journey, where can they find you?

Ceora: So, I'm most active on Twitter, my username on Twitter is @ceeoreo_. So, that's C-E-E-O-R-E-O-underscore. And I talk a lot there. I probably spend more time there than I should. And I just created my blog ceoraford.com where I'll be posting more articles and sharing my thoughts on various things. And yeah, that's basically the only two places that I post a lot of tech-related things if you want to keep up with what's going on with me and keep up with my journey.

Corey: Excellent. I will of course include links to that in the [00:32:30 show notes]. Thank you so much for taking the time to speak with me today and, of course, for letting me take the time that anyone is listening to this off so I can watch the next generation learn to effectively cry itself asleep.

Ceora: [laughs]. Of course, I'm so excited for you and excited for this opportunity and all that kind of good stuff.

Corey: Thanks once again, Ceora Ford, learner advocate, and instructor. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts whereas if you've hated this podcast, please leave a five-star review on Apple Podcasts along with a comment telling me what other jobs you should not accept while failing the first time.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Tim Banks
Tim started his 20+ year career in tech quite non-traditionally. After joining the US Marine Corps to be a musician, he was reassigned into an avionics specialty based on the results of standardized testing.

After learning about and working on electronic equipment in the military, Tim went on to work for hardware manufacturers and defense contractors as a civilian. Specializing in systems administration and operations for large Unix-based datastores, Tim left the government contracting world for the private sector, working both in large corporate environments and small startups.

Today, Tim leverages his years in operations, DevOps, and Site Reliability Engineering to advise and consult with engineering groups as a Technical Account manager. Tim is also a husband and a father of five children, as well as a competitive Brazilian Jiu-Jitsu practitioner, and the reigning American National and Pan American Brazilian Jiu-Jitsu champion in his division.

Links Referenced

  • Company Site: https://www.missioncloud.com/
  • Twitter: https://twitter.com/elchefe

Transcript
Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Catchpoint. Look, 80 percent of performance and availability issues don’t occur within your application code in your data center itself. It occurs well outside those boundaries, so it’s difficult to understand what’s actually happening. What Catchpoint does is makes it easier for enterprises to detect, identify, and of course, validate how reachable their application is, and of course, how happy their users are. It helps you get visibility into reachability, availability, performance, reliability, and of course, absorbency, because we’ll throw that one in, too. And it’s used by a bunch of interesting companies you may have heard of, like, you know, Google, Verizon, Oracle—but don’t hold that against them—and many more. To learn more, visit www.catchpoint.com, and tell them Corey sent you; wait for the wince.

Corey: nOps will help you reduce AWS costs 15 to 50 percent if you do what tells you. But some people do. For example, watch their webcast, how Uber reduced AWS costs 15 percent in 30 days; that is six figures in 30 days. Rather than a thing you might do, this is something that they actually did. Take a look at it. It's designed for DevOps teams. nOps helps quickly discover the root causes of cost and correlate that with infrastructure changes. Try it free for 30 days, go to nops.io/snark. That's N-O-P-S dot I-O, slash snark.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by someone who's a great engineer, but frankly, an even better human. Tim Banks is currently a TAM or technical account manager at Mission Cloud. Tim, welcome to the show.

Tim: Thanks, Corey. Glad to be here.

Corey: So, I always feel the need to explain the TAM acronym because, as everything in technology, we like to overload things. It can mean, depending upon who you're talking to, either technical account manager or total addressable market. And if you get them confused, it really cuts across a little strangely, when you're dealing with AWS, for example, it's, “Oh, my TAM is enormous.” It's like, “Are you fat-shaming someone, or you're talking about marketing?” And the answer could easily be both. But you have a different direction on time sometimes.

Tim: Yeah the technical account manager can vary in it’s meaning—almost like DevOps does—from organization to organization in what they do. And a lot of times I found, what happens is they end up just being relationship curators, which is fine in a lot of cases. But I really think the whole point of having a technical account manager is technical accountability, both to the customer and to the company that you represent. And the reason I say that is because you are essentially an evangelist for the products you represent, for the company you represent, and for you the work that they do, somewhat of a developer advocate.

But it also goes the other way. You are also an advocate for the customer back to the company, back to the product teams, back to the business, back to making sure that your leadership is doing the right thing by them. And so I think that accountability is what's important. Even when some people just say, “Oh, all you're going to do is escalate this ticket,” or whatever, which is often one of the recurring functions of a TAM is to make sure support escalations happen, that's still a form of accountability in the end; making sure that the bugs get addressed is accountability. Making sure that the strategy that you came up with with your customer in a meeting or in an architecture meeting is being implemented and the roadblocks are trying to be removed on both sides; that's accountability. And if you're not doing that, you're not really doing the damn job right.

Corey: To be clear, you used to work at AWS as a TAM. And as it turns out, as I've walked through the world, I've met an awful lot of TAMs, mostly at AWS, but again, it's hardly a principle that they have pioneered. But one of the problems that I’ve found, especially during an era of rapid expansion, was that between account to account, the TAM's quality was markedly uneven. And I'm not saying that as far as not caring; they all cared. But it was their level of ability to achieve resolution, to get back to customers rapidly, effectively the core skill set of the job was very uneven.

And at the time, it felt to me like it was a symptom of rapid growth on the part of AWS itself, as it was talking about its millions of customers. It seems like you can't scale up a support org and maintain a consistently high-quality bar, though Lord knows they tried. And again, if you're listening to this and you are attached to AWS, don't worry, I'm not talking about you. I'm talking about those other crappier TAMs who don't listen to the show.

Tim: Yeah, those are the ones you always got to watch out about.

Corey: Oh, absolutely. When you say, “Corey Quinn,” and they say, “Who?” Ugh, warning sign.

Tim: [laughs]. So, I think it's interesting, especially in representing AWS. AWS is essentially a collection of about 500 startups that are operating under one flag. Every org has its own culture. Even within that org they have various factions within their own culture.

TAMs are regional at AWS. So, you have Dallas TAMs, you have Austin TAMs, you have Boston TAMs, you have Silicon Valley TAMs, and that culture tends to cater to the customers that they serve. And ideally, when everyone works in their little section, their little region, it's pretty uniform because you have that kind of unified approach and those standards within that region. What happens is in the reality is that TAMs work across regions and across cultures. For example, I worked out of Austin, Texas, but my customers were in New York, Seattle, Portland, the Bay Area, Chicago.

And so when you have that type of spread across the regions and across, especially, the various areas where TAMs worked, you would see differences in how they operated. And so I think also, there's a difference in experience people bring to the table. My experience as a TAM was based on 20 years of being an engineer, with a large part of that being an AWS customer, whereas sometimes TAMs come from more of a background of—not customer service, but more traditional account management and consulting, which is a different animal than being the hands-on engineer. So, I identified very well with engineering groups, and with the developers, and with the SREs that were banging their head against the wall trying to get things moved, and then taking those concerns up to their leadership and holding them accountable for how they wanted to implement their strategy.

Corey: My approach to TAMs has always been to not view them, necessarily, as technical resources, but rather their role was more or less to do traffic management between getting the customer problem to the person at AWS who’s empowered to fix it. With almost 200 services now, I think it is unreasonable to expect any one person to know what all of them do, and how they work, and the intricacies of how they break. Lord knows I'm trying, but I admit I have a couple of gaps.

Tim: One or two.

Corey: [laughs]. It’s a communication role. The number one thing to focus on is driving the ability to foster the right conversation with the right people, and not be afraid to escalate. But you've left AWS, and now you are same job, different company: Mission Cloud, which is, to my understanding, not a cloud provider as such, but rather a consultancy.

Tim: At Mission Cloud, we do some things that are essential for a lot of people: cost optimization, architecture reviews. We do a lot of work on ProServe. We do manage DevOps, which means different things in different people, and we do some monitoring. What I think the biggest thing that we do that helps people out is that we offer them insight and experience.

We have a very senior group over there. It's some of the smartest and, honestly, the most caring people I've ever worked for as an organization, across the board. People have different levels of experience who have different strengths, but the one thing they all are is the embodiment of the AWS principle of customer-obsessed. We actually care about the people and the customers and the individuals with which we work, which I think is a luxury of having a smaller company; we are more invested in the success of our customers. And that comes across.

And I think people who've gone from very large companies to very small companies see that. One thing that I've noticed is that the smaller companies, they have agility and they can move quickly and things like that, but what they also have the ability to do, it seems almost counter-intuitively, is the act on doing the things that they believe are right, or the things they believe are ethical, or the things they believe are moral which, like I said, seems counterintuitive because you would think a larger company with more money and more leverage and more influence would be able to do whatever it needed, and it would be able to take good stands, and moral stands, and ethical stands rather than doing maybe unsavory business with people or having unsavory business practices. But I found that the larger the company is, the less likely it is going to take the moral or ethical stance.

Corey: It's interesting to me that when I ask people to show up on the podcast, they often bring very different assumptions around what the show is going to cover, like, “Oh, do you want to talk about X, Y, or Z? The answer is always, in fact, “What do you want to talk about?” So, it's interesting to see some folks want to take it directly into, “Oh, let's talk about technology,” And I have to back people off from, “Well, let's not get into the intricacies of how an API works. No one is going to be staring at a computer and listening to this.” They're usually commuting—back in the before times, and that was a thing people did—and I don't want to put them to sleep and have them ram a bridge abutment. I don't want that on my conscience.

So, it has to be something that you can follow while driving around. And some people go in the technical direction, talk about capabilities, others talk about business, but occasionally I wind up with someone like you who talks about, I guess, working in technology in varying capacities, either as a business or as a person while retaining your soul. And that's always been something that I think you've been extraordinarily good at centering.

Tim: I will say that I've always been as good and I think being an underrepresented minority in tech—I'm a half Black, half Mexican pansexual man—that's not always been my luxury, especially starting out. The tech industry, especially where it was pre-Y2K when I joined into the industry, was hostile toward people of color, and toward women, and toward our LGBTQ friends. And a lot of us left. A lot of us washed out. We couldn't take it.

And the ones that did had to work in environments and in cultures, in offices and around people that were very hostile towards us. And the freedom I have to be a whole person and who I am now and to follow my conscience, and to do those things that I believe are right, within the confines of a professional and business relationship only come because I've had to endure what the industry was like before. The industry has come a long way; we have a long way to go.

Corey: Yeah. Well, we've come a long way. Therefore, we can plant a flag, declare victory, and go back to talking about the computers, right?

Tim: Don't we wish. I relate it to having worked in the kitchen. So, it can be said that I'm a bit of a renaissance man. I've had several careers, many running in parallel, and one of them was working in kitchens, and working in kitchens around the time that the big economic downturn of 2008 hit, where you saw a lot of very fine dining, very highly rated capable and talented chefs leave these well-funded, very expensive and very traditional and regimented kitchens, and go into doing things like food trucks, or delivery stuff, or very small things where they can be themselves and be creative. And what you saw that was the democratization of food and a democratization of fine dining, where it wasn't turned from a place where you go in and you sit down and some person in a suit comes to you with a bottle of wine, another person comes to you in a suit and fills your water, then yet another person in a suit comes and takes your order that was prepared by toque wearing, check wearing chefs that were trained at a culinary school.

That's not what most people equate with fine dining now. That fine dining is far more accessible than ever was because it has been democratized. You're seeing that process starting now in tech. It's starting. We’re at the beginning. We aren’t at the end. We're probably not even at the middle.

But you are seeing the democratization of tech, and as more and more people come into tech, more people of every background, and of every ethnicity, and every gender or sexual persuasion and of every religion, when tech is truly democratized, what you will see is tech companies will start to make more ethical decisions. Tech companies will start to have more diversity and inclusion and equity and leadership. Tech companies will be able to start to lead the way in doing the right things. They will be able to start influencing governments, they'll be able to start influencing—you know, whether it's local governments, foreign governments, even, you know federal governments, then we don’t start influencing them because they are made up of a representation that is not currently reflected in those governments, and they will be a source of power. For me, tech is a source of generational wealth that my family has never been able to achieve.

Truth be told, I don't love doing tech. I'm good at doing tech, but tech is my way to ensure that my children have something to inherit, that we have a house, that we have something to pass down, right, that's more than just a family Bible. And so you'll find that there are more stories like that. And when you see people that have that background when you see people that have the varying types of experiences, you're going to see better software, you're going to see better code, you're going to see much more smoothly operating practices because you have now an environment that is inclusive of everyone, that caters to what everyone's needs are, and what everyone's wants are. And when you do that, you're going to have just a much better product all around. There's no two ways about it.

Corey: One thing that I think I got wrong a lot was I would see these various conference submissions that I applied to—back when conferences were thing that we did—and there was a diversity and inclusion track, and I never submitted anything to those tracks because I looked at this, I looked at myself, I looked around, and saw how wildly over-represented my demographic was and thought that I had nothing to say. And this caused two problems. One, it is very easy for silence to be mistaken for complicity, and I never want to fall into that trap, and two, it puts the burden of doing the work on the people who this current societal structure has expressly disadvantaged. It's, “Oh, you're not being treated with equity in the same way that the overrepresented demographic is, so, therefore, we're going to make you do a whole bunch of unpaid extra work.” That's shitty.

Tim: It is.s, I think, one of the positive things that has come out of the current climate, especially since George Floyd’s murder, was that white people, white males, well-to-do and privileged white people have joined the discussion, and joined the discussion in an impassioned manner. They're out there in the streets, and they are protesting, and their kids are protesting, and they're finally really lifting where they stand. And I think that's what's been the biggest change from protests of the past. The things we saw in Ferguson were still primarily Black people out there protesting. And it was very easy to paint them a certain way; it was very easy to ignore.

But in 2020, it's a much more diverse group. And when I say, “Much more diverse group,” I literally mean, there are more white people protesting, and there are more white people with money who are protesting. And maybe it's fashionable, maybe it's something that corporate can latch on to and try to make money off of it. And sure, that has some negative connotations; it has some negative effects. But in the end, the more people that join this discussion, the more people that say, “This is enough and we need to make a change,” the better, especially when those people are white men.

Corey: In what you might be forgiven for mistaking for a blast from the past, today I want to talk about New Relic. They seem to be a relatively legacy monitoring company, and I would have agreed with that assessment up until relatively recently. But they did something a little out there: they reworked everything. They went open source, they made it so you can monitor your whole stack in one place and, most notably from my perspective, they simplified their pricing into something that is much more affordable for almost everyone. There's even a free tier with one user and 100 gigs per month, totally free. Check it out at newrelic.com.

Corey: I think part of the problem is that from the white guy perspective, there are two different directions to go in. One of which is that you deny that there's a problem and that it's just people who want things handed to them that—the all lives matter crowd. And that's a shortcut to being a dumpster fire of a human being. And the other side is that, well, I don't want to take up space. It's too easy when we have these conversations for me as a white man to inadvertently—or advertently—center myself.

And I want to make sure that I clear space for it. The problem is, is from the outside until you get to extremes, they look identical. And for me, my wake up moment was, if I don't start saying things out loud, I'm not going to be distinguishable from some of the worst people in the world. My argument had always historically been, well don't get political on Twitter because I run a business and I want to make sure that I don't inadvertently put half of my prospective customer base off.

And now it’s, “No. If I need to do business with some of the worst people in the world, I'd rather shut the company down first.” At some point, no. We do have an ethics policy of who we’ll work with at The Duckbill Group, and it's squishy by design, but all it comes down to is when I tell my daughter and future—as of this recording—unborn child, where their college funds came from, I don't want to be ashamed. That's really all it is. And by sitting here, shutting up and not talking about these issues, I would be ashamed. I am ashamed that I sat here quietly and didn't say things for the past few years.

Tim: I think as we look at where we are and where we have been, I don't think it's necessarily bad for us to feel some sort of shame--at least contrition—for the way we were. Lord knows I have. I may be Black, but I'm still a man and I was definitely either unknowingly or unwillingly, but certainly, a vehicle for misogyny when I was younger. And I feel shame for that. And it's not in the context, “Oh, because I had daughters,” because that's irrelevant. And not because, “Oh, I have a wife.”

You know, none of that is rel—it's because I've done things that were not beneficial. They were harmful in and of themselves, without the context of whether or not I have female children or a spouse. And I think once we realize that, and you talked about talking to your daughters and not being ashamed to tell them where the college fund came from. And it's not even that context, right? You have progressed, and you have learned from your mistakes, and you are not the person that you were before.

We talk about it like in software development: like you're iterating on it, and you're improving each time. You're fixing your bugs; maybe you have some new bugs, you're going to fix those as well, right? That is a natural progression. That is the thing we're supposed to do as people. The problem comes when the people don't want to do that.

When they're going to sit there and say, “There is nothing wrong with Windows Vista.” Which is essentially what that is, right? That whole notion is the Windows Vista of personalities. Where “This is good, I don't need to change anything because this is the way it's always been.” Or, “This is the thing that I believe in, and why would ever change that?” It's just not healthy, and the lowercase p ‘progressive’ it's not progressive.

And who wants to stagnate? Who wants to sit there and do nothing? Who wants to just exist and never change? So, that's good. And I think we all need to do that. And I think when we look at where we are at Mission as a company, one of the things that we talked about is trying to be much more aware of how we are inclusive and how we are trying to encourage diversity.

But I think the key that you talked about, and I've mentioned about is, you know, the shame and, kind of, guilt for where you've been before, but we have to have compassion in there. One of the things we're talking about is how to have inclusive language, how to be more inclusive in what you say to people, how to be more inclusive of gender, religion, sexual orientation, race, or whatever it is, but when you have to say to somebody, “Hey, that person’s preferred pronouns are ‘they’ not ‘she?’” You have to do that with compassion because we're trying to learn. We're trying to progress. And I think you've mentioned this before, I know sometimes with me, I'm still defensive, right, for stuff like, “Oh, I'm not wrong,” or, “This is the way [00:21:14 crosstalk]—”

Corey: Well, my immediate response when I get called out for something—maybe it's a joke that punched down in a way that I didn't see it, or I didn't realize, but what I was talking about on Twitter, for example, is an expression of privilege that I didn't see, and my immediate response whenever I'm called out for that is a flash of defensiveness. And I've learned to walk away for a minute, or not indulge that because growth always starts off with that feeling for me, in my experience, and taking that as it's easy to shortcut that into someone has called me out, politely or otherwise, for saying something hurtful that I did not intend to be hurtful, my immediate shortcut defensive reaction is, “Oh, they're wrong because I didn't mean it that way.” That's never how it works.

Tim: No.

Corey: To be fair, a few times this has happened, and I thought about it for a little while and I went back and no, I don't believe that that's accurate—I forget what it was, but recently someone responded on Twitter to a pull quote from one of these podcast episodes with that is a problematic term. And I went back later and said, “Okay, I've done some research on this. I don't believe it is, and I can find no corroboration on it. How so?” And their account was deleted. And it turned out that it was effectively a bot or troll account. So, oh, I did a whole lot of work, and it turned out that no, I was just successfully baited into something. But, you know, I'll do the work every time.

Tim: I think the work is what's important. And this is what I was talking about: when we have compassion, when we say to somebody, “Hey, maybe you should say that differently,” or, “This is what that person prefers.” That's the beginning of the work. The work for receiving that correction, or call out is to look in yourself, reflect, maybe educate yourself a little bit, and then say, “Hey, I'm going to make this change.”

But that's not a switch, right? Bugs don't get fixed the second you think about them or the second somebody mentions them. You have to do work to fix them. You have to iterate on yourself in order to correct that. And so I try to approach that with compassion. Obviously, some things are very blatant. Some things are very immediate in their need to maybe—sometimes are a little more inflammatory, but especially around people whose companionship, whose friendship, whose opinions that I value, I will try to be as compassionate as I can, especially to strangers.

Compassion is important. Compassion, I think, in tech and in the world, is what makes a huge difference. Whether you express it in the form of empathy, where you express it in the form of letting people have their time, or where they need to self-care, making sure they don't burn themselves out, making sure that you're speaking to them inclusively, making sure that you are receptive to being told that you're not being inclusive. All that is compassion.

I think really what we see as a larger thing in the world today is a lack of that. It's either lack of compassion or lack of maybe taking the difficult stand. And I think it's interesting to pivot to that a little bit—using, kind of, the difficult stand part—one of the things that you brought up earlier today that I thought was very important was the actions of large cloud provider who maybe was meeting with companies, and then—under the guise of advisement, and then making competing products. I think you've mentioned that before, and it's something that I've seen as well in my time, and that kind of punching down is horrifying.

Corey: I see it also with certain large cloud providers with their non compete clauses in their contracts. I'm sorry, I have remarkably little sympathy for a company that is valued at over a trillion dollars, when they start, effectively, continuing to maintain their level of supremacy on the backs of their employees who effectively have zero power comparatively. Now, collectively, their employees have a lot of power. But oh, that's union-type words. And we don't want to have those conversations.

Tim: Oh, heaven forbid that we have to treat our employees fairly.

Corey: Oh, it's maddening. And at some point something changes, I think. I was independent for a while, I was an employee—badly—for a long time, and then when I started hiring people, I kept waiting for the switch where I would suddenly want to screw over my employees in favor of my own advancement, and I haven't found it yet. Maybe it happens at some point. I mean, I firmly believe that you don't get to build a trillion-dollar company and keep your soul. I think that you have to make compromises that are objectively evil. But on balance, I'm not trying to build something that's going to turn me into a billionaire. I have no desire to ever get there.

Tim: I think the important thing is to do what's right. There's an LP with Amazon, saying that, “Leaders are right a lot.” And I think that should change.

Corey: Yes, meanwhile here, a followership principle at The Duckbill Group where we tell people we're right a lot.

Tim: [laughs]. I think instead of, “Be right a lot,” it needs to be, “Do the right thing.” That difference is huge. It’s a huge difference between being right and doing the right thing. It's like you said before, you know that you're not a racist, you know that you're not a bigot, you know that you are doing these things—

Corey: I am not a bigot, but I will say that we are all racist based upon the aspects of society that we have grown up in, just—not to interrupt you on that. But I have been shaped by the racism I have grown up in. There is a world of difference, though, between saying, “I’m subject to racist tendencies, and thus have them myself,” and saying, “I am a bigot.” There’s a world of difference, and I think in common conversation, people conflate those two terms. They're not alike.

Tim: I agree and appreciate that correction. But I do think what's important is when you said that it wasn't enough that I simply wasn't a bigot, but what needed to happen is that I needed to then take action; I needed to not stand idly by, and be right. Instead, what I needed to do was do right. And I think if it was instead of, “Be right,” it was, “Do right,” then some of these decisions that you see being made, would not be made.

You would not see companies leveraging the technology or the ideas of their customers to compete against them and make more money. They would not be leveraging their positions against their employees to keep those ideas and that talent from being able to make them money at other places. Because it's not enough to think the right thing or to say the right thing: it is important as a leader, and for someone that wants to have a good impact on the world to do the right thing.

Corey: Doing the right thing is super easy to do when it has no consequences and it's low impact. It's one of those, “Oh, do I wind up buying a carton of eggs that explicitly says cage-free on it for another 50 cents instead of the one that doesn’t?” All right. I'm going to do that, I’m going to feel virtuous. I’m basically buying a good feeling for 50 cents at that point.

When it starts turning into something that has the potential to impact business, to cause controversy, to take a stand on something, that's where it gets scary, and that's where it gets fraught. One way I could see having to stop talking about these things at some point because I feel it now when it was just me I could sound off about anything I wanted in tech with zero consequence. If I get it too terribly wrong and turn myself into a cautionary tale, well, now we're 10 people. That is a lot of people's livelihoods I'm gambling with.

And one thing I still do and I make it a point, whenever we have a final round interview, is I make it a point to be on that call. So, I can say, “This is a risk factor.” It has to be me that says that because if anyone else on the team says that about me it sounds like, “We can't stand Corey and he's terrible.” But it's one of those things that is important to me that, first, I remember and I honor, and secondly, people are not blindsided by the fact that I'm occasionally going to go on Twitter and say something dumb.

Tim: I really think the advantage of doing that, and really, the thing that makes it appealing is that people get to say, “All right, well, I want to hitch my horse to this wagon.” Or I guess, “My wagon to this horse.” Because you can do a little bit of research about Corey Quinn and, kind of, find out where you stand on a lot of things, and I think that goes into making the choice of where you want to work. Because you are small enough to where your values—you know, the things that you hold dear and believe to be important—are reflected in what the company does. And there's an advantage in that in attracting the people that want to work in that culture. I don't think they're going to work for The Duckbill Group to become billionaires, sad as I am to say.

Corey: No, not until we finish our acquisition of Gartner.

Tim: Oh, in that case, well then yeah, I will gladly buy your pre-IPO stock, then.

Corey: Excellent. Excellent. Thank you.

Tim: But when you work for companies that are very large, and now it's the name you want to work for, not the culture of the person, it becomes a different animal altogether. It's like, why would people buy Nikes? Why would people pay so much for Nikes? It’s because they say ‘Nike’ on them. There are maybe shoes in a similar quality or better quality that are less expensive and less well known, but you're going to buy these Nikes because they say Nike on them.

A lot of people want to have Amazon on their resume. A lot of people want to have Google on their resume, or Apple, or any of these other large players in the tech industry. It's good to have in your resume. A lot of people want to work there because they said, well, I’m—you know, when was the last time you heard of somebody that worked at Google that didn't tell you that, right?

Corey: Yeah, they are very practiced telling the story of why you should work somewhere. It's good for the resume. It turns out, it's going to help you out quite a bit. I don't have any big names on my resume, not in tech, not that anyone would ever notice or care about, and that's okay, but I can definitely understand the appeal. You see this sometimes the people who used to work at Netflix, for example, and every conference talk, they get is about, “Well, back when I worked at Netflix 20 years ago…” at some point, it's, “Yeah, you can only trade on what you used to do for a while.” People are going to ask, what are you doing now?

Tim: I think that's what becomes important. Going back to, like, doing the right thing. You can say, well, I did this back then, I did this back then, and I made this decision back then it's like, “Oh, great. What are you doing now? What is the great thing you're doing now?”

Like you say, “What influence do you have now? What is the difference you're making now?” And if you're just standing up there, essentially reliving your glory days in high school on a conference talk, I don't think that's going to impress anybody. And it almost becomes sad because you're using this platform now to sing about something that you did twenty years ago, ten years ago, five years ago, three companies ago, instead of trying to make some difference now. And now, more than ever is when we need the difference to be made.

Corey: I think that's probably the best place to leave it. If people want to hear more about what you have to say, where can they find you?

Tim: I'm on Twitter @elchefe. Hang on for the ride because I'm very much the authentic me on there. But that's the place for now. I might appear on another podcast later on, and if so, I will put it on that Twitter account.

Corey: Excellent. Some of my best Twitter accounts that I follow make me uncomfortable from time to time, and that's exactly why I follow them. Bring your whole self to work, or at least to Twitter.

Tim: Absolutely.

Corey: [laughs]. Thank you so much for taking the time to speak with me. I appreciate it.

Tim: Thank you, Corey. I appreciate it.

Corey: Tim Banks, TAM, and excellent human being at Mission Cloud. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts, whereas if you've hated this particular episode, please leave a five-star review on Apple Podcasts as well as a comment explaining why you're not a human trash fire.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Alex Chan
Alex is a software developer at Wellcome Collection, a museum in London that explores the history of human health and medicine. Their role primarily focuses on preservation, and building systems to store the Collection’s digital archive. They also help to run the annual PyCon UK conference, with a particular interest in the event’s diversity and inclusion initiatives.

Links Referenced

  • Wellcome Collection: https://wellcomecollection.org/
  • Twitter: https://twitter.com/alexwlchan
  • Blog: https://alexwlchan.net/

Transcript
Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Catchpoint. Look, 80 percent of performance and availability issues don’t occur within your application code in your data center itself. It occurs well outside those boundaries, so it’s difficult to understand what’s actually happening. What Catchpoint does is makes it easier for enterprises to detect, identify, and of course, validate how reachable their application is, and of course, how happy their users are. It helps you get visibility into reachability, availability, performance, reliability, and of course, absorbency, because we’ll throw that one in, too. And it’s used by a bunch of interesting companies you may have heard of, like, you know, Google, Verizon, Oracle—but don’t hold that against them—and many more. To learn more, visit www.catchpoint.com, and tell them Corey sent you; wait for the wince.

Corey: nOps will help you reduce AWS costs 15 to 50 percent if you do what tells you. But some people do. For example, watch their webcast, how Uber reduced AWS costs 15 percent in 30 days; that is six figures in 30 days. Rather than a thing you might do, this is something that they actually did. Take a look at it. It's designed for DevOps teams. nOps helps quickly discover the root causes of cost and correlate that with infrastructure changes. Try it free for 30 days, go to nops.io/snark. That's N-O-P-S dot I-O, slash snark.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Alex Chan, who among many other things that we will get to is, most notably, a code terrorist. Alex, welcome to the show.

Alex: Hi, Corey. Thanks for having me.

Corey: So, you've built something wonderful and glorious. Well, that's my take on it. Most other folks are going to go in the opposite direction of that, and start shrieking. Namely, you've discovered that AWS ships a calculator, something the iPad does not, but AWS does, and the name of that calculator is, of course, DynamoDB. Tell me a little bit more about how you made this wondrous discovery.

Alex: So, I was watching one of your videos where you were talking about some of the work you were doing in AWS land, and you were talking about how you were starting to explore DynamoDB. And DynamoDB is the primary database that we use in my workplace, on AWS. And I knew you could do a little bit of mathematics in AWS; you could do sort of simple addition using things like the update expression API, and you can also get conditional logic using conditional expressions, and then I decided to string all that together with a series of Python and see if I could assemble those calls to get a basic working calculator.

Corey: Some folks would say that you could just do the calculator bits without the DynamoDB at all. Python has a terrific series of math functions and other libraries you can import. What do you say to those people?

Alex: The thing about running your code in Python is that you have to have somewhere to run it, and that's probably going to be something like a server. Whereas if you run it in DynamoDB, that's serverless. And as we know, that's much better.

Corey: Oh, absolutely because that's the whole point of modern architectures, where we wind up taking things that exist today, and then calling them crap, and reimagining them from first principles, just like we would on Hacker News, and turning them into far future ways that doesn't add much value but does let us re architect everything we've done yet again and win points on the internet which, as we all know, is what it's all about. And I'm extremely impressed by this, but my question now is, so as you've figured this out, what got you to a point where you looked at a database and said, “You know what? I bet I can make that a calculator.” How do you get there from here?

Alex: So, in this specific case, I’d done just a little bit of work with DynamoDB already, and I sort of had a vague notion that you could do something like this, but I'd never really understood. It was this API that I knew existed, and so I just started working through the documentation, pulling apart examples, and discovering, “Oh, yes, actually, this API can do this,” And in the process, actually came up with a much better understanding of the API than I had before.

Corey: One of the interesting pieces to me is that there's a certain class of people—and I refer to y'all as code terrorists, and I aspire to be one on some level—I look at something and my immediate question is, how can I misuse this? Kevin Kuchta, for example, was able to build a serverless URL shortener with Lambda—just Lambda; no data store. It was a self-modifying Lambda function, which is just awesome, and terrible, and wonderful all at the same time. And you've done something similar, and I have, in the sense of getting Route 53 to work as a relational database, sort of. And whenever I describe these things to people, they look at me in the strangest and arguably saddest ways that are just terrible, absolutely terrible.

I love the idea, I love the approach, I love the ethos, and I want to see more of it, but there's a long time that goes between me coming up with something like Route 53 as a database, and then there's a drought, and then eventually someone like you comes up with now we're going to use Dynamo as a calculator. How do we get more of this? I want to have these conversations more than once every six months.

Alex: How do we get more of this? I think, first of all, just coming up with the ideas and actually putting them into people's hands. We don't necessarily need to be the people who do them. Obviously, this DynamoDB as a calculator came about because of a throwaway comment you made on a video that I watched. So, if we can think of ways to, I don’t know, use SQS as persistent storage, or use SNS for two-way communication, and then you put those ideas out there and it will fit in someone's head—and hopefully someone who knows a bit more about these things—and they will start to think about it, and they will know what the edge cases are, and what the little bits of API they can exploit. And that's how these things come about, I think. But we've got to go out there and plant the idea in somebody else's head, and then let them sit on it three months, and then finally go, “Aha. I know how I'm going to do that.”

Corey: So, I didn't realize I was one of the proximate causes of this, and I take a little bit of joy and, I guess, pride in the fact that I can inspire that. So, there are a few different directions to take it in now. One, do you think that Texas Instruments has woken up yet to the fact that they have a new competitor in the form of AWS?

Alex: I mean, I've not heard from any of their lawyers yet. So, I have to assume, no.

Corey: I mean, the TI-83 hasn't changed since I was in high school 20 years ago, and it feels like maybe Texas Instruments is not the best-suited company, as far as ‘prepared to innovate rapidly’ goes. I kind of hope we’ll finally break their ridiculous calculator cartel.

Alex: I mean, certainly, I think that's one of the greater injustices in the tech industry today is that, you know, Texas Instruments remains the undisputed king of graphing calculators, and obviously we look forward to DynamoDB breaking that hegemony.

Corey: Well, that's the real question. How far does this go as a calculator? Basic arithmetic is sort of a gimme; what does it do beyond that?

Alex: I implemented the basic arithmetic operations, addition, multiplication, subtraction, and division. And then along the way, as I was trying to build those, I ended up coming up with a series of logical operators. So, first, NOT, then OR, and then an AND, and then finally, a NAND gate. Now, I don't have a computer science background, but I can read Wikipedia, which is almost as good, and Wikipedia tells me that if you have a NAND gate, you can essentially build a modern processor. And since we can do NAND gate in DynamoDB, there's no reason we can't build virtual processes on top of DynamoDB as well. And once you can simulate a virtual processor, then really it's a few short steps from there to having the next EC2.

Corey: Well, what I was wondering is, now we have the scientific calculator stuff potentially taken care of, the next step becomes, clearly, graphing calculators. Is that something that DynamoDB is going to be able to do, or are you going to have to switch over to Neptune AWS’s graph DB, which presumably will be needed for a graphing calculator?

Alex: I confess I haven't used Neptune. I have thought a bit about, though, how you might—

Corey: Well, you and everyone else, but that's beside the point, really.

Alex: But I have thought about how you might use DynamoDB for a graphing calculator, and really the answer here, obviously, is that we're going to have to turn to the console. The console will show you the rows in your DynamoDB database, and so we just use very narrow column names, and then we fill them with ones or zeros, and then that will allow you to draw shapes in the console.

Corey: I'm thinking through the ramifications of that. If you're able to suddenly start drawing shapes in the console, that would put the database system—now a graphing calculator—significantly further ahead than other AWS services like, for example, Amazon QuickSight, which ideally is a visualization tool, and in practice serves as a sad punchline that we're hoping they improve faster than Salesforce can ruin Tableau. But right now, it's like watching a turtle race.

Alex: Exactly, and I think this is the flexibility of having a general-purpose compute platform. You can do arithmetic operations, you can do your scientific calculator, but then you can build on that to do more sophisticated things, like visualizations. And I think it's a shame that more people haven't tapped the power of DynamoDB already.

Corey: I keep hoping to see further adoption. So, changing gears slightly on this, you have a, apparently, full-time day job that is not building things like this, which first, may I just say, is a tragedy for all of us. But secondly, what you do is interesting in its own right, tell me a little bit about it.

Alex: So, I work from for an organization called Wellcome Collection, which is a museum and library in London—

Corey: And that is Wellcome with two L’s.

Alex: Wellcome with two L's. It ruins your ability to spell the greeting.

Corey: I'm waiting for AWS to acquire them. Anything that has a terrible name with extra letters, vowels, et cetera, seems like it is exactly on-brand for them.

Alex: I couldn't possibly comment.

Corey: Of course not. So, what does the Wellcome Collection do?

Alex: So, we are a museum and library, and we primarily think about the human health and medicine. So, obviously, we've got a lot to think about right now. And one of the things we do is we have a large digital archive so that's a significant quantity of [00:10:13 born] digital material. Somebody actually gives us their files, their documents, their presentations, their podcast recordings, and also a significant amount of digitized material where we've got some book in the archive, we take a photograph, we can put that photograph on the internet, and then people can read the book without actually having to physically come to the museum. And what I work on currently is the system that's going to hold all of those files because it turns out that if you have 60 terabytes of stuff and you just put it on a hard drive and leave it in the closet, that's apparently not so great.

Corey: It's great, right up until magically it isn't, or so I'm told. But yeah, you're right. Every time I talk about long term archival storage, it seems like there's a difference in terms. “Oh, some of our old legacy archives are almost five years old,” is a very different story when we're talking in the context of a library. People who believe the internet is forever, my counter-argument to them is, “Great. Do me a favor, what is the oldest file on your computer?” And that tends to sometimes be an instructive response?

Alex: Absolutely. I mean, our digitization program goes back about a decade at this point. So, we've been keeping files for that long. But then we're looking very far into the future with the archive in general.

And in fact—so rules vary around the world, but in the UK, certainly, the standard rule is that if you've got an archive about a living person, you typically close that archive for 70 years after their death. So, that means they're gone, anyone who remembered them was gone, and also probably their children—and maybe their grandchildren as well—are gone because, particularly when you're dealing with medical records, say people may not want it known that their grandfather was in this particular hospital. So, we are planning very much decades or even, in some cases, centuries into the future.

Corey: So, when you're looking at solutions that need to transcend decades, does that mean that cloud services are on the table, off the table, part of the solution? How do you approach this?

Alex: So, for the work we're doing, we very much do use cloud services. We’re mostly running in AWS, and then we're going to start doing backups into Azure soon because you don't really want to rely on a single cloud provider for this sort of thing, in case AWS hear what I'm doing with DynamoDB and close our account, and in part, because organizations like AWS, like Microsoft Azure, they have much more expertise than we can have in-house on building very robust, reliable systems. So, if you imagine a 60 terabyte archive, and you're holding that locally, that's probably at the limit of your personal expertise on how to store that amount of data safely. If you give 60 terabytes to the Amazon S3 team and say, “This is a lot of data.” They will laugh at you because their entire job is around storing large amounts of data, ensuring it lasts a very long time, ensuring that when disks fail they get replaced and the data is replicated back onto the new system. So, for us, we've really embraced using the Cloud as a place to put all this stuff because a whole lot of problems around, “Is this disk still going to work in two years?” Are solved for us.

Corey: There's another challenge, too, in some ways. As you look at larger and larger datasets and looking at cloud providers—one of my favorite tools in Amazon's archival toolbox has been this idea of Glacier Deep Archive where the economics are incredible. It's $1,000 per month per petabyte, which is just lunacy-scale pricing. But retrievals can take 12 to 24 hours depending upon economics and how you want to prioritize that. And that works super-well in scenarios where you're a company and you need to keep things around for audit or compliance purposes, and your responsiveness to those requests is going to be measured with a calendar rather than a stopwatch, but for a library where you need to have a lot of these things available online when people request them, 12 to 24 hours seems like an awfully long time to sit there and watch a spinner go around on a website.

Alex: It is and it isn’t. Twelve to twenty-four hours is actually, in the context of some library things, is quite fast. If you're requesting physical objects certainly, there are a lot of things in the library you can't just go and pick up off the shelf. London real estate is expensive, so about half of our physical collection actually lives in a salt mine in Cheshire, and if you want to see it, you make a special request to us, and a van drives to the salt mine and picks it up for you.

Corey: In a digital context, we generally just refer to the ‘salt mine’ as Twitter.

Alex: Exactly. The way we actually handle this for most of our work is we have two copies of everything because you never want to have just one copy of it because then you're one fat-fingered delete away from losing your entire archive. So, we have one copy that lives in standard IA. That's the warm copy; that's the copy that we can call up very quickly if someone wants to look at something on our website, and then we have a second copy that lives in Glacier Deep Archive, and then that’s separate. Nothing should be reading that; nothing should be really touching that, but we know there's a second copy there if something terrible happens to the first copy.

Corey: It comes down to the idea of what the expected tolerances and design principles going into a given system are. You're talking about planning things that can span decades into centuries, and I'm curious as to how much the evolution of what it is you're storing is going to change, grow, and impact the decisions you make. For example, if I'm starting a library in the 1800s, the data that I care about is effectively almost entirely the printed word, and ideally some paintings but that's sort of hard to pull off. As we continue to move into the 20th century, now you have video to worry about, and audio recordings that are becoming a technology. And nowadays we're seeing much higher fidelity video, and larger and larger objects while the cost of storage continues to get cheaper over time, as well. So, I'm curious as to how you're looking at this. Today's 60 terabyte archive could easily be 60 exabytes in 20 or 30 years. We don't necessarily know. How are you planning around that looking forward?

Alex: Well, it's very hard to predict the future. And if I could, I would have made some very different life decisions. So, what we've done instead is just try to build it in a quite a generic way in a way that doesn't tie too strongly to the content that we're storing. The storage service we've built mostly just treats the files entirely as opaque blobs: it puts them in S3, it puts them in that Glacier Deep Archive, it checks they're correct, but it doesn't care if you hand it a JPEG, or movie file, or Word document. It's just going to make sure that file is tracked, and is sorted vaguely sensibly. And we're hoping that that will give us the flexibility to continue to change the software that supports it as our requirements change.

One of the things we were very conscious of is any software that we write in 2018 and 2019 to do this sort of thing is going to be obsolete and thrown away, probably within a decade, certainly by 2040. And so we want to design something and store the data in a way that was not tied to a particular piece of software, and that somebody could come along in the future, and pull it back out again and understand how it was organized, or start adding their own files to it if our software has long gone.

Corey: If you’re looking to wind up standing up infrastructure but don’t want to spend six months going to cloud school first, consider Linode. They’ve been doing this far longer than I have, they’re straightforward to get started with, their support is world-class—you’ll get an actual human, empowered to handle your problem rather than passing you off to someone else like some ridiculous game of ticket tennis—and they are cost-competitive with any other provider out there, with better performance in almost every case. Visit linode.com/morningbrief to learn more. That’s linode.com/morningbrief.

Corey: Historically, when I was doing this stuff with longer-term, “Archival media,” quote-unquote—you know, those special CDRs that are guaranteed to last over a decade. Now the biggest problem is finding something to read them because technology moves on. Bit rot became a concern; the idea that the hard drive that you stored this on doesn't wind up working, or there's media damage, or it turns out that there was a magnet incident in the tape vault. Whatever it is, the idea that eventually the media that holds that data winds up eroding underneath it, rendering whatever it stores as completely unrecoverable. How do you think about that?

Alex: We think about that a lot because we still have a lot of that magnetic media. We still have Betamax cassettes, and VHS tapes, and CD ROMs, and one of the big things we're currently doing is a massive AV digitization project to digitize as much of that as possible before it becomes unreadable. I think if I'm remembering correctly, Sony stopped making Betamax players a number of years ago, so the number of players left in the universe—and the number of spare parts—is now finite, and is only going to get smaller. And even though those types might be good for another 10 years in our temperature-controlled vaults, we and a lot of similar organizations are really having to prioritize digitizing that and converting it to a format that can be stored in something like S3 because otherwise, it's just going to be lost forever.

Corey: Do you find that having to retrieve the data every so often, and validate that it's still good, and rewrite it is something that is viable for this? Does that not solve the problem the way it needs to be solved? I've dabbled looking into a couple of options at this stuff years ago, and never really took it much further than that. So, I'm coming at this with a very naive perspective.

Alex: We've never looked at this in detail, but we have done exercises where we pull out large chunks of the archive, and completely re-checksum of them, and validate the SHA-256 of the thing we wrote six months ago is indeed still the SHA-256 of the thing that's now sitting in S3. We did this for a significant chunk of the archive recently; it was pretty cost-effective. We were able to run it on Amazon Fargate, scaled it out massively, ran in parallel, it was very nice. The biggest cost was in fact, the cost of all the GetObject calls we had to make against S3, but it was a couple of hundred dollars at most. So, this sort of money where if we felt it was important to do, we’d just do it again.

Corey: It feels like on some level, that's what things like S3 have to be doing under the hood where they have multiple—like this idea of erasure coding or information dispersal—the idea of you can have certain aspects of it rot and it doesn't tarnish the entire thing, [00:20:55 unintelligible] some arbitrary percentage. And we've played with this on Usenet in years past with parity files: download enough objects and you have enough to reconstruct the whole.

Alex: Exactly. And there are people at Amazon whose entire job revolves around making sure S3 doesn't lose files, which is part of why we use it because they're going to think about that problem much more than we can. And we basically trust that if we put something in S3, it's probably going to be fine there. The biggest thing we're worried about is making sure we put the right files into S3 in the first place.

Corey: And that's always the other problem, too, which is a—if you'll pardon the term—library problem. You have all this data living in various storage systems. That's great. But how do you find it? It feels like it's that scene from Indiana Jones and—one of those movies, I don't know what it was—Indiana Jones and the Impossible Cloud Service, where they have a warehouse scene at the end where everything for miles is just this enormous warehouse and everything's in crates. How do you find it again? Which system did that live in? That always seems to become the big challenge. And we see it with everything, be it Lambda functions, DynamoDB tables—still a great calculator—and other things. Which account was that in? Which region was it in? Expanding that beyond that to data storage feels like, unless you're very intentional to beginning, you're never going to find that again.

Alex: Yeah, so one of the things we did that we made quite a conscious decision to do early on, was we tie everything back to an identifier in another system. So, all of our files will be associated with at least one other library record. So, they might be associated with a record in our book catalog, they might be something in the painting records, it might be something in the archive catalog. And then that's the identifier we use just to hold the thing in the storage service.

So, you can look at a thing in the storage service and say, “Ah. This has the identifier B1234. I know that's a book number, I can go and look in the library catalog and find the book that's associated with,” and vice versa. So, essentially, we're pushing out the problem of organizing back to the librarians and the archivists because they have very strong opinions and rules about that sort of thing, and it's much easier just to let them handle it in one place than to try and replicate all that logic again in a second place.

Corey: Do you think that there is a common business tie-in as far as—a library looking to store things on that kind of timeline seems like it's a very specific use case and problem space that any random for-profit company is going to take one look at and say, “Oh, that's not really our area. We don't know what next quarter is going to hold, let alone the far distant future.” Do you think that's a mistake? Do you think that there are lessons that can be learned here that map to everyone, and where do those live?

Alex: Certainly a lot of what we've been doing, I think, is more widely applicable than just the libraries and archives space, and I've been writing about a lot of it; we've been sharing a lot of what we've been doing, both for the benefit of other libraries and archives, but also for people who might find some of this stuff useful. And I think one of the big decisions we made early on that I think was really valuable and would serve a lot of companies well, is that idea that we assumed all of our software would eventually be thrown away; that at some point, all of the code we've written is going to be thrown in the bin, and someone is going to have to do a migration to a new service, whether it's using JavaScript, or it's running on DynamoDB, or we’ve progressed past cloud computing and we're into nebula computing. Whatever it is, we assume that the software will become obsolete, and we very intentionally optimized to make that easier for whoever's got to do that in future. And a lot of the time, I think, I see companies build something that works great right now, but the moment it breaks or it needs to be rewritten, it's going to be a huge amount of work. And just thinking a bit about that earlier on would have made it much easier to move away when they eventually need to do so.

Corey: Part of me wonders on some level though, that when I'm building a company that doesn't necessarily know whether it's going to exist in a week, it feels like that is such an early optimization. Like, the things that I worry about, even in the course of my business—which is fixing AWS bills—is, “What if Amazon dries up and blows away?” Is fundamentally very core to my business as far as disruptive things that might happen, but if that happens, everyone is going to be having challenges. It's going to be a brave new world, and building out what I've done in a multi-cloud-style environment or able to be pivoted easily to other providers just hasn't been on the roadmap as a strategic priority. Maybe that's naive, but I honestly don't know at this point.

Alex: No, and we’re still—you know, in the grand scheme of human history, we're still very early in these things. And yeah, maybe AWS will go away next week. I certainly hope not, but when we were doing this, we didn’t—obviously, we thought about this a lot more than a lot of people would because we really are expecting to optimize for that very long use case, so what I'm suggesting is not that you prepare yourself to pivot multi-cloud, that you prepare to be able to run workloads anywhere, you be able to shift your workloads around dynamically, but it's just taking a little look at your decision saying, “Is this decision going to lock me in, in a really aggravating way? And is there just a slightly simpler way that I can do this that is going to be much easier to unpick from later?”

One of the big ones for us was, for a long time, we were looking at using UUIDs to store everything because UUIDs are brilliant. You never have to worry about uniqueness or versioning. It's just handled for you, but then we thought about what it would take to unpick those UUIDs later and work out what they meant, and we realized, “Well, alternatively, there's a great identifier over here sitting in the library catalog. Why don't we just use that instead, and throw away all these UUIDs?” And that wasn't a huge amount of work, right? It was just a case of deciding which of these two strings do we put into the database. But I think long term, that's going to make a massive difference to how portable the system ends up being.

Corey: That's one more topic I wanted to get into before we call it a show, that ties together the two things that you've been doing, maybe. Namely, how did you get into looking at systems like DynamoDB and seeing a calculator, and possibly the archival stuff, too, in such a weird and unusual way? It's not common, and it is far too rare of a skill. How did you get like this is the question I guess I'm trying to ask, but without the potentially insulting overtones?

Alex: No insult taken. So, I think like a lot of people, my first community on the internet was primarily fan-ish. I grew up on the internet, reading fan fiction, and for people of a certain age they will remember sites like fanfiction.net, Wattpad, and the big one at the time was LiveJournal. And huge fan history discussions were conducted on LiveJournal.

And I got to know a few people there, and a friend of mine was friends with the head of LiveJournal’s trust and safety. And if you've never come across it, trust and safety is this fascinating role where you have to look at every aspect of a system and think about how terrible people will misuse it to hurt people. And we're talking about things like stalkers, like abusive exes, like that coworker who doesn't know what boundaries are, and you've got to work out how, for example, a social media site is going to be completely ripped to shreds by these people and used to hurt users. Because if you're not doing that in the design phase, those people will do that work for you when you deploy to production, and then people get hurt, and then you're extremely sad.

And so that was the thing I was thinking about very early on on the internet, was I was talking to these people, I was hearing their stories, I was hearing how they design their services to prevent this sort of abuse, and I got into this mindset of looking at the system and trying to think, “Well, okay, if I wanted to do something evil with this system, how would I do it?” And in turn, when I'm building systems, I'm now thinking, “What would somebody evil do with this, and how can I stop them doing it?”

Corey: It almost feels like an InfoSec-style skill set.

Alex: Yeah, there's definitely a lot of overlap there, and a lot of the people who end up doing that sort of trust and safety work are also in the InfoSec space.

Corey: I think there's a lot of wisdom buried in there, and I think that, frankly, we've all learned a lot today. If nothing else, how to think longer-term about calculator design. Thank you so much for taking the time to speak with me. If people want to hear more about what you have to say, where can they find you?

Alex: I'm on Twitter as @alexwlchan, that’s W-L-C-H-A-N, and I blog about brilliant ideas in calculator design at alexwlchan.net.

Corey: Excellent. And we will throw links to that in the [00:29:30 show notes]. Alex Chan, senior software developer, and code terrorist. I am Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts, whereas if you've hated this podcast, please leave a five-star review on Apple Podcasts and a comment telling me exactly why I'm wrong about Texas Instruments.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Jason McKay

Jason is responsible for leading Logicworks’ technical strategy including its software and

DevOps product roadmap. In this capacity, he works directly with Logicworks’ senior engineers and developers, technology vendors and partners, and R&D team to ensure that Logicworks service offerings meet and exceed the performance, compliance, automation, and security requirements of our clients. Prior to joining Logicworks in 2005, Jason worked in technology in the Unix support trenches at Panix (Public Access Networks). Jason graduated Bard College with a Bachelor of Art and holds several AWS and Azure Professional certifications.

Links Referenced:

  • Logicworks: http://logicworks.com
  • LinkedIn: https://www.linkedin.com/in/jasonhmckay/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. This week's sponsored guest is Jason McKay, the current, and as it turns out, first CTO of Logicworks. Jason, welcome to the show.

Jason: Thank you, Corey. Great pleasure to be here.

Corey: Well, you say that now, we'll see how you feel in half an hour or so.

Jason: [laughs].

Corey: But you have a fascinating career trajectory, where you started off as an engineer at Logicworks in 2005—back in the dawn of time. That was before the iPhone, to give people a sense of history here—and then you became a senior engineer, and then a director of engineering, and then the VP of engineering, and now you're the first CTO. So, you're basically the computer equivalent of someone who worked their way into a leadership role but started off in the mailroom.

Jason: Yeah, pretty much. I think there's places to go from here. You know, I have my eye on marketing, but we'll see.

Corey: [laughs]. So, there's something to be said for folks who have done the role at the company that they're working at, and now in a position of leadership that it really tends to lead to, I guess, an interesting sense of perspective across the board. But that's really what I want to talk to you about today is perspective. To begin, Logicworks is a MSP—or managed service provider—for several cloud providers, now. What the heck is a managed service provider?

Jason: [laughs]. Well, managed service provider, the role and function really predates the Cloud as it exists today, but our job is really to provide sort of guidance, best practice, and guardrails around operating critical applications in the public cloud today.

Corey: So, when people talk about choosing an MSP when it comes to negotiating cloud deals or doing a cloud migration, from my perspective, most of my client base has always been the type of folks who are working directly with a cloud provider. If you take a look at the broader ecosystem, what does an MSP do for its customers?

Jason: If you're talking about a market where the application is critical to the business—so maybe it's the bulk of their revenue or their focus—they want to be able to put the focus of their business on that business driver. So, if they have an application or a suite of applications, they want to focus on the development of those, improving those, they need to ensure maximum uptime, availability, scalability, operational best practices, et cetera, and a managed service provider will cover a lot of those things outside the application development itself and lets you really focus on where the meat of the business is.

Corey: Gotcha. So, on some level, it means that your company doesn't necessarily need to be focused on a relationship with a cloud provider, but rather focusing on what your company sets out to do. It, more or less, streamlines a lot of the sharp edges of interacting with hyper-scalers, for lack of a better term.

Jason: Absolutely. Obviously, necessarily part of that is being very, very familiar and integrated with the different hyper-scalers out there—the ones that we support—and knowing them like the back of our hand so we can provide that best value to the customer and not require that the customer be intimately familiar with every service and every offering from the public cloud provider itself.

Corey: On some level, how prescriptive does that become? In fact, let me ask that question in a far more confrontational and offensive way, aren't most MSPs kind of shitty? I mean, it feels like on some level, “Great. Oh, you're going to manage my cloud environment. That means you're going to slap me into a virtual machine that comes in small, medium, or large, and once I wind up getting that done, if I take a step back and squint hard enough, what you've built is functionally the exact same architecture that I had in my data center, except now I get to pay by the hour for it instead.” That's in many ways, my common historical perception of what the problem with MSPs are. Am I right and this is a really short conversation, or am I missing something key?

Jason: No, you're not. I mean, those MSPs exist, absolutely. And we enjoy very much eating their lunch every day. We view the role of an MSP very differently than that. Being prescriptive, where it matters, is important, but being flexible and adapting to our customer requirements is far, far more important.

And also our focus on the public cloud in general, it's because we're very much appreciative of the rate of iteration and innovation on the public cloud, and the last thing that we want to do is get in the way of that for our customers. So, we certainly have identified some, kind of, best practices that we will generalize and apply across our customer base, but we absolutely stand firmly by being flexible with the customer applications as they come in. And for us, that's one of the major reasons to move to the public cloud is things are moving quickly, new services are introduced; if we're not allowing our customers to leverage them—and even, I would say, going a little bit further and acting as the R&D arm on behalf of our customers with new services and offerings from the public cloud, then we're not doing our job. So, if our competitors are doing as you described, that's great. Don't tell them about this.

Corey: Oh, yeah. They're too busy trying to figure out the best way to build their own custom load balancer because the ones the providers give them aren't meeting their needs, for whatever reason.

Jason: [laughs]. Right. Exactly.

Corey: A lot of folks have other business concerns that they're trying to aim themselves at.

Jason: Yeah.

Corey: I mean, again, it sounds like it's a mixed bag, there are some MSPs that sound like they're really on the cutting edge, keeping up with the technologies, doing what's right for their customers, and then other MSPs are Rackspace. So, first, I guess, how would someone wind up picking a MSP when they're looking at this assortment? Because I know that if I Google for ‘MSP’ I have to click through five or six pages, at least, before I see something that isn't an ad in one form or another. Every time I log into the AWS Marketplace, for example, and type in MSP at that point, someone knocks on the window. It’s, “How on earth did you find out where I live?”

It's very much a competitive market, and it seems on some level—at least from the 10,000-foot view of marketing and seeing the industry that way I do—there's something of a struggle to differentiate. What do people look for—successfully—in an MSP? What do they look for that they shouldn't really be paying attention to? And, honestly, what are folks getting wrong in that assessment process?

Jason: I think what they should be looking for, and this is certainly an area we've tried to focus on, is if it's the right MSP, and this is ‘right’ from my point of view—and I think, based on what you said about the pitfalls of most MSPs—from yours, as well, you're looking for one that is engaged in thought leadership and forward-thinking approach to public clouds, and you're looking for one that has a focus and a cultural embrace of innovation, and that's why they're in the space that they're in, as opposed to trying to factor things down to common generalities and make a cookie-cutter approach to things. So, you're looking for an MSP that's comprised of people that are focused on innovation and fast-moving technology advances in the public cloud. And you'll see that in less, kind of, cost-based marketing or positioning, and more thought leadership where the MSP is focused on new trends and new capabilities in the public cloud.

Corey: One of the problems that I tend to see, even among my customer base of AWS environments with significant billing challenges, you think at some point that starts to normalize and everyone more or less has the same architecture. Yeah, every time I think I've seen it, all it takes is one more client, and then I'm seeing something else different. You wind up with a huge breadth of workloads even on a relatively small client basis. How do you handle this without going insane?

Jason: That's a challenge. We see everything. I mean, the good thing is that there's a virtue in being in that position where you have to deal with the breadth of the spectrum of applications along several axes, like cloud-readiness, cloud-nativeness, how modern the application architecture is, et cetera. So, one of the interesting things that we get to do is because we're dealing with everything from very traditional risk-averse industries, like insurance, or healthcare regulated industries, at one end of the spectrum, and then at the other, the really early adopter, bleeding-edge be damned, I will run the thing that's in beta and was released yesterday.

We get to deal with both of those, which, for us, and particularly from an engineering background, that new cool stuff is great. So, by the time it becomes mainstream, a little more mature, more stable, ready to be adopted by those folks at the other end of the spectrum, we're already experts. So, it's nice that we get to cover that whole ground with a focus on the bleeding edge, and the sort of regulated maximum uptime all changes are risky crowd at the other end.

Corey: So, when I wind up taking a look across your history, you obviously started out in a time before there was Cloud, to speak of.

Jason: Indeed.

Corey: Yeah. Then AWS came along—I think they went GA, what, in 2006 or so—

Jason: Yeah.

Corey: You waited a while on that. A while? Huh. Yeah. You wound up signing up as a partner with AWS and focusing on managing workloads in it, it was back in 2012.

Then five years go by, it's 2017, you became a partner with Microsoft Azure. And that's it at the moment. There's no third answer as far as large hyper-scalers you're working with. So, it's clear that, one, you're not looking to partner with anything that holds still long enough, which is kind of a refreshing change if I'm being perfectly honest, and it speaks to a certain thoughtfulness in how you decide to support a given platform. At least that's my perception. Am I right? Am I wrong? Am I hilariously naive, or all the above?

Jason: Well, but I hope you're right because it's certainly complimentary if you are. So, first of all, we're customer-driven. So, we were driven to AWS by our customer base, we recognized it as a force in the market pretty early. I made some of the same missteps that other folks did, including our own public cloud offering for what feels like a couple of weeks there. And then the move into Microsoft Azure recently was largely customer-driven, where customers were saying, “We need to have some workloads running on that platform, can you move there as well?”

And we're certainly not ruling out expanding our supported cloud stable, but it really has to get to kind of critical mass because, having done this, we know that it's not an easy thing to bring in a whole new cloud paradigm alongside the rest of your customer base. And you have to do it right, so our standards are pretty high there. So, there will come a point where we're adding GCP, or—gulp—Oracle Cloud, but it's not there today. So, yeah, customer-driven. And if you take a step back, our job is really to look at overall trends in where workloads are running. And I'd say we got to the Cloud fairly early—it's, maybe, late by AWS early adopter standards—but we'll make the move to the next paradigms, hopefully at the right time.

Corey: So, you clearly have experience migrating folks from data centers of varying quality, ranging from first-rate, everything's run super well, to, well frankly, what a lot of us descend into, and have to fight to keep the raccoons from carrying our servers off into the wilderness before they get decommissioned. And you have experienced moving people from on-premises into one of the cloud providers you work with. Do you have a lot of experience migrating folks from one cloud provider to another?

Jason: You know, I would have expected if you'd asked me a few years ago, whether that was going to be something that we would be called upon to do, I probably would have guessed, “Yes.” But in truth, it simply doesn't come up, or it hasn't for us. We have customers who run workloads, in both clouds. Certainly have customers who are in one cloud or the other.

We have not yet—to my knowledge unless I missed it. I'm pretty sure I didn’t—but we've not yet done any kind of large scale public cloud to public cloud migration. We certainly have, back in the old days, pre-Cloud, we had our own data center footprints, and we were happily cannibalizing our own business from effectively on-prem data centers into the Cloud, but public cloud to public cloud, it's not really come up.

Corey: I see the exact same type of thing, where I thought there was going to be a lot of people leaving one cloud provider for another, but I don't see it. And more importantly, I also don't hear about it. Because if you win a customer away from one provider and they pick you instead, you're never going to shut up about it. That's going to be the highlight of your keynote for the next five years because, “See? Someone went through the incredibly painful migration between cloud process to use us,” and I just don't see it happening. I see people threatening to do it all the time, people exploring it, but it's expensive, and it doesn't add direct business value, so why would you do that? It's a negotiating tactic to be sure, and sometimes a workload might move, but I don't see it being done large scale where, today we're all-in on AWS, and tomorrow we're all in on Azure instead.

Jason: Yeah, we don't see it either. We definitely see customers running on both, and there's several reasons for that. Some of them are commercial. And—

Corey: —oh, I definitely want to get into that with you.

Jason: A little bit risky. But—

Corey: No, well it's true. I mean, a common refrain that I've been taking has been that I don't see that multi-cloud is a best practice for individual workloads. I can see a story of, “Oh, I'm a big company and I acquired another company that's on a different provider. Yeah, great. Migrating them may not make sense.”

Jason: That’s—

Corey: See previous diatribe. Yeah, that can happen. Sometimes it does. Sometimes it doesn't.

Jason: And it's also sometimes commercial situations where a customer has end customers who potentially compete with one of the other cloud providers in some way, and that may govern where the workload ends up as well.

Corey: Right. I can be much more specific about that, where Walmart has very publicly said that they don't put anything on AWS—although let's not get ourselves they do have a disturbing number of employees with AWS specific skill sets on LinkedIn, but who am I to judge or cast stones?—and they've also said, “Great. They don't want their vendors to wind up having their data living on AWS either.” Which, like it or not, it's a thing.

Now, it's one thing to say that I don't want to necessarily target retail as a industry I'm after, so building on top of AWS makes sense, but what you're also intrinsically saying is, I don't want to target retail or companies that will do business with retail. And suddenly it's, “Oh, that's a much bigger customer base than I would have expected.” And people say, “Ah, well, what about that?” As if I'm somehow some kind of AWS partisan? It’s no, if I'm building something that wants to sell into those markets today, I probably wouldn't start on AWS, if I'm going greenfield. And the whole problem goes away.

Jason: Yeah. I mean, that would be a rational approach.

Corey: Yeah. But I would pick a provider—maybe Azure, maybe GCP, maybe Oracle Cloud, maybe IBM Cloud, if I've taken a sudden, sharp blow to the head—and then I'm going to go all-in or whatever provider that I've picked. And for better or worse, I'll be able to leverage some of the undifferentiated heavy lifting. What I don't see in my practice, and I'm curious if you do, is single workloads that are built to be deployed in a provider-agnostic way. I found a few, but they are the exception far more than they’re the rule.

Jason: Yeah, we don't see that in practice. I will say there's a slight nuanced difference that not only do we see, we practice it ourselves. So, we actually have internal applications that will leverage both clouds, where we'll have a portion of the application function running in AWS and an ETL pipeline to a SQL service on the Azure side, and some business analytics around it, and then sending that data back to an application front end in AWS. So, that's a multi-cloud application, if you will, but the same functions are not served by each cloud.

Corey: Right. You're effectively using best of breed or best economic story for technologies—

Jason: Exactly.

Corey: —on each end of that, and that makes perfect sense. But it also flies against the pro-multi-cloud nutters out there who are, “Ah. So, now you're apparently going to say that if one of your cloud providers goes down, your application will too, so you have to scale across multiple providers.” With what you just described, if either side of that takes an outage, that application is no longer working correctly.

Jason: That's true. And that's absolutely something that we consider, and we have various advocates inside our organization for one platform or the other. And I often find myself in the bad cop position of saying, “If there were a problem with this environment, how would that affect the entire thing?” And that makes us revise our architectural choices there.

Corey: Oh, yeah.

Jason: In general, everything is a trade-off. So, you make your risk assessment and you put your money down.

Corey: And let's not kid ourselves either. None of the hyperscale providers today have outage stories that are so significant that, “Well, the problem we run into is that that cloud provider is down one day out of every three.” At that point, they would not be a cloud provider anymore.

Jason: Yeah, I’m knocking on wood just for the entire digital economy’s sake right now, but yeah, you're right.

Corey: People ask me, “Oh, so let me get this straight“If newsletter pipeline that I build that goes out every week and makes fun of what AWS has done, well, that's only in a single region—us-west-2—if all of Oregon region goes down for a while, does that mean, you're not going to be able to write your newsletter?” And the answer is, “Look, if AWS takes an entire region down for a protracted period of time, I'm writing a very different newsletter issue that week, and—not for nothing—I can do that one by hand, it's fine.”

Jason: And yes, you'll have plenty of compelling material, absolutely.

Corey: Exactly. And it comes down to also understanding the model of your business. And this obviously doesn't apply to everyone, but if you're selling socks, for example, and your site takes an hour-long outage, according to some of the metrics and studies that I read, well, “Oh, that means that people are not going to buy those socks, and you've lost that opportunity forever.” Very often people come back an hour later and buy it then. That said, it said, if you're an ad tech, people are not going to come by an hour later to click an ad. So, there's a question of aligning what your risk exposure is to your business model.

Jason: Absolutely. And it's one of the things that I get a chuckle out of sometimes when you're meeting with a prospect who's coming in and saying, “I need five nines, four at a minimum.” And you're like, “Yeah. I don't think you're like medical life-saving equipment in a real-time sense. Do you know the dollar cost per nine here you're talking about?” People have some pie in the sky ideas of what they actually need versus what they can actually do.

Corey: And that's part of the whole problem I see is that you wind up with folks who get carried away. We saw this after the big S3 apocalypse back in 2017, where S3 was down for most of an afternoon in a single region. And a bunch of engineers that I spoke to were all talking about how, “Oh, now they're going to back up all of their data to multiple regions and maybe other providers as well.” And it's, “Okay, first off, you're talking about doubling or tripling the raw infrastructure costs, plus whatever ancillary things around that are going to factor into that, you’re increasing complexity significantly to avoid a black swan event once every seven years that already will never recur in quite this same way.” What does that workload actually do? Oh, it turns out that there's actually no serious business impact past being embarrassed when that thing goes down. Furthermore, great. You read the news, it was never about your site being down. It was, “Amazon had a problem and the internet has melted.” It's a very different story than engineers will often see from their particular position within engineering. There's a larger context here that businesses generally have to weigh because everything involves trade-offs.

Jason: Yeah, absolutely. It's a very interesting change, and if you think about the carefully preserved and bragged about uptime, when everything was under much, kind of, tighter, more local control in terms of infrastructure and application, now that's weather, right? And if your site is down, so is everyone else’s, most likely, and it's a non-issue.

Corey: So, changing gears slightly, people love to comment on various episodes of the podcast where, “Oh, you asked them a bunch of softball questions. Why didn't you ask them difficult questions?” Because insulting people with questions they're not going to answer is always the way to lead to a terrific show. But if I asked you a question, such as, “Who's your target customer?” Great. That sounds sales pitch-y and it sounds like I've set you up for it, so I'm going to turn it around the other direction here. Who would a terrible customer be for Logicworks? Who would come in the door, at which point you would turn them around, send them on their way because they're not a fit for your model, how you view things, their architecture is complete nonsense—for God's sake, they use Route 53 as a database—who would use such a thing? Who's a terrible fit?

Jason: Certainly, we've had our share of them. And too often, you find out far too late.

Corey: Yeah, some of their logos, I'm sure are on the customer wall because that's always the way it works.

Jason: Yeah, exactly.

Corey: But I digress.

Jason: But mostly, it's about their mindset and how willing they are to, kind of, take guidance, and if they come to us with some wrong ideas, but they're willing to take guidance on that, that's one thing. And, you know, that's a great thing. But if you run into that sort of set-in-a-ways approach, that's going to be a challenge. So, the misunderstanding of business requirements, of availability, and data durability, and things like that, and as we just discussed about the cost of a nine, that's one. Another one is there's a common sort of misunderstanding that I can take my crusty old application, which is still, kind of—my deployment method is like artisanal handcrafted servers that takes six weeks to get the code onto them, and registry edits put in, et cetera, I can take all of that and virtualize it in the Cloud, and now I'm Netflix.

That isn't the case, so that's a common scenario run into where, just because you're in the Cloud, you're not cloud-native now, necessarily, and you might even be making some mistakes about how you deploy. So, for example, I've seen large applications with many tiers employing auto-scaling, where there's no way any of these things could actually survive an auto-scaling event. And they don't need it: they don't scale up or down at all, even after running for a long period of time. And you look at something like that and you'll say, “I know you wanted to leverage all the cloud-y things; your application isn't ready for it. Let's give you something that actually works and gives you high availability and uptime, and as you mature the application, we can move over and take advantage of some of those cloud-y things.” That's one.

So, it's incorrect expectations is typically one. Obviously, there's both ends of the extremes there. The really old applications that they'll try to wedge into the Cloud, that's going to be a challenge on the other end. They might have the complete, early adopter situation where everything is containerized including their own load balancers and layer seven load balancers with HAProxy, and they're not leveraging any services because they need to be fully cloud-agnostic.

Corey: Sorry, somewhere, listening to this, a VMware rep is drooling. Please continue.

Jason: [laughs]. Well, VMware certainly has its own things going on. I did have an account on their short-lived cloud, and I remember using both the features and seeing that it was satisfactory. But yeah, I mean, it's really the extremes, but most importantly, I think it's those that are not really willing to have a dialogue with us on how to adjust their use of the Cloud for their application as it is today and not what they hope it will be.

Corey: That's a pretty honest assessment. And it's got to be a difficult message to deliver because what you're functionally saying is, is that you've almost fallen for the hype of what cloud is, and it's not going to serve you well to continue down this path until you've made some additional changes. It's always tricky telling a prospect that they are not a good fit for what you do in ways that could be perceived as negative. But I've got to say, every time a vendor has done that with me, I come away—even though it may bruise at the time—with a almost begrudging respect for them, just because, “Wow, they could have taken my money and completely left me high and dry.” It’s nice when that doesn't happen.

Jason: Instead, they told you the baby was ugly, right? Sometimes you have to.

Corey: Yeah. There's really a lot of value to that. So, going back to my naive approach of the whole MSP space, it seems to me from where I sit, never having gone deep into that universe, that the cloud providers must hate you on some level because you're effectively stepping into the relationship between them and their customer, and I would naively think that, oh, that means that you're perceived as a direct competitor. But with a little digging, it turns out that you and a lot of other MSPs are all partners with all of these providers. There's clearly something I'm missing. Help?

Jason: That's an easy one, actually. Thanks for the softball. I will call that out as a softball.

Corey: It wasn't intended to be.

Jason: [laughs]. Most of the cloud platforms—hyper-scalers—they're obviously running at this huge scale with a huge customer base. What they want is those customers to stay on their platform, have a successful experience, and not be tempted away to some other 0.002 cent per hour cheaper service. So, our job is to ensure the success of these customers on that platform. So, from that perspective, they very much view us as partners: we're going to ensure that there is a successful outcome for this customer on that cloud. And that's really what that comes down to.

Corey: Gotcha. When I look at your case studies and the customers you showcase, by and large, it feels like it is heavily slanted towards migration stories. Is there a viable path for a greenfield company that has just received funding, they're starting up today their Twitter for Pets, but they're going to be an S&P 500 component in a decade? Is there a path, if they're born in the Cloud, where working with an MSP makes sense? And if so, where's that onboarding?

Jason: Absolutely. So, marketing focus aside, it may look like a lot of migrations, but I will tell you that the majority of our customers, certainly in the first three or four years, were greenfield customers. They were potentially large organizations or small business units with software development teams that were moving into the Cloud, and we effectively built them from the ground up, including account provisioning, and organizations, and linking of accounts, et cetera. So, we actually see that quite a bit. It's not uncommon at all.

And certainly, there are a lot of migrations as well, especially as you go, kind of, upmarket. You're talking about bigger orgs, but much bigger software portfolios or application portfolios. And you will be migrating those, sort of, sets of applications at a time. But we do still get net new greenfield applications today.

Corey: I guess one question I have is—I’m assuming that there's a pricing model where you charge more money to the customer than they would pay going directly to a provider, which makes sense. There are stories around doing it as a percentage, there are stories around doing fixed fee, depending on what that looks like, but what I'm wondering is, how do you wind up telling a compelling story for a company that just getting started, where their cloud bill might be, I don't know, 15 bucks as they're building these things out? Because whether you're charging a percentage, there's no reasonable percentage in the world where it's even worth the cost of the phone call—at least up front—or you're charging a fixed rate. Cool, I'll pay thousands of dollars a month for my $15 cloud account. It seems like there's a challenge as far as getting those early adoption customers, particularly when there's no guarantee that any given customer is going to grow.

Jason: That's absolutely true. And you gave an example that is—that's below our market, right? We would not serve that $15 a month in usage customer. There's a certain calculation of risk that we have to take ourselves that we don't necessarily want to play the startup lottery. So, our market is going to be much more, kind of—there's a higher chance of success, where if it's a new greenfield application, it might be an application being developed by an independent team within a larger org that we're not worried about going away overnight.

So, we have to do a little calculus there ourselves. That extremely low end is probably wouldn't make sense for them to go through us. Especially, if they're just getting started like that, they can afford to make some mistakes at that low end now. If they get to a point where there are investors or backers, and the uptime, the correct operational health, and the cost controls, and security around all that are required, then they're probably at a point where they want to engage in a managed service provider.

Corey: Yeah. It feels like on some level, it's going to be a challenging discussion, at least that early on in the process. I know that when I was getting started my big questions were, “Will I be able to eat this month?” And less to do with the longer-term story of building relationships. And Lord knows, if I had gone, instead, of the bootstrap path through some VC or whatnot, I would suddenly have a funding announcement, and suddenly, every vendor on the planet is reaching out trying to solve problems I didn't realize I had, some of which are actually legitimate problems, and others are, “That’s snake oil, you're selling me. I can tell by the label.”

So, it feels like it's one of those challenging markets, and I'm not particularly surprised that it doesn't seem—at least in the broader industry—like there's a big push to pick up those very early stage companies because again, 98 percent of them are never going to turn into them, and every once in a blue moon, you wind up with effectively the next Google.

Jason: Yeah. Yeah, exactly. We talked about this when we moved into the public cloud space back in 2012. We actually had a lot of internal discussion about what markets and made a conscious decision that down at that very low end, it doesn't make a lot of sense for us or them, until they're at a point where it's a viable business, it's proven, there are people that care enough that they want to ensure there's operational security and cost controls around it, at which point, they're ready for us.

Corey: Got it. I think that's a very transparent and honest approach to it. I want to thank you for taking the time to speak with me today. If people want to hear more about what you have to say, what you do, how you do it, or read some glorious marketing case studies. Where can they find you?

Jason: Oh, they can find information on logicworks.com, or they can look up any info about me. I'm on LinkedIn. It's really the extent of my social media, but I can be found there.

Corey: It must be nice to not have to deal with the Twitter mobs on a day-to-day basis. My God. Thank you so much for taking the time to speak with me. I feel like I know more about a sector that I've always just sort of passingly acknowledged as we pass like ships in the night.

Jason: Well, happy to help Corey, and happy to talk anytime.

Corey: Well, thank you. Jason McKay, CTO of Logicworks. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please rate it five-stars on your platform of choice, whereas if you hated this podcast, please rate it five-stars on your platform of choice and leave a comment why I should never consider Logicworks, and instead go with your crappy MSP instead, that is staffed entirely by raccoons.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Veliswa Boya

Veliswa Boya is a 2x certified AWS Cloud Engineer currently working in financial services. She works with application teams on cloud migration strategies and cloud architecture designs. Veliswa has been in the IT industry for 20+ years, starting her career as a mainframe developer working on critical systems for car manufacturers, insurance companies, and banks.

Veliswa is a member of Indoni Developers, which is a platform for African women in coding/tech. She speaks at meetups and was one of the speakers at the inaugural AWS Community Day Cape Town in 2019. She especially enjoys speaking and connecting with those who are new to tech and specifically new to AWS.

Veliswa mentors young people who are looking to embark on AWS certification journeys, she shares her own experiences, gives guidance and support. Veliswa also likes to write about “what she’s learned so far on AWS” and publishes on her Medium blog.

For fun Veliswa enjoys the outdoors, she regularly goes hiking and loves road running.

Links Referenced:

  • AWS Community Hero: https://aws.amazon.com/developer/community/heroes/veliswa-boya/
  • Twitter: https://twitter.com/Vel12171
  • LinkedIn: https://www.linkedin.com/in/veliswa-boya/
  • Dev.to blog: https://dev.to/vel12171

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: nOps will help you reduce AWS costs 15 to 50 percent if you do what tells you. But some people do. For example, watch their webcast, how Uber reduced AWS costs 15 percent in 30 days; that is six figures in 30 days. Rather than a thing you might do, this is something that they actually did. Take a look at it. It's designed for DevOps teams. nOps helps quickly discover the root causes of cost and correlate that with infrastructure changes. Try it free for 30 days, go to nops.io/snark. That's N-O-P-S dot I-O, slash snark.

Corey: This episode is sponsored in part by Catchpoint. Look, 80 percent of performance and availability issues don’t occur within your application code in your data center itself. It occurs well outside those boundaries, so it’s difficult to understand what’s actually happening. What Catchpoint does is makes it easier for enterprises to detect, identify, and of course, validate how reachable their application is, and of course, how happy their users are. It helps you get visibility into reachability, availability, performance, reliability, and of course, absorbency, because we’ll throw that one in, too. And it’s used by a bunch of interesting companies you may have heard of, like, you know, Google, Verizon, Oracle—but don’t hold that against them—and many more. To learn more, visit www.catchpoint.com, and tell them Corey sent you; wait for the wince.

Welcome to Screaming in the Cloud. I'm Corey Quinn. I’m joined this week by the Veliswa Boya who, among other things, is a Community Hero based in Johannesburg, South Africa. Veliswa, welcome to the show.

Veliswa: Hi, Corey. Thank you so much for hosting me.

Corey: So, there's a lot we can dive into here, but let's start with the context of at the time that we're recording this, I'm planning to go out on parental leave in about a month. And I've as people are aware, presumably by this point—depending on ordering. Maybe this is the first, I don't have that figured out yet—I'm having people act as guest authors while I'm out so I can actually take time off to spend with my family. As a result, that means that effectively I’m giving the platform to people for a week and giving them a chance to tell their story, mostly because I don't want to actually write anything down. So, first, thank you for letting me spend some time actually focusing on things that, believe it or not—don't tell anyone—more important than cloud computing.

Veliswa: Sure. This is awesome. So, thank you for giving me the opportunity, as well.

Corey: Of course. There's several things you do: you work in financial services, but more interesting and relevant to what people here are generally listening to. You are a Community Hero for AWS. What does that entail, and how did you become such a thing?

Veliswa: So, I remember when I got this invitation to accept the nomination as Hero, I remember thinking how awesome it is that you go through an entire career—I've been in it for 20 years now, and that entire time feeling like you’re invisible, you never get recognition, no reward for anything that you do. And then this comes along, and it is at such a global stage. It means everything to me. And how I got it, is I like to share a lot with the AWS community. I started working with AWS about three years ago, and what I started doing is to share my learnings.

I don't come from a strictly—your traditional tech background, so I've had to learn a lot of things. It's been a huge on-ramping for me, learning AWS. And as I learn things, I like to share. I like to share with those that maybe are in the same space as me starting out, trying to figure this out. As I learn things or share with them.

I think that, kind of, got me the attention. I speak at meetups, and I'm always speaking as someone starting out:, “I’ve done it, you can do it as well, you should starting out.” I think that maybe got me the attention that got the recognition as an AWS Hero. And what's been so awesome is first woman out of Africa has to be named a Hero. It's so awesome, but of course, I don't want to be the only woman out of Africa forever that is an AWS Hero. I'm hoping to see more come on board as well.

Corey: Wouldn't that be a wonderful change? So, in my home office, I have a map above my desk. It has a bunch of pins in it. And people come in here and look at it and say, “Oh, is this a list of all the places you've been to?” And the answer is, “No, it's not.

There's a pin in there for every AWS region that has been deployed and announced, and a different color pin for all of the CloudFront Edge locations.” And, I tell people, like, “Oh, that explains why in Virginia, there's a pin of a dumpster fire.” Like, “Yes, you get it now.” But I look at this map, and for a long time, it was—I look at Europe, and right now there are one, two, three, four, five, six regions that are there, and an additional one that has been announced in Spain, and there's a crap-ton of CloudFront edge locations right next to it as well. It looks like you can't walk a block without tripping over AWS infrastructure in Europe.

And then I look at Africa. And it's one of those, “New phone: who's this?” Type of moments where, as of April of this year, there's in the reserve region in South Africa—Cape Town, specifically—and there are two cloud front edge locations that I see. One of them in Johannesburg and the other in Nairobi. And that's it. I look at this, and it’s—these two things are in no way alike. Why is it that, I guess, Africa has not been on the map in the world of cloud computing moreso?

Veliswa: I think it's getting there. I think it's getting there. I think it's adoption, and speed of adoption as well. But I think we'll get there. The community is growing quite aggressively, and are, sort of, vibrant.

And all since becoming a Hero, on Slack we've got our AWS community in Africa workspace, and that's where we all meet people from all over Africa. And it's really filling up. And there's a lot of excitement around AWS adoption. But then again, there's always an issue around—because here we also—the struggle around job opportunities, as well. Is always the issue around, I go and I learn all of this AWS and then I sit on—well, I got a certification, but now I end up sitting on it because I'm not getting any jobs.

So, I think the whole job opportunity still has to align, also, to the level of learning and on-ramping that people are doing. So, many people are learning and getting very excited, but they're also looking to get opportunities to use what they've also learned, and to make their lives better as well. But I think it’s getting there.

Corey: It absolutely is. I have one of my consulting clients with a strong presence in Nairobi, and that was, sort of, my introduction to a lot of, I guess, the cloud infrastructure stories here. And sure, I can see pins on a map and all and understand how that works in the abstract, but then talking to people about their actual experiences networking, as we go through the project and start having social conversations with people, it's a different world. I mean, I live in San Francisco, which is its own world, and then some. This is the land of people who believe that if their website doesn't work well on the latest Apple MacBook Pro on a gigabit connection plugged directly into the server, then you should upgrade your systems.

Yeah, the rest of the world does not work that way, and there's a vast disconnect in many respects. And for me, one of the more frustrating pieces that I found is I didn't understand the idea of latency, where people need to be close to where the website is hosted so that it loads in a reasonable period of time. Because naively, I assumed, why does that actually matter? How impatient are you? If it's 300 milliseconds, between me and the server, okay, great. That's going to be less than a second, round trip. So, what's it matter how quickly the webpage loads? Because I still live in the 1998 style of thinking where a web page is a single static file. Yeah. Then I look at how most websites are built. And it's, “Oh, you have 300 requests, sequentially.”

Veliswa: Yeah.

Corey: And I'm looking at this and it's, oh, yeah, suddenly I understand why CDNs are a thing and why infrastructure matters.

Veliswa: Yeah, yeah. What's also been so awesome for me since becoming an—because I'm based in South Africa, and I think my world was kind of South Africa and interacting with people here. But then becoming a Hero, I started meeting all these people from throughout Africa. And I speak to people from Zimbabwe, from Kenya, from Nigeria, and the amount of innovation, the entrepreneurship that actually exists, always being ready to look for opportunities and ways to learn new things and make your lives better. I'm so in awe of everything that I'm learning since I’ve started interacting with everyone throughout all of Africa.

Corey: It's fascinating from a perspective of whenever I travel places, and I start experiencing, I guess, the internet from a different part of the world—which I know it sounds ridiculous. Like, it's one of those, “What the hell do you think travel is?” It’s, “You’re viewing it through how things work on, from the internet, what is wrong with you?” The honest answer in my case is that I travel a lot for work, and what really brought it home for me in a way that nothing else had before, was when I went to Australia a couple of years ago. I had never been before; I thought the place was a myth.

Like, “Oh, I'm going to Australia.” “Great. Are you taking a connecting flight through Narnia?” No, it turns out it's a real place. People live there; who knew? And I wind up doing all the things I normally do: I spin up an EC2 instance in the Australia region, and I get a data plan for iPad, and I get there, and I get into this thing, and it is unusably slow. It is garbage internet, awful. And it's, “Wow, I knew that infrastructure in this part of the world was not necessarily great—people love to complain about Telstra—but how do people get any work done?”

So, cut into the endpoint of the story of something that I didn't know at the time—if you're listening to this, you might learn from my side of it—it turns out, at least in the United States, when you buy an international data plan, it backhauls the data to the United States through whatever discount rate carrier they go with. So, I was sitting in Australia and tethering to data back in the United States, which then would reach back across to Australia to talk to the EC2 instance. And it’s, wow. Other than going the entire span of the globe from one side to the other.

What's dumber than that? Doing it multiple times. So, it was, yeah, it turns out that bouncing back and forth across the Pacific Ocean a few times does add latency. And when I told versions of this story to people I know in Africa, they had similar stories, in many respects, where it's a lot of sites are just designed with this idea that all of our customers or target market are going to live in these particular locations that are right next to various AWS regions, so it'll be fine. My question for you is, how does that experience manifest? What is it like to live somewhere where the typical VC tech bro types don't consider? When they picture their customer, it’s very often not folks in Africa, it's folks in, oh, if they're really far away, they'll be 45 minutes away in San Jose, not in San Francisco itself.

Veliswa: So, I think for me, personally, and I think, actually, for a lot of people, how it actually really stood out was when we got the Cape Town region because up until then we were Ireland. And then Cape Town came up, so we could now spin up instances there, and at [unintelligible], we've got a region now, here. And the difference when you're uploading onto an S3 that's same region as where you are, I think that's when it actually stood out for us. The difference, the old latency, and appreciating all of that. That's when the huge excitement actually came in for a lot of us here. So, now people can't wait to spin up instances in Cape Town because it's right here, it's not in Ireland. But we were fine with Ireland because that's all we had at the time. But now we have Cape Town, and it's right here, and it's very exciting for a lot of people.

Corey: It is. What really always takes me aback is whenever there's a new region announced. I hear it, and—I mean, I see every announcement that comes out of AWS. Spoiler: never do that. But as I go through and weed through all of them, it’s, “Oh, another region announcement.”

And my instinct is to treat it like I would a service expansion to a new region, namely, I don't particularly care. That is my immediate reaction, and I've learned to suppress it because when they announced a new region somewhere, particularly in places that are not overly saturated—this does not apply to, “Yay, we're launching an eighth region somewhere in Europe.” Yeah, yet no one actually cares about those. But when something comes to the Middle East, or something comes to Africa, or something comes to South America, there's such a groundswell of excitement of people in the community being enthusiastic about this. Sometimes, admittedly, it's people who are, I don't know, fortnight players who are super excited, “Yay, there's going to be a region close here so I get better latency when I'm playing games.”

But they're also people with, shall we say, more serious business concerns, who are super excited about that. And I have to say, the South Africa region, back when it was announced in April of 2018, was met with a tremendous amount of excitement, just from what I could see on the internet. Was there a similar level of excitement in the community, when you were talking to folks about that, and the announcement came through.

Veliswa: There was such a lot hanging on this. There was so many businesses that wanted to migrate to AWS, but with some of the areas here, there was a lot hanging around, having the region here because of compliance and data cannot be outside of where you are, things like that, were actually governing everything. So, when we actually finally got a region here, there was so much excitement because finally, those businesses could start their migration journeys to AWS. I think just the community at large as well. All this excitement that we've got our own region here, it’s just so exciting. But for businesses, it was a huge thing for businesses because we were waiting on that for a lot of data privacy, and all of those issues, they kind of needed to have a region here. So, it’s a huge thing when it actually finally did land.

Corey: Azure was the first hyperscaling cloud provider to have a region presence in Africa, weren't they?

Veliswa: They were. They were here for a while before AWS actually got a region. Azure was here for a while, but you have those people that just wanted to go on to AWS, and so you had those few—well, not few, you have people trying Azure already, but you had those that wanted to migrate to AWS. And while migrations were happening, but having a region here was such a big thing for some.

Corey: Some regional expansions seem like they're focused on regulatory compliance. It's why there are so many regions in Europe, for example. Germany was always famous for this, and one of the reasons that they had the Frankfurt region spin up so early was that there were requirements for many German businesses that if you wanted to use Cloud, your data had to live within German borders. And a number of countries are enacting regulatory requirements like that. So, you see the typical regulated industries getting excited about that, but you also see folks getting much more genuinely excited when this is relatively underserviced market where you have no good latency options, and suddenly, now you do.

One of the weird announced regions that is not generally available, at the time of this recording, has been the Jakarta region, where, on the one hand, yeah, great. There hasn't really been a lot super close to Indonesia. There are also regulatory requirements that increasingly require data to remain within the country—and people love to skip over this part—Indonesia has a quarter billion people; it's not exactly a small country. So, when Cape Town was announced, clearly there's a sense of, “Yes. Now there's finally something local that I can talk to,” But was it also aligned with a regulatory story? Or is that less of a concern in a lot of African businesses?

Veliswa: From AWS side, there probably was alignment, but definitely for businesses that were waiting on that for regulatory compliance alignment. I know banks, definitely huge, and some telcos as well—you know, telecommunications and all of those—big for them to have the data reside within the borders. So, it held back. So, what I used to find, especially in the financial services is that they would migrate applications that are not as critical, but you have those that are very critical, where there's huge customer data in them, those kind of, were hanging back waiting, and they were just migrating not very essential applications to AWS. So, it was a big thing that we get that region here. And then they could start, you know, in a big way, doing their migrations.

Corey: This episode is sponsored in part by our friends at Linode. You might be familiar with Linode; they've been around for almost 20 years. They offer Cloud in a way that makes sense rather than a way that is actively ridiculous by trying to throw everything at a wall and see what sticks. Their pricing winds up being a lot more transparent—not to mention lower—their performance kicks the crap out of most other things in this space, and—my personal favorite—whenever you call them for support, you'll get a human who's empowered to fix whatever it is that's giving you trouble. Visit linode.com/screaminginthecloud to learn more. That's linode.com/screaminginthecloud.

Corey: One of the things that I find fascinating from an almost journalist style perspective, is when I look through the analytics for the announcements that I wind up including in the newsletter—to be very clear, I've built a custom analytics system on top of this: I don't see whether you personally have clicked any given link because to be perfectly honest with you, I could not possibly care less what you click in the newsletter. What I do care about is the aggregate. So, I learned that, cool, I don't care about IoT, but an awful lot of people do because a lot of people click links to articles about IoT, so I include it. I don't care about Windows and it turns out most of the readership doesn't either, so I don't spend a whole lot of time covering it. It helps inform the direction I go in when I talk about things.

Something that I've never been able to get a real sense of when curating the newsletter about what to include and what not to has been the story about regional expansion of services. For me, it always felt like, eh, great. They've announced this service that, I don't know, talks to satellites in space. Good for it. Maybe it has customers, maybe it doesn’t, maybe I don't care, but okay, now we've expanded that ridiculous service to an additional region in another part of the world. I've always mostly discounted that on the theory that people who are really going to be using that service will probably be using it either where it exists, or have a regulatory reason. Is that the wrong direction to look at this? I mean, is there excitement in the community when a service expands to a region where it previously did not exist?

Veliswa: I think it depends on exactly that: people in that region, do they care? Do they care about that service? I know when Capetown launched, EKS, for example, was not available here. And when it was announced, I remember the excitement because a lot of people here in containers, actually using that service, it was a big thing. So, I think it depends on the service; do we actually care about that service day in that region?

If we do care, then it's a huge thing when it actually does actually expand to that region. I think it always depends on the people there, are actually using it? Do they care much about it? So, it always depends on what actually gets announced. I think it definitely depends on the demand in that region.

Corey: What I don't understand, personally, is how the hell some of these regions launch without these services in place. And I'm sure I am inadvertently insulting some incredibly hardworking, incredibly talented people. But the South Africa region launched in April of 2020 and didn't get EKS until the beginning of August, four months later. Now, I know that the folks at AWS aren't sitting there going, “Huh. Maybe people might want to use Kubernetes. Oh, we never thought of that.”

Of course, they're going to want Kubernetes. But it seems to me, naively, that getting the buildings built out, getting all the power and interconnectivity set up, is a massive, heavy lift. Getting all of the equipment stuffed in is probably more complicated than frigging re:Invent. But, “Oh, we should really just turn on this service and push that code out to somewhere else,” seems like a relatively trivial slash solve problem. Do you have any idea why that's not true?

Veliswa: I don't know. Okay, I think this is where it crosses over to speaking on behalf of AWS.

Corey: [laughs]. Oh yeah. Oh, again, AWS Community Heroes are not AWS employees, in many respects. I've asked people before under the auspices of various NDAs, and the answer I've always gotten has been, “Well, it's really hard.” Which I get. Believe me, I get. And I know that people can't tell me the real reason that things don't roll out.

Veliswa: [laughs]. Yeah.

Corey: From my naive fly-on-the-wall perspective of needling AWS on Twitter, it's one of those things that I just don't understand. And the one reason I would actually be seriously interested one day working at AWS is not for any of the reasons people think, but mostly so I could just see underneath the hood all the things that annoy me about the company, the things that, “Why on earth, are you doing it this way?” I am certain that for every one of them, it's going to be, “Oh, that's why. That makes a lot of sense.”

Veliswa: Yeah. I know, we actually—because I know where I work, the web services that we were actually waiting on to have in Cape Town. So, we constantly kept following up on those, and there’s a roadmap that, kind of, guides these services being launched. So, just like any service becoming available in any region, it kind of follows the same roadmap, I think, from what I understand. So, I think it was a roadmap, and just waiting for things to actually, you know, be in line for being launched in your region.

Corey: The more than I talk to AWS employees, the more sympathy I get for the challenges that they have because every time I've seen some ridiculously easy, trivial thing that, “Why don't you support this one very simple, very basic thing?” I have never yet had a scenario where I talk to the service team who builds the thing, and their response is, “Holy crap. We never thought of that.” It's always, “Yeah. We would like that, too. Here are the list of constraints that were—just the ones we're allowed to tell you—that are causing a problem for us deploying that.”

And the response to the end of it is, “Holy crap. This is so complicated, how does anyone ever get anything done?” The answer, as it turns out, is they hire people who are way smarter and better than me at all of the stuff. But every time I start thinking, “Well, that sounds easy. Why don't they just do that?” It means that I haven't thought deeply enough about the problem, or there's a tremendous amount of context that I'm missing. And I have no doubt whatsoever that this is absolutely one of those issues.

Veliswa: Yeah, there's always a reason there that makes sense once you actually dig in, I think, yea, from what I've seen, as well.

Corey: So, you said you've been working with AWS yourself in the community for a number of years now. And you were recognized as an AWS Hero for doing that. What are you seeing as the community continues to grow, and expand, and evolve? I mean, when I got started with AWS for my first outing, it was 2009, give or take.

And my response was, “Holy crap. This is so many services, I'm never going to be able to learn them all. This is overwhelming.” There were 12. There are now almost 200. I can't imagine getting started today. What are you seeing as far as people who are discovering Cloud, or new to the tech industry, or new to their own careers, as they wind up discovering what this somewhat ridiculous, borderline nuts world is?

Veliswa: I see people who are being incredibly overwhelmed. It's actually so weird because it was only three years ago that I started learning AWS, and I didn't feel as overwhelmed as I feel today. And people who are learning it today feel that very same level of being overwhelmed. It's also, “Where do I start?” But I think what I'm starting to see, which I think is quite awesome, as well, there’s starting to be this focus, like, people will go in and they'll say, “I just want to focus on machine learning. That's going to be my thing on AWS.” Or, “I just want to focus on serverless, that's going to be my thing on AWS.”

I've met people who say they've just learned serverless, they don't even know how to spin up an EC2 instance. So, I think that also is kind of helping people with the adoption. If you just pick your area of interest, and you just go for that. I know people who are all about IoT, that's all they focus on; they don't really bother about anything else. And I think it helps them.

But definitely, you can very easily feel overwhelmed; I know I do. I’ve gone for all three of the associate level certifications, and to learn opers now, when you're going to study for a certification, you have to, kind of, touch everything. You can’t say, “Oh, not going to worry about EC2s, I'm just going to focus on this.” You have to worry about everything. And it is such a huge amount of work, and resources to consult, and everything to go through just to prepare for those exams. So, actually, it's very overwhelming. But I think if you just pick your area of interest, and you zone in on that, I think that helps as well.

Corey: One thing that always amazes me is that whenever I talk to AWS employees who are incredibly deeply technical on a particular area, it feels like I'm going to school on some level of, “Wow, I know absolutely nothing about anything you just told me.” And, “Wow, I'm an idiot and a giant fraud because you people are way better than I am.” They’ll be talking about some deep internal aspect of AWS networking or something, and then it's like, “Huh. That's the end of it. So, thanks. I appreciate your help on that. Oh, by the way, I have a question for you about S3.?” “Huh. What's S3?” It's, “Oh, that's right. You people can be very deep on things, but it's harder to be broad in many respects, too.” We've long since passed the point where I can talk incredibly convincingly about AWS services that aren't real and not get called out on it by AWS employees because who in the world can hold all of this in their head?

Veliswa: [laughs]. Yeah, and it's actually amazing you say that because I joined all these webinars, especially now that we've been at home and I always walk away feeling like, “Oh my word, I'm not smart at all. All these people out there are so smart.” The level of detail with a thing that we may be discussing on that day. So, yeah, but I mean, there's such a lot to actually get through. And I think at some point, you just pick that one thing, and that just becomes your thing and you just focus on that. Unless you want to [unintelligible] a certification, of course, then yeah, then you have to worry about everything.

Corey: Yeah. There's absolutely so much to keep up with, and it doesn't get any easier. By the time that this episode sees the air, you'll have had the dubious joy of sitting through a week of curating the newsletter, drinking from the firehose of everything that gets released. There are 40 distinct RSS feeds, now, from AWS. There's no aggregation system that they provide, so I had to build one.

And again, an awful lot of what they put out is not particularly relevant outside of a particular sector or market. So, it comes down to always being a question of what do I find personally interesting? There's no right or wrong answer about what gets included and what doesn't. But keeping up with all of those announcements really has forced me to understand exactly how many moving parts there are in something like this. I think it's impossible to curate a newsletter like this for any appreciable length of time, and not come away profoundly impressed by just the sheer scale of everything.

Veliswa: Yeah, for sure, for sure. And that's been another thing, you know, when I've mentioned on other platforms that I mentor people starting out, and it’s also the same response. It's, “Where do I even start?” There is so much coming out. I used to follow AWS This Week news every week. And at some point, that think I don't even follow regularly because the amount of speed that everything gets released and gets announced is quite a lot, and to try and keep up with everything is quite a thing. For sure.

Corey: I am extremely interested to see what winds up changing and how. It's going to be just an absolute mess—but a fun one—when we see how this continues to scale because at some point the pace of innovation, as they say, is only ever increasing. They're hiring more service teams, more product teams, they're shipping more than ever before, and at some point, no one's able to keep up with it all. The newsletter is a bit of a bulwark against the tide coming in, but the tide does come in, eventually. There has to be a change and I can't wait until one day they put me and my sarcasm completely out of business.

Veliswa: Yeah, let's see. Yeah, would be awesome. And I know I [unintelligible] as well. I'm like, “Wow, this is going to carry on for a while. Let's see.” But I think also because there's such a huge interest in new things coming out in technology update, it's almost like all these tech companies have to try and keep up with this huge demand because people want so much, and there is so much innovation that people are, kind of, waiting to do with all the services that come out. So, tech companies kind of have to try and—I think they also trying to fit this list, that is us actually wanting all these things.

Corey: There's no other way to frame it other than just it's absolutely massive. It's, again, I wanted to point out that Community Heroes are not paid by Amazon. These are people who are recognized for their contributions to larger community. Notably, I have never been invited to join the program because I am effectively the AWS community antihero--, but I'll take it. It's a tremendous honor and I think that whenever you see someone with a Hero title, one of the best responses is to listen to what they've got to say, I've almost never been disappointed by doing that.

Veliswa: Also because now you become a Hero, you join this community of other Heroes out there. And these are really impressive people out there. So, many of them, I was following on social media for a really long time. And then I get announced as a Hero, as well, along with them? That's crazy. And I'm on these calls with them, I’m like, “Is this really happening?” So, it's a huge honor to get to be part of it, for sure.

Corey: One of the things that continues to just absolutely blow my mind has been just the willingness of people to help others up. It's the idea of ‘send the elevator back down.’ It's incredibly encouraging to see. In that vein, what advice would you have for people who are new to this space on how to come up to speed, how to land a career that they enjoy and is lucrative for them? What guidance can you give the next generation?

Veliswa: Mm. Okay, let me go by what I did. And I'm not saying that what I did is what will work for everyone, but for me, I think just being open to learn things, just being open to that things are changing at such a fast pace, as well. Try—as much as you can to try and keep up. Just being open to learn, and being able to adapt to change around you, as well.

But definitely, that willingness to learn will take you quite far. People never believe that I'm from a mainframe background, and here I am today: I'm a cloud engineer. It's just that I've always looked for opportunities to learn new things and see what can I learn is the next thing that will help me remain relevant because you know the world of tech changes all the time. So, just always being open to learn new things, and just see. Yeah, just follow the right conversations, and conversations that grow you as well. And there's so many communities, as well, that are just there and willing and open to helping and teaching. Just finding those and following them. It helps quite a lot.

Corey: Thank you. If people want to hear more about what you have to say, follow your exploits as you continue to talk about the larger AWS community, or reach out for guidance as far as what they're seeing and see what you would advise them to do if they see echoes of themselves in your story, where can they find you?

Veliswa: I am on Twitter, I am on LinkedIn, I'm very busy on LinkedIn. I've just joined the Dev.to community as well, that’s where I do my blogging. So, those are the main platforms that you’ll find me.

Corey: Excellent. And I will include links to all of those in the [show notes]. Thank you so much for taking the time to speak with me. And again, thank you for letting me take some time off for the first actual vacation from the newsletter that I've had in the three years I've been writing it. Because, you know, a peaceful relaxing vacation involving a tiny newborn.

Veliswa: Yeah. Of course [laughs].

Corey: Yeah, great. That’ll be easy. I'm sure it'll just be like going on a vacation.

Veliswa: [unintelligible], for sure. [laughs]. This was fun. Thank you so much for inviting me. Thank you for hosting me. Thank you for the chat. I really enjoyed it. Thank you so much for it.

Corey: Thank you. Veliswa Boya, AWS Community Hero, and terrific mentor for folks who are looking to get into the space. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts, whereas if you've hated this podcast, please leave a five-star review on Apple Podcasts followed by a complaint after you've routed the complaint back and forth between Australia and wherever you happen to be, several times.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Stuart Miniman

Stuart Miniman has been an analyst and co-host of the online video interview program theCUBE for a decade. Stu's background is in networking and virtualization, he focuses on cloud and disruptive technologies. Stu has interviewed thousands of guests on theCUBE and written on wide variety of enterprise technology topics. His past positions, including sales, product management and strategic planning provides him with perspective on how to focus on the needs of customers. Stuart's previous employers include EMC (with a primary focus on storage networking and virtualization), Lucent Technologies (now Avaya) and American Power Conversion. Stuart holds a BS in Mechanical Engineering from Cornell University and an MBA from Bryant University.

Links Referenced:

  • theCUBE: http://theCUBE.net/
  • Twitter: https://twitter.com/stu
  • theCUBE Bio page: https://www.theCUBE.net/theCUBE-hosts

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Catchpoint. Look, 80 percent of performance and availability issues don’t occur within your application code in your data center itself. It occurs well outside those boundaries, so it’s difficult to understand what’s actually happening. What Catchpoint does is makes it easier for enterprises to detect, identify, and of course, validate how reachable their application is, and of course, how happy their users are. It helps you get visibility into reachability, availability, performance, reliability, and of course, absorbency, because we’ll throw that one in, too. And it’s used by a bunch of interesting companies you may have heard of, like, you know, Google, Verizon, Oracle—but don’t hold that against them—and many more. To learn more, visit www.catchpoint.com, and tell them Corey sent you; wait for the wince.

Corey: nOps will help you reduce AWS costs 15 to 50 percent if you do what tells you. But some people do. For example, watch their webcast, how Uber reduced AWS costs 15 percent in 30 days; that is six figures in 30 days. Rather than a thing you might do, this is something that they actually did. Take a look at it. It's designed for DevOps teams. nOps helps quickly discover the root causes of cost and correlate that with infrastructure changes. Try it free for 30 days, go to nops.io/snark. That's N-O-P-S dot I-O, slash snark.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Stu Miniman, who is a senior analyst and co-host of the online video interview program, theCUBE. Stu, welcome to the show.

Stu: Thanks, Corey, great to talk with you.

Corey: It's always great to talk with me. Can you answer something for me that took, basically, me hosting the show to really finally grasp fully? What is theCUBE?

Stu: Well, Corey, you know, most people out there are probably a little brighter than you, so they understand who we are, what we do. So, we are an independent media company. Our primary medium is video, and the way that we did things before 2020 was, I'd say, 80 percent at shows. So, big shows, little shows, the conferences, the ones that we all went to, collected swag, saw our peers, loved the hallway track. That's where theCUBE went.

So, shows like AWS re:Invent, VMworld, many other shows out there. Last year in 2019, we did about 100 physical events. Plus, we also have studios on the East Coast and West Coast and, order magnitude, we've done thousands of interviews. I personally have done a couple thousand interviews in 2019. I think I did, like, 470 interviews out of the about 2,000 that we did as a team.

So, we are part of SiliconANGLE Media, which is the umbrella brand, and that came between two companies that merged. One was a research analyst firm, Wikibon, which is where I had started about 10 years ago, and the other one is SiliconANGLE, which is an online Silicon Valley-based media company, which also still exists. So, we have those three pieces: there’s an analyst firm, there's a journalist team, and there's the video program, which is theCUBE. And, Corey, what’s tough is ‘theCUBE’ is a rather generic name and we are not in a box-like structure, so sometimes it's a little confusing on theCUBE. But we're an online video program.

Corey: I just figured you were a bunch of squares.

Stu: [laughs]. Very cute, Corey.

Corey: Thank you, I do my best. So, you had me on as a host a few times over the past couple of years. First, you had me on as a guest, which at some point is, “Hey, Corey, do you want to opine on this thing that we think a lot about,”—but in actuality, I don’t—“and get your picture, and face, and voice on a live videocast from a variety of different conferences?” And I am in love with the sound of my own voice, so of course, I'm going to say yes to something like that. And that was on me. I accept that. But you kept inviting me back and eventually let me host, and that, in turn, was on you.

Stu: Yeah well, Corey, one of the things I love doing is working on our programs. So, I have a few different titles, wear a bunch of different hats, and one of them right now is on the general manager of content. So, number one is I’ve always focused on putting on really good content. So, I'm as deep in the community as I can be, the technologies I know, the communities I know, the social media circles where I am, so I'm always looking for good guests, whether that is the founder of a cool new startup, some people in the industry—call it thought leaders if you will, sometimes—out there. And then I also like people with a good sense of humor, and that can dig into the topic.

So, I like to pull some of those in to guest host, and had you on as a guest a couple of times, and you did pretty well with it. And if I remember right, you said that the first time I interviewed you was really the first time that you’d done videos. So, that was surprising to me; you picked it up really well, and you obviously talk—

Corey: Well, thank you. I have a face for radio.

Stu: [laughs]. Corey, I've been doing this 10 years, I don't have much in the way of [air]. I'm a technologist by training; sitting in front of a video camera and having people watch me was one of the last things I thought I'd be doing. But I try to forget that there's all these people online watching, and dig into the conversations and enjoy it.

And I've always enjoyed talking with you on an off-camera about these things, so it made perfect sense to pull you into some areas where you're well known, like have you on at the end of AWS Summit, as well as some areas where you were poking fun at, like the Kubernetes community and it was lots of fun to do that show with you.

Corey: Oh, KubeCon was fantastic. It was a couple of days of, effectively, back-to-back-to-back interviews with a whole bunch of people. And it was fun because it helped me get away with some of my nervousness because on podcasts like this, if I say something stupid, we can cut that out and I can repeat what I said. Live video doesn't work that way, so I had some nerves. But the nervousness at being on video was completely swept away and buried because of the show we were at. And instead, I was focusing more on not being too insulting towards the nonsense that is known as Kubernetes to the people who built Kubernetes. And I think I pulled that off because no one actually mailed me poop afterwards.

Stu: Obviously, there was a little concern. You have poked quite a bit at the idea of multi-cloud, you have been insulting of some of the things there. But yeah, I mean, Corey, there were a couple of the analysis segments, I said, “Please go full Corey on it. Let's have this discussion. Let's go out in the open.”

It's still something that I've seen you do plenty. And I try to help people understand what is good about these technologies, what are the challenges about these technologies. Just because it's something that Google built doesn't mean that your team of five people are going to fall in love with it. So, absolutely you need to have an understanding of there. With my job as an analyst, I try to share information and get to the heart of things, and it's interesting space to look at, absolutely.

Definitely Kubernetes, boy, you look at some of the developer tooling out there, it is complex. And I feel for the teams that have to sort some of these things out, while constantly learning technologies and answer to whatever management is pushing down at them.

Corey: It's always fun talking to you about these things because you folks also take these small little clips from videos and put them out to promote various shows you were at. And one of them was, I think, me saying at the Google show that, “I’m not saying that Google is focused more on what they're building than what they're shipping, but even their conference is called ‘Next.’” and that, in turn, wound up generating a fair bit of traffic. And it's always interesting, when I wind up pulling out the list of things that seem to do super well, I always come in a strong second place because invariably, I wind up at the same shows where Abby Fuller is on, and she not only has a personality but also has something intelligent to say about technology. Whereas, again, I have a personality. So, she always does way better. I come in, usually second to her. But across the board, it's always been really useful just to see a combination of fun engaging dialogue on theCUBE, but also some hard-hitting analysis, that you wouldn't expect coming from a particular vendor’s show, you're not always particularly gentle; it's not paid access, which I'm envious of, to be perfectly honest, people generally don't invite me to things so I can make fun of them. At least not more than once.

Stu: I thought that is what you do, Corey, for quite a bit. But look, our team here, one of our earliest clients called us the ESPN attack, and that's the way we think of it. If you take your favorite sports team, there are the local stations that cover the team, and they tend to be a little bit biased. And then at its core form, what the ESPNs of the world did is, “I want to cover the sport, but I will poke at—here's the good things and here's the bad things about that team.” And if we're doing our best selves, that's what we are doing.

So, I am a fan of new technologies, helping people be more efficient and take advantage of these new things, but it does not mean that we should get too excited about the next shiny thing. And I am not based in Silicon Valley, so I try to have a balance of understanding the potential of new things as well as understanding the reality of what it takes to roll these out. And I've been in the industry long enough now that I feel I have a little bit of balance on that. But yeah, it's looking at these technologies; it's looking at all of these vendors, and everything that they're saying, but one of my favorite things, really, is talking to those practitioners: the people that don't care about what your slides say they care about what their business does out the various solutions help them do their job and make their businesses more successful.

Corey: So, one of the things that I've always found fascinating about you in that—far more so than I do, even here on a relatively agnostic show, I do have a bias towards the things that I work on, day-to-day. Given that your day-to-day involves going to a whole bunch of different events—you know, back before the dark times—and having conversations with basically everyone, you have a much better bird's eye view on this entire ecosystem than I do, just due to the nature of my own biases. So, as an analyst, what are you really seeing? Let's start, I guess, with the ever-popular topic of multi-cloud. Where do you see it?

Stu: So, it's an interesting thing. I've actually done podcasts talking about the quote, “analyst role,” because, Corey, when I came into this role, analyst was not something that I strive to be, and what I always—

Corey: No, analyst is something that happens to you.

Stu: [laughs]. Yeah. It is, what is your background? Why should I listen to what you're saying? When I turn on the television shows where they have people opining about it, you know, here in 2020, everybody is both an epidemiologist and a constitutional scholar, where most of the people talking are neither of those things.

So, my background is I have lived in technology. I’ve spent most of my career living on the vendor side, but even when I did, I actually never personally had a marketing role. I've lived in engineering, I actually did a stint in sales early in my career, and my job usually was to be, hopefully, the best part of consulting that you do there, it’s to help understand what you're doing, help you understand how the technologies fit. If it was a product that I was attached to, that was great, and if it's not well, we move our separate ways, and we go on from there.

So, coming to today's analysts’ world: multi-cloud. I think back to when the first cloud stuff happened. At that time, Corey, I was working for EMC. So, EMC storage company: most people consider storage kind of boring. I was sitting down the row from the guy that came up with the term—the first time I've seen it—private cloud and had arguments with him for months as to how private cloud wasn't really cloud, you know, I was having conversations online with the clouderati about how Amazon was going to take over the world. But back then, we came up with—when private cloud came out, there were these terms of hybrid cloud and multi-cloud, and the words made sense but if you look back a decade of what we were talking about versus where we are today, boy, a lot has changed.

So, in the last couple of years, I've looked at hybrid cloud, we’ve looked at multi-cloud, we’ve talked with the vendors and the customers. Thing one is when I talk to customers, Corey, they don't use those words. [laughs]. They might use cloud as a word, and they come up with a cloud strategy, but their deployment mechanism as to whether it is hybrid or multi, they don't really think about that. They have their data, they have their applications, they have their infrastructure, and they have the services that they use, and they kind of find themselves in a solution.

You talk from a multi-cloud standpoint, they usually have that. And it was a—one group, I had this really cool thing I wanted to use from one vendor, one other group we might have acquired, and they were using this other technology, and oh, yeah, are they using Amazon? Of course, they're using Amazon across these environments. So, unfortunately, like you said, you just become an analyst. Well, most customers kind of wake up one day and say, “Oh, I have multi-cloud.”

The definition I put in place is if multi-cloud is to be successful and a real thing, multi-cloud should be more valuable to the customer than the sum of its parts alone. And for the most part, I don't think there are many people that are there today. It is more a composite cloud, or I have pieces, or I basically recreated the multi-vendor world in today's multi-cloud environment.

Corey: This episode is sponsored in part by strongDM. Transitioning your team to work from home, like basically everyone on the planet is? Managing gazillions of SSH keys, database passwords, and Kubernetes certificates? Consider strongDM. Manage and audit access to servers, databases—like Route 53—and Kubernetes clusters no matter where your employees happen to be. You can use strongDM to extend your identity provider, and also Cognito, to manage infrastructure access. Automate onboarding, offboarding, waterboarding, and moving people within roles. Grant temporary access that automatically expires to whatever team is unlucky enough to be on call this week. Admins get full audit ability into whatever anyone does: what they connect to, what queries they run, what commands they type. Full visibility into everything; that includes video replays. For databases like Route 53, it’s a single unified query log across all of your database management systems. It’s used by companies like Hearst, Peloton, Betterment, Greenhouse, and SoFi to manage their access. It’s more control and less hassle. StrongDM: Manage and audit remote access to infrastructure. To get a free 14-day trial, visit strongDM.com/sitc. Tell them I sent you, and thank them for tolerating my calling Route 53 a database.

Corey: One of the biggest problems that I've seen has been that people always have these different definition in terms around what something means. You have companies that have a hybrid strategy; well, what that really meant, from my perspective, is that they had a data center, they were going to move at the cloud, they got halfway through because they were sold on a beautiful vision that didn't actually come to pass, and realized that, wow, some of these workloads are super hard to move, if ever, so we're going to give up, plant a flag, declare victory. And, “Yay, we're hybrid cloud now.” On some level, it feels like multi-cloud is aimed at that. In other cases, we see, “We're multi-cloud because one of our divisions is on Azure, and the other division is on GCP.” And those things all work for me. What doesn't work and has always driven me nuts has been this idea that I'm going to set out on day one to build something brand new—greenfield—and I want it to run seamlessly on every cloud provider. And I've yet to see something like that become successful. Am I just missing something?

Stu: Yeah, no. Corey, that's a fundamental question that we have about this whole Kubernetes space. So, I’ll relay what I've heard from customers. If I talk to a typical customer that running Kubernetes, one of the top reasons that they want it is that flexibility.

Say, the typical customer that's running Kubernetes, they're on AWS. And you say to them, “Well, are you looking to run on other clouds?” And they say, “Well if I have to, I want to know that it's not going to be too hard for me to do that.” I think of it more like I want the eject lever in case the plane is going down. So, if Amazon gets difficult on pricing if there's something else I need, but I think we're mostly beyond the, “Oh hey, I'm going to wake up every day, and let me just look at what the pricing is out there, what services are out there, and I'll just move things around the cloud.” Data gravity is very difficult, skill sets are not necessarily transferable, and even if you're using a handful of services from one cloud, going to another cloud, it requires some translation to go from it.

So, even using Kubernetes, I've talked to customers that have moved from a cloud to another cloud leveraging Kubernetes and can be done, but it is not trivial, it's not something like flipping a switch, it's not like the old days of [Emotion]. So, you're right, multi-cloud, the question I have is what are you solving for? If you want to have this layer there so that you can, if needed, make a change, I understand that. You have said are we, therefore, go to end with least-common-denominator cloud? I don't know that that's necessarily the case. You can still leverage services from a cloud, knowing that if you need to make a change that it's going to require changing, might not get all of the functionality. So, Corey, I understand some of the concerns and points you have about multi-cloud. The reality is most customers today are ending up with more than one cloud, and I don't see that changing.

Corey: Oh, I agree wholeheartedly. This gets into that same issue of people using different terms to mean different things. Where, for example, right now, I have all of my company email on GSuite, for a variety of reasons, not least of which being Google is freakin excellent at it. I use GitHub—or GifHub, depending upon pronunciation because, let's not get ourselves, I'm not ridiculous—and I use AWS for a lot of my serverless things because, well, they were sort of the first to market with one of the best offerings that integrates with everything else I do.

Now, each of those is a very different workload, so I wouldn't consider myself multi-cloud in the sense that some vendors do, but others will look at that and say, “No, you definitely are multi-cloud.” But then I see statistics that talk around things like that being used to prop up these, “Oh, we're going to give you a single dashboard into all of your different cloud providers. Look at what percentage of companies are using multi-cloud. It's 90-something percent.” I don't need my AWS spend, or status, or metrics on the same dashboard as I use to admin GSuite email. That is ludicrous.

Stu: Yeah. So, if I can tease out [unintelligible], Corey, there's one thing—number one is, I think we tend to—you know, I gave you the disclaimer, my background is infrastructure, and we tend to look from the, “What is the platform that I'm building on and what am I doing?” Well, if I flip this and say, “What applications am I using?” Well, if I'm using lots of different SaaS services, I don't even think about what cloud most of them are running. We've got Salesforce, we're using GitHub, there's lots of different places they could live, and if that vendor makes a change, it shouldn't impact what I'm doing.

Number two, there's a difference between saying, “I have a bunch of clouds that I'm using,” versus, “I have a good strategy on multi-cloud.” And then you bring up a really important one is, “What do I want to manage across these environments?” So, one of the things I said we're looking at is, what's the difference between multi-vendor of a decade ago and multi-cloud today? Corey, you've been around long enough. Remember, multi-vendor management, and does it give you cold sweats just thinking about the companies that came and went and tried to solve that problem? And it never worked well. So—

Corey: Oh, yeah. We wound up in a world full of finger-pointing, as vendors would each blame each other, and you wound up effectively having to go and yell at people on a consistent basis, just to get even the most baseline of errors debugged?

Stu: Yeah. When I think about, if you could get a tool that manages whatever device or software you're using, alone that was usually the good thing. Let alone, managing and changing everything on the way. So, management is obviously a big concern. Something I've been watching in the last year is all of these Kubernetes, multi-cluster managers, everything from what Google does with Anthos, and Azure Arc, Tanzu from VMware, and many others that are driving at this space.

And I've been trying to dig in and understand whether we are going to repeat the sins of the past, or if we’ve learned from these, and will they truly be valuable? Because of my background and networking, when I talk to my peers there, the joke always one is, if you want a single pane of glass, we know that that pane is spelled P-A-I-N.

Corey: Oh, yes. A single control pane of glass is also something that people have always been talking about. You see this, too, with companies asking for, “Oh, I want a developer dashboard where I can go ahead and see just the things I care about, and not have all the extra cruft, and do the things that are currently difficult and annoying, and do that in a super streamlined way.” People want them even for single providers, and it's virtually impossible to come up with something that works outside of one company's use cases, let alone something that will do that, and then span multiple cloud providers. It's always a question of trade-offs, and people don't value good user experience the way that they claim they do when it comes down to write checks.

Stu: Yeah, Corey, and you bring up a good point. You talk about the developer world and the traditional enterprise world, it is still early days of those coming together. I've spent a bunch of time over the last few years really trying to understand those community needs more. We've been expanding our coverage in this space. You saw serverlessconf.

Last year I was at AnsibleFest. My line I've used many times is, “Yes, I do have a closet full of hoodies that I’ve gotten at all these shows, but I don't try to fool anybody. I am not a developer.” But I believe I understand enough of those communities; I've read enough, I’ve talked to enough people, and taking what they are doing and driving in a company is something that big companies are still trying to get in line with and help move forward because really the objective of the day is for companies to be able to move fast and be more agile, and therefore they must work closer with their developer [unintelligible].

Corey: Absolutely. One of the biggest challenges that I'm seeing when I talk to people about this is the—I guess, how do you distill and consume all of the information that's being firehosed at you? You mentioned that you did—what was it—how many interviews last year?

Stu: About 470.

Corey: About 470. Each one of those more or less with a different company. We'll round it and say 450 different companies talking about things at a bunch of different events. And of course, theCUBE is larger than just you, unless you eat, sleep, and basically die in airplanes. And having met several of your team, I can attest the fact that node is not just you; sorry, to burst your magic bubble on that.

But how do you consume all of the different stories, all of the different narratives where everyone's vying for attention? It feels like it is a full-time job, and then some, to keep up with all of the products in this space and then come out with something that works for your use case, without, ideally, selling your soul to a particular vendor.

Stu: Absolutely, Corey. That is super difficult. Number one is trying to have a general understanding of some topics, let alone a little bit of depth. Absolutely, I am definitely broader than I am deeper on most technologies. There are certain ones that I've got plenty of background and history on, but there's many things that I'm learning about it through doing the interviews. And that kind of journey, I find there's usually a good appetite for.

So, yes, there's people that are trying to learn the 300 or 400 level on it, but there are way more people that can always use that 101 guide. From interviewing a person or company for the tenth or twelfth time, I'm usually going to be quite a bit deeper than the first time I'm learning about some wonderful cloud-native, networking, long-distance solution. But there is work that I do to get ready for it. You're absolutely right that it can be complete, overwhelming firehose of information at many shows. Especially, you know, take something like the Amazon ecosystem.

What was really just eye-opening to me is you get access to some of the smartest people in the industry. When I talked to James Hamilton, I talked with Adrian Cockcroft at the Amazon show, and you ask them a question about something at that show, and they're like, “I have no idea what that is.” And you're like, “Oh, good. If these people, other than Andy Jassy maybe, has visibility into everything and knows enough about it, and probably one or two, and you know way more about these things than the average bear, too.”

Corey: Well, I ran into the same problem. I long ago figured out the only way to credibility is exactly what you just said. Is to be one of those senior people who, when you're asked about something you're not familiar with, don't fake it because it becomes super obvious and doesn't look good. Now, on some level, you have to do the work. If someone asked me what my thoughts are about Kubernetes. And my response is, “What's that?” In 2020, that's not a great look for me keeping up with the entire ebb and flow of the ecosystem around us.

Instead, I have to do at least a baseline level of work. But when someone comes up with something that I haven't heard of, I find that saying, “Oh, what is that?” is the right answer. And honestly, only a complete jerk or someone with something to hide is going to pull a, “You haven't heard of Dingus X? Oh, you must be an idiot.” That's never true. There's always more stuff to pay attention to than any of us can keep up with.

What I look for when I'm having these conversations with different folks—and I imagine you, too, given the sheer volume of them—is recurring themes. The first time you hear about Istio, or Envoy, or one of the other various service mesh, you can dismiss it. But once you start hearing the terms again and again and again, you realize at some point, “Okay, there's something here, and I need to do a dive into what that might be.” At least that's how I operate. Is that similar to you?

Stu: It is similar, Corey, and it is challenging because since last year I did about 25 shows. There's certain shows that I've done for years, and therefore don't need to do a ton of prep, and there's others where—Kubernetes one, I think, for example, but there's a couple of hundred vendors there, and there are so many open-source projects that there's no way you're going to keep up on it. So, there's the areas where I spend more time on it, there's areas where I understand I'm going to stay pretty high level on [a thing]. I'm not going to become a Java expert out there. It's just not my thing, and there are other analysts out there that are always going to know more about it. So, I really liked what I heard at, like, the Microsoft show last year is I really want to be somebody that is in a constant state of learning because we don't like people that think that they are know-it-alls. I know enough in my career to know that I know nothing.

Corey: So, as you look across the ecosystem, I mean, obviously we're now in a surprise pandemic for the year, at least, where suddenly people aren't doing conferences and even once we start to see the restrictions relax, it feels like there are very few conferences that are worth risking death to attend. So, I don't imagine that it's going to go back to normal for another year or two. Are you seeing the narratives change as people are moving towards webinars, webcasts, on-demand events, difference in conversations where people are maybe emboldened, now that they're having a conversation when they're not standing close enough to get slapped by someone, or is it more or less returning to a business as usual story?

Stu: It's still pretty early, Corey. So, one of the challenges right now is every week, the narrative changes, some. So, I've seen some large companies that really put a freeze on things, they focused on essential businesses, they're helping their customers, everybody in IT, I'm sure, helps out some essential businesses, whether that be medical specifically, or all of the underlying infrastructure pieces that go. Some companies are relatively fast returning to business as usual. They're doing the product launches, they're beating the drum how they're the best thing ever, and ended up leader in some analyst’s report. So, you're slowly seeing a little bit of change for some, but not necessarily for others.

Corey: For the record, I do want to point out that we're recording this on April 27th, which is—I have a bit of a backlog, so it'll be a bit before the show airs. And given how quickly that whole chain of events is unfolding, I want to make sure that people aren't sitting there listening. “Wow, they didn't talk about the giant meteor at all.” Yeah, exactly. We never know what each one of these progressively worse-feeling weeks is going to have in store. So, as of right now, that is the state of things.

Stu: Yeah.

Corey: Three months from now, when you say these are early days, I just want to make sure people aren't blaming you, but rather my slow-ass production process.

Stu: Yeah, Corey. Thank you for making that disclaimer. That's a really good point. It is really hard to believe that, yeah, as you said we're sitting here at the end of April, two months ago, I was on a family trip outside the country. The joke we have is that there used to be the line, “I’m disappointed how much changed in the year but amazed what happened a decade.” And boy, March 2020 was an insane decade. And we'll all remember that month, to be sure.

Corey: Oh, yeah. Like, to peek behind the scenes a bit of this podcast, I wound up with at one point having a one more week ready to go and nothing else, and my podcast production team was screaming at me. So, I sort of overshot and now have a bunch of items in the hopper. That usually all well and good because we talk about broader themes. We're not talking about last week's news on this podcast, but an awful lot of these were recorded before shelter in place and social distancing. So, relistening to some of these things, we sound like ridiculous Pollyanna types, where it's, “Oh, yeah, we're looking forward to a great conference season.” Yeah, not so much as it turns out.

Stu: Oh, yeah. Trust me, Corey. We had a product launch we did and got lots of people online. “Why do you have two people sitting next to each other?” It's like, “Because it was recorded three weeks ago.”

And absolutely, we're doing 100 percent of our interviews remote right now, watching and listening to all the governors talk about when social distancing will be changed. But I'm not ready to make any statements about what the long term viability of conferences are for 2021. You know, sitting here at the end of April, yeah, I wouldn't be shocked if I'm not at a conference in 2020 and, as somebody that goes to a lot of them, boy, we're watching at the ripple effects of this.

Corey: Absolutely. And these things are moving super quickly. Now, it's one of those great—I always tend to avoid video because one, it's expensive and difficult to get right, and I don't believe in doing things in a crappy, mediocre way if I can avoid them, so whenever I did video, it was always, “Oh, great. I'll have a team of people who actually know what they're doing, and I stand here and look great.” Now I'm forced to learn how all of these things work, and, whew, there's a lot that goes into it. It definitely makes me rethink my choice not to participate in A.V. club back in school.

Stu: [laughs]. Absolutely. But Corey one other thing, you talked about change here. How much will the current global situation stop certain projects, or accelerate certain projects? I’ve seen certain parts of the world that might have been a little bit slow on adoption of public cloud, and now they're saying, “Well, hey, I can't go touch my gear, really, we need to spin these things up in the Cloud.”

I'm curious to hear the operators of the hyperscalers what's really going on inside there. Here in the US, we hear what happened in some of the large meat-producing companies; there's concerns about the supply chain. I haven't heard much about—you think about how many servers the Amazons, Googles, and Microsofts of the world spin up on an average week, and they still seem to be clicking on all cylinders, and moving forward.

Corey: Yeah, it really is interesting, on some level you feel for some of the product and service teams are expecting a big reveal at various events and looking forward to getting out and telling these stories, and now it's, “Nope, nope, nope, just kidding. You get a blog post.” That does definitely sting, but on the other hand, it feels like if you go down the list of all the inconveniences this situations having for everyone, that is so far down the list, it's almost not worth mentioning about, except it really does reflect the culmination of 18 months of work for a lot of people.

Stu: Yeah, I guess my point was that the person that's still expected to open up the new data center and keep capacity growing because the Cloud is always auto-scaling, and all of these new companies that meet online are concerned less about that new feature, but we're about, “Wait are we going to have some outage in the Cloud that’s going to be related to COVID-19?”

Corey: Wh0o, boy. Yeah, or failure to invest appropriately a few years ago. For example, we saw the Azure capacity constraints. I looked through Microsoft's old security filings and identified risk factors I never saw, “People might take us seriously when we say we have a hyperscale cloud and try to use it that way,” as being identified as a risk, but here we are. On the plus side, they do get to say they have the most regions. On the other side, it looks like each region is a couple of racks in some random data center, and oops, it's full.

Stu: Wait, I can't have a cluster with a single node, Corey? Is that what you're saying?

Corey: Well, that’s what virtualization is for.

Stu: [laughs]. Fair enough.

Corey: So, if people want to learn more about who you are, what you do, and the various other things you have to say—some brilliant, some insipid—where can they find you?

Stu: All right, Corey. First of all, the Twitter is still a good place. It has not completely imploded in 2020, at the recording of this interview. I think Corey, do I have the best Twitter handle of any guest you have? Because it's just S-T-U, my name?

Corey: Well, that is a terrific Twitter handle, but if, and only if your name is in fact, Stu. Otherwise, it really doesn't fit.

Stu: Yeah, it was funny, I actually—one of the only times I’ve really trended on Twitter with, I think I had 10,000 likes, and all these responses, the person that owns @stuart and myself had a Twitter conversation, and it seems both of us would have originally preferred to have the other handle, just something about the way Twitter handled it, and the like. So, that was a funny one. So, Stu is my Twitter handle. theCUBE.net is the website for where we do everything with the video. And if you hit one of those sites, you can find much more about me. My bio is [unintelligible].

Corey: Excellent, Stu, thanks again for taking the time to speak with me. I appreciate it. I'm sorry, we couldn't do it in front of a camera like we normally do.

Stu: Yeah, Corey, hope to see you again in person in the not too distant future. And yeah, please stay safe.

Corey: Stu Miniman, senior analyst and host of theCUBE. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts, and if you've hated this podcast, please leave a five-star review on Apple Podcasts and a ten- to fifteen-minute video of you debating exactly what I got wrong.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Angela Andrews
Angela Andrews is a solutions architect at Red Hat. Prior to that, she was a systems administrator in higher education for over 15 years. With many interested in technology, she’s dived into areas like cybersecurity, where she was a substitute teacher and teaching assistant (TA) for a cybersecurity boot camp. She attended a full-stack coding boot camp and also taught and TA’d classes teaching people how to code. Angela is the organizer for #PythonForAll, an online meetup called where people learn how to program in Python. She’s given talks on topics like self-care, WordPress, AWS, and also is a contributor to Women Techmakers YouTube video series.

Angela is married with two sons and a dog whom she adores, named Scout. In her spare time, she likes to read, lift weights, SPIN, swim, learn new technologies, and occasionally blog.

Links Referenced:

  • Connect with Angela
    • Twitter: @scooterphoenix
    • LinkedIn
  • Red Hat
  • Sponsors
    • Catchpoint
    • StrongDM
    • Linode

Transcript
Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Catchpoint. Look, 80 percent of performance and availability issues don’t occur within your application code in your data center itself. It occurs well outside those boundaries, so it’s difficult to understand what’s actually happening. What Catchpoint does is makes it easier for enterprises to detect, identify, and of course, validate how reachable their application is, and of course, how happy their users are. It helps you get visibility into reachability, availability, performance, reliability, and of course, absorbency, because we’ll throw that one in, too. And it’s used by a bunch of interesting companies you may have heard of, like, you know, Google, Verizon, Oracle—but don’t hold that against them—and many more. To learn more, visit www.catchpoint.com, and tell them Corey sent you; wait for the wince.

Corey: This episode is sponsored in part by strongDM. Transitioning your team to work from home, like basically everyone on the planet is? Managing gazillions of SSH keys, database passwords, and Kubernetes certificates? Consider strongDM. Manage and audit access to servers, databases—like Route 53—and Kubernetes clusters no matter where your employees happen to be. You can use strongDM to extend your identity provider, and also Cognito, to manage infrastructure access. Automate onboarding, offboarding, waterboarding, and moving people within roles. Grant temporary access that automatically expires to whatever team is unlucky enough to be on call this week. Admins get full audit ability into whatever anyone does: what they connect to, what queries they run, what commands they type. Full visibility into everything; that includes video replays. For databases like Route 53, it’s a single unified query log across all of your database management systems. It’s used by companies like Hearst, Peloton, Betterment, Greenhouse, and SoFi to manage their access. It’s more control and less hassle. StrongDM: Manage and audit remote access to infrastructure. To get a free 14-day trial, visit strongDM.com/sitc. Tell them I sent you, and thank them for tolerating my calling Route 53 a database.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Angela Andrews, who's currently a solutions architect at a small company called Red Hat. Angela, welcome to the show.

Angela: Hey, Corey, how are you? Thank you so much for having me on.

Corey: No, thank you. And what makes this a little bit strange is that I'm standing, sort of, in two points for this conversation temporally. At the time that we're recording this episode, everything is still in a planning stage, but by the time people are listening to it, you are this week's guest author of Last Week in AWS because I am out on parental leave, where probably at the time people are listening to this, I'm surrounded by screaming and diapers, and then when I'm done with my morning routine, I have to go deal with an infant.

Angela: I'm so happy that I get to help you out in your wonderful time in your life, and I'm very excited to be doing this. So, thank you.

Corey: Of course, it's always a challenge because the first year that I went down the path of building out a newsletter, and then, “Oh, hey, turns out I'm expecting a kid.” I took one week off, but then I could find time to wind up doing the newsletter on the side. It's not really a side project anymore in the way that it once was, so all right, how do I wind up not having to, effectively, give family time in service of basically entertaining the masses about cloud computing, which compared to those two things doesn't really went out as far as priorities go. So, thank you for making it possible for me to take, I think, what is my first real time off since I started this thing.

Angela: I'm really happy to do it. Thank you.

Corey: So, you are currently a solutions architect, but before that, you were a systems administrator in higher ed for over 15 years.

Angela: Yes, I was and it's a very interesting transition. It's actually a very logical transition. Being a systems administrator is challenging, and it never stays the same, and just moving into this space where you are now not beholden to a whole community of faculty, staff, students, parents, administration, but you're working at a software company where your reach is global and you have the opportunity to touch not just one university, but many universities that I support. So, it was a very logical shift. I didn't know it at the time, but I am so happy to be here. This was the best thing that could have happened to a diehard first Windows, then Linux/Windows, and then systems administrator. And now I am a solutions architect. So, the transition is phenomenal.

Corey: So, the path that you've taken is somewhat familiar to me. My first job that—let's not kid ourselves—I bluffed my way into as a Unix administrator was at a university. And at least at the small university I was at, their computer science department was, eh, roughly six people. So, it wasn't at a scale where it had massive technological progress. So, yeah, it turns out that someone who's super loud, and it just so happened—I'm not making this up—the technical reviewer was out sick that day, and I sounded confident, was enough to basically land the job. I crammed like absolute hell and spent 16 hours a day learning the things I should have learned, for the first three months I had the job.

I don't recommend this, by the way, but that's what happened. And then I realized, huh, somehow I’d learned enough to surpass a number of people that I'd already been working with. And I was already learning early in my career that my personality and higher ed, don't particularly get along. I'm a terrible student for similar reasons. But I was only there for a year before the siren song of private industry wound up coming out, luring me away into places where I could, you know, have my own personality defects manifest differently. You stayed there for 15 years. What was that like?

Angela: Well, it was interesting. I stayed so long because working in a higher ed institution, your children get to go to school for free or on a discount. So, there's that draw that keeps you there longer, probably longer than you should be, but it is a really interesting draw. And as luck would have it, I think most of us have bluffed our way into higher ed, wouldn't you say? So, yeah, I started out in corporate and decided to pivot into higher education.

And it was an interesting transition coming from corporate to go to higher ed. You really need to loosen up, and slow down, and realize that things ain’t the same. Nothing is the same anymore. There's bureaucracy, but it's a different kind of bureaucracy. And it took a while to, like, unbutton the top button, metaphorically speaking because they just ride roughshod in higher ed. I'm joking, but I'm not joking, you know what I mean? It's just different.

Corey: One of those, “Haha, only serious,” type of comments.

Angela: Exactly.

Corey: Yeah.

Angela: JK, JK. [laughs].

Corey: One thing that really shattered my faith in higher education, on some level, was working within it. And I'm not trying to be rude on that. I come from a family of teachers, and I was always brought up to lionize higher ed, which is why, basically, having an eighth-grade education on paper was such a disappointment to my parents. Now, I'm a disappointment in other ways. But going through that it was always higher ed is the way and the light.

And then I started supporting a bunch of professors with PhDs. And it turns out a common failure mode there is in this one extremely narrow segment of a very deep field, I am one of the world's leading experts in this thing. Which means I'm probably a world-leading expert in everything else that you do, too, because they don't even offer PhDs in making the printer work. How hard could it possibly be? And I joke sometimes about Googlers being condescending, but they've got nothing compared to tenured professors who want the computer to go.

Angela: Nothing. Absolutely nothing. That's funny you say that because I always championed higher education myself. So, much so, I wanted to guarantee that my kids had the opportunity to go, and I wouldn't go into the poorhouse doing it. But once you're in there, you realize that it's a bit of a sham. Like, not a sham.

I think because there's so many careers nowadays, with the advent of the internet and so much information out there, you don't need a college degree to be successful in a certain field. I think putting people in debt for hundreds of thousands of dollars for a job that they really can't pay off their student loans seems like a horrible purpose in life. And the fact that the cost of education has increased 400 times—and I've seen numbers over the years, but people could actually work their way through college in the 60s and 70s, and things like that. You cannot work unless you worked at Google to be able to pay your way through college. How can that be sustainable?

And I think that question is—those chickens are coming to roost. I hate that phrase, but COVID is really going to expose how a higher education is effectively educating our next generation. So, with internet information being so prevalent, COVID, so much online learning where you don't have to spend $50,000 a year to learn something because there's just so much information out there. I think we're in a very interesting time for higher education. And I'm interested in seeing what the next three years, next five years flush out. Things are going to be different after this, and I hope it's for the better because people go broke trying to get a higher education. And I think that's very unfair.

Corey: We see it beyond that in recession times where you have this entire generation—at least I consider myself part of this generation—that was sold a bill of goods in that, “Oh, go to college and get a bachelor's degree. It doesn't matter what the degree is in, just get a degree. It's a credential, it guarantees you a job.” Yeah, funny story. Turns out it doesn’t.

And so people didn't know what they wanted to do, so they picked degrees and things, and okay, fine. I'll get a degree in English or philosophy—which was one of my majors once upon a time, but I was trying to get through college and failing—it was okay, as long as I have the degree, that'll guarantee me a job. And then people graduate with this with a hundred-twenty grand in debt or whatnot, and, okay, now what? I can't get a job. And very often the response is, ah, you need another degree.

And at some point, it's one of those, stop: you're going through this completely backwards. When I talk to folks who are having trouble finding a job and considering going back to school, my approach is always the same. Stop for a second, figure out what job you want to have, find people who do that job, take them out for coffee, meet with them. Figure out what would they recommend you do here? And then oh, they say, “Oh, yeah, if I were going to try to do what I'm doing now, yeah, you're going to need to get a degree—or not get a degree.”

If a job you want is you want to be an anesthesiologist, yeah, you're going to need to go back to school. But if it's going to be running systems at scale for web properties, maybe you kind of don’t. Figure out what it is that's holding you back, you are going to need a piece of paper that says you're capable of doing things. But a degree is increasingly not about that, and was, arguably, never intended to be. A certification might be, but the right answer really is a resume showing the things you've done this before, which winds up conveying, I guess, a sense of competence. The hard part is, how do you bootstrap that when you don't have it to begin with, and you're starting from scratch?

Angela: That's true. Well, I agree with you about being sold a bill of goods. I mean, I fell for the okey-doke myself. It was, “Oh, you have to get a college degree. You have to because how else will you get a job? How else will you be able to find something that you want to do? You don't want to work as a—fill in the blank—for the rest of your life.”

So, I fell in lockstep. I tried a bunch of different things and decided that I didn't like any of them, but I was able, [laughs], I was able to get through college, and at the end of it, I realized that I really didn't need this degree to do what I'm doing right now. I work in IT. I really don't think that four-year degree helped me be a better systems administrator. It did help me in other ways: I think it maybe helped me be a better thinker, a better planner, a better time management person.

I think it helped in some respects, but I didn't think I had to pay thousands of dollars to learn those skills. Now, the funny thing is, for some jobs if you want to work in IT—you have to have a college degree to work in higher ed. I think their motto—or mantra—could be, “Yes, this person works in IT, but we have this many people with bachelor's and this many people with master's degrees working on our staff, so we value education.” And I don't need to see my name with my degrees after it in a staff catalog because who cares? You're going to call me when your printer doesn't work, you know what I mean?

It's really interesting how they bolster that up. And you said something else that I thought was interesting. People want to degree their way into a job and you said, stop. Well, I learned something—I'm going to say it—I learned it in college where if you want to find something that you want to do—because you go to college, sometimes you don't know.

There's something called an investigational interview. So, you find someone, you look them up on LinkedIn, you get a referral, and you ask them for coffee. And I've done this. I wanted to be an attorney, so I was studying for my LSATs, I was doing paralegal work for the ACLU, and I'm like, “Do I really want to take this next step?” Mind you, I'm an adult, and I would be an adult learner going to night school, basically, to become a lawyer.

So, I spoke with people who had done this, and everyone's story is different, and I found out that you do not have what it takes to work all day, to be a mom, to be a wife, to be a pet parent, and then try to go off and go to law school at night. So, I only did one semester in law school and I just threw in the towel. I said, “this is not for me.” So, I didn't waste too much time.

But I did at least investigate. I discovered what kind of lawyer I could possibly be. Like, where my strengths were, where my weaknesses were. And those conversations, I thought were very helpful. Now, of course, you cannot get in Google and be an attorney. That's not something that you can do. So, I get it.

Corey: See you say that, but giving legal advice on the internet does not the subject itself to the same credentials as you would expect.

Angela: It sure doesn't.

Corey: I swear, I'm not making this up: I am one of the moderators of Reddit’s legal advice subreddit for proof of that.

Angela: You're giving legal advice?

Corey: No. Mostly it’s how to be an adult advice of, “Stop talking to the police. Yeah, you're going to want to hire an attorney. And no, you can't sue a dog.”

Angela: You can't?

Corey: Well, not successfully.

Angela: Oh, okay. [laughs].

Corey: No, and in practice, though, I thought something much the same once myself of, “Huh, maybe someday, if I ever learn to become a good student, I could, I don’t know, become an attorney.” It was a great example. And it turned out I married one, and when I said, “Would you think I'd be a good attorney?” And she laughed and laughed, and then realized I was seriously asking, and started screaming, “No, don't do it.”

And yeah, when you see folks who I know have gone through the gauntlet, get that thousand-yard stare of no, as they remember all of the pain and tribulation of getting somewhere, it's, “Huh. Maybe my understanding of this thing is naive.” I have saved myself untold hundreds of thousands of dollars over the course of my career, just by asking people who've already done a thing, whether it's anything like I imagined it would be. Yeah, it turns out that being a lawyer is nothing at all like it appears on Boston Legal.

Angela: Nothing. I was an LA Law fan when I was younger, and I just knew that that was the bee's knees, “Oh, this is amazing. Championing for the little people, that'll be awesome.” It is nothing like that. It is so much work.

It's nothing like you would expect it to be. But again, that is one of those things where you don't go into it lightly. And that's not the only career that you don't go into lightly. You should do your homework, you should talk to people. And I think talking to people is a vehicle for many things, not just career advice, but for just getting to understand people, and getting to know people, and being able to relate to people that are not like you are not in a position like you. So, I'm an avid… Tweeter? Twitterer? I don't know what you call it, but—

Corey: Thought leader.

Angela: Oh, there. I am an avid thought leader on this platform called Twitter, and I enjoy talking to people that aren't like me because I think I've grown—should I say that? Should I say that I think I've grown since I've become a thought leader on Twitter? Please don't quote that. Because that sounds so pompous. But—

Corey: Oh, that's going to be a pull quote when we wind up—

Angela: Oh my God.

Corey: —doing the tweets about this.

Angela: Oh, my God. But I do. I think I've learned a lot from people by interacting with them on social media. And it's not always things that I like about them, or things I like about myself, or things that I like about the world that I live in, but there's a lesson in it. And I think that has allowed me to be a much more empathetic person, a much more thoughtful person.

So, I think that in and of itself has helped me become a better person overall, away from the keyboard, away from the Twitter app. It just helps me interact with people better because not everyone agrees with me. Not everyone has lived my experience. And I think sometimes social media can be the great equalizer—you know what I mean—if you're willing to put in the work and listen to people. And it's also a great place to build communities.

There's something that binds you to someone else somewhere on social media. And if you're lucky enough to find that, and then find a group of people who believe and like the same things you do, that community can be very encouraging. It can be very fulfilling. And that's what stemmed from a group that I started earlier this year because I was learning Python. So, I dabbled in web development over the years, and programming and I wanted to start using Python, especially at work.

So, as a systems administrator, Python is a great tool for automating all the things. So, I'm learning it with a guy that I met on Twitter—shoutout to Will—and we decided, “Wow, this would be great. More people probably want to learn Python too. We should probably put a post out and see who wants to learn with us. I mean, it's just two of us. Let's see if there's other people interested.”

Man, my Twitter blew up. I didn't realize so many people wanted to learn Python. So, we started this little group called Python For All. And it means Python for all the things. You know, Python touches everything, it touches Cloud, it touches automation, it touches cybersecurity, it touches ML, data analytics. So, Python For All, that's what I meant by that.

And luckily, we were able to get people—who wanted to be group leaders—manage these online meetups. So, we have four that go on different time zones, different days, and people just get together once a week and learn Python. So, that is just an example of the good that can come out of social media and extending yourself and building communities. So, it's not the cesspool we all think it is all the time. There was a little glimmer of something special about being a thought leader on Twitter.

Corey: This episode is sponsored in part by our friends at Linode. You might be familiar with Linode they've been around for almost 20 years. They offer cloud in a way that makes sense rather than a way that is actively ridiculous by trying to throw everything at a wall and see what sticks their pricing winds up, being a lot more transparent - not to mention lower - their performance, kicks the crap out of most other things in this space, and, my personal favorite, whenever you call them for support, you'll get a human who's empowered to fix, whatever it is that's giving you trouble. Visit linode.com/screaminginthecloud to learn more. That's linode.com/screaminginthecloud.

Corey: There really is. And one thing that I found that works, that is certainly an expression of privilege is getting something wrong on Twitter means I will, in fact, be corrected by people on the things that I got wrong. And mostly, it's not going to be calling me a trash goblin. Now, that doesn't apply, depending upon followers, and the rest, and what circles it winds up going big in, but I've been a big fan of the idea of learning in public on some level, I do worry that that's incredibly problematic for people who don't look like me.

Angela: That's true. So, you usually get that lesson pretty immediately.

Corey: Well, to be clear, when I talk about learning in public, I will sometimes ask questions such as—like, at this point, I built up enough of a reputation for knowing what I'm talking about that I have no problem admitting out loud the things that I don't know. In fact, when I was a hiring manager, I was always viewing the differentiation between senior and not, was being willing to admit you didn't know something. And, “Yeah, I don't know how this thing works. What's the best way to do this thing?”

And I get remarkably few people telling me that I'm a moron and a fraud for not having the answer to that particular thing. I get a lot more people chiming in asking how they can help, which is encouraging. Other approaches that are also high risk that have done… a bit of a mixed bag results for me were talking about how Git works in a conference talk. I submitted a talk called Terrible Ideas in Git, and it got accepted. So, yay, I can use this to teach people Git. It's in four months, what do I need to do first? Probably learn how Git works. And that's sort of a forcing function that inspires you to learn things, but that has a failure mode that is painfully embarrassingly public.

Angela: A much higher threshold than just learning in public. That's totally different, I agree with you.

Corey: Oh, but if ask on stage, and someone raises their hand and—rightly or wrongly—calls bullshit on everything you've just said, the only response that I've ever needed for that is, “That is a great point. A little out of scope. Come and talk to me after.” Spoiler, no one comes and talks to you about that stuff after. They just want to look smart in front of the rest of the room.

Angela: I have to use that. Let me write that down.

Corey: Oh, yeah. It works super well because there are three things that are not questions, and when I've had enough of people's nonsense—back when we would speak on stages—my last Q&A slide didn't say, “Questions?” It said, “Three things that aren't questions, reading your own resume, comments, and calling bullshit on my entire presentation.” You can, in fact, condense all three of those down with a great compression algorithm into one bullshit onion called, “That's not how we did it at Google.”

Angela: [laughs]. That’s—I’m writing this down.

Corey: Oh, by all means.

Angela: Three things that aren't questions. I've given a couple talks in my time, and most people have great questions, and I've been lucky enough that I think I've been able to answer some of them. But yeah, you're right about—it gives people the opportunity to tell you how much they know, more than what you know because they know more than you know. So, I've had that, and I think that's very interesting. It's like sticking your chest out like you're a chicken and you're walking around.

So, I think that's a very funny analogy. I think learning in public is a great opportunity. And you're right, people are very helpful. I find I've had issues where I'm a person that really—it's like, “Oh my God, I have to ask someone for help.” And I'm not really sure what that's about, but it pains me to have—let me figure it out. Let me figure it out. It pains me to have to ask because it's like, it's probably easy, and someone's going to say, “Well, you didn't know that?” That's one feeling that I really don't like, but—

Corey: A question like that tells you a lot more about the person asking it than the person they're asking it to.

Angela: That's true. And I have no problem most of the time admitting what I don't know, and you don't know what you don't know. So, I've put myself out there on social media quite a few times. I wanted to learn something, I wanted more insight into something and people that don't even follow me have reached out to me and been amazingly helpful with something that I'm learning. And I felt as if, on the other side of it, wow, I don't think I would have come at it at this angle had I not spoken to X, Y, and Z.

So, I think there's an opportunity to put yourself out there sometimes, even though sometimes it might pain you, you don't know what opportunity awaits, you don't know who's going to read that post, you don't know if someone feels like they're the perfect person to give you that little gem of information that you're looking for. So, I think it's a great idea. Don't be afraid to learn in public because you really don't know what you're going to get, but it usually is something positive and worthwhile.

Corey: That's sort of how I wind up seeing a lot of things. It's the, learn in ways that help inspire other folks. I mean, I've got to admit, I am not someone who learns well when I sit there and read a blog post on something, or I sit in a class, or I go to a workshop. For me, I've got to have a problem that I need to solve, and then I need to stumble my way through it, as far as solving it.

And I wind up learning things that I would have caught earlier if I had the ability to learn more effectively in a classroom. Like, “Int and a string? What the hell is the difference between that? Aren't they the same thing? One’s numbers and one has the numbers and other stuff?” What's the problem and I sound like, effectively, Fisher-Price teaches programming, more or less, as I stumble my way through awful implementations that eventually I batter into working. But the online meetup you talk about of Python For All sounds like it would have been something that worked super well for my type of learning style.

Angela: I think it has helped quite a few people. I mean, again, I've never met, maybe 98 percent of the people that are a part of the community. And what I've found is when you get on these meetups, and you hear what activities they've done during the week, or how they've solved a certain problem and you get to look at their code, some people get those aha moments just by looking at someone else's code or talking through a problem. And again, those are different learning methodologies, different learning tools that may be us assigning something; it really didn't click for you. You get the assignment, and you look at it, you're like, “Oh, wow, okay.”

But you come to the next meetup, and you hear the conversation that people are having about how they worked through it, or how they were able to solve it. And you get to see different solutions for the same type of exercise, and that is interesting. I love seeing that aha moment because that opens up to more understanding; that opens up to being less afraid, and those are the people I find who as the weeks go on, they are emboldened because they've seen the light. They've seen that there's a different way to approach a problem. And those Python For All people, they shine, they shine.

And as the weeks go on, I'm amazed at sometimes you just need to see something. It can be the same way multiple times, but like you said, depending on how you learn, it doesn't matter. Once you get that aha moment, you're in great company, and you're unstoppable. And then who knows what's next? Who knows what you learn next? And who knows what you can help someone else learn next? So, yeah, that's great.

Corey: I love that you're having this conversation with me now before you've written your guest issue of Last Week in AWS because the entire nonsense production system is beaten together with Python, learning the way I just described, and it is monstrous. And I'm going to bet that when this publishes, you will immediately take to Twitter repudiating everything you have just said on this. Like, “Yeah, you know I said it's important to be empathetic and help people learn the way that's best for them. Yeah, I take it all back. Corey is rubbish at code, and should never be allowed in public or within 200 yards of a computer ever again.” I'm assuming that that's the outcome we're going to get to.

Angela: Well, yeah. I'll probably say, “Scratch that I didn't mean any of it.” You know—

Corey: Exactly.

Angela: [laughs]. Well, we'll see.

Corey: Like, “No, it's very important to help people learn in the way that they learn best. Except for Corey. Never teach him things.” I'm willing to be the exception that proves the rule.

Angela: I'm going to put you on the ‘do not learn’ list.

Corey: Exactly. [laughs]. Thank you, once again, both for being on this podcast, as well as taking the burden of the newsletter off of me this week. It’s really appreciated.

Angela: Well, I'm really happy to help. I am a follower of the newsletter. I was introduced to it by a former boss, and it has been very helpful in teaching me new things and opening my eyes to new services, so it really was a great introduction to meet you, and to talk to you. I actually think I met you briefly. I think it was at re:Invent in 2019. I think I kind of said, “Hi,” when you were putting stickers down on this table, you probably don't remember it because there was—

Corey: That is entirely possible. It turns out that at re:Invent, I wind up getting firehosed half the readership of the newsletter in about a 20-minute span. So, it’s—

Angela: I bet you do.

Corey: My entire brain winds up completely short-circuiting. It's like, “Really? We met and we had lunch, and we wound up signing a business deal, and we did a bunch of sponsorship work with you? Couldn't pull you out of a police lineup, but thanks for saying.” It is such a high impact in such a short period of time, number of people. Do not recommend.

Angela: Definitely. It was my first one. I hope it's not my last. I mean, COVID is kind of changing the game, but it was a bit of a firehose. So, if you've never been to re:Invent, I think everyone should go at least once, and then hate it, and never go again. So, hopefully, I'll see you there in five or six years, maybe? I don't know.

Corey: I will be there whenever they start running it again and I'm not risking my life in it. But until then, where can people find you if they want to learn more?

Angela: Well, like I said, I'm on Twitter. And my Twitter handle is @scooterphoenix. And I tweet about all sorts of things. So, if you like Beyonce, you should definitely follow me. If you want to hear about Black Lives Matter and Brianna Taylor, people not being arrested, you should definitely follow me. And if you just like cheesy jokes and memes as a response to anything you tweet me, you should definitely follow me. So, yeah, I’d love to hear from you.

Corey: Can endorse. Your Twitter feed is constantly surfacing new gems.

Angela: [laughs]. Thank you.

Corey: Thank you once again for taking the time to speak with me. It really is appreciated.

Angela: Thanks. It was a pleasure.

Corey: Angela Andrews, solutions architect at Red Hat. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts, whereas if you've hated this episode, please leave a five-star review on Apple Podcasts along with a comment telling me exactly what I got wrong about tenured professors.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Spencer Viernes

Spencer Viernes has been practicing law for over 15 years. His practice consists of local, national and international clients. He has advised clients regarding digital infrastructure, including cloud infrastructure, data centers and telecommunications infrastructure, technology transactions, energy and natural resources, and economic development initiatives. Spencer has provided strategic guidance, both legal and practical, related to diverse geographic business and legal issues.

With a practice that includes deals across almost most major continents, Spencer has a demonstrated track record of closing deals with a savvy business mind and a sharp legal perspective. He works side by side with business and project teams to drive results that benefit his clients in a meaningful way and allow room for mutual benefit that solidifies lasting successful business partnerships between otherwise opposing parties.

Links Referenced:

  • Vierness ESQ Website: http://www.viernesesq.com/
  • Spencer Viernes Email: sviernes@viernesesq.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Catchpoint. Look, 80 percent of performance and availability issues don’t occur within your application code in your data center itself. It occurs well outside those boundaries, so it’s difficult to understand what’s actually happening. What Catchpoint does is makes it easier for enterprises to detect, identify, and of course, validate how reachable their application is, and of course, how happy their users are. It helps you get visibility into reachability, availability, performance, reliability, and of course, absorbency, because we’ll throw that one in, too. And it’s used by a bunch of interesting companies you may have heard of, like, you know, Google, Verizon, Oracle—but don’t hold that against them—and many more. To learn more, visit www.catchpoint.com, and tell them Corey sent you; wait for the wince.

Corey: This episode is sponsored in part by strongDM. Transitioning your team to work from home, like basically everyone on the planet is? Managing gazillions of SSH keys, database passwords, and Kubernetes certificates? Consider strongDM. Manage and audit access to servers, databases—like Route 53—and Kubernetes clusters no matter where your employees happen to be. You can use strongDM to extend your identity provider, and also Cognito, to manage infrastructure access. Automate onboarding, offboarding, waterboarding, and moving people within roles. Grant temporary access that automatically expires to whatever team is unlucky enough to be on call this week. Admins get full audit ability into whatever anyone does: what they connect to, what queries they run, what commands they type. Full visibility into everything; that includes video replays. For databases like Route 53, it’s a single unified query log across all of your database management systems. It’s used by companies like Hearst, Peloton, Betterment, Greenhouse, and SoFi to manage their access. It’s more control and less hassle. StrongDM: Manage and audit remote access to infrastructure. To get a free 14-day trial, visit strongDM.com/sitc. Tell them I sent you, and thank them for tolerating my calling Route 53 a database.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Spencer Viernes, who is an attorney; not typically the sort of profession that we see a lot of on this show, but one of the open secrets of working with large cloud providers is that a lot of companies have custom pricing deals and custom commitments with all of them. But because those contracts are themselves bound by NDA, people don't generally talk about them in public. Well, this is largely what Spencer does for a living. Spencer, welcome to the show.

Spencer: Hey, thanks for having me on. Corey, I really appreciate it.

Corey: The pleasure is mine. There's always been this sense of, if we talk at all about the fact that we have a custom deal with our cloud provider in place at all, they're going to come crashing through our windows and sue us to death. That is not matched, in my experience, of the reality of it, but there's definitely a chilling effect. Have you found that companies are generally reticent, in your experience, to talk about these things, or is this more of an engineering-side affliction?

Spencer: You know, I think it cuts both ways. Certainly, if you look across the lawyer spectrum—and as you said, it's a bit of an open secret. Lawyers are happy to discuss that they can get a deal. Certainly, the other side of that is, they're never going to tell you who their client is, and there are professional obligations that would limit them from doing so. So, it becomes practical, professional knowledge that you can get more customized deals, and you'll find folks who will make it a benchmark of their practice to identify that they are able to do those things, but the specifics are never things that they're going to be willing to share in public.

Corey: Exactly. Do you notice that there are major differences between the different cloud providers as far as when it makes sense to begin talking about custom deals when it makes sense to reach out, or just not even bother; get turned down flat, when companies get approached by their providers for these sorts of things?

Spencer: I wouldn't say that I recognize a ton of material differences in the way that cloud providers might approach a customer, or customers generally approach those individual cloud providers. I think one of the interesting things that—and I don't think this is a secret either, but one of the interesting things that I've recognized across my practice as a lawyer, working with mid-market or startup enterprises and even large multinational enterprises is, the most telling indicator of when you're going to be ready to approach or when somebody is going to be approaching you, whether it's a cloud provider or an enterprise approaching their service provider is going to be the economics. If you've got the specific bargaining leverage to go and approach your cloud provider and there's enough of a potential loss to them for having lost a customer as a result of not being able to engage in that level of flexibility, it certainly makes sense for them to come to the table. On the flip side if you're a startup and you're trying to get after a cloud provider and say, hey I'd really like a deal on this, it's going to be a lot tougher.

The better or worse reality of business and contract negotiation is that as much as we lawyers like to think that we can rule the deal, and we're these expert negotiators, it's the numbers that really run businesses. And at the end of the day, those are the things that help guide whether a contract is going to be reasonable. And certainly when you have these large organizations that are cloud providers, very seldom are they going to engage in deals where they just do something that's economically unreasonable.

Corey: The business has to make sense, regardless of what provider you're working with, you're—

Spencer: Yeah.

Corey: —never going to see a scenario where one provider—for example, data transfer—is going to offer a phenomenally better discount than another provider for the same type of traffic. They all generally round to the same baseline level, which in some cases is sort of strange because their retail pricing in many cases is nowhere near the same number, even though after negotiating for a significant scale, that does become effectively the same thing. It almost feels like it's a marketing problem in some respects, more so than it is a, “Well, don't bother talking to them. They're not going to negotiate for anything.”

Spencer: Right. I think that's true. The way that companies—and I actually don't think cloud providers are that different from many other businesses in the way that they look at their economics. Certainly, there are complexities and the number of layers from varying service level agreements for differing services or the economics that go into how those individual services are rendered to a customer, and then the consequent or resulting potential flexibility and negotiating leverage that they might employ as they enter into some engagement or come to the table to renegotiate something, those things can become more layered but, at the end of the day, it's, “How do we make sure we maintain a specific margin?”

And there are teams upon teams that try to identify what this is going to be, and whoever comes to the table has a spectrum within which they can negotiate and that's what they do. And that spectrum doesn't even open itself until you become level X so that they feel like you're a big enough player that it makes sense for them to want to hold on to you because then your customer lifetime value is exponentially greater than the initial startup that's just getting free services, right?

Corey: Absolutely. The strange way that I've seen this manifest in some of the negotiations I've been a part of is it's almost always positioned as, “Oh, well, the provider is investing in you.” And I love that term. “We're making an investment by deigning to give a discount.” Which feels like salesmanship and not particularly good salesmanship at that.

But there is an element of, yeah, in order to get larger discounts, there has to be more of a commitment made, either in terms of dollar spent or spiritually if nothing else. None of these providers want to find themselves in an undifferentiated race to the bottom, but the realities of the market do in turn dictate that this is something that is going to have to be competitive, or there's no point in even having a conversation.

Spencer: Yeah, I think that's right. So, I'll provide what I think is—I don't know if I'd say it's exactly an analog because Cloud has moved forward into a place that colocation and wholesale data center leasing used to occupy. In many of those spaces, the wholesale folks, the colocation folks, that really has started to become a consolidated industry and a commoditized race to the bottom. And so many of those players, especially the bigger ones, are really looking at how they can develop regions either for large enterprises that have just such a large footprint that they don't really use Cloud in the same way that the general use case would apply, or they're trying to prep so that they can get one of these big cloud players to come in and say, “Oh, great, you built a building, you build all the infrastructure. We'll occupy it.”

I think that same thing may become true for Cloud eventually. I don't know that it has hit that point exactly across every user type, given the various services and software, kind of, add-ons that cloud providers can provide to potential users, but I think you're absolutely right on how, potentially, goes because as you start to see services become more and more similar and offer similar use cases and similar performance metrics, you're right, it really does become kind of a race to the bottom, and then you're starting to look at—to your earlier point—is it the marketing that gets people? Because it can be really tough to get off of one cloud service and move on to another.

Corey: Oh, yeah, the idea of even if you approach this from a perspective of, oh, we're only going to use the common services between all of these providers so that we could migrate easily, it never works out that way. “Cool. How long is it going to take you to migrate those 18 petabytes of data from one provider to another? How long is it going to take you to repoint everything, and the basic assumptions of what provider you’re in that have been baked into your app?” In order to do that, there has to be a compelling story, and that's usually not going to make a pure economic argument or even close to it. So, there has to be a strategic win for it.

Spencer: Yeah, I think that's absolutely right. There's a lot that goes into cloud services, and once you start engaging in more than one service, as we talked about before, the complexity just increases over time. And unraveling that can be really difficult.

Corey: One thing that's on top of everyone's mind these days has been, when you sign one of these contracts, it is almost always tied to a certain level of committed spend over a period of time. Sometimes you pay upfront, in return for a small discount. Other times, it's just over the course of various time spans, and there are different ways to structure that. But they all distill down to you will pay us X dollars over Y time period.

When people were making those projections during times that were not these uncertain times—I miss the more certain times we used to live in without realizing it—everything was up and to the right. Of course, we're going to build in a bunch of growth. Of course, we're going to be able to achieve those points. Now, it turns out that in pandemic times, when suddenly 80% of your user traffic has evaporated, you're not going to hit that commit anymore and people are basically panicking in order to have conversations with cloud providers about, “So, what are you going to do for us?” And the response, that I've seen at least, has been scattershot at best so far.

Spencer: Right. So, I'll give an interesting example, especially during the pandemic time. I spent a lot of time as a young lawyer in the great financial crisis between 2008 and 2012. And there were properties being foreclosed on, there were tenants that couldn't make their rent payments—and these were commercial tenants—and so they were going back to their landlords and saying, “Hey, we're not going to be able to do this. We need some forbearance.”

You're right, I have not seen a significant amount of flexibility that would be similar to this concept of forbearance in a real estate industry, across the cloud industry. I think even before you get to that, many of the IT professionals that I've done engagements with, a lot of them think, “My IT spend is here, and my expectation with Cloud is that there’s—” I don't know, “Less than 5% of wasted expense.” I actually think that's underestimated to a significant degree, generally, and then you add the pandemic in, and the delta there becomes really extreme.

Corey: That's part of the challenge, too, is that a lot of these providers have approached building things out from an original perspective of, “Okay, this is the Cloud. It's going to be on-demand. Use exactly what you're going to need and no more, and then turn it off when you're done—” Spoiler; no one ever turns anything off. But when dealing with custom enterprise-scale negotiations, that is a much more bespoke, hands-on process. And some companies—[cough] Google—tend to have challenges in reaching out at the human level and negotiating these deals, whereas others have gotten much better at speaking to enterprises, or in the case of Microsoft have 40 years experience of doing exactly that.

But now that suddenly the same thing as hitting globally—everyone all at the same time—it feels like whatever processes and systems all of the large providers have put in place for handling these enterprise deals are getting oversaturated enormously because it's not part of the normal negotiation cadence. It's everyone all at once, all asking the same thing. And they can't deal with the sheer volume. Is that something that you're seeing signs of yourself, or is your practice a little bit more removed from that?

Spencer: I would say on the one hand, I think my practice is maybe a little bit more removed from that sheer volume. So, I have a small boutique practice that specializes in the digital infrastructure side. Everything from the construction negotiations for developing an actual white box data center that might have cloud deployments in it, to helping people understand what their cloud agreements are actually saying and what that's going to mean to them as they move forward, and trying to translate that into what actual business people refer to as ‘real numbers.’ And so I probably haven't seen the same volume that some folks have, but I would say from my experience with my own clients, I do know that there is a significantly increased volume of folks, even at the bespoke level saying, “Hey, we need to go back and talk about this.” And generally the in-house legal teams for these types of businesses, they're not really staffed to handle people coming back and renegotiating really frequently. They're really staffed to handle longer-term agreements and have templatized agreements on the sales side, which is why you see so many clickthroughs.

And so this is proving pretty difficult for them, and my understanding is that many of them are having to look at their outside counsel—which gets really expensive really quick—and try to see if there are other efficiencies, like using outsourced talent or some type of automated process to look at how you would check this box, or check that box—give in on this or, hold firm on that. And I mean, that's a tough process to implement really quickly, especially within a matter of months. My suspicion is that most of them are still struggling just to figure out what their policies are going to be, even now, after folks are already starting to bombard them with things.

Corey: It doesn't, also, help that the way that you deal with some vendors does not work with others. For example, I've heard multiple stories now of purchasing departments reaching out to their top five largest IT vendors or working down the list of we—during these uncertain times. Once again, pining for certain times—we'd like to get a 20% discount off of what we're paying you, and none of the cloud providers are budging on anything remotely approaching that. Instead, it's, “Oh, just turn things off and then talk to us.” And people, there's always a giant pile of waste in various accounts. During better times, it was always amusing to me watching people negotiate for months over 2 or 3% difference, whereas they can save 15 to 20% just by turning a few things off that are no longer in use. I digress. It always seems like there's this weird misunderstanding and misalignment between engineering and the rest of the business in these conversations.

Spencer: Yeah large enterprises also struggle with this. They're certainly not immune to the disconnect between engineering, and ops folks, and other strategy or, say, sales folks. And I use my own experience having been in-house counsel for years with one of the large tech enterprises in the Bay Area. One of the things that we had guys complain about all the time were, you'd have folks who were talking about hardware, and nobody was talking about connectivity.

You'd have folks speaking about operations and execution, and then you'd have other folks on the coding or developer side saying, “Well, you can't do what you're doing.” But they were talking at the same time, and I don't think that's something that is specific to, or unique to any one enterprise. I think many enterprise clients that I've dealt with deal with those same challenges, and it certainly rears its ugly head as you start to look at implementing some type of renegotiation, or trying to get to addressing some of the issues that this unique, kind of, time in our history is certainly create.

Corey: One of the strangest things is that no one seems to know what this is going to do longer term for their business, what the, even, mid-to short-term is, is this something that they should start planning for recovery in three months, six months, one year, two years, ten years? And one of the big challenges is that in some cases, it seems that companies don't even know the questions to ask of their providers. Is that something that you have noticed increasing, or is it always, frankly, been a little bit of a, the two companies tend to speak past each other and not often going in the right direction?

Spencer: So, with respect to some enterprises, certainly, so some of your large enterprise clients that are really used to having bespoke agreements, right, because they had bespoke agreements with the colocation guys, and they would come back to the table, and they'd renegotiate every, I would say, at most every 36 months because they’d probably end up seeing a new deployment as they were growing in capacity if their business was moving along a very typical trajectory. And so those companies are kind of used to this and used to going back to their providers and saying, “Hey, we need this. Hey, we need that.” And they would do that with lots of different vendors. And it's not to say that they were always fantastic at it, as we talked about before, but they have some experience in doing this.

What I found even before the pandemic timeframe, and certainly now, is that there are so many—and I would say mid-market companies that it's hit or miss as to whether they really fall economically in that zone where they were getting bespoke agreements or they were getting a lot of flexibility from their providers. The larger category of folks that are really struggling with this, were the folks that they heard the phrase ‘digital transformation’ but they didn't have the in-house expertise with which to create a strategy and a plan for execution, and they also didn't really have the in-house expertise to know, “Who do I talk to? Do I talk to Corey Quinn? Do I talk to KPMG? Do I talk to Deloitte? Can I even afford Deloitte and the expense of their reports?” And not to say that they’re poor reports, but certainly, they have a lot of resources that they can throw at things. And so I think there's been a lot of confusion around this that isn't necessarily specific to the pandemic but is really indicative of a more significant desire for executives to move into a more digital existence, but a lack of knowledge around how specifically they should do that, and whether they're actually choosing the right way to do that once they've made a decision.

Corey: That's increasingly something that I'm finding is less understood—at least in my experience—with the smaller, more nimble, shall we say, startup companies that didn't exist five or ten years ago, versus the enterprise blue-chip style companies that have been around for a century or so. The ladder has certain expectations of how an enterprise sales process is going to go and that, on some level, has been something that providers have all had to scramble to catch up to. Whereas startups are basically making this up as we all collectively go along to some extent, where it's, “Oh, this is the contract you're going to give us? It's fine. You want me to sign now?” “Uh, maybe you might want to read it first, just as a thought.” There's a maturity difference on some level.

Spencer: Yeah, I think that's right. So, I would say a classic example of your typical enterprise customer that's been around for a while—think insurance companies, hospitals, those guys, they have internal departments that have been doing this—you know, banks—for a long time. And as they move into something, they tend to have what we would consider a traditional enterprise cycle, which means not fast, usually. And then to your other point, you've got the startups and they've got a lot of really smart folks that have learned to develop amazing software or build a really interesting product, and now they need cloud services. And these guys haven't spent their time, or I doubt haven’t spent their time as engineers back in the old AT&T days, right?

They're not the folks that came out and were grunts having to think about what a mechanical system does, or a water loop is in a big data center, or even how to physically configure servers, or figure out how to route switches from your garage to wherever you wanted to connect to. And so they're used to having things on the fly, and they are looking at that, kind of, upper right quadrant all the time and thinking, “Yeah, that's how we're going to grow and so we just need to make sure we've got everything on-demand to get there because what's most important is our growth rather than this type of management and trying to, kind of, be Peter Prudent with his pocket protector, and pencil, and thick-rimmed glasses looking over contract provisions.” What you realize when you go from that small startup to raising your round at $150 million—and say you're a FinTech company—is that you might be overspending by millions of dollars every year and you don't really know that. Or really know if you know how to decipher that, or what you can turn off to make sure that you can save that.

Corey: In what you might be forgiven for mistaking for a blast from the past, today I want to talk about New Relic. They seem to be a relatively legacy monitoring company, and I would have agreed with that assessment up until relatively recently. But they did something a little out there: they reworked everything. They went open source, they made it so you can monitor your whole stack in one place and, most notably from my perspective, they simplified their pricing into something that is much more affordable for almost everyone. There's even a free tier with one user and 100 gigs per month, totally free. Check it out at newrelic.com.

Corey: One thing, that’s always struck me as odd in the latest consumer space has been when people will call and complain, or try and negotiate, and they enter into it with the same what feels like very amateur-hour mistakes, specifically, that they go into the conversation without having any clear idea in their own mind of what it is that they're asking for. Now, that does manifest differently in the context of a conversation like this. Usually, it's, “Oh, I want to get a bigger discount on that one service.” “Yeah, but you don't use that service, so even if you were to get it as a talking point, does it move the needle on anything?” No, it lets you feel like you've got one over on the cloud provider, but that's not really the way to success.

Spencer: Yeah, I think that's right. I don't want to say—it's certainly not rocket science. And I'm definitely not trying to put myself out there as being someone who is ahead of the curve as far as intellect because I'm certainly not. What I do think and what I have learned over my experience—and this doesn't just apply to cloud agreement negotiations or other Software as a Service negotiations, but negotiations generally—there has to be—and you talked about the disconnect, sometimes, between engineering and other groups, there has to be a connection between the real business objectives and why you're trying to commence—or why you're trying to get into that negotiation, and then how you approach that negotiation. If there's a disconnect between those two things, and we talked about this before, numbers run the business, and so understanding what your objectives are from a numbers basis and having had the discussion beforehand, or before you start to focus on negotiation around why you want to have that negotiation?

What your objectives are? What things can you do internally that are levers that will help you either save money, or be more efficient, or drive more revenue, it is really best to have an understanding of things. I'm not saying it's always possible, but it's best to have that understanding because it means that you go into a negotiation better prepared to both ask for things and be able to have your own flexibility on your side to say, “Okay, look, I can live with that because I know I can turn this lever over here if I can do that thing over there.” And it's hard to do that if you haven't done some time, kind of, researching how do you make the most efficient use of whatever it is that you're already paying for, or identify those things that you should be willing to turn off.

Corey: You'd sure like to hope so anyway, otherwise it just seems like it's going round and round in circles to some extent. One thing that I've always found that is just bizarre to me has been that this, I guess, almost ridiculous idea that you're going to yell at a vendor when you're trying to negotiate a deal with them and at the end of it things are going to be fine. Well, it's not like buying a car. You don't get to walk away at the end of having squeezed every penny of margin out of the deal, and it doesn't matter—in the car context—because you're not going to have an ongoing relationship with the dealership in many cases. Whereas when a cloud provider runs your entire infrastructure, they're your partner, whether you like that fact or not, and effectively salting the earth during your ham-fisted attempts at negotiation, usually has consequences or how much we try to hope otherwise that just doesn't make it worth pursuing.

Spencer: Yeah, I think that's absolutely right. I'll tell you two things that I feel like, over the course of my career have been really significant when I think about negotiations. The first one that I think is really important is, it doesn't matter if you feel like you were able to get more off the table than the person on the other side. It doesn't matter what agreement your lawyer puts in place on paper. If you negotiate a deal that's really bad for one party, the greater likelihood is that somebody is going to break it.

It's just a reality because economically—or it's not economic, but it's some other thing that just so hamstrings the other party and makes it completely unreasonable for them, they're going to break it because the economic consequences are probably less than the consequences of complying with that agreement. On the other side, I know a guy who, we were sitting in a negotiation and he leaned over to me—and he was a senior lawyer and he said, “I’m going to stand up, I'm going to throw my pen against the wall, and I’m going to walk out of the room. I'm going to come back in five minutes, and I want you to tell me what they said.” And I just don't even understand the logic behind that because, in my mind, it's looking for some type of win that doesn't really have any value. So, I've always kind of negotiated from the place of, “Look, I can tell you where we stand here, I can tell you what's actually going to work for us. And if we can't get to a place of it being reasonable, we may end up just having to break the contract anyway.”

I think that is more of an economic threat to the provider to know that at this place, it becomes a deal killer without having to get up and make any shows of yelling or screaming, or irrationality. And I also think, to your point, if you're going to be in these relationships longer term, which you necessarily are because of how these types of services actually work within your enterprise, it doesn't make sense to create animosity across the table. I just don't see the benefit in doing that.

Corey: So, here's the question that is the other side of what I always say, which is, “I am not an attorney; consult your own when you're having the legal discussions around terms, conditions, et cetera.” But from your perspective, when is it time for someone to engage with an attorney during a cloud provider or other vendor negotiation?

Spencer: So, I think there are probably multiple answers to that. That's probably the attorney answer that everyone was expecting. But I think it's time to consult an attorney when you really know that you're going to be spending consistently on something. If you've got this free service and you have no idea whether your company is going to make it, you're just struggling to find time to work on your project. Usually, when you get your first real round of capital, it’s a good idea to take a review and say, “Hey, here's what we're doing.”

And it doesn't have to be your guy that just does your securities, and you probably won't even have a general counsel by that time. But it does make sense just to sit down with somebody and get them to help you understand what you're actually signing on for so that you can incorporate that into your longer-term strategy. For your enterprise clients that have agreements in place, if they're really starting to struggle. So, a great example is [unintelligible], a midmarket hospitality entity that is seeing a significant reduction in revenues right now, it might be a good time to look at what you've been spending on Cloud and understand whether you can go back and make that request. And it's not to say that it's always going to be in the agreement, but it might be time to get counsel and go and talk to your account manager and say, “Look, we need to have a talk about this.”

Because otherwise you're just going to see less money on the provider side, and there are bound to be a wave of insolvencies, and that's also no good for the providers. The cloud providers don't want to just see a huge vacuum of revenue going out the door, so if they can do things, and partner as well, it may be in their interest. Now, there are varying degrees of flexibility among cloud providers, given the way their enterprises operate outside of even the pandemic, but it certainly doesn't hurt to ask and I think it's worth it to do so if you can find someone who understands the Cloud and digital infrastructure space, and can actually speak to those provisions or to those agreements intelligently.

Corey: What winds up creating an opportunity for a boutique law practice such as yours to step into this space, as opposed to a company going through their typical law firm they use for more traditional general counsel style activities or in-house general counsel for that matter?

Spencer: Yeah. So, I think the way that I'm going to interpret your question is when is it ideal—or when does it make more sense for an enterprise that may have in-house counsel, and even if they don't have a large group of in-house counsel, may have their own general counsel to use someone like me who really does specialize on the digital infrastructure space rather than their typical group? And I would say what is increasingly necessary for businesses in today's economy—certainly the pandemic is starting to really raise that flag for many enterprises—is a transformation or a strategy around how to move into the digital realm. Not just, “Hey, we've got an app or we use some software,” but, “We're efficient, we understand what our IT operations look like and how they coordinate with our revenue-generating services, but we may not have the folks in-house that are experts on cloud agreements or specific digital infrastructure nuances.” SaaS is one thing, real estate is another thing, Cloud and digital infrastructure and data centers are this amalgamation of IT, SaaS, real estate, and managed services provider relationships. And so it can be helpful to have folks in that space. I often would say it's the digital transformation that if you find yourself or feel like you're kind of lagging behind in that, that's where you start to seek out a person or a boutique firm like mine, primarily because we've just been in this space, and there's no need to reinvent the wheel on that.

Corey: Yeah, it's your first time seeing one of these contracts, whereas it's your third time this week.

Spencer: Right. I mean, you go to Google Cloud, or you go to AWS, or you go to Azure, and, you know, you've got these clickthrough agreements, which they seem—I don't know if they do, or not, because I'm a lawyer, and I probably read things differently than everybody else in the world, but they don't necessarily seem like they are particularly onerous.

Corey: No, in fact, the enterprise agreements are almost always nearly identical, just with more generous terms around things like shutting off for non-payment or whatnot. The termination provisions are slightly different. But yeah, there's remarkably little that shifts meaningfully. Sometimes it's a whole bunch of nonsense around indemnity, but that's basically people bring you in for.

Spencer: What you may not be seeing is the 50 different links to SLAs or CCPA terms—which are the new California Privacy Act—or the global data protection regulation, the GDPR terms, or the different terms that apply to this particular service versus that particular service. And so some of those things are nuances that you just wouldn't catch. Most agreements are going to have, even if they’re Software as a Service agreements, they're going to have one service level agreement. They're going to, in most cases, have one product that they're offering you. The very interesting thing about cloud services suites is it really is a suite: it's an entire platform upon which you can build an architecture.

And so there are so many different products that have different terms, it can be really tough to navigate that. And it really is something where you have to have the knowledge to bring in—because let's say you're an early stage, or let's say you're even past early stage and you've gotten to your B round, your CTO may still be a developer. And that, I don't know, is always going to be the expansive knowledge set to really understand how you deal with the economics on the vendor side. Because that's, I think, really where you get to the ‘yes’ or the best negotiation is being able to understand where the other side is coming from and how you can reach a reasonable agreement with them.

Corey: Thank you so much for taking the time to speak with me, especially around something that is not often spoken about. If people want to learn more, where can they find you?

Spencer: Yeah, so I maintain my law firm’s website www.viernesesq.com, and happy to have people reach out to me. They can also just shoot me an email at sviernes@viernesesq.com.

Thanks for having me on. Corey, this has been great. And it's interesting because it's not often that folks on the business side really want to hear much of what the lawyers say, and so this has been a great opportunity.

Corey: Well, if they don't listen early, they're going to definitely have to hear about it later. So, [laughs]—

Spencer: [laughs]. Right, right.

Corey: It's one of those, if you have to have these conversations, maybe do it when you still have the option to do something productive about it. It just seems to make the most sense.

Spencer: Yeah, that's absolutely right. And hopefully, the folks that do a really good job on the legal side, understand that as well. They want to provide you with an ounce and prevention for that pound of cure, as opposed to trying to front-load all the stuff because, again, for the best lawyers, it's all about having partners as well, as opposed to just calling them clients.

Corey: Absolutely. And that's the key takeaway here, above all else. So, thank you once again for taking the time. I appreciate it.

Spencer: Thanks, Corey. I appreciate it.

Corey: Spencer Viernes, managing member of viernesesq.com. I am Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts, whereas if you've hated this podcast, please leave a five-star review on Apple Podcasts anyway, after clicking through the shrinkwrap agreement.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Ev Kontsevoy

Ev Kontsevoy is the CEO of Gravitational, where he and other engineers build open-source tools for other developers for securely delivering cloud apps to restricted and regulated environments. Besides computers, Ev’s obsessed with trains and old film cameras.

Links Referenced:

  • Gravitational website: https://gravitational.com/
  • Gravitational GitHub: https://github.com/gravitational
  • Teleport GitHub: https://github.com/gravitational/teleport

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Catchpoint. Look, 80 percent of performance and availability issues don’t occur within your application code in your data center itself. It occurs well outside those boundaries, so it’s difficult to understand what’s actually happening. What Catchpoint does is makes it easier for enterprises to detect, identify, and of course, validate how reachable their application is, and of course, how happy their users are. It helps you get visibility into reachability, availability, performance, reliability, and of course, absorbency, because we’ll throw that one in, too. And it’s used by a bunch of interesting companies you may have heard of, like, you know, Google, Verizon, Oracle—but don’t hold that against them—and many more. To learn more, visit www.catchpoint.com, and tell them Corey sent you; wait for the wince.

Corey: This episode is sponsored in part by strongDM. Transitioning your team to work from home, like basically everyone on the planet is? Managing gazillions of SSH keys, database passwords, and Kubernetes certificates? Consider strongDM. Manage and audit access to servers, databases—like Route 53—and Kubernetes clusters no matter where your employees happen to be. You can use strongDM to extend your identity provider, and also Cognito, to manage infrastructure access. Automate onboarding, offboarding, waterboarding, and moving people within roles. Grant temporary access that automatically expires to whatever team is unlucky enough to be on call this week. Admins get full audit ability into whatever anyone does: what they connect to, what queries they run, what commands they type. Full visibility into everything; that includes video replays. For databases like Route 53, it’s a single unified query log across all of your database management systems. It’s used by companies like Hearst, Peloton, Betterment, Greenhouse, and SoFi to manage their access. It’s more control and less hassle. StrongDM: Manage and audit remote access to infrastructure. To get a free 14-day trial, visit strongDM.com/sitc. Tell them I sent you, and thank them for tolerating my calling Route 53 a database.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. This week's promoted guest is Ev Kontsevoy. Ev, you are the CEO of a company called Gravitational. First, welcome to the show, and secondly, what's Gravitational do?

Ev: First, thank you for having me, and second, Gravitational enables developers to have secure access to all of their production environments, and it also allows you to run applications on any environment anywhere in the world.

Corey: It's definitely the right direction to go in, compared to, I don't know, giving developers insecure access to other people's production environments, which seems to be incredibly in vogue in some circles these days. But—

Ev: Absolutely.

Corey: —it's important to solve problems like this, where I want to access different environments that, if there's a just and loving God, are actually separated from one another. It's the old saw of everyone has a test environment; some people are lucky enough to have it be separate from their production environment. It becomes a challenge that everyone struggles with on some level. And everyone who talks about having solved it is basically lying, in my experience. “Oh, yeah, we have a completely distinct test environment that's a high fidelity copy of production.” And then you scratch a little bit, and it turns out the, “Well, you know, except for the continuous deployment server because that thing talks to everything, and please don't ask us about our security posture there.” So, there's a lot of directions you can go in when we're talking about securing access into environments. Where do you folks start, and where do you stop?

Ev: Well, it's interesting thing that you used this word environment so many times. We believe, fundamentally, that having to maintain computing environments is just fundamentally a major limitations for us—software developers today. Like, back in the day, I could just build my software, put it in a box, give it to you, and that would be the end of transaction, right? You would put it on your laptop, or put it on your server in the basement, and you will just run it yourself, or software will just run by itself. But as we transition to this cloud world, so now in addition to building your software, you also need to build and maintain the environment.

You know, we celebrate everything is—infrastructure as code movement, but what it leads to is just enormous complexity. So, essentially, we're not just building our applications, we're also building computers to run our applications on because that's what these environments are all about. So, ultimately, we want just to enable you—engineer—to not even think about the environment. Don't you think that engineers should just do a git commit, git push and go home, and have software magically run anywhere in the world where users need it? So, that's really where we're going. But we're starting with access. Access today is the problem that we solve really well. And if you accessing your servers using something else, you’re probably either not doing it well, or you’re probably very inconvenienced.

Corey: I talk from time to time about the ridiculous serverless system that I built that is my newsletter publication system, and aspects of that do have a, everything that hits git automatically gets deployed. Now, it sounds awesome—and it is from a developer perspective—but let's be real for a second here: there's no test built into this. I'm a terrible developer. Security? Yeah. There's one user account in this thing: it's mine, it has full access to everything, and my default security posture is, sure. It's not so much a posture as it is an unfortunate slouch.

Ev: You know why is that?

Corey: Because, again, what you talk about is right, it's a colossal pain in the butt. When I don't have to worry about things like insider threats; this is a back-end system, at the end of the day, so if things break, there is no user that is impacted beyond me; and it gets everything else out of my way. There's no value to me in increasing the security posture because, in my perspective and perspective many other folks, security is a continuum between, effectively, being, it completely wide open and being completely unusable. Where on that spectrum do you wind up wanting to fall? I bias for getting things done quickly.

Ev: Absolutely. If you don't mind, I will be quoting you in the future. You basically said, there is no value in security. This is why I don't like when people call Gravitational a security company. We don't give you security, we give you instant access to whatever you need to be productive right now.

So, one of our open-source project is called Gravitational Teleport. Why we call it Teleport? Because it creates this illusion that you can teleport any computing device your company owns into the same room with you, so you can instantly access it using protocols like SSH, Kubernetes, and support for other protocols coming. It’s on GitHub; go check it out. But it completely erases this mental partitioning that we apply to all these environments, like test, production, Amazon, GCP, VMware, on-prem, basement, satellite, self-driving car.

All of those are going to be instantly accessible to you without compromising security. Security, it's almost like when you building a bridge over a river. Is security benefit? Of course, you don't want that bridge to collapse, right? So, it's secure in that sense, but the benefit is getting there, getting on the other side of the river. So, that what we believe we enable with Teleport, specifically.

Corey: Unplug the computer and it's way more secure, but your customers are probably not going to be happy.

Ev: Yes. You're losing access.

Corey: You're right when you, I guess, tweak my quote to say that there's no value in security. There's a lot of truth to that because I'm very reassured constantly by a wide variety of companies that security is incredibly important to them. They always tell me this in their announcements of data breaches holding my data, where it was very clear that security was not very important to them. It was a backburner priority, and now they're doing damage control.

It's, “Oh, what is your intrusion detection system?” “The front page of The New York Times. We check that thing every day.” It becomes an afterthought because you're not going to improve security and get to your next business milestone by doing that, in most cases. It is not directly aligned with the stated goal of most businesses to improve security posture. It's something has to be done; it's not a value add.

Ev: Look, it's just expensive, too, and there is always talent shortage to remember about. If you think of major tech companies here in the Valley, you know, Google, Netflix, Facebook, go and ask them, “How do you SSH into anything?” And you will see that most of these companies, they built quite sophisticated internal products to do that. They do not use off the shelf completely unmodified components, like OpenSSH, for example. OpenSSH is just the building block.

Who else can do that? That's really my question. I went to a couple of dinners with CTOs of other companies here in the Valley, and they all confessed to me that they're all struggling with it, that it's basically—it's wild, wild west out there, how access to infrastructure is implemented. And you could even do it yourself. Go and ask your engineer friends, “Hey, do you still have access to production environment of your former employer?” You know, how many of them will say, “Yes, I do.”

Isn't that scary? Like, there are companies out there running applications holding data of their customers, and their former employees still have access to it. Yeah, because of this trade-off of security and convenience. If you don't make it convenient for engineers, they are not going to be productive, or they are going to build backdoors. So, I've seen that happening.

Corey: One of the more, I guess, amusing anecdotes from earlier in my career was when I was unceremoniously fired from a company—that kind of happened a lot based upon, you know, my personality, and everything I say, combined with everything that I do—and for the next day or so, there were a repeated series of outages on the company's service. My perspective is and always has been, once I no longer work here, I don't care about you anymore; I'm certainly not vengeful, or wrathful. But what I strongly suspect happened—because remember, I ran the ops teams at these places—where suddenly they're having to rotate credentials across the board of the shared service role stuff, which is absolutely the right move. “Eh, we wound up letting someone go. Maybe they're bitter, maybe they're vengeful but, oh crap, that person had access to all of these secrets that are difficult to change. We'd better get started right away.” And frankly, as a former employee, I want them to do that. Two years later, if there's a data breach there, I don't want to be even a remote consideration—

Ev: Exactly.

Corey: —as the cause of having done that. Because it's, “I don't work here anymore, please lock me out.”

Ev: Absolutely. I would even say the rotating credential is an anti-pattern. Like you're not even supposed to have long-lasting credential, so that removes the need the rotation. Technically, this benefit is called reducing operational overhead of implementing proper security. Well, it starts with using the proper tools.

I have a somewhat controversial statement, for example, to say that, “Hey, if you're using SSH keys today to access your infrastructure, even if you're storing those keys in a secure vault, you're doing it wrong. You're not supposed to be using SSH keys to access servers. You're supposed to be using certificates, and certificates need to be issued with automatic expiration for you every day.” So, it's just as convenient to use, but it removes the need to rotate anything. So, if you just don't show up for work the next day, your access will be automatically revoked.

So, that's the way you do SSH security, for example. And this list goes on and on and on, if you do it right, or if you are using a solution that does this right, by default, without you having to configure thousands of config files all over the world, then the operational overhead and pain kind of go away. And that is something that we, with our open-source project, to trying to promote and enable.

Corey: The challenge I run into is, I agree with everything you just said from a high-level perspective, but then it turns into almost conference-ware on some level, where the idea of what you say versus the real-world consequences slash implementation of that winds up breaking down. For my exact use case, for example, I have an EC2 instance that I use as my development environment. When I'm traveling, I only take an iPad with me and I use an SSH client called Blink to get into this thing. It also speaks Mosh, which piggybacks on the SSH handshake, but that's neither here nor there. It does not support certificates.

And it's an iOS app, so getting it wedged in is going to be a little bit of a challenge. So, I could see in an environment where I'm doing that, yeah, we don't use SSH key pairs, except for that thing. Same story with GitHub repositories. To my understanding, they also don't support SSH certificates and whatnot. So, this turns into the edge case exception territory pretty easily, where, yeah, we generally believe as a best practice that you should not be SSH key pairs, but here's the long list of things that require it, so we do it anyway.

Ev: Absolutely. You could also talk about network equipment. How many routers you've purchased even for your own house or apartment that have a baked SSH server that only supports keys?

Corey: Someone has fancy network equipment. Mine is still stuck on Telnet.

Ev: [laughs]. Oh, yeah, that's, that's even worse. But, seriously, this is why we are doing it in the open because at the end of the day, if you have this vision for how future is going to look like, and that's a better future than the present, so you need to find as many like-minded people as possible. And the best way to do it is just to put everything you're building in the open out there, and make it easily accessible. So, if someone is working on the next generation of iPad SSH client, they could just go, and take, and use our code to make it support certificates.

Or if someone is struggling to set up certificate authentication for SSH with existing open-source tools, here's another open-source alternative, and it does certificates by default, so there is no complexity, kind of, associated with it. So, by doing it in the open, and interacting with the community, and sitting here chatting with you, that's how I think we will proceed moving forward. It's just like, slowly reshape the future, towards simplicity, first of all; ease of use, first of all; and then security compliance and everything else that come with it.

Corey: So, tell me a little bit about what the getting started process looks like. It's one of those ideas where, on some level—we see this on conference stages all the time where—the one that really stuck with me was I was reading an AWS blog post about how the whole point and value of Kubernetes—which anything they say after this that isn't, “It's good on your resume,” is kind of a lie. And that's a hill that I understand is unpopular but also completely correct—and they say part of the value here is that you never have to SSH into your environment ever again. And that was great. When I finished reading the blog post, I checked what else would come out that day, and oh, now they've launched this new integration with Systems Manager Session Manager that lets you get a shell inside your containers so you don't have to SSH into them anymore. It’s, oh, that's right, the things we say versus the things we do. What's the getting started process look like that helps make the ideal city on a hill version a little bit closer to reality.

Ev: So, before we're going to getting started, I think that the question of either you should or should not SSH into pods—or SSH into machines that Kubernetes is running on—it's really up to you. It's up to your organization; it's up to your operational philosophy. I don't think there is a single answer, like, or industry best practice that you just go out and say, “You should always do that.” And when companies come out with these messages, you're right, it just feels not genuine. There is a way to do it simply and securely.

So, and if you want to SSH into the same infrastructure that Kubernetes is running on, go ahead and do it. However, make sure that you use exact same credentials, make sure that they are consistent. So, for example, Kubernetes has role-based access control. SSH at its core does not have role-based access control, so you should use SSH implementation that enables you to set the same kind of roles and same permissions. For example, you could say, developers must never touch production data.

So, your SSH layer needs to understand what is the production and what is staging, right? And then when you accessing a machine, your SSH layer needs to be aware if there is any customer data on that machine, or that machine gives you access to customer data. Traditional open-source SSH tools, they're just too low-level to understand these modern cloud complexity. And this is why some companies just say, no, we're going to disable SSH completely. They should only use Kubernetes. And then they run into other issues when they do that.

So, going back to our project, how did you get started? Well, go to github/gravitational/teleport and look at the readme. It's very small readme, and it tells you that Teleport is a single binary, same as sshd. So, you put it on your servers, and then you give it a very little configuration. And it gives you all of these things: it gives you the proxy that you use to access all of your infrastructure, it gives you role-based access control—what's coming up in open-source version, by the way—it gives you integration with single sign-on so you could do like something like GitHub authentication to get into infrastructure or Okta or whatever. And it does it in the same way as you access Kubernetes.

So, if you are member of a developer's team on Kubernetes RBAC, you are going to be member of developer’s team on your SSH RBAC. And the same rules that you set for developers will be applied for both protocols. So, you could have your cake and eat it, too. But yeah, then you click download link to play with it, and then there is documentation on the quickstart. It's basically the same experience as we're all accustomed to when playing with well-packaged open-source solutions.

Corey: A lot of marketing on your site talks about using this to get into clusters of machines or, of course, Kubernetes, which is the third rail I am not touching at the moment because, oh dear stars, do I get letters whenever I do. And that's great and all, but what my primary use for most of what I do with EC2 these days is using it as a developer environment. I don't have a cluster because it turns out that some of the EC2 instances I use are really big and they keep making bigger ones, so the problem gets way easier.

Ev: [laughs]. It's kind of interesting, you mentioned this. All right, let me step down from, you know, CEO of Gravitational role here and, as an engineer, I do find it quite interesting that we are now getting these enormous machines with 64 cores and terabytes of RAM—thanks for AMD stepping up their game. So, it is kind of questionable, do you even need an entire environment with this kind of auto-scaling stuff if computing is now so cheap? And engineering your application to run on a single box is actually much simpler.

So, part of me looks at this whole thing, and I'm thinking, how many startups are out there, how many, just, web applications would do way better running on a single box? And you could have the other one for failover, but the point is, that I'm quite fascinated with the progress we're making on computing, again. But going back to your question on, what is a cluster? Do I even need a cluster? This is basically a question about the language.

It's the word ‘environment’ I like to use. Like, you have an environment you want to go to: it could be a single machine, it could be two, it could be two thousand. But then you have these things like regions, right? You could have a single node in one region on AWS, then you could have another node in another region. So, how do you call those?

And then you have, for example, systems like Kubernetes, and they do use word ‘cluster.’ So, we try to use language that is as agnostic as possible. Just think of cluster as just a collection of machines, and a single node is a cluster. And that's another problem that we're struggling with. How do you call a server these days?

If use the word server, some people will say, “Oh, no, no, no I don't need access servers. I need to access VMs.” [laughs]. Or they will say, “No. I don't need to access VMs I need to access computing instances.” So, what is it this thing you're accessing? Or is it an instance, or maybe it's a Kubernetes pod.

So, we try to use this language that's kind of neutral. So, we use ‘cluster’ to describe a collection of any computing devices you may have, And we use the word ‘node’ to describe anything you can SSH into. It could be a pod, it could be a VM, it could be an instance, it could be a server, it could be Raspberry Pi or a self-driving vehicle. So, it's all node from Teleport’s point of view.

Corey: In what you might be forgiven for mistaking for a blast from the past, today I want to talk about New Relic. They seem to be a relatively legacy monitoring company, and I would have agreed with that assessment up until relatively recently. But they did something a little out there: they reworked everything. They went open source, they made it so you can monitor your whole stack in one place and, most notably from my perspective, they simplified their pricing into something that is much more affordable for almost everyone. There's even a free tier with one user and 100 gigs per month, totally free. Check it out at newrelic.com.

Corey: To be clear, it has its own standalone installer, or am I in NPM hell, if I want to be using it?

Ev: It's a single binary. You don't need installers. Again, we as a company, we are addicted to comple—to simplicity.

Corey: Oh, you almost misspoke there as, “We're addicted to complexity.” And yeah, I actually have a list about 15 companies, I could absolutely put that as their tagline.

Ev: And you understand why this happens, especially in the open-source world. If you make your product so easy to use, and so dead simple, then people will just start saying, “We know what, why would I pay you money?” It's open-source, it's just this magic dust, I could sprinkle it all over my infrastructure and call it a day. So, then the open-source companies, they're basically forced to make the product more complex. And they say, “Well, these 57 features that exists only in enterprise version, and you really need some consulting help to set them up.” So, I could see why that is the case sometimes.

But in our case, I think it kills the value proposition. If you have stamina and talent to deal with complexity today, you could build yourself a fantastic access solution using OpenSSH. Go ahead and do it. But if you want something that requires as little as possible time commitment, and even expertise, you want the right thing by default, so go and download Teleport and see how easy it is. It's a single binary. It's the simplest thing we could possibly think of.

Corey: Getting up and running quickly and easily is helpful. I absolutely agree with that. This is one of those stories where I am in no way shape or form an expert in the area that you have built an entire company around, however, I'm an overconfident white guy on the internet, so of course, I'm going to make unfounded wild speculation, and present it as fact. My experience with the open-source world has always been that people are thrilled to pitch in on open-source projects and get them to a point where they scratch their own itch, but a lot of what makes software usable and approachable by various folks is accessibility. It's UX, it's polish. And that's not really fun for people to pitch in on, in most cases, on a volunteering basis. So, I've always sort of taken the perhaps overly charitable position, that the reason that so much open-source software is crappy is because the stuff that makes it easier to work with isn't the fun stuff to build. You need to start paying people to do those things.

Ev: Oh, that's absolutely true.

Corey: And I do understand that I'm conflating a bunch of things that don't necessarily agree. Open-source does not mean volunteers only: a lot of people are paid to work on open-source; there's a variety different governance models, et cetera, et cetera. Please, please, please don't write me letters on this one.

Ev: I was expecting you to say something more controversial, but honestly, everything you said, I think most of us will agree with. That, yes, it is true that what motivates people to begin open-source projects is to scratch their own itch. For example, why we started Teleport? So, the previous company, the same team here, we built was Mailgun: email delivery, you probably heard of it.

And after Mailgun acquisition by Rackspace, so we joined Rackspace, this much, much larger cloud company. And Rackspace, understandably, they told us, well, you have to migrate Mailgun from SoftLayer, which is now IBM, to Rackspace data center. So, think about it. So, you have this cloud environment that you set up, and you just need to take it and move it somewhere. How long do you think it took us? Wild guess?

Corey: I'm going to guess… well, that's a problem. There's two different directions I could take that in. I could come out with something, “Oh, 30 seconds.” And then it like, “Well, no. We're not that good.” Which is never a good thing. Or I could go the opposite direction, “18 years?” And the answer is, “No, no, no, we're a defense contractor. It took us 40.” So, there's no good answer as far as how to come up with that. So, I would hesitate to even hazard a guess.

Ev: But why didn't you say 18 seconds? Because if you think about it—so someone asked these guys to move a bunch of software. So, what is software? Software is just text files. So, how big are those files?

I don't know, like a megabyte, five megabyte. How much code can you type in the three years? So, even if it, let's say, five megabytes, you take the speed of internet and you just divide that by five megabyte, that's the speed of transfer that software can travel between data centers, so why isn't it seconds? That's the same question to ask. And my non-technical friends, when they asked me, “So what are you doing post-acquisition?” And I said, “Well, we're moving our software from one data center to another.” And they said, “Well, how long does it take to copy a bunch of files? Why is it a project? So, how long it took you?” And I said, “Six months.” And people is like, “Wow. Why is it six months?”

It's because of all of this complexity that we've attached to all these environments. And that's why we started to work on Teleport once we were out of Rackspace. Because even setting up similar infrastructure security takes you a while, okay? So yes, we did scratch our own itch. And the second open-source project we also work on, it's called Gravity. Gravity allows you to take all of your software, like your entire Kubernetes cluster—so that's one important limitation, that Gravity only works with Kubernetes clusters—so it takes your Kubernetes cluster—

Corey: And I personally really hope that the next words out of your mouth to finish that sentence are, “And throw it in the garbage.” But I have a sneaking suspicion, that is not the case.

Ev: You know what? It makes it easier. But let me finish that sentence.

Corey: [laughs]. Of course.

Ev: So, it takes your Kubernetes cluster, and packages it into a single file, similar to a Docker image. It says, “This file is your software.” So, if you want to throw it into garbage, you can literally drag and drop it into a garbage on your desktop. But if you want to drag and drop it into a different data center, you could do that, too. So, now when CIA comes to you, and they say, “We want to use your software, but we don't want your SaaS, we want yourself there on our own top-secret cloud, and we're not going to let you your DevOps people touch it.” So, then you will use Gravity to give it to them and say, “Here is file. And that file is the software you want.”

This level of simplicity is what I've personally been missing since I moved from, kind of, more traditional server programming to this cloud world, where modern cloud applications, they don't even feel like software, sometimes. They feel like it’s an advanced form of configuration for your environment. It's this thin layer of stuff that you spreading across many, many instances on Amazon something. And then you can't even tell where my software. It’s just, like, it's everywhere. It's 15 different repositories, and a couple of Docker registries and no one really knows how to collect it together. So, that's what Gravity does, it allows you to say, “This file is my software,” and then it will just run anywhere by itself.

So, yes, going back to your original question, we did build these things just scratch our own itch, but the ease of use is probably the most important internal metric that we share when we work on this project; simplicity. And that is because it just so happens that making things simple and making it management free is also scratching our own itch. Just think of Gravitational engineers. We've all done our share of DevOps and system administration in the past, and we are also engineers, developers, so we don't really like babysitting hardware. Because when you are babysitting your environment, that's really what you're doing.

So, my ideal version of any kind of software product, it needs to be unmanaged. You know what people say, “Oh, buy managed Kubernetes, managed database, managed this, managed that. It’s like, “Why do we need to manage software?” We dreaming about having self-driving cars. You see the irony here? So, we want cars to be self-driving, but we want to manage software? Why can't we make software that's self-driving first, before attempting something even riskier?

Corey: Oh, absolutely. It's the ideas of things we claim to want, and things we actually want our worlds apart. Great, we have this complete CI/CD system. We had to add a step where a human being can click a button for audit compliance. Yeah, that's not great. And then they automate something that hits a key every 20 seconds on a keyboard, with some physical IoT robot to get around that it’s, okay, at this point, you've built a ridiculous Rube Goldberg contraption. But that's in every CI/CD system on a long enough timeline.

Ev: Let me tell one little example. When sometimes investors would try to book a meeting with us, and they would be asking this questions like, “Do you have any plans to add this advanced user interface or user interface for this or that?” And sometimes I just give them a straight answer, that if you make a feature robust enough and it just works, so you don't really need to manage it, so you don't really need the user interface for it, but I like to give them this joke. Like, look. You have SSD controller in your laptop, right now. Every single employee at your company has that SSD in their machine.

It says these have controllers. Controllers run complex piece of software on them. Do you look for a single pane of glass to manage SSD controllers across all of your employees? Of course, you don't look for that. It's silly. Why? Because it just works. So, why don't we make our cloud software to be like SSD controllers, so then we don't need to have this massive control panels with knobs, and charts, and graphs, and logs? That is the future that we're optimizing for, and I can't wait for this to happen. I'm personally tired managing environments.

Corey: That's, I think, something that everyone wants to say: that they're tired. They're tired of doing the stuff that is garbage and doesn't add direct value. Going with AWS Organizations—which is where I tend to focus on—really emphasizes this. I spend more time setting up subordinate accounts for isolation of workloads, or teams than I often do working within those accounts. It's painful process, it's boilerplate, and solutions are slowly evolving in that direction, but if I didn't have to do a lot of that stuff, I wouldn't. Which means that this ties back to the whole problem of, what are we trying to achieve? And if whatever I'm working on now doesn't directly align with that, I'm probably going to have acid as soon as humanly possible.

Ev: Yes. I honestly, I thought that Heroku-like environments are going to be the future. It's been now, what, almost 10 years since Heroku was launched. I don't feel we have made enough progress on that era. Some people call it serverless, but what I see, like, this serverless movement is being hijacked by cloud providers by simply saying, “Hey, we're going to manage this serverless framework on top of servers, basically, for you.”

But ultimately, that's something I don't want to think about. I want to think about this entire environment: AWS, a region, access point. That's my computer. So, going to push my software into it, preferably as a single file. It just makes it easy for me to reason about this way. And just have it run there by itself. That will be the dream. And I don't want to know about what load balancer is. I don't even understand. If my application has a defined entry point, why can’t this thing make it accessible for me? Why do I need to understand different types of LBs, and how to scale in groups? I think we're pretty close to closing this gap in our abstractions. So, I think it's about time, and Gravitational is working on it.

Corey: Before we call this an episode, there's one more thing I want to talk to you about. Now, for folks who have not been on this podcast before—which I'm told is still more people than have been because I don't have that many episodes yet—one of the things I do on the background process is, I start having a quick conversation before I whack the record button. And I asked a very small list of questions. One of them—which leads to fascinating answers—is, what do you want to make sure that we don't talk about because this is not a story of, I only tell the stories people want to have told, but if I sit here, and I beat someone up on PR missteps that their company has made, for example, it's not a good episode. It's awkward, it's uncomfortable, and no one likes it, so I want to avoid those things, if possible. And I've gotten a range of hilarious answers over the years of asking that question. But you gave me a great one, which is specifically, don't ask you to shit on other company’s technology. I love that answer. Tell me more.

Ev: Well, I believe that we are much better off when we build on top of lessons that we learn from each other. No single company—even as powerful as Microsoft, Google, or Amazon—is capable of solving every single problem in the best possible way. So, if I make a mistake, I don't mind that some other open-source project will come in and correct me for it. That keeps me honest. And I will do the same to other companies who are working in open-source space.

And also, as an engineer, I love stitching the best open-source tools from different authors, from different vendors, to assemble solution that works for me. So, by criticizing each other's projects, we’re not really helping anyone making these choices. But what I do think is fair is criticizing approaches. For example, what I earlier said about SSH keys. It's just not very scalable way of doing it, and I could argue, using very technical arguments for it. But genuinely, I do believe that every open-source project, every product out there deserves some attention, and ultimately we should be trying to integrate with each other, hopefully, using open-source, open standards that leads to this outcome that I'm dreaming about.

Corey: People think that a lot of my brand is built on crapping on company’s technologies, and they're kind of right, but I'm careful to punch up. There's a reason that I own twitterforpets.com. If I make fun of someone's actual small startup, I'm a jerk because that's people's hopes, dreams, et cetera. Most of the companies I make fun of, other than a few very egregious examples, are either multibillion-dollar entities or are publicly traded.

At that point, you've opened yourself up to scrutiny and criticism, and frankly, you can weather my slings and arrows in a reasonable way. If people are listening to what I say and feeling bad as a result, I've failed somewhere. I also think that when companies try this—our entire marketing brand and persona is going to be built on crapping on other people's work—it doesn't look good at all.

Ev: I agree. Look, even going after larger companies, you have to remember that the companies are not monsters. Take, for example, Amazon famous for the two pizza teams. So, if you picking like a particular Amazon offering, and you going after them, there is basically, like—what—10, 15, 20 people behind it. So, that's really the group you having an argument with; it's not all of Amazon. And they all have feelings, and we all work hard, and I do believe that the at the end of the day, that we love technology, we like computers, we like what we do, so creating drama, necessarily, that's not something I could be excited about.

Corey: Yeah. At that point, the feud becomes the story, rather than the actual value of what it is that you've built.

Ev: [laughs]. Exactly. It's just too much of that is happening in open-source space, and that is unfortunate.

Corey: It's the rage-fueled equivalent of we don't have much useful to say, so we're going to throw a big party at a conference.

Ev: Yeah.

Corey: So, one last question to get the slightly back to topic before we wind up calling it a show. Do you think that, as we look at what's happened over the course of the industry, progressing from running things in mainframes to the whole data center story, to the thin client, thick client, et cetera, back, now we're in a world of cloud, do you think this is the end state of, I guess, the pinnacle of computing? Nothing to go beyond this; shut it all down, we're done. Where do we go from here?

Ev: I think there is definitely going to be another thing that will come to replace the Cloud. And I hope that soon we will be able to say that, “Oh, if you're doing cloud-native computing or cloud-native application, that's the legacy way of doing thing.” I don't know how this next thing is going to be called, but I do like to think about it a lot, like, what it will look like. And I think one area for improvement is for us to close this gap.

We keep saying the data center is a new computer. Like, data center is a computer. Data center is a computer. Kubernetes is an operating system for data center as a computing device. A DC/OS from Mesosphere, remember? But it just hasn't happened yet.

Kubernetes still feels like a collection of primitives to manage a bunch of containers, and plus the million of other stuff—it just feels like you're dealing with drivers. If you compare Kubernetes to an operating system, I say that drivers are too exposed. If I'm building an application for Linux, or Mac, or Windows, I don't think about USB drivers if I wanted to take sound from USB mic. But modern cloud environments, they make us think about load balancers, volumes, and all those stuff that is just too low level. So, I believe that we will arrive at this post-cloud world where a data center truly becomes a computer; where the process of creating and publishing an application for Mac OS, or iPhone, or AWS will be extremely similar, where you say, “Here's my file. This file is this image, it's my application. You could put it into a AWS account, and it will run.”

That's what I believe the post-cloud world is going to look like, and it's almost like Gravitational vision for the serverless. Because today, when we talk about serverless, it's just basically another framework on top of something like Kubernetes. But we believe that serverless is when you—the process of building an application ends with a single build artifact: this file is my software. Where it will run, I don't want to care. I don't want to know. That's a true serverless to me. So, hopefully, once that happens, we could finally say goodbye to cloud-native world. Welcome to this post-cloud world that Gravitational is trying to enable.

Corey: And you're doing a better job of articulating that vision and that story than the certain large company that did a cloudless hashtag that resulted very quickly in being, effectively, cyber-bullied off the internet because the entire premise was ludicrous. I think that you're right. I think there is definitely something that has to come next. If not, then what are we all doing here? We could not have imagined 15 years ago, a lot of the things that we take for granted today. And I imagine that that trend is not likely to slow down anytime soon. I feel like that is tempting fate to make that observation in 2020, but I mean it from an optimistic point of view.

Ev: Agreed. Agreed. I'm not going to talk about those other companies because that's specifically something we, I asked you not to ask me about.

Corey: Exactly. We're not going to name names, it's fine. But if it helps anything, they’re worth tens of billions—sorry, my apology. They’re worth hundreds of billions of dollars. So, again, I don't consider it punching down, which is really I think shorthand for what I view this as.

It’s, I learned, you can punch down at big companies when they're trying something new and you're crapping on them, or they're talking about their journey and how they wound up going somewhere. I got that one wrong in the early days. And well, “I’m sorry, I'll do better,” is sometimes the only thing you can say.

Ev: Sounds like a plan.

Corey: So, if people want to learn more about what you have to say, what you have to show them, what do you have to sell them, in some cases, where can they find you?

Ev: They could go to gravitational.com. They can click on Teleport to learn how we do this magical access to everything, or they can click on Gravity to learn about packaging applications as a single file. Or they could go on GitHub and just dive straight into the code. It's github.com/gravitational.

Corey: Excellent. And thank you so much for taking the time to speak with me. I really appreciate it.

Ev: Thank you for having me, Corey.

Corey: Ev Kontsevoy is the CEO of Gravitational. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts, whereas if you've hated this podcast, please leave a five-star review on Apple Podcasts along with a comment explaining why everything Gravitational is building is overly complicated and unnecessary because all you really need is Telnet.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Fredick “Flee” Lee

Fredrick "Flee" Lee is the CSO at Gusto, the platform that helps 100,000+ small businesses nationwide pay, insure, and provide benefits for their teams. Flee has more than 15 years of experience leading global information security and privacy efforts at large financial services companies and technology startups, most recently as Square's head of information security. He previously held senior security and privacy roles at Bank of America, NetSuite, and Twilio.

Links Referenced:

  • LinkedIn: https://www.linkedin.com/in/fredrickdlee/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Catchpoint. Look, 80 percent of performance and availability issues don’t occur within your application code in your data center itself. It occurs well outside those boundaries, so it’s difficult to understand what’s actually happening. What Catchpoint does is makes it easier for enterprises to detect, identify, and of course, validate how reachable their application is, and of course, how happy their users are. It helps you get visibility into reachability, availability, performance, reliability, and of course absorbency, because we’ll throw that one in, too. And it’s used by a bunch of interesting companies you may have heard of, like, you know, Google, Verizon, Oracle—but don’t hold that against them—and many more. To learn more, visit www.catchpoint.com, and tell them Corey sent you; wait for the wince.

Corey: This episode is sponsored in part by strongDM. Transitioning your team to work from home, like basically everyone on the planet is? Managing gazillions of SSH keys, database passwords, and Kubernetes certificates? Consider strongDM. Manage and audit access to servers, databases—like Route 53—and Kubernetes clusters no matter where your employees happen to be. You can use strongDM to extend your identity provider, and also Cognito, to manage infrastructure access. Automate onboarding, offboarding, waterboarding, and moving people within roles. Grant temporary access that automatically expires to whatever team is unlucky enough to be on call this week. Admins get full audit ability into whatever anyone does: what they connect to, what queries they run, what commands they type. Full visibility into everything; that includes video replays. For databases like Route 53, it’s a single unified query log across all of your database management systems. It’s used by companies like Hearst, Peloton, Betterment, Greenhouse, and SoFi to manage their access. It’s more control and less hassle. StrongDM: Manage and audit remote access to infrastructure. To get a free 14-day trial, visit strongDM.com/sitc. Tell them I sent you, and thank them for tolerating my calling Route 53 a database.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Fredrick Lee the CSO of Gusto, but the world knows you as Flee. Welcome to the show.

Flee: Thanks a lot for having me, Corey. Yes, the entire world knows me as Flee, and as I've told other people before, pretty much the only person that calls me Fredrick would be my parents, so Flee is what I'm most comfortable with, and I appreciate you taking the time out to actually chat with me today.

Corey: No, by all means. So, a few things to go into. But first, you're the CSO—again, three syllables, not one. That stuff matters—or two, whatever it is, again, it's always three. Three is the right number of syllables as longtime listeners of the show are well aware, but that differs from CISO: C-I-S-O. What's the deal?

Flee: Yeah, so it's literally in the abbreviation itself. A CISO is Chief Information Security Officer. In this [unintelligible] I think a lot of people are more familiar with that as, kind of, that role when they think of cybersecurity. I am the chief security officer, which means that I have a much broader view and a much broader mandate with regards to security. So, part of my responsibility is also the physical security and the safety of Gusto, our assets, as well as our people, along all those lines.

So, the nuance there is if you're a CISO, you get to spend a lot more time just focused on just the cyber side, and punching keys on the keyboard. When you're a CSO, you have to think about things like, “Hey, is this building going to be resilient towards an earthquake?” And some of those other fun topics. But it's all, at the end of the day, it's still risk management, and it's just a matter of, like, how broad do you actually want to go with that risk management. And there are always pros and cons of which roles you want to have.

I have a lot of fun being the CSO because it also just really brings so many other things into our toolkit with regards to just security. One of the classic things that I think people often forget about is that bridge between our cyber world—I mean, like, information technology, et cetera—and the physical world. A common example of that is badge access, right? So, you think about all the things you can do as a CSO when you can combine information and telemetry from your badge access for buildings into telemetry regarding how users are actually logging into endpoints, et cetera. So, it's a really, really great role. I'm very, very thankful for all the people that have helped support me to be successful here. I have a phenomenal team. And yes, it just a lot of fun being a CSO.

Corey: So, in the interest of full disclosure here, here at The Duckbill Group, we are Gusto customers, which is great. You folks are not sponsoring this episode, so basically this is just my way of getting trouble tickets escalated. But for those who aren't aware, so we can have a conversation that doesn't leave everyone scratching their heads, Gusto provides, effectively, payroll and benefits as a service, mostly. I have better things to do with my life then sit there and click through payroll, and running all of that myself. It's awful, and you never get it right, only wrong, so I pay a third party—in this case, Gusto—to handle a lot of that for me. Given that we have employees in multiple states, it handles a lot of the administrivia, that is kind of awesome.

Now, security is one of those areas that really matters for something like this because, yeah, my business partner and I got to pick who we use as a vendor, but our employees do not. So, they show up, and, “Hey, we're going to be giving all of your deeply personal information: your social security number, the validation of your legal right to work in the country, et cetera, to a third party, and if you don't like it, you can quit.” is more or less the position that people are confronted with. So, everyone has to care about security, but you folks have to care about it in a way that goes beyond what for example, my Twitter for Pets side project does.

Flee: Oh, definitely. And we take it really, really personally as well. It's fundamentally part of our DNA. Without strong security and without strong privacy, we can't really deliver the Gusto service in the way that our customers deserve. And I love one of the things you actually pointed out, Corey, which is, even though we want to make Gusto as delightful as possible across the board—but both for you as an employer, but also for your employees—we recognize that we have an oversized burden to make sure that we're extra cautious due to the nature of the data, and also because of how people get onto our platform.

Philosophically, we view ourselves as data custodians, as opposed to data owners. And there definitely is a difference with how we treat your data as opposed maybe how some other companies would treat your data. When you view yourself as a data custodian, you're always are thinking about the human at the end of that data, and what would that human want done with their data? So, for the example that you mentioned, Corey, for the people that are your employees, you're exactly right. They didn't opt into Gusto, so I feel like we've been would have a higher responsibility to these individuals to make sure that we're doing everything possible to make sure that their data is not only secure, but also respectful of their privacy, and that includes things like not selling their data—which I wish was just an obvious choice for other companies—all these other kinds of things that you would actually kind of want to have as an individual when you would hand over something that's precious to you to a third party.

You want to make sure that Gusto or any of your other suppliers are always acting in your best interest, and putting your interests first, as opposed to putting the interests of the company, or some random advertiser, or some marketing firm that just wants to collect the data. So, it really is, it’s critical for us to nail security. And it's also critical for us to nail privacy because this is ultimately some of the most sensitive data that people have in their life, and we recognize the responsibility that actually comes along with that. And yeah, I mean, it's definitely a privilege to have people trust Gusto to not only deliver them a great employment experience, but to also do that in a secure, private, and ultimately respectful manner.

Corey: One of the biggest challenges I find in security is akin to what I deal with dealing with the AWS bill and the pain that it causes. Everyone really cares about this and tells you how much they care about this right after they really should have cared about this. Now, the stakes are lower when it comes to cost than they are with security. There's headline risk, there's the actual real harm that breaches can do to people, when in some companies where a security breach could actively lead to people being killed, depending on the circumstances. Now, Gusto is generally not at that tier, but you are the custodian of other people's information. So, I guess this is a question that I normally ask with a very different tone, but how do you sleep at night? How do you sleep at night?

Flee: How do I sleep at night? It’s an interesting thing, and actually I want to flip that question over. Obviously I have years of experience in the security industry, and it's easy for people in security to maybe get a little… I don't want to say necessarily lazy, but jaded because of all the things that we actually see. For me, I focus more on what gets me up in the morning. And the reason why I say that is exactly what you said, Corey.

The nature of the data that we deal with is extremely sensitive, or as you articulated that, yes, maybe Gusto, we're not handling nuclear secrets, we are handling people's lives, though. Those names, those addresses, insurance information, people's ideas about how they actually want to save for retirement, et cetera, those are human lives. And I take that very, very personally and very seriously, and so I wake up every morning energized about what else can we do to make this safer for people, but also easier for them to take advantage of that safety. But you're exactly right, it is a serious thing, but it's one of those things that I believe that we actually built a great infrastructure and good team so that we are more energized by this problem we need to solve, as opposed to maybe scared and apprehensive, staying up at night. I feel like if you're staying up at night, maybe you don't have great plans. I do feel like we have actually have some really good plans here at Gusto, as well some as already great things we've already implemented from a security and privacy standpoint.

So, I'm not a huge fan of actually being in the waiting game either. And this goes somewhat into my philosophy about how to actually build and construct security teams. Yes, obviously, traditional security teams do spend a lot of time quote-unquote, “waiting,” But I think that's not the best utilization of a security team. Your security team should be much more proactive: they should be finding new things that they could be building, they should be finding new ways that he improved the infrastructure, they should be finding new ways to make it even easier for our internal employees to build great secure products by default. And if you're in that scenario, if you have built a security team that's kind of like an enabling team, it's really all about engineering, building great security controls, building great security frameworks, then you're constantly having your security team with a mission to do, and not just a wait and see what the attackers are going to do.

We should always be feeling like we need to be out-competing attackers at every given second, and we need to be energized—not just a Gusto but the industry overall—and move away from this previous model of just waiting for some blinking lights on the screen, but being more to a model of building things to either automatically address those blinking lights/alerts, or automatically remediating issues as they pop up.

Corey: Before you wound up a Gusto, you did a number of senior privacy and security roles at small companies without a lot to lose, like Bank of America, Twilio, and Square. So, on some level, it feels like this is, okay, compared to some of those things, you have all the same personal information of a large number of people, but at least this time, you don't generally hold their actual money, presumably. Although then again, we do have bank transfers going to you folks, which are then paid out to tax authorities and to employees, so maybe I'm speaking too soon on that.

Flee: Yeah, I mean, obviously, we deal in finance as well because we do move money. And we do help people manage their money as well. There's a product that we have, Gusto Cashout, which allows employees to have more control over their paycheck and when they're paid, and so there's a lot of the same problems. You indicated I’ve kind of been in the business of protecting people's really, really sensitive data for quite some time, even at those small tiny companies like Bank of America or Square, but it is one of the things, we're dealing with a lot of the exact same things here at Gusto, but there's another interesting element, I believe, with the nature of the data we deal with here at Gusto is that we're also doing this on behalf of a lot of really small and medium-sized businesses. And to that extent, I feel like there's maybe even an extra obligation that we have because not only do I need to be a great service security person for Gusto, I need to make sure I'm also being a great security person and a great security team for our Gusto customers, and that means thinking really, really proactively around some of our security controls that end-users touch.

Now, at Bank of America, that's kind of like a different situation because Bank of America thinks that their customer and their profiles are maybe a little bit different. Bank of America has a lot of B2B customers as well. Twilio is primarily a developer API type company, still sensitive data, but the interactions with our end customers is a little bit different. We do know that at Gusto, that it would be unfair of me, as a CSO of Gusto to expect Corey to have a full-blown security team and it should be part of my job to help Corey solve security problems, and to make sure that your company is secure at a minimum when you’re actually interacting with Gusto, but ideally if there are more things we can actually do along those lines, are there other things that we can do proactively to help Corey ultimately just have a successful business? And I do believe as security practitioner—and I hope the security practitioners believe this—it’s part of our job and part of our calling as security professionals to make the world more secure. Not just our company, but also our customers and other people that are impacted by us in the ecosystem.

Corey: You've been focused on something that you're calling lovable security, which sounds like a ridiculous oxymoron from where I sit, like military intelligence or IBM Cloud. Tell me more about this.

Flee: Yes, yeah. Loveable security. [laughs]. I personally don't see it as an oxymoron. Partially, I love security. I've been in this, I guess, industry—it may be better actually described as I’ve been in this discipline of cybersecurity since I was a teenager. I mean, I've always loved all the problems and challenges that we get to solve.

However, how people experience security has not always been great. And that's either in the form of how they interact with security, maybe at their company, maybe how they interact with security at a provider that they have, or how they interact with security even in their personal life and their personal computer, et cetera. The stereotype of the security team is, “Oh, it's a bunch of grumpy, mean people that, every time you want to buy something, or every time you want to install a new app, or every time you want to use a new technology, they immediately say no, and they're always grumpy, and they're constantly trying to find some way to essentially stop you, to essentially slow you down, to make things more difficult.” You know, they're the kind of security team that says, “Hey, you need a 26 character password, and you have to change that password every 90 days.” Things that ultimately just aren't good experiences.

And we know that in order for security to really be successful we need to drive behavioral change in people. And you get people more interested in collaborating with security when your security team when your security philosophy when your security discipline is much more approachable. And that's where the lovable comes in. It’s this whole notion that security control, security people should manifest in such a way that the end-user doesn't try to avoid it. And a great example would be some of the things we're actually seeing now in modern computing when it comes to things like authentication. So, for example, if you use an Apple device, for example, then Apple has encouraged people to have strong passwords on their mobile devices because you get a benefit of being able to use Touch ID or Face ID, which is a much more pleasant experience, but it actually gives you even better security.

When it comes to people inside of a workforce, if your security team is a team that partners with other business units inside of your company and helps them actually build to find great solutions for security challenges, that becomes lovable. This whole idea of having a security team that you actually go to proactively because you see them as a source of help, as opposed to a source of fear or a source of potential trouble. Your security team should look a lot more like your personal trainer or your fitness coach in the gym, as opposed to the mean, angry bouncer at the nightclub. Actually, I don't know that much about nightclubs, but I know that they have bouncers. But, you know, that kind of thing.

Corey: So, I did something for the first time earlier this year, before these unprecedented times—do you remember, back in precedented times, which we all miss—and I went and walked the expo floor at RSA and looked at the program, and all the rest, and I've learned a few things. One, you're not allowed to give a talk at RSA without the word firewall in the title. Two, there's an awful lot of companies sitting there at the expo floor selling rewarmed versions of the exact same thing, and none of them seem to be focused on actual security improvements. Instead at all seems to be directed at, more or less, box-checking, and ensuring that once you have your data breach, you can point to having done all the right things in a fruitless effort to prevent it. Is that an overly cynical take, is it spot on, or something else entirely?

Flee: [laughs]. I’m laughing because I might be one of the worst people to ask this question of because I don't think it's overly cynical. I do think for the majority of it, it is spot on. This is an area where the security community should and could do much better. One, there's tons and tons of companies that are trying to sell you things that say, “Oh, this is going to solve all of your problems,” or, “You're going to magically make your data loss issues go away,” or, “We're going to prevent all attacks from this particular vector.” when the reality is, we know that that's not a real thing.

We know that a motivated attacker can definitely circumvent all kinds of various different controls. We know that there are human elements that can make some of our security controls fail as well. And I do agree with you, a lot of these things are almost like rehashes of each other or just clones, but even more so, they're not really trying to help you manage and reduce risk, they're really just trying to help you feel better about your security posture, and to give maybe your execs, or the board, or an auditor something to look at. And there's a lot of products that are also focused on trying to solve problems that probably aren't the best risk reduction return for you. You go to RSA, you walk the expo floor—or any of these other security events, or just vendors in general—and you always hear that the classic acronym that makes me cringe every single time, APT—Advanced Persistent Threats—some people actually do that, most people don’t. You hear all these things around zero-days. Yes, zero-days are an issue, but that's not generally how people get compromised. They’re getting compromised by vulnerabilities that have been living with them for 90, 180, 360 something days, et cetera, and you don't see a lot of vendors addressing these fundamental security issues.

Some of these things that are actually more process-related, or just ultimately aren't as sexy, and even more so, they don't believe they can attach a six- or seven-figure price tag to. Some of the best security tools that I have found and have used have all been open-source and free. And if they weren't free, they were ridiculously cheap. And there's so much power we can actually get out of that. It is also one of my frustrations with some of the cloud providers, with some of the security tooling and features that they provide where it's actually just overpriced.

And so, when you think about the vendors in general, I just think security vendors can do a much better job, but it also requires us in the broader ecosystem of technologists and consumers, to force them to be more reasonable with their pricing, but even more so, force him to be more transparent and honest about what their tools actually do. And it's perfectly fine for a tool to only solve one particular thing, or to be targeted at a really, really small scope. In fact, I actually would argue that may even be better. So, I just think that we have to start holding our vendors a little bit more accountable and asking for a lot more transparency. But that’s also going to require some of us in the security industry to be more honest with ourselves.

Some of the reasons why the vendors can get away with some of the things that they do is that we in the security industry have often forgotten our technical roots, and we haven't put our technical hats back on, and aren't keeping abreast of modern technologies as well as we could. And because of that, yeah, vendors can kind of pull the wool over some people's eyes, and promise a product that we know from a technical standpoint either can't work, or won't work that well. I think that's actually one of the areas—or again, another area that there's almost a shared responsibility, shared blame: blame on both the vendors for being way, way, way, way too shady, but also blame on us in the security industry for not staying as technical as we need to be, to not staying on top of technology in the same way that we need to be.

Corey: In what you might be forgiven for mistaking for a blast from the past, today I want to talk about New Relic. They seem to be a relatively legacy monitoring company, and I would have agreed with that assessment up until relatively recently. But they did something a little out there: they reworked everything. They went open source, they made it so you can monitor your whole stack in one place and, most notably from my perspective, they simplified their pricing into something that is much more affordable for almost everyone. There's even a free tier with one user and 100 gigs per month, totally free. Check it out at newrelic.com.

Corey: So, one of the things I saw somewhat recently was a list of slides for all of the different security services that AWS has, and I know more than I probably should about basically every AWS service because I have a trick memory for stuff like that, and I fix the AWS bill, so I'm most familiar with all of their various pricing models. And I'm doing the mental math, figuring out how much all these things would cost, and yeah, it's clear that just having the data breach would be less expensive overall, so go ahead and do instead, past a certain point. I kid but not by much. I've always been very down on the concept of charging extra for security in a variety of different platforms. If Gusto, for example, charged me an extra fee to add multi-factor auth to my login, I would not be a Gusto customer, for example. Where do you land on this?

Flee: You and I are cut from the same cloth in that, Corey, because yeah, I mean, I would never want to work at a company that was charging users more to be secure. I feel like that should be a default. We shouldn't charge you for 2fa, and with AWS, it really is somewhat of a pain point and a frustration for me. Obviously, I do want great security, and that means at the end of the day, we're going to pay what we need to pay in order to get that. At the same time, I feel like a lot of things should be almost just included in platforms at almost, like, a bare minimum.

When you think about some of the great things that AWS has as services, we should want every AWS customer to have those by default. You think about things like Macie, for example; it's helping people discover if they have sensitive data that they have in S3 buckets, or just in other areas in AWS. That should be something that we just give to people for free because at the end of the day, my hope is that by making AWS or other cloud providers, or just software in general, secure by default, having those security features easy and free, or really, really, really cheap, we get more people to adopt it. So, more people would want to be on AWS if they realized that they could actually have the exact same security controls or better security controls in the Cloud. And then there's probably, actually, maybe some edge cases where AWS may want to charge some additional prices, I don't know Amazon's complete business model, or how they actually price things out [crosstalk]—

Corey: —anyone knows Amazon's complete business model. My product strategy is, “Yes,” So, who could say?

Flee: Yes. [laughs]. I’m not even sure if Amazon knows themselves. I think their business strategy is making money. But there's so many just small, bare minimum things, and I think ultimately it can drive greater adoption. If we made it even easier for people to run good tools to understand, are they under attack? Do they have a misconfiguration? And I think everybody knows that IAM is this great, extraordinarily sharp knife, that's there for developers to cut themselves on, make security mistakes—

Corey: And when you get it wrong, it's your fault.

Flee: Yes, yes. [laughs]. And wouldn’t it—

Corey: Haven't you read the shared responsibility model?

Flee: Yes. [laughs]. And it's like, you know, I can have a shared responsibility model also with my five-year-old nephew, but it doesn't really seem to quite equivalent. But we think about some of these other AWS services, you think about things like GuardDuty, think about all these other things—you think about, actually, just even their Security Hub, and that Security Hub should just be free, and it should just be for everybody. AWS should actively encourage and want people to have better security because ultimately that could even be a competitive advantage for them.

It makes it even easier for people to actually get into and drive adoption. And the reality is a lot of people were attracted to move into the Cloud because they thought that there was a promise—whether real or not, or articulated or not—that the Cloud was going to make it easier for them to operate at speed, with agility, and with safety. And a lot of really, really small businesses got their starts in AWS, and it'd be even better to encourage other small startups and other small businesses to operate in AWS, but they need a security team, or they need a security expert and they need, like a really, really, really well tuned-in DevOps engineer, in order to be successful doing that. That definitely is not lovable, and I think one of the things that would be great to see from Amazon, or even other cloud providers because Amazon's not the only one out there, is to make all those security features—or at least think about some core critical ones—make those free, make those extraordinarily cheap. Are we in a day and age where people even need to pay for certs anymore?

We should make TLS just prolific and extraordinarily easy for people that are actually in AWS. Should we need AWS customers having to think about DDoS protection, and trying to actually figure out, well, how much should I pay for? Should I pay for it? That seems like the kind of thing that you would just want by default by moving to the Cloud and taking advantage of that. Do we need a small startup of four or five brilliant little engineers worried about how to configure Amazon’s Shield and Amazon's WAF, or if they should even pay for those things? I would hope that those kind of things would at some point become free, or much, much cheaper than what Amazon currently offers.

There's so many great things that are in Amazon, that so many security engineers—or just security-conscious people—want to take advantage of, but there is a ceiling of the cost. And AWS bills can already be really expensive to start with, and for the uninitiated or for the people that maybe aren't as sophisticated, it's easy to say, “Well, this security thing looks really, really expensive, and I don't quite understand what the benefit is.” Because often security tools are targeted at preventative controls or helping to detect certain black swan events and so you don't always realize the value of it until after an attack has occurred, even though it's actually something we want everybody to have; it's kind of like fundamental things. There's so many great things, I think Amazon even has vulnerability scanning was one of their offerings, and there's AWS Inspector—the list of security services, actually, as you said, Corey, it's really, really broad and it’s massive, and I think maybe an even more interesting experiment would actually be to go out and see how do these services that Amazon offers from a security standpoint, how do they actually cost people in the real world operating at scale? So, when you're operating at quote-unquote, “Netflix size” or something like that, and how does that actually compare to actually buying some other commercial alternative? Or how does that compare to using an open-source project? Or also—

Corey: In some cases, you can pay for that particular security service, or staff a team of eight people and save money.

Flee: Yes.

Corey: It's ridiculous on some level. And I would also argue the competitive threat story. If there's yet another S3 bucket leak that makes the headlines in The New York Times, then the response from the world is not, “Oh, Amazon is insecure, I should move to Azure instead.” Instead, it's, “Cloud isn't secure,” So, whenever there's a setback like this it affects the entire industry. And the S3 stuff drives me up a wall because it's pernicious, I understand how it got there, but if you're managing data that you have to be wary of, it is not that challenging to ensure you get there.

The Capital One breach, for example, was not an S3 bucket problem, it was a series of various small misconfigurations chained together. And that's still not okay, given that they are a bank, however, it was a sophisticated attack that exploited those small misconfigurations, not, “They forgot to check the box somewhere.”

Flee: Yep. Definitely agree with you. And I love one of the analogies and the fact that you also mentioned in a bank at the end. So, previously, we were talking about a small company I used to work for called Bank of America, and this is a podcast, so people can't see all the gray hair that I have, but earlier, several decades ago, online banking was a new thing, and one of the things that we recognized in the banking industry was that it was important for all banks to be secure because we wanted customers to feel secure using online banking. So, we recognized that a customer wouldn't go and say, “Oh, Bank of America is insecure, so I'm going to go use Wells Fargo.” A customer, instead, would just say, “Online banking is insecure.”

And we're still somewhat in that situation when it comes to cloud providers. The Capital One situation was really, really interesting. As you said, it was very, very nuanced, but it still had somewhat of a chilling effect. And there are still companies where mentioning the word cloud or trying to get a migration to the Cloud is still taboo. And these little small things at up, and it adds up into the decision-making of other executives within those companies.

So, when you think about a CFO or CEO, they may not be following Amazon tech in the same way that we do, as a CSO, or CTO, or even a CIO, and because of that, all they hear is what they saw on the news. They heard that, hey, a bank was compromised, and they ran in the Cloud. They don't pay attention to the rest of nuance there. And when there's all these stories that are related to compromises, as you said, things like S3, things like some of the concerns around AWS’s previous metadata service implementation, when you hear all these problems, you can easily see how that could cause fear and this chilling effect of these larger corporations that want—or at least have individuals that want to migrate and move to the Cloud. And it’s a really interesting thing because I feel like Amazon definitely figured the story out with regards to scaling compute, but they had the same opportunity to scale security.

And it's not just Amazon, obviously all the cloud providers. Some do better, some do worse, but I think overall as an industry, the cloud providers can do a lot more for their customers and a lot more for the ecosystem when it comes to security and making it very, very, very easy for advocates of cloud technology to be able to explain to other people inside their organizations why moving to the Cloud can be done and also be done safely.

Corey: The increasing challenge that I see is that there are so many different levels and layers to security. Yes, this is a complicated, sophisticated space, but at the same time, as you were alluding to, you cannot have a scenario where you're a four-person startup and you need to hire three full-time security people in order to keep all of it straight, and even then potentially get something wrong. It feels almost like a losing battle.

Flee: Hopefully, it's not a losing battle, but it definitely is more complicated than it needs to be. There are some really, really good resources for tiny companies trying to get started. And free resources. Like obviously there's OWASP if somebody just wants to be a self—motivated engineer to go learn some things. I think everybody who's interested in Cloud should be reading every single thing that Scott Piper ever writes, that helped them also be a little bit more self-service.

But again, today, you're still trading time, so, instead of actually working on those features that you need to work on as a small startup, your four-or five-person company, one of your biggest risks is that you're not building fast enough, and you're not building the right thing, and so to have to spend some of that time, also worried about actually solving security feels unfair, or at least not as helpful as we could be—or the industry, as security professionals could be, and the industry for cloud service providers could be. There are some good additional resources. I know that there is a startup security organization—there's actually a couple of CSOs that got together, like he wrote some guidance to help startups but still requires at that time investment.

What I would love to see instead is Amazon being extremely opinionated—and not just Amazon, but other cloud providers as well—and making it really easy for people to make the right security choice by default. And I know that Amazon has made some improvements there with regards to how S3 buckets work now with regards to public versus private, a lot of other things they do on the network side, but I think there's a lot more they can do so that all these tiny companies have security by default if they're kind of on this golden path. And this notion of building a golden path, I think is fundamental to this idea of supplying lovable security. You have to make the right thing to do the easiest thing to do. And Amazon, GCP, Azure, et cetera, they still have quite a ways to go before the right thing to do really, truly is the default and easiest thing to do, in particular for unsophisticated users.

Corey: I think one of the hardest parts of all of this is when you look at the entire landscape of what you can do, there is always an infinite amount of work that could be done to improve the security posture. There's going to be an inherent trade-off, in some cases, between usability and security, and it's easy to look at that entire thing and realize, “To hell with it. I'm just a small company, no one's ever going to discover my open S3 bucket.” And not do any of it. It feels like there's an 80 percent that takes 20 minutes, and then you can spend a career getting that last 20 percent.

Some of you folks, for example, that handle the payroll for people like me, absolutely should. For the rest of it, Twitter for Pets, not too many people want to see dogs DM’ing each other and what they're talking about. It's just a lot of angry noise. So, I think understanding that that spectrum exists—there's always going to be more work to do but that shouldn't prevent you from starting is one of the biggest challenges I'm seeing.

Flee: Definitely agree with you. Hopefully, people don't get overwhelmed and believe that they can't get started at all. And that's where this notion of pragmatic risk comes into play. You're exactly right: the risk profile of Twitter for Pets is very different than the risk profile of Gusto. The assets that we keep at Gusto are different.

And really what companies need to focus on is, are the security controls appropriate for the nature of the assets that they manage and control? So, if your assets are just cat pictures, or I guess in your case, puppy pictures on Twitter for Pets, yeah, the controls you want to put in place are very different than if your assets are payroll information, other finance, people's personal details, et cetera. And so, taking time to dive into this kind of this concept of pragmatic risk is really where people have to go. And if you run a business, you're already you’re familiar with risk management anyway because so much around running a business is trying to make trade-offs.

Okay, what is the most important thing to work on? What is the risk if I don't roll out this feature? What is the risk if I don't service these kinds of customers? You also have to just include, well, what are the risks of certain kinds of security events happening to my company, and what am I willing to pay to mitigate that risk? It’s not a one-size-fits-all kind of scenario. The nature of security that you want at a Twitter for Pets or cat picture website company is definitely very different than what you would want even at actual Twitter.

Corey: Well, after having this conversation, I am a little bit more confident, I guess, in my choice of how I'm handling payroll at the moment. Flee, thank you so much for taking the time to speak with me. If people want to learn more about what you have to say, where can they find you?

Flee: Yeah, I'm pretty easy to find on LinkedIn. That's probably where I'm the most active and it's actually a great area for me to actually get connected to other security professionals or just anybody who wants to talk about technology, talk about Gusto, talk about diversity in tech, or just talk about why I have the best opinion on barbecue. So, that's always a good place to actually find me.

Corey: Excellent. And we will of course put links to that in the [show notes]. Fredrick Lee, better known as Flee, CSO at Gusto. I am Cloud Economist Corey Quinn and this Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts, whereas if you've hated this podcast, please leave a five-star review on Apple Podcasts as well as a badly spelled comment telling me that I'm doing it wrong with outsourcing payroll and that Gusto is just somebody else's payroll specialist.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Corey Sanders:

Corey Sanders is the Corporate Vice President for Microsoft Solutions, an organization dedicated to partnering with customers as they transform into successful digital businesses.

He is responsible for sales strategy and corporate technical sales across Solution Areas and Teams that include Azure Applications & Infrastructure, Azure Data & AI, Business Applications, Cybersecurity Solutions Group, and Modern Workplace. His focus also includes selling the full value of Microsoft cross-cloud solutions and advancing the technical depth of the Microsoft Solutions team.

Prior to this role, Corey was Head of Product for Azure Compute and the founder of Microsoft Azure’s infrastructure as a service (IaaS) business. During that time, he was responsible for products, strategy and technical vision aligned to core Azure compute services. He also previously led program management for multiple Azure services. Earlier in his career, Corey was a developer in the Windows Serviceability team with ownership across the networking and kernel stack for Windows.

Corey joined Microsoft in 2004 after graduating from Princeton University and resides in New Jersey.

Links Referenced:

  • Microsoft: https://www.microsoft.com/
  • Twitter: https://twitter.com/coreysanderswa
  • LinkedIn: https://www.linkedin.com/in/corey-sanders-842b72/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Catchpoint. Look, 80 percent of performance and availability issues don’t occur within your application code in your data center itself. It occurs well outside those boundaries, so it’s difficult to understand what’s actually happening. What Catchpoint does is makes it easier for enterprises to detect, identify, and of course, validate how reachable their application is, and of course, how happy their users are. It helps you get visibility into reachability, availability, performance, reliability, and of course absorbency, because we’ll throw that one in, too. And it’s used by a bunch of interesting companies you may have heard of, like, you know, Google, Verizon, Oracle—but don’t hold that against them—and many more. To learn more, visit www.catchpoint.com, and tell them Corey sent you; wait for the wince.

Corey: Normally, I like to snark about the various sponsors that sponsor these episodes, but I'm faced with a bit of a challenge because this episode is sponsored in part by A Cloud Guru. They're the company that's sort of famous for teaching the world to cloud, and it's very, very hard to come up with anything meaningfully insulting about them. So, I'm not really going to try. They've recently improved their platform significantly, and it brings both the benefits of A Cloud Guru that we all know and love as well as the recently acquired Linux Academy together. That means that there's now an effective, hands-on, and comprehensive skills development platform for AWS, Azure, Google Cloud, and beyond. Yes, ‘and beyond’ is doing a lot of heavy lifting right there in that sentence. They have a bunch of new courses and labs that are available. For my purposes, they have a terrific learn by doing experience that you absolutely want to take a look at and they also have business offerings as well under ACG for Business. Check them out. Visit acloudguru.com to learn more. Tell them Corey sent you and wait for them to instinctively flinch. That's acloudguru.com.

Quinn: Welcome to Screaming in the Cloud, I'm Corey Quinn. I am joined for a third time by Corey Sanders, corporate vice president—which Twitter tells me all three of those words are bad—at Microsoft. Corey, welcome to the show.

Sanders: Thank you. It's great to be here. And it's great to be on with a common-named host.

Quinn: Absolutely. Whatever we say, Corey, it's great: it makes it really easy on the people doing the transcription. Just put, Corey, and that solves the problem neatly. So, this is your third time on the show, second time as the corporate vice president for Microsoft Solutions. Before that, you were deep in the weeds of Azure itself and now have gone into a bit of a broader remit.

Sanders: That's right. That's right, yeah. So, I am responsible for enabling customers, helping them deliver solutions across our Azure solution—so that's certainly the infrastructure side, which is again, as you said, my bread and butter, as it were, and also the data side—and then expanding out to our business application. So, Dynamics and our Power Platform, and security and, and Modern Work. So, that would be Teams and Office 365. And so, running the full range of capabilities for customers.

Quinn: So, it's always fun to compare episodes with the same guest next to each other. It started off as first, “Oh, wow, this Azure thing. What's it about?” Then last time, we had this conversation about Build. Now, of course, world changes, and we're talking about a global pandemic. I'm hoping next time we don't talk about the meteor, but you know, we have these hopes on this. On slightly happier news, you apparently have a new child.

Sanders: I do, yeah. I have a young child here, actually born at the very beginning of the pandemic, so she has really only seen me and my wife without masks on. And so I sometimes wonder what the result of that will be, that we're the only two faces that she's actually seen top to bottom since she's been born. So, it's kind of an interesting psychological experiment, which is typically not a good thing to run on your daughter, but we don't really have a choice, I guess.

Quinn: This is something I hadn't considered yet. I have a kid due myself, in a couple months. So, this is going to be an interesting experiment, myself.

Sanders: We can compare results.

Quinn: Exactly. It feels like something we really should have gotten an ethics sign—off from someone on first.

Sanders: That's right. That's right. [laughs].

Quinn: Let's talk a little bit about what you folks are seeing in the context of COVID, what it's doing to the business. Which again, even saying it like that feels like a very cavalier way of addressing a global crisis, but there are very few companies that are in Microsoft's position, whereas you have cloud offerings, you have communications offerings, you're sort of across the software stack. Of any company, you folks are in the best position from my perspective to get a holistic view of what customers are seeing what customers are doing during these times. What have you noticed?

Sanders: Yeah, absolutely. And to start off, obviously, as you mentioned, just the entire—the impact, the sickness, the challenges, the social implications that we've seen have just been very, very difficult. And so I'm hoping that we can get through this as fast as possible, and continue to make improvements every week. As part of the effort and sort of response, our biggest focus has really been around, how do we help customers in this time? And that runs a pretty wide range of capabilities, solutions, and expectations.

And so you've mentioned a few. A good example is Teams and enabling customers to be able to work in a remote environment. And what I think's interesting is that I think that the result of this is certainly customers learning, and understanding and better appreciating the needs and the capabilities to be able to work remotely, but also, I think, fundamentally changing the way people work from now on. I think the expectations of being able to work, I think the success that customers have seen, when leveraging Teams to be able to work in this remote way, again, has enabled them to approach their entire work culture in a different way.

But it doesn't just end with Teams. I think the need for Remote Desktop, and being able to do secure and protected work, but without necessarily having to have everything loaded on a local laptop that may not have the same level of security control that a customer may look for, and then you get into the broader range of security, just that customers needed to reevaluate a lot of their security principles and reassess the way in which they were approaching their security environment, the amount of VPNs [laughs] that ended up failing in the process of this change has been quite numerous. Where customers were pin-pricking everything through a VPN that was outside of their corporate environment, well, suddenly, when everything's outside of your corporate environment, the VPN struggles. And so, we—

Quinn: [unintelligible] for 10 to 100 users, and now we have 10,000 on it, and it turns out that TCP now terminates on the floor, and we have a problem.

Sanders: [laughs]. You got it. Exactly right. I mean, it's just the ability to scale, the ability to handle that. And then when you think of Teams and running something like Teams, all of it through a VPN device, it becomes mind-boggling just how hard and challenging that becomes to a network. And so there's just changes across the board, Corey, just in how people are thinking about it, and responding, and all of it was to try and enable them to be able to work in this new environment.

Quinn: Tell me a little bit about Teams. I've used it a few times myself, and the sharp edges that I've had with it, to be very honest with you, feels like it has more to do with my understanding and my contextualization of how these things work. I'm an old grumpy Unix administrator—because it's not like there's a second kind of Unix administrator—and I come from an IRC world where everything is just a text chat and response. Threading, I find it offensive. I felt that gifs add nothing to the conversation and make things worse. I'm a grumpy old man standing on my porch shaking my fist.

When did Teams come out of it? It feels like it's this weird hybrid between SharePoint, between Microsoft’s somewhat document-centric approach, and then a chat system bolted on top of that. But again, this is from an outsider who's used it for all of three days in the course of my life. It's very obvious I do not have a good holistic view of this.

Sanders: Yeah, I mean, I think that you've captured it pretty well, which is that it ends up being a single consistent collaboration environment that allows customers to bring together in a single pane of glass, all of the collaborative work experiences that they expect to have. And so the interesting thing that we've seen, especially in this sort of everyone's remote environment, is the ability to both have a meeting, whether it's scheduled or impromptu, to be able to be chatting as part of that conversation, and then being able to sharing, and editing, and modifying docs all together in one experience, it's actually quite powerful from a productivity perspective. And we've had customers even come back and say they've seen their productivity go up in this environment because these are all in a single experience. And certainly with gifs, obviously, that increases productivity. I’m, you know, astounded—

Quinn: Oh, you’re a soft g person on gif.

Sanders: Oh, I know, I know. You know what? It's funny. I should have prepped on this one.

Quinn: What is it with cloud providers and pronouncing things badly? I don't know what it is.

Sanders: I think it's an East Coast thing, actually. I'm going to blame New Jersey as the pronunciation challenge here.

Quinn: Kid, I don't think you get to pull that excuse.

Sanders: [laughs]. My team has yelled at me about this. It's so funny that I just recently had this fight, and I went online and went searching for it, you can find all kinds of articles going both ways, so I'm going to leave it at that. ‘animated pictures.’ how about this? How about I say it that way?

Quinn: We will accept animated pictures.

Sanders: [laughs]. So, the collaboration’s just been pretty amazing. And a great example of this is I've now been in meetings where conversations proceed, people have their videos on, they’re chatting, you can see the emotion from people and so on, and then the splinter conversations spin up in chat. And there are times where I'm like, “Oh my gosh, that's kind of annoying.” But then there are times where you take it in, you're like, “Wow. That is a separate conversation, tied in with the main conversation,” but from people who maybe weren't comfortable raising this because the conversation was ongoing, or people were dominating in the main thread.

And so in some ways, the part that I love about that is that I feel like it's opening up different avenues of collaborating all at the same time. And sometimes those chats are then brought back into the main conversation. Sometimes they just close there as follow-ups. But either way, the person was heard in a way that I think, would actually have been lost in an in-person meeting, if you can believe it. And so I'm actually pretty excited about the way in which those collaborative experiences can happen. Sharing documents live, just all of those aspects are just huge.

And now we're starting to see, Corey, people building as a platform on top of it. So, it's no longer even just our solutions. It's no longer just sharing on SharePoint, and Word docs, and Excel. But we're starting to see partners and even customers deploy their own experiences on top of it to enable those collaborative solutions. So, now you've got Azure apps, you've got Power Apps deployed and exposed through Teams as the single pane of glass. It's secure, it's enabled securely on people's phones and laptops and that's now how people get work done all as a single experience. So, I'm pretty bullish on it. I think it's the new way people work, and people who have fully embraced it, I find that they've actually found new ways to be productive.

Quinn: Okay, I'm going to challenge you slightly on that.

Sanders: Do it. Gifs.

Quinn: Exactly. I am not—to be very clear—a team's user myself, other than a couple of strange edge cases. But when I work with Slack a fair bit, everyone talks about the apps that you can build integrations, and then I scratch beneath the surface and they are fundamentally two different types of things. The first is a notification from something else. I think calling that an app is a little bit of a lofty descriptor. But the more advanced version—“Oh, now we're in the future. You can click a button in that notification and make something happen.” And I feel I have now captured the sum totality of the integrations and apps. Teams, are these still early days? Does it go more fully-featured than that? What's the story?

Sanders: Well, so think of it this way. I'll give you one example, perhaps. You're in a meeting and you're tracking action items. So, you're in a meeting, you're chatting, you're tracking action items, and the ability to easily say, “Hey, I want all the people in this meeting to have access to this tracker. We're going to capture action items. We're going to assign it to people in this meeting,”

So, we already have the scope of who is going to be assigned to what. And we can see it happening live. So, while you're in the meeting, you can see they’ll pop up and say, “Hey, you've been assigned this action item.” And so it becomes, again, a very collaborative, engaged experience. So, that's one example, perhaps where again, these apps can become an integral part of the workforce.

The other aspect that's been interesting, and I've had a couple customers say this, which is we have some of these Power Apps that have gone out. So, one is a crisis management app, one that we're working on right now is a return back to work app that we're seeing customers deploy, and helps you understand what buildings are closed, or opened and safe, and et cetera. And the feedback that we've heard is, “Look, we've gotten Teams installed on every single one of our customer’s phones, we've gotten Teams installed on every single one of their laptops. We don't want to go install yet another app. We've already done the work, we've secured it, we want this to just be an experience through the Teams app, and expose through it, you can install through it.”

And then of course it gives you, as you said, notifications through it, chatbots, et cetera, so it's all integrated. And so, I think both of those are kind of the primary value props that you have. One is just, it's a single pane of glass and so you don't need to install yet another thing. You can secure it, you can build the environment through it. And then, two, there are actually apps that are very integrated with the collaboration experience.

Now they're different. That Return to Work app is not something you'll pull up in a meeting and work together on. It's something that's just a part of your environment. But I think both are pretty relevant in how customers are looking at this new work pane as it were, I don't know if I've convinced you. Have I convinced you? Are you going to start saying, gif?

Quinn: Well, I don't know about—I’d go that far. I mean, I still have principles and standards here, and so I gotta say, there are sponsorship packages available, but I don't know if there is a high enough tier 1 for start changing pronunciations of words on me. I will say that it seems like it ties into something you've mentioned a few times: Power Apps. And Power Apps are interesting to me, mostly because I only discovered them about a week or two ago when a certain competitor of yours launched a no-code/low-code solution. I'm like, wow, this is kind of amazing.

Nice to see companies getting into this, and everyone else looks at me like I'm a fool. Well, Power Apps have been around for a long time. Oracle's Apex has been around about the same length of time. It turns out that no, no, it's only new and exciting when Amazon releases things. Other than that it's just boring and crappy—which is the narrative, and it turns out that's completely untrue. So, tell me a little bit about Power Apps for those of us who have lived in a world of Infrastructure as a Service for a long time. It feels like it's something from another universe of the Microsoft ecosystem. Tell me more.

Sanders: And explain it so that it's not boring and crappy? Is that sort of the—that's the starting point that I think I've got here.

Quinn: Crappy version? Instead of having you on, I would have invited one of the many glossy brochures that you folks put out on things.

Sanders: [laughs]. Got it. Okay. So, look, I mean, I think the principle of a low-code solution is pretty clear. I think what we've done with Power Apps, the way to think about it is it's sort of the combination of PowerPoint and Excel. Where you've got the PowerPoint experience around building apps and UI, and, like, for anyone who's built a pretty complex application, sometimes the UI can be some of the hardest things to get right, get placed, get organized, and so on.

So, the ability to have PowerPoint as this experience of controlling your UI. But then the key thing is, is that it's got the Excel-like experience for bringing the low-code part of it. Writing the formulas, writing the ways in which you want the UI to interact with the end-user, and then layer on top of that the full extent of data sources that could be pulled, to be brought into it, whether it be Excel, whether it be SharePoint, but then also, whether it be Salesforce, or Dropbox, or Twitter, or Facebook. All of these data sources can be brought in in a very simple and easy way. And this is really the secret sauce, I think, with Power Apps is that it's not only that easy app building an easy low-code experience to make a pretty powerful application, but you can bring in all these data sources that then allow you to really expand well beyond the power that we see from some of the competitors to bring in a really comprehensive application.

And so it's a pretty exciting trend that we're seeing. And some of the things that we launched in the response to COVID to help customers get going quickly, Crisis Response app, which is basically we pre-built an app that allows customers to go through how they're going to notify their employees on potential issues, how they're going to communicate out challenges, or risks, or places to avoid from an office perspective, and be able to track their employees, right, in case they needed to respond, or get help. And so all of that was pre-built, and it allowed customers to modify and update in a fairly simple and easy way. I mean, it's just been a huge, huge opportunity for customers to get started quickly and build on top of it.

Corey Quinn: In what you might be forgiven for mistaking for a blast from the past, today I want to talk about New Relic. They seem to be a relatively legacy monitoring company, and I would have agreed with that assessment up until relatively recently. But they did something a little out there: they reworked everything. They went open source, they made it so you can monitor your whole stack in one place and, most notably from my perspective, they simplified their pricing into something that is much more affordable for almost everyone. There's even a free tier with one user and 100 gigs per month, totally free. Check it out at newrelic.com.

Quinn: I will admit that I was something of a skeptic of the entire space because I'd never really had it be something I cared about. But I build my sarcastic newsletter every week through a whole bunch of different Lambda functions tied into various API gateways, and I run it through scripts because front end is something I've never understood. I was finally about to pull the trigger and hire someone to build out a front end for all of this for me, and a buddy of mine who's a terrific developer—now, as it turns out, one of my employees—popped up and said, “Hey, what about Retool?” “Well, what is Retool?” And the answer is it well, effectively it's Visual Basic for APIs and integration.

It speaks to arbitrary APIs, random data sources, but lets you drag and drop an interface into place. Which was a sort of fascinating because it needed a little bit of code, not a tremendous amount, and suddenly it unlocked the ability for me to iterate rapidly without having to spend untold amounts of money on front end folks. And as an added benefit, I did some digging underneath the hood. It turns out it runs on top of Azure. So, yeah, you folks are everywhere at every layer of the stack. And I was very dismissive of the entire space until I started using this to solve a problem. And now it's very hard to get me to shut up about it.

Sanders: I don't believe that about you.

Quinn: Yes, I do have an ongoing love affair with the sound of my own voice; there's a reason why I have a podcast. But it's just such a fascinating approach to me of the idea of it acts as a force multiplier. On some level, the idea of you don't need a developer at all anymore is a bit of a red herring because—

Sanders: I agree with you.

Quinn: —having a developer to work on these things and help get stuff set up. Yeah, but they can drag, drop, get things set up in a few hours, and then go back to the thing that they're normally working on, and that is such an unblocker for business users. And in the context of front end, I am the exact opposite of whatever a developer looks like.

Sanders: Well, I agree on all your points. I mean, I think that to your key point that the statement of, “Hey, you guys with Power Apps, you no longer need developers,” I agree that that's actually an incorrect assessment of the power of the tool. And when you look the combination of Power Apps with some of our Azure application services—let's say API Management or some of our security services, and so on—the combination of bringing together strong developer skills to lay the groundwork, and then the Power Apps experience to be able to extend in easy ways things like those UI experiences, things like those fast and easy tweaks, such that the developer doesn't need to actually do everything. And even if the developer is going to do everything, it's easier to do some of those top-level functions while focusing perhaps on the deeper parts of the platform.

And this is why the integration with data is so key because you can start seeing a world where the hard development work is around creating the experiences and getting the data shaped up in such a way that then the Power App sitting on top, it's just taking advantage of that corporate data that's been built out with the developer and with the data scientist working hand in hand. So, I think that's really the future where we're seeing this. It's in, like you said, an add-on. It's an extension to the power of developers, not a replacement in any way. And this is why we're even seeing integration with things like GitHub, and so on, where governance and bringing these codebases together, the combination of Power Apps with that underlying code and development work that's being done is becoming the expectation from customers.

Quinn: So, tie this into, I guess, another area of Microsoft that’s a giant mental question mark for me: Dynamics. I keep hearing the term, I keep seeing Microsoft folks getting incredibly excited about it. What is it?

Sanders: Well, so Dynamics or D365, is effectively our business applications platform. So, it offers a set of solutions, whether it be solutions around customer engagement, so things like being able to understand who your customers are and help you categorize, segment, and then deliver marketing content to them. It delivers customer support, so it enables you—one of the things that we've seen pick up a lot of steam during this COVID crisis is a capability called Remote Assist, where it allows you to engage with, let's say, someone working on the factory floor from a remote location and help them respond to some incident or some outage, and in fact, the combination of that with HoloLens has become a really interesting, powerful solution where now you can directly be told through the HoloLens experience, how to go respond to something on the ground. But then it's also finance and operations. So, being able to support areas like commerce, so we've seen an outpouring of eCommerce space solutions as an example, and curbside pickup solutions. So, I think that these are the areas and some of the power that we're seeing with Dynamics and Dynamics 365 specifically, with our cloud-based solutions.

Quinn: I have to ask, you call the D365. And among its other failings, 2020 is, of course, a leap year. Are you planning to shut the whole thing down for a day so you don't have to rechange the name to 366?

Sanders: [laughs]. You know, at this point, I'm going to say I think we'll probably keep it running the full 366, but I would need to check back with that engineering team and just make sure.

Quinn: Exactly. One wouldn't want to over-promise availability.

Sanders: I can’t. Yeah, exactly, exactly. I don't want to. I know on the Azure front, we've done a lot of work around leap day because I remember leading a bunch of that work. But yeah, let me go back and check on the Dynamics front.

Quinn: Other things that are interesting and possibly related to you possibly not. Let's talk about Xbox. Is that something that you deal with at all yourself? Is that viewed as a completely separate division? How do you see it?

Sanders: So, it is definitely a separate division. It is not explicitly in my purview, per se, but obviously, we work closely with them, one because we work with gaming customers out there, and certainly they run quite a bit on Azure. So, we have a bunch of understanding and engagement there. Seeing a crazy—I mean, I think they've seen, like, a 50 percent increase in multiplayer gameplay since the crisis which, I guess for many of us who do play games, it's not surprising necessarily, but that's resulted in, of course, a bunch of growth on the platform and then partnership with some of the other gaming companies out there as we look at continuing to support this growth. It's a pretty exciting field overall. But the actual business, if you ask me about specifics on games, and when they're coming out, and what titles look like, and so on and so forth, those I would probably not be either capable or willing to answer. Let me put it that way.

Quinn: Honestly until they get around to remaking TIE Fighter, the best game ever created, I don't care about games.

Sanders: Oh, we've got so much to talk about now about TIE Fighter. Oh my gosh. So, often, I constantly say to my team, “Mission critical craft under attack,” and they never understand what I'm talking about.

Quinn: Best game ever. All these gaming companies wasting our time rather than remaking TIE fighter. I don't understand it.

Sanders: How far to the Emperor’s circle did you get? Did you get all the way in?

Quinn: All the way.

Sanders: Me too.

Quinn: Some people had friends in the ’90s. I didn't have that problem at all. I had TIE fighter.

Sanders: Well, and here's the key. I used to play with my brother. He used to fly and I used to be watching the monitor to see when red dots were popping up behind him. And so this tag team was good. I considered myself the force and he was the actual pilot. So, that was the way I made myself feel better that I wasn't actually playing.

Quinn: All power to shields.

Sanders: [laughs]. Indeed.

Quinn: Got to play those games with people.

Sanders: Oh, man, it was a fun game. It was much more fun than X-Wing, by the way. That was, uh, yeah.

Quinn: There needs to be at least a spiritual successor if nothing else.

Sanders: I agree with you. I agree with you. Anyway. Okay. Back to other topics at hand.

Quinn: Yes.

Quinn: Talk to me a bit about Windows Virtual Desktops.

Sanders: Yeah. So, this capability—it's funny because it's a recently launched capability on the platform, and as I mentioned earlier, we've seen strong momentum, actually we’ve seen a lot with financial services, a little bit with manufacturing, retail as well. So, a lot of that momentum has been around being able to host full Windows client-based experiences in the cloud, and the key thing has been for a lot of customers, the multi-session support, so you can end up really utilizing the hardware in a much more optimized way than on some of our competitors. And it allows, of course, cost savings around it. And so we've seen a lot of use of this, people actually using it to run some of their M365 basic capabilities, or office experiences, even Teams through their Windows Virtual Desktop experience to enable people to have that single pane of glass, to log into and have a zero-trust environment on their local machine. The other nice thing is we've got great partnerships with both VMware and Citrix. For customers who have those management experiences that they'd like to continue, they can deploy onto a Windows Virtual Desktop and enable VMware and Citrix as part of it.

Quinn: Every time in the past, I've tried to look into the world of getting Windows Virtual Desktops or something like that up and running. It was always A) extraordinarily enterprise-y when all I really needed was a Windows machine to run, I don’t know, the proper version of Excel or it wound up going down this rabbit hole of licensing for Terminal Services and the rest. Is that still the case? Admittedly, it's been 15 years since I played in this space with any serious attention.

Sanders: Yeah, I think you'd find it easier. I will admit, I think it's definitely skewed toward solving enterprise problems, although I guess how we define enterprise runs a full range of spectrum. But we've seen, actually, quite a bit of uptake in small businesses as well leveraging it for a single experience to log in no matter where you are. And so sometimes when you've got small businesses that are on the move, it's a fairly easy thing to get set up, and you can deploy and launch into it. So, I do think it's a lot better, I think—certainly, the licensing story is actually a lot cleaner as well since it's all tied into the Azure consumption motion.

So, if you can understand how VMs work and are built, then you can understand this as well. So, I do think that's been a much greater improvement, the licensing story is definitely. You don't have to worry about the RDS and the server hosting and blah, blah, blah, blah, blah. You can dig in and just get going. So, yeah, I’d give it a try and let me know. Reach out, tell me if you feel like we still got work to do.

Quinn: Don't offer if you're not serious.

Sanders: Oh, I’ll listen. I'm not necessarily going to commit to solving your—if you give me problems that other people say I will solve them. If it's just your problems, Corey, we'll have to have a conversation.

Quinn: Sounds good. To be clear, at the time of this recording, it's somewhat open-ended as far as how it's ultimately going to shake out, but I do want to talk to you for a minute about JEDI, specifically, the Department of Defense contract that you folks won, which I think was something that not a lot of folks saw coming, and I'll admit—to be very honest with you—when I saw that, it recast how I was considering Azure in a competitive light, in that, okay, there is something here that I am not seeing historically. And again, given that I tend to specialize in born-in-the-cloud, cloud-native companies, my side project Twitter for Pets has almost dozens of customers. Azure was never something that we considered because that was always for big enterprises and not aimed at the technological capabilities of where the world was going. That is provably untrue now that you—given the access that you wound up competing on and winning.

Sanders: Yeah, absolutely. I mean, look, I would argue we believed this on our side for a long time. And we've got a good range of large enterprise customers that we've won over the last couple years. But certainly, the announcement with JEDI has been exciting. And I think it's a combination of things.

I think, certainly the platform and the support on the platform, but I think it's also the deep integration with security, I talked a little bit about it already, just the identity support that we offer that spans the services, the growing networking capabilities and security capabilities that we have built into the platform, and then certainly our hybrid story. I think our hybrid story is really just amazingly strong. And this is something that we just hear from customers all over the place, whether it be manufacturing, whether it be retail, the ability to take this split world where you've got computation that needs to run local, it needs to run right near those end customers, those end experiences, those end actions that are happening, and then being able to use the Cloud for the scale motions, for the broad data analytics, the predictive expectations, and so on. And this seamless capability and platform, taking that hybrid story, being able to run it local, whether it be with Azure Stack Hub, or leveraging something like Azure Arc, to be able to create this experience that spans both with the same—back to my governance, and identity, and security conversation, it creates this really nice fluid opportunity that I think is quite unique in the market. And that's, I think, certainly a big part of the conversation and certainly something we hear from customers no matter what the industry, but particularly in government.

Quinn: It's definitely recast my understanding of the entire cloud landscape. At this point, it's become pretty clear, especially with some of the larger enterprise deals that I have been working on with my existing consulting customers, that Azure is very much in play. Really, I've got to say, it still remains—I know we talked about this every time, but it definitely remains one of the business school case studies of the future there's going to be highlighted. Just a complete cultural and perception turnaround in the past decade. It's really something to see. To be blunt I counted Microsoft is down and out, and on a long decline into irrelevance back into naughts and early 2010s. That feels like an incredibly naïve and out of date perspective now.

Sanders: [laughs]. Yeah, I mean, I should hope so. But yeah, I mean, I think that's right. I mean, look, probably the strongest point that I'd make on that regard is the focus that we've had over the last few years, certainly bringing the platform into a competitive place, and now, I would argue, in many places exceeding our competitors. But I think the key point is, and we've talked about it, and I've weaved in a few points around Dynamics, and around Power Platform and around Teams, and certainly on Azure, it's all about helping customers get to that next level of transformation.

This is something that we are just seeing all over the place, and I gave those examples around Edge and hybrid with, like, a retail store. They need to rethink how they engage with their customers. They need to rethink how they are selling to their customers. And being able to bring together a solution like a new eCommerce platform to sell remotely, combined with something running on the Edge, that's an application to be able to bring insight and knowledge around the customers that are shopping locally, and having that all come together into a single picture. This is just commonplace now for a retailer, it needs to be. But that's a big shift for a lot of customers.

And so this is where I think when you look at the overall spectrum of engagement that we have with customers, it's really around helping customers get to that next level. How are we enabling and supporting them to grow, build their business, and achieve that next level of opportunity for their end customers? That's really where I think we, with Azure, and the progress and growth that we've made there, with hybrid, as I mentioned, with IoT, with Dynamics, with Teams, it's all around that principle. And I think that that's really resonated with customers and something that I also think is pretty unique.

Quinn: I would wholeheartedly agree. If people want to learn more about what it is you have to say, what you're working on, how you view the world, where can they find you? The easy answer is microsoft.com, but I'm wondering if there's another place?

Sanders: Me particularly, you're saying?

Quinn: Oh, well ideally, yes.

Sanders: [laughs]. Not just Microsoft. Yeah, the best places to find me, I’m relatively active on Twitter, I probably should be more active on Twitter, and I've started to pick up a little bit on LinkedIn as well. So, I think those are probably the best places. I try and go through Twitter comments pretty regularly. It's tough though. It's tough. I've been reading your Twitter pretty religiously, of course. So, those are probably the best places. In fact, it's probably easier to get me on Twitter than it is in my email, just given the scale of email that we do.

Quinn: Oh, yes, I can well imagine. Corey, thank you so much, once again, for taking a third half-hour out of your life to speak with me. As always, it's appreciated. Thank you.

Sanders: You bet. Thank you. It's been fun.

Quinn: And enjoy the rest of your parental leave. When that comes up. I will. I’m looking forward to it.

Quinn: Corey Sanders, corporate vice president at Microsoft, I am Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts, whereas if you hated it, please leave a five-star review on Apple podcasts anyway, along with a comment correcting other Corey on the proper pronunciation of GifHub.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

Links Referenced:

  • Honeycomb: https://www.honeycomb.io/
  • Personal Blog: https://charity.wtf/
  • Honeycomb Blog: https://www.honeycomb.io/blog

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored in part by Catchpoint look, 80% of performance and availability issues don't occur within your application code in your data center itself. It occurs well outside those boundaries. So it's difficult to understand what's actually happening. What Catchpoint does is makes it easier for enterprises to detect, identify, and of course validate how reachable their application is. And of course, how happy their users are. It helps you get visible and to reach a bit availability, performance, reliability, of course, absorbency. Cause we'll throw that one in too. And it's used by a bunch of interns and companies you may have heard of like, you know, Google, Verizon, Oracle, but don't hold that against them. And many more. To learn more, visit www.catchpoint.com and tell them Cory sent you wait for the wince.

Corey: This episode is brought to you by Trend Micro Cloud One™. A security services platform for organizations building in the Cloud. I know you're thinking that that's a mouthful because it is, but what's easier to say? “I'm glad we have Trend Micro Cloud One™, a security services platform for organizations building in the Cloud,” or, “Hey, bad news. It's going to be a few more weeks. I kind of forgot about that security thing.” I thought so. Trend Micro Cloud One™ is an automated, flexible all-in-one solution that protects your workflows and containers with cloud-native security. Identify and resolve security issues earlier in the pipeline, and access your cloud environments sooner, with full visibility, so you can get back to what you do best, which is generally building great applications. Discover Trend Micro Cloud One™ a security services platform for organizations building in the Cloud. Whew. At trendmicro.com/screaming.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by someone who needs no introduction, so I'm going to give her as little of one as possible. Charity Majors founder and CTO of Honeycomb, but famous on the internet long before that. Welcome to the show.

Charity: Thank you, it's nice to be here. We've been trying to do this for a while.

Corey: We have, and it seems it's scheduling never quite works out until suddenly everyone is trapped at home, and eventually we both run out of excuses.

Charity: Yes. Well done.

Corey: So, what are you up to these days?

Charity: Well, I feel like I've turned into a den mother for a bunch of very anxious camp-goers. Through this whole COVID thing, I have always hated repeating myself, and that's something that I feel like once you go into management, you kind of have to get over. But during the pandemic, repeating myself, honestly, has come to feel like my one and only job.

Corey: In what sense?

Charity: “Everything's going to be okay. You're doing everything we could have asked of you. You're doing all right. Everything you're doing is enough. You are enough. You are beautiful. Please go home, get some sleep, have some soup. Take care of yourself. Take care of your loved ones. Yes, you're working hard enough. Yes, you are doing your best. Your best is enough.” You know, that kind of thing.

I feel like something that I kind of said off the cuff in our all-hands a few weeks ago was just that brain weasels love authority figures and they love reassurances from authority figures. And it's so true. We're adults, we can be as anti-authoritarian as we want to be, it's still—there's something very reassuring—when we're very anxious about everything, there's something reassuring about somebody in a position of authority just telling you that you're good enough and everything's okay. And for the first time in my life, I find myself awkwardly in that position of authority. And so, just trying to use it to reassure people as much as possible that what they're doing is good, they are fine and that they can take a break, seems to be my main job right now.

Corey: Management seems to be one of those interesting areas of expertise because, to be blunt, a lot of the people who write these books and tell you how to be a good manager—

Charity: I hope “expertise” was said with scare quotes.

Corey: Oh, absolutely it was because all of these people you read, “Oh, you're going to give me advice on how to manage? Great, what do you do?” “Oh, you haven't managed anyone in 20 years. Instead, you've been writing books on how to manage people.”

Charity: Yes. Or, or better yet, they are a VP of something-something, and then you hear around the grapevine that they are terrible at their job.

Corey: Exactly. I can think of several people off the cuff that we are not naming because—

Charity: No, no, of course not.

Corey: —eh, lawsuits are passé.

Charity: Yes.

Corey: But you're right. I see something similar to this when I was getting started in freelancer circles, where people are focusing on specific areas that they want to tackle. It’s, “Oh, I'm going to learn how to do positioning as a freelancer,” and it turns into, “Huh, it seems I can make more money teaching other freelancers how to do positioning,” for example, “Or how to do sales than doing those things myself.”

Charity: I wish that people were legally mandated to present context with their advice. This advice is true within the narrow circumstances of these things that I've experienced before. Because people sound very authoritative. They're like, “You should do this.” And they don't bother to mention that they come from some completely different context, or we're talking to completely different people. Context is everything when it comes to advice.

Corey: Absolutely. I mean, one of the things I've stayed away from for a long time has been management advice, simply because my only management training was having a bunch of really terrible managers, and then doing the exact opposite—

Charity: Mm-hm.

Corey: —which it turns out gets you further than you might expect. But then I started this thing, and when I took on a business partner, Mike, I stopped managing anyone because what I do requires me to be more than a little bit self-promotional, and that is the recipe for a terrible manager in my experience. You have to build other people up, but I have to be the loud obnoxious one all the time.

Charity: And you hate that, don't you?

Corey: I can do it, but it's one of those things where I can't do it while simultaneously lifting up others.

Charity: You do it and you're very good at it.

Corey: Well, thank you, I appreciate that. I am the loudest and most obnoxious out of the entire flock.

Charity: Hearts.

Corey: Exactly. So, one of the questions I see is given that you have been able to walk a tighter line than I have because you still manage teams and you tend to focus on the empathy piece quite a bit, but you also have never been one to shy away from sharing your opinion online in various ways that I hold a deep personal affinity for. Some might call them bombastic; I call them enthusiastic.

Charity: Sure. We only speak in grand statements. Yeah, well, I mean, part of it is I don't really manage a team right now. I manage three people: Liz, Emily, and Megan. They've graduated, they don't need my bull-[BLEEP], you know?

So, I honestly feel like if we're talking to engineering managers here—context, right?—I would say that your job is to craft a team. Your job is to craft and coach a team. You're not hiring individuals, you're not hiring people. You're curating a team. And the kind of just emotional intensity that you need to have with that team, you're in it with them every day. And I'm not doing that right now.

And so most of the, when I'm giving my quote-unquote, “advice,” it's mostly me thinking back to the days where I was doing that and drawing on my experiences there. That plus I don't really like to talk about the people that I’m working with now. It feels kind of rude and kind of non-consensual to show our dirty laundry. I'll talk about them someday, but right now, I honestly try not to talk very much about people at Honeycomb.

Corey: Well, you also have entered one of those very interesting and rarefied positions where the people who report to you are themselves senior enough and have the context to be able to need guidance in a very different way, versus junior people or folks who are new to the workforce in one form or another, often don't have that. And the way to manage those people, in my experience, has been radically different.

Charity: Yeah, because I don't really have to think about the careers of the people who report to me much. I trust them for that. They're in a good spot. It's much less of a here, help them grow in their skill set to get to senior engineer, and then help them be effective in a large [unintelligible], and it’s much more just, they're grown. They're baked. They're cool, so I need to just give them the information that they need.

But managing managers is very different than managing engineers or managing ICs of any sort. The visual that I always get is, when you're doing the work, you're assembling a basket with your hands. When you're managing people, you're walking up and down watching people assemble baskets, and you're inspecting, and peering. When you're managing managers, you're standing over on the other side of the pier, squinting and trying to see how the baskets are coming out, and then waving at them with flags to try and signal what they should be telling their team to do. You're so far removed from the actual being able to do the thing that it's terrifying, honestly.

It's terrifying and it's a whole different skill set that involves a lot of letting go of the ego. We spend our whole career trying to get good at this skill set, right? Being good at engineering, being good at getting things done, making things, and then we get drafted into management, and it's not even, like, a sidestep career-wise, it's like you're starting over at zero. But you still get to sit there and enviously watch the people doing the thing that you love to do. And you're trying to wall that part of yourself off, and I don't know, it's weird, man. I don't understand why people want to be managers.

Corey: My honest assessment of why people want management is that they're told that that's the thing you do next. It's what when you've conquered the IC scale, then you get promoted to management, but it's a parallel track. It's not something in most jobs that should be considered a promotion.

Charity: Or even more perniciously, they feel disempowered. You know, what did Daniel Pink say that we want out of life? We want autonomy, mastery, and meaning. That’s what we want from our work. So, many people show up to work every day, and maybe mastery, maybe they're getting mastery, but they may or may not have meaning, but autonomy, they don't have it. And they want it. They crave it because we all do.

And they're told, or they intuit for themselves, that the way that you get more control over what you're doing, you should go into management. And this is so… I mean, this is why I went into management. I wanted a bigger say over what I was doing. No shame. I'm not sharing to shame anyone who's done any of these things.

I think that even wanting more power, I think that it’s better to be honest about these impulses within ourselves and deal with them in a straightforward way than it is to learn to puppet that, “Well, I just want to help other people, or—” it's better to be honest with yourself about why you want something because I don't think there's a wrong reason.

Corey: I would like more money and the ability to tell people what to do.

Charity: You know, that's fine. That said, you're going to be in for a rude awakening when you realize that managers don't actually get to tell people what to do, at least not in healthy organizations.

Corey: Oh yes, but what made you think that most organizations are healthy ones?

Charity: Oh, there's the rub. And so there's always this back and forth between why do you give people advice for how to be a manager, but it very quickly segues into how healthy is your organization? How healthy can you make it? How many other people are there who feel like you? Because here's the thing, your company doesn't exist: you show up every day and create it with your colleagues. There's nothing that exists other than that. We have all of the power that we need in our hands to change things radically if you just convince people that it's worth doing, and enough of the people show up and create it this different way tomorrow, right?

Corey: Absolutely.

Charity: Which is a terrifying and awesome thing to realize.

Corey: That's quite the thought to wrap my head around, but you're right.

Charity: And so many organizations where the only way to proceed is to go into management. Well, what if there could be parallel tracks? What if there could be technical tracks that go hand in hand with managerial tracks. But the way that you need to sustain this is it can't just be managers giving power to engineers. It has to be engineers embodying that power and taking it for themselves, too. Have you ever been on teams where managers want to give engineers more power and they won't take it? Because I have.

Corey: It's fascinating to watch that dynamic play out and how the organization responds to that tells you an awful lot about that company.

Charity: It really does. Now, in my mind's eye, I have this ideal, the way that engineering organizations should operate. This is a creative profession. Yes, there's a high basic skill-level for entry, but we create things, and managing a large codebase is fundamentally a creative act, which means that we're going to do our best work when we are empowered, and inspired, and motivated. So, I feel like power given to managers should be limited and enumerated. Less King George and more articles of the Constitution, where it's like, okay, we're not supposed to bow and scrape to our leaders anymore.

They serve us, and we are giving them these articulated powers, and they have authority and power and responsibility to do these limited things insofar as they are doing them on behalf of the company as a whole. I think that's the appropriate healthy attitude towards management because there is this drift, this power will tend to drift towards management over time unless you're consciously pushing it back out to the ICs. Because so much of power is information flow, and managers, they are absolutely privy to private information, and they spend their days talking to people instead of heads down in code. So, power is going to drift towards them unless you actively push it out.

Corey: And something you said just there's there does resonate in that we have a high basic skill level required for these jobs, but if you look at a lot of historical companies, by the time that someone is coming in making, say, six figures at a manufacturing plant, they've been there for 20 years, they’ve worked their way up. Now, a new grad gets that and a lot of assumptions in a lot of these cultures is baked in that by the time you're paying someone a certain amount of money they, for example, have the insight, wisdom, and political savvy not to call their boss an asshole in the all-hands meeting. That does not always play out the way that one would hope.

Charity: No, it doesn't. Something I've been thinking about a lot lately is the way we humans, we’re hierarchical beings. If there's a ladder there, we instantly want to climb it. And there's this woman, Molly, who I worked with, she was an engineer, she graduated, she got offered a manager's job pretty quickly, she climbed a ladder, she had a series of high profile executive jobs.

Twenty years she worked this industry before she suddenly realized she hated it. She was miserable. All she wanted to do is write code, and she was jealous every day of the people that she worked with who were working software engineers. She's like, “Why did I climb this ladder for 20 years? I hate everything about this.” And I find that so sad. And there's so much pathos to that because we have this assumption—just like you said—that if there's a hierarchy than the person at the top of the hierarchy, is superior to all who are below them, in knowledge, in judgment, in experience, and all these things when in fact, I can now say from experience, it's just a different kind of work.

Yes, when you have an organization of a certain size—it's like a human data structure, basically, that you need for an org chart because you need some people to have their heads up scanning the horizon and thinking about that one to two-year, five-year journey, and you need people down in the mid-range with their heads at that 6- to 12-month range, and you need people down in the weeds, but they're just different kinds of work. They're not better or worse than each other. And some of us are made happy by some types of this work, and not other types, and if we could just strip it of all of the baggage of power and hierarchy, if Molly could just work on the things that brought her joy, this is the amazing thing that we take for granted this industry. We get to work on things that bring us joy, every day. This in and of itself, sets us above 99 percent of all humans who have ever lived, and yet we throw away this joy and meaning that we found to chase the being better than the people around us. And I feel like happiness should be given a greater weight.

Corey: In what you might be forgiven for mistaking for a blast from the past, today I want to talk about New Relic. They seem to be a relatively legacy monitoring company, and I would have agreed with that assessment up until relatively recently. But they did something a little out there: they reworked everything. They went open source, they made it so you can monitor your whole stack in one place and, most notably from my perspective, they simplified their pricing into something that is much more affordable for almost everyone. There's even a free tier with one user and 100 gigs per month, totally free. Check it out at newrelic.com.

Corey: You make excellent points, and I think it's time for a second rant on a different topic if you'll indulge me yet further. Namely, this for a lot of people relatively new to the workforce has been their first downturn. The last one was 10 years ago and change, and there have been a lot of people entering the workforce since then. What are you seeing as far as the downturn’s impact on Cloud and cloud-like services?

Charity: Yeah. So, a little bit of background. I love running mail servers.

Corey: I started my career that way, and I'm right there with you. And then it—ehh, Postfix was great.

Charity: Right? At Linden Lab, I was the only one who could get IMAP working. We ran our own trainer and spam assassin, and ClamAV, like antivirus—

Corey: Dovecot, or one of the older ones like Courier?

Charity: Yeah. All that.

Corey: Excellent.

Charity: The first job I had in industry was actually writing QMail spam filters, the first spam filters in QMail.

Corey: Ohh, you were on that side of the world.

Charity: Yeah. So, I love that [BLEEP], and when it was—oh, [BLEEP], how many years ago? A lot of years ago. I was still a teenager—and they wanted to move our mail to Google, and I was like, “No.” And I wrote this long and passionate email about why it was terrible to outsource our email to Google. And fundamentally, I still wanted to be able to grep my mail spool. Anyway, TL;DR we moved to Google, I quickly saw the light, and the world has been outsourcing more and more ever since.

Corey: Right. Personally, I finally shut off my mail server in my rack and moved it to Google, and suddenly I had hours a week back that I was spending playing whack-a-mole with spammers.

Charity: I know, it's amazing. When you free up time from the optional things to work on the things that actually matter to you, it's a good trend, I'm all for it. So, the last time we had a downturn, 10 years ago, we saw a lot more of this. And the way I feel about it is, God knows I have tried—I see people out there going, “I’m going to hire an observability team, and they're going to build it from scratch for us.” And I'm just like, “Oh, you sweet summer child.”

But because we are not very good at quantifying the costs in human terms, this keeps happening. It’s kind of like a vanity project, almost, where people would just be like—you know, they want to have more teams underneath them, or they want to feel like their workload is so special, and so one of a kind, and so bespoke, and so they need to do all these things themselves. And they don't think critically about the cost involved. Well, during downturns, all of that changes, and all of a sudden, people are looking at the fact that they have a hiring freeze for the foreseeable future, if not letting some percentage of their team go, and they're still going to be responsible for the same amount of work, if not more, and suddenly the value of focusing on their core business initiatives—the things that are differentiators for them as business—becomes manifestly clear. And the argument for outsourcing major parts of their business—like, email is still very core to everyone's business, right?—becomes very, very compelling.

And so these great waves happen of retrenchment, where people move things to the Cloud, move things to other providers, and then when the money comes back, they don't get reversed because it was always a stronger argument all along. It's just that we don't really know how to assign actual dollars and numbers to it, so it's still a very manager-driven instinct, intuition-driven process of allocating people to effort. But what we've seen at Honeycomb is very much this: that a lot of companies were, kind of, in the—like, large, mid-size tech companies large enough to maybe start to preen and go, “Oh, we're so special,” but actually not—those people were planning on going and hiring teams to do observability, and instead, they're coming to us. And they're especially coming to us because we're very much associated with the next way of doing things, and new things get a lot of heat—sometimes deservedly. God knows I've [BLEEP] all over serverless enough—but one thing they do tend to be is cheaper.

Corey: By a lot.

Charity: By a lot. Orders of magnitude.

Corey: Oh, yeah, the most expensive thing at any company never shows up in the AWS bill, and that is the engineering time gone into wrangling these services that show up on your AWS bill.

Charity: Yes, thank you. Absolutely true. And the newer ways of doing things are cheaper from a financial perspective. And it's almost like, we don't have to think about that while times are good. We think about all the other things—because think about all of the factors that go into making a technical or architectural decision. There's so many. What does your team like? What languages do you know? What stuff do you have provisioned? What other stuff do you already have written? All this stuff, and it can become very muddled. And in downturns, the pure cost of doing business, and writing code, and shipping value to customers, stands out.

Corey: And that's what it comes down to is value to customers, on some level. Most of us have this insane thing beaten into our skulls, where—we talked about learning these things, and how we start out going down these paths. We’re generally hobbyists playing around, and our time is free; computers cost money. That's no longer true at all in any professional context.

Charity: No. But because we love what we do, it still feels free to us.

Corey: Exactly. So, “Oh, why would I pay extra for that managed service? I can build it myself on top of EC2.” And then people set out to do it.

Charity: “And I enjoy it. And I know all the ins and outs,” and blah, blah, blah, blah, blah. Yeah, exactly.

Corey: Right. And half the time, I come in and I see an environment like that, and I understand intrinsically why it is the way that it is, and how it came to be that way, but it's not a good idea. And trying to convince folks of this is sort of an uphill battle because people have to learn this on their own.

Charity: I mean, yes, and no. I'll give the old guard a little bit of credit in that the costs of learning a new system can be quite high. I do think these need to be, kind of, step functions. I think it's fine to invest in technology, and a architecture model, and a set of technologies for a while until they start to—you know, you kind of want to standardize on a golden path until they start to show their age, which could be three years, right? [laughs]. But every three to five years, you're going to need to renew and refresh it radically from scratch.

Corey: And that's an awful lot of…

Charity: It's kind of the ‘optimize locally versus optimize globally’ tension.

Corey: Yeah. And the problem, of course, is I look at this, how in the world, am I going to be able to intelligently advise companies to do this on a one-off basis when I see this pattern happening again, and again, and again, everywhere? It feels in some cases like I'm shouting into the wind.

Charity: Yeah, well, I think you have been because they've had more money than sense. And now for a time they won't.

Corey: And we've always been waiting for the next correction. I don’t think we expected it to look quite like this. But now it's a question of, “Okay, if we have to reduce headcount or stall on hiring, which things can we transition to a managed provider that yes, will cost us more in raw infrastructure terms, but reduce the expensive engineering time by 10x?”

Charity: Yes. Some reporter was asking me, “Are people going to be backing away from the Cloud?” I'm like, “No,” [laughs]. “No, the sticker cost it is, like, a 10th of the actual cost.” Most of the cost is an engineering time and an opportunity time: you not doing the things that could have made your business succeed because you're doing the [BLEEP] infrastructure.

Corey: Exactly. “Will people be backing away from the Cloud?” “Well, not smart, people.”

Charity: [laughs]. Not smart people. Exactly. And serverless is an order of magnitude or more cheaper than even containers. Now, you can't use that for everything, and maybe it didn't make sense for a long time to rewrite parts of your infra in that, but we wrote our database in serverless, basically. It’s S3 files and Lambda functions, and it fell in cost by an order of magnitude.

Corey: And that is something I'm finding that’s—it—just people look at cost through the wrong lens. They're worried about lock-in. Well, great. So, you're going to build this thing—

Charity: It's easy to get hire.

Corey: Oh, yeah.

Charity: It’s easier to get headcount. So.

Corey: And finally, we're starting to see that change a bit—I wish it were happier circumstances—but we see Kubernetes, too, where it feels like all these people—like, that's been sort of the refuge for engineers who want to build their own cloud provider, but don't actually work at a cloud provider, so here we go.

Charity: I have been calling Kubernetes resume-driven development because nobody needs it but all the engineers think they need it on their resume in order to get jobs. So…

Corey: And if you're working on Kubernetes all the time, who's minding the store?

Charity: I know. What the [BLEEP] are you even doing? You know, funny story. Actually got leaked a couple of slide decks from startups: founders who are out there pitching basically Honeycomb for Kubernetes, or observability for Kubernetes, which is missing the point in such an enormous way because the point is supposed to be that it's boring. The thing that serves the code is supposed to be as boring as possible. It's inside your code and inside the systems that is supposed to be interesting.

Corey: Exactly. If you get excited about the things that Honeycomb does, for example, past a certain point, you probably should apply to work there.

Charity: Yeah. Well, lots of people do. But anyway.

Corey: [laughs]. You should not try and roll it yourself from first principles inside of your own company.

Charity: Yeah.

Corey: There’s, like, five companies in the world that might want to do something like that intelligently, and the rest are, again, trying to reinvent things from first principles that they best should not.

Charity: Sadly, all the people who are out there hiring, quote-unquote “observability teams” are hiring to staff them with the last generation of time series database folks who are just like, “Oh, this is a nail? I have a hammer.” And I think they're going to be not too happy with the results.

Yeah, I don't know. I think that some downward pressure is—you know, creativity loves constraints. And I feel like for too long in engineering, we haven't had constraints in terms of the resources. And I don't want to minimize the amount of pain that's out there right now, or people losing their jobs. That sucks a lot, and I feel for you.

Engineers, I think you're going to be okay. I mean, I think the other part of this downturn is folks are going to distributed teams in a real way, which is exciting. Engineers are going to be able to get jobs for the most part. And I think it's good to have some downward pressure on costs. But every crisis is an opportunity, and I think that's it companies that have learned these lessons well, and have not forgotten them between downturns have been the companies that you really look to who are doing things well, who are doing things in ways that are sustainable, who don't have to make many corrections when these things hit.

Corey: And I hope that that winds up, ideally, breeding a culture that's a bit more aligned around this going forward.

Charity: I would hope so.

Corey: I don't want to go back to the way things were and seeing these ridiculous things that are pitched with absolutely no business model.

Charity: I know. I know. At some point engineers, we have to come into the fold as a part of the business. Engineering exists to serve the business. We exist to build things for users. We don't exist for our own entertainment. There are plenty of ways you can write code for fun if that's what you were trying to do, but at work, we have a responsibility to not just take these things into account but to do better at learning how to verbalize them, how to spread them as our values, how to make them something that is not just passed along as an oral tradition the way I feel like it is right now, but actually, professional standard.

Corey: Wouldn't that be the dream for all of us?

Charity: We're moving in that direction. It's just a question of how quickly.

Corey: We certainly seem to be. So, thank you so much for taking the time to chat with me about this. If people care more about what you have to say, where can they find you?

Charity: Oh, gosh, I have no idea. [laughs].

Corey: Right?

Charity: I'm all over the internet. You can find me, my personal blog is at charity.wtf. Honeycomb is at honeycomb.io. Our blog on the Honeycomb site has a lot of observability stuff. And basically I'm talking [BLEEP] all over, so.

Corey: Which is delightful and entertaining. If you enjoy the theme of my nonsense, I definitely recommend you check out Charity, should you not be fortunate enough to be already aware of it.

Charity: We have similar nonsense themes, it is true.

Corey: Well, thank you once again, I really do appreciate it.

Charity: Thanks, Corey. It's been fun.

Corey: Of course. Charity Majors, CTO and co-founder of Honeycomb. I am Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts, whereas if you've hated it, please leave a five-star review on Apple Podcasts along with a comment telling me which system you're building yourself from first principles.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Stephanie StimacStephanie is a Design Technologist and Program Manager for Microsoft Edge Developer Experiences. She comes from a background in design and after initially spending 6 years focusing on a career in web design, has spent the last 4 years working on Microsoft Edge to improve developer tools and the browser. Currently she helps run an initiative called the Web We Want that focuses on identifying problems developers face in their day-to-day work and is passionate about HTML, CSS and inspiring a new generation to get involved in the web.

Links Referenced

  • Microsoft company website: https://developer.microsoft.com/en-us/microsoft-edge/
  • Corey’s websites https://gaslighting.me/ and https://stop.lying.cloud/
  • The Web We Want: https://webwewant.fyi/
  • Smashing Conference: https://www.smashingmagazine.com/events/
  • beyond tellerrand: https://beyondtellerrand.com/
  • An Event Apart: https://aneventapart.com/
  • webhint: https://webhint.io/
  • Stephanie’s Twitter: https://twitter.com/seaotta

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is brought to you by Trend Micro Cloud One™. A security services platform for organizations building in the Cloud. I know you're thinking that that's a mouthful because it is, but what's easier to say? “I'm glad we have Trend Micro Cloud One™, a security services platform for organizations building in the Cloud,” or, “Hey, bad news. It's going to be a few more weeks. I kind of forgot about that security thing.” I thought so. Trend Micro Cloud One™ is an automated, flexible all-in-one solution that protects your workflows and containers with cloud-native security. Identify and resolve security issues earlier in the pipeline, and access your cloud environments sooner, with full visibility, so you can get back to what you do best, which is generally building great applications. Discover Trend Micro Cloud One™ a security services platform for organizations building in the Cloud. Whew. At trendmicro.com/screaming.

Corey: Normally, I like to snark about the various sponsors that sponsor these episodes, but I'm faced with a bit of a challenge because this episode is sponsored in part by A Cloud Guru. They're the company that's sort of famous for teaching the world to cloud, and it's very, very hard to come up with anything meaningfully insulting about them. So, I'm not really going to try. They've recently improved their platform significantly, and it brings both the benefits of A Cloud Guru that we all know and love as well as the recently acquired Linux Academy together. That means that there's now an effective, hands-on, and comprehensive skills development platform for AWS, Azure, Google Cloud, and beyond. Yes, ‘and beyond’ is doing a lot of heavy lifting right there in that sentence. They have a bunch of new courses and labs that are available. For my purposes, they have a terrific learn by doing experience that you absolutely want to take a look at and they also have business offerings as well under ACG for Business. Check them out. Visit acloudguru.com to learn more. Tell them Corey sent you and wait for them to instinctively flinch. That's acloudguru.com.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Stephanie Stimac, who is currently a Microsoft Edge developer experiences program manager, which almost but not quite fits in a tweet. Stephanie, welcome to the show.

Stephanie: Hi, Corey, thanks for having me.

Corey: So, who are you and what do you do exactly?

Stephanie: Currently, I am a program manager for Microsoft Edge developer experiences. And so my journey has been a little bit different. So, I've been at Microsoft for four years, and it's only been within the last year or so that I've transitioned into actual program management work. And on the developer experiences team, what that means is I am looking to engage with developers and find out what makes working on the web hard. What obstacles are you encountering, and what's preventing you from also using Edge or switching to Edge? And really trying to dig into their problems and make things easier for them.

Corey: So, in the interest of full disclosure, I, many, many moons ago, had a foot firmly in the Windows world. This was back in the Internet Explorer 6 days, which should explain entirely too much about why I'm angry all the time. And in those days, there was Internet Explorer for Mac, which was eventually killed. And since I wound up moving over to the Mac universe and the Linux universe, I haven't really paid much attention to the goings ons of Microsoft.

Then they had their sort of Cloud-focused renaissance that changed the focus to, “Oh yeah, you actually do get to care about Microsoft again.” And people started talking to me about Edge. And when I asked what that is, the answer that basically can be distilled down to, “Oh, it's their web browser.” And whenever I asked, “Oh, so it's a rebranding of Internet Explorer?” Everyone got super angry and stopped speaking to me. So, what is Microsoft Edge for those who followed, I guess, a similar track to my ridiculous one?

Stephanie: Yeah, let's go back. So, there was Internet Explorer, and that was running on a browser rendering engine called Trident. And, let's see, I want to say six or seven years ago, a new version of Edge was released. And that was the Edge HTML rendering engine, and that was actually based on Internet Explorer's rendering engine. So, it was still Trident, but it had just gotten rid of all the old legacy code and all that legacy baggage that made Internet Explorer a little bit difficult to work with. And so that was Edge HTML.

And I don't want to say it was a rebranded Internet Explorer because it wasn't quite, but it was a new browser. And then within the last year, we ended up switching our rendering engine to Chromium because one thing that we were having a hard time with was just keeping up with all the new features that developers wanted in a modern browser. And when we adopted Chromium, what that also meant was—so the old Edge was tied to the Windows operating system and only got updated, like, every six months. And that just wasn't great for developers when you're waiting for a fix in something in your browser, and your website's not rendering right. You can't wait six months for that.

And so, with Chromium, that only caught us up with all the modern web features that developers had been asking for, but it also allowed us to change our update cadence. So, now we have four different channels, so we have our Dev channel, our Canary channel and our Beta channel. So, Canary gets updated every day, and then we have a stable release that goes out every six weeks. So, we're not tied to Windows anymore.

Corey: Pardon my ignorance because I tend not to play in these waters too much, but it feels to me like originally, it really, really mattered what browser you were using for anything in particular. And this was, of course, back in the early nascent days of web apps, where you had a whole bunch of fully-featured applications running on a computer, but you had websites, and then you would pull up things that were basically your local newspaper, and it would say things like, this website works best in Internet Explorer four, or Netscape Navigator two-dot-whatever-it-was, and you sort of stare at that and look at it for the longest time and it felt like it was a very prescriptive approach. These days it seems that whenever I encounter a website that doesn't work super well on a particular browser, that's an aberration and it is certainly not the norm to the point where it is almost tweet-worthy whenever that happens. So, would you agree with that statement?

Stephanie: Oh, absolutely. We're in an era now where there are so many different tools available to developers to test different browsers to make sure that things are working. So, yeah, there's no reason a website should not be working in a browser today. Something should not only be working in Firefox or Chrome. The only reason I could see that happening that would be an acceptable answer is if maybe you're trying some new experimental web tech, and just letting people know this may not work because I'm using this new thing that maybe hasn't been standardized and isn't really common yet.

In that case, I can see it being acceptable. But so yeah, there are some companies that internally still rely on old technology like ActiveX, or just things that aren't supported anymore, and so you can get away with that more if it's internal and not external facing because for a lot of these companies, it's really, really expensive to try and migrate to a new system and move all these old legacy apps off this legacy tech. And so, again, in that situation I'd say that's acceptable. But otherwise, if you've got a website that you just didn't take the time to test in another browser, I don't think that's super acceptable.

Corey: I will occasionally grant affordances for this. For example, the web app that I use to record this podcast with folks who are not in the room—which is basically everyone now that that becomes a deadly risk—only claims to support specific versions of Chrome, the end, and I sort of understand that given the weird intricacies it has with making sure it does high fidelity recordings on both sides, setting up a VoIP call, having the right permissions model. I don't like it, but I tolerate it for the recording piece. I completely failed to forgive them entirely for the fact that I can't modify my account, check billing, change plans, book recordings, et cetera, also unless I'm using Chrome. That's just inexcusable.

But the fact that it at least only biases for one browser for a very specific app that requires that much systems integration, cool, I'll allow it, angrily. But I had no tolerance for this in a world of my bank demands I run a specific browser. Well, sounds to me, at least in this decade, like I need a new bank rather than I need a new browser. And it does seem that the release of Edge, in many ways, goes along with that in a similar timeline, with Microsoft itself reinventing itself for a cloud era. It went from a company that, to be honest, I despised into a company that I've come to begrudgingly admire on a whole bunch of different axes.

And using Chromium, for example, rather than its own custom internal rendering engine that it decides is better than everyone else's and they shove it down people's throats, it felt like for a while—and please don't take this the wrong way—that its biggest weakness was the word Microsoft at the beginning. Even now, that is not even a concern given some of the, frankly, stellar moves that we've seen Microsoft undertaking in the past five to six years.

Stephanie: Oh, absolutely. I think one of the cool things about Microsoft right now is just all the teams that I work with, and their willingness to really be an open-source, and work on those projects. And that's one of the cool things about working with the Chromium project because it is open-source, and all these conversations that used to happen—before my time, so I'm making an assumption here—but I would assume all these conversations occurred in secret and private and trying to come up with new competitive features that would further your market share and then sort of lead to gaps within the web platform and cause those interoperability issues and now we're working so closely with the Chromium team, a lot of these discussions happen in the open and they happen where people can participate, not just people on browser teams, but where the public can see it. And I think that is really cool to watch browsers sort of work in this way and advance the web platform by listening to the community and not just kind of deciding, oh, this is what we're going to do without talking to anyone.

Corey: So, forgive the, I guess, sheer ignorance that is packed into this question. I am not a front end person; if you ever seen any of my code that is not front end, you can understand why that is. It's bad there; it's worse when it actually impacts users. So, I always thought that Edge was not going to be something that I could ever participate in because when I am on the road—remember back when we used to go places? It was great, maybe someday we'll do it again—the only computer I ever took with me was my iPad.

But just checking this in preparation for this show, it turns out that yeah, Edge is available for iOS. Now, in my naive idiot corner of the world, I thought that every browser that you could spin up on iOS was forced to use Safari under the hood. So, it just effectively is a different window dressing for the exact same browser. Is my understanding dramatically misunderstanding something? Am I mostly on the money or something else entirely?

Stephanie: I'm going to have to fact check this, but yes, I believe here mostly on the money. Yeah. We might have some more, like, Microsoft-y privacy features, but I am pretty sure that yeah, it is just WebKit under the hood.

Corey: The fact that I can sit here and say, “Well, I don't like using Chrome and I'll use Edge instead because I prefer the—from the privacy story is way better with Microsoft than it is with Google.” The fact that I can say that and it's not sarcastic, would have absolutely blown my hair back 15 years ago to hear me say—I feel like I want to—like, time-traveling angsty me from my childhood who wants to travel forward and just slap me so hard the candy comes out. But you're right. Everything you say is absolutely aligned. I would have no problem running Microsoft Edge in a way that I struggle mightily with the idea of running Chrome. In the interest of full disclosure. I switched a couple years ago over to Firefox, now that it seems to not be eating RAM for breakfast anymore, preferring instead to leave RAM-devouring to Slack.

Stephanie: Yes. Can I just give a shout out to Firefox really quick because I have a great set of developer tools that… I still do front end code in my spare time, and they just have some great front end tools that aren't rivaled in any of the other browsers right now. So, I just want to give a shout out to Firefox because they're doing some great work there.

Corey: I went through the terrible mistake, I suppose, of doing some front end manipulation stuff for a couple Lambda@Edge functions I use on the AWS side. For example, I took a Chrome extension, with the help of people who are good at things, and shove that into a Lambda@Edge function because I could never remember the URL for the AWS status page. And it always lies and tells you things are fine even when they're not, but now I have a quick URL that points to that called gaslighting.me, or stop.lying.cloud, depending upon your tastes.

And the latter of those winds up doing a whole bunch of dynamic JavaScript manipulation to remove a whole bunch of the ‘all is well’ green field of dots and actually tell you what's broken from that, which is super handy. But looking into how this works, and getting that up and running was an exercise in—I don't think frustration goes far enough. The whole idea of asynchronous callbacks in JavaScript? That was a complete head-scratcher for me. Wait, that piece is further down the page than that, why is it loading before the slow thing up above? And the more I worked with it, the more I realized that, A) I have not kept up with the current state of technology, and, B) front end is very clearly not for me. But you're right, Firefox did make it somewhat easier to wind up troubleshooting those things once I got back to a computer. Doing this troubleshooting front end from an iPad: pro tip, don’t.

Stephanie: [laughs]. You know, one thing I want to say about staying up to date with all the technology out there, someone actually messaged me the other day, and they're trying to become a developer, and they were just struggling with, “What do I focus on? And how do I just stay up to date? And how do you stay up to date?” And one thing I've learned in working on the web platform and just seen the breadth of the web and all the different niche areas is there is no way to be an expert in all of it.

You really have to pick your area that you love. I admittedly do not know that much JavaScript. I can go StackOverflow a couple things, or find some CodePens and probably do what I need to do, but HTML and CSS are my bread and butter, and I have a team of people who know JavaScript: they’re JavaScript wizards. So, accepting that, yeah, that's not my area, but accessibility and HTML and CSS, that's still valuable and that's my niche area. I think developers can get a little bit caught up in… there's always a new framework, and they're always arguing over a framework, and there's always a lot of noise on Twitter about that. But I think you just need to use the tools that you find useful, and just focus on that and what you enjoy doing.

Corey: And that's, I think, part of the issue is it's too easy to wind up being angry and opinionated about things that, A) don't really matter that much, and B) even if they do, it's about others people's choices in ways that don't directly impact you. I care about a number of things when logging into a bank's website, for example, but which framework they've chosen has never been anywhere near the top 500 items on that list. It just doesn't matter all that much from a customer-user-facing perspective. If you're working on something and building it yourself, sure, I can see making those arguments, but past a certain point, I just don't care.

Stephanie: Absolutely. I think one thing we do have to be mindful of, and there was just a whole discussions about React this last week and performance, and one thing that I try to be mindful of now, especially—and Alex Russell has kind of—I've heard him talk so much, I've kind of been drilled into my brain about using frameworks or even just building something from scratch, caring about performance and making sure that you're testing on low-end devices for people who don't have access to a brand new iPhone and fast devices, that is still something I think more developers need to pay attention to. And I don't know how much of picking a framework actually affects performance, but I do think it's something developers just need to be aware of when they are building.

Corey: In what you might be forgiven for mistaking for a blast from the past, today I want to talk about New Relic. They seem to be a relatively legacy monitoring company, and I would have agreed with that assessment up until relatively recently. But they did something a little out there: they reworked everything. They went open source, they made it so you can monitor your whole stack in one place and, most notably from my perspective, they simplified their pricing into something that is much more affordable for almost everyone. There's even a free tier with one user and 100 gigs per month, totally free. Check it out at newrelic.com.

Corey: So, one thing I'm curious about is you've transitioned from working in design to working in program management aimed at developer experience. First, that feels like a strange transition. Can you tell me a little bit about that?

Stephanie: Yeah. So, I went to university for design. My major was Digital Media Design, and I really fell in love with the web, but I never ever considered myself a true web developer. I could code some basic HTML and CSS when I graduated, but I was also sort of that era that learned how to use Flash to build websites, and so the only code that I needed to learn was ActionScript to make things animate and provide interactivity that way. And right after I graduated, Flash basically died as a [laughs] web design medium—

Corey: Yes it did, along with the batteries that it killed.

Stephanie: [laughs]. Yeah. And so I was pretty proficient in this technology that was suddenly not relevant at all. And I had a couple jobs where I was a graphic designer and did some print work. And that was fine, but the thing that I really loved to do was, like, work on wireframes for web apps, and sort of design web experiences.

And so my first job, I ended up at a startup in Bellevue and ended up designing the user interface for their 2.0 release of their app and I really, really enjoyed that. And I was only there about a year because it was a startup, so pay wasn't that great, and hours were not maintainable. And I ended up at a communications agency as a designer, and this is where I sort of fell in love with the web, and really gained all the skills that I had for building websites. And Microsoft was a huge client of this agency, as they are for a lot of agencies in the Seattle area.

But after a couple years, I became the digital expert at this agency, and so I was responsible for every aspect of the web design process. I would do the user research, and then I would do the wireframes, and go through the client review with that, and the information architecture, and then I’d do the visual design, and then I would code it, and then if we needed some deeper functionality that I didn't know how to do, we would bring on a true developer. And I was at that agency for three and a half years, and I got a message on Twitter—of all places—from a PM on the Microsoft Edge team who was looking for, he was looking for a PM but I was going to be doing design work. And so I didn't want to turn that opportunity down at all. I was looking for something new because it was just kind of time to move on from that agency.

And I ended up getting hired onto the Edge team, and for the first three years, a lot of my job was kind of like what I did at the agency, but with developers being my customer. And so I had my coworker, Melanie Richards, who's still on Edge and has also sort of made the same transition as me, we were both the designers for the web platform team and focused on building things for developers, and we would do projects around our developer portal and design code demos for new web platform tech that would only be an Edge. And then Melanie transitioned into a traditional PM role, and I was really struggling with making that transition for a while because there was some fear in my mind that if I became a PM, I wouldn't be a designer anymore. And that's actually not the case; I'm still a designer. So, about a year ago last May, I made the transition to not designing anymore, and have been in this PM role where I focus a lot on finding those developer problems. So, my main focus has been an initiative called The Web We Want that I run with my coworker Aaron Gustafson, and that has sort of taken up all my time.

Corey: Is it a conference? Is it a website? Is it something else entirely?

Stephanie: So, yeah, it’s something else entirely. There's sort of two components to The Web We Want. So, The Web We Want is a open platform to gather feedback from developers about problems they encounter on the web. Or maybe not even problems, just feature gaps that they've been struggling with and they just think there should be a native solution for. And so the question that we asked developers to answer, “Is if you could wave a magic wand and change anything about the web platform or dev tools, but would it be?”

And that all gets posted up onto our website, which is webwewant.fyi. And the thing I love about this initiative is it's not Edge specific. So, you're not just giving feedback for Edge. You're giving feedback for the whole of the web platform. So, we're working with people on Chrome, Firefox, Igalia, and Samsung Internet, and I think, maybe one other partner, I can't remember. But all these people are looking at what is being submitted to this website, and we're starting to look at these things and assess how to move forward with some of them.

But the cool thing about The Web We Want is there's this whole online component, but there's a focus on the web community; it's about what they want. And so we've partnered with Smashing Conference and beyond tellerrand and An Event Apart. And we actually run a forty-five- to hour-long session where people who have submitted their ideas to The Web We Want actually have a chance to present either in person, or if they can't attend in person they can do a screen recording, and it's them just pitching their idea in a quick three to five-minute lightning talk, and walking through this problem or feature that they've encountered on the web and why they think we should go fix it. And so, it's been really fun to engage with the web community that way, and I love running that session at events because you really get to see how passionate people are about the web and what they're doing. I've seen some detailed case studies—I mean, detailed for a five-minute lightning talk—about this problem, and they get into the nitty-gritty, and it's really inspiring to see so much passion. And so that is The Web We Want, and that's been taking up my time, and I absolutely love it.

Corey: That brings us to one more topic I wanted to cover with you where historically sessions at events and whatnot—it seems sort of antiquated, now that we're all locked down—to be clear, at the time of this recording it's the very end of April—so given how quickly these events tend to outpace ourselves, if you're listening to this wondering why we didn't talk, I don't know, about the giant meteor that is now bearing down on us—who knows what we'll be dealing with, that is why. But as of this time, there's a pandemic on, and we aren't going to events anymore. You've been killing it on Twitter with the live streams of bartending every day, of teaching us a different quarantini recipe. That's amazing, and I'm wondering how much of that ties back to your stage and speaking engagements.

Stephanie: So, yeah. I'll just give a little bit of backstory about my relationship with public speaking because two—oh, I guess it's almost three years ago now, I was supposed to co-present with a co-worker on the Edge team about a tool called webhint that I worked on. And I was absolutely terrified. I never—I hated public speaking. I hated talking in meetings, just someone who never wanted to stand in front of people and give a presentation.

And so that all changed about a year ago. I was down in San Francisco for Smashing Conference; the Edge team had a booth and I was just there to sort of promote, download our new Chromium browser. And I was at the speaker dinner and ended up talking to a couple of the speakers about, “Oh, yeah. No, I've never spoken but this call for proposal came through, and the conference theme is about luck and how has luck sort of played a part in your career.” And they encouraged me to submit the talk, and I felt somewhat heartened by the fact that they had told me that, “Yeah, if you don't get nervous before you go on and talk, that's not something to be proud of. Everyone gets a little bit nervous.”

Corey: Right, if you're not nervous, it probably means you're about to give a pretty crappy talk.

Stephanie: Right. And so I was like, okay, I have all these world-class people in my industry and telling me, “Yeah, do it.” And so, I ended up submitting the CFP, and a week or so later, I got an email and was accepted to the conference. So, I was like, “All right, this is happening.” I spent, oh, three or four months practicing my talk every single day because I wanted to make sure that I was prepared. And it was my first conference; it was in Scotland; I ended up giving the talk in front of 200 people and afterwards I was like, “Oh, I'm still alive.” And so—

Corey: You get high on the adrenaline. You're thrilled, you submit for other ones, and then oh, no, the whole process repeats.

Stephanie: Yeah, exactly. And during this time, I had actually been running these Web We Want sessions at conferences, so I was getting used to standing up in front of people. It was only 20 or 30 people sometimes, but doing the live streams kind of helps me prepare a little bit for public speaking down the road. I was supposed to give, like, six talks this spring, but because of the pandemic, they all got moved. But before the pandemic, I gave two talks, and kind of got that out of my system.

And I almost feel like my quarantine cocktail hour is me sort of getting that out of my system, and still practicing talking in some way because generally I'm not very good on the fly. I want to know what I'm saying. And so it's been fun to do those every day, and not only just have a fancy drink, but there's all sorts of interesting history around different cocktails, and so it's fun to go research all the origins of these things and the different alcohol and whatnot. So.

Corey: It's definitely worth tuning into if you haven't already. Hopefully, by the time this airs, we'll be back out going into real bars, but in the event that they're not, where can people find you for this?

Stephanie: So, I usually stream on Periscope, which is connected to my Twitter account. So, I'm seaotta, S-E-A-O-T-T-A, on Twitter. I don't have a set time, but if you don't want to catch the live streams, I've also uploaded most things to my YouTube channel, so I think if you just search ‘Stephanie Stimec quarantine cocktails.’ It'll come up. And the whole thing about my episodes of quarantine cocktails is I'm in quarantine; usually I have just done a workout and I'm coming straight over to my home bar to make a drink and just kind of like a little bit of a hot mess. I'm not scientific about it. Occasionally we'll pour too much alcohol into one of the drinks and then be like, okay, well, I guess we're improvising, and so it's kind of a hot mess express situation, but it's a lot of fun. [laughs].

Corey: Excellent. [laughs]. It's nice to bring a little personality to these things.

Stephanie: Oh, absolutely.

Corey: It's hard to see sometimes. But it winds up at least telling stories and reminding folks that there's a human at the other end of the line. And that's increasingly difficult when we can't actually go out and see said human.

Stephanie: Yes, absolutely.

Corey: Stephanie, thank you so much for taking the time to speak with me today. I appreciate it.

Stephanie: Thank you for having me, Corey.

Corey: Of course. Stephanie Stimac, design technologist and program manager for Microsoft Edge developer experiences. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts, whereas if you've hated this podcast, please leave a five-star review on Apple Podcasts, and then a lengthy diatribe ranty comment about which favorite JavaScript framework you have.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Mark Nunnikhoven
Is this system safe? Is my information protected? These are hard questions to answer. Mark Nunnikhoven works to make cybersecurity and privacy easier to understand.

A forensic scientist and security leader, Mark has spent more than 20 years helping to defend private and public systems from cybercriminals, hackers, and nation states. A sought after speaker, writer, and technology pundit, his message is simple: secure and private systems are a requirement in today’s world, not a luxury.

Links Referenced:

  • Trend Micro: https://trendmicro.com/
  • Twitter: https://twitter.com/marknca
  • Mark’s website: https://markn.ca/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored by a personal favorite: Retool. Retool allows you to build fully functional tools for your business in hours, not days or weeks. No front end frameworks to figure out or access controls to manage; just ship the tools that will move your business forward fast. Okay, let's talk about what this really is. It's Visual Basic for interfaces. Say I needed a tool to, I don't know, assemble a whole bunch of links into a weekly sarcastic newsletter that I send to everyone. I can drag various components onto a canvas: buttons, checkboxes, tables, etc. Then I can wire all of those things up to queries with all kinds of different parameters, post, get, put, delete, etc. It all connects to virtually every database natively, or you can do what I did, and build a whole crap ton of lambda functions, shove them behind some API’s gateway and use that instead. It speaks MySQL, Postgres, Dynamo—not Route 53 in a notable oversight; but nothing's perfect. Any given component then lets me tell it which query to run when I invoke it. Then it lets me wire up all of those disparate APIs into sensible interfaces. And I don't know frontend; that's the most important part here: Retool is transformational for those of us who aren't front end types. It unlocks a capability I didn't have until I found this product. I honestly haven't been this enthusiastic about a tool for a long time. Sure they're sponsoring this, but I'm also a customer and a super happy one at that. Learn more and try it for free at retool.com/lastweekinaws. That's retool.com/lastweekinaws, and tell them Corey sent you because they are about to be hearing way more from me.

Corey: This episode is brought to you by Trend Micro Cloud One™. A security services platform for organizations building in the Cloud. I know you're thinking that that's a mouthful because it is, but what's easier to say? “I'm glad we have Trend Micro Cloud One™, a security services platform for organizations building in the Cloud,” or, “Hey, bad news. It's going to be a few more weeks. I kind of forgot about that security thing.” I thought so. Trend Micro Cloud One™ is an automated, flexible all-in-one solution that protects your workflows and containers with cloud-native security. Identify and resolve security issues earlier in the pipeline, and access your cloud environments sooner, with full visibility, so you can get back to what you do best, which is generally building great applications. Discover Trend Micro Cloud One™ a security services platform for organizations building in the Cloud. Whew. At trendmicro.com/screaming.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Mark Nunnikhoven, currently a VP of cloud research at Trend Micro. Mark, welcome to the show.

Mark: Thanks, Corey. Long time listener, first time screamer.

Corey: Excellent. So, you work at Trend Micro, an antivirus company, which seems like a very hip and relevant thing for 1998. Now, it seems like, so, you're effectively you're the equivalent of John McAfee, only they didn't put your name on the company? And yes, I understand that comparing you to John McAfee may very well be nastiest thing anyone has ever said to you in your life?

Mark: Well, while my hair is as crazy as John McAfee is just generally, yeah, Trend definitely has that reputation. The good news is, and this may be eye-opening for the listeners, antivirus like it's 1998, that was old for Trend in 1998. Trend’s been around 31, 32 years now. Started with the early products like PC-cillin for those of you who have been around long enough. And now, anything that needs to be secured Trends got products or research that does it, whether that's in the Cloud—relevant, obviously, to this audience and to you and me—or smart cars, smart factories, anything like that. Really, really cool. But yeah, that company built its power base, its customer base on antivirus, and that’s still a big strong part of the business.

Corey: Which I wouldn't doubt. The problem, though, is that I remember using Trend Micro. It was a solid product. There was a lot of, to be honest, crap in that market, where the systems were worse than the problems that they were designed to prevent against, and I don't talk about this too often, but I used to be a help desk person turned windows admin. And at that point, dealing with antivirus was part and parcel.

Now, credit where due, Trend Micro’s offering was great, but that's not really the marketing angle anyone wants to care about these days because, “Yeah, we were awesome 20 years ago,” is really much more on-brand for IBM than it is for companies that are still, you know, relevant.

Mark: Yeah, fair. Totally fair. And that's why at this point, I would say antivirus is probably the smallest portion of our business. One of the biggest, and growing by leaps and bounds, is actually helping cloud builders secure their deployments in the Cloud, whether that's servers and instances or all the way through to serverless architectures. So, the one thing about cybersecurity is you can't stay resting on your laurels.

If you were good six months ago, that's not even as relevant as what do you have that viable today, and on top of that, the way I'd like to describe Trend, jokes aside, is really, we're a company that has a huge amount of knowledge around cybercrime, computer security in general, the products that the company sells are just a manifestation of that knowledge, and they're going to change constantly because technology is changing constantly. What doesn't change is that history and that continuing quest to keep learning more, to keep pushing more. And that's why, blissfully in my job, I'm on the research side, I don't really have to worry about the products too much.

Corey: Yeah, that's one of the nice things, I think, about working on the research side is, for better or worse, you are a VP. You're not going to get on a podcast like this and then immediately savage the company, nor should you. That says something about a person, none of it good. But I've never gotten a sense from you that you were there to push a narrative or a company line, and we've hung out in passing it an awful lot of various AWS events. Which I think leads me to my next question for you. You by all appearances are a very traditional looking security person. Why do you focus on Cloud?

Mark: Yeah, so I'm going to take that reference to traditional-looking because I've been dealing with the diversity angle when trying to help initiatives to get diversity in tech.

Corey: Well, you do have earrings, at least one of them—

Mark: I do—

Corey: —so—

Mark: —I have two.

Corey: —it’s clear you’re not a culture fit for IBM.

Mark: Sure. Though ironically, I worked for IBM, and at the time when I was a much, much younger man, I had, at that point, a multitude of piercings and tattoos while I was working for IBM in an externally facing role. So, I've toned it down as I've gotten older. But yeah, I'm a very traditional security person. My background and my training is in forensic investigation, I worked for the Canadian federal government for a decade working on nation state-level security stuff.

And if you've said, “How do you get to be a security professional?” I think a lot of the stuff I've gone through and a lot of the certifications and training really fits there. But the reason why I've been focusing on Cloud for the last eight or nine years is because it's a huge opportunity to fix everything that is wrong with cybersecurity and information security, and there is a mountain of things that are wrong with the approach for security. Go no further than talking to any team in a large organization and ask them what they think of their security team, hopefully, you'll get a disclaimer about how they're nice people, but then you're just going to hear the complaints roll out about how they're grumpy, how it's the team that says no to everything, they slow everything down, they make things way more complicated than they need to be. And I get it, I understand how things got to that point because it's one of those baffling things where if you unravel it and look at each decision in turn, they make perfect sense.

Back in the 90s, we used to have strong perimeters around everything, so everyone had a big, big firewall, and we used to use antivirus everywhere because that was the only tool we really had. And add up 30, 40 years of these kind of decisions, and you end up where we are now, where more often than not, security just sucks; it just slows people down when it doesn’t have to be that way. And we're starting to see that now that Cloud is far more mature than it was 10 years ago, you're seeing security really be an enabler, really push things forward—and I hate that word enabler, but it is accurate. Security's there to make sure that whatever you're trying to build does what you intend and only what you intend.

Corey: That's an incredibly nuanced and difficult thing to do in a world of Cloud. Again, this was never easy working in on-prem environments, but at that point, you at least knew, for example, the number of computers you had running in the building, give or take. Now, at the time of Cloud, it's not just understand in some ways. You have a bit of an unbounded growth problem, but you also have the things that are running today could behave in different ways tomorrow, not just in terms of performance, but in terms of capability of, “Hey, the provider has now launched additional aspect of functionality.” A classic security example of this is tag-based access control which you can do what IAM now. And that's awesome and exciting, and it's more than a little confusing, but remember that for 10 years or whatever it's been now, that was not the case.

So, everything was set to be able to set tags with reckless abandon, it wasn't scoped, people were encouraged to do it, and now as soon as you wind up doing anything that's using tag-based access control, you wind up accidentally creating security risks out of every single one of those things that are out there, perhaps unknowing, perhaps not. But it used to be that the worst-case scenario for tags was someone could wind up changing allocation rules or fill your logs with garbage. That was about it. Now, “Oh, wow. There's security problems here.” It's a new world. That feels like one of the dark side sharp edges to the amazing capability story that is Cloud.

Mark: Yeah, absolutely. And that example, I think, is the prime example because it hits on a few things. Permissions are hard for a lot of people. So, one of the things I keep getting frustrated at and I keep talking to the AWS teams about—not to call them out specifically—but in IAM, in the Identity Access Management console, there's a bunch of managed policies, which are great. They help you get up and running, they help you figure out what permissions are linked to what actions, but there's a whole bunch of them that end in the two horrible words, “full access.”

And while these are great for maybe a limited development environment, you should never see them in a production environment because I have yet to really honestly truly hit a case where a role or a user needed full access to a server or a service in production. So, that's a fundamental problem. Now you add tags on top of this, and like you said, for years, we've been adding them and using them primarily to route billing. And now you're saying, “Well, now you can accidentally grant production rights to somebody because the tag is wrong.” That is an absolutely significant challenge, and definitely, a downside.

It actually came up in a discussion I was having with some development teams I was talking to today. They were discussing how they were leveraging the well-architected framework. They're just getting started with it, which is great. They're on their learning journey, but they were discussing it like it was a one-time thing. And they're saying, “Well, this is our architecture, and we've done a well-architected review. And we're good now.” And I brought up your exact point, which is, “Hey, wait a minute. You may be done with your architecture, but that's built in the Cloud, and you're using all of these different services, and those service providers are making changes, which then have implications to your architecture, whether that's security implications or not.”

And that's a huge shift for development in the Cloud, but for security, it's one of those sort of brain breaking things where you go in the traditional model, we used to love to put our arms around it and go, “This is mine, I can protect it. I know where the boundaries are.” And that is fundamentally the history of cybersecurity. “I can put my arms around it, I can protect it. I'm good,” even though that's just sort of a false belief that a lot of people live with to make it easier to sleep at night.

When you're in the Cloud, yeah, you don't know what your environment looks like now versus 20 minutes from now, if you start to see a surge in traffic and things scale. And then the advantage of having all of these new features come out from your cloud provider is that they also potentially do open up these avenues, like with the tag example.

Corey: That is one of the best aspirational descriptions of cloud, the idea that the entire environment is highly dynamic, however—and maybe this is biased by the customers that I tend to work with—very often, that's not the case. I take a look at the big, expensive environments, and it's always stuff that was provisioned in 2016 or earlier. It's a lot less dynamic than people often believe that it's going to be. And that's not, incidentally, just an artifact of large accounts. I've been dealing with a lot of legacy cruft in my AWS account that I launched four years ago, and it's all serverless stuff, but I have a whole bunch of roles, I have lambda functions now using deprecated runtimes, that do ridiculous things I don't understand.

I have a series of escalating challenges in that account where at some point, I want to burn it from orbit and start over, but that's kind of hard to do, and there's still library management problems and the rest. I'm spinning up new accounts constantly, like I work at Wells Fargo. It's awful. That stuff always accumulates in accounts, and fundamentally, no one understands it. I'm the only person that builds stuff in this [era’s] account and I still don't understand what all of the moving parts are. And look at how much time I spend on this stuff.

Mark: Yeah. Hundred percent. The challenge and the balance I always have is ‘nerd me’ is like, “Look at all the possibilities. Look at the amazing, crazy autonomous self-healing deployments that we can make.” This stuff is great. More pragmatic, older ‘I’ve seen some things’ me is exactly what you said. There's legacy, people are slow-moving in general, the vast majority of cloud usage is not the cool stuff.

So, Cloud generally suffers from the same thing that security does from its public image in the respect that if you look at security conferences, 99.9 percent of the material is about the latest zero-day vulnerability and exploit; it's about the latest cool hack. The reality of cybersecurity is that you spend your day worrying about patching threat assessments, a whole bunch of stuff that's never talked about publicly. The same challenge exists in Cloud is that if you look at what happens at conferences, whether they're physical or virtual, all the latest blog posts, a lot of this stuff is like, “Look at this cool stuff you can do on the cutting edge.” It's not, “This is the reality that you have a business app that's working, and there's no way your boss is ever going to let you completely rebuild it from the ground up to make it just do the exact same thing, but in a better way because at the end of the day, it's doing what you need it to do.”

So, the reality is very, very messy. But I think the possibility is there. The biggest challenge is the lack of tooling in general, not just for development, but tooling in general, to help people shift their mental model. And this is one of the unfortunate things that I see every year at re:Invent—not to call re:Invent out, it just happens to be the biggest single source of announcements—is that a lot of those announcements—

Corey: You’re telling me.

Mark: Yeah, exactly. I mean, you probably burned through a keyboard in November alone. And the challenge is, is that a lot of those new service announcements are stuff that are designed to help people bridge that gap and to get more cloud-y in their thinking, but not pushing cloud forward. And it's the Cloud coming back to bring them along, which is a very good thing, but also a lot of the time reinforces some of these older mental models, these older approaches because you're very right, in the Cloud deployments are not nearly as dynamic as they could be, and I think that's not a lack of the Cloud backing or the services not being able to do that, it's the people can't think that way. It's like, if you sit a programmer down and go, “Okay, write me a script that does XYZ.” And then if you tell them to do the same kind of thing across multiple threads, people's brains don't work in multi-threading mode, they work in, sort of, just serial mode, one after the other. And while parallelizing everything would be better for a lot of cases, it's just not how our brains work. And that's the same thing that we see in Cloud.

Corey: We talk about Cloud a fair bit, but let's make it a bit more specific. In addition to your apparently easy, minor day job as a VP at a major company, you also are an AWS Community Hero, which I interpret to mean that you looked at the vast landscape of what causes you could volunteer for, and picked a trillion-dollar company. What's that about? What's it like? What do you do?

Mark: [laughs]. Yeah, so the Hero program’s been going for a number of years now. I've been in it for a while, and it's basically some folks within Amazon, recognize what you're doing out in the community or in a specific technology stack—so there's Serverless Heroes, Machine Learning Heroes, Data Heroes, that type of thing—and it's just an extra recognition from AWS. The advantage for us as Heroes is it gives us some opportunity to speak at AWS events, to publish on AWS properties, but also to get access, and to give feedback to the product teams a little more frequently than the customers normally do. So, it's a nice little recognition; it feels good to see that the efforts for me specifically was based on a lot of speaking that I'm doing, a lot of community outreach to help people understand security, it's a nice recognition there.

But yeah, volunteering for a trillion-dollar company is not a totally incorrect way, but the good news for me is that's in addition to volunteering to run the youth basketball, and scouting, and stuff like that. But yeah, it's a good program. It gives us some of the insights that you get a glimpse of being a prominent influencer, so that when you speak people listen. The Heroes program is a little bit of that endorsement for the company itself.

Corey: That’s, I think, a fair way to say it. One of the areas that I found the Hero program to be incredibly useful from my perspective—again, I am not a Hero for reasons that should be blindingly obvious. I have never been invited to be an AWS Hero because, let's not kid ourselves here, Amazon hires smart people who can read the room, but I've also only ever interacted with folks on the periphery of it. But one of the values that I find coming out of the Hero program is the fact that it helps me contextualize what a release means in a different context. Because the Heroes don't work for Amazon. There have been a few people who are no longer Heroes because they took jobs at AWS. At least one other is no longer a Hero because they took a job at Microsoft, but that's a different problem.

And I see that this is a group of people who are not beholden to AWS for anything other than a bit of publicity, so if AWS does something egregious, they're in a much better position to sound off about it, or highlight things that may not make much sense. At least it feels to me like there's not any constraint on saying things because you don't work there; what are they going to do, take away your birthday? You can say what you think. And through that lens, I find that the stories that come out of the Hero program are a lot more authentic, for one, and also a lot more bounded to use-cases you might find in companies that aren't either Amazon or their specific target customer case studies. Would you agree with that?

Mark: Yeah, I think that's accurate. I think what you're saying there in a very eloquent way is that the Heroes tend to provide a little more perspective. So, I'll give you a very pragmatic or very practical example, every year for re:Invent, registration is a disaster. Trying to get a reserved seat in a session is always frustrating because the third party app goes down, and people who paid good money to go to this conference to see good content, have a really hard time getting into those talks, and that's frustrating for absolutely everybody, AWS included. As part of the Hero program, I think three years ago, they actually briefed us all ahead of time because they said, “Listen, we know you guys are looped in on a bunch of social media posts, and people getting really frustrated. Here's what we've done behind the scenes. Here's what we're trying to do to make things whole.”

And they actually gave us a contact during the week of re:Invent as well to make it easier to answer community people's questions. So, they said, “Look, we know you're going to give the best answer you can if somebody is asking for some help, but here's how you guys can make that simpler for everybody, and here's a direct line into somebody on the events team who can handle that.” But I think for me, that was a good thing in that they didn't tell any of us to stop complaining about broken sessions and registration, they just said, here's how we're trying to help. And that rolls out to services, you'll see that Heroes are not necessarily going to be disparaging the company or taking it down, but we’ll speak our mind, too. So, for me personally, Amazon Neptune, the unheard-of one of many data stores—

Corey: Yes, their giraffe database, named after giraffes, which are of course themselves made-up animals that don't really exist.

Mark: Fair. Totally fair. Because what sound does a giraffe make? Nobody can tell me.

Corey: Exactly.

Mark: So, for Neptune: great service, super cool possibility, and totally wasted for a huge amount of customers because it's traditional spin up an instance and pick out your capacity, no serverless capability in it whatsoever where having a graph database, fully serverless, would have been amazing. And that's a frustration, so I mentioned that when it was launched, I mentioned that a few times before. So, there's definitely that context, and I think the advantage for the Heroes not being employed by AWS, but also having sort of the insider track on some AWS stuff is when new services get released, most of the time, we've either used them already or had a chance to provide feedback and help shape them, so we can help people understand the better context because that continues to be sort of a weak spot. When a service comes out, everybody looks at it through their own lens and goes, “This solves my problem.” And it's like, “Yeah, probably not, but here's the problem it does solve.”

And then it also, like you said, shows those additional use cases. So, from the Heroes, you're going to see here's how you can build serverless applications using Auth0 instead of trying to shoehorn something into Cognito. Here's how you can leverage DynamoDB while doing your compute in GCP. That kind of stuff because we're not restricted, and I think—not to speak for every Hero, but I'll speak for every Hero—we're just trying to help people because we find this stuff interesting, and we just like to help people.

Corey: In what you might be forgiven for mistaking for a blast from the past, today I want to talk about New Relic. They seem to be a relatively legacy monitoring company, and I would have agreed with that assessment up until relatively recently. But they did something a little out there: they reworked everything. They went open source, they made it so you can monitor your whole stack in one place and, most notably from my perspective, they simplified their pricing into something that is much more affordable for almost everyone. There's even a free tier with one user and 100 gigs per month, totally free. Check it out at newrelic.com.

Corey: One of the common threads that I see with the Heroes that I've dealt with has always been a willingness to help others. It's not about, “Look at how awesome I am,” It's, “Let me help lift other people up and teach people things,” and they do this on a volunteer basis, in all seriousness, which makes the naming of it, in typical Amazon fashion, terrible. If you call someone an MVP at Microsoft, that tends to be something that you can self-describe: when you're elected MVP for a game, you can say that; it's on your resume. When you call yourself a Hero though, it always feels weird, at least in the common baseline, English language perspective. It's like ‘entrepreneur.’ it's one of those things other people call you far more than it is something that you call yourself. If you say, “That person's a Hero,” then yeah, you feel good; that person did something great. When someone says, “I'm a Hero.” Oh my God, this person is so full of themselves, I can only assume they're a venture capitalist.

Mark: Yeah, so I'm just thinking in the back of my head. Did I use the word Hero in my own intro when this started? I hope not.

Corey: You did not. I was listening for it—

Mark: Okay. [laughs].

Corey: —because I would have needled you. Instead, I had to go to my fallback, and needle you about your employer instead.

Mark: Fair.

Corey: Aren’t you glad you went that way?

Mark: Iffy. Depends what my boss says when he hears this. I agree, and Hero is a loaded word. I actually liked, or preferred—in Australia, there is a smaller regional program that's somewhat similar, that's the ‘Cloud Champions’ and that was a little more straightforward I think because these are people in the country who champion cloud usage and cloud technologies. And I thought that was simpler and easier to describe because I've been in a number of channels and outlets like the How to re:Invent series that AWS runs, and a few other venues, trying to explain the Hero program, and it the explanation of what we do is pretty straightforward but using the word Hero is uncomfortable because especially what goes on in the world on any given day, to call somebody who likes to teach people technology a Hero, I think that's a stretch, but I do very much like the program.

Corey: Especially during a time of pandemic where you have people literally risking their lives to bring you groceries, for example, and you look around it's like, “What do you do?” “I'm a Hero.” It rings a little hollow I guess is probably the best way to put it. I do love what they're doing in other geographical locations in AWS. They have a similar program I believe called Cloud Warriors, which just sounds incredibly badass.

Mark: Yes, for sure. But then I think the problem with warriors is you get this mental image, and then if I come on camera, you're like, “Oh, I didn't know it was nerd warriors. Not what I was expecting.”

Corey: Oh, yeah, in my case, that's all about fitness is really what I'm about, namely fitness entire burrito in my mouth.

Mark: You got it.

Corey: So, we've talked a bit about AWS, but what do you focus on? Do you spend time with the other cloud providers, specifically GCP and Azure—when we say cloud, that's generally what we're referring to—or are you more of an AWS focused person?

Mark: Yeah, I actually split my time among all three. The AWS designation and the reference point tends to go just to market share, and also the fact that they were first out so I've had the longest working relationship with them, but from Trend Micro’s perspective, we're officially partners with all three, and I've been involved since the beginning of all those, but just as a technologist I enjoy working with all three. They all have their ups and downs, but when I generally talk publicly about Cloud, it's Azure, GCP, and AWS.

Corey: So, you're in a great position to wind up answering this in a probably more objective fashion than I do, given the fact that my entire business revolves specifically around AWS. Compare and contrast the big three: what do you like, what do you hate about each of them? Superlatives, I guess. Go.

Mark: Nice. GCP, I think from a technology perspective, once you get your head around it, I like the basic structure. So, as a security guy, the fact that you need to turn everything on explicitly is a big win for me. Their project focus—as opposed to account focus—makes it a lot easier to implement some interesting boundaries. I also like the way that things like virtual machines are just sliders. Yes, they have templates for instance sizes in AWS, but it's easy to just say give me more CPU, give me less RAM, that kind of stuff. So very cool there. What I don't like about GCP is, when you're accessing it through the SDKs, you need to be Google-y. If you are not Google-y, it is going to be a very frustrating time. So, I do a lot of work in Python, the GCP Python SDK is the most un-Pythonic thing I have seen. It is tough to wrap your head around that, but once you realize it's really just all Java and Go primitives wrapped in Python, it works okay.

For Azure, I think the breadth of services are great. There's some really interesting things like Cosmos DB, which is trying to be all DBs to all people. Not quite getting there, but more than anything, I find the interesting balance where especially GitHub being under the Microsoft umbrella, there's a big merging happening there with some of the online coding tools that are happening in GitHub. So, I think there's some really good dev tooling in Azure. That being said, what the heck is Azure DevOps? Like, you can't take a philosophy and make it a service. That's not even really one thing, but combining thing among a bunch of other services. Really challenging there, very Microsoft-y. Same with the SDK for them. You better like having extra attributes for no reason in your properties just because that's the way they do it.

For AWS, it's the beast in the room, for sure. Very good at what they do. The biggest challenge I think, for AWS is trying to figure out what service at this point; there's just so many of them. But then also because of their rapid iteration in services, if you don't have the exact use case for a service when it comes out, it gets really frustrating really quickly because you see what it could be, and it will get there in a year or two, but not out of the gate, so whether or not you push through—and then I think lining up with with your opinion that you're never shy about, all three of them have a really, really hard time naming things.

Corey: I have noticed that, and I can't believe I'm going to say this out loud, but I have a sympathy for that problem because it turns out that it is way easier by four orders of magnitude to make fun of a name than it is to come up with a good one. I mean, for God's sake, my first version of this company was the Quinn Advisory Group. And it took us two weeks to come up with that.

Mark: Yeah. Yeah, it is hard, for sure—

Corey: Names are super hard.

Mark: —the challenge ends up being consistency. So, you look at AWS, and half of them seem to be named as nicknames, some of them are named very deliberately in what they do, others, you know, the whole AWS versus Amazon thing is just… okay. And then in GCP, it's better than their public-facing Google Meets, meet on Google, Google Hangouts, hang in a Google Meet, meet up in a Google Hang product fiasco, but it's not much better. And then Azure is actually surprisingly direct with the exception of the DevOps stuff. But again every once in a while, something more aspirational like Cosmos sneaks in. But at least they shoved a DB in there to make it make a little bit more sense.

Corey: You're absolutely right. And one thing that I think is happening is that, on some level, the folks using Azure and the folks using AWS—on the typical customer side of things, not necessarily the partners, or the folks that interoperate between the two—there tend to be almost two camps that don't talk very often. And when I started playing around with Azure, I find some things to be actively impressive about—same story with GCP as well. And it sounds kind of weird, but I mean, the best analogy I've got is when I learned French, I found that I understood English a lot better as a result. When I learned about Azure or GCP, I find that I start to understand concepts about AWS more effectively as well. For example, I still maintain one of the absolute best descriptions of what a lot of what AWS services are, are Microsoft's list of analogies and comparisons, were just lists a bunch of AWS services, the Azure competitor, and then—and this is critical—it describes what each one of those services does. And wow, how come AWS marketing hasn't gotten in front of this?

Mark: Yeah, [French]. The end of the day, as much as we poke fun at the naming, does the service work is really the key thing. But what frustrates me on the naming, you know, and again, I was having a discussion last week with some development teams, where they were having this revelatory experience because they found a service that addressed their need in one of these clouds. And I just kind of stepped back and went, wait a minute, that service has been out for, like, four years. But because of the name, it never clicked into them what problem it actually solved.

So, they were trying a couple different workarounds, looking down different avenues, and then when they found this, they were like, “Wow, it's super easy. I just had to click a button, create the primitive in the service, and I'm done.” And that's where the naming thing—besides just being a fun pastime to poke fun at the names—that's where it really frustrated when it's hindering somebody to solve a problem because they don't understand at first glance what this thing does. That's a serious issue.

Corey: It is. And I also take that a step further. Some people think I focus too much on names, and they may be right, but the more I make fun of cloud services, and the more opportunities I have—let's call it that—to speak to the teams that are building these services, and when you start tying a human face to these things, you develop—at least I develop—a sense of sympathy because it's hard to build these things; no one claims otherwise. And the challenge, as a result, is I don't want someone to release a new feature that they've been working on for months or years, and the first thing that happens I make fun of it. But no one spent 18 months naming Systems Manager Session Manager, and if they did, they should feel bad.

So, names are a safe thing, and another secret angle to that as well is you don't need to have 10 years of experience as an infrastructure engineer to appreciate a joke about a bad service name, whereas if I can start making esoteric references to XML as applied to the S3 API, yes, I can: you’ll want me to absolutely wash my mouth out with SOAP after that one. Those kinds of jokes are not going to be nearly as accessible, and it feels like it's inside jokes among friends. And that's never as engaging, and it acts as a gatekeeping aspect more than it is about building a bridge towards inclusivity.

Mark: Yeah, and I think there's another angle to that as well. I think using a joke to break down that barrier to be more inclusive also dispels the myth that cloud services are for the wizard on the hill, that you need this crazy level, that you need to be an expert before you just start experimenting and trying things. And that's one of the things that I think is amazing about Cloud is that you can just sort of stumble out of the gates, and have something working, and see the fruits of your labor. I spend a lot of time, especially now during the pandemic, looking at and helping kids learn to code. And one of the things that I find out amazing now vs way, way, way back when I learned was that there's a lot of tools, and kits, and toys that have a physical aspect that links to the digital so that kids are manipulating something physical with their code or with their hands, and it's changing the code.

So, there's a linkage there that makes it easier to understand. And I think there's a parallel there to being able to make fun of some of these names that some of them are, they're fine, it's okay, but it's still fun to make fun of others, like System Manager Session Manager, that deserves to be shredded to the ground every chance we get, but it does make it a little more accessible as part of that inclusivity. It kind of demystifies them a bit to say, “Yes we're joking around about this stuff.” It's not only, “I need to be doing serious work with this stuff.”

You can be having fun with it, which is where we see a lot of work around Alexa skills. There's a mountain of them that aren't ever going to be run more than once, but there's a lot of just little fun stuff that people use as an entry-level to start learning how to code, to start learning how to take advantage of cloud services, and that I think is really, really important because I don't think enough people experiment with technology and playing around with it not only just for the career-wise, but we're surrounded by this stuff. The more we understand it, and demystify it and make sure that it does break, but we can fix it. I think that's a win for everybody.

Corey: And I think that that's really what this comes down to. And I've seen that trend in the AWS ecosystem by itself, but I see it across the others as well, where once upon a time when EC2 first came out, you basically needed a doctorate to get it up and running. There was an entire cottage industry of companies like RightScale that made that interface something a human being could use. Over time, it's become more broadly accessible.

Look at things like Lightsail. You don't need to know much about anything within AWS’s very vast umbrella to get started with that. And we're seeing it now move even further up the stack in other arenas, too. I'm excited to see what the next five years hold in that context. The idea of low-code and no-code movements and that type of engagement means that suddenly you can have a business idea and not have to go spend the next few years learning how computers work in order to enable some of that.

Mark: Yeah, and that's where I think Azure, may be the dark horse, especially with GitHub being under the umbrella, but as much as—

Corey: Microsoft gets developers, credit where due.

Mark: This is the thing. As much as I like to go at Microsoft, they have a long history of enabling development for outside of the traditional IT umbrella. So, Visual Basic, huge win. In the early windows days, you could throw a couple buttons on a form, double-click on it, and write a tiny bit of code that you could, you know—not Google at that point because Google wasn't a thing, but you could read through the Docs or kind of stumble your way through, and you had a working Windows application. You didn't know anything about the Win32 Library set, you didn't know anything about how this worked under the covers, but there was. You typed in something in your text box, you clicked a button, and it took an action.

And I think we're starting to come back around to that with cloud services. We're seeing more and more of that in the data stack on Google, where Bigtable BigQuery, just shove a whole bunch of data in here, it's going to make some intelligent decisions, and then you can single-click or drag and drop things to get some really interesting business results out of that, and that's a huge win. Anything we can do to make it easier for non-technical people, or non-deeply-technical people, the better off we are. So, if we go back to our earlier part of our conversation, people like the AWS Heroes, people like yourself, people who are interested and really immersed in this stuff, we're going to figure it out no matter what. We can wade through the documentation, we can look at the mountains and mountains of bad JavaScript code and eventually make sense of something. But that's not who we need to worry about. It's a similar challenge we have in security is where we've made security this obscure art when really it's just one concern of many of everybody who's building technology, and we need to make that more accessible just like we need to make building the technology more accessible.

Corey: I don’t think that I've ever regretted making things accessible to a broader audience. One of the early versions of my “Terrible Ideas in Git” talk was aimed at inside baseball on Git developers. And yes, I had to learn how to Git worked before I could give that talk. And the three people who understood that in the audience thought it was great, and the rest didn't know what the hell I was banging on about. By instead turning it into something that was this is what Git is, and this is what it does, and here's how to use it by not using it properly, is fun, it's engaging, but it means that you don't have a baseline level of you must already have X level of experience in order to begin using this new and exciting capability. And that is the biggest challenge for all of the cloud providers from where I sit. And it sounds like I'm not alone, from what you have to say.

Mark: Yeah, I a hundred percent agree with that. And we're seeing it get better and better, with better interfaces, a better understanding that as far as—yes we need API first, but that also has a big assumption of technical level. If you're saying here, you can only interact with it via an API. But I have a similar experience. Last year—so 2019—one of the talks I gave at several of the AWS summits was about advanced security automations made simple, and it was targeted not at security people, but at developers.

And the whole idea was breaking down this myth that cybersecurity was super hard, and it was something that they had to hand off to another team, and it was just showing them, “Look, this is the basic thing. There's a whole bunch of big language around this that doesn't mean anything. Here's what you're trying to accomplish. You're trying to make sure that if you write a function that creates a TPS report, it can only print TPS reports, and not HR, or finance, or inventory reports.” Which I always say because no one knows what a TPS report actually contains, so I just randomly name other things that it shouldn't be doing, and nobody's ever called me on it so I keep getting away with it.

But the concept comes across pretty crystal clear for developers or people who are building something is that I want this to be an apple, and if it is an orange, that is bad. And that's really just cybersecurity in a nutshell. But I don't think many people present it that way. So, as far as making this more approachable—which may be a better word than accessible—more approachable, easier to understand, that's a huge thing for me personally, with security, with privacy, but with cloud in general. And I think the advantage—early in our conversation, we talked about environments not being nearly as dynamic as they could be, or as that we think they are, and I think a lot of that comes down to this tooling and that approachability.

Once we get there, and we make it super easy for somebody not to worry, that's really going to pay off. I saw that this week. I was doing a virtual meetup for Cloud Security London, and they had deployed a polling app, like an open-source Kahoot! And it was interesting because the cloud security meetup is run by twin brothers, one works for AWS and one works for Azure, and the third brother is not—

Corey: Thanksgiving dinner has got to be awkward.

Mark: It gets worse. The third brother, who's not a twin, is a .NET developer who works for AWS. It's a complicated relationship, to say the least.

Corey: And an NDA violation waiting to happen.

Mark: Oh, massively so. I'm sure even just the recipe for the yams at Thanksgiving is probably under one or more NDAs. But they deployed this open-source version of Kahoot! which was this multithreaded sort of polling app, and the nice thing is that they didn't have to know, even though these are really technical folks, they didn't have to know the ins and outs of it. The tooling was simple enough that they actually just ran one script—in this case, I think it was a Terraform—that deployed the whole thing for them, and it worked.

And to think about what actually happened behind the scenes is now you have a-real-time highly scalable, dynamic audience polling system in place after just running one command was pretty amazing. And I think that possibility is extremely exciting. We just need to continue to relentlessly chip away at the barriers to make it more and more approachable.

Corey: That's, I think probably the best place to leave this. If people want to hear more about what you have to say, where can they find you?

Mark: You can find me on social @marknca, M-A-R-K-N-C-A, or at my website, Markn.ca, as in Canada because that's where I'm at.

Corey: And we will of course throw links to that into the show notes. Mark, thank you so much for taking the time to speak with me today. I really appreciate it.

Mark: Thank you. I appreciate the opportunity.

Corey: Mark, Nunnikhoven, vice president of cloud and research at Trend Micro.k I am Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts, whereas if you hated this podcast, please leave a five-star review on Apple Podcasts anyway, along with a comment after you update your antivirus software.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Adithya Reddy

I do everything from web and mobile to backend architecture at Branch Insurance. I started out as a purely frontend developer but after joining Branch, which was 100% serverless from day 1, I expanded across the stack and probably write as much CloudFormation today as I do JavaScript.

Links Referenced:

  • Branch Insurance Main: https://ourbranch.com/
  • Twitter: https://twitter.com/TheTallpants
  • GitHub: https://github.com/tallpants

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is brought to you by Trend Micro Cloud One™. A security services platform for organizations building in the Cloud. I know you're thinking that that's a mouthful because it is, but what's easier to say? “I'm glad we have Trend Micro Cloud One™, a security services platform for organizations building in the Cloud,” or, “Hey, bad news. It's going to be a few more weeks. I kind of forgot about that security thing.” I thought so. Trend Micro Cloud One™ is an automated, flexible all-in-one solution that protects your workflows and containers with cloud-native security. Identify and resolve security issues earlier in the pipeline, and access your cloud environments sooner, with full visibility, so you can get back to what you do best, which is generally building great applications. Discover Trend Micro Cloud One™ a security services platform for organizations building in the Cloud. Whew. At trendmicro.com/screaming.

Corey: Normally, I like to snark about the various sponsors that sponsor these episodes, but I'm faced with a bit of a challenge because this episode is sponsored in part by A Cloud Guru. They're the company that's sort of famous for teaching the world to cloud, and it's very, very hard to come up with anything meaningfully insulting about them. So, I'm not really going to try. They've recently improved their platform significantly, and it brings both the benefits of A Cloud Guru that we all know and love as well as the recently acquired Linux Academy together. That means that there's now an effective, hands-on, and comprehensive skills development platform for AWS, Azure, Google Cloud, and beyond. Yes, ‘and beyond’ is doing a lot of heavy lifting right there in that sentence. They have a bunch of new courses and labs that are available. For my purposes, they have a terrific learn by doing experience that you absolutely want to take a look at and they also have business offerings as well under ACG for Business. Check them out. Visit acloudguru.com to learn more. Tell them Corey sent you and wait for them to instinctively flinch. That's acloudguru.com.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Adithya Reddy, the first developer hire at Branch Insurance. Adithya, welcome to the show.

Adithya: Hi, Corey, nice to meet you.

Corey: Well, always a pleasure. So, you're working somewhere fascinating, specifically, Branch Insurance, which despite all assurances to the contrary, I refuse to accept is not what developers need to get for when they really screw up a merge. So, is that roughly what you do, or is there a bit more to it than that?

[laughs]

Adithya: I would love if that was actually what we did. I would personally be the highest paying customer for that. But no, we are an [00:03:12 insurance startup], and we sell bundled home and auto insurance in a bunch of US states, list growing. And what sets us apart from a technology sense is that we've been serverless from day one. We started out servers, we are still today entirely serverless, we have never had a instance running. And it's an interesting perspective, just technically, about how that works in a industry like insurance, where that’s definitely not what you expect.

Corey: It's not, in fact. When you hear about a company that is, “Oh, we are all serverless, never had an instance running from the very beginning.” You think, “Oh, so what do you do?” “We're Twitter for Pets, which is just like regular Twitter, only somehow eighty times less racist.” But you're a [00:04:01 insurance startup], which is synonymous, for most of us, with regulated industry, which means not—shall we put this—the most avant-garde technology selections, if for no other reason than explaining it to regulators and auditors becomes a more than full-time job.

I mean, I had this problem back when I was working in FinTech startups, where we'd have auditors rock in and ask, “Great, so where's the Active Directory credentials?” “Yeah, we don't really have one of those.” And they looked at us like we were out of our minds. For all I know, maybe we were, but what has your experience been when talking to other folks about, “Oh, we're in the [00:04:39 insurance space], but we're also full serverless.” Is there a disconnect there?

Adithya: Most people, when I say that we work in insurance. It's just like, “So, you write PL/SQL queries, and you have stored procedures, and the best you could possibly hope for is lots and lots of Java.” But it's that assumption, but it's far from the reality of what needs to be done, and—

Corey: Oh, let's be fair, I'm sure there are some [00:05:05 insurance companies] that are all about the .NET instead.

Adithya: Oh, definitely. There's a lot of diversity in the space with technology.

Corey: [laughs]. It's one flavor of big enterprise software versus a different flavor of big enterprise software. But again, as you mentioned, you're an [00:05:18 insurance startup], and the idea of approaching an industry like that, that is-frankly—so incredibly rife, or in need of disruption, with a—frankly—bleeding edge architecture, I’m sorry, I just sort of look at this from a perspective of you can't help but admire the audacity of that.

Adithya: Yeah, I mean, we've had our own challenges, but I would say overall, I am pretty happy that this is what we've chosen, and just the benefits that we've seen from this, I don't think we would be able to replicate what we've built so far if we were just using the architecture that you would read about on Twitter, which is, like, a giant Kubernetes cluster and lots of Java apps. And just for some perspective, we’re pretty small. I think when we launched, we were two engineers inside the company, plus three people helping us out from Uruguay. And right now, we've just added a bunch more people, but we're still about 10 software developers, and we've managed to launch a complete insurance product that you can buy from five states. Everything works and people, kind of, don’t—they're not able to reconcile insurance plus a tiny team that builds stuff faster. And I would say probably the biggest thing that's enabled that for us is just not having to worry about a lot of things that serverless takes care of for us.

Corey: Now, as you mentioned, you're 10 engineers, and you have an established track record. Your founder, Joe is a known entity on the stage constantly at Serverlessconf. The hardest part is getting him off the stage long enough to have a conversation with him. And now it's a known thing, but you were the first engineer hired at Branch, and what was that like walking in the door? Because my immediate knee jerk response to, “Hey, we're doing this very legacy, very controlled thing in a brand new way, architecturally,” it sounds like Hacker News broke loose somewhere and is now dictating architecture. Today, it makes an awful lot of sense, but grinding back the clock a few years when you first started, did it seem as nuts then as it sounds like?

Adithya: So, there were a lot of very odd things about how I ended up here. So, for some backstory, I was in my final year of college in 2018, and I was like, “Okay, I need to find something to do now or I'm going to be broke.” So, there was a thread on Twitter—I forgot who it was—and it just said, if you’re a new grad looking for a job, put your resume here, we’ll critique it for you. So, I put it up and I got a DM from Joe who said, “Yo, I'm looking for an intern, a front-end development intern, would you be interested?” So, the whole thing was very suspicious because that's not the kind of offer you get on Twitter DMs, and then as I found out more about it, it's like, “Oh, we're going to build an [00:08:10 insurance startup].” And everything was just very fishy about it. But I thought, I'm out of college now. If I'm going to do something stupid, now’s the time.

Corey: Yeah, ridiculous job beats no job.

Adithya: Yeah. So, I said, “Yeah, let's go for it.” And I was an intern at Branch for six months. And at that point, I mostly just [00:08:30 unintelligible] the front-end stuff. So, I would build the front-end UI, I did some React Native stuff for our mobile app, and I did some weird UI sketches in Balsamic, and all that convinced me was that I shouldn't be designing UI.

And it just grew into a proper company, and everything just sort of came together, and that wasn't really something that I expected. So, I mean, now I'm really happy that's how it's all come together, and we have a great company running right now. But it was just very strange to me that we're building an [00:09:02 insurance startup], and also when I joined it’s like, “Oh, we're all serverless.”

And the day I got there, I asked, “Okay, so I'm going to build your front-end now. Where's the API?” And he said, “Oh, there's no API. It's, like, GraphQL.” And I said, “Well, what's GraphQL?” That was a big rabbit hole. And then I said, “Okay, so where is the back end running?” And he’s like, “It's not running anywhere. It's serverless.” And that was another six-month journey for me to discover that all this stuff that I had learned in college was not as mandatory as they made it sound.

Corey: I had to go through that same transformation myself, only instead of college—because some of us never went, or at least never graduated—it was, “Well, just unlearn the last 15 years of your career. It's fine.” And that 15 years of the career was spent on things that are now largely irrelevant in a serverless world, but they were the entire core of what I did. So, for an awful lot of folks looking at this with suspicion and distrust, it's not even so much a problem with the architecture or the constraints itself, but rather than it sort of cuts at the core of our identity. For the longest time, I was a Unix/Linux server administrator, and now there are no servers for me to administrate. What does that make me? It's the evolve or die dinosaur type of paradigm. And I won't lie, it's scary.

Adithya: Oh, yeah, definitely. I mean, there's still a lot of [00:10:19 unintelligible] questions that I get from people who I talk to, and I tell them this is what our architecture looks like. And, “Well, where’s your DBA, where’s your on-call, where’s your DevOps person? Who manages the infrastructure?” And it's just… Amazon does it because they do it better than I could, so why would I want to take on that load for myself?

And a lot of people seem to have just gone in the opposite direction of I want to [00:10:43 unintelligible] everything, and to me now with the perspective that I have at Branch, that seems just a tremendous waste of time, and also just a tremendous amount of stress and responsibility you take on yourself. Why would I try to do what Amazon has been doing for a decade now, and they have people who are paid much more than I would be able to pay someone to do it, and they're going to do a better job than me? So, why would I just not leave that up to them, and I will do what I'm good at, and I will do what my company needs? So, the big focus, the big shift, for me, has been coming out of school as I'm an engineer, and I do tech stuff, and that's what I care about, and it's all about the technology. And it's changed to, “What does the business want?” and, “What do we want to build for our customers?” And if we can pay someone else to take care of this stuff for us, and we can focus on the product that we're building and what our customers want, then that's a net win for me.

Corey: It really gives you the opportunity to focus, on some level, on the things that differentiate your business, that lead to success. I spent untold weeks or months of my life in aggregate now, configuring and building load balancers, but I never worked at a company whose job was to build the load balancer, it was always to do something thing else, be it, I don't know, help with expense reports or help dogs tweet to one another. It was never aimed at, I’m working at a load balancer company. Like, yeah, if you work at F5, you probably should spend a fair bit of time building load balancers, but for the rest of the world that is no longer a differentiator or internal core competency that you need to have in most companies. There are always exceptions, and they are always going to be these expressions of those core competencies.

I'm a big believer in the idea of the T-shaped engineer, where you should be broad across a wide variety of different things and deep on one or two areas. I started off being deep on email systems because every company back then needed an email administrator, and I thought it was fascinating and the closest thing I'd ever seen to magic. Now I look at this through a lens of, huh, it was pretty clear in retrospect that fewer and fewer companies ran their own email servers, so this was not going to be a job for most companies. Instead of every company needing dedicated email people, there were a couple dozen and that was it. I wanted to go with something that had a bit more mass-market appeal. And one of the neat things about serverless is that it seems to me at least, that it's making every engineer, to some extent, step through that gateway.

Adithya: Oh, yeah, definitely. I mean, I think one of the things that won me over about this whole architecture is that I came out as a front-end developer, right. So, that's what I do: I want to write react apps. And what was interesting to me is that Joe was like, “That's fine. You don't have to learn this stuff.”

And all of the engineers that we've hired have also just been front-end devs primarily, and we've always had the confidence that they'll just be able to pick up whatever they need to do for the backend, just because there's really not that much to be done. There is no setting up a server, there is no load balancers, there's no how do we set up this relational database, and how do we make sure that the replicas talk to each other? There's not been any of that. What has to be done? Can you write it in Javascript? Cool, so let's put on a Lambda, and then it'll run. And what do you need to store it? Can you put in JSON? Sure, so let's throw it onto DynamoBD.

And it's just been… it kind of moves the barrier to entry quite a bit back. I wouldn't say it removes it entirely; there's still like—you know, you have to be an engineer. It's not a no-code thing, but it makes it so that you will, as a front-end developer suddenly have a lot more power, so to say, about what you can build without having to rely on I need the DBA, and I need someone who knows Java really well, and I need someone who can deploy things to clusters and knows about Linux. It's just removing a lot of those things has made it so that we've hired engineers who have just been purely front-end, and then in two months, they're writing CloudFormation and setting up SQS queues and reading messages off it. It's just been really, really refreshing in that sense.

Corey: One of the things that you're almost certainly going to hear from the Hacker News ‘well actually’ brigade is, “Ah, but by going with serverless, you have now locked yourself into AWS. If you were to want to port Branch Insurance over to GCP, or Azure, or to your burning dumpster fire of your own data center, you're going to have an awful lot of work to do that, so you've made a grievous error, and now you are beholden to whatever Amazon decides to do with you.” How do you feel about that?

Adithya: I mean, initially, I kind of bought into it. You find new young developers and put them on Hacker News, and just, they’re going to absorb everything, and the end result is not going to be pretty. Luckily, I had a lot of people who—yeah, no, just don't fall into that clique, let's say, of Hacker News. And the realization for me about lock-in has been, yes, you're not going to leave Amazon anytime soon, but then I have never seen a company who were like, “Yeah, we're going to leave Amazon for Google because Amazon has decided to screw us by raising their prices.”

And as far as I know, I don't think Amazon has ever raised prices on anybody. And that's always been a constant: how do we make things cheaper, and sometimes far cheaper. And also the lock-in, there’s this assumption built into that statement of, “If we don't use serverless, we are not tied to Amazon at all, and we can just leave at the drop of a hat,” which is not true. There's a lot of stuff set up, implicitly, maybe not documented very well, but there’s just a lot of stuff going on with the way you set up your clusters on Amazon, on the way you use EC2, or the way you set up your load balancers, or your network, or DNS, and you're not going to be able to port all of that to Google, or Azure tomorrow. So, we're just talking about degrees of how hard is it going to be to migrate, notwithstanding the fact that it's very, very unlikely that you will ever have to migrate.

And then also, it becomes… if you work and build apps with the mentality that I need to be able to switch this away to other providers at the drop of a hat, you're locking yourself up in saying that I will only use the lowest common denominator of the services. And you're just making life harder for yourself, so why would you make the 99 percent case much harder to solve for the 0.1 percent case of, “Oh my god, Amazon is screwing us and we need to leave and go to Azure right now.”

Corey: In what you might be forgiven for mistaking for a blast from the past, today I want to talk about New Relic. They seem to be a relatively legacy monitoring company, and I would have agreed with that assessment up until relatively recently. But they did something a little out there: they reworked everything. They went open source, they made it so you can monitor your whole stack in one place and, most notably from my perspective, they simplified their pricing into something that is much more affordable for almost everyone. There's even a free tier with one user and 100 gigs per month, totally free. Check it out at newrelic.com.

Corey: I agree with everything you just said, but by hearing someone who's actually built something on top of these technologies, it comes across as way more genuine than if I sit here on my pedestal of thought leadership, which is, turns out, not a real thing, and then go and tell everyone that, “Oh, this is what people should or should not do.” Hearing people who have made actual business decisions around this is one of the most authentic ways to make the point. I keep looking for people to take the other side of it, “Oh, no, we actually had to build something in a way that was portable to other providers.” But you never see them taking advantage of that capability. It always, for whatever reason, comes down to a baseline model of, “Well, we built this optionality in, and kept everything agnostic, and then we feel better about running it all in AWS,” or GCP, or Azure, wherever they happen to be running.

But it's an optionality that people never take advantage of. So, my position is, go all-in on a provider—I don't care which one—and see what happens. If you have to move later, at least you've survived long enough to get to that point of having to make the pivot, rather than spending all of your time working on these agnostic shim layers, rather than building out features that will move your business, and get to the next milestone.

Adithya: I mean, I agree with that. I think one thing that people don't realize is that you might think you're building an agnostic application, but I can almost guarantee that at 99 percent of cases if you have to move your cloud provider, you're going to find out it wasn't really as agnostic as you thought it was.

Corey: That's one of the things that I find so… I guess ridiculously strange. As far as these people talking about, “Well, what if? What if everything changes, and we have to pivot massively? Well, what if AWS shuts down?” Okay, let me back up for a second. If AWS shuts down, are any of us going to keep our existing jobs or, far likelier, are we going to instead pivot to helping all of these companies with big problems and bigger pocketbooks migrate off as fast as they can? That is such an unrealistic outcome, that I can't quite fathom. It comes down to where I want to spend my time. I don't really think that hedging the risk of this Cloud thing being just a giant fad is the best use of my time anymore.

Adithya: That's true. It also just comes down to the probabilities. What is the probability that you are going to run out of money trying to build everything yourself, versus the probability that AWS is just going to decide that, “You know what? Screw you, we're going to raise all our prices.” And it’s just—I feel like a lot of people just want to tinker with cool stuff: they want to tinker with load balancers and play around with Linux, and that's fine, but it's the fact that I'm going to justify this with a business reason that I really don't get. That's very strange to me.

Corey: So, one of the things that you put in your biography, that we always ask for, “So, who are you, and what do you do?” When people sign up to record an episode here, is that you started off as a front-end engineer, and now you write at least equal amounts of Javascript—or ‘yahva-script’ as some pronounce it—and CloudFormation, which is interesting on a couple of levels. We'll start with the easy one first. Why CloudFormation instead of Terraform?

Adithya: So, I think there's a couple of reasons for it. The primary reason that we started off with CloudFormation is that we have a really excellent partner in Stackery, who support automation really well, and they give us a lot of help with our tooling. And a lot of times it's been, we need this feature, and we ask Stackery like, “Hey, how do we do this?” And they say, “Oh, no, we can't do that yet, but thank you for pointing that out. We will build it for you.” And we have it two weeks later.

So, that was the reason we started with CloudFormation. And then, as time has gone on, we realized that it's a lot simpler to do things with CloudFormation—if you're primarily being serverless—on AWS, just because of SAM and the way that AWS is supporting it. And they're doing a lot of work by creating these specific serverless resources of an AWS::Serverless::Function. So, having that kind of higher-level abstraction over the base CloudFormation, which encapsulates all these common patterns of here's one resource you can declare that will create IAM roles for you, and also our CLI is going to package up your code and upload it, and then set up the function running.

And you have things like an HTTP API resource, which would be very, very hard to set up manually, but because of this kind of thing where Amazon has seen the patterns that are generally prevalent with serverless applications and made it much easier to use and they seem to keep iterating on it, that's been a big plus for us. And also just having the state of your application be managed by Amazon itself, like the rollbacks, and updates. And Terraform’s kind of coming there now with their Terraform cloud offering, but that's how we started, and now it’s… maybe Terraform can do some of what it does, but I would much rather bet on CloudFormation in the sense. Of course, if you're doing this on not AWS, you probably might have a different answer, and you should do your own research and come to your own conclusions.

Corey: One of my favorite things is watching people, I guess, incorrect each other on oh, well, you should only ever use one or the other, or some other system because of this one weird edge case, that it turns out hasn't been true in three or four years. But as a counterpoint, I have problems personally with JSON. Specifically, it doesn't parse as well as YAML because I have a Python background and whitespace messing things up for me has really been my entire career. So, back originally, CloudFormation only supported JSON. And then one day, bam, YAML showed up. And this was awesome. And this was great.

And one of the problems I saw for the longest time was that all of the examples, all of the blog posts, all of the stuff people had written, it didn't really matter because all these examples in this giant history with CloudFormation had been written in JSON. So, it took time for there to start to be public examples of what the YAML version of it would look like. And I'm wondering, first, if you've encountered that, and secondly, if that seems like it's a bit of a carry-forward into other areas where a new feature comes out, for example, but all the old examples don't demonstrate how that feature works.

Adithya: I mean, I've seen my fair share of—I Google, how do I do this? And then it's a bunch of JSON. And I personally think the JSON representation of CloudFormation is pretty hard to read. There's no comments, there's no lots of flex spacing, you can't really differentiate things for the reader. So, luckily, if I just copy this JSON and paste it into a JSON to YAML converter, it does have to work.

It's not going to be accurate, but it's useful enough that I can write my own YAML now, looking at that. Or you could just write your own YAML looking at the JSON. But I think in general, this is something that most developers have to pick up as a skill of how do I differentiate old information, and how do I not fall into the trap of reading old blog posts, or old books and thinking, yeah, this is how it is now, just because this field changes quite a lot. And this is just the general skill that you need because a lot of stuff on the internet when it comes to anything related to software is going to be outdated, and you need to figure out what resources that you need to use, and also trust the authoritative source of Amazon's documentation first, and only then go around looking for supplemental information, let's say. And also be very suspect, look at the date of the article before you’ll start to read it, and don't take anything you read as gospel because it's quite possible that the thought leader whose blog post you're reading also might have been wrong about something.

Corey: Oh, my stars, yes. One of the hardest challenges for me when I'm putting together my newsletter every week is I see a lot of things come across the wire in terms of blog posts that touch on what the community is doing. I have to vet these because some of them are not just poorly written but actively harmful. I saw one somewhat recently that’s, “Oh, you can save money by moving all of your S3 long term stuff into Glacier Deep Archive.” Well, yes, you'll save some money, but there are some serious constraints around that, around access times, how it's built in certain ways, and it didn't even hint at some of those. If you blindly go ahead and click through and, “Sure that sounds awesome.” It's not going to go well for you. You want to make sure that the thing you're applying is relevant, and it takes your constraints into consideration. In some cases, it feels like people write things that don't even take their own constraints into consideration. It's, more or less, Hacker News fanfiction.

Adithya: Oh, yeah. I mean, you see this a lot in Hacker News specifically. You go look at a post about, “Oh, look at this cool new thing I built.” And you look at the comments and it's all, “Oh well, doesn't this have this issue? And doesn't this have this bug?” “Oh, this was fixed in 2006.”

But it's just, you read it once, and now you just—I mean, it's a general thing in software, where you read it once and you just assume, yeah, this is a constant. This is not going to change, but stuff changes all the time, and you always have to approach everything you read in this industry with a lot of skepticism about, is this still up to date? Does this person know what they're talking about? Even if this person does know what they're talking about, have they made a mistake with this? Do your own research.

And also, it applies to you, too. Know that your own research isn't infallible, and you're going to make mistakes, and when someone else points that out to you—in a blog post, let's say—don't go to the Hacker News comments for that blog post and point out the person writing it is a really stupid idiot and they don't know what they're talking about.

Corey: Oh, absolutely. People don't know how to, in many cases, ask questions or answer them. I tend to assume good faith, and I find that I am pleasantly surprised way more often than I am negatively surprised when I do that. I assume that people are asking a question that they've tried the basic things, I assume that they're asking in good faith, not trying to prove a point, and I assume that when someone says, I'm looking to do X with Y technology, that the correct answer is not, “Why technology sucks.” There has to be a reason that they're asking this question. Helping, in some cases, clarify and refine the question is a better path than assuming they're a moron and lighting into them. I wish more people took that perspective if I'm being perfectly honest and grouchy about it.

Adithya: [laughs]. I couldn't agree more.

Corey: So, if people want to hear more about what you have to say, where can they find you?

Adithya: You can find me on Twitter. You can follow me on GitHub. I should start a blog sometime soon, but I just never get around to it.

Corey: The story of our lives. You are @TheTallpants on Twitter. And we'll put a link to that in the [00:28:49 show notes], of course. So, Adithya, thank you so much for taking the time to speak with me today. I really appreciate it.

Adithya: Great. Thank you. Thank you so much. Thanks for having me.

Corey: Of course. Thank you for taking the time. Adithya Reddy, first engineer hired at Branch Insurance. I am Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts, whereas if you've hated this podcast, please leave a five-star review on Apple Podcasts and file your complaint in both JSON and YAML.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Christina Warren

Christina Warren is a Senior Cloud Advocate at Microsoft, where she helps shape the overall video and broader content strategy for Channel 9, Docs.Microsoft.com, and the greater CA team. In this role, she hosts shows on Channel 9, Microsoft’s video channel for developer content, creates technical content snd demos, speaks at events, and interviews people within the developer community. Prior to joining Microsoft, Christina spent a decade in digital media as an editor, senior reporter, and commentator, with a focus on technology, business, and, entertainment. As a journalist, she appeared as an expert or commentator on ABC, NBC, CBS, CNN, CNBC, Fox News, Fox Business, Bloomberg, the BBC, Marketplace Radio, The Today Show, Good Morning America, and many more outlets. She also co-hosts Rocket, a popular tech news podcast, which has the distinction of being one of the only tech podcasts with an all-female hosting team.

Links:

  • This Week on Channel 9
  • Rocket Podcast
  • Microsoft Build
  • Microsoft Developer YouTube
  • Screaming in the Cloud Episode 68

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored by a personal favorite: Retool. Retool allows you to build fully functional tools for your business in hours, not days or weeks. No front end frameworks to figure out or access controls to manage; just ship the tools that will move your business forward fast. Okay, let's talk about what this really is. It's Visual Basic for interfaces. Say I needed a tool to, I don't know, assemble a whole bunch of links into a weekly sarcastic newsletter that I send to everyone. I can drag various components onto a canvas: buttons, checkboxes, tables, etc. Then I can wire all of those things up to queries with all kinds of different parameters, post, get, put, delete, etc. It all connects to virtually every database natively, or you can do what I did, and build a whole crap ton of lambda functions, shove them behind some API’s gateway and use that instead. It speaks MySQL, Postgres, Dynamo—not Route 53 in a notable oversight; but nothing's perfect. Any given component then lets me tell it which query to run when I invoke it. Then it lets me wire up all of those disparate APIs into sensible interfaces. And I don't know frontend; that's the most important part here: Retool is transformational for those of us who aren't front end types. It unlocks a capability I didn't have until I found this product. I honestly haven't been this enthusiastic about a tool for a long time. Sure they're sponsoring this, but I'm also a customer and a super happy one at that. Learn more and try it for free at retool.com/lastweekinaws. That's retool.com/lastweekinaws, and tell them Corey sent you because they are about to be hearing way more from me.

Corey: Normally, I like to snark about the various sponsors that sponsor these episodes, but I'm faced with a bit of a challenge because this episode is sponsored in part by A Cloud Guru. They're the company that's sort of famous for teaching the world to cloud, and it's very, very hard to come up with anything meaningfully insulting about them. So, I'm not really going to try. They've recently improved their platform significantly, and it brings both the benefits of A Cloud Guru that we all know and love as well as the recently acquired Linux Academy together. That means that there's now an effective, hands-on, and comprehensive skills development platform for AWS, Azure, Google Cloud, and beyond. Yes, ‘and beyond’ is doing a lot of heavy lifting right there in that sentence. They have a bunch of new courses and labs that are available. For my purposes, they have a terrific learn by doing experience that you absolutely want to take a look at and they also have business offerings as well under ACG for Business. Check them out. Visit acloudguru.com to learn more. Tell them Corey sent you and wait for them to instinctively flinch. That's acloudguru.com.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Christina Warren for a second time. Christina, welcome to the show.

Christina: Hey Corey, it's great to be back with you again. I guess it's a little over a year later since we talked. Things have changed in the world a tiny bit.

Corey: Just a smidgen. Some things haven't though. You are still at Microsoft, and you remain a senior cloud advocate.

Christina: That is also true.

Corey: Whenever I hear someone say that they're a cloud advocate, senior or other appellation, I just tend to assume that that job basically entails people asking you, “So, what do you think about Cloud? And your response is, “Well, frankly, I'm for it.”

Christina: [laughs]. Yeah, I mean, you're not wrong. I mean, so I'm a part of Developer Relations, and we talked about this on the show that we did last year, and so I encourage listeners to go back and listen to that one if you want to know more about my career transition and things like that, but I think that, yeah, you're not wrong. Part of it is absolutely saying, “Why yes, I am for the Cloud.” But I think the bigger thing, the way I view my job is that I'm advocating for the users, and we kind of act as this bridge between the people in the product teams and people who are using our products, as well as the people who are marketing, kind of creating content for those things. So, we try to sit in that middle space where we're really advocating for the users and doing what we can to improve the products so that people will build more stuff on our platforms.

Corey: So, there's a lot to admire about this decade’s Microsoft, but one of the most admirable things I've seen recently, just in terms of achievement, not in terms of necessarily impact in the world was, with only a couple of months notice, you were able to turn Build not just into a digital event, but rather, sort of, the definition of what a lot of online events could aspire to one day become.

Christina: Yeah, I know. The team did an amazing, amazing job, and really, it came together in about five weeks because the decision was made in March to move it online, there were still some other things that had to be figured out, but it really was about a five or six week period where it all came together. And I'm with you; I mean, the team and there were so many people involved. They did just a tremendous job, not just from creating the content, and how do you move what the sessions look like, and what's the formatting, and how is that different when it's virtual versus when we do things in person? But then there were also a lot of technical things that had to be done with making sure that Microsoft Teams would work the right way and that we will be able to have an interface for people to be able to select the sessions that they would want to watch at different times, and working with people to make sure that their setups that they're using from home will be robust enough, and that they can deliver their content. What we also did for Microsoft Build this year, in addition to being completely virtual, is that we did a 48 hour live stream so it was across time zones, and that was a massive undertaking because traditionally we do it on the west coast time zone, you go, you start, maybe, at eight o'clock in the morning, you end at five, and then people have their side events, and stuff and certainly content is often streamed online and people can tune into those things, but it's not set up for people who are in different parts of the world to be able to experience in a real-time way. And in this case, it was, and so we had the key segments were available on a replay for the ideal morning time zones for different parts of the world, but what we also did is we had presenters who were often presenting three different times, sometimes other presenters helped out live for a specific time zone. Meaning that when—for instance, I was hosting a lot of the Build live content, and I was holding down the desk from Redmond, but we also had hosts in the UK because that was the time zone that my shift was on. And so it was midnight to 9 a.m. in the United States, but that would be the equivalent of [crosstalk]—

Corey: Oh, I saw that on twitter at one point. So, one of the nice things about the pandemic is it lets me replace some of my bad habits with better ones. For example, I took my bad habit of sleeping, and replaced it with a good habit of lying awake in the middle of the night and worrying—

Christina: Right.

Corey: Which is also known as Tweeting. So, I saw it scroll past where suddenly I see a pic of you in a mask—good for you—in a car going somewhere. Well, that's not something we do these days. What's the story here? And you were going to do your segment at Build.

Christina: Exactly, exactly. So, we had hosts remotely, but we also wanted to tie things back to our studio, and also just frankly, in case there were technical problems and somebody’s internet went down, we didn't want the stream to go down. We didn't want dead air. But people were doing presentations live throughout the 48 hours, regardless of what timezone they were in, meaning that sometimes the presenters, it might have been three o'clock in the morning that they were doing a breakout session with people and we're taking Q&A and it was actually in real-time.

So, it wasn't a situation where I think—it would be easy to think, okay, everything is pre-recorded, and you just make it available to people, maybe you hit a play button and people are able to interact with it that way. That's not what the experience was, and so there were a lot of moving parts. I am so proud of the team who did all the work. I had a small role but I was very, very proud to have had even any role in it because I really think that it was a terrific event, not just for virtual events, but I really feel like based on the feedback we heard from the community, people had a really good experience, and that was really gratifying.

Corey: It really was a, forgive the Amazonian language, but it really was a bar-raising experience for online conferences, in the sense of this is what they need to become. It was not just turning a typical event you'd see in person into now a bunch of videos you can go and watch. The talks have to change, the interaction model has to change. I know I mentioned it when I spoke to Jeff Sandquist on the show in a previous episode, but I'll say it again, one of the most impressive things I saw was Emily Freeman giving a talk that was obviously pre-recorded. And then at the end, she just starts answering live Q&A from the audience and, “Oh my god, it wasn't recorded at all. She did it live.” Which is, from my perspective, a little on the silly side because what is the value of that, but on the other, she did it so well, it was flawless. And that was the first time of three that she gave that talk during the conference.

Christina: No, you're exactly right. And you're dead on. Emily—and so many other people were pros—Emily did such a great job that yeah, you would think that it was pre-recorded. But when it is live, when you can answer those questions, that does add a little bit of a different dynamic. And what I'm hoping comes out of this is that when we are able to return to having in-person events because I'm actually still very pro in-person events.

I think that there's a place for both. I personally get a lot of value from meeting people and being around people; I really like that. But what I'm hoping comes from this, and I've had this conversation with a number of people at Microsoft and other places too, is that what this will mean is that in the past the online component for various events, whether it's a developer conference, or something else is always been seen as an also-ran. You know, just been ad hoc.

Corey: It’s an afterthought. “Well, we have all the trouble of getting these people here. May as well slap a crappy camera and some bad audio in the back of the room, and we’ll put something up.”

Christina: Exactly. “Okay, we want to be able to have some on-demand stuff for later if you can't attend the sessions.” But it's an afterthought. It's not considered the same experience. And what I'm hoping comes from this is not that we get rid of in-person events because again, I think there's tremendous value there, but that we don't consider the virtual component a second class citizen, and that we start to see them as equals because they should be, just based on how our world works, and how our world is going to work going forward. I think that's really important. So, to me, I think that's what I'm most hopeful about is that we will not go all-in on one or the other, but we'll say, “No, these are both things we need to do, and we need to be thinking about both of these experiences, in whatever capacity we're doing them in.”

Corey: One of the hard parts for me looking at this is, I don't want to call any individual company out, and fortunately I've been to enough of these things where I don't have to. There are an awful lot of these online events that are frankly terrible, where you’ll have some people who are extremely good at giving talks and are used to working a crowd, and now you're going to put them in their home, you're not going to have any custom lighting setup, you're going to have them using their potato-quality webcam, and it looks bad. One thing that I found that I'm doing is instead of speaking engagements for some of the sponsor work that I do, we've switched to doing some of these online, and after the first couple it was, this isn't going to work. So, first, I upgraded all of my equipment—and we can get into an equipment rant in a minute—

Christina: Yes.

Corey: —on my end, so now things are effectively flawless from an AV quality perspective. And then the problem there was, “Okay, I look great and they look like a potato.” So, now I have a pelican case that I ship out to the other end when I do these webinars, and it’s, oh, now we both look good; everything's there, and then we can do a lot of editing of that work in post, because it almost never needs to be done live, and then you can fix any embarrassing bloopers, you can have smooth transitions, you can make the video work. And you can also avoid the trap of, yeah, for a 45-minute talk, no one is going to get up and leave from the front row, at least in large numbers because they feel it's rude. Closing the tab is way easier, too.

Christina: It’s way easier. It's way easier. And if you have bad audio, or if it's hard to see, or if it's not keeping you engaged, you're not going to continue watching. We've all learned that. I mean, I think that was one of the reasons why we made a point where our segments that Build this year were 30 minutes because we were trying to be very intentional about how long can you expect people to focus in on something? And some people really like longer content, and there's definitely a place for that. But you have to be thoughtful about the fact that just because people are viewing this from home doesn't mean that you suddenly have unlimited time to be able to sit someplace, and doesn't mean that your attention is going to be the same.

Corey: I thought that Build worked for this masterfully. The longer form content, people appreciate that is great. I'm not one, personally, I have attention span issues, which is pretty obvious by anyone who knows me. I sometimes struggle to have the attention span to complete writing a tweet. But there's a lot of people for whom that is very much not true, and that is the way that they absorb things best. I love the fact that there are different sizes of content, different formats, and different ways of meeting people where they are. I agree with you, as far as what you said earlier, where there's tremendous value to me in going to in-person events and meeting people and forging connections. That's where we met a year ago at build.

Christina: Sure is.

Corey: The problem there is there has to be a business value that can be articulated to get that critical mass, and if we can say that an online event is 90 percent as good as being at an in-person event, then great. Is that extra 10 percent worth the additional expense to the business? And by expense, I’m not talking to the hotel and airfare. That's usually irrelevant. Now that I run a business and see the economics, I believe that more than ever. It's the opportunity cost of you're not going to get anything else done for one to three days while you’re at this event. So, is that cost worth bearing? For an awful lot of conferences I've been to, the answer that has been, not really.

Christina: Yeah, no, I mean I think it's a really good question. And that's going to, I think, really impact the event business going forward more than just the traditional, can people travel, and do people feel safe being in large groups, and what is the appropriate thing to do there? I think you're right, people have to evaluate the business benefits, and what the cost is. And in some cases, I think you can say there are certain conferences where, yes, it is worth it. It is worth whatever the cost is going to be, whether it's the hard costs, or if it's going to be something like the time and what goes into it, what you can potentially get out of it is going to be worth that. And for some things, that's maybe not. And so I think a lot of events are going to have to really start thinking about what is our value and how do we make sure that we can get that across, whether it's virtual or it's in person.

Corey: One of the things that people are getting wrong across the board right now, is this idea that whether your talk is good or not is going to depend on what kind of microphone you have, what kind of camera you're using, how well you wind up looking at the camera versus off to the side. And all of those things are definite value-adds, but the thing that's going to make or break it is, is the content good, or is it crap? Because people will suffer through an awful lot for good content. I would talk to you on this podcast over a rusty set of tin cans and a string if I need to. But it doesn't matter how good the production quality is if you're boring.

Christina: Yes, that's a very, very good point. And you'd mentioned something earlier that said people are used to giving talks in person, and now they have to do things online, and I do think that's actually something that's important to point out and it's something that changes a little bit in my job because historically I've done a lot of in-person talks, that's been a big part of my job. I'm fortunate in the sense that I have actually more experience, and in some ways, I'm more comfortable doing things either pre-recorded like a podcast or even live video to a camera. I have more experience that way than I do in person. And if I'm being honest about my own strengths, I’m probably better in front of a camera than I am in front of a crowd. I think I'm a very good public speaker in front of a crowd, but I think I’m better—

Corey: You are, in case you wondered.

Christina: —thank you, but I think I'm better in front of a camera. And the reality is, is that they are different skill sets. Now, that doesn't mean you can't achieve both, but you have to think about it a little bit differently because when you're giving a talk in person, you do have the immediate feedback loop of the audience. You can feel that energy, for better or for worse, you can see what the feedback is, and you can riff. a lot of people are built off of that, and it can really change the dynamic of what the presentation that they give can be when you see really fantastic live speakers, they are, in my opinion, usually people that are completely feeding off the energy of the audience, and then the audience in return is feeding off of their energy, too.

And it's different when you're presenting virtually or to a camera. It's just a different concept because you don't have that feedback loop. And I think that a number of people who are really, really good public speakers aren't necessarily as comfortable on camera or on microphone because they don't have the experience. They watch themselves back, and they're like, “Do I sound like that? Does my voice really sound this way? Is my movements, are those things correct? What is my eyeline like?” People can become obsessed about little things, and that—maybe they feel more stilted—and that can affect the experience.

And so I think this is something that a lot of people are going to have to start playing around with and getting more comfortable with on their own about how they can do the right things and make that content, to your point, interesting. So that it, regardless of the quality of your camera or your microphone—obviously those things can help—people want to continue to engage and stay tuned because you're right: we will pull up with a lot, if the content is good enough. But the minute that the content is anything other than just exceptional, those other things, in my opinion, like the audio quality and video quality, that then starts to really just become a bigger and bigger issue, and makes you just that much closer to [crosstalk]—

Corey: [crosstalk] sound like a jerk, so that's going to be a problem.

Christina: Yeah.

Corey: I, fortunately, wound up doing a little bit of video work a couple years ago when I took over some of the release review segments for the A Cloud Guru video training series. And that was, again, doing video right, to some extent. They had a production studio there, or they would have a video crew that would come in and do the recording here, so I had people dealing with things. Remember back we could have people working on video, and—

Christina: Yeah.

Corey: —[crosstalk] ourselves?

Christina: I miss that.

Corey: And it's a very different experience because I spent a lot of time on stage, and I flatter myself perhaps, but I believe I give good talks. And the reason I give good talks is because I gave a lot of terrible talks for a while, first. That is sort of the progression that it takes. But it's a completely different skill set. I normally will have a few bullet points on a slide, and the presenter notes at most, and I'll get up on stage and riff off of it. That doesn't work; there's nothing to riff off; there's no energy; it's you and the camera. So, I started using a teleprompter and writing these things out. It helps that I write the way that I speak, so that makes it easier, but using a teleprompter is its own skill.

Christina: It is. It is. I use a teleprompter for the show that I do, This Week on Channel 9 and I'm really good on a teleprompter. I was fortunate that I had teleprompter experience before I joined Microsoft, but it's interesting because I've worked with a lot of colleagues who they've never used teleprompters before, and then they do for the first time and you think it'll be easier than it is. And it's not. It takes time to get the timing and to get the other things down and even writing how you write your script for the teleprompter and making sure things are spaced out enough and that you've got things moving at the right speeds.

That all takes work, and you've got to figure it out. But to your point, I'm the same way: when I give live speeches, I tend to speak more extemporaneously, and I tend to riff more because you can. But if you were doing something in a recorded scenario, you don't have that same luxury because you have to stay consistent and on track, and it's one thing if we're having a conversation like you and I are right now. You can have your outline of your notes, and we can note things we want to say, but it actually works better if we’re not scripted.

Corey: You can use notes? That would make this way easier?

Christina: Well, you have notes but, you know, you have [crosstalk]—

Corey: Well, I would hope I do, but you’d be surprised how unprepared I am these days.

Christina: But what I mean. You have kind of an idea. You don't have anything scripted out, but you maybe have an idea of what you want to talk about. That works for this sort of scenario, and for these sorts of conversations. It doesn't work if I'm presenting how to do something if I'm going to talk about how to create an Azure static web app. I need to actually—this is going to be a recorded thing. I need to have it as scripted as well as possible so that I know that I'm not missing anything because people are going to be consuming it in a different context.

Corey: Right. I'd be curious to hear your teleprompter story because my experience has been that there are traditional, extremely expensive, professional teleprompters, yada, yada. What I use are these metal and glass things that sit on a tripod, you put the camera inside of it, and then it has a holder at the front where you can put a tablet or a phone there. I [unintelligible] for a dedicated, cheap iPad that has a teleprompter app off the store.

Christina: Same.

Corey: And that sounded great. Now, the first version of the app that I looked at, great. It could hear what I was saying and automatically scroll to match where I was in the script. It's sucked. It was terrible. There was no good way to do it. How did you solve that problem?

Christina: Yeah, so I use an app called Teleprompter Premium, I believe is what it's called from JoeAllenPro. This is for iOS, and they actually have an Apple Watch app, too. I believe it might be able to do that thing where it can hear you, but I don't do that. Instead what I do is I set the way that it scrolls, and then I set that to a cadence that I can keep up with and talk to. And so it took a lot of testing and training to know this is the right speed of scrolling for me.

Really a big part of it was me learning to write my scripts the right way and knowing that I need to space things out in certain sections and not have really long blocks of text, to be able to have things that I can get through. And that was the big thing, knowing how to write the scripts correctly so that when it's scrolling, I don't have to worry about it adjusting based on me talking. That would be great, and some of the professional teleprompters can do that, but most broadcasters, how most of their work works is that the teleprompter is controlled by an actual technician. So, they are actually manually adjusting the script based on how the anchor is speaking, and they can either speed things up or slow it down, or go to a completely different section if that's what they need to do. So, that's how it works in broadcast.

So, to approximate when you're doing it yourself, I do like the app that I use because it has an Apple Watch app, meaning I can stop, or speed up, or pause, or whatever, if something is going too fast because what'll happen occasionally is I'll be recording something, I'll be like, “Oh, I've gotten ahead of myself here,” or, “This is going too fast, I need to go back,” And rather than having to stop, go to the camera, scroll back, I can just use my Apple Watch to do that, which is useful.

Corey: In what you might be forgiven for mistaking for a blast from the past, today I want to talk about New Relic. They seem to be a relatively legacy monitoring company, and I would have agreed with that assessment up until relatively recently. But they did something a little out there: they reworked everything. They went open source, they made it so you can monitor your whole stack in one place and, most notably from my perspective, they simplified their pricing into something that is much more affordable for almost everyone. There's even a free tier with one user and 100 gigs per month, totally free. Check it out at newrelic.com.

Corey: On of the ways that I got around that particular problem myself because I was already learning a bunch of new skills and didn't really want to have to learn script writing and being able to plan the cadence of what I was going to say in that way is, I put a little thought into this, and because I've used apps that have the remote control with something you hold, an iPhone, for example, but then you have to make sure it stays on, make sure you're tapping it in the right way, sometimes it's not as responsive. So, what I did next was I wound up spending 100 bucks on Amazon, which sells pretty much everything, and most of them useless, but this was pretty great. It's a pair of foot pedals that wind up acting as a—you can remap it in their app to a bunch of different keystrokes whatever makes sense, so I can either have it control the speed, go back, go forward, I could have it advanced line by line, or page by page—

Christina: Oh, that’s genius.

Corey: —because I'm not doing videos where anything in my lower half is exposed, so it's not noticeable when I'm tapping something with my foot.

Christina: Right, right.

Corey: I’m sitting here smiling, looking calm and composed, meanwhile, I’m frantically tap dancing underneath the table.

Christina: Okay, that is a brilliant idea. Okay, I'm going to steal that because that's even better than my Apple Watch solution, honestly.

Corey: It also lets me to have my hands on camera, and not look like I'm sitting there scrolling.

Christina: Yeah, I mean, well, usually what I would do is I would have it scroll, and I have my hands, and then if I needed to stop and go back, that's when I would use the Apple Watch to scroll back if I needed to. But I really like the foot pedal idea. That's brilliant.

Corey: Yeah, these are these little things I've iterated on, again. Normally—please learn from my mistakes so that you don't have to wind up making the same—

Christina: No, this is great.

Corey: —sort of path that I do. I spent extortionate amounts of money on equipment that is sitting unplugged in the closet, because, “Well, that was generation three. I'm now on generation six of my podcast setup,” and it's ridiculous, but it works, and it iterates forward. The problem is everyone has an opinion on this stuff, and opinions are terrible unless they're your own.

Christina: [laughs]. Yeah, no, I mean, that's true. I mean, we were joking before we started talking that we can kind of—and I think this is regardless of what field you're in, and I know this isn't as cloud-centric, but I think that whether you're an implementer or a developer or whatever, we're all now in this space where we've had to become AV experts overnight because it's become a crucial part of our job, and of course, Amazon sells everything except they don't have capture cards, and certain equipment in stock. To try to find a webcam these days is still over to kill us challenge because all of a sudden, we have to have these really high tech setups at home if you want things to be as good as possible. And you can fall down a rabbit hole in that way, too, where maybe you go too far.

Like for what you're doing, I think that it makes sense for you to iterate and to continue to find the best solution possible. I do think that sometimes, you know, regular people, if you're just in Zoom calls, you might not need to be spending $1,000 on the camera and getting the pro microphone and whatnot. You could do much, much better at a fraction of the cost and still improve your quality significantly. But it does change things if you're now going to be connecting with people and creating content online, it’s—I've kind of been joking with people, it's like, we're all YouTubers now, and we're all having to set up these home studios, and learn these tricks, like setting up pedals to control a teleprompter, or using a stream deck to have macros that will control the front parts of your screen as you're capturing things and sharing code that you're working on. These are components that are different from what would be the case normally if you could go into a studio and work with professionals who would handle a lot of that for you, or if you're presenting something in a live setting, which is just different.

Corey: One of the things that surprises me a bit is when I talked to people about what I've done, they said, “Oh yeah, after this pandemic, you're going to just stay home and do all this in your AV studio, right?: Hell no. I'm going to go back to doing what I should be doing now, and having the sense to hire professionals to do these things. Because it would be great if I didn't have to advance the teleprompter, for example, or someone else could work on the light balance. Or I don't do what happened once already, where I sit there and record a 20-minute video and then send it off for editing, and the response I got back for my video was, “That's great. But let's try one more take, and this time, maybe don't mute the microphone.” It's the, going through the iterative, dumb mistakes that everyone does. Having a team of professionals who are good at things is absolutely worth pursuing. It is worth paying for expertise, full stop because there's never enough time in the day to do everything, so being able to delegate to subject matter experts is absolutely worth doing. Sort of the whole premise is, I would argue, very directly aligned with Cloud.

Christina: I completely agree. Actually, that's a perfect analogy, you're right. It's about figuring out, focus on what you're good at and what you want to do, and letting other people handle the rest of it because doing all of it, as we're learning, is a lot. I mean, my background, I studied film and video production in college. I spent a number of years—my career before Microsoft—doing tons of stuff on camera and actually creating content.

I have a tremendous amount of video experience. I still would much rather have a professional camera people having their lighting setup, having their infrastructure, do what I'm doing because yes, I can set it up and I can probably do a pretty good job of making things work the way I want, but it takes a lot of effort, and it takes time away from doing things that I would really like to be focusing on. I would like to do some of the other things, you know? And what we've actually been doing as we've been recording things for Channel 9 or Microsoft Developer, our YouTube and online video presence for various things, is we've actually been working with remote producers, which has been really nice, where we’ll connect with somebody over Skype, and then they'll use a tool called OBS, or the Open Broadcast System, to capture and composite things, and the producer will do a lot of the compositing, and the graphical work, and then some of the editing so that the content creator doesn't have to do that, which is really nice. But there's still elements you have to do to kind of make sure okay, what does my camera framing look like? How does my audio sound? [crosstalk]—

Corey: Great video. Your fly was undone the entire time. [crosstalk]—

Christina: Yeah, which that's a mistake I've made, too. We were doing some of our free content, kind of our teaser content for Microsoft Build, and I recorded these videos and then realized that I wasn't recording audio. I mean, this was a case where I had a remote producer who was listening and was watching—so basically, they were almost in real-time, they were watching me as I was looking into my camera, and they were saying, “Okay, let's try this again.” They were giving me prompts, and they were listening. And we did this, we had two hours and thankfully it was in the first hour, it's after the first hour ended, I realized I wasn't recording my audio.

Corey: One thing I do want to call out is that before this sounds like we are just incredibly overprivileged, which we are—

Christina: Yes.

Corey: —but that's beside the point, I want to call out that in both of our cases, these are business expenses that are aimed at a goal. I mean, this podcast, for example, is sponsored. I make money by doing these podcasts. If you're trying to build a personal brand, and no one is paying you for, and you're doing it for the love of it, whatever you've got is fine. You can do this on your phone, start out and see if it's viable first.

When I started this podcast, I had a series of checkpoints—same with a newsletter and the rest—that if I hadn't hit certain goals with them, I was going to wind them down. I didn't want to be someone where I’d been running this podcast for seven years and had almost 60 whole subscribers. At that point, why bother? There needs to be at least a critical mass of audience members, and it has to resonate, it has to be something that catches on. The other podcast I do, the AWS Morning Brief for example, back before they got acquired by Cisco, ThousandEyes sponsored a twelve-week miniseries on that called, “Networking in the Cloud.” And it was more or less a network primer introduction in the time of Cloud. I thought that would be great. I'll turn to an ongoing running series. Yeah, after episode eight, I was running pretty low on the list of ideas, and at this point, I don't want to talk about it ever again because it turns out it's not interesting enough. It's not a broad enough topic from my perspective, to come up with interesting and creative conversational topics every week. So, that didn't pan out. But always have a plan. Start small, iterate forward. My first newsletter was written in Google Docs. Now I have a whole production system, but I didn't start that way.

Christina: No, I think you're exactly right. And I think, as engineers, a lot of times are impulse, I know—maybe I'm just speaking for myself, maybe I'm projecting, but I think a lot of times our impulse is to just buy the best. We read all the reviews, and we just want to go immediately to the top to the high end: I'm going to get all these things. And certainly, when I was starting out when I was making movies when I was a kid, that was the thing. Saved all this money for a Mini DV camera, and I got the best one that I could get, and I was more focused on the equipment than I was on the actual skills itself.

And the camera was great, but it doesn't matter if what you're doing with it either isn't done correctly or, to your point, if people aren't seeing it. And so, I mean, yeah, I think that if this is something that is a business expense and something you can iterate over time, you can invest more, but you're right people's phones at this point, the front-facing camera on your iPhone is better than any webcam you can get first of all, but it's also going to be in many cases, the video quality is better than a lot of DSLRs that are a couple of years old if you haven't spent a tremendous amount of money on them. And so for me, the advice I always give to people is that audio is the thing that you should look at upgrading first. Getting something like a $100 Blue Yeti USB mic for your computer can do a tremendous amount of work. If you don't have $100, spending $30 or $50 on a quality headset, or even looking at—there are some other USB condenser mics that you can get that will really improve your sound quality. That's going to help tremendously because, my perspective, video quality matters a lot, but I think audio quality is more important because a lot of us—and look I'm ADHD as well, so maybe—I don't know if you are, but I'm ADHD—

Corey: Extensively.

Christina: —I multitask all the time. I listen to things more frequently than I'm watching. You know, if I'm watching, maybe it's in a smaller window, or it's on my iPad, I've got six other things open. So, video quality is important; it's more important if I can see if you're, like, showing a screencast, okay, capture that in as high of a resolution as you can, but what that person looks like, it's good, if there's good lighting, and other stuff that makes it more compelling, but really I need to be able to hear them. So, I always say to people, yeah focus on getting the Blue Yeti mic is great, or there are some options that are even less expensive than that, or even if you're just doing regular conversations with people in meetings, business meetings, just a proper headset will go a long way. And if things become—

Corey: AirPods are terrific these days for that sort of thing.

Christina: Yeah, AirPods, I love my AirPods. I use a different microphone, but I often use my AirPods to listen to. So, I have a microphone that I might talk through, but my AirPods are fine. But you know what, when I was able to go into an office, I would often, on conference calls, use my AirPods. They're fantastic. That's a tremendous solution.

So, oftentimes you can reuse stuff you might use in other areas. But yeah, your phone is a great start. I mean, the cameras on smartphones these days are really, really good. And including the front-facing cameras. And so you don't have to spend a ton of money and go down that rabbit hole. You can, if it's something that's part of your business, and if you think that there can be value in it, and if you enjoy it, then iterate over time.

But it's not something where—I mean, this is something that I have to tell people on my team, I have to caution against them because I see [unintelligible] the company’s massive equipment list, and, like, “What do you think of all this stuff?” And I'm like, “Well, it's great, but what are you doing, and do you really need all this, and couldn't we do this for 20 percent of the cost?” And in most cases, you could. And then it's like, okay, if this becomes something you enjoy, that's when you can look at—as you have, upgrading your setup, where you’re on, now, like, version six.

Corey: One of the things that we'll do periodically is ship out a USB microphone to people.

Christina: Yeah, we do that, too.

Corey: They cost 60 bucks a pop, and it's great. It works super well. There are a few different models in that price range, depending on what's in stock right now, it's a bit harder than other times, and that works super well. But the next year is the tier that I'm recording in now, which is about 400 bucks a microphone, and it makes sense. There's another tier as well beyond this, that's around $4,000. And I have no interest in getting that because, at this point, it wouldn't make any meaningful difference until I—

Christina: No. If you’re not recording music, if you're doing voice, and you've got to think that for the sort of content that we're doing—okay, for you, it's a podcast, meaning that it's going to be compressed down to, like, 64 kilobits, maybe 32, more than likely mono. It's going to be an MP3 file that someone listens to. You do not need a $4,000 microphone. You don't because you're not recording instruments, you're not recording vocals, you don't need it. So—

Corey: And it's worse than that because before you get any of the benefit from that microphone, you have to effectively turn the room you’re recording in into a sound studio with sound deadening and very specific acoustic things. I have room noise here that I've done a fair bit of work to muffle, but this is my home office as well as my podcast recording studio, and I've always wanted to have it be comfortable to work in. I will admit now it's getting less comfortable to work in, just because of the video equipment. There are portions of the office that are no longer open for my use in pacing. I have to thread my way through microphone booms, and lighting racks, and the rest.

Christina: God, I know. Welcome to my life. I got a green screen, and I got one of the Elgatos and we had one that—

Corey: I had two screens because the first one, it turns out I now know how tall 10 feet is.

Christina: [laughs]. Oh no. Oh no. Well, see, this was a similar case for me where I had to actually—I mounted mine on the wall because my ceilings are too high, but it's one that I can pull down from rather than raise up, but I wanted to get one of those because it was like, okay, I could get the muslin, and the stands and set it up, but that's the whole thing. And I just don't have the room, frankly, to be able to tear that down and put it back up when it's needed.

You know, another point, yeah, you bring out, you have to have the right room connections, the right dampening or whatever the case may be, but you also need the right amps, and configuration and equipment to actually be able to use that. That $4,000 microphone is going to need a really expensive amp, and you're going to need to have somebody who can really understand those levels so that they can get the best out of it. And also, if you were to go into any radio station in the world, like the top tier radio stations, you usually see them on probably a $400 mic maybe—

Corey: Yeah, Joe Rogan is constantly on, I think SM7B in pictures I’ve seen.

Christina: Yeah. But I've been in iHeartRadio, I've been in the place where Ryan Seacrest does his show. He was not there, but I was interviewing Bob Pittman, and we were in Ryan's space. And I felt kind of good about myself because I looked at the microphone he was using. I don't remember the model, but it wasn't that much better than the HEiL that I have that I don't even use. And then the headphones that they use—and this is true. Any recording studio that you will go to, you just see those 7506 Sony's that have been around for 30 years. That's the standard. So, to me, if people who are doing radio and are doing these things professionally, if they're not spending $4,000 on a microphone, then you at home, absolutely have no reason to, unless you are actually doing musical recording, and that’s a completely different thing.

Corey: Yeah, Taylor Swift probably has a $4,000 microphone.

Christina: Oh, I’m sure she does. [crosstalk] does.

Corey: And you can always throw money at this. Some of the RED cameras I was looking into, it's like, “Oh, what if I just buy the best?” Well, it turns out the best starts at $25,000, and I have a lot of other things I'll buy first.

Christina: Right. And it turns out that to get the best out of that, you need to have a lot of experience. A guy that we've worked with before at Microsoft is a guy who has a RED, and is really good with it, and does a lot of work, and actually paid it off because he's really good at operating his RED, and so will do work for people using it. But if you don't know how to use that, if you don't know how to get the most out of that, that RED camera that you've spent all that money on, is not going to be of any benefit to you. It's like, look if you're MKBHD, Marques Brownlee, and your brand is to have these beautifully shot and composed kind of tech porn videos of gadget reviews, awesome, right?

But most people, that's not the case. And again, you have to think about how are people viewing this? It's going to be in a small window, maybe on a phone, maybe on an iPad, maybe just listening to the audio, and so you think about there are things you can do with your content that is better. I feel like lighting and audio are the two things that are the least expensive to really significantly improve, but are the things that have the biggest impact on keeping people engaged, making you feel comfortable because has a quality experience coming out of it. Get some sort of key light, or some sort of other major light source, and step up your microphone.

To your point, when we have guests on Rocket, a podcast that I do, we send people USB mics. We started doing that a couple of years ago because the quality otherwise just wasn't reliable, and it was such a small expense, assuming you can find things in stock, and it was such a small expense, considering how good the output could be otherwise. So, I mean, we would even do some certain Plantronics headsets if we needed to. We were really trying to budget for people and turns out great. It's a great investment to make to really improve the final product.

Corey: I would agree with everything you've said. It's strange how if someone had told me six months ago that, well I spent a lot of time this next year thinking about audio equipment, “Oh, am I going to be famous on YouTube?” No.

Christina: No. I’m just going to be on calls.

Corey: —I am not going to famous on YouTube. We're going to be trapped indoors for a year. It's like someone wished on a monkey's paw, and here we are.

Christina: Yeah, yeah. And it's so weird for me because I'm somebody who I've spent the majority of my career doing podcasts, doing video, and even for me, it's different. Doing it from home and doing it yourself, and in just the times that we're living in, so many people have made this comment that this is not normal times; this isn't normal working from home; this isn’t normal production from home. Things aren't available as much, and so we all have to adapt to the different changes. But it is very weird. You know, a year ago, I fully expected that you and I would be hanging out at Microsoft Build in person again. I'm so glad that I'm back on the podcast. I'm so glad that you are able to enjoy the event virtually, but—

Corey: You're welcome back anytime. And for some reason, apparently, people didn't take a lesson last time and invited me back to Build. I'll be in digital format this year. I imagine they'll fix that for next year. But if they don’t, I’m thrilled.

Christina: Yeah, no, I think we like having you there. You give it to us honest. You give us the good feedback that we need to hear, frankly. And I would say that for anybody who's listening, we didn't get into a lot of technical things in this conversation, but if you have feedback for us, positive or negative, let us know, or at least let me know and I can do what I can to get the feedback to the right people. I have no problem tracking people down and yelling at the right people. And we're listening. I mean, I think that's the biggest thing that all of us can take from this, is just, at least for me, it's reaffirmed how important it is to listen to people. Do a lot of talking, but it's really really important to listen.

Corey: Yes, funny thing, we talk so much about microphones and never about headphones.

Christina: Well, headphones are a whole other thing. Like I said, for audio stuff, I stick with my Sony 7506s, but when it comes to music fidelity, I have many many other opinions. But, yeah.

Corey: Yes. Which is fodder for another time.

Christina: It is. It is. Well, we could turn this into kind of an offshoot of Accidental Tech Podcast.

Corey: There we go. Yeah, no, see, I'm not cool enough to hang out with famous people. That's still you. I mean, basically, you sometimes decide to go slumming and hang out on other, lesser podcasts like this one.

Christina: No, no, no, no. And look, those guys—look, they're way more famous than me, but they're also the biggest nerds, which I say in the best way. So.

Corey: Of course. Thank you so much for taking the time to speak with me once again.

Christina: Thank you, Corey, I really, really appreciate it.

Corey: Christina Warren, senior cloud advocate at Microsoft. I am Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts, whereas if you've hated this podcast, please leave a five-star review on Apple Podcasts and a comment telling me why my audio setup is garbage.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Charles Fitzgerald

Charles Fitzgerald is a Seattle-based angel investor, with a focus on developer platforms and infrastructure. Previously, he spent 20+ years working on platform businesses at Microsoft and VMware. He can see the cloud from his house.

Links Referenced:

  • Main company site: https://www.platformonomics.com/
  • Twitter: https://twitter.com/charlesfitz/
  • LinkedIn: https://www.linkedin.com/in/charlesfitz
  • Sponsors
    • Trend Micro Cloud One
    • A Cloud Guru
    • New Relic

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is brought to you by Trend Micro Cloud One, a security services platform for organizations building in the cloud. I know you're thinking that that's a mouthful because it is, but what's easier to say? "I'm glad we have Trend Micro Cloud One, a security services platform for organizations building in the cloud" or "Hey, bad news. It's gonna be a few more weeks. I kind of forgot about that security thing." I thought so.

Trend Micro Cloud One is an automated, flexible, all-in-one solution that protects your workflows and containers with cloud native security. Identify and resolve security issues earlier in the pipeline and access your cloud environment sooner, with full visibility, so you can get back to what you do best, which is generally building great applications.

Discover Trend Micro Cloud One, a security services platform for organizations building in the cloud. Whew. At trendmicro.com/screaming.

Corey: Normally, I like to snark about the various sponsors that sponsor these episodes, but I'm faced with a bit of a challenge because this episode is sponsored in part by A Cloud Guru. They're the company that's sort of famous for teaching the world to cloud, and it's very, very hard to come up with anything meaningfully insulting about them. So, I'm not really going to try. They've recently improved their platform significantly, and it brings both the benefits of A Cloud Guru that we all know and love as well as the recently acquired Linux Academy together. That means that there's now an effective, hands-on, and comprehensive skills development platform for AWS, Azure, Google Cloud, and beyond. Yes, ‘and beyond’ is doing a lot of heavy lifting right there in that sentence. They have a bunch of new courses and labs that are available. For my purposes, they have a terrific learn by doing experience that you absolutely want to take a look at and they also have business offerings as well under ACG for Business. Check them out. Visit acloudguru.com to learn more. Tell them Corey sent you and wait for them to instinctively flinch. That's acloudguru.com.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Charles Fitzgerald, who's a Seattle-based angel investor, historically with a focus on developer platforms and infrastructure, but before that in the depths of prehistory, he spent over two decades working on platform businesses at Microsoft and VMware and, as his bio states, he is very proud that he can see the Cloud from his house. Charles, welcome to the show.

Charles: Thanks for having me.

Corey: Thanks for being had. So, you run the platformonomics.com website, which it’s appropriately named after your company, Platformonomics. What do you do?

Charles: Well, my usual answer to that question is, “Hey, I'm on Twitter,” which leaves people dumbfounded. It does leave me some spare time, and it's probably worth going through my background. I joined Microsoft out of school a long time ago, and basically worked on every platform effort there from 16-bit Windows, to 32-bit Windows, to COM, OLE, ActiveX, served in the trenches during the browser wars and was one of the very first people to work on .NET. So, I've seen a lot of iterations of platform gambits in the marketplace.

Left Microsoft end of 2007, kind of dorked around with a startup, ended up at VMware for about two and a half years, where I helped launch Cloud Foundry, and by about 2013, I was kind of done working for the man. And since then, I have been angel investing primarily in Seattle, my nominal focus is on developer infrastructure and platforms, but done lots of investments in other stuff from FinTech to crypto, to VR, a lot of boring SaaS companies will pretty much look at anything. And during that time, I've also, kind of, couldn't but help myself watch the evolution of the cloud race. I joke I can see the Cloud from my house because I live in Seattle and there are plenty of clouds to be seen. You know, if I twist my head and open the blinds, I could probably see Amazon right now, and continue to have flashbacks from Microsoft. So, I have a pretty good perspective on what's going on in the Cloud. So, to the degree I do anything, that's what I do.

Corey: So, the way that I feel became familiar with your work is when I started my consulting company three and a half years ago it was, “All right, I have a background as an AWS architect/systems engineer type. What expensive problem can I align this around?” And the AWS bill was the easy answer. So, I started doing some research, as some people apparently don't do.

And I was, “What do I call myself? ‘Cloud Economist’ sounds like the right direction to go in because it's two words no one understands.’ ‘cloud’ is a bunch of other people's computers, kind of, and ‘economist’ is someone who claims to know about money, but dresses like a flood victim. Smash it together, and no one knows what the hell I'm talking about.” But when I was doing some research to see, does this actually exist? Platformonomics came across my radar, and it was one of those things where, “Huh, I’d better stay away from this guy. He might actually be able to call me on my nonsense.” So, I was scared of you for the first couple of years until I started reading more of what you've done. And the more I read, the more I like.

Charles: Yeah, familiarity just undermines that intimidation factor completely.

Corey: Oh, absolutely. It's bred an awful lot of contempt. It's been spectacular, but one of the things that I found that you talk about, along with basically nobody else, has been a recurring series that you're calling Follow the CAPEX. What is that?

Charles: CAPEX is capital expenditures. And the origin story, I was—I mean, this was probably 15 years ago, I was still at Microsoft, I was sitting in a meeting where somebody was reviewing Microsoft's efforts to build data centers. And if you grew up in the software business, and ‘grew up’ may be a strong term in my case, CAPEX was, “Hey, we have some buildings, we have some desks, some new laptops.” That was about it. There really was very—it's a very capital-like kind of business.

And so, sitting in this presentation, and listening to people talk about there's concrete, there’s steel, there are vast racks of hardware. This is before the big cloud companies started doing their own transoceanic cables, but I was just fascinated because it was so counter to what a software company had traditionally done. And then in the early days of the Cloud, I'm sitting here watching everybody posture and tell us how great they're going to be, and I would just look at the CAPEX numbers. And the numbers are just, they're obscene. I mean, I have been updating these every year for a few years now, and if you look at Amazon, Microsoft, Google, they've spent almost $350 billion since the year 2000. We've gone from this incredibly asset-light industry to the cloud vendors are now on par with the biggest CAPEX vendors in the world: the energy companies, the telcos.

And that's not all cloud, for sure. Google is mostly search and YouTube. GCP turns out to be a very small tail on what is a very large infrastructure dog. I used to joke that there must be a space elevator in there as well. That was, kind of, the acceptable societal trade-off for accepting the ad surveillance network. Doesn't look like we're going to get that now. If you look at Amazon, and Amazon's numbers have just gone absolutely ballistic in the last couple of years.

They spend on AWS, it's probably a fraction of their overall spend. They spend on robots, and trucks, and planes, and warehouses. And now, they really do have a space program as they start to do this internet satellite constellation. Microsoft has probably the cleanest CAPEX in that they don't have a whole bunch of those other extracurriculars, but it's just an obscene amount of money and turns out to be a huge barrier to entry. I mean, imagine somebody gave you $100 billion and said, “Go try to catch these companies. Good luck.”

Corey: That wouldn't be enough.

Charles: It would take a long time. So—

Corey: All three of these companies had other businesses that financed the build-out of their cloud environments. No one has done a, ‘we're going to build a cloud,’ and succeeded at a hyper-scale way. There's always been something else, be ad sales, e-commerce or giant software ecosystem that has financed this build-out.

Charles: Yeah. You definitely had to have a big high margin, big balance sheet in order to go do that. So, anyway, I'm sitting here watching this stuff, and then you start to look at the rhetoric coming out of different companies and some companies, the CAPEX was up and to the right, and you could be pretty optimistic about them. Some companies were talking a big game, but their CAPEX was actually declining. And at some point, I decided I would separate these different vendors into, we have the ‘clouds’ and the ‘clowns’ and that resulted in a bunch of interesting interactions with some of the clowns.

Corey: Oh, yeah. And one of the hard parts about this too is all right, you're going to throw enormous piles of CAPEX at your cloud buildout. Great; that's not going to pay dividends for three to five years at the earliest. It isn't something that you can flip a switch and you're good to go. It takes a lot of time to build these things.

Charles: I mean, one of the other lenses on these companies is some of them are engineering engineering companies, some of them turn out to be financial engineering companies and either weren't willing to spend the money on concrete and servers and fiber and everything else, or they really didn't have the money. You know, the dividend was far more important to them than investing in the next round. And IBM in particular is the poster child there. I wrote a post in 2013, arguing it doesn't look like IBM was going to make the cloud transition, and the post was pretty well-timed. I think I missed IBM's all-time stock high by about 16 days, and they proceeded to reel off—what was it 24—straight quarters of revenue declines.

But I had all these IBM people coming at me waves, they kind of made three arguments. One, we're IBM; we always make technology transitions. And my response to that was, “Well, you actually have to do something if you want to make that transition.”

Argument number two was, “Hey, Wall Street loves us.” Well, it got to the point where even Wall Street was telling IBM not to focus on the financial goals, they need to focus on their survival.

And then the third argument was, “Well, at least we're not HP,” and I had to give them that. But I've sort of instituted a policy where if you don't say ridiculous things, I'll leave you alone, but if you keep saying ridiculous things, I'm going to call you on it. So, it's led to a very interesting set of interactions with IBM over the years, where they continue to say preposterous things about their relevance to the cloud space. And that's in contrast to somebody like Oracle, where Oracle talked some ridiculous things, but their CAPEX is—I mean, it got to the point where the big cloud guys were spending more on CAPEX in a quarter than Oracle had—or maybe was more in a year than Oracle had in the entire history of the company. But they've kind of stopped saying ridiculous things. Larry occasionally slips out, but they've dialed it way down.

Corey: They hunt him down rather effective. They've gotten much more efficient at it.

Charles: Yeah. You could tell they're there about three straight quarterly earnings calls where they did not mention cloud. Was very clear they'd gotten the discipline about what they were going to say and what they weren't going to say. So, it's turned out to be a super useful lens where you just—who's spending big on data centers, and transoceanic cables and millions and millions of servers, they're probably going to do all right if they can meld that to software, and go to market, and everything else. And the people who are talking a good game and issuing press releases and not spending any money, they may have a couple of closets here and there with a couple of racks in them, but that's the furthest thing from a cloud. So, it turned out to be a great lens. I mean, I karmically shorted IBM; I wish I'd shorted them because they dropped a third of their value and have just been a financial disaster over the last 10 years.

Corey: One of the hard things for me is, as difficult as it is to interpret all of AWS’s 200-some-odd services and put them in their proper place and keep track at all, that is a trivial easy exercise compared to reading their financial statements and making sense of them, just because you want to talk about real engineering versus financial engineering, it feels like they're heavily involved in both where they're building out data centers by financing it through both capital leases as well as these bespoke custom lease instruments that really obscure what's going on and how. It does all kinds of weird things to their free cash flow that really lets them plow money into this at an unprecedented rate. And it's also super hard to disambiguate this across different business units as well, in some cases. So, it’s like, “Is that an office building that they're spending all that money on? Is it data center regions? How does that break down?” And at the end of it, I basically shrug, give up, and leave it to you.

Charles: Don't forget the big biospheres. I mean, the things they're spending money on—I mean, they have a very clear strategy to take all that free cash flow and reinvest it in something, and they've kind of run out of things to invest in. I mean, a satellite constellation may be the last gasp after hundreds of thousands of trucks, and a fleet of airplanes, and distribution centers in every city, and convenience stores on every corner. But yeah, I mean, on one hand, I think they've been smarter about leveraging their balance sheets, they actually had to call out and explain to people that they were using those capital leases because people were not aware of them.

Corey: Yep.

Charles: It was just another complicated footnote in the financial statements. But I'm pretty sure that those capital leases are the server investment behind AWS, if you look at the lifespan, it really correlates to the useful life of the server.

Corey: And they've also been talking on earnings that they've been increasing those lately.

Charles: Yeah, they just bumped it up to four years, and the amount of new capital lease has actually flattened out pretty dramatically, right about the same time. So—

Corey: Yeah, you take a look at their two generations back systems, no one in their right mind is going to be building net new workloads on top of those, but if you look CPU capabilities and do a one-to-one comparison, ah, this is almost definitely what things like Fargate are running on top of and presumably Lambda as well, so repurposing things like that from an engineering perspective where you don't guarantee specific performance characteristics is a great way of doing that, and it's clearly working. Those services are growing like wildfire. It's a brilliant play.

Charles: Yeah. Being able to reuse that stuff and get more value out of it for as long as you can, is obviously super key. One of my projects for the summer—we'll see if I get around to it—is to try to put some scope around just how big AWS is. So, from a how many servers, how many cores, how much storage? I’ve sort of got some ideas on ways to go after that. So, if I get motivated, I'm hoping to tease that apart because AWS is definitely a small part of the overall Amazon CAPEX spend, but I have some theories on how to triangulate it.

Corey: I would be very interested in hearing what you come up with on that front.

Charles: As would I. [laughs].

Corey: One thing that I've noticed among cloud providers, everyone looks for superlatives to some extent. And Microsoft’s strategy, which I can't help but see as a blunder, has been to focus on having the most global regions, and very proud of that fact. They’re trumpeting that during Build this year: Satya Nadella called that out explicitly. And yeah, it's important to have regions near where your customers are, both for performance reasons, as well as for regulatory purposes, but by doing that, rather than focusing on fewer, more robust regions, instead going with a more far-flung approach, the capacity of each one of those regions is not able to absorb sudden influxes, such as, oh, I don't know a global pandemic changes the way that everyone interacts with their cloud services. And there were serious capacity shortfalls for weeks as a direct result of this. That is my analysis of the situation. Do you think that that's spot-on, completely misses the mark, or something else?

Charles: They’ve definitely had some capacity issues, and I think all the big cloud vendors I'm sure are furiously scurrying, given the disruptions on the supply chain side, combined with increases in demand. I think it's a good indication that Microsoft has a different set of priorities than Amazon or Google. Microsoft is much closer to those enterprise customers, it's those regulatory demands where customers say, “For data sovereignty reasons, we've got to store our stuff in the same country.” And Microsoft has been more responsive to that. Contrast that to the Google model, which is far fewer data centers that are of enormous capacity, but don't align well with those sovereignty needs. Microsoft's probably at the other end of that extreme, and AWS is somewhere in between, but all of those companies are going to continue to build out their local capacity over time and hopefully, the pandemic will recede and it'll be easier to do your capacity planning going forward.

Corey: Oh yeah. At some level, I'm reminded, for some reason, of the dot-com bust where everyone was building up in the dot-com bubble, they're convinced that this is going to be the way and the light of the future. And it was, just not for a bunch of those companies, and one thing that we saw happening was the telcos laying enormous quantities of fiber, assuming it was going to go up and up forever, and what they inadvertently wound up doing—from a certain point of view—is investing their own way out of a business model, where suddenly there is so much capacity for high-speed transmission and the rest, long-distance rates plunged to effectively zero. And suddenly you're seeing this whole world of, “Oh. We’re the provider of pipe, and suddenly there's way more pipe than anyone is going to directly need. We can't charge a premium for that anymore.” I'm wondering if we're going to see something like that this time around.

Charles: You didn't have Moore's Law type phenomenon on the telco side, right? I mean, a quality phone call takes about the same amount of bandwidth, and so it's really kind of more Moore’s Law and—

Corey: But it takes way more bandwidth to get a crappy one through a whole bunch of things. WebEx looking at you.

Charles: Yeah. It's been fun dealing with the network issues in the last month or two, I think we're going to continue to see the demand curve is different. The fact that you had lots and lots of fiber didn't result in you making a lot more phone calls. But as compute gets more and more pervasive, people can do more and more things. What we've seen is that demand has continued to go up.

The other thing that's going on with Cloud is we still have a huge amount of substitution to do. I mean, one of the things I'm hearing in the last couple months is a lot of the last holdouts, the people who were very proud of their on-prem data centers and continued to talk about their security concerns with the Cloud, a lot of them have suddenly decided maybe it's time to start to move, and to the degree that sub 10 percent of workloads are in the Cloud or whatever your favorite number for that particular metric is, there's still a phenomenal amount of upside. And I think we're going to continue to find new things to do with compute cycles for the foreseeable future. So, I'm not too worried about that. I think it's a different dynamic.

Corey: I think that it's easy to predict the future after it's already happened.

Charles: Indeed.

Corey: It's one of those stories of, “Oh, well, that's what happened last time, so, well, obviously, that's going to happen again.” As they say, “Economists have successfully predicted five out of the last three recessions.” The challenge, of course, is that history doesn't repeat. It rhymes, but it doesn't repeat. What do you think's going to be potentially different this time around?

Charles: Yeah, I've spent a bunch of time thinking about that. And we really want to go back to the models of what happened after the global financial crisis, or what happened after the dot-com bubble collapse. And some of those elements—I mean, I tend to focus on early-stage investment ecosystems, and we've been living in this world for the last five-plus years where zero interest rates have pushed people to look for more yield. It's pushed a lot of money into other asset classes, more speculative asset classes, including venture.

So, we've had all kinds of crazy money floating around. I mean, it's just insane that I'm a little angel investor in Seattle, and I'm running into sovereign wealth funds that are traipsing around in the same spaces that I am. So, one of the big questions is, a lot of this money's been allocated big venture funds, sovereign wealth funds, all the hedge funds, or many of the hedge funds have gotten into venture, you've got corporate venture, what's going to happen to these big pools of money? Are they going to dry up and go away? Are they still allocated?

If you've raised a venture fund, you've got the money, you probably need to go allocate it. And one of the things that I noticed in 2000 is everybody assumed things would dry up immediately. And what we saw is that people continued to invest like it was 1999 for the first year. So, I think we're going to have to wait until some of these investors need to go back and raise their next fund, and I think it's going to be a lot tougher for them to raise it. I think we still are going to see a lot of money sloshing around in the near term.

You continue to see all these headlines about companies raising big rounds, tens of millions of dollar rounds. So, the money's still out there for companies that have the right story. How many of those deals were done, pre-pandemic and are just now getting announced so that people can look like they're still active, I don't know. But figuring out what happens next, which segments, which opportunities face very real headwinds, which pick up tailwinds is a super interesting question, and everybody's trying to figure that out right now. Generally, I'm not a consumer investor.

So, I think I've got probably one company of the set of companies I've invested in that has had real headwinds, and they've done a phenomenal job pivoting, but most of the others companies I've got are either net neutral, or in some cases, they've picked up a tailwind in this environment, which is great to see. But I think it's going to be a long trek to get back to normal. So, we'll see if—you look at the stock market today, and it's kind of oblivious to the idea that there's anything going on in the public health world. To the degree it takes us a really long time to get to the other side of that, things may get a lot more dire.

Corey: In what you might be forgiven for mistaking for a blast from the past, today I want to talk about New Relic. They seem to be a relatively legacy monitoring company, and I would have agreed with that assessment up until relatively recently. But they did something a little out there: they reworked everything. They went open source, they made it so you can monitor your whole stack in one place and, most notably from my perspective, they simplified their pricing into something that is much more affordable for almost everyone. There's even a free tier with one user and 100 gigs per month, totally free. Check it out at newrelic.com.

Corey: Speaking of the healthcare world, which is always something interesting to cover, one announcement at Build this year was that Azure was going to have a offering focused specifically for the healthcare market. If AWS were doing something like that—targeting a specific industry vertical, I would make fun of that until I was blue in the face because it would be ham-fisted and the complete wrong direction to go in because AWS is terrible at this. What stops me from making that assessment about Microsoft doing this is that they have 40 years of history in being phenomenally successful at targeting specific industry niches and verticals. What's your take on that move?

Charles: Yeah, I saw that headline. And I thought, “Oh boy, this is going to be another niche cloud.” IBM and Bank of America announced a financial services cloud and never got beyond the press release. It doesn't look like it's a separate cloud. It just looks like it's a set of solutions that run on top of Azure, which makes perfect sense given that 40 years of experience working with different enterprises, different industries, and they’re probably are lots of things that have a greater urgency in healthcare to get to the Cloud than they did three months ago: telemedicine, drug discovery, whatever it may be, so I think that makes sense.

But I'm surprised we don't see more niche clouds. As I think about where the opportunities are going forward, you've got these, sometimes I describe it as the three big hyper-clouds, sometimes I talk about the two and a half big hyper-clouds—and maybe we can talk about GCP and whether they're serious about this in a minute. But normally, you'd expect to see a power-law distribution. And so we've got these three huge players. And then there's kind of nothing, right?

You've got a few little companies that they’re regional hosters who may be a little bit ambitious in describing themselves as clouds, you've got some people who cater to cost-conscious developers, and that's never a great way to build enough margin to invest. But I would expect to see more niche cloud offerings, and they really just don't exist today. So, it's something that I’m—definitely have my ear to the ground on as I look at what's happening next in the cloud world.

Corey: One thing that is interesting, too, that I see in these global market share analyses is folks are saying that Alibaba has surpassed Google for number three. In fact, they just announced this year a $28 billion investment in their cloud offering. But by everything I can see, all of that is focused on Mainland China and servicing businesses with significant presences there, I tried to sign up for their cloud and give it a test run, and I could not get that far because it wanted me to upload a whole bunch of documentation in order to qualify for a couple hundred bucks in credits to test stuff out, and all I really wanted to do was spin-off a Linux VM for half an hour and play around with it, and kick the tires. It was very clearly not aimed at a business that is not used to the Chinese market. And I'm curious as to see, do you think that that is likely to change? Do you think that there's going to be a breakout of Alibaba Cloud into the broader cloud ecosystem, or is it going to remain a—I hesitate to say ‘niche player’ because the Chinese market is still over a billion people, but I'm curious to hear your take.

Charles: I really don't think we're going to see any new hyper-scale challengers to AWS, Azure, and GCP. They have such huge moats, they have such huge investment, and if you look at companies that have the resources and the skills to do it—maybe an Apple, maybe a Facebook—they really don't have any interest in serving an enterprise customer and you can't play in this space if you're not going to go after the enterprise customer because that's where the majority of the dollars are. So, I don't think we're going to see generally new hyper-cloud competitors. I don't think we're going to see the Chinese hyper-clouds get much traction outside of China or places where China has a lot of sway. This deglobalization trend that was happening prior to the pandemic has really only been accelerated by it.

So, I think we're going to see a world where the big hyper-clouds from the west are one set of vendors, and then we're going to see a different set of vendors inside China. I don't think the Western companies are going to make progress inside of China and vice versa. I tried to dig into Alibaba’s CAPEX spend a couple years ago, and I discovered that if you're a Chinese company and you're listed on American stock exchanges, you don't actually have to report any financial numbers. I mean, it's really crazy.

Corey: That sounds like an interesting edge case that is not broadly known.

Charles: Yeah. So, if only American companies could get the same treatment. Financials? Eh, who needs them?

Corey: Yeah, just take our word for it. It's fine.

Charles: Yeah. It's pretty tough to see what's going on there. Somebody told me that they have recently gone back and started to report—they’re unaudited numbers, but they're more detailed numbers, so I was going to go take a look at those and try to get a sense of which of their CAPEX expenses might be much cheaper. They're probably paying less for concrete and labor, their Intel processors, I'm sure they're still paying the same prices as anybody else. So, that's another thing that's kind of on my list to go look at, but last time I looked at it, it was very, very difficult to make any headway on what's really going on there and how big they are.

Corey: It's sort of a giant open question mark. And I see these analyst reports saying that they're number three there, but I've got to say, I've never seen them in the wild. As I wind up walking through the world and talking to big customers. There was one company I talked to that was global and did a small workload there, but they were running on everything. I mean, I certainly see Azure, I see GCP a lot, too, and then there's a couple of weird exception edge cases.

I’m increasingly starting to see Oracle Cloud, of all things, around data transfer deals, where maybe having a retail price that is not eight times as expensive is starting to win those deals over. And yes, of course, everyone at size pays discounted rates on every provider, but when the list price is an order of magnitude or so larger than someone else's, people aren't even going to call and request a quote past a certain point.

Charles: Yeah, it's funny to think of Oracle as the low-cost leader.

Corey: Oh my god, I've always operated under the assumption that if I recommend—“You should pick Oracle,” for an economic story, great, that would get me thrown out of the room, possibly via the window. But it is economical in this case.

Charles: They're in their ‘we try harder’ phase. We'll see how long that lasts.

Corey: Yeah, number two tries harder. And yeah, Oracle's has a lot of number two, historically.

Charles: Number four, or five. But, who's counting?

Corey: One thing that is fascinating to me is that it's clear that they're making some aggressive moves in the place. They've hired some amazing executives out of AWS—and other places—and given them free rein. The biggest weakness I see an Oracle Cloud today is the word Oracle at the front of it, but I would not necessarily count them out. The problem is there's so much bad blood around Oracle that there has to be a massive reputational and perception change. I would say it's impossible except for the fact that Microsoft did it. Where’s their Satya Nadella moment? It [unintelligible] require having Larry Ellison sent to a home, but that's neither here nor there.

Charles: It's more than Larry. If you're a distant challenger in the market, you really need customers who are willing to take a chance on you and make a bet on you, and it's hard to see a lot of people are willing to do that with Oracle. They have hired a lot of people up here in Seattle. This is their second, third, fourth go around on that. I'm not sure that strategy has been super successful for them, the kind of people they've hired, they've been more motivated by big salaries and big titles, necessarily, than building something great. And so I don't think they're going to be able to hire up a bunch of people and successfully challenge the big clouds because it's just too late in the game. They needed to do it 10 years ago when Larry was still poo-pooing Cloud.

Corey: I don't know, I mostly try and stay out of the predict the future market, which I know makes me something of a rarity, but the reason behind it was always that my first year doing this stuff in a ‘predict what's going to be released at re:Invent’ capacity was also the last year I did it because I remember the day before that newsletter went out that I had a customer who, through their TAM was getting roadmap emails of selected things that were coming out, and I double-checked that, and it turns out, I’d gotten three or four things, right, so I had to quickly go through and remove it, and I realized the risk was too high because I had access to inside information from time to time, and when I got it right, no one would have believed that I didn't break NDA. So, it just easier for me to stay away from forward-looking speculative things around specific services or products.

Charles: Yeah. I mean, you don't have to look at the specifics. You just have to look at the magnitude of the gap that exists between—I mean, this gap is true. AWS is probably twice the size of Azure, Azure is growing 50 percent plus faster. The gap between Azure and GCP is probably 4x, and we don't know what the GCP growth numbers are, but they're probably growing marginally faster than Azure, and arithmetically it's extraordinarily hard to paint a scenario where they can catch up in any interesting timeframe.

If you then go look at some of the also-rans, the Oracles, the IBMs, they need to be growing many, many, many times faster than the existing leaders in order to catch up because they're starting from such a small base, and it's just impossible to pencil out the math that makes that work. And you start to look at cumulative CAPEX spend, you start to look at number of employees, you start to look at portfolio of services, it's really, really hard to catch up. I mean, even today, GCP is super challenged. I mean, they have lots of talent, they have lots of money, they have phenomenal infrastructure, and they're really having a hard time keeping up with AWS and Azure. And in the end, I think it comes down to the fact that GCP is a hobby, and for the other two companies, AWS and Azure are existential for those two companies.

Corey: The challenge, too, is that there's a very much a virtuous flywheel effect here, where when you're one of the big providers you are drawing customers in, who are A) asking for new and creative things that give you ideas to explore next in ways that satisfy their needs, and they're paying you an absolute mountain of money, so you can afford to make those investments in building those things out. Whereas, if you're trying to enter the stage as a cloud provider, you're still struggling with operational concerns that the big players have already mastered. Things like, “How do we run Linux VMs in various parts of the world reliably?” There's a gulf that's only getting larger. I'm sort of worried on some level that Google is completely going to screw this up, and Microsoft is going to remain relegated to the enterprise world leaving us, more or less, in a world of an Amazonian monoculture, and I don't want that.

Charles: Yeah, I think we're going to see Amazon continues to be the leader, I think Microsoft's going to be a strong challenger. I mean, the thing to remember is, as the enterprise opens up, an awful lot of the available money for the cloud market is on the enterprise side, and that explains why Microsoft continues to outgrow AWS. They're just a lot more mature on the enterprise side. The challenge for Google is Sundar wakes up in the morning and his Google Cloud, all up—not even just GCP, but combination of G Suite and GCP is, at best, the fifth-largest business.

Corey: And it's the first largest business, as far as something that is not directly or indirectly driven by advertising.

Charles: Yeah. The other four businesses above it on the list are nicely intertwined in terms of who the customer is, the skill sets, the talent inside the company. I described GCP as it’s a hobby, and that annoys my friends at Google, but that's the reality. And we'll see what they decide to do with it.

Corey: Yeah, I'm curious to see what happens. Thank you so much for taking the time to speak with me today.

Charles: It's been fun.

Corey: It really has. If people want to hear more about what you have to say and how you say it, where can they find you?

Charles: Well, I'm on Twitter at @Charlesfitz, charles-F-I-T-Z. And my website is www.platformonomics.com, which I sort of treat as long-form Twitter. You'll get really long posts from me, very sporadically.

Corey: Well, thank you so much, once again, for taking the time to speak with me. I've really enjoyed this.

Charles: You bet. It's been fun.

Corey: Charles Fitzgerald, managing director at Platformonomics. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts, whereas if you've hated this podcast, please leave a five-star review on Apple Podcasts along with a comment detailing the exact CAPEX breakdown of AWS’s last earnings release.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Bill Staples
Bill Staples is New Relic’s chief product officer, responsible for driving the company’s market-leading platform strategy and leading New Relic’s product management, engineering and design functions. Staples is an execution-focused product and engineering leader who loves to build and scale cloud-based businesses. Previously Staples was at Adobe, where he led the 1,500+ employee global engineering team behind Adobe’s market-leading Experience Cloud. Prior to Adobe, Staples spent 17 years at Microsoft, including his last role as vice president of Microsoft’s Application platform. At both Microsoft and Adobe, he successfully led transformative product, culture and technical innovation agendas, helping both companies expand multi-billion dollar cloud portfolios with developers and IT as the customer.

Links Referenced

  • Main company site: https://newrelic.com/
  • Twitter: https://twitter.com/bstaples
  • LinkedIn: https://www.linkedin.com/in/williamstaples/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Bill Staples, New Relic’s chief product officer. Bill, welcome to the show.

Bill: Hey, thanks, Corey. It's great to be here.

Corey: So, before this, you were over at Adobe, and before that, you spent a brief 17-year stint at a small company called Microsoft.

Bill: Brief 17 years, that's correct.

Corey: And now you've been at New Relic for a whopping six months, which in technology terms is close to about seven years, and given this whole pandemic era that we're living in, it kind of feels like it's close to that even without the joke.

Bill: It's true. I mean, starting to not feel like the newbie anymore here at New Relic. It's been an awesome six months though. Wow, what a ride.

Corey: New Relic is one of those interesting companies in that it feels like—at least from my perspective—that it's been around forever. It's a product that I used, at first begrudgingly, and then happily, and then I, you know, like any shiny thing, I wound up doing my raccoon act: seeing something shiny somewhere else and wandering off into the weeds. It was a really innovative product for a while and it seemed that there was nothing else in the market quite like it, and what surprised me was how long it seemed that that lasted.

Bill: Innovative product for a while, Corey. We got to get you back into New Relic. We need to win you back.

Corey: That's the problem, is I still have this mental model of New Relic being, “Oh, that's the thing you use for application performance monitoring.” “What is that?” “Well, it's a buzzword, but what it really means in practice is you open up a webpage, and it tells you why your website is loading slowly, or your application isn't working anymore.”

Bill: Right.

Corey: That has always been the way I have contextualized New Relic. One of the things I see across the industry, though, is that we have these—you have companies like DataDog, for example, that have a monitoring strategy of, yes. Whatever it is that you want to do in the world of monitoring that you could possibly do, that's what they do. And every company, for whatever reason, in the monitoring space seems to start off as doing something that is niche. With New Relic, it was the application performance monitoring piece, a bunch of other companies are focusing on the log analytics piece, other folks are doing synthetics. There's the idea of the Pingdom style of ‘is the site up, or down, or not?’ but over time, it seems that every monitoring company tends to expand into a, I guess, a broad-based platform of, “Yes, we're going to start doing all of those things as well, and spreading out from the core competency.” Why is that?

Bill: Well, therein lies the problem. I mean, the reason for that is engineers choose the best tool for the job. And increasingly, the world has gotten more complex. Like the use case you described of, you deploy an application, you want to see what's going on, so you deploy New Relic agent, open up a webpage and see if it's running. That was so avant-garde, 2008.

In today's world, you've got, increasingly, complexity at the infrastructure layer, with [00:04:39 finally], you know, virtual machines and infrastructure as a service from public cloud providers, but container-based systems like Kubernetes. You've got increasingly number of digital clients. Not just your mobile phone, but TVs, and other wearable computing, and in-store advertising, and all kinds of devices that deliver an experience to you. And then in between, you've got this explosion of distributed services, services calling services, inside your company and outside your company. And whether the application or, I guess, the customer experience is being served or not is not a simple question of is the application up and running, anymore.

It's a question of all those layers working in concert to deliver the end-user experience. And you need a more sophisticated approach to understand is it up and running. And so increasingly, as you said, vendors and open-source systems are striving to expand the capability set to cover that full-stack view. And that's the journey New Relic’s been on, and that's what we're doing with the New Relic One platform and the announcements we made last week.

Corey: Credit where due, the entire world has changed. When you have a system running on a single computer. It's pretty easy. Is it up or down? At some point that narrative changes, and as we look into this world that we live in today of distributed applications, it's not, “Is the service up or is the service down?” It’s, “how down is it?” And there's a monitoring evolution. I still tend to live in a world where you could say, okay, users are complaining, what is the monitoring story? And my honest answer is, what do you mean, you just told me my monitoring story.

Bill: [laughs]. Yeah, exactly. And the other factor that's, I think, playing into it is developers, they want to move faster, they want to innovate, and with that complexity, the reluctance, sometimes, on behalf of companies to make changes in productions, and put lots of barriers to ensure that the customer experience doesn't regress, slows down innovation. So, it's as much about dealing with the complexity of the application and experience architecture, as it is the need to unlock developers, and moving fast, and being agile. Again, the only way you can do that as if you measure everything, understand how things are performing now, and then understand how they appear to be performing after, say, a deployment, with the ability to roll back if things aren't going well. So, both factors play a role in terms of companies looking harder at monitoring, both to deal with the complexity and to speed up or accelerate innovation.

Corey: I've got to say, one of the challenges I had with New Relic for a long time was, I was at a couple of different companies now where there was a fixed contract price for the year, things were great. And then, surprise apps do what they do: they get bigger because it turns out that your infrastructure footprint and, by extension, your AWS bill is never a function of how many customers you have, it's a function of how many engineers you employ. And it seems like there's a natural state of truth to that problem. And midway through the cycle, it was, oh, you have the salespeople reaching out with, hey, you're using more than we anticipated. Let's settle up again mid-cycle, which was just a terrible experience from where I sit.

It's “oh, great. Now I get to go and tell my boss that A) I'm bad at managing vendors, B) I need to get more money,” which is always a compelling story to tell someone who has already done their own budget. And it basically led to a lot of unfortunate outcomes as a direct result. One thing that I find really interesting about what you folks are doing this year is you've completely revamped the pricing model in such a way that even your messaging around this is pretty much burning the boats behind you. What led to that? I think it's something I haven't seen much of in this industry for a while, and I'm actually really excited about it.

Bill: Yeah, I mean, you can disrupt lots of ways and we're disrupting ourselves from a messaging standpoint. I think I've said, “As the company who invented or created APM, we're going to be the ones to end it.” We no longer sell our APM product as an independent product anymore. With the announcements last week, we established a new packaging and pricing structure for the New Relic One platform, and the intention there is to just make it simpler for companies to adopt observability. The number one problem as I spoke with customers and their ability to take advantage of observability, is the complexity and the cost of it, frankly.

They're left with this expanding set of tools and vendors to deal with, and even from New Relic. You could look at any of our competitors, it's the same thing. Up until a week ago, we offered a dozen-plus different products with lots of different pricing meters. And any VP of engineering or CIO, who looked at that and tried to predict, “Geez, how is my infrastructure going to expand across all of these dimensions, and how can I forecast the budget to ensure that my developers have the best tools and are covered in what they need to do?” It was almost impossible. So, it led exactly to what you described in terms of surprise bills. And that led then to companies deciding, hey, we can't afford to monitor everything. Let's only deploy agents on the most critical layers of the application stack.

Corey: Or you start sampling, like, okay, one in every 10 servers is going to have this on it. And that gives us a rough statistical overview of these things.

Bill: But then that leads to blind spots, and it leads to surprises. And no engineer wants to wake up in the middle of the night because, oh, that service didn't get instrumented, or, oh, there's an outage in part of the production degradation service, and part of the production cluster that didn't have the monitoring agents until it was caught too late. And so the changes we made are directly in response to that. We believe everything should be instrumented, and the groundbreaking product we introduced last week is telemetry data platform. It's been our back end technology for a lot of the applications that we've built up over the years, but we're now exposing it as a first-class platform. It handles metrics, events, traces, and logs.

We've open-sourced all of our agents and integrations. So, we're doing those roadmaps now in the open and accepting code contributions from the community to make it easier and richer than ever to capture telemetry from everywhere. We're embracing open-source systems. So, we've announced as well support for Prometheus, and its remote write capability directly into this telemetry data platform so you can get extended retention, enhance security, increase scale of your Prometheus systems, PromQL as a query language on top of it, so you can ask questions of that metrics data and use Grafana dashboards on top of it, all for incredibly low price. The multi-tenant nature of our architecture and the high scale of it allows us to pass those savings on to customers so that they truly can instrument everything now in the telemetry data platform across data types, and visualize and analyze that data in ways they've never been able to before, functionally or because of the economic barriers.

And it's in direct response to those challenges you mentioned. It shouldn't be so hard to instrument, your digital infrastructure, and be able to get answers to questions that you didn't even think about when you were building the system. I mean, that's the major shift between monitoring and observability in my view, is monitoring you deploy after the fact to detect when something goes wrong, observability you instrument upfront so that you can understand why something's happening as it's happening, or even investigate and understand how the system’s performing in ways that you didn't think about given the complexity when you were building it.

Corey: One of the things that struck me when I looked into a demo of New Relic about a year ago was when New Relic One was first announced, it ties directly to what you're saying. I ran into enough challenges along the way when I was writing my review that I eventually tabled the thing because, and I don't say this lightly, it was to mean. I tend to have a bit of a brand for punching up at publicly traded companies when they sell things that I don't like. And it was just a—to be blunt—terrible experience start to finish. It was confusing, the documentation on how to install it on the serverless side disagreed with itself, talking to the support folks to get up and running, it seemed that some of them didn't know that this product existed, and it was a very disjointed story.

I knew what it was going to cost at least because it was a fixed fee, but, A) it was a lot of money, and B) the annoying part is after the demo was done, it took my normally $2 serverless app and add another $40 in AWS charges on top of that for the way that it was set up to pull the monitoring and all the rest. I mean, again, it was this is not the end of the world stuff, but it was frustrating, and it led to a point of, “Oh, I guess yeah, I'll just continue to consider New Relic from an APM perspective and move on to other things in the rest of the world.” Now you've relaunched the telemetry platform it’s, okay, this is unifying this in a way that how I try to view a monitoring product from a vendor—monitoring offering from a vendor—and now it's really coming across as, “Oh, okay. This is a cohesive product where someone has taken the bold step of introducing various product managers to each other.” And now you finally have relaxed the NDA that prevented internal service teams from speaking to one another.

Now it feels like there really is a sense of this is an actual product offering that is aimed at user problems, not a component by component microservices product offering strategy that each one goes into different bits and pieces. I mean, I was, I've got to say, a little concerned at the IOpipe acquisition on those grounds of, “Oh, dear Lord, this is going to be yet another attempt at a second or third Lambda service monitoring system.” No, I'm seeing integration stories coming out of this, I really should have given those folks more credit. They're great. I was a paying IOpipe customer, and it's nice to see that those folks are clearly being listened to, as well as many other obviously intelligent folks at New Relic.

Bill: Yeah, absolutely. And apologies you didn't have a great experience the first time around with New Relic One.

Corey: Oh, it was ancient history. It was more than six months ago before you were there.

Bill: [laughs]. We'd love to have you back and give it a shot and give us feedback on if we've addressed the issues you hit.

Corey: More to come on that for sure.

Bill: Oh, good. Good. I definitely think we heard the feedback from you and other customers as well. And I believe we've solved some really fundamental problems with the announcements last week. As you said, starting with experience, we heard that the user experience was too disjoint. We had a lot of our existing applications and monitoring tools in a separate experience from the New Relic One experience. We heard that it was too complex to get started and to learn to use new features. And then, as we've already talked about, we heard that the economic barriers just didn't make sense, didn't help with adoption.

Corey: Yeah, and other engineering stuff that was irritating, too, of, “Oh, how do I install this?” “Oh, grab this janky script off of GitHub and run it.” “Okay.” I read the script because you can only trick me into doing things like that once before, “Okay, I'm going to learn on this,” and, “Okay, read things where you run them.” That's one of those things that’s extremely valuable to know right after you should have known it. But it was fine. It did everything I would expect except there was no uninstall option at that point.

It's, “Oh, great. Then I get to go through, see what it did, and manually remove resources if the trial doesn't work out.” I have checked, that's not the case anymore. It’s, “Okay. Someone is paying attention to this stuff. It's not just a repositioning/cash grab,” if you'll pardon the cynicism. It's, “No, this is something that is actually envisioned through a lens of what a customer wants.” So, I've got to say, I'm an inherently cynical person, it's nice to see it.

Bill: I can tell. And speaking of cash grab, you probably saw we're offering now a very generous free tier. I think it's 10x more generous than any other vendor out there. Hundred gigabytes, free in just telemetry per month.

Corey: That's great. I'm going to start using my logs as a database.

Bill: Totally.

Corey: I've already been using DNS that way for years. So, now it's time to upgrade.

Bill: Take all your logs, put them in New Relic, 100 gigabytes free per month, or metrics, or events, or tracing, your choice. And you get our full-stack observability product on top of it, which is the complete set of APM, infrastructure, digital experience monitoring, synthetics, you get full access to all of that, one user, free forever. So, no more money grabs. And I'll say to the experience question that you had earlier, one of the explicit focus points for us over the last several months as we prepared for this launch is the out-of-box experience and the minutes to do it. We had a goal of five minutes to join.

We want any engineer to be able to sign up for our new offering, get into the experience, automatically we provide a built-in data set. It’s actually kind of an interesting data set. We've aggregated all of the public API calls that our customers are using across all of our customer base, and we actually provide analytics on all of those public API endpoints and how they're performing. So, you get an out-of-box data set that’s kind of interesting to explore, built-in dashboard, so you can kind of see the capabilities of New Relic dashboarding and visual realization, explore that, and then begin to ingest your own data from more than 300 different sources of telemetry, and start to create your own dashboards and dive in to understand how your own systems are performing. Like that, five minutes to joy, that simplicity of experience has been a big focus for us. And we hope people take advantage of the free tier to check it out.

Corey: So, I have to ask, because this is the sort of thing that no one is ever going to ask in a formal CNBC style interview because they have this thing called, you know, tact. I'm going to come out and ask it. How the hell did that fly when you're sitting in an internal strategy meeting, and someone says, “Now listen here. You know how we normally charge people a whole bunch of money and put them into contracts so it’s, we A) charge them a lot, B) know what it's going to cost, and C) be able to book that revenue for several quarters out? Yeah, we're going to get rid of all of that, and just charge people for what they use instead,” and still be employed 20 minutes after saying that. How did that happen?

Bill: What is the New Relic values—and I have to say, I've been some pretty great companies, but New Relic lives its values more than any company I've been at. One of our values is ‘bold.’ and this was a really, really bold decision. If you need further evidence of that, just look at our stock price since we announced this strategy. Every investor I've spoken with this week has said, “We like your strategy.” It's very bold, and it's very disruptive, and it's pretty exciting to see it happen in this space. But we want to see it play out. [laughs] which is normal. Like, you know—

Corey: So, does the rest of us.

Bill: Yeah. Well, we're very bullish on it. Yeah, we spent a lot of time not only coming up with that idea, but validating with customers. So, before we ever rolled it out, we did, as you can imagine, a lot of customer studies, we've done pilot deals, we made sure that what we were doing was actually solving customer problems as opposed to making a great PR move. And that's what gives us the confidence to take these lumps, in terms of uncertainty in the market to really serve our customers because we believe when we serve our customers in the right way, it'll pay off in the long term. So, that's what we're doing.

Corey: I have a complicated relationship with usage-based pricing models, just because it comes off as being in general, yes, it can be a lot less money. I mean, I look at it through the lens of AWS as a general rule, but the problem is, is it is virtually impossible to predict, and you wind up with thinking you're doing everything right, and then finance as a minor question about why month-over-month spend is up 20 percent, and the answer is always, “The data warehouse.” But looking at that from a perspective of, especially during these unprecedented times, as we've taken to calling them, it does have a very customer friendly aspect to it that I hadn't really appreciated until confronted with a worldwide pandemic, where I have a number of customers who have a infrastructure strategy where, all right, their user traffic has dropped off of a cliff, and their infrastructure scaling and spend has not because of how things were structured and how they were built. We're hearing horror stories from other companies in the space where their entire model is that they are not letting customers do any contract adjustments whatsoever. And because you signed the line, you get to pay for it. The end. Usage-based pricing, like what you're doing solves for that. And that's really kind of a neat lens to view this through because it winds up becoming such a, I guess, a friendly story of you pay for what you use. If you don't like what you're paying, use less and the problem solves itself.

Bill: Absolutely. And there's an extra dimension of the way we structured our usage-based billing that I think it’s worth calling out because you pay for what you use, but one of the challenges even in that model is if the meters that you're being charged for are hard to predict, it's really hard to budget for. And if you look at the meters in any public cloud provider, and also maybe some of our competitors, you can see examples of that. One of the choices we made early on is to say, one of the most predictable things that every company has is their people costs, their headcount. Typically those don't go up and down very wildly.

And so as we looked at a pricing model for our full-stack observability product—which for most companies will be the majority of their spend with New Relic—we intentionally chose that as a metric, A) because it's used very commonly in other industries. You think about, for example, sales teams with, say, Salesforce as a software service. They charge based on the number of users, the number of salespeople. Very predictable model, very common. If you look at, say, the design industry with, say Creative Cloud from Adobe, again, you pay by the user for your enterprise so all your designers can access the Creative Cloud tools.

Same thing now is true with observability. You can now essentially cover your engineering team with access to a full stack of observability tools that we provide for one price. It's extremely predictable, and you only pay for what you're using. So, if your engineers are getting value out of it, then you pay for it if they aren't, and you don't pay for it. And we love that kind of mutually beneficial or reinforcing value proposition, right? If we're not giving you enough value that incents your engineers to want to use it, then you shouldn't pay us. And we're motivated to create more value and in return, customers are willing to pay for it. So, that's an exciting change.

Corey: It resonates. I mean, my least favorite software that I use right now is Tableau. And the reason is because it's a per-seat license, and it's on a year or multi-year contract, which means that every time I'm sitting here saying, oh, we've hired someone new, do they need access to Tableau, or not? The tool’s incredibly useful, and it's good, and there's really nothing else quite like it, but it's a investment decision every time I have to weigh the trade-offs of those things because the price is not small, especially as you continue to scale up. Getting away from models like that, where, “Well, I have this new application I'm launching. Do I want to wind up splurging on licenses for New Relic with it?”

I mean, I will say a part of me, I'm cynically going to miss the days where it was charged per instance because in the early days before everyone fully understood how AWS and Lambda worked, I could spin up 20,000 concurrent executions and have it reporting into the New Relic library, and then later that day, which is superhero speed underpants-outside-the-pants style for enterprise software people, you'd have one of the salespeople calling me and the hardest challenge was not making a cash register sound with their mouth. And now it's a very different story. The world and market’s understanding of how software is continuing to evolve has itself evolved. And now it's just this… it's a much different universe. It seems that now rolling out something like New Relic can also increasingly become a bottom-up strategy, as opposed to being something has to go through procurement in the traditional style of ‘determine your usage before you roll it out for anything more than a POC.’

Bill: Absolutely. Shouldn’t observability just be part of the basic engineering process? We take for granted now vast array of developer tools, source code control systems, issue tracking. There are many open-source versions, many commercial systems that are—tools that are free, GitHub, wildly popular. Shouldn't observability be the same way?

Engineers can adopt this basic part of their tool chest, use it where they want, and then companies pay for what kind of coverage that they want to have for their employees, and deploy it everywhere and instrument everything so that they're not left with surprises and waking engineers up in the middle of the night because they couldn't afford to monitor everything.

Corey: One question I've always heard from folks, whenever you talk about a software product or anything that aligns with selling something that is consumption-based—SaaS particularly—well, what happens when AWS decides to compete with you? It's always a fun story. I mean, in the monitoring space, I don't hear it nearly as much because the easy obvious rejoinder is let's get them to update their status page in a timely manner first, then we'll worry about how they're going to magically fix all of my other monitoring and observability challenges. But I'm curious to get your take on that. Is Amazon going to try to compete with you? It's probably the right question to ask.

Bill: You know, they've got CloudWatch, and they've got monitoring and other tools. Same as Microsoft. I think every smart public cloud provider will provide out of box monitoring tools for their services. Again, it's an essential part of providing a service nowadays. However, customers are often multi-cloud, and customers are not all in the public cloud, and so you need a system that can span clouds, and span your on-premise data centers, and your public cloud instances.

In fact, the very nature of observability—of measuring things—is you want to measure the differences between those environments, right? And especially as you're moving applications or services between environments, you need to understand is the performance going up or down, or being less reliable? And so we think there's a very strong need to have an independent cross-cloud observability vendor, and New Relic strives to be the best one.

Corey: I'm really excited to see where it winds up heading from here. The market seems to like it so far, but the challenge is that it's super easy to get positive responses and everyone to say nice things when the engineering effort involved is more or less a press release. I think it's really going to come down to how the market responds once people start seeing how this works for themselves. I will say, having done an analysis of the pricing model, I don't see too much to complain about there at all. So, I'm eager to see what happens.

Bill: Well, I am as well, and we hope you give it a shot. We'd love you—check it out. Let's do a future podcast on your experience.

Corey: Oh, if it goes well, then you'll be [00:28:15 laud] as a visionary. If not, you're going to be invited one of those fun meetings, where someone from HR you've never met before is sitting there, and they don't offer you coffee. That's always a warning sign. Last question for you before we wind up calling this an episode. Right now, as you look at New Relic across the spectrum, and analysts being analysts, and all the rest, what do you think is currently being the most misunderstood about the company?

Bill: I think some of the sentiment—and maybe some of it’s earned—is I've heard the ‘Old Relic’ sentiments, that New Relic is the APM vendor of yesterday.

Corey: You can't call it Old Relic, that's not an anagram of the founder’s name.

Bill: [laughs]. Well, that's true. That breaks that pattern, but there's some sentiment that New Relic’s best days are behind it. I am definitely a believer that the best days for New Relic are ahead, and it's been so exciting to be part of the reimagination of not only the New Relic One platform but the company. I mean, we talked about the bold pricing changes that impact our business model. Not many companies would be willing to take that kind of risk. And New Relic, like I said, lives our values, ‘bold’ being one of them. We've made some really big changes. And it's just the start of really exciting things to come. So, I'm excited by the future of New Relic.

Corey: Well, I’m looking forward to seeing how it shapes out. Again, if it goes well—or if it doesn’t—I'm going to be here, either way, casting stones from my particular corner of the world in which I produce remarkably little.

Bill: [laughs]. Well, have fun doing that, and stay in touch because we want to hear how it's going along the way.

Corey: Oh, you'll hear, unfortunately. That's what Twitter is for. If people want to hear more about what you have to say and what you're up to, where can they find you?

Bill: I’m @bstaples on Twitter, and you can search for my name on LinkedIn. I’m not hard to find.

Corey: Perfect. And of course, newrelic.com is where people can go to kick the tires on your new offering if they want to try it out for themselves, rather than listening to me cast slings and arrows from a place of relative impunity.

Bill: Yes, go newrelic.com, click on one of the many free buttons to try it out. Again, it's not a trial, it's free forever. And we'd love to hear what you all think. Thank you so much.

Corey: Thank you. Bill Staples, chief product officer at New Relic. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts, whereas if you've hated this podcast, please leave a five-star review on Apple Podcasts and a comment, and one of our enterprise sales reps will reach out to you in three to five weeks.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Rodrigo Flores
Rodrigo Flores is Managing Director of the Accenture Cloud Platform (ACP) in charge of Architecture, Product Management and Engineering innovation. Currently he leads a team of 400 people delivering the service to 3000 clients and $400M of cloud spend. The Accenture Cloud Platform is a hybrid cloud service that delivers a variety of providers such as Amazon, Azure, Google, VMware and Microsoft-based private clouds. Additionally, Cloud Management Services such as patching, security, backup, hardening, monitoring and security. This includes cloud brokering, security and cloud optimization services.

Additionally, Rodrigo frequently works with Global 2000 enterprises to help them in their cloud transformation and migration projects. These engagements include DevOps, financial, security and program reviews as well as one-on-one coaching and consulting with C-suite executives.

Prior to Accenture, Rodrigo worked at Cisco in the cloud software business group as CTO. He founded newScale (acquired by Cisco) the service catalog and cloud management platform pioneer from 2000--2011.

Links Referenced:

  • Accenture main site: http://accenture.com/
  • Rodrigo’s blog “Working Class CTO”: https://workingclasscto.com/
  • LinkedIn: https://www.linkedin.com/in/roflores/
  • Twitter: @rfflores

Transcript
Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored by a personal favorite: Retool. Retool allows you to build fully functional tools for your business in hours, not days or weeks. No front end frameworks to figure out or access controls to manage; just ship the tools that will move your business forward fast. Okay, let's talk about what this really is. It's Visual Basic for interfaces. Say I needed a tool to, I don't know, assemble a whole bunch of links into a weekly sarcastic newsletter that I send to everyone. I can drag various components onto a canvas: buttons, checkboxes, tables, etc. Then I can wire all of those things up to queries with all kinds of different parameters, post, get, put, delete, etc. It all connects to virtually every database natively, or you can do what I did, and build a whole crap ton of lambda functions, shove them behind some API’s gateway and use that instead. It speaks MySQL, Postgres, Dynamo—not Route 53 in a notable oversight; but nothing's perfect. Any given component then lets me tell it which query to run when I invoke it. Then it lets me wire up all of those disparate APIs into sensible interfaces. And I don't know frontend; that's the most important part here: Retool is transformational for those of us who aren't front end types. It unlocks a capability I didn't have until I found this product. I honestly haven't been this enthusiastic about a tool for a long time. Sure they're sponsoring this, but I'm also a customer and a super happy one at that. Learn more and try it for free at retool.com/lastweekinaws. That's retool.com/lastweekinaws, and tell them Corey sent you because they are about to be hearing way more from me.

Corey: Normally, I like to snark about the various sponsors that sponsor these episodes, but I'm faced with a bit of a challenge because this episode is sponsored in part by A Cloud Guru. They're the company that's sort of famous for teaching the world to cloud, and it's very, very hard to come up with anything meaningfully insulting about them. So, I'm not really going to try. They've recently improved their platform significantly, and it brings both the benefits of A Cloud Guru that we all know and love as well as the recently acquired Linux Academy together. That means that there's now an effective, hands-on, and comprehensive skills development platform for AWS, Azure, Google Cloud, and beyond. Yes, ‘and beyond’ is doing a lot of heavy lifting right there in that sentence. They have a bunch of new courses and labs that are available. For my purposes, they have a terrific learn by doing experience that you absolutely want to take a look at and they also have business offerings as well under ACG for Business. Check them out. Visit acloudguru.com to learn more. Tell them Corey sent you and wait for them to instinctively flinch. That's acloudguru.com.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Rodrigo Flores, whose professional affiliation is a bit of a story. Originally he was the founder of the newScale platform which was sold to Cisco, and after a time at Cisco, he then went to become a managing director at a small consulting company called Accenture. Rodrigo, welcome to the show.

Rodrigo: Hi, happy to be here.

Corey: So you're coming to this show from a slightly different perspective of the majority of our guests because you don't work for a cloud provider itself, and you don't work for one of those cloud-native type of companies. You're coming from a service provider perspective—which is rare in its own right—and specifically one that's focused on the global 2000 largest enterprises, which is a bit of a departure from the typical types of conversations that we tend to have here. Given that you've listened to at least one or two episodes, and put up with my slings, arrows, and other forms of various sarcastic misfortune on Twitter. What's different about the market that you work within?

Rodrigo: Ah. The global 2000—and government—have a tremendous amount of legacy because of the success. In our industry, we tend to think of legacy as a bad thing. In fact, in other countries legacy is a good thing. So that means that there are hundreds of thousands of applications that were built way before even the advent of Unix, in some cases.

So the notion of adopting public cloud, for many of them, is not and cannot be an all or nothing. So there will be data centers, there will be mainframes, there will be—still. And so the question of how to adopt Cloud, it's both a transformation program, it requires new operating models, which means thinking about roles, new skills, probably the biggest challenge to public cloud adoption today is the lack of skills in general, that are available, and experience in public cloud. And so, on the other hand, most of them are adopting public cloud in some fashion or another. So whether it's the adoption of SAAS, Microsoft Office 365, Salesforce, Workday, ServiceNow, and of course, the big three.

Many of them advertise as the majority of advertising dollars flow through online. In fact, you'll see their marketing operations and sales operations be actually quite knowledgeable about cloud-native stuff. So at different stages of maturity, even within the same company. I've run into companies that have very old mainframes from brands that are not IBM, that's how old, and then they have bought other operations that are completely cloud-native, and they have to somehow make both operational models work together. So they need a lot of help refactoring applications, migrating applications, re-conceiving the operating model, as I said, creating centers of excellence. It's a lot of work.

Corey: My approach towards multi-cloud has always been that it is not a best practice that anyone should strive for. I've given talks on this. I've written blog posts, I've been chased out of conference halls. But my point has always been that from a position of best practice and a per workload basis, that's where I stand. Examples where it does make sense are exactly what you just described, where you have a multi-divisional company that acquires something.

No, migrating them to a new cloud provider is almost never going to be worth doing. And dealing with companies as they are, who have an established track record and have a serious revenue base for which they have to express concern means that suddenly you don't get to throw something away just because it's not using the absolute latest JavaScript framework that Hacker News is falling in love with this week. There becomes a serious business consideration and a serious concern around this, and watching those concerns make their way into a world of public cloud has, from someone who doesn't participate overly much in that sector recently, has been fascinating for me to watch.

Rodrigo: Yeah. I think I can agree with you because of the [00:05:58 caveat] you put on a per workload basis. I do not see much value in trying to have the same thing across a couple of clouds because of the knowledge and the skills that you need to have, the security models et cetera, et cetera. But it is a multi-cloud world because of M&A, because different divisions have different needs, and then you have the other issue which is, outside the US, Windows servers are the data center. And to be frank because of licensing and other reasons, if you have a lot of Windows workloads Azure is probably a pretty good choice for you.

Corey: I would argue that Azure is not a terrible choice, even if you don’t. There's a startup that provides the no-code, low-code solution I use to put my newsletter together called Retool, and I was seeing an outage from them one day, and I checked—and they were public about this—they run on top of Azure completely, and there was an Azure issue, that, oh, that is fascinating to know. Now, the fact that it came to my attention through an outage is a little, eh, but by the same token, if you choose Azure, I do not believe you are fundamentally making a mistake.

Rodrigo: Yeah. Absolutely. So as I mentioned another time, Accenture itself has been for, now almost three years, 95 percent of our internal workloads are on public cloud, and somewhat divided between AWS, Azure, and now emerging Google. And so we've had a lot of experience running and managing multi-cloud. And we have to because we manage what our clients tell us to manage. We don't get to dictate the terms of where we host things. The client does that.

Corey: Oh, I bought a small services company myself, I absolutely appreciate and respect the need not to come into an engagement and tell the customer that everything they're doing is wrong, and they should fix it immediately. That is the hallmark of a junior, crappy, and short-tenured consultant.

Rodrigo: [laughs]. That's not going to work.

Corey: Well, not a second time.

Rodrigo: Not a second time. [laughs]. Yes. So the notion here is that it will be a multi-cloud world, it will be a bit messy, but yes, if you're starting one app, for example in my environment, Accenture cloud platform—which is a combination of commercial off-the-shelf software. We don't write our own monitoring, we buy monitoring tools. We don't write our own backup, we buy backup tools. But to do cloud discovery at scale, and security at scale with thousands of accounts, we actually had to create our own tool back in the day, which was all built on servers. So part of my architecture has been entirely serverless. I mean, like, dictatorially serverless, meaning we didn't want a single VM because VMs in that case, would engender patching, and we need to work with another team in operations, and we wanted to do it all DevOps.

And so, for that part of the platform, it’s entirely serverless. And what the challenge has been managing the economics, teaching the information security folks how do you secure serverless because almost all the protocols were for VMs. And that part is entirely on AWS. Yet for our billing operations, we use a tool that was acquired by Google, so eventually, to do our analytics for cost, we basically are on Google. BigQuery, et cetera, et cetera, because our tooling ended up there, not by us doing it because the provider got acquired. So that's an example where you have to say, “Well, I have to secure the stuff, manage the stuff. I did not plan for it, but that's how it went.” And I can't really replace the billing tool overnight. It takes about a year due to the complexities of cloud billing, which you—right.

Corey: Let's not get ourselves. We talk about cloud billing, we're really talking about the complexities of Microsoft Excel.

Rodrigo: Well, when you have a bill that is over a billion lines you can't even open it in Excel.

Corey: Exactly. That's when you move to a big data problem, also known as Excel for Workgroups.

Rodrigo: Yeah. Big problem. And the complexity of different contracts of that because, in our case, we are a consumer cloud—so, like a regular user—and that generally happens when there’s something called enterprise agreement. But we also resell. So that happens, and there are different contract with different terms, different liability, different discounts. And so for every single account, and every single account on the bill, we have to figure out under which contract, what discounts, and what terms and conditions. a little different than just being a user.

Corey: Just a bit. But one thing that's been fascinating to me is watching Accenture—from the outside, obviously. I fit in large companies about as well as I fit into other people's shoes—it's fascinating watching first, of course, you have to wind up being conversant across every cloud provider. There's no real alternative for you folks. You have to go where your customers are, and spoiler.

As someone who's dabbled with enterprises once or twice, they're everywhere, there is virtually no provider that a sufficiently sophisticated enterprise is not running at least some workload on. So watching what happens in that space is really neat. There's some kind of joint venture called the Accenture AWS Business Group. I don't know whether that is staffed by people from Accenture or people from AWS, but by calling it the Accenture AWS Business Group, it definitely shows the trademark AWS creativity and naming.

Rodrigo: Yeah. Accenture has a very strong ecosystem alliances model, where we work with the top providers in the world, across the entire ecosystem. We’re on SAP, Oracle, those are very big business groups. So a few years back, I think around four years ago… well, more than that, Accenture was a very early adopter of public cloud. That’s something people don't necessarily think about that. Very early adopter. I joined in 2013, and public cloud was already in use at Accenture.

And the notion became, it's like we have to go big, and so we partner with AWS in creating a joint business group—not a joint venture, but its joint business group, that has an executive from Accenture, an executive from AWS, and then personnel underneath to essentially create a workforce that could help Accenture clients move to Cloud, whether migrate to Cloud, build cloud-native, I believe the last numbers I saw, there are over 3000 people in that business group. Now, there's also one for Azure, and there's one for Google. So each of them has been staffed by a separate executive with a separate workforce. So there are literally tens of thousands of people dedicated to public cloud at Accenture. It's quite a big scale, Actually.

Corey: It’s always interesting to me watching companies adopt these cloud providers at scale, because, yeah, we're used to seeing—more or less—conference-ware: people do a quick demo online, or an open-source project of here's this thing I beat together in 20 minutes. Looking at what it takes when every person at a small startup becomes an entire division at an enterprise, and the political consequences thereof of just the interdepartmental communication, it's an organizational beehive, for lack of a better term, and finding ways for those processes to continue to function in a cohesive, responsible, and compliant way, is a massive, massive challenge. So a lot of things that come out of enterprises that they show to the world about their architecture can look, from a startup perspective to be hilariously overwrought. Like, there's somehow this, “Oh, someone decided to build as complex a pipeline as they possibly could in order to make enough work for everyone to have something to do.” Whereas in practice, it's just the opposite. This is a significant streamlining of the thing that came before, and I'm not sure that's well understood.

Rodrigo: Yeah. Well, the big thing that drives big companies is liability, right? As I said earlier, I spent most of my career in small companies and startups. In the year 2000 I started newScale—which pioneered the service catalog concept, was the first company to ever have this notion of a service catalog. We had to explain what a service catalog was. We had to prove that services could be cataloged—and I started, literally, with four people, just some friends of mine. We had nothing, so you could get very, very risky, and cut corners because who's going to come after you? You have nothing.

Big companies, though, have a lot to lose, so you end up with a lot of processes that are driven, if you will, by security concerns, by the lawyers, and then because of the scale, what you would have—in a small company, one person do, all of a sudden becomes, there are 20 different people with 20 different roles that are all involved in the same darn thing. And trying to stitch all that together when you do DevOps is the big, big challenge for big companies. That’s what I call the issues with the operating model. For example, I was with a large company here in the Bay Area, in the healthcare space, and the whole issue of procurement and budgeting for IT just completely gets blown up by cloud. You know, people are spinning up instances left and right. And somebody says, “Well, hold on a second, how do we control that? Do we need approvals?” And I go, “Well no, you can't have approvals because you won't be able to do CI/CD.” He says, “Yeah, but then how do we ensure that the budget stays in place?”

Corey: That auto-scaling group scaled up without a change approval request.

Rodrigo: Right. And the VMs may scale up, but the wallet doesn’t.

Corey: Sure. And well, at least at enterprise scale, the numbers are always a little bit different, but in some cases, those conversations could be hilarious, particularly for people who don't seem to understand a lot of these things that go on. When a CTO is asking, “Why did that scale up without a change approval request?” The easy answer is, “Who are you, and why do you think that, uh, that's how this should work?” The honest answer, in some cases, is, “Hi. I'm the person that will go to jail if we don't have the proper compliance controls in effect.” That's not really how startup-land tends to think most days. Move fast and break laws.

Rodrigo: Hey, look, I was talking about a serverless app, right, so I actually have a tool that we built so I can keep watching the spend. So every month we get the bill, and I call up the chief architect for that app, and I say, “Hey, dude, how come this thing went up 20 percent?” And, “Well, we're managing some stuff, so basically, as we manage more and more instances and cloud resources, of course the cost will go up naturally, organically, and that's fine.” But then we saw a spike. Spike is something that we've identified as this went up vertical—like we've seen the bill, right, 90 percent, and basically said, “Whenever things go spike, we got to watch them.”

So we actually created a tool for ourselves, we call ‘Spike Watch’ just to see, is that a legitimate spike? Like, for example, one of our accounts does a lot of end-of-month processing on Redshift, so you see the Redshift go for the first week of the month, just, you know, go up and up. And then after about five, six days, with the batch processing done it goes down to zero, which is exactly what you want. But initially, we didn't know, so the question then is, how do you get visibility on that, and management on that and say, “No, no, that's a fine spike. That is exactly working as intended, versus what happened here in Spike Watch.”

And the first time we ever saw that, what it was, it was one of the providers that moved us to the wrong [00:17:56 rate card]. And if you have a bill that is over a billion lines, there's no human being that can watch them. You need tools and systems, but you also need processes because the people watching that need to be able to communicate with the operators and say, “Hey guys, there was a massive spike in costs here.” And the operator needs to say, “Hey, we haven't touched a thing.” “Okay, root cause analysis.” a couple of weeks later, “Aha, we found the root cause.” Wrong rate card, not out of scaling, not anything. So when you think about going back to this large client here in the Bay Area in the healthcare space, they're really challenged for how to think of procurement, financials, and operator, and how those things come together.

Corey: In what you might be forgiven for mistaking for a blast from the past, today I want to talk about New Relic. They seem to be a relatively legacy monitoring company, and I would have agreed with that assessment up until relatively recently. But they did something a little out there: they reworked everything. They went open source, they made it so you can monitor your whole stack in one place and, most notably from my perspective, they simplified their pricing into something that is much more affordable for almost everyone. There's even a free tier with one user and 100 gigs per month, totally free. Check it out at newrelic.com.

Corey: Oh, yeah. In the world of AWS billing, I find that talking to people on Twitter, one of the biggest misconceptions people have between the folks I talk to about their personal accounts and their bill shocks versus the large customers, people live in fear on their personal accounts of having a bill surprise that leads to some extortionate number. The enterprise story is very different. If I wind up compromising your account and spinning up a bunch of bitcoin miners, for example, if you're spending $30 million a month on AWS, that isn't even going to register on the bill. So let's face it, the bill is in many cases, the best security reporting tool people have set up in their accounts. It's a very different mindset. And, okay, the bill is now 10 percent higher this month, why? Understanding why the bill has increased, or in some cases decreased is a very non-trivial exercise past a certain point of scale and attendant complexity.

Rodrigo: One of the things I noticed in my serverless accounts, one day I was like, I said, “Look, this thing's going up quite a bit. And I don't think the number of items that we're managing and discovering has gone up.” And the chief architect—I said, “What happened?” He says, “Oh, new infosec demands about encryption and security, so we are having to lock this.” And so next thing it's like, “Oh, I see. 40 percent of the spend now is with security issues.” So hey, you get an API gateway, and you have KMS, and all of a sudden KMS has been invoked every single time, and now you have a hefty KMS bill.

Corey: Yeah, it's little things like that. And again, people looking at the bill who receive the bill in accounting have no idea what these various sub-services are. In fact, the first learning experience for a lot of them is when they get this enormous bill from Amazon Web Services, they see Amazon and wonder how many books you've bought this month because they didn't notice you reading that much when they walk past your team in the elevator. It becomes a learning process and, okay, KMS is now super expensive. Is that normal, or is it not? Okay, it sounds ridiculous to people who are up to speed on what the various services do, but that's every bit as legitimate from their perspective as, huh, the EC2 bill is awfully high, should that be a major component of the bill? And understanding the nuances of this, it’s—I keep describing what I do as less about engineering these days, and more about marriage counseling for engineering and finance: getting two very different groups who, believe it or not, are aligned, communicating in a common language is a fascinating problem.

Rodrigo: Now you take it to a bigger company, and now you have three groups, or four groups: you need to be talking to finance, which includes both budgeting and procurement; and you need to talk to engineering, of course, the developers; but then operations is usually a separate group, and so you got to talk to them; and in the middle of all of this, there's also security. And so bringing, now, those four or five different areas together, it's what I mean by the operating model. How are we going to tackle this new reality? We want the benefits of public cloud, but we also are now going to have the problems of public cloud, and we have to have some way of bringing all this together. And that requires new roles, new skills, new operating model. And next thing you know, as you said, is you're doing marriage counseling rather than architecture.

Corey: It's easy to sit here and think from a perspective of, oh, well, I would just solve it by X. But that misses the entire point. There is so much else that has to happen, and how this stuff winds up working that even looking at this from a perspective of, oh, I know what the right answer is, is a lie.

Rodrigo: Mm-hm. And if you haven't done it before, if you haven't been living cloud-native, the question I have is, how would you know? And this is the challenge for a lot of clients, which is, where do they get started, right? I was with the VP of ops of a large and well known financial services company here in the States, and they run almost everything off of a mainframe and Unix. And so they wanted to know how we could help them with cloud transformation, what they need to do, what the business case is.

And at the end of the meeting, he said something that stuck with me, he says, “Look, we're running this big operations, we're all in working long days, we've had two outages in the last 60 days that have completely drained all our time, where do I get the time? Where do I find the skills and the know-how for us to start moving to public cloud?” He says, “I think it’s the right thing to do, but our situation is we don't have the money, we don't have the people, and we're all working long hours already.” So there's a real human story there, right, Corey? That means that it's a real issue.

Corey: Oh, absolutely. And again, all the hilarious things companies do generally can boil down to a lack of understanding or context. A question for you that I like to ask sometimes is, what do you think is the most misunderstood aspect in the general market about Accenture? But before you answer, I'll give my answer for this because I turned this into a conference talk. I saw a billboard at an airport once for Accenture. “New isn't on the way we're delivering it now. New, applied now.”

So I made hideous fun of this, as I am want to do. And then a friend of mine works in marketing who's way smarter at these things than I am sat down with me and said, “Oh, by the way, what's going on with that really, right?” And, “Sure I do. Of course I do. But why don't you tell me because I like hearing the way you say it.” That's how one lies convincingly. And the answer was, “Yeah. Look at what they sell. These are large scale implementation projects, in many cases, and there’s some strategy work too. But they're at a point where no one is going to go to accenture.com, punch into a shopping cart, “One consulting, please,” and click the buy button. Instead, it comes down to brand awareness, so when they wind up pitching to the decision-maker or the board, the question is not, “Who the hell is Accenture?” It's a brand awareness campaign, and that is its entire purpose. In fact, the fact that you talked about them on stage would be considered by that marketing team to be an absolute win. Sure, it was funny and it wound up taking it out of context, but that name is going to stick in people's minds, and it gets them one step further removed from, “Who the heck are you?” That was my early introduction to marketing and I thought the idea was fascinating but curious to hear your answer.

Rodrigo: Let's see. Having been here seven years, I think that one of the things that I appreciate about the Accenture culture is the intense focus on client success. And I know that sounds silly because what business could afford to not focus on client success, right? But in the sense that we value the relationship, we build long term relationships with our clients, the biggest insult that I heard one managing director one time in Australia say about another, she said, “He's way too transactional. He's not relationship-focused.”

That's why people work with Accenture. They trust the relationship. They trust they're not going to get burned. I won't say things don't go bump in the night, and everything's perfect all the time. What I'm saying is the culture here absolutely is client-centric at all costs. And I say that because most of my career was working in software companies, whether it's small ones, like my own or big ones like Cisco. And there, it's always focused on the sales numbers, right? Take the money off the table. That's not the way people think here.

Corey: One thing that has always stuck to my mind is that I come from a very small business background. To me, a big company has 200 people, and it's common for me, in that case, to say, “Oh, that company sucks.” “Well, okay, why?” “Because Ted works there, and Ted's an asshole.” And that's accepted when we're talking about a 10 person company. Accenture has half a million employees, at least, and what is astonishing is that, yeah, at that point, it goes well beyond any one person, team, or project, or even a client for that matter.

Accenture has been around for what, 31 years, and it's still one of those areas where, yes, they get things wrong, and sometimes it makes headlines. Every company who has been around long enough will have stories like that, or they haven't really affected any change at all. And it's easy to wind up coming at this from a dismissive thing of, “Ah, see. Here's a bunch of bad headlines going back 20 years.” Yeah, but no one writes an article on how that engagement went super well, and it makes The New York Times.

Rodrigo: And also, a lot of the things we do, part of the service is confidentiality. There are many things that people use and work with, that they're very happy; in fact, Accenture's behind it. The best description of what Accenture does actually was done by Jimmy Kimmel. I don't know if you ever saw that Jimmy Kimmel skit.

Corey: That depends. Jimmy Kimmel has done a couple of those.

Rodrigo: On Accenture?

Corey: Oh, I did not see that one specifically, no. He's done a lot of work, and for some reason, I've missed his management consultancy send-ups.

Rodrigo: Yeah. [laughs]. Well, if you recall, when healthcare.gov was having a lot of trouble, right? I don't know if you recall that, and it was all over the headlines—

Corey: Oh, yes.

Rodrigo: —and the government was using a particular [00:27:56 set of] service providers, and they decided that they need more help, and they called Accenture. And basically, we stepped in, and nobody knew that. Nobody knew. I think this came out later. And we put, literally, thousands of people to work on that. And so, Jimmy Kimmel comes in, “Who's Accenture?” So he did a little thing, and it’s, in the modern, complex world, sometimes [BLEEP] gets [BLEEP]-ed up. And that’s—

Corey: Always.

Rodrigo: —when you call Accenture. So anyway, if you find it, it's pretty hilarious. But that's the notion. There's a lot of services that we provide, and that we do that are essentially not seen by anybody. And confidentiality is part of the package, which is why I can't talk about them. [laughs].

Corey: Yep, that's the worst part, in some cases, about doing really awesome things for customers who have fascinating billing issues. It's one of those, “Hey, can we tell the story publicly? a lot of people would benefit from it.” And the response is, “Absolutely not. Anything involving what we spend on any given cloud provider, or in some cases, the fact that we're that cloud provider's customer at all, is a complete non-starter. If you say that, we will sue you to death.” “Cool. So that's a no, then, or…?” And it's very much a no.

I mean, at some point, even getting logo rights is like pulling teeth. We certainly don't get them all the time. But it's a challenge, in that some of your best stories you can't talk about, which from a branding and marketing perspective makes it sound completely like you're making this stuff up. The usual way that I wind up splitting that difference is all right, you can't put it publicly, but if anyone's on the fence, have them call me and I'll tell them a story over drinks, which is ideally sometimes the best you can get to. But it's an ongoing balancing act.

Rodrigo: Yes, it is absolutely a big challenge to get client references, always, because of the delicate nature, some of the stuff that we do, or how it can actually influence stock prices in many cases.

Corey: Oh my stars, yes.the publicly traded company, and the markets, and Matt Levine of Bloomberg saying constantly that everything is securities fraud. Anytime a company does anything, they get sued for securities fraud for not disclosing it, or disclosing too soon, or in the wrong way. It becomes a challenge. And that, for better or worse, is not something I generally have to deal with in most cases.

Rodrigo: Right. So that's the story.

Corey: It is. So at the time of this recording, you are winding down your time at Accenture. The plan then is for you to go and take some time off, something that will have transpired by the time people are listening to this. So as of this moment, looking forward into the ‘yay, I get to sit down and not do anything for a while’ what are you looking forward to?

Rodrigo: Well, I'm looking forward to having a bit of a rest. When I started newScale in 2000, and that was a small startup, raised money, run it. We get acquired by Cisco, you have to hit the ground running, and learn a brand new company. When I left Cisco, my last day it was on a Friday, and on a Sunday I was on an airplane to my first day at Accenture in New York. So I haven't really had a rest a true long vacation since before the year 2000, so well over 20 years ago.

So I plan to spend the next few months recharging the battery. But I'll still keep in Cloud. And I advise startups, so I'll probably continue to do that, and do a little more. My sister has her own company, so she needs help. But, I'll be deciding whether I go back and join a big company, or whether I go and join a small company, or start another company.

But it's time, and I'm blessed to have had the success and the opportunity to work at Accenture and Cisco, that I can afford to do it. And so this is a good time, in the midst of a pandemic, to step back and say, “Okay, I'm of a certain age. I got probably one more big thing in me. What do I want that to be?” No regrets, right? No regrets.

Corey: Exactly. It might be a digital transformation, or a cloud migration or, heaven forbid, a new startup. One never knows.

Rodrigo: Yeah. Because there are some interesting problems in Cloud that I think could use some new approaches.

Corey: I agree. And some of them are attached to people with job titles. [laughs]. Rodrigo, thank you so much for taking the time to speak with me today. If people want to hear more about what you have to say, where can they find you?

Rodrigo: I have a blog. And I have Twitter. So my Twitter is @rfflores. In case you wonder, that stands for Rodrigo Fernando Flores. And I have a blog called Working Class CTO, so you can find my writings there. So I'm easy to find.

Corey: Excellent. And we will of course put links to that in the [00:32:33 show notes].

Corey: Rodrigo, thank you so much for taking the time to speak with me. I appreciate it.

Rodrigo: You're welcome. It's been delightful.

Corey: It really has, hasn't it? Rodrigo Flores, managing director at Accenture Cloud Platform. In the closing days of that role. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts, whereas if you've hated this podcast, please leave a five-star review on Apple Podcasts, and in the comments include your RFP acceptance criteria.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Blake Stoddard

Blake is Senior System Administrator on Basecamp’s Operations team who spends most of his time working with Kubernetes, and AWS, in some capacity. When he’s not deep in YAML, he’s out mountain biking.

Links Referenced:

  • Basecamp: https://basecamp.com/
  • Twitter: https://twitter.com/t3rabytes
  • Signal v. Noise blog (author page): https://m.signalvnoise.com/author/blake/
  • Signal v. Noise blog (main site): signalvnoise.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored by a personal favorite: Retool. Retool allows you to build fully functional tools for your business in hours, not days or weeks. No front end frameworks to figure out or access controls to manage; just ship the tools that will move your business forward fast. Okay, let's talk about what this really is. It's Visual Basic for interfaces. Say I needed a tool to, I don't know, assemble a whole bunch of links into a weekly sarcastic newsletter that I send to everyone. I can drag various components onto a canvas: buttons, checkboxes, tables, etc. Then I can wire all of those things up to queries with all kinds of different parameters, post, get, put, delete, etc. It all connects to virtually every database natively, or you can do what I did, and build a whole crap ton of lambda functions, shove them behind some API’s gateway and use that instead. It speaks MySQL, Postgres, Dynamo—not Route 53 in a notable oversight; but nothing's perfect. Any given component then lets me tell it which query to run when I invoke it. Then it lets me wire up all of those disparate APIs into sensible interfaces. And I don't know frontend; that's the most important part here: Retool is transformational for those of us who aren't front end types. It unlocks a capability I didn't have until I found this product. I honestly haven't been this enthusiastic about a tool for a long time. Sure they're sponsoring this, but I'm also a customer and a super happy one at that. Learn more and try it for free at retool.com/lastweekinaws. That's retool.com/lastweekinaws, and tell them Corey sent you because they are about to be hearing way more from me.

Corey: Normally, I like to snark about the various sponsors that sponsored these episodes, but I'm faced with a bit of a challenge because this episode is sponsored in part by A Cloud Guru. They're the company that's sort of famous for teaching the world to cloud. And it's very, very hard to come up with anything meaningfully insulting about them. So I'm not really going to try. They've recently improved their platform significantly, and it brings both the benefits of A Cloud Guru that we all know and love as well as the recently acquired Linux Academy together. That means that there's now an effective hands on and comprehensive skills development platform for AWS Azure, Google cloud and beyond yes and beyond is doing a lot of heavy lifting right there. In that sentence, they have a bunch of new courses and labs that are available. For my purposes, they have terrific learn by doing experience that you absolutely want to take a look at. And they also have business offerings as well under ACG for business, check them out, visit acloudguru.com to learn more. Tell them Cory sent you and wait for them to instinctively flinch. That's acloudguru.com.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Blake Stoddard, senior site reliability engineer at a company that has been in the news a fair bit lately: Basecamp. Blake, welcome to the show.

Blake: Thanks for having me, Corey.

Corey: So, Basecamp was always this sort of aberration in tech circles. You rather famously did not take VC funding; you were originally called, I believe it was 37signals. Campfire was a early Slack-alike, only less awful in some ways than Slack has become; and recently you launched a second product, for lack of a better term, HEY, a controversial email client. Controversial because it doesn't let people track everyone who uses it to read email while they're sleeping.

Blake: Totally. That is exactly how it went.

Corey: So, it's been an interesting ride, most notably where people would have heard about this, is—in many cases—in small publications like The Wall Street Journal and The New York Times because post-launch, Apple decided that it was inconceivable that anyone could make money without giving Apple a 30 percent cut, and one of your f—your two founders, Jason Fried and David Heinemeier Hansson—did I pronounce that correctly?

Blake: Mm-hm. Yep.

Corey: —or DHH as we all know him, more or less went nuclear against Apple, which, given that my entire world professionally revolves around kicking a trillion and a half dollar company right in the shins, it really resonates with me. But for my case, I do it out of a labor of love, not out of a higher principle, necessarily, or a business perspective. Again—in fact, many of my business advisors urged me constantly to stop making as much fun of Amazon as I do, so then I make fun of them until they leave me alone. What's it been like? What has the experience been, from someone who more or less starts off with the, “Yeah, I'm a senior SRE; my entire job here is to keep the server's going, and oh, there's our company in The New York Times.” It's got to be a trippy experience.

Blake: It's been fun, yeah. I feel like HEY, the infrastructure behind HEY has been under my wing for a long time, for probably nine months to a year. I was really the sole infrastructure engineer working on HEY, so I’ve really seen it from the beginning. And as we got closer and closer to launch, we did all this load testing, we predicted like, okay, in the first couple months, we want to see, I don't know, x hundred thousand customers, and we knew the resource amounts needed to meet that. And we expected to get there in, I don't know, maybe six months. That sounded great.

And then leading up to the launch, and the day of the launch, everything just goes viral, and it's in the news here and there. And then all of a sudden, we're looking at traffic graphs, and we're seeing—the numbers that we expected to see six months into the launch within, like, two weeks of the launch. It's been a wild ride to scale this app up, and to be able to do it with very little customer-facing effect. Everything has gone to plan, which is interesting to say from an ops perspective. It's all worked great. We’re proud of how it is.

Corey: To be clear, in the interest of full disclosure, I am a HEY customer. I am a huge fan of the idea of what HEY is built around, specifically the idea of it calls out and shames tracking pixels, various sketchy things that have—I’m just going to say it—have infested email for a very long time, and it is terrific seeing a company take a stand against this.

Blake: Tracking pixels are something that we hold near and dear to heart in a way that we want to extinguish them from the earth. And HEY helps us do that.

Corey: Oh yeah. To be very transparent, the Last Week in AWS newsletter does have a tracking pixel at the end that tracks opens in aggregate. There's also a custom link fuzzer that I have built on top of this, which tells me, in the aggregate, cool, I wound up sending 28 links last week to—I don't know, what is it now—21,000 people at time of this recording, and I want to know how many people clicked any given link, and show me a list of what are the top-performing links? What are the top five? I care in the aggregate that some number of individuals have clicked the link, but I could not possibly care less about which individual person clicked a link. I want to know what's resonating with the audience versus what isn’t, in other words. And frankly, no one ever hits reply, so there's no real other good way to get that data.

Blake: Yeah, and I think Last Week in AWS does this perfectly. And it sets the bar between the ideal use case of no tracking pixels whatsoever, and the current use case of every email marketing company including a tracking pixel that's linked back to the user.

Corey: What I'm hoping, personally, is that HEY, takes off and launches a bit of a revolution. a bit of a revolution; that sounds like a weird, weird way of min-maxing in the same phrase, but I'm hoping that it sparks something where it becomes acceptable to—so our media kit can say we send this to x thousand people, and that's all we've got. But we already have that level of lack of transparency around podcasts, where when we put this podcast out, we have no earthly idea who's going to listen. The single metric we get is, oh, this many downloads. And from there, it's all a great mystery.

And that's kind of fun in some ways because there's no good way to track people and the only other way to do it is horrifying. For the first three or four months I was doing this podcast, I thought I'd forgotten to turn the microphone on because I got no feedback. Then I went to a conference, and more or less got swarmed by people. “I love the podcast.” “You listen to that?” And it became this really interesting journey of discovery for me. Turns out, it's easier to hit reply to a newsletter—although almost no one does that either—than it is to, “I’m going to pull over while I'm commuting, pull out my phone, and yell at whoever is on this podcast right now.” Turns out you have to have a really, really bad take for that to be someone's response.

Blake: Podcasts are interesting, too, because I feel like—I'm not very old; I’m 23 at the time that we're recording this, and back when I was, I don't know, in middle school or high school, I read about podcasts, these like cool things that anybody could do. And I really wanted to do one, but the audience just wasn't there. But now we've come around to, like, podcasts are the new wave of things that influence the way people think, and it's been a wild thing to watch change. Back on the tracking pixel thing, I feel like we've started to make a dent in the worldview there. We were talking to MailChimp recently, now any new MailChimp campaign doesn't come with tracking enabled by default. Email marketers have started writing blog posts about, like, what are we going to do now that pay is blocking how we're able to get information? And then we've seen other blog posts from companies that they say, “Okay, HEY is blocking or things now, here's all the sleazy things we're going to do now.” And then we get to go back and block all those to make their life even more fun.

Corey: I will say, I sometimes look at the feed—because again, when you build a platform, at some point, you're sort of interested to see, “Oh, wow, who's signing up for this stuff?” “How many people are signing up when I do this, this week?” Or, “How many sign up when I do Y?” I have no idea where these people are coming from, but there's been at a pretty steady organic growth of about 100 net new subscribers every week, almost since launch. And I smile, I nod, and it's like when I'm in an airplane. I don't think too hard about the physics of how this works because if you question it, it stops working. That is how I believe these things work.

And it's nutty, but I started seeing a bunch of people signing up from hey.com email addresses, which is awesome. I signed up myself from my HEY account to see how it went. And there were remarkably few scary shame-y things around what I send, which is kind of awesome, and all in all, it's been a great experience. Now, of course, I have a laundry list of feature enhancements, things that annoy me about HEY. I mean, it is a piece of software, and let's not kid ourselves, the purpose of software is solely to piss people off for not doing things as they would have them done.

Blake: Yes, HEY is definitely one of those pieces of software where you have to follow the way that we envision the software being used, or you're going to have a bad time.

Corey: I will also say, since I had not been tracking the development super closely, it's Imbox: I-M—as in Mike—or Mansy, for those who have watched a certain show—box. I look at this, and my immediate response when that pop up showed up was, “Oh my god, they launched this thing after all this work, and they had a—”

Blake: With a typo.

Corey: “—egregious typo at top of the screen.” And then it was, “Oh, it's not a typo, it's cutesy. I hate it. Thanks.” And then, okay, now I've just become accustomed to it. Mostly.

Blake: Well, I think our corporate stance is that your inboxes for important things, that's what email should be for. And so, that's the take we have, is that it's a box for important mail. It's actually gone as far as that we've had customers that have started creating Chrome extensions that find every reference to the word ‘imbox’ in the app and change it to ‘inbox.’ to appease their—

Corey: Oh, I've got to get me one of those.

Blake: [laughs]. Appease their brain.

Corey: Yeah, I like it quite a bit. It feels like it needs a few more features before I start taking it incredibly seriously as a mail client. In that, right now it only easily works with a hey.com address, and as trendy as it is these days and as much as I appreciate the long-term perspective that Basecamp has brought to all of its stuff, it feels like it's only a half step removed from, “Oh, you can email me at flyingdingus@aol.com,” where it’s, this is tied to an ISP or provider that very well may not be around for the long haul.

I was checking the other day, my vanity domain for my personal stuff, sequestered.net, was registered back in 2001, and I have gone through so many life changes, iterative steps forward. At one point, the domain for that lived running a postfix in a rack in downtown LA, that I was down there at least once a month fixing things that I'd horribly broken because it turns out that remote access out-of-band was not something I figured out in those days, all the way now to, it lives at Gmail. I'm not super thrilled with it, but it works well enough for the time being. It's been this iterative process through, but the addresses remain the same. Trusting that whatever happens in the future to the hey.com email domain, it makes me reluctant to give it out to various companies where I'm going to need to continue to have an ongoing relationship from that contact point in perpetuity.

Blake: Sure. And that's one of the policies that we have at Basecamp is that everything we create will live until the end of the internet, and we've exemplified that, actually. The very first product for release, Ta-da List, we still run it. It's running on AWS; we moved it from on-prem, it's, like, shuffled around through all of the different iterations of how we run our infrastructure, but it's still going today. And that's the plan [unintelligible] is we're going to run it forever because that's the policy: we'll run things until the end of the internet.

Corey: That's a very AWS-like policy as opposed to GCP where someone shakes their car keys just off-screen and, oop, forget this. I'm going to go chase that fun noise with the shiny thing.

Blake: What is it? SimpleDB that's been around for—I don't know—since the dinosaurs were here, and AWS still runs it.

Corey: Andy Jassy, CEO of AWS, once on record in an interview with the press as calling it a failed product, but you can still get it. People say, oh, but it's not in the console anymore. Spoiler: it never was. And I checked the other day—the job posting is currently filled, apparently—they still hire for the SimpleDB team, which feels on some level, like, wow, I didn’t realize you'd hire people directly into it. I assumed someone was screwing up, you'd put them on a pip and that was digital Siberia that you would ship them off to.

Blake: [laughs]. Oh, that's a great take.

Corey: That's got to be the saddest team there, just because sure everyone is going to insult your products on the internet, but when the CEO of your company does it in a press interview, that can't feel great.

Blake: No, not at all. And I feel like, on the other hand, I'm wary of putting anything on a Google product because I don't know if six months from now it will still be available.

Corey: And that brings us to what I really wanted to talk about by having you on the show. You spoke with the A Cloud Guru folks somewhat recently, and you alluded to some things that I wanted to dive into a bit. Specifically, you've mentioned that you are historically an on-prem shop but talked about the launch for the infrastructure of HEY running on top of AWS—which is kind of awesome—and using Kubernetes—which is the exact opposite of awesome. So, tell me a little bit about what would take you from an on-prem environment, where you are currently happily living with—by all accounts—no intent to leave, into launching something on top of a public cloud provider? What's your strategy around that? What's the story?

Blake: So, the current status of our infrastructure is that we are, I guess, technically a hybrid-cloud company. We still have on-premise data centers; we have two of them. We have several racks in those, and in fact, we run several of our large revenue-generating apps there still. In fact, we've actually run applications with their front-end compute in a major cloud provider with their database still on-prem because it’s cost and performance prohibitive to do that in the Cloud. So, we had a mandate to explore the Cloud as an option to see if we can run the same workloads that we run in our own data centers in a cloud provider’s environment where we can do the same thing at the same or cheaper price, with access to additional managed services to allow us to do more with the same size operation [unintelligible].

Corey: So, it acts more or less as a capability store slash force multiplier, in other words?

Blake: Sure. Yeah. And in fact, since we've moved to the Cloud, we have only grown the team by two. So, while growing the number of applications that we run by. One, so I guess it's not a great ratio. [laughs].

Corey: [laughs]. True, but at the same time, depending upon the actual percentages, well, okay, that's a little sketchy at scale, but with small numbers, and you look at the amount of time it takes to do these things, versus what the alternative is, that's not bad at all.

Blake: No, not at all. And I think the typical Silicon Valley VC startup backed thing to do would be, “Oh, we're going to launch a new product. Let's hire 300 people just because we can.” And in Basecamp’s case, that is totally not what happened at all.

Corey: In what you might be forgiven for mistaking for a blast from the past, today, I want to talk about New Relic. They seem to be a relatively legacy monitoring company, and I would have agreed with that assessment up until relatively recently, but they did something a little out there: they reworked everything. They went open source, they made it so you can monitor your whole stack in one place. And most notably from my perspective, they simplified their pricing into something that is much more affordable for almost everyone. There's even a free tier with one user and a hundred gigs per month, totally free.

Check it out at newrelic.com.

Corey: One the thing that I see periodically, when you have an on-prem environment that then decides to expand into the Cloud for some aspects of it, is—how do I put this politely? I guess I don’t. You wind up with a VMware model. The payday lender of technical debt, where you're going to just run a bunch of VMs, but now also in the Cloud. You're not really leveraging cloud in that story so much as you are making it look like a version of your data center. You are effectively worsening the cloud environment in order to slightly improve your capability, data-center-side, not that that's inherently a bad thing, but it's not what I would call cost-effective either.

Blake: No, not at all. That's one of the things that I think should be a prime tenet of any corporations looking to move to the Cloud. It should not be seen as a way to obfuscate CAPEX to OPEX. Because if you're going to the Cloud, you should be doing it to gain additional value. In our case, we decided to—let's explore the Cloud as an option. The mandate was explicitly do not lift and shift. In fact, we had a typo recently internally where we said ‘lift and shift,’ and I think that's a pretty accurate representation of how some corporate cloud moves go.

So, when we started looking at the Cloud, we knew that we wanted to use containers as a way to orchestrate the way that we run our apps. When we—the things that we've run on-premise, we do so on bare metal, we deploy them with Capistrano, we use Chef to manage the boxes. We don't use containers on-premise, but by going to the Cloud and being able to use a managed container orchestration service, that gives us a great chance to look at containerizing our apps, and running them in containers, not just because it's the cool thing to do, but also because we gain value in being able to binpack them better and use compute more efficiently than we could if we were just running them on a fleet of t2.nano instances.

Corey: So, tell me a little bit more about the Kubernetes decision. Is this something that predated your HEY build-out? Is that something you've been dabbling with for a long time? I've got to say that Basecamp as a company has seemed relatively immune to hype-driven development in many respects, so seeing that you folks were on Kubernetes was a little bit surprising.

Blake: Yes, it did predate HEY development. When we first moved to the Cloud and decided that containers were the way forward for us, we started out using AWS as Elastic Container Service, or ECS as they prefer to call it. ECS is fine. It works well enough, but our take was that the service was not gaining features at a good enough pace, and we were running into problems with it being sometimes a very big black box, where things would happen, and you didn't know why, and your fix was to open a support ticket and hope they responded to you quickly and helpfully. And our experience with support is that neither of those two things happened.

Corey: Right? You just shame them publicly and loudly in increasingly public places.

Blake: Yes. And sometimes it helps a lot. [laughs]. Yeah, beyond that, when we decided that ECS wasn't the way forward, but we still wanted to use containers, Kubernetes was the thing to do here. And around the same time, we also started looking at leaving AWS in favor of Google Cloud Because if you want to use containers, Google's GKE product is the way to go. I mean, for a project that came out of Google, using their managed version of the product is the way to go.

And in fact, we did do that. Basecamp 2, we actually ran on GKE for several months with minimal issues until Google started having a few bad days a lot. And at that point, we decided to move Basecamp 2 back on-prem, but we didn't want to leave Kubernetes as a whole. By this point, we had already started work on HEY, and HEY was living on GKE, too. But when we moved Basecamp 2 away from GKE and moved it in on-prem, we decided that we were going to not use Google at all. And so since we already had the infrastructure, we were able to make use of the flagship feature of Kubernetes, and move it to another Kubernetes platform with very little additional infrastructure work. And from there, we moved to the EKS, and it's been living there fine since then.

Corey: What made you decide to go EKS instead of rolling your own control plane?

Blake: The price of a managed Kubernetes cluster is less than the price of the number of engineering hours it would take to run Kubernetes on bare metal, or on EC2.

Corey: Nope, very fair. To be blunt, I wish more people accepted that. Something that also struck me as interesting about your exploration of this, was the idea of using Kubernetes on top of Spot. That's something I've been advocating for from an economic perspective for a long time, but in practice, it feels like you talk about the Cloud as oh, it's elastic, you can scale up, and load increases, and scale down when it doesn't need to, which makes everyone feel super good about not doing exactly that. Everything winds up at the same baseline level of usage. And oh, we'll get to it next sprint, as if suddenly you're going to stop making poor decisions right after this one.

Blake: Using Spot requires a bunch of additional thought and processes around getting workloads onto Kubernetes, but once they're there, you are able to reap the benefits of not having committed compute capacity just sitting around when you don't need it, you're able to use all of the instances that AWS offers you, all while saving money while doing it. Now, on the other hand, Spot looks great on paper, it looks great from a billings perspective, but I'd argue that Spot is one of the stickiest services that AWS offers because you can't replicate that on-premise. Google's Preemptible Instance [unintelligible] only gives you a thirty-second window, versus Amazon's two-minute window. Amazon's implementation of Spot works really well for us because our workloads, we’re able to know—well, we're able to have two different versions. We know that some workloads are okay being terminated with the two-minute warning, and then things that we know aren't okay being terminated with a two-minute warning, we run them on on-demand instances, and they're able to work well.

Corey: So, you said—or implied at least—that Spot is one of those things that is not able to be replicated on-prem; you're fundamentally not going to get there in the same way. Tell me a little bit more about that. What do you think the Spot market does in terms of lock-in? Do you think that this is something that's going to drive people, once they learn to adopt to something like the Spot market, that makes it in some ways harder to leave AWS than any pure technology buy?

Blake: No, I don't think that it's a product that causes you to say like, “Oh, this is absolutely amazing. What will we do without it?” I think that it makes you feel really nice inside when you see the cost savings. And when you see the ability that you have to pull from vast resource pools that you can't replicate on-premise.

Corey: It's neat to start seeing capability stories like that around the things, to be very direct, that people are using in AWS. Because they get on stage, they talk a lot about their ridiculous nonsense, “Oh, it's a machine learning musical keyboard.” Or, “Hey, it's this ridiculous thing that winds up leveraging 18 different implementations of blockchain.” But if you look at what people are actually using/spending the money on, it’s EC2, it’s data transfer, it’s S3, it’s database store, it’s disk. It's the boring building blocks that no one wants to talk about in keynotes because it's not interesting or exciting anymore the way that it once was, but it's what the world runs on. So, things like Spot seemed like they are aligned with that vision of the future.

Blake: Totally. And using Spot isn't just, click a button and your workloads will be fine. It requires taking the time to look through them, seeing what can deal with being torn down with a very short amount of notice, and accepting that maybe your workloads aren't good for Spot. Maybe you don't want to use this. But for a lot of workloads, you totally can, especially for front-end web workloads like ours, we have no reason to not use Spot.

Corey: One of the things that I find compelling is something like that does force you to refresh your instances—or at least have a plan to refresh them—at virtually any moment. Whereas in practice, we talk—in many environments—oh, we believe in cattle instead of pets, and then you look at their environment, like, “Oh, great. So, I can turn any one of these things off?” “Absolutely. Except for that one, that one, that one, that one and, oh my god, that one.” And it comes down to the, ‘what we say versus what we do’ story.

Blake: Yeah, totally. And that's actually one of the things that Kubernetes helps us with a ton by being able to use Spot instances, is that we can let things come and go into the auto-scaling groups and actually treat these instances like cattle because Kubernetes can take care of scheduling things in other nodes when they become available, it can take care of what's running now, what needs to be running, we have controllers in the cluster that notice that, “Oh, we've lost an instance. We have pods that need to be scheduled, but we don't have that capacity. Add it to the cluster.” Kubernetes helps us a ton there.

Corey: So, when you take a look across the landscape of what cloud providers offer, what you're able to achieve on-prem, do you think the future is pure public cloud for the sort of things you do? Do you foresee there's going to be Basecamp data centers for the next century or something else entirely?

Blake: I think hybrid cloud is going to continue to be the thing that we see more and more of. I think Google Cloud is pushing it a ton with Anthos. We're doing it ourselves. We've done it before, where we run front-end compute on Kubernetes, and we keep databases on-prem, and we connected them over direct connects and TCP interconnects. Some things are just not feasible to do in the Cloud, whether that's from a cost standpoint, whether it's from a performance standpoint, whether there's some service that you want to run that isn't offered in a managed form in AWS if you want to use it as a managed form. I think hybrid cloud is going to be the end goal. And I think we've even seen companies like Netflix that start out being all-in on AWS decide, “Oh. This is kind of expensive,” and decided to run their own data centers with their own hardware sometimes. So, hybrid cloud, I think, will end up being the end goal of what we do, but I think the implementation will be different for everybody.

Corey: What do you think is currently the most misunderstood thing about Kubernetes in the larger ecosystem?

Blake: That you're not required to use it. [laughs].

Corey: [laughs]. I like that quite a bit. This is coming from Basecamp. Again, one of your founders was the creator of Ruby on Rails. There's a very anti-trendy-JavaScript-framework philosophy there that, frankly, I wish the rest of the industry shared. So, it's easy to dismiss this as, oh, those are just a few countercultural folks who are trying to stir up trouble. I don't believe that that's true. I think that there's a lot to be said for using technologies that have been shown to work, for focusing on parts of the story that are, I guess, more in line with what sensible people with a business interest are concerned about. And it feels like on some level that the number one project everyone's trying to solve for remains their own resume.

Blake: Especially for a company the size of Basecamp. We won't do things unless we see value coming out of them. We didn't look at the Cloud and decide to do it just because it was a cool thing to do. We didn’t look at Kubernetes and decide to start using it just because it was a cool thing to do. We started using Kubernetes because it gave us a path to accomplish our end goal, which was making our compute more efficient, being able to run it in ways that we can't do on-prem, and being able to do more with the same number of operations staff members. Kubernetes and public cloud are great for those things.

For a product like HEY, if we were running that on-premise, we would have massively over-provisioned the hardware, spent hundreds of thousands of dollars on hardware for a product that we can't guarantee is going to perform the way that we think it will. We would have had to hire additional people to be in charge of racking and stacking that year, and that's just not something we have to worry about when we’re running on a managed cloud product, and when we're using something like Kubernetes.

Corey: It seems increasingly like solving for business value rather than for hype is taking a renewed focus at a lot of companies. I suspect that the longer this pandemic drags on, the better enterprise tech is going to become, just due to lack of executive exposure to ads in airports, which historically seems to have been driving an awful lot of very strange decision making. Now we're seeing, “Well, what is the business value for this?” Type conversations coming out. And I'm optimistic that this is going to usher in a new era of good decision making. But, on balance, these are human beings we're talking about and well we have some track record of showing how that is not true.

Blake: My hope is that we start seeing executive sponsorship lean more into how does the operations team want to run things? How can they run things efficiently? Let the people who do the work make the informed decisions about how they want to work, how they can work most efficiently, how they can meet the goals of the business, rather than just handing it down from the top with no discussion to the implementers.

Corey: If people want to hear more about what you have to say about this and other topics, where can they find you?

Blake: I do a lot of ranting about Kubernetes and public cloud on twitter at @t3rabytes—the word ‘terabytes’ with a 3 instead of an E. And I've also been writing more on our company blog, Signal v. Noise, which is signalvnoise.com.

Corey: And we will put links to those in the show notes, of course. Thank you so much for taking the time to speak with me today. I appreciate it.

Blake: Absolutely. Thanks for having me, Corey.

Corey: Blake Stoddard, senior site reliability engineer at Basecamp. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts, whereas if you hated it, please leave a five-star review on Apple Podcasts and a tracking pixel in the comments.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Sachin Agarwal

Sachin Agarwal is the Worldwide Product Management Lead at IBM Aspera where he is responsible for leading up the portfolio's strategy and roadmap. Prior to joining IBM, he held various product ownership positions including Principal Product Manager at LaunchDarkly, VP of Product and Operations at Nylas, and Director of Product Management at Oracle. Sachin's post-career goal is to retire in Hawai'i, where he plans to open a great Chicago-themed gastropub/cocktail bar.

Links Referenced:

  • IBM Aspera: https://www.ibm.com/products/aspera
  • Twitter: https://twitter.com/sachinag
  • LinkedIn: https://www.linkedin.com/in/sachinag/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

This episode is sponsored by a personal favorite: Retool. Retool allows you to build fully functional tools for your business in hours, not days or weeks. No front end frameworks to figure out or access controls to manage; just ship the tools that will move your business forward fast. Okay, let's talk about what this really is. It's Visual Basic for interfaces. Say I needed a tool to, I don't know, assemble a whole bunch of links into a weekly sarcastic newsletter that I send to everyone. I can drag various components onto a canvas: buttons, checkboxes, tables, etc. Then I can wire all of those things up to queries with all kinds of different parameters, post, get, put, delete, etc. It all connects to virtually every database natively, or you can do what I did, and build a whole crap ton of lambda functions, shove them behind some API’s gateway and use that instead. It speaks MySQL, Postgres, Dynamo—not Route 53 in a notable oversight; but nothing's perfect. Any given component then lets me tell it which query to run when I invoke it. Then it lets me wire up all of those disparate APIs into sensible interfaces. And I don't know frontend; that's the most important part here: Retool is transformational for those of us who aren't front end types. It unlocks a capability I didn't have until I found this product. I honestly haven't been this enthusiastic about a tool for a long time. Sure they're sponsoring this, but I'm also a customer and a super happy one at that. Learn more and try it for free at retool.com/lastweekinaws. That's retool.com/lastweekinaws, and tell them Corey sent you because they are about to be hearing way more from me.

Corey: Normally, I like to snark about the various sponsors that sponsored these episodes, but I'm faced with a bit of a challenge because this episode is sponsored in part by A Cloud Guru. They're the company that's sort of famous for teaching the world to cloud. And it's very, very hard to come up with anything meaningfully insulting about them. So I'm not really going to try. They've recently improved their platform significantly, and it brings both the benefits of A Cloud Guru that we all know and love as well as the recently acquired Linux Academy together. That means that there's now an effective hands on and comprehensive skills development platform for AWS Azure, Google cloud and beyond yes and beyond is doing a lot of heavy lifting right there. In that sentence, they have a bunch of new courses and labs that are available. For my purposes, they have terrific learn by doing experience that you absolutely want to take a look at. And they also have business offerings as well under ACG for business, check them out, visit acloudguru.com to learn more. Tell them Cory sent you and wait for them to instinctively flinch. That's acloudguru.com.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Sachin Agarwal, currently at IBM Aspera. Sachin, welcome to the show.

Sachin: Corey. It's great to be here. Thanks for having me.

Corey: So, before we get started, one thing that you just called out during the pre-show is something I always do, but it never occurred to me to talk about publicly, and that is going back and forth with a guest to make sure I pronounce your name correctly. I always assumed it was one of those baseline decency things that doesn't need to be called out, but you raised a very interesting point, specifically that call these things out and normalize it because it's important. There's sort of this spectrum of how important it is to get someone's name right. At some points is, “Steve. Call me Steve.” I'm putting an order in at Starbucks, this is not a long term relationship. The other failure mode, though, is welcome to a podcast or a keynote, I'm going to guess at your pronunciation of your name and wing it, which is awful.

Sachin: Yeah, we've all seen that, right? Like, people have gone up there and they've completely butchered it. I've had teachers butcher it, and every Indian person in America has a Starbucks name. Mine is ‘Sachin’ which is different than ‘Suchin,’ but my middle name is actually Dave, which is not short for David; it’s just Dave. And if you're Indian, Sachin Dave is a reasonable name. When you put it in Roman letters, and it's Dave, that's hella weird. And so, thanks dad for doing that to me.

Corey: In the before times when I used to go to Starbucks, my Starbucks name that I gave always was Spartacus because I had hoped one day of 15 people all trying to claim that drink like the old Spartacus movie. So, far, no luck and everyone looks at me strangely, but what I have learned is that there is an entire universe of different ways to spell Spartacus.

Sachin: I can only imagine. Do they at least all start with an ’S?’

Corey: Mostly, I think someone tried at once with a ‘C’ but I don't want to go too far down that path, my God. So, back to the actual reason that you're here. So, you have an interesting career path. Historically, you were at places like [Nylas] and LaunchDarkly—in fact, Heidi from LaunchDarkly was the number one inaugural guest of this show, once upon a time—and you also worked at places like Oracle and SAP and now you are the worldwide product management lead at IBM Aspera, which means you're pretty much Benjamin Button-ing your career. And I can only assume your swan song will be taking a role at the DMV someday.

Sachin: The DMV is a good place to be. And the California DMV is incredibly under-appreciated for the hard work that they do.

Corey: Truly.

Sachin: I was very happy to get my RealID during these times from the California DMV. It is possible online.

Corey: That must have been an experience.

Sachin: You upload your files online and you get your stuff. It works.

Corey: That is revolutionary. But back to the topic. Snark aside, you are at IBM, this is presumably an intentional choice on some level?

Sachin: It is.

Corey: What do you do there, and what was appealing about IBM specifically, Aspera—am I pronouncing Aspera correctly? I ask about people, never the company.

Sachin: I believe it's also pronounced Aspera, so if we're wrong, we can be wrong together.

Corey: We could go Aspera. We could wind up pronouncing IBM is basically one word instead of an acronym; all kinds of different things we could do. But what are you doing there? What brought you to IBM in this decade?

Sachin: Yeah, so it's a really interesting thing. So, Aspera is a startup that IBM acquired about five years ago. And the founders of Aspera were solving a very deep, hard, technical problem, which is moving data over bad internet connections, or bad network connections more generically. And essentially, Aspera has this technology called FASP, which stands for Fast Adaptive Secure Protocol. And they solve, essentially, the TCP and the UDP joke problems. So, everyone who listened to this is familiar with the TCP joke and the UDP joke and blah, blah, blah, blah, blah, and Aspera solves that. And IBM scooped them up and made them part of the broader IBM behemoth.

And so for me, it's really interesting because I'm a startup guy, but I'm also comfortable in big companies. And hey, there's this startup that lives inside of IBM. And they have really cool technology; the people there are really smart and really passionate and amazing. And the office happens to be walkable from my house. I can go in, I can do a bunch of good stuff, I can foot in both sides, and leading product for really cool, technically amazing stuff is kind of my wheelhouse right? Nylas, LaunchDarkly, those are dev tools companies. And thinking about Aspera as a dev tools company is really interesting and exciting for me.

Corey: It's easy to sometimes get lost in the weeds of the snark in the various things that I do. I am snarky and sarcastic about every large multinational company out there because my belief has always been punching up is fine, punching down is not. There's a reason I'm not here making jokes about LaunchDarkly or Nylas in the same way because those are small, scrappy startups that are doing well and doing great things, but me crapping on them in public does not lead anywhere positive in the same way that making fun of a company that is over a century old, in IBM’s case, does. That said, you can make fun of companies, but not people. I don't believe in inviting folks from these companies onto the show and then blaming them for their company’s failings.

To be very clear, this may change if I wind up getting your CEO on, for example, or a board member, but almost no one I speak to here is generally in a position to unilaterally set multinational company strategy. So, if people are wondering about why I'm not really putting your feet to the fire on any of IBM’s various missteps, well, it's because I'm not naive enough to fall into my old pattern of believing that a big company was 200 people, and everyone had an impact on everything the company did. So, you'll forgive me if I don't blame you for every failing, real or perceived, that IBM has about it in the past century and a half.

Sachin: That is a reasonable and measured take. I appreciate that Corey.

Corey: We do our best to make this a friendly, welcoming place. But, so let's talk about the somewhat recent IBM Cloud outage that went around the internet three times. Now, I want to be clear that first, Aspera is not—to my understanding—part of the IBM Cloud group, but you do use aspects of it. And secondly, of course, you are not speaking on behalf of IBM here, you are speaking on behalf of yourself.

Sachin: I am speaking indeed on behalf of myself. And indeed Aspera is a customer of IBM Cloud, same as any other customer would be. So, obviously, yes, that happened. It was unfortunate for everyone involved. You know, we went down, everyone else went down. The folks there really do their best. I appreciate the hard challenges of doing this. The Cloud is just someone else's computer, and in this case, it was IBM’s, or our computer. And I think the team learned a lot from it.

There are changes that have happened because of it, but I think the team actually performed reasonably well. They were really on top of it. It wasn't that anyone was left unawares. The machine worked the way that it was intended and supposed to work, and the machine is better off today than it was in the past. You know, one of the things that's interesting about IBM from the outside, or being new to IBM, is that that client-service mentality is really deep and ingrained.

And so when you have a client-service mentality, and you really care about each and every individual customer as a client and not just as a set of customers, it's very familiar to me as an ex-investment banker—way, way back in the day—is that there is some basic set of decency and some basic set of being circumspect. And so IBM did a very good job of communicating with its customers—Aspera internally, or other folks externally—for those who asked questions, those who wanted to know, keeping us up to date. And IBM’s culture is very much better at doing that one-to-one when someone has a question or talking to clients. IBM has gotten better at doing things a little bit more publicly. You can see—you know, there was a small, minor incident earlier this week, and the IBM Cloud Twitter account correctly went out there, the status page was very quickly updated, folks had a level of certainty as to what was going on.

Corey: For those who are wondering we're recording on Friday, July 17 of 2020. It will be a bit of a production delay before this makes it out during these unprecedented times. That's right, these times are no longer precedented.

Sachin: So, from my perspective, I feel heartened by what I've seen as a result of it, right? All you can ask is that when things go badly, people do their best, they make things work, they do the root cause analysis, the post mortems, they ask themselves, “How can we do better?” And the folks at IBM Cloud honestly have hit every single one of those. And again, I'm speaking for myself, I'm not here to be speaking on behalf of other folks in the organization, but I am proud of what they've done.

Corey: I disagree with aspects of your experience, as it does not match what I'm hearing from some other IBM customers but again, the point of this show is not to drag you, get you in trouble, or call you to account for frankly, a division for which the only thing you have in common with is the name on the paycheck and remarkably little else. There's a whole universe of criticisms, arguments, user experience stories, et cetera I could levy against IBM Cloud, but again, you don't speak for them, I wouldn't expect you to speak for them. And if the only thing I knew about you was your IBM.com email address, it's easy for that vision to get lost in the weeds. What I would rather talk to you about—if it's all right—is, tell me about the idea of a startup getting acquired by IBM, and specifically, what it is you folks do because when we first met, I was skeptical, but we started talking and it's, “This is kind of amazing. How come we aren't hearing more about it?” So, tell me what it is you do?

Sachin: Sure. So, at Aspera, we move data over the public internet faster than anything else on Earth. So, that protocol FASP stands for fast, stands for adaptive, stands for secure. And one of the things that we're really excited about is that we actually have a multi-tenant, hybrid multi-cloud thing called Aspera on Cloud where we actually run our software in all four of the major public clouds. Yes, IBM Cloud, but also AWS, Microsoft, Azure, and Google Cloud.

So, we're in all four of those places where we run compute, so that way customers of ours can literally just attach their storage to our servers, and move data into and out of public cloud storage, super-duper fast, you know, tens, hundreds, even sometimes thousands of times faster than TCP-based protocol, so your standard stuff. You can imagine that when that happens, that you end up having implications of workflow changes and business model changes. You go from doing batch processing to real-time processing; you go from everyone must be co-located to people can be wherever they want, because I can get this file to you across the world as fast as I could across the office. And so it's really interesting, sort of, things where, as you end up in a more distributed environment, we end up in these, people are doing compute in one place but running their ML models on the copy of that data in a different cloud. You want to be able to have that information at the place where it matters as soon as possible, and Aspera kind of like magically handles that for folks. And so it's really interesting to have this really cool technology and be able to leverage it and present it to folks. It's interesting when you say, “Why haven't I heard about it?” Well, thank you for having me on, so at least this way, the few dozen folks who listen to this podcast will have heard about it. But that's my job.

Corey: Oh, those are just the ones that angrily dial in and respond.

Sachin: Ah, do they take their answers off air? That's really the important part.

Corey: Yes, sometimes they're delivered by brick through window.

Sachin: Oh, there's a Microsoft Windows joke here and I can't find it in time.

Corey: Yeah, there's a Databricks meets Microsoft Windows story in there somewhere, but who has time to look for all of the puns we could come up with on those?

Sachin: I'm just not at the dad level that I aspire to be.

Corey: So, one of the challenges, I think, whenever you start talking about multi-cloud stories around, oh, there’s the thing that improves data transfer stories, it fixes some of the economic stories about it, but what it inevitably does is it takes an already extraordinarily complex story, and then adds an additional layer of complexity on top of it. Whether that takes the form of an abstraction layer that purports to simplify it or not, it does act as a underlying complexifying factor, and that has always been a source of aversion for a number of companies, at least at a scale where they still imagine that one person could theoretically hold their entire software stack in their head; at enterprise scale that falls by the wayside. So, I'm curious as to how do you view that, and who do you end up targeting from a market segmentation perspective?

Sachin: Yeah, sure. So, how we handle it, God bless SREs, God bless our engineers, God bless our networking folks. Like, it is a hard, complex topology, right? One of the things, Corey, that I've heard you say is that people should pick a cloud, stick with it and take advantage of the uniqueness of whatever cloud they pick. For us, we want to move data to anywhere from anywhere. We're on those four major public clouds, but Aspera is also embedded into Akamai storage and a bunch of other things that are out there. We wanted to make sure that we were as agnostic as possible about where folks wanted to store their stuff, so the topology that we run is crazy. And so it's very, very difficult.

And from a market segmentation perspective, it's actually really interesting. So, yes, we're part of IBM, but we actually have a [paygo] option where you can get up running on your own, without talking to an IBM seller. Aspera on Cloud is actually in the Google Cloud Platform marketplace, and we're working very hard to get into other cloud marketplaces as well. And so from a market segmentation perspective, everyone says, “We want everyone.” I'm actually working very hard with our team to be everywhere that people are looking to solve these sorts of problems so they can choose to buy, get started, trial, whatever they want—Aspera in the way that they're comfortable with.

So, for folks that are big, large enterprises, who really want to have a sales conversation with a seller, and a technical expert, and go through the topology and read through all the reports, we've got that. For someone who's just like, “I’ve got 10 petabytes on my laptop, and I need to move this up to S3, and I got to get it there now.” We've got to click through, push-button, do it right now, sign up, five minutes, attach your S3 bucket, go. We try to do it all. Obviously, when you try to do it all, things fall through, you have to make trade-offs, and that's actually kind of the fun part of product management is really being intentional about what you choose to do, what problems you choose to solve, who you choose to solve for, the emotional benefit states that you end up solving when you do your job correctly, and where you want to be. It's a really interesting product management challenge. IBM calls it ‘offering management’ for these reasons, which, honestly is not a bad term, now that I've come around on it. So, yeah, they're hard problems. They're fun problems. That's kind of why I’m excited.

Corey: It's not a lot of fun for anyone to sit here and only solve easy problems for the rest of one's career. Which really brings me to the next point I want to cover with you is, you and I have something in common in that we have never stayed as an employee of a company for longer than a little over two years, in your case, slightly under in mine. I'm curious to get your thoughts on that because there's a certain vocal majority who believes, when they look at people like us, that this says terrible things about our character, about our entire approach to different things, that we are somehow just bad people.

Sachin: Yeah. Obviously, it's an intentional thing for both of us is that we do that. It's not that, oh, we get bored at a place and we leave it in two years. As a people manager, one of the things that I always feel is it is my job to earn your presence at this company on a day to day basis. It is my job to make sure that we are challenging you, and giving you the opportunities to succeed, compensating and recognizing you fairly, and if we don't shame on us. And so it's not that companies that I've worked at haven't done that, but sometimes you find more interesting places in other things. Sometimes circumstances change.

One of the things that I mentioned is that I can walk to work—you know, back in the old days where we actually went to offices, but I can do that. And so for me, one of the things is when I first moved out to the Bay Area, I was commuting from the city down into the peninsula. Well, as I got older, and my knees stopped working, and my back stopped working, I wanted to be closer and closer to home, and so the combination of things changed for me—my priorities changed—but also I'm intentionally looking at my career and what's best for me and my family. You know, those natural breaks tend to happen, and from my perspective, no one owes a particular loyalty to their employer. That's something I believe strongly, and as a manager, it is your job to make sure the people who work for you are recognized, challenged, and appreciated.

Corey: I would argue that you do owe a particular loyalty to your employer for as long as they remain your employer, but you don't owe them a duty to remain employed there. Just to, I guess, add some nuance to a statement that has the potential to be taken wildly out of context.

Sachin: Thank you, Corey, for helping me keep my job.

Corey: I do my best.

Sachin: You said it better than I did. You're right. You do have a particular loyalty to your employer when you're employed. Exactly.

Corey: Oh, I've been very vocal about my distaste for non-compete agreements, but I haven't been quite as vocal about my support for confidential information clauses, for non-disclosure agreements. I have a sarcastic number of non-disclosure agreements with a wide variety of companies to the point where I don't keep track of them anymore because my default response is without explicit and enthusiastic permission to tell someone's story tied back to them, I can't do that. Just because—even if it's okay, and if there's even the sense that enters the industry of I tell other people's stories, and I can't be trusted, my entire business dies. I take this stuff incredibly seriously just from a—first, a branding, but more importantly from a morality perspective.

Sachin: Yeah, I think that's exactly right. Like, you do the best you can where you are. And you have to be good to the folks that you work with. You have to think about your career, and you have to think about your reputation, and doing what you just said is the right way to handle that long term. And short term. It's the right thing to do.

Corey: In what you might be forgiven for mistaking for a blast from the past, today, I want to talk about New Relic. They seem to be a relatively legacy monitoring company, and I would have agreed with that assessment up until relatively recently, but they did something a little out there: they reworked everything. They went open source, they made it so you can monitor your whole stack in one place. And most notably from my perspective, they simplified their pricing into something that is much more affordable for almost everyone. There's even a free tier with one user and a hundred gigs per month, totally free.

Check it out at newrelic.com.

Corey: It really seems to be. It's an interesting world, and we're seeing it continue to get more interesting. So, now that I've done some help—you know, helping you keep your job, let's go back again to putting it in peril, again. So, occasionally, I make questionable, weird decisions that, in hindsight, weren't the best. Specifically around my finances, if I'm being honest. I know, this is a weird thing for a cloud economist, but if Apple releases a new product, I'll rush out and buy it, and then I'll look at it and kind of regret overspending. Like, did I really need to spend another thousand dollars on an iPad? Similarly, IBM went out and spent $34 billion buying Red Hat. Can you add any color or context to that acquisition?

Sachin: Sure I can. Obviously Arvind, Ginny, Jim have answered those questions a million times, so there's no reason for me to rehash that. I'll tell you my perspective as someone who works at IBM and interfaces—oh God, I use a technical word. That's 20 bucks—as someone who talks to folks on the Red Hat side, I am very, very happy that Red Hat is a part of IBM. IBM has done, from my perspective, a really, really good job here. And I'm not speaking for IBM, I'm speaking for myself. IBM has done a really good job of letting Red Hat be Red Hat; letting Red Hat be special.

Red Hat is scrupulously neutral, and that has continued to be there. But IBM biases towards Red Hat. Like that's, I think, a reasonable way of saying it. Like, Red Hat is true to itself, and IBM is biased towards Red Hat. You know, Red Hat Enterprise Linux is really, really good. Red Hat OpenShift is really, really good. The people at Red Hat are really, really good. It is awesome to have those people; those technologies; that customer base; that knowledge; that deep love for open-source; the deep understanding of communities, and network effects, and ecosystems all of that inside of IBM. I think Red Hat is actually influencing IBM a hell of a lot more than IBM is influencing Red Hat, which, frankly, from my perspective is as it should be. So, for me, I'm thrilled about it. That acquisition actually closed before I joined IBM, that's how recent I am to IBM and to Aspera, but I'm thrilled about it just at a personal level because it aligns with, I think, where everyone sees what's happening 2020 and beyond.

Corey: That's an interesting glance into the future. There's a lot of good things that have come out of Red Hat. Lord knows, I have friends who work there, and friends who work at IBM—I know that biggest disclosure here, that's the bombshell is that I actually have friends—but what I find interesting about these companies that are targeting market segments that are not as exciting and as appealing to me as some others, namely large enterprises that are still not fully invested in the idea of cloud, then it's easy to be dismissive and look at these companies as a lumbering dinosaurs that are just not suffering from the asteroid impact yet. And that's unfair. It doesn't do me any credit to have those opinions because the people I talk to there are absolutely some of the best in the world at what they do.

There's a tremendous upside to working at these companies. These are not people making bad decisions, and these aren't people who have just given up on things, despite my snark level at you toward the beginning of this episode. There's something that is incredibly fulfilling, in some ways, about being part of a large company where you are able to focus on a fixed set of scope of a problem, and really dive into that in a way that at smaller companies we never really get to do. So, it's just a different approach, a different way of solving the needs of a different customer base. So, my slings and arrows completely aside, there's a lot of good things, and there's clearly a lot of synergy here, and despite my snark, again, you don't get to drop $34 billion on an acquisition without having a sound strategy aimed at the future. Because I can't see it from the outside, it's in fact, the likelier option that I just don't have the vision or perspective to see it, rather than they didn't have one, or it didn't work out. I think it's far too soon to say.

Sachin: Yeah, I think the larger the acquisition, the longer it takes time. And honestly, Aspera has been acquired for almost five years at this point, and IBM’s still digesting it correctly. It's a phrase called bluewashing, and again, they haven't really bluewashed Red Hat. They've done the right thing by Red Hat, and they've done the right thing by Aspera. Bringing us in where we make sense, letting us to be by ourselves where it makes sense, too.

One of the things that drew me to this job, and I'll be frank about it is at an individual level, how do I make IBM more nimble? Lou Gerstner wrote the book, Making an Elephant Dance or something, something like that. Everyone understands that IBM is large. Everyone understands that IBM is complex. Everyone understands that IBM is global in nature, has hundreds of different product lines, and all the rest of it, but the folks there are trying to do their best job and trying to make strategies that make sense and, honestly, the caliber of folks at IBM is really, really good. It's really, really high. We hire really good folks; we appreciate them; we invest in them in a way that a lot of other large organizations don't. We appreciate them. There's a lot of good stuff here. And so I think part of it is, for the first year, IBM and Red Hat have been relatively independent. I think, obviously, the digestion happens over time. And I think you're going to see a lot more Red Hat influence in IBM than the other way around, which I think everyone listening to this would be like, “You know what? That's a good idea.”

Corey: One of the things it’s easy to lose sight of—to be blunt—is what does IBM do, exactly? I take a look at the things that were foundational parts of the IBM experience in my career, and they all seem to either been discontinued, or no longer where the interesting parts are. Easy example is I'm a sucker for buckling spring mechanical keyboards; the type Ms and type Fs were great. The whole IBM Selectric typewriters, those were awesome, too. And moving up the stack into the actual things that were serious, and not spun off, necessarily, like the Lenovo sale of the ThinkPad line, we get into things such as the AS/400, also known as I series.

Well, I started my career selling tape drives into the I series slash AS/400 market. And pro tip: don't do that as a small scrappy startup trying to compete against IBM on backups with an argument of price being the better way to go. Like the old saw, no one ever got fired for buying IBM. Yeah, you absolutely will get fired for scrimping on your backups, my God. Like that is the one thing where you spend all the money if you have a mainframe.

Sachin: Yeah.

Corey: I digress. But these days, you look at these things and it feels like it's legacy technology aimed at a past that sparkles more brightly than the future or the present do. How do you see that? I believe profoundly that you don't take that perspective, or you would not work there.

Sachin: Yeah. Obviously, I don't. And thank you for setting me up with a softball question.

Corey: It wasn't intended to be a softball question. It's one of those—but I said so much in there that you need to—that you technically should push back against, I'm telling you, you don't need to. I'm setting this up as a bit of a gathering the criticisms around this. Again, you know someone in legal or PR is going to listen to this at some point, and those people have no sense of humor of which they are aware. So, I want to be explicitly clear for them, and for the rest of the audience as well, that I am setting up a bit of a straw man for you to knock down. But these are sentiments that have been floated by the larger ecosystem, and me not talking about them doesn't mean that they're not there. But I also don't expect you to be able to refute all of those points, either. Take it away.

Sachin: Look, Corey, everything you've said in there is not anything that's new to any one of us who work at IBM, has heard. IBM has been around for a very long time. As a result, IBM has a very large installed base. IBM does a really good job taking care of the folks who have invested in IBM, not over the years, but over—honestly—the decades. IBM’s really proud of the fact—and this gets back to that client-service mentality I talked about earlier of taking care of each and every one of their customers individually. I think that's really a huge credit to IBM.

And when you talk about other companies, you snark about other folks don't understand enterprise sales. That's the insight. There's no secret sauce, it’s, you solve for the customer—you solve for the problems of the human that's in front of your face that’s asking for your help. And so what does IBM do? IBM tries to scale that, right? We solve hard technical problems that are complex: they require some level of expertise, and whether that's integrations, or making systems work together, it's solving business problems with technology, and that's what IBM is really, really good at.

And so what IBM tries to do is it tries to build these integrated experiences that have a lot of that. So, you can pick and choose what you need, but everything works together, and then you've got the power of Big Blue behind it which means that if you buy it from IBM, you expect some base level of quality customer support, some base level of security, some base level of documentation, and some base level of future-proofing, meaning you're not going to invest in something that's going to get discontinued without an off-ramp that's gentle or whatever else. When you buy something from IBM, you're buying it for a very long time. And IBM recognizes that because they've been doing this for a very long time. For me as a product manager, it's actually really interesting because I can't go and pivot, and do things, and A/B test API calls. I can't do any of that. Anything that I build, I know I'm building for a very long time.

And so as a product manager, you have to be really intentional about what it is you choose to solve and what it is you choose not to solve, and be able to make sure all of the things that you do fit in with the IBM customer base, fit in with the way that IBM goes to market. It's actually a really challenging sort of thing. So, I know I took a question about IBM strategy and brought it back to, like, day-to-day life as a product manager or product leader, but it's actually a really interesting set of resources, and capabilities, and artifacts, and benefits, and a fair set of challenges as you pointed out, Corey.

Corey: There's no good answer in some cases for, well, how does this company wind up revolutionizing everything, and become the glamorous darling startup of the ages? You’re IBM; your existing customer base would not stand for that. When you bring in IBM as a vendor, you are aware of a few things, chief among them for many folks is that they're still going to be in business in 20 years and supporting whatever it is they sell you today, whereas the other end of that spectrum, Google Cloud is selling you something that may not still be supported by the end of the day. So, there's a definite—you can trust IBM to continue to have the long-term relationships with you, and support what it is that they sell you far after, let's be frank, you really should be still on whatever it is that they've sold you. Everything has a shelf life, and some companies forget that, and we're still supporting 60-year-old technology.

Sachin: At least they're supporting it, right? And that's kind of an interesting sort of thing.

Corey: Truly.

Sachin: When folks are supporting that stuff, they don't—doing it because they got on the treadmill and forgot to get off. They're doing it because there's a demand for it.

Corey: I've said it before, but legacy means that it makes money.

Sachin: Yes.

Corey: It's one of those, “Oh, that thing's ancient. Why don't you replace it with something new?” And the answer is, “Because it makes $4 billion in revenue for us every year, so unless you have a plan to mitigate that risk, maybe you should shut up and keep supporting that thing.” [laughs]. There's always this story that we somehow fool ourselves into believing, that newer is always better. It's not the case.

Sachin: Well, and the flip side of that is, think about it from a user, or a customer, or a human point of view. No one accidentally pays $4 billion in revenue and invoices. They're doing it because they're getting some benefit.

Corey: Having sent fake invoices for that amount to companies, you're right, they don’t. It only takes one before my entire world changes, but so far, no luck.

Sachin: Unfortunately. If only life were so easy.

Corey: So, thank you again for taking the time to speak with me today. It's deeply appreciated. If people want to hear more about what you have to say, where can they find you?

Sachin: They can find me on Twitter. I'm @sachinag on Twitter, LinkedIn, whatever. And I think if you Google for that, you'll find me, kind of, anywhere else. Otherwise, we'll find me at a bar here in Oakland somewhere sitting outside with a pint.

Corey: Yes. Someday when I leave my house again, I'll have to join you for that pint. Thanks again for taking the time to speak to us about your thoughts, what you're up to, and make a variety of comments that ideally will not get you censured.

Sachin: Perfect. Thanks, Corey.

Corey: Of course, Sachin Agarwal, lead product and platform for Aspera at IBM, I am Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts. Whereas if you hated this podcast, please leave a five-star review on Apple Podcasts along with a comment about how you will in fact get fired for choosing IBM.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Simon Elisha

As Head of Technology and Transformation at Amazon Web Services, Simon Elisha is sought after by C-Level Executives who want deep insights into combining modern technology innovations with the pragmatic lessons learned from hard-earned experience. An expert in Cloud Computing and Organizational Change needed to get the most out of it; Simon is able to demystify how technology innovation is best applied to enable organisations improve customer experience, reduce costs and adapt quickly.

Simon was a leader in cloud well before it was mainstream. As the first technical staff member for Amazon Web Services in Australia, he led the charge to Public Cloud. Bringing over 30 years of industry experience in software, infrastructure and business consulting to the “brave new world”– he has guided start-ups, digital businesses, Government agencies and Enterprises alike on their journey to the cloud. He now leads the AWS Solutions Architecture team located in capital cities across Australia and New Zealand.

A noted industry speaker and communicator; as host of the AWS Podcast, Simon speaks to a global audience of technology leaders and practitioners on a weekly basis. Simon has held senior roles at organisations including Pivotal Software, Cisco, Hitachi Data Systems, VERITAS Software, PriceWaterhouseCoopers and EDS. In addition, Simon earned an Honors Degree in Information Technology from Monash University and holds eight patents.

Links Referenced:

  • AWS Podcast: https://aws.amazon.com/podcasts/aws-podcast/
  • Ensuring Rollback Safety During Deployments: https://aws.amazon.com/builders-library/ensuring-rollback-safety-during-deployments/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Simon Elisha, head of technology and transformation, Australia and New Zealand public sector for a little company called AWS. Simon, welcome to the show.

Simon: Hey Corey, long time, first time.

Corey: Exactly. Which I imagine is one of your colloquial expressions, which means it's great to talk to me. And you're right: it is. So, what is it that you do at AWS because I understand that you run the AWS podcast, which is a subject near and dear to my heart. Feels a lot like this one, except you probably spend less time trying to figure out who should sponsor it.

Simon: [laughs]. Probably true because that was one of the considerations I didn't have to worry about. So, I wear a few hats here at AWS. So, firstly, I lead a team across Australia and New Zealand who work with our public sector customers on building, deploying, and getting the benefit from really cool technology for citizens. So, kind of nice and very, very rewarding.

But also, as you mentioned, I do host the AWS podcast as well, which was a crazy idea I had back in 2012 when I actually went to search for podcasts about AWS, going, “I'd like to listen to one,” and there wasn't one. So, I did that crazy thing of saying, “I'll make one. How can it be?” And now I learned that having a podcast is like having a puppy. It's for life, not just for Christmas. And I really enjoy making it.

Corey: Oh yes, having spent a little bit of time, myself, around AWS folks, I know that if you play the popular drinking game of taking a drink every time someone in an AWS meeting uses the word customers, you will die. So, you just mentioned that working in the public sector, you work with citizens. How often do you wind up accidentally using one term instead of the other and having to self-correct? Or is that something you've been able to train yourself out of doing?

Simon: It's more the case that it’s a different mindset when you think about citizens versus customers. And our customers are the departments and the agencies that we work for, but their customers are the citizens. And the distinction I like to make is a citizen can't choose the service they choose. So, you have your tax department, and that's the one that you'll be providing your tax details to. You have your immigration department, your defense department, your health and human services department. You don't get to choose.

A customer gets to choose, a citizen get a service by their government. And so my goal is always to help out agencies provide the best possible citizen service. So, it's an interesting nuance, but it's an important one. And the nice thing is a lot of governments are speaking about citizen-centric services, which ties very much our own customer-centric thinking.

Corey: I like that quite a bit. One thing that I find interesting is in my own mapping—mentally—of big companies, my impression has always been that the bigger a company gets, the more inherently narrowly defined various employee roles become. So, it's interesting to me that you're not the, for example, head of the AWS podcast on the following four days out of the week. Instead, you have another full-time job that isn't speaking into a microphone. How did the podcast come to be, and why you, I guess, is probably the rude version of that question?

Simon: [laughs]. It's because I have an absolutely beautiful head for podcasting is what I would describe it as.

Corey: Yes, a face for radio, is what I have on this end.

Simon: [laughs]. But look, I think it's interesting is that one thing that we do when we hire folks at Amazon, and AWS in particular, is we want to hire builders and let them build. And so what that means is we hire, looking for people who the leadership principles really resonate with them. You know them just as well as anyone else, so I won't list them out here, but what those leadership principles do is keep us these broad running rails and decision making filters to go do stuff that is really, really good on behalf of our customers. And so what that means for me is when I sort of saw a need, which was, “Hey, there's no podcast. Maybe people would be interested in this.” I could go build it.

And yeah, I had to talk to some of the right folks and say, “Hey, this is something I'm thinking of doing. Can I do it?” And I [unintelligible] trust with those people. They said, “Yeah, it sounds like a great idea. Give it a go.” See what it is, it's what we would call a two-way door. So, it's a decision we can unmake later on if we want to. And history tells the title, that a lot of people found it useful. And probably one of my favorite things is especially people that I meet at conferences, saying, “Hey, I really love the podcast.” That’s really gratifying. But more importantly, is when people say, “Hey, I got a job and your podcast helped me prepare.” That’s, like, awesome. Very happy with that.

Corey: Of all of the feedback that I get around the different conversations I have with folks from this podcast, and newsletter, my being obnoxious on Twitter, the most common is, “What's wrong with you?” And other various forms of insult. But the ones I like the most are where I've helped someone do something next in their career, where it's getting someone from where they are to someplace else. Maybe it's career based, maybe it's solving a pernicious problem. But that's always meant a lot more to me than the slings and arrows I get, which, let's face it, I definitely invite that criticism onto myself with basically everything I say and do. But there is something to be said for having been in a position to impact people's lives in a positive way.

Simon: It is humbling, and not a sort of twee way, but if I think about it—like, I've been in IT for 30 years now, I'm an old person, now, relatively speaking, but I think back to those folks who helped me when I was a young whippersnapper and gave me guidance, gave me an opportunity, took a punt on me and said, “Hey, this person can do something. Let's let him do it,” type of thing. And you’ve got to pay it back. There's so many amazing folks coming into the industry for many different aspects. One thing I'm really conscious of is not just saying, hey, he's an IT graduate, that'll be perfect.

Well, what about someone who's had a completely different career trajectory, but he's really interested in this domain? How do we get them into that? Giving that assistance is huge. And the weird thing is it has a really outsized effect compared to the input you put in. So, much like yourself, you put work into the podcast, you do it, but you kind of get it done. It's done; it's out there. And what you don't think about is there’s people listening, sometimes years later, saying, “God, this is really useful to me. This is inspiring me, getting me to that next level.” Can’t ask more than that.

Corey: A realization I was somewhat late to was the, I guess, dawning awareness that a lot of what AWS was releasing—things like SageMaker, or the DeepLens, or the DeepRacer, or the DeepComposer—also known as Dr. Matt Wood’s piano recital at re:Invent last year—is, sure, on some level this stuff is fun and goofy, and an easy target to mock, but on the other, what what you're fundamentally doing with these things is making new fields available in a fun and engaging way to people who might otherwise never go in that direction. And that's a very powerful thing.

Simon: It is. And that's one of the things, particularly DeepRacer as an example is something I've seen so many people grab that and use it to learn. And they've sort of done it because, yeah, writing remote control cars with computers, who wouldn’t want to do that? But afterwards, they're like, “I learned so much about how I can apply this, and importantly, where I can't apply it as well.” So, just because you have a tool doesn't mean you should use a tool.

So, one of the things about DeepComposer and DeepRacer and DeepLens is to learn where's the best fit. In my toolkit of things what should I use, and how should I use them? And by making them more available and fun—fun being a relative term—it means that people can get their hands on these technologies, and that's the big shift. If I reflect on, sort of, what's changed a lot over the last 30-odd years, 30 years ago, if you want to use a new technology, you had to get on a project that used that new technology, and that was hard. Whereas today, it's like, “Oh, I’ll just spin up my account on AWS and I'll just get a DeepComposer, or I’ll do DeepRacer. I’ll spin up some SageMaker, or whatever.” The barrier to entry is much lower, and that's really exciting.

Corey: That really is the differentiator here. And I'd argue that is the real transformational power of Cloud, where I know I'm a little late to the observation here by about 15 years, but the fact that I can have an idea, something I want to experiment with, and more or less run a single command, and, these days, within seconds—once upon a time, within many minutes because, hey, the provisioning plane needed to evolve from time to time—I could have an environment set up and ready to go. And when I was done, I could then, in theory, just turn it off and never be billed again. In practice, I would then spend 22 cents a month for the rest of my life for ancillary resources that wound up being spun up that I could never fully track down. Which brings us to, I think, a topic that is near and dear to both of our hearts. Specifically, the idea of—not necessarily reducing costs because that topic gets done to death, but more about understanding and attributing it to different aspects of the environment. Tell me what you're seeing.

Simon: Yeah, I think you raise a good point, is that really the high-value discussion, the discussion you want to have with your CFO, or COO, or whoever pays the bills is, “Is what we're doing valuable to our organization as a function? Is it worth spending the money? And is it worth spending this amount of money?” And the beautiful part of Cloud and AWS is that you can track down to the cent, what you're spending on a transaction, on the system, on an operation, on a development process, et cetera. And that puts you in the absolute box seat to make decisions about whether it's worth even doing.

And then secondly, you can think about, from an architectural standpoint, what are called economic architectures. And this is where you sit down as an architect or as a developer, and you make informed choices about design patterns you apply that effect in price is a consideration. Now, this is really different to how the model used to be in my day. In my day, we’d specify all the hardware, software, storage, network, et cetera, that we thought we needed upfront before we even built the software, and hoped that we got it right, which we, spoiler alert—

Corey: Oh, yeah, we pushed the purchase request uphill both ways.

Simon: Yeah, absolutely. Never got it right. So, we either had too much or too little. And really, what's important here is to think about, am I using the best possible tool? Is it the most efficient? Does it deliver my functional and non-functional requirements? So, am I getting availability? Am I getting security et cetera? Am I avoiding some operational cost in the future?

Well, honestly, my projects at the moment, I [audio break] think I've spun up an EC2 instance in months because I just do everything serverless now because I just don’t like patching systems, and I don't have to. So, that has an ongoing benefit. And so looking at holistically, really, is the difference here. And one of the trends that's happened is DevOps and DevSecOps—or call that domain whatever you will—but the better understanding of developers in terms of how these systems operate in the long term, and the operations teams in terms of some of the decisions that get made early on in the development lifecycle that have long term repercussions is enhancing and improving substantially. And the organizations that start to get that right tend to be much happier places, and they deliver systems that have better value and can show back that they're running as they should run.

Corey: That's an increasing challenge. And I find that there are two reasons that that's challenging, at least in my experience. One is the obvious of there's an awful lot of moving parts, and not everything is easily attributed back to a single cost center—shared services make that a challenge—but the other is that when you're having this weird cultural dynamic where there isn't a culture of cost attribution, of understanding how to transform legacy processes into something that is more dynamic, it seems like a win when once upon a time, it took us six weeks in a good day to wind up getting a server provision, and now with the Cloud and the company's processes, it only takes four. But it feels like there are opportunities to optimize that. And that, in turn, leads to other weird anti-patterns, such as if it takes that much work to get something spun up, you'll never give it up because you might need it again, and who has that kind of time to kill?

Simon: Yep, yep. Let me, maybe, tell a story that relates to that, and answer that question because I think it's a good point. And this happened a few years back with a customer that was all-in on AWS. And they had a big developer team, like hundreds of developers. And they had set a EC2 limit of about, it was about 500 concurrent EC2 instances, and they were always at that limit. And when they dived deep on it, they discovered that the reason was exactly what you just said: the developers would not release the systems because they were so close to the limit that they were worried that they wouldn't get one back.

So, counter-intuitively they upped the limit to 750, and the actual usage dropped to 400. So, they saved money by increasing the limit, and it was a really interesting study of human factors of that supply and demand, and concept of scarcity. If there’s scarcity, human beings tend to hold onto things, it's just kind of how we’re wired. If there's abundance, then we're not worried about it. And so to your original point of how do we take the time to understand cost? Who cares? Et cetera.

At some point, someone is paying the bill, and what I find works really well is getting as good a handle as you can into your environment—and tagging is a big part of that. And there are some great tagging strategies you can use to get a view of cost at various levels—but the other thing is having the right conversation with the right people in the organization. And this is where—you know, IT and finance don't often talk in detail. They tend to talk at the high-level numbers, the CIO goes and submits a budget, gets some budget, applies it, et cetera. What I've found is that when we bring the finance department into the conversation and show them how much granularity we can see around our spend compared to our utilization, we're suddenly speaking their language, and they're like, “Well, I'm invested in this. I want to know more. What can you show me? How can I understand more about what we're doing?”

And then, once I understand the visibilities that I have, that then informs reducing or lightening the governance frameworks around it. Because once I can see something, and I can see it in near real-time, I don't have to put these heavy gates in front of me. So, like you said, I don't have to have a four-week delay to get an EC2 instance if I know—if you spin it up and it doesn't heed the policy, it'll just get turned off automatically within five minutes. I'm good with that. That changes that mental model, and having those high-level conversations back with data is really changing some of these organizational structures and governance processes.

Corey: This is a difficult thing to achieve in most corporate environments. I can't imagine how much more difficult it is when you have a lot of these practices enshrined in the law, or process that is so documented and required by so many other aspects, it may as well be law. How do you begin to pivot the culture around, I guess, a transformative opportunity that was never even considered when a lot of these laws and processes were written?

Simon: I think the thing to remember is that there are many stakeholders for whom change is their goal. They don’t come to work going, “You know, I want to keep it exactly the same as always been, and it shall always be ever thus.”. I’m not saying everyone comes to work not thinking that, but there's a big component of people who like, “Hey, how can I make these better? How can I improve this? I don't want to sit in an architectural review board meeting every single week for six hours to approve something I know is going to be ok.”

So, tying into that cultural change, and again, finding those right stakeholders, giving them the information, and having the right conversation is really important. And let me give you a for instance, or an example of this. I remember many years ago, many years ago, early on in the days of Cloud, I met with a significant bank. And I had a workshop with the security team. This is 15 of their heavy-hitting security architects that basically got the opportunity to say, “Hey, Simon’s going to come in, you got an hour and a half with him. You can ask him any question, throw anything at him, have fun.” And they put me through the wringer for an hour and a half, and at the end of that session, they said to me, “Everything's going to be on Cloud in the future. This is fantastic. What can we do to help?” And I was like, “Okay. That's an interesting lesson learned, which is that there are people who—they are there to protect and apply governance and requirements, but also to evolve thinking based upon new data. And one of the biggest challenges I find is people aren't always aware of what is possible today. You work around AWS and Cloud all the time, as do I, so we just assume, “Oh, spin up an instance, create a VPC, or spin up a database or—”

Corey: Or talk about new service, and then have AWS employees look at me and wonder if I'm making the service up because who can possibly keep track anymore.

Simon: [laughs]. But that's a mental model that we're very comfortable with because we live it all the time, but for a lot of people, they've never worked, or lived, or really been exposed to that kind of velocity, or opportunity, and so they get delighted when they can do that. So, there's a lot more latent change opportunities than you may think.

Corey: And that’s always a challenge because looking at, I guess, the trajectory of companies as they go through various ‘digital transformations’—which I despise the term, but don't have a better one, so I begrudgingly use it but always with a cynical tone of voice—and you see that they go from this idea of data centers, where everything is, effectively, very fixed in terms of cost, and then as they move through the, I guess, spectrum of how Cloud-y or not something is, it becomes a lot less expensive—in theory—and it winds up instead billing more upon usage. And at some point, you wind up hitting the pure serverless ideal of transaction-based billing, where it’s, oh, I can now trace, as Simon Wardley says, the flow of capital throughout my organization. I haven't really seen that yet because almost nobody is full-on serverless to that degree, and the few shops that are—for example, A Cloud Guru is famous for this—but during the Serverlessconf, they get up and show their AWS bill and it's something like 500 bucks a month. It doesn't actually mean enough money to matter. So, what does transaction-based pricing start to look like?

Simon: So, one of the things with that is, firstly, it's often really, really low. And that's scary, as you said because we're not used to dealing with those types of numbers. But what it looks like is an evolution in certain systems within organizations versus the whole thing. So, again, this is a continuum. You're right, for a lot of organizations, this is hard to change; there's an inertia that builds up over time.

But much like planting a fruit tree, the best time to plant is 10 years ago, the next best time is today, people need to start. And so we have lots of customers who have saved huge amounts of money, like millions and billions of dollars just by changing their operational model. Now, they haven't necessarily had to say we're going to go all-in on Cloud, or we're going to move everything across. They've picked their nastiest, most expensive, problematic, non-scalable system, what have you, the one that's the most opaque, and chosen to attack that and deliver that much less. Now to your story, sometimes it's so much cheaper that people don't necessarily want to talk about it.

I was working with a customer once, a little while back, who was replacing a major system, this system probably cost them—just thinking—it was about $10 million a year to run this system traditionally. And we said, “Okay, we can lift and shift, make a few changes, cloudify it a bit, run it serverless, and it'll cost you $1 million dollars a year.” And they said, “We can't go back to our leadership and tell them that.” I said, “Why not?” They said because it's too cheap, and we're embarrassed that we spent $10 billion over the last few years on the same system. And I said, “I hear you, but that's not actually the problem here. This is a big win. It's not about the decisions you made in the past with the best knowledge you had at the time. It's about the decisions you make now, knowing what you know.” And that mental model shift is really important. And doing it on a case-by-case basis is really important.

In what you might be forgiven for mistaking for a blast from the past, today, I want to talk about New Relic. They seem to be a relatively legacy monitoring company, and I would have agreed with that assessment up until relatively recently, but they did something a little out there: they reworked everything. They went open source, they made it so you can monitor your whole stack in one place. And most notably from my perspective, they simplified their pricing into something that is much more affordable for almost everyone. There's even a free tier with one user and a hundred gigs per month, totally free.

Check it out at newrelic.com.

Corey: And that's something that I found is that one of the hardest parts. The technology is relatively straightforward. The getting people to a point where they can have that pivotal moment, that shift in how they view these things, is always the hard part. And even now, going from the idea of—for me—running on a bunch of virtual instances that are running everywhere. Great; now moving to containers, heaven forbid, or serverless is still a whole ‘nother sea change.

But backing up, how do you even, in today's world, go from a time of understanding the world through a lens of data centers in computing to moving to Cloud? Because to be blunt, when I did this, there were a lot fewer services in Cloud that I had to worry about. I looked at the AWS Console, and, oh my god, I'm never going to be able to learn about all of these services. And there were 12. Now, there's significantly more than that, and I don't know where to even begin in a modern era. Where do you stand?

Simon: Yeah, so let me tackle that from a few places. So, firstly, one thing that I think is really important in a long term IT career is continual learning. And I think many of us start our careers off loving to learn new stuff, and then we kind of stagnate after a while because life gets in the way, but if we're not relearning all the time, we're not going to be able to deliver the best for our stakeholders, for our customers, for our sales, what have you, take advantage of the latest and greatest. Now, I'm not saying always use the cutting edge, the bleeding edge, et cetera. I'm saying use the things that make sense given the best of what you know.

The reason why I mention that is the way I look at all the services we have for customers is, it's kind of like a painting palette. You might not use every color in the palette, you just use the ones you need at the time. So, the mental model I use is to say, “Think about what you're trying to achieve, and then pick the services you're using.” So, for example, you just mentioned, “Hey, when I host my website, how would I do that?” It’s like, okay. Hosting a website; that sounds like some content: lives on S3. Sounds like I need to give it to some people: CloudFront probably fits the bill there. Probably need to have some security, HTTPS: that's Certificate Manager, strap that in there. And then maybe I'll do some logging, some CloudFront logging, maybe I'll report on it using Athena, and I might call it good. If I want more functionality, I can add more functionality, but I can just stop there. And I don't need to go super deep in each of those services, I just need to know they’re kind of there. And I can use them. Because I want to tie into your drinking game, Corey, so I can have a drink. Is, because our roadmap is 90 to 95 percent driven by customer requirements and feedback. They're telling us what they would like us to build, and take off there [unintelligible].

Corey: That 5 percent gap, the things you release that no customer asked for, is amazing.

Simon: [laughs]. Well, you know, the thing is, is that we don't always know what we want. So, sometimes we need to see things that we didn't think of. But in terms of your question, “We've got this wall of services, what do I do?” It's about thinking about what you're trying to solve for today. And assuming that, well, I'm sure other customers have asked for this, too. Let me go check if Amazon has tried to solve this on our behalf. And usually, the answer is yes.

Corey: One of the problems I have is that invariably AWS has this incredible knack for releasing a service that solves those global pain points directly after I've implemented something badly to work around it. I think one of the terms that Forrest Brazeal likes to use is ‘spackle punched’ where, “Oh, I built this ridiculous janky thing. And now AWS has spackle punched me, and solved this problem globally.” Which is awesome, but it would have saved you a couple weeks worth of effort, if it had just come out a little bit sooner. Is that just me with my terrible timing, or is that one of those universal moments?

Simon: I think it's just a feeling that you get with all technologies in a way. But I think if you think about it, tying into it say, “Well, If I'm having this problem, and I'm spackle-filling, as you say, then there's probably literally thousands of people at the same time all around the world doing that.” And then, as you say, it becomes a timing issue. So, whereas you may have already built it, and now you get to replace it, there will be many, many, many more people who will never have to build it in the first place because it's already there. Now, related to that, I might make an interesting counterpoint here is that experience is really important, and understanding how things work versus how things are, gives you a relative concept.

And you talked about people not necessarily understanding how hard it is to get a server racked up, et cetera. Now, I've a got large cohort of solution architects who work in my team, and many of them are much younger than me, and they've never seen a data center; they've never had to order a server after a PO process, et cetera. It's always been on demand. So, we actually created a little learning series to say, “Here's how we used to do it. Just so you understand what was involved.”

And that was fascinating. When you sit there describing connecting to a storage area network using HBAs, you watch people's eye roll in the head and go, “Well, that's not a career I would have chosen,” compared to now just going, attach EBS volume, get on with my day. So, understanding what came before helps you understand how much better things are now, but it's no excuse not to continue to make them better. And so, we will continue to feel that spackle wherever we can.

Corey: And again, I'm not sitting here saying that we should stop the march of progress, or well, it might upset someone who's already built something janky, so you should never release a service. But there is that moment of, at some point, I take a step back, and I realize that, huh, I'm spending an inordinate amount of time solving what feels an awful lot like a global problem. Maybe that's not the best path forward. Last time I did one of these things in earnest was about trying to get replication working with encryption on RDS between AWS regions. Today that's, click a button, but at the time, it was not.

And that was painful and challenging, and I'm looking at that, and it felt like I am probably not the only company in the world that has this problem. Maybe there's a better way forward? I can't shake the feeling that by going down this path of either cloud agnosticism with going multi-cloud or building everything in your own data center yourself with an eye towards I want to avoid lock-in at all costs, you're effectively having to build those solutions that can be done for you by an organization that focuses on solving those global problems for you, like AWS, case in point. I feel like it's so easy to wind up getting wrapped around your own axle of, “So, what are you doing right now that adds business value?” It’s, reimplementing a load balancer doesn't really seem like the right answer, unless your business is load balancers.

Simon: Yeah, yeah. I think the concept of undifferentiated heavy lifting—[take] a drink—is really important because, as IT people we know, we've built stuff that’s really hard to do and you're kind of proud of yourself. Hey, I made this work, but my goodness, that was hard, and no one really cares that I did it. And really what it's about is doing as little as possible to get the outcome you need. And I remember early on in my career, I studied software development, and I got to work with some really, way smarter than me developers who really knew their stuff, and one of them took me aside one day and said, “You need to learn to be a lazy developer.” And I'm like, “What do you mean? Lazy? That's bad.” “No, you need to learn to do as little as possible to get the outcome you're trying to get for the software. Write as little code, be as efficient as you can, use as few services as you can, reduce the complexity, reduce the time it takes to get done.” I was like, “Aha. That's a really interesting insight.”

And that was 30 years ago, when I spin forward to today, what it means to me is, when I'm building a system, I'm going to work really closely with the cloud provider of my choice, I'm going to write as closely to their APIs as I can because I know if I want to make a change, I'm just going to change a few APIs, talk to a different service provider, use a different service, drop-in replace it, whatever, but that's going to get me to my outcome quicker, which means I get my solution in front of my customer quicker, which means they can tell me whether they like it or not, and I can either make a change, or double down on it. And so, it's that mental model of building as little as possible as quickly as possible that, really, has helped in this current environment.

Corey: So, a question I have for you that is a common refrain on this show—it’s, sort of, our themes go—if you were starting today and you're at the beginning of your career—because let's not kid ourselves, you and I have been doing this for decades at this point—where would you start? And at some point, are you done? I mean, is there a place where, “And now I have learned the cloud. Box checked.”

Simon: [laughs]. Check. It's a good point. And to give you some context, I mean, when I started, I started on mainframes. So, I learned CICS, COBOL, DB2. I can talk about JES queues, and ISPF till the cows come home. I still miss it. But then I had to learn client-server, which was the brand new hotness at the time, and then the web came out, et cetera. Really what the lesson is, is that we're always learning, so never get too fixed in your mental models around technologies.

Now that said, if you're starting today, there are some really good foundations to build upon; so lucky. And so, one of the things I point to customers do all the time, now that's available is something called the Amazon Builders’ Library, which is really a set of blog posts and explanations about how we build and operate software at scale. And this is not saying this is the only way to do it and this is how you should do it, but this is saying, well, this is what we do at global scale, and it seems to work pretty well. You might want to learn some things about this. So, now one of things that I really like talking to customers about is how to deploy software and rollback software.

Super hard problems to solve, but we do it all the time, many, many times a day as you can imagine. How do we do it? What does testing look like in that environment? What is validating that's going to work? How do you roll features forward, roll them back, maintain compatibility? All this stuff, it’s already there. There's a lot that exists to learn upon. Even simple things like you mentioned, “I want to get started in machine learning. What do I do?” Well, even if you jump onto the AWS Console, there are pre-built videos, labs, instructions on how to get up and running. So, you can at least learn and stand on the shoulders of giants to get going. What is existing now, we'll look back on in 5, 10 years ago, and say, “Well, how cute. You're doing this, you're doing that. Look how far we've come.” But you can do that at any point in time in computing.

Corey: Now, it's six lines of YAML.

Simon: Yeah, exactly. You can do it at any time in your career. I mean, now, I joined AWS back in 2011. I think about some of the stuff I built back then that, like you said, I can build now with literally a few lines of code, but back then I was going, “Wow, look. I didn't have to spend six months doing this. I did it on my kitchen table.” Just maintaining that mental flexibility is super important. And you’re right, you’re never done, and I think that's really hard for all of us as technologists because we tend to be quite mathematical in our mindset, which means we like complete proofs, and finality, and a solution, and it's solved, and it’s done. And I can, like you say, I can tick that box.

Corey: I've gotten to the end of AWS. The final boss was super hard.

Simon: Exactly. Whew. All done. And you just can’t. And it's hard doing that. And I can tell you that that's something Amazonians struggle with all the time because we'd love to know everything. And I'll share with you a personal story, there was a brief shining moment where I could, hand on heart, say I knew AWS. And, like you said, it was early on. We didn't have that many services. I knew them upside down, two ways from Sunday, inside out. I knew them all the way. But you can't now. There’s, like, over 175 services. [laughs].

Corey: Meanwhile, one of the early engineers who's there and still there, and is now Distinguished Engineer or something, is just sitting there mumbling, “No, you didn’t,” Because there's always another level to go down to.

Simon: It does depend what level you go down to. But as far as I was concerned, I had done it. I was on top of my game, I knew everything I need to know. And then we released another five services, and another ten services, and just [unintelligible].

Corey: Programming, I got it. I have both languages: JSON and YAML. Yeah, so one of the challenges we see, too, is that at some point, you can't go by experience anymore. When I was starting out in my career, and I was trying to sound like I was someone a company would want to hire, there was a point where I wanted to add as many years as I could to experience.

Now, on some levels, I kind of want to shave some off because your skill set is what it is, and how long it took you to get there is in some ways, an interesting metric. I guess the depressing end of the spectrum is I've met people who've been working in tech for 30 years, but they don't have 30 years of experience. They have one year of experience repeated 30 times, and that always really depressed me because at some point, the tide rises, the thing that you do winds up getting washed away, and there aren't very many opportunities to continue doing that one thing. It feels like tech is one of those areas where you have to reinvent yourself the entire time that you have a career.

Simon: Yeah, I think what we say at Amazon is we say yeah, we want to be stubborn on the vision, but flexible on the details. So, being stubborn on the vision is, “Hey, I want to be a technologist, I want to build systems, I want to solve problems, that's what I want to do.” But whether it's COBOL, or C, or Python, or PowerBuilder, or Delphi, or Visual Basic, I don't really care. Like, I care at the time, but I'm not going to be bound to it and say, “Well, there is only one true language from here on out.”

Corey: Oh, I was so angry about things like that back then. Oh, god, picking fights about programming languages, and systems, and architecture. That was one of my favorite things. It turns out, I was a terrible person. Some of us evolved past that, hopefully.

Simon: You've recovered. You've recovered. But it's true, but that's the thing is, we tend to get into these, kind of, almost philosophical arguments about this chipset versus that chipset, or this operating system versus that. It just doesn't matter. What matters is the outcome you're getting, how easy it is to run, easy to manage, easy to learn, et cetera.

And when the time comes to replace it—because we're all building the legacy systems of the future—how easy is it to replace as well? That's what you want to think about. And that'll set you up for a long satisfying and invigorated career, versus just fighting these battles that you’re just going to lose eventually. You can’t win that conversation.

Corey: And that's part of the challenge is, it's even hard to talk about because it’s, on some level, you definitely don't want to be coming across as saying evolve or die, dinosaur, but at some point that, that's kind of what you do. I mean entire jobs that were things when I started my career—like firewall engineer was a six-figure salary, if you had that skill set—now, more or less—I was going to say that, wow, any basic network engineer should have that skill set, but even beyond that, today, it's kind of pretty much—do you know how security groups work? Spoiler. No one knows how security groups work but roll with me here. And that gets you where you need to be, the baseline level of experience is necessary. How do you find that the fundamentals—the things that I guess we had to learn at one point because there was no other option—manifest today? Are they still necessary?

Simon: Well, I think the detail is not necessary, but I think the fundamentals are. And the fundamentals don't change. Now I need to build a system that's resilient. Well, what does resilient mean? Well, resilient means a lot different now than it did 30 years ago. I need to build something that's user friendly. Well, when I was at study at university, I remember clearly my lecturer telling me, “For a good user experience, when they hit enter on the console, it should take no more than four seconds for it to come back. That's the baseline of good user experience.”

Corey: At this point, it can go to the moon and back.

Simon: I know. It changes a lot. But understanding that you need to care deeply about user experience will get you a long way. Understanding that people don't really care about the details of technology—depending on their perspective—so as IT professionals we care deeply we're all about it all the time. That's what we love, we're passionate, we're enthusiastic. See, most organizations doesn't care. They just want to know, does it work? Is it fair value for money? Does it put me ahead of the competition? That's it. That's it.

Now, whether it's all singing, all dancing this, or an all singing, all dancing that, they don't really mind; they just want to get it up and running and done. So, learning how to speak to non-technical stakeholders in an accessible and meaningful way is a super valuable skill. And let's face it, Corey, you’ve made a career out of it. And I'm sure it wasn't a course that you did, or an instructional video you followed. It was realizing, “Hey, I need to talk to these people that don't look like me from a technology perspective, but need to hear what I'm saying.”

Corey: One last area I want to talk about with you before we call it a show was announced last year at re:Invent—back when we would all gather in the same place, and not worry about our lives—was the Amazon Builders’ Library. And that is ‘Builders’ Library’ not ‘Billers’ Library’ which presumably is a compendium of all the different API calls that you can make that will cost you money in embarrassing ways, obvious only in hindsight. Talk to me about that because I love it personally, but I want to get your take on it.

Simon: Yeah, it's one of those great things that we have a lot of really experienced engineers building software here at Amazon. And Andy Jassy often says there's no compression algorithm for experience, and he’s right. And the other way I like to look at it is, I'd much rather learn from other people's mistakes than my own because then I don't have to suffer from the pain of that. And this catalog, the Builders' Library, really showcases what we've learned. How do you build a system that is distributed, that can recover from an outage without the thundering herd of new transactions coming in, thus storing it again and creating a repeatable loop of [unintelligible]? How do you build a Continuous Integration/Continuous Deployment Pipeline that genuinely works at scale and you can easily roll back from? How do you deploy technologies like Shuffle Sharding, or leader election, or all these other interesting things that are highly difficult problems, but have been solved and solved effectively at scale?

And what this library does is gives you all that information for free. Like, if I was a graduate student today, I would just sit down and read that for a day, and understand it deeply, and go, “Wow. I've just saved myself 20 years [laughs] to learn all that.” And so I'm really excited that that's publicly available and continues to grow and have new content placed on it because it is genuinely—now, this is not about Amazon. It's not about AWS, it's just about building software at as high quality as you can.

The lessons you learn, even simple stuff like introducing jitter in the way you handle workloads because that improves your ability to handle it at scale. Like, just stuff that you wouldn't necessarily think about in your day to day work that can inform your own design decisions and the way you choose to build software and may give you ideas to improve the way we build software in the future. So, it's a big one, to go to the Builders' Library.

Corey: Do you think that it has the potential to cause problems for folks, in the sense of architecture as imagined by Hacker News where someone is building out something to work internally at their company, and if they follow every tenant laid down in the Builders' Library, if assuming such a thing was even possible, they would be building this world scanning system that more or less would run payroll once a week for 200 people. It feels like it could lead to scenarios of stupendous overkill, and more or less writing code for the joy of it, rather than to solve a business problem. Do you think that that's a realistic concern, or am I not giving people enough credit?

Simon: I think it comes down, again, to that mentality of building only what you need and evolving the architecture over time. And that's where this information is baked into that. And there's actually a talk we do at re:Invent—I did it many years ago, my colleagues do it now, we evolve every year, which is the scaling to 10 million users—or it could be 100 million users now, but really talks about moving from, I've got one EC2 machine, now I’ve made it highly available, now I'm scaling out, et cetera, those thought processes to go through. You're not building the end state. You’re building the start state and evolving. And I think a lot of the lessons in this can help inform when you might need to reach for that technique in your tool belt based upon the scale you're at. It's like, “Ah, now I've got this problem. I understand what they are talking about. I know how to solve it,” versus, “I didn't know this would be a problem. Now it's a problem, and I don't know what to do.” So, it positions you better to tackle the future.

Corey: If you were to dive into the Builders' Library today, what would you start with? Because it turns out that there's an awful lot—I wouldn't say an awful lot. Not by Amazon service terms—but there are enough documents in there that it could be challenging to pick which one to start with, and on some level, they get very deep very quickly. Is there something that you would say is the most accessible way to get started?

Simon: If you had to read one, it would be the one that's called Ensuring rollback safety during deployments. That's the one. Because that's all about risk management. It's about saying, “Hey, how can I deploy software frequently and safely?” And it solves a lot of problems. What if I moved from XML to JSON? What if I change my protocols? What if I change a database? How do I test that I can roll back? What does that really mean? And there's a comment in that particular article that I had a chuckle about because it says, “At Amazon, one of our leadership principles is frugality, but we don't believe in frugality when it comes to testing.” And I thought, “Yes, [laughs] that is correct. That is the way to think about it.” So, that article, I think, is an absolute go-to. If you read one article, that's the one to read.

Corey: Excellent. We'll throw a link to that in the show notes. Simon, thank you so much for taking the time to speak with me today. If people care more about what you have to say, for some ungodly reason, where can they find you?

Simon: Well, they can indeed find me at the AWS Podcast. It's available where all good podcasts are caught. So, the podcast catcher of your choice, and if you search up AWS Podcast, it will be the first it webpage-wise as well.

Corey: One day I will beat you on that listing.

Simon: Challenge accepted. [laughs].

Corey: [laughs]. Thanks so much for taking the time to speak with me today and suffer my egregious slings and arrows.

Simon: Always a pleasure, Corey.

Corey: Simon Elisha, head of technology and transformation, Australia and New Zealand public sector. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts. And if you've hated this podcast, please leave a five-star review on Apple Podcasts, along with a detailed comment filling out exactly why you felt the need to dislike it, in triplicate.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Jaana Dogan

Jaana Dogan is working on Spanner at Google to make state not your problem problem. She has 15+ years of experience in building infrastructure, developer platforms, and tools. Jaana's current work is focused on storage systems, observability and performance tools, and helping customers with architectural design tradeoffs.

Links Referenced:

  • Recommended book: https://www.amazon.com/Designing-Data-Intensive-Applications-Reliable-Maintainable/dp/1449373321
  • Twitter: https://twitter.com/rakyll
  • https://jbd.dev/
  • https://rakyll.org/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: It's, at least as of this recording, morning on the West coast, which means there's no better time than to inflict a homework assignment upon you in the form of a 42 page ebook from StackRox. Learn about the dancing flames of EKS cluster security, evade the toxic dumpster of the standard controls, and tame the wild beast of best practices for minimizing the risk around cluster workloads. Become renowned for your feats of daring, as you learn the specific requirements for securing an EKS cluster and its associated infrastructure. To learn more, visit snark.cloud/stackrox. That's snark.cloud/stacROX.

Corey: This episode is brought to you by Trend Micro Cloud One, a security services platform for organizations building in the cloud. I know you're thinking that that's a mouthful because it is, but what's easier to say? "I'm glad we have Trend Micro Cloud One, a security services platform for organizations building in the cloud" or "Hey, bad news. It's gonna be a few more weeks. I kind of forgot about that security thing." I thought so.

Trend Micro Cloud One is an automated, flexible, all-in-one solution that protects your workflows and containers with cloud native security. Identify and resolve security issues earlier in the pipeline and access your cloud environment sooner, with full visibility, so you can get back to what you do best, which is generally building great applications.

Discover Trend Micro Cloud One, a security services platform for organizations building in the cloud. Whew. At trendmicro.com/screaming.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Jaana Dogan, staff engineer at a small company called Google. Jaana, welcome to the show.

Jaana: Hi, how are you?

Corey: I am very well, and I'm better, now that I get to talk to you. One of the—I guess, not one of—the best database in the world is DNS, and that is a hill I will die on. Almost as impressive is a product that you work on, namely Spanner. What is Spanner? And why would someone care about it?

Jaana: Spanner is our relational, transactional, and globally scalable database. So, historically—or even today—it's just actually really hard to make transactional relational databases scale. Google actually has humble background in databases as well. Lots of people are thinking about Google as this large company that only cares about large scale problems, but in the beginning, it started very small. There's this very typical story around our MySQL usage, specifically, AdWords—our ads business—had been heavily dependent on MySQL, and they got to this point that there was, like, 90 shards of MySQL instances. And they’ve been dealing with [library] sharding things, it's was causing outages, and so on.

And around this time, people decided to maybe take another look at the storage in general, and they figured out, we definitely need something transactional because, you know, we were doing a lot of money transactional related things. Consistency is really important for us because—you know, you want to be consistent, especially if it's about money. And they needed relational capabilities because there are a lot of relational problems they had. So, Spanner came as a result of these problems, but it didn't appear in a day or two. It took them, like, six years of experiments to figure out the right thing. And it's been largely in use in a lot of systems at Google. And one of the things that I really like about it is, it does a lot of work on behalf of you. It gives some, sort of, promises, and you as a user don't have to think about those problems that much. We can talk a little bit about maybe some of the higher level promises it makes.

Corey: One of the interesting things that came out of the original Spanner paper was that in the world of databases, there are the idea of CAP theorem where you have either consistency, availability, or partitioning. It's one of those good, fast, cheap; you can only ever have two. What made Spanner so groundbreaking was that, yeah, we've decided that we can actually cheat and hit all three of those things, which normally one laughs and makes fun of and then goes back to doing serious work, but this wasn't Twitter for Pets announcing this. This was Google, you folks generally tend to hire smart people who are right about these sorts of things. So, that was definitely eye-opening. I guess first off, how is that possible? And how does it do it?

Jaana: Firstly, maybe I should explain to you my mental model about the CAP theorem. Because according to Eric Brewer, this is a way that you think about these problems. You think about, like, compromises in distributed systems, and there are three things that you care at the very extreme cases. Consistency, availability, or network partitioning. You just pick two of these, right? But according to him, when you're getting closer to 100, things are changing so much. You can’t do all of those compromises.

When Spanner was launched, they come up with this idea that we are almost 100 for all of these, but not hundreds. What Eric Brewer was telling was CAP theorem is really great if you're very close to 100, but it depends on how close they are. For example, Spanner says, we have five nines of availability. Is that close enough to 100%? Or is it, like, you know… eight nines that we should consider to get to what Eric Brewer is talking about? So, I think the controversy or the discussion around whether Spanner is actually breaking the CAP theorem or not, is what does CAP theorem actually means, or what Eric Brewer was trying to achieve by talking about those extreme cases?

Corey: It also feels like this is much more of a, you must be at least this far along the maturity curve before you begin to worry about these sorts of things. I know, for example, when I build databases, what compromises do I make on a CAP theorem? I hit none of those three points because everything I build is fundamentally awful. At some point, you start caring about this sort of thing only after you've gotten to a point of your site doesn't fall over on its own every 20 minutes.

Jaana: Absolutely. And lots of people don't need actually five nines, right? Lots of large businesses, just, are on three nines. And as a random business, you probably don't need that much availability as well. Like, three nines is still at a level that cloud providers are trying to achieve. So, five nines or beyond is just really extreme.

So, I think what Eric Brewer was trying to explain in the CAP theorem mental model that, like, he was trying to give people a way to think about these extreme problems and extreme sacrifices you have to make. So, maybe in practice, it doesn't necessarily fit well because we never can achieve—or there's no real reason for us to achieve that level of extreme availability, or consistency, or network partitioning problems.

Corey: Feel free to opt out of answering this question, but Spanner was always one of those who were doing something at Google that we can't generally talk about, similar to Borg. But then one day, there was an announcement out of Google Cloud—a division, I tend to spend a fair bit of time tracking—and announcing that you were releasing something called Google Cloud Spanner. Which, ooh, Spanner that word sounds vaguely familiar. What is the relationship between Spanner as we've been talking about it, and Cloud Spanner that is something that I can go out and buy with somebody else's credit card?

Jaana: So, I think one of the differences—a lot of people keep asking me this question; is it like the case where we open-source Kubernetes, which looks like an equivalent of Borg but it's actually, like, completely different systems. What is the relationship with the Cloud product and the internal product? What Cloud Spanner does is it packs the internal solution we have and deploys it to the user/s nodes. In order for us to have isolation, we have to make sure that we are not sharing the same deployment environment. So, there are a lot of cases that Cloud Spanner is completely behaving similarly, but it's running on our cloud stack. So, networking-wise, maybe, it might be going through some different hops. But it also is trying to achieve something similar, which is dedicated networking, similar operational model, and so on. So, they are way more similar than what people think.

Corey: Of course, there is always the difference between the fact that I can buy Cloud Spanner, whereas if I want to buy actual Spanner, I probably have to acquire Google, which is currently not on the roadmap for at least the next 18 months.

Jaana: Yeah, there are cheaper options, probably. [laughs].

Corey: One or two. It's interesting because whenever I've worked with various environments where I was running ops teams or, heaven forbid, being an operations-engineer-type, the database was always the root of all problems in the sense of, okay, we're doing disaster contingency planning; we want to have multiple availability zones so we could wind up taking up rack loss, or building loss, and then expanding beyond that into going multi-region. Oh, now you have a problem because, sure, you can read from databases from anywhere, but when you start having something that has a lot of writes, and you want those writes to be consistent, now you're having to make a whole bunch of determinations that all come down to something you talk about in your bio, which I will quote from now: you spent a lot of your time, “Helping customers with architectural design trade-offs,” and everything that I've ever seen around databases—and most other things as well—are built from trade-offs. So, how does that inform how you see the world?

Jaana: A while ago, a coworker of mine said this very useful quote that I can completely relate to. He said, “Any useful system has some state.” And in any architecture, when I was working with any customer at large, I realized that there is no way that you can ignore data problems. Even in systems where data is not the intensive work, a lot of things are designed around data problems.

And databases world is far more complicated than anything else, even though it becomes such a fundamental thing. Lots of people are joking about, like, it’s just as complicated as JavaScript frameworks, but that's such an underestimation of the problem. We see outage, data loss, revenue loss on a daily basis, everywhere in the industry. Just because you don't understand, or you necessarily can’t pick the right solution, you see a very complex application layer, poor maintainability, low velocity, not being able to open to change, declined morale among developers, and everything. It’s just at the core of the system problems.

I realized that over the couple of years, I see myself recommending Martin Kleppmann’s book on data-intensive applications for people who are asking for architectural catalog of problems. I just realized that there's such a big overlap in terms of hard system problems, as well as data problems. And one of the biggest problems with databases is databases don't keep their promises. There are edge cases—the way they implement some of the features have a lot of edge cases, and they don't necessarily are transparent about what's out there. And there are so many choices; if you think about the whole spectrum of databases, we have relational databases, and then we have NoSQL, we have key-value stuff, we have different storage engines, we have different persistency options. And then you have niche databases where you have document DBs or graph DBs, whatever. Basically you have to know a lot about your problem, and a lot about databases in order to get things right at the first time.

So, I've seen that if I can go and explain people the overall trade-offs, and give them some guidance about data, it really reflects on their progress on their overall system design. Because data keeps being always the bottleneck. I'm really surprised that we're talking a lot about, like, this infinite scalability when it comes to Kubernetes, or Lambda, or whatever. But in the end of the day, your biggest bottleneck, you're going to hit that bottleneck with your data system. And the way you handle data from your modeling, from the way that you operate your database is just really impacting the whole design of the system.

Corey: For me, one of the reasons I always stayed away from databases, to be perfectly honest, is that if I screw up a web server, well, that's funny: everyone gets the point and laugh, we’ll redeploy it, and things are back to normal. If you screw up a database, in some cases you don't have a company anymore. And I am whatever the digital version of accident-prone is. So, first, this taught me to do very good backups and, two, it taught me to hire professionals to wind up handling anything with persistence, by and large, which has led to some very interesting beliefs and structures in my world around, for example, DNS as a database. What do you find that—from a customer perspective—the biggest misconceptions are that require architectural trade-offs?

Jaana: I think the biggest problem is—especially with Cloud—they believe that, like, resources are infinite. And it's easy to auto-scale. Some of the customers are coming from this really dynamic workload type of environments, and they believe that over on Cloud, we have no capacity issues, plus we can just auto-scale and we can dynamically resize our pool. And I think most of our compute products are sort of making this a bigger issue because we made auto-scaling too easy, without necessarily considering what it means to the overall limitations of the design. So, I see a lot of people coming from that and realizing that that's not the case. And then they start to see everything more holistically, maybe they realize that they need to start about understanding the limitations at the database layer.

And that's also a very complicated problem because the things that they are looking at, like latency, and throughput—and these are very superficial numbers to take a look at—they still have to realize they have to still identify large specific operations and how they're going to work against their database, particular loads, and so on. They're kind of getting lost because the spectrum is really high in terms of what to measure, and the existence studies around standard benchmarking or standardized stuff just doesn't really help their particular use cases. So, they have to do a lot of prototyping, they have to evaluate a lot of things before they are somewhat happy about their overall initial design. And this is if you're building things from scratch. If you're migrating over, it's just getting much harder.

Corey: In some ways, it feels like working at Google puts you in a position where something that the rest of us have to struggle with, but Google doesn't. Specifically, whenever I build something, it probably doesn't need to scale until suddenly, it's absolutely going to have to scale because it turns out, I built something that resonates with people. That doesn't seem like it's a Google problem because if you slap the word Google on a product that gets launched, on day one you'll have millions of people using it, so anything you build has to scale. Therefore, it removes the entire debate of do we build this right or do we just slap something up there and go back and refactor it later, I would think. Am I right or am I wrong?

Jaana: It's true that we design for scale because we expect, let's say, this number of millions of users on the day first. This is mainly true for large products that we are going to release. So, Google is a very large company, we have, like, all these different systems that doesn't do any consumer market things, so there's actually a variety of different scales. But for consumer market problems, yes, it's true. We have this large XX million expectation on the first day, that's why we specifically pick this type of trade-offs, pick this type of solutions, and everything is more in an [unintelligible] way, but there's a large spectrum of other problems inside Google that doesn't necessarily need that type of scale. And internally, for example, we have a lot of database solutions, a lot of general storage solutions, and there's a huge—also a decision chart internally that—which one you have to pick. And it really depends on the type of problem you have. So, even at Google, it's true that—product teams especially—are more biased towards very large scale. There are a lot of small scale problems, too.

Corey: In what you might be forgiven for mistaking for a blast from the past, today, I want to talk about New Relic. They seem to be a relatively legacy monitoring company, and I would have agreed with that assessment up until relatively recently, but they did something a little out there: they reworked everything. They went open source, they made it so you can monitor your whole stack in one place. And most notably from my perspective, they simplified their pricing into something that is much more affordable for almost everyone. There's even a free tier with one user and a hundred gigs per month, totally free.

Check it out at newrelic.com.

Corey: In many cases, what's right for Google is absolutely not going to be right for other people, but at a certain point of scale, certain things change. And if you take a look at all of the big tech companies out there, they've all built their own programming languages. For example, Microsoft has a whole bunch of them: .NET, ASP.NET, C# et cetera; Facebook came out with Hack, their custom PHP thing; Apple came out with Swift; Amazon came out with CloudFormation; and Google came out with Go, which is something you were deeply involved in before a relatively recent shift over to work on Spanner. What did you do for Go, and what made you decide it was time to go stare at databases instead?

Jaana: I was working on Go after the 1.0 release. So, I started, I think, around 2012. The funny story is, when Go was released, I was not working at Google, and I was working in telecoms. We were working on message parsing systems. These are highly concurrent systems, so we were just basically looking after what else is coming—especially in the languages and runtime space—to make our jobs much easier.

And I was looking at Go around that time. And no, I didn't truly understand the type system or anything, and I felt like this could be something that I can consider in the long term, maybe, but I don't really feel like this language is really the best choice for my personal things that I like in a language and so on. So, I just really didn't do much work. But after I joined, I was—by chance—sitting right next to the compiler team in the Zurich office in Switzerland. And I was just, kind of like, you know, in the conversations because they were all language enthusiastics around me.

And I started just taking a look at things, and I at that time, I was working on Google Drive, so we had about bunch of migration projects with lots of networking and an I/O, and I started just kind of writing small things, and trying out things, and as a result of that, I started to publish some of my open-source tools, and so on, and realize that the community is just really amazing in Go. And I realized in a couple of years that maybe I should do something as a part of my full-time job. And I joined to the Go team to work on, generally, our external API, client libraries, some tooling around them, gRPC—gRPC was just coming around at that time, so we did a lot of work on gRPC as well. We had some sort of project to unify our stubby internal API stuff with gRPC. So, I did some Go specific things. I did all these reviews for all these cloud products who wanted to support Go. I actually initiated one of the earlier projects for our cloud to support Go as a first-class language.

So, back then nobody was interested in Go. This was back in 2002. Go was still kind of like a smaller language and a community. So, I initiated a lot of, like you know, the bunch of small things, and they got funded, luckily. Now there are teams actually working on those things that I initiated as a 20 percent, finally. That's how I started my journey with the Go team. I was necessarily just handling more of the cloud support related things, and then I switched, some sort of like, my interest. There was a project that was trying to make the Go runtime working on Android and iOS. I briefly worked on it, also contributed some of the tooling.

And recently, before my switch, I was really interested in instrumentation, and performance, and debugging tools, and that sort of—and there was a small subset in the Go team that was handling a lot of performance-related stuff. I'm not sure if you're familiar with Go has pprof support, we have an execution tracer. We were thinking about maybe establishing some sort of primitives for distributed tracing, we were thinking about some metrics APIs now. There were a bunch of small things, so we were trying to see what is the overall larger picture; what else we can do.

And so I worked on that team for a while and then left that team to work on instrumentation at Google at large because I realized that a lot of things that I was trying to do in Go was actually larger than just Go problems. So, maybe I should just go and work on the instrumentation team to get some more exposure to that problems, and then I can go back to Go and apply them. But then I ended up being on the Spanner team because, you know, instrumentation team at Google was sponsored by the storage system, so I was really involved in a lot of storage problems as well as networking problems. That was a really gradual switch from Go to other things, but it's funny, sometimes you have to do what you can do, and what's important, and the most priority thing.

And I like to be able to switch back and forth at Google. We have this very loose way of collaboration, and so, I mean, we don't have to necessarily go through interviews or anything, you can just switch projects and contribute, and sometimes overlapping different skills are very useful for the project that you are going in. Like, you're bringing a completely different background. On the Spanner team, for example, I have some experience before coming to Google. I actually left my previous company because of the database problems that we have. So, I had a lot of experience migrating us to different systems, designing, and evaluating databases. That was my previous job, and at some point—I don't want to name a database—but we were losing data, and it turned out to be a very fundamental issue at a database I don't want to name, but we spent weeks, and I spent—

Corey: Yeah, I will fight you if the answer to that database that you don't want to name is DNS.

Jaana: [laughs]. It is not DNS, thankfully. DNS is a better database than that database, I'm pretty sure.

Corey: [laughs]. The funny thing is I honestly don't know the answer, but I can think of at least five that fit that profile.

Jaana: Yeah, yeah. Probably you can tell, if you have maybe a shortlist of two, you will be able to tell which one it is. I don't want to tell it; I don't think that I want to be in that fight. [laughs]. The thing is, you know, I just gradually ended up leaving my previous company just because I was so tired of the storage problems. And I joined Google because they gave me an offer, first, and the second thing is I can actually learn about storage, maybe, at this company because they seem to know what they were doing.

Corey: Yeah, I mean, again, there's a lot of criticisms that you can lobby at a whole bunch of different companies. Google is right in that list too, and I do have a counted list that I don't have the time on this episode to read and blame you personally for all of them, but one thing that Google has always gotten very right has been fantastic technical solutions to incredibly hard problems at scale. It's easy to bag on companies, but there's a lot of hard work that goes into making these global world-scanning systems, and I think that that's something that often gets forgotten. I mean, there was a time where Google was lightyears ahead of absolutely everyone else. And now it seems that, oh, well, what do they do? They built this world-spanning thing that's super fast no matter where you are on the planet. Yeah, here's my five lines, of YAML and 20 cents a month, and I could do something like that, too. We all stand on the shoulders of giants, and it's easy to forget that Google's 25 years old; they've built a lot of these things that have been transformative to the entire industry. But now it's, well what have they done for me lately?

Jaana: I agree with this, but I feel like we still need to do a better job in terms of understanding what it means to scale to small, right? We have these aspirational experiments, maybe, we're running on behalf of other people because we had some unique problems in the past and we built these systems that works for us. Some of those experiments could be aspirational for other things, rather than completely translating—the interesting thing, before joining to the spanner team, I had this concern: is this like—we don't want to be this new, shiny new thing that nobody cares about that only works at Google scale. A lot of people on the team have been telling me they had the same concern. They didn't want to actually release Cloud Spanner.

The more they were talking about customers, largely about database systems, they were asking them to release Cloud Spanner because they wanted a solution. They didn't want to deal with consistency issues the way they used to do. Like, in traditional databases, especially relational ones, it's so hard to scale writes, for example. This is such a fundamental problem, right? You can’t scale your writes, but you want to launch this large game and you want to focus on your game, you don't want to deal with your database layer.

So, just because they've been so hard on them, on the team, if you have an end solution to this, you should share as a product. And then that's how Cloud Spanner actually comes around. I was very impressed by the fact that Spanner team is thinking too much about the customers. This is something that I have to tell. At Google—maybe this is the only team that I've been working at—and at every meeting, I think, we keep talking about customers and customer issues all the time. And like, that is such an incredible thing, how much we actually care about customers in necessarily prioritizing what they want.

Corey: So, before you run Spanner, you worked on Go, and you mentioned that you did telco work before that. What is the story? Generally speaking, Google's staff engineers do not spring fully formed from the forehead of some god. What was the journey that got you to where you are at your career?

Jaana: I started actually—before telecoms, I started at a small company based in London—it’s headquartered in London called Multimap. They were online mapping company that, sort of like, was Google Maps before Google Maps was Google Maps. And they’d been acquired by Microsoft. So, my first experience in life was actually working at a company that was acquired by Microsoft. And it was very interesting for me because I've seen two phases of the same thing, right?

Like, you have this small company that cares a lot about their business problems in a smaller scale, as well as there's this giant coming in and trying to see what kind of differentiated value that acquisition can bring. And they have a completely different scale of systems, and different trade-offs, and so on. So, that was like a really eye-opening thing. And just going through a lot of discussions about how things work, or why things don't work in the new scale, and so on, really helped me a lot. After that, I started to work mainly in telecoms.

And, as I said, I was working on this highly concurrent message parsing rule engine type of systems. They're actually very boring, but they really helped me a lot. Just prototype, and evaluate, and understand the overall system problems. What are some of the limitations that I should take a look? What are some of the failure modes that I should care? And it was in a very, actually, fast-paced environment that I was able to prototype, build, change significant things, push things to production, see some outage, iterate.

So, I've seen a lot of interesting things, being able to touch a lot of different things. We were also trying to modernize parts of our stack, so there was a lot of work in terms of going and discovering new things, and, like, stress testing a lot of new tools, or a lot of new libraries, or language runtimes, and so on. As I said, I was looking at Go as a part of that work. So, that really helped me to see everything in a broader sense. And I really like the fact that I've worked so much more outside of Google than years that I've spent at Google because I've seen problems of different scale, and I've seen different levels of flexibility when it comes to introducing new technology, and I've also seen a lot of different types of organization with different types of problems in terms of scale and the organizational issues.

That really is helping me right now to help our customers because, I mean, I can relate to a lot of things large majority of the customers are going through. After telecoms, by the way, before coming to Google, I've worked for two small companies. They were trying to bootstrap; they were just at the initial design phase for lots of things, but they also had some sort of established business going on. So, I had the chance to see the both sides of the things in more of a playground area where we can go and try out new stuff, as well as a lot of established problems, and legacy issues, and scale problems, and large organizational problems that we have to tackle.

Corey: It's always interesting to me to see how people come to where they are from various different places. So, my last question before we wind up calling it a show is what advice would you give to people who are looking to go from where they are in their careers to a job that looks a little bit more like yours?

Jaana: My biggest concern when I was earlier in my career was, I was like thinking, I am wasting too much time. I was feeling like I am wasting too much time all the time; I'm working on these problems that actually doesn't make any sense in the larger scheme of the things, and so on, and I was very frustrated. I was feeling very demotivated. And if I was able to go back to that person and tell her something, I would say that, “Just don't worry,” because at some point, all of that little experiences really just gets you to a point that you can overlap different experiences, and bring some different perspective.

What I realized is, over time that—especially on the Go team; there were a lot of very senior engineers—but in some cases that I realized that my particular background in all of this weird stuff gave me a huge, very niche thing, but at the same time, a very general perspective that I can apply anywhere on some of the topics. And in some cases, I was the only one in the room that actually have any experience in all across, and was able to say something when we're thinking about design, or some of the goals, or some of the trade-offs. So, I would say people shouldn't worry too much, and especially if they think that they are learning, there is nothing tedious about learning, trying out.

Going after new shiny thing—most people think going after the new shiny solutions is the only way to kind of have any job security. I also disagree with that. Just work on the tedious things. Most of the problems in the world are very tedious and become the person who can recognize and identify the tedious parts, and the common patterns of problems. Tinker through them, and that's going to contribute to your career or your growth more than just going after every shiny new thing.

Corey: Thank you so much for taking the time to speak with me today about basically all of this. If people want to hear more about what you have to say, where can they find you?

Jaana: I have a Twitter account. I usually am trying to be very public about what I am working on, which helps me to hear other voices. So, you can find me on Twitter. And I try to write a lot. Nowadays, I'm not writing that much, but I have two blogs. So, I'll give you the links. You can probably link them.

Corey: Excellent and I will put links to those in the show notes. Jaana Dogan, staff engineer at Google. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts. Whereas if you hated this podcast, please leave a five-star review on Apple Podcasts, and then leave a comment telling me what I got wrong, written in Go.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Jon Myer
A Partner Solutions Architect for Cloud Management Tools at AWS. Jon Myer is an evangelist for all things AWS and passionate about educating, teaching, and connecting with others about new or existing services AWS releases.

Links Referenced:

  • The AWS Blogger: https://www.theawsblogger.com/
  • Twitter: https://twitter.com/_jonmyer
  • LinkedIn: https://www.linkedin.com/in/jon-myer/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

It's, at least as of this recording, morning on the West coast, which means there's no better time than to inflict a homework assignment upon you in the form of a 42 page ebook from StackRox. Learn about the dancing flames of EKS cluster security, evade the toxic dumpster of the standard controls, and tame the wild beast of best practices for minimizing the risk around cluster workloads. Become renowned for your feats of daring, as you learn the specific requirements for securing an EKS cluster and its associated infrastructure. To learn more, visit snark.cloud/stackrox. That's snark.cloud/stackrox.

Corey: This episode is sponsored in part by our good friends over at ChaosSearch, which is a fully managed log analytics platform that leverages your S3 buckets as a data store with no further data movement required. If you're looking to either process multiple terabytes into petabyte scale of data a day or a few hundred gigabytes, this is still economical and worth looking into. You don't have to manage Elasticsearch yourself. If your ELK stack is falling over, take a look at using ChaosSearch for Log Analytics. Now, if you do a direct cost comparison, you're going to save 70 to 80 percent on the infrastructure costs, which does not include the actual expense of paying infrastructure people to mess around with running Elasticsearch themselves. You can take it from me or you can take it from many of their happy customers, but visit chaossearch.io today to learn more.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Jon Myer, a partner solutions architect for Cloud Management Tools at AWS. Jon, welcome to the show.

Jon: Thanks, Corey, for having me. I really appreciate it.

Corey: So, you're a partner solutions architect for Cloud Management Tools. I feel like there's someone at AWS who gets bonused per syllable that they put on business cards some weeks. Can you distill down what exactly that is?

Jon: Yeah, you think they get bonus points, try writing it out. That's why I shortened it to PSA-CMT.

Corey: Exactly. We do love our acronyms.

Jon: [laughs]. Yeah, exactly. All right. So, what a PSA for Cloud Management Tools does is—you know, partners are my customers. I work directly with partners to really help them out. So, if you think of some of the top partners out there for cloud management, CloudCheckr, CloudWire, Cloudability, Turbo, Turbonomics, those are the ones that I directly deal with. And I'll work with them for doing webinars; blog posts on their applications; maybe they're bringing out a new UI, I'll test it out.

And I'll take some of the AWS products or services that might be up and coming, and I'll work directly with them to try to get it integrated, or try to figure out some caveats for it, so that they can tell it out and make sure that their application works with it, so when we release new services, or they get launched, they're already kind of jumpstart on in the help their customers out because we don't want customers to be in a tough situation when a new service is released, and the partners can't support them.

Corey: So, the hard part that I’ve found has been anytime you talk to someone who does solutions architecture in anything even tangentially approaching AWS, let alone someone who actually works there is, how do you figure out where to start and stop? And that is a sincere question, by the way, because we have long since passed a point where I could talk to you right now—as an AWS employee, about various AWS services that don't exist. And we're at a point now where there's no way you'd be able to, on the fly, disambiguate between which ones exist versus which ones I'm making up to be funny. I don't generally tend to do that because, ‘haha I made someone look foolish’ is not generally my brand, unless that someone is a multinational company. But the problem is that it's so hard to be a specialist across AWS. Whenever someone says I'm an expert with AWS, my immediate response is you're lying because no one knows in all. Where do you draw the lines?

Jon: That's actually a good question. We come across that a lot. The great part is that we have a strong, vast majority team behind us. And customers and partners know that and understand it. So, if I'm in a deep conversation with my partner, and they want to talk about AI/ML, which I'm not an expert in: I understand it in the basics; I'll tell them that I have to bring in another specialist or somebody who can work with them. He'll join the call, that person will not only educate the partner, but I'll start to educate myself on bits and pieces of it. We have, like, what 190 some services and don't quote me on that one because I'm sure that there are plus or minus a couple—

Corey: You're forgetting a few depending upon how you slice it. I was slicing apart Boto and came up with I think around 220, but some of those are questionable. Yeah, that's the problem. There's not even a clear answer. We could have a whole debate about the number of services.

Jon: Yeah. So, that usually gets a little tough for us to-back then, when we had just a little majority of services, it was all right to be, “Yes, I knew AWS,” and somebody’d be like, “All right, hey, they know at all.” But with so many services, I try to put my focus on the ones that my Cloud Management Tools partners are really focusing on and using, and if they need help with another service, I have no problem reaching back to my team and saying, “Hey, listen, I need some help. Who's available?” And there's always 10 or 20 people like, “Hey, I can help you. I can jump in there.” So, it really makes you feel comfortable that you have the full support of your team behind you.

Corey: That's got to be fun. Now that we have a staff on our side, I'm starting to learn how some of these challenges begin manifesting because the company, at least on my side, has grown to a point where I can't hold it all in my head anymore. There are consulting engagements in flight right now on my side that I have never interacted with, which feels bizarre to me, and it's like I'm careening out of control and losing it. That's something, for better or worse, at a big company you've already made your peace with as soon as you effectively walk in the door. How long have you been with AWS?

Jon: I've actually been with them about 10 months now. That's probably about, what, 20 years in AWS terms?

Corey: Especially given the last two months of pandemic, where at that point at, well, we're absolutely at this point, “Wow, March lasted 10 years.”

Jon: Yeah, so the 10 months that I've been with those, I felt myself grown, I've had a lot of fun learning, and onboarding, and working with various people doing all kinds of cool and crazy things, and this year, I have so much more planned, and it feels like I've been here probably I'm going to say at least five years already. I feel like a veteran on this, and I'm just really engrossed in AWS.

Corey: One of the funny parts is that Amazon, by and large, is generally not a remote-first company. For an awful lot of roles, you can work from anywhere you want, as long as it's Seattle. And you on the other hand are out near Allentown, Pennsylvania, a city near and dear to my heart, given that I was born there. I kid, I kid, I was not actually born: I sprung fully formed from the forehead of some God. But you have been doing the remote thing at AWS for a while. Tell me about that.

Jon: Yes, so that's actually one of the benefits. I've been working remote, or from home, probably for the last five years. So, I've kind of had my routine, and when this role came up, and I'm still, kind of, based out of New York, I traveled to New York occasionally when they need me, and in fact, I've traveled quite a bit, but through my own options, and I had that ability to pick and choose some of the things that I'm able to do and support my partners, which is great, but being able to work, I want to say it's a small town compared to, or a really small town compared to Seattle, but it's nice because we're still connected to our teams. We are using various tools to do our communications. We use Chime for our video chat, and it still feels like I'm connected to the team because we're constantly interacting in one way or another. I love working from home, and I like the ability that I can travel occasionally when I need to; go to New York to work with some of the colleagues. It's really made it flexible.

Corey: One thing I really want to emphasize as someone who's been working from home myself for the last three and a half years, as of the time of this recording is that this is not normal. I have had so many different conversations with folks where they feel like they're not being productive; it turns out that they actually suck at working from home; they need to be in an office. No, these are not normal times. My productivity has taken a dive, not because I don't know how to work from home, but because we are now staring down a pandemic, and for most of us, it's sort of hard to compartmentalize that off from being the most efficient corporate drone we possibly can. So, talk to me a little bit about, first, what you've seen change from a remote culture, and then let's talk a little bit about workflows.

Jon: That's actually a really good point. Since I've been working from home, this is definitely not the norm. We are not dealing with a normal situation, I would have a normal routine. Get up, get the family ready, kids are off to school, I’m engrossed into work. You know, lunchtime, work out, afterwards meetings. Really cutting off around 5:00, 5:30 of my day, and that's great, and I have flexibility throughout the day if I need to run somewhere, and do something, and handle something.

But now you're not allowed to go anywhere, and you can't do anything, and everybody is stuck in a routine where you're only doing things around your house, and occasionally going out and doing various trips for essential items. For me, the first two weeks of dealing with not only the family being home, but colleagues, interacting with them, and they're like not sure what to do, and how to handle this office-to-home transition, they didn't have the setup that I have to get things done efficiently. They're adapting, they're learning new ways and tools to do it, and then in the two weeks after that people are like, “All right, I need to chat with people,” So, you have a lot more chat conversations, video chat, people want to jump on it. I've offered some virtual coffee time for people who just want to download, and need some of that social conversation. Now, I've been working from home, so I don't always need that conversation. I've kind of adapted to that, but you're helping colleagues out who aren't used to that portion of it. And that's really where it comes down to, is that the office workers who are working from home now need help in that transition. And I think working from home might be, kind of, a new norm as people get into that routine and figure out new and inventive ways to do it.

Corey: That's something I found that has been—people are both focusing on in a way they never did before, but they're also in some cases focusing on the wrong part of the story. Personally, on a selfish level, I'm very eager to see what happens when we start going back into offices, now that people know exactly how little it costs to build out a really nice home office from an equipment perspective. It's let me get this straight: you pay me [clears throat] thousand dollars a year, but I'm not allowed to splurge on the expensive $800 standing desk that I spend more time with than I do my own family? It becomes one of those, “Wait a minute…” I think you're going to start seeing a backlash—I hope—against this open plan office mentality where everyone works in these giant tables, cheek to jowl with a bunch of other developers. I'm optimistic. Maybe I'm wrong on that, but that one feels like an easy win for me.

Jon: I think you're right. I think more the people transition from that office to work from home, where working from home will be the new norm, and when you go to the office is really when everything's going to be quiet, or you're out of office will be on, and maybe saying, “Hey, listen, I'm in the office, and I'll get to you when I get a chance,” But when you're at home, you'll do more of that communication. I think it's going to be a total reverse effect.

Corey: Talk to me a little bit about services that AWS offers that you're finding to be particularly helpful now that everyone is, surprise, working from home. And I do want to say, incidentally, that this is not a valid work-from-home test. Surprise, with basically no notice, everyone is now working from home when they didn't expect to be, during a pandemic, is not an adequate test of how good are you at working from home? Is working from home for you? Is your company structured to work from home? The answer to all of that is that this is awful. So, bearing that in mind, what have you found that AWS offers that makes this a little bit more manageable from a day-to-day perspective?

Jon: I think you made a valid point—before I jump into any services they offer—I think of any companies takes this pandemic and working from home, a way to gauge or measure on how well people will work from home as you start to do that as a norm, it would be a total reverse. It's not a valid analysis of the environment. There's no way that you can say, “Hey, listen, during this pandemic, we had 100 people work from home, and productivity was down.” No, there's no way that you can take this—you got to take all the data from people being able to work from home, and not compare it. You have to take into an equation on this pandemic, and that they're working for home. Efficiency—people are unsure right now, so efficiency might be down or up for certain people as they start to learn new ways and to do things. And it all depends on the person itself.

Corey: One thing that I am loath to say when there's any chance that he can hear me, but there's no way he has the spare time to listen on a podcast, is Sid Rao is the general manager of Amazon Chime, and one of my favorite things to do on Twitter is, basically, complain about how awful Chime is, mostly because I find that inspires Sid to do some of his best work. Here, in reality, Chime is really, really good. I've used an awful lot of different messaging services over the years; they all have their faults, and those faults become extremely apparent because a call drops, or suddenly there's a sync issue that doesn't work right. You're always using these apps, and it's incredibly frustrating and noticeable. It's either terrible or it's invisible. Chime has really gotten good. Please don't tell him that because I have a reputation to uphold. That's why it's just between you, me, and several thousand listeners.

Jon: I'll have to accidentally send him the recording of this one.

Corey: Uh-oh.

Jon: Oh, so you mentioned Chime. I like Chime, and in the last two weeks, I've learned some cool features about Chime that I didn't even know existed. One of the other ones that you found out is that I can invite external users to communicate with AWS folks for Chime. Now, please don't share that out globally because everybody will be emailing and trying to reach out to me on Chime, but it's a really good way for me to communicate on one platform with another user that is not part of AWS and to keep in contact, and those communications and to share something really quick. It's better than dropping that email and hopefully, they read it. Chime does that one.

The other feature that I just found out two weeks ago, is that Chime has an event mode, almost like you have a webinar series going on. And guess what? I'm going to record this, I want to make this a webinar right here and now. I can actually turn my group or my meeting into an actual webinar or event, and you can do things like mute to participants, they can’t unmute it, they can ask questions, you can record it, you can share the link out with multiple people. One of the things I'd really like to see is a registration page, I can send it out. I mean, these are features that are just completely awesome about Chime, and the chat portion, groups—something that I know, Corey, you just found out is that you can do a group chat with Chime that you weren't aware of. So…

Corey: Yeah, I've been using Chime for years now. All I needed was an Amazonians email address, and I could message them on Chime, which scared the living crap out of a number of people, which is fine. But it’s how I tend to communicate; it works super well for me and the way that I process things. If we start lobbing emails back and forth, we're lucky we can go three rounds, where it dies in my inbox. I mean, I'm a big believer, these days, in inbox zero [BLEEP] given.

So, if you email me, you're absolutely not going to get a response super well. In fact, when I took on a business partner, having me pass email communication to him for client stuff was basically job zero, we dove directly into that because getting me out of the critical path was great. But messaging people on Chime for asynchronous questions is awesome. I have other problems given my nature of how I tend to operate with AWS and various things that I do, where people think, oh, wait, are you coming at this from a perspective of a journalist, I have never been a journalist. They're good at things and I'm not. Are you coming at this from a customer perspective? Am I coming at this as an analyst? Am I coming at this from just trying to be funny? Am I coming at this to help reprioritize that Chime needs a block feature? Et cetera, et cetera.

Jon: I completely agree. I actually—I’ve been using Chime before I was with AWS, and that's how I communicated with my APN partner, my PDM, and how I would message that person and get immediate responses. Now, just like you on sending emails to not only that person—or I might have sent some emails to you as well, that go on answered—no comment there. And it's easier to get a comment or a response back quickly Chime because you quickly return back a one-liner and it's done. You don't have to click the send button as quick—and your email read through a whole bunch nowadays is going to be clearing your chat inbox versus your regular inbox.

Corey: Yeah, I mean, again, I could think of a lot of feature enhancements. I mean, Slack just got the ability to group channels, which is awesome. I mean, I'm in a company Slack myself with roughly 20 people once you add in our contractors, not to mention all of the shared channels we have with folks, and being able to categorize those by project or whatnot is super handy. With time right now, it's just a last conversation you've had, so that becomes a bit sticky. But by and large, it's a decent product. What else have you folks got? What else have you released that is making the pandemic remote work story bearable from folks who are discovering these services, and what services should people look into if they haven't figured this out yet?

Jon: Yeah, the other service I really like to touch on is Amazon WorkSpaces, which is your Desktop-as-a-Service solution. It's been around for years. In fact, working for another company, I kind of pioneered it internally. Nobody had the expertise. It was new and exciting and deployed it for a large enterprise-level company, got it out there and I was just learning everything, the ins and outs about it. Now think about this, this is, like, four or five years ago. The workspace client that connects to your Amazon WorkSpaces works on any device. I mean, I had it deployed out onto an iPad, and logged into it and actually did some configurations while I was about to take a cruise, a family vacation, and I was able to log in there and get things configured within 10 minutes and resolve issues.

Something that Amazon is doing now, during this, is they're simplifying the delivery of it and reducing the cost of Amazon WorkSpaces during this. They have a couple of free-tier offering or standard-tier offering that allows you to deploy and quickly deploy out for your environment or remote users, and what you're doing is you're eliminating the need to overbuy your desktops, your laptop resources by providing an on-demand access to cloud desktops. And in fact, it's something I'm going to work on is I'm going to work on a video recording and get it posted out there, I do have a webinar in I think about two weeks around AWS WorkSpaces, and how you can set it up, and get it quickly deployed for not only customers but remote workforce.

Corey: And because we wind up having a bit of a delay, that webinar will be out by the time that we wind up publishing this episode, but I will throw a link to it in the show notes, which is kind of weird because you can go watch that webinar right now if you're listening to this, but at the time, we have no idea how it turns out. Personally, I'm hoping there's a hilarious pratfall in which Jon winds up tripping over himself and falling backwards out of his chair. But hope springs eternal; only the future will know.

Jon: I'm going to have to add that to my recording and say, “Hey, Corey, that one's for you.”

Corey: So, something else I want to cover with you is that you recently—as it turns out, I only found this out while researching this show—have launched a personal blog. And in a hilarious coincidence, your first post was on how you did it, and you use the exact same WordPress hosting architecture that, after you had already launched and gotten this up and running, I excoriated on Twitter as being hilariously over complicated. And I wanted to talk to you about that, first, to talk about your blog, and secondly, to add nuance to my observations about WordPress architecture that don't eloquently fit into 280 characters.

Jon: So, let's go on to my blog and why I did it, and let's talk about the architecture. Why did I do it? Well, one of my ultimate goals is to become an evangelist, so I enjoy reaching out and connecting with the audience. I like to work with how-to videos, basically anything AWS, I really enjoy it. I really enjoy sharing what I've learned and very small, quick things. I started writing some blog posts. I wasn't really huge on writing them, but all of a sudden, I found so much in my head that I just wanted to dump down on into a blog and share out with people, and I just thought it was really cool.

So, I was hashing out the name, a buddy of mine, him and I were going back and forth, and he's like—I gave him a couple characters and he gave me back, and then he came up with this and I'm like, “Oh, that domain’s available.” So, I bought the domain right away. And I was like, “Okay, I guess I should go get some approval before I start writing this since it has the word AWS in the domain.” And I'll tell you what, I got the full support of AWS, for us to really do social activities, and share this. Obviously, we have to have our little disclaimer, “This is all me, this is my beliefs.”

Corey: Speaking of someone who owns lastweekinAWS.com who has never been affiliated professionally with AWS, that's an interesting point that you bring up right there. In fact, right before I launched this newsletter, three and a half years ago now, I bought lastweekinthecloud.com with the assumption that if I ever wind up hearing something from AWS legal, I can spend 20 minutes, pivot it to something else, and then talk about AWS as competitors. There we go; that winds up being a nice hedge. Now, I want to say, to be very clear, in three and a half years, despite some periodically incendiary takes on what AWS has done, I would argue when they deserve it, I have never heard a negative word from AWS about this. And that's to their credit.

Jon: Yeah, I love the support. In fact, I brought it to my manager, who's a great leader, and I'm like, “Oh, hey, I did this.” And he's like, “That's awesome.” He goes, “Just make sure you have the approvals.” I sent them in. And he's like, “Thanks, you're good to go.” So, I just like it. I throw in as much as I can on there. In fact, I'm due for another post, which I am working on. I'm going to get that up there. I'm working on a UNDERGROUND DeepRacer—or a Deep UNDERGROUND RaceTrack, as per Corey. So, that will be shared out on there; you'll see that in an up and coming—or by the time this gets out, you already follow me and see that post.

Corey: In what you might be forgiven for mistaking for a blast from the past, today, I want to talk about New Relic. They seem to be a relatively legacy monitoring company, and I would have agreed with that assessment up until relatively recently, but they did something a little out there: they reworked everything. They went open source, they made it so you can monitor your whole stack in one place. And most notably from my perspective, they simplified their pricing into something that is much more affordable for almost everyone. There's even a free tier with one user and a hundred gigs per month, totally free.

Check it out at newrelic.com.

Corey: One thing that amuses me very much about your website—and again, this is not your fault directly at all. It is the entire Amazon social media policy because you mentioned this, so I'm bringing it up, which makes it fair game. At the footer of your website says, “This website represents my own viewpoint and not of my employer, Amazon Web Services.” The problem, from my perspective, given that I'm cynical and more than a bit of a has always been that I don't think anyone is ever going to read someone's blog or social media postings and think that they are speaking on behalf of their employer. Instead, the problem is, “Wow, now we know that these are your thoughts, and yet we continue to employ you, despite them.”

And I’m being a little sarcastic here; there is nothing objectionable that I've ever seen come out of you. But—well, I take that back. We have had some conversations about a few AWS services which I feel you are deeply and profoundly wrong about. But I've never seen anything controversial or problematic come from you. And frankly, any AWS employee for the most part.

But the idea of I don't speak for my employer, I understand why that statement has to be there. I absolutely understand the risk mitigation, but on some level, I can't help roll my eyes whenever I see it. And I strongly suspect that failing is on my part because, again, I'm a small company that for the longest time was just a corporate wrapper around my crappy excuse for a personality. If I say, “Oh, I don't speak on behalf of my employer.” “Yeah, sure you don't. It's the question of which pocket is your snark coming out of this week?” It's a different story. I understand having it there. But I've got to say it is through no fault of your own, or frankly, anyone at AWS’s: in their shoes, I would almost certainly do the exact same thing. But I chuckle every time I see that disclaimer.

Jon: I actually like the disclaimer, only because there will be that one person that says, “Oh, look what Jon from AWS might have posted.” Or something like that. And you're right, I don't usually post anything negative. I know there's flaws in almost anything we do or another service outside of AWS in your daily lives, and I've come to, “Okay, that service doesn't work, but how can I reach out there and try to make that service better?” So, I take that feedback, and I add that to it.

So, if I find something negative, I don't like to just throw it out there, and point it out to everybody and whatever, I'll try to keep it internal, and try to make it better. There might be a roadmap to actually doing that already, so what’s the sense of me actually making a negative? I tried to think of positive, or what I would like to see. I know a lot of people like some of the negative portions, and everything like that. Maybe that's down the road, but right now I like sharing the positive. And the disclaimer just helps it so that one person doesn't think that Jon is commenting on this from an AWS perspective. I'm just commenting on, personally. That I work for AWS, please don't confuse the two.

Corey: I will say it's always amusing to me to see various social media profiles like AWS, VP of corporate communications, I don't speak for my employer. You want to bet? At some point it becomes parody. Credit we're due. Andy Jassy, CEO of AWS does not have, “I don't speak for AWS,” in his social media profile, which is good because, frankly, it is impossible for him not to, given who he is and what he does.

But there are occasionally humorous moments, usually on the marketing side, where I'll see things like that. What I want to talk about, too, is the architecture that you picked for your WordPress blog. You followed the best practices highly available architectural pattern guide, which looks hilarious on some level. It's easily $800 a month in services if you build out the full thing. And people love to dunk on it, which basically when I say people, I mean me. But it's a great glimpse into what it takes to scale these various web applications out in a controlled, highly available way.

It's not because WordPress is what people should be running themselves in this context, but rather because here is an architectural reference guide that you can use, where WordPress is basically the classical three-tiered application. Here's how that can work if you want to put that in a cloud-friendly architectural environment. Now, I have angry and negative reactions to WordPress because one of my jobs a while back, was at Media Temple before they were acquired by GoDaddy, and we ran a WordPress grid that had something like 100,000 customers on it, so I have scars that run deep for WordPress. That said, lastweekinaws.com runs on WordPress, but I pay WP Engine to handle it for me, which means hilariously, it's now hosted on GCP. But it's fun. It's great because I have the good sense to pay someone else to run the blog for me, and I haven't built Frankenstein's monster in order to serve out traffic.

Jon: So, when I first started off, I was actually starting it in another provider, a free one, whatever it was to actually get there. It was Squarespace when I was actually starting off, and I was like—

Corey: The first iteration of my corporate website was also on Squarespace because I could easily spend 40 hours a month playing around with the website that generates zero business.

Jon: Exactly. I was playing around with it, and I was like, “Okay, you know what? I work for AWS? Why am I not deploying this in AWS?” And I started searching, I was going through, and I’ve used Elastic Beanstalk before and I was like, “You know what? That's perfect. Let's do that.” And I was going through it, and I'm working on a blog post for that one, an iteration of the current one that I've actually put out there on my website on, I started off with Elastic Beanstalk.

And then I came across an article, and it was CloudFormation stack on GitHub, and it was like deploying out a WordPress site, and highly available, and I loved the architecture. I actually loved the deployment of it. Now, do I care if it's WordPress or not? Well, I actually can use WordPress; I am not a website developer, or designer, or image person by any means. I just wanted something that I could get up, and get out there, and focus on the content itself, but I also wanted it hosted on AWS, and I felt that me as an AWS person, that I could talk more about it, and share some of my architecture with people, and some of the use cases on why I did it. I did obviously lower down the instance sizes on there because I didn't think I was going to be needing a major X large instance size right now, but it is highly available.

I'm using EFS, which is awesome. I've used it in the past, but in this context, all my data is on a shared storage, and no matter what happens, my WordPress site is deployed on them, it is configurable, I can put it on one, it goes on another one, I don't have to move EBS volumes around, I don't have to have to attach or reattach. It's just readily available, and I just love utilizing the RDS amongst it. And my next step is to throw a CDN right in front of it, and I'm working on that portion as well. Now I can go back to the CloudFormation stack and add it, but I think I'm going to add it myself, and then I'm going to write a document on how I added for the WordPress site. And by the way, my first version of this really sucked. The images were bad because I'm not a UI person. So, I needed help on that front, and I think WordPress was just a lot easier to get it out in time for a webinar that we did.

Corey: This sort of ties back to what the beginning of our conversation started with, with the idea of working remotely, and we're seeing all these services now, and people are doing webinars left and right, which, incidentally, I was always very down on webinars: than I did a few, and the viewer numbers are lower compared to this podcast, or the newsletter, but people take action based on webinars way more frequently. I wind up with outreach based upon them. I've started to understand why enterprise salespeople swear by them. But the reason I bring this up in the context of writing a blog is that when building out a video studio, or spending thousands of dollars on equipment and getting everything set up, is similar in some ways to building out a blogging engine. Namely, people spend so much time focusing on those implementation details that they forget the most important part, which is you have to be able to tell a story that's interesting; that brings people forward. Otherwise, it doesn't matter what your production quality is. No one's going to care.

Jon: I completely agree. I mean, and just some of my blogs, a lot of mine are more personalized, or just my thoughts, random thoughts dropped down on there. And I actually find out, looking at some of the statistics, that more people are reading it, or there's more actions upon it, webinars that are happening, a lot of more people are reaching out to me for communication, and some feedback and I try to provide that as best as possible for them. But ultimately, I agree with you on webinars. I think they are obviously the latest thing that we can do during this certain times that we can share with our community, but we can also touch everybody on a personal level, as well.

Corey: Yeah, when I was looking at your blog, it was oh, wow, it looks like you only have five or six blog posts. Oh, another abandonware blog, and then I checked the dates and they're all within the past month. Okay, someone, at least, as being prolific and that's sort of the important part. You find your own voice and you see what the traffic starts to look like, and what resonates, and what doesn't. I mean, I have that problem, incidentally, with my newsletter.

It's one of the reasons that I still have some weak form of tracking built into the newsletter. I don't actually care whether you click on a link or not. What I care about is across the entire board. For example, what links do well, and what don’t. That in turn informs what I should be putting more of in. For example, I don't care at all about IoT. But it turns out the audience really does, which means I need to spend more time focusing on IoT based upon reader responses that I'm getting. But I don't want to be creepy. I don't care who you are. I don't care what you're clicking on. I just want to know, in general, what do people like versus what don't they like? Because feedback is a very hard thing to get objectively because you either get people who love everything you do or people who hate you, and neither one of those is what we would call unbiased.

Jon: No, I completely agree with you in that.

Corey: So, one last thing I wanted to cover with you before we wind up calling it a show. Talk to me a little bit about certifications. One of the reasons that I only have the Cloud Practitioner, among several other reasons in fact, is because it is challenging and annoying for me to find a place to go to a testing center. Well, now, all of these certifications can be conducted remotely, which means it's less of a pain. Tell me a little bit about that.

Jon: So, the certifications and the testings that started out—if you look back, and I believe it was probably January, February, on the timeframe that the Cloud Practitioner was allowed remote or online, that you didn't even have to go into a testing center and actually take the test. Due to the current state that we're in, the pandemic, those who were studying for an AWS exam couldn't go to a testing center. So, if they had stopped their exam studying, and then they'd have to pick it back up that was not really friendly for them. It does take a lot to study each of these domains, and the areas of expertise. AWS released your certifications and your exams proctoring at home online, you can proctor at home at any time available. Somebody will be watching you, obviously, through your camera. I've got some good feedback on that, on how they're being handled and monitored. And—

Corey: My CEO went on a bit of a rampage when he got his Cloud Practitioner on it, and credit where due, they fixed an awful lot of the problems that he wound up highlighting.

Jon: Yeah, exactly. I mean, so AWS, obviously, his customer feedback is the biggest thing for them. So, there's something that you see that doesn't work right, they take that feedback. Working and doing these exams at home have allowed people to kind of stay focused during this time, and continue their learning and training. So, when we are finally allowed to kind of go out, and things have kind of settled back down—I don't want to say normalize in a way—but things start to go back to normal, that they've got their certifications and they can continue on the path of their career, or their goals.

Corey: I used to be very down on certifications, and it turns out that my reasons for being down on them were almost completely invalid, and were also steeped in privilege, to be very honest with you. Because by God, I'm Corey Quinn, I don't need a piece of paper from Amazon saying that I know how Amazon works. I can just point at my own body of work, I can point at my resume of working on these things, therefore certifications are a useless money grab. And this was patently untrue, and extremely unfair because when you're starting out, particularly, there needs to be something that says, you have some baseline level of experience with a system. As your career progresses, that becomes a list of things you've done before, but early on you don't have that option, so certification is a terrific signal that you can at least have a discussion around the topic, and you can be conversant with the vocabulary.

Second, when you're running a large scale company and you want to get everyone trained up and validate where they are, it's a terrific thing to work with at large scale from an organizational point of view. There's also some partner program stories behind it that, frankly, I find a little silly, particularly in light of we're trying to keep business afloat but now we have to jump through partner hoops to maintain various positions within the partner program. It's frankly one of the reasons I never joined the AWS Partner Network. So, I've really come around on certifications. And at some point, I would probably consider getting more of them.

One idea I had was that I would love to try speed-running my way through all of them, but because you can't livestream these things, you can't share the contents of it, and there's not a lot you can do, the way I learn, that’s going to be helpful for other people. There's no story for that behind, “Look how smart I am.” And it turns out that that is not a compelling pitch, and there are better uses of everyone's time. But I still toyed with the idea because I have a trick memory and I test well.

Jon: Actually, I think you touched on it very well. It's not only the validation within your company but as personal validation. Testing yourself and challenging yourself, that you understand the basics, you understand some of the caveats that you're going to run into, and it will not only help you, but it will help customers and partners that you work with, your company that you work with, and some of the way you see things within AWS. And the second point that I have on that one is, I was always afraid to share this, but I failed my Networking Specialty exam. I found it really challenging in an aspect that I will complete it, I will pass it.

I actually even wrote a post on that one and shared it out on my failure, but it was a failure at this time. I didn't fail: I just failed to pass this time, or I didn't pass this time is the correct verbiage. I find it as a challenge, and I will continue to strive towards that. I had some cool and interesting questions in the exam, and the good part was, I didn't understand what that question was, I thought it was a certain answer. And I took it back to the team and I was able to grab insight on why it couldn't be that way, or that wasn't the answer. And I'll tell you what, I'll never fail that question again, or I'll never run into that issue when I work on a project and design something.

Corey: Yeah. And the problem, too, is that when you start going around the certification dance and trying to gloat and compare notes with folks—like, there were a few folks that I'm not going to name that when the Cloud Practitioner came out, they went out and got this, and talked about how quickly they were able to speed run through it. And sure, it's fun and all, but these are folks who've been working with AWS for 10 years. And when you say something like that, all you're really accomplishing is making people who aren't as steeped in the ecosystem feel crappy. That's not good for anyone, and it doesn't come off super well because you're not the target of that certification.

I mean, I have a similar story myself. I got one question wrong on it, and I know exactly which one I got wrong because I was honest, rather than following the guide, namely, how long does it take to recover an RDS instance from snapshot, or something like that, and I have some horror stories around that. So, I said weeks. Yeah, it turns out that's not the right answer. But I'm not the target market for that sort of certification. If we want to see the other side of it, have me go take the machine learning or database certification that apparently does not include Route 53, which is a travesty. And you'll watch those certs sweep the floor with me. We all have our areas of focus, and I think that comparison really is the thief of joy when it comes to covering these extraordinarily complex services in a way that is cohesive, accessible to a lot of people and, let's not kid ourselves, fits the format of a multiple-choice quiz.

Jon: No, I agree.

Corey: So, if people want to wind up hearing more from you, where can they track you down?

Jon: Corey, great question, I actually wanted to share this that if you follow me, check out my website, theAWSblogger.com, you can follow me on Twitter—and unfortunately I have to say, “At underscore,” because if you search for some reason, it just doesn't pop up—so @_jonmyer, J-O-N-M-Y-E-R, or look me up on LinkedIn. I actually post a lot more on Twitter now thanks to Corey, and I kind of share a little bit more information. Those are great places where I'll be constantly posting and sharing stuff out for the community, and everybody that wants to digest and regress some more AWS information.

Corey: Excellent. Jon, thank you so much for taking the time to speak with me. I appreciate it.

Jon: Yeah, and thanks, Corey, for having me on here. It's always a blast talking with you. I have lots of fun. The conversations are casual and nothing scripted. It just comes out.

Corey: Absolutely. Jon Myer, partner solutions architect for Cloud Management Tools. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts. If you've hated this podcast, please leave a five-star review on Apple Podcasts, and leave a comment explaining what's wrong with it in multiple-choice format.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Tim Bray:

Tim is a general-purpose Internet-software geek. He specializes in Web, search, writing, speaking, business. He founded Textuality in 1996. He is available for consulting on issues of technology leadership, software construction, and distributed systems. You can follow Tim's musings on his blog ongoing and on Twitter.

Links Referenced:

  • “Working Effectively with Legacy Code”: https://www.amazon.com/Working-Effectively-Legacy-Michael-Feathers/dp/0131177052
  • Twitter: https://twitter.com/timbray
  • Blog: https://www.tbray.org/ongoing/
  • Article Tim wrote about his PR FAQ: https://www.tbray.org/ongoing/When/202x/2020/06/21/A-Cloud-PR-FAQ
  • Tim Bray’s “Split AWS from Amazon” PR FAQ he wrote: https://github.com/timbray/a-cloud-prfaq
  • Tim Bray’s “Break up Google” article: https://www.tbray.org/ongoing/When/202x/2020/06/25/Break-Up-Google

Corey’s Fake PR FAQ: https://www.lastweekinaws.com/blog/introducing-aws-elastic-beanstalker/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: It's, at least as of this recording, morning on the West coast, which means there's no better time than to inflict a homework assignment upon you in the form of a 42 page ebook from StackRox. Learn about the dancing flames of EKS cluster security, evade the toxic dumpster of the standard controls, and tame the wild beast of best practices for minimizing the risk around cluster workloads. Become renowned for your feats of daring, as you learn the specific requirements for securing an EKS cluster and its associated infrastructure. To learn more, visit snark.cloud/stackrox. That's snark.cloud/stacROX.

Corey: This episode is brought to you by Trend Micro Cloud One, a security services platform for organizations building in the cloud. I know you're thinking that that's a mouthful because it is, but what's easier to say? "I'm glad we have Trend Micro Cloud One, a security services platform for organizations building in the cloud" or "Hey, bad news. It's gonna be a few more weeks. I kind of forgot about that security thing." I thought so.

Trend Micro Cloud One is an automated, flexible, all-in-one solution that protects your workflows and containers with cloud native security. Identify and resolve security issues earlier in the pipeline and access your cloud environment sooner, with full visibility, so you can get back to what you do best, which is generally building great applications.

Discover Trend Micro Cloud One, a security services platform for organizations building in the cloud. Whew. At trendmicro.com/screaming.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Tim Bray, formerly a VP and distinguished engineer at AWS, and now an unemployed bum. Welcome to the show, Tim.

Tim: Delighted to be here.

Corey: So, no matter what side of the world you live on and what your perspective is, it's pretty clear, there's one thing everyone can agree on, and that is that you are basically just the side of a war criminal. You were one of the original voices behind the XML spec, and for those rare few of us who really like XML, you also participated in the JSON atrocity.

Tim: Well. So, what's plan C if neither XML or JSON are going to work for you?

Corey: Uh, YAML. It's how we build our safest, whitest spaces.

Tim: [laughs]. Yeah, you know, YAML’s so convenient and pretty looking. Most senior engineers hate it for the simple reason that in YAML when you have a message, there's nothing that marks the end of the message, so if you screw up and drop some off the end, or the network breaks at the wrong time, well, you just go ahead and process YAML as if you got it all. But you didn’t, and that can cause some really horrible things. It's worth talking about XML and JSON a little bit though.

So, we did XML in the late 90s; it was completed in ’98. So, the really astonishing thing was before that, there just weren't any data formats that were vendor-independent, whether it's computer vendor or database vendor, or programming language independent, which is shocking because we'd been doing computing and networking for 20 years at that point. And so since XML was the first one, it got picked up and used for everything, which didn't work out that well. XML was dreamed up by a cabal of publishing technology bigots, who really cared about deeply nested document structures and tables of contents and index cross-referencing and things like that, which turns out not to be the greatest choice for using in REST APIs. Having said that, there's still a thriving XML community doing things like legislation, large-scale technical documentation—like for airplanes—humanities computing, so on and so forth. So, if you're in the publishing space, hey, that's what you want.

As for JSON, you know, JSON is irritating in some respects, but we've made it work. It does most of what you need. We all know what the next few things we would change in it, if we could, to make it better are, but we're not going to because it's too late. You know, all the software is written and you can't be changed anymore. The question of what the best way to package up data is to ship it back and forth between heterogeneous systems is interesting, and you can have a lot of fun. And these days, there's things like Thrift, and Avro, and protobuf, and people, for very poorly worked out and verified reasons, seem to think those are the universal choice in the future. Eh, I’m unconvinced.

Corey: [unintelligible] you can go schemaless with it, and have all kinds of fun ways of packaging. I mean, SaltStack back in the early days, with Python Pickles on top of that.

Tim: Yeah, that's absolutely true. And actually, that was one of the most interesting things I was working on there at AWS, is we did a schema repository last year. And that means that, from the AWS point of view, you now have data types with ARNs. And I think that there's way, way too many AWS services that interchange—

Corey: You can stop the sentence there if you'd like. Way too many AWS services. Full stop.

Tim: [laughs]. That's another interesting discussion you can have. You know, Andy loves to get on stage and talk about how many new things they shipped every year. And—

Corey: And everyone else gets this sinking sense of dread in the pit of their stomach when they see it, of, oh, dear Lord, at least 20 of those would have helped solve problems that I have, but I'm too busy doing my job that isn't keeping up with Amazon services.

Tim: Yeah, you know, having said that, are you actually going to say it's wrong because, you know, it seems to be going pretty well. Yeah, yeah, there's more there than any one human being can stay on top of, and in a lot of cases, people pick the technology based on what they already know, as opposed to what might be optimal.

Corey: I suspect that's the case on almost every scenario. People talk about, on stage, about the exciting far-future stuff they're doing, but if you look at actual cloud bills and the environments that the serious, huge reference customers have, it's all a bunch of, to be frank, the boring stuff. And we pick that boring stuff for the simple reason of, we know how it breaks. If you don't know how something breaks, it's hard to trust it with critical workloads.

Tim: Sure, fair enough. But in 1982, you could have said, well COBOL’s what everything runs in, so why would we be poking around looking at other stuff? And you need to throw a bunch of new stuff at the wall to see what sticks. And some of it is sticking. I mean, if you look at the numbers for things like Lambda, and EventBridge, and so on, you know, a lot of people are using that stuff these days.

And it is non-traditional, and the way things are done in 15 years is going to be quite different from the way things are done today, and the only way we're going to get there is by introducing new things, and finding out which ones actually meet a demand.

Corey: So, there's a lot of things we can talk about in the technical space, but first I want to get to the—well, I wouldn't even call it an elephant in the room, because everyone acknowledges it, talks to it, it's wearing a name tag, and often gets its own introduction, and slide deck on stage at the podium—you were a VP slash distinguished engineer at AWS. There are less than 20 of them at all of Amazon. It is the highest pinnacle of technical achievement. At other companies, they will be called ‘fellows’ in many cases. You resigned on the basis of ethical issues earlier this year. Talk to me about that, since that is virtually unheard of—at any company—for someone at that tier of technical achievement, and that level of seniority within an organization, to leave publicly citing anything.

Tim: Well, on the other hand, I didn't see, really, any other realistic choices. When you achieve VP rank, you're expected to be on the side of the company consistently, and be willing to get behind what the company is doing, and not be a [BLEEP]-disturber in public, and I respected that. And when I found the company was doing something that I just couldn't live with, I made that clear going through proper channels internally. And having done that, what other choices were there? It's not a thing that one can just get on stage and say, “Well, firing whistleblowers is okay, because—” because there's nothing that comes after the ‘because.’ so… resigning seemed like the only possible thing to do at that point. And it didn't make me happy. You know, AWS is a fantastic organization. I loved working there: the people are great, the customers are greater, the work is fantastic, but, you know… it just didn't seem like there were any other options.

Corey: I hear you. Specifically, for those who have not been following that particular saga, there were a number of high profile firings of warehouse workers who just so happened to be involved in speaking out specifically around aspects inextricably linked to labor organizing, which, heaven forbid, we wouldn't ever want the people who work in our warehouses to be able to bargain collectively; that could end badly for us.

Tim: Well, except for that's empirically false. I mean, there was a really great case study that just happened in May, I believe, where the Amazon warehouse workers in France were concerned about the handling of the COVID and didn't think the company was doing enough to protect them, and so they started raising their voices. But, you know, Amazon doesn't talk to unions. Only, in various jurisdictions in Europe, the law says you have to. So, the union took them to court and won a judgment that Amazon had to talk to the union, and ship only essentials.

Amazon reacted by shutting down all the warehouses—which doesn't seem like the smartest thing to me—and appealing the court judgment, which they lost on appeal. And then, having been backed into a corner, they went and talked to the union, in a matter of a few weeks, worked out a deal and got all the warehouses reopened and working. And, you know, Amazon's doing okay in Europe, so it is empirically true that Amazon can, in fact, work in a unionized environment and still be successful.

Corey: To be clear, my statement earlier was dripping in sarcasm, which doesn't always come through, especially if you're reading the transcript of this episode, rather than listening to it so you can pick out the sarcastic overtones of my voice. But yeah, it absolutely makes sense to me that if you are working in virtually any employee role, that organizing makes an awful lot of sense. And I say this as a business owner. I don't view the fact that employees can come and talk to you on a collective basis to in any way be a negative direction for society to evolve in. Now, I understand that Amazon has its own position on this, and that's fine. To be very blunt. I don't know what it takes to run an 800,000 employee company. I don't ever expect or hope to find out. I feel that if I've gotten to that point, something has gone egregiously wrong.

Tim: Well, it sounds trite to say, “It's just too big,” but I think it's just too big. And it's not just an Amazon problem. Amazon is a symptom of the problem. I think that it's not controversial to say that there is an overly high concentration of wealth and power in the economy. And that's specifically egregious and visible in big tech where you have Google, Microsoft, Facebook, Apple—who am I missing—the Big Five that dominate their sectors and behave in classically monopolistic ways.

And I don't think it's controversial to say that that is dysfunctional regardless of your ideology, even if you are a firm believer in capitalism first, last, and always, capitalism doesn't work that well when you get overly high degrees of monopolization and concentration. And I'm a little bit even more fundamentalist than that. I think that it just doesn't work that well when the companies get, just, too big. And there's some line in the sand when you get bigger than that it just becomes hard to exhibit any aspects of humanity in dealing with either your employees or your customers. And I think objectively, you can look back the last couple of decades and say, “No, a lot of things about the economy just aren't working that well.” The experience of dealing with businesses, as a customer or as an employee is not getting better, it's getting worse. And we need to explore some, I think, fairly radical measures to sort that out.

Corey: Again, I run a 10 person company and believe it or not, we do have an ethics policy of who we will and will not do business with. And to be direct, the way that I've structured that has been that in 15 years when my now three-year-old winds up going to college, when I tell her where her college fund came from I don't want to be ashamed by the answer. And that means different things to different people, and it's super squishy, but because I'm small, I have the privilege of being able to say that. I don't want there to be a logo on The Duckbill Group's website for customer wall that is the same logo you're going to see on bomb parts scattered around a village somewhere. It's just a baseline level of who I will and will not do business with. I think you don't get to be a certain size of company and still hang on to that. I don't think there's any large scale cloud provider in the trillion-dollar range of valuations where they've kept their souls. I don't think you can and still get to that level. Do you agree with that? Do you think that there's a way around that, where there was a better path forward, or is this just the trade-off you have to make for growing to that scale?

Tim: Well, the world is complicated, and nothing is white or black. I will say that in my view, AWS as an organization is run pretty ethically. There were very few things we did they gave me, personally, heartburn. And I noticed that things that are stinky do get addressed, like the recent change of policy around recognition. I would be much happier if the company walked away from a whole bunch of policing and security sector businesses. But would just making it smaller address that? I don't know.

I think that, in general, if you dislike the way that larger organizations are behaving, the only effective solution to that is to change the rules because Amazon and its competitors—AWS and its competitors—by and large, sort of, kind of, do play by the rules. And as John Adams says, we still are mostly an empire of laws rather than of men, although if he’d said that these days, it would be people. And so change the laws. I mean, politics is boring, and icky, and slow, and tedious, but it's really I think, the only hammer we have to drive this nail. So, if you don't like the fact that the company is selling to businesses that you think are awful, well, maybe we should change the rules to make those businesses not possible anymore.

If you don't like the way that ICE behaves, well, change the way that—the legislative framework so that ICE is simply not allowed to do all the egregiously brutal things they're doing these days. I hate to say it; I mean, I would like to be able to talk to, say, Facebook, the way I talk to my kids and say, “Now you play nice, or no dessert.” But that doesn't work. It has to be done, I think, through traditional legislative frameworks.

Corey: Now, and it comes down to a lot of power imbalance, too. We saw recently, AWS filed yet another non-compete lawsuit against a former employee, in this case, Brian Hall, who went from AWS to GCP, and on some level, it’s, yeah, okay. It's a contract that you signed, and you should live by that contract. I'm sympathetic to that argument. The counterpoint is, though, is the way it is scoped is so hilariously overbroad. It’s scoped to all of Amazon—which, first, can you name a single industry that you could safely guarantee by the time you leave that company, Amazon will not be doing anything with?

And it covers anything you may have had access to, by a strict reading, which is really the only way to read something like that before you sign it. And it's the only company I've seen that enforces these things this aggressively and at this scale. And I've got to say more than almost anything else, things like that, where Amazon is, I guess, unkind and punching down are the few points where I lose respect for them. Sure, they can ship a bad technical product, and I can make fun of it, but we're all still friends at the end of the day. You can't start punching down at your own employees, which is really what ties this all together, and still be someone I look at with an, “Oh you,” look on my face.

Tim: Well, yeah, and obviously the thing that drove me from the company—the firing of the activists—is you can describe it using similar terminology: ‘Punching down.’ Having said that, it isn't just Amazon. Microsoft has a history of doing this, too, bringing up the non-compete cannon. And what Amazon and Microsoft have in common? Well, they're both Washington State companies.

Non-competes are non-enforceable in California, and that's one of the big reasons why a large part of the technology industry is still in California. I mean, the birth of Silicon Valley was Intel, and Intel was an act of treachery when a bunch of people from Fairchild Semiconductor went off and decided to take what they learned and build a new company around it. And that's fine. I've got no problem with that. And I think Washington state would do its citizens and its employers and its employees a service by striking down the ability to enforce non-compete clauses.

Corey: I would absolutely hope so. I think that there is so much that goes wrong, and could be better than it is. It's frustrating. I think that non-competes are one of the most aggravating things in the world because when it comes time to go somewhere else, there is such an imbalance of power.

I mean, I'm a small business. If I'm trying to hire someone and Amazon reaches out with, “So, just so you're aware there's a non-compete issue here,” It's very easy for me to say, “Oh, I don't want to go into a legal battle with Amazon. So, I'm just going to go with my second choice candidate instead.” It provides an incredible chilling effect. Now, personally, because I am who I am, my response is, “Bring it; I will burn this place to the ground fighting you if I need to,” just because I don't know when to pick and choose my battles. But it is the common case that it is incredibly chilling. And the cavalier attitude, and the leaked memos that come out of Amazon about these firings has just been bizarre to me. It’s, what are they so afraid of?

Tim: Well, yeah. At the end of the day, it was the ethical dimension of the whistleblower firings that drove me out, but you also have to point out that that was, like, really stupid, egregiously stupid. I thought, “Maybe I'm missing something.” As you say, I don't run an 800,000 person company, either, but firing whistleblowers feels like, “I never heard of the Streisand effect. I'm putting up letters of fire 50 meters high in the sky saying ‘hey, there's something here we really don't want you to look at.’”

I don't see how anybody with any maturity could think that doing that would have a good result, creating a climate of fear in response to activism. And the people who are activists had no thought of gain for themselves. They weren't trying to make money, or advance their careers, or anything like that. They were doing something that I thought was wholly admirable. I mean, they could have done something like say, “Okay, you guys are upset about that. Great, you got a new job serving on the task force to make the warehouses safe, and prove we've done it.” There was a certain amount of loss of respect in my mind for leadership around that.

Corey: I would agree. I think that there's so much that Amazon could do in this space, and demonstrate real leadership and, for whatever reason, it's choosing not to. I mean, I consider myself an Amazon fan. I think that they do a lot of good things, and I think a lot of what they do is admirable, but I'm believed when I say those things, because I also call them out when they do things that are awful, and this is one of those awful things.

Tim: In terms of all the places I've worked in—I'm old guy: I’ve been doing this for 40 plus years—AWS was by far the best managed place I ever worked, and also the asshole density was really low compared to anywhere I've worked. And, you know, there would be whole weeks at a time when I never got mad at anybody. Which, as you know, in most jobs is not the common state of affairs. So, yeah, they're doing a lot of things right, and that's why this kind of stuff stands out in such stark contrast.

Corey: One thing that—especially in the wake of the non-compete lawsuit, which I've been getting more traffic on than I have the warehouse firings—I've gotten a bunch of outreach, both from people considering working in Amazon, and people who do work at Amazon expressing concerns about their non-competes, and what should they do? And I have my own thoughts on it, but where do you stand?

Tim: So, I got the same thing. When I quit, I got a lot of outreach from Amazon employees saying, “Oh, gosh, now you quitting makes me feel bad. Can I go on working here?” And to be fair, I had to point out that I am elderly, and nearing retirement age, and financially secure, so this was a lot easier for me than it would be for a lot of other people.

Corey: Oh, we are both dripping in privilege in this context.

Tim: Absolutely. And I think that at the end of the day, that the problem isn't really Amazon. The problem is the Reagan/Thatcher consensus, 21st-century capitalism, and the imbalance of power and wealth in modern society. And if you're going to boycott anything remotely related to that, you're going to have a hard time making a living. And there are lots worse places to work. Amazon is—by a wide margin—not the worst operator out there. So, I've actually been advising people, “No, don't do what I did, unless you're absolutely unable to sleep at night, and got a financial plan.” But that's all I could say.

Corey: In what you might be forgiven for mistaking for a blast from the past, today, I want to talk about New Relic. They seem to be a relatively legacy monitoring company, and I would have agreed with that assessment up until relatively recently, but they did something a little out there: they reworked everything. They went open source, they made it so you can monitor your whole stack in one place. And most notably from my perspective, they simplified their pricing into something that is much more affordable for almost everyone. There's even a free tier with one user and a hundred gigs per month, totally free.

Check it out at newrelic.com.

Corey: So, let's pivot a little bit towards the thing that, again, you were one of the best in the world at, which is sort of funny, because originally, you were one of less than 20 people who had reached the pinnacle of the individual contributor engineering ladder at AWS, then you resigned in protest and suddenly, “Oh, that guy had terrible judgment. We never liked him anyway. It was probably an aberration,” was the entire thrust of the messaging around you. That's some impressive level backpedaling, but let's not kid ourselves here, you can be as modest as you want, but I'm not going to be.

You’re one of the best engineers in the world. That is why you were in the role you were in; it's why you've done the things that you've done, and I want to talk to you a little bit about technology. So, from a high level, let's talk about cloud stuff, for example. I think most notably in a combination technical/political venue to transition us, let's talk about fear of lock-in. a lot of people are scared of being all-in on a provider that reduces their choices in the future. Where do you stand on it?

Tim: I think it's a really important issue, and not one that simple or straightforward. You know, AWS is a $40 billion business these days, which is to say bigger than a lot of the customers. And in the IT world historically, there has been a trend to platform lock-in. For example, if you are a large company, you no longer get to decide what your budget for desktop technology is; you pay whatever Microsoft says you're going to pay; you have no bargaining power. Similarly, on your database spend, you pay whatever Oracle says you're going to pay; once again, no bargaining power.

And you just know that in a lot of industries that have some decades of experience in IT, the leadership, the CEO, the board, even, is going to the CIO and saying, “Now, don't you come back here in a couple of years, and tell me that you've lost control of our cloud spend.” And, okay, that's a very reasonable concern to have, and I can't see how you could argue against that point of view. Well, you can because in practice it turns out to be really difficult to be cloud-agnostic. I mean, if all you're going to do is rent CPUs, and storage, and operate databases, fine, you can do that. But even at that level, you find that there's a bunch of really annoying semantic gaps between the way the major cloud vendors think about what you mean by ‘instance’ and what you mean by ‘object’ and so on and so forth.

So, you can in fact build out your technology stack in a cloud independent way, but you're going to pay a lot more, and you're going to go a lot slower. So, I think the choice for a startup is a no brainer. A startup should pick one of the cloud vendors, jump into bed, aggressively make use of all the hot stuff they're shipping in that particular year, and that way, they'll get the most velocity; they'll spend less; they'll move faster; they'll have a better experience. You know, startups are short of time and short of money, so that's what they should do. Now, if you're a bank, or a pump manufacturer, or a book distributor, or something like that, can I honestly say that, “Ah, blow off that stupid lock-in concern, just go jump into bed with a cloud vendor?” I don't know.

I think that is a more nuanced discussion and one that you have to have. How much are you willing to pay? Now, a lot of people say, “Well, the cloud lock-in isn't really at the level of the APIs. It's at the level of data gravity.” Once you've got several petabytes of data sitting in Google Cloud or AWS, well, you're just not going to move out because it's too hard. I’m not sure that's true. I think you can in fact do that. And in fact, these days, you can go on a case-by-case basis. I mean, are you willing to use a proprietary database like Dynamo, or BigTable, or something like that—

Corey: Or Route 53?

Tim:—[laughs] in the knowledge that, then you really pretty well are going to have a hard time getting off that vendor. Maybe not. On the other hand, if you are going to decide, “I’m going to standardize on Postgres. Am I willing to use hosted Postgres, and use somebody else's control plane to create databases?” Well, that's a much less severe degree of lock-in.

And this is a very complicated discussion, and I would be willing to have an opinion after a discussion with any individual business that is getting into the cloud, but would I lay down a dictum saying, “Oh, fear of lock-in is dumb, or always resist lock-in?” I would not say either of those things. It really is a case-by-case situation.

Corey: Very much so. My railing about multi-cloud is always from a best-practices position. I think that as a general best practice, pick a provider—I don't care which one—go all-in. Now, there's a litany of exceptions to that, where it does not make sense, where you cannot do that specific thing, and that's fine. My argument has been against hearing this from a conference stage from some second or third rate vendor saying, “Oh, multi-cloud is absolutely what you want to be doing,” and accepting it blindly. That has been my issue.

Tim: Mm-hm. Fair enough. And I think this is a super hot area of technology, and I think there's lots of opportunity for vendors to facilitate the process of being multi-cloud where necessary. I mean, HashiCorp is now a unicorn, right? And we should all be paying pretty close attention to what they're doing because I think their existence shows that there's a hot button out there that they've pressed, and a lot of people care about a lot.

You know, at the moment, I can only speak to AWS because that's where I worked, but AWS is a really pretty safe place to jump into bed with at the moment because people like Andy Jassy, and Charlie Bell, and so on, are just really customer-obsessed. They spend a huge proportion of their time trying to figure out what's making customers unhappy and just fixing it. And they're super aggressive about cutting prices, and things like that. And that's great, so that's the kind of vendor I want to have as a customer. But 10 years from now, when I'm retired, and Charlie's retired, and so on and so forth, is it still going to be the same, or is it going to be a bunch of ex-private equity MBAs running the thing and trying to figure out how they can extract the maximum rent? It's a thing that's reasonable to worry about.

Corey: Well, always a disturbing concept. One thing you said a minute ago about HashiCorp being aimed at the multi-cloud story. When I had Mitchell Hashimoto on the show previously, we talked about this. His idea was never that Terraform should be used in this multi-cloud way. It's not about workload portability, it's about workflow portability.

And having seen that play out at multiple customers in the past, it's always seemed to align in that particular direction, where, okay, you're always going to have to do a lot of rework to get something that works on one provider in Terraform to work somewhere else, but the workflow, how you interact with your codebase, how you get things from your head into production, that's what you can wind up porting between providers. So, I want to be very clear on that, or oh, my stars, will I get letters. But let's talk about something else that I think has areas of lock-in commonality. Let's talk a bit about, for example, the idea of event-driven architectures. It feels to me like the event model is one of the least portable things you could have, depending upon what it is you're doing. Am I wrong in that?

Tim: It depends. I mean, each of the cloud vendors has their own event pub/sub, routing, invocation frameworks, and it's shocking how similar what they do is. If you care about all the different semantics that you can have around messaging and eventing services, there's not that many. And you can make up a laundry list of things that you care about. And you say, “My requirements are A, B, and C. And yeah, I can do that at Google Cloud, or AWS, or Azure.”

And given that, it's kind of annoying, that things are, relatively speaking, incompatible. There's a lot of people I talked to who were using things like Apache Camel, which is this big, hairy, complicated API, but what it does is provides an abstracted send message, receive message thing. And I'd see people using that, and then wiring it up to SQS, or RabbitMQ, or something like that, and then the person who's running the Java production code didn't have to care about their messaging framework. So, there are some tools for making it less proprietary. But having said that, that's what you kind of expect, because the event-driven stuff is sort of the new shiny at the moment, what with Lambda, and EventBridge, and that kind of stuff.

That is where we're pushing back the boundaries and figuring out new ways to build applications. So, you go explore the new territory first, then after you figured out what you use, you write the rules for it. You know, after you figured out what works, you write the rules for it. And we're still in the stage of figuring out what works, and the initial signs are that, by and large, event-driven software works pretty well, and is highly applicable in a lot of situations. But do we have a set of commonly agreed on best practices and standards yet? No.

Corey: That's part of the issue is that as standards fail to emerge, and they start turning into these, I guess, de facto standards rather then ones that are imposed, there's an awful lot of companies jockeying, for whatever it's worth, to start being the voice that determines what those standards are going to look like. And increasingly, it seems that those standards have been driven less by the community and more by whoever has market share and is the first mover in the space. Does that align with your experience?

Tim: Well, this is the classic thing is you get a market leader who becomes the de facto incumbent, and then all the market trailers get together and form a standards organization. [laughs]. And, whereas they would never come out and say that, one of the goals of the standards organization is to slow down the incumbent to give the up-and-coming smaller parties a chance to compete on a slightly different playing field. And that's okay. That's fine.

I've been in both roles now, as the market incumbent and as the upstart trying to get a standard put in place. It's part of the organic flow of how we do things in this technology. In the space at one point, SQL was seen as a highly proprietary technology, and these days, it's regarded as a pretty level playing field. I'm not smart enough to predict how this is going to shake out, but I think I’d put my stake in what I said a couple minutes ago, which is that, you can either use the cutting edge, sharp new stuff that does magic, or you can play on a level, well-understood, standards-driven playing field, but you really can't do both.

Corey: Yeah. It definitely seems that for better or worse, there's going to be a shakedown at some point. It feels to me--this is aligned in a lot of ways, with the idea of complexity gets built on top of complexity, and it almost like a sawtooth graph because eventually people look at it and say, “This is nuts,” And it collapses down into something that a human being has a hope of understanding. Kubernetes, right now, feels like it is incredibly complicated and overwrought. In a few years, it won't be because something's going to come along with, huh, okay. Maybe a team of engineers who cost a quarter-million bucks apiece is not something every company is going to want to run its application architecture, and there's going to be some evolution there. I guess cynically, I tend to view things like that let companies cosplay as cloud service providers themselves when they don't work for one.

Tim: Yeah, I don't want to diss Kubernetes too much because it does some—

Corey: Oh, I do.

Tim: [laughs]. It—I don't actually use it, either, but I don't want to diss it too much. It does some remarkably clever things that do make operator's lives easier once you've got everything set up. But it fails one important test in my mind. I think that technology should make easy things easy, and difficult things possible, and Kubernetes clearly fails the first half of that. It's too hard to understand, and bootstrap, and it costs too much, as you just said, in terms of engineering to keep it going.

We've seen this movie lots of times before. There was ISO networking, and seven-level stack, and so on. And then the internet geeks came along with TCP/IP, which only did 20 percent as much, but it turned out to be the right 20 percent, and [unintelligible] rest of that stuff away. Same thing happened with the web. There were increasingly elaborate, complicated application framework, starting with Visual Basic and, you know, and Adobe's offering, and Novell’s offering, and so on and so forth. And then the web came along with a vastly simpler and impoverished user interface paradigm, and swept all that crap away because the 15 things it did really well turns out to be the ones that matter.

And I would be deeply unsurprised if that didn't happen to Kubernetes. Kubernetes does a whole bunch of stuff, and all of it is not equally important. What I frankly expect to see is something come along that does the things you actually really care about, and you can read about it Monday morning and have an app up Tuesday afternoon, which is incredibly important. And if you look at the adoption of each real game-changing technology, it's almost always had the characteristic that it's easier than what came before. And this profession is hard enough as it is, without making things unnecessarily complicated.

Corey: So, at a high level, now we spent some time looking back, as you said, you've been doing this for 40 years. What technology trends are we seeing that interest you the most? What's the future look like?

Tim: At the moment, well, it's one that we already talked about a bit, which is, are there interesting tools that are going to be truly useful, and enabling you to avoid vendor lock-in without grossly increasing the cost, and decreasing the velocity? We can all be a little bit cynical about that now, but these things have happened, and the most classic example of that was the World Wide Web when it came along. Before the web came along, there were all these competing technology stacks, and you had to figure out whether you were going to be on the Microsoft stack, or the Sun stack, or the IBM stack. And then the web came along, and it was a platform without a vendor. That was a profoundly important thing; it was the first time we'd had a platform without a vendor, and is there going to be a public cloud platform without a vendor out there? Any steps towards that would be super interesting in my mind. So, that's one.

Another one moving in a completely, totally different direction is, I'm all excited about Augmented Reality, AR. How many years has it been, since a new mobile device changed anybody's life? Come on, they’re just not interesting anymore? Or, how long has it been since a new programming framework has really rocked the industry? Well, serverless in 2014, maybe, but they're few and far between.

So, AR I think has the ability to really dramatically enrich the life of humanity at large, and right at the moment, the technology isn't good enough. What you would like to have is technology so you can just hold your mobile phone up and look through it and see a decorated version of the world. You know, you're walking in a park at night, and you can see fiery dragons climbing up the trees. You walk into a superstore and say, “Where's the deodorant?” And a big arrow appears in front of you pointing towards that. There's so many applications.

Technology isn't good enough yet, but when the technology does become good enough, it's going to be a huge new area of human creativity, and innovation, and new business sectors that we can't even begin to think about. There's little birds here and there telling me that Apple's got some great stuff coming along. I’ll believe it when I see it, but the problems we're trying to solve are basically computer vision problems, and ML problems, and so on, and they'll get solved all right. So, that one's got me all excited.

Another area that I'm super interested in these days is public sector procurement. In my career, I got, twice, mixed up in large government procurements of technology, and oh my God, what a cluster-[BLEEP]. It's done very, very badly, and I think this is, sort of, generally universally true across the governments of the world. And the blue suit consulting companies that specialize in government technology procurements are just not working properly, and we have endless disasters. So, I'm aware of some efforts to try and bring sanity to that process, and that's super interesting to me as well. So, we could dive down any of those rabbit holes.

Corey: I think that the problem with rabbit holes is that in some cases, there's a rabbit at the bottom and, be careful, that rabbit’s dynamite; pointy teeth, and whatnot. One of the problems that I see is, as I look across the landscape, everyone's talking about the future, about the trends of what is coming. But if I look at my customer base, and I look at what is driving the actual spend—where the money's going—it seems that a tremendous constituency is treating the Cloud as someone else's data center, where they run VMs steady state, there's no auto-scaling, and it seems like the primary problem that they're trying to solve is that they suck at running data centers themselves. Some people give up halfway, and call it hybrid, and it feels on some level to me like they are improving their data center environment at the cost of leveraging the Cloud for what it's really capable of doing. And I find that depressing on the one hand, but on the other, I understand that they have what we'd like to disparagingly call legacy workloads, which means they make money. How do you feel about that?

Tim: Well, to certain extent, that's simply an inevitable result of arithmetic. So, AWS is, what, $40 billion a year now, and maybe the whole rest of the cloud industry together is that again? I don't know, something like that. And so $80 billion sounds like a lot of money. Last time I checked, the enterprise IT spend globally is like $2.7 trillion per annum. So, the whole public cloud infrastructure is a small sliver of the business.

Corey: Oh yeah, everyone talks about AWS being a $40 billion annual run rate, but Dell’s revenue was $60 billion in 2012. There's a tremendous—and that's one company that basically was selling hardware and some services back then. We’ve only seen the tip of the iceberg.

Tim: And for that very reason—as you say, running your own data center sucks. The cost, admin, velocity, security advantages of moving to the public cloud, are—I think everybody agrees now—are pretty high. And so I think that whether we like it or not, the revenue growth and core business of the public cloud is basically going to be doing exactly what you said: getting the people the hell out of there data centers, and improving their security posture, including their availability posture, that kind of stuff, just by moving existing stuff to the public cloud, your classic lift and shift. So, okay, that's fine. There's nothing wrong at all with that. It's a little bit boring, but it's an important big money, big deal.

So, let's look at what comes next. If you look at your typical large enterprise, they'll have an inventory of apps—a couple of hundred is typical—apps that their business depends on, in any given year, they maybe do work on a single-digit number of apps. You know five or ten kind of thing, and the rest of them just tick tick along. And so what's actually interesting is the new stuff that gets built. And so, for new apps that are getting built, are they being just done on classic monolithic vert architectures, you know, web server in front and database engine behind? Well, yeah, some, and there's nothing wrong with that, especially for smaller-scale departmental ops.

But anything that's being built that's new, and ambitious in scale, I think the technology uptake of the new stuff that I was working on is pretty good, actually. And incrementally as we do these five apps this year, another five apps next year, eventually the present and future start outweigh the past, but at the moment, they don't. The past is winning, we're moving the past from data centers into the cloud, and realistically that's where the revenue is coming from right now in its largest bulk, and that's okay. I don't see that there’s anything wrong with that, but obviously, those of us with an eye to the future are really, really worried about making event-driven, and serverless, and container, and all that stuff, technology, irresistibly attractive for the construction of the new apps.

Corey: The idea of throwing away the old in favor of the new is one that feels pervasive, and is something that everyone loves to talk about, but it feeds into a common deception that I'm seeing where everyone seems to believe that just after this next sprint is over, then, then we're suddenly going to start making good decisions instead of the bad ones we've made to date. And, oh, I'm as guilty of this as anyone; this is not me casting shade. I wish that oh, as soon as I get a little chance to breathe, I'm going to go back and fix all of my crappy broken code. Spoiler: he would not. But it's the hope that springs eternal. And you see that companies never seem to quite outgrow this.

Tim: Well, as a former principal engineer and distinction engineer at AWS, one of the things that principal engineers spend a lot of their time doing is stopping people from doing that. You know, respect the past is a core engineering principle. And you may hate the existing code base that's running your business-critical applications, or your business-critical AWS service that was launched prior to 2010, but it works. And part of the problem is that a lot of developers hate reading other people's code, and don't want to learn how it actually works, they just want to rewrite it themselves. And once you get to be in a position where you've done this for 20 or 30 years, you realize that, you know, that isn't as easy as you think it is.

And embodied in that crufty old codebase is a huge inventory of decisions that were made to meet particular weird situations and corner cases, and achieve non-obvious behaviors to turned out to be correct, and there's no way to know that by looking at it. Now, things are getting better. There's this great book called Dealing With Legacy Code. And it defines legacy code, interestingly—nothing to do with age or anything like that—as code that lacks unit tests.

And since unit tests became pervasive sometime between 2010 and 2020, things have gotten better because in many cases, the unit tests realistically represent the contract between the codebase and the outside world. And they make it much more thinkable to replace the codebase with something that's more modern, runs faster, runs cheaper, runs cleaner, emits less carbon, and if it still passes the unit test, hey, it's probably going to work. So, yeah, respect the past. Don't flippantly decide that you're just going to rewrite the system because you're smarter than the people who wrote it, because you're not. But on the other hand, we should be super dogmatic and make sure that we never do anything that isn't fully and completely unit tested to enable our successors to be a little bit more courageous in replacing the past where that's the right thing to do.

Corey: I think that's a really good takeaway here. If people want to learn more about what you're up to these days, how you're thinking about things, where can they find you? Now that you’re an unemployed bum, there's no corporate blog, you can drive people to. Where do they find Tim Bray?

Tim: Just type my name into Google. There's lots of stuff that comes up. I talk to the world on Twitter and on my blog, and I enjoy doing that, and I'm not going to stop doing that as long as I can lift my hands up to the keyboard. I published a PR FAQ for how to spin off AWS from Amazon, which is something that should obviously be done, but I also published an article about how Google should be broken up, which is another thing that should obviously be done. And I'm not as funny as you, Corey, but I do try to be kind of entertaining.

Corey: Oh, I was a big fan. I took inspiration from that to publish my own fake PR FAQ about Elastic Beanstalker. I love the format; I think it's an underappreciated comedic medium.

Tim: [laughs]. It's definitely an underappreciated medium. I've had a few advisory consulting gigs since I left, and a lot of people want to learn about what's this six-pager thing that Amazon talks about, and how does that really work? And it does really work, and I think I've given them value by talking to them about that. But I think the comedic six-pager is an underappreciated opportunity, so keep running with that.

Corey: That's a good way, I think, of framing it, and a decent enough place to leave it. Tim, thank you so much for taking the time to speak with me. I know that your days are full of unemployment these days—

Tim: That's right.

Corey:—but again, your time is always appreciated.

Tim: Oh, well, thanks. It's been delightful and entertaining, and I love talking about cloud stuff, so give me a call anytime.

Corey: Don't offer if you're not serious. Tim Bray, former VP and distinguished engineer at AWS, now unemployed bum, I am Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts, whereas if you've hated this podcast, please leave a five-star review on Apple Podcasts, and a comment telling me what's going to get you to leave your employer on ethical grounds.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Katie Bullard

As President of A Cloud Guru, Katie leads the sales, marketing, customer success, and partnership teams for the world's largest and most trusted cloud training platform.

Links Referenced:

  • Main company site: https://acloud.guru/ and https://acloudguru.com
  • A Cloud Guru Twitter: https://twitter.com/acloudguru
  • A Cloud Guru LinkedIn: https://www.linkedin.com/company/a-cloud-guru/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Katie Bullard, newfound president of A Cloud Guru. Katie, Welcome to the show.

Katie: Hi, Corey, thanks so much for having me. I'm really excited.

Corey: So, at the time that we're recording this, how long have you been with ACG?

Katie: You know what, I think it's been almost exactly three months. It feels both like three weeks and three years.

Corey: Yes. And, surprise, a pandemic showing up in the middle of that seems to have made our sense of time all go a little bit wonky. Like oh, ACG, haven't you been there for eight and a half years by now? And no, no March just felt that way.

Katie: It is definitely a very interesting time. I was talking with my team last week, just in general, about how we're all feeling, and the word that we used was just off-balance. To that point I think, time has taken on a different meeting, our routines have taken on a different meaning, but the positive thing is, we're literally all in this together.

Corey: So, A Cloud Guru has been one of those companies that an awful lot of people not only know about, but have strong positive feelings for because, for better or worse, that's how an awful lot of us have learned various things about the Cloud. It's basically taught the world to Cloud. And what's always surprising when talking to people about the company is that, “Wait, you mean there are people who work there who aren't in front of the camera, teaching things eight hours a day?” It becomes this weird perception that when you listen to people who are teaching you things in an instructor setting, to forget that there's actually a company behind this. It's not just the people who are on camera, or who are teaching a particular course, but also there's an entire logistical operation. Historically, you had the founders, Sam and Ryan Kroonenburg, who were themselves involved in various video properties as well. So, it had a very small company feel. How big is A Cloud Guru?

Katie: So, we're now approaching 400 employees across the globe. About a quarter of those are our instructors, as you mentioned, but three-quarters of them are all of us who are behind the scenes. And what's really kind of funny is when I first started talking to Sam and Ryan about six months ago, the company had 100 employees. So, that gives you an idea of how quickly it's growing and scaling.

Corey: It's one of those business problems that there's virtually no one on the other side of. It's, “Wow, we're teaching people how to do a thing that is very clearly transforming an entire industry—or basically the entire series of industries.” It's hard to name a single sector that is not being radically transformed by Cloud to some extent, and the idea of teaching people to work within that new paradigm is one of those very rare things where there's really no one on the other side of that issue. Where in some cases, it’s, “Oh, buy this particular product to do monitoring.” “No, that person is doing monitoring the wrong way. Buy this other competing product.”

In this case, the story of learned work in this new world, more effectively, it really then just comes down to a discussion of what is the most effective way for different people to learn. And an awful lot of folks respond super well to the idea of video instruction. Personally, that's never been the way that I tend to learn things best for me. It’s, I’m going to build something and get it hilariously wrong, and through doing that, that's my learning approach. But I've come to find that, talking to people both on this show and as I walk through the world, I'm very much in the minority. There's an awful lot of folks who find the idea of having an instructor-led lab-style session: here's what we're doing and how, to be a much more appropriate way of learning things correctly. And looking at my half-baked understanding of a lot of these technologies. I have to think that right.

Katie: Well, you know what's really interesting is I think this notion that we all have different styles of learning covers any sort of education spectrum, whether it's cloud education or math education, and I've only been here three months, but my first month here was the acquisition of Linux Academy. And as we've really dove into the two very different styles, actually, of the two companies and the two online education experiences, what we heard over and over again from customers and students was, we really love the video learning, the short, kind of, 20-minute sessions that it A Cloud Guru has really become known for, and/or we love the hands-on labs, sandbox environments, like, let me go play, focus that Linux Academy has had. And so, from our perspective, what we're trying to do is deliver as much value to our students in as many different forms—to your point because we all have different learning styles—as possible. I think that's one of the things that I got really excited about in my first week here was the opportunity to really do that to a level that both A Cloud Guru individually and Linux Academy individually had not done before.

Corey: What were you doing before you would up in A Cloud Guru?

Katie: So, I was actually the president at ZoomInfo, which was previously a company called DiscoverOrg. We acquired ZoomInfo; it was a data and content business as well, but we sold to sales and marketing professionals. And for the last four years, I've been president over there leaving marketing, product, engineering, actually IT, and got to see that company grow from about 200 employees to 1200 employees in a four-year period from 40 million in revenue to 400 million in revenue. And how you build a really great, quite frankly, content business in a software that delivers a ton of value to people who in many cases aren't sales or marketing professionals, that are entry-level. They're just trying to figure out how to do their job better. And that was one of the things I really liked about ACG was so many correlations. Totally different industry; different student profile, but in many ways, we're all trying to do the same thing, which is, figure out how to either be better at the job that we're doing now or learn a new job.

Corey: And that's always a scary thing. I learned to speak publicly, insofar as I'm capable of speaking publicly, years ago as a traveling corporate trainer for Puppet, which was an interesting series of coincidental experiences all wrapped into one. First, it's software, and you're doing live demos and software does what software does, by which I mean: break. It's also, at that point, was perceived by the students, who were generally systems administrator types, that this was the software that was coming to automate them out of a job, so they were not super thrilled from that perspective. And when people are paying top dollar for this, and they're already upset, and demos break, they don't have a lot of patience. So, you learn to deal extemporaneously with a lot of stuff, you have to be able to find a way to build commonality for people to relate to you because otherwise, you're never going to be able to teach them anything.

And that was a heck of a learning curve that I had—learning cliff really, that I had to surmount and I somehow came out the other side somewhat intact. But it really gave me a keen appreciation for just how difficult education is, not because the technology or the subject matter itself is hard, though it is, but because there's so much of a human piece that is built on top of that, that I think it's incredibly overlooked by folks who've only ever been on the student side of the education system.

Katie: It's so true. And like you said earlier, it's scary. For somebody who's trying to learn something for the very first time, for any of us, that's a very vulnerable place to be. And so you want a safe environment to do that in a way that allows you the space and the grace to fail, to succeed. And I think that's what people gravitate to, honestly.

I'll say for me personally, I don't come from a tech background. Actually, the first half of my career, I was in architecture and real estate—very different—and I ended up getting into a real estate software company, which is how I landed here. But in every job I had, I was having to do something that I had never done before. I remember when I got first asked to actually lead an engineering department. Honestly, Corey, if you had asked me what API stood for, I would have no idea.

Corey: It wouldn't have helped if you’d asked the question at Amazon because presumably, they pronounce it AH-pee the way they pronounced AMI is AH-mee; same story. Yeah.

Katie: [laughs]. I remember when I first got introduced to ACG I was like, “Oh, man. I sure wish I had known about this back when I was first starting to take on one of these organizations because I would have felt way more competent talking to my team at the time.”

Corey: So, I have to ask. Growing up one of the seminal moments of my childhood was my dad, who demonstrates the same sort of impulse control that I do, decide it’d be fun one day to go out, and buy a trampoline for us. And my mom came home and saw this trampoline, and given that she has relatives who work in emergency rooms said, “Absolutely not. What do you mean this thing just showed up here? This was not the arrangement we had.” Long detailed story short. How similar to that story of, “Wait, you bought what?” Was you showing up and finding, “Oh, you know that whole Linux Academy thing? So, funny story: we bought that.”

Katie: [laughs]. Well, to be fair, I did get a preview of it; not before I said yes, but before I actually started.

Corey: Gotcha. It's one of those, like, “So, you'll laugh about this later but…”

Katie: But I will tell you a funny story. So, I was actually in Costa Rica when I accepted the job at ACG this was back in October—

Corey: Oh, vacations. I remember we used to be able to take those.

Katie: I know, I know, it was crazy, back when we could travel. So, I was in Costa Rica. And I remember telling Sam, “I'm so excited. I'm so, so, so excited, but I've got to move from Portland to Austin. I'm on vacation. Let me start in January.”

So, we agreed on this date in January. And it was literally six days later that he called me, and he was like, “So… uh, you can still officially start in January, but there's this thing that might be happening,” and I had actually just gone through a really similar acquisition with DiscoverOrg and ZoomInfo, like, literally, deja vu scenario. He's like, “and we would really just love your expertise. So, even if you're not full time, just come help us. You don't know the business yet, but you at least know the questions to ask.” So, I was like, “Yeah, yeah. No problem. I'll do it ‘part-time’—” I'm putting this in air quotes right now, “—I’ll do that part-time until I start in January.” Of course, that it was a full-time job from November on until we actually announced the acquisition and did it. But yeah, honestly, it was really, really exciting. I couldn't complain.

Corey: So, one thing that's been interesting to be is Linux Academy has always been, I guess, the names in the space of, “Learn to Linux,” which again, at one point sounded like an awfully good idea. I should learn to Linux. And then containerization and serverless took off, and now it turns out that most people don't really need to know the intricacies of a given operating system to build a functioning application. This is a good thing. But looking at this now, how aligned are Linux Academy and in A Cloud Guru before the acquisition, and what is the reason other, “Than, hey, we found this thing on sale in the impulse buy aisle.” I don't get the sense that that's the kind of purchase it was. What drove that acquisition?

Katie: So, it's really interesting. It's interesting that you asked that question because that was the problem that Linux Academy had is that there was a perception that their training was Linux focused, which it had been when the company first started, but actually, the majority of Linux Academy’s training is cloud training: AWS, Azure, GCP, just like A Cloud Guru’s, but they had taken a slightly different approach, I think because of their underpinnings and their foundation. They were really, really strong on hands-on labs, cloud playgrounds for those cloud service providers. Where ACG was, I would say more geared towards the novice. So, the—me—new person coming in trying to learn cloud and maybe get their first job there. Linux Academy was much stronger with intermediate and advanced DevOps professionals, but it wasn't actually Linux content focused anymore. And so, it's interesting, as we were talking to the CEO at Linux Academy, he was saying, “We were probably going to rebrand anyway this year because we kept running into that issue.”

Corey: So, we look at Linux Academy, and at the time it was founded, that was absolutely the right name because learning Linux was what people needed. Do you think that there's a future story where A Cloud Guru suffers the same problem because, “Oh, ‘Cloud’. That sounds about as contemporary as ‘mainframe.’” At some point, the terminology changes the way people view computing, and their relationship to it changes. Do you think that this is going to be one of those enduring brands that lasts a century and still has high relevance or do you think at some point, you're going to be looking down the possibility of a rebrand based upon cultural shifts?

Katie: Yeah. I think typically when that happens, two different paths confront a company. One, yes, is that you rebrand. And companies rebrand, especially in the B2B space—less so and B2C, and we're 60 percent B2B, 40 percent B2C, so we have to consider both—but in the B2B space a rebranding is relatively common every five to ten years, or as market dynamics change and as those things happen, so certainly that's an option.

The other thing, though, that can happen is that if you do a really good job, the perception of a company is not so much tied up in the name of it as it is the value that it delivers. But there are a lot of consumer brands that have a very specific name, that now we just use the name even though the name doesn't actually reflect what they—

Corey: The Kleenex or a Coke.

Katie: —yes, exactly. Because there's such a strong identity associated with that name. I'm not saying I know which of those is what is going to happen, or when, or if that's ever the case, but I think as business leaders, you're always looking at that and thinking about that, and the good thing is, you have options.

Corey: One thing that I find a little bit perplexing, that you just mentioned, is that it seems that you have two very different target customers. One is the B2B or business-to-business story, where you're selling training services to companies, and building out a training program with them, and the other is the B2C: the typical consumer, or as I sometimes disparagingly think of some of them, the angry children of Reddit. And it seems like those are two very different constituencies, despite the fact you're teaching the exact same material with the exact same curriculum style to both of those groups. How you approach them and how those businesses look winds up being something very different. I mean, I think originally when I first started talking to the ACG folks years ago, there was no real story that was beyond just selling directly to individuals. That has very clearly changed. How is the company, I guess, differentiating between those two groups, and is there a target future in which one of those constituencies gets left behind?

Katie: Yes, originally A Cloud Guru started really directly to the student, selling individual licenses, courses directly to the student. Over time, what we found was we would find multiple students who worked at the same company, and who were using their own funds to learn, let's say, AWS, or Azure because it would help them at their current job. And from that sprung the idea of the business product. So, the first part to your question is actually that they are two different products.

So, our B2B product and our B2C product are actually different. The courses are the same, but the features, the functionality, and the ability to use that content is different. If let's say, I'm a CTO of a large organization that's trying to understand across the board all of the skills across my team. Where I'm strong, where I'm weak, and where they should focus learning. Those are features and functionality that we have for our business customers that we don't necessarily need for an individual student.

So, that's the first part to your question, which is, it's actually two different products that we're selling for two different audiences. The second, though is, I think at the end of the day, the goal is the same. And I think that's what allows us to have a consistent message and brand in the market, which is whether I am an individual trying to better myself, or I am a CTO, trying to upskill my team, what I am trying to do is improve my modern tech skills for some different purpose. But that's it. I'm trying to improve my modern tech skills.

And so, as we have grown, as we've gone from, as I mentioned, 100 to 400, what we're starting to think about is not only how does the product become specialized, which is already in place, but how do our teams actually become specialized around each of those audiences? How does the messaging become specialized for each audience, but also under this single umbrella of teaching the world to Cloud? And it's actually—it's really exciting to me. There are a lot of companies who've done this really well. Atlassian has done this really well. Twilio has done this really well. There are companies out there that forever maintain a really strong B2C and B2B base. And from our perspective, it's actually the student, the B2C consumer, that is the foundation of any business product that we have because that's the core person, at the core—individual who is benefiting and getting value from the specific content.

Corey: So I want to talk about New Relic. I know you're probably thinking I should talk about other monitoring companies these days that are a bit more in the news. And a month ago, I would have agreed with you, but New Relic did something a little out there. They reworked basically everything. They went open source, they made it so you can monitor your whole stack in one place and they simplified their pricing dramatically. There's even a free tier with one user and a hundred gigabytes per month. Totally free. Check it out at newrelic.com. Observability made simple.

Corey: Enterprise software says a lot, but none of what that says generally is a positive experience. So, coming at it from the student-first perspective, and then finding something that scales upward is, I would argue, at least from my perspective, the right direction. In the interest of full disclosure, I feel like I should point out that at The Duckbill Group we offer ACG subscriptions to all of our staff who want them, to learn how all of these cloud things we're talking about works, all the time. I personally haven't sat through most the lectures because (a) my attention span issues and (b) again, as mentioned, my preferred way of learning is to build something hilariously wrong, authoritatively state that is correct, and then wait for the internet to correct me and teach me what I should do instead. That's not as scalable as one might hope.

Katie: Well, we'll work on that. We'll work on that next. Hey, you know what? Have you tried Linux Academy's platform as well as ACGs?

Corey: Not yet. To be honest, it comes down to the entire model of the constrained environment, which, again, I appreciate, but my approach has always been a… I piece together 15 different blog posts, some of which were last updated in 2008, and at the end of it, I have something that, let's not kid ourselves, monstrous, but it kind of works. And then I show it to people and their response is, “Oh, I'd love to see what you've built. Oh my god, where did you come up with this?” And they come up with a variety of euphemisms for, “Who hurt you?” And that's how I tend to learn best. I don't recommend this to anyone. In fact, I would recommend doing anything other than this approach.

Which does bring me to a somewhat interesting point. I talk to a lot of folks who are eyeing, not just A Cloud Guru, but learning Cloud in general, from a perspective of not so much wanting to learn the underlying fundamentals of how it works, and why it works, and what makes this good, but rather their outcome, and the goal they have in mind is to get certifications from one or more providers. And that's a world that I've never spent much time in. Whenever I see certifications, it's down to people trying to qualify for the next tier of partner status by having enough certified staff or, alternately, people relatively early in their career looking for ways to demonstrate that they know this technology and get a jump on things. But once there's a certain level of baseline knowledge and experience, and a piece of paper that says that you are experienced with these things is less of a certification, and more of a resume showing a bunch of similar projects that you've worked on historically, it feels to me like the value of certifications becomes a little bit oversold. I don't know that I'm right on that at all, but I'm curious to get your thoughts on it.

Katie: I think you're right, in that the value of certifications varies depending on your stage, both your stage in your career, but also the stage of the company. So, in one of the examples you said, let's say—and this is an actual example of a student we had—I was somebody who used to work at a bowling alley, and I realized that I want a new career. I see a path, tech is the place of the future for me to go, but I don't know where to start. And I need to learn, and to position myself to get that first job in the industry. Certifications have a very different and very important value in that scenario.

Now, let's say I work for a Fortune 100 company; I'm a senior architect at a Fortune 100 company, and we are migrating—and don't take any of this personally—from AWS to Azure. Or we're bringing on Azure as one of our providers. And so now I've got to learn a new cloud technology that I didn't know before. In that scenario, the certification isn't really in and of itself what I'm after or what's valuable. It might be a symbol that I've learned this, but when we're selling into enterprises like that, who are trying to skill their teams, certification is actually not the primary thing we're focusing on.

Corey: For me, certifications have always been useful insofar as the actual test: pass, fail, whatever, for me at least, didn't make much difference, but the things you had to learn to make a serious attempt at that certification was where the value laid. The challenge, too, is in some cases, you see people teaching to the test, which means okay, now you're entirely going to succeed or fail based upon how adequate a job that certification itself does of encapsulating the knowledge someone needs.

Katie: Yep, I totally agree.

Corey: Something you just touched on is the idea of expanding into a variety of different cloud providers, which from a training institution like you're doing is absolutely the right move of being able to address whatever it is in the world of cloud someone wants to do, you being the de facto place to go to learn that makes sense. I personally don't have a strong opinion, as far as which cloud provider someone should use. This entire show and most of what I do tends to be more agnostic than people are led to believe. I just have been drawn, in a business sense, to where the expensive billing problem is, and that's historically always been AWS. Now, as that changes that may have changed on my side, too, but right now my approach is and remains, pick a single provider, I don't care which one, and go all in, whether that's you learning something, whether that's a company deciding what to build on top of, and there are some exceptions to that, but that's the general course I tend to take. What are you seeing as you look at both the individual learners expressing interest in multi-cloud—and companies as well going down that path—what is the actual state of what people are interested in learning today.

Katie: So, it's very dependent on the maturity of the organization as it relates to cloud adoption. So, if I am a legacy enterprise software company, I'm trying to get everything from an old on-prem software to the Cloud, I'm typically going with a single provider; I'm going all in. It's a brand new skill set for everybody to learn; we want to make this as low-risk as possible, and we will tend to see those go with a single provider. For larger and more mature organizations, what we're actually seeing is much more of a move to multi-cloud. So, specific cloud technologies have specific strengths in different areas, and as the infrastructures being built out, they're leveraging each provider. And so we're seeing demand from our students go up across all cloud providers, and we're seeing demand, especially in the enterprise space, for multi-cloud training going way up.

Corey: What is driving that? Is that for workloads that are going to be spanning multiple providers? Is it for different divisions or different groups are using different providers? In other words, is this the same people that they want trained on multiple clouds? Or is it different teams that they want trained on different clouds?

Katie: It's a little bit of both. So, in some cases, yes, it's different teams being trained on different clouds. In other cases, what they're finding is, one cloud provider is really exceptional at one thing, and another cloud provider is really good at another thing, whether it's, like, machine learning and AI, versus core infrastructure, and so they're leveraging pieces of the various clouds to build a best of breed solution.

Corey: That tends to be a relatively reasonable approach to take. Different cloud providers do specialize in different things. Some specialize, for example, in giving things terrible names, others specialize in turning things off when you're becoming dependent upon them. But none of these things are immune from sarcasm and stark, but at some point, you have to put the jokes away and actually get down to work.

One last topic I want to talk to you about. What was it like to come in as the president of a company that until now has been founder led—the people who envisioned this, dreamed it up, spent all the sleepless nights building this—now suddenly, you've come in and you are effectively running an organization that, until now, has been this organically grown thing? What does that like from a cultural perspective?

Katie: Well, it was not quite as black and white as that. Sam and Ryan obviously founded the company back in 2016, I think now, but they actually opened their US operations in late 2017 in Austin and brought a COO on at that point in time, to really build out the US operations—Sam is based out in Melbourne, Australia, and Ryan's up in London—and so to be honest, they had really started to think through how to blend their exceptional passion, and understanding, and entrepreneurial spirit that created ACG with bringing in some outside expertise for things that they hadn't done previously. So, I cannot say that I'm the first one who has come in and done that, in any way shape or form. What I love about working with Sam and Ryan, and this was really why I agreed to come on board for—well, there were lots of reasons I came on board. Mainly because the customers love the product. And that's a special place to be, but the other reason was, I think we all really understood what our strengths were. Like, I'm never going to be an instructor. You will never see me teaching somebody Kubernetes. [laughs]. I'm never going to be developing the software itself, but when you can work with a team that are really egoless in every situation and understand where strengths are, that's really amazing, and I could not imagine coming into a company and being able to be as successful as I can be if I didn't have a CEO or founder who had that passion for the business that they had built.

Corey: That's an interesting perspective to take on it, and it's absolutely something that resonates with me. I am much, much, much earlier slash smaller in the scale of running a small business. I went from being just me to taking on a business partner. Now we have a number of staff, but we are nowhere near the scale of bringing in entire, basically, executive teams to run these things. It's one thing that's always been extremely obvious just in the interactions I've had with ACG, was that Sam and Ryan were very clearly running around with their hair on fire at every possible opportunity. I don't think I've ever been to a conference with one of them, where there wasn't a line of 200 people waiting to get a selfie with Ryan and shake his hand. He's the Mr. Rogers of Cloud is the expression I’ve heard.

Katie: Oh my gosh, can I just tell you, I went to re:Invent before I had officially started—Sam said you have to go to re:Invent—

Corey: Oh dear.

Katie: —and just watch the Ryan phenomenon, and I did, and it was insane.

Corey: There's nothing quite like it. He's the Mr. Rogers of Cloud; everyone loves him because he's the person that taught them the thing that got them their career.

Katie: Yeah, which is so special. Yeah.

Corey: Oh, yeah. And it's nice to see them actually taking a step back because it was either that or be dead of a heart attack in five years because there's so much work that was going into this. A while back, I was doing periodic release reviews. Every month there would be a video for ACG of some release Amazon did, and I would, in my typical fashion, make fun of it. And for the first one, Sam flew out here to San Francisco for the day, and we all wound up having a video crew, and we wound up getting it dialed in just right.

And it was, “How do you have time to do this?” And he said, “Well, it's busy, but we make it work.” This was over lunch afterwards, and midway through that question, we were interrupted by someone who had walked up and said, “Excuse me, this is going to sound bizarre. Are you Sam Kroonenburg.” And, yep, here we go. It was the ‘recognized in public problem.’

Katie: And they're so humble and they're so meek, I love it. But you know what? You think about, whether it's like your third-grade teacher or somebody likened it to your Peloton trainer that you watch every day. You become so connected, emotionally connected to these people that you know through a screen, who honestly change your life, or help you reach some goal, whatever that goal is, and I think that's what, honestly, makes this company so special.

Corey: I think that's probably something that I would not question in the least. Normally whenever someone says, “And that's what makes this company special.” My immediate response is, “Well, let's go ahead and snark on it.” But there's so little to snark about with respect to ACG. It's a positive mission; I’ve never met an unhappy ACG customer, and I don't disagree with anything I've ever seen come out of you folks.

It's absolutely aligned with something the world needs, especially right now, in this time of pandemic. And that's one last topic I want to hit before we call this an episode. Obviously, no one wants to turn the pandemic into a marketing story, but what have you seen as far as business changes to ACG since suddenly no one's allowed to go outside anymore, we're not allowed to hire video crews to come into these places, so surprise, everyone's their own one-person video crew? What have you seen on the production side? And what have you seen from the customer demand side?

Katie: Well, first, let me just say that I think as a company ACG, we're dealing with all the same stuff that every company is dealing with. There's so many things that are just literally not in any of our control right now. We are all trying to find balance in a very off-balanced world where we don't see our co-workers the same way we did, we don't see our family, we don't have the same outlets that we have, so that is a struggle for everyone. I think where we have tried to focus is there are a lot of people who right now, for terrible reasons, are in a position that they either have more time because they either don't have a job or they have more time because they aren't commuting into their job, and we can offer something to them in that time period that will help them be positioned for an even better opportunity when we all get through this.

And I was talking to my sales team last week because, honestly, we all have this discussion. It's a really weird time to try to sell. It's a really weird time to try to tell somebody to spend any money at all. And so what we've really tried to focus on is how do we have empathy for what our students are going through, and what the businesses are going through? And how do we simply provide value in an even more important way now?

So, you asked: “what did we see?” I will say right when everyone had to start working from home, to be honest, we saw a big spike in demand because one either had the time or in the case of businesses, they couldn't do in-person trainings anymore at all. And so they needed to get some sort of online training for their students. What that is going to be down the road, and what the ripple impacts of all of this is, I think none of us really know. And so what we're trying to do is just to make sure that for our customers, and for our students, we're continuing to be there, we're continuing to give them an education that is going to help them when we all get through this.

Corey: These are scary times, and having ways to upskill remains important. So, if people want to learn more about what you're up to, what you folks are doing, or basically hear your thoughts on a variety of things, where can they find you?

Katie: Sure. The website for the company is acloud.guru. If you put in acloudguru.com you'll also get to the same place, but acloud.guru is the best place to get information. We've got a really active Twitter account, LinkedIn account, and Instagram account. So, encourage everybody to check us out on social media.

Corey: I did not know you had an Instagram account.

Katie: Yes, we do.

Corey: That one's definitely new, and I will put links to all of those in the show notes.

Katie: Sounds great.

Corey: Thank you so much for taking the time to speak with me today. I appreciate it.

Katie: Thank you, Corey. I really appreciate it.

Corey: Katie Ballard, president of A Cloud Guru. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts. If you hated this podcast, please leave a five-star review on Apple Podcasts and a comment telling me which cloud service I should learn next.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Liz Fong-Jones

Liz is a developer advocate, labor and ethics organizer, and Site Reliability Engineer (SRE) with 16+ years of experience. She is an advocate at Honeycomb for the SRE and Observability communities, and previously was an SRE working on products ranging from the Google Cloud Load Balancer to Google Flights.

Links Referenced

  • Company Site: https://www.honeycomb.io/
  • The Duckbill Group: https://www.duckbillgroup.com/
  • Honeycomb Liz: https://www.honeycomb.io/liz
  • Personal site: https://www.lizthegrey.com/
  • Twitter: https://twitter.com/lizthegrey

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Are you better than the average bear with AWS? If you're listening to this podcast, the answer is almost certainly yes. Want to turn those skills into money? If you're US-based and have an AWS certification, sign up as an expert on AWS IQ today, and help customers with their problems., visit snark.cloud/IQ to learn more.

Corey: This episode is sponsored in part by our good friends over at ChaosSearch, which is a fully managed log analytics platform that leverages your S3 buckets as a data store with no further data movement required. If you're looking to either process multiple terabytes into petabyte scale of data a day or a few hundred gigabytes, this is still economical and worth looking into .You don't have to manage Elasticsearch yourself. If your ELK stack is falling over, take a look at using ChaosSearch for Log Analytics. Now, if you do a direct cost comparison, you're going to save 70 to 80 percent on the infrastructure costs, which does not include the actual expense of paying infrastructure people to mess around with running Elasticsearch themselves. You can take it from me or you can take it from many of their happy customers, but visit chaossearch.io today to learn more.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week once again by Liz Fong-Jones. Liz, welcome to the show.

Liz: Thank you for having me on.

Corey: So, you were one of a very small group of people who was on this show before it launched: you generally want to make sure you can have more than two or three conversations before launching what purports to be a weekly podcast, and you were generous enough to put up with my bumbling, dropping the microphone, et cetera, when I had no idea what I'm doing.

Liz: And to add a complication, I was working for a giant multinational cloud provider at the time that closely scrutinized everything I said.

Corey: Yes, which makes it extra fun. These days, however, I'm still fumbling but I don't drop the microphone because now I can afford a microphone stand. Other than that, a few things have changed. You are now no longer at said large company with incredible degrees of scrutiny. You're instead a principal developer advocate at Honeycomb, a company that everyone loves, or should love if they've heard of you, which again, they really should have by now.

Liz: We're pretty polarizing. Some people hate us.

Corey: So, tell me more about that. Who would hate Honeycomb, and on what axis?

Liz: I think that the large APM vendors are very, very scared of us, and therefore they are trying desperately to undercut our messaging. We also have a very colorful CTO who loves to make provocative statements on Twitter that push the envelope, and people tend to give her a lot of hate for that, as well. So, we're kind of in this interesting position of being now this scrappy startup that everyone loves to cheer for or hate, as opposed to the giant cloud company that people mostly hate on.

Corey: So, one thing I want to call out about Honeycomb is that I've worked with you folks in a variety of different ways, all public. One, you folks have sponsored this and other various productions over the last few years that I'm involved with, which is always appreciated. We love our sponsors; that's how that works. Second, you folks have been a public reference client for us at The Duckbill Group. You had some fantastic stuff going on in your AWS account that to be very honest, we learned an awful lot from as we went through it, and it really is nice to have folks who will get up and say, “Yes, in fact, we will talk in-depth about the things that we were able to do.” It's surprisingly challenging to find that in this space.

Liz: Yeah, definitely. Honeycomb has a culture of transparency with our customers, which means that we can talk about things like, “Hey, here are our outages. Here are cost savings.” These are things that we don't necessarily have to play super close to the chest.

Corey: And the last thing, which is not nearly as well known, but I've never for a minute forgotten it, is back when it occurred to me, “Hey, I should write a sarcastic insulting newsletter about AWS every week,” Charity Majors, your founder and CTO, wound up tweeting about it before I launched. And before that, I was on track to have a couple of people who I wasn't a blood relative of, sign up for the newsletter. And after she did it, over 500 people signed up. The first issue went to exactly 550 people. Now it's 20,000, but I have never forgotten that, and that was one of those transformative inflection points. And it didn't take much work on her part. It was a tweet that was it. It's the sort of thing that people will do without don't even really thinking about it. But it mattered a lot to me, and I've been effectively trying to repay that forward for the last couple of years as a result.

Liz: That's so awesome to hear when things that we do have a really positive impact upon people's lives, for sure.

Corey: So, let's get in a little bit to your perspective on, well, for lack of a better term, the cloud wars. You were at Google Cloud for a long time, and now you're in an AWS shop. You were one of the folks early on who left Google based upon, effectively, its labor practices. Is that right?

Liz: Yes, both its labor and its product ethics practices. And unfortunately, I can't say that AWS’s practices are much better in that department: they do the same kind of collaboration with the federal government on defense matters, they definitely do a lot of the supporting indirectly of Palantir, so there's no big winner among the cloud providers for ethics issues. But a matter of how I spend my precious labor versus who I pay for commodities and services, I think those are two very different axes.

Corey: My opinion on this is it's easy for me as someone who owns a small business that has no illusions of becoming a venture-backed monstrosity that takes over the world, and a $500 million exit is a bad exit. That's not the track that I'm on, so I have the luxury of being able to pick and choose who I do business with. So, when people like to rejoinder things I say with, “Oh, so your company's hands are completely clean?” Well, yeah. We don't do unethical things, but we're also, at time of this recording, seven people. So, I understand that I have the luxury of being able to take those positions. I'm not convinced it's possible to become a large, multinational publicly-traded company without, for lack of a better term, committing atrocities along the way.

Liz: This is the joys of existing under late-stage capitalism.

Corey: Indeed, and we hold this up as the one true path forward.

Liz: Yeah, right. Like, you know, you kind of can't throw stones from glass houses, but you can at least try to make those glass houses better. That's how I tend to approach things.

Corey: One thing I find really interesting is looking over time, as people's bios change as they change companies, now the one that you submitted for the show says that you're a developer advocate, labor and ethics organizer, and SRE with 16 years of experience. Labor and ethics organizer is one of those interesting things in a bio, specifically because you normally don't see it. With most companies looking at most folks who are on the job market, having something out there, you may as well put a racist slur into your bio, for as hireable as that's going to be in an awful lot of shops given our current approach. The fact that you're out there with something that is more or less poison from a lot of these big companies’ perspectives, is fascinating, and I really wish we saw more of it. The ethics and labor part, not the racist slur part.

Liz: Yeah, no kidding. I think that part of why I do that is to set an example using my financial privilege. I definitely had the financial privilege of when Google actually valued my labor and ethics organizing, which was five years ago it did, they paid me a lot of money to continue doing it. And therefore, I can choose to say no to any job that I want. I'm also very highly in demand, right? So, that kind of combination of financial privilege, and also having this profile in the community means that I have that capability to be this walking advertisement for, “Yes, I am going to make your workforce more ethical. Yes, I am going to organize your workforce and you have to be okay with that if you're going to hire me.”

Corey: On that note, if it's not too far of a bridge, how are you finding the, I guess, labor organizing effort at Honeycomb, given that you are not that much bigger than The Duckbill Group, last time I checked?

Liz: I think that it's definitely—companies reach an inflection point when they get past 50 employees, where it's no longer everyone closely knows each other; everyone no longer necessarily closely trusts management. At the moment, we are in a position where people feel like they can get issues addressed by talking directly to executives. And I think that part of my goal is to ensure that the right structures exist to support that feedback loop as the company grows. So, that's kind of the overarching lens that I'm looking at it through, with the understanding that we saw what happened with Kickstarter, we saw what happened with companies like Lanetix, we saw what happened to NPM. God bless the hearts of the employees at NPM who tried to unionize. The instant that your board or funders get wind that you are trying to unionize, using the word union, they will have no qualms about crushing your company. And I think that's something that I'm trying to be very careful around.

Corey: It turns out that a lot of what I thought I stood for isn't actually what I stand for, once I was able to refine that, and have a little bit of space. Easy example: when I thought that I was starting my company, my entire position was, well, I can't ever be political on Twitter—

Liz: [laughs].

Corey: —because if I do that, all I'm going to do is potentially alienate half of my potential customer base. And on some level, there is validity to that line of thought. On the other, things that I'm currently political about, for example, kids in cages. That's one of those things where if you're on the other side of that issue, I don't want your business to be very honest with you. I'm not saying that there's this very high moral bar that every company has to wind up passing. I'm not sitting here calculating out what the pay differential is between their CEO and their low paid employees—thought maybe there's an argument that I should be—there's a spectrum there. But at some point, you have to take a stand for something, or you're ultimately willing to do business with absolutely anyone that pays you. And down that path, lay things that are fundamentally monstrous.

Liz: Yeah, I think it's this interesting dilemma where boards certainly want maximum profitability unless you're a B Corp. Executives are torn between their own personal values and what the board is mandating that they do, but employees can exercise their conscience in a way that executives are not necessarily as beholden to. And you think that means that employee organizing is the best bulwark against companies doing an unethical thing because they kind of have a lot more flexibility to push back against the board in a way that executives can't do by themselves.

Corey: One of the most interesting aspects of smaller companies is how accessible everyone has to be out of necessity, for lack of a better term. And when I was at large companies, as one case through acquisition, I took that with me, and when someone who was a SVP or equivalent, said, “Oh, I have an open-door policy,” I took them at their word and I would come in and talk to him about things that were on my mind—

Liz: Ohhh.

Corey: —and they were always surprised that I’d done it. Because it’s, “Well, don't offer if you're not serious.” Now, yes, there's a tremendous amount of privilege, me as a cishet white dude being able to walk into a room and worst case, people think I'm being an aggressive go-getter and not stepping outside of my lane, as much. I definitely had that come back on me, once I did start stepping outside my lane, but I took advantage of that, and it turns out when people say that they don't actually mean it.

Liz: No, they get upset that you went over their head. Or they get upset that you're bugging them about something trivial.

Corey: Yeah, it's psychotic in some ways. So, moving back a little bit to the area of cloud, what's your take these days on AWS and/or GCP, since I'm going to bet that those are both companies and offerings you have opinions on?

Liz: Yes, I feel like I have a reasonably—well, no one has a good understanding of AWS’s is offering, let's be clear—but I have a competent understanding of both AWS and GCP’s offerings, and, you know, we are an AWS customer; we are going to be an AWS customer for the foreseeable future, but I also continue to follow what's happening over in GCP-land. And I think what's particularly interesting is incumbent effects are really, really serious. I think that, kind of, price is an area that AWS is putting pressure on GCP on, and I think that no amount of technical superiority of the GCP offering is going to be able to compensate for those factors, at least in the near term. So, what do I mean by these things? Well, I think that on the cost dynamic front, what we’re seeing is that the introduction of Gravitron2 on AWS has been a game-changer because it has meant that people can run the same workload at 40 percent off, 50 percent off.

And that was one of the ways in which GCP had previously been able to win things was that, “Hey, Google has more efficient data centers, and therefore has a ability to offer more competitive pricing.” But that, kind of, no longer is something that is possible, as long as people are willing to recompile their applications for ARM. Google has no answer to that. Google only recently has introduced POWER on GCP. Which, you know, “Would you like a $20,000 per hour instance?” “Why, yes, you can have your IBM Power of $20,000 per hour instance.” Like that’s, kind of, been their only foray into non-x86. Google has made forays into, you know, you can run your giant SAP database. Yes, you can run your giant IBM mainframe power architecture workload on GCP. But Google has not done anything that, at least it's publicly visible, on the front of non-x86 architectures for general use, for general compute workloads.

So, that means that unless they catch up, they're going to be at a cost disadvantage for people who care primarily about cost. The other area that I was talking about was incumbency effects where we are a data ingest provider. We ingest a lot of data, which means it's super cheap for us to ingest data because it's free in a lot of cases, although you helped us discover some cases in which it wasn’t, but it means that someone who is on GCP and is trying to send telemetry data to Honeycomb is paying a dramatically increased cost in order to egress that data over the public internet to Amazon. And I think that if we were in the inverse situation of Honeycomb being hosted on GCP, AWS customers would balk at the idea of paying eight cents per gigabyte to egress the data over the public internet article. And I think that that is a significant issue where these walled gardens in the form of eight-cent per gigabyte taxes really, really add up, and mean that you as a service provider, have to locate yourself where a majority of your customers are, or else you're going to wind up running multiple environments or having your customers foot the bill for paying through the nose.

Corey: There's a lot to unpack there. So, first, I think that Gravitron2 is fascinating. When I first heard about it, I thought, “Oh, yeah, that’s not for me. All the stuff I use is x86, why in the world would I want to use a different architecture?” Yeah, it turns out that there were exactly two command-line tools that I used that did not offer a already compiled ARM binary, and it was mostly due to oversights in both cases. One of them was an actual AWS-provided tool. Both of them do now, and from my perspective, there is no functional difference for anything I've ever written or built, and that's neat. So, it becomes a straight cost optimization story, as well as using something that's nifty and from the future. To be blunt it gives me optimism that at the time that we record this, there's been no formal announcement about Apple switching to ARM, but that no longer seems as far away as it once would have, to me.

Liz: Yeah, definitely all of the tech rumors are saying that it's going to happen. It will be a surprise if it doesn't happen.

Corey: At this point. Yes. Bloomberg does not generally make things up, unless they're talking about the big hack story a couple of years ago.

Liz: [laughs].

Corey: For those who weren't paying attention, that was that Amazon, Apple, and a couple of other companies had been compromised by microscopic implants in their mainboard. The sort of stuff that your racist uncle will send you on Facebook. And it was never substantiated; every named source repudiated what was said, and Bloomberg has never retracted the story. So, it’s, “Cool, I like to go on the internet and write fanfiction, too, sometimes, but I usually do it on Twitter, not Bloomberg, the cover story.” It was one of the most bizarre things, and it really damaged my impression of Bloomberg as a journalistic outlet at the time and still has echoes today.

Liz: Yeah. So, we were talking about how ARM is a game-changer, and it definitely is. We definitely were some of the earliest adopters of Gravitron2. We’re definitely cited in the press release, and it's been super, super exciting for us as a startup that now cares very, very much about our runway, right, series A in the middle of a pandemic recession. These things really, really add up and matter.

Corey: One other thing that we talked about a second ago was data transfer and it's ridiculous edge cases. One of my favorite things that has happened this year was both Zoom and Eight by Eight have done large public deals with Oracle Cloud. First, Oracle Cloud’s public data transfer pricing is one cent per gigabyte or less, depending. Whereas AWS’s starts at nine cents, so there's a massive cost differential there. Now that's fine. First, would I trust Oracle Cloud for something? Well, I don't know, but an 8 to 10x cheaper data transfer story, well, for some workloads, you have my attention, and I'm willing to do an experiment and see.

Liz: Yeah, for some workloads, the data transfer is the majority of the cost. It's not the compute.

Corey: For me, at least the funny part, though, was watching Amazon lose its collective mind when this happened. It's strange because I don't get the distinct impression, so far, and maybe I'm wrong on this, that Oracle is a serious, large-scale competitive threat to AWS, but for some reason, with that particular company, Amazon loses its mind. Andy Jassy—the AWS CEO—always takes a swipe at Oracle during his keynote talks. It feels like it's a also-ran company in the modern era. I do not understand this level of pathological hatred for Oracle. Now I get it from Oracle customers, don't get me wrong, but by the same token, airing that in public seems like a very odd choice. And whenever I see AWS get that animated about something, I start paying attention. So, I started digging into Oracle Cloud, and please don't tear me to pieces for this, Internet, but it's not all terrible.

Liz: I personally, because of my distaste for Oracle's business tactics, I've never tried Oracle Cloud. But I think in the broader networking space, definitely my experience from having worked in the past with Google Cloud’s networking, say what you will, it’s had some prominent ingest layer outages in the past couple of years, but overall, Google Cloud’s network, from a latency perspective, is significantly better than Amazon's Ingress, than Oracle's Ingress. I will take Google Cloud’s user-facing network over any other cloud. There's a reason that that is a premium feature. And then they offered, hey, by the way, if you want this same shitty level of reliability you get out of every cloud provider, we’ll discount it by, like, half. So, I think that that's this interesting thing where, if you care about quality of service, you're going to pay the same price for egress that you would at Amazon, and you'll get a much better latency experience. Or you could pay less, and get the same experience as on other clouds.

Corey: This episode is brought to you by Trend Micro Cloud One™. A security services platform for organizations building in the Cloud. I know you're thinking that that's a mouthful because it is, but what's easier to say? I'm glad we have Trend Micro Cloud One™ a security services platform for organizations building in the Cloud, or, “Hey, bad news. It's going to be a few more weeks. I kind of forgot about that security thing.” I thought so. Trend Micro Cloud One™ is an automated, flexible all-in-one solution that protects your workflows and containers with cloud-native security. Identify and resolve security issues earlier in the pipeline, and access your cloud environments sooner, with full visibility, so you can get back to what you do best, which is generally building great applications. Discover Trend Micro Cloud One™ a security services platform for organizations building in the Cloud. Whew. At trendmicro.com/screaming.

Corey: With AWS, it's, “We have one network, it's awesome, and you will pay a premium for it.” Well, great; I have workloads where I don't necessarily care how reliable it is, or even necessarily how fast it is. I want to get this data from here to over there, and I don't care if it takes you a month to do it. It just needs to get there at some point. Don't charge me through the nose for it. And unlike almost everything else in AWS, there is no dial on that.

Liz: Yeah, I think it's this dilemma of some things are latency-sensitive user-facing interactive service workloads, and others are not. Making that kind of prioritization is an important thing that people are realizing they have the need for, that I think GCP was ahead of the game there. Although GCP originally released the premium only and then went to premium and standard.

Corey: There's a lot to admire technologically about GCP.

Liz: There's a ton to admire technologically, it is built so solidly. I really, really respect the engineers who built it.

Corey: I feel like they're being let down on some level by the business leaders, to be very direct with you. Every time I talk to a customer who's eyeing GCP, they raise the same concern: “What if Google gets bored, or changes the pricing model in a way that dramatically impacts our business?” And every time I bring this up to people at GCP, generally, the response is usually something delivered in an accent heavy with condescension. And it's, oh, I clearly don't understand the grand vision. And maybe; maybe I don't, but I'm passing this on from customers I talk to who likewise don't understand it, and yelling at people does not turn them into customers. Believe me, I've tried it.

Liz: I think that from the strategic perspective, Google has invested so much into GCP that it's a “can't fail” project. I think that it is what Google's leaders are looking to as its next several billion-dollar business line, you know, separate from ads. So, I'm less worried about the company deciding it’s bored and it going away. When I am worried about is speed of execution. As of when I left Google, there were several hundred people working on Stackdriver.

And it's like, Honeycomb is running circles around some of this with a dozen engineers, and its own dedicated sales and marketing team. Anything that this kind of velocity penalty of having everything be big company, whereas you can trust it with—Amazon, sometimes left hand doesn't talk to right hand, but they're at least a little bit more agile about launching stuff, and I feel like Google's kind of mired a lot in bureaucratic shit these days and kind of overbuilding these teams, and I worry about the squandering of resources, there.

Corey: It's one of the strangest things I've noticed about Google, is it's difficult to understand what it is that drives them in, I guess, anything beyond the, “This is technically fascinating,” or, “We found a new way to show ads that people don't want to see to them.” There's a lot there. I joke, conversely, that AWS has a product strategy consisting of a post-it note that says, “yes” on it. I will say that AWS is absolutely leaving Google in the dirt with respect to excelling at suing former employees under the auspices of non-compete agreements. It's clear that if Google wants to be serious about being a world-class cloud provider, they've got to start suing a lot more people who got them where they are under incredibly onerous, overbroad agreements, especially to keep their existing employees in line.

Liz: Yeah, in terms of keeping existing employees in line, like, going back to the first subject that we talked about. What Diane Greene saw was the potential to have a multi-billion dollar business doing business with the US Department of Defense. And employees said, “Hell, no, we're not doing that.” Google did retaliate against employees, but Google also did show Diane Greene in the door, in the end. She's no longer an employee of Google. Thomas Kurian is leading the organization now.

But I think the problem was that all those eggs were in that one basket. That Diane Greene put all those eggs into that one basket—or I guess, into two baskets: defense business, and online e-retailers who don't want to pad Amazon's margins. Those are the two big investments that Diane Greene made as leader of GCP. And I think that Google now is having to pursue, “Okay, what is a plan for revenue look like, that is a little bit more diversified?”

Corey: That's really a sad thing to see on some level. It feels like, on the one hand, everyone's competing to see who can come up with more ridiculous things that don't really seem to move the needle. Serverless space is ripe with this. Kubernetes, eh, kind of. But if I look at—

Liz: Kubernetes, I think, was a strategic coup. It was successful, right? It forced people to cloud-neutral their workloads. Sure, you might still be tied to RDS. But—

Corey: I’m talking the ecosystem. Everything tied into Kubernetes, here so you can run Kubernetes on your Raspberry Pi. While you're already up, make—

Liz: Oh, goodness. [laughs]. God, the proliferation of Amazon product names, it's hilarious.

Corey: My opinion on Kubernetes on a cynical level is that it was, in some ways, an effort by Google to help the rest of the world write software in ways that were more Google-y.

Liz: In ways that were more Google-y, and that were less Amazon-specific, right? I think it was a successful project in that way.

Corey: You're not wrong. And a lot of the tenets that it gets at are extraordinarily valuable for companies to consider. My concern is it's still finicky, it's difficult, it’s an awful lot of painful things that need to be aligned just so, and it doesn't make it a fit for every workload, but just serverless one of its biggest challenges is its entire, I guess, hype-monster that lives behind it, pushing it for, whatever your problem is that I'm not listening to, here's the technology that will solve it for you.

Liz: And indeed, right? Like, Honeycomb is not a production Kubernetes user. That's why because we focus on utilizing whole machines, and running on whole machines because that's how our workload is. It's a very homogenous workload. And I think that the things that people under-discuss about Kubernetes is, when is it right? When is it wrong? And where Kubernetes is right is where you're trying to bin pack a heterogeneous workload. And there are a lot of organizations that do not fit that mold, and I think that those organizations are trying to adopt it just because it's flashy, and they're seeing mixed results.

Corey: One thing that I find neat about Honeycomb is on some level, you could be accused of some of the same things, whereas you're fundamentally trying to get people to change how they think about workloads in production. Now, it doesn't have the same force of a foundation that can more or less vote people off the island if it doesn't want to, behind it. But it's definitely gaining mindshare among people who are generally known for having made a career out of making good decisions. Tell me a little bit more about that.

Liz: Yeah, I think the discussions about test in production, and discussions of push on Fridays and discussions of instrument your code rather than monitoring what the CPU usage is on one machine, we're moving towards our opinions being a lot more mainstream rather than heretical, so I think that that's been really, really exciting to see, but it's gotten to the point where the large incumbent vendors are now kind of bandwagoning, and saying, “Oh, we do that too,” by bolting products together, by remarketing things, by trying to redefine observability. I think that that’s… imitation is a form of flattery, and I have to remind myself of that every single day. But I think the way that you get there is by demonstrating to people so they see it with their own eyes, with their own hands-on keyboards, with their hands on their braille readers. Is that observability is a better strategy that gives meaningfully better business results? If you don't have the results to show for its hype, I think that getting out of that hype-cycle into here are the concrete results that people are seeing, I think that's what's actually moving the needle is not just us talking about it, but people talking about their own experiences using Honeycomb; with their own experiences using distributed tracing, using wide events using ask-any-question-based observability tooling, even if it's not Honeycomb’s.

Corey: At the time that we're recording this, I spent some time last night doing a Twitter thread, as happens when I get bored, and I've scrolled to the end of the internet. It’s, “Okay, give me an AWS service and I'll explain it sensibly.” Someone said X-Ray, and it's someone who saw what Honeycomb was doing and wanted to do the same thing as an Amazon product. Unfortunately, their understanding was limited to yelling at people to do things differently, and not understanding the why behind it. And when I look through a lot of the tracing stories that are out there, it feels like there's a conceptual gap between where people are, and understanding the value of it.

I've long said that one of Honeycombs biggest challenges is describing what they do to someone who does not themselves have an SRE background because if you wind up explaining this to someone who is making your coffee because, I don't know, you insisted on paying for your coffee with Bitcoin, and during the 15-minute transaction settlement process, you're now going to yell at them about observability, then it doesn't wind up having any impact because it sounds like easy problems to anyone who's not trying to solve them themselves.

Liz: Yeah, I think the challenges of explaining why the previous approaches didn't work, that's the real challenge is explaining to people why—oh, isn't it just as simple as doing x? The soundbite for why it matters is easy. The soundbite is, “We help your developers waste less time, and help your site be down less.” But the problem is people are used to doing it the old ways and don't understand that it's too costly to do it the old ways, that it's too slow, or that you just flat out can't understand your systems anymore. And I think that that's really been the challenge because people look at these disjointed tools that people have put together to try to solve the problem using their existing solutions, and they're like, “There's no way that you can make this work.” People have been burned by years and years and years of really bad distributed tracing, really bad instrumentation, with high overhead that's hard for your developers to understand, and then really poor analytical tools. And then when you're like, “Hey, if you just use the right backend data store, that's a column store, all these problems reduce to being very simple.” And people are like, “What planet are you from?” I think that's the challenge is explaining to people that this isn't science fiction, this is real.

Corey: One of the hardest parts for me is, I guess, getting people to internalize that the things that we talk about on Twitter do in fact have bearing in reality. And these lessons are all hard-won, but a lot of people will not appreciate the value of a lesson until they experience the pain themselves.

Liz: I know. I really wish people would learn from other people's hard experiences, rather than having to go through that pain themselves. I watch people try to scale up systems based entirely off of metrics, and fail. And it's like, “Can you not learn from the 100 other people who also fell down this rabbit hole? Like, seriously?” People repeating the same mistakes over and over. And some lessons, I guess, have to be learned the hard way. But other people are more receptive, and those are the people that are going to have more success because they listened early.

Corey: One last thing I want to call out because I generally refuse to be on podcasts with you unless I'm able to tell this story because it was such a transformative moment for me. One of the first times we ever spent time together was at SREcon EMEA—that was when I was just learning to live-tweet events in a humorous fashion before that became a standard bit—and someone got up and gave a talk from some giant multinational insurance company, or bank or whatnot, about their DevOps transformation, and I started dunking on what the company was doing, and you called me out for that and said that I was punching down. And sure the usual response is getting called out, at least for me at least, a flash of defensiveness that I've learned to ignore and move past because it's not productive, but then it was, how is that even possible? How could I be dunking, or punching down at a giant company? You were right, it is possible to do that. And when I dunk on companies, I try and do it not from a position of power. There are times when I will be actively insulting: AWS suing former employees, I will go after them with a vengeance for that because honestly, the people who work there deserve better.

Liz: That's what it is. It's about the difference between the corporate entity and the employees. I think that you can definitely wind up accidentally punching down at the employees when you're meaning to punch up at the company. And I think that's where we have to come at things from, is like, when you look, for instance, at Google having this reputation for being condescending. I think a lot about the fact that my colleagues who are in Google DevRel all wanted the best for users. That the engineers on product engineering teams all wanted the best for users. And that overall the thing that we shouldn't be criticizing is not the individual employee’s intentions, it's the system that, for instance, prevents feedback from being heard, or people feeling like it's hopeless to raise the issue of fixing individual bugs.

Corey: It's easy to lose sight of the fact that there are people behind these things. As my platform has broadened and gotten bigger. I have to be more constrained in what I say about new AWS service launches. Making fun of names is usually safe because no one spent 18 months at Amazon naming Systems Manager Session Manager. If they did, maybe they should feel crappy, but that can't be said for building the service. If you put blood, sweat, tears, passion into something and the first thing that happens is some jackass on the internet starts dunking on it, and everyone thinks they’re being hilarious, that doesn't feel good, and I don't want people to read my material and feel bad as a result, with a few very specific exceptions.

Liz: Yeah. Humor; not being a jerk.

Corey: Yeah, it’s, if the joke requires you to crap on someone, maybe the joke’s not that funny. So, where can people find you if they want to hear more about what you have to say about life, the universe, ethics, labor organizing, or anything else that crosses your mind?

Liz: People can find me at honeycomb.io/liz, at lizthegrey.com, or @lizthegrey on Twitter, spelled with an E rather than A, although I should probably register that domain name, too.

Corey: That would not be a terrible idea. As always, thank you so much for taking the time out of your day to speak with me. It's appreciated.

Liz: It's always fun talking to you, Corey.

Corey: It really is, isn't it? Liz Fong-Jones, principal developer advocate at Honeycomb. I am Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts. Whereas if you've hated this episode, please leave a five-star review on Apple Podcasts along with a comment telling me why observability is overrated, and prefer dead reckoning instead.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Forrest Brazeal

Forrest is an enterprise cloud architect, speaker, and community advocate. Currently a senior manager at A Cloud Guru, he spent years designing applications for the cloud at Infor and Trek10. One of the original AWS Serverless Heroes, Forrest was also named one of Jefferson Frank's Top 7 Global AWS Experts in 2019. His first book, "The Read-Aloud Cloud", is coming from Wiley in September 2020.

Links Referenced

  • Book: https://www.amazon.com/Read-Aloud-Cloud-Innocents-Inside/dp/1119677629
  • Cloud Irregular: https://cloudirregular.substack.com/ works. Note that this is the link behind some text that is “forrestBrazeal.com/mailinglist” which was not a working link.
  • Twitter: https://twitter.com/forrestbrazeal
  • LinkedIn: https://www.linkedin.com/in/forrestbrazeal/
  • Personal Website: https://forrestbrazeal.com/
  • Article “Why Central Cloud Teams Fail”: https://info.acloud.guru/resources/why-central-cloud-teams-fail-and-how-to-save-yours

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Are you better than the average bear with AWS? If you're listening to this podcast, the answer is almost certainly yes. Want to turn those skills into money? If you're US-based and have an AWS certification, sign up as an expert on AWS IQ today, and help customers with their problems., visit snark.cloud/IQ to learn more.

Corey: This episode is brought to you by Trend Micro Cloud One™. A security services platform for organizations building in the Cloud. I know you're thinking that that's a mouthful because it is, but what's easier to say? I'm glad we have Trend Micro Cloud One™ a security services platform for organizations building in the Cloud, or, “Hey, bad news. It's going to be a few more weeks. I kind of forgot about that security thing.” I thought so. Trend Micro Cloud One™ is an automated, flexible all-in-one solution that protects your workflows and containers with cloud-native security. Identify and resolve security issues earlier in the pipeline, and access your cloud environments sooner, with full visibility, so you can get back to what you do best, which is generally building great applications. Discover Trend Micro Cloud One™ a security services platform for organizations building in the Cloud. Whew. At trendmicro.com/screaming.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Forrest Brazeal, cloud bard at A Cloud Guru. Forrest, welcome to the show.

Forrest: Well, thanks, Corey. It's always a pleasure to chat with you.

Corey: It really is. I'm delightful. So, you've been a lot of things: an enterprise cloud architect, a speaker, a community advocate, and now you're apparently a senior manager, which means, dear Lord, people work for you. Or with you, technically. But yes.

Forrest: Yeah. So, I came to A Cloud Guru specifically to just help them engage with the community. And I do a lot of things that you've probably seen floating around there. Some of them have pictures attached to them, some of them have words attached to them, some of them may even have music scored along with them. Basically, my goal is to tell the story of the Cloud any way that I can, and that's why I sometimes call myself a cloud bard.

Corey: And some of the songs and other, shall we say, artistic endeavors you come out with are, frankly, astonishingly good. It's wonderful. It's similar to stuff I want to create, except for the fact that I don't have that little thing called talent.

Forrest: Well, Corey, I think that you're underselling yourself there. But one of the things that I've learned in the years that I've been doing this is, you're going to have a lot more fun in your job if you can find ways to pull the things you enjoy right in alongside the things that you are required to do and figure out a way to get paid for the things you enjoy. And so, over the years, I've spent more and more time figuring out how to do things like write, and sing, and draw, and just pull that right into my job, and somehow that seems to be my thing now.

Corey: And it seems to be working out really well for you. The fact that you can play the piano was definitely a recent, shall we say, revelation to many of us. Your love ballad to S3 was absolutely my kind of ridiculous. Good work.

Forrest: Well, thanks for that. You know, I've got a few AWS related songs up my sleeve, so maybe we can collaborate on something down the road.

Corey: Oh, we absolutely should. There are things in the pipeline on this end, as well, that I think will be absolutely appreciated. So, tell me a little bit about your journey. The last time we really spoke in any depth, you were at Trek10, which now that you've left, of course, Trek9. And now you're at A Cloud Guru. Why?

Forrest: Yeah, well, first of all, I don't think it's fair to call them Trek9. I think I was pretty quickly replaced over there by amazing folks like Alex DeBrie, who I think has also been on the show. So, if anything, they've leveled up. They're Trak10 plus plus now. But yeah, so I was at Trak10 for a couple of years.

I was doing a lot of consulting, a lot of enterprise consulting, and so I was going into these large organizations, and I was helping them figure out, essentially, how to successfully adopt the Cloud, trying to lead these cloud initiatives. I learned a lot doing that. It was really fascinating. And what I came to realize is I wanted to find ways to kind of scale that even farther. A Cloud Guru has this really fascinating opportunity where they are sitting on well over a million students who are learners on the platform right now. And a lot of those folks are coming from enterprises, from businesses large and small.

And one component of my role of A Cloud Guru that I really enjoy is this ability to hear from a lot of these customers, understand where they're succeeding or failing with rolling out cloud competency, not just to, like, that little central group of experts, but to this broad team of folks—many of whom have years of technological experience, but it's all very legacy-based—and figuring out how to design and scale cloud training initiatives that will work for those groups.

Corey: Something I've appreciated about you—well, there's not much I appreciate about you, at least that's the public like because you're one of those people where if I'm too flattering to you, you become completely insufferable—but one of the things I do appreciate about you if I can be sincere for a second, has been that you're less aimed at solving the esoteric technical challenges, though you certainly do that than you are at helping people understand and wrap their heads around what's happening in this space in an intelligent way that doesn't assume that someone has just spent four years with getting a computer science degree first.

Forrest: You know, I appreciate you saying that. And I think that the word that we have to keep in mind here is empathy. It's easy for those of us who've been in the Cloud for a while, and I've been building in the Cloud for well, at least six years now. And some folks even longer than that. And it's easy for us to think, “Oh, well, everybody's doing this, now. Cloud has won, and we're in that late adopter phase,” and that's really not true.

I mean, what is it, something like two, three percent of all IT spend is in the Cloud, even today? I was just on a call the other day with someone in the public sector, a large group of people actually in the public sector, and was stopped about halfway through my presentation on the Cloud by someone who said, “I just really don't get this term ‘cloud.’ I, in fact, I have just barely gotten to understand the term ‘server,’ and now I'm really having trouble with this new abstraction.” And this is a person who's super successful, has been a professional for a long time. And they're just now getting to the point of even beginning to engage with what this whole abstraction of the Cloud means.

And so, what I try to do when I'm creating content, when I'm speaking to audiences, is step back and remember, how did I feel back when I first encountered this beautiful, dangerous idea of the Cloud. This meeting of the internet and cheap, compact hardware that I consider to be one of the greatest technological innovations of the last half-century. It's super exciting. There's so much color, and magic, and mystery to what we're looking at here when we build on the Cloud. And I want to help people understand that. I want to get them excited. I want to send them off to build the next generation of great cloud applications.

Corey: I think that might be the transformation that we're seeing, and that people are taking far too lightly, where it's less about needing to have this deep background in technology to deliver outcomes. That is what these cloud services are empowering as they move further up the stack. They're all built on these primitives that we've all come to know and, if not love, at least tolerate. I mean, you love them, you write ballads to them. But I'm not trying to build a service anymore that leverages how to store data in vast quantities on a cloud provider. S3 has nailed that for me. Instead, I want to do something interesting that happens to require that functionality. And the things that move up the stack and leverage that, but mean I don't need to spend three weeks learning what S3 is and how it works, that just sort of abstract that all away from me, are incredibly powerful.

Forrest: That's exactly right. And somebody asked me a question on Twitter the other day, and I wish I could remember who this was, or I would name-drop them here, but the person said, “Hey, Forrest, why is it that I have to still go into the AWS console and choose compute and persistence options separately? Shouldn't we be past that point? Shouldn't I just have something handed to me that packages that all up in one section?” I had about 10 answers pop into my head of why that doesn't make sense, and then I thought about it a minute longer, and I was like, “You know what? There's going to be more and more people asking this question over the next five to ten years as we have more folks that just want to build.” They're not so much interested in okay, well, how do I match the exact compute obstruction to the exact right storage abstraction.

For the 80 percent use case, that really should eventually be something that's packaged up for me, and I can focus just on plugging services together and building on top. That's been the serverless dream. That's been the cloud-native dream. That's the quote-unquote, “low code, no code dream,” and we're going to see more services like that coming out. But I think we're talking about hundreds of millions of developers who are going to be practicing professionally by the end of the next decade, and it's up to us to create experiences, to create best practices, to create guided solutions that are going to get these folks where they need to be with as little fuss as possible.

Corey: To that end, you have just recently released a book. It comes from Wiley publishing. What ridiculous title have you given it?

Forrest: So, the title of the book is The Read Aloud Cloud, subtitled An Innocent's Guide to the Tech Inside. And you say, “Why would you create a tech book with a title that rhymes?” And the answer is be—

Corey: The entire thing is in verse.

Forrest: The entire thing is in verse. You're absolutely correct. Ev—

Corey: I was kidding.

Forrest: No, I'm not kidding. It is a hundred percent written in verse. There are some prose essays at the end of the chapters that, you know, unpack things a little bit. They're called “Word to the Nerd” sections. But that's like a “To the parent” section. You can read that if you want to.

Corey: Well, I've heard verse ideas.

Forrest: You've heard verse ideas, yes. Well, you'll hear a lot more verse ideas if you happen to read this book. So, basically the idea is, it's a book first and foremost that you as a technical professional, a cloud professional, can hand to your non-technical family member, friend, coworker, someone who's always asking you, “Hey, what is the Cloud? Can you explain that to me again?” And they just feel like they can't get their heads around it.

That's not a dumb question. It's not that these technologies are super hard to understand. They're just really abstract. It's easy to understand why a doctor exists, why a lawyer exists. We all have some intuitive understanding of that. It's not so easy to visualize what a cloud architect does. What exactly does that person build? And this book is designed with, yes, verse, and with lots of pictures, lots of cartoons, and goofy images, and insane visual metaphors to help you understand what the Cloud does, what people in the Cloud do, and why this is something that's important for your life, whether you're actually looking at building on the Cloud yourself or not.

Corey: So, to be clear, is this aimed at children? Is it aimed at adults? Who is the target audience for this because when I see a read-aloud book that rhymes, I'll be honest, I buy an awful lot of those lately, but they usually are for my toddler.

Forrest: So, look, people who buy books tend to be folks who are adults, not children. So, in that sense, the book is for adults, it's definitely something that is.

Corey: Good point. Step one: target an audience that has money.

Forrest: You know, that is marketing 101. But it’s certainly something that's age-appropriate; you could read it to a child if you wanted to. My suspicion, certainly based on the folks who've seen it so far, is that it's one of those things that you just kind of want to have for your desk if you're an engineer. It's something that we can participate in, that we can be in on the joke with, as you’re, for example, paging through a chapter that's called “Evolution of the Cloud: A Prehistory,” and is showing you an IBM 401 mainframe fighting a triceratops. It's showing the Cloud is this explosion coming out of a volcano, right as the—

Corey: Yes, the meteor in the sky that's labeled AWS/400. And here we go.

Forrest: There you go. Exactly. I think there's a diagram somewhere in there that shows all the parts of the computer as they relate to the anatomical parts of the dinosaur as it's explaining mainframes. So, it's just very zany. There's a lot of visual humor in there. And I think that engineers will get a kick out of it. But it's definitely something that you can hand to either a child or to a non-technical adult—hey, maybe even your CEO; they like pictures. And you’ll walk away from it with some idea, some mental map for what exactly the Cloud is.

Corey: So, it's fun and relatively straightforward for me to explain concepts of the Cloud on Twitter in pithy short statements. For example, someone said, “What is the Cloud? Explain it using small words.” And my answer was, “Oh, cloud means that you used to run programs on computers. Now you run them on money.”

And that's fun, and it's great, and it's pithy, but it doesn't actually improve any understanding. And in fact, it does, in some cases, lead people in the exact opposite direction in service of a joke. What I like about your approach has been that you've never gone in that direction. The jokes may suffer for it, in my own personal opinion—but I'm reasonably certain that's envy speaking—whereas you go the extra mile to make sure that regardless of the joke, it's in service of education. Your priorities are reversed from mine, and I think that is admirable.

Forrest: Well, I’ve learned a lot from you, Corey, and I know that a lot of folks have, and I think we all have different approaches to—

Corey: I'm a terrific bad example.

Forrest: Well, I think at one point, you referred to me as ‘Safe For Work Corey Quinn’ and I feel like ‘Safe For Work Corey Quinn’ wouldn't have much value. So, you're your own thing. But what I'm trying to do is to create something that sticks with people. And you add a little bit of that magic, a little bit of that, as we sometimes call it‘the sidecar of delight’ to what you're creating. And that brings people back for more; it gets them nerding out about it.

I mean, about a year ago, I had created a rap battle between two versions of myself around serverless and containers, and it was ridiculous and goofy and silly. But it really was designed for a serious purpose, which was, without a large amount of snark, without a large amount of disdain for either side of that ongoing battle, just to help people understand what are the trade-offs here? Where might I choose to use one of these technologies over the other? And then ultimately, aren't they both in service of the same goal? Isn't it both about crawling up the stack and trying to minimize undifferentiated heavy lifting, as our friend Werner Vogels would say? So, that's the reason behind why I create these things. If I fail on that, and certainly I do, I appreciate it when people let me know, so I want to create things that help you understand and then meet you where you are.

Corey: And I say an awful lot of snarky things, but I'm quite sincere when I say that there is something incredibly admirable about that. My approach has always been that—my theory of adult education, and I learned to give talks by being a corporate trainer for a very interesting product at the time, was first you've got to get people's attention if you want to be able to teach them anything. So, my approach was always on getting people's attention and then leaving the actual heavy lifting of teaching people things to, you know, others who care about things and are good at them. You have sort of condensed those two into the same step because the jokes that you do are not the setup for the education, the jokes that you do tend to contain the education. And I think that that's a key distinction that really changes the entire format of the message.

Forrest: Maybe that's the case. And not to get too far down a rabbit hole here, but one of the things that always frustrates me about certain folks who do comedy today—and I'm thinking about folks like John Oliver at Last Week Tonight, is there's a serious message that they have, and then they try to sort of flavor that with jokes, but the two things are always discrete, and it feels like they're using a spoonful of sugar to make the medicine go down, but the sugar and the medicine are two very different things. I want the sugar to be the medicine. Getting back to the idea of the Cloud, I personally find this subject fascinating. I think it's amazing to learn about. I think it's mind-blowing what we can do.

And I want to make that process of learning itself enjoyable. And that's another reason that I came to A Cloud Guru, is that’s something that they've had in their DNA since day one. With their instructors, with their features, and everything about the way that they present their content. I started working with them years ago when I started being involved more heavily with Serverlessconf, and right away, kind of, got that that was a big part of what they were all about. And so, I'm excited to be able to do a little bit more of this kind of work on the clock now.

Corey: It's always nice to find a place to be where the thing that you're doing aligns directly with the business needs. I've struggled in my entire career, where me doing the snarky thing, the conference talks, the brand-building was always viewed as a distraction from the thing I was supposed to be doing. The fun thing is, though, was I was better at the fun things that I enjoy doing than I was at the actual real job I was supposed to be doing. I didn't recognize it for what it was at the time, but what meant was I was in the wrong job.

Forrest: Yes, there you go. And of course, one solution to that might be just to sort of create your own job description and, in fact, an entire company around your exact skills, which is something I admire a lot about you, Corey.

Corey: Oh, absolutely. There's a lot to admire about me. I'm delightful. But one of the tricks I found as I went down this ridiculous path the way that I did, has been that I've been figuring out as I go, the things that work, and at least for me, it was the everything I do now is directly aligned with the things that got me fired from a whole bunch of companies. And I finally figured, all right, there's one of two scenarios that's correct: either every employer, boss, friend, mentor, casual acquaintance, person on the bus, et cetera who told me what my problem was, was right, and this means I'm going to starve to death.

Or I see something other people don’t, and embracing this in a constructive way could turn into something better than the jobs I've had. And I was always going to wonder if I didn't try it. All right, let's give this a shot and see what happens. And here we are four years later, I seem to have come up with an answer that, I guess, solves for that problem. I haven't gone out of business yet. I keep looking every month, and nope, still in business. It's worked, for better or worse, and it turns out that when you own the company, you can't exactly get fired anymore without some serious work.

Forrest: Well, that is true to one extent. Being in consulting, as you know, you're still sort of employed by your clients to some extent, but it seems to have worked out.

Corey: For better or worse, yes. And to be fair, there are different definitions of consulting. Very often people have one or two clients that are 80 percent of their revenue. The goal is to never have a single client that's more than 20 percent of revenue, and I crossed that milestone a while back. So, it now means in order for me to get quote-unquote, “Fired,” there either has to be something systemic like I don't know, a pandemic that hits a little differently, or I wind up instead doing something disastrous myself like, wow, and today, Corey went on a complete racist tirade on Twitter, which is generally not my failure mode.

Forrest: Yeah, it's interesting to have such a large portion of your career tied to how you present yourself publicly to everyone. And that's something that I've struggled with at times because it can feel like it's very all-consuming. And a lot of times people will ask me, “How do you make time to create, and how do you find a balance between that and the rest of your life?” And I've not always been great with that. If anything, it's been a struggle to extricate myself, and have a persona that's different from what's out there publicly. And I'm curious to know how you balance that.

Corey: It's a weird thing. In fact, you're at A Cloud Guru now. One of the founders, Sam Kroonenburg, was out in San Francisco a year and a half or so ago, back when I was doing a few of the release review episodes to basically show me the ropes, and we do the recordings, and it went pretty well. Sam kept cracking up, which is generally a good sign for my type of humor.

And then we go out for lunch. And we get to talking, and we're interrupted by someone, “Excuse me, this is going to sound incredibly rude. Are you Sam Kroonenburg?” And it turned into this whole, “Huge fan of what you do,” et cetera, et cetera. And I realized that if I'm not careful, that turns into me, where I'm not able to go walk down the street in San Francisco without getting stopped and said, “Excuse me, are you that open-mouth jerk from Twitter?” At which point, that's my immediate cue to duck because it's probably someone who works at AWS, and I'm about to get punched in the face.

But it tells me that at some point, there's a Rubicon that gets crossed, and you can't go back from it. Now, this sounds like a complete unrealistic thing to worry about or complain about, but if you talk to actual celebrities, not these crappy Twitter celebrities, of which I am barely one, this becomes an actual problem. When you have paparazzi following you all the time, you can't go out in your sweatpants to get a gallon of milk without having pictures taken and showing up on the internet. You always have to be prepared to deal with the public. Now, I'll never get there, but there's definitely a spectrum between no one knows or cares who you are, and everyone recognizes you on sight and possibly wants to stab you with a pitchfork. I don't know where I'm going to wind up on that, but I am starting to realize that it's sort of a one-way door. You can't go back to obscurity once you start getting recognized.

Forrest: I just want to say I think it's hilarious that you get stopped on the streets of San Francisco. I feel like, only in Silicon Valley would that happen, would you be recognized for your Twitter profile picture. That's all well and good. I mean, I'm not particularly concerned about reaching a point where I have to think twice before I go out to the store for milk.

Corey: In fairness, it does happen more when I'm walking inside of Amazon buildings in Seattle. The spit takes people do, even though they're not drinking coffee, at the time is pretty impressive.

Forrest: Yeah. We're just waiting for an Amazon building to actually be named the Low-Flying Quinnypig, or something like that. I think that's probably your ultimate evolution. But no, I where I was really going with that question was more just it's hard to set a balance between the time that you spend putting yourself out there and the time that you spend recharging. And I think that's something that even if the people that I'm talking to you—you know, you and I are probably never going to have Corey Quinn levels of notoriety, but if we're creating, there's always that challenge where you're trying to find places to just be yourself, places to recharge, and places to not just always be focusing on putting some version of yourself out there to the world. And I haven't cracked that problem yet, but I think that one thing that has helped me is just really trying to move more of that, actually, into my day job. And again, going back to what I do now, that's why I'm grateful to have the ability to do it from nine to five.

Sponsorships can be a lot of fun sometimes. ParkMyCloud asked, “Can we have one of our execs do a video webinar with you?” My response was, “Here’s a better idea. How about I talk to one of your customers instead, so you can pay to make fun of you.” And turns out, I’m super-convincing. So, that’s what’s happening. Join me and ParkMyCloud’s customer, Workfront, on July 23rd for a no-holds-barred discussion about how they’re optimizing AWS costs, and whatever other fights I manage to pick before ParkMyCloud realizes what’s going on and kills the feed. Visit parkmycloud.com/snark to register. That’s parkmycloud.com/snark.

Corey: And I think you're doing a fantastic job of it, too. I've got to be perfectly blunt with you here. There are times I feel a stab of envy. When Forrest posted something on the internet, and it's blowing up again, it’s, “Dammit I've been outshone again,” but you deserve it. And again, this is not a zero-sum game by any stretch of the imagination. What you do is impressive, incredible, and I'm a huge fan of what you do. It's just always interesting to realize, oh, I'm not the only funny person in the world. And I've got to say, in some of the meetings I've been in, I kind of start to believe I am.

Forrest: Well, that's the thing about being in tech is—and I said something about this the other day—in tech, to some extent, code is table stakes. Folks are technical; that's why we're in tech. And having some of these other skills in your stack, like speaking, like writing, like a sense of humor, as surprising as that might sound, it's honestly not that common. And being able to layer some of those other things in, it really makes you stand out.

And so I always encourage people to think about, what can I do that sets me apart? What makes me not just a pair of hands on a keyboard somewhere? What makes me someone who can actually provide unique value? And you may say, “Well, there's nothing.” But there absolutely is. And it may just be that you're still exploring and discovering that about yourself, but what else do you do? What do you know? What are you able to put out there?

For me, I mean, I like to draw pictures, and I like to engage with people and explain concepts in ways that makes sense to me. And that's something that you can do over time. And some of the people that I really enjoy who are really good at this, people like Julie Evans, who does those beautiful, beautiful zines on topics like containers and Linux kernel topics, just absolutely amazing stuff. And she's so sincere about how she does it. That's the thing I love about her: she's not creating in a way that’s saying, “Hey, look at me,” or, “She is all about me.” It's just, “This is something that I really love, and I'm excited about it, and I want you to be excited about it.” That's contagious, that's infectious, and that's what brings people to you and it causes what you create a stick.

Corey: And that's really what it comes down to. People say that, oh, you're sort of riding the lightning, whatever you're on social media, or putting yourself out there like this because one wrong move, and people are going to jump on you and tear you to pieces on it. And I don't find that to be true. Disclaimer: I am a straight, white man in tech. I am absolutely not the typical target for harassment, and I have a very different experience than many other folks do.

But, I have found that when I get something wrong—and yes, I do—and I get called out on it, as happens from time to time, the trick I've learned is that, first—and this is a psychology thing that is challenging for everyone, and I go through it every time—I suppress slash ignore that initial flash of defensive irritation, of oh, but it was a good joke. Someone's complaining about it. And I have to force my way through that and realize, okay. Let me put myself on the other side of this for a minute. Did I just potentially make certain folks feel crappy?

And if I did, I've gotten it wrong in almost every case, unless that other person happens to be the person that named Systems Manager Session Manager, in which case, yeah, you probably should feel a little bad about that. And if I need to resort to making people feel bad in order to make a joke land, then I probably need a better joke. And when you apologize for getting it wrong, sincerely and full-throatedly, only jerks continue to beat you up on that. But the trick, of course, is you have to do better. You can't keep making the same mistake.

Forrest: That's right. And when we conduct so much of our lives online, it's inevitable that we're going to put a foot wrong. I've certainly done it. I mean, I think there's probably a normal range of ways to put your foot wrong without just bringing malice aforethought to it. But I think there needs to be an understanding for us as a community that when someone does that, there's a reasonable window of time for them to apologize and say, “Okay, I should have thought about that more before I said it. I can see that came across wrong. That's on me. I apologize.”

And that's normal, and we move on. And it always makes me a little bit sad when I see the Twitter mob doing the two minutes hate on someone who they've decided is “It” for that day. Nobody ever wants to be “It” on a given day on Twitter. I try to stay out of those piles wherever I can, just because I don't think it's very constructive, but especially when that person has said, “Okay, I get it. I was wrong. You've educated me.” Then that's the time to celebrate that someone learned.

Corey: It really is. And I think that people forget that at their own peril. It's a common problem, and people also I found, tend to see some of the things I do, and completely misinterpret all of the thought and work that goes into it, and figure oh, Corey is just being a jerk on Twitter. I too can be funny by being a jerk to people on Twitter, and it goes the exact wrong direction that a human being would want to see it go.

Forrest: Exactly right. And I've actually said this to people before. You're a unicorn, Corey, in the sense that you can calibrate this snark in a way that it lands in the right way, and it's directed toward the right folks, and it actually ends up with a constructive outcome. That is not something to try at home. That's something to leave to the experts.

You're going to have much better success in terms of joining a community and contributing, and having that public presence that you may want to have if you focus on being positive, and finding the things that you love that you can celebrate. a feature, or a tool, or a service that you've used, that is just making you really excited, has made your life easier in some way? Go ahead and write a blog post about that. Do a tweet thread about it or something; share it on social and tag the creators. They're going to be so happy that someone saw what they did and enjoyed it. They're going to share it, they're going to welcome you into their community.

This is how I got started in the serverless community back when that was very nascent, is just by writing about what I was doing, and really celebrating services like Step Functions when that came out. I was so excited; it solved a huge problem for me. And that led to Serverlessconf and meeting a ton of great people, some of whom are now my coworkers in A Cloud Guru. So, it really can happen if you are looking for something to be your shtick, to be your thing, you know, being Corey Quinn is not for everyone. It may not be for anyone other than Corey Quinn, it's a better idea to just become a person who is genuinely excited about the things that you like, and that will pull people in.

Corey: And that's something that I found is often overlooked. If people are trying to get out there to be the next Forrest Brazeal, which I've got to say I attempt to do from time to time. Like, “I could learn the piano and figure out a way to sing.” Not nearly as well, and work on growing a beard, because I still look like an angry 14-year-old trying to prove a point to Mommy and Daddy in the form of rebellion that isn't going super well. But if I succeed all of those things, I'm only ever going to be the second-best version of you. I've got to find the thing that works for me instead. And that's something that I think people also forget. They pattern themselves after a specific person, or group of people and try and become them. But you have to walk your own path.

Forrest: Well, that's exactly right. And I mean, that could get us on a whole rabbit-hole about the self-help industry, and folks hold themselves up as, hey, this is the one true way, and in fact, this is something that worked for them. You've got to look at your background, and your environment, and who you are, and where you came from, and find a way to make that work for you.

Corey: So, one last topic before we call it a week, I guess, is you had a blog post somewhat recently called ‘Why "Central Cloud Teams" Fail.’ What drove that, and what are you seeing? Because I have angry opinions on it, but you probably have data.

Forrest: Yeah. So, this is something that I've been struggling with for a number of years, back to when I was working in-house on central cloud teams, and if you've been on one of these, you know what I'm talking about. These are the teams who usually give themselves a cute name like the Cumulonimbus Team or something like that. I mean, they're often established right at the beginning of an organization's cloud transformation. So, these are the folks, they're usually self-professed experts, they're folks who are, you know, the quote-unquote, “10X engineers,” whatever that means.

They're the people who have a lot of ability to learn on their own, and pick up technologies quickly and so, when the organization decides, okay, we're finally going to do this cloud thing, you take this group of people—it’s usually smaller number; it's going to be less than 10—and you say, “Okay, we're going to form a team of you folks, and you are now going to establish the standards, you're going to establish the best practices, the automation, the tooling, all of our other legacy product teams are going to use to migrate to the Cloud over the next X number of months or years.” And usually, that's something that's kind of an ego trip for the central cloud folks, they say, “Okay, I’m finally going to be able to shape the direction of this company the way I always believed it should be, and I'm going to build on these services,” and they think to themselves, especially, “Well, I don't need any kind of formal training here because I'm really good at googling stuff, and I'm just going to make this happen on my own.” And unfortunately, what happens is you fast forward six, twelve, eighteen months, and that central team is burned out, they're frustrated, they're not having the impact they expected to have, to the point where a lot of them are leaving the company. And the reason for that is that they're getting clobbered with support tickets. I've seen this in-house.

I saw it as a consultant on multiple occasions, and now that I've been talking to a lot of customers at A Cloud Guru, which is a great way to see a really broad spectrum of folks at different stages of their cloud maturity, I've seen it over and over and over again. These are teams that are getting—instead of pushing standards, and tooling out as they were hoping to do, they find that all of the work is coming inbound, and it's folks from these legacy teams who are saying, “Well, what is that S3 thing again? What is EC2 stand for? I was trying to log into the AWS console, but this secret key doesn't seem to be going into my password box, or whatever.” And that's what causes, again, that central team to burn out. Those folks leave, they move on to greener pastures, and then the organization's cloud adoption stalls.

And the light bulb moment that I had around the time that I was coming to A Cloud Guru was—because I'm one of these engineers, who is relatively high functioning, high performing, and has been in the past. And I was one of those folks who thought, “Well, I'll just figure out what I need to know as I go along. I don't need to sit down and do a bunch of training.” And what I realized is—this kind of blew my mind—you need on that central team, training more than anyone. Your experts need training more than anyone else, but it's not for themselves, it's for the rest of the organization. You have to have a way to scale your expertise and turn yourself into a Center of Enablement, where you go out and allow those other legacy engineers to skill up on their own terms. And that's where having an actual formal training program will help. It'll help them, and it will keep you from drowning under a billion JIRA tickets that just say, “What is AWS?”

Corey: And that is a common pattern. The problem that I've had is every time I talk to companies doing the central cloud team thing, they make a few, from my perspective, perceived blunders. First, they call it a Cloud Center of Excellence, which to me implies that oh, that's the Cloud Center of Excellence—and, let’s face it, who doesn't love a team that calls themselves excellent to your face—that implies that there's a counter team somewhere else called the Data Center of Mediocrity. And obviously, it's clear where the top-flight high performers go and where the folks who we’re just trying to replace with shell scripts live. It sets up an incredibly toxic environment due entirely to terrible naming. Some of the aspects that it brings in can be debated, but if you start off with I have this idea and the name of it enrages people, you're really fighting an uphill battle you don't need to fight.

Forrest: Yeah, I think that's true. I like the Data Center of Mediocrity. I tend to call it the Legacy Donut of Dreadfulness, which is what surrounds the Cloud Center of Excellence. And you're constantly trying to take bites out of that legacy donut, but it tends to sort of engulf you. So, some of this sets itself up for failure, and I think that going from the Cloud Center of Excellence model to the Center of Enablement model is something that we've been hearing a lot over the last couple of years as more and more folks at the enterprise strategy level have started to realize that this is not going to work long term.

Where the Center of Enablement thing even can fail is if you are not understanding that you actually have to go out and embed and engage with these individual teams. I mean, enablement doesn't just mean putting up a page in Confluence that says, “Here's how we're doing things now.” That's not going to get adoption all by itself. I don't care if you have executive sponsorship, or people pushing down from the top, or whatever. What it's going to lead to is shadow IT. It's going to lead to people going off and doing their own thing in the shadows because they don't have the relationships, they don't have the incentive to go and engage with you, and all of your fancy-schmancy guidelines and guardrails.

So, you've got to work on building those relationships and taking some of your experts and rotating them around between these teams. I was talking about this with a coworker the other day, and they pointed me to an example from the air war in World War II, where you had the US forces actually taking their best pilots off the frontlines, they’d take their aces out of the planes, and they ship them back to America, and they’d put them into flight school where they were training the next generation of recruits. The , on the other hand, the Germans and the Japanese, they are leaving their aces on the frontlines until they get shot down, and that's the end of the competency of that Air Force. And it's the exact same thing in the Cloud, you want to make sure that you are putting your best people in a position where they can skill up others. And tying it all the way back to the very first thing we talked about with what it means to move into managing as opposed to just being an individual contributor engineer. That's one of the things that's really hard to get your head around is yeah, my primary value-add here is in helping other people skill up. And that can be extremely rewarding, extremely satisfying, but it’s definitely a mindset shift, and it takes time to get good at that. So, I'm still working on that.

Corey: And I think you're actually doing a pretty good job of it. I don't want to come across in any way on this recording as having a problem with anything you're saying, anything you're doing. I love when you send out your emails. In fact, that is what inspired me to start doing a weekly long-form email. My biggest problem with your Cloud Irregular dispatches is that they are, in fact, irregular, and I prefer seeing your comment more frequently rather than less.

Forrest: Yes. Well, you and me both, unfortunately. I do also have a lot of other things that are happening, and I thought it was wiser to commit to something that I could put out on my own schedule rather than tying myself to something that I knew I wouldn't be able to keep up cadence-wise. Corey, I don't have the advantage of having a newsletter be kind of my main thing, so I have to do that on the side. But yeah, I'd love to be able to write more. I'm actually thinking of doing some different things with that newsletter. So, if you're interested in checking it out, forrestbrazeal.com/mailinglist, and you can get the Cloud Irregular.

Corey: Excellent. Sounds good. So, thank you so much for taking the time to speak with me. If people care more about what you have to say, where can they find you? And for folks with media presences, like you, this is always a loaded question. I have 15 platforms, which one do I mention?

Forrest: Right. Obviously, on Twitter and LinkedIn. You can certainly connect with me there. I do have a website, forrestbrazeal.com. You can find my newsletter there. It is called the Cloud Irregular, so it comes out every month or two. And then if you're interested in the book that we mentioned, The Read Aloud Cloud. We’ll hopefully have a link to that here in the show notes. You can check that out. It's available wherever extremely weird books about the Cloud are sold. And yeah, certainly look forward to any comments or feedback you may have about that.

Corey: Excellent. Thank you again, for taking the time to speak with me. It is deeply appreciated, as always. Forrest Brazeal, cloud bard at A Cloud Guru. I am Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts. Whereas if you've hated this podcast, please leave a five-star review on Apple Podcasts, and a comment that must rhyme.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Melanie Cebula

Melanie Cebula is an expert in Cloud Infrastructure, where she is recognized worldwide for explaining radically new ways of thinking about cloud efficiency and usability. She is an international keynote speaker, presenting complex technical topics to a broad range of audiences, both international and domestic. Melanie is a staff engineer at Airbnb, where she has experience building a scalable modern architecture on top of cloud-native technologies.

Besides her expertise in the online world, Melanie spends her time offline on the “sharp end” of rock climbing. An adventure athlete setting new personal records in challenging conditions, she appreciates all aspects of the journey, including the triumph of reaching ever higher destinations.

On and off the wall, Melanie focuses on building reliability into critical systems, and making informed decisions in difficult situations. In her personal time, Melanie hand whisks matcha tea, enjoys costuming and dancing at EDM festivals, and she is a triplet.

Links Referenced:

  • Twitter: https://twitter.com/melaniecebula
  • Melanie Cebula’s website: https://melaniecebula.com/

Transcript
Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Are you better than the average bear with AWS? If you're listening to this podcast, the answer is almost certainly yes. Want to turn those skills into money? If you're US-based and have an AWS certification, sign up as an expert on AWS IQ today, and help customers with their problems., visit snark.cloud/IQ to learn more.

Corey: This episode is brought to you by Trend Micro Cloud One™. A security services platform for organizations building in the Cloud. I know you're thinking that that's a mouthful because it is, but what's easier to say? I'm glad we have Trend Micro Cloud One™ a security services platform for organizations building in the Cloud, or, “Hey, bad news. It's going to be a few more weeks. I kind of forgot about that security thing.” I thought so. Trend Micro Cloud One™ is an automated, flexible all-in-one solution that protects your workflows and containers with cloud-native security. Identify and resolve security issues earlier in the pipeline, and access your cloud nvironments sooner, with full visibility, so you can get back to what you do best, which is generally building great applications. Discover Trend Micro Cloud One™ a security services platform for organizations building in the Cloud. Whew. At trendmicro.com/screaming.

Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Melanie Cebula, staff engineer at Airbnb. Melanie, welcome to the show.

Melanie: Thanks for having me, Corey.

Corey: So, let's start at the very beginning, I guess. What is a staff engineer, and how did you become such a thing?

Melanie: To understand a staff engineer, you probably need to understand that there's levels with engineering. So, a lot of people when they're new to engineering, they just know that there's software engineers or developers. But when you join most large companies, they need a way of tiering people based on experience, and the kind of work they do. So, there's essentially junior engineers, software engineer, senior, staff, principal, and it goes up from there.

And it varies from company to company, but essentially, the kind of work that you do at different levels does change a lot. And so it is helpful to frame them differently. So, staff engineers usually are given less scoped problems; they’re sort of given problem spaces, and they kind of work within that space and provide direction for the company in that space.

Corey: Most companies that aren't suffering from egregious title inflation, which I can state Airbnb is not, it tends to mean an engineer of a sufficient level of experience and depth, where folks are more or less, as you said, trusted to provide not just solutions, but insight and effectively begin to have insight into strategic level concerns more so than just, what's the most elegant way to write this particular code section? Is that a somewhat fair assessment?

Melanie: I think that's fair.

Corey: Okay. So, Airbnb, which we are not here to talk specifically about your environment, but rather, in a general sense, is an interesting company because they tend to be, from the world's perspective, a giant company with a massive web presence, but unlike a lot of other folks that I get to talk to on this show, you're not yourselves a cloud provider. You're not trying to sell any of the infrastructure services that you're using to run your environment. Instead, you're coming at this from a perspective of you're a customer, just like most of us tend to be customers, of one or more cloud providers, with the exception being that you just tend to have bigger numbers to deal with in some senses than the rest of us do. Not me necessarily. I mean, I tend to run Twitter for Pets at an absolute world-spanning scale because I know it's going to take off any day now. But most sensible people don't operate at that level.

Melanie: Yeah, so I think what's so empowering about being a user of all these technologies is you can be really pragmatic about how you use things. You don't have to look at, “Well, this is what Google does, and this is what this other big company does.” Or, “This is what this vendor is pushing this month. This is AWS’s, newest latest technology.” You look at what you need, and what the problems you have are, and what are the solutions out there, and you can actually try them out and find what's the best one, and then you can share that with everyone else. “Hey, for our scale, for these kinds of problems we have, we’ve found this technology works for us.” Or, as the case usually is, “With a lot of work on our end, we've made this technology work well enough.” And I think that's really refreshing. I love talking to other users of technology and coming to what is actually the best solution for your problem because you just don't get that from vendors. They're trying to sell you something.

Corey: Right, it comes down to questioning the motives of people who are having conversations around specific areas where they have things to offer. Apropos of absolutely nothing, what areas of problem do you tend to focus on these days?

Melanie: In the last few years, I’ve worked a lot on our infrastructure platform. And so what makes our infrastructure more usable, more easy to operate, make it more functional for developers? And then the past few months, instead, I've started working on cost efficiency and cloud savings.

Corey: a subject that is near and dear to my heart. When I started my consulting company a few years back, the big question I had was, “Great, what problem can I solve with a set of engineering skills in my background, but I want it to be an expensive business problem? Absolutely nobody wants to see me write code. And oh, yeah because of some horrible environments I’d worked in previously. I refuse to work in anything that requires me to wake up in the middle of the night, so this has to be restricted to business problems.” The AWS bill was really where I landed on when I was putting all those things together. And for better or worse, it seems to have caught on to the point where I fail to go out of business pretty consistently, every single month.

Melanie: Yeah, it's a really hard problem. And I do agree: I've had on call before, and I've worked on different problems, and I think cost is actually a really great problem to work on. And when I worked on cost, too, I was wondering, hey, is this actually a problem worth working on? Because when you think about some of these other problems, there's more sophistication around companies that have reliability orgs or developer tooling orgs. And most companies just don't have that sophistication when it comes to cloud savings, so it is a new and exciting place. And so I'm happy to be working on it, and the problems do not stop. So, I think it's quite interesting.

Corey: I would absolutely agree with you. It was one of those things that when I first got into—which it turns out was surprisingly easy to get into because when you call yourself an expert on something like the AWS bill, no one is going to challenge you on such a thing because who in the world would ever claim such a thing if it weren't true? And I figured it would be a lot drier and less technically interesting than it turned out to be. The more I do this, the more I'm realizing that it is almost entirely an architecture story past a certain point. It's not about the basic arithmetic story of adding up all the bill items and make sure that the numbers agree. That's arithmetic that's not nearly as interesting or, frankly, as challenging of a problem.

The part that's neat to me, is that past a certain point of scale—and that point is not generally in someone's personal test environment—spending significant time and energy on not just reducing the bill, but understanding and allocating portions of the bill to different teams, environments, et cetera, is something that companies begin to turn their attention to. At a certain point of scale, as it's clear that you folks are at, having an engineer or engineers focusing on that problem makes an awful lot of sense. It seems that some folks try to get there a bit too soon by hiring an engineer to do this who cost more than their entire AWS bill. That seems like it might be an early optimization, but I'm not one to judge. So, my question for you that I want to start with is what do you see as being the most interesting thing that you've learned about AWS billing in the last 6 to 12 months?

Melanie: That is a big question. [laughs]. I’d have to say that the most interesting thing that I have learned has been around architecture and architecting for, basically, efficient compute and cost savings. And so things like the way that we send data between services, and configuring retention for that isn't always straightforward. And another big one has been data transfer; I mean, that's been huge.

When people think about availability and being available in multiple zones across multiple data centers around the world, I don't think cost goes into that equation. And what I found in my own experience so far is that it definitely should because to solve that problem, you're looking at building enough provisioning into your compute layer—so having enough so that if one data center goes down, that the other ones can spin up fast enough, and get that compute in time to handle an outage. You're looking at changes in your service mesh to send traffic to different sources within the same availability zone. I mean, the list goes on: Kafka clusters needing to send traffic, setting one up in every AZ. And for me, that's just been one of the most fascinating things is that when I first started working on cost savings, I didn't think that that much architecture work would be involved, and certainly there's a mixture of things: I’ve had to build little tools, little scripts, some automation. I've had to do some of the, oh, let's just get rid of that manually. But it goes from these basic, easy wins to, we need to really rethink this entire piece of infrastructure, and so that's kind of exciting.

Corey: The hard part about data transfer pricing is that it's inscrutable from the outside, and it's not at all intuitive as far as understanding what makes sense from a billing logical perspective. It costs the same, for example, to move data from one availability zone to another, as it does from one region to another for most workloads—there are exceptions to virtually everything that we talk about in this, which is part of what makes this fun. And, general rule of thumb—this isn't quite right, but if you're looking at this as gross estimation perspective, storing data in S3 for one month is roughly the same cost as moving it once between availability zones or between regions. So, if you're passing the same piece of data back and forth four or five times, maybe just store it more than once and stop moving it around, where the processing and reprocessing of data. I mean, you talked about Kafka. There's always the challenge of historically, compression wasn't as great as some of the newer versions that have been—some of the newer pull requests have merged in new forms of compression that tend to offer a better ratio. There's a pull request that was merged in somewhat recently where you can query the local follower.

But you're right, you have to have things like your service mesh understand that you can now route those things differently, and what your replication factor looks like becomes a challenge. And a lot of, at least in my experience, has always become a more strategic question where it's a spectrum that you have to pick a point on that you're going to target between cost efficiency and durability. I mean, things are super cheap if you only ever run one of them compared to running three of them. But if you accidentally fat-finger the wrong S3 bucket, you don't have a company anymore, in some cases. So, aligning business risk and technical risk with something that is cost-efficient is a balancing act. And anyone who tells you that stuff is simple is selling something.

Melanie: Yes. For me, what I found is framing the problems almost along a matrix of, well, we could go all the way on making this as cheap as possible. That's not ever anyone's preference. People want some amount of durability, and availability, and redundancy, and I think that's great. What I've seen in a lot of strong engineering organizations—this is normally a good thing—engineers really want to do the best thing, at least in my experience.

And so they're very optimistic on some of the engineering work they do, and how available they want things to be, and the pricing and some of these things just need to be considered and architected for. So, I don't think you necessarily have to make a choice on being this available and this durable is prohibitively expensive, but the ways that people do it naively can be. And so having to think through the ways to solve the problem I think it's really interesting. And another example of that is EBS versus EC2 costs. We recently discovered if you're trying to run a certain kind of job on these instances that need to have—so storage on them, there are multiple ways to solve this problem, and so what I found is, what we're really looking at is different ways to solve it with different AWS resources. And the pricing can matter on those kinds of things.

Corey: Oh, absolutely. And there are edge cases that cut people to ribbons all the time. Almost every time you see io1 EBS volumes, my default response is, “That's probably not what you mean to be doing.” You can get gp2, which is less expensive, to similar performance profiles up to a certain point, but before you hit that, an awful lot of the instances will wind up having instance throughput limits. And that's the fun part is, no matter what AWS service you look at, by and large, there are going to be interesting ways to optimize once you hit a certain point of scale.

The hard part in some cases is finding an environment that's using a particular service in such a way where you get to spend time doing some of those deep dives. For example, everyone loves to use EC2—or rather, they use it. Whether they love it or not is a subject of some debate—but it turns out that Amazon Chime: maybe there's ways to optimize the bill. We wouldn't know; we've never seen an actual customer. So, finding things that align with everyone, and hitting the big numbers on the bill before working on the smaller ones is generally an approach that I think, for some reason, sails past people because the bill is organized alphabetically, but we're also seeing that folks tend to wind up getting focused on things that are complicated and interesting to solve from an engineering perspective rather than, step one, turn things you're not using off because the cloud is not billing you based upon on what you use so much as what you're forgetting to turn off. But that's not fun or interesting, so instead, we're going to build this custom bot that powers down developer environments out of hours. And that's great, but development in some cases is 3 percent of your spend, and you haven't bought a reserved instance in two years. Maybe fix that one first.

Melanie: It is so interesting that you say that because I do know an engineer who has built a bot to spin down development instances and, actually, I think it was really quite effective because the development instances were so expensive, and a lot of developers were not using them. So, in that case, it was a really great tool. What they didn't know is that a lot of the development instances were actually at one point spinned up as CI jobs. So, someone had an interesting idea where we could run integration tests on these development machines, so they glued everything together and it didn't work that well, and then they forgot to turn them off and the bot didn't account for those. So, over time, what we saw is that the bot was running, but machines weren't going down, and it was because there was so many of this other kind of machine up.

And so it really came back down to is, what are you actually running? Or, what EC2 instances do you have running that you're not using? And even in this case, the idea that there's a lot more development instances than we think are being used, it still didn't get root-caused at the right level. But I definitely have found that a lot of the tooling I've built and a lot of the solutions that I worked on, at least initially, were just low hanging fruit. There are a lot of things that I think companies don't realize they're not using. Like, if they knew they weren't using it, they wouldn't be paying for it, but they just don't know. And I think some costs I've seen around that is S3 buckets not setting lifecycle policies, and you have a lot of data being stored there that is never ever removed, and you're just not accessing it; you're not using it. And another example I've seen is with EBS volumes becoming unattached, and then never being cleaned up. And so that also can cost a bit over time.

Corey: Oh, yeah. Part of the problem, too, is a lot of tooling in this space claims to solve these problems perfectly. The challenge is, is all of them lack context. Hey, that data in that S3 bucket has never been accessed, so we can get rid of it is probably accurate, if you're referring to build logs from four years ago, probably not if you're referring to the backups of the payment database. So, there's always going to be a strange story around what you can figure out programmatically versus what requires in-depth investigation by someone who has the context to see what is happening inside that environment.

I think a lot of the speeding things up and never turning them off in some respects is a culture problem. First, people are never as excited to clean up after themselves as they are to make a mess, but in some companies, this is worse, where back in the days of data centers you wanted a new server? Great, if you have an IT team that's really on the ball, you can get something racked and ready to go and only six short weeks. So, once you've run your experiment, would you ask them to turn it off? Absolutely not. If you have to run it again, it'll be six weeks until they wind up getting you another one, so you keep it around. I've seen some shops where they run idle nonsense, like Folding@home on fleets just to keep utilization up so accounting doesn't bother them. It's really a strange and perverse incentive, but this idea of needing to make sure that people aren't spinning things up unnecessarily can counterintuitively cause more waste than it solves for.

Melanie: The basic idea is that engineers want to hold on to the things they spend up to?

Corey: They want to hold on to things if it's painful to turn them off.

Melanie: Okay, yes.

Corey: Or rather, if it's painful to get it spun back up where, if it takes you three hours of work to get something up and running, once it's up and running, you're going to leave it there because you don't want to go through that process of spinning it up again. Whereas if it's push-button and receive this thing that you were using almost with no visible latency, then people are way more willing to turn things off.

Melanie: Where I have seen hesitancy in turning things off, it generally comes from a state where they know that it's painful to spin up again, and they really don't want to ever do that again, and in cases where people just aren't sure if it's safe; they just don't know. And they're afraid of there being consequences. And so, I think, our old school development environments, I mean, that was another problem with that, is that people didn't want to spin them down even if they hadn't used them in quite a while because of the way those machines were configured. It's not using containers, there's a lot of stateful data, it means that people want to hold on to it because spinning it up again is so painful. And what I’ve found with us moving to a lot of containerized technology and stateless things is that, at least in those cases, spinning things down to the right utilization has been a lot less controversial.

And another strategy has been making it not the developers problem. So, when you don't have autoscaling, and you don't have sophisticated capacity management in place yet, a lot of developers tend to over-provision things because that's how they handle traffic spikes. They just try to make sure they always have enough compute to handle the traffic spike, but that's not necessarily efficient. So, when you make it not their problem and you just have the say, okay, well, this service, it can target 50 percent CPU utilization, let's say. And then we can, behind the abstraction layer, sort of spin things up and down—or in this case, Kubernetes is doing it. You can also use auto scaling groups with EC2 or other solutions. I have found that, in general, trying to get rid of that problem and, sort of, get rid of the attachment is one way to solve it. But you'll always find cases where you have to kind of deal with that, sort of, perverse incentive.

Corey: Right. And you also find that there's this idea as well, that, oh, we're going to build tooling to solve all of this. But it turns out that mistakes on turning things off can show, and the first time in most shops that a cost savings initiative takes production down, you're often not allowed to try to save money anymore because, “Well, we tried that once and it ended badly.” There often need to be better safeguards and people trying to dive into these things with the best of intentions, but not the real-world experience that—or at least the scars that come from real-world experience having tried such things in the past.

Melanie: Yeah. I think if you're an inexperienced shop trying to work on cost savings, be really careful, I guess would be my advice. What you're doing, really, with cost savings, is you're trying to run—like with compute, you're trying to run things more efficiently. I mean, there's services and applications that just have never run this hot before, and they might not perform well under those kinds of circumstances. I can imagine cases where you don't think anyone's using this S3 bucket and you delete the bucket and, well, now you're in that situation.

And so I think when you're looking at cost savings, you are—every operation with cost savings is a risky one. And so for me, taking my reliability background and applying that to this problem has been really helpful. I mean, having run books, having operation plans, having the needed metrics, and introspection to make these changes. So, one of the biggest changes I've done here, at least at Airbnb was there were places where we just didn't know what was being used. There just was no observability.

And Amazon offers products for this, so S3 object analytics and metrics on usage and things like that, just enabling those for the buckets where it made sense has been really helpful. I will say that, like you've said, there's an edge case for everything, so there are certain buckets that have such an extraordinary number of objects that enabling this kind of observability would be very expensive. And that's the other category I've seen, is people not understanding where the bill can become exponential, and so S3 buckets with a lot of objects is one of them.

Corey: Oh, yes. I saw one once that was just shy of 300 billion objects in a single bucket.

Melanie: Yep.

Corey: You try and iterate through those, it'll complete two weeks after the earth crashes into the sun. And their response, when you ask them about that was—what is this? And they had an answer to, they tried to build some custom database style thing, and they said, “This may not have been the best approach.” I'm going to stop you there. It was not. But at some point, things become so big, you can't instrument them using traditional methods and have to start looking at new and creative ways. Things that are super easy when you do a test case on a small handful of resources explodes in fire, ruin, and pain when you get to a point of scale.

Melanie: Yeah, and the other thing I noticed is—so AWS’s billing doesn't necessarily have these safeguards out of the bat. So, one of the first things I would do is implement some safeguards so that you don't shoot yourself in the foot and make it worse. And the other one is to build an understanding of what observability you need, and what you can enable. And so that's been helpful.

This episode is sponsored in part by , fellow worshipers at the altar of turned out [BLEEP] off. ParkMyCloud makes it easy for you to ensure you're using public cloud like the utility it's meant to be. just like water and electricity, You pay for most cloud resources when they're turned on, whether or not you're using them. Just like water and electricity, keep them away from the other computers. Use ParkMyCloud to automatically identify and eliminate wasted cloud spend from idle, oversized, and unnecessary resources. It's easy to use and start reducing your cloud bills. get started for free at parkmycloud.com/screaming.

Sponsorships can be a lot of fun sometimes. ParkMyCloud asked "can we have one of our exacts do a video webinar with you?" My response was "here's a better idea, how about I talk to one of your customers instead, so you can pay me to make fun of you?" And turns out, I'm super convincing. So that's, what's happening.

Join me and ParkMyCloud's Customer Workfront on July 23rd for a no holds barred discussion about how they're optimizing AWS costs and whatever other fights I managed to pick before ParkMyCloud realizes what's going on and kills the feed. Visit parkmycloud.com/snark to register. That's parkmycloud.com/snark.

Corey: Something else that I think is not well understood by folks who are used to much smaller environments is, if I were to check my AWS credentials into GitHub—or GIF-huhb, depending upon pronunciation choice—then I would notice that I had done so pretty much immediately when my $200 a month bill is now $15,000. Past a certain point of scale, even incredibly hilarious spin-ups of all kinds of instances that are being exploited, or misconfigurations that are causing meteoric growth disappear into the low-level background noise because it takes a lot to have even a 10 percent shift in big numbers, versus in my case, if I have a Lambda function get out of hand, I can have a 10 percent shift.

Melanie: Absolutely. And I think that is what makes cost such a snowballing problem. People don't understand why the bill is increasing at this rate, and it's because the more crazy the bill gets, the more things get hidden by the crazy bill, the harder it gets to go after and fix all these things in a way that is systematic and prevents it from happening again. And what I found is we had to build a lot of custom tooling, and so one of the most important ones is not necessarily, show alphabetically what's the most expensive or anything like that, but show the difference. Like, these costs that we have tagged, their delta over the last day or the last three days is this big, and so that's going to be at the top of the list, is actually that delta and pricing change.

Corey: I would like to point out, just for the record, at the moment that in the event that someone else is listening to this thinking, “Oh, I'm going to go build some custom cost tooling myself.” Don't do that. You don't want to do that. You want to go ahead and see if there's something else out there first before you start building your own things. I promise, having fallen down that trap myself, please learn from my mistake.

Something I want to talk to you about that is, well, how do I put this in whatever the opposite of least confrontational way possible is. Okay, so at Airbnb, you run an awful lot of really interesting, well built, very clearly defined awesome technologies, and also Kubernetes. What have you found that makes Kubernetes interesting—if anything—from an AWS billing perspective?

Melanie: From an AWS billing perspective, what I would say is—so when you work with Amazon Web Services, a lot of it, you're working with the different services that they define, and so their billing can show you how you use their resources. When you run your own infrastructure on top of EC2—in this case, we run our own Kubernetes clusters on top of EC2—they don't get the same insight. And so when you're looking at cost, what you get is the cost of your Kubernetes clusters. So, it's not that helpful to know that this Kubernetes cluster got more expensive from day one to the next day. What's helpful is to know which namespaces or which services essentially got more expensive and why. And so having to do that second level attribution, as I call it, is necessary to understand your compute costs. And so, I do think there is a lag from when you use some of the latest technologies, and they're not necessarily AWS services, then you take on a lot of the maintenance of owning that technology and running it, and then also the cost savings for that technology.

Corey: Part of the challenge too, is that folks who are really invested heavily in Kubernetes are trying inherently to solve infrastructure problems, or engineering problems, and that's great. No one is setting out to deploy Kubernetes—I could stop that sentence there, and probably a decent argument—but no one is setting out to deploy Kubernetes from a cost optimization, or more importantly, cost allocation perspective. So, whenever you wind up with a weird billing story on top of Kubernetes, a lot of things weren't done early on, and now there's a bit of a mess because, from the cloud provider’s perspective, you have one application that is running on top of a bunch of EC2 instances, or otherwise. And that application is called Kubernetes, and it is super weird because sometimes it does all kinds of weird data transfer, sometimes it beats the crap out of S3, sometimes it winds up having weird disk access patterns, but figuring out which workload inside of Kubernetes is causing a particular behavior is almost impossible without an awful lot of custom work. Today, I'm not aware of anything generic that works across the board from that perspective. Are you?

Melanie: I am not aware.

Corey: I was hoping you'd have a different answer to that.

Melanie: Well, I can say that because of some standardization and implementation details of how we implemented Kubernetes and how we hosted it on AWS, we were able to come up with a strategy for tagging different namespaces and getting them attributed it to the right services and there for the right service owners, but I think we did have some insight about standardization, and very opinionated usage of Kubernetes. I think if you didn't have that, you would be in a much tougher position. And I also think that's probably why it is tough to find a solution out there so far. And I think it's because you can use these super-pluggable, flexible infrastructure, you can use them a lot of different ways, and so it's hard to build a tooling that kind of just works. I mean, how you define namespaces, I think would be really huge for how you know what is increasing the Kubernetes data transfer costs, or whatever it is.

Corey: A further problem that goes beyond that, too, is every time I've looked at workloads inside of Kubernetes, as we talked about earlier, there's not a lot of zone affinity that is built into this, where if it's going to ask a different microservice—because everything's a microservice because why shouldn't every outage become a murder mystery instead? It reaches out to the thing that's defined with no awareness of the fact that that very well might be someplace super expensive versus free. And again, AWS helps with approximately none of this, because data transfer with AWS is super expensive because bandwidth is a rare and precious thing unless it's bandwidth in the AWS, in which case it's free; put all your data there, please. Have you found that there's any answer to that other than just building more intelligent service discovery on top of it, and then having to shoehorn it into various apps?

Melanie: Well, what I will say is that I think we were probably one of the first companies to truly try to run AZ aware workloads on Kubernetes on AWS. And so we did run into interesting problems and ways of solving this. So, right away, on the orchestration layer, we found bugs in the Kubernetes scheduler that prevented AZ balance from happening. An engineer on one of Airbnb infra teams actually made some changes to the Kubernetes scheduler upstream. Once we had that fixed, it was possible to have pods be balanced across AZ zones, so we could do AZ aware routing in a way that didn't just blow up one zone because it was not balanced.

And then, because we've been working on using Envoy, we—Envoy has quite good AZ aware routing support, and so we've been able to use that to route traffic. But there's this other idea in Kubernetes, about, like, scheduling preferences, so scheduling pods in such a way that the same AZ is preferred, but not required, so that—you want to prefer this because it's much cheaper, and it's a better solution, but if you choose the required option, what you'll get is actually an outage if you have problems with AZ zone. So, that's a little bit in the weeds, but what I have found is we had to solve it at multiple layers.

Corey: And that's sort of the problem because it feels like Kubernetes was sold, perhaps incorrectly, as this idea of you have silos between Dev and Ops, but that's okay because you don't have to have any communication between those groups. That was never true, but this is one of those areas where that historical separation seems like it's coming home to roost a bit. One of the whole arguments behind containerization early on, was that, oh, now you can have developers build their application, they don't have to worry at all about the infrastructure piece, and then they throw it over the wall, more or less, and let operations take it from there. When you have to build things into applications to be aware of zonal affinity, and the infrastructure absolutely has to be aware of that, it feels like by the time that becomes an expensive enough problem to start really addressing, there's already a large enough environment that was built without any close coupling between Dev and Ops to build that type of cost-effective architecture in the base level. So, it has to almost be patched in after the fact.

Melanie: Yeah, so I think Kubernetes actually does bill itself as DevOps empowerment, which is interesting. The idea that you can sort of create your own configuration and apply it yourself. And in practice, I'd say Airbnb has had a really strong DevOps culture historically, so a lot of our engineers are on call for their own services, their own outages, essentially, and so when services do have a configuration problem, like a Kubernetes problem, it is generally that team that is paged. There is a problem, definitely, with the platform that we've built, where if you have a general issue with supporting AZ, that's going to fall to the infrastructure team to solve, really, because at that point it's just so in the weeds that I think a regular application developer would be probably horrified if you asked them to try to solve AZ aware routing in Kubernetes for their service.

Corey: I am the opposite of an application developer and I'm still horrified at the idea. It's one of those complicated problems with no right answer that also becomes a serious problem when you're trying to have this built out for anything that is not just a toy problem on someone's laptop. Oh, wow—because not only you have to solve for this, but you also have to be able to roll this out to something at significant scale. And scale adds its own series of problems that tend to be a treasure into the light for everyone experiencing them for the first time.

Melanie: Yeah, and I think it's interesting because I think infrastructure is in a Renaissance period right now, where people are really excited about all these technologies, but it's just, the whole category is kind of immature. And so there's these growing pains, and I think when you're in my position, you see those growing pains. I think a lot of people are starting to acknowledge that now. And you can still be really excited about all these technologies and possibly be willing to run them in production, but when we use these kinds of technologies, there will be trade-offs, there will be growing pains. When there's these paradigm shifts and how infrastructure is used, they come in waves, and so for example, service mesh, I think, is probably the biggest example. When you are dealing with all these microservices, these other technologies become kind of crucial to running them at a certain scale.

Corey: Yeah, there's really not a great series of stories that apply universally, yet. And I think you're right: Things come in waves, where you wind up with things getting more and more layers of abstraction; the complexity increases to the point where something happens, and that all collapses down on itself into something a human being can understand again, and then it continues to repeat. It's almost a sawtooth graph of complexity measured over decades. I think that Kubernetes is one of those areas now where it's starting to get more complex, and it's worth, when at some point, you look at all the different projects that are associated with it under the CNCF, and you look around for the hidden camera because you're almost positive, you being punked.

Melanie: [laughs]. A lot of these technologies, I'm really excited by all the development of them, but I can also acknowledge that it's far too complex for, not just the average use case, but any use case. No one wants to be running anything this complex. And it's also fair to say that it would be hard to build something that supports all of these use cases and not end up as complex. But every time we go through this development cycle and iteration, we learn things, and we build it better the next time.

And so what I'm actually seeing is a proliferation of opinionated platforms being built. So, Kubernetes is actually one of the older ones at this point, although it's surprising to say that with Borg being its predecessor, and now these other opinionated platforms that are really quite new. I mean, you look at the serverless movement, AWS Lambda, Knative as another iteration on Kubernetes. And so I think what we are seeing is people trying out different opinionated platforms and building tooling around it, and I think we are moving in the right direction, but we'll see these waves of complexity until we get there.

Corey: I think that's probably a very fair assessment. And I wish I could argue with it, but everything old becomes new again, sooner or later.

Melanie: Yeah, and I think what we'll find is in some areas we’re really insightful, in other areas we kind of just went too far, and we’ll course correct over time. And, yeah, I think that's what we're seeing right now is I think people are agreeing that it's quite complex, and it has all these implications all the way down to cost. And yeah, I think there'll be—there already is a lot of development there. I think people are actually hopeful that there'll be one platform that comes out, and everyone just tells them to use the platform and it's great. I don't actually think that's going to happen. I think what we'll get is a few specialized platforms that are really good at what they do. So, I'm excited for that future.

Corey: I am too. I'm looking forward to seeing how it shakes out. Melanie, thank you so much for taking the time to speak with me today. If people want to hear more about what you have to say, where can they find you?

Melanie: So, you can reach out to me at @MelanieCebula on Twitter or on my website.

Corey: Excellent. We'll throw the links to both of those in the [show notes]. Melanie Cebula, staff engineer at Airbnb. I am Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts. If you've hated this podcast, please leave a five-star review on Apple Podcasts, and then leave a comment incorrectly explaining AWS data transfer.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

Links Referenced:

  • Twitter: https://twitter.com/substitute

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Are you better than the average bear with AWS? If you're listening to this podcast, the answer is almost certainly yes. Want to turn those skills into money? If you're US-based and have an AWS certification, sign up as an expert on AWS IQ today, and help customers with their problems., visit snark.cloud/IQ to learn more.

This episode is brought to you by Trend Micro Cloud One, a security services platform for organizations building in the cloud. I know you're thinking that that's a mouthful because it is, but what's easier to say? "I'm glad we have Trend Micro Cloud One, a security services platform for organizations building in the cloud" or "Hey, bad news. It's gonna be a few more weeks. I kind of forgot about that security thing." I thought so.

Trend Micro Cloud One is an automated, flexible, all-in-one solution that protects your workflows and containers with cloud native security. Identify and resolve security issues earlier in the pipeline and access your cloud environment sooner, with full visibility, so you can get back to what you do best, which is generally building great applications.

Discover Trend Micro Cloud One, a security services platform for organizations building in the cloud, whew, at trendmicro.com/screaming.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Conrad Heiney, a principal cloud engineer at Glidewell Dental, but I much prefer his Twitter bio, which simply states, “I would prefer not to.” Conrad, welcome to the show.

Conrad: Glad to be here. Thank you.

Corey: So, tell me a little bit about who you are, and where you came from because my first introduction to you was somewhat unorthodox. I saw you get retweeted a few times by the Pinboard account—use Pinboard—and you had a certain acerbic style that really resonated with how I tend to view the world; intelligent, snarky, sarcastic, but also very clearly clued into current zeitgeist in modern cloud. Who are you, and how’d you get here?

Conrad: Well, I'm not a computer person originally. I was a humanities type. I was raised in a household full of writers and teachers, and I took one computer class in college, and then the second class caused people to die. It was a weed out, so I went to back to English. And computers were a hobby for some time. And I guess around 1989 or 90, I started getting into BBS stuff again. And then I made the mistake of getting Prodigy.

Corey: The web service, not the band, to be clear?

Conrad: Correct. No “Fat of the Land.”

Corey: No twisted “Firestarter” there?

Conrad: None of that, no. And, of course, Prodigy was profiting in a number of ways. One of which had ads. And there was an ad for this new thing called America Online because PC America Online had just launched. So, I was a so-called “charter member.” And they had a trivia club online. And I was always good at that: College Bowl and all that. And I enjoyed that, but AOL at that time, off-peak was four bucks an hour and I didn't have any money. But if you won the trivia game, you got your hours comped. That almost worked.

Corey: Back in those days, that was amazing. I mean, now if you win the trivia game, you get a job offer as a developer.

Conrad: Right. Right. So, the I got a little bit too into it, and ended up being one of those AOL guides, which was a combination of tech support and chat patrol. And at the time, I was a manager in medical records, and I was futzing around with computers, and the two sort of dovetailed, that I had to manage—wait for it—a Novell NetWare 286 install on a 1990s healthcare. And I ended up leaving healthcare, taking a pay cut, and working for a company called Hollywood Online, which was AOL’s only entertainment provider. And so that’s, in ’95, how I ended up in that field. And then everything after that was sort of accidental. It was kind of a backwards introduction into computing.

Corey: That's one of those interesting stories, in that you talk about coming up through computing in an era where a lot of the, well, we'll call the modern generation of developers, who have been doing this for almost three whole years, didn't get to experience in quite the same way. Which on the one hand, is kind of sad because they didn't wind up getting the same deep level of exposure and experience with some of these things, and then the other is good because we generally frown on child abuse. So, there's definitely a strong sense of the world modernizing.

But something that's always bugged me about that entire approach has been that I look at why I'm capable of answering weird questions about cloud computing, and invariably, I go back to fundamentals that I picked up back when I had to actually care about things like, is the SAN currently getting a full disk? Does it need to be partitioned further? What's the temperature look like? Did someone actually pee into it? And now that I don't have to think about those things at a general basis, it's great because there are enough abstractions that have made it to a place where I don't have to. But the counterpoint is, is that I feel like I would be way less effective without that grounding in Unix.

Conrad: Yeah, the progress of abstraction is a thing. It's this train, and where do you get on the train? And so, I knew people when I was a kid who were working on computers when you crawled inside them, and when everything was a big chunk of metal, and later on when things got more abstract, they were very, very good at it. My brother is a scientist, but he learned C because he had to write device drivers in order to make science go. And those people, to me, are sort of godlike because when they end up say writing a software, they know what's going on all the way down into the electrical stuff.

And I've always noticed that the people who were really good at computing were not CS graduates necessarily, because that varies a lot, but they were electrical engineering graduates. Those people really know their stuff. So, yeah, I mean, I joined at a time when, so to speak, the modern internet as a product was being invented, which was great because you didn't have to have a certification. We were inventing it, which was great. But we were still about five steps down from the originators. And as you're right, it's good and bad because when you get all those levels of abstraction, you keep moving up, you can get so much more done. I don't have to worry about anything, but at the same time, eventually you're going to have to go to somebody who knows that stuff, and it'll take you longer if you didn't have that background.

Corey: One of the things I find that's so compelling about vendor selection in terms of managed services, or cloud provider, or whatever phrase of the decade we're using to describe it, is I always tend to bias for the one that I've worked with before. And some would say that, oh, that makes me a stick in the mud. And I'm not willing to embrace progress, but my counter-argument to that has always been that if I don't know how something fails, then I don't trust it in production because, spoiler, it's going to fail. And given my luck, it's going to be at two in the morning on a weekend.

Conrad: Right. You know, one of the things that happens in this field in our technical field is people tend to carry around the other people they want. Like, a new person comes in, and their crew comes with them. Yeah, and I think technologies and products are that way, too. It's like, “Yeah, I know the failure mode of this thing.” As opposed to, “Well, they say it doesn't fail. I wonder what it will be?” So, yes, I agree.

Corey: One of the strange things that always tweaks my notice, is that you see folks who have resumes in the cloud engineering, slash DevOps space, slash SRE space, slash production engineer, slash senior systems administrator—again, we are all systems administrators, some of us have just found the magic incantation that winds up leading to larger salaries. But there really tend to be two clear paths that people tend to go down before they land one of those roles. One of them is the path of the developer: the deep dive into I write code all day, every day, and now I'm learning how operations works because, holy crap, it turns out that when I write nonsense, there are consequences to it. And I love working with folks like that, who have, I guess, seen the light. And then there are folks who generally tend to come from the other side of the world. And that's the one that I tend to come from, of being the systems administrator type, circa 2000s era. The reason I call out the decade, is because back in the 90s, to be a decent systems administrator, you had to know C pretty well; you had to have an in-depth knowledge of system calls and, “you don't write code as a systems administrator,” was not something that anyone sensible would ever say out loud more than once.

Conrad: Right.

Corey: How did you come down this path? Where did you start? And do you see that being a real division? Or do you think that that's a bit of a false dichotomy?

Conrad: I'm very much in the sysadmin end. I mean, I started out as sort of jack of all trades, doing everything from plugging in the modem, to bed web design, to actually doing sort of entertainment-related writing and things like that. And sysadmin was where I gravitated because I found that the human factors end of it was terrible, but the technical part was easy. But I don't like sysadmin culture. I learned around that the job itself was negative, but it was a bad idea to be negative. And that developers had more the right idea of being optimistic. There was a guy I worked with who was old school. Like, he hired the guy who wrote POP3 now [unintelligible] hanging out. And then he gave me this lecture about your entire job is negative, now. Your entire life is negative.

You're just going to be the person who tells people it won't work, and they can't do it that way, and you'll never create anything; you'll never be the hero; you'll never make money; they'll never do the product, and your entire task for the rest of your career is to spin negativity and make people like you. And I thought, “Yeah, he's absolutely right.” And so my strategy after that was to not be the BOFH, right? To be the nice sysadmin, be the one who says, “We can't do that, but I’ll help you get it done.” And I think I learned both from developers, and from my—in the medical field—from people like physicians, that you have to combine that conservatism with an optimistic feeling, and a desire to make things better. I, like you, really appreciate the developers who learn ops—they're rare—because I'm not good at programming, and having somebody to work with who can do both is tremendous. And I think that as sysadmins, we criticize devs too much, but that it's a good idea to take some of that confidence, and positivity, and put it into your work.

Corey: Lots to unpack there. One thing I want to start with, I think, is how do you view this whole world of serverless? I mean, the idea is incredibly compelling. You write code, you throw it over the wall to your provider, and they handle all of the operational bits for you. And then you say, “All right, great, what if it fails? How does it break? What does that mean for my application?” And everyone just kind of stares at you blankly until you stop talking, then they try to pretend that they didn't hear you. I feel like there's something that's getting lost there as far as being able to say, “All right, I'm going to write all of my stuff in Lambda functions, and now it is purely Amazon's problem,” because that's not going to break at all until suddenly, one day it does. And we don't necessarily know how that's going to look. Does that make you uneasy?

Conrad: Yeah, it does, mostly because that way of looking at things which—you go back to containers, to JavaScript technology being used for lots of things, or whatever, all of that whole trend really grows out of startup web development, and a very developer-centric world, where basically everything should move fast and break things. The consequences for breaking things aren't huge. And everything is focused around getting something done and running. That anything else, like planning, or sort of the conservatism of ops stuff is seen as overhead. Kind of the negativity thing that I mentioned before. And the thing about that is it gets stuff done real fast.

Like, if you're a developer and you can run a command line, essentially, take all your code and then it's just running. That's awesome, but the trend basically doesn't just abstract out all of this ops stuff: it denies it. It basically says that's not happening. That's not a problem. If you've got a two-year exit strategy on the web service, it just says, “Yo,” this is awesome. If you've got a manufacturing line, if you've got financial stuff, if you've got stuff with real-world consequences, you have to look at the ops world. And you have to either have a really good vendor or have people who understand what's going on behind the curtain and can help with it. Because otherwise, at that point, it really is sort of a near fraudulent startup dream that you can just basically copy paste a bunch of code, pay someone with the fake money that you got from venture capitalists, and it will just work. And that's a cultural or financial problem that I think has worked its way into technology to an extent where, yes, I love it, you can get stuff done really fast, but you really have to either not care about it being a black box, or hire people who know what's going on in that black box.

Corey: The problem that I see with this is on some level, it's always going to be a black box, and hiring someone who understands what that black box looks like, is impossible at some level of abstraction. No one can hold everything in their head anymore. If you want to try this at home, listeners, all you have to do is wait for the next person to come interview who bills himself as a full stack developer, and then ask them questions about device drivers to see how quickly people's knowledge erodes. No one is good at everything. There is no such thing, in my mind, as a full stack developer. I'm the closest you get because I'm a full Stack Overflow developer in which I copy and paste code that other people write, and wind up doing that until it compiles.

Conrad: Right. You’re right. There's something that one of the things I’m a bore about is, sort of, engineering disasters and their implications for us. And there's a very, very good book called Normal Accidents—Charles Perrow. He defines a situation where if there's a system that one person cannot comprehend and that is also very tightly linked—you know, so that a change over here causes a problem way over there—you will have accidents. His example, unfortunately, his big example is nuclear plants, but you run into this a lot with what we're doing now. But yeah, even something like, oh, we have an app that, let's say, assesses risk for something. Well, the risk modeling person knows that part, and the ops person knows this part, and the dev knows this part, but nobody knows the whole thing, and so it will break.

And I think the nearest thing we can get to dealing with that in that serverless world facility you were talking about, is to educate, first of all, that people actually understand on some level what's happening back there, and then get people with experience who’ve seen things because you develop a heuristic for things after a while. I mean, you do notice, like, this looks like there might be a disk full somewhere, or this looks like network latency of this kind, and I don't know how to get that, aside from A) making sure that devs do understand the basics of this stuff and what they look like, and then paying attention to ops, and getting ops people who are able to penetrate that enough to do diagnostics, rather than just have to know how it all works.

Corey: It's a terrific ideal. The problem is—and this is, I think, the real role of systems administration, however you want to term it, however you want to lock that down and define it, it all comes down to fundamentally the idea that at some point, the buck has to stop somewhere. There's no longer anyone else within your company, to technically escalate issues to, so the answer to almost everything that you wind up dealing, once you've automated the banal away is, “I don't know, but I'll find out,” and that role has always had a bunch of different terms. In my first job where I found myself in that role, we were the “network engineering group,” or “network administration group,” depending on who you asked. And that was great. 90 percent of what I did didn't really touch networks, but rather, had to do with digging into interesting problems and solving weird issues.

And it was pretty apparent that what we found, and how we wound up solving problems was in some ways dangerous because, easy example, someone on the support desk was terrific at following instructions. He'd been there for 20 years, and he was great at things, and I was making a change and I reached out proactively to him—what a concept: collaborating with the support desk because they are your phone firewall—and saying, “Look, for the next 24 hours if anyone has issues with this particular system, here's the command you give them to clear their local DNS cache because I'm going to be updating and repointing this, and it may cause some weird inconsistencies,” because in fairness, I didn't plan ahead and reduce the TTL a week before. Oops. So, the changeover went well, and things are going reasonably decently. And a couple weeks later, I walk past his desk, and I listen to him telling someone to clear their DNS cache for a completely unrelated issue. And we saw some of this where it’s, whenever you learn something new that applies to a specific problem, we often tend to have an incomplete understanding.

It's why when, if you wind up calling for support, assuming that you do that for desktop support these days—I don't know how that works anymore. I throw it away and buy a new one. But the idea was, “Oh, you call in and yes, I've rebooted it. Yes, I've reinstalled it. Okay, I'll defragment my hard disk and then call back.” And there's this litany of, I guess, theater that you have to go through before you can get to the underlying issue in some cases. And because they tend to clear up occasional cases, it just gets added to the flow chart list of things to try before actually diving deep into stuff. I feel like every company has a group that is responsible for coming up with these offbeat, deep dive solutions to esoteric problems, where the buck has to stop with you. Invariably, those people get very well acquainted with things like Stack Overflow, Reddit, IRC, various forums, insulting people on Twitter, etcetera, etcetera because that's, in many cases, where you find the answer to things. It's also why, even if you don't write code yourself, you find yourself reading an awful lot of source code as a step of last resort.

Conrad: Right. There's a level at which there's lore that accumulates with somebody, which is terrible. And then a level at which somebody is… yeah, kind of a detective, and the gap between that role and the support people is huge. I trace it back, essentially, for the bad attitude system, and thing of having contempt for users. I have to say we often see the support people as our buffer against users, and we don't want to hear it, or whatever, and so we give these incomplete instructions to support people as a way of shielding ourselves, rather than a way of helping people out. I mean, ideally, you should have run books for these things, where the support person can intelligently look through steps to do, or things to check, or whatever, that are explained so they don't end up with who you're talking about, like, “Oh, this is a solution thingy.” It's like people who don't understand how medicine works giving injections because injections work, and then it turns out they're injecting dirty water. So, I think that, yeah, there does need to be a person or people with that escalated role of, “Hey, I'm going to go find out what to do about this.” But if you don't take that extra step of then rationalizing that into something that your perfectly capable and not inferior support person can then communicate to your also not inferior user, then you might as well be back in the 90s, manually typing every single thing that needs to be done, and telling everyone else to get away from your desk.

Corey: This episode is sponsored in part by ParkMyCloud, fellow worshipers at the altar of turn-that-[bleep]-off.

ParkMyCloud makes it easy for you to ensure you're using public cloud like the utility it's meant to be. Just like water and electricity, you pay for most cloud resources when they're turned on, whether or not you're using them. Just like water and electricity, keep them away from the other computers.

Use ParkMyCloud to automatically identify and eliminate wasted cloud spend from idle, oversized and unnecessary resources. It's easy to use and start reducing your cloud bills. Get started for free at parkmycloud.com/screaming.

Corey: There's this idea of the “Department of No,” and I've heard it applied at different times to systems administrators, network administrators, security engineers, one very ornery barista, etcetera, and it seems like there's this adversarial historical culture where there's this natural tension that tends to exist between development and operations, whereas development wants to release features more rapidly. Operations, on the other hand, wants to maintain stability, and if things are working right now, well change has the potential of disrupting that. And there's a natural tension, and a requirement to work together that I feel is part of the DevOps movement came out of. Is that something that you see?

Conrad: Yeah, and it's not at all specific to tech. And I think tech people can learn from how this works in other environments. My very first job—I was very lucky I got a professional journalist job and I wasn't even finished with college—and I got to see how a circa 1986 newspaper worked. And so, you had editorial, so we write, and we get photos and we edit things, we produce what we call nowadays, the content; and we had advertising, where they sold the ads, and then they would end up with a spec for an ad and how big it would be; and then you had production, who had to take the editorial stuff, and the ad stuff, and make a newspaper out of it. And we all hated each other. The editorial people didn't have what anybody needed: it was too long, or too short, or angered the advertisers, and it was always late. And the salespeople were always coming in at the last moment with something stupid, or dictating what to be read, or saying, “Please move the coffee cup on this thing. 0.3 inches to the left.” And the production people had to take all of this stuff and make a newspaper out of it, and it never worked, and so they were, so to speak, the BOFH sysadmins of this triad, and it was terrible.

It was a weekly paper, and then we would have these screaming fights, we’d put out the paper, we'd all go out drinking, and then we were friends again the next day because everybody understood that tension was going to happen. It’s not going anywhere; it’s how these things work; it's how the paper gets put out, and how we survive. And you have to manage that tension, and do your best not to be a complete jerk about it, but it's not going away, and the best thing you can do in that situation is to be less of the stereotype you want. So, as a sysadmin, I would, like, not be the “Department of No,” I would be the “Department of Hang On, Let's Get This Thing Right And Get You What You Want.” That I'm not saying, “No.” I'm saying, “Please don't do this by pouring gasoline everywhere and striking the match, but I'm not going to hold you back. I'm going to get done what you want to do.” And I think every job I've had along the way has had some variant of that tension. The difference in tech is that I think that everybody's a little immature and they think they can win: the bosses, the devs, the marketing people, the sysadmins all want to win the argument, and that's not how anything gets done. What you have to do is have those arguments, get things nailed down, and continue to cooperate, while knowing that this tension, in fact, isn't going to go anywhere; nobody's going to win.

Corey: I feel like there's been a forcing function, at least, because you have development and operations. Both are required to work in relatively close proximity in order to get things out the door. Every time there's a release, that is a forced interaction. You see the same thing with the media interaction that you just described. This, I think, is the root of my actual problem with the term FinOps: the idea of finance and operations working hand in hand. There's nothing that forces that tight coupling. One of them is definitionally going to be a leading function, and one is going to be a trailing function, and which one you pick to go in which spot has a lot to dictate what culture you're building. I don't think that there's ever a story where you're going to have those people working in lockstep together on an ongoing close-quarters basis in the same way. I think that the answer to the cultural story of FinOps, if you will call it that, as a direct result is simply distilled down to make sure that you keep different business departments in the loop, and remember that we're all here to work together collaboratively, rather than being massive jerks for the fun of it.

Conrad: Yeah, the question of why are we here, and what are we doing? And when the tech people don't seem to want to answer, you have to—I mean, if you're working for a business, what you're doing there is the business's goals, which are in some level to make money. If you're working for highway patrol, you're there to protect people's safety, and blah, or whatever, and you have to start with that and end with that. And FinOps—I hadn't heard that one—is terrible. I mean, sure, ops, if you want to call it that, has to do with the business does just as everyone from the janitorial staff, to the drivers, to anyone else in the organization does. It's not special. Tech people aren't special. And I've had these ridiculous arguments in the past where someone wanted to do something that was expensive and difficult solely because of some kind of tech purity reason and became enraged when someone asks, “Well, what is this going to do for the business?” So, yeah, you do have to work with—tightly—what the whole organization needs, whether that's finance people, marketing people, the person who sits next to you.

Turning it into a buzzword for a particular thing, I think obscures that essential problem that tech people need to get a little bit out of their gamer chair and realize that what we do all day is for the end of somebody else, that someone else is either paying for it or needs it, and that we're there to provide them something that works, rather than them begging us for something that we have that’s precious. So, yeah, I don't like that concept of FinOps. I just think, “Right, well, everything is BizOps or OrgOps, right?” And you do have to know what your organization is about, and that you're not rowing the opposite way from them, or whatever buzz word you want to put on it. And that is not easy culturally for us.

Corey: That's one of the problems that we see with, I guess, junior consultants, where it's easy as hell to come into any environment and say, “Oh, what have you folks built? Draw it out, or describe it for me.” And at the end, “Well, that's stupid. What complete moron built this?” Invariably you ask that question to said moron, and in virtually every case, there is something that that person is missing, and that thing is context. No one, unless they work at Twitter's verification program, shows up in the morning hoping to do a crappy job at work today. Everyone is trying to get something out the door and achieve a business outcome. There may be constraints you're not aware of. Yeah, I can look at architectures from 20 years ago and, “Well, why don't you use some service that didn't exist back then? Oh.” Or people may not have been aware of an alternate way to do things, or there was a strategic imperative to maintain some level of agnosticism, for example, there are a bunch of reasons that things could be radically different, but that's also with the benefit of hindsight. Remember, however broken and terrible it is, it's gotten them this far to the point where they're hiring said consultant to come on in and take a look at things. There's not a lot of value in just insulting the current state of the art. a child can draw a whiteboard diagram that shows an architecture that works super well. The problem is how you get from the current state to that future state, or something like it, without shutting everything down for 18 months.

Conrad: Yeah, again, this is a cultural problem on some level, again, about tech being special. When I worked in healthcare, I worked with physicians, and they have the worst tech support job on the planet. Awful people show up and lie, and they smell bad, and they've done everything wrong, and you still have to do your job—if you're a nurse, a doctor, anybody in healthcare—and you don't do that by being a jerk, and assuming things, or whatever. Everybody arrives with their story, and you have to deal with it. And with technology, it’s the same. When you say Junior consultant, you walk in, and somebody has this horrible thing you need to fix, well, you're going to go into a room with their friends and cuss about it later, right?, “Oh my god, what are these people are doing? They bolted something onto ColdFusion to run Java, what year is this?”

But you can't do that with them, or even, you can't keep that attitude while you're working at it. You have to say, “Oh, wow. Okay, this is a big challenge. There's something here that needs to be improved.” And we had to find out how much can be improved, and what we can do for these people, and try to suss out what the brokenness is in the organization without saying the organization is broken, we can get something done. Sometimes it means you have to turn something down. So, the essential thing there is, you arrive quite often with negativity in a mess. You know, the times when you have a greenfield thing that you can do yourself from the ground up, and what you think of is right, are very rare. And also, when you look back on what you've done, quite often you think, “Oh, my God, some unfortunate person is probably fixing that right now.” So, I think yeah, the thing is, you can't get rid of that challenge, but just as the way that everybody from nurses, to waiters, has to go in the backroom before they start cussing, about how idiotic that thing is, I think that as tech people, we should do our best to treat people as we would patients with a problem or people who want their fries done extra well, and suck it up to a certain amount, and treat it as something to improve rather than something to be disrespectful about.

Corey: I think that that's the attitude to take with these things. As much fun as it is to be snarky and cynical and, I guess, difficult on the internet, we work with people, and everyone's trying to do the best they can with the tools they have available. And I think that we still struggle as an industry with empathy.

Conrad: Yeah, it is. And I mean, a lot of it is just unearned arrogance from the past. I mean, some of the people who originally came up with this stuff had to deal with, really, weaponized stupidity, and people who didn't understand anything at all, and didn't want to, and so forth. It is a somewhat different world now that your, so to speak, your customers do have some idea of what's going on, or whatever. And also, you're not all that. I mean, we're not the people who invented the internet. We're more like our customers than we are like Eric Allman, or whatever, you know? So, I think, yeah, that empathy and a certain amount of humility is a very important thing and attitude. There's a thing that I call the fish food syndrome. So, let's say you have some friends, and they have a goldfish, and they don't know what the hell they're doing. They're feeding the goldfish like hotdogs, and oatmeal, and gasoline, and dumping stuff in the bowl, and they kill four sets of fish. And they don't know what they're doing. Well, what you really want to say, “You are ridiculous. Goldfish get fish food, and not that much of it, and you have to clean the bowl. What are you even doing? You’re a fish murder.” It won't work. They'll get mad at you. They won’t be your friend anymore, and they'll keep killing their fish. You come back to them and say, “You know what? I heard about this awesome thing. There’s special food for fish; only for them. You just give them a little of it; it works great; your fish won't die. What you want to do is go to the pet store and say I got some goldfish. I need some fish food. They'll help you out. It's not expensive and it's going to make your fish really great.” And then it'll do it. But you kind of have to be next to them. Like, “Oh, we have this thing that we're both looking at,” rather than, “Are you completely high?” You’re shifting into reverse at 60 miles an hour. Not only will you be hurtful to people, which you really shouldn't be, but at the same time, you won't get anything done, and people will hate you. So, you do have to go to that level of, I'm not all that either, but I've got a secret I can share with you, and then I can help you take care of it. And together, we're going to get things done better.

Corey: So, Conrad, if people want to hear more about what you have to say, where can they find you?

Conrad: I don't know. I mean, I use Twitter, but like most people, I use Twitter in a way that is not particularly handy for technical people. Mostly, it's me yelling about things, sometimes profanely or—

Corey: I don't know what you think technical stuff is, but that's where I live.

Conrad: Right? Right. But there's only political stuff, and shitposting, and pictures of cats, or whatever. I used to write on the internet, a lot more. I mean, I did my decade of LiveJournal. And there is a blog out there, but it's old and not that great. So, I don't think I really have a place. So, yeah, my Twitter handle is @substitute. You can go there and look, and you might or might not like it, but it's not exactly a fount of technical information.

Corey: Yet. There's always hope for a brighter tomorrow.

Conrad: [laughs]. There you go. Maybe I'll reinvent myself and start producing useful information. It’s always possible.

Corey: Thank you so much for taking the time to speak with me. I appreciate it.

Conrad: Great to be here. Thank you.

Corey: Conrad Heiney, principal cloud engineer at Glidewell Dental and @substitute on Twitter. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts. If you've hated this podcast, please leave a five-star review on Apple Podcasts and a comment explaining why it is either the fault of development or operations.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Jessy Irwin

Jessy is Founder at Amulet. She enjoys the challenge of translating complex cybersecurity problems into relatable terms, and is responsible for developing, maintaining and delivering comprehensive ecosystem security strategy that supports and enables the needs of the people who depend on Tendermint and the CosmosSDK.

Links Referenced:

  • https://twitter.com/jessysaurusrex
  • https://jessysaurusrex.com

Transcript
Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Jessy Irwin, who today—doesn't matter at all what she does today because she used to have the best job title in the universe: security Empress at 1Password. I just want to let that sink in for a minute. Jessy, welcome to the show.

Jessy: Hello. Thank you so much for having me.

Corey: So, it's got to be challenging to know that when it comes to job titles, you have peaked, not just as a person but for the entire industry a couple years back, and it's all sort of downhill from there. But what are you up to these days?

Jessy: Yeah, I'm sad that I'm no longer the Empress in my previous place, but I actually have decided that I wanted to set up my own organization to work on security problems. So, these days, I'm not sure if I'm technically Supreme Ruler of the Amulet Universe, but I'm working on my own project where hopefully I can help make security stick better for people. That's my catchphrase: “Help make security stick better.”

Corey: I like that. You’ve become famous in small circles which is, I guess, probably the best way to put people who are big deals on Twitter. But what's always been interesting to me about your approach to security has been the human-centric piece of it, where it's not about trying to talk about the far-future advanced persistent threats, although you certainly can do that, but more along the lines of how you effectively raise the security bar for day-to-day, folks. What got you to focus on that?

Jessy: So, this is the part where I get to tell you that majoring in art history in college was basically the best life decision I've ever made. And I'll tell you why. Art history is interesting because you have to study objects and images, you have to be able to do analysis, especially technical analysis. But when you step back, you are looking at objects that represent societies, and cultures, and lives. And what I remember most about my art history classes and my time as an archaeologist in college is really that people have been engaging in security behaviors pretty much ever since human settlements started.

We've had to protect ourselves from each other, and from external threats in so many different ways, and risk management is something we've done long before computers ever happened. Unfortunately, computers make everything easier to do, especially remotely. So, the same problems we had a very long time ago—keeping our coin hoards safe, for example, we still have those, and it's easier than ever for somebody who wants to separate you from your identity, your data, something that is valuable or important to you that, online, to do that. And I just really think that a lot of times the focus is too heavy on the technical side.

If we're talking about PGP and ZTRP, and we're throwing the alphabet soup together, we're really forgetting the part where somebody just wants to pay their online power bill, or somebody wants to log into their bank account, and know that they're not giving another person all of their money, or all of their personal information in a way that will harm them. And I think that's way more important, and really the core of what we should be doing, instead of engineering these perfect invisible systems that nobody understands, and everyone has to become an engineer to use.

Corey: And that's always been, sort of, the weak spot of security. It's not the advanced super deep-dive breaking into things. It's the fact that someone isn't trained and falls for a spear phishing attempt, and emails the company payroll to someone. It's the human side of people entering their credentials into the wrong website, and it's always seems like it's never the big stuff. In the world of cloud, we see this all the time, whereas you have the S3 bucket negligence story of people failing to secure their S3 buckets, and instead exposing company database backups, people's social security numbers, etcetera.

Then you also do see the more advanced attacks like the one that Capital One was subject to, where there was effectively four or five different misconfigurations that were then chained together in order to result in something kind of neat. But to the outside world, those two things look the same, but they're very much not. It comes down to fixing usability. I've spent an awful lot of time trying and failing to find a legal way to patch humans, and I've never been able to do that. Is this problem ever fixable? Is this something that we're going to continue to see iteration on, on the human side, without getting anywhere? Or is there light at the end of that tunnel somewhere?

Jessy: I'm a little optimistic about this. I hope that after realizing that we can only create so many protocols and so many new whiz-bang code things, that the code is not the answer is now really starting to hit people in the face, or in the feels, or wherever they need to be hit to change their point of view. But ultimately, we have two problems to solve. All security is actually behavioral economics and policy that you have to stick together and align towards a specific outcome. And I think right now, every company is essentially its own little nation-state with its own little national security stance, whether they've got a thousand security engineers keeping one of those many-numbered threats out of everyone's email in the morning, or whether they're a small business down the street.

And it's our job to make security into something that is part of your launch checklist, or your productivity tool, or so normal and so mundane that it's like the cyber equivalent of vacuuming the house. A lot of people refer to what we should be doing as setting up cyber-hygiene programs. That's cool, but we also need to make sure we are thinking about what the people abiding by those programs, or following them, would actually you need to do. You're going to get, realistically, 30 seconds of attention from someone. Even on YouTube, someone bounces from a video after 12 seconds, if it looks sort of boring, so when you think about this problem overall, and this war for attention that we've created with technology, plus all of the new products that come out and all of these sneaky side menus and configurations you have to know, there's always something more to do. And there's always another way to spend more hours of your life trying to secure something that you should. It would be nice to just have 10 commandments that we focus on. And for those of us who are in a position to build products, and to work with product teams or product managers, to just take the core security stuff, put it at the top of the list, and get it done with as early as we can so that we're not all having to freak out and become firefighters and incident responders, with or without tons of resources.

Corey: The challenging part that I found across the board with infosec as a whole, is part of the reason that I've always found you to be such a refreshing voice in the space, is that by all perceptions, from everything you say online, you have an incredibly rare skill in the infosec space, by which I mean, you are not a massive jerk to everyone. There's definitely an asshole problem in the world of infosec. And it's something that you have never exhibited that I've ever seen. How is it that it is, first, so difficult to find people who aren't being obnoxious in the world of security, and, two, how have you avoided it?

Jessy: I think that everyone has an opportunity and a choice about whether they want to be an asshole or not. I tried really hard not to be a giant one but, more than anything, an attitude that has been exhibited to me over the past 10 years I've been playing in security, and the past seven years where I've had direct jobs in security, there's a lot of gatekeeping going on. I mean, I come from a background with lots of humanities, and creativity, and writers, and I love that, but ultimately, the world is a better place when we have more people thinking about these problems, not less. And the attitude that I've seen come from the community around security, and a lot of the industry around security has been to use some of the stupidest things you could ever come up with as a way to intimidate someone from taking a first step into learning more or getting interested because if you have more people who aren't like you join the industry, people who've been around the longest, or people who feel like they get power from their roles, lose that.

And I get it. That's scary. But this is a specific problem where we need to be making friends. Like, we should be in a land grab to make everybody think that two-factor is the coolest thing on the planet, and we've got to be creative about it. Instead of two-factors of authentication, maybe you need two raptors running after an attacker who tried to log into your account if they can't authenticate correctly. That's way more fun to think about. But everything is so serious and end-of-the-world all the time and, I don't know, that just doesn't seem like a group of people I want to hang out with, and it certainly doesn't seem like the way that we recruit and onboard the entire planet into making better passwords and changing their behavior online.

Corey: One of the most transformative things I've ever done for my own personal security was getting a password manager that I could just shove everything into, and then eventually spending a very long few days at previous jobs—when I didn't want to do my actual job—of going through and rotating everything to a unique password. The benefit there is I don't have to remember anything except that initial password, and when one place gets compromised, I now don't have to worry about changing that password everywhere in creation across 400 different websites. What’s surprised me is how easy it's been to get other folks who do not spend their lives working in the world of high technology on board with that. Things like Apple supporting password managers natively in iOS where it just auto-fills from a lot of those apps that are out there has been transformative for an awful number of people. It definitely demonstrates your point of the tide rising. What's next along that axis? I mean, you can always talk about how to get folks who are deeply technical sorted out, but how do you fix this for, I guess, the people who have, you know, real jobs?

Jessy: Yeah. So, this is the hard part. Essentially, we have been very good as technologists at developing broad appeal among our own, but this is actually a consumer branding problem. We have to build a culture around the technology that we use, and we build to make these things that take a lot of extra time on setup, cool, and fun, and worth it. And what that really requires is for us to know our user, and to know our audience.

The conversation that I have with my 72-year-old dad versus the conversation I have with my 27-year-old little brother-I think he's 27, at least—those are two totally different conversations. And when I have to talk to them about why this is important, their incentives are hugely different. If you are talking to your family, and they can remember a time in your life when you were in diapers, they're probably not going to listen to you very often, especially on technology advice. It's why Thanksgiving and all the other holidays can be tough when somebody pulls out a phone and wants tech support.

But we should be able to tell someone who doesn't want to be an engineer why they should do this. If it's for a mom, more often than not, it's a way to take care of your family. Grandma likes taking care of families, too, let me tell you. If you're talking to a student who's never had anything completely terrible happened to them in their life, but you're sending them off to college, and they need a plan for taking care of Social Security cards, and identity documents—really important stuff to do—not for right now, but making it about investing in their future, and making sure that no one else can hurt them when they're on their way up in the world and finding their footing. We have to be able to talk to everyone about why all of this technical security stuff is worth it, even when it fails, even when it's a pain in the butt, even when you really just want to reuse that one stupid password because the password manager is not working and it won't generate or fill the right way because some mean developer made your fashion blog website completely unusable in mobile.

We have to be able to at least incentivize people to do more of the right thing, and maybe not even the right thing all of the time so that we make continuous progress, and we see continuous improvement, rather than trying to get everyone to be perfect at everything at all times for exactly the reasons that we think it's important. Our threat models, as technologists, are completely different than the threat models of everybody else around us. But often, because if you're on a security team, you've seen these advanced threats, maybe you do get a point of view that the password manager doesn't work because the attacker is going to hack into your operating system and steal the plaintext out of your memory, blah, blah, blah. But that's not a reality most people face, ever.

Corey: Especially in a world of Cloud. I mean, it seems to me that a lot of the best practices, like encrypt everything at rest, that made an awful lot of sense, for example, your laptop, or for your data center where there have been numbers of stories of people breaking into improperly secured office data centers, or driving a truck through a wall and grabbing a rack into the back of it and taking off. Good luck doing that with one of the hyperscale cloud providers. But it's hard to get those edge cases explained across the board. So, in many cases, just saying do it all as a best practice seems to be the path forward. It's giving people fewer decision points. If you give people three things to do, they'll generally do them. Give them 100, they'll do none of them.

Jessy: Exactly. And something that I see a lot actually—so I think too much about this—but when you look at health advice, when you look at wellness advice, sometimes you will run into an article written by a doctor who knows to tell you five things because you're only going to do two, and realistically one of them will stick within a month of reading an article. Other times there's a checklist of 283 supplements and things you should be eating, and blah, blah, blah. And at the end of the day, what all of these security problems actually boil down to are lifestyle choices in the same way that some of our issues with healthcare are also lifestyle choices.

I can talk to a small business owner and ask them security intake questions, and just like any other survey, they're going to tell me all of the nice answers. When I actually get in to do the work, I'm going to see where they've done the technical equivalent of having cake for breakfast, and fries for lunch and dinner every day. And it's okay. I mean, that's reality. But unfortunately, I think a lot of us are happy to portray an ideal lifestyle, and we don't actually talk about the lifestyles that we actually live on the security front, and again, it sets this impossible standard, and it makes it really, really hard when you essentially have people writing fanfiction about the 283 steps they take to secure their home and family, when in reality, it's probably, like, 15 or 20. And maybe one or two of them you're a little lax on.

Corey: And that seems to be the, I guess, the message that gets lost, where unless you're doing all of these different things, you're going to be in incredible danger. There was a talk I saw once—I wish I could recall who gave it—where there are two threat models to be concerned about Mossad and not the Mossad. If it's the Mossad after you, you will die. I think was James Mickens that may have made this observation, but please correct me if I'm wrong. If it's not the Mossad, well, there are things you can do.

It's about raising the bar for what it takes to compromise someone. At some point, if people invest enough resources, they're going to wind up breaking into anything you've got. The question is, is what is that bar? If your password is the same thing everywhere, and it's just the word kitty—sometimes an exclamation point at the end—then maybe you should try and raise what it takes a bit further. But past a certain point, it winds up mattering less and less. An easy example of that would be—I’ll ask you this: is there a material difference between whether I have 40 characters in my password, or 50?

Jessy: I would say yes, but half of that difference depends on what is going on on the back end of a web service you might be using or server infrastructure, and how that's been designed. So, maybe you don't actually know as a user, and it never comes up to you. On the other end, maybe it's no because your extra 10 characters are all zeros, or they're the same word twice. That's really easily undiscoverable. There's a quality piece there that is really difficult to judge. There is a numerical difference between the password strength of 40 characters versus 50 characters, but on the other hand, if you're reusing the same word five times to get to 50 characters, maybe it doesn't matter. Maybe there's really nothing there, and it's going to get broken easily anyway.

Corey: This episode is sponsored in part by ChaosSearch. Now their name isn’t in all caps, so they’re definitely worth talking to. What is ChaosSearch? A scalable log analysis service that lets you add new workloads in minutes, not days or weeks. Click. Boom. Done. ChaosSearch is for you if you’re trying to get a handle on processing multiple terabytes, or more, of log and event data per day, at a disruptive price. One more thing, for those of you that have been down this path of disappointment before, ChaosSearch is a fully managed solution that isn’t playing marketing games when they say “fully managed.” The data lives within your S3 buckets, and that’s really all you have to care about. No managing of servers, but also no data movement. Check them out at chaossearch.io and tell them Corey sent you. Watch for the wince when you say my name. That’s chaossearch.io.

Corey: An argument that I heard once, when botnets were on the rise, was that at some point you have to begin assuming from a security posture that whoever it is that is attempting to break in will, more or less, have infinite computing resources to throw at this. So, the answer starts instead becoming things like two-factor auth, or as you correctly pronounce it two-raptor off. Tell me more about that.

Jessy: Yeah, so the main goal of security, it's not to keep everyone out all of the time. If that's our goal, we're going to fail at it, and we should never take any of these jobs or even bother, quite frankly. But the main thing that's the most important to do from a security perspective—whether you're a farmer 50,000 years ago, or you're the guy holding the keys to the Vatican art galleries—I promise I'm going somewhere with this—it's important to raise the cost of an attack. What's really interesting, especially online, is we figured out that passwords are the weakest link. They are a huge privacy problem. They're a huge security problem. So, essentially, we need something else.

We all came together and decided that we would make sure every computer on the planet had two velociraptors that were trapped inside, and in the case of an attacker coming to try to steal your passwords, they'd be unleashed and they'd go eat his face off. Or at least that's how I like to explain it. In reality, we needed something else. We needed another layer of defense, and the best thing we came up with was, I guess, a rotating 6 to 10 digit code that lasts for anywhere from 30 seconds to 5 minutes. It's very hard for an attacker to steal from you or to take away from you, especially if you're using a physical security key that produces those numbers automatically.

I kind of joke that with two-factor authentication, instead of just one password, now we have two because it is basically a one-time-use password in most cases. But we're always going to find that we need to add an extra layer or a something else. It's just a question of how much of something else will our users be willing to sustain, and to do? And when should we remove the burden of doing a something else from them and integrate it into what we're already building?

Corey: And that's part of the problem. It's not always about how to build something new, and exciting, and revolutionary, but how to retrofit that back to real world problems that people have. Now, that brings us, of course, to blockchain. My argument for a while is that it's a neat technical trick that is struggling to find anything approaching a business model for the past decade that doesn't revolve around speculation or scamming people. Where do you stand on this given that you actually work with it in capacities that aren't just making fun of it on Twitter?

Jessy: Honestly, I agree with that assessment. One of my most endless frustrations for the past two years of working in the blockchain space has been watching people pay more attention to coins, and their value, and their worth, versus some of the fun things that we're actually engineering with code. And the hype machine is, frankly, incredibly annoying. There's some really interesting things we get to play with in blockchain that I don't think anyone really realizes. We get to play with virtualization; we get to mess with encryption; we get to do all kinds of exciting encoding and decoding things.

I feel like in some aspects, there's actually been a huge contribution from blockchain-land, to web application security, just based on the fact that we offer built-in bug bounty programs for any code holding coin. There's also—non-joking—playing with distributed networks; playing with the concept of decentralization, which really gives security properties resilience and redundancy to networks, which is really exciting. But overall, the big disappointment is just watching people try to slap a blockchain on any problem that exists, rather than stepping back and thinking about some of the properties that you can get from what exists, or how you can tweak what exists to get something new, and refreshing, and exciting we haven't really seen before.

Corey: And that's what's interesting is the idea of exposing new and exciting things that until now, we did not have the capability to solve, is incredibly promising. And I love the idea. The problem is, is that so far, most of those examples revolve around well, there's no central authority that we can trust. Well, we tend to live in societies where, for better or worse, there are parties out there that even if don't we don't agree with them personally, they are sovereign entities who are empowered to enforce what they want to do, so you, sort of, have to trust them. There are stories for removing middlemen from transactions—or middlepeople from transactions—that tend to be compelling on one level, but in practice, there are a lot of entrenched interests that are going to fight explicitly against that. I think that rather than looking to necessarily supplant existing structures with a complicated by-in story, finding new and exciting things that this empowers is probably the right answer, but I see less of that than I would like.

Jessy: I totally agree. Personally, I might not be salty enough to work in security when I say this, but eventually you have to trust something, somewhere. There's going to be a Root of Trust in anything you build, whether it's cryptographic, whether it's reputation based, it's just going to happen. You make the decision to trust, or to buy-in based on a something somewhere. These things don't develop in a vacuum.

So, trying to watch people basically deny that trust is required, and then go build a trustless universe has been this Kafkaesque adventure that I've been on. But more than anything—okay, the first exciting application everyone came up with for blockchain that was sort of meaningful was supply chain. We see all of these case studies from IBM, and we see case studies from Walmart, and major retailers, and major regional grocery chains, who are using blockchains so that when you scan a QR code on a product, you can see where your lettuce was grown, and all the facilities that it went through, and how it made its way to your table. That's really exciting.

Do I know if it requires a blockchain versus a transparency tree? I'm not totally sure, but okay, fine. Keep going with that. I'm totally down with more technologies that can empower transparency, especially in an end to end situation for food, or medication, or agriculture, where the choices we make impact the future of our planet, and the health of our bodies. That's fine, but what I think is really being missed in all of this blockchain cryptocurrency hype is the opportunities we have in some other places.

Microsoft has done some incredible work on distributed identity, and decentralized identity. There are so many opportunities for tamper evidence, and resilience, and even sharding identities, that blockchain technology can give people to play with. And especially given how hard identity problems are to solve in security and in computing, it would be nice to see more people besides just Microsoft get laser-focused on where the opportunities are and what we could build out of that playground. On the other hand, one of my favorite applications of blockchain has been watching all of these different mesh networking technologies essentially plug themselves into blockchains, and to enable an entirely new decentralized infrastructure for the internet. From a security perspective, watching major protocols get hijacked.

Watching all of these cloud providers have massive DoS attempts, just get thrown at them all of the time, I would love to see more resiliency and more decentralization of our network, so that they can be tamper evident and more fault tolerant. But I don't see enough people getting excited about where else we could go with this that has nothing to do with money and more to do with resilience and making sure that we're not just all making these internet companies and these services that are too big to fail, and when they get attacked the right way, they fall over and that's it. I mean, everyone freaks out when Twitter goes down or Slack goes down. Or Zoom, by the way. It's like the end of the world. But how cool would it be to say maybe this internet we've been building since the 70s needs a rethink. Maybe we should play with mesh networking more. Maybe we should think about how we are connecting the rest of the world. And maybe that model doesn't need to look like what we've been doing previously.

Corey: That was a fascinating question is, what do those new models look like? There's a question of, we should build out these new formulas, these new structures, these new approaches. It's just hard to find people that are genuinely doing it. Blockchain has fallen into almost the category of punchline, in the same way that AI and machine learning is, to the point where, when I see someone talking about a blockchain-derived-machine-learning-powered-serverless organization, oh, you're trying to scam money from VCs.

Why didn't you say so? I know the secret handshake, too. And it seems that very little transformative winds up being derived as a result. It winds up, from my perspective, tarring a lot of good faith efforts with a somewhat ridiculous brush. I mean, one of my more obnoxious tweets on this was, if I had somehow come up with a terrific, transformative, legitimate usage for blockchain, I would go significantly out of my way to avoid referring to it as blockchain so that people would take it seriously. It's an ongoing problem in the space, to a point where it's almost impossible to have a serious conversation about it without some subset of the population rolling their eyes and tuning you out.

Jessy: Yeah. This is a problem I've dealt with for the past couple of years. When I want to talk about what I'm working on, I don't use the B word. It's a bad word. In the blockchain industry, and in the communities around all of these different coins and network protocols, people double down on that blockchain culture that we've heard all kinds of stories about, which is really difficult because it's hard to create broad appeal for people who do want to work on some of the engineering problems in this space that are super interesting.

On the other side, it would also be great for people to just suspend the jokes for five seconds, and think about when we've seen this behavior before. I remember, what, a decade ago when people started really getting into the Cloud, every security person on the planet was just like, “Oh, it's not the Cloud. It's somebody else's computer.”

Corey: Yeah, that is a tired trope at this point.

Jessy: It is. And I think about it this way: the point of security, and I think the point of technology, is supposed to be that we make cool, creative, amazing stuff happen, even if it's sort of wild, and a little nutty, and you have to suspend some disbelief like it's some movie. But on one side, I remember everyone criticizing the Cloud out of existence or so they thought. And I remember the race to the bottom for the jokes, and, “Oh, who's going to use that?” I also now look at the environment in the space, and I see security engineers pulling their hair out because instead of running to the front lines and trying to figure out how to get involved, and how to move things forward, and to get security in at the very core, they just made a bunch of jokes, and dug their heels in, and thought that saying no was going to be enough.

And from a security perspective, I think this is a huge industry problem, but also, you're not going to criticize something out of existence. On the Cloud side, look at the market cap of Google Cloud and Amazon. Look at all of the Cloud bills people pay. I think I even pay two different cloud providers right now—

Corey: That you know of.

Jessy: Yeah. Two that I know of, technically. But I feel like in the blockchain space—not the cryptocurrency part, but the blockchain part—there's billions of dollars hanging out over there. People are funding research, and trying to at least, have a creative, experimental place where we're trying to figure out how to make things better, and play around a little bit. When that used to happen, when people did it in their garages in the 70s, it's totally cool. And now, it's kind of bad, and awful, and evil, and we shouldn't do it, and it's a huge joke, and I don't quite get it. There's a bit of a disconnect there.

Corey: I would agree. But I think this also gets to one last point that I want to talk to you about, which is, how do you, I guess, evolve the mandate from the way that security currently is from this idea of being top-down—command and control everything—to being something that helps people get further, faster? I mean, how many people do you know who wind up effectively having a second computer to run the antivirus suite that their company mandates that they run, and they have something else that they do their actual work on, then copy it over? I mean, it completely defeats the point.

Jessy: So, something I think a lot about is how, in development environments and engineering teams and even among product teams, the Cloud coming after us all has essentially changed how security teams have to work with the rest of an organization. Theoretically, I think it's a major failing that we take all of the riskiest, hardest, most complicated things away from our developers day-to-day, and we shove them onto this team from the side that's usually understaffed and probably under budget and never going to be able to get ahead of an entire organization, and we make it their problem, and nobody else has to think about it. That is so wrong because what are we supposed to do? Have a 10 person security team in a 500 person company, essentially split up and be in charge of enforcing X number of employees across Y things? It's a recipe for burnout.

What we have to stop doing is looking at our jobs as control, and power, and doing the things that are not great. So, many security teams that I've interviewed with have said, “Hey, actually, we're so powerful, we can stop a product release.” And I don't know if that sounds like power, or if that sounds like being a jerk, because frankly, if somebody worked for two years, or a year, or however long your product cycle takes, to ship something, and you had no involvement in it, or you had minimal involvement in it, but at the last minute, you can throw your foot down and stop it, nobody's ever going to want to work with you again. What we have to get better at doing is building coalitions, making friends, aligning our incentives with one another.

And we also, a little bit, have to get to know the people around us. We can't just all hang out with our hacker buddies at conferences, and think that we're changing the world. We need to go hang out with marketing because they control reputation, and reputation is kind of a big deal. It's a huge asset for a company. We need to go hang out with operations; we need to go hang out with finance, and all of the critical functions of business, who maybe need it support, or who are going to always be looking for shortcuts, because they're understaffed as well, and we have to learn how to advocate for them.

We have to learn how to make things easier for them. And we have to not shame the crap out of them if they get something wrong. Security teams get it wrong all the time. That's why we have all of these issues with data leakage, and data breaches from AWS buckets. Yeah, they might get shamed by their colleagues on Twitter, but everybody else is probably too afraid to speak up, and to try to advocate for change with them. We have to be better ambassadors, and we have to be better at building teams and collaboration, or we are going to fail so miserably, that at some point, it's just going to be too expensive for anybody who's not a company with their own cyber military to actually go online and do business. And that's not okay. That's not what the internet was for.

Corey: What's the role of your CSO? Uh, mostly to sit in an office and play with a desk toy until the next data breach, and then they get ceremonially fired and replaced. That's not a viable outcome, even though that seems to be some company's actual strategy.

Jessy: Most people that I know who have been in a CSO role, especially actually in blockchain space, they get asked for policy all the time. As if writing down a bunch of rules is going to be what protects you from an attacker who doesn't give a crap about any of your rules.

Corey: As my primary IDE PowerPoint, that's usually not the right answer for a lot of these things.

Jessy: Exactly, exactly. And it's just really unfortunate that, again, we take these people who do security work, and we shove them all in one team instead of embedding them, or putting them in a position where they can educate, and advocate, and also build in technical reinforcements, and technical support, and monitoring, and metrics, and visibility. The answer isn't to shove everything into a black box security team; the answer is actually to make the creation process for whatever you're building and whatever your business is, more internally transparent, so that when an issue pops up, the humans who are good at identifying risk—security team or not—have a place to go and can voice that because more often than not someone closest to a business process, or closest to critical code knows where the bugs are going to live anyway. And maybe they don't know all of our vocabulary words. Maybe they shouldn't have to go memorize an infosec dictionary to make a point, or to surface something. And that's really the mindset that I think more security people need to have. We shouldn't force everyone into our worldview. We should be able to have someone describe a situation the way that they might describe an ache or a pain to a doctor, and work from there to diagnose what's actually going on, not just throw a hissy fit and tell them to stop, I don't know, looking at cats with cheeseburgers on the internet, because that's where the malware comes from.

Corey: I think that it always comes down to meeting people where they are. And we see that in Cloud, we see it across the board with user behavior, and these problems aren't getting smaller. They're definitely getting larger. If people want to hear more about what you have to say on this and countless other topics, where can they find you?

Jessy: I will resume yelling on Twitter again soon about these topics. I've taken a bit of a hiatus because I've been in creative build plans, and conquer part of the world again mode. But I blog at my website, jessysaurusrex.com, though I don't do it often. And every once in a while a cool person will invite me to their podcast, and I can rant for a while. Another place to look would be to look for some of the keynotes or talks that I've given at previous conferences, especially if you don't travel for conferences quite a bit. That content I try to make evergreen and helpful.

Corey: Thank you. And I will throw links to that, of course, in the show notes. Jessy, thank you so much for taking the time to speak with me today. I appreciate it.

Jessy: Thank you so much. And don't forget to turn on your two-raptor authentication, if you don't have it on already. The dinosaurs will thank you. They are hungry.

Corey: Jessy Irwin, former Security Empress at 1Password. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts. If you've hated this podcast, please leave a five-star review in Apple Podcasts, and then be sure in the comments to leave your date of birth and mother's maiden name.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Dwayne Monroe
I've been a technologist, in some form, for most of my conscious life (starting with a Sinclair kit computer). I work as a cloud architect, focused on Azure and spend a lot of time thinking and writing about that (particularly controlling spend). Besides that, I enjoy a good Bordeaux or martini, travel and my life as a transplant to Amsterdam.

Links Referenced:

  • https://retool.com/
  • https://twitter.com/cloudquistador
  • https://www.linkedin.com/in/cloudquistador/
  • http://monroelab.net/projects/cmdlet-daily/
  • Books on Amazon

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is brought to you by DigitalOcean, the cloud provider that makes it easy for startups to deploy and scale modern web applications with, and this is important to me, no billing surprises. With simple, predictable pricing that’s flat across 12 global data center regions and UX developers around the world love, you can control your cloud infrastructure costs and have more time for your team to focus on growing your business. See what businesses are building on DigitalOcean and get started for free at do.co/screaming. That’s D-O-Dot-C-O-slash-screaming and my thanks to DigitalOcean for their continuing support of this ridiculous podcast.

This episode is sponsored by a personal favorite: Retool. Retool allows you to build fully functional tools for your business in hours, not days or weeks. No front end frameworks to figure out or access controls to manage; just ship the tools that will move your business forward fast. Okay, let's talk about what this really is. It's Visual Basic for interfaces. Say I needed a tool to, I don't know, assemble a whole bunch of links into a weekly sarcastic newsletter that I send to everyone. I can drag various components onto a canvas: buttons, checkboxes, tables, etc. Then I can wire all of those things up to queries with all kinds of different parameters, post, get, put, delete, etc. It all connects to virtually every database natively, or you can do what I did, and build a whole crap ton of lambda functions, shove them behind some API’s gateway and use that instead. It speaks MySQL, Postgres, Dynamo—not Route 53 in a notable oversight; but nothing's perfect. Any given component then lets me tell it which query to run when I invoke it. Then it lets me wire up all of those disparate APIs into sensible interfaces. And I don't know frontend; that's the most important part here: Retool is transformational for those of us who aren't front end types. It unlocks a capability I didn't have until I found this product. I honestly haven't been this enthusiastic about a tool for a long time. Sure they're sponsoring this, but I'm also a customer and a super happy one at that. Learn more and try it for free at retool.com/lastweekinaws. That's retool.com/lastweekinaws, and tell them Corey sent you because they are about to be hearing way more from me.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Dwayne Monroe, Senior Cloud Architect at Cloudreach. Dwayne, welcome to the show.

Dwayne: Thank you, Corey.

Corey: So, you come from a world that I don't understand in the least. I'm not talking about Amsterdam, which is where you apparently live, but rather that you're a cloud architect who's focused on Azure. I tend to operate in a world where I deal with obviously a lot of AWS, I do see some GCP from time to time, in similar environments, but while I do have customers that are doing significant spend with Azure, I never encounter it in the course of what I do with them. So, what I brought you here to ask you about is, you don't work for Microsoft. To my understanding, you've never worked at Microsoft.

Dwayne: That's right. Yeah.

Corey: And what I want to know is, tell me about Azure customers, please. As someone who does not have a particular horse in the race that depends upon you selling Azure to answer the question.

Dwayne: To answer the question, I think, let me tell you a little bit about my own, to use the term of art now, cloud journey; forgive the phrase. I started my career, as many of us of a certain age did with general geekery. And then when I graduated from university—in which you have computer labs, and all that kind of stuff—I was floundering for a job, and I was working at a boutique bank in Philadelphia. And that bank found itself facing some challenges from the FDIC. That led to me becoming—because I was the youngest person there and the person who seemed to have a handle on client/server technology, that led to me becoming the person who deployed the first network.

Kind of fast forward a bit. I'm working in enterprise-scale data centers, like pharmaceutical firms, and this sort of thing, and at one of these particular situations, I was introduced to AWS because there was a problem that the organization had, which was scalability. We were selling books, and every season, the type of book that was being sold, the VMware infrastructure that we were using would be stretched to their limits. A very smart colleague said, “Hey, how about EC2 instances, Elastic Beanstalk?” I’m like, “Well, what the hell are you talking about?” So, I dove into that. So, I began my cloud journey with Amazon, as I think many people did—

Corey: For a while, they were the only option. I mean, you wouldn't call it Cloud if you spent a long time working on, I don't know, VPSs offered by some fly-by-night web hosting company. That, it would sort of look like cloud today, but we never called it that back then.

Dwayne: That's exactly right. And Microsoft at the time under, I think it was still Mr. Ballmer, Azure was just Windows Virtual Machines, but it was very weak. So, I didn't pay attention to what Microsoft was doing in that space because what Microsoft was doing wasn't compelling. Fast forward a couple of years and leave out a couple of details, I'm working for a firm out of New York that had pivoted to Office 365, and one day—and this may make me sound like a genius or may make me sound like a goofball, I don't know—but one day, I was sort of poking around at the base of 365, I noticed that there was an Azure AD tenant. And I said, well, let me just log on to this Azure AD tenant, and voilà, there's all this going on. And I’d been so Amazon focused, and Microsoft up to that point—and this is going back four or five years—up to that point Azure had not been compelling, but it was becoming compelling.

So, I pivoted to Azure because I'd already had kind of a deep commitment to Microsoft technology, to the Microsoft stack, and it seemed to be a logical progression for my career. To get to your question, and that's a long-winded way of getting to your question, the typical Azure customer, from my point of view, is an organization that has a deep commitment in Microsoft stack, and sees the need to modernize that stack into the Cloud, and Microsoft has provided a bridge. So, you have SQL Server on-premises, and you can modernize that SQL Server, you can do it on-premises. But then that SQL Server on-premises, the code has been put into SQL server to see assets in public cloud Azure.

So, Microsoft's hybrid story, I think, is very well realized. And so the typical Azure customer, from what I've seen, are customers that say, “We have these investments in databases and so forth, storage that are Windows-based, built around the Microsoft stack, and we need to get out of the data center business; we need to get out of the infrastructure business. Let's move that stuff to the public cloud. And Microsoft has already built the code that allows us to do that, if not easily, at least you can see the direct path.” You can cross the Rainbow Bridge, and go from where you are within your messy data center to your cloud estate. It's very logical.

Corey: I have a standing policy of not insulting various customer choices, workflows, environments, etcetera, unless A) they are very clearly egregious or B) I'm trying to make a larger comic point. And I want to revisit that because there's absolutely nothing wrong with what you have just described. The whole Silicon Valley model is built upon more or less sneering condescendingly down your nose at anything that was written more than 18 months ago, but that is very clearly not how the world works. We are focused on building new, but everything that we're building on top of has been around for ages. And just because there's a new or different paradigm for developing these things, does not mean you get to sweep away the last 30 years of development work. So, there's a tremendous need for an awful lot of these workloads to migrate out of the data centers in which they find themselves. So, far, the worst environment I've ever seen, from a physical risk standpoint, was where the data center was in a company's office. That office was located on a boat. That boat was holding the data center, and data center was below the waterline.

Dwayne: Oh, my God.

Corey: So, there was a migration to Cloud, and not surprisingly, there was a challenge as soon as that was completed. People who were working on the boat were reporting severely increased latency.

Dwayne: Yes, of course.

Corey: This stuff becomes a problem. Stacks aren't written 15 years ago to take advantage of the new paradigm that we find ourselves in today. How do you get someone to go from where they are to where they should be without A) insulting them; B) disrupting what they're doing, or C) forcing them to re-architect everything instead of developing new things for the next couple of years. And Azure has, to be honest, a terrific story around that.

Dwayne: I think so. And I'm not a Microsoft partisan because I definitely see strengths and weaknesses in all of the major CSPs, Cloud Solution Providers, including Alibaba, and even Oracle within a very narrow frame. But yes, I think that—just from what I've seen, and again, I'm speaking as a person who admires what Amazon has achieved, but from what I've seen dealing with enterprises—and let's just take an example: you're making paper clips. You're not a sexy Silicon Valley company. You're not Twitter for Pets.

Corey: Well, you're not Twitter for Pets.

Dwayne: Right? You're not building satellites that beam lasers amongst themselves or whatever. You're making paper clips and you're a multi-billion-dollar enterprise, and you have manufacturing concerns. You have all the classic concerns of companies that do things. And you're saying I understand that my estate of Windows 2003 Servers and Exchange 5.5, all these things. I understand that all these things are a mess. I know that. What do I do to modernize, and the data center refresh is coming up, a hardware refresh cycle is coming up. And my god that bill, it looks pretty bad. Looks pretty bad, am I'm getting yelled at because the CIO went golfing with someone and said, “Well, we're in the Cloud. What are your guys doing?” and was like, “Well, we still have a bunch of servers in our data center.”

I think that the way that you would help customers bridge this gap is number one—and this is I think, where Microsoft does well relative to its competitors, and also where those of us who have been in this business for a while do well, is you, first of all, acknowledge people's legitimate concerns. You do not sneer; you do not condescend. I was at a client that's a manufacturing firm in North America, and the guys who run—and it was all men. That typical—the guys who run the manufacturing facility had legitimate concerns about plant operations going wrong, and things exploding. I mean, literally exploding if valves didn't open, and various things didn't happen. So, the systems control and data acquisition, SCADA network that they use in their plant operations could not be relocated to the Cloud. Full stop.

So, in the push for the Cloud, you have to listen to that, and then architect a solution that allowed them to participate in what was a larger cloud migration strategy without being left behind, but at the same time, which acknowledged their challenges. And so, in the case of this particular project, it was that there was tolerance for the data that was gathered to be on Azure SQL, and Azure SQL Data Warehouse, and eventually Cosmos DB. There was tolerance for that because that was not part of the control plane. It was part of the analysis plane, so if there was a little bit of latency in doing analysis, that was fine. And that took a huge load off their shoulders because they were hosting both the control plane and the analysis plane within the plants that this particular organization has.

So, listening and crafting a strategy that is sensitive to the real-world concerns of organizations. And again, people are actually—they're running vast enterprises, or even mom and pop shops, whatever it may be, they're running a business, and they have legitimate concerns. Sometimes they make mistakes. Sometimes people are stubborn. We are error-prone creatures, so there's that. But often at the base of it, there's a legitimate concern—and usually, there's legitimate concern. That's what I found to be successful.

Corey: I think that this turns pretty easily into a cloud migration story. The problem is, is that it seems that every Azure customer I wind up speaking to, in some way turns into having been a migration story like you're talking about. Do you see net new being built on Azure?

Dwayne: I do. And one of the things I see that is very popular on Azure, and I've seen this multiple times, is IoT. I've seen this in customer after customer after customer. And also net new databases. So, they're building new solutions that are not necessarily customer-facing, by which I mean, they're not the sneaker store. You might build that using Lambda or some other platform, but all of the infrastructure that supports that revenue-generating function might be developed on Azure. And I've seen this an awful lot because I think people—it’s particularly companies that have an acute need for security and compliance to meet those needs, they turn to Azure because I think that the role-based access control story, some of the methodologies such as blueprints and policies, and so forth, it's very, very strong. And it also makes the life of individuals who have to be concerned in, say, pharmaceutical, and finance, and so forth, it makes their lives easier because the reporting from the platform about what's going on is very robust. So, those are the kinds of net new solutions I'm seeing being built, particularly as I said, in those areas in which customers have a very strong concern about compliance and security. I've seen customers turn to Azure for that.

Corey: One thing that I’ve found that I think is, I guess sort of the outlier—it was odd enough that it's worth commenting on—was, I made a tweet a while back, that whenever Azure has issues, it seems like nobody's website goes down. And that was sort of a facile, off-the-cuff response, but what made that interesting to me was that it was basically true. I know it's a weird, and it sounded borderline insulting way of framing it, but in the tweet, I said, “Now, I'm sure there's a bunch of SharePoint servers that are broken.” And a lot of people came in and weighed in on that—oh, did I get letters on that one, all saying, “That's not true.” But I wasn't getting any factual correction or substantiation for it other than, “Well, we don't really run SharePoint anymore. That's kind of what Microsoft Teams is.” Ah-ha. I knew it.

But the lesson that I took from this, though, was that okay, so what are those workloads? IoT makes an awful lot of sense and to be honest, I really should have connected those dots sooner. I had Dr. Galen Hunt from Microsoft, working on Azure Sphere—which is an IoT security platform—about a year ago now. And he had a great conversation about what this looked like. But you generally don't see a whole lot of net new web apps, but I found one. In fact, it turns out, I'm the customer of one of them, and that, sort of, threw me for a loop. I've been somewhat public about using Retool—that's retool.com—not, at time of this recording, a sponsor of anything that I'm working on. But I'm working on them to fix that because I talked enough about how awesome they are that they frankly should be paying me by this point.

Their service runs on Azure. They’ve positioned themselves as a basically Visual Basic for tying web API's together. Though, that is my phrasing, not theirs. I think frankly, it's a better phrasing than theirs because it makes it intuitively clear what it is. I don't have front-end skills, so I can slap everything I need together to build my newsletter, and click a few buttons and it generates what I need. I can pass this off to other folks who I've hired to do some internal work on, for example, getting the sponsor copy in where it needs to go. And this all ties back together and it's transformative. But it's never been intended—from their perspective to be something that is public: it's for internal applications. But it is itself a web app being sold to the masses. So, they built a SaaS product, but even the SAS product that they built on top of Azure is still aimed at back office workloads, and I don't think there was anything intentional about that. I mean, I don't have insight into the decision, but even when you have built something that is effectively public-facing that public-facing thing is just for back-office stuff. If, for example, all of Retool and their hosted version went down because of an Azure outage, that would not cause any disruption to the websites of its customers.

Dwayne: Exactly right.

Corey: Of Retool’s customers. Now, Retool itself may very well have its marketing site down, but it's not going to cause, for example, the front page of an e-commerce website to crash.

Dwayne: That's right. Yeah, that's right. And you make an interesting point because to return to the example of the manufacturing firm I was mentioning earlier, they did build net new web applications using, at the time, Azure web apps. Facing customers, but in their case, the customers would be other organizations. Companies that were ordering their product and could order millions of units of the items that they manufacture. But you and I wouldn't see that if it went down, but many, many large enterprises, including McDonald's and so forth, would know that because they use it to order some key parts of their logistics chain. And so that's the kind of thing I am indeed seeing with Azure that people are building these sorts of things.

I think that—this is just my opinion on this, obviously—I think that it might have to do with who has the mind-space of developers. Who do developers think is worth pursuing, or is a cool organization or something—and not just cool. That’s kind of pejorative and dismissive—but who developers feel is listening to them and offers them the tools that they prefer. And I believe that Microsoft is extremely developer-friendly, but I think other organizations such as Google, and also Amazon, I think that they have, kind of, a panache about that, that perhaps Microsoft doesn't have, or at least perceived as not having because you could very, very easily, or at least you certainly could build a customer-facing web application. It could even be serverless; it could be using Azure Functions, and talk to Cosmos dB, and it could sell your really cool sneakers, but for interesting reasons, I think, that maybe are deserving of delving into, a lot of developers are not pushing for that; they're not going in that direction.

But I think that management from the other side is saying, “Hey, when we do our work, our really key business glue work, it's going to be in Azure.” Again, this is what I've seen many many times. When you ask people, “Well, so why did you choose Azure for this particular workload, these particular series of projects?” “Well, there was advocacy within the organization for people who are advocating for Azure. And also, management said this meets our business requirements.” It doesn't have to do with what you like, we did an analysis, and this is meeting our business requirements, again, getting back to the compliance and security, and also the fact that they can modernize existing skill sets for public cloud.

Corey: This episode is sponsored in part by ChaosSearch. Now their name isn’t in all caps, so they’re definitely worth talking to. What is ChaosSearch? A scalable log analysis service that lets you add new workloads in minutes, not days or weeks. Click. Boom. Done. ChaosSearch is for you if you’re trying to get a handle on processing multiple terabytes, or more, of log and event data per day, at a disruptive price. One more thing, for those of you that have been down this path of disappointment before, ChaosSearch is a fully managed solution that isn’t playing marketing games when they say “fully managed.” The data lives within your S3 buckets, and that’s really all you have to care about. No managing of servers, but also no data movement. Check them out at chaossearch.io and tell them Corey sent you. Watch for the wince when you say my name. That’s chaossearch.io.

Corey: I think you're absolutely right as far as what you just said. There's a tremendous groundswell of developer energy that is pushing for a lot of the developer-first platforms, by which I'm talking specifically about AWS because people have more experience there, and GCP because their developer experience is phenomenal. I mean, that sincerely. Azure has really caught up in a bunch of different ways, and things like Visual Studio Code are transformative.

Dwayne: I love that app so much. I didn't in the beginning, but now I’m like, “This is really fantastic.”

Corey: Get it working on an iPad, and I'll be a convert for life, and I am deep into the vim weeds. Their acquisition of GitHub was transformative, and they are leveraging that in an absolutely major way. I think that they absolutely are positioned to go super well. And I think that saying, “Well, they're not necessarily going to be something where developers flock to,” I think that's wrong. If you take a look at the industry of development across the board, there's an awful lot that doesn't get well represented in typical developer surveys.

These are folks that spend 40 hours a week writing Java at some large company, and then they go home. They don't tend to engage in community as much as some of the avant-garde front end development types might. You see people doing the same thing with a whole lot of .NET and earlier Microsoft technologies, and there’s this certain type of developer out there—and I say this with no condemnation, and I’m not trying to sound pejorative at all—but there's a subset of folks for whom this is not a passion or part of their core identity.

It's a job similarly to the way that an awful lot of accountants don't go to accounting conferences, or hang out on accounting backchannel Slack rooms. They show up, they do their job, and then they tend to go home and live their lives. And there is absolutely nothing wrong with that approach. And Microsoft has been phenomenal at addressing that constituency historically. Just because we don't see the new whiz-bang startups built on this, does not mean that there is not a tremendous groundswell of Azure uptake. I do not believe that Azure is not in second place, as most analysts tend to say that they are. I think that their cloud numbers are largely real. I think that there's a tremendous groundswell of enterprise support for this, but enterprises hate anything that remotely looks like publicity, particularly around something that carries perceived risk.

Dwayne: That's exactly right. When the JEDI announcement came down, the Pentagon selection, I recall on Twitter, there was a lot of outcry from Amazonians who said, “Well, this could not possibly have been driven by technical excellence, but it was entirely political.” And I have to say, as a person who has enjoyed and built things on both platforms, I found that to be a little bit astounding, because can you admit that there is competition? But also, it was very clear to to me that the [00:24:02 hyper division]—Azure Stack, which we haven't talked about, but I think Jeffrey Snover’s work and others’ work on that has been really quite strong—and the compliance and security model, if I was the person in the Pentagon, looking at public cloud, Azure would make a lot of sense to me, because here I have Azure Stack, and so I can have isolated Azure, that is entirely secure; it can be off the grid; it can phone home when necessary. It can be in remote places, like bases and so forth.

Or the classic example that Microsoft gives, which is the cruise ship. And then I have Azure, which I have all of these tools available to me to track, to control, to monitor, to know exactly what's going on. Now, that doesn't mean that these tools are effectively used all the time, but they're all there. And so this cornucopia of command and control tooling, which is well designed, probably built upon—not probably—definitely built upon what Microsoft learned from on-premises Active Directory, and LDAP over the years. Getting that wrong, getting increasingly right with each iteration, and then finally, sort of, migrating that to the Cloud, and then making mistakes and learning.

Those are all extremely compelling, and that's why I think that the audience for Azure, I'm just trying to find the words to say this without coming off as a jerk, but the audience for Azure, I think, tends to be people who are more concerned about foundational operational things than the very, very flashy things. Which is not to say that you can't do these things on Amazon. Of course you can. But I think the audience, I think that the mind space or the attraction, I would say, to Azure is very clearly for people who say, “I have these problems. I'm an enterprise. I want to take advantage of Cloud, but I need to be secure. I need to be compliant. What's the best platform for that? I've looked around and I think Azure is it.” It's not a mystery to me why organizations would be choosing the platform for those things.

Corey: Yeah, and I think that that tends to be a very different story than one that we've been seeing specifically. That brings us to another topic I wanted to discuss with you. Let's talk about cost. What does that look like in the world of Azure? How do you wind up handling cost prediction, cost allocation, what does negotiating these deals with Microsoft look like?

Dwayne: So, there's a couple of elements to cost with Azure. There is, of course, the classic negotiation with Microsoft. And that is an area that I have more experience with than I care to admit, having worked with enterprises to help them do that. But once you get past that particular hurdle—because it's not as complicated today as it was in years past. I mean, it still can be complicated, but it's not as nuts as it was in years past with the licensing, and so forth. Microsoft has done a lot to simplify that. It still just comes down to consumption. And yes, there's reservations and all these things, similar to Amazon and, I believe, GCP as well, but it just comes down to consumption and control. So, on Azure, you have the classic problem that you have on all cloud platforms, which is what people are able to deploy, and they have to deploy solutions. They do so. And they do so sometimes—especially when they're really cooking with wild abandon, and they're not turning things off—it's all the problems you have on all the cloud platforms. They're over-specing virtual machines, or they're choosing virtual machines instead of choosing platform services. They're doing things that are generating a lot of cost that could be avoided. And so the thing that I think Microsoft has offered customers that sets them apart from their competition in this area—and I've actually written a book about this—

Corey: And we'll throw a link to that in the [00:27:53 show notes].

Dwayne:—thank you—is the Azure cost management and cost analysis tooling.

Corey: Now, before we begin to continue down that, did that grow organically, or was that what they did when they acquired Dyne, or DYN or Dyn, or whatever the hell they call it?

Dwayne: Yes, Cloudyn. Cloudyn. Yeah—

Corey: Cloudyn. That’s right. CloudDYN, CloudDYNE. Cloud… something.

Dwayne: Yeah. Which, as I recall, was an Israeli firm, which had a really good story to tell; really strong product. So, the cost management platform existed prior to that, but I think it really became much better once they built the Cloudyn code base into what they were offering, and I think they've also enhanced it by leveraging Power BI because they're definitely eating their own dog food as the saying goes. And what that gives you the ability to do, which is similar to Cost Explorer on AWS, but in my experience, looking at the two, it's a bit more logical and a bit better organized. It still lacks a few things that are business sensitive, but it does give you a tremendous amount of data.

Your metrics, where the cost is coming from, and so forth. Tagging is key, obviously. A lot of organizations are not doing that properly. But what it gives you on the Azure portal is a really lovely business intelligence interface that at a glance can tell you exactly where your costs are coming from. And it builds upon some of the compliance and the management tooling that I talked about earlier, like Azure Management groups, which is an organizational container for your subscriptions and resources that are contained within subscriptions. You can point cost analysis towards aggregated resources if they're properly tagged. And particularly, you can get very, very granular data about where your spend is coming from, and you can create views at the portal for say, the finance officer who can then log into the portal, and get his or her view of what's going on.

There's ways of enhancing still further. You can use the Consumption API to get that data in a way that maybe is more to your liking or more to your needs than the reports that are available from the portal, but the fact that Microsoft has even thought about this, I think is rather impressive because what I'm seeing is that a lot of organizations are not thinking about cost, or rather, they are thinking about cost, but they're not sure what to do about it. And they're not aware of the linkage between architecture and cost. They're not aware or rather, they're not taking advantage of tagging. They're not doing a whole bunch of things that you can do—not using budgets.

There's a lot of things that they're not doing, and I think that it's providing fodder for some of the people who are still trying to sell you boxes on-premises because they're saying, “Oh, my God, your cloud bill is so high, and there's nothing you can do about it.” I'm always frustrated whenever you can article, like, on LinkedIn, or some other place in the tech press. Someone saying we repatriated our stuff from the Cloud, because the cost was too high, and no one ever digs into a couple of things like, well, was the spend generating revenue, and what's the tie to that? But also, what did this company do to try to get a whole other spend that they simply built, as I've seen, like, “Hey, wait, we have 3000 VMs on-premises. Let's build 3000 VMs in Azure.” And, “Oh my god, [laughs], the cost is crazy.” Like, well, yeah. It certainly can be a problem if you didn't refactor to the greatest extent possible, even after you lifted and shifted.

So, I think that as in other areas, what Azure is offering to customers is what I'm going to call a business-friendly way to get an understanding of where your costs are coming from. And even though the recommendation engine, which is a bit of ML—machine learning—some recommendations right there in the portal, about how you can save some money. They’re not always good, because it's an algorithm, obviously, but at least there's a kind of, an understanding of what your patterns are and how you can save money. And this is another area in which I think that, as I said, I think Microsoft is doing a very good job of talking to the business audience; they're talking to developers; they’re talking to infrastructure people, but they're also talking to the business audience, and it's that third part of a tripod, I think that is missing, often, from their competitors is that there’s those conversations with business people as to what the business actually needs and how the technology can help.

Corey: I think that's something that gets lost quite easily when you see engineers talking about cost. I belabored that point to death on this podcast before, to the point where it's not necessarily worth revisiting. But I think you're right. There's a very strong story here around Microsoft, specifically, having a much more mature view of how—I don't want to say legacy companies because that's unfair—but historical companies who have existing and exhaustive physical estates and enormous investments in how they do things that you don't see in a company that didn't exist 20 years ago. That is actually one of their strongest assets, and I think it's easy to dismiss that unfairly.

Dwayne: Yeah, I think that does happen. And what I think is happening amongst partisans for cloud providers is I think we're talking past each other because, I mean, as I say to my colleagues, what is this technology for? What is it for? It's not for you to do cool stuff. It’s to get something done. And that something done might be cool. It might be landing a probe on Mars. That's very cool. Or it might be helping your grandmother get something that she needs. I mean, that's also very cool, and also very necessary. There's a million other things that the world needs that technology should be helping us achieve.

And so, your code could be sweet and all that could be good, but if you're not accomplishing anything with it, what's the point? And I think Microsoft today—in the past, I think there was some stumbling and some problems—but under—certainly under Nadella, I think that Microsoft understands how to have that conversation with their customers. And they keep building things that make it very easy for me to advocate for the platform because I'm not advocating for a platform that you can't use to solve your actual existing problems. It doesn't mean it's perfect, it doesn't mean that I don't get frustrated. It doesn't mean that some things are not in preview for way too long, and I'm like, “It's been a year. Why is this still in preview?” It doesn't mean that you don't encounter problems when you try to deploy stuff. But it does mean that you know what the direction is. And the direction is to help you get stuff done. And I think that's very compelling.

And that's why I think some of these enterprises that I've worked with tend to be large scale organizations that are building things themselves. And they're not keen on showing off a technology. They're keen on getting things done, and they've made their decision about what platform will best help them do that. Although there is a multi-cloud story as well in which you're using not multiple platforms for the same application, but different platforms for different things, which I think we've seen a bit of that as well.

Corey: I think that's probably a good place to leave it. It winds up being a convoluted complex story, and I don't think that migrating to any cloud provider from any other cloud provider is going to materially change the complexity of the billing problem. This is a systemic problem. It is not a provider problem, for better or worse.

Dwayne: That's exactly right. That’s exactly right. Yeah, I 100 percent agree with that, but it's also—speaking to my fellow techies out there—it's also a perceptual problem. When you design something, cost should be on your mind. In the Azure case, there’s a pricing calculator; there's all these tools available to you; there's the API. You can say to your organization, “I have architected the following solution. And I can provide an estimate of how this design costs versus that design,” always keeping foremost in mind what the goal is.

So, one of the things that frustrates me is when your organization is surprised by their bills, because I'm like, well, you could have predicted that if you were paying attention, and it's now kind of everybody's job to pay attention. I mean, to return to Amazon for a moment, the CCOE—which I think is an Amazon creation, the Cloud Center of Excellence—I think that idea is fantastic. I think it should be implemented. I think one of the things that a CCOE would do for organizations is help keep everybody's mind focused on what you want to do and how much it’s going to cost, and also whether or not it helps you achieve your goals. And if you spent 100 million, but you made a billion because of that, then okay, that's fine, but if you spent 100 million and you're not making 100 million because of that infrastructure spend, then you've got a problem.

And you as a technologist should be able to predict—if not what the eventual product will make, at least what the runtime costs of the thing that you're proposing is, and also monitor and remediate. That's one of the beauties of the Cloud. You can change, you can change your architecture if you do things correctly. So, yeah, if anything, I would say to my colleagues, billing is not boring. It's a signal of what you're doing, and it's a signal to you to do it better.

Corey: I think that's probably one of the most astute things that we've heard today. Do things better. But it tends to be one of those simple bits of advice that people don't tend to take nearly seriously enough. Dwayne, thank you so much for taking the time to speak with me today. If people want to hear more about what you have to say, where can they find you?

Dwayne: So, I'm on Twitter. Twitter at @cloudquistador. And also I'm on LinkedIn. I'm Roberto Dwayne Monroe. That's my middle name comes first, and find me there. And I also have a blog that I have to revitalize called The Azure Cmdlet Project in which I talk about things Azure.

Corey: Excellent. Thank you very much for taking the time to speak with me today. I appreciate it. Dwayne Monroe, Senior Cloud Architect at Cloudreach. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts. If you've hated this podcast, please leave a five-star review in Apple Podcasts and a comment explaining your reasoning in triplicate for our procurement department’s consideration.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Wes Miller

Wes Miller analyzes and writes about Azure infrastructure services, including Azure Virtual Machines and Azure Active Directory, and Microsoft systems management technologies.

Before joining Directions on Microsoft, Wes was a product manager and development manager for several Austin, TX, start-ups, including Winternals Software, acquired by

Microsoft in 2006. Prior to that, Wes spent seven years at Microsoft working as a program manager in the Windows Core Operating System and MSN divisions.

Wes received a B.A. in psychology from the University of Alaska Fairbanks.

Links

  • https://twitter.com/getwired
  • https://www.directionsonmicrosoft.com/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is brought to you by DigitalOcean, the cloud provider that makes it easy for startups to deploy and scale modern web applications with, and this is important to me, no billing surprises. With simple, predictable pricing that’s flat across 12 global data center regions and UX developers around the world love, you can control your cloud infrastructure costs and have more time for your team to focus on growing your business. See what businesses are building on DigitalOcean and get started for free at do.co/screaming. That’s D-O-Dot-C-O-slash-screaming and my thanks to DigitalOcean for their continuing support of this ridiculous podcast.

Corey: This episode is sponsored by a personal favorite: Retool. Retool allows you to build fully functional tools for your business in hours, not days or weeks. No front end frameworks to figure out or access controls to manage; just ship the tools that will move your business forward fast. Okay, let's talk about what this really is. It's Visual Basic for interfaces. Say I needed a tool to, I don't know, assemble a whole bunch of links into a weekly sarcastic newsletter that I send to everyone. I can drag various components onto a canvas: buttons, checkboxes, tables, etc. Then I can wire all of those things up to queries with all kinds of different parameters, post, get, put, delete, etc. It all connects to virtually every database natively, or you can do what I did, and build a whole crap ton of lambda functions, shove them behind some API’s gateway and use that instead. It speaks MySQL, Postgres, Dynamo—not Route 53 in a notable oversight; but nothing's perfect. Any given component then lets me tell it which query to run when I invoke it. Then it lets me wire up all of those disparate APIs into sensible interfaces. And I don't know frontend; that's the most important part here: Retool is transformational for those of us who aren't front end types. It unlocks a capability I didn't have until I found this product. I honestly haven't been this enthusiastic about a tool for a long time. Sure they're sponsoring this, but I'm also a customer and a super happy one at that. Learn more and try it for free at retool.com/lastweekinaws. That's retool.com/lastweekinaws, and tell them Corey sent you because they are about to be hearing way more from me.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Wes Miller, who is a research analyst at the interestingly named company, Directions on Microsoft. Wes, welcome to the show.

Wes: Thank you for having me, Corey.

Corey: Now, let's begin with the fact that I have absolutely no leg to stand on because I wound up once starting a newsletter called Last Week in AWS, and I've been dealing with the low-grade number of people who seem to think that I work for AWS, ever since. So, what is Directions on Microsoft, and how did you get there?

Wes: Sure. So, Directions on Microsoft is actually remarkably old. It’s, I think, 28 years old, And it was started by Rob Horwitz, in 1992. He started as a developer, and he went to business school, when he came back from business school, worked a little bit in marketing. And he just, sort of, came up with this realization that what developers were telling customers the software could do, and what marketing was telling the customers the software could do, neither one was really true.

And so he wanted to split the middle. And he came up with this idea of building a newsletter, and he prototyped it internally. Threw it out and saw how many subscribers would be interested. Most of this was within Microsoft itself, and he got enough people interested that he basically self-bootstrapped it with a buddy that he'd met at business school, and we've been around ever since.

And it actually started as Microsoft Directions and at the request of a certain local company, the name of the company changed, but it's actually been Directions on Microsoft almost since the beginning. And what we do is really explain to Microsoft's largest customers, partners, vendors, all sorts of people, what the company does, ideally where the roadmaps look like and help people in Microsoft’s sphere make decisions.

Corey: Excellent. I once had a Last Week in AWS affiliation on a badge at an AWS conference, and I show up; the people at the badging booth looked at Last Week at AWS presumably assumed that I worked there as the director of, “Take this job and shove it. My last day is Friday,” and gave me an employee lanyard, which was hilarious, but it's not the same level of confusion as from you, where people you talk about licensing if people believe that you are a Microsoft employee talking definitively about how licensing works, you'd almost be speaking ex cathedra, to some extent. It seems like the failure mode there is perilous.

Wes: Definitely. I think the reality is, a lot of people—first of all, a lot of people have never heard of us because—I wouldn't say we're boutique, but we're a small company. We're only interesting if you're a customer of a certain size in Microsoft's world.

Corey: What is that size, give or take?

Wes: Well, I can't pin down a number. But realistically, if you don't have, like, five people in your org, who are focused on Microsoft as a key part of their day job, it's not going to make sense. But you reach a certain point where if you're spending hundreds of thousands or millions of dollars per year, you'll realize, “Hey, it might be interesting to have an ally on my side who actually spends their day just focused on understanding the intricacies of Microsoft software,” And, as you and I've discussed, licensing.

Corey: So, one of my origin stories that I don't usually go into, but before I became a grumpy Unix systems administrator, which aged me 40 years overnight, I was a Windows admin for internal desktop support style stuff at a bunch of companies. And one of the things that drove me away from that early on was the joy of helping a company through a Microsoft licensing audit. And at the time, I had problems with authority because some things never change, and the problem that I experienced was I already had what felt to me at the time, like a very difficult job of making sure that all the computers kept working. But in addition to that, I had to sit down and become more or less an accountant and keep track of all the varying license arrangements and the rest.

And I found this awful enough that it drove me into the world of Unix and Linux, and I, sort of, never went back after that. Looking back now, I marveled at how naive I was because we're talking basically a few Small Business Windows 2003 Servers, and maybe 20 or so desktop computers, and that's it. I didn't have problems back then. It felt like I had problems I didn't. At scale, this becomes an absolutely massive approach, but one thing that hasn't changed for me is that sense of why am I playing slap and tickle licensing agreement deals when there are actual, does the system do what I needed to do in a capability perspective? It always bothered me, and it felt like unnecessary busywork. Am I alone in that?

Wes: No, I think there's a certain point that a lot of businesses hit with if we go talk to a member of the press, or we talk to someone in the general public, and we explain what we do, and you say, “Well, I talk to customers who are spending—” and you show them your significant numbers of zeros, and they're just in awe. But these companies don't spend that money—I don’t want to say they don’t spend money carelessly, but they don't spend it trivially. They spend it because Windows Office and this infrastructure that Microsoft has created has become a key component of these businesses. And so the reality is that you keep spending that money in order to get the latest version of it to stay supported, to stay secure, and occasionally get some features, but in general, all that's being pulled towards services rather than software. But as a whole, I think you're absolutely right.

The frustrating part of about this, both for me before I started here—I mean, I started working with SQL Server licensing in the last millennium, and it was not fun then and it's only gotten more complicated since then. For people who aren't in the business, the best way to visualize it is if you take accounting and chess and you combine them at high velocity, that's what software licensing is. It's what you get with Microsoft. It's what you get with Oracle. It's what you get with IBM. It's what you get with Corel, etcetera, etcetera. They all do the same thing. How can I make my software a revenue source that I can keep tweaking and modifying, tweaking and modifying and grow into new businesses, and keep my revenue going in a positive trend line? It’s what everybody does.

Corey: So, one of the interesting pieces that I'm seeing is whenever I talk to customers about their AWS bills, and oh, do I see enormous AWS bills, optimizing them becomes an exercise in planning, and prediction, in architecture. But licensing usually doesn't enter into it. The closest thing in AWS land to licensing concerns with multi-year concerns is things like commitments for spend, in varying ways: reserved instances, and savings plans just being two of them, but there are a lot of others. So, what I'm trying to wrap my head around those, whenever I have glanced into Azure bills, with customers who also have spent over in that side, the first time that happens, “Sure, I'd be thrilled to take a look and give you some thoughts.” And it was so intricately tied to custom enterprise agreements, there was license portability between what they had in their on-prem environment, and what they had in various cloud environments, the Software Assurance stuff added a whole nother level of complexity.

And I quickly realized that A) I have no idea what I'm doing, and I can't responsibly respond to this, so I don't talk about it. Secondly, the problem that I was seeing was that there was so much complexity here that if I were to give advice on this without having either an attorney or someone who is very well versed in the intricacies of Microsoft licensing on my side, that there was going to be a great chance of me getting sued, and rightfully so. And thirdly, even if we do have someone like that on our side who gives absolutely correct advice, if Microsoft comes in and says, “No, your interpretation is wrong, you owe us more money instead.” Even if I'm completely right, we're going to be spending huge money litigating that if they decide to force the issue. Is my understanding of that correct? Am I basically sitting here fear-mongering without realizing it? I know that FUD has usually been the area that Microsoft liked to focus on circa 1996, but we're long past that. Am I spreading my own, now?

Wes: No, I think there's a lot of truth to it. So, for example, my colleague and I, Rob Horwitz, Rob spends a lot of his time focusing on programs like EA and other, basically, volume licensing programs. As I spent my day job focusing on technology around identity, and security, and systems management. When I talk about licensing, usually I'm talking about technologies like SQL Server, or services like Office 365 or, as you mentioned, Azure, which is this weird—I don't know the best way to describe it. It's just this cube, this complicated cube, which I'll circle back to.

But I think what you've said is really important. You know, obviously, I'm pretty active on Twitter, and I answer questions related to licensing. Some of them are really gross, some of them are actually pretty trivial. The problem is that it's really easy, as we can see if you go look on Reddit, for licensing answers, any person who has a concept of what the answer could be will come in and chime in and say, “Oh, no, no, no, that's wrong. All you need to do is…” and beware of anybody who ever says, “All you need to do is…” because usually, that's wrong. The answer is always the most expensive. Whenever you're looking at whether it's Microsoft, IBM, Oracle, very rarely is the, “Oh, I think I can do this and save a little bit.” That's the thing that will come back and bite you.

And I think you're exactly right with trying to analyze—this is where I think you and I have these different universes, and it's like comparing Marvel and DC in some way because you've got this whole sphere you're used to, and AWS I look at from the side, and I think I get these pieces, but there's a lot that I don't get. And then Microsoft, you're absolutely right that what the company has done—and we can look at a lot of the revenue growth and scratch that and say, “Well, that's very interesting,” But a lot of it is taking their on-premises advantage, and then using that to go towards the cloud. How can I get these Enterprise Agreement customers to spend a little bit on Azure, spend more on Azure? And a fair amount of this is back-scratching. In many ways, what Microsoft will do, it's not outright, but what you'll see is a little bit of discounting here in order for Azure spend there. So, it becomes a Gordian knot if you actually were to try and untie it because understanding what's Azure, what's volume licensing, and what's the spend inside of Azure, there's so many things moving at the same time, it's very hard to actually process and untie the savings.

Corey: And that's what gets strange for me from an analyst perspective—which is what I call myself when I want no one to know what I actually do, and cloud economist for whatever reason doesn't seem to fit the bill in that particular conversation—but we saw that Microsoft, at least as of the time of this recording, does not break out Azure revenues as a line item in it’s earnings reports. But they give percentage growth numbers, which is effectively useless in order to figure out how much it's actually making. Turns out it's super easy to get huge growth numbers on small numbers versus big numbers. So, first, it tells me that the numbers are probably something they don't want to disclose for a variety of reasons, but okay, great.

It does make me wonder how much of Azure’s growth is effectively financial engineering of people signing agreements that have a portability benefit to them, expiring credits that are included that wind up being booked as revenue despite the fact that no workloads are being used that consume those credits. It really makes me wonder. Now, I do want to caveat this with I don't believe that the actual dollars going into any of the big cloud providers actually matters. If you're picking any one of the big three, you're not making a bad decision. I want to be very clear on that. So, who's further ahead and by how much? It's navel-gazing.

Wes: Absolutely. And I think especially when we're talking about number one, and number two, which no matter what anybody says to me, AWS is number one, and Azure is number two—

Corey: Fully agree.

Wes:—and it's because you have a first-mover advantage to begin with, but also AWS focused in on services first because that's where the company started. It was just one giant service. So, sort of stepping back and taking a look at Microsoft. One thing that I like to do when I'm looking at moves the company makes—and I do tend to apply chess metaphors to a lot of what the company does—everything operates on a triennium basis. It's a three-year basis because that's what an EA is in most cases. It can be longer in certain cases, but usually it's three years.

And so, when you sign that EA, let's say, just for interesting scenarios, let's say June of 2020 because that's when a lot of them come up: at the end of Microsoft's fiscal year. You signed that, and it wraps back around until June of 2023. And so what are you agreeing to in that? You're agreeing to a bunch of things. You're agreeing usually to how widely you'll use Office, Windows, Client, System Center, maybe Microsoft Servers, Microsoft System Center, etcetera.

So, you got all these things tied up in there, and the problem comes in when you start bringing in things like let's take two products, in particular, SQL Server and Windows Server because there's now benefits—they are literally called the Azure Hybrid Benefit for, and then insert product name. So, Azure Hybrid Benefit for Windows Server. And for Datacenter Edition of Windows server, you actually get a certain number of cores worth of Windows Server that you can run in Azure at no cost. It's not for free: it's actually a debit off of the cost you'd normally pay for pay as you go. So, it's a complex set of machinations.

When you look at it, it's very interesting because it's an on-premises benefit of the software you license that's in Azure. And so the question I've always had is, “Okay, so if you have the rights to that, and you're always running it, and even if you didn't go in and check the box to say, use these cores, where's that accounting happening?” Because technically Windows Server sits in the Azure area. Your point is really valid that because you can't untangle it, and untwine it from Azure versus on-premises stuff, it's really hard to know what was sold, and what was spent, and who's using what versus, “Hey, we made a number go up.”

Corey: If you're like me, one of your favorite hobbies is screwing up CI/CD. Consider instead looking at CircleCI. Designed for modern software teams, CircleCI’s continuous integration and delivery platform helps developers push code with undeserved confidence. Companies of all shapes and sizes use CircleCI to take their software from bad idea to worse delivery, but do so quickly, safely, and at scale.

Visit circle.ci/screaming to learn why high-performing DevOps teams use CircleCI to automate and accelerate their CI/CD pipelines. Alternately, the best advertisement I can think of for CircleCI is to try to string together AWS’s CodeBuild/Deploy/Pipeline suite of services, but trust me, circle.ci/screaming is going to be a heck of a lot less painful and it's where you're ultimately going to end up anyway.

Thanks again, to CircleCI for their support of this ridiculous podcast.

Corey: That is always the big challenge with any cloud provider, to be very clear. That, even without the license shenanigans, what's happening in the AWS space is at any non-trivially-sized enterprise where we're talking tens or hundreds of millions of dollars a year in spend, it's usually not one person spending all of that up, hopefully. So, it winds up being cross-division, cross-account, cross-functionally. And, as a result, when finance hears that, oh, the bill is 20 percent higher this month, is that the new normal? What does this mean for our projections? By the time that filters through corporate telephone, it translates into you're spending too much money. Stop it.

And the big problem always comes down to it's not how much money, it's how do you attribute it? What caused this? Was it a fleet that was spun up that didn't get turned off? Was it a billing mistake on AWS aside—which never, ever happens, except when it does—or is it something else? What is it that drove that? The visibility into what is driving costs is super challenging. And strong credit were due to AWS, I have a crap ton of problems with how their billing system manifests in the real world, but everything that shows up on your bill, with remarkably few exceptions is you have a thing that is currently running. If you don't like paying for it, simply turn it off, and it goes away from that point forward. It doesn't feel like it works that way in a licensed software world.

Wes: It actually doesn't work that way in the actual licensed software world. And that's where Microsoft, sort of, gets the best of both worlds because we've got the company moving from a world where, again, everything's licensed on a three-year basis, and it doesn't matter whether it's shelfware, whether you buy it and you actually don't deploy it, which happens a lot with, like, Office 2019. We anticipate a lot of customers buying it, but sitting still on it because it's got such a limited support lifespan. And then, when you look at things like Azure, you're right. It's purely a services play and just—as I refer to it—Doctor Whovian pricing. It's really about space and time.

Corey: And the pricing is always bigger on the inside.

Wes: Exactly, exactly. And so if you're not using the thing, if you kill the VM—often Microsoft case if you spin it down and de-provision it correctly, you're not paying for it. You're obviously still paying for space, and space is often the thing we found will surprise customers. Like, “Wow, I totally didn't anticipate that A) that wouldn't go away and B) that we use as much space consumption on things as we actually wind up using.” To me, I think that's the fundamental problem with every cloud.

And I think I've mentioned this to you before, that this thing that I refer to as the cloud paradox, the fact that if somebody comes to you and says, “How much will it cost me to do this thing in the cloud?” And you literally can't tell them. But what you can do is say, give it to me for a month, I'll run in the cloud. And I'll come back on the 31st, 32nd day and tell you how much it costs me to run it for a month. That isn't necessarily a good prediction of what it will cost to run for a year. I need it for a year to tell you that.

And there's this whole problem, and so we wind up with—is exactly like you're saying, these projects that get spun up and spun out, and they're really broad across the organization because it’s not centralized, both in terms of development processes, and actually organization in our case, within the Azure world. People wind up spinning up things, either they forget to spin it down, or grows and become successful—Heaven forbid you actually used what you paid for—then you wind up with the opposite problem, which is how do I get cost control to work? And it's the same problem in Azure, AWS, I have to believe it's the same in Google Cloud, really anybody's cloud because you can't wrap it up and put it in a building, there's no way to actually go in and count it, and make sure that people are using things wisely, and not wasting money unless you really dedicate the resources to doing that. And we're just getting started as practitioners, or as teachers, trying to understand how to teach people to fish in this regard.

Corey: That is one of the biggest challenges. To some extent, there always becomes a question, at some point, of scale, where, when does this become an internalized core competency for any given company? So, one of my, I guess, strange questions for you is what is the value of what Directions on Microsoft does to company that’s spending giant piles of money on Microsoft every year? I mean, at some point, doesn't it make sense to have someone or several someone's whose full-time job is working through the intricacies of licensing?

Wes: Well, I'm obviously biased, but I absolutely say so. I think that the reality is, if you're spending a certain amount of money with any vendor, you should have somebody whose job is—a core competency is understanding the value that a vendor provides you, and really the negative value that a vendor provides you. Here's the tendencies of vendor X to try and extract more revenue from us year by year, so maybe we should be trying to control the spend on that. So, I guess when you look at what we do, there's a couple of things: in particular, focusing in on our update and our roadmap, we attempt to really take what Microsoft is doing, and distill it down so you can, as a reader who doesn't necessarily have time to digest all of these press releases and all of the things that the company is pushing out without press releases, understand really what's relevant to you, what the general direction is—we spend a lot of time on our roadmap to give you an idea of, every six months, here's what Microsoft 365 is going to look like; Azure is going to look like, and what but legacy on-premises software—and that's important: the legacy on-premises software because there's not a lot of love coming on-premises anymore—what are these things look like?

And then the other piece is looking at licensing as a whole. We have our reference set, which is basically Wikipedia for people who have a licensing problem, as disturbing as this may sound. And the intent of all of it is to help you answer a question because we have these weird questions that people come up with, what's the answer? And there isn't a single great resource. And so that's what we've tried—we spent 12 plus years trying to create that, is to answer questions for people. And then, our boot camps are really about two things: teaching you how the rules work, so you can stand a chance of complying with them, and understand where they're going; and also, as you're, usually, an EA comes around, how can you maximize your investment that you're making, and make the best decisions for your business, which may or may not reflect the directions that Microsoft might want you to go.

Corey: And that's part of the challenge, is on some level, it feels like these cloud providers have all had the same problem, and I am sympathetic to this, where they want to push customers in a particular direction. And that doesn't always work because customers are where they are, and having a story of how they can get from where they are to a better place is all well and good, but it takes years, in some cases, for companies to even begin making moves in that direction. So, deprecation cycles become important. Microsoft has been terrific about having support for legacy things. And by legacy, I of course mean it makes money.

Amazon, too, is famous for never turning things off. You can build businesses with assurance on top of virtually anything that they launch. Google is a separate story, but we'll get there down the road. The problem that I see with this is, when you're doing things like that, it's very difficult to get customers to move in a direction that aligns with your business when the strategy shifts. So, strategy shifts take, in some cases, decades at that scale.

Wes: Mm-hm. No, and that's what we see with legacy software is, again, when you look at what software people are running, I will often run into, on Twitter, people will say, “Well, why would a business buy Office 2019 and still be running 2013 or, Heaven forbid, Office 2010?” Because it's a sunk cost, and people don't get paid more money by rolling out new versions of software, necessarily. So, whether we're looking at Office or we're looking at moving from even in the cloud, from technology A to technology B or in Microsoft’s world, region A to region B where if you move to the new region, you can save money. It's crazy, but it's true.

You have to strategically A) know that that's something you could do, B) know that it's something you, maybe, should do, and actually C) both take the initiative yourself, and get buy-in up the executive chain to say, “Why do we need to do this again?” And everybody along the way—no matter which org you look at today, in almost every organization, you're going to hit executives up there saying, “Why are we spending money on this again?” So, you, sort of, lose that will to push against it, and to try and make change. I think that's where we wind up with that stagnancy, that customers will sit still, whether it's on-premises, or they'll chuck it to the cloud, and leave it running in the cloud.

And that's the thing that terrifies me about lift and shift, and Microsoft with their extended security updates for all the 2008 and 2008 R2 servers because you're reinforcing the worst habits. You're telling people, “You'll just take it from on-premises and put it in Azure, and we'll give you free ESUs.” And then people have maybe pulled it into Azure to do that, and they're getting free security for a maximum of three years, but what then? These are businesses that sat still for twelve years, or eight years. Do they have an actual plan to get off this software in the next three years? Because most of the cases, I would bet the answer's no.

Corey: One of the things that really rubbed me the wrong way—and I wonder how much of this is me just being hopelessly naive, is that last year, I was a bit of a champion for Microsoft. They invited me to the Microsoft Build conference—and again, turns out when you invite me, I wind up thinking more favorably towards you as a general rule, just because my standard rule of thumb is not I’m not sure how much I try and separate out editorial from other work and keeping it independent voice, it's very human that when someone does something nice and invites me, that I want to wind up treating them well. And it doesn't hurt that their products were legitimately awesome, and it was fun being able to look at what they're doing. I also have a personal standing rule of if you invite me to a thing that you're doing, and all I can do is trash you, I'm not going to go because that is unpleasant; it's not the brand I try to build. There are rules of snark, and that tends to run in the complete wrong direction.

The challenge that I saw, though, is that they had this terrific story about their transformation. They can never quite come out and say it, but my strong belief toward the end of last year was that this is a new Microsoft. That the Microsoft that I hated, that drove me into using FreeBSD, and Linux, of all things, was dead and gone, and there's a new friendlier, cuddly Microsoft. And I believed that, and then they announced this whole Dedicated Instance licensing change at the end of last year, and I felt a little bit betrayed. First, can you explain what that license change was? And then let's talk about that.

Wes: Sure. And actually, I think it's really important to step all the way back and think about where this came from. The best example to understand is VDI: Virtual Desktop Infrastructure. If we look back years ago—first of all Microsoft's whole world, their whole realm is based on traditionally per-device licensing. And we've actually gotten to the point where cores come into play for servers, but it's really all about licensing the thing that that's going to run the software. And years ago, when Microsoft was selling Windows only on a per device basis, there really was no way to do VDI in the cloud. And Amazon came up with a couple of crafty ways with workspaces, one of which was to run Windows Server as a workstation, which is legally permissible. It's kind of funky, but if it does the job, who cares if you can make the revenue work, and it answers customer questions?

But a lot of customers wanted to run Windows Client, and the deal was that if you wanted to run Windows Client on infrastructure, you had to have the hardware dedicated to you. And Amazon decided—they ran the numbers and figured out okay, well, what we'll do—and several other cloud vendors did the same thing—then what we'll do is we will lease the hardware to you, and it will only be your hardware, and because of the legal structure of that, you can now run Windows Client software on it as if it was your own hardware on-premises. And Microsoft has been pivoting and shifting a lot of the licensing rules over the past several years to honestly to monetize VDI, and we look at the whole Windows Virtual Desktop system they built now, you can see that direction.

And so when we look back at last fall, there was this whole series of dedicated licensing changes. And what Microsoft did was, they laid down a series of rules that said, if you're doing dedicated hosting of Microsoft software, here's the new set of rules. And basically, they said, if you're a major cloud vendor, including Azure, it's no longer available to you. And the including Azure was, sort of, an asterisk because what they had done at the same time was add the Azure hybrid benefit, which didn't give the portability to other clouds, but did give Microsoft a bit of an advantage. So, I think it's important to look at your point that the new Microsoft—a lot of people talk about the new Microsoft. I'm famous for being cynical, as you may know, from following me on Twitter, and I think my challenge has always been that a lot of people are really bullish that the company has changed.

The company hasn't really changed. What they've done is they've become much more strategic about messaging, much more strategic about open-source, and I think in general, a lot of this has been, for the customer, good. They've focused on open-source, they've opened up a lot of things. Not the key jewels, but they've opened up a lot of things. And that's an important piece because if you look at what they have open-source versus what they haven't, you can see the strategy inside of it. And it's the same thing with these dedicated rules, that what they did was they changed the rules in a way that, honestly it disadvantages every other cloud, and it advantages Azure. There's a distinct advantage to going to Azure now for Windows Virtual Desktop for your client, for Microsoft, for Windows Server, or SQL Server for your server operating system or your database. And it’s important because they did something that gives their own cloud a pretty significant advantage, and it's a forced move.

A lot of customers can sit still in AWS, and AWS is working really hard to make sure that customers can continue to run the software as long as they believe it's permissible, and it is permissible for some period of time for some set of software licenses, but as you can see, with the report we published, the rules are very intricate and it affects everything from Windows Client, Windows Server, SQL Server, Office 365 Pro Plus—the Office client that you get as Software as a Service—as well as Office 2019 Professional Plus—the on-premises legacy perpetual software. So, there's all these things in motion. And every time we hold one of our boot camps every two months, we have to now explain to people usually it's, “Hey, we use this on AWS, what does this mean to us?” And I have to even stop and think, “Okay, you're using this variant of that, there. Here's your rules.” So, it added a lot of complications, especially if you're trying to use Microsoft software on AWS as a lot of customers still are, and will for the foreseeable future.

Corey: And that's the thing, too, is that very few people are migrating off of Microsoft for reasons that they used to. I don't get the stories of aggressive BSA audit enforcement the way that I did in the 90s and early 2000s. Is that gone, or is that just that they're getting better at shaping the narrative?

Wes: It's gone in the sense of—I think there are probably still BSA audits, it's not something I hear about. However, the way enterprise agreements work, agreeing to get audited is a part of the equation. So, it’s, by definition, it's there. The other thing that we are getting some sense of is that there's more interest in auditing the cloud, both directly and indirectly. And as a whole—like with some of the infrastructure Microsoft's put in place for SQL Server, and for the extended security updates in particular, in order to do that, you have to give Microsoft a fair amount of insight into what's running. So, in a way, you're giving them the ability to audit the un-auditable. Well, things that at least used to be something you couldn't audit, or audit easily.

Corey: And to be explicitly clear on this, people often lack nuance and context. I am in no circumstances advocating software piracy in any way, shape or form. That is not what this is about. But spending, effectively, person-months proving that you're in compliance when you've made a good faith effort to comply is what drives me nuts. I am not talking about lawbreaking; I am not talking about violating contract terms. I'm talking about people have very difficult jobs. We're all trying to do more with less, and spending all the time proving that your good faith efforts to comply got it right, is what's awful.

Wes: I totally agree, but one of the funny things is—so one of the sessions that I teach at our licensing boot camps, in addition to the Azure session, I teach our SQL Server session. And I always use an analogy from when I was growing up in Montana that we used to have a garden in the backyard. It was split between our house and the backyard neighbors. And every year we would try to cut back this one set of plots that were there so we could grow some strawberries and by, like, mid-summer, it was overgrown with rhubarb. No matter what you did, it was overgrown with rhubarb.

And I actually always start with the same analogy for SQL Server because whether it's SQL, Oracle, or really anybody's database or anybody's office software, it's the same problem. It proliferates like rhubarb. So, if you don't actually have a process of constant consolidations—I was going to say remediation; in some ways it is—but constant consolidation for a lot of your key spend software, you're going to wind up with these problems. And so an audit will come along, and an audit will surprise you, and an audit will be very unpleasant. The other interesting thing that I feel is that a lot of bad licensing, again, it's one thing to understand what you're running, it's a whole other thing to say I'm actually living within the rules that Microsoft, or Corel, etcetera, provide or require. But it's a whole other thing to say we have no idea what software we're running because if you don't know the software you're running, you can't be license compliant.

Corey: Well, not with that attitude. No, I hear you.

Wes: Obviously, yeah. But the problem is, that also means that you can't be secure because if you have no idea what versions of software you're running, who's patching it? Who's maintaining your access control lists? It's a very similar problem, and in fact, I encourage people to have the security team and the asset management team working more closely together because they're actually working on a similar, or at least an interrelated problem.

Corey: That is, I guess, the challenge here. I mean, the last thing I'll say before we wind up wrapping up here, is that I have a lot of sympathy for many of my friends at Microsoft, specifically those who work on the Azure Advocates team. They work tirelessly to tell stories about how Azure can solve business problems, about the capabilities of the platform; they make the community better through their work, and I am in tremendous awe of what they're able to do every week. And when this change came out, one of my comments was that it, more or less, seemed like it turned much of what the team had said, into untruths or lies. And I don't blame the team at all. I guess I'm mostly just mad at myself for believing the transformation narrative without taking a more critical look to it. And I guess I'm curious as to whether or not you've seen anything in this space that would shed light on that.

Wes: You raise a key point, which is that a lot of those Azure advocates, a lot of the people I know as well, a lot of them are really good-hearted people who have the best intentions. And I think that, as a whole, everyone in Microsoft still wants to build software that makes the world more effective at what they do. And when we look at what Azure does, that's still the case. However, there's often this detachment from the people who are telling that marketing story—which Azure advocacy may not feel like marketing, but in some ways, it is because I'm helping you build a solution. That's actually where NT evangelization started, way back in the day.

I think as we look forward, the reality is that people need to maintain that Microsoft is doing a couple of things at the same time, the first of which is, they're trying to make sure that Azure is the platform that people want to go to first. And part of that is by breaking norms that we would have expected in the past: you're just as Welcome to run Linux-based technology as you are Windows technology, but also that the company has levers that they need to push and pull. And at the end of the day, a key metric that the company both measures itself on and gets measured on by investors, and pundits is revenue. And how do you make revenue grow? It's by either inducing more customers or getting existing customers to pay more for either new technology or for the thing they already buy. So, the reality is, it's going to get more expensive to buy everything. It was E3; now it’s E5. For Azure, it was basic, now it's premium. These things will continue to change, and it's one of those things that I think you have to keep in mind that this person talking to you, they're an advocate, but they're not necessarily going to hold your business needs at the same tier that you need to.

Corey: That is probably one of the best ways to frame it. If people want to learn more about what you do, how you do it and hear the wise things that you say when opining on a wide variety of topics, where can they find you?

Wes: Well, they can learn more by going to directionsonmicrosoft.com. So, just directionsonmicrosoft.com—

Corey: Yes. Motto, “We're not Microsoft.”

Wes: We're not Microsoft. Although you can call us for Directions to Microsoft, as people have done in the past. The staff loves that.

Corey: Indeed, it's the Microsoft equivalent to Apple Maps and Google Maps.

Wes: Absolutely. We’ll also tell you whether a Microsoft employee is a good hire or not.

Corey: No, that's Directors on Microsoft.

Wes: Oh, sorry, my bad. So, the other thing you can do if you're really feeling like it, is you can follow me on Twitter, I’m @getwired on Twitter, and I talk a little bit both about licensing, and politics, and a few other things. So, that's also one opportunity.

Corey: Excellent. Well, thank you so much for taking the time to speak with me today. I deeply appreciate it, and I feel like you've shed at least a little bit of light on a very dark and confusing place.

Wes: Thank you, Corey was great to be here.

Corey: Wes Miller, research analyst at Directions on Microsoft. I am Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts. If you've hated this podcast, please leave a five-star review on Apple Podcasts and then prepare for your licensing audit.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About David Boeke

David Boeke is the CTO and VP Services for Turbot. David has 25+ years of experience in IT and is recognized as a transformational leader that has enabled some of the world's largest enterprise organizations to make the transition to public cloud. Prior to joining Turbot, David was the Global Head of Enterprise Architecture and led the cloud transformation for a Fortune 50 life sciences company."

Links

  • Turbot: https://turbot.com/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by David Boeke, CTO, and VP of services for Turbot. David, welcome to the show.

David: Hey, Corey, great to be here. Thanks for having me.

Corey: Of course. So, this is a promoted episode as I tend to do from time to time and as a result, we are talking about Turbot, as opposed to the other approach of, “So, where'd you grow up?” And, “What do you do for school?” Let's talk about the company. Turbot’s interesting to me because it targets an area that is very near and dear to my heart, specifically helping organize cloud resources in a way that you can understand them. But it does way more. What do you folks do?

David: We believe that we're really the first true governance platform. There are a lot of tools out there, there's a lot of things that will audit what you're doing, and give you reporting on kind of what you're not doing well, but we really designed Turbot from the ground up—as a governance platform—as tooling for people who like to build things in the Cloud. And one of the things I always like to tell customers is, if you love cloud-native, you're going to love Turbot because we're kind of multi-cloud-native. We're a single platform that you install, and run, and can discover resources across all of your clouds, and then we have a ton of built-in automations within the tool that help you do interesting things with that data. And once you have a complete CMDB and change history of all of your cloud resources, that's updated in real-time.

The idea of then writing automations against that data and understanding it better, maybe for the purposes of security, maybe for the purposes of finding where there's cost and waste, maybe it's something simple, like tagging resources. There's so much that you can do in that space once you have the data. And that's really what we excel at.

Corey: One of the problems I've always had with the idea of building a tool or getting any tool that purports to solve these problems, is it's not a tool. It, more or less, is a few hundred or thousands of tools. So, when I was digging into Turbot, it’s, oh, you actually have 6000, I think was the number that your marketing pages quote, number of things that can be adjusted or monitored inside of a given cloud environment. Am I roughly accurate with that?

David: Yeah, so there's a, I think, roughly right now about 6000 policies where we're adding a few hundred every week. So, a lot of those based on customer requests, some things that we see, a lot of it driven by the actual acceleration of the Cloud. So, you've got to Amazon, you've got Azure, GCP, all pushing new things, new services, new capabilities for existing services on a daily basis now, so it takes a lot to keep up with that. And then think about that in terms of an enterprise mindset, and how might an enterprise want to use that service, or do something with that service?

Corey: One thing that I find compelling about tooling that solves for this is first, it's not something that's exciting. I've got to be blunt with you. It's just not to most people. I get excited about it. When I talk to people, they look at me like I've skipped a circuit somewhere. The idea of governance and seeing what's going on across your estate is not resonant with people who are building these small apps, and they're just setting out for their first time in the Cloud.

You generally only appreciate the value of this, either after having done it a few times or in the more common case, you are an enterprise who has compliance obligations placed upon you, where understanding what is in the environment is of critical importance. And I've got to be honest, humans suck at this. It's something that you need tooling to solve for you. It's neat to see that there's a vendor out there, in your case, who at least seems to understand aspects of this. Rather than, oh, we built the same dashboard everyone else is, but we're going to compete on better customer service, for example.

David: Right. No, that's totally true. I have led large enterprise IT organizations in the past. One of the things that always stuck in my craw, because early in my career, I went from being the business unit IT guy, where I was building solutions, or purchasing solutions, and deploying them, and configuring them for customer need. And then at some point in my career, I decided that these guys at corporate don't know what they're doing, and they're preventing me from doing what I want to do all the time.

So, I'm going to go up to corporate and I'm going to fix those things. So, I shifted my career and moved into the corporate environment, and then I became like the Borg and started preventing people from doing the things that they want to do because in a large organization, you have a million people spending hundreds of millions of dollars all doing the same thing in different ways. So, I've seen the problem from both sides on the enterprise IT side. One of the things that always stuck with me, and I think very early on when we were originally looking to move to AWS as a platform for large enterprise applications, was we would spend literally hundreds of millions of dollars a year on new IT capabilities, whether that's in the data center deploying disk, deploying virtualized compute, networking equipment, etcetera and—as well—apps, and as soon as you deploy them and put them in place, you would lose them forever. No one could ever find that asset again if it killed them. And then you'd spend hundreds of millions of dollars more implementing something like ServiceNow, so you could put human processes around tracking all those assets and what changed on them, etcetera and the thing that jumped out at me very early on—and really who brought it to my attention is our CEO, Nathan Wallace who used to work with me in that enterprise environment—was that with the Cloud, you have an accessible, you can search and find those resources. It's a programmable environment, and you now have this ability to discover things, but also query your existing environment and see what you have. You don't have to lose those assets anymore. And so, one of the things I'm really proud about Turbot is that we do that discovery, we monitor all the real-time events. I have a customer who has 1000 AWS accounts. And if you imagine the Kinesis Stream of event flow from 1000 AWS accounts of all the changes happening across all those accounts.

Corey: Oh, yeah, it takes your entire microservices architecture and makes it worse.

David: [laughs].

Corey: And let’s not kid ourselves, a lot of AWS services, have trouble even communicating with instances of themselves in other regions, let alone different accounts.

David: Yeah. No, totally. And the idea of now, okay, I have 1000 different accounts, and I have a dashboard that I can go in and I can search for an instance ID. I can search for a bucket name, and instantly find out what that is; Who deployed it? How did it change its configuration over time? What is it currently doing? All of that information is at your fingertips.

And if I had that 15 years ago, as an enterprise IT guy, we would have felt like we were ruling the world. You know, you always go back to your 15-year-old self, when we were working on Apple IIs and things like that, with 16 K of RAM, and now looking at how much power you have on your phone: it's kind of the same thing, which is that this migration to cloud and moving all your workloads there, gives you this real visibility to what you're doing, and how you're spending, and what are your assets that are deployed in the environment. So, I really love that aspect of it.

Corey: One thing that resonates with what you folks do is that it's separated from the typical case in this arena, which I have an unfortunate amount of personal experience with. Hi, we're a new startup and we're going to tell your bank how to properly run your IT department. It's condescending, it sounds like a bunch of people who have no experience in this space who are basically storming in to tell companies they're doing it wrong, and here's what they should do instead, without any conception of what life in that environment is like. And what makes you folks interesting is yeah, if you squint hard enough, you look like that. You're you've been around six years, you're definitely a startup. But the people who started your company have come out of that world. You were at Johnson & Johnson for years yourself, for example.

David: Yeah, and all of our core team, both engineering and on the management side, all came out of large IT, either in the financial industry or in pharmaceuticals and the life sciences industry. And those are two very highly regulated industries that spent a lot of time early thinking about these problems of if you move to the Cloud, how do you do that in a regulated environment? And so, there was actually a ton of really great people out there that were frustrated because their day job was essentially doing PowerPoint presentations to people that really didn't understand what they were talking about. And we have, over time, come out ourselves and said, “Hey, look, we want to do something different. We want to take this experience that we have in doing these large scale cloud transformations, and move that into tooling that we can then accelerate other organizations.” By ourselves, we can only ever make an impact to one organization at the time, but collectively as a team, and building a product, we can have that same impact across, hopefully, hundreds if not thousands of other organizations that are going through that same pain.

Corey: That's partly I think, where a lot of folks get it wrong. The things that work for Netflix, quote, unquote, where they get on stage and talk about these amazing transformative things that they've built. There are a few problems there. One, you folks stream movies, let's not kid ourselves. Two, you have enormous amounts of resources that, to be very blunt, most traditional corporate IT departments do not have, full stop. So, trying to retrofit solutions that work in a radically different culture, into an environment where you don't have the same freedom and you have serious regulatory burdens that may not exist elsewhere, mean that a lot of the first attempts at things like this come across is, to be honest, kind of a sad joke.

David: They do. And I really feel for—you know, a lot of people that are in those situations understand the trade-offs understand a lot of the technical approach that needs to be taken in order to solve these problems. But unfortunately, because their management chain is focused on a different business—they're not focused on a technology business—it's difficult to raise those things in a way that can get you the resources that you need to do it effectively. So, many of our customers are working in a resource-constrained environment, where they have a few internal people that really understand it and get it, but there's no way that over time, they could actually hire and maintain and keep the talent that would be necessary in order to build the tooling that they need.

One of the things that was really intriguing to me—because I've been at re:Invent almost every year since it’s initial inception—is that Capital One—you know, large financial institution—has been a platinum sponsor of re:Invent for a number of years, mainly because they were looking for people to bring into their organization. It's essentially a recruiting mechanism to go find people who love this stuff and that they can bring in because, as a company, even a huge financial institution, a huge life sciences institution like Johnson & Johnson, it's hard to find and attract technical talent that wants to come to a large IT organization, a traditional IT organization, and do this type of work. And I believe that's changing. I do believe that the years and years that a lot of IT departments have spent outsourcing more and more technical work, that is coming full circle. The ability to move workloads to cloud and cloud transformation, along with the transformation of data science and companies seeing the value in technology again, means that IT departments are going to be re-funded business leaders are hearing it from the Mackenzies of the world, etcetera, that they need to reinvest in these areas and build out the technical expertise that they have.

And so, I think there's going to be an opportunity, and over time, we'll see larger internal IT teams going. But I work with customers that have one guy who gets it, and he has no ability to hire any other resources to help him with this problem. And he needs automation and tooling that's already done, so he can roll that out and achieve value. And then we also work with enterprises that have two dozen software developers who write their own guard rails and their own controls on top of Turbot, and deploy those, and manage large fleets of AWS accounts and Azure subscriptions and Google projects. So, it's really a mixed bag right now, but I see more people moving in that direction, especially those that are getting large return from their cloud investments and going more cloud-native in their workload migrations.

Corey: One thing that is a universal truth—and I don't care how big your company is, or small—is when you look at how things are set up from a governance perspective, from a tagging perspective—which is my personal hobbyhorse—it's always the same answer, which is: you haven't been doing it properly, and there's a giant gap between where you should be and where you are today. And I don't know that that gap ever gets closed, but you can always do better than you are. But it leads down the path of the first thing people do when they see these things is feel bad. I hate that model because first, it's not your fault. There's a lot of things that go into this. And secondly, you're never going to get to 100 percent. But automated systems are the only way forward. Telling humans to tag things the right way all the time does not work because we're humans, we suck at doing the same thing over and over repeatedly and consistently. So, how do you approach that? How does that inform how you go to market?

David: Yeah, I love that analogy because it is a microcosm of, kind of, what happens in a real environment, which is that some smart people get in a room early on and say, “Okay, we need to have tagging standards.” And you come out with the tablets of the eight tags that every resource must be tagged with, and you kind of go down that path. And the first thing you do is you just you publish that. You write it up, you send it to everybody, you say, “Okay, by—whatever it is—July 1, all project teams must have all the resources tagged, and we're going to go forward.”

And of course, it doesn't happen. And you do some reporting on that. You pull out data, and then you're frustrated. Like you were saying, you feel bad. The management feels bad that hey, we spent some time, we decided on this, we came to consensus across the organization, we rolled it out, we communicated it, and no one did anything. And no one did anything, not because they don't want to, or whatever. Everyone's got a full stack of work on their plate, and they're keeping their applications running etcetera, and running around and tagging things are probably the least sexy thing to do in that space.

The next phase of that, that I normally see, then people go down the road of, “Okay, well if we can't get people to do this stuff manually, let's automate that tagging process.” And then that sends you down a software development path. We're going to build tooling to do that we're going to build serverless lambdas that are going to run in each region in each account, and look for resources and tag them based on a central repository of metadata that we're keeping etcetera. It's a lot of work, you get all that done, you're super happy. You've got that running across all 25 Amazon regions, and you go to re:Invent, and you sit in the keynote, and then Jassy comes out and goes, “Oh, we're launching a new outpost in Los Angeles. And it has a different ARN than every other region that we have.” And it is kind of like you're constantly in this software development mindset in order to achieve any value there. And, like you said, you're still not feeling good about it.

I think one of the things that Turbot does in that space, or a few things Turbot does in that space is that the automation to tag the resources, but to discover that the resources are created, that tags have changed over time, etcetera, that's all built into the product. So, without having to do anything, just turning us on and running it, all the event model of all those events, all the resources that exist, what their current tags are, etcetera, is flowing in your system and you have a place that you can interrogate that, and report on it across many, many accounts, across subscriptions, across clouds, etcetera. So, all that's in, but I think the other thing that's interesting is the way that we approached our dashboarding, which is, we avoided the trope of the red, green, yellow stoplight. There's so many things that will just tell you, everything is wrong. You have one resource that's out there that isn't tagged and your entire environment is red. Your entire AWS environment is now a stoplight red, and everyone feels bad because of that.

One of the things that we did is we turned all of our dashboard reporting into shades of that. So, our standard view is a bar chart of how many resources are red versus how many resources are green, so you can roll up and instantly see, “Hey, 1 percent of my environment is red and 99 percent of my environment is green. And that's good. And I'm getting better over time, and reducing that.” And I think that's one thing that we can do to kind of help drive that, which is that it becomes more about how do I solve these things over time? How do I get better? How do I get less tickets this week than I had last week? I think that's our key mantra, which is, kill the ticket. Do that, and measure your performance over time.

Corey: It's the idea of continuous improvement. I think that's something that people tend to forget sometimes because you look at this monstrous amount of work. I mean, we talk about Turbot, there are 6000 some odd controls that it can wind up applying or influencing. And great, where do you even start with something like that? The idea of it is almost overwhelming and it’s, great, pick one. It doesn't almost doesn't matter which. Pick something that solves a painful problem you have now, and get started. Again, I tend not to view this through the lens of any particular tool. I've done similar things in the past with—you talked about Capital One, their cloud custodian open-source project that came out of it is, in some ways, able to do some of these things. And that wound up being a decent way to get started, too. Find something, anything, but don't just sit there and feel sorry for yourself. Fix it.

David: Yeah, absolutely. And I think that whole idea of, do one thing, and get that going, and feeling good about it, that is key to the developer mindset. It's like, you're programming something; it's 2 a.m., and you're on your fifth Diet Mountain Dew, or whatever it is, and you get that thing to run, and it passes the unit test that feels good. And you go, “Okay, what's the next? What's the next? What's the next thing?” And I think for the cloud operations team, it needs to be the same thing.

If you look at the entirety of—I have companies that come to me and say, “Hey, I want to enable NIST 853 for my environment,” which is six years of work of constant work and driving there, and you'll never reach that goal line. You need to pick up something, whether it's tagging, whether it's public access, whether it's encryption. Pick one thing; pick one control; roll that out; have that be successful. And then if that's automated, you don't have to think about that anymore. So, I can build that out, I can get all my accounts compliant with that. Occasionally, someone's going to do something off the beam, but those are going to be onesie-twosie.

But if you put that to bed, you can really feel good, and then roll in the next day and start, “Okay, now we're going to start on the next thing. We're going to look at making sure that our route tables are not being changed over time,” or whatever it is. That's the whole point there is that I love that philosophy. And I can almost tell when I work with a new customer, whether they're going to be successful or not based on that mindset. Are they coming in saying, “Hey, look, I've got a laundry list of 300 different controls in a spreadsheet that I want you guys to go implement in the next 90 days,” or am I taking an approach of, “What's the minimum viable product that we're going to deploy here? And then let's improve over time.” It's a real litmus test as to whether you're going to be successful or not.

Corey: One thing that I really appreciate about Turbot is, on the one hand, yes, you are positioning yourself as a multi-cloud offering. My argument has always been that multi-cloud as a strategy is dumb. That said, there are ways to approach it that isn’t, and you tend to offer a solution that works regardless of platform, which from your perspective is genius. One, it meets customers where they are—which is always a good thing—rather than yelling at them that they're doing it wrong. Two, the fact that you are able to service customers on different clouds absolutely broadens the appeal. And three, which is was I really what I want to dig into is, you have solved one of the big problems of building a software product aimed at the Cloud because everyone worries about AWS releasing something at re:Invent.

Andy Jassy gets on stage and, in my fantasies, he will start talking about tagging and how to do governance. In practice, I think that's not the interesting sort of stuff that gets keynote time, and then you don't have a business. But instead, it wouldn't be a feature release that causes you problems, it would have to be thousands of feature releases that most of us are growing old and are pretty much expecting are never going to happen within our lifetimes. And you're across multiple providers as well.

David: Yeah, it's interesting. I think everyone that sells tooling in this space is anxious going into re:Invent time about what's going to be announced or whatever. But it's funny, Turbot has been around for six years. So, we're a, “Startup,” I'll put in quotes, but it's a fairly long time for this space and Turbot as tooling pre-existed before AWS Organizations existed. And so, I had customers at that time telling me oh, AWS is coming out with this thing. It's called Organizations, you guys should be worried, and doing all these things.

And the bottom line is, is that the more the cloud providers deploy, you know, the more services that are created creates more things that the enterprise has to do, and that means that there's more value that we can bring to the table. Now, sometimes things that we were doing—like we had, for one of our big early controls, was the idea of finding unused resources and auto-stopping instances, things like that. So, a big thing some of our early customers were using was the ability for Turbot to, essentially for dev environments, to turn off EC2 instances at the end of the day. So, everyone goes home at six, seven o'clock, and then they basically had Turbot configured to stop those instances at the end of the day, and then start them back up around 7 a.m. before everybody got back in the office, which is a big cost savings.

It's even a cost savings over Spot, in some cases, just depending on what kind of instances you're using. Now, AWS came out with that capability and built it in the last year, or the year before. That's great. Always lean into cloud-native. We're not about abstracting you. Far from the opposite.

I think one of the key things that Turbot does is we say we never want to abstract you from using the Cloud. If you like using the AWS console, if you like using the CLI, if you like using Terraform, or Cloud Formation, or whatever the tooling is that you like to interact with your environment, we don't want to intercept. We don't want to be in the middle of that with you because that's what makes you productive, and what makes your team productive. What we really want to do is give you great tooling on the back end to evaluate those changes that are occurring in real-time, regardless of how they're doing that, and do that in a way that doesn't take away value from the developer, doesn't slow the developer down in any way.

Corey: That's part of the value, where the idea of, oh, there's now a native way to do it. Terrific. There are two problems with this. One, not everyone keeps up with every release. In fact, most people don't know that you can stop these events, and two, great; the fact that you have that as a capability also exposes ways for you to either extend it yourself or leverage it for some of the things that your tooling does, or ultimately have it replaced by that native offering, and that's great because it's not the only thing that your company does. And that's where people tend to get stuck. It’s, great, oh, you're just a thing that turns down my development environments at night. Well, that's kind of a limiting experience and offering that really helps folks out.

I do have a question about that particular use case, though. Often I'll talk to companies who really care about exactly that problem, but then analysis shows that 3 percent of their spend is development environments, and they haven't bought an RI in twelve months, so that doesn't seem to be the biggest area of concern. Secondly, there's always the outliers that seem to cut things like this to pieces, and this probably is a larger concern across the entirety of what you do. It works out super well until one day, there's a big release and people are trying to get something out the door; they're working late for a deployment or whatnot, and then the system turns off their instances. Because well, it's quitting time: go home. And it causes a disruption. How do you handle that?

David: Yeah, it was a use case that we had to worry about early on. I actually had a situation where Turbot was competing with an auto-scaling configuration. And so, the auto-scaling configuration kept launching instances, and Turbot kept terminating them because they weren't encrypted or something. But in terms of starting up instances and starting them and things like that, we took a very Turbot-y approach, I'll say, to that, which is that we only try to stop the instance once. So, if you say stop all instances at 8 p.m. And we stop it at 8 p.m. And you start it back up at 8:15, we don't try and stop it again. So, that was how we landed on that to create some value, but not get in the way, in terms of someone that's trying to do productive work where the automation might be fighting with them.

Corey: And that's, I think, the biggest challenge. People wind up fighting with the automation. At least in the world of cost, the first time that your cost savings attempt causes an outage, you're not usually allowed to try to save money anymore. In the world of governance, where you're trying to make sure that there is something holistic across the board that feeds into compliance programs, “Yeah, it caused a problem, so we're turning it off,” That's not really an option unless you're basically Enron. So, at some point you need to learn to live with it, and building out things that are user-friendly and not user-hostile in that approach is one of the things that I think, to be blunt, differentiates you.

David: Yeah, very much so. Early on, early versions of the product wouldn't even terminate things if Turbot didn't see them spin up and they were older than x minutes, like say, 30 minutes. So, the guardrails that would do destructive things only operate on resources that were newly created, so that we didn't have a situation where an internal bad actor could configure Turbot in a way that would destroy someone's entire multi-account environment and just wipe out everything. Over the last couple years, as cost has become more and more of a factor, we started to encounter use cases—I have a customer who—really cool use case, which is that they give every internal IT employee, I think it's $50 or $100 in AWS credits, and they're allowed to spin up an account, and take whatever training classes that they want, and try out resources, try out automations and things like that: basically a sandbox for each individual. They go out to a website internally, they sign up for it, their manager approves, they get this environment, and it's deployed for them, and they can play around with it. Great use case, I love it. I love the fact that they're giving their employees tools to learn this space, and to do things, and to transform from old skills to new skills.

The problem was, is that those $50 sandboxes, the person would stop using it, and then they would still sit out there and still churn cost. So, we actually built specific controls in for them that would take enforcement action—in this particular instance—on the sandbox environments. And so, when they reached their cost limit of spend, or when they reach a specific time value—90 days, etcetera—Turbot would actually go and start terminating all resources within the account and shutting it down, getting it ready for the next person that needed a sandbox. So, basically reset the account and put it back in the hopper to be available for the next person to come along. That's a very interesting use case. It's something that initially we didn't consider because we were thinking of those enterprise use cases: dev environments, QA environments, etcetera, but that idea of kind of time-limited sandboxes, and doing that is a unique proposition, and we had to do some custom things for that customer in order to achieve it.

Corey: It really is one of those areas that is under-appreciated, the idea of building out sandbox environments. That's one of those things I've been asking for—for AWS—for years. I would love to be able to spin up an account and say, “Great, everything in this account is fine. But after a certain date, I want the whole account terminated,” to solve that problem. And then have a positive pressure style system, where I have to keep explicitly, proactively renewing that account as a developer to continue working on things is the right answer, especially if you look at where this world came out of. When you used to have to wait six weeks to get an environment spun up, and now, because it’s Cloud, you can do it the only four weeks of approvals, once you get that, you're never going to give it up because it's painful to do that. If you can automate that and get out of people's way—you're never going to succeed there by chasing people to turn things off by hand. It has to be done automatically. Building tooling that makes this happen as the default, rather than something that a human has to remember, is the only way forward.

David: Totally. And I don't know if you can hear me smiling on the other side of the microphone here, but two of my really favorite controls that we have are associated with this. So, Turbot has its own built-in CASB solution that is very similar to AWS Single Sign-On, where you can grant someone access to an account, and they have that access; they go to a website, and they log in with their SAML identity, and their two-factor authentication, and then they can just click and get single-button access to any AWS account that they have access to. So, if you're on the security team, and you've got metadata level access across every AWS account in your environment, you have a single page you can go to and just login to it. A lot of people have this: there's been some default out of the box solutions from AWS professional services for a while, and then they rolled out Single Sign-On recently, which is a great product and very much aligned to what we do there as well.

But one of the things that we do that's kind of fun with that is, you can grant access for a specific amount of time. And when I was at Johnson & Johnson, one of the things that was always a pain point for us was periodic review of access. So, your manager has to review whatever access that you have on a yearly basis—some organizations do this every 90 days but, every 90 days, every year, whatever it is, your manager has to go through and review all your access for all systems and re-approve it on that yearly basis. And if you don't do it, you're not meeting Sarbanes–Oxley requirements, and your financial reporting system is deemed insecure.

So, if you think about that use case, there's two ways to do that. One is that you provision all access centrally through a central system, and you record all that, and then you have some sort of website that you have to build that shows the manager all the access that you have in a way that they can evaluate whether it's appropriate for you to still have that or not and click that it's okay to do, etcetera. And it was just, kind of like, a very minor tweak that we did, which is, within Turbot, you can grant access for a specific amount of time. So, you can grant access for 364 days. And so, you know that if an auditor comes in and looks at anyone's access within the environment, every bit of access within the environment has been granted within the last x days because if it's over that, it will have been revoked.

And it's a great kind of hack for not having to do periodic review. If you never grant access over that amount of time—whatever your periodic review time is, so 89 days or 364 days, or whatever it is—you can be assured that your audit-ready 100 percent of the time because there is not going to be anyone in the system that has access that's growing. And in the same way, we do the same thing with our policies. So, of those 6000 policies that we were talking about, if you say, “Hey, I want to enforce encryption on all my S3 buckets, and I want to use this very specific customer-managed key,” regardless of the policy-type setting that you're doing, the setting can be time-limited. So, for that team that wants to try out that brand new AWS service that just came out, but you don't want to give them permanent access to that, you can essentially enable that service for them for 30, 60, 90 days, whatever it might be, and then at the end of that basically Turbot revoke their access, so next time they go to the console and try and use it, they're going to have to re-request extending that period of review time for themselves in order to continue to have that access. So, both those use cases in terms of doing things in a time-limited way, are really great way and hacks to solve some of these problems of, there's not enough people in the world to manually go up and follow up on all these things. But if when you're providing the thing, you're limiting the time frame for that, it just kind of solves that problem.

Corey: Yeah, it feels to me like it's the only sensible way to solve these issues. I do want to thank you for taking so much time out of your day to speak with me. If people want to learn more about Turbot. Where can they find it?

David: Yeah. Best place to go is turbot.com, our website. We have a ton of resources out there, and getting started guides. If you're interested in kicking the tires and getting things going, there's call to actions on that page to sign up for a free trial account, import one of your AWS accounts, hey run the entire CIS benchmark for AWS, Axure, GCP, whatever you want to do, against your account to report out, see how things are working. Maybe better, just do one thing. Maybe, you said, let's pick out one thing and do that within Turbot, and then see if it helps you.

Corey: I think I might actually try that myself. But that is a topic for another time. Thanks once again for taking the time to speak with me.

David: Thanks, Corey really enjoyed it.

Corey: David Boeke, CTO, and VP of services for Turbot. I am Cloud Economist Corey Quinn and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts. Whereas if you didn't enjoy this podcast, please leave a five-star review on Apple Podcasts along with a comment explaining why it's better to do all of your governance and tagging work by hand.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Chris Hill

Chris is a Knoxville, TN native and owner of the podcast production company, HumblePod. In addition to producing podcasts for nationally-recognized thought leaders, Chris is the co-host and producer of the award-winning Our Humble Beer Podcast. He also lectures at the University of Tennessee, where he leads courses on podcasts and marketing. He received his undergraduate degree in business at the University of Tennessee at Chattanooga where he majored in Marketing & Entrepreneurship, and he later received his MBA from King University.

Chris currently serves his community the American Marketing Association in Knoxville, where he is currently the President-Elect. In his spare time, he enjoys hanging out with the local craft beer community, international travel, exploring the great outdoors, and his many creative pursuits.

Links

  • http://www.humblebeerpodcast.com/
  • https://www.humblepod.com/
  • https://twitter.com/christopholies

Transcript

Corey: This episode is sponsored in part by ParkMyCloud, fellow worshipers at the altar of turn-that-[censored]-off.

ParkMyCloud makes it easy for you to ensure you're using public cloud like the utility it's meant to be. Just like water and electricity, you pay for most cloud resources when they're turned on, whether or not you're using them. Just like water and electricity, keep them away from the other computers.

Use ParkMyCloud to automatically identify and eliminate wasted cloud spend from idle, oversized and unnecessary resources. It's easy to use and start reducing your cloud bills. Get started for free at parkmycloud.com/screaming.

Corey: Do you know what your cloud is doing when you're listening to this podcast? Turbot does. Give Turbot permission to record every configuration change, tag every resource, remove security threats, and delete unused, not to mention costly, resources.

With Turbot watching over your cloud you can enjoy my rants more and worry less. With Turbot you can discover everything and remediate anything. Tune in to their upcoming episode on June 24th to learn more.

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week—for a bit of a different episode—specifically, I’m here with Chris Hill, the CEO, and founder of HumblePod, the company that produces this and other lesser podcasts. Chris, welcome to the show.

Chris: Thanks for having me, Corey.

Corey: So, I've been fielding a few questions here and there about how I wind up producing the podcast, and the honest answer that I try and weasel my way out of is, “I don't know.” I invite people to a recording link, much like the one that we're on. I say, “Hey, come on, and let's talk about something vaguely related to the world of Cloud.” Some people—read as white dudes—say yes before I get the full sentence out. Other folks need a little bit more convincing, but by the time I'm inviting someone, I already have a rough idea of the stories I want to tell. I already know what notes I'd like to hit, so I have a rough idea in my mind already. Once I've convinced other people that yes, they do, in fact, have a story that they should tell, that I want to hear, that I want to help them tell, and here's a platform: let's go. I throw a recording link, we record. And then, at the end of it, we say goodbye, and that's the last I really deal with it. The rest is all in your hands. So, to field some of these questions, it seems like you're the right person to bring on to, I guess, answer this reader's choice style of listener mailbag questions.

Chris: Awesome.

Corey: Before we dive into that, tell me a little bit about HumblePod. It doesn't seem like a typical Harvard Business School business case where, “I’m going to start a podcasting company,” is a common request. How did you get where you are?

Chris: So, I started podcasting, well, it goes back several years. If we go all the way back to 2006, I was turned on to a podcast called Letters to America. And Letters to America was a podcast about an ex-American patriot living in Ireland. And where the story gets interesting is that ultimately, the guy doing this podcast moved to the United States. I'm from Knoxville, Tennessee; this guy moved to Nashville. And at the time, I was actually in school in Chattanooga.

And the guy, the host—his name is Jett Loe—he said, “Hey, I'm looking for things to do. I'm looking for people to interview, things to do around Nashville and East Tennessee, and just this whole region that I'm in, what would you suggest?” And as a listener, I was like, “Wow, this is really cool. He's near me, and he wants to talk to people.” So, I pitched him an idea for the show. And to my surprise, he said, “Yes, I'm going to come on. I'm going to come down to Chattanooga, Tennessee, and we're going to go on this adventure in Chattanooga together.”

So, that was my first introduction to podcasting in real life, and the crazy thing was he showed up at my house where I was living, at the time, with a bunch of other guys because I was still in college, and he just burst in the door. It had been raining outside; he was soaked, and he had a recorder in his hand. And he was just, like, talking nonstop, just describing everything around him and everything he was doing. And I was like, “A) this is really cool, and B) holy cow, this is all it takes to podcast, is just a handheld recorder? That's all this guy is doing?” And so it really, really just impressed upon me that anybody could really get into this, and it really excited me for what the future could be. So, that was a really fun experience, and then from there I just kind of caught the bug. And so over the years, I had the opportunity to do another business, actually my first startup. We got a trade for service with a radio station, and we did a radio show on social media marketing—yes, you heard that right—social media marketing on, of course, podcasting—or we did it as a podcast as well, I should say. So, that was really exciting and really fun to do.

And then, from there, it led to, “Okay, well, we've done that, I want to do something else.” That led into me ultimately starting a craft beer podcast with another one of my friends. Again, just as a hobby, just for fun, no intention on really starting it as a business other than, “Hey, if we get some free beer out of this, we can, and I think that'd be really cool.” And so, we ended up getting to the point where even during our first year, we got to be a pretty big craft beer podcast for the region and got to the point where, in East Tennessee—if you're familiar with Tennessee, or the southeast and the craft beer scene here, we ended up in front of Highland Brewing Company, actually, at their corporate headquarters, interviewing the founder, who’s—many people call the godfather of craft beer in the southeast, and he's one of the original people to really bring it, and one of the oldest breweries in the southeast right now. So, we got to interview Oscar Wong, his daughter Leah, and their female head brewer at the time, which was really cool.

So, we got to do all that, and I was just like, “Wow, podcasting has led me here in less than a year. That is crazy.” And then, from there—the name of that podcast was Our Humble Beer Podcast. We started calling it HumblePod for short, and from there, because I had been doing it, and because I've been doing it now—actually as of April, it'll be five years—since I've been doing it for so long, people started asking, “Hey, could you help me create a podcast? Could you help me do this?” And having my background in business, and marketing and everything like that just led me to go, “You know what? Yeah, I could. I could actually help you do this, and I could help you make it better because I understand not just the editing side of it, the technical side, but I also understand some of the marketing, and the promotion side as well.”

So, that's ultimately what led me to decide, hey, look, this podcast business could be a thing. And I started with a few friends, I started doing it pro bono so if I screwed up, nobody was asking for their money back; they couldn't be mad at me. And ended up doing all this, and telling it to your good friend and business partner Mike Julian, going, “Hey, look, I've got a couple people working for me.” And Mike said, “Hey, it'd be really cool if maybe you could do a podcast for me.” And so we talked about what that was, and that led to us starting his podcast, and ultimately led me to meeting you.

So just, kind of, has grown. And from there, it's been just a wild ride this past year as the business has grown. But yeah, I mean my goal isn't just, “Hey, let's make podcasts and have fun doing it.” My goal is to make it as easy and simple on someone who wants to do podcasting as possible. And so, like you said when you were prefacing this, we help you to where you don't really have to do or touch anything after we've got that recording. We take it, we do the transcript, we do the content writing, we do all the editing, we clean it up, we make it sound professional, and we put it out for you, so that you don't have to do anything other than have great conversations with people.

Corey: So, it's an interesting and winding road as far as getting to the point of doing the thing that you do now. And I've got a confession to make on my side in that I was never much of a podcast listener. I was for a few years, and then I started working from home, and it turns out that when I'm not driving anywhere, I don't have the availability to listen to them anymore. I can read sarcastically, quickly, so I don't absorb information auditorily in the same way that I can read it. So, that becomes, at least for me, a bit of a bottleneck. So, I'm not too up on what most quote-unquote “the kids” are doing in the world of podcasts. But I was told by a bunch of folks, not all of whom worked for podcasting companies, that it made sense for me to look into a podcast. So, I assumed that people were going to be right and that I was not my target audience, and I gave it a try. And I learned some interesting things along the way.

Chris: Well, what did you learn, Corey? [laughs].

Corey: [laughs]. Among other things, that when I send out a newsletter and I get something wrong, oh boy, do I hear it. I get emails, I get tweets, I get people responding en masse. For podcasts, I don’t. People do not reply to podcasts. There's no ad for them to click on; many of them don't wind up typing in the code for whatever it is that we're selling them. And I was assuming for a while that oh, I probably forget to turn the microphone on.

Then I started going to conferences and getting mobbed by people who were big podcast listeners. It turns out that the feedback model is radically different. People are, for whatever reason, not likely to respond by email; they're not going to tweet at me, but when I'm there in person, people mention the podcast far more than they do the newsletters, the tweets, the blog posts, the ridiculous stunts I do, etcetera, etcetera. And it's either a different audience, or there's a tremendous sleeping giant out there of folks who consume this stuff but don't feel that they can hit reply in the same way that they could with other media.

Chris: Yeah, yeah. I mean, I definitely notice. It feels like there's a void sometimes when you create a podcast, but at least in my experience, I've noticed you'll be just out at the bar talking to somebody and all of a sudden it's randomly like, “Hey, yeah, you're the podcast guy, right? And you did this.” And all of a sudden people know very intimate details about you that you're like, “How do you know that?” And it's like, “Oh, yeah, I guess I said that on a podcast.” So, they're listening, but yeah, finding that way to get feedback is one of the biggest challenges of podcasting, for sure.

Corey: For me, at least, it was also a lot of learning experience. And again, now that I'm 100 some-odd episodes in, I can admit to the ulterior motive I had when I started. It turns out that if you're trying to sit down and have a conversation with folks who are giants of our industry, “Hey, do you want to grab a cup of coffee?” Doesn't get the same reaction as, “Hey, would you like to appear on my podcast?” And it gets me into conversations with people that, let me be blunt with you, I have no business speaking to. And yet, it works.

It's something that people are glad to be a part of. It's built friendships that I didn't see coming. It wasn't entirely business-driven, but it was also an approach of, get me out there. Get me talking to people who are outside of my bubble. I spent half my day crapping on, I don't know, Google Cloud to pick an example. And then, I'll have VPs from Google on who, first, are not going to crap on their own service, surprise. And secondly, are very willing to have a thoughtful conversation. Many of those conversations have in turn shaped my view on the companies in discussion.

Chris: Yeah. Well, I've always said that podcasting is a bit of a magic wand. Like I was mentioning earlier, when I was doing Humble Beer in that first year, getting to the point where we were in front of Highland Brewing Company was a big deal for us. And the free beer thing, while I joke about it, it would just kind of naturally happen. And I know those are small examples compared to, maybe, what you're describing, but that's a big deal. When you've got a podcast, just saying, “I’ve got a podcast,” makes people go, “Oh, yeah. You want me to be on the show?” Or people start really opening up in a way that they don't for other types of media. If you did the same thing and said, “Hey, I want to come do a video of you.” You might get people a little more hesitant to be involved in that for whatever reason, either they're afraid of being in front of the camera, they're not sure how they'll present themselves, those sorts of things, but for some reason, a podcast seems to be a more open forum for that. And yeah, I've definitely noticed that.

Corey: Oh, I have a face for radio. I'm right there with you.

Chris: Dude, I got my hair cut today for this. I'm not even kidding.

Corey: Excellent, excellent. Did you trim the beard as well?

Chris: Um, not yet. [laughs].

Corey: One of the strangest things for me, when I started down this road, was everyone has freakin’ equipment recommendations, and prejudices, and nonsense. I don't have an ear for this. I am the exact opposite of an audiophile. And back when I started and I was talking with you about what equipment do I get? You fell into the audiophile trap, initially, of, “Well, it depends on…” and you gave me an enormous list of variables. And my answer was, “Assume I have no budget. What should I get?” And the response was, “You don't want to say that,” because it turns out you can spend as much as a house on microphones.

Chris: Oh, yes.

Corey: And that was not something I cared about. But we had to dial it in because on the other end of the spectrum, in the very early days again, you were saying, “Well, this microphone works. This other one's better, but this one saves $20, so it makes sense.” And this is part of a business strategy. This is effectively the marketing arm of my consultancy, like it or not. When people know who I am and hear what I have to say, business follows. We don't have a direct relationship between me going on the podcast and business coming in, but when I talk to existing customers, they first heard of me through the podcast, through the newsletters, through Twitter, and they don't really remember which they encountered first, which is super difficult for attribution.

Chris: Yeah. So, I mean, I think when it comes to good quality equipment, the important thing there is it's a very subtle thing, but especially—you've just kind of proven this out for me as we've built this podcast—has been that quality does matter to people even if they say it doesn’t, and they say you can get away with lower quality gear and things like that. The minute you get into the ins and outs of gear, just having good audio and really just honestly, for those of you all out there thinking, “How am I going to get my podcast started? How am I going to do this?” You don't have to start with the most expensive equipment out there or even the fanciest. Even the stuff that both Corey and I are using today. You can start with something simple and low cost. I still think at the end of the day, with podcasting, it's all about just getting out there and executing. But beyond that, once you get to the place where we're at, where you're looking at advertisers, and you're doing that, having good quality equipment allows you to present yourself at a higher level, more professional way. And I think it's just one of the subtle things that we do that helps bring better quality and drive more interest in the show because how many shows have you listened to where the quality is not good? You might listen for a little while but chances are you're not going to be a long term dedicated listener to a show with low quality for every episode.

Corey: If you're like me, one of your favorite hobbies is screwing up CI/CD. Consider instead looking at CircleCI. Designed for modern software teams, CircleCI’s continuous integration and delivery platform helps developers push code with undeserved confidence. Companies of all shapes and sizes use CircleCI to take their software from bad idea to worse delivery, but do so quickly, safely, and at scale.

Visit circle.ci/screaming to learn why high-performing DevOps teams use CircleCI to automate and accelerate their CI/CD pipelines. Alternately, the best advertisement I can think of for CircleCI is to try to string together AWS’s CodeBuild/Deploy/Pipeline suite of services, but trust me, circle.ci/screaming is going to be a heck of a lot less painful and it's where you're ultimately going to end up anyway.

Thanks again, to CircleCI for their support of this ridiculous podcast.

Corey: It gets worse than that. At one point, I was listening to a podcast somewhat recently, I think it was Federico Viticci was interviewing a Craig Federighi of Apple about something and the audio quality was just absolute garbage, and I could barely stand to listen to it. It was awful, and I don't understand how they let that audio quality out. And then, near the end of the show, I was playing with my podcast player, and it turns out that there's an audio setting for arena that I was listening to in the audio playback. So, it turns out that yeah, with booming echoes and sounding like you're in a stadium, the audio quality is awful; I turned that off and suddenly it was clear as a bell, and I felt like a complete idiot. And of course, I should never have doubted Federico and his amazing work. But my god did that absolutely resonate. I didn't really understand what you were getting at until that terrible experience; entirely self-inflicted.

Chris: Yeah, yeah. And there have been other un-self-inflicted ones. I remember one recently with—what is it—Reid Hoffman, Masters of Scale where he interviewed Bill Gates, and for some reason, I guess it was just a timing thing, but whenever they recorded him, it sounds like their interview for him is on an iPhone and they repeatedly bring that up on episode after episode, and it's just like, for everything that they do, and the quality they have on the show, why this quality? But yeah, it makes a difference for sure.

Corey: Yeah. Other questions people have asked, “If I'm podcasting and my child walks in, do I roll with it or stop the recording? Remove the child in an orderly manner, etcetera.” And the nice thing is, this isn't live. My apologies for those who believed otherwise, I can say, “Oh, hang on a second, remove said child,” and pick up right where I left off. And we just—well, I don't cut that out, you cut that out, and that's what makes this awesome. That's the entire point of having guests on a podcast is people were initially surprised by the fact that this wasn't a sarcastic show, where I dragged my guests. It turns out that unless you're very careful in how you do that and you have guests who can hang with it, you're beating someone up. And when you start down that path, it's very hard to get another guest ever again. And that's not really the person I wanted to be.

Chris: Yeah, that is definitely for sure. You want to make sure that things are clean and professional from an editing standpoint as well as just, definitely very important to make sure the editing is done well and done professionally, too. Because I mean, there is to some level, a degree to which you want people to see behind the curtain; you want to see some of that transparency, but if it's your child running and screaming every five seconds, then you're probably not going to be wanting to continue to listen to a show like that. You want to hear something that's nice, clean, professional.

Corey: Exactly. And, “Sorry, I'm expecting a package, I’m having a dog barking in the background,” is said occasionally on some shows, and you'll never hear it. It's taken out super well. You also have it during a brief moment during the recording, not the entire episode. People will tolerate momentary disruptions, but at some point, you try and create a good experience. Not to mention it's super distracting for guests.

So, another question I've periodically gotten that in fact, I don't have a good answer for, I'm hoping you do. There's a thing in the world of podcasts called dynamic ads where it can apparently, and correct me if I'm wrong on this, insert different advertisements based upon where the listener is located.

Chris: Yeah, so they're getting there. And dynamic ads are pretty cool from the perspective of they're allowed to insert ads into spots on the podcast that are pre-determined by the editor. And for those who want to know the bare bones behind that, basically what you do is you create a marker on the back end of the podcast. And you actually literally just insert that digital marker in the show and then put it up on the podcast host, and the host itself will actually pick that spot and say, “All right, this is where we insert the ad here.”

You see that most commonly on people that do podcasts on Anchor. That's a really common one, and Libsyn is another common one that does dynamically-inserted ads. And you can do it at any level: you don't have to have a billion followers or a million downloads an episode to qualify for these, but a lot of times with dynamic ad insertion, the challenge becomes that you really do need a decent-sized audience if you want to make a living at being a podcaster in that way.

And that goes for a lot of different ads when they're broken up into what I call a CPM ad, which is cost per millia or cost per thousand, meaning that the advertiser is paying you for every 1000 downloads you get to your podcast. And typically that rate is somewhere between $8 and $35 for every thousand downloads that you get. So, you end up in a situation where you do have to have twenty, thirty, forty thousand, hundred thousand, two hundred thousand to really start making serious money and consider podcasting as a career as opposed to just on your own. What we find though—and Corey, I know you guys are great here at Screaming in the Cloud, but listeners probably noticed we're not using dynamically inserted ads in this show.

Corey: No, generally I tend to speak personally about the sponsors that have sponsored the show. And that leads to an excellent question as well. Why don't we sell mattresses on the show?

Chris: Well, that is a great question. We don't sell mattresses, like Casper—they're going to pay us for that later, I guess—but we don't sell mattresses on this show because we are not really focused on being a CPM show. I mean, there's this little thing in marketing called positioning. And it's how you position your business, and where you really say that your focus is. And with every podcast, I consider every podcast to be, really its own independent business unit, especially in your all's case—or for some of my clients, it is their job. For some people, it's just another extension of their brand, and positioning is really all about where you stand in the mind of your customer, or in your listener.

And with positioning for a show like this, you're not trying to reach the masses. I'm not going to have my wife, who's a schoolteacher say, “Hey, listen to this episode of Screaming in the Cloud, I'm sure you'll find it super relevant to your sixth-grade students,” Unless you do one that has to do with history in data cloud storage, they're probably not going to be a listener to that audience. And the advertisers that you want for that show, like a mattress seller like Casper, are not going to be the type of people you really even want to have advertising to that audience. So, what we find in cases like you all is that really a flat fee advertising structure is really better for you guys. So, you say we're going to charge X dollars per episode, one-time flat fee to be an advertiser on the show because you're reaching this niche audience that's a dedicated listener base to me. Like you were talking about and like we've talked about on-air, off-air: you have listeners, you have fans that come up to you and talk to you and engage, and they're all at the conferences you go to and all where you are. So, you're an influencer, if you will, in that space and I think we—

Corey: I prefer the term Quinnfluencer but I’ll take what I can get.

Chris: Quinnfluencer, yes, yes sorry you’re a Quinnfluencer. I will use that from here on out. But as the Quinnfluencer, I think you're the—I saw somewhere—the number one cloud influencer online. I don't know how verifiable that is, but that's awesome. But that said, having that level of clout, being able to influence like that, being able to say that people listen to me when I talk is a big deal. And people do listen to you when you talk, and people are responding, and are engaging with you in that way.

And so you have an audience there that someone who's just talking about video games, or just talking about politics, or sports isn't going to have, and it doesn't reach as broad of an audience, but it's still a very important audience. And I think that's really why you all don't sell mattresses.

Corey: Attribution is always one of the strangest parts of podcast advertising because when you think about it, people are listening to this show when they're commuting, they're listening when their hands are full, they're usually not going to stop and punch in a URL for something, but it is brand awareness. I had a sponsor, who I'm not going to name, that in the early days, sponsored a bunch of episodes and then decided not to renew—this is normal, incidentally. No one sponsors the same thing for years on end in most cases—and what was atypical was that they came running back about three or four months later, trying to buy out everything we'd sell them, and, “Well, this is interesting and welcome, but may I ask what changed?” It turned out that they'd spoken to a couple of their larger customers and hearing about the product on this show was what in turn inspired them to start trying it out. And, “Oh, yeah, that's right, there's that thing I heard about,” and it turned into a business deal.

The hard part is it's very challenging to tie listenership or any particular ad back to specific customer actions. There's this idea of effective frequency that you have to see an advertisement or a brand something like 20 times before you buy. And we have this attribution problem across all of advertising, where the last thing that they saw that they click on and buy the thing gets all of the credit and everything that happened before that point was money wasted. It never actually works that way; we find this as much more akin to brand advertising and getting awareness out there.

It's challenging to get people to sign up for things on a podcast. It's challenging to get people to punch in a discount code. I had one sponsor very early on, they wanted me to read a very long URL, including a bunch of tracking parameters. No one in the world is going to click on that. That's why I built the snark.cloud redirector for some of these things because people want to figure out how well is this message resonating? Are people taking action based upon what it is that they're hearing? And the answer is they are, but it doesn't always take the form that folks would expect. This is the dark art of demand generation, really, is understanding how what you're doing interacts with any particular audience. Yeah, if we were using CPMs and effectively going after mattress sales, we would make far, far, far less money than the show cost to produce. But in this case, because of who the audience is, what the listenership is comprised of and what the topics are, it instead turns into a surprisingly lucrative endeavor, even though that wasn't the initial intent when I set out to do this.

Chris: That's a very, very good point is that it takes multiple touches for people to become customers. The sales process—any sales cycle is never a one time, close deal, you're done. I mean, that can happen when you get to a certain point in your business, but they've typically either done their research beforehand, they've looked at other options, they've talked to other people, and then they've made their decision. It's never a—you're introduced to it, you immediately buy it with something like the products that are advertised and marketed on this show. And so it's a very long, long process.

And also, what you're building through this show, too, is essentially a long tail for the content that you're creating as well, because these ads will live on, essentially, in perpetuity. Until this show or these episodes are no longer hosted on iTunes, anybody can go back and say, listen to an old episode with Jessie Frazelle and go and listen to that from months and months and months ago, and it'll still have those same advertisers on there that were there when you originally recorded it. So, there's a lot of advantage to that in that regard, too because, I mean, I just saw this on the stats for our Humble Beer just in the past day or so where I had somebody listening to episode number one. I had three listens to episode number one just in the past few days. And I'm like, that was five years ago. People still want to go back and listen to that content, and I'm sure people are listening to your content as you build up that library too. So, there's a lot of opportunity for advertising in that way, and it grows, and it opens up opportunities for developing those opportunities after the fact even after that episode initially airs.

Corey: I will also admit freely that I am not particularly up to speed on the business of running this podcast, let alone others. I have you doing all of the production pieces where I just record, I yap into a microphone for half an hour and then throw it over the wall. But on the editorial side, which is kind of neat, we have a full time employee, Caroline, who winds up handling the sales for all the sponsorship stuff. So, very often I will not know—for example, right now I have no idea at the time of this recording which is, sponsors are going to be sponsoring this episode, so as a result, I don't need to worry about what I say, how I say it. The same story with the AWS Morning Brief, my other podcast, as well as the newsletter. The sponsorship things are the last piece of all of it that I do, so I don't have to worry, necessarily, about whether what I say is going to upset a sponsor.

It might; it hasn't happened yet. I'm sure it's inevitable, but I don't tie the two things together in such a way that I wind up having to apologize for it. And I don't let who is paying for any given episode influence what happens. The only restriction I put out there is that if you are sponsoring this podcast, you also can't have a guest from your company on unless it's a promoted episode because, at that point, it just becomes super weird. Conversely, we also aren't going to let you put a sponsor in that's your direct competitor, because that's not great. This is somehow challenging with AWS because they compete with basically everyone and everything, but they've smiled, shrugged, and come to the acceptance phase of bargaining.

Chris: Yeah. And I think that's a really good place to be with yourself as an advertiser, because, and especially by just articulating that to your audience, I think what you're showing too is, “Hey, look, we're not being manipulated,” if you will. Not that you would be, but you're not being influenced by your advertisers to say things or to do things. You're still maintaining that level of autonomy. And I think that's important, especially when you're in the position you all are in, talking about the Cloud and talking about Amazon, and things like that, that you're not doing it from an internal company, “I’m paid to say these things,” perspective. So, I think that's actually a really good way to have it structured.

Corey: The challenge, of course, is always making sure that we're not ambushing people. When I wind up doing a sponsored podcast with a guest that is a sponsored guest—which we call promoted episodes—I’m always going through the same process with them than I am with everyone else, which is, what do we want to talk about? What do we want to make sure that we don't hit on because that's going to be overly sensitive? And what questions can I ask that give you the opportunity to tell a story you don't get to tell very often? It gives people an opportunity to shine light from a different angle than a lot of other folks do. I don't have any particular agenda when I sit down with someone and start having a conversation here. My goal is to have a good conversation and keep it entertaining. Anything else and I get bored and zone out, which is a question someone asked me on Twitter about podcasting techniques. How do I keep myself focused? I have good guests, and I entertain folks that are going to have a conversation that keeps everyone engaged. Without that, it's just not a very good episode.

Chris: Absolutely. You need to stay engaged, and stay active, and listen actively, I think is one of the important things for those, again, out there that might be listening to this going, oh, I'd like to start a podcast myself. Active listening, taking notes, doing things like that, staying engaged with the guest is important. And of course, having interesting guests is good too. So, yeah, I totally agree there.

And by the way, just a side note, you mentioned having an agenda when you talk to somebody. I had one episode where we actually had an advertiser for my personal podcast that we were doing, and they had asked us to ask us one very specific question of an advertiser, and I won't go into details, but it did not go well. So, I would advise staying away from trying to have an agenda when you interview someone unless that agenda is very much agreed to before you start the interview.

Corey: And making sure that you wind up setting the expectations and then living up to them is critical. I think without that we wouldn't have the audience, we wouldn’t have the guests that we do, and we wouldn't have the voice that we've accidentally stumbled upon and built for ourselves.

Chris: But it's a great voice.

Corey: Well, thank you. So, if people want to learn more about your thoughts, generally probably as related to podcasting and/or craft beer, where can they find you?

Chris: You can find me at HumblePod.com. That’s the easiest place to find out everything about me. You can find out the other shows we edit. You can find out a little bit more about the services we offer, and by the time this goes live, there’ll also be some stuff up about gear and equipment. You'll even be able to see what Corey uses on his podcast as part of that, as well. So, that's where I would go to check things out. If you want to find me personally, I'm on Twitter at @christopholies. And yeah, that's really about it.

Corey: Chris, thank you so much for taking the time to field the slings and arrows of Twitter on this podcast. It's appreciated. And as always, thank you for the fine work that you and your team do as far as making the nonsense that comes out of my mouth borderline intelligible.

Chris: Alrighty, well, thank you, Corey, and it's a pleasure editing the show for you.

Corey: Chris Hill, CEO of HumblePod. I'm Cloud Economist Corey Quinn fixing AWS bills in San Francisco, or wherever I happen to find them. And this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts. If you've hated this podcast, it's probably the editing and Chris's fault, but please leave a five-star review on Apple Podcasts regardless.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Miles Ward

As Chief Technology Officer at SADA, Miles Ward leads SADA’s cloud strategy and solutions capabilities. His remit includes delivering next-generation solutions to challenges in big data and analytics, application migration, infrastructure automation, and cost optimization; reinforcing our engineering culture; and engaging with customers on their most complex and ambitious plans around Google Cloud.

Previously, Miles served as Director and Global Lead for Solutions at Google Cloud. He founded the Google Cloud’s Solutions Architecture practice, launched hundreds of solutions, built Style-Detection and Hummus AI APIs, built CloudHero, designed the pricing and TCO calculators, and helped thousands of customers like Twitter who migrated the world’s largest Hadoop cluster to public cloud and Audi USA who replatformed to k8s before it was out of alpha, and helped Banco Itau design the intercloud architecture for the bank of the future.

Before Google, Miles helped build the AWS Solutions Architecture team. He wrote the first AWS Well Architected framework, proposed Trusted Advisor and the Snowmobile, invented GameDay, worked as a core part of the Obama for America 2012 “tech” team, helped NASA stream the Curiosity Mars Rover landing, and rebooted Skype in a pinch.

Earning his Bachelors of Science in Rhetoric and Media Studies from Willamette University, Miles is a three-time technology startup entrepreneur who also plays a mean electric sousaphone.

Links

  • SADA.com
  • LinkedIn: https://www.linkedin.com/in/milesward/
  • Twitter: https://twitter.com/milesward
  • Cloud N Clear Podcast: https://sada.com/insights-events/video-podcasts/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is brought to you by DigitalOcean, the cloud provider that makes it easy for startups to deploy and scale modern web applications with, and this is important to me, no billing surprises. With simple, predictable pricing that’s flat across 12 global data center regions and UX developers around the world love, you can control your cloud infrastructure costs and have more time for your team to focus on growing your business. See what businesses are building on DigitalOcean and get started for free at do.co/screaming. That’s D-O-Dot-C-O-slash-screaming and my thanks to DigitalOcean for their continuing support of this ridiculous podcast.

Corey: This episode is brought to you by Spot.io, the continuous cloud cost optimization platform, saving businesses millions of dollars each year on their cloud bills used by some of the world's largest enterprises and fastest growing startups like Intel and Samsung. Those are enterprises and duo lingo. That's a startup. Spot.io delivers the optimal balance of cost and performance by leveraging spot instances, reserve capacity, and on demand. Give your workloads the infrastructure they deserve. Always available, always scalable, and always at the lowest possible cost. Visit Spot.io to learn more.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by the CTO of SADA, Miles Ward. Miles, welcome to the show.

Miles: Hey, thank you so much, Corey. I'm super excited to be here.

Corey: So, until today, my previous interaction with you has largely been busting my chops anytime I say something unflattering about GCP. Now, in your defense, I tend to phrase things in the most obnoxious way possible. But a little background on you first. You were at GCP for a while; did an awful lot of things there. Before that you were at AWS and launched the first version of the Well-Architected Framework, better known as the Well-Actually Framework. But now you're over, doing CTO style work at what I believe to be the largest Google partner out there, if I'm not mistaken.

Miles: That's exactly the gig. You know, I helped build the solutions architecture practice at AWS. I was the fifth member of that team, that's now well over 1000 folks. And when I came to Google, I came to build solutions architecture inside GCP.

But in both of those roles, there was always this little gap at the end where customers actually go use those solutions, and implement them, and run production on top of them, and make promises to their customers based on them, and because of the requirements of large scale companies, there's just always a friction between really being deeply accountable to a customer and frankly, interacting with them in those complex scenarios. So, I took a role as CTO here at SADA to plug into customers and do the last mile to deliver the solutions that I had been spending a decade designing.

Corey: So, let’s—I guess, rather than looking into the ancient history story, because again, any cloud provider 10 years ago looks very different than it does today—

Miles: Yes.

Corey: —the modern landscape is very interesting to me. When I'm talking to customers—and my perspective is one of, believe it or not, being relatively impartial. I focused on AWS because that is where my customers tend to live. But we do see an ever-growing constituency representing Azure, representing GCP, Oracle—if you read Oracle press releases—and the entire space is growing.

So, trying to race clouds against one another is, I guess, an interesting hobby, but doesn't seem to get you very far. At least, that's my personal perspective. So, I guess my first question for you is this: multi-cloud/hybrid-cloud; is it a thing, or is it just something analysts make up as fan-fiction?

Miles: So, I think there's a couple layers there. First—I mean, it's Andy Jassy—he was at one of the re:Invent conferences—who described that the biggest impediment to AWS growth was the lack of a viable competitor. And so, I took that as a bit of a personal challenge, and I'm happy to see the degree to which GCP certainly takes individual customers and has great growth, but I think is firmly cemented as one of the companies that you can evaluate when you're thinking about an infrastructure that you don't have to manage yourself.

You know, I think the concept of hybrid is something that's been screwed up by everybody involved. The idea that all of us are not persistently in a hybrid operations mode is bizarre, really. I know of no company, anywhere that consumes the entirety of its technology infrastructure from a single vendor. I actually think such a thing as impossible. How do you get a phone from Microsoft, or a desktop operating system from our friends in Amazon? You need all of the building blocks that are involved.

So, that we are in hybrid all the time and will be in hybrid forever, suggests that, for our friends in the procurement department, the adults who are responsible for the care and feeding of all this, sort of, fun toys that we get to play with, you don't have to be a double-masters in procurement to have gotten the, probably, first line of the first page of the book that talks about best practices for procurement, which is you shall have more than one vendor for things on which you are critically dependent. So, I think that kind of thinking drives a bunch of behavior from customers who are eager to be able to run in more than one environment, to play these providers off each other, to retain power in the negotiation, and to do so at a cost that's not totally exorbitant and insane. Because you can imagine if you were doing this all from scratch—by hand—on your own, you implement against three different SDKs; you learn effectively three different environments, nuances, and all the details of their products. That's just going to be particularly complicated. So, we're watching businesses work really hard to reduce that impediment to the best practice, or that impediment to what the business people in the building require.

Corey: And that is, I think, an interesting story. I view, at least in the hybrid space where, “We're hybrid-cloud,” means that we started doing a full-on cloud migration, realized halfway through it's super hard to migrate some workloads, gave up, planted a flag, declared victory, and now we're hybrid. Multi-cloud feels a bit different, in that everyone wants to have a different story as you go down that particular path. So, a clear question: you were at Google for a while, after having done an awful lot of AWS. And again, nothing good can stay at Google because it gets deprecated. Why did you leave? Or you weren’t—you didn’t so much leave, as you were Google Reader-ed?

Miles: [laughs], no quite the opposite. I received what I can only consider a totally shocking retention offer from the Google people. So, I wasn't Reader-ed. I had a great time there, and the people that I worked with I love very deeply and I say that in the most human and personal way possible. There are a lot of people that will be friends of mine for life, and it was an incredibly hard decision to leave, but this requirement, right? I mean, you cannot learn more about what you're doing if you don't actually do it. Hands-on is the way for me.

I suppose you could do synthetic research about the effectiveness of individual solutions writ large across whole sets of customers, but I just knew we were missing data about this last step, where we actually go out and implement them and hold customers hands and onboard them to the details and bear the risks together with them for those deployments. So, a big driver for leaving was really needing to be involved in that way, hands-on, with the deployment of customers.

Another big driver, another important thing was frankly, I have seen a bit of this movie before. I was at AWS, I helped put together the Partner Programs, the systems integrator and reseller programs at Amazon, together with Dorothy, who's rad. And I watched a bunch of those partners capture big markets and do incredible things as Amazon passed these sort of operational thresholds: they got to double-digit billions in revenue, and they started to do a regional sales model, and they got to capacity management problems because all of a sudden they're getting customer demand that they didn't expect in different areas, they started to have bunches of regions available so you can do a really global deployment.

Google has crossed all those same thresholds that Amazon crossed, in 2014. So, being able to participate at that scale, I thought there—was really clear that the work that was going to happen in systems integrators was going to be some of the most interesting work of this generation, or this phase of the expansion of the Google environment. And then, the last area, the place where, I think, some really hard thinking to do was to unpack what kind of gig I wanted to have. I had started to manage. The solutions architecture team on the Google side had gotten over 80 people, which doesn't sound very big in comparison to AWS. But think of it more like the—

Corey: It's both larger than two pizzas, no matter how you slice them.

Miles: It is substantially larger than two pizzas. It is also more than one calibration meeting, and the performance management, and personnel management load had become fairly high. So, I'm a dweeb, and I wanted to do stuff with my hands, and it was much easier for me to do that when I wasn't also managing 80 incredibly smart, challenging, hard-working people.

Corey: So, this does lead to an actual question: is solutions architecture the same thing between the two companies—or basically everywhere? Does SA work differ based upon the culture in which you find yourself in? That does sound like a leading question because I can't imagine the answer being no, but tell me about it.

Miles: Oh, sure. So, we were trying to figure out what to call it on the Amazon side, and we had this meeting and it was Rudy Valdez who—his suggestion was solutions architecture because he had met some Oracle folks that had that title, and he thought that just sounded cool, was basically as much aggressive thought that went into it, which is a little wild for however many thousands of people do that role now in the public cloud context. The guide for those, and I think the dimensions on which engineers consider a role that has as much customer-facing time and as much of a communication requirement as I think all solutions architecture gigs share, is where you sit in the org, relative to things like quota or not. Are you really in the sales org, or are you, sort of, an overlay that doesn't bear the individual deal responsibility? Another is how much access or interaction you have with product and product engineering. Are you an advisor to product management? Have folks moved from solutions architecture into product management and vice versa? How is that balanced struck? And then, many orgs differentiate—Google certainly thinks of product management as different than the core technical leadership for an individual product, and so how much access do you have to the—like, literally, to the software developers that are building the services that you're out representing? These things change rapidly enough that it's really critical, I think, to be able to make a positive impact on customers that you have that kind of access.

So, one of the reasons that I pushed very hard in the structural definition for solutions architecture on the Google side was to resolve several problems that I thought made solutions architectures choose between recognition and high performance inside the AWS business, and a successful outcome for their customers, or the very best recommendations to them. If you bear quota, if you have any kind of tie to this individual customer's outcome, I think it can be hard to always carry the highroad with every employee about telling them the right thing to do, as opposed to maybe the slightly more expensive way of doing a given thing. It's also very difficult, I think, to step back and think about higher-level recommendations or higher-level structures that you might need to build to enable every one of the solutions architects to be well.

So, early on, as we were typing up this, what was really a checklist for Reddit, to make it so that they would get the heck out of a single zone, and stop going offline when Amazon did exactly what it promised to Reddit that it would do, by having zones go offline. The Well-Architected Framework was one of the attempts to try and help all the solutions architects do well. Not just to individually succeed with the customers that I was working with. And so, I wanted to build a program on the Google side that was deeply oriented in that way, where you took more responsibility over scaled impacts over multi-participant positive benefit, as opposed to the individual successes with individual customers because those come and go.

Corey: You definitely sound like you have a deep and abiding knowledge and perspective on customers, but let's do a quick spot check here. A lot of people like to claim that they worked at Google. But that often just sounds to me like one of those lies you tell on your resume. So, let's do a spot check here. If you really worked at Google, what product or service did you kill?

Miles: [laughs], yes, that's true. If you're going to be in a leadership position over any substantial period of time, and you don't slay something, can you really stick Google on your resume? I posit that perhaps you can't right? I mean, there's just an incredible number of services that have been turned off, which, the only number that is bigger, obviously, because, well, that's a tautology. The number of services that have been created has to be a lot bigger than that. So—

Corey: But only by a small margin.

Miles: [laughs]. Right. Necessarily, but only by a small margin. So, I spent a bunch of cycles with marketing and PR leadership, as well as with product management and senior executive leadership, speaking very specifically to this issue. That when we, in the public's eye, seemingly arbitrarily, turn off products, or services, or features, or capabilities, or worse, change them after the fact, or change them as they're expanding and growing, we erode public trust, we make us seem unprepared for the real world expectations of enterprises and major customers. I know that only because of the personal feedback I received when I, yes, did turn off a Google product. And I think it's probably worth exploring what that looked like, and my decision-making as it went into it because I think in many, many cases, it's the same kind of analysis that other product managers and product owners are making as they decide to pull that trigger on the Google side. Is that worth taking the time?

Corey: I think so.

Miles: Okay. So, the product is called the TCO calculator, or the Total Cost of Ownership.

Corey: Oh, yes, every cloud provider has one of these, and it's always very polite fiction—

Miles: Yeah, oh no—

Corey: —because, it’s just, what do you include? What do you not? What story do you want to tell? Tell me the conclusion I can get your data to back it up.

Miles: Well, the story we wanted to tell was that Amazon is hideously expensive and that Google is better. And so, in order to be able to tell that story, we had to be able to unpack the real-world pricing differences, just the raw unit cost differences between the two environments, but we also had to show the differences in pricing model. So, for example, AWS has this thing called a Reserved Instance, you may be familiar with those. Reserved instances fix values like which operating system you're running, or which individual instance family you've chosen, or which zone in which region that you intend to consume. Google's model has—

Corey: Since supplanted by savings plans; same discount, none of the restrictions.

Miles: You, which actually solved a whole bunch of these issues at the time—

Corey: Oh, god, yes. I’d been complaining about the Reserved Instance limitations for years. I’m glad it finally got fixed. Credit where due. I want to call out when they have fixed a thing.

Miles: Yeah, and the savings plans are dramatically better. I think there remains some advantages in the simpler parts of Google's model. One example of those being sustained usage discounts, where without any action on your part, if there is any way for us to infer that you have consumed any amount of virtual machine resources that can be considered one thing—for the purposes of calculations—over any more than 25 percent of a month, you are automatically getting a discount. You cannot opt out, you cannot click the wrong button, you literally can't screw it up. And that, as a difference from AWS, was one that was very difficult for modelers to include in their analysis.

So, the TCO tool did include it and was able to incorporate that, as well as things like the committed use discounts, the higher throughput networking, a bunch of other building blocks that were advantaged on the Google side. But that's not the story I'm trying to tell. I'm trying to tell the story about killing the thing. So, we had built it, to be totally honest, in a mad rush to respond to Gartner feedback that said you don't have a real cloud if you don't have one of these. I blame that squarely on very smart folks like you who are working very hard on this pricing analysis stuff, and all of the great customers who really do want to understand these problems as much as there is always going to be assumption and fictionalization and madness built into those calculations. So, we built it in a week. Literally Urs Hölzle and Ben Traynor—Urs is the de facto CTO that SVP for—

Corey: Oh, we’ve spoken once or twice on Twitter. He loves calling me out when I go a step too far, which, please keep doing this. That's what I'm there for. If no one calls you out, have you gone too far is always the question.

Miles: That's right. That's right. Urs, I have an incredible amount of respect for the impact on the planet that Urs has had. He is an incredible person. And then, Treynor is no small part of that. He is the head of operations for all of Alphabet. So, his LinkedIn says, “If Google's down, it's my fault.” And the two of them, to have them working in a Google Sheet with me is a little [laughs] bit of a terrifying afternoon, and you really double check your multiplication.

But we were able to pull together a model that they found viable, and that several of the customers we were working with thought was useful, and get that out and published on the web, and then connected to Google’s—the cloud.google.com website for exposure to customers. And that all went great for about four and a half months until which point, as both Amazon and Google had changed some of their pricing. So, I said, “Hey, I think I'm going to make some updates and changes to this thing.” And because it had now been online for a while, it was a part of the production management and review cycles for all normal products on the Google side. I was like, “Great, I would love input and feedback. Help me make sure that I'm doing this stuff the right way.”

And the core of that was a requirement for the operations story for this product. And I was like, “Oh, that's no problem. I will get up in the middle of the night and I will track when there's been a change in pricing, and then I will adjust the model myself to make sure that it reflects reality, and then I'll publish the changes.” The whole app sits on App Engine, so there's no operations of any kind because App Engine is great, and everything will work slick.”

And they're like, “Oh, no, no, no, that's, that doesn't work at all. You need to be staffed for a full-time member of the site reliability engineering organization, to observe and manage that product, and to ensure that your releases happen in a timely manner, to hold things on the right stead.” I was like, “Awesome. So, you assign that person?” They're like, “No, no, they come out of your team. You have to staff that person out of your resourcing.” I was like, “Well, then that's me. I'm the person who's the SRE.” And they go, “Well, no, it has to be a different person than the developer who's in charge of the business.” I was like, “Well, I'm the only one on my team,” at the time, so—

Corey: Oh, yes. These are common problems. And the challenge is, is that a lot of them seem to start from how things have to fit into a matrix internally rather than what is best for a customer in this story. Now, there are reasons that a company would do things that suit internal needs, but the customers a) don't care and b) don't love it when whatever it is that gets implemented deviates from a successful outcome to their eyes.

Miles: Oh, yeah. No, I mean, I really see how a customer who had set up a meeting with their boss to walk through the decision-making for buying GCP versus something else and was expecting to go to cloud.google.com/products/tco and see this tool and then have the thing not show up. So, I wrote long-form papers describing why it is this was valuable, and here's our customer traffic, and this is the sort of support and responsibility that I need from a shared group to be able to participate in this. And I'm pretty persuasive, and so it ended up living for about a year based purely on my ability to cajole others into supporting this piece that we had onboarded.

But as new leadership and new people came through, I wasn't persuasive enough. And so, I really get how, as a product owner, there's this balancing act and granularity between, do we only take bets that we can sustainably resource for 1000 years, or are we allowed to take bets that we don't have a clear worldview on how we would sustainably resource for a thousand years? Because all of Google is a bet that cannot be sustainably resourced for 1000 years. That's what the company is. It is a super dynamic driven business looking at the market, identifying new ways to be able to organize and make useful the world's information. And so, I don't think it has the kind of institutional will that would be required and, frankly, the cost overheads and the resource commitments that would be required to be able to take the thousands and thousands of experiments—and it absolutely thinks of them as experiments internally, which because Google is so big, become things that businesses depend on.

Corey: You'll notice the word experiment never appears in marketing dialogue.

Miles: My product had a big asterisk at the bottom that said, “We are providing this as a service of our communications and technical teams. We reserve the right to take it offline as pricing changes or other requirements adjust.” So, I don't know if I quite called it an experiment, but I certainly described the fact to which that it was provided as a best-case scenario kind of a basis.

Corey: This episode is sponsored in part by ChaosSearch. Now their name isn’t in all caps, so they’re definitely worth talking to. What is ChaosSearch? A scalable log analysis service that lets you add new workloads in minutes, not days or weeks. Click. Boom. Done. ChaosSearch is for you if you’re trying to get a handle on processing multiple terabytes, or more, of log and event data per day, at a disruptive price. One more thing, for those of you that have been down this path of disappointment before, ChaosSearch is a fully managed solution that isn’t playing marketing games when they say “fully managed.” The data lives within your S3 buckets, and that’s really all you have to care about. No managing of servers, but also no data movement. Check them out at chaossearch.io and tell them Corey sent you. Watch for the wince when you say my name. That’s chaossearch.io.

Corey: True, but you've been at both shops. I mean, AWS could be said to have all of these same constraints around being forced to sustain whatever it is that they release indefinitely, but to their credit they have. You don't see deprecations, effectively, ever on the AWS side of the house. Which is why it's interesting because both GCP and AWS have the same language around deprecation timelines in their terms and conditions. No one ever brings it up in a serious context around AWS. They do constantly with GCP, and it's no small part based to people's experience on the consumer side of the house with Google deprecating things that are beloved. Reader, I am still avenging you.

Miles: [laughs]. Yeah, no. And I think I think it's important for customers to do the balancing act analysis, to say, “Okay, I'm interacting with a provider that has a different worldview than AWS, or Oracle, or IBM or any of these other competitors.” Google is a different company. There are some really positive benefits of that worldview. It is experimenting on your behalf a lot more, in my view, person to person or dollar per dollar of revenue, than those other businesses are.

Then the downsides are that I think that it has, by way of policy, absolutely structured itself in a way that means that more of those experiments get turned off. We think there is not nearly enough work done today to capture the feedback and interest from customers who would characterize those experiments as critical where right now, there's not nearly enough of that sort of color or anecdote or communication with customers to be able to get that final detail that says, “I need to have this. This thing is now a critical part of my workflow.” If they had more of that view, I think, it would be easier for them to make the case internally.

I know really quickly, I ran to be able to provide basically a user feedback form as a part of the pricing—or in the TCO calculator, because I wanted all of those anecdotes to be able to identify for my stakeholders, no really, people are in this thing all the time and using it. I can show you the traffic and you won't care about that because it's a lot smaller than YouTube, but you certainly should care about which kinds of customers are in there and participating. So, I think Google could go further, to spend more time thinking in that way, and push product managers harder to pay very close attention to that. I think it's worth a million dollars, at least, in marketing cost every time they turn a product off at this point, because—

Corey: I would argue more than that, depending on what side of the fence it's on.

Miles: Yeah, no, I agree. I think that there's culpable risk there, and that's a lot less than the cost of an SRE, part-time.

Corey: Oh, yes. Not to belabor this one, but when GKE went from free to just kidding, it's going to start charging per control cluster now. There was a lot of response, correctly so, that $73 a month is not going to break the bank for anyone's Kubernetes cluster. But first, there's the problem of trying to dictate to customers, well, it's fine, not that's not a big amount of money, you'll get used to it. “Well excuse you, who are you to tell me that?” Secondly, it's the moving of the goalposts where when you were building things out and your cost model was this thing is going to be free, and the workload on it would be costing us, suddenly it changes how that is perceived.

And given that Google already has a strange reputation for changing the deal after people are using things and loving them, it makes for a very dangerous question for a lot of the Big E enterprise, folks. If, for example, Microsoft Excel changed the way that they wound up running formulas at random and pushed it out to some subset of their users, they would be burned alive out of their offices by people who are; you don't change anything we're doing a Big E enterprise without minimum 18 months of roadmap notice, and even that it's tight for some shops. It's a different culture than exists internally at Google and exists with a lot of the born-in-the-cloud startups. And if that is where Google is serious about competing, it needs to be able to assuage those customer concerns, even if they're not directly articulated to them.

Miles: Yeah, I think there is a balance there. I think it does need to get better, and be serious about the places where there is a real impact to customers in the changes that they make. And not just on the impression of it, but the actual outcomes for those customers. I work together—the one I'm thinking of is in Google Maps, they made a substantial change in the pricing structure for maps, and there's a bunch of customers for whom that's just, sort of—they take it as a shot across the bow, and they, sort of, evaluate if there's some sort of alternative there, and now they have to do this in situ reevaluation of the business structure. And that's, that's not something you should do with thousands to hundreds of thousands of customers, even if it costs a percentage or two on your side, or maybe more than that, to be able to keep things going.

So, I say yes, and sitting as a partner outside of Google, working together with them every day, we advocate and push on that issue all the time. We are one of the signals into them describing the criticality of that issue with our major customer opportunities and the places where we're interacting. I think the far side of that, the other end, though, is big enterprises do have to weigh, how much value they get out of an experimentative culture and what are the kinds of things that maybe they should be working with Google and others to figure out the best practices for interacting with something that is more volatile than they're used to because, I don't know, what I saw was Google was able to wrap itself around and solve really gnarly problems really rapidly because of this dynamic internal behavior. So, I don't want them to lose that for the benefit of keeping a couple of products going a little longer than maybe they otherwise would.

Corey: Absolutely, and this is the challenge too, is that a lot of the reputational risk factors are not themselves directly aimed at things that are quantifiable in the traditional sense. When you wind up trying to decide what cloud provider do we go with, no one realistically, or at least what rounds to no one is going to say, “Not Google, because they turn stuff off.” But it will be used as a bullet point in the list of a larger argument. It shores up the question of would we be able to wind up having these conversations in anything approaching a realistic way if that's how it were to play out with this thing that we cared about? It winds up adding wood to the arrows that people use to take down a certain cloud provider, and it's ground that I don't believe GCP needs to give up as easily as it does.

Miles: Mm-hm, yeah, and I think that's another driver in this overall multi-cloud/hybrid-cloud issue. If you must pick a single provider for a decade, and it will cost you a billion dollars to change providers, these issues are paramount, and you really need to make a good bet. And you really can't afford the kind of complexity or overhead that a switch might maintain. If, on the other hand, these things become increasingly commoditized, it is trivial to move workloads between them.

Corey: Ah, but not data. The Achilles heel of every public cloud is the cost of data transfer out. Well, that’s, sort of, a two-part problem. One is the actual cost. Secondly, the sheer level of complexity and modeling what that's going to be before trying it to see.

Miles: Oh, yeah, as one of the guys who helped design the 800 gigabit a second connectivity from Twitter's primary data centers to Google Cloud, which is, I believe, more than the throughput of the public Internet—

Corey: I don't think I can post nonsense nearly that quickly.

Miles: [laughs], it was an incredible amount of shitposting per millisecond flying back and forth between those two environments. So, the bandwidth is a thing. Dedicated interconnect and direct connect and the private linkages between these different environments and their customers are certainly a way to offset some of that, but I also think it's a model component where Google—following Amazon in structure, bluntly, was really, I think smart in trying to land on a model that, that enabled customers to be able to pick and choose. So, they landed on this egress system called Premium Pricing that lets you use Google's primary network to do all of the interregional transmissions—all that stuff is built in—all that cost more, and then there's the lower cost one if you want to use a dirty internet connection like Amazon would provide to people as where it only comes out of a single regional provider ever. And if that outbound linkage goes down, it gets worse.

Lots of customers that I talked to tolerate in their own environments, even riskier network setups and that's the place where I think none of the clouds have really landed on a great offering is super unreliable, super high latency, weird, but great throughput and low-cost connectivity in and out of their systems. I don't think they do that, as much as I think it is easy to assume, to say, oh, they make the egress prices higher, so you can't take the data out so they can lock you in. I think Amazon landed on that model early, and the others followed them into it because it was the standard structure for this stuff. And it’s—I didn't see either on Amazon or on Google’s side, them talking about, oh, sweet, we've got them all locked in. Now, they can't get their data out of there. I mean, the whole point of putting data anyplace is to be able to use it.

Corey: Ideally. Yeah, the fun part is is that I feel like a lot of the pricing things are outgrowths of how things used to be in people's historical use cases, which means that for some things, it is not a fit for certain modern patterns, and having to play those games and fight it out is challenging. I don't know that there is any, I guess, golden path forward for all folks, I think everyone's use case is different and making sure that every provider has a story that resonates with that person's particular use case is going to be paramount. Whether or not various providers are able to deliver on that really remains to be seen.

Miles: Yeah, I wonder if there's a model available, I'm thinking in the same way, as Google offers sole tenant nodes, or Amazon has whole hosts are bare metal instances, things like that, where you are getting lower down in the stack to make longer-term commitments for bigger chunks of what are the real world building blocks. I mean, as a customer, if I were able to purchase whole connectivity links and provide them to a Google, or an Amazon, or an Azure, or whoever, and be responsible for a bunch of the operational overhead and structural complexity of maintaining and managing that link, I would expect that you'd also then be able to match that up with some lower fixed rate for the internal networking costs that are now currently zero on cloud providers when you move inside of a single zone around that zone.

That stuff's free, but it sure isn't free to build. So, I think there's some kind of a model there where, rather than—I think this is a negative externality—rather than embedding the cost of internal networking in the egress rate, if instead, they were able to pry that back out and allow you to go do the legwork to purchase egress in whatever way you think you're going to do a better job of the Google side or the Amazon side to provide, I think that's a model that would be attractive, especially for customers that are themselves telecommunications companies, or customers that have unreasonable access to this kind of bandwidth.

Corey: I think that's very fair. So, if people want to hear more about what you have to say on this, and many, many other topics, where is the best place they can find you?

Miles: Sure. I'm not a deeply organized person, I think if you interact with me—

Corey: Guilty as charged.

Miles: —oh, yeah. So, I spent a lot of cycles working together with Google marketing. I delivered about half of Google Cloud’s keynote sessions over the course of the last five years, so we're close buddies. And so, SADA is Google’s—literally their partner of the year. So, we work together with them to participate at Google events all over the place. So, if you come to a Google Cloud Platform event, there's a pretty good chance that me, or somebody from my team will be there presenting and answering questions and plugging in.

At SADA.com you'll have a whole running list of the small format events that we do that are closer to customers, that allow people to interact with us one on one, that’s got my travel schedule going. I participate mostly at LinkedIn and Twitter, so both of those are just my full name, Miles Ward, so that's pretty easy to hunt down. And then, we also spend and have on our side, our own version of this same sort of thing. We call it Cloud N Clear. Both the CEO, Tony Safoian, and I have produced sessions of that. And so, maybe the next place for people to hear about us would be when we're interviewing you there. It’d be super fun.

Corey: Well, we'll certainly see if that happens to take place, especially because at the time of this recording, it's questionable whether anyone will ever attend a conference event in person again—

Miles: Yeah.

Corey: —due to a pandemic.

Miles: Yeah, that's right. No, we're actually doing some really interesting work together with Google marketing to think through because they have canceled Google Next in person. And, you know—

Corey: I was going to be there with bells on.

Miles: Oh, yeah, no, I had my multicolored Converse shoes and my electric tuba, ready to go. So, we're working with them now to do a bunch of the emergency creative thought about how do you translate a bunch of this stuff, although we still haven't figured out a way to have the, sort of, hanging out at the lobby bar at two in the morning by way of a Google Hangout. There's just something that doesn't quite translate. Although we have been working quite a lot with some of our partners and customers, like, Domino's is a customer of ours. We're trying to figure out how to ship everybody pizza, and we've been hanging out with DoorDash quite a bit. So, maybe we'll Dash people a bunch of chow while we're on a Hangout or something. We'll see how it all goes.

Corey: Excellent. Sounds like a plan to me. One way or another, we'll make it work. Thanks again for taking the time to speak with me. I appreciate it.

Miles: Thanks, Corey. Good to see you, too.

Corey: Myles Ward, CTO at SADA. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts. If you've hated this podcast, please leave a five-star review on Apple Podcasts and a separate five-star review Google podcasts to embrace a multi-cloud strategy.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Alex DeBrie

Alex is an author and self-employed AWS trainer and consultant focused on serverless & cloud-native technologies. He has been recognized as an AWS Data Hero for his community work with DynamoDB and other database technologies. In a previous job, he worked at Serverless, Inc., creators of the Serverless Framework. If you go even further back in his employment history, you'll see Alex had a brief stint as a corporate lawyer.

Links Referenced

  • The DynamoDB Book: https://www.dynamodbbook.com/
  • The DynamoDB Guide: https://www.dynamodbguide.com/
  • Twitter: https://twitter.com/alexbdebrie
  • Blog: https://www.alexdebrie.com/

Transcript
Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Are you better than the average bear with AWS? If you're listening to this podcast, the answer is almost certainly yes. Want to turn those skills into money? If you're US-based and have an AWS certification, sign up as an expert on AWS IQ today and help customers with their problems. Visit snark.cloud/iq to learn more.

Corey: Do you know what your cloud is doing when you're listening to this podcast? Turbot does. Give Turbot permission to record every configuration change, tag every resource, remove security threats, and delete unused, not to mention costly, resources. With Turbot watching over your cloud you can enjoy my rants more and worry less. With Turbot you can discover everything and remediate anything. Tune in to their upcoming episode on June 24th to learn more.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Alex DeBrie, who's an author and self-employed AWS trainer and consultant, with a focus on serverless and cloud-native technologies. But most importantly, he's the author of the brand sparkling new book called The DynamoDB Book. Alex, welcome to the show.

Alex: Thanks for having me on, Corey.

Corey: No, thanks for taking the time to speak with me. It's nice seeing folks in the community who are putting out books around AWS services, especially some of the more venerable services that are not iterating quite as rapidly as something newly released. It feels like if I were to write a book about anything released in the last three years, it would almost certainly be out of date before I ever got it to my editor, let alone shelves. But DynamoDB has been with us for a while and it seems to be relatively stable. So, let's start at the beginning. What the heck inspired you to write a book about Amazon's second-best database. The first, of course, being Route 53.

Alex: Well, yeah, I knew you already had Route 53 cornered, so I had to choose a different tack there. But really, for me, it was just using DynamoDB, and especially using it incorrectly over a number of years was how I got into DynamoDB. So, I've been using it, probably, for about four years now. And mostly been using it with serverless technologies like AWS Lambda, because it fits really well with Lambda. But modeling with DynamoDB is quite a bit different than modeling with a relational database.

So, for my first two years, I thought I could do it pretty well, and I was mostly modeling in a relational way. And then, two years ago, I watched Rick Houlihan’s re:Invent talk about how to model what they—DynamoDB in, and it kind of blew my mind. As a result of that, then I made a different site called dynamodbguide.com, which was just basically dissecting Rick's talk and taking it from a 600 level talk down to maybe a 2 or 300 level talk, and walking through how you think about DynamoDB, how to use it, what the API is like, things like that. I would say even at that point, I didn't really understand DynamoDB.

But since that time, I've had a lot of work with the community and folks at AWS and things like that, to where I've continued to learn more about it. And I still feel like there's a gap in how to model with DynamoDB, so I wanted to fill that gap and write a book. And really what you were talking about was one of the main considerations, where DynamoDB has stabilized to some extent to where it's not going to be changing quickly to where, if I do release a book, it's not going to be totally irrelevant by the next re:Invent. You might need to tweak around the edges, but I think it should be something that will work for a couple years at least.

Corey: It's always interesting to me to see where things tend to come from, and as it turns out, about I want to say what was it 20—math is hard—in the fall of 2018, is where I first became acquainted with Rick, and as it turns out, I was given the opening keynote of Latency Conf in Perth, Australia, and he was giving the closing keynote. And then, we wound up on a panel, which was interesting, but I had never heard of him before that moment. And oh, some exec from AWS. Great, I'm sure this will be an interesting talk where they go on about leadership principles or something. Yes, it was very much not that. The first half of the talk was awesome. And the second half of the talk, I had to be excused for, because my brain was full.

Alex: It's so true. I mean, he’s just pretty amazing, just in the amount of stuff that he knows and how quickly he conveys it, and then spits it out. But yeah, he's really become, sort of, a cult hero over the last couple years due to his re:Invent talks.

Corey: A lot of different directions we can go with this. But let's start, I guess, with a common one that I've gotten, where I'm a fan of DynamoDB, in that it's a great NoSQL database. You could call it a key-value store, you could call it a bunch of things. I'm calling it a database because I know it offends some of the most pedantic people in the world. But it is a lock-in story.

There's nothing that has Dynamo’s model anywhere else that I'm aware of. So, what is it that makes it, I guess, okay to go for a level of lock-in around a datastore, when we've seen from Oracle what happens when you have a datastore that is locked into one particular proprietary provider? And how that can wind up causing constraints down the road? What makes Dynamo exempt from that? Or isn't it?

Alex: Yep, I think that's a good point, you know, and I think a lot of people are worried about lock-in, and especially the ones that have the battle scars from previous generations like you're saying, I don't have those battle scars, so maybe I'm naive, but the big thing for me, I think is, I think AWS and really, Amazon more broadly, has shown such a commitment to delivering value for the customer and not trying to squeeze them for everything that they're worth. So, you don't see price increases on AWS services. You generally see things get better over time. So, I'm a little less worried about the pricing aspect. And for me I've mostly opted into the AWS Serverless ecosystem, so things like Lambda and API Gateway and all that.

And all that is stuff I couldn't move to another cloud anyway, so I'm going to be rearchitecting my system if I move at that point, anyway. So, I think at that point, it really makes sense to go all-in and use what works best. And I just think the other things you get from DynamoDB are so much clearly better than the other databases out there, whether that's the pricing model, or how well the connection model works with Lambda, or how easily it works with infrastructures, code, all that stuff. That makes it worth it for me to opt into that ecosystem.

Corey: So, I guess one of the big questions from a technical perspective is I built something on top of DynamoDB, how painful is it to migrate that to a different datastore to migrate that to I don't know, Cassandra, for example?

Alex: Yeah, that's a great question. I’ve—

Corey: Or possibly Redis, or possibly Route 53.

Alex: Yeah, exactly. That's a really good question, and as I've been working on the book, and communicating with people about it, the big question I always get is around migrations. And I'd say more of the question has been around migrating within DynamoDB because you usually design your DynamoDB table to work with specific access patterns. And then, so people are asking, “What if my access patterns change? How do I migrate that?” And I have some content in the book around that. In terms of migrating to another database, I think it's going to be pretty similar to any time that you're migrating across different types of databases.

So, if you're migrating from something relational, like MySQL or Postgres, into MongoDB or DynamoDB, you're going to need to do an ETL process that reshapes your data to make it work for the other datastores. So, the same is going to be true for Dynamo, right? There's nothing, really, that has the exact same data model as DynamoDB. Although Cassandra is pretty close. But yeah, really, if you're moving out of DynamoDB to anything else, it's going to involve some sort of ETL process where you're pulling that data out, you're reshaping it in a shape that works for this new database that you're using, and then putting it in there. So, I think it's not really any different—the one thing that it might be harder than, is if you're moving from one relational database to another relational database, that's going to be pretty straightforward. But if you're really moving across data models like that, it's going to be about the same.

Corey: One of the, I guess, questions then becomes, is it better to start treating it as the most naive possible data model that you can come up with, in the event that you want to potentially migrate off of it someday, or once you've made a decision does going all in make sense in a way that it may not have historically with other things?

Alex: I would say go all in. I mean, I'm a big cloud-native advocate, I would say, and really take advantage of the features of your cloud provider and treat your cloud provider as a partner, rather than maybe an adversary or someone to be wary of. So, I would say go all-in with it. You know, I don't think it's going to make it that much easier if you have a very naive DynamoDB implementation as compared to a complex one and you want to take it out. You're still going to be doing the ETL stuff, which is going to be making sure you're writing the right script and all that stuff. It'll be a little easier with the naive model, but really don't see a lot of people actually migrating clouds very often, or even switching databases very often, so I think it makes sense to optimize for the common scenario of—that you're going to be on this database for a while, and making sure you're building something that really works well for it.

Corey: Before you started writing this book, you—or maybe when you did start writing this book, you were at Serverless Framework, which is where I first started encountering you, in that whenever I have trouble with open-source projects, my default support method is to whine about how dumb I am on Twitter, and hope someone takes pity on me. Much to your everlasting chagrin, you did.

Alex: Yep, exactly. Yes, I was at Serverless for about two and a half years. I just left there at the beginning of 2020 here. But yeah, I did a few different things at Serverless. I started on the growth team, first as a data engineer, and then as head of growth, where I was working with some great folks there to just make content, mostly, and engage with the community because I think serverless, and the Serverless Framework made it a lot easier to use AWS. But you still needed to learn quite a bit of AWS to get the value out of that. So, we were working to make that available to folks that were starting to dip their toes into the AWS ecosystem. So, did that for a while, and then actually was on the engineering team at Serverless for about a year, year and a half there, as well. So, yeah, I had a great time there. I really loved working on the Serverless Framework, and just the serverless community, and all that stuff.

Corey: It seems like it's a fantastic way of dipping your toes into something that is very finicky to put together manually. I started using it, and personally, I never went back. It was the sort of thing that just, sort of, made sense and, I guess, aligned with my entire half-baked understanding of a lot of these concepts and it papered over things that I didn't really need to understand super deeply. Something that was similar to that, in some respects, I guess, was the ill-fated two weeks I spent playing around with the Amplify CLI at Amplify Framework—or Amplify library, whatever it is that we're calling it that AWS has released.

That effectively you wind up importing it into a React or View project, and you tell it what you need, it builds a bunch of backend services, and then spits out the JavaScript component that interacts with those services; which, on the one hand, is awesome, and on the other is the exact opposite of what I needed. Because I tend to understand the infrastructure pieces, but not how to actually do anything useful in frontend. But what was interesting about that is how much it was using things like Dynamo under the hood, which in some cases, it never even bothered to tell you.

Alex: Yeah, it is interesting. You're really seeing Dynamo, I think, be the database of choice for a lot of these serverless applications. And part of that is because they're very infrastructure as code first. So, whether that's the serverless framework, or SAM, or the Amplify CLI that you're talking about, they're all using infrastructure as code. And DynamoDB works really well for that, where you can provision your table and it's all ready to go, and start interacting with it, as compared to something like a relational database where maybe you can provision that database but now you need to have something else outside of your normal infrastructure as code workflow where you need to be creating users in that database, creating tables, indexes, all that stuff, which can be tricky. So, I think that infrastructure as code friendly feel is a big part of it.

But also, and really the big thing, is the connection model where if you're using Lambda, like a lot of these serverless projects are, and you have this hyper ephemeral compute where it's spinning up only when a request comes in and then it goes away. It's faster if you can use something that has a stateless HTTPS connection model rather than making sure you're in the right private network to access your relational database, and then setting up that connection pool, making that request, and making sure you're not having too many connections as well. And I think that's the big reason why you're seeing so many people use Dynamo with serverless is that connection model problem. And really, I was hopeful for a while that the relational database problem would be solved in serverless. I've, sort of, given up on that, and I think for most people, it's going to be much faster for you to learn how to model in Dynamo than it is for someone else to build a relational database that really works with the serverless compute.

Corey: One of the things I've always appreciated about serverless was the idea that you weren't paying for instances to sit around idle, you effectively pay for what you use and that's it. With Dynamo, now it winds up having an on-demand capacity model where you only pay for things that it uses and it scales to zero. But even before that, the capacity model it’s always had has never been instance based. It's been about reads and writes, specifically throughput. And that pricing model is as best I can tell, unlike virtually any other database on the planet. Am I missing anything?

Alex: No, I think it's really awesome and pretty amazing, right? Because when you're thinking about your database, usually you're building out this new application, you think, “How big do we need to scale this database?” and you're thinking, “Well, how many users and requests are we going to be getting per second?” and then you try and map those requests into CPU, and RAM, and network, and all that stuff, which doesn't really convert well. So, mostly you're just guessing, and hand waving, and buying some instance and being wrong, and either overpaying or throttling and jamming up your application.

Whereas with Dynamo, if you can actually do that capacity planning, you can provision the reads and writes that you need to handle it, and you won't be overpaying. And if you don't have a good sense of your capacity, you can provision a bunch of capacity, and overpay for a little while, and get a sense of your baseline usage, and then scale it down to smaller read and write capacity as needed. So, yeah, like you're saying, even before that on-demand came out, it was pretty flexible and much easier to do capacity planning than with other databases. But now with on-demand, I mean, that's a big game-changer where you don't have to specify any capacity upfront.

You just say, “I have a database, and you handle reads and writes, and I'll throw them at you,” and it'll handle it for you. So, that's pretty incredible. The on-demand pricing is more costly if, in the theoretical case, you get full utilization, but most people just aren't getting full utilization. And that's true whether it's DynamoDB, or a relational database, or anything else because usually, you're way over-provisioned so that you don't jam up your database. So, I find out for most people, I think it's a pretty good pricing option. And in the worst case, at least, it's something you can turn on at the beginning and not have to think about capacity planning for a while until you've set that baseline.

Corey: If you're like me, one of your favorite hobbies is screwing up CI/CD. Consider instead looking at CircleCI. Designed for modern software teams, CircleCI’s continuous integration and delivery platform helps developers push code with undeserved confidence. Companies of all shapes and sizes use CircleCI to take their software from bad idea to worse delivery, but do so quickly, safely, and at scale.

Visit circle.ci/screaming to learn why high-performing DevOps teams use CircleCI to automate and accelerate their CI/CD pipelines. Alternately, the best advertisement I can think of for CircleCI is to try to string together AWS’s CodeBuild/Deploy/Pipeline suite of services, but trust me, circle.ci/screaming is going to be a heck of a lot less painful and it's where you're ultimately going to end up anyway.

Thanks again, to CircleCI for their support of this ridiculous podcast.

Corey: The strange part for me has always been in trying to predict capacity. Now, of course, like anything, DynamoDB offers auto-scaling these days which, like most forms of auto-scaling gives you exactly what you need, 20 minutes after you needed it. So, you wind up in these sudden burst scenarios, where suddenly you'll wind up with a bunch of things being throttled as a result. That, for better or worse is a lot easier to manage with something like Dynamo but it's still there and it's still obnoxious. What have you found that makes for a reasonable approach to start minimizing the impact of that sort of thing?

Alex: I think my first recommendation is really evaluate whether auto-scaling at all works for you. I think it works pretty well for EC2, or different things like that, but for DynamoDB, I found it to not work quite as well because you do need more of an even scale-up period because, like you're saying, it takes a little time before you, sort of, trigger your threshold, and then it actually starts scaling up. And that might be too late for you, and now you're getting throttled, and you have unhappy users. So, first of all, that scale up might not be as fast as you want. And then, number two, you probably have a target utilization of, maybe pretty low, maybe 30, 40, 50 percent, to where, as soon as it hits that, now it's going to start scaling up to make sure you're doing that far enough in advance.

And at that point, if you're looking at a 30 or 40% target utilization, well, I think at that point, you might as well use DynamoDB on-demand because now you're being pretty competitive with that on-demand price pricing if you're only getting sub 30% utilization out of it. So, I think auto-scaling works for some people that have pretty predictable scale-up patterns, but I would advise against it in most cases, I think.

Corey: One of the, I guess, more interesting pieces that I tend to see with any AWS service—and frankly, I’m not going to bound that AWS, any service at all—is how people can misuse it in weird and creative ways. So, I guess, I have two questions for you. First, what do people misunderstand the most about DynamoDB? And secondly, what is the most egregious misuse of DynamoDB that you have seen to date?

Alex: Oh, man, that's a good question. So, I think first, a couple of egregious misuses. The classic one is people that try to implement what I call “Faux SQL,” where you're trying to implant a relational pattern on top of NoSQL. So, you're not really using SQL there. You know, I see that happen. I actually think it's not the end of the world in terms of patterns. If you do have a highly relational model, and you're pretty new to Dynamo, and you want to dip your toes in the water that way, I think it can work. But you're not really getting all the scaling benefits, and you're not getting the true benefits of DynamoDB.

I think the second and more pernicious one is thinking that NoSQL and schemaless means very flexible data modeling and data access patterns. And DynamoDB is pretty unforgiving that way. You really need to model your database ahead of time to handle the queries that you want. So, NoSQL and schemaless doesn't mean freeform, whatever. And I think part of that comes from MongoDB because I think at smaller scales, Mongo is more flexible and more developer-friendly if you do have a pretty small application. There's a lot of flexibility around query patterns, and just query and data access in general to where people thought NoSQL meant very, very flexible data access, and that's just not going to happen with DynamoDB. So, I think that's where people really get in trouble there, is not knowing their access patterns upfront and not designing for those access patterns. Now, second question, the worst abuse I've ever seen. Let me think on that one a second.

Corey: Because for better or worse, I don't see DynamoDB cropping up on most of my cost consulting engagements as a tremendous driver of waste. Sure, there's an optimization series of steps you can take around significantly large, DynamoDB environments, but it's not one of those scenarios where while, “Welp, 80% of your Dynamo spend is being wasted, knock it off,” in virtually every case. It's hard to screw up, in an economic sense, with Dynamo without really trying.

Alex: Yep, that's true. If I think about, sort of, the worst patterns I've seen, I'd put it at two ends of the spectrum. One is, sort of, relying on the scan operation in any hot path. So, if you have a HTTP request, where a user is actually waiting on the response, it's almost a bad idea to use the scan operation. People do it because they think that's the only way they can get across their data that's across different partitions keys, which is a big mistake. And maybe it works when they're developing, and they only have 50 items in their table, but then when they get some serious usage in there, then it gets pretty slow.

The second example I have, I think, is when you read about these fancy access patterns, and especially access patterns that are designed for very high scale tables, and you try to apply it on a low scale table. So, the example I'd use here is there's a pattern called write sharding or read sharding, where a partition in DynamoDB, which is a very small segment of your database—think, like, a particular user’s data items. There's a limit of 3000 read requests or 1000 write requests per second on that given partition, which again, is not your entire table as a whole, it's just the single partition. So, if you're going to be more than 3000 reads per second, which is a pretty high volume—now we're talking about amazon.com and Lyft, and Uber and all those very high scale applications, but if you're going to be over something like that, then you might need to shard up a particular partition—so a particular user, or something like that—into different partitions, and then do, like, a scatter-gather query on those, where you query all those partitions and put it back together.

That pattern is—remember, most people are not going to be in that use case of 3000 accesses per second on a particular user, or something like that. Most people aren't going to have those. But I see people try to implement some of those uber high scale design patterns at pretty low volumes when they don't need them. And it really complicates your logic, and it's probably a little bit slower as well. So, I think that's the other place where I see people make mistakes.

Corey: It's weird when you start off designing something—and this, I think, is probably one of the failure modes that DynamoDB has—back around two and a half years ago, two years ago, somewhere in that range, I decided that it was time to actually build something custom to write my Last Week in AWS newsletter, and all of the links live in DynamoDB. As it turns out, I had made a number of decisions then that didn't actually wind up working for the current generation of how I generate these things. So, the schema that I came up with, such as it was, when I was designing this hasn't lent itself super well to be extensible in a modern way. And I'm looking at this thinking, there are now almost 4000 links that live in that database and doing a modification on all those would be incredibly challenging and very expensive.

And then, I realized, wait a minute, it's a lot in the context of all the material that I've written that now lives inside of that database, but it's only 4000 items in this thing. It is a megabyte or so on a disk. I have continuous backup turned on for this, and the total Dynamo spend on this thing is nobody-cares money. We're talking pennies a month. And I'm realizing I don't actually have a big data problem, or even an old data problem compared to virtually anyone else out there. It's just I'm stuck in the mindset of seeing things a certain way. Taking a step back and getting some much-needed perspective has been incredibly valuable on that. And heck, you’ve helped me with that through a series of conversations we've had over the past year, starting to reimagine how things could work differently.

Alex: Yep. Yeah, exactly. I think people do, sort of, think, “Oh, man, that's a lot of data I'm going to have,” and they think about the pricing—or they're not thinking specifically about what that pricing actually means. But yeah, 4000 items seems like a lot, but then you go look at even the on-demand pricing which, like we said, is going to be more than your provision capacity. And I believe it's a quarter per million reads from a table. So, you're not going to be anywhere near that, and even so it’ll only be a quarter. So, yeah, the pricing really works out for a lot of folks and I find out is, for a lot of people, especially building smaller serverless applications, which you see a lot, the pricing is not going to get you as much as other things like API Gateway might, or different other parts of your application.

Corey: I think the hardest part is that people have a hard time—at least I have a hard time, because I'm the most common type of developer, by which I mean bad at it—and it's hard for people in my position, at least, to step outside of our own use cases at times and imagine what other people's use cases look like. Being a consultant is definitely helpful in that regard. Every year that I do this, I feel like I'm getting three years of experience just because of the different variety of environments that I tend to see. You're an independent consultant yourself. What have you, I guess, learned from your customers as you've gone through it?

Alex: Yeah, I still think there's just an education gap around DynamoDB generally, and I think that's getting better. I've got dynamodbguide.com out there, I've got the book that's out, and also the re:Invent talks from Rick Houlihan are great. And I think you're just seeing more of a community develop, but there's still not enough out there, and one thing I've been really grateful for over the last couple years, as I've put some of this stuff out there, and then had people reach out to ask for help around DynamoDB is, like you're saying, it really accelerates your learning. And one thing I see is there's so much uniqueness around how you have to model your DynamoDB table where, if you have a relational database, there's likely one way to model it. You have a different table for each of your entities, you have the relationships between them as modeled by foreign keys, and you're pretty good to go.

But with Dynamo it's a little more art than science to where you need to design that data model for your application, and you need to know all these little tips and tricks. So, I've had a great time working with clients and just other folks in the community over the last couple of years figuring some of those out, most of those are encapsulated in the book. So, I have a thing that I call strategies, where you really just need to know all these different micro-strategies about how to model in DynamoDB. And those are just different tools in your toolbox, as the saying goes, and when you're attacking a particular access pattern, you just think about what are the contents of that access pattern and then how can I put together a few different strategies to help handle that? So, I'm hopeful that the book will have some of the learnings that I've had to do the hard way and other people have had to do the hard way, to where you can pick that up and accelerate your learnings in a similar way.

Corey: Yeah. DynamoDB the hard way is always fun. It seems like DynamoDB the easy way is to call you.

Alex: I like that. I like that, it's a good catchline.

Corey: So, origin stories are always interesting. And you're deep in the weeds of database design and architecture. So, what's your backstory? Computer Science degree? Where’d you go?

Alex: Yeah, yeah, not not a computer science degree at all. I actually have a liberal arts degree; economics, polisci, that sort of thing. And then, I went to law school. So, my wife and I, we got married right after college, and we both went to law school together. And I was doing law school, and during your summers of law school, you usually go work at a law firm. I was doing that.

I met a friend there who, he was a lawyer, but he was also into the tech scene and, sort of, got me into that. Actually started working on a startup with him and my brother-in-law where we were putting soil moisture sensors into farmer's fields to help them figure out when they need to irrigate, things like that. So, sort of, an early Internet of Things application, we thought anyway. So, they were starting to work on that. I asked how I could help. And they said, “Well, you could write a business plan.” So, I did that, which was a lot of wasted writing for no purpose, really. But as I, sort of, got into the project, I was enjoying it, and asked how I could help. And they said, “You could learn to code if you wanted to.” So, this was my last year of law school.

I had about a semester left, and your last semester of law school is pretty casual, you're mostly just running out the clock. So, I spent most of my time in class hanging out on Stack Overflow and trying to figure things out. Trying to build this application, but one thing I did is I would just hang out on Stack Overflow and wait for someone to ask a question about Django, which is the popular Python web framework. Someone asked that question, I had no idea how to answer it. So, I would just go Google it and search around and look through the source code and try and figure it out and try to answer it for them to the point where I was, like, the number two Django guy over a six month period, or something like that, for a while, even though I'd never done any real, genuine production.

But I think it was a really useful way for me to figure out how to debug code, and read source code, and figure problems out. So, I really enjoyed doing it, and it was a good learning experience. I graduated law school and worked as a corporate lawyer for almost a year, not quite a year, and then decided to make the jump into tech. So, I was fortunate to get a job at a startup that was based in Nebraska. Did some AWS and data engineering stuff there, and yeah, and the rest is history. So, yeah, I used to be a corporate lawyer. I only made it about a year and now I tell everyone, I'm a suspended lawyer because I have no longer paid my legal dues.

Corey: Phenomenal. I think that—talk about unicorns. You are probably the only lawyer I've ever heard of who went to go work on Databases but didn't take a job at Oracle.

Alex: [laughs] I love it. Yeah, they weren't hiring for me in that case, yep.

Corey: [laughs] Now, it's always interesting to see in the path that people walk, and the assumption that everyone—oh, if you're doing thing X, you must have educational background Y. That really only matters the first few years. Then it's a question of where has your professional path led you. I don't see too many options where it makes sense for someone mid-career to go back and get an undergraduate degree in a particular area of study. It's easier in most cases for them to effectively lateral from wherever they happen to be towards whatever it is they want to be doing next.

Alex: Yeah, I totally agree. I mean, there's so many resources out there to where you can really pick this stuff up if you want to. And with something like programming, it's easier than most fields to, sort of, prove that you know something. So, even though I didn't have a CS degree or anything like that, I had worked on this IoT-ish application with a web front end, all that, for a year. And in my interviews, I was able to speak coherently about different aspects of AWS, things like that. So, it was clear that I knew something. Also, at that same time when I was deciding to go into tech, I actually had another offer from a startup that would have been remote, and really how I got in with them is I was using their product, and I realized their documentation and examples weren't very good. So, I wrote some blog posts and examples for them to the point where they then offered me a job. So, I would say in something like tech, it's easier than most places to, to prove your worth and get that start, even if you don't have the traditional credentials. Whereas I couldn't have made that other switch. I couldn't have gotten an undergrad CS degree and then tried to convince a law firm to hire me on as a lawyer. That just wouldn’t happen.

Corey: Yeah, for some things like regulated professions, there's really not another option.

Alex: Exactly.

Corey: sort of like no one yet has let me try my hand at being an amateur surgeon.

Alex: Yep, [laughs] exactly. I don't know why Corey, I think you can do it. But yeah there's a lot of materials out there. You mentioned DynamoDB the hard way earlier, and really the resource that got me started was Learn Python the Hard Way, by Zed Shaw. That was one of the first resources I did, and I just worked through all those examples, and then just trying to build something myself, and also help others on Stack Overflow. So, yeah, that path is definitely possible for people that don't have the traditional background, which is great.

Corey: So, if people want to learn more about what you have to say, where can they find you? And where can they buy your book?

Alex: A few different places. If you want to find me on Twitter, that's probably the best way to get in touch with me. So, that's @AlexBDeBrie. I also blog at AlexDeBrie, so that's confusing because they're different. But if you want to Google Alex DeBrie, I'll probably come up in those places. If you're interested in DynamoDB more specifically, go to dynamodbbook.com. That's got a link for the book where you can buy that. And feel free to reach out if you have any questions on that stuff. I've also got dynamodbguide.com, which is more of an intro to DynamoDB, more about interacting with the API a little bit. So, if you want to look at some free stuff before you take a look at the book, go ahead and check that out as well.

Corey: Excellent. Alex, thank you so much for taking the time to speak with me today. I appreciate your time as always.

Alex: Thanks, Corey. Thanks for having me.

Corey: Alex DeBrie, author, AWS Data Hero, independent consultant. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts. If you hated this podcast, please leave a five-star review on Apple Podcasts and leave a comment explaining why you should never use a single table pattern.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Stephen O’Grady

Stephen O'Grady is a Principal Analyst and co-founder of RedMonk, the open source industry analyst firm. He focuses on infrastructure software such as programming languages, operating systems and databases, as well as covering horizontal industry trends such as open source and cloud computing.

Before setting up RedMonk, Stephen worked as an analyst at Illuminata by drawing on his real world expertise in architecting and developing applications for leading systems integrators. Prior to joining Illuminata, Stephen served in various senior capacities with large systems integration firms like Keane and boutique consultancies like Blue Hammock.

Regularly cited in publications such as the New York Times, BusinessWeek, the Boston Globe, and the Wall Street Journal, and a popular speaker and moderator on the conference circuit, Stephen's advice and opinion is well respected throughout the industry.

Links Referenced

  • The New Kingmakers book: https://www.amazon.com/New-Kingmakers-Developers-Conquered-World-ebook/dp/B0097E4MEU/
  • RedMonk: redmonk.com
  • Twitter: https://twitter.com/sogrady

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is brought to you by DigitalOcean, the cloud provider that makes it easy for startups to deploy and scale modern web applications with, and this is important to me, no billing surprises. With simple, predictable pricing that’s flat across 12 global data center regions and UX developers around the world love, you can control your cloud infrastructure costs and have more time for your team to focus on growing your business. See what businesses are building on DigitalOcean and get started for free at do.co/screaming. That’s D-O-Dot-C-O-slash-screaming and my thanks to DigitalOcean for their continuing support of this ridiculous podcast.

Corey: This episode is brought to you by Spot.io, the continuous cloud cost optimization platform, saving businesses millions of dollars each year on their cloud bills used by some of the world's largest enterprises and fastest growing startups like Intel and Samsung. Those are enterprises and duo lingo. That's a startup. Spot.io delivers the optimal balance of cost and performance by leveraging spot instances, reserve capacity, and on demand. Give your workloads the infrastructure they deserve. Always available, always scalable, and always at the lowest possible cost. Visit Spot.io to learn more.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Steve O'Grady, Principal analyst and co-founder of RedMonk. Welcome to the show, Steve.

Steve: My pleasure.

Corey: So, let's start at the beginning. What exactly is a RedMonk?

Steve: [laughs] It's an excellent question. So, a RedMonk is somebody who, in my case obviously, works for an analyst firm. In this particular instance, RedMonk, the analyst firm. We're a small firm, we're developer-focused. We have been doing this for much longer than James and I would care to admit. So, yeah, I think, like a lot of analyst firms, we do research analysis, and we look at trends and what's being used, what's not, why, all sorts of fun things, we just take a little bit different, yeah, like I said, lens or angle to the question, in the sense that we are big believers in the practitioner. So, yeah, that's what RedMonk is.

Corey: So, let's back up a second because Lord knows I didn't know the answer to this. What's an analyst firm?

Steve: Yeah, that's a really good question. And it is one that, frankly, if you asked my parents or probably any of the parents of any of our other analysts, none of the parents would be able to give you a clean answer. So, you are in good company if you don't know the answer to that. The short answer is that, so we research. That's our primary work activity. And that means we talk to all sorts of people. We talk to developers, we talk to engineers, designers, operators, admins, take your pick. DBAs, and an on and on and on. We perform quantitative research. So, we'll go out and look at, oh I don’t know, GitHub, Stack Overflow. Anything we think is going to give us a sense of quantitative trends from a developer perspective. And then, we talk to companies, lots and lots and lots of companies. So, the nut of it is that we talk to—so we have all these conversations, we do all this research, primary and otherwise, and we synthesize that. We look at it, sit from a distance. And we make judgments in terms of, okay, we think this technology is going to go up and we think this technology is going to go down. Here's why. And perhaps more importantly from a commercial standpoint, companies and businesses alike then want to understand. Okay, so given this trend, how do I apply this to my business, right? So, an example of this would be I wrote a piece last month, I think—at least a couple weeks ago—talking about Amazon, and how one might go about competing with Amazon. That piece is itself a product of a lot of analysis on our part, looking at conversations and also having a long history in this market looking at, okay, what has worked and what hasn't, and why. So, you put the history together; you put the research together; you come out with this piece that says, okay, if you’re going to try to compete with Amazon, which is obviously very difficult, here's one way you might do that. And so, a lot of people read that and we're fortunate to have this to be pretty widely disseminated. And then, basically, companies will come back to us and say, “Okay, well, what does that mean for my business?” And our job is to answer that as best we can. So, yeah, I mean, we do lots of different things for lots of different companies, and it depends on what people need. Sometimes it’s, tell me why my messaging sucks. Sometimes it’s, how do I reach developers? So, there's lots of different questions we answer, but that's probably the shortest version I can give you.

Corey: [laughs] It's fun, in that I've inadvertently been dragged into analyst events and a few analyst engagements so far, which is fun because I have no earthly idea what an analyst actually does, but in my expression of it, it always seems that oh, the sarcastic snarky things you say to people in public that you have to apologize for later. Well, can you come say that to us about what we're working on? Which was interesting. It was a source of insight to some extent, and I didn't realize this was actually an entire profession. And surprise, come to find out that it, sort of, is.

Steve: Yeah, not everybody can quite bring the snark the way you can, but for sure, I mean, one of the things that all analysts, to some degree, are paid to do, at least by—well, I should say there are companies that basically want to pay analysts to come in and say that they're the greatest thing in the world. And if anybody listening to us is an analyst that does that, more power to you. That's not what we do. We don't believe in that. So, we're there to talk about are companies that genuinely want to improve, want to get better, and want to know, all right, where are my blind spots? What are the things I'm not thinking of? Because however smart any one of these companies are, the simple fact is that we have access to many, many more companies than they do. Because—look, whether you're selling technology, whether you're buying technology, most of your competitors probably aren't going to talk to you, right? They do talk to us. So, we're in a, we're in a very unique position to be able to have a lot of conversations that very few companies are in a position to have on their own. So, we can pair that and add that to all these performance research that we do. And, hopefully, in any given conversation tell you something you didn't know before, we showed up.

Corey: Insight, it's one of those valuable things that people tend to, at least in my experience, value almost exactly as much as they pay for it. Free advice, turns out you can get that on the internet and most people don't tend to play those games. Whereas, when you have someone who sits there—especially in your case, where you provide a quantitative aspect to something that is otherwise a qualitative discussion—it starts to be a much different story. I made fun of the whole Gartner Magic Quadrant dance for many years, back when I was an engineer, and then I started doing more, as I always say, business-level work, speaking with decision-makers at levels that went beyond just code, and I suddenly saw the value of it where I never had before. Which brings us to a topic I've been meaning to pick your brain on in a scenario when you couldn't possibly say no, the book that you wrote circa 2012 if memory serves—

Steve: Yeah. Was it 2012? I think it's 2013. I don't even know at this point.

Corey: My awareness of the passing of time continues to get fuzzier.

Steve: Yeah, I hear you. Yeah. So, what do you want to know? What do you want to know about The New Kingmakers?

Corey: Yes, The New Kingmakers was a relatively short read that talked about developers as being the determining factor whether or not any particular vendor would succeed or fail at driving adoption. Is that a reasonable tweet-size synopsis? Or have I missed a whole bunch of salient point?

Steve: Well, so I think the—so you mentioned the length. The length is interesting because that is—I joked about this on Twitter at one point, if anybody is suffering from a lack of humility, the simplest solution for that is to write a book, and then to go read the reviews of your book, particularly the one- and two-star reviews, because it is one of the major complaints and criticisms of that book was that it's too short. Because I don’t remember what it was, I think it's, I want to say it's, like, 18, 19,000 words which is, say half the length, maybe a little more than a third the length of regular business text. So, the funny thing is, I did that intentionally. It's not for lack of case studies. It's not for lack of evidence, it was more, I had read a book by Erik Brynjolfsson and Andrew McAfee. And basically, they did the same thing in their book. And it was 75, 80 pages, I want to say. And it's I realized after reading it, I was like, okay, look, you can communicate about a topic, introduce it, provide the evidence, provide recommendations in a relatively limited span of time, because essentially, from my perspective, I didn't have time to spend one hundred fifty, three hundred, four hundred pages, whatever it might be on that subject for a book. I did have time to invest, all right, 75 pages, I can knock that out. So, yeah, I took the essentially same approach with The New Kingmakers and thought, all right, there's probably a whole bunch of people who are not willing to invest three, four or five hundred pages in concept of developers, but if I keep it shorter maybe they'll go out and read it. And for the people that worked for, it worked like a champ. People were like, “Oh, this is great you can, in many cases, absorb it one sitting.” And then, for the people who it doesn't work for them, well, they're the people who are leaving one or two-star reviews. More power to them.

Corey: Never read the reviews, never read the comments.

Steve: Right. Now, some of the reviews are good and they'll provide you with feedback. I could say that though, as a white male. Much easier for me to do that, than if you were a person of color, or women or, or something like that. So, yeah, you're going to get to own and acknowledge your privilege where it exists. Anyway, back the point. One of the really funny things about that book for me was that I had a lot of conversations with engineers after it came out, and most of the conversations went something like, hey, it was really good, it was interesting, and there are definitely some use cases in here that I hadn't heard of, or some facts and some figures and so on, but I mostly knew this. This isn't anything super new to me. And my reply to all these engineers was, it's not supposed to be, this is not for you. Basically, the purpose of the book, in a nutshell, was to the importance of developers in it, as we term it, the new kingmakers, had been apparent to us for quite some time. And I kept having the same conversation over and over and over with senior executives at vendors, enterprises, and everything in between, and thought, “All right, I can keep having this one on one level conversation, or I can package this up and potentially recruit developers into upselling the book, one for their own benefit and two, so I don't have to keep having the same conversation over and over.” So, basically, the object was to take this, put it in developers hand, and be able to put them in a position where they can hand it to their boss and say look, this is only, like, 70, 80 pages, depending on the printing, whatever it is, you could read this quickly, and then you'll realize why these teams are important. And in many cases, it was a success. There are certain folks, as I said, who are less than thrilled with one or more aspects of the book, but I think for the folks that, particularly for the engineers whose bosses have read it, I think, hopefully, it made a difference, just in the way that they're valued within an organization, which is, look, if I could accomplish nothing else in our role at RedMonk, that's a good thing.

Corey: One thing that I found refreshing about it is the length, where there are too many business books that I've read over the course of my career, where it seems that, yeah, this really is about an 80-page book that is struggling to get out of the 300-page book that the publisher made them write. Where you wind up with this tremendous amount of fluff that belabors the same point, all the case studies, all the rest, and yeah, I just needed that core of it.

Steve: That's right. Yeah.

Corey: Yeah, the overall thesis of the book where developer experience was driving acquisition decisions, as opposed to terrible software that was aimed at the buyer, who was very far removed from the person implementing it, that rising tension and that rising approach has been proven to be correct, in an awful lot of spaces. What I find challenging from a business perspective, given that what I tend to do is purely advisory around the AWS bill, is that engineers, in my experience, for what I do, are terrible customers across the board.

Steve: [laughs] Yeah, yeah. So, like I said, we definitely had some interesting and some challenging conversations with engineers afterwards, either of the type that I noted, which is, hey, I knew this already. Or people who wanted to bike shed it, and basically say, “Hey, here's this one nit in this one area and let me tell you all the reasons that this is wrong.” So, that's fine, right? That comes with the territory. That's not that big a deal. I think, like I said, the difference, I think, at least in my experience, as I have—well, literally as I directly experienced, in the wake of the book is, essentially, that even the people who wanted to nitpick it, they recognized that it was in their best interest to distribute this. Because it's difficult for us to remember now, because we exist in a world where developers are very highly prized, but certainly in the years running up to the publication, and even the years directly afterwards—we're going back again to, whenever it was, 2012, 2012 certainly when it was in the drafting forum, either published in 2012 or 2013, I can't remember—the importance of developers within organizations was certainly not the centrally agreed-on valuation that it is today. So, what we're trying to do is put developers in a position where they could—help them make a business case to their boss, or their boss's boss, or their boss's boss's boss or, in a couple of cases, we had conversations with developers who had said—where was it? I think it was a Red Hat summit—there's a gentleman by the name of Kieran Broadfoot, who is a senior executive over at Barclays, and he gave a talk at the summit and whipped out the book and was like, “Hey, this is great, and we issue this to our engineers.” Very, very nice guy. I had the opportunity to chat with them afterwards. The interesting thing to me is that I later had conversations with engineers, whose CEO, VP of engineering, pika, in the CIO, pika senior leadership term, was at that talk, and their senior executives went out and bought the book because, hey, somebody on stage at Red Hat recommended it and came to them and said, “Hey, you guys are really important. What can I do to help?” so I think it is certainly true that engineers are not always your best customer, but I think when you're in a position where you are going to directly advance their interests, and where your interests are very much mutually aligned, then your life is, I want to say, much easier.

Corey: This episode is sponsored in part by ChaosSearch. Now their name isn’t in all caps, so they’re definitely worth talking to. What is ChaosSearch? A scalable log analysis service that lets you add new workloads in minutes, not days or weeks. Click. Boom. Done. ChaosSearch is for you if you’re trying to get a handle on processing multiple terabytes, or more, of log and event data per day, at a disruptive price. One more thing, for those of you that have been down this path of disappointment before, ChaosSearch is a fully managed solution that isn’t playing marketing games when they say “fully managed.” The data lives within your S3 buckets, and that’s really all you have to care about. No managing of servers, but also no data movement. Check them out at chaossearch.io and tell them Corey sent you. Watch for the wince when you say my name. That’s chaossearch.io.

Corey: To be very clear, one of the challenges I've had with engineers as customers is—I should probably qualify that before I wind up getting letters—has been that engineers are always very passionate about what it is that I'm focusing on. I will sit down and have a 90-minute conversation easily with engineering types. And they will talk all about how I do what I do, the different areas I can wind up addressing. And at the end of it, it's almost like I'm talking with Hacker News, surprise, surprise. Well, that doesn't sound hard. I could build that in a weekend. Cool, great. But before you do that, and go tilt at that windmill, can I can maybe talk to your boss? And from there, that conversation generally tends to pivot into introductions, because it turns out if you dig a little bit beneath that surface layer as well, engineers are very challenging to win over, and by the time that you wind up, effectively getting them in your corner, it turns out their signing authority caps out somewhere around 50 bucks a month. So, they're not in a position to be your buyer. At best they can be your champion.

Steve: Yeah. Yeah, I think it helps when, first of all, I was helped out because the product was initially sponsored by New Relic, later sponsored by—I can’t remember who the second one was—PayPal. These are all through O'Reilly. So, developers didn't have to pay anything. They could just go buy it. And in those cases, if I remember right, I think it was a PDF. So, they could just take this free PDF that they’d been given and send it to their boss. So, in these cases, I don't have any issues with, you know, certainly selling the book. If you can't sell a $8, or $10, or $12, or wherever the price is now, item to a developer one time, then it may be that the product is not worth $8, or $10, or $12, or wherever. And as far as being a loss leader for RedMonk sales, like I said, in our case, we wanted it to lead to many conversations, which it has, certainly from a commercial standpoint, but really, if nothing else—set the commercial relationship aside—the writing of the book, for me was really almost a exercise in essentially simplifying the conversations or, I guess, improving the conversations that I would have externally. So, prior to The New Kingmakers becoming more of a common framework that people discuss, I would have to sit there and make the case, talk to them about developers, and here's some of the trends, and here's why, and so on, and have that same conversation over, and over, and over, and over. And, look, there's nothing wrong with repeatability, certainly in the consulting profession. It can pay bills and so on. But at some point you want to be able to say, “Okay, look, the basics are established. We all agree on this. Now, let's go have the more interesting conversation in terms of, what does this mean to your business? And how do we change this? And how do we fix that to take advantage of the situation, or mitigate your liabilities, whatever it might be?” So, yeah, it was honestly more of an attempt to shift the type of conversations that we had then spin up entirely new ones, in spite of the fact that it is most certainly done that for us over the life of the title.

Corey: Do you find that the book had its intended effect, in that it did change the tenor and character of those conversations, that it got you to a point where you're able to have those discussions in a way that resonated more?

Steve: For sure. Now, I will say that—well, let me answer that two ways. So, the first way is that in cases where people read the book, or frankly even in some cases, haven't read the book, but have seen presentations by, like I said, at Red Hat summits or SAP has mentioned it on stage, IBM, I think Microsoft has. Anyway, so it's been picked up and talked about by many of the largest vendors in the world, who then go out and propagate it to all of their customers. So, in cases where people have read the book, and come to the table saying, “Okay. Yeah, I've read this, I understand these pieces,” then great. It has certainly shifted the conversation. But the second, probably more important context is that, I think, to a large degree—and we say this all the time at RedMonk—we're a small analyst firm, and we have always tried to recognize and to be self-aware enough to understand what we can and can't do. And what RedMonk, as a small analyst firm, can't do is shift the market. What we can do is basically say, “Look, this shift is occurring. Let me tell you what we know about it.” In some cases, like in The New Kingmakers, let me give you a term for it, but, frankly, in a world where The New Kingmakers never gets written, does the market still shift? Yeah, for sure. Because basically, we're talking about massive trends that are in flight: open source, cloud, Software as a Service, democratization of access to educational resources, on and on and on. So, these things were going to have an effect one way or another, whether we knew what to call them, or whether we came up with a name for it or not. So, yeah, certainly in a very narrow and specific context, the book was absolutely able to help improve those conversations, but in all likelihood, they probably would have improved one way or another.

Corey: So, here's the dangerous question. You've had an entire internet of people telling you in a variety of different contexts for the past, well, let's call it seven years. If you were to rewrite the book today, what changed

Steve: [laughs] Oh, yeah, James, put you up to this, didn’t he? Yeah, James, for listeners who are unaware, has been badgering me for several years to write a second—

Corey: And who is James for those others who are unaware?

Steve: I should mention that as well. James is the co-founder of RedMonk. James Governor and I founded the firm way back when. Much longer, again, than I would care to admit. Anyhow, so my co-founder has been after me to write a follow up for a number of years, and who knows, it still could come to pass. It turns out that writing a book when you don't have a child is a lot easier than writing a book when you do. But books are not written solely by people with no children, so who knows, maybe at some point, I'll find the time for it. The short answer is that, I think in many respects, it's a scenario where I sit down to write the second edition tomorrow, it's less debating the concept and proving the concept—which is largely what the first edition was about, but rather—okay, let's take for granted that much of what was predicted has come to pass, and certainly spent some time on what didn’t, maybe, and why, but largely thinking about, okay what's the impact and what does this mean moving forward? Because the world, again, that that book came out in from a landscape standpoint was really different than the world today. And that is just one example. We have seen written; talked about extensively, problems of fragmentation. So, what we mean by that is that if you go back 20 years—frankly, if you go back even 10 years, the volume of available solutions that developers and the businesses they work for have available to them is much smaller. So, one of the things that has happened is that developers essentially took over the world, and they kept producing software, which, at first, it’s great. Anything you could possibly want to do, there's a library for, there's a piece of software for somewhere. But then, you start thinking about that problem at scale, which is, okay, I'm going to need a couple different pieces of software to solve any given problem. And now I have maybe a dozen, two dozen different credible choices within each one of those areas. And oh, by the way, there isn't just one or two or three ways to do things now, there might be four or five or six. So, generally speaking, if we were to think about a second edition, to think about what a follow up might look like, it would honestly spend less time on the mechanics of how developers were empowered, as the first did, and more trying to understand the implications of, okay, what does that transition empower? What does it mean? What are the practical implications of that for both the developer and the businesses alike?

Corey: One of the hardest parts of all of this is you wind up effectively, in some ways, disturbing established orthodoxy, where people push back not because you're inherently wrong, but rather because you are saying things that run counter to the established narrative that they are invested in preserving. Have you seen a lot of that?

Steve: [laughs] Oh, way back in the day for sure. I mean, I can tell you—like I said, RedMonk’s been around for a long time. And some of our earliest roots were looking around—and open source is one example in terms of, okay, hey, we're going out and talking to all these businesses and more importantly, engineers working there. So, we have a pretty good idea of, hey, there's a lot of open source in these organizations, and then you go and talk to some of the analysts at the time, or read some of the reports. And because that can't be measured in the same way because it used to be, all right, we're going to trap unit shipments, we're basically going to track the finances because that was, once upon a time, that was the only way to get software, you had to pay for it. And we just looked at this and said, “This is insane.” You're basically missing a whole class of software that is widely in usage at scale. We saw the same thing later on with the cloud. And I can remember looking at some of the reports, and they were saying things like, “Hey, these x86 hardware segments are leading the market and their future is bright,” and so on, and you're like, okay, what about this cloud stuff? Or what about these ODMs? We don’t have a good way to track that, so we're not counting that. So, when we would point these things out, from time to time, you would certainly get a lot of pushback. People are like, oh, this is insane. Why would I ever care about a developer if they don't have any budget? This cloud stuff is just a toy. The toy thing is a recurring theme. It's one of these—I’ve joked about this on Twitter at one point. When some established technologist calls some other technology a toy, like my ears perk up. I'm like, oh, okay, I’ve seen this play out before. So, yeah, the short version is that we have seen this many, many times over the course of our career, where we're saying one thing that the larger established analysis firms are saying something totally different, and we look insane, frankly. And we've been fortunate that some of the bets that we have made, largely around developers and things that developers use, have proven to be, I think, fairly accurate over time. So, what that means is that for a little while, when you're an unknown and people have never heard of you, then it's much easier to dismiss you out of hand. Once you've been doing it for a little while, and people have some background in terms of, okay, I've seen this stuff before. These people are maybe not totally insane, so maybe I'll listen. So, yeah, I think the short version is that we got a lot more of that a while ago. But we still see that from time to time. People will—oh, I’m trying to think of things that have been more controversial lately—certainly some of the things I have said about open source in APIs and so on, are not popular, and certainly violate the orthodoxy if you will. But as I said, I think we've been doing this long enough, we have a track record that people can look to, and it's certainly not that we're 100% right, or right all the time, or anything like that, but I think people at least are willing to listen in ways that maybe they weren't 5, 10 years ago,

Corey: We can hope that people evolve their listening skills as they evolve their careers, but that's not always a guarantee. And it turns out, there's always a new generation who has to learn things from first principles, that's called Hacker News.

Steve: I see a lot of that in the open-source space. I’ve seen a lot of that in [laughs] I won't name the company for obvious reasons, but there was a company that did pretty well with a particular class of infrastructure software, and we talked to them a couple times; we’re like, “Okay how are you going to make money here?” And at the time, they were super, super successful and, from a visibility standpoint, was not terribly preoccupied by the revenue. And we had said, “Hey, look, we've seen this pattern play out our number of times. And this is typically where people make money. So, is that of interest?” And by and large, it wasn’t. And yeah, that didn't end up particularly well for this company. It's honestly, I mean, you probably go through this, I'm sure, yourself. And I'm sure many of the listeners have this experience, too, which is, one of the single most defining characteristics of really successful people, in my experience, it doesn't matter who they are or what they do is just, just listen. Listen to what people have to say. When we talk to our clients, we tell them this all the time, hear us out. If you think we're wrong, and you can build a case for it, by all means, lay it out, and if we need to update our opinions, we will. But if you come in and think you know everything already then, maybe you do, but you're probably going to miss some things along the way just because you can't listen.

Corey: I think that's probably the best lesson to take from a lot of this. If people want to hear more about what you have to say on this and other topics, where can they find you?

Steve: The simplest way is to go to RedMonk.com. That's where all our stuff is. You can follow me on Twitter @sogrady, S-O-grady at Twitter. And yeah, website and Twitter are probably the best means.

Corey: Yes, showing up at your house is a distant third.

Steve: I would—yeah. Yeah. Office, too.

Corey: [laughs] That works. Well, thank you so much for taking the time to speak with me today. I appreciate it.

Steve: Not at all my pleasure.

Corey: Steve O'Grady, principal analyst and co-founder of RedMonk. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts. If you've hated this podcast, please leave a five-star review on Apple Podcasts, and then go talk to someone higher in the corporate hierarchy about why it is the way that it is.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Sid Rao

Sid Rao is the GM of Amazon Chime. He has over 25 years of industry experience, having worked at Infosys, Nortel, Microsoft, and CTI Group.

Links Referenced

  • https://chime.aws
  • “Chime after Chime” by Tim Leehane and Spencer Johnson

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Sid Rao, the GM of Amazon Chime. Sid, welcome to the show.

Sid: Thank you, Corey, for the opportunity. It's great to join you today and talk about Amazon Chime. Happy to be here.

Corey: It's easy to make a joke about you’re a hard person to get in touch with, but you're really not. And because you're on Chime, you sort of have to be accessible to some extent. But for the longest time, you were engaging with me on Twitter through a user account that just had your pets up there, and I've got to be honest, I thought for the longest time I was having conversations with your dogs.

Sid: Well, let's be clear: Sam and Max and Hawk—well, Sam, unfortunately, passed away last year, but Max and Hawk, they’re much better spokespersons and or spokes-dogs then Sid is, so I tend to use them in that capacity from the Amazon Chime team. They're very capable members of the team. One of them just got promoted from Puppy Level One to Puppy Level Two, and we use them routinely as mascots and spokes-dogs for various different activities. And you know, I'm not very familiar with Twitter, I will tell you, I'm not the best Twitter user. I've just started using it as of a year ago. So, I'm learning and it was better to learn with my dogs than learn alone. So, that's why you see Max and Hawk on my Twitter profile.

Corey: Excellent. I'm sure that they rank highly on the dig deep leadership principles.

Sid: Yes, they absolutely do. Especially Hawk. He's very good at diving deep. He definitely thinks big when it comes to food. And he definitely shows customer obsession in terms of licks, and you know, asking to be pet all the time. So, Hawk is definitely a role model on Amazon leadership principles applied to a puppy.

Corey: So, we started talking when I wound up making fun of your service, which it turns out is a terrific way to meet people. If you want to get to know what they're doing, insult their work, and oh, they come at you in a serious way. In all seriousness, don't do that the world has enough jerks in it. But my running joke about Chime for the longest time was that it has no customers because it was easy to fall back on. You had this, sort of, giant war in the messaging apps, originally of Slack versus HipChat. And then, HipChat, really, discovered that failing to innovate for a decade wasn't the best plan, and Atlassian sold the HipChat stuff over to Slack. Teams was then on the rise significantly, and they tend to talk about all of their daily active users because it likes to open itself, which is, yeah, if you can bundle something in, it works out super well. I'm not a fan of Teams, because I tried to use it once. And that really leaves Slack in that space. So, Slack has always been the yeah, that's adorable. We're going to be using Slack, but Chime is there just more or less as a placeholder until an actual competitor comes along. And now you're announcing that you're actually doing business with Slack. What's the deal here?

Sid: Correct. Amazon Chime has a number of different features and different service capabilities. We have a messaging platform in there. We have a meetings capability. We have a business calling and PSTN connectivity capability. And so, it's very easy to be, kind of, misunderstood as to is Chime a chat app? Is it a meetings app? Is it a calling app? What is it doing? And if you look at the world of unified communications today, that's actually pretty common across all of these services. Even if you look at Slack today, they have a chat capability. They have a meetings capability. They recently introduced a calling ability that actually works with Microsoft Teams. So, it's very easy in our industry to get confused about what is the primary focus of an app? Where are they going? And is Chime supposed to be a Slack competitor? And I'm going to answer Corey that usually these applications have all the functionality that a customer needs to communicate within their enterprise context, but they tend to focus on one area where they're going to do well. In the world of Slack, they've had a focus around messaging. Its world-class, best team messaging, and collaboration, and a platform for hooking in various different sources into a common channel. And they're really good at context management and how to keep context within that particular channel. They do also have a meetings capability, but it does have some limitations. It only supports up to 40 participants. You know, there's some limitations in terms of the audio and video capabilities that it supports such as PSTN dialing.

Corey: Oh yeah, they bought Screenhero which was fantastic, and then, from all accounts from the outside, immediately set about ruining it. And on one level, it's easy to cast stones, on the other, it turns out this stuff is super hard to scale. The idea of being a household brand, especially in a time when surprise, everyone's doing video conferencing in some form now when they weren't expecting to, is a hard problem.

Sid: Correct. And you know, it's not, I would say, the primary focus of Slack. If you look at how Slack presents itself to customers, it's very focused around messaging. So, now we are actually doing some work with Slack. And what we're doing is helping Slack extend the capability, security, global distribution, low latency audio and video services into the Slack application as a native capability of Slack meetings. So, when you use Slack meetings, starting now, you're going to start using Amazon Chime’s infrastructure and global deployment to power those meetings; improving the quality of those meetings, improving the resilience of those meetings; and also allowing Slack to continue to focus on their core competency around messaging, and leaving the problem of how to get video streams across the globe to talk to each other in real-time to Amazon Chime. So—

Corey: It splits the world. You’re right, it leaves Slack to focus on their core competencies of messaging, and changing their UI at random and confusing all of their user base. So, this has been in the works for a little while. So, that leads to two questions. One, what was that now that I'm allowed to ask? And two, what does that mean for the future?

Sid: Sure. So, what Slack was doing there was an experiment using one of the capabilities of the Amazon Chime SDK to determine what AWS region was closest to the Slack end-user to determine what region to use to host a meeting or to stream video or audio to, and what's the closest region to use to limit latency for customers. So, we provide an API. It's backed by Route 53 and a number of other AWS primitives, which allow our customers to discover regions that are optimal to host meetings. So, what Slack was doing was running an experiment across their entire user base to determine where the regions that are optimal for certain end-users, store that information, perform data analytics on that information, and then apply a meeting hosting algorithm to optimize audio and video latency using our SDK. So, that was what Slack was up to. They did it—they didn't want to obviously reveal to the world that we were working together, yet, and that's why we asked you politely, Corey, to not see that and be quiet about it. But—

Corey: Believe it or not, I am capable of being quiet. Not my core competency, but I am capable of it.

Sid: You are capable, and thank you for doing that. But that was the focus of Slack. They were trying to basically get ready for the rollout and determine what regions should be used to host a meeting. And so, they ran that experiment, they have their data, but that, again, highlights a benefit of Amazon Chime's SDK and infrastructure. We're in 14 regions; we're expanding that count every day; it would not surprise me in the fullness of time that we're in every AWS region that's available for customers to use our services. Latency is a very important problem in the audio/video streaming space. Limiting that latency is what drives a very high-quality audio/video service. So, what Slack was doing there was actually capitalizing on our infrastructure and our global reach using capabilities of the SDK that allow Slack to optimize what regions are used to host meetings. It's actually a very interesting problem. So, if you have, for example, a user who's based out of the East Coast of the United States, and they want to collaborate with a user who's in, for example, Taiwan or India. It's a very interesting problem as to what AWS region should be used to host that meeting. The natural choice a lot of application developers would do is just host it in either India or host it in the East Coast of the United States. Actually, the optimal spot, due to network and fiber routes, is actually in Paris. In France, actually, is the middle point that's most optimal for hosting that meeting. And you know, what the Amazon Chime SDK allows you to do is start that meeting in France, and then you can spin up another meeting, say, as participants start to join in and the majority of participants are, say, in India, you can start up another meeting closer to India, like Singapore, and start migrating users between meetings, optimizing on the fly for audio and video latency. So, there's some pretty exciting things you can do with our SDK by the flexibility we allow you to start up multiple meetings, migrate users between meetings, optimize audio and video latency based on region selection. And these are things that are really hard to do unless you have an SDK like Amazon Chime's SDK. If you were trying to build all that on your own, it's a quite a bit of operational load, it's a lot of effort, you have to manage a large amount of infrastructures across 14 different regions. And you're responsible for all that.

Corey: Well, aren't you overstating it just a little bit. I mean, I've been reliably informed by randos on Twitter, that Slack is just IRC as a Service and only fools would wind up falling for it, and here we are as a society that has been completely taken by the scam. Thank you random egg-looking person on Twitter with a bunch of random numbers after your name.

Sid: I think it's easy to think that chat is easy, or audio and video streaming is easy. After all, all you have to do is put up a box on EC2, and stream some packets between two endpoints, web RTC is already part of the browser. Shouldn't this be easy? And yes, it's easy to do that for a one-on-one conversation. When you're trying to scale this to hundreds of thousands of users across the globe with varying different connectivity and access types: customers coming across LTE, and 5g sessions, and DSL still a thing, and various different connectivity types. Managing the quality, managing the latency, managing the bit rates, reducing echo, these are all really challenging things to do. And in fact, we apply artificial intelligence and machine learning techniques to our streams to start to optimize them. We do a number of things that are really hard to do on your own. You can definitely go and try to do it on your own. The WebRTC is a standard, there's plenty of open-source packages that support WebRTC, and you can definitely get your initial deployment to 100 users done in the cloud, on your own, but when you want to scale that, or managing that over the long term, or supporting end users over the long term, it's a little bit harder to do that just on your own. So, that's a core problem we're solving with the Amazon Chime SDK.

Corey: So, this is always gotten to a bit of a, well, I'll call it mental challenge for me, where I look at things that AWS does, and if I can steal a borderline insulting analogy from the Git project, they talk about porcelain versus plumbing, and nowhere is this more evident than in Chime. We have the plumbing underneath: the infrastructure for communications that takes messages, routes them appropriately and securely. And that is awesome. AWS does plumbing super well. The app that we install on our phones, our tablets, our computers, etcetera, probably on an Echo device, because why not? That tends to be the porcelain portion of it, and it feels like AWS’s user interface pieces, the applications that end-users interact with has never been as robust or as polished as the plumbing. How do you feel about that?

Sid: That’s definitely fair feedback. We obviously can always improve our user experience, the customer experience of our applications. And Amazon Chime, as a service owner of Amazon Chime will tell you that it makes me sad every time an end-user has a problem with their applications. We definitely need to do better on the application side of Amazon Chime. And we regularly take feedback, focus on our customers, stack rank it, and try to fix the user experience problems in a iterative and consistent way. In the last year we shipped, I would say, at least 20 or 30 features that are end-user focused. For example, we supported the ability to share content just natively from the browser without requiring a plugin. We support a new Outlook plugins for better integration with Windows and Mac desktop users and Outlook. We supported Slack, actually, as an application and plugin mechanism as well. So, we continue to make end-user enhancements to Chime. I think there's two things to keep in mind though. It's evident from our launches that we've been a little bit more focused on meetings than on the messaging side. And that's what confuses a lot of end-users. They expect us to rise to the end-user capabilities of an application like Slack or Teams, and we're not doing that. And they get confused because they expect us to do so because we have a messaging function. And I think the second thing that folks need to think about is team tends to take an approach that until the plumbing is good—not just good. Great. Fabulous. You really shouldn't focus too much on the porcelain. And we basically have taken a intense amount of focus, Corey, over the last year about making sure the plumbing is good. Now, we have demonstrated the plumbing is good. We have customers who are using that plumbing. So, for example, Mindbody is using our plumbing for virtual yoga and health services on May 15, supporting a number of virtual capabilities that are focused around how to keep people fit and healthy during this time. And you know, the plumbing works. And the plumbing is actually not just good, it's fabulous, and that's what our focus has been on. If you look at how the plumbing has had to scale over the last couple months, Amazon itself had to double its utilization of Amazon Chime. So, they're, of course, a rather large customer in our customer list. And they doubled their utilization of the platform and we had to scale for that, and we scaled without a hitch. There were no capacity problems. There were no availability problems when we had to go through that scaling exercise. We then subsequently onboarded an insane number of customers onto our platform, with over—thousands and thousands of percent of growth. Let's just put it to you that way. Multiple orders of magnitude of growth and the plumbing did fantastically well. So, focusing on the plumbing is an important thing, and we wanted to make sure we got that right. We have started to show signal that we've got that right. We're now going to use that to start to add functionality to the application: improve the meetings experience, and start to focus more on that porcelain, and make that porcelain shiny. But the first thing to do is to make sure that the food is good. As you know, if you use the restaurant analogy, you can have the best friend to the house. But if the food tastes like crap, you're not going to eat dinner there. We've made sure the food is good, and now we're going to start to think about the front of the house a little bit and improve the meetings of application and the meetings user experience. And I think that's how we'll focus on the porcelain. So, it's a two-step process for us, and it does confuse people because they expect the porcelain to be good at the same time as the plumbing. And Amazon Chime has been launched and been out in the market for a couple years, and we had to go and focus on the plumbing and make sure it was good. And now we can come back to the user experience.

Corey: And you got to be able to focus on the plumbing because if you look at what a messaging app is, effectively every user has a monitoring system sitting open in focus on their computer most of the workday, and as soon as there's a slight blip in the messaging—maybe it's their terrible ISP, maybe there's a systemic problem, maybe there's a capacity issue somewhere along the line—but your app is going to get the blame for that. So, making sure the blame falls where it needs to is going to be incredibly important.

Sid: Yeah, I mean, blame is a very interesting topic for us. When we think about blame, we actually don't think blame is a good thing. What we'd like to do is point out when there's a problem, whether it's the ISP, whether it's your WiFi network, whether it's the Bluetooth headset you're using, all the way up to maybe we do have a problem with our infrastructure or capacity. You know, it's important to point out where the problem happens and exists, but then automatically help the customer with that problem. And that second part is the hard thing to do. The first part is actually relatively straightforward to do. We have monitoring hooks across the entire chain. We can tell what your signal strength is on your WiFi while you're in the middle of a call and realize that you're going to have a bad experience. So, there's definitely a ton of hooks we have, and monitors and alarms we have, but it's actually fixing the problem that's the hard thing to do. Yes, some things can only be fixed if human being gets involved. If your Comcast cable modem is having challenges and needs a reboot, you give it a reboot. But generally speaking, there are things we can do that actually optimize the end-user experience. So, for example, if we see that your ISP is having packet loss, connecting to a meeting that's hosted in say, us-east-1, we can test and proactively probe as to what the network telemetry looks like to connect to Ohio-based region, for example, and transition the meeting to Ohio. And what the SDK does is it enables a tremendous amount of power for the app developer to do those kinds of activities. One of our customers has built a town hall application, actually, where they have multiple different meetings set up for various different use cases. One is like a green room where people who want to contribute content to the town hall first get vetted and interviewed before they're put on air with the VIP who's actually doing the town hall, and then there's another meeting for the VIP to talk to people in their constituencies, and then there's another meeting for the popcorn gallery to interact about the town hall that's occurring. Well, what's really neat is they have multiple meetings being managed within a single application that's used by content moderators, the VIP, and participants, and they can have meetings that are, for redundancy purposes, set up in various different regions so that if they see problems, they can basically automatically switch users and optimize to a meeting that's better suited for their needs at that moment. And that flexibility is what the SDK empowers. If you want breakout rooms, if you want redundant meetings, if you want the ability to do green rooms, and various other different use cases, you can do that with the SDK. If you look at the meetings in collaboration world today, a lot of us are used to these very static monolithic applications, whether it's Chime the application, or Zoom the application, or Teams, or whatever it might be, and they're limited in terms of flexibility. When you're on a Chime meeting, you're on a single meeting. When you're on a Zoom meeting it's the same thing, you're on a single meeting. You can do breakout rooms because it's a feature, but there's a limit to that as well. And those limitations go away if you're using our SDK. So, what we've done is allowed JavaScript developers, iOS developers who like Swift or Objective C, Android developers who love Java, or Kotlin, to basically—we've given them a really powerful SDK that allows them to do basically whatever they need to for their collaboration needs, and customize both the security context, the quality context, and the usability context to the application that they're adding collaboration to, versus being constrained to a fixed application and trying to kind of modify it to meet their needs. So, for example, if you want to add virtual stand-ups to a coding tool or a development lifecycle tool, well, you can do that with a couple hundred lines of JavaScript added to that web app. And you can control it to be an all-day standup so there is no dial-ins or anything of that sort. You just drop in when you want to give a update and you drop out. So, that's very powerful and allows collaboration to be customized to the context that a user is living in, whether a developer living within a development tool, a financial services operator focused on an accounts payable or accounts receivable system, a yoga instructor who's living within their yoga scheduling app and that particular environment, or a doctor who's focused around the electronic health record. So, the final thing I'll say about the electronic health record is when the health crisis started, the first thing that almost every healthcare institution did was allow for automatically creating meetings for every doctor-patient interaction. So, everybody went and started their Chime meetings, the Chime link or Zoom link, or a Teams link. People would click on this link and join a regular meeting and perform their telemedicine activity. What they learned though, as soon as they rushed that out is they lack security controls over that meeting. It required the doctor to basically flip to another application so that when they're trying to update the electronic health record, they aren't able to see the patient at the same time. It reduced the amount of context brought into the meeting as to the medical record that's in play versus the actual communication that's ongoing with the patient. So, what the Amazon Chime SDK allows that electronic health record company to do is basically bring in the conversation with the patient into the context of the medical record, the prescriptions the patient’s on, the follow-up actions that are required, they literally can see the patient in the top right of the application they're using to update the electronic record. So, that context switch between a voice and video app and the business application the customer is trying to use at the same time, that's not a good context switch. You want the doctor to be constantly looking at the patient while they're updating the medical prognosis for that patient. And so, we're really happy about how a number of different electronic medical providers and telemedicine providers have started to adopt the Amazon Chime SDK to provide that level of integration and seamless experience.

Corey: If you're like me, one of your favorite hobbies is screwing up CI/CD. Consider instead looking at CircleCI. Designed for modern software teams, CircleCI’s continuous integration and delivery platform helps developers push code with undeserved confidence. Companies of all shapes and sizes use CircleCI to take their software from bad idea to worse delivery, but do so quickly, safely, and at scale.

Visit circle.ci/screaming to learn why high-performing DevOps teams use CircleCIto automate and accelerate their CI/CD pipelines. Alternately, the best advertisement I can think of for CircleCI is to try to string together AWS’s CodeBuild/Deploy/Pipeline suite of services, but trust me, circle.ci/screaming is going to be a heck of a lot less painful and it's where you're ultimately going to end up anyway.

Thanks again, to CircleCI for their support of this ridiculous podcast.

Corey: So, I have a couple of questions. First is when I have a Chime call set up, it tells me which region it's routing through. Now, if I were a regulated entity, I would probably care about that more. But from my perspective, I just want it to work, and I want it to be quickly. In fact, the only slight concern I have is some of these meetings are so freaking boring, that I want to make sure that they're handled domestically because sending it across borders is almost certainly a war crime. It's not something that tends to matter to me on a day-to-day basis. Am I weird like that, or is that a very specific edge case for a very specific customer profile?

Sid: So, when we added their region label to a meeting, there's a number of things that went into that thought process. The first thing is, there are actually a number of customers who are very paranoid about where a meeting is being hosted—

Corey: And rightly so.

Sid: —and said, “Look, we're taking meetings and making automatic decisions about where customer content is going to be processed.” Now, obviously the customer has control of this, they can go into the AWS console, and select what regions you want to use to host a meeting. So, first of all, I want to state that right upfront. Customers are in control of their content. They get to choose what regions are used to host meetings. But by default, we're going to automatically optimize that. And that's a great thing to do, to optimize audio and video latency by picking regions automatically for customers, but you better tell customers what's going on. So, that was the first reason why we did that. The second reason why we picked that was to also highlight to customers that this is a global service. We are in a number of different locations, and we're using the power of AWS to ensure that your latency is optimized, and your audio/video quality is great. And we're trying to keep that multimedia as close to a customer as possible. The first reason was the primary driver. Look, it's pretty scary. If I'm working on a very sensitive project with sensitive intellectual property considerations, I may not want meetings to be hosted in particular regions, and if that happens, I want to know about it so I can adjust my behavior accordingly. So, a lot of this was about making sure customers realize they were in control, and they knew what was happening to that media as it was processed. And some of the substitutes and alternatives in the market have had challenges around that. They haven't been very open and transparent about where meetings are being hosted and processed, and customers are rightfully annoyed about that. So, we wanted to make sure that didn't happen when we rolled out our solution to this problem, and that's why we put a specific label around it.

Corey: It's one of those things that's useful for some folks, and for the rest of us it more or less, it's just one of those things that shows up there. And I always feel like I should be doing something when I see something tell me what region is going through. Like I should be sitting there and [unintelligible] measurements to see which supported region is the closest for all people, and I don't think that's what it was telling me to do, but I always view those things as something actionable.

Sid: But, this is a very interesting thing to raise, which is happening, I think, in the collaboration space. So, collaboration spans so many different use cases. Like right now where you people are using video to talk to their friends and family. They're using Amazon Echos to drop in on various different family members, and they're using FaceTime for talking to their families on their iPhones and iPads. They're using Zoom for consumer needs, even though it would start out as very much a business application. They're using Chime to do that as well. Everybody's got all these tools, and they're using them within consumer context and enterprise context. And this is going to be one of the interesting things that happens in the tool space, is that Chime does something that's very focused on a business use case, which is presenting the region where your content is transiting through what specific region. It’s very important to do that in a business context. Doesn't make sense in an end-user or residential context. And so, this is why, again, we felt the SDK was the right posture. What the SDK allows you to do is make sure the customer experience that you’re vending, whether it's a telemedicine experience, or maybe it's a party app. In fact, we have customers today that are doing virtual tours and virtual consumer activities together, using Amazon Chime—games as well, that are also in that world—well, guess what? They don't need to put the region up. They don't need to present the region. Region doesn't matter. And in fact, you can capitalize and use the entire display to render video, and you don't even have to put any other labels out there whatsoever. So, the SDK enables that level of flexibility, and it puts it within, again, the application developer’s hands, so they can select what the right customer experience is for the context that the customer’s using video, whereas these tools are very rigid. Like even Amazon Chime the app is pretty rigid. It has a very specific way it goes about starting an audio/video call. It involves pins, it involves lobbies, and moderated rooms, and all these other things that really aren't relevant in a consumer context. So, rather than taking a square peg and trying to stick it in a round hole, which is what happens when you try to take these multi-party video, enterprise video collaboration tools, and then use them within consumer context, rather than trying to do that, you should use the SDK and customize the user experience for the context that your end customer is actually going to be using your application for, whether it's a real estate app to home showings, or it's a telemedicine app for helping doctors connect with their patients.

Corey: Is that going to happen on the actual Amazon Chime native app itself, too? I mean, given the fact that you're now signing deals with Slack, does that mean the application itself isn't a priority? I mean, to put it bluntly, does Amazon care about the app anymore? Or is that more or less a, you don't know how to deprecate things, so it's going to sit there in limbo forever?

Sid: Oh, we absolutely care about the app. There's a number of reasons why we care about the app. Obviously, we have customers who are using the application, and we are here to support customers, and I want to state that right up front. But the second reason is, the app is the perfect vehicle for exercising the features of the SDK. We are our own customer now. So, we use our SDKs to add features to the application. And so, there's a symbiotic relationship there. If we have a customer who asks us for a feature in the application, for example, there's a customer who's asked us recently for better noise suppression in the application. Well, that feature now gets built into the SDK, and the application then consumes that feature out of the SDK. So, everybody wins: the app developers who are betting on our SDK and using it routinely within their applications now get the benefit of features we build within the app; the users of our app get the features that we build within our SDK. And so, it's a virtuous cycle of being able to build features into the SDK, get benefit in the app, get feature requests in from the app, which lead to features within the SDK, which our application developers can now use. So, everybody wins. And the app is a very important part of that lifecycle. If we just built an SDK, we don't get that feedback loop on better supportability mechanisms, better audio and video quality things we could be doing. We just don't get the learnings from our customers to improve the SDK. And that's actually a differentiator, we're an owner/operator, we have to operate the primitive and also be an operator of the application. And by being an operator of the application, it helps us bar-raise the software development kit in a way that no other competitor who provides that SDK can do. So, it's a very important thing to be an owner/operator, and we are an owner/operator: we operate the SDK, and we also have to operate the app. They both feed into each other, and we're going to continue to do that in the fullness of time. Think of the app as the ultimate example code of how to use the SDK. Does that help, Corey, provide some context on the app and why we're going to continue to invest in it and support it?

Corey: Oh, it does. It absolutely does. And it sets me up, I think, pretty well for my final question for you, since we're coming to the end of our allotted time, much to the relief of pretty much everyone at AWS once they realize I'm talking to you. And that is that Amazon is a company that is famously willing to be misunderstood for long periods of time. So, my question for you is, what about Chime is being the most misunderstood right now?

Sid: I think the number one misunderstanding customers have about it—and this is our fault, and you know, I think it's an area of improvement for our service is we’re misunderstood to be a messaging app first, with a audio and video capability on it. We do it to ourselves in terms of how—if you launch Chime, the first thing you get is our messaging service. And so, people just naturally are assuming we are a messaging app with audio/video capabilities versus, actually, we're an audio/video app with a pretty rock solid messaging service as well. And that’s a misunderstanding that the team is going to work on, we're going to iterate on, and we're going to improve on over the coming time. So, that misunderstanding sets up for customer disappointment because they immediately assume we're going to have all the features that a Slack would have, or they assume that we're going to have all the features that Teams would have. And that's just not our core focus. If you look at our roadmap from last year, yes, we did improve the messaging service. We added a number of features to it, but the majority of our work went into improving audio and video capabilities within the application. So, this is a misunderstanding the team is going to work on, and we're going to try to help educate customers on what our position is, and what we're going to improve on. I think the relationship with Slack is a very good first step in doing that. By providing audio and video services in Slack, we're basically telling our customers where our focus is and where the majority of our investments are going, and where our roadmap is heading. So, I think that's a very important first step. But that's an area of misunderstanding that we would like to work on as a team.

Corey: That is one of the hardest parts about—I would imagine—being in AWS is when someone walks up to you, “Oh, AWS, what do you folks do?” “Well, pull up a chair, son. It's going to take a while.” Because the answer is yes, there always is. And Chime is one of those lesser-known services externally. In the Amazon ecosystem, everyone uses it and is very familiar with it, and I scared the crap out of people who didn't realize it was globally federated, and I would pop up asking them questions. It's given rise to a parody song: Chime after Chime. But it's felt always, for now, like an Amazon-specific ecosystem tool. I'm curious to see if announcements like this will begin to shift that narrative.

Sid: I think it will. Look, again a couple weeks ago, Mindbody used Amazon Chime in a way that nobody would have ever expected. I would have never expected to see virtual yoga sessions being hosted by Amazon Chime within the Mindbody application to be a thing. So in a way, this health crisis and the need to add video collaboration into various different contexts and support various different customer experiences with video and audio has just basically changed the entire way Amazon Chime is received by customers, by the overall tech industry, and by actual end-users. I got a direct message on Twitter from a student the other day who's using Amazon Chime for their learning, and education, and classroom activities. And she loves the app, and she was actually just like, why didn't I know about this? Well, there's a reason why you didn't know about it, which is we really had not found a way to introduce the great things we had done to the broader market. The SDK was an excellent solve to that problem, and now people are consuming us as an application, as an SDK, and we're getting broader awareness. And I think sometimes you have to do small things to fit within that broader construct of AWS. Surfacing our primitives and vending them to customers made us a more natural fit in the world of AWS, and I think that was a singular biggest improvement we've done as a service to help resolve that question of is Amazon Chime an external service? It absolutely is. It absolutely has use cases outside of the general AWS community, and we now have people from online learning, to telemedicine, to virtual classrooms, to virtual yoga sessions, and actually virtual experiences is the new thing I saw the other day, all using Amazon Chime. So, we've now managed to introduce Amazon Chime into all kinds of user experiences from consumer, to business, to enterprise, and it's been an exciting journey over the last couple months.

Corey: I'm looking forward to see where you folks go next. Because honestly, I think it's time for me to come up with a new joke for Amazon Chime.

Sid: I think so, too, Corey. I'm sure you'll find one, and help drive us because those jokes—we laugh at first, then it hurts a little bit, and then we internalize it, and we work hard to make sure you have to find a new joke about Amazon Chime. So, Corey, I think you should find a new joke other than we don't have customers anymore. We have plenty of them. And they're all across the globe in all kinds of use cases. So, time for you to find another joke, Corey.

Corey: You're right, and I don't know what that's going to look like yet. But at least for this episode, I think that winds up taking care of itself. There was, as I mentioned, a parody song was written by Spencer Johnson and then performed by Tim Leehane. I will play you out with that song, for those who have not had the dubious pleasure of a ridiculous song about an Amazon service.

Sid: Okay, there we go.

Corey: So, thank you once again, for taking the time to speak with me today. I really do appreciate it, given that you have literally everything else that would be a better use of it than entertaining my nonsense.

Sid: No, thank you, Corey, you're helping us get our message out as well, and we definitely can do a better job at that. And you're definitely have been a helpful vehicle for doing that, and we thank you for your time and spending the time to learn about Amazon, and Amazon Chime, and AWS. We really appreciate it.

Corey: No, and I appreciate you as well with all the hard work you do in suffering my egregious slings and arrows. If people want to hear more about what you have to say, where can they find you?

Sid: Just go to https://chime.aws and you'll land on our product detail page, and that also has links to our GitHub repos, and example codes, and customer references as well. So, just go to chime.aws and you'll land on our information pages.

Corey: Excellent. And we will, of course, throw links to that in the show notes.

Sid: Thank you.

Corey: Sid Rao, general manager of Amazon Chime at AWS. I am Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts. Whereas if you hated it, please leave a five-star review on Apple Podcasts anyway, and a comment telling me exactly what redeeming feature you seem to think Microsoft Teams has, even though you're wrong.

[“Chime after Chime” by Tim Leehane and Spencer Johnson]

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Jeff Sandquist

Jeff leads Developer Relations for the cloud at Microsoft, leading the team reinventing Microsoft's relationship with software developers around the globe. Their team is maniacal about making the world amazing for developers of all backgrounds.

They are excited to support and contribute to open source platforms, tools, and processes. As Developer Advocates, they’re spreading awareness of Azure and enabling developers to do what they love; write, code, and learn. Great online content (docs, demos, videos, code) is the foundation of everything they do.

They create global developer online experiences for Microsoft like docs.microsoft.com, Channel 9, and dev.microsoft.com. They connect with developer communities through their programs including Microsoft MVP, Microsoft Regional Director, their annual Build conference and third-party developer events around the globe.

Links Referenced

  • Microsoft.com/learn
  • Microsoft.com/learn/tv
  • bronconamedsue.com

Transcript
Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is brought to you by DigitalOcean, the cloud provider that makes it easy for startups to deploy and scale modern web applications with, and this is important to me, no billing surprises. With simple, predictable pricing that’s flat across 12 global data center regions and UX developers around the world love, you can control your cloud infrastructure costs and have more time for your team to focus on growing your business. See what businesses are building on DigitalOcean and get started for free at do.co/screaming. That’s D-O-Dot-C-O-slash-screaming and my thanks to DigitalOcean for their continuing support of this ridiculous podcast.

Corey: This episode is brought to you by Spot.io, the continuous cloud cost optimization platform, saving businesses millions of dollars each year on their cloud bills used by some of the world's largest enterprises and fastest growing startups like Intel and Samsung. Those are enterprises and duo lingo. That's a startup. Spot.io delivers the optimal balance of cost and performance by leveraging spot instances, reserve capacity, and on demand. Give your workloads the infrastructure they deserve. Always available, always scalable, and always at the lowest possible cost. Visit Spot.io to learn more.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Corporate Vice President at Microsoft of Developer Relations Jeff Sandquist. Jeff, welcome to the show.

Jeff: Hey, thanks for having me, Corey. Great to be here.

Corey: It's always a pleasure to talk with you. So, you are a corporate vice president: three words, each one of which tells me that Twitter means something bad. But the developer relations part is the rest of it. What does that mean? Where do you start and where do you stop professionally?

Jeff: I work in engineering. I work in Scott Guthrie's organization and we see developer relations as an engineering discipline. And we believe that's really something, kind of, simple. It's about helping developers. So, we build the services that run documentation, our learning platform, localization. Ensure that our products get to developers around the world, and are really able to be used by many.

And a big part of that is advocacy, our cloud advocates, and really developer relations, to me, is really about going where developers are and being able to connect with them authentically. Advocate about a community, Node, Python, and really bringing them inside of Microsoft and helping them understand why we make the decisions we do, and what will it take for them to be able to use our products. At the end of the day, we're just here to help.

Corey: One thing that I found that was super interesting, at least from my somewhat naive perspective, is that I know an awful lot of people who work for Microsoft Developer Relations, or specifically over in Azure, and they say a lot of things, oh, do they say things. But one thing that they don't say is nothing that has ever come across my radar as an explicit sales pitch for Azure. So, the old trope about Developer Relations—DevRel, being a Devreloper, whatever term you choose to use—it does not translate in your interpretation of it into a sales force with credibility, as best I can tell.

Jeff: We help first. And I think—maybe even just, kind of, step back is that a few years back, we made a decision to, in a lot of ways, rebuild, build-out, really our advocacy efforts at Microsoft. And we didn't start there. We actually said—the first thing we did was, “Wow, we need to go build great docs.” And that was really one of the first problems we had to solve. How do we build great docs, not just for .NET Windows, but for Node, Python, Go? Azure is a cloud that really any developer should use. And it has to start with docs.

But our advocacy work is really founded in the principles of, everything starts with great content. We go where developers are, and we bring them back to Microsoft. And our approach, when we went to restart things, was really about aligning with communities. So, I’m sure at one point in time at Microsoft, we probably had a Windows start menu advocacy, or evangelism team or a team that was from Windows Server. And when we started developer advocacy at Microsoft, we built our teams around communities.

So, I have a team that focuses on Node, and those people come from the Node community and live with that community. And they spend all of their time advocating on behalf of that community, in engineering. How do we need to be looking at our docs? How do we make this product easier? How do we make it five minutes to “Wow?” And around that—of our people being from those communities. It gave us a connection to those communities, but I actually think it actually changes the tone of the conversation because we're not taking someone from Microsoft that may have grown up through experience with .NET or Windows or Visual Studio and saying, “Hey, be a Linux person,” or, “Be a Node person.” That's almost like asking somebody to be a poser. We hired people from those communities that are genuinely part of it, and we embraced that. And I think that's where that comes through. We also—these people aren't on quota. We're not in marketing; we're in engineering, and we run this as an engineering team, with sprints. It’s really about connecting with those customers.

Corey: And to that end, we're having this conversation right after your Build Conference finished, which I got to say was nothing short of astonishing from my perspective. We went into worldwide lockdown about three months ago, give or take, which means you didn't have a whole lot of time beyond that to transition the event into a fully remote digital experience, to borrow from your marketing's overly corporate phrasing, but you really pulled it off in ways I was not expecting. Oh, great. It's going to be 48 straight hours of content—and it was—but it was structured in such a way that made sense. A certain competitor of yours just announced the same week that they're going to be doing their online conference for eight weeks. I imagine that they could easily spend that much time talking at their customers, but I'm not sure anyone wants to hear it. This was, it felt like, the perfect amount of time, and it leveraged the remote aspect in a way that I was not expecting. I frankly expected it to be a terrible half-assed version of an in-person conference and it was very much not that at all. How did this happen?

Jeff: How did this happen?

Corey: Explain yourself.

Jeff: Yeah. Sometimes I wonder how it's happened, too. But how did Build happen? Build happened because of, just, a phenomenal group of people, like, an absolute team sport. I had a great partner on this, a guy named Bob Bejan, who's been working on events forever. And about eight weeks ago—probably shorter than that—we knew that Build was not going to happen. And you saw events being canceled and so forth, and we made the decision in the company that we're going to still have Build. We were going to make it free, and we're going to make it online.

Corey: And both of those make sense. But I expected the third bullet point there, that you didn't hit, was, “And you're going to postpone it.”

Jeff: Yeah, and we did not postpone it. And I would say—I mean, I use a quote that Scott Hanselman was quoted on The Verge in an article he said, “Hey, we didn't just pull this out of our butt. We created an all-new event.” It was quite eloquent.

Corey: Oh, yeah, there was some joking about that going around on Twitter, but I’m in favor of quotes like that. It’s, “Heaven forbid that due to that quote, one of our folks sounded like a human being.” Yeah, that is not a failure mode in any realistic sense, for any customer you want to deal with.

Jeff: And being a human—what a great way to put it—that's what this was all about. So, eight weeks ago we rewrote the event. We started from the ground up and said, “We cannot run this with somebody for three hours in front of a podium going through product announcements.” We have to design this for, really, the attention of the internet. We have to make this entertaining, and really even being careful because COVID-19 is going on, people are going through all sorts of different things.

But we knew we had to bring the community together. We knew that we had to go connect with the community, both ourselves, and we knew, if we were going to do this, we had to do a few things, which was, one, really start and build an event and adjust the overall flow and timing of it. So, we didn't do those long keynotes. We shortened them up as small segments. We did Imagine Cup, which is something we've done always at these events with our students. That was a shorter, smaller bit.

And we really started writing the show. When we work on an event like this—we always do this—we treat it like a movie. We have a writer’s room, three times a week, we get together and we start adding people into that group, and we start writing the story of the event. Now, this isn't like the story of what product news we're landing and exactly what that announcement is, yet. This is about talking about the overall story of the event, and we start writing that out.

And I think it's really important to know is that we started this eight weeks ago from the ground up. We were planning an event called Build. And so, we decided that no, we're going to do a whole new session process. We're going to shorten these to small segments. There's a lot of us at the company. I started a thing called Channel Nine with a group of people quite some time ago, and it was progressive at the time. It was really about—

Corey: Oh, yeah. Big fan.

Jeff:—getting people with video cameras and allowing people to meet the people behind our products. We have a lot of people that work with media, much like you. I'm a kind of hacker at it as well, but love the space and connecting with people on social. And so, there's a lot of people that got it. We cannot just go up and push sessions out. So, as we crafted the event, we started really working with our speakers.

And we ran speaker training. And the speaker training was really about how do you hold an audience? How can we ask the questions differently? What are the things that streaming does to bring an audience? And how do you do these demos? And that became this area of the event you saw Scott Hanselman, do a keynote where he completely changed it based on the modality. Bringing people into meetings, via Teams, and bringing people into the overall keynote. And I think really what happened is it, kind of, became a life of its own. You can imagine a company after you do an event, and you [unintelligible] look the press out there.

There's been phenomenal engagement numbers around the event. What was so special about this is it truly felt like it was just a solid core group of people that cared about developers. And it was just this like, “Hey, how are we going to do this?” The answers were always, “We'll figure it out.” And I'll be honest, there were some areas, if you go around the event, you're going to see where it was rough around the edges. That's because we're repurposing infrastructure for session videos and live that maybe we built five years ago, we weren't quite using anymore. And as the time started leading up to the event, and we got ready to announce registration for it, could not believe what was happening for, really, just sign up, and getting people to attend the event. Let me look for some of these things here as numbers.

Jeff: Okay, so, so, couple things. We know, right, there are people around our company, we live and breathe community. We're part of those communities. We live there and we wanted to get out and get with them. We were hearing from people, too. Our customers were quarantined at home, too. A lot of people looking for online content to consume. And let's just, kind of, talk about the event.

At the start of the event, we were well over 200,000 people that registered. I want to put that in perspective. I don't even—[unintelligible] weren’t quite sure what to do with that number, Corey, because this is a free web streaming event. You didn't have to register, but 200,000 people decided that they would. And we made sure that the streaming was everywhere. The idea we wanted was we wanted no friction. If you want to watch the keynote, great. Watch it or watch the sessions, you can do—so, via Twitter, you can do on our own properties.

Corey: And they don't appear to be edited either, so if someone had spilled a cup of coffee over themselves, you can still grab that and just go loop that as your new Zoom background.

Jeff: You could, and I'm sure you'll find some, some happy—

Corey: Sorry, Teams background. I forgot who I'm talking to for a moment there.

Jeff: Maybe a Team's background. But hey, we go where developers are. So, sometimes we might be using Zoom, we might be using many of those things. But, we're at 200,000 registrations. Guess what? People still kept registering, even when the event’s going on. Tens of thousands of people, and they were in chat rooms, and attending those sessions, and it was going on on Twitter.

Guess what, though? You know what the average viewing duration was of our users? 161 minutes. 95% of our presenters were remote. Heck, we had several hundred people delivering their program from their bedrooms, or home offices, in their kitchen. There's a global audience. 80% historically was US attendees for events, 20% around the world. 65% of our audience was from around the world, and we made the decision—and I was so proud of the team—we're going to go 48 hours straight. And there's a very small number of people, probably—what is it—like, 0.01 percent of our customers, really, of any company that ever get to a major developer event.

And we didn't want to do something where we just showed up to the US and, hey, did something great and ran pre-recorded materials through the night. We wanted to go around the world, and we wanted to follow the sun. When we did that, sessions that we did during the day, when it became the evening, we ran them again. Scott Hanselman did this great talk on really this .NET futures talk. They delivered it last night at two in the morning.

Now, sometimes we did pre-recorded talks. And we’d run them again. So, we showed up in the chat room. We're upfront about it, “Hey, this is pre-recorded. That's why I couldn’t answer your questions, but I'm here to answer that.” And it was really about making that connection with the community and paying it off for our customers that are all around the world and giving them an outlet. I'm really proud of our team, too because what we did was we did a live—a sports test, so to speak. And that was my advocacy team. They were out there and they're great online. They're great at making things interesting.

Corey: One thing that really stuck out to me was my favorite talk of Build was a pre-recorded talk that Emily Freeman did, where she talked about building remote DevOps cultures. And she's always a good speaker; always gives a good talk. And at the end of her pre-recorded talk, that turned into a screen split with someone that she was talking to who asked questions from the chat that people had asked just then, and she answered, and oh, my God, it wasn't pre-recorded. This was done live. And that completely caught me off guard at the end of that just because it was so well polished, and no glitches whatsoever that I had just assumed that, “Oh, yeah, of course, it's going to be a video thing. Why wouldn't you pre-record it?” It really forced me to reassess how, I guess, proficient at this, both she was personally, as well as the entire team.

Jeff: Doing the right thing at the right time, kind of like having the right Lego blocks. a little bit further back, probably about 12 weeks ago, we started seeing some of our key flagship events having to be canceled due to COVID-19. And our team, we live in media. We're doing videos and short bits, and we build an overall skilling platform that's part of this. It's pretty interesting.

But, we started, kind of, building on small nuggets of content. And so, what we tried to look at is like, “Hey, what's the best use of time? And how do we make the most of this?” And I think when you go to an event, if you're like me, it's hallway conversations. I want to get a question answered that I can get answered nowhere else. So, maybe for Emily, it was like, “Hey, let's do a pre-recorded talk because of timing, but let's be there. Let's be with a chat with our customers.”

Corey: Just to be very clear, it wasn't a pre-recorded talk. It was so polished, I assumed it was just absolute clarity on that point.

Jeff: Okay, that is clarity because I was at, actually, Emily’s talk, and watched [crosstalk]—

Corey: Yeah, and it was amazing.

Jeff: Yeah. And I'd say, you know what, a lot of the people that are across our team and company, they're comfortable with doing online. The advocacy team that I lead is about online outreach. We [unintelligible]. We go to third party events, and communities, and even our first-party ones. But it's about connecting with developers where they are, and guess what? We go where developers are, they’re online right now, and so that's where we are. That's not just video. That’s GitHub, contributing to open source projects, but I do believe at our company—and this goes way back—Microsoft, in many ways has probably, and easily the most liberal social media policy of any company around the planet.

We trust our employees to get out there and communicate authentically, and we say, “Come as you are and do what you love.” And we mean that. We mean that; when you're speaking at a talk be Emily. Be Jeff. Be Corey. Be an engineer, and be that person that a developer wants to talk to. Even how we build up the events. Satya will say this over and over, “Hey, come work at Microsoft. But make sure you use Microsoft as a platform to be able to do what you believe in, in the world.” And he really means that, and so you saw things around build for good, or areas that we did certain things with students. Those are areas that people at our company will have a passion for, and it manifests in the event.

And so, when I was talking about that writing room and how we write the story of the event, that is a lot like when people write a TV show, we're in that room, we're coming up with ideas, not all of them do we go do, thank goodness, but that is really about how do we deliver on something great for our customers, and something for the community, and we will really mean that and that comes through. It's paid off with what I believe we've got some of the best presenters in the world, in the industry, in our company, in our team. And guess what? They're all making the adjustments. Not all are just advocates that are used to being online. Through this, we purposely reworked our speaker training and dry-run approach to be designed for this medium. And we helped one another. By the way, lighting, this is the Achilles heel.

Corey: Oh, my God, yes. Same problem in my world. The camera, I can make that work. Audio, I do podcasts and I sound amazing. But lighting is the bane of my existence for these things.

Jeff: What's awesome, though, and I mentioned a guy named Bob Bejan, who was really my partner in crime on the event. He's got years of TV production, and it's really this fun dynamic. We have an amazing group of people that are pro-style, they produce our shows, our events, and videos, and it's just a phenomenal team. And I've tweeted some behind the scenes pictures of things that were happening around the event. Now, they're pros at a lot of this stuff, from lighting and lighting sets. And I was talking to one of them and said, “Hey, I got a TV in the back of my office that I want to run and play videos. I don't know how to do it, I have these key lights, and they're showing up. And I'm literally chatting with one of the producers. He says, “Oh, you need to get some gaffer tape.” And I didn't even know what gaffer tape was until this, and he helped me understand how I build, kind of, a crown with a gaffer tape around the light.

And this is actually just a fun part of the event is, I guess, [unintelligible] COVID-19. We're all trying to figure out how do you lead teams? How do you connect with developers and do so online? We're all figuring it out together. And so, you see many of us, kind of like, literally decking out our battle stations. A year ago, we probably had a nice, simple, minimal desk, and now we're in this battle station with microphones and Stream Decks.

But what's been fun about is we're learning together. People are helping one another. We're learning how to even use things like OBS. We’re pretty nerdy about it. And we're building ways so that we can really connect with people, because right now—the world needs community now more than it ever has, I totally, completely see that and believe it. It's lonely. We're at home. We're trying to connect. We're not getting up at conferences to meet with the people that we care about, too. And we knew we had to go create something with that. And what's been fun is watching people learn new ways of doing that and how excited people get when we're able to connect to them. But it's seriously like when Satya says, “Hey, we did two years of evolution in two months.” Totally, with an event like this.

Corey: One thing that I think you nailed that is a common, if not the most common failure mode, is you get the lighting finally mastered. You get video taken care of. You get the sound done. The upstream is great. And all of the production quality becomes first class and the content is garbage. Where it's boring, it's crappy. It’s nothing anyone cares about. It doesn't matter how well-produced it is if it's crap, whereas people will forgive an awful lot of production snafus for content that's engaging and fun. Ideally, you hit both. And in your case, you did.

Jeff: Oh, thank you. You know, everything starts with great content. And I think when we start building the event, it's by developers for developers. The people building this event are developers, or we're a developer. And it's really about what do developers want to hear? How do we help explain what we're building at Microsoft? We're excited about what we're building. We want to bring people inside and we want to let them understand why we're building things.

We want to be able to share with them why certain things would have bugs, or certain ways of using it, and where we're headed. And we're out here to listen. And really how do we make this product better? How do we hustle? How can we make it that people want to go use it? A lot of the workaround this one feature, it's this Azure Static Web App. We just came up with that, and this is an example of an area where our advocates, especially for people that work with the Node community. How do you make it really easy to deploy static websites? And I lived in the weeds with an event and I lead a big size team that is doing this, you imagine a lot of times I'm doing anything but writing code.

Jeff: So I was talking to John Pop on my team. And John is from the Node community, he's a cloud advocate. And I run this website that's basically static HTML, and it's running in Azure. Really simple. And I use it just to keep up with deploying it. And there was an area of Azure that John and team really, really, really wanted to make sure that we made it easier for the Node community. And so, we released this thing called Azure Static Web Apps.

And I was talking to John and I said, “How hard is it for me to move over to it? Should I?” And he said, “Which website are you doing?” And I said, “Well, it’s Bronco Named Sue. It’s for my Bronco. And I want to move it over to this because I want to start doing some more things with Node and React with it.” And he goes, “Where's your GitHub repo?” And I said, “Okay, here it is.” I gave him a link to my GitHub repo for it. And one minute later, he came back and said, “You're live.” And literally, based on that repo—because it was public—in two minutes, he was deployed and running on Azure Static Apps.

And for me, one of the things we want to really enable and really get for developers five minutes to wow. Okay, I didn't know what this thing was. What are the docs? We are grounded. We live in docs, everything starts great technical docs, period. That's what developers are. And if we can take them from the docs, to deploying something, that's what we want to try and do. Now, we don't do that all the time. But that's what it's all about for me is how do we help?

Corey: One challenge that I had as I was, I guess, digesting the firehose of announcements that came out of Build was I consistently felt, to be very direct, lost. Where you folks were talking about a service I was either directly or tangentially familiar with and then seamlessly transitioned into talking about things that are very, to be direct, Microsoft ecosystem, which is not a world I am particularly well versed in for the past decade and a half. So, on some level it felt like I was either missing obvious things, or I was not up to speed where I needed to be. In truth and in practice, it was aimed at folks who are much more aligned with the broader Microsoft ecosystem than I tend to be in large part. But I'm wondering, for someone in my position, what is the best on-road to, I guess, learning more about this that doesn't involve: step one, go work somewhere that's steeped in the Microsoft ecosystem and spend five years learning all the ins and outs.

Jeff: Well, first, you got to hang out with us a lot more.

Corey: Oh, there we go.

Jeff: That's number one. But I think we have at Microsoft, some of the best ways to go learn about our platform. And I think it starts with something called Microsoft Learn. Microsoft.com/learn, and really, where the [00:26:13 unintelligible] is think about, like, TryRuby for the cloud. And it's really about the fact that how do you learn today? People have all different ways that they want to learn, and we believe that you want to make it so that people can do 10 minutes here, or 15 minutes there. And so, I think one of the best ways for you to start spending time going through Microsoft Learn.

And what's, kind of, unique about it is, it’s typical training, but it's small bite-sized chunks. It's really built around learning paths where you can do five minutes here—and guess what? If you need that little section, say something around identity, and you come down to another concept that you want to learn later? Guess what? It's checked off and you don't have to do it. And as you work through and answer questions, and assess your skills, you get to do a few things.

One, we have this thing called Cloud Shell. It's basically our command-line interface, right in the cloud. When you need to deploy a VM, our Cloud Shell pops onto the screen of Microsoft Learn, you start typing command-line commands for deploy a VM. And guess what? It's free. That's a subscription that I run. We set some group policy around it, and you get to go use it free. And so, in your company, you don't have to worry about somebody accidentally deploying some Hadoop cluster, not knowing what they're doing, and probably driving costs up. It's probably something that you know a bit about of—

Corey: Oh, maybe once or twice.

Jeff: But it's really about giving that way to go do it, and really learn the platform. Now, we don't just sit there and go do simple if/thens, we actually look at the deployments of the users against that, to give them points. And so, if somebody takes the default settings in the learning, and applies it as is, they get a certain set of points, but hey, maybe they deploy to a different region or different data centers. They basically build up those skills. And that's something that we've been building up for about the last two years, and it's been unbelievably successful for us.

We've had about 72 million monthly active users to our technical docs and learning sites. But on the Learn platform alone—and this is relatively a new thing for us over the last couple of years—about 3.9 million registered learners now. And you go from February to March of this year, it's like 25% month over month. Frankly, we're about 272% year over year. But this Microsoft Learn: get started there. But you know how developers learn, and I think how we think about our developer relations work is it's two in the morning, inspiration strikes, you're a developer, you don't go to the marketing pages. You don't go to reading the glossy brochure. You sure as heck don't go to your procurement manager and say, “Hey, can I get a license of this?”

You go to Google, maybe a few percentage of you go to Bing, and you start typing in search terms, you start typing in codes, things like this, and where do you end up? You end up at your docs. And that's why we really focus on having great docs that are localized, at least across the 17 languages that we localize Azure, but maybe up to 65 locales around the world. And we care deeply about linguistic quality. We have humans, both in the community and outside the community, that makes sure that is of quality, and we're maniacal about our docs. And we wanted that to be one of the first things that a developer sees, because that, to me, great docs is about the ultimate source of empathy. James Governor said this a while back, and it's true. How do you get somebody started? And it really starts with that great content. And we care deeply about that total renaissance of technical documentation at Microsoft over the years.

Corey: This episode is sponsored in part by N2WS. You know what you care about? Many things, but never backups. At least until right after you really, really, really needed to care about backups. That's what N2WS does for your AWS account. It allows you to cycle backups through different storage tiers; you can back things up cost-effectively, and safely. For a limited time, N2WS is offering you $100 in AWS credits for setting up their free trial, and I encourage you to give it a shot. To learn more visit snark.cloud/n2ws. That's snark.cloud/n2ws.

Corey: Oh, the documentation is nothing short of spectacular, to be very clear. I wound up pulling up the page you folks have, which I think is a great page to have, by the way. Explains Azure services through a lens of what AWS’s equivalent are. And I don't know Azure services for beans, but I can quote chapter and verse in the AWS side, and I was all setting up to just tear it apart and spend some time dunking on you folks. And I couldn't, it was really, really well done. The other only one or two minor things I saw and they were more stylistic than anything else. This is an actual legitimately good resource. In fact, the only thing that causes any skepticism around it is the fact that it’s Microsoft on the top of it. It was very even-handed and it got it all right. I don't rave about documentation all that often, but you folks have really hit it out of the park.

Jeff: It's hard work and it is a team sport, I'll often say. This didn't happen overnight. Our workaround documentation and really putting a focus on it has been something that we've been after for many years, and we'll never be done. Our docs at Microsoft at one point in time, I think they were scattered across 17 different websites around the company, all varying areas of quality. And probably were very emphasis on .NET Windows.

Now, we love those, but we're the cloud for Node, JavaScript, Go, Kubernetes. And so, we had to go really, kind of, rebuild our technical documentation from the ground up and really build a platform for it. We went and talked to the people at Stripe, the Twilio; people that really nail it for developer experience, and we started small. I remember this one night—and we made the decision that we're going to go after this, and at any big company, it isn't just a top-down, hey, we're going to do this and it happens. Maybe that happens over another cloud, but you really have to go work with the community and the company to make this happen—and I remember it was late at night.

It was about two in the morning, and I was walking around with my team, and we were trying to figure out how we're going to go build a doc site for our company, and somebody said to me, “Hey, Jeff, what are we going to build for a CMS?” And I said, “You know, everybody I've ever met who's built a CMS either failed or is fired.” I said, “Let's not do this. What if we built her on GitHub?” And this was years ago, and this was one of the best moves we did. We really standardized on GitHub for our documentation platform. And that gave us a couple things that were really wonderful.

One, you're an employee and you want to make an update to our docs? It's a pull request. Docs are just code. You’re a customer, and you want to update a doc because you say, “Hey, this is not accurate,” and you want to do it. It's just a pull request. The format and the tools that we really use, it's just markdown. So, we're able to simplify overall the tooling that we use: it's just markdown. We're able to make it that, to contribute to our docs, whether you work at Microsoft or in the community, it's just a pull request, docs or code.

And then, we just really were able to build, frankly, a very simple platform that was around GitHub where developers are, and make it run great for people to be able to update it and participate on it. And it's been a number of years working at it. And it's not just about the platform. It's also working across the content in itself. And this starts at the top. I mean, literally the entire top of the company. And absolutely Scott Guthrie, who runs our overall cloud and AI, he absolutely is a champion of docs, and really, we'll even run through product reviews. And as we're going to launch said, “Okay. Let's start going through the docs. What do they look like?” We spend time on it.

Corey: I have no trouble believing that because this level of documentation does not come from someone saying, you know, we should really improve the docs one of these days. This has to come from the top. It has to be a strategic initiative and one that has paid off handsomely.

Jeff: I don't know if a strategic initiative, actually, will do it. It has to be part of the lifestyle. And sure, we have an amazing team, they report to me, there are technical writers. To me, it's one of the most underappreciated disciplines and crafts of our entire industry. And I think some companies do a disservice to that role. Not at Microsoft. But documentation is not the responsibility solely of a DevRel team, or a docs writing team. Docs are about building the product. And so, as we build docs—we will write the docs before we write a line of code. Our docs can be written by our PMs, our engineers, and frankly, it's a badge of honor to write great docs because it's hard.

And you know what? If the docs takes 70-some pages or 17, or 7 to write, and that's too long, it probably was not a problem with the writer. It's probably a problem with the user journey. So, why wouldn't you want to start writing that out from the beginning? And so, it's not just having a great docs team, or not about just building on GitHub. This is cultural. And Microsoft is a developer-first company. That's how we're founding it. So, not only do we talk about this at Scott's level, I've been in Satya’s leadership team meetings where we've talked about docs, and Amy Hood will talk about, “Oh, my gosh, I sent this to a customer and it works so well. And these docs are great.”

We are talking about documentation at that level because it matters so much in our company. And there's no point—this is the very first thing that Scott talked to me about when I was thinking to come back to Microsoft. He said, “There is no point Jeff in going out and doing evangelism, advocacy, or developer relations if you cannot go on stage or go online, and after you've finished a talk, say ‘hey, you want to do this, go to ak.ms/this and get started.’” You have to be that way, and it's cultural and it's lifestyle. And you can tell we're really proud of it, but we're never going to be done. We're never going to be done with our docs. We're always going to be updating them. And that's where we're lucky to have our own GitHub because it makes it fairly easy.

Corey: One thing that was challenging for me is shifting my mindset away from the lens that I normally look at the cloud space through and coming to a Microsoft specific one. I was given early access to a lot of the announcements through the analyst program, which was appreciated and also useful because it turns out I had a really bad take the first time I saw a particular announcement: namely the Azure for Healthcare offering. And my immediate thought on that was that, wow, that's really dumb. It doesn't make any sense whatsoever. It's bifurcating the market and more or less just distracting people from the things that are truly important.

And that's the right perspective to take for other cloud companies. But then I got to thinking, wow, think twice, then write. And what a concept, embargoes help with that. And I realized that Microsoft has done exactly this—the industry specificity—for many decades now, and it has worked out profoundly well for them across the board. So, I'm looking at this and realizing no, that's not a terrible idea at all, that is the right differentiation direction to go in. But having time to think and absorb something that goes beyond the 30 seconds, it takes me to write us a crappy tweet was extraordinarily helpful, and it led me to wonder what other things are being perhaps viewed unfairly through the lens of, well, if another cloud company did this thing, it would be terrible, therefore, it must be terrible if Microsoft is doing it, too.

Jeff: Some companies might call that being customer-obsessed. But really, just on—

Corey: Careful, they may have trademarked it by now.

Jeff: They may have. But I think that's the first thing on the healthcare is really about listening to customers and really about—look, we've been out in the enterprise, and how we really are able to deliver solutions and things for its customers is basing them on what they're looking for. And that there is based on experience, and listening, and learning our customers. Where are we misunderstood in other areas like that, I think it’s—what I would want to make sure that if somebody listening to this podcast and they're saying, “Look, why would I trust Microsoft?” Or, “Why would I want to go learn more about something of us?” Or, “What do I misunderstand?” Is we're a developer-first company. And I've worked at other companies, but we're not founded in retail, social networking, digital advertising, any of those, and founding moments for companies matter.

And don't be confused about this, we were founded with two nerds that were basically building tools for developers. BASIC for the Altair: it was Bill and Paul. It was the very first thing that we were doing as a company. That's our founding moment. And what was the first thing that they did after they got done? They went off to the Homebrew Computing Club and went off to go share what they built. They did it at an event. They did it about trying to share something that they truly believed in and were excited about, and wanted to share it with other developers.

And if you look at us at Microsoft, and you're wondering what makes us tick, it's that founding moment. And I think what I love about the Microsoft that I came and returned to is we care about developers because we are, too. And because of that moment, understand us that when we are trying to ship products, we're iterating like you. We’re trying to make it better. And frankly, we get really excited about the work that we're doing. And maybe we can do a better job of, kind of, giving more context, but you ask anyone from Microsoft, “Why does this work?” Or, “How can I make this work in my environment?” We are hungry for the business. We're here to hustle, and we're here to learn from you. And that, combined with our founding moments is really what drives us.

Corey: One other thing that really is, I guess, challenging for me, again, not being steeped in the ecosystem, is I took a look at the things that were announced, and it turns out there's kind of a lot. What are the highlights? Some of these things are very clearly aimed at, if you're using this product, and using it with this other product, it is very clearly a win for you, but the rest of us are sitting around trying to figure out if those are real products or things that got made up. I have it on good authority that something called Dynamics 365 does exist. And there's this whole Power thing, as well, that is not made up just to troll me, but is in fact viable business in their own right. But looking at this from the outside in, what are the interesting key takeaways? What are the easy on-ramps and what are the notable changes that were announced?

Jeff: Great question. Okay. Let me think through a couple things. We have some amazing big announcements over at Azure. You read the blog posts, you walk that through, but let me give you a couple things that I think people should pay attention to, where I think there's opportunity. Number one, you made some comments about Teams. But if you're looking at building something in your company, and so forth, or maybe even building, kind of like, the next thing, or doing a startup, go look at Teams. Now, Teams is not Zoom, and it's not Slack. It does those things there.

It's a platform. And it's a platform that is real, that you can go build on. You can write Node to go be part of Teams, and it is an overall platform. And I would say Microsoft Teams is probably the single largest developer opportunity that's new. On the planet right now, period. And I use Slack, and I've used all the different products over the years, but there's something very unique right around with Teams that you have—and M365 that you cannot ignore, due to the growth of that overall product. And A lot of it's due to the unprecedented times, but I think that as people and organizations really get used to the way that probably like a lot of people who listen to this podcast work, that is going to be a platform where there's going to be opportunity. The second one is just developer productivity. VS Code that is used around the planet. We are all about how do we make it easier for you to do your work. How do we make developers write less code? How do we make it quick and easy? You know, we're showing different ways for updating, even iOS apps are built through .NET. We talked about all the different workaround code spaces, and how do we make it just easier for developers to do what they love? But there's another aspect of it is Power Apps.

Corey: Tell me more.

Jeff: Power Apps is an absolutely—it’s no-code, right? And I know there's lots of people in the valley that are talking about no-code, and it's all the new, new thing—

Corey: I'm a huge proponent of the whole no-code movement. So, please, you have piqued my interest.

Jeff: Yeah, we've been doing it for a while and we have a real project there. And it is so important. What I love about Power Apps is, one, I remember back years ago, when I was an IT admin, and you wanted to have somebody in line of business area do something, to go build an app. I think you probably gave him an SQL Server password and the SA accountant said, “Go to town,” and they built it around Excel. Power Apps—you know Excel, you can build an app. And it's going to be GDPR, it's going to be really about something that you can enable with all the controls. And your developers are not going to have to build those apps, because you're going to enable people within the business to go do that. And it goes beyond that.

Developers, it isn't just about them saying, “Hey, great, I don’t have to go build this app.” There's certain areas where our templating and different things that we enable through Azure—our portal—that are built around this, and the developers themselves should be looking at Power Apps saying, “You know what, I really don't want to write code for this. I want to spend my time writing code for something that's really going to need this. And where can I go and make it so that we can have these no-code solutions?” Because there's so many apps that need to be built around the world, that there's not enough developers for it. There never will be.

And I absolutely adore Power Apps. Think of it almost like our VBA in a good way, that connects our cloud together. Think about it, how we actually bring together really, we say, M365, but that's our productivity cloud. How do we bring things together and connect that back to Azure? And so, really, the next one is, look at Power Apps. In these times, do you really need to build custom code for every app that you're doing? No, please don’t. For the things that can and you can do some really compelling applications with it. That's Power Apps.

Next, look at all the things that we did about building the best, really, developer workstation. Go watch Hanselman’s keynote. We talked about all of the work that we're doing around Linux and really enabling all of that, from GPU to our new terminal. We want Windows to be the best darn developer box that you can have, and we want to make it so you don't spend two days setting up your dev environment. We want you to go five minutes to wow. And really, that is about developer productivity again. Those are a few things that stood out to me.

And they're just areas that we're more even personally of things that I can go use, because I don't just run an awesome advocacy team. I run a service engineering team. I've been really living the move to the cloud, like anybody. I've got legacy systems that I'm bringing online and hundreds of engineers that build services for Azure, as well. And I lead a dev team, too, so I'm looking at these not just as somebody at Microsoft to share to the developer community, but as a leader of a developer teams as well, too.

Corey: So, I've been fairly public about my love of the whole no-code world. I write a sarcastic newsletter every week that people really should be subscribed to and if they're not it's called Last Week in AWS. Slap a dot com on the end, and there you are. But the way I do this is with a bunch of lambda functions, specifically at last count, 27 of them behind four API gateways. And tying this all together with scripts was not workable. So, I don't know how front-end works. I'm terrible at it, and I get more confused when I end than where I start. But I found something called Retool that got me pretty far down that path, where it just hits API endpoints.

And it's drag-and-drop for a web interface for internal apps, which was effectively life-changing for my perspective of these things. It feels, to be blunt, like Visual Basic for Web Apps, which was exactly what I needed, and is the fun cherry on top. I did a little digging into oh, what am I talking to when I connect to this website? It all runs on top of Azure, which is fascinating to me. It's oh wow, I accidentally trip over an Azure customer in the wild. It was really just a glorious thing, start to finish. Later in time, I wound up bullying them into sponsoring a couple of things, and I’m just a fan of what this unlocks. Suddenly you don't need to go to cloud school, or developer school to learn how all these things work, you can have a business idea and put that together quickly and easily.

Jeff: So, let me tell you a story. And this is one of my favorite ones about Power Apps, and it’s from a while ago; there's great video out there. And there's a fellow, he was at Safelite Auto Glass. And he worked in the claims adjusting group, and so he saw what you'd see in many companies: hey, somebody has a window auto glass that needs to be fixed. They have a mobile adjuster, comes and looks at it. They fill out a PDF. I think then that was uploaded somewhere where somebody turns it into another PDF—the ongoing story of just inefficiency. And there's a fella, he was not a dev. He worked in the claims department, and he went there's got to be a better way.

And on his own he got to Power Apps, and he went, “Wait, I can use this. Literally, I can build a mobile app where my friends that are claims adjusters can literally just bring it up on the phone, and we can make it that we don't need a PDF. I can do this PDF, we can actually bring it right into the systems.” And so, he did that. And the app was used around the UK. There was a great video on YouTube around this. And guess what he did after that? He started building more apps, and really, his career totally changed. And now he is deploying these apps and building them out for the company for all of Safelite. And so, not only did he totally help change the company from how they're doing tooling and how they're automating, he was on Microsoft Learn, he was actually able to build more skills and invest himself and it's been a game-changer, both for him professionally, and in his company.

And those type of stories again, and again—I don't know how you started, but I started on a Commodore 64 and in my basement and it was that discovery approach of trying something. Did it work? Constantly pecking at it. And this story, I tell it again and again, where there is somebody that was in a department, no IT resources, that completely changed through Power Apps, and automated something in a way that would not have been possible, and did so where he was able to totally forever change his company. And there’s story and story like that again, and again, about Power Apps that it's unbelievable. If you can mess around in Excel, you can build a mobile plus web app, and you can do it in around your company. Frankly, Teams; we announced Teams, you want to build apps around Teams? Great. Go do it through Power Apps, too. And I think there is a lot there, and it is a place to pay attention to.

Corey: It'll be interesting to see how the messaging continues to evolve because from what you're saying, it sounds like this world of Power Apps is also accessible to folks who aren't already deep into the ecosystem. It sounds like a very reasonable on-ramp for folks who might be working on other platforms who are, sort of, across the map picking best-of-breed things from here and there. It feels like it's a very easy on-ramp from what you're saying. But if you look historically at things called power, it always felt on the other hand like it was one of those, oh, this is only for very Microsoft-y companies.

Jeff: No, I mean, it's for companies. We aspire—and in a lot of ways are—we want to be the platform for every developer. And that's everybody from no-code, all the way to an architect that's pulling together, to a data scientist. You should go take a look at Power BI if you're not. You're crazy if you haven't been on Power BI, they have—

Corey: Oh, I’ve looked at Power BI, that's a whole separate kettle of nonsense.

Jeff: That's actually part of Power, though, as well too. Power Apps, but that's over that power suite. And Power BI is so essential Because of how quickly you can put together dashboards together and get that data into the companies. It's all of these things combined. Don't just look at us as Azure South Lake Union, of course. It's about VMs is the canonical unit for everything. But for us, look at us as an entire platform. Sure there's Azure and there's our work around AI, but it's that combined with Teams and Microsoft 365. It's a cloud. It's a cloud on the enterprise.

And all of these pieces together, from Power BI to things that you can go do around Teams—go look at Fluent. Go look at what we talked about—data build is another example of something that's super interesting. Very sexy demos, really a modern kind of canvas that you can basically build next-generation documents that individual items are addressable from a developer. You have to look at us as Microsoft, I'd say the thing to make sure, don't be confused of, is the platform is Azure. But the platform overall is our productivity cloud.

It's that combined with Azure and what you can go do, and so you have to look at all these pieces together and don't feel like you got to learn it all. Go pick up a small little bit, go on Microsoft Learn and hey, deploy a VM. Once you go do that through command-line and go do that on Azure for free and go, oh, wow, you can deploy Linux VMs yeah, it's real. We do that. We do so much more than that. And I think you want to look at us as a company holistically, that is this entire set of clouds, it’s from productivity to our Software as a Service like to what we go build there.

Corey: I would also just like to point out as well that when you say, oh, go ahead and deploy a VM on Azure and do it for free. This is real free, not pretend free, where surprise! Here's a $700 bill you weren’t expecting. It is a legitimate gateway between a free account and a chargeable account. There are no billing surprises here.

Jeff: Microsoft Learn; no credit card required. Get started, and you'll be deploying VMs. We have a sandbox environment that you're able to do. We're not going to email you afterwards and do a sales call and say hey, thank you for signing up for this. Can I get you to buy X Y and Z? No, it’s about learning and you can go do that for free.

Now, if you get further along—and it's much further along—in modules, and certain things like that, you may set up a trial account and so forth, but we want to get you started. We don't think you should have to pay to go and learn our platform. And we want to make that as easy as possible. And that's what my team does every day. How do we go help the community? And how do we help arm them with great technical content, and a service that really makes that easy for them to get started? We're just here to help.

Corey: That I think is the probably best way to wind up wrapping this episode up. It really is a brand new Microsoft. I know I've said that before in previous years with other guests from Microsoft and various aspects, but you've successfully been able to navigate from a company that everyone—including me—hated more or less to one of the most admired companies out there. And the folks that are very anti-Microsoft these days are, in some ways, living in the past for a lot of the reasons that they are. The fact there's now a Linux kernel built into Windows—an actual full-on Linux kernel—means that it finally took Microsoft of all frickin’ people to bring the year of the Linux desktop here, and that year is apparently 2020. So, now people are going to learn a second joke that's going to be challenging for some of them. And I understand wanting to live in the past, but it really is a whole new ballgame, and it's one of the best transformation stories out there. This is going to be a case study in Business School for the next hundred years.

Jeff: You know, we're not your grandparent’s Microsoft. But somebody on my team said the following—they were at Build—“Today I moderated Microsoft dev conference, Build. I was on Twitch. We're an app using [unintelligible] for the front end, and Node.js on Azure Functions was demoed. We connected it to Kubernetes, and running a kubelet written in rust-lang that was compiled Wasm.” This is why I wanted to work here. Welcome to the new Microsoft. That's the company who we are. We're a company that loves developers. And thank you for having me on the show. I really appreciate it. And folks, we’re hungry for your business. We want to help. Come give us a try, and we'll be here to help you. Thank you so much Corey, for having me. And I hope people were able to stay listening and learn something.

Corey: Oh, yeah. And careful what you wish for. I will be trying to do my typical experiment of live-tweeting spinning up a VM on top of Azure. It's been a year or so since I did it. And if it works, well, great, that gets tweeted. If it goes poorly, that gets tweeted, too, and we all learn something from it. So, good luck, we'll see how it goes.

Jeff: Thank you for having us.

Corey: Thank you.

Jeff: And thanks for joining us at Build. It was a wonderful week.

Corey: It really was.

Jeff: We're really proud of the work that we did. But the last thing I'd say is Build’s still going on. As Build finished, I talked about—what eight weeks ago—we decided to build the event. Five weeks ago, we were like, you know what? When this thing's over, people aren't going to want to go home. People are going to want to be able to—well, they are at home but they're going to want to make sure they connect with the community and we tried something new. We've launched a called Learn TV; Microsoft.com/learn/tv, and our advocates, as the credits rolled for Build, we're still online. We went live with, kind of, a fun thing. We're doing live programming on-demand, Q&A, and the show must go on. And so, it's like our own little TV channel. And we're learning there as well, too. So, make sure you join us over there, and maybe one of these days we'll have you on as a guest as well.

Corey: Uh-oh, I think that's one of those things that would cause minor heart attacks through at least a decent portion of the organization.

Jeff: I don’t think so. I think we'd love to have you and we'll have you on someday, for sure.

Corey: All right. Thanks once again for taking the time to speak with me. Jeff Sandquist, corporate vice president of Developer Relations at Microsoft. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts. Whereas if you hated it, please leave a five-star review on Apple Podcasts along with a comment listing no fewer than five minutes Microsoft Power BI implementation ideas.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Colin Percival
Colin is the founder of Tarsnap, a secure online backup service which combines the flexibility and scriptability of the standard UNIX "tar" utility with strong encryption, deduplication, and the reliability of Amazon S3 storage. Having started work on Tarsnap in 2006, Colin is among the first generation of users of Amazon Web Services, and has written dozens of articles about his experiences with AWS on his blog.

Colin has been a member of the FreeBSD project for 15 years and has served in that time as the project Security Officer and a member of the Core team; starting in 2008 he led the efforts to bring FreeBSD to the Amazon EC2 platform, and for the past 7 years he has been maintaining this support, keeping FreeBSD up to date with all of the latest changes and functionality in Amazon EC2.

In his spare time, Colin serves as an alumni representative on the Senate of his alma mater, Simon Fraser University, where he frequently brings a perspective from the world of startups to the ivory tower.

Links Referenced

  • Company site: https://www.tarsnap.com/
  • Twitter: https://twitter.com/cperciva
  • Blog: http://www.daemonology.net/blog/
  • Patreon: https://www.patreon.com/cperciva

Transcript
Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is brought to you by DigitalOcean, the cloud provider that makes it easy for startups to deploy and scale modern web applications with, and this is important to me, no billing surprises. With simple, predictable pricing that’s flat across 12 global data center regions and UX developers around the world love, you can control your cloud infrastructure costs and have more time for your team to focus on growing your business. See what businesses are building on DigitalOcean and get started for free at do.co/screaming. That’s D-O-Dot-C-O-slash-screaming and my thanks to DigitalOcean for their continuing support of this ridiculous podcast.

This episode is sponsored in part by N2WS. You know what you care about? Many things, but never backups. At least until right after you really, really, really needed to care about backups. That's what N2WS does for your AWS account. It allows you to cycle backups through different storage tiers; you can back things up cost effectively, and safely. For a limited time, N2WS is offering you $100 in AWS credits for setting up their free trial, and I encourage you to give it a shot. To learn more visit snark.cloud/n2ws. That's snark.cloud/n2ws.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Colin Percival. Colin is the founder of Tarsnap, which is a secure online backup service, as well as having been a staple in the EC2 history for the one true operating system, FreeBSD. Colin, welcome to the show.

Colin: It’s good to be here.

Corey: So, let's start at the very beginning. What is a FreeBSD, for someone who might never have encountered such a thing in the wild?

Colin: FreeBSD, for people who have very little computing background I often say it's like Linux, but it's not Linux.

Corey: Oh, I bet that irritates some people.

Colin: I'm sure that does irritate some people, and I don't like it when people refer to FreeBSD as being other Linux, which EC2 still does in some places. But for people with somewhat more of a technical background, I say FreeBSD is Unix, and it's about as close as you can get to the natural successor to the original Unix.

Corey: So, once upon a time, back when I was first starting out in my career, I found myself at a university and FreeBSD was what I wound up single-handedly deploying, because of a few different failure modes. One, it turns out that when you have someone who pretty much bluffed their way through the technical interview, and then you give them carte blanche to deploy whatever they want, you get some strange things happening. Not that this was necessarily a bad decision. But, years later when the statute of limitations has run its course, I can now say the reason that I went in that direction was because I had a mentor who was very anti-Linux and very pro-FreeBSD and quite simply, he would help me if I had a FreeBSD question, but he would look down his nose at Linux. Therefore, I was pretty much in a position of, well, beggars can't be choosers. So, I made a full-throated endorsement of FreeBSD, rolled it out, and ran it for a year. Then I moved on to other jobs and haven't touched it in anger or in production ever since. But I still miss it 15 years later, or so.

Colin: That makes sense to me. To be honest, the reason that I started using FreeBSD was it was easier to install than OpenBSD.

Corey: The problem I ran into was that, I guess—how to frame this for someone who hasn't done a whole lot of work with either one? Because I tend to assume that you don't need to have a background as a Linux or Unix administrator to listen to this show and get something out of it. But from my perspective, it felt like FreeBSD was an environment where everything was very clearly ordered. Everything belonged in a certain place. There was a right way to do things. It didn't have a manual. It had a handbook that told you how to go through any aspect of the system. There was a start to it, a middle, an end, and it was great. Going from that to Linux felt like suddenly I'm living in the middle of barely organized chaos. And I kept waiting for that feeling to fade. It hasn’t. I’ll let you know if it ever does.

Colin: I don't know if I would say that in FreeBSD there is always one right way to do things. We have binary packages you can install, or you can build things in the ports tree if you prefer. You can even try to build your own binary packages if you want to build them, and then it's all done, just to make your life more complicated, for instance. I would say that FreeBSD is developed by people who try very hard to make sure that the options that they offer are good ones. And, sometimes that means there is one good option and we tell people this is what you do. Sometimes it means there are several good options and we tell people pick one of these options. But I will agree that in some other platforms, there are no clear good options, or there are many options which are not good ones, and people flounder and end up with things that are not good.

Corey: That's probably a fair way of assessing it. Now, back in the day when I was playing with these things, it was all on-premises hardware, I would say servers, but that's putting it generously. It turns out that when you have—and in this era, this was not an unreasonable operating system choice, but a bunch of desktops that were running Windows XP. And then, after three years, they were deprecated off the books and users wouldn't tolerate how badly they performed, you could then repurpose them, install a different operating system on that, and put it in your giant shelf of badly maintained servers.

So, we had 15 mail servers running like that. One of the earlier projects that I had during my year, there was to rip a lot of that out, but I got to experience an awful lot of janky hardware, interestingly supported on a variety of operating systems. And a few years later, I wound up encountering you, but not knowing it at the time ‘till many years after the fact, because you were effectively one of the driving forces behind getting FreeBSD working on EC2 in the early days.

Colin: I would say in the early days, I was the person that decided this is something that should happen.

Corey: And happened it did. But I guess my question is, is what does it take to look at an existing offering like EC2, where you could ask what operating systems they supported, and back then the answer was, oh, both kinds, Windows, and Linux. And from there, okay, how do you go from having something like a, I guess, a Linux or Linux-y operating system, and then effectively doing a wholesale replacement of the OS, I guess, first, in a way that works, and secondly, ideally in a way that doesn't offend the purest sensibilities of a number of Unix aficionados?

Colin: Well, I want to just correct one thing you said there you said both kinds, Linux and Windows. In fact, when I first decided I wanted to get FreeBSD working, there was one version of Linux that was supported on EC2. And then, later they were both kinds, being CentOS and Ubuntu. It was actually a few years before Windows came along. And by that point, I was already trying and failing, in a wide variety of different ways, to get FreeBSD working.

Corey: Amazing. I guess that's one of those history lessons that I wound up avoiding by virtue of not being at all involved in the cloud back in those early days. But wow, it's not often that I wind up getting exposure to a trivia fact on AWS I didn't already know. Good work.

Colin: Well, I mean, there's a lot of interesting trivia from back then, like the fact that in the early paravirtualized days of EC2, you didn't just have a machine image you also had a kernel image and a RAM disk image. Because you couldn't just say whatever is on this disk, you had to give the paravirtualized Xen the kernel it was going to run, and Linuxes, at the time, needed a RAM disk with… something, I don't know exactly what, on it before it could load everything else off of the filesystem on disk.

Corey: Back in those days, there was the RAM disk, you had to pick what kernel you had to run through. I don't want to say that it was complicated or Byzantine, but there was a company, RightScale, back before they were acquired and no one heard from them ever again, where their entire business value was wrapping the EC2 API's into a front end dashboard that a human being could understand, and then charging a percentage of whatever you ran through it, which sounds ridiculous today, but pretty much everyone I knew, to a large part, back in 2008, 2009 was running through this just because it was so complicated to get up and running. The documentation wasn't there, and the folks who were super involved with it largely were themselves AWS employees, or effectively the next closest thing. How did you dive in and get started with something like that back in those days, where, effectively, it was the digital equivalent of rubbing two sticks together to make fire?

Colin: So, in some ways, it was easier to get started back then, because AWS was really small, and pretty much everybody inside Amazon, or inside AWS at least, knew what everybody else was doing in there. And they didn't have huge numbers of customers asking them for help. So, I sent an email to Jeff Barr. I said, “Hey, I want to get FreeBSD working on EC2.” And he wrote back to me and said, “Here's some people at Amazon you should talk to you.” And for the first few years of trying to get things working, pretty much all of my contacts at Amazon were going through Jeff Barr.

I talked to him on Twitter and so on, but if I needed somebody, Jeff knows everybody. I just sent Jeff an email; he connects me with the right people. And the Amazon engineers were always incredibly enthusiastic. I got the feeling that Amazon as a corporate entity didn't really appreciate FreeBSD. The managers didn't know what FreeBSD was, except that they could tell they didn't have any customers using it. But all the engineers, they loved the idea of getting a different operating system running on their platform.

Corey: It seemed like it was almost a great hobbyist direction back then, in that people were excited to see what the potential use cases of the platform were because this was back in the day before you really had giant companies going all-in on this. Now, for that same level of excitement, people instead have to settle for watching me misuse Route 53 as a database or similar. But back in those days it was, is this even possible was the burning question in everyone’s mind. You proved that it was. And what astonishes me is now, years later, there is still a thriving FreeBSD offering on top of AWS. Is that entirely you? Is there a larger community behind it now? Is it officially supported by folks at Amazon?

Colin: So, the FreeBSD offering on AWS now is officially supported by the FreeBSD project. So, in the early days, it was me building disk images. And, at one point, about a two hour round trip time to test any things I needed to upload a 10-gigabyte disk image before I could boot it, and see where it failed to boot.

These days, the FreeBSD release engineering team is doing all the builds, just as part of our standard release building process. Now, that's the actual building of the images. Making things work, that's a completely different matter. And there's been a lot of work, a little bit by me, but largely by other FreeBSD developers, in the early days working on Xen, dealing with new sorts of Xen devices we needed to handle. More recently, a lot of bug fixes on our NVMe driver because all the Nitro instances expose NVMe disks. Amazon—I was very happy to hear when they launched the Elastic Network Adapter a few years back, they had a Linux driver, and I got worried that they were going to pay people to port their Linux driver to FreeBSD. And, I mean, in FreeBSD, we’re used to, there's a Linux driver over there, go ahead and try to port it. Sometimes there's a Linux driver and here's some documentation for it. It's absolutely wonderful when we have a company actually present us with a FreeBSD driver, and Amazon paid for that to be done, and in fact, has been having people maintain it for us ever since.

Corey: So, other than hobbyists and you, who's using FreeBSD on top of AWS these days? Are there public reference customers? Is this mostly a bunch of hobbyists building interesting things? I mean, you're running an entire business on top of this, which is not nothing. But who's playing around in the space these days?

Colin: To be honest, I don't know exactly who is running FreeBSD on EC2. I can tell you that from the EC2 marketplace, we have around 3 or 4000 instances running that were launched through the Marketplace. I'm sure more far more than that, that were just launched by somebody copying and pasting the AMI ID into the console or onto the command line. There are companies that we know use FreeBSD, like NetApp. I would assume that some of their cloud offerings also run on FreeBSD because why would they not use the same platform for their cloud offerings? But large companies have been very reticent to talk to me about what it is that they're doing with FreeBSD on EC2. It's one of the things I really regret, not hearing from these large customers.

This episode is sponsored in part by ChaosSearch. Now their name isn’t in all caps, so they’re definitely worth talking to. What is ChaosSearch? A scalable log analysis service that lets you add new workloads in minutes, not days or weeks. Click. Boom. Done. ChaosSearch is for you if you’re trying to get a handle on processing multiple terabytes, or more, of log and event data per day, at a disruptive price. One more thing, for those of you that have been down this path of disappointment before, ChaosSearch is a fully managed solution that isn’t playing marketing games when they say “fully managed.” The data lives within your S3 buckets, and that’s really all you have to care about. No managing of servers, but also no data movement. Check them out at chaossearch.io and tell them Corey sent you. Watch for the wince when you say my name. That’s chaossearch.io.

Corey: That's always been one of the challenges I've seen in the FreeBSD universe, to be fair, is that because of the license is such—the BSD license is you can use the source code, you can do whatever you want to and you don't have to re-contribute any changes back, it becomes very uncertain to be able to attribute who is using this in any meaningful way. In fact, the only way that I was able to find FreeBSD jobs when I went looking was by looking specifically for the term FreeBSD in job descriptions. That's how I learned companies like for example, Juniper were big proponents of FreeBSD. But there was remarkably little representation in the common community-style circle.

Colin: That that definitely is an issue. The license being more generous and also, to be honest, the fact that the license is so brief. It does make it harder to identify who's using BSD code. If you buy a TV, and it comes with a copy of the GPL, it gives you some idea of what software is running it. BSD license, you might not even notice because it's half a page rather than 10 pages.

Corey: Yeah, when the entire license fits in a tweet, one starts to wonder.

Colin: Exactly. I don't think BSD license is quite that small, but same idea, yes. So, yeah, it is harder to tell who's running FreeBSD. As far as large companies using FreeBSD now, if you look at FreeBSD developers and where they work, I mean, it's clear—there are companies like Juniper—I don't know if they have FreeBSD developers right now, but they certainly had many in the past. Netflix, of course, has many FreeBSD developers and goes to conferences and talks about the work they're doing on FreeBSD, getting 200 gigabits per second of TLS throughput from their movie streaming devices. So, there certainly are large companies out there using FreeBSD, being open about the fact using FreeBSD, and contributing changes back. But I'm sure there are others out there that are quieter about it.

Corey: Which is very fair. So, let's talk instead, for a minute, about the company you actually run, because it turns out that volunteering your spare time to get FreeBSD working on EC2 is not, in fact, your primary vocation these days, you run a service called Tarsnap. What is Tarsnap?

Colin: Tarsnap—well, the slogan is, “Online backups for the truly paranoid.” It is an online backup service with a tar command-line frontend. And so, you type in a command that—actually it could just be a tar command except with the word Tarsnap in front instead of tar. And say you want to create an archive containing certain files or directories, it bundles those all up, it deduplicates them, it compresses them, and then it encrypts everything before it uploads it to the storage service, which is ultimately backed by Amazon S3.

Corey: Gotcha. So, you say that their backups are for the truly paranoid. Everyone likes to think of themselves as being paranoid with backups, but in my experience, everyone cares an awful lot about backups right after they really needed to care about backups. And even then, they are diligent about making sure that things back up but they never test a restore. So, it leads you to a fun place where backups for the truly paranoid mean different things for different folks. What does it mean for you?

Colin: So, I started this when I was a FreeBSD security officer, and as a FreeBSD security officer, I would get advance notice of security vulnerabilities that affected FreeBSD. So, problems in Sendmail, problems in bind, problems in OpenSSL. And at a certain point, I was looking at all the vulnerabilities I had sitting on my laptop waiting to be fixed because, of course, we always coordinate these disclosures. We pick some date so that everybody can log a patch at the same time. And I was thinking to myself, if somebody got their hands on my laptop, bad things could happen because they could exploit that one, and they could exploit that one, or they could [00:23:28 unintelligible] the vulnerability.

And I thought to myself, “Wait a minute, if somebody's got their hands on my backups, we would be in trouble as well.” Now, like most people, I didn't really do very good backups at the time. But I then was thinking, “Well, if I start doing backups more often, than that means there's more copies of all this scary information sitting around somewhere. How do I do this securely?” So, I looked around at what I could find online in 2006, and they're really wasn't anything out there that I could trust to do backups securely. I was a FreeBSD security officer, I had quite a background in security and cryptography at that point. And just based on my expertise in the fields, I didn't trust what was out there.

So, I asked around, some of my friends and posted on my blog, and I asked, “If I build this myself, would anybody else want to use it?” Lots of people said, “Yes, this is something we would pay for.” So, it happened I was looking for work at the time. I had a job offer from Google to go down to San Francisco and do research. But I wasn't [00:24:36 unintelligible] offer for a few reasons. So, I decided well, okay, I'll build it myself and see how it goes. So, it turns out it is very much a startup in the open-source tradition of scratching your own itch. I had a problem. Some of the people said that they have the same problem, so I decided to fix it.

Corey: And fix it you did. It's been around for a while. You have some very impressive name-brand customers who are publicly referenced on your site. Stripe is a, I guess the canonical example. If you take a look at how much they care about the sanctity of their backups, I don't feel like there's really enough words to express the answer to that question. If you get effectively the internet's payment systems data, there is disaster, and hellfire, and brimstone, and nothing looks the same tomorrow if that happens. So, it's obviously validated and tested by folks who take their workload seriously. I guess the question is, is why did you go down the path of A) using FreeBSD for this, and B) building it on top of EC2 instead of a bunch of different options that you could have potentially gone with?

Colin: So, in 2006, when I decided I wanted to do this, I knew—I mean, I'm a software guy. I knew that I did not want to be dealing with physical hard drives and I definitely didn't want to be driving down to the datacenter to swap out failed hard drives. So, I wanted something out there that could store the data for me and not lose it. S3 launched earlier that year, so I said to myself, “Okay, S3 sounds like the backend I want to use for this. And then, well, I need to have some code running in front of that. Oh, look, here's this Elastic Compute Cloud service that lets you have servers that are really close to S3 and can push bits in and out of S3, without paying any bandwidth costs.” So, it was just a natural connection there, EC2 was what I needed to be able to use S3 efficiently.

Corey: And one thing sort of leads to another. And, I think, as anyone tends to learn sooner or later, they, kind of, wind up staying wherever they wind up initially building something out barring a tremendous strategic reason to change providers. So, one thing that I found interesting that I saw a while back and [00:27:05 I'll link to it in the show notes] was Patrick McKenzie wound up doing an entire analysis of Tarsnap and writing—an essay doesn't really encapsulate the entirety of what he wound up writing—it was more or less a day-long tear down of your entire market positioning and effectively giving a laundry list of things he would change if he were doing the marketing piece for Tarsnap. What led to, first, him doing that? And secondly, what was your response when you wound up going through all of the copious detail that he wound up putting in there?

Colin: I can't remember the exact history leading up to that, although, I mean, we had exchanged comments about Tarsnap on Hacker News for a couple years leading up to that. He did ask me, by the way, was I okay with him doing this, and I was very enthusiastic and I still am very enthusiastic. That blog post actually brought Tarsnap more customers than anything anybody has ever written by, probably, a factor of 10.

Corey: Just wait until this podcast goes out and we'll see if we can beat it.

Colin: Well, that would be fantastic. [laughs] But as far as my opinions on what he wrote about Tarsnap, I think his view of Tarsnap is somewhat different from mine. He, and also Thomas Ptacek, who shows very similar opinions to Patrick about Tarsnap, have said that the worst thing for a small business to be is a utility. This idea of pricing Tarsnap the same way that you pay your power bill is just terrible. As far as they’re concerned. My view is exactly the opposite. I think backups should be a utility. And if people can pay their Tarsnap bill the same way that they pay their AWS bill, that is not bad in my opinion.

Corey: I would say that there's definitely an argument that could be made in either direction. The joy of looking at things from a utility perspective is that, okay, great, you wind up paying for things that turn on, turn off. And we've seen companies move away from this. I mean, remember back when Dropbox instead of being a bloated monstrosity that failed to work in most respects and beat your CPU to death whenever something touched a disk when it used to just be a folder that would sync between various computers and have the same contents in it at all times, like magic, that felt like a utility. Now, of course, it's a platform and it certainly worked for them. They've gone public and done super well. But they clearly have departed from their routes of being, do one thing, do it well in a utility fashion, so maybe that means that Tarsnap is not fated to become a publicly-traded company worth billions of dollars.

Colin: That is quite possible and honestly, I don't really mind if Tarsnap fails to become a publicly-traded company. Being a publicly-traded company is an awful lot of work, and I don't think I really want to do that.

Corey: No, no, there's certainly a list of things I want to deal with versus don't want to deal with, and paperwork is very clearly in the second category. That's why the company is never just me. It's always good to have people who are better at things that I suck at. So, something else you've written recently that I wanted to talk about was imds-filterd. That's Indigo, Mike, Delta, Sierra, dash filterd. And one thing that I love about imds is that it the first sentence in the readme tells you how to pronounce it which is a rarity around anything that touches AWS. Usually, it leads to warfare, character assassination, actual assassination, and I still stand by my AMI pronunciation, but what is imds-filterd?

Colin: So, imds-filterd is a filtering daemon for the instance metadata service. Amazon refers to the instance metadata service as IMDS. I'm not quite sure why metadata gets two letters instead of one, but maybe they think meta and data are different words, I'm not sure. In any case, they call it IMDS, so I call it IMDS. And imds-filterd is a daemon which restricts access to whichever parts of the instance metadata service you would like to restrict access to. And it does this based on rules that you provide with user IDs, and also group IDs if you want. So, this means that you could say, this web proxy should not be accessing IAM credentials. We do not want people to use this web proxy to get the credentials to access S3 and steal all of the information on 100 million credit cardholders. Or you could tell it, user nobody should not be accessing things. So, that privilege separated SSHD that you've got running, the preauth process, if there's a vulnerability in there, somebody should not be able to exploit a preauth vulnerability in order to steal those same IAM credentials.

In general, I would say you probably want to let root access everything, because well, we’ll just root it and turn off the filtering daemon if it wants to anyway, but it's essentially a way of fixing the fact that, in the early days of the of IAM, they decided the right way to expose credentials was via the instance metadata service, which is accessible via HTTP from any process on the system. Honestly, I think that was the worst security mistake Amazon has ever made in AWS, but they haven't fixed it so I figured, well, I need to step in and I need to fix it.

Corey: And step in and fix it you did. They wound up releasing the v2 endpoint, but that's going to take, as I believe you've mentioned, ages for that is globally supported to the point where the v1 endpoint can be turned off. That's going to be a painful thing for a lot of shops.

Colin: Getting to the point that people can turn off Version one access is going to be painful because you need to have code that supports v2 before you can block Version one. Also, v2 doesn't solve the problem completely. It solves the problem of, I have misconfigured proxy. But it doesn't solve the problem of somebody managed to break into my server but within a sandbox. They're running as user nobody, or they found a bug in Apache, so they're able to [00:33:47 uncode] as the www user. Those users should not have access to IAM credentials unless there's some credential you need to have, but in general, they shouldn't have access to those credentials. And even with Version two of the instance metadata service, right now they do have access because they can make the necessary requests.

Corey: Excellent. And I will throw a [00:34:11 link to that in the show notes] as well. Last question before I let you go. You wind up doing an awful lot of work for the larger community, in order to make FreeBSD on EC2 run. If people want to support you, how can they do that?

Colin: So, a couple years ago, I set up a Patreon. The original idea, the way I set it up was just, this will be a way that people can cover things like my travel expenses because there have been times I've considered going to conferences and said, “You know, it might be useful for me to go somewhere like Amazon re:Invent, but I don't really want to pay for that out of my own pocket.”

As it turns out, now I'm a Amazon Community Hero. So, Amazon pays for me to go to re:Invent. But there have been other events I've considered going to and decided not to because I didn't want to pay for it myself. And at this point also it would be nice if the community could pay for some of the time I spent working on this because I do have a day job and the more time I spend working on getting FreeBSD working on EC2 and fixing issues as they arise, the less time I get to spend on Tarsnap.

Corey: That's absolutely something that is worth supporting. I think that we take people doing what amounts to volunteer work in the open-source space far too much for granted. So, absolutely thrilled to [00:35:30 want to throw in a link into that].

Colin: Great.

Corey: Colin, thank you so much for taking the time to speak with me. If people want to hear more about what you have to say, where can they find you?

Colin: They can follow me on Twitter, @cperciva, or follow my blog daemonology.net/blog.

Corey: Excellent, and we will absolutely [00:35:50 toss links to that in the notes] as well. Thanks once again for taking the time to speak with me, I appreciate it.

Colin: Great to talk to you.

Corey: Colin Percival, founder of Tarsnap and, effectively, one-man force of nature in the FreeBSD ecosystem on AWS. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review and Apple Podcasts. If you've hated this podcast, please leave a five-star review in Apple Podcasts, and a comment explaining why FreeBSD is your favorite distribution of Linux.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Prashanth Chandrasekar

Prashanth Chandrasekar is Chief Executive Officer of Stack Overflow and is responsible for driving Stack Overflow’s overall strategic direction and results. Prashanth is a proven technology executive with extensive experience leading and scaling high growth global organizations. Previously, he served as Senior Vice President & General Manager of Rackspace’s Cloud & Infrastructure Services portfolio of businesses, including the Managed Public Clouds, Private Clouds, Colocation and Managed Security businesses. Before that, Prashanth held a range of senior leadership roles at Rackspace including Senior Vice President & General Manager of Rackspace’s high growth, global business focused on the world's leading Public Clouds including Amazon Web Services (AWS), Microsoft Azure, Google Cloud Platform (GCP) and Alibaba Cloud, which became the fastest growing business in Rackspace’s history. Prior to joining Rackspace, Prashanth was a Vice President at Barclays Investment Bank, focused on providing Strategic and Mergers & Acquisitions (M&A) advice for clients in the Technology, Media and Telecom (TMT) industries. Prashanth was also a Manager at Capgemini Consulting where he managed Operations transformation engagements and consulting teams across the US. He holds an MBA from Harvard Business School, an M.Eng in Engineering Management from Cornell University and a B.S. in Computer Engineering (summa cum laude) from the University of Maine. Prashanth is married and has two children.

Links Referenced

  • Twitter
  • Rackspace
  • Stack Overflow
  • Stack Exchange

Transcript
Corey: This episode is brought to you by DigitalOcean, the cloud provider that makes it easy for startups to deploy and scale modern web applications with, and this is important to me, no billing surprises. With simple, predictable pricing that’s flat across 12 global data center regions and UX developers around the world love, you can control your cloud infrastructure costs and have more time for your team to focus on growing your business. See what businesses are building on DigitalOcean and get started for free at do.co/screaming. That’s D-O-Dot-C-O-slash-screaming and my thanks to DigitalOcean for their continuing support of this ridiculous podcast.

Corey: This episode is brought to you by Spot.io, the continuous cloud cost optimization platform, saving businesses millions of dollars each year on their cloud bills used by some of the world's largest enterprises and fastest growing startups like Intel and Samsung. Those are enterprises and duo lingo. That's a startup. Spot.io delivers the optimal balance of cost and performance by leveraging spot instances, reserve capacity, and on demand. Give your workloads the infrastructure they deserve. Always available, always scalable, and always at the lowest possible cost. Visit Spot.io to learn more.

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by relatively recent Stack Overflow CEO, Prashanth Chandrasekar. Prashanth, welcome to the show.

Prashanth: Thank you, Corey, a pleasure to be here. Thank you for having me.

Corey: Your history is fascinating. Before you joined Stack Overflow as the first non-founder CEO, to my understanding, you came out of a relatively lengthy tenure back at Rackspace.

Prashanth: That's right, yeah. So, Rackspace, just a tremendous company and experience, just, I'm so grateful for that as part of my career and, just really having, kind of, worked with a tremendous number of amazing people at the company down in Texas, and around the world. Yeah, so in my journey at Rackspace was all about how do you redefine the company in the context of a fairly competitive landscape in the cloud? If you remember, Rackspace, originally it started in the context of a managed hosting company well before I joined there.

And when I joined in 2012, we had our own public cloud, the OpenStack Public Cloud, and going head to head against Amazon Web Services, which was our primary competitor, and as we know today, obviously AWS is the leading public cloud capability and still growing. Just a massive, massive success by Andy Jassy and team. And so, back in 2014 or so, our company decided that we wanted to effectively not compete based on infrastructure and just based on a price war, which is obviously going to go to zero, which didn't make any sense. And that—

Corey: Yeah, who could lose money the slowest and be the last person standing.

Prashanth: Yeah, indeed. And then, Amazon being the company that it is relative to Rackspace, it’s relative size, it didn't make a lot of sense for us to compete based on that dimension. So, what we had to do is some soul searching to determine what we were truly good at. And we decided that we were really good at what we call fanatical experience—or fanatical support, what we used to call it back then—and that was all about providing services. And so, that's what we ended up doing, saying why were we so specific about providing services only on our cloud or our infrastructure? Why wouldn't we do it on any kind of infrastructure, much like we would support Linux or Windows, in the operating system context?

So, that's when we made a strategic decision in the company, one of the most important decisions in the history of the company to support Amazon Web Services, and then soon after, Microsoft Azure, and Google Cloud. And so, that was just a, I would say, a huge pivot for the company. And I was part of the founding team that did that, and then also helped lead that business to a multi-hundred-million-dollar business over the course of just a few years. And that ultimately was called the Managed Public Clouds business at Rackspace.

And that was, I would say, was a tremendous experience in building something very rapidly, scaling it by rapidly, enabling a lot of people both on the product engineering side plus also the go-to-market side. So, just an all-around, a great experience that I'll never forget.

Corey: Oh, absolutely. It's fascinating talking to you, just because for those who aren't privy to all of the ins and outs of my various career twists and turnings, you and I went to undergrad at the University of Maine in Orono, at the same time. And that is the last time we ever really had anything in common. You wound up graduating summa cum laude with a BS in Computer Engineering. You got a master's in Engineering and Engineering Management from Cornell, and then an MBA from Harvard Business School.

Whereas I failed out of undergrad, discovered my high school diploma was not accredited, so on paper, I've been walking around with an eighth-grade education that no one can ever take away from me. The path not taken as it were. But you've had a fascinating career trajectory and now have landed as the first non-founder CEO of, whether we think of them as a cloud company or not, I would argue that the cloud would not exist without Stack Overflow in its current form.

Prashanth: Yeah, no, by the way, on the background, yes it's just a fantastic coincidence that we both went to the same undergrad schools. And obviously, listen, everybody has their own path to their own, kind of, finding their own true purpose, and what they end up doing in life. So, nothing's right or wrong. It just happens to be what the paths that we chose. So, I would say, I think that yeah, Stack Overflow throughout my time at Rackspace was just such a loved entity and a loved community. All my technical teams always historically used Stack Overflow. We've always known about it. I’ve known about the company for a long time. And then, more towards the end of my time at Rackspace, I was approached about the role, saying that, “Hey, this great foundation exists.”

So, this community still has something like 50 million users that come to the website every month, just on Stack Overflow. If you include Stack Exchanges, another 70 million folks that show up at the Stack Exchange websites. That's about 120 million people that show up every month. We have something like 150,000 new signups on Stack Overflow every month. So, this is a tremendous amount of scale at, 12 to 10 years after founding the company. And that foundation is just, I would say, there aren't a lot of companies that actually have that, sort of, a phenomenal foundation.

And on top of that you had these great products that the company was beginning to build and had built over the past few years. Talent product, which is our job listings product where big companies can post their jobs on our website. And obviously, we have the world's developers on our website, about 25 million, all of them are—they pretty much use Stack Overflow on a daily basis. So, they're able to match jobs and applicants. And then you've got the ads business that basically is an ability for big companies to showcase their developer-centered products. So, think about, for example, in the case of Google Cloud it could be BigQuery, and how ads on that so that developers can leverage the product.

And then, finally, the latest SAAS product that we have is called Stack Overflow for Teams product, which is just a tremendous way to share knowledge and collaborate within companies across their development teams, their product teams, their security teams, IT teams, etcetera, to make sure that they have the latest and greatest content internally around their feature releases, and code snippets, etcetera, and all be available for Teams like go-to-market teams and other teams that need to be very close to the core information. And that product has really been a tremendous growth engine for Stack Overflow. So, we've seen that business double year over year. And we've got companies like Microsoft that have something like 70,000 users on Stack Overflow for Teams.

So, that whole story and why Stack Overflow it’s kind of an inflection point in its history. It's helped build the cloud, to your point, by enabling developers around the world to really rapidly build out capabilities across AWS, Azure, and Google, but also now helping them become even more efficient as part of their development workflow, and adding value like in helping them find jobs and awareness about various products. So hopefully, that's a helpful overview for you.

Corey: Oh, it absolutely is. It's always interesting when you see a company that is effectively a household brand has a management or leadership shakeup in different ways. For example, when Google acquires something and the immediate knee jerk reaction is oh, no, they're going to kill it. When Stack Overflow effectively changed leadership, it was—the biggest reaction I had, as a full Stack Overflow developer myself is, “Oh my God, there might turn off copy and pasting, and then where will I be?” So, just for the record, there is no current plan to disable copying and pasting on Stack Overflow, Stack Exchange, or any of the affiliated halo sites.

Prashanth: Oh, my goodness, no. I think that we want to do more things for our community, not less things, and we want to make life easier for our community members versus the opposite, so absolutely not. And so, just if you think about a lot of what we're planning on doing this year, one of our top priorities is community engagement and inclusion. So, we really are trying to make sure that we get more and more developers, and hobbyists, etcetera, and people are writing code earlier and earlier in their lives.

We were trying to make sure that they feel extremely welcome to leverage the community, to utilize the resources there. How do we have even more productive ways to intersect our public community Q&A platform and our private Stack Overflow for Teams product, so that there's even more developer workflow integration, so people don't context-switch, etcetera. All those things are in the spirit of making sure that life is easier for community members like you. Because today, most people like you go to Google, type in a question about, whatever, could be a Python question, or Amazon Web Services question, and you land on Stack Overflow—

Corey: Or copy and paste an error message directly in and cut out the middle person. For whatever reason, that seems to be a less common approach than it really should be, given how many problems it seems to solve for.

Prashanth: Exactly, yeah, so that happens more often than not. And we just believe that it should be a lot more seamless of an experience for users like you. And that's why we're integrating in the community we integrated with GitHub last year. For our Stack Overflow for Team product, we have integrated with Slack and Microsoft Teams, and we're about to announce our integration with JIRA and GitHub Enterprise. So, there's just a lot that we're hoping to do to make sure that the developer workflow is highly integrated, and ultimately, we’re indispensable as part of that.

Corey: I would say that, for better or worse, whether or not companies know that they're dependent upon Stack Overflow, they absolutely are, in that it is saved so many person-years of time, in the past decade or so—however long it has been in business, it feels like forever, but I understand that it's not the awareness of the passing of time is a slippery thing—but it has been a transformative place to solve things. I mean, I will say that I've dabbled a little bit as a part of the community and found it wasn't really for me, and for the best of all reasons, namely that giving an answer should not be a facile one- or two-line response. There's expected to be some depth, and some exposition, and explaining not just what the right answer is, but why it's the right answer. And when I'm trying to do this late at night, from a phone in a hurry, that doesn't really lend itself to the same thing. I mean, I rarely have the attention span to finish writing a complete tweet, let alone a full deep technical dive into something. So, that high community standard is absolutely one of the differentiating virtues of the entire community.

Prashanth: Yeah, no, spot-on on that. There is a very—you've got to give up to founders just a tremendous amount of credit. Both Joel and Jeff Atwood have built this system that is sustained for so long, and despite what could be perceived maybe harsh realities of downvotes or, kind of like, a very binary experience or a very specific experience, it works. It just works. And that's the whole point of it is to make sure that people get the right answer very quickly, and they're on their way. And that does come with some downsides with regards to, sort of, the friendliness element and with, if you're a newcomer, you might feel—may not be, kind of like, the easiest way for you to get going.

I remember when I was dabbling myself in the community, even before I joined the company, I would say, it was, sort of, a harsh reality of actually participating as a brand new member. It's not the easiest way to get started. So, we are testing and iterating on several ways to make sure that we are a lot more welcoming, and making sure that people don't feel intimidated. One of the more recent things we've done is actually make sure that people asking great questions also get as many points relative to answering questions. And then, we also are making sure that we remove negative comments and we've actually made some tremendous progress over the past quarter on that. Almost half the negative comments relative to the prior period of measurement. So, we are doing many things to make sure that, despite the very objective binary kind of system that's in place, we are making sure that we are making it more welcoming for new users.

Corey: You've also made significant changes in the past year around how you're handling diversity conversations, representation from different groups in the community. And what’s, I guess, fascinating to me is how public you've been about some of this. It hasn't been something that you give lip service to it once a year, and then it vanishes. We're starting to actively see changes in the community happening and percolating out. You haven't been trumpeting it from the rooftops, but you also haven't been secret about it or working non-transparently to get there. It's fascinating and I'm curious as to how that's, I guess, how that came to be, and how we can start seeing that in more companies.

Prashanth: Yeah, no, I think, we just believe that Stack Overflow, given, kind of like, the level of influence that we have around the world, we are very much a platform, a community platform that represents our users, the user base across the world. So, the world's developers leverage our platform, but we have a long way to go in terms of even making sure that everybody signs up for an account and is active, etcetera. So, there's going to be work to do, and part of when we do our developer survey, for example, a yearly developer survey, it's just very, very interesting for us to see some of the results there. Some of them stood out to us, whether that’s, hey, most developers are—80 percent of the developers are coding as a hobby.

And there's a diversity element there. Like how do we make sure that even newbies or people that are new to programming feel included as part of this? Or the fact that you have people that are less than 17 years old, 18 years old, we're talking about at least 50 percent of the people that we polled were, globally, less than 18 years old. So, we’re talking about an age demographic inclusion metric that's very important to say how do we welcome youngsters into the flow of making sure that they are participating?

Or even more specifically around gender diversity. And if you look at that there's, it's heavily lopsided in terms of males and females. And we say, I think approximately in only 11 percent of our US survey respondents were women this year, it was a slight improvement from last year, about 9 percent last year, but still very, very small. And if you think about that, it's not representative of the world's developers. I grew up in India, and just within that country alone, the number of women developers that are emerging into the workforce, this is a fascinating number. And obviously, I moved here for college and joined you in Maine when I was 17, but the point is, it’s just an important statement and, kind of, a piece of work that we need to make sure that we are representing what's truly in the world.

And so, we are very committed to making sure that diversity and inclusion is a top priority for us. This starts, by the way at our own company. Like the people that we hire, my own leadership team. So, we’re making a very conscious effort to make sure that we are very much reflecting our community in a way that is fully representative of the true user-bases out there, to make sure we lead by example. And you will see us do many things this year to make sure that we made progress on this key initiative that will keep you posted on.

This episode is sponsored in part by N2WS. You know what you care about? Many things, but never backups. At least until right after you really, really, really needed to care about backups. That's what N2WS does for your AWS account. It allows you to cycle backups through different storage tiers; you can back things up cost effectively, and safely. For a limited time, N2WS is offering you $100 in AWS credits for setting up their free trial, and I encourage you to give it a shot. To learn more visit snark.cloud/n2ws. That's snark.cloud/n2ws.

Corey: How do you wind up effectively balancing between two very important but almost inherently oppositional goals. One, to wind up continuing to drive change within the community and make the product offering and discourse, frankly, better, versus the other side of avoiding the Digg or what seems to increasingly be the Reddit trap where a redesign drives the community members away, or—anytime you start changing a community, there's a large group of people out there, and often I'm one of them who despises change, or anything that smells of change because we delude ourselves into thinking that the world will hold still long enough.

Prashanth: Yeah, phenomenal question. I think that a couple things to note here. I think one is, we are very much convinced that it is important for us to make sure that we seek input from a very broad set of folks, just like our developer survey that I mentioned. We've got new mechanisms in place, like The Loop survey, which we've launched, which really expands the number of voices that we have in terms of the voice of the community, so to speak, so we really understand what people want, to make their experience even better. And so, that change or set of changes—because historically we have completely relied on just mechanisms and forums like Meta, which is home to our tremendous power user base and the folks that—such highly valued members of our community.

But it does represent a very specific portion of our community and doesn't necessarily include the entire population that we want to really hear from. So, that is important. We want to make sure that we are making sure that we make changes to the question asking process, Questions Asking Wizard, Unfriendly Robot, and others. And this is really led to some really strong results. So, more people are asking questions without seeing a dip in the question quality. And we cut the number of negative comments nearly in half without seeing any sort of reduction in the overall comments. And ultimately December was our best month ever in the history of our company in terms of new signups. So, there are things that we are doing that are resulting in, I would say, positive traction.

Now, to your point, not everything is easy to be managed and terms of change. So, we have to articulate—we need to do a much better job, by the way, as a company—of articulating how and why we're including these new mechanisms, what is the role of existing mechanisms like Meta? Because they're obviously very valuable for us, as we continue to evolve the community. How do we make sure that our power users are part of the journey of helping us get to the overall goal of making this community even more impactful? Because clearly they care about that. So, there's a lot we’re going to manage around that. But I think that we are very committed to making sure that we have a robust listening process, to make sure we listen to our community, and also to have bi-directional feedback. And to make sure that folks like me and others are going to be very active, to make sure that the community is always in the room and we have a key conversation. We're going to make that happen.

Corey: One thing that—I don't know if it was a problem or a change that I was somehow tripping over based upon asking stupid questions, but it seemed for a while that every third time that I googled for a particular error message, the top result was always something on Stack Overflow. And that was always a conversation that was closed as off-topic or marked as a duplicate, yet somehow it was always at the top of Google search results to the point where it became a recurring joke. I don't see that happen nearly as much anymore, to the point where when I do it feels like it's almost an oddity in its own right. Is that something that was a change on your end? Was that just I learned to Google for better questions? Or was there a flare-up in that that suddenly just wound up getting fixed, but everyone thought that was how it was supposed to be?

Prashanth: Yeah, this is interesting. This is about basically improving our code of conduct and educating our moderators and power users. So, it's making sure that we are very specific about how we improve on a day to day basis. We make sure that people are evolving how they actually operate in the community and how we bring along the community for that ride. So, that's effectively what you're seeing there. And to make sure that our moderators are enabled with the right sort of information, and they are really trying harder to make sure we help new users and beginners, and not to be too zealous about our gatekeeping so that people are more welcoming here internally.

Corey: It definitely seems to have had a meaningful impact. Now, I have to assume that there's more to your company, as far as revenue models, as far as what you folks do, than providing the answer to me frantically googling while on a phone screen, “How do i FizzBuzz?” What are you folks beyond the community, which is the way most of us tend to interact with you?

Prashanth: Yeah, thanks, Corey, for that question. I think that Stack Overflow is mostly known for the community that we have established over the past decade. So, the 50 million folks that show up to our website every month. There's the additional 70 million that to our Stack Exchange websites, and then 150,000 people that sign up every month for new accounts. But what's not well understood and perhaps not as recognized is the fact that we are a true company, a SAAS company. And so, very much like how GitHub had its own enterprise journey, we've been on that journey for a few years, it's a couple years. And that's one of the reasons why I'm also on board, which is to make sure that we really build a sustainable and long term and successful business at Stack Overflow in addition to, obviously, the core mission of making sure that we really help our developers and our technical workers. So, in terms of the products, we have our Talent business, which is really—we have something like 40,000 jobs were posted on Stack Overflow for jobs in 2019. Just matching against close to a million searchable profiles of developers who are interested in being contacted by job on Stack Overflow Talents. So, that's our first business, the talent business, it's a big area.

And we have a second business which is called the advertising business, which is something like a million developers found new and useful tools, after seeing a company advertise on one of our sites this past year. So, think about Microsoft advertising about Microsoft Azure, or Google advertising about BigQuery, those sort of examples. And then, our third product is our true SAAS product called Stack Overflow for Teams, which really allows developers and product folks, etcetera, to really collaborate and ship products really fast within their organizations. And so, we've had hundreds and thousands of engineers who have leveraged our product in 2019. As an example, Microsoft has something like 70,000 developers on Stack Overflow for Teams.

So, those products are just the beginning of, I would say, our commercial journey, and some of those products, especially the Stack Overflow for Teams, is growing at a very rapid rate. And we have all the fortune 100 companies, really some phenomenal logos like Bloomberg, and Microsoft and a whole bunch of other fast-growing startups like Expensify, etcetera, that are all leveraging Stack Overflow for Teams, and are just amazing use cases internally in these companies. So, we're really tagged against the future of work, if you will. So, we integrate with Slack and Microsoft Teams, and very soon with GitHub Enterprise and JIRA, to make sure that we're really part of the developer workflow so that we help developers are faster and faster as they ship products.

Corey: As you take a look across the larger ecosystem and community, what do you think is the most misunderstood about the company in the common case?

Prashanth: Yeah, it's a great question. I would say, it's probably the notion that Stack Overflow is a community, and it's really not—people don't truly understand that—so the other products that we have that are also tremendously valuable. I think most people just know us about the community and assume that we don't have a revenue model, when in reality, we have a very robust revenue model that's, in many ways, just about to really explode in a good way. So, that's just, I would say, the biggest, I would say, lack of understanding. And I think the other part is, I think people just don't realize the scale at which we operate. I think people just don't understand that we are pretty much, I would say, the life-force of the internet. I mean, everybody—[laughs] there is not a single developer in the world that does not use Stack Overflow. I mean, that is not a kind of, a brash claim. I think that it's hard to come across a developer that has not used our website, one way or the other. It is the new way, or if you will—

Corey: Well, many of us may not admit it, but absolutely we use it.

Prashanth: Exactly. And so, I think that is, I would say, it's the unsaid, kind of, statements and most people—I don't know if everybody realizes that, realistically.

Corey: Yeah, well, what would you do if you had this problem in an interview? Well, I'd google it. Well, what if Google is down? That's okay. I already know Stack Overflow’s address. We're good. It becomes one of those fun exploratory things. I guess, what are you seeing or concerned about as a CEO who is not a founder of the company, effectively their first external CEO, what's top of mind as you step into what are admittedly some very big shoes.

Prashanth: Yeah, so I've had—a lot of what I've done so far as to really listen to all our employees, our community members, and our customers, of our products, and constantly, every week, I’m talking to a large number of them on a regular basis, just to really make sure that I have the full context to really make the right decisions in terms of the types of products that we build, what we need to evolve as a company, etcetera. So, a lot of work where I think what we're trying to accomplish, I think, even in my first 90 days of the company, one of my observations is that I think a lot of, while the external community and the tech industry has very rapidly evolved if you think about the cloud, and you know this better than anybody, which is between Amazon Web Services, and Microsoft Azure, and Google Cloud, and concepts like Kubernetes, container orchestration, Lambda with Serverless.

There's just a whole world where infrastructure and DevOps and software engineering, software frameworks, all of those things are, sort of, very much colliding, and really the lines are blurring. And so, even for, sort of, our experience on the community or our community websites, I don't think have evolved as much as, the end customer, the end community member has evolved in terms of what they do on a daily basis. So, there's a lot of what we have to improve in terms of the experience between our Stack Exchange websites and our Stack Overflow website, and to make sure that people have all the things at their disposal, whether that's DevOps concepts or cloud concepts or Stack Overflow concepts, all questions and answers at their fingertips.

As an example, we have 182,000 questions on Azure. And we have 87,000 questions on AWS and 20,000 questions on Google. But it's actually not well understood or well recognized because it's actually sitting in Stack Exchange, sort of in a corner with a bunch of other websites that we had. Their phenomenal properties, but it's away from Stack Overflow. And most people do go to Stack Overflow. And we also have, by the way, communities like Information Security, and DevOps, and Data Science, and so on. So, there's a lot that we can do to make sure that we bring those to the forefront, so people actually are able to leverage all these great resources. So, that's really, I would say, a lot of what we're thinking about. Make sure—one of the things that I've noticed that we have not necessarily evolved, just like the external developer [unintelligible] outside our company, has evolved very rapidly, and we need to do a much better job of doing the same on our end.

Corey: Fantastic. Last question before you go, and I just wanted to get your take on this. A while back, Joel had a blog post—I think it was a blog post, it could have been a bunch of other different media, I follow basically everything he says sooner or later—and what struck me was the way that you folks focus on employee experience, it's sounded, hands down, like one of the best companies in the world to work at. And this is going to sound petty or like I'm being sarcastic, but I swear I am not. The fact that every engineer there gets the option of having their own private office is mind-boggling to me. I have worked in too many startups where you're sitting cheek to jowl with everyone next to you that holds still long enough. And just the idea of a private office is hands-down one of the best opportunities to differentiate yourselves. I don't know why other people don't do it, but it is one of the most aspirational aspects of Stack Overflow culture that I think I’ve ever heard of.

Prashanth: Yeah, no, absolutely. I obviously can't take credit for that. That's all Joel and his philosophy and his—which I totally agree with, by the way, which is really making sure that we provide an environment for our engineers and our developers to make sure they're highly productive, in the same vein of what I've discussed about various products, etcetera. But it's the physical environment, how do you make sure that they actually have the space for them to be, sort of, in the zone, so to speak, to be able to really contribute at their highest level? And so, I think it's very distinctive about our culture. One of many things, by the way, that makes Stack Overflow such a great place to work for. And so, we welcome—we're always hiring, by the way, so if anybody's listening here, we are always looking for new and talented Stackers as we call them, so please reach out to us.

Corey: Thank you so much for taking the time to speak with me today. I really do appreciate it.

Prashanth: Same here, Corey a real pleasure, and great to reconnect with a fellow Black Bear.

Corey: Yes, eventually, those scars will one day fade as soon as the Maine chill gets burned from my bones. Prashanth Chandrasekar, CEO of Stack Overflow. I'm Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts. If you've hated this podcast, please leave a five-star review on Apple Podcasts and tell me exactly what my problem is.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Bryan Liles

Bryan Liles is a Senior Staff Engineer at VMware. He leads the Developer Experience group, which creates solutions to help developers be more productive in Kubernetes. When not working, Bryan builds and races cars and drones.

Over the past 20 years, Bryan has worked around cloud technology and distributed systems. His approaches to technology are: simplify with fidelity and technology should give access to all.

Links Referenced:

  • https://vmware.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is brought to you by DigitalOcean, the cloud provider that makes it easy for startups to deploy and scale modern web applications with, and this is important to me, no billing surprises. With simple, predictable pricing that’s flat across 12 global data center regions and UX developers around the world love, you can control your cloud infrastructure costs and have more time for your team to focus on growing your business. See what businesses are building on DigitalOcean and get started for free at do.co/screaming. That’s D-O-Dot-C-O-slash-screaming and my thanks to DigitalOcean for their continuing support of this ridiculous podcast.

Corey: This episode is sponsored in part by N2WS. You know what you care about? Many things, but never backups at least until right after you really, really, really needed to care about backups. That's what N2WS does for your AWS account. It allows you to cycle backups through different storage tiers so you can back things up cost-effectively and safely. For a limited time, N2WS is offering you a hundred dollars in AWS credits for setting up their free trial, and I encourage you to give it a shot. To learn more, visit snark.cloud/n2ws. That's snark.cloud/n2ws.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Bryan Liles, Senior Staff Engineer at VMware. Bryan, welcome to the show.

Bryan: Thanks for having me.

Corey: We've been trying to set this up for about a year or so, now, since we first crossed paths, I want to say, at KubeCon in Barcelona in 2019.

Bryan: Yeah, so I have this thing where people ask me to do podcasts, and they send me DMs, and DMs is like purgatory for me. I get just enough, where if I don't respond to you in about two hours, I'll never see it again. So, what happened is I promptly forgot, and Corey just pinged me, and here I am.

Corey: I do my best. The hard part is I find myself in the same exact, I guess, position when it comes to dealing with losing track of DMs. Twitter's automatic sorting of randos into the spam folder is always interesting, but I also wind up with people—I follow them, they follow me and they're still, for some reason, sitting there as if I've never spoken to them before. It's great.

Bryan: Yeah, yeah. So, that's about what happens. And really, it's about how I use Twitter. I mean, I hate to be selfish, but I don't like the DM feature. If I really want to talk to you, I'll give you my email.

Corey: The problem is some folks take the exact opposite approach where Twitter is the public face, email is good lord, you could spam me with that. And I think we've all wound up on enough marketing lists to understand how that winds up having a failure mode that's awful.

Bryan: I know. But real talk; I could ignore you on any platform. So—

Corey: Oh, absolutely.

Bryan: I was trying to be nice.

Corey: Even in person, or on a podcast, it works super well. So, in your day job, you work on developer experience at VMware, which to my understanding is aimed at helping developers be more productive in Kubernetes. If I were to be uncharitable, I'd say that doesn't sound like a particularly high bar, given how much time I see people investing tilting at that particular windmill, but tell me about what you do before I start savaging it undeservedly.

Bryan: Oh, and it's fine. I mean, those who do, do those who can't talk about it, so I'm always good with that. But what do I do? Right now I'm looking at Kubernetes as a platform for deploying applications. And with that comes a whole smattering of challenges; something that we at VMware are trying to call the secure software supply chain. And what is that? Well, I’m not quite sure yet. But really, here's what it comes down to. As a application developer, why do you need to be a Kubernetes expert? And I don't want anyone to answer that question, because the answer is you shouldn't be. You could deploy applications to Linux, and not be a Linux expert. Tell me 15 syscalls. You can't do it, but you can write apps that run on Linux with Ruby or Python or even C for that matter, and I want that experience for Kubernetes.

So, a lot of my day is spent doing research. How do we get applications into clusters? How do we make sure that they can talk to other things? How do we manage what we have? How do we make it easier for others to deploy applications into Kubernetes? How do we package them up? What do the workflows look like? This is what I'm actually looking at. And really what it is, is it's more research because anyone can come and say that they have the grand plan, but as we see right now, in 2020, that's not true.

Corey: It feels to me that Kubernetes offers an awful lot of value for the developer side of the world, when it effectively from my perspective, at least, sort of realizes a lot of the initial promise that Docker had, where you can write your application however you want, and then more or less, throw it over the wall, and it doesn't really matter where it runs as far as what provider it's on top of, as far as who's going to be supporting it in which environments. As long as you wind up structuring your application correctly, which is a minor detail, we'll leave for the implementation folks, you should largely wind up being okay with it. And from my perspective, it seems like, from that development point of view, it has succeeded. Would you disagree with that?

Bryan: No, I would not disagree with that. So, Kubernetes will be six years old this year. There's been lots of success. If you think about it, I worked at Heptio, and Heptio got bought by VMware, and that's is a success in itself. We have companies like Red Hat, they got bought by IBM, and a big—I’m guessing, I have no insider baseball on this one—but I bet a big piece of that was what is happening in OpenShift. And then, you have companies like Weaveworks and Rancher, and they're all doing something interesting. There's a very interesting problem here, it's just a big problem, and it's taking us some time to, basically, pull at the thread to get to the real answer.

Corey: Part of my, I guess, objection to it, for lack of a better term is that I was never an application developer. If you've seen any of the code I've written, it is blindingly apparent why that is. FizzBuzz is, sort of, the pinnacle of problems I don't know how to solve. But I do come from the ops world where I was a grumpy Unix systems administrator because it's not like there's a second kind of Unix systems administrator, I automatically became dour and sarcastic and aged 40 years in the afternoon that I was first anointed. And my challenge with what I see of Kubernetes—and to be clear, it has been a couple years since I last played with it in any real depth—that the implementation and rolling out a stable Kubernetes platform on top of a single cloud provider, let alone multiple cloud providers was basically the sort of work that took months and then by the time you were finished, the level of observability into what was going on inside of that platform when you started seeing intermittent issues was, back then at least, either non-existent or incredibly complicated. That's still the case?

Bryan: Well, I don't think Kubernetes is solving that problem. So, that would be like saying, “I deployed this fleet of Linux on bare metal, and how do I determine what the motherboard temperatures are? How do I determine what the [00:08:18 unintelligible] disk or the SSD temperatures are?” you have to have a framework for that. And what Kubernetes is doing is giving you a framework. It's basically giving you the base of an operating system with which you can build on it. And fortunately, and this is actually the greatest thing, in a lot of cases, Kubernetes doesn't have a lot of opinions on things like monitoring. You can get pod and node CPU memory statistics.

But there's other tools out there, like what Prometheus has done, that needs to be applied, that you can actually get the fidelity of what you want. The problem is, is that there is a lot of tooling out here and you might not know. If you do know, you might not be able to get it installed right. And that's still a problem, but I don't think it's a strictly a Kubernetes problem. Kubernetes is—it's like blaming Linux for—here's a great one. It's like blaming Linux for ls not working. And you might think, “Well, what are you talking about? My ls command doesn't work.” Well, guess what? That's not Linux. Kubernetes is the kernel. And we need to really work on user-land.

Corey: Yeah, no, you're absolutely right. If the stack call isn't working, that I lost calls, then we're having a different conversation. But until then, it's user-land.

Bryan: That’s right.

Corey: My arguments—and I made this a little over a year ago now on Twitter, and it caused a stir—that in five years from now, no one is going to care about Kubernetes was my hot take, and I just re-upped it back in February and said four years. And I maintain that I'm probably right on this, but from the perspective—not that it's going to dry up and blow away—I think that only a moron would suggest such a thing now, but rather that it's going to slip below the surface of awareness. How many people need to be in-depth Linux systems experts today versus 10 years ago? It's still going to be there, but we don't really care how applications or workloads get orchestrated in the cloud, or in our environment, it just happens.

Bryan: Well, that's the whole point. I think that you are correct. If someone asked me about that, I will deny saying that, but Kubernetes will go away. So, here's a good example: Let's see, about three years ago, or four years ago, how do you install Kubernetes? That was the buzz. How do you install Kubernetes? And throughout the years, where we have cloud providers, and then we have vendors like Rancher, what we did at Heptio, and Red Hat, and others, installing Kubernetes is not really a problem anymore. Even on my development machine. I have a very large Mac, and right now I'm running six instances of Minikube, all 12 gigs of memory apiece. No problems. Spin them up, spin them down.

What's going to happen with Kubernetes over the next few years is that these things that we consider hard and consider difficult, we're going to solve them. And we're going to solve them in a way where they actually do disappear. And we're going to find a new set of problems because, I'll tell you what, I've been coding for money for 25 years now, actually a little bit more. And when I started out, like you, I was a sysadmin. And there was no—I mean, I had Linux on my desktop at home, but we did Sun, and we did HP-UX and we did AIX. And guess what? We've moved past all that right now. If you gave somebody an HP-UX box, or no, even better yet, you gave somebody an AIX box with their crazy Korn shell, people would look at you like you're crazy now. Where's my bash? Where's my sed shell? Where's my fish? So, we'll move past it.

Corey: And I think that that type of shift becomes largely inevitable. It's just challenging as we go through the period before we get there. I mean, even from my somewhat removed perspective, as a cloud economist looking at companies’ very interesting AWS bills, Kubernetes still causes challenges where, from the infrastructure point of view, when you have a full Kubernetes environment, there's a single application workload that's running from the cloud provider’s perspective called Kubernetes. So, when you see interesting data transfer patterns, interesting disk access patterns, it's just a very, very strange, single-tenant application.

Now, once you get a level of visibility into what namespaces are running, what workloads have been migrated into Kubernetes, and how those things interact, suddenly you see an entire ecosystem of an entire microservices application in most cases, that is working and humming away. But it is abstracted far enough away from what's happening underneath, that it becomes almost incomprehensible. I'm not saying that's necessarily a problem from the user or operator experience, just from a billing optimization perspective, which, surprise, was never anyone's first goal when building pretty much anything. Can I build this for the least amount of money possible, is not usually the clearest path to success for most products. Who knew?

Bryan: Yeah, you know, something interesting is so I built a cloud, I work for clouds, I work with things that became clouds, and now I work at a vendor. And I'll tell you what, from what I've seen, VMware is Project Pacific and the new way of thinking about running Kubernetes on vSphere. If you're still running your own data center, I think that some of the things that you brought up, like making Kubernetes that opaque thing are going to be a thing of the past. And I don't work on that team, so I actually can't tell you what they're doing, because I have no idea, but from what I've seen of the demos at various places, is that I think that once we've learned—so, it took us a while to actually learn what Kubernetes was, and how it felt and how it tasted.

And now that we have that, what we can now do is start molding it into our environments. And we're seeing that with VMware and Project Pacific, we're seeing that Google for better or for worse. GKE is actually a great Kubernetes experience. And I think Microsoft and maybe even Amazon are thinking about this in a little bit different way now. So, I think what's going to happen is, we expected that this thing was going to be a panacea and solve all of our problems. But us being well-adjusted adults should know that there's no such thing and that new tech takes a long time, especially big tech takes a long time to integrate, and we're just going through that cycle right now. But people are seeing successes and I enjoy that piece.

Corey: You mentioned Project Pacific a minute or two ago. What exactly is that?

Bryan: So, what Project Pacific is, and once again, disclaimer, I do not work on that team and I am not really familiar with how vSphere works. So, you have your vSphere cluster. You can enable Project Pacific on it whenever it comes out, and I don't know when, and then what happens is it gives you this concept of a supervisor cluster that can actually manage other customers or clusters. So, you can give all your developers their own clusters, and it becomes a real part of vSphere. So, if you're using VMware as networking technology, whatever NSX that you happen to be using, I think it works well with that. And then, you also get access to, like, vSAN. And you have real access to disk. And it just becomes another way to run your workloads on vSphere. And one neat piece about this, and I think AWS and Microsoft are able to do this right now as well, is that whenever you boot up a pod, and a pod in Kubernetes is just a container, or actually a set of containers that run your applications, they boot up in a VM. So, now you get all the same benefits of that, but running in a VM, so you get your own set of isolation and things like that.

So, I'm enjoying seeing that we can take this thing like Kubernetes and we can apply it to different domains. And, of course, everybody doesn't run VMware stuff, but most people do, if you're running in the cloud, and this is the crazy part, hold with me here for a second is that we always talk about moving between clouds, and it doesn't work. Yeah, it doesn't work, because it's a impedance mismatch. But being able to run Kubernetes in more than one place, and actually being able to abstract away disk with CSI that's built into Kubernetes and abstract away network with CNI, does get you closer to that place where you can move workloads. Well, they work all the time? Probably not. But many times it will.

Corey: This episode has been sponsored by CHAOSSEARCH. If you have a log analytics problem, consider CHAOSSEARCH. They do sensible things like separating out the compute from the storage in your log analysis environment. You store the data in S3 in your account. You know where it lives, you know what it costs, and then they compress it heavily while indexing it, and then they query that data using a separately scalable fleet of containers. Therefore, the amount of data you’re storing no longer is bounded to how much compute you throw at it, as well. It’s broken that relationship, leading to over 80 percent cost savings in most environments, and being a sensible scaling strategy while still being able to access it through the API’s you’ve come to know and tolerate. To learn more visit CHAOSSEARCH.io.

Corey: One of the, I think, common misconceptions about how I tend to approach things is I speak to the general case and people tend to assume as a result that I'm speaking to individual specific cases, multi-cloud is a terrible best practice, rant is a great example of this. I think that if you're designing something today, greenfield, that you're aiming to deploy seamlessly into multiple clouds. Barring a few edge cases, it's probably not the right direction to go in. That said, if you take a look at existing workloads, where there is a requirement to do this, then full speed ahead.

It's important that there be solutions that address all of these problems. I try and guide towards the common case of what people are considering as they're just dipping their toes into this cloud world. Big E Enterprise has so many interesting workloads and so many fascinating stories within it, that I feel like it's not getting the kind of attention that it deserves. Oh, that's legacy code. By legacy we, of course, mean revenue-generating. And what does that nonsense do? Only about $8 billion a year in revenue, why do you ask? Oh, you should move it to serverless. You should move yourself somewhere else because that's not ever going to be a viable story. Being able to meet customers where they are and give them the ability to migrate from where they are to where they want to be, to be able to address emerging technology trends and be able to effectively speed the time from a developer having an idea to that code running in production is really what it's all about. And I feel like, right now, a lot of the arguments in this space, have lost sight of that, and instead focused on the best way to achieve that goal, but have lost sight of that goal itself.

Bryan: Well, that's the problem. Whenever you look for the best way to solve something, that only works for you in that individual time and probably will never work again. And right before Heptio, I worked for an enterprise large bank, and we did multi-cloud. And I'll tell you what, we were not moving workloads between cloud one and cloud two, we had workloads that ran in cloud one, and we had workloads that ran in cloud two, and maybe we backed up some data between them. But even with us—and literally had an army of developers and engineers that could do all these things, and more money than you could shake a stick at, and we could not do it. But what we did is that we just found that best practices work everywhere. And I bet—and the crazy thing is that, you think that moving to the cloud is your biggest concern. I bet your CI and CD pipelines aren't very good. I bet the way that you actually deliver software to production isn't very good. Maybe we should spend more time focusing on those things, so whenever we have to move our platform—because we will one day—it'll be easier, we can just re-platform it. And I'm waving my hands now like it's magic because to the people who don't prepare, it does look like magic, but it's just that we work hard, and that we have an actual goal, and we have a plan to get to that goal, and then we use that and form our strategy, and do the right thing.

Corey: It feels to me on some level, like one of the greatest challenges that VMware has as a company is in their name. It's progressed so far beyond, hey, I want to virtualize a machine somewhere, how do I do that? Once upon a time that was super hard and expensive and, believe it or not, back then I was very anti-virtualization. I think I called it a flash in the pan. Surprise, I was wrong. Further surprise, as a futurist, if you get it wrong, no one ever calls you to account for it, which is super awesome. Obviously, I was very wrong on that. And now there's so much more capability stories, VMware has exploded into the modern cloud era in a way that honestly, I don't think most people saw coming. And on some level, it just sounds like it's aimed at the problems of yesteryear, even though it's very much not.

Bryan: So, VMware is way older than my tenure there. They did a whole bunch of interesting things in virtualization, and then networking and storage, this whole concept of a virtual data center, software-defined data center, you know, think they do a pretty good job at that. So, I think that the executives, they realize that and that's why we have this new brand called Tanzu. It's our cloud-native brand. And what we'll see over, I actually don't know, over the next few months, over the next few years, we will start shipping a new breed of software on there. And whether it be things like Tanzu Mission Control, which is helping development teams, or actually helping enterprises manage the Kubernetes clusters they have, or if it's something in actually booting up clusters, whether it be on vSphere or not, or is it something that I'm doing where we're looking at the developer experience? I think that the powers that be decided that that's a whole brand new experience for VMware, so let's give it a new name. And it's TAWN-zoo, not TAN-zoo. TAWN-zoo.

Corey: We do care intimately about pronunciation of various things on this show. That has been a bastion since the beginning. In fact, we were still very happy with the story that you folks got out of your acquisition of BIT-n-ay-em-eye. Some folks call it BIT-nah-mee. Those folks are wrong.

Bryan: You said BIT-n-ay-em-eye. [laughs]. I'm sorry, that’s funny.

Corey: Oh, yes, both of the founders of BIT-n-ay-em-eye basically facepalm. My personal theory is that Erica Brescia left to get away from my mispronunciation of companies, and as a result, that's why she's now the COO of GIF-huhb.

Bryan: Right. Well, of course. That's actually pretty funny, and I actually work with the other founder on projects, so I will be sure to use that pronunciation.

Corey: Oh, the best jokes needle at people to the point where they're lying awake at night just angry about them, but they don't punch down at anyone. That's the trick.

Bryan: Yes, definitely. And that's an interesting segue if we were to take it.

Corey: By all means.

Bryan: So, I think that one of the reasons why I do podcasts like this is people see me on Twitter, and they see what I say, and I want them to be able to hear the voice and then hear that whenever we talk about things that are bad, we're just sharing what we see from our point of view, and we're not punching down on it. If I say that, you know, I'm not underrepresented anymore, and you're overrepresented. That's not the punching at you. That's me changing the conversation so I don't sound like I'm less than you. And I think that we need more people doing these things out there, speaking their voices, who come from various backgrounds and have varied experiences. Because, you know what, we've been doing this whole white guy thing for a couple hundred years in this country. And I don't know, we might need to hit the reset button and try over again.

Corey: But things are going oh, so very well for everyone right now. What could you possibly mean? No, I'm right there with you on that. I try mightily to reach out to folks who don't look exactly like me. A weird dynamic that I've noticed has emerged, in the general sense, where I invite people from all walks of life onto the show. Invariably, when I wind up reaching out to folks who look like me, I can't even get the word podcast out completely before I get a hell yes, let's do it. Other folks, because of the nature of the world in which we live, have to be more guarded with that. There's a whole series of questions that people want to know of am I a secret trash goblin? Am I trying to do this to push a narrative that's going to be awful? I hope the answer to that is no, and I continue to get folks from a tremendous variety of different backgrounds and different employment situations onto the show, but no one knows their own reputation. I hope I'm doing the right things, but at some point, all we can do is guess and do our best and try again tomorrow.

Bryan: Yeah, when I was growing up, my dad used to say, “Be the change you want to see.” And that's all I'm trying to do. I'm trying to be forever positive, even though I see all the weirdness around me, and show people that if we can take control of what's going on now we can actually take control of our futures. And never let anyone tell you that no, you can't do that. Unless it's something wrong, then you shouldn't be doing that. But for our careers, and our careers in tech, and our careers in tech-adjacent things, I want to go, two years from now, I want to be the best JavaScript developer in the world. Put in the work, let's go do it.

And I think that if anyone should be able to have those kinds of thoughts, if they can put in the work, and they should go do it, and they shouldn't have to be scared that someone else is going to steal their thunder, or someone's going to look over them because they are a certain way. And that's not right, and that's why I use myself. I mean, I'm, I don't know, third-highest level of engineer at VMware, I am the probably, I think I am the highest-ranked engineer of African descent at VMware. I have a lot to lose, but I also have A lot to gain by helping you out. So, that's why I'm out here having these conversations with people.

Corey: From my perspective, I mean, I'm a white guy in tech. My failure mode is a board seat and a book deal. So, I feel like I have virtually nothing to risk. So, what I did is, as soon as I had the platform, was start attempting, with mixed success, to get this to a point where other people can use the platform to tell their stories. I've learned an awful lot along the way, and Lord knows I'm not done yet. One thing that continues to surprise me, and some folks may have noticed this, and I don't think I've talked about it yet on this podcast, but I've started making fun of things like service names a lot more than the actual substance of newly launched services. Because it turns out that sure you have these giant companies out there that are launching these things. And yeah, I can make fun and it's punching up at giant multinationals. But there's usually a small team internally who lost blood, sweat, tears and weekends, building a service, and if the first thing that comes out of it is me making fun of it, that doesn't feel too great. So, making fun of service namings, which people did not spend 18 months on, I hope, is the sort of thing where it's still fun, it's still jocular, but no one hears what I have to say and feels crappy as a direct outgrowth of that. Some days I get it completely wrong, but I am trying.

Bryan: So, that's a actual interesting way of thinking about it. I think about things that I say sometimes because if I typed or said what I thought, then people wouldn't have the context. So, I started doing this whole whiteboard series, and it's all satire, and it's all poorly done, of me sitting in front of a whiteboard with a t-shirt and something on the whiteboard. And that's actually how I do it. It's it basically de-arms my message, but my message still gets across. And I think that we all need to find our ways with which we can share our message and get those feelings out because I don't want to sit on some of my angst, go crazy. Again.

Corey: It's different messages resonate in different ways. I find that sarcasm was my first language growing up. So, it was always a easy path for me to go ahead and alright, that quiet voice inside, why don't you try using that and see if it resonates with anyone else? It turns out it does. But now it's a question of great, what do I do beyond that? Just being snarky and crappy and sarcastic about everything doesn't get very far. I mean, having you on the show, for example, and then we'd spent the last half hour of me making uninformed but belligerent comments about Kubernetes, that doesn't build anything and it just cements the idea that this entire industry has to be oppositional, where in order for one thing to succeed, other things have to fail. And I don't believe that that's the case, nor has it ever been.

Bryan: No, it's not, and you know what, I envy people who can be sarcastic all the time. I mean, I'm mostly cynic, but I can't be sarcastic because, in a workplace, people get scared, and they take me too seriously, so I have to really be direct with what I'm saying. So, I'm envious of that, but on the other side, I respect it. And I also respect people who can inflect them themselves and understand that this is how they are being presented in the world, and how they are presenting themselves in the world. So, it takes all kinds to make a village, so we can't say how people should be. But as long as you're not doing anyone dirty, not punching down, or really not even punching up, you shouldn't be punching up. You should be like, kind of like—I don't know, maybe we should punch up, but don't ever punch down.

Corey: That's the rule. And if it helps anything, my sarcastic nature was the number one contributor to the reason that I spent most of my 20s getting fired energetically from all kinds of companies. It takes time to season it and find a path forward. I don't recommend this for anyone. I just do it because I don't know how to do anything else.

Bryan: I've only been fired twice, for my first two jobs. And I learned very important lessons about the world. But ever since then, usually what happened is, oh, I'm not going to get promoted. Well, I'm out of here. And to the point of, I would left on—one day, the guy said, “Well, I don't think you're going to get promoted this cycle.” And I said, “Well, that's cool. If I don't, I'm taking me and the rest of the sysadmins with me today.” And he laughed. So, I took him—I took the rest of the sysadmins with me today, we walked out of that office. But over the past few years, it's been getting better, where I find that, at least at first, people are pretty welcoming. And now that I have this position at VMware, and I think it's it's definitely been a lot better. So, hey, kids, it can't get better. And then, also don't speak your mind all the time. Sometimes just keep it to yourself.

Corey: I have no ability to do that.

Bryan: Oh, I wish. I mean, this is why I have a wife.

Corey: Yes.

Bryan: I save it all up for her. And then, whenever she's done working, I tell her all the things that I was going to tell somebody else, and she does the same to me. And then, we're both better for it.

Corey: I made the interesting life decision of marrying a corporate attorney. And suddenly I have a different level of scrutiny on some of my more energetic online stunts. But it's definitely been an ongoing process. We do the best that we can. If people want to find you on the internet to hear more about your philosophy on these things, about how the Kubernetes developer experience is continuing to evolve, and other thoughts and commentary, where can they find you?

Bryan: So, I have a Twitter account. It's @bryanl. And then, I also have an Instagram account that I'm not going to tell you about because I can't understand Instagram. I don't understand the whole picture thing, and so we'll just stick to Twitter.

Corey: Thank you. I thought it was just me. The closest I get to Instagram is mispronouncing it as QUINN-stagram, and then everyone groans, and won't talk to me anymore.

Bryan: Yeah, I think I have like, eight posts over many years. And that's about where I am with Instagram.

Corey: Well, we will put the link to the Twitter account in the [00:32:16 show notes]. Bryan, thank you so much for taking the time to speak with me today. I appreciate it.

Bryan: Oh, thank you for having me.

Corey: Bryan Lyles, senior staff engineer at VMware. I'm cloud economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts. And if you hated this podcast, please leave a five-star review in Apple Podcasts and a funny comment so I know where to improve.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Jill Rouleau

Jill Rouleau is a member of the Ansible engineering team, focused on maintaining AWS and other Cloud modules. Prior to Ansible, they worked on OpenStack, using more than a decade of operations and SRE experience to improve deployment tooling for cloud operators.

Links Referenced:

  • Jill Rouleau's Twitter
  • Ansible
  • RedHat

Transcript:
Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is brought to you by DigitalOcean, the cloud provider that makes it easy for startups to deploy and scale modern web applications with, and this is important to me, no billing surprises. With simple, predictable pricing that’s flat across 12 global data center regions and UX developers around the world love, you can control your cloud infrastructure costs and have more time for your team to focus on growing your business. See what businesses are building on DigitalOcean and get started for free at do.co/screaming. That’s D-O-Dot-C-O-slash-screaming and my thanks to DigitalOcean for their continuing support of this ridiculous podcast.

Today's episode is sponsored by Springboard. If you want hands-on experience, getting deeper into machine learning, check out Springboard's machine learning, engineering career track. This program is for existing software developers who want to get deeper into machine learning without driving themselves mad. With springboards one-to-one mentorship that is closer than ever. They also offer a job guarantee in which you pay nothing until you're gainfully employed. To learn more, visit springboard.com and apply for free. Use the phrase "AI Springboard," and the first 20 students will get a $500 scholarship. That's springboard.com code "AI Springboard."

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I always like talking to people in the open-source community, and that goes double if they're not lunatics. From that perspective, welcome Jill Rouleau, a member of the Ansible engineering team focused on maintaining AWS and other clouds, if that actually existed, which I'm not convinced that they do. Your day job is a senior software engineer at Red Hat, but you identify more as a member of the community. Jill, first, welcome to the show.

Jill: Thanks for having me, Corey.

Corey: Thank you for agreeing to tolerate my various slings and arrows, by which of course I mean, stupid puns. So, your day job is working at Red Hat, but you do identify as a member of the Ansible engineering team, which at first struck me as really discordant until I remembered, oh, that's right, Red Hat bought Ansible, and suddenly everything made sense again. Every once in a while I drop something off of the mental stack. But, now that I'm up to speed on this, talk to me a little bit about what is it that you would say it is you do in the context of the Ansible community?

Jill: So, the AWS modules for Ansible are entirely community-made content, all of the modules have been written, and submitted, and are maintained by various members of the upstream Ansible open source community. So, my job is mainly to, kind of, oversee that community, make sure that things are moving, help review pull requests, maintain CI, and just generally make sure that everyone has what they need to get those modules, kind of, submitted, improved, update new features, and maintained so that people can make use of them.

Corey: So, Ansible was what I like to think of as, well, I'll call it second wave, but no one is going to agree with that, so let's just get that out in the open right now. Originally, if we go—we'll ignore the early era of Bcfg2 and other terrifying things—the real first broadly adapted wave of configuration management systems was Puppet and Chef. You can decide the order on your own. I am not a cloud historian; I don't care. The next phase of, “Great, we're going to go ahead and do something better.” And the two shining lights in that space at the time were SaltStack and Ansible. I was one of the early developers behind SaltStack, which is why it, sort of, isn't the huge thing that it could have been because everything I touch withers and dies. Ansible, on the other hand, was something that I didn't touch and now has become a household name in the infrastructure automation space. How have you seen that, from someone who's actively been involved on, shall we say, the winning team?

Jill: [laughs] I think they both win in different ways. I was actually also an early SaltStack user. I, kind of, got into this space from a background as a sysadmin who was using these tools, and eventually, I wanted to start contributing back to them. I think one of the big differences you find between maybe some of the earlier tools and what we see now is the ease of use and the ease of modification. Salt and Ansible are both written in Python and maintained through a horrible series of YAML files.

Corey: Oh, yes, we should instead all organized around XML as the way in the light of the future.

Jill: Yeah, no, let's not with that. But I think Python is a very accessible language if people need to patch something, modify something, write something custom to deal with some homegrown internal service that they have. And YAML, for all of the horrible things you can do with it, is somewhat more readable than things like XML or having to learn a new DSL. It's at least approachable from looking at a playbook or a state file for the very first time and trying to read in, kind of, human language, what those things are doing. And I think that helps with adoption and with getting people to contribute back when the bar is lower because it's more familiar tools or languages.

Corey: I absolutely find that that's one of the better aspects of YAML. With JSON, it seems like you get lost in a sea of braces, parentheses, commas, missing commas, commas that need to be there, and ultimately feeling like you're trying to talk to a computer rather than something a human being can consume. For better or worse, I really do think that YAML is the right answer for an awful lot of these things. But that's not really where I tended to see most of, let's call it, the mistakes get made. For me, it was never the configuration management piece. It was the—I mean, if we look at it across the board, there was a shift that I think none of the players that I've mentioned saw coming, and that was this giant embrace of immutable infrastructure and early on that was crap. Oh, you want to do a one-line change, great. We're going to build a new AMI, we're going to ship a whole bunch of new systems out there, or VMs, or whatever the hell insane thing we're building now, and that'll be great. Your code will be in production in just 45 short minutes. And this was laughable. Then Docker, Docker, Docker came along, and suddenly, the world shifted as developer workflows started to impact operations. Everyone was angry about this. I was certainly angry about this. But you can't deny today that if you're starting something greenfield, that configuration management does not have the center place that it once did in the ecosystem. Would you agree?

Jill: I would agree with that. I think with Ansible especially, I'm not sure as much with SaltStack, But Ansible definitely had a lot of very early adoption, not from operations people but from developers. You know, when I first heard about Ansible, it was fairly new. I want to say it was maybe 2013 at, actually the Southern California Linux Expo. There was a panel and they had all of these config management tools up there. And someone from Ansible was there talking about it. And I looked at that, and I'll admit, I laughed. Oh, that's just a thing for developers who don't want to go through process. I don't want developers SSHing directly into my servers as root. Haha, I would never want something like that going on. And here we are today. But I do think that easing that developer friction caused a lot of adoption that sysadmins, maybe traditional sysadmins were not expecting and had to, kind of, catch up to.

Corey: One of the things that I think sped Ansible adoption that they got very right, and on the SaltStack side we got wrong, was the idea that every communications model happens over SSH. There have to be keys in place, it has to be able to address the right thing at the right time. But we had already—from our legacy of managing systems—we intrinsically understood SSH. We knew how it worked; we knew the security model; we knew the problems, and pitfalls, and caveats that went into that. Whereas with Salt, it was, oh, we're going to run on top of ZeroMQ, then we’re going to use some compression on top of it—it originally was Python Pickles, then became MessagePack.

And you had to learn how all of this protocol stuff worked, and there were new ports to care about in firewall contexts, and suddenly it looked like an uplift, even though it really, kind of, wasn't in some ways. Ansible nailed that in a way that I don't think we understood early on in the Salt world, that this mattered. It resonated because it fit the mental model people had of how systems were going to work. Back in, I want to say the very early 2010s, I had a boss who decided to effectively imagine a configuration management from first principles, Hacker News-ing his way through it. He was kind of right because he conceptually built Ansible in his mind, only crappy because Hacker News first principles.

Jill: You know, I would love to see Hacker News actually build something that they say they can build in the weekend, and then maintain it for eight years. Ansible has been around for almost eight years, and that's the maintenance of that thing that you build in a weekend via Hacker News, is always, kind of, the kicker. But, there are some ideas there, of having something that's simple and easy to use, and that doesn't require shoehorning into a lot of your infrastructure, that just kind of works with whatever you're doing. Ansible doesn't have to take everything over. It doesn't have to be the only answer. And you can use Ansible to deploy your Salt stuff and deploy your Puppet things if you really wanted to do that. But you're right, yeah, it fits in using the models we already have of SSHing onto a box and hopefully not doing things by hand; doing things by YAML instead, and doing a hundred boxes all at once.

Corey: What always astonishes me is no matter how dyed in the wool anyone is, they're adamant that everything we do is immutable infrastructure. It is cattle, not pets. It doesn't take more than about 30 seconds to find the exception case and, “Oh what do you use for that?” “Oh, Ansible. But don't tell anyone.” It is super easy to get up and running with. It's a great tool. The challenge, as well, is at first getting people to admit they're using a thing that for some reason the cult of Docker has convinced them that they should be deeply and profoundly ashamed about is the first problem. The next challenge is figuring out how to get people involved. I mean, you say that you're a community engineer. What does that look like? How do you get people to care about a configuration management system in this year of our Lord 2020?

Jill: Most of the people, I think, that I see coming into Ansible and then sticking around are trying to scratch some itch that they have, you know? Oh, AWS released a new feature and the modules don't support it yet. Well, I want to use that feature, so I'm going to write a patch, and submit it up on GitHub, and see if I can get it submitted for inclusion. And then, getting those people to stick around is one of the hard parts that we have to do is encouraging them that, “Hey, that's awesome. Thank you for submitting.”

Making it a good experience so they want to keep coming back, so they can see all of these things that I need to do. You know, if you're doing anything significant on cloud for all that we say we want to automate it anything away at some point, you're probably writing something. You're writing some amount of code, Why not submit that back upstream so that you can pull it back downstream and use it via Ansible or so that other folks can. It's getting people who are already doing these things in isolation to want to come back and give it back to the community.

Corey: One of the early community attributes in SaltStack, which really got me into contributing code open source is something that in recent years Ansible has adopted, and I love it. If you go back into the mists of time, from my earliest pull requests against SaltStack, you saw that Tom Hatch, the guy that built SaltStack and founded the company, would accept whatever I proposed, and then immediately there would be another pull request that was merged from him fixing my horribly broken idioms. Now, what people don't see is the fact that, first, Tom Hatch was and remains the nicest guy in the world. And he never said a word of criticism about anything I did, even though, honestly, it kind of deserved an awful lot of criticism.

It was the welcoming aspect of the community that really inspired me to continue sticking around and participating in this. And recently Ansible has definitely gone down that path as well. It gets away from some of the old traps of overly corporate software in some respects where it's well, we need to have what effectively looks like a change advisory meeting on every pull request that goes through, and you can see the governance gone amok. It seems that Ansible is largely avoided that. It's still welcoming for folks who are, “Well, I don't really know how to code, but it's Python, so how hard could it be?” as I once said. And it's there in a way that a lot of tools and projects simply aren't.

Jill: So, that's awesome. I'm really glad that you see that and feel that way because we do work hard—Ansible is a really large really, really busy project and it is challenging to scale that type of feeling for people when they come into the project. There are literally thousands of modules in Ansible on top of the actual core engine code, and everything else that we have to maintain to make Ansible work. There are hundreds of people putting work into it, and only some of them are actually core engineers or Red Hat people. The majority of the contributions that we get are from the community and that balancing act of figuring out—we're not maybe going to merge everything that folks sent in and then come along and clean it up later, but if you open a pull request against Ansible someone is going to review it and give you feedback and hopefully be welcoming and let you know that, “Hey, thanks for the submission. We have, maybe, some feedback to get it into shape.” We require, you know, CI tests so that we don't merge broken code, but we work really hard to make sure that folks have a good experience and want to come back.

And I think one of the things that has helped is we've empowered a lot of our community to be that person. You know, you don't have to wait for myself or one of my team to review your code. We actually have community members that are subject matter experts on the different things that we do, like AWS, or various other modules, and they're empowered, once they've been a member of the community for a while, to pay that forward to other people the same way that they were welcomed in and trusted to contribute to the project. So, hopefully, that's a part of it.

Corey: Absolutely. Again, if you're on the other side of this, and someone is new to the project and starts contributing, and your immediate response is, “Listen idiot—” If that's how you're starting your comment, maybe reconsider about whether that's the impression you're trying to give, even if what they're proposing is patently ridiculous, as is almost everything I wind up submitting, either intentionally or accidentally. Ansible lives on [00:15:48 GIF-huhb], or GitHub as some people choose to mispronounce it, and one of the great features that GIF-huhb offers is inside of a given project, you can tag various issues as good first issues for someone looking to get into a project. What I haven't seen yet and really wish they would put out there, and Ansible would be a great fit for such a thing is good first projects to contribute to. One of the challenges of another common player in the space is Terraform out of HashiCorp. I constantly have things I want to improve, and I periodically go over and start to build something that might address the problem, and I got about 10 seconds in before I realize, “Oh, wait, that's right. It's written in Go.” Go is for smart people, and I can only stumble my way blindly through Python, Bash, and a little bit of Perl due to previous life choices that went awfully. So, that's not really available to me.

But being able to say, “Great, I have moderate Python skill and I'm looking to get involved in an open-source project. What can you recommend?” It turns out it's super hard to get a good solid recommendation because asking any person for this, you get a giant pile of bias back of whatever project they love, whatever problem they're trying to work on this week. It isn't the most, I guess, accessible onramp for folks who are very easily overwhelmed by the sheer variety of what they could be working on.

Jill: Yeah, so and with Ansible, especially because right now, today, we have a single repo that contains the Ansible engine code, and all of the modules. This is the batteries included model, where we have everything from modules that can control your Cisco switches to your AWS Cloud to your Linux boxes, OSX machines, Windows hosts, security appliances. Someone who shows up and just says, “Hey, I want to help, what can I do?” There's almost too many things that they could do. So, some of that, for us, is asking, “Okay, what are you interested in? What are your skills, what are you as subject matter expert on?” and then maybe getting them paired up with either a working group or an interest group for that specific area, like cloud, or AWS, or network appliances, and then, kind of, moving down. But that does require them to make that initial showing up and asking.

It's a lot harder to look at just showing up to the repo and saying, “What can I work on?” because there are so many thousands of things that could be worked on, which is actually a scaling challenge that we've been working on right now, for the last year. How do we scale the size that we've gotten to and the breadth that GIF-huhb—so to say—dot com/Ansible/Ansible covers? How do we scale the management of that project, and the community onboarding, the community management of that? So, in the future, later this year, we will actually be splitting that out. Ansible Collections will be a new packaging feature for how content that goes into Ansible can be, kind of, split up and managed from a repo and packaging perspective.

So, at least for the AWS side of things, one of my plans, once we split that out into its own repository that lives on GitHub, we can have things like project boards that make sense, and wikis, and different things using some of those GitHub tools so that people can show up, just look at the repo that interests them, just look at the content that their expertise makes them a good fit for and say, “Oh, here are some projects. Here's a board that has some ideas, some open issues that needs to be worked on,” and make it a little easier for people to get directly involved with just the things that they care about.

Corey: This episode is sponsored in part by N2WS. You know what you care about? Many things, but never backups at least until right after you really, really, really needed to care about backups. That's what N2WS does for your AWS account. It allows you to cycle backups through different storage tiers so you can back things up cost-effectively and safely. For a limited time, N2WS is offering you a hundred dollars in AWS credits for setting up their free trial, and I encourage you to give it a shot. To learn more, visit snark.cloud/n2ws. That's snark.cloud/n2ws.

Corey: The problem is, is that we all have—at least those of us who've been around long enough—have experiences with the exact opposite of what you just described as far as welcoming and encouraging and enthusiastic community. [coughs] Debian [coughs]. Sorry—

Jill: [laughs]

Corey:—that's not—something in my nose here. So, as a result, some people were driven away from this. What's curious to me is your background. This is very much not your first open-source rodeo. Prior to this, you were heavily involved with OpenStack, which was a fascinating project across a wide variety of different things. And it was interesting watching that evolve. For a long time I really, really wanted to see that succeed. And for one reason or another, I get the sense largely due to governance, it didn't fulfill the promise it had laid down. And I still feel that lack to some extent. Now, it's obviously still seeing adoption in certain sectors; telcos love it. But that was interesting just watching from the outside. Can you tell me a little bit about how the community piece of that worked?

Jill: That has been an interesting challenge for me moving from a project like OpenStack to something like Ansible. OpenStack is also a very large and very complicated project or set of projects. One of the things that happened there, though, is there was a lot of hype, or, you know, like you said about, “Oh, it's going to take over the data center. It's going to be the new way of everything. You'll be able to run your own cloud, your own data center, and everyone is going to do this thing.” For all of the hype that there was, though, and all of the excitement—and there was money being thrown around; there were huge parties; there were lots of excesses—which surely aren't happening in any other communities at the moment, that hype train hasn't just moved on to any other projects—but the people that showed up to do the work, were the people from the telcos or from organizations like CERN, people that had specific use cases that weren't being solved, primarily were the people that showed up and said, “Hey, this sounds really cool. I have engineers that I want to put on this.”

It ended up being a very, very corporate-sponsored project where you had all of these different organizations showing up and saying, “I have use cases to offer. I have engineers. I have test environments that we can use, I want to do work.” And not in isolation. Public Cloud certainly had a part of it in easing barrier for smaller people to just get things going without having to deploy a private cloud. But I think that was a big part of how OpenStack ended up moving into this niche where it's serving a couple of really specific verticals for which there is almost no other alternative on the market, but a lot of that was driven by these corporations showing up and saying, “I want to commit developers to this. I want to contribute engineers. I'm going to send my operators to come bring their use cases.” And that ended up being a big part of what drove OpenStack in that direction.

With Ansible, it's a lot easier for people to make what you might call drive-by contributions to do that. Hey, I scratched an itch I had. I wrote a module or I patched a module, have a contribution, and move on. It was much more complicated to do something like that in OpenStack, where you're dealing with really complex infrastructure. There had to be, kind of, more context that you had to learn to get involved in contributing to understand how do I actually manage a hypervisor? How do I actually manage software-defined networking? So, you ended up with a different, and much more static, group of corporate and specific use case backed contributors. That pattern doesn't apply as much to something like Ansible, where people are using it for so many different things, and it's much higher up the stack, you can just have one-off, two-off low barrier to entry contributions.

Corey: You can wind up saying an awful lot critical at every big company in this space. Well, maybe you can’t. You have an actual, you know, job and employer, but I can. But there's precious little to fault with how Red Hat and its various associates, affiliated projects, and acquisitions, and divisions and whatnot, deal with the open-source community. Not everyone likes the outcome of every decision, but I don't know too many people who are going to sit up and say that they felt that their concerns were not heard. That they couldn't communicate with the rest of the community, etc. It's strange in that Red Hat feels almost like a unicorn, where they are more or less the success story about open-source companies going public, and they've been the edge case exception in many respects for 20 years. It's really interesting watching the journey continue to evolve.

Jill: Yeah, I mean, and like all open-source companies, Red Hat is full of people who care passionately about the community and what we do. We're just all, kind of, lucky enough to get to do it as our day jobs instead of side projects. But, we all really do care that much about the community and what we're doing.

Corey: So, if you were giving advice to I don't know, the people that we were back when we first met, what was it now 15 years ago almost, looking back, the road that we walked is very clearly closed. How would you go about finding the path forward into a world of contributed to open source in a day when there are so many different directions to go in, and it is always increasingly murky to find a path to get somewhere sensible? Where does the next generation come from?

Jill: So, this is actually something that I end up trying to answer a lot. I'm also super involved in user groups, I help run my local LUG. Linux User Groups are still a thing, it turns out, in 2020. And we get young folks coming in all the time, whether they're recent grads or maybe they can't afford to go to college here in America, where that's the cost of a small mortgage. And they want to know, how do they get their foot in the door? How do they get that first job? How do they get started? I got my start—I’m a career changer. This was not my original plan. I do not have a CS degree. I got really lucky that I, kind of, got in approaching the tail end of when you could just show up and say, “Hey, I know stuff about computers, you should give me a job.” And people would do that. Which is a little bit bonkers if you think about it, but that's—and I think you, kind of, joined the industry at a similar time as I did.

You can't do that anymore. You can't just show up and say “Hey, I've been playing around with—I got Linux running on a PC in my spare bedroom. You should give me a job managing your servers.” And it's hard to be in that spot now. There aren't a lot of great answers. You have the [00:26:06 unintelligible], well, throw a bunch of stuff up on GitHub, and build a website as a portfolio, and hope and pray, but the market right now is so flooded with people trying to do that, that it's hard. And the only advice I can really give people is show up to things. Show up to meetups, and meet people, make connections, network, go to conferences, if you can afford to. There are lots of really great local conferences that tend to be affordable. Show up to things and talk to people because right now it really does feel like you don't have that, the advantage of a CS degree and an internship, getting your foot in the door right now is almost entirely about networking, and meeting people, giving talks at meetups and saying, “I know stuff. I'm willing to work hard. I'm willing to learn things. Can you help introduce me to someone, give me a referral, walk me through making a contribution, connect me with a project that is not going to exclude me, or crap on my work, or not pay attention to my pull requests?”

And having that kind of personal touch, helping other people, helping the next generation get in, it's, kind of, on us to help them. There’s some paths. Outreachy is a really good one that OpenStack is a big member of, getting people paid internships that don't have to go through a CS program, and matching them up with open source projects where they're going to get one-on-one mentorship and help making those first contributions, learning their way around a project, learning the 21st-century nettiquette, as it were, for getting involved. I love Outreachy. I think they're great. I wish more projects would support them. But it's things like that, that I think we need to be doing more of.

Corey: That's the big problem that I think we see almost industry-wide is we seem to think that only the super senior people have something valuable to contribute, but that is very clearly not true, even at the company level. I lose sight of the sheer number of companies out there who I can ask them, “Great, explain to me what you do.” And they talk for a minute and, “Okay, I get it. Now, explain to me what you do if I don't have a decade of experience as an engineer.” And they have no idea where to begin. Spoiler; a lot of junior people are terrific at being sounding boards for telling these stories, or coming up with ways to make it more accessible to a broader group of folks. It's not just the people who can think about something hard enough, and it starts smoldering. Everyone has something to contribute, and I really wish that there was more of a broader awareness of that.

Jill: Yeah, I mean, if everyone that was working on software, came from the same background; had the same use case, things would be really boring. We wouldn't end up with a lot of tools and things that we have now because everyone would be using the same editor, on the same platform, with the same hardware, in the same configuration. There would be a need for less of us. We all have different needs, we have different backgrounds, we have different ways of approaching problems. I've had the good fortune to work with some really amazing junior people over the years who have caused me to learn new things and question things that I know, and I'm so grateful for that.

I can't imagine why anyone would not want to bring in more voices and more people. If we're not bringing new people in, I don't understand what we think is going to happen to our projects in the next couple of decades. Eventually, all of us will age out of being able to make massive contributions. Who is going to be maintaining these tools in these projects or building new better ones if we're not bringing new people into the industry and into our communities? Just survivorship of projects. There's not going to be any more gray-haired neck-bearded old school Unix hackers. We've got all of them. Everyone that was on IRC in the 90s that wants to be doing this stuff mostly is. We have to let new people come in and do stuff.

Corey: There's no room anymore for gatekeeping.

Jill: There never was. I mean, it was happening and it's still happening, but it was never okay. There was never room for it. I mean, we're suffering now from all of the people that we have either kept out of the industry or pushed out over the years. That's equally as much of a problem that I don't think we talk about nearly enough. How many people have been driven out of his industry because of gatekeeping that had something of value, that had experience and knowledge and we drove out? Yeah, that's not okay.

Corey: No, it's really not, and you're correct. It never has been. For some reason, I think a number of us deluded ourselves at the time into, well—believing all the tropes of, “Well, the good people will be able to put up with a toxic, crappy environment.” And what a broken, weird thing to say, or even to believe. But I was advocating for such things, many moons ago. It's was always come from place of insecurity. It's, “Well, I'm not anything special, so if we let anyone in, they'll learn that there's not anything special. So, I have to cling to this thing that makes me be unique and different, and of course, better than everyone else.” And there's room for more people at the table. My God, it's strange seeing how so many of those conversations played out, I look back at some of the things I said in the early naughts, and I'm ashamed.

Jill: And I think most people have something like that. We've all had to do learning or growing about something at some point in our lives, whether it's gatekeeping to open source communities or something else. The idea that people can't learn, and grow, and become better people is just patently false. Everyone has had something at some point in their life that they were mistaken on, or that they maybe were a bit of a jerk about it to someone at some point. If we can't just say, “Hey, oops, I made a mistake, and I'm not going to make it anymore.” And be big enough to own those things and move forward, I think that's a sign of a better person. Being able to say, “Hey, I've screwed up in the past, but I'm going to do better going forward,” than someone who just professes to be perfect all the time because no one is perfect all the time.

Corey: Oh, absolutely.

Jill: Except maybe you, Corey.

Corey: Well, hardly. I mean, when I say I did some things that I regret and feel ashamed about, I'm not talking anything horrifyingly egregious that would get me voted off the island. I'm talking about answering beginner questions on IRC with RTFM. Go read the freaking manual. I'm not talking about, “Well, you're not actually a person.” No nonsense like that. My god, I was never that big of a dumpster goblin. But even now it's everything is new to someone, and rather than viewing your castle as being under siege by newcomers, you get to share this awesome thing with other people and watch them learn. And fun story: If you teach someone something, you get to experience it again through their eyes. Also, I can't think of a better way to learn something than explaining it to someone else.

Jill: Yeah, I totally agree. I was fortunate enough a couple years ago to launch a new NOC, or a Network Operations Center, at a job that I was working at—

Corey: The home of the original Knock Knock joke.

Jill: Hahaha. [laughs], but it was at an organization where we'd never had a junior team before. And it was, “Well crap, we're going to hire these people and put them in charge of really important systems. We're going to have to train them. We're going to have to figure out how to actually work with people who aren’t super senior engineers.” So, I put together a three-week training program, went out, executed it for all of the onboarding new hires. I had a blast. I had so much fun watching these folks learn how DHCP works because you know, DHCP… We didn't have to troubleshoot all of the DHCP problems on new-hire laptops.

That was spectacular. I would love to do that again. But watching them get it, watch them figure out, how do I troubleshoot this? How does it work? And then making connections between things and then going off on their own and working, and it was like, “Oh, my little NOC techs are all independent now, and I trained them, and it was awesome.” And getting to see them grow was one of the best feelings. Seriously if anyone has not done that before, has not mentored someone at a significant level, or trained someone a totally new skill and watched them grow and learn and become independent with it, do it. It is the best natural hot you can possibly get.

Corey: Fantastic. And I think that's probably a good place to wind down this conversation. If people want to learn more about you, your thoughts on various things, or how to get involved in the community, where can they find you?

Jill: Well, you can track me down on Twitter. That's @jillrouleau. That’s J-I-L-L R-O-U-L-E-A-U. Or on the Freenode IRC network—we are still alive and well on IRC—I am jillr, and you can come track me down in any of the Ansible channels on IRC.

Corey: Thank you so much for taking the time to speak with me today. I appreciate it.

Jill: Thanks, Corey.

Corey: Jill Rouleau, member of the Ansible engineering team and senior software engineer at Red Hat. I am Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple Podcasts. If you hated this podcast, please leave a five-star review on Apple Podcasts and I will do my darndest to pretend to care.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About the Jeff Smith

Jeff Smith has been in the technology industry for over 15 years, oscillating between management and individual contributor. Jeff currently serves as the Director of Production Operations for Centro, an advertising software company headquartered in Chicago, Illinois. Before that he served as the Manager of Site Reliability Engineering at Grubhub.

Jeff is passionate about DevOps transformations in organizations large and small, with a particular interest in the psychological aspects of problems in companies. He lives in Chicago with his wife Stephanie and their two kids Ella and Xander.

Jeff is currently writing a book “Operational Anti-Patterns with DevOps Solutions” with Manning publishing due out in Spring of 2020. He’s also the chapter president of Blacks in Technology

Links Referenced:

  • @DarkAndNerdy
  • Article: The First DevOps Hurdle
  • BITCon
  • Blacks in Technology
  • Jeff's Book: Operations Anti-Patterns with DevOps Solutions

Transcript
Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is brought to you by DigitalOcean, the cloud provider that makes it easy for startups to deploy and scale modern web applications with, and this is important to me, no billing surprises. With simple, predictable pricing that’s flat across 12 global data center regions and UX developers around the world love, you can control your cloud infrastructure costs and have more time for your team to focus on growing your business. See what businesses are building on DigitalOcean and get started for free at do.co/screaming. That’s D-O-Dot-C-O-slash-screaming and my thanks to DigitalOcean for their continuing support of this ridiculous podcast.

This episode is sponsored in part by N2WS. You know what you care about? Many things, but never backups at least until right after you really, really, really needed to care about backups. That's what N2WS does for your AWS account. It allows you to cycle backups through different storage tiers so you can back things up cost-effectively and safely. For a limited time, N2WS is offering you a hundred dollars in AWS credits for setting up their free trial, and I encourage you to give it a shot. To learn more, visit snark.cloud/n2ws. That's snark.cloud/n2ws.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Jeff Smith, the Director of Product operations for Centro in Chicago. Jeff, welcome to the show.

Jeff: Hey, Corey, thanks for having me.

Corey: So, there are a lot of things we can talk about here. But really, I think the most interesting to me personally, and therefore for absolutely everyone else, is that you are writing a book, the book that I first heard about and did two things. One, scheduled this podcast recording, and two, became angry and annoyed that it hadn't occurred to me to write something like this myself, but fortunately, you are. Tell us about your book.

Jeff: Sure, yeah, so the book is titled [00:03:02 Operational Anti-Patterns with DevOps Solutions]. There's a funny story behind that title that I'll get to as well. But, when I was originally approached about writing the book to Manning, the first thing I thought of is all of the DevOps books that I read were really sort of targeting the corner office. There were a lot of conversations about organizational restructuring, and seat shuffling and responsibilities. And when I first started looking at DevOps and what it could do for me in my organization, I knew that I didn't have that sort of backing. So, I really started with a groundworks guerrilla campaign to sort of prove out the value of DevOps without having the entire organization behind me. And when I did that, I quickly realized that there were all these different patterns that were contributing continuously to the problems in the organization. And I talked about that a bit in the book, but the paternalist syndrome. Everyone knows about the grumpy operations group that says no to everything. Well, why is that? Why are they like that? And how do we go about helping teams and organizations correct those things? So, as I started to collect these patterns, I realized this is an approach to DevOps that I can take piecemeal, that I can do myself as an individual contributor and later as middle management, and I could affect change without having to necessarily reorganize how people sit and reporting structures. So, that's really where the genesis of the book came from, was a manual to folks like me who want the DevOps change, but can't necessarily get senior leadership on board right away.

Corey: Whenever I hear anti-patterns, I get this joyful feeling in my cold, dead, salt-packed heart, because I hear that as business-speak for you're doing it wrong. And I always found that to be a method of storytelling that resonates. If I want to get people excited about a technology or aspects of a technology for me, I find that as soon as you start saying, this is how you don't do it, people sit up and pay attention for something like that. When you say this is how you do a thing, oh, it's another tutorial and people go back to staring at their phones. There's something about the, “Don't do this, what I'm about to show you,” that draws people in. Maybe that's just me and my ridiculous excuse for a personality, but I don't think I'm entirely alone.

Jeff: I don't think you're alone at all. And one thing that we've seen during the early access portion of the book is that it is exactly what resonates with people. And I think part of it is, as a reader, the author is building some credibility with you because as individual contributors, a lot of times, on the front lines, we recognize the problems. We know what's going on, and we may not be able to articulate to management or make it a priority, but we see this stuff. So, to read in the book, and see this pattern that you know that your higher-ups think is working and you know that it's not, it sort of creates his connection and you're like, “Yes, I'm bought in. He understands that problem.” Now we can get into how you go about solving it. And that was some of the early feedback that we got in the book was that the examples that were given, people were like, “This resonates with me. This is going on in my company right now. I feel these pains.” And it's a way to build credibility with the reader. So, I don't think you're alone at all.

Corey: I think the entire idea of approaching things through a story of, “This is how it looks like when it goes wrong,” absolutely builds credibility. It builds authenticity. Whenever someone gets on stage and talks about a transformation or a technology solution, and it's all a glowing story about how it solves everything and now we have world peace, I know I’m listening to a sales pitch. The really bad versions of that are where the person who's evangelizing the technology doesn't seem to realize they're giving a sales pitch. But that doesn't change the fact that it's still challenging to listen to and sit through. So, let's start at the beginning, I guess. Do you have an example of an anti-pattern that you're highlighting in the book, sort of a sneak peek, as it were?

Jeff: Sure, I guess one of the huge anti-patterns that still exists in a lot of organizations is change management. Like, change management can sometimes be the biggest farce on the planet, because it quickly devolves into a rubber stamp committee with a group of people that are asking the same people that have actually requested the change to happen if it's okay to do the change. It ends up slowing everything down, and it doesn't really manage risk the way that the process was originally intended to do. So, when you look at it, and you step back, you say, “Well, what are we really doing here, and what value is it adding? So, looking at that pattern, we can take it and break it down and say, “Well, what is it exactly that we're trying to do?” Most times, we're talking about notification to the appropriate parties, right? So, there are tons of automated ways that we can do that. Sometimes it's ensuring that a process has an appropriate rollback process. We can do that through automation. There are mechanisms within the ITSM framework for change management to have these changes become what we call standard changes, where you don't necessarily have to go through the full change management process. But, you know, a lot of organizations that I talk to don't really take advantage of that. So, what happens? You end up creating an emergency, so that you can push this change through faster because again, you're not getting the actual value out of the process. It's really just a rubber stamp that you have to jump through. So, in the book, we talk about how can we leverage automation to make that portion of change management just go away? How do we take the things that an approver would be looking for, and how do we turn those into qualitative signals that an automated process can read and decide, “Yes, this is okay to continue,” or, “No, it's not.” And it sounds scary at first. People tend to freak out like, “Oh, there's no way you could automate this.” Believe it or not, nine times out of ten, when you're taking a particular action, especially with a simple or complex task, there is a playbook that your mind is going through. And while you think you're adding a lot of nuance to it, a lot of times you're not. You're saying if A then do B, if C, then do D. And those are all things that you can codify and put into an algorithm. So, how do we do that, free ourselves up to make those processes go faster, skip that change management process, and then how do we leave change management for the truly monumental things that need to go through that process? And then how do we do it in a way so that if you're in a larger organization that has very strict change management processes, how do you fit that within the ITSM framework, so that you are still meeting all of those requirements? And it's really not that difficult, it just takes a little planning and a little forethought. So, that's one of the big anti-patterns that I see that I think resonates with a lot of people when they read the book, especially in large organizations.

Corey: Oh, it absolutely does. And even hearing you say that I immediately come up with a bunch of objections and potential challenges to that. Can we go into them for a little bit?

Jeff: Oh, yeah. Absolutely.

Corey: Excellent, let's fight. One of the joys of dealing with a lot of different companies as a consultant is that I get to see an awful lot of different cultures, and culture comes from the top. I don't think you can ever change the culture as an outsider, even though there are a bunch of consultants who swear they can do it. Good luck. Have fun. Send me a postcard if you ever get there. Very often as companies go from being small, scrappy startups to being larger entities that have actual assets and things that they could be risking, the idea goes away from build fast, move fast and break things, and into the idea of let's not break the thing that makes the money. So, it seems that, left unattended, these companies often—every time that there's something that breaks, it causes an outage, they build a check, or a system, or a process to prevent that specific thing from ever happening again, which makes sense on its face. But in time, suddenly you have so many different things that have to happen for anything to get done, that it ossifies the entire organization, and they're incapable of moving rapidly. Now, I love the idea of what you're talking about, of being able to move a lot of the process and approval into automated systems. The challenges is, as you're rolling out an automated system like that, I can imagine a world in which something gets through that doesn't get caught—not that it would have gotten caught by human review, either, for that matter—but at that point, it seems like it turns into a blame festival of, “Well, why didn't this system catch this? We're canceling the project and moving back to the human approval method.” How do you overcome that?

Jeff: So, a lot of times, in my experience, I've overcome a lot of that with data in some respects. So, one of the things that I've done in the past is like, well, let's point out this particular issue wouldn't have been caught by our other process either. That's not necessarily to give it a pass, right? You're not going to walk in the office like, “Well, the old process wouldn't have worked either.” But you have to continue to make the point that the old way wasn't delivering the value either. But, the new way is at least giving us the advantage of speed in a lot of scenarios, of flexibility. There's always going to be incidents that slip by, whether it's going to be an automated process or a manual process. What you can do in those respects and with regards to automation, is you can tier the level of automation that you're doing for a given process. And what I mean by that is, we always think of automation as being this single automated process.

I click a button, and we go from beginning to end and no one has to do anything. Well, maybe there's a process that's a little riskier. Maybe the automation around that requires certain levels of confirmation, whereas I move through the process with various steps, I'm confirming that information with a human. Is it a bit slower than the whiz-bang-y way? Yes, absolutely, but you still might take a 60-minute process and boil it down to 10 or 15 minutes and you're allowing that opportunity for a human to verify at these critical juncture points. The other option is what I always say is fail fast and fail often, where if you have this sort of happy path, if anything looks even remotely suspicious you bail out and you alert a human. So, I think the way to approach it—and without a very specific example, it's hard to get into the details—but I think you can slice a problem up based on its risk and determine the level of automation that you're going to give a thing.

If there's a human involved, where they're sort of prompting along the way, that's a lot more palatable to some people. And maybe you do that just to build confidence in the automation. Like, “Hey, we've been doing this for three months, and Bob's index finger is getting tired from hitting yes all the time, and we haven't had an issue. What can we do about rolling that safeguard back? Or, like I said, sticking with data and saying, “Hey, we've done this change 95 times now. 95 times it's worked, maybe it's time to take the guard rails off.” But I think building confidence through that is really the only way to convince people like, hey, I know that we were gung-ho about the old human method, but it wasn't great. This process is getting stronger and stronger, every time we do it.”

Corey: I find that very often you wind up—without having a transformation like you're discussing, you wind up with a very oppositional approach, and the cycle repeats and reinforces itself, where you'll have a company that says, “Okay, any deployments to production now need sign off by a VP.” And the way to overturn that is you have engineers who then decide, “All right, every change I make, I'm going to call the VP at two in the morning to approve this change,” and with the malicious compliance approach of, “All right, I'm going to effectively make sure that person never has a minute's peace until this policy gets rolled back.” And it's effective, but I don't know that it necessarily sparks the kind of transformation that most companies tend to find themselves aspiring to.

Jeff: Yeah, I mean, I absolutely agree. I think the scorched earth approach is, is probably not great, but I hope you didn't think that's what I was advocating for—

Corey: Oh, no, not at all. That's what I'm advocating.

Jeff: Oh, you're advocating for that. Oh—

Corey: Oh, I am the source of most of these terrible anti-patterns. I want to be very clear on that right now.

Jeff: Yeah. So, I don't know that—I mean, I've never tried it. I don't know that that method is particularly effective, especially at the VP level, but who knows. That's the approach that we're basically taking with alert fatigue and on-call, right? You give it to the dev because the dev is the person that's capable of making that correction. And the theory being if they're getting woken up in the middle of the night, they're going to jump on in the next morning and solve it. So, who knows, maybe—

Corey: Pain is instructive. If you have people who are experiencing pain, sharing that pain with the people empowered to fix it is not a half-bad strategy, snark and cynicism aside.

Jeff: Absolutely. I mean, there's this whole concept of skin in the game, where we transfer a lot of risk to other areas, and as long as that risk doesn't reside on me, the way that I'm going to interact with the process is very different than if I've got skin in the game. So, you're right, sort of spreading that pain around does bring a certain level of forbearance to the problem. I just don't know how effective it gets when you go up the chain.

Corey: Right? And that's where it starts to break down. To that end, what you're describing, at least in my mind, I envision a certain type of company or organization where these types of problems start to resonate. But I'd rather ask in your case, are you envisioning this being aimed at a particular type of organization? Is this something that you think applies universally? I mean, who is the target audience for something like this?

Jeff: So, the target audience for me is these small- to medium-sized businesses. I think there are some lessons to be learned from larger organizations. But you know, to be perfectly honest, these extremely large organizations have all types of other issues that they need to deal with; compliance issues, regulatory issues, legacy cruft in terms of policy, so I think some of the ideas will resonate. I don't know how effective some of the methods of implementation might be. So, where I'm going is really the small- to medium-sized companies, particularly medium-sized companies that have a bit of organizational cruft. They might have some old school processes in terms of how they're dealing with things like change management, like deployments, like access to production. Every company that I've been with seems to go through this growth phase where they're small, and they're scrappy, then they get a little bit bigger, and they start hiring a different type of person, different types of engineers and leaders, and then suddenly, they're overburdened with process, but they really don't need it for the stage of the game that they're at, but they find themselves trapped there. So, I've come into a couple of organizations in my lifetime that have been in that state where it's like, “Hey, we are doing this because Frank, who was hired last year as the VP of something came from a finance background, with a company with 250,000 employees, and this is what he's used to. But we can really roll some of these back.

Corey: That's incredibly helpful to hear. One of the biggest complaints that I have, and I suspect I may not be entirely alone on this one, is you'll go to a conference and you'll have someone getting on stage and talking about their incredible journey, their wonderful transformation, or even their technology, it doesn't really matter what the subject is, but you see them get on stage and you know, before they even open their mouths, that they're about to tell you a story that has questionable relation to what actually happened. And this has happened to me repeatedly, where I'll be in an audience sitting there watching a talk, and I'll—say—turn to the person next to me and say, “Wow, I'd love to work in a place that did that.” And their response is, “Yeah, me too.” And they work at the company that is presenting. So, you're always getting glimpses and sanitized stories, and condensed summaries that skip over the painful real-world messy stuff. The world is messy. And things like architecture diagrams or conference-wear don't tend to embrace or reflect that messiness very often.

Jeff: Oh, I so agree with you.

Corey: So, it sounds like what your book is aiming at is talking about the messy parts.

Jeff: Absolutely. That's one of the things that I think is a common theme in a lot of my talks when I go to conferences, is I throw the dirt out there, because exactly what you said, no one wants to hear the rosy red picture of how perfect everything was now that you're running Docker and Kubernetes, and there was no pain. It's like—you're right. It's messy. It's dirty. And the thing is, it's not always a complete win. So, I did a talk, “DevOps, the Good, the Bad, and the Ugly,” where I sort of detailed, “Hey, we went through a bunch of different organizational models. And as great as it sounds to have ops embedded in dev teams, there are some things that suck about that.” Mainly, you're scaling your dev team much faster than you're scaling your ops team, or this ops guy is sitting around in sprint planning twiddling his thumbs because half the stuff you guys are talking about, he doesn't really care about, doesn't have a direct interest or need to work on. So, I think it's important that when we're giving these talks, we try to be as truthful and honest as possible. And I know that some companies might take umbrage with that. Like, “Oh, we don't want you out there airing our dirty laundry.” But that's where the learning happens. No one learns from a perfect story. Where I learn from is when you get up and tell me how your Kubernetes cluster fell over because you only had one DNS pod and it got overran and it took the entire cluster down. Those are the stories that I'm interested in. If you're going to sit there and tell me oh, yeah, we migrated to Kubernetes, and everything was great. It was a little hard, but now we're saving all this money. That's kind of hogwash, useless, and you're right, I sort of turn my nose up the minute I hear it, like, “Oh, boy, here we go.”

Corey: Yep. Oh, here we go again. It's one of those very narrow talks. It sets off my sense of this is a sales pitch, and the only question left in my mind is, do they know as a sales pitch, or are they falling for their own mythology?

Jeff: Yeah, and it's great when you hear a speaker—when I was at re:Invent 2019, I was listening to the Pokemon team talk about their migration, and they were talking about these very common bad architecture choices that they had in their organization. And while everyone groans, everyone is also like, yeah, I got three of those in my organization as well. And to have that sort of realness, I was like, “Alright, I'm bought in.” I want to hear where these guys are going because they were being authentic, versus, “We don't really make any mistakes and architecture is perfect, and everything's beautiful, in Pokemon World.”

Corey: Exactly. You can't go ahead and just assume, in the absence of anything else, first, that what someone is saying on stage is true. But secondly, taking it a step further from the person who's on stage, I feel like a lot of times these talks—and I know I'm picking on conferences, but that's okay. They deserve it. They're on stage, that's the point—it feels like they don't set ground rules and context for who the talk is for. These—the canonical example that built a talk that I wound up giving was, I watched someone on stage talk about how Netflix does things and how everyone has access to root and production, yadda, yadda. The person next to me is super excited about that, and I glanced at their badge and they work for a bank.

Jeff: Right.

Corey: There's a difference between production for the infrastructure streaming thing, and production is in the thing that runs the ATM network. You kind of need to have those things treated differently, but the lack of context-setting in the talks that the audience can internalize is often missing. So, I'm a stickler for, who is this for? And sometimes when you look at some of the more poorly executed marketing events, you realize it's not really for anyone.

Jeff: Right, Corey, if I could wave a magic wand, I think I would ban Google and Netflix and Apple from writing any blog posts.

Corey: And the internet. Or is that a bridge too far.

Jeff: Right, so, it's like just stop writing blog posts, because people read this stuff, and they think that everyone is doing that. And it’s a large source of where all this imposter syndrome comes from. You read these blog posts about how super perfect their architecture is and you're like, “Well, man, I'm trying to just stop having to restart my web servers every eight hours because they're running out of memory.” So, you read this stuff, and you internalize it, and you feel like, “I’m just not adequate.” When you get into those shops, though, when you talk to those people, you realize it's the same dumpster fire, it's the same set of bodies. They just polish them up and keep them buried where the smell doesn't waft up too high.

Corey: I hate when people leave a talk feeling crappy. It’s, “Oh, you built a globe-spanning CDN that talks to everything in the world simultaneously. Cool. Cool. Yesterday, I spent three hours hunting down a misplaced underscore.”

Jeff: Right, right. Exactly, exactly.

Corey: The idea is, people should never leave a talk feeling crappy.

Jeff: Yes, absolutely. People should feel energized, they should feel charged, they should feel like, “Wow, I'm not alone, and I'm capable of doing this.” And I'm sort of hoping that's what the book does for folks. When they read it, they say, “Wow, this is my organization sort of reflected in this. And this guy found a way out, and it's not perfect by any means. But it's a lot better than where I'm at. Maybe I can get there too.” And that's how you should feel leaving a talk. Like, “Okay these guys have the same problems. They have the same set of skills. It's not a bunch of rocket scientists that they recruited from NASA to somehow run a Kubernetes infrastructure. They're just regular tech people that are doing a job that are making mistakes and they're sharing their learnings.”

Corey: Yeah, and if you're up there telling the story, then during the Q&A section. It's a well, “How do you handle this problem?” And the response is, “I can't talk about that.” Then why are you even having Q&A? If you're not allowed to talk with anything other than it was expressly approved, get off the stage.

Jeff: Right? Exactly.

Corey: Or just don't set it up that—it stinks of, that's only for special people to know. When I hear that, I just assume, “Oh, you've done something Byzantine that wouldn't work anywhere else, and, frankly, is probably more than a little terrifying.” I've talked to people at all of these big tech companies, most of them on this podcast, and something I've learned having done this long enough, is that there is no mythical magic company where everything is done correctly. Everything is a dumpster fire under the hood, but most people don't talk about it.

Jeff: Absolutely. You know, one of the best moments I had at a conference was I gave this talk and people were pretty impressed. They were giving me all types of credit that I don't deserve, and then someone says well, how do you go about handling your infrastructure testing? How do go about automating that? And my response was simple. I was like, “Poorly, we do it poorly. We're not good at that. And we're working to get better. Here's what we do. It's not the greatest. You're probably in a very similar situation.” But to see people's heads nod and say, “Okay, alright. Sweet. So, I'm not alone.” You know, that's just, I don't know, it's a good feeling. And the thing I always try to iterate to people is like, just keep getting better. That's all you can do. Just keep getting better. And if you're getting better, you're in a good spot. As long as you're not sitting around and saying, “Well, this is just the way life is. And I guess we're stuck here.” That's a place of defeatism. But when you're saying, “Hey, this is wrong. We're trying to make it better, and we're making progress.” That's the place I want to be. I told my boss—my boss said, “Jeff, what would make you leave?” And I said, “Well, if we were ever in a situation where we were doing something really stupid, and we were just completely okay with it.” I'm okay if we're doing stupid, and we're saying, “Well, we know this is stupid. We got to make it better,” but if we're like, “This isn't that dumb, and we're okay with this,” that's when I'm like, “Alright, it's time to look for another job.” So, I just feel like you just got to keep making progress and just recognize when you're doing something dumb and try to make it better.

Corey: Right. And it's disturbing how often people tend to fail that basic test. This sounds like the type of book that I want to shove down people's throats before they're allowed to give conference talks about transformation.

Jeff: Well, I will gladly do that, especially if you're going to buy a copy before you shove it down their throat.

Corey: Oh, in fact, we’ll do more than that. We'll give people a discount on the early copy of the book. You can check it out at manning.com in their Manning early access program with the discount code of PODSCREAM20 and get a few bucks off.

Jeff: Because if you're going to shove a book down someone's throat, you should at least get 40% off.

Corey: Exactly. That's the entire point. The reason there are retail prices is so people can feel good about getting discounts on things. So, let's talk a little bit about that writing a book process.

Jeff: Oh gosh.

Corey: I've never written a book. If you know what my personality type is that makes sense. I don't have the attention span, most days, to finish writing a tweet. So, that's the exact opposite of what I tend to do. What inspired you to do this? Was it something you had floating around for a while? Was it something you'd half-written and then you pitched it to someone? Did they come talk to you out of nowhere and, “Hey, you should write a book,” and then you wound up stumbling into that, because you didn't know what you're getting into? How did it happen?

Jeff: So, Manning had reached out to me to write a book on Puppet A number of years ago, and at the time, I just had a kid. And I was like, “Look, I'm just not in a place where I can write a book right now.” So, we sort of kept in touch. And then I wrote a blog post for CIO magazine, called “The First DevOps Hurdle”, and just sort of talking about some of the early phases of a DevOps transformation and one of the publishers reached out and said, “Hey, would you be interested in writing a DevOps book?” And I'd been thinking about this—I’d been sort of noodling on this for a while, and the timing was right, so I was like, “Yeah, sure. I'll gladly do this. Let's do it.” And then it started and I'm like, “Oh, man. I really didn't think this through. This is kind of tough. You're working 8-, 10-hour days, you're coming home, you're playing with your kids, and then you got to put two to three hours into writing a book? Like it's been a slog, it's not easy. And then you've got all of this, like I said, imposter syndrome sort of floating around, you're reading all these tweets, and other people are producing books, and you're like, “Oh, man, did I include this? Should I include that?” There's a pretty heavy cadence that Manning puts you on to make sure that you're delivering content pretty regularly. And then, when you go through what they call the first review process, where you take the first four chapters of the book, and you send it to some internal reviewers. Oh, man, you want to talk about just biting your nails, talking about Kermit nervous, you're just like, “Oh, goodness, oh, goodness, oh, goodness, what's the feedback going to be?” Because you're always going to focus on the negative feedback. Doesn't matter what it is, doesn't matter how much positive feedback you get. The negative feedback is always the stuff that jumps out at you. So, it's been a bit of a process but it's been fun in some cases, not so fun in others, but it's definitely an undertaking.

Corey: Yeah, the line I've heard is, “No one wants to write a book. Everyone wants to have written a book.”

Jeff: Yes, exactly, exactly. One of the funny stories, we—titling a book as an extremely painful process, I've learned. So, when we first started, we were doing this—the title was Demystifying DevOps. And the publisher was a little lukewarm on it, but I really liked it. And a mutual friend of ours, Emily Freeman, she wrote a DevOps book, and as soon as it came out, I was like, “I got to get this. I'm just so excited to read this.” So, I get it, and her very first chapter is titled “Demystifying DevOps.” So, I'm like, “Well, I'm going to put this on the shelf. Not going to read this until I finish my book.” And then we went through the process of retitling again.

Corey: You don't want to accidentally rip people off. That never goes well.

Jeff: Right, right. So, I was like, “All right, so I'll have to read that when I'm finished with mine.”

Corey: One last question before we go here. You are one of the organizers behind BITCon.

Jeff: Yeah. So, I'm the Chicago chapter president of a group called Blacks in Technology, and we've got chapters all over the US, but the national body runs a conference every year, and for the last two or three years, it's been in Minnesota. We did an event here in Chicago with the CIO of Illinois, and he pulled me aside, he was like, “Jeff, I was at BITCon last year. What do I got to do to get this in Chicago?” And I was like, “I think you just did it.” So—

Corey: Oops.

Jeff:—the powers-that-be started talking about it, and Illinois has been great in offering us a lot of support. So, BITCon will be coming to Chicago, we’re still working out the exact details as we finalize the actual location, but if you go to blacksintechnology.net, feel free to keep an eye out there. We'll be posting updates as they come along.

Corey: So, I have to ask one more question, though, then, I guess. So, you're organizing a conference and you're writing a book? Do you ever make good decisions?

Jeff: No, actually, I don’t. I think—when was the last time—2015. I think it was a Tuesday, I was going to get Taco Bell and instead I got Subway. And I think that's the closest I've gotten to a good decision probably in the last three or four years.

Corey: Excellent. It’s just both of those things are such strong demands on the time and it's invisible, at least until the event happens. But so many people think that a conference just magically springs up the day before. It's talking to people at big conference events, they have a ground team in place weeks or months beforehand for some things. It is such a tremendous amount of effort to pull something off that looks effortless.

Jeff: Yeah, this will be my first time organizing. Luckily, I've got some friends that have done this a few times, so, I'll be leaning on them. But you're right, here it is—I mean, we started talking about this in December of last year, what it was going to take getting the ground crew together to be able to hit the ground running as things heat up, and it feels like the work just comes in waves. So, yeah, anytime you go to a conference, you just have no idea how much work went behind that, even for the crummy ones, right? Even ones where you were like, “Wow, this is an organizational disaster.” Let me tell you, there was a lot of sweat and tears behind that organizational disaster. So, yeah looking forward to a kind of sort of, you know…

Corey: Again, similar to the book thing, it feels like it's going to be better to have done it than to be doing it.

Jeff: Yes, absolutely. And that's what everyone says. They're like, “But when it's over, man, you're going to feel so good.” And you know, just sitting around waiting for those royalty checks that are never going to show up, it's going to be exciting. So, I'm just looking forward to being able to have something that my kids can look at and say, “Oh, my dad did that,” because being able to show something to your kids and be like, “Hey, listen, you can do this. You can be a part of this world.” Where I grew up, like, I grew up in Upstate New York in a town called Schenectady. And it felt like the world was out there. It wasn't in Schenectady, everything happened out there. So, I've always thought it is so great to be able to show my kids you can participate in this larger world. It's not beyond you. You can do it. So, I think that's probably one of the biggest drivers that's been going for me is the kids.

Corey: Kids are one of those things that just changes every, I guess, priority calculation you have. Everything that I wound up doing is somehow influenced by having a kid even though it doesn't look like that. It completely reshapes your entire worldview, I think it's the best way to frame that.

Jeff: Dude, you have no idea. I didn't understand it until—so, I've got this thing about priorities, and I don't believe in multiple priorities. I believe you have one priority. You have a bunch of other important stuff, but there's one priority at any one given moment. And I don't think I really thought or learned about that until I had a kid because when something happens with them, everything else takes a backseat. That's the priority, and you realize like, I'm not juggling multiple priorities. This thing is a priority. Another thing is on the back burner. Now, once I handle that I'll elevate something else, but kids really did help me look at the world in a completely different light.

Corey: Absolutely.

Jeff: And they make me drink.

Corey: Yeah. Oh god, yes. I'm looking forward to the point where she's old enough to drive me to drink, like in a car. That would be handy. Then you have your built-in DD, things’ll go super well. So, if people want to hear more about your thoughts on the world, obviously, there's the [00:35:20 book with the discount code]. But where else can they find you?

Jeff: You can find me on Twitter is probably the best place. I'm @DarkAndNerdy on Twitter. That's usually where, if I happen to get around to making a Medium post, that's where I'll do it. I've been slowing down a lot with that because every time I sit down to write a blog post, I say, “You know what, you should be writing a chapter in your book instead.” But following along there is probably the best place to catch me.

Corey: Fantastic. Thank you so much for taking the time to speak with me today. I appreciate it.

Jeff: Thanks for having me, man. Had a great time chatting.

Corey: Excellent. Jeff Smith, Director of Product Operations at Centro I'm Corey Quinn. This is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review in Apple Podcasts. If you've hated this podcast, please leave a five-star review in Apple Podcasts.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Yulan Lin

Yulan is a data nerd with experience in everything from bioinformatics to NLP. She’s currently working as a Developer Advocate for Google’s Data Studio. Prior to Google, she worked in research, event management, and government data science. When not computerating, she can be found reading, hosting dinner parties, and making music.

Links Referenced

  • DigitalOcean: https://www.digitalocean.com/
  • CHAOSSEARCH.io
  • http://do.co/screaming
  • DataStax: http://www.datastax.com
  • https://www.datastax.com/accelerate
  • Twitter https://twitter.com/y3l2n

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is brought to you by DigitalOcean, the cloud provider that makes it easy for startups to deploy and scale modern web applications with, and this is important to me, no billing surprises. With simple, predictable pricing that’s flat across 12 global data center regions and UX developers around the world love, you can control your cloud infrastructure costs and have more time for your team to focus on growing your business. See what businesses are building on DigitalOcean and get started for free at do.co/screaming. That’s D-O-Dot-C-O-slash-screaming and my thanks to DigitalOcean for their continuing support of this ridiculous podcast.

Corey: This episode has been sponsored by CHAOSSEARCH. If you have a log analytics problem, consider CHAOSSEARCH. They do sensible things like separating out the compute from the storage in your log analysis environment. You store the data in S3 in your account. You know where it lives, you know what it costs, and then they compress it heavily while indexing it, and then they query that data using a separately scalable fleet of containers. Therefore, the amount of data you’re storing no longer is bounded to how much compute you throw at it, as well. It’s broken that relationship, leading to over 80 percent cost savings in most environments, and being a sensible scaling strategy while still being able to access it through the API’s you’ve come to know and tolerate. To learn more visit CHAOSSEARCH.io.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’m joined this week by Yulan Lin, a developer advocate at a small company called Google. Yulan, welcome to the show.

Yulan: Thanks, Corey. Thanks for having me.

Corey: Of course. So, you self describe yourself as a data nerd. With experience in everything from bioinformatics to NLP? I have to look up what some of those words even mean. So, backing up a sec, what do you do and how did you get there?

Yulan: Yeah, so I’m a developer advocate for a product called Data Studio at Google, which is a business intelligence and dashboarding product that we have. And, how did I get here? Well, I studied chemistry. I actually thought I was going to go down the research route, and I had a bioinformatics research project, which is basically like computational genomics kind of stuff. And I was looking at RNA sequencing and macular degeneration. But the interesting part of that was, I had a dataset that crashed Excel, and I was like, “I don’t know what to do.” And so, what ended up happening was in the process of learning to analyze that data, one, I learned all sorts of statistical techniques that were completely new to me, but I also learned how Python scripting worked, R scripting worked, learned a little bit of SQL along the way, and I realized that I was much better at picking up those data analysis skills that were transferable than I was at keeping cells alive. And kind of went from there.

Corey: Yeah, you found your way through to Google of all places which, if you’re working with data, it seems like a decent place to go. I’ve been told they have a bit of that there. But your path sounds like it diverged from mine almost immediately. When I wound up early on in my career with data sets that had problems in Excel, I threw the computer aside and a huff, and instead of doing data analytics, I just figured I would talk qualitatively instead and tell interesting stories and indulge my ongoing love affair with the sound of my own voice. It didn’t occur to me that there might be better ways to solve these problems. They’re the path not taken as it were.

Yulan: I actually think there are really incredible stories to be told with and about datasets which is, I think, what compelled me about data analytics in the first place. So, in research, when I was looking either at chemistry education, or at bioinformatics stuff, what I always loved was that I could ask a question, get some data about it and tell a compelling story that, I think, I would argue mattered to the world. And I think the same is true for a lot of data sets. Because after Google, I ended up at a nonprofit doing event management and ops work. And so, I was doing a lot of the statistics around a large, about 15 thousand person event, and so in the process, I learned a lot about what different stakeholders wanted to get out of the stories inside our data sets about who was attending and how registration was going—

Corey: Out of all our talks, which irritated people the most. I mean, fun things like that.

Yulan: [laughing] I mean yes, but also really interesting things like you find users who take different paths through the registration system. And so, you end up with really interesting technical issues, like that the tags associated with someone’s registration don’t make any logical sense, because they found a bug that allowed them to take multiple routes through it, all these data management and data quality issues as well. And I loved going in and figuring out what didn’t look right, why it didn’t look right, was there something we could do about it that would stop that problem?

Corey: That sounds like it’s, first, incredibly difficult. And secondly, it sounds like it’s the sort of thing that no one knows exists. No one pays attention to that sort of thing at all. At a conference, it magically happens, it sprung fully formed. And sure they started setting this up yesterday evening or something. People don’t realize all of the heavy lifting and tremendous amount of work that goes into putting something like that on. I’ve talked about that previously with other folks on this show. But it never occurred to me to figure out the other end of that; of alright, with all the data that’s thrown off and stuff like that, how do you analyze that? How do you turn that into something that is usable by, you know, humans?

Yulan: Yes. And I think the same question of given the data that you have, “How do you turn it into something usable by humans?” is a question that applies across a lot of organizations in really small ways. So, everyone’s talking about big data. But I think a lot of quick wins are to be found in the spreadsheets that are on people’s local devices, or just one analyst or person is maintaining month after month and converting into some kind of presentation or doc or report. Because those are often human-curated, often a little bit messy. But if they were regularized, they were shared with the right people, you could link them to different data sets, all of a sudden, you have a wealth of information at your disposal, and context as well. And the ability to present things to different stakeholders and tell the right story. And that’s really cool to me.

Corey: One of the things that I always found somewhat challenging was the idea of when do you have a big data problem? The rule of thumb was if it’s on a thumb drive, it’s not big data. It’s little data, medium data at absolute most if you squint hard enough. And then the other argument that became, “Oh, if it fits in RAM, it’s not big data.” And then I started seeing instances in cloud and whatnot with many terabytes of RAM in there. So, is there, I guess, a clear line differentiating what separates big data from medium data? Or is it more of a -ish type of soft boundary?

Yulan: I think in general, it’s an -ish boundary. But I think the framework I use is less why do you care about how big the data is? Do you care about it for reasons of data engineering, and you want to know what the best kind of technical ways to manage and process your data analysis pipelines are? Or are you interested in what statistical techniques are valid on the data? Because the definitions of quote-unquote, “big data” differ across those things.

Corey: That doesn’t occur to me to think of it in terms of domain-specific. I mean, on some level, log data could be enormous if you log everything forever from just a simple web service. But it also winds up being awfully repetitive. Oh, wow. 98% of our data in the logs is the load balancer checking to see if the thing is still okay. Maybe there’s a transformation that makes this a little bit more usable as you start filtering that through. And again, I am not a data person at all. It turns out that stateless stuff is way more aligned with how I tend to operate because if I break that, I can push a button, build a new one and no one notices or cares. When you lose the data, very often you don’t really have the company anymore after that.

Yulan: Yeah. I think the other thing too, with longitudinal data or data over time is that definitions can change too. And so, within the same organization, even if it’s been collecting a particular piece of data forever, the original reason might have been to answer question X, and then at some point, they realized question Y might also be kind of relevant to this data set, so I’m going to add a couple other fields to capture those things, as well. And tracking that metadata and the evolution of the whys behind why a database exists or why a field exists in a table, I think really can inform the questions that are valid to ask about the data set.

Corey: I think one of the challenges with data, at least one that I experience myself is that I don’t know what questions to ask the data can effectively answer. I mean, so from that perspective, it’s always challenging to figure out what questions does data visualization solve for me.

Yulan: I think jumping to shiny visualizations before understanding the data set in the domain is actually going too quickly. At my last job, I sometimes described it as I was playing data therapist, because I talked to different people about what data sets they had and what questions they wanted to answer. And whether or not those datasets could effectively answer those questions. We also talked about what are the best ways to answer those questions? Is it some kind of analysis? Is it some kind of Visual Dashboard? And so, that’s something, I think, that has to be done in partnership with a domain expert. And also just time spent in the data, right? What’s the distribution of things, what do null values look like? What are things I should know about what different codes mean? All of these questions really should be thought through, in partnership with a domain expert who then also has a better idea what are the things that they want to track, or that would impact their day to day work?

Corey: This episode is sponsored in part by DataStax. The NoSQL event of the year is DataStax Accelerate in San Diego this May from the 11th through the 13th. I’ve given a talk previously called the myth of multi-cloud, and it’s time for me to revisit that with… a sequel! Which is funny given that it’s a NoSQL conference, but there you have it. To learn more, visit datastax.com that’s D-A-T-A-S-T-A-X.com and I hope to see you in San Diego this May.

Corey: Well, let’s back up a second here just to clarify something that I may not be entirely clear on. One of Google’s core competencies is taking words and putting them after the word Google as a product. In this case, they’ve done that with Google Data Studio. What is Google Data Studio?

Yulan: Yeah, so Data Studio is a in-browser Data Visualization BI kind of dashboarding product that connects to all sorts of data sources. So, the way we describe it is, if it has an internet-accessible API, you can probably get the data into Data Studio. So, it allows people to integrate data from different data sources into the same place so that it’s easy to have a at-a-glance look or analysis of whatever metrics you care about. And it’s also really easy to make sure that it’s shared with the right stakeholders.

Corey: So, it winds up visualizing data for human consumption, not machine consumption?

Yulan: Yes, it’s for human consumption, and it’s also structured in such a way that it’s relatively easy to get started with it because the product itself, it’s a click and drag kind of product. It’s a GUI based thing, even though I work on the developer features, which is kind of this separate box.

Corey: So, I guess my question for you then becomes as a developer advocate for something like this, what developer advocacy around data visualization look like? Who are the people you’re talking to? And what challenges do they have?

Yulan: Yeah, I think to answer the question, it might be useful to talk a little bit more about my job. So, my job is to actually support this API called Community Visualizations, which allows people to build their own custom visualizations and integrate them into Google Data Studio dashboards. And so, the reason to have a Developer Advocate around data visualization is really to show people the power of different kinds of visualizations, or different kinds of solutions, and storytelling around their datasets, and how to build them with Data Studio. And so, it’s things like, is there a chart that somebody made in an academic paper that actually would be really great for your use case, but it’s super specific and you have to have everything configured a certain way. And when does that chart work, when does it not? Is it for a particular dashboard or infographic, or is it something that’s generalizable? I think these are all questions that I’m hoping that my work helps people to answer a little bit.

Corey: It’s always difficult, I guess, from my perspective, to figure out how to structure any sort of visualization of reasonable data. It’s easy once you have a dashboard, or something that shows the relationship you’re looking at. Oh, yeah, that’s incredibly valuable and helpful. For whatever reason, I don’t know if it’s just who I am, or this is something a lot of people struggle with, but I personally have trouble figuring out even how to begin structuring what I might represent data as in a visual context. Is that common? Am I just crap at this thing and I should accept that? What is the, I guess—what are you seeing in the world as far as people’s level of comfort with this sort of thing?

Yulan: Yeah, that’s a great question. I think that it’s actually a really hard problem, and it’s deceptively hard. And the reason is because I think the right visualization or the right structure of a dashboard depends so heavily on what you want that dashboard to do. Because there’s a difference between some of the key metrics you want to have on a TV display, in your lobby or in an open office, then something that you want an analyst to be able to interact with and find trends or interesting things in, and it’s different than another dashboard that summarizes particular metrics for an executive. And so, everyone cares about different things, so I think my first question is always, what metric do you care about? Who is looking at it, and is it meant to be kind of a display kind of thing? So, a dashboard in a lounge, or a infographic kind of thing or is it meant to be something you can interact with, and a means of exploration and analysis? Because that tends to help me start deciding how complicated things should be. Should they be scorecards? Should they be pie charts or bar charts? Do I want to bring in something really complex because it actually represents something like the number of people transferring in and out of certain regions well, or the genome data well? Should I be bringing in domain-specific things like that?

Corey: That’s an area where it seems to be extraordinarily challenging to, I think, articulate to folks who aren’t steeped in areas of this. I mean, it becomes the popular question that I think a lot of us who work in anything that even remotely touches technology has to answer when we deal with folks who are not in that space, usually at holidays with family, explaining what you do for a living to people who have no touchpoints for it. Do you have a go-to that you wind up using for that?

Yulan: Yeah, I talked about the New York Times data visualization team, partly because it was their work that inspired me to care about data visualization and see how powerful it was in the first place. And because that tends to be a good point of reference, so even if people aren’t familiar with that team, if I pull up a map that they’ve created, or pull up some charts that they’ve created to go with some of their stories, people immediately understand like, “Oh, seeing this in a chart instead of a table actually makes it click in a different way or I asked different questions. And that starts the conversation around data visualization.

Corey: Let’s go down a path that I love to explore that most people often don’t, generally because it’s a terrible way to teach people things, but I find it entertaining. What are some of the most egregious misuses of data visualization that you’ve seen? Or, I guess, bad data visualization. This is an audio podcast, so showing people crappy charts is not going to be as compelling when you’re just describing a crappy chart, but have you seen anything that is horrifying?

Yulan: It’s hard to say things are actually horrifying, but I think there are some cases where there’s just lines everywhere, it’s incredibly complicated, and there’s no explanation or walkthrough of what the different icons mean, and why lines are moving in certain directions, and whether or not things were stylized, or whether every angle and motion of the line or color variants means something and what that maps to, because I think at some point of complexity, my brain personally just kind of shuts down. The other thing I found, and I’m guilty of this, too, is just making decisions that look kind of pretty but have no meaning. So, arbitrary color changes because it matches a particular palette, even though colors have absolutely no meaning, that ends up being very confusing. So, those are some of my own pet peeves. Oh, that and I also really dislike low contrast color palettes, for accessibility reasons but also just readability reasons. It’s like, “Cool, you used this very uniform palette that looks great with your branding, and I cannot tell the difference between your different categories.”

Corey: One of the things I’ve always found is, for whatever reason, and I see this periodically in various state-of-the-cloud-style reports, where they’ll have a whole bunch of different providers or services or offerings that they’ll wind up trying to visualize. And this isn’t even a data visualization issue as such, but it’s always we’re going to represent each one of these different things, five or 10 of them, in different shades of blue. Maybe there’s another color or million that you could use that would show a little bit more contrast. At some point I look at that wonder if I’d suddenly gone colorblind. No, it’s just graphic design is hard for everyone.

Yulan: Yes. Yeah, and I think there’s also the sense of making something clear and easily readable might be at odds with some kind of sleek visual identity that certain infographics or reports want to attempt to follow. So, it’s this like, do you pick readability? Do you pick your brand palette? What if they’re at odds?

Corey: And that’s always a question. I dare not tread down that path. I found that it is best not to walk down into the den of corporate communications, and branding and, oh, no, no, no, no, you wound up not quite centering that, or the font isn’t quite right, throw it away, start over. And if you do it, again, you’re being censured. I may deal with big companies too much at this point in that context. So, changing gears slightly, you are a developer advocate. What exactly does that look like in your particular scenario? Very often, I’ll find that developer advocates spend the bulk of their time arguing with other developer advocates about what developer advocacy is.

Yulan: Yeah, that’s a great question. I will say that I think a definition most of my colleagues and peers can agree on is that we want technical practitioners to be successful with our products. Ultimately, if that happens, then I feel like I have succeeded. In my particular case, I think my goal is to build an ecosystem around this API. I want people to know what’s possible with it, and I want to help people solve problems with it. And so, to understand, you know, why they should care and also have a clear path to success once they figure out, “Oh, I want to build something.” So, it involves, for me, everything from talking to developers and understanding their use cases so that I can make sure their concerns are addressed as I write the documentation, or make videos, or give talks. And that tends to be the bulk of my work is just thinking through, how do I make somebody successful who thinks, or wants to build something using this API?

Corey: Do you find that the bulk of your developer advocacy work, it looks like blog posts, like one-on-one conversations with customers or developers in the community? Are you giving conference talks? Are you writing API examples and documentation-style stuff? Or other things entirely. There’s so many different expressions of the whole [inaudible 00:21:36] world that I learn something new every time I talk to someone who does this full time.

Yulan: Yeah, so the answer to your question is yes, I do most of those things if not all of them. It is a lot less speaking than I thought it would be. So, my time, at least right now, is spent split between a couple things. One is content creation. So, making sure the documentation is there, making sure there are examples, some blog posts, some social things. I also spend some time talking to developers and companies who are developing against this API. And the other thing I do is I am writing API examples and developer tooling that makes the developer experience easier. And as I’m writing these examples, I’m also collecting my feedback and other people’s feedback about the API and bringing it to our internal teams, and saying, “Here are things I think would help the future developer experience. These are ways I think we could make it easier,” and then trying to address it either from my end, or talking to our internal teams to see, can we solve this problem for future developers?

Corey: I remember back when we first met, you had given a talk at a conference and we wound up catching up at the event afterwards and got to talking. A few speakers started gathering together, and I think you were even asking me, “How can you start doing more conference talks as a part of a career path?” And I think the default response from everyone who’s done that was, “No, don’t do that. It’s awful. It’s drudgery, and misery, and horrible.” And, as I recall you, you hadn’t gone through to that side yet and thought that it was going to be fun and amazing and worth doing. Where do you stand on public speaking now, now that you have found the job where that is part and parcel of what you do?

Yulan: There’s a couple things. One is, I still absolutely love public speaking. And I do wish I did it a little bit more because there is something—especially in a small- to medium-size audiences about being in front of people and sharing things, but also just reading and reacting to the energy of the room, and helping people understand something new, or hopefully learn about something that hadn’t heard about before or thought about before. At the same time, I think there is a sense of, maybe not the talks themselves, but travel for conferences, has become less shiny to me, even though I actually do it less as a developer advocate than I did before. And part of it might just be because it’s part of my job, it seems less shiny to do it for fun. And I think part of it, too, is I think it’s different to nerd out as Yulan being a data nerd because data is super cool, versus when I’m representing a company or a product because they’re just different considerations. And that’s not an aspect of it I had ever thought about.

Corey: Yeah, it’s always very interesting to me seeing how people’s evolution as they walked down the path of doing whatever it is that they’re involved in tends to modify itself and, I guess, express itself in different forms. It’s strange when you think of someone who’s has a background in data engineering, at the path that you talked about going down, the idea of, oh, and then pivoting to becoming a speaker and someone who helps other people understand these things, it’s always interesting seeing the different routes people take to get there. I mean, you see people who look an awful lot alike on stage sometimes, but the paths they took to get there are incredibly varied.

Yulan: Yeah. And it’s also, I think, that the same skills and things that people enjoy can be expressed in so many different ways throughout a career or throughout a job. And so, when I was not a developer advocate, I still loved helping to organize and speak at meetups because I just loved seeing people learn more and seeing knowledge sharing within the community. And also, I was really excited to just show people things I thought were cool. I’m still very excited to show people things I think are cool. There’s a part of it to where, when it’s my actual job, I’d have to think not only about like, how do I tell people about something I think is cool is this question of how am I good at telling people about that thing? And kind of content creation and technical communication is its own set of skills in addition to kind of having the technical knowledge of whatever it is I’m trying to communicate.

Corey: Do you have any advice for people who are looking to get started with data visualization to where they can go to learn more? How can people dip a toe in this water if it’s something that they’re unfamiliar with and want to learn more?

Yulan: Oh, so many places. One is I would just start looking for the places that are building charts that you respect and trying to figure out what you like about them. So, there’s several news outlets that are really good about that. And there’s also some independent data visualization experts, designers that I really respect, people like Shirley Wu and Nadieh Bremer—I hope I’m pronouncing her name right—and so, that’s one piece, to learn about the design aspect. And then technically, I would say, get started with either Python or JavaScript, and just get into a data set and figure out, how do I put something on a page? How do I make a chart? How do I explore this data? Don’t worry about the bells and whistles, that takes time, and that will come, but just trying to figure out what is that conversion from data to pixels on a page look like? And oh, also draw data visualizations on graph paper, because it’s a fantastic way to get an intuition for what you’re trying to map out and why.

Corey: That’s a great starting point that I think people will appreciate. If people want to learn more about what you’re doing and the various things that you have to say about a variety of topics, where can they find you?

Yulan: They can find me usually on Twitter, I’m @y3l2n, and I talk about all sorts of things from women-in-tech rants, to data visualization, to posting videos I make for work, so that’s a great place to look.

Corey: We will definitely throw a link to that in the [show notes 00:28:03]. Yulan, thank you so much for taking the time to speak with me today. I appreciate it.

Yulan: Yeah. Thanks for having me, this has been great.

Corey: Yulan Lin, developer advocate at Google, specifically on Google Data Studio. I’m Corey Quinn. This is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star rating in Apple Podcasts. If you’ve hated this podcast, please leave a five-star rating on Apple Podcasts and tell me what my problem is.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Hiro NishimuraM.Ed. in Special Education from University of Maryland. Five years experience working as an IT Engineer in New York City. Now, a “Freepreneur” (Freelance Entrepreneur) in the DC suburbs.

Technical Course Instructor at LinkedIn Learning. Founder of AWS Newbies and Cloud Newbies. Founder and CEO of 24 Villages, LLC. – a Writing and Consulting company.

Links

  • DigitalOcean: https://www.digitalocean.com/
  • http://do.co/screaming
  • http://CHAOSSEARCH.io
  • AWSNewbies.com
  • Twitter: https://twitter.com/hirokonishimura
  • Intro to AWS: IntroToAWS.com
  • Cloud Newbies: CloudNewbies.com
  • Screaming in the Cloud: ScreamingintheCloud.com

Transcript
Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is brought to you by DigitalOcean, the cloud provider that makes it easy for startups to deploy and scale modern web applications with, and this is important to me, no billing surprises. With simple, predictable pricing that’s flat across 12 global data center regions and UX developers around the world love, you can control your cloud infrastructure costs and have more time for your team to focus on growing your business. See what businesses are building on DigitalOcean and get started for free at do.co/screaming. That’s D-O-Dot-C-O-slash-screaming and my thanks to DigitalOcean for their continuing support of this ridiculous podcast.

Corey: This episode has been sponsored by CHAOSSEARCH. If you have a log analytics problem, consider CHAOSSEARCH. They do sensible things like separating out the compute from the storage in your log analysis environment. You store the data in S3 in your account. You know where it lives, you know what it costs, and then they compress it heavily while indexing it, and then they query that data using a separately scalable fleet of containers. Therefore, the amount of data you’re storing no longer is bounded to how much compute you throw at it, as well. It’s broken that relationship, leading to over 80 percent cost savings in most environments, and being a sensible scaling strategy while still being able to access it through the API’s you’ve come to know and tolerate. To learn more visit CHAOSSEARCH.io.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’m joined this week by Hiro Nishimura, founder, and technical writer at AWSNewbies.com. Hiro, welcome to the show.

Hiro: Thanks for having me, Corey.

Corey: So, thanks for taking the time to speak with me. So, let’s start at the very beginning. What is AWS Newbies? Where did it come from, and what’s the origin story?

Hiro: Yeah, so the origin story of AWSnewbies.com is that I was an AWS newbie, and I had absolutely no idea what this whole AWS thing was. And this was a year and a half ago, actually almost two years ago. And I needed to pick a direction in my career in IT, and I needed to see what I wanted to do, because I was working as a help desk engineer, and I decided, “Hey, this cloud computing thing sounds cool. Let’s check it out.” And I really tried pretty hard to check it out, and I could not understand any of it. And I had told my manager, I’m going to take the cloud practitioner certification exam, and I signed up for it. And it had gotten to the point where I am two weeks out before the exam, and I still had absolutely no idea what any of it was, what the services were, and why all these names were so confusing, and why one of them just sounds like a street name.

Corey: Yeah, it is not at all obvious. I mean, I remember back when I was getting started with AWS, it was overwhelming. There were too many services in the console to keep track of. I didn’t know where to begin, and there were 12 of them. Now that problem is an order of magnitude worse.

Hiro: Yeah, this was in I think, 2018. So, I think since then, it’s even blown up bigger than it was when I initially was really confused about it. Now they even have satellites going on. But, I needed to pass a certification exam. I still didn’t know anything about those EC2 instances, and I have a background in special education. So, my degrees are in special education; I just made a pivot into IT as my career. And I decided hey, the best way for me, personally, to learn is to teach and regurgitate the content back in my own words. So, I started AWSNewbies.com as a study blog for myself to pass the certification exam, and I just regurgitated all the information that I needed to know to pass the exam. And it took me, I think, around nine days to get all that information out. And then I studied my own study notes for a couple days and then took the certification exam. I passed it, and I was like, “Okay, I guess I’ll just leave it up.” You know, I had it on AWS at that point at the free tier. So, I was like, “I’ll leave it up for a year until they start charging me, and I’ll take it down. And if one or two people find it useful, that’s great.” And turns out more than one or two people found it useful. A few months later, I think, I was contacted by a content manager at LinkedIn Learning, asking if I’d be interested in creating introduction courses to AWS for non-engineers. And that just started this whole entire, I guess, new pivot in my career going like, “Wait, people are interested and wants someone who has absolutely no idea what they’re talking about has to say?” And it’s been a fresh breath, I guess, of, “Oh, hey, this is actually something that’s needed.” If I, as someone who’s pretty good at googling can’t find the answers, it means that it’s a niche that needs to be filled. So, I started—

Corey: There’s—absolutely it’s a wonderful place to start targeting. There’s always this assumption based into anything that we’ve got in any field of technology that things are super hard to learn, and then once you’re able to learn something, well, it must have been easy and everyone obviously knows it. So, it tends to lend itself as a contributing factor to a lot of the toxic oh, just read the manual, everyone knows this part. We’ll talk about the hard stuff over here. All of this is new to someone and we all have areas around which we are completely ignorant. It’s important to be able to be accessible so that other people who don’t have the same background can embrace that stuff.

Hiro: Exactly. One thing I noticed, I finally was able to deconstruct in the past, maybe a month or two, is when I was trying to study for these certification exams and figure out what even the heck this cloud computing thing was. I was taking courses and reading manuals that said, “Hey, this is for complete beginners. We’re teaching like you’re five,” you know, “This is completely fundamental beginner newbie-friendly stuff.” And I’m sitting there going, “I don’t get any of this. Am I not even at an intellectual level of a five-year-old?” And it was like, okay, one demoralizing, but two there’s something going on here that’s not correct. And I finally came to realize it’s because they are for cloud computing beginners, but not for IT or engineering beginners, so they still take for granted that you have a certain level of engineering and technical expertise and knowledge, which isn’t who I was and isn’t what I think a lot of my audiences is. A lot of my audience are people who have absolutely no technical backgrounds or came from very non-traditional technical backgrounds. So, they do have some technical backgrounds and work experience, but there’s a lot missing that’s, like, the fundamental building blocks. Which makes a lot of the technical knowledge bases and resources inaccessible because we don’t understand half the words that are written there.

Corey: Oh, absolutely not. It feels at some point, like, we’ve long since passed a Rubicon, where I can talk convincingly about even just service names in AWS for services that don’t exist, and not called out on it by AWS employees. And to the point—I’ve made that joke a couple of times now, and it’s gotten to a point now where whenever I start talking about real services, Amazon employees get nervous with the, “Are you messing with me?” look on their face, but no one wants to ask the question, because no one wants to jump out there and say, “Hi, I don’t get it,” because that’s scary, and there’s this attitude that this is probably a stupid, obvious thing that everyone in the world already knows, but I don’t. And that’s such a self-defeating way of looking at things. All of this stuff is complicated. None of it is straightforward, and we’re all trying to get through it as best we can.

Hiro: Exactly. And I’m very anti-technical jargon because I consider it gatekeeping to people who don’t have that vocabulary. And it’s just so prominent in this industry where someone would try to beat you down or try to say, “Hey, this is what you’re worth, because you know, this word, this word and this word.” But I don’t believe that’s the way you can really evaluate someone’s potential or knowledge. So, we try very hard to remove the technical jargon and keep our text and resources as, I guess, plain text as possible.

Corey: That’s really, I think, one of the best points to do this is it’s easy to fall into jargon and acronyms and the rest. The problem is that even internally, the acronyms have multiple explanations and expansions. EBS, we talking Elastic Block Store, or Elastic Beanstalk? We know it’s elastic, but we’re not sure what kind exactly.

Hiro: Yeah, I was wondering that too, sometimes.

Corey: I’ve never found that when I’m building a talk or putting something together that making it more accessible has served me wrong. I’ve given a few talks where it winds up being incredibly detailed and very arcane technically, and maybe four people in the room understood the nuances of what I was talking about, and the rest of the room looked at me, like I was completely out to sea. And they weren’t entirely wrong. I had completely done a swing and a miss on some of this. But that wasn’t useful for anyone. Making this much more broadly applicable and accessible to everyone was way more engaging for the audience. And, to echo something you said earlier, there’s nothing that teaches you something better than teaching it to someone else.

Hiro: Yeah, exactly. And when you have to sit back and read what you’ve written to say, and then really ask yourself, “Hey, is this word as accessible as I think it is?” Like, am I taking for granted this vast amount of knowledge that I’ve accumulated over the past 10 years of my career and using this word because it’s easy for me, and it means I don’t have to do the hard work of describing it and explaining it? I think a lot of people would really benefit from taking that few extra minutes and really going, hey, this word, is it really as, “obvious,” quote-unquote, as I think it is? Because a lot of things, I feel like we take for granted a lot of things. I guess, humans are inherently lazy, but we want to do as little as possible. And if that means we can lean on words and concepts that maybe isn’t as accessible as you think it is, I think a lot of talks and a lot of documentations would benefit a lot from that.

Corey: I think that that is something that needs to be socialized a lot more. Instead, it’s felt, for example, that, “Oh, you’re giving an intro talk, you must obviously not be quite as good with this stuff as the rest.” Nonsense. I find that even the areas in which I’m a, I guess, relative expert, although I hate the term expert. I find that by actually explaining it to someone who knows nothing about it, I learned more about that topic than I do when having discussions with someone who’s super deep in the weeds, because you don’t really know something until you have to find multiple ways of explaining the same concept. And if you don’t have a very solid grounding, you’re going to find yourself quickly stammering, and out of your depth.

Hiro: Yeah, they do say the best way to show that you actually explained something is to explain it, you know, like your grandmother or your five-year-old. If you can’t do that, you’re not actually an expert. And I think that’s very true. It’s very hard to deconstruct it to that fundamental building blocks, and really get in the nitty-gritty and explain it out to people. So…

Corey: Well, when I gave one of my talks, Terrible Ideas in Git one time, I said, “I had to still get down to the point where my mother could understand it.” And that wasn’t to speak to sexism or ageism, but rather because my mother was sitting in the front row to support me in that talk. Huge laugh from the audience and it went well and at the end of it, my response to my mother was, “Great, so did you understand what I talked about?” She said, “Not a clue. Computers aren’t my thing, but you sound like you know what you’re doing. And I’m very proud of you.” Which, awesome, terrific. Great. Then she asked me why I couldn’t have been a doctor. But that’s a whole separate argument.

Hiro: Well, I guess you didn’t really pass that test. But, um—

Corey: No, no it’s a parent’s primary job is to be disappointed in you, as far as I understand it.

Hiro: It’s very true.

Corey: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the Enterprise (not the starship) on-prem security doesn’t translate well to cloud or multi-cloud environments. And that’s not even counting IoT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IoT devices, detects these threats up to 35 percent faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at ExtraHop.com/trial.

Corey: Changing gears slightly, I know that I’ve had a number of questions about this over the past few years myself, but talking to someone else who’s done something very similar, well, let’s talk a little bit about what it’s like to effectively go independent in the world of AWS. Why did going down that path rather than joining one of the cloud education companies, or becoming a trainer through some third party service or whatnot, but instead of going down that path, what made you decide to strike out on your own?

Hiro: So, I quit my full-time job as a sysadmin at a tech startup in New York half a year ago back in June. And, surprisingly enough, the options you just mentioned of going into a training company or somewhere else full time just didn’t even cross my mind, because I felt like one of the biggest pros of me going independent and freelance was that I will finally have control over my own time and my own, I guess, direction of what I wanted to do what I thought was important and where I wanted to take my time and efforts. And, when you work in a company, it’s not up to you really what you create, or how you create it, when you create it. It’s up to the big boss, and I felt like if I’m going to take the leap and get out of corporate, then I’m going to get out of corporate and do what I thought was important. One of the most important things to me right now is that time constraint by other entities is really not a thing. And I have a choice in what my priorities are. And my priorities are to, kind of, give back to the community and create resources and areas of, I guess, networking and asking questions that didn’t exist when I was trying to get into cloud computing. So, with a corporate job, it’s not as easy to do that, and you have to, I guess, balance, the stability, and the income, and the benefits of corporate versus going freelance. But to me, the ability to direct my efforts where I think is important was really important.

Corey: That echos is my experience as well. Back when I was an employee other places, having a conversation with management was always entertaining. It’s like, “So, Corey, we have a challenge around some things you do.” “Oh, really? Like what?” “Well, basically everything you say to anyone at any time in public or private.” Like, “Oh, great, good, good, good. My personality works exactly the way you’d think it would in the context of a large company. And I always found myself being extremely unemployable and being told that the way I approach things was going to hold me back. It was a severe limiting factor. So, I finally, I snapped and finally said, “All right, either this is true, and I’m going to try it and prove it, and then I can always go by my middle name or something, once I’m completely unemployable under my real name, and then go back to a job, but now I know.” Or there’s something to it, and maybe I see something other folks don’t. And so far it seems to be working. So, I don’t pretend to be able to say I predicted this, it was just an experiment and I expected it to fail. I built success criteria and failure criteria, and figured, all right, let’s see what happens. And it turned into a consulting business, a newsletter, this podcast and another and, more or less, the Corey Quinn show as I like to talk about it sometimes, where it’s just me being out in public and doing my zany, sarcastic snarky observational thing.

Hiro: Yeah, you could totally trademark your appearances. We all know exactly when you’ve started your snarking.

Corey: Yeah, it’s one of those weird things where it’s finally—Twitter took me a long time to crack that nut, just because I was on there for seven or eight years and basically had zero traction. And then I just, sort of, found a way to make it work for me and it’s, sort of, been growing ever since. It’s weird, but it turned into a marketing vehicle where people hear about me on Twitter, and then come and talk to me about the actual serious, expensive business problem that they have that I talked about. I would not have predicted that that pipeline or funnel would have worked.

Hiro: Yeah, Twitter is honestly where I get most of my business, too, of people finding out what I do and the content and things I write, and—off my Twitter because I post basically everything I do on Twitter, and companies or individuals or small businesses will come through Twitter saying, “Hey, so would you be interested in writing this for us?” Or, “Would you be interested in making these courses for us?” And people are like, “Oh, no, what are you doing on social media all the time?” I’m like, “No, no, no, this is for business. I swear I’m not just posting another cat meme.” Which I definitely am, but I’m sure you know it’s for business, [laughing] net positive.

Corey: Oh yeah, a few years ago, I was online, looking up something, I think I was on Reddit at that point, and my boss walked past and looked at my screen and said, “Is that really the best use of your time right now?” At which point I realized I was not going to thrive in that type of environment, I needed to be able to embrace, like, various paths and go down things. And doing work doesn’t always look like typing code into an editor forty hours a week, at least I hope it doesn’t. Finding things that resonate with the way I saw the world was important to me. For a big part of it, it was—I built my company around, I guess, my own personality defects. There are things I didn’t want to do on an ongoing basis, So, I built out ways to not have to do them. I didn’t want to write production code for people because, honestly, I’m terrible at it, and I also don’t enjoy it. So, cool, advisory work only instead of writing code, that opened up the approach. One thing that I really admire about what you’ve done is you’re also in the space of selling products, whether it’s videos or courses, where you get to front-load a fair bit of that work, and then you’re done with it, And you can continue to sell that to various customers on an ongoing longtail basis. With the services delivery and consulting, I’m only as good as my last project.

Hiro: Yeah, that’s actually one of the biggest reasons why I do enjoy creating courses or ebooks, or whatever, to front load. Though it, I guess costs more in time and effort that you’re not getting paid for until that reaps benefits. But I have a couple disabilities, I have a lot of illnesses and whatnot, which require me to really be cognizant of how I use my energy and time because I’m not guaranteed to be able to work tomorrow or next week or next month. Heck, like, a whole month could go and I hadn’t written anything because of whatever health issues I’m going through. And I began side hustling and trying to figure out if there’s a more sustainable way for me to work when I was diagnosed with rheumatoid arthritis a couple of years ago, and the permanent disability rate of that is very high. And being unable to work was one of those big scary things for me, especially in this economy.

Corey: Oh, yeah.

Hiro: And so at the time I was doing full time, commuting to an office 45 minutes each way in new york. And so being able to move, being able to commute, and then stay there for eight, nine hours and do physical work, and then come home and then do everything else that life needs you to do. It was getting really, really hard. And I was thinking to myself, if I can’t sustain this, then I won’t be able to work. And if I can’t work, I’m pretty much screwed. And so I started thinking about how can I work and change the way I work so that I can continue working, even if I develop permanent physical limitations. And beyond that, what are the kinds of work that I can do that I need to get started on developing so that when something like that happens—or I need time for myself, so for example, I was taking care of a family member for a month last month, and I was gone basically the whole entire month, but no one noticed because I’m completely remote, and it’s one of those things where I didn’t really do any actual work. But the amount of income that I had was almost the same, other than the months that I had really buffed up income because I wrote something big or produce something big. It was the same because I have passive income coming in, which doesn’t matter if I’m working or not. And so, personal finance and financial independence was a really important topic for the past couple of years. And that’s one of the biggest reasons why I really like this idea of creating courses that can be done when I have enough energy and then just go for as long as it goes.

Corey: That’s, I think, a really important thing to consider too. I wound up having terrible timing, and almost everything I’ve done in business where I started this company—I left my last job and started this company two weeks before we found out we were expecting our child. Protip: don’t do that. And, something that I found the first year, I mean, it was not a good year financially, my first year in business because it turns out that someone no one has ever heard of pops up and claims to be a subject matter expert in the world of fixing bills in cloud computing. Yeah, how plausible. Yeah, that sounds like something that could happen. It took time to build traction and the rest. And for me, one of the hard parts was living in San Francisco. At any point, I could have said, “The hell with this.” Walked down the street, gotten a job at any random tech company, and made many times what I made that first year and had done a lot less work. So, there were days that that was sorely tempting. It was hard to, I guess, continue to muster the energy to go and do it and do this thing before it found any form of success. That was the hard part for me.

Hiro: Yeah, I think one of the reasons why I felt this level of okay—I mean, I don’t make that much money, but I make this, like, very base level income without basically doing work because of the passive income. And that, plus the fact that I’ve been in this space for, I guess, a year and a half now, over a year before—at the point where I quit my full-time job, I had been in this AWS Newbies space for almost a year. And even though it was only a year, it was, like, a very, very, I guess, deep year where a lot of people and a lot of, I guess, companies found out about what I do, and were interested in my work. So, when I said, “Hey, I quit.” A lot of people reached out going, “Hey, so we’ve been waiting. Let’s work together.” And that to me was really exciting because, for the first time, I wasn’t the one that had to be begging for scraps. I could be the person that gets to pick and choose the projects that I think is cool, that I want to work on. And to have that tiny security in the first half-year of going freelance, I think was extremely big in keeping me from wavering too much that, “Oh, maybe I made this really, really deep mistake and I need to go back to full-time job.” It also helps that I am not expecting a child and I just have a kitten to take care of but—

Corey: That does change things around a bit. It definitely is a priority shuffle at some point.

Hiro: Just a tiny bit.

Corey: It is also a—on the other side of things, it did light a fire, where it’s I’ve got to turn this into a success or alternately, I need to find a way to do something else in short order just because it’s, at some point, you have to understand that to folks who are not as deeply engaged in cloud computing, or the aspects of marketing of a small business, it looks an awful lot like you’re not really doing work so much as you are aggressively shitposting on Twitter all day. And it’s difficult sometimes to articulate the direct business benefit of such behaviors, so there has to be some demonstrable success tied to that, or people tend to lose interest and patience with waiting for the thing to hit.

Hiro: Yeah, definitely. A lot of people really don’t understand what I do. When I was about to quit my job, I told my coworker that I put in my notice, and they’re like, “Oh, are you going to be okay? What are you going to do?” And I said, “Oh, I’m not—” “Where are you going next?” And I said, “I’m not going anywhere next. I’m going freelance.” And they’re like, “Oh, so you’re going to do, like, Upwork or Uber.” And I’m like, “No.” But I guess that’s—

Corey: That sounds like a good way to starve to death slowly. Anytime you start doing things that involve billing by the hour or providing commoditized services or whatnot, it always turns, sooner or later, into a race to the bottom. It’s applying expertise and solving expensive problems seems to be the path to at least reasonable success for small independent consultancies, at least, or small product-focused companies. I don’t know how well that would scale up. I mean, VCs think that my entire model is ridiculous. They’ve dismissively told me, “Yeah, you can make I don’t know 10 million bucks a year on this eventually, sure, but we don’t see a $500 million exit for you.” At which point I stared at them and had to ask what planet they thought I lived on. Wow, if I don’t make half a billion dollars in my life, I will be a failure it—I guess I evaluate myself by a different rubric.

Hiro: How could you Corey? You’re not going to make half a billion dollars? What a failure. Only 10 million a year. I don’t know how you’re going to live with yourself. Your mother will be so disappointed.

Corey: Oh my god. Yeah. Oh, I’ll never be forgiven for that. It’s a great—like, a VC comes by like, “Hey, if we invested $2 million in your company, what could you do? “And I have a laundry list of things that could happen. They’re like, “Cool. What if we invested three and a half billion dollars? What could you do then?” And the only acceptable answer is something monstrous that is terrible for everyone in society. There’s no good answer for a company like this at that level of capitalization.

Hiro: No, and I honestly—scaling is something I’ve been thinking about—scaling is another one of those technical jargons, but I do like the word scaling—and scaling is definitely something I thought about because a lot of stuff like writing or producing content, it’s what I can get done with my limited resources and time. And I’m like, this is not that scalable as a business, but at the same time, I’m also wondering, does it need to really become that multi-million dollar company? It just needs to feed me, my family, my cat and produce this system where I don’t have to worry about money to pay my medical bills, and God knows my medical bills are off the charts. So, it’s one of those things where I want to balance my quality of life, versus how much I really put into this work. And also, I think when I start thinking, too, oh, how do I make this bigger, I lose sight of what I’m trying to build and start trying to build something that will get me more income perhaps, but I don’t think in the end it is what my brand is or what I want to accomplish, which is to make tech accessible for more people, and unfortunately, a lot of the people who I want to access my content, they are not that well off and that’s why they want to make a career change into tech. It’s like one of those, oh, which one do I pick the money or my, I guess, my goals or ambition.

Corey: Yeah, there’s this poisonous approach in tech where if it’s not a SaaS product with the potential to scale to millions of people with minimal intervention on the business side, that it’s not worth doing, I find that ridiculous. I don’t need to be, effectively the next Uber for cars or whatever the heck it is people are doing next. That’s not how I evaluate success. Solving an expensive problem is important and being able to reinvent myself from time to time—I’m going to be incredibly depressed if AWS doesn’t fix their billing situation before I retire. That just means that they have completely failed their customers. I want to focus on other things. I want to move on and do new and interesting things. I think you’re perfectly well positioned in that context, too, where this stuff doesn’t get easier over time and it doesn’t get less broad and people are not going to suddenly lose interest. Once you find a way to teach complex concepts to people in a way that’s accessible and approachable, that can be applied to a lot of different things. And that’s a skill that never goes out of style.

Hiro: Yeah, I definitely feel that, I really feel that personally because I like, okay, I’ve talked about AWS for the past two years now. And I’m sure I can talk about it for another couple of years, maybe, but then I’m going to want to move on. But if this whole entire, “Hey, let’s introduce Amazon Web Services, and cloud computing to people without traditional tech background is not fixed in, I guess, the five years I was trying to do something about it, I would also be a little depressed about the way tech is going because clearly it’s a problem, and I’m trying so hard to tell people that this is a problem, and if some big fix that’s not a band-aid doesn’t come up in the next couple of years, it’ll be very disappointing.

Corey: And there’s always going to be other options to pivot. Easy example is explaining at a high level what all these things mean and what the consequences are to business-level decision-makers. Those people have money, they need to be instructed and what these things are, they don’t need to go down into the how to work with it level, but they need to understand enough of what’s involved to be able to evaluate for themselves whether something is worth pursuing, and by being an external voice that doesn’t have an agenda or a vested interest in a particular outcome, there’s tremendous value in something like that. And there are countless other examples there for you to continue expanding to if it makes sense.

Hiro: Yeah, yeah, definitely. Or maybe you know, by that point, I would have pivoted to another career, because I have a very bad track record of keeping to one thing for a very long time.

Corey: You’re telling me. Before I started this thing, I’d never stayed at a company longer than two years and that one that I stayed at for two years; a consulting company where I was changing clients every few months.

Hiro: Yeah, yeah. No, I mean, my background in special ed and the longest I was at a company was for two and a half years. But, um, yeah, and then I quit. So, there’s that.

Corey: So, if people want to learn more about what you’re up to, how you view the world, etcetera, where can they find you?

Hiro: Yeah, so one place, that like you, I am always living in his Twitter. So, my Twitter username is @Hiroko, H-I-R-O-K-O, and then my last name, N as in Nancy, I-S-H-I, M as in Mary, U-R-A [ed: @Hirokonishimura]. You can find the same name if you go to twitter.com/AWSNewbies, then my username is on the profiles, if that’s easier to find.

Corey: And we’ll put a link to that in the [show notes 00:33:45] as well.

Hiro: Yeah, great. And then AWSNewbies.com, and IntroToAWS.com are where I list my resources, and my courses, my ebooks, stuff like that, everything to get you started with cloud computing and AWS, minus all the technical jargons. And I recently started a community of both seasoned pros and the cloud newbies called CloudNewbies.com and it’s, like, a Discord of people who both want to learn and also are interested in teaching or being the sounding board for beginner questions to learn about cloud computing, or AWS, or Microsoft Azure. It’s cloud-agnostic, so you can come in with any cloud question, or if you’re studying for a certification exam you can study together with other people in the community who are studying. And it was like a community was something I’ve been wanting to build since the beginning, but I just didn’t have the time or the technical headspace to execute it. And I finally did earlier this year, so I’m very excited about that, where people who are completely new can come in and be like, “Hey, I don’t get this one concept. I tried so hard to understand it, but all these technical documentations aren’t helping me.” So, that’s one project. And yeah, I mean, Twitter is where you can probably find out a lot of what I do and all the dumb things I talk about every day, and too many cats.

Corey: I just signed up for the Cloud Newbies society myself, as well, on Discord. I’ve never used Discord before, so we’ll see how that works out. But yeah, I’m always a fan of trying to find places people can chat and talk about these things. There needs to be more of them.

Hiro: Yeah, and the focus here is—I joined a lot of Slacks and Discord channels, where we talk about AWS and a lot of the problems or questions, but they tend to be a lot more engineer-y and nitty-gritty and higher level than what the level of questions that I would want to ask or what I can contribute as a member. So, I really wanted to create a place where newbies can come in and ask their questions and follow conversations that are more inclusive. So, I’m really excited to see where this goes because I would very much love to have had a community like this when I first started and tried to figure out what the heck cloud computing is.

Corey: Yeah, I think that’s going to be a fantastic way for people to like, engage with it. I’m always a fan of finding different ways of reaching people. Some folks love videos, some people love podcasts, some people love asynchronous chat, some people like real-time conference calls, it really just depends. Everyone learns differently, so having different ways to embrace them in ways that resonate with them is critically important.

Hiro: I agree. Yeah, it’s one of those things where you have to reach your audience where they are at, and a lot of times, they’re not going to come to this brand new different thing, they want they’re used to, so it’s like, “Hey, if you’d like to chat, come chat with us. If you want to read about different career opportunities in the cloud, here’s a blog. If you want a course I’ve got courses too.” It’s just like everything.

Corey: Excellent. And I will throw links to that in the [show notes 00:37:14] as well. Hiro, thank you so much for taking the time to speak with me today. I appreciate it.

Hiro: Of course. Thanks for having me.

Corey: Hiro Nishimura, founder of AWSNewbies.com. I’m Corey Quinn. This is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five star review in Apple Podcasts. If you’ve hated this podcast, please leave a five star review in Apple Podcasts and tell me what my problem is in the comments.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About xssfox

Just a dumb fox

Links Referenced:

  • DigitalOcean https://www.digitalocean.com/
  • CHAOSSEARCH http://CHAOSSEARCH.io
  • Big Buck AWS https://github.com/xssfox/bigbuckaws
  • Corey’s talk, “Terrible Ideas in Git” https://www.lastweekinaws.com/blog/terrible-ideas-in-git-by-corey-quinn/
  • Twitter: https://twitter.com/xssfox
  • Screaming in the Cloud http://ScreamingintheCloud.com

Transcript
Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey Quinn: This episode is brought to you by DigitalOcean, the cloud provider that makes it easy for startups to deploy and scale modern web applications with, and this is important to me, no billing surprises. With simple, predictable pricing that’s flat across 12 global data center regions and UX developers around the world love, you can control your cloud infrastructure costs and have more time for your team to focus on growing your business. See what businesses are building on DigitalOcean and get started for free at do.co/screaming. That’s D-O-Dot-C-O-slash-screaming and my thanks to DigitalOcean for their continuing support of this ridiculous podcast.

Corey: This episode has been sponsored by CHAOSSEARCH. If you have a log analytics problem, consider CHAOSSEARCH. They do sensible things like separating out the compute from the storage in your log analysis environment. You store the data in S3; in your account. You know where it lives, you know what it costs, and then they compress it heavily while indexing it. And then they query that data using a separately scalable fleet of containers. Therefore, the amount of data you're storing no longer is bounded to how much compute you throw at it as well. It's broken that relationship leading to over 80% cost savings in most environments and being a sensible scaling strategy while still being able to access it through the APIs you've come to know and tolerate. To learn more, visit CHAOSSEARCH.io.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’m joined this week by Michael, who is, similar to a previous guest, an Australian code terrorist. Michael, welcome to the show.

Michael: Hey, thanks for having me.

Corey: So, you wrote something a while back that really took some of the serverless world by storm. It’s a GitHub repo that you called Big Buck AWS. What does that do, exactly?

Michael: So, it’s like a tech demo, I guess, of how you can abuse some of the inner workings of AWS, specifically Lambda. And it allows you to stream content, so like a mp4 HLS stream, without really paying for much data. And it does that by abusing the ability that you can pull your code back out of Lambda. Yeah, so, the thing is, when you upload code into Lambda, they also let you download it again, to view it in the online editor, or if you have some internal tool, you can use it. But the way it works internally in AWS is when you upload it, it ends up in an S3 bucket. You can grab that data back out from the S3 bucket through a signed URL, but that bucket is Amazon’s bucket, not yours, so you don’t get charged for it. So, that’s kind of a key part of it.

Corey: And that always struck me as something very strange. I’m sure there’s a technical reason behind it, but you get 75 gigs of Lambda storage per account, it’s an Amazon bucket, there’s no requester pays or anything like that, and it just never occurred to me that you could pull data back out of it, mostly because, I guess, I lack the appropriate level of code-terrorism-style imagination, but also because it’s something that has always happened under the hood. Oh, that’s where the—the code gets uploaded there, then it magically runs as your Lambda function. I guess, starting at the beginning, what got you thinking down this path?

Michael: I guess the key thing is, I was waiting for woefully large Lambda function to upload, and I was just looking at how big of a Lambda function can I actually upload, and what sort of limitations there are. So, they’re got me looking at the limits on that. And that’s when I started thinking down the path of, why is there a 75 gig limit to Lambda functions? There’s probably a good reason for that, and from there, I thought, “Oh, it’s because it’s not hosted in your bucket.” Unlike CodeDeploy or OpsWorks where you’ve run the infrastructure, it’s actually in Amazon’s bucket, so that’s why they’ve put that cap on it. And then I thought about ways that that could possibly be, I guess, abused to store your own stuff for free.

Corey: Oh, it could be abused awesomely. This is also, I guess, this is not the first time that I’ve looked into that 75 gigabyte free storage option for, I guess, effectively stealing resources that they didn’t think anyone would actually use in order to build something horrifying. Ben Kehoe and I have talked a bit about building a PackratDB, which is how much of a database can you actually shove into an AWS account without paying for anything. And we’re continuing to explore what that might look like, and this was absolutely one of the single largest points of data store you can get. Everything else is about free tags on resources that don’t cost anything, build an awesome key-value store, etcetera, etcetera. But this one blows away everything else we found so far, as far as, what can I get massive storage-wise without having to pay for any of it.

Michael: Yeah, certainly, and, the other thing is thinking about transfer costs. So, a lot of places you can store data, but you somehow get charged for requests to pull it back out. And so that’s what was a little bit weird with the Lambda functions. And, I guess, the only tricky part is making it usable for an end client without having to have a Lambda function that pulls it out, does some data mangling it to get it into the right format, and then send it back to the client. So, that’s the only tricky part with using Lambda functions like that.

Corey: Okay, so let’s continue through the demo. You figured out that you could have 75 gigs of data just hanging out there and you could pull it back at no charge. Where do you go from there?

Michael: Right. So, I remembered a blog post, a Medium post by Laurent. And that blog post was about using Google Docs to store video, and they used the HLS format to basically skip through to the section that actually contained the video content. So, in their example, they’re uploading the videos as inside a PNG file, and then they could skip through to that. Their purpose is they wanted to hide it away from Google. So, Google thought that it was just a picture, didn’t try to do any copyright data matching on it, and from there, I thought, “Hey, I could probably use that inside Lambda to make use of the storage.” So I guess the key problem for Lambda is you need to upload a zip file, so you need to somehow have the video content in the zip file, but still accessible by the client. So, for that, what we do is we take the zip file, and we actually compress it with zero compression. So, depending on your zip utility, there’s lots of ways of enabling that, but basically you say zip it up, but don’t compress it. So, the entire file is there, uncompressed, you just need to jump to the right byte offset. And that’s where the HLS stream comes in. You can say, “Hey, just jump to this part of the file, and you can skip all the zip header stuff.

Corey: Right, so, you have the header itself that’s there, but effectively, when people are saying they’re looking for a great compression algorithm, you were looking for the exact opposite of that.

Michael: Correct, yeah, we want something that’s not compressed, that way that—because the video client doesn’t know how to deal with the compression. So, yeah, if we have it not compressed then that works better in our favor.

Corey: Gotcha. So, you wind up with an effect of having one giant object sitting there but you can also do the byte offset to tell the video player where to get it from?

Michael: Yeah.

Corey: Is that using the byte-range stuff that is in S3’s GET API, or is it using something else?

Michael: Yeah. So, if you do a normal GET to S3, it uses the bytes range header as part of the request. So, it will skip through that.

Corey: Yeah, I think that’s one of those things that people aren’t generally aware exists, where you don’t need to pull the entire object down. You can just say, give me this very specific portion of it.

Michael: Yeah, it’s a very handy feature for more production-like workloads.

Corey: So, you wound up then putting this behind an API gateway, and then hooking that up to a Lambda function?

Michael: Yeah.

Corey: Have you looked at all into their new HTTP API option, which, I think, is now in beta? They talked about it a lot at re:Invent, but I haven’t had the chance to play with it myself yet.

Michael: Yeah, so, I actually tried too—because I thought this would be a brilliant demo of testing that out. And I tried to set that up and I followed all the steps, and I just could not get the API to return anything but, a 403 or a 500 or something. So, I clearly—

Corey: That sounds like most of my early explorations with API gateway from start to finish until I started just using something like serverless framework to wrap it for me. I feel like, for a long time, that most of what the stuff I was getting from API gateway was just a comedy of errors. It is not the most intuitive thing to learn. And I’m disheartened to hear that potentially is what we’re seeing from the new version as well.

Michael: Yeah, I’m not sure. I feel like I’ve probably done something wrong. There’s probably some key part of documentation or possibly I just seem to wait long enough for DNS to propagate or whatnot for the new API. But I quickly—

Corey: Like, who has the patience for that?

Michael: Yeah. Yeah.

Corey Quinn: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the Enterprise (not the starship) on-prem security doesn’t translate well to cloud or multi-cloud environments. And that’s not even counting IoT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IoT devices, detects these threats up to 35 percent faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at ExtraHop.com/trial.

Corey: Okay, so once this video is up and running, and you give someone a URL that’s fronted by an API gateway, that’s awesome. But what, I guess, viewers or clients could wind up understanding that?

Michael: Yeah, so one of the problems you have with this is because it’s hosted in S3, in that bucket, they haven’t enabled calls for us. And you know, that’s not surprising. But you can play it in basically every modern HLS player, so VLC, MPlayer, I even got it running in Windows Media Player. The problem you have is you can’t embed it into a webpage because, of course, so you can link someone to it and they can open it in their player, but you can’t open it in a browser.

Corey: Forgive me if I’m remembering this backwards, would it be possible to adjust that by modifying the payload response that winds up being returned maybe through the Lambda function itself?

Michael: So, yes, but the problem is, the only thing we’re doing in the Lambda response is providing a 302, and it’s the S3 bucket of where all the Lambda functions are stored. That’s where the calls needs to be, and we don’t [crosstalk 00:12:23]—

Corey: Gotcha. And if we want to continue to make AWS eat the bill for it, there’s not really a great series of answers around it.

Michael: Yeah. So, in theory—and this is where my knowledge falls apart—in theory, I believe there’s like a same origin, or there’s an HTML tag that you could potentially put in the video tags to get it to work, but it seems that none of the out of the box HLS streaming JavaScript libraries that I looked at support doing that, and I don’t know if it is possible, but it seems to imply that it is.

Corey: What other limitations do you see in something like this?

Michael: So, you are limited in your chunk size. So, you can only have a 50 meg file per function, and that poses a problem if you’re doing 4k video because you won’t have a very large, like, a long duration of video content in that, and that can cause some issues with some of the clients trying to buffer because they only try to download the next couple of files not based on time, so you end up with stuttering stuff. So, for 1080 video it seemed to work okay. I haven’t tested it with 4k. I did this all by hand by manually writing the HLS stream file. It’s trivial to automate, but if I was to do a 4k one, I would have hundreds and hundreds of Lambda functions I need to upload

Corey: I’m guessing this is probably not intended for anything remotely resembling production use?

Michael: Most certainly not. I imagine Amazon is going to do something to block this somehow. There’s a few ways that I’ve thought of how they could do that, but I don’t know what approach they’ll take, because they do have to remain having backwards compatibility with—

Corey: Right, they view APIs as promises, as they love telling us, so the question then becomes, what could they change to make this a non-workable solution, and of that list, which of those are viable without having to break existing functionality?

Michael: Yeah, so the two that come to mind is they could put request-to-pays on the S3 bucket. And that would basically just mean that you’d start paying the cost for downloading that. I’m not sure if that’s something they would do. The other one that comes to mind is if they did the zipping themselves as well so that it is actually compressed, I’m not sure internally how that would go if that would break things. It would fix using this for HLS, but it could still be used for other data purposes.

Corey: Yeah, it’s one of those interesting areas where it would solve this particular use case, but there’s nothing preventing someone from building out something like this that just grants, Oh, run this one-liner, and you’ll get 75 gigs of storage per account.

Michael: Yep.

Corey: Yeah, it’s interesting to see how this might wind up. I guess, influencing AWS. I mean, there is always the option where they just decide that this is an acceptable loss. They’ve made something like what $32 billion in revenue last year. Yeah, if people want to go through these kind of hoops, okay, they won, will let them. Unless they start seeing widespread abuse of this, which frankly, I kind of have a hard time envisioning, I don’t know that this is necessarily going to be on the top of their list of things to chase down.

Michael: Yeah, I’m not certain about that because, I guess, one of the use cases for this is video piracy. So, it could potentially be used to pirate movies and stuff as they come out. And it means that the person doing the pirating or hosting the pirating, apart from having their Amazon account shut down, they don’t really risk spending a huge amount of money. But, the other use case I kind of thought about—I mean, I just did this for fun, I had no no use for it—but the other use case I thought of after I built this was those times where media outlets have this huge story, and they want to release it to the world, but they know that they’re going to take a pretty big hit in terms of hosting costs for it. They could just quickly do this. And that would save them like probably millions if everyone’s looking at the same video content.

Corey: Right, effectively a freestyle CDN to some extent. The counterpoint is I can’t really see any reputable media organization going down this path. It just seems like it would be a little bit too far towards the, “What do you folks think you’re doing?” model.

Michael: [laughing], Yes. Yeah, certainly. Yeah, I can’t imagine anyone doing it, but I don’t know, any way to cut costs, I guess. [laughing].

Corey: As soon as one person does it at that point, I feel like they’re no longer able to ignore this as just some weird proof of concept someone on the internet threw up.

Michael: Yeah, certainly, yeah. I imagine this will last a while. It’s definitely not—I imagine they’ll monitor it, maybe, run some analytics on this. And then at some point, once it gets to a tipping point where it’s worth changing, then they’ll look into it.

Corey: Right. And there’s always the customer-unfriendly approaches where they can just solve this entirely in their terms of service. And after they find egregious users, and start effectively turning off their AWS accounts, the message would probably get out. It feels like something that a company that isn’t Amazon would be likelier to do.

Michael: Yeah, yeah. I imagine like based on some of the previous examples of this sort of, I guess, code terrorism, it seems like they’re more likely to just eat up the bill until it becomes a huge problem.

Corey: I feel like I need to highlight yet again, this is not something that people should use for production use. For a while, I was giving a talk called “Terrible Ideas in Git,” and I had a Docker container that was published and ready to be used for this, just because resetting a whole bunch of Git repositories after you’ve mangled the hell out of them is obnoxious. Just run a Docker container every time you give the talk, and things are great. The container was called Terrible Ideas, and I’m sure someone was using it for something in production because people do terribly stupid things without any rationale, similar to the time where I started making jokes about using Route 53 as a database, and I started getting people responding with, “Well, that’s not the worst idea in the world. What if we did it like this?” And it’s no, no, no, no, no. At some point, the job takes on a life of its own, but you kind of want to at least keep the sharp edges away from people who may not understand what exactly it is they’re doing.

Michael: Yeah. And that’s why I tried to put a fairly decent write up on how it works, and also the limitations. Sort of a disclaimer to say, “Hey, this probably won’t work in the future.” But I am very scared about how many stars this has gotten on GitHub. And there’s apparently three forks of it already. So, hopefully, no one’s actually using this in the production sense.

Corey: One would very much like to hope. The counter-argument though, is that people will always surprise you with what ridiculous things they’re doing. Back when I was doing open-source development work on SaltStack, I figured that, oh, the problem clearly is that everyone who’s used Puppet, or Chef, or anything like that was just, oh, they weren’t very good at what they do, and the tool was not adequate. We’ve built this thing, it’s going to be amazing. And that lasted right until I saw my first customer use case where, oh, it turns out that anything is a hammer if you hold it wrong. It’s difficult to get people to see the vision, and I feel like the things that you build never survive encounters with other people’s horrifying use cases.

Michael: Yeah, a lot of the things I have built in the past have been, I guess, horrible, horrible things. Mostly just for fun, just seeing how far you can take a tool to work. I guess an example of that is I have built in the past a Lambda function that works as a custom resource in CloudFormation that starts a Mechanical Turk instance question and provides access keys and secret keys. So you can free text, create your CloudFormation and just say, “Hey, can you build me an S3 bucket? and it will fire off a Mechanical Turk request and ask someone on Mechanical Turk to build the S3 bucket for you.”

Corey: On some level, you have to wonder at what point they just automate a lot of these common solutions into something that is AWS Solutions option, or a Quick Start, versus how much of it is something like AWS IQ where you can effectively pay people a few bucks to do common or uncommon things, as the case may be. I would not put this past being wrapped around an official AWS offering at some point.

Michael: Yeah, yeah, I can imagine.

Corey: So, I’m assuming that this is not the thing that springs fully formed from, “You know, I’m going to go online today, on my first day on the job and go ahead and build something like this.” Where does it come from? Where did you wind up, I guess, starting down the path of thinking about creative use of services like this?

Michael: I always think about limits and how can I, at least, make use of the limits, get to the boundary? That limit is set, how do I get right up to that edge and make the most out of this? So, I guess, every time I look at something, if I see a limitation—I guess the prime example is S3 when they first released it, and I’m pretty sure they only charged for data transfer. When they first released it, they didn’t have any billing for GET requests, or HEAD requests, or options and all of that stuff. All those requests weren’t billed. So, I heard the story, a long time ago, about someone that essentially used that as a database. Because those requests were free, they weren’t really grabbing any data out of it. And that’s when Amazon had to add that limitation, I’m like, “At some point, I really want to be doing that. Like I want to be the reason why Amazon puts in that limit.” So, every time I look at a new service when it’s released, I look at the limits and try and work out like how can I use this to its fullest potential that Amazon never actually planned on it being used that way.

Corey: Right. I want to be the exception case, how do I make that possible?

Michael: Yep, exactly.

Corey: And, you always think, well, no one would actually go to the trouble of doing that stuff. Well, have you met me? Your Lambda function is a whopping 36 lines, all in, in Python. This is not a massive amount of code. It’s not anything that is overly complex. I think it just requires looking at these things from a certain point of view that, very often, the people building it never considered.

Michael: Exactly. And I feel like there’s probably a few other cases in Amazon where this approach, like not this exact code, but this approach can be applied to. I haven’t been able to see those. It’s just, I happen to use Lambda enough that I have worked out the inner workings a bit more. But, I’m sure there’s other places where you can upload data and get it back down for free that could be abused, maybe not for video streaming, but at least as a free database.

Corey: You almost start to wonder, okay, what is the upper bound of data you can attach to an AWS support ticket?

Michael: Yes.

Corey: Because it does have an API. One other thing I thought was kind of neat, too, right around the same time that this came out, was someone did a whole write up about how anything outside of the handler function in a Lambda ran with the full 3 gigs of resources and 2 vCPUs, and wound up not charging—or wound up not billing, or something on the order of that, were until it entered the handler, that it was either unbilled or build at a small fraction of what it was that was being charged. Do you remember what I’m talking about?

Michael: Yeah, I’ve seen that and read that, and so… yeah, I think, if I remember correctly, it’s just, as soon as—on that cold start, it has full power, full memory, to just set everything up. It’s designed for big Java applications and whatnot, so it can quickly get started, so that cold start time is really short and then run, but you can abuse that by running your code outside the handler and do that. And a lot of the stuff I do professionally, we do a lot of stuff outside that handler for the—just setting everything up. But I never really thought about using that extra time. But I have wondered, could you expand that to be more useful as a clustered, distributed computing system? It’d be really cool to see that expanded on because Amazon gave that tick of approval to say, that’s fine. So…

Corey: Yeah, they said, “Have fun,” on multiple folks who are in a position to be authoritative on Twitter, said, “Yeah, go for it. See what you can build.” So, all right, challenge accepted. This is the danger of goading people on who have very little sense of, “Oh, I shouldn’t do that. They wouldn’t like it. Oh, no.” At some point, you’ve got to get portions that AWS bill back.

Michael: Yes, Yes, certainly. Amazon is definitely not losing out on this. If you’re using any of their services, they’re winning.

Corey: Absolutely. I’ve yet to see a single exploit like this that didn’t result in, “Yes, and it winds up causing a slight discounting on my bill that is already a phone number.” It’s… this is very much a rounding error, even a per-user account for most of these things.

Michael: Yes.

Corey: So, we’ll see. So, if people want to discover more of your various acts of code terrorism, where can they find you?

Michael: They can find me on Twitter. So, @xssfox. X-ray, Sierra, Sierra, Foxtrot, Oscar, X-Ray.

Corey: Excellent. Thank you so much for taking the time to speak with me today, I appreciate it.

Michael: No worries. Thank you.

Corey: Michael, code terrorist at undisclosed location for obvious reasons. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave an excellent review on Apple Podcasts. If you’ve hated this podcast, please leave an excellent review on Apple Podcasts.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Sandy Carter
Sandy Carter is the AWS Vice President, Public Sector Partners and Programs. In her new role, she is responsible for driving next-generation partnering. Her responsibilities include evolving partner models to intensify partner innovation, AWS cloud adoption and creation of mission critical cloud solutions with partners across public sector. Her impact will be growing the partner ecosystem as a major driver for public sector and contributing significantly to the success of Public Sector customers.

Prior to this role, Sandy built an enterprise workload team as the Vice President of Windows and Enterprise Workloads at Amazon Web Services (AWS) focused on helping companies innovate using their current technology and assets with migration and modernization through containers and serverless. She led the team to overall growth with AWS now hosting nearly two times as many Windows Server instances in the cloud as Microsoft, per IDC. She led her engineering team to optimize SQL Server on AWS which exhibited 2X+ better price/ performance than Azure per ZK Research. For her leadership on the VMware Cloud on AWS business, she led the team to deliver results of 4x the number of customers year over year, with those customers having deployed 9x the number of VMs now vs 1 year ago. Finally, she grew the number of competency partners by 3x and the number of ISV validated solutions by 4x in the last year.

She is the author of Extreme Innovation: Three Superpowers for Purpose and Profit, built on her research with Carnegie Mellon. Sandy was named Lifetime Achievement Winner, 'Excellence in Cloud Achievement' for 2019, AI Innovator of the Year Nominee in 2019, Top 10 AI Influencers for 2019, Top 10 Cloud Computing Influencer, Top 39 Engineers by Business Insider 2018, Top 50 AI influencer by Onalytica 2018, Top 10 Future of Work influencer in 2018, and Top 10 Women in Technology by CNN.

Sandy is the Chairman of the Board of Girls in Tech, and an adjunct professor at Carnegie Mellon University Silicon Valley. Last year, Girls in Tech had over 125K women participate in their “Hacking for Humanity” initiative, and trained over 90K women globally on coding through boot camps, and workshops. She was honored two times with the AIT United Nations Member of the Year award for helping developing countries with technology. She is an Advisor to startups in AI, IoT, and AR/VR.

Links Referenced

  • DigitalOcean: https://www.digitalocean.com/, http://do.co/screaming
  • Observe 2020 Virtual Conference: snark.cloud/observe
  • re:Invent: https://reinvent.awsevents.com/
  • ExtraHop: http://ExtraHop.com, http://ExtraHop.com/trial
  • Coding for America: https://www.codeforamerica.org/
  • Twitter, #techforgood: https://twitter.com/search?q=techforgood
  • Twitter, Sandy Carter: https://twitter.com/sandy_carter
  • LinkedIn: https://www.linkedin.com/in/sandyacarter/
  • Email: mailto:sandyct@amazon.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey Quinn: This episode is brought to you by DigitalOcean, the cloud provider that makes it easy for startups to deploy and scale modern web applications with, and this is important to me, no billing surprises. With simple, predictable pricing that’s flat across 12 global data center regions and UX developers around the world love, you can control your cloud infrastructure costs and have more time for your team to focus on growing your business. See what businesses are building on DigitalOcean and get started for free at do.co/screaming. That’s D-O-Dot-C-O-slash-screaming and my thanks to DigitalOcean for their continuing support of this ridiculous podcast.

Corey Quinn: This episode is sponsored by the Observe 2020 Virtual Conference happening on April 6th. Are you looking for a germ-free way to learn all about observability and possibly the new open-telemetry project? If you visit snark.cloud/observe, you can be part of this one day virtual conference happening on April 6th. With talks and workshops packed tighter than a multi-stage flattened nano-based Docker image. $50 per person, but group rates are available for distributed teams, too. Ticket purchases go towards providing free learning for folks experiencing economic hardship and underrepresented communities in tech. Visit snark.cloud/observe to join them on April 6th that's snark.cloud/observe.

Welcome to Screaming in the Cloud. I’m Corey Quinn. I’m joined this week by AWS’s VP of Partners and Programs. Sandy Carter. Sandy, welcome to the show.

Sandy Carter: Thank you, Corey. I’ve been so excited about joining the show, so, I appreciate your having me on.

Corey Quinn: Of course. So, when we first started speaking, you were effectively running AWS’s Windows operations. And then when your bio came through, it said that you’re now the VP of Public Sector Partners and Programs, which, okay, I guess that might be Windows-oriented, sort of if you squint hard enough, but what’s the story behind that?

Sandy Carter: Yeah, so I came into AWS about three years ago to run enterprise workloads. And so I focused on our Windows business because obviously lots of enterprise workloads run Windows, as well as VMware Cloud on AWS. A lot of our enterprise clients today run VMware. SAP as well as Salesforce, and now IBM Flash Red Hat. So, I really honed and focused on those workloads. And then I was asked to come over and broaden the perspective, so not just those five workloads, but what could we focus on globally from a partner perspective? And so we’re now looking at public sector, in not just the government but state and local, education, healthcare, a lot of healthcares are nonprofits, outside the US, a lot of telecommunications, a lot of transportations, and then, in several regions of the world like the Middle East and some in Asia, the entire country, essentially is public sector, and a lot of that growth and a lot of that scale will come through our partners. So, I am super excited to take that Microsoft role, VMware, SAP, etc, and expand it out and about, across the globe.

Corey Quinn: What’s fascinating about your background to me is you come from pretty much the exact opposite of everything I’ve ever done in my career, for lack of a better term. You come from extremely large companies, big e-Enterprise style approaches, and I come from a small business world where, to my mind, a big company has 200 people. So, I assume that’s roughly how big AWS is. So, it’s, “Oh, you work at AWS? Do you know Jordan?” And the answer is always “No.” It turns out you folks do have more than a couple hundred people working there. Who knew? It’s just a different mindset, and it’s really its own skill. Unfortunately, it seems to be one that I lack, more often than not. How have you found that, I guess, transitioning from some of the largest companies in the world to, I guess, AWS in its, I guess, practically small team mentality, is really—I guess, what's that been like? Has it been a change? Has it been more or less the same?

Sandy Carter: Well, here’s the way I would think about it. So, I was at IBM, and I was actually running our partner ecosystem for the Cloud, IoT, and machine learning, and AI. I was asked to move out to California—which I love, the valley is just so energetic around tech—and to go out there and really develop that ecosystem. When I made the decision to leave IBM, I decided, I had been working with all of these startups, and when you’re in the valley, you get the startup bug. So, when I left IBM, which is this huge company, I went and started my own company. We were working on some research about how a company’s culture needs to match what they’re doing and innovation. So, Corey, I would see all of these huge companies and countries come into Silicon Valley. They would spend six months to a year, and they would learn everything about how all the top innovative companies there did business. And then they would take it and they try to take it back to their country or their company, and it would literally fail. So, what I worked on as my startup, is I worked on the ability to do kind of a Myers-Briggs for a company and match innovation tactics. So, needless to say, it was very small. So, I went from big, big, big IBM to really small; my own company. And that actually eased the transition into Amazon. So, once I had built my MVP, I had done some work with Carnegie Mellon, Silicon Valley, on a piece of research, I was able to sell that—those assets off because I only had done an MVP, so I can’t even say product. I was presented with an opportunity to come to Amazon. And so when I came, I had kind of already transitioned to being my own assistant, doing my own two slides and narratives and documents and thinking differently, being very scrappy. And I think that that transition really helped me a lot as I came into Amazon because as you know, Amazon’s a bunch of two-pizza teams. So, you are working in a start-up when you enter AWS.

Corey Quinn: That’s always been the sense I’ve gotten. It’s—I’ve never been the type of person that would ever fit in any large company. But of any large company that I’ve ever spoken to, it seemed that AWS was the closest to the startup mentality that I tend to deal with, without some of the problematic parts that, generally, startup culture in San Francisco is best known for.

Sandy Carter: Yeah, I would say, I’ll just give you one quick example. When I started, an intern that worked for me came up with this idea. He did a six-page narrative which, of course, you know we do at AWS. He brought it in to me, I was like, this is a magnificent idea. We had, you know, of course, a team of people sitting reading the document in silence and then evaluating it. And so, I went to my boss and I said, “This is a great idea. How do we get this going?” And he’s like, “Well, you go. Just take your team and you just do it.” Which was very different from every other company I’d worked at. And so we did it, and it’s grown to be a sizable business, an idea that came out of an intern, which is typically a story that you would hear at a startup in the valley. And I have many stories like that from Amazon. So, it’s quite a fun place, and a very innovative place, and a very much of a small feel. I really love it.

Corey Quinn: Changing gears slightly, something that’s been very interesting is watching, I guess, almost a transition in what AWS has been focusing on and speaking on. At re:Invent last year, one of the things that was very notable throughout the course of the keynotes and the releases was an emphasis on the idea of hybrid. There’s always been some options, Direct Connect, or the Snowball Edge with the computer variant as well, but this was the first time that we really saw more or less the entire theme of the conversation revolving around hybrid. Things like VMware Cloud for AWS, Outposts being released, it really felt like it was a change in tone. Is that an accurate assessment? Or is that just, from where I sit, I’m just starting to notice things I didn’t used to notice before?

Sandy Carter: Well, I would say this. You know, I came in with VMware Cloud on AWS being one of my big missions, and that is hybrid cloud. So, we’ve been speaking a hybrid, if you would, for quite some time, but kind of with different motions. And what we were doing is, as we looked at our VMware Cloud on AWS offering—great offering by the way, if you are a VMware shop today, you want to get to the Cloud, it really helps you quickly migrate, integrates networking, both NSX-T with Direct Connect or vSAN storage with our EBS, it’s a great solution, and a great way for customers to kind of get to the Cloud, but in a hybrid way. And also, many of our customers love the fact that they can VMotion applications back and forth as well. As we started understanding that solution, we started listening to a lot of our customer feedback, which is what we’re known for, right? Working backwards from the customer. So, I think what AWS has done on hybrid cloud is really amped up the messaging, amped up the product portfolio. I mean, Outposts is an amazing offering, that gives you not something like AWS, but gives you AWS on-prem and has that ability to enable customers to now leverage hybrid in a way that you haven’t been able to before. Same with Wavelength or our Far Zones concept. It’s just an incredible way to take this to the next level. So, I would say we’ve been talking about hybrid computing for different segments of customers. I think what Amazon released, and Andy talked about a lot at re:Invent, was really fulfilling more promises, working backwards from our customers.

Corey Quinn: That’s one of the things that always struck me as being a truism about AWS, and Amazon in a larger context, and that people like to say, “Oh, it’s a corporate talking point.” And I’ll be perfectly honest, I used to think the same thing. But increasingly, every time I would wind up having a conversation with someone at AWS about something, it doesn’t matter how ridiculous my ask was, or what I wanted to do—if it was part of the setup for one of my ridiculous ideas, like using Route 53 as a database—the answer is always the same. “Interesting, can you tell me more about your use case?” And sometimes that use case was ridiculous, and other times it was, “Oh, well, you might not think of it that way. What about if you use another service or offering?” But I’ve never yet been made to feel dumb when talking to someone from AWS, despite the fact that oftentimes I was being extraordinarily dumb.

Sandy Carter: You know, it’s really interesting, because when I came into AWS, and when I developed—I had an engineering team and product management—when we developed our first—my first service, I will say that, Corey, I talked to 141 customers myself, got all kinds of use cases, and no question or comment was dumb. In fact, it all played into, you know, working backwards for the customer. You know, we always say 90 percent of what AWS builds is what AWS customers ask for. The other 10 percent, of course, is strategic interpretation of that need. I think that that is really where the power is, is listening to customers, you know, having builders talk to builders and really understanding. In fact, I kicked off one of my sessions once with a quote from Stacey King. I don’t know if you like basketball, but I’m a big basketball fan. And Stacy King was a rookie, and he played with Michael Jordan who, in my mind, was the best all-time basketball player ever, and he said, “I’ll always remember this day as the day that Michael Jordan and I combined to score 70 points.” And Stacey King scored one point.

Corey Quinn: Excellent.

Sandy Carter: Whenever I see us announce something, I always give credit back to the customers. I say, “I scored the one point because I engineered and did the product management on it. But you customers, you did that 69 points because you gave us the ideas. You gave us the nuances, you gave us the use cases that enabled us to build something new and something that works for you.” And that’s why we have such fast adoption. In fact, that service I told you about was License Manager, which is now used by thousands and thousands of customers, and was only announced less than a year ago. And the reason the adoption is so fast is because we do listen to our customers.

Corey Quinn: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the Enterprise (not the starship) on-prem security doesn’t translate well to cloud or multi-cloud environments. And that’s not even counting IoT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IoT devices, detects these threats up to 35 percent faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at ExtraHop.com/trial.

Corey Quinn: So, let’s pick a fight for a minute. [laughing]. I have a privilege of seeing you speak at a couple of summits last year. And one of the things you said, both times—and I figure if you say it on stage in front of the world, it’s a fair thing to ask questions about—and what you said was that AWS was the best platform for Windows workloads. Fair enough. I’m not going to dispute it, but I will ask why?

Sandy Carter: Well, first of all, AWS has run Windows for over a decade. We host nearly two times as many Windows server instances in the Cloud, as does Microsoft. So, you’re like, “Okay, so what?” Well, if you have all that experience, you know them Andy Jassy quote that says, “There is no compression algorithm for experience.” It’s really true. Our experience running Windows workloads has earned us our customers’ trust. We have tens of thousands, hundreds of thousands of customers bringing their workloads over. So, whether it’s Autodesk who’s been running Windows on AWS for 10 years, or Salesforce, who has over 10,000 Windows instances, they run it on us. So, the question then is why, right? Why is that the case? And I would tell you, there are I think three or four core reasons. One is, we continue to innovate. So, whether it’s you know, Windows Server with things like License Manager I just talked to you about, which enables you to manage your licenses. Or SQL Server, being able to innovate with insights for SQL using machine learning, we can help you debug SQL application bugs on AWS faster, or .NET support. We were some of the very first support Lambda for .NET, and use of machine learning with .NET, or application modernization, we just innovate faster and have that agility that enables our customers to move quickly. It could also be in innovation and optimization. So, right now, the performance of running SQL Server on AWS is two to three times faster than running SQL Server on Azure. And that’s because we’ve taken the time to optimize our networking and our storage to make that happen. The second big reason is reliability. And some people are like, ah… but reliability matters. A lot of the Windows workloads are mission-critical. And reliability really begins with our global infrastructure, of which underpins all of those Microsoft services. We have 61 availability zones, 20 geographic regions around the world and the latest—it’s only from 2018, but they published it—it said that the next largest cloud provider had seven times more downtime hours than us. So, we get a lot of customers coming to us saying, “Hey, I tried running my Windows workloads on Azure. It’s just not reliable enough, I can’t count on it, so I need something better.” And I would say the third big reason is security. Our security, as Andy likes to say, is our priority zero because it’s so important that we prioritize it above everything else. And just one example is, if you look at FedRAMP, which is really important, not just for the government, but for regulated industries, we have over 90 Solutions today that are FedRAMP authorized, and that is four times more than Google and Microsoft combined. So, we take this stuff very seriously, and it’s because of our innovation and our reliability, our performance, our security, that people are choosing us and it’s in the numbers. It’s not a debatable thing. It’s definitely in the numbers, Corey.

Corey Quinn: There’s a lot in there to unpack. And I’m sure that some people are going to be sitting up there taking angry issue with almost everything you said, in which case, cool. Email you, not me. But one thing I have noticed is that Microsoft’s cloud offering is one of the, to my understanding, the only major cloud provider that offers an SLA around individual instances. Which, first, is terrifying about what it says about those workloads. Windows workloads have a tendency to be, shall we say, tetchy, but it does speak to the idea that this single pet of an instance can’t die, and that is something that I know that I’ve spoken with customers where that idea resonates. As much as we love talking about the idea of, “Oh, there should be no single points of failure in your environment.” Well, here in the real world, we’re often nowhere near lucky enough.

Sandy Carter: That is true. And I would say that, you know, as you look at our customers, there are different use cases for each customer as they’re deploying, and based on what their needs are, we really try to make sure that they’ve got the right level of security, reliability, performance. So, for example, one of the things that we announced most recently is FSx for Windows, which is a fully managed native compatibility with Windows, the SMB protocol, the industry standard. And, you know, a lot of people were like, “Well, why are you doing that?” It’s Windows native, it’s built on Windows Server, but who’s going to really use it? We had such uptick in the use of that even though it was really about performance and cost primarily, we thought, in the media industry, but then we had customers like Neiman Marcus, Ancestry, Cube Research really take advantage of using this feature that drives wicked fast performance with lower TCO. So, I think that you can find a use case or something that fits every different customer scenario, and that’s one of the things that Amazon is so good at is trying to figure out what does each customer need? What is the level of security? What is the level of performance? What are they really looking for? And trying to fill that need.

Corey Quinn: There’s definitely a value in meeting customers where they are. Back when EFS, and then it’s Windows equivalents came out, I was fairly harsh to the idea, because from my perspective, if you’re building something net new greenfield, you don’t necessarily want to use a shared file system in a cloud environment, more or less the effective way of doing these things, object stores, etc, etc. And I sat here in my ivory architecture tower, and I wound up getting some feedback from the GM of that product—Hi Wayne—and the entire point was, “Okay, you can take that position if you’d like. But understand that not everyone has the option of throwing everything that they’ve built away, and starting over from scratch.” That seems like it is the, from my perspective, the San Francisco startup disease of anything older than 18 months is ready to be thrown away. But that’s just not how the business works, it’s not how the world works. And if you’ve got to run NFS or the SMB/CIFS equivalent in a cloud environment, there’s no better way to do it. And I started digging into the services and realized, yeah, I really should have had a more informed take on that when it first came out. There’s something to be said from learning from our customers.

Sandy Carter: Absolutely. And, can I pick up on another point, Corey?

Corey Quinn: Please.

Sandy Carter: That you said that you can’t throw everything out and start over. I mean, those are the customers that I’ve been working with for the last three years. They’re enterprise customers, so they don’t have the luxury of being a cloud-native startup, which has advantage because they’re beginning with native cloud services on Linux. And that’s also why I’m so excited about another announcement that we made, which was BYOL on Dedicated Host, and making that more cloud-like, right? Because we have so many customers today who are still running Windows Server. And they need a way to get that to the Cloud. And instead of telling them they have to just throw it out and get rid of it, we took a different approach, and the approach we took was, you want to continue to bring those licenses to AWS but a dedicated host is kind of a non-cloud-like experience, right? You have to have a dedicated machine to run that application on. So, how could we take away you know, that physical server that’s dedicated to the single use of the customer, and make it appear or seem like it was running an EC2 instance. So, that’s what we did. We enabled our customers to be able to create a BYOL instance just like they do any other instance. They can do that in the command line, it can do that in the GUI. To me, this was one of the most exciting announcements, again, leveraging what you were just talking about. I can’t throw everything away, how do I get that to the Cloud so I can start my experiments?

Corey Quinn: That’s—it’s always easy to sit here and tear things down that are existing. What are you doing with that ancient piece of crap? Oh, about eight billion dollars a year in revenue, why do you ask? It feels like there’s an aversion in tech to anything legacy, by which we mean it makes money. There’s a lot of pushback in from the quote-unquote “thought leaders” of the world who are, “Oh, this is now the way to do it, the way and the light, and here’s our beautiful whiteboard architecture.” I’ve never yet seen one of those that was a realistic representation. Production, the real world, it’s always messy. There’s no way around that, and my biggest concern now is when people—one of them, there are many concerns—is that when people see these things on a stage somewhere at a conference, they’re convinced, rightly or wrongly, that, “Oh, this company has it all figured out, whereas ours is garbage.” But I’ve been sitting in conference halls where someone from a company is talking about their migration or their beautiful architecture. And I turn to the person next to me and say, “I’d love to work someplace like that.” And the person next to me says, “Yeah, I would too,” and they work at the speaker’s company. It’s conference-wear for lack of a better term. We all tell stories that get to a point, but nothing is perfect. Nothing is idealized and the world is messy, and I think accepting that is important.

Sandy Carter: I agree with you, I love that—conference-wear. You know, and these are the companies that like I said, I’ve been working with for the last three years. Great companies, great IT shops that are looking to, you know, mix that legacy with the new world, and so I was just so excited to work with them, not just because they’re big companies, you know, like Capital One, but because they have this need, and they really want to get to the new world, and they’re developing training programs and education to get there. And I would tell you another great idea that came up from one of those companies was single sign-on. So, most of our—most companies overall, not just most of our companies—they use Active Directory. And so one of the things they kept asking us is, “Look, you launched single sign-on and we can work with our existing AD on-prem, but because we use AD, we need to also be able to interface with Active Directory on Azure, because every Office 365 needs that, so how can you help us? So, one of my other favorite announcements, again, came from this need of leveraging that legacy. We provide support for Azure Active Directory. And you can do that through standard protocols, so you don’t have to get locked into anybody, gives you more flexibility in your overall cloud. But we also added in flexibility and multi-factor authentication so you could use Google Authenticator or Microsoft Authenticator. And then we integrated it into the command line, to make it easier for someone to sign in and stay in securely without having to change, and with an automated short term credentials to eliminate the hassle of staying secure, and freeing up that time for the developer, yada, yada, yada, all those great things. So, this is one of the reasons I really love my job, because my job was to help those legacy companies, and even more so now, help those companies figure out how to move and get there. And they want to get there, and we’re seeing them get there, and I think that’s probably one of the most, you know, exciting things. Other than seeing the CEO of Goldman Sachs DJ, at re:Invent last year. [laughing].

Corey Quinn: Yeah, that was definitely something. I heard the story, I knew that he effectively did that in his spare time. I did not expect to see him doing it on stage and that was fantastic.

Sandy Carter: I agree.

Corey Quinn: I will ask you about the re:Invent keynote, though. It’s one of my recurring jokes has been that the head of Amazon’s product strategy is a post-it note that says “yes” on it. And that is never more true than during the three-hour barrage of Andy standing on stage, talking about new services and features. So, we’ve long since passed the point where I can talk incredibly convincingly about services that don’t really exist, but no one at AWS would dare question it, because maybe it’s just not in their group. Maybe they haven’t heard of it yet. So, I promise not to do that. But I will ask you, what were some of your favorite releases that came out?

Sandy Carter: Oh, wow. So, I would tell you that one of my very favorites wasn’t in my group, but the IDK for machine learning. So, part of my passion—when I went to college way back when, I majored in computer science and math, and one of my minors was in artificial intelligence, which has changed a lot. And then I told you about my career at IBM where I spent time developing training modules for Watson, which is IBM’s AI. So, this was so exciting to me that we now have SageMaker, which is essentially now a true IDK. So, I just thought that that was a neat announcement to make, as a coder. You know, one of my favorite quotes is that programming is the closest to magic that you’ll ever have. And for me, that is so incredibly true because coding, you can create anything, you can really drive anything, and now with SageMaker, I’m just thinking, you know, SageMaker Studio that enables everything to be integrated for machine learning; or Notebooks, that gives you that notebook experience with a Quick Start; or the Experiments, enabling you to organize and track all of your experiments that you want to run; Debugger; the Model Monitor, which I love, because I love creating models, detecting the quality of your model and helping you to take corrective action. Really excited about that entire announcement, which I thought was amazingly cool. And then if I switch back to my Windows hat—I guess Windows and Linux hat—another one of my super cool announcements that I love was Image Builder. I don’t know if you caught that, but when I first started at AWS, the number one complaint my customers gave me is I have to build a compliant Windows or Linux image for AWS. And when I do that, I have all this work, right? I have to create it, I have to manage it, I have to deploy it, I have to keep it in compliance. And so, my team started working on this, really thinking about it, because it is hard. So, what we were able to release is Image Builder that automatically produces that new image and distributes it to AWS Regions after you validate the test on it. And your test could be your test, could be some of the Amazon best practices, it helps you reduce—the ability to build those compliant and up to date images. It was really one of my very favorite announcements. I got to do that announcement on stage. And when I did that, I think it got the most applause, and I even had a couple people stand up, which was super cool because we worked on it for so long. And it enables both Windows and Linux, we started building it just for Windows, but we made it for Linux as well. It enables customers to save so many different things. And I would say, my last one, I know I have a lot here but the last one that I really loved, and I don’t know how many caught it was, again for our Windows users, was the ability now to take Windows Server 2003 and 2008, and enable them to move up to the next version, but to encapsulate those 2003, 2008 API’s, move that up to 2006 or 2009, without refactoring.

Corey Quinn: Yes.

Sandy Carter: And we do that through a program, but also through some code as well.

Corey Quinn: When you say a program, I assume that is not in the sense of, it’s a bunch of code that compiles, and we hand it to you. The program as in a organized system of people.

Sandy Carter: That’s right, you got it. We announced some partners who are able to perform that, like Smartronix and Accenture, that we also announced—I think it’s important that it’s not just a marketing program. It’s actually technology that enables you to package these legacy apps up, so they can run on new versions of Windows Server without having to refactor.

Corey Quinn: That struck me as being something that is going to be transformative for an awful lot of companies because it—technical debt of—sorry, the term sounds derogatory, but let’s not even talk about in that sense—the problem of having workloads that are not current or haven’t been touched in a long time, is the people who built them are often gone. No one knows all the edge cases of it. I mean, the worst programmer I ever met, heck, was myself six months ago, when I do git blame and find out what moron wrote this, and I’m said moron, it’s, “Well, that tracks.” It winds up being a terrible problem, and how do you wind up solving this? If you don’t have to tear an entire application apart and rebuild it just to move it, suddenly that unlocks an awful lot of opportunity. I would further agree with you, with what you say about the EC2 Image Builder. In hindsight, I should have seen it coming, I think is probably the best way to frame that. But the easy way to figure out of what AWS is likely to do is, when I was back doing engineering work, and I was transitioning jobs every year or two, because my primary employee skill set is getting fired, what is one of the first things I would always have to do when showing up in a new environment? It’s, you set up certain things like the SFTP endpoint that then copies files to S3. That came out a year ago, which I was super thrilled about. Now you also set up something with Packer equivalent, an entire build pipeline to generate the Golden AMI that you’re going to use to spin things up in your environment. It’s the same undifferentiated work, that—it shouldn’t require a three-day project with bringing in a whole bunch of external dependencies to do that. So, that is one of those things I’m very thrilled to see come out and for better or worse—it’s a little late for me just because A) I’m not doing engineering anymore, and B) I have jumped on board the serverless train, which means that I don’t have to run computers anymore, or so they tell me. But, man, is that transformative for an awful lot of shops.

Sandy Carter: I have to tell you, I had so many customers come up to me and say, “Thank you so much.” I mean, literally thanking me for doing that, and talking to me about how many people that was going to save for them. I just think that that was, it was one of my very favorite announcements, for sure, as we talked about it, just because it’s not sexy, right? It’s not ML, because that was where I started. It’s not that sexy announcement, but it’s a real valuable announcement because it saves customers time, and then they can put that time and those people into other things that are a little bit sexier, but it enables them to do their jobs better. And that’s what really makes me happy is that personal connection, and being able to help our customers do things at lightspeed, if you would.

Corey Quinn: It’s the sort of thing that never makes a headline, and never makes the front page of the keynote, but it’s always worth looking at. One last thing I wanted to ask you about, you’ve been a passionate supporter of women in tech projects and movements for a long time. Can you tell me a little bit about that?

Sandy Carter: Yes, I would love to talk about that. So, I believe and AWS believes that—we’re all about innovation, and the way that you innovate is you bring in diverse ideas. And if your employees look like your customer, you’re going to be able to come out with more innovations that are right for them. So, first of all, I’ll just say that diversity and working with women in tech and girls in tech, it shouldn’t be just a nice to do, it really needs to be viewed as the business case of diversity. There is value in your innovation. There is value in just the different ideas that are generated. I’ll just give you one case in point. If you look at data that comes out around diversity, companies who have a diverse leadership team, are 45 percent more likely to grow their market share, and 70 percent more likely to capture a new market. And that’s directly from a McKinsey report. So, the data is there to support the value of diversity. One of the reasons that I’m really passionate about it is one of my mentors, Corey, always told me, when you make it, and everybody makes it every day of their lives, you have to reach back and pull someone along. So, part of my mission has been, how can I help women who are already in tech? How can I help them do more? How can I help girls who are exploring tech? How can I help them get really excited about the potential of technology? So, for example, one of the things that we had looked at, at our recent conference was a drone show that was led by an all-female drone team from Intel. And they were supporting girls in tech. Girls need role models, and here’s an all-female drone team that did—I don’t know if you saw it, but the spectacular drone show that was there at our conference. So providing not only funding and money to get them into coding school and classes, but setting that great example. So, I’m really passionate, very excited about helping others to see the value in technology and the fun of it too. It’s super fun, right?

Corey Quinn: It really is. It’s also important to be as inclusive as possible. It’s—there’s no value in having people who all look and think alike in the same room trying to build products. At that point, you wind up with hilarious misfires. The more voices, the better. And I think that there are problems in this space that are significant. We’re starting to see some advancement on there, but oh, there’s so much work yet to be done.

Sandy Carter: I agree. And I don’t want anybody to think I’m just talking about gender diversity. I think there’s diversity across the board. So, one of the interesting new facts I found was that more than one in three new businesses today are started by someone over the age of 50. So, you know, it’s generational. It’s cultural. It’s everything. In fact, one of our teams—I have to tell you this story. One of our teams was not innovating as fast as possible. And if you looked at them, you’re like, wow, it’s culturally diverse, it’s generational diverse, it’s gender diverse, what’s going on? And we found out that they all graduated from the same school. So, their frameworks and their mental model had all been trained by one school. So, diversity is really about diversity of thought, just so you can gravitate to hearing all those ideas that you talked about.

Corey Quinn: That’s, I think, where a lot of the work remains to be done. There’s—it’s a complicated series of issues. And it’s one of those things that I regretfully suspect that we’re not going to solve this in my lifetime, but I think we can definitely make strides and get far further than we have.

Sandy Carter: I completely agree. And I must say, if I could just kind of add one thing, too, that I think is so powerful, it’s another just thought about helping others, Corey. One of the things I also love about Amazon—and I know we saw some of this, we’ve seen it a lot—but one of the reasons why I love Amazon Web Services, as an example is, we have this—in my new role, we have this whole nonprofit group, and yesterday we were just reviewing some of the impact. You know, it’s not just about making money, it’s also about doing good. I really believe in tech for good. And we just went through yesterday, how Amazon technology is helping to identify human trafficking rings and child exploitation that’s going on. We saw yesterday a video from a skin disease called EB, which I had never heard of, and how Amazon is using machine learning and AI to help discover patterns and trends to help, hopefully, one day cure this disease. So, we’ve talked about a lot of things today. I think the bottom line, though, is that as you’re starting the New Year out, you’ve got to think about technology as a great enabler. It does allow you to do magic, but you should share that magic not just with your company, for earning money, but get your company to support diversity inclusion. Some of these other things that are tech for good—like what I love that you did at last year’s re:Invent conference with the shirts—you raised, how much money for St. Jude’s right?

Corey Quinn: $18,000 last year, yeah.

Sandy Carter: $18,000, which is so—

Corey Quinn: It’s surprisingly challenging to find a charitable cause that no one is on the other side of, except complete monsters. There’s always going to be someone who said, “Why this cause and not the other cause?” “Well, that one is a good cause, but is it good enough?” But kids with cancer is one of those things where it’s hard to find something more serious, and it’s hard to find something that everyone can get behind in quite the same way. There are options, but it’s always—it’s one of those slam dunk things that I personally feel passionately about. So, it was one of those things that just sort of tied together across a whole variety of box-checking, as well as a cause that I’m, again, profoundly passionate about.

Sandy Carter: Yeah, and if we jointly, Corey, could leave one message—I know everybody listening is all about tech—I think what I would challenge you to do, you probably already set your New Year’s resolutions but set one to do something, like what you did at our conference, where you just said, “Hey, buy a shirt and we’re gonna give the money here,” or Coding for America, or joining one of these efforts for helping human trafficking or curing disease. It’s really about giving back. So I would hope that everybody listening would also take a challenge in 2020 to do something good. To do hashtag #techforgood.

Corey Quinn: Thank you so much for taking the time to speak with me today, Sandy, I appreciate it.

Sandy Carter: It’s always cool. I wish we had a picture, though, that we could both open our mouths really wide, right? [laughing]

Corey Quinn: Oh, we can probably make that happen one of these days. If people want to learn more about what you have to say on this and many other topics, where can they find you?

Sandy Carter: Well, I would love for them to go to either Twitter: it’s @Sandy_Carter. I’m there with Corey almost every day. LinkedIn is another great way to catch up with me. It’s just Sandy Carter. I’m also on Instagram and Facebook, and I will also give my email out; it’s sandyct@amazon.com, if you’ve got a great suggestion for something on Windows, or a new partner thing, or just anything, your tech for good idea, email me as well. I’ll be happy to get back with you and chat with you as well. And Corey, I’m really honored to be on here, really appreciate all the great work that you’re doing, both on tech for good as well as just teaching tech and educating as well. So, thank you for having me.

Corey Quinn: Of course, thank you once again, Sandy Carter, VP of Partners and Programs at AWS. I’m Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on iTunes. If you’ve hated this podcast, please leave a five-star review on iTunes and a fun comment.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Ben Sigelman

Ben Sigelman is a co-founder and the CEO at LightStep, a co-creator of Dapper (Google’s distributed tracing system), and co-creator of the OpenTracing and OpenTelemetry projects (both part of the CNCF). Ben's work and interests gravitate towards observability, especially where microservices, high transaction volumes, and large engineering organizations are involved.

Links Referenced

  • OpenTracing: https://opentracing.io/
  • OpenTelemetry: https://opentelemetry.io/
  • Twitter: https://twitter.com/el_bhs
  • Email: bhs@gmail.com
  • This podcast: http://ScreamingintheCloud.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud. I’m Corey Quinn. This promoted episode is brought to you by LightStep and, as a result, I am speaking with Ben Sigelman, founder and CEO at LightStep. Ben, welcome to the show.

Ben: Hi, Corey. Thanks for having me.

Corey: You have an interesting backstory. We’ll get to the whole modern LightStep story, but originally, some folks are born in the cloud, you instead were born at Google. You were a co-creator of Dapper, which is my understanding their internal distributed tracing system, and you’ve done a lot of open source work, too, OpenTracing then OpenTelemetry, both part of the CNCF, so you’ve been focusing on the monitoring/observability/don’t ever get those two words confused movement for a while now. What’s your backstory? Where do you come from?

Ben: Well, my mother and father looked a—no let’s see, where did I come from? I was in college, and I started off freshman year with all of the seniors getting a thousand job offers in 1999. And then I graduated in a very different environment, and all the internet busts had happened and things were looking a little grim. And I just barely was able to get an offer anywhere, but I was lucky to get it from Google. And I went there and worked on ads originally, which I, frankly, didn’t enjoy at all. My first couple of months there, I was pretty unhappy actually. I didn’t like the work I was doing, I didn’t like the product. And then I had a meeting with this woman named Sharon Pearl, who was, at the time, she was working on five or six different computer science research projects. She had come over from Digital Equipment’s Research Lab along with a bunch of the other old-guard people at Google, like the first 100 employees. And she was super, super, super—well she is super, super, super smart.

And I just—we had this, literally this serendipitous, completely random conversation, and she asked me what I was doing. And I said I didn’t like it that much. Asked her what she was doing. And she rattled off a list of several projects. One of them, I remember was this distributed blob store, kind of like a S3 or GCS type of thing. There’s another one that was a global identity management system for all Google end-users, etc. But there’s this one called Dapper that she had prototyped with Mike Burrows and Luiz Barroso, who also came from these research labs in the late 90s, early 2000s. And it just sounded fascinating. Unfortunately, it wasn’t done, so no one could really use it, but they’d realized that it was possible to trace requests across what we would now call microservices at Google. They didn’t call it that, but you could watch a single request go from a web user all the way down through thousands of services and back to the user in 100 milliseconds, or whatever it was. And I just thought it was fascinating. And at the time, my direct manager had 120 direct reports. I don’t mean that his organization was 120 people. But he had 120 direct reports, one of which was me, which is to say he had no idea what I was doing, because how could you? And I just started working on Dapper instead, and I thought it was awesome. And it started to work, actually, I got it to work in pre-production environment, and built a team around it, and then put that into production. And that was 2005, and I’ve just been pretty mesmerized by this overall problem space, and I don’t think that’s ever really going to let up, so I just keep on working on it.

Corey: It’s strange, in that I had the privilege of working with a Google VP many years ago who had left Google and was talking about some of the same principles of tracing. Specifically, every system should expect a event identifier in it, and if it doesn’t wind up getting one of those, first it should add one, and secondly, it should raise an exception, so that that can get caught as the fact there’s something that is not participating in this event tracing system. Now, what made this unique at the time was this was circa 2011 or so, and we all looked at him like he had just grown a second head, because how big do you think this website is, buddy? Maybe that’s fine for Google. But here in reality, that’s not how the rest of us tend to conceptualize these things. Well, then we went into a microservices direction, which turns every outage into its own version of a murder mystery. And now having something like that is no longer optional for reasonable troubleshooting perspectives. It’s sort of suffered on some level from the curse of being too early to the market. It seems like you folks are right on time.

Ben: Yeah, at the moment, it does seem that way. When we started the company, that was my biggest concern, actually, I wasn’t worried about whether this would be necessary, but I didn’t know when, and I think in retrospect, our timing was right on target. There are other products that came before LightStep’s that were in a similar vein that I think we’re actually too early, that started four or five years earlier, and they built a great product, but all that you could install it on was like a PHP website or something, and it’s just, not like a Facebook kind of thing, but just like a blog, and it’s just you don’t need distributed tracing to manage your personal blog. So I think we did get the timing right on the nose, but frankly, that was an accident, [laughing], just good luck.

Corey: One of the things that I’ve always found to be a challenge for the distributed tracing set has been in trying to articulate the value of what you do. For example, I’ve gone round and round with this, with the Honeycomb folks in previous venues and different folks. And I know, for example, that you are legitimately in this space because whenever I refer to you as being observability-focused, Charity Majors doesn’t punch me in the face. So, first, you have the Charity-not-screaming-at-people seal of approval. So good job on that, this is legitimate, not a branding exercise.

Ben: I’ll put it on my tombstone. Yeah, Charity is my best frienemy in the industry, we obviously compete at some level, but I think Honeycomb does great work, and she doesn’t suffer fools, so I’m glad that so far she hasn’t called me out, or something like that. Yeah, to be honest, I don’t really like positioning LightStep as a distributed tracing company. That’s not really how I think of our mission, or even really exactly our product. I think our technology under the covers certainly is all about distributed tracing. But that’s, in my mind, an implementation detail. We do see a lot of people in the market that have heard about distributed tracing, can recognize at some instinctual level that being able to follow requests across services is going to be part of the solution, and then they start looking for distributed tracing. And, frankly, if someone comes to our door and says, I want to buy tracing, and we have a tracing based solution, we can sell them that product, and I think there’s a lot to be said for that. Both parties benefit from that, but it’s not really the way I think about the space, and I do think that for distributed tracing going forward, it’s important that we talk about what it does and not what it is, if that makes sense.

The point of distributed tracing, for me, is just to satisfy the same old requirements we’ve always had for observability or monitoring before; we need to deploy our software faster, we need to understand why it’s performing slowly and where, and then we need to reestablish regular performance if there’s been some kind of emergency. Really, those three things are the driving forces behind every observability product, and tracing just happens to be necessary if you want to do any of that in a distributed environment, like microservices or serverless.

Corey: One of the challenges you have is that historically describing what it is—an offering like this does—presupposes A) that someone has a extensive background in running large scale applications, and doing that in a very public, very global fashion. And secondly, that they have spare 45 minutes to sit there and listen to the in-depth exposition that describes what the heck your thing does. So has that problem gotten better? In other words, is it easier to describe to people today what you folks do, then it was a few years back?

Ben: It has gotten a bit easier. I think you hit the nail on the head, though. We were chatting before we started recording and I was explaining that I have no interest in turning this into a product pitch. And this question, it risks us going in that direction, which I really don’t want. But part of the reason that’s gotten easier is that products like LightStep’s product, they solve problems, right? And I think it’s much easier to explain these problems to people when they’re actively feeling a lot of pain around them, as opposed to it being a theoretical exercise. I think before people moved to microservices, we could draw diagrams of, this is what it's gonna look like in a year or two years, when you make all these transitions and when your system is distributed and ordinary transactions no longer exist in only a single host. It was a theoretical exercise at that point.

Now, it’s a much more visceral thing where we can say, “Have you ever had an experience where you have two teams shouting at each other because they can’t decide which one is the root cause of the problem? And they both have dashboards saying they’re healthy, yet it’s clear that one or the other is actually responsible?” or, “Have you ever had Kafka just totally on fire, because you have one of the 10 tenants is suddenly sending more traffic, and you can’t figure out which?” Or, “Have you ever had a situation where you’re dealing with a P0 emergency, and the one person who actually understands how to debug it is on vacation?” These sorts of things are symptoms of microservices and deep multi-layered systems, and once people can identify those problems, it’s much easier to say, “Well, let me explain how the sort of technology we’re bringing to bear is relevant to those problems.” So it has gotten easier over the last couple of years, frankly, because the level of active pain has gone way, way, way up with I think that the credible migration towards more distributed architectures in the last couple of years in ordinary mid-market enterprise companies.

Corey: Well, let’s go back in time a little bit, if we may, to originally, I don’t believe LightStep was aimed at the monitoring/observability space at all, to my understanding you were something of a social media company, and then you had one a heck of a hard pivot. Did it turn out that you just had—you sucked at telling a social media story, and then well, we’ve raised this money, we may as well do something fun, because we’ve made ourselves unemployable along the way, or is there something more nuanced to it?

Ben: That’s, yeah, it’s not a well-known fact. It’s not a secret, it’s just—feels, I don’t know, it’s not something that I expect people to ask about, and I forget to tell them. So LightStep, per se—when I left Google, I really had a bee in my bonnet about Facebook actually, as a product. I thought it made people miserable. I actually still think it makes people miserable, and the observation was that most people, certainly including myself, are complicated, and if you compare your inner experience as a human being to other people’s carefully curated vacation photos, it doesn’t feel very good. And this is, at this point, a well known—it’s almost a trope at this point, but when I left Google in 2012, that wasn’t as well known. I wanted to create a social media product where people were encouraged to be more candid and to be themselves, and then we would connect to each other—they could, I don’t know, find some common ground. The funny thing about the product—so I managed to raise a seed round around that idea, which I’m forever grateful for, I mean it was a really fun thing to build. And I had a very small but very high-quality team, and we built a prototype of this, and got it out on the app store and so on, and the surprising thing is one, it actually kind of worked.

Like, we had a bunch of people that love this product, they really loved it, and you’d say, “What do you think of this product?” And they’d say, “Oh, this is the most important app on my phone. This has gotten me through really hard times, that sort of stuff, which is great if you’re building a social product to have that kind of zeal.” And then you ask the same people, “Okay, well, who would you tell about this product?” And they would say, “I would never tell anybody about this. It’s way too personal, way too private.” And so I—after about a year, after having the product in market, I decided that we had built almost exactly what I wanted to build. The vision had been achieved, and it was a total failure.

The people that we retained, were, I would say 90 something percent of them were depressed introverts, and I love depressed introverts, a lot of my friends are depressed introverts, but they’re a terrible, terrible audience if you’re trying to build a viral social media product, they just won’t talk about anything. Like it’s impossible to get them to talk about it. So, at that point, I told the investors—I wrote a board deck that one of them is actually anonymized and used with her other portfolio companies that won’t admit that they’re failing. And I basically explained why this is never going to work. And I said, you can have your money back if you want it, but like I’m done, because I’m not interested in running out the clock, or I’m gonna do something totally different. And I do think it was relevant because prior to that I had been working on this observability type stuff at Google. And I actually really enjoyed it intellectually but had this idea that I wanted to work on something that was more meaningful to society in a direct kind of human to human way. And that experience building that social product was quite sobering for me. First of all, I’m really bad at it. I mean, really, really bad at it. I think my intuition around what’s going to work, what’s not going to work is not that good compared to how it is in other areas, like in the [inaudible 00:13:45] realm. Second of all, I think to win in those games, you have to play dirty and I don’t like doing that.

And the funny thing about enterprise software is that when people are paying significant amounts of money for a product, they don’t just take your word for it. You actually have to deliver value, and it kind of goes back to high school economics, where it’s a mutualistic thing for all parties. A vendor can exist by amortizing the cost of developing something really powerful across many customers, and the customers win because they could never build something like this, or if they did, it would cost way too much money. So it’s this thing where you have this really clean, honest sales process, and at the end of it, both parties feel like they’re winning, because they are, and I find that, after working with consumer, which I thought was frankly kind of depressing, I found it to be a huge relief. And the reason that we’re working on this is just that it’s an area where we think we’re contributing actual value in a way that’s tangible. Like, you can tell that you’re doing something valuable because people pay for it and they want to renew year after year. And to me that’s a more validating feeling, then trying to sneak a couple seconds of people’s time, while they’re in the bathroom or something like that, which is like how it felt on social media, frankly. It just wasn’t that gratifying when it was all said and done.

Corey: So a common problem that you’re going to see with a lot of companies that trend into the monitoring/observability/yelled at by Charity Major space is the propensity to wind up going broad, where, yeah, today you do, for example, distributed tracing. Tomorrow, you’ll do log analytics, the next week, you’ll do alerting, and suddenly you’re trying to be Datadog, Jr., but we already have a Datadog. And as you look at these companies continuing to expand to all of these different coverage areas, it becomes very challenging to differentiate any of them other than that one area that they excel at. First, you haven’t done that, so how have you avoided it? And secondly, what do you think drives that?

Ben: Well, we’re not as old as Datadog, so one way to not to do it is it’s not to be around long enough, right? But there’s also why we wouldn’t do that. I don’t want to throw too many stones at Datadog, either I mean, they’ve obviously built a—

Corey: To be clear, I’m not trying to insult Datadog with that comparison. They’re fantastic, but they are the best of breed in this space. So everything that’s the newer generation trying to become the next coming of Datadog, well, why? I can see the story around individual components being awesome. What I’m not loving is this idea that everyone needs to be a broad platform for all use cases.

Ben: I think the problem in my mind, it really comes back to what I was saying earlier about whether LightStep is a distributed tracing company. Again, we use distributed tracing, and it’s the core of what we do. I do not think of us as a distributed tracing company. Nor do I think the problem that we solve—the problem we solve is not distributed tracing, or it really shouldn’t be. And I don’t consider it to be dishonest, but I do think it’s confusing for the market to have large vendors, whether it’s its Datadog, or Splunk is also acquiring their way into similar position, right? I don’t think it’s helpful for the market to position the problem space in terms of these technologies. Not just because it’s confusing, because tracing is not a problem, it’s a solu-, it’s not even a solution, it’s just a technology, right? Like, it solves nothing on its own. It’s partly that and, I think more importantly, that you don’t want to solve problems by having three or four different tabs open. Like having a tab open to the logging, and tracing, and metrics portions of some product suite is not a useful workflow. In my mind, the things that people are trying to do on a day to day basis are to deploy software with more confidence, to improve performance, and/or to recover from errors with haste, like some kind of on-call firefighting workflow. That’s like—you know that, of course, we can drill down into the ontology below that, but those are the main things you’re trying to accomplish if you’re actually an end-user of say, Datadog, and my issue with the Datadog’s product strategy is not so much the accumulation of all these different data sources, which I think actually is totally reasonable, but the fact that they’re positioned—

Corey: Oh, that’s what you want if you’re a Datadog customer, absolutely.

Ben: —right, but they’re positioned almost separate SKUs. In some cases, they literally are separate SKUs you pay for separately, but they’re actually not integrating that data from a workflow standpoint in a way that I feel like is very beneficial for their end-users. I think because it’s hard to do that, it’s not that they don’t want to, I think if you watch their keynotes and stuff, I think that’s what they’d like to do, as well. But they haven’t been able to do it because there’s too much gravity and too much velocity around the products they do have for them to execute on that. So I think the way you become—let me say one more thing. I definitely hear vendors talking all the time about building this platform or that platform. When you go and talk to buyers, especially at larger organizations, nobody is saying that we want to have one vendor for everything. I talked to a buyer at one of the major investment banks once, and he was saying, we already buy 57 different monitoring products, so don’t tell me you’re gonna sell me one tool to rule them all. Of course, I asked him why he couldn’t buy 58, right? Like, that’s the—

Corey: Oh, that’s the question you got to follow up with.

Ben: —right, exactly. But seriously, it’s not—maybe it’s some fantasy level, they would like that, but it’s completely unrealistic because you’re dealing with maybe four or five different generations of application technology. And so, if nothing else, Datadog is great, but it doesn’t really do a lot for your mainframe, right? I think you’re going to, at the very least have to integrate generationally, and then I think you also have to do some integration across different business units and pieces of the org that buy different tools for whatever reason. And so, the one platform thing isn’t something I really hear from the market as much as I hear it from vendors, for what it’s worth. Now, going back to the heart of the question, though, I think that the only answer, in terms of the company that wins in any of this stuff, whether it’s LightStep or someone else, is to actually to focus on specific jobs to be done and to try to solve them end-to-end within a single tool. I do think that it’s necessary to bring other forms of data to bear on that problem, which is why LightStep, frankly, by the end of the year, I don’t think will be thought of as a tracing company, per se, as much as we are right now. I do think other forms of data are necessary.

But I think it’s a mistake to position the product as a series of modules that are tied and tightly coupled to specific types of data. For instance, metric data is mostly totally unused and totally useless in any given investigation. There’s a very small subset of metric data that’s actually relevant. In order to figure out what subset that is, you have to, I don’t mean you should, but you have to be able to understand the relationships between different services on a per-transaction basis. The only way you can do that is to look at traces in the aggregate. So in my mind, the tracing data, it’s not the product. And LightStep’s product, you spend a very small amount of time actually looking at traces. You spend a much larger amount of time looking at statistical aggregates drawn from those traces, either directly or used to inform the interpretation of other data, whether it’s metrics, or logs, or whatever. So in my mind, the user needs to tell you what they’re trying to do. Are you trying to deploy software? Are you trying to solve a performance problem? Are you trying to resolve some kind of page? That context is enough to take all of this telemetry data, which is what we’re talking about here: metrics, logs, traces, etc, are telemetry data, that context is enough to interpret the telemetry data in a way that doesn’t require some kind of advanced degree or dozens of years of experience with tracing systems. And I don’t think that Datadog’s strategy is a bad one from a data acquisition standpoint, but from a product standpoint, I think it ultimately, it leads to a really fragmented end-user experience. And I find that to be kind of problematic. So that’s, that’s how I see it. And LightStep’s overall strategy is just to focus on specific workflows and to be the best at that. That’s what we’re trying to do, and not to be terribly distracted by the portfolio of telemetry data types that various other companies are integrating or acquiring through partnership or otherwise.

Corey: And I think that that’s a very fair assessment. To be clear, I have no problem at all with Datadog providing this. It’s something that is, I think, the right move. The problem I have is that you see so many companies that specialize in one thing, and they’re founded and they do that one thing so well, that then it feels like they’re suddenly veering into that we’ve got to do everything story. And for example, LightStep does a phenomenal job of effectively instrumenting observability into microservices applications, I don’t know that you would necessarily do nearly as well with log analytics, for example. The idea of—it’s the loss of focus on the one differentiating factor that makes everything work. It’s the same story is why no one has ever bought a multifunction printer that they liked. It’s, do one thing, do it well, and leave the rest for other folks.

Ben: Yeah, I think I was talking to, this is early on in LightStep when we were just in some customer discovery mode, we didn’t really have a product, but I talked to someone who bought—well, I’ll just say it—they bought New Relic. This is like in 2016 or something. And I asked them if they liked it, and they’re like, “No, not really.” And I was like, “Well, why do you buy it?” And they said, “Well, it’s B minus at everything.” And I think that was supposed to be sort of a good thing, right? It’s like, they didn’t need to buy—they wanted to have fewer vendors, not more vendors, and that helped them in that goal, but they weren’t particularly enthusiastic about it. And to your point about LightStep, if someone has a three-tier app like if they’ve got some Java app sitting on top of Oracle, and that’s the whole application, we would immediately walk away from that. You said log analytics. To me, it’s less about that particular data type and more about the architecture. LightStep is very focused on organizations that have incorporated some microservices. I mean, 100 percent of our customers still have a monolith as well, but the point is that they’re actually doing microservices in some capacity, and that opens up a whole set of problems that we’re designed to address. So, I think it’s funny when vendors claim that they’re the right thing for everybody. Everyone should be focused on a particular part of the market, and that’s the part that we’re focused on.

Corey: And I think that that’s very fair. Now, what makes this interesting is you’re also involved heavily with the OpenTelemetry project, which is a CNCF open-source project. Do you find that that is a, I guess, either a diversion of focus, or a conflict of interest, given that you have a private company that’s aimed at something that is very similar, if not identical, from a naïve third-person point of view?

Ben: I really don’t think that LightStep and OpenTelemetry have that much overlap, actually. Any of these companies—or open-source projects, if you were to look at things like Jaeger, Prometheus, that sort of stuff—the problem can be segmented pretty neatly between the acquisition of data, which in this case is telemetry data for observability, and the interpretation and analysis of that data. LightStep has long believed that the acquisition of data really should be something that’s done in the commons. This has a lot of benefits for customers in that integrating with LightStep or anything else that supports OpenTelemetry is a matter of binding yourself to a portable, non-vendor-specific, I don’t want to use the S-word, standard, because it’s not like an IEEE thing, but the point is that it’s a portable framework that you can use to integrate any number of potential downstream analytical tools. That decision is completely decoupled from which of those analytical tools you want to use.

I think the fact of the matter is that OpenTelemetry doesn’t actually do anything. It doesn’t present you with a UI. It just gathers data in a way that is vendor-neutral. And that’s the point of that project. The reason that LightStep pursued that is partly just, frankly, trying to be customer-focused. That’s what we think is best for our market. And so we want to bet on that technology. And then the other piece of it was that, if you go and talk to people who worked at New Relic and AppDynamics, and their glory days when they were ascending very quickly, they’re spending like 80 or 90 percent of their engineering resources on agents, which are actually not even differentiated anymore. For a while there, APM agents were the thing you really were buying, and then the analytical tool was pretty basic. That’s kind of flipped over at this point, where the analytics are getting much more elaborate, mainly because of the rise of deep multi-layered systems, and microservices, and so on. But the collection of telemetry data, people expect it to be automatic at this point. No one has patience for manual instrumentation.

And the idea, with things like OpenTelemetry and auto-instrumentation OpenTelemetry, is to make that shared responsibility of everyone who’s trying to do observability rather than having every vendor repeat the cost of building all that in a proprietary way. Because that’s how things were several years ago. And the irony of all this is that the vendors that did that work, they’ve spent a lot of muscle marketing those agents, but privately, they’re very excited about getting out of that business. It’s not differentiating for them, and it’s still a big cost center. It’s not something that their customers really benefit from anymore and it takes up a lot of the resources. So there’s pretty wide alignment around the value of something like OpenTelemetry, and with so many different competing vendors involved in the project—at this point, I think they’re like eight or nine of us—it’s difficult for anyone to kind of run away with the ball. there’s a governance structure and so on. So, I don’t think there’s much of a risk of it turning into a mechanism for any one vendor to win or lose. The main thing I see is a potential for us to have some kind of rising tide that our customers benefit from as well. And that sounds like B.S., frankly, but I mean every word of it. [laughing]. You’re welcome to call me on it.

Corey: No, no, I accept that. One thing as well, that I think has been extremely valuable from the perspective of looking at LightStep and understanding what it is, is the interactive sandbox that really takes you by the hand through using it in a production style environment. It’s handy to help folks really understand what it is, it sort of walks around the, “How the hell do I describe this and an elevator pitch to someone without the same level of experience.” But it is useful and, credit where due, it only demands an email address and not a 15 field form to start playing around with it. So, if people are curious about what LightStep does, I would encourage them to go and take a look at the interactive sandbox at LightStep.com.

Ben: Yeah, I think that’s a good idea, mainly because people often ask us to describe what we do and how we’re different, and we can describe it in words, but we realized what people want to understand is how it’s useful, it’s not how it’s different, and the sandbox environment allows you to walk through scenarios like deploying software, or finding that the root cause for an error or performance anomaly in a way that lets you do all the clicking. And you can explore the whole product from there if you want, but it gives you some guardrails just to actually solve the scenario. And people have found it to be quite educational, I think, just in terms of how we built it, and we’ve actually had folks from much larger organizations, like the kind of Googles and Facebooks of the world have been using it as well, to help develop their own internal approach [inaudible 00:29:22]. So, even if you have no interest in LightStep, I think it’s still a worthwhile thing to check out, because a lot of the stuff that we’re showing in there, I think, is actually somewhat novel and just kind of fun. So, many people have told us, that’s been a really helpful thing for them just to understand the space better, independent of LightStep.

Corey: Excellent. Well, Ben, thank you so much for taking the time out of your day to speak to me today. If people want to hear more about what you have to say, where can they find you?

Ben: Yeah, you can find me on Twitter. @el_BHS, like the Spanish article, el BHS, and you’re welcome to look me up on the internet and send me email, or whatever. I’m actually pretty good about responding to that. And of course, LightStep is at LightStep.com, and the sandbox really is the best way to understand what we do if you’re an engineer, but you’re always welcome to reach out to any channels to me if you want to provide feedback or ask any questions or any of that stuff. I love talking with folks.

Corey: Thank you so much for taking the time to speak with me today, I really do appreciate it.

Ben: Thank you, Corey. It’s been really fun.

Corey: Ben Siegelman, founder, and CEO of LightStep. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on Apple Podcasts. If you’ve hated this podcast, please leave a five-star review in Apple Podcasts, and make sure to instrument it appropriately so that we can trace where it entered and exited.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Farrah Campbell

After 10 years of working in healthcare management, a serendipitous 20-minute car ride with Kara Swisher inspired Farrah to make the jump into technology. She has worked at multiple startups in many different capacities, eventually working her way to being the Ecosystems Director for Stackery in Portland, Oregon.

As the Stackery Ecosystems Director, Farrah has managed the Stackery relationship with AWS including Stackery as an Advanced Technology Partner, achieving the AWS DevOps Competency, a launch partner for Lambda Layers and is an AWS Serverless Hero. Farrah has cultivated the serverless community as an organizer of Portland Serverless Days, the Portland Serverless Meetup, along with numerous serverless workshops and the Portland tech community events from Techfest to bringing multiple luminaries to Portland.

Links Referenced

  • AD: DigitalOcean
  • AD: DataStax
  • Portland Serverless Days
  • Portland Serverless Meetup
  • Twitter: @FarrahC32
  • LinkedIn: https://www.linkedin.com/in/farrahcampbell/
  • Email: farrah@stackery.io
  • Personal site: https://medium.com/@FarrahC32
  • Company site: www.stackery.io

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey Quinn: This episode is brought to you by DigitalOcean, the cloud provider that makes it easy for startups to deploy and scale modern web applications with, and this is important to me, no billing surprises. With simple, predictable pricing that’s flat across 12 global data center regions and UX developers around the world love, you can control your cloud infrastructure costs and have more time for your team to focus on growing your business. See what businesses are building on DigitalOcean and get started for free at do.co/screaming. That’s D-O-Dot-C-O-slash-screaming and my thanks to DigitalOcean for their continuing support of this ridiculous podcast.

Corey Quinn: Welcome to Screaming in the Cloud, I’m Corey Quinn. I’m joined this week by Farrah Campbell, ecosystems director at Stackery. Farrah, welcome to the show.

Farrah Campbell: Hey, thanks, Corey. Thanks for having me.

Corey Quinn: So, let’s start at the beginning with what is an ecosystems director?

Farrah Campbell: [laughing], well my job, basically, is to connect with people across AWS serverless ecosystem, so that we can increase more serverless adoption. Making connections with partners, customers, community members, and then strategic partners like AWS.

Corey Quinn: So, it sounds like you’re viewing this as something that is distinct from, effectively, community. Where’s the dividing line?

Farrah Campbell: That’s a good question. I think right now, it encompasses both. It’s a new technology—well, it’s not necessarily so new now—but I think that—I feel like I kind of represent the community, but also our customers and our partners are part of that community. I need to be wherever they’re at, focusing on making sure that we’re in this together.

Corey Quinn: So, one thing you mentioned a minute ago was that you’re focusing on serverless. And then in almost the same breath, you tied it back to AWS.

Farrah Campbell: Mm-hm.

Corey Quinn: Do you view serverless as being primarily AWS driven? Is that your area of focus? Is that just shorthand because AWS is such a, I guess, gargantuan presence in the market that it’s easier to contextualize it there? Where do you stand on that?

Farrah Campbell: Well, I guess I should back up a little bit and say, working at Stackery, Stackery is—definitely—everything we do is built on top of AWS. We speed up the application development, delivery, helping teams to securely deploy and build these serverless applications. At this moment, we don’t have a system that works with Azure or Cloud Run. And so, when I talk about serverless, I guess I talk about the people that are right in front of me. [laughing].

Corey Quinn: I’m right there with you, you may want to back away from that, I want to charge directly into it. The reason that I fix AWS [builds 00:03:25], it’s where the big expensive problems are. When I’m doing serverless development work on a bunch of different ridiculous things that I then throw up on Twitter for mockery, I find that I’m always tripping over AWS things in the wild. When I just Google serverless, and then whatever I want to do, URL shortener for example, I’m not seeing a whole lot of options that aren’t built on top of AWS, for better or for worse. In my experience, that is where the community tends to circle around. Even when I step outside of my, I guess, bubble, I’m still seeing that AWS is definitely carrying the torch in this space.

Farrah Campbell: Yeah, I think so. But there’s a lot of good people that are doing good work at Microsoft and Google that actually I wish I could work more with, but we’re just not there yet.

Corey Quinn: Of course, I’d even take it a step further and invoke the great Satan, in this case, the Oracle folks when they wound up acquiring a lot of the iron.io people, they had a fantastic technology and a fantastic story. And it’s just overshadowed by Oracle’s business practices, so, for better or worse, I don’t get to play with that stuff in any meaningful way. Yeah, it’s always about the—when I talk about companies and I guess, throwing muck, it’s never about the individual people. It’s about the corporate cultures, by and large. I sometimes do a poor job of articulating that, but, “Oh, I work at a company he hates, therefore he must hate me,” is never accurate. It tends to not—separating the person from the employer is challenging at times, but I sometimes feel like I stumble over doing that well.

Farrah Campbell: I think you do a pretty good job. I’ve never felt like you were targeting a specific person. I think that when you have, it’s always to give them Happy Birthday videos or something cool.

Corey Quinn: Yeah, the only exception to that rule that I have is Larry Ellison, because he’s not people. He just thinks these people and he has no friends or people who love him to take offense on his behalf. So, therefore, it’s a free target and it’s never punching down.

Farrah Campbell: The funny thing about Larry Ellison, in my past insurance life when I was a broker here in Portland, Oregon, his ex-wife, Barb Ellison, who owns a—it’s called Wild Turkey Farm, which is a horse studding ranch, she was one of my clients.

Corey Quinn: That is amazing. And thank you for naming her because it is Larry Ellison, I would have to ask you which ex-wife, there are three to choose from. But I digress. Sorry, this is not a dump on Larry episode, though it’s somehow turned into that, but all right, moving on. So, how did you get to the place you are? I don’t think I’ve ever been talking to companies who are looking to hire and they say, “You know, what we really need is an ecosystems director.” That feels like a role that emerges, rather than something that is actively recruited for. And maybe that’s just because I’m not looking in the right places. But how did you get where you are?

Farrah Campbell: Well, how did I get where I am? Well, Nate created this role for me, here, at Stackery. I had known him from the community. I’ve worked at a number of startups here in Portland. In fact, actually, I got my start, because when I moved to Portland, I knew I wanted to do something different and I was trying to get a job. It took me about—well, for a year I applied, got no responses. I ended up volunteering at a conference called TechfestNW, and I ended up being the speaker handler. Well, my first job was to pick up Kara Swisher. I didn’t have any idea on who she was, and they just told me she was a very important person. Anyhow, I was super incredibly nervous that entire time, but it was actually pretty amazing because I had the opportunity to talk to her, and she actually inspired me to quit my full-time job and take this startup that had got in front of my face, which was four hours a week.

Corey Quinn: That is a heck of an origin story.

Farrah Campbell: [laughing].

Corey Quinn: I feel like Kara, based on my impression, I’ve never even met her myself, would argue that everyone’s important, but she’s one of those inspirational people that I really love just seeing almost everything that she does. Even when I don’t agree with her, I just love the way that she frames and does things.

Farrah Campbell: Yeah, she’s pretty awesome. She even sent me the Steve Jobs commencement address and he told me to watch that, and then she like, “Next time I hear from you, I want to know what you’ve done to try to get into tech.” And it was the next day I had this offer. It was pretty crazy. I’m incredibly thankful to her for that. She’s not that lady behind the glasses that everybody thinks.

Corey Quinn: No, it’s very clear just following her on Twitter that she has definite personality that she does not bother to hide, which from my personal perspective is just something that I adore. I spent my first part of my career trying and failing to hide my personality, as anyone who’d ever met me can attest, and finally, this thing turned into me embracing my crappy excuse for a personality. And here we are. I scream into microphones about cloud computing.

Farrah Campbell: I’d say you have a pretty good personality, Corey. I really love that picture with the motorcycle helmet on the horse.

Corey Quinn: Yes, the dating profile photo I had when I met my wife. So, she knew exactly what she was getting into 11 years ago now.

Farrah Campbell: I mean, that takes a lot of personality.

Corey Quinn: Yeah, that was an experience, let’s put it that way it’s been… I tend to be, if nothing else, refreshingly direct. So, something else has happened recently that I wanted to get your thoughts on or as much as you can tell us; you’ve become an AWS Serverless Hero, and Hero of course, is the name of their program, which I’ve always found to be kind of weird. It feels like “hero” or “entrepreneur” or a few other things are things other people call you. Calling yourself that sounds weird. “Thought leader” is similar in that space. But I’m curious as to what it was like to, first, become a hero. Secondly, what does being a hero mean?

Farrah Campbell: So, what is being a hero mean? Actually, I guess, heroes are early adopters, spirited pioneers of the serverless ecosystem. I think that the announcement of me being a serverless hero was pretty—I thought it was impactful because it shows that, I guess, experts in the field, that it requires more than just the ability to code. It requires the ability to have communication and to lead with empathy. And I think it’s important that they’re helping to change the way that we’re looking at leaders in our industry.

Corey Quinn: One of the, I guess, I love the terminology, if for no other reason than I heard that they had an AWS Heroes program, and I immediately set out to become an AWS Villain instead—

Farrah Campbell: [laughing].

Corey Quinn: —and I’ve mostly gotten there. If you look at some of the things that I’ve celebrated, even on this podcast, I’ve had a couple of, as I termed them, code terrorists who have shown up and demonstrated horrifying architectures and patterns to achieve reasonable outcomes. But by and large, I find that the folks in the Hero program tend to be, at least from the outside, never having been one myself, those are the folks that we hear an awful lot from in their respective subject areas. They, more or less, are the torchbearers for the rest of us who are stumbling our way, more or less blindly, through a lot of these things. And when you have people who can be held up and demonstrated as this is one possible way to do it, here’s someone willing to tell stories about this, it becomes incredibly compelling and for whatever reason, those are folks who can tell stories with a greater degree of honesty and transparency than you’ll often get from either a marketing apparatus, or a developer advocacy, or developer relations program because none of these folks work for AWS. They’re all community members or work at other companies in the space, but no one is carrying a quota from Amazon of, you must get at least X people to sign up or you’re not a hero anymore.

Farrah Campbell: No, in fact, I haven’t seen any requirements at all, to be honest, and I wasn’t setting out to be a hero. I never ever imagined that I actually would, just because if you look at everybody’s profiles, everybody’s written books or had all these open source contributions, tons of medium posts, and had just had a lot more experience than me when it comes to building applications. I guess it reminds me of having that serverless mindset and focusing on what matters, and I feel like that’s kind of what my career has been since I’ve gotten to tech. My life was always—had to be like that, I had to focus on what matters just because I was a single mom, but it’s pretty powerful when you can start to translate that into your work and see growth and affect change.

Corey Quinn: Can you tell me a little bit more about that? Having a single child right now with my spouse and I, that seems like an impossible amount of work. I can’t fathom what it would be like, having to do that alone. But how does that map directly into the work side?

Farrah Campbell: From very early on, I was the queen of making poor decisions when I was younger. I ended up divorced and had two kids and I was 24, and on my own in a place I didn’t have any family near me, and I guess my relationships were suffering, too, with my family. But I had two little guys to take care of, and so, when it came down to it, everything I did, the time that I had, I needed to be focusing on something that mattered. And so, I even think about that even to—day-to-day being a single mom. If I have a number of things to do, if I can outsource housework or picking up groceries, whatever it is, time is important. I made sure that we didn’t have a lot of TV, I wanted to make sure that we had a lot of face to face time in those moments that I did have that time with the boys. And I think taking that and applying that into your work life, there’s a lot of stuff that happens in our day to day that’s just noise. And if I take that time to focus on that, I’m now wasting time on things that don’t matter, and things that actually might take me to a place that’s actually not productive. I guess that’s how I would say translates from being a mom, applying that mindset of focusing on what matters in your life, and then translating that over to your work. Did any of that make sense or am I just rambling? [laughing].

Corey Quinn: No, it absolutely does. There’s a tremendous amount of, I guess, gatekeeping in the infrastructure space of, “Oh, you’ve only been doing this for a couple of years? Well, unless you were doing this back in the Solaris days of Big Iron and the rest, well, you’re not really involved in the infrastructure world,” which is, first, nonsense and secondly, not usually relevant most days, when you’re talking about getting something out the door. That’s sort of the entire premise of serverless is the stuff that everyone has to learn, well, a lot of that is now handled for you. Sure, there are times where it’s very useful to have specialist insight into how a lot of those subsystems work, but you maybe don’t need that to build a quick API that does something that demonstrates business value.

Farrah Campbell: No. In fact, for somebody, I mean for me, I created a webhook, I authenticated with Google, and I was having it look for specific information through my emails that would drop into a Google Sheet and then I could then act on that. I’ve built a language translator app with my co-worker, Danielle. I just created my own website, built on top of AWS services using Route 53, CloudFront, Certificate Manager, S3, and using the Lambda function to deploy that, and I’m doing this stuff by reading these tutorials, and in all honesty, I have really no idea what these things are doing. Like Route 53—

Corey Quinn: I maintain it’s a database. I’m told it’s not but I don’t even care.

Farrah Campbell: [laughing]. So, but I can get these things all up and working, so one of the things that I hope to help people do is just take a step back and forget about all the things that you know. We’ll bring that back later, but forget all the things that you think you know, and then just go try it, and see what happens. You’d be amazed but, with all the information that they must have after building systems, how they can apply that to building new systems, what they would be able to do.

Corey Quinn: This episode is sponsored in part by DataStax. The NoSQL event of the year is DataStax Accelerate in San Diego this May from the 11th through the 13th. I’ve given a talk previously called the myth of multi-cloud, and it’s time for me to revisit that with… a sequel! Which is funny given that it’s a NoSQL conference, but there you have it. To learn more, visit datastax.com that’s D-A-T-A-S-T-A-X.com and I hope to see you in San Diego this May.

Corey Quinn: I think part of the key to this whole, I guess, unlocking the power of serverless is, first, it’s not even that old. So, if you’ve been using these technologies for 20 minutes, congratulations. You’re one of the elder statespeople who has been using this for a long time. It’s hard to remember sometimes, but 15 years ago, there was no such thing as cloud in any real sense. So, everyone who’s been doing this since the beginning, well, yeah, they don’t necessarily need to be that old to get there.

Farrah Campbell: No, you don’t. And I’m shouting this from the rooftops because there’s a moment right now to come up and level up all of your skills. That’s what everybody is doing. Everybody else is going to be doing it in the next couple of years, three to five years later, and if you do it now, sets you up in a pretty good place. You just have to be willing to put in the work.

Corey Quinn: Oh, yeah. And there’s always time to go back and learn the underlying things that you need, when you need them. Well, sure, using Lambda, and API gateway, and that’s fine for you right now, but that’s not going to be economical and scale once you wind up with 8 million users in a typical month. Oh, heavens, seems to me that might come with some other benefits, where, huh, I can afford the time or the resources to delve more deeply into that arena when I need to, but optimizing in advance of something like that, you’re almost never going to have that happen. Why bother?

Farrah Campbell: Yeah, and I think a lot of times you hear about people just try to, like, want to rework the same things that they’ve already done, or just having a preconceived idea about why something won’t work, and they’ve tried it in the past, or they had used it for some part of application it didn’t work, or maybe they were on a completely different team. That doesn’t matter. It doesn’t mean it doesn’t work for anything else ever again. There could be all kinds of factors that were a part of the reason why it didn’t work at that time. But anyhow, I think a can-do attitude can enable a lot specifically for developers working in the serverless ecosystem.

Corey Quinn: Something I’ve learned about serverless is that I’ve been using it for a couple of years for different things, and I was working on a system the other day, and my immediate response was, this is awful. Whoever wrote this obviously had no idea how best practices are supposed to work. There’s at least 16 errors I can see on this. What moron did this? And of course, I whipped out git blame, and of course, the answer was me. And at which point I just quietly fixed it up and there was no need for me to bring that story up any further. But it always seems like no matter what the tool is, and how perfect it can be, you can always either, first, learn new things about it and makes the person you were yesterday look a little bit less bright than the person you are today. And, secondly, any tool can be used or misused. Look at me using Route 53 as a database. It’s a recurring joke for a reason, in that it’s something you can do, but probably shouldn’t, but it could actually be forced into service as I wind up evolving that joke forward. You can misuse anything, and that doesn’t mean the tool itself is bad. It just means the architecture is more than a little bizarre. I always try to provide a positive example whenever I bring that up these days, and sometimes I forget to do it, but every once in awhile I pay off my guilt for an entire generation of people listening to this who have now built CMDBs inside of DNS, but I at least hope that people are taking the right lesson from these things.

Farrah Campbell: [laughing]. Yeah, I actually kind of like—this conversation kind of made me think about my role a little bit. I started thinking about, like when we think about somebody that’s evangelizing, or even the HERO Program, when we think about evangelizing a new technology, we tend to think about the cool tricks, the new features, the integrations, but actually, like, think about evangelism is much more about that. You have to gain trust from the people that you’re talking to. You have to inspire them, and then support them, make them want to follow and to be a part of what you’re doing.

Corey Quinn: Part of what I think is so interesting about all of this is that everyone winds up in their own silos, and we see this with communities in a lot of different ways, but look at what you’re doing. You said at one point you were the queen of poor decisions. Well, now you’re currently an organizer for Portland Serverless Days, you help run the Portland Serverless Meetup, you do a bunch of service workshops, and you’re doing a bunch of other Techfest things as well. That’s an awful lot of volunteer work, which doesn’t, I guess, directly tie back to, for most companies anyway, I don’t know how Stackery is structured internally, that doesn’t confer any direct benefit immediately to the company. That’s, you’re effectively giving an awful lot of yourself to the community.

Farrah Campbell: Yeah, I don’t know that I know how to do it any other way, to be honest. When I talk about bad decisions, this is like, I was the queen of negativity, and everything around me, something wasn’t going right. I was the queen of blame and victimizing myself. And it’s been pretty amazing to see what my life has been like when I stopped that. And, again, that happened when I entered into tech, and I’m like, I’m only focusing on what matters here. And it’s amazing what a better place I am with my career, what a better place I am with how I feel about myself, with my friends, with my kids.

Corey Quinn: It’s the idea of bringing your whole self to work.

Farrah Campbell: Yeah, for better, for worse, I bring my whole self to work.

Corey Quinn: Given that you’re involved in all these different things and talking to all these different folks about various aspects of it, I have to ask, what do you think, right now, is being the most misunderstood about serverless?

Farrah Campbell: So, serverless, it’s a lot more than Lambda, or Knative. It’s a mindset. It’s an innovation strategy for enterprises. But, I also think that—the technical aspects of serverless are important, but I think that there’s a paradigm shift in engineering, and the work that has generated so many tech careers is now hidden behind this growing menu of cloud services. And I feel like that there’s this host of fears, and then a focus on negativity, instead of a focus of how to be involved and to start modernizing these systems.

Corey Quinn: I always come away from serverless approaches of being of two minds. One is that it’s a fantastic approach for greenfield development of new experiments. The other is the exact opposite, where this is perfect for modernizing legacy things. And I go back and forth as to which it’s a better fit for. I started to come away with the idea that it may very well may be perfect for both use cases, depending upon other constraints.

Farrah Campbell: I think that is totally true, and it doesn’t have to be one or the other. It can be used with containers, it can be used with on-prem solutions, or—I actually, I don’t even know the old ways actually building all these applications, to be honest with you, but that doesn’t have to be—serverless is not an end for all, like, everything has to be serverless. It’s something that I think that companies are striving to get to, and again, I see it as more of the way you’re building applications.

Corey Quinn: Out of curiosity is Jamstack tied closely to the serverless movement? I keep hearing things about it, and I thought it was something one company was doing, and now I’m not entirely even sure what it is. Have you come across that at all?

Farrah Campbell: So, Jamstack is, I think, it’s marketers that, I think that a lot of them are using serverless. I actually haven’t been a part of it in any of their meetups, but we did actually host one here at Stackery, I think a year ago, and I know a lot of great people in the space like Jason, I think he works at Netlify now. I can’t remember what his last name is, but he lives in Portland. But I think that marketers are leveraging the power of serverless. If I think about it from how Stackery is using it, I can run my own A/B tests by adding in new functions to the usage event [fork 00:25:01] trigger that our team built out. I don’t need to bother anybody from our engineering team to get that done.

Corey Quinn: Yeah, a lot of times I’ll start to see people arguing against the use of managed services or SAS platforms of, “Oh, that’s not real serverless.” I am trying to solve a business problem here not strain for ideological purity, and I tend to get relatively fed up relatively quickly, with folks who tend to let the, I guess, religious movement that they’re trying to embrace, for better or worse, overtake the actual business value the entire thing is designed to deliver for customers.

Farrah Campbell: Totally. I see it’s a technological goal for all these vendors to start deliver hosting computer power to their customers. Maybe one day, we’ll call it something else, because the name just doesn’t make sense.

Corey Quinn: That’s the problem is it’s easy to sit here and cast aspersions on various names of things. I mean, Lord knows I do that enough with AWS launches, but naming things it turns out is super hard.

Farrah Campbell: It turns out also that naming things is super important when you’re setting up your environments when you’re installing certificate tokens on your computer because you need to be able to figure out what those are later, and default, and dev, and then one that just has a number isn’t helpful. These are all things that I learned last week, so, I thought I would share.

Corey Quinn: And typos, of course, are always fun and things that matter. Computers: extraordinarily literal to a fault.

Farrah Campbell: Naming just seems to be a completely hard thing in technology, period, no matter if it’s setting up your environments, keeping your tokens, knowing what AWS account you’re building in, and naming the services, again. So, what is Route 53 by the way?

Corey Quinn: Managed DNS, that’s all it is. It’s effectively the internet’s phonebook, how it converts, www.twitterforpets.com into an IP address: 1.2.3.4.

Farrah Campbell: How does one use that for a— I’ve been, I can’t stop thinking about using it as a database and—because nowhere, when I was setting anything up in there, did it look like it could be used for that.

Corey Quinn: Ah, great question. You can use text records or txt records that you have arbitrary strings up to I want to say 4k. I haven’t tested the limits of my database lately. And then you can just query it with a standard DNS query and it returns the value. You teach your code to ask the question through DNS, take the response, and do something with it. And you’ve built fundamentally what is a database in its purest sense. Now people are going to send me letters over that description but I’ll take it. And it’s not the best service for the job. I mean if you wanted a key-value store that can scale, that’s designed for this, something like DynamoDB makes an awful lot of sense. But you can misuse DNS this way.

Farrah Campbell: Well, we did talk about people misusing technologies earlier.

Corey Quinn: And I don’t think that’s going to stop anytime soon, and I don’t think the current generation of serverless proponents and detractors invented the concept, for better or worse.

Farrah Campbell: Well, maybe Alex DeBrie can help you. We could get that migrated to DynamoDB, I heard he’s really good with it.

Corey Quinn: Absolutely, there’s got to be some sort of weird failover between DynamoDB and Route 53 simultaneously. So, if people want to wind up hearing more about what you have to say, where can they find you?

Farrah Campbell: They can find me on Twitter. I am @FarrahC32. You can find me on LinkedIn. And then you can always find me by email. I’m just farrah@stackery.io.

Corey Quinn: Excellent. Farrah, thank you so much for taking the time to speak with me today. I appreciate it.

Farrah Campbell: Yes, thank you so much for having me, Corey, I appreciate it. It’s been a lot of fun. And I just repeated what you said.

Corey Quinn: No, that’s fine. That’s called being a thought leader when people repeat what you say, you’ve made it.

Farrah Campbell: Awesome. Woohoo!

Corey Quinn: Farrah Campbell, ecosystems director at Stackery. I am Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave an excellent review on Apple Podcasts. If you’ve hated this podcast please leave an even better review on Apple Podcasts, but at least make the review entertaining.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Tobi Knaup
Tobi Knaup is a Co-Founder & the Chief Technology Officer of D2iQ. Knaup is an experienced software engineer focusing on large scale systems and machine learning. Previously, he helped scale Airbnb to millions of users worldwide as technical lead. Tobi’s research work is on Internet-scale sentiment analysis using online knowledge, linguistic analysis, and machine learning. Tobi also co-founded his first company at the age of 15.

Headshot

Links Referenced

  • Twitter Username: superguenter
  • LinkedIn URL: https://www.linkedin.com/in/tobiasknaup/
  • Personal site: https://tobi.knaup.me/
  • Company site: https://d2iq.com/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey Quinn: This episode is brought to you by DigitalOcean, the cloud provider that makes it easy for startups to deploy and scale modern web applications with, and this is important to me, no billing surprises. With simple, predictable pricing that’s flat across 12 global data center regions and UX developers around the world love, you can control your cloud infrastructure costs and have more time for your team to focus on growing your business. See what businesses are building on DigitalOcean and get started for free at do.co/screaming. That’s D-O-Dot-C-O-slash-screaming and my thanks to DigitalOcean for their continuing support of this ridiculous podcast.

Corey Quinn: This episode is also sponsored in part by logz.io. logz.io hears it loud and clear from many engineers specifically that they prefer using open-source tools for observability. They get an easier onboarding experience, a built-in community ready to go, and a lot less vendor lock-in. But what about the Enterprise-grade scale, support, and security you need to improve the performance of any given cloud environment? That’s where logz.io comes in. They offer a fully managed observability service, log management, and Cloud SIEM based on ELK, and infrastructure monitoring based on Grafana. The open-source you love at the scale you need. Sign up today for a 14-day free trial at logz.io/screaming, and for your chance to receive a free logz.io T-shirt.

Corey Quinn: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’m joined this week by Tobi Knaup, the co-founder and CTO of D2IQ, which you probably have not heard of, and what used to be called Mesosphere, which you most assuredly have. Tobi, welcome to the show.

Tobi Knaup: Thank you for having me.

Corey Quinn: Of course. So let’s start with the, I guess, burning question that at least is on my mind, if not a bunch of other folks, Mesosphere was a company that everyone in the infrastructure space at least had a vague awareness that there was that thing over there. And last year, I think it was last year, time is speeding up, the company rebranded. What was behind that?

Tobi Knaup: Yeah. So, what’s behind that is, in hindsight, it wasn’t a very good idea to put a technology name into our company name, to be honest, because, technologies change over time. And we obviously started the company, Mesosphere, in 2013 around Apache Mesos. That was the core open-source project that my co-founders and I had been using at Airbnb and Twitter. And we wanted to start a company around that to help every Enterprise out there to adopt Apache Mesos. But very quickly, we actually started helping people with other technologies from the cloud-native ecosystem. We help folks automate things like Kafka, and Cassandra, and Spark, and build these data pipelines on it. And very quickly, actually got involved in Kubernetes, as well, actually in the first year when it was announced. And so, over time, the name Mesosphere as a company name became sort of a stumbling block for us, because we always had to explain that, yes, we are the Mesos company, but we also do all these other things. We help you build data pipelines, and we help you with Kubernetes, too. And so it kind of became this anchor, and we decided, it’s maybe not a good idea to have a specific technology in our company name, and so we decided to rebrand and we wanted to pick a name that kind of expresses what we really do, what we help our customers with. And that is, we help them on Day-2, we help them be successful on Day-2 and be smart about the day to operations. So, Day-2 in the sense of, the DevOps concept of Day-2, so the ongoing operations and maintenance of production systems. So that’s what’s behind that.

Corey Quinn: You always could have gone down the path that I did, where I started with a newsletter, Last Week in AWS, a consulting company that had no bearing on any of it, and this podcast, Screaming in the Cloud. There were three brands instead of one, which means that whenever anyone asked me, “So what do you do?” My answer is always “Well, it depends. Can you contextualize that question for me a bit more?” It winds up effectively having to lead us down this weird path of branding things very differently. And then, of course, I started another podcast with a completely separate name on top of that, called the AWS Morning Brief, and it’s at this point, I just sound like I’m professionally confused. Naming is hard, especially once you have a name that is no longer accurate in some ways, but it’s something that people have a definite affinity for, you have brand recognition. We had a guest on previously, from palantir.net, which predates a terrifying Palantir in the Valley by about 10 years. And it seems like their tagline is become, “We’re Palantir. No, not that one.”

Tobi Knaup: [laughing]. That’s lovely. Yeah. Obviously, like you said, naming things is hard. And renaming a company is hard, too. We built up a lot of brand equity over the years. And so, what was important to us is actually that we don’t give that up. And so the name Mesosphere actually lives on. It’s now the name of our product family around Mesos. So, name lives on, just the company has a different name.

Corey Quinn: So are you finding that—I guess it’s, obviously, from the time that you started Mesosphere back when—when was that?

Tobi Knaup: 2013. So we’re almost seven years old.

Corey Quinn: Forever ago in Internet time. There’s been some, let’s say upheavals in the infrastructure space. Back then, I would have frankly bet the farm on Mesos. It seemed like the right answer. A lot of the big shops were doing that. And today, whenever you suggest that to people, they look at you a bit strangely and say, “Yeah, if we’re doing anything net new, it’s probably going to be on top of Kubernetes, which I have a laundry list of complaints about. But I’m curious to get your take, how have you seen Mesos’s rise and fall through the eyes of what you do for customers?

Tobi Knaup: Right. So, I think what we’ve seen with Kubernetes is really the power of community. When I talk to folks and ask them, “Why Kubernetes?” That’s the thing that people most commonly mention, it’s the community in the broadest sense, meaning there’s a place online where I can go to learn about Kubernetes and related technologies. There’s a place I can recruit talent from. There’s people that want to have that on their resume. And obviously the community is so much bigger than any single vendor could ever be. And so, that’s where a lot of innovation happens. And innovation happens much faster in that community. So that’s really the most common reason we hear. Mesos started as a abstraction layer for large compute clusters. And, while we do a lot with Kubernetes now, and we have an entire product line around it, we also still have our Mesos product line. And it is still the platform of choice for those large scale deployments. So we have customers with hundreds of thousands of nodes in production, and they’re running Mesos, and they will be running Mesos for a while. So it’s really a best tool for the job kind of situation. If you’re a small shop, you’re getting started with cloud data, you have maybe a 10, 20, 30 node cluster—20, 30 nodes is where we see most clusters out there in the industry—Mesos may be not the right choice because it is built for scale. And so what we said is, “Hey, let’s offer our customers what they want. Let’s give them Kubernetes. Developers want that. And we still keep the Mesos platform for those large scale deployments.

Corey Quinn: Are you seeing net new activity around Mesos in 2020?

Tobi Knaup: We do actually. So one thing that we built, that we invest a lot of time in over the years, is helping customers automate data services. Building end-to-end data pipelines with Kafka, Spark, technologies like that. And the experience that they get around that on Mesos doesn’t exist the same way yet on top of Kubernetes. We’re working on making that happen. And obviously there’s a lot of activity around building Kubernetes operators. There’s various different approaches to building operators. We started an open-source project about a year and a half ago called Kudo that aims to make building operators very easy. It’s based on our learnings on top of Mesos. And so, the ecosystem is going to get there but it’s not quite there. And so we’re actually seeing a lot of people still start new projects around these data infrastructure projects on top of Mesos.

Corey Quinn: It’s interesting you bring up releasing open-source offerings around, I guess, anything in the infrastructure world. Lately, it seems that there’s been a bit of a pretty persistent narrative around the danger of open-source as a business model because then someone like AWS comes in and launches effectively what you do as a managed service. Is that something that’s currently on your threat radar? Is that something that you don’t see as being particularly credible? Or am I missing something entirely?

Tobi Knaup: It’s definitely on our radar. And I think there is, while this is a threat that everyone’s facing, there are also opportunities to build differentiated product for maybe a different use case for a different customer demographic. So what we see a lot, these days, is folks wanting to run any combination of hybrid or multi-cloud scenarios. So they want a public-cloud-like experience like they can get from AWS, but they want it on the infrastructure that they choose. So we see a lot of activity, we work with a lot of customers that have industrial IoT use cases. So let’s say they have a manufacturing plant, a factory where they have thousands or tens of thousands of sensors that produce data in real-time that they need to process and do things like predictive maintenance, finding outliers in the sensor data, and things like that. Those factories are often in areas where they don’t have a high-quality connection to the cloud. So it’s not feasible to send all that data in real-time to a public cloud, you have to kind of process it locally. And so essentially, what those customers need is they need a mini Edge Cloud. They obviously don’t have highly skilled cloud-native engineers in every one of their manufacturing plants and some of those people have over a hundred of these plants. So what they need is really a public-cloud-like experience sort of in a box that they can deploy on the Edge. Now, they also want to run a bunch of infrastructure on the public cloud, and they also want to run a bunch of infrastructure on their existing data centers. So how do you do that? How do you operate, pick Kafka as an example or Spark, in a consistent way across all of these platforms? That’s one thing we’re focusing on and where, yes, you can go to AWS and you can get a managed Kafka, but you can’t get it in a manufacturing plant or in an air-gaps case, right where you don’t have any Internet connection. So there’s still a lot of these use cases out there, and that’s how we differentiate, or it’s one way we differentiate.

Corey Quinn: I’ve always said that one of the most effective attack ads you could come up with about running Kubernetes would be to send someone who’s considering it to a three-day Kubernetes workshop, and by the time they come back, they will understand that here be dragons. And that has sort of continued to be the case, as far as talking to anyone who’s doing anything at significant scale in the Kubernetes ecosystem, is just the sheer level of abstraction built upon abstraction that fundamentally turns into something that is incredibly difficult and opaque to understand what’s going on underneath the hood. So, it’s not the Day-1 experience it’s the Day-2 experience, as you alluded to earlier in the recording, that once you have something running and then you see a degradation or an intermittent failure, it becomes super challenging to figure out what’s causing that issue and why.

Tobi Knaup: That’s absolutely right. And the typical journey that we see a lot of people go through is something like, they decide to do cloud-native, they decide to do containers, or their boss tells them to. They go on the Internet, they go on Stack Overflow or wherever and they find Kubernetes. They try it out, they download it onto their laptop and have a great experience. The first touch experience with Kubernetes is really great. You can get a container up and running quickly or get your guestbook example up and running quickly. And so, too many people assume that putting it in production is going to be a similar experience. And the first common mistake we see is that people assume that Kubernetes is all they need. That Kubernetes gives them all the tools that they need to put a container stack into production at an Enterprise. And that’s just not the case. You need a bunch of other tools from the cloud-native ecosystem around Kubernetes. You need a monitoring stack, you need logging, you need networking, load balancing, all of those things. And because people kind of take this fairly agile approach where they try it out, and then when they hit a wall, they figure it out. Let’s say I start a container, a stateless container, and that’s a great experience. Now, I need to add state to it. How do I do that? How do I get volumes? They kind of take it step by step, and that’s where we see a lot of cloud-native projects failing is because, like you said, at some point, they face the complexity and they’re like, “Oh, wow, there’s actually a lot of things that I hadn’t thought about.” What we like to do there, is make sure that people are educated about that. So we say, “Hey, when you need to go to production, these are all the things you should pay attention to. Make sure you have proper monitoring, make sure you have proper logging, you need a networking layer, and so on.” And that’s part of what we teach in our Kubernetes trainings, too. So, we do these free trainings, in the field, in various different cities in the world, to just highlight these problems because, like you said, a lot of people just aren’t aware of those.

Corey Quinn: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the Enterprise (not the starship) on-prem security doesn’t translate well to cloud or multi-cloud environments. And that’s not even counting IoT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IoT devices, detects these threats up to 35 percent faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at ExtraHop.com/trial.

Corey Quinn: One of the, I guess, arguments in favor of Kubernetes historically has been the hybrid story, which I’m sympathetic to, and the multi-cloud storage which I’m slightly less sympathetic, in that on paper, it looks fantastic. In practice, it means that you’re not just dealing with one cloud provider’s deficiencies, you’re dealing with all of them. And that’s been a recurring subject of some debate on this show for a while now. Where do you stand on the idea of multi-cloud as a best practice?

Tobi Knaup: Yeah, it’s one of my favorite topics. So I think multi-cloud is where everything is going to move. To me multi-cloud—it also includes hybrid, because every large Enterprise has massive workloads that they want to keep on-prem for various reasons, whether it’s they want to protect their data, whatever, and at a certain scale to actually running your own gear becomes more cost-effective, too. So I think ultimately, every Enterprise is going to be there. Now their reasons for why they want to do multi-cloud vary, and I think a lot of folks, when they hear multi-cloud, the first thing they go to is, “Oh, I’m gonna have this abstraction layer, Kubernetes, or whatever it may be. And I’m going to dynamically move my workloads around, and I’m going to look at where the costs are optimal, or, I’m going to optimize for other things.”

That’s typically not the main reason why people do that, although we are working with some customers that are fairly sophisticated, that are literally doing that, they’re watching the spot instance price market on all the different cloud providers and then hour by hour, decide where things should go. But that’s only a handful and they’re ahead of the pack. They’re fairly sophisticated customers. For most folks, the reasons are something different. We work with a lot of companies that work globally, that work in a lot of different countries and jurisdictions. And so they need to take a look at data privacy laws and regulations around that. So when the infrastructure they stand up in China, that data that they process there, for the Chinese customers, can often not leave the country, so they need to run on a Chinese cloud provider. They may be operating in Europe, so they need to run on European infrastructure and within Europe, in each country.

And so, in this case, when we say multi-cloud, it's often not actually one of the big three cloud providers that they’re thinking about. This may be some fairly small infrastructure as a service provider in one specific country that they need to run on top of, for these data privacy reasons. And so, in this scenario, multi-cloud makes a lot of sense, because you want to architect your stack once. You want to build it on top of an abstraction layer like Kubernetes, and then be able to stand that up in multiple countries on different IS, for those reasons. That’s a common one that we see. Obviously, that’s with companies that act globally, that are working in a lot of different jurisdictions. Another reason we see for multi-cloud, often, is that they want to handpick certain cloud provider services that they like. So they may want to go to Provider A, for their machine-learning stack. And they want to go to Provider B because they have the better managed databases. So it’s more of those reasons, I think, and not so much what most people go to immediately, which is dynamically moving the workloads around.

Corey Quinn: And the dynamic movement of those workloads seems to be what people put up as “Oh, it would be great to be able to magically deploy our entire application anywhere we need to at any point in time.” Except data gravity always makes it a bit of a challenge.

Tobi Knaup: That’s right.

Corey Quinn: The joy of trying to get even that baseline fundamental, consistent experience working between two providers, even when one of them is on-prem and you control virtually every aspect of it, is non-trivial. An argument I’ve enjoyed for a while now has been, “Great, take your primary cloud provider, whichever one it happens to be, I don’t care, you probably care, I don’t care, and triangle multi-region. Be able to span to multiple regions of the same provider and see what breaks.” It’s a good baseline story for the things you’re going to have to start thinking about, and then some, when you start going multi-cloud. Now there are workloads that justify that level of work, and experience, and stress. But it’s certainly not, I’d say, worth an awful lot of companies time and effort to do it.

Tobi Knaup: Yeah, you’re absolutely right. That experience. is very similar to what you’re going to have to do in multi-cloud. And there’s one more use case I forgot to mention earlier, and that is, people in certain industries that are regulated, they actually have to go with multiple vendors. They have to, for regulatory reasons, pick two or more cloud providers, and so that they’re kind of forced to do that. One of the main things you’re going to have to build your own, or it’s actually something we help our customers with, is replicating your data. Like you said, data has gravity. And so the people that we see that are successful at doing multi-cloud or multi-region, they do things like using Kafka to replicate their data, or using Cassandra to replicate the data asynchronously, between different infrastructures. So that’s not something the cloud provider offers, but we help folks manage Cassandra, manage Kafka, so it makes that a little easier.

Corey Quinn: So tell me a little bit about where you came from. Most people don’t decide to spring fully formed from the forehead of some ancient god in the form of a co-founder of a company in the infrastructure space everyone has heard of. Where were you before Mesosphere, if there can be said to be a time before the Mesospheric era.

Tobi Knaup: There is definitely a time before the Mesospheric era. Yeah, my exposure to the Internet and infrastructure and HA basically started as a teenager. So my co-founder, Flo Leibert and I, we grew up in Germany in the same town. And when we were teenagers, we started building websites. And we started building some adventure games and things like that. So we knew how to build websites. And we grew up in a fairly small town, 50,000 people, and this was the late 90s. And even in that neck of the woods, companies started to hear about the Internet. And so they’re basically wondering what this thing is. Someone told them, “Hey, you need to be on the Internet, you need to have a website, and you need to have an email address as a business,” but they had no idea how to think about this and how to approach it. And at the time it was really hard to get a website because you basically had to work with three different companies. You had to find someone to design it for you, you had to find someone to program it, and then someone to host it. Those were typically three different companies. And so what Flo and I did is, we said, “Hey, we know how to build websites, and we know how to run Linux servers.” We just dabbled with that on the side. And so we actually convinced my mom to register a company, and so we could program websites for people and host them. So that was our first experience with infrastructure, and even back then we did HA things. We bought two servers, not one. One would have been enough to host all of our clients, but we wanted it to be highly available. So got some experience with that, running Linux servers, running production infrastructure.

And then when you grow up in Germany, or anywhere outside of Silicon Valley, and you’re in tech, then you hear the stories. You hear about Silicon Valley, and I’ve always imagined it to be the super-futuristic place, and I kind of wanted to check it out at some point. And so in college, I found an internship at this startup down a Redwood City, and join them for three months, and helped them build their websites, PHP and Ruby on Rails, at the time. So that was my first exposure to Silicon Valley. And I just loved the energy, the people that are full of ideas, and the speed at which things get built. And so, after I finished college, I joined that same company where I did that internship, worked full time and built the infrastructure there, built the website.

And then my next job was with Airbnb. I joined them pretty early on as engineer four, and so wore a lot of different hats there. And one of the things I did there is also design and build the infrastructure for their massive growth. Hired the engineering team and did some machine-learning work there too. That’s my other passion besides infrastructure. And at Airbnb, that’s when we started using Apache Mesos. We built data infrastructure there based on Apache Mesos which my third founder, Ben, was working on in Berkeley at the time. I should have mentioned Ben and Flow and I, the three founders, we’ve known each other for a long time. Flo and I grew up together, and Flo did a student exchange, and stayed with Ben’s family in high school too. So, we all love computers, we talked about Mesos and that’s how you know Twitter and Airbnb ended up using Mesos. And to us as the people running the infrastructure there, and the people with the pagers that would go off at three in the morning, sometimes, Mesos really felt like magic. It was a 10 times better solution, because we could automate a lot more things and the pager wouldn’t go off as much in the middle of the night.

And so that’s when we decided hey, this is a great opportunity to start a company, because the problems we were solving there with automation, they were not unique to Twitter, or Airbnb, or any Silicon Valley tech company. They were infrastructure challenges that every company would face at some point. And, this was around a time when this idea of software is eating the world, that Marc Andreessen wrote about, I think in 2009, that was still fairly new. But, we saw that every company, whether it’s a bank, or an insurance company, or a car manufacturer, will have to run large scale cloud infrastructure at some point. And in fact, in order to stay competitive in the future, they’re going to have to use some of the same technologies that the best software companies in the world are using. And so that’s where we saw the opportunities. We saw that we had this tool that automated a lot more things, and making infrastructure more robust, and scalable, and cost-effective. The only challenge at the time was it was an open-source project. There was only a few people in the world that knew how to use it, and so we decided, let’s form a company around it. Let’s build an enterprise product around the open-source core. And that's how Mesosphere was formed in 2013.

Corey Quinn: I have vague recollections, back in the dawn of my version of the era of computing, we would configure a core switch at the office I worked at, then we rented a van and a few of us on the tech ops team drove it down to the data center about 30 miles away and did the installation. And we learned a few things. One we are super crappy movers. Two it is vaguely disturbing the company decided not to spring for professional insured bonded movers for this. And thirdly, there’s something very surreal about loading a piece of computer equipment that fits in a rack, two or three of you can lift it up into the back of a van, and that van costs less than the switch does. It was still such a strange and surreal experience. You don’t get to experience that in the world of cloud in quite the same way. But it’s more than made up for it with the other hilarious and sarcastically disturbing things that it has exposed for us.

Tobi Knaup: Yeah, absolutely. I think interacting with real hardware and a real data center, it’s an experience that really, really shaped how I think about stuff. And one big way is that—always expecting failure, because I could tell so many partially funny, partially painful stories, of things that went wrong in the data center in the physical world. That, I think, what that taught me is to just expect failure, always. Everything can fail at any point in time and then when you build software, even if you build it on top of a bunch of layers of cloud computing, you have to expect that. And even in the cloud, machines will fail, and you get that email from AWS that says, “This instance is now broken.” And I think if you don’t have that experience, racking and stacking gear, and seeing a bunch of physical failures, you might not think about it the same way. So we don’t want to miss that experience. But yeah, like you said, there’s all kinds of other funny behavior that happens in the cloud. It’s just abstracted away by a bunch of layers, but I’ve lived through some funny cloud outages there, too, where packets went in a circle, and then, that caused EBS to go crazy, and all kinds of fun stuff.

Corey Quinn: Oh, the cascading dependencies are always the story, the stuff of legend after the fact. And it makes sense, in hindsight. Every failure does, to some extent, but when you’re in the middle of it, you’re wondering if you’ve lost your mind, if the old rules no longer hold, this behavior is completely inexplicable, what happened? And I guess figuring that out, and living through that a few times is really, I think, the best way to learn to approach those things in a more methodical way. But, ouf, some of those early failures were not fun. Seeing aspects of that manifest in cloud environments is absolutely something that is definitely reminding me that old things are new again.

Tobi Knaup: That’s true. And one thing we shouldn’t forget, too, is that by using these cloud services, you give up a lot of control, too, because when things do fail, there’s only so many things you can do. APIs may all of a sudden be read-only, and you cannot restore your database from a backup all of a sudden, or you cannot promote your read-only database to a master instance all of a sudden. So there’s definitely that aspect, too, which that was much easier when we were running our old gear is, you’re in full control. You control the whole thing, and if you want to do something crazy to try and fix a problem, you can do that. Can’t do that on the cloud.

Corey Quinn: So last question, I suppose, before we wrap this up. I made a prediction about a year ago, I said five years, now, so we have four years to go, where I argued that in four years now, nobody is going to care about Kubernetes. And my argument was not that it’s going to dry up, blow away and be replaced with something else, but rather that it or something like it, is going to slip below the surface of awareness. Just like we don’t have to worry about what kernel version we’re running on an operating system anymore, we won’t care what’s handling orchestration in our various data centers and cloud providers. Do you think that that is an accurate prediction, or I’m going to be eating some crow?

Tobi Knaup: No, I think that’s absolutely accurate. It is a substrate, it’s becoming a substrate. And unless you’re directly involved in adding features to Kubernetes, or you’re using it in some other way where that requires you to make changes to it directly, you’re probably going to use some other higher-level API. You’re probably going to be interfacing with a CI/CD system, as a developer. And maybe you know that it’s Kubernetes under the hood, just like you know right now that it’s Linux under the hood, but you’re not really interacting with it directly that much. Or, if you’re a data scientist, so you work in the data infrastructure world, you’re much more likely to use a tool like Kudo to deploy that service, versus trying to piece together your Kubernetes primitives in order to stand up that service. So I absolutely agree with that. I think that’s a trend. It’s sort of this abstraction layer wave that’s always behind us or the rising tide of abstractions. And so, I think the same thing will be true for Kubernetes. I think most folks out there, individual developers or end-users off a platform, they’ll know that it’s there, but they’re going to be talking to other APIs at different levels. And we’re seeing a lot of activity around CI/CD right now. I think things like Argo and Tekton are super exciting. There’s a lot of activity around that, and people wanting to use GitOps approaches to deploy their software. So I think those are some of the signs that we’re seeing of the abstraction layer rising. And then, of course, serverless, too.

Corey Quinn: So if people want to hear more about your thoughts on these and other topics, where can they find you?

Tobi Knaup: So, they can find me on the usual places. I’m pretty active on Twitter. I’m on LinkedIn. I give talks at conferences sometimes. Those are some of the first places to find me.

Corey Quinn: Excellent, thank you so much for taking the time to speak with me today.

Tobi Knaup: Absolutely. Thank you so much for the opportunity.

Corey Quinn: Tobi Knaup, CTO and co-founder of D2IQ, formerly Mesosphere. I’m Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave it a great rating on Apple Podcasts. If you hated this podcast, please leave it an even better rating on Apple Podcasts.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Gene Kim
Gene Kim is a multiple award-winning CTO, researcher and author, and has been studying high-performing technology organizations since 1999. He was founder and CTO of Tripwire for 13 years. He has written six books, including The Unicorn Project (2019), The Phoenix Project (2013), The DevOps Handbook (2016), the Shingo Publication Award winning Accelerate (2018), and The Visible Ops Handbook (2004-2006) series. Since 2014, he has been the founder and organizer of DevOps Enterprise Summit, studying the technology transformations of large, complex organizations.

Links Referenced

  • The Phoenix Project: https://www.amazon.com/Phoenix-Project-DevOps-Helping-Business/dp/1942788290/
  • The Unicorn Project: https://www.amazon.com/Unicorn-Project-Developers-Disruption-Thriving/dp/B0812C82T9
  • The DevOps Enterprise Summit: https://events.itrevolution.com/
  • @RealGeneKim

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey Quinn: This episode is brought to you by DigitalOcean, the cloud provider that makes it easy for startups to deploy and scale modern web applications with, and this is important to me: No billing surprises. With simple, predictable pricing that's flat across 12 global data center regions, and a UX developers around the world love, you can control your cloud infrastructure costs and have more time for your team to focus on growing your business. See what businesses are building on DigitalOcean and get started for free at do.co/screaming. That's D-O-Dot-C-O-Slash-Screaming, and my thanks to DigitalOcean for their continuing support of this ridiculous podcast.

Corey Quinn: This episode is also sponsored in part by Logz.io. Logz.io hears it loud and clear from many engineers, specifically, that they prefer using open source tools for observability. They get an easier on-boarding experience, a built-in community ready to go, and a lot less vendor lock-in. But, what about the enterprise-grade scale, support, and security? You need to improve the performance of any given cloud environment. That's where Logz.io comes in. They offer a fully managed observability service. Log management and Cloud SIEM based on ELK, and infrastructure monitoring based on Grafana. The open source you love at the scale you need. Sign up today for a 14 day free trial at logz.io/screaming and for your chance to receive a free Logz.io t-shirt.

Corey Quinn: Welcome to Screaming in the Cloud. I’m Corey Quinn. I’m joined this week by a man who needs no introduction but gets one anyway. Gene Kim, most famously known for writing The Phoenix Project, but now the Wall Street Journal best-selling author of The Unicorn Project, six years later. Gene, welcome to the show.

Gene Kim: Corey so great to be on. I was just mentioning before how delightful it is to be on the other side of the podcast. And it’s so much smaller in here than I had thought it would be.

Corey Quinn: Excellent. It’s always nice to wind up finally meeting people whose work was seminal and foundational. Once upon a time, when I was a young, angry Unix systems administrator—because it’s not like there’s a second type of Unix administrator—[laughing] The Phoenix Project was one of those texts that was transformational, as far as changing the way I tended to view a lot of what I was working on and gave a glimpse into what could have been a realistic outcome for the world, or the company I was at, but somehow was simultaneously uplifting and incredibly depressing all at the same time. Now, The Unicorn Project does that exact same thing only aimed at developers instead of traditional crusty ops folks.

Gene Kim: [laughing] Yeah, yeah. Very much so. Yeah, The Phoenix Project was very much aimed at ops leadership. So, Bill Palmer, the protagonist of that book was the VP of Operations at Parts Unlimited, and the protagonist in The Unicorn Project is Maxine Chambers, Senior Architect, and Developer, and I love the fact that it’s told in the same timeline as The Phoenix Project, and in the first scene, she is unfairly blamed for causing the payroll outage and is exiled to The Phoenix Project, where she recoils in existential horror and then finds that she can’t do anything herself. She can’t do a build, she can’t run her own tests. She can’t, God forbid, do her own deploys. And I just love the opening third of the book where it really does paint that tundra that many developers find themselves in where they’re just caught in decades of built-up technical debt, unable to do even the simplest things independently, let alone be able to independently develop tests or create value for customers. So, it was fun, very much fun, to revisit the Parts Unlimited universe.

Corey Quinn: What I found that was fun about—there are few things in there I want to unpack. The first is that it really was the, shall we say, retelling of the same story in, quote/unquote, “the same timeframe”, but these books were written six years apart.

Gene Kim: Yeah, and by the way, I want to first acknowledge all the help that you gave me during the editing process. Some of your comments are just so spot on with exactly the feedback I needed at the time and led to the most significant lift to jam a whole bunch of changes in it right before it got turned over to production. Yeah, so The Phoenix Project is told, quote, “in the present day,” and in the same way, The Unicorn Project is also told—takes place in the present day. In fact, they even start, plus or minus, on the same day. And there is a little bit of suspension of disbelief needed, just because there are certain things that are in the common vernacular, very much in zeitgeist now, that weren’t six years ago, like “digital disruption”, even things like Uber and Lyft that feature prominently in the book that were just never mentioned in The Phoenix Project, but yeah, I think it was the story very much told in the same vein as like Ender’s Shadow, where it takes place in the same timeline, but from a different perspective.

Corey Quinn: So, something else that—again, I understand it’s an allegory, and trying to tell an allegorical story while also working it into the form of a fictional work is incredibly complicated. That’s something that I don’t think people can really appreciate until they’ve tried to do something like it. But I still found myself, at various times, reading through the book and wondering, asking myself questions that, I guess, say more about me than they do about anyone else. But it’s, “Wow, she’s at a company that is pretty much scapegoating her and blaming her for all of us. Why isn’t she quitting? Why isn’t she screaming at people? Why isn’t she punching the boss right in their stupid, condescending face and storming out of the office?” And I’m wondering how much of that is my own challenges as far as how life goes, as well as how much of it is just there for, I guess, narrative devices. It needed to wind up being someone who would not storm out when push came to shove.

Gene Kim: But yeah, I think she actually does the last of the third thing that you mentioned where she does slam the sheet of paper down and say, “Man, you said the outage is caused by a technical failure and a human error, and now you’re telling me I’m the human error?” And just cannot believe that she’s been put in that position. Yeah, so thanks to your feedback and the others, she actually does shop her resume around. And starts putting out feelers, because this is no longer feeling like the great place to work that attracted her, eight years prior. The reality is for most people, is that it’s sometimes difficult to get a new job overnight, even if you want to. But I think that Maxine stays because she believes in the mission. She takes a great deal of pride of what she’s created over the years, and I think like most great brands, they do create a sense of mission and there’s a deep sense of the customers they serve. And, there’s something very satisfying about the work to her. And yeah, I think she is very much, for a couple of weeks, very much always thinking about, she won’t be here for long, one way or another, but by the time she stumbles into the rebellion, the crazy group of misfits, the ragtag bunch of misfits, who are trying to find better ways of working and willing to break whatever rules it takes to take over the very ancient powerful order, she falls in love with a group. She found a group of kindred spirits who very much, like her, believe that developer productivity is one of the most important things that we can do as an organization. So, by the time that she looks up with that group, I mean, I think she’s all thoughts of leaving are gone.

Corey Quinn: Right. And the idea of, if you stick around, you can theoretically change things for the better is extraordinarily compelling. The challenge I’ve seen is that as I navigate the world, I’ve met a number of very gifted employees who, frankly wind up demonstrating that same level of loyalty and same kind of loyalty to companies that are absolutely not worthy of them. So my question has always been, when do I stick around versus when do I leave? I’m very far on the bailout as early as humanly possible side of that spectrum. It’s why I’m a great consultant but an absolutely terrible employee.

Gene Kim: [laughing] Well, so we were honored to have you at the DevOps Enterprise Summit. And you’ve probably seen that The Unicorn Project book is really dedicated to the achievements of the DevOps Enterprise community. It’s certainly inspired by and dedicated to their efforts. And I think what was so inspirational to me were all these courageous leaders who are—they know what the mission is. I mean, they viscerally understand what the mission is and understand that the ways of working aren’t working so well and are doing whatever they can to create better ways of working that are safer, faster, and happier. And I think what is so magnificent about so many of their journeys is that their organization in response says, “Thank you. That’s amazing. Can we put you in a position of even more authority that will allow you to even make a more material, more impactful contribution to the organization?” And so it’s been my observation, having run the conference for, now, six years, going on seven years is that this is a population that is being out promoted—has been promoted at a rate far higher than the population at large. And so for me, that’s just an incredible story of grit and determination. And so yeah, where does grit and determination becomes sort of blind loyalty? That’s ultimately self-punishing? That’s a deep question that I’ve never really studied. But I certainly do understand that there is a time when no amount of perseverance and grit will get from here to there, and that’s a fact.

Corey Quinn: I think that it’s a really interesting narrative, just to see it, how it tends to evolve, but also, I guess, for lack of a better term, and please don’t hold this against me, it seems in many ways to speak to a very academic perspective, and I don’t mean that as an insult. Now, the real interesting question is why I would think, well—why would accusing someone of being academic ever be considered as an insult, but my academic career was fascinating. It feels like it aligns very well with The Five Ideals, which is something that you have been talking about significantly for a long time. And in an academic setting that seems to make sense, but I don’t see it thought of or spoken of in the same way on the ground. So first, can you start off by giving us an intro to what The Five Ideals are, and I guess maybe disambiguate the theory from the practice?

Gene Kim: Oh for sure, yeah. So The Five Ideals are— oh, let’s go back one step. So The Phoenix Project had The Three Ways, which were the principles for which you can derive all the observed DevOps practices from and The Four Types of Work. And so in The Five Ideals I used the concept of The Five Ideals and they are—the first—

Corey Quinn: And the next version of The Nine whatever you call them at that point, I’m sure. It’s a geometric progression.

Gene Kim: Right or actually, isn’t it the pri—oh, no. four isn’t, four isn’t prime. Yeah, yeah, I don’t know. So, The Five Ideals is a nice small number and it was just really meant to verbalize things that I thought were very important, things I just gravitate towards. One is Locality and Simplicity. And briefly, that’s just, to what degree can teams do what they need to do independently without having to coordinate, communicate, prioritize, sequence, marshal, deconflict, with scores of other teams. The Second Ideal is what I think the outcomes are when you have that, which is Focus, Flow and Joy. And so, Dr. Mihaly Csikszentmihalyi, he describes flow as a state when we are so engrossed in the work we love that we lose track of time and even sense of self. And that’s been very much my experience, coding ever since I learned Clojure, this functional programming language. Third Ideal is Improvement of Daily Work, which shows up in The Phoenix Project to say that improvement daily work is even more important than daily work itself. Fourth Ideal is Psychological Safety, which shows up in the State of DevOps Report, but showed up prominently in Google’s Project Oxygen, and even in the Toyota production process where clearly it has to be—in order for someone to pull the andon cord that potentially stops the assembly line, you have to have an environment where it’s psychologically safe to do so. And then Fifth Ideal is Customer Focus, really focus on core competencies that create enduring, durable business value that customers are willing to pay for, versus context, which is everything else. And yeah, to answer your question, Where did it come from? Why do I think it is important? Why do I focus on that? For me, it’s really coming from the State of DevOps Report, that I did with Dr. Nicole Forsgren and Jez Humble. And so, beyond all the numbers and the metrics and the technical practices and the architectural practices and the cultural norms, for me, what that really tells the story of is of The Five Ideals, as to what one of them is very much a need for architecture that allows teams to work independently, having a higher predictor of even, continuous delivery. I love that. And that from the individual perspective, the ideal being, that allows us to focus on the work we want to do to help achieve the mission with a sense of flow and joy. And then really elevating the notion that greatness isn’t free, we need to improve daily work, we have to make it psychologically safe to talk about problems. And then the last one really being, can we really unflinchingly look at the work we do on an everyday basis and ask, what the customers care about it? And if customers don’t care about it, can we question whether that work really should be done or not. So that’s where for me, it’s really meant to speak to some more visceral emotions that were concretized and validated through the State of DevOps Report. But these notions I am just very attracted to.

Corey Quinn: I like the idea of it. The question, of course, is always how to put these into daily practice. How do you take these from an idealized—well, let’s not call it a textbook, but something very similar to that—and apply it to the I guess, uncontrolled chaos that is the day-to-day life of an awful lot of people in their daily jobs.

Gene Kim: Yeah. Right. So, the protagonist is Maxine and her role in the story, in the beginning, is just to recognize what not great looks like. She’s lived and created greatness for all of her career. And then she gets exiled to this terrible Phoenix project that chews up developers and spits them out and they leave these husks of people they used to be. And so, she’s not doing a lot of problem-solving. Instead, it’s this recoiling from the inability for people to do builds or do their own tests or be able to do work without having to open up 20 different tickets or not being able to do their own deploys. She just recoil from this spending five days watching people do code merges, and for me, I’m hoping that what this will do, and after people read the book, will see this all around them, hopefully, will have a similar kind of recoiling reaction where they say, “Oh my gosh, this is terrible. I should feel as bad about this as Maxine does, and then maybe even find my fellow rebels and see if we can create a pocket of greatness that can become like the sublimation event in Dr. Thomas Kuhn’s book, The Structure of Scientific Revolutions.” Create that kernel of greatness, of which then greatness then finds itself surrounded by even more greatness.

Corey Quinn: What I always found to be fascinating about your work is how you wind up tying so many different concepts together in ways you wouldn’t necessarily expect. For example, when I was reviewing one of your manuscripts before this went to print, you did reject one of my suggestions, which was just, retitle the entire thing. Instead of calling it The Unicorn Project. Instead, call it Gene Kim’s Love Letter to Functional Programming. So what is up with that?

Gene Kim: Yeah, to put that into context, for 25 years or more, I’ve self-identified as an ops person. The Phoenix Project was really an ops book. And that was despite getting my graduate degree in compiler design and high-speed networking in 1995. And the reason why I gravitated towards ops, because that was my observation, that that’s where the saves were made. It was ops who saved the customer from horrendous, terrible developers who just kept on putting things into production that would then blow up and take everyone with it. It was ops protecting us from the bad adversaries who were trying to steal data because security people were so ineffective. But four years ago, I learned a functional programming language called Clojure and, without a doubt, it reintroduced the joy of coding back into my life and now, in a good month, I spend half the time—in the ideal—writing, half the time hanging out with the best in the game, of which I would consider this to be a part of, and then 20% of time coding. And I find for the first time in my career, in over 30 years of coding, I can write something for years on end, without it collapsing in on itself, like a house of cards. And that is an amazing feeling, to say that maybe it wasn’t my inability, or my lack of experience, or my lack of sensibilities, but maybe it was just that I was sort of using the wrong tool to think with. That comes from the French philosopher Claude Lévi-Strauss. He said of certain things, “Is it a good tool to think with?” And I just find functional programming is such a better tool to think with, that notions like composability, like immutability, what I find so exciting is that these things aren’t just for programming languages. And some other programming languages that follow the same vein are, OCaml, Lisp, ML, Elixir, Haskell. These all languages that are sort of popularizing functional programming, but what I find so exciting is that we see it in infrastructure and operations, too. So Docker is fundamentally immutable. So if you want to change a container, we have to make a new one. Kubernetes composes these containers together at the level of system of systems. Kafka is amazing because it usually reveals the desire to have this immutable data model where you can’t change the past. Version control is immutable. So, I think it’s no surprise that as our systems get more and more complex and distributed, we’re relying on things like immutability, just to make it so that we can reason about them. So, it is something I love addressing in the book, and it’s something I decided to double down on after you mentioned it. I’m just saying, all kidding aside is this a book for—

Corey Quinn: Oh good, I got to make it worse. Always excited when that happens.

Gene Kim: Yeah, I mean, your suggestion really brought to the forefront a very critical decision, which was, is this a book for technology leaders, or even business leaders, or is this a book developers? And, after a lot of soul searching, I decided no, this is a book for developers, because I think the sensibilities that we need to instill and the awareness we need to create these things around are the developers and then you just hope and pray that the book will be good enough that if enough engineers like it, then engineering leaders will like it. And if enough engineering leaders like it, then maybe some business leaders will read it as well. So that’s something I’m eagerly seeing what will happen as the weeks, months, and years go by.

Corey Quinn: This episode is sponsored in part by DataStax. The NoSQL event of the year is DataStax Accelerate in San Diego this May from the 11th through the 13th. I've given a talk previously called the myth of multi-cloud, and it's time for me to revisit that with... A sequel! Which is funny given that it's a NoSQL conference, but there you have it. To learn more, visit datastax.com that's D-A-T-A-S-T-A-X.com and I hope to see you in San Diego. This May.

Corey Quinn: One thing that I always admired about your writing is that you can start off trying to make a point about one particular aspect of things. And along the way you tie in so many different things, and the functional programming is just one aspect of this. At some point, by the end of it, I half expected you to just pick a fight over vi versus Emacs, just for the sheer joy you get in effectively drawing interesting and, I guess, shall we say, the right level of conflict into it, where it seems very clear that what you’re talking about is something thing that has the potential to be transformative and by throwing things like that in you’re, on some level, roping people in who otherwise wouldn’t weigh in at all. But it’s really neat to watch once you have people’s attention, just almost in spite of what they want, you teach them something. I don’t know if that’s a fair accusation or not, but it’s very much I’m left with the sense that what you’re doing has definite impact and reverberations throughout larger industries.

Gene Kim: Yeah, I hope so. In fact, just to reveal this kind of insecurity is, there’s an author I’ve read a lot of and she actually read this blog post that she wrote about the worst novel to write, and she called it The Yeomans Tour of the Starship Enterprise. And she says, “The book begins like this: it’s a Yeoman on the Starship Enterprise, and all he does is admire the dilithium crystals, and the phaser, and talk about the specifications of the engine room.” And I sometimes worry that that’s what I’ve done in The Unicorn Project, but hopefully—I did want to have that technical detail there and share some things that I love about technology and the things I hate about technology, like YAML files, and integrate that into the narrative because I think it is important. And I would like to think that people reading it appreciate things like our mutual distaste of YAML files, that we’ve all struggled trying to escape spaces and file names inside of make files. I mean, these are the things that are puzzles we have to solve, but they’re so far removed from the business problem we’re trying to solve that really, the purpose of that was trying to show the mistake of solving puzzles in our daily work instead of solving real problems.

Corey Quinn: One thing that I found was really a one-two punch, for me at least, was first I read and give feedback on the book and then relatively quickly thereafter, I found myself at my first DevOps Enterprise Summit, and I feel like on some level, I may have been misinterpreted when I was doing my live-tweeting/shitposting-with-style during a lot of the opening keynotes, and the rest, where I was focusing on how different of a conference it was. Unlike a typical DevOps Days or big cloud event, it wasn’t a whole bunch of relatively recent software startups. There were serious institutions coming out to have conversations. We’re talking USAA, we’re talking to US Air Force, we’re talking large banks, we’re talking companies that have a 200-year history, where you don’t get to just throw everything away and start over. These are companies that by and large, have, in many ways, felt excluded to some extent, from the modern discussions of, well, we’re going to write some stuff late at night, and by the following morning, it’s in production. You don’t get to do that when you’re a 200-year-old insurance company. And I feel like that was on some level interpreted as me making fun of startups for quote/unquote, “not being serious,” which was never my intention. It’s just this was a different conversation series for a different audience who has vastly different constraints. And I found it incredibly compelling and I intend to go back.

Gene Kim: Well, that’s wonderful. And, in fact, we have plans for you, Mr. Quinn.

Corey Quinn: Uh-oh.

Gene Kim: Yeah. I think when I say I admire the DevOps Enterprise community. I mean that I’m just so many different dimensions. The fact that these, leaders and—it’s not leaders just in terms of seniority on the organization chart—these are people who are leading technology efforts to survive and win in the marketplace. In organizations that have been around sometimes for centuries, Barclays Bank was founded in the year 1634. That predates the invention of paper cash. HMRC, the UK version of the IRS was founded in the year 1200. And, so there’s probably no code that goes that far back, but there’s certainly values and—

Corey Quinn: Well, you’d like to hope not.

Gene Kim: Yeah, right. You never know. But there are certainly values and traditions and maybe even processes that go back centuries. And so that’s what’s helped these organizations be successful. And here are a next generation of leaders, trying to make sure that these organizations see another century of greatness. So I think that’s, in my mind, deeply admirable.

Corey Quinn: Very much so. And my only concern was, I was just hoping that people didn’t misinterpret my snark and sarcasm as aimed at, “Oh, look at these crappy—these companies are real companies and all those crappy SAS companies are just flashes in the pan.” No, I don’t believe that members of the Fortune 500 are flash in the pan companies, with a couple notable exceptions who I will not name now, because I might want some of them on this podcast someday. The concern that I have is that everyone’s work is valuable. Everyone’s work is important. And what I’m seeing historically, and something that you’ve nailed, is a certain lack of stories that apply to some of those organizations that are, for lack of a better term, ossified into their current process model, where they there’s no clear path for them to break into, quote/unquote, “doing the DevOps.”

Gene Kim: Yeah. And the business frame and the imperative for it is incredible. Tesla is now offering auto insurance bundled into the car. Banks are now having to compete with Apple. I mean, it is just breathtaking to see how competitive the marketplaces and the need to understand the customer and deliver value to them quickly and to be able to experiment and innovate and out-innovate the competition. I don’t think there’s any business leader on the planet who doesn’t understand that software is eating the world and they have to that any level of investment they do involves software at some level. And so the question is, for them, is how do they get educated enough to invest and manage and lead competently? So, to me it really is like the sleeping giant awakening. And it’s my genuine belief is that the next 50 years, as much value as the tech giants have created: Facebook, Amazon, Netflix, Google, Microsoft, they’ve generated trillions of dollars of economic value. When we can get eighteen million developers, as productive as an engineer at a tech giant is, that will generate tens of trillions of dollars of economic value per year. And so, when you generate that much economic activity, all problems become solvable, you look at climate change, you take a look at the disparity between rich and poor. All things can be fixed when you significantly change the economic economy in this way. So, I’m extremely hopeful and I know that the need for things like DevOps are urgent and important.

Corey Quinn: I guess that that’s probably the best way of framing this. So you wrote one version that was aimed at operators back in 2013, this one was aimed at developers, and effectively retails and clarifies an awful lot of the same points. As a historical ops person, I didn’t feel left behind by The Unicorn Project, despite not being its target market. So I guess the question on everyone’s mind, are you planning on doing a third iteration, and if so, for what demographic?

Gene Kim: Yeah, nothing at this point, but there is one thing that I’m interested in which is the role of business leaders. And Sarah is an interesting villain. One of my favorite pieces of feedback during the review process was, “I didn’t think I could ever hate Sarah more. And yet, I did find her even to be more loathsome than before.” She’s actually based on a real person, someone that I worked with.

Corey Quinn: That’s the best part, is these characters are relatable enough that everyone can map people they know onto various aspects of them, but can’t ever disclose the entire list in public because that apparently has career consequences.

Gene Kim: That’s right. Yes, I will not say who the character is based on but there’s, in the last scene of the book that went to print, Sarah has an interesting interaction with Maxine, where they meet for lunch. And, I think the line was, “And it wasn’t what Maxine had thought, and she’s actually looking forward to the next meeting.” I think that leaves room for it. So one of the things I want to do with some friends and colleagues is just understand, why does Sarah act the way she does? I think we’ve all worked with someone like her. And there are some that are genuinely bad actors, but I think a lot of them are doing something, based on genuine, real motives. And it would be fun, I thought, to do something with Elizabeth Henderson, who we decided to start having a conversation like, what does she read? What is her background? What is she good at? What does her resume look like? And what caused her to—who in technology treated her so badly that she treats technology so badly? And why does she behave the way she does? And so I think she reads a lot of strategy books. I think she is not a great people manager, I think she maybe has come from the mergers and acquisition route that viewed people as fungible. And yeah, I think she is definitely a creature of economics, was lured by an external investor, about how good it can be if you can extract value out of the company, squeeze every bit of—sweat every asset and sell the company for parts. So I would just love to have a better understanding of, when people say they work with someone like a Sarah, is there a commonality to that? And can we better understand Sarah so that we can both work with her and also, compete better against her, in our own organizations?

Corey Quinn: I think that’s probably a question best left for people to figure out on their own, in a circumstance where I can’t possibly be blamed for it.

Gene Kim: [laughing].That can be arranged, Mr. Quinn.

Corey Quinn: All right. Well, if people want to learn more about your thoughts, ideas, feelings around these things, or of course to buy the book, where can they find you?

Gene Kim: If you’re interested in the ideas that are in The Unicorn Project, I would point you to all of the freely available videos on YouTube. Just Google DevOps Enterprise Summit and anything that’s on the plenary stage are specifically chosen stories that very much informed The Unicorn Project. And the best way to reach me is probably on Twitter. I’m @RealGeneKim on Twitter, and feel free to just @ mention me, or DM me. Happy to be reached out in whatever way you can find me.

Corey Quinn: You know where the hate mail goes then. Gene, thank you so much for taking the time to speak with me, I appreciate it.

Gene Kim: And Corey, likewise, and again, thank you so much for your unflinching feedback on the book and I hope you see your fingerprints all over it and I’m just so delighted with the way it came out. So thanks to you, Corey.

Corey Quinn: As soon as my signed copy shows up, you’ll be the first to know.

Gene Kim: Consider it done.

Corey Quinn: Excellent, excellent. That’s the trick, is to ask people for something in a scenario in which they cannot possibly say no. Gene Kim, multiple award-winning CTO, researcher, and author. Pick up his new book, The Wall Street Journal best-selling The Unicorn Project. I’m Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you’ve enjoyed this podcast, please leave a five-star review on Apple Podcasts. If you hated this podcast, please leave a five-star review on Apple Podcasts and leave a compelling comment.

Announcer: This has been this week’s episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Jeremy Bowers
Jeremy Bowers is an Engineering Director for the Newsroom Engineering team at The Washington Post. Previously, Jeremy was the Senior Editor for News Applications on the Interactive News Team of The New York Times, where he led a team focused on writing software for elections, Congress and the Supreme Court. Jeremy was also a news applications developer on the NPR Visuals team and a Senior Newsroom Developer at The Washington Post.

Links Referenced:

  • Twitter: @jeremybowers
  • LinkedIn: https://www.linkedin.com/in/jeremyjbowers/
  • Personal site: jeremybowers.com
  • Company site: https://www.washingtonpost.com/

Transcript

Corey: Hello and welcome to Screaming in the Cloud with your host, cloud economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

This episode is sponsored by Cloud Zero. Cloud Zero doesn't believe me when I say that trying to make your developers care about cost is a waste of time, yet they still sponsor my podcast, so here we are. Their platform is designed to help SaaS companies understand how and why their costs are changing and lets you easily measure by things SaaS companies care about, like cost per product or feature. Also, they won't make you set up rules or budgets in advance, but will still alert you before your intern spends $10,000 on SageMaker on a Friday afternoon. Go to CloudZero.com to kick off a free trial. That's Cloud Z-E-R-O dot com. And my thanks to them for sponsoring this podcast.

This episode is brought to you by DigitalOcean, the cloud provider that makes it easy for startups to deploy and scale modern web applications with, and this is important to me: No billing surprises. With simple, predictable pricing that's flat across 12 global data center regions, and a UX developers around the world love, you can control your cloud infrastructure costs and have more time for your team to focus on growing your business. See what businesses are building on DigitalOcean and get started for free at do.co/screaming. That's D-O-Dot-C-O-Slash-Screaming, and my thanks to DigitalOcean for their continuing support of this ridiculous podcast.

Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Jeremy Bowers who's an engineering director for the newsroom engineering team at The Washington Post. Jeremy, welcome to the show.

Jeremy: Thanks, Corey. It's good to be here.

Corey: It's great to have you on the show, but it does feel a little surreal. It's oh, you work at an actual journalistic outlet. Cool, do you want to come on my nonsense podcast and talk about computers for a while? And first took a fair bit of courage to ask that question. The response was, "Absolutely." Which oh great, you folks are human beings, too. You put your pants on two legs at a time just like the rest of us.

Jeremy: You know, you keep making these assumptions about pants and I'm not 100% sure that you've ever put them on ... OH, yes. Right. Of course we put them on two legs at a time.

Corey: Oh, yes. True thought leaders jump into their pants.

Jeremy: Exactly.

Corey: So you're an engineering director, which most people can wrap their heads around what that means if they're listening to a show like this. But for the newsroom engineering team, what is that?

Jeremy: Well, any newspaper of sufficient size, or even media company will run into these problems where they have reporters who are working on beats that are really heavy in data, or they have folks who need tools built for them to help them get to stories that they can't get to. And many organizations kind of deal with this on ad hoc basis. When I used to work at The St. Petersburg times in Florida, now the Tampa Bay Times, don't ask. We would handle this with like a small group of like one or two people who were like self taught programmers. But The Post being a large news organization has an actual legitimate engineering team. This is the kind of thing we deal with. So when we have reporters who are trying to make sense of a voter file, which is like a gigantic spreadsheet, that's 384 million rows long and about 600 columns wide. Or if they're looking at Federal Election Commission data that has campaign finance disclosures.

It's just not something that they can look at with their eyes and make sense of it. They need software written for them, and so that's what the team does.

Corey: That sounds like it's almost the quarter of the size of a typical month's AWS bill.

Jeremy: We won't even talk about the size of our AWS bill.

Corey: One can only imagine. It's interesting because it's sort of a business bubble problem that business analysts and BI folks deal with all the time. But given what you do and how you present it to the world, you take it a step beyond. It's not just about sorting through vast quantities of data and turning that into pretty reports in PowerPoint. But you then have to take it a step further and present that as something that someone who's not in a business environment or has the context to absorb with potentially an MBA equivalent level of understanding, but people in middle schools, back when ... once upon a time when, I'm dating myself now, we would have to take out clippings from various newspapers for current events day and go in and give a small presentation on it.

And I don't know if calc kids do that today, but there was always a, "Have something data oriented" was periodically something they tried to wind up having us talk about. But being able to deliver complex distilled outcomes from data to a middle school reading level has got to be a phenomenal challenge that most business folks don't I guess have the luxury of not having to worry about.

Jeremy: You know, this is one of the things that we struggle with a lot. Early attempts at data visualization from the 1800s were using relatively complex stuff.

Corey: Oh, like back in the 1800s?

Jeremy: Yeah. Legitimate 1800s.

Corey: Oh, early days of Microsoft Excel.

Jeremy: Exactly. Yes, right. So they were writing their VB script. William Playfair was writing Visual Basic to try to generate these complex charts and stuff that, you know, for people who had exceedingly low reading levels for the general population. He was trying to get rather complex things into their heads. I think the thing that we have now is we understand that our readers might be sophisticated if they can put their full brunt of their attention against something. But we understand that we're competing for attention with so many other things. And so the majority of what we're trying to do isn't just to distill it down to make it easier to understand, but to make it quicker to understand so that you can instantly get the thing that you need to know about the story. And typically that's what we're using the data to do.

Corey: Do you have an example of a story or a series of stories that you spent a significant amount of time on? I mean, I have to imagine you're not sitting there doing the deep dive distillation of data for, "And that's who won the Grammy's this morning." Or actually, that's a terrible example, because I'm sure there's data that goes into it.

Jeremy: There's excellent data.

Corey: Yeah. And that's the enormous pile of data that talks about, oh, I don't know, something dumb some political figure said this morning.

Jeremy: You know, I'll tell you, the thing that we've been spending a lot of time with, so my team focuses mostly on elections and it means that we spend a lot of time with our politics staff, our national politics staff. And a thing that we're learning is that this is a lot of reporters, it's their first time really working with large data sets on their beats. A lot of them come from other beats inside the paper and they're working on the campaign for the first time and they suddenly have to start explaining things like how the shape of the American electorate. Or what are these trends that we should know about campaign finance donations?

And these are legitimately difficult questions to answer, even if you're a real nerd and good at this. And it's particularly difficult if you've just come from sports or from features and now you're writing stories about how campaigns are operating. And so the vast majority of the stuff that we end up working on is trying to find ways to integrate the sort of complex stuff that we get from these data sets and into reporters like Daily Work Flow. So a good example of this would be we have reporters who want to go write about the state of Texas. They would like to write a story about what's happening in Texas and is it likely that Texas is going to turn blue or purple in the 2020 election. And so in typical years, like our reporters might go down and write a story about Austin, because there's a lot of democratic voters in Austin. But we do a quick look through the voter file and we can point out that a story about Texas turning blue shouldn't focus on voters in Austin, because there's not a lot of new voters in Austin.

If you want to talk about Texas changing colors in the next election, you want to talk about the Houston suburbs where there's been a huge influx of Latinx and democratic voters that are relatively new to the state and they are not voting republican. So these are trends. It's basically updating the priors that reporters are holding to help them write better anecdotes when they go to write an anecdote. So if we're going to send a reporter to Texas, we won't send them to Austin, we'll send them to Houston.

Corey: So how does that data get intelligently surfaced to the newsroom? I can see having to just run the following set of SQL queries is all well and good, but having had conversations with a fair few reporters over the course of my career, most of them did not spend most of their time dealing with databases. And if they did, they were really sad all the time.

Jeremy: Right? I mean, this is .. Actually, it's probably the hardest part of the job. You would think that maybe the hardest part of the job would be data ingestion and cleanup or the sense making steps that we take to sort of find trends. But if the only thing that we can build for a reporter is like a dashboard that we have to ask them to come look at and like stare at pie charts for a couple hours, they're just not going to include that information in their story and then we're all kind of ... it's a lesser version of the story that surfaces.

So like the first thing that we think about is all right, so when we find something interesting in this data set, how are we going to get this to a reporter so that they can make, so that this can be a first class part of their story. So the majority of the time, we spend a lot of time trying to figure out how to turn it into words rather than into something visual, although we do have a pretty good relationship with our graphics desk, but that's a whole other conversation. But for our reporters, we've focused on these little newsletters that we can send them. We have a lab that's working on the elections. This professor, Nick Diakopoulos from Northwestern University is hanging out with us for the fall. And one of the things that he's working on is lead generation. Not like what you would think from business intelligence, but literally helping reporters write leads to their stories.

So he has this little piece of natural language processing, but like reads in a bunch of data and produces like essentially a tip sheet for a state that shows off things like here's places where a lot of new voters have registered, or here's places where that voter registration is weird or interesting in some way. Here's places where the vote history has changed significantly in the last two years. And these are really great ways to sort of get a reporter a little fact or two that they can use when they go out and do their reporting. And so it's not going to be a story that you look at and go, "Hey, that's a story about data." It just is a standard newspaper story that might feel anecdotal to you, it's just that it's the correct anecdote instead of the wrong one.

Corey: That sounds like an incredibly challenging problem, but it also feels like it's directly aligned with a lot of the modern day snake oil. Sorry, it is more upscale than that. The modern day serpent grease-

Jeremy: Yeah. Appropriate.

Corey: ... that is artificial intelligence/machine learning/math if you're not trying to scam VC money out of people.

Jeremy: Yes, right.

Corey: How, I guess, when I think about a newsroom and its technology, my immediate mental leap is to typewriters and notebooks, and people who buy pencils by the gross. And it's a very old timey type of mental image that's largely informed by comic books and cartoons when I was a kid. I'm going to assume the technology has evolved somewhat if for no other reason that I listen to, for example The Washington Post podcast in the shower most mornings. I don't wind up opening the window and getting smacked in the face by some paperboy hurling a ... sorry, paper child, paper youth as the case may be, hurling a newspaper into my face. So there's obviously been significant technical changes that have hit the journalism profession in the last 30 years. But it's strange I think on some level, the collective consciousness hasn't really caught up with that.

Jeremy: It's funny that you should bring that up about having a street urchin hurl a wadded up piece of newsprint at you in the shower, because it sort of illuminates this problem that we have in journalism, which is that there was a lot of technological thinking and engineering thinking that went into changing how we present information to our readers, but comparatively little that goes into how we change the way our reporters actually report on stories. And so you're joking about typewriters, and honestly a lot of reporting is not that different than it was in the 1980s or 1990s. The advent of the Internet is a thing that we definitely have to have in like that reporters use Lexus Nexus and things like that to do searches on people.

But truthfully, there's a whole lot of change that's happened in say in other fields where like reporting has just sort of lagged behind. So a thing I think that I'm pretty excited about doing is like attempting to take some lessons that we might've learned at other places, minus the serpent grease, of course, and try to bring some of that sense making to our reporters. Now critically, I think that it's basically impossible or a poor task for us to try to replace our reporters with say a fleet of robots who do all of the reporting work, produce a story and then put it out on the Internet. That would be-

Corey: That's called a content mill.

Jeremy: Exactly. I still think that there is, I still think that reporters have these pattern matching skills that are more or less ineffable. And I think it would be remarkably difficult if even impossible to replicate. But there are certain things that they do on a daily basis that would make your head spin if you saw the wasted skills. When I worked at The New York Times, we had this Pulitzer Prize winning legal reporter and every morning he would wake up and he would refresh the supreme court website to see if any new audio transcripts had been published yet. It took him about 10 or 15 minutes to crawl to all the different pages on the supreme court's website to see if anything new had popped up, and then he would brush his teeth and come into work, and then he'd check again. After a couple meetings, he would check again, do a couple phone calls, check again.

And I just found this to be like a breathtaking waste of an incredibly highly powerful mind to spend those extra minutes pressing F5 on SupremeCourt.gov. So we built him a little bot. And all the bot does is tell him when a new transcript has been filed and it drops it off to him in Slack, and then he can click on it and read it.

Corey: RSS. You invented RSS.

Jeremy: Exactly. That's the thing is this is the lowest of low hanging fruit. It's practically touching the ground. There's so many of those cases where a lot of the tooling that we build, it does not feel technologically superior. It feels like some real lightweight stuff but in truth we get a lot of mileage out of just getting the lowest fruit. And in particular, helping reporters out with little things like that means that they will be more trustworthy, our team will be more trustworthy to them when we start working with them on things that are more sophisticated and require honestly more trust on the part of the reporter that the changes that we're making to the reporting process aren't bunk.

Corey: I have to ask given this is the Screaming in the Cloud podcast, how has cloud technology impacted what you personally I guess, and newsrooms collectively do, if at all? I feel like I need to throw the "if at all" in there, but let's not kid ourselves. I don't think there's any industry that hasn't been touched by this.

Jeremy: It's a pretty standard story. I worked at The Post in 2011 and 2012, then took a short break at The New York Times and at NPR in between before coming back in the last April.

Corey: Oh, a good detox. Small publications.

Jeremy: You know, I did get to tell my friends at The New York Times that I was leaving to go back to my hometown newspaper and I don't think any of them found that very funny.

Corey: No, I imagine they would not have.

Jeremy: The post in 2011 and 2012, when we were running election results, we were running them on physical servers that were located inside the building, like just around the corner from my desk. I could actually go to the server room and look at it if I wanted to. We have five servers that we ran millions of page views off of, and we did not have root access to those servers, because a kind soul in New Jersey had decided that we did not deserve to have root access to them. So on the night of the New Hampshire primary in 2012, we were running varnish on our own servers because we didn't have access to the Post CDN at that time. We ran out of file handles, so our Apache instances that were running on those servers slowly throttled themselves to death, because they couldn't open any new network sockets.

And as a result, we were down for about 15 or 20 minutes and not showing results pages to the world. And that was the time when, basically right after that election was over, I slept in the next day. I didn't go into work, but the Thursday after that, I walked in and I said to my boss, "That's it. We're going to the cloud. I am tired of all of this physical server baloney. We can rent servers. We can buy as many of them as we want for like six hours and then we can just turn them all off."

Corey: Less on the election side, but it also seems to me just from a perspective of what actual real newspapers do with protecting sources and whatnot, you're one of the few people who can say that information security does have people's lives on the line.

Jeremy: Yes, absolutely. And there are definitely cases where we step back from doing things in the cloud. We do a lot of document analysis, which requires us to do OCR. And a lot of the good cheap OCR in the world is cloud based, which requires us to upload potentially hundreds of gigabytes of PDFs up to either Amazon or Google to apply their OCR software to it. So one of the things that I worked on at The Times and that I'm working on here at The Post is a large local system for OCRing and transcribing documents, so that if we have things that are really sensitive, we don't have to think twice about where we're putting them.

And sure, we could solve some of those problems with cryptography, but that really feels like and now you have two problems sort of situation, where just having like a beefy server here with like 96 cores that can just rip through a bunch of tesseract and turn a bunch of PDFs into actual texts that our reporters can look at, that feels like a thing that is okay for us to say, "Maybe this one doesn't have to end up in US East One."

Corey: It seems like it's one of those areas where there's a number of companies that, "Oh no, what we do is so secret. It's our secret sauce and our problems are so special and beautiful and unique that no one can handle these as well as we possibly could." And oh, so what did you spend the most time on last year? Oh, replacing failed hard drives. Good, good. That sounds like a differentiated thing that everyone should be focusing on.

Jeremy: A core competency.

Corey: One question that I think is probably going to be on a number of folks' minds is The Washington Post is owned by a patron for lack of a better term, Jeff Bezos, the founder and CEO of Amazon. Is there any business relationship between The Washington Post and Amazon other than one might expect for, "Oh, we buy pencils off of them" or "We use them for cloud services"? Is there editorial oversight? Is there, "Oh, you can use whatever you want as long as it's AWS, because that is owned by the same owner"?

Jeremy: You know, I have to say, early on even before Jeff Bezos bought the company, our need to go to the cloud in some way predated that quite a bit. And actually, the guy who is now our CIO used to be the CTO. His name is Shailesh Prakash. He had come over from Microsoft via Sun Microsystems, so he's a computing OG. And he basically got here to The Post, looked at this litter of hardware that we were running in multiple data centers including one in our own building, and he had said, "This is ridiculous we can't be doing this," to your point about changing out had drives, like isn't really a core competency of ours to pull servers out of racks and make network cables all day long. Probably not, right?

So there was this big push just about the time that I was getting ready to leave The Post in like 2012 to move basically everything to the cloud. And at the time, basically only AWS existed. I think there were some other sort of lightweight solutions that were available. I think there was Linode and like a handful of other things like that, but AWS had basically scaled, like large scaled kinds of things that we could use, like tooling that we could use. So we made the call, we moved to AWS. It was grand. I left to go to The New York Times. The Times of course made a huge move from AWS to Google while I was there, I think in 2017 or thereabouts. And so that was like the, I have to tell you that was one of the most breathtakingly difficult things that I have ever worked on. It's like trying to learn all the new idioms. That's basically the same thing only slightly different, between two huge cloud computing giants.

You know, when I come back to The Post, it was basically like going home a little bit, getting comfortable again with AWS. But you know, to answer the question more specifically, I think that if we ever felt the need to move along, I don't know that I would ever have a problem with that. Our relationship with AWS is basically, "Please fix this problem. We're a really good AWS customer," and not "Please fix this problem. We're going to call Jeff in the middle of the night and have him nuke it from-"

Corey: You can threaten almost anything you want, it turns out.

Jeremy: Only I could, yes.

Corey: You can threaten anything. It doesn't mean you actually have to be able to follow through on it. I mean, I swear. My entire escalation point for most of these things is simply Twitter. Start tweeting something obnoxious and there you go.

Jeremy: Well, we beg our AWS reps when they come to town. We have our punch list of things that we're desperate to get fixed. Web sockets and a handful of other little things like this, but you know, they know when they come here, they're basically just going to deal with us the same way that they deal with any other customer. Except I think we're slightly nicer to them.

Corey: Well, from my perspective, that seems like a relatively low bar.

Jeremy: Yeah, fair. Fair. I don't think that they get like the ... I don't think that that's a job that I could ever do, travel between clients. Basically all you get to do is hear them, hear folks be upset about like what isn't working this week.

Corey: And you have no direct impact on anything that needs to be fixed, because effectively, "Hi, I'm your account manager, I'm here for you to abuse," more or less, and it turns out that my understanding of how large companies work is fatally flawed. I'm a five person company, so I assume every other company is, too. I assume when Jeff Barr, AWS's chief evangelist isn't frantically writing blog posts, he's building the service he's about to write a blog post on. I assume everything he writes about he built himself single-handedly, because why wouldn't he. It turns out companies don't actually work that way. So you're very often having to cajole folks internally to get status updates, to get information about what's going on in a timely manner to customers and in a way that doesn't inspire blind panic.

Yeah, so it turns out that regions down and we're not entirely sure why, and when I called the team all I heard was screaming and then the line got disconnected, so I don't really know what to tell you. Huh, why's our stock down 40%? Yeah, it's one of those areas where messaging is important, message discipline is important. And outcomes are always challenging. On some level, I feel like that is aligned in some ways with the role that journalism has to play. In many cases, the journalist does not get to become part of the story. It's about understanding the other players and telling a story about what's going on. Storytelling is one of those, I think disappearing arts as we sort of descend down the well into page views and click bait and trying to drive outrage more or less instead of actual journalism.

Given the hot button issues that a lot of your reporting tends to focus on, collectively as a group and what you're working on specifically, that extra care has to be taken not just to avoid conflicts of interest or anything in that vein, but rather the appearance of conflicts of interest.

Jeremy: Yeah, absolutely. A thing that I ... I wasn't a journalist in high school or college. And honestly until maybe even a few years ago, it was really hard to even think of myself as like working ... I worked for journalism organizations, but I did not self identify as a journalist. And even now I think like my team, even though they literally sit in the newsroom and we literally work on election results, there still feels in my head like there's a little disconnect. But in truth, we're working on something that is what I would consider to be like the family jewels of The Washington Post. Elections coverage is practically sacrosanct around here.

I have this, I'm lucky, I was blessed with this team of engineers, many of whom have never worked on elections before, but who are just working on side projects that were so clearly demonstrated an interest in this that I had no choice but to go thieve them from the directors that were running their teams. But they have all decided that, like as a team we decided that we were going to follow the newsroom standards for how we conduct ourselves during an election year, which means everyone on the team, we don't vote in primaries. We don't give money to political campaigns. We don't put up yard signs. And we do this not because it's like ... because we think that that'll make us fairer. We do it because it's critical that we understand our place and how people perceive the work that we do. And so it's not fair. It's not fair to say to my engineers, "You can't have an opinion about how our world works."

But it is fair for us to say about ourselves, "You know, we're going to do the same. We're going to treat ourselves the same way that like say a political journalist would.. and we're going to use those same standards of objectivity." I think that it is really great on the part of my engineers to sort of take this on and I think that it's kind of great on the part of The Post to trust us with those parts that are so clearly important to the organization.

Corey: This episode is sponsored by ExtraHop. ExtraHop provides threat detection and response for the enterprise (not the starship)! On-prem security doesn't translate well to cloud or multicloud environments, and that's not even counting IOT. ExtraHop automatically discovers everything inside the perimeter, including your cloud workloads and IOT devices, detects these threats up to 35% faster, and helps you act immediately. Ask for a free trial of detection and response for AWS today at extrahop.com/trial.

Corey: One of the hardest things that I tend to see on the Internet is whenever The Washington Post or The New York Times or someone else does, has a columnist or an op ed or writes an article in a way that the Internet, by which of course I mean Twitter, it's all the same thing these days, perceives to be a bad take, then suddenly the entire world goes nuts with, "I'm canceling my subscription." Okay. I get the outrage and I get wanting to send a message, but by the same token, there's an awful lot that's good and everyone writes a bad article or has a bad episode once in a while. Not this one, mind you. But at some point, okay, wow. With all these people canceling all the time, where do they ever have any customers? They must be out of business.

Then you do a little digging, wait a minute, this is the third time this year this person claimed they're canceling their subscription. So on some level, it feels like it's a big histrionic and it also feels on some level that you folks can't actually win. Where no matter what you write, you're going to wind up irritating some folks. Part of that is the element of speaking truth to power and part of it also increasingly feels like we're all living in our own bubbles, however we tend to surround ourselves. And I'm as guilty of that as anyone else. I tend to assume every random person I pass on the street has an active Twitter account, that they absolutely care about weird nomenclature from giant web services companies, and they have an unhealthy aversion to people mispronouncing acronyms.

It turns out that a lot of those are not factually true, and every time I encounter someone that shakes me out of that viewpoint, I have to stop and reevaluate. I do wonder how that winds up evolving over time, because it doesn't feel like it used to be quite this, I guess partisan. Then again, I feel like we've been saying that forever too and maybe it's just the world isn't changing, I'm just getting older and my perspective's shifting.

Jeremy: Right? I was just reading Poisoning the Press, which is a book about Nixon and his relationship with the press in the 1970s. Because you assume that everybody that you meet has a Twitter presence and cares about the mispronunciation of acronyms. I assume that everyone that I meet has a pile of unread books that's approximately three and a half years old and that they're barely making their way through them. So I have finally made my way through this book and the thing that struck me about it was it described a time that felt not super unfamiliar to our current one. It felt like adversarial relationship between the president and the press corps. It felt like real similar levels of like discontent. And one of the things I thought that was really intriguing about that was that those times in the mid and late 1970s was also like a really difficult time for newspapers.

There was like an advertising bust around that time that made things really difficult. It was like leading people to sort of question their business models. I recall that The New York Times around that time made the decision to break out a real sports section and focus on their features desk after a series of reader surveys indicated that these are things that people were interested in. And so the thing that I ... I don't think think I disagree with your general feeling about this world. A thing that I am sort of intrigued by is the sort of instant feedback that we get from people, histrionic or not. It's definitely a feeling that somebody has. And especially when it comes to things like elections. I really enjoy getting more or less instant feedback on things that we're trying out in various election nights. We had this Virginia election very recently where the state of Virginia just held elections for its house of delegates and its state senate, which are now controlled by democrats for the first time, the house, the senate, and the governorship are controlled by democrats for the first time.

This was an interesting opportunity for us to get to test out some new elections features and it wasn't the whole Internet that was excited about it. It was a lot of Washington Post subscribers in the Virginia and DC area, buy still it was really nice to get that sort of instant feedback about what people liked and didn't like. We take some of it with a grain of salt, like obviously there's folks who are going to be pretty upset. But on election night, it's not like we're getting people who are like, "We're going to cancel our subscription because your tables were right aligned." The stakes feel like they're a little bit lower for us, honestly.

But it is interesting and it's a difficult ... Elections are one of those places where a newspaper has like a real opportunity to kind of reach out and connect to readers. Everybody wants to know what happened or what's happening and who's winning and what's happening in our democracy. It's one of those places where we just have like a real opportunity to show trust and reward people's trust in us. And honestly, for me it's like the most exciting time to work is like on an election night when the results are flowing in and you're sitting in that room watching all the stuff happen, it's just like it's great. It's like being at the nerve center of a democracy.

Corey: It really is interesting seeing how some of the sausage gets made to some extent. I happened to be in New York City last weekend at the time of this recording and I walked past The New York Times building and holy crap, that thing's big. And then you remember there are bureaus scattered around the world as well and huh, it never occurred to me that that newspaper that shows up or the apps that constantly get refreshed have more than a couple dozen people working there. Turns out it's kind of a massive undertaking for any of this stuff at scale.

Jeremy: That's absolutely the case. You know, and the funny thing is like even here at The Washington Post, we had like an election night for us involved something like 30 people. About 15 or 20 of us huddled in a room, about 10 or 12 working remotely or via Slack. And that's just for an off year election in our backyard, a Virginia election and a Kentucky governor and a Mississippi governor. And then if you can scale that in your head to what super Tuesday or the New Hampshire primaries are going to look like, then you start to get a picture of like what it's like to be working in a place where the entire organization is just like laser focused on a single night and a single experience. It's just crazy. The energy in the building is palpable.

I mean, you wouldn't want to drag your feet on the carpet because you'd probably set off a thunderclap. It's a little crazy. But honestly, to me it's just like one of the reasons why I enjoy working for a newspaper, because it's rare ... Normally if you were going to go work at a newspaper, you would have to work at a place where you were not using a whole bunch of your technological savvy so to speak, a thing that I thoroughly enjoy about working at The Post is that I get to do, like work on hard engineering problems, but I also get to do it at a place where on an election night, basically everybody stops and is watching to see what just happened. Who did we just elect? What's happening? How are we going to explain this to readers? It's a good time.

Corey: It does feel like it's something that has a little bit more permanence than a lot of other things. But I'm not trying to talk smack about various software as a service companies or startups or whatnot. But The Washington Post, The New York Times, these are societal institutions that have been here since before most of us were alive and will presumably be here long after we're gone. We hope. I mean, you folks do have the permanent headline to the top of your website, "Democracy dies in darkness," which is just, you know, it's taking a while and the article isn't done yet, but I'm sure it's in progress.

Jeremy: It's so goth. Basically, The New York Times is "All the truth that's fit to print." And The Washington Post's is "Democracy dies in darkness." It feels like these are the two sides, the yin and yang if you will. I reward my employer for having the guts to go with the slightly grumpier, slightly goth-er tone. We're scrappy. Underdogs. Democracy dies in darkness.

Corey: Exactly.

Jeremy: Get your t-shirt.

Corey: Jeremy, thank you so much for taking the time to speak with me. If people want to hear more about your thoughts on these and other topics, where can they find you?

Jeremy: Well, I am on the tweets, because I am a person that you have bumped into before. So naturally I have a Twitter presence. I am @JeremyBowers on Twitter.

Corey: Excellent. And we'll put links to all of that in the show notes as well. Thanks again for taking the time to speak with me when you could've been doing literally anything else.

Jeremy: That's okay. When I'm done with this, I'm going to go to a meeting about making sure that we have the right campaign finance data available for reporters when the filing deadline hits in January 31st of next year.

Corey: Exciting. And I have no doubt you'll get there on time. Jeremy Bowers, engineering director at The Washington Post. I'm Corey Quinn. This is Screaming in the Cloud. If you've enjoyed this episode, please leave a five star review on iTunes. If you've hated this episode, please leave a five star review in iTunes.

This has been this week's episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com or wherever fine snark is sold.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Kelsey Hightower
Kelsey Hightower is a principal developer advocate at Google, the co-chair of KubeCon, the world’s premier Kubernetes conference, and an open source enthusiast. He’s also the co-author of Kubernetes Up & Running: Dive into the Future of Infrastructure.

Links Referenced

  • Twitter: @kelseyhightower
  • Company site: Google.com
  • Book: Kubernetes Up & Running: Dive into the Future of Infrastructure

Transcript

Announcer: Hello and welcome to Screaming in the Cloud, with your host Cloud economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored by CloudZero. CloudZero wants you to know a few things. One: they realize it's not that original to put a random noun after the word cloud to name your company. Yet somehow they did it anyway. Two: most cost management solutions weren't designed for engineering teams, and three, spending $50,000 on something by accident would be bananas in virtually any other industry. That's why CloudZero built a product that helped engineers correlate their costs with the engineering activity that caused it, so you can quickly detect anomalous span, investigate the source of any cost, and slice up your AWS costs by the metrics that matter to you. Go to cloudzero.com to kick off a free trial that's cloud Z E R O.com and my thanks to them for sponsoring this episode.

Corey: This episode is brought to you by DigitalOcean, the cloud provider that makes it easy for startups to deploy and scale modern web applications with, and this is important to me: No billing surprises. With simple, predictable pricing that's flat across 12 global data center regions, and a UX developers around the world love, you can control your cloud infrastructure costs and have more time for your team to focus on growing your business. See what businesses are building on DigitalOcean and get started for free at do.co/screaming. That's D-O-Dot-C-O-slash-screaming, and my thanks to DigitalOcean for their continuing support of this ridiculous podcast.

Corey: Welcome to Screaming in the Cloud, I'm Corey Quinn. I'm joined this week by Kelsey Hightower, who claims to be a principal developer advocate at Google, but based upon various keynotes I've seen him in, he basically gets on stage and plays video games like Tetris in front of large audiences. So I assume he is somehow involved with e-sports. Kelsey, welcome to the show.

Kelsey: You've outed me. Most people didn't know that I am a full-time e-sports Tetris champion at home. And the technology thing is just a side gig.

Corey: Exactly. It's one of those things you do just to keep the lights on, like you're waiting to get discovered, but in the meantime, you're waiting table. Same type of thing. Some people wait tables you more or less a sling Kubernetes, for lack of a better term.

Kelsey: Yes.

Corey: So let's dive right into this. You've been a strong proponent for a long time of Kubernetes and all of its intricacies and all the power that it unlocks and I've been pretty much the exact opposite of that, as far as saying it tends to be over complicated, that it's hype-driven and a whole bunch of other, shall we say criticisms that are sometimes bounded in reality and sometimes just because I think it'll be funny when I put them on Twitter. Where do you stand on the state of Kubernetes in 2020?

Kelsey: So, I want to make sure it's clear what I do. Because when I started talking about Kubernetes, I was not working at Google. I was actually working at CoreOS where we had a competitor Kubernetes called Fleet. And Kubernetes coming out kind of put this like fork in our roadmap, like where do we go from here? What people saw me doing with Kubernetes was basically learning in public. Like I was really excited about the technology because it's attempting to solve a very complex thing. I think most people will agree building a distributed system is what cloud providers typically do, right? With VMs and hypervisors. Those are very big, complex distributed systems. And before Kubernetes came out, the closest I'd gotten to a distributed system before working at CoreOS was just reading the various white papers on the subject and hearing stories about how Google has systems like Borg tools, like Mesa was being used by some of the largest hyperscalers in the world, but I was never going to have the chance to ever touch one of those unless I would go work at one of those companies.

So when Kubernetes came out and the fact that it was open source and I could read the code to understand how it was implemented, to understand how schedulers actually work and then bonus points for being able to contribute to it. Those early years, what you saw me doing was just being so excited about systems that I attended to build on my own, becoming this new thing just like Linux came up. So I kind of agree with you that a lot of people look at it as a more of a hype thing. They're looking at it regardless of their own needs, regardless of understanding how it works and what problems is trying to solve that. My stance on it, it's a really, really cool tool for the level that it operates in, and in order for it to be successful, people can't know that it's there.

Corey: And I think that might be where part of my disconnect from Kubernetes comes into play. I have a background in ops, more or less, the grumpy Unix sysadmin because it's not like there's a second kind of Unix sysadmin you're ever going to encounter. Where everything in development works in theory, but in practice things pan out a little differently. I always joke that ops is the difference between theory and practice. In theory, devs can do everything and there's no ops needed. In practice, well it's been a burgeoning career for a while. The challenge with this is Kubernetes at times exposes certain levels of abstraction that, sorry certain levels of detail that generally people would not want to have to think about or deal with, while papering over other things with other layers of abstraction on top of it. That obscure, valuable troubleshooting information from a running something in an operational context. It absolutely is a fascinating piece of technology, but it feels today like it is overly complicated for the use a lot of people are attempting to put it to. Is that a fair criticism from where you sit?

Kelsey: So I think the reason why it's a fair criticism is because there are people attempting to run their own Kubernetes cluster, right? So when we think about the cloud, unless you're in OpenStack land, but for the people who look at the cloud and you say, "Wow, this is much easier." There's an API for creating virtual machines and I don't see the distributed state store that's keeping all of that together. I don't see the farm of hypervisors. So we don't necessarily think about the inherent complexity into a system like that, because we just get to use it. So on one end, if you're just a user of a Kubernetes cluster, maybe using something fully managed or you have an ops team that's taking care of everything, your interface of the system becomes this Kubernetes configuration language where you say, "Give me a load balancer, give me three copies of this container running." And if we do it well, then you'd think it's a fairly easy system to deal with because you say, "kubectl, apply," and things seem to start running.

Just like in the cloud where you say, "AWS create this VM, or G cloud compute instance, create." You just submit API calls and things happen. I think the fact that Kubernetes is very transparent to most people is, now you can see the complexity, right? Imagine everyone driving with the hood off the car. You'd be looking at a lot of moving things, but we have hoods on cars to hide the complexity and all we expose is the steering wheel and the pedals. That car is super complex but we don't see it. So therefore we don't attribute as complexity to the driving experience.

Corey: This to some extent feels it's on the same axis as serverless, with just a different level of abstraction piled onto it. And while I am a large proponent of serverless, I think it's fantastic for a lot of Greenfield projects. The constraints inherent to the model mean that it is almost completely non-tenable for a tremendous number of existing workloads. Some developers like to call it legacy, but when I hear the term legacy I hear, "it makes actual money." So just treating it as, "Oh, it's a science experiment we can throw into a new environment, spend a bunch of time rewriting it for minimal gains," is just not going to happen as companies undergo digital transformations, if you'll pardon the term.

Kelsey: Yeah, so I think you're right. So let's take Amazon's Lambda for example, it's a very opinionated high-level platform that assumes you're going to build apps a certain way. And if that's you, look, go for it. Now, one or two levels below that there is this distributed system. Kubernetes decided to play in that space because everyone that's building other platforms needs a place to start. The analogy I like to think of is like in the mobile space, iOS and Android deal with the complexities of managing multiple applications on a mobile device, security aspects, app stores, that kind of thing. And then you as a developer, you build your thing on top of those platforms and APIs and frameworks. Now, it's debatable, someone would say, "Why do we even need an open-source implementation of such a complex system? Why not just everyone moved to the cloud?" And then everyone that's not in a cloud on-premise gets left behind.

But typically that's not how open source typically works, right? The reason why we have Linux, the precursor to the cloud is because someone looked at the big proprietary Unix systems and decided to re-implement them in a way that anyone could run those systems. So when you look at Kubernetes, you have to look at it from that lens. It's the ability to democratize these platform layers in a way that other people can innovate on top. That doesn't necessarily mean that everyone needs to start with Kubernetes, just like not everyone needs to start with the Linux server, but it's there for you to build the next thing on top of, if that's the route you want to go.

Corey: It's been almost a year now since I made an original tweet about this, that in five years, no one will care about Kubernetes. So now I guess I have four years running on that clock and that attracted a bit of, shall we say controversy. There were people who thought that I meant that it was going to be a flash in the pan and it would dry up and blow away. But my impression of it is that in, well four years now, it will have become more or less system D for the data center, in that there's a bunch of complexity under the hood. It does a bunch of things. No-one sensible wants to spend all their time mucking around with it in most companies. But it's not something that people have to think about in an ongoing basis the way it feels like we do today.

Kelsey: Yeah, I mean to me, I kind of see this as the natural evolution, right? It's new, it gets a lot of attention and kind of the assumption you make in that statement is there's something better that should be able to arise, giving that checkpoint. If this is what people think is hot, within five years surely we should see something else that can be deserving of that attention, right? Docker comes out and almost four or five years later you have Kubernetes. So it's obvious that there should be a progression here that steals some of the attention away from Kubernetes, but I think where it's so new, right? It's only five years in, Linux is like over 20 years old now at this point, and it's still top of mind for a lot of people, right? Microsoft is still porting a lot of Windows only things into Linux, so we still discuss the differences between Windows and Linux.

The idea that the cloud, for the most part, is driven by Linux virtual machines, that I think the majority of workloads run on virtual machines still to this day, so it's still front and center, especially if you're a system administrator managing BDMs, right? You're dealing with tools that target Linux, you know the Cisco interface and you're thinking about how to secure it and lock it down. Kubernetes is just at the very first part of that life cycle where it's new. We're all interested in even what it is and how it works, and now we're starting to move into that next phase, which is the distro phase. Like in Linux, you had Red Hat, Slackware, Ubuntu, special purpose distros.

Some will consider Android a special purpose distribution of Linux for mobile devices. And now that we're in this distro phase, that's going to go on for another 5 to 10 years where people start to align themselves around, maybe it's OpenShift, maybe it's GKE, maybe it's Fargate for EKS. These are now distributions built on top of Kubernetes that start to add a little bit more opinionation about how Kubernetes should be pushed together. And then we'll enter another phase where you'll build a platform on top of Kubernetes, but it won't be worth mentioning that Kubernetes is underneath because people will be more interested on the thing above.

Corey: I think we're already seeing that now, in terms of people no longer really care that much what operating system they're running, let alone with distribution of that operating system. The things that you have to care about slip below the surface of awareness and we've seen this for a long time now. Originally to install a web server, it wound up taking a few days and an intimate knowledge of GCC compiler flags, then RPM or D package and then yum on top of that, then ensure installed, once we had configuration management that was halfway decent.

Then Docker run, whatever it is. And today feels like it's with serverless technologies being what they are, it's effectively a push a file to S3 or it's equivalent somewhere else and you're done. The things that people have to be aware of and the barrier to entry continually lowers. The downside to that of course, is that things that people specialize in today and effectively make very lucrative careers out of are going to be not front and center in 5 to 10 years the way that they are today. And that's always been the way of technology. It's a treadmill to some extent.

Kelsey: And on the flip side of that, look at all of the new jobs that are centered around these cloud-native technologies, right? So you know, we're just going to make up some numbers here, imagine if there were only 10,000 jobs around just Linux system administration. Now when you look at this whole Kubernetes landscape where people are saying we can actually do a better job with metrics and monitoring. Observability is now a thing culturally that people assume you should have, because you're dealing with these distributed systems. The ability to start thinking about multi-regional deployments when I think that would've been infeasible with the previous tools or you'd have to build all those tools yourself. So I think now we're starting to see a lot more opportunities, where instead of 10,000 people, maybe you need 20,000 people because now you have the tools necessary to tackle bigger projects where you didn't see that before.

Corey: That's what's going to be really neat to see. But the challenge is always to people who are steeped in existing technologies. What does this mean for them? I mean I spent a lot of time early in my career fighting against cloud because I thought that it was taking away a cornerstone of my identity. I was a large scale Unix administrator, specifically focusing on email. Well, it turns out that there aren't nearly as many companies that need to have that particular skill set in house as it did 10 years ago. And what we're seeing now is this sort of forced evolution of people's skillsets or they hunker down on a particular area of technology or particular application to try and make a bet that they can ride that out until retirement. It's challenging, but at some point it seems that some folks like to stop learning, and I don't fully pretend to understand that. I'm sure I will someday where, "No, at this point technology come far enough. We're just going to stop here, and anything after this is garbage." I hope not, but I can see a world in which that happens.

Kelsey: Yeah, and I also think one thing that we don't talk a lot about in the Kubernetes community, is that Kubernetes makes hyper-specialization worth doing because now you start to have a clear separation from concerns. Now the OS can be hyperfocused on security system calls and not necessarily packaging every programming language under the sun into a single distribution. So we can kind of move part of that layer out of the core OS and start to just think about the OS being a security boundary where we try to lock things down. And for some people that play at that layer, they have a lot of work ahead of them in locking down these system calls, improving the idea of containerization, whether that's something like Firecracker or some of the work that you see VMware doing, that's going to be a whole class of hyper-specialization. And the reason why they're going to be able to focus now is because we're starting to move into a world, whether that's serverless or the Kubernetes API.

We're saying we should deploy applications that don't target machines. I mean just that step alone is going to allow for so much specialization at the various layers because even on the networking front, which arguably has been a specialization up until this point, can truly specialize because now the IP assignments, how networking fits together, has also abstracted a way one more step where you're not asking for interfaces or binding to a specific port or playing with port mappings. You can now let the platform do that. So I think for some of the people who may be not as interested as moving up the stack, they need to be aware that the number of people we need being hyper-specialized at Linux administration will definitely shrink. And a lot of that work will move up the stack, whether that's Kubernetes or managing a serverless deployment and all the configuration that goes with that. But if you are a Linux, like that is your bread and butter, I think there's going to be an opportunity to go super deep, but you may have to expand into things like security and not just things like configuration management.

Corey: Let's call it the unfulfilled promise of Kubernetes. On paper, I love what it hints at being possible. Namely, if I build something that runs well on top of Kubernetes than we truly have a write once, run anywhere type of environment. Stop me if you've heard that one before, 50,000 times in our industry... or history. But in practice, as has happened before, it seems like it tends to fall down for one reason or another. Now, Amazon is famous because for many reasons, but the one that I like to pick on them for is, you can't say the word multi-cloud at their events. Right. That'll change people's perspective, good job. The people tend to see multi-cloud are a couple of different lenses.

I've been rather anti multi-cloud from the perspective of the idea that you're setting out day one to build an application with the idea that it can be run on top of any cloud provider, or even on-premises if that's what you want to do, is generally not the way to proceed. You wind up having to make certain trade-offs along the way, you have to rebuild anything that isn't consistent between those providers, and it slows you down. Kubernetes on the other hand hints at if it works and fulfills this promise, you can suddenly abstract an awful lot beyond that and just write generic applications that can run anywhere. Where do you stand on the whole multi-cloud topic?

Kelsey: So I think we have to make sure we talk about the different layers that are kind of ready for this thing. So for example, like multi-cloud networking, we just call that networking, right? What's the IP address over there? I can just hit it. So we don't make a big deal about multi-cloud networking. Now there's an area where people say, how do I configure the various cloud providers? And I think the healthy way to think about this is, in your own data centers, right, so we know a lot of people have investments on-premises. Now, if you were to take the mindset that you only need one provider, then you would try to buy everything from HP, right? You would buy HP store's devices, you buy HP racks, power. Maybe HP doesn't sell air conditioners. So you're going to have to buy an air conditioner from a vendor who specializes in making air conditioners, hopefully for a data center and not your house.

So now you've entered this world where one vendor does it make every single piece that you need. Now in the data center, we don't say, "Oh, I am multi-vendor in my data center." Typically, you just buy the switches that you need, you buy the power racks that you need, you buy the ethernet cables that you need, and they have common interfaces that allow them to connect together and they typically have different configuration languages and methods for configuring those components. The cloud on the other hand also represents the same kind of opportunity. There are some people who really love DynamoDB and S3, but then they may prefer something like BigQuery to analyze the data that they're uploading into S3. Now, if this was a data center, you would just buy all three of those things and put them in the same rack and call it good.

But the cloud presents this other challenge. How do you authenticate to those systems? And then there's usually this additional networking costs, egress or ingress charges that make it prohibitive to say, "I want to use two different products from two different vendors." And I think that's-

Corey: ...winds up causing serious problems.

Kelsey: Yes, so that data gravity, the associated cost becomes a little bit more in your face. Whereas, in a data center you kind of feel that the cost has already been paid. I already have a network switch with enough bandwidth, I have an extra port on my switch to plug this thing in and they're all standard interfaces. Why not? So I think the multi-cloud gets lost in the chew problem, which is the barrier to entry of leveraging things across two different providers because of networking and configuration practices.

Corey: That's often the challenge, I think, that people get bogged down in. On an earlier episode of this show we had Mitchell Hashimoto on, and his entire theory around using Terraform to wind up configuring various bits of infrastructure, was not the idea of workload portability because that feels like the windmill we all keep tilting at and failing to hit. But instead the idea of workflow portability, where different things can wind up being interacted with in the same way. So if this one division is on one cloud provider, the others are on something else, then you at least can have some points of consistency in how you interact with those things. And in the event that you do need to move, you don't have to effectively redo all of your CICD process, all of your tooling, et cetera. And I thought that there was something compelling about that argument.

Kelsey: And that's actually what Kubernetes does for a lot of people. For Kubernetes, if you think about it, when we start to talk about workflow consistency, if you want to deploy an application, queue CTL, apply, some config, you want the application to have a load balancer in front of it. Regardless of the cloud provider, because Kubernetes has an extension point we call the cloud provider. And that's where Amazon, Azure, Google Cloud, we do all the heavy lifting of mapping the high-level ingress object that specifies, "I want a load balancer, maybe a few options," to the actual implementation detail. So maybe you don't have to use four or five different tools and that's where that kind of workload portability comes from. Like if you think about Linux, right? It has a set of system calls, for the most part, even if you're using a different distro at this point, Red Hat or Amazon Linux or Google's container optimized Linux.

If I build a Go binary on my laptop, I can SCP it to any of those Linux machines and it's going to probably run. So you could call that multi-cloud, but that doesn't make a lot of sense because it's just because of the way Linux works. Kubernetes does something very similar because it sits right on top of Linux, so you get the portability just from the previous example and then you get the other portability and workload, like you just stated, where I'm calling kubectl apply, and I'm using the same workflow to get resources spun up on the various cloud providers. Even if that configuration isn't one-to-one identical.

Corey: This episode is sponsored in part by DataStax. The NoSQL event of the year is DataStax Accelerate in San Diego this May from the 11th through the 13th. I've given a talk previously called the myth of multi-cloud, and it's time for me to revisit that with... A sequel! Which is funny given that it's a NoSQL conference, but there you have it. To learn more, visit datastax.com that's D-A-T-A-S-T-A-X.com and I hope to see you in San Diego. This May.

Corey: One thing I'm curious about is you wind up walking through the world and seeing companies adopting Kubernetes in different ways. How are you finding the adoption of Kubernetes is looking like inside of big E enterprise style companies? I don't have as much insight into those environments as I probably should. That's sort of a focus area for the next year for me. But in startups, it seems that it's either someone goes in and rolls it out and suddenly it's fantastic, or they avoid it entirely and do something serverless. In large enterprises, I see a lot of Kubernetes and a lot of Kubernetes stories coming out of it, but what isn't usually told is, what's the tipping point where they say, "Yeah, let's try this." Or, "Here's the problem we're trying to solve for. Let's chase it."

Kelsey: What I see is enterprises buy everything. If you're big enough and you have a big enough IT budget, most enterprises have a POC of everything that's for sale, period. There's some team in some pocket, maybe they came through via acquisition. Maybe they live in a different state. Maybe it's just a new project that came out. And what you tend to see, at least from my experiences, if I walk into a typical enterprise, they may tell me something like, "Hey, we have a POC, a Pivotal Cloud Foundry, OpenShift, and we want some of that new thing that we just saw from you guys. How do we get a POC going?" So there's always this appetite to evaluate what's for sale, right? So, that's one case. There's another case where, when you start to think about an enterprise there's a big range of skillsets. Sometimes I'll go to some companies like, "Oh, my insurance is through that company, and there's ex-Googlers that work there." They used to work on things like Borg, or something else, and they kind of know how these systems work.

And they have a slightly better edge at evaluating whether Kubernetes is any good for the problem at hand. And you'll see them bring it in. Now that same company, I could drive over to the other campus, maybe it's five miles away and that team doesn't even know what Kubernetes is. And for them, they're going to be chugging along with what they're currently doing. So then the challenge becomes if Kubernetes is a great fit, how wide of a fit it isn't? How many teams at that company should be using it? So what I'm currently seeing as there are some enterprises that have found a way to make Kubernetes the place where they do a lot of new work, because that makes sense. A lot of enterprises to my surprise though, are actually stepping back and saying, "You know what? We've been stitching together our own platform for the last five years. We had the Netflix stack, we got some Spring Boot, we got Console, we got Vault, we got Docker. And now this whole thing is getting a little more fragile because we're doing all of this glue code."

Kubernetes, We've been trying to build our own Kubernetes and now that we know what it is and we know what it isn't, we know that we can probably get rid of this kind of bespoke stack ourselves and just because of the ecosystem, right? If I go to HashiCorp's website, I would probably find the word Kubernetes as much as I find the word Nomad on their site because they've made things like Console and Vault become first-class offerings inside of the world of Kubernetes. So I think it's that momentum that you see across even People Oracle, Juniper, Palo Alto Networks, they're all have seem to have a Kubernetes story. And this is why you start to see the enterprise able to adopt it because it's so much in their face and it's where the ecosystem is going.

Corey: It feels like a lot of the excitement and the promise and even the same problems that Kubernetes is aimed at today, could have just as easily been talked about half a decade ago in the context of OpenStack. And for better or worse, OpenStack is nowhere near where it once was. It would felt like it had such promise and such potential and when it didn't pan out, that left a lot of people feeling relatively sad, burnt out, depressed, et cetera. And I'm seeing a lot of parallels today, at least between what was said about OpenStack and what was said about Kubernetes. How do you see those two diverging?

Kelsey: I will tell you the big difference that I saw, personally. Just for my personal journey outside of Google, just having that option. And I remember I was working at a company and we were like, "We're going to roll our own OpenStack. We're going to buy a free BSD box and make it a file server. We're going all open sources," like do whatever you want to do. And that was just having so many issues in terms of first-class integrations, education, people with the skills to even do that. And I was like, "You know what, let's just cut the check for VMware." We want virtualization. VMware, for the cost and when it does, it's good enough. Or we can just actually use a cloud provider. That space in many ways was a purely solved problem. Now, let's fast forward to Kubernetes, and also when you get OpenStack finished, you're just back where you started.

You got a bunch of VMs and now you've got to go figure out how to build the real platform that people want to use because no one just wants a VM. If you think Kubernetes is low level, just having OpenStack, even OpenStack was perfect. You're still at square one for the most part. Maybe you can just say, "Now I'm paying a little less money for my stack in terms of software licensing costs," but from an extraction and automation and API standpoint, I don't think OpenStack moved the needle in that regard. Now in the Kubernetes world, it's solving a huge gap.

Lots of people have virtual machine sprawl than they had Docker sprawl, and when you bring in this thing by Kubernetes, it says, "You know what? Let's reign all of that in. Let's build some first-class abstractions, assuming that the layer below us is a solved problem." You got to remember when Kubernetes came out, it wasn't trying to replace the hypervisor, it assumed it was there. It also assumed that the hypervisor had APIs for creating virtual machines and attaching disc and creating load balancers, so Kubernetes came out as a complementary technology, not one looking to replace. And I think that's why it was able to stick because it solved a problem at another layer where there was not a lot of competition.

Corey: I think a more cynical take, at least one of the ones that I've heard articulated and I tend to agree with, was that OpenStack originally seemed super awesome because there were a lot of interesting people behind it, fascinating organizations, but then you wound up looking through the backers of the foundation behind it and the rest. And there were something like 500 companies behind it, an awful lot of them were these giant organizations that ... they were big e-corporate IT enterprise software vendors, and you take a look at that, I'm not going to name anyone because at that point, oh will we get letters.

But at that point, you start seeing so many of the patterns being worked into it that it almost feels like it has to collapse under its own weight. I don't, for better or worse, get the sense that Kubernetes is succumbing to the same thing, despite the CNCF having an awful lot of those same backers behind it and as far as I can tell, significantly more money, they seem to have all the money to throw at these sorts of things. So I'm wondering how Kubernetes has managed to effectively sidestep I guess the open-source miasma that OpenStack didn't quite manage to avoid.

Kelsey: Kubernetes gained its own identity before the foundation existed. Its purpose, if you think back from the Borg paper almost eight years prior, maybe even 10 years prior. It defined this problem really, really well. I think Mesos came out and also had a slightly different take on this problem. And you could just see at that time there was a real need, you had choices between Docker Swarm, Nomad. It seems like everybody was trying to fill in this gap because, across most verticals or industries, this was a true problem worth solving. What Kubernetes did was played in the exact same sandbox, but it kind of got put out with experience. It's not like, "Oh, let's just copy this thing that already exists, but let's just make it open."

And in that case, you don't really have your own identity. It's you versus Amazon, in the case of OpenStack, it's you versus VMware. And that's just really a hard place to be in because you don't have an identity that stands alone. Kubernetes itself had an identity that stood alone. It comes from this experience of running a system like this. It comes from research and white papers. It comes after previous attempts at solving this problem. So we agree that this problem needs to be solved. We know what layer it needs to be solved at. We just didn't get it right yet, so Kubernetes didn't necessarily try to get it right.

It tried to start with only the primitives necessary to focus on the problem at hand. Now to your point, the extension interface of Kubernetes is what keeps it small. Years ago I remember plenty of meetings where we all got in rooms and said, "This thing is done." It doesn't need to be a PaaS. It doesn't need to compete with serverless platforms. The core of Kubernetes, like Linux, is largely done. Here's the core objects, and we're going to make a very great extension interface. We're going to make one for the container run time level so that way people can swap that out if they really want to, and we're going to do one that makes other APIs as first-class as ones we have, and we don't need to try to boil the ocean in every Kubernetes release. Everyone else has the ability to deploy extensions just like Linux, and I think that's why we're avoiding some of this tension in the vendor world because you don't have to change the core to get something that feels like a native part of Kubernetes.

Corey: What do you think is currently being the most misinterpreted or misunderstood aspect of Kubernetes in the ecosystem?

Kelsey: I think the biggest thing that's misunderstood is what Kubernetes actually is. And the thing that made it click for me, especially when I was writing the tutorial Kubernetes The Hard Way. I had to sit down and ask myself, "Where do you start trying to learn what Kubernetes is?" So I start with the database, right? The configuration store isn't Postgres, it isn't MySQL, it's Etcd. Why? Because we're not trying to be this generic data stores platform. We just need to store configuration data. Great. Now, do we let all the components talk to Etcd? No. We have this API server and between the API server and the chosen data store, that's essentially what Kubernetes is. You can stop there. At that point, you have a valid Kubernetes cluster and it can understand a few things. Like I can say, using the Kubernetes command-line tool, create this configuration map that stores configuration data and I can read it back.

Great. Now I can't do a lot of things that are interesting with that. Maybe I just use it as a configuration store, but then if I want to build a container platform, I can install the Kubernetes kubelet agent on a bunch of machines and have it talk to the API server looking for other objects you add in the scheduler, all the other components. So what that means is that Kubernetes most important component is its API because that's how the whole system is built. It's actually a very simple system when you think about just those two components in isolation. If you want a container management tool that you need a scheduler, controller, manager, cloud provider integrations, and now you have a container tool. But let's say you want a service mesh platform. Well in a service mesh you have a data plane that can be Nginx or Envoy and that's going to handle routing traffic. And you need a control plane. That's going to be something that takes in configuration and it uses that to configure all the things in a data plane.

Well, guess what? Kubernetes is 90% there in terms of a control plane, with just those two components, the API server, and the data store. So now when you want to build control planes, if you start with the Kubernetes API, we call it the API machinery, you're going to be 95% there. And then what do you get? You get a distributed system that can handle kind of failures on the back end, thanks to Etcd. You're going to get our backs or you can have permission on top of your schemas, and there's a built-in framework, we call it custom resource definitions that allows you to articulate a schema and then your own control loops provide meaning to that schema. And once you do those two things, you can build any platform you want. And I think that's one thing that it takes a while for people to understand that part of Kubernetes, that the thing we talk about today, for the most part, is just the first system that we built on top of this.

Corey: I think that's a very far-reaching story with implications that I'm not entirely sure I am able to wrap my head around. I hope to see it, I really do. I mean you mentioned about writing Learn Kubernetes the Hard Way and your tutorial, which I'll link to in the show notes. I mean my, of course, sarcastic response to that recently was to register the domain Kubernetes the Easy Way and just re-pointed to Amazon's ECS, which is in no way shape or form Kubernetes and basically has the effect of irritating absolutely everyone as is my typical pattern of behavior on Twitter. But I have been meaning to dive into Kubernetes on a deeper level and the stuff that you've written, not just the online tutorial, both the books have always been my first port of call when it comes to that. The hard part, of course, is there's just never enough hours in the day.

Kelsey: And one thing that I think about too is like the web. We have the internet, there's webpages, there's web browsers. Web Browsers talk to web servers over HTTP. There's verbs, there's bodies, there's headers. And if you look at it, that's like a very big complex system. If I were to extract out the protocol pieces, this concept of HTTP verbs, get, put, post and delete, this idea that I can put stuff in a body and I can give it headers to give it other meaning and semantics. If I just take those pieces, I can bill restful API's.

Hell, I can even bill graph QL and those are just different systems built on the same API machinery that we call the internet or the web today. But you have to really dig into the details and pull that part out and you can build all kind of other platforms and I think that's what Kubernetes is. It's going to probably take people a little while longer to see that piece, but it's hidden in there and that's that piece that's going to be, like you said, it's going to probably be the foundation for building more control planes. And when people build control planes, I think if you think about it, maybe Fargate for EKS represents another control plane for making a serverless platform that takes to Kubernetes API, even though the implementation isn't what you find on GitHub.

Corey: That's the truth. Whenever you see something as broadly adopted as Kubernetes, there's always the question of, "Okay, there's an awful lot of blog posts." Getting started to it, learn it in 10 minutes, I mean at some point, I'm sure there are some people still convince Kubernetes is, in fact, a breakfast cereal based upon what some of the stuff the CNCF has gotten up to. I wouldn't necessarily bet against it socks today, breakfast cereal tomorrow. But it's hard to find a decent level of quality, finding the certain quality bar of a trusted source to get started with is important. Some people believe in the hero's journey, story of a narrative building.

I always prefer to go with the morons journey because I'm the moron. I touch technologies, I have no idea what they do and figure it out and go careening into edge and corner cases constantly. And by the end of it I have something that vaguely sort of works and my understanding's improved. But I've gone down so many terrible paths just by picking a bad point to get started. So everyone I've talked to who's actually good at things has pointed to your work in this space as being something that is authoritative and largely correct and given some of these people, that's high praise.

Kelsey: Awesome. I'm going to put that on my next performance review as evidence of my success and impact.

Corey: Absolutely. Grouchy people say, "It's all right," you know, for the right people that counts. If people want to learn more about what you're up to and see what you have to say, where can they find you?

Kelsey: I aggregate most of outward interactions on Twitter, so I'm @KelseyHightower and my DMs are open, so I'm happy to field any questions and I attempt to answer as many as I can.

Corey: Excellent. Thank you so much for taking the time to speak with me today. I appreciate it.

Kelsey: Awesome. I was happy to be here.

Corey: Kelsey Hightower, Principal Developer Advocate at Google. I'm Corey Quinn. This is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple podcasts. If you've hated this podcast, please leave a five-star review on Apple podcasts and then leave a funny comment. Thanks.

Announcer: This has been this week's episode of Screaming in the Cloud. You can also find more Core at screaminginthecloud.com or wherever fine snark is sold.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Rob Zuber
Rob Zuber is a 20-year veteran of software startups; a four-time founder, three-time CTO. Since joining CircleCI, Rob has seen the company through its Series B, Series C, and Series D funding and delivered on product innovation at scale. Rob leads a team of 150+ engineers who are distributed around the globe.

Prior to CircleCI, Rob was the CTO and Co-founder of Distiller, a continuous integration and deployment platform for mobile applications acquired by CircleCI in 2014. Before that, he cofounded Copious an online social marketplace. Rob was the CTO and Co-founder of Yoohoot, a technology company that enabled local businesses to connect with nearby consumers, which was acquired by Appconomy in 2011.

Links Referenced

  • Twitter: @z00b
  • LinkedIn URL: https://www.linkedin.com/in/robzuber/
  • Personal site: https://www.crunchbase.com/person/rob-zuber#section-overview
  • Company site: www.circleci.com

Transcript
Announcer: Hello, and welcome to Screaming in the Cloud with your host cloud economist, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored by CloudZero. CloudZero believes that, despite what some would have you think, the health of a company isn't measured by how much you spend on EC2. That's why they'll help you measure your spend in terms your business may actually care about. Like whether you should sunset that feature that no one uses because it's costing you a fortune and eating away at your margins. CloudZero will also alert you right away when you push out a feature that, say, unknowingly doubles your S3 costs before you get hit with that enormous bill. Go to cloudzero.com to kick off a free trial, and my thanks to them for sponsoring this episode.

This episode is brought to you by DigitalOcean, the cloud provider that makes it easy for startups to deploy and scale modern web applications with, and this is important to me: No billing surprises. With simple, predictable pricing that's flat across 12 global data center regions and a UX developers around the world love, you can control your cloud infrastructure costs and have more time for your team to focus on growing your business. See what businesses are building on DigitalOcean and get started for free at do.co/screaming. That's D-O-Dot-C-O-Slash-Screaming, and my thanks to DigitalOcean for their continuing support of this ridiculous podcast.

Welcome to screaming in the cloud. I'm Corey Quinn. I'm joined this week by Rob Zuber, CTO of CircleCI. Rob, welcome to the show.

Rob: Thanks. Thanks for having me. It's great to be here.

Corey: It really is, isn't it? So you've been doing the CTO dance, for lack of a better term, at CircleCI for about five, six years now at this point?

Rob: Yeah, that's right. I joined five and a half years ago. I actually came in through an acquisition. We were building a CI/CD platform for mobile, iOS specifically, and there were just a few of us. I came in an engineering role, but within, I think a year, had taken over the CTO role and have been doing that since.

Corey: For those of us who've been living under a rock and recording podcasts, CI/CD or Continuous Integration/Continuous Delivery has gone through a bit of, shall we say, evolution since the term first showed up. My first exposure to it many moons ago was back when Jenkins was still called Hudson, and it was the box that you ran that it would wait for some event to happen, whether it was the passing of time, a commit to a particular branch, someone clicked a button, and then it would run a series of scripts, which sort of lent itself to the idea of the hacker news anthem, "That doesn't look hard. I can build that in a weekend." Now, we've seen a bit of growth in that space of not just, I guess the systems you can run yourselves, but also a lot of the SaaS offerings around this. That's the, I guess, the morons journey from my perspective to path through CI/CD. That's almost certainly lacking nuance. What is it, I guess in the real world with adults talking about it?

Rob: Yeah, so I think it's a good perspective, or it's a good description of the perspective that many people have. Many people enter into this feeling that way. I think, specifically when you talk about cloud providers in CircleCI, we do have an on-prem offering behind the firewall. No one really runs anything on-prem anymore. But we have an offering for that market, but the real leverage is for folks that can use our stuff, multi-tenant SaaS cloud offering. Because, ultimately it's true. Many people have start with something simple from a code based perspective, right? I'm starting out, I've got a small team. We have a pretty simple project, maybe a little monolith Ruby on rails, something like that. Actually, I think in the time of the start of CircleCI. Probably not too many people kick off the rails monolith these days because if you're not using Kubernetes and Docker, then you're probably not doing it right.

Corey: So, the Kubernetes and Docker people tell us?

Rob: Yeah, exactly. They will proudly tell you that. We'll come back around to that point if we want to, but so you have simple project and you have simple CI, right? You may just have a simple script that you're putting in a Jenkins box or something like that, but what ultimately ends up happening is it gets complicated, and as it gets complicated, it becomes a bigger and bigger distraction from the thing that you're really trying to do, right? You're trying to build a business to ... I don't know, to do ride hailing, to do scooter sharing, what's big these days. You might be trying to do any of the ...

Corey: Oh, my project is Twitter for pets. We're revolutionizing the world of pet communication.

Rob: Right. And do you want to spend your time working on pet communication or on CI/CD, right? CI/CD is a thing that we understand very well, we spend our time on it every day, we think about some of the depths of it, which we can go into in a second. One of the things that gets complicated, amongst others, is just scale. So you build a big team, you have multiple projects and you have that one box under your desk where you said, "Oh, it's not that hard to build CI/CD. Now, everybody's waiting for their stuff to run because someone else got in there before them and you're thinking, okay, well how do I buy ... maybe you're not buying more boxes, you're building out something in a cloud provider and then you're worrying about auto scaling because it starts to cost you too much to run those boxes, and how do you respond to the amount of load that you have on any given day?

Because you're crunching for a deadline versus everybody's taken a week off. Then, you want to get your build done as quickly as possible. So you start figuring out how to paralyze the work and spread it across those machines. The list goes on and on. This is the reality that everyone runs into as they scale their work. We do that for you. While it seems simple and ... I said I came in through an acquisition, we were building CI/CD for iOS, and I was that person. I said, "This seems really simple. We should build it and put it in the market." It didn't take us very long to get that first version to build, and it had to be generic to support many different types of customers and their particular builds.

It was a small start but the ... we started to run into the same problems, and then of course as a business, we ran into the problem of getting access to customers and all those things and that's why we joined CircleCI and that became what is now our iOS offering. But there is a lot of value that you can get quickly, to your point, but then you start focusing time and energy on that. I often refer to it, others in the industry refer to these sorts of things as undifferentiated heavy lifting. Something that becomes big and complex over time and is not the core of your business. Then as you start to invest in it, as we invest in it, then we build capabilities that most people wouldn't bother to build when they write that first bash script off a trigger or whenever, around helping you get your project set up, handling the connection into hooks, handling authentication so that different users only have access to the code they should have access to, maybe isolating access to production secrets, for example, if you're doing deploy.

The kinds of things that keep coming up over and over in CI/CD that people don't think about on that first pass but ended up hunting them down the road.

Corey: What do you think that people tend to misunderstand the most about CI/CD as you take a look at that throughout the ecosystem? From my perspective, when it was a box that you ran, behind the firewall as you say, the problem was is that everyone talked about, "Oh yes, we use cattle, not pets, except the box that does the builds. Of course, that box has a bunch of hand-built stuff on it that's impossible to replicate. It has extraordinary permissions into production environments and can do horrifying things, and it was always the star of various security finding reports. There are a number of us who came up from an operation side viewing CI/CD as, in some ways, a liability, which I understand is a very biased and one sided perspective. But going beyond that, what are people missing? What are they not seeing about the CI/CD landscape?

Rob: One thing that I think is really interesting there, well, one thing you call that was just resiliency, right? We think about that in the way that we operate that system. We have a world of cattle because we've managed to think about that as a true offering. So, as you scale and you start to think, "Oh, how do I make this resilient inside my operation?" That's going to become a challenge that you face. The other thing that I think about that I've noticed over the years is, I want to call it division of labor or division of responsibilities. Many of those single instance or even multi-instance self-managed CI/CD tools end up in a place where, past any size of team, honestly somebody needs to own it and manage it to make sure it's stable.

The changes that you want to make as a developer are often tied to basically being managed by that administrator. To be a little clear, if I have a group responsible for running CI/CD, then, and I want to start building a different type of code or a different project, and it requires a plugin or an extension to the CI/CD platform or CI/CD tool, then I need to probably file a ticket and wait for another department who is generally not super motivated to get my code out into production, to go make a change that they are going to evaluate and review and decide ... or maybe creates conflict with something somebody else is doing on that system. And then you say, "Oh well actually we can't have these co-installed so now we need two systems." It's that division of responsibilities. Whereas, having built a multi-tenant cloud offering, we could never have that. There is no world in which our customers say to us, "Hey, we want this plugin installed. Can you go do that for us?"

Everything that is about how the development team thinks about their software and how they want their build to run, how they want their deploys to run, etc, needs to be in the hands of the developers, and everything that is about maintenance and operation and scale needs to be in our hands. It has created a very clear separation out of necessity, but one that even ... I mentioned that you can deploy CircleCI yourself and run it within a team, and in large organizations, that separation really helps them get leverage. Does that make sense?

Corey: It really does. I think we're also seeing a change in perspective around resiliency and how this works. I once worked at a company I will not name where they were. It was either CircleCI or TeamCity. This was years and years ago where I don't recall exactly what they were using, but it doesn't matter because at one point the service took an outage, and in typical knee jerk reaction, well, that can never happen again. So they wound up doing all of the CI/CD work for some godforsaken reason on a Raspberry PI that some developer brought in and left in the corner of the office. Surprise, it took an awfully long time for tests to run on basically an underpowered toy project. The answer there was to just use less tests because you generally don't need to run nearly as many.

I just stared at people for the longest time when it came to that. I think that one of the problems that we still see, I know when I write code myself, I'm as guilty of this as anyone, I am a terrible developer and don't believe in tests. So, the CI/CD pipeline that I tend to look at is more or less a glorified script runner. Whenever I make a commit to this branch, go ahead and run the following three lines script that does a serverless deployment and puts it where it needs to go, and then I'll test it manually, or it's a pre-production environment so it's not that big of a deal. That can work for some use cases, but it's also a great thing that no one actually depends on the stuff that I write for day-to-day business operations or anything critical. At what point does it stop being a script runner?

Rob: Well, to the point of the scale, I think there's a couple of things that you brought up in there that are interesting to me. One is the culture of testing. It feels like one of these areas of software development, because I was around in a time when no one really understood what it was to do automated testing. I won't even go into TDD, but just, in general, why would I do that? We have this QA team, it's cost effective to give it to a bunch of people. I'm thinking backwards or thinking back on that, it all seems a little bit well, wrong. But getting to the point where you've worked effectively with tests takes a little bit of effort. But once you have that, once you've sat and worked on something and had the feedback loop of, oh, this thing's not working. Oh, I'll just change this, now it's working.

Really having that locally, as a developer, is super rewarding, in my mind and enabling I guess I would say as well. Then you get to this place where you're excited about building tests, especially as you're working in a team, and then culturally you end up in a place where, I put up a PR and someone else looks at it and says, "I see you're making an assumption or I believe you're making an assumption here, but I don't see any way that that's being validated. So please add testing to ensure that is actually true." Both because I want to make sure it's true now, but when we both forget that you ever wrote this and someone else makes a change, your assumptions hold or someone can understand that you were making those assumptions and they can make appropriate changes to deal with it.

I think as you work in a team that's growing and scaling and beyond your pet project, once you've witnessed the value of that, you don't want to go back. So, people do end up writing more and more tests and that's what drives the scale at least on the testing and CI side in a way that you need to then manage that. Going the opposite direction of what you're describing, which is, hey, let's just write fewer tests and use cheaper machines, people are recognizing the value and saying, "Okay, we want that value, but we don't want to bottleneck everyone with an hour long build to run all these. So how do we get a system that's going to scale and support that?"

Corey: That's what's fascinating, is watching that start to percolate beyond the traditional web applications with particular blessed languages and into other things. For example, in my copious spare time, I'm the community lead for the open guide to AWS, which is a GitHub project that has 25,000 stars or so, so you know it's good, where it's just a giant markdown document that lists the 10,000 tips and tricks that we all wish we'd known when we'd gotten started with AWS, and in a format that's easily consumable. The CI/CD approach we have right now, which I believe is done through Travis, is it just winds up running a giant link checker in parallel across the thousands of links that are ... sorry, I wanted to say 1,200 links, that are included within that document.

There's really not a lot else we can do in that type of environment. I mean, a spellchecker with all of the terms of art involved would more or less a seg fault itself to death as soon as it took a look, but other than making sure we don't have dead links, and it feels like there's not a lot of automation or testing opportunity in something like that. Is that accurate? Am I completely wrong and missing something?

Rob: I've never built that particular site so it ... I mean, it sounds reasonable. I think that going the other way, we often think about, before we kick off a large complex set of testing for a more complex application, maybe then a markdown document, a lot of people now will use things similar to what you're using, like maybe part of my application is a bunch of links to outside docs or outside sites that I'm referencing or if I run into a problem, I link you to our help site or something and making sure all that stuff is validated. Doing linting on the structure and format of code itself. One of the things that comes up as you scale out of the individual script runner is doing that work in parallel. I can say, you know what? Do the linting over here, do the link checking over here. Only use very small boxes for those.

We don't happen to have Raspberry Pi's in our infrastructure, but we can give you a much smaller resource, which costs you less if you're not going to be pushing the limits of that. But then, if you have big integration tests or something which need more space than we can provide that as well, both in a single channel or pathway to give you the room to move faster and then to break that out and break up your work. At an extreme example, and of course, anyone who's done parallelization knows there's costs to splitting up work in like the management overhead. But if you have 1200 links, like you could check them all at the same time. I doubt that would be a good use of our platform, but you could check 600 in one and 600 in another, or 300s at a time or whatever, in find the optimal path if you really cared about getting that done more quickly.

Corey: Right. Usually, it's not that big of a concern and usually it winds up throwing errors on existing bad links, not something that has been included in the pull request in question. Again, there's nothing that is so awesome that I can't horribly misuse it for something ridiculous. It's my entire stock and trade. It's why I believe route 53 remains the best database option for everyone, but it's fun going through this space and just seeing how things have evolved. One question I do have since you come from a background, by way of acquisition, that was aimed squarely at this, historically, it seems that running a lot of testing on mobile devices, specifically iOS devices, was the stuff of nightmares because you couldn't really run that in any meaningful way in a virtualized environment. So, it generally required an awful lot of devices. Is that still the case? Has that environment changed radically since I last worked at a mobile shop?

Rob: I don't think so, but I think we've all started to think a little bit differently. We got started in that business because we were building iOS apps and thought, wow, the tooling here, it's really frustrating. To be clear, at CircleCI and at that business, we were solving the problem of managing the machines themselves, so the portion of the testing that you would run effectively in a simulator, not the problem of the device farm, if you will. But one of the things that I remember, and so this is, maybe 2014. Late 2013, early 2014 as I was working on mobile apps was people shifting the MVC layers a little bit such that the thing that you needed to test on a device was getting smaller and smaller, meaning putting more logic in, I forget what the name was specifically, but it was like the ... I don't want to try to even guess.

But basically pulling logic out of the actual rendering and down into what we'll call state transitions I guess. If you think about that in modern day and look at maybe web frameworks like React, you're trying to just respond with rendering on top of a lot of state change that happens underneath that. In that model, if you thin out the user interface portion, you make a lot more of your code testable, if that makes sense. The reason we're all trying to test on all these different devices is often that we've baked a lot of business logic into the view layer. Does that make sense?

Corey: Yeah, it absolutely does. Please continue.

Rob: Instead of saying, well, all our logic's in the view layer, so let's get really good at testing the view layer, which means massive device farms and a bunch of people testing all these things, let's make that layer as thin as possible, and there's analogies for this in even how we do service design these days and structure the architecture of systems, basically make the boundaries as thin as possible and the interaction with the outside world as thin as possible. That gives you much more capability to effectively test the majority or much larger portions of your business logic. The device farm problem is still a problem. People still want to see how something specifically renders on a particular screen or whatever. But by minimizing that, the amount that you have to invest in that gets smaller.

Corey: You mentioned device farm, which is an app choice, given that that is the name of an AWS service that has a crap ton of mobile devices that you can log into and it's one of my top candidates for the, did I make this service up to mess with you competitions? It does lead us to an interesting question. CI/CD has gotten an increased amount of attention lately from pretty much everyone. AWS, as is typical for Amazon, tends to lie awake at night worrying that someone somehow is making money that isn't them. So their product strategy distills down to, yes. So, they wound up releasing a whole bunch of CI/CD oriented products that at launch were, to be polite, terrible. Over time, they've gotten slightly better, but it's still a very confusing ecosystem there.

Then we see things like Azure dev ops who it seems is aimed at a very similar type of problem and they're also trying to challenge Amazon on the grounds of terrible names of services. But we're now seeing an increased focus from the first party providers themselves around the CI/CD space. What does that mean for existing entrenched players who have been making a specialty out of this for a lot longer than these folks have been playing with it?

Rob: It's a great question. I think about the approaches very differently, which is probably unsurprising. Speaking of lying awake at night or spending all day thinking about these things, this is what we do. You've the term script runner a few times in the conversation, the thing that I see when I see someone like AWS looking at this problem is basically, people are using, the way that I think about it, is maybe less the money, although it translates pretty quickly. People are using compute to do something, can we get them to do that with us? Oddly enough, a massive chunk of CircleCI runs on AWS so it doesn't really matter to them one way or another, but they're effectively looking to drive compute hours and looking to drive a pathway onto their platform.

One thing about that is it doesn't really matter to them in my perspective, whether people use that particular product or not. As a result, it gets the product investment that you put in when that's the case. So, it's a sort of a check the box approach like, hey we CI and we have CD like other people do. Whereas, when we look at CI and CD, we've been talking about some of the factors like scaling it effectively and making it really easy for you to understand what's going on. We think about very much the core use case, what is one of our customers or users doing when they show up? How do we do that in a way that maximizes their flow? Minimizes the overhead to them of using our system, whether it's getting set up and running really quickly, like talk about being in the center of how much of the world is developing software.

So we see patterns, we see mistakes that people are making and can use that to inform both how our product works and inform you directly as a user. "Hey, I see that you're trying to do this. It would go better if you did this." I think both from the, honestly, the years that we've been doing this and the amount that we've witnessed in terms of what works well for customers, what doesn't, what we see going through just from a data perspective, as we see hundreds of thousands of builds running, that rich perspective is unique to us. Because as you said, we're a player that's been doing this for a really long time and very focused on it. We treat the experience with, I guess I'm trying to figure out a way to say this that doesn't sound as bad as it might, but a lot of people have suffered a lot with CI/CD.

There's a lot that goes into getting CI/CD to work effectively and getting it to work reliably over time as your system is constantly changing. Honestly, there's a lot of frustration, and we come in to work every day thinking about minimizing that frustration so that our customers can go spend their time doing what matters to them. Again, when I think you sort of ... a lot of these big players present you with a runtime in which you can execute a script of your choosing. It's not thinking about the problem in that way and I don't see them changing their perspective. Honestly, I just don't worry about them.

Corey: Which is a very fair tack to take. It's interesting watching companies and as far as how much time and energy they spend worrying about competition versus how much they focus instead on customers. To turn it around slightly, what makes what you do challenging in some respects, I would imagine is that a lot of your target market is themselves, developers. Developers, in my experience, are challenging customers in that, first, they tend to devalue their own time to the point where, oh, that doesn't sound hard. I'll build that overnight. Secondly, once you finally win them over to the idea of paying for something, it's challenging to get them to have the necessary signing authority. At best, they become champions. But what you do has to start with developers in order to win widespread adoption and technical buy-in. How does that wind up manifesting as approach to, well, some people call it developer relations, developer advocacy. I refer to those folks as developers because I have problems, but how do you folks view that?

Rob: Yeah, it's a really insightful view actually because we do end up in most of our customers, or in the environments of our customers, however you want to describe it, as a result of the enthusiasm of individual developers, development teams, much more so than ... there are many products certainly in enterprise software and I don't really think purely in enterprise, but there are many products that can only be purchased by the CIO or the CTO or whatever. Right? To your question of developer relations, we spend a lot of time out in the market talking to individuals, talking at conferences, writing content about how we think about this space and things that people can do. But we're a very product driven company, meaning both, that's what we think about first, and then support it with these other things.

But second, we win on product, right? We don't win in the market because you thought the blog posts that we wrote was really cool. That might make you aware of us, but if you don't love the product, I mean, developers, to your point, they want to use things that they really enjoy using. When developers use the product and love the product and they champion it and they get access because they might work on a side project or an open source project or maybe they worked in another company that used CircleCI and then they go somewhere else and they say, "What are we doing? Life is so much better for you Circle CI, those sorts of things. But it very much comes from the bottom up. It's pretty difficult to go into an organization and say, "Hey, you should push this down to all of your developers."

There's a lot of rejection that comes from developers on mandated tooling. We have to provide knowledge, we have to provide capabilities in our product that appealed to those other folks. For example, administrators of our tooling, or when it gets to the point where someone owns how you use CircleCI versus just being a regular user of the product. We have capabilities to support them around understanding what's happening, around creating shared capabilities that multiple teams can use, those sorts of things. But ultimately, we have to lead with product, we have to get in into the sort of hearts and minds of the developers themselves and then grow from there and everything we do from a marketing, developer relations myself, I spend a lot of time talking to customers who are out in the market, is all about propping up or helping raise awareness effectively. But there's nothing that we can do if the product doesn't meet the needs of our customers.

Corey: That's what it seems like it comes down to a fair bit. It's always weird to consider that, at its heart, developer relations is marketing. The folks I talk to who argue against that, it seems that it comes from a misunderstanding of what marketing actually is. It's not buying ads in airports, it's not doing podcast advertisements. That's a subject near and dear to my heart. It's not about annoying people by showing up at their office with the sales team. It's about understanding what their challenges and problems are and then positioning a solution that ideally solves them in a place that and in a way that they can be receptive to. Instead, people tend to equate marketing to this whole ridiculous statistics driven nonsense that doesn't really resonate with anyone and I think that that's unfair to everyone involved.

That said, I will say that having spent a fair bit of time in this space, I've yet to see anything from CircleCI that has annoyed me to the point where I would have remembered it, which is awesome. I don't see it in flight magazines, generally. I don't see it on obnoxious people try to tackle me as I walk through an expo hall and want to scan my badge. It just seems very well executed and you have some very talented people working for you. To that end, you are largely a distributed company, which is fascinating. Did it start that way? Did it happen that way by a quirk of fate?

Rob: Yeah, I those two things probably come together. The company, from very early days, now I wasn't there but I think some of our earliest engineers were distributed and the company started out basically entirely as engineers. It's a team solving problems of other engineers, which is ... it's a fun challenge. There were early participants who were distributed. Mostly, when you start a company and no one has ever heard of you and no one knows if you're going to be successful, going and recruiting is generally a different game than when you're, certainly, when you're where we are now. There were some personal relations that just happened to connect with people around the globe who wanted to participate.

We started out pretty early with some distribution, and that led to structuring the org in a way, both from a tooling and process perspective. A lot of that sort of happens organically, but building a culture that really supported that. I personally am based in the Bay Area, so we have headquarters in San Francisco, but it doesn't really make a difference if I go in versus just stay and work from home on any given day because the company operates in such a way that that distribution is completely normal.

Corey: We accidentally did the same thing. My business partner and I used to live across the street from each other and we decided to merge a week before he moved out of state to Portland. So awesome. Great. We have wonderful timing on all of these things. It's fun to build it from that way, build that way from the ground up. The challenge I've always seen is when you start off with having a centralized office and everyone's there, except this one person who, no matter how you try to work around it, is never as involved. So it feels like the sort of thing you've absolutely got to be building from day one, or otherwise, you're going to have a massive cultural growing pain as you try to get there.

Rob: Yeah, I think that's true. So I've actually been that one person. I, at some point in my career prior to CircleCI, was helping out a company founded by some friends of mine based in Toronto. I grew up in Toronto. I kicked off a project and then the project grew and grew until I was the one person out of maybe 50 or 60 who wasn't in an office in Toronto. It got to the point where no one remembered who I was and I was like, "Cool, I think I'm done. I'm out." I was fine with that. It was always meant to be a temporary thing, but I really felt that transition for the organization. I would say in terms of growing, I mean, yes, if you start out, it goes both ways, if you start out distributed, you're going to remain distributed.

There are certain things that get more challenging at scale, right? If everybody is sort of just in their home all over the globe, then the communication overhead continues to increase and increase in just understanding who people are, who you should be talking to. You need to focus-

Corey: There's always the time zone hierarchy.

Rob: Ooh, the time zones are a delight, yes. I would say like we talk a lot about, in this industry, Dunbar's number and sizes of teams and the points at which things get more complex. I think there's probably a different scale for distributed teams. It takes fewer people to reach a point where communication gets challenging, and trust and all the other things that go with Dunbar's views. You kind of have that challenge and then you start to think, oh well, then you have some offices, because we actually have maybe six physical offices, partly because in our go to market org, we've started to expand globally and put people in regional offices.

There's this interesting disconnect. I don't know about disconnect, but there's a split in how we operate in different parts of the org. I think what I've seen people ... well, I don't know about succeed, but I've seen people try when you start out with one org, or sorry, one location is, let's not jump to that one person somewhere else and then one person somewhere else kind of thing, but build out a second office, build out another office, like pick another location where you think you ... it's often, certainly where we are, in the Bay Area, it's often driven by just this market. Finding talent, finding people who want to join you, hanging onto those people when there are so many other opportunities around tends to be much more challenging. When you offer people alternatives, like you can stay where you are but have access to a cool and interesting company or you can work from home, which a lot of people value, then there's different things that you bring to the table.

I see a lot of people trying to expand in that way, but when you are so office-centric, a second office I think is a smoother transition point than just suddenly distributing people because, especially the first and second one, unless you're hiring in a massive wave, are really going to struggle in that environment.

Corey: I think that's probably one of the more astute things that's been noticed on this show in the last couple of years. If people want to hear more about what you have to say and how you think about the world, where can they find you?

Rob: I would say, on our blog, I tend to write stuff there as do other people. You talked about having great people in the organization. We have a lot of great people talking about how we think about engineering, how we think about both engineering teams and culture and then some of the problems we're trying to solve. So, off our site, circleci.com, and go to our blog. Then, I attend to is to speak and hangout on podcasts and do guest writing. I think I'm pretty easy to find. You can find me on Twitter. My handle is z00b, Z-0-0-B. I know I'm not super prolific, but if someone wants to track me down and ask me something, I'd probably be more than happy to answer.

Corey: You can expect some engagement as soon as this goes out. Thank you so much for taking the time to speak with me today. I appreciate it.

Rob: Yeah, thanks for having me. This was a ton of fun.

Corey: Rob Zuber, CTO at CircleCI. I'm Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave a five-star review on Apple podcasts. If you've hated this podcast, please leave a five-star review on Apple podcasts along with something amusing for me to read later while I'm crying.

Announcer: This has been this week's episode of Screaming in the Cloud. You can also find more corey@screaminginthecloud.com or wherever fine snark is sold.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Andreas Wittig
Andreas Wittig and Michael Wittig are freelancers, entrepreneurs, and authors. As freelancers, they are training, coaching, and consulting their clients on all things Amazon Web Services (AWS). In their role as an entrepreneur, Andreas and Michael are building SaaS products. The brothers have published two books Amazon Web Services in Actionand Rapid Docker on AWS and are blogging at cloudonaut.io.

Links Referenced

  • Twitter: @andreaswittig
  • LinkedIn: https://www.linkedin.com/in/andreaswittig/
  • Company site: https://cloudonaut.io
  • Books:
    • Rapid Docker on AWS
    • Amazon Web Services in Action

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud, with your host, Cloud Economist, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored by CloudZero. CloudZero doesn’t believe me when I say that trying to make your developers care about cost is a waste of time. Yet, they still sponsor my podcast, so here we are. Their platform is designed to help SAAS companies understand how and why their costs are changing, and lets easily you measure by things SAAS companies care about like cost per product or feature. Also, they won’t make you set up rules or budgets in advance, but will still alert you before your intern spends $10,000 on SageMaker on a Friday afternoon. Go to CloudZero.com to kick off a free trial. That’s CLOUD Z-E-R-O dot com, and my thanks to them for sponsoring this podcast.

Welcome to Screaming in the Cloud, I'm Corey Quinn. I'm joined this week by Andreas Wittig of cloudonaut. Andreas, welcome to the show.

Andreas: Thanks for having me, Corey.

Corey: You and your brother Michael are fascinating folks. You're freelancers, you're entrepreneurs and notably to the time of this recording, you're authors. You are the voice behind cloudonaut, which is fundamentally the thing that resonates throughout the community side of the ecosystem. But you're also doing a fair bit of consulting work through your company Widdix or Widdix. That's a “w” in front, but I'm not sure how it's pronounced.

Andreas: Yeah, so it's Widdix. Yeah, that's absolutely true. Our story actually is we started as software engineers and Michael and I ended up working in the same company, and we were looking for a way to get our software out to the world, to deploy it in a professional manner. That was actually hard, that time. Of course there was a single machine and some server rack, and deployments go wrong every now and then, more often than they actually deployed successfully. So we were looking for a solution and that was when we found out about AWS. This is now six years ago. Since then we have gotten deeper and deeper into that ecosystem and learned a lot over the years.

Corey: I first started doing my ridiculous newsletter about three years ago and one of the first things I started seeing when I was doing all of my, I guess, dog and pony show was a lot of references to the stuff that you and your brother had put out. It was fantastic to continue putting it into the newsletter. You were releasing a lot of interesting things. Midway through, about a year and a half in, I want to say, I realized that there were in fact two of you. As opposed to just one person who was super-productive, moving back and forth very quickly. So yeah, having two people does seem to make it easier, but if more or less whenever you folks wind up releasing something, it's shortlisted for inclusion in the newsletter, just because I very rarely see you writing anything that isn't useful or actionable to at least some subset of the population.

Andreas: That's a very great feedback. Thank you for that. I think what we try to do with writing on cloudonaut.io is we write about things that are out of interest for our self that are coming from the projects that we are working on. So I think that is actually the secret source behind all of that, is we write about the stuff that matters to us, to our clients and yeah, write blog posts out of that.

Corey: I guess I've sort of started doing the same thing and I don't think that I've been as, I guess, upfront about it as I probably should have been. I don't pretend that my stupid newsletter and accompanying podcasts are the authoritative list of what happened last week in AWS. For me, it's what happened that mattered. What happened that was interesting. There are times where I will cut three quarters of what's officially announced out just because it doesn't resonate with me personally. My biases absolutely factor into that.

I don't write a whole lot about Microsoft SQL Server. I don't write a whole lot about Kubernetes, largely because that doesn't resonate with what I'm working on. I don't spend time in those worlds for the most part. I guess from my own perspective, I don't see that that puts me in the place to be saying positive or negative, and too much commentary about anything that comes out in that space. Every time I dabble into it, I am immediately corrected loudly and violently by Twitter. So I try and stay in my own lane of things that matter to me. It sounds like you've been largely doing something very similar.

Andreas: Yes, that's absolutely true. But there's also funny stories, so sometimes I think I would never write about a certain topic. One of the topics for example is Oracle. That's a funny story. Oracle was really a topic that I was not interested at all, but recently I was actually migrating an Oracle APEX application to AWS and we managed to put that into a container and run it on Fargate. Now I'm so excited that we found a way to lift and shift this legacy application to a very cool platform on AWS, so that now I really want to write about Oracle on our blog. So sometimes the experience from our day to day work also changes what seems to be interesting to talk about in our writing.

Corey: I agree wholeheartedly. In fact, at the time of this recording, I am currently trying to figure out the best way to do something kind of obnoxious architecturally. The short version is I have a script that I run in various client accounts through assuming roles, but this script does run in my own account and it can take, in some cases, an hour or two to wind up completing. So that rules the traditional serverless story out of it, but I have to run it in every client account in every region because of course you can't globally query regions with AWS. Why could you? So it turns into a ... as you wind up with more regions being launched all the time and additional client work that needs to happen across a variety of different accounts ... like some customers are on a couple of hundred accounts easily ... that needs almost a step function sort of automation story.

I've been looking at Fargate with increasing interest, because it sounds like it's something that supports massive parallelization. As long as I can get this stuff stuffed into a container, and given that it is a single one-line script invocation with a few parameters to change, it seems like the right path. But other than that, I haven't gotten into it in any depth yet. So from that naive perspective, since real problems are always better than talking about theoretical ones, how would Fargate start to solve that problem? What is implementing something like this on top of Fargate look like? If that's a fair question. Because you know, getting free consulting under the guise of podcasting is always the right answer.

Andreas: Very clever. Very clever. Yeah, but I think that's a really cool thing to talk about. We are very big fans of Fargate since it was announced, and it has become a very convenient tool for deploying applications on AWS since then. Also integrations have become better and better. So the service maturity also increased a lot during the last years. Actually, I think of Fargate as a serverless compute service, like I think about Lambda but with less limitations. So we don't have the 15 minutes limitation. We don't have any limitations regarding our programming languages or libraries that we want to load. So Fargate is really easy to use.

The thing that you are talking about for example, so you have mentioned step functions and everyone knows that there is a good Lambda integration for it. So if you want to build workflows, you can use step functions, trigger Lambda functions and then build everything together, wire it together, with the state machine, and actually the same is possible with Fargate as well. So you can also trigger Fargate tasks or actually spin up containers with the help of a step function as well. And then the step function's state machine is taken care of, handling timeouts, retries and all of that in a similar way that it does with a Lambda function. So I think that's a very powerful concept to think about not only Lambda, but also Fargate as a serverless compute engine.

Corey: That absolutely resonates with my understanding of it. I think the piece that makes the most sense from my perspective when you say that, is the idea of, I have a thing, as long as I can fit it into a containerized task definition, I can throw it to Fargate and not worry about it. The premise of Lambda is similar, but there are enough, I guess, constraints around that where that doesn't map itself directly to every random bash script or Python script that I've beaten together. So as long as it runs in a container and I can hand it to Fargate, I don't have to think about it much beyond that. Is that more or less the way you see it?

Andreas: Yeah, that's absolutely true. So I watched a talk, I think it's four years ago, and it's still resonating with me. The guy in front of the audience said it doesn't make sense to run containers on AWS as long as our smallest resource that we can virtualize in the cloud is a virtual machine. So it doesn't make sense to operate your own cluster management on top of EC2 just to deploy containers to the cloud. Why are we doing that? I think he was absolutely right because a lot of people, us included, spend a lot of time and money on building clusters, container clusters on top of EC2, and that changed completely with the launch of Fargate.

Because now you can launch a container as you launch an EC2 instance, with the benefit that the container includes everything you need to run your application. You can run that on your local machine very simply, and your deployment pipeline as well. So it's a very effective tool to get your application, get your code, running on AWS. I think it's sometimes even a little bit more convenient than Lambda actually, because you can have your code running very easily on your machine deployed on AWS. That's the promise of containers, right? That's really working now because we don't have to care about the underlying infrastructure anymore.

Corey: From my side of the world focusing on cloud economics, I've noticed that Fargate has gone from interesting toy, but super expensive for anything non-trivial. There've been a couple of iterations on that. First we saw that it wound up getting a almost 40% price cut across the board, if memory serves, and suddenly it went from stratosphere down to the realm of reasonable. Then we took a step further and saw that the savings plans, their new RI replacement dingus, applies to Fargate as well. So, okay, great. If you commit to having compute in some form, whether it's on EC2 instances, it doesn't matter now which family, which region, or whether it's even EC2 or Fargate. As long as you're using compute, it offers it at a discount, and suddenly this becomes incredibly compelling just from a story of being able to execute applications that don't need to be shoved into artificially constrained environments.

Andreas: Yeah, absolutely. I think also savings plans really made Fargate much more competitive compared to running your workloads on EC2. So now we pay a premium on top of the EC2 price of about 20% to 30%. You could say that is a fair price for not having to deal with virtual machines anymore. So I absolutely think that this is a game changer for Fargate workloads because many of our customers told us, "Yes, Andreas, that's a very great idea. Fargate looks really cool, but why should we pay that much money for our infrastructure?" So I think that that argument is getting less and less important and so Fargate will in the future, I think, become more and more interesting. Also maybe let's see if ... Things that are missing is, for example, spot markets or spot instances is something that we don't have an equivalent with Fargate, but yeah, reserved instances, savings plans, we have that with Fargate now and I think that's a really cool thing.

Corey: Yes. We should also highlight that we are doing this recording of the week before re:Invent. So much of what we're about to say may not apply by the time this episode is published, but I have a sneaky suspicion that some of it will at least still be relevant. One challenge I've seen for some customers with Fargate as well is when they look under the hood at what it's running on from looking at CPU information, it appears at least to be running largely on C3 and M3 style instances, which is fine for an awful lot of workloads. But anything that's particularly compute-intensive, there's not the capability of using some of the newer generation of CPUs under the hood.

So if you have anything that's time-bound, it may not be the right fit either. But for just "run this script" I don't particularly care upon what you run it, and put all the answers into a queue somewhere," It's fine for that use case. What I think people also lose sight of is our time as engineers is never free. So the idea of, "Well, I could knock 30% off of that bill by running it on top of EC2 instances," and sure, if you're talking about multiple millions of dollars a year in spend, there's a terrific model for that and a market that makes sense, but for my toy example that I gave you, we're talking about, on a per-customer basis, something like $12 versus $20. Would I pay that marginal eight bucks in order to have this all handled for me? Hell yes, I would.

Andreas: Sure. I absolutely agree. So you don't have to convince me with stuff like that, but a lot of customers still need to be convinced...

Corey: Yes, yes. I probably should have warned you, we are in fact recording this. So some of my arguments are not aimed directly at you personally, but rather the people who are eagerly preparing to tweet at me about the conversation we're having right now.

Andreas: Okay, perfect.

Corey:

So you took it a step further. Rather than just mucking around with Fargate and other fun ways of getting dockerized things into production, you wrote a book on this that is going to be for sale well before the time that this winds up being published. So tell me about your book.

Andreas: Michael and me, we wrote our first book, Amazon Web Services in Action, and this year we decided it is about time to write our next book. We picked an area that we really liked and this was Fargate, and this was then why we started writing Rapid Docker on AWS. The book is, as the name implies, the focus is on rapid. So we want to make sure that the reader is getting into that as fast as possible, and can spin up the infrastructure that is needed to run an application. So a Ruby, PHP, Java and OTS Python application, as far as possible.

So we wrote a small book, it's only 120 pages. So it's to the point, it's only the relevant parts are in it. We also built a whole infrastructure as code example that really spins up the whole infrastructure that is needed. So it's consisting of EZS, Fargate, RDS, Aurora Serverless, and even the deployment pipeline. So everything bundled together. Of course monitoring is included with CloudWatch as well. We bundled that together so that as many people as possible can easily deploy their workloads on Fargate. So that is the intention of the book.

Corey: One of the challenges that I've seen in the past is that when you have books like this that are written around the easiest way to get from problem a customer has into solutions, whenever it starts talking about a particular service or path to getting to that outcome, it comes off, on some level, just at a cursory glance, "Oh, it's a sales pitch for service X." I understand why people think that. But having read through pretty much everything you folks have written over the last three years, you're not shills, you're not trying to sell services that aren't a fit for the task.

Your approach largely seems to mirror my own of, "What's the business requirement? Great. Now let's find the best way to get you to an outcome that satisfies that business requirement." So anyone who's thinking, "Oh, it's just a puff piece about an AWS service that they're trying to drive attention to," I disagree strongly. It's one of those services that suffers from not being the easiest to understand. Until people try something with it, it's very easy to overlook.

Andreas: Yeah. So you know, Michael and I, we always try to be also skeptical about the technology and the services that AWS announces. You might've read our reviews about services like Aurora Serverless or Global Accelerator, but we think with Fargate and a very simple setup with a load balancer, a Fargate run into containers, and an RDS Aurora database, we really think that this is a modern architecture to build on AWS that just focuses on simplicity. I think that's a very powerful tool if you're looking for something to get your stuff running as fast as possible without spending a month of engineering just for the infrastructure.

I think that's what a lot of people, a lot of developers actually are looking for. That's the focus of our book. Of course, it's not going into each and every detail, but we try to cover everything that is needed to also get the full picture and to understand how things work together. We also spent some pages on how to monitor and debug your infrastructure. So how do you find out about problems, how do you locate them, and how do you fix them? That is part of the book as well.

Corey: You bring up a couple of other interesting things that I've largely forgotten about. First, whenever I'm trying to deal with anything that even remotely touches on IAM. For example, one of my first resources that I go to is iam.cloudonaut.io. Because instead of playing hunt and check for two hours with the official documentation, I can drill down very quickly and see exactly what permissions are around a given service. It's incredibly useful and I can only assume that someone internally at AWS proposed this years ago and was immediately shot down on the basis of being too friendly to customers. Because this is hands down one of the most valuable resources whenever I'm playing around with IAM. The second most valuable resource of course is just great access to do everything, and I'm sure I'll remember to go back and fix it later.

Andreas: Yeah, I think we built the iam.cloudonaut reference when the documentation from AWS was not existent at all. It has gotten a lot better since then over the years. So it's getting less and less relevant. But I think still it's a very fast overview of all the IAM actions and resources and conditions that are available.

Corey: You also referenced something else that I thought was fantastic, which was your view of Global Accelerator. For those who aren't aware, Global Accelerator ideally improves latency and designing for failure around having traffic enter AWS's backbone, as close to the customer as possible, instead of waiting until they smack into the region they're going to, which is their default behavior. My AWS Morning Brief Thursday Networking in the Cloud show did have a segment on this that got some feedback, but your review of it was incredibly evenhanded.

It talked about the things it was great at and things that it wasn't terrifically great at, until I wound up reading this. My initial approach, that Global Accelerator was simply ... I'm not sure what it does so I imagined it makes the world spin faster, because Global Accelerator makes sense. Which in turn has got to be doing terrible things for climate change. At which point I was asked to please stop making that joke immediately because no one else found it nearly as funny as I did. You instead wound up doing some actual experimentation with it. Tell us about that.

Andreas: We are doing reviews of AWS services from time to time because I think it's interesting to look into the technical details of a service instead of just reading the marketing pages that AWS is putting out. One service that was announced at re:Invent last year was Global Accelerator, and it was on my to-do list for services to dive into a little bit deeper for now almost a year. Finally I found some time to actually look at it a little deeper. The thing is the promise of Global Accelerator are actually two promises. One promise is Global Accelerator will reduce the latency, the network latency from your clients to your AWS infrastructure. The second promise is you can actually build multi-region infrastructures and use Global Accelerator to route traffic to the closest and also healthy region that is available.

So I did some testing, and so the thing with multi-region deployments and routing traffic based on health checks, that was actually very easy and very simple, there was not much to have a look on. But the part of, "We're making it faster so we reduce latency," this was interesting because I actually wanted to measure what that means in reality. So what I did is I searched for a tool that allows me ... or actually a service that allows me to measure the network latency from all over the world to the AWS infrastructure, and then measure the latency to an application load balancer with and without Global Accelerator. And also compare that to CloudFront, which is actually a similar, or not very similar but another solution that is useful when you try to reduce latency and use that to dive a little bit deeper into the latency reductions that can be achieved with Global Accelerator.

Corey: When you take a look at that, do you think that it's a service that's worth using by default to speed up end-user traffic, or do you think that it's better reserved for specific use cases or specific problems that customers might see manifesting in their environment?

Andreas: I think the default service, if you run a web application that is doing something like HTTP, then you should probably have a look at CloudFront because there, you have the ability to cache requests at the edge. But if your workload is not HTTP-based or there might be other reasons to not use CloudFront, then I think it's actually very interesting to have a look at the Global Accelerator to just reduce the latency to your endpoints. As we all know, each and every millisecond that you can reduce typically helps in ecommerce, in gaming, in finance and so on. So it's probably valuable for a lot of scenarios.

Corey: What other services have you been seeing that you think don't have enough of a light being shone upon them? Again, we are recording this before re:Invent. So let's look back in time or I guess the previous year, before re:Invent hit. What do you think doesn't get enough attention, that people are taking far too lightly? It's very obvious from our conversation that Fargate was one of them. What else are you seeing?

Andreas: Actually today I played around with Amazon Connect, so the contact center solution, and I was really impressed about this service. I did know that this existed but I never used it before, and today I had a small use case to try it out and to build a small project with it. I think that's a very, very mature service with a really handsome UI. That's really something noteworthy because the UI is so easy to understand, you are creating a contact center within minutes and it's very easy to connect to different actions, everything. So I think that's another interesting service. Of course it's very niche, but definitely something that is interesting to have on the roadmap. Another one that I really, really like is Amazon Athena, the service that you can use to query data, unstructured data actually, that is stored on a tree with a SQL-like query language.

Corey: That is a fantastic point where the idea of having a SQL query language is, there are so many better ways to query things, et cetera, et cetera, et cetera. Yeah. Here in the real world though, everyone already knows it and no matter, even if you're not in the technical division, if you're a business analyst, you still know SQL or something very similar to it. So being able to have a language that is commonly spoken by the business as well means that suddenly you have an AWS service that's available to far beyond just the engineers or the "builders" as AWS likes to call them.

Andreas: Yeah, absolutely. I think Athena combines that, so the SQL engine and also the possibility to just store the data on a tree which is very convenient and very easy, also possibly cheap to do. Then query that data and analyze the data without, I don't know, spinning up a Redshift cluster or Elasticsearch cluster, something like that. So instead of that, you get a really easy way to analyze the data and you only pay for it on demand. So only when you run a query, you pay for the processed amount of data. So that is very handy. So I'm using it, for example, for the latency benchmark, so analyzing the results for Global Accelerator benchmark. I also use it for analyzing log files. So HTTP access logs, for example, or I'm using it a lot for also the EC2 network benchmark, for example. So there are many use cases where you just throw up some files on a tree and then use Athena to get some insights out of that.

Corey: I've been playing around with Athena a fair bit myself lately and there is some latency concerns there. The counterpoint is, the only thing to really get around that at significant scale that I played with has been Redshift. That apparently runs on burning piles of money because you need a lot of those to wind up paralyzing the queries appropriately.

Andreas: Yeah, absolutely. So Athena, and then another interesting service that relates to that is QuickSight. I know you always say they don't probably have any customers, but they have at least one.

Corey: I wish it were better than it is because I am burning a fortune on Tableau licensing instead.

Andreas: Yeah. They have at least one customer and that's me, and I'm building dashboards with QuickSight. I know it's not the greatest experience that you could imagine, but I think the big benefit to you is it integrates very well with Athena. And QuickSight has that in-memory cache that you can use to cache queries to your data, so you don't rely on a round trip to Athena and your objects on a tree every time. So that speeds up at least having a look at dashboards and visualizing your data a lot. So that helps sometimes with the latency for the queries running on Athena.

Corey: One other thing that I've ... Yeah, QuickSight is fantastic and I think that it is a service that has an awful lot of potential. If I'm being fully honest here ... because I imagine that people read my snark in the newsletter but no one listens to me when I talk into a microphone ... that I'm hoping that by my constant goading of QuickSight that they finally snap and a) fix it and b) bring me in to have a conversation. "All right. All right, jerk face. Here's exactly what we're doing. Tell me what ... how again this doesn't work for your use case," and then we will have an honest conversation and one of us will fix something. I'm not entirely sure whom, but it'll be a fun conversation.

Andreas: Yeah, that sounds great. I have something to add to that list as well, so...If you need some input...

Corey: Exactly.

Before we go, one more question for you. What do you do as a consultant? Because most of the stuff that I'm familiar with involves the things you're putting out into the world publicly. For example, the books, the write-ups, the blog posts, the reviews. What do you do that generates ideally the money that you use to feed yourselves? What's the consulting niche? What's it look like?

Andreas: Yeah, very good question. It comes down to maybe three main topics. It's technical coaching. We coach individuals or teams to get started with AWS or with a specific part of AWS. For example, containers on AWS or Serverless. That is one part that we do, we call it technical coaching. The second thing we do is infrastructure bootstrapping. We mentioned already that we have some CloudFormation templates open-sourced and we use them to spin up infrastructures very quickly. Usually we can get maybe 80% to 90% of the infrastructure that a project needs up and running with our base templates.

Then we only need to do 10% or 20% of the work to adopt these templates or to write new ones for the specific needs of a client. That is what we call infrastructure bootstrapping. This is typically something we do for something between one or four weeks. At the end we hand over the infrastructure, the documentation, and trained a team that is there in the future and taking care of that infrastructure. That is the second part. The third thing we do is we also ... I told you at the beginning, we are actually software engineers from the beginning. So we also write software or build stuff on AWS. So we do a lot of Serverless applications for our clients as well. That is the third thing that we typically do in our consulting business.

Corey: If people want to learn more about the various things you do, where can they find you?

Andreas: I think the best place to go is cloudonaut.io. This is our blog and it actually links to everything we do. So you can find our podcasts there. You find our blog posts also linked to our consulting business, to the book Rapid Docker on AWS. So that's basically the point that you'll find everything that we do, cloudonaut.io.

Corey: We'll definitely throw a link to that in the show notes. Thank you once again for taking the time to speak with me today. I appreciate it.

Andreas: Thanks for having me. It was a great pleasure.

Corey: Andreas Wittig, founder, entrepreneur, author, gadfly, et cetera, at cloudonaut.io and Widdix. I'm Cloud Economist Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this episode, please leave it a five-star review on Apple Podcasts. If you've hated this episode, please leave it a five-star review on Apple Podcasts.

Announcer: This has been this week's episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Thomas Hazel
Thomas Hazel is Founder, CTO, and Chief Scientist of CHAOSSEARCH. He is a serial entrepreneur at the forefront of communication, virtualization, and database technology and the inventor of CHAOSSEARCH's patent pending IP. Thomas has also patented several other technologies in the areas of distributed algorithms, virtualization and database science. He holds a Bachelor of Science in Computer Science from University of New Hampshire, and founded both student & professional chapters of the Association for Computing Machinery (ACM).

Links Referenced

  • Company site: http://chaossearch.io
  • Twitter: @ThomasHazel
  • LinkedIn: https://www.linkedin.com/in/thomashazel/

Transcript
Announcer: Hello and welcome to Screaming in the Cloud with your host, cloud economist, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Thomas Hazel, founder and CTO of everyone's favorite company, CHAOSSEARCH. Thomas, welcome to the show.

Thomas: Thank you for having me.

Corey: So, let's start at the very beginning. Most people don't spring into existence having founded a company. There's usually an origin story. What were you doing before CHAOSSEARCH?

Thomas: So I had a background in big distributor systems, made a career out of building God boxes back in the day in telecom. And then for the last 15 years, I'm working on new computer science with respects to database, data analytics, really trying to be inventive in the way of the force of computer science.

Corey: And for those who have been living under a rock and not been paying attention to any episode of anything I've ever done until this one, you folks have sponsored a number of different things that I've been involved in, including this episode. But for those who have not been paying attention, what does CHAOSSEARCH do?

Thomas: So at a high level, we created some innovative technology that allows customers to do analytics, text search, and both relational directly on their office storage, particularly Amazon S3. And we do it at a scale and cost and price point that is quite unique, quite disruptive in the market. And so, for the last four years we've been building out a new indexing technology as well as associated architecture as a service that customers log on to our service, within five minute registration, they're up and running doing terabytes of analysis, let's say for log analytics, within minutes and using their favorite API, Elastic API or coming out in 2020, SQL API.

Corey: Before you folks had ever sponsored anything that I was involved with, I was aware of you because you employed Pete Cheslock, everyone's favorite DevOps thought board, as the VP of product. When he started talking about what you folks did, I said, "It seems implausible that it's as amazing as you say, but all right, I'll suspend disbelief. Tell me about it."

And it turned out that everything that I was told was actually in fact correct. I recommended you folks to clients of mine, not because there was any business relationship, and that is still something people can't buy, as turns out. Credibility, it matters. But because for a certain subset of use cases, it is an incredibly cost effective approach to handling things. So, I guess this is probably an idiotic question, but I'm going to roll with it anyway. What is it that makes this such a unique thing? There's nothing else in the market like this, but there should be.

If we take a look at cloud, all the different providers took a vote and the storage technology that won by orders of magnitude in terms of price, has always been object store. Why is there nothing that provides a simple searchable API for a rapid response for data that lives in S3?

Thomas: Great, great question. So Mike Backrum as I mentioned, is in distributed computing and computer science and there's existing technology out there, Lucene, it's an inverted index, really driver force behind log analytics. There's comm storage for warehousing, B-trees, LSM-trees, all these classic computer science data structures and algorithms that have been used for the last 30 years. And the real issue is, is that they are at the breaking point of the scale that we're seeing today. Moore's Law says one thing, but machine generated data is outpacing it.

And about five to seven years ago, I wanted to crack that code of creating a new data structure and algorithm that could really provide the next level of scale, hundred thousand magnitude scale, without having to use massive amount of compute and a team of engineers trying to erect a system of terabyte and petabyte scale. So that was what I wanted to set out and go do.

And so I reached out to people like Pete to say, "Hey, I have this idea. What do you think if I solve this problem, will people care?" He kind of laughed a little bit. He says, "Everybody wants this problem solved and people want unicorns and a pot of gold. Call me when you figure it out." And a couple of years ago, I reached out to him. I said, "Hey Pete, I think I figured it out." And we went to market with that solution.

So really the essence of it is, I created a new representation, really a compression of them that's both a database index that can uniquely support text search, like the classic hunting and log analytics that a Lucene index would support, as well as relational analytics that you think warehousing technology with the same representation, without having to have siloed databases that you'd say, "I'm going to store data necessary for archival, but then move it out, ETL into say, a relational system or say Elasticsearch cluster."

You can imagine that is of great complexity and at scale, these systems start to break. And so I said, "What if you created a service that you leave the data in your storage?" Really, the idea of storage analytics convergence. And the power of object storage, amazon S3 was a great place to start where it's infinitely scalable, wonderfully inexpensive, secure and durable. But the problem is technology that was built 30 years ago cannot access it in a high performance way.

So people move it out of S3 and then into one of these classic databases. So I thought with this innovative technology that we have patent pending papers on that said I could take our index and create a new architecture that leverages distributed compute to the essence, unlock that storage that I had to move it out of the system.

And that is what we did. CHAOSSEARCH, as you can imagine, search in the chaos is why we came up with the name, because there's so much data being stored in S3 it's like all data roads lead to S3 and object storage and now every cloud provider has one. And so that's what we did over the last four years, build out a service that takes the customer's S3 account, they provide read only access, we provide the compute, we index the data and write these indices back into their account and then they can do a text search for say, log analytics as well as relational force, a business intelligence analytics.

And so for the last several years we've been building that platform and we're super excited. So Pete said, "Hey, I got to get on board in this thing because it seems like you guys cracked the code and this is what I see everyday as a problem in the market."

Corey: It's a great approach if for no other reason that it finally does what I think everything should be doing, is separating out the compute and the storage layers. With anything that's legacy you're playing within this space and you're running clusters of these things, oh, you need to store more data in it, add more nodes. You're adding compute when you maybe or maybe don't need that and your storage is going in lockstep with that. Conversely you need better performance, well add more nodes that add to the storage burden.

What I like about your entire approach is that there is no management overhead. The data lives in S3 and you don't have to touch it again once that index is created, it's just there available for querying.

Thomas: Absolutely and the separation was obviously extremely important where if you have one terabyte a day of data to a hundred terabytes, we just spin up additional compute and that nightmare of trying to create shards of a task storage compute is the nightmare that people deal with today, particularly say, with Elasticsearch and their clustering technology. And so the ability just to have infinite storage with dynamic compute, and so you can have one node represent the entire dataset, or a hundred nodes.

The ability for us to allocate to make indexing faster or query search faster is all on the fly and all dynamic. But this technology and this index does it at a price point that is very, very unique and very disruptive. Our costs to do indexing at scale is quite low and the ability to do on the fly compute allocation for your tech search, for your aggregations across terabytes of data, is extremely cost effective and you can see that in our pricing today.

Corey: One of the things that I find so interesting about this entire approach is that people generally already have something that's using Elasticsearch out there. I run the numbers myself, I didn't need you to tell them to me. It's one of these things of working from an economic point of view, nothing personal, I generally try not to take too much of what the vendors tell me on faith. I'd like to test these things out myself and it was right. It was knocking a majority of the spend off in virtually every case and that was incredible. Especially when you apply it to something like log analytics.

You take a look at any of the log analytics companies and people can think through a wide variety of log analytics companies, it doesn't really matter which one, and at some point at scale you have to begin kidnapping princesses for ransom in order to pay for the bill. So it's absolutely one of those challenging problems. And then when the bill gets too high, you talk to that vendor and their response is, "Oh, log less stuff," which sort of cuts against the entire premise of a log analytics company because you can start noticing relationships and data if you have the data. But if you don't, that door is shut for you. It really has seemed like a disjointed, fractured industry for an awfully long time. I'm still trying and failing to find examples of other approaches that solve this problem the way that you have. There's nothing else like it.

Thomas: You know, you're exactly right. And what's interesting to me is every single time we hear a customer wanting to increase their retention, they double their cost. So every time you double it, your costs doubles and your complexity doubles. And so the idea that we've created that the storage is infinitely scalable, S3 has been quite proven on that, and our ability to elastically scale up the compute and deliver the compute where the storage is for those query's, it seems so obvious, it's so natural. The problem is that the separating of storage compute is not the rocket science. It's the technology, this index that has allowed us to do it at a price point all pure on S3 or object stores in general.

Corey: We take a look at Reinvent last year and they announced enhancements to the Amazon Elasticsearch service that they run and I thought, "Okay, this is it. They're finally going to do the same thing," and they launched their badly named UltraWarm tier that had an accelerated performance approach, or I guess a lower cost of being able to stage old data out. Even they went in the exact wrong direction for this problem as I understand. It seems bizarre.

Thomas: So the funny thing is, they solved a problem like everyone else was solving. They solved the Lucene problem via band-aids, if you will, to be so bold. The system was never to do this type of scale, right? Elasticsearch was surprised that the log analytic community adopted this technology and UltraWarm is just another way to provide a caching layer to make it a little bit better, a little bit more cost effective. And so what we wanted to do is in essence make S3 ultra hot. And the idea is that with this technology, we don't have to play those bandaid caching games. It's pure access on S3 and it feels like it's a hot cluster, but it's on S3 with R compute.

And that's where our technology, our index technology has cracked that code. How to make S3 high-performance, make S3 a data lake ground that Amazon's pushing but not make it swampy. We provide actually data discovering and catalog what's in there, but to ultimately index the data to provide log analytics via the Elastic API and Cabana or SQL say through a Looker visualization for BI workloads or Athena workloads. And again, the key thing is to make it performant, make it extremely scalable without having to worry about charting as well as high performance with a very low cost.

I know that those are a lot of what ifs that we started out this company, but we cracked that code as I mentioned, and we're super excited about what we've built because we're seeing at our customers that we take their bill and literally cut it in half if not a third.

Corey: So one of the, I guess, constraints I have is that when you first learn about something, it's difficult to go back and relearn it as something else. I was introduced to CHAOSSEARCH as you effectively, whenever you're using Elasticsearch as a part of something, maybe an Elk Stack, maybe something else, the API is equivalent to a drop in replacement of CHAOSSEARCH. And that's how I've always conceptualized. Anytime I see a big Elasticsearch bill, this is one of the things that I tend to think about.

The challenge, of course, is that it sounds like you're going beyond being effectively just that. What do I misunderstand in that oversimplified description of what you folks do?

Thomas: I hate to say it but we are building the next generation database that really rethinks how stored analytics converge, fuse together. We created and distributed fabric and we have this ability to export compatible APIs. We don't run any Elasticsearch underneath the hood or Lucene index, but we support a open standard Elastic API that people know and love with the Cabana integration for all that great visualization that people do in log analytics. And the idea that you can do the Elastic API with say, an index pattern, and the same index is a table and say a Presto dialect SQL interface with Looker without having that cost and complexity that standing up a relational system like a warehousing or standing up elastic cluster that it's almost hard to believe because it hasn't been done before.

There are some companies on the fringes of trying to solve this. Snowflake has done some separation storage compute, though they still have the cash out on a physical disc. We are 100% pure object storage and the difference is it's your storage. You own the data. We are just the distributing compute that manages the index and the query execution.

So it's another way of delivering the idea that you dump data and then what our service does is we support what we call this refinery within our service. So you index it once, you index everything. And then with our refinery, you can create virtual transformation or views that look like index patterns in Elasticsearch or tables in say, SQL, all in that same representation on the fly.

So imagine if you had a hundred terabytes of data and get a physically ETL, typically folks who do EMR to take it out of S3, transform it and put into say a warehousing solution or Elasticsearch. What if you just brought up the wizard, created a view, created your transform and it's available immediately without having to do anything physical. This is the power of this index technology created. It is distributed, it supports text search and relational. It's uniquely compressed to save on costs, but allows for these virtual transformations late materialize that allows you to do all that variation that each department maybe in a company wants to read and analyze that data.

Corey: So putting the shoe on the other foot a bit. In what use case is someone going to be using Elasticsearch and for when CHAOSSEARCH is not going to be a fit?

Thomas: Yeah, so a classic case is I call the Elk use case where, let's say you have logging for denial of service attack and you want to know what's going on. So often people stand up, CDN logs in say, Elasticsearch or Elk, and this can get pretty big. One terabyte up to 10 terabytes a day, maybe even more of log data. And they typically see that there's a denial of service attack and they want to figure out what's going on. The problem with the Elasticsearch, it has more denial of service attacks come in, more logs come in, and now you're querying more often. This is classically what makes Elasticsearch sad where the cluster comes down because you've only provision so much capacity.

With our system, you dump your data into S3 and we index it, we allow for dynamic scale to do that exact same security ops type use case, denial of service attack, maybe CloudFlare logs, app logs. Constantly people come to us and they have really messy data for their app logs. They're dumping, it's an S3, and they want to know what's going on.

Again, CHAOSSEARCH is a great way to do a app log analytics. Hey, what does my application doing? Is it running slow, were there problems? So app logs, security ops, DevOps, those type of use cases are really a sweet spot because almost everyone we've talked to when we built out this idea, they say, "Do use S3?" And they almost would say, "Duh, of course we do." And do you use the ElasticSearch? "Of course we do." And they typically stored in S3 first and then store it to the Elastic cluster and we just say, "Keep it in the S3, we'll index it, and you can have that exact same Elk functionality that you've had without the cost and complexity."

Now what we're not good for, there was a time where we were staying away from real time functionality where you wanted a sub-second type performance for a short window of data. And we were going after the big, big data where customers that had 10 terabytes a day up to 100 terabytes a day of analytics that was just breaking the bank at that scale. Actually in Q-1 of this year, 2020 we're coming out with the real time. So just as you would put data into Elasticsearch for instance, you can put data to us, it's available real time. We'll write this data into your S3 account and as we come out with our SQL interface, we'll support, create updates and upcerts as well. So it's really turning into a full fledge data platform that can handle the real time and obviously that real scale.

So one of our limitations was in the real time because we were not focused on that. But as we go into more of a BI real time use case, we're adding that feature out. Other limitations I would say parody with the Elastic API, we're really focused on log analytics and coming out BI analytics. The last API is quite broad and we don't support all the classic low level texture type capabilities like fuzzy type searching. Not that we won't, but that's not really where our wheelhouse is. But we do support all the classic log analytic, text search, wild carding, etc.

The limitations, I don't know. We have some big ideas and some pretty powerful technology and architecture. So if there are limitations today they won't be limitations tomorrow.

Corey: One of my personal guilty pleasures is pointing out the terrible business practices of others, and I've been vocal about this for a little while where there's a giant slap fight between Elastic and pretty much anyone who is selling anything that looks like Elastic. It seems like their approach to open source by and large is, "Use our open source software. No wait, not like that." And so now there's a trademark lawsuit among other things. There's a slap fight that's going on between AWS and Elastic where it was also launch their open distro for Elasticsearch. Elastic was doing their whole only some of the code in the same repository and some of the same commits is free and open source, the rest comes with a commercial license. CHAOSSEARCH bypasses all of that, correct? Effectively it's sit on the sidelines and watch popcorn. There's no Elasticsearch under the hood here.

Thomas: Yeah, yeah, yeah. No, we're Switzerland in those battles. I mean, the open source community is so important to all of us and Elastic has done some great work and Amazon has done some great work. I know Amazon gets some bad raps about taking open source and making it a service and making a business and open source companies, once they have to start making money, may make some software proprietary. We were using the open source Cabana out of Elastic and to be frank, when Elastic started close sourcing, for lack of better term, or making a license different than than the Apache2 for some of their more advanced features out of Cabana, this past summer we adopted the Amazon's open distro for the alerting the timeline, the role based S control because it was open and it was keeping with that philosophy where a whole community was built based on the idea that this would be free and open.

Now I understand business and I understand that we have a business, we're creating a service to make money. But the idea to have open APIs, it's really something that it's harder to fight that. And I do believe that open API is the way to go. I understand why Elastic is doing what they're doing. I understand what Amazon is doing as a business. We're in that service of solving customer problems and you get paid for it as a service. So we're kind of in that mode where let's keep the API is open, let's make a business on offering a service or support like open source used to always do.

Corey: Well I wish you well, but I kind of think you're going to struggle until you learn to do what real companies do and threaten to sue your customers.

Thomas: Yeah, yeah. Like I said, I'll play Switzerland on that one. But I mean, listen, it's amazing what open source community has done and we leverage open source. The idea that you open source something and then in essence close in the future, that's a tough scenario for people who bet into this API that now everyone uses today. So how it all plays out, I'm not sure. There's a lot of money involved with that. And our viewpoint is if there's tooling and APIs that we can leverage from the community, we'll use it and bring those to our customers.

Corey: I love making fun of companies doing different things. That is my part and parcel and I've got to say, I'm sorry, you're not immune either. I have to ask the burning question that's been on my mind since I first heard of you folks and was corrected on this. Why is CHAOSSEARCH all in caps?

Thomas: It's not as well thought out as you might think. I had the idea of naming a company chaos-something. And originally, before we were CHAOSSEARCH we actually were called Chaos Sumo, wrestling the chaos. And as we were going after log analytics and we knew Sumo Logic was out there, we didn't want to be confused with them. So I thought, "CHAOSSEARCH because we're searching the chaos, is really where the value is," the searching analytics. So I said, "Ah, let's rename the company CHAOSSEARCH." Do we call it Chaos Space Search? Do we call it CHAOSSEARCH? Do we do all lowercase? Do we do all cap case? Really we did all these variants really quickly on a piece of paper and we looked at all three or four variants and we like, "Oh, the capital looks pretty good. Let's go with that." And it was nothing more than that.

And so what we've been doing to get around the all caps concept is making the first part of chaos bold. It seems to be working for us. But yeah, we get teased about that all the time and it was more just a looked good in the font we used and that's why we chose it.

Corey: And sometimes that's all it takes is going down that path. But it, of course, opens you up to all kinds of criticism from the peanut gallery, by which, of course, I mean me. I mean, you have search in the name. Search a little harder, you can find the caps lock key to turn it off. And the counter response, of course, is that, "Oh wow, there's a caps lock key, that makes it way easier to type the name of the company. Great. Just great. It's cruise control for cool."

I know that asking for what's coming next is always perilous because the best laid plans, et cetera, et cetera. What's next for you folks? What are you focusing on in this year of our Lord 2020?

Thomas: Our big vision is to build out a new type of data platform and as I mentioned, we went to market last year going after the big scale of log analytics. Support in the Elastic API and Cabana interface and solving those type of problems. But the vision is to have a true multi-model, real time database that the idea that we deliver on that data lake philosophy and you have database tooling that is natural, you're used to it. So once we come out with this multi-model capability, we're going to go after the Athena use case. We hear a lot of complaints about the costs and the scale and the idea that you have one data source or multiple data sources via the Elastic API or SQL API. 2020 is going to bring out really the first true, true multi-model functionality, our database that we're really excited about.

We had customers asking us all the time, "Can you support the Athena use case or the SQL Presto dialect?" And that's where we're going to do. We're going to first offer it to our law customers and then start growing the business into BI and ad hoc analytics all on S3.

We do have a plan for later this year to go multi-cloud. We've been pulled both from Microsoft and Google and which one we do next, I'll keep that a secret, but we will be coming out with a multi-cloud thing later in 2020.

Corey: One problem I see in the world of cloud billing historically has been that it's when you switch to a consumption based model, people don't know what something's going to cost. And sure they'd like to whine and cry and complain about it. But with a lot of systems with significant storage volume, you could run a query that costs tens of thousands of dollars without knowing it in advance. Driving down the overall cost, acquiring these things is, of course, incredibly valuable and helpful. But what are you seeing in the space as far as addressing that from a larger perspective? When you have an internal application that, run a query here and it costs $20,000, when someone hits submit, first, even attributing that back to that one query is a super hard problem. But the gold standard people are going for as a pop off of, "Hey, if you run that query, it's going to cost you a giant pile of money. Continue, yes or no?"

So there is that problem of doing the cost attribution of querying interestingly large datasets. Is that something that's on your roadmap at all? Is that not something you're seeing in your customers?

Thomas: So I was holding that one back. So yes, we actually have something that we're coming out with to make our system both be upfront storage type pricing as well as consumption based. And that's a very novel and unique in the logging space, it's not so much in more of the BI.

And a part of that offering we start supporting consumption based model is we actually know the cost of a query. And so we're going to have tooling in our user interface that you can say, "Oh this is going to cost me X amount of dollars over these hundreds of terabytes. Do I want to do it?" Or maybe Susie user can do that type of query, but maybe Bobby only can do short time query's for this amount of cost and have a whole billing construct within our user experience to keep those costs down.

Now to your point earlier, we've cracked the code on reducing that cost dramatically low. But imagine if your costs of indexing was virtually free and it's all based on queries, but your queries are cost effective as well as intelligent. And I'm coming out and I'm saying we are coming out with a feature that will be consumer based and the user will know and control how much data and what's going to cost.

Corey: To be clear, when you say that it predicts the cost and tells you what it runs in advance, is that the cost for CHAOSSEARCH? Is that the infrastructure cost underlying what's going on? Or is it both?

Thomas: So clearly we'll have a margin within-

Corey: Well, of course, I'm not suggesting you should.

Thomas: Yeah, it'll be cost of the query. So we want to not only be disruptive in the classic log pricing model where everything from $100 per gigabyte and up, where we're currently $10 to $15 per gigabyte, which is very disruptive in this market. We're going to be instead of $5 per query per terabyte or $1, I'm not going to say what we're going to be, we're going to be dramatically cheaper than that and the idea that, "Oh, these 10 terabytes of one query is going to cost me X," you can control that. You can set policies so that makes sure that you only use what you want to pay for.

And not necessarily a credit based because that can get complicated. I know other vendors do a credit base that you put up front a cost and then you eat into that. That's hard to deal with. This is going to be a lot more driven by your controls and and what you want to do. So it's going to be significant, it's going to be disruptive and hopefully the customer's are going to like it. And they kind of a la cart, maybe you want to choose full upfront, by storage, or maybe uses base as the way you want to go.

Corey: Having a variety of different options is always a good direction to go in as far as meeting customer requirements. Everyone has a different use case and everyone wants to express that in different ways. There few things more frustrating than when a vendor's pricing model doesn't align with how you are intending to use the service.

Thomas: Yeah, I mean, here's a good example. We have people that come to us, say, "I have a hundred terabytes. I only have to query maybe once a week. It doesn't make any sense for me to stand up a huge system to do that because consumption base makes a lot of sense," right? Or when there's a denial of a service attack, then you want to really hunt and figure out what's going on. But the rest of the time, the system's pretty idle.

Now if you're doing some more real time where you're doing dashboarding or it's built into your application, okay, maybe the consumption base is not the right pricing, but when you're doing a lot of ad hoc or an investigation or you need it when you need it, but you don't need it when you don't, there's really no good solution out there in the market, particularly in log analytics.

Corey: So if people want to learn more about what you folks are up to, continue to follow your exploits, get annoyed at your unnecessary capitalization, where can they do that?

Thomas: Come look at the CHAOSSEARCH.io. We have a whole bunch of material that talks about the platform. We're actually updating our content later this month on a whole bunch of detail use cases and an ebook coming out. So come to our website, ask for a free trial. It's fully automated, you're up and running within five minutes on your S3. You can also set up a larger POC where we allow you to test out 250 gigabytes of data, which is actually pretty big for the free trial. And if you're a really big account we have that we call big POC where you can call us up and if you're looking to test out terabytes of data per day, we can work with that with you.

We have a lot of good blogs and a lot of good documentation, but sometimes just kicking the tires is the best way to learn and our free trial is probably the quickest way to learn what we do.

When you first log in, it's quite unique. We're the first company that starts with your storage and not just the idea that you dump data into them and then you start playing with the product. The product starts when you first log into your storage. We have a lens into your S3, we have a refinery to create different viewpoints. And then we have Cabana, your favorite visualization tool and Elk to do your analytics. And we've automated the process from raw data to insights really best of breed.

Corey: Well, thank you so much for taking the time to speak with me today. I appreciate it.

Thomas: Thank you.

Corey: Thomas Hazel, founder and CTO of CHAOSSEARCH. To learn more, visit CHAOSSEARCH.io. I am Corey Quinn, this is Screaming in the cloud.

If you've enjoyed this podcast, please leave it a five star review in Apple podcasts. If you've hated this podcast, please leave it a five star review in Apple podcasts and tell me what my problem is.

Announcer: This has been this week's episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Dai Wakabayashi
Dai Wakabayashi is a tech reporter for the New York Times based in San Francisco whose primary focus is all things Google. Prior to joining the Times, Dai covered Apple and Japanese tech companies (e.g., Sony, Nintendo, Panasonic, and Sharp) for The Wall Street Journal for almost eight years. He also worked for Reuters for nine years, focusing on Microsoft during his time there.

Links Referenced:

  • Dai’s recent AWS article
  • NY Times hires Wakabayashi to cover tech
  • Twitter: @daiwaka
  • LinkedIn: https://www.linkedin.com/in/dwakabayashi/
  • Personal site: https://www.nytimes.com/by/daisuke-wakabayashi
  • Company site: nytimes.com

Transcript
Announcer: Hello. Welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on this state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored by Influx Data. Influx is most well known for Influx DB, which is a time series database that you use if you need a time series database. Think Amazon time stream except actually available for sale and has paying customers. To check out what they're doing both with their SaaS offering as well as their on premise offerings that you can use yourself because they're open source, visit influxdata.com. My thanks to them for sponsoring this ridiculous podcast.

Welcome to Screaming in the Cloud. I'm Corey Quinn. Last December, an article came out in the New York Times titled Prime Leverage: How Amazon Wields Power in the Technology World. Joining me today is Dai Wakabayashi, the journalist who wrote that article. Dai, welcome to the show.

Dai: Thanks for having me.

Corey: First, thanks for taking the time to speak with me. You are a journalist's journalist, not a journalist in the sense of it's a polite term that we're going to appropriate for someone who works in corporate comms. You are a reporter. That is what you've done for your entire career as best I can tell.

Dai: Yeah, that's definitely ... ever since I left college, that's all I've been doing. It's a job that doesn't particularly pay well. It's not one that gets a ton of respect in the broader world these days. But I do think it's an awesome job in the sense that I get to be on the front lines of interesting things happening and gets taught to a lot of interesting people. It's one of the few jobs where you can stick your finger in someone's eye and you're allowed to do it and not only are you allowed to do it, you're sometimes encouraged to do it.

Corey: It's refreshing to talk to someone who isn't trying to sell anything, which is I guess a depressing commentary on the state of the world today. So this article came out, I want to say December 15th. In fact, yes, it was December 15th. It's a lengthy article. The general thesis is that in the technology world, AWS more or less winds up having a product strategy that distills down to yes and effectively strip mines open source projects and companies for a lot of their innovations and then rolls it into first-party services. Is that an effective summary? Have I missed some of the salient points?

Dai: Yeah, I think so. I mean I think part of it is that what we saw as one of the main ideas here is that Amazon has created this incredible platform, which is now essentially taking over the way people buy and spend and use technology. Now instead of just being a platform and being the thing that everyone just builds on top of, they're also now offering everything else that runs on top of it. So all these companies and many software companies we spoke to who felt like maybe five, six years ago thought, here's an opportunity for us, a new frontier for us, is now increasingly finding that, well, this frontier may not be as lucrative because we're competing with the very company that is the platform. So I think that notion of the responsibilities of a platform is something that generally we have been kind of looking into at the Times about big tech companies, whether it's Google or whether it's Facebook or whether it's Apple. Certainly this is our look at what Amazon is doing in the cloud.

Corey: What's interesting to me is seeing the community response to this in various corners. There's of course the immediate knee jerk response of, well, that isn't accurate. There was an agenda, yada, yada. It comes off as the people immediately saying, "Well that article is great, but there's 40 years of nuance go into the whole open source world that was largely passed by in the article itself." My position on that is the nuances of the open source world are something I don't fully understand despite having been actively involved with it for roughly 15 years. The fact that some of those nuances and edge cases didn't make it into a front page story of the New York Times isn't the most surprising thing in the world to me. What is your perception of the feedback then since this article was published?

Dai: Yeah. I knew that that was coming. Look, I don't think we ever thought we could fully capture all the nuance of something like open source. I mean, it is akin to like a religious war on some level, that you will never fully capture all the history and the back and forth of the industry. But we're trying to give a flavor of it and we're trying to give enough of a flavor of it that people can get a base level understanding of what's at stake and then read the article for what it's worth. I do understand that people feel like there was some nuance missing. But I think one of the things, one of the challenges with journalism, is always writing something very nice for a very general audience can be the challenge.

It's how far back do we pull back the lens? Do we pull it back so far that it becomes almost impossible to distill what specific thing you're talking about, or do we go too narrow and sort of alienate a broad part of the audience? Look, there's never a right answer to exactly how far back to pull back the lens. We tried what we thought was the right level of altitude from which to write about this.

Corey: That's one of the more interesting parts of I guess seeing articles on things that I tend to be relatively well traveled in. I see certain shortfalls are papering over of complex issues. I guess as I went through the maturation process in so far as I ever did, it really was an eye opener for me as far as realizing that, okay, if this is an area in which I have subject matter expertise, well, there are shades of nuance and journalism aimed at the mainstream doesn't necessarily pick up on all of that nuance. The next logical step for any was I wonder with all these other things in which I don't have subject matter expertise, is this going on there as well? The answer of course is obviously. Nothing is going to act as a in-depth primer to a field of study. Reading a newspaper does not make you a subject matter expert on anything other than reading that day's issue.

Dai: I mean, I think one of the things that's really eye opening about the process of writing a fairly long story like this is the collaborative nature of it. I speak to a ton of people, right? 40, 50, 60 people for this. I get that deep perspective. When I go to write it, I feel like, well this is very important. I need to explain the nuance of the licensing deal and why this licensing agreement is different than that licensing agreement. Then my editor looks at it and says ... I mean, she looks at it and says, "If we're talking about licensing, we've lost. Like the specific different licenses. We've lost people." That's the collaborative effort, right? I go and I get really steeped in it and I get really granular, but as I write it, I think, okay, I need to pull this back a little bit. Usually the editor pulls it back even more. We have a fantastic editor. Her name is Pui-Wing Tam.

Her job almost always is to make us be ... We're looking at something at 5,000 feet. We need to be looking at it at 10,000, 15,000, 30,000 feet. Pull it all the way back and tell us what it is about the industry. For us, this story was always about power and influence. Here's a company with enormous power and enormous influence now in the technology world and how do they wield that power and influence? She always kept us on that North Star, for lack of a better phrase. But that's part of the collaborative effort. So that's always really interesting to me where I feel like sometimes I'm in there saying, "Well, this is losing too much nuance." And she says, "Well, what is the real nuance here that you're trying to make?" Then we'll rephrase it. It's a back and forth process.

Corey: Well, what's interesting to me as ... The response that I heard from this article from my friends at AWS was it mostly came back to the position of that's not accurate at all. My argument in response to that is I agree with what they're getting at in that I do not personally believe that there's someone at AWS, or most people at AWS for that matter, are sitting around figuring out, okay, how can we completely undercut various people that do business with and rely on us? I don't believe that is the starting point that anything reasonable or rational comes from.

But when you're building that out, you don't have the luxury of understanding how actions get interpreted in the broader marketplace. At this point, Amazon is give or take a trillion dollar company. They deserve a increased level of scrutiny as a result. While that isn't how we intended this to come across, well then you should've done a better job of telling a story around it because people don't see the world the way that you see it internally at your company. The actions you take reverberate throughout all of society at your scale. It's important that people are aware of the outweighed impact their words carry.

I was having this conversation recently when I turned to a peer and say, "Hey, I wonder what this line on the graph means." That's just an idle question. If I say that as someone's manager, it can very easily be interpreted as you should find out what that line on the graph means. The fact that there is that power disparity completely changes the context of the exact same statement.

Dai: Totally. One of the things that I thought was interesting, I mean Amazon clearly, they put out a blog post. If your listeners haven't gone and read it, they should.

Corey: We'll put a link to that in the show notes.

Dai: Yeah. In response to the article. One of the things companies are very good at is if they find a story that they especially don't like and they find an error in it, they will gently prod you to correct that error in the sort. There were no requests for corrections. Granted they did not like the way I interpreted a certain set of facts. That I understand. They might think that we had a certain agenda. That I understand, but the notion of accuracy is something that we take very seriously, obviously. When a story goes out like that, the thing that I'm sitting there also very worried about is, well I wonder if people are going to find a mistake in the story. We did not get a request for a correction.

The one thing we did get, which I thought was interesting, and I guess I should make the disclosure now, I didn't even think to check to see if the New York Times is an AWS customer. It turns out, yes. Yes, we are an AWS customer along with the GCP. That disclosure was not in the story. In hindsight maybe we should've had it in there, but it was the farthest thing from my mind to even think about that, which probably is a shortcoming.

Corey: Well, the other side of it too. Whenever the Washington Post has any article that even touches on Amazon, they have to disclose, "Oh by the way, the founder of Amazon owns the entire paper." That doesn't stop them from being overwhelmingly critical from time to time, but those disclaimers in there from a journalistic integrity point of view are incredibly important. I would argue that I don't think anyone thinks that you took it easy on AWS because the New York Times is an AWS customer based upon the response I saw to that article.

Dai: Yeah. I don't think that's the concern. I do think that disclosures are incredibly important. Look, a lot of people talked about the reaction then and everyone ... Part of the privilege of working for the New York Times and part of the responsibility is that we have sort of a bullseye on our backs and that when we write something, a lot of people read it. We have a huge platform, and that's really a privilege. But with that comes an incredible responsibility that we have to be able to weather the criticism of that story.

For us, I was not surprised that a lot of people had criticism for the story. I think that's totally fine. Amazon's response, I think it's totally within their right to do so. I don't take it personally. I don't think Amazon took the story personally either. They realize, I think at some level, that as they get bigger, people are going to start looking at them a little closely. Certainly this year we've seen a ton of great work, not only from the Times, but also from the Wall Street Journal and other publications. I think Buzzfeed also had a great story recently that looks at the reverberations of the Amazon world that we live in that that goes beyond AWS and obviously on eCommerce and this notion that everything should be delivered to you the second you want it. We live in an Amazon world, and so I think it's only fair that reporters get more and more interested in looking very closely at that world. Amazon is totally within their right to be very aggressive in responding back.

Corey: I find their entire blog post that responded to your article to be a little on the interesting side. I mean they of course have the trotting out the hostages of all of their large partners who are active in the open source work. "These people are great examples who love working with us." Well, yeah. I mean, what else are those companies going to say? You can't ever speak out against one of your largest partners and expect to live. But I guess pushing back on this of, oh, companies aren't actually afraid that Amazon is going to move into their space. No one has actually said that statement aloud because it's provably untrue. Whenever I talk to people even on this podcast and occasionally in the pre show discussion I'll have, "Oh, are you worried about AWS? I'm releasing a product next week to put you out of business." And their response is, "Don't even joke about that."

People are incredibly sensitive to that. There's a palpable sense of tension in partner meetings when people are preparing for large conferences at Reinvent, for example, when they have their big annual event and their giant keynote where Andy Jassy more or less gets on stage for three straight hours and recites new services. It's just what it feels like sometimes. There's a palpable sense of tension of is this going to be the thing that more or less fundamentally drives our business in an unexpected and unwelcomed direction?

Dai: I mean I've covered technology now for almost 20 years. I've covered Microsoft in the 2000s. I've covered Apple. Now I'm covering Google as my main sort of beat. I've never seen companies more scared to talk out about another company than I saw with Amazon and AWS. It was really fascinating. There were a bunch of companies who off the record or on background would really just tee off on the company on Amazon and AWS's practices. But if you tried to push any of them to talk either on the record or even the on the record things you ask them to make that ... Or when they were willing to talk on the record, it was a total whitewashing of the things that they would say privately.

That's what I thought was really eye opening to me is that a lot of companies are just deathly afraid of even saying the most innocuous thing about Amazon. The trotting out of the partners I found to be less useful. Obviously Amazon tried to get us to talk to the partners who were happy. Certainly I don't doubt that there are some partners who are genuinely happy about working with AWS. But I also doubt that even the happiest partners would voice their concerns to the New York Times in a frank way. That was one thing I thought was sort of interesting.

The other interesting thing that I found about their response was the way a large story like this works, Amazon is not blindsided by that story. We worked with them for a week before the story ran going over every detail of the story. They voiced their concerns, some of which we put in the story. Their rebuttals and such. But companies also play games. There is a lot of, well, you can talk to our executive here, but it'll be on background and you can't use it. You can use the things we talked about, but you can't attribute it to that person. Then after the fact, they have this very long and exhaustive blog post. The notion that somehow that we didn't give them a chance to really talk and rebut is sort of misleading I think. They have ample opportunity to respond to the story, and their response was in the story. This notion that somehow that they were caught off guard, if that was the intent, that's I think misleading.

Corey: I think that it's easy to fall into the trap for a lot of these companies as well. I mean certainly I was in this when I was starting my business and it's, well what happens if I upset Amazon? I have nothing even remotely resembling a survival instinct. So I started a newsletter called Last Week in AWS giving them even trademark grounds if they wanted to come after me on that perspective. Then more or less made fun of them every week in an email newsletter. I expected that it was going to basically make me persona non grata as far as anything AWS oriented. What I learned from that, which was shocking was that first they love that. They have a bit of a corporate beat the crap out of me fascination, so okay, we'll roll with that.

But also there is no one person who is I guess responsible for the viewpoint of an entire company at that scale. I'm used to small business where a big company has 200 people. There's no Ted Amazon sitting there deciding whether they like someone or hate them. It's a bunch of very small teams that are largely independent more so than at most other companies. Some of those folks love what I'm doing. Some of them despise everything about me. It's one of those things where I've learned that you can't make people happy all the time. As long as I continue to have people take my calls, I assume I haven't gotten it entirely wrong.

Dai: No, I think you walk that very fine line. I think you are incredibly fair. I'm a big fan of the newsletter. I think you have a very ... I'm also a big fan of your Twitter account. You have a very good tone to it. I think the reason why, I imagine the reason why, people don't dismiss you is because you come from an incredibly informed place. You know what you're talking about and it's obvious. So that makes it a harder to dismiss you as some kind of kook or someone with an agenda because I think it's obvious that you've done the homework. Now I think when someone like us stops in and writes something critical, it's easier to dismiss us. Especially when we compare cloud computing to a coffee shop. Sorry, open source to a coffee shop. I think it gives them more fuel for the fire.

But I don't think they take it personally. I mean I think that this is part of the back and forth and they'll become better for it. I think their message will become sharper and they'll think about, hopefully, how their partners will respond to certain decisions they make about products hopefully. I think that that's all part of learning the responsibilities of being this major tech platform. They're not up and coming anymore. I mean I think it was very clear [inaudible 00:22:05] that they are now the big dog. I think this year more than other reinvents that I watched on YouTube, it was clear to me that they saw sort of themselves as the big dog in a way that I don't think they did in the past.

Corey: I still think internally they struggle with that. They tend to pride themselves in their two pizza teams. So you have a team of 20 people or so, give or take, that are launching an entire service. I'm not sure that they realize that this is something that millions upon millions of people are going to see. I am big enough to admit to more than a little professional jealousy of you at this point. I have never had a same day blog post come out from AWS, penned by a VP to the theme of screw this guy over something that I've written. We all have bucket list items. That is definitely one of mine.

Dai: As a reporter, nothing makes me more uncomfortable than when someone that you wrote a story about or a company that you wrote a story about comes back to you and tells you, "Oh, we love that story." That probably means I wasn't critical enough, that I wasn't analytical enough. So getting a response like that, well look, it's not the thing I was hoping for. It also doesn't bother me at all. It probably means that I was sufficiently critical or sufficiently analytical.

Corey: I remember that you and I had spoken a couple of times when you were putting this article together, but that was the last that we'd really communicated. The next thing I knew is it's Sunday morning and I wake up leisurely. My toddler let me sleep until 7:30 in the morning. I wind up looking at my phone and it's exploded. I have messages from a whole bunch of people of, "Great quote in the New York Times." Wait, that was sarcastic? Oh, you're a relatively high level executive at Amazon. What the hell did I get quoted on? My only quote in the article was after the observation re-invent was called AWS red wedding. My comment was, "Nobody knows who's going to get killed next." First, of all the things I've ever said, that's the objectionable thing that you think is the problematic piece. Oh my God, don't ever look at my Twitter account.

But more than that, it's fascinating just from my perspective seeing the sheer reach that this has. The quote was near the end of the article. It refused to call me a cloud economist because the New York Times style guide and me, my God. The kerning was wrong on the AWS. The period after the A and the W and the S as well. It's just where you have gone, I cannot possibly understand it. But I still had hundreds of newsletters signups that day on a Sunday, which is okay. Lesson learned. If I ever want to pick some good place to do some advertising, the front page of the New York Times is not the worst in the world.

Dai: Initially, we went back and forth a little bit if you recall about your title. In one of the drafts of the story, I had you quoted as cloud economist. I believe my editor wrote, "Really?" I said, "He insists that's his title." Then I said you sent me some links about how there was a PhD student who ...

Corey: Oh, full PhD. His name is Owen Rogers, at 451 research. He has a PhD in cloud economics. There are entire cloud economics divisions at companies including AWS. It's a real thing. I made it up when I first started my consulting company, and then I look around and I realized, wait, other people have used this term too. When I met Owen Rogers with his PhD in cloud economics, he was very excited to meet someone else who is in the space and wanted to talk about it. I had two paths. I could either be honest and tell him I made it up or I could bluff and probably wind up with a book deal out of it. I picked the more noble path. Part of me always wonders if I'd gone down the other direction how that would've played out.

Dai: Right. Basically, it started with ... I did gently push back and I said, "Look, this is what he says his title is." But basically that was a line in the sand for my editor and she was just like, "No, we're not giving someone the title of cloud economist." I think she wrote around it basically.

Corey: Society is always one of those challenging things to wind up evolving. One of these days I'm sure we'll get there. Now I would sooner fix the period after A W and S. That is the hill I would choose to die on long before my ridiculous nonsense title.

Dai: I am not a fan of the periods either. I had to go back and write each, fix every AWS reference to include those periods. So it was laborious and time consuming. But like even when we do CEO. That's C period E period O period. For the lack, I mean that's just our style guide. At some point you you just can't fight so many battles. You've got to pick and choose and that one is not a hill I'm willing to die on.

Corey: No, and it shouldn't be necessarily. I've done a lot of interesting writing, some of which has circled the internet three times, it feels like. But this was a whole nother level of attention and people being angry about it. If you go on Twitter, professional advice, never go on Twitter, and take a look at any given day. People are raging about the New York Times. I mean I'm a subscriber, I have been for years. The argument is always this is a terrible egregious breach of public trust, journalism ethics, et cetera. I'm canceling my subscription. Well isn't that the third time this month you've canceled it? If a paper only prints things that you like and agree with, that's propaganda. I don't think that serves us well. I think that we almost have a subversion of journalism by, to be blunt, corporate comms people in some respects, but there are also larger societal challenges around that too.

Dai: I mean, I think it is a very worrying time for journalism. For us covering, for example, the tech sector, which is now where the most powerful and the richest companies and the most influential companies of the world exist, it's not crazy for us to have one beat reporter. Like I have one beat reporter. I'm the one beat reporter covering Google. There are hundreds of communications people at Google, right? It's impossible to match them in body count. It's important for us to keep these powerful people to account. There's ways to do it that's fair. A lot of these companies do incredible things. I don't discount that at all. Google has made the internet usable. AWS has launched hundreds of startups because it's a lot of whole new way of companies to buy and use technology. I don't discount the importance of that, but that also doesn't give them the right, I think not them as in AWS, but companies in general to do whatever they want.

I think that it's important for us as reporters and journalists to look at these companies with a critical eye, especially when they are so influential to society at large. The things that they do make a huge difference in the day to day lives of people who are really far removed from the seat of power. We can do a service for a lot of readers and a lot of people who rely on these companies for many, many aspects of their lives.

Corey: I think that that is one of the most impactful things that you could say at this point. You're right. Holding power to account is incredibly valuable. It is the entire purpose behind a free press. I think that that is something that we protect at almost any cost. That said, I do have sympathy for some of the corporate comms folks and the rest. I mean, on the scale of Amazon tier companies, I'm a nobody and I'm perfectly okay with that because I can shoot my mouth off about whatever I want. Worst case, I have to apologize. If you wind up saying something as a representative of Amazon that gets misconstrued and it moves the stock price 1%, that's about $10 billion. That is orders of magnitude more than it costs to have me killed. The higher you rise, the less you can say off the cuff and the more everything has to be scripted and rehearsed. They have that challenge. I empathize with them. Truly I do, but not at the expense of the larger good of society.

Dai: I mean, I think the thing that really bugs us a lot is for ages we had a really hard time getting comms people to put their names on statements. A lot of times companies will have statements that are very innocuous, just we believe in helping customers live a better life. Then when you go to the comms people and you say, "Oh, can we put your name on that?" Then it's like, "Oh, Well." It's like, well part of your job is to deal with the press. If part of that is to put yourself out there ... I don't know. I just sometimes feel like there's this whole group of comms people who see their job not necessarily as informing the process, but more in managing the press and wrangling the press. I don't know. I guess I have less sympathy for corporate comms than you. But I recognize that I'm-

Corey: It is the absolute opposite of a job that's appropriate for me. I mean, their entire role is to say no comment, and I have a comment for everything.

Dai: Yes. That's why I as a reporter enjoy talking to you.

Corey: If nothing else, I certainly stonewall. Do you have anything to say on this? Well, of course I do. I know nothing about it, but I'm thrilled to shoot my mouth off anyway. If people want to learn more about what you have to say, follow your exploits, where can they find you?

Dai: I have a Twitter account @DaiWaka. D-A-I-W-A-K-A. Obviously I publish in the New York Times, but that's probably the two key ways.

Corey: Excellent. I will put links to those things in the show notes.

Dai: Here's what I will say. I appreciate the feedback. If anyone wants to get in touch with me, wants to talk to me about the story and feels like I misunderstood something or just kind of wants to help me be better, which I can always be and I'm always open to, they can email me. That's Dai.Wakabayashi@nytimes.com.

Corey: You are a braver person than I. You have enough reach saying, "Oh yeah, anyone who wants to can wind up getting back to me." I can't imagine what that's like. I mean I wind up telling people they can hit reply to my stupid newsletter and it hits my inbox. The reason I can do that is because no one ever does. On a busy week, I'll get a dozen email responses. At your scale, I have to imagine that it'd be thousands.

Dai: No, but I get my share of enthusiastic people. That's a euphemism. I think it's great. I love talking to readers and I love responding to people. I respond to people on Twitter sometimes. That can be a dicey game, but often it usually ends up in a positive experience. I always love the feedback and even if it's negative, I'm happy to have the conversation.

Corey: Excellent. Thank you so much once again for your time.

Dai: My pleasure, Corey. Thanks.

Corey: Dai Wakabayashi, New York Times reporter focusing on technology. I'm Corey Quinn. This is Screaming in the Cloud. If you've enjoyed this podcast, please leave it an excellent rating on Apple podcasts. If you've hated this podcast, please leave an even better rating on Apple podcasts.

Announcer: This has been this week's episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Leon Adato
Leon Adato is a Head Geek and technical evangelist at SolarWinds®, and is a Cisco® Certified Network Associate (CCNA), MCSE and SolarWinds Certified Professional. His experience spans financial, healthcare, food and beverage, and other industries.Before he was a SolarWinds Head Geek, Adato was a SolarWinds user for over a decade. His expertise in IT began in 1989 and has led him through roles in classroom training, desktop support, server support, and software distribution.Funny:In my sordid career, I have been an actor, bug exterminator and wild-animal remover (nothing crazy like pumas or wildebeasts. Just skunks and raccoons.), electrician, carpenter, stage-combat instructor, American Sign Language interpreter, and Sunday school teacher.Oh, and I work with computers.Since 1989 (when you got a free copy of Windows 286 on twelve 5¼” floppies when you bought a copy of Excel 1.0) I have worked as a classroom instructor, courseware designer, desktop support tech, server support engineer, and software distribution expert.Then about 16 years ago I got involved with systems monitoring. I've worked with a wide range of tools: Tivoli, Nagios, Patrol, ZenOss, OpenView, SiteScope, and of course SolarWinds. I've designed solutions for companies that were extremely modest (~10 systems) to those that were mind-bogglingly large (250,000 systems in 5,000 locations). During that time, I've had to chance to learn about monitoring all types of systems – routers, switches, load-balancers, and SAN fabric as well as windows, linux, and unix servers running on physical and virtual platforms.Full LengthLeon Adato is a Head Geek and technical evangelist at SolarWinds®, and is a Cisco® Certified Network Associate (CCNA), MCSE and SolarWinds Certified Professional (he was once a customer, after all). His 27years of network management experience spans financial, healthcare, food and beverage, and other industries.Before he was a SolarWinds Head Geek, Adato was a SolarWinds user for over a decade. His expertise in IT began in 1989 and has led him through roles as a classroom instructor, courseware designer, desktop support tech, server support engineer, and software distribution expert.In the early 2000s, Adato got involved with systems monitoring and has since worked with a wide range of tools including Tivoli®, Nagios®, Patrol, ZenOss®, OpenView, SiteScope, and of course SolarWinds. He has designed solutions for companies that were extremely modest (approx. 10 systems) to those that were mind-bogglingly large (250,000 systems in 5,000 locations), through which he gained experience monitoring all types of systems – routers, switches, load-balancers, and SAN fabric – as well as Windows®, Linux®, and UNIX® servers running on physical and virtual platforms.His career includes key roles at Rockwell Automation®, Nestle, PNC, and CardinalHealth providing server standardization, support, and network management and monitoring.

Links Referenced

  • Twitter: @leonadato
  • LinkedIn: https://www.linkedin.com/in/adatole/
  • Personal site: www.adatosystems.com
  • Company site: www.solarwinds.com

Transcript
Announcer
: Hello and welcome to Screaming in the Cloud with your host, cloud economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored by InfluxData. Influx is most well known for InfluxDB, which is a time series database that you use if you need a time series database. Think Amazon Timestream except actually available for sale and has paying customers. To check out what they're doing both with their SaaS offering as well as their on-premise offerings that you can use yourself because they're open source, visit influxdata.com. My thanks to them for sponsoring this ridiculous podcast. I'm Corey Quinn. I'm joined this week by Leon Adato. Leon, welcome to the show.

Leon: Hello there.

Corey: So you're a lot of things. You're a head geek at SolarWinds. In some circles, that would be considered an insult. In your case, it's a job title.

Leon: Absolutely. I took the job almost sight unseen when I found out that was the title. I'm really happy about having it.

Corey: We all gravitate for different things. You're also the cohost of the Technically Religious podcast.

Leon: Indeed.

Corey: You do identify as an Orthodox Jew.

Leon: Also indeed.

Corey: Excellent. I myself am such a good Jew I don't need to practice, which is neither here nor there. But we met a long time ago, I want to say, at one of the Devops Days events and got to talking and for some reason, never really stopped talking about a variety of different things despite the fact that you tend to operate in a world that I don't spend a whole lot of time in, namely on-premises and in data centers.

Leon: Yeah, I mean, the fact is is that it's not just on-prem, but it is a style of IT work that I think is particularly unsexy or seen as being unsexy. But the fact is is that plumbing is... Anyone who's ever done plumbing, and I do a lot of home repair, and that's my Achilles heel, plumbing is not sexy. There is simply no way to... Despite what pornography movies might have you believe, plumbing is not fun nor sexy, and yet routers and switches and the data must flow. And that's an area that I tend to focus on, and I tend to work with people to improve.

Corey: I think that the interesting part of this too is that we've set up a false dichotomy where, oh, you're either in data centers with the ancient buffalo, or you're entirely cloud native. In practice, almost no one is. I mean, as much as I want to start casting stones like that, I look at my NAS in the corner where I do local backups and things and realize, "Huh, okay, maybe there's some unfortunate truth I don't necessarily want to acknowledge here."

Leon: That's absolutely true. And it's more true for most of your medium and especially the large businesses that simply can't get a catapult and fling everything into the stratosphere as it were. First of all, as I mentioned before, the data must flow, that at a certain point you have people sitting at desks with compute devices, and those bits have to get there somewhere. So, you still have an investment of your networking gear if nothing else. And with the networking gear comes funny little things like firewalls. And you probably have a few proxy servers and some load balancers, and all of a sudden, you have a nice little stable of on-prem equipment that still has to be cared for.

And then even when you're talking about stuff in the cloud, again, the techniques still matter. You still have to know how to subnet. You still have to know how to set up a VPN. You still have to know how servers and applications and services and processes and events work regardless of whether they are in the cloud or, as my fellow Head Geek Thomas LaRock likes to say, earthed. Regardless, the same techniques work.

Sure, many applications are transitioning to full cloud stack. They're ephemeral, and they use containers, and they are microservices and all this. Yeah, many things do that. Many things are also simply lifted and shifted into an AWS instance, and they just sit there. And they still have to be managed the way we always manage them. It's just they're not there.

But here's the funny part, in the last... So, I've been in IT for 30 years. In the last 20 of those 30 years, I have never been permitted, not just that I haven't had a habit of it, but I've never been permitted into the room where the machinery that runs the stuff that I do was kept. It was remote. Remote in Amazon's data center was as remote, or perhaps in some ways, closer to me, than the data center down the hall. So I tend to work in that space of those same, I'll say it, old, techniques, processes, standards, systems, et cetera.

Corey: I would like to point out that those are your words, not mine.

Leon: Yes. Fair enough.

Corey: So, I periodically encounter you in your natural habitat, which is of course at a booth at an expo hall at a conference purporting to talk about technology, the conference, not you. The you, in turn, talk about what SolarWinds does, and this is not a sponsored episode. We're not going to go into the realm of everything your company does, but that's what I want to mention though is it seems like you folks do kind of a lot. It almost feels like a choose your own adventure story where no matter what you're into, there's something that you folks have to demonstrate that aligns with who the customer is.

And originally, this was, okay, great. So, I see you at some of the, shall we say, legacy vendor shows, fine, whatever. But then I see you at the cloud native shows, and you've still got interesting things to talk about. What did your, I guess, progression look like? You have been doing this in data centers far longer than I ever did. But now, you've migrated to be able to have reasonably intelligent-sounding conversations about things that touch on cloud. What did that transition look like for you?

Leon: Well, I think full disclosure, my degree's in theater, so I can act like I know what the hell I'm talking about most of the time and pull it off fairly well. But as I said, I've been in IT for 30 years, so hopefully, I've earned some credibility along the way. And I actually worked my way up the IT food chain. I started off in classroom training, teaching people how to use MultiMate and WordStar and continuing for about five and a half, six years as technology progressed from there and then desktop support to server support to network support.

And then I dovetailed into this really cool bizarre corner of the IT universe of monitoring and using a variety of the usual suspects, big blue and big red and all the other colors of the rainbow, I guess, doing monitoring. Basically, my shit would watch your shit and wait for it to break and tell the right person. That mindset has helped me transition because I really do believe that monitoring engineer... Okay, go ahead and laugh. Monitoring engineer is a thing in the same way that 10 years ago, infosec professional was not a thing. You just had people who really, really enjoyed ACLs or really loved to dumpster dive through log files on systems. And now, you have infosec professionals and red team, blue team, purple team, et cetera.

Right now and for the last, say, five or six years, people have been focusing on monitoring as a discipline, understanding the failure modes, not just, oh, I monitor storage, or I know how to monitor virtual, or I know... No, no, no. It's how things fail overall and how they all fit together. And when you're focusing on that in the way that I have for, like I said, about 20 years, you start to care less about the platform, "I'm a Windows admin." "No, Linux forever," whatever. And you start to focus on this is an interesting system, and this is the way it interacts with other systems. And this is, again, how it falls down or how it gets wobbly, and here's how we can catch it at this point.

So that combined with the fact that as Ecclesiastes says, "There's nothing new under the sun," you say containers, and I think LPARs running on AIX, which are literally exactly the same thing. You can talk about virtual storage, and I think DASD or shared drives or even timeshare on a mainframe. That's a little bit beyond my time. But the point is is that a lot of the things that are being used new are really reinventions of old techniques or technologies. In fact, I was just reading this morning about how Sun was trying to invent basically cloud-based desktop compute a couple of decades ago, and it never took off because people weren't ready for the concept. But these things keep coming around, and if you've been around long enough, you see those patterns.

So with all that said, I don't care what it is that you have. I don't care whether it's physically sitting on the ground, or it's hovering in the cloud, or whatever it is because I'm just interested in what it's doing and how I can get added and into it and figure out if it's happy or sad or somewhere in between.

Corey: The hard part I think is bridging the fundamental humanity of what this industry really is about with the technology bits. And for a long time, I think back in the early part of your 30 some odd years doing this, that wasn't really as, I guess, front and center. It seemed to be at that point, there was a serious jerk problem in our industry. Whether there still is or not is a subject for a different episode and an entirely separate debate.

But right now, I think that there's a certain empathy required to get in the front door, at least any place we'd want to work. Back then, that was far from true. If you knew how to write code and offend everyone around you, well, I stopped listening after the right code. Come on in. And what's strange is that you're one of the most empathetic people that I've had the privilege of speaking to about a number of things over the years. And I guess I'm wondering, how did you become that person? Or have you always been that person, and you just learned to walk amongst the jerks?

Leon: First of all, I think that you're right. In the '90s, which is when I started getting into IT, the lone genius was still very much the model. It was also the lone owner. Remember that in the '90s was when the single guy who invented PKZIP created his utility and sharewared it until somebody bought it, and he got his, forgive the euphemism, but his fuck you money. And he's out. And lots and lots of people did that with lots and lots of software.

So the idea of the brilliant loner who could do it all himself except run the business or hire people or even talk to people, but it wasn't necessary because it was, again, that loner mentality was very much in vogue. In corporations, you had the crystal tower. You had the mainframe operators or the data center or the data processing center. And again, this was a cobble of arcane wizards who could speak their own language but couldn't really relate to anybody else in the cafeteria.

But also remember that the flip side of it, that at that time, people who worked in IT were pretty heavily mocked and kept at arm's length as outsiders. Nowadays, for the last, let's say, 15, 20 years, it's become eminently evident that the business can't run without IT, and IT has no purpose without business. The second piece is taking a little while for IT people to grok, but we're getting it.

But businesses have realized that they really live and die based on how well they technology. And that's brought the jerk problem to the fore. And I think it is getting better. I don't think it's perfect everywhere. And to your point, the places where we want to work are the ones that have solved it. Sometimes, companies don't advertise, "We employ jerks here." So yeah, sometimes it's trial and error.

As far as how I got there, like I said, my degree's in theater, and I got into tech because at the time, the definition of a trainer was the person in the room who knew what was on the next page of the manual. They didn't have to know much else. I give a really funny DOS basics class, really thigh-slapping funny, and Windows basics and all that stuff. And I did a lot of that.

So when you're in a room with five to 10 people who used to be part of the typing pool on their IBM Selectrics, and they now know that if they cannot learn what they need to know in the next six hours, they will no longer have a job, and that's what comes in at nine o'clock in the morning when class begins, and it's on you. You do need to develop a bit of empathy. You do need to make sure that they know that you know what's on the line for them. And that continues. Again, desktop support, you're walking desk to desk, or you're working a help desk, you really do have to know what this person's going through, what they need, or else, you're going to hate them, and you're going to hate your job pretty quickly.

So I like to think that I'm a naturally empathetic person, but it hasn't hurt me. Also, being an extrovert has made me a flying pink unicorn of IT, especially back 30 years ago when it was populated by a lot of what I call green sock, blue sock people, the ones who didn't match their clothes and really didn't relate to other people particularly well. And as you can tell, I'm really shy. So that also has helped me to be in the kind of role I'm in now, which is, really, a storyteller and a story listener, listening to people and finding out what their challenges are and how they overcame them, or they haven't yet overcome them and putting them in touch with the right people, or helping to develop the right solutions that will get them on their way to fame, fortune, and success, or at least getting home on time.

Corey: I frequently said that multi-cloud is a stupid best practice, and I stand by that. However, if your customers are in multiple clouds and you're a platform, you probably want to be where your customers are unless you enjoy turning down money. An example of that is InfluxData. InfluxData are the manufacturers of InfluxDB, a time series database that you'll use if you need a time series database. Check them out at influxdb.com.

And that's really the trick is it seems that finding a way to meet people where they are is critically important. People sometimes ask how I wound up becoming the monster that I am now. And the piece that I think that I, I guess, give some of the most credit to has been I started off giving a bunch of terrible freaking conference talks in 2012, 2013 talking about SaltStack, which was a fantastic technology that soundly lost the configuration management wars. And that was fun for a while, and it turns out that putting your documentation on a slide and reading it to people is a terrible freaking way of building a reputation and learning how to give a fun talk. So in time, I asked people, "So how did it go?" And their response was, "Well, that was great. Now, here's how you make it even better." Yeah, people had the grace to approach me in that way.

And a year later, I wound up getting a traveling trainer job under contract for Puppet, which lost the configuration management wars in a radically different way. And that was a fantastic learning experience for lack of a better term because again, I knew what I was doing, but I also wasn't so confident about it that I felt like I could teach it. That changed rapidly because it turns out when you teach the same thing week after week, you learn how things break. And because of some of the dynamics we've already touched on, I was, in many cases, dealing with people who thought that I was the representative of the company that was coming to take their job away. And that was a interesting story. Then the demos would break periodically because hey, computers, and good luck, and oh, if they hate you enough, you don't have this job anymore.

So it was a really interesting story as far as learning to A, think on my feet, B, present to a potentially hostile audience, and C, amuse myself. And there was a weird progression there as the first one, I was freaking terrified. The second one, third one, was okay. I got it. The fourth one, no, I don't got it. But over time, I got better at it, and it went from exciting and terrifying to rote to routine, and then it was a matter of finding something new. But that was the training perspective.

What was interesting to me, it was talking to a lot of the students and seeing their transformation, their progression as they started wrapping their head around how a lot of this works. Now, that was Puppet, which was interesting and had its own DSL style of approach, but okay, fundamentally, all it was doing under the hood was managing some baseline primitives, some files, users, permissions, run random executions there and install packages, and you're more or less done. I can't imagine what it would look like today to be a trainer to, "Now, I'm here to teach you, pick your cloud provider of choice, how to use that in the world that you're operating with," and everyone just stares. How do you begin? How do you even start with something like that? I think it comes back to empathy. I think that trying to put yourself in that situation, even if you're not, is critical. I think the best teachers are great at walking that journey with their students.

Leon: You're a lovable, fuzzy, cuddly monster whom no one should ever cross because you will eviscerate them on Twitter if nowhere else.

Corey: Those are your words, not mine.

Leon: I understand. And you had a lot of what I think a lot of us call character-building experiences. Those are always good. But I think you're right. I think, first of all, you never understand something quite as well until you've had to teach it to somebody else. And the process of teaching not only forces you to build what I call technical empathy, the ability to see something that is familiar to you from a completely foreign perspective, but also, it gives you a chance to hear how others are using something with which you are somewhat familiar.

Again, once you get past that first class you taught that was terrifying, you were still familiar with it. You just didn't know what was going to happen. But you're taking something that you are familiar with and seeing it through their experience. And sometimes it's like, "Oh, wow, I've been using this completely differently than anybody else," which can either be, "Well, I'm a genius. I found this new use for dragon's blood," or it can be, "Wow, nobody understands what this product is supposed to do." So then you become an evangelist of showing people the one true way to Puppet or to whatever.

But at the same time, one of the best parts about training isn't, "Hey, I know this. Let me show you." The best part about training is, "Wow, that's a really interesting question. No one's ever asked that before. Let's find out."

Corey: Those were the best moments. The fun thing too is at some point when you're teaching the same curriculum pretty consistently again and again and again, you can find the common failure modes so frequently that you don't even need to walk around and look at their screen. You can tell them they dropped a comma, and they look at you like you can read their minds. And you can't. You just see the exact same thing manifest itself enough times that by the time it comes around there, it's old hat. The first time you see something, it's new and revolutionary. The sixth time you see something, oh, well, that's just how it works. I've seen it. There's no magic to it anymore. I've seen how the sausage gets made.

Leon: Right, so I, again, talking about that experience both from your side of the teacher's desk and the other side of the teacher's desk, one of the things I love about the Orthodox Jewish educational system called yeshiva is that the highest compliment you can be given by a teacher is not, "That's a really good answer." That's not the highest compliment, the highest compliment, and it's always given in Yiddish, Du fregst a gute kashe, you asked a really good question. Now, the funny part about it is that it may not be an original question. Remember that Jewish thought and Jewish conversations have been going on for at least 3,000 years.

Corey: Jewish arguments but only holidays.

Leon: Well, question, answer, argument about why that's the answer, et cetera, the dialogue, the debate has been going on for quite some time. So the odds of you asking an original question are fairly low. That's not going to happen. But the fact that you asked a question that somebody from 1100 or whatever also asked means, wow, you are thinking along the same lines as these great minds. That's a really good question. Now, let's go into the answer. Now, let's dig into where that line of thinking takes you.

And again, I think that's a very IT practitioner way of thinking. We don't get into this business because we want to stay in a steady state of I already knew that. None of us... Because if that's what you need, please, accounting, or I don't know, almost anything else. And you can say, "Yeah, I know that. Things have changed very little," but in IT, oh, it's Tuesday. Everything's about to change again.

So we revel in the idea of let me find out. I think that's our sweet spot as folks who build our careers in IT. So yeah, when you're teaching, what else is there, right? You've seen all the common failure modes. You can read their faces, if not their minds. Occasionally, they ask an interesting question. That's the part that you're looking for, for that spark.

Corey: I think that that is the single biggest divergence that I've seen from engineers and IT folk and the rest from the start of my career to the present day where... I'll use myself as an example. I started off my career as a subject matter expert in large scale email systems. I considered myself an expert in Postfix of all things because, not to toot my own horn, I was because it's a sad thing that not many people had to care about. But I did, and it was fun.

But I looked around the ecosystem and saw that things like managed email services were very clearly on the rise, and given at the time I was 23 years old, give or take, and I didn't have the confidence that this was going to be something that I could build an entire 40 some odd year career out of, I felt the world changing out from under me. And my answer to it for better or worse was, okay, well, if this isn't going to work, then what can I do instead? What do I pivot into? As we've already covered, my pivot was into configuration management, and then that changed, and now, I claim to know a lot of things about cloud, and I don't. But no one calls me on it because I'm confident and I'm loud on the internet as you've observed.

The fun part though is that you have to reinvent yourself. Let's say that AWS crashes itself into the sea through its own terrible naming of services that drive it out of business. I have two choices. I can either continue to shake my fist at a platform that is eroding out from under or I can pivot what I know into something that has more business relevance. And I say this now, and it sounds ridiculous, "Oh, how would Amazon ever go away?" Yeah, go back in time. How would Solaris ever go away? And you start to see that these shifts happen all the time. I don't think there is anything safe in the world of technology that you could start working on today as a new graduate and still be doing something even remotely related to it 40 some odd years later. And that's going to be a fascinating evolution point.

The question is, how will you reinvent yourself? Now, it's scary. Don't get me wrong, but it's also not what people think it is because I don't think you're coming back at this from a perspective of starting over at entry level. No, you take the things you know and parlay that laterally into something else that's tangentially related to the thing that you were working on and bring the baseline fundamentals with you. And it doesn't take long to develop some semblance of mastery in this new arena given a baseline level of experience and exposure where you started.

Leon: Yeah, it's amazing how many things dovetail, and it's also amazing how many things that are tangential to the work you do today become front and center later. But you realize that that tangential experience really gave you everything you need. And here's a simple but powerful example, as I said, one of the things that I taught early on in my career was word processors, and WordPerfect for DOS was a thing, right? And WordPerfect had a function called reveal codes. Bolding wasn't just a matter of highlighting because there was no mouse, and therefore, there was... You had to go in front of the word, and you had to say, "Make it bold," and then you go to the end of the word and make it turn off the bold.

And there were these little brackets, the greater than, less than symbol, and it would say... Sorry, no, they were square brackets. Let me go back there. And there were these little brackets, just the square brackets. And it would say bold at the beginning, and then you'd move to the end and say bold at the end and those kinds of things. And to see to make sure that you hadn't accidentally bolded an extra space or something like that, you would turn on reveal codes.

And so you'd see in reveal codes, words and these brackets, bold, paragraph, heading, underline. Sounds kind of like HTML, right? Because it kind of is. And when HTML came around six, seven years later for me, it took me zero time to get up to speed and start doing nascent web development because it's like, "Oh, it's just like reveal codes in WordPerfect." Now, reveal codes wasn't the be all and end all of WordPerfect. There was plenty of other things that that word processor did. And I am proud to say that I was a WordPerfect certified resource. And I fought really, really hard to get that certification, and I'm never giving it up. But reveal codes was not the most important piece of it. It was tangential. And yet, five, six years later, there was something where I was pivoting from doing desktop support to building a web design company on the side, part of my side hustle. And that one experience allowed me to really pivot quickly and smoothly into that new area.

So the mindset for those people who are listening, and they're thinking, "I don't have it in me. I can't do it," the mindset is not, "I'm going to reinvent myself. I have to completely transform myself." The mindset is to just remember that as IT practitioners, we have committed to being lifelong learners. And we learn things, whether it's to paraphrase an XKCD cartoon that one weekend I spent screwing around with Perl, or that one YouTube video I listened to, or that week I really dug deep into XML, or what have you. Those are the things that feed the next thing you're going to do. You will be amazed at where those tangential experiences crop up, and they end up becoming like a secret superpower that you didn't know you had.

Corey: If people want to hear more about the wise things you have to say, whatever they might be, where can they find you?

Leon: So you can find me... My blog is adatosystems.com, and there, I list out all the things I've written and all the videos I've done. But the other thing is I run a podcast also because I am a middle-aged white man, and we all have our podcasts.

Corey: A group of middle-aged white men is collectively known as a podcast.

Leon: Right, exactly, that's the grouping, right? Fish, they're in schools and-

Corey: Developers are merge conflicts, et cetera, et cetera.

Leon: Right, right. Exactly. Yeah. Head geeks, by the way, move in chaos. There's a chaos of head geeks. So anyway, so I have a podcast. And myself and Josh Biggley, who is Mormon, and let's see, Keith Townsend, who is Baptist, and Al Rasheed who is Muslim, and you're starting to get the theme here, we've all been in it for decades each. And there's a few more of us. And we've been in IT for decades each. We all have a very strong religious, moral, ethical point of view. We've got evangelical folks. We've got atheists. And the podcast is called Technically Religious, and we talk about the intersection between our strongly held religious and moral ethical views and our work in IT and how we make those two things, in fact, not conflict, but become synergistic between the two. And so you can find that at technicallyreligious.com or anywhere that fine podcasts are peddled to the masses.

Corey: Excellent. Thank you again for taking the time to speak with me today. I appreciate it as always.

Leon: Well, thank you for having me. This was amazing. I can't wait to hear future episodes of your stuff and to see you at another conference with old fogies with retired technology.

Corey: I don't believe we'd be able to recognize each other if we weren't wearing conference badges at the time.

Leon: Exactly.

Corey: Leon Adato, Head Geek at SolarWinds and all-around decent human being. I'm Corey Quinn. This is Screaming in the Cloud.

Announcer: This has been this week's episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Adam Jacob
Adam Jacob is a co-founder of Chef Software and the creator of Chef. He has over a decade of experience designing, building, and managing large production systems. Adam is Chief Executive Officer & Co-Founder of The System Initiative.

Before Chef Software, he founded HJK Solutions, an automated infrastructure consultancy where he built production cloud infrastructures. Adam has been responsible for large production systems, internal corporate automation, and Sarbanes-Oxley compliance efforts.

Links Referenced

  • Twitter: @adamhjk
  • LinkedIn: https://www.linkedin.com/in/adamjacob/
  • Personal site: https://sfosc.org
  • Company site: https://www.systeminit.com

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host Cloud economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud. Thoughtful commentary on the state of the technical world and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored by InfluxData. Influx is most well known for InfluxDB, which is a time series database that you use if you need a time series database. Think Amazon time stream, except, actually available for sale and has paying customers. To check out what they're doing both with their SaaS offering as well as their on premise offerings that you can use yourself, because they're open source, visit influxdata.com. My thanks to them for sponsoring this ridiculous podcast.

Welcome to Screaming in the Cloud I'm Corey Quinn. I'm joined this week by Adam Jacob, best known, I suppose, as the co-founder of Chef Software and the creator of Chef. These days he's the CEO and co-founder of The System Initiative, but today we're here to talk instead about a fight we picked on Twitter.

Adam, welcome to the show.

Adam: Hi, Corey.

Corey: So as of the time of this recording, which in November 15th, yesterday it was announced, or possibly the day before, I saw it yesterday, that beloved password management in the Cloud, there's the Cloud angle, company, 1Password received from Accel partners, or Accel, whatever they call themselves these days, a $200 million investment. Have I nailed the salient facts so far?

Adam: Feels good to me, yeah.

Corey: Excellent. I'm sure it feels good for them. And my immediate shoot from the hip take on Twitter, which is why they call them hot takes and not lukewarm takes or reasonably room temperature takes.

Adam: More like incomplete takes.

Corey: Exactly, it was extremely incomplete and... But my take was, "Okay, that's unfortunate, because now to receive the typical VC 10X return on something like that, they're going to have to build something monstrous." And your criticism of that, or I guess, response to that, which is valid in some ways, is that it was a lazy take, and it was. I was lying on the couch when I tweeted it.

Your position was that the same leader who was there last week is still there this week, and that if it becomes something horrible, it's not VCs that are directly to blame, so much as the leaders of the company. And that was a fascinating enough position that I couldn't immediately refute it and figured, let's have a talk.

Tell us more. I'd love to hear where you're coming from on this.

Adam: Well, yeah, so look if you don't know who I am, and Corey gave you a little of my background, but I wrote thing called Chef and then I co-founded a company with some lovely human beings called Opscode, later called Chef, and now I've done it again, I have a company called The System Initiative. All of those companies, both of them... Took a bunch of capital, I started a couple of businesses before I started ones that took venture capital. I started a consulting business, we had a body shop for a while, long ago, and the thing for me about right this second in history, is that it's a very, very hip moment, in particular, at least in my Twitter feed, it's particularly hip thing to just be like, "Venture capital ruins companies."

It just happens sort of instantaneously, kind of consistently. And so I was kind of reacting to that part of your tweet, whether just sort of ambiently I had seen a lot of that sentiment. In this particular case, it's... I think it's... there's a couple of pieces there, right? So one is the idea that when companies that venture capital, they are somehow, I don't know, ruined. And most often, I think that take comes from the idea that venture capital requires a return, which is true. So they're not giving you money for nothing. The deal is they give you money on what is actually pretty great terms, most of the time. Certainly, it's not a loan. On the hope that you create a business that's worth more than they invested.

And ideally, they'd like that number to be a lot more than they invested, but it usually isn't, right? So their own math, and the way they talk to their own customers, has that number being relatively small in terms of the ones that win, right? What people tend to believe that there is... Once you take that venture money there's some kind of Svengali like backroom that gets created, where they take otherwise pristine and virtuous founder/managers and turn them into groveling, money hungry shit birds, who are going to destroy the thing that they loved or that they cared about, because there's this person that gave them money, who wants a bunch more money. And I just...

My experience is that that's just not what happens, and not how it happens, and that's not only my experience, personally, but I talk to a lot of people who are in startups, who are starting startups, I do due diligence for the venture capital folks that I like, I talk to a lot of people. And when we fail, we most often fail because management screws it up, because the people that you like and love, they just made bad calls, and usually, it's more than one bad call, and it was sort of a series of them. And I'm deeply sympathetic to the complexity of those decisions, but ultimately, the deal is venture capital gives you money, in the hope that management builds a great business. And, yes, in some cases, certainly, they take a board seat, sometimes there's so much money in that they have control, in terms of being able to decide who the leadership of the company is. But, ultimately, the job of the founders or the job of the CEO or of the executive team or whatever you have it, is to go create a good business.

And if they do that, then everybody's aligned and there's now of the bad behavior that we think goes on. And then if there is, usually, it's very rare to have a lot of drama in businesses that are high functioning, right? So if the management team's doing a good job, you don't tend to see a lot of the founder displacement or executives being forced to do bad things. If you think about it, if you were on a board and you really believed that the company's future was at stake if you made the... continued on a course of action that you were taking. And your choice was to force a person who doesn't want to do the thing you think is so foundationally important to do your bidding, or fire them, what would you do?

Corey: It's a terrific question. And I've done a little research, since my tweet. It turns out that, wow, oh, wow, people are talking to me about this and it's not always the hee-hee you're funny feedback that I have grown addicted to on the Twitters. It was critical feedback, not just from you, and it was interesting, because it seems like my feed is pretty evenly divided. I've done some homework since then, and it seems that by all reports that they currently have a 174 employees, which to me is, I guess, it changes my perspective ...

Adam: Pretty big.

Corey: ... on the company.

Yeah, I was thinking this was five or 10 people sitting in a room. Now, that's still over a million dollars raised per employee, and it's also, according to Crunchbase, the arbiter or such things, the first time that they've taken funding-

Adam: That's right.

Corey: ... from an outside source, which the concern I have now is that I... The reason I care about this is not because I enjoy crapping on things for fun.

Adam: No, no it's because you love 1Password.

Corey: I love 1Password.

Adam: 1Password is my most beloved thing.

Corey: Exactly. I take a look at my dock or my menu bar on my Mac, and that is the thing that I am the most accustomed to seeing that automatically, if you'll forgive my term, sparks joy. And if I look at the-

Adam: It's the first thing I install.

Corey: Exactly. It needs to be, and I look at the rest of the dock, and I look for great applications that were developed by software companies that took a lot of VC money, and I don't see them. I see things that started off terrific and beloved, and have turned into something kind of monstrous, Dropbox. Sorry, something caught in my throat there.

Adam: Oh, but come on. On mind, right now, Slack, iTunes. The idea that it was the source of the money is crazy person talk, and if you were the manager of 1Password, you have this great application... The problem with 1Password is my mom doesn't use it. So if you... I bet that business has crazy good unit economics, so for people who don't know what I'm talking about, it's can you produce the goods for less than you sell it? It's the key feature of telling you you're going to have a good business. And so every copy of 1Password they sell, I bet they make a bunch of money on. And the side effect of that is that if they could just get more reach, if more people knew what they should be doing is running 1Password, they don't have to do anything different than what they're doing already. They just need to reach 10 times the audience.

At which point the odds that that audience would love it just as much as we do is very, very high, right? There's no question my mom could use 1Password, she just doesn't even know it exists.

Corey: It sounds like you're not as good of a son as I am, I got my mother on to 1Password a couple of years ago, and suddenly the, I forget my password calls stopped, which was-

Adam: Oh, that's amazing.

Corey: ... amazing. Oh, yeah.

Adam: Yeah, I should have thought it through. Well, now that I'm doing it, I like, "Oh, I'm not a good son." But from a business point of view, the alignment that says what we should do is take $200 million and just drive those unit economics as hard as we can. Just put as many dollars into getting the word out that 1Password exists, that you should have it, that this is what it does for you, this is how you use it, that you can support that massive influx of users. It's not rocket science to tell you that that's what you should do. And the odds that the people who took that money are smart people who do good things, pretty high, because there we are talking about how much we love 1Password.

Corey: And that's part of, I think, the problem, because people have seen this as a small, scrappy, bootstrap company, where I believe in them, I vote with my wallet. A number of people complained heavily, oh, did they complain. When they switched away from their historical model of each version costs money, you buy the software, it's yours. And they switched away from that over to a subscription model. I personally, have no problem with that. I'm a big fan of, I love your service/product, charge me for it, and whatever allows you to build a sustainable, scalable, reasonable business.

And I have no issue whatsoever with them deciding that subscriptions were the way forward. Is it my first choice? Maybe, not, but I don't know what it's like to run a company like that. I can't imagine some of the pressures they must be going through.

And that's kind of the challenge I have, is your narrative, where they take that $200 million and plow it directly into marketing, and that's the last money they'll ever need to raise, until they go public in a few years. With the right emphasis on the correct things, especially, if, according to legend, they were profitable to start with. The-

Adam: Yeah, I mean, you'd have to do it really wrong. With $200 million, at a 150 people, or whatever they are, I don't know their revenue, but let's assume that they're halfway to a 100 million in ARR, and that they're growing, whatever, 20% quarter over quarter, year over year. 20% year over year, right. It's a pretty solid bet that if that's the shape of, or anywhere close to the shape of the curve that with $200 million and an investment in sales and marketing, they're done. They'll be a public company that has great economics and they'll IPO, and their investors are going to be happy at the incredibly safe, relatively safe investment they made.

Corey: Absolutely. The other piece of it though, is I migrated off of LastPass, after some very weird customer service experiences there, and some implied data breach stuff that wasn't taken particularly seriously, over to 1Password. Migrating wasn't the hardest thing in the world. There's a great importer that 1Password spun up. Would spinning it over to something else be possible? Sure. Would it be painful? Definitely, it would kill an afternoon at least. And the question is is what would it take to get me to completely turn my back on 1Password? The only realistic concern that I've had until now, has been a data breach, where if I can't trust my password manager to securely protect the passwords, given that is the single, sole job that I expect them to do for me-

Adam: Sure.

Corey: ... then you're right. I don't see that there's too much of that they can get wrong, as long as they keep the passwords safe. Now, the other side of that though is with enough money, let's say it wasn't 200 million, let's say it was, I don't know, the Vision Fund 2.0, where they need to take, yet, more investors who have bone saws when things go wrong. And at that point they're hurling piles of money into it, let's say it's two billion. At that point you'd almost have... There's no way they're going to spend that on sales and marketing in any responsible way, so you'd look at that and think, "There's got to be something else that they're going for." Is that also a naïve approach?

Adam: I mean, yes and no, right? So look you can take too much money, so in the grand scheme of venture capital and business, you take capital in at a valuation in the hopes that you then exceed that valuations terms and everybody ends up happy, right? That money comes with terms, there's a lot of complexity in those terms, but one good example there is preferences. So the odds that that $200 million came with a 1X preference is pretty high, right?

So what that means is, when they go to sell 1Password or 1Password becomes liquid, so either it goes public or someone buys it, the $200 million that Accel put in, comes out off the top first, and then everybody goes together, as if they were starting from scratch, right? So let's say, they sold it for a billion dollars, or they went public at a billion dollars, so the $200 million comes out first, then the 800 million gets split up according to whatever the shareholding percentages were, right?

So how you get in trouble for taking too much money, is either take too much money at too high a valuation, or take you just a little bit of money at a crazy valuation. It's really the valuation that gets you in trouble, right? So when you take any money and you say that it's worth a billion dollars, you have to live up to those evaluations, right? You have to actually get to a place where what you build is worth more than the money that the valuation bank was investing.

I think that, so in the case where they took in $200 billion, look any company who's taking in billions of dollars in venture capital, the only reasonable way to do that would be at a crazy valuation, right? So if you took $2 billion and you were a 180 people and you had sub a $100 million in revenue, that's a big valuation, right? Even if we assume that you were only expecting like a 2Xor a 3X return, which is likely what someone like Accel is accepting. We tend to talk about venture as a 10X return, but that's when you're doing what I'm doing with the System Initiative, where I don't have a product. I'm empty, and you're just filling me up. You're filling up the tank with money, and then hopefully, I turn it into some real stuff.

But, ultimately, for them if they took a $2 billion influx of cash, the concern would be, it was so much money on so little business evidence, that there's no possible way they could live up to their own obligation. And that sucks too. That situation is also a management problem. So, yes, it's an investor problem, you would love to have investors who just say, "No." And lots of people will say, "No," to those crazy terms. WeWork was a good example. Lots of people said, "No," to WeWork. Lots of people would have said, "No," to 1Password, if 1Password said what they wanted was $2 billion. So that's the trick, is that what kills you isn't that they have $2 billion and now they need to spend it, and therefore they have to do some crazy shenanigans.

Everybody is aligned, investors are aligned, 1Password management are aligned, 1Password's employees are aligned, but what they should do is build a great product that people continue to love, and sell it to more people. It's not that hard. Now, if you set the goal for that alignment in this unrealistically, humongous way, then sure bad things will happen, because you'll likely fail to live up to those obligations. And so as a manager, as a person who takes the investment and decides what to do, you're balancing both kinds... You're balancing the upside with the downside risk. What if I don't make it? So 1Password takes in $200 million, let's assume that they didn't give up control of the company in exchange for the $200 million. So the worst thing that happens to them, if they didn't give up control, and I bet they didn't, means that they have an unhappy investor.

That's pretty much it. If they fail to live up to their obligations, and the business is generating revenue, and it makes more than it spends, it's kind of status quo-y, right? They're very little risk.

Corey: I've frequently said that multi-Cloud is a stupid best practice, and I stand by that. However, if your customers are in multiple Clouds, and you're a platform, you probably want to be where your customers are, unless you enjoy turning down money. An example of that is InfluxData. InfluxData are the manufacturers of InfluxDB, a time series database that you'll use if you need a time series database. Check them out at influxdb.com.

The challenge I see, I think, is that I believe very strongly in what 1Password does, and as their core focus of that single thing that they do, namely keeping my passwords secure, safe, and available for me, but not to others, that’s a terrific business. I don't necessarily see a path to going public. I mean, the blog post that Accel put out announcing this, they mentioned three other companies that are 1Password customers, Slack, PagerDuty, and Dropbox. All three of those have gone public in recent years, and they all are sort of faced with a similar problem, from the outside, which is, okay, you're public, but right now all you do is this one core function. PagerDuty they wake you up in the middle of the night. Slack, it's IRC for work as a service. And Dropbox is one folder on your computer that syncs to all your things.

And the challenge is is all three of them are now self-describing as platform companies, where they're expanding well beyond what they initial niche that they needed to fill. So now we take a look at this and we see that the market for better or worse seems to push these component companies into becoming platform companies. And I guess, I'm concerned ...

Adam: Let's be clear they're all public companies now though, right? So-

Corey: Oh, absolutely.

Adam: And so the market we're talking about isn't venture capital Svengali, do you see what I'm saying?

Corey: Yes.

Adam: You're talking about public market Svengali's. You're talking about everybody else who wants to put money in them is saying, "Where are my growth dollars? Where's the growth honey?" And that's a different question. If the answer is, I'm upset that the market wants growth in public companies? Yes, sure-

Corey: Well, that's the question, is it possible to raise $200 million for a company like 1Password in its current position without the hope and expectation of taking it public?

Adam: Of course not. No, of course, they're going to go public. Of course, they're going to. No, and it would be ludicrous to take it, if that wasn't your plan. So to be clear, the only reason to take that capital is because your intent is for that capital to become liquid on a time horizon that everybody agrees on. If we don't all agree that that's the point, then don't take the money. Talk about bad management insanity. If you take someone's money, that explicitly says, "Hey, I exist in order to become liquid. I exist because I take pension fund dollars from Ford or from Teacher's Unions, and I take some sliver of their pension and I stick it in a high risk asset and venture capital, and then I took some of that money, and I didn't intend to ever return it? That's crazy."

You see what I'm saying? That's nuts.

Corey: Oh, yeah.

Adam: Yes, of course, their intent is to go public, but the idea that a thing that you love, that they people who built 1Password, that they shouldn't go public, that instead they should be like, I don't know, content with their knitting or they should go drive race cars like DHH, because it's more morally pure to build a business that doesn't take dirty, dirty venture capital. Fuck all that. None of that is pure. None of that is. These are just... The people who built that business, and who decided to do that thing, the idea that what they want to do is convert that wealth into liquidity, that is not insane. It's just not.

And the best path for them to turn that into liquidity while remaining the 1Password you know and love, is public market, because the alternative is private markets. And in the private market it means somebody else is going to run that business, right? Somebody else is going to take that thing and they're going to shove in some direction that you wish it didn't go. And maybe it'd be okay if Apple bought them or whatever, but I don't know, Apple's pretty bad at running services, right? So if I had a choice about do I want those people who built this product, who I trust, to continue to have that trust, do I continue to trust them, and do I believe that what they're going to be able to do is go build a public company? Absolutely, I do. Right?

Does that mean I'm going to agree with every product choice they make? No, but PagerDuty is a great example. Do I trust the leadership at PagerDuty, who's the same as the people before they went public, are going to make reasonably good decisions about PagerDuty? I do, right?

Corey: Oh, yeah, they're great people over there.

Adam: They are.

Corey: I agree wholeheartedly, but even now you see that... I love PagerDuty, as a great example here, or what they were, which is the thing that wakes me up when I have a system that wakes me up, and they are spectacular at that. I've given conference talks, where I have their logo on a slide with a speech balloon that says, "Wake up, asshole." Because that is the function that PagerDuty provides, and they have the grace not to sue me to death over that.

Adam: Sure.

Corey: But as now they're talking about, oh, expanding into some of the higher level... They're becoming a platform, to use their terminology. I see it not being great for the core mission of the product as such, and that's sad and I can get over it to some extent. I can grouse about it significantly more with Dropbox, but the concern I have is that if that pattern is true, and it does point to something systemic, with my password manager now needing to grow beyond managing passwords into other things, that's a scary, scary prospect. And maybe I'm just too early to the fear mongering place.

Adam: Look I don't think you're too early to it. I think that in one of these categories we have human beings who in general we trust and like, or at least want to believe that we do. I couldn't name a single person who works at 1Password, but I can name the people who work at PagerDuty. But on the other, we have this big bucket of mushy bankers, right? Who as far as people can tell, have no idea how they work. What we know is that they have a lot of money and they put it in stuff that we think is dumb. And then we hear about the ones that were spectacularly done, WeWork, and we're like, "Man, those guys are jackasses."

And sure, okay, sometimes they're jackasses. Also, the one thing that's always true, in all of the things that mess up in the way that we're talking about, is management made terrible decisions. If PagerDuty becoming a platform company is going to destroy the value of PagerDuty, so that their core business gets eaten away, they're going to get destroyed in public markets. And that will mean that their leadership made terrible decisions. It's not because they were having a board meeting, and the board made them do it. They wouldn't make her do it, they'd fire her. And they'd put somebody in who would do the thing that they believed they needed to do. And that doesn't mean that she's doing it against her will, I don't believe that for a second, because if she was doing it against her will, they'd have fired her already.

Do you see what I'm saying?

Corey: Absolutely. Oh, absolutely.

Adam: There's no fight here like that. And so sometimes I'm sure that happens, but way less often, and so in the chance of 1Password they, so far, have made decisions that make a product that I love. Do I believe that they will continue to make good decisions with the product that I love? All evidence tells me I should believe that, because they've continued to do it over the course of years and years and years, which is why I'm a loyal customer and I love the product. And so the idea that they're choice of capital infusion so that they can get liquidity, is a risk, that's not the risk. The risk is are they going to be jackasses? Which was the same risk you had before they took the money, see what I mean? All it takes is one person deciding that the most important thing for 1Password isn't keeping your password safe, and next thing you know, you're moving off 1Password. Right?

Corey: Oh, yeah. Or they start doing the Google approach of selling customer data, and given that these are passwords, that does raise some eyebrows. I'm kidding, I'm kidding. I have no-

Adam: No, but it's a great example.

Corey: ... imagination that they're going to start doing that.

Adam: But it would be... that would be a terrible business decision, and if they decided to do it, the backlash would be swift and intense, right?

Corey: It absolutely would.

Adam: They'd be real bad.

Corey: The question I have is, and I see-

Adam: And they probably won't do it.

Corey: I do see what you're saying, but isn't there a counterargument, where you take a look at the leadership that you have invested in, that you trust, who are still running the company, until now they really, I guess, from a naïve perspective, were in charge of their own destiny. They could make whatever decisions they thought were best, and they still are in charge of the company, they can still make decisions, but now they have someone who just wrote them a $200 million check tapping their foot and saying, "Well, I have some suggestions." And at some point those suggestions have a very heavy weight, that would almost certainly influence how the management winds up addressing things.

I mean, you took VC money back when you were at Chef. When investors said-

Adam: I've taken a lot.

Corey: ... "We think you should..." Yeah, when people said, when your investors said, "We think you should do X." Whether you did X or didn't do X, you certainly owed them a response including your thoughts on X, I would imagine, no?

Adam: Sure, but let's back up. If Chef... Most businesses aren't like Facebook, or even 1Password, they don't just become awesome and then stay awesome. Do you know what I mean? We draw charts that are like the curve up and to the right. That's not the truth. The truth is it's like a weird stair step, and at every moment it flattens out, and you're like, "Oh, my God, we're all going to die." Right? And depending on how long it stays flatish as opposed to sort of tight and towing to the curve, the more everyone gets nervous, because you don't know what's going on. And you try to fix it, and you're not sure, and you're just throwing stuff at the wall.

This is true of every business. It's not unique to venture capital focused businesses. When the business is great, when the economics are good, when your numbers are good, when everything is up and to the right, no one is standing behind you tapping their foot. The only thing that they do is say, "Really good job. Keep doing that work." That's it. Board meetings are the best, because all that happens is it's great, and they go, "Keep on being great." And you're like, "Sure will." And everything's cool. And if they have suggestions, it's about how you can be cooler, how you can be greater. Not how you can could fix the mess that you're in.

Now, business stops being great, things get weird, you start selling the metadata from the password vaults, which causes everyone to flee the service. Do you see what I'm saying? Bad things happen that change the fundamental shape, absolutely, you investors care about what's going on, right? In those moments they tend to not be... it's not like they're hostile, where they're like, "Oh, we better make you do bad things." Their alignment is the same as yours. What they want is the business to grow and to work, so that it can become liquid, so that they can get their returns that were the reason they were there in the first place. So, yes, there are bad venture capitalists who freak out, who drive all kinds of bad behavior, who do bad things, because they're bad board members, and it's a real pain for the people who have to manage those companies.

In general though, the thing about it is that even a bad board member is a good board member when the business is killing it. Does that make sense?

Corey: It does and-

Adam: There's just... There's no... it's really easy. So 1Password, today they have one board member, they probably have a minority stake in the business, and maybe they have some terms around liquidity. They probably have some terms around a right of first refusal on a transaction or something like that. Other than that, they're going to come to a meeting once a quarter, maybe, and they're going to talk about how the business is doing, and the answer's going to be, "We're killing it." And everybody's going to shake hands, and they're going to go away, and then they're going to have conversations like bankers do about whether now is the time to take it to a public market, yes or no?

And what are those metrics? And how will we do it? And it is not a Svengali product level design conversation, not at all. Not in my experience and not in the experience of the people I know, who've been through it.

Corey: I think that you have an excellent point. I think that the next step on this is likely going to be that we do another podcast recording, whenever they publish their S1. And then we can take a look back and see how things wound up progressing over the intervening time, and how this impacted the trajectory of the company, the numbers they disclosed, the risk factors. And I think that history's going to be the judge here, until then, as much as I hate the term, I'm just being a pundit.

Adam: Yeah, I mean, we all are, right? One last thing before we go that is one of the reasons that I'm so... I feel so strongly about this point is that one of the worst things in Silicon Valley is the degree to which we believe that founders are infallible, or that all the good things that happen in the companies or products you love come from the people who founded those companies. And don't get me wrong, often they do, but-

Corey: I can't imagine every problem I ever had with Chef, when it wouldn't compile, when I had weird issues that I couldn't understand, I assumed they were all your fault, personally.

Adam: I mean, mostly they were, right? But I think the thing about that founder focus and the thing about the reasons... one of the reasons I said that it was a lazy take to blame the venture capitalists is that we let bad founders and bad managers off the hook by saying, "Hey, if they just... those venture capitalists. Some people in some backroom made people do the blah, blahs." And man, it's just not usually the way it went down. And usually, the truth is, that leader was not a very good leader, and they failed to do the things that they needed to do. And it's okay that they failed, that doesn't mean they're bad people, it doesn't mean that they'll fail next time, it doesn't mean they shouldn't have another chance.

All that stuff is real, and also I've seen more entrepreneurs and more great businesses and more great ideas die because they failed to realize that their job was to build a great business. And everyone around them told them that their job was culture, or their job was product, or their job was any number of other things. And what they wind up with is maybe good on those things, but definitely not good businesses. And then once that business has to be... Well, somebody has to deal with the mess. Then we say it's the venture capitalists fault. Then we go, "Oh, man the mess here that that venture money forced them into these bad contortions."

And no way, no way. It as a mess, somebody had to clean up the mess, maybe they did a good job, maybe they did a bad job, but don't make a mistake about who built the mess, right? Don't let the people who built that mess off the hook, that's a mess. And it wasn't because of the money they took that it was a mess. Even if it was, even if the situation is that you took crazy money at a ridiculous valuation that you were never going to be able to reach, you took it, management took it. You negotiate those terms, you negotiate those valuations. You negotiate how much money you take. You negotiate those terms. That's a management choice.

And that, when we say, "Oh, VC ruined it." We take the actual people with actual control, whose decisions decide everyone's fate, and we just let them off the hook. It's insane.

Corey: To be very clear, my, I guess, initial distaste for this wasn't nearly as specific as blaming the VCs, blaming the management, or even blaming people on Twitter, which I know is my preferred direction to take things in, but rather the entire situation is, I guess, disheartening is probably the best frame for it. Because I like them when I thought of them as a small company. I liked them when I didn't really have to think about, will they still be doing the exact same thing in three years?

Because, like it or not, picking a password manager with all the different sites that live in there is a commitment for a period of time. Do you ever notice the number of things you have in your 1Password vault never gets smaller?

Adam: Sure, but I love Heavy Metal. I love Heavy Metal so much, and when you start getting into the specifics sub-genres of Heavy Metal there's always a band where you're like, "Ah man, I loved them when they were cool. I loved that band when they were true cult black metal. Before the keyboard showed up." And I get it that's true, and also it wasn't the record label. You know what I mean? They didn't make them take their guitars and write a pop rock record. They did it.

Yes, there were incentives. Yes, the machine drives them in a direction and you're on the conveyor belts of whatever, but fundamentally, they did it. That was their choice, and in the end with something like 1Password, I look at the people who built a cashflow positive business. They bootstrapped their way to 150 people. I bet they have fabulous unit economics. I bet they're raking in cash money, and in the end do I believe that those people will likely take $200 million and turn that into a public company that they can be proud of? Yeah, yeah, I super do. Because it's really hard to get to a 150 people, even when you have venture capital. It's even harder when you don't, because there's all that cashflow you have to balance. There's just a lot you have to do.

And do I have faith that they're going to figure it out? Absolutely, I do. I also, to your point at the... when we started talking, I have a metaphysical bone to pick in his whole conversation, which like things you never said. Do you know what I mean? It was a tiny tweet that said, "I hope VC doesn't ruin this thing that I love." And I was like, "Come on, VC doesn't ruin it.

Corey: Oh, yeah.

Adam: ... Leaders ruin it."

Corey: I look forward to it. I love the people charging into my mentions to have discussions about this stuff, it's great. I mean, because I'm a man on the internet and people don't normally call me out on things when they read things into what I say. Normally... And that means, of course, our experience on Twitter is radically different from that of most folks. The difference in this case though, is it's nice to have this conversation with someone I know and like. We know each other in real life. It wasn't some rando jumping out of nowhere just to go, "Well, actually." It was, "Oh, well, Adam's usually a very thoughtful, kind person, and I imagine he's being... Now there's something I don't get, let's dig in to it."

Adam: Yeah, but I mean, there are plenty of people... That conversation is spun. There are plenty of people, who I'm pretty sure are really disappointed with my point of view on this thing. I'm having a Twitter conversation before I got on this podcast, where people were like, "Look the venture destroys these companies. It's turns their employees into meat." People have a lot of real feelings about this stuff. And I just... The thing that is most common, it's not universal, but is most common, is that the folks who feel that way, haven't been the people who have to make the decisions about what their business does or doesn't do in order to live and die.

Do you know what I mean? And I remember what it was like before I had had that experience and what I thought it would be like, and what I thought my day to day decision making would be and what I thought it would be like to be in board meeting. Or what I thought it would be like to have investors, and I just was wrong. And most of the time I really do think, in the large, that shape of seeing venture capital as bad and venture investment as bad... I'm not saying it's always good, it's not. And there are situations where people do it badly or they're bad actors, or they take too... You know what I mean?

There's a million ways it can go wrong. And also at its most fundamental it's a pretty, pretty good thing, right? And the number of people whose livelihood exist where, it's one thing to say you'd rather that a company didn't take venture money, but if I'm one of the 150 people at 1Password, if they were getting equity all along, which is unclear that they were, right? But let's assume that they have some now, they have a much better window into being liquid than they had yesterday. Now, maybe that turns out, maybe it doesn't, but yes, it's a lottery ticket, and also, before yesterday they didn't have a lottery ticket.

Do you know what I'm saying?

Corey: I do.

Adam: And I think the... Especially, when you think about it from a perspective of trying to start a business or run a business, it is a fucking miracle that venture capital exists in the way that it does. That you can raise money, and look, some people have an easier time raising money than others, there's a bunch of systemic problems. We don't give nearly enough money to women lead startups or people of color. There's a million problems, but fundamentally the mechanism is that people give you money with no expectation that that money will ever see the light of day again, other than that you try your best to do it. That is magic. Right? That's crazy, but that's precisely what happens.

And so when you're someone like 1Password, I think look it's not the end for 1Password, so they still have to convert into a larger business that can be liquid, and be in a public market, but man, I am thankful that that's a thing they get to do. I'm thankful that that's a thing that the people who built that thing that I love can take, in the hopes that it gets them what it is that they desire. And I don't agree and don't think that the folks who look at that and go, "It's this horrifically corrupting influence that ruined the universe." I don't buy it. Facebook when it was under venture capital control and not Mark Zuckerberg pure control, not that it ever was, but a lot of the things you hate happened when it was in public markets, not private ones. Right?

Corey: Yeah, you're right. I think there is a easy, lazy shortcut to just blaming, "Oh, whatever this market model is is obviously the root of all things that I hate." Because the alternative is blaming people who built something you love for what you perceive to be a poor decision.

Adam: That's right.

Corey: On something else that feels more personal. In this case it's-

Adam: And there's people who like-

Corey: VC is faceless to me.

Adam: Yeah, exactly. And they experience of it is faceless. You've never dealt with venture capital, you haven't sat in a room, you haven't negotiated those things. You haven't lived with ...

Corey: Yeah, it turns out that it's super hard to do a raise when, "So what does your company do?" "Well, we have a cartoon platypus, we fixate over bills-

Adam: We make fun of AWS.

Corey: ... and I shit... Yeah, I shit post on Twitter, what about it?" And suddenly the phone-

Adam: I would argue that it's-

Corey: ... is talking to dead air.

Adam: That's venture capital at work, that you didn't A-raise.

Corey: Exactly.

Adam: It's good you didn't raise. That would have been a mistake. That would have been bad management of your business Corey Quinn, and if you had asked me, I'd have told you, "Do not raise money on that. That's a terrible idea."

Corey: Excellent.

Adam: But, yeah, I really think that there's a lot more nuance and more, to me, I worry more about founder cults than I worry about venture capital. I know my own mind day in and day out. I know what decisions I have made. I know what my positions have been on all sorts of issues that the world at large doesn't know and never will. And also I know there are a lot of people who put a lot of their life and opinion and caring into me as a human being, sort of. Like the totem version of me. Do you know what I'm saying?

Corey: I do.

Adam: And I'm so appreciative of that, there's a lot of really excellent, lovely, beautiful things that that has brought into my life that I love, and I'm so, so thankful for. And also the totem version of me is not me, and... Does that make sense? It's not the same. The same with you, there's a totem version of Corey that people love, and if you stay that person, then they're happy, and if you ever do anything that doesn't map to their totemic version of Corey, then they're not going to be happy with you. And-

Corey: My personal favorite is when I get accused of being too hard on AWS and of being an AWS shield, in response to the exact same tweet.

Adam: Right, at the same time. Yeah, that's right. And one of the things that I'm most proud of is I think I did a good job and continue to do a good job of building actual real businesses that take in money, revenue and solve problems for customers and makes people happy. I would rather that when people think about what it is that is good about the companies you love or bad about the companies you love, we should be having conversations less about where their money comes from and more about how they're being managed. Because we can actually have a reasonable chance of changing people’s direction, if it makes sense to change. Do you know what I mean?

Because most of the decisions that you make as an executive, are pure judgment calls. They're not... Very rarely, are they decisions where you are the sole person who had the depth of knowledge to make a choice. Instead, it's that nobody around you could make the choice. Do you know what I mean? If nobody else could make that choice, then it winds up in your lap. And so by the time it gets to that place, it's usually because the pros and cons pretty much met out on any given choice, and you have to just make one. And it has to be what you hope is the right decision. And, hopefully, you are a good leader who is willing to reevaluate your position when you get new information.

But maybe not.

Corey: If people want to learn more about what you have to say about this and other topics, where can they find you?

Adam: They should just follow me on Twitter, I suppose. I'm AdamHJK. I also talk a lot about open source and sustainability stuff. I wrote a little, it's not a book, but it's longer than an essay, and you can read that at sfosc.org, sustainable, free, and open source communities.

Corey: Excellent. Thank you so much for taking the time to speak with me in a format that extended beyond 280 characters.

Adam: It's my pleasure. You're a delight.

Corey: I really am, but so are you. Adam Jacob, co-founder of Chef and creator of Chef as well, now the Chief Executive Officer and co-founder of the System Initiative and internet gadfly. I'm Cloud economist Corey Quinn, and this is Screaming in the Cloud.

Announcer: This has been this week's episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

Links Referenced

  • Oxide Website
  • On The Metal Podcast

Transcript

Announcer: Hello and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on this state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey Quinn: And this episode is sponsored by InfluxData. Influx is most well known for InfluxDB, which is a time-series database that you use if you need a time-series database. Think Amazon TimeStream except actually available for sale and has paying customers. To check out what they're doing both with their SAS offering as well as their on-premise offerings that you can use yourself because they're open source, visit influxdata.com. My thanks to them for sponsoring this ridiculous podcast.

Welcome to Screaming in the Cloud, I'm Corey Quinn. I'm joined this week by not one guest, but three, because you go ahead and get Jessie Frazelle, Bryan Cantrill, and Steve Tuck together in a room and then telling them no, they can't all speak. Together, they're Oxide Computer. Welcome to the show folks.

Bryan Cantrill: Thanks for having us.

Steve Tuck: Glad to be here.

Jessie Frazelle: Yeah. Super exciting.

Corey Quinn: Let's start at the beginning. You're called Oxide Computer, so obviously you focus on Rust. What does the company do?

Bryan Cantrill: It is a bit of a tip of the hat to Rust actually. I am falling. I mean I did fall in love with the Rust but it's not weird. It's not only about Rust. So just you want to-

Corey Quinn: No, at some point, you've got to stop talking about work and actually do work.

Bryan Cantrill: That is true.

Jessie Frazelle: Yeah. I guess our tagline is a Hyperscaler infrastructure for everyone else. If you think about how Amazon, Google, Microsoft, the hyperscalers, built their internal infrastructure, we are trying to do that for everyone else with hyper, with like a rack scale server based off open compute project and going all the way up with the software layer to deploying VMs. You can basically plug in a rack and then you are able to deploy VMs within like a few minutes of it booting. It's the goal.

Corey Quinn: I'm sure you've been asked this question once or twice, but I say the quiet part out loud. Why in the world would you build a hardware computer startup in the time of Cloud? Did no one hug any of you enough as children, perhaps your parents were cousins? What inspired you to do this in the year of our Lord 2019.

Jessie Frazelle: Well, a lot of people still honestly run on-premise. If you could talk to companies really like a lot of where the hype is with containers and Kubernetes and a lot of other things like that, that's super, super forward facing. A lot of people run on-premises for good reasons. Either strategic or the unit cost economics of them actually running in the Cloud, it's too expensive. But if you're a finance company and you are hyper currency trading, you need good latency or you really care about the security of your infrastructure because you saw it with the Capitol One breach that a bank running in the Cloud. Like it's actually super horrifying if they get popped.

I think that there are really, really good reasons for running in the Cloud and a lot of that market has been neglected for so long. I spent a lot of my fun employment talking to a lot of these companies and honestly like they're in a great deal of pain and we even saw it during the raise that people don't seem to believe that they exist but they do. We're here to disrupt that and build them something that is actually nice.

Bryan Cantrill: Corey, this is the life that Steve and I lived. We worked at a Cloud computing company actually. We're very pro-Cloud just to be clear. We love the Cloud, but we also know that there are economic reasons in particular that you don't want to, you may not want to actually rent your compute. You actually may want to buy a computer or two as it turns out if your Cloud bill is high enough.

Corey Quinn: Well, looking at the Cloud bills of large companies, there's a distinct break from what they talk about on stage at re:Invent where, "Oh, we're going to talk about these machine learning algorithms and all of these container things and Serverless is the future." And you look at the bill and it's all primarily just a giant pile of EC2 instances. The stuff that they talk about on stage is not the stuff that they're actually spending the money on. I'm curious as much as we hear about Cloud, is there still a sizable market running data centers?

Steve Tuck: Yeah, there is. I think your question is one that we've heard quite a bit in the market of wasn't everything going to the Cloud and having run lots of infrastructure ourselves on-premises and then talking to a lot of the enterprise market. There is still an enormous amount of infrastructure that is being ran and that they are expecting to run for years and years and years. Frustratingly for ourselves, for them in the past, that market has been neglected and there has been this false dichotomy that if I'm going to run on-premises, I can't have the kind of utility or elasticity or just ease for developers that I would have in the public Cloud. You're right, 80%, 90% of the bulk of usage for folks today is more like easy too. Let me be able to easily spin up and provision a VM via an API of various sizes with various software running on it and store some data from that and connect to that over the network.

Bryan Cantrill: Corey, when was the last time you bought a box? Like a to you box or whatever and stood it up?

Corey Quinn: Oh, probably 2012.

Bryan Cantrill: Yeah, right. We're here to tell you that the state of the art hasn't changed that much since 2012 or since-

Steve Tuck: ... Or 2008.

Bryan Cantrill: Yeah, right. Exactly. It's basically still a personal computer. If you are in the unfortunate position where you are actually Cloud aware or you know what the Cloud is and then you have to, for one of these reasons that just mentions: economic reasons or strategic reasons or so on, actually stand up your own Cloud because you actually want to own your own computers. It's very dismaying to discover that the state of the art has advanced so little and then really heightening the level of pain, Google and Facebook and others make it very clear the infrastructure that they are on and it's not the infrastructure that you just bought. In fact, it's much, much, much better and there's no way you can buy it. This is the kind of the fundamental issue. If you want to buy the Facebook's open compute server, which is, you've got like Tioga Pass or Bryce Canyon, these are, it's a nice box. Or if you want to buy Google's warehouse size, the data center size computer, they're not for sale. If you're a customer who's buying your own machines, this is really frustrating.

Steve Tuck: Yeah. For six or seven years, there's been innovation going into these hyperscalers infrastructure to deliver out the services we all consume, Amazon, Google, Microsoft, Facebook, et cetera. And none of that innovation has made its way into the larger enterprise market, which is frustrating.

Bryan Cantrill: And then, you're told they don't exist and they go into outrage.

Steve Tuck: ... Because it's all going into the public Cloud.

Jessie Frazelle: Yeah.

Steve Tuck: So what does it matter?

Corey Quinn: Yeah. If I go and buy a whole bunch of servers to put in a cage somewhere at a data center, you're right, my answer hasn't really changed in years. I would call up Dell or HP or Supermicro, I would be angry and upset by regardless of which direction I went in an no one's going to be super happy. Nothing's going to show up on time and what I get is going to be mostly the same except for those one or two edge case boxes that don't work the same way that slowly caused me to rip the remains in my hair out and eventually, I presumably have a data center up and running.

Cloud wasn't, for me, at least this thing that was amazing and solved all of the problems about instant on and on-demand compute. But, it got me out of the data center where I didn't have to deal with things that did nothing other than pissed me off. I guess my first question for you then becomes what makes you different than the rest of those, shall we say, legacy vendors? And then, I want to talk a little bit about how much of what the hyperscalers have done is applicable if you're not building out rows and rows and rows and rows of racks.

Jessie Frazelle: I think honestly, what makes us different is the fact that we have experience with Cloud compute and we also have experience in making tools that developers love. We Also have the experience with Bryan and Steve of running on-prem and all that pain. What you get, like you were saying when you buy like a few servers, is you get a kit car and so what we are giving people is not a kit car. It's like the Death Star all built together.

Bryan Cantrill: ... Is that a Margaret literature where the sun?

Steve Tuck: We have to update that slide.

Bryan Cantrill: I know Death Star. It feels a bit like what's Alderaan in that?

Jessie Frazelle: I know. I should-

Bryan Cantrill: ... We should go to that one.

Jessie Frazelle: ... I should have used that in the raise honestly.

Corey Quinn: You'll get attention.

Bryan Cantrill: Actually, that was not the point I was trying to make. Actually. It should be said Corey that I'm not sure that Jess believed how bad the existing vendors were. I think she heard Steve and me complain but then she acted too surprised when she talked to customers.

Jessie Frazelle: Yeah. No, this is very true because they did complain and then I was like, "Okay, I'm going to go talk to some people." And so, I spent a lot of time actually tracking down a bunch of fuel and getting on the phone with them. Honestly, then I was like, "Whoa, Dell is really bad. I didn't realize that everyone has the same exact problem." Not only that, you get it by these vendors into thinking this problem is only yours.

Corey Quinn: You bought the wrong rails. No one has ever done that before. You must be simple.

Bryan Cantrill: We've never seen this problem before. You're the only ones to report this problem. I cannot tell you how deeply to the marrow frustrating it is. Honestly, this is... Corey, you made this point too about AWS and the Cloud taking this pain away. That is one thing that AWS really does not do. They do not try to tell a customer that they are the only ones seeing a problem. Logic is they can't get away with that. Right? They can't get away with saying no. Frankfurt is only down for you.

Corey Quinn: Yeah, Twitter would disagree.

Bryan Cantrill: Exactly. But, you don't have a way of going to Twitter and understanding that "Hey, is anyone else seeing a level of dim failure that that seems to be too high to take an example?" We were seeing a problem guys. Steven was at that in 2011 when we were seeing a pandemic of reliability issues. We were being told we're the only ones that are seeing this problem and it's a problem that we couldn't possibly create so it made no sense.

I asked a room full of, I don't know how many were in there, like 300 people, "Could you raise your hand if you've ever seen a parody error on your rate controller?" All of a sudden, 20 or 30 angry hands shot up. It wasn't very many people, but it was we were not the only ones-

Steve Tuck: ... All these people discovering they weren't the only one. They all have the same message.

Bryan Cantrill: They were not the only ones. I was looking around and I'm seeing like, "Wait a minute." My hand isn't the only one that's up and all of a sudden you have like an instant therapy group. That shows you the level of frustration and that level of frustration has only grown in the last decade.

Steve Tuck: A big part of the frustration stems from the disconnected hardware and software. To your example, Corey, you go buy a bunch of hardware and if you can get it all uniform and get it stood up and racked and tester drives and now you have a collection of hardware, but yet now you have to go add software.

Bryan Cantrill: But Steve, we can just run Kubernetes on bare metal, right, Jess?

Steve Tuck: Stop it.

Jessie Frazelle: Don't even. Do not. That was not the third rail that you're supposed to touch by the way.

Bryan Cantrill: I love that third rail. That's one of my favorite third rails to touch.

Steve Tuck: But, tightly integrating hardware and software together is going to be one key differentiation point for us, Corey. And owning both the hardware and the software side so the customer doesn't have to so they can effectively roll a rack in, plug it in and expose the APIs to the developers that they love and operational tools that they love. This is the kind of stuff you would have expected to come out of hardware 20 years ago, 10 years ago and so we're finally going to do it.

Corey Quinn: I will say that what you're saying resonates as far as people having anger about what's going on in this, I guess, world of data centers. I gave a talk at scale beginning of 2019 called The Cloud is a Scam. Apparently, someone ripped it off and put it up on YouTube a few weeks ago and I just glanced that as 43,000 views and a bunch of terrible comments and someone's monetizing it who is not me. Great, awesome. At least the message is getting out there. But the goal, the entire talk is in five minutes, is just a litany of failures I had while setting up a data center series in one year at an old job. Every single one of those stories happened and people resonated. Each one resonates with people. The picture of the four different sizes of rack nuts that look exactly the same, that are never actually compatible with one another, a Schrodinger's GBIC where you send the wrong one every time and there's no way out of it. Just going down this litany of terrible experiences that everyone can relate to.

Steve Tuck: Was that GBIC a Sun GBIC by the way?

Corey Quinn: No, I think it was a Cisco GBIC at that point, just because we all have our own personal scars.

Steve Tuck: Serious GBIC problem. That was a century back.

Corey Quinn: Yeah, and the fact that it resonates as strongly as it does. It's fascinating.

Bryan Cantrill: Yeah. In terms of what a kit car, the Exton data centers, It's not good. The problem is that you're left buying from all these different vendors. As Steve says, you've got this hardware/software divide and what we think people are begging for is an iPhone in their DC. They're begging for that fully integrated, vertically integrated experience where if there is a failure, I know the vendor who is going to actually go proactively deal with it. Just as you've got that at Google and Facebook and so on. A group that I think Jess and I should be coined or not, but I love the term infrastructure privilege, which those hyperscalers definitely have and we want to give that infrastructure privilege to everybody.

Corey Quinn: Well, that does lead to the next part of the question, which is okay, if I want to build a data center like Google for example. Well, step one, I'm going to have an interview series that is as condescending and insulting to people as humanly possible. Once I wind up getting the very brightest boys, and yes they're generally boys, who can help build this thing out, super, that's just awesome. But now, most of what it feels like they do requires step one: build or buy an enormous building, step two: cooling power, et cetera, et cetera. And suddenly like step 35 is and then worry about getting individual optimizations for racks. Then going down the tier even further talking about specific individual servers, how much, what scale do you need to be operating at before some of those benefits become available to, shall we say, other companies who have old school things like business models?

Jessie Frazelle: Yeah. Honestly, in the people that I was talking to, and I won't name names, there were numerous who contemplated doing that. But, the problem is of course obviously that's not their business. Their business is not being a hyperscaler. You very rarely see companies doing it unless there is this huge gain for them economically. A lot of the value add that we're providing is giving them access to having all the resources almost of a hyperscaler without having to hire this massive team and have an entire build out of a data center for it. You can still host in a colo.

Bryan Cantrill: Yeah, I think that is that you start with a rack, right? That you can do with a single rack, especially if that rack represents both hardware and software integrated together, you can get a lot of those advantages. Yes, there are advantages absolutely when you're architecting it at an entire DC level or an entire region level. But, we think that those advantages actually can start at a much lower level.

Now, we don't think it starts much lower than a rack. If you just want a 1U or 2U server, no, that's going to be a two smallest scale. The smallest scale that we will go is to a rack. It is a rack scale design.

Steve Tuck: To your question on what is the applicability, how can people use this? Do they have to reimagine their data centers? Most modern data centers today have a sufficient power capacity to support these rack scale designs. Especially if you're a company that is using some of these rates like Equinox or Iron Mountain or others, those service providers are beginning to and have really come a long way in modernizing their facilities because they need to reduce, they need to help customers who want to reduce PUE. They need to help customers who are out of space and need to get much better density and utilization out of their space. You can take advantage of benefits on one rack and these racks can fit in many modern data centers today.

Corey Quinn: Is there a minimum size or is this one of those things where, "Oh cool, I will take one computer please and it's going to be just like my old desktop except now it sounds like a jet engine taking off every time someone annoys me"?

Bryan Cantrill: I think a rack is going to be the minimum buy.

Corey Quinn: Okay. So it needs to be a much larger desk is what I'm hearing?

Steve Tuck: Reinforce the desk.

Bryan Cantrill: Yeah, exactly. Reinforce the desk and it's going to be it'll fit on a 24 inch floor tile and it will fit in extant DC within limits, but a rack is going to be the minimum design point.

Corey Quinn: I frequently said that multicloud is a stupid best practice and I stand by that. However, if your customers are in multiple Clouds and you're a platform, you probably want to be where your customers are unless you enjoy turning down money. An example of that is InfluxData.

InfluxData are the manufacturers of InfluxDB, a time series database that you'll use if you need a time series database. Check them out at influxdb.com.

As you've been talking to, I guess, prospective customers as you went through the joy of starting up the company, did they all tend to fit into a particular profile where they across the board, for example, it feels like a two person company that is trying to launch the next coming of Twitter for pets is probably not going to necessarily be your target market. Who is?

Jessie Frazelle: I really tried to get a diversity of people who I was talking to so it's a lot of different industries and it's a lot of different sizes. There's super large enterprises there with numerous teams internally that we'll probably only interact with maybe one of those teams. There's also kind of smaller scale, not to the scale of only two people in a company, but maybe like a hundred person company that is interested as well. It's really a variety of different people. I did that like almost on purpose because you can easily get trapped into making something very specific for a very specific industry.

Bryan Cantrill: In terms of Twitter for pets, they should start in the Cloud. I would encourage anyone, any two person software startup go start in the cloud, do not start by buying your own machines. But, when you get to a certain point where VCs are now looking not just for growth but are looking for margin say, and you're kind of approaching that S1 point or you're approaching the point where you've got product market fit and now you need to optimize for cogs, now is the time to consider owning and operating your own machines. That's what we want to make easy.

Jessie Frazelle: Totally. We're not convincing anyone to move away from the Cloud for bad reasons for sure.

Steve Tuck: Well, I think everyone's going to use both. They're very few and far between companies that we've spoken to so far are not going to have a large footprint in the public Cloud for a lot of the use cases in the business that have less predictability or new apps they're deploying have seasonality, you have different usage patterns. But for those more persistent, more predictable workloads that they can see multi-year usage for, a lot of them are contemplating building their own infrastructure on-premises because there's huge economic wins and you get more control over the infrastructure.

Corey Quinn: There's a lot of opportunity to, I guess, build baseline workloads out in data centers, but the whole beautiful part of the Cloud is I can just pay for whatever I use and it can vary all the time. Now, excuse me, I have to spend the next three years planning out my compute usage so I can buy reserved instances or savings plans now that are perfectly aligned so I don't wind up overpaying. There starts to be an idea towards maybe there is an on-demand story someday for stuff like this, but I'm not sure the current configuration of the Cloud is really the economical way to get there.

Steve Tuck: Yeah, it is still much more the hotel model. If you're going to stay in the city for a week, a hotel makes sense. But as you get to a month and then three months and then a year, it's pretty difficult to get to the right unit economics that are necessary for some of these persistent workloads. I think bandwidth is part of it. There's just between compute storage and bandwidth, when used in a persistent fashion, are pretty expensive.

Now, if those work with one's business, I think when there's not an economic challenge or a need for control or being able to modify the underlying infrastructure, the Cloud is great. It's just we're hearing from more and more companies where it is a pretty big spread for that infrastructure that they're running. Even with reserved instances, it's tough to put a three to five year value on it.

Bryan Cantrill: We've been convinced that this post Cloud SAS workload is coming. In part because Steve and I lived it in our previous lives, so we've been convinced it's coming economically. I think it's a little bit surprising the degree to which it's arriving. We've had increase increasingly been surprised by conversations that we've had with folks who are like, all right, well, this company is going to be, they're going to die in the Cloud. They would never contemplate going to the Cloud.

And then, we learned that... I know actually we've got a high level mandate to get off the Cloud by 2021 or what have you.

Steve Tuck: Or move X percent off.

Bryan Cantrill: Or move X percent off or what have you. It's surprising. I think it shows that they've got high level mandates to either or to move to that hybrid model that Steve described. Especially when you look at things like bandwidth. Bandwidth by the way, as when that Steve talked a bingo card for this little panel.

Steve Tuck: Oh.

Bryan Cantrill: Yeah, I know there's no way you're going to let Corey get through an entire episode without mentioning bandwidth costs because bandwidth costs are outrageous on AWS.

Corey Quinn: The problem isn't even that they're expensive because you can at least make an argument in favor of being expensive. The problem I have is that they are incom-freaking-hensible as far as being able to understand what it's going to cost in advance. The way you find out is the somewhat titillatingly called suck-it-and-see method. Namely, how do you figure out if a power cable in a data center is live? Well, it's dangerous to touch it. So you hand it to an intern to suck on the end of the cable and if they don't get blown backwards through the rack, it's probably fine. That's what you're doing except with the bill. If your bill doesn't explode and take your company out of business, okay, maybe that data transfer patterns okay this month.

Bryan Cantrill: That's it and I think that I've always said that, you know who loves utility billing models? Utilities. People don't love it. It's like in a... Corey, I know you've got teenagers, but-

Corey Quinn: She's two at the moment so it's almost a three nature though. We'll see.

Bryan Cantrill: ... but okay, even like you know-

Steve Tuck: Oh, four, five.

Bryan Cantrill: Oh, bandwidth overages do not start at too young at age at this point.

Steve Tuck: Yeah. Five years old [crosstalk 00:24:14].

Bryan Cantrill: Five years old and like, you're scared as hell that they're going to be on the cellular network when they-

Steve Tuck: Watching YouTube channels.

Bryan Cantrill: ... watching, exactly. Being on the Cloud is like having thousands of teenagers that are on YouTube videos all the time and you're praying that they're all on the wifi. They all claim they are but of course half of them are accidentally on the cellular.

Corey Quinn: As you were talking to a variety of VC types, presumably, as you were getting the funds together to launch and come out of stealth mode, what questions did you wind up seeing coming up again and again? They can be good questions, they can be bad questions, they can be hilarious questions. I leave it to you.

Steve Tuck: What a spectrum of questions.

Jessie Frazelle: I actually think what was most fascinating is that nobody really asked the same questions. A lot of people kind of had differing opinions almost so much so that if you were to combine all the VCs we talked to, they would all not align.

Bryan Cantrill: Right. You get one kind of super Voltron VC that understands everything and then one just total idiot VC that understands nothing.

Corey Quinn: The thing that was-

Steve Tuck: And everything in the middle.

Corey Quinn: Everything in the middle. The thing that was also interesting and I don't know if you two felt the same thing, but I felt it was the thing that people would often have the most angst about or the most questions for us about are the things they understood the least.

Jessie Frazelle: Yes. There were other VCs who they understood those things the most and then all the other things they were like, wait, but what about? So if you combine them together, you would almost get our ideal scenario where they understand everything or you would get the worst scenario where they understand absolutely nothing.

Steve Tuck: And ask the most question.

Corey Quinn: Or the best questions that is not entirely clear which side of that very large spectrum they're on. Well, what about magnets? You're sitting there trying to figure out are they actually having a legitimate concern about some weird magnetic issue that you hadn't considered or did they have no actual idea how computers work or is it something smack dab in the freaking middle?

Steve Tuck: Or is it employee to try and throw you off?

Bryan Cantrill: We've got to talk with some of the dumbest questions we got.

Jessie Frazelle: Okay. Can we do?

Bryan Cantrill: Yes.

Jessie Frazelle: Okay. There was one question where we were trying to raise enough money to hire enough people to go build this basically. A lot of the questions that would come up would be like, "What can you do with like-

Steve Tuck: A lot less money?"

Jessie Frazelle: A seed thing. Just a seed thing, something like $1 million.

Corey Quinn: Fail mostly.

Jessie Frazelle: One rack, can you build them?

Steve Tuck: Get some proof points and get off.

Jessie Frazelle: One rack?

Bryan Cantrill: Right. There's one VC in particular and I had said, "Can you shrink the scope of the problem a bit so that you can take less investment, prove it out and then scale the business? What if you were to say shrink this down to like three racks?"

Jessie Frazelle: Three racks.

Bryan Cantrill: If you just did the three racks, would that then prove out things instead of going big and building a lot of racks, what if you just built three?

Steve Tuck: Corey, to clarify, most of what we're building is software so that first rack is actually the expensive one.

Bryan Cantrill: ... 90% of the time.

Steve Tuck: The second rack is a lot cheaper.

Jessie Frazelle: Eventually you caused the chasm and you have a rack and then you can go sell that rack to everyone.

Corey Quinn: And then, you go pitch to SoftBank on the other hand like, "Okay, we like your perspective but what can you do with $4 billion?" And the only acceptable answer that gets the money is something monstrous.

Bryan Cantrill: You know what's funny is when we were first starting to raise, we're like, "What would we do if SoftBank came to us?" And then, we realized very shortly after starting our raise like, "SoftBank is not going to be coming to us or anybody else."

Steve Tuck: A long time.

Corey Quinn: They're hiding from an investor who has a bone saw.

Bryan Cantrill: Exactly. [crosstalk 00:27:53] God, that got dark well.

Jessie Frazelle: That was really dark.

Steve Tuck: That got really dark. Yeah, it was.

Corey Quinn: This is why they don't invite me back more than once for events like this. But now, it's easy to look from the outside and make fun of spectacular stumble failures and whatnot. It's also interesting to see how there are serious concerns like the one you're addressing that aren't, shall we say, the exciting things and have the potential to revolutionize society. It doesn't necessarily, and correct me if I'm wrong, sound like that's what you're aiming at. It sounds more like you're trying to drag, I need to build some server racks and I want to be able to do it in a way that doesn't make me actively hate my own life. So, maybe we can drag them kicking and screaming at least into the early 2000s but anything out of the '80s works.

Bryan Cantrill: Yes. Actually, I think we actually are changing and do view ourselves as changing things a bit more broadly and that we are bringing innovations that are really important, clear innovations that the hyperscalers have developed for themselves and have actually very charitably made available for others.

The open compute project that was originally initiated by Facebook was really an attempt to bring these innovations to a much broader demographic. It didn't exactly, or hasn't yet, I should say, hit that broader demographic. It's really been confined to the hyperscalers. These advantages are really important. It was really interesting, we were talking to Amir Michael who we had on our podcast and had a just a fascinating conversation with him about. Amir was the engineer at Facebook who led the open compute project. A big part of the motivation for Amir was the ecological efficiency that you get from designing these larger scale machines, rack scale designs and driving the PUE down. That was a real motivator for him. A really deep earnest motivator and an earnest motivator for Facebook and the OCP was allowing other people to appreciate that.

I don't think it's too ridiculous to say that as we have a changing planet or much more mindful about the way we consume energy. Actually delivering these more efficient designs to people, it's not just about giving them a better technology but a more efficient one as well.

Jessie Frazelle: Yeah, the power savings are actually huge. It's really good.

Steve Tuck: Yeah. Corey, I know you had someone on your podcast talking about comparing Cloud providers and who are focusing on sustainability both in terms of utilization and also offsetting. You've got large enterprise data centers that are even further behind some of the hyperscalers that aren't, aren't maybe scoring as highly against a Google or another. This gives them the ability not only to increase density and get much a smaller footprint in either their colo or in their data center, but also get a lot of these power efficiency savings. We don't want it easier for people to go build racks. We actually want to make it easy for folks to just snap new racks in and have usable capacity that is easy to manage.

Bryan Cantrill: Also, by crossing that hardware/software divide, give people insight into their utilization and allow people to make the right level of purchase. Even if that means buying less stuff from us next year because we know that's going to be a lifelong relationship for us.

Corey Quinn: What is the story, if you can tell me this now, please feel free to tell me you can't, around hardware refreshes? One of the challenges of course is not only at this technology get faster, better, cheaper, et cetera, but it also gets more energy efficient, which from a climate and sustainability perspective is incredibly important. What is the narrative here? Frankly, one of the reasons I like even renting things rather than buying them in forms of a cell phone purchase plan is because I don't have to worry about getting rid of the old explodey thing. What is the story with Oxide around upgrading and sustainability?

Steve Tuck: Well, first, I think is you got to take it from a couple angles. From when should one replace infrastructure, this is something that hardware manufacturers do a very poor job of providing information for one to make that decision. What is the health of my current infrastructure? What is the capacity of my current infrastructure? Basics that should just come out of the box that help you make those decisions. Should I buy new equipment for three years, four years, five years? What am I warranty schedules look like? Being a former operator of thousands and thousands of machines, one of the things that was most frustrating was that the vendors I bought from seem to treat those 8,000 machines as 8,000 separate instances, 8,000 separate pieces of hardware that I was doing something with and no cohesive view into that.

Steve Tuck: So number one, make it easier for customers to look at the data and make those decisions. The other element of it is that your utilization is more environmentally impacting in many cases than the length of time or the efficiency of the box itself. How can I be smarter about how I am putting that infrastructure to work? If I'm only using 30% of my capacity but it's all plugged in drawing power all the time, it's extraordinarily wasteful and so I think there's a whole lot that can be gained in terms of efficiency of infrastructure one has. And then yes, there's also the question of when is the right time to buy hardware based on improvements in the underlying components?

Bryan Cantrill: Yeah. I think the other thing that's happening of course is that Moore's law is slowing down. What is the actual lifetime of a CPU, for example? How long can a CPU run? We know this the bad way from having machines that were in production long after they should have been taken out of production for various reasons. We know that there's no uptake in mortality even after CPUs had been in production for a decade. At the end of Moore's law, should we be ripping and replacing a CPU every three years? Probably not. We can expect our RDRAM down state a level. We can expect our... Obviously our CPU clock frequencies have already leveled, but our transistor that season, the CPUs are going to level.

And then so, how do you then have a surround that's designed to make that thing run longer? When it does need to be replaced, you want to be sure you're replacing the right thing and you want to make sure you've got a modular architecture and OCP has got a terrific rack design that allows for individual sleds to be replaced. Certainly, we're going to optimize that as well. But, we're also taking the opportunity to rethink the problem.

Corey Quinn: When people look at what you're building and come at it from a relatively naive, or shall we say, Cloud-first perspective, which let's face it, I'm old enough to admit it, those are the same thing. What do you think-

Bryan Cantrill: Can I first do the nationalist movement though. That makes me feel kind of uncomfortable honestly.

Corey Quinn: It really does feel somewhat nationalists and then we talk about Cloud native and oh, that goes nowhere good.

Bryan Cantrill: ... Oh.

Steve Tuck: Oh, God.

Corey Quinn: I'm not apologizing for that. But, what are people-

Steve Tuck: There are more Cloud nativist is what I understand.

Corey Quinn: Exactly. What are people missing when they're coming from that perspective looking at what you're building?

Bryan Cantrill: I don't think we're going to try to talk people out of that perspective. That's fine. We're actually going to go to the folks that are already in pain, which Jess knows many.

Jessie Frazelle: Yes. Definitely a lot of people in pain and also we do understand the Cloud and the usability that comes the interfaces there. I also think that we can innovate on them, but yeah, I think that we're not opposed.

Bryan Cantrill: Well, just like someone who's using Lambda doesn't necessarily need to be educated about, well actually it's not sort of a lesson, there's something you're running in. Or someone who's running containers, it doesn't necessarily have to be educated about actually a hardware virtual machine and what that means. Someone's running a hardware virtual machine doesn't need to necessarily be educated about what an actual bare metal machine looks like. We don't feel an obligation to force this upon people for whom this is not a good fit.

Corey Quinn: If they don't want to learn more about the exciting things that Oxide Computer is up to, now that your post stealth, where can they learn more about you?

Jessie Frazelle: Head on over to our website oxide.computer and also, we have our own podcast called On The Metal. It's tales from the hardware/software interface, so you can subscribe to that as well.

Steve Tuck: Kind to say in Corey, we obviously, we love your podcast. Our podcast is awesome. It is so good.

Jessie Frazelle: Don't start a competition right now.

Corey Quinn: ... No, it's not a competition.

Steve Tuck: Come on.

Corey Quinn: It's like we can both be-

Steve Tuck: We aspire to-

Corey Quinn: No, we can be... We're different. We're not in the... One is Screaming at the Cloud. The other is tales from the hardware/software interface. These are very different-

Steve Tuck: They do talking two different domains.

Jessie Frazelle: And to be fair, we copied Corey on his entire podcast set up.

Corey Quinn: Oh, it's absolutely fine. I've made terrible decisions in the hopes that other people will sabotage themselves by repeating them. Deal it.

Bryan Cantrill: Well, one step ahead.

Steve Tuck: Too late.

Bryan Cantrill: I actually think, and I think we all think, that we made the podcast that we all kind of dreamed of listening to, which was folks who've done interesting work at the hardware/software interface describing some of their adventures. It's amazing. It's a good reminder that no matter how cloudy things are, the hardware/software interface still very much exists.

Steve Tuck: Yeah. Software still has to run on hardware.

Bryan Cantrill: Software still has to run on hardware.

Jessie Frazelle: Computers will always be there. Servers will always exist.

Bryan Cantrill: Corey, you surely must have done in emergency episode when HPE announced their Cloudless initiative. I assume that you had a spot episode on that. I think that they would track that.

Corey Quinn: Oh, they shut that down so quickly that it wasn't even there for more than a day, which proves that cyberbullying works. It's abhorrent when you do it to an individual, but when you do it to a multibillion-dollar company like HP, it's justified and frankly everyone can feel good about the outcome. Hashtag Cloudless is now gone.

Bryan Cantrill: Yeah, and it did not last long. It feels like Microsoft Tay may have lasted even bit longer. God, it was within... It's because it's stupid. It's stupid to talk about things that are Cloudless and things that, and even Serverless, like we have to be very careful about what that means because we are still running on physical hardware. That's what that podcast is all about.

Corey Quinn: Well, even you're defining things by what they're not an Oxide is definitionally built on something that is no longer the absence of atmosphere. You're now about presence rather than absence. Good work.

Jessie Frazelle: ... Wow. Okay. That is the worst.

Steve Tuck: That was meta.

Bryan Cantrill: That was meta. We've got a lot of good reasons for naming the company Oxide. Oxides make very strong bonds. They're very... Silicone is normally found in its oxide. But, that's very meta, not thought of that one.

Corey Quinn: Oh, yeah. Wait until people start mishearing it as oxhide, something you skin off a Buffalo.

Jessie Frazelle: We'll have to cross that bridge when we come to it. But, thanks for planting the seeds.

Bryan Cantrill: Yeah. Why are we settling so poorly in the Great Plains states?

Corey Quinn: Sustainability. Thank you all so much for taking the time to speak with me today. I appreciate it.

Bryan Cantrill: ... Corey, thank for having us.

Steve Tuck: Thanks, Corey. It's great.

Corey Quinn: Likewise. The team of Oxide Computer. I'm Corey Quinn and this is Screaming in the Cloud. If you've enjoyed this episode, please leave it a five star review on iTunes. If you've hated it, please leave it a five star review on iTunes.

Announcer: This has been this week's episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com or wherever fine snark is sold.

This has been a HumblePod Production. Stay humble.

View Details

About Rachel Stephens
Rachel Stephens is an industry analyst with RedMonk, the developer focused industry analyst firm, covering a broad range of developer and infrastructure products. At RedMonk she has worked with vendors such as Amazon, Google, IBM and Microsoft.

Rachel arrived at RedMonk with a background in finance, including an MBA with a Business Intelligence specialization, along with broad exposure to a variety of enterprise database systems. Her analysis and work leverages a variety of programming and statistical modeling languages including Python and R. At RedMonk she has covered everything from Infrastructure-as-a-Service pricing patterns and trends to explorations of serverless definitions and usage.

In her free time, Rachel enjoys skiing and spending time in the mountains of Colorado, where she lives with her family.

Links Referenced

  • Twitter: @rstephensme
  • LinkedIn: https://www.linkedin.com/in/rachelstephens/
  • Company site: redmonk.com/rstephens

Transcript
Announcer: Hello and welcome to Screaming in the Cloud, with your host, Cloud Economist, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: And this episode is sponsored by InfluxData. Influx is most well known for InfluxDB, which is a time series database that you use. If you need a time series database, think Amazon Timestream except actually available for sale and has paying customers. To check out what they're doing both with their SaaS offering as well as their on-premise offerings that you can use yourself because they're open source, visit influxdata.com. My thanks to them for sponsoring this ridiculous podcast.

Welcome to screaming in the cloud. I'm Corey Quinn. I'm joined this week by Rachel Stevens, who's an industry analyst with everyone's favorite analysis firm RedMonk. Rachel, welcome to the show.

Rachel: Thanks for having me.

Corey: So let's start at the very beginning because six months ago I had no answer to this question that wasn't actively insulting. What is an analyst firm and what did they do?

Rachel: This is a great question and I've been doing this for three years and I'm still not sure I have the correct answer, but I can tell you what RedMonk does. How about that?

Corey: Perfect.

Rachel: So RedMonk, I would say we help technology companies understand how developer preferences can impact their business strategy. So sometimes it's working with them on things that are technical. A lot of the times it's working with companies on things that are decidedly not technical. So it will be things like marketing and messaging and developer personas and things like that. So it can take on a lot of different flavors depending on what the company is looking for and what help they need. But generally it's just trying to help people understand how developers drive their business.

Corey: The strangest thing that I find whenever I'm in analyst rooms or attending analyst summits at various events is people look at me and say, "Oh Corey, good to see you. Glad you're here." And my response is, "I think I'm in the wrong room. I don't think I'm an analyst." And the response is always, "Well what do you do?" It's like, well I have no idea of what I do exactly. You are an analyst is the usual punchline there. But I talk to companies that are building things. I talk to their customers that are building things and everyone smiles and nods and then where I deviate is, and then I take the service that they built, they released and I build something with it myself. And everyone stares at me for the longest time because that apparently is the one step too far where most analyst firms break off.

So it feels like everyone has their own pocket definition of analysis of being an analyst. There's a bunch of snark and sarcasm in the rest that I could hurl at you, but I won't because you know, insulting people's professions on the air, generally a bad look.

Rachel: Oh, no worries.

Corey: In the wide world of analysis, what is your specific area of expertise?

Rachel: So I come to the technology field with a fairly atypical background because I actually started my career in finance. So I started out running people's budgets and crunching numbers and typical bean counting kind of things that most developers do not resonate. It doesn't resonate with developers really at all. When I'd kind of talk about my background and then I was a DBA, which is usually like a developer enemy. So that's always fun. So usually I try to steer clear of talking about how I came into the analysis field when I'm with a developer audience because usually they don't really care about anything that I've done in my past. But I have found that it can be really helpful to have both of those backgrounds around how a company makes money and what a company does with the data. Like developers may not think they care about that, but they actually do.

Corey: It's strange. I mean, I started my company of fixing AWS bills three years ago and at the very early days I assumed, "Oh, it's an engineering problem." Turns out it's not. Engineers don't care about the bill. It's a distraction from what they got into their jobs generally to do, which is usually feature releases, building exciting things, not cool. We built this awesome thing. Now let's make it cost a lot less money as as a primary function and in my experience it seems that finance and engineering are wonderful at talking past one another. How do you find that relationship between the two groups generally tends to play out?

Rachel: I think you are 100% correct. I think they are departments that have historically not been aligned very well. I think that's changing a little bit. I think especially as the role of technology evolves and more and more businesses are seeking this, like ever elusive digital transformation. It's changing the nature of what an IT shop is actually in charge of. Because IT is no longer just a cost center. It's starting to be this like driver of value and that was maybe how engineer's always saw what they were doing. But I think that that value is now being recognized more widely across all of the departments of an organization, the realization that you really need to have a strong technology shop to figure out how you're going to compete in evolving markets.

And so I think that has changed and maybe leveled the relationship a little bit more where I think that IT has maybe more of a seat at the table than they have had in the past. And I think finances started to become more involved in how to actually have IT costs come together.

Corey: While we're talking about the joy that is the... I guess this hearkens back to Gartner's model of Bimodal IT, which has been pretty soundly debunked on a variety of different levels. But something I noticed in talking to different companies and how they structure things is that some companies view it as corporate IT. When you say IT, people hear, oh, the mouse monitor and printer people and other groups think of themselves as engineering, the things that are directly aligned with the various line of business application. I guess the most straight forward example be a SaaS company where what those engineers are building is directly aligned with the thing the company does in order to make money. The strange part is that the way that those two groups even at the same company are treated in my experience and how they're perceived are worlds apart. Does that map into any of the stuff that you're seeing in your part of the industry?

Rachel: I think definitely. I think an engineer's reluctance to be called like an IT person it's often and frequently something that the IT moniker is not something that they have adopted for themselves. It's usually more related to their title or their role or the technologies they work with. And very much they don't want... they're happy to be the person who can fix your printer when you go to visit your grandma at Thanksgiving. But they don't really want to be the person who is seen as somebody who just can fix your tickets. You don't want to be somebody who's help desk.

I just think that one of the things that makes that tricky in terms of that way that most of the organization interacts with the engineering group is that most of the interactions cross departmentally are going to be people who are filing those help desk tickets for things like, "Oh no, I can't log in, I forgot my password." So it feels like maybe some of the deeper work that is maybe more of that value add is more opaque to the people in the finance departments.

Corey: My customers tend to buy us for being SaaS companies and they also tend to largely be the board in the cloud type of companies. That said, that's certainly not all of them. And the strangest part that I've seen as I look at companies that are undergoing a, please pardon the term 'digital transformation' is getting the rest of the business on board with what that transformation tends to look like. I gave a talk about this somewhat recently at the DevOps Enterprise Summit. Where when you have a data center too that you build out to serve an entire environment. Once you've spent an enormous pile of money on building out that data center, you can amortize out what the cost of that is going to be to service your workloads for the next three years almost to the penny.

Whereas when you're paying only for consumption and as you become more and more cloud native, whatever that buzzword means to you, and you get into a serverless style of world, every time I mentioned the word serverless, I get a dollar. Then you wind up having what... what it costs to provide your service is a function of number of users that you have. And the numbers are always way less when done properly, but it's also increasingly challenging to accurately predict and that's not inherently a problem. What is the problem is how that has been communicated back to finance teams, the, I guess, partnership between finance and the engineering side of the house fundamentally has to shift. How are you seeing that play out?

Rachel: Well, I come from a finance world and so I've seen head's role in the finance department when the groups that they are in charge of budgeting missed their budget by a significant amount. Like it's a fireable offense for finance people in certain shops anyways. And so the motivations for people in the finance world can not always be aligned with this whole digital transformation goal because one of the things that they like is that stability and the predictability and the ability to budget and forecast things.

Those are core parts of what these people are tasked with doing as their careers and part of what they are measured on. And so the shift to making their job inherently harder and more opaque and less predictable can sometimes be a hard partnership to communicate about why it is that you are moving to kind of pay as you go pricing versus something that was more known and more understood before.

Corey: The hardest part I see is not that finance can't wrap their heads around this. I mean for God's sake, the entire purpose of finance is to understand the interplay of money, but rather in to be very direct engineering's lack of ability to articulate what is changing in a way that finance can understand.

Rachel: Yeah, so if you think about... Okay I'm going to throw some finance words at people and it won't get scary, I promise I'll explain as we go.

Corey: Let's hit them with that.

Rachel: All right, so if you think about five to 10 years ago, our hardware was mostly capitalized in a data center and that means that we would account for it on our balance sheet and then just pull over pieces of it every month. So we kind of just say, "All right, we're going to pay a whole bunch up front and then we're just going to count little pieces of that towards our profit or towards our costs every month." Software was licensed, same general thing where we would take the entire cost of that license fee upfront and then we would break it out into chunks and so we could really understand what our monthly approach would look like in terms of there was just more known quantities as we would accrue like that.

And to a person who has made their career on understanding the exchange of money. Like those are things that are... accruals and amortization are concepts that they've learned to ever since they were in school. Trying to understand how to price in API gateway and trying to figure out the number of API calls an app is going to make and try to do predictions about that. It's a little bit farther outside of the wheelhouse that they may have been trained in. It's not to say that they can't figure it out, it's just to say that the drivers and the cost drivers are changing and as they change, there needs to be that increase of partnership to help your business partners understand what are your actual variables that are driving usage and helping them understand how to come up with a way to think about how to actually predict and forecast things out.

Corey: Part of the challenge I think is that back when I started my company around, "Oh, I'll fix the AWS bill." That was an expensive problem that my boss would always come screaming into the office on some random date. It's was suspiciously close to the beginning of the month, freaking out about it and fix it, fix it, fix it, fix it, fix it. And I did that as a distraction from what I was normally doing because well my boss told me to, I assumed I knew what was going on. It turns out I didn't. What really happened in those scenarios was my boss was suddenly talked to by her boss and intern he was talked to by his boss.

And you trace it back to the organization and eventually it goes back to someone in finance who has zero context on what engineering is doing. They see a big Amazon bill and their first question is, "Well that sounds like a lot of boxes showing up the office that I haven't seen lately. What's the story? They having them shipped to their houses." The concern is not even from finance. Wow! We overspent over spending too much money. It's okay. It was 20% higher last month, is this the new normal? Is this going to change our projections? Is this just an aberration? Is it a mistake? Tell me the story here.

But playing the game of corporate telephone, by the time it landed on my desk, it was a, "You're spending too much fixed the bill," and it was a very different message that was received than what was transmitted partially because on the engineering side, individual contributors don't think about this, but also because I was never given context to understand the actual problem that the business was facing, which was not spend. It was predictability.

Rachel: Mm-hmm (affirmative). I agree. And I think you also touch on a great one, which is there's just so much more seasonality and variability in a pay as you go pricing model. And so a finance team trying to look at things like they care about things like year over year numbers or month over month numbers. And so trying to understand how those drivers are changing and what's impacting their costs, because they're going to have to sit down and write a variance explanation that's going to go up to the CFO. And so they're just looking for some context on what's happening here. Because all of a sudden I have to explain this huge cost increase to somebody who cares about it and I have no idea what's happening. And if somewhere in that game of telephone, like you said, it comes not so much about understanding why but the message gets morphed.

Corey: Well one of the questions I have too is, is this purely a problem for executives and managers or is this something that individual contributors need to weigh in on? I keep going back and forth on that one from the engineering side, but I don't have the experience of working in a finance department to be able to articulate anything sensible reasonably. You do. Where do you land?

Rachel: I feel like I also am not sure I have the answer there. One of the stories that I loved was from a customer who shall not be named, but they were doing a road show around all of their offices and the C-suite was kind of giving strategy day presentations to the entire company and getting everybody on board with what they were doing next. And then they took Q&A and one of the senior engineers in the front who is absolutely brilliant at building the product raises their hand and goes, "But why do we need to make money?" It's like I heard this story and my finance heart just kind of shriveled a little bit. I was like, "What do you mean? How can you be so smart and not understand that that is like a core part of the business."

But from the perspective of an engineer and you've worked in primarily growth based startups that are venture backed and all you have cared about is getting product out the door and you haven't ever had to care about business model and you really haven't ever had to worry about any of that product market fit or profit, any of those things... then yeah it can be like when you start to present these strategies of like, "Oh like why do we care about revenue?"

It can be a fair question, but it also in my experience then means probably most individual contributor engineers don't need to worry about this. That's kind of what this story tells me. You should probably have a basic understanding of what your company's business model is. Probably you don't need to get super in the weeds on the finance teams, and I think it gets more important the higher up the chain you guess. To me, I kind of view it as a manager and above collaboration area, but I'd love to hear what you think.

Corey: First. I absolutely want to say that that resonates. When I was an engineer, I hated, hated doing anything that looked like cost savings. I think that it wound up coming down to an unfortunate reminder that I was somebody's cost center and working in startups with ping pong tables, which, oh my God, are those the most useless pieces of crap in an open plan office? Why are there a ball under my desk when I'm trying to work? I digress. Tick, tick, tick, tick. But it's not consistent. It just, it drives me up a freaking wall. I don't do open plan offices right out. If you disagree with that, please feel free to go on Twitter and keep your freaking opinions to yourself. But the problem that I had was it was an unfortunate reminder that I had to generate more value than I cost.

And the fact that I wasn't able to work on a project like that while remaining comfortably removed from the fact that I'm super expensive. So ideally the thing I'm doing should be throwing off more money than that was... It was a nice affordance to live in a fantasy land and being removed from that and reminded that, "Oh yeah, there is a P&L involved here," and if you eventually cost more than the value that you deliver, you're not going to be here anymore. It was unfortunate.

The second concern that I see very often when talking with other engineers who have worked on projects like this, they talk about how they decided to spend a couple of weeks fixing the AWS bill and there are two ways that story plays out. One of them was that they wound up playing around it. Oh yeah. They'd knocked out 200 bucks a month off their bill and only two short weeks and they get dinged on their annual review for it because yeah, there was better uses of their time.

Okay. I talked to someone else, similar story and they found 10 million bucks and they were dinged on their review because there was a better use of their time. They should have been working on a different feature. And the fun thing from my perspective is that I am not convinced that those aren't the same story because I was talking to those engineers in a vacuum. Now saving 10 million bucks a year. Maybe that's valuable, maybe it's not, but it does come down to what else would you have done instead, what feature could have shipped sooner because there is a theoretical upside of how much money can I save off of someone's cloud bill of 100% of their cloud bill.

When you're talking about speeding time to market or releasing the right feature at the right time to the right folks, you could double or triple revenue, so cost savings and cost optimization is always going to take a back seat to accelerating what a company is working on from a strategic point of view. The problem I've seen is that I don't see too many people telling these stories in venues where engineers listen.

Rachel: Yeah. And I think that also just it speaks to the way that people run their performance reviews. And if you do want cost optimization, if you want those $10 million savings, which are great, you actually have to reward them. So like I think it depends on like obviously you don't want to spend two weeks of engineering time saving $200 that's suboptimal. You'd love to find the multimillion dollars of savings. That would be great. But you also do need to balance that out with what could I be building at that time. And I feel like performance reviews in particular are not always nuanced enough to capture that. And people are like, if cost savings is something you care about, it has to be something you reward. And the way you reward it is what you incentivize performance review time.

Corey: Oh, most companies seem to do a terrible job performance reviews. And now that I wind up managing people myself over here, I'm starting to understand why. These things are nuanced. There's no manual for doing it. And it's incredibly easy to say something offhand that is interpreted in a way that you didn't intend. So it's... again, from an engineering point of view too, I had a very different relationship with money than I do now that I run a business. The reason behind that is I was making decent money. Sure, don't get me wrong. And I had root access to everything in production and I could with the wrong command take down the entire thing. Remember that business we used to have, whoops, a doozy, but I needed approval from my boss to buy a $50 book that would make my life better.

So in the context of having to worry about costs like that, it seems perfectly natural that, "Oh, my development environments costing 600 bucks a month, I should really spend the time to save that $200." It doesn't actually matter to the business. It's just process and lack of context shining through. I mean, I talk about engineers making poor decisions and doing dumb things in a lot of my talks and stories. The part that I don't say is that over 80% of the time the idiot moron engineer was me. I have my own favorite punching bag when it comes to these things because I didn't see it back then and now that I do see it I think, "Oh, how could I have been so silly and naive," but I wasn't, I just didn't have context. So from that point of view, what should engineers know in order to partner effectively with finance?

Rachel: I think one of the things that's really tricky about partnering with finances, the finance team, often if, especially if you're a public company, there's going to be an inherent lack of transparency because there's a tendency to kind of hold financial results close to the chest because the more people that have them, the riskier it is that there will be a disclosure of non public information. I can make it really hard to actually share that context. And I don't know, at least nowhere that I've ever worked has I ever figured that balance out correctly. But I think it can be useful to start expressing that desire of like, "I'd really love to understand what is my group," the number of people who don't understand what their group's budget is for the year, who have no idea where they kind of sit to see year to date results of how the company overall is doing.

I think all of those things can be shared safely to some degree in a lot of cases, but it's not something that I think most businesses are inherently good or willing to do. You kind of have to ask for that collaboration. But I think that that context in terms of the numbers can be helpful. I think one of the things... so I think that's something finance should know to kind of help how to partner with engineers and give context more effectively. I think one of the things that engineers should know is just understanding those constraints that are hitting their finance team.

I guess just everyone should know. Finance should understand the importance of developer velocity and trying to understand like, oh is your team doing sprints? Do you have to understand what people are trying to build and try to understand how projects are flowing in and out just to have a general sense of what is being built and how that would be great if finance had any sense of that and any sense of importance over why the developers are trying to go quickly.

I think engineering on the other hand should understand the pressure of trying to create a public facing budget in particular. So if you have to give guidance to the street on what you're predicting for the next year, if I'm giving 2020 guidance, I probably as a company started that process the August of the year before, trying to put those numbers together and roll up everybody's individual budgets and going back and forth with all of the groups. It's a process. And so I realize it's fully ridiculous to say that you can have any sense of what's going to be happening 18 months away. Your finance team recognizes that as well, which is why it's always a fun exercise, but it's just one of those nature of the beast things that I think nobody really understands the pressures that are driving the other group.

Corey: I frequently said that multi-cloud is a stupid best practice and I stand by that. However, if your customers are in multiple clouds and you're a platform, you probably want to be where your customers are unless you enjoy turning down money. An example of that is InfluxData. InfluxData are the manufacturers of InfluxDB. A time series database that you'll use if you need a time series database, check them out at influxdb.com.

It always astonishes me when by default, this is how it's set up. Even in AWS accounts. Where I can go in and I can spin up $50,000 a month of resources with zero approvals, zero oversight, it's an API call away or let's face it. I'm terrible in many ways. I'll do it in the console. I click a button. So that's done. It's up and running, but I'm not allowed to see what the bill is. And in the console for those who do not spend their lives there, or even in the API, there is no price tag next to everything. Click this button, it'll cost you X money.

So you are incredibly far removed. So as a result, every time I hear about something I've done having financial impact, it's always A, a trailing function and B after the fact so that I'm yelled at for this thing that didn't wind up that was not surfaced or visible to me. It's ideally there's an alert or an alarm that goes off a day or two later. In practice, it's once the next month bill comes in and then someone kicks the door off the hinges.

Rachel: Yeah, I think that's something that the clouds in general struggle with. And it seems to me that both the individual developer and engineer doesn't have a great sense of what things cost and the finance team is struggling with that too. Like the lack of visibility into what is running based strictly off the tools that the cloud providers give. It's really a lot of people flying fairly blind in terms of trying to figure out what their primary cost drivers are and what's creating problems. And so I think, I would say that both groups struggle with that.

Corey: The hard part for me is how did this get fixed? I mean we talk about... Usually you wind up seeing policies and procedures in organizations that feel like scar tissue from that one time a bad thing happened in advance and now whenever that specific bad thing happened again. So you talk to large companies and well back in our data center days, it took six weeks to get a server provisioned and now that we're in the cloud, it only takes four weeks to get that same server provision. And the challenge you run into there is you're actively incentivizing terrible behavior. Shadow IT is a corporate credit card away before someone can spin something up and get working on it right now. You also will find that if you have a human being in line, that as a provisioning delay, people will bug that person every 10 minutes until they get their resources spun up, but they'll forget to turn things off.

And why would they turn things off when the alternative is to have to go back to that incredible process? Good Lord, last time it took me six weeks to do this again, I might need it again. And they just add it to the collection. Sometimes they'll run something in a loop. I've seen this where it boosts utilization metrics higher than you would expect just so when people do a quick scan, it doesn't look like those systems are sitting there idle and that's messed up. The problem is that as you build out controls, they seem to become gates rather than anything that resembles a guardrail that makes things better.

My firm belief, as I've said, a number of talks now has been that when you build guardrails, they have to be easier to do things the right way than the wrong way because otherwise you're never going to get people on board. Governance inherently has to be a trailing function in this sense, in this to some extent, and you have to accept that there's going to be a fudge factor to budgets. This is often difficult for people to hear when they come from a very hierarchical, regulated command and control style environment. But it's what I keep seeing again and again in the market. Does that map to what you're saying?

Rachel: Absolutely. I feel like you articulated that very nicely, but I think that's one of the things that needs to be communicated to the finance teams is that importance of self service and automation in processes and as a path for developers to do things the correct way and to go through all of the controls that you want to set up, you have to make sure that that self service component stays in place. I think that would be for me, the number one recommendation for making sure that your financial controls are actually effective and people don't just go around them.

Corey: And in some cases you have the luxury of, "Oh, if you wind up exceeding budgets in too far of a... in one direction or another, you'll wind up getting centered or fired, in the federal government, you exceed the budget in the wrong direction you're going to prison." So it's nice to be able to have that level of stick, but I think it's the wrong approach. I think it comes down to understanding what people should be spending on. Very often in my client engagements, it turns into less a story of, "Okay, you should do this, this, and this to save all the money." Very often it's, "Yeah, do those things, but then over here spend more," because previous cost-cutting attempts have durability to the point where, yeah, good work. You're saving a lot of money. You also have no backups. Maybe that's not the right answer.

Now for some data, that's perfectly fine. If you can reconstruct it, no problem. If it's you don't have a company anymore, if that data goes away, it's a different story. It all comes down to context and I still maintain that there's no API for any of this. There's no replacing RedMonk and there's no replacing what I do with software. Sure. Software can help an awful lot of this. But having a discussion with people and being able to understand localized context is critically important.

Rachel: I agree completely. I think there are processes that will help everyone. So like really good tagging and labeling. And that's not something you could necessarily automate but you can kind of build processes around how to do things, but eventually everything's going to come down to some level of human judgment. So I think the goal is to automate what you can and then to have reasonable policies in place to kind of guide decisions the rest of the way. And hopefully that that leads to some policies that make sense for everyone.

I think as much as possible, you want things to be open on the front side, so kind of guardrails rather than gate keeping and try to flag things via templates or solutions that people can have on the front side. Rather than coming at them on the back end after you've overshot. So I think there are some general principles, but really it's going to vary a lot based off of the organization. It'll vary based off of your capital structure. It'll vary based off of the policies you have in place, the size of your teams, how dispersed you are geographically. There's a lot of factors that can come in. So there's really just no clean cut way to put all of this together.

Corey: And there is no easy, straightforward answer. I mean, I'm somewhat disheartened by all of the SaaS companies that are springing up around, "Oh, optimize your AWS bill, save money." But that's not the real business pain. It's not something that people are focusing on as the business objective. Of the successful companies are in fact not even emphasizing the save money. They're emphasizing the story around visibility, around transparency, around being able to know that what you're spending is correct or at least allocate that to various business units. But it's step one. Step two requires conversations. It requires partnerships between engineering and finance. It requires executive stake holding and investment in understanding this brave new world we found ourselves in.

Rachel: Yeah. So question for you. When you work with organizations on their AWS bill, what part of the organization are you usually interacting with?

Corey: Great question. The short answer is yes, it usually starts at an engineering side or I'm inflicted upon someone by finance, but every time someone reaches out with our bill is too high, the question always becomes great, why? Why do you care about the bill? And very often the answer is, "My boss does," great, let's talk to your boss or, "Oh, someone over in finances yelling at us," or "I'm in finance and I don't understand how many books they're buying with that Amazon bill," or they say, "They're optimized and they swear up and down."

When the bill kept climbing and I put a restriction in place, suddenly they cut the bill in half overnight. So what do I do? Do I just set arbitrary targets? Where do we wind up collaborating? So it really is a mixed bag. Historically, it is someone who on some level has P&L responsibility for a division where the bill impacts the performance of their business group and they want to understand it if not correct it.

Rachel: Got you. Fascinating.

Corey: Yeah. I embedded the name titles just because sometimes the smaller companies that C levels, other times it's a VP where that's an important title. Other times a VP means you've been out of college for two years, here you go. Welcome to Bank of America. So there's a very different... like titles are zany across the entire spectrum of things.

Rachel: Mm-hmm (affirmative) Fair enough. No, I think that's a good descriptor. So thank you. It's always nice to hear how people are approaching things. Have you seen any patterns or anti-patterns and how finance and engineering departments have interacted?

Corey: Oh, absolutely. Screaming or we're going to allow they fire someone as an example to the rest is always a strange one. Very often though... one of the things that drives me nuts is I will see the exact same problem being solved the same way in different companies. But no one talks about it publicly. So you're solving these global problems that are not in any way shape or form even remotely resembling competitive advantage, but no one talks about it because if you talk about your bill or how your budgeting works, "Oh, that's destructive and could lead to the downfall of everything we know."

The other big anti-pattern that I see surprisingly is from finance where they'll sit down with me and one of the first questions they'll ask is, "Okay, so it costs us X dollars per monthly active user to service them. How does that benchmark across the world?" And I understand why they ask that question, but it is completely meaningless because if your customer is, let's say a user on Twitter for pets, it's going to cost you next to nothing per active user. It turns out Twitter for pets has no users because dogs don't tweet nearly as much as I thought. So the cost is surprisingly high.

Counterpoint. If you are dealing with B2B and each of your customers is a fortune 500, it could cost you millions of dollars for each monthly active user to service them. There's just no way to... It just depends. Even in the same sector, you still wind up with very different application architectures and controls. So there is no real benchmark, but I'm getting tired of telling people that. So my default position now as a thought leader is that the MAU benchmark is 32 cents. Now everyone can hate me in ops because that is completely unrealistic for almost everyone go with it anyway because I said so on a podcast.

Rachel: Oh good base. I think what that reminded me of is I feel like one of the things that we don't often talk about publicly that is true at every company I've ever worked with in terms of big enterprise organizations is that groups might not have the same understanding or business glossaries for terms. So like does your finance group have the same understanding of the user as your engineering group? Because oftentimes engineers will maybe think about roles a little bit differently. So like a user has someone who has root access and can edit things and they are in the software all the time versus like a VP who is viewing the dashboard from his iPad.

The manager who is trying to go through like the high level dashboards for her team, things like that. So I think that's another one that is tricky. So one of the places I've worked in the past didn't have a unified definition of what a transaction was and there's a whole story behind it. This could be its whole own podcast, but I think that's one that is another area that is a struggle that nobody ever talks about and it's not really one that you can globalize because everyone is going to have a unique struggle, unique to their business. But I do think that it's a problem for teams to just be on basically the same page, let alone doing this deep collaboration.

Corey: It all comes down to people problems. From my perspective, those are the only interesting ones because making a computer do the thing you want, it'll work. It won't work. Eventually you throw it out the window in a fit of rage, but you can't do that with people. Well, not more than once anyway, and I think that that's the interesting part where regardless of how often I see the same things in bills from company to company, the dynamics are always slightly different. The stories and why people care always varies slightly. I get bored working on computers all day. It's nice to be able to go out and talk to humans.

Rachel: Yeah, I feel like there's a lot of engineers who bristle at that a little bit. I'm like, "Oh, you try doing all of this impossibly difficult things that I am making the computer do and I freely admit that I cannot do those things." You do a lot of black magic with the computers engineers and we respect that, but I do think that, I agree. When it comes down to it, it's the people, it's what are your customers trying to do and how is your business trying to actually solve those problems for them? And then how are all of the people in your business interacting to actually solve those problems? That's where it comes down to.

Corey: I started my technology career more or less at in higher ed, running the network for a school. I have an eighth grade education, but dealing with PhDs and faculty with various networking issues was fascinating. Where it's like, I have a PhD, I am the world's leading expert in this very narrowly defined field and that is an incredibly hard field and I am incredibly gifted at this. Therefore, I'm very good at everything else too. And you see that pattern come up with engineers too, they don't even offer a PhD in networking? Well, back then they did. It was called Cisco CCIE Certification, but I digress. It wasn't at all accurate, but we see people who write code for living very often diminishing folks who do not code.

But I think the list of problems that people can solve by writing code is a very small subset of the world we live in. There's so much more out there beyond that and finding ways to apply skills like writing code to other problems is the hard part and I think it's the right answer.

Rachel: Agreed. I think your next podcast needs to be, can technology save us?

Corey: That is a wonderful question. If people want to hear more about what you have to say, where can they find you?

Rachel: You can find me on Twitter at rstephensme, R-S-T-E-P-H-E-N-S-M-E.

Corey: Excellent. And we will throw a link to that in the show notes. Rachel Stevens analyst at RedMonk. Thank you so much for taking the time to join me.

Rachel: Corey Quinn, Analyst at Duckbill Group. Thank you so much for having me.

Corey: Please. I'm a Cloud Economist. At least no one knows what that is. I am Corey Quinn Cloud Economist at The Duckbill Group. This has been another episode of Screaming in the Cloud. If you've enjoyed it, please leave a five star review on Apple Podcasts. If you've hated it, please leave a five star review on Apple Podcasts.

This has been this week's episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

Announcer: This has been a Humble Pod Production. Stay humble.

View Details

About Pete Cheslock
Pete is Professionally Unaffiliated, but spends his time consulting and advising companies such as CHAOSSEARCH and CloudTruth.

Previous he was the VP of Product for CHAOSSEARCH, and before that he has been running large scale infrastructure on Amazon Web Services since 2009

Links Referenced

  • re:Invent Expo Nature Walk Twitter Thread
  • Twitter Username: @petecheslock
  • LinkedIn URL: https://www.linkedin.com/in/petecheslock/
  • Personal site: https://pete.wtf
  • Company site: https://pete.wtf
  • CHAOSSEARCH

Transcript
Announcer
: Hello, and welcome to Screaming in the Cloud with your host, cloud economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Pete Cheslock. Pete, welcome to the show.

Pete: Hello. It's awesome to be here yet again.

Corey: Oh yes. This is now a tradition annually where we do a reinvent, recap. And conveniently, this episode is sponsored by CHAOSSEARCH, a company that you until recently were the VP of product of.

Pete: Yeah, CHAOSSEARCH. I fixed my caps lock key by moving on from CHAOSSEARCH. I'm just an advisor, consultant, helping them out on a couple of projects, and very happy fan of the company.

Corey: Which is how I started, and then they started paying me to sponsor things, so now it's cool. It's always nice when you're paid to promote something you actually like. CHAOSSEARCH, for those who are not aware or are crawling out of hibernation, which you're a few months early for, is Elasticsearch except it doesn't suck, which is not what they say in their branding because they have to be professional, whereas I can just be honest and funny at the same time.

Pete: CHAOSSEARCH, I love the product. I love the world that they're working in. We talked to so many customers storing months and years of data, or wanting to, and trying to do that with Elasticsearch, unless you have unlimited money of which, of course, there's one or two companies out there with that kind of money, there's really just no other way to do it. And it's a really awesome technology. I love, love what they're building.

Corey: As do I, and we will get to what they do in more depth soon. But the short answer is, of course, that they wind up separating compute from storage with an Elasticsearch compatible API, data lives in S3. Now, that seems a good jumping off point for what we saw at re:Invent. I think the dumbest name of the show award goes to something similar in that Amazon Elasticsearch is painful, annoying, et cetera. It's also expensive because it runs on burning piles of money, either yours or someone else's. They now have a lower cost storage tier called UltraWarm, apparently named after a type of personal lubricant.

Pete: I thought it was actually called LukeWarm, but UltraWarm, I think they might've missed the boat on that.

Corey: They tried that. It got a tepid response. It was an experience. Again, I feel like you have amazing engineers at AWS, amazing product visionaries who build these things out, and then they have a naming committee at the end that picks the best name in the list, and then they go drop the bottom on that list, and get it backwards and pick the name at the bottom, on the, "For God's sake, never name something this," section.

Pete: Well, I mean, I think they roll a dice. I think there is a bunch of dice with a bunch of names on it, and they toss it out there and one dice maybe flips up, says, "Ultra," and one dice flips up says, "Warm," and they say, "Ship it."

Corey: I feel the need at this point to be very clear that you speak for yourself and not me in this. I, on the other hand, would never disparage machine learning in such a way.

Pete: That's going to come out as the new deep dice machine learning, and it's only priced by $10 per thousand rolls.

Corey: God does not play dice, quoth Albert Einstein. Yes, we freaking do, quoth Werner Vogels.

Pete: But really, I think Amazon is known for building things their customers ask for, or maybe just a random person on the streets of Seattle ask for. But this is, again from my time at CHAOSSEARCH, this is a problem. People want-

Corey: You saw that with the DeepComposer. Someone asked for a keyboard with machine learning built into it, and it turns out that person was probably Jeff Barr. But they didn't ask enough follow up questions to learn it was a keyboard you could type in, as opposed to the kind of keyboard that you could make music out of, and now we're stuck with DeepComposer because Amazon never turns anything off.

Pete: Well, now you can make some pretty cool EDM music I assume, and maybe next year, you'll be on stage at Replay.

Corey: Right now, I think they're more concerned with how to keep me from attending Replay, but that's neither here nor there. I make nuisances out of myself. It's kind of what I do. One of the things that did surprisingly well was, and I'll throw a link to this in the show notes, was my Twitter thread on my re:Invent nature walk through the Expo.

Pete: That was actually one of ... There's obviously a lot going on on Twitter, hard to keep up on everything, but as I was catching up on threads and things at the end, that was one of definitely I think the better threads that I saw, and something that ... I personally do it, where I will walk the Expo hall and honestly look at just a lot of bad marketing or mistakes people make or how aggressive the Datadog BDRs outside the booth are. But I think you really brought it to a new art form as you are wont to do.

Corey: Yeah, for those who are unaware, I acted as if I was on a safari, guiding a tour group through various booths, pointing out various companies and making up compelling backstories about these animals and their ecosystem. Remember, don't let them scan your badge or they'll grow to depend on re:Invent attendees to survive, and we don't want that. I will say that I heard rumors, though I couldn't confirm them myself, that logs.io has taken the mantle of aggressive badge scanning from Datadog.

Pete: Interesting. I actually walked by it a few times and I never got approached from them, although I did stop by the Datadog booth to say hi to a couple of friends, and of course got the requisite scan. So I feel like I should win an award for the furthest away you can stand from a booth and still get that scan.

Corey: Oh, you'd be surprised. The fun thing's going to be the people you never saw coming. Some of the better Datadog BDRs can scan a badge clean of 200 yards.

Pete: It's a real talent. I mean, I was looking out for the ... In the old days of the Pringles can WiFi thing. I'm just waiting for one of these, for booths to have the little secret Pringles can pointing at people as they walk by.

Corey: Absolutely, although they're going to wind up doing it in something far less obtrusive than that, because the Pringles can, come on. They're going to at least wrap it with their branding.

Pete: Yeah, they're going to put the dog on it or use those socks for something. They'll find a way to get that dog on it.

Corey: One of my favorite responses when approached by someone from Datadog is, "Hi, are you familiar with Datadog?" is to get this outraged look on your face and come back with, "Date a dog? That's disgusting!" And watch how quickly they backpedal.

Pete: Well, last year I talked to a friend of mine who's at Datadog, and we were chatting and I was talking to him, and Datadog, huge booth, they have multiple booth because they're so big and spend so much on re:Invent. And I just said to him, "It must be pretty amazing. People come up ..." And I've only worked at start ups, so when people come up to me, it's like, "Where do you work? Have you ever heard of my company?" And the answer is, almost often always, "No, I haven't heard of you. Tell me what you're doing." So I said to him-

Corey: Are you sure? We're AWS. They get super offended when you say that there.

Pete: Yeah. What do you guys do again? What is Amazon about?

Corey: Our product strategy is yes. Next question.

Pete: So I said, I was like, "It must be very interesting to just, people come up and know exactly who you are," and this is where I realized that we live in the bubble of technology and that, he said, he's like, "Most people that came by had no idea what we do." And I just kind of was blown away by that statement. I'm curious if that's still the case post IPO, but I'm guessing so. There's still a lot of new people into the space and really, I mean, how would you hear of Datadog if you've never been to an industry event and are just getting involved in cloud?

Corey: Hey, if I had a SaaS product, I'd almost certainly pick them to monitor it. It was easy to make fun of them when they first came out. It's, "Look at this nonsense. It's easily understood, it's like Fisher Price for monitoring." And you look through it, and yeah, that's exactly what people needed. "Ha, look at those fools making a product that's super complex accessible to people. Ha." And it was kind of amazing going through that evolution. I mean, to be honest, the best sales pitch for Datadog for a long time was trying to use one of their competitors.

Pete: Yeah, absolutely. I mean, I was a very early Datadog user a very long time ago, and I put the agent ... And this isn't even an advertisement for Datadog. I was just a happy user. What I always said is their time to value was incredible. You could deploy their agent, click a couple of boxes, and it was like, hey, here's some dashboards. And as a monitoring snob as I often am, I was one of those people that said, "Oh, it's too simplistic, it's too basic. I don't have any of the flexibility I need, compared to other solutions in the market." But again, the people coming into the space, they're so new that they actually don't know what they need, and oftentimes, I just want to see a dashboard with a line on it, and I can point at it and say, "That looks weird." Not have to go to Harvard or MIT to get a doctoral student-

Corey: Right, and of course-

Pete: To help me monitor.

Corey: They're public. They can't have fun with it, but if I were doing some of their branding, it would be, Datadog. It's monitoring for people with real work to do. You don't have to mess with things like that, and I'm a big fan of making things more accessible to people no matter where they are, which in turn gets us to some of the announcements at re:Invent. What did you think of the various keynotes?

Pete: So I loved the keynotes. I mean, so I've actually been to every single re:Invent with one exception where I worked for a DNS company that didn't actually use any Amazon services. So seeing the evolution of the different re:Invents has been amazing. The very first one honestly was quaint by comparison. I think there were only 4,000 people at that first re:Invent. But the first few re:Invents, every single one was price drops and more price drops. The rate of new services was a lot slower and now, you can definitely see the rate of services is very high. The rate of new kind of hardware things that come out and the innovations that they're doing there are very impressive. But what I think what I saw, especially in the new product releases that came out, is just how pervasive ML and AI is within Amazon and just how it is touching so many of their services. I mean, I haven't done the actual count yet, but my guess is is that if you were to bucket together kind of announcements, ML-related announcements probably was the number one spot.

Corey: I think that you're probably spot on. To me, what really wound up striking me was the tone changed. It wasn't aimed at start ups at all at this point. At this point, it seemed like it was aimed much more at big E enterprise, where people who are just starting the transformation journey, it's like oh, you have an ancient piece of crap architecture. What do you do with that? About $8 billion of revenue a year. Why do you ask? And I think there's a lot of validity to that. There was a lot more Goldman Sachs, a lot less Netflix on the keynote floor.

Pete: Well, that is an extremely good point. Goldman Sachs I think had a booth at the event. I mean, just think about that for a minute. But you're absolutely right, which is, Amazon has a lot of customers. There's no doubt about that, and they say numbers like a million customers. But mostly, I always think that's like me and my 12 Amazon accounts for various one-off testing, so I don't know really how true it is. But-

Corey: Oh yeah. I wonder how they disambiguate that. It's like, well, we claim to have 20 million customers, at least going by the number of e-mails we send everyone.

Pete: Exactly, exactly. So, but I don't disagree. Obviously, there's a lot of people using it. But there's still a huge contingent of people that operate in the data center space, and they have hundreds of millions of dollars invested in existing data center spaces. And if you think about insurance companies and airlines and these huge enterprises that are, maybe they're playing around with cloud, and in talking to some enterprises recently, a lot of them have been using things like Lambda and Serverless, which has been a little bit of a surprise. But yeah, I think there is still a lot of market to be won in cloud for those, I don't even know if you should call them, ultra large enterprises or old stodgy data center enterprises.

Corey: Yeah, and it's easy to sit here and make fun of those companies, but there's a lot of business there. That's a trillion dollar industry or two, and rounding errors, and it really throws into stark relief how the big cloud providers have been fighting over a very small slice of the pie when it comes to interesting start ups, doing fascinating and far future things. But heck, I work on AWS bills for a living, and something I see is that the majority of spend is on EC2. It's fascinating.

Pete: It's something where if you think about the large enterprises, and this was actually a conversation that I had with multiple people this year at re:Invent which is, I had asked a friend of mine who was at a large retailer, and they, data centers, the whole deal. I mean, everything super-antiquated. And they were very aggressively using Serverless on public cloud. And I had said, I was like, "Wow, that's really innovative, and you think you would hear smaller companies adopting Serverless more than larger companies." And they just said to me, "Listen, finding software engineers is a lot easier for us than finding people with operational expertise. Those are a lot harder to find. And we can also hire software engineers from a variety of sources that normally you wouldn't be able to kind of find those ops folks. We can hire people from code boot camps and things of that nature, and so, hiring people out of college."

And they're talking to me about how they're replacing a mainframe that does all of their ERP stuff with a series of functions, and they just don't have the operational expertise and I think they're now realizing, especially those big companies, they've realized that they need to move to the cloud and now they're just trying to figure out, what are the right workloads and what are the things that they have or don't have that they need to ask for and all that fun stuff. But it's pretty amazing, and I really wonder too especially with the Amazon Outposts stuff that you can actually now get, if again that's another classic, get this technology into the hands of mega enterprises.

Corey: And that's what's fascinating to me. Before we get into Outposts, you know what else was awesome?

Pete: What was awesome?

Corey: CHAOSSEARCH. That's right. This episode is sponsored continuously by CHAOSSEARCH. Fundamentally what they do is they take an Elasticsearch compatible API and slap it on your data. The data itself lives in S3, queries take place through containers, and even after their licensing charge, it's still something like 80% less expensive for most workloads than running your own Elasticsearch cluster on top of it. The additional benefit of course is that you don't have to muck around with managing Elasticsearch. And increasingly, you're not going to be threatened to be sued by Elastic the company who is apparently pivoting business models in Oracle's footsteps and becoming a law firm.

Pete: It's like Lambda except it actually works and you can search upon it.

Corey: And it isn't defined by its constraints.

Pete: The other part too is you can go try it out now. I think that was one of the greatest features that they had been working on over the last few months was go sign up right on the web site and you can get into CHAOSSEARCH, you can point it at your S3 bucket and start indexing some data and run some queries and can get equally as mind-blown as I was when I saw that product for the first time.

Corey: Yeah, it's absolutely worth paying attention to. They were generous enough to provide a space to pass out my stickers and help boost my ego still further on the re:Invent Expo floor. So I do feel a certain sense of kinship with them. And they feel a certain sense of, for God's sake, get our name out of your mouth. Do you have any idea how disruptive you can be? And I figure that's a good mutual relationship, so my thanks to CHAOSSEARCH for their continuing and baffling support of my nonsense. So we're talking about Outposts, and I think those are fascinating because it's, how do you make cloud accessible to people who are scared of the cloud and possibly their own shadow?

Pete: Yeah, I remember last year, right, they announced it, and it was ... I don't even know if they announced it as preview or what the specific announcement was, but they just kind of said, "Hey, we're doing this Outposts thing." And I talked to a friend of mine who works there, and actually works on the Outposts team, and I just said to him, "What's going on? And where's this going? Can I buy it? Can I get an S3 in my data center? Can I get a this, can I get a that?" I think that was one of the things that I heard from the keynote is how S3 is hundreds of microservices and so it's like, "Can I get an S3 in my data center?" And he gave me the best, most Amazon response of, "Yeah, you can have whatever you want. Whatever you want." Now it is generally available, right? And you can apparently go to the portal, go to your Amazon management console, and spec out in Outposts.

What I did notice, and I'm not sure what kind of you saw with it that really caught your eye, but what I did notice is that it is fully managed. It's not just, here's a rack of servers with some software on it. It is a fully managed service, so I assume that means is that Amazon engineers are managing these remotely disparate Outposts kind of wherever they're at.

Corey: And there are workloads that are not suited for running in the cloud due to a variety of different constraints, so that does become an interesting option. What I like is the pricing is not completely out of reach. The development racks start at seven grand a month on a three year commitment. I'm still not going to put one in my living room, yet, but it at least, the option is there if I need it, mostly because I want to watch the look of befuddlement on the AWS team that comes out to do the installation. And yes, it does come with one.

Pete: Would you cook them dinner when they came over?

Corey: I don't know if it works in the opposite direction, but personally, I tend not to eat at AWS events just due to the risk of poisoning. For example, in my hotel room, it was super nice of them, they left me a live wolverine as a welcome gift.

Pete: I would've assumed it would've been a snake or something quieter that could sneak up on you.

Corey: No, no, I'm not a subtle creature, so they didn't give me a subtle creature in return.

Pete: But you survived. You survived a re:Invent all, I don't even know, six days of it?

Corey: Ugh, but it felt about three times longer, at least.

Pete: Yeah, I think we both actually had the misfortune of booking our travel such that we had days on either side, just to account for complexity of travel and other things, and so I was there pretty much all day Sunday through all day Friday. And there were things going on. I mean, it's truly a, let's say Monday through Friday with hours on Sunday. I mean, it's a truly, a five plus day marathon of cloud.

Corey: I love our annual tradition where we grab breakfast on the last day, and we just sit there and barely talk, because at that point, we're socialed out.

Pete: I actually, I feel like this year we both went really all in with meetings and workshops and whatever, all the different activities we were doing, kind of left it all on the field, because when we sat down at that breakfast and you came over and just had this look where you were full of energy, and I've seen you on stage and do that thing. But you had this look, and then you looked at me and I had this look, and we just almost were like, "What if we didn't say any words, and we just sat here quietly?"

Corey: And it would be beautiful. Unfortunately, it turns out you cannot give a session track talk that way. I tried.

Pete: I saw your session talk. That is actually a great segue into that, which was about responsible disclosure. I really loved that talk.

Corey: Yeah, that was fun. I wound up giving the talk twice. It went vastly better the second time, because my co-speaker and I hadn't the chance to rehearse in person together, but, and also, as an added bonus, the second time, the cameraman didn't fall off the podium midway through the session and have to be attended to. We're going to take a five minute break before we go back to making lighthearted jokes about cloud. Can someone make sure he's still alive please? Yeah, it turns out that it's super hard to recover tonally from something like that.

Pete: Did he fall over from laughing? I mean, I assume that's what happened.

Corey: I assume he fell asleep, because we were talking about ISO standards at that point.

Pete: Well, what I thought was most interesting was a couple of things. So I had the opportunity to get over to the Aria. I had the opportunity to get to the Mirage, which was where your redo was. I never got down to the MGM, but I think this was really the first year that the scope of how many locations had conference talks. And it wasn't like a side room at the Mirage. I was deep into the bowels of the Mirage conference center, and it was, I think the first time I'd really seen this, again I've gone to conference talks at re:Invent where it's a room and there's someone on stage and everyone is in the same room for the same talk. And the room at the Mirage you were in, there were three talks going on at the same time and we all had headphones on. Now if you were in the earlier rows, or the closer rows, you could probably just hear fine. I was a few rows back.

But I actually wanted to ask you, as a speaker, did you find it challenging, because obviously you could hear the other speakers, maybe that was popping out. But also, did you find it challenging because I actually found myself not laughing like I should have because I had headphones on. It was a weird dynamic.

Corey: That's absolutely a real challenge. I've had the benefit of having given a few of those "silent disco" style talks before, and people are not generally willing to laugh when wearing headphones unless they're sitting on a city bus and not wearing headphones. So it's challenging to wind up getting the same crowd reaction. You can power through, but it has to be a little bit over the top. It also adds a certain sense of intimacy, because now instead of me speaking to a booming projecting microphone, I'm speaking and it is whispering sweet nothings directly into your ear, just like I'm doing right now on a podcast. Shh, go to sleep, go to sleep, watch out, there's a bridge. Yeah, it's fun.

Pete: Yeah, I was very surprised to see that kind of setup in there, and I think there were actually, that room that we were in, was actually six specific talks going on. And they were sizeable. I mean, they were rooms that held a couple hundred per talk. I mean, they're cavernous rooms.

Corey: And that's sort of a challenge too. When you have 200 people attend a talk in a room that seats 600, it feels sparsely attended. Whereas if you give a talk and there are 50 people in a room that comfortably seats 20, it feels like it's a way more successful talk based upon audience size. It really changes the dynamics of the room. The worst talk I ever gave was at a puppet conference, years ago. And it was the last talk of the conference, they'd opened that free bar 20 minutes before the talk. It was a room that would've seated easily 400 or 500 people. 12 people showed up, no two sat next to each other, and it was at that point, hell with it. Let's just have a fireside chat. It was not a very well thought through process.

Pete: That actually reminds me, not re:Invent related, but I love sharing this story about room sizes for conference talks. I understand how hard it is to put on a conference and, knowing the room sizes in advance and who actually wants to go to which talk, that is a challenge I can't even begin to imagine. But a friend of mine side, "I hope you'll come and speak at my conference. I loved such and such talk that you gave." I said, "Great." And at the same time, they had actually someone drop out of the conference, and just said, "Hey, is there anyone that could help?" I directed them to a friend of mine who actually had a lot of research around the DDoS that happened at Dyn, the DNS provider, a few years ago. And I said, "Oh, you should talk to this person," who ended up talking at the conference.

Of course, as fate has it, they scheduled their talk directly against my talk, which is fine. I mean, two different types of things. But also due to a snafu, the room that I was in easily held 2,000 people. And because my talk was against something that was admittedly a lot more interesting, there were maybe 12 people that showed up in the 2,000 person room. But what was great is how sparse everyone decided to sit, and there were people sitting in the back that I just was like, wave to me if any of this is making sense. And I still give my friend a hard time about that. He feels terrible, but it was still hilarious.

Corey: And that's sort of the entire point of giving conference talks is to get out there, try new things, see how it resonates. Often I ... What always disturbs people at the more, shall we say, thoroughly produced conferences like re:Invent is I often have very little on my slide. Now I improv most of what I say, which is good, because most of what I say would never pass slide review of, "You can't actually say that in front of other human beings on our stage." Yeah, but I don't know I'm going to say it until I'm up there and it pops into my head and I go for it. And it either lands super well, or it goes over like a lead balloon.

Pete: I totally agree. I think if, especially too, I like giving talks multiple times just because you think of things on the fly. Although I practice in advance with talk notes because if I think of something, it pops in my head, I start talking about it, and then I'm like, "Wait, where was I at? I need, where." And then the notes kind of bring me back a little bit. But you're right, if you need your slides to be approved by the approaching Oracle-sized legal team at Amazon ... I mean, just kidding. Nothing can approach Oracle's legal team size, but you have to have real legal people review it for things, just keep it light. No one really wants to read slides anyway. They really are there to hear you talk.

Corey: That's something I always wanted to do was try and give a talk with no slides. I think that's called improv, but I'll have to do it at some point, just to see what happens. To some extent, that's kind of what a podcast is.

Pete: Absolutely. I think everyone listening to this can just imagine the slides we would create for this talk.

Corey: So other releases of re:Invent that were neat. They had a crap ton of announcements around Redshift, which is awesome, except for the part where Redshift runs on burning piles of other people's money. So the price just to get started is not nothing. So I need to start looking at some larger level data analytic shops and start seeing what they're doing with Redshift. It does change things a bit. You can now query Redshift and hit a whole bunch of other data stores with federated queries. It winds up having a whole bunch of export options. You can, again, have data live in S3 rather than live online, CHAOSSEARCH. That tends to be the direction that the world is going in. But there really is something for everyone. What announcement do you think made the least waves when it should have made more?

Pete: So, made the least wave. That's actually a really good question.

Corey: Yeah, what do you think is the most underrated release so far?

Pete: Well, I think it's underrated just because no one knows what to do with it, is really that DeepComposer release, which is, I think there is something there that will ... This will end up powering, right, that I don't think anyone realizes yet. And so I think in many ways, people saw this as a, it's a MIDI keyboard and a very early UI where you can play eight bars of music. At least, that was, I did the workshop for it, and it maxed out at eight bars. So that's not a lot of music that you can create. But they talked about this generative, adversarial networks. They're called GANs, and I am not an ML person. I find this stuff to be super fascinating though. And they really talked about how you can train these GANs against different musical types.

Where I think this is actually super interesting is that I've actually recently, in the last couple of years, started to learn the piano. I wanted a hobby that wasn't computers, and mostly my daughter gave up the piano and I had a piano in my house. So since I had to look at it every day, I felt like-

Corey: They tend to accumulate, don't they?

Pete: They really do. And so I'm looking at a piano every day thinking, "I'll go and learn this thing. For people that want to make music, the accompaniment around it is challenging. You don't have an orchestra or whatever and maybe you don't have all the other additional equipment, but what if you could have some ML that you can train based on a certain style of music? The demo they showed, I think, was super impressive but only to classical music nerds is, they had a training of countless Bach recordings that are all public domain, and you could play some music, and because there was so much source data that they trained, you could listen to the results at all of the epoch at each time it improved it, and you could hear, let's say, Twinkle Twinkle Little Star or something at five epochs, and then at a thousand, and at two thousand.

And at the end of it, which I'm not sure how many people actually made it that far in the workshop, just from a time constraint, at the end of it, it was mindblowing. It was like you were listening to Bach play this individual song. So I think, again, I don't know what people are going to build with this, but the technology underneath it, these generative adversarial networks, that technique of which there are papers recently published on, I think could be something that we don't even know yet. It's just so far out in the future.

Corey: I think there's a lot of neat things happening there. I guess what I found interesting was a bit more prosaic on some level. You see the ... There are a lot of things that came out, but one of the notable things in the world of hardware was the Graviton2 processor. It's been a banner year for chip manufacturers. I mean, AMD has launched their whole second generation of Epyc, AWS launched their second generation custom ARM chip called Graviton2, and Intel threw an awesome party at Replay.

Pete: So I totally agree. I actually have that tab open on my laptop because I also want to talk about that. I'm not a hardware nerd. I got started in technology when it was all hardware. I worked in a data center for a hosting company that was-

Corey: The magic smoke gets out and you have to buy a new one, that gets expensive. Software, you can get just reset your git repo.

Pete: Exactly, right? And I thought I remembered reading around the size of the, I don't even know the words, but of the circuits or the size of the transistors I suppose. I don't ... I'm sounding like an idiot because I just don't know anything around hardware, but I thought there was a lower limit that we were approaching and something they called out was this seven nanometer manufacturing technology. So it appears as though Amazon, in addition to creating everything from hardcore software services and MIDI players for music, also seem to be moving into this working with chip providers to improve how they make their hardware. It's pretty crazy, the scope they're in right now.

Corey: I think that Intel's also falling massively behind. I think that their public roadmap, they weren't able to meet, has pissed off pretty much everyone. I mean, you see Apple laptops being hamstrung. I'd be shocked if we didn't seen an ARM MacBook in the next couple of years, just based upon needing to control their own destiny. At this point, it's ... Apple laptops are not cheap, and the fact that you're getting a processor that a couple years old at this point is not super awesome. And I understand this stuff is hard. You don't do agile chip development in most cases. But at the same time, it seems that this game was entirely Intel's, and now they're really alienated most of their distribution channel based upon that. And I had AMD down as all but dead, and look at them now.

Pete: I think it's amazing to see. It's a good place to be where we have diversity in chip manufacturers. You don't want it to be just Intel everywhere. But also too, something that has happened in these recent years is security issues around specifically Intel chips, in design choices that they have made in this race for faster and faster. And my mind was appropriately blown away when I watched at DevopsDays Chicago, Jess Frazelle talk about the kind of underlying bits of hardware. You think about Ring ZERO of an operating system, but she was talking about Ring -1 and Ring -2, where you're just talking about everything that happens before your OS starts up. It was fascinating, and also frightening how complex that code is that no one really understands. And so-

Corey: Yeah, how do computers work? Nobody really knows.

Pete: It's truly amazing. And so I think it's great to see AMD and these other chip creators coming out and building new and innovative technologies as they are.

Corey: One of the things that I also think was super neat that isn't getting nearly enough attention, and it certainly deserves more, is S3 access points. Part of this is due to Amazon's inability to tell an articulate story on a keynote stage. They talk about that almost entirely like it's aimed at data lakes. Here's the trouble with that approach. No one thinks of themselves as having a data lake, even if they have near exabyte scale of data living in S3. Secondarily, S3 access points give you very granular access control to the same S3 bucket. So you can provide different end points and different applications that are scoped appropriately. That doesn't require you to have hundreds of petabytes of data. You can do that with three megs or less. It's not a big data feature. It's an access control feature, and by talking about it in a data lake context, I'm worrying that's sailing past people who could really benefit from it.

Pete: Yeah, data lake is such a polarizing term that I wish we could come up with something better for. I actually gave a talk where I didn't want to use the term data lake in talking about kind of the future of monitoring, so I called it a data bagel, and mostly just because I feel like if you can create a dumb term for something that gets people to stop and be like, "Well, what does that mean?", versus a data lake where people are like, "Oh, I know what that is," right? The data lake concept is, it's one of those things I think you and I spoke about this earlier, was, no one has said, "I want a data lake." I mean, there's probably some suit-and-jacket person saying, "Oh, we need a data lake strategy in our company."

But I think like you said just a minute ago, there are companies out there pushing petabytes and exobytes of data into places like S3, they wouldn't not call that a data lake. Even thought they have a multitude of services pulling data from there and doing different things with it, they still wouldn't call it a data lake. So I think by branding it as such, I totally agree, it's going to ... People are going to think that I'm not big enough or they just don't like the term. It could be a little bit weird.

Corey: Something else that I thought resonated with me at least was the computer optimizer. I don't think that people are going to use it for much. It tells you, "Oh, with this instance family, this is the performance profile you'd see." It will help people figure out what they should be doing, but having spent three years in this space, even if you tell people exactly what they should be running instead, very often they won't move due to inertia, due to fear, due to the fact that it's not easy to move some workloads to a different size or family of instance. And I get it. I think it's a great tool. I don't for the life of me understand why it's a top level console feature rather than buried within Cost Explorer, but we'll see.

Pete: Yeah, I'm going to be intrigued to see how many people actually use that technology. As an operator for so many years, I've wanted to resize, and I had done, this was my time at Threat Stack. We ran a cloud security platform, a lot of data. Elasticsearch was one of the many databases we used there. We continually optimized the size of it, moving from I2 instances to R4 with EVS. The NVMe instances came out and we moved it to there. But moving data is expensive, and you know this probably better than anyone about moving data within Amazon is extremely expensive, just the network transfer.

So what ... You might save money by moving to a different instance type, but the ROI could take some time if you have to transfer hundreds of terabytes of data over the network, or for people who really play in the big data world, you might have a petabyte or a petabit of network transfer. That's real money in Amazon. So you can go through all this work and the people work and the testing and whatever else. You could have a 12-month ROI on an instance type change, depending on what you're building.

Corey: Yeah, and oh, the data transfer as well. How much does your database speak to your application server? Most people don't know that, but they're about to know that in a big way.

Pete: Yeah, right about the time you start line-iteming your Amazon bill and said, "Wow, our data transfer last month was $60,000. Does anyone know why?" And I don't even know if Amazon knows why, other than they might say, "Oh, you must run Cassandra," which is a common joke within-

Corey: Oh yeah, which they now have a service for. What I found was fun was that when I was asking people at Amazon internally for the pricing of Cross-AZ data transfer, no one would answer me. And I thought they were stonewalling and trying to be insulting. It turns out, no one actually knows. So once I ran some experiments with DD and Netcat, suddenly they update their documentation almost instantly to explain what happened. It wasn't that they were stonewalling me, it's that no one knew.

Pete: I remember when you started talking about on that Twitter thread and we did the math on it, that again, I don't recommend that you build an application this way, but if you wanted cross-region availability, it actually makes more sense, and again, you don't have latency requirements of Cross-AZ. It makes more sense if let's say you were building an Elasticsearch cluster or a Cassandra cluster to basically put it in one AZ and us-east-1 and one AZ and us-east-2, and that-

Corey: CHAOSSEARCH gets around that. However ...

Pete: And that appears to be half the price of Cross-AZ traffic.

Corey: It is half the price. That's why it appears that way. And I would challenge your assertion that you wouldn't recommend building an application that way. If you could build it that way from day one, you're already designed for multi-region, and that is potentially a far more durable architecture if it needs it. And not every application does. I mean, a lot of the business stuff that I do can sustain an AZ being down for a day or two because it's not in line with what customers need. I mean, worst case, some of my AWS newsletter creation stuff lives in a single region. And if that region, us-west-2, is unavailable for a few days, well, not for nothing, I have a much more interesting story to write about freeform in the newsletter that week.

Pete: You've got to know what you're optimizing for, and you're right, which is if you can say you're multi region, you're actually probably beating out a lot of companies in the space where they are still stuck in single AZ. Because let's be honest, cross-region is hard. It's not an easy challenge to get there, and it takes time.

Corey: Oh my stars, yes. What else did you see at the event that was notable for you?

Pete: Well, there wasn't as many people as I expected. And there were more than last year-

Corey: Yeah, only 60,000 or so.

Pete: Only 60,000, which I believe was still more than last year. I think last year, I don't know if you remember-

Corey: You're right though. Most of my clients did not have a presence there this year.

Pete: And I wonder, there's a lot more local events with the summits. I've been to a few. They're all kind of mini re:Invents, single day. They're honestly, I think they're really great if there's one near you. They're free too, which is amazing. You can attend these local events for free. And also, it looks like they're splitting off some other events as well with re:Inforce, their security side. I feel like they had an event that was specific to AI and ML. But I also wondered too if that there wasn't a 50% growth or anything, but they still kind of took over the city with talks at so many locations. Again, I have no data points to support it, but it does feel like talks were spread around much more. And as at least I navigated around to different locations and places, I didn't find it to be as cramped and anxiety-inducing of so many people as it had been in previous years.

Corey: One other thing I found super interesting was the higher level analyzers, the abstractions built on top of abstractions. We see this with the IAM Access Analyzer, the S3 policy analyzer. AWS, I'm sorry, Amazon Detective, which is a security roll up above things like GuardDuty, above Security Hub and the rest. I think that that is super interesting in that it's distilling it down further and further to something actionable. Because how many security breach stories have we heard about where, oh, they installed the alarms and the alarms were false positiving all the time, so people started ignoring them, and then they missed the actual important thing. It feels like chatty monitoring systems are only there to be able to blame people with once something breaks.

Pete: I think any business that gets involved in helping to separate the signal from noise is going to have a really big success. I don't want a thousand alerts per day of security stuff. Even if I'm getting attacked a thousand times, can you distill them down to let me know what's the three most actionable things I need to do? Because at the end of the day, security people, ops people, we're all so busy, we just don't have the time, and that's where mistakes happen. So if anything can just say, "Hey, that thing right there? Go for it," that's going to be-

Corey: Right. Random IP address on the internet, port scanning me, that's one tier of problem. My firewall node is now enumerating S3 buckets is something else entirely.

Pete: Exactly. And so, that ... and if you're at a level of scale where you've got trillions of these events happening, how do you find ... Again, it's the needle in the haystack, it's the unknown. I don't even actually know what I'm looking for, but if something strange is happening of a high priority, please let me know about that one, and maybe not so much let me know about the other stuff.

Corey: One other thing that was announced in the keynote, or at least not announced but pretty obvious, and I think it's unfortunate. I think Amazon struggles across the board with empathy, and it shows, because yet again, I'm sad to report that the re:Invent house band was not put to sleep. They are clearly suffering. They are not happy. Please, put them out of this world.

Pete: I unfortunately missed the midnight madness, but I only missed it-

Corey: Oh, this was Andy Jassy's keynote, and then Andy Jassy recited the lyrics. I want an Alexa skill that is just Andy Jassy reciting song lyrics to me. I think that would be phenomenal.

Pete: Yeah, he could do it onstage as part of his keynote for the version two of DeepComposer where he just says the words, and DeepComposer is composing the music behind it.

Corey: Oh yes, I think the “Andy Jassy Sings” album would be phenomenal at holiday time.

Pete: That's a double platinum if there ever was one.

Corey: Absolutely. Anything else you noticed that was fun, exciting, worth mentioning?

Pete: Honestly, this was a weird event for me. I was helping CHAOSSEARCH with the re:Invent-

Corey: CHAOSSEARCH.

Pete: With the re:Invent event, and helping them make sure ... We did a lot of planning for it, so I wanted to make sure that they got through it and were all set. And I really spent a lot of time going to workshops. I mean, I always think about, what's the best way to do re:Invent? If you're getting certified, that's a great opportunity to go, do a boot camp and get through the certification. The workshops, they don't redo those. All the conference talks are online after the fact, so spending time trekking to different talks I find to be a lesser activity for me at least, just because I can watch them later. But I really try to optimize and spend a lot of time talking with folks, talking with friends of mine who are in town, talking with even venture folks, and just try to hear and listen, what are people building and what challenges there are.

And I think the thing that I noticed, and the thing that I keep thinking about that really resonates is that kind of sys admin and dev ops expertise, this is an expertise that is going away. I mean, more and more people are abstracting away what we used to deal with earlier in our careers. But it definitely seems to be happening more and more, and the folks that are still around doing that type of work are going to become more and more valuable, but also going to become busier and busier. So always trying to chat with people and see what are the pain points and challenges.

And this is where these public clouds can really have a big impact in helping take over something that is just not part of my business, and I just don't want to deal with. I'm not going to run Postgres unless running Postgres is part of my business. I'm just going to use RDS. And that's the obviously very simplistic way of thinking of it, but as you see all these new services come out, they're really trying to abstract more and more, even abstracting the abstractions they've created from previous products.

Corey: That is fascinating in that they're now building services that in turn roll up other services in the context a human can understand without a deep subject matter expertise on something. And in large scale environments, that's huge. I am seeing a lot more of what they're releasing that is clearly built for scale. And that's a term that gets thrown around way too lightly. But take the simple stuff. You look around on GitHub. There's an awful lot of open source tooling that does all kinds of things, cost analysis, tagging of things, enumeration of various systems. And they all assume that you're running in a small scale environment. Well, when you have 15,000 nodes in a single account in a single region, and you try and do anything to all of them at once, you hit massive API rate limits. And that is something that a lot of tooling naively does not take into consideration, because why would you? If you're building something to work on your 200 instances, you're never going to hit those rate limits that you will at 10 or 20 times that scale.

Pete: Absolutely, and I think it goes back to what you said, which is, they've been messaging to the enterprises more recently, and I actually wonder if some of these enterprises that have adopted it, let's say three years ago when they really started making a push, are now reaching scale points. Or maybe it's the enterprises that are left are like, oh, they can't handle my scale, and this is them trying to show no, actually, we really can, and let us show you how that's even possible.

Corey: That's something that I think is overlooked on the AWS side. Whenever you build something in the world of AWS, there is no way around it. They can't order freaking pizza without having a massive logistics challenge just based upon the scale that they operate at. Part of me worries that my sarcasm and my snark get overlooked or taken too seriously in that they're focusing on all the things that AWS does that I find annoying or humorous, rather than the very hard work that some very intelligent and very dedicated people do. I worry that people at AWS who build these services read my snark or sarcasm and then walk away feeling crappy as a result. That's never been how I want to come across, and if you're listening to this and you're one of those people, first, I'm sorry. Secondly, please reach out. I'd love to hear more of the stories about what it takes to build these services, because I certainly don't know. All I know is what it takes to make stupid jokes that make people laugh. It's a bit of a different scale of a problem.

Pete: I'm the same way. I've worked in a lot of start ups, and scale is relative. I mean, there are some places I've worked where it's large scale, but in comparison to who, is really the question. There's a fascinating world within Amazon, just how they even construct these services, from the product management side to the engineering side and everything else. I have to imagine that some of these services that they create go from zero to 100 miles an hour once the APIs are around, and so it's really true interesting scale. But I do like to say that when it comes to scale, and I remember the talk a couple re:Invents ago, where they basically talked about their power generation facilities that they were designing. And that's where I was like, you know what? That's when you know you've made it, is when you are designing more efficient power generation facilities in order to actually power your data centers. That's truly next level scale.

Corey: That is, I think, the most interesting and useful thing to remember. Everything that AWS releases is important to someone. None of it is important to everyone.

Pete: Absolutely, absolutely. I mean, you don't have to use everything, and I think your bank account would groan if you really did. But it's the-

Corey: That's what other people's accounts are for.

Pete: It's the optionality, which is what makes it great. And obviously they're setting the tone, but these other cloud vendors are coming up behind with different ways and different technologies as well, and I think it's the golden era of computing. I mean, we've got some amazing things. You can start a company for almost nothing and get access to a level of scale and growth that was just unheard of five years ago, 10 years ago. It's really shocking. And I've talked to a lot of start ups, and I'm doing a lot of start up advisory and consulting right now.

And the thing that has been repeated most often is how many of these companies are just consuming either Amazon or other cloud providers' services as a first level citizen, right? They're not deploying an EC2 instance to run their app. They are going with Fargate. They're going with Serverless only. They're using the hosted database services. They're using the suite of tools, because they realize that that's not their business model, and they don't want to deal with it right now. And that's amazing. It's allowing these companies to go from zero to a real product and focus on their product and get it out to market a lot sooner. It's truly remarkable to watch.

Corey: It really has, and I think that what's the most impressive of all of it to me, I guess if I had to sum it up, was that Andy Jassy has to walk a razor's edge in these keynotes, because he's got to tell stories to a very diverse audience with all kinds of challenges, and in doing so, speak to everyone while putting no one off. And that is incredibly difficult as far as needle threading goes. And if the worst commentary I have on that is, "Well, I thought the re:Invent house band was a little corny," then I don't think we have much to complain about.

Pete: No, we really don't. I mean, when we think about things to complain about at re:Invent, it usually comes to some of the things they're doing, not the stuff that they're announcing. It's like the house band or the food or something, something absolutely silly like that, which I always laugh at as well because my wife works in education. She goes to conferences too. Her conferences usually cost $50 or $60 to go to, there's no sponsors, and it's a box lunch. And she's like, "Well, what about your conference?" And I'm like, "I think Foo Fighters is playing at it? I'm not sure." We're in a far different world than pretty much everyone else.

Corey: I think that's probably the best place to leave it at this point. Pete, thank you so much for taking the time to speak with me on this podcast today. If people want to learn more, where can they find you?

Pete: So I actually updated by blog for the first time in many, many years. It's pete.wtf. I feel like it's the best use of the WTF vanity domain. And I'm actually trying to blog a little bit more. There's a couple of posts on there about stock options and things if anyone has any interest. But I am on the Twitters, petecheslock on Twitter. And that's where I post pictures of various meats that I'm smoking and veggies that I'm smoking and cooking food, and hot takes on technology when the time calls for it.

Corey: Excellent. Well thank you for your time, and thank you to CHAOSSEARCH for sponsoring this ridiculous podcast.

Pete: Thank you a ton for having me back again. This is two for two, right? So next year-

Corey: And we'll do the third one next year, assuming you still go to re:Invent.

Pete: Absolutely. I'm going to put it on my calendar right now.

Corey: Excellent. Pete Cheslock, unaffiliated start up advisor. I am Corey Quinn, cloud economist, and this is the Screaming in the Cloud re:Invent recap. Thank you all for listening.

Announcer: This has been this week's episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production.

View Details

About Chloe Condon
Chloe is a Bay Area based Cloud Advocate for Microsoft. Previously, she worked at Sentry.io where she was an advocate for their open-source & hosted error monitoring tool, and created the award winning "Sentry Scouts" program. Her unique demos and projects with Microsoft Azure have ranged from fake boyfriend alerts to Mario Kart "astrology", and have been featured in VICE, The New York Times, as well as SmashMouth's Twitter account. Chloe holds a BA in Drama from San Francisco State University and is a graduate of Hackbright Academy. She prides herself on being a non-traditional background engineer, is likely one of the only engineers you'll meet who has played an ogre princess, crayon, and the back-end of a cow on a professional stage (a true "triple threat"), and is passionate about bringing folks with non-traditional backgrounds into tech.

Links

  • Twitter: @ChloeCondon
  • LinkedIn URL: https://www.linkedin.com/in/chloecondon/
  • Personal site: https://dev.to/chloecondon
  • Company site: https://azure.microsoft.com

Transcript
Announcer: Hello and welcome to Screaming in the Cloud with your host cloud economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: And this episode is sponsored by InfluxData. Influx is most well-known for InfluxDB, which is a time series database that you use if you need a time series database. Think Amazon Timestream except actually available for sale and has paying customers. To check out what they're doing both with their SaaS offering as well as their on-premise offerings that you can use yourself because they're open source visit influxdata.com. My thanks to them for sponsoring this ridiculous podcast.

Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Chloe Condon, a Bay area based cloud advocate for Microsoft. Welcome to the show, Chloe.

Chloe: Thank you for having me. This is such a long time in the making. I'm so excited.

Corey: It has been a while. Originally I started trying to get you on the show way back when you were at Century where you were doing a bunch of open source work with their hosted error monitoring tool. You were doing the Century Scouts Program, which I always thought was just cheesy enough to be interesting.

Chloe: Yes.

Corey: And then you went past that, went to Microsoft. You started doing a bunch of unique demos and interesting projects with Microsoft Azure that are just phenomenal. You have a fake boyfriend alert, which is awesome because when he came over for dinner he was totally convincing. I would never have guessed he was fake.

Chloe: Oh, he's so convincing, that smug look and we brought my skeleton son over as well. It was a family affair.

Corey: It was a blast across the board for everyone. People have known you from all kinds of interesting places. You've been featured in Vice, the New York Times, Smash Mouth's Twitter account somehow.

Chloe: Yes.

Corey: What's the story there? That's what I haven't seen in a bio before.

Chloe: Oh my goodness. Well, it's so funny because a lot of the demos and examples that I use at Microsoft aren't really the traditional, not your dad's Microsoft as maybe one would say. My background of course is in theater performance, so my brain works a little bit differently than other people. And then I also have ADHD, so the left side of my brain is very creative.

So when I think of these ways and examples and demos for using Azure, my mind always goes to these weird places, but I can't take credit for this amazing Smash Mouth scenario. Because I did a workshop of my fake boyfriend workshop, which is a Azure functions workshop that integrates with Twilio.

And essentially it was built for those awkward situations at conferences or parties or events where you want to leave early, but you need that like fake call to come in to save you from the conversation you maybe want to leave. And I was giving this workshop in New York at the New York reactor. And one of the women in the group decided to repurpose it. I usually have a little MP3 of Rick Astley's Never Going to Give You Up that plays automatically.

Again, can't take credit for that. That's a Twilio docs Easter egg that I'm absolutely obsessed with. But this woman Khalila, She added Smash Mouth's All-Star to it. So it's just a never ending MP3 of, hey, now you're an All-Star and you press this button and you get a text that says, "Somebody once told me so." I love to see people, especially with these open-sourced workshops and projects, people taking it and running with it and making it a fake daughter call or a fake, you know, "Oh, Beyonce is calling me." It's always fun to see what people repurpose with it.

Corey: I will admit, in years past I've had an emergency panic button like that myself at conferences. Not that I had to use them very often, but every once in a while when someone follows you into the bathroom because they want to hang out with you but don't have anything to say, it gets a little on the strange side. And it's, "Okay, now I'm actively uncomfortable." Doesn't happen often, but when it does, it is a lifesaver.

Chloe: And it's one of those things where with the nature of our roles and specifically with developer advocacy, a lot of our job is being not only just like a kind of spokesperson and voice for the company and to give advice, but to also be at a nice human being all around. And I'm an ambivert, which for those who don't know as an introverted extrovert. And I cap out my social interaction after a couple hours.

And I'm too nice. So if someone says like, "Oh, I must be bothering you." I'll be like, "Oh my gosh, you're not bothering me." When they are a hundred percent bothering me. So I have to socially check in with myself. My boyfriend has a higher tolerance for human interaction than I do. So I always strategically go, "Okay, 30 minutes from now is when he's going to want to leave. So I'm going to start bugging him now."

So it's very mathematical. There's a whole algorithm to my social interaction. So this Azure functions app solved a lot of problems for me at conferences and even just house parties really.

Corey: It's a way out of awkward situations without making people feel bad.

Chloe: Yes.

Corey: If you don't have a way to extricate yourself from a situation like that without stepping on people's toes, so to speak, people don't really want to have you around anymore. I don't want to say that there's a requirement to be a public figure. Even framing it like that sounds ridiculous and self-aggrandizing, but when people know who you are more than you know who they are, it changes the dynamic in a way that I'm not sure people can necessarily understand if they haven't done that in some form or another.

Your background in drama is another great example. You used to do a lot of interesting things on stages of a different kind.

Chloe: Yes. Yeah. I did many a musical in my day. Actually we just moved apartments and I stumbled upon, I've never had a child, but a photo of me looking very pregnant in Jerry Springer, the opera has a pregnant teenager. So I had a lot of very kooky credits on my resume.

Corey: For some reason I feel like there's a few of us in this space who tend to go down that particular path. I can't speak for anyone else, but my personal reasoning behind bringing humor into things as much as I do is without that, I've got to say the entire world of cloud computing is kind of dry. If it's a just the facts story, if it's a just the facts style of retelling.

Instead, I much rather would see folks getting the audience engaged. Because you need people engaged or they're just going to stare at their phone and miss the whole point of what you came there to tell them.

Chloe: Oh my gosh. Absolutely. And I think that Smash Mouth repurposing of that demo was such a great example. One else is Smash Mouth going to retweet an Azure project. Never, really, right? Smash Mouth probably doesn't even still know what Azure is, but that was sort of the brilliance of what Khalila had done with this repurposing of the app is like make it a pop culture, funny, often named and referenced thing and make Azure relevant, which I thought was just so cool and such this weird hybrid. I've done some hardware hacking unrelated to work on things like Furbies and tapping into that humor, nostalgia. There's really something there. I see a lot of people doing this like April Vogan code on Twitter. People who really notice like, "Oh there's something about this throwback thing."

Clippy is a really great example that gets people interested and engaged or even... You and I make a lot of puns on Twitter. I mean we're probably two of the biggest offenders. I'm sure you're about to say a pun. And you're developing one as I'm speaking right now. But I truly do... My background is in musical theater. I told my boyfriend when I first started my journey into learning how to program and going to a boot camp that if I bring anything to this industry, "I'm going into it for the puns," I told him.

Because I just kept thinking of all these awful, the more you know, things like that in my early days of programming and it's gotten worse. My plans have gotten better and worse in a lot of ways.

Corey: Well, that's the trick is it's very easy to come at this world from a perspective of intense cynicism. And increasingly I've been actively trying to lift people up more than I tear them down. My rule has always been that I make fun of large successful companies. So for better or worse, congratulations Azure, you win. You're now in that club.

I own twitterforpets.com on the other hand, because making fun of a real small startup is not going to win friends, influence people or have the impact I want the joke to carry.

Chloe: And I think back on being a performer and doing musicals and if you look up a picture of me, I'm this quirky large-eyed, blonde, five foot two girl, which is every girl who does musical theater, spoiler alert. Also quick PSA, if you're a man looking to have a musical theater career, Broadway is calling, they're looking for you.

It's the exact opposite in theater and technology. It's all women, no men. And coming into tech was quite a culture shock, but this is all to say that I would always get typecast as the ingenue because I had these big doe-eyes and long hair. But I really, really loved playing the quirky sidekick. I played penny and hairspray once and it was just so much fun. I got to do 25th annual Putnam County spelling bee, which is literally just like a long form improv musical that's about adults playing children in the spelling bee, need I say more?

But I noticed towards the end of my acting career that that's what I really enjoyed. I really enjoyed the working of the crowd and finding the humor in things. And being able to find that in tech has been so fascinating. I think first of all, I'm a very weird anomaly to most people in tech. I don't look like the traditional engineer, although I do own one Patagonia jacket and I do own several pairs of Allbirds and I did drink Soylent exclusively for a couple of weeks, but we'll just like forget that ever happen.

I'm very much this quirky pink loving. I'm wearing an extra-large Target shirt that's a little girl's shirt from target right now. I don't look like a traditional engineer. But I think a really, really big part of my aesthetic is showing this other side of tech, this humorous side of technology.

And RIP, I'm going to pour one out for my favorite tweet I've ever tweeted that I accidentally deleted, me giving a thumbs up in a women's bathroom. Just wearing this pink shirt, had this crazy look on my face and I was like, "Hey women, where are you?" Like the bathroom lines don't exist. Come and join tech.

And I think what made that tweet go so viral was people were like, "Oh my God, we've never seen this before." And this is so funny that it's true. And I think back of like looking at my pin tweet right now and my highest performing tweets, they're always some sort of pun. Like I think I had one for a while that was like Google in the sheets, Excel in the streets.

And my pin tweet right now is I like my coffee like I like my browser tabs in excessive amount to the point where it causes me mild anxiety. And as much as I would like to think that's humor, that's just the truth for me. But I think people, Danny Donovan gave a great keynote about this at anxiety tech, which people really started to respond to her illustrations and visualizations of ADHD because it's this collective consciousness of like, "Oh my God, that happens to me too."

And we see people like Cassidy Williams doing this on, oh my gosh, she's hilarious on-

Corey: She's amazing and on my short list of people I'd love to meet someday.

Chloe: Oh my goodness. I didn't know I was in the presence of greatness when Sarah Chip's took Cassidy Williams and I out to dinner and then I discovered if you haven't watched this woman's videos on TikTok or Twitter, she is one of the funniest people in technology. But that just goes to show, I mean I'm looking at her Twitter right now and she's got 54.4 K likes on her very funny video about when your code works on the first try. There's really something to be said for that ironic humor that we all face every single day as engineers.

Corey: Oh my stars, yes. I frequently said that multi-cloud is a stupid best practice, and I stand by that. However, if your customers are in multiple clouds and you're a platform, you probably want to be where your customers are unless you enjoy turning down money. An example of that is InfluxData.

InfluxData are the manufacturers of InfluxDB, a time series database that you'll use if you need a time series database. Check them out at influxdb.com

I think that people who come into this space from alternative backgrounds for better or worse tend to bring a unique perspective that occasionally lends itself to interesting storytelling. I mean, my primary skill is wearing a suit, although if I'm being perfectly honest most days it wears me instead.

That for some reason lends itself to a particular way of presenting myself on stage that was early on in my speaking career, grabbed attention in a way that I wasn't at all expecting. And it's fascinating seeing other folks like you, for example, that remind me, oh wait, I'm not the only person out there who doesn't lead with this is my code editor and here's some code I wrote and wow, that was a fast 45 minutes. Why is everyone sleeping? I mean, power to people who can make a talk like that engaging. I never could.

Chloe: Yeah. And I think that so much speaks to... There's so much talk these days about like soft skills are important and it's important that you're able to command a room and things like that. Obviously I have this very bizarre leg up. Especially when I was entering the industry, I had to position myself as like, "Look, I'm a more junior engineer but you're not going to find an engineer who's been doing public speaking for 20 plus years of their life."

I've been on stage since I was four essentially. So, I think that's a really interesting place to come from to be up-leveling your tech skills. But having those person skills because they ask you in an interview question, "As a junior engineer with a theater degree have you ever dealt with any difficult people?" And it's like have you met an actor before?

Corey: Right. Well, I managed to keep a straight face in response to that question. So, yeah that should do it.

Chloe: And that's why I'm such a huge advocate for these transferable skills. I have this background in theater, there's so many stage managers. Oh my gosh, stage managers could be some of the best PMs. If you've never done a show before, a stage manager essentially is doing the job of multiple product managers, like balancing all the schedules and the scripts and making sure everything goes well. So I'm a huge advocate for...

I'm working on a project right now that's a wine bot with a former sommelier who is now going to a bootcamp. And I've met botanists and principals and bringing those perspectives into the industry, it's just so valuable. Not only does it create better products, but it often goes hand-in-hand with diversity because, spoiler alert, most... I come from a... I'm a white, to give you a picture, paint you another picture. I'm a white woman from Sacramento, California. Middle class family.

I had computers around, but I never saw anybody who looked like me or acted like me doing anything with computers. There would be nothing to push me that way. So I think it's not only important to show people like, "Hey, here's what engineers can look like and here's what engineering can look like." But finding all these little niche pads that it seems that both you and I have found for ourselves and Cassidy has found for herself as well of almost this, it's like a tech entertainer, but it's in no way acting.

It's being a personality. But Deverel is such a weird field, right? Because we're ourselves, but we're also representing a product. And we also want it to be realistic and genuine. And I think that's why I lean more towards these more humorous, quirky, funny examples because if I took it too seriously, I don't think I could do it.

Corey: I am right there with you. I mean, I take a look through my news feed right now apparently the day that we're recording this is the first day of Microsoft Ignite. And there have been a bunch of announcements out of there and some of them are fascinating and others are just a little out there. For example, Microsoft pre-announced this morning Azure quantum, which is apparently a quantum cloud computing service that is going to be rolling out in the coming months.

Now, there's a lot of deep math in something like that that I'm sure someone smarter than I am could talk about. But my immediate knee jerk response to something like that is, "Great. How can I make fun of it?" Because if I'm making fun of it, well at least we're having a conversation about it. And generally speaking, I tend to put Microsoft into that bucket of companies that are large enough that if I make fun of them, they're probably going to be able to withstand my slings and arrows.

Chloe: So I got to know them. Where did your mind go with quantum? Was it some sort of like Avengers villain or something?

Corey: It feels like it's one of those new names for a technology of some sort. Maybe it's a new kind of CPU. Maybe it winds up being a throwback to quantum leap, but it was a great show and we could do a whole skit based upon nothing other than that. Only instead of jumping randomly from time to time to time trying to get back home, we're crashing into various meeting rooms trying to figure out where the hell we're supposed to be to have a conversation because nothing is labeled sensibly in this entire building.

Chloe: It sounds like a Hulu, Hulu show maybe. Yeah.

Corey: Yeah. It really does. Or taking a meeting at any one of our number of companies I'm not allowed to name.

Chloe: Exactly.

Corey: And other times it's, for example, sure I can talk about some of the upcoming stuff that Microsoft might be doing for example. But it's way more engaging if I start building protest signs and marching downtown outside of the Microsoft reactor urging Microsoft to turn it off before it melts down and kills us all.

Chloe: I know! What is in there? There's all that like...

Corey: It is glowing at weird times and strange noises coming out of it. And it's not just Microsoft, I think that it is unconscionable that AWS has launched the AWS global accelerator. It is doing terrible things for climate change. And anyone who knows what any of these things are looks at me as if I've blown several IQ points out the back of my head the last time I sneezed.

It's not based out of not having anything else to say. It's the getting people to do a double take and suddenly engage. You meet people where they are because I'm sorry, press releases inherently are so watered down that even the people writing them by the time they go out just don't care anymore.

Chloe: And I think such a great example of that is, well there's been all this talk lately that Microsoft has really changed, especially in the last couple of years. I think such a great example of this, which I just have a lot of pride in because this is just like a silly thing that I love, but this resurgence of Clippy, which kind of in a weird way started from me making very silly business cards, but there's a whole definitely Google it type in or Bing it, I should say I'm a Microsoft employee.

Go to bing.com and type in unauthorized autobiography of Clippy. It's about an hour long video by two folks at Andreessen Horowitz. And it was all inspired by the business cards that I made, but specifically around, you know, Clippy was viewed as such a failure back then.

And to give everyone context, I just turned 30, I don't come from a computer science background. I have a theater degree. So, as a 30 year old, newer to the industry, I had no concept that Clippy was a failure. In fact, I thought Clippy was a very, very cute nostalgic thing, which I'm coming to find out is a pretty common sentiment among people my age. We were not in the software engineering field at the time that this was viewed as... that was such a strange thing for me when I started tweeting. People were like, "Oh, this guy is back?"

Corey: I know he was great. He was taken before his time. Microsoft took him out behind the shed and Google retired him.

Chloe: Exactly. And I love Clippy and I view it even to this day. I mean, obviously as I sit here at my desk, I have a Clippy coin purse and a Clippy, re-foldable straw. I love that little guy and it's been really interesting to learn about the history of Microsoft while working at Microsoft because there's a lot of little nuances that I didn't know. I had to watch the developer's, developer's, developer's video because why would I have watched it before this time in my life? I had no awareness of it.

So it's a weird kind of catch up that I have to do not only from a technical standpoint but also from a pop culture standpoint of understanding who the heck these people are that get referenced on Twitter often and like especially when I started in the industry, I worked at a Docker CICD company called Code Fresh and I'm like, "What the heck is a Kubernete?" Like, I go, "Okay, how do you spell that? K8."

Corey: If you're going to answer that one, please let me know because I'm still trying to wrap my head around it.

Chloe: Same. Same. Same. Maybe I need that children's book. Maybe that'll help me a little bit.

Corey: I'm right there with you. I mean, I spent this morning for example, talking with my new best friend on Twitter, Microsoft Excel. I made a reference a couple of days ago and someone mentioned, "Oh, don't joke spreadsheet Twitter is small but passionate." I'm like, "Yes, but you shouldn't date from it. Never hook up where you VLOOKUP." And Microsoft Excel chimed in on the conversation and we're best friends now. We're going to hang out.

Chloe: I love it. And also, and mind you, I don't know who runs these accounts, personally since Microsoft's really big, but I want to say it was either the Windows dev account or one of those verified major Microsoft accounts on Halloween said something like have a spooky Halloween on the web and it had all these spider emojis. And I'm like 10 years ago in what world would we see an official Microsoft account tweeting like the Wendy's account, you know what I mean?

Corey: And you could see what that was like today just by looking at any large banks, Twitter account or investment firm where they have no personality of which they are aware, and I'm certain if whoever was controlling that account expressed one, they would be immediately fired.

Yeah, that's what people tended to think of. It's the old school marketing voice, for lack of a better term. Now it's engaging with people who care. And sometimes it works, sometimes it doesn't. As long as you don't wind up getting dragged by a social media mob, it's generally okay.

Chloe: Yeah. Yeah. It's an interesting time to be on Twitter as a brand and specifically as a tech company, I feel, because there's sort of this... And it's totally what we were talking about earlier about this idea of incorporating humor and having it be successful, which is a whole other science, right?

Like what is funny to a particular demographic. When you look at the demographics of the people who follow me on Twitter, it's mostly tech people and my tech jokes do not land on Facebook. I'm not really on Facebook anymore, but all my Facebook, Instagram followers are people from my theater life. And I tell I a [inaudible 00:24:03] joke and then obviously it lands very flat, but it's been really interesting to find the different humors of these different sides of things and to really channel that. To really make that the thing that helps sell the product or make the product more interesting.

It's exactly what you said. You go to these, you tag a brand and you're like, Norwegian airlines a few my plane was delayed and you get the autoresponder, it's going to be very, very different than a quick back at you. Especially with someone like you, Corey, who, who can respond really well and they're going to have a hard time if-

Corey: Well nowadays, I mean again, everyone... Well, I think I go back two years now and I had fewer than 1500 Twitter followers and it took me seven years to get there. And some of my early tweets since deleted were pretty much me yelling at various companies about perceived customer service failures, which everyone can enjoy. I was the worst kind of Twitter user.

Wait, I take that back. There are several worst kinds of Twitter user, but I was one of the unpleasant to listen to types of Twitter user, not the actual horrifying type of Twitter user.

Chloe: This was early on. Yeah, and a funny story, I don't know if I've ever told you this, but my boyfriend Ty is a Android dev and when we first started dating, he was working at Twitter. He was an android dev at Twitter. And in theater actresses for the most part, at least in musical theater in the Bay Area, we mainly used Facebook and Instagram for social media.

So I was like, Twitter's weird. Who uses Twitter? And flash forward to today and Ty my boyfriend's like, "Can you please get off Twitter? Can we have a conversation face-to-face?" And I'm like, "I got to just draft this joke real quick." So it's been interesting to not only learn how to, there's a whole like learning process of knowing how to interact with people.

I'm 30 and trying to learn TikTok right now, shout out to Tyranny and Julian, Emily for you know, trying to get this old lady to learn a new platform. But yeah, there's a whole subculture to every, you know, LinkedIn is a, well, it's its own platform in itself to engage hashtag.

Corey: Meanwhile, I come from an era where our social network of choice was communicating with one another via BGP route announcements, which is, that's an inside joke for some folks. So what was the face mask you used? That must have been part of that.

Chloe: It was a eucalyptus. Yeah, it's a whole skin regimen treatment. I'm off to do a blog post about it.

Corey: I do have to thank you as well. One of the nice things about knowing people such as yourself who have a background in theater are that whenever there's a... there are periodically certain referrals I can benefit from. And the one that's stuck most notably in my mind was the person you sent me to get my headshots done.

Chloe: Yes. Oh my gosh. He's my neighbor now. Yes. Ben Cramps.

Corey: And we'll throw a link to that in the show notes as well. So, when he suddenly wonders where all these ridiculous people are coming from, don't tell him, just show up and get head shots done.

Chloe: Thanks for doing my headshots. Since I started doing theater professionally in the Bay area and a funny story, there's a woman named Lauren who works at Microsoft. And when I was working at century we had a call together and we were the first two people on the call. And you know when you have a teams meeting or a Skype call, if you're not doing video chat it just shows your image, which is usually your headshot.

And we were waiting for other people to join and I said, "I'm so sorry I have such a random question for you. Are you an actress?' And she said, "How did you know? I'm an opera singer." I have such a radar for headshots and photos from theater performers and actors. Because it's so different.

There's this sort of like, it's not a smize but it's like, "I'm acting." So I'm so glad you're the perfect person to go get a headshot shot from Ben, Corey.

Corey: And it worked out super well for the first three quarters of the shoot. And then I was like, "Yeah, let's do one for fun." And I do the happy with my mouth open face. And Ben is a professional. Professional enough not to ask what is wrong with you? He just smiled, took the picture and it worked out super well.

Chloe: He was like, "Okay Crazy." Yeah, he's great. And also Easter egg about Ben, he's also a really amazing performer. I saw him in Drowsy Chaperone once. I used to only know him as a photographer and he was a producer when we did Jerry Springer the opera. And then I saw him on stage and he blew my mind with his voice. Another multi-talented individual working in several fields.

Corey: Yeah, he was a consummate professional. That's BenKrantz.com. We'll throw that into the show notes and see if we can get him a few extra gigs for people wanting to look amazing on Twitter profile pictures, on conference speaking circuits, or even on their dating profile.

So, I've done a couple of ridiculous music video style things recently where I'll write song parody lyrics. I have other people perform them. Sometimes I'll perform them and then my staff laughs and won't let that get into the light of day. But one of these days we should collaborate on something.

Chloe: Oh my gosh.

Corey: I think that there's a certain affinity for interesting music, making fun of things in tech and more or less just, I think both of us were born without that part of our brain that experiences shame. So, there's no problem with either one of us going out there and being actively ridiculous.

Chloe: Oh my gosh, yes. There are so many cool people who are doing specifically music parodied stuff. My coworker, Cassie and I for a while we really wanted to do this twitch stream where we would essentially put up a Microsoft learn lab and then work with our audience to come up with puny songs and lyrics and things like that. But the other day I was programming with my friend Kimberly. And she's newer to programming and I told her, "Oh, for the blog post, just use a GitHub gist."

And she was like, "Oh, what's that?" And I was like, "Oh here, just use this gist." And I recently while out to just give you an insight to how my brain works. While out to dinner with my boyfriend was eating Lobster Bisque and was singing This Bisque by Faith Hill. Like this bisque, this bisque, this bisque.

So obviously Kimberly and I were like, "Oh wow, we have to do a parody of this gist. So coming to the billboard charts soon, maybe a collab with Kimberly Corey. We need to also the CCK CPK. We'll think of a better band name. But there's just so much-

Corey: We certainly can.

Chloe: Cassie and I came up with if I could deploy instead of If I Were a Boy from Beyonce. There's just so much there. And I'm saying it out loud on this podcast so you guys can steal my ideas. But yeah, Waffle JS always has an open call for performers. So, I think we've got to get in there, Corey.

Corey: I really think there's another opportunity around the, if I could turn back time only about how to use Git properly.

Chloe: Oh my gosh. Oh my gosh, that's so good. In a full drag share look. Absolutely.

Corey: Ah, so if people want to learn more about the exciting life that you lead, what you're up to next, and of course get to find out the latest on your album before it drops, where can they find you?

Chloe: Yes, of course. Well, you can definitely follow me on Twitter. Just my name @ChloeCondon. I also have a website. If you go to Chloecondon.com. By the time this is live, I'll have a cute little website that links to all my videos and upcoming talks and some of the fun projects that I'm working on lately. And if you want to check out how to get started with Azure very simply and easily, you can go to aka.ms/screamingwithChloe.

Corey: Excellent. Thank you so much for taking the time to speak with me and of course for tolerating my ridiculous sense of humor.

Chloe: Well, thank you. And I'll probably be sending you a Twitter DM running a pun by you soon as per usual.

Corey: Yes, that doesn't tend to differentiate itself. We have so many of those back channel puns. Is this funny to anyone besides me? Don't care. Post in it.

Chloe: YOLO.

Corey: Chloe Condon, cloud advocate at Microsoft Azure. I'm cloud economist Corey Quinn and this is Screaming in the Cloud. If you've enjoyed this podcast, please leave it a five star review on iTunes. If you hated this podcast, please leave it a five star review on iTunes.

Announcer: This has been this week's episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com or wherever fine snark is solved.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Cody OgdenA product designer at heart, Cody has been crafting experiences for the web since he was ten years old. He’s best known for his open source website and its cheeky Twitter account, Killed by Google, which Fast Company called “an informational fever dream,” and one netizen praised as “an ignorant meme.” His project tracks news of Google’s product decisions until they are laid to rest in the Google Graveyard.

Cody works remotely as a software engineer at Cannabiz Media. He's a fan of hard cider, winters in Minnesota, and sees himself moving into a product design role at some point in the future.

Links

  • Twitter: @killedbygoogle
  • LinkedIn URL: https://linkedin.com/in/codyogden
  • Personal site: https://codyogden.com
  • Company site: https://killedbygoogle.com

Transcript
Corey: Hello and welcome to Screaming in the Cloud with your host, cloud economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

This episode is sponsored by AWS Solutions, which is the exact opposite of AWS Problems. AWS Solutions are vetted technical reference implementations that are designed to help you solve common problems and big faster. They themselves are free, though occasionally some of the products they stand up are not. But it's a great way to click a button, wind up receiving a technical solution that's implemented that ideally solves a problem you have. Visit snark.cloud/awssolutions. Again, that is snark.cloud/awssolutions. And my thanks to AWS for their generous sponsorship of this episode.

This episode is sponsored by AWS Solutions, which is the exact opposite of AWS problems. AWS Solutions are vetted technical reference implementations that are designed to help you solve common problems and build faster.

They themselves are free, though occasionally some of the products they stand up are not, but it's a great way to click a button, wind up receiving a technical solution that's implemented that ideally solves the problem you have. Visit snark.cloud/AWSsolutions. Again, that is snark.cloud/AWSsolutions. And my thanks to AWS for their generous sponsorship of this episode.

And this episode is sponsored by InfluxData. Influx is most well known for InfluxDB, which is a time series database if you need a time series database. Think Amazon Timestream except actually available for sale and has paying customers. To check out what they’re doing both with their SaaS offerings as well as their on premise offerings, that you can use yourself because they’re open source, visit InfluxData.com. My thanks to them for sponsoring this ridiculous podcast.

Welcome to Screaming in the Cloud. I am Corey Quinn. I'm joined this week by Cody Ogden who, among other things, is the Mortuary Assistant at Killed by Google, that is killedbygoogle.com. Cody, welcome to the show.

Cody: Thanks for having me, Corey.

Corey: Well, thank you for what you do. For those who have never heard of it though, let's start at the beginning. What is Killed by Google?

Cody: Killed by Google is an open source project that tracks just kind of the status of Google's product history, and keeps track of current products and products in the future that may end up being deprecated or killed.

Corey: It really drove home to me that when I first pulled the site up months and months ago, where it has this counted list of all the various services that Google has announced an end-of-life date for. Maybe one or two are questionable on this, but the last time I counted it was over 180, wasn't it?

Cody: Yes, I think we're up to 190 now.

Corey: Which means that Google has announced more products that they are end of lifeing then AWS has launched, period.

Cody: It seems to be true.

Corey: It's always interesting because Google has sort of built up this sort of reputation now, where whenever they announce an update on fill-in-a-service-here, the immediate response from the entire universe is, "Oh, crap, they're turning that one off too." Whenever I make that observation on Twitter, I don't know if you've seen any of it, but I periodically get, well actually get to death by a whole bunch of Google fans and oh, that's not true. It's only consumer stuff that gets turned off, it's for the better, yada, yada yada.

There's always the same lines that get trotted out, but I don't necessarily agree with the criticism in that it's building an expectation that when Google releases something, its days are inherently limited, so fun and jokes and poking large companies with sticks aside, where do you tend to fall on the Google-has-announced-a-thing-How-long-is-it-for-this world spectrum?

Cody: That's a great question. I think for me personally, I've grown much more skeptical about adopting especially a new software service that Google puts out into my daily routine, especially if it's something that's very like productivity oriented. Because if I get so used to having it in my life as a piece of me that really helps me do something in my life, there's no one to guarantee it's going to be available to me down the road.

Corey: The problem that I've always had is it does definitely color my opinion of the company. The notable starting point for a lot of this and most people's consciousness was Google Reader. It was for those who weren't around then or weren't paying attention or simply have better things to worry about than shaking your fist angrily at something years old now, Google Reader burst onto the scene and more or less eviscerated an entire nascent industry of RSS aggregators because it did it so much better than the rest of these and it was free.

You could just shove all your various RSS feeds into Google Reader and it was glorious. And in the meantime it also wound up put this eviscerating a decent portion of the market for companies that were already playing in this space. But these things happen. It was amazing and I recommended it to people left and right, and I used it more or less for my entire news outlet.

That was where all of the, I guess, things I cared about would show up. I would muscle memory, type it into any browser I was in idlely, when I was waiting for something else to show up. And then they announced that they were end-of-lifeing it. And there's some... I've heard different stories about why: some said that it was because it was running on a bunch of legacy systems, that once this was removed, they could finally turn them all off. Cool.

The other was that they were building Google Plus as their new social network and wanted to clear the way with anything else that vaguely resembled a social network. Awesome. And so they announced this about six months in advance or so, and people built a few replacements that are mostly there. I use Feedly; everyone else is using something different, but it's not the same. And that was the first indication I really had brought to my attention that suddenly this thing that I'd built my workflows around, this thing I built my digital life around wasn't guaranteed to be there in coming years.

So anything I started using online should have a backup plan and I started going down that path pretty easily and okay. Now, as a result, anything that I tend to use for the most part, I have at least an alternative, if not as good, will at least get the job done. And I do credit Google with helping teach that lesson though. I don't think that's the lesson they were trying to teach the world. What are your thoughts?

Cody: Yeah, my experience was similar. I was a Google Reader user myself, and when it disappeared, I was very upset. I had invested a lot of time and really fine-tuning my set up in that one piece of software, and when that service just wasn't available to me anymore, I had to go searching out for alternatives.

For me as a developer, I really leaned into things that are more self-hosted, so I spun up my own RSS aggregator that was like open source and you could self host yourself and I've been using that ever since just because I didn't want to take the chance that relying on another company that could be acquired or go under or just have their product shut down or make a pivot would affect me again with my fine-tuned set of news, things that I want to have access to in the future.

Corey: It's one of those stories where the old Baader Meinhof effect is the term for it where once you see something, you hear a new term or learn a new concept, suddenly you start seeing it everywhere and or the easy non tactical example would be, oh, if you're considering buying a particular make and model a car, suddenly you see them on the road everywhere when before you'd never notice them.

And once it started happening, I started seeing more and more products that Google had released or acquired suddenly being turned off with little to no warning and some of these mattered way less than others. One that I thought somewhat recently was interesting was their Hire by Google product. Do you care to tell that story?

Cody: Sure. Hire was a pretty interesting product from an external standpoint. They'd really paired it. They're ripe the set to create products that are focused on enterprise users, especially for that small to medium business range. You can't afford to invest in those, like a large applicant tracking systems that are already available.

And they are ready to disrupt the industry. So they launched Hire back in, I want to say like 2017 or 2018 maybe and they quickly got a lot of the early tech companies, a small startup tech companies onboard, which of course creates word of mouth. And then they decided to shut it down. They announced a one year for their SLA, they announced that they would shut it down a year from the 1st of September. So next year for 2020 it will be going away.

And I think that puts a lot of weird tastes in the mouth of people who may have been considering or still use Google's G Suite Platform because those were heavily integrated pieces of software because they really combined G Suite into hire to make it a product that really worked proactively for that hiring process. And now they're forcing all of these companies to finding an alternative to figure something else out. Go back to spreadsheets possibly.

Corey: Yeah. And it was not cheap either. This required some work for companies to integrate with. And once you went through the pain and hassle of doing this, suddenly whenever someone randomly would search your company name and then careers boom, right there were all your job postings at the top of Google's organic search results, which is incredibly powerful for how most people look for jobs these days.

But also kind of damning in that you're telling me that you can effectively get top placement above ads and still not find a way to make money out of this thing. And it seems like it really cuts against a lot of their own stated value propositions. And that in turn you have these companies who have paid the money for this, who've invested in it, and that really sort of takes the wind out of the counter argument of Google launches and kills consumer things all the time. And that's fine, but you absolutely can trust Google Cloud and anything enterprise because it's always going to be here. Well, that's sort of a thing that was aimed as solely at companies and look where we are.

Cody: Yeah, I agree and I think it's important to also realize that the people who follow Google product launches, they're probably not your average consumer. They're probably more tech minded or they work in an industry where they work adjacent to the industry. These are people that are going to be eventually making the decisions about what platforms to adopt, what services to buy into building that reputation, especially at the enterprise level, but also at that prosumer level. I don't know, it feels like that the opportunity is just being missed to build a decent reputation there.

Corey: It is because you have to increase the only disambiguate between what is the product name that starts with the word Google that I can trust and build a business around versus what product, starting with the word Google, would I be a complete moron to trust? We'll still be here five years from now. And the answer to that always presupposes that someone has an in depth knowledge of Google's organizational structure. I don't.

Cody: Yeah, same. It was interesting I was preparing for this, I was reading back through their letter to investors when they did their IPO. I found it really funny. One thing that they called out was that they want to think long term but then they quantified long term is three to five years.

And I found that interesting because the people who've looked through this lesson don't run some interesting calculations from a data they've found that the average for those products is about four years. So if one term is four years, like I just think there's a disconnect between what Google might consider long term as far as the tech world and what the rest of the world considers long term.

Corey: I’ve frequently said that multi-cloud is a stupid best practice and I stand by that. However, if your customers are in multiple clouds and you're a platform, you probably want to be where your customers are unless you enjoy turning down money. An example of that is InfluxData. InfluxData are the manufacturers of InfluxDB, a time series database that you'll use if you need a time series database. Check them out at influxdb.com.

Our biases tend to inform how we wind up looking at different tools and how they might factor into our own use cases. So rather than continuing to assume that everyone would use a time series database like I would, because I have assisted in background, what kind of customers do you have? What are they doing with a time series database that isn't, for example, just monitoring whether a computer is up or not?

Corey: Personally, if I take a look over the AWS side and do a comparison here, there's a very clear differentiating line where AWS will announce something that is absolutely ridiculous. Its Looney Tunes come to life and I may think it's ridiculous. I may think it has no market. I may not understand its target market, but I would not hesitate to build a business on top of it simply because they don't turn things off full stop.

And on the consumer side, sure, the Amazon fire phone for the longest time I would be getting boxes shipped to me by Amazon prime and they had the fire phone tape on it and it was always this moment of sheer terror. Oh, no. Has someone sent me a fire phone? I don't want one of those before they blissfully killed the thing.

They don't do that over in the AWS world. They'll call something classic like EC2 classic or ELB classic, which means it's not getting new features, but you can still use it. And simple DB was one of their first products. You can still use that you probably don't want to, but you can. So as a result, based on the 10 years and change of history with them releasing products, I wouldn't hesitate to believe that anything under the AWS umbrella is going to be something I can trust with my business. I don't have that certainty with Google on the list.

Cody: Yeah, I just think that timeline is such a weird disconnect for even consumers, but enterprise too like that, that cloud computing as well. Like that timeline of what is actually longterm is either not clear to people in their pitch or it's just fundamentally a culture difference between what they consider long term versus what everyone else would expect long term to be.

Corey: Oh, yeah and reputation matters so much in this space. People don't realize this, but AWS and GCP have the same equivalent contractual requirements around notice before deprecating a product. And that's fascinating when you're trying to do an Apples to Apples comparison. But here in the real world, they certainly don't have the same Mindshare equivalent thereof. I know theoretically, yes, Amazon could decide to turn off AWS in a year, but they're not doing that. I don't see any future where they would.

Cody: Yeah, I can't speak enough to AWS to really speak well to it, but there's definitely a difference how the reputation between those two really plays out.

Corey: Conversely, if you were to sit here and ask me, do I think that GCP is going to be turned off when Google loses interest? My response to that is, of course, no, I don't believe that's true. I think they're going to be in this for the long haul. But, imagine for a second that suddenly they did. Suddenly there's an announcement later today after we're done recording, where they're announcing a three year sunset period where GCP is getting turned off. Customers would be up in arms and screaming on Twitter because that's what customers do, but the response would be it was a Google product. What did you expect? And that right there is the problem.

Cody: Well, heartedly agree like that is the exact reaction would be, what did you expect? Just look at the history, look at how they've treated their products in the past. Look at how they just cut them out. They said, we're done. We're bored with it. We don't want to do this anymore. And, yeah, the reaction from all spheres would be the same thing. It's no surprise almost at this point.

Corey: At the time of this recording, they recently announced that they were acquiring Fitbit, which, okay, great. Fitbit has been sort of struggling for a while, but you want to place any bets as to how long Fitbit is going to be around. If someone's listened to this podcast in three years, will the term Fitbit mean anything to them?

Cody: Yeah, it's funny. If we look at the history of hardware acquisitions from them like Nest and I think before that they had to revolve, which were all internet of things, acquisitions the history doesn't look bright or at least the future doesn't look bright for Fitbit, both as a brand and as a product. And for consumers, they may end up with hardware that just doesn't work anymore.

Corey: And there was a similar story years ago where Sony installed. root kit on various CDs, audio CDs that people would buy that would effectively subvert their computers to make sure they weren't copying these things. And that caused a massive kerfuffle.

And I went a decade after that without buying Sony equipment and whenever I talked to people about that, they looked at me like I was nuts because well, that was clearly a radically different division of Sony than the one that makes cameras and laptops and I agree with them. They're right.

However, the reason that companies do these acquisitions and have all these different divisions is that that one division that has a particular product line generates good feelings toward that and companies want those good feelings to convey to other aspects. It's not just good feelings that do this. I almost wonder as a result, if there would be a consumer brand for Google and an enterprise brand for Google that they could launch to start differentiating this and get away from some of the, I guess user pain that folks have had every time something like inbox gets turned off.

Cody: Yeah, I definitely think that they have at least Google has already started this with how they've structured Alphabet. You're starting to see a lot less, recently they announced touringbird.com it was like a travel informational point of interest site that tried to help get you deals and information about places of interest.

They just announced that it was turned down. I had never heard of this before. I didn't even know it was a Google brand. Apparently it was under this umbrella of something called Area 120, which is another company that's under the umbrella of Alphabet, not of Google, but now that technology is being absorbed into Google. So, yeah, it's interesting the way they're structuring things now is almost to off you skate the impact or that or protect or create a barrier for that reputation so that they can't just keep contributing to that one brand that's most well known.

Corey: One thing that's strange to me is that I don't see people getting confused by Amazon doing the exact same thing where they have early versions of echoes and a bunch of devices that they put out and then wind up replacing and sun-setting as they do with many consumer electronics. The only time I've ever heard someone talk about, well, what does that mean for AWS? Inherently, comes down to someone trying to defend Google and making that point. But I see zero customer confusion about that.

Cody: I think the interesting thing that might contribute to that is just the approach to the customer experience that Amazon does versus Google does too. Even at a consumer level, Amazon has really focused on making sure they have great customer service.

You can reach a real person and actually talk to them. Amazon not necessarily AWS unless you want to pay a lot of money, but as just a normal every day buying an echo customer, I can get ahold of somebody Amazon. At Google, it's not been my experience with any Google product I've ever owned and especially not with any of their software services. It's always empty form that never gets a reply or a help article is all you can really find that doesn't actually solve your problem.

So I think that there might be that at least at a consumer level, this perceived idea that Amazon really does have that approach to customer service that makes the product more valuable because they provide that. Whereas Google, it's like, here's a product, that's all we got for you.

Corey: It also feels like Google has so much of their revenue coming from ads that anything that doesn't directly lead to ad sales means that it's just a hobby project for them. That they don't really see any downside to turning off on a whim when the person who championed it moved on or got promoted, it really does feel like there's something systemically strange with the culture.

Cody: It's funny that you mentioned that because I've read accounts from people who claim to be ex Googlers on sites like Hacker News on other people who've recounted these same stories that they feel like part of the issue with, especially the consumer product shut down is that the promotion cycle internally is so focused on launching and not maintaining.

So you launch a product and you get a promotion and you product and you get a promotion. So the person who is supposed to be like the lead advocate for whatever the product is. Like once, they get that promotion, they move on and there's no one else to take that helm. And this has been recounted in multiple different places that I've run online, but it's, yeah, it could definitely be at least partly driven by that internal promotion culture.

Corey: Yeah, it's really one of those unfortunate areas. We saw this at Google next where they're talking about things from all across the enterprise side of the business and then they put the Google voice logo up on the slide and the entire room just falls silent. And they mentioned the rolling into G Suite and there was a little bit of like scattered applause like, "Oh, we thought we'd get a better response to that." Yeah. "Because we thought you were about to kill it on stage in front of everyone." At this point, whatever, Google brings up a service that's been seeming neglected for a long time. It's never good news. I'm cautiously optimistic about this.

Cody: Yeah, I think Google voice is absolutely a great example of that. It got left behind. It had like one of the oldest iOS updates. I'm an Apple user, so I had one of the oldest iOS updates, had some really old design technique on it and had been updated for like almost a year or so before they finally gave it a refresh and then it just went dormant again and didn't receive any updates.

And then when they run it into G Suite, people were like honestly surprised. Like the momentum behind that pivot almost with Google voice coincided with them bringing project Fi out of project Fi and turning it into Google Fi. And it felt like Google had ignored Google voice for so long. And that's typically their MO.

They ignore the app until they're just like, finally, yeah, let's just be done with it and kill it. And I think that there's an opportunity there though, where they can start being more transparent with roadmaps around the products that they put on because that's not only going to make sense to, the pro-consumer levels who kind of worked in industry and understand the product life cycle, but also really be more transparent with consumers so if they can make more informed choices.

Corey: And that's really what it comes down to is just telling customers what to expect and then meeting those expectations. Now, the expectation is they release something. Who knows if it'll still be here in a week. In fact, a lot of that iOS refresh update wasn't even intentional. They were using internal certificates to do a few things with research groups that Apple didn't allow and revoke the certificate.

So they had to rebuild their applications to get them up and working again with the latest version of X code, which forced compatibility with modern iOS devices. So suddenly everything, the Google docs app suite of applications suddenly were fitting the screen on some of the newer iPads. And the fact that, that's what it takes to get them to update their stuff when it's just push the button and republish it is nuts to me. It feels like Google has been engaged in an ongoing war against its own users for far too long.

Cody: Well, And it's not necessarily that they're like anti user, it feels like there's not a culture emphasis on maintenance is just as important as new product. Maintaining a good product that we know people use isn't as important as creating something new or watching something new or getting something new out there to try to see what hits the wall and sticks. And I think that's a really... It's a difficult thing and because like you said earlier, the majority of their revenues coming from advertising and they have the privilege of almost not having to care because it doesn't undermine their bottom line.

Corey: It must be nice, but at some point public sentiment starts to turn against them and this is the biggest damage that they're doing. I mean, a lot of their uptake and their cool factor came from people who are early adopters who were driving their friends and family to use these things.

So in some cases, some extreme cases, you start seeing things such as you get your parents on board with something, you get your nontechnical relatives on board with other things, and suddenly Google takes it away. In box being a good example of this, what are the ads that you're going to recommend a Google product again to those folks? Because suddenly you're the jerk who got them into a thing that's now being turned off and ripped away from out from under them. It's not a great feeling to be one of Google's most devout fans when it increasingly makes you look like an idiot to your social circles for having believed in them.

Cody: Yeah, I think that goes back again to that the messaging piece towards, whoever is using their product, I said earlier, it's the tech is the people who are around the tech industry and they're interested in new technology that want to try the new stuff and want to really dig in and figure out how it works and what it can really do for them, right? But if they don't have a clear message on, hey, this was just an experiment, we don't think it's really going to be a longterm, they're going to make those recommendations. Other people get them involved, just like you said and then it's on them when something when they kill it.

Corey: And that's the interesting part is at some point you feels like we're all being Google's beta testers, for lack of a better term. They're even blurring the line now between, well, if you're not paying for it, you can't expect it to be around next month. Well, yeah, except the messaging has always been this is great, you should use it and then suddenly it's not.

Cody: But if we look at the history of even their most successful prospects like the Gmails, Gmail was in potent beta for years and it was never guaranteed to be an actual thing right now today. But how's that? I think the bigger question was has that like impression of what beta is changed throughout the years and I think that... Also as Google driving that impression of what beta actually means.

Corey: Funny, you mentioned Gmail about a year or so ago. There was an announcement about G Suite now is going to have a price hike from $5 per month per user to six and the internet lost its collective minds over this. But I was thrilled because I'd rather you charge me more for a service that costs you money to run than turning it off out from under me. Can you imagine having to replace Gmail for effectively every company that's been using it? And the how of it is, I'm not entirely sure I'd put it past Google at this point to try it.

Cody: Yeah, I think that... It was crazy. Like the first time they've raised that price for basically since G Suites inception and it's almost like counter to the argument a lot of people make about some of these Google products that get killed is, you weren't paying for it,

Corey: Right. Like that's somehow makes it okay.

Cody: Yeah. But you can't have it both ways. You can't always have the same price and never see that price increase even though you're getting a lot more value now out of the services that they're offering you as part of G Suite. So I don't know, it's such a silly thing to think about because people do bring up the argument all the time. Well, name me something that people paid for that Google killed and our people paid yet people, sorry, name me something that people paid for that Google ended up killing and you can just go down the list and you can find things that people paid pretty significant amounts of money to get their hands on and then Google turn around and shut it down.

Corey: Well, like the Google maps API changes where surprise your bill is now going to be 14 times what it was before for the same usage pattern that tends to break the implicit contract you have with customers.

Cody: Yeah, absolutely.

Corey: As we record this, it is November 4th, 2019 what do you think the next service to die is going to be?

Cody: I get this question a lot. I never liked to speculate too much because at the end of the day I want to be an optimist about it. Like I want the products that they create, they can really be foundationally changing for people and really improve their lives through technology to be great to work out for them. But at this point I'm not even sure.

Corey: You also don't want to give them ideas?

Cody: That too. I think that could be a very good take on it.

Corey: So have you gotten any feedback on the killedbygoogle.com site?

Cody: Feedback from?

Corey: Oh, the entire internet. Let's start there. I imagine Google would not go on record unless it was with a cease and desist, but what are they going to do? It's all true.

Cody: Yeah, there's been tons of feedback. I think I first described the feedback is that people take away what they want from killed by Google. They either are awestruck at, I can't believe Google has killed this many things. Some people visit, they want to just take a trip down memory lane. It's nothing they can do. These are cool things I got to use at one point.

Some people look at this and they're like, yeah, this is why we need better policies internally for our company about data retention. Like, how can we export things that people are doing on google's platforms to make sure we have an archive for our company or for our team or whatever.

And then some people have a lot of criticism. They're like, oh, like I said for, well, people didn't pay for these products, so it's Google's right to shut them down and they want to, and I accept that criticism. I think that's important more that we have that discussion about adopting technology into our lives that may not last forever and what that looks like for people from the consumer level to the enterprise level.

Corey: Yeah, it's easy to turn this into a roasting Google story, but in practice it does inspire longer term thinking around some of these things. I think that there's a strong balance that needs to be struck. There are a lot of services that people found near and dear to their hearts that aren't there anymore. They had to massively shift their own workflows and a timeline, not of their own choosing and that stuff leaves scars. Now I have to ask, I know it's hosted on GitHub, but are there any Google services that could be killed that suddenly you have to redo how Killedbygoogle.com is hosted?

Cody: Yes, potentially. So Killed by Google is hosted on netlist fly at the moment and they use a "Multi-cloud platform" which as far as I can see, they're using at least Google Cloud top form for a lot of the static sites that they host. So my site is literally hosted on a Google server, which I find kind of funny and ironic, but at least I've met with fide to back me up and moved me over to 80 worth AWS if I ever need to. It's a transaction.

Corey: I have a sneaking suspicion in some cases it's not a question of if but when.

Cody: Yeah, no, I can't say I disagree.

Corey: So other than killedbygoogle.com where can people find you if they want to hear what you have to say about this and other topics?

Cody: Well, they can find me on Twitter @killedbygoogle. They can also find me on my own personal website, codyogden.com

Corey: And we'll throw up links to both of those in the show notes. Cody, thank you so much for taking the time to speak with me today.

Cody: Yeah, absolutely. Thank you for having me.

Corey: Of course. Cody Ogden, mortuary assistant at @killedbygoogle. I'm Corey Quinn and this is Screaming in the Cloud. If you've enjoyed this episode, please leave a five star review on iTunes. If you've hated this episode, please leave a five star review on iTunes.

Corey: This has been this week's episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com or wherever fine snark is sold. This has been a HumblePod. Production. Stay humble.

View Details

About Russ Savage
Russ Savage is a Product Manager at InfluxData where he focuses on enabling DevOps for teams using InfluxDB and the TICK Stack. He has a background in computer engineering and has been focused on various aspects of enterprise data for the past 10 years. Russ has previously worked at Cask Data, Elastic, Box, and Amazon. When Russ is not working at InfluxData, he can be seen speeding down the slopes on a pair of skis.

Links Referenced

  • russ@influxdata.com
  • https://www.linkedin.com/in/russellsavage/
  • https://www.influxdata.com/

Transcript

Announcer: Hello and welcome to Screaming in the Cloud, with your host cloud economist, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Russ Savage, the director of product management at InfluxData, which I am under no circumstances to refer to as Quinnflux Data or Russ will presumably hit me with a belt. Russ, welcome to the show.

Russ: Thanks Corey. Glad to be here.

Corey: About a month or so ago at this point we had, effectively one of your co-founders, Paul on the show and talked for about half an hour about some of the ins and outs of Influx and how one might use a time series database, what a time series database might be. And that got a decent enough reception that we decided, "Hey, who else can we have on?" So congratulations. You're in the hot seat. But unlike some software products, people generally don't emerge into the world fully formed like from the forehead of some ancient God. How long have you been at Influx and where were you before?

Russ: Yeah, I've been at Influx for over two years now. I've been in the data space for a long time. I came from a small Hadoop startup that got acquired by Google and then I was previously at Elasticsearch.

Corey: That sounds like an interesting and winding road. It sounds like you did fundamentally what everyone was trying to do now and get out of Hadoop, and you were able to do it by changing companies. Some folks, not so lucky. But from my naive understanding, I mean for God's sake, my favorite database is Route 53 so I have other problems that I have to work through. But Hadoop seemed like it was a big thing and then suddenly it wasn't a thing anymore. And I don't see new Hadoop projects. They all feel somewhat legacy. Does that match your experience or am I just talking about things I don't understand or quite possibly both?

Russ: Well, I think my personal opinion is I think Hadoop solves Hadoop scale problems really well. I don't know if as many companies have Hadoop scale problems as they think. What's happening is computers are becoming more powerful. Individual computers are becoming more powerful. So you can do more with less. And so I think for a ton of the problems that people are facing now, you can run that on much smaller systems. You don't need a 50 or a 500 node Hadoop cluster to do that. Also, the complexity of setting up and maintaining Hadoop was very high. I think with these smaller systems, it's much easier to get them set up and keep them running a lot less maintenance overhead than you have with a Hadoop cluster.

Corey: It always was interesting watching people come out with big data problems and "Oh great, so where's your data?" And they pull a thumb drive out of their pocket, and you don't have a big data problem. Alternately, you can create big data problems if you're a terrible enough programmer. I mean my only debug strategy is print lines. So my logs are enormous because I never go back and fix anything. So every time someone has a request, I get 8,000 lines of logs. So yeah, there's always something people can do to turn it into a big data problem.

Russ: Yeah, I think a lot of those problems stem, they actually turn into larger than Excel problems or larger than what I can run on my local machine problems. But that doesn't mean that they're Hadoop problems.

Corey: That seems to be the point that the world has reached consensus around as well. So you wound up leaving a Hadoop startup and finding yourself effectively selling open source time series databases, which step one, selling open source software has always been a bit of an interesting challenge for some folks. Can we talk a little bit about that, and I guess first, what inspired you to come to InfluxData? And secondly, how are you, I guess engaging with the community when you do have a service or a software that you are attempting successfully, it turns out, to sell?

Russ: Yeah, great question. So my background is as a developer, and I love to write code, not production code, but a lot of code regardless. And I think my passion for open source, I love the community that evolves around open source projects. And so when I was looking for a new role, I was looking specifically for companies that were strong in the open-source communities. And so InfluxData is a great, great company for that.

So we have a set of products, open-source products that we use as the basis for our cloud and enterprise offerings. I really love the fact that there's so much capability in our open-source tool that the individual developer or the small team can really, it's not just demoware it's not just shareware. You can actually get real workloads done in the open-source tool. And then when you're getting more value out of open source or you need some support or you want to create a team dedicated to managing it, we can help you do that easier.

Corey: And I do want to point out, otherwise Lord knows, I will get letters, that as you're mentioning is I pull up the InfluxDB open-source repository on GitHub or GitHub, depending upon pronunciation choice and yes indeed it is licensed under the MIT license, an actual open source license. None of these nonsense "Oh you can use the source for anything unless you make money on it. And then we're coming for you" as some companies in this space recently have been doing as a, I guess a hedge against cloud providers. It's interesting that you folks A, haven't felt the need to do that. And B, have continued to engage in good faith with your own user community to work with people who are very clearly passionate fans of what you're doing.

Russ: Yeah, I agree. And I really, I applaud our leadership for doing that. Paul Dix, our CTO, is very passionate about open source and very passionate about licensing and he tends to write a lot of blogs on that topic. That was also really important that we're a true open source company, not one in name only.

Corey: I really think that can't be stressed enough. If I were going to be building something to give to the world or put out there as a product, I have a lot of decisions to make. And on the one hand, there is some compelling value to having something to be open source, but that does come with costs. And I think these ideas of just try to get the good parts of the open-source world - community engagement, free promotion, et cetera, et cetera - but not any of the downsides of well no one can ever offer this as a service except us, it just rings hollow and feels transparent. I mean, I'm sympathetic to business model challenges, but suddenly moving the goalposts on existing communities really breaks the social contract that open source communities have come to expect.

Russ: Yeah, I agree. And the problem that I see, so I've been on the other side and I've done evaluations of different software and looked at... Long term, you're thinking, you're bringing in third-party software into your company. Your company is going to be around for a long time and you want to make sure that you're not bringing in something that you're going to have to replace in a few years. And I think the real key with open source is you can take that technology and really embrace it and really take it to the next level, really expand upon it.

And when you're bringing open source software into your company, you're thinking about the long term horizon. And if you have companies that are making changes or deciding to suddenly change the way their licensing or change that, it really calls into jeopardy your long term vision and it's a risk.

Corey: I think that understanding that open source is a double-edged sword in some ways and it is a approach to solving certain problems, but it's not a strategy in itself. It is something that a number of somewhat naive founders have gotten trapped in in the past. Of course, here we are now where you're very clearly doing well. I've talked to a number of Influx customers that we'll talk about in a bit, but it's very clear that you're doing something right. And believe me, when I have people on this show, one of the first things I do is start Googling, all right, who hates them? And I can't really find anyone saying negative things about Influx. Believe me, I've looked. Am I just looking in the wrong places or does the community actually seem to like you folks?

Russ: Well, I talk to the community every day and I have a ton of positive interactions with them and we're always working to bring them closer and bring them into our discussions as we move as a company. And so I think that community engagement is really important, and it also drives a ton of transparency in the organization. Right. So Telegraf is one of our projects. It's a data collection agent. And all of our discussion happens in the public with our community. I think that's really important and it makes people want to be a part of that community. And we want people to want to be a part of our community and that's what we could strive for.

Corey: That tends to be a hot button that gets a lot of people riled up. Let's change gears a bit and talk about something else that I've been using to successfully annoy people for a couple of years now. I have been a long fan of the idea that multi-cloud as a best practice, is stupid. If you're building something and you want to be able to magically deploy it to AWS and GCP and Azure and Oracle Cloud because you lost a bet somewhere, then you're slowing yourself down and effectively trading feature velocity for a level of faux agnosticism that you're not necessarily going to ever take advantage of.

And I stand by that statement, but what people then don't listen to is paragraph two after that, which is, here's a list of exceptions. First and foremost among that list of exceptions is who are your customers? If your customers are attempting to service their own customers in a variety of different locations, they have to meet them where they are.

"Our service is awesome, but you have to move over to this cloud provider" is absolutely not going to happen. You're not going to be able to serve their needs. And in turn, telling customers that they should migrate their own work, their own stuff and become multi-cloud themselves, if they listened to previous episodes of this podcast, they're going to say, "Oh wait, we've heard about this. It's stupid."

Now because of you being who you are and having to service a wide variety of different customers, you fundamentally need to have an offering that spans beyond a single provider.

Russ: No, I completely agree with you. I think one of the things that you look at is we are a platform company. And we want our users to build on our platform. And when you're running a platform company and you want users to build on top of it, you need to make that platform available where your users are. And the reality is is that different companies use different clouds for different reasons. I don't think that one company is going to... It makes sense for one company to run in multiple clouds, in most cases. We have some customers who are running in Azure Cloud for specific reasons that we're not really... We don't really care what those reasons are. We just know that they can only run in a specific cloud and that's where we want to be. And so, yeah, we 100% want to run our infrastructure in as many clouds as our customers need.

Corey: And that's sort of a piece that often gets lost in nuanced conversations about this. When customers have needs, you've got to be able to meet those needs rather than condescendingly telling them that they're wrong all the time, unlike certain cloud companies we can all think of, but not name. The fun part where that also expands beyond just the idea of multiple clouds, it's easy to also sit here and say, "Oh, and hybrid is dumb, too. You shouldn't have any on-premise physical data centers." Oh, that's adorable. Great theory. But here in practice, we have this thing called legacy, which engineers here is old and broken, but here in the real world I hear as things that make money.

So I don't see hybrid going away anytime soon. And unlike a lot of different platform companies we could name, you have viable options for folks who are running on-prem. Was that something you planned from day one? Did that accidentally happen along the way as you were building this stuff out?

Russ: Yeah. So we have on-prem and we have cloud. And I think the interesting thing there is a lot of companies will tell you that they're all cloud. But if you actually dig down into their engineering organizations or dig down into the development areas, you'll actually find a ton of quote unquote “on-prem development.” We’ve got people writing applications against on-prem versions of software and they want to make sure that those applications run, whether it's on-prem, whether it's in the cloud, wherever.

And so I think one of the big things that we've focused on is, again, the idea of a platform and being able to develop against that platform and having a common set of APIs, whether you're running this stuff on-prem and you're building out development as a proof of concept to make sure that it's working for you, or you're running this on a cloud and you want to service a global customer base, or if you're an IoT shop and you're running in an oil rig in the middle of the ocean and you don't always have cloud connectivity. I think, again, it's another question of where are our customers running their infrastructure and since we're an infrastructure company we want to be where they are.

Corey: But something you just mentioned is fascinating. The idea of oil rigs for instance. People tend to not necessarily have a lot of exposure to that type of environment, but to my understanding, a lot of them tend to live in the middle of the ocean where internet connectivity, shall we say, comes at a premium. And having something cloud-hosted just simply isn't an option there for anything that needs to be even slightly performant. So there really is no alternative short of asking AWS to build a region in your house, effectively. I've asked, they won't. US Bathtub One will not exist this year, but we're optimistic for 2020.

There are use cases that the cloud cannot, with current technology, address. And I think that there's this certain willingness of the part of the cloud native sort to turn our noses up at that type of use case. But things like that are important and customers are there with problems. Making sure that you can address what those customers are up to and meet them where they are is something that I think an awful lot of companies just sort of tiptoe past and don't really investigate in the name of ideological purity.

Russ: Yeah. I think you're right. And I think we see at least in our customer base, basically where the application or where the infrastructure is deployed, kind of serves for different use cases. And so you look at running something on the edge or running something on-prem, on an oil rig for example, right, you're basically... Who's your customer in that regard? It's the individual operator of the system on the platform looking at real time data of how different gauges, how different pressures are operating. Right? And that's valuable and it gives you a lot of capabilities and insights.

But in order to get long term trends or in order to do large scale analysis across hundreds of platforms all over the world, you want to send that data into the cloud for processing. And that data that's in the cloud is going to be massive aggregated data across many different platforms. And the audience for that is going to be a data scientist who's building out a machine learning algorithm that can then be pushed down into the oil rig. Right? And so you get this system where you're actually, you're running your software, you're running the same software, but you're addressing two different needs and two different business needs.

Corey: I guess I have the advantage on some folks because though I've never worked on an oil rig, I did grow up in rural Maine. And you have about the same level of internet connectivity in some of those places as you would on an oil rig. So running things in my living room was always what I did growing up. There was always the one room that was 20 degrees hotter than anywhere else because that was my makeshift data center with crappy old computers. And it was terrible, but it was also a half step above IBM cloud, so there's that.

And eventually the world modernized. Internet came to rural Maine. But I remember those days where you're generating more data locally than you're ever going to be able to shove into, well, however large a pipe is to get it to a cloud provider. So, especially when you're in the business of doing data collection and aggregation, which fundamentally is what a lot of databases are, whether they're time series or not, it needs to be as close to what's generating those inputs as possible, I'd imagine.

Russ: Yeah, and especially true for time series data where time is important and latencies are important and milliseconds matter, right? And so we get into a scenario... Our database is one of the few that can do at the nanosecond level. And while we're not recommending that people monitor everything at the nanosecond level, it does help if you can trigger an event and then you start collecting data a lot more frequently than you normally would so that you can see exactly what's going on during that event. And then you go back to one second intervals after it's finished.

And so I think putting your database close to where you're located and reducing that latency gives you an advantage in your response time to the specific problem that comes up.

Corey: Oh, absolutely. Anyone who wants to experience this themselves can wind up trying something at an AWS account, and then just start a stopwatch and see how long it takes that event to show up and either CloudWatch Logs or far worse Amazon CloudTrail. They put the eventual in eventual consistency. There's value to getting rapid response that could inform what's going on. I mean there are use cases where that simply won't do. Imagine having significant latency and a self-driving car for example. At that point, oh, you should have stopped at that light four lights ago. Not a great plan.

Russ: Right, yet, the latencies on on-prem and when you're close to the source are very different than the latencies than in cloud. And so again, it's different use cases that you're solving and it's different dimensions where you're analyzing the data.

Corey: Our biases tend to inform how we wind up looking at different tools and how they might factor into our own use cases. So rather than continuing to assume that everyone would use a time series database like I would, because I have a SysAdmin background, what kind of customers do you have? What are they doing with a time series database that isn't, for example, just monitoring whether a computer is up or not?

Russ: Yeah, great question. So to be clear, a lot of our customers are monitoring whether a computer is up or not. That's just a huge use case for us, and-

Corey: It seems a common pain point.

Russ: Yeah, it turns out it's hard to keep those things running. No. So, we have a ton of use cases around time. And our belief is if you actually look at the data that's being collected and the data that's out there, most of the data that you're seeing is actually more valuable when you look at it through the lens of time and you look at it, how things trend over time.

So even if you think of traditional data stores, if you add a time element to it, you suddenly start getting more insights than you could without. And we've got use cases everywhere. So for example, we were talking about oil platforms earlier, right? We have a company named Equinor who is using our platform for monitoring those oil platforms in the Norwegian continental shelf. And it's really exciting and really interesting to hear some of the stuff that comes out of there. But basically the ability to gather that data in real time and make decisions on it was really important for them.

And when you think about... We consider this an IoT use case. A lot of people when they think about IoT, they think about you're looking at temperature sensors or you're looking at different air quality measures. But a ton of our industrial IoT customers are monitoring large, large, massive oil rigs or drilling platforms or things you wouldn't typically come to the top of your mind when you're talking about IoT, but very, very important use cases.

Corey: So what use cases do you have that are outside of the IoT space? I mean as much fun as it is to talk about A, oil rigs and B, refrigerators that tell me when the milk is expiring, what other use cases do you see from customers?

Russ: Yeah, IoT use cases are really fun to talk about because they're so varied and widely used. But we have people using us for a ton of different things. So as I mentioned before, we have customers that are managing data centers with tens of thousands of machines. And we even work with NASA who's monitoring their infrastructure for launching satellites.

Corey: And at least until it launches, it presumably has a better connectivity to the internet than it does once it's in space. Although, one starts to wonder.

Russ: Exactly. I never know, you never know what Elon's up to.

Corey: The problem I run into always is that I tend to think of these things just in the context of what I've used time series databases for before. Now, the challenge of course is I was a network admin once upon a time when the closest thing we had to a time series database was RRDtool or MRTG that wrapped RRDtool. And as a result, oh time series databases. Those are those things that are terrible and cause all matter of problems and its primary use case is to be embedded in cacti, whose intern primary use case is to sit in the corner until someone breaks into it and uses it as an attack platform. That may not be entirely hypothetical.

But increasingly, the capabilities have dramatically changed. And being able to see such, I guess, different capabilities now, I look back at the outages of yesteryear, for lack of a better term, and it would have been so much easier to diagnose some of these outages with the right tooling, as it turns out. Unfortunately, it is a time series database is not a time traveling database, so I don't think you can go back and necessarily help me with those problems in the past.

But for other people who I guess who are stuck in that outmoded view of time series, what's changed? How would you describe the advancements? Because I'm sort of going to assume just on a lark here that you didn't declare InfluxDB feature complete in 2013 and the rest has just been maintenance releases.

Russ: No, definitely not. I think the key fundamental difference between a generic general database and a time series database is you can turn a generic database into a time series database with enough manpower, with a big enough team. I think the question-

Corey: Oh, with a dumb enough perspective, you can turn DNS into a database as well. I've done it. But everyone looks at you with this horrified look on their face and then you get asked to leave the interview immediately.

Russ: Right, right, exactly. It's basically, it's about what you ultimately want to accomplish and where you want to put your resources. Right? Everything, and I'm a product manager, so I'm always thinking about resources and priority. But if you're a company and you are solving a specific business need, does it make sense to actually spend 50% of your time taking a generic database and turning it into a time series database so then you can solve your problem? Or would it make sense to go out and get a best-in-class time series database so that you can focus on the thing that's actually going to make to drive capabilities of your business? Right?

And so when I think of the reasons why you would choose a time series database, behind the scenes, we've done a ton of legwork and a ton of the optimizations that you would have to do yourself if you were returning a generic database into time series. And so that gives you a great starting point and allows you to focus on the business problem that you want to solve instead of these infrastructure issues with creating a time series database.

Corey: What's interesting to me, if I take a look through the Influx suite of tooling, for lack of a better term, you have a hosted version that people can run wherever, you have an enterprise offering, you have cloud-hosted versions. What is the decision matrix for people who are trying to figure out which version would make sense?

Russ: Great question. So we talk to our customers a lot. And basically we kind of break it down into two different paths. Between those paths are different decision points that you can make depending on what you want to solve to cross the divide, right? And so on the on-prem scenario, you've got the open-source instance that you can run locally and build your application against. You have an enterprise instance of on-prem that you can host yourself if you need more power or high availability or security. Then we also have the equivalent in cloud. So in cloud, we have a free tier that lets you get started quickly and build your application against. And then if you want to host all your infrastructure in cloud, we then offer that pay-as-you-go service in cloud.

And so basically it depends on your use case. It depends on your infrastructure. It depends on how you want to run things. But either way, whether you're on-prem and want to manage all the infrastructure yourself or you don't want to manage any infrastructure and you want us to do that for you, we have offerings for both.

Corey: One of the, I guess, surest signs of a fanatic is when they have absolutely nothing negative whatsoever to say in any context about a given topic. You don't strike me as someone who falls into that category. So I have to ask, from a high level, what unsolved problems are there today in the world of time series databases?

Russ: Oh, that's a tough one. Unsolved problems. So the thing that I see when I talk to customers is, and you're starting to see this a ton in the marketplace, is I think everybody had this dream and bringing this back to Hadoop is this like this data lake, which actually turned out to be a data swamp as people realized, but you've got these specialized databases and I 100% agree that you should use the right tool for the job. And so I think the specialized databases make a ton of sense.

But I think the issue that you're seeing is you're actually starting to create these individual silos across different storage mechanism, different databases, different tools. And one of the struggles that people have is connecting all of that stuff together, right?

And so you see a ton of companies coming up that are basically just connecter companies that bring data from A to B and C to B and all of these things together. And so I think that's a huge struggle in specialized databases and in time series. And so that's something where you'll see a lot of development in our data manipulation language called Flux that actually will make it much easier to bring that data from those disparate databases, from those disparate locations and run analysis that you couldn't really do before.

We see a ton of people who in the past have basically had to build applications on top of multiple platforms to bring data from all of those platforms and combine them, and our goal is to bring that application logical closer to the data store and so that you can then leverage the capabilities of the community and not have to build everything yourself. And so we're starting to see areas in Flux where people are connecting to different data sources, bringing that information that has typically been siloed in individual tools into the same place and running that analysis.

Corey: I have some level of contempt for companies that are entirely built around solving a particular pain point or a problem that are effectively one feature release away from another service or product that puts them completely out of business from a model perspective. I'm not going to necessarily name names, but if your entire company's job is to take data in one format and translate it to another, if the source company releases one day of feature of, "Oh by the way, you can now query this and get a JsonResult instead and that destroys your company," maybe when you're raising giant piles of VC money, thinking through how unassailable of a moat you've built really winds up being worth pursuing.

It definitely seems, from conversations I had with you and with others, that this is a serious... It's not a problem that you've built things around. It's an entire field of inquiry where you might have heard of time series database and are scratching the surface a little bit, but the further you get into it, the more you realize there's an entire sector here. There's an entire ecosystem that is dealing with painful problems.

I mean, we've gotten to a point now where, I wouldn't have believed this was possible, you can commit the Cardinal sin of posting a graph on the internet and not label your axes, but you can generally assume with a time series driven graph that X is time and it moves from left to right. That alone is an industry win. Congratulations. You have finally gotten statisticians and high school math teachers to stop screaming about that one. That alone says that there's a lot more to it than that.

I think my rubric for determining whether something's important or not based upon a 10th grade experience I had might be slightly flawed, but I'd also think that this is something that is starting to rise in awareness. People are starting to understand that there's something there. And again, I keep looking for dirt on you people and I'm having a lot of trouble finding it.

Russ: I appreciate it. I think the one thing that I like to focus on again, I really love the community that we've built over the years and the trust that we've built in that community. And so I think if you're honest and you're doing good work and you're being transparent about the work you're doing, then it'll attract good people. And so I'm again going back to our original conversation, that's one of the main reasons I came here in the first place, and it's the reason I'm still here now, so.

Corey: Well thank you so much for taking the time to speak with me. If people want to learn more about what you're up to, where can they find you?

Russ: If you want to learn more about InfluxData, influxdata.com. We've got a free signup for our Cloud 2 offering. You can take it for a spin and see if it meets your needs, and that's a great place to learn more.

Corey: Terrific. Thanks again for taking the time to speak with me today. I appreciate it.

Russ: Thanks, Corey. Thanks for having me.

Corey: Russ Savage, Director of Product Management at InfluxData. I'm Corey Quinn, and this is Screaming in the Cloud. If you've enjoyed this episode, please leave it a five-star review in Apple Podcasts. If you hated this episode, please leave it a five-star review in Apple Podcasts and in the comments, ask Russ to hit me with his belt.

Announcer: This has been this week's episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com or wherever fine snark is sold.

Announcer: This has been a HumblePod Production. Stay humble.

View Details

About Tiffany Farriss
Tiffany is the CEO and co-owner of Palantir.net. Along with George DeMet, she provides the vision and values for Palantir. She has over 20 years of internet consulting and development experience and extensive experience providing information architecture and usability consulting for a wide variety of clients. Tiffany has a BA in Mathematics from Northwestern University, where she focused on mathematical modeling and human-computer interaction, and was a member of the Drupal Association Board from 2009–2017.

Links Referenced

  • Sponsors
    • AWS Solutions
    • Influx Data
  • Twitter: @farriss
  • LinkedIn URL: www.linkedin.com/in/tiffanyfarriss
  • Company site: Palantir.net

Transcript
Announcer: Hello, and welcome to Screaming in the Cloud with your host, cloud economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored by AWS Solutions, which is the exact opposite of AWS Problems. AWS Solutions are vetted technical reference implementations that are designed to help you solve common problems and big faster. They themselves are free, though occasionally some of the products they stand up are not. But it's a great way to click a button, wind up receiving a technical solution that's implemented that ideally solves a problem you have. Visit snark.cloud/awssolutions. Again, that is snark.cloud/awssolutions. And my thanks to AWS for their generous sponsorship of this episode.

And this episode is sponsored by InfluxData. Influx is most well known for InfluxDB, which is a time series database if you need a time series database. Think Amazon Timestream except actually available for sale and has paying customers. To check out what they’re doing both with their SaaS offerings as well as their on premise offerings, that you can use yourself because they’re open source, visit InfluxData.com. My thanks to them for sponsoring this ridiculous podcast.

Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Tiffany Farriss. The CEO and co-owner of Palantir.net, also known as no, not that Palantir. Tiffany, thanks for joining me and welcome to the show.

Tiffany: Thank you so much for having me, Corey.

Corey: Of course. So, given that you do work for a company called Palantir, I think we have to first open with the question ... I will be torn to pieces if I don't ask. Where exactly does your company stand on putting children in cages?

Tiffany: We're firmly against putting children in cages.

Corey: That is a bold, political stance in this day and age. So given that you are not the monster that Thiel built, what does Palantir.net do?

Tiffany: We are a digital consultancy that uses open-source tools to help solve our clients problems. We work primarily with nonprofits, institutional nonprofits like higher education and healthcare as well as state and federal government.

Corey: It's easy to sit here and say, "Well, that's sort of a derivative name. Why would you name your company after an existing company that already exists?", until you realize that your company's been around for 23 years.

Tiffany: We're very proud of that. Yes. 23 years. Well, and the name is a throwback, right? So for those Tolkien fans out there, 23 years ago, if you think back, it was all about the information super highway and when my founder, George DeMet kind of named the company, he thought that was a terrible metaphor and that really the promise and the potential of both the Web and the Internet was as a nodal communication device. So thinking about it as more of the network of Palantiri that existed in each of the cities of the realm was a much better way to approach it.

Corey: Yeah. We've turned into something of a dystopian future where it has much more ominous overtones than it did. I don't think looking back to the late 90s that any of us saw the Internet and culture surrounding it evolving the way that it has. It never occurred to me that it would be a commonplace, day-to-day thing that everyone knew about and that it would have dramatic social implications.

Tiffany: I think if you go back and read some Neal Stephenson, I think there were some people who really did anticipate it. I think Snow Crash kind of got some hints of it, and even Ender's Game had a lot of overtones of the kind of discourse and disinformation potential in the Internet. So I think there were people ... there were harbingers out there. But I certainly didn't see it coming like this.

Corey: So on a happier note, on a happier story, I suppose, let's talk a little bit about things that are, I guess, more cloud adjacent than talking directly about the cloud itself. Something you've been interested in for awhile and talking about in different fora have been sort of the issues that cloud has created for open-source, especially recently with some of the moves Amazon and others have made. And the open-source ecosystem goes back far beyond cloud was ever a thing, and watching the evolution of that ecosystem has been fascinating from both the inside and the outside.

Where do you stand on it?

Tiffany: I think it has been such an amazing journey. I think that the cloud has opened up a lot of opportunities for open-source. But, at its very core, the GNU Public License, the GPL, is fundamentally a distribution license. And I think that the cloud introduced the possibility of creating software for others, but not distributing it. And I think that we're still wrestling for that soul of open-source in that context. What does it mean that you can take open-source, create derivative works on it that you sell to others and then not have to contribute back to the community whatever innovations you've added to it. I think that's the interesting challenge of our time.

Corey: And that's sort of the interesting question that sort of ties all of this together to some extent. At what point is someone obligated to give back to an open-source project? I mean, in an absolutist term by reading of the license, you're not. There's no requirement that you give back other than if you make changes to the source code you're required to make those changes available and republish it for some but not all licenses. But there's never been an obligation of, oh, if you take this and turn it into a product, then you're obligated to buy a different tier of licensing or the new licenses that do speak to that are definitionally not open-source as determined by the OSI.

Tiffany: Correct. Yes. I think that there are several different levels of obligation. You have your ethical obligation. You then have a self-interested obligation. And I'm far more interested in talking about the self-interest side. I'm a bigger believer in carrots than sticks, and I think when we start going down the ethical path, it tends to feel like you're bashing people. And I've had the best success in trying to align the interests of the open-source communities that we work with at large with the interests of the ecosystem that relies on them. And really making sure that there’s sharp relief about the alignment between what the business interests are and what the community's interests are. And I think we've done a fairly good job of that in Drupal so far, but I think all communities are somewhat vulnerable to this.

And the sooner that open-source communities take seriously how they influence their own economies, essentially, I think that the better off we're going to be.

Corey: And you talk about self-interest from a perspective of doing what's effectively what is best for a company. And so the question naturally arises, if I'm an Amazon or another large cloud provider, how is contributing back to open-source in my self-interest in a direct sense?

Tiffany: I think for those hosting providers that choose to offer as part of their suite open-source products, either there's two different levels, right? There's the open-source on which their cloud platforms are built, and then there are the services they may provide as kind of turnkey services or self-start services. And I think those are two different sets of incentives and alignments that are possible. For those companies that have built their product, their cloud product on open-source, I think it's a little bit harder to be able to find the incentive, because you do have the question of competitive advantage in that way. And I think it falls to each of the providers to be able to decide how much of that bespoke software do they want to carry? It's a question of technical debt. It's a question of what advantages you gain by giving that back to the community versus your assessment of what potential harms you may have if your direct competitors were to also enjoy those advantages.

And a lot of that comes down to the culture of the open-source project as well and how they enforce the norms and what that particular project sees as its bar for participating and first-class citizenship. Then you have the case of the cloud providers that provide kind of self-service where they make it so that you can install Word processor, install Drupal in a very easy way. And I think we've seen perhaps better traction, at least in the communities that I'm closest to, with those hosting providers, cloud hosting providers, in making a lot of what the contributions that they make back to the community. And so I think that becomes a little bit easier because it's not about the product itself. It's really more adjacent. And the better a cloud provider's reputation in certain communities, the more likely they are to win business. So, that's what I mean when I talk about the communities need to think carefully.

They need to be quite thoughtful about what level of influence they have over the economies within their own communities. And what are the incentives that they have to offer, and what are the benefits that all of the companies in the ecosystem could realize and really find those pockets of the win-win.

Corey: One of the things that I think that hasn't really been discussed much, and I'd like to maybe go into that a bit with you now, is one of the, I guess powerful aspects of cloud has been that it's been lowering the barrier to entry. And that's a challenge in not just computing, but also the open-source movement have been struggling with for a long time. What's the easiest way to get someone on board? In my own case, I'm a terrible developer, as you can tell, by looking at any code I've ever put on GitHub for anything. So my involvement was for the better part of a decade, I was network staff on Freenode helping to provide a sense of fostering community in an IRC world. But even then, getting someone onto IRC in the first place was challenging and difficult and finicky and required a certain level of technical know-how, that until you figured that out, you weren't in a position to ask for further help.

A similar example that I think you mentioned to me at one point was Eternal September back on the Usenet days. Where every year when September hit, Usenet, an old school I guess message board for lack of a better term, roll with me on that one, would get a whole bunch of new users coming in as they went to university and had access to this in a way that you didn't when you were a home user. So it took time over the course of September into October, for the newcomers to understand the netiquette, for lack of a better term. The way you behave. The way you comport yourself. And then one year, AOL hooked up all of their users onto Usenet and that was known as the start of Eternal September.

Have I gotten the story mostly right?

Tiffany: Yes, you have.

Corey: So the interesting thing there, there's still is people that are claiming that today, for example as we record it, we think of this as October 23rd, 2019. But through the lens of Eternal September, it is now September 9,549th, 1993, because this is the September that never ends.

Tiffany: Painful thought.

Corey: Exactly. And that was an example of the Internet culture having to adjust. Having to adapt to a new normal. Suddenly, people who would never have had access to these things, did. And indelibly changed the culture. And that's been viewed as a sort of condescending way of approaching it, but cloud has done the same thing, too. Once upon a time in order to spin up a company or any random item, you'd need to spend hundreds or thousands of dollars on collocated equipment. On getting a whole bunch of things bolted together. And now it's an API call or a button click away. The barriers to entry have dramatically been lowered in terms of cloud, as well.

Do you think that there's, I guess, an alignment or that there's been even a direct relationship between that barrier lowering in open-source, as well?

Tiffany: Absolutely. I think that's what we experience in our open-source communities all the time. If you think back about how the Drupal Association was created, Dries Buytaert had founded the Drupal Project from his dorm room in Belgium and was hosting it and open-sourced it, and people started to use it. And then at some point the servers melted down and so he put out a call for help. And so they needed to buy more actual physical servers, so donations poured in from everywhere. He ends up with tens of thousands of dollars for the purposes of supporting this open-source project. This little nascent open-source project. He needs somewhere to put it so that he doesn't end up getting all the taxes himself on it and so he creates the Drupal Association.

And over the years, the Drupal associated was tasked with really supporting the Drupal community primarily through the maintenance of Drupal.org. We had physical, actual servers that hosted Drupal, that hosted the SVN, Subversion, code repository that we used. And even that started to create, I think, a barrier for people to participate in it. You had to learn SVN. It wasn't as accessible. You had to be able to run your own local environments. You had to be able to set those up. And as we have enjoyed the benefit of all of the cloud providers, you could spin up a Drupal site with a cloud provider really through clicking buttons now. So it's very, very different from the days of IRC and SVN, right? We have GitHub. We have Slack. And so there are people who can really stumble quite into the Drupal Slack. And again, they don't know really what's expected of them. And I love that term, netiquette. But I think as we look more broadly at open-source, it's really a question of citizenship. It's a question of what is expected of me? What can I expect from you? And what can we expect collectively from this environment that we're in?

And so I think that's one of the places where we have this opportunity now to look back at something like Eternal September where how did Usenet responded? Well, I mean in my experience and again, I'm kind of revealing my age at this point, but there were a lot more community managers. Community moderators, who kind of got involved and for better and for worse, at different times. And it started to create ... the expectation started to shift based on what they could do and how they could extend it. And I think that that's what we need to reckon with in open-source as well is that the barriers to using open source have dropped significantly and we have an opportunity to think differently about how we might impact the barriers to contributing back to open-source. All of these cloud services make it easier in some ways to contribute, but we need to create those incentives. We need to create those expectations. We need to create the mechanisms by which people who are newer, and who maybe are not indoctrinated in the culture, who maybe got there because the open-source product is the best product, not necessarily because it's the best open-source product. Which, I think, is quite honestly a victory for open-source.

But we have to recognize that these folks need to be successful onboarded, not just into adoption but through adoption into the contribution journey, moving them along the engagement ladder within open-source to really embed the sustainability that we need.

Ad: I’ve frequently said that multi-cloud is a stupid best practice. And I stand by that. However, if your customers are in multiple clouds, and you’re a platform, you probably want to be where your customers are, unless you enjoy turning down money. An example of that is InfluxData. InfluxData are the manufacturers of InfluxDB, a time series database that you’ll use if you need a time series database. Check them out at InfluxDB.com.

Corey: I think that there's also been a shift in many open-source projects, or at least as I like to think of them, the successful ones, where you take a look at the projects that have thrived and the projects that have not. Largely, there's going to be a difference as far as how easy it is for someone who is new to it, how easy it is for them to get started. One of the things that started my entire rampage on this was Jordan Cecil. The guy who wrote Logstash Networks over at Elastic-- I saw a conference talk that he gave very early on in my speaking career, and one of the comments he made was that if a new user has a bad time, it's a bug.

And every successful project that I'm engaged with these days is following right along in its footsteps. I mean, I wouldn't consider it code, but I have a bit of an open-source project these days. I'm the community lead for the Open Guide to AWS. It's a giant markdown document that lives in a GitHub repository. That's the 10,000 tips and tricks that you or I would trade with one another over drinks, be that via coffee or tea or whatnot, where one of us is getting started with AWS and the other was giving insight and guidance into how it works. And now it has over 25,000 stars on GitHub, so you know it's good.

185 contributors. There's a Slack team tied to it now with 9,000 people in it. It's the largest Slack team that I'm aware of for the AWS universe. But notice that it is a Slack team. It's not a channel on FreeNote because the world has moved on, because IRC was never friendly and welcoming. We've now gone to this Balkanized environment, which, by the way, I think that Slack is terrible for this, because it does not offer anything approaching full text history, any sort of controls or whatnot for open-source projects. I would love to pay them, but the last time I did the numbers on active users, it would cost me over $60,000 a year. I'll pay hundreds, but not thousands for that.

And that's something where I think we're losing something from the old school, open-source mentality and methodology. But as a counterpoint, if we had an IRC channel instead, we wouldn't have 9,000 people in there. We'd have maybe 90.

Tiffany: Right. That's right. I think that this is another case where open-source is a victim of its own success. As we make the tools easier to use, we really struggle and find the edges of what is sustainable by the old models, right? We need to align the interests of all parties involved. The open-source communities themselves. The users. The contributors. The companies that build successful businesses on and around open-source. We need to make it really easy to get these on-ramps to connect the time, the talent, and the treasure, to the areas of greatest need. Right now, I think we often settle in the places of the least friction.

There's nothing wrong with that. That's not a judgment statement. Slack is a place of low friction. But as you know, there are limitations to it. It's not going to be sustainable for an open-source community to invest in a commercial license to be able to have the full text history of it. And so we lose things all the time by not rolling around. But I think what the flip side of it is that all of these tools have accelerated that adoption and accelerated the community's ability to really focus in on what it is they do best. So, in some ways, it's helped us all get off the island a little bit. But I think that there are some strategies. There are certainly some tactics that open-source communities need to be thinking about right now about how to meet then needs of not only where they are now, but where they'll be in three, five, 10 years. Particularly with respect to succession planning.

Corey: Yes. And that's something that I think people don't take seriously enough. We take a look at tech and how rapidly it moves. So things you'd have to generally care about in other ecosystems and other spaces is how do you build something that out lives me? Whereas, the idea of me writing code today that someone is still using five years from now is horrifying, and my lifetime planning extends beyond a five year time horizon. Ergo, I will always be around as long as this code needs to be maintained. So I don't really need to worry about who's going to take it over once I'm dead and gone. That isn't really accurate or fair any more, but that's the mentality people are approaching this with. This benevolent dictator for life nonsense that we see in some projects suffers from exactly this failure mode.

Tiffany: Yeah. And I think coming from the Drupal community, I am still a defender of the BDFL. Benevolent dictator for life. And what I think that in our case and the Drupal community's case, Dries has done really well is he started to build a team of people around him. And to start to distribute decision making, right? I mean, this is really where open-source projects move from just doing agile to starting to be agile. And I think that's the next place where we're going to start to see improvements. And even at scale. Even at the scale of a project like Drupal. But I also think that the succession planning is threatened in some ways by the impermanence of the cloud.

It's both freeing because you don't have to worry about maintaining project issue queues like Drupal does. I mean, that is certainly a huge expense. It's been a super power of Drupal's for awhile, the way that they manage their issue queues and the way they manage contribution. I think really allowed our project to be very, very successful five years ago. But at the same time, maintaining that, the level of technical debt that we also have to support just around the periphery of our project um, is a significant drain on resources, or a significant use of our resources.

That said, moving the hosting of Drupal.org into the cloud, away from the servers, has been a good thing. But I do worry about the ancillary services that we use like Slack because it is so impermanent by design, that we lose that kind of institutional knowledge. Which was always, I think, a bit tenuous anyway in open-source. Those who were there remember, and they have a bit of a long memory. But it used to be able to be passed down from person to person. The culture of open-source was really a very personal one. And I think that the cloud has enabled open-source to scale so fast and projects to blow up so quickly, that that culture doesn't necessarily have a chance to catch hold and spread in the way that it used to.

So we need to adapt the ways that we do this and it's going to involve some change management. Some new thinking around what it means to operate at this kind of scale, even beyond the questions that arise because you have the distribution and the licensing questions. I think there are just even bigger questions posed by the cloud in terms of how are you onboarding people? What is your obligation to someone who hasn't built the karma yet in your community?

Corey: And that was part of the challenge I had with a number of open-source communities. Every project seemed to have a slightly different vibe to it, but there were several that ... well, I'll name it. Debian was a good one back in the day, where I would ask a question and the answer was that I should shut my fool mouth and read the freaking documentation before asking any further stupid questions and educate myself first. The lesson I took from this was don't use Debian. Instead, go in any other direction here. And again, not to steer this into the weeds necessarily, but that's the experience that I had. I found it incredibly off-putting and I am a cishet white dude in tech.

The entire open-source ecosystem feels ... and along with the rest of society, was built to cater to my precise demographic. Obviously because we're better than everyone else. Please! The point is, if I found this off-putting, I cannot imagine what it would be like for someone who didn't look and sounded like me. And that was terrifying and horrifying. And it was not in any way, shape or form inclusive. And that is where so many open-source communities have fallen. Others have stepped up and absolutely thrived as they wind up instituting a don't be an asshole policy.

Tiffany: I think it's absolutely key to make sure, for the success of open-source, both as a culture and a community, but also as a product and a project, to be able to invite more people in. We need it for adoption but we also need it for innovation. I mean, the research is incredibly clear that diverse teams create the most successful projects and products. But then it does create this question, right? Open source really did start as a monoculture. And your experience is not uncommon. You were part of the dominant culture. And yet, it still could be a very hostile place, right? And I think that the projects that ended up being more successful from an inclusion perspective are the ones that ... at least initially in my experience they were by accident. I mean, I think that there were women and people of color who got involved with those communities and either because the people that they interacted with were very welcoming as in the case of Drupal.

You have webchick, Angie Byron, who was just out there welcoming everyone for whoever they were. And wasn't shy about who she was and what she brought to the table very, very early on in Drupal, that representation mattered. And so when you started to get pockets of people who felt welcomed, who felt landed in a community, it continued to spread. And that's what we see. We see that in companies but we also see that in open-source as well. I think what I've always found really interesting is this notion that in a lot of the projects that are online only as opposed to those that have an IRL component, whether it's through meet ups or through conferences, you do see the opportunity for those who have non-gendered handles to be successful.

And I think that that was very interesting. That was very freeing to some folks. And as that became ... I think as that became accepted, I think those are the communities that we saw who embraced that diversity as one of their core values. But it takes a tremendous amount of work to do the community management as you shift from this originally kind of as you noted, white, male cis monoculture where, yeah, there was some diversion. But everybody was fundamentally the same and you kind of had a baseline assumption of what the norms were going to be and how you were going to be treated, over to a truly multiculture environment. And if we look at how this has been handled most maturely, so far you look to corporate America, right?

And so and what you find is that most of them move from a monoculture to a multiculture through a compliance phase, right? And this is where I see a lot of open-source projects right now. They are with varying degrees of success, struggling with what community governance actually means. And the ways in which that governance should be structured, should be supported. And as well as tasked with accommodating, supporting and encouraging diversity, equity and inclusion. But, that phase is really essential. And I think the more that open-source projects can lean on each other, not only for experience but also for literal support. I would love to see communities get together and have ... serve on each other's appeal boards for their community working groups or whatever the analog is in their community.

Because there's a lot of expertise that is being built and there are a lot of successful experiments going on in different projects that could really catch on if they were amplified. And so there's a lot of opportunity for knowledge sharing. There's a lot of opportunity for literal bodies and experience to be shared between the open-source projects as we all navigate through this kind of compliance phase where the leadership has bought in that it is important that we make this an inclusive and welcoming space. But you may have little microcultures within your community that aren't yet there, or are afraid about what that means or afraid that they might inadvertently say something wrong.

And so I think the more we are able to provide community governance structures to support that, the better it's going to be, right? I mean, I think it's really a question around helping communities understand that mistakes are going to happen, but that all mistakes are recoverable if you approach it with honesty and openness. So, creating that kind of a structure particularly at scale is daunting and it's a place where I see a certain economy of scale for the projects that they start to work together.

Corey: One of the ... even if you take a step back for a second, and go back to a pure, self-interested perspective, assume that even back in the days when you had someone showing up and your first insult when they asked a dumb question was to insult them. What, out of pure self-interest, what is the shortest path to get effectively that person who knows nothing and contributes nothing into someone who is actively contributing to the code base? To the environment? To the community? And if you drive people off, the short answer is, they won't. They're going to go find somewhere else to be.

An early, formative experience for me was I was, I think, developer number 15 behind SaltStack, and I am about as good at Python as you probably would suspect I am from this conversation. But they merged every pull request that I put in. Sure, 10 minutes later, there was another pull request from one of the founders of the project immediately fixing all the things I just broken, and they reached out and talked to me about this. But it was such a welcome environment that oh my god, I can do this! And I couldn't do it, but I thought that I could. And over time, I became able to do it. Because I wasn't driven off for not being good enough yet.

From that perspective alone, I became a champion of the project. I would talk about it to people constantly and thankfully for everyone involved, I stopped contributing a lot of code to it. But instead, I started contributing in other ways and being a bit of a cheerleader for it. That's value to a community for an open-source project that lives or dies based upon what people think about it. I just don't understand the shortsightedness that goes into driving newcomers away.

Tiffany: What you're describing is a learning culture. It's a culture where anyone, regardless of where they are on their journey, is welcome and that it's a mistake making place, right? I mean, that's one of the brilliant innovations of the pull request ... this kind of pull request culture, I think. If you look at it from that perspective, you can allow people to both make mistakes and feel like they have this contribution, and then show them what that mistake is and show them how you correct it, right? I think that's a ... it makes such a difference to people, and I think it's crucial especially as we acknowledge that unlike maybe 20 years ago or 25 years ago when we had a lot of hobbyists and enthusiasts who were playing around with open-source, the folks who are asking to play, to work with open-source now, are probably doing it for their jobs.

And so, the environment itself whereas it used to kind of have this almost extracurricular feel to it, is largely professionalized. And certainly, we see this in the third generation of open-source projects. They were really born to serve a need, either because they have a corporate sponsorship or that they are a consortium of companies really backing it and working together to create and put it out there. And I think that's another key piece to consider is that the element of choice and agency in working with open-source has changed incredibly now. And that evolution changes the dynamic around contribution. People will do what they need to do for work and that can be consumption.

If you have to use open-source because is is the right tool or there was a strategic decision, great. You're going to do that. But the real question is, how are you going to take that person who maybe is required to do it for work and either get them to have contribution policies back from those companies themselves, get some time to contribute patches back, whether they're incremental or more major? Either way, I think a lot of that depends on how folks are treated. Because a company, I can't ask someone on my team to work in a toxic environment. We have a code of conduct that we expect when folks are working at my company and I can't ask them to go somewhere where they would receive lesser treatment. So I think that starts to create that pressure because I can ask somebody to use something. Anybody can go to a website and download whatever tool it is.

Go to GitHub, get the code you need, and then just use it. That's really where the opportunity costs really starts to stack up when open-source communities don't create this kind of welcoming on-ramp to take that person who is doing it professionally, because it's a part of their job, and really get them indoctrinated and onboard with the expectations of the community and to be that evangelist within their own companies for why it matters, how their company will benefit, how the team themselves will benefit from having more people reviewing their code. That certainly was our experience, why we went to open-source, was that Drupal had more people on the security team than I had in my company.

So I knew that my efforts to develop my own CMS were going to be swamped. And if I put those efforts toward Drupal and really helping Drupal succeed in the areas where our own CMS had succeeded, we were all going to be better for it. But I was also going to gain the things that they did much better than we did. Whether it was security team review or whatever that was. So I think acknowledgement of the fact that open-source has professionalized and is now a routine part of people's jobs, it's something I'd like to see more folks talking about and why that matters when we start to look at some of the toxicity that can happen in communities and how that is handled.

If it happened within your company, HR would be expected to step in and to take care of it. If your open-source community doesn't have the equivalent of a community HR or community working group, you may run into some issues as a result.

Corey: It really comes down to being inclusive. It leads to better outcomes for everyone on all sides of the table. I just don't understand how there are people on the other side of this particular issue.

Tiffany: I try to approach it from a position of compassion and really what is driving that? And I work from a place of positive intent, that it's not about excluding people as much as it is the fear of being excluded. Particularly, if you've been part of the dominant in-group for a long time, it's important to you. open-source is incredibly important to folks who've been involved in the community for a long time, or certainly some newcomers who just find their place and feel like they belong. And I think if you've ever had a place where you belong and you are going through what you feel is an existential threat where by including people who don't look like you, that you feel like your place is at risk, I can understand why you'd get concerned.

I also think that it's incumbent on those who are working on DEI, diversity, equity and inclusion initiatives to help folks who have always been there find that sense of belonging and help them understand that there is such a systemic and institutional conditioning that we are all subject to. That everyone, everyone is going to make a mistake. Everyone is going to, at some point, just be transgressive, unintentionally, right? And accepting that position, like, “You know what? I might say something sexist, or I might say something racist, but what matters is how I recover from that and how I make it right.” And modeling that in your community. That it's not about perfection.

It's not about everybody coming from the same point of view. In fact, that would undermine it if you only had folks who shared the my way of thinking or my ideology. It's not about that at all. But it is about making sure that it is a safe place. That our open-source communities are places where I can come and I can be curious about what those around me have to contribute and they will be generous about sharing their knowledge with me. And likewise, I will be generous about sharing what I bring to the table for them. That's how we create an environment of safety for everybody.

And there's never been a perfect place. And there has never been a perfect person. There's not a person among us who has not made a mistake in dealing with anybody. And so I think that the more we talk about that and the more we create systems that help call people in when their behavior isn't meeting the accepted norms, if they're not understanding how we do it here, give them that opportunity to reflect on it. Give them a space to be able to process it. Give them a space to be able to change. If they want to. If they don't want to, then yeah. Okay. You can choose to leave. But it is fundamentally your choice. It's not about having to be perfect to stay. It's about understanding and buying into the norms that my presence here can't make someone else's presence invalidated or erased or just unwelcome.

Corey: I think that's something that needs to be internalized by people a lot more than it currently is. I want to thank you for taking the time to speak with me today. If people want to find more about what you have to say and get your thoughts on these and other matters, where can they find you?

Tiffany: I'm on Twitter at @farriss, F-A-R-R-I-S-S. You can also see things that I blog about and usually come through the @palantir handled. P-A-L-A-N-T-I-R.

Corey: Yes. Official motto, no, not that one.

Tiffany: That's right! This is not the Palantir you were looking for.

Corey: Indeed! Thank you so much for your time. I appreciate it.

Tiffany: Thank you so much Corey. I enjoyed it.

Corey: Tiffany Farriss, CEO and co-owner of Palantir.net. I'm Corey Quinn. This is Screaming in the Cloud. If you've enjoyed this episode, please leave us a five start review on iTunes. If you've hated this episode, please leave us a five start review on iTunes.

Announcer: This has been this week's episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

Announcer: This has been a HumblePod Production. Stay humble.

View Details

About Sasha Rosenbaum
Sasha is a Program Manager on the Azure DevOps engineering team, focused on improving the alignment of the product with open-source software.

Sasha is a co-organizer of the DevOps Days Chicago and the DeliveryConf conferences, and recently published a book on Serverless computing in Azure with .NET.

Links Referenced:

  • Sponsor: Snark.cloud/AWSsolutions
  • Twitter Username: @DivineOps
  • LinkedIn URL: https://www.linkedin.com/in/sasha-rosenbaum/
  • Personal site: https://www.sasharosenbaum.com/
  • Youtube channel: Azure DevOps
  • Company site: microsoft.com

Transcript

Announcer: Hello and welcome to Screaming in the Cloud with your host cloud economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on this state of the technical world and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored by AWS Solutions, which is the exact opposite of AWS problems. AWS Solutions are vetted, technical reference implementations that are designed to help you solve common problems and build faster. They themselves are free though occasionally some of the products they stand up are, not but it’s a great way to click button wind up receiving a technical solution that’s implemented that ideally solves a problem you have. Visit snark.cloud/AWSsolutions again that is snark.cloud/AWSsolutions and my thanks to AWS for their generous sponsorship of this episode.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Sasha Rosenbaum. Sasha, welcome to the show.

Sasha: Thank you. Thank you for having me.

Corey: So you're a senior program manager or thereabouts on something called the Azure DevOps engineering team where you focus on improving the alignment of the product with open-source software. That's a wonderful, I guess synopsis of what is almost certainly a much larger story, but let's begin at the beginning. What is it you'd say it is that you do?

Sasha: Okay, so thank you for this question. It's kind of a complicated question. So I've held many titles over the years, right? I started off as a developer and I did that for a bunch of years. Then I was a consultant. Then I became a cloud architect, which means I helped a bunch of people move to the cloud, coincidentally happens to be Azure Cloud.

Sasha: And this was the first time I'm a program manager and I'm on something that's called a community team for Azure DevOps, which means a lot of my job is around connecting people and so helping us align on different messages and help each other. Because we often find it that different people across the industry or even across the company are working on the same thing and they just don't know that the same effort is going on. So sort of helping people stay aligned is where we are.

Corey: Which I guess leads to an interesting question of, and again my apologies for not having done enough homework to be able to answer this myself, but what precisely is an Azure DevOp?

Sasha: Okay, so obviously I did not name the product, and as our friend Matt Stratton likes to say, you know "You can't buy DevOps but I can sell it to you." So this is kind of maybe along these lines, but also, you know words tend to evolve. So the folks who coined the word DevOps meant one thing, and right now we're seeing it all over, you know all over the industry where people apply it to all sorts of different things.

Sasha: Azure DevOps is a product that is a SAAS product based on Azure that helps you implement CI/CD and also helps you with a whole host of other things related to development, such as you know managing the project, you know planning and [inaudible 00:03:07] and all sorts of things. So kind of the whole package, as a SAAS product.

Corey: Okay. There are two things I feel that urge to, I guess one thing is a statement. The second is a question. The first is a statement in that this is an actual product name. So if you see someone in the wild with Azure DevOps on their resume, that's not someone who has no idea what words mean, but rather someone who is effectively dealing with a terrible product name that makes a resume look like random buzzwords thrown together for SEO purposes. Nope, it's real. You're not allowed to make fun of someone or reject them as a candidate for having this on their resume unless they're responsible for having named it.

Corey: Secondly, it sounds like the service itself is very similar to what one of your competitors does, except for the fact where you own this competitor, namely GitHub or jif-hub, as I insist on pronouncing it, seems to be focused fairly heavily on expanding beyond just a place to host your Git repos, or jit repos, we're all things to all people here. And increasingly giving insight into these things of being able to handle CI/CD natively rather than just integrations with third parties that provide these things with actions, with workflows, with a bunch of, I guess, almost a pipeline style approach.

Corey: Where every time I look at the GitHub landing page when I log in, there's new capabilities I didn't know were there. So now this is sort of the inherent problem of building anything that looks even remotely like a platform. What's the difference, I guess, between what GitHub is doing and Azure DevOps or isn't there one?

Sasha: So, first of all, thank you for not holding titles against people, because I know a whole bunch of really excellent people who are you know Azure DevOps professionals, and they are doing great work out there. So it's awesome that they can get you know hired. It's funny because I track Twitter feed for Azure DevOps and I see a whole bunch of job postings with that in a title. So it's kind of funny to me.

Sasha: But so on the GitHub question, so this is, and I feel like I've answered every single one of your questions, but like a preface of it's complicated but it is somewhat complicated. This is one of those things where humans like to have a very cohesive story. And we started to have a little bit of shades of gray answer to the question. So we, the Azure DevOps is a product that evolved out of TFS. So it's been in development for over 10 years, and we've modernized big parts and pieces of it.

Sasha: But, again, this is sort of a progression of an enterprise facing journey that we've been on. And now recently we've acquired GitHub. So GitHub is now part of the Microsoft family. But at the same time they're a standalone company. They manage their own decisions, right? They're cloud agnostic and all sorts of different things like that.

Sasha: So as far as this journey, back a year ago, I guess, when GitHub released Actions, the biggest requests for them was to implement CI/CD with Actions, right, because originally they did not intend to do that. And so once they saw the overwhelming requests from the community, they decided to go after that workload. Which means that, yes, you're right, we kind of in a Microsoft family now have two products that can solve the same questions, the same workload.

Sasha: It's also true for repositories, right? Because obviously GitHub is the biggest source code on... you know hosted product in the world and Azure Repos also has Git Repos available. And so we get asked a lot of times on which one should I choose, and the thing is technically there is some differences between the two and one of them might be more applicable to you than the other. And I'm not going to get into the technical differences.

Sasha: I think we're going to eventually publish documents, like publicly available documentation on those. But the other thing is, so I think GitHub in terms of CI/CD is kind of starting out, right? So they are ready, like just from releasing the preview of Actions, they've seen amazing adoption and lots and lots of users on the platform. But also in terms of being enterprise ready, it's going to take a while. So again, if you are looking for a SAAS product that's enterprise ready, you might be better off with Azure DevOps, but if you are sort of... In the long, long run, we're probably all looking at GitHub.

Corey: Got you. And at some point is there a meaningful distinction between them? Because fundamentally it is all one company and I know that's sort of anathema to a number of large environments where their biggest competitors are other internal groups. But you'd think that at some point when everyone's email address ends with the same domain, that you tend to be aligned on things. Where when an internal group offers a better solution for a particular part of the problem, then well why not just merge those things in rather than having two completely different tracks moving forward?

Sasha: Right, so this is not quite the same story because, like I mentioned, GitHub is maintaining its own leadership and its own roadmap and it's not going to become sort of the Azure first platform, right? Because GitHub is going to continue to serve AWS customers and GCP customers and whatnot, you know other things out there. Whereas Azure DevOps is obviously you know sort of tied into Azure, although little known fact, we have lots of integrations and we have like AWS integration that's maintained by AWS. So it's kind of cool.

Sasha: So basically again, we're never going to merge them. You kind of have that, and again, I don't know how visible it is outside of Microsoft, but you kind of have that example with LinkedIn, where Microsoft does not control LinkedIn as a company or the roadmap for their products. So sort of, we do obviously create more integrations around that stuff, but we don't set the goals for the these companies.

Corey: That's an interesting perspective. I suppose that having GitHub or jif-hub retain its independent leadership is absolutely critical to maintain the culture that it's built and you know maintaining the value that you folks paid for it. But on the other hand, I can't shake the feeling that there's a certain inefficiency when company divisions start competing with other company divisions for public offerings.

Corey: And again, this is not in any way a Microsoft-specific problem. I think we see it with every large cloud provider period. We see it with even relatively early stage companies, where they start to offer different product offerings that are differentiated, but also start to step on each other's toes. I, I guess it's one of those things where it's easy for me to sit here as a five person company and say, "Ah, you're getting twisted around an axle here."

Sasha: Wait, you have a five-person company? I didn't know that.

Corey: Oh yeah, there are two of us who own the thing and then we have three employees and growing. It turns out that we are larger than people thought. I'm just personally scaling horizontally as I eat my way to becoming a 10x engineer.

Sasha: Very, very cool. Yeah, we all are, right?

Corey: Indeed.

Sasha: Unfortunately working with computers doesn't help. Yeah. So, it's a great point and I've seen it a lot you know in the past where, especially in large companies, you kind of start implementing the same thing over and over again. So it's actually an interesting fact that like one of the things that Microsoft did as part of like the DevOps transformation that we had was that we mandated everyone to be on the same platform for CI/CD, right?

Sasha: So whenever you were building and releasing software, we wanted you to be on Azure DevOps, which made us our own biggest customer, which means that we can improve much faster, right, and we can test our own stuff, because we do this progressive release, right? And by the time the release hits our actual customers, it's been released to all the internal Microsoft employees, and so between... You know we have 100,000 internal customers, so in the 100,000 hopefully someone raises a red flag if there is some issue with a release.

Sasha: So, that's a big powerful thing because otherwise what you see is, like you said, people are competing with the same product, and it gets kind of complicated. Unfortunately here, we can't do that, because you know we still have sort of different constituencies because... So with GitHub, like GitHub has to, has certain set of priorities that will not align with potentially Microsoft internal customers and vice versa, right? We have to serve our internal customers, you know sometimes it doesn't quite align with what GitHub wants to do.

Sasha: And so we will definitely maintain so two products for a while. We do share some of the same code base, and we do try to bring features along to both of these platforms and we're going to continue to do that. But again, we still will have a separate roadmap.

Corey: Which makes a lot of sense. I think that there's a definite validity to keeping those streams separate. It just does become occasionally confusing when I think from a customer perspective where you're trying to figure out what is the best path forward? Because increasingly as I look down any technical path that transcends virtually any vendor out there, how do I wind up doing thing X?

Corey: And it doesn't matter what that thing might be, but it's not that there's an answer to it, it's that there are hundreds, and they're all equally valid at some point. Where it's as soon as you start looking to different blog posts and then starting to combine them together, it turns out that you've built an insane Frankenstein monster that no one fully comprehends and, surprise, you are in unicorn territory where all of your problems are ones of your own making and no one can help.

Corey: So the idea of having a blessed golden path to go down is tremendously compelling. I'm just not convinced that it actually exists here in the real world.

Sasha: I think, you know and in my personal opinion, the technology landscape just continues to get more complicated, which is really interesting, because if you look at like user experiences for things around us, we make it simpler and simpler as we go. But if you look at the technology experience, like a building technology, it gets more complicated progressively. And part of it is you know open source is a great thing, but it also means that a lot of people are building software you know sort of on small teams or even on weekends.

Sasha: And I've seen you know large companies stake dependencies on software packages that were built by someone as a hobby on weekends. And you can imagine that it you know doesn't always turn out so well, you know and unfortunately, I don't know that there is an answer to that. I think it's just going to continue happening and we continue to see companies pop up and solve different problems, at the different levels of the stack, right, and because like you know nothing is a complete solution, you're going to continue having like these little startups innovate in particular area.

Sasha: One thing that personally I think is that if you can go with a SAAS product, it's always, like I would always have a preference for that, because, hey, you know we're pretty good at running products at scale. And if you don't trust us that means that you trust your... Like you hire internal people to manage the same compute and to match the same availability and to match all of these things, and honestly it's probably going to cost you way more. And you might not even end up with a result as good as that. Right? So not everyone needs to build their platform from scratch. If you can consume some SAAS based product that will solve most of your problems, just go have that.

Corey: Oh yeah. Down this path is the same problem that we see with multi-cloud across the board. This is somewhat challenging for companies that are their own cloud providers to articulate because it sounds incredibly self-serving. But from my neutral third party perspective, I don't care what cloud providers someone picks. If it's Azure, it's AWS. If it's GCP, if God forbid it's Oracle Cloud, great. Pick the thing that you want that makes sense, and then go all in with whatever offerings they've got. Because trying to stitch together 500 platform as a service offerings together to build your own Franken-platform is almost never the right answer for almost everyone.

Sasha: I will say, so in terms of going crazy, like I have seen people try to do sort of multi cloud deployments in terms of like you know trying to do the same architecture across the AWS and Azure. Never quite works, right? But at the same time there are certain things that you could consume, right, if you're calling Cognitive Services API kind of it doesn't matter which cloud you call it on, it still works. It responds to you and stuff like that.

Sasha: So you could, and again, Azure DevOps is a SAAS product, like we have customers' using it with on prem and with other clouds and stuff like that, right? So, in the end it kind of doesn't matter where it is, but certainly if you are trying to deploy VM skill sets across cloud, I wish you the best of luck.

Corey: Yeah. So changing gears slightly, I encountered you at a few different conferences now, which is interesting. I had to triple check your title before we started recording this. Expecting to see advocate or evangelist or some other form of devreloper type job. But instead, no, you are a senior program manager. You've been a fair, you've been in this space for a long time at fairly senior roles. You have a hands on engineering background as well, but you've never had a job where, at least from my reading of LinkedIn, that everything was aligned around storytelling. It seems like that has been incidental, at least public storytelling. What's that like?

Sasha: Yeah, so I think that's you're spot on. So for me, this storytelling has been a hobby rather than a primary job. Now granted in my current job, I actually apparently do this more than I used to, but it's never a sort of primary thing that we work on, that I worked on. I do like talking to people. I do like being on stage, even though I... You know usually the half an hour before the talk, I deeply regret that I've you know done it to myself but, and in the moment I'm really enjoying it.

Sasha: I think you know part of it is just wanting to contribute to the community. You know I also am an organizer of DevOps Days Chicago and I've been there for I think six years now, and I, we just started a new conference which I want to talk about, which was called Delivery Conf, and again the biggest mission for all of these things is just connecting people and sharing knowledge.

Sasha: Because that's one of the things that we, well I don't know if it's as an industry or as humanity are not particularly good at, is sharing experience and telling you how to instead of you going and building a Frankenstein based on blog articles from all over the place. And you know that some of them are not even going to be accurate, which was unfortunate. You can talk to real practitioners who are building the same type of software at big companies, at small companies, right, and sort of learn from that experience.

Corey: I liked that you said small companies on there just because one of the, I guess, challenges is that very often we'll see the same collection of companies getting on stage, talking about their journey, how they're going to go ahead and build stories for the future, how they're structuring what you should do as a best practice, etc., etc. And inherently you see the same companies, you see the Netflixes, the Googles, the Capital Ones, the same folks who are already significantly far down whatever transformation or revelatory discovery that they've made and what they're saying doesn't look an awful lot like too many other shops.

Corey: Sure, there are lessons there, but the context doesn't seem to filter through very often. And that may be an unfair criticism in some cases. I'm certainly not saying that every talk by someone from those companies is inherently terrible, but it does often come with shades of nuance that are lost on certain audience members. Where, well, this is what Netflix does, so we're going to do it. Ignoring entirely the fact that they stream movies and you're a bank, is a common failure mode that we'll see. You know-

Sasha: I will say, Corey, I'm going to interrupt you and say this is my favorite example, right? I mean the Netflix versus a bank, because so when I go to my Netflix queue, it routinely doesn't start where I left it off and I'm fine with that, right? I mean they're running some containers, distributed operation and skydive. I can find the time spot in my movie, right?

Sasha: But if my bank didn't start at the place that I left it off, I would be very severely disappointed, right, not to mention that would face legal problems if they did that. So, eventual consistency as well and good, but it doesn't work for everyone. And so we are solving different problems in different industries. And definitely, again, there's a big difference between startups and big companies.

Sasha: I will say one thing that is kind of exciting for me at Microsoft is that we are a unicorn, but we're also... Like when I talk about the Microsoft DevOps journey, there is a real transformation there because we didn't start that way, right? We're a 40 year old company, and we did not release software on a three week cadence. I mean we release software on a two year cadence, right?

Sasha: And so there was a big, big difference and big sort of understanding a lot of things and evolving in how we structured teams and how we do releases and how we like even get to live site management and stuff like that. Right? So there's a lot of lessons learned for sort of enterprises that are trying to evolve towards that goal.

Corey: Absolutely. The trick is of course, making sure that context carries through, and to some extent that responsibility does clearly fall on the audience. Where, huh, if you know you're a bank and your failure mode is somewhat different than a large streaming media company, you have to go in with that expectation. The idea that everyone gets the root and production for example, is not really one that your auditors will listen to without throwing you out of the building. That, that has to come with the audience.

Corey: And I think that it's not potentially fair to say, well we shouldn't hear these stories because they won't apply to everyone. That said, I do like the idea that you're going to be hearing at Delivery Conf from folks who are not the usual suspects, that they're not going to inherently just be the same five people that we've all heard from before and telling the same slightly modified version of the talk they gave at the last conference. So I am looking forward to that. I think that there will be a lot of interesting stories coming out of it.

Sasha: So I hope so. And so the reason we started, and again, there's a lot of conferences out there, right? And so me and a couple other folks, so Ken Mugrage and Jason Yee, we just sort of came up with this idea and it was, it started with Ken of like, hey, we've been doing DevOps Days for awhile. And DevOps Days is absolutely great, but because it's a single track conference and it caters to a wide audience, we can never get deep into technology, right?

Sasha: So we keep talking about principles, we keep talking about culture, we keep talking about high level, hey, you know DevOps transformation and stuff like that. But we can't really get into the nitty gritty, because when you are getting to the nitty gritty you have to show me the tool, right? You can't abstract a way that you're using the particular cloud and a particular automation tool and stuff like that. And for people really to be able to learn from you, you should be able to demonstrate that.

Sasha: And we kind of noticed that all of the industry conference, which were technical, were very sort of vendor specific and we wanted to give a stage to a wide variety of vendors. So like, hey, if someone from AWS and someone from Microsoft and someone from cloud be used and whatever it is, wants to go on stage and talk about their stuff. Again, we're not encouraging product pitches but like show me how you solved a problem with a real tool is totally cool. So I am excited for this and we are going to you know try as best we can to select practitioners for giving the talks. So I think the content is going to be really good.

Sasha: We're also doing one more thing that is different and we'll see how it goes, but it sounds like a great idea that could work out really well. So we all really appreciate the DevOps Days in terms of having open spaces, and if folks are not familiar with an open space, it's just kind of a space where someone proposes a topic and then a bunch of folks can get into a room and have a discussion on that topic. So it's an open discussion, not a preset you know PowerPoint or whatever, and you can actually have a conversation.

Sasha: And these conversations sometimes are the most valuable thing that comes out of any conference. And so, but the problem is that these conversations, like if you're not in that specific room, then you're not going to be a part of it, and you're not going to hear from other practitioners and whatnot.

Sasha: So, what we're doing at Delivery Conf is actually we're going to record, so after each session we're going to have a 20 minute, and 20 minutes is not enough but hey, that's what we got, a 20 minute discussion on the same topic. So kind of not around the speaker but around the topic itself. So I'd say, you know you presented on ML and I totally disagree with your conclusions. I can still participate in that discussion, right?

Sasha: Whereas usually, you know when we ask for Q and A, we asked for Q and A that a question ends with a question mark and it's not an argument. So you totally, you're welcome to bring your arguments to this discussion and we are going to record the discussions. So it's going to be first part, you know first, oh man... Anyway, it's going to be the content that's available on our website. Once we released the talks, we will also release those discussions so they can become those like mini podcasts or whatnot and people can hear from other practitioners.

Corey: That sounds like it's a compelling story. I'd be very interested to hear how it works out just from a what works with conferences and what doesn't perspective.

Sasha: Right. Well, you're welcome to attend and participate. Then you can talk from a first party you know experience.

Corey: That's the hard part is it seems that if I let myself, I would never be home anymore because I'd be attending too many conferences. I made a concerted effort to knock that down significantly this year. And even as I started to plan early this morning before we recorded, the my schedule for the next year at conferences, I'm still slated to go to at least 16.

Corey: And that becomes a bit of an ongoing challenge for me personally just because there are things I have to do professionally that don't necessarily lend themselves to me spending a week going up hill and down dale for various conferences, despite the fact that there are so many I want to go to.

Sasha: Well, see you beat me on conferences, you know talking about devrelish things. But so we did schedule Delivery Conf in January, January 21st and 22nd, which was off season for most conferences. So we're kind of hoping that will help offset you know how many events are out there. There's certainly a lot of events out there. You know as I'm venturing into the speaking space a little bit more, I'm just noticing how many are out there. Just kind of, you know you could go to three conferences every week if you wanted to.

Corey: Yeah, that's the thing. You don't even need to have an apartment anymore. You can just go from conference to conference and eat nothing, but the buffets and the happy hour drinks and the rest and there's probably an entire subsistence living you could eek out just by going from conferences and migrating around. Problem is, is that is for someone in a very different stage of life than I'm at.

Sasha: Well, I do know a couple of people who became sort of digital nomads and like not necessarily conferences, that also happens to be true of like if you are on a consulting team that travels every week. You might arrive at a stage where you don't actually need an apartment. But yeah, if you have family at home, that certainly doesn't work so well.

Corey: Yeah, that's exactly the entire point. It's, you've got to at some point be there or people will wind up hurling you into the dumpster of history. It just doesn't go well.

Sasha: Certainly.

Corey: So if people have enjoyed listening to you, which they should have, if they're paying attention even slightly, where can they go to hear more of your thoughts?

Sasha: I don't have me, Sasha, I don't have exactly a channel. I do have a website which is Sasharosenbaum.com and it's sort of, I just post links to recent talks that I've done. So it's not really a blog, it's more of a collection of links. And then we do record weekly you know sprint videos, which are about every three weeks, so we do recorded video updates.

Sasha: So if you want to watch me on video, which is always torturous on my side to watch myself on video, I still haven't gotten used to it. So there's a YouTube channel for Azure DevOps that you can follow. Yeah, and then I am on Twitter. I now live on Twitter so you can hear my thoughts on Twitter @DivineOps.

Corey: Excellent. We'll throw a link to that in the show notes. Thank you so much for taking the time to speak with me today. I appreciate it.

Sasha: Yeah, I appreciate you having me and it's always been a pleasure. I'm looking forward to more snarky tweets from you.

Corey: Oh, I don't think we get away from it at this point. It's one of those things that's as certain as the tides these days.

Sasha: Okay.

Corey: Sasha Rosenbaum, senior program manager at Azure DevOps. I'm Corey Quinn. This is Screaming in the Cloud. If you've enjoyed this podcast, please leave it a five star review on iTunes. If you've hated this podcast, please leave a five star review on iTunes.

Announcer: This has been this week's episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com or wherever fine snark is sold.

Announcer: This has been a HumblePod Production. Stay humble.

View Details

About Tanya Janca
Tanya Janca is the co-founder and CEO of Security Sidekick. Her obsession with securing software runs deep, from starting her company, to running her own OWASP chapter for 4 years and founding the OWASP DevSlop open-source and education project. With her countless blog articles, workshops and talks, her focus is clear. Tanya is also an advocate for diversity and inclusion, co-founding the international women’s organization WoSEC, starting the online #MentoringMonday initiative, and personally mentoring, advocating for and enabling countless other women in her field. As a professional computer geek of 20+ years, she is a person who is truly fascinated by the ‘science’ of computer science.

Links

  • Twitter Username: @shehackspurple
  • LinkedIn URL: https://www.linkedin.com/in/tanya-janca-60ab0998/
  • Personal site: https://dev.to/shehackspurple
  • Company site: https://securitysidekick.dev
  • Sponsor: www.manifold.co

Announcer: Hello and welcome to Screaming in the Cloud with your host, cloud economist, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode of Screaming in the Cloud has been sponsored by Manifold. Manifold powers marketplace infrastructure that connects millions of developers to the best APIs, tools, and services in the fastest-growing communities and also Kubernetes. They offer a complete toolkit that allows you to deliver your API first product to millions of developers. Check them out at manifold.co. Again, that's manifold.co.

Corey: Hello and welcome to Screaming in the Cloud once again. My name is Corey Quinn. I'm joined this week by Tanya Janca who is currently the co-founder and CEO of Security Sidekick. Tanya, thank you for joining us today.

Tanya: Thanks for having me, Corey.

Corey: So last time we met in person briefly at a conference, as I think we both were sprinting past each other like ships in the night, you were employed at Microsoft doing something that sounded vaguely securityesque. Again, we were sprinting past each other at a conference. Now, you've started your own company, presumably no longer at Microsoft?

Tanya: No longer at Microsoft.

Corey: Wonderful. So let's start, I guess in that timeline at the beginning. So you started off at Microsoft. First, what org were you in? Microsoft is a big company these days and it turns out that my mental model of the 10 people I know there isn't really a representative sample.

Tanya: So, I was a cloud advocate or a developer advocate. And basically, it was my job to create contents and get feedback from the community in the industry about what works and what doesn't work to help them change their products so they're what people actually need and want as opposed to what we think they need and want, and then create a ton of content so that people know how to do anything they want to do with it.

Tanya: So I specialize in application security and cloud security, so I would create a lot of content about how to create a secure app or how to verify that your app is secure in, for instance, Azure DevOps pipeline.

Corey: Yes. Azure DevOps of course being an ill-fated product name and not a thing that one does that is culture-oriented, correct?

Tanya: Yes. That's the name of the product. Yeah.

Corey: Excellent. I find periodically I have to remind people that that is a product, so if you see it on someone's resume, they're not smacking a bunch of words together. That is the actual name of a product. It's rare that we see a service or product name that is so bad that it can negatively impact someone's career just by mentioning it, but we've done that. Usually, the way to get a higher score I think is to come out with something bigoted.

Tanya: Oh, my God. Well, I wasn't in charge of naming anything.

Corey: Excellent. No one wants to accept responsibility for those.

Tanya: Definitely.

Corey: So we have it really?

Tanya: No responsibility here.

Corey: Exactly. And how long were you at Microsoft for?

Tanya: I was there for two years and I cannot tell you how much I learned. There's a lot of people-

Corey: That's eight years anywhere else, my God.

Tanya: It's basically like a thousand years anywhere else. Like it is... I learned a lot of stuff from a lot of people. It's really cool. So there's a lot of traveling around, speaking at conferences, writing blogs, making videos, things like that. But yeah, I want to just start my own company, I guess. You know when you sit down with your manager and they ask where you want to be in three to five years and then you realize it's that you want to work somewhere else. Like even though you're having fun where you are, you're like, "Oh, I want to do even bigger things." And then you tell them and they make that frowny face. They're happy for you, but they're also like, "Oh, that's not where I thought this conversation was going to go."

Corey: I've heard this story, but personally, whenever, I sat down with my manager, it was always a conversation that started off with, "you know what your problem is" not always from me, not always from them, and it sort of devolved from there. So for me, starting my company was more or less coming down to the fact that I'm unemployable at this point and well, it's either that or starve to death. And other people it turns out have options and the ability to have a employer-employee conversation. I just never excelled at that, but I kind of imagine what it might have been like.

Tanya: I do think that I am a little bit hard to manage because I have really big ideas and if a manager says not to do them, I take that as advice not as, not as like a commandment.

Corey: I view feedback as one person's opinion and that's fair.

Tanya: Mmhm.

Corey: Depending on where it's coming from, it has different weights on it. But it turns out that a lot of management types, specifically crappy managers could frame it as I have feedback for you, and when you have the response, thank you for your opinion is the wrong answer. It just comes down to I think an impedance mismatch sometimes when you have managers who are not great at managing, dealing with people who are, as you say, difficult to manage.

Tanya: Yeah. I also think that like management as a whole, because we're getting pretty off topic but just that-

Corey: Oh, of course, that's the point of a podcast. We can talk whatever we feel like.

Tanya: But there's leaders and there's managers and sometimes they're expecting managers to lead and sometimes they're expecting leaders to manage, and it's not always like those two unique skill sets and the same person.

Corey: Oh, absolutely. In my case though, I found that wow, every manager I had for a while was a jerk. And wait a minute, the only consistent feature here is me. So maybe it's not everyone else's fault that that realization came to me later in life than it should have. But you're right, we are getting slightly off topic.

Corey: So you left Microsoft doing the devreloper thing for lack of a better term.

Tanya: Mmhm.

Corey: And yes, I do call them devrelopers because I have problems. And from there, you decided to start Security Sidekick.

Tanya: Yeah.

Corey: What did the Security Sidekick do?

Tanya: So we do real-time web application and vulnerability, inventory and discovery. Let me explain what that means.

Corey: Please do.

Tanya: Basically, we sit on your network and we're one hop after your DNS. We're an invisible proxy, so everything just goes through us and then we can just recognize, Oh, that's an API. Oh, that's a web app. Oh, that's a SaaS product. And then we just make a list of them for you. We do a passive scan every single time you visit. So we're like, "Oh, a security header got removed." Or "Oh, you've never had this security header." And so you can actually see all the apps you have and a baseline of what's wrong with them.

Tanya: And a lot of people say to me, "Well, isn't inventory kind of boring?" It is. But it's actually one of the most difficult things to get right when you work in an application security engineering role, is that developers do not necessarily tell you because they released a new API and they're like, "Oh, it's number 72." Like they don't care if there's another one that does this slightly different thing. Yes, I do care, and I really want to know. I really, really want to know about every single one of them that's living on our network. Thank you very much.

Corey: Yes, and that that does have significant value. But something I found when I started my consultancy aimed specifically at fixing AWS bills is that there's a lot of affinity to the security space in cost optimization, where it's easy to wind up dumping the billing equivalent of a NASA SCaN on someone. Here's the 8,000 things you can change in your environment and then that rots on the shelf and 95% of them are tiny and no one cares and Oh, do these other few things and you'll cut your bill in half.

Corey: You see that with security, too. And this is one of the recurring stories we see in tales of security breaches where when you have tools that identify security problems like this, there's an awful lot of noise and the signal is buried in them where it almost seems that no one implements something, and the only value these tools bring is being able to make headlines after a breach. And well, it was right there in the logs. Why didn't anyone do anything? Ignoring the fact that there was half a terabyte of logs for someone to go through and that was no one's actual direct job responsibility. How does Security Sidekick get around that?

Tanya: So we basically just make a list of all of your apps on this dashboard that we've created. And then you can click on the app and it tells you all of the things that we have found wrong with it. And then from there, once you know the app exists or if a new app comes out, we'll alert you. Oh by the way, this wasn't previously on your list of things, but did you know it's living on your network? Right?

Tanya: And then you can actually apply your processes to it. You can actually apply your policies to it. Like a lot of places I've worked, we've hired a person to do our application portfolio management, and this very fancy consultant will come in and spend a year or a year and a half interviewing people and asking them which apps they have. But if we could tell you in like 24 or 48 hours, like these are all the apps that people visited that are on your network or in your cloud or wherever it is that's within your domain, oh, okay. So great. Now I actually know what I need to look at. Right?

Tanya: I feel like if you can have a complete picture of what you're looking at, you know what I mean? If you're like, Oh yeah, I have 32 apps and 10 APIs, but then we come in, we're like, "You have 40 apps and you have 25 APIs." Okay, great. So now I can actually look at this up to date list and seven Excel spreadsheet that someone made four years ago that probably only has a third or half of your apps listed on it, and it has some apps that were actually taken offline that they still think exists for some reason. And then you can actually put pipelines around those things. Or you can, you know for instance, we can find... I guess at this point, we're in beta so we can find seven types of vulnerabilities. But we are building that process out.

Tanya: But so you have like a list of things that are wrong. Great. If you can see, if you can look at your analytics, like look at the reports we make and say, "Okay, so it turns out we have like a really big problem with doing direct object references in our URLs, like in our URL parameters. So the vulnerability is called IDOR, like an indirect object reference. But the idea is in the address bar, it's like bank account number equals one, two, three, four, five, right?

Tanya: So if we see that happening a lot, there's clearly like a developer or a group of developers that doesn't see this as a problem. They think it's fine. So then you can make a lunch and learn or address this with you know some training. You can address this with a new... Sorry. You can address this by going to that team and explaining the relevance and why this is important. And then you can try to eliminate classes as a whole because you finally have a complete picture.

Tanya: I've worked at a lot of places doing... like I do a lot of consulting and then I also have been an employee for a long time, and basically, like I would come in places and they'd be concentrating on a thing because you know a pen tester came and they could only afford a pen tester like twice a year, let's say. The pen tester would be like, "I found this injection, vulnerability injection. It's the worst thing ever. And it's awful and it is awful injections, bad." But they found one and it turns out in all your apps there was only that one. But what's really problematic is that everyone is doing cross-site scripting in every single possible input field everywhere in every single app.

Tanya: And you would actually do better to do a deep dive into cross-site scripting and teach everyone about that and then just address the one injection vulnerability like uniquely rather than making everyone sit through training for that, right? Because when you give training and you let's say, you pay a trainer $5,000 to come in and spend a day, everyone's like, "Oh, it was $5,000." "No, it was not." If you had all your developers sitting in a room for an entire day, that probably was $100,000 because developers costs a lot of salary dollars.

Corey: Oh, absolutely.

Tanya: If you have room full of them, you're wasting time. And it's condescending too if you're a senior developer and you're like, "Yeah, I know injection inside and out. That was you know a student that we hired or whatever." Right? Like you want to spend your time on the things that matter and if you don't have a complete kind of higher level picture of things, it's harder to decide what you actually want to do with your time and your limited budget.

Corey: And it gets worse than that. A lot of times, compliance requirements dictate you have to send people through the same ridiculous training.

Tanya: Oh, my gosh. Yes.

Corey: And it doesn't add a whole lot of value. It's the, we had to go and check the boxes and the rest. I see the same thing with this being an ongoing challenge where for example, in the world of cost, which is the one I know best, a lot of companies will come at this from a perspective of we want to train all of our engineers, and my response is, "really?" Because most of what they need to know about AWS billing can fit on an index card. You don't need to have a three-day training for every engineer in the building.

Corey: Sure. Someone should probably know the nuances in this environment, but that is a far cry from everyone having to think about this all the time. Because in almost every case, people cost more in compensation than they spend in infrastructure.

Tanya: Mmhmm.

Corey: And it's... You see the same thing when you have all these trainings on all of these different attack vectors. At some point, yeah, you should have every engineer know how to sanitize inputs, but maybe every engineer doesn't need to be a fully qualified pen tester in most companies.

Tanya: Oh, my gosh, Corey. There's so much training that I see teams go through. They're like, "Yeah, we're going to get..." I'm not going to name the trainings, but where it's like how to hack some random version of Unix or something. I'm like, I don't need a software development team to know that. I just don't. And so I know hackers are cool and you want to put E's and threes instead of E's in your name or whatever because you want to... Because you saw the movie hackers and you're very excited.

Tanya: It's like, what I actually want you to know is just like our secure coding guideline. I just want you to know these are the security headers I want you to use. You know here's an overview of why. If you want to get deep into it, come to my office, but like please just use these headers and these are the settings I'd like. If those don't work for you, come to my office and we'll talk about what we can do to make sure you get your business things done. Right?

Tanya: Like yeah. I feel there's a lot of money to be made in things that are cool and hacking is cool. And just like physical penetration testing, oh my gosh, you do not need the average person to learn that.

Corey: Right. It's the reason you can hire a specialist who do nothing but this all the time.

Tanya: Yeah.

Corey: It's strange and I've always felt somewhat aligned with iInfo Ssec [inaudible 00:14:51] folks just from a perspective of no one cares about the AWS bill and no one cares about security until right after they really, really needed to care about both of those things.

Corey: It's always a trailing function and there's never a great time to come in in advance and say, "Ah, but if you pay me now, you'll save orders of magnitude more in the future." And the response is generally, "Yeah, but we could also spend that time working on feature development instead, and the company is still in business later." And they're not wrong. They're absolutely not wrong.

Tanya: Oh yeah.

Corey: It's, there's a spectrum on both of these sides of things where you can be so good at it, you never get anything else done, and then the company dies. It's always a series of trade-offs and I think that that is something that is not always well understood by folks, especially in the C suite where it's, "Oh we just want to be secure, check the box please, and call it good."

Tanya: Yes.

Corey: There's always going to be tradeoffs. At what level of risk are you comfortable with? And having those conversations is always a difficult discussion to have with various stakeholders.

Tanya: Yes, I cannot agree with you more, Corey? There's a PCI compliance rule that you have to do continuous security testing and it is not explained what that means. And our tool works in real-time and every time you visit something, it tests it. Right? So we wanted to put continuous security testing, but I have been told that CSOs will literally start crying if we say that word.

Tanya: I mean as vendors, all of them apparently are saying that even if it's like you actually have to manually turn on the tool. And so I guess it's the most used word for CSOs at this point and I've been told they're allergic to the word continuous. I should just not use that word at all. I'm like, "Oh, okay, thank you. This is good information to know."

Corey: Don't get me started on the obnoxious challenge that seems to be using the same terms again and again, meaning different things. It's, oh, you sweet summer child. Let me explain to you what that term means you babe swaddled in the cashmere blanket of ignorance. It's always... people use these terms in a bunch of different ways. I mean we see that with definition of terms like cloud native for example, where everyone has a different definition that just happens to align perfectly with the thing they're selling in another market, but there's no broad consensus.

Tanya: Yes. Can I give you like... Since I don't work for a cloud vendor anymore, I'll give you my idea of what cloud native is.

Corey: Please do.

Tanya: No don’t. Then I was hoping you'd make fun of it after.

Corey: Oh, I got that’s for free. If not here then certainly on Twitter.

Tanya: Cloud native is the tools made by that vendor for their cloud that they want to sell you. They made it on purpose for their cloud. It's not going to work in the other cloud. Cloud native.

Corey: I liked that quite a bit, but what about multi-cloud? Remember, you have to be able to go between cloud vendors seamlessly

Tanya: Hmm.

Corey: And effortlessly despite the fact that no one in the history of time has ever done this because if not, we have nothing left to sell you. That's a different definition of cloud native,

Tanya: Oh yes.

Corey: Which means who has contributed enough money to our foundation.

Tanya: Oh, that's such a good point. Yeah. Multi-cloud strategies sound really... Although they're becoming more and more popular, they're very painful looking. Like yeah, it's a lot of tooling that you have to buy that has to get along very well and when you have multiple clouds and you have on-prem and all of these things, how do you keep track of all your stuff and where it is and who's in charge of it and has it been looked at?

Tanya: Definitely that like that is a thing we're trying to do and a lot of other vendors are trying to do, trying to actually give you visibility into all the things. I don't know what's going to happen, Corey, when there's like 50 cloud providers or a hundred or 200.

Corey: I'm not sure there will be. I feel like we're seeing consolidation in that space. You're going to have the big four for lack of a better term.

Tanya: Mm.

Corey: Which four I'm talking about is left as exercise for the reader. But after that you're not going to see much other than the very distant second place folks where they're pushing a strong multi-cloud narrative because if you go all in on one provider, it will certainly not be theirs and then there's going to be a long tail of specialist folks or small operations that target very specific use cases.

Corey: And that in turn is going to be a challenging market. I don't think that we're going to see too much more than that in the platform as a service space. Now, where we will see differentiation is going to be higher level software as a service offerings that solve very specific business problems

Tanya: Oh yeah.

Corey: That don't fit in a single Lambda function or two. And so therefore, they're no longer a trivial exercise for the reader to solve. Instead, it becomes an actual company.

Corey: I use the example of this that I've loved for a long time is PagerDuty where they will... They've solved for the problem of when a thing breaks, wake me up and it sounds like an easy thing to build yourself until you try it and realize, wow, we don't route between this many different providers to get to you across multiple paths in the event of any particular piece of infrastructure dying in a way that they do because they've tackled that entire problem space. You're not going to build a better version of that in your weekend's 20% time.

Tanya: No, definitely not. There are so many kick-ass SaaS tools coming out. Like I have a friend that's a massage therapist and she was showing me that there's... So I live in Canada and I'm from Ontario and I live in British Columbia now, but there's different rules for massage therapists in different provinces just like in America there's different states.

Tanya: And there's a person that has this SaaS tool that I guess like, I don't know how they know all the rules of... Maybe they took massage therapy in school but then also took computer science, but they've made this perfect tool and basically almost every single massage therapist uses it and it's really reasonably priced and it just does every single thing according to like how to book their appointments, how to make sure the taxes are charged correctly, that they you know have a place to put the exact things they have to do to obey all of the rules of their, you know of their certification.

Tanya: And she's just like, "Oh, yeah, everyone uses it. Why?" Like there's literally no point, like the amount of effort you'd have to do, and I think he charges like 130 bucks a year. It's like nothing. And then that person has a full-time job based off of that, and it just... you know and you can talk directly to him if there's a problem. She's like, "Oh yeah, he's a dream." And I feel like SaaS is coming out in a way where it's like making people's lives just so much better.

Tanya: Could you imagine before something like that,

Corey: Oh.

Tanya: Like you'd have to install it on your computer and then you know you're a massage therapist, you're really awesome at what you do, but you're not a technologist, right? And it's like, "Oh, but I didn't back it up and then now everything's gone." No. He does that for you. He does everything, SaaS cloud, awesome."

Corey: That's... It also has really reduced the level of friction to running businesses. I mean, I can't imagine having to build my own payroll system, for example. I'd pay another company to make that go away.

Tanya: Oh, yes.

Corey: Every single piece of noncritical in line with what my company actually does, if I can farm that out to someone else, that becomes a terrific story and an uplifting narrative for all of it. Which is interesting coming at this perspective that you are where you're building a SaaS offering that effectively saw, or not necessarily SaaS, but a tooling story

Tanya: Yeah.

Corey: Around security where your customers need to understand on some level that they're able to outsource work, but they cannot outsource the responsibility. And that's where it feels that companies get wrapped around their own axle.

Tanya: Yes. Oh my gosh, Corey. It's so true. Yeah. I feel like a lot of companies don't know where to start in regards to application security because traditionally, we just we protected the perimeter and then we just walked everything down inside like enterprise security. No administrative rights for you. No installing stuff on your desktop, et cetera. Right. And then now we have all of these old guard security people where they're really good at intrusion prevention, intrusion detection, things like that.

Tanya: But now the weakest point is software, right? That's how if you look at the Verizon breach report, the past three years that they've issued the report, unfortunately, weak application security is like the winner of the cause of the most breaches everywhere, hands down by a landslide every single year, which is bad news, not good news. And but like we have all of these people that are slowly coming towards security that are learning about AppSec, but because it's not being taught in schools really, and it's it just... I guess it's not new, but it is, if that makes sense.

Tanya: Like it's been a problem for a while, but just in the past few years, it's become the weak point because like the security industry or InfoSec industry is really kicking ass in regards to protecting the perimeter and they're really kicking ass in regards to enterprise security and discovering threats. But we're not kicking butt yet in regards to securing our software. And yeah, we were hoping to help. That's our goal.

Tanya: Basically, I only wanted to join a company if we're going to do something brand new. And my friend Aaron and I just kept going back and forth about what we felt the biggest problems were in our space. And some of the biggest problems are, you know, developer education. So we're going to release, so everything that our tool can find, we're going to release videos for free to everyone about how to fix the thing it found. I don't know if you know, but most apps that companies actually charge you extra if you want to learn how to fix the things that it found, if you want to... [inaudible 00:25:00]. .

Corey: We found these things, we won't tell you how to fix it unless you pay us extra.

Tanya: Yeah.

Corey: I've always hated the I know something you don't know, but if unless you pay me, I won't tell you what it is model of pricing.

Tanya: Yes. Yes.

Corey: I have a tee shirt that get printed that I love, which is, it says on it quite simply "teach everything you know," and I try to do that myself. I can talk about any particular aspect of the AWS bill for free and I will, but it turns out doing a deep dive analysis on someone and seeing exactly which things apply to their various environments, that's a whole different series of conversations and I'm not doing that for free.

Tanya: Yeah. That is service.

Corey: That's where I draw the line.

Tanya: Mmhm. But that's you. That's a service, right? So I was saying to Aaron, "Well, okay, so if our tool finds something, I absolutely insist that we're going to teach our customers how to fix the thing that we found. Right?" He's like, "Of course," and I'm like, "And if I'm going to work really hard to like write you know blog posts about it and documentation and videos, it costs us nothing to put it on YouTube and give it away to everyone as opposed to just giving it to our customers." And he's like, "That's a good point." And I'm like, "And then some of those people will see it and maybe they want to be our customers, but for everyone else we'll just have... That means when we surf the internet we'll be safer."

Corey: Yeah.

Tanya: That's what I want." And so he's like, "I'm in, let's do it." So...

Corey: You wouldn't think you'd be asking for a lot, but there you have it.

Tanya: I, well I mean part of, I don't know. Part of me wanting to start a company is so that I can do good and that is grammatically correct the way I'm saying it. I want to do good like Superman does good.

Corey: Oh, yes. You want to do good and do it well.

Tanya: Yeah, exactly. And I feel that one of the things, one of the ways that I could do good is by using my expertise to help the most people possible but without just constantly working for free and being exhausted. Like you were saying, like you know you can share all of the stuff that you know like in a wide range, but if you're going to go into someone's company for the day, you have to charge for your time, right? Because you have a mortgage and bills to pay. So I wanted to calculate ways that I could do good with my life, but still you know pay all the bills. And so this is our compromise. Have you heard of the Effective Altruism movement?

Corey: I have not.

Tanya: So a whole bunch of computer scientists decided they wanted to perform good and they're like, "but we want to do the most good we possibly can." So for instance, like if you donate you know a can of beans to the food bank, right? That's not as valuable as if you give them the money that you paid for that can, that's more effective. But also you can be infinitely more effective by, for instance, giving $34 to the Anti- Malaria Foundation because then they can buy X number of bed nets from people that live in areas where there's lots of malaria. And then with $34 approximately, you can save a person's life because if you give that money away, on average, one person will not catch or however many people will not catch malaria and one of them that would have died will be avoided. Right?

Tanya: And so they've like taken math and statistics and all the information and then they've found a bunch of charities that are the absolutely most effective. And so I am an effective altruist, and so I'm like, "I want to do good, but I want to make sure I do the absolute most good." So yes, I could volunteer to go to the food bank and I could like move cans for them all day, let's say, or drive you know two hours a week delivering food. Right? But I have so much more value that I could deliver to a much bigger audience, if that makes sense. Right?

Tanya: So like I mean, as much as I like to think I'm strong and fit and I could definitely carry a whole bunch of canned foods, that's not like the best use of all my skills of how I could help people. And so yeah, I wanted to work that into our company so that I feel good. If that makes sense. Yeah. But if you do... I don't know. Check out the Effective Altruism movement. It's pretty interesting and it's just it's almost like 95% computer scientists and programmers, people like who for whatever reason all think the same way. It's like, "Let's tackle this head on.

Tanya: Like the Bill and Melinda Gates Foundation would be an excellent example of like effective altruism. Like they look at really big problems as a whole and then attack them strategically as a whole. Like he could just give away all of... or they could just give away all of their money to, I don't know, the food bank as an example, but instead they're trying to tackle like really big systemic problems, and I admire that quite a bit.

Corey: It's a common thing that I think people don't tend to fully grasp, where nonprofits can do the most good is with money. They already have optimized, streamlined pipelines for this. That's why for the tee shirt drive for last week in AWS, I wound up raising money rather than trying to go in and volunteer at a hospital or something like that. It's it just turns into the most effective way to start combating these things is to pick a decent nonprofit aimed at the problem and then give them money. I think that's something that people overlook. You don't want to go volunteer at a soup kitchen. They'd rather have money so they can start to build sustainable programs.

Tanya: Yes. Unless you have a very, very specific skill set. So let's say, Corey, you were going to go into a hospital and volunteer, but what you did was analyze their, their billing for their cloud and help them optimize it so they could save money every month from then on. Assuming the cloud or the hospital was so forward-thinking that they were in the cloud, which is unlikely, but let's pretend. Right? Then that would be a thing that you could do that saved them so much money in the future that is even better than the money that you raised. Right? That's another way to do effective altruism is like if you have a super special skill set.

Corey: With the caveat that not as many people do as think they do.

Tanya: Yes.

Corey: It turns out, for example, that, I don't know, going in to help nonprofits fix their AWS bill, not as big of a problem as you might expect it to be.

Tanya: Oh, I had no idea.

Corey: I've tried it. And that's the challenge is that it seems that very few nonprofits have a significant spend on these sorts of things compared to other drivers there because of donations and the rest and how they wind up doing things. It's a very different market and I was very surprised by that.

Tanya: That's really, really interesting.

Corey: Yeah. There are always exceptions to everything, and if you're listening to this and you're one of those exceptions, hi, get in touch. But there's that's always the weird thing to me is figuring out that the world is never exactly like I expect it to be, but that keeps it fun.

Tanya: I definitely could not agree more with that statement, Corey.

Corey: So where can people learn more about you and what you're up to? Where can they hunt you down, for lack of a better term?

Tanya: Well, you can hunt me down on Twitter, YouTube, Twitch, Medium, dev.to. But basically, if you just look up, SheHacksPurple, you're going to find me. Or if you go to my new company's website, securitysidekick.dev. We have a YouTube channel and a Twitter handle, secsidekick@SecSidekick. And basically, yeah, I am online a lot. If you follow me on Twitter, it's where I announce all my things. I even have a mailing list now. So I'm going to, I'll send you some links after for the podcast notes if you do that.

Corey: Excellent. By all means, you can find them in our show notes.

Tanya: Thank you. But basically just look up SheHacksPurple and that's going to be me with the purplish hair.

Corey: Excellent. Thank you so much for taking the time to speak with me today.

Tanya: Thank you so much for having me, Corey. I'm sorry I got so off topic. I'm really passionate about philanthropy and I guess it-

Corey: It's an important area.

Tanya: It just spills out sometimes. Sorry.

Corey: Of course. No apology needed.

Tanya: Thank you.

Corey: Excellent. Tanya Janca, founder and CEO of Security Sidekick. I'm Corey Quinn. This is Screaming in the Cloud. If you've enjoyed this podcast, please leave it a five star review on iTunes. If you hated this podcast, please leave it a five star review on iTunes.

Announcer: This has been this week's episode of Screaming in the Cloud. You can also find more corey@screaminginthecloud.com or wherever fine snark is sold.

Announcer: This has been a HumblePod Production. Stay humble.

View Details

About Josh Doody
Josh is a salary negotiation coach who helps experienced software developers negotiate job offers from big tech companies like Google and Amazon.

Links Referenced:

  • Twitter: https://twitter.com/joshdoody
  • LinkedIn URL: https://www.linkedin.com/in/joshdoody/
  • Personal site: https://fearlesssalarynegotiation.com/
  • My Coaching Site:: https://fearlesssalarynegotiation.com/coach/
  • A detailed article on how to answer the salary expectations question: https://fearlesssalarynegotiation.com/the-dreaded-salary-question/

TranscriptAnnouncer: Hello and welcome to Screaming in the Cloud with your host Cloud Economist, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode of Screaming in the Cloud has been sponsored by Manifold. Manifold powers marketplace infrastructure that connects millions of developers to the best APIs, tools and services in the fastest growing communities and also Kubernetes. They offer a complete toolkit that allows you to deliver your API first product to millions of developers. Check them out at manifold.co. Again, that's manifold.co.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Josh Doody of fearlesssalarynegotiation.com Josh, welcome to the show.

Josh: Hey Corey, thanks for having me.

Corey: Thanks for being had. So this is a bit of a radical departure from the typical guests that I tend to have on this show in that you don't work for a tech company much less one that's in the Cloud. You're a self-employed consultant from sort of the same school that I tend to come from of self-employed consultants. Namely you find a specific expensive problem and then you go after it whole hog.

Josh: Yeah, I think that's a great description of how I got into what I do and what I do. Yeah.

Corey: Yeah. Mine, as listeners to the show are probably well aware by now, is that I fix the horrifying AWS bill. What problem do you solve?

Josh: I help uh experienced software developers negotiate job offers from big tech companies, so I guess that's not really a problem being solved. The problem that I solve is that a lot of experienced software developers um frequently accept far less pay than they should for their valuable skills, and so my job is to help them arbitrage that gap, close it a little bit and hopefully get paid what they're worth.

Corey: Let me level set this conversation a little bit. Um toward the beginning of this year, Sonia Gupta and I gave a talk a couple of times called Embarrassingly Large Numbers, Salary Negotiation for Humans. I consider myself not bad at salary negotiation. If I were to take a job tomorrow, I would hire you as a coach. Not optionally. That is one of those concrete things that I know I would do. I consider myself good. You are worlds beyond what I'm able to achieve.

Josh: That's um really flattering, really kind of you to say and although I hope that you don't need to hire me anytime soon because what you're doing continues to thrive,

Corey: Absolutely.

Josh: I would welcome the opportunity to work with you. I think it would be a lot of fun.

Corey: My business partner and I, who are co-owners of the company, uh keep going back and forth joking that we're going to hire you to help get an advantage over the other one and then we realize, wait, we just get partnership percentages here so all we have to do is sell more. We can't really do anything other than that. Which it's sort of a shame because on one end I love working with you on the other I kind of like not having a boss because my biggest barrier to employment is of course my personality.

Josh: I feel the same way. I was just having a conversation with some friends last night about day jobs but yesterday was four years to the day since I quit my last day job um and my reasons that I gave them, because most of the people that I know, actually I know a lot of people who are sort of untraditionally employed but the reason I give to them that I quit my day job was I hate meetings and I do not like having a boss. So I can fully sympathize with what you just said.

Corey: So let's start at the very beginning here. When we talk about salary negotiation, when does that take place in the course of, "You know what, I finally hit my rage quit position, I'm going to go find another place to work." Starting from that point, when do I start thinking about compensation in my next role?

Josh: Yeah, you should start thinking about it pretty much right away. And so I'm going to answer this question and maybe take like a quick tangent and then come back. So I'm putting a pin in that sentence. The time when most people think about salary negotiation happening is when they get, a lot of people kind of imagine they're going to get like a formal offer letter on company letterhead or you know, electronic version of that that says, you know, "Welcome to Google. Here is your offer. Sign here to accept." And that does sort of happen eventually. But that actually often will happen after you have fully negotiated everything. And so you kind of have to back up quite a bit from where most people think negotiations happen. They don't happen when you get that formal letter. Usually the formal letter is more of a document of what you have negotiated or what you've agreed to.

Josh: Um they even um begin before maybe the next stop point, if you're moving backwards in time, which would be, you know get an informal offer. I get a verbal offer from a hiring manager or recruiter or maybe I get an email with some bullet points in it that says, "Here's what your comp package will look like. Do you accept this?" And in fact, that's the most common thing that people think about. And I think you and I will talk about that. But now I'm going to go pull that pin out that I put in just a minute ago and zoom all the way back to when you are actually beginning conversations with a company and talking to a recruiter. Frequently the salary negotiation will begin right away and when I say how it begins, most people say, "Ah, that's happened to me and maybe I didn't realize I was negotiating."

Josh: And that is the recruiter will say to you, "Hey, this is exciting. We're really glad that we found you and that we're going to be interviewing you soon for this opportunity. I just have a few questions for you before we start. And one of those is, you know, what are your salary expectations here? What are you hoping to make if this opportunity works out for you?" And that is actually the beginning of the negotiation. There are large mistakes that you can make or avoid in that moment and in conversations that are very much like that. So I think that you should have some idea kind of what you want.

Josh: And I also think that you should not disclose that information in those early kind of screening type, early stage conversations with companies because that's when the negotiation has actually begun. Even though they're usually clever enough to disguise it as though it's like a pre-screen question or an interview question that you might perceive as some sort of a gate that you need to get through to get to the real interview and then the real negotiation. Uh in fact, it is the beginning of the negotiation and they just sneak it in early sometimes.

Corey: Oh absolutely. "Oh thanks for your interest in our company. What's your current compensation or what are you hoping to make now?" It feels like it's the socially accepted attempt at screwing someone over and it's one of the last uh questions someone should ever answer.

Josh: 100% agree with everything you just said.

Corey: At least that's my current approach. It's fine. If you Google salary negotiation for engineers, stuck to the very top of Google is Patio 11 Patrick McKenzie's blog post on this and you're mentioned in it.

Josh: Yeah.

Corey: As someone that he actively recommends in this space. This is something that is almost universally held to be true and people don't believe it or if they believe it, it's an intellectual discussion. It's not something that they internalize.

Corey: From my perspective, the reason to bring you in on a salary negotiation piece isn't because you're necessarily going to tell me anything that I don't already know, but because you see this far more frequently than any job searcher does. When you're talking to a recruiter at the beginning of those conversations, "Well how much are you looking to make?" That's a question that they get to ask four times a day. Whereas you get to answer that question basically once per company that you talk to. They are far more practiced at it. How do you start leveling that distance between your skill level and theirs? From my perspective, you bring in an expert of your own and that's what you do.

Josh: Yeah. I think you hit the nail on the head. I mean, I love the way that you framed sort of even with numbers what's going on in the other side. And I think this is something that makes a lot of my clients maybe a little bit more at ease, but also maybe a little bit more prepared to go in and actually advocate for themselves. And that is when I tell them, "Look, you know this recruiter that you're talking to, whether it's avoiding the salary expectations or current salary questions or it's counter offering or it's asking for more equity, even though they already gave you a little bit of a bump, that kind of thing, that you're working with someone who's a professional salary negotiator, at least as part of their job, and they do this many times a day. And you, although this is a unique opportunity for you, it may be one of five or 10 opportunities you'll pursue in your lifetime."

Josh: That recruiter is probably working with five or 10 people right now. They're going to hang up the phone after they talk to you and they're going to call another person and they're going to say to that person, "Hey, super excited for you to be pursuing this opportunity. You know, what are you hoping to make if you come? They're going to keep asking that question over and over and over again. And they're going to keep answering counter offers over and over, over and over again.

Josh: And I think that it's important to realize that you're not talking to a friend or a family member or a buddy at the bar, you're talking to a professional negotiator and a professional who gets paid based on the number of people they place and a bunch of other metrics. So it's important to understand that you are talking to a professional. I think it's important in that context to comport yourself as a professional, especially with the information asymmetry that you mentioned where they have a lot more data, a lot more reps than you do. And so you're already at a disadvantage. And that's where I come in.

Corey: One of the biggest problems that I suspect that you may see is that, well, let's back up a second. Have you ever met a software engineer? Generally the approach tends to be from their perspective that I'm really good at this very hard thing, which is objectively true. So that probably means I'm good at everything else too. And from there we get hacker news. The problem with that entire approach is that it's provably false.

Corey: But very often when I've been recommending your services to other folks, and just as an interlude by the way, this is not a sponsored episode. This is something that I believe in, but you haven't paid me a dime for this. But sorry. So just to be clear, because it sounds like I'm effectively turning this into a sales pitch for you and yeah, when I believe in something strongly enough, that's what I do. But I've sent people to you before and they've always been a little skeptical of, "I don't know, it might cost a few thousand dollars to have a coach come in and help me with this." And my response has been every time that if they are not happy by the outcome, I will pay their fee for them and no one ever has.

Josh: Yeah. It's really rare that I get to the end even. Sort of for most of my clients I would say kind of you know worst case scenario is that your job offer goes away. That doesn't happen with my clients. But the next worst case scenario is, "Well, I negotiated, I did everything I could and they didn't budge," which is actually, I mean it's not what we want, but it's the kind of outcome where at least you can look at it and say, "You know what? I did everything that I could. I'm going to be really confident that when I go start this job on the first day that I did get the best compensation possible. And I don't have to wonder if the person at the desk next to me is making more than me or if I got shafted or anything like that."

Josh: And so I think that one of ... Next to almost always improving a compensation package and sometimes substantially, the next thing that I offer is the peace of mind of knowing that you had an expert along for the ride and you did the best you could in that one of the few opportunities that you'll have in your career to maximize your salary. And by the way that you now know the process that you can use to negotiate future salaries and that you can help your friends and family use to negotiate their salaries if you like. And so I think just having that knowledge is super valuable. In addition, you often get a bigger paycheck as a result.

Josh: I think that's why you know it's very rare that somebody will come to me afterwards and say, "I'm disappointed in this." Um usually it's, "Well, you know we did our best and we wish we could've gotten a better result, but we got the best result that we could," or, "Wow, I can't believe that much money was available to negotiate because I was willing to just accept it as is because I thought it was a good offer."

Josh: So you know I think that there's a lot of value just in following a process and like you said, sort of outsourcing. You know I was trying to think of the name of ... There's a fallacy, like a name for that fallacy where you're good at one thing and therefore you assume you're good at all things. Um but you know, I am a software developer and I know how hard it is to write software and I also now know how hard it is to negotiate job offers and become an expert in that area. And I think that something I actually kind of admire about a lot of software developers is they do understand that idea of specialization and comparative advantage and they are usually willing to put an economic value on it.

Josh: It's one of the things I like about working with software developers is that usually, although there are many who populate ACRA news and other sites, many of them understand, "I'm really good at this thing. Josh is really good at that thing. I don't want to get really good at that thing, so I'm going to hire Josh to do that thing for me so I can keep focusing on my area of expertise."

Corey: When you take a look across the landscape, what is the number one thing people screw up on the most?

Josh: It's the same-

Corey: Let's bound that to salary negotiation because I have a laundry list of other things that are mostly personal failure.

Josh: I'll just close the book here. Um so the biggest thing that people screw up the most is where I started and it's why I talk about this so much. I've spilled more ink, I've said more words about this than I think any other topic which is do not tell companies what you would like to make or what you are currently making. It is not their business and it will definitely trip you up. So that's by far the biggest mistake people make and it's one of the most difficult to sort of unring as a bell.

Josh: Um it can be un-rung and there are ways to work around it, but it makes your life a lot harder. Whether you work with me or you try to negotiate on your own or even just, you know you move from company to company, you will continually find that you're behind the pay curve because you answer the salary expectations question. So I think the next biggest mistake ... So I think we're taking that one for granted. I wanted to emphasize it because it's so important and because it's so common, the next biggest mistake that people make is also maybe a little obvious and that is they just don't negotiate.

Josh: So very frequently somebody will come to me and say, "I was thinking about hiring you, but I got this offer and it's a really good offer. So you know I'm just not comfortable. Like I don't want to ruffle feathers or kind of build a bad reputation. So I talked to my family about it and they don't think it's a good idea to negotiate. So I think I'm just going to take it. It's really good. It was better than I was expecting. So I'm just going to do it. You know I was thinking about hiring you but I think this is good enough." And for me the reason that's a mistake is really non-obvious. And this reminds me of like the idea of the seen and the unseen and what the seen is, yes, you got a good offer that exceeds what you're expecting.

Josh: The unseen is you don't know how much better that offer could be and you don't know why that offer exceeds your expectations. And frequently the reason that offer exceeds your expectations are that your expectations were miscalibrated. That you had bad data or you didn't do your research or you went with your gut instead of looking at levels at FYI or Paysa or Glassdoor or talking to colleagues in this space for really niche kind of opportunities and you underestimated your value to that company. And therefore the offer they made you reflects maybe the lower end of the value that you actually bring to the company and may also exceed the value that you anticipated you would bring to the company.

Josh: And so anytime I hear someone say that, I say, "You need to realign your baseline. Your estimate was off. Their data is better than yours. And they're telling you your estimate of what you should be paid is too low and that's all the more reason to negotiate." Usually what that means for me as a salary negotiation coach is there's actually more room to negotiate than there normally would be in an opportunity because you underestimated your value so much and they probably are offering a lowish number that just seems really high to you.

Corey: My default response when people ask, "So what would you like to make in this next role," is, "At this point about $80 million a year plus a company helicopter." At which point the worst thing anyone could ever come back to that with is, "Okay." Because then I know I should have asked for more money. There is no right answer that doesn't have the potential of completely undermining your entire argument. It's better not to play those games.

Josh: Yeah, you're describing, I call it the bad yes. Um and I think a lot of people have experienced this actually in their career. And that is where they say, "What do you want to make," or you throw a number out there, maybe you are negotiating, but you say a number, you know they offer you 90 and you say, "How about a 100," and they go, "Okay." And then you're like, "Yes, I got it. I got exactly what I asked for. This is amazing. I negotiated so well. I went from 90 to a 100," and then usually within a matter of minutes you think, "Wait a minute, they said yes on a 100. What would they have said at 101 or 5 or 10. How high could I have pushed them?" And what you realize is like you said, "$80 million and a company helicopter, maybe 81 million and a company helicopter was on the table and I left a million dollars sitting out there. Could have done a lot with that million dollars." Right?

Corey: And you've got to be careful because they're going to give you the crappy last year's model of the helicopter.

Josh: Right. You didn't even specify what kind of helicopter and maybe you should have gone for a Gulfstream or something.

Corey: Exactly. There's always room for negotiation. One question I do want to ask you though, um as one white guy in tech to another who's no longer in tech but now tech adjacent, what about people that don't look and sound like us? How much of what you do maps to members of underrepresented groups or honestly let's phrase that differently, how much of this maps to folks who are not a member of our very specific over-represented group?

Josh: Yeah, it's a good question and so I can give you my answer to it and then the asterisk that's right next to it is I'm working from a limited data set and that is from many, people that I've coached relative to me as one person. I think I'm at I don't know, 60 to 80 clients that I've coached now, something like that. And so that's big for me. But relative to the size of tech, it's nothing. And so that's kind of my caveat is the answer I'm giving you is based on my data and what I've seen and then maybe some opinion that I'll throw in. So your question was how much of what I do maps to other people who are not necessarily like me. And the answer is as far as I can tell, 100%. And in fact I think it may map more because a lot of times those people are the same people who haven't negotiated before for different reasons.

Josh: And I think you know we would have to like hone in on different groups to talk about what those different reasons are. But I think you know if we take you know us and then kind of everyone who isn't quite like us that we could then segment the not quite like us folks into lots of different groups with many reasons for not negotiating in their past. And so they end up sort of behind a pay curve sometimes. And so the things that I do, the process that I follow, which I follow with all of my clients, sort of regardless of background and that sort of thing um is effective for most people. And I think more effective for people who for whatever reason, find themselves currently behind the pay curve. And so the process works for everyone. I've worked with lots of different people from different nationalities, from different countries, male and female, everything across the board.

Josh: I've worked with all kinds of people and I find that the process that I follow works because at the end of the day, I think from the business side, certainly in terms of salary negotiation, it's a business transaction. A negotiation is a business transaction that they are uh working with you on. And so for them it's usually driven more by a spreadsheet with numbers and financial data on it than it is anything else. And a lot of the other things that could factor in are sort of personal biases on uh the client side or on the company side and that sort of thing. But I think those are usually swamped by just negotiating effectively and following the playbook that I put together, which is designed to sort of work within the playbook that most of the big tech companies are using to find and recruit good talent.

Corey: And that's really I guess what I want to talk to you about for I guess the second part of this episode, which is what flexibility, what wiggle room do you have when negotiating with large tech companies? I'm going to pick on Amazon for a couple of reasons. One, they're the largest player in this space so they can suffer my slings and arrows. Two, I've gotten a job offer from Amazon that I turned down, so I know how some of it's structured. And three, it mostly amuses me to make fun of them for the following thing. They're a company that's valued at about a trillion dollars, but one of the 14 leadership principles that they love is frugality, internally, to my understanding. It's referred to much more frequently as frupidity. But we're an enormous company, but we like to save money so we don't pay extravagantly it turns out is not one of the most compelling value propositions to work somewhere.

Josh: No, it's not. And for whatever reason, what you just said reminded me a lot of like our safety is always on his soapbox about how silly it is to not drink frappuccinos if you want to get rich one day. Um and I think it's kind of one of the modern versions of you know penny wise and pound foolish. And that sounds a lot like what you're describing um with Amazon, who we happen to be picking on or any other company that says, "You know what, we're going to try and save money on salaries." Um for me, even if before I was a salary negotiation coach, that's a really short sighted way to look at things, especially in a market ... You know so I'm thinking of the U.S. mostly right now, but like our unemployment rate rate is lower than it has been in a very long time.

Josh: Uh it's a very hot labor market. There's a lot of demand for employees in general and even more demand for technical people in general. And then you keep going up that pyramid to like more specialized people, software engineers, machine learning folks, SREs, dev ops. You go up the pyramid and there's a ton of demand for these people. So even if you're Amazon and you save a few bucks hiring someone by you know not giving them an optimal offer or not playing ball as much as um maybe you could when they negotiate, that person's going to find a home at another competitor very soon. Especially, you know we've talked about Cloud companies. I mean Oracle, huge in Cloud, Google, Microsoft, those companies are trying to find people too. Um and so I think it's just really shortsighted.

Josh: So that's not really kind of related to what you're saying but that's my soapbox is how silly it is to do that um and to try and save a few bucks and be frugal on salaries for something that is not a commodity, but is a very you know highly demanded, specialized skill set that lots of companies need. That's the time that you need to pay the premium to try and get good people in and avoid turnover.

Josh: Um how much room is there to negotiate? It's a good question and most of my clients will ask me this before we work together. My answer is always the same and that is somewhere between, not at all and a lot. And it's hard for me to know whether it's not at all or a lot until we actually negotiate. Um usually I can get a bead on that by just talking to the person, either in the intro call that I do or in our kickoff call, which is the first thing we do after we decide to work together and I'll find out more about their background and how they came to find out about the role and what their resume looks like. And you know what team they're going to at Amazon and that sort of thing.

Josh: Um and even then, even if I have an intuition, I'm like 60% you're going to get a good, a really good result here. A big result, a lot of money. Um there's that 40% where I just don't know. And there seems to be sort of a lot of randomness involved there. Um and so sometimes it's nothing. I had um earlier this year, this isn't Amazon, but I had a client who went to Google and I thought, "My goodness, she's going to clean up." I mean she got a good offer as an L5 but I know from the numbers that I see that there was a lot of room above that in the L5 pay band and they wouldn't budge. And I still don't know why. It's very confusing to me. Um and I don't think it had anything to do with her. I think it had to do with maybe departmental budgets or some other nonsense that I couldn't see. But it was very surprising.

Josh: Whereas I've had other clients who their resume looks pretty good and then we negotiate it and wow, there was a lot of flexibility there. And so it's not random from the Amazon side, but from my side it can be perceived that way. And so the best that I can do is follow a process, listen to what the recruiter is saying to my candidate or my client, listen to my client, get their history and negotiate as well as possible.

Josh: Um so the low end is zero. The upper end really depends um and like you said, is obfuscated very much by Amazon's weird pay structure. So upper end for salary is often capped at Amazon depending on, I think it's geography, basically. It's somewhere between 165 and 185k and then once you cap that base salary, then you move over to a combination of equity and sign on bonuses that are designed to have um a very weird vesting structure and schedule on the equity. But that is designed to look like a sort of flat annual total compensation number over four years.

Josh: Um sometimes there can be a lot of flexibility there, um especially in terms of the equity and then the sign on bonus trails that for sort of esoteric reasons about the way they structure their offers. So I would say a really good result for the typical experienced software developer would be if they got 50% more equity, that's usually a pretty big win. Um and I would say five to 10% more equity is an okay when it's not great, but it's not bad.

Corey: My default advice that I tend to give people who are looking at equity offers is I tend to value equity at zero and this is either wonderful or terrible advice depending upon the perspective. You work for big tech companies that are publicly traded. It is pretty easy to get an approximate value of an RSU since they trade in the open market. If I'm offering you options in my bullshit Twitter for pet startup, then you can't exercise those. You can't do anything with them until there's an exit event or it's unlocked on the secondary market, which is not a guarantee. So at that point you're more or less getting paid in lottery tickets. Am I mistaken on any of that?

Josh: No, I think everything that you just said is 100% correct. And I am struggling to think of anything that I would add to it. That sounds right to me. That's the way I look at it too.

Corey: So it seems to me that again, this is San Francisco, and I understand this is not necessarily the typical common case, but it feels like anywhere north of $200,000 a year in base salary starts to get harder and harder and harder to find where everyone seems to want to make that up with equity grants. And those equity grants can be massive, but you're still making at most maybe a quarter million bucks a year in cash. There are a couple of notable exceptions to that, but that's the perception I'm getting from talking to people in my peer group. Is that accurate?

Josh: Yeah, I would say it is. So I've seen, um like I said, I mean at Amazon they explicitly cap pay salaries, so that's an easy answer. But even for the other big companies, Google, Apple, Facebook, Microsoft, um I've found that even for really experienced software developers that salaries tend to cluster in that sort of 150 to 200k band um and they will sometimes exceed that for special people. But I've worked with more than one client, more than one of those companies, so this isn't just one company that I'm talking about where um I've seen a quarter million dollar base salary. This was for basically a unicorn for one of these companies that had a very special skill set um that was coming from another company where they used to use that very special skill set and they managed to get 250k as an individual contributor base salary and then a ton of equity on top of it. That's an outlier.

Josh: Um whereas you know at Google, it seems like the salaries are sometimes lower, 150 to 175. For other companies it seems like they'll push that 200 number. Apple will push that 200 number. And then they start to just pile on more and more and more equity. So it's a very weird thing where you can't just look ... If you show me a job offer and all you show me is the base salary, I can't do a whole lot with it in terms of telling you what kind of person got that offer, what level are they, how experienced are they? Because I've seen machine learning PhD's get an offer that's 175 base and I've seen experienced sort of you know full-stack software developers get offers that have about 175 base.

Josh: But then when you flip the page and you look at the equity say, "Ah, there's the difference." And so yes, those companies right now are very interested in piling on equity and their base salary numbers tend to cluster between 150 and 200. I know we haven't said this out loud and I've alluded to it, so I'll just say it. So Amazon is an outlier because they do the same thing, but their vesting schedule is not a flat four year vesting schedule like the other big tech companies. So most of the time you get, let's say $100,000 in equity, it's going to vest 25% a year for four years, 25k in the first year, at the end of the first year, 25k the second year. You know usually monthly or quarterly or six months or something like that. And it goes on until you vest all of it.

Josh: Whereas at Amazon you get a 100k equity. The first year you only vest 5%, the second year vest, 15% and then the third and fourth years it's 40%. And what they do, I mentioned earlier that they try to still show on paper that you're making essentially a flat total compensation every year. And the way they do that is in that 5% and 15% year one and year two they also add a sign on bonus that closes some of the gap. And so you might get a 20k sign on bonus and like a 15k sign on bonus in year two from Amazon. But all of them are basically doing the same thing, which is we're going to pay you a pretty good salary. We're not really going to pay heavily on salary. We're going to go between 150k and 200 and then depending on essentially kind of how specialized you are, how experienced you are, what team you're going to, we're going to pile on equity and maybe sign on bonuses to really convince you to join our team as opposed to going to a competitor.

Corey: I'm a big believer that pay transparency among colleagues is incredibly important because companies already know how much everyone makes. Not just amongst the people they hire, but among other companies too. Because people do tell what their compensation is when they're confronted with that question. So to that end, I want to throw into the public domain I guess the offer that I got from Amazon last year. I wrote a blog post about why I turned down an AWS job offer. Did pretty well. Specifically, it was along the lines of not wanting to suffer their 18 month post-employment non-compete.

Corey: But I didn't negotiate anything on comp because that wouldn't get moved and rendered the entire thing moot. But for Amazon folks, this was an L7 job offer. Their salary was capped, I think at 160 and then the first year a cash comp because of the bonus came in just a hair north of $350,000 which sounds ludicrous and nuts to anyone who's not in a tech city doing these things longer term. But I know people at Amazon who are making two, three, four times that. There's always a bigger fish. And historically my being embarrassed to talk about numbers like that didn't serve employees super well. So I'm actively getting out of that mode now. So if you hear that and think that that's boasting, great. Good for you. That's not the way I'm intending that to come across. It's a data point that people who very often don't look like us, will be able to use internally.

Josh: Yeah. And I'll say it's funny because I hear those numbers right. And of course I have a reaction to that based on all the other people that I've coached and I've seen AWS offers and other Amazon offers. And so my response to that just for anybody who's listening and wonder like, "How good is that offer?" And my response it was, "That's pretty good." So I'm not like, "Wow, I've never seen anything like that." And it's definitely not a low offer. But for an L7 it's pretty good. And I think it is important to have that kind of transparency. So I'll use ... I think I wrote about this in my book, but maybe not. I think it is in my book. This is one of my favorite hacks to kind of ... Especially in the U.S. we have this social stigma with discussing salaries that you just talked about.

Josh: You're trying to overcome that a little bit by publicly writing about an offer that you got, which hopefully will at the very least people will see that offer and they say, "Well shoot, I should be negotiating them. They've got a lot of money to spend on salaries and equity." Um but the hack is, obviously it's considered sort of rude in the U.S. To say, "Hey, what salary are you making," either to somebody that you currently work with or to like maybe a future colleague or somebody who referred you to a position at one of these companies. Referral programs are really big.

Josh: And so what I suggest that you do, and this is almost entirely transparent and yet it frees people up to be honest. And that is to say, "Hey, if somebody were to be hired in, you know, I'm looking at a position that's a lot like yours. If we hired somebody right now or if your company hired somebody right now with your resume and experience, you know your experience able to do the kind of job that you do, like what do you think that they would pay them coming in to work for your company?"

Josh: So basically what you're saying is like, "What's your salary?" But you're saying it in a way that's like, "Let's anonymize this. Let's pretend we're talking about an anonymous third party or hypothetical or whatever and then you can tell me." And so I really liked that it takes the pressure off and that way the person sort of has this plausible deniability that, "Well I'm not telling you my salary, I'm just telling you what I think somebody who has the same experience and resume as me in the same position might make. And by the way, the most pertinent data point that I have to share would probably be my own compensation. And so I'm probably going to tell you either my actual compensation or something in the ballpark." So that's something I like to do. I find that people are very comfortable using that for both sides um and it's a good way to get some of that data.

Corey: Anytime you can contribute what you make to a larger database and helping other folks, it only helps. There's this weird perception that, especially among the folks I guess who come from a background like I do, where I'm not worth even a small percentage of what I'm being paid. So I'm embarrassed and I don't talk about what I make. Well, what if I am in fact being underpaid compared to other folks? What if, more to the point, people who are my peers who are in many cases better than I am, aren't making anywhere near as much as I am? Wouldn't that hurt me? No, it's going to help them. That's the entire point. It is not a zero sum game.

Josh: Yeah, and I liked that you mentioned like kind of contributing to databases, so a site that ... It's funny because people will come to me and they'll say, "I have this offer from you know AWS or from Google, Microsoft," whatever it is, and, um you know "How good do you think this offer is?" And what I do inevitably is I go to levels at FYI and I look and I say, "Well, is it an L6? Okay, let's look at the," because there's a lot of data that is publicly available for people who uh subscribe to what you just said, which is let's just get that data out there. You can use Paysa, you can use Glassdoor. You have to take them with, weight them appropriately. There's some selection bias and some other types of bias that go into the data that you see there, but at the very least, look at Levels.fyi and um contribute to it when you can.

Josh: Also some companies, I know at Google there's a giant spreadsheet that has lots and lots and lots and lots of internal salaries and equity and all that kind of stuff in it. And I don't think it's very hard to find. Um if I know about it, it's definitely not hard to find. And so I suspect there is similar you know Google sheets or you know spreadsheets that are floating around other companies that do something similar. And so you know finding that, contributing to it if you can, I think is really helpful um to try and level the playing field.

Josh: In a weird way I'd like to see a world where my job isn't necessary. It is necessary. And unfortunately, you know I work with one person at a time and so I'm really not going to make that big a dent in the universe from my chair. But at the very least, if what I do encourages people to negotiate but also to share stories of negotiation and to share data with other people and to kind of start overcoming that social stigma with talking salaries, then maybe we can at least start kind of normalizing and leveling the salaries that people are paid and stop some of those people from being sort of low end outliers where they have no idea that they're underpaid by you know 50k a year or a 100k a year or more.

Corey: And I think that's probably all we have time for today. But there's so much more to say on this topic. If people want to learn more, where can they find you? I'm going to assume that there are options on a spectrum between never think about you again and throw a giant pile of money at you for coaching.

Josh: Yeah, that's probably true. So probably the easiest way to find me in an accessible way where I'll kind of respond right away to you is on Twitter. I'm Josh Doody on Twitter and then my knowledge base that I've built on salary negotiation and other career things is at fearlesssalarynegotiation.com. Lots and lots and lots of articles from interviewing, asking for raises, negotiating offers, lots of content that's specifically written for software developers and experienced software developers to help interview better and negotiate. Um and then I do have my coaching offering, which is sort of at the other end of the spectrum that you mentioned, which is fearlesssalarynegotiation.com/coach um and that's where I describe uh my salary negotiation for experienced software developers going to big tech companies, a process, what that looks like, pricing and all that stuff. And there's an application there that you can fill out if you'd like to talk to me about an offer that you have or that you're anticipating later on.

Corey: Terrific. And we'll put links to all of that in the show notes. Josh, thank you so much for taking the time to speak with me.

Josh: Thanks for having me on, Corey. This is a lot of fun. I love talking about this stuff and I hope this has been useful to somebody.

Corey: As do I. Um Josh Doody, Fearless Salary Negotiation, salary negotiation coach, and frankly just a good person to pay attention to if you'd like to make more money.

Corey: I'm Corey Quinn and this is Screaming in the Cloud. If you've enjoyed this episode, please give us five stars on iTunes. If you've hated this episode and want us to go back to talking about technology, please give this show five stars on iTunes.

Announcer: This has been this week's episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com or wherever fine snark is sold.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Paul Dix
Paul Dix is the creator of InfluxDB. He has helped build software for startups, large companies and organizations like Microsoft, Google, McAfee, Thomson Reuters, and Air Force Space Command. He is the series editor for Addison Wesley’s Data & Analytics book and video series. In 2010 Paul wrote the book Service-Oriented Design with Ruby and Rails for Addison Wesley’s. In 2009 he started the NYC Machine Learning Meetup, which now has over 7,000 members. Paul holds a degree in computer science from Columbia University.

Links Referenced:

  • Twitter Username: @pauldix
  • LinkedIn URL: https://www.linkedin.com/in/pauldix/
  • Personal site: pauldix.net
  • Company site: www.influxdata.com

TranscriptAnnouncer: Hello, and welcome to Screaming in the Cloud with your host, cloud economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. This week's episode of Screaming in the Cloud is sponsored by InfluxData, makers of InfluxDB. As a part of that sponsorship, they have generously provided one of their co-founders to have this conversation. Paul Dix, welcome to the show.

Paul: Thanks, Corey. Glad to be here.

Corey: Thank you for taking the time to entertain my ridiculous nonsense. It's always appreciated. I guess where I want to start on this is for... Let's begin at the very start of all of this. You are makers of the premier offering in the world of time series databases. For those of us whose platonic ideal of a database is Route 53, what is a time series database, and why might I need one?

Paul: Yeah. So, uh I mean, a time series database is basically just a database that's optimized for a specific kind of workload, which is time series data. Now, the thing that makes time series data different than, say, reference data that you keep in a relational database uh is that it's largely an append-only workload. Right? You have new data arriving all the time. You're not updating previous records, uh and when you query the data, you're frequently uh creating large ranges of data to compute like summaries of you know what was the min value in these five-minute increments for the last four hours or something like that.

Paul: Um, so you can certainly use other types of databases to store time series data. You can use relational databases or other NoSQL databases, but for... for time series data specifically, there are optimizations that you can make to deal with the very high rate of ingest, the query workloads, which are very, very different, and one other... A few other things.

Paul: Uh one, which is it's very common in time series to have your high-precision data that you keep around for a limited window of time like, say, I'm going to keep all my raw data for seven days, and then you want to uh summarize it or downsample it and, say, keep those summarizations around for longer periods of time like three months or a year, whatever.

Paul: Uh and a good time series solution database will basically uh handle that data management life cycle for you automatically, evicting the high-precision data, downsampling the other data. Um so the eviction I think is actually really interesting from a database perspective uh when you think about relational databases. Right?

Paul: So in the naive case, if you're going to evict your time series data and, you say, say you want to keep it around for just a day, what that means is the naive way of doing this is every time I do a write, I have to delete the oldest data point, right, because if I'm ingesting at a fixed rate, then I know for every write that goes in, there's a delete happening.

Paul: And regular databases aren't designed for this workload. They actually assume that you want to keep most of your data around for all time, so deletes actually are expensive. So time series databases do things like that. They optimize the ingest, the eviction of high-precision data, the downsampling, and the summarizations that you might want to do in real time.

Corey: So I played a little bit with things like this once upon a time in my first life as a network admin. Uh we ran Cacti to manage a lot of these things. Whose, if you've never used Cacti, the primary purpose of that software is to sit on your network, be written in PHP, and be exploited, and use as an attack platform for the rest of your network.

Corey: It displayed graphs using our RRDtool or RRD-based uh tools out of the can of MRTG and a bunch of other similar products, and it's exactly what you described. As you look back further in time, the data gets less and less granular under the baseline assumption that you won't need to have that level of insight and visibility into things that happened a month ago as you might yesterday. Is that aligned with the same principles?

Paul: Yeah, that's similar for the most part. Um there is an important distinction with RRD that I like to make, which is um when I think of time series data, I think of two types of data. Uh what's called a regular time series, which is samples taken at fixed intervals of time like once every 10 seconds, or once a minute, or once an hour. And then, there are irregular time series, which are basically event streams. That could be individual requests to an API and their response times. That could be a container spinning up or shutting down, uh any sort of an exception and application, any sort of event.

Paul: Now, RRD, and uh I guess it's kind of like a spiritual successor, which would be probably Graphite. Those are based around storing regular time series data. So regular time series data is basically a summarization of some underlying like distribution, or raw event stream, or whatever. So for example, if I store the response time for every request made into my API, that's an event stream. That's an irregular time series, but I can query that time series and say, "Give me the 95th percentile in 10-minute intervals for the last eight hours of time." And what you've done there is you've created a regular time series from the underlying irregular event stream.

Paul: So essentially, when you think about putting data into RRD, you're summarizing your data before it ever goes into the database, um and with Influx, what we wanted to do from the very early stage was we wanted to be able to store the raw event stream as well as the summarizations.

Corey: If we go back to when you said you made your first commit to uh what became InfluxDB back in 2013. Uh if I, I didn't notice it at the time, and if I had, I would have made blistering fun of it. It's, "Oh, you're building your own database engine. I know you, you're that guy from Hacker News come to life." And it sounds like something that only someone who's deranged would do it except for the fact that it worked. You've built a successful company. You are a name brand in the time series space, and you were clearly right about this. I mean, now we have other entrance to the space, which we'll get to in a little bit. But at the time, did it feel like you were potentially making a catastrophic mistake?

Paul: Uh it really didn't. Uh and that's based on my experience from a few other times. So in 2010, I was working for a FinTech startup, uh and we had to essentially build a you know "time series database" or "time series solution" for tracking uh you know real-time pricing for like 200,000 different uh financial instruments. Uh and the solution I built was basically Scala web services uh using Cassandra as the underlying longterm data store and Redis as a real-time indexing layer.

Paul: Um and then uh, later on, when I built uh the first product that we built actually, which was a SaaS platform for real-time metrics and monitoring in the server monitoring space, uh I had to build a time series solution again so... and this was basically two completely different problem domains, but the solution was the exact same. Like in that case, for the first version of this, I did Scala web services on top of Cassandra with Redis, and then I built a next version of that API, uh which I used Go for, and this was in... I guess late 2012 is when I started development on that. So I used Go, and I used the LevelDB, which is an open-source library that Google had at the time, and I built this whole thing. So I essentially built a database for that API.

Paul: And the thing that I realized through the process of doing this was that for solving this, this time series use case problem, you could use a general purpose database. You could use Cassandra or you could use MySQL or Postgres, whatever. But the thing is I had to write this like mountain of application code of web services code in Scala to make the whole thing work. And my feeling in 2013 was there's nobody focused on this exclusively.

Paul: Graphite at that time was largely an abandoned project. Uh everybody you know who was using it was complaining about the fact that it wouldn't scale, uh so I thought, "Okay. Nobody is focused on this, but here I am." I've had to solve this problem multiple times in the past few years. I saw people at large companies trying to solve the problem themselves, and I saw all the monitoring companies doing the same thing, so I basically thought, "Here's you know here's something... Here is a need that isn't being served, uh so I might as well go do it."

Paul: And I think the other important thing is like to think of the timing like... So in 2013, obviously, NoSQL was a big thing. And you know you have the different players in the space, and it wasn't obvious how that was going to shake out. Although, MongoDB was obviously already very popular, um but it wasn't like... I feel like over the last like two or three years, there's literally like a new time series database or a new database of some kind every other week that's on the front page of Hacker News. So I, I feel like there was a little bit less like new database fatigue in 2013 than there is now.

Corey: Yeah. It's, it's one of those areas where there are so many different database options that... uh to rip off the ancient JWZ quote, it's, "Oh, I have a problem. I'll use regular expressions. Now, that way, I have two problems." It's, "Oh, I'll just write my own database." And invariably, in 2019, it feels like that's exactly the wrong direction to go in, but counterpoint, you folks have recently released Influx 2, so what's the story behind that?

Paul: I mean, I guess... So first off, I should probably say that I'm uh, I'm a firm believer in what I call polyglot persistence, which is you use the right tool for the job and not every single persistence need is the same as others. Right? For some, you absolutely will need uh a relational-transactional database, and for other things, you won't need that. Um and you know all of this would be kind of like a moot point. If relational-transactional databases were infinitely scalable and infinitely performed, right, then we would just use those. But that's not the case, right? You make trade-offs and optimizations based on your needs and the specific use case.

Paul: So initially, you know with InfluxDB 1.X, that's what that was about. But we started with the database, and then we saw that there were other needs that people had in this time series use case. Right? And for, for time series, like what I realized is it's an abstraction that works well for solving problems in multiple domains. Right? Server monitoring, I mentioned. Financial market data is one, uh real time analytics, but also, sensor data of all kinds, be it industrial, oil and gas, wearables, consumer tech. All that kind of stuff.

Paul: So people had other needs to solve these problems and to build applications on top of like this time series abstraction. They had to collect it. They had to store it and query it. They had to process it for either doing ETL for enrichment or for doing monitoring and alerting. Uh and then, finally, they had to summarize the data for human consumption either through visualization, or reporting, or other kinds of things.

Paul: So as I built the company, uh you know I raised capital and hired developers to build these other pieces, uh and we learned a lot over that period of time over the last six years. Um and the thing I realized is the... What I wanted is a platform that was easy for developers to use that kind of encapsulated all of that. Right now, we have... In 1.X, we have four separate products. With 2.0, what we tried to do is combine them into like one cohesive whole where there's a single API that is you know consistent, that's easy to use. Uh you know there's a swagger definition for it, and there's a user interface on top of it.

Paul: And then, the last bit with uh 2.0 is... You know with InfluxDB 1, we had a query language that looks very much like SQL, uh and that's because I thought it would be easier for people to pick up. And it certainly was, but what we found is there were more complex like analytics and processing tasks that people wanted to accomplish that they wanted to push down into the database.

Paul: And because they couldn't do that, a common pattern emerged where people would write code in you know whatever language they choose, right, like Python, or Ruby, or apparently your favorite, PHP. Uh then, uh they would query data out of the database, and do some post-processing, and then write data back into the database so that they could get it back into the tool chain for monitoring it, for visualizing it, and all these other things.

Paul: So when we created 2.0, we decided, "Let's create a new language called Flux, which is not just a query language, but it's also a query planner, a query optimizer, and it's a it's a scripting language so you can push down this kind of complex processing into the data platform." And as a language, you know we want it to be turning complete and generally useful, but we also want it to be able to pull in data from sources outside of InfluxDB.

Paul: As I mentioned, I'm you know I'm a firm believer in polyglot persistence, and what that means is... You know Influx is great for time series data, but it's not good for reference data, so we want to be able to pull in you know data from Postgres or MySQL, or from any sort of third-party API that you want to pull data in from. Right? You could hit GitHub for data that you could mix and match with your time series data. So basically 2.0 is the realization of you know collapsing those four components into one cohesive whole and creating a language that allows people to define really complex analytics and processing that the platform will just do for them.

Corey: It seems almost like you're going through Hacker News to some extent and picking all the terrible ideas at once. So you just mentioned, for example, that you built a new query language called Flux. Um two issues.

Paul: Mmhmm.

Corey: One, writing your own language is always one of those things that's fraught with peril, but based upon what you demonstrated, I will absolutely extend credence to that that you're probably doing the right thing.

Corey: But I will say that from my experience, working in tech for entirely too long, which is where I guess this bitter cynicism all has root, I find that whenever I have to learn a specific language to use a particular tool, it means tears before bedtime, and I want to wind up calling out a bunch of different companies that have done this, but it's unfair because it seems that every time I've dealt with this specific DSL, you have these problems.

Corey: I'm still going to maintain that Kubernetes wrote their own custom DSL called YAML, which is so historically incorrect that I don't even know where to begin, but that's why I like saying those controversial things. What, I guess what made you decide to do this uh I guess in the face of historical terrible experiences with these?

Paul: Yeah. By the way, I think Kubernetes was... That's not original to create a YAML DSL. I think they're just copying uh Spring and Struts who created a DSL and uh that...

Corey: Once upon a time, we wound up adding Jinja to YAML and called... well that were effectively turned into SaltStack for its configuration language. Again, we're all code terrorists in our own way.

Paul: Yeah, yeah. No. I mean, it's absolutely fair. I agree. Like uh you know generally speaking, like why would you create your own language? There, there are countless other languages out there. And uh so you know so basically, like one option is we just go with SQL. Right? Well, one, creating an actual SQL-compliant SQL is really, really hard. It's a lot of work. Two, SQL isn't turing complete. It's actually not a programming language. It's a declarative scripting language.

Paul: Now, Microsoft's version of SQL has extensions that makes it turing complete. Uh Oracle's version of SQL has the same so... But then, again, like you're not using standards-compliant SQL, and really like even when you get down to it, every single major database has differences between what their versions of SQL are. So you know there's a standard, which is the lowest common denominator, and then when you get into more powerful query functionality, you end up getting into the specific database implementation.

Paul: And as you mentioned, like there's so many tools that have their own languages. Basically, like I think any analytics tool in existence, whether it's log analytics, user analytics, business intelligence, marketing analytics, they all have their own custom query languages, uh and that hasn't, certainly hasn't stopped them from becoming popular, but let me speak to uh one thing about our specific journey to Flux that I think is relevant, which is you know in 2013, I created InfluxDB with this language, this query language that looked like SQL, but it wasn't actually SQL. It was different in ways that are actually frustrating if you're a SQL expert and you try to use it. Um but at the time, like tons of people picked it up because you know most people actually aren't writing SQL day-to-day. They are using their ORM.

Paul: Now, I personally had a viewpoint probably in the fall of 2014 that the SQL style of writing queries was maybe not the best way to work with time series data, which I basically viewed as just like ordered streams of data coming through. And I thought a functional style language would actually be the better, the better way to represent the query style.

Paul: Now, I wouldn't make... I was too afraid to make that change at the time, but when we introduced our processing agent capacitor, which is there for like background ETL and monitoring, alerting, and real-time processing, uh when we introduced that in September of 2015, it had a language that was more functional in nature, so we we again like made the foolish mistake of creating a language, and we actually made not just that mistake, but the other mistake, which is we created a platform that now had two separate languages that were custom. One for interactive querying and one for background processing.

Paul: Um and uh the language itself called TICKscript actually looks like nothing else that uh you've probably ever seen. It's very, very strange. Um but over the last... Was it three and a half, four years? Uh a surprising number of people have actually adopted it, and a surprising number of people have written very complex TICKscripts despite very serious gaps in the functionality they should provide as a language uh and some gaps in what I call like developer ergonomics, which is the experience of actually writing code in it, and developing and testing things.

Paul: Um but it has this like fan base that uses it, and they get a lot of value out of it, so I thought, "Well, there must be something there because if those people are willing to suffer the pain of using this thing that I can see all sorts of like horrible warts on, there's something worth you know putting more effort into." And when we went to create Flux, it's... You know we wanted something that could be used for background processing as well as interactive querying, and the, the choices then were... You know we knew we couldn't use SQL for the reasons I mentioned, so at this point, it's either, "Do we use an embedded language like Lua, or do we create our own?"

Paul: Now, Lua obviously like there are very mature implementations of it. It would have been way easier to just use that. My problem is like I don't think Lua has enough popularity and that people are familiar enough with it. I think the learning curve is too high for people to adopt Lua like it's, it's just not that easy for regular developers to use.

Paul: So the other thing we wanted was we wanted to be able to control the tooling around the language. We want to, ultimately, we want to create an experience that has a UI in it that allows you to create these Flux scripts without actually writing Flux code. Right? So point-and-click interfaces that describe like data flows of different time series data that you're collecting that output you know monitoring/alerting rules or all sorts of other things. Um so we want to be able to control the language.

Paul: And then, the next thing I did was I thought, "Okay. How, you know if we're going to do a new language, it has to be easy to use. It has to be easy to pick up," so we intentionally made it look like JavaScript, which... You know plenty of people hate on JavaScript, but the fact is it's probably the most widely used programming language in the world. Even people who don't write JavaScript day-to-day are usually pretty familiar with it.

Paul: You can look at the code and kind of understand what's going on, um and we said like... The truth is like the learning curve in this thing is going to be the API, and the API learning curve would be there regardless if we had written... you know if we had used Ruby as the starting point, Lua. All of those other things. Like the API is the biggest surface area. The surface area of the actual syntax of the language, that can be covered uh in 15 minutes by reading a getting-started guide that shows you the basic pieces of it.

Paul: So that's you know that's the bet we're making. Obviously, you know we just launched 2.0 as a cloud product. The open-source product is still... The open-source build is in alpha right now. We've just released a new alpha release, so it's really too early to tell what's going to happen, but the thing I... The joke I like to make is I'll either be spectacularly wrong or spectacularly right, but there probably won't be a middle ground.

Corey: No, it's fair. And I think we're going to see one way or another. It's, it's an interesting space, and we're seeing a lot of emergence coming out of it, which I guess gets me to um one issue that has been a recurring theme on this show, has been the idea that multi-cloud is generally not a terrific direction to go in. Pick a vendor and go all in.

Corey: The challenge is that you're already going to be locked in to whatever it is you choose, unless you're spending an awful lot of time working around that to no real benefit, but understand where that lock-in comes from. From that perspective, if something gets built on top of Influx, is that fundamentally locked in from a data model perspective? Is there lock-in being driven from a "once you start paying, you never ever, ever get to leave" and an Oracle-esque model? How does that story play out as far as adoption-implementation go?

Paul: Yeah, so, so we have the open-source InfluxDB, which is basically a single server. You can use that obviously. It's MIT-licensed uh with no restrictions, so you can do whatever you want with that. If you want to make your own new version of Influx and fork it, go for it. That's up to you. Um our cloud product is basically the exact same API, the exact same user interface. Um we don't yet have bulk export and bulk import of data, but our goal is to have seamless data transitions from open-source nodes at the edge and our cloud product uh running in whichever provider you choose. Right?

Paul: Right now, we're adjusting AWS, but soon, we'll be in GCP and Azure. Um and ultimately, like we don't want to hold your data hostage. The data model of InfluxDB is simple enough that you can represent it pretty easily. Like trivially, you can represent it in any relational database, and you can also represent it in Cassandra, or HBase, or whatever. Uh but basically, that open-source component is ideally the thing that gives you the feeling that you don't have lock-in, but I agree with you in the sense that once you've invested a certain amount of time and effort into a piece of infrastructure and particularly into a provider that's hosting your data for you, there's kind of, you know there's lock-in just by virtue of the fact that you don't want to spend the time and money to move off of it.

Paul: And the other thing about data is that it has gravity. It's not free to move from one place to another. So, and particularly in our use case, we're talking about large amounts of data, so that becomes that becomes a thing that you actually have to pay attention to. So you know ultimately, like people want to feel like there's no lock-in, but you know if it comes time to like, say, switch cloud providers, like are you really going to haul all feature development for six months while you do you know this lift and shift over to another cloud provider that provides zero customer value? Right? Like the main thing you need is the threat of moving to another cloud provider to give you you know a pricing package.

Corey: Yeah, and there are mixed reports as to how well that actually works, but it does raise an interesting question. Um one of the easiest jobs in the world has got to be running product strategy at AWS because you're just a post-it note that says, "Yes," on it there. There's really no thing that I would put past them building at this point in time.

Corey: And to that end, they have announced their own time series offering called Amazon Time Stream, which... It sounds almost like it can manipulate time itself, which it probably should because it was announced at re:Invent last year, and we're about to hit re:Invent this year, and it still hasn't been released. So it's like Influx without those whole pesky customers. So I don't, I don't know what the story there is, but more interesting to this conversation and germane to what you're doing and what you're building, what is it like when Amazon enters your market, when they come to crush you for lack of a better term?

Paul: Yeah, so it was... When that got announced last year, it wasn't actually entirely unexpected. Uh I just knew it was a matter of time, just when. Um you know it's obviously concerning. That's always the question is like, "What if so-and-so comes to build your product?" Uh I mean, the things I take comfort in are basically that Amazon isn't guaranteed to win in every single market they enter, and it's not necessarily, necessary that um you know it's a winner-take-all situation for every single product. Right?

Paul: So for example, Amazon entered Elasticsearch. Right? They have Elasticsearch hosting, and by all accounts, they make far more money uh at it than Elastic does, but last I checked, Elastic's market cap was pretty big and they're doing pretty well despite the fact that Amazon has come for them and is trying to crush them, right, and has forked their distribution, even though they don't call it a fork.

Paul: So I realized that there are things we can still do to try and deliver customer value that's outside of what Amazon does. Right? Like we're not going to be able to buy server time, and memory, and network bandwidth cheaper than they can, but hopefully, we can provide a developer experience that's better. We can provide a user experience that's better.

Paul: Um so you know when I think about Time Stream, which is their you know their soon-to-come time series database in the cloud, that's just one component of what Influx 2.0 does. A query in storage is just one piece. If you are going to try to cobble together what Influx 2 does by yourself, you'd have to take Kinesis paired with Lambda, paired with Time Stream, paired with S3, paired with uh some sort of visualization engine. I forget what Amazon's is, but they seem like...

Corey: Uh QuickSight, and that also has no customers because it's like Tableau, but crappy, and it turns out that's not the most compelling market to butcher.

Paul: Yeah, so that's the other thing that... I've heard you say and I've heard plenty of other people say, which is uh Amazon competes very fiercely on basically infrastructure, and cost, and scale, and stuff like that, but they, for one reason or another, just don't see fit to build user experience as they're compelling and UIs that are compelling. So that's one thing that we you know continue to invest heavily in is the UI, and the API, and how those things work together.

Corey: Yeah, and it seems that a number of the higher level differentiated services, the user experience lacks a certain polish. I think that's something that only comes with time for starters, but it also seems that it's not a high priority. And when you're dealing with a tool like this where you're going to be spending not inconsiderable amounts of time gazing into it, that experience should be reasonable and polished. And the idea that someone should be able to go from, "I've never heard of this thing before," to using it effectively should be measured in hours, not weeks, and I think that that's a lesson that sometimes gets lost.

Paul: Yeah. I mean, our goal is to measure it in minutes and hopefully seconds.

Corey: Exactly. Pictures are worth a thousand words as they say, and graphs on the other hand lets you figure out exactly how many words each picture is worth if you wind up getting your axis and calibration done appropriately.

Paul: Mm-hmm (affirmative).

Corey: So what's coming next if you have anything to share as far as what Influx is doing? What's interesting that people should keep an eye out for? What does the future hold?

Paul: Uh so we recently launched uh basically monitoring and alerting features inside our cloud 2.0 product, uh and that basically turns you know Influx 2 into a full monitoring/alerting platform in addition to a time series database and an agent that can collect data. Within Flux the language, uh what I'm most excited about is basically packages. Right? The ability for users of the system to create their own packages and share them with other people, and those packages could be bits of you know Flux source codes, so they're shared like you would on NPM, or RubyGems, or Crates.

Paul: Um and then, the other piece is you know packages that allows people to share essentially entire application experiences, which could be dashboards that you see. It could be drill-downs that you can do within your time series data, and all of that is kind of scope to the structure and schema of what data looks like inside of InfluxDB, um so that packaging thing is something I'm excited about.

Paul: And then, the last bit is... You know obviously, like InfluxDB, all the core components are open-source, and we really need to drive towards getting open-source InfluxDB 2.0 into beta. And for that, what we need is the basically the GA of 1.0 of Flux, the language, and we need the compatibility layers that users of InfluxDB 1.X can point to InfluxDB 2 and work with it as though it's a 1.X server, and we needed the migration tooling. And then, after we're in the beta, it's all about performance testing, robustness, and getting to the point where we can get open-source InfluxDB into to 2.0 and to general release.

Corey: Got it. Well, it sounds like there's going to be some interesting stuff coming up, and I'm very curious to see how that winds up manifesting in the marketplace and seeing it in increasing numbers of environments. I want to thank you for taking the time to speak with me today. If people want to learn more about Influx, about you, your sage thoughts on things that people should and absolutely should not do, where can they find you?

Paul: Uh so Influx, you can find at InfluxData or on Twitter as @InfluxDB, and I can be found on Twitter as @PaulDix.

Corey: Thanks so much for taking the time to speak with us today. I appreciate it. Paul Dix, founder of InfluxData, makers of InfluxDB. I'm Corey Quinn. This is Screaming in the Cloud.

Announcer: This has been this week's episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com or wherever fine snark is sold.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Paul Chin Jr.
Paul Chin Jr. is a curious human who likes to work with new technologies. His day job is at Cloudreach as a cloud solutions architect, working with enterprises to modernize their applications in the cloud. On the side, he’s a prophet for Nicolas Cage and is called to spread his message.

Links Referenced

  • Twitter Username: @paulchinjr
  • LinkedIn URL: https://www.linkedin.com/in/paulchinjr/
  • Personal site: https://www.paulchinjr.com
  • Company site: www.cloudreach.com
  • Talk hashtag: #praisecage

TranscriptAnnouncer: Hello, and welcome to Screaming in the Cloud, with your host, cloud economist, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on this state of the technical world and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Paul Chin Jr, who's a lot of things. He's a curious human who likes to work with new technologies. He's a cloud solutions architect at Cloudreach, helping companies modernize their applications in the cloud. But most importantly, and the reason we're having this conversation is that he is a prophet for Nicolas Cage and has been called upon to spread his holy message. Paul, welcome to the show.

Paul: Welcome. Thank you so much. I'm so, so happy to be here with all of your listeners, my brothers and sisters as I like to call them, in the church of Cage.

Corey: So tell us a little bit more about this. To be clear, we are speaking on Nicolas Cage, the somewhat washed up actor who could never turn down a role, correct?

Paul: That is correct. I wouldn't say he's washed up, but yes. I do fancy myself at a profit of Nicolas Cage. I started off uh where I had been doing a bunch of different startup things, not even remotely connected to technology. Um I've had restaurants and food trucks and T-shirt businesses and photography studios. And the whole time it was all converging into this singular point where I was being called in a different direction. Um and that point I knew I had to learn how to code, how to write it, um how software was actually made, instead of relying on somebody else to make it for me. Um and it's a pretty big, daunting task when you're just like, I don't really know how to code. I'm okay with computers. I've used spreadsheets a lot. Um the logic in it seemed to make sense. But as you're starting out, as I was starting out, I really needed something to guide me and I needed something to really focus on because you get lost in sort of the tutorial madness of building the same thing over and over again.

Paul: And so when I looked almost jokingly deep down inside me, I needed to find like a muse or a source material and it was it's the internet that really has given me the superpower of learning how to code. So I said to myself, how am I going to really give back to the internet? And the best and the only answer was to worship Nicolas Cage as a God of the internet. So all of my projects focused on creating things or exploring Nicolas Cage as the one true God.

Paul: So I started off building IOT projects in order to worship him, I started building react components to pull in and modify GIFs of him. I built all kinds of crazy things uh and they were all themed around Nicolas Cage. And as somebody starting off with no previous industry experience, it was the best way that I knew how to one, learn core concepts of programming really well, apply it in a way that kept me motivated and then gave me a really compelling reason to show other people, so that I can get myself in front of them to try to jump into a brand new industry.

Corey: Got you. So can you give me an example of one of the talks you've given? Because if you just say an isolation "Yeah, I give a whole bunch of talks about Nicolas Cage." You sound insane. So having a story around, give an example how this might manifest so our listeners can visualize this as they are frantically trying to find any other podcast to listen but this one.

Paul: Sure. So um one example being that this is a cloud podcast, uh cloud technologies are so new all the time, coming out with stuff, AWS just cranking all kinds of features and services out. And um whenever they come out, I have to play with them. One that came out at the time was Step Functions. And I had been using lambdas For a little bit and we knew that step functions was going to serve a very interesting spot in order to help orchestrate how the lambdas get run. And I thought to myself, well this is a perfect example of when you really need a to bring some extra control to a certain process. And I thought, what is a better process than stealing the Declaration of Independence? And so I started exploring this idea of breaking down the heist in terms of singular functions for each thing that Nicolas Cage has to do as he goes through the process of stealing uh the Declaration of Independence.

Paul: In my step function, I have all the different states that he goes through in the movie. I compared each you know function to what he actually does. And then I gave a talk about step functions and where I teach people about the step functions, how they're supposed to work, how you move pass state around. But the thing that we're actually doing is calling to different parts of the movie, uh just in code. I thought that was a really a cool and fun way to show off one thing while still talking about our one true God, Nicolas Cage.

Corey: Excellent. So I tend to have a certain affinity for the ridiculous when it comes to conference talks, from using illustrative points to tell a story. A while back I started talking and evangelizing I suppose about my favorite database, Route 53. And that was hilarious in so only as far as that no one could quite tell if I was serious or not. Spoiler, I'm completely serious. And someone built an entire system in Ruby, on top of Route 53, so you could query it, you could update records, create tables, etc.

Corey: And someone on Reddit wound up posting later in time that they weren't sure if I was serious or not and this actually sounded like a good idea, should they do it? At which point I started to realize, wow, people are actually listening to the nonsense that I say. That becomes something of a problem. I have spoken with some of the people on the Route 53 team who are in equal parts amused and horrified and it it turns out that it's not the worst idea in the world for some applications. So my question for you becomes, does at some point when building these talks, does the actually stealing the declaration of independence begin to seem like a good idea?

Paul: All the time. Uh I've done another talk where it was all about ETL pipelines and I compared it to compared it to Gone in 60 Seconds when they have to steal all the cars. I map each of the AWS services to a different character in that movie and their role. And it's all about you know doing the pipelines and you know moving data from one warehouse to another, which is essentially just moving the cars from one warehouse to another.

Paul: And yeah, after I give these talks, I really do feel like I have this extra super power to go achieve these outlandish things like stealing a hundred cars in a night or taking the Declaration of Independence. Or it just keeps going on forever. There's 103 different movies that Nicolas Cage has done uh and I only have six talks so far, so that's still at least another hundred talks I could probably give on these movies.

Corey: He is closing in on the number of AWS services.

Paul: Haha. I should do that. Yeah, that's definitely a good talk idea. All the different services as Nicolas Cage movies.

Corey: There are so many opportunities in there too. You give a talk on systems manager, Nicolas Cage's manager.

Paul: Haha. So people do come up after as the talks um to me and really questioned whether or not I believe in Nicolas Cage as a God. Um and I definitely feel like the performances that he puts out belong in the realm of this like higher level thing. He does so many different things that there is room in there for normal people to pull interpretation from it. That is art but that is also you know maybe part religion, it's up to you however you feel about it. But when people ask me about it, I've definitely seen ... uh I haven't seen all of them. It's very difficult to see them all. But I've seen many, many of them.

Corey: No, I'm not wanting to mock other people's beliefs. Some belief there's a UFO trailing behind a comet and that ritual group suicide is the best way to catch it. Other people believe Nicolas Cage is a good actor and who am I to crap on anyone's religious beliefs regardless of how outlandish they may be.

Paul: Exactly. Haha.

Corey: So on a slightly more serious note, you don't have a background yourself as a developer, so you wound up picking up learning how to code as an outgrowth of your job and you went in a serverless direction with it. Can you talk to us a little bit about what that journey looked like?

Paul: Sure. Um yeah, like I said before, my original upbringing was a lot of um small business. I grew up in a restaurant, my parents had a Chinese restaurant here in Norfolk, Virginia. And I learned at a very early age about business and making sure that you know you have people coming in the door and you do your operations and all that. So I got a very long crash course in how to run businesses. And then when I decided to learn how to code, I knew that software was going to be the way that I can scale any business, whether that's a restaurant or a T-shirt company, like I have to have some level of understanding with software development.

Paul: And I came into serverless, it was very much a natural thing. Right like when I got started, I use a lot of different cloud services because like I didn't know how to start an actual server. I didn't have a computer to um build a local machine on. I never learned any of the networking stuff at first. It was all very much, uh what is the tool that's going to get me an application the fastest?

Paul: And I started learning in 2015. And at that time, serverless really wasn't like a big name thing yet. It was almost about two years later, uh all the blog articles would come out about it and it was everywhere. But when I first started off, you know I used a lot of cloud native tooling. I didn't know that's what it was called at the time. I just knew that there was a free tier. I could use this to get up and running and I had an application. Right like I was able to host files without ever configuring anything. Um I was able to make APIs without uh worrying about Linux or installing it or running a shell command. Like I never had to deal with any of that.

Paul: That gave me a both a good advantage and disadvantage. Um when I look at customer situations now, I have to be empathetic and mindful of previous technologies, you know what they call "legacy systems". They did the best that they could at the time, but I had to go back and relearn all these challenges that they faced so that they could build the solutions that they needed. Now that you know I come in and try to help them modernize that stack, um it's both looking at it with fresh, fresh eyes, as the possibilities never end in cloud technologies, but also knowing how much they had to pull and push to make that application work you know 15 years ago.

Corey: It's interesting to hear you say this, where you don't come from a background of operation central focus. For example, I was a grumpy CIS admin for years, so running the Linux box was always the easiest part of everything else involved. Writing code that worked was a whole separate story and that was something of a challenge.

Corey: The part of the story that resonates though is the idea of having a larger goal that isn't getting the baseline stuff up and running and just being able to move past that directly into doing the thing that you actually want to do/need to do for whatever the outcome you're chasing is. So what's fascinating to me about the whole serverless ecosystem is that you have people coming from such a wide variety of different places when they start, and they wind up all gathering around over a shared goal of getting something done without uh I guess the ideological purity of having to spend four years getting a CS degree first.

Paul: Yeah, totally. Um I'm constantly trying to back myself into a CS degree. Every time I dip into the theoretical waters of computational sciences, uh I ended up questioning reality again, like is this real? Is this bit real, where does it exist? Uh I fall down this hole that doesn't ... it's very cool. Like it's very, very cool that we can learn how the code gets executed on a machine. And for some workloads, it's very essential that you know how performant your stuff is. But in business, 90% of the time uh it's not going to affect most systems I would say. Um what's really affecting the bottom line is how quickly you can get this feature out uh and get it validated in the marketplace.

Paul: And so I thankfully got to leave behind a lot of the grumpy CIS admin stuff uh and just focus purely on writing this code that's going to let me do uh what I need to. I truly believe that making this technology more accessible to more people so that other folks like me who grew up not like never thinking that they could make a computer program, to make to to really empower other folks to say, "You know what? I can do this. Um I can build a computer program. I can make the computer do what I want and use it as a true tool."

Paul: Um it's something that I'm really passionate about. I really look out for tools that have little to no configuration. I look for tools that are going to have very clean interfaces. Whether that's programmatically with an API or you know with a browser based GUI. Um I believe in this so, so, so much. I I help volunteer for local Great Computer Challenge that's been running for a really long time and we actually give kids, uh like first and second graders, iPads with Scratch on it. They create these amazing stories and amazing interactive art pieces using nothing but drag and drop tools. And they understand the logic that goes into it, that's freaking amazing. I do the same thing with my little four year old daughter and anybody I meet on the street really, I tell them you know you could learn how to control a computer. I don't tell them code because that may or may not be too scary for them. But I tell them, you can learn how to control a computer that is not outside of the possible.

Corey: What also interests me is that you mentioned you grew up with your parents owning a restaurant and being able to, I guess see the logistic side of that, of being steeped in a business that historically has always had to focus on things like making payroll and being able to handle like this silly, outmoded concept that here in tech we don't care about anymore, but legacy businesses do, known as profit, and being able to make sure that you can stay afoot every month. It it feels like it gives you a grounding in reality that you don't always have when you're coming at this from a more theoretical point of view. Is that an accurate assessment?

Paul: Definitely. Um so the big leaping block from doing like I guess more main street brick and mortar businesses into understanding the importance of software, was this small window of time here in Norfolk where like the startup scene was really buzzing and everyone wanted to be a part of an incubator or have that you know million dollar Facebook idea and everyone's got an app idea. That whole ecosystem started cropping up around that time and I knew that um the business people who were able to understand the technology we're going to you know make themselves be able to perform better.

Paul: And I saw way too many people who think that it's the technology that drives the business. Um and I saw a lot of folks who are very, very talented engineers, fully capable of creating anything their minds imagined but it doesn't necessarily mean that it's going to work well in the marketplace or that you know that idea was going to be able to scale to be able to support themselves, their family, employees and all that stuff.

Paul: So depending on what a business' goal is, I've actually, side note, I think some of these businesses that come out and list as technology unicorns, I don't think their actual goal is to make money. I'm not sure what their goal is. Uh I'm not really playing at that level yet, but you know my essential upbringing is how many egg rolls can I sell? What did it cost me to to put into that egg roll and how many can I move an hour? And then I know how much money I can make, every day that's what I think about. What is the situation? What are my inputs? What is my output? Then I have what I have at the end of the day.

Corey: How do you think that this is going to shape what, for example, you alluded to teaching kids how these things work and having a daughter yourself, who's a couple of years older than my daughter. What world do you think they're going to grow up in? How is this going to shape what education looks like?

Corey: I mean from my perspective, when I learned this stuff in school, uh in seventh grade I had a typing class and I was always getting poor grades in that for two reasons. One, I skipped ahead and got the entire assignment finished in the first five minutes. And secondly, my typing form wasn't perfect because for me, at least at that era, I had a mental map of the keyboard, so I didn't hit anything from the appropriate fingers or whatnot. I just put my hands on the keyboard and then words came out correctly and that was the end of it. So it's it's weird that typing is the least interesting part of any of all of this, but that was as far as computer science education went back when I was in school. I don't think that's the case now, but I'm curious as to what it's going to look like, especially with the advance of things like serverless technologies in the education space. Thoughts?

Paul: Yeah, uh man, Stem is such a buzz word, just like startup. There's tons of initiatives and very passionate people about making Stem happen. Um and I feel like we have to come back just a little bit more fundamental of regular problem solving. And in education, I really want to see these serverless and cloud technologies enable people to control a computer, build a program without ever realizing that they're doing it. To them it's just like making a game or creating some outcome that they want to have happen.

Paul: And you know I show my kids, they're very little, they're one and four, how to use an Alexa. And they can communicate with it. They can utter things, add it. I joke that my second daughter, her first word was going to be Alexa you know because we use it to do everything. All of that is powered by, obviously AWS and serverless technologies. And now on my iPhone I can build, drag and drop integrations with the Alexa without ever having to write any code. I had built Alexa skills, you know writing JavaScript. But now I can also do it with my fingers. So for education, I think that we have to take a different look at what it means to be using technology as a tool and what it means to look at computer science as a discipline, as a different craft that underlies the implementation of the tool.

Corey: That really resonates. It hadn't occurred to me that the idea of voice first was going to be an issue. But you're right. When my daughter was a couple months old or damn near it seems, she wasn't verbal yet, but whenever we spoke to Alexa, there would be an immediate, she would look exactly at the speaker and wait for it to respond, which was, oh, she's smart. It didn't occur to us this was a real first interaction with technology.

Corey: I think the world that she's going to grow up in is going to be radically different to the one that we're living in today, just from how pervasive technology can be. And that is a double edged sword. I'm not one of those stars in my eyes idealists who's convinced that this is nothing but a net positive. I think that there are serious questions about that, but the fact that it's more accessible, that it's no longer going to be a bunch of ivory tower types who are the keepers of the flame when it comes to working with technology is going to be transformative and I do believe that serverless is a step along the way towards whatever comes next. Now what that is, I don't know, I'm not a futurist. I don't consider myself a digital prophet, but there is clearly something brewing. I just don't know what it is yet.

Paul: It's funny, you mention your daughter understanding that she has to look at the device. Right my daughter would also just sort of look at it. She knows that it's a focus point for control. And I am not one of those parents that really regulate screen time very hard. What I regulate is their intention on it. Um I tell my daughter like, "You can be on this device as long as you want, as long as you're creating, as long as you're trying to solve a problem, you can be there. Um but if you're going to be passively watching a cartoon or something, then yeah, there's going to be some time to take a break, exercise your mind a little bit, let it wander on its own."

Paul: And um I think that that's a really powerful thing to think about as we continue to integrate with technology, in our lives. For so much of it, it is about consumer technology is about making things easier for yourself, about turning your mind off. Well, now I truly believe that the tools, the serverless tools um and the new abstractions that other people are building um can give us the ability to really flex our minds now with technology, not just consume.

Corey: I think that's an excellent point and one that people tend to skip over far too frequently. Now, before we go, I do have one more question for you. Uh out of however many movies he's been in now, I think you said 103.

Paul: Yeah.

Corey: Which is your favorite Nicolas Cage movie?

Paul: Oh, everybody always asked this and I feel like I should be better prepared for it, but every single time I try to sit down and think of of the best answer, it's so difficult because it's like if someone asks you what's your favorite anything. You have different answers for the different moods. So I'm going to take a cheat out and say that the movie that I like watching over and over again, is probably Raising Arizona. Uh I feel like that's just a great film, regardless of uh Nicolas Cage being in it or not. Him being in it is even better.

Paul: Uh I think that he has of course an amazing track record of being that quintessential 90s action star, so I really love Con Air. I'm trying to work it into another talk. I am not quite sure exactly how I'm going to do it. I kind of want to do the new AWS event, um event [inaudible 00:26:31] with Con Air somehow, not quite sure about that yet. Um but yeah, between Raising Arizona and Conair, I just think those are timeless movies. Then he's in some really great brand new ones like Mandy and his voice acting in Into the Spider Verse. It was it was so great.

Corey: He's a very versatile actor.

Paul: Yes.

Corey: Can't take him to an Italian restaurant though because he gains 80 pounds cause he can't turn down a role. But other than that ...

Paul: Haha. I'm going to make a believer out of you yet. I can still hear a little bit of skepticism, but I believe that you will come around to it.

Corey: So where can people find more of you, your antics, etc, if they wished to learn more, both about your philosophy to serverless, your journey you're on and of course our one true prophet, Nicolas Cage?

Paul: Yes, uh they can definitely follow me on Twitter. That's where I'm the most active. Paul Chin Jr, all spelled out. J-R at the end, P-A-U-L C-H-I-N J-R, Twitter. Um I use an awesome hashtag called #praisecage when I give my talks and different concepts that I'm working on. They can also check out my GitHub. The GitHub handle is P-C-H-I-N-J-R, so P Chin Jr. I release a lot of the example code that I have out on GitHub.

Paul: Uh and what's fun on Twitter is when you use that hashtag I can can see when other developers are also working on Nicolas Cage projects. In fact, I have one pinned uh where another fellow node botanist uh created a Nicolas Cage face that follows you around the room like on a servo, with face detection technology. So definitely some Face/Off nods there, but yeah, we'd love to have more followers and more folks coming into the Church of Cage.

Corey: Excellent. It it sounds like something that's well worth having a I guess a following built around.

Paul: Yes.

Corey: Paul, thank you so much for spending the time to speak with me today. I appreciate your being so generous with your time.

Paul: No problem. This was a ton of fun and I want to give one last shout out to all my uh mentors and community members here in in the Norfolk, Virginia area, Linda Nichols and Travis Webb and Kevin, and all the folks uh who have helped me and welcomed me into the technology space, so thanks.

Corey: And of course, oh wait, you can never thank Nicolas Cage enough either.

Paul: Also our one true God, Nicolas Cage. Praise be to him. Look up the three cats of Cage, so that you knew how to fight our devil, John Travolta. Thank you.

Corey: Paul Chin Jr, evangelist for Nicolas Cage and all around decent person. I'm Corey Quinn. This is Screaming in the Cloud. If you've enjoyed this episode, please leave us five stars on iTunes. If you've hated this episode, please leave us five stars on iTunes.

Announcer: This has been this week's episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Paul Johnston
Paul Johnston is an interim CTO, CTO and strategist who has particular interests in serverless, cloud, startups and climate change. Formerly, Paul served as a Senior Developer Advocate at AWS for Serverless and CTO of multiple startups, including one of the world’s first serverless startups. Paul’s also a keynote speaker, tweets a lot at @PaulDJohnston, and blogs a lot on Medium. Right now, he may also be working in stealth mode on something (it’s probably serverless)…

Links Referenced

  • Twitter Username: PaulDJohnston
  • LinkedIn URL: https://www.linkedin.com/in/padajo/
  • Personal site: https://medium.com/@PaulDJohnston
  • Company site: http://roundaboutlabs.com/
  • Sponsor: Manifold

Transcript

Speaker 1: Hello, and welcome to Screaming in the Cloud with your host, cloud economists, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on this state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey Quinn: This episode of Screaming in the Cloud has been sponsored by Manifold, Manifold powers marketplace infrastructure that connects millions of developers to the best APIs, tools, and services in the fastest growing communities, and also Kubernetes. They offer a complete toolkit that allows you to deliver your API first product to millions of developers. Check them out at manifold.co. Again, that's manifold.co. Thank you for sponsoring this ridiculous podcast.

Corey Quinn: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Paul Johnston. Paul, welcome to the show.

Paul Johnston: Hello. Hi. Very nice to be here.

Corey Quinn: So once upon a time you were an interim CTO, CTO and strategist, and you were focused very heavily on Serverless. In fact, when I first met you, you were a senior developer advocates at AWS for Serverless.

Paul Johnston: Yep, that's correct. It was an exciting time at AWS, and there were only two of us at the time, myself and Chris Munns. Yeah, it was an awful lot to do, and Lambda was a little bit younger back then. There were an awful lot of companies to get around, an awful lot of organizations to talk to, and I was based in [inaudible 00:02:07] Chris was based in North America. So yeah, it was a fascinating time to be around AWS back then.

Corey Quinn: It certainly seems so. Lately though, it seems that you haven't been speaking nearly as much about Serverless as you have about climate change. And what makes you interesting is that I can find any number of random yahoos to come on this show and talk about Serverless in a variety of different tones of voice. But no one in our space seems to be talking in any meaningful way about climate change except for you. And based on conversations that we've had, I know you to be a thoughtful, sincere person and you're not a ridiculous crank. So let's talk this half hour about climate change and how it impacts cloud computing.

Paul Johnston: Well, thank you very much for the eh saying I'm not a crank. I hope I'm not so. It'd be a great thing to talk about, I think.

Corey Quinn: About climate change or being a crank?

Paul Johnston: Haha. About climate change. Let's talk about that. Hopefully, I'll come out a lot better at the end of it, and not a crank.

Corey Quinn: So you've given a presentation on this. You've written on Medium, which is fussily named because it is neither rare nor well done, and you wrote a white paper on energy usage in data centers contrasted between the large providers. How are we doing from a climate change point of view in the cloud computing industry?

Paul Johnston: So the cloud computing industry, it's not doing too brilliantly, although there are some real bright spots and some real low spots. The industry itself is ... If you look at the whole data center industry, and there are different opinions on this and it's difficult to look at, but the whole industry appears to be actually a pretty bad pollution in terms of greenhouse gas emissions. So some will say it's below 1%, others will say it's 3%, 4%, 5%. I take a roughly middle view of about 2% of greenhouse gas emissions are worldwide. This is our down to data centers.

Paul Johnston: Now, if you think about the fact that actually we're all streaming Netflix and we've got all of these things stored in the cloud and we've got all of this storage and we're using cloud providers, there's 2% of ... one 50th of the greenhouse gas emissions that are going on in the world are basically down to us using computers as techies. Then you look at where the growth is in the market, and the growth is in cloud computing sector. It's, yes, we've still got all of these corporate data centers but they are all beginning to move, and over time we'll move and shift to cloud computing.

Paul Johnston: So the the fact is that we've got to look at cloud computing and we've got to look at cloud computing as, is it a good example of ... is it good at looking after resources and looking at whether or not it's emitting well or badly, or what what's it doing? Is it offsetting? Is it doing all of these things? So myself and Curry created a white paper last year. We did a little bit of research and we took the top six cloud providers, and we looked at them and we basically said, "Who's doing well? Who's doing badly? What are they doing, and how are they doing it?"

Paul Johnston: We effectively came up with ... and just take the top three because it's the easiest. The top three of AWS, Microsoft Azure, and Google. Google are the exemplar. They are doing really well. If you have a look at their sustainability page on Google Cloud, they try to offset all of their usage with renewable energy, and they do a really good job of it. It's not perfect. So if you see their China usage, you will find that they are struggling to find renewable resources to offset, but they do as much as they possibly can.

Paul Johnston: It offset in other areas of the world to do that. But they do a very good job and they offset a 100% of their usage as much as possibly can with local renewable resources, but otherwise by renewable energy credits. Microsoft are, I think if I remember rightly, over 50% now. I think over 60% possibly, and then buying renewable energy credits for the rest. But Amazon's AWS is the one that we're all concerned by, and at the moment, they are stating 50% and they're not actually saying what what they are now. They said that in the beginning of 2018, and they've grown since then. We don't know what their electricity usage is now. They don't release numbers.

Paul Johnston: So we're not getting updated information and we're reaching the point where we're constantly asking, "Let's have some updated information, let's have some updated information," and they're not giving it to us. But they have four regions which are Oregon, Montreal, Ireland, and Frankfurt, and those regions are a 100% sustainable. So if you can put workloads that are non production or not needed all the time, then put them in there and you're 100% sustainable, you can be green and still use the cloud and still use AWS. But otherwise, us-east-1 is basically, just don't use it if you can at all avoid it, for environmental reasons if nothing else.

Corey Quinn: There have been many reasons not to use that particular region, but this is the first time I've heard this one.

Paul Johnston: They are very, very many, and it's it’s basically a very simple thing to consider. We've looked at things like machine learning workloads. You could move those out. You don't need to run those in us-east-1. You can move them anywhere. Just all of these kinds of things. We just automatically use us-east-1. Well, I used to work for AWS and there are a lot of people there that I really ... I have a lot of fondness for, and I remember, a lot of my times there.

Paul Johnston: Well, you know it's still a big company. It's got a lot of money. It could buy some renewable energy credits to at least offset to 100%, and get there and then start investing in the renewable infrastructure. It could do these things, and actually make a big difference in the world, but it appears to not be doing that, and we're not entirely certain why.

Corey Quinn: So other than I guess people who can think further ahead than next quarter, there are a number of people who are clamoring about climate change across the spectrum. Why don't we see that in tech? Or if we do, we have probably people in tech who are sounding the alarm on this, but they're not talking about cloud computing. Why?

Paul Johnston: I'm not sure. I think we've got a bit of a problem within technology for a number of reasons. But I think we see a technology as not really a part of the problem. So one of the things that I've found, looking into this, is that technologists don't really see technology as a whole as creating a carbon footprint. One of the CTOs that I've talked to recently said to me, "I don't know why you're talking about all of this stuff. My company doesn't have a very big carbon footprint," and this was a company that has a huge data center.

Paul Johnston: I was looking at him, and just trying to understand how he didn't see the correlation between the amount of electricity he uses in his data center and the power stations that create that electricity, and it's because in our world ... There are a lot of conversations that go along with this and I'm not trying to agree with any political side. But in our world, because we have created this globalized world, we've actually abstracted away the carbon and abstracted away the emissions so far away from the turning on of the light or the turning on of the computer and the use of our electricity to the point where we don't see the impacts.

Paul Johnston: That's especially true for the culture that we live in. So when you talk about the climate impacts, there's a lot of ... I mean this is true for technology, but it's true for the culture that we live in in terms of the West. So when you talk to climate activists, what they all talk about is the Global South and the Global North, because actually the impacts of climate change is being felt in the Global South and they are being felt now, just because of the 1.1 degrees we've already had of warming. Whereas, the Global North doesn't feel those changes quite so much. Even though we've had a few hurricanes and they've been really bad, we don't feel them, especially in the UK and North America.

Paul Johnston: We haven't felt them anywhere near as badly as somewhere like India or Africa has felt them in terms of droughts or in terms of extreme weather events. So we've got this, we’ve got this because it isn't my problem, it hasn't affected me, and because I don't see the problems that are coming down the road, and everything's abstracted away, and all these things don't really matter. I think we have a problem in tech in terms of taking responsibility, and in terms of how do you say I am responsible for something and taking that responsibility forward. I think we've really got this problem with convenience and responsibility there.

Corey Quinn: Do you think that this hits on the entire idea ... I guess, we'll go back to Serverless for a minute, of sure there are servers there, but I don't have to think about them or care about them. If we take a look at cloud computing, there are things that I obviously have to care about. We can pervert the shared responsibility model, as an example here. I don't have to care, for example, about what hard drive vendor AWS uses. That is so far below the level of things I need to care about.

Corey Quinn: It feels deceptively compelling to be able to say, "I don't care where they buy their power, because I shouldn't need to." That’s it feels like going down that path is the only realistic stance someone can take because they're not empowered to directly change it by most reasonable perspectives that I've heard, and they're trying to solve business problems that don't directly align. This has been something that gets a surprising amount of pass, and I don't think it's accurate, but I suspect that might be where that attitude comes from.

Paul Johnston: Yeah, and I think I would agree. I think you know since the cloud computing ... the convenience of it, the convenience of being able to click a button or even just send an API call, and you have this huge amount of computing power at your fingertips is a world away from having to provision or even buy servers, and then provision them and set them up and make sure you've bought your electricity, and then you set up things called power purchase agreements.

Paul Johnston: You know have to buy power from a power company. You know that’s the kind it's so far removed from all of that. You abstract all that complexity and you have this convenient service that gives you all this compute power. It's so far removed. Then I can completely understand how a technologist would feel disempowered in that scenario to say to someone like AWS, "Oh yeah, you need to change your power usage. You need to change the way you do power. You need to change the way you think about your electricity."

Paul Johnston: And one of the CTOs I spoke to about this recently said, "Well, why don't we just get everyone to move to Google?" I said, "Well, how's that going to work?" One company moving to Google is not going to make one you know... an AWS, or you know if we're talking about Oracle, for example. Let's not just demonize AWS here. I'm not trying to do that. I'm just using them very much as an example. Oracle's only at 33%, for example, in terms of renewable energy usage.

Paul Johnston: You can't just say, "Well, one company moving to Google," which is the gold standard here. You can't just say, "Let's move everyone to there," because one company moving will make no difference. We've got to have awareness across the organizations who are using cloud, and then get them to combine and get the industry to understand itself. This is what activism is. It’s not activism in this scenario is not about getting everyone to vote with their feet and move away.

Paul Johnston: It's about actually getting the industry to realize its responsibilities, and start to have the conversation and create the conversation. And I think the industry should be having the conversation that says, "Right. I think we need to do something about this. I think our responsibility is to everybody in the world," and this this industry is huge. It's got a huge amount of money involved.

Paul Johnston: And actually we should care about the fact that our cloud provider is is providing all this convenience but it’s not convenience without a cost. And that cost shouldn’t be at the expense of our childrens’ future And it shouldn’t be at the expense of you know cities in India in 2050 being unlivable. I mean essentially that’s what we’re talking about is partly and while I know that there are people going well that’s not me using cloud computing is not going to do them. I understand that one person doing one thing is not going to change it. This is about changing perceptions across the industry.

Corey Quinn: The other challenge is that people have a sense of immediacy, where I wind up talking to people who are big in the cryptocurrency space and they have the temerity to yell at people who travel too much of, "Well, you understand that the jetliners have a massive environmental impact." Yet, what do you think mining your Dunning Kruger is causing to the environment? It's one of those incredibly wasteful things, it almost feels like it's too surreal to exist, but it does.

Paul Johnston: Yeah. And I think the problem is that ... I mean, you look at the Bitcoin specifically. Bitcoin and Ethereum have a proof of work, and proof of work is essentially a massive lottery that everyone takes part in. And if you take part in the lottery, you use compute power to try and guess a number to win some Bitcoin or Ethereum. That's essentially what it is, and everyone joins in this lottery and they win some money at the end of the day, and then 10 minutes later they get to do it again.

Paul Johnston: Huge amount of electricity to win some Bitcoin, and the the research is showing ... I think I did it recently. The research showed something like driving a car 800 miles is the same as one Bitcoin transaction, and that in terms of global emissions. So when you start talking about it in those terms, you start to really get scared about, is this really worth it? And I think we have a real problem in terms of when we start to point fingers with somebody's individual behavior. An individual's behavior is actually really not that important in the grand scheme of things, but when we gain together, when we do something together as an industry, we will make a difference.

Paul Johnston: Things like basically saying, "Proof of work is a terrible thing to do. We should not be doing that," and saying, "Yes, our cloud computing should be 100% offset, and we should be investing in data centers that are completely renewably powered and we should be investing in battery technologies, proper battery technologies. We should be creating data centers such that we are not using on demand compute, not creating servers that are over-provisioned and under utilized. We should be creating technologies like this as an industry stating that our aim is not to waste the electricity that we're given."

Paul Johnston: But actually we don't do that. We just create technology, and the way that we create technology because effectively it's free, and there's an individual's choice. I have cut down on the amount of travel. I certainly don't fly in Europe, anywhere near as much as I used to. I try to use the train in Europe. Transatlantic travel is very, very difficult as Greta Thunberg has shown. Um and so you usually just have to use a plane for that point, so you've learned to offset.

Paul Johnston: So you learn to use offsetting various offsetting companies, and there's a gold standard for offsetting which you can use. So I think when the when the finger is pointed, I think you just have to turn around and go, "Well, what else are you doing? I'm doing all of these other things. What else are you doing?" And yes, no one person's life is ever going to be completely free of all of these problems, because our countries and our systems of government and our ways of life, our globalization is so full of sunk carbon, for want of a better phrase, that we can't get away from.

Paul Johnston: It's virtually impossible to get away from having a carbon footprint that isn't sustainable for the long term. So unless we change our ways of looking at life and the universe and everything, then I don't think we're going to do anything. There was a talk yesterday that I went to, where someone said, "Effectively, unless governments and a few very big companies change, then the only other way of doing this is to eh is to get the entire world to go to put solar panels on their roofs and to stop driving cars and to stop flying and to effectively all go vegan."

Paul Johnston: That's pretty much what we've got to do, and we've got to do in the next five to 10 years, and I'm pretty sure that's not going to happen. So we've got this problem where we need to fix it, and we don't have the tools to do it. So when someone points the finger, I think we just have to say, "Well, do something but don't do nothing." I think that's the best way of looking at it.

Corey Quinn: Migrating between regions, let alone between different cloud computing providers, is a massive undertaking for almost any-

Paul Johnston: Absolutely.

Corey Quinn: ... environment that is more than trivial to scale. It almost feels like that is far less likely to happen than enough people shaming Amazon into stepping up on the climate front.

Paul Johnston: Yeah. I completely agree with that. I've seen you know I’ve seen data centers that are basically stuck in buildings that are never going to get moved, or at least ... because they'd been there 20 years, and the person who's there is basically ... there's only one person in the company that really knows how it works, and the cost of moving it would be an absolute nightmare but they're in production.

Paul Johnston: I can imagine there are cloud environments now that have been sitting on ... EC2 instances, or moved between EC2 instances and various different things that have been there, what? 10 years now? The people can't move because they don't really understand how they work. And I can imagine all of these scenarios for the non trivial things that we're talking about. We can't just go, "Oh, I'll just move. I'll just do this." It's too simplistic. We've got complex technology, and I think building a movement around this, I think is the best way of going forwards.

Corey Quinn: So what can people listening do?

Paul Johnston: So we’ve got a number of different things, but the ... I think probably the easiest thing is to go and learn, go and have a look at what's out there, go and read some blog posts, go and follow some climate scientists. I think before you do it, and I think this is especially true in America, I think leave your preconceptions at the door. I think we have the same in the UK. I think it's all around the world. If you go and read the IPCC report from last year, which was the special report on global warming to 1.5 degrees and above, I think you'll find that that was signed off by every nation of the world.

Paul Johnston: It's not a political report. It's a report that basically gives you the information. Go and read that, learn something, don't listen to the politicians who are using words that are mired in alternative meanings. Go and figure out what those words are basically hiding, and then go and do the reading for yourself. I've got a Medium post on that I wrote a few days ago actually, which was about stop being a techie and start being a human. I think the tech world has a tendency to hear something like this, realize that there's a problem and then go, "Right. How do I fix this? What's the solution? How can I build a website, an app? And then that will be the thing I need to go and do."

Paul Johnston: And I would just caution against that simply because the problem is so huge and complex that that's not really going to solve anything. You might solve one problem but then there are going to be 100 other problems out there. So I would suggest probably going and sitting and listening to whoever your nearest climate activism groups are. So your 350.orgs, or your ... whoever they are within your group. So with Greenpeace, or Extinction Rebellion, or I guess the Sunrise Movement, and all of these organizations, whoever it is.

Paul Johnston: Just listen. You don't necessarily have to agree with them. Just listen. There are huge number of amazing climate people. There's a climate Twitter that is just full of amazing content. Just go and start to absorb this content, read it, take it on board, and you will start to find that there is an undercurrent of hope that we can do something incredible, but it's an organizational start, at present. We're not at the point of really doing anything. We are gathering, and I think we're in a gathering phase at the moment. But I would just caution against too much doing and not enough thinking and waiting and listening.

Paul Johnston: I know that for techie people that's quite a difficult thing to hear, because we all want to just build something and try and 10-X it and get some funding. But I don't think that's what the world needs right now. I think the world needs people to start spreading messages and talking. I think that's a very different thing to the way that we're used to do in the world now.

Corey Quinn: This is usually where we try to have something uplifting, and oh, here's how it's all going to be okay, and I feel like I'm grasping a little bit to find that narrative.

Paul Johnston: I know. And okay. So let's put it this way. The tech industry is full of people with money, with huge brains, and with a huge amount of love for the world that we live in. We have the opportunity to use those skills and the skill sets of running, building businesses, of taking out ... of building organizations and communication and technology, and taking those skills and putting them to use in the right ways, and in the right places.

Paul Johnston: I think there is the scope for turning is often the industry that people turn around and laugh at and see an awful lot of bad things. The disaster capitalists that effectively that ... A lot of people look at tech and see an awful lot of money and not a lot of goodwill, and I think we've got an opportunity to say actually there's an awful lot of good people in tech. There's an awful lot of good things that could be done. I know that it feels like there's not a lot of hope, but there is an awful lot of good stuff out there.

Paul Johnston: For example, if you've never come across Project Drawdown, so one of the things that we need to do to essentially turn back the clock on climate change is we need to stop emitting carbon. But we don't just need to stop emitting carbon, we need to also be taking carbon out of the atmosphere, so carbon capture and storage. We need to be reducing the amount of carbon that's coming out of various different other places around the world.

Paul Johnston: For example, someone did a plan for doing this and it's called Project Drawdown, and it's got a whole list of things that would help in the climate change movement. Most people would think, "Well, it's basically decarbonizing the grid so it's clean and making everyone vegetarian and vegan," and actually it isn't. There are things on there like educating women in Africa and Asia, and there's things in there around better chemicals in air conditioning units, and there's things around ... It's a completely different set of skills and technologies and understandings.

Paul Johnston: It's also the wind farms, and it's also the batteries, and it's also the things that are obvious. But there are things in there that are tech mind and a set of tech companies could pick up and go, "Actually, I can do something around that. I could build something. I could create something in that space." It's not going to be your venture capital funded, you know 100 billion unicorn company necessarily. But it might be one of the bricks in the wall that helps us do something about climate change in the world.

Paul Johnston: You may get to have a more fulfilling life than helping one of the big companies make you know another 100 million of people playing a game on an app, on a mobile phone that they bought for $1,000 in an Apple Store. We can do better with this industry. I have a huge amount of hope that this industry, it's attracted all the clever people. Maybe all the clever people can go and do something to fix the world we're in now. I have hope.

Corey Quinn: I guess that is the uplifting narrative that we look for in this. I'll include links in the show notes of the article that you wound up writing. The white paper as well. Where can people go to learn more about you?

Paul Johnston: So I’m on Twitter, @PaulDJohnston, and on Medium at Paul D. Johnston. Yeah, just come and find me there, ping me a message if you want to know more. There is also the ClimateAction.tech Slack group, which is well worth getting involved in. There's an awful lot of tech people on there who are trying to figure out how to use their tech skills for this kind of thing, for climate change and for good. So I’ll be on there as well. Come and find me.

Corey Quinn: Excellent. Of course, you are one of the organizers/founders of ServerlessDays.

Paul Johnston: Yeah, I am. If you want to get involved in doing the Serverless thing, which I still enjoy doing, then at Serverlessdays.io, there's going to be one somewhere near you. I'm pretty sure. So go and have a look, and if you want to become an organizer or get involved, get in touch by the website as well. While we've talked about climate change, Serverless is still a big thing. I've written a lot of Medium posts on that as well, but we haven't talked about that at all.

Corey Quinn: Excellent. Thank you once again for taking time out of your evening to chat with us about this.

Paul Johnston: Thank you.

Corey Quinn: Thank you for that-

Paul Johnston: Appreciated.

Corey Quinn: ... uplifting story.

Paul Johnston: I've really enjoyed it. Thank you for allowing me to speak and to have a bit of time to talk about it.

Corey Quinn: Paul Johnston, interim CTO, and currently working on something stealth mode, and not a crank. I'm Corey Quinn. This is Screaming in the Cloud.

Speaker 1: This has been this week's episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

Speaker 4: This has been a HumblePod Production. Stay humble.

View Details

About Nicole Forsgren, PhD
Dr. Nicole Forsgren does research and strategy at Google Cloud following the acquisition of her startup DevOps Research and Assessment (DORA) by Google. She is co-author of the Shingo Publication Award winning book Accelerate: The Science of Lean Software and DevOps, and is best known for her work measuring the technology process and as the lead investigator on the largest DevOps studies to date. She has been an entrepreneur, professor, sysadmin, and performance engineer. Nicole’s work has been published in several peer-reviewed journals. Nicole earned her PhD in Management Information Systems from the University of Arizona, and is a Research Affiliate at Clemson University and Florida International University.

Links Referenced

  • Twitter Username: @nicolefv
  • LinkedIn URL: https://www.linkedin.com/in/nicolefv/
  • Personal site: nicolefv.com
  • Company site: cloud.google.com/devops
  • Manifold: https://www.manifold.co/

Transcript
Speaker 1: Hello and welcome to Screaming in the Cloud with your host Cloud economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on this state of the technical world and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode of Screaming in the Cloud has been sponsored by Manifold, Manifold powers marketplace infrastructure that connects millions of developers to the best APIs, tools, and services in the fastest growing communities, and also Kubernetes. They offer a complete toolkit that allows you to deliver your API first product to millions of developers. Check them out at manifold.co. Again, that's manifold.co.

Hello and welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Dr. Nicole Forsgren, who's a VP of Research and Strategy at Google Cloud.

Nicole: One of these days that title is going to stick.

Corey: I insist it always will stick. You may argue with me that you're not a VP there and I will well actually you into saying, yes, you are.

Nicole: I like it.

Corey: Thank you. You're also the coauthor of the Shingo Publication award winning book, Accelerate: The Science of Lean Software and Dev Ops, and you're also best known, the reason we're having this conversation again, is that you're known as the lead investigator on the largest dev ops studies to date. You've been a successful entrepreneur with an exit to Google, a professor, a performance engineer, and for your sins assist admin. Your work has been published in multiple peer reviewed journals, whereas my work has been published on Twitter. Nicole, welcome to the show.

Nicole: Thank you. Thank you.

Corey: So when last we spoke, we did a podcast recording about the Accelerate State of DevOps report and it was glorious because people really engaged with that episode. They listened, they were excited and right when we got to the Cloud part, and we'll pick this up next time. Well, when is next time? We didn't tell anyone.

Nicole: They were angry.

Corey: Oh, the natives were very restless. But now, when is this going to happen next? That's right. Now.

Nicole: We're here. We have arrived.

Corey: We are. We are. So if you don't know what the State of DevOps report is and you might be forgiven for that, go ahead and listen to the first episode with Nicole a few weeks back and then come back to this one because we're not going to cover the same ground. We're exploring new ground and we begin with Cloud. What did the State of Dev Ops report have to say about Cloud?

Nicole: Well I mean, I would be polite except this is your podcast. So I'm going to dive right in. I mean we discovered a few things.

So last time we talked about how the elite performers are developing, delivering software with speed stability and they're optimizing on all dimensions. We also found that low performers are on the struggle bus now. So often people are like, well, how do I get better? I'm like, well, there are several, so there are several things you can do. I want to say, the professor in me is going to say, do your homework. You can read the report, you can read six years of the report, you can read the book. It basically falls into a few categories, right? There's continuous delivery and things like technology and automation. There are process-based things like working in small batches. Having good visible systems, viewable systems, right? That also falls into metrics and monitoring. There's also having a good culture. There's also the Cloud, right?

At that point, executives like scrunch up their face at me and they're like, well we went to the Cloud, we didn't see performance improvements and I'm like, wait, so did you go to the Cloud or did you like write the check for the Cloud? Because I mean like I got myself a gym membership, but if I didn't do the work, I'm not going to get any better. Right?

Corey: But that's not what the Cloud sales people told me.

Nicole: I know because they want your money. So like that's really ends up being the big difference in the big problem is that so many people keep redefining Cloud in a million ways. Now, the very smart, astute listeners of this program are going to say, well then Nicole, how are you going to be able to tell me that using the Cloud is a statistically significant predictor of performance if no one knows what they're talking about?

Astute listeners, I love you. So here's what we do. What we do is we come to a clear definition, and y'all, I did not make this up. I look to NIST, right? National Institute of Standards and Technology. They have a NIST framework. And using this framework, there are five essential characteristics. And what we find is that for people who are using the Cloud and executing on these five characteristics, elite performers are 24 times more likely to have met all of these essential characteristics. So that's what it means when I say, if you're in the Cloud and if you're doing it right and if you're not cheating, right? Like I bought the gym membership and I actually went, we see performance gains. So, and however, right? Yes, and. Of the people who say they're using Cloud only 29 percent are actually using Cloud infrastructure and executing on all five capabilities.

So I think this also helps explain this huge disconnect that we see all over the place in industry when people are like, well, I'm in the Cloud, or, but I'm in the Cloud, and I'm in hybrid Cloud, in multi-Cloud, and on-prem, and so many different definitions of Cloud is that we're all talking past each other. We don't always have the same definition, right? We don't have a shared language. That's so important. I mean, I don't know if if you've seen this?

Corey: I see it constantly when you, especially when people try to participate in that least productive of all activities, which is ranking the Cloud providers. The cloudwars.co run by Baghdad Bob Evans insists the number one Cloud vendor is Microsoft. Number two is Amazon and number three is Salesforce, which is a little on the nutty side because until you dig into this and realize, oh, if you're including things like SAS and whatnot and a bunch of other stuff, SAP is beating Oracle is beating Google.

Okay? But at some point that just means software that runs on someone else's computer and it becomes a list without distinction. I would argue it doesn't even matter beyond that of, well is this, let's look at number two for those of us who live here in the infrastructure world of, number two is Azure almost certainly, and then potentially Google, but wait, what if that's backwards? What if the other two, what if it goes the other way around? And the answer is who cares? If that is your decision framework for picking a provider... Who's the biggest or where exactly does it fall in the ranking? Is it number six or is it number seven? Then I'm afraid I just don't care anymore. I don't see how that factors into anything meaningful.

Nicole: Honestly, I love that you brought this up, right? It fundamentally matters how you define Cloud and so we have defined it according to NIST, and these NIST five I'm going to get to in just a hot second because I want to take a half step back.

It matters according to how you define it and how you define it in terms of what is important to your organization. Right? So many people just say, oh, it's just the Cloud. Right? Okay. Does that mean it's a SAS? Does that mean it's highly available? Does that mean it's highly reliable? Also, how does your organization define and depend on highly available, highly reliable? Because right now, so many organizations and so many executives who might not be super technical are treating Cloud, Cloud the Cloud, like it's a commodity and it might not necessarily be a fungible commodity. They aren't necessarily the top two, top three, top seven Clouds for whatever definition they are, are not necessarily completely tradable, right? They're not necessarily the same thing.

And unless you understand what your decision criteria is, unless you know what your definition of what's important to you is, you can't just say, I'm going to trade out this Cloud provider for this Cloud provider, for this Cloud provider, because they actually have different characteristics depending on what you're looking for. And so many organizations and decision makers in organizations don't fully grok that just yet.

Corey: Oh yes. I mean, at this point, I'm running a five person company and we are depending on who you ask multi-Cloud across the board, because all of our infrastructure is on AWS. But we use G Suite for our email internally and yes. Oh, that's right. We are paying GitHub customers or GitHub depending on your preferred pronunciation. The second one is correct. Oh, and of course we have Office 365 floating around. So yes, we are full Cloud customers and we're just not talking about apples to rutabagas here.

Nicole: Right. Okay. So I promised I would talk about these five essential characteristics. These are the five. The first one is on demand self surface, right? You have to be able to automatically provision your compute resources without human interaction. You can't be putting it behind a service now ticket where you wait for someone to approve it, that doesn't count because it introduces delay, it introduces air. It's got to be fully automated and honestly like that's the biggest mistake and the biggest problem we ever see. The second is broad network access. So can you access it from multiple devices? The third is resource pooling. So basically are the provider resources pooled in a multi-tenant model where resources are kind of like dynamically assigned on demand.

The third is rapid elasticity. Most do okay with this, right? So bursting like magic can can you handle like a black Friday situation? And then the fifth is measured service. So systems can automatically control and optimize and report resource use and that's all you're paying for. So, those are the five. Now, something that I think is interesting and important to note here is that these are also architectural outcomes. They're design outcomes, they're automation outcomes. Whether you're on a public Cloud or private Cloud or even let's say you're not on a Cloud. Let's say you're in a main frame environment, you can improve your software delivery performance by architecting your infrastructure with these outcomes in mind.

That's all you have to do. I mean, I say it like it's easy, right? You still have to put in the work.

Corey: Implementation is left as an exercise for the listener.

Nicole: Yes. So I mean, well it's interesting you point that out. Some people get real mad at me because I do this research and I do it with capabilities and practices in mind. It's an evaluative criteria or it's an implementation criteria. I love that you say it that way. I don't do it with tools, which means that I haven't told you which tool to buy. Instead I tell you which things to implement. So now you got to go figure out what tool that is or you got to figure out which practice that is or you got to figure out which bash scripts to build. Right? But now that means that you can select any tool you want. You can select any Cloud you want, you can select any bash script to build. But now you know what things are going to be successful so that if some tool comes out with a new release, like that's fine, just figure it out.

Corey: Which makes an awful lot of sense. Did you do any analysis in the report on multi-Cloud hybrid Cloud? Anything that involves smashing two Clouds or more together? I think that's called a thunderstorm.

Nicole: So we ask what types of trends people are seeing, but it's because people are talking past each other so much. We don't do analysis there because I can't run real analysis until I know I have consistency in measurement, consistency in language, consistency in terminology. And so all we do is basic trends there. We do see an increase in people who are reporting using more multi-Cloud and hybrid Cloud. So I do think that's really interesting. And anecdotally, and by anecdotally I mean like I work with dozens of companies and I go to like way too many conferences, and we are seeing lots and lots of people who are saying that they're using more and multi and hybrid Cloud solutions for lots of different reasons. Right. Sometimes it's for flexibility. Sometimes they say it's because of regulatory concerns. Sometimes they say it's because of risk, right? Like they don't want to be locked into one Cloud that they want to be able to have it for like fail over reasons, like resiliency reasons, so we are seeing more or like more are being reported.

Corey: It's interesting to see that people are going with multiple providers for, from two different angles and one of which I agree with, which is the idea of picking best of breed individual services and putting a workload that makes sense into there. That works really well. The second is, oh we to have resiliency, so we're going to put the same thing on two different providers. Yeah. Then you find that you have more outages caused by heartbeat failures or split brain or some other weird interconnection problem then you'll ever have running on a single provider in a decent HAA setup.

Nicole: Okay. Set up. Yeah. I mean it's really interesting to see that the types of performance applications that you have when you're split versus single, right. Although I will say that that oftentimes slash most times we do see stronger and better performance when you're in the Cloud versus on prem, particularly for SMBs and corporate clients because you can rely on people who are already experts in this space. Right. And by that I mean unless you are very, very, very large, it can be difficult to have to maintain several nines of availability and all of the regulatory requirements surround most of your businesses, right? Like the availability, the uptime, the servers, the physical access, the power requirements, right. Power space and cooling around that. So in general, and I think most of, I mean I almost feel bad repeating this until I actually run into someone who is still trying to argue that like Cloud is not as secure but honestly for the most part, Cloud is more secure.

Corey: It absolutely is from a perspective of a variety of things that people have forgotten about, of having to check the ID of people walking into the data center, of having the low level baseline stuff handled by someone who does nothing but this at all times. The availability improves because when you have an array fail, instead of waking a couple of people in your team up, they're swinging 200 of the best in the industry into place to fix this thing at scale. ... There's also some cover when you're down and everyone else is down, why... People aren't going to be yelling at you in the headlines. It's the provider that gets thrown under the bus. If that's the analogy we're going with this week.

Nicole: Yeah. Well and even so you just said checking the ID of the person who enters the data center. Many times people don't even think about checking the ID of the person that has access to the power lines. Right? Like once you're operating at a high enough scale, you suddenly need to be worried about access to power or what happens if you want to build out a larger data center? I worked, I have worked at large companies that want to upgrade their data center and suddenly they're told that they need to also install a new power center or power grid for the entire city because they have now maxed out power. And that had not even occurred to them because power space and cooling, you're not just paying for the machines, you're paying for the cooling of the machines. Like there's so many implications here.

Corey: There's something to be said for making as many of these problems someone else's, as long as they're not aligned with the core competency of what your company does, it seems to be the right answer.

Nicole: Exactly. Exactly. To a point of scale. Right? I mean there are cases where once you're very, very, very large, having on-prem for certain types of workloads can make sense. But for the vast majority of organizations, Cloud is kind of the smartest move.

Corey: The only thing that's better than paying someone else to solve your problems is tricking chumps into solving them for you for free. Which brings us to open source.

Nicole: Yes, yes, yes.

Corey: One of the findings that I found interesting in the report was that the lowest performers use the most proprietary stuff. Is that an oversimplification?

Nicole: It's actually not so, I mean a tiny bit, but also not really. So when we take a look at the data across tool usage and open source, what we see is that the lowest performers are using the most, what we would kind of term as like fully proprietary software, right? So it's primarily developed in house, it's proprietary to the organization. And what that ends up doing and what that ends up meaning to the organization is you have the fully burdened cost completely dependent on you. You have to build it, you have to maintain it, you have to document it. You are also incredibly limited in terms of recruiting. You have very, very high risk in terms of turnover.

It's really difficult, right? And then when you look at the contrast, when we look at open source, open source ends up having lower costs. Now I'm going to put an asterisk there, right? Because open source, the joke is like open source is like a puppy, right? Like it's free, but you have to pay to maintain it. But those costs end up being, many of those costs and that being very distributed. If you have a question, you can source that question internal to the company. If there's a proprietary piece of the question that's related. But you also have a large community from which you can draw solutions. You have a large community that is actively developing and testing that solution. You have an open community from which you can draw for recruiting and even excitement, right? So excitement in terms of like training. They're actively developing a workforce.

You don't have as many concerns about aging out a workforce, whereas fully proprietary solutions, you alone are responsible for training up and remaining current and understanding what the state of the art is. No one else knows what the state of the art is. The entire community for open source knows what the state of the art is. They are currently discovering bugs and side cases. And what happens if you develop something in this weird, bizarre, complex distributed system, right? Like you and your company... In contrast when you're looking at, in house developed proprietary tools, your organization is the only one doing that. So it's really hard to discover what those types of things look like.

Corey: The challenge that I'm trying to wrap my head around is you talk a bit about the report about COTSS or commercial off the shelf software, that feels like it is the antithesis of open source software. Everyone's open source trying to be smashed into a business model, notwithstanding, and it seems to me that there is definitely a correlation, but what drives it? What, I understand that you're doing the research and showing the data and correlation does not indicate causation, but do you have an operating theory as to why the crappiest of teams often tends to use the commercial-est of software?

Nicole: Yeah. So it's interesting you bring this up because when we take a look at the COTSS solutions, they actually look relatively even, right? So we're seeing who 14 percent, we see the lowest use of COTSS, like just like almost no COTSS or like the lowest use of standalone COTSS in low performers. But we see COTSS heavily customize their low performers like 17 percent there. And then we see pretty consistent use of COTSS at medium, high, and elite performers, 21 percent, 18 percent, 20 percent respectively. Now if I look at COTSS heavily customize at medium, high, or low, we see eight percent, seven percent, 10 percent, so it's not really used. So people come to me and they're like, what's this COTSS thing? Right? Like why? I thought COTSS wasn't good, right? Like commercial off the shelf software, like big, big installations, like ERP systems, right? Enterprise resource planning systems.

Corey: It even sounds sleepy.

Nicole: Right? It sounds like enterprisey like Salesforce or like SAP or whatever. And they're like, but I thought we were supposed to be doing the software development delivery. I thought software was eating the world. So this is a case where if you only look at the data and you don't dig in, it can be, I'm not going to say misleading. It can be misleading or it can be confusing or you cannot know what's happening. But if you take another look, if you take a step down, here's what we see. Martin Fowler has this great article, and we link to it in the report on utility versus strategy. The elite performers are incredibly strong and disciplined at this. The low performers, God bless them, they usually kind of suck at this. Here's the difference. Utility. Think about utilities, commodities. You turn the lights on, right? I turned my lights on, I got electricity.

I'm not going to make my own electricity. I am going to buy it. I need to spend minimal resources to acquire things that are the same thing for everyone. I will spend my resources on my time on strategic things that deliver value that no one else can get. So the things that make me money, that differentiate me, that make me special, that's where I spend my resources and my time and my developer's time. That is where software eats the world. Anything else I buy cheaply and I do not spend my my time on.

So let's start with low performers. Low performers are doing some COTSS, but they're also heavily customizing it. Low performers are not as good at differentiating between things. They're strategic and things that are not. They also love the idea of solving hard problems and deriving everything from first principles. And so they might go ahead and roll their own HR system and roll their own accounting system and heavily customize everything because they are a special special snowflake and they have amazing processes that make them unique yet no, no, you are not.

Corey: That's what you get when you hire off of Hacker News.

Nicole: Ah, y'all, it doesn't matter. You are not that special. Especially if you're a smaller company. Now, if you have, if you're the People's Republic of China and you have like 20 million employees, okay, then it might be worth customizing. But keep in mind, every time you customize a COTSS system and you have to upgrade that COTSS system, you have to un-customize COTSS in order to upgrade. And then you have to re-customize. That is epic amounts of resource that you are spending. Okay. Now in contrast, look at the elite performers. They are disciplined, they are rigorous, they are ruthless. They will say, is this something that delivers core value and competency to the company? Let's say, okay, let's make up a company. Corey, what are we building and delivering?

Corey: Twitter for pets.

Nicole: Okay. We are Twitter for pets. Okay. We are making-

Corey: It's like regular Twitter only 80 times less racist.

Nicole: Yes, and full of cute, cute puppers and adorable kittens. I like this. Okay, so as part of our core value, that is absolutely what we build. We build Twitter for pets. We optimize algorithms to bring the cute puppers and the great gifs and the animated gifs and the cute animations, and it will make sure that the volume is not too loud because my ears don't like it late at night and if the volume is low enough, it will also get past my boss, when my boss walks by. I like this. Okay, we will not roll our own HR systems. We will not roll our own accounting systems. We will not roll our own talent acquisition systems. We will buy those. That is what we will buy and we will be ruthless about it. That is the difference. So that is why our COTSS numbers will look similarly high because yes, we will buy systems even though software is eating the world and then we will spend all of our resources building the perfect Twitter for pets. That's where some of those numbers may look similar.

Corey: Gotcha. So take it one step further for me, please. If I'm reading the report and looking around, and what you've just said resonates in really uncomfortable ways. And I look around and realize that my organization is in fact crap or low-performing is the polite term for it, what can I do about it? Do I start polishing my resume? Do I take it personally in that, oh my God, maybe I'm the problem and no one ever told me this before. What is the next step? I generally imagine that the next step is not to sit there and feel sorry for oneself, but what is?

Nicole: I mean I've been there, I usually do that for about a day and then I drink diet Coke and I eat ice cream. That always makes me feel better. Okay. And then there's a plan of action, right? And we can think about what our next steps will be and it kind of depends on where we are in the organization, right? So if we want to improve our performance, like our speed and stability, there are, like I said, there are categories of things that we can do to get better. There are things like technology and automation. There are things like process, like working in small batches, improving our change approval process. There are things like improving culture.

Now where we sit in the organization changes the types of things that we can influence. So if I'm an IC, if I'm an individual contributor or practitioner, some things that I have great influence over and a lot of amazing power to change are things like implementing test automation, helping with deployment automation, right?

Those are things that have shown to have like great predictive ability over improving software development delivery. Those are great things I can do. If I am at an executive level, right? Some things that only I can change or things that I have oversized influence over are things like improving the change approval process because there's some bureaucracy tied up in there. Right? So taking a look at our process, removing some heavyweight change approvals. By the way, check out the report. We've got a great section on how the cab can move to being a more strategic body. But even just taking a look at the change approval process to make sure it's more clear. It doesn't even have to be automated yet, but if it's clear for everyone to know how to get from start, submitted to end and approved.

Now some ICS can do that too in part by documenting it and understanding what those steps are, the variability in those steps, so that you can highlight it and pass it to the top. Now some things you kind of want coordination, right? Like the CI process. That is a an IC process but it may take coordination among groups because that CI process can span several teams. So understanding where you are and where your impact is going to be the strongest can be a huge win and a huge benefit to organizations. And by the way we do outline this in the book as well.

Corey: Excellent.

Nicole: Oh sorry. In the report.

Corey: Some of our listeners may not have done the homework.

Nicole: I know, it's okay though. That's why we're talking about it because like let's be real. It's kind of long. I'm sorry I got so excited. But we will include a link to the report and the part where I just outlined, that there are some things that happened at the organizational level, some things that happened at the team level and some things that happen at both the team and the organization level, that's on page 39. So now you can just like skip straight to that section.

Corey: Excellent. But if you decide not to read any of these things and instead only pick one page, I highly, highly, highly recommend that you pick, I believe it's page 78 the acknowledgements page because I'm listed there.

Nicole: Okay, so if you want to check out the, that's actually page 79. I'm awful. I'm going to correct you. Okay. But you know what? The part that Corey, well Corey helped with a bunch of stuff, but the part that he helped with like a lot, the part that's probably my favorite is the snark. Do we have time to sneak this in Corey? It's got snark.

Corey: We do indeed.

Nicole: It's real. I think we should and it's a nice throwback to what we were talking about earlier with Cloud. So Cloud helps us with availability, Cloud helps us with reliability. Cloud also helps us understand how to align incentives so that we can reduce costs and have better transparency. Now, the thing I loved about this is I invited Corey to be an early technical reader and I did this for a couple of reasons. First, because Corey knows what he's talking about. Second thing is-

Corey: Secondly because you make poor decisions.

Nicole: I make poor decisions. Third, because like, oh, what's up? What's the polite way to say this? You have a reputation, you were brutal.

Corey: The good kind or the bad kind?

Nicole: I mean you were brutal. You gave me some solid feedback that like, I think I had one sentence in there that was like, oh yeah. Also by the way, like Cloud can help you with cost-savings because we had some really great data around that. And then Corey, I think you included a comment that was by a comment, I mean like four paragraphs long where you pointed out, yes, Cloud can help you with cost savings, but how do you explain this to the CFO? And so we include what this incentive structure looks like and how the shift to the Cloud gives you amazing, amazing opportunity to align incentives and to drastically reduce costs because of greater information transparency.

But organizations have to align incentives better in order to do this, otherwise it won't happen. And so like this is totally sincere. Corey, thank you so much for like helping me understand that I had to get that out of my head because it was in my head. I used to be an accounting professor. I used to do custom managerial accounting, so I knew it, but I had not fully articulated it until you called me on my BS. And so now this is much more clearly articulated in this amazing sidebar on page 37 so thank you.

Corey: Of course. Thank you for appreciating it and asking me. It's, to be clear, none of this is intuitively obvious. It's the you only, I spent three years going through environments where I look at AWS bills and the spend and how that's driving change organizationally and I don't really find a lot of costs savings. That's not the big driver for successful Cloud stories. It's capability enhancements.

Nicole: Yeah. Well it's interesting though because the two can go hand in hand. So I know that the Broad Institute does research. They happen to do research on Google Cloud, but they took an iterative approach, right? So they initially moved to the Cloud and then by understanding that and having this greater transparency, then they were able to change how they were architecting their work and then like gradually reduced their footprint. Right. So you're right, like initially it comes down to increasing your capabilities, but it is possible. You just have to make sure that the transparency is there for your systems engineers and then have those mechanisms available because if it just remains completely invisible to everyone, it just comes down to like a change in how you're building out your Cloud. So it's super interesting.

Corey: It's something that I don't think is well understood even among companies that have gone through it because they take a look and they see that things are attributed much more granularly. You pay for what you use, et cetera, et cetera. And it feels like you're spending less or accounting isn't bothering you anymore because they are so overwhelmed by the fact that you have a massive Amazon bill and they just assume you're buying way too many books and they don't tend to think of it any more closely than that, but whatever the reason it feels more removed and if you start looking at some of the higher level strategy pieces that isn't always born out. Now, does that mean that going into Cloud is a dumb expensive move, maybe, maybe not. That depends on the specifics. But understand what you're going to get out of it. If we're going to recoup our expenses in year one, are you really? One wonders.

Nicole: I mean I do think it is long term, even sometime short term it ends up being a really smart decision. As long as everyone understands what the architectural outcomes should be and everyone's aligned, right? It just, shocker, just requires communication, right? We just have to be aligned.

Corey: Which is harder than you'd think. If there were an API for fixing people, it would need some serious rate limits.

Nicole: I know, I know.

Corey: Ah, if people want to hear more of your wise thoughts, where can they find you?

Nicole: Well, I am on twitter @nicolefv and my website is nicolefv.com.

Corey: Excellent. Nicole, thank you so much for taking the time to speak with me again.

Nicole: Absolutely. Oh, can I throw out one more quick note? Dora's research is all online now at Cloud.google.com/devops

Corey: And I have gone up one side and down the other. It is not partisanly slanted towards Google.

Nicole: It's all of our research.

Corey: I will say it is slightly biased on the first page because they include Google's logo, but not other Cloud competitors logos because apparently those other companies did not have the foresight to sponsor. Frankly, that shows clear thinking on Google's part.

Nicole: I mean some people are smart.

Corey: Exactly. Thank you once again for your time, patience, and continued tolerance of me.

Nicole: Thanks for having me.

Corey: This is Screaming in the Cloud. I'm Corey Quinn and this was Dr. Nicole Forsgren, VP of Research and Strategy at Google Cloud. Thanks for listening and I'll talk to you next week.

Speaker 1: This has been this week's episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com or wherever fine snark is sold.

Speaker 1: This has been a HumblePod production. Stay humble.

View Details

About Andrew Peterson
Andrew Peterson is the CEO and Cofounder of Signal Sciences. Under Peterson’s leadership, Signal Sciences has become the #1 and most trusted provider of next-gen WAF and RASP technology and one of the fastest growing cybersecurity companies in the world. As CEO, Peterson is responsible for overseeing all business functions, go-to-market activities, and attainment of strategic, operational and financial goals.

Prior to founding Signal Sciences, Peterson has been building leading edge, high performing product and sales teams across five continents for over fifteen years with such companies as Etsy, Google, and the Clinton Foundation. In 2016, O’Reilly published his book Cracking Security Misconceptions to encourage non-security professionals to take part in organizational security. He graduated from Stanford University with a BA in Science, Technology, and Society.

Links Referenced

  • Twitter: @ampeters06
  • LinkedIn: https://www.linkedin.com/in/andrewmarshallpetersonSignal Sciences
  • Sponsor: X-Team

Transcript
Narrator: Hello, and welcome to Screaming in the Cloud with your host, cloud economist, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud. Thoughtful commentary on the state of the technical world. And ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This week’s episode of Screaming in the Cloud is sponsored by X-Team. X-Team is a 100% remote company that helps other remote companies scale their development teams. You can live anywhere you like and enjoy a life of freedom while working on first-class company environments. I gotta say, I’m pretty skeptical of “remote work” environments, so I got on the phone with these folks for about half an hour, and, let me level with you: I’ve gotta say I believe in what they’re doing and their story is compelling. If I didn’t believe that, I promise you I wouldn’t say it. If you would like to work for a company that doesn’t require that you live in San Francisco, take my advice and check out X-Team. They’re hiring both developers and devops engineers. Check them out at the letter x dash Team dot com slash cloud. That’s x-team.com/cloud to learn more. Thank you for sponsoring this ridiculous podcast.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Andrew Peterson, CEO of Signal Sciences. Welcome to the show, Andrew.

Andrew: Thanks for having me Corey.

Corey: No, thanks for joining me.

Corey: So, let's start at the very beginning. What is a Signal Science, and given you have several of them what do you folks do?

Andrew: Yeah so the marketing term that we call ourselves is a next generation web application firewall and/or a runtime application self protection tool or RASP. You can thank Gartner for that one. But both tools essentially are about, how do you protect your web applications, your APIs, your microservices that you're running, basically all layer seven type of traffic and across any type of platform that you're using it on? But that's essentially what we do.

Corey: In order to do the disambiguation between, "Oh, a security vendor. I've never seen one of those before." I guess every security vendor in most cases tends to go in a bit of a different, differentiating direction if for no other reason than it's very sad when they don't. But I guess what is it that makes Signal Sciences different than, I guess the typical run of the mill, endless sea of folks at RSA, all independently trying to sell me something with the word firewall in it.

Andrew: Yeah, so that's, I'll start with this. I think a lot of it, it just comes from where we come from and our background. We, for better or for worse, didn't wake up some day dreaming to be a security vendor. So we're the accidental security vendors in some ways. Our background was actually building technology and products and security tools in-house before. So we, me and my two co-founders, we worked at a company called Etsy, that's about 10 years ago when we first started working together. And Etsy, a lot of people know it's E-t-s-y, it's a big retail marketplace based out of New York. And their backstory is actually really interesting from a technology perspective because they were really on the forefront and one of the pioneers around the DevOps movement.

And so our challenge and how we started working together and coming up with some of the lessons learned that has turned into the vendor that is Signal Sciences now, is we were trying to build a security program there that was really counter to I think a lot of the kinds of security programs that we had seen before. Which, the old model of security was, "Look, we're going to be grumpy. We're not going to like dealing with engineers. We're going to blame engineers for all the bad things that they put into our code all the time that makes it insecure. And we're going to tell them 'No, they can't do anything all the time.'" That doesn't really work when the goal of the entire business, and especially the engineering program there, was about how do we empower people to launch code faster, to make changes quicker, to make our systems more resilient and more reliable. These are all the tenets of DevOps and doing that in a culture where you're getting these siloed teams to really work together.

So, in many ways as we built the security program there, it was probably one of the first dev-sec-ops types of, you know I hate using all these buzzwords here, but it was really about how do you get these three teams to work better together? And the lessons that we learned in the context of that were, not only is it really helpful when the security teams can not just say "no," but they can say, "yes," and think about how they can really contribute to making these teams better.

But when you actually start thinking about, if we can build products as a security team that are not only incredibly easy for people to use but also make the engineering teams feel like they are able to learn things about the behavior of people using their applications or using their software in ways that they never were before, that's actually helping them do their job. They actually are going to want to pull and actually use those tools. And so I think that that's been the unique approach that we've had to becoming a vendor, is to say "Like, look, if we're going to go to the dark side and go to that other bad place of security, which is the world of security vendorship, we're going to do it with a lot of empathy for understanding what actually works from a practical perspective, because we were in-house before building a lot of this stuff." But also that our philosophy was that the only way to scale security effectively is to scale it actually through the engineering teams. So we sure as heck better be working with and taking their feedback into account the entire time.

Corey: Absolutely. The hard part I think when you're running an application or a website or any significantly scaled-out service or product is, security has always been one of those things that is inherently an afterthought for most of us. Because everyone likes to say, "Oh, security is job zero," or, "Security is the most important thing." Well, a quick look at what companies spend research and development budget on proves that is not true. It always is something, it's like insurance. Most people should have some form of insurance but you don't expect your house to burn down. So it's never the number one thing you think about when you're when you're setting up something new. But it does need to be something that I guess folks care about. I mean I come from a similar perspective where I look at cloud costing. It's never job one, it's always a trailing function. How does that manifest for you both among your clients as well as for running a company yourself and having a good security posture internally, given that you are a security company and security issues would be problematic?

Andrew: I mean it's a great question. This hearkens back to, and I'll use our initial experience when I was at Etsy in-house before, because you're really struggling over and wrestling over these issues of, every security person on the planet wants to say, "Hey, security is the most important thing." Right? And, "That's all that matters and that's what we should be prioritizing first."

But, I was running product teams before. And our goal of developing products and features and software were really way-more business-related to say, "Hey, we have to get these features out because we're trying to help the business improve, right? We're trying to make money, we're trying to help our customers. We're trying to help people actually get things done."

The initial work with our security team was they were like, "Hey, you have all these potential bugs or vulnerabilities in your code so before you actually can ship this to production, you have to solve these things." Yeah. So that doesn't really work. Because guess what, like the business is going to move forward with or without security. So that's the former relationship that we've seen.

The thing that we've seen both with our customers now, but our big aha moment when we were doing this stuff in-house before was, look, if you're going to go and you're going to talk to an engineer and tell them that they have security flaws in their code, they're going to come back and say, "Well, yeah, that's one of many bugs that I have. I know I have bugs in my code. My question is, why should I prioritize working on this one over other functional bugs that I could go solve that are actually going to help our product actually be better and help our customers actually use the technology better??

And in the past, I think a lot of the response from the security team was because they're like, "Well, because security is important and don't you not want to get hacked?"

Look that's, I get where that's coming from, but it's not terribly productive and it's certainly doesn't speak to the way I think engineers and especially modern engineering organizations are thinking about this stuff. They need data. You need to have some data behind why these things are important. And so for us, what really changed our conversation around that type of thing specifically, right? Like about why should we build in security initially? Why should I even be fixing some of these bugs that I know are security bugs in the first place was, well when we could set up monitoring on being able to track actually what attackers are even attempting to do across, especially, let's just use the application itself, across the different parts of the application, it really changed the conversation.

Like, before I think engineers really thought, and when we would have these conversations they'd be like, "Well, I just don't think that we're actually even being attacked right now. So, security guy with the tinfoil hat over there that's super paranoid about everything, of course they're going to be screaming that we're going to be getting attacked all the time, but I just don't really think it's happening."

So, the easiest way to be able to respond to that was to say, "Okay well we set up monitoring to track different types of attack behavior that's happening on different parts of the app. And at least we could show, look this is the sub-directory or the mobile site or whatever part of the application that you're working on and here's the actual attacks that are happening on that right now." That really made it not only real, right? Like it's like, "Okay, this is real data that we're looking at right now and this is actually really helpful."

But then it immediately actually got alignment internally in the organization to say, "Hey, developer team, I'm not actually fighting against the security team who's just on my back all the time trying to get me to fix things. Security team and development team are now aligned against the real problem, which is the attackers on the outside who are trying to get in." So that data and that visibility slash/ability to have that detection on that type of behavior, it just completely changes the conversation that you can have between your security and engineering teams in-house.

Corey: I think there's also the, we're seeing an emerging, I guess, class of vulnerability as far as... When people wind up going to sleep at night and they work in a company, their prayers before bed are, "And finally, dear Lord, please don't let me be subject to a breach. But if I am, at least have it be something incredibly convoluted and clever, not something stupid like an open S3 bucket." Or whatever it is that winds up... Effectively there's this narrative that's entered the public consciousness, that when a company suffers a data breach, that they are obviously idiots who did not invest at all in cybersecurity and they failed a very basic thing.

And I don't think that narrative works anymore. I think that there is a lot of nuance to this. I think that there is a tremendous number of interesting attack vectors that need to be defended against. Despite what we tell ourselves, it's never going to be the top-most job for a company to care about. But this stuff still happens. And yes, it is a failure, especially when it's not your data that gets breached, but rather the data that you've been entrusted with. But in the public consciousness it's still, "Oh, you got breached, you must hire morons." Isn't true. It simply isn't. Do you see that narrative changing at all in the public awareness, or is that a losing battle from the get-go?

Andrew: I do. And I actually think it's a really important question because there's two sides to this, which is, one is, is it a losing battle for companies to try to change how they're protecting themselves in the first place and try to change their security posture. I think the second question that I think a lot about as it relates to just security professionals overall is like, "Is there any way to win at security? At your job? Are you basically just sitting there waiting to lose?" Which I think by and large it is, or at least it has been for a long time. But I think the thing that's changing and the hope that I have for the industry that's changing a bit is... I like to use the example of a lot of how operations has changed. And how success for ops teams have changed.

And I think in the past, you look 10 years ago about when ops teams and/or DevOps teams were a lot more immature, the expectation there was, "Look, we have to have a hundred percent uptime. We will never go down. And it's a binary concept, right? We're either up or we're down and the goal is 100%." Very similar to security, right? The goal is either "We're breached or we're not breached and there's no middle ground and nothing else matters. We should just try to be never breached ever."

Trying to, and the realities of if that can actually be happening in the more and more complex technology world that we live in where as you said, Corey, there's more and more nuanced ways where people can actually get access to data and what a breach looks like. This is going to be totally different in the future. I think we need to really see a maturity the same way we've seen it on the upside.

I think now when you look at really great ops teams and you look at how the success of those ops teams is even measured in the first place, is that, it's not about uptime and downtime necessarily, but if you do go down, it's only a small functional component of your application. Or a small functional component of the infrastructure. You're also doing a really good job of being able to identify when those things go down and communicate that back to your consumers. You're doing a good job of actually defining and fixing those things faster. And so that the success metrics are not, are you up or you're down, but it's how fast have you identified it? How small can you contain the impact of that service outage? How fast and how well you can actually communicate that back to your customers? And then ultimately you're going for how small of an impact can you actually have on their business and/or their lives or their use of the product.

And I think that that's really where I'm hopeful and starting to see the security community go to. But also, I think I'm also starting to see this from the consumers expectation, is that so many consumers that I talk to, or just friends and family even, are saying, "I feel like having my data get breached on various companies is kind of inevitable." And their judgment on how that breach actually happens... I hate picking on specific breaches, but I think the Equifax breach for example, has continued to stay in the limelight because of how poorly it was handled and not necessarily because of the exact breach itself. I mean, how many other breaches have come and gone in the last few years, and the Equifax one keeps coming up, I think because of a lot of the ways in which the management team and/or the communications around it was handled.

So I think that's the stuff where it's like, "Look, if you have really good communication, we can start scoping out our actual architecture and infrastructure, such that we can reduce the surface area or the amount of data that actually gets breached in a given attack." Those are going to be things that I think are bigger success factors for security teams and security people.

And I'd like to think that's the future of what consumers are going to look at to say, "Hey, this company really handled this well." And not just saying, "Oh, they're just another one that got breached. They must all be dumb. But wow. Of course they got breached because it's inevitable to have that happen to some extent. But I really feel like they were on top of their game. They really communicated this well to me and I actually feel in some ways safer knowing that they're so well informed and were so fast to take action on it."

Corey: I see an awful lot of companies with the mistaken idea that, "Well, we're paying a large cloud vendor to run all of our infrastructure and they have a bunch of services that they offer of varying degrees of utility. What do we need partners for? Why can't we just have everything be first party and that's the end of it?" And the honest answer to that is, "Well have you tried it? That's why." But you can't exactly say that to customers in some cases. How do you find those conversations tend to unfold?

Andrew: So there's a bunch of different things to unpack with this because I think it's, yeah, there's a bunch of angles to that. I think the first is, one of the things I've heard from a lot of customers is... Let's use AWS as an example and... Look, let's actually compare AWS and Azure, right? As two different platforms here. The one thing that folks say, "AWS has a lot of features, right? They have launched a lot of different types of functional features around security." And one of the biggest challenges that I've heard people have using using AWS is, especially if they're... Look, let's say it's a development group, they actually have all the intentions of doing the right thing by setting up the right security features in the first place but they're not security people themselves. And when they talk to their let's say network-focused security teams, the network focused security teams don't actually give them a great roadmap for what to use in the first place. So they're on their own to try to figure out what to select to use from a feature perspective.

And, they're not going to take one of all of them. They're not going to be like, "Okay, I'll turn on a hundred different features." They're trying to figure out what are the basic ones that they start working with and turn those on. And they're not really getting a whole lot of direction I think, from the Amazon folks right now.

So this is one of the areas that I've heard, like Azure in some ways is actually more preferable because it's a bit simpler and a lot more well-defined about like, "Hey, here's a reference architecture from the security feature component of what you should use them when you're using this, right?

So that's, that's sort of step one is... I think folks need a little bit more guidance on what they should be using or not. Then step two would be, I think to your point, Corey, when they start using these features, the question is, "Okay they're there, but are they actually good and are they solving real problems and can I automate these things and are they helping me to actually stop real problems? Or are we reverting back to the, 'Okay well if I just turn it on and I have it there, then I've covered my rear and I'm not going to get in trouble from a compliance perspective or something.'"

I don't like this, right? I think there are certain people that are like, "Okay, I have some of these pieces in place. I'm just checking boxes," because to me that's a reversion back to compliance-based security rather than security that's really focused on solving problems. But this gets back into this issue where it's really hard to find people that have a lot of not only, let's call them cloud and application development skills, but then also have security skills. Most of the people that we have in the security world have a network-focused background and most of the application developers really know applications but they don't necessarily know security.

So that cross-section between the two I think is really... It's hard then to set up systems to be able to say, "Hey, here's the features or the functionality that we're expecting from these different types of products that we're going to add on in our cloud environments," so that they can actually take some type of objective view on the value or the efficacy of that feature or that function.

Corey: Something you said just really resonates with specifically the idea of treating security as something beyond the checkbox, for the compliance dance. For anyone who's ever listened to me for more than 30 seconds, this will come as no surprise, but I have challenges when it comes to checking off box items and doing things for the sake of bureaucracy. I have zero tolerance for that, which makes me not a great employee, but that's beside the point. It tends to make me not the sort of person you want in the room dealing with auditors and dealing with compliance. Because I tend to see those check boxes and get at, "Okay, what is the actual intent behind this control? What is the problem it is attempting to solve for?" And you step down that path and try and solve the actual issues, auditors want the box checked. They want to make sure that you're rotating your API credentials and your IM users every 60 days, for example.

Even NIS doesn't recommend that anymore and the real world that we live in here, well, if you compromise a credential by checking into a repository at GitHub, the time between that happening and the time you start to see it being exploited is less than a minute. It's a 90 day rotation or 60 day rotation, does nothing to stop that. In many cases the alarm that goes off that shows that that's been compromised, is the bill: "Surprise! You've been mining a whole bunch of Bitcoin this month!"

That's where it really tends to fall to, I guess, fall by the wayside. But you can't, as a company, ever bypass compliance and say, "Yeah, it's a stupid requirement so we're not going to do it." You don't get the beautiful shiny certificate that you need to remain in business if you go down that path. But how do you reconcile that?

Andrew: Well, in general, I think the more Coreys of the world that can be running security programs, the better, I think for most everyone. So we are fully in the camp of... Look we, like, as a product category we help to check compliance boxes for a lot of our customers. But we have from the very beginning basically told people unapologetically, "We are not in the business of solving compliance for people. We're in the business of solving security problems." And if we can do both of those things at the same time, great. But the people we work with and the people that we're really seeing start to take over the security industry are really those that are highly focused and highly engineering-focused on exactly what you were saying. Like, I'm looking to understand what the actual problem is I'm trying to solve and then come up with solutions to those problems.

So I think there's probably a series of security vendors out there that are terrified about this movement that's happening where you're getting more and more, sort of less and less auditors controlling security programs. Although there's certainly still compliance and audit programs within every company, including our own. And there's an absolute... I think there is a world where those things are actually still valuable, but splitting compliance and security I think is actually quite important to the future of being able to solve these problems.

So yeah, as it relates back to the original question around like how do we separate... Checkbox compliance is not actually doing anything from real compliance. One of the positive movements I've really seen is that the actual compliance standards, the people that are writing those compliance standards are actually becoming more pragmatic about being able to solve these problems instead of just having a checkbox for a checkbox sake. So that's one of the things that I've seen is, is actually a much more relaxed definition of different types of solutions. So on the actual security engineering side, or let's call it the security side that's focused not just on checking the check box. They really can start to say, "Hey, this functionality that we have here that's really solving the core problem, that the spirit of what the compliance checkbox was trying to check, like we're able to actually still check the compliance checkbox even if it's not falling into that exact definition." Because they're either changing the definition to make it more relaxed, or I think the actual auditors themselves are starting to understand and get smarter about being able to be lax on those things.

So I think that's been a really great change to that part of auditor versus security. I think the other thing that you brought up and even in some of those examples, which is like... Look, I don't actually care necessarily about rolling creds every 60 or 90 days. I really care about when someone is actually compromised those credentials because that's ultimately what is the root of the problem that you're trying to identify. So focusing then and trying to get capabilities around detecting when that happens and then ideally having some sort of automated response to be able to actively respond to that issue, that's really where the technology-based or the engineering-based security group goes immediately to saying, "That's how we identify that problem and that's how we solve it."

And that is heads and shoulders or light years ahead of where we were five years ago of just being like, "Oh well. We have this basic change control in place so everything's good."

Corey: Yeah. We're also seeing security, from my perspective at least, emerge in different directions as far as you have, I don't know, a system that's designed to do one thing. But you take a look at what it's permissions are scoped for and it has the capability to do an awful lot of other things. Now, on the one hand, there's the first approach of, "Hey, how about we alert when it does any of those other things," which is great and handy and useful. But in some ways the better approach might almost be, "Why don't we take away those excess powers that it doesn't need?" The principle of least privilege seems to have in some respects fallen by the wayside. And I don't think it's intentional. I think it often starts as, "Oh, we're going to make it work so we're going to start with a broad scope and we'll come back in step two and narrow it down." But we never get to step two. It gets dropped and we move on to other burning fires.

Andrew: Yeah. This is a tough one because we've lived this in practice from again, sort of previous lives where... Look, if you are living in this DevOps world, which to me, I think a lot of it is about developer empowerment and really being able to actually change who the power groups within these ultimately political organizations are, which is like... The folks that ran hardware used to have a lot of that power because they had huge budgets to buy big hardware pieces, and now a lot of the investment money is actually going into the development organization and actually building software. And so guess what? The power is going over there as well. So the default attitude from a lot of those groups is to basically say, "I should have access to everything to be able to do anything I want at any time, because if I don't have access to everything, then it'll slow me down and I can't do anything."

So A, yeah, you want to be able to empower people to do things and move fast and be able to get access to things. But like you got to have a responsible conversation around that, which is... I really think things like GDPR are really lending themselves to saying, "Okay," well especially if we're thinking about this from a data access perspective, let's really think about privacy and data privacy by design being something that we implement at the beginning, such that we can not only limit access for different people internally to different types of data sets, which I think is just a great thing to do from a security hygiene perspective in the first place. But it's also actually falling into this compliance standard that we need to follow now because of things like GDPR.

So this is where these I think good changes that are happening in the industry right now as it relates to how we're thinking about implementing new types of compliance standards. I think the new compliance standards, I think they give a lot of people headaches, I think sometimes. But I think the intent of what they're trying to do is good, not only for consumers and access to the data around that, but I think it's also good just as a basic engineering practice to make it so that not everybody has access to all different types of data internally.

Corey: No, I think that you're absolutely right. It's an evolving question about what the right security posture is and how that winds up mapping to an individual organization's needs and requirements. The hard part is figuring out where people fall on that spectrum. And then of course, figuring out why we were going to invest in that before you get to the point right after you really, really, really should have been investing in this.

Andrew: Yeah. Well and it's, to be totally fair it's not an easy conversation. It's not an easy change I think for people to make. Because there are meaningful trade-offs between access to data and speed, versus privacy and security architecture or responsible security architecture. And I am actually not in the camp of saying, "I'm going to dictate this is exactly how it should be, this way or the other." But I think at the very least, things like GDPR again, I think are forcing people to have these conversations and it's good to just have the conversation.

Because let's put it this way. If you want to go down one road and you're going to say, "Hey, this is going to be our philosophy and we're going to make this decision," at least you're making that consciously as to say, "We understand that this is a more risky path to go on because we are making a lot of these tools or a lot of this data or a lot of these systems, whatever you want to call it, like a lot of these things available to more sets of people internally than we would on a decision that would actually be more optimized around less people having that data. But you've made that consciously. And you've actually had that conversation internally.

Where I think, in the past the default was just to be like, "Look, we don't even need to have that conversation," because it just wasn't even something that people were thinking about at the beginning. And they probably would have made different changes, or they might've made different decisions on that architecture or on those internal policy decisions if they had had that conversation in the first place.

Corey: I think that it's always a hard part to wind up getting buy-in, and to some extent a company's security posture is almost entirely going to be dictated by how effectively information security leadership is at articulating a vision and telling a story. If we want to be cynical about it, we could even extend that to spreading fear, uncertainty and doubt around what could possibly happen and scaremongering in order to drum up budget. I mean hey, whatever it takes.

Andrew: I think it comes back to, again, these are actually, it's nice to have some recurring themes I think in what we're talking about. But those teams and security teams that I've seen have way more success at being able to either create a culture of security that's more embraced internally, and/or just creating tie-ins with other different business units internally are the ones that are able to show and use visibility. Like, basically making investments into visibility around what actual attackers are doing across their system versus just again, sitting there and saying, "The sky is falling, the sky is falling. We need to focus on security."

And if people aren't, then they just revert into this, "Well nobody ever cares about security and we're never going to get anything done unless we have buy-in from the top." You know it's, it's just not, I don't think that's an effective route to do things and I don't think it ever will.

But being able to use data and use real-time information and visibility, again, that you can point to, to all these teams internally to say like, "Look, this isn't a theoretical thing. This is a real thing, that we are being attacked in these different places all the time and what we're going to do is we're going to be smart about how we set up our security programs, to make it so that it doesn't hinder your job. Ideally, it would actually help you do your job better, but at the very least we're going to make this stuff so easy and really understand your goals as different business units internally to make sure that we're not impacting those goals."

Like, yeah, that is a completely different way to approach that discussion rather than just being the guys that say no to everybody all the time.

Corey: Absolutely. Andrew, thank you so much for taking the time to speak with me today. If people want to hear more about what you folks are up to, where can they find you?

Andrew: Yeah, for sure. I think everybody runs, at this point everybody's building some type of software and everybody's running some type of web application or service. All these themes that we're talking about today really fit into what we're talking about, which is we help give you visibility over, yeah, the people that are trying to impact or attack those different layer 7 architectures that you guys have. You can come find out more at signalsciences.com. I promise we won't brow-beat you with too much vendor speak.

Corey: We will hold you to that. Thanks again for taking the time to speak with me today. I appreciate it.

Andrew: Yeah. Thanks so much Corey.

Corey: Andrew Peterson, CEO of Signal Sciences. I'm Corey Quinn. This is Screaming in the Cloud.

Narrator: This has been this week's episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

Credits: This has been a HumblePod Production. Stay humble.

View Details

About AJ Stuyvenberg
Aaron Stuyvenberg (AJ) is a Senior Engineer at Serverless Inc, focused on creating the best possible Serverless developer experience. Before Serverless, he was a Lead Engineer at SportsEngine (an NBCUniversal company). When he's not busy writing software, you can find him skydiving, BASE jumping, biking, or fishing.

Links Referenced:

  • Twitter: @astuyve
  • Serverless.com
  • Serverless Blog

Transcript
Speaker 1: Hello and welcome to Screaming In The Cloud with your host cloud economist, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world and ridiculous titles for which Corey refuses to apologize. This is Screaming In The Cloud.

Corey: This week’s episode of Screaming in the Cloud is sponsored by X-Team. X-Team is a 100% remote company that helps other remote companies scale their development teams. You can live anywhere you like and enjoy a life of freedom while working on first-class company environments. I gotta say, I’m pretty skeptical of “remote work” environments, so I got on the phone with these folks for about half an hour, and, let me level with you: I’ve gotta say I believe in what they’re doing and their story is compelling. If I didn’t believe that, I promise you I wouldn’t say it. If you would like to work for a company that doesn’t require that you live in San Francisco, take my advice and check out X-Team. They’re hiring both developers and devops engineers. Check them out at the letter x dash Team dot com slash cloud. That’s x-team.com/cloud to learn more. Thank you for sponsoring this ridiculous podcast.

Welcome to Screaming In The Cloud. I'm Corey Quinn. I'm joined this week by AJ Stuyvenberg, a senior engineer at Serverless Inc. Welcome to the show, AJ.

AJ: Thank you, Corey. Thanks for having me.

Corey: So, we've had Austin Collins, the founder and CEO, if I'm not mistaken, of Serverless Inc. on the show before, but enough has changed since then that it's time to have a different conversation ideally with a different person. So, we've at least validated now that two people work at Serverless.com.

AJ: That's correct. And there are at least two of us.

Corey: Excellent. We have not yet proven that you folks stay longer than 15 minutes. But that's a Serverless joke. Ba-dum-tiss!

AJ: Love it.

Corey: So, let's start at the very beginning. What do you do at Serverless Inc?

AJ: Yeah. So I am a senior platform engineer and I'm working on some of our new features that we launched on July 22nd that we call the Serverless Framework Dashboard. It's sort of a sister product that is launched along with our Serverless Framework CLI that everyone kind of knows and loves already and the goal is really to offer a full life cycle Serverless application experience.

Corey: Got you. Let's start at the beginning. I'm sure you've told the story enough, so I'm going to take a whack at it about what the Serverless Framework does and please correct me when I get things hilariously wrong.

Once upon a time we would build lambda functions and things like it in AWS or their equivalent in other lesser five providers and you would wind up seeing that it was incredibly painful and manual to do yourself. There were a lot of boilerplate steps, a lot of things that were finicky and you'd tear your hair out. Then we sort of evolved to the next step, at least in my case, of writing terrible Bash scripts that would do this for us. From there we iterated the step further and thought, okay, now I could do a little bit of this with Python and hey there's this framework. Now I write in my favorite configuration language YAML where I wind up effectively giving only a few lines telling it what to do. This interprets it in Node, the worst of all languages, and then spits out effectively under the hood cloud formation then applies it it with no interaction from me into my AWS environment. Have I nailed the salient historical points?

AJ: Yeah, I think you did along with your beautiful commentary as well.

Corey: Well thank you. And that's effectively what this did a year ago when I finally made the switch, saw the light, et cetera, et cetera. My ridiculous newsletter, for example, has about a dozen of these things all tied together under the hood for the production pipeline and hey, using Serverless made this a lot easier. But everything I've described so far, first has sort of been in a bit of stasis for a little while, and secondly is entirely open source and available to the community. So first, have I just been not paying attention? Or has there been a bit of a lull in feature releases until recently?

AJ: That's a great question. We've been making a lot of headway supporting a lot of different runtimes just in general, along with lots of, you know, supporting lots of new features that AWS has launched. So specifically, I would point to the recent launch of EventBridge. Almost two weeks after that was launched we actually had a support for it inside of the Serverless framework. So a lot-

Corey: As of time of this recording, it's been out for about a month and there is still no cloud formation support.

AJ: That's correct. We had to implement it using a custom cloud formation resource.

Corey: Because everything is terrible.

AJ: Yeah. And to get something done in life, you have to suffer a little bit.

Corey: Well, NDAs are super important and whenever you're building something at AWS, you sign one that agrees that you won't tell the rest of the world what you're doing. Then you start building a new service and you sign a second NDA saying that you won't breathe a word of what you're building to either the CloudFormation or the tagging teams.

AJ: That doesn't make any sense to me, but I don't work there.

Corey: I can imagine that none of that is actually true, but that's my head narrative of why things come out without baseline support for these things.

AJ: It does feel like that sometimes and a lot of what we've done over the last year is really try and support the breadth of the services, not only inside AWS but elsewhere. Because what we found is to make a compelling offering on a framework level, we really have to have everything that people want to do, right? If people end up going back to that sort of a Bash script, deploy pipeline kind of world you described earlier, we've really failed. Right? So in that time that maybe some people perceive as us going dark, we've really been working on supporting a lot of different services and making, making improvements to the framework that maybe aren't as big as far as like a big splashy product on Hacker News kind of launch.

Corey: Okay. I would also point out that I am the absolute worst kind of customer in that I'm not really a customer at all. Everything I've been using the Serverless framework for is freely available as open source. I pay you nothing and I complain incredibly loudly. I'm like the internet brought to life.

AJ: Absolutely.

Corey: So with that in mind, what is effectively the business model here? At what point do people start paying you? I imagine it's not an out of the goodness of their heart situation and I don't think that the investors behind you folks have suddenly decided to turn this into the weirdest form of philanthropy we've ever seen.

AJ: Yeah, that's a great question. So alongside with the framework and open source contributions we've made over the last year, we've also been really hard at work on this dashboard product and that's what we actually do sell. There's commercial offerings. It is completely free to try. Free up to 1 million invocations. We'll track them, we'll give you all sorts of insights onto what your services are doing. You'll be able to use things like our secrets functionality, which will allow you to encrypt secrets and then decrypt them at run times so you can pass them between your service without actually having been floating around in plain text or in get repo.

You can use safeguards, switches and policies, code framework. It allows you to control what your team can and can't do, like which regions you can use, which AWS accounts you can employ to, when you can deploy, et cetera. All this stuff is completely free to use. But we do have paid plans and that's where we do make money. So after you go past a million invocations in a month we'll charge $10 per million of invocations and $99 per seat. And then we do have Enterprise plans that are available, which allow you to run this entire thing on your own cloud infrastructure.

Corey: Awesome. Okay. So let's break down some of the releases in a bit of a ... I guess order that they were released. Correct me if I get any of this wrong. The big one that really caught my attention was that the Serverless framework is apparently now going full life cycle around offering things around testing, deployments, monitoring and observability and probably a few pieces I'm missing.

AJ: Yeah, that's completely correct. On July 22nd we announced this kind of new and expanded Serverless framework, which includes the real-time monitoring, the testing, the secrets management, other security features. They live inside the Serverless framework dashboard, which is kind of integrated into this Serverless framework CLI that we already know. And again, this dashboard is completely free for you to use, up to a million invocations per month.

Corey: Got you. So the way that I always thought of Serverless framework is it wound up ... I would run effectively a wrapper command around a whole bunch of stuff deep under the hood and it would package up my Lambda functions. It would deploy things appropriately. It would wire up API gateway without forcing me to read the documentation for API gateway, which was at the time impenetrable. It's like a networking Swiss army knife. It can do a whole bunch of different things and the documentation was entirely in Swiss German. So you'd sort of get it wrong, get it right by getting it wrong a bunch of times first. But now it's added a bunch of capabilities that go beyond just pushing out functions and keeping them updated. What have you added?

AJ: Yeah, absolutely. So the biggest thing would be the kind of monitoring and observability capabilities we've added on this dashboard. So we'll get you insights into things like hey, a brand new error was just detected. Here's the full stack trace pointing to where the error was thrown. Here's how many times it was thrown in the last X amount of time. Here's the complete reconstructed logs from Lambda, it kind of allows you to immediately diagnose and describe the issue to your coworkers or yourself to go off and patch and figure out and solve the problem.

That, those types of insights are sort of available kind of in aggregate where you're able to see ... Okay, so let's say for example, during the average week I might do 5,000 invocations per day and then one day I might do 10,000 or a hundred thousand invocations. We'll trigger an automated insight that says, "Hey, this function is now doing a lot more invocations than it was doing previously. This might be something you want to look into."

So it's the sort of the full life cycle of your application more than just the packaging, more than just the configuration of your services that you're interacting with inside of your favorite cloud provider, but also bring it all together and experience that, you know, someone who's not necessarily traditionally familiar with Serverless would be able to understand and grapple with.

Corey: Understood. So if I take a look across the ecosystem now, I think that the biggest change is that historically I would use the Serverless framework to package these things up and get my applications up and running. I'd use a different system entirely for the CI/CD stuff that I was doing. I would pay a different vendor to handle the monitoring and observability into it. And now it seems like you're almost making a horizontal play across the ecosystem. What's the reasoning behind that?

AJ: Yeah, that's a great question. So we think we're in the best position to offer the best experience using the Serverless framework. We don't think that anyone should be forced to cobble together their own solution using multiple providers or writing their own custom log pipeline to do analytics or any of the sort. We think that we should be able to offer something compelling out of the box and easy. After all, that's kind of the Serverless promise. Like get up and running, very little configuration, scale to zero, scale to infinity, paper execution.

And that's the type of thing we're trying to bring to the entire lifecycle of your app. Because once it's running in production you need more than simply a really ease of access of different services and really an easy way to package and deploy your application. You need to monitor it, right? You need to handle secret management. You need to make sure that proper safeguards are followed and things are done according to your company or your group's policies. And you need to be able to keep an eye on things. And that's what we're trying to do. We're trying to be the one stop shop for all things Serverless.

Corey: Got you. It's interesting because historically in order to get all these things done responsibly with best of breed, you had to go build a microservices architecture by stringing together a microservices vendor strategy where you have a bunch of different companies doing the different components for you and then tying that all together into something that vaguely resembled some kind of cohesive narrative. Now it seems like that's no longer the case.

AJ: Yeah, absolutely. You know, the downside was sort of that approach and experiences that you end up with this really sort of fragile ecosystem surrounding your application. And these applications don't live in a vacuum. They have to interact with other services, other applications. So to have this sort of really immense configuration alongside of it simply to monitor your applications isn't really a solution anymore. So we needed this way to have one place to go and look and see what is my service? What is my application? What is my function doing at this time? And why is it broken? Let me get there and fix it quickly.

Corey: Right. And that tends to lead to some interesting patterns where you effectively have to pass through a whole bunch of different tooling in order to get a insight into it. Which I guess raises the real question I've got. Again, this is not a sponsored episode. You're here because I invited you here. It's ah, nice of me, wasn't it? But it also means that it's not a sales pitch. So you get to deal with the fun questions. Namely, if I'm going to effectively pick a single vendor and go all in with them for all of my Serverless needs, why wouldn't I go with, for example, AWS themselves, if that's what I'm doing? I mean they have services that do this, they have services that do everything up to and including talking to satellites in orbit. So if I'm going to wind up going down that strategy, why pick you instead of them?

AJ: Great question. So the answer is simply that we think we offer the best experience on top of the Serverless framework that you're already using. We understand everything that's going on in that Serverless EMO file that you're configuring. If you have multiple Serverless apps, we are understanding how they're talking across things like API gateway or SQS or SMS. So it's a lot simpler for us to give you a perspective, you as the customer, a perspective of your application that mirrors what you understand it and not simply a bunch of little services linked together.

Now I think there's competing offerings all over the map here. And if you still want to go through the joy of creating your own log pipeline or all of your own metrics or ingest system or monitoring or what have you, you still can. The Serverless framework is still completely open source. You're free to do that. But if you're looking for one place to get up and running quickly, to get started and get your code out the door to production as simply as possible, I think we offer the best solution there.

Corey: Got you. I've got to say that's ... as much as I like to challenge you on this, I obviously agree. I've been using you folks for a while now. So what came after the full life cycle release? There was something else.

AJ: Yeah. Just a couple of weeks after we finally announced the Serverless components, which is sort of a new take on using Serverless services in your entire application ecosystem. The idea is you should be able to deploy an entire Serverless use case, like a blog or a registration system, a payment processing app or an entire full stack page. Should be able to do that on top of whatever you're doing in the cloud without ever managing that configuration. Right? That's kind of the vision behind components. And the idea is that you can define these use cases, these Serverless use cases as components and you interact with them in a way that you would be familiar with if you are using React, for example.

Corey: Got you. Didn't you release something that was called Serverless Components a year or so ago?

AJ: I think it went into beta officially a year or so ago and then we finally released a GA.

Corey: Okay. So is it fair to view this as effectively, I need a portion of an app to do a thing. Maybe it's an image re-sizer, it's a classic canonical Serverless example. And normally you might consider using something like AWS' Serverless application repo, but maybe you don't hate yourself enough to wind up using SAM CLI instead. So this winds up meaning you don't have to make that choice.

AJ: Yeah, you can sort of pick and choose what aspects of Serverless use cases you want. And like you had said, the image re-sizer is like a super, super common example, but there's so much more than that. Right? If you want to run a really, really simple monolith out the side of ... or I'm sorry, on top of your application, you can. There are examples for how to do this where you might have like a ... what we would call like a mono Lambda structure where you have a rest API that's routed under the root domain, right? And this entire application can simply be deployed with one command using Serverless components.

Corey: Got you. So as you look at this across the board, what inspired you folks to build this out? What customer pain was not being met by other options?

AJ: Yeah, that's a great question. I think the biggest was reuse, right? When we talk about developer practices, things like Solid, what we really want to do is reduce the coupling between aspects of your software. And we're trying to do the same thing for Serverless use cases. So instead of having ... You might have that image re-sizer could be part of one application in one aspect or one area of your microservices architecture, but you're going to want to use it somewhere else likely. And sometimes that means either redeploying it or other times it means simply has to go around a route to do that. Either way with components you can package these things up in a really easy to use way. Include them just like you would any other piece of code, right? And then inject it into your service. And I think that's where that really came out of. That was kind of the inspiration.

Corey: So how much of what you've built out as a part of Serverless Components is, I guess, tied to your enterprise offering versus available to the larger community?

AJ: Yeah, it's 100% open source right now. There are a few last steps we have to complete before we'll tie it into the enterprise offering. Like I said, we did just launch it. However, I don't think that road will be very long.

Corey: Yeah, there's some of the things you've said are compelling. At various times in my evolution of what I built, it would have been useful, for example, to have a payment gateway that I could have dropped in rather having to build my own.

AJ: Totally.

Corey: For better or worse, I'm not irresponsible enough to try rolling my own payment system or shopping cart or crypto. So I smiled, nodded and paid someone else to make all that problem go away. But there's something valuable about being able to take what other people have built and had audited and done correctly and just drop it in place.

AJ: Precisely. And that value extends to more than simply the use case inside of that generic image re-sizer or payment gateway. But put yourself in a position in a larger corporate environment where you might have several teams working together and let's pretend that their application has specific API contracts that kind of bind the different services together and you want to just deploy that middleware layer anywhere you want inside of your application. You could write that as a component and then simply share it so you have one source of truth for that sort of interoperability and you can share that between all the different teams. And now it's very simple to get started instead of kind of each team implementing their own flavor, which I'm sure you've experienced at different parts in your career.

Corey: So something else you launched recently was called Safeguards. What is that?

AJ: Great question. So Safeguards is a feature that is built into our Serverless dashboard. What it is really is a policy as code framework, which means you can define different policies for your Serverless applications in code. Now we include several for free that you can try out. Some really simple examples are Whitelist. You know, AWS regions you can deploy to. For example, you can Whitelist specific accounts that you can deploy to. You could restrict things like, for example, people often wildly over-provision IAM roles with wildcards. So we can easily restrict things like no wildcards in your IAM roles. You can also-

Corey: Is that done via config rules? Service control policies? Something else?

AJ: It's done by ... it's actually ingesting your Serverless YAML file. So because we understand what you're trying to do and we are the ones who are ... like, our framework is responsible for translating your YAML into cloud formation and then we can actually, we can use safeguards to digest that and appropriately allow or deny those configuration changes. But it's more than just configuration management. It also allows you to control when your application could or couldn't be deployed. For example, if you're one of the many groups that has a no deploy on Friday policy or you say no deploying on Friday afternoon.

Corey: Careful, say that three times and you'll end up summoning Charity Majors. They yell at you.

AJ: I personally believe that we should deploy forever and always, right? As frequently as possible. But I understand that some people don't.

Corey: Well that depends, too. Are we talking about code that you trust or code that someone else wrote?

AJ: Absolutely. I mean we're at the size at Serverless Inc. thankfully where the answer is both. We have seven people who've been working on this Serverless dashboard offering, so the group is small and the knowledge is tribal at this point. But we're still, we're growing fast and we're making lots of good changes as we go.

Corey: This week’s episode is sponsored by CHAOSSEARCH. If you’ve ever tried managing Elasticsearch yourself, you know that it is of the Devil. You have to manage a series of instances, you have to potentially deal with a managed service. What if all that went away? CHAOSSEARCH does that. It winds up taking the data that lives in your S3 buckets and indexing that and providing an Elasticsearch compatible API. You don’t have to manage infrastructure, you don’t have to play stupid slap-and-tickle games with various licensing arrangements, fundamentally, you wind up dealing with a better user experience for roughly 80% less than you’ll spend on managing actual Elasticsearch. CHAOSSEARCH is one of those rare companies where I don’t just advertise for them, I actively recommend them to my clients because, fundamentally, they’re hitting it out of the park. To learn more, look at CHAOSSEARCH.io. CHAOSSEARCH is of course all in capital letters because despite CHAOSSEARCHING they cannot find the caps lock key to turn it off. My thanks to CHAOSSEARCH for sponsoring this ridiculous podcast.

Corey: How do you wind up building something like this, I guess in the shadow of AWS, because they kind of cast a large one? Where this started gaining traction and then it felt like they realized what was going on. Shrieked, decided they were going to go in their own direction and started trying to launch the SAM CLI, which despite repeated attempts, I can't make hide nor head or tail of, and it still feels to me at least like it is requiring too much boiler plate and it doesn't make the same intuitive level of sense that the Serverless framework does. That's just my personal opinion, but it seems to be one that's widely shared. You take a look at the rest of the stuff that you're offering and they are building offerings around that stuff as well. At some point, does it feel like you have diverged from them in a spiritual alignment capacity?

AJ: I think we diverged from the very beginning. I mean our goal is to let you build Serverless applications on whatever cloud provider you want. We support AWS, we support Azure, we support Google Cloud Platform and we support IBM OpenWhisk. So that's something that SAM is never going to compete with on a philosophical level. They would not build tools for their competitors and that's where I think is kind of the ideological separation of the two. It's really-

Corey: Yes. But not for nothing. I mean that's valuable in a tool, sure. But at the same time, how many people do you really see using the Serverless framework and then deploying from the same repository, for example, into multiple providers?

AJ: Yeah, that's a great question. I haven't personally seen it. I would expect that that will probably come up a lot more as different vendors kind of continue to either dominate or introduce new features that people want to use. Obviously it's all about capabilities, right? It's all about using services that these vendors provide and that's something that I think we have the most compelling offering on right now.

Now our question was, why would you build something like this kind of in the shadow of AWS? The answer was we needed it. We weren't getting enough from the services to do what we needed to do. So you know, Serverless Inc. is a big believer in dog feeding our own product. All of our entire dashboard application is all built using the Serverless framework. A lot of aspects of our development are monitored using the Serverless dashboard. So we're using it every day. And we think that that sort of mentality really can put us a step in front.

Corey: I would agree with you and I think there is value in a tool being able to speak to anything. As far as any individual customer, I get the sense that they probably don't care. For example, I care profoundly about your support for AWS functions, but I don't use Serverless technologies from other providers so I could not possibly care less about the state of your support for those things. I feel like it's one of those things that matters in the aggregate, but on the individual customer level it's pretty far down the list of things anyone cares the slightest about.

AJ: Yeah. And concerns about that type of thing vary depending on who you are. Right? Like a developer or an individual contributor like yourself doesn't care about a service they're not using. But a chief information officer really does care if they have the capability to move aspects of their Serverless application from one vendor to the other if needed. So it really depends on the target audience.

Corey: Got you. So next, normally the way that one contributes to an open source project is they open issues on GitHub, which is how I insist upon pronouncing it, but I don't have to do that because I have a podcast. Instead, I'm going to give you a litany of complaints about Serverless for you to address now. That's right. It's ambush hour.

AJ: Let's do it, I'm ready.

Corey: All right. For starters, I have to use NPM to get it up and running, which exposes a lot of things under the hood, namely NPM. Does require NPM in the first place?

AJ: A great question. Right now our framework is published on NPM. We are experimenting with publishing binaries on our own...

Corey: Now in theory, I could wind up just rolling it myself without ever touching NPM and just use the Java script and compile it manually, but that sounds like something a fool would do.

AJ: It does sound like something a fool would do. Yes. And we are, like I said, trying to work through a point where you can download this binary on your own.

Corey: Right, because invariably I find that everything wants different versions of NPM, so I have to use NVM to manage NPM versions and now I'm staring down at a sad abyss that annoys me. I want to be able to do things like brew install Serverless. Or I don't know, app get install Serverless. Or if I'm using Windows I just go home and cry for a while. Then I get a Linux box and then I can Yum install Serverless.

AJ: Yeah, absolutely. I think we see that vision, too, and like I said, that's been on the roadmap and that's one of the things we're really working towards is being able to do binary drop in installations of our framework.

Corey: Okay, next complaint. It feels like it is fighting ideologically with SAM. AWS is a Serverless application model. Part of this is SAM's complete and inability to articulate what it's for in any understandable capacity. You read the documentation, you are more confused than when you started. This feels like it's an artifact of AWS' willingness to be misunderstood for long periods of time and that being interpreted as licensed to mumble.

AJ: Yeah. I mean I'm not going to comment necessarily on your interpretation of SAM, but a big part is buying into sort of the ethos and the vision of the tool you're using, right? Like our vision is to let you just deploy use cases simply and really focus on writing your business logic in the form of a Lambda. You should not be responsible for going out and trying to figure out how to wire your Lambda up to API gateway or SNS or SQS. That's not something that any developer wants to spend their time on. And that's what we're trying to do. We're trying to abstract away the configuration of these services and let you as a developer focus solely on the experience of building your business logic.

Corey: Fair enough. Next complaint. It seems like you try to be all things to all supported Lambda runtime. So there is, of course, the whole story of running your own custom layer, which generally is not a best practice if you don't have to, but it does definitely feel like there are favorites being played. For example, it is way easier for me to build a function in Python than it is in COBOL, which is probably as it should be. But do you find that the experience is subpar depending upon other lesser widely deployed languages?

AJ: Yeah, it really depends. Clearly if you look into the large ecosystem of Serverless plugins that are available, you'll notice a trend towards things like Python and node JS. I think that reflects just the reality of the world we're in right now in the modern web development age. If you do really want COBOL, I mean I know the head maintainer of the Serverless framework and we can talk with them about it, but I don't expect you to get much traction because I don't think it's really being demanded.

Corey: Fair enough. Last time I played with this in significant depth in the wake of the Capital One over scoped IAM rule issue, it was ... you could set an IAM policy ... Sorry an IAM role within a service and it would apply to all functions in that role. But scoping that down further on an individual function basis, for example, that function needs to be able to write to DynamoDB, but none of the others do, was painful. I'd have to wind up rolling a fair bit of custom CloudFormation myself. So I just shrugged and over scoped IAM rules because I'm reckless. Is that still the case or is that changed and I just never noticed?

AJ: So all of the functions take an IAM role statement that you can actually give individual IAM role access control to on an individual level. But at the same time, that still creates a rule. It doesn't prevent you from inheriting it in another function. Cloud security's a really tricky thing and like the Capital One breach and et al. We see that routinely it gets mis-configured and that's kind of a big part of what we're trying to do around Safeguards is to kind of define these and limit their access.

But for your specific question, the answer is you can, right underneath the name of the function, it takes a parameter called IAM role statement, which then takes an array of IAM role statements, affect, allow action, whatever resource, whatever.

Corey: Fair enough. Another one is I need to use the universe of plugins to do a lot of common things that tends to cause a few problems. One, I have to wind up installing them and then NPM shrieks whenever it can't find them and I get to go back into node hell, which isn't really your fault. But then the quality of some of those plugins is uneven. There are plugins out there that let me integrate with CloudFront and Route 53 for domain management, but they in turn then, oh we're going to update your CloudFront distribution. Not going anywhere for awhile? Grab a Snickers. And that's painful, for example, when you're doing this part of a CI/CD pipeline because you pay by the minute in code build. So that feels like it's one of those, I never quite know whether I can trust a plugin as something that is a first class citizen or something that someone beat together in the middle of the night. Is there any curation around that?

AJ: That's a really good question Corey, and the answer is yes. If you go to Serverless.com/plugins we have a full plugin directory that you can search. There are check boxes for certified and approved and community, so they're kind of different levels. Starts at community then there's approved and certified. And that'll be what I'd suggest going to as your first resource, to kind determine-

Corey: And the ones without the check marks install Bitcoin miners?

AJ: I can't guarantee that, I haven't read the code, but it is open source and I would encourage you to do that.

Corey: Excellent. I encourage me to what? Read the code or install Bitcoin miners on other people's systems?

AJ: Read the code and then funnel the Bitcoin funds to me. Thank you.

Corey: Absolutely. It turns out as we said on the show before, it is absolutely economical to mine Bitcoin in the cloud. The trick is to use someone else's account to do it.

AJ: Yeah. It's even trickier with Lambda.

Corey: So one of my personal favorite things to make fun of is naming things and this is no exception, as well. First, the fact that you called it Serverless at all. Are we talking the architectural pattern? Are we talking about the Serverless framework? Are we talking about the company? And it's sort of very challenging, disambiguate that from time to time. First, awesome SEO juice, but on the other side, it feels like that tends to cause a fair bit of customer confusion.

AJ: Yeah. Austin touched on that actually in your first episode with him and I would echo his sentiment that-

Corey: But he hasn't renamed the company since, so we're touching on it again.

AJ: I think we're really pushing the Serverless Inc. to brand the actual company and Serverless to define the framework, would be my answer to that question.

Corey: Understood. And that's fair. We're not done with naming yet. You have plugins and you have components and it's going to become increasingly challenging, at least for me, to keep straight which does which. Am I the only person that's seeing issues with that, I guess, overlap between which side of the fence one of those things would go on? Or is that something that is ultimately designed to be aligned along the same axis?

AJ: I don't think you're the only one confused about that. I think that'd be a stretch to say. Serverless components are really about reusing Serverless use cases. Right? And Serverless plugins are really about enhancing the Serverless framework to do other things on top of the open source offering. So that would be how I would kind of delineate between the two.

Corey: That's fair and understood. So I think that really runs out my list of things to complain about. What haven't I complained about that I really should?

AJ: I think we're all still waiting for NVS Lambda MVPC cold start time to go down. I don't know about you, but I come from a very relational database background using my SQL or PostgreSQL and right now-

Corey: My relational database of choice remains Route 53.

AJ: Oh wow. That's one option. You do send out-

Corey: You can use things that are not databases as databases and it's a lot of fun and scares people away from asking further questions.

AJ: It's true. Anything's a metadata service if you try hard and believe in yourself.

Corey: Exactly.

AJ: One of the things that I would really like to see out of the Serverless ecosystem is a reduction in the cold start time of AWS Lambda functions inside of VPC. That would really allow us to start to utilize all the services that Amazon includes. Things like RDS, right? Databases, your relational aspects, relational databases that you can't use right now and you're kind of stuck to using HTP implementations. Obviously we've seen Jeremy Daily's blog posts about Aurora getting a lot better over the last year and I think it's a great step. But ultimately, I think for me, the biggest thing that I'd love to be able to interact with inside of a Serverless application is a relational database. I think that's kind of the last big piece before all the services that developers are frequently using become available. Things like you know, Reddis or Memcached or Postgres can actually be utilized in an efficient way because right now that that cold start time is just killer.

Corey: Understood. One last question I have for you around this, and it's a bit of a doozy I suppose, is if I take a look across the ecosystem, and as a cloud economist I tend to see a fair bit of this, there doesn't seem to be any expensive problem around Serverless technologies, if we restrict that definition to functions, yet. Sure S3 you can always store more data there and it gets really expensive. DynamoDB if you're not careful, but Lambda functions always seem to be a little on the strange side as far as no one cares about the cost. For example, last month my Lambda bill for all this stuff I do was 27 cents before credits. And if you take a look at other companies, whenever you see hundreds of dollars in Lambda, you're seeing many thousands or more in easy to usage. Are there currently expensive cost side problems in the world of, let's say Lambda functions?

AJ: Yeah, that's a really good question and I'll actually answer that in a couple of ways. The first, we should admit that compute is a commodity at this point. Would you agree?

Corey: Absolutely.

AJ: Right. So like any commodity, the providers are finding more and more efficient ways to provide them. Lambda is sort of the natural evolution of that provision. They previously AWS was, you know, selling their EC2 instances, virtualized instances on top of machines. But they were guaranteeing a certain amount of memory, a certain amount of CPU power. Really like a certain amount of compute was being sold to you. Lambda takes that a step further by not guaranteeing you anything and just saying, we'll run your function when it gets called, which allows them to really pack more of these Lambda functions and runtimes into smaller servers, really at the end of the day. I mean we say Serverless, but somewhere down there are servers, I just don't care about them.

AJ: So from that standpoint, you're correct in saying that the compute bills are generally cheap now. The expensive part depends on which services you interact with and how they're set up. I've read several blog posts about people getting burned by a ridiculous DynamoDB and API gateway builds. There's been popular blog posts that made the circuit discussing how you can save, you know, save 30 or 60% of your AWS bill by switching from API gateway to Application Load Balancer, which I think is all true.

I think a lot of people getting burned on the cost front comes from not recognizing essentially what they're provisioning or not necessarily using the correct data model or data access pattern for their use case. That being said, it makes sense that your Lambda bill will be cheaper than your EC2 bill, for the most part. Right? Your EC2 bill is like your house with the air conditioning running all day long versus your Lambda bill is more like just stepping into your car with air conditioning running and then turning it off when you're done. It's just going to ... it's night and day.

Corey: It absolutely is. The concern that I have is that it's always challenging to wind up convincing a company that's spending, I don't know, $300 a month on their Lambda bill to spend even at least as much, if not more, on the tooling around Lambda. I was using a monitoring product for awhile that would tell me in big letters on the dashboard that this month's Lambda bill is 22 cents. That is 7 cents higher than it was last month. Maybe look into this. Yeah. How about I spent my time doing literally anything else because it's more valuable than that, and it continually almost eroded the value proposition I was getting. I was thrilled to pay more for that than I was for my Lambda functions, but the focus on cost and cost optimization in that scenario felt like a hell of an anti-pattern.

AJ: Yeah, absolutely. I mean there's a price for your time and at 7 cents it seems like it's a little cheap.

Corey: A little bit.

AJ: That being said, you know, scale to zero and paper execution are what Serverless is all about. Only paying what you use for are what it's all about. And I think we're at that point now where it's really enlightening for people that have been AWS customers for five, ten years who are used to paying hundreds of dollars in compute bills to see new services cost pennies, right? Now is that worth an alert in your inbox? I don't know. I would guess that a person in your position doesn't read too much email anyway-

Corey: I email 15,000 people a week. I assure you I read more than you'd think.

AJ: Oh man, that sounds awful.

Corey: People have opinions on the internet.

AJ: Yeah, and they have to be heard. And that's why we follow you on Twitter, Corey.

Corey: Exactly. Wait, people read that? I thought that was a write only medium?

AJ: Nope. No, you'd be shocked. It's a multiplexing system.

Corey: Oh dear.

AJ: One to many.

Corey: Okay, so if people want to learn more about Serverless or what you're up to in particular, for some godforsaken reason, where can they find you?

AJ: Yeah, you can check out what we're doing at www.Serverless.com. You can catch us on GitHub, github.com/Serverless. If you have any interest in following me whatsoever, I don't recommend it for the same reason you don't recommend people follow you, you can follow me on Twitter. Reach out.

Corey: Well, thank you so much for taking the time to speak with me today. I appreciate it.

AJ: Absolutely, Corey, thanks for having me.

Corey: Of course. If you're listening to this show and you love it, please give us a positive rating on iTunes. If you're listening to this show and can't stand it, please give us a positive rating on iTunes. I'm Corey Quinn. This has been AJ Stuyvenberg. This is Screaming In The cloud.

Speaker 1: This has been this week's episode of Screaming In The Cloud. You can also find more Cory screaminginthecloud.com or wherever fine snark is sold.

Speaker 1: This has been a HumblePod production. Stay humble.

View Details

About Josh Stella

Josh Stella is co-founder and CTO of Fugue, the company delivering autonomous cloud infrastructure security and compliance. Previously, Josh was a Principal Solutions Architect at Amazon Web Services (AWS), where he supported customers in the area of national security. Prior to Fugue, Josh served as CTO for a technology startup and, for 25 years, in numerous other IT leadership and technical roles.

Links Referenced

  • Twitter: @joshstella
  • LinkedIn: linkedin.com/in/josh-stella-949a9711
  • www.fugue.co

Transcript
Announcer: Hello and welcome to Screaming in the Cloud with your host, cloud economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This week’s episode of Screaming in the Cloud is sponsored by X-Team. X-Team is a 100% remote company that helps other remote companies scale their development teams. You can live anywhere you like and enjoy a life of freedom while working on first-class company environments. I gotta say, I’m pretty skeptical of “remote work” environments, so I got on the phone with these folks for about half an hour, and, let me level with you: I’ve gotta say I believe in what they’re doing and their story is compelling. If I didn’t believe that, I promise you I wouldn’t say it. If you would like to work for a company that doesn’t require that you live in San Francisco, take my advice and check out X-Team. They’re hiring both developers and devops engineers. Check them out at the letter x dash Team dot com slash cloud. That’s x-team.com/cloud to learn more. Thank you for sponsoring this ridiculous podcast.

Welcome to screaming in the cloud. I'm Corey Quinn. I'm joined this week by Josh Stella, founder and CTO of a company called Fugue. Josh, welcome to the show.

Josh: Thanks Corey. It's great to be here.

Corey: Let's start at the very beginning. What is a fugue? My awareness of the company starts and stops with a T-shirt I got at an event a couple of years back. It's glorious, people have general trouble reading it, it's very nice once you look at the abstract design and then eventually realize it says the word "fugue" in a very stylistic way.

Josh: There is a backstory there. A fugue is a compositional style in music and particularly fugues are constructed out of relatively simple musical phrases that evolve over time and interleave and have produced some of the most sophisticated, beautiful pieces of music ever written. What made me think of it for the company in the software we were writing was a book published in the 1980s that made an impression on me as a young person called Gödel, Escher, Bach: An Eternal Golden Braid by Douglas Hofstadter, which I highly recommend. It's about the nature of complex systems arising out of simple systems. And I wasn't quite sure it would be a good company name, but my colleagues, when I brought it up kind of insisted, so Fugue we are.

Corey: Excellent, and Fugue you shall remain. At a high level what your company does to my understanding, and please correct me loudly and energetically if I'm wrong, but you focus on cloud governance specifically with an eye towards security and compliance.

Josh: Yes, that's true. I would say cloud governance focused primarily as you said, on security and compliance. The way we do that is very different than others who are thinking about security. From Fugue's perspective, the cloud is itself software defined, which means that security is a software engineering problem, not a security analysis problem. We actually create a complete model of the entire application or system that's being run in the cloud and then we can compute against that. We can compute things like, is it in compliance with certain compliance regimes like HIPAA or CIS or NIST? But we can also compute things like has it changed over time? Has it mutated in dangerous ways from one moment to the next? That's the space we're in, our approach is a little different.

Corey: Understood. And security in the cloud is something that no one really pays attention to and it's not particularly interesting because nothing much ever really happens in that space. In a completely unrelated topic, as at the time of this recording, a couple of weeks back, there was a Capital One breach that puts a lie to everything that I just said. And you wrote a technical analysis of how that attack may have been pulled off that aligns almost perfectly with my own assessment of it. Can you take people through the high level of what happened and what you suspect occurred?

Josh: Absolutely, and I think the keyword there is may. We only have a certain number of data. Largely I focused on the DOJ complaint because that's an official document. And I did look at the screenshots of the attacker's Twitter feed, although only after I tried to recreate the attacks. I'm gonna describe it from a technical perspective, and if you want me to, I can go back and describe it more in layman's terms. I think I've come up with a decent analogy for that. And this is a maybe. Some things we know about the attack from the DOJ complaint to the degree that that's accurate, but it's not enough to really piece the whole picture together. I made some assumptions and I'll try to highlight those as I touch on them. And let me start by saying I have really good friends both at AWS and Capital One that are absolutely brilliant engineers, and I think both as organizations do a phenomenal job. This is stuff that can happen to just about anyone.

Corey: I would absolutely agree with that. In fact, taking a look at what we've seen far, there's been a lot of noise around this that doesn't necessarily seem to bear itself out. And a lot of people of course, because it's the internet speculating wildly and passing it off as fact.

Josh: Oh yes. Most of what I'm seeing out there I think misses the most important points. And I hope your listeners get some value out of at least the things as somebody who's worked in cloud security and integrity for years now, I think there's some important things to notice that are largely getting ignored. My theory of the attack... And the way I did this is I came up with a theory and then I recreated this attack in my own environment against my own infrastructure. Not production Fugue infrastructure, structure, it actually would not have worked there, but against my own development environment. What we know from the DOJ complaint is there was a misconfigured firewall. We don't know what kind of firewall, we don't know what the misconfiguration was. And I think a lot of folks are conflating misconfigured firewall with the fact that an IAM role with WAF, which stands for web application firewall, was used. That may be true, but it also may not be true.

In my theory, there is another firewall, a traditional IP firewall that has a bad port. And we have heard from I think fairly credible sources that the attacker was scanning the internet for vulnerabilities. And that's the first point I'd like to make is in the old days, attackers would target organizations specifically and go look for vulnerabilities. Now that still happens, but what's happening most of the time is attackers have automated their looking for vulnerabilities and then they pick targets from those who are vulnerable. In this case it was Capital One, but it could have been Josh's Auto Repair, they just found a gap. They found the back door open in a firewall and started working from there. That might've been a security group that had a bad port open, it could have been another kind of firewall, but it appears that the attacker got access to likely an EC2 compute instance. And I say likely because there is a need to collect metadata off the metadata service and the hypervisor and to assume IAM roles. And that's most likely EC2.

I think that the attacker got in the back door, looked around and found exploitable EC2 instance, and then in her Twitter feed she says she used assume-role and this is where things really get interesting. Assume-role means taking on a different IAM identity. You've got this server and maybe it has an IAM identity that only allows it to do things like, oh I don't know, connect to a database. But once you're on that server, if that EC2 instance has itself IAM permissions to look at other IAM roles, you can go shopping for an identity that allows you to do more destructive things.

Let's assume for a moment, because we know Capital One are really good at this stuff, that that server's native IAM role, that EC2 instance's assigned IAM role didn't have things like list S3 buckets because that's a bad idea, right? You should know what S3 buckets you need to talk to and listing them is like getting a phone directory to where the data safes are. But if you could get into that EC2 instance and shop for an IAM role that did have S3 list and then assume that role, now you get a listing of those S3 buckets. I believe that was the next step.

And then once those buckets were identified in the DOJ complaint, they use some interesting and specific language. They say that there were four commands executed and in the fourth command, they describe it as a sync command. Now S3's API does not have a sync API end point, but the AWS CLI has a utility function called sync for S3. I suspect what happened is the attacker got into EC2, looked at the metadata service's credentials, used those escalated identity and permissions through identity, listed S3 buckets, and then synced those S3 buckets to another collection of S3 buckets. And of course. That would not send a lot of network traffic. It actually wouldn't send any network traffic over a VPC for things like VPC flow logs to catch it with all be S3 to S3 and largely invisible to traditional security tools. That's in a nutshell my theory of what happened.

Corey: I would almost take it a potential step further and wonder if once you wind up assuming a new role, you still wind up having credentials that are now able to do that. There's no restriction that I'm aware of that will only permit that credential set to be used from a certain location. Running this from somewhere completely removed from Capital One's portion of the internet from a host somewhere on the other side of the world, for example, would potentially have been able to do every bit as much of this without having full on access to an instance inside the Capital One environment. Is that your understanding?

Josh: Yes and no. Generally true. However, let's again assume that Capital One are highly competent. We know they are. When you define S3 permissions, you can actually limit that to VPC CIDR blocks that you choose. And assuming that they limited it to certain VPC CIDR blocks, the commands would have had to have come from those CIDR blocks. But I think the point here is that that protection didn't help. Those S3 buckets that things were synced to could have been in another account and that account needed nothing to do with that IAM role, it just needed public S3 buckets to shove data into.

Corey: Everything you say is completely plausible, that's more or less where your article starts and stops. That said though, from my perspective, it feels like there are additional parts of the story. For example, 127 days elapsed between the time that this data was exfiltrated and the time that Capital One was made aware of this by an external security researcher. There was apparently no auditing going on for strange behaviors such as listing 700 S3 buckets and then massive transfer of the contents of sensitive buckets. Am I mistaken on something there?

Josh: I don't think you're mistaken, but one of the hardest problems in cloud security is finding signal in noise. Let's break this down, what might have they detected? A listing of S3 buckets? Yeah, you'd find that command in the logs and so on. But the sync command is doing an S3 to S3 copy. Unlike a traditional data exfiltration, which would show up in the a VPC flow log or in the old world it would be hitting your perimeters of your network, that doesn't hit any perimeters. That is just a command and the data transfer happens behind the scenes, not in the virtualized network that you would be monitoring. Well, my educated instinct, let's say, is that this would be hard to detect. And given that she used both Tor and IPredator to cover her tracks, I strongly suspect the only reason she got caught and this gotten noticed is because she bragged about it.

Corey: And stored the exfiltrated data and tooling she used for this apparently in a GitHub account linked to her name.

Josh: Yeah. This feels like somebody who... I don't like to play psychologist, but somebody who wanted notoriety and I think that was the undoing. She mentions that in the Twitter thread. But what that means is if you are out there operating on the cloud and you're not very certain of how you have configured IAM and S3 and particularly IAM permissions to S3 on production instances, I suspect there are a lot of folks with this vulnerability. It's like in the lateral movement kind of attack that we've seen or evidence of lateral movement, but the lateral movement is through identity on the cloud because identity isn't just user identity, it's system component identity. And it almost forms a new kind of network. And I think it's that complex and that big a problem. We're going to have to come up with security tools. And I'm biased in this because this is part of what we do that considers IAM to be more like a network than just a set of usernames with authorizations.

Corey: One other area for intercepting this, even if it didn't stop the attack but would've flagged it before someone else had brought it to their attention, wouldn't in a relatively well managed environment have this level of sensitivity given that they are in fact a bank, isn't it sensible to say that they should have had something alarm whenever an EC2 instance role called an assumed-role API.

Josh: Well, sure. Doing that is easier said than done unless you have that complete picture of the infrastructure and you can detect all drifts. That's why we take that approach. Otherwise you're tending to look through big, long logs of change of mutation and trying to pluck out from that what looks scary, and that's a really hard problem. Let's rewind for a sec and talk about the fundamental benefits of cloud. It's not a data center, it's a moving surface. It is less like a bank vault, more like an aircraft carrier. You're building something that is constantly in motion and that means there's a flood of API calls going on all the time, every day if you're using it effectively. And that's a blessing because it allows us to build a systems that scale automatically, that operate at high speed to compete with others in whatever your businesses. I was talking to a customer the other day that they do five to 10 production pushes a day.

That really wasn't like that in the old days. I think it comes down to should things be caught? Yes. But what should be caught from that flood of information that is coming across the wire and the old approaches of looking for things that appear scary, just fail. You have to have an understanding of what the known good state of that infrastructure is and the ability to catch any drifts that occur to that infrastructure to elevate the actual changes above the normal functioning of the infrastructure so that you can put eyes on it. Or even better yet, our belief is those things should be automatically healed. That the attackers are automated, the defense needs to be automated. That's a long winded wandering answer. I apologize for that.

Corey: No, no trouble at all. Ostensibly, isn't the value proposition of both Macie and GuardDuty in different ways to identify anomalous behavior such as a EC2 instance that, "Oh, it's always been a firewall until now that is now listing S3 buckets and causing a whole bunch of object gets inputs?"

Josh: Yeah, I can't really speculate as to whether they were using those services and something got lost in the noise or if they weren't using those services. I mean that's what those services try to do-

Corey: Yeah, they could forgiven for not using Macie, they are a bank, but even a bank runs out of money sooner or later, and Macie is nowhere near affordable for any reasonable workload.

Josh: Yeah, fair. Yeah. I think there are other approaches. But I'm gonna go back to things like Macie and others, they're still trying to infer what matters. And I personally believe these systems can be made much more deterministic than that. That you don't need like fancy logic to pluck out the signal from the noise if you know what correct looks like. Things that alter from correct are always suspicious unless they're handled through a proper CICD tool chain. And again, I'll say the cloud is actually a big software computer, it's not a data center. And therefore you can use software engineering approaches like infrastructure as code and policy as code. And I think that's a much more successful way to deal with this kind of thing than trying to find needles in haystacks of data flying by.

Corey: This week’s episode is sponsored by CHAOSSEARCH. If you’ve ever tried managing Elasticsearch yourself, you know that it is of the Devil. You have to manage a series of instances, you have to potentially deal with a managed service. What if all that went away? CHAOSSEARCH does that. It winds up taking the data that lives in your S3 buckets and indexing that and providing an Elasticsearch compatible API. You don’t have to manage infrastructure, you don’t have to play stupid slap-and-tickle games with various licensing arrangements, fundamentally, you wind up dealing with a better user experience for roughly 80% less than you’ll spend on managing actual Elasticsearch. CHAOSSEARCH is one of those rare companies where I don’t just advertise for them, I actively recommend them to my clients because, fundamentally, they’re hitting it out of the park. To learn more, look at CHAOSSEARCH.io. CHAOSSEARCH is of course all in capital letters because despite CHAOSSEARCHING they cannot find the caps lock key to turn it off. My thanks to CHAOSSEARCH for sponsoring this ridiculous podcast.

Corey: There is another position to take as far as blaming people goes. Cause whenever something like this happens, the first thing everyone wants to know is exactly whose fault it was and how irresponsible and terrible they were. And I don't tend to give much credence to that. It's a natural human reaction, but having punitive responses inspires people to hide things. But there is an argument to be made that the current state of cloud security is such that there's more than any one person can hold in their head as a full time job, let alone the fact that most people are not security engineers.

They have a thing they're trying to do and security is part and parcel of that, but it's not their core objective in what they're doing. Understanding all of the nuances of how all these things interplay feels like it's an awfully heavy lift. Maybe not as much for a bank, but given that we've seen this spate of cloud security issues and Capital One is almost certainly not the only company out there as susceptible to something like this, it really does make someone wonder at what point do the providers themselves bear some level of responsibility for simplifying the stupefying complexity that is the security model?

Josh: Boy, you covered a lot of ground there. I'm going to go backwards. Why is it stupefyingly complex? It is stupefyingly complex because there are tons of features of the cloud. I remember back in the in the very early nineties when I was a a new Unix system administrator and I first got root access, I blew up my machine. I did. I got the arguments to a tar command backwards and I replaced the contents of the kernel file with an empty tape. Why could I do that? Should Sun Microsystems have prevented me from doing that? At the time I felt like it. But looking back, no. In fact, the beauty of these very rich and powerful systems is they allow humans to make lots of decisions and be clever. And with that comes risk. With that power comes risk. Will the cloud providers get better at showing people the sharp edges? I'm sure they will, they have over time. But I think this is a more fundamental problem than the cloud providers should do better or cloud customers shouldn't make mistakes.

I think this is actually a physics and biology problem. Human beings typically can only remember about seven discrete pieces of data. This is why phone numbers sans area code in the US were seven digits long. The average person can remember seven things. We are bad at specificity, we are bad at detailed memory. And when you look at one of these cloud environments, let's say, one of our customers might have 50,000 or so cloud resources, a resource being something like an EC2 instance or an S3 bucket. And when you look at all the ways you can configure those, each resource can be configured in thousands, tens of thousands, maybe hundreds of thousands of ways. And multiply those together.

That is not the kind of problem humans are good at solving. It just simply isn't. But we have these handy things called computers and this sixty-year-old practice called programming and software engineering that is very well-suited to this problem. I think the blame lies on our collective imagination to understand that the cloud is actually a big general purpose computer and it needs to be programmed like one and it needs to be automated, not just in terms of its scaling functions and business functions, but in terms of its security functions. That might've been a little too geeky and down in the weeds. I'll take another shot at it if you like.

Corey: No, I think that's an absolutely fair assessment. With great power does come great responsibility. I think that there's a responsibility to defend against sophisticated attacks like this. One thing that I've noticed across the internet in the wake of this has been, "Oh, she worked at Amazon back in 2016 on the S3 team, she must have used insider knowledge to pull this off." Well, you left Amazon. To my understanding, you used to be a principal solutions architect and you left before 2016 and in 2017 they had their S3 apocalypse and then rebuilt the entire system from the ground up. Plus, you've laid out a very convincing and very plausible way this could have been exploited that none of it required inside knowledge. It just required deep familiarity with the platforms and publicly exposed utilities.

Josh: Yeah. I can't know for sure. Apparently she worked on the S3 team, but my instinct is that that is bullshit. If that's not okay to say, I'll do something else but-

Corey: No, please. We'll keep it.

Josh: Okay. I think it's bullshit. I recreated this in about five hours, my theory of how this worked. I've been out of Amazon since 2013, I used no insider information to recreate this. I looked at APIs. I thought about it for a minute, actually for a few minutes, maybe a few hours over the course of doing it and pieced together a way to do this kind of attack. I think she was creative and you know what? So are a lot of people, we should be worried about that. The notion that an identity system, and most people when they hear identity, they think about my active directory user when I log into my machine, that's not what this is. This is the identity of components of the system that can completely circumvent the traditional network boundaries and that becomes a vector again for lateral movement. I don't think there was insider information used. I don't know, but none was needed.

Corey: Absolutely. This is the sort of sophisticated and clever attack that, for example, I would dream up. Maybe not to the same degree, certainly not with the ethical lapse, but there's nothing that's required from an inside baseball perspective on this and saying that, "Oh, it's obviously a failure of how Amazon hires people, that someone who worked there many years ago now did something awful." Theoretically if you were to go and turn evil and do something like this, I think there would be a whole hullabaloo made about the fact that once upon a time you worked there, so it must have been with inside knowledge you did all of these things, and I just think that's crap.

Josh: I think so too. Look, people like things to have a bow on them. They like to have something to point out and blame. Human beings in my experience are very uncomfortable with the idea that we're facing issues that are truly complex, that require a lot of thought, creativity and hard work to solve and instead look for some simple explanation. And I think that's one of a number I've heard. When actually for us to do something productive about this means really understanding what happened in an honest way. And I'm not claiming that my theory is perfectly correct or even, mostly correct. It's the best one I could come up with, but if nothing else, my theory and experiments have shown that that is a massive attack vector and people need a solution to that. I think it's pretty cheap to say, "Oh, this is an ex-AWSer," it could have been just about anyone.

Corey: It's also a bit insulting. Oh, no one could possibly understand AWS unless they worked there for years. Nonsense. Humans can understand anything. Nothing's impossible for a person to wrap their brain around. It just requires dedication and effort more than almost anything else.

Josh: Yeah, absolutely. I'll point back to the original use of the term hacker. It wasn't somebody who broke security walls down, it was somebody who did clever things in C or Lisp. People who are compelled to do creative things with computers get creative with computers. They use them in unintended ways and if you are a bad actor, this is what that looks like. If you're a good actor, you might make the next great application or secure an existing one or what have you. There's no easy answer to this.

Corey: And here's I think the most lingering question that we're faced with. Capital One may have their faults, but they don't hire stupid and they care and they pay attention to this because they know what's at stake. If they were in a situation to fall victim to this, how many other companies are too?

Josh: Every company that's running a digital computer, whether it's in a data center or on the cloud, Capital One does hire excellent people. One of my best friends and most brilliant programmers I know works there. We haven't spoken about this at all by the way. We talk more about things like Rushton Haskell. I know the quality of their team. Asking people to be perfect is unreasonable and I think the real question we have to be asking ourselves as an industry is how do we become resilient? Because perfection is not an option.

Corey: Exactly. And it can't be M&M security where once you break through the hard outer candy shell, everything inside is soft. Defense in depth is critical.

Josh: I would even go beyond that. I completely agree. There is no perimeter. Forget about that. That's gone. It's never really been there. Security is a collection of architectural decisions, it's not a technology you can layer on. And to get security right means understanding the system as a whole. I would argue that not just defense in depth, but defense at every level, all the way up and down the stack where you're doing your best to eliminate these vulnerabilities. And pretty clearly in this case, once that firewall was penetrated and the IAM role-assumption was possible, that was a pretty a soft middle. But I think that we need to think about this in terms of every layer of the stack and baking insecurity as an architectural practice and that is directly at odds in many cases with speed and efficiency. And again, I'm going to come back to, computer science has pretty good answers for this, people just aren't thinking about the problem that way.

Corey: I think you may absolutely be onto something here and I'm beginning to understand why it is that you started a company aimed at solving these problems. If people care more about what you have to say and want to see your thoughts, where can they find you?

Josh: Our company is at Fugue, F-U-G-U-E, .co, dot C-O. I'm joshstella on Twitter and if you want to reach out to me, I'm josh@fugue.co.

Corey: Thank you much for taking the time to speak with me today. I appreciate it.

Josh: Thanks Corey. It's been fun to talk to you.

Corey: Likewise. If you've enjoyed this episode, please leave us a positive review on iTunes. If you hated this episode, please leave us a positive review on iTunes. I'm Corey Quinn. This is Screaming in the Cloud.

Announcer: This has been this week's episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com or wherever fine snark is sold.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Matt Broberg

Matt loves working with technology communities to develop products and content that invite delightful engagement. He’s a serial podcaster, best known for the Geek Whisperers podcast, is on the board of the Influence Marketing Council, co-maintains the DevRel Collective, and often shares his thoughts on Twitter and GitHub @mbbroberg. He’s also a fan of tattoos and cats, though remains unsure of Schrödinger’s.

Links Referenced:

  • Twitter Username: Mbbroberg
  • LinkedIn URL: https://www.linkedin.com/in/mbbroberg/
  • Personal site: Mbbroberg.fun
  • X-Team
  • CHAOSSEARCH

Transcript

Speaker 1: Hello and welcome to Screaming in the Cloud with your host, Cloud Economist, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world and ridiculous titles for which Cory refuses to apologize. This is Screaming in the Cloud.

Corey: This week’s episode of Screaming in the Cloud is sponsored by X-Team. X-Team is a 100% remote company that helps other remote companies scale their development teams. You can live anywhere you like and enjoy a life of freedom while working on first-class company environments. I gotta say, I’m pretty skeptical of “remote work” environments, so I got on the phone with these folks for about half an hour, and, let me level with you: I’ve gotta say I believe in what they’re doing and their story is compelling. If I didn’t believe that, I promise you I wouldn’t say it. If you would like to work for a company that doesn’t require that you live in San Francisco, take my advice and check out X-Team. They’re hiring both developers and devops engineers. Check them out at the letter x dash Team dot com slash cloud. That’s x-team.com/cloud to learn more. Thank you for sponsoring this ridiculous podcast.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Matt Broberg. Welcome to the show, Matt.

Matt: Thank you so much, Corey. It's a pleasure to be with you.

Corey: So, your formal bio's a delight. "Matt loves working with technology communities to develop products and content that invite delightful engagement," is how it starts, which is just ... It feels like someone has workshoped that to no end.

Matt: I mean, you can just do the kind of kissing motion while doing ... Just tastes good when you say it.

Corey: Exactly. But it continues. He's a serial podcaster, best known for The Geek Whisperer's podcast, is on the board of the influence marketing council, despite the fact that somehow no one really knows who you are. I may have added that.

Matt: That part's on the bio, yeah.

Corey: Yeah. Co-maintains the DevRel Collective, which I'm sure we'll get into, and often shares his thoughts on Twitter and GitHub.

Corey: Do you share your thoughts about GitHub on Twitter or vice versa?

Matt: Oh, dear God. I think you're really just reminding me I'm due for a revision on this.

Corey: Well, absolutely. There's no time like preparing for something until right after the point you really need to do, which why not? Game on. I might have you back some day.

You currently work at IBM Hat as a technical advocate and editor for Opensource.com

Matt: In fact, I am employed at Red Hat, which is still Red Hat, and I do. I am one of the editors of Opensource.com, which is a lovely site that you can contribute your ideas of what open source means or share tutorials about it. It's all under Creative Commons and the code is based on anything that meets the open source licensing. So, it's really great place to share and engage in the open source culture and to do so in a way that doesn't focus on code contributions since we all know that's not the only way to contribute back to open source.

Corey: Oh, yes. You read an article of mine, which I'll throw a link to in the show notes at some point, and have had-

Matt: I did.

Corey: ... the good sense never to invite me back.

Matt: Well, yeah. I mean, it doesn't come to memory, so I guess it wasn't on purpose, but yeah, maybe next time, Corey.

Corey: Exactly. And no, you are at Red Hat, which at the time of this recording, has just recently been acquired by IBM. So, the hats are red, but the suits are IBM blue.

Matt: Yeah, I mean you're absolutely right and it's a pretty exciting time to be there. I think I'm more impressed than ever that I was actually interviewing while IBM was in the process of acquiring Red Hat. The news came through during the interview process, so it really caught my eye. And to everyone's credit, it still went through completely smoothly. I just continued on with the interview process because the company was still exactly the company I'm working for now.

So, I feel really good about that and I think it's a real credit. There's a lot of stark to be said about how companies change after acquisition big and small, and this is the first time I haven't seen it completely freeze all hiring processes.

Corey: It seems that everyone I've talked to there has continually nice things to say, which tells me that their legally required to. Which is fine, and that's not the point of having you on.

Matt: No, sir.

Corey: Instead, it's mostly to talk about what you've been up to lately. What have you been doing? What excites you? What interests you?

Matt: You know what, Cory, the thing that's been on my mind is I just finished a stint getting a ... what we might call colloquially, a DevRel, Developer Relations organization off the ground in a technology company. My last three jobs have all been focused on DevRel. And on the side I've been really fortunate to participate in the DevRel Collective, which is this brilliant little brainchild of Dave Josephsen who wanted to get other people that were firmly developer evangelists, talking to each other on the side and getting each other some support because it was so frequently teams of one in the early days.

Well, we've definitely evolved far away from that timeline into it being in the mainstream and every organization, almost to a degree of cargo culting, pulling that term into their organization. And I'm fascinated by this. I see myself more as like a Jane Goodall character in DevRel, curiously watching what's happening from the trees, trying to blend in. And I have some observations and I know you have some opinions, so I wanted to banter with you on that for awhile.

Corey: Oh, I generally have opinions on most things, but to clarify, you would not consider yourself to be a developer advocate or developer evangelist or a DevReloper as I insist on calling them?

Matt: Not at the moment. No, I definitely am good at playing one on TV as need be. I definitely think of it as like a bit of a caricature you have to play sometimes, but in order to be in the DevRel space at a given time, you need to have a particular tribe to be talking to. I think that's an essential element of it. And your goal ... My take on the overall goal is to inspire by example. That is fundamentally what a developer advocate's role is to be.

Google has a definition as well from 2015 that, if I can pull that quote, is, "Developer Relations role is to create a vibrant ecosystem of third party developers by bringing the interface between those developers and their platforms product, engineering and design teams," which is kind of a fun idea of being the glue between these different experiences. At the moment I'm very happy to be just focused on writing articles and talking to people. So, not at the moment but I'm fascinated by it. I like to get into conversations about it on Twitter and I think we're still at a place where it's just so ambiguously defined what that role is and what it looks like depending on where you're getting hired, that you have to really peel back that onion and know what you're signing up for or you have no idea what to expect from your job.

Corey: I have mostly stayed out of those debates intentionally, largely because I know I'm going to irritate a number of people regardless of what I have to say.

And regardless of what side of offense I come down on. So, let me preface the rest of this conversation with a disclaimer, namely that I'm not sure who's on what side of what war is being fought over this, because honestly I have other things I'm focusing on rather than outreach to developers. For my business, developers are generally a champion at best and not someone that I'm trying to sell anything to. So, from my perspective, I don't need developer engagement. That said, let's start at the beginning. What is Developer Relations?

Matt: Well, so I gave you the textbook definition by-

Corey: Yes, but now I want to hear it applied.

Matt: Yeah. By Google's definition, and then ... I really do think that Developer Relations is a category of positions that is ... its main goal is to inspire by example for a given audience. The three disciplines I think that really are enveloped in this new strategy of DevRel are community, which isn't traditionally its own organization, but at sometimes community management is. Engineering, some function of engineering, and some function of marketing. If you smoosh those together and you sprinkle some unicorn dust, a rainbow pops out and then a developer advocate is born.

Corey: Wonderful. And what is, I guess, functionally the role of one of these people? Because I will start off by saying the external crappy view of a developer advocate or evangelist ... We're not sure which they call themselves by because there's strong feelings about either side of that and they're too busy quibbling among themselves, in my experience, to articulate a clear difference. So, cool.

Matt: I thought you weren't going to have an opinion on this, Corey.

Corey: Oh, I said I was going to have an opinion. I just don't have a horse in the race. There's a difference.

Matt: Ah, I see. I see.

Corey: So, I see the quibbling happen and talk to people about what they see the role being, but let's take the prototypical poorly understood crappy example where you have someone who looks like a developer except for the part where they cost way more and their job description, more or less, is to fly around the world to exotic places where a conference is being held. And then they get on stage and they say something like, "Yes, I currently work at Twitter for Pets. Twitter for Pets is hiring, let me know." And that's the last time in their talk that they mention the company because now they're talk about how to choose the proper standing desk for your home office is what they're focusing on and that's not in any way relevant to what Twitter for Pets does. And then they flit off to the next conference and continue to go on, effectively, drinking parties with their friends. Now, this is probably not the complete and capsulation of the role, but to someone not steeped in it, it doesn't sound like it's that far from it.

Matt: Yeah. That can be a pretty tough pill to swallow if you're in the profession right now and you hear that description. That is the most annoying stereotype of what the work is. That you fly to exotic places, you get to hang out on a beach and then you are paid to drink beers with your favorite people because they're are also DevRel and they're in the same space.

Corey: Forget annoying description. I'm pretty sure that was a verbatim job posting two years ago.

Matt: I am not going to advocate or argue against that, but I will say that if we want to talk about what it feels like in practice, I've done the job a few times now for a number of years and one of my favorite things is talking to people in the DevRel Collective and to people trying to break into DevRel. So, what I found is that ... I just want to start by saying there's a wide variance in what the job is depending on the size of the company, the subject matter the company cares about and basically how well funded they are.

I think early startups to mature cloud vendors have wildly different expectations. But in its best form there are a few key aspects of it that tend to come up over and over again, and it has to do with winning hearts and minds by having somebody who is, first off, a member of the given tribe that they're trying to influence. So, you, Corey, if you ever were to dare, could easily be a great dev advocate for a cloud vendor if you wanted to try. Would you ever try that?

Corey: I have the sneaking suspicion that I would be entirely too honest. Let's also not escape the fact that I am a terrible employee for everything that my personality suggests I would be.

Matt: Fair enough. Fair enough. Okay. So, if we just focus on that aspect, you would be good to go. But there are other aspects, right? There is the angle of what is the value that company is expecting to get out of this business investment of a developer advocate or Developer Relations professional? So, in that way, it really depends on whatever executive coughed up the hundreds of thousands of dollars to millions of dollars of investing into this project. And to give some concrete examples, if you look at the top three cloud vendors of Google, Amazon and Microsoft, you see really deep lineups of developer advocates and each of them are focused on some form of public speaking, some form of listening to users in a very strategic way and then trying to provide a better experience for their particular cloud by in whatever way they think makes sense to them.

You recently had Jess Rozelle on the show. I think Jess is one of those secret superpowered dev advocate types who fundamentally is an engineer at heart, but she can't help but go fix things for actual people in real ways. So, it's not just about what's been put on the backlog on Jira. It's what are people piss about? How can I go find the person that can fix that and hold their head towards the tweet that says how bad it is until they actually fix it? That is one of the superpower kind of elements of a DevRel

Corey: And we're calling those individuals a DevRel as a singular now? That's the terminology and nomenclature we've settled on?

Matt: No, I'm terribly sorry and I would love to provide an alternative. So, Developer Relations would be ...

Corey: Oh, I wasn't criticizing. I think every term I've heard about this has been awful, but I'm just willing to figure out what the appropriate term is. I'm not one of them. I use the term that they tell me to use.

Matt: I don't want to put my wood behind the fire of DevRel. It's a terrible term. But Developer Relations is the umbrella organization, if you will, if we're going to thinking about this from a business standpoint. And then within a Developer Relations organization, if you're creating that, you tend to have developer advocates, which are your general purpose unicorns as we're talking about right now. You might also have developer experience engineers in that organization that focus more on SEKs and library support. And then in the larger organizations, they have an entire staff of sort of product and program management teams that help support the initiatives that the DevRel organization is uniquely responsible for.

So, my understanding over at Microsoft is that there's a DevRel organization that has program managers that support their own band of certain events and activities and sponsorships of podcasts and things that differentiate from the core brand of marketing that that company is doing that are more focused on giving back to the community at large and less about immediate conversion for the sake of marketing goals. Which marketing's goals is to convert people into sales opportunities.

Corey: And yet, as you mentioned earlier, they put millions of dollars behind a DevRel program at these companies and they want to see some value in return for that, otherwise you may as well send people you like on far away trips to expensive places. But the question then immediately emerges, how do you measure that value that you're getting from it?

Matt: This is my happy place. This is the thing no one has still cracked the code on and it's absolutely amazing that it's so consistently problematic that ... We're talking about organizations. So, org charts will always have those steady pillars in organizations of marketing, sales, engineering, and the only other essential one is finance because you got to make some money and know what to do with it. But being so new and not being so one to one aligned with KPIs like these four organizations are means the DevRel organization will always look a little alien when we start talking about return on investment.

The ROI of a Developer Relations program has to be aligned with some sort of strategic strategy or else it just looks like the sort of general purpose, we just make things better for the company type statements, which don't tend to land well in a boardroom scenario. But to double down on that for a second-

Corey: Right. Because that always in the rock, paper, scissors game falls to, "I make things less expensive for the company." So, there has to be a value clearly articulated.

Matt: Yeah, exactly. So, I kind of keep going back to business theory in this and I'm like, "You're either a cost center or a profit center." And DevRel, unless communicated in a way that it is providing some sort of coherent value based on the other organizations we're used to, it's going to eventually be a cost center and eventually cut because they can be wildly expensive due to the activities. But why are people investing in them in in droves still? It's that the organizations that currently exist are fundamentally broken at communicating with the target audience.

The idea of most organizations, both their marketing and engineering teams having a coherent conversation with an end user are just busted. You have engineers that will produce whatever their backlog tells them to, not because they want to or it's been validated, but because that's what they have to do. And then you have marketing organizations that are very good at times at talking to the C-suite or talking to the person they perceive to be the money maker, but we're living in the wake of developers being the new king makers and that whole idea-

Corey: Nice reference.

Matt: Thank you, thank you. And not being able to talk to developers is seen as a huge liability. So, we're left in this little bit of a primordial ooze of, what do you do in this state where we don't have the right resources in the right organizations?

Corey: This week’s episode is sponsored by CHAOSSEARCH. If you’ve ever tried managing Elasticsearch yourself, you know that it is of the Devil. You have to manage a series of instances, you have to potentially deal with a managed service. What if all that went away? CHAOSSEARCH does that. It winds up taking the data that lives in your S3 buckets and indexing that and providing an Elasticsearch compatible API. You don’t have to manage infrastructure, you don’t have to play stupid slap-and-tickle games with various licensing arrangements, fundamentally, you wind up dealing with a better user experience for roughly 80% less than you’ll spend on managing actual Elasticsearch. CHAOSSEARCH is one of those rare companies where I don’t just advertise for them, I actively recommend them to my clients because, fundamentally, they’re hitting it out of the park. To learn more, look at CHAOSSEARCH.io. CHAOSSEARCH is of course all in capital letters because despite CHAOSSEARCHING they cannot find the caps lock key to turn it off. My thanks to CHAOSSEARCH for sponsoring this ridiculous podcast.

Corey: From my perspective, and I'm curious to see what your thoughts are on this, I view DevRel as an extension of marketing. Now, I know somewhere listening to this, someone just sat bolt upright and started swearing and possibly almost caused an accident if they're commuting, but I don't mean that in any way as an insult. I consider myself more than a little bit of a marketer. Because the way I see it is that it aligns so clearly with the larger picture of marketing.

When people have a, I guess, angry opinion toward marketing, from my perspective, it's the crappy implementation of marketing. Easy example for that is we're on a podcast. As people have heard at the beginning of this episode, it's sponsored by a company. Huzzah.

Now, people are generally not going to be pulling over to the side of the road, causing a second accident on some of these freeways around here and pulling up long URLs and punching in discount codes necessarily. But there is value as far as brand awareness goes, to affiliating a brand with something like this that people, at least ostensibly, like.

And measuring the ROI on something like that becomes super challenging. So, companies instead try and game the system. Well, if you make them use this code for some discount, then we'll be able to loop back around and track the exact number of impact, but it never works that way. I mean, there's problem of effective frequency. If you see a brand 15 times and the 16th you see the sign, you decide to buy it. Well, that 16th impression winds up getting all the credit and everything else was just a waste of money, but it's not how it works.

Matt: 100% correct. I'm with you all the way, Corey, that we're now really getting into the guts and really the deeper level than most people want to about what is marketing about? If you can tell you're being marketed to, it's because it's bad marketing. If you start from that premise, you can think about how Developer Relations could be a bandaid over the current status quo of marketing in most tech organizations.

We need developers talking to developers about things developers care about and keeping them engaged, excited and contributing back to something larger than themselves. So, you see this a lot in open source communities because it's easy to see that the more I engaged with contributors and maintainers, the more they contribute back and thus the better the code gets. In sort of API driven companies you see that DevRel, when they do a tutorial at a conference that gets people, their first push to their API, they sign up for a new account. So, it's an easy ROI conversation.

When it gets quite tricky is when we focus on where our hearts are here in the cloudy space or the sort of products that are more regular releases of things that people go and run on premises themselves. You have to have a bit of a leap of faith that when I get somebody you trust to talk in front of you, you are going to be more likely to use this software or hardware or whatever it may be than you were before you talk to that dev advocate.

So, there are things that are easy quantitatively to measure, but that doesn't make them qualitatively correct. Finding the line between what's easy to measure it and what's useful to measure is something that we're still way off on as a principal. If I could tell you the one thing that has to be measured across all organizations, I absolutely would, but I have found that there are so many variances in how people define their DevRel program's goals that you have to align the KPI, the outcome that we're pushing towards, with the business initiative or you're just trying to whitewash this idea in a way that doesn't work.

Corey: Absolutely. And I think that if you view marketing through the lens of crappy marketing such as people who are statistics or metric driven around it all comes down to the clicks and getting in front of people in obnoxious ways, then yeah, I wouldn't want to be associated with that either. But as you say, that's crappy marketing. That's not what marketing should be about.

Another direction I've heard people take DevRel in is that it becomes the voice of the customer or the developer back in to inform the product and make it better over time. And if you ever want to see someone struggling to keep a straight face, I would urge you to use that exact phrasing to someone who runs a product org at a company. And then pour six to eight drinks into them and then let them loose and see what they have to say. Because fundamentally, user interviews are an integral part of product teams. Conducting those is its own skillset and making sure that you wind up asking the right questions in the right way is incredibly complex. I mean, I hire someone specifically to do testimonial work with my existing customers just because I don't have that background to be able to ask the probing leading questions to get the story that I'm looking to get from a consulting engagement.

So, from that perspective, a bunch of people with some form of engineering background who are now giving talks, writing blog posts, et cetera, the idea that they can walk out and into a random conference room, have a conversation for 20 minutes with someone that they like and come back with meaningful product feedback that carries statistical weight to product seems a little farfetched.

Matt: Yeah. I mean, you're scratching on that itch again of just, we're asking too much from a single new idea. I do want to give a shout-out, not that he will necessarily listen but maybe he listens, but Kelsey Hightower, I think is a wonderful example of what this looks like in practice where DevRel is adding value in a unique way as opposed to trying to replace what product management would do.

His work with empathy sessions, this idea of gathering users into a room using his star power and expertise in the community space to grow people and then to aggregate them into a place gives him an opportunity to get people using a new tutorial he put together, but then product team members can join that exact same space and get to know the same people and have that one to one relationship or one to few relationship that DevRel is just better at. Because to date, many, many product organizations do not actually talk to their end customers, because to date, many product organizations struggle to get that same level of relationship with people because they're not engaged on Twitter like a developer advocate would be and they don't show up to the conferences regularly enough to be just invited into a room of people.

So, there is a value add, but we have to start talking about this as additive and less as being a replacement for other organizations. We still need marketing, we still need engineers who are writing code and pushing bug fixes and we still need product management that's doing what product management does uniquely. But we're also saying that that's not enough anymore and in the current state of affairs, we need an organization like Developer Relations to provide that social tissue, that connectivity with our target audience. And then to lead by example, by participating in the ways that we want other people to participate.

So, I'm fully with you that there's expertise in each of these nuances to each of these roles, but DevRel has an additive value in this current state of affairs in the current time that we're in, in the current market we're in.

Corey: Absolutely. And one thing I want to address is why-

Matt: It's kind of open for you to argue with me. I'm not going to lie.

Corey: I don't know. I don't know if I agree with you or not. I think that it's a nuanced issue and again, I'm trying to be very clear on this. I'm not trying to call anyone in particular out on anything. Some of the most inspirational people I've ever had the privilege of speaking with are Developer Relations professionals and I'm a big fan of the work and I definitely think it's needed. My concern has been largely around the way that it's talked about in some circles, and I'm not necessarily trying to get hate mail, but I'm also not trying to paper over a lot of, shall we say, unfortunate behavior that I'm seeing in this space.

For example, a lot of companies are effectively trying to significantly beef up their Developer Relations function where they have a tremendous roster of people who are known and respected throughout the entire industry. And the problem that you see is, great, so these people are using their personal credibility to convince me to try something I otherwise wouldn't try? Well, absolutely. I like these people. If you tell me something's good, I'm going to believe it. And then it turns out it's crap, and 12 months goes by and, "Hey, try it again. It's not crap anymore." I'm not going to trust you the next time that you say this.

So, I fear that people are, in some cases, burning personal credibility because they feel they either need to or they actually need to based upon what their employer is measuring them based upon.

Matt: Ooh, that's some tough love right there. But I can't disagree with the premise that when being in a Developer Relations environment, when that is your role, part of what you're doing is you're putting your credibility out there as a source of value for the company you're working for. That's part of the idea of the top-of-funnel marketing draw that you provide that other people in marketing can't provide with their latest version of a webpage or new place where you can add your email to get a cookie.

You're totally right that people are putting themselves out there, and I think being on the other side of that more than a few times, I think the hope is that you aren't too blind to the fact that it has flaws as well and you can honestly advocate for the right way to use something and the wrong way to use something and when it is good and when it isn't good. And you can request the kind of feedback that's appropriate to the life cycle of the product.

Sometimes I've had to push back and be like, "I know you want me to get people to go buy this, but I think I actually need to tell people that it's in beta and here's an open source release they should download." So, you're talking business strategy at that point and talking about maturity of the audience and whether they're willing to accept something where it's at even though it's not where you want it to be yet. And having that heart to heart with people that are well above your pay grade, but a good dev advocate has those tough conversations when they need to.

It's part of the job for sure, to put your yourself out there and ideally you can stay honest in your job. I have a core value that I don't want to be at a place I need to lie. I would rather quit or get fired before I'm lying for somebody and that's resulted in both.

Corey: Absolutely. And I think there's also a strange spectrum that we're seeing in the DevRel space. Easy example of this would be ... The canonical example I always like to go back to is Matty Stratton at PagerDuty. He's been doing this for 20 years, so if he sits down in a chair at a stage and spins it around, straddles it backwards, because for some reason in my mental image, that's always how he would do this. And he would say, "I have seen some stuff. Let me tell you about it. You're doing paging wrong."

Now, I don't necessarily believe that I am. I've been doing this for not quite as long but not a short period of time either. But if Matty's going to get on stage and he's going to say something that provocative, he's got something to follow it up and he has my attention.

When you have someone who's relatively new to the industry on the other end of that spectrum and maybe one to three years of development experience, I'm a lot less likely to believe that necessarily, because I've been through the cycle enough times to see tomorrow's amazing shiny technology turn into yesterday's garbage and that leads to a certain cynicism at some point. Oh, you've got a new way to do ... I don't know, logging or paging or whatever it may be. I'm cynical. I'll hear you out, but I'm not going to suspend my disbelief.

And what's strange to me is that as the more I talk to people in this space, the more it seems like there's almost a bimodal distribution of people who've been either doing this for well over a decade or people who were relatively new to the industry. Is that something you're seeing or is that a selection bias on my part?

Matt: That's an interesting observation. I want to explore that with you. I see it more as related to organizational size and expectations for the role that on the younger startup space, people are like, "Oh, everyone's saying dev advocate and DevRel, we should probably hire that too." And these individuals join an organization, it's unclear who their leadership is. They're probably rolling up into marketing or into engineering and then actually rolling up into marketing. And then they're asked to do everything from blog posts every week, write as many talks and get them accepted around the world at first-class events, and then also maintain the documentation, maybe a few dozen GitHub, GitLab repositories and then just do some more on the side because that's not enough. And they kind of have the trial by fire version of what it means to be in an organization.

They didn't want to go into engineering directly because maybe their coding isn't quite to that level and they like to socialize, so they wanted to be more on the social side of things. That's what I see a lot of startups hiring for in the DevRel space. While larger organizations are really looking for subject matter experts with deep expertise that are talking about a specific topic in depth, and that is more of the fundamental thing.

So, I see more as an organizational struggle and maybe a budgeting struggle if we're going to get honest about this, that it's easier to hire people earlier in their career in smaller organizations because it more aligns with a startup budget as opposed to larger ones.

Corey: So, if we take a look across the ecosystem around people who have successfully transitioned into becoming developer advocates or whatever ridiculous term we want to use for them. I still go with DevReloper and I will die on that hill. What is it that, I guess, that separates them out from someone who continues down an engineering path or goes in a different business direction?

Matt: It's a fantastic point to mention that DevReloper ... Why DevReloper? Wait, can we have a quick aside? Of all the things, not DevRelian, not Developer Avocado or something like that?

Corey: Because when I say DevReloper, people first think they misheard it and then it finally sinks in and they just stare at me with this disappointed look on their face and it reminds me of growing up.

Matt: That's fantastic. Okay, so what does it look like from here? Okay, I think Developer Relations in practice is ... it's engineering that empathizes with its users. It means spending more time on documentation and discussions with end users than most people get the opportunity to do in their day. It's also mostly good marketing. It's the kind so good you actually give a damn when people say the words in front of you that sound like the words in the right order and not like word salad. And it's also community leadership. It's people that are uniting those that care about a given technology in a given space and providing empathy and leadership to do so.

So, what I love about it is that anyone who is really good at any of those skills has every right to be in the Developer Relations space and to lead as a DevReloper and whether that is a forever position or whether they decide, "You know what, I want to do this more on the product management side," or, "I want to work back into engineering and lead developer experience efforts." Or they want to be a technical marketer or a product marketer, there's a wide breadth of expertise that they can bring to any organization if they do a stint in DevRel.

Corey: Sounds completely reasonable. Matt, thank you so much for taking the time to speak with me. If people want to hear more of your sage advice, where can they find you?

Matt: I'm always happy to banter on Twitter at M.B Broberg or you can hit me up at mbbroberg.fun to reach out.

Corey: I'm sorry, did you say dot fun?

Matt: Absolutely.

Corey: Do you ever get ... Well, it's just perfect setup pitch to make fun of someone and your mind goes completely blank and all you can do is point and do a ha-ha.

Matt: Yeah, I was really hoping, but I think I've set you up so much for my willingness for you to make fun of me that it's almost kind of surprising to you and you seem to be [crosstalk 00:35:20].

Corey: At some point I start to feel bad for you. Are you enjoying this? That's kind of sad. Not to kink shame, but my God.

Matt: No, this is lovely. No, Corey, I can't thank you enough. It's a pleasure to be here and I'm happy to talk to anyone about this stuff or anything else related to my absurd profile.

Corey: Matt Broberg at Red Had/IBM Hat, depending upon who you ask. I'm Corey Quinn. This is Screaming in the Cloud.

Speaker 1: This has been this week's episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com or wherever fine snark is sold.

Speaker 4: This has been a HumblePod production. Stay humble.

View Details

About Nicole Forsgren, PhD

Dr. Nicole Forsgren does research and strategy at Google Cloud following the acquisition of her startup DevOps Research and Assessment (DORA) by Google. She is co-author of the Shingo Publication Award winning book Accelerate: The Science of Lean Software and DevOps, and is best known for her work measuring the technology process and as the lead investigator on the largest DevOps studies to date. She has been an entrepreneur, professor, sysadmin, and performance engineer. Nicole’s work has been published in several peer-reviewed journals. Nicole earned her PhD in Management Information Systems from the University of Arizona, and is a Research Affiliate at Clemson University and Florida International University.

Links Referenced:

  • Twitter Username: @nicolefv
  • LinkedIn URL: https://www.linkedin.com/in/nicolefv/
  • Personal site: nicolefv.com
  • Company site: cloud.google.com/devops
  • X-Team: x-team.com/cloud

Transcript

Announcer: Hello and welcome to Screaming in the Cloud, with your host cloud economist's, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on this state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is screaming in the cloud.

Corey Quinn: This week’s episode of Screaming in the Cloud is sponsored by X-Team. X-Team is a 100% remote company that helps other remote companies scale their development teams. You can live anywhere you like and enjoy a life of freedom while working on first-class company environments. I gotta say, I’m pretty skeptical of “remote work” environments, so I got on the phone with these folks for about half an hour, and, let me level with you: I’ve gotta say I believe in what they’re doing and their story is compelling. If I didn’t believe that, I promise you I wouldn’t say it. If you would like to work for a company that doesn’t require that you live in San Francisco, take my advice and check out X-Team. They’re hiring both developers and devops engineers. Check them out at the letter x dash Team dot com slash cloud. That’s x-team.com/cloud to learn more. Thank you for sponsoring this ridiculous podcast.

Corey Quinn: Welcome to Screaming in the Cloud. I'm Corey Quinn. I am joined this week by Dr. Nicole Forsgren. Nicole, welcome to the show.

Nicole Forsgren: Thanks so much for having me.

Corey Quinn: Thank you for joining me. So, you work at Google cloud these days as a VP of Research and Strategy.

Nicole Forsgren: I mean, let's call that aspirational. I'm not a VP just yet.

Corey Quinn: I understand the Google's org chart is not caught up with your magnificence. Other people are willing to cut them slack. I am not. You are a VP to me. You will remain a VP, and eventually the business cards will reflect that very bright reality.

Nicole Forsgren: I'll take it. Yeah. Right now, my title is research and strategy.

Corey Quinn: Yes, you've done so much that it's difficult to start out, to even figure out where to start with what you've done and who you are, but so let's take it in stages. You've somewhat recently wrote the book Accelerate, The Science of Lean Software and DevOps, which is a fascinating book. I recommend that people check it out if they're at all interested in, I guess, putting a little bit of data to anecdata, but that's not where we you really began. To do that, let's go back to the very beginning. Who are you?

Nicole Forsgren: That involves me just starting out in a small farm town in Idaho, but maybe we want to go farther than that. I actually started out, it's interesting because some people are like, "Oh, you're just a researcher. You're just an academic." But I'm glad you asked this because I started out as a software engineer. Well, I guess I started as a programmer. I was on mainframe systems, but then, that was a software engineer at IBM. So, I was developing systems, and then, as I swear this happened so often, I had to maintain my own systems. So, then I was sysadmin. I was running my own systems, and then I ended up doing some consulting a bit because I wanted to help other people run their systems, and build their systems, and solve more interesting problems.

And then, I actually ended up in hardware for a bit. I was running RAID, which that's kind of a blast from the past, right? We don't do RAID the same way we used to do RAID.

Corey Quinn: Well, not on purpose anyway.

Nicole Forsgren: I know, right? And then, I ended up going to get my PhD, because I realized that kind of cycling through some of these consulting problems and even solving some of the problems in larger organizations, because I'm just bouncing back and forth between consulting and IBM for a couple of those last several years. It felt like many of the problems I was solving and many of the complex problems and organizational problems felt like I was answering some of the same problems in the same way. And in particular, when I was going to management and suggesting solutions, many times they were saying, "Oh, well that won't work here." Or, "Well, I know that worked there but that won't solve this problem."

Nicole Forsgren: And I was thinking, "Well there has to be some type of way to solve this, in some way that's more generalizable." I wonder if there's a classic problem that can be solved similar ways. So, that kind of led to the PhD and doing some research.

Corey Quinn: What is your PhD in?

Nicole Forsgren: So, my PhD is in MIS, it's Management Information Systems. And the reason I chose MIS as opposed to computer science is, I liked the fact that I could link technology and computering things with business outcomes, right? So MIS is inherently an interdisciplinary field, and back in the day, it was unique because it really specifically was linking and tying computer science concepts to business outcomes. That really is what I've done for over a decade now, is find ways to deliver business outcomes, or organizational outcomes, or team outcomes from computer types of things, like capabilities and practices. So, I was, this is such a hipster term, but like, it's like, "I was doing it before it was called DevOps."

Nicole Forsgren: And really, I kind of was, so I started doing my research in this area in '07, which is pretty parallel to a lot of the DevOps movement. And then, I finished my PhD in '08.

Corey Quinn: Excellent. So, one could say almost that you've brought ivory tower academia into the streets?

Nicole Forsgren: Actually, yeah. In many ways I did. And also that was in parallel with a handful of other academically rigorous research. So, there were a handful of people about the same time I was doing my research that were at IBM Watson Labs, right? So Cadigan, and Maglio, and Haber, a handful of people there, they were studying sysadmins specifically in some of their work practices. I started a bunch of my research with sysadmins as well, going to the LISA Conference a few years later, I chaired LISa. Then, I expanded my research to include developers and other engineers, software engineers, and a bunch of my work was focusing on how capabilities and practices in tooling, or automation, or process, or culture had impacts at the team, individual, and then organizational level, which if we think about it, that kind of is how we think about and define DevOps now, right?

Nicole Forsgren: It's tooling and automation, it's process, and it's culture, and how that has impacts at largely the software development and delivery, and then organizational level, how we deliver value.

Corey Quinn: All of that is made manifest in this year's State of DevOps report, an incredibly thorough academically researched paper except that a human being can read it. It's probably the best way to frame that from my perspective.

Nicole Forsgren: Yes. I often joke that I speak two and a half languages, English, academic English, and a little bit of Spanish.

Corey Quinn: Also, add math to that list.

Nicole Forsgren: Yes, yes, a little bit of math, more statistics than other types of math. And what we try to do is we try to take this really academically rigorous work and translate it, not just translate it, but also make it very, very accessible to people so that they can use it. Right? So, I've been leading, and running, and conducting the State of DevOps Reports for six years now, starting in 2014, now through 2019, so these reports are super accessible. I joke it's like an adult picture book, right? Like we have large type, we have graphics, we have pictures. It's very easy to flip through. It's about 80 pages, but it's like very large print. This is not like dense text.

Corey Quinn: Oh, and it's so gorgeously designed. I had to triple check to validate that you folks were still part of Google.

Nicole Forsgren: I have to say my copy editor and my designer are fantastic. Cheryl Coupe and Siobhan Doyle are unbelievable, unbelievable to work with. I will say the last couple of weeks of copyedit and design are a little intense. They're a little rough, but they turn around the most gorgeously designed work and they really helped me. We worked very closely together to make sure that it's very accessible. It's easy to read, it's easy to navigate. We're working to put out a couple of pages of an executive summary as well. So, if you just want to like flip through and find something that's really quick, that's available as well.

And then, in addition to this, like you mentioned, me and my co-authors for the book, Jez Humble and Gene Kim, also pulled together the first four years of the research into something that's a little more detailed, right? That includes additional descriptions about the capabilities we've researched, additional information about the outcomes that we've measured, more detailed information on the statistical methods about what it means, and the methodology and where the data comes from, and why we choose the statistical methods that we do. And then, part three included a contribution by previous Shingo winners, Karen Whitley Bell and Steve Bell on a case study out of ING Netherlands. And then, the book itself just to Shingo. And I will take them out of-

Corey Quinn: Congratulations.

Nicole Forsgren: Thank you. It's the first time as far as we can tell that a Shingo has ever been awarded to anything in technology. Now, I will say that came out of 2014 to 2017, so we have two more state of DevOps reports, research projects that have been published since then. So, my editor keeps pinging me, asking for a second edition. So, as soon as I take a few naps, we will work on that. And I did want to mention really quickly I highlighted the authors for the book, the authors for this year's report. So, I led, I was first author, Dustin Smith is a researcher who joined this year's report.

He was fantastic. He has a PhD in top stats for five years. So, he was wonderful. And joining this year's report, Jez humble was third author, and then Jessie Frazelle joined as an author this year as well. She was wonderful, wonderful to work with.

Corey Quinn: She's been a previous guest on this show, and we'll absolutely have to have her back to talk more about some of this.

Nicole Forsgren: Yeah, I think she's going to join us on another podcast where we will dig into all sorts of cloud and open source excitement that we covered this year.

Corey Quinn: Excellent. Excellent. So, before we dive in too far into the intricacies of this year's report-

Nicole Forsgren: Oh, there are so many things.

Corey Quinn: And there are, but the problem I've seen in most reviews and most discussions around the State of DevOps report is that no one starts off with a primer for someone who's never heard of it before. So, from that perspective, guide me through it. What is the Accelerate State of DevOps Report? Where did it come from? What is it for, and why do I care?

Nicole Forsgren: So, what would you say it is you do here, Nicole?

Corey Quinn: Exactly.

Nicole Forsgren: So, the nice thing about this report and the thing that makes this so unique and so different is that this is not just another vendor report, right? We're not selling a technology, we do not talk about vendor tooling or products anywhere in the report. I think there's one line that lists a whole bunch of tools as an example, right? What we do instead is we investigate the capabilities and practices that are predictive. So, if someone says, "I'm doing the DevOps," or whatever you want to call it, find and replace, whatever your company is doing, whether it's technology transformation, or digital transformation, or DevOps. If you say, "I want to know what types of things are actually impactful, which things are actually predictive of success in a statistically meaningful way.

Now, go back a little bit, right? Hit rewind on this podcast. Remember how I said I used to do consulting or I used to do these things in my organization, and my manager always said, "Ah, that's not gonna work here." Well, this helps answer that. Like it says, in a statistically meaningful way, these things will actually have an impact. There's a high likelihood this will work. So, this research takes an academically rigorous approach. So, I designed this from a research designed, PhD level standpoint. We designed this research to test a bunch of hypotheses to say, "According to the research, according to existing literature, according to lots of other things that suggest, these types of things have a good likelihood of having a difference in lots of different types of organizations. What will actually work?"

Then, we collect a bunch of data, and then, we see, "Okay, what works in a statistically meaningful way? What does the evidence show? Then, I'm going to break that down a bit. I say capabilities and practices, but we don't test tools. The reason we don't test tools is because, well first of all, there's a million different tools, That's going to be too hard. Also, tools change, right? Feature sets change, capabilities change, lots of different things change. So, instead what we do is we test capabilities and practices, because then, what that does is it gives you an evaluative framework. So, then you can go back. You can go back to your organization.

You can go back to your team, whether you're an IC or you're a leader and you can say, "Okay, these types of things will work. These types of things have a high likelihood of working." Okay, so when I'm doing CI, CI has a high likelihood of meaning that you will be more successful in developing and delivering software with speed and stability. What does CI mean? Also, everyone like redefined CI to be their own special thing. What does CI, in order for CI to be impactful, what does that mean? It means when you check in code, it results in a build of software. When you check in code, automated tests are run. You need to have automated builds and tests running successfully every day, and developers need to see those results every day.

Those four things need to be happening. Now, anyone can go back to their CI tool set of choice and they can say, "Are these four things happening?"

Corey Quinn: What I find fascinating about all of this, as I read it, it's very, again, first you brought the data. So, every time I see someone starting to argue from ... make a point, an anecdote, or pull a well actually against anything that you ever list in these reports, it's screamingly funny to me. I just immediately cringe and hide behind the tarp because there's going to be a bloody red mist where that person used to be by the time you're finished with them, metaphorically speaking, you bring the data and they're [crosstalk 00:15:58]-

Nicole Forsgren: I can be polite.

Corey Quinn: You are.

Nicole Forsgren: But, yeah, I've got data.

Corey Quinn: Yes, and what you say is right.

Nicole Forsgren: And we retest many things every year. We revalidate things several ... Some things have been revalidated for six years. Now, not everything needs to be revalidated every single year. We rotate them in and out, but we also do the revalidation thing, right? So, it's like this really has been revalidated several years. You can fight with me if you want, but if it's not working for you, maybe you're not actually doing it. Maybe it's not actually automated. Maybe it's hidden behind a manual gate. Like you're putting it in service now and you're waiting for a person to click it. I love you, but I award you no points, may God have mercy on your soul.

Corey Quinn: Exactly.

Nicole Forsgren: Like, citing Billy Madison. What part of this thing is not actually working. What part does not match?

Corey Quinn: Right. What I like about this before is that I did a lot of digging into it last year when I saw this, and really paid attention to it, is you come up with this idea of performance profiles, where you talk about high performing teams, elite performing teams, low performing teams, and I always wondered, didn't get the time-

Nicole Forsgren: People get real defensive, people get real defensive.

Corey Quinn: Well, that's what I wanted to ask you about, to some extent. Very few people self identify as, "Yeah, as far as performance goes, are company is complete crap. Thank you for asking." People like to speak aspirationally about their own work and unless you wind up working at Uber, generally you don't show up hoping to do a crappy job today at most companies. So, there's a question around, how do you wind up assessing whether a team is high performing, low performing, et cetera. Since this is all based on survey responses, you don't get to actually look at output of teams other than what people self-report. Correct?

Nicole Forsgren: Right, or do you know what also is interesting is occasionally these bands change, and the people are like, "Why did it change? How did it change? This should be a static low, medium, high elite performance category. I need to have a goal to point to because then I can arrive and I could be done." I've had people tell me that, and I'm like, "But that's not how the world works. The industry is changing, the industry is moving. We don't make software today like we made software 20 years ago. Why would that make sense?" And so, I love this question because what we do is we collect data along four key metrics. These have been termed the four key metrics. So, we've been actually collecting this data for six years now, and it's interesting.

ThoughtWorks actually started calling them the four key metrics, and enterprises around the world, across all types of industries have started tracking these and using these as outcome metrics to track their technology transformation. Now, these four metrics fall into two categories, speed metrics and stability metrics. Now, I'm going to come back to these but I'll explain the process really quickly and then we'll come back. What I do every year, so what I said is, I don't just arbitrarily decide this is low performance, this is ... like here's a line. This is medium performance and here's where you are, and this is high performance, and here's where we are.

And then, it's like set it and forget it, and let everyone decide where they are, because the industry changes. So, why would it make sense for me to just make something up and let everyone set themselves according to that? We are very data-driven. We want to see what's happening. What's important is for us to set and collect the metrics that are outcome metrics. So, we use speed and stability. The reason we choose speed stability is because they are system level outcome metrics. We're talking about the DevOps, right? We're talking about pulling together groups with seemingly opposing goals. Developers want to push code as often as possible, which introduces change and possibly instability in the systems.

You have operators, sysadmins, who want to have stability in systems, which means they might want to reject changes. They may want to reject code. So, can we see how does it make sense, Corey? How we may want to have these two metrics, because the goal of an organization is to deliver value, but you also want to have stable systems. So, we want to have both of those metrics in place, right? It's like a yin and a yang. So, we capture both of these, because if you're only pushing code, that doesn't help. But if you only have stable systems, if I only ever say no, then I never get changes. It's not just features, it's things like keeping up with compliance and regulatory changes.

It's keeping up with security updates, keeping up with patches. So, I capture these four metrics, and what I do ... Okay, I'm going to tell you what these four metrics are. My speed metrics are deployment frequency. Okay, so, we'll keep talking about these four metrics. Here are my four metrics. I've got deployment frequency, how often I push code? This is important to developers. It's important to infrastructure engineers, right? I also have lead time for changes. How long does it take me to get code through my system? I measure this is code commit, to code running in production. Now. from the stability point of view, I've got time to restore service. So, how long does it generally take to restore a service?

Anytime I have any type of service incident or a defect that impacts my service users, like unplanned outage, a service impairment, and then I've got change failure rate. That's my fourth metric, might my other stability metric. So, what percentage of changes to production result in any kind of degraded service, anytime it requires someone's attention. So, a service impairment, a service outage, anytime it requires remediation, like a hot fix, or rollback, a fixed forward, a patch. So, what I do is I take, like I mentioned, a very data-driven approach. I take these four metrics, I throw them in the hopper and I see how they group. It's called the cluster analysis because I want to see how they cluster.

And what I have seen for the last six years in a row is that these four metrics cluster in distinct groupings. This year, they fell into four distinct groups. So, you've got a group at the high end, where all four metrics group well, I'll say, where they group well together. And when I say they grew up well together, that means deployment frequency is fast. You're deploying on demand, your lead time for changes is less than a day. Your time to restore a service is less than an hour. Your change fail rate is low, between zero and 15%, so your elite performers are optimizing for all four, right? So, you're going fast and your stability is good. Okay, so I've got a group up there.

Then, I've got a gap. Then, I've got a group. Then, I've got a gap. Then, I've got a group, a cluster. Then, I've got a gap. Then, I've got a cluster. By the way, all of these groups, these clusters, were all statistically significant. They're significantly similar to each other and different from the other groups. So, what that tells me is that speed and stability don't have trade-offs. You don't have to sacrifice speed for stability, or stability for speed. Now, that's not necessarily what we heard for a long time. We used to think that in order to be stable, you had to slow down, but that's not what we see and that's not what we've seen for six years now. The low performance group, their deployment frequency is between once a month and once every six months.

Lead time for changes to get through that pipeline, the same thing, between once a month and once every six months. So, their time to restore service is between once a week and once a month. And then, that change fail rate is in that area between 46% and 60%. Okay, so now, I'm going to get back to a question you just asked me. How can people answer these questions for me when they're survey questions? You'll notice that I'm asking things in ranges. I'm not asking for millisecond response times. I'm asking for things in a scale, in a log scale. People can tell me if I'm deploying on demand, or they can tell me if I'm deploying about once a week, or if I'm deploying about quarterly, or if I'm deploying just a couple of times a year, right?

People can tell me that, or they can tell me when things go down, how long it takes us to restore a service. About a day, about a month. So, what I'm asking in those time increments that go up on a log scale, people can answer those questions. Does that answer?

Corey Quinn: No, that absolutely does. The question that I have then is, when you assimilate all of that and you read this, there's an awful lot of data in here and there's an awful lot that, shall we say, inspires passion in people who are reading it. For example, last year there was a kerfuffle that generally low performing teams tend to outsource an awful lot of technology. This was hotly debated and found to be completely without merit by outsourcing companies.

Nicole Forsgren: By outsourcing companies.

Corey Quinn: Exactly.

Nicole Forsgren: Now, I will say that it was highly correlated and we did make a careful distinction that that was outsourcing by function. And so, what happens there is it's outsourcing if you take an entire batch of something and you throw it over a wall, and you let them disappear for a while, and then, throw it back to you later. So, if you take all of development and you let them go do something and come back later, or if you take all of operations and you throw it away and you never ever see it. That is not what happens if you have a vendor partner that operates with you at the cadence of work, because what often happens then is you have introduced delay. Introducing delay, I love that you brought this up here, what we've seen is, introducing delay can introduce instability.

Because what happens then is when you have delay, it causes and leads to batching up of work. Batching up of work leads to a larger blast radius, a larger blast radius when you finally push to production leads to greater instability. And when you do have that higher likelihood of downtime, that higher likelihood of downtime also means that larger piece of code or something you have pushed makes it harder to debug. So, it's harder to restore a service.

Corey Quinn: You used to be a programmer, as you said at the beginning of this show, so it's always easier to think about what the bug could be in the code that broke the build three minutes ago instead of that code you wrote three weeks ago.

Nicole Forsgren: Yup, exactly. And now, you've got this giant ball of mud that you pushed instead of this nice tiny little tight package that you pushed.

Corey Quinn: Exactly. And this is really, I guess, the point that I'm getting to here, is if people want to read something and then feel bad and not change anything, we have something for that already. It's called Twitter. What impact do you find that these reports have in the world? What changes are companies making based upon these findings?

Nicole Forsgren: So, we've seen huge impact. As I mentioned, we're actually seeing several organizations using these four key metrics as a way to guide their transformation. The nice thing is that it's actually really difficult to fully fully instrument a full metric space platform to capture and correlate metrics that reflect your full instrumentation tool chain. People are like, "Oh, we'll capture system-based metrics." That can be a two to four year journey. Capturing in broad strokes your four key metrics of deployment frequency, lead time for changes, main time to restore, and change fail rate can be at least relatively straight forward.

Nicole Forsgren: You can capture these on a team level to see how well you're doing, and if you're at least generally moving in the right direction. So, that helps. And then, what you can do is you can say, "Okay, what types of things should I be focusing on to improve?" And then, you can identify the capabilities that have generally been shown to help improve, come up with that list. We actually outlined in this year's report like, what types of things ... It's sort of choose your own adventure, right? So, in this year's report we have the performance model, which this is, helping you improve your software delivery performance. And then, we have a productivity model, but start with this model. If this is software delivery performance, and that's what you want to improve, great.

Then work backwards. Which types of things, which capabilities improve it? Start with that list. Once you have that list, no, that does not mean that you start working on every single capability that improves that, because that list is like, after six years of research, that list is 20 or 30 capabilities long. But that's your candidate list. This is the list of all the possible things you could improve. But, you look at that list and you say, "Which things are my biggest problems right now? So, adopt a constraint space approach. What's my biggest constraint? What's my biggest hurdle right now? Pick three or four. Devote resources there. Now, I see resources. That doesn't always mean money, although money is nice. That could be time. That could be attention. That can be anything, right?

Focus there first, spend six months there, and then come back and reevaluate, "Is this still my hardest challenge?" It can be automation. It can be process like, "Am I having a really hard time with WIP limits? Am I having a really hard time breaking my work into small batch sizes? Can I deliver something in a week or less? It could be that.

Corey Quinn: Well, ask any software engineer, "Oh, I can build that in a weekend." You can deliver anything in a week. It's easy. Just ask them.

Nicole Forsgren: But can I do it without burning myself out?

Corey Quinn: Oh, now you're adding constraints.

Nicole Forsgren: I know heaven forbid.

Corey Quinn: You've been doing this for six years. As you look at this year's State of DevOps Report, what new findings, or I guess old findings for that matter, surprised you the most?

Nicole Forsgren: We had a couple. So, an additional thing that we asked this year was about scaling strategies. What types of things are you seeing in your organization to help you scale DevOps? That's a big question I get constantly, how do I scale? What's the best way to scale? A couple of things aren't big surprises, right? Centers of excellence, not great, big bang, not great. Big bangs are used most often by low performers. It doesn't necessarily mean that it's a bad thing, it's just that that's usually only used in the most dire of circumstances. When you really have to wipe slate clean, start over, you need to be most prepared for a longterm transformation. Something that was a bit of a surprise, but also not, can I answer it that way? A surprise, but also not a surprise is that dojo's aren't well used, aren't commonly used among the highest performers.

What we see is that the highest performers, so those that are high performers and elite performers, so the top 43% of our users, focus on structural solutions that build community. So, what does that mean? What that means is that those types of solutions focus on things like building up communities of practice, building up grassroots efforts, and building up proof of concepts, because these types of things will be resilient to re-orgs and product changes. We don't see things like dojos, like training centers and centers of excellence because they require so much investment. They require so many resources. We do see them, but we only see them 9% of the time. When we share this finding with a handful of people, they're shocked because they hear about it so much.

The thing is though, they only hear about it among a handful of cases that have been successful and those successful cases had tons of resources. They had entire buildings set out, they had entire education teams, they had curriculum teams, they had training teams. They also had an amazing PR.

Corey Quinn: Absolutely.

Nicole Forsgren: I think that was something that like at first was surprising, because I'm like, "It's so low. But then I realized I've only heard about it in a couple of cases and it's the cases where they have immense, immense resources.

Corey Quinn: One of the things I always found incredibly valuable about the reports is if you go to conferences and listen to people talk about whatever it is they're talking about, they're doing at their own workplaces, everything sounds amazing and wonderful, and it's all a ridiculous fantasy. Everyone's environment is broken, everyone works in a tire fire and there's not a lot of awareness, I think, in some circles that that's the case. So, whenever someone looks at their own environment and compares it to what they see on stage, it looks terrible. This starts putting data to some of those impressions and I guess contextualizing that in the larger sense. A question that I do have, I don't know if the study gets into this in any significant depth, is it possible for different organizations to simultaneously be high performing and low performing, either along different axis, or in different divisions?

Nicole Forsgren: Oh absolutely, and I'm glad you asked that, and we try to highlight this and we never do a good enough job. We do reiterate it throughout the report. The analysis and the classification for performance profiles is always done at the team level. That's because, particularly in large organizations, team performance is different throughout an organization. As I'm sure you've seen, because when you go to really large organizations, some teams are working at a super fast pace and other teams are at a very, very different place. And so, we always do the analysis at the team level.

Corey Quinn: There's an entire section in the report that talks about cloud computing, which is generally what people tune into this podcast to talk about, and we're not going to talk about it today. We're going to have a second podcast episode about that.

Nicole Forsgren: It's so good though. It's so good. Is this where I get to tell people that like you read, you did a pre-read on the report for me and you're like, "Hey, Nicole, you missed this whole section of nuance that you talk about in one sentence, but you have to expand it because otherwise people are gonna scream at you and I get to thank you for it."

Corey Quinn: I don't think that I framed it quite that way, or if you want to say-

Nicole Forsgren: It's not polite, but it's real.

Corey Quinn: Or, take it the other direction. I practiced that whole statement, "Well idiot," and then went from there. Yeah, you've got to double down on those things.

Nicole Forsgren: By the way, thanks.

Corey Quinn: No, thank you for asking my opinion on this. I'm astonished that anyone cares what I have to say, that it isn't a ridiculous joke or a terrible pun.

Nicole Forsgren: I mean, it's real though.

Corey Quinn: Well thank you so much for taking the time to speak with me today.

Nicole Forsgren: Yeah.

Corey Quinn: There will be another episode.

Nicole Forsgren: Can I get a quick teaser on the cloud stuff though?

Corey Quinn: You may indeed.

Nicole Forsgren: Okay, so cloud's important and it does help you develop and deliver software better, but only if you do it right. You can't just buy a membership to the gym and then not go to the gym and expect to be in amazing shape. That's what we find.

Corey Quinn: Excellent. And I'm sure that the correct answer to solving that problem is to buy the right vendor tool instead.

Nicole Forsgren: Something like that.

Corey Quinn: Yes. So, I will put a link to the report in the show notes so people can download this wonderful work of art/science, I consider it both, and go from there. Thank you. If people care additionally beyond that, of what you have to say and how you say it, where can they find you?

Nicole Forsgren: So, they can find all of DORA's research at cloud.google.com/devops, and if they want to snark on me, I am online at nicolefv.com.

Corey Quinn: Excellent. Nicole, thank you so much for taking the time to speak with me today. I appreciate it.

Nicole Forsgren: Hey, thanks so much.

Corey Quinn: Thank you for listening to screaming in the Cloud. If you've enjoyed this episode, please leave it five stars on iTunes. If you didn't like this episode, please leave it five stars on iTunes. I'm Corey Quinn and this is Screaming in the Cloud.

Announcer: This has been this week's episode of Screaming in the Cloud. You can also find more of Corey at screaminginthecloud.com, or wherever fine snark is sold.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Richard Campbell

Richard Campbell wrote his first line of code in 1977. His career has spanned the computing industry both on the hardware and software sides, development and operations. He was a co-founder of Strangeloop Networks, acquired by Radware in 2013 and was on the board of directors of Telerik which was acquired by Progress Software in 2014. Today he is a consultant and advisor to a number of successful technology firms and is the founder and chairman of Humanitarian Toolbox (www.htbox.org), a public charity that builds open source software for disaster relief. Richard is also the host of two podcasts: .NET Rocks! (www.dotnetrocks.com) the Internet Audio Talkshow for .NET developers and RunAs Radio (www.runasradio.com) which is a weekly show for IT Professionals.

Links Referenced:

  • Twitter Username: richcampbell
  • LinkedIn URL: www.linkedin.com/in/richjcampbell
  • Personal site: https://rcampbell.me/
  • Company site: http://runasradio.com

Transcript

Announcer: Hello and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This week's episode of Screaming in the Cloud is sponsored by LightStep. What is LightStep? Picture monitoring like it was back in 2005 and then run away screaming. We're not using Nagios at scale anymore because monitoring looks like something very different in a modern architecture where you have ephemeral containers spinning up and down for example. How do you know how up your application is in an environment like? That at scale, it's never a question of whether your site is up, but rather a question of how down is it. LightStep lets you answer that question effectively. Discover what other companies including Lyft, Twilio, Box, and Github have already learned. Visit lightstep.com to learn more. My thanks to them for sponsoring this episode of Screaming in the Cloud.

Welcome to Screaming in the Cloud. I'm Corey Quinn. I am joined this week by Richard Campbell, a man who is more enigma than almost anything else. Richard, welcome to the show.

Richard: Thanks for having me on, friend.

Corey: Thank you for being had. Your bio says you've wrote the first line of code in 1977.

Richard: Mm-hmm (Affirmative)

Corey: Your career has spanned the computing industry, both on hardware and software sides, development and operations. You were a co-founder of Strangeloop Networks, acquired by Radware in 2013, which around then was about the time I was playing with one of their load balancers it turns out. You were on the board of Telerik, which was acquired by Progress Software in 2014. Today you're a consultant and an advisor to a number of successful technology firms. You're the founder and chairman of Humanitarian Toolbox, a public charity building, open source software for disaster relief. But the way I found you is that you're the host of two different podcasts .Net Rocks!, which I'd never listened to and I would argue the point, and the one that I encountered you with, RunAsRadio.

Richard: Right.

Corey: Which is a weekly show for IT professionals. Stop me if you heard this one. You're a Microsoft MVP and you are more or less the... It seems like a strange version of me who, let's admit it, you're not quite as well dressed, but that's okay.

Richard: That's true.

Corey: I forgive you, but over in the Microsoft ecosystem rather than the Amazonian one.

Richard: Yeah. I do grow a better beard than you though.

Corey: Yes. I still look like an angry 14 year old trying to prove a point to mommy and daddy and-

Richard: Well the show is called Screaming in the Cloud. I think you're supposed to be angry.

Corey: Oh, absolutely. I do my best. But I encountered you a couple years back when I attended my first Microsoft Build, when they sent me an invitation to go and record an episode and I thought, "Wow, that's amazing. They've gotten me confused with someone who knows what they're talking about. Awesome. I'm going to act as if, because as a mediocre white man, that works really well for people who look like me."

Richard: Right. Sooner or later they're going to figure it out, but if I move fast enough, maybe I can get the show in.

Corey: Exactly. And it worked and I had Corey Sanders, at the time the corporate VP of Compute. Now he's ascended somewhere else and we had him on again recently at Azure... at Microsoft Build this year. The interesting part though was that this year you reached out and invited me up to record a bunch of shows and people who've been following this on a week to week basis have noticed since then. "Wow. There've been a lot of Microsoft guests there." Yeah, that was recorded during a three day span when I'm up there trying to I guess maintain focus after the seventh podcast recording in a row.

Richard: You can get punchy after a while when you knock them out back to back like that.

Corey: Absolutely. So I did a little more digging into you and it turns out that as of the time of this recording, you and I have recorded a few days ago, an episode of RunAsRadio, your show, and you mentioned that that would be episode 650.

Richard: Right.

Corey: So I'm not sure on the scheduling, this may come out before that one. But what's fascinating to me is you've been doing this about 10 times longer than I have, so, "Oh, this is what I'll be if I grow up."

Richard: I'm still trying to figure out what I'm going to be when I grow up. All I know for sure is I'm pretty persistent.

Corey: Yeah, and that that seems to work out really well, and I've looked through some of your back catalog of various episodes you've had. We've had some people in common. You have a list of people that are high on my list of folks to wind up dragging and interrogating, I mean interviewing. And it's just been... It's really neat looking at what someone can turn this into. I don't view this as a deviation from the typical theme of this show, which has been the business of cloud computing because you very much have a business that is running around the world of cloud. I'm intentionally vague. I don't say a technology business. I say a business. So you've turned this into a RunAsRadio company around this. You have a very streamlined production pipeline that puts mine to shame. It's just fantastic watching you work.

Richard: Well it's been 12 years. So we did migrate to the cloud, you know? The original incarnation, you're talking about the early days of podcasting. Carl Franklin, my cohort on .Net Rocks!, he recorded his first show in 2002. The word "podcast" didn't even exist yet. And I started out run as of 2007 so the word now existed but it was still very early days. Of course I started an IT podcast that was Microsoft centric right after Vista shipped because I'm clever like that. And so you know, those were tough days the those first couple of years.

But the machine already worked pretty well and it was not hard to just keep every Wednesday, you know, since April 11th, 2007 make sure a show publishes, just keep doing it and the numbers sort of pile up. It's way more fun now that the cloud has grown up and become a thing and we're still exploring what it can do. That's a nontrivial area of conversation, you know, week over week along with topics like DevOps and still, you know, dealing with the reality of do you run a mail server, do you own a domain controller? Like all of those kinds of problems that IT folks have and are just figuring out the right ways to do things.

Corey: Right. Whereas I come from the Linux and startup world here in San Francisco where our entire approach to development is laughing derisively at anything that's more than six weeks old and throw everything away. When everything's greenfield, you can build amazing stuff that never looks bad because you won't be running it long enough for it to become legacy.

Richard: Yeah. By that point, you've pivoted.

Corey: Exactly.

Richard: Where I think, yeah, a non-trivial chunk of my listeners are folks that are still paying the price for decisions made in the early aughts.

Corey: Isn't that the truth? When people say legacy, I hear it makes money and, again, I'm a three year old company. There's only a handful of us here now, but everything we've built is largely serverless and I look back and I still see that I have piles of technical debt. "Ooh, I would do that differently. There's a better way to have done that. Oh, my account management was crappy, et cetera, et cetera, et cetera." But it doesn't actually solve any business problem for me. If I fix that, it just assuages my sense of something not being right in the world.

Richard: It's also a balancing act on those things as well, right? You have it pivot moments like I'm about to renew this contract for the next two years or five years. This is the moment to make a decision like how much will I regret this when that contract expires that I did this again? Shouldn't I make a move to the code or do a rearchitecture, retire or a piece of technical debt? There's only certain moments when it makes sense to actually pursue that. They slowly cost you, right? They are hammers in that sack of hammers you're carrying around and, you know, the chance you get to shake a few off. It's pretty compelling.

Corey: It really is. But what's also interesting to me about you is that you and I have both gone to the same direction where to some extent we've stopped writing code day to day, 40 hours a week and started doing other things. And the podcast is fascinating. I view you as someone I can reach out to from time to time and get good advice and figured it's probably time to wind up putting you on the show as well. But from my perspective, when I started this, it was, "Well, what do I do with the podcast? Like how do I structure this show in a way that makes sense for people?" I know I wanted an interview show for starters for two reasons. The aspirational of trying to help other people tell stories and effectively give them exposure to a new audience. And the real answer, which was, "I'm a nobody who no one should ever listen to seriously." But when you say, "Would you like to be on my podcast?" People say, "absolutely" instead of, "Who are you? Get out of my office."

Richard: How'd you get in here?

Corey: Exactly.

Richard: I also... I've discovered how powerful it is just as a research vehicle. You have these amazing conversations where ultimately you're... You represent the audience while just helping to try and understand what's going on, what they're talking about, and if you can clarify that, you'd probably clarifying it for the listener as well. But you ended up in this sort of state, this gestalt state, this, "I have a sense of a view over this platform because I've talked to all these different people and then gotten feedback from my listeners." And so I sit in a sort of interesting analytical space where it's like, these things seem to be important for the companies that are selling them, but not that important for the users that don't really care about them.

Corey: Right. And there's so much coming out across this entire environment, sort of cloud ecosystem, that no one can keep up with it all by a landslide. Picking and choosing different aspects of it. And when you get someone in who's well connected in the environment and being able to ask them what parts of this matter versus what parts of this can be safely ignored or even not even framing it that way. It's cool. Like when I had Corey Sanders back on or Scott Guthrie, it was great to be able to talk to them about the entire Azure ecosystem, something I don't know much about, and see what areas they focus on. Because there's so much they could talk about, but instead they pick one or two areas that they're excited about, that they want to tell a story around. And that helps contextualized and shape how I start to view that ecosystem as a result.

Richard: Yeah, absolutely. And I think for your listeners as well, just to be more aware of this larger cloud world we're living in and say like, "Who's moving where?" Because it certainly... Amazon and Microsoft are paying close attention to each other.

Corey: Oh, absolutely. I tend to feel like I personify my listener reasonably well because, for better or worse, through most of my career, I have always been usually the first person in any given room to stand up and say, "I don't get it. There's something that I'm missing. There's something that is unclear to me and it's just not clicking." And invariably, whenever I do that, other people chime in. "Yeah, me too." I did an-

Richard: Right. That sort of relief that grows across the room. Oh, it's not just me who's like, "This makes no sense."

Corey: Right. Today we're recording this on July 25th of 2019 and late last night, right before I went to bed, I tweeted that I just spent a few hours wrestling with S3, trying to get URL redirects working, gave up rage quit, did it in 10 minutes in Nginx on an EC2 Instance, and went to bed. And the replies to that that I woke up to, massive, fall into two camps. One of them is, "Well, have you considered..." And then it gives some Archaean, Byzantine serverless solution. It's no because I have a job.

Richard: Right.

Corey: And the other is, "Oh, thank God. It's not just me. This stuff is super complicated." And that's why I mention, "This makes no sense to me and I don't get it" as often as I do because when I do that it reminds people that they're not alone. So if I don't understand a lot of things about a product or an ecosystem, and I call that out, I feel like in many ways I'm speaking for the users when I do that.

Richard: Well, I mean particularly in this space, you're often working alone anyway. There's lots of people depending on you, but you are the "expert", you know? You're handling our cloud migration, right? And you're like, "Yeah." Or when you get that pop up dialog that says, "This didn't work. Contact your administrator for more help." And like I'm the administrator and I have no idea what you're talking about. So I think having a show, these kinds of conversations, because you're often alone, it's like this is your one connection that reminds you you're not alone. That there are other people struggling with this, that this is a normal part of us building a new part of civilization, which I mean I'm not exaggerating. I say the cloud is reshaping what humanity is going to look like and our small roles in it, whether you're building it, utilizing it, or talking about it all are part of moving that ball forward.

Corey: One of the things that I guess astonishes me is just how well this has resonated. I thought this was going to be something I do for a couple months and then shut down and here we are well over a year later. Well not shutting it down now, but what's strange is you know, no one ever knows their own reputation. What you mentioned to me is that getting guests for the podcast initially when you were working behind the scenes where I knew you were doing that was challenging based upon the name of the show.

Richard: Yeah. Well you know, first impressions and visceral responses, but also non technical people. I mean when you're talking about working around an organization in a conference as large as Build, there's a lot of people who have says in different things and so they come at it from different perspectives. And if there's PR folks involved, well they care about the optics more than anything else and and they're going to go on the simplest, easiest response. They got to triage hundreds of things and what's the quickest way to do it? It doesn't make it accurate.

Corey: Right. The show is called Screaming in the Cloud. Obviously it's a ranty thing that's looking an awful lot like you just sit there and berate someone for half an hour.

Richard: And unless you actually listen to the show, in which case you know, that's not even remotely true.

Corey: Right. But when I started this, I didn't know what it was going to turn into and I learned a lot as I did this in the first dozen episodes or so. For example, I thought this was going to be sarcastic, snarky, and bantering. It turns out when there's an interview show and you have a different guest on every week, you can't do that in the same way because if you and I banter and we're snarky and we're sarcastic and you smack it back every time I toss something your way, it's a great show.

Richard: Sure.

Corey: But the next week I have a different guest where they are not as good at this and it looks instead to all the world like I'm beating them up. Someone listens to that, no one's going to come out of that thinking, "Oh, I'm going to be on that show." Hell no.

Richard: And it's not fun to listen to either if someone's confused or struggling. You know, we have the advantage that we've known each other for a couple of years now. I mean not every day, but they got a couple of events, you know, and so I'm pretty sure I know your snark when I hear it.

Corey: And we both do podcasts, so we definitionally have the ongoing love affair with the sounds of our own voices.

Richard: Well, I do sound amazing.

Corey: Yes, I have a face for radio. It works out well, but what's fun too is that as I was looking down this path, I realized about, at the time of this recording, a couple months back now that I wanted to have a podcast where I could be snarky and sarcastic. So I launched another one, the AWS Morning Brief and that's now about two months in where it effectively right now is on Monday mornings. It's a 15 minute show and it just talks about the releases that AWS has done over the past week like I do in the newsletter. And then I make fun of them as I do in the newsletter, but I often go into more depth. I talk about why it matters. I mention things that don't fit in the newsletter for whatever reason. I go off on tangents from time to time and I'm snarky and I'm fun and it's sarcastic and I don't have a guest so I don't need to worry about making anyone look bad.

Richard: Yeah.

Corey: There are plans for You Heard it Here First to look at expanding this to a multiple day a week show.

Richard: That's a lot of commitment, friend. Like because all that stuff is super timely.

Corey: It is. But that's the trick. I don't think I could do a daily show, for example, about what they released yesterday because, spoiler, no one really cares. People want to get a roundup, but they don't need to have it as breaking news. So there need to be other things to do, like talking about services, doing some sort of deep dive or whatnot, talking about things that are relevant to the ecosystem. Because frankly, just talking about releases, I get bored, let alone other people. I can't imagine what that would look like. You've got to keep it fresh. You got to keep it interesting. You got to keep it snarky.

Richard: Yeah. Analysis of an outage is always interesting. Like what really happened with Cloudflare the other day. You know, I know we've been making fun of their mistake, which is lovely, you know, because-

Corey: I have such little tolerance for that. Yeah. It turns out all this stuff is super hard. Anyone who's sitting there saying, "Oh, you went down because you employ morons who are terrible at these things." Shut up and go away.

Richard: It's not even close to true.

Corey: Yeah, it's very clear that whoever's saying that has never had to run anything of any reasonable sense of scale or importance. So if you've... Like one of the questions I always liked to ask people when I was an ops manager in the interview was, "Tell me about an outage that you were at least partially responsible for." And is always is partially responsible for. And people's reaction to that is telling. I don't want to see people beating themselves up, but I want to see people telling a story about it and if the answer is, "Oh, I've never taken production down."

Well there are a few things that are possible, none of them are compelling. First, either you're lying because you think it's an interview and you're never supposed to have broken anything, which cool, we're done here. Or no one has ever trusted you with production, which, okay, that's interesting as a data point. Or you are so methodical and so careful that you need to quadruple check everything. We're at launching space shuttle style of software rather than something that's web facing. And that's also useful to know and we'll uncover that as we have the further conversation. But-

Richard: Yeah. You don't even know if the space shuttle is a good comparison because they didn't... That didn't always work out well for them either.

Corey: Oh yeah. It was great. I got to tell that story once from an outage that happened in a team I was running. Someone on our team wound up doing a misconfiguration that led to something else and an edge case hit it. That's always how these outages tend to happen these days. And the discussion was "Okay, so that misconfiguration really... Like any one of those things not happening would have meant no issue. But that misconfiguration..." And the interviewer leaned in and said, "Was it you?" And my response was, "It was the team I was managing. So yes, of course it was my responsibility, but who actually made the mistake completely irrelevant and it could have been any one of us." It's the truth. You don't blame people for individual mistakes. The one that irritated the living hell out of me was the S3 apocalypse a few years back where... "Oh, it was a typo on the command line."

Corey: No, that effectively meant someone stepped on a landmine that many other people had a contributory effect in burying. If one wrong command typed into the wrong place can take down all of S3 for an afternoon, well that is a process and systems failure and it's a learning experience. I've never figured out who it was and the only reason I'd want to is so that I could call them and say, "You know what? We've all been there. Can I buy you a drink?" Because man there but for the grace of God, any one of us could have had something like that happen and it's certainly not their fault and I just want to validate that there was never any direct consequences to that person. I don't imagine there would have been, but-

Richard: No, not that you would presume in these big organizations, they're mature enough now to know that this was a failure of process to even be put into that situation where it's possible to do that.

Corey: Yes. Andy Mai-Lan Tomsen Bukovec, the VP of S3, was on this show very beginning of this year and we touched on that topic as well. But before the system had been restored, there had already been changes pushed out that would've prevented that mistake from ever happening again. They've also gotten onstage since then, various folks from her organization, and talked about how it's now running, I think on 235 microservices that drive S3. They've completely rearchitected a lot of it under the hood and we're seeing performance and stability enhancements come out of it and contextualizing that's important, but telling the human story around these things is always interesting because a root cause analysis...

I feel like companies are sort of becoming victims of a situation they've created for themselves where marketing and PR only ever released very carefully crafted statements so people are used to having to sort through those for subtext but root cause analysis or cause of error reports or whatever term we want to use so J. Paul Reed doesn't yell at me are all very... Those tend to be very honest and... But try to dissect that into the, "who screwed up?" Is counter productive and unhelpful.

Richard: Yeah, no understanding how we can get into a situation like that and how we're not going to get into it again is the real root cause analysis. That we get to a place where now we... Analysis isn't good enough if there's no action to take at the end, right? The whole point is to get to a place that says, "If we take this action, we decrease the likelihood of these kinds of problems."

Corey: Yeah. So I'm a consultant these days. I fixed the horrifying AWS bill. I opine on architecture a lot, but I don't really run production infrastructure anymore. So one of the things that I worry about, and I'm curious how you've addressed it, has been now that I spend the bulk of my time, at least as far as the public sees it, of doing podcasts, writing newsletters, and being obnoxious on Twitter, how do you avoid losing touch with the technology to the point where you just become a talking head who hasn't ever used the thing that they are taste making about?

Richard: Yeah, I do think you have to use the thing, right? And it may not be large scale websites. It's your own stuff, but that's fine. You know, actually spending time in the tool matters a lot. I laugh because of course doing RunAs now for 12 years, I continued to run my own exchange server. I was allowed to. I have the license for it. It makes no sense. It should be in the cloud, but I can also look an exchange admin in the eye and say, "Dude, I feel your pain." Right? Like running mail services is a thankless job. Well most of IT is thankless, right? You can only get a C or an F. If it works, nobody can tell. And if it doesn't work, well you're an idiot. There's nothing... There's no win. It's just you made it through another day. Have a beer. Keeping the infrastructure up and running for the show is part of that.

What's the right way to do this? When do we improve it? How do you think about the sort of next stages on that? At the same time you're having conversations with the people that run these larger installations and, you know, pursuing this case studies. In the early days of Azure, I stopped talking about it on the show because the only people I could get to talk about it were the people building it and the folks that were talking about the people who are building it. I just hit a point where I'm like, "Until you have a case study, until there's somebody running it that it's meaningful, I don't think this is worth talking about anymore." Because that's, I think, what people actually wanted to hear. "Talk to me about a regular group of humans who use this tool successfully and what they liked and what they didn't like." That to me is a useful conversation. All too often you're caught in the... You can be stuck in the bubble of the people who want to promote it more than the people that actually are using it.

Corey: One of the things that astonished me as I did this because I started off building all the infrastructure myself because, "Hey, I run infrastructure. It's what I do." The newsletter and the podcast and the rest have gone from toy projects to viable businesses and at that point I felt a sense of responsibility to stop keeping this all with this Archaean architecture I built in my head. I migrated all the web properties to WordPress. WP Engine hosts that. I'm giving a conference talk about moving it off of serverless and onto WordPress over at Serverlessconf in New York later this year called Benjamin Buttoning Serverless. And I expect to mostly be shouted off the stage because it's something that flies against common wisdom and true believers. Which, okay, I get, but by the same token, I want to make sure that we're actually doing things that are right for the business.

What astonishes me is I... Whenever I find myself talking to analysts, they look at me and they think I'm a kindred spirit because I talk to the people building these things. Yes, I do. Then I talk to customers using them. Yes, I do. And then I build something with it myself to validate what I've heard. And invariably when I tell that to an analyst, I get a flicker of terror in their eyes as if somehow-

Richard: What would you do that?

Corey: Well not even that. About am I about to be the kid that points out that the Emperor has no clothes? Because you have entire analyst firms that never touch this stuff that are very good at talking about it. And I never want to be that.

Richard: No, and I think it's a conscious choice that you also make in your work, right? I mean, the good news is, especially when we're talking about the cloud space, is that we're all using it to some degree. It's just how aware are you and how responsible are you for it? It'd be way harder to be in a space that was much narrower, more specific where you really wouldn't have any excuses to utilize it. But I think we also make a choice, right? It's interesting that you pulled away from a serverless architecture over to WordPress Solutions. Sounds like very practical engineering and I think people-

Corey: Oh, absolutely. I could hire someone to handle all this stuff for me in the development side. I can pay a company probably too much money to host this all for me and I wouldn't have to think about it or touch it. Whereas finding someone who can do that same sort of work on the bleeding edge of serverless stuff, well, okay, that's not a basket full of money. That's a dump truck full of diamonds. So it gets incredibly expensive and for what business value? Because again, "Well you'll have better availability and utilization, et cetera, et cetera, et cetera." Yes, true. However-

Richard: Will I? If I actually have to take care of it myself because I'm awfully busy.

Corey: Right. My time is worth more than the infrastructure cost at this point because it's the opportunity cost of not doing something that adds business value.

Richard: Well I think you get more uptime out of having... You have people who can take care of it, even if it's a bit more fragile. Like the people factors on all these things matter. The availability of resource matters, right? You have to be able to scale and it's not just, you know, scaling the site, it's scaling... Having people around that you can trust to make sure those things keep working.

Corey: Right. And if WP engine falls over, they have an entire team of people who are paid to do nothing other than make sure that that doesn't happen and they're going to get it up and running and notice it way more effectively than I am. Yeah.

Richard: Long before you do. And that's all the point of the cloud is we're concentrating the best and brightest around these infrastructures. So they do-

Corey: And my use case isn't what people tend to... I think that people lose sight of that. If I'm an ad network, for example, people are not going to come back if my site is down and try and view an ad later. Every second there's down, there's a very real revenue cost of that. With last week and aws.com, which has my blog, the podcast, the newsletter stuff, if that's down, someone's not going to be able to sign up for during the duration of that happening. But it doesn't have a direct revenue impact on me. I mean, I'm still averaging a hundred signups a week on the newsletter, so if I'm down for 10 minutes or so and I wind up losing at a high side 12 of those, it's not good. But it's certainly not the end of the world economically.

Richard: Yeah. It's also a challenge to measure that too, right? You don't have a-

Corey: Right.

Richard: Equation to that. Yeah, will those 12 that couldn't get there come back later? Like those are hard questions to answer, but it's all still the same basic point, which is, you know, the value prop here for the amount of effort necessary and costs necessary is high enough that this is the best way to do it.

Corey: It really is.

Richard: There's no ideal logs, right? I mean, if you really cared to optimize everything, you'd still be handcrafting C++ and CGI calls for this, right? We don't care that much. We don't need, you know, the difference between a second and a 10th of a second. It's just not that big in this scenario.

Corey: Exactly. And at some point, Charity Majors had a great line that she's said a few times that "Nines don't matter if users are unhappy."

Richard: Yeah.

Corey: And that works in the direction of the experience needs to be good. Just technically saying you're up doesn't count. But by and large, if people don't care that my site is down at any given point in time, then is it really a problem?

Richard: Oh yeah. Right up until the cashflow stops, you know?

Corey: Well, there's that. Then I have a problem. So as we take a look back over the sweeping landscape we just covered, we've talked a lot about where we've come from and how we built these things. What's the future look like for you?

Richard: Well I you like... The podcast model, I think, appeals because it isn't your principle focus. I don't think a lot of people sit down and just listen to a podcast. They tend to listen in their car on the way to work or their transport of choice or when they're working out and so forth. So you know, I'm big on the in the ears model. I think it's very interesting. The question is always are we telling stories that people care about? Are we staying focused on those kinds of things? And are we cross pollinating? You know, here we are crossing the streams, you and I, between my show and your show and certainly the same thing we were doing at Build with all those podcasts.

We're all way more alike than we're different. And the more times that we can sort of mix that up or remind each other that everybody's working hard, trying to do the right things and that all of these technology choices are viable, right? There is no one right way. There's what you know, how well you can manage those costs, and how much value you get from it. I think the more times we understand that with everyone in this particular circuit, the happier we're going to be.

Corey: I think you're absolutely right. The very honest truth, one of the big reasons I have a podcast that does interviews, is every week you get to borrow someone else's audience. What always drove me nuts is when I would guest on podcasts and they wouldn't promote it. It's... What's the point of having me on your show if you're not going to ask me to tweet about it or leverage my audience? Because that's the entire point.

Richard: Yup. That is absolutely the machine and more sources of good information, people want to know about them. They're valuable. So it's certainly, it's worth putting energy into putting that out in the world. We still have Engineer's Disease, right? Which is, "If I make it, that should be sufficient. If you don't value it, well that's on you." But that's not the truth, right? There's plenty of great things that have been made over the years that nobody got to know about because it didn't put the effort in to making it visible.

Corey: I think you're absolutely right on that and I think that's probably a good place to leave it.

Richard: For sure.

Corey: Richard, thank you so much for taking the time to speak with me today. Where can people find out more?

Richard: Well, for this particular audience, come and listen to RunAsRadio, which is exactly the domain name as well. And that website is filled with CIS admin jokes. So if you take a look around, you'll see little things that if you're a CIS admin, you think that's funny. And we put out a show every Wednesday since 2007 so you know, you can count on us to be there. And otherwise you can reach me on Twitter. I'm Rich Campbell on Twitter and always happy to chat and snark and have some fun with this great career that we're in.

Corey: I'll be sure to throw a link to RunAsRadio.com in the show notes, but that link will be largely symbolic.

Richard: That's how all of them --

Corey: Yeah. Welcome to the CIS admin joke pool. It doesn't get better from here. Richard Campbell, Microsoft MVP, founder of several podcasts, RunAsRadio.com, and raconteur for lack of a better term.

Richard: Nice.

Corey: I'm Corey Quinn. This is Screaming in the Cloud.

Announcer: This has been this week's episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com or wherever fine snark is sold.

Announcer: This has been a HumblePod Production. Stay humble.

View Details

About Mike Warren

Mike Warren is cofounder and CTO of Descartes Labs. Mike’s past work spans a wide range of disciplines, with the recurring theme of developing and applying advanced software and computing technology to understand the physical and virtual world. He was a scientist at Los Alamos National Laboratory for 25 years, and also worked as a Senior Software Engineer at Sandpiper Networks/Digital Island. His work has been recognized on multiple occasions, including the Gordon Bell prize for outstanding achievement in high-performance computing. He has degrees in Physics and Engineering & Applied Science from Caltech, and he received a PhD in Physics from the University of California, Santa Barbara.

Links Referenced

  • @m_warren
  • https://www.linkedin.com/in/mike-warren-3a0439b1/
  • https://lightstep.com/

Transcript

Announcer: Hello and welcome to Screaming in the Cloud with your host cloud economist's, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world and ridiculous titles for which Cory refuses to apologize. This is Screaming in the Cloud.

Corey Quinn: This week's episode of Screaming in the Cloud is sponsored by LightStep. What is LightStep? Picture monitoring like it was back in 2005 and then run away screaming. We're not using Nagios at scale anymore because monitoring looks like something very different in a modern architecture where you have ephemeral containers spinning up and down for example. How do you know how up your application is in an environment like that? At scale it's never a question of whether your site is up, but rather a question of how down is it. LightStep lets you answer that question effectively. Discover what other companies including Lyft, Twilio, Box and Github have already learned. Visit lightstep.com to learn more.

Corey Quinn: My thanks to them for sponsoring this episode of Screaming in the Cloud. Welcome to Screaming in the Cloud, I'm Corey Quinn. I'm joined this week by Mike Warren Co-Founder and CTO of Descartes Labs. Welcome to the show, Mike.

Mike Warren: Thanks Corey, happy to be here.

Corey Quinn: So you have a fascinating story in that you had a 25 year career at the Los Alamos National Lab and then joined, not joined, started a company Descartes Labs about four years ago now.

Mike Warren: That's right. I kind of considered, you know, it was a 30 year long education that provided me kind of the right skills to go out and start a company that used computing and science to help our customers understand the world.

Corey Quinn: And it certainly seems that you've done some interesting things with it lately. At the time of this recording, about a month or so ago, you wound up making I guess a bit of a splash in the cloud computing world by building a effective supercomputer on top of AWS and it qualified, I think, what was it, place 136 on the top 500 list of the most powerful supercomputers in the world. And that's impressive in its own right, but what's fascinating about this is you didn't set out to do this with an enormous project behind you, you didn't decided to do this with a grant from someone. You did it with a corporate credit card.

Mike Warren: That's right. I guess it wasn't such a surprise to us. We knew the capability was coming and it just happened to be the time was right and I had some time to compile a HPL LINPACK and run it. But that's the way the computing industry is going. And anything you can do in a data center or in a super computing center is eventually going to be possible in the cloud.

Corey Quinn: So one of the parts of the story that resonated the most with me was the fact that you did this using Spot instances, but you didn't tell anyone at AWS that you were doing this until after it was already done. And I confirmed that myself by reaching out to a friend who's reasonably well-placed in the computer org. I said, "Wow, congratulations. You just hosted something that would wind up counting as one of the top 500 super computers in the world," and his response was, "Wait, we what now?" Which was fascinating to see where it used to be that this was the sort of thing that would require such a tremendous amount of coordinated effort between different people, different stakeholders across a wide organization.

And it turns out that because it's neat, wouldn't have been a sufficient justification to do approximately any of it. And now apparently you can do this for 5,000 bucks, which is less than a typical server tends to cost, and about what a typical engineer embezzles in office supplies in a given year. And the economics of that are staggering. How long did you plan on doing this before deciding just spark it up?

Mike Warren: Not Long. And it's really why the cloud is so attractive for businesses like ours. You don't have to wait for resources, you don't have to coordinate, you don't have to spend time on the phone, you don't have to worry about support contracts and all of this. You know, even a top 500 run on a supercomputer requires an enormous amount of coordination. You've got to kick all the other users off. You know, it's a big deal and it may only be done once on one of these nation scale supercomputers, but this on AWS, any given day when there's a thousand nodes available, you can pull those up and run at a petaflop.

Corey Quinn: Which is astonishing. I mean, you did some back of the envelope calculations on what it would cost to do this in hardware. And the interesting part to me wasn't the dollar figure, which, what was that again?

Mike Warren: We figured, you know, $20 to $30 million in hardware to get to this near to petaflop level.

Corey Quinn: Yeah. And the amazing part for me was that looking at this, you also mentioned in that article, and I'll throw a link to that in the show notes, but the fact that it would also take your estimation six to 12 months just to procure the hardware and get it set up and get everything ready just to do this.

Mike Warren: Yeah, I think people don't realize the sort of unmentioned overheads in all of these HPC sort of applications. You know, you've got to procure the hardware and make sure you have the budget authority and get the right signatures, and that's after, assuming you have a data center or a building and enough cooling and power and all those sorts of infrastructure.

Corey Quinn: So one thing that rang a little bit, I guess we'll call it false, in the narrative and not to poke holes in your legend has been that we didn't tell AWS that we were going to be doing any of this, we just decided to go ahead and run it. But anyone who spent more than 20 minutes swearing at the AWS console knows that, okay, you start up an account and the first thing that you need to wind up doing is request a whole bunch of service limit increases while you're already running one of those instances in that region. Why do you need a second one? Justify it. And they're pretty good about granting those service limit increases. But I'm curious to know, did you already have sufficient limits to run this from other workloads that have been done in that account or did you wind up opening the most bizarre limit increase service ticket that they've probably seen in a month?

Mike Warren: As I recall, we had about half the resources in US east where we ran this. So we had kind of this scale of resources spread across different zones, but there's also different quotas between Spot and not Spot. So, there was a specific request to up the quota for Spot in US east to this level. And there's another interesting API limitation, which is a single call to allocate nodes for Spot can only allocate a thousand at a time. So it's a little extra workflow. You got to break it into two parts and there's a throttling on the API calls, so you can't immediately allocate, say 1,200 nodes, you've got to allocate a thousand, wait a minute, and then allocate the remainder.

Corey Quinn: Take us through this from the beginning where effectively... what is the architecture of this look like? I'm not entirely sure what the supercomputer in question was chomping on, so in my mental model, I'm going to figure that you were just mining Bitcoin, which it turns out is super financially viable in the cloud, you just need to do it in someone else's account. But on a more serious note, what was the, you mentioned LINPACK, what does that do?

Mike Warren: Well, an important distinction to make is how much communication needs to happen among the processors. So mining Bitcoin, the processor doesn't need to know about anything else going on in the world, it's completely independent. So you can start up a thousand different nodes mining Bitcoin and all of the clouds do that very well. The top 500 benchmark is based on the inversion of a very large matrix and it uses a piece of software called the LINPACK. And in the solution of that problem, all of these processors have to talk to each other and exchange data and they have to do that with very low latency. So, you know, starting up a thousand nodes and if any one of those fails in terms of computation or network communication during that process, the whole thing falls over. So these tightly coupled types of HPC applications that are typically written in something called MPI, which is the message passing interface, are a lot more challenging to do in the cloud than the task parallel sorts of applications like a typical web server.

Corey Quinn: One of the things that fascinates me is that whenever you're using any sufficiently large number of computers, some of them are intrinsically going to break, fail out of the cluster, et cetera. That's the nature of things. That sort of goes double when you're using something like on Google's pre-emptable instances or an AWS Spot Fleets where it turns out that subject to the available supply, things are generally not nearly as available as you would expect if, for example, someone else in that region decides to spin up a supercomputer to take a spot on the top 500 list. So how does each one of those things checkpoint what it's working on in some sort of central location so that it can die and be replaced by something else without losing all the work that node has done, or it doesn't it?

Mike Warren: No it can't. There's no check-pointing these sorts of problems in any easy manner. So everything's got to work perfectly for the three or six hours that this benchmark is running. So the reliability is very important. And I've heard anecdotal reports of, you know, some of these very fast top 10 supercomputers needing to try to run the benchmark several times before they can get it to run reliably.

Corey Quinn: So when you wound up doing this, did you just over-provision by a bit and assume you would lose some number of nodes along the way?

Mike Warren: That also doesn't work. Once you've sort of labeled each of these processors with a number, you can't have one disappear. It's got part of the state of the problem in it. So if you start with 1,200 nodes, you need them all to work the whole time. So that's why these tightly coupled applications are a lot more challenging to scale up.

Corey Quinn: So when you wound up doing this, you effectively wound up with how many instances that were a part of this? I mean I saw the statistics on on petaflops, but that's hard for me to put into something that I can think of.

Mike Warren: Yeah, it was a bit under 1,200 nodes and these were a 36 core processors, or rather a 18 core processors with two dies per node. So that gives you a 41,000 odd processors and these are the hardware cores, not the abomination of virtual counting.

Corey Quinn: Gotcha. Yeah, it looks like you did this entirely on top of C5s, which is... It's the performance of those things is impressive and the economics of it are fantastic. For me, I think the hardest part to wrap my head around is the fact that you had 1,200 of these things running without a hiccup for three straight hours on Spot, which historically was always extremely interrupt prone. And it's surprising to me that none of those wound up getting reclaimed during that window. Just at that scale, I would expect even traditional computers, one of the 1,200 is going to fall over and crash into the sea because I'm lucky like that.

Mike Warren: Well there's, I think two things you're talking about there. One of those is solved by Amazon's Spot blocks. So you know, you say you want a certain number of processors for some number of hours between one and six and then Amazon guarantees it's not going to optionally take one of those away from you. So those are a bit more expensive than the normal Spot, but a lot less expensive than the non Spot instances.

So what becomes important is just the inherent failure rate of the hardware and the network. And in our experience the cloud resources we've used have been incredibly reliable to the point where, you know, certainly we've seen in Google their predictive task migration where they can sort of understand that a node is about to fail and then migrate that whole kernel and processes to another piece of hardware so that you never know about it. And they can pull that defective hardware out and and fix it.

Corey Quinn: The idea of being able to do something like this almost on a lark on some given afternoon for what generally falls well within any arbitrary employees corporate spending approval limit is just mind boggling to me. I know I can keep belaboring this, but it's one of those things that just tends to resonate an awful lot. Can you talk to me at all about how this winds up manifesting compared to, I guess the stuff you do day to day? I'm going to guess that the Descartes Labs doesn't have a lot of interesting stuff. I mean you folks work on satellite and aerial imagery to my understanding, but how does that tie back to effectively step one, take a giant supercomputer, we'll figure out step two later?

Mike Warren: Well, it's really democratizing super computing and it's an extension of the power that software gives an individual. You know, we're able to do things now with good software and a smart person that used to take an entire group of people with lots of infrastructure to do. So, it's always been something I've been very interested in and goes back to our building Beowulf clusters back in the 90s. That was really democratizing parallel computing for people who couldn't afford the state of the art super computers. And there were untold times that, you know, people collared me in and told them stories about their group in college who had a cluster that they built in the closet that allowed them to do the research that they were doing.

Corey Quinn: So on a day to day basis, what does I guess your computing environment look like? I mean, you spun out of a national lab, which for starters we know is going to be something that is relatively, shall we say computationally impressive. Mike Julian, my business partner used to help run the Titan Supercomputer at Oak Ridge National Lab about 10 years ago. So I've gotten absorbed into it by osmosis almost. And I always view that as here goes smart people, I just sit here and make sarcastic comments on the side. But what does that look like on a day to day basis of what Descartes Labs actually does?

Mike Warren: Well, the last big computation I did before we founded Descartes Labs was a on the Titan machine. We had 80 million CPU hours to calculate the evolution of the universe and we did that with a trillion particles. So a lot of that kind of thinking has carried over to our environment at Descartes Labs, and I spend a lot of my time in the command line writing Python scripts, but our interaction with with customers is much more focused around APIs and making it very easy to interact with these peta scale data sets. You tell us where on the earth and what time over the last 20 years you want to see an image and we can deliver that to you in less than a second.

Corey Quinn: And Are you building this entirely on top of AWS? Are you effectively for all comers as far as large cloud goes, is there a significant on-prem component?

Mike Warren: There's no on-prem at all. Descartes Labs, IT infrastructure consists of a laptop for everyone essentially. The bulk of our platform is implemented in Google Cloud. We do have data input processes that are running in AWS, but we've tried to keep most of the platform cloud agnostic so that, you know, we would move to another cloud if that made sense in terms of the economics.

Corey Quinn: Often a great approach, especially with what you're talking about, it sounds like the way that you're designing things requires some software that seems highly portable to almost wherever the data itself happens to live. You're not necessarily needing to leverage a lot of the higher level services in order to start effectively chewing on mass quantities of data with effectively undifferentiated compute services. It feels like the more you look at something like this, the economic story starts to be a lot more around the cost of moving data around. It turns out that moving compute to the data is often way cheaper than moving data to the compute.

Mike Warren: Right. And our philosophy has been to work at a fairly low level. You know, one of our big successes has been essentially implementing a virtual file system on top of the cloud object store. So we can take any number of open source packages and they see a POSIX file system to interact with and we don't have to spend a lot of time rewriting the io routines there.

Corey Quinn: A lot of those open source packages, are they relatively recent? Are they effectively coming from, I guess I want to say the heyday of a lot of the HPC work, but at least what felt like the heyday of it back in I want to say 2005 to 2010 that may not actually be the heyday and just when I was looking into it, but I'm assuming that my experience mirrors everyone's experiences?

Mike Warren: No, it spans the whole range. We've go all the way from you know, 20 year old million line, four tran packages to, you know, the latest convolutional neuro network, which was written a month ago.

Corey Quinn: So as you take a look right now at getting this project up and running, what could AWS have done to make it easier for you, if anything?

Mike Warren: Well, I think it's this environment that they offer is what I worked in 20 years ago. It's Linux plus Intel plus a fast network. And you know, the real key in my mind is the real expense is the software engineers that you need to write this code and deploy it. And now Amazon has eliminated all the friction around the hardware capacity to actually execute that and bring it to reality. So that's kind of the magic.

Mike Warren: Now, any improvements we can make in terms of making our programmers more efficient are not overshadowed by the fact that they can't get access to a petaflop machine to test it or help develop it. You know, it's remarkable, now you can probably get more capacity in AWS with five minutes notice than you can on most of these supercomputers who have to keep their queues very full to keep their utilization high, to justify the cost of that dedicated hardware.

Corey Quinn: When you take a look at doing this again, I mean effectively at first, is this something you would ever consider doing again just as a neat proof of concept? I mean arguably it got more attention when I linked against it in last week in AWS than most articles I link against. So it's clearly resonated with people.

Mike Warren: I mean sure maybe we'll help our local university do the next one or you know, do it in a couple other clouds at the same time. It's a big community and it's not... You know, the top 500 is kind of a... it doesn't provide any utility in itself, it's kind of just demonstrating what could be possible with a set of hardware. So I'm kind of more interested in going to the next level of lets run more codes than we already are in this HPC environment.

Corey Quinn: So I guess the big question here is what inspired you to do this on AWS? You mentioned that the bulk of what you're building today is on GCP. What triggered that decision from your perspective?

Mike Warren: I think AWS has been the first to have the network infrastructure that then makes this possible. You know, in the same way that you need all of the CPUs to be fully available during this computation, you can't have any network bottlenecks. So the new architecture and scalability of AWS's network, not having other users interfering with the network bandwidth available and having low enough latency in these messages. AWS was just the first to get there, but Azure has also demonstrated this sort of performance in their hardware that has a dedicated low latency networks. And you know, I would imagine Google is not far behind.

Corey Quinn: It's interesting and I think that a lot of people, myself included, don't tend to equate incredibly reliable networking as a prerequisite for something like this. But in hindsight that is effectively what defines supercomputers. Things like InfiniBand, how quickly can you get data, not just a across the bus inside of a given node, but between nodes in many cases. It's one of those things that sounds super easy until you look into it and then realize, "Huh, the way this is currently architected from a network point of view, we are never going to be able to move the kind of data that we think we're going to be working on to all of the nodes in question." It's one of those, I think very poorly understood aspects of systems design, especially in this modern world where everyone is going to more or less wind up in a place of, it's just an API call away. The number of people who have to think about that is getting smaller, not bigger.

Mike Warren: Definitely. And there's I think some research that needs to be done in the range of latencies. You know, InfiniBand's down at the microsecond level. AWS can now do things at the 1520 microsecond level and that's a lot shorter than the generation of 10 gigabit Ethernet, which a lot of applications showed didn't have a low enough latency for their needs. So a lot of these important sorts of molecular dynamics, seismic data processing, all these big HPC applications, we really don't know if they're limited by the current implementations or if you really need to go down to the microsecond latency network.

Corey Quinn: What's next? I mean you've been doing this an awfully long time. You've been focusing on HPC, working on computer to scale that frankly most of us have a difficult time imagining, let alone working with. What do you see as the next evolution down this path? I mean, HPC historically has been more or less the purview of researchers and academics. Now we're starting to see these types of things move into the commercial space in a way that I don't think we did before.

Mike Warren: Well that's, I think... I saw a factor of a million improvement in performance from when I started writing our gravitational end body evolution code in graduate school. And you think about that factor of a million in anything else in your experience, it's just never happened. So parallel computing has been a unique experience over the last 20 years and where it's at now is just being available to tens or hundreds or thousands times more programmers. So I think finally we'll get to an era where the investment in hardware has not disadvantaged all of these programmers who could really make some breakthroughs, but it's hard to do that when your code goes 10 times faster every five years without you having to do anything.

Corey Quinn: There's something to be said for the idea of a cloud computing provider where you just throw an application into their environment and you don't touch it again and over time the network gets better, the discs get more reliable, the instances it runs on gets faster. If you try that in your data center raccoons generally carry it off by the end of year three and that says terrible things about my pest control, but also about, I guess the way that people tend to on some level of reduce these hyper-scale cloud providers down to, "Oh, it's just someone else's computer. It's a different place for me to run my Vms," and I think that you're demonstrating that it has the potential and the opportunity to be far more than that.

Mike Warren: Absolutely. I mean there are now clear economies of scale for HBC, whereas before these were all very specialized systems in a not very big market. So the real democratization puts this power in the hands of anyone who can write the right type of software to take advantage of it. And it becomes a true commodity that is really just distinguished by its cost.

Corey Quinn: Wonderful. If people want to learn more about how you've done this and more about what you folks are up to over at Descartes Labs, where can they find you?

Mike Warren: We're at DescartesLabs.community. We've got a good series of blog posts around what we're up to and what we're doing and we've got big plans to grow our data platform beyond geospatial imagery into a lot of other very big data sets that are relevant to what's happening in the world.

Corey Quinn: So I'm going to go out on a limb and assume that you're hiring?

Mike Warren: We definitely are and it's a great place to work for a scientist. I kind of took my experience at a national lab and having worked at universities and I think we've put together the best of the history of research and engineering and made it into a really great place to work and think about software and solve the biggest problems that the world is facing.

Corey Quinn: Thank you so much for taking the time to speak with me today, Mike. I appreciate it.

Mike Warren: Thanks Corey. It's been fun.

Corey Quinn: Mike Warren Co-Founder and CTO of Descartes Labs. I'm Corey Quinn. This is Screaming in the Cloud.

Announcer: This has been this week's episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com or wherever fine snark is sold.

Announcer: This has been a HumblePod Production. Stay humble.

View Details

About Chris Vickery

Chris Vickery is Director of Cyber Risk Research at UpGuard. His research has protected over two and a half billion private consumer and account records which would have otherwise remained at risk of malicious exploitation. He has been cited as a cyber security expert by The New York Times, Forbes, Reuters, BBC, LA Times, Washington Post, and many other publications. Some examples of his high profile data discoveries involve entities such as Verizon, Facebook, Viacom, Donald Trump’s campaign website, branches of the US Department of Defense, Tesla Motors, and many more.

Links Referenced:

  • https://www.upguard.com/
  • Twitter: @VickerySec

Transcript

Voice: Hello and welcome to Screaming In The Cloud, with your host cloud economists, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud. Thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming In The Cloud.

Corey: This week's episode of Screaming In The Cloud is sponsored by LightStep. What is LightStep? picture monitoring like it was back in 2005, and then run away screaming. We're not using Nagios at scale anymore because monitoring looks like something very different in a modern architecture. Where you have ephemeral containers spinning up and down, for example. How do you know how up your application is in an environment like that? At scale, it's never a question of whether your site is up, but rather a question of how down is it.

Corey: LightStep lets you answer that question effectively. Discover what other companies including Lyft, Twilio, Box, and Github have already learned. Visit lightstep.com to learn more. My thanks to them for sponsoring this episode of Screaming In The Cloud. Welcome to Screaming In The Cloud. I'm Corey Quinn. I'm joined this week by Chris Vickery of UpGuard. Chris, welcome to the show.

Chris: Thank you for having me. It's an honor to be here.

Corey: Oh, likewise. I've been a follower of yours for a long time. Trying to keep abreast of the interesting stuff that you do, but we'll get there. First, what do you do at upgrade?

Chris: Well, my title is Director of Cyber Risk Research. I do a lot of things. Probably the thing that I'm most well known for is for leading our breach site team, which is a platform where we give the white glove approach to enterprise level clients that have large network footprints, and want somebody to shepherd over them, and make sure there aren't any obvious glaring problems that can be potentially taken advantage of by actual bad guys.

Corey: I first became aware of you folks because it turns out there's a number of security companies out there. But I became familiar with you when I kept seeing the same type of breach announcements coming out. Well, breach is sort of a lofty term, but I'm sure we'll unpack that at some point. Where companies had not properly secured S3 buckets, and they had exposed varying amounts of customer data. These are effectively brand name companies in many cases, not random hole in the wall taxidermists. Every time I kept seeing, discovered by UpGuard, discovered by UpGuard and I was waiting inevitably to finally snap say, "All right. What's UpGuard?" And the response to be, "Not much, what's up with you?" But it turns out it's not a pun. You're actually a real company.

Chris: Yes, UpGuard is a real company. It was started about six years ago I believe. I've been with UpGuard since 2017, and it was started by a couple of Australians. Our main CEO is Mike Baukes. He moved the company over to the US after getting it started over in Australia, and we're incorporated in Delaware, and all nice and legal and headquartered in Mountain View. Things are picking up quite well.

Corey: So I've become familiar with you folks as the research company that finds publicly exposed S3 buckets and then writes about them. But I'm going to guess that when you're more than a two person company, that probably isn't where you folks start and stop.

Chris: UpGuard has about three, well right now three platforms that we offer. One we call core, and it's the internal kind of on premises, watch your configurations and make sure everything's hunky dory within your environment. Can discover all the random stuff that you have plugged in, that you maybe don't remember is plugged in, and keep everything going well. Then we have the cyber risk platform, that watches your vendor risk ratings, and aggregates the whole total for your company, and scores your vendors on a scale from zero to 950. we say anything under a 600 is pretty bad, and I could probably find some exposed data for them, if I were to have enough time in the world to look at everybody that intensely. Then we have breach site, which is the thing that I am in charge of the team for, and that's where we give the white glove approach to the network footprint of large enterprise customers, and make sure they aren't accidentally exposing anything that bad guys can take advantage of.

Corey: I'm assuming that most of the public ones, where you're cited in the newspaper articles, is stuff that you've discovered as you walk through the Internet. It's not one of those stories where, "Oh yeah, we just do this for our customers and then we write news articles about it that basically publicly shamed them." Is that accurate?

Chris: Well the take that I have on it, and we do have automated systems these days, always looking, finding this stuff, alerting us to do more manual scans and looking at things more intently. But I don't think of it as much shaming as it is raising public awareness of the problem of exposed data. Whether it's involving Amazon S3, or just open rsyncs, or anonymous FTPs, or as your Google Cloud or any number of other hosts out there, people are going to mis-configure their stuff. It's just a statistical probability and human nature. So we would like the public to take that into consideration a little bit more when they're trusting companies with their data, as well as for companies when they're hiring people to work with these platforms, and their customers' data.

Corey: This stuff is complex. I don't think it's unfair to say that no one wakes up knowing this stuff, and it's easy to understand how a lot of these mistakes got made. At least in the realm of S3 buckets a couple of years back. There was a default setting in the web console from AWS, where any authenticated user could read the data. Sure, that makes sense. This is just company confidential. What people didn't realize was that was any authenticated user globally. They since fixed that and then in turn made it increasingly difficult to accidentally do this with, are you sure dialogues, and scary labels in the console, and series of emails that go out, and entire services that are designed to stop this.

So my perception, at least from the outside public world as I've found some of these myself over the years, has been that it feels like this is tapering off. You're not seeing open S3 buckets in the same volume as you used to, and I mentioned that on Twitter and then you commented in which is what started this entire podcast recording and said, "Well, that's not what we see." You are way better positioned to see how this industry is doing across the whole, what am I not seeing?

Chris: Well, for starters, the global authenticated users setting is still an issue. I don't know what you're referring to with they've fixed that, but we notified a fairly large entity just last week of a bucket that had been open for quite a while with that exact setting being the problem to it.

Corey: Yes, for clarity, when I said they fixed that, I meant that in the console it's no longer a checkbox just sitting there waiting there as a trap for the unwary. It's no longer there in the console. You can still set it, but you have to do it explicitly by an API call.

Chris: Okay. That makes more sense to me. Yeah, that's still an issue. But we're seeing plenty of buckets exposed. There's not as many low hanging fruit hanging as low as it used to be, perhaps. Because I'd like to think efforts such as our own have made people and systems administrators more aware of the dangers of leaving data just exposed, or using publicly accessible buckets. But there are still quite a number of them out there. My team at UpGuard has been focusing a lot more recently on supporting our clients, and taking care of basic responsibilities that we have there as well as the advanced new stuff that we're always finding. So we haven't been writing as many reports, but there are still quite a few out there that could fuel a lot of coverage of the issue still.

Corey: Gotcha. For a while here in my office I really only had two pieces of art, because I have a very, well, let's say crappy aesthetic sense. But one of them is a map on the wall of all the announced and active AWS regions and cloud front edge locations, mostly because I want to keep the small map pin industry in business. The other, it was a monitor, just having a consistent ongoing scroll from the certificate transparency logs of S3 buckets that had been opened, and announced as far as there'd be an announcement of a new bucket. Great. Okay. Now there's an automated system that checks and sees if there's a quick list bucket call against this. And this is not a tool I built. This is something that was available on the larger internet, and it would continue to scan this and if it was available, it would flag it.

It also did a similar check for the authenticated user approach as well. But it seemed over time from my perspective that that was became almost entirely noise, as opposed to anything that was substantive as far as being clever and discovering these things. It almost seems like there ... And again this is also tied into the further problem where, for many use cases having an open S3 bucket is a desirable trait. That's something that people want to do, and there are financial reasons why they should continue going down that path. The problem is, is that that's not the bucket you want to start your database backups in, or your user database or credentials to access things that are expensive and important to the company.

Chris: And I completely agree, there are plenty of good use cases. If you just have an assets bucket that's just a graphics creatives, or whatever the heck, and you don't really want to mess around with authentication too much to have random web browsing behavior, be able to pull them. There's no huge problem there. You can do that. You can even make them listable, who cares if people know they're there. They're just little images or whatever the heck. Your transparency log anecdote is along the lines of what I was getting at when I said the low hanging fruit isn't hanging as low anymore. As in I don't know if Amazon changed some way that bucket name registering occurs. But I agree there's not as much of a stream coming from easy feeds like that.

I am familiar with the scripts, and the tools you're talking about there that that have kind of made the rounds. But there are still plenty of them out there as well as plenty that were discovered years ago that are still exposed. Mostly in other countries that don't speak the same language that I do, or anybody that I know does. So I still see it as a big problem, but you know, perhaps the 13 year old sitting in his parents' basement whatever, wouldn't be able to find him quite as easily.

Corey: Absolutely. I think you're right when it's about raising the bar of low hanging fruit. There is an argument, where at some point if a state level actor is working against you, you're probably going to lose for most values of you. Some people in my experience have taken that as like, "Oh, so why even try? Security's impossible." Well, not really. It's a spectrum. Most of us are not getting breached by the Mossad. We're getting breached by some random person running a script they found on the internet, because you forgot to change the default password. It's raise the bar at least enough so that you're not one of those low hanging fruit companies where it's just an easy mistake to make.

Chris: Yeah, that's the whole idea behind security in my mind. It's about resiliency and making yourself not an easy target. Raising it to the point that they're going to go after the next guy. Maybe Mossad is targeting you, but there's other targets that are equally juicy that are less secure. So they'll go after them first, and maybe you'll get ahead of the game. That's the whole security thing. It's not about being 100% impenetrable. Anything that uses electricity, just that blanket statement there, anything that uses electricity can be manipulated in ways that you and I would not anticipate, I'm certain.

Corey: One thing that I've seen with a number of these breaches that have come to public awareness, has been that the company will admit the breach as they're legally required to do. Sometimes they dragged their feet, sometimes not. But they're always very quick to say that it was, oh, a third party contractor did it. And I understand why they want to emphasize that, but on the other side of the coin, they picked that contractor. I do business with a company, I don't vet who they have business relationships with. Given that you do this for a living, more than as someone who's just sits there on the sidelines like I do and angrily observes things, where do you stand on that responsibility breakdown?

Chris: I don't think that you can contract away the liability. I don't like that argument that companies try to toss out there and obfuscate things, where they say, "Oh, look at this clause here, it says we're indemnified against mistakes that are our subcontractor makes," or whatever the heck. You can write anything you want in a contract, but doing business with a third party to handle your cloud stuff, if the third party screws up, you're not absolved of any responsibility there. It's a natural human reaction to try to put the blame on the other guy. But I'm not a fan of that argument, and I don't think there's much legal precedence to hold it up either.

Corey: No, and that becomes a somewhat serious and questionable concern as far as companies think, oh well that's okay. I'm just going to punt the responsibility to someone else. You can't. I don't think you can. It feels to me a soup to nuts that you can outsource work but never the responsibility.

Chris: Yeah, I completely agree with that. I've advocated for a long time about creating ... You can write anything you want in the contract still. But if you were to specify in contracts with third parties of where the work is going to take place, and make it a neutral zone where let's say the names of the buckets that will be used are known and written in the contract, and those are the only ones that'll be used, and you being the first party can check, and see if they're open and exposed to the public anytime you want. It's verifiable. It's what Reagan said about the Soviets. Trust, but verify. If everybody would start doing that sort of approach where anybody can check it, it would possibly keep some of these problems from happening, I think.

Corey: I strongly suspect you're right. It is possible to get this done properly. You remember that article in the Wall Street Journal about how the Pokemon company winds up inspecting the security practices of it's business partners, and the reason that jumped out to me was that it called out in the article that a vendor they were debating doing business with, had improper security controls around an S3 bucket. And their response was, "Cool, we'll use another vendor." That is I think only public example of something like that coming to the forefront. I sent them a polished bucket engraved with S3 Bucket Responsibility Award on it to their office and to my understanding, it's still on display there. They have the good sense not to let me in the building, but that's neither here nor there. But it is possible to make smart decisions. It just requires not assuming you'll double back and fix things later.

Chris: Yes, it is certainly possible to do. There's a certain level of human competency, and human nature that goes into the equation, and that really shines a light on the importance of hiring the right people. Making sure that people you have get the right training, and not just going with something because of sales or marketing something demonstration looked fancy and cool to you. You really got to have the right people that understand, and can integrate with this great new technology. Otherwise you're potentially in for some surprises.

Corey: I think one of the hardest things to get across to folks who are new to the world of cloud, at least the world of AWS, has been their vaunted shared responsibility model. Which is an incredibly boring and droll way of saying that, there are some things that AWS is responsible for, and there's other things that customers are responsible for. An easy example would be ensuring that the application doesn't have a bunch of bugs in it. That's the customer responsibility. Ensuring someone doesn't drive a truck into a data center, grab a bunch of drives and take off. That's AWS' responsibility. I'm curious where in that divide you find S3 bucket permissions.

Chris: My stance on that for a while has been that I agree with the premise that there are certain responsibilities that you just can't strap on being Amazon's liability to worry about. Things that you customize and upload to their cloud space, that you've rented from them. They have no control over what you're putting up there. So bugs in your application. Yeah, they have no control over that. Where there's a fuzzy line is in the concept of, okay, did they develop this platform in a way that is proper for the way that it's been marketed?

If they're making it sound really, really super easy, anybody with a credit card can sign up and upload data, and bam and go. But it really does take a little bit more knowledge than that to do it right, and not risk clicking on the wrong box or whatever the heck. There's an argument to be made that maybe that could be better architected. But that's a continuing goal of any business to improve their product, and make it more user friendly and less mistakes be Made. So that's where I see the line getting fuzzier.

Corey: I would absolutely agree with that. Interesting, at the time of we're recording this, a couple of days ago there was a very public breach on the part of Capital One. It initially to some people looked a lot like an S3 bucket permissions problem. A little more digging turned out that it wasn't. This has been all over the news now and I imagine most listeners have heard about it, but at a very high level, can you give a quick summary of what happened?

Chris: Well, the details still are a little fuzzy when you get down to the nitty gritty. But in essence there was an individual named Paige Thompson I believe, who through some digital trickery was able to enumerate certain information about a lot of cloud accounts. But Capital One is the one that is in the news right now, and was able to list buckets and access the data within them. at least to my knowledge, not because they were improperly made public or anything, but because there were some side doors, and little channels that you can gain information from, and use that information to get a little bit more information. Then if somebody mis-configured part of the chain, you may be able to gain some privileged information that allows you to get through the authentication wall.

It was not a simple thing that anybody on the street could probably do. It was something that required a little bit more advanced knowledge. Interestingly enough, not that this has been reported as a cause of the situation, but this person, Paige Thompson was at one point an Amazon web services employee, a little while ago. So I've brought up the concept of, we need to ask the question, did this person already know how to do the types of things that we're used to get to the data, in this situation? Or did this person have experience from being an employee at AWS, and then corrupted that knowledge into using it in this way, or what? It's not answered right now, and the affidavit that the DOJ filed with the charges doesn't do much to further illuminate that question.

Corey: I would agree. Everything that I've read so far to my mind is the sort of thing that I would do if I dropped my sense of ethics and decided, you know what, let's see how much damage I could possibly do. These are all things that don't require any insider access, and I would be in fact very surprised if it came out that there was any insider access that even remotely came into play here. But you raised the excellent question of, how much of this came from a baseline level of exposure, and experience from working there? And that's one of the fun questions that I think that a lot of companies haven't really asked, is who are the people building services at large cloud providers? Far and away almost all of them are decent, ethical, intelligent people. But as we see, it generally only takes one person going in a strange direction, to start raising uncomfortable questions like this one. The real answer, at least in the world of AWS is, we don't even know publicly how many employees of Amazon work in AWS, let alone the rest.

Chris: Yes, that is absolutely true. As I brought up before we started recording here, when you throw in the contractors and subcontractors as well, it just throws a bunch of wrenches in the machine, and people are not taking this into account when they decide to move their data center into the cloud. There's no way that Capital One has done background checks on all the AWS employees that have access to administrative level things that could be abused. Not that that was the case in this situation. But it's just a good example of, there should be a concern there that I don't think is being addressed very well.

Corey: Yes. To my understanding, we've never yet had a public case of an insider at a cloud provider causing problems like this. I still would argue that we haven't in this case as you mentioned, she left a couple of years before this wound up happening, and since then we saw the giant S3 outage in 2017. AWS is very publicly rebuilt, massive swaths of S3 in a customer transparent way. So even then some of the knowledge around how the system functions internally is going to be out of date. It just comes down to, the question now of even though we haven't seen this in years past, is this a vector for the future?

And AWS is very front and center about how most of their employees, in fact in some cases any of their employees, won't have access to customer data period provided that the customer configures all of the various security apparatus correctly. That's kind of what leads us to where we are now. It's very hard even for a company as incentivized to do that as a bank, to wind up getting all of the edge cases nailed down.

Chris: Yes, that's very true. It illustrates the kind of balance here, the seesaw of, the cloud services provider is going to higher reputable, well meaning people that aren't going to do anything to cause problems. But there's also a responsibility on the client side, the Capital One in this situation side, to configure everything correctly, and have the employees on their side that are knowledgeable enough to configure things correctly, and not expose data or vectors or side doors into data storage areas. So it's going to be a constant back and forth on where the responsibility, and the liability lies in these situations. Like you said, I don't think there's a lot of, if any public cases, or situations where that sort of thing has been really sussed out to this point.

Corey: Absolutely. The fun thing though, that from everything that's been read and reported so far, has been that the attack vector was more or less someone tricking an Edge device of some sort, to make a web request against its own local metadata endpoint. Which if you know where to look, it'll spit out a set of temporary credentials that are bounded at six hours of validity, and then you can grab those credentials, start exploring what else those things have access to. In this case it turned out there was an overly broad role. Okay, great.

Now it will list 700 buckets, and the contents of it, and transfer them out. There's a lot of things that have to happen first to be able to pull off something like that. But also in order to permit that level of oversight, why does a role assigned to a firewall have access to talk to 700 S3 buckets, is sort of the big one that I don't think anyone has a good answer for.

Chris: Yeah. The initial genesis of the techniques that were used here, that initial, how was that request made, that went to the internal facing area, that's still up in the air. Like you said, it's some sort of trick we're assuming. But until it's known more widely, and concretely how that was done, it's hard to say where the initial blame lies. If it was just a completely mis-configured open something or other, that Capital One had either run incorrectly or configured incorrectly, then the blame would lie more on that side. If this was something that was very cryptic, and hard to catch, and maybe affects a lot more clients of AWS, then that may be something that Amazon wants to take a look at. At about maybe putting a little sign, a little flashing sign saying, "Do not click this button," or something. Unless you want to expose things. But we just need to know more details at this point.

Corey: Oh absolutely. I guarantee you that there are other companies that are vulnerable to this, because let's not kid ourselves. As easy as it is to make cheap shot jokes at a company post breach, Capital One Has an awful lot of very intelligent technologists working there. They don't show up in the morning assuming they're going to do a crappy job today. If they can get this wrong, I assure you there are way more companies out there who have gotten it far more wrong.

Chris: Yeah, that's probably very true. I wouldn't have a job if these sorts of things were not widespread.

Corey: It's weird, I'm in the same boat. I fixed AWS bills. You handle cloud security. In an ideal world, neither one of us would have any sort of job that remotely resembles what we do. We'd have to go build things rather than fixing things other people have built. For better or worse here in reality, that's not the way the world works. It's one of those in theory versus in practice stories. In theory, there's no difference between theory and practice, and in practice there is.

Chris: Yeah, and if either one of us were very, very good and godlike at our jobs, we would put ourselves out of business. But it's human nature, always fighting back against that.

Corey: Oh, it's generally a requirement that this type of function be reactive. I think that cloud economics and cloud security are both in the same boat, as far as it's a number one priority for a company immediately after. It really could have benefited from being a number one priority. It always feels like a trailing reactive function, just because it's super hard to invest in this upfront. An argument I made on Twitter a couple of days ago, was that someone could have charged Capital One a million dollars to go in and just fix all the scoping on their IAM roles, and they would've been laughed out of the room if they'd proposed it. But now, because they didn't do that, according to their own statement, they're assuming this year's charge for this will be between a 100 and 150 million dollars. The ROI would've been instant and immediate. But you need to have the pain first before justifying anywhere near that spend on a project like that.

Chris: Yeah. A couple of years ago I read an article that claimed ... Some reputable polling company had talked to a bunch of chief technology officers, and the consensus was that the CTOs would pay out of the company pocket I don't know, $160,000 or something just to not deal with the smallest of breaches. They just, without even thinking twice would toss that kind of money just to not deal with any type of breach whatsoever, no matter how small. So yeah, I agree that if somebody had proposed a million dollars to go in and fix all this stuff, most executives would have laughed him out of the room. But time has told the truth, that they would've been better off doing something like that.

Corey: Oh, absolutely. I will say that it's easy to be angry and blame Capital One for this. They're a bank. They need to take responsibility and handle these things. But looking at everything we know so far, I'm not seeing this as someone just sort of phoned it in one day when they were going about their job. This is a sophisticated attack that understood deeply how all of these systems work together. This is not generally speaking, someone random off the internet who is bored in their dorm room somewhere. This is someone who has expertise in this area, and a deep knowledge of how these parts all interplay together. This is the sort of thing I might come up with, but I'm almost 20 years into my career at this point, and I've been staring at this exact problem space for an awfully long time. It's not in the same realm to me as someone just inadvertently left all of their user data sitting around in an open S3 bucket, despite the increasingly frantic warnings from AWS over the last year or so.

Chris: Yeah. This was a bit more complicated, but it raises the question of, how honest are the marketing and salespeople being when they go and they demo a how great AWS, or any other cloud provider is, and they say, "This is totally secure, as long as you don't mis-configure it etc. etc. There's no way anybody can break in." The executives or the people on the other end of that presentation may take it hook, line and sinker without any grains of salt, and believe that.

But if there is something that a sophisticated attacker can chain together if they're dedicated enough, that needs to be at least part of the fine print. Part of the, "We do as best as we can, but nothing's foolproof. You're taking a risk here, blah, blah, blah." But I get the feeling that's not being represented as realistic as it should be.

Corey: I absolutely agree with you. It's one of those things where, "Oh, don't worry about it and move it into the cloud. It'll be better." But it does raise the question that, if a company hadn't been in the cloud, would this exposure have been worse? Would it have been something, would they have a better security posture if they'd never gone into the cloud in the first place. And sure for this particular use case probably. What would they have exposed instead, by not having effectively some of the best technologists in the world at a public cloud provider, building these things out?

Chris: And how much profit would they have lost for not being as nimble as they can be by using cloud services? It's kind of an apples to oranges comparison. People ask me that all the time, would they have been better off hosting it in their own data center? But it brings up a host of other problems that you've got to deal with then, and un-optimized issues. So it really is, whether you prefer the taste of apples or you prefer the taste of oranges here. They're hard to compare, but you have a preference for one or the other, and they each have their goods and their bads.

Corey: I think that you've absolutely nailed the salient point on that. Chris, thank you so much for taking the time to speak with me today. I Appreciate it.

Chris: Thank you for having me.

Corey: If people want to hear more of your sage thoughts on these and other matters, where can they find you?

Chris: Well, you can always check out the latest blog postings at upguard.com. Or you can go to my Twitter handle, that's Vickerysec, V-I-C-K-E-R-Y-S-E-C on Twitter, and read my various musings there.

Corey: Thank you so much. Chris Vickery, UpGuard. I'm Corey Quinn. This is Screaming in the Could.

Speaker: This has been this week's episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

View Details

About Tara Walker

Tara is a Principal Software Engineer on the Azure IoT product group primarily focused making services for IoT and Intelligent Edge great on Azure. While she now primarily focuses on IoT, Tara has additional expertise and interests in Serverless, Artificial Intelligence (AI) cloud services, and Mobile Development solutions. Over her 20-year career, she has been employed by Amazon Web Services, Turner Broadcasting/Time Warner, Georgia Pacific, and various other Fortune 500 companies.

She holds a Bachelor’s degree from Georgia State University, and currently working on her Master’s degree in Computer Science (MSCS) at Georgia Institute of Technology.

Tara is passionate about technologies including: Artificial Intelligence/ML services & Deep Learning frameworks, Mobile/Game development, Cross-Platform development, and proficient with different programming languages. While primarily focused on IoT services engineering, she also leverages her knowledge and expertise with the aforementioned topics in speaker engagements, as well as, engagements directly to developers & software engineers with OpenHacks and Engineering Code-With activities with Microsoft customers throughout the global community. Her self-imposed goal is to help developers of all walks of life realize they can leverage their tech skills to not only become great engineers but Inventors of the next "Big Thing" that may change the world.

Links Referenced

  • https://lightstep.com/
  • https://azure.microsoft.com/en-us/overview/iot/
  • https://twitter.com/taraw
  • https://www.linkedin.com/in/taraewalker/

Transcript

Announcer: Hello and welcome to Screaming in the Cloud with your host, cloud economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This week's episode is sponsored by LightStep. What is LightStep? Picture monitoring like it was in 2005 and then run away screaming. We’re not using Nagios at scale anymore because monitoring looks like something very different in a modern architecture where you have ephemeral containers spinning up and down for example. How do you know how “up” your application is in an environment like that? At scale, it’s never a question of whether your site is up, but rather “how down is it?”

LightStep lets you answer that question effectively. Discover what other companies including Lyft, Twilio, Box, and Github already learned. Visit LightStep.com to learn more. My thanks to them for sponsoring this episode of Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Tara Walker, Principal Software Engineer at Microsoft with an emphasis on IoT. Welcome to the show.

Tara: Thank you. Thank you so much.

Corey: You are one of those people that I have wanted to get on this show for a long time because when I first heard about you, you were writing blog posts, among many other things, at AWS, and I wrote up summaries of at least a few of those in my sarcastic newsletter, and eventually it was always ... Everything you wrote was well written. The code made sense to the point where someone who's not great at writing code, hi, could wind up making sense of this.

Corey: Talking to you on the record about some of this stuff was always near the top of my list, and it never worked out, and then one day I got the sad news that you were leaving. Now I encounter you again working at Microsoft, and I can talk to you. First, thank you for coming. It is so great to finally get you on the show ...

Tara: Thank you for having me.

Corey: ... a year after I tried the first time.

Tara: Oh no, thank you for having me. Wow, it's really flattering that you even wanted me on the show so I appreciate it.

Corey: Oh yes. So you ... I misunderstood originally because I saw the blog posts that you put out. They were well written, they were well researched, and my default response was, "Oh, she must be on the blog team." Yeah, turns out not so much. You have a deep background in software engineering, specifically again IoT historically, but also you were doing, are and were doing, a bit of serverless stuff.

Tara: I was, yeah. So my role was never official, you know a blog person. That just kind of happened, if you will. It was like, "Oh, you know, we're really swamping our blog team. We would love for our evangelists, our engineers, to come in and have some bandwidth to write about some releases."

So especially being an IoT and serverless SME, and working very closely with the engineering teams, it just made sense for me to be able to speak toward what was happening in those spaces.

Corey: So that's what the past looked like. Can you talk a little bit about what you're working on these days?

Tara: Yes. So, I am extremely excited. This is ... Okay, I got to tell you. So first I work in commercial software engineering. It is a great group, and it's kind of the intersection between the engineering group and some of our big customers that we're doing some of these cool, new things you saw here had build, and actually putting it in real life. Kind of putting the rubber to the road, making sure we're not only solving their problems but making sure our solutions actually are viable.

So it's a great group because we go out ... by the way, as I said I'm focused on IoT, so I get to work with some of the top companies, and some of the companies were actually on the build slide of the keynote. I will not say which ones they are. Some things that we do from our engineering perspective are great. It works wonderful. The customer's happy. But like any other software, there are things where the customer's, like in real life, "You forgot this." Or in real life, "We would love if it did this."

So we will either build interim solutions to help them get to where they need to go, or we'll work directly with the product group that says, "Hey, let's kind of rethink this," or, "Hey, let's tweak this," or, "Hey, if you guys are working on this," because they can only work on so much, "we'll fill in the gap and build that SDK or build that whatever it is, do that pool request, and get things done."

And actually, believe it or not, my first foray whenever I was in my other life in the other company, one of the first things I did was work with the product group, and they didn't have bandwidth, so I built an SDK for one of the services. So it's kind of like back in my engineering world, I've always been an engineer, but now I'm fully in an engineer, and I'm not kind of straddling between, "Oh right, this code. Now go talk about it in the blog." So it's back to my engineering roots, and it's where I love. So it's IoT, it's robotics, it's getting to what we were talking about on edge, that's my passion.

Like, I've always had a bunch of devices. I've always been the person, to the chagrin of my parents, that would tear down my toys and try to rebuild them. Especially electrical ones. So this is ... I'm back in not only my space, but this is a place I'm really passionate about.

So what I'm really focused on now will be IoT edge, things on the edge, implementing things on the edge in interesting projects like with robots and manufacturing that I cannot talk about, but talk about the customer, but how do I get this robotics data to do what it needs to do in the cloud and actually affect, based upon changes, to deliver something to customers that is really cool that I can't say what it is, but I'll just say a few of the customers that I've been working with, to make sure not only we understand the product is right but doing things in production ... we're on the build side, so I was excited.

Corey: I was never that into the whole IoT devices and robotics world for a few reasons. Primarily, that when you break things in the real world it's expensive to replace them. When you break things in code, you just restore to your last save point and try again. And given my proclivity for doing things incredibly wrongly, even a success story means, "And now I have a robot that hunts me through the streets," which let's not kid ourselves, I've had that coming to me for a long time.

Tara: No, no. What's good about IoT now, it's really evolved though where it was before. IoT when it first came out wasn't as connected as it is now. It was a bunch of liking to tinker with things, "Oh, GPIO, I can make this relay go," and things of that nature. But now we're in a space where connected devices are something that's becoming more mainstream.

Now that it's more mainstream, you now see that there are nuances to develop a software for, and I don't want to say the average developer because that sounds kind of remiss, for any developer, you know, can get into IoT, use their skills, your software skills that you have, to now build a device that does things.

So you don't have to be an EE. I'm not an EE. My background is ... I have a CS degree, and I'm getting a CS Master's so I'm not ... Yeah, I'm at Georgia Tech. I would advise all your listeners, do not work full-time at a tech company and go to grad school, especially at someplace like Georgia Tech that actually wants you to be a grad student.

But yeah, so my background is not an EE, it's as a software engineer, but this intersection now where we're really making IoT available for any developer, and you can now solve solutions whether they are complex or simple by the simple fact that we're democratizing the ability to get into internet of things.

I get it, a lot of people say, "Oh my God, the soldering iron, the everything else," and they're a little bit afraid of diving into it, but IoT now is truly, in my mind, for everyone. And this, the fact that we now have machine learning that is now also becoming more democratized, that marriage between IoT and ML, you can do so much.

Even with, as you say, the robot hunting you, you can at least hopefully program the robot and through the cloud push the button, have him stop hunting you, shut down his operations, and protect yourself and save yourself. But you know, IoT now is really for everyone. It's not just for the geeks who like to tinker with the soldering irons.

Corey: It's always interesting to talk to people who have interests that lie in different directions than what I spend my time on. Because at that point, every time they open their mouth I'm learning something. As opposed to winding up in an argument about the best way to structure a web app, and that's why Twitter For Pets is going to be the social media network that takes over the world of pet communication any day now. We're waiting for traction, we're waiting for traction.

So changing gears a bit, let's see if we can have this conversation in a way that doesn't result in angry letters to either of us, but you worked at Microsoft for give or take 11 years, and then you left and went to AWS for five years, and now for almost a year you've been back at Microsoft which first, is a fascinating boomerang story, and it means additionally that you left in such a way that you didn't torch one of the many buildings to the ground to the point where at least, well-

Tara: Everything was still standing, yes.

Corey: Or it was the building that had the records in it, and no one could understand why, but I guess the question I have for you is, "Why did you leave, and then why did you come back?"

Tara: I did a myriad of different things at Microsoft the first time, and I loved Microsoft the first time, but it was going through a transition. A transition in management, in leadership. It was going through a transition in trying to find its direction. It was going through a lot of transition points that people that were really passionate Microsoftees were like, "No, let's not go this direction."

One of those for me was our direction in forcing people, and forcing is such a strong word, and encouraging people very strongly to only use Microsoft tech. And so I had interest in things like Mono and Linux and building not just for the Windows Phone, God bless and rest its soul. Also for iOS and Android. I really was digging Xamarin which was formerly Ximian and everything else, and that was where I was starting to really get passionate about.

Especially around ... That was during the time where the world was just opening up to mobile, and I just thought this Xamarin thing was so cool. I could take ... Instead of me writing all this code in Java and then changing my mind, well not changing my mind, but then redoing it again in Objective-C and then redoing it again, maybe, for the Windows Phone in C#, what if I could actually write something that could be compiled down natively?

So for me, Xamarin, and this is way before Microsoft and Xamarin came together, was fascinating. Then we were doing things like dropping support for XNA for people who were C# developers, not code and everything, but I just thought that was wrong. You know, C# developers were your bread and butter, and now you're saying they can't build games?

So that's when I got into the open source project, MonoGame. That was super cool. So I wrote a bunch of tutorials of how you would build ... I kind of just dived in. My boss at the time, Bob Familiar, who was fantastic hi Bob, was like, "Hey, let's just dig into this and figure this out."

So while he was very open to that, as a culture during that period of time, it was very much, "Why are you doing that? You should not be doing that. You should be focused solely on our own products, only Microsoft 7." I just thought that was really shortsighted. Culturally, it was very much changing from the Microsoft I joined.

Then my management changed, and I lost a support who understood that looking at these other technologies was important, and I really just wanted to do something different. I wanted to get out in the world where it wasn't all Microsoft, I wasn't "very strongly encouraged," I won't say forced, to only use Microsoft products, and that's when I started to think maybe I actually would leave, and that's kind of how that story came.

I'll be honest with you, I never thought I would come back. Like never, ever thought I would come back.

Corey: Hey, I hated Microsoft for the longest time, and 2006 was the last time I used Windows in anger. I swore never again, and today if my current venture were to collapse out from under me, I think Microsoft would be the first company I would call as far as places to work, and I don't know how they did that.

Tara: I mean, I have to give credit ... I had to watch that transformation from afar, so I have to give Satya the credit. He was actually, in my opinion, really looking at where we used to be as what I call the Oh Microsoft where we were going during the period of time I was looking to leave, and then was like, "But we're not this company." This is one of the places I have met the most smartest, I mean like brilliant, genius level people in my life, and genius level people that actually can form a sentence and are really amenable to talk to you.

Corey: Their first language was not math.

Tara: Right. It's amazing, right? So it really hurt, honestly, to leave. When I left, to see him make that transformation, not just back to what Microsoft used to be as far as innovation and send vision, but also to embrace every other technology out there whether it's open source, whether it's Linux, whether it's anything because we wanted to like developers, and for developers' whole purpose in this world, and this is why I like Microsoft's new vision, is to really empower you to build and change the world.

That was what the initial mission was, and I feel like we kind of got a little bit away from that for a period of time. When Satya took over, the transformation of him getting back to that was amazing, even to watch from afar.

So I just happen to see a former colleague that didn't leave, and he just like, "I'm telling you, it's amazing. It's a different place." The good management is really important, and he was like, "I have the most fantastic manager," and he was right because actually we're in the same group, and this guy, oh my God I was like, "How great can you be?" So he kept trying to convince me to really look back at Microsoft and that was, for me, just crazy because I was like, "But that's going back. I'm supposed to go forward," kind of thing.

It was one of the best decisions I think I made, given that I have been not only embraced amazingly coming back, but it is truly a new company. I don't feel like I've come back. I feel like I just joined a new company because it is completely different than when I left. Both in tech, both in culture, and everything. It's not the same place, and it's wonderful actually.

Corey: You are an incredibly well respected engineer. They don't pass out that principal title like candy at any of the major tech companies out there. So with that in mind, if you don't mind the question, why are you pursuing a Master's in computer science?

Tara: Okay, so that's a good question. So it's for two reasons. Really, you're going to laugh, but two reasons. First of all, my family is huge on education. I mean, huge. Every last person, almost, in my family, especially immediate but even outside of that, has multiple degrees. It is completely understood that you are going to get multiple degrees. In fact, my parents tricked me for so many years. I did not know people didn't go to college. I thought ... I don't know what I thought. Looking back on it I was like, "You were so daft," but I didn't know people didn't go to college. I didn't know that was an option because they tricked me. They did trick me. Anyway, so.

Corey: For you, it was not an option.

Tara: Right, right. I didn't know the rest of the world had options out there. But so it was always ... I came out of school, and I got some really cool opportunities as far as job wise, and then just the ball kept rolling. I got a few certs here and there. Like, I got my ISSP and some other things just to go for it or whatever, and I was just doing good progressing and learning in this craft.

My parents still were like, as I've been doing this journey, "But you are going back to school, right?" Like, I would get at the Thanksgiving dinner, "Yeah, we're just waiting for Tara to go back to school," and so Georgia Tech is ... I live in Atlanta so Georgia Tech is a great engineering school. I was like, "Okay, well, let me just bite the bullet. I'm not getting any younger," and there was just a great opportunity to go back to Georgia Tech.

What was great about it was this is also when they were offering their online program, and so I went first just as a traditional student, which I'll be honest was kicking my butt because trying to work and the travel and everything else, and that's when they had this option that you can also take this exact same degree and do it virtually. So the streaming classes, the you take all the same exams and everything else. That's when it really opened up to me that I could actually do this because actually going to classes was not ... I mean, I was making it, but I was sleeping 0%. It was always expected for me to go back. I think I've been forced, but no, so that was kind of why I went back.

I think the biggest catalyst for me finally taking the bullet was, believe it or not, I was at Georgia Tech doing a talk about, you know, I like to go back and try to inspire students. So I was doing a talk about something, some coding, something going on, and one of the professors there ... You know, I don't want to even put that on them. Maybe he's not a professor, but one of the guys that was at the talk came up to me and just said, "Oh, that was great," asked me to what school I went to.

He was like, "Yeah, you know, because we don't ... People like you," these were his exact words, "don't do well getting grad degrees here at Georgia Tech. So it's probably good you got into your career when you did," and so when he said that to me, it was a great, "Okay, I got to back anyway. This is a challenge. Did you just say I can't graduate from Georgia Tech at the grad school because I don't look what you believe like an engineer should look like?" Or someone who can matriculate at Georgia Tech, and so I am horrible about being challenged and then just going for it.

So that was kind of the catalyst. I knew I could stop the conversations at Thanksgiving about, "When is Tara going back?" And also this guy challenged me to say, "We don't really have your kind to graduate here with graduate degrees." It was on. I was like, "Oh really? I could show you." So I foolishly took that. I mean, I'm in it now, and I'm excited. I'm just taking my time going through it, and I'm sure I'll be glad when it's over, but I'm in it now.

Corey: Until then you get to enjoy the journey.

Tara: I don't know about if enjoy is the right word. I get to suffer through the journey, but I'm sure ... Everybody keeps saying it's going to be worth it. But yeah, you're right. I don't think per se that it'll, maybe or maybe not, affect my career, but it's just a checkbox that I just need to go ahead and get out the way.

Corey: Which makes an awful lot of sense. Tara, thank you so much for taking the time to speak with me. If people want to hear more of your wise words of wisdom, where can they find you?

Tara: My wise words of wisdom, oh my. So if you want to chat with me, and I'm really responsive on Twitter believe it or not. Maybe it comes from my old evangelism days, but you can always follow me at @taraw on Twitter. Also, while I've been here at Build, we've done a series of IoT ... I was a dev in IoT for ... dev and seat for IoT here so I've done a lot of conversations with some of the execs around IoT and some of the sessions of IoT so you can follow me there also on the Windows Developer Dev Collective, and they have a section just for me, and we're talking about IoT and some of the enhancements that have happened with IoT, the advancements, the announcements, etc, etc. So you can also follow me there.

Tara: I'm also on LinkedIn. I suck at LinkedIn. So I'm going to ...

Corey: I think we all do.

Tara: Okay. Because I've had people fuss at me that they're mad at me. I'm like... I suck at LinkedIn, but I will give you my LinkedIn because I have no problems people reaching me out to it. It's Tara E Walker. Of course do the LinkedIn.com, but-

Corey: Of course, and they would like to add you to their professional network on LinkedIn.

Tara: I'm sure that'll happen. That's fine.

Corey: Perfect. Once again, thank you for taking the time out of your day to speak with me. I appreciate it.

Tara: No, thank you for having me. This is fantastic.

Corey: Tara Walker, Principal Software Engineer with an emphasis on IoT who has boomeranged to Microsoft. I'm Corey Quinn. This is Screaming in the Cloud.

Announcer: This has been this week's episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com or wherever fine snark is sold.

Announcer: This has been a HumblePod production. Stay Humble.

View Details

About Ken Collins

Ken Collins is a Staff Engineer at Custom Ink focusing on DevOps and eCommerce architecture with an emphasis on emerging opportunities. Custom Ink is approaching its 20th year in business and is entering its second phase in Cloud adoption where Ken helps an increasing growing engineering team succeed using AWS-first well-architected patterns. Ken lives near Norfolk, VA and organizes the area’s Ruby User Group.

Links Referenced:

  • Twitter: @metaskills
  • Custom Ink

Transcript

Speaker 1: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey Quinn: This week’s episode is generously sponsored by Digital Ocean. I’d argue that every cloud platform biases for different things. Some bias for having nearly every feature you could possibly want as a managed service at varying degrees of complexity. Others bias for, “Hey! We heard there was money in the cloud and we’d like it if you would give us some of that!” Digital Ocean is neither. From my perspective, they bias for simplicity.

Corey Quinn: I wanted to validate that so I polled a few friends of mine about why they were using Digital Ocean for a few things, and they pointed out a few things. They said it was very easy and clear to understand what you were doing and what it took to get up and running when you started something with Digital Ocean. That other offerings have a whole bunch of shenanigans with root access and IP addresses and effectively consulting the bones to make those things to work together. Digital Ocean makes it simpler. In 60 seconds they were able to get root access to a Linux box with an IP. That’s it. That was a direct quote except for the part where I took out a bunch of profanity about other cloud providers.

Corey Quinn: The fact that the bill wasn’t a whodunnit murder mystery was compelling as well. It’s a fixed-price offering. You always know what you’re going to wind up paying in a given month. Best of all, you don’t have to spend 12 weeks going to cloud school to understand all their different offerings. They also include monitoring and alerting across the board and they’re not exactly small-time. Over 150,000 businesses and three-and-half million developers are using them. So give them a try. Visit do.co/screaming, and they’ll give you a free $50 credit to try it. That’s do.co/screaming. Thanks again to Digital Ocean for their support of Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud, I'm Corey Quinn. I'm joined this week by Ken Collins, who is a Staff Engineer at Custom Ink. Ken, welcome to the show.

Ken: Thanks, Corey. It's great to be here. Long time listener, first time caller.

Corey: Well, it's funny you mention that, because it turns out that I'm relatively familiar with Custom Ink. For those who were around a year or so ago, or a little less than that, I did a fundraising t-shirt where I had a picture of US East 1 as a burning dumpster on the front of it. All proceeds went to benefit St. Jude's Research Center, which benefits kids with cancer, an issue that virtually no one who was isn't monster is on the other side of. You folks were the house that handled fulfillment and printing of those t-shirts. So I've kept a loose eye on you folks ever since, and it turns out you have some interesting stories to tell about Cloud as well. So this is a fun and interesting combination of something that happened in the real world as well as something that interacts with my own ridiculous nonsense.

Ken: Yeah, everybody wears t-shirts.

Corey: You'd be surprised. I'm known for wearing suits for a living.

Ken: Under your blazer.

Corey: So let's start at the beginning here. What is your background? How did you enter into this whole world of Cloud-native development?

Ken: First, thanks for running that campaign and picking that fundraising platform that we do. I was glad to see that you were able to do that successfully and also contribute money to a needy cause. I've been at Custom Ink for about maybe six years now, maybe seven. I think six. It's a very old company, they've been around for 20 years now, I think. It's a great story. We've been working on getting our Cloud adoption skills up over the past two or three years. That's what my role plays into, I'm the Staff Engineer, I help make technology and architecture decisions across disparate teams and functional groups. I find as a longtime Rubyist, I've been retooling my career, which I think a lot of Rubyists may end up doing as we learned to program the Cloud more.

Corey: Something that I find interesting is that the streams crossed where I tripped over you, not because of any of the t-shirt stuff, but because you wound up releasing something called Lamby. That's L-A-M-B-Y for those who are sitting and not reading the transcript. What is that exactly?

Ken: Lamby is an open-source project that is basically a thin wrapper around your Rails project to help ship your Rails application to AWS Lambda.

Corey: Gotcha. We'll get into that in a little bit, but I found that just to be a fascinating story in its own right. I've had the privilege of walking around these parts of the Custom Ink office, and you folks are bigger than I think most people would think. It's a t-shirt company, that's effectively, what, three people with a screen printing press and that's the end of it. It turns out you have multiple floors in a non-trivially sized building. How big are you folks?

Ken: I think we're about 1,700 employees now. We maybe have about 60 engineers, and those are split up mostly between our Fairfax DC office and a prod team that runs a company called Represent as well. And maybe about half a dozen web ops people. It's a nice, large size company that still got some smaller roots. But we definitely have a lot of production facilities. We've got a huge production facility in Texas, Reno, so that's where most of the employees end up sitting, is in sales, service, and production facilities.

Corey: You have been in business for just shy of 20 years at this point. I'm going to go out on a limb and guess that you weren't a Cloud-native company, given that there really was no Cloud 19 years ago. If we take a look at that, what does your architecture look like?

Ken: Let's see, so a lot of these things went back before I joined. But mainly, I think Custom Ink was one of the first large companies that had a Rails app in production. We still deal with that today, it's a big monolith. We have traditionally two monoliths, like most companies do, where you have a front end and a back end architecture. After about 19 years, we started breaking that up more. We now have hundreds of applications and services, some of them on AWS Lambda, but most of our stuff looks like EC2. About four years ago, we finished a a major shift from your co-located facilities and actually moved into AWS. But it very much looks like a lift and shift. We had a lot of employees back in the day, like, say, Seth Vargo or Nathen Harvey, who used to run a lot of our config management. They've since moved on to bigger and better things, but our Cloud infrastructure, as we moved it into there, it looks very much like a bunch of config management and Rails running on EC2.

Corey: Seth was a guest on this show in its early days, and it's always interesting to me to see just how cyclical everything is, how everyone ties back into the same thing. The world is never quite as big as we seem to think that it is. As you've wound up progressing, back in the heyday of Ruby on Rails, which was what, I want to say early to mid-aughts, and now it seems to have fallen into something of disfavor. I understand that's something of a controversial statement, but there was an awful lot of stuff that wasn't written in the last six weeks that runs powerful companies that do interesting things. These aren't just companies that are doing the Twitter for pet style nonsense, rewriting their stack every month because of something they read on Hacker News. There's validity to software that solves business problems, and a lot of that is written in Ruby. How do you square working on a stack that is, in many cases, perceived as old news because it's stable, it doesn't crash every three minutes, so therefore it's not the programming language of the future, with having to solve real world business problems? I mean, you print t-shirts, it's hard to get more tangible than that.

Ken: I think when you say stack, I hear the words Rails, and the newer stack I'm looking at is Lambda. It's easy to get distracted and buy into technologies that may solve problems that are interesting but don't deliver business value. I'd like to think that one of the things that I do at Custom Ink is always keep that thumb on the scale when we're doing things that reminds people about what we're eventually doing. Customers don't care if we're running in EC2, if we're running in Lambda, or any other thing like that. They want a website that works. They want to be able to design on a t-shirt and have it look good. You may be able to solve some interesting problems in some smaller pieces of architecture, but a lot of times, we're exploring React in different ways. Rails is an easy choice, and I think there's going to be a second coming of Rails after V6 is released. Using new technologies like Lambda, what I like to call full stack serverless, to borrow a term, and pushing bigger, larger things into AWS Lambda.

Corey: So what does that adoption curve look like for you? When you go down the path of building out something that historically has been something of a monolith, and breaking that apart into individual components? First, you have the technical piece, and that's, I guess, an interesting topic in its own right. But I think what resonates a bit more with folks who are looking at that themselves is, what does that cultural transformation start to look like?

Ken: Yeah, I think we're going to hit two sides of that. One of the benefits of Custom Ink is that we have this well understood lay lines of where our architecture and our platform needs to be. After so many years of working, it's easy for us to see where this monolith needs to be broken up. Over about the last year, we've done quite a few high profile lambdas that are just very small microservices. They do one thing very well. Some of them may compose designs on product images, others may resize clip art when you're working in the lab. But those can be solved and you don't really have to think too much about the architecture, they do their one thing well. I think the thing that we're looking at now with how to move the mindset into more Cloud native, especially around Lambda, not getting into containers yet, is just thinking how much you can actually put inside of that lambda, what it can actually do.

Many people will build, say, a node framework, like what are some of the popular ones out there? I think it was called... Oh, I can't remember. But you could put a node framework that does a web application framework, and you could shift it off, and it would be small, per se, some people would think it's small. But I almost guarantee you that if you think of Rails on Lambda, you might think that it's a big thing, and it's a big mind shift to move something that monolithic. Rails is quite small when you look at comparison to node projects.

I think the mind shift is in two things. Realizing when you're dealing with a microservice, and realizing when you're pushing into full stack serverless with Rails and Lambda. And each of those have interesting problems to solve and topics to bring up.

Corey: So why go in a Cloud-native direction? Obviously, what you've built has been working for you. It's gotten you this far, for lack of a better term. What was the advantage of moving off of that legacy architecture?

Ken: Well, I don't think we're going to know for another few years or so. But the bet is that we can better adopt Cloud-native technologies, the full offering of AWS, and ship faster. Many of our applications now, some people might find this horrific, we're deploying applications through Capistrano and that legacy manner. Shifting to Lambda is a way that you can get inside the ecosystem of AWS. From there, you can branch out and say, okay, now let me go learn about DynamoDB or Route 53 for latency based routing, or these N number of other offers that they have out there, and start to make good, isolated platform decisions, where larger amounts of our teams can make these decisions and get things out without having to be tied down to our ecosystem. Where you don't have to hook things into Capistrano for deploy. In fact, you don't even have to deploy at all. You can start bringing up more modern things like code pipelines and automatic deployments and stuff, green/blue deploys, et cetera.

Corey: Gotcha. That's not necessarily to say that there's a right or wrong path here. It's one of those areas where there is absolutely a capability story of taking advantage of new technological developments. It's just always interesting to me to see how that manifests at a company that isn't venture funded, where engineers generally don't tend to have carte blanche to do whatever they want. Seeing how companies that are, I guess, doing things here in the real world with a solid business model behind them, are thinking about these things.

Ken: We always have the new apps that we have to build or the new services that are tied to business OKRs, and we're generally left free on how to do that from an engineering side. We have good architects that help makes decisions for us on what investments we want to make in certain technologies. But the way to do that is really left up to the team. We even have some groups now working on a new container based system that might be ECS, that might be Fargate, it might be Kubernetes, who knows?

Corey: It shouldn't be Kubernetes. It shouldn't be Kubernetes, but that's okay.

Ken: Yeah, we don't want to give money to the Greek god of spending money.

Corey: Exactly. No, it's always fun to wind up seeing how this stuff tends to come to be. I'm not entirely sure there's any right or wrong direction as far as how most of that goes. But I'm curious about your personal transformation, as well. You started off as someone who was a longtime Rubyist, and now you've been transitioning into what looks at an awful lot like a Cloud-native developer. What's that been like?

Ken: Well, I think I've been putting it into three phases here. One is learning about. The other is choosing new tools and then building pipelines around those tools and those technologies, and those run concurrently and synchronistically together, and even sometimes recursively where one dovetails or other way back into the beginning.

For the learning side, that's been really hard. I've, myself, have gone out and got AWS certified from, I believe it's called the Developer Associate. Being someone who hasn't gone through academic training, I don't have a CS degree, I'm very much self taught, I was very hesitant to get certified. I feel in some ways, certification is just to learn enough to pass a test. But learning about and working in the tool has more value. But I've ended up finding that certification is a good way to learn about something. It's like why you go to conferences. You don't go to conferences to learn things, you go to learn about them. Learning about AWS, it's just going to be a full time job. That certification gets you a kickstart into that, and hopefully gets you hooked into always keeping up with AWS offerings, hopefully with things like Last Week in AWS or any other avenue. But that learning is the first part, is the key. You have to know about all these things in AWS because they're certainly asking you to architect solutions around cobbling and all of these things together. And if you don't know the right stones to reach for, then you're not going to do well.

Corey: I think you just hit on something very important. That the purpose of conferences is not to learn something new, it's to learn about them. I feel like an awful lot of folks go to the big AWS conferences and others with a mind towards attending specific technical sessions, to learn specific implementation details. At least for me, I've never found that to make a whole lot of sense. You want to wind up figuring out what those tools are, of course, and you want to see how those things wind up being used in the real world. But especially at a conference where all of the sessions wind up on the internet two days later, I'm not convinced of the utility of flying halfway across the country to do that. I feel like people don't tend to come into conferences with the right strategy in mind.

Ken: Maybe it's because when they go to the conferences, their only dialogue is that one-on-one they're getting from the conference speaker. They're not maybe going up and engaging the speaker before or after the session. Maybe they're not having enough hallway conversations throughout the conference about those topics. I tend to agree. But I also like to go and learn about how other people are doing very minute problem solving and learn if I could fit that in my tool or not. I'd say if you're going to a conference to figure out if Kubernetes is right for you, then that's probably not the right solution.

Corey: I think that a lot of people also go the other direction, view it as just an excuse to go party somewhere for a few days. And sure, you can do that. But I don't find it to be the most productive use of anyone's time when we're talking about being around some of the smartest people on the planet when it comes to a particular technology.

Ken: One of the things we've done at Custom Ink is we pick out these new technologies and how we'd like to work with them, and we have a roadmap of keystone projects that help us understand the architecture. So we're not going to do a lot of meetings and make decisions about, all right, everybody, is Fargate or Kubernetes? Hands up or hands down. We're going to maybe look at a few projects, we're going to spin some up, and we're specifically going to call them out as these are exploratory projects that we're going to learn if we like something or not. Sometimes they may end in something being set down in stone and have a ways of working, but we're really cognitive about, hey, this is something new. This is something we want to try. We don't even know if we know enough about it to make a call or not. Let's just explore that and let a few people go deep into the tea and just drop down really low, go learn all this stuff, and when they surface back up, we'll share that laterally with the org.

Corey: I think there's a lot of value in being intentional in the technologies you select rather than just what seems to have the most stars on GitHub or something similar to that. It winds up almost, instead of trying to follow the herd, you're trying to solve for specific technical problems that you have. That said, you can take that too far, and reinventing everything from first principles and going off in your own direction, that is completely at odds with the rest of the larger industry. If nothing else, it makes hiring a lot harder.

Ken: Yeah, and this is one of the things. I do like GitHub stars. If you do star the Lamby project, I would love you to death. But I think Lamby, for me as a Rails developer, and I hope this is true for other Rails developers or anybody looking to do Ruby specifically on AWS Lambda. It's this incredible thing where I can take this ecosystem that I'm familiar with, I can ship it off to AWS Lambda, and it forces me to reconcile the other things that are going on at AWS and how to be more excellent in that ecosystem. I don't think there's any better way to learn if something is right for you. I'm not even worried about containers right now in my career, but I will be in a few months.

The one thing that I think that will fit together with Rails developers, and I hope they adopt with the Lamby project, is that once you get into this ecosystem and once you have this capacity to ship Rails to Lambda, you're forced to learn about these other things. I'm having a ball every day learning more about AWS. I think last week, I recently started putting some deep thoughts into how we deploy our applications across multiple regions. Not for disaster recovery, but for availability and redundancy. I would not have thought of that if I was spinning up a single dyno in Heroku or just putting things on the EC2 and filling them over the the DevOps wall, if you will.

Corey: Well, getting to that end, how do you wind up hiring effectively into an environment where you're starting to do something that is relatively unheard of in the larger industry? By that, I do mean Lamby, in that until last reinvent, when they announced a Ruby runtime for Lambda, if you wanted to do this, you were using things like Traveling Ruby and having to shove a whole bunch of things into place. And frankly, Rails itself never really seemed like much of a particularly Lambda friendly environment when everything is built in a monolithic sense where there's startup latencies and the rest. You've made that work with Lamby. So I guess a twofold question on that. The first is, how do you wind up effectively hiring people to start wrapping their heads around this who have the experience that you need? And secondly, how do you wind up seeing this evolving over time as best practices among Cloud technologies continue to evolve?

Ken: I think you have to take two approaches. One is, you have to hire people that you think have the right ability to learn and solve problems, and just train them into your ways of working, whether that be Ruby or Node or what have you. The other is maybe start to look a little bit remote for your hires. Traditionally, Custom Ink has been very centric to hiring in the Fairfax and northern Virginia area, but we're looking at maybe changing that. I think problem solving is the most important thing a programmer can do, besides, say, communication or something else. Rails is not so hard. It's easy to train somebody into Ruby and stuff. But it is a hard problem to solve. Rubyists aren't coming out of the walls anymore. They're just sitting there delivering business value softly in the background.

Corey: For better or worse, that's probably the right answer. I think all of us need to deliver business value, and it turns out that rewriting everything every three weeks is not necessarily the best way to do that for most shops. As you decided to start down this path of converting what you had running in Rails into Lambda functions or series, you wound up having to build your own framework to do this, called Lamby. Tell us a little bit more about that please.

Ken: Sure. I'd like to think it's only several dozen lines of code, and a lot of it's inspired from AWS themselves on how to get a Sinatra app running. Right now, Lamby basically just acts as a small wrapper around your code in Lando to send the handler to a rack compatible app, which a Rails app is. It doesn't do much more than that. Everything else that the project talks about on GitHub is mainly tooling around using AWS SAM, which is the serverless application model command line interfaces, to package and build the app and deploy it out to production. If anything, it's more deploy scripts than a framework for Rails.

The one thing that I really like about the project is that it actually, other than the internals of your app and how you use, say, S3 or DynamoDB or anything like that, the project does not care about your Rails app. There's no difference between a Rails app for an EC2 instance or a Rails app for Lambda. It's basically just your Rails app and you're delivering it. I like that approach a lot, and I think it's going to afford us some certain decisions later on in the project's life cycle when we eventually write wrappers for application load balancers versus API Gateway. But for now, mainly all the gem does is, it acts like a rack up file where it mounts your application for the Lambda handler and just sends direct messages to it all day long.

Corey: How constraining do you find a Lambda environment for the way that Ruby on Rails historically thinks about applications?

Ken: The only one that I've really hit upon now is, I've not really done any database connections. I know that there's the VPC cold start issues, so as soon as you attach a VPC to your Lambda, you're going to incur a penalty of maybe some 15, 20 seconds of startup time from cold. That's true if you use a Lambda with, say, ElastiCache, like Memcached as well. So there's those things that we haven't really touched on. A lot of the Rails apps that we've been prototyping in it are isolated image processing. We're taking advantage of DynamoDB versus, say, PostgreSQL or MySQL. And I have not found that too terribly discouraging. I'm an RDBMS guy, I love databases. It's been a hard mind shift to start thinking of DynamoDB, but I'd say it's database connections.

Corey: I think that historically, the challenge I always had was that Rails tried to do everything for you under the hood. And that's great, right up until it wasn't, where it started doing things where I didn't agree with the data model, active record, and a bunch of issues with databases under the hood. Effectively getting in, unwinding some of that here be dragons. It's always felt like it was terrific for rapid prototyping, something I needed up and running quickly. But as that application grew and scaled, breaking it apart, even upgrading from one version of Rails to the next, was hell. There's really not a nice way to put that. Finding ways to modernize that onto something that resembles a serverless stack feels like it would be something almost insurmountable. I wouldn't know where to begin. And yet you didn't just begin, you did it.

Ken: Maybe it's my proficiency in Rails, but I find it incredibly easy to... To me, it was an easy decision to go, hey, if I don't have an active record model, let me use the AWS record jam, which is a thin wrapper around DynamoDB. It even encourages you to mix into things that active record does, so out of the box Rails will give you active model validations. The same patterns are requested in the AWS Record gem, which basically will bring DynamoDB to the table for you. You mix in active model validations, and pretty much you just feel at home again.

It does require a little bit of tool sharpening and bringing things to bear on it. But I think the the gems out there to get things done, if you're looking at a full stack serverless, which I definitely recommend people do, there's definitely uses for small applications, where a Rails app just does one thing through a couple gems. But the problems are easily solved. If you're doing things like image processing or if you're not moving your oracle database over to serverless, then I think it'd be just fine.

Corey: So what's next for the project? What is it that's keeping it back from the next step on its path to greatness?

Ken: I'd like to do a few regional comp talks about it where we do some workshops. We've done a few internally for Custom Ink where we walk engineers through creating their own Rails app. Very much like a 201 class where you get a little bit past hello world. I think for me, it's basically going to be moving to, what it it, the Application Load Balancers for the event, versus API Gateway. I read a recent blog article that said that was a lot of cost overhead. Not a lot. Everything at Lambda is going to be cheap, but that you could probably get a little bit better performance and a little bit more cost squeezed out if you switch to the Application Load Balancer.

Other than that, it's really just building a community around it and finding out stories and people talking about how they solve certain problems. People can pull more resources together and understand that, hey, if I had an application that did one small thing, what's the pattern that you might choose from, or, my Rails application might be using more of the the Java script stuff. What's the best way to build the assets for that and ship them off to a CDN, et cetera.

Corey: How do you see the future of the project evolving now that Ruby is a first class runtime in the world of Lambda?

Ken: I think it's going to grow. I think there's going to be a lot of people, a lot of Rails programmers, that are, to date, shifting their careers, doing things like migrating Heroku instances to AWS, that are doing this type of consultancy that need to learn more about AWS. With tools like Lamby out there, they're going to be able to make some really good decisions for them on how certain Rails applications go into AWS Lambda and how it uses more of their ecosystem.

I'm excited to learn more about people's stories about what they can share on the the Lamby website. We can contribute more blog articles and also technical content for the product website on how other people can take those learnings. And that's really where I'd like to see it go. Just more success stories and more adoption and more documentation.

Corey: If people want to learn more, where can they find out about Lamby in general and you specifically?

Ken: Lamby is located at Lamby, L-A-M-B-Y, dot Custom Ink tech dot com. I'm pretty much metaskills on every other thing, whether that be Twitter or Xbox or LinkedIn or what have you.

Corey: Ken, thank you so much for taking the time to speak with me today. I appreciate it.

Ken: Thanks, Corey. I've had a good time.

Corey: Ken Collins, Staff Engineer at Custom Ink. I'm Corey Quinn. This is Screaming in the Cloud.

Speaker 1: This has been this week's episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

Speaker 2: This has been a HumblePod production. Stay humble.

View Details

Dr. Galen Hunt founded and leads the Microsoft team responsible for Azure Sphere. The mission of his team is to ensure that every IoT device on the planet is secure and trustworthy. Previously, Dr. Hunt lead the Operating Systems Group at Microsoft Research and pioneered technologies ranging from confidential cloud computing to light-weight container virtualization, type-safe operating systems, and video streaming. Dr. Hunt was a member of Microsoft's founding cloud computing team and helped build Microsoft's first cloud operating system. Dr. Hunt holds 98 U.S. patents, a B.S. degree in Physics from the University of Utah, and Ph.D. and M.S. degrees in Computer Science from the University of Rochester.

Links Referenced

  • https://azure.microsoft.com/en-us/services/azure-sphere/
  • https://twitter.com/galen_hunt

Transcript

Announcer: Hello and welcome to Screaming In The Cloud, with your host, Cloud economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming In The Cloud.

Corey Quinn: This week’s episode is generously sponsored by Digital Ocean. I’d argue that every cloud platform biases for different things. Some bias for having nearly every feature you could possibly want as a managed service at varying degrees of complexity. Others bias for, “Hey! We heard there was money in the cloud and we’d like it if you would give us some of that!” Digital Ocean is neither. From my perspective, they bias for simplicity.

Corey Quinn: I wanted to validate that so I polled a few friends of mine about why they were using Digital Ocean for a few things, and they pointed out a few things. They said it was very easy and clear to understand what you were doing and what it took to get up and running when you started something with Digital Ocean. That other offerings have a whole bunch of shenanigans with root access and IP addresses and effectively consulting the bones to make those things to work together. Digital Ocean makes it simpler. In 60 seconds they were able to get root access to a Linux box with an IP. That’s it. That was a direct quote except for the part where I took out a bunch of profanity about other cloud providers.

Corey Quinn: The fact that the bill wasn’t a whodunnit murder mystery was compelling as well. It’s a fixed-price offering. You always know what you’re going to wind up paying in a given month. Best of all, you don’t have to spend 12 weeks going to cloud school to understand all their different offerings. They also include monitoring and alerting across the board and they’re not exactly small-time. Over 150,000 businesses and three-and-half million developers are using them. So give them a try. Visit do.co/screaming, and they’ll give you a free $50 credit to try it. That’s do.co/screaming. Thanks again to Digital Ocean for their support of Screaming in the Cloud.

Corey Quinn: Welcome to Screaming In The Cloud. I'm Corey Quinn. I'm joined today by Galen Hunt, a distinguished engineer and the managing director of Azure Sphere. Welcome to the show.

Galen Hunt: Thank you, Corey. It's great to be here.

Corey Quinn: So Azure Sphere is a lot of things and I'd like you to tell us what that is, but the most compelling part that I saw was in a single sentence on the website: "Our goal is to make IOT safe for society." Through the lens of that very inspiring statement, what is Azure Sphere?

Galen Hunt: What is Azure Sphere? So Azure Sphere is an end-to-end solution for addressing the security needs of IOT devices. Okay. It consists of three pieces. There are Azure Sphere compatible chips that are built by our silicon partners and incorporate intellectual property from Microsoft into them. There's an operating system that runs on those chips and then there's a Cloud service that works with the chips and the operating system to keep the devices based on them secure. And that's fundamentally what we're trying to do. We're trying to make sure that any device manufacturer can build a device based on Azure Sphere and ship it out and know that for the lifetime of that device, it is going to remain secured.

Corey Quinn: And that's, I guess, from a very naive perspective, not having much of a background in IOT myself, but I think of the internet of things, this entire world of devices that are living in my house. I go out and I buy a scale or something and it talks to the internet. In other words, I don't know, maybe it posts on Twitter to shame me whenever I gain weight. And that's awesome and I keep that thing for years on end.

So instead of focusing on a real device... For example, my Twitter for Pets company, my side project, decides to get into the IOT space. We're going to build combination toaster-refrigerators and it turns out that the product does not see a lot of market success because physics. And after selling a whopping three of these, we pivot. We post, "Our amazing journey has come to an end" on Medium and raise another round, because that's apparently how failure works today.

And we still have those three that are out there, and at that point, the Cloud services that we were paying for have been turned off. There's nothing for the other end to talk to, and assuming that there isn't a failure mode where we have just bricked the expensive thing that people have bought from us, you now have this thing sitting there, unpatched, in perpetuity...

Galen Hunt: Sitting on the internet.

Corey Quinn: Exactly. And now, one day, someone, maybe a state actor... Oh, sorry, in InfoSec, we call them nation state, which irritates a lot of people, the same way that on-premise instead of on-premises does. So we're just going to call this episode On-Premise Nation States, just to irritate everyone.

But once you look at that and it starts attacking things, there's a responsibility issue and there's a how do you even identify that that is a thing that your device is doing? If you think about that, it feels like an incredibly large-scale problem with no easy answers.

Galen Hunt: It's a huge problem, because in the old days, if you made some new device like this toaster-refrigerator combination... Gosh, I'd really love to have one, okay?

Corey Quinn: Oh, yeah. Saves so much space in the kitchen.

Galen Hunt: Yes, exactly.

Corey Quinn: And the unexplained fires have not been proven in court.

Galen Hunt: In the old days, you could build one of those and you could sell it to your customers and basically, your engineering job, your hard work, was done the day you shipped it, because you never saw that thing again. The problem is, when it's an IOT device, your hard work begins the day you ship it. It's the day that it goes into a customer's home or into an office or another environment and it gets connected to the internet. That's the day the Internet -- and the hackers come. And from then on, until that thing is disconnected permanently from the internet at the end of its life, it is at risk, from a security perspective.

And this is the fundamental thing. You know, IOT is super powerful, because it creates a connection between that device and the manufacturer or the customer and the manufacturer. It creates a connection, but every Internet connection's a two-way street. Right? And what that means, "Hello, hackers." So and what we're trying to do with Azure Sphere is we recognize it... This company that builds this refrigerator-toaster, they know how to build a refrigerator-toaster, let's hope. Knock on wood.

Corey Quinn: In theory, yes.

Galen Hunt: In theory, okay, but in practice, almost none of them know anything about internet security and it is a hard place to be. The internet's a very scary place. I have a former colleague who's a professor at Harvard, James Mickens, and he likes to say, "The internet is this cauldron of evil." Nation states, professional hackers, whatever you want to call it, and what we try to do is say, "Well, how can we package up the experience that Microsoft has?" Because, by the way, we've been doing this for a really, really long time. I've been at Microsoft 22 years. My entire career has been spent working on internet security, one form or another, trying to keep the hackers out.

And we said, "Okay, is there some way that we could take all of this expertise and experience that Microsoft has and package it up so that we could give it to device manufactures and then actually keep giving it to them so that we could help them keep building secure devices?" And that's what we've fundamentally created.

Corey Quinn: It seems to me, looking through... The way that I've historically seen Cloud services tend to manifest is... There's an economic challenge here where people are going to pay for a ridiculous IOT product like a toaster-fridge or a scale that fat shames you or whatever it is that you wind up buying but they're generally not going to want to pay a subscription for that because it doesn't tend to comport with our mental model of how services work. So people will go and they'll spend money, sometimes a lot of money on something like that, but they're not necessarily going to want to sign up for a recurring subscription model.

So the challenge then becomes you just need to be able to provide secure Cloud services for things that, in all likelihood, are going to be talking to the internet way longer than anyone thinks they will. It's, "Oh, I'll just get that scale for two or three years" and mine is coming up on 10 years old. I'm sure it's an attack vector for something now but I'm irresponsible. There's an economic story where if you have to pay at a monthly basis or per API call that thing makes to a Cloud provider, that, at some point, you are now spending more on the long-tail Cloud service than the thing made you in profit and you are losing money on every sale. How does that wind up tying into how customers are approaching IOT today from a security perspective?

Galen Hunt: Well, so one of the things we looked at is how do we make it... You want to make the security decision be a one time decision. Do I want a secure device or not? Okay. Hopefully, the answer is... The answer should always be, "Yes." Particularly, you don't want people asking on a month to month basis. "Do I want security this month or am I feeling lucky?" And in fact, the business model we came up with, the Azure Sphere, is it's a one time transaction. When the manufacturer decides to buy an Azure Sphere chip, they get with it from their distributor the chip and the license to our operating system and the license to our security service. And that includes the ongoing security work for both through a period of the expected lifetime of that device.

So let's say, a 10 year period, and it's a one time so nobody's paying money... 10 years in, seven years in, you're not paying more money to keep that device secure. And the other thing that we did with the Azure Sphere is we've actually separated out... If you typically looked at an embedded device, the device manufacturer takes an RTOS and they take their code and they put it together and they're responsible for everything. And what we've done is we've actually broken up the way the code is factored so that we can keep updating the operating system.

So there's new security vulnerabilities and new security threats and attacks come out, we can update the operating system. In fact, we will. We will update the operating system and the security features on the devices out in the field. So let's say, to use your refrigerator-toaster example... Let's say they go out of business or they say, "We're not going to support this thing", if it is based on Azure Sphere, Microsoft is going to keep supporting that and we're going to keep updating it and addressing security vulnerabilities until that thing is done.

Corey Quinn: To be honest, it wasn't even until this conversation where we look back at things like HeartBleed. When that came out, I was doing a fair bit of consulting with a number of different customers and talking to them, making sure they were patched, making sure my own stuff was patched. But not until now, did it occur to me, "You know, I wonder if that stupid scale of mine at home wound up getting patched or not." Almost certainly not because the company got acquired twice and who even knows at this point. It's basically a hazard to all around it in an emotional way and a physical way now but it's... This is not something anyone, even people who think about this stuff in a security context, are generally going to think of intuitively.

Galen Hunt: Yeah. Because you want to just buy that thing and install it and forget that you have to worry about it, right?

Corey Quinn: Yeah.

Galen Hunt: And that's exactly what we're trying to address here.

Corey Quinn: To be clear, there are remarkably few companies that could make a statement of, "If your company goes out of business, that's fine. We're going to continue to maintain security updates for the infrastructure for this IOT stuff." But if anyone's earned that, it's Microsoft at this point. The long tail legacy support for fascinating and varied use cases is borderline legendary. And for anyone who’s had to write code around some of this, it's kind of obnoxious to have to still work around, "Well, people are technically still using Internet Explorer." They announced that in the keynote at Build where the next version of Edge now has built in Enterprise support and two or three people in the audience just lost it, cheering.

And you look around, like, "Oh, those are the sad people." Because we lived that life. We know what that pain looks like. But the idea of being able to have a perspective of looking long term back at... This is important and it needs to be able to support this from a business continuity perspective. It's powerful and I think Microsoft gets that, arguably better than anyone else that...

Galen Hunt: We have been doing it a really long time. Like I said, I've been at Microsoft 22 years. And I remember when the Slammer and Blaster viruses came out and us having to figure out... I was on the task force at Microsoft to figure out, "Okay, how are we going to address these class of things and make sure that they don't happen again?" And all the skills that have to... And if you think about building a highly secure device, and that's the term I use... So highly secured is something I can just really depend on the fact that it's secured. There's a lot of skills that go into that. There's a lot of engineering up front that you have to do to get all the pieces together right so that you don't have... If you use a really bad random number generator so that even if you have this amazing crypto, well, it doesn't matter because you've thrown it away, the random number generator.

So it's a bunch of engineering and then there's this ongoing work that you have to do of, every time some new vulnerability like... What's the one you used?

Corey Quinn: HeartBleed.

Galen Hunt: HeartBleed. Okay. Like HeartBleed or the crack vulnerability in the WPA2, the Wi-Fi protocols, okay.

Corey Quinn: Oh, that brings me back.

Galen Hunt: Yeah, a year ago, or when these new things come out, somebody's got to look at that and say, "Does this apply to this advice and what are the changes we have to make?" So you've got to have an ongoing security expertise. And then you figure out a patch. Okay, you can say, "Oh, well here's how we're going to mitigate that. We're going to fix the patch." And then you've got to have this expertise of, "How do I actually roll it out to every fat shaming scale on the planet and make sure that everybody's device is actually updated? Do I roll trucks? Do I send emails or the devices automatically update themselves?"

So you have to have this operations logistic expertise on top of this ongoing security analysis expertise on top of the engineering expertise you have to have. And what we're basically trying to do is take all of that and offload that to Microsoft.

Corey Quinn: Well, where are the bounds of Azure Sphere in that sense, where if I build a device and I put this solution into it, it obviously controls the firmware. It winds up controlling the version of RTOS patching. Does it control, for example, the Wi-Fi aspect of it? Is that in bounds for this, assuming there's another Wi-Fi WPA2?

Galen Hunt: So it's pretty extensive because we own the entire operating system. And so for example, with Wi-Fi, if there was a crack vulnerability. Let's say someone's coming up with a new vulnerability... Actually, let's talk about the crack vulnerability. What had happened, it was a little over a year ago... Almost a year and a half now.

Corey Quinn: Why does it feel so much longer ago?

Galen Hunt: There's a lot of IOT security news out there, right? It just keeps coming. We had a fix, a verified fix, for that available within 24 hours of the vulnerability because one of the things we've also learned how to do very good at Microsoft is figure out what is the fix that we have to do for a particular vulnerability and how do we test our systems so that we actually know that the fix is correct, et. cetera. And then we had the deployment technology to build and deploy that out within hours to billions of devices.

Corey Quinn: And none of the customers who manufactured these things even had to think about this? It was simply done for them.

Galen Hunt: So if you were using Azure... If you had an Azure Sphere based device. Say you're a manufacturer and you build an Azure Sphere based device and you get woken up with this headline of crack vulnerability. If you're using Azure Sphere, what's your responsibility? Go back and go to bed. We got your back. It's our problem. And that's a key thing. We own the entire operating system stack on the device. On, not just the bits that we give you as a manufacturer, but literally the bits on the device. So that we're going to fix them out on the devices on the field and we also own the security services providing so that all the bandwidth for the updates and we do the updates both for the OS... We also provide an update channel for what we call the application. The OEMs code.

So the device manufacturer... Let's say the toaster-refrigerator, they come up with a new feature... I don't know. It's a thing that's going to shoot the ice cubes out into the toaster because everybody wants toasted ice, right? And it turns out that's just a software update. Well, they can... They can and they want to get that software update out to all their customers because who doesn't want toasted ice? Well, they create the new update and they turn it over the Azure Sphere security service and say, "Hey, deploy this out to all our customers", and we do the heavy lifting for that as well.

Corey Quinn: One thing that I've always found aligned with the security mentality is the way that I tend to approach Cloud economics. Specifically, in that, no one sets out to build a product or service for the least possible amount of money so waste creeps in, in the same way that almost no one sets out to build a product from day one to be the most secure thing in the world. They want to build a thing that ideally gains market traction and people buy it and security as the number one bullet point doesn't move almost any of these things unless it is a security device itself. So there's something to be said for using this service... And effectively at that point, you are taking the entire security issue and more or less outsourcing the work if not the responsibility to a provider that it just works.

And everything handles itself. That's compelling. That's the sort of story that I think is going to win the security wars, for lack of a better term. And I'm not talking about competitors' security wars. I'm talking about the ongoing battle against the cauldron of evil. It's how you wind up getting somewhere that you don't have to go out of your way to do the right thing. You've built a guardrail path where doing the right thing is easy, straightforward, and is, in some ways, much easier than doing the wrong thing.

Galen Hunt: Well, that was the objective. I launched this thing five years ago. Got it started, building the initial prototypes and everything. And that was the objective, was, "How do we make it so that security is so simple that everybody uses it?" Okay. And it was really critical to do that because, as you said, people don't immediately recognize, "Oh, why do I need security?" or "How much do I... " It's like, "Oh, I just want to do just enough security." Well, the problem is that the internet's a really, really dangerous place.

Corey Quinn: And it's not getting less so.

Galen Hunt: And it's not getting less so and just because you're new to internet security doesn't mean that the hackers are new to internet security. And so there's a pretty high bar of what it takes to build a device, even today, even if you just build for what are the known security issues right now. It's a really high bar and it really, really requires a lot of expertise and so we're trying to address that. The other thing I'll mention is you talk... People tend to say, "Oh, nobody's going to be willing to pay for security." We believe security is the differentiating value prop of IOT. Okay? Because when it really comes down to it, nobody wants their refrigerator-toaster that creates botulism or that blows up their house and the line between an IOT device and a dangerous device is really, really thin without security.

Corey Quinn: Oh, absolutely. But putting on the front of the box, "Won't burn your house down", in big letters is one of those, "huh!" That's selling a breakfast cereal, it's like, "Contains no rat poison." Well, it wouldn't have occurred to me to ask that question until you bring it out there. That's the marketing problem.

Galen Hunt: Yeah, well one of the things we have found... We've done a lot of looking at this. One of the things we did is we did a security survey with consumers across the United States and Europe. We interviewed somewhere about 3,000 individuals. We actually went and had face to face meetings and talked with them. And what our data showed is that most people, the vast majority of people... If they knew that a device was secure, they would buy a secure device over an insecure device, and they would pay more money for it.

Corey Quinn: From that perspective, is security framed as won't attack the underlying DNS infrastructure of the internet or is it contextualized more as privacy? I make a joke about a fat shaming scale but having it leak your personal information is, I think, a lot more resonant with people than some ephemeral, "Well, one day, the internet's going to be slow and broken" and my failure mode is, "I'm going to have to go outside for a little while."

Galen Hunt: Yeah. You kind of have to make it personal and one way I try not to scare people but if you just kind of step back and think of it like... So one of the things we're trying to do with Azure Sphere is make it even approachable for micro-controllers, the very cheapest class of computers. And to make it really personal, if you go into your home, it's a micro-controller that is keeping your furnace from creating carbon monoxide and poisoning your family. It is a micro-controller that is keeping your gas stove from exploding. It is a micro-controller that's keeping your dish washer and your washing machine from flooding your house. And today, those things are safe because they're not on the internet at all. But when they come on the internet, they have really got to be secure.

Corey Quinn: It's... We've talked a lot about ridiculous IOT approaches. Do you have an example of a customer or two that's doing it right? As much fun as it is to sit here and talk about terrible ideas that should never have been built, I'm more interested in a uplifting story. Who's using Azure Sphere today and making the world a more safe... making society a safer place for IOT?

Galen Hunt: We have a company in Europe called Eon that is doing home energy management systems and they've got car chargers and batteries in homes and solar power systems and you think about... There's a lot of electricity running in those things. They could actually... Those things could be dangerous but then Eon said, "No, want to make sure that these are trustworthy systems and are metering right and everything else." And so they've chosen to use Azure Sphere.

Corey Quinn: It's fascinating to see just the different verticals that these things tend to get used within. You talk about in almost the same paragraph, you talk about a retail establishment that sells coffee and a solar power company. And we're starting to see that the entire world is, in fact, becoming more connected. And it's... There are a lot of people who hear something like that, and I confess I'm generally one of them, who thinks, "Is this all good? Is this going to be something that leads to a better society or does it lead to a story where suddenly every bit of information about me is for sale on the dark net to the highest bidder?"

And that has been an area of growing concern. At this point, I've started thinking, "Oh, well, how many devices do I have on my internet connection at home?" And I realize, as I just think mentally, last time I looked at that to update something in my mobile app for the Wi-Fi, there's over 40 devices connected. There are three humans who live there. That seems a little excessive but everything starts to wind up being connected and this is going to be an area that is absolutely not going to go away anytime soon and it's getting safer.

Galen Hunt: Until Azure Sphere. We're going to make it safer. It is not going to go away and it's just going to keep coming. Security is necessary for privacy.

Corey Quinn: Yes.

Galen Hunt: Okay. Because if your devices aren't secure, it's like, OK. Well, if they're secure, then there's the question of, "What's my relationship with the manufacturer then? What's their privacy policy? Et. cetera" and things like that. But if it's not secure, hey, that stuff's open to any hacker that wants to come in. We've seen... We've seen headlines, IOT security headlines, fridges sending spam and baby monitors being used to spy on families or project messages into families. And so you really, really want these things to be secure.

Corey Quinn: As compelling as this sounds, it doesn't work, generally speaking, to think of security in a context of absolutes, like the idea of M&M security's always a challenge. You wind up breaking through the perimeter and now you have everything there. How does Azure Sphere tend to address that particular threat model, if at all?

Galen Hunt: Okay, so when we think about security... We actually published a paper that I co-authored about two years ago called The Seven Properties Of Highly Secured Devices, particularly to help explain to people how they should think about security because as we'd go out and talk to device manufacturers early on but a couple of years ago, we were just getting to kind of the prototype proof of concept stage. They would... Sometimes we'd have this conversation, they'd say, "We have some security, is it good enough?" So we'd try to help them frame that. And one of the topics we talk about in that paper is defense and depth and this is, "Do you have multiple layers of defense so that when something goes wrong, if somebody is able to circumvent one layer of your security, you've got another."

Give you a just kind of physical example. You think about, if you go into a fairly secure building, like, say, a courthouse or something like... A Microsoft office. Some of them or things. Or a bank. You'll go in and there will be locks on the door and there will be a guard and there might be a metal detector and there's video cameras and there's a safe. Okay. And that's because someone might be able to figure out how to break the lock on the door but then you've got a safe and then... Or you've got cameras so that you can figure out who it was. And you've got all these different layers and that's because, well, if you have only one layer of defense, you have a single point of failure and that means if something goes wrong either intentionally or accidentally in that piece, you don't have any security at all.

And the thing we found, most IOT devices that are out there today have really been built with... It's the M&M, hard on the outside, soft on the inside, security model instead of this defense and depth. And what we've done with Azure Sphere is we have multiple layers of defense and depth so within the hardware itself, we have three layers of defense. In the truest way, in the operating system itself, there are four layers of defense and depth in the operating system. And that's so that if hackers are able to find a vulnerability, get into one piece, they can't just keep going and be able to... In fact, we can actually... We, detect that they've gotten into a device and we can kick them out and renew the security on that device.

Corey Quinn: Fascinating. That's one of those areas that, I guess, makes a lot more sense once you get into the space but coming from an outside perspective, it would never have occurred to me to start thinking at that layer of complexity. It's a war that's probably never going to be won but you can absolutely embrace the stakes.

Galen Hunt: Yeah. And it's the... It's what's required out on the internet today.

Corey Quinn: I think we'll want to hear more about your thoughts on this. Where can I find out?

Galen Hunt: So I'm on Twitter. Galen_Hunt on Twitter. We also... They can go to the Azure Sphere website and find out more.

Corey Quinn: Thank you so much for taking the team to speak with me today. I appreciate it.

Galen Hunt: Thank you, Corey. It was a great conversation.

Corey Quinn: Galen Hunt, distinguished engineer and managing director at Azure Sphere. I'm Corey Quinn. This is Screaming In The Cloud.

Announcer: This has been this week's episode of Screaming In The Cloud. You can also find more Corey at screaminginthecloud.com or wherever fine snark is sold.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Christina Warren

Christina Warren is a Senior Cloud Advocate at Microsoft, where she helps shape the overall strategy for Developer Relations in Azure. As an advocate, she hosts shows on Channel 9, Microsoft’s video channel for developer content, speaks and creates content at events, conducts on-camera technical interviews within the developer community, and liaisons with product teams across the company.

Prior to joining Microsoft, Christina spent a decade in digital media as an editor, senior reporter, and commentator, with a focus on technology, business, and, entertainment. As a journalist, she appeared as an expert or commentator on ABC, NBC, CBS, CNN, CNBC, Fox News, Fox Business, Bloomberg, the BBC, Marketplace Radio, The Today Show, Good Morning America, and many more outlets.

She also co-hosts Rocket, a popular tech news podcast, which has the distinction of being one of the only tech podcasts with an all-female hosting team.

Links Referenced

  • Twitter: @film_girl
  • http://christina.is/
  • Rocket on Relay FM

Transcript

Announcer: Hello and welcome to Screaming in the Cloud with your host cloud economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This week's episode is generously sponsored by Digital Ocean. I'd argue that every cloud platform biases for different things. Some bias for having nearly every feature you could possibly want as a managed service at varying degrees of complexity. Others bias for, hey we heard there was money in the cloud and we'd like if you would give us some of that. Digital Ocean is neither, from my perspective they bias for simplicity. I wanted to validate that so I pulled a few friends of mine about why they were using Digital Ocean for a few things and they pointed out a few things. They said it was very easy and clear to understand what you were doing and what it took to get up and running when you started something with Digital Ocean. That other offerings have a whole bunch of shenanigans with the root access and IP addresses, and effectively consulting the bones to make those things work together.

Digital Ocean makes it simpler. In 60 seconds they were able to get root access to a Linux box with an IP, that's it. That was a direct quote, except for the part where I took out a bunch of profanity about other cloud providers. The fact that the bill wasn't a who done it murder mystery was compelling as well, it's a fixed price offering. You always know what you're going to end up paying in a given month. Best of all you don't have to spend 12 weeks going to cloud school to understand all their different offerings. They also include monitoring and alerting across the board, and they're not exactly small time. Over 150,000 businesses and 3.5 million developers are using them. So give them a try. Visit do.co/screaming and they'll give you a free $50 credit to try it out. That's do.co/screaming. Thanks again to Digital Ocean for their support of Screaming in the Cloud.

Welcome to Screaming in the Cloud, I'm Corey Quinn. I'm joined this week by senior cloud advocate, Christina Warren. Christina welcome to the show.

Christina: Thank you so much, I'm so excited to be here.

Corey: I first found out about you years ago when you were a guest on the Accidental Tech podcast. And now you're on Screaming in the Cloud, which is undoubtedly a brand new career low for you. So first my condolences.

Christina: Well first of all thank you that ATP episode was one of my favorites, so we actually swapped hosts that week. So John Siracusa went to Rocket and I went on with Marco and Casey. And I still, it's funny, this just shows the reach of their podcast, I still have people, you hear like, "Oh I first knew you from that." And I'm like, yeah I will probably never reach an audience like that big.

Corey: I am still astounded when I get stopped at cloud conferences invariably and I'm asked, "Are you the person from the podcast?" And it's, oh wait, how do you know what I look like? I have a face for radio, there's a reason I do this and it's always strange seeing where people come in, it's weird. Then you start getting loud on Twitter or do a newsletter, or go fall in love with the sound of your own voice and give conference talks as I've done for entirely too long. And it's, you never quite know where people first hear of you and start-

Christina: Totally.

Corey: ... following you.

Christina: Yeah, no that's great that ATP was your entry point, that's awesome, I did not know that, that's really cool. That's still one of my favorite podcasts.

Corey: It really is. Whenever I get the chance I listen to that but it turns out there are not enough hours in the day for all the various podcasts out there.

Christina: This is becoming a real problem, that we're like now adding to obviously. But, so kind of like peak Netflix, we're not a peak podcast but it's one of those things where I'm like, it used to be something you'd do on your commute and now you're like, oh but if my commute were four times as long, then I could get through my podcast but.

Corey: Increasingly it seems like one of those valuable commodity, we all have his attention. And because that is finite we can't make more of that, and some day I'm sure the most valuable commodity will be water. But until that societal collapse, attention seems to be it right now.

Christina: I think you're right, yeah.

Corey: So once upon a time you were a journalist.

Christina: I was.

Corey: And now you're a cloud advocate.

Christina: I am.

Corey: That is atypical as far as career progressions tend to come.

Christina: It is, it is. And it's funny because I'm actually going to be giving a talk tomorrow, kind of about how I switched careers. And, the interesting thing though is that I was kind of an atypical journalist in terms of how I got into that too. So even though it's a very uncommon trajectory, it sort of makes sense in my life. So as long as I can remember the two things that I, well the three things I guess that I've loved the most have been computers, technology, gadgets, whatever, pop culture and writing, and story telling, whatever. And when I was, I guess when I was 12 years old I was in a bike accident. And I ran into a tree and messed up the side of my face and broke my jaw and all that good stuff, fortunately no scars. And when my mom, when we were coming back from the ER from getting the X-rays and stuff she was like, "You can have whatever magazines you want." As like solace for me breaking my face.

And, I had already had that month's Teen Beat or Seventeen, or whatever, and I'd already decided that I'd wanted to make it like my summer project to learn about computers. And so I picked up two issues, one of PC World, one of PC Magazine, I didn't understand really anything that was in them or on the covers, and I just started reading and going to the library. And this was before I had the internet at home, and just really nerding out and loved it, just loved it instinctively. And, so I did web development stuff and I didn't study CS in college because when I took programming in high school, I didn't really like the people that I was in the classes with and I was like, I don't want to spend four years with this.

And, my whole career, even as a technology journalist, a lot of times I was way more technical than a lot of the other writers. But I had to explain things to a new audience. And then I come over to Microsoft and I'm still a very technical person but there are people who are way, way more technical than me. So I kind of go from being like the person who is always having to kind of dumb myself down to being like oh man, I need to like step up. But, if I think about it, the things that I did as a technology journalist, and the things that I do in dev role aren't that different in so far as, I would go to conferences like we're at Microsoft Build, right now, I would go to conferences as a journalist, and I would talk to developers.

And I would hear about the things that they were excited about, or the things they were mad about, and I would try to advocate for them, and I would try to write articles that I would hope that the companies were reading saying, these are real issues. This is why this platform is taking off, or this is why it's not, and really kind of get the feel for those things. Because those people are my people. I've always been, friends with developers, whether it's like web, or app, or whatever. And now that's still kind of my job, right? It's just, I can now go directly to the product teams, rather than having to scream on the internet, I can just scream in person.

So it's still it's a uncommon career change, but a lot of it really does come down to liking the audience, liking the content, and then wanting to advocate for those people.

Corey: It also speaks to a mature advocacy organization that, doesn't approach you and say, "Wow, we absolutely love your ability to tell stories. We love the credibility you've built up in the space. But in order for us to hire you, you have to solve an algorithm challenge on the whiteboard." I've gone through interviews like that in years past for developer evangelist style roles. And the question was always, well what is the purpose of this?

Christina: Right.

Corey: Well, we need to make sure that you have technical credibility in front of an audience. Cool. How did you hear about me and consider me for this role in the first place? "Oh, you have a massive following and massive credibility in front of the audience." Did you hear those two sentences put next to each other? It becomes a very strange story. And that's what it's about for me. Is at least the storytelling aspects of it. And that's one of the things I find compelling about Microsoft as a whole, is a very practiced approach to telling stories. Satya's keynote on day one of the conference was pitch perfect, probably one of the best delivered keynotes I've ever seen. Every word was very intentional, emphasized correctly. And when I tweeted about that, I got a few responses of, "Well, was he reading from show notes to do it?"

Yeah, I don't care what tool you want to use, I don't care if you have an earpiece and someone whispering what to say in your ear, stand in front of 2,000 people and give a talk, you're going to be nervous, and a spoiler, you're going to do a terrible job, at least the first few times you do that.

Christina: Yes.

Corey: It doesn't matter, it's just the ability to tell a compelling story and deliver that well, is fascinating.

Christina: I agree. And I also, okay, so does he have a teleprompter? Does he have show notes? Do you know how hard it is to read off a teleprompter? I do it all the time, I'm become very good at it, it is very difficult, and there are people who are so much better at it than me. And so people who say that the first thing I think I'm like, okay, well, you've never done this. Because if you did you know that in many cases, it's actually harder to read off the teleprompter, than it is to be extemporaneous, or to memorize.

Corey: Most of my speaker notes when I'm giving a talk on the bottom of slides are either bullet points, or things not to do.

Christina: Same.

Corey: Don't make this joke because it's coming in three slides. Good. Because after you do that a couple of times and realize you've completely neutered your own joke, you kind of don't want to do that again.

Christina: No, I'm the same way. And when I give talks and whatnot, my speaker notes are usually bullet points, usually not scripted. There has been, I've been part of Microsoft Ignite the tour for the last few months, and those talks which are more structured, and I'm not always the only one giving them, are a little more rehearsed and practiced. And at this point I've done them so many times that it's pat. But in general, my whole public speaking style usually is tended to be more extemporaneous. And I don't say that as a brag, that's the easy way, right? Like, to me, people who have something really well practiced, and rehearsed, and written out, that's a skill that I don't have. I mean, unless it's on a teleprompter that I can read from.

Corey: One of the best things I ever did to learn to give a decent conference talk, was to give a bunch of really terrible ones first. I started this way back in the early days of 2012, in the dark ages. And there was a project I was working on, Salt Stack, as it turns out, which did not do as well in the market as I would have wanted it to, it's sort of not really where the zeitgeist is. But, this is amazing, why is no one talking about this? Good question. Why aren't you talking about this? So I put documentation on slides, I read my slides to people, I mumbled into my own shoes, and eventually people like, "That was great. Now here's how you could do that in a way that isn't objectively terrible."

And it got slowly better. But for me the breakthrough was when I was a traveling trainer for Puppet. Going into different cities, sitting down in front of people who'd paid a fair sum of money to learn how this worked, in some cases, a very hostile audience because they perceived, rightly or wrongly, that this was going to automate them another job. And they were sort of forced to be there by their employer, and then periodically because computers, demos would break. So the first week, I'm white knuckling this the entire time, by the third week it's okay, I sort of have this and by the seventh week it's, this is super boring, I'm giving the same exact talk for three days and doing the same jokes, and you mix it up a little bit. But past a certain point, you get tired of telling the story the same way and you want to start improvising.

Christina: Yeah.

Corey: And having the flexibility to do that when giving the same talk repeatedly is important.

Christina: I agree.

Corey: So from your perspective, when you're on Ignite the tour and traveling to all these different places, and giving what is fundamentally the same talk that you've given a bunch of times, but is new to the audience, how varied do you make that?

Christina: I haven't been varying it too much this time. I think when we do the tour next year, I might mix things up a little bit more. Part of that has been because, we have the slides that have already been localized for whatever language they might be in. The other thing is that even though, and this is a challenge that I haven't ever had before. In many cases, English might not be the primary language of the country that I'm speaking in. So that means that I might have a live translator, where if someone will be translating, as I'm speaking in people's ears, or we're having the closed captioning translating happening. And so for those reasons, I don't want to play too much with the format. But you're right, because that is something I always think about. I'm like, I would really like to mix this up more and maybe make this more innovative.

But sometimes it's a challenge when you have other things to consider. And so in those cases, I'm like, all right, you know what? I'm not the goal of this right? That's I think the important thing too, is recognizing in some ways, and this isn't always the case but I think this is important to know. When you are the focus of the talk, when it is kind of about you and when it is not, when it is about the content. And in these cases, it's not about me. They want to get the information and I want to present it the best way that I can, but it's not the Christina show. It is about getting the information to people and explaining things, and then hopefully being able to answer their questions and provide them with more resources at the end.

Corey: It's about the audience, it's never about the-

Christina: Completely.

Corey: ... presenter.

Christina: You're right. I mean, there are a few times when you might be giving a more personal talk, or might be something about your background where you can take that on more, but you're exactly right. But sometimes I think that, especially those of us who speak a lot, can kind of get caught up in that idea of oh, it's about me and what I want to project, and what I want to do, and it's like no, it's not always. And sometimes you really need to, even if it's boring for me to do the same talk over and over again. And to be clear, it hasn't been because again, these other challenges of, do I have a translator? How is the audience going to know this and whatnot? Then that's enough of a kind of a difference to not make it the same old, same old. But even if it were, I think for this particular thing, and like I understand the context, it's like this is for them, this isn't for me.

Corey: One thing that you've talked about previously in other shows and things that you have done, has been talking about how you view developer advocacy. And I think a common misconception industry wide is that, to do that, you need to be a public speaker. The way you've defined it, it sounds like you cast yourself almost as more of a public listener.

Christina: Yes.

Corey: Of talking to customers, and more importantly, listening to customers, as opposed to some companies that talk at customers, but that's a separate problem.

Christina: Right.

Corey: So I guess one of the questions I have is when you're out there talking to people who are using Azure for various business tasks, for setting up new things, for exploring, what have you learned from them?

Christina: I learn all the time what pain points are. I learn maybe what we're doing well, what we're not. And those are the things that I really want to take back to the product teams, and then help maybe if I see something strategic that we might be able to do. I want to help improve that. I learn all the time about whether our message is getting across or not. Because there are many times I think, and this is why I think advocacy is important, where, and it's not the product team's fault, I am no way dismissing them. But you know, engineering, you have certain goals and you are in sprints, and you are kind of in your own head, and you're getting your project done. And you don't maybe always have the best insight into how are people really using this or not using this.

Like of course you have internal telemetry and analytics, and you can see things on Twitter, but if I'm talking to someone, and they're saying, "It's really hard for me to, why can't I take a custom image and move it to a different region or a different subscription? Why do I have to copy things to a new VM, and a new storage blob, and do this process and it's convoluted and I do this repeatedly, over and over again, and I have scripts that will do this for me, but this is not a seamless process." And I'm going, you know what you're right. And I can talk to the product teams and know the technical challenges around it, and I can empathize with that, and I can try to express that to the customer. But I can still say, okay, but this is still a really big problem and we need to continue focusing on finding the solution, because this is something that's impeding the way they're using things and what they're doing.

Corey: You're talking about something that started to, I guess, tease at the back of my mind. So at the beginning of, I'm recording the show at Microsoft Build, and they're in the first day, waiting for the keynote to open, I had half an hour to kill. So I spun up my first Azure account, and I decided to spin up a VM. And I hit a, I was live Tweeting the entire experience because I don't have an internal filter, and the entire world cares about everything I'm having for lunch. And I started Tweeting the experience. There were a lot of great things at it, and then spinning up a single VM, it defaulted to picking a Debian release and oh my stars, that brings me back.

Christina: Yes.

Corey: And it was fun and exciting, and it took 19 minutes and 22 seconds to provision. And, I'm huh, this seems kind of slow. And people chime in and start looking into the problem under the hood, there was an issue with the storage layer in that region at that time. And I have this knack for finding corner cases by blundering into them. And I think one of the directors of compute wound up, one of the directors in container services wound up chiming in and digging into this. And it was, their customer service experience was exceptional. But I'm starting to realize now that by being loud and noisy, even though my cloud bill personally is 20, 30 bucks a month, I do not at this point, have the typical customer experience because I am perceived, rightly or wrongly as being loud and influential.

Christina: Right.

Corey: I am loud, I would argue whether I'm influential.

Christina: You're influential.

Corey: I try, my mother tells me so and I try to believe it, but you know how mothers can be. The problem I see as I look at this though is, I don't have a bearing anymore on what the "typical customer" experience is, I don't know. For example, is that something that anyone who had that problem and Tweeted about with Azure, would they have experienced that? I'd like to hope so because it's scale.

Christina: It doesn't scale. I don't think that you would see the director of compute responding to somebody with 200 Twitter followers, right? Like, if we're being honest. I just don't think so just because I don't think it would hit their radar. But what has impressed me, and I can't say this for all the times is, how many times I will Tweet things, and obviously, I'm influential in so far as any of us are, and I'm loud, and have a following too. But I'll Tweet things without tagging people or just might be innocuous. And the Tweets will be found by some of the QA teams, who will then comment and reach out. So people are actively looking for these things all the time, and gathering that feedback, and I think trying to help when they can. And it's not always going to be that direct kind of hand holding, let's investigate what the issue in this region is.

But I think that the goal is to hear as many inputs as we can, and even though it doesn't scale where every single person is going to be able to have that type of experience, is you want to be able to, when I'm talking to people, I don't know or care what their cloud bill is. If they're having feedback, or if they're having an issue, I want to connect them with somebody who can solve their problem period. Right? Like that, to me is is the goal. And that whether we succeeded at that or not, that's for other people to determine. But I think that's always the goal is to get people's questions answered. And I would hope that the experience would be realistically, you're not, someone else who doesn't have your following is getting the same level of maybe hand holding. But I would like to think that there would be things in place that people would pick up on and be able to say, "Hey, this is going on, or have you tried this?"

Corey: One thing that I will say and go on record with, is that this is not at an individual level. Something that I see as a differentiator between any of the public cloud providers. Every individual employee I know at all of the major players has something very similar to this attitude. If a customer's having a bad time, that's a problem and needs to be addressed. The challenge that I see and where I think Azure is excelling in this space is, in making it very clear in communicating to its existing 40 year history of customer install base, that they are very interested in talking to them, and learning from them, and supporting them wherever on the curve they tend to be.

And sure the future is going to be pure, not purely but, largely public cloud driven. And if you're not there yet today, that's okay. There's no sense of we will begrudgingly go ahead and give you an offering that sort of works for your current environment, but you really have to move, clock is ticking. It's much more of a, we are here when you are ready to go here. As opposed to taking a data center and slowly flooding it. And as the water rises higher and higher up the rack-

Christina: Hahaha, you better come.

Corey: Exactly.

Christina: Told you, told you, we told you. No I mean, I think you're right. And I think part of that is, when you have, as you said like a 40 year old business that has differentiated business aspects right? So which is going to be unique in our space, where you have lines of business that have existed for a long time where we know A, firsthand how important some of those workloads might be, and how difficult it might be to migrate them as much as we might want them to. And also just the fact that some people might not want to do that. And, I think it's kind of, legacy is always a difficult thing to kind of figure out how long do you to support something? When do you kind of force people to move on?

But Microsoft has a very, very well known history of supporting legacy things. Some would argue too much, right. So I think that that is part of kind of the goal is to, yeah, we're here for you when you're ready, but we are going to continue to offer other offerings too. Obviously the benefits of cloud and an iterative development are that you get updates and things more quickly. Like I think like Office 365 is a great example of that. There's still a standalone version that it's kind of locked in time, and you'll get some maybe security updates and whatnot, but if you want to get the frequent feature updates that are pushed all the time, you need the cloud accounts. And, I think that for a lot of people, that can be a compelling reason to do that, but there might be some people who are like, "No, you know what? I don't care. I just want my perpetual license." And that's it.

Corey: There's an entire class of users out there who if you move an icon, the thing is broken, they open a support ticket. I've been accused of being insulting when I say this before, and it's never intended that way. But, Microsoft has 40 years of experience in apologizing to its customers for software failures. And you need that expertise when you're working with cloud. I'm sorry, it's computers, it breaks, that's what they do.

Christina: Absolutely.

Corey: And communicating that to the customer in an effective way that first expresses empathy with them. But secondly, it says that yes, we know this is a problem, we are working on this, thank you for continuing to do bear with us on this and communicating that well, and not leading to the various gossip rags tearing the company down in tech browse, is a skill in its own right. And it's one that I think Microsoft can foundationally teach a master class in.

Christina: I think you're right. I mean, and because you're right. I mean with that history and with products that are used by as many people as use them, even if you had the highest level of QA in the world, you're still going to have issues, they're computers.

Corey: And even direct honestly doesn't work here. Well, if your software had been written differently, our outage wouldn't have taken you down. While true, and while something useful and valuable to communicate, if you're not very careful with the phrasing, it sounds-

Christina: Right it sounds you're blaming them.

Corey: ... blaming the customer.

Christina: Exactly, I mean, that's the thing. And when we have outages and that happens you know at every company. I've actually been impressed with how well some of the write ups, some of the technical write ups are, because this used to be what I would do as a reporter. AWS would go down and usually it'll be like an S3 bucket and like North Carolina or whatever, right? Like it'd be like Eastern Region and it would go down. And then the entire consumer internet goes down. And you have to explain to mom and dad, why Pinterest isn't loading. And you know, they don't care about the various esoteric things that oh, well if they'd had multi region, or this or that, they don't care and this is down.

And sometimes, and I always feel so bad for the dev ops people and the SREs in those situations who are getting the calls, and are trying to get things up and running, because you never know what's going to happen. But sometimes getting the information about what it is, can not always be easy. It's, I think Amazon has actually done a really good job as of late of making those, communicating better when those things happen.

Corey: In some ways they've been forced to by their own customers.

Christina: Right, right. But that's something that I have, when we've had outages I always look for because, people usually aren't coming to me, they're not yelling at me about it. But I do a weekly show on our YouTube channel that's Channel 9, and if we have like a major outage, I need to talk about that, right? I need to actually be able to explain this is what happened, and this is what we're doing about it. Because otherwise I don't think that that's authentic or helpful, right? Like if this is, we're publicly communicating with people, we need to do that. And so, I've been really impressed with how well some of the write ups have been, like the postmortems after the fact, right? Like you have the log that says it's happening, but the postmortem is after the fact I'm going, oh, that's interesting.

And then when you look at that you can kind of understand, wow, these are massive challenges, and these are like maybe a set of events and circumstances that would not be anticipated. and that would be whatnot. And in some ways almost, I almost wish that we would have some of our instances be used as almost like a class or talks for our users to be like, this is what we've learned from this, and this is how you can learn from the same things that we have because the challenges aren't wholly unique, right? Like a data center going down and maybe, or part of it having an outage because of something being configured the wrong way or maybe like a weather condition or something else and not having made decisions for whatever reason to architect things that might have made that not happen.

When you learn that, when you look at what are the trade offs, those are interesting discussions to have with your end users too. Because they're going to be, they're not worrying about the data center itself, but they might be having to worry about their own places a fault might happen and how they're going to recover from that. I think that it's not unrelated, and there are lessons to learn from that, period.

Corey: And one of the most understated parts of the cloud as a whole is, let's say that I write a reasonably terrible web app, because that's the level of developer I am. And then as soon as it's done, I make the final commit, I host it on a cloud provider, and then I sail around the world. 10 years later I come back because I'm as good at sailing as I am at writing code. If I'm running this in a cloud provider, things are great. It is, the storage has gotten faster, the reliability has increased, whereas if I leave this running in my own data center, the raccoons have taken it down by year three. Things inherently get better as opposed to succumbing to bit rot. And that's something that I think is not talked about nearly enough.

Christina: I totally agree, I totally agree. You know, the one thing you have to keep in mind though is that if you're settling on a role for 10 years, your software and stuff, your VMs, they've all been updated, right to keep up with security. But like your code might stop working.

Corey: Or you're assuming my code ever worked in the first place, that's beside the point.

Christina: Okay, that's my point.

Corey: It's my code, come on now.

Christina: Well, you know what I mean? I mean, I think that's something that is different, that is also, I mean, this is great, but this is also I think sometimes a challenge and a thing with educating people saying, okay, yeah, once you have this, this is great that these things are going to be staying up to date and get the patches, but if something goes, if like a version of PHP finally has been like out of support for a long time, and we're going to cease it and force an upgrade on the library to that. If your code hasn't been updated for that, it might break.

Corey: This is also part of the advantage of a server less story as well, where it's okay, my code might still be crappy, but I don't have to worry about the-

Christina: I don't have to worry about the other stuff.

Corey: ... TLSPs, or the web server, or the patching or-

Christina: Yes, I 1,000% agree. I think that is like the wonderful part of the server less story, and hopefully we can get some more of that where, because ultimately, that would be a great thing to not have to worry about that stuff. Where something's upgraded underneath me, and now my stuff breaks. Because that's like, and then that's a hard challenge too. How much do you just let people use out of date, non updated things, right? You could conceivably let them do it forever, but is that the right move?

Corey: It's always a problem. I used to work at a web hosting company running a giant WordPress grid. And that's awful the plugin problems, the outdated version of PHP that everyone was using, and people are paying 10 bucks a month for it, you can't break all their sites and uplift them, you'll lose the customers. How do you solve that problem? And this sort of gets at one other point I wanted to talk to you about. You have entered technology as a profession, through what I would classify as a non traditional path.

Christina: Definitely.

Corey: And everyone I talked to at this stage of the game, seems to have come from a few different pathways that are increasingly closed to people who are entering the workforce today. Well, how do you get really good at cloud? Well spend a couple of decades building out data centers and things like that, oh, spoiler, you're going to be woken up in the middle of the night and have to fix something at 2:00 AM while crying, and it builds character, because character is what you get when you don't get what you wanted the first time. And you iterate forward from there, well, today, you can leapfrog over that entire thing. My daughter is two years old and I'm not going to be teaching her Q basic, as how to use a computer. She instead turns and talks to the robot lady who lives inside of the phone, or the speaker on the counter, or something and that is her interaction with technology.

And how do, I guess not that I'm trying to plan for the two year old case, but people who are graduating college today in 2019, or considering career in tech, where does the next generation of cloud oriented developer types come from?

Christina: I think they come from people, the way they've always come, from people who tinker, people who play with things, people who are getting excited about things. It's just the way you're tinkering that's different, right? I mean, I think we were kind of talking about this before we started recording that, I think it's interesting to think about the kids that are graduating from college now, because public calls have officially been a thing for long enough, they're literally like a cloud first generation. And the training and the teaching has gone level two. So they're not going to have the experiences dealing with the the old, shared, web hosts, or dealing with having to maybe like rent their own servers, or co locate, or whatever.

They've always known a world where the cloud has been what they use, what they deploy to, what they develop on. How they deal with any of their networking stuff. And it's always been flexible, they can always add more power and, or take things away. And I think that's interesting, because in some ways that's kind of forced the rest of the world to maybe catch up. But I also think that it's exciting because those are the people, like when I look at like Visual Studio Code, which is, even though I didn't work at Microsoft, I would be such a huge fan of that product. I think it's such an amazing editor.

Corey: It is incredible, and the one thing I'm waiting for, is I want it to work on my iPad Pro.

Christina: Yes. Which hopefully with the preview that they announced this week of the online hopefully-

Corey: They're dancing around it, there's-

Christina: I know.

Corey: ... just one missing piece.

Christina: I know, I know. Because I want it on my iPad Pro as well, that's the dream. That's ultimately the dream. And I think we're getting super close, especially with some of the remote stuff that they're doing. Like we're almost there, but-

Corey: I'm old and grumpy, I'm an old grumpy UNIX sysadmin, which we call sysadmin because there really is no other kind in that world. And my editor of choice is Vim. I've been using it that way for a long time. But oh my stars is Visual Studio Code pretty, it just works. There are things I don't have to worry about and-

Christina: Right. And but my point too, is that, you look at the ecosystem around it, you look at the people who are building some of the best, most inventive plugins, you look at people who have even before we announced Visual Studio online, were doing their own kind of implementations of Monaco, which is the open source, you know source of the browser engine, or the text engine, I guess, for Visual Studio Code, where they're putting that in containers that can be running a web browser. And you will use on your iPad, right? A lot of people doing that work, are like 21, 22, 23 years old. You see so much stuff happening in JavaScript that is in containers in general, but serverless, right? Are young people who are just playing around and excited. So-

Corey: But my interview question still involves, what level of raid is the right one for this particular use case?

Christina: Right, which-

Corey: Why?

Christina: Exactly right, which is wrong. And I mean, I think that that's one thing where we're going to have to catch up on how we interview and what questions we're asking, and how we're assessing people's skills. And I also think sometimes you have to figure out, what skill is important for this particular role? Because it can change and it can alter. I mean, I always think the biggest thing whenever I interview people is, I'm wanting to know not what they know, but what they're willing to learn. And, because to me, that's the biggest thing. I love to learn, I love it. I'm obsessed with learning and knowing as many things as I can. And I'm always amazed of all the things I don't know, you know?

Corey: Looking for curiosity, looking at the, tell me about something you're the best at. And it's great to have those conversations instead of, well, you're super good at X, Y and Z, let me find a crack in the knowledge where you don't know the answer to a problem that I had to look up to write. That's not helpful, that is looking for absence of weakness instead of hiring for strength, and that's philosophically something I've always been opposed to as a hiring manager.

Christina: No, I'm totally in agreement and I hope that we can get away from some of those. You have to do this on a whiteboard to get here. Because, look in some cases that might be valid and completely necessary, and in some cases it might not. Because like for instance my job, it really doesn't matter if I can solve the problem on the whiteboard. What matters is can I communicate with people? Can I help tell the story that's going to help people get what they need to get done? But even more importantly, like you said, can I listen? And can I then synthesize that feedback back to people who can actually make the changes?

Corey: The last topic I'd like to cover with you is, the divide between ops and dev.

Christina: Yeah.

Corey: How do you see that?

Christina: I mean, it's so interesting. Because before I joined Microsoft, I really, I mean, I knew about operations, but I'd never really considered it two separate disciplines. Because, a lot of the people that I know, and people who work at startups, that line ceased to exist a really long time ago, everybody kind of does everything. And so I didn't ever know like how, I guess like church and state for some people it was. But I think it's disappearing over time. And I think it's good for a couple reasons. And I think that it's been disappearing for a really long time, it's just becoming more visible now. And I don't mean that that DevOps is the solution. What I mean is that, I talk to so many operations people who will say to me before I talk to them, "Well you know, I'm not a developer." And then they'll show me all of their answerable or, power shell scripts, and all these things that they're doing, and all their tool.

And I'm like, okay, I'm sorry, this is code, this is a lot of code. And this is doing the exact same things that somebody else would do, what are you talking about, I'm not a developer? And sometimes you talk to developers, who might say, "Well, I'm not an operations person," but and then they have these amazing build pipelines, and they have tests, and all kinds of things going up. So I think we kind of have to get over this idea that, I think that some developers need to stop being so precious and accept that there are all kinds of things that can be code, and not just whatever they think. And I think some ops people have to kind of maybe open themselves up to the idea that, this identity I have, there's nothing wrong with it, but I can do more and I'm actually capable of doing these other things, I don't have to just be responsible for this one piece.

Corey: From my perspective, it's always been a strange reality of our entire industry, that whatever tool you're using today, whatever thing you're great at, assuming that you are more than five years away from retirement, what you will do, by the end of your career will bear almost no relation-

Christina: Absolutely.

Corey: ... to what you're doing today. And I've talked to a number of friends about this, and I'm still gathering research, but I gave a talk with Sonia Gupta, a couple of months back on embarrassingly large numbers, salary negotiation for humans. And someone came up afterwards, because we talked about different folks from different underrepresented groups, and how that was communicated, and how negotiation strategies differ. And someone said, "That was great, but I didn't notice you talking about ageism at all in this."

Christina: Totally.

Corey: And first, holy crap they were right. But the follow up question, I don't have an answer for this is, how much of ageism is based on actual bias and based upon things that have no bearing in reality, versus how much of that is based upon, well, I've been doing this for 40 years, and I continue to insist on writing Perl for Solaris, at which point that the industry has moved on and the jobs simply don't-

Christina: ... those jobs simply don't exist. Yeah, no. I mean, I think it's probably a little both. I think more of it is probably bias, but I think some of it is people. Because I talk to a lot of IT pros who are trying to, they're freaked out by the cloud, they're freaked out by their jobs going away. I totally understand.

Corey: I was that person.

Christina: But yeah, you've transitioned and you've seen the opportunities, but you've taken like, you're the perfect example to me. Because you took the stuff that you knew as an IT Pro and translated it into the cloud. And let me ask you genuinely, how different are the fundamentals really?

Corey: The fundamentals are incredibly important. Learning the fundamentals required exposure to environments that don't exist in the same way today. But I started off my career being very deep in the world of large scale email administration. Unless I want to work for maybe five companies, that's not really something everyone needs to do.

Christina: But I guess my point though, is that those skills though, like how hard was it for you to then transition and maybe apply those things to what you do and to setting up maybe a large scale email organization in the cloud?

Corey: It was applicable in the sense of going back to the idea of the T shaped engineer. Where you have to be effective in something approaching an operations style role. You need to have a broad baseline of relevant experience. And then for marketing and self promotion purposes, you want to have one or two areas in which you're deep. And I started off being very deep on email, I then transitioned into configuration management when it seemed like that was not going to be a longer term success play. And that seemed like a great play at the time, and then this whole industry sort of pivoted around immutable infrastructure, and that writing was on the wall. And my most recent pivot as an engineer was into this world of cloud, and only somewhat recently have I realized that if I have to pivot again and do something else, I'm almost certainly not going to find my next role as an engineer somewhere anymore.

I've turned into something different that's really hard to define but, honestly the answer is, I'm not happy sitting there writing code for a living anymore. I used to love that, and now I find that I want more and spoiler, my code's really not that good.

Christina: No, you want to at the analysis, you want to look at the big picture things, you want to talk about those trends, which I think we need that and that's important. But I guess what I was trying to get at is, when you made kind of that transition from doing kind of the old world things to going into cloud, I mean, obviously, it takes a lot of work and I'm not trying to say that you can do it overnight, but it's not as if you had to learn a brand new language.

Corey: Oh, absolutely, it's a series of half steps. It's the same way as when you talk to people in their career and they, for example, you were a journalist before this.

Christina: Yes.

Corey: And now you moved into, effectively what is a technical role.

Christina: Yes.

Corey: And-

Christina: I'm technically a content engineer, so yes.

Corey: Exactly. And to do that, it was a lateral shift from taking things you were good at in your previous role, and applying them in a new context to the role you're in. It wasn't what a lot of people think of a career transition being of, well, now you have to go back to square one and take an entry level job, and a massive pay cut and-

Christina: Yes. Well this is kind of my overarching point, right? And I think that this is what sometimes I think people of all ages, this isn't just saying to the gray beards, right? This is something that people need to know that you don't have to start back at ground zero. If you are an IT pro and you are having to move to the cloud, you are not starting from scratch. Is there a lot of things that are going to be different and that you need to learn? And maybe that you need to brush up on? Sure. But a lot of things, it's not as if you're starting from from nothing. And a lot of the things that you do, and the experiences you have will come in handy, and will be applicable when cases you don't even know it.

Case in point for me. When I was interviewing at Microsoft, my whiteboarding experience was, how do you break down a complex problem? And I used an example of a story that I'd recently written that was incredibly complex, and had a lot of moving parts, and that I had had to learn about in very, very, very little time. And then I had to not only learn about it, synthesize it, then I had to write it down in a way that the audience could understand. It was a complicated, ridiculous thing. When I was in the process of that exercise I was realizing, oh, I can do this, this job I can completely do. But it took being in the interview room, and doing that process for me to realize, oh, these skills that I had that I didn't think would be applicable here, are completely applicable.

And these things that I already know, it doesn't mean that I don't continue learning, and I don't continue improving, because obviously you do and that, you should be doing that even if you were writing the same Perl code in Solaris for 40 years, you should not be, in my mind, you should always be curious and always be trying to figure out, well how do I write this better? How do I optimize this better? What else is going on? But it does mean that, I think there are a lot of things that people already have in their knowledge base, that are going to give them a step up if anything. And so does not mean afraid of, just because it's different means that I'm zero.

I think you put it exactly perfectly. I'm not starting from ground zero, I don't have to start school again, and take a lower wage job, and go to another place. It's like, no, I just need to pivot and think about how I can apply these things in new ways, and also take the time to learn more, and to continue growing. And I think the big thing too a lot of people don't realize is that, when you join a new role, a new company, nobody is really expecting you to be operating at 100 the first day. You're allowed to learn and people are going to train you, and people are going to help you.

Corey: I have a friend who joined Azure after being deep in the AWS weeds, and I was speaking to this person, and they mentioned that they've been there for six months and were still learning a lot of the Azure stuff, despite the fact that this person has forgotten more about AWS than I will ever know. And it's, you're always learning, there's always a ramp. And being able to do the entire job written in the job description means it's probably going to be a boring job. There needs to be some stretch room.

Christina: Completely. I mean, that's something that, women especially, often we look at job descriptions, and studies have shown this, and we think that we have to have all the requirements. But you're exactly right. If you have all the requirements, you're overqualified.

Corey: Oh, absolutely. My first job, I looked at the list of requirements that they wanted for a UNIX admin, and I'm looking at that going, I don't have any of these second half of the list. And then I found out what the pay range was, and it was effectively a seventh of what it would have been if you were someone who had all of that stuff. And looking back now 15 years later, I still don't have everything on that list, it's an aspirational shopping list, not a whole list of hard requirements.

Christina: Definitely.

Corey: So if people want to hear more about what you have to say, where can they find you?

Christina: So I'm on Twitter, I'm @film_girl on Twitter, and I'm on Instagram in the same handle. I do stories from time to time, not too often.

Corey: You're on the gram.

Christina: I am on the gram. But I actually do a weekly tech news podcast called Rockets at relayfm.com/rockets, and that one's more consumer tech focused. We, and a little bit of pop culture. But it's really fun.

Corey: Thank you so much for taking the time to speak with me. I appreciate it.

Christina: Corey, thank you for having me.

Corey: Christina Warren, senior cloud advocate at Azure, I'm Corey Quinn, this is Screaming in the Cloud.

Announcer: This has been this week's episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com or wherever fine snark is sold.

Announcer: This has been a HumblePod production. Stay humble.

View Details

About Mitchell Hashimoto

Mitchell Hashimoto is Founder and CTO of HashiCorp. He is the creator of Vagrant, Packer, Serf, Consul, Terraform, Vault and Nomad - a set of open source tools that each individually are downloaded and used millions of times per year. At one point Mitchell was in the top 5 most active users on GitHub. At HashiCorp, Mitchell is helping define a remote-first culture with over 500 employees spanning dozens of countries. He loves open source, automation, and working from home.

Links Referenced:

  • Twitter: @MitchellH

Transcript

Announcer: Hello and welcome to Screaming In The Cloud, with your host cloud economist, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world and ridiculous titles for which Corey refuses to apologize. This is Screaming In The Cloud.

Corey Quinn: This week’s episode is generously sponsored by Digital Ocean. I’d argue that every cloud platform biases for different things. Some bias for having nearly every feature you could possibly want as a managed service at varying degrees of complexity. Others bias for, “Hey! We heard there was money in the cloud and we’d like it if you would give us some of that!” Digital Ocean is neither. From my perspective, they bias for simplicity.

Corey Quinn: I wanted to validate that so I polled a few friends of mine about why they were using Digital Ocean for a few things, and they pointed out a few things. They said it was very easy and clear to understand what you were doing and what it took to get up and running when you started something with Digital Ocean. That other offerings have a whole bunch of shenanigans with root access and IP addresses and effectively consulting the bones to make those things to work together. Digital Ocean makes it simpler. In 60 seconds they were able to get root access to a Linux box with an IP. That’s it. That was a direct quote except for the part where I took out a bunch of profanity about other cloud providers.

Corey Quinn: The fact that the bill wasn’t a whodunnit murder mystery was compelling as well. It’s a fixed-price offering. You always know what you’re going to wind up paying in a given month. Best of all, you don’t have to spend 12 weeks going to cloud school to understand all their different offerings. They also include monitoring and alerting across the board and they’re not exactly small-time. Over 150,000 businesses and three-and-half million developers are using them. So give them a try. Visit do.co/screaming, and they’ll give you a free $50 credit to try it. That’s do.co/screaming. Thanks again to Digital Ocean for their support of Screaming in the Cloud.

Corey: Welcome to Screaming In The Cloud. I'm Corey Quinn. Once upon a time in 2014, I was a traveling trainer for a configuration management system that need not be mentioned here. I was on the way back, and the person next to me on the plane saw my shirt. Oh, great. A plane talker. Doesn't everyone love talking to one of those? He started talking about a system he was in the process of building and about to release that would handle infrastructure management at scale programmatically. Cool, great. We've all heard those stories before. I didn't see it going too far, but wished him well. Now we're having this conversation almost five years later. Mitchell Hashimoto, founder of HashiCorp, author of Terraform and a bunch of other things, welcome to the show.

Mitchell: Hey Corey. Nice to sit next to you again, virtually.

Corey: Oh yeah. I will freely admit I was wrong on that one. It's turned out that you can have a 98 or so percent success rate by talking to people who are building something new and saying, "Oh, that'll never work." Then you're going to get it hilariously wrong for conversations like this. But by and large, it always seems like being a pessimist is a viable strategy or a path forward.

Mitchell: It's a great way to lose nothing, but also to gain and very little too.

Corey: That's the fun thing of being a futurist, too. You can always make outlandish predictions. If you happen to be right, you'll be lauded. If you're wrong, well, everyone's wrong and no one really holds it over your head anymore.

Mitchell: Yup. Yup. To be fair, I wouldn't have been optimistic about me either, from the third person.

Corey: It seems to have worked out though. You folks have now become a name brand in the space. It's been years since I had to mention what Terraform was when giving a talk around this stuff.

Mitchell: Nice.

Corey: That said, for those who are unaware ... Because not everyone knows everything, as it turns out, can you describe what Terraform is for folks who may have never heard of such a thing?

Mitchell: Of course, yeah. So Terraform is quite literally infrastructure as code. So you describe servers, switches, DNS records, anything you would imagine. I like to say anything that would be in a "data center" to run an application. You put it into a text file, you tell Terraform to make it for you, and it does that using ... By stitching together a variety of APIs from cloud providers and SaaS providers and so on.

Corey: For those of you in the AWS ecosystem, you can think of it almost like CloudFormation done properly, which is, I'm sure, a battle that we will get into later and no one would ever disagree with. Meanwhile, you just hear people starting to seethe in Seattle at this point.

Mitchell: Yes, yes. I won't say you're right, but let's see where this goes.

Corey: Absolutely. But first, let's start with something else that'll irritate people. I noticed a while back that you had a post on Reddit, which is always a great place for a conversation to start, talking about multi-cloud. Given that I've spent roughly 18 months now talking about why multi-cloud as a best practice is almost always the wrong decision to go in, I would love to get your thoughts on this.

Mitchell: Yeah, yeah, so multi-cloud has been ... Let me start by saying that the HashiCorp's sort of core mission or vision has been enabling people to adopt multi-cloud. So that has the underlying assumption that multi-cloud is real. For a number of years early on in the company, that was pretty laughable, I would say. When we founded the company in 2012, I think that it was pretty common for people to think that there was no possible future reality, other than a fully Amazoned world.

That has changed over time, at least from my biased point of view. But yeah, what I always tell people is sort of like I've learned ... I bet this early on, but I've learned in practice since then that multi-cloud is somewhat inevitable. It's kind of like entropy in my mind. It's like you could fight it, and you could fight it, and there is good reasons to fight it. I won't disagree with you there, but at a certain point it's inevitable.

So if it is an inevitability, then how do you best equip your business or organization to be multi-cloud or ready even if you're not? So I want to make it super clear that I never tell people, and I don't advocate for people being multi-cloud early. I think you should 100% focus on a single cloud, get really good at it until you're just forced away from it for a good reason. But while you're single cloud, you should just try to adopt tooling that will make that a little bit of a smoother path.

Corey: Before we dive into this and potentially start tilting at the same windmill from opposite sides, it feels like we should define terms. When you hear the term multi-cloud, what does that mean from your point of view?

Mitchell: Yeah, that's really important. So for me, it means what I like to call workflow portability. I define multi-cloud in my mind as four different things. It's workflow, portability of work load portability, data portability and traffic portability. I've experienced when most people say multi-cloud, they immediately snap to work load portability, which is sort of the idea that any application you write could run on any cloud. For me, that's absolutely not the case.

I think you should take advantage of all the best-in-class, first-class services that make applications explicitly not portable between clouds. What is more important to me is workflow portability, which is sort of the process by which you get infrastructure, deploy applications, secure them, network them. Those sorts of core things should be multi-cloud.

Corey: I would agree with you that on a per-workload basis it tends to be something of a red herring. I would also say that, "Oh, you should never use a second provider for anything", that's nonsense. I mean, easy example of this is PagerDuty. There's not a whole lot else like that in the market and certainly none of them from any one large cloud provider.

I would also strongly suggest people use ... Even if you're on AWS, I'd still recommend you use GitHub, which is a Microsoft product, or Office 365, or G Suite. There's always going to be higher-level differentiated services that make sense to consume. I think that people tend to hear multi-cloud in conversations I've been in, and where their minds tend to go is almost a VM-centric, where you're coming from an on-prem environment where you've got this thing sitting running in VMware most likely. I should be able to seamlessly have that flow to AWS, to Azure, to GCP, to Oracle Cloud in some very strange conversations, and that nothing else should have to operate around that in any way. I just haven't seen that in the real world.

Mitchell: Exactly. Yeah, and I've seen ... I think I alluded to this in that Reddit comment, but I've seen the opposite now, where it's been long enough with AWS in existence that I have now seen a company go from startup to public pure AWS, like actually AWS poster child in a way, and then be forced to adopt another cloud and the almost tragedy that fell out of that. I think that's very interesting.

The abridged form of that story is really that there is a darling startup, which I won't name, but it's real. They went 100% into AWS, blogged about how they automated the heck out of it, and really good thought leader at AWS became a multibillion dollar public company, full AWS again. They ended up acquiring another company for a couple of hundred million dollars, and this other company was pretty large themselves. I mean, if you're getting acquired for that much, they're pretty large themselves, pretty successful. 100% on GCP.

It became this thing where ... I've talked to the executives, and it was one of those things where you're not going to block an acquisition that's good for the business of a public company based on their tech stack. Like it's just not going to happen. If the strategy makes sense, the partnership makes sense, it rarely is going to happen. So they decided to acquire this company no matter what. They bought them, and then what ended up happening was they basically churned their whole infrastructure team. Pretty expensive. It was either convert everyone to AWS, which they were violently against and the people internally were against from a cost perspective, or just figure out how to do multi-cloud.

Again, this wasn't workload portability. They didn't want to move anything over to AWS, but just how do you do governance? How do you do cost reporting? How do you do like these higher level things across clouds? They ended up basically rehiring an entire infrastructure ops team to handle this, and it was messy.

Corey: It's a form of lock-in that people generally don't tend to think about. It's not just technical or architectural, but if you have an entire team of people that are great on cloud A and then, "Oh by the way, we're migrating things to cloud B as well", very often people don't want to do that. They'd rather continue to focus on their area of expertise rather than start to dilute that across multiple providers and confuse themselves. There's also any higher-level service story starts to break down.

I think that when there's a compelling business reason to go multi-cloud, I'm all for it. What I'm agitating for has not been, "Oh, pick a provider and go all-in at the cost of your business." That makes no sense to anyone. But building something Greenfield on day one with this idea that you're going to run it on AWS, or GCP, or Azure, or wherever else you want to put that, maybe you don't do that. Maybe pick a provider. I don't care which one. They probably care which one. And just hope for the best until you have a compelling reason to move.

But trying to build something that can be seamlessly deployed to everything ... It doesn't work. If you do go down that path, you are going to find that you're constrained to some very low-level primitives, because nothing beyond more or less a bunch of VMs and maybe some databases and block storage tends to be the same across providers. Even database provisioning calls tend to start breaking down.

Mitchell: Yeah. So yeah, that's an important part of the way we view multi-cloud. One of the first things I tell anyone that is adopting Terraform for the first time is ... They read CloudFormation, but for anything or they read something like that somewhere on the Internet and they think, "Oh, this is the right once run anywhere dream." But I make it really clear, and I'm going to make it clear right now that it's not that. Like the Terraform you write is AWS specific, or Azure specific, or Google specific. You cannot rerun that and target a different cloud platform. That's not what we're doing. That would be workload portability. What we're doing is workflow portability.

So any configure write for any cloud, it's always the same command to create it, the same command to see what would happen, the same structure to investigate for cost analysis, things like that. It's higher level than the actual tech itself.

Corey: I spoke to a single tenant customer that was trying to build out a deployment on AWS and GCP. They were already a Terraform shop, so they were very conversant with the tool. But getting the networking and the security groups to coexist between providers was massively painful. That is not anything I would consider having added differentiated value. For their business case, it was worth the experiment. Didn't really pan out longer term, but the challenge that you'll see is, at that point when you try an experiment and it doesn't work out, you have to sort of stop riding two horses. Which one do you pick? Invariably, I find that data gravity tends to carry that argument pretty quickly.

Mitchell: Yeah, yeah. I would agree with you completely there. I think that's why applications tend not to move. I mean it's very rare for our multi-cloud believers, so to speak, to move an application or ... Especially because of the data associated with it. What's much more realistic is a new application or a totally different team or business unit decides for really practical and good reasons that they want to use a different platform.

Corey: When I'm doing Greenfield architecture and looking at something that, "Okay, I want to go ahead and build this on a particular provider", I generally don't bake multi-cloud requirements into that. But I will say that I do try to avoid, I guess, strategic architectural lock-in. There's a whole bunch of stories around that, but an easy one might be using something like Cloud Spanner or Dynamo DB Global Tables, just because those tend to be things that don't manifest in the same way as anywhere else. So migration isn't just a painful move a thousand things over. It's also and rearchitecture entire applications data model.

Mitchell: Yeah. But again, I don't worry too much about that because I don't think that your applications necessarily need to move or be portable. It's going to depend on your use case. But if ... Cloud Spanner is quite expensive, but if you look at Cloud Spanner, and it's going to save you a ton of time in order to get a technically correct application to market, then I think you should do it. I think that if your application is successful, it's unlikely that it'll have to be across multiple cloud platforms. So it's worth it. But at the same time, it's just a question of whether you want to be locked in technically to a provider or not.

Corey: I would agree with you, but this question often comes from higher-level strategic business decision makers. The question of, "What happens if Amazon decides to compete with us?" To which my response is always, "What do you mean "if"? Their product strategy is yes."

Mitchell: Yeah, we've seen that more in the recent years. We've seen that more and more, for sure. It's been really, really interesting watching entire industries basically create plans to shift providers overnight due to major acquisitions like that.

Corey: It's always challenging watching people go through that process and, "Well, we've decided that we're not going to put anything on AWS, because some of our customers demand it." You take a step back and you look at some of those customers who are demanding that, and those customers themselves have workloads living on AWS that aren't small. So it's one of those fun explorations of, "Do what I say, not what I do." It feels, to some extent, like a little bit of tail chasing going on in our industry.

It also feels to me, to some extent, like multi-cloud on a workload basis is actively being pushed by a number of vendors in this space. Where if you take a step back and say, "We're going to go all-in on any given provider", well at that point there may not be a story for those vendors who benefit from the multi-cloud approach to have something to sell you.

I'd like to ask from your perspective, given that Terraform's entire ... One of the reasons that Terraform exists is to support infrastructure across all providers. If someone goes all in on a particular provider, do you see that you have something to sell them?

Mitchell: Yeah, so I think we're different because we sit at that workflow level, that even these passionately single cloud companies, we're seeing them use things like Terraform in many cases. Not always, but in many cases. That's because Terraform is a lot more than just the infrastructure. I think you brought this up. It could glue together so many different levels of things. It's DNS providers, it's GitHub source control, it's CDNs, it's things like that that a copywriter alone doesn't often provide all of them.

So Terraform is a good way just to do that. Even if you say all our compute networking storage, like the traditional things are all on a single vendor, but we want to pull the latest commit from GitHub and use that as a way to tag things or something. I don't know. There's all sorts of use cases sort of like that where Terraform is super useful.

So I think that it doesn't affect us as much. We explain the benefits we view of multi-cloud from a workflow perspective, but the ironic thing is we have a lot of paying customers that are both single cloud, and we have a lot of bank customers that are actually not cloud at all. They're all just private data center. And so I think we build technical tools that work with both, even though we think that ... We use multi-cloud is a way to basically explain how workflow portability is extremely important.

Corey: I think that's the narrative that makes sense. I think that's what we're starting to see resonate in the space. I also do want to special case the idea of hybrid cloud, where you have on-premises data center environment that tends to be there for a while. Having a hybrid strategy makes sense. You can also see it as a transitional state. For a lot of these companies, that transition looks like the better part of a decade through no guilt of their own. And that's okay. One thing I want to make very clear is that I'm not trying to be architectural shaming of anyone here.

The idea of, "Well, we built a thing for certain constraints, the constraints have shifted and changed. 20 years later, we have this thing that makes money." We can't just wave our hand and turn it off because something on a whiteboard looks more appealing now. I have a lot of sympathy for that and I don't think anyone has intentionally gone out to make a series of poor decisions on this.

Mitchell: Yeah, and I mentioned it in our keynote at our last conference that ... I'm sorry if this is a negative thing, but I think we've made those transitions a lot longer for a lot of people because our tooling has made the "legacy environments" a lot more productive. So Terraform is actually one of those places we see it pretty often where people that opt Terraform before their cloud adoption story and then realize like, "Oh hey, Terraform works pretty great with something like V-sphere", and "Oh wow, we can automate our entire on-prem set up end-to-end much better. Let's just actually move that fence post of decommissioning this off another couple of years."

So again, one of those ironic things, we've signed a lot of customers where they're saying, "Okay, we're completing 100% transition to the cloud in ...", they'll say three years. Now we're talking to them two or three years later, and they're like, "Maybe we could have done that, but we're happy. We're going to push it off another three years." I'm either sorry or like kind of excited that people could get more value out of what they already paid for.

Corey: Oh, I think the idea of being able to stretch an existing investment further is compelling to everyone. I don't think anyone has any particular problems with that. You see people moving from on-prem to public cloud as well, and it's just almost a comedy of strategic errors that you see getting made along the way. Where first they'll spend eight months figuring out what it's going to cost them. Spoiler, they're going to get most of it wrong, and it's effectively an educated guess. Then they're going to start the process. It's going to take longer than they thought. It's going to be more difficult, and they're not going to save money during the initial phases until they start teaching applications to dynamically scale, which many of them don't.

But it's worth doing in most cases anyway, from a time-to-market and feature velocity story, even if the cost savings won't be captured for a longer timeline than they would expect them to. It feels like there's almost an entire cottage industry that sprung up around making those migrations take longer than they need to.

Mitchell: Yeah. We're still super in the ... It's kind of sad in this way, but we're still super in the beginner cloud usage days. Or I would say the majority of large companies right now ... You brought up auto scaling. I mean auto scaling is still so rarely used. Actually, Microsoft research published a really interesting paper about doing an analysis across, I believe it was Azure, I'm sure. But it's sort of an academic paper. I mean you could read it publicly, but they did this study on how many people are doing auto scaling and what the benefit would be if they did it theoretically. The number was laughably small. If you exclude people that use auto scaling as simple way to just make sure a single VM stays up and running, or a fixed count of VM stay up and running, it's basically zero.

So I think as we get closer and closer to actually getting to more advanced resource usage like that, then it'll get better. But really, it's more or less like the ... The conclusion of that paper and the conclusion I'm getting to is cloud adoption is more or less lift and shift still at the moment.

Corey: One of the problems I have with it is the idea that that's somehow a terrible pattern. Because the alternative is you transform as you go. "Okay, we're going to take this application and the data center, rewrite it to take advantage of cloud-offered primitives and then deploy it." Then it doesn't work, or heaven forbid, it's intermittently bad. Then you get to start going through this entire analytics process. I'm a fan of lift and shift, but you have to follow through on the second phase. Expect that to take longer than you think it will.

Mitchell: Yes, it's going to take a very long time. But I'm saying that optimistically, because if it weren't for ... And I always am super clear about if it weren't for cloud existing, then companies like ours wouldn't have even had a chance. Because the existence of cloud is forcing companies to open their minds to the idea of maybe adopting a new way to do infrastructure, or a new way to think about security or networking or things like that. So if cloud didn't exist, then I don't think something like Terraform would be nearly as popular or Vault. Like the way we do security would be different. I think software is getting better, and the infrastructure around it is getting better, even if we're still in this beginner phase of using what the platform actually gives you.

Corey: I think that you're onto something there. The idea that you can take wherever you are and have tooling that winds up meeting you there and helping you get to a better place or getting more out of what you've already bought is compelling. Now, not to get too annoying for people on this, but let's go with the exact opposite direction of making everything better, and talk a bit about CloudFormation.

So when Terraform came out, CloudFormation already existed. Was it just the multi-cloud story that inspired you to start building something that stood in its place?

Mitchell: No. So, let's talk about the history of Terraform, because there's a truth in the history of Terraform ... Which I'm sure I've said it somewhere publicly, but I'm not sure if I have, which is that when we first made Terraform, we actually didn't see the multi-cloud use case. We built Terraform as what we felt was just a better way to do infrastructure as code on single providers at a time.

It wasn't AWS only or anything. It was either you wrote Terraform code that worked just with AWS or just with Google and couldn't overlap them, basically. The multi-cloud, like multiple providers and a single configuration sharing data and stuff, that was a coincidence prior to the release of Terraform, which is fun. I think that coincidence was the thing that made me actually realize like, "Oh, I think we created something really cool here", but that wasn't actually the original motivation. The original motivation was that I felt we could do a better job.

Corey: Taking a look at it now, I would argue that by any objective standard, there's a strong argument to be made that you had. We still see scenarios where Terraform providers for new services come out before CloudFormation does, which is just ridiculous to me. I don't pretend to understand what goes on behind that. But very often, people I've seen go directly into a Terraform world for pure AWS environments, just due to a wide variety of factors, it all distilled down into it offers things that CloudFormation simply doesn't.

Mitchell: Yeah, I mean Terraform definitely offers things that CloudFormation doesn't. I'm providing resources. I think the open source community is a big help in making us go faster, but at the same time, someone did an objective analysis of CloudFormation versus Terraform and found that it was more or less a wash, in terms of resource speed. There was a bunch of CloudFormation was first at, there was a bunch of stuff Terraform was first at. It was kind of a wash.

But yeah, I mean, there's a bunch of other features. Like the ability to run it as a CLI on your own machine and not use a dedicated service. For a long time, Terraform Plan was completely unique. I think CloudFormation has something like that now. But I think the biggest thing overall is, I think, the ergonomics of the language are better in a lot of ways. I think that the fact that you could glue together different providers and enhance it by writing your own providers is extremely attractive. But you know, I'm speaking from the most biased point-of-view.

Corey: Well, absolutely. I mean you even got support for a CloudFormation provider to my understanding, where you can take any arbitrary CloudFormation template you want and do that within Terraform.

Mitchell: Yeah, and there's practical pragmatic reasons for that, which is you're transitioning to Terraform and you've converted a bunch of it, but you still have that one CloudFormation template that you have. We'll run it for you.

Corey: Oh, it also absolutely empowers my preferred developer model, which is copying and pasting directly from stack overflow, which is why I consider myself a full stack overflow developer. It winds up making a fair bit of sense for me.

So one decision you made early on that I would love to ask you about is you picked writing your own domain specific language for Terraform-

Mitchell: I knew it was going in this direction.

Corey: Right. At the time, CloudFormation was only JSON. Now it supports JSON or YAML, but I'm holding out for XML, just to ensure that I'm irritating everyone with this question. Why a DSL and would you do it again?

Mitchell: Yeah. So this always hits different people with different cords. So one thing to, I think, put into the context about the history of our company and our tools is that we've tried both ends of that spectrum. So Vagrant was the first thing we made. It's almost 10 years now, at the time that we're recording this, and it was all Ruby DSL, like Ruby config.

So you can do anything if you're proficient with Ruby in a Vagrant configuration. Some people love that, and some people hate it. Then, before we go deeper into that, then we did Packer after that. Because of what I learned from Vagrant, and sort of the sharp edges around Vagrant, I decided that the obvious solution with Packer was to use a pure data structure language. So I use JSON. JSON, a lot of people love that. The idea behind JSON was any modern language can generate JSON.

So if you wanted to use Ruby, you could actually use Ruby and just generate JSON. There you go, you have your DSL. That's not strictly true, but it solved a lot of problems. JSON was good, but then I realized that there's a lot of people that hated it. The joke I tell people is that there's ... About once or twice a year, there is a single week where I'll get two emails of someone saying, "I love HCL", which is the name of our config language in Terraform. "I love HCL." Then there'll be someone who says, "I hate HCL", or someone that'll say like, "I love Ruby", or "I love Vagrant's Ruby config", and then someone else that says, "I hate Vagrant's Ruby config."

I think when you get these totally split spectrum points-of-view, what it represents to me is that it's an opinion. There's something here that you can't really win. When we took a look at HCL, Terraform's language, we were trying to balance these two things as best as we could. So what we did with HCL was the actual native HCL syntaxes as human-friendly as possible. It's meant to be human-readable, human-writeable. It's not meant to be machine-friendly in a big way, but what we did was make it a one-to-one mapping with JSON.

So there is an exact JSON syntax that every HCL could make maps to and vice versa, and that is super useful for machines. We did this thing. I think you can't really win. We get some people, like I said, that still love it. We actually ran some like independent studies across users who had never heard of us, users who had used some parts of our tools, users who have used, all of them. HCL performed the best in terms of what people liked the most. But yeah it's a hard thing to say like right or wrong.

Corey: If nothing else, you've given people something to focus on that isn't core and central to the ethos of building the tool. It's, oh no matter what you pick, someone is going to have a problem with it somewhere on the internet, loudly and usually not particularly coherently.

Mitchell: Yeah. Yeah, and the problem is like even if they direct their hate, or whatever it is, towards me or whoever for doing this abominable thing, I understand the frustration because you know I prefer some things to other tools. It's there but it's hard to satisfy everyone.

Corey: Well, last question around the, I guess, glorious dissatisfaction of various people. Originally when Terraform came out, it was coupled fairly tightly with Packer console, Nomad at the time. This was all wrapped together in an enterprise offering called atlas that you folks offered, which was fantastic, having played with it a couple of times. The challenge that I saw with adoption, when I was talking with various clients back then, was always that it required an all-in investment, where suddenly you're using all of the HashiCorp Stack or else you're sort of on your own.

Given that you no longer sell or reference Atlas, I have to assume that that wasn't just something that came out of my own experience. Can you talk a little bit about what that tight coupling looked like, and I guess why it isn't around anymore?

Mitchell: Yeah, it's gone. The idea was that ... What's the best way to build a business out of a product? It's like ask the customer. What do they want? If you give them what they want, the theory is that they'll pay you for it. That's the dot, dot, dot, and then you're on step three profit and you're happy.

So we did that. We asked the customer what they wanted. The problem was that the "customers" we're asking were HashiCorp fanatics. When you asked that segment of user what you want, they are going to say, "We have all these HashiCorp tools. We love them. We would love if you just tied them together for us to make it work." So we did that, and we had a small group of pretty happy users of what we built.

The issues really came around with the not-fanatics, with the people who either didn't know who we were or adopted like one or two things. Because exactly like you said, there was no incremental adoption, and even worse there was not really incremental pricing. I mean, that's a fixable thing but it just made it hard to reason about. Things were expensive because there was a lot more value that they didn't want out of it, things like that.

So what we realized is what we built as an enterprise product as our business was sort of counter philosophical to how we viewed the world in open source. That was a problem, both philosophically and in practice, because our open source point-of-view has always been, "We have six open source tools that do dramatically different things", whereas one open source vendor might have put them all in a single mega tool, we've always been really clear that we split them up because it makes them easier to adopt, easier to learn, easier to interface with other software, things like that.

We completely threw away that philosophy originally for our enterprise products. Saying it now in hindsight, it's so obvious. But what we ended up doing was just breaking it down in that same way, so we could let people adopt an enterprise version of both, separately from Console, separately from Terraform, etc.

Corey: No, I think that it's one of those lessons that only really makes sense after you've either gone through it once or seen someone else go through it. I don't know if it would've been possible to predict that, going in. I mean sure, with the benefit of hindsight.

Mitchell: It seemed like it.

Corey: Everything feels so easy to predict in hindsight, but in the thick of it, you never know quite how the market's going to react, how the market's going to shift. I mean, at the time we met, I was speaking largely around configuration management. It seemed like that was a "full steam ahead" business unit. Now, it seems like the entire industry has taken a turn toward immutable infrastructure. I can either sit here, and shake my fist at this, and refuse to change, or I can find something else to be good at that isn't imperative or declarative configuration management. Instead start looking at something else. Infrastructure as code has clearly carried the day.

Mitchell: Yeah, and I was actually, coincidentally prior to recording this, I was just talking to another founder who has pretty cool, very hip, very edgy technology. One of the things I was telling that founder was, "Don't try to be revolutionary in too many things, because if you attach yourself to something that's already like rocket shipping up, their marketing dollars lift you up too. But if you fight it and you say, "No, our way is the true way", then you're just fighting something that you don't have the resources to fight.

I think we didn't try to do that, but that's what we were happening to do, was saying, "Oh yeah, you adopted these three tools but really the one true way is our mega platform." Cast aside the fact that it worked, and it was pretty cool, and things like that ... Of course, a lot of people didn't work but for the right type of user, it did work. But even if you throw that aside, it didn't matter because we were fighting what the rest and sort of the industry was telling people to do.

You kind of have to be pragmatic about that and say, "Okay, they're going to adopt this one thing we have because they see the value in that, even though we want them to use something like Terraform, let's say they're going to use something like Azure Resource Manager or something. We just have to be okay with that." So we switched to that point-of-view.

Corey: The one thing that continues to stick in my mind a year and a half after I did it was after the first year of running my newsletter, I wound up conducting a reader survey. I got almost a thousand responses and asked people, "What are you using to manage your infrastructure?" CloudFormation and Terraform were within the margin of error of being equal to each other.

Now in hindsight, I could have constructed the survey a bit better, because let's not kid ourselves. An awful lot of shops are using both in different ways. But it did turn into very clear this is as widely deployed among an AWS user base. It is large enough to be statistically significant, that this is something that's real. This is not just some random project on GitHub that someone stumbled over. You've tapped into something. You're seeing adoption across the enterprise and in startups as well, and everything in between.

Mitchell: Yeah, and I already shared this publicly, but I won't name the exact cloud, but there's a major cloud provider where they've shared with us that Terraform drives a double digit percentage of all infrastructure that they spin out. That's not just Compute, that's like the databases, the load balancers, things like that. It's not just statistically significant, that's financially significant, that's industry significant and that's really cool. But I think that goes to show that there could be an independent infrastructure as code thing that can rival the first-class, first-party solutions that the cloud providers provide.

Corey: That's just a fantastic story. It's nice to see that even in this world where it feels like the cloud providers are eating everyone's lunch, that it's provably not true. You just have to offer something that the market needs.

Mitchell: Yep. Yep, and I think we did that, and I hope so.

Corey: It would seem that the entire industry has agreed with you at this point. Mitchell, thank you so much for taking the time to speak with me today.

Mitchell: Yeah, nice talking to you.

Corey: If people want to hear more, where can they find you?

Mitchell: I don't hide. Just hashicorp.com for our company or @MitchellH on Twitter. I tend to respond to as many people as I can. I don't hide my email, so you'll figure it out through those avenues.

Corey: Sounds good to me. Mitchell Hashimoto, founder and CTO of HashiCorp. I'm Cory Quinn. This is Screaming In The Cloud.

Announcer: This has been this week's episode of Screaming In The Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod Production. Stay humble.

View Details

About Sean O'Dell

Sean is a troublemaker living on the bleeding edge of technology and innovation. As a member of the VMware Cloud Services - Solution and Technology team, Sean is responsible for Evangelism, Developer Relations and assists in many GTM functions. Sean joined VCS in February of 2017 and helped shape and launch the set of SaaS solutions at VMworld 2017. Prior roles include Global Technical Lead for Network Insight (vRNI), Sales Engineer Leader for Arkin, VMware Cloud Management SE and CUSTOMER.

Links Referenced:

  • Twitter: @theseanodell
  • VMware

Transcript

Announcer: Hello and welcome to Screaming in the Cloud with your host cloud economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on this state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey Quinn: This week’s episode is generously sponsored by Digital Ocean. I’d argue that every cloud platform biases for different things. Some bias for having nearly every feature you could possibly want as a managed service at varying degrees of complexity. Others bias for, “Hey! We heard there was money in the cloud and we’d like it if you would give us some of that!” Digital Ocean is neither. From my perspective, they bias for simplicity.

Corey Quinn: I wanted to validate that so I polled a few friends of mine about why they were using Digital Ocean for a few things, and they pointed out a few things. They said it was very easy and clear to understand what you were doing and what it took to get up and running when you started something with Digital Ocean. That other offerings have a whole bunch of shenanigans with root access and IP addresses and effectively consulting the bones to make those things to work together. Digital Ocean makes it simpler. In 60 seconds they were able to get root access to a Linux box with an IP. That’s it. That was a direct quote except for the part where I took out a bunch of profanity about other cloud providers.

Corey Quinn: The fact that the bill wasn’t a whodunnit murder mystery was compelling as well. It’s a fixed-price offering. You always know what you’re going to wind up paying in a given month. Best of all, you don’t have to spend 12 weeks going to cloud school to understand all their different offerings. They also include monitoring and alerting across the board and they’re not exactly small-time. Over 150,000 businesses and three-and-half million developers are using them. So give them a try. Visit do.co/screaming, and they’ll give you a free $50 credit to try it. That’s do.co/screaming. Thanks again to Digital Ocean for their support of Screaming in the Cloud.

Welcome to Screaming in the Cloud. I'm Corey Quinn. I feel the need to point out, at the beginning of this show, that guests are always invited to speak. Sometimes people suggest themselves, but it's always something that's based around whether or not it makes sense, whether there's a story I think is there. You cannot buy your way into being a guest on this show. Bearing that in mind, please welcome Sean O'Dell, developer and cloud advocate from VMware. Welcome to the show, Sean.

Sean: Thanks Corey. Glad to be here and I'm looking forward to the banter today.

Corey: Absolutely. And the reason I feel the need to caveat your presence here is that lately VMware has been on a little bit of a buying spree. You've acquired a bunch of interesting companies that we'll get to a little bit later and it's very strange. It seems like a company that is moving very determinedly in a particular direction, but it's not clear what that direction is. So, what is VMware and what do you folks do?

Sean: Yeah, absolutely. So VMware, historically ... right? I think everybody knows VMware is vSphere or ESXi, right? And it's prime 86% of the hypervisor market. You know, VMware owned it. Like it was the highest number. Kind of the crazy number that we saw. And then, from that, we built an ecosystem around, obviously with plenty of partners. But then obviously VMware went into the hybrid cloud management space with the v realized stack, went into the end user computing space in mobility with Airwatch and some of our, you know, our own solutions. Obviously, Airwatch was another acquisition we made. And then I think kind of the two pivots before ... or I guess maybe the last pivot before we talk about what we're going now is, into the networking and security space, right, with the acquisition of Nicira and a few other things around them. And really, you know, VMware solidified kind of the next generation data center, if you want to call it that, and then ultimately the hybrid cloud.

And now, I think if you go back to VMWorld two years ago and then solidified last year, both in US and Barcelona, it is becoming a multi-cloud company. So no longer focused purely on the vSphere hypervisor. And so, that's a little bit nervy for a lot of folks. But at the same time, we firmly believe, because our customers are telling us this, is that they are becoming multi-cloud companies, right? So we need to be able to support AWS, Azure, Google just as well, or treat them as first class citizens just like we do vSphere. So that's kind of the future direction. And obviously we can talk a little bit about Kubernetes as well.

Corey: So if I take a look back, I've been seeing VMware in the marketplace for about 20 years, give or take. I consider you folks the beast of many products. And instinctively, that sort of got me stuck somewhere where I think of you folks as cutting edge technology from 2002. If I'm being less charitable, I've referred to you as the payday lender of technical debt. And I wouldn't have had you on the show if I thought that this was a truism and there was no value in anything that you were going to say that was not going to change my mind. And even before having you on this, I've started to find that my own view of your company is starting to change. I will admit that you folks do have a serious problem of taking random words and putting the word ... putting the Letter v in front of them.

That's a little, okay ... we're ... I get it. Branding. It's a thing. But let's start with what we ... I guess what you just alluded to. There've been a number of interesting companies that you've acquired. Most recently, you have wavefront, Hepti-I-O, or Heptio, however they want to pronounce it, Cloud Health and Bitnami, which some people mispronounce as Bitnami. So, the easy and facile answer is that you have acquired these things just to ruin them. And you're this decade's Symantec. I imagine, based upon conversations with folks at those companies, that there's something else in mind. What is, I guess, the vision of the future as you see it?

Sean: Yeah, 100% and I would also argue we are not a ... We're not a mainframe company. So, we're not that old in legacy ... No, just kidding.

Corey: No, that would have been the seventies or sixties.

Sean: Exactly. Right? So, you got the 30 year advance, right? No. To be very clear Corey, you know, VMware solidified itself, right as you mentioned, in the, you know, private cloud, hybrid cloud, in The vCloud, if you want to call it that. Right? And I think as an organization, obviously top down, we realized our customers were moving to the public cloud, right? They were moving to Kubernetes or cloud native architectures in a lot of ways. So, it was either stay in the past and only support vSphere, or move to the future while continuing to improve vSphere. Right. And I think we probably won't get into it a lot today, but we now have partnerships with Amazon, Azure, Google to run the vSphere hypervisor in their public cloud. Right? So, that's-

Corey: Well, yes. Your partner's strategy looks a lot like AWS's product strategy. It's a post it note that says yes on it and that's about it. I mean, and again, given what you do, that's not a terrible decision. The challenge that I see is that the selling point for VMware in a modern day cloud, and maybe I'm misinterpreting this, has been that you can take whatever VMs you're running today in your data centers, and without changing a thing, move them seamlessly into public cloud providers. And that is a thing that occasionally is extremely useful as a component of a larger vision, as a transitional step, but rarely as an end goal in and of itself. And it almost feels, to a number of folks out here, that that story is more or less kicking the can down the road yet again on modernizing a legacy architecture that fundamentally will not do well in a cloud environment.

If you have a specific virtual machine that if that goes down, so does your application, that's an architectural problem that is going to become much more pronounced in a public cloud environment than it likely will in your on prem environment. And a lack of awareness around that, or trying to paper over that idea, seems to be doing, in some respects, the industry a bit of a disservice, because people believe what you say, your VMware. Your words carry weight.

Sean: So let me answer this in kind of two parts. I'll actually address the kick the can down the road. Right? You know, I'm not a huge fan of the term. I think I understand where everybody, you know, is attempting to drive with that. But to be fair, our customers are asking for ways to utilize the resources, the knowledge, the historical utilization consumption, you know, understanding of the VMware ecosystem in the public cloud. Right? And that's only part of the picture, right? If we only focus there, sure, maybe you could say, "Well yeah, they're just kicking the can down the road. They're trying to be, you know, the same thing they've been for 20 years and not really modernize."

But what most people don't understand, or I guess we haven't done the best job of articulating, is while we're continuing that strategy, we've invested heavily in this idea of instead of using the vSphere hypervisor everywhere, why don't we provide things like operations management, visibility, cost analysis, security on native public cloud workloads, right? So consuming those eight in public cloud workloads and making sure, really based upon customer desires, needs, concerns, that they are meeting enterprise standards. You know, oftentimes we talk with customers. You know, given one customer specifically, obviously without naming names. But in their case they have a vSphere team or a VMware team that runs their private cloud hybrid cloud on vSphere. They have an entirely separate team that's running AWS, and an entirely separate team running GCP. Right? And in those scenarios, there's no cohesiveness.

There's no knowledge that transfer back and forth, and they end up managing independent systems. And so, we're really kind of at the beginning stages of that. And that's why you see acquisitions like Cloud Health, right? Or Bitnami. And then we ... you know, we'll talk about Heptio as we get into it. But to summarize that, there's really a two prong approach, continue to enhance the hybrid cloud and bringing, you know, on premises data center into the public cloud in that fashion. But secondarily, and just as important, is expanding to becoming a true multi-cloud company, truly seeing all workloads, whether it's you know, Amazon, EC2, rds, red, you know, red shift, whatever you want to name it, Azure AKX or GKE, treating everything as a first class citizen, which means something that doesn't run a vSphere hypervisor all the time. Right?

So, it's definitely a transitional shift and I think you'll continue to see that as we move forward.

Corey: And I'm never going to be sitting here saying that the best approach is to be nonresponsive to customer needs and customer pain. I'm ... I think that if you start down the path of correcting your customers when they tell you they have a pain and need to do a thing, that leads down to a very condescending place that I'm not sure benefits anyone. The counterpoint too, I think that also strengthens your argument, is that a lot of your customers are not Twitter for pets born in the cloud companies that are three years old. These are enterprise companies that predate either public cloud entirely or responsible use of public cloud. And they have very different needs, very different workloads. I ...

When you start talking to that caliber of customer, a lot of my ranting, for lack of a better term, against multi-cloud changes form slightly. When you have a bunch of different lines of business, when you have different acquisitions, for example, running a different cloud providers, you're right. There's very little reason to start merging those to a single provider. That's a lot of work that generates a unclear business value. And until you get clarity around exactly what that value looks like, you probably shouldn't do it. Where I've been ranting about the idea of multi-cloud being dumb is that you take a given workload and be able to seamlessly deploy that workload anywhere. It doesn't care where it lives.

And that, to be clear, is something that VMware is very good at doing. That was, in many respects, something you were doing even in on prem environments where we want to be able to have vMotion and move a VM off of a failed instance onto something else without dropping a packet. And that can be done, which is amazing. And then the world changed. Now, people are trying to do that between providers and it very often doesn't make a whole lot of sense. Because as soon as you start doing that, you wind up in this unfortunate world where you are reduced to using the baseline primitives of any ... of what all of these lowest common denominator offerings tend to be, which means you're not going to be using a message queue that you get for effectively free.

You're going to have to roll your own in a bunch of vms running in VMware and that has to migrate too. And it just leads down a very unpleasant architectural path. It's strange. Because, on the one hand, that is something that your company is ... effectively was built on top of and continuing to evangelize that is in the company's own best interest. The other side is with these ... this new breed of acquisition that you've been doing. You are painting a path to a brighter tomorrow. And it ...

I no longer .... When I hear that someone is considering engaging with VMware as a part of their public cloud strategy, I no longer immediately have the knee jerk reaction to crap all over it. It's, "Well, let's talk about that more. Oh, cloud health. Okay. That's a reasonable path forward. Heptio great, awesome." We don't know what the Bitnami a merger is going to look like ... or acquisition is going to look like yet. I have ... I'm still biting my time on that one.

Sean: Sure.

Corey: But there's a lot of this stuff that makes ... that makes it a lot of sense. It's you're starting to diversify not just your product lines but also your messaging. And that, I think, is one of the most interesting things out of you folks.

Sean: Yeah. And to be fair, right? When you go into this, as you mentioned, you know, the idea of moving a workload between clouds, I would argue that's lift and shift. And in most cases, lift and shift is more expensive in the public cloud, right? When you literally take an over-provisioned workload or a, you know, shrink wrap application and attempt to move it to the public cloud ... which we've seen organizations do it. Right. I can name several off the top of my head. The problem is-

Corey: I'm an advocate for that, but that's a separate argument.

Sean: Correct. But there's literally no optimization, right. Or at least it's the initial stage and then they have to clean up. And-

Corey: Right.

Sean: And it just-

Corey: To be clear, I'm a fan of it as a transitional step.

Sean: Correct.

Corey: Because shifting it while you move it means that now we have no idea where this problem is coming from.

Sean: Correct.

Corey: It's nicer to be able to move it first, take the financial hit, and then optimize as phase two. But you've got to do phase two and it's going to take longer than you think.

Sean: Correct. And so some organizations have decided, "Okay, fine, I'm not actually gonna do a lift and shift. I'm going to do a complete refactor." Well, the initial cost of that tends to be greater than actually doing the initial lift and shift. Right? We can get into ... And by the way, we're speaking in generic terms here, every organization is different. The one thing I would say though is VMware is doing this kind of two step, if you want to call it that, right? We absolutely are continuing to improve on what we've built over the past, you know, 18, 19, 20 years, right? But at the same time, we truly are embracing organizations, right? With ... I'll use Cloud Health as an example. We looked at Cloud Health and pretty much most of the Cloud Health customers, when they purchased Cloud health, it wasn't the traditional VM ware team that was also buying Cloud Health, right?

It was two separate parts of the organization. It was two separate buyers, in some cases, very unique buyers. And so that gives us a completely different view into this public cloud world, right? We now have a better understanding of how consumers ... you know, how many users are consuming the public cloud, why they moved to the public cloud in the first place. And we're not going to take Cloud Health and be like, "All right. Well, let's go prove, you know, if you take all of your Azure workloads and you put them back on vSphere ... "

We're not even doing that. That's not even in the question or in the equation. We are absolutely giving our customers options. If you want to run native GCP, go for it. If you want to run native VMware, go for it. You know, there was some talk that, "Oh, we're gonna, you know ..." that organizations were getting out of the VMware business, or out of the data center business. And that was a couple of years ago. What we've actually seen is those organizations who came to us and said, "Yeah, we're going to slowly transition out VM-ware and maybe move to a 80/20 model or a 75/25 model public cloud versus on prem," they're actually shifting back a little bit, right? Maybe 50/50, 40/60, and so on.

And so, now what we have is the ... the really the desire of VMware. And as I stated earlier, two years ago when we announced this focus on multi-cloud, or the focus to support multiple workloads, that was really that entry point. And we're just two years into this now. And it's truly just starting, right? That's why we ended up making additional acquisitions, whether it's Bitnami ... I love the Bitn AMI, by the way, so keep it going. It's actually perfect.

Corey: Oh yes. Their CEO continues to roll eyes whenever I say that. Their former COO actually tries to hit me. It's great. It just goes great.

Sean: No. And I actually saw Daniel this week. So, just to be clear, the acquisition is final. Pat actually announced it on the earnings call last week. And we're excited to have the Bitnami team onboard. We've got some fun things that we're doing. It really is a portfolio enhancement, right?

Or, helping our customers transition to public cloud. Or even if you want to go back to the idea of moving a workload from on prem to public cloud or between public clouds, having an enterprise catalog to do that makes sense, right? And then you only have to shift the data. But that's way in the weeds. But we are absolutely excited about the direction. And thank you. Obviously, you have plenty of snark and comments about VMware historically. And look, if we can chip away, whether it's monthly or yearly, whatever it may be, to change the perception, we're happy to do it. Right? I think for us the next phase is the developer space or the cloud native space, so lots opportunity to to improve and learn from our customers, and then go from there.

Corey: Largely the cloud native and developer space is something of an untapped market for VMware historically, other than things like VMware workstation or VMware fusion so that you can run different operating systems on your Mac or windows or Linux box. And things like Docker have more or less taken a chunk out of that where that's no longer the workflow that people go with. They aren't going with license fees for enterprise commercial software for this problem anymore. And of course, with the rise of public cloud, click button receive actual instance in environment, which means that even local Dev is something of a controversial approach these days.

Some people want to do their work entirely in a cloud environment. Cool. I don't tend to be particularly prescriptive around workloads. The challenge though turns into how do you parlay the experience you have working and speaking to giant enterprise companies into a relatively, smaller definitionally, cloud native type of environment where they're not cloud migrants. They're cloud natives and they have a relatively small team. What they start working with is likely what they'll continue working with as they become those larger corporate entities. How do you speak to them? How do you reach them?

Sean: Yeah, 100%. I would say this VM ware is a software company. At the end of the day, as an organization, we've been developing software for a long time. You know, one could argue the VMware hypervisor kind of spun this idea hyperscalers and all that good stuff. And so we, as an organization, right ... You know, traditionally we would build software in a very waterfall, I called it aged and that's just my fun term. Because it would take, you know, 18 months to get a brand new vSphere release, right? Well with this transition as becoming a SaaS company and becoming a modern, if you want to use that term, or a cloud native type, we started actually doing that with vSphere, right? We do that in the VMC model, where in some cases we're ... you know, we can push out code every day if need be or every other day.

And then ultimately, that ends up pushing back into our ... you know, back into the on premises and hybrid cloud solutions. Right? So we are a software company. We use a lot of open source software, right? There is no ... I think one thing a lot of folks struggle with in this developer enterprise space is, "Well, you're going to either use all shrink wrap or you're going to use all open source, right? If you're a true cloud native company."

Well, the problem with that, when you get in the enterprise, it's really a conglomerate. And we have organizations that are using part us, part IBM, part Red Hat, part open source in some fashions, right? So, it's finding the balance in all of that. And I think where my team, my organization, to be clear, right? We started the developer and advocacy program here at VMware focusing on developer needs, public cloud needs. And then, obviously we'll bring that on, you know, back to the on prem private cloud. But we need to be able to encompass multiple aspects, right? Everything starts with a pipeline. It should live and die in the pipeline, right? Our group, we run something called cloud journey. And it's kind of our take on this approach of public cloud into the developer agility space, if you want to use that term.

And so, even for us, as cloud journey continues to evolve, our team, very small today, we're trying to do things like bringing cost analysis into the pipeline, right? We had a meet up just this past week where we were highlighting how to take Cloud Health and the cost of an AKS cluster in Azure, and could I provision to it based upon my budget of that particular workload or project and so on? That's really where this changes. And then the second thing that we have, and we've kind of looked at from a ... not only from a continuous verification perspective, but as looking at the security aspects, and making calls to open source solutions like Clair to look up CVEs inside of our containers, or, you know, all the way to, "Let's make configuration changes to our open S3 bucket that may be okay, but it's unencrypted on the back of that."

There's so many things that we get into. What we wanted to do is just start the path and start the journey. And so, as we've done this, really developers internally are starting to provide expertise, right? VMware has run an SRE team for about two years now, right? Internally on our own solutions. So we are a software company, we are going to continue to be a software company and we've gotten rid of some of those pieces along the way that we failed up. But we want to make sure that we're listening to those customers. If it's open source or on, you know ... or a shrink wrap, whatever it may be, we're here to bundle it together and really be an independent third party. Some cases, it may be our tools, and some cases it may not be. And that's absolutely okay. That's a little bit different than, "Hey, you have to run our hypervisor." And so this transition's been a little bit fun and we're just getting started there.

Corey: I think that there's an interesting future ahead for you folks, whether or not you're able to successfully navigate this, I guess, complete landscape transformation is something I don't think we're going to have an answer to for another 20 years. It's ... There's so much going on in this space. And the companies that, well, I guess you deal with primarily as well as the company that you are, is at such a scale that even small changes take forever to wind up percolating through. There's an inertia problem where steering the ship takes years, if not decades. Can you talk at all about what that cultural transformation looks like internally? Is it something where it feels like you're swimming against the tide? Is the company relatively aligned on where it's going next? Is it a topic we don't want to get into?

Sean: No. This is a fun topic. So, this is in a couple parts. I think if you look at it from a leadership, from an investment perspective, right, at the c suite and the broad strategy as an organization ... look, we've been on this journey for a couple of years now, we've made multiple acquisitions. We've made changes internally to that. But the thing I would argue is there are still people buying mainframes, right? So when we get into that mainframe conversation, there's these, what we call, main frame huggers. We had server huggers right before the hypervisor kind of took over. And now, we have vSphere huggers, right? And that's okay. To each his own or her own. But what we have to understand is while there be still some mainframe huggers or vSphere huggers out there, there's actually a growing community of folks who realize that not everything is going to sit on a single hypervisor or inside a ... you know, one methodology from a cloud perspective. Right?

And internally ... obviously leadership, as I mentioned before, but even across the board, right? Teams are fully embracing the idea of workload portability, if you want to use that, to allow for workloads to be on Kubernetes or on, you know, even OpenStack, right? We have one of the leading distributions of OpenStack. And so, we are able to, I think, navigate that in a proper fashion. Are we perfect? No, not at all. But I think we're learning, we're growing. And I think that chasm has already been crossed, obviously from a leadership perspective, from an investment perspective, but even in the weeds, right? VMware has absolutely focused on public cloud and on cloud native where we had not been traditionally.

Corey: So, one other thing that I mentioned somewhat recently that of course set off a tremendous reaction from folks, and I suspect that given where you work, your reaction will be no different than most, was that I wrote a blog post titled Rightsizing your Instances is Nonsense. And in that post, I made the argument that yes it does wind up saving a lot of money, and yes, it's absolutely a good thing to do. But whenever you look at an existing workload, modernizing it to a newer instance family type or resizing the one that you've got is nowhere near as trivial, in many cases, as people like to pretend it is. And I got, "well actually" to death, but I stand by it. Hit me with it. What are your thoughts on that?

Sean: Yeah. You know, I wish you would not have put that out late on a Friday night, or I probably would've commented too. But to be very clear, right? We absolutely work day in and day out with organizations who are doing right sizing, whether that is in the private cloud or in the public cloud. I can say personally, when I was ... you know, six, seven years ago when I was focused on the private cloud, we absolutely ... of the hybrid cloud, the VMware cloud, we absolutely work with customers to improve usability of resources and the consumption of said sources, right? We have organizations. You know, maybe we go back to the lift and shift methodology, where they do need to do some rightsizing, right? There needs to be an improvement. There are some workloads, right sizing doesn't make sense. And I think this really goes back to the customer question.

One of the things that I know I have tried to do and our team and our organization tries to do is just really give customers choice. I mean, I've got example example where rightsizing saves money, or the purchase of RIs or convertible RIs saves money. So, you know, there are absolute technical challenges. I think we could go into those into individual pieces and parts. While I would say rightsizing is not insane or unimportant, I would say it really depends on the scenario. It depends on the organization and the model. But I can tell ya if we could, under NDA, I would absolutely show you some really good examples of our customers do this, how they do it, and ultimately the savings that they gain from it. So, to each his own. If you put me back in an organization where it didn't make sense, I'd probably say it doesn't make sense. I just think that it comes down to options and it really depends on context.

Corey: Absolutely. And that, I think is the entire crux of the argument with less of a click bait title.

Sean: Yes.

Corey: Where the idea being that, yes, there are tremendous financial and performance gains to be had. And when it makes sense, absolutely do it. But if it winds up having to re-certify an application.

Sean: Sure.

Corey: That is a process that was going to take months to do. And it winds up saving you good work at a $300 a month, that's not generally worth doing. It's if you can do it and it's a trivial change over the course of an hour or so, and you just ... you have an auto scaling group that now re-provisions it, and there's no other changes that are workload impacting. Great. Awesome. Go for that. That's absolutely fine. I'm just ... I guess by saying that, what I'm trying to achieve as an outcome is this idea that if you wind up going down this path, just make sure you know what you're getting into.

Sean: Correct.

Corey: It all comes down to context.

Sean: Correct. And look, if you de-couple ... I mean, let me just take the easiness of the public light. If you've done this right, and we'll pick on Amazon, you can change from one class to another. As long as you've decoupled the data and you've not written everything to the OS partition, then you're probably fine, right? It just depends on how it was architected, you know? And in the context, if it is an application that the organization developed, that probably is easier to do this. But if you get into some of the shrink wrap applications where they make sure you ... that dictates certain parameters and installation types and how it was deployed, probably don't support auto scaling, right? That's where you get into that gray area of does this make sense?

And I think, you know, depending on the organization, maybe it does, maybe it doesn't. Right? That comes down to the dollars and cents question that we all should look at in a much bigger picture. But maybe, cause I know my Cloud Health team would hate me if I didn't say it, we absolutely do have customers do it and we're always here to kind of go back and forth, and work with them to find those gaps. It's not just instance types. It could be RIs. That's another area. I mean, even VMware is a customer of Cloud Health. We've saved significantly with RIs and CRIs. And so, it's just getting into the conversation past the clickbait. Clickbait's all good man. That ... I actually laughed at it. And so, you know, let's just get into the context.

Corey: Sounds good to me. I think that a lot of the challenges that we'll see in this space is that, to some extent, coming from the type of company ... like, for example, Cloud Health. Now, that's your problem, your responsibility, where you wind up ...

At least, as I go through talking to companies with large budgets, challenges, the dashboards that get spit out by any of the tooling in this space are effectively working from the same starting point, where they're established APIs, at least in the AWS space that you can wind up pulling this data from. And then what you do with that is up to you. And very often, it serves as a useful starting point, but it doesn't get you there as far as it doesn't provide context. As soon as you wind up with that tool suggesting that, "Okay, now it's time to turn off those idle instances. And what were those? Oh, the DR Site."

Sean: Yeah.

Corey: Well, people start to take it less seriously. Two or three outings like that and people tend to view it as largely useless. That's largely unfair, but it does wind up providing, I guess, a psychological reinforcement that tools are not to be trusted in this space. It lacks context and it lacks business awareness. And to be blunt, I don't know how you fix that from a SaaS perspective.

Sean: Well ... So, to be fair, there is an evolution or a journey that our customers are on, right? Everybody in the industry, you know, unless you're just a cloud native company, like a Netflix, where you've done this, you've built this. I mean, you don't don't need to go back, right? You don't have the history to go along with it. If they're doing it right. And I'll use VMware as an example 'cause we don't mind talking about it. If you tag your workloads properly, if you are doing the necessary Beta data layers to your applications, and when you provision them, if it's terraform or Ansible, you know, whatever solution you're using in the industry or even our own, if you don't tag those workloads, you're already setting yourself back, right? So context comes from the organization itself.

And what we have is we have customers who use our perspectives. And they have tagged properly or they go back and tag properly to provide that context. And then, that's where the slicing and dicing of the data comes in. And I will say it's totally ... It's not only about cost savings, right? We want to get into governance, right? Should Tim on my team be able to provision a $10,000 a day workload in Azure, right? Or should he be able to, you know, make changes to a security group or some firewall within one of the public clouds, right? So, we're absolutely into expand it beyond cost. But the cost conversation is one that always comes up first. And then, we kind of see the security. And then, we kind of get into some other areas around governance and so on. So, this is a journey that everyone is on. If they want help with some of the context, let me know. I can give some tips, right? There's not only our stuff but others out there who attempt to articulate the idea of using the Beta data in your favor. But some organizations didn't do it at the start, right? So they obviously have to do a little bit of homework.

Corey: Absolutely. And I don't think that you're ever going to get away from having to do homework. As much as we want to pretend that it's push button receive cloud. I don't think that there's a particularly clear path to get there. And we can get close, but I don't think we're ever going to get there. And it's almost a question of do we find that answer here? Do we find it there? Do we find it VM ware? And no, I'm not apologizing for that.

Sean: Pun intended. Right? I think everyone in this industry, right? We all want to come in and have our, you know, extremely opinionated view. And we know better than everyone else. But just like any net new technology or any, you know, truly differentiated organization, there's a series of learning and a series of opinions, right? And I think obviously I'm very biased in this fashion. I do think VMware has an opportunity because of our history and because of what we were able to do in the on Prem space, right? We had hardware providers galore that we were able to work with from an ecosystem perspective. But now, we want to be able to do the same thing in the public cloud. Do we have all the answers today? Not at all. Are we getting answers? Yes.

And we're continuing to work and harbor the questions back and forth with our customers and non customers alike. There are some done VMware customers that are now coming into VMware say, "Hey, I've got a challenge in the public clouds. How can you help me?" Right? So, we're here to do that. We're here to help in that fashion. And it's a fun journey to be on so far.

Corey: I would agree with you. Sean. Thank you so much for taking the time to speak with me today. Where can people find more of your wise thoughts?

Sean: Wise thoughts? Well, thank you for the kind words. Thank you for the opportunity to be on here. Follow me on Twitter. I am the Sean O'Dell, @theseanodell. Yes, that's a gay like version, S-E-A-N. And then, find our team, my fearless leader, Bill Shetty, Melinda Psi, teammates, Dan, Tim, Wia and Tree, we're all on cloudjourney.io.

We're about to do a quick refresh of the page here in the next couple of weeks. Its going to be fun. It's exciting. We write things that have nothing to do with vSphere. And we write things that oftentimes have nothing to do with VMware solutions, right? So, that are focused on public cloud, cloud native, and so on. So give us a shout out, reach us tweet us, you know, beat us up. Whatever you want to do. We're here to engage and always exciting to work with the community and just see what everyone's doing.

Corey: Sounds good to me. Thanks once again. This is Sean O'Dell, developer and cloud advocate at VMware. I'm Corey Quinn and this is Screaming in the Cloud.

Announcer: This has been this week's episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever find snark is sold.

Announcer: This has been a HumblePod Production. Stay humble.

View Details

About Corey Sanders

Corey Sanders has 15 years of experience at Microsoft with 13 years of managerial experience. In the last 9 years, Corey has been in the Azure team building the Azure Compute service, and he recently moved into a new role as Corporate Vice President for Microsoft Solutions.

Links Referenced

  • Twitter: @CoreySandersWA
  • Microsoft Azure

Transistor

Announcer: Hello, and welcome to Screaming in the Cloud, with your host, Cloud Economist, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey Quinn: This episode of Screaming in the Cloud is sponsored by N2WS. There are a number of backup solutions that are available in AWS including the recently announced AWS Backup. Well, AWS, back the #$*% up. Backups are incredibly easy. Restores, however are absolutely not. You want to find out whether your backups worked well in advance instead of the way that most of us do: when they don’t work quite right immediately after you really, really, really needed them to work correctly. Check them out at n2ws.com. That’s n2ws.com. Thanks to them for supporting this ridiculous podcast.

Corey Quinn: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined, for a second time, by Corey Sanders.

Corey Sanders: That's right. I think the second time is supposed to be better, but we'll see. No promises.

Corey Quinn: So they tell us. The last time that we chatted a year ago, here as well at Build, you were the Corporate VP of Azure Compute.

Corey Sanders: That's right.

Corey Quinn: You're now the Corporate VP of Microsoft Solutions. What is that?

Corey Sanders: You know what, when I figure it out, I'll let you know. No, just kidding. Well, one important aspect is I moved from the Engineering Product Team into the Sales Team, into the Technical Sales Team, so that's a pretty big shift. I'm responsible for the four big solutionaries that we sell as a company. That includes data and AI, apps and infrastructure, which together become what we know as Azure. But then also Modern Workplace, which is Office 365 and Windows Client and then also business applications, which is our dynamics product. I'm sort of responsible for all of those. That's the new gig.

Corey Quinn: It sounds like it's a lot of work, to be very blunt.

Corey Sanders: It's a lot more work than I want, let's be honest.

Corey Quinn: You're doing Microsoft solutions, other people are causing Microsoft problems.

Corey Sanders: That's right.

Corey Quinn: It winds up sort of this wonderful balance.

Corey Sanders: That's right. I get called in ... I'm sort of the fireman, as it were, but without any running water, so we'll see what happens.

Corey Quinn: This morning there was the Imagine Cup World Championship, where you were on video from the expo hall, which looked like it was 30 feet away from where the rest of us were sitting. One thing that you mentioned was that it turns out as the Microsoft mission statement, which I didn't realize companies still had, but you folks do. "To empower every person and every organization on the planet to achieve more." Now, normally, I tune out on those things, but you were wearing a t-shirt in the video, which is a bit of enough of a departure from what everyone else was wearing. "Okay, I'm going to pay attention to this guy."

Corey Sanders: "He must know what he's talking about."

Corey Quinn: Exactly. And if he didn't know, let him dress like this. I'm giving him the camera.

Corey Sanders: That's right. Yeah, well, the wearing of the t-shirt ... I'll answer that question first and then I'll dig in because I think ...

Corey Quinn: Start with the easy one.

Corey Sanders: Oh, man, yeah. I mean the wearing of the t-shirt, this has become sort of, in some ways, my brand, which is a little bit of a weird thing that you'd day your brand is that you wear t-shirts. But as I moved over from engineering to the sales organization, it was sort of like I continued to just wear t-shirts. And then it became sort of a thing that, "What funny, weird t-shirt is this guy going to wear?" That now I have to wear t-shirts. I don't actually even have an option anymore, because everyone's like, "You're wearing a dress coat. What are you doing, man? You look like an idiot." And so t-shirt all the time.

Corey Quinn: "Are you here to fire someone? What's the story? Oh, you're interviewing for another job? What's going on?"

Corey Sanders: Yeah, exactly, "I hope the interview went well." That's the thing with my t-shirts. I have to continue to buy cool t-shirts. So if you've got any ideas, let me know.

Corey Quinn: Well, today was a plain gray shirt with no logo on it.

Corey Sanders: It was, which is like a marketing thing. I had logoed shirts, and they were all going to risk people suing us, so we ended up wearing blank.

Corey Quinn: Got you. So you fundamentally gave up on whole NASCAR approach to business models?

Corey Sanders: That's right. That's right, that's right. Apparently no one owns plain gray, so that's good. So that's what I wore.

Corey Quinn: Yeah, I'm sure someone out there is currently...

Corey Sanders: Someone out there's already filing suit. So then going back to your question about now that we finally got through that we can talk about the mission statement. I find the mission to be super inspiring. What an interesting mission statement to be so focused on what we can help others to do versus what we ourselves are doing.

Corey Quinn: Right. Other folks' mission statements tend to be somewhat inward looking. Like, "We're going to categorize the world's information and then sell it to people."

Corey Sanders: I have no idea who that is.

Corey Quinn: Or, "We are going to not rest until no one else can make money doing anything."

Corey Sanders: That's right.

Corey Quinn: And it's great. This is a little bit more aimed at helping other people achieve.

Corey Sanders: That's right. That's right. One of actually the more exciting parts about moving in this transition. When I was in the product team, I was out talking with customers a lot, figuring out what they needed, what they were looking to do, how our products could help them. And this shift to the sales organization has been pretty exciting because when you think of sort of a modern sales organization, it is not about, how do I get in there and try and get you to sign this paper and give us a check? It's all about, how am I helping you solve problems? How am I sitting down and understanding what issues you have and how we can bring the right resources from our side to partner with you.

There's been some really exciting examples of this. I think we may have talked about it a year ago with Walmart. Some of the work we did, we've got sort of the Joint Development Center with Walmart. We're doing really cool things with stores and an IoT-based solutions with tracking sort of the health of their refrigerators. There's been a lot of these really interesting opportunities for us to learn more about retail and them to take advantage of some of our deep technical understanding for the Azure services. When you think of the modern era of selling in the cloud, it really is around enabling and less selling, if that makes sense.

Corey Quinn: Which is interesting, just from a perspective of meeting customers where they are without feeling the need to be what your customer is. More or less, providing services that help empower them to continue doing the thing they're already doing without a lot of the toil goes into that.

Corey Sanders: That's right. Yeah. And in some cases even doing things better. Or doing things more interesting. I think, a good example also sort of retail is Kroger's, they've added as part of their shelves, they have this product called Edge and they add sort of an additional advertisement underneath the chips to help you sort of decide which chip you actually want to go by, which turns out as actually a really important problem that we have to go work through.

Corey Quinn: If they can solve that one for me, it saves at least 40 minutes every I go shopping.

Corey Sanders: Exactly. This is a partnership with them to sort of solve this experience, this selling-buying experience in a new way. Not only is it making existing problems solved in an easier way, but creating new solutions to problems that sort of they didn't even realize they had. It's been really exciting to be a part of that and to be sort of at the front line of those conversations from the sales side.

Corey Quinn: Talk to me a little bit about how you're viewing the, I guess, world of hybrid, where people start off on premises, generally speaking, there aren't too many born-in-the-cloud companies that have scaled out today on Azure that have been in Azure their entire existence and are now multi-billion dollar companies. Everyone has something legacy. There's always something that's vaguely greenfield. Last year, we spoke a little bit about Azure stack, 12 months later, how's that going?

Corey Sanders: It's actually very exciting. I mean, I think we've seen a huge amount of interest and growth on the Azure stack side. What we're seeing is that there's, not only a lot of opportunity for these hybrid deployments for customers who come in and say, "Great, I'm going to need both the cloud-based solutions, but also something deployed locally. And it's not just because I used to have something locally, but because there's something that requires me to stay local.

I'm running a manufacturing plant and I really can't depend upon the network to always be there. I'm going to have something local that will keep the manufacturing plant running, but then use the public cloud for additional data analytics, additional analysis and so on." That combination of sort of intelligent cloud, intelligent edge has become a really interesting cornerstone of our overall platform.

Corey Quinn: Which seems critically important. We all have a story from somewhere in our past, where we have a dependency built into something that's far away and remote, and invariably, the fiber line leading their encounters it's natural predator, the backhoe. Suddenly, the entire factory is down for want of a single fiber connection.

Corey Sanders: Right.

Corey Quinn: And we have these agonizing stories that take 24 or 48 hours to get resolved during which time nothing happens. And the story of, "Oh, everything should live in the cloud. It should just be okay," simply isn't tenable when you're talking about significant volumes and significant scale here in the real world.

Corey Sanders: Right. That's right. Absolutely, and we're seeing this, whether it be network connection, whether it be local proximity, whether it be security reasons, having this combined solution is something that, I think, we really invented with Azure and we're starting to see some of the other cloud providers actually come out with similar solutions. Although, I would still argue ours is both the best and sort of hardened, but ...

Corey Quinn: Well, let's not kid ourselves, of anyone who plays in this space, I don't think you'd be able to find someone who understands what it's like to deal with on-premise customer workloads for the past 40 years at Microsoft.

Corey Sanders: That's right. Exactly. We've got a little bit of history in this space.

Corey Quinn: Oh yes. Everyone remembers those days with a smile and a wince and it's ...

Corey Sanders: Yeah, exactly. Maybe a whimper, but hopefully, that experience is turning into something very valuable to customers today.

Corey Quinn: Well, it does. People think I'm being sarcastic when I say this, and I may have said it to you last year, but Microsoft has 40 years of experience in apologizing for software failures to customers because in the cloud, things break, computers fall apart. It's what they do. It's in their nature, and learning how to tell that story in a way to a customer that is first, sympathetic and also aware of the fact that they are in pain and not blaming them for it-

Corey Sanders: That's right.

Corey Quinn: ... is absolutely critical.

Corey Sanders: Well, and then always making the service better. This is the thing that I feel really passionate about. The opportunity to learn from both what customers are doing, how they're using our services and even the problems that we have and how to consistently make it better and better and better. That is one of the exciting parts of the cloud. In the history of on-premise software, it was a three-year cycle, three years later things got better and now it's three days later. The opportunity to sort of have that turnaround is really pretty exciting.

Corey Quinn: Yeah. The faster you can iterate forward in speed time to market, the more valuable it is for everyone.

Corey Sanders: Absolutely.

Corey Quinn: Increasingly, although not for everyone, of course, there's not as much business value as there once was in running your own data centers effectively. Let's not kid ourselves, if you can't run a data center more effectively than I can, with your resources versus my Twitter for pets company of four people, there's a serious problem for everyone.

Corey Sanders: That's right. Well, and let's be clear, it's not me personally, I'm not actually in there running it. I don't think you'd want that. You may actually be able to do that better than me, Corey.

Corey Quinn: Well, Microsoft solutions does feel a bit like a catch-all solution. We don't know.

Corey Sanders: That's right. They call me in when the plug gets pulled out, someone tripped on it and I plug it back in. But-

Corey Quinn: Then you're the hero.

Corey Sanders: I am frequently the hero. That's right. As any good Corey would be. Let's be honest.

Corey Quinn: Absolutely. It's all in the name.

Corey Sanders: Yeah, that's right.

Corey Quinn: It's our Corey competency, one way or another.

Corey Sanders: It is, exactly.

Corey Quinn: Microsoft, in general, and Azure in particular, have an awful lot of services. During the keynotes today, first Satya's, and then I got to see Scott's as well, it went from, "Oh yeah, that's interesting. I've heard of that. I've played with that, some. Oh, that one seems better," and then it sort of drifted into the realm of, "I'm not entirely sure if you were having a joke at the audience's expense or not," where there's so many service offerings, it felt like I had gone across the street to the Cheesecake Factory instead of flipping through their menu with all of the different options you can go through. It's almost analysis paralysis.

Corey Sanders: Like a New Jersey diner. Yeah.

Corey Quinn: Exactly.

Corey Sanders: Yeah, yeah, yeah.

Corey Quinn: It's overwhelming. And I've been using at least a few of these things for almost 30 years myself. From your perspective, again, bounding into Azure, what are the major tracks of Microsoft offerings?

Corey Sanders: Yeah, yeah, yeah. When we think about, and we talk with customers about it, we do split it up into two big categories. One being migrate, one being innovate. When you think about what, again, coming back to, what is the customer trying to accomplish? Are they taking deployments and they're trying to reduce some costs? Or trying to reduce some of the energy of maintaining it? Then migrate's your path and you're likely using something like infrastructure as a service and then a bunch of the surrounding services, security, identity, management to sort of make sure you can run that infrastructure in a healthy and clean way. Then there's innovate. A lot of the Build talk track is around innovate, for obvious reasons.

These developers were sort of building new things, but then you sort of have a little bit of the data side, and a little bit of the application side. Application side, we talked a little bit about two distinct services, app service and AKS, or Kubernetes service, and really focused on those being sort of the cornerstone of the application side of the house when it comes to innovate. Then, of course, data. Quite a few services on data. One of the challenges with data in general is just how many different types of data opportunities there are in the world, whether it be NoSQL, whether it be SQL-based solutions and then NoSQL there's like a dozen of different one to ... You find anyone on the street, and they'll tell you, "No, Mongo is the best." Then the next person would be like, "No, Cassandra is the best."

Corey Quinn: Then you have people saying snarky things like, "No, Mongo's great for your production data. Not my production data, that stuff's important, for yours it's awesome."

Corey Sanders: That's right.

Corey Quinn: It feels like the number one thing people love in that space is arguing about whether other people are wrong.

Corey Sanders: That's right.

Corey Quinn: They'd be, "One thing that's better than all else is being right when other people aren't."

Corey Sanders: That's right. That's why we do our jobs, so we can relish in that experience. This is where, some of the services that I think are really exciting, something like Cosmos DB where it ends up being multi-model. Cosmos DB comes in and says, "Hey, look, we're a NoSQL solution." You choose the model you want. You want Mongo, you want Cassandra, you want Gremlin, you can use it on top of this Mongo solution and it all works, and globally distributed, et cetera. It's really pretty, pretty powerful.

Corey Quinn: Do you think there's an architectural lock-in concern there?

Corey Sanders: Well, this is what's so beautiful about using those open-source models on top of Cosmos DB. You can come in and you can code to Cassandra, which is not locked. I mean, it's not locked in any way. You can go and run it in any cloud. You can run it on premises. In fact, we have customers who are running on-prem and Azure using cosmos DB as part of it, but you don't have to worry about the management. It becomes sort of, in my mind, the best of both worlds. You're not locked in, you've got this open-source model that you're using, but you don't have to worry about management when it's run in Azure. In some ways we're winning you over, hopefully, with the ease of use versus this sense of, "Once you deploy here you don't have any choice."

Corey Quinn: Right. One thing you mentioned a minute ago, there's a lot to unpack in what you just said.

Corey Sanders: I say a lot.

Corey Quinn: We'll take it piece by piece. You mentioned the build is aimed at being a developer conference.

Corey Sanders: Yeah, yeah.

Corey Quinn: Which is likely to raise an eyebrow or two from people in, basically, all of the tech cities that live on the coast that we all live in and we all know and love, in that, well, look at the customer stories you told. These were retailers, these were auto manufacturers, et cetera, et cetera, et cetera. I think that something that gets lost in a lot of the conversation is that IT or engineering, writing software, is not writing Twitter for pets in the middle of San Francisco where we've taken a job we can do from literally anywhere and build a land crunch in eight square miles on an earthquake zone.

Corey Quinn: Instead, it's now about things like a hospital in Duluth, it's an insurance company in Omaha. It's companies that are doing real world things that aren't just creating this new data manipulation or tying APIs together and calling that a service. These are companies that do things that have business models our grandparents might understand mainly make more than you charge. These are business model that our grandparents might understand where you make more money than you spend and that's called profit, which apparently is a dirty word.

Corey Sanders: That's right.

Corey Quinn: It's neat to see developers who are writing "enterprise" software are not being forgotten, if anything, they're celebrated...

Corey Sanders: That's right. Yeah. I mean, I think that's exactly right. Especially, when you look at the breadth of different customers we had up there. Obviously, had a lot from Starbucks, then Virgin Airlines. The breadth of these types of customers, and the problems that they're solving. One of the things that we talk about a lot inside the company is every company in the world is becoming a software company, because every single company is now thinking about, what is the software that we need to build to be able to deliver services to our customers?

Whether that be retail, whether that be manufacturing, whether that be financial services, they are all building and developing solutions. That's why a developer conference includes folks from all of those industries across the world. It's a very exciting time to be in technology, frankly, because you're just seeing this really blossom no matter what industry you're in. It's all about the tech.

Corey Quinn: Absolutely it is. Although, I do question the validity of some of those demos. For example, you had Starbucks up there doing a whole demo talking about what they were doing with Azure, and they didn't mispronounce a single service name. They're Starbucks, getting people's names wrong is their entire take. I have to wonder how many takes it took to get them there, "No, no, it's called Azure. No, that's not how you pronounce it. No, wrong company. Try again."

Corey Sanders: Now remember that, part of that problem may be the handwriting on the cups. Maybe when someone else wrote it, maybe they didn't have the same challenges that they have inside their stores. So that could be, "Maybe we figured something out here."

Corey Quinn: Cache invalidation, naming things and renaming things. Hard problems.

Corey Sanders: In fact, that's right. Coming up with new names for sizes is definitely something that they needed a machine learning model for.

Corey Quinn: You talked a bit about Azure serverless databases as well, on stage today, you collectively. Which was fascinating, in the idea of pay per second in return for performance, so we can scale down to nothing.

Corey Sanders: That's right.

Corey Quinn: Out of curiosity, if something is stopped, and you just start it up again, is that going to have a cold start issue? Is that going to just suddenly be there and ready to go through some sort of interesting caching layer? What's the story there?

Corey Sanders: Yeah, the thing about both of the serverless product than the hyperscale product, both announced today, very, very exciting, is they effectively separate the storage from the compute. And then it allows a lot of things to happen. One on the hyperscale side, it allows you to scale horizontally as needed. When you look at some things that you'd normally would be concerned about based on how much data you have, like taking a backup. Normally, when you're thinking about a database, you take it back up, you're suddenly like, "Oh gosh, how much data do I have? How long that's going to take? Is it going to slow things down," et cetera.

With the new ability to split and scale, suddenly that becomes a non issue. Similarly, with the serverless solution, it allows you to effectively separate out your data storage from what's going to run on top of it. Things like backup or things like computational queries and so on. It allows you to split them apart and run them as needed. There won't be a cold start problem per se because it is still compute nearby, but the key thing is the separation and then the scaling as needed. You can scale both tiers independently, which you think of a classic horizontal, excuse me, vertical database. You have that sort of scaling problem where one may scale more than the other.

Corey Quinn: It also allows you to scale without downtime.

Corey Sanders: Yes, that's right. Exactly. Exactly. One, on the hyperscale side allows much, much better performance, and on the serverless side can result in much better cost efficiency.

Corey Quinn: There was also an announcement today about aspects of your database offering being able to run at Edge and on ARM, which is fascinating. Can you tell me a little bit more about that story?

Corey Sanders: Well, this is the ability to take SQL and run anywhere, I think is a key aspect. One of the things that we launched, what, two, three years ago, was SQL running on top of Linux, which is a big step forward for the product. There's taking that product and be able to run in any form factor, being able to run really, really small, being able to run really, really large, being able to run on Linux, being able to run on Windows, that opportunity to take SQL, it makes it much more portable. It allows you to avoid, again, this lock-in point. You can take it anywhere you need it. You can take it and put it on an Edge device, you can take an and deploy it in the cloud. It'll work anywhere you go.

Corey Quinn: It seems like it's going to be unlocking a lot of interesting stories. Do you have any customer stories you can relate today, or is that still too early to answer?

Corey Sanders: We have a lot of really exciting Edge-based stories, maybe not as many today yet that we're ready to talk about, using sort of this new database offering with ARM. But quite a few examples where we have people doing super interesting things with Data Box Edge, being able to do computation on the Edge, being able to take sort of information running from drones and take cameras and being able to sort of do visual representation of those cameras. I think we demoed that last year as built-in Edge-based functionality on those IoT devices. There's a lot of opportunity here, and we're really just sort of scratching the surface at this point.

Corey Quinn: You also have live-announced today an early preview, a Bata Box Edge Heavy, I didn't catch the exact numbers other than 650 pounds, which is a strange way for me to measure data storage, but I'll take it. What is the story there?

Corey Sanders: Yeah, I mean, we had Data Box, which ends up being a way for you to take data, ship it to us. We have folks who are doing that, and it's portable, you can pick it up, a human can pick it up and it's ruggedized. You drop it, it's not going to be a problem and so on. But then we've now added this Data Box Heavy solution, which is one of the more interesting aims, I guess, we've chosen, which is not something that a human can pick up. It ends up being...

Corey Quinn: Well, not with that attitude.

Corey Sanders: Yeah. If you get in the gym more, maybe.

Corey Quinn: Exactly.

Corey Sanders: But it's just a significantly larger device for storage and being able to basically transmit a ton more data right into the cloud. It ends up being a great opportunity for other solutions where the network isn't as rich or as open or just the amount of data's so prolific that it just takes that type of device to bring it up there.

Corey Quinn: How much data does that 650 pounds actually let me transport?

Corey Sanders: It allows you transfer up to one petabyte of storage and it's secure, so it ends up being encrypted as part of that transfer. So a lot. That's a lot.

Corey Quinn: Right. Not only is it going to be encrypted so someone can't steal data off of it, there's no way any reasonable person's going to be able to even lift it in the first place.

Corey Sanders: That's right. The one petabyte actually weighs a lot apparently.

Corey Quinn: Yeah. Absolutely. It's-

Corey Sanders: Is that not a measure of weight? I don't know.

Corey Quinn: It feels almost like it's a mind bender. How much does a petabyte weigh?

Corey Sanders: That's right.

Corey Quinn: And I betting that it's decreasing line over time.

Corey Sanders: But is it in a vacuum or not? I don't know. It's like-

Corey Quinn: That would be mass not weight.

Corey Sanders: That's right. A petabyte of feathers versus a petabyte of stones. I think it's still a petabyte, either way.

Corey Quinn: I think you're right.

Corey Sanders: That's right.

Corey Quinn: Something else that was released-

Corey Sanders: We've done some really deep thinking here today.

Corey Quinn: We really have.

Corey Sanders: It's been lovely.

Corey Quinn: Truly. Something that was mentioned about a month or so ago as best I can tell is the premium pricing for Azure functions. Specifically, you pay a little bit more on a per function basis and in return you don't get cold starts. Things are pre-provisioned, which I'm of two minds on. First, it sounds great that you're able to pay a premium for a tier of offering that doesn't have a cold start problem, and suddenly it's right there ready to go whenever something hits it. On the other, it does irritate me a bit to hear people devolve ...

So the serverless function discussion down into always talking about cold starts or not, because there are so many interesting use cases for this sort of thing that go well beyond someone is clicking in a browser and watching a spinner go until the site finishes loading. For that use case, yes, absolutely. This is awesome. Based upon what you've seen with people adopting Azure functions in the wild, are you seeing that those are the primary use cases? Are you starting to see people use these as more backend processing, where if it takes an extra couple hundred milliseconds to spin up, it's irrelevant?

Corey Sanders: Right. I would say, we are seeing a lot of both. I think this is why actually having the two offers is so important for the customers who, to your point, don't really care about that immediate reaction time and really don't need it to be warmed up when someone clicks on it at all times, having that sort of ability to just react and do it casually, casual functions, as it were, which is not the product name, but should be, I think. The backend processing, there's a lot of things, even response to IoT-based actions, where there may not need to be an immediate response, some sort of signal or some sort of information that that has popped.

Given there's typically a large amount of time just to get the information to you, a couple, like you said, even a couple of seconds is not going to actually make a big difference, but in the cases where it is a human interaction, UX clicking on buttons, et cetera, where it actually will change, fundamentally experienced, by adding one or two seconds, then you can pay a little bit of a premium and basically have that warmed up.

From our side, it's really just comes down to, "Are we starting it cold, quite literally, from nothing and we're just going to go spin up that function, or is it ready and waiting to run and just perhaps taking a little bit less of that processing power?" That allows us to have that different approach. I mean, I think we are seeing a lot of both. Certainly, a lot of the cold start is being used today, and I'm excited about this offering, because I think we'll see a lot of that warm start now come through.

Corey Quinn: One last topic that I wanted to cover before calling it a show, is the idea of lock-in. People talk a lot about it, and often in some of the stupidest imaginable terms. One of the undercurrents that I saw, through both keynotes I saw today, has been a repeated effort for Microsoft to reassure people that lock-in is not a concern. Everything is open, whatever you put in, you can take out. To me at least, it always seems like a bit of a red herring concern, because even if you build everything in the most open possible way where you can take it anywhere you want in a day, great.

No one is going to move $50 million worth of resources overnight anywhere. Plus, dealing with staff retraining or turnover as a part of that. Plus, while that's in flux, what happens? You can't generally take a multi-week outage, for most use cases, to do a migration. Everyone talks about going significantly out of their way to avoid lock-in, but in practice we're already locked in, and in many cases, by our own data gravity, by the choices we make around technology. How does Microsoft view lock-in?

Corey Sanders: Yeah. It's very interesting, I think, when you really sort of dig into customer motivations and customer experiences, there's probably two different tiers of the way customers approach lock-in. One is at the infrastructure level and, to your point, there's a little bit of, no matter what. Customers are locked into their own premises database today in many cases. They've got the tooling, they've got the powershell scripts that work, they've got the processes that work. And so, even moving to the cloud is breaking from that lock-in.

Customers wouldn't probably call that lock-in, because infrastructure management is infrastructure management. It's going to change and you're going to have to learn it no matter where you go. And so, to your point, there's sort of a baseline set of challenges just to move anything, period. Now, obviously, we want to try, and make it as simple as possible. Solutions like Terraform can enable sort of a little bit of that mobility, but there is still a management aspect. There's a monitoring aspect, there are going to be some deltas, even between clouds, even using something like Terraform to sort of offer a layer above.

Corey Quinn: And you're still looking at that point only using baseline primitive services. The higher level platform offerings are never going to be one to one compatible.

Corey Sanders: That's right. This is where, I think, when I think about lock-in and some of the approaches that we've taken with our platforms to minimize this as much as possible, the open-source capabilities that we have on top of our past services dramatically improve your ability to move if you want to. I think this is really where when I talk with customers about lock in, it's less around, "I want to move every week back and forth and back and forth and back and forth." No, it's more, "Hey, I want the option to move if things go sideways, if it turns out negotiations don't go well or if I don't like you anymore." Hopefully, that's not true for us in any case, but the opportunity to say, "And I don't want my developers to have to completely rewrite the app in those situations.

So yes, it's going to take work. It's going to be a migration cost. It's going to perhaps be a period of downtime to move things, but I don't want to have to go completely rewrite things." When you look at some of the sort of classic examples of lock in that really gets customers frustrated, and I'm not going to mention those customers by name here, the biggest concern is, "I wrote an app, that developer has moved on and now I have nothing that I can do to fix it other than perhaps hiring a whole bunch of new developers for." When you think of like Cassandra support on Cosmos DB, Postgre support on SQLDB, Mongo, all those examples, they offer this portability that is an escape hatch. That's actually really important for customers.

The ability to say, "Look, if you guys really screw the pooch on this, we have the option." And that is why I love those open-source capability. Even AKS, Azure Kubernetes Service, I love those open-source options because something like AKS, fully downstream compatible. If you love our service, stay. If you want to use our serverless capabilities, stay. If you decide you actually want to move to somewhere else, it's fairly easy to be able to just say, "Great, I don't have to completely rewrite everything. It's going to be a very similar experience." There's variances to this lock-in point. I think our open-source focus for a lot of our platforms, and then portability focus, even SQL is very portable, gives customers this option if they need it.

Corey Quinn: Yeah, it definitely seems like there's a long history of learnings that have helped shape a lot of the decisions that Microsoft has made.

Corey Sanders: Indeed.

Corey Quinn: Just from talking to customers, seeing the pain and the suffering and the triumph and the tears of running infrastructures in the 90s and the early 2000s. It's strange in that, first, that's an incredibly valuable learning field. But, secondly, I've got to say, I don't recognize the old Microsoft in what I'm seeing coming out of you folks today. I think that's a compliment, but-

Corey Sanders: I will take it as a compliment, whether you meant it that way or not. Yeah, I know. Going back to the beginning, even as you talked about sort of our vision, it is a very new style of vision that really all we focus on is how we're enabling others. It's very exciting. It's a great time to be in tech and a great time to be with Microsoft.

Corey Quinn: It certainly sounds like it. If people care to hear more of your wise words of wisdom, where can they find you?

Corey Sanders: Oh gosh. Twitter. If you want to hit me up on Twitter at Corey Sanders WA, W-A. That was originally Windows Azure and then we rebranded and I didn't change my name. Then it became Corey Sanders, Washington state. Then I moved to New Jersey and now I don't know what it is. So Corey Sanders WA on Twitter and questions, comments, people hit me up. We have a show, actually, Tuesdays With Corey Show that you can find there as well.

Corey Quinn: Wonderful. We will do our best to come up with a backronym for that Twitter handle these days.

Corey Sanders: Yeah. Please, tell me what it should be and I'll change it.

Corey Quinn: You heard him, Twitter. Thank you once again. Corey Sanders, corporate VP of Microsoft solutions. I'm Corey Quinn and this is Screaming In The Cloud.

Announcer: This has been this week's episode of screaming in the cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

Host: This has been a HumblePod Production. Stay humble.

View Details

About Austen Collins

Austen is an entrepreneur and software engineer located in Oakland, CA. He is also the founder and CEO of Serverless, Inc. and the creator of the Serverless Framework, an open source project and module ecosystem to help everyone build applications exclusively on Lambda, without the hassle and costs required by servers. He describes himself as a product-obsessed, software engineer who is focused on making meaning, business value and great customer experiences.

Links Referenced

  • Serverless.com
  • twitter.com/austencollins

Transcript

Speaker 1: Hello, and welcome to Screaming in the Cloud with your host, Cloud Economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey Quinn: This episode of Screaming in the Cloud is sponsored by GoCD from ThoughtWorks. The worst part of trying any form of CI or CD tooling is getting it up to speed to test it out in the first place. You’ve got to configure it. You’ve got to hook everything together. Somehow now you’re 80 steps in and for some reason step 81 is that you need to tame a wolf.

GoCD knows this better than most of us. That’s why they’ve got a new test drive option to get up and running with their tooling in seconds. Download their binary, run it locally, and set up your first pipeline in your browser. Think of it like a demo, but rather than some arbitrary “Hello Word” app that looks nothing like what you’re running, instead it’s being done with your own application. Give it a try. Stop getting bitten by wolves. Learn more at gocd.org. That’s G-O-C-D dot org. My thanks to GoCD for their support of this ridiculous podcast.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Austen Collins, founder and CEO of the Serverless Framework and also AWS Serverless Hero. Welcome to the show, Austen.

Austen: Hi, Corey. Thanks for having me.

Corey: Thank you for joining us. One of the things I've always found fascinating about serverless was the learning curve on what it takes to get up and running. When I built my first Lambda function, it was a two-week long process of stumbling my way blindly in the dark with an ax hoping I wasn't going to lop off a major limb that I cared about. It was a bunch of guess and check, a bunch of iterating forward, and it never really got easy for me.

Then, months go by. It becomes a little bit more understandable, and then I discover a few bits of tooling including the serverless framework. My last Lambda... My last application that I built using serverless took about 15 minutes also because I'm a terrible developer who doesn't know how tests work, but it's really sped development fantastically, so first, thanks for building it.

Austen: Yeah. Thanks for the positive feedback. Your story is the same story I had early 2015 right after I re-invent... right after they announced Database Lambda. I stumbled through it myself. in those days, it was still pretty raw. I mean, it just was very limited when it comes to features, and use cases, and documentation, of course. All that kind of inspired me to build the Serverless Framework in the early days, so yes, I empathize with your pain, and hopefully, we're building a solution that can help people get through that more easily.

Corey: For folks who have not yet dipped their toes into the serverless waters, I feel, first, we should probably disambiguate a few things. AWS Lambda is a function as a service offering, which has been a part of a movement that has been defined as serverless. That is not to be confused with the Serverless Framework, an open-source project that you started, and Serverless Inc., the company that you now run that does things around the Serverless Framework. It's a bit of a namespace collision, but I've got to say. Great SEO for you folks.

Austen: Yeah. It's certainty interesting position to be in these days. I won't deny. It certainly has its benefits, but it also comes with a lot of confusion. Going back, it's a bit hard to... for people to understand now, but in the early days of 2015, there wasn't a serverless trend. There wasn't a huge buzzword. It wasn't like a major category, and it wasn't clear too that this was going to be such a big deal. Right?

To be honest, I was in the Bay Area. I was working as an AWS consultant at the time, and I was trying to raise awareness about Lambda just within my personal network saying like, "Look, this could be the future of how we run compute in the cloud." It's got everything that I've always personally wanted. It's like this convergence of great ideas of our time. It's microservices, paper execution, auto-scaling, event-driven, but there wasn't as much excitement kind of back then because there wasn't a big story around how you could put a lot of use cases around it, and certainly, there wasn't this great buzzword that developers love, and so in the early days, it was all a lot of speculation like, "Hey, this might be... This seems like the thing we should call it."

All I knew at the time was a serverless... Not technically accurate. It's not as technically accurate as the cloud, right, as a term for describing something, but when you say that to a developer, and I had the same experience when I saw the word for the first time, the emotional response that you get out of the developer and that I had when I first read it is significant. As a developer, it's just music to their ears. It's everything they've always wanted, they've always wanted in a compute platform. Right?

Yeah. Now, it's a bit early. Now, it's a bit confusing. There's a lot of other people trying to cram a lot of other things under the serverless definition, but in the early days, it wasn't the case. It was kind of a big speculative bet, and it's just surreal how things change over time, especially if you've been there from day one.

Corey: For me, one of the big value propositions behind serverless when I first got into it was... It was sort of a dual-revelation for me, the framework itself. Namely, that, first, I could build something programmatic without clicking around in the console and creating disasters for myself in the future, and secondly, I never even had to look at CloudFormation. Although, you do distill down into CloudFormation templates, which is how you run programmatically under the hood.

I mean, that alone is worth a lot to me because CloudFormation, at least for me, doesn't resonate with how I tend to think about infrastructure or how I tend to think about configuration, and that's probably my own failing, so don't send me emails folks if you're a big CloudFormation fan. First, I'm sorry. There are pills for that, and secondly, it's just not how I tend to operate.

From my perspective, the value of Serverless Framework was that it provided an intelligent wrapper that got rid a lot of boilerplate, a few entries in a YAML file, and I would wind up effectively being able to compress down a 200-line CloudFormation template in about a dozen lines of YAML that were extremely straightforward and every line did something that I could understand. That was transformative for how I started to think about this. API Gateway even more so, given that that thing is impossible to understand when you're looking at, but it winds up being three of those dozen lines, and it just works.

Austen: Yeah. It's a powerful abstraction, and I don't... Looking back at the whole trend, and all the movement, and how people are talking about it, I'm not sure... I think Amazon absolutely all credit kind of goes to them, but I think it's that configuration file in the Serverless Framework that really helped kind of prove out or at least show the world like how to think about a serverless application or a serverless architecture in general.

Again, back in the early days, there was this great new compute service. They were kind of pitching it. Amazon is pitching it as a venture of encode as glue code, and when we saw that, we're thinking like... You can build all types of applications and use cases on this because the value prop is so compelling, and that is deliver a software and applications with radically less overhead, but there wasn't a clear story as to how you could put all those applications and use cases on top of this, and that is something that I spent a long time kind of scratching my head over like, "What is the developer experience you could wrap around this great new infrastructure as a service to make it accessible to developers and maybe people who don't necessarily have a lot of cloud knowledge?" Which has worked pretty well now.

We actually surveyed our user base, and about 25% of our users have never even used AWS before, but we could chat about that a bit later, but it's that simple kind of... It's a simple story of functions and events. At least that's how I think of it and I... When you go into that configuration file, you're really just thinking about, "Okay. What's my business logic? What's my task? Let me declare that in the least meta-configuration as possible, and then wire it up to events that trigger it to run."

The framework, our focus is, "How do we make that configuration minimal, keep it minimal, and have it actually kind of remain focused on a workflow rather than the infrastructure?" I think that simple story of functions and events kind of helped propel the Serverless Framework and the serverless architecture to the mainstream, and it still resonates incredibly well with developers today because they're not thinking about all the infrastructure components behind the scenes that need to get provisioned and wired up in order to work correctly together. They're just thinking about a workflow. "Here's my logic, and when this thing happens, trigger this logic to run."

Corey: That does become something that is profoundly transformative. It effectively lets you run... just worry about the code and the application that you've built, not the infrastructure underneath it, not patching, not continuing to have to worry about things that are frankly not central and core to what your company is attempting to achieve.

Austen: Absolutely. No company sets their Q1 objective as to get a database provisioned, right? They want an outcome, and that's one of the strengths of the framework. I mean, you're really focused on more of the outcome level, which I think is a big reason for its success, and increasingly, as I look at the cloud and the direction the cloud is going, and it seems like more of a outcome-focused direction rather than infrastructure-focused.

Corey: I think that's probably the right answer. I think the tide is rising, and as I've talked about on the show repeatedly with other folks, as cloud tends to wind up, removing a lot of the things that we used to have to do in every company, there's a need to focus on the future. What are we building? What is the point of our companies? That's not replacing hard drives anymore. Increasingly, it feels like it's not even going to be provisioning instances. It's going to be about writing application code, and everything else gets handled for us.

Austen: Totally agree, and I'd question the amount of application code we'll be writing in the future too at the same time because it's... One thing we see a lot... The way we think about serverless architecture is you're basically building applications on top of managed services and kind of gluing them together using this... Yeah, using functions and using events, and most companies are... It's usually never just about the functions. It's about the other managed services that you could use with the functions, and we see every day the majority, vast majority of Lambda functions are being used with other managed services, and it seems like people are relying on more and more these managed services rather than building and maintaining their own versions of those increasingly at an unprecedented rate.

Based on kind of how we're using the cloud, the big public cloud providers react to all this, it seems highly likely that they'll hit the market with a whole bunch more managed services in the future. Higher levels of abstraction where you won't be asking Amazon for a virtual machine anymore. You're going to be asking Amazon for a solution to a business problem like, "Hey, Amazon. What's in this image?" or, "Can you take this audio file and transcribe it?"

This whole model of building applications where you're just kind of gluing all these higher order services together via functions and events is fantastic. I mean, you're able to produce more meaningful outcomes and use cases as a result and of course, write less code because you're outsourcing that code to more and more services, so it will be interesting to see how this unfolds, but I am kind of questioning how much code people will be writing in the future because already, we're seeing the percentage of custom code that people are writing in their application shrinks significantly.

Corey: Oh, absolutely, and thanks to Serverless Framework and a few other things, I wound up having a developer recently working on something for me, and what all I said... because I didn't want to go through all the pain and process of teaching a contractor how necessarily to work with the intricacies of serverless stuff. I just told him, "Cool. Edit this script, and then commit it and push it to master." Whenever it's pushed to master, it will leave 20 seconds later. Then, it's going to work at this end point, and I gave an API gateway URL, and it worked almost like magic.

Under the hood, that was just a code build tied to Serverless Framework doing a quickly deploy to a particular branch, and that's a test environment. Worst case, I can go in and make some changes in there pretty easily if anything got broken, but over a two months' span, nothing ever did. It just worked.

Austen: Yeah. That's the beauty of the serverless architecture. It's not quite there yet, but it is getting closer and closer to... It's sort of a set-it-and-forget-it type of develop and model.

Corey: Absolutely. One thing I do and I asked you about is... Well, there are two directions to take this in. The first is we've talked about AWS, but one of the distinguishing characteristics of Serverless Framework is that it works across multiple providers serverless offerings. GCP, Azure, and I'm sure there's one or two other things in there that I'm not thinking of. Personally, I haven't played much with any of those other providers. I've been doing all of the serverless stuff that I do in the AWS ecosystem. Is that fairly typical? Are you seeing broad adoption industry-wide of multiple providers?

Austen: First off, AWS. Strong credit goes to them for being such innovators in the cloud space in general. Then, all over again in the serverless space. Right from the beginning, I was fortunate enough to work with just like the Lambda team in the very early days, and it was very clear how focused they were on investing in this, and maturing it, and kind of pushing it forward as the next kind of mainstream cloud architecture, and they were just all in from the beginning, moving really quickly in that direction. It was pretty incredible to watch.

As a result, they've really come out with some of the best in class serverless offerings right now that are the result of years of investment, and fine-tuning, and hardening, and so they have that going for them. Now, the other cloud providers are certainly catching up, and they even have some serverless services that are... have some compelling features over what AWS offers, but for the most part, absolutely, we see the whole industry largest is kind of favoring AWS in general. I mean, they already kind of had that going for them, and then they invested so aggressively in the serverless stuff. It's been pretty easy for people to just adopt serverless on AWS in general because a lot of the pieces are there. You need to build a lot of use cases, so we definitely see a lot more AWS usage in the framework.

Now, when it comes to other vendors and multi-cloud, questions around portability and stuff, we get a lot of questions about this and people ask like, "Hey. What does this look like? How are people leveraging different clouds, different cloud providers?" We also have all this kind of thinking coming from the container era where you just lift and shift stuff around. That doesn't necessarily fit within the serverless era because serverless is a totally different beast. I mean, you're building applications on top of proprietary cloud services because you made a conscious decision to just take advantage of those cloud services because they offer a great managed experience out of the box so they reduce your overhead and you could deliver something to the market much faster with the least amount of overhead and cost.

Now, so we don't see people kind of lifting and shifting serverless architectures. It's very rare that we see that. We see a lot of people looking to be able to easily basically have a similar application experience across the providers and across different services, so we see a lot of interest in kind of deploying stuff as easily to land as they... to Fargate as well or a lot of interest in just like, "How do we just deploy a function to Google Cloud and Azure without having to change our mindset and become cloud platform experts?"

Then, also, we see a ton of interest in what I call vendor choice largely, and I think this is... This might be the interesting trend to watch in the serverless era. Some people call it serverless. Some people call it serviceful because you're composing all these services, cloud services together to make meaningful outcomes. I think in the serverless and in the serviceful era especially, a lot of people don't want to be limited to specific platforms in order to take advantage of the services they need to build the best possible product.

I think if you're a developer, if you're an organization, you want to build the best products. You want to be able to leverage all the services, the best possible services to do so, and so what we hear almost every single day now is yes, we're using a ton of AWS. They've got like some of the mature serverless infrastructure services around, but we're really looking at Google's machine learning stuff, and now we're kind of looking at this spanner thing, and is there a way that we could have the same application experience and build serverless use cases on top of Google, on top of Azure just to take advantage of some of these interesting services that they might need to use for one reason of the other without having to, again, become like a platform expert because obviously, platforms are super complex and there's just so many dimensions of complexity each single platform.

That's the beauty of like a great vendor-agnostic abstraction. Right? Without having to necessarily become a platform expert, you could build serverless use cases. Take advantage of super powerful services on each cloud provider, again, without becoming that expert, so I think that is probably the thing to watch in the serverless era. Lifting and shifting these architectures is a huge challenge. Again, it's not just about the functions. It's about the services you use with the functions and all... Most functions are being designed with proprietary code or they're being written around proprietary APIs, and so... but this thing about being able to use services on different platforms I think is certainly interesting.

We're hearing the demand for it, and I... As a very product-focused engineer, I empathize with it, right, because when you want to build great stuff, you don't want to be limited to a specific platform, and furthermore, if these cloud platforms keep kind of hitting the market with more and more manual services that are higher levels of abstraction that solve every single type of business problem you have, I say you especially don't want to be limited. Right? You want to be able to use all those things to build the best possible outcome.

In fact, I would say a lot of the serverless applications that are using our framework today, I'd probably say the vast majority of them aren't exclusively AWS. Now, that does not mean that they're AWS, and Google, and Azure, but it does mean that they're... highly likely that they're AWS and some OS-Zero, AWS and maybe some Twilio, AWS and some Stripe.

Corey: Oh, absolutely. I take a look at everything I've built. There's almost always going to be some other third-party API pinboard or pocket for a couple of things that I've built. There is also integration on the other side with SendGrid because SES is still not quite what I'd want it to be, and you go down this entire list. There's always going to be something else.

When we talk about multi-cloud, there are... You see things such as integrating things like PagerDuty. I mean, there's no under-the-hood service that AWS would offer something like that today, and so at some point, being multi-cloud does not mean that it becomes a suicide pact. It just usually takes the form, at least in the experiences that I've had, of being focused on, "This is the central core of what we do, and then we loop in other things in an ancillary basis where appropriate."

Austen: Absolutely, and that is where we see the most of the demand on the framework on our end. It's like, "How do you give us this great vendor-agnostic abstraction so that we could use all these things, again, without having to become platform experts because those platforms are huge?" Right? They want to be able to use all these things. Okay?

You want to build the best possible products? You should be free to use the best possible services, and where we're going, this serverless/serviceful era, it seems like it would be a wise decision to leverage something that gives you access to all that, but on our end, it's a ton of work, right, because to build that vendor-agnostic abstraction, to deliver that experience, it's not easy, but we've got some pretty cool plans around that in a potential Serverless Framework version two, which we've been working on for at least a year.

Corey: One thing I do want to cover though is I guess I've started to feel a little uncomfortable recently with the level of hype around serverless as the architecture of the future. Now, I'm very much a proponent of using it. I think that it makes an awful lot of sense for an awful lot of use cases. The concern I have is where you start to see this with any technology appropriate part of Gartner's hype curve where we are now seeing people half-listening to part of a problem description, and then chiming in with, "Oh, serverless is the answer. Serverless is the answer," and I find that to be worrying. Do you see any of that?

Austen: Of course. Here's kind of how I think of it. We certainly see it. It's par for the course with any major exciting technology trend. Right? There's always kind of these devout believers, and we see people cramming use cases in the Lambda. Tons of use cases that aren't appropriate for Lambda all the time. Right? It is incredible, the amount of kind of hacks and workarounds that they're implementing just to be able to use Lambda. Right?

I'll never get over just the... Yeah. It's so astonishing to see all the creativity that's going into this, but on one hand, it's like, "Okay. This is kind of crazy. You shouldn't be having to warm up all these functions all the time." Right? On the other hand, I look at it and I get it. Right? I mean, going back to the why, why is this such a big deal. It's the promise to deliver software with radically low overhead, and everybody wants that. It doesn't matter if you're a large enterprise, you're an SMB, you're a startup, or you're kind of a lone hacker in your mother's basement. Right?

Everybody wants to be able to deliver software that has like... requires the least amount of maintenance, and the fact that there's all this hype around it I think is... I get it for that reason because everybody wants that. On one hand, I want to say, "Yeah, it's crazy that there... that everybody is implementing all these workarounds and these hacks to warm up their functions and get around database performance issues." On the other hand, I think you could also interpret it as possibly a signal of the pent-up demand that is just waiting to take more and more advantage to this stuff, waiting for the services to become more mature, waiting for them to get more hardened so that they can appeal to a wider variety of use cases, and I think that's going to happen.

Based on our conversations with AWS, and Google, and Azure, everybody is working to almost transform a lot of their cloud services to be entirely serverless, and I think today, yes, it looks a little bit crazy, but I understand the motivation, and I think it's a good signal for how the market will move in the future.

Corey: I agree with you, but I also wind up seeing people pushing back against various decisions that I've made, for example, whenever they learn what's under the hood. Recently, you may or may not have noticed that suddenly, my crappy 1998 style web design now has a giant platypus on the top of it and has I guess a more cohesive visual aesthetic for lack of a better term.

Austen: Mm-hmm (affirmative).

Corey: As a part of that, I migrated off of static site hosted in S3 with the land view Lambda@Edge functions and a bunch of other nonsense over to WordPress. Just standard WordPress hosted by someone else because I am still a responsible grownup and running WordPress is something I'll never do again myself. On the one hand, it feels like a bit of a regression, and people get very upset when they hear that I've done that because I'm moving backwards, I'm not paying attention, et cetera, et cetera, but the reasons that I had to do that are very in line with the underlying theory of serverless.

When I'm paying someone else to run the entire thing for me, I don't care that it's WordPress. I have people working on that. I know that when I need to bring in a developer to build a thing, I can find WordPress developers for far less effort, time, and better availability than I can finding someone who's up to speed on the latest serverless technologies. For what I do and the business needs that I have, it's very much the right answer, and that's one example of a workload that was not appropriate for serverless, given the constraints that I tended to have. Where do you stand on that type of thing?

Austen: I think you should focus on the business problem first and foremost. Right? I'm a big fan. I think techno- It's not about technology. It's about vision. It's about the outcome. It's about the problem you want to solve. Right? The reason I got into the serverless stuff and why I was so obsessed with Lambda and working nights and weekends on this framework in the very early days is because I think technology should enable vision, and then get out of its way. Absolutely. Right?

It also depends on how you define serverless. I mean, is kind of WordPress as a service, is that serverless? Absolutely. Right? Serverless is kind of a notoriously vague definition, and I think it will get perhaps more vague here, but I'd say WordPress is a service solution. It's almost as serverless as it gets. Right? In general, absolutely. Solve your business problem first and foremost, and then second, of course, like Lambda... The serverless architecture. It's not the best fit for everything, and it may not be the best fit for everything.

It's naïve to think that a single technology will be able to accommodate all possible use cases. Right? There are a lot of people who really value and need to do the regulatory requirements for just policies. They need to have full control over everything. I completely understand that and empathize with that. I absolutely focus on the business problem, and I still think of a hosted WordPress offering or that outcome as a service is perhaps even more serverless than Lambda at the end of the day.

Corey: Yeah. I think you're very right as far as looking at this from a larger context of solving business problems rather than focusing on the hype cycle... rather than focusing on the true religion that is serverless. That tends to be one of those areas that is fraught with peril, and continuing to barrel down that. In some cases, it's not going to be viable.

What I do for a living is I fix horrifying AWS bills, but there's never going to be a story where I look at what someone has built and the response becomes, "Well, you can save 80% bill, which I could, by rewriting it all using serverless services," which they could, "And it will only take you 18 months of doing nothing else to get there." Also true. Also, something virtually no company in the world is going to do with an asterisk next to it because that's not easy. That's not straightforward. That's not something that is going to add significant differentiating business value for an awful lot of companies out there given what other constraints they have to work with them.

Austen: Yes. We see that a lot. The majority of serverless architectures being built today are greenfield. There aren't a lot of migrations because it's a lot to handle at once or at least currently to migrate to a serverless architecture. You're not just kind of dealing with this whole new architectural pattern, but also, there are other things like microservice design, just adopting... You have to have a lot of knowledge about the cloud, and so it's a lot to take on currently.

In the future, I think that there will be a better migration story, something that looks like... You should just be able to pour it over your previous outcome on a serverless infrastructure without even really having to know about it. Right? But until then, until that stuff comes out, I think it's certainly trying to figure out how to migrate these things over to a serverless architecture. It can be fairly time-consuming because there's a lot to take in at once.

Corey: It very much is. One thing that I wanted to bring up with you... Normally, at this part of the show, I would be asking, "Is there anything you want to promote?" but you are effectively the CEO of a startup, which means you have... always have something to promote and it's always generally the same thing, so rather than asking you, I want to turn this into a question on my end. Namely, you've raised two funding rounds publicly that have gone reasonably well. You've raised a fair bit of money and certainly, have a whole lot of buzz going around this, but my relationship with your company comes down to the open-source community Serverless Framework, and that's awesome. You've been talking a bit lately about the Serverless Framework Enterprise or the Serverless Enterprise Framework. There are three words in there. I'm not sure quite the order they go in. First, what is the order? Secondly, what does it do?

Austen: Yeah, great question. Yes. Serverless Framework Open-Source came out 2015. I think it was July 2015, and Serverless Framework Enterprise just came out about a month ago, and here is why we came out with Serverless Framework Enterprise. Putting aside, yes, we're a business, we've been commercialized. We've been focused largely on community growth for the first few years, but at the same time, when we walk into kind of like an organization and they're serious about adopting serverless, which is where I'd say a lot of the market is at right now.

I think the past few years have been around doing some POCs, kind of playing around with it. Usually, a developer will kind of bring Serverless Framework database Lambda in their organization. They will start doing a few kind of minor use cases, automating some part of their devops pipeline or something by putting some raw... some low-risk task in a Lambda function, and then they'll start doing a couple more functions. They'll start to attract attention from maybe the team lead, or director, or something, and then they start looking at their first serverless use case.

I'd say like the market has kind of done all that now, and they get the value prop of serverless. Again, deliver software with radically less overhead, and they want to do two things. They want to bring on more serverless developers. They want to kind of build out the number of developers on their teams doing serverless development, and they want to focus on more mission-critical use cases with the serverless architecture.

That's kind of where we're seeing the market is at right now, and when we kind of walk into organizations, and I talk to at least two organizations every single day that are doing this, they all say the same thing. They say, "Look, Serverless Framework Open-Source is fantastic kind of development, deployment, kind of testing tool for serverless architectures. It really helped our developers get our first few serverless use cases up in the cloud, but now we want more developers doing this, and now we want to focus on more mission-critical use cases. Can you help our entire team get into production as easily as you helped kind of those initial couple developers build out those first serverless use cases with Serverless Framework Open-Source?"

They said, "In order to do that, here's what we need. We need monitoring. We need a better monitoring solution. We want a better testing solution. We want to make sure that our developers are following organizational policies. We want to make sure that there's oversight for all this and that this jives well with the ops team." Right? They say, "Could you give us all this stuff out of the box with the Serverless Framework?"

We thought about that for a long time. We've kind of experimented with a few different products over the years, and this one certainly makes the most sense I think because the serverless architecture, yes, incredible value, but it's a different architecture. Right? It is kind of a whole new thing, and it introduces changes across the entire application lifecycle. This is not specifically a deployment problem. It's not specifically a testing problem. It's not specifically a monitoring problem or security problem. Right?

It's all these things at once because you're talking about building out this microservice architecture-based and all these functions as a service. Every single one of those has infrastructure dependencies. A function needs something to trigger it to run. A function needs some infrastructure to complete its business logic like a database, for example. How do you monitor all these things? How do you test them all? How do you collaborate across them? How do you share access to all them? How do you monitor the security landscape? Right? Every single one of these things has its own permission model and you want to make sure that none of those are overly-permissioned. Right?

Again, we look at all this stuff, and after kind of trying a few things behind the scenes, Serverless Framework Enterprise, a single solution that could really handle every single phase in the application lifecycle in one nice integrative solution made a lot of sense, especially for developer teams that want to take this more seriously. A lot of them are usually trying to build a lot of this by themselves on AWS like at their... a better serverless platform for their own team or they're trying to kind of cobble together their own solution to this via like five other vendors. One for security. One for monitoring. All that stuff.

Our mission is like, "Hey, we just want to give you all this stuff out of the box right after you run a serverless deploy," and we think, especially with the serverless architecture, that this... This is a bigger opportunity than ever, right, because the whole premise of serverless was like, "Yeah. We need to lower overhead and reduce maintenance as much as possible in the software we deliver," and you should actually need... If we're successful at this, I'd argue you don't need to have a separate monitoring product for this or a separate security product.

In fact, I think that if we do end up needing all these things for serverless, I'd say that something has gone horribly wrong, and I'm the first one... I do drive our team nuts with this because we are doing serverless monitoring features and stuff right now within Serverless Framework Enterprise. Every single time you do a serverless deploy, you get monitoring, you get metrics, you get alerting all automatically configured for you out of the box with zero config, and we're talking about like, "What does that monitoring experience look like? What do the metrics look like? What do the alerts look like?"

We're so tempted to go kind of build this dashboard just filled with chart junk. Right? Just a whole punch of like flashing lights and lines zigzagging left to right, and I am... On the team, I'm kind of pushing them and say, "You know what? Like I personally got into serverless to get away from dashboards like this, to get away from all this stuff. Like how can we simplify..."

Corey: "We have a metric. We have to expose it to everyone."

Austen: Yes, exactly. I'm really prompting the team hard. I'm like, "How do we simplify this? How do we..." We'll go through some creative prompt exercises like, "What does serverless monitoring look like without any charts? Like what if you just had to take all that stuff away, like what would that look like?" I think the answer that's in there and in those questions I think is more true to the spirit of serverless than having to go through this exercise of like bringing all these other tools and like making your own kind of... cobbling together your own solution for this.

I just think, again, if we have to do that, something has gone horribly wrong in the serverless trend, so we're on a mission again to simplify all that stuff and provide a single integrative solution for developer teams out of the box that's Serverless Framework Open-Source paired with Serverless Framework Enterprise, and in about one to two weeks here, Serverless Framework Enterprise will just be included in Serverless Framework Open-Source, so every single time... within a freemium tier. Every single time a developer deploys for the first time their serverless architecture, they'll have the metrics, the alerting. They'll have testing capabilities, all that stuff just out of the box in a dashboard right after you run serverless deploy.

Corey: That is absolutely going to be a significant value. It almost feels on some level like it's a bit of a misnomer when you say... When you put the word "enterprise" on something, I hear expensive which... Great. Awesome. We all need to make money, and I don't think any of the VCs behind your open-source project are doing it as a charity, but by the same token, that's something that's incredibly valuable for folks all across the spectrum from small businesses like mine to giant enterprises like companies that actually make money. As we spend to scale, it looks like a phone number and everything in between.

Austen: Yeah. Absolutely. I'm pushing as hard to change the enterprise naming. I think that there's a place for that, but right now, we're very much focused on developer teams, and so we might adjust that name in the near future here, but yeah, this stuff will all be given to you out of the box. This is kind of our vision for how serverless development and operation should be, and we'll... We've got like 10% of our product roadmap completed, but right now, there's a ton of interest. I mean, everybody like if we... As soon as they see that you get all this stuff out of the box right after your first serverless deploy, it's a pretty magical thing.

Corey: Absolutely. Austen, thank you so much for taking the time to speak with me. If people want to hear more about what you're up to, where can they find you?

Austen: Serverless.com. That's everything that's going on with our company, and there's just a ton of material in general on serverless architecture, on the trend. A lot of great learning resources, of course. Tons of examples. We've got a brand new example to explore. I'm a big fan of starting with examples rather than going directly at the docs, and then I'm on Twitter, @austencollins. Austen is spelled A-U-S-T-E-N, so it's a bit different than usual, but yeah, those are the two places to watch, and we've got a ton of announcements coming out this year.

I mean, we're both trying to cater to the serverless architecture challenges today. Right? That is, again, brand new architecture. It introduces differences in every single phase of the application lifecycle, and we're determined to solve all those with Serverless Framework Open-Source and Serverless Framework Enterprise, and then we have a handful of kind of more progressive ideas that we're playing with like serverless components. We have event gateway. We have an event specification that we kind of pioneered back in 2017 called Cloud Events, which has since been adopted by the CNCF, and a lot of people are working on that. Yeah. I'd say serverless.com and my Twitter handle is a great way to stay in touch with all these things.

Corey: Perfect. Austen, thank you so much for taking the time to speak with me today. I appreciate it.

Austen: Likewise, Corey. Take care.

Corey: Austen Collins, CEO and founder of Serverless Incorporated. I'm Corey Quinn. This is Screaming in the Cloud.

Speaker 1: This has been this week's episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com or wherever fine snark is sold.

Speaker 4: This has been a HumblePod Production. Stay humble.

View Details

About Emily Freeman

Emily Freeman grew up in the “swamp” as Trump lovingly refers to it. With politics in her blood, she chased after her dream of living out an episode of the West Wing. After four years of arguing — pretty much sums up a PoliSci degree — she left school disappointed that campaigns are more about recruiting 20-year-olds to live in poverty than it is to wine and dine Koch brothers.

Her dreams of Aaron Sorkin-level dialogue and Michelin-star dinners dashed, Emily took up ghostwriting. No, those bloggers you read with millions of followers don’t write their own articles. Sorry to disappoint.

After many years of typing, Emily had a slightly-older-than-quarterlife crisis and made the bold (insane?!) choice to switch careers into software engineering. With no experience at all, she packed her six-month-old daughter, blind dog and a few boxes into her anti-mom mobile of a sports car and drove across the country to attend a seven-month code school.

Emily completed seven grueling months of code reviews, pair programming and learning Ruby on Rails. After falling in love with Denver, a city as vibrant as she is, Emily decided to stay.

Emily is the author of DevOps for Dummies and the curator of JavaScript January — a collection of JavaScript articles which attracts 30,000 visitors in the month of January. To learn more about Emily's story, visit Growth in Fear.

Links Referenced:

  • @editingemily
  • emilyfreeman.io

Transcript

Speaker 1: Hello, and welcome to Screaming in the Cloud, with your host cloud economist, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode is sponsored by Scaylr. Because all kids hate their logs… doo do do doo do doot.

When your site doesn't go

Or maybe it's slow

And people can't load up your blog

Where do you start

what's the state of the art

With logs logs logs

Logs, logs

Full of repetitive noise

Logs, logs,

Awk and grep? Sorry they're toys

Everyone hates the logs

Nobody can read their logs

Improve the state of your logs

Scalyr can help with your logs

logs logs logs

Logs. From Scalyr dot com

Everyone wants a log

You're gonna love it, log

Come on and get your log

Everyone needs a log

log log log

Logs! from Scaylr.com.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by your friend and mine, Emily Freeman, who's a senior cloud advocate at Microsoft. Welcome to the show, Emily.

Emily: Thank you. Thanks for having me.

Corey: No, thanks for coming. So you are a cloud advocate at Microsoft, presumably Azure?

Emily: Yes.

Corey: And that's a cloud-ish.

Emily: Hey.

Corey: So tell me a little bit, first of all, what is it you do as a cloud advocate?

Emily: Very similar to any kind of developer relations role, as a cloud advocate, I am sitting in that space between the developer communities and the product teams at Azure. And so, day to day, it can vary but generally, I focus on creating high quality technical content. That could be writing, speaking, video but also relaying feedback from those developer communities back to the product teams so that we actually build Azure in a way that solves the problems you're currently having.

Corey: I'm having many problems, and I'm not sure how much any one company can do to solve them for me. Most of my problems have long names and they're expensive to fix.

Emily: I feel like that's more like therapists?

Corey: Exactly. That feels very much like cloud is more about transitioning a group therapy into a corporate setting.

Emily: Yeah, that's not wrong.

Corey: You could argue that's what DevOps is due to some extent.

Emily: It really is.

Corey: As of the time of this recording, you have put a nearly final draft of DevOps For Dummies to press. Tell me about that.

Emily: Yes. So I'm currently in the review. They've already marked it up with plenty of red. And now, I'm kind of going through and making sure that I actually agree with my past self. But it's been an adventure. I wrote the book. It's a high level overview of DevOps as a methodology and how to actually implement it as a practice.

So, with DevOps For Dummies, I tried to hit all of the stages of DevOps including the start where like, "Okay, I've heard of this. What is it? How do I implement it? How do I transform the company and get people on board?" I talked a lot about persuasion, and showcasing how DevOps can improve your team's velocity, whether to executives or other engineers, all the way through injecting DevOps into every piece of this offer delivery lifecycle.

Corey: So, I'll admit when you first mentioned that you were writing that book, my immediate response was that DevOps is such a broad and vast topic that how could any one person be qualified to write a book on this? And then about a year and a half ago-ish, you wound up giving a keynote at DevOps Days Indianapolis. And it was the first talk given that it is a keynote, I would say, I would say it's almost a value note more than a keynote. But that's okay.

And you open with the line, "The dumpster is on fire, and then the fire alarm went off." You couldn't have timed that better. I don't know who you bribed to pull that off, and that was the single greatest display of DevOps mastery that I've never seen. So the question is not who are you to write the book, it's who the hell am I to question you.

Emily: That's amazing. Yeah, that experience was hilarious. I felt like, "Okay, I've peaked as a speaker. I will never surpass this moment. I have to stop speaking now." For the record, I did not bribe anyone at the hotel. But-

Corey: Oh, yeah, your intermediary is cut out to do that for you. I understand. I understand-

Emily: Plausible deniability, Corey.

Corey: Exactly. So how long did it take you to write the book?

Emily: So this deadline was insane, in my opinion. I think overall, it took six months to write it and complete it, send it off which is really fast like that's a lot of typing at night, drinking wine in your bed, silently crying to yourself.

Corey: Oh, yeah. Write, drunk, edit, crying.

Emily: Yes. There's a whole cycle to it and it's pretty similar to coding actually. Like when you are solving a problem and building a feature and working through a nasty bug, you hit this sort of same peaks and valleys emotionally. But it's definitely more intense in a book because I think it is so final, like I am canonizing my thoughts on DevOps into a book that I cannot change in the future unless I earned a second edition.

And that's very scary that I have to like really nail it so that little mistakes don't follow me for the next five years.

Corey: Here's the $64,000 question, would you do it again?

Emily: Good question. I'm of two minds about this. The instinctual response is like, "Hell, no. I'm never doing this again." That was such a crazy experience. It's so draining. Emotionally, it's just incredibly draining. And energetically like you're just empty. But that was my first book, and Orin, my colleague, actually brought up the fact that the first book is your hardest because you don't actually know that you can finish a book.

You approach the project, you're like, "I think I can write a book," but you don't know that you can actually complete such a Herculean task. And once I finished it now, I feel almost like I finished the training course on writing books, and I feel like the subsequent books would go more smoothly for me. But I also feel like this is that sort of like post-childbirth high where you're just like, "That wasn't so bad."

Corey: And that's when the acquisition editors pounce and get you to agree to a second one.

Emily: 100%, yes. I'm definitely in that sort of blissful, "Yeah, I kind of like this moment."

Corey: Let's say that I get a wild idea that today, I'm going to sit down and pitch a book and it gets accepted. Now, not about the dummies books or any other publishers you heard of because they all have editorial standards, and well, have you met me? So I'm going to self-publish a book or get someone who doesn't know what they're getting into to agree to green light cloud for morons. And I'm going to go ahead and write that book. What should I know going into it?

Emily: Okay, what should you know going into a book? One, I think it's really important to have a unique point of view. And that can be your past experience. It can be a strong opinion for how you think the industry should continue to evolve. But it needs to be a strong point of view and that you are uniquely telling a topic in your specific story format. Like you are forming your story from your point of view.

Emily: The second thing is that I think you have to understand just how much energy it's going to take this, and with no like emotional feedbacks or little wins. There's no, like when you write a blog post, you write it. People read it. People respond. This is even a faster cycle with social media, so if I tweet something, 100 people like it, I feel good. Like that little dopamine drop-

Corey: It's very short cycle time.

Emily: Exactly. The book takes you ... It's a marathon, it's an absolute marathon, and for someone like me who's very much more a sprinter both psychologically and physically, it was a real challenge to just kind of keep pushing through.

Corey: It sounds like the exact sort of thing that I would sign up for naively, and then immediately regret that decision, and spend the next 18 months sobbing.

Emily: No, there's a lot of crying. The other thing, my editor had this great line when I turned in my last chapter and I was like, "It's done." He was like, "Congratulations on completing what more accomplished people than you could not." And I was like, "Damn, yeah."

Corey: That's heavy.

Emily: That's very heavy and it made me feel good. I was like, "Yeah, I finished a book. I wrote a thing." But yeah, I think a lot of people signed up for a book, and then figure out that they're either not the right person to write the book, or they need a co-author or they just simply can't do it.

Corey: For me, it always seems that it would be a terrible idea not because I know what it takes. I don't. I sometimes lose interest halfway through writing a tweet. But I talked to people who've written a book. You, Mike Julian, who's my business partner and a few other people. And whenever I say, "Should I write a book?" Immediately they get a look in their eyes that can only be described as haunted.

"Don't write a book," is usually what they say. Like the same tone of a four-year-old girl out of nowhere who appears to you at an airport when you're boarding who just looks at you with this thousand-yard stare and says, "Don't get on the plane."

Emily: You're like, "I'm going to die. I'm definitely going to die."

Corey: Exactly. I'm training my daughter to do that incidentally so that I get to clear the upgrade list faster so I get the upgrade every time rather than just sometimes. Let's talk about the substance of what you wrote, not the ins and outs necessarily of what you wrote, because I don't think you want to continue talking about that given you're about to go on a book tour of some sort. If not, surprise, you're going on a book tour.

But the bigger question I have is, is DevOps still relevant in this year of our Lord 2019?

Emily: Yeah. This is a debate, I think, is starting to surface and I think a lot of people have different opinions. Mine is that it is still very much relevant, but the sort of adoption cycle has shifted. So the adoption curve, we're further down on the adoption curve. I think when you and I talk about it, or people like Jason Hand, Chris Short, when we talk about DevOps, we've been in it for a really long time. Many of us have actually been involved in DevOps for a decade.

And so we've seen how it kind of ... In our experience, and talking to people specifically on that sort of cutting edge of technology with start-ups and such, we see it sort of tapering off. We see the holes in DevOps and how it needs to evolve into something else. I think part of the answer for that is SRE, site reliability engineering, but that is a little bit more prescriptive and there's some polishing that I think needs to be done around that.

But that said, I still think that DevOps is incredibly relevant for people and companies and engineering teams that are just a little bit later to adopt it. Like we take DevOps for granted, but for many people, DevOps is still a very new concept. It's still something that they're not super familiar with, and they certainly don't know where to start as far as implementing it. Is it just a CICD practice? Like what exactly is it?

And so for me, writing this book, I wrote this for a CIO in Hartford, Connecticut for an insurance company. Someone who has to think about security and compliance, someone who can't move as quickly or take on the latest and greatest technology without first vetting it. And so this is the person I had in mind writing this book for.

Corey: That is absolutely a valid approach and don't let my snark in any way detract the value first of what you've done. And secondly, from the value of transitioning your culture at an organization into something. That second one was to the general you, not you personally here. I'm not terribly worried about people worrying that I'm being too snarky or sarcastic to Microsoft's internal culture.

Corey: I think you've got that. It's been 40 years. You're a trillion dollar company. You got it. This is more for folks who are still finding their way.

Emily: Yeah, absolutely. And I think we're all sort of finding our way in different paths. One of the things I talked about a lot in the book is I don't want DevOps to be this sort of sexy thing because I don't think that really answers the question. When you or I get up on stage and we talk about a greenfield environment, and it's all unicorn and sunshine and DevOps is the solution to everything, I think that's bullshit.

And it cheats the actual methodology of stretching and beginning to solve problems in the nitty-gritty, in the legacy systems that have existed for 15 years where there's that one piece of code that no one touches because it works and no one knows how it works, and obviously, we can't mess with it at all.

And I think a lot of code bases, probably the majority have those sorts of problems and tendencies. I've never met a code base that isn't a complete shit show that's longer, older than like two weeks. And so-

Corey: Even when writing things in the middle of it, it's a lambda function. It's four lines and I'm on line two, and I'm looking at this going, "This is a tire fire." And I know it going into it. That doesn't make it better, but at least, I'm not blind.

Emily: Yeah, there you go. There you go. So I think it addresses some of those nitty-gritty problems and it kind of gets into the more difficult challenges, which I think are more interesting and fun to solve.

Corey: One of the challenges now is that you see that the terms have most been co-opted by people who are driving transformational change and grifters. And it's sometimes challenging to figure out what's what. You see a bunch of sysadmin teams that are just rebranded as DevOps, which if that's you, frankly, good for you because we did a survey a few years back. I think it was DevOps Days Austin, but don't quote me on that. And it turns out that with DevOps in your job title, even though the job is functionally the same, it winds up being something like 30% pay bump.

So yeah, if you want to call me something that pays 30% more, I don't really care what that is for the most part. I'd prefer nothing scatological but I'm willing to flex if you can go to 40%.

Emily: I like that. I like that resume-driven development.

Corey: Absolutely. Where do you think Kubernetes came from? But now, we're starting to see this branched out into other things with DevSecOps and FinOps, and I'm sure people put QA into it but there's no way to get that without coughing up half a lung midway through the word. QAOps is not really a thing. And at some point, it starts to take on a parody of itself element where it's no longer about having these different cross-functional teams, collaborating and communicating in a way that makes sense for everyone so much as it is, talk to other teams and do your job.

And that seems to get lost somewhere along the way in a profound way. It feels that by slapping a new shiny DevOps-ish label on something, that is going to magically fix all the broken processes and lack of cultural advancement going on in any given company, and you're not going to fix anything but you're going to feel better for doing it.

Emily: Yeah, I mean absolutely. I think you bring up quite a few of the problems that I have with the current sort of status of DevOps in the industry. I think it's suffering from a lot of the same issues that Agile has and is suffering from, which isn't surprising. DevOps was born out of Agile.

And because it is such a ... It's not a prescription. There's no like, "Do this, this, and this, and then you'll be fine." It's more, "Yeah, work together, assholes. Do your job. Talk, communicate." If you're a manager, be a good one. Protect your engineers. Give solid market salaries, incentivize the team in the same direction so that they're not fighting each other.

These all seemed like very basic principles of just good business, but the truth is, and this goes back to my power-lifting days. When I was lifting, some of the best advice I ever got was it's very easy to do the complicated thing. If you're doing like, let's take diets. A complicated diet is in some way so much more easy to actually go through it than to experience because you have like your exact prescription for what to eat, and what not to eat, or you're just on a seven-day fast or something.

It's much harder to do the simple thing, and I think DevOps is the embodiment of doing the difficult simple thing. The actual principles of DevOps are pretty simple and straightforward. But implementing them, it takes a lot of energy and a lot of collaboration. And those things like the human problems. The human problems of tech are the most complicated, in my opinion.

Corey: Absolutely, and I'm reminded of all the early days of DevOps in the serverless environment that we're seeing today. I keep hearing echoes of a blog post as it lurks around the internet. And I'll throw a link to this in the show notes called DevOps is a poorly executed scam. It's from March of 2011. So it's by a guy named Ted Dziuba, and his entire approach was that it felt like the second coming Agile, but no one was asking for money. There was no conference. There was no books. There was no training. You've gotten a bunch of buzz but no one is selling anything to you in this space.

Well, in the DevOps world, I think we fixed that. And it just took a little longer than people expected it to. But now I'm starting to see some of the same hype style stuff with serverless, where people are going significantly out of their ways to build things that turn into a somewhat insane and ridiculous approach, where everyone is trying to build things in Byzantine ways, as the best expression of ideological purity. And often, they blame each other for not wishing hard enough to build the thing and being traitors to the cause.

And it just winds up being something awful and unsightly. I don't think that we're quite there yet but I'm starting to worry that that's where we're headed.

Emily: No, I think that's fair. It's funny you say at the beginning DevOps didn't have a ton of vendors trying to sell things. I am actually coming to the perspective that engineers as a whole are not extrinsically motivated with money. I think we all want to be paid fairly and we all want to be able to support our family and do things that we care to do.

But that aside, like throwing extra money at an engineer doesn't typically gets you a ton of long-term motivation.

Corey: Repeating that again will get you highly motivated engineers to silence you who want large piles of money thrown at them.

Emily: That's fair. That's fair, yes. But that said, I think-

Corey: Paying me large piles of money probably won't solve your problem but let's try it.

Emily: Yeah, exactly. What was that study in like a couple of years ago they found that after like 75,000 or something, extra money doesn't bring you any of their happiness is something like that?

Corey: True. That is normalized as a baseline load. It tends to vary based upon major metro areas, obviously. But you're right. Past a certain point, you are going to be measurably happier with two yachts instead of just one.

Emily: Yes, exactly. Throw that second yacht, I mean. Oh, my god.

Corey: You're truly wealthy when you can walk in and pronounce it phonetically as "yatch" and no one corrects you when you're making the purchase.

Emily: God, that would be special. I would like to see that. Can we do that? I feel like that would go viral, Corey.

Corey: This is the entire problem, is that whenever you and I hang out, it immediately leads to hilarious and disastrous encounters for everyone around us. We were at Build somewhat recently, and we happened to show up at the same off book meet-up near at the same time. And you walked over and, "Corey," you start seeing the new pin on my lapel and start fidgeting with it. And someone just asked, "Do you know him?"

And the obvious immediate response was, "No, why?" Yeah, it went super well. And everyone just sort of looks at us and stares and we're terrible influences on one another which tells me we absolutely need to collaborate more and inflict us on the world somehow.

Emily: Yes, I think both of us are very comfortable in awkward social situations, and then it becomes compounded when we're together. And then we kind of like ... I get joy out of making things a little weird. I just think it's funny.

Corey: If we're not in an awkward situation, we make it that way. That's who we are as people. That is fundamentally our nature.

Emily: Yes, which I think a lot of people just looking at me don't think about. They don't see this sort of mischievous side but it's definitely there.

Corey: Oh, absolutely. It's one of those things where it's the kindred spirit approach where we're effectively ... We see the world in the same bizarre way. One of these days, we have to find some kind of collaboration opportunity. I don't know if that's you give a talk, or something and then I give the rebuttal to that talk. Picture the political state of the union things. We have the opposition party go up and do a "well, actually" afterwards, only that with more swearing and screaming.

Emily: Yes, but we should do it in like a British parliamentary way, like where we can kind of just yell at each other.

Corey: Oh, absolutely and go round and round and round and round, forget a deadline is approaching and the continually push it back which is I'm told how you write a book.

Emily: I want to come back to that. But I wanted to do like I have an idea for a conference where it's basically debates for charity. And so you get people like I want two people that sign the Agile manifesto, like Kent Beck and someone else to debate whether they think it still holds or something like that. And then we pay to attend and donate all the money to charity. I think that'd be fun.

Corey: I love that entire idea. I think that's something that is phenomenal. I'm always a big believer as well in doing things like that for charity because let's not kid ourselves, no one is going to pay money to me to watch my ridiculous horse shit. But they might tolerate my ridiculous horse shit if it benefits a good cause.

Emily: Definitely. And I think a lot of us struggle with, okay, many developers and engineers didn't come from well-off families. A lot of times, especially with code schools, you see someone being the first person in that my family to make X amount or six figures, or be able to support their family and bring them out of poverty.

And so I think a lot of us struggle with, "Okay, we're here. We make a really solid wage, how do we make sure that we're giving back and we're not just living in a sort of self-centered and selfish way that we actually embed altruism into our everyday lives?" That's something I struggle with, like I don't always know the answer to that.

So, yeah, there's a lot of opportunities for charity in this industry.

Corey: Absolutely and there's a reason that every year, I do the charity T-shirt where it's something snarky and sarcastic for last week in AWS, and all proceeds go to a random charity. Last year was at St. Jude Medical Center for kids with cancer research. That was fun. I didn't know who it's going to be this year but I have some ideas.

But doing something like that that's larger than just my own ridiculous nonsense/corner of the internet, first, it feels good. And secondly, it helps justify, if only to myself, that I can take this ridiculous love affair I have with my own personality/sound of my own voice and use that as a force for good beyond just insulting things all day long.

Emily: I think that's wonderful, yes, insults for good. I want to loop back to the procrastination in writing because this was fascinating to me. If I had to write a book on writing a book, it would be lessons I learned from staring a carpet fibers and it's because you struggle. Like you sit there and you stare at the computer, the blank page, the word document because I had to type in Word. Feel sorry for me.

And then you're like, "Ugh," you go through this phase where you're like, "I'm not the right person to do this. I have nothing to say on this. I'm the worst writer in the world." And then you like roll around in the carpet a few times. If you have a partner, you'll talk to the partner like, "I don't know. I shouldn't be the one writing this. This isn't right."

Corey: And then the crappy partner, they come back with, "You're absolutely right. You're good at nothing."

Emily: Don't stay in this those relationships, please.

Corey: Get out.

Emily: Yes. Talk to me. I will help you.

Corey: Exactly. It's a DevOps metaphor.

Emily: There you go. So, yeah, and then you kind of roll around and after five hours of moaning, you pull your shit together and then you type some words. And about three pages in, you're like, "Oh, my God. This could be a book in and of itself." It's a fascinating sort of process.

And procrastination, I think, actually plays an important role and that you're not actually just sitting there moaning. I mean that's sort of the output. But I think the process of that and the journey of procrastination is you're truly ruminating on what you're going to say. And it doesn't feel productive but I think your brain actually needs that time to sort through random thoughts before it can be coherent and clear written.

Corey: This aligns very heavily with a pair of conflicting Twitter takes I saw going around a few months ago where on the one hand, it was, "Well, we just paid a lot of money for this conference. So, thanks for telling us you threw the talk together that were paying to see the night before. Awesome, thanks." And the counterpoint to that is, sure, they finalized their slides the night before but they'd been thinking about this talk for months and months and months and months and months. And then finally had a forcing function of building slides because you're not paying for the slides, you're paying for the viewpoint, the perspective, the story that goes along with it. So, I'm extremely sympathetic to both sides of that.

Emily: Yeah, and I think it's funny this is apt timing. I had a nightmare last night that I haven't ... I'm keynoting DevOps Days Toronto in like two weeks, and I haven't actually started the talk. And I know what I want to say but I had a dream that I didn't finish the deck and I had to talk with just like a six-slide deck and it was bad. That said, we will not let that happen.

Corey: The sneaky way around that incidentally is to give a talk with no slides. If you bring in other couple devices to do this, and talk to me after the show, I'll give you a link for these, that you can plug into a projector and it turns the thing into a paper weight because it lets the magic smoke out of the back. "Oh, no, the projector exploded. I guess I won't do my slides and I'll just do it without slides." Then all you have to build is a title slide.

Emily: Fascinating. Okay, yeah, that's one approach.

Corey: A way of the weasel.

Emily: That's not my approach. We've talked about this.

Corey: Yes, because you're good at things, and I'm just really, really, really good at distraction.

Emily: Well, and you're so comfortable sort of riffing on stage whereas my personality doesn't work like that, like I'm contextually funny but I'm not like a ... You could be a standup comedian. I wouldn't be a good standup comedian.

Corey: What do you mean could be? I go to Seattle once every month or so and do Tonight in AWS. I make fun of cloud computing for 40 minutes. It goes super well, all proceeds to charity.

Emily: There you go. And most of us just we don't have the skillset or confidence to do that.

Corey: The only skillset I'm convinced that you need is a complete lack of self-awareness and/or shame. And as long as you've got that, everything is great. Like the crowd boos. They pull you off the stage and then people ask you like, "How did it go?" "It killed. Everyone was cracking up," because you're always a legend in your own mind.

Emily: It's true. It's true. And related to that, I think writing a book is interesting. When you said you're asking your friends like, "Should I write a book?" I think of the interesting byproducts of writing a book is the credit that other people give you for free. Like I'm more or less the same level of intelligence and filled with the same level of knowledge that I was six months ago. Now, I just wrote it in a book.

But when you're an author, people give you a lot of credit. They're like, "You must know what you're talking about. You're an expert."

Corey: Really. What book did you write on that because I wrote a book.

Emily: Exactly. That's exactly it. And so it's this sort of ace card or this trump card you can kind of holds if you need it.

Corey: It's definitely one of those appeal to authority things. And there's also, not for nothing, but again, I've written no book of any stripe, color, et cetera. At this point, I might be able to write a matchbook and that's as far as it goes.

But it also has its own gravitas when you're with the known publisher as well. I mean the For Dummies series was transformative. This dates me, not like anyone else will, but back in the 90s, I got started with Linux For Dummies. That was my first outing to the idea of a Unix-like operating system. And it was captivating.

Emily: I love that. Yeah, I'm excited about The Dummies. And I think it doesn't get the credit that it deserves and that it fills a piece of the market where you can be a complete beginner. You don't actually even have to know much about tech to read my book, like you can glean a bunch of things about management in there, about being a leader without a management title, about energizing and galvanizing your team no matter what your job is.

And so I really tried to write it and meet people where they were at to make sure that it would be applicable to them no matter where they are along the process, what their background is, what they look like, where they came from, how they got into tech. I didn't want any of that to matter. I wanted it to be completely open and to answer the questions that anyone coming to it would have.

Corey: And I think that's absolutely the sort of thing that is important for people to hear. We've been joking a bit and snarking at one another as we do. It is our way. But there's extreme value in making things like DevOps and modern technological practice and culture shifts accessible and approachable to people on a wide variety of spots on the spectrum.

And I don't want our sarcasm and joviality to wind up occluding that.

Emily: Yes, absolutely. And I think tech in general could do better at this. I think, because we are such an intelligent-focused industry where we measure ourselves and each other based on what we know. It's very difficult for us as engineers to say, "I don't know that," or, "I've never heard of that," or, "Explain that to me."

Corey: Or the mistakes things that are common to the industry as prerequisites, "Well, I don't think I'd be very good at doing DevOps style of work because I'm not very good at coming up with puns." Wait, what?

Emily: No, absolutely. We think that you need to have X factor on your resume to do whatever else. And I think that's bullshit. And something I've been focusing on at Microsoft is IaaS, usually stands for infrastructure as a service. And I apply it as idiot as a service.

So, I am open to looking like a moron and asking the dumb questions so that everyone's on the same page because if you don't have that foundation of clear communication and context, and understanding why you're working toward a goal, the whole thing is going to fall apart at some point. It's a house of cards.

Corey: And something I'm trying to focus more on is giving more airtime to my stupid questions. Most of the time, it turns out it wasn't a stupid question at all. There's something there that I'm not getting and I need to understand. That's the point of asking a question.

But it's easy to lose sight of it when people aren't doing that in the open, where they realize they spent three hours messing around with something and realized that they missed a comma or they're using the wrong version of a thing and the documentation says in giant bold red letters "Do This Thing First". And if you don't do that, surprise, it doesn't work.

And everyone has those experiences of going back and forth rapidly cycling between, "I am a genius and how am I allowed outside unsupervised." It's back and forth so quickly.

Emily: I know. I talked about this in my Dr. Seuss talk where there's this rhythm to software developments. And it's like, "Oh, this should be easy," no problem, you dig into it a little bit and then you're like, "Uh, that doesn't seem so easy or straightforward and then you're like, "Help! I need help! I've no idea how to do this. I'm an idiot."

Emily: And then when you actually solved the problem either individually or as a sort of team, you feel like a god, like a god of all things machine.

Corey: So, if people want to hear more of your sage thoughts/scintillating wit, where can they find you?

Emily: You can pre-order DevOps For Dummies on Amazon. You can find me on Twitter, @editingemily. @editingemily was my writing business before I moved into tech. And you can watch my talks at emilyfreeman.io.

Corey: And we will put links to all of those in the show notes. Emily, thank you so much for taking the time to have a chat with me today. I appreciate it.

Emily: Thank you. I love talking to you.

Corey: Emily Freeman, Senior Cloud Advocate at Microsoft, of which Azure is a cloud. I'm Corey Quinn. This is Screaming in the Cloud.

Speaker 1: This has been this week's episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com or wherever fine snark is sold.

Speaker 4: This has been HumblePod production. Stay humble.

View Details

About Anna Spysz

Anna Spysz is a writer turned software engineer at Stackery in Portland. When not software engineering, she likes to travel, play music, and kung fu fight for fun and profit.
Links Referenced:

  • https://www.stackery.io
  • https://medium.com/@annaspies

Transcript

Speaker 1: Hello, and welcome to Screaming in the Cloud with your host, cloud economist Corey Quinn. This weekly show features conversations with people during interesting work in the world of cloud, thoughtful commentary on the state of the technical world and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this way by Anna Spysz, a software engineer from Stackery. Welcome to the show, Anna.

Anna: Thank you.

Corey: You're one of those people with a fascinating and "non-traditional" path into tech. What did you do before you were a software engineer?

Anna: Well, a lot of things. I started my career mostly in the sphere of writing and journalism with some translation on the side. I was a liberal arts major and you know how well that pays, so I did a few stints as a journalist. I went freelance for a long time, which was nice but it was a hustle. It was very much a hustle. And the last actual job I had, I would say, was as a tech journalist and that kind of brought me in the sphere of startups and tech and all of that.

Corey: It's interesting watching people who have come from alternative backgrounds. Namely, the liberal arts seem to be a terrific breeding ground for this where you don't have the same exposure in many cases of the, I guess, traditional paths for startup software engineer types, such as ... Which usually looks something a lot like, well, I got my first computer when I was five years old. I was always playing around with that. I went to an impressive college everyone has heard of because I've mentioned it three times in this sentence already. I've studied computer science because of course I did and that's why, after all of this, I'm making the world a better place now at Twitter for Pets.

And it's a common story to the point where we almost start to think that's the only way to get there; where if you didn't have that background and oh, you didn't start using a computer until you were in your teens or your 20s? Oh, forget it. You're never going to be a success. And that is very clearly not true.

More to the point, something I've always seen is that when you come from a non-traditional background into this world of technology you bring things from alternate fields that really tend to shed a light on things. How do you notice that your background has impacted what you do?

Anna: Well, first I kind of want to add on to that because while I completely agree, I feel like I've always been a little tech-adjacent. Even as a freelance writer, I was still coding my own website and just hacking enough together CSS and HTML to make things work. You know, make little sites for friend's bands and stuff like that. I had a GeoCities page, not to date myself, back when I was a teenager.

Corey: I remember those days. Whatever happened to those?

Anna: Right? I know. I believe there's a little thing called the internet archive that might help you with that, but yeah, so I feel like it's totally possible to go from a non-technical background to a career in tech but I also feel like that interest has to be there. Like if you're doing it just for the money, just because you're like, oh, developers get paid well. I want to be a developer. I don't think that'll work in the long term. But if there's a hint of interest for a long time, I think that at least from what I've seen in the colleagues, they'll be a better indicator of long term success in this field.

Okay, that was diatribe on that. What was your actual question?

Corey: No, I like the diatribe. Let's stick with that for a minute. I think that's ... The problem that I had with that approach and ... First, I don't think you're wrong. I think there has to be some kind of affinity for technology or you're not likely to go very far in this, regardless of who you are. If you don't have a passion for it you don't enjoy it and you just show up and drag yourself in front of a computer for 40 hours a week because hey, I hear it hears well, it doesn't go super well. You want people to have an interest and a passion for this.

The challenge I see is when that starts translating into job requirements. Where it's you most absolutely be in love with technology. Oh, you work a 40 hour a week job doing software development. Cool. Awesome. What get out projects are you doing on the weekends? And that isn't something that's particularly fair to folks with, I don't know, families in my case or lives in other people's cases. And it just seems that it turns very quickly into something that's not particularly pleasant.

Anna: Yeah. I completely agree. I think I'm very lucky at Stackery. Half of us ... Half the people there have kids and work/life balance is very important. And if I code on the weekend it's for some pet project that I'm really excited about and feel like doing, but most of the time I'm just taking a hike on the weekend or watching Game of Thrones. You know, having a life.

Corey: I think you're right. There's an awful lot of value to having interests that go beyond tech, which brings us back around to the question of how do you find that having a non-tech professional background has impacted what you're doing professionally now?

Anna: I think the most significant part of my background is being having been a professional communicator. It definitely helps in what I've been doing. A lot of what I've been doing is on the documentation side of things as well as actually writing code for our app. Part of this is because we're still a small startup. There's not that many people. It's something I fell into and I feel like I have a passion for communicating how to actually use the things that we're working so hard on. How to make it as successful to as many people as possible, how to have somebody come in that might be new to tech and be able to read a tutorial and actually succeed at the thing they're trying to do rather than spend hours fumbling through the docs trying to find the one thing they're trying to do and it's something so absurdly simple that the person writing the docs didn't even put it in there because they're like oh, everybody knows that.

But it turns out not everybody knows everything about your app like you think they do, so yeah. I think it's definitely helped in that. Just being able to talk about the product, being able to make tutorials and communicate what we're doing in a way that brings people in.

Corey: Once upon a time, back in the year of our Lord 2012, I watched a conference talk that was borderline transformative. It was before I was on the speaking circuit, myself. It was … from Jordan Sissel, the author of Logstash, which later became part of the ELK Stack. He works at Elastic now, but that was one of the inspirational talks for me that made me think wow, I'm not super good at computers but maybe I could learn to give talks one of these days because his entire theme of his talk wasn't about this is how the code works, this is what it does under the hood, look at how clever I am. His thesis, instead, was if a new user has a bad experience, that's a bug. Your documentation is at least as important as your code.

And I was sitting there and my first response was holy crap, that's transformative. That's important. That goes beyond any individual technical implementation. And secondly, right on the heels of that was holy crap, again! I'm not a moron; this stuff is just badly documented. And once you start seeing that, it's almost impossible to not see it. And generally speaking, in the modern cloud ecosystem, that is getting better but I don't think it's there yet and I still see there are things that wind up not quite getting where it needs to be.

I will say when looking through Stackery's documentation, it's very well written. Your impact is felt.

Anna: Aw, thank you. At least traditionally, it feels like most documentation was written by engineers and the traditional computer science degree engineers you spoke of earlier and it kind of shows because it's ... Well, I won't mention some certain cloud giants who make a possibly deliberate puzzle of their documentation, but yeah, it's not intuitive. It doesn't tell a story. It doesn't guide you from one step to the other. It's very, oh, here's how we do this thing that we know all about and you know nothing about but here's just some raw data.

Corey: Yeah, well one ... I will throw stones at a large cloud provider. Where recently, AWS ... Recently, last year, AWS dropped all of its documentation into a bunch of GitHub repositories and said hey, if you find a bug in the docs you're welcome to fix them, which I'm conflicted on. Because on the one hand, that feels like an awesome way of opening things up and letting us actually get things that annoy us fixed. On the other, your company keeps flirting with a trillion dollar market cap and you're asking me to volunteer for you? What? It feels like they're almost ... And I don't think this is there intent, to be very clear ... But it does feel in some ways like they're trying to put the burden of good documentation onto the community rather than hiring people who are, themselves, excellent writers.

Anna: Right. I mean, I think it's not just them. I think it's just docs are an afterthought. Let's build our thing and then eventually we'll teach people to use it maybe or they'll figure it out. But that's ... I don't think that's the correct approach.

Corey: Once upon a time, because I have many sins and I was forced to pay for them, I was heavily involved with the Freenode IRC Network. I was volunteer staff for about a decade, which, pro-tip, don't do that. It was fun, it was enjoyable in ways, but it also had challenges that came with it. But as a part of that, I learned to ask questions in textual formats pretty effectively because, you know, whenever you hit a bit somewhere in a software program or you get stuck, no one has the context of everything you've been working on in your head. And spending four hours getting them up to speed on that context just isn't going to happen.

So, stripping it down to a skeleton case of I'm trying to do X, I'm seeing Y. When I look at the documentation, this thing should be happening and oh, dear, I found a bug. And by framing the question that way, you either wound up realizing that you misread something and ... More often than not, in my case anyway ... And you were wrong and you never have to ask anyone that question as your building that. Or you discover there's a bug either in the code or in the documentation, more often.

Once you keep running into that, it's easy to assume malice where people are writing crappy documentation because they want to be exclusive keepers of the fire, I guess. And I don't think that's ever been anyone's intention. I think, instead, it's largely been around just ... You're right. It's an afterthought. How do we fix that?

Anna: Yeah. Yeah. It's just simple neglect and it makes sense in startups where there's five people and they're working really hard just to build the thing. But I think, yeah, how it fix it is put a priority on that and hire people who are not just engineers for the documentation. You almost need this combination of somebody with some sort of communication background that also understands the technical aspects of the product.

Corey: That seems like a good point for us to segue. As important as documentation is and having a background where you can speak to folks who do not, themselves, come from a steeped in code background where all they speak is in ones and zeros. Let's talk about actual software development for a minute, here.

You've been working lately, to my understanding, on the idea of local development, particularly for Serverless. Tell me a little bit about that.

Anna: Okay. Yes. This is a pretty exciting thing we've been working on at Stackery. It stems from this problem that Serverless changes local development. So, in a traditional workflow you had your code, ran local hosts, you did your thing, send it off to DevOps when it's ready to be deployed. But in Serverless it's very different because you have cloud resources that are part of your application. You have to run your code against those managed services. So, this is the problem we've been trying to solve lately and it's something we've run into ourselves because we're a serverless shop so ...

Corey: To throw my own bias out there, and I'm absolutely not contradicting you, your company, et cetera, but my approach has always been that local development, for me, has been largely a bit of a boondoggle when it comes to anything that approaches cloud because you're going to wind up first building crappy versions of whatever those cloud services ... Endpoints look like. There's going to be inconsistencies that creep in. It turns out that I'm super good at mocking cloud services in a sarcasm sense as opposed to doing it in code. That misunderstanding led to my entire current career path.

But what I don't understand, personally, is the value of local development for most work flows. Can you help me get there?

Anna: Okay. It's about speed, basically. So, yeah, like he said, mocking ... First of all, it takes time. You have to mock out all your services, get the events, et cetera. And it's not faithful to your actual cloud of resources. Your permissions probably won't match, you know? You have to hard code your environment parameters and that also takes time.

The other option is just to code and deploy, code and deploy. But that takes a few minutes between every deploy and it's like the waiting to deploy ... Waiting for your deployment is then you waiting to compile, so that's also wasteful.

Our solution is-

Corey: I consider that my coffee time.

Anna: Well, true, but if you're doing that 10 times a day then that's not-

Corey: Oh, yeah. I start shaking around two in the afternoon.

Anna: Yeah. And you get distracted and you open your email and you're like oh, that thing's ready. What was I doing? So, our solution to that right now is you have your local machine, you have your Lambda code running on that and it is actually interacting with deployed cloud services. So, say you're working on a new app, right? You first architect it. You get some API gateway, going to a Lambda going to a Dynamo Table. And you architect that, you push it up to the cloud, you deploy it out, it's on AWS, you have the endpoints, right? You have all the permissions set in your Template Dynamo or whatever your using. And it's out there.

We've got a command on our CLI that just came out called Stackery Local Invoke. And when you run that, you are testing your actual code on your local machine against those cloud resources. So, you can read data from your table. You can write data to said table. You can trigger events from that endpoint. And what you're doing is iterating and then running against those resources with the same permissions, same environment variables in real time and it's like a three to seven second delay rather than a few minutes for every deploy.

I mean, in my own workflow I've been testing this extensively for the past few weeks, just building my own pet apps at work for our change log and stuff. And it just ... It's a game changer. I can actually see what I'm doing interacting with the cloud. I can catch like, oh, I missed a ... I forgot to close my brackets and the whole thing crashed and I didn't have to deploy this out just to find out I forgot to close a bracket. It's a huge thing for me, anyway.

Corey: This is probably going to wade into territory where I start getting angry emails, but it sounds almost like this is not as much local development as it is almost hybrid development, which I don't know I've ever heard the term before-

Anna: Yeah.

Corey: -Please don't tell me I just coined that.

Anna: You probably didn't but you may have.

Corey: Uh-oh. It feels like whenever I talk to people historically about, well, I'm not a big believer in local development; I prefer remote development, especially given that I tend to travel with an iPad. I tend to do all kinds of different things and we all talk in the SRE/DevOps/test admins/whatever world we're calling it this week about you don't want cattle, you want pets but your laptop is the most precious pet of all because your development environment takes four days to build then okay.

So, I want to be able to build all this stuff on my laptop when I'm on a plane and there's no Wifi. Cool. That's the use case people bring up and I understand and respect that, but I fly 110 thousand miles a year. I don't find myself mid-air without Wifi needing to write and deploy code all that often. It feels like that's happened maybe once in the last three years. And, admittedly, I'm not a software developer as a primary function but it is something that strikes me as a bit of a constrained use case.

So, what your describing does, to my understanding, require access to the internet.

Anna: Yes. Yeah. Very much so. And that's the thing. This speeds things up to the point where hopefully you finish what you're working on before you ever get on the plane and then you can enjoy your in-flight movie.

Corey: Absolutely. Or everyone has a plane drink. Mine has always lately been the gin and tonic out of a-

Anna: Yes.

Corey: -Mistaken belief it will protect me from whatever the rest of the mouth breathers on that plane are infected with. It doesn't work but I pretend it does and it makes it better.

Anna: Oh, yeah. Yeah. I mean, especially on overseas flights. Just keep the red wine coming and that's it.

Corey: Exactly. And it's great because it leads to some hilarious commit messages right around drink four.

Anna: Yeah. And that's probably not the optimal way to work on your app. I just want to put that out there.

Corey: You'd think that. I get incoherent but my code gets better. To be fair, it's my code.

Anna: Really?

Corey: There's really nowhere for it to go but up. But usually baseline is it's hit rock bottom and started to drill.

Anna: Well, I mean ... It's like ... Was it Hemingway said, "Write drunk. Edit sober." So, as long as your code reviewer is sober you're probably good.

Corey: That's a whole separate argument for another time. I guess from ... What you're describing is local development is invariably done on your laptop? Or, I guess to some extent for what I've been doing lately, I've been having the "local" portion of that living on an EC2 instance just because that's somewhere in a constrained environment, I don't have to worry about data leakage, I don't have to worry about dropping my laptop into the bay like the last time. And it gets to a point pretty easily where at that point it's "local development" but even with a hybrid model, everything lives in the cloud somewhere else. Is there a use case where I want to go ahead and be able to do that where something has to be run locally?

Anna: Well, I mean one way or another you're committing to git all your stuff somewhere safe hopefully. Yeah. If you're committing the git then you can drop your laptop in the ocean and it will still be there. Yeah. I don't see why you would want to go fully local when ... At least when you're relying on managed services and you want to make sure not only your code works but all the permissions are correct; that you're able to interact with the resources you need in the same way you would in a production environment.

Yeah. No, what you said about hybrid is completely accurate. The actual code you're working on, the Lambda you're writing, that is on your machine at the moment but it's also in git, so it's not fully on your machine. But everything else is in the cloud and the only reason you have your Lambda code locally is just so you can iterate a lot faster.

Granted, I have not tried your EC2 method but it sounds a little bit more complicated; just going to throw that out there.

Corey: Sort of. I mean, not to be prescriptive. I think everyone has a development workflow. But, again, I come from a grumpy sysadmin background when we call sysadmin because there isn't really a second kind. And my editor of choice for 15 years now has been VI or VIM. And that's one of those things you can use anywhere.

I would agree with you. If I were doing something like Visual Studio Code or something more complicated with a full IDE that has a visual representation, yeah that would be a living nightmare. I'd probably want to do something like run it on ... If I wanted to let's say model, I'd use something like an Amazon Workspace over a tool desktop that lives in the cloud. But for what I do and the way I do it, probably incorrectly for modern iterative purposes but retraining me takes work and time, it doesn't matter. I need something that I can SSH into and that's all I need for any form of development that I tend to do.

I am absolutely not suggesting that that is a best practice. I'm not suggesting that would necessarily work for other folks. But for me, at least, it seems to have worked out okay.

Anna: Yeah. No, that makes perfect sense. Like, I am a VS Code girl. I like my IDE. I like having all my folders neatly organized, et cetera. I like being ... I'm also a very visual person. I just like seeing everything running in the terminal along with my code. I think that's why prefer this way. It also ... I mean, to be fair it's not as simple as I make it out to be in every case because say you have an RDS behind a VPC, right? How are you going to interact with that? You need to run a tunnel most likely. It can get more complicated for sure. But that's also doable.

Corey: Yeah. I started playing around in the last six months or so with Visual Studio Code and oh my stars. This is one of those things that is ... I can see how it has the potential to be transformative. I need to un-learn a whole bunch of habits for that to start being something that I use at an ongoing basis and please, please, please, please, please if you work at Microsoft and you're listening to this, get me a version of this for the iPad. I know. It's not going to be able to compile a bunch of stuff down locally. I mostly write in Python. I don't need it compiled. Sure, it's going to be limited but it's so pretty and it's so much nicer. Please build that.

Anna: Yeah, that would be sweet.

Corey: The more I use it, the more I'm starting to realize that there is a whole other world out there as far as using a tool like this that lets you have a bunch of things side by side, being able to take a look at exact runs, without getting this all set up in a bunch of Tmux panes or a bunch of deep-dive VIM configurations.

Again, I do not care prescriptively what editing environment someone uses. One of the worst interview questions I ever heard was what text editor do you prefer? Which is more or less an invitation to yeah, let me either agree with you or call you a moron. It doesn't work that way. Some of the best developers I've ever known have used JOE, which is an ancient code editor that goes way back or Nano, which everyone likes to make fun off because it's just a what you see is what you get, more or less. And everyone likes to taunt these folks. It's like what are you? Not good at computers? What could you possibly build? And the answer at one point was non-trivial parts of Google. Okay, then.

Maybe judging people by the tools they use isn't necessarily the best path forward, particularly in a career context.

Anna: For sure. For sure. And there's also the context ... Like I would think most current CS grads and definitely code school grads, at least from everybody I've met that went through code school, they're probably going to use an IDE. It might be VS Code, it might be Webstorm or something high term, whatever. But that's the way that I've seen that development is being taught these days. So, it's probably becoming more and more common and VIM ... You know, if anybody ever figures out how to exit out of it, it's going the way of the dinosaur maybe.

Corey: The one I've always used was that my wife uses Emacs and I use VIM and we've been together a decade because neither one of us knows how to quit. And it feels like there's a bit of truth to that. I did hear once that someone was debating whether they should hire someone because that person used Sublime Text and everyone else used other stuff and well, Sublime I think was 75 bucks or 100 bucks for the editor for a year. At which point I just sort of stared at them. What are you planning on paying this person, buddy? They're going to embezzle more than that in office supplies because all of us do and the coffee budget will be multiple times of that. Are you sure you're focusing on the right part of this story?

It's weird having these conversations with people who just don't seem to get the larger picture.

Anna: Right. And it's like you can change your workflow if everybody at the office uses Webstorm. Okay. That's fine. You're not used to it? You can learn it in a few days.

Corey: Right. I don't care if your primary means of writing code is on a legal pad with a pen. As long as you can solve the problem and deliver the value that you need to deliver, I don't care what the tool looks ... At some point, it turns into an okay, are you actually going to do any work? But if you're trying to judge people, for example, how quickly they can type that has never been the limiting factor of any developer I've ever known or at least wanted to work with because that's not writing code; that's data entry at some point.

Anna: Yeah. For sure.

Corey: Before we go, is there ... I guess, do you have any words of wisdom for folks who might find themselves teetering on the brink, shall we say, of coming from liberal arts or writing background who want to wound up potentially dipping a toe into tech? What advice would you give them?

Anna: Sure. Well, my journey was I went to a code school about a year and a half ago now, but that was only after I had been doing stuff like Free Code Camp and Code Academy, et cetera, online courses for three, four years off and on. Getting the basics of HTML, CSS, JavaScript down. I feel like by the time I decided on code school ... It's also a perfect turn of events that I finished a big contract job. I had a cushion. I could take three months out of my life and not have a job and do this full-time.

But it was also the point where I realized okay, I do enjoy this. I'm serious about it. And I'm the kind of person that needs a teacher over me saying do your stuff but also just people I can ask questions of when I get stuck. So, that's when I decided on a code camp. But, yeah. I would say before you spend a whole bunch of money on some sort of a in-person course, take the time to go through something like Free Code Camp and see if this is actually what you want to do. See if this interests you, see if you keep going. Or are you just making yourself get there? Because then maybe it's not for you or maybe there's some sort of tech-adjacent role. Maybe you want to work at a startup but in a different role than an actual engineer.

Yeah. That would be my advice. Make sure this is what you want to do and then when you do think that's where you want your career to go, invest in it. Go through some sort of course. If you're the kind of person who needs in the classroom instruction like I do, do it. It'll be worth it in the end and, yeah, the job hunt sucks at the very beginning but once you get that first role ... And hopefully you'll get as lucky as I did in mind. But then you can only grow and keep learning.

Corey: If people want to hear or possibly read more about how you view these things, where can they find you?

Anna: I occasionally contribute to Stackery's blog. I have a Medium. I can't remember the link to it.

Corey: We will throw a link to it in the show notes.

Anna: Yeah. That's fine. That's about it. And, yeah. Stackery Docs. There's a big chunk of my stuff in there. Do a tutorial just for the literary value, sure.

Corey: Exactly. Yes. Reason this is a good ... English classes years from now will look at documentation as an example of the modern era.

Anna: Oh, God. I hope not.

Corey: Anna, thank you so much for taking the time to speak with me today.

Anna: All right. Thank you, Corey.

Corey: Anna Spysz, software engineer at Stackery. I'm Corey Quinn. This is Screaming in the Cloud.

Speaker: This has been this week's episode of Screaming in the Cloud. You can also find more Corey at ScreamingintheCloud.com or wherever fine snark is sold.

View Details

About Scott Guthrie

As executive vice president of the Microsoft Cloud + AI Group, Scott Guthrie is responsible for the company’s computing fabric (cloud and edge, including cloud infrastructure, server, database, CRM, ERP, management) and Artificial Intelligence platform (infrastructure, runtimes, frameworks, tools and higher-level services around perception, knowledge and cognition).

Prior to leading the Cloud + AI Group, Guthrie helped lead Microsoft Azure, Microsoft’s public cloud platform. Since joining the company in 1997, he has made critical contributions to many of Microsoft’s key cloud, server and development technologies and was one of the original founders of the .NET project. Guthrie graduated with a bachelor’s degree in computer science from Duke University.

Links Referenced

  • https://azure.microsoft.com/

Transcript

VO: Hello and welcome to Screaming In The Cloud with your host, cloud economist, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming In The Cloud.

Corey Quinn: This episode of Screaming in the Cloud is sponsored by GoCD from ThoughtWorks. The worst part of trying any form of CI or CD tooling is getting it up to speed to test it out in the first place. You’ve got to configure it. You’ve got to hook everything together. Somehow now you’re 80 steps in and for some reason step 81 is that you need to tame a wolf.

GoCD knows this better than most of us. That’s why they’ve got a new test drive option to get up and running with their tooling in seconds. Download their binary, run it locally, and set up your first pipeline in your browser. Think of it like a demo, but rather than some arbitrary “Hello Word” app that looks nothing like what you’re running, instead it’s being done with your own application. Give it a try. Stop getting bitten by wolves. Learn more at gocd.org. That’s G-O-C-D dot org. My thanks to GoCD for their support of this ridiculous podcast.

Corey Quinn: Welcome to Screaming In The Cloud. I'm Corey Quinn. I'm joined today by Scott Guthrie, Executive VP of the Cloud and Enterprise Group at Microsoft. Welcome to the show.

Scott Guthrie: Thanks for having me.

Corey Quinn: Always a pleasure. So there are a number of large public cloud providers today, but you obviously are Azure more or less there. I don't know that there's too many people above you in the corporate pecking order as I understand things as far as running public clouds go.

Scott Guthrie: Sure, yeah, no, I am responsible for Microsoft Azure as well as our Dynamics 365 offering. And then also things like Power BI, PowerApps and Flow, which we call our power platform cloud. Then also GitHub and then Visual Studio and Visual Studio Code and then the core Windows operating system. So kind of a broad smattering.

Corey Quinn: Oh, is that all? Oh good, good. And in your spare time... Yeah. This is sort of a high level question then, but what is the Azure philosophy from your point of view, what is it that differentiates that from the competitive offerings out there? I don't believe the easy explanation of, "Oh, we see another company building a public cloud and doing super well with it. We're going to do that too." I think to even suggest that is a bit of a disservice to the very intelligent people who work here.

Scott Guthrie: I think there's several things that I think probably differentiate Azure in some unique ways. I think the first is really our philosophy around hybrid, which is something that we've really had since the very beginning of the product and the service, and really thinking about hybrid kind of very holistically. It's not just about infrastructure because that is, no one just does infrastructure. It's really thinking about tooling. It's thinking about your developer tools as part of that tooling. It's about thinking about your operating systems, your data platform, your AI capabilities. It's about thinking around security and management and it's really about thinking about having a control plane across all that.

I think for large organizations or really any organization that's been around more than five or six years has existing investments that they've got that run their business and making it easy for people to take advantage of those investments as they move to the cloud and do it as seamlessly as possible is probably the single biggest differentiator that we've had for a lot of customers that have moved to Azure.

I think the second thing that we really differentiate ourselves around is thinking about Azure and cloud platform in a much broader context. Meaning it's not just about customs systems and cloud infrastructure. It's really thinking about how do you integrate that with your end users inside your company or the end users that are accessing your systems. That's where things like Office 365 and Dynamics 365 and our power platform capabilities come into play.

Certainly for large organizations as well as small organizations, that ability to more easily integrate the custom apps that you build, which where you're going to want a cloud provider like Azure or AWS, with the end users in a much more holistic way, is one of the key things that everyone struggles with. We make it really turnkey and easy, whether it's around identity, whether it's around data access, whether it's around compliance, all the way through accessibility.

I think the third thing then is, I've talked a lot about enterprises and some of the listeners on your show are probably thinking, "Well I'm a start up, but I'm building apps. Is Azure just the cloud for enterprises or for businesses?" The other third angle that we really differentiate is through what we call our co-sell program. That basically means if you're a startup or you're software vendor that's selling to businesses, we have a unique set of offerings both from a technology perspective with Azure, but then also with what we call co-sell, where we give you the access to the world's largest enterprise sales force, which is the Microsoft sales force. They all get quota relief when your solution that's built on Azure gets sold to their account.

Then we also have the world's largest channel partner network. So we can help put hundreds of thousands of people to bear who basically a compensated when your solution gets sold. That is another unique thing that we have in addition to that end to end cloud, in addition to the hybrid, that together has really powered Azure's success.

Corey Quinn: When you lie awake at night, right before you fall asleep or try to fall asleep, and you think about the entire ecosystem of Azure. What keeps you up at night worrying, assuming you're the worrying type?

Scott Guthrie: Well, you know, I think the thing that's unique about cloud all up is just we're at an inflection point as an industry where cloud has gone from being a small part of the market to really being the future and being the thing that's not just the future but the here and now. So thinking through as we become more mission critical for the world, how do you make sure you deliver that robustly? How do you make sure you deliver that securely? And recognize that we've got hospital operating rooms. We talked about some of them in my keynote today that run using Azure. We've got airlines that can't fly with Azure. You've got banking systems that can't transact.

So how do you deliver that rock solid reliability knowing that the whole world depends on it? That's something that I don't think you're ever done with. All the cloud vendors, I'm sure not just me, but my counterparts in other clouds as well, that's probably the single biggest thing that keeps us up at night. Then how do you do it at the scale that we're all operating at, which is increasingly now measured in the millions or tens of millions of servers, hundreds of data centers and I can't even count how many miles of cabling that we all have.

Corey Quinn: All of them. Approximately all of them.

Scott Guthrie: Yeah.

Corey Quinn: It's strange in that if you take a look at what's been happening the few years, among other things, you've managed to hire some amazing and incredibly talented people who all throughout Azure. I mean you take like Emily Freeman, Chloe Condon, and Tara Walker, Ashley McNamara, and insert 5,000 other names here. How do you do that? I mean these are people who generally do not have credibility for sale, so it's not as simple as just back a truck full of gold bricks up to their driveway and drop it off. There has to be a narrative and a story that's compelling. How are you, I guess, hiring some of the most influential names in tech?

Scott Guthrie: Well, you know, I think one of the things that we've tried to do, I mean there's obviously the tech, what we're building. I think there's a lot of really cool, amazing things that we're building from a tech perspective. I think the other thing that we've really tried to do as a company, and we've been kind of public with it the last five years, is really reboot our culture and really reboot the company in a broad way. Some of that is around tech. People say, "Oh, Microsoft's become a cloud company as part of that reboot." But a big part of it comes down to culture. We like to kind of say, "Culture eats strategy for breakfast." So you can have the best strategy, you can have the best tech, but if your culture is not right, at the end of the day the culture is going to decide your success.

We've tried to have a culture where we say let's put our customers at the center and really change the way that we partner both with our customers, but even collaborate internally, and create what I think a lot of people think of is a pretty exciting place to work. I think ultimately that has been, for a lot of the names that you mentioned, why people have joined us as they kind of see that culture change. They're hearing about it from some of their friends that work at Microsoft. Then more importantly, when you have a customer centric culture, the great thing is you get to celebrate with your customers when they are successful. So in my keynote today, we had BMW on stage, we had Asos, which is a big retailer in the UK, Kroger, Coca-Cola, Virgin Atlantic Airlines, and a whole bunch more. So being able to kind of celebrate and deliver on some of the transformation that all those had, it's super fun and it's kind of addictive. It brings some of the best talent that wants to work on those hard problems.

Corey Quinn: It really did seem that all of the companies mentioned in your keynote were very aligned with solving real world problems here in reality. These aren't companies that are eight or 10 years old that were born in the cloud. These aren't the Twitter for pet startups out of San Francisco. These are household brands that have been around for a century or more in some cases, and it was a stark reminder that corporate IT or production engineering doesn't always look like a startup in the bay. It very often, much more frequently is a hospital in Duluth. It's an insurance company somewhere in Omaha. There's a whole world of serious businesses out there that I think Microsoft's doing a fantastic job addressing.

Scott Guthrie: Yeah, man. I think one of the things that sometimes in tech, and we're guilty of this too, so this is not a statement about others. It's easy for us to kind of assume everyone is like us or everyone knows what we know or everyone does it the way we do it. I do think that's a common thing that I see my own team struggle with at times. Certainly I see other companies struggle with is kind of assuming, "Oh you know the latest edition of this and you're a member of the repo and you've seen all the check ins" How do you scale that out to an organization that's using it that might be a hundred people somewhere else in the world? They might be 100,000 people somewhere else in the world. That's when the stuff gets hard.

Part of what we've tried to do throughout our history, I mean even if you look at where the company was founded 40 years ago, is really try to democratize technology and not require you to be a rocket scientist to be able to deliver great solutions and be successful. We've done that with operating systems. We did that with development tools. Our very first product was a development tool. A lot of people forget that about Microsoft. Even you see that in the roots today with VS code and be VS and GitHub, that's still very core to our DNA and it's really around empowering people to do more.

I think when you do that, then the great thing is more people do stuff and that's been a core strength at Microsoft for 40 years. I think it's definitely, as we think about cloud transformation, it's empowering people to use leading edge tech. So you can also just make it point and click and simple. It's really around how do you use Kubernetes? How do you use the latest AI models? How do you use edge computing? How do you really push the boundaries of the tech but do it in a way that you can accelerate your success because between the tools and the end to end focus, hopefully we make it easier to adopt and ultimately be successful with it, which is what matters. That that earns you fans. That's been the key part of our success.

Corey Quinn: It strikes me just from a marketing and messaging perspective that Microsoft's approach to hybrid has been very different than what I'd say almost anyone else's has been. Again, Microsoft has a history of back when hybrid and the only option and there's a 40 year business case here. So for what you're doing and with the customers you're seeing here is having something on prem is not a new way of thinking. This has been around for a long time. The idea of burn the boats and the data boxes behind you and go all in and move everything to the cloud is simply not tenable for an awful lot of these use cases, nor should it necessarily be that way.

I get the sense you're one of the only providers that seems to be approaching this from a position of that's okay even as a mid to long-term strategy rather than as a temporary pause until they set fire to the data centers, move everything into a cloud and then yay, everyone's going to be happy and a success here.

Scott Guthrie: I mean, I think there's two angles to think about, three angles maybe to think about that. One is just as a... We certainly believe in a big way on cloud. So it's not that we are saying hybrid because we're trying to kind of cloud wash on prem and say, "Oh everything's hybrid now." We do believe that the majority of workloads in the limit will be running in public cloud data centers. Whether that's 60%, 80%, 90%, 100%, I don't know. But it's going to be the majority. I don't think we have any dissonance on that and we've been very clear on that. The part that I think we've been able to really provide something that helps our customers accelerate that is this hybrid approach in recognizing the fact that, let's say even if you do decide to turn off your on prem data centers, you're not going to do it overnight.

Even if you've decided all these workloads are moving into the cloud, you have this challenge, which is, okay, I'm going to migrate over the next year 100% of my workloads. I need to kind of keep all my systems running every single day of that year. I also not only need to make the tech work in the new environment and integrate with the pieces that haven't already moved, but I also need to make sure that I'm leveraging the skillset of my employees because I'm also not going to go overnight from how I used to manage things to becoming a complete DevOps automation expert.

Corey Quinn: Oh your entire career has been what? Replacing hard drives? Cool. You're a JavaScript developer now, good luck. Have fun. Here's a book. It doesn't work that way.

Scott Guthrie: It doesn't work that way. Today roughly 60% of all databases running in the enterprise are running using Microsoft SQL Server. They might not be always the tier one databases, but if you look at certainly by tier two workloads, you know SQL servers, the dominant share in, we have a lot of tier one workloads as well. That ability to say, "Hey, let's move that to the cloud," and rather than bundle it up from bare metal servers to a VM, which gives you some benefits, but at the end of the day it's still managed as kind of raw infra, hey migrate it to be a managed service and we'll do backup, we'll do high availability, we'll let you run in a serverless option. We'll do built in threat detection. That's suddenly gives you a heck of a lot of capabilities.

Then when you can say to someone, "And by the way you can migrate it without changing a single line of code in any of your apps," that's gold. Because it allows someone that would otherwise spend weeks to have to test and re-update every app to suddenly build an app migration factory that lets them migrate many, many apps per day and again, use the same tools and the same languages and the same apps against it when they're done.

So you know we've done that whether it's with our own stuff like SQL server, we're doing it with Postgres, we're doing it with SAP environments, we're doing it with net app systems. We even support and give you the ability to run Cray supercomputers in our data center. We don't do a lot of those but a lot of, your in your oil and gas space, you still have a Cray somewhere that you're using for some of your HPC workloads.

Corey Quinn: There are dozens of us.

Scott Guthrie: You need to find a solution for it before you can turn off that data center. So you're thinking about it holistically has been key. Then most recently we announced our partnership with VMware and/or Dell Technologies, of which VMware is part of. We're the only cloud vendor that offers a first party VMware service meaning you can actually buy VMware directly from us and buy our VMware service directly from us. The beauty there is you can use the same VMware technology you use on prem, you can use that to migrate the infra. But then where we're differentiated say versus AWS, which has-

Corey Quinn: Yeah, because VMware has also been doing a fair bit with other providers too.

Scott Guthrie: The difference is you have a single throat to choke with us. So in the AWS case you've got to buy VMware from VMware, and then you're buying the cloud from AWS. The majority of operating systems running in VMware are Windows Server, and the majority of databases running in VMware today or SQL server. So having a single support model that you can then migrate that stack and know that both at the VMware level, the OS level, the database level, the identity level, all the way up through some of the office 365 work we've done with VMware as well, you can kind of have kind of a single support model and single integration, that's compelling. Again that's all about trying to make it easier for people to leverage the cloud by reusing the skills and the investments they already have.

Corey Quinn: Something I think that people take far too lightly about Microsoft is it's four decades of experience in having conversations with businesses about a wide variety of things and they speak that language fluently and that's something that the other players in this space tend to struggle with by and large. Somehow during the cultural transformation you talked about recently, that has gotten better, if anything. I can't deny it's been there, I mean 15 years ago I despised everything Microsoft and would never dream of even coming to a conference. Today if I were to admit defeat, shut down a company and go take a job somewhere, Microsoft be absolutely top of the list and I don't even know how you pulled that transition off, but it's widely recognized, widely understood and it seems to have only made your enterprise relationships better as you've done that. What happened and how did that look from the inside?

Scott Guthrie: Well, you know, it's a combination of things. I mean first, thank you for saying that. I think part of it is, a big part of it is really just that focus on putting the customer first. It sounds trite, but really if you put the customer at the center and work everything backwards, it's just so much easier to actually then talk to the customer when you kind of have your strategy be that. I do think that was a blind spot that we had in the past. Sometimes we'd say, "Here's our strategy, now let's go talk to a customer," versus the inverse.

So I think that is fairly deeply profound in terms of the transformation. Then I think the other thing that we've gotten better at is in what we build. Again, taking that customer focused approach. Let's be pragmatic even around how do we partner, what do we do to differentiate, versus how do we stop doing things that are just different? So if you look at our embrace of Linux, or if you look at our brace of Kubernetes or you look at our brace of chromium or electron with VS code, and chromium most recently with our own browser, those are places where I think the Microsoft of old would sort of say, "Oh, we're going to fight this sort of battle on licenses…”

Corey Quinn: And charge a license for every one and here's the terms. It's not being played like that anymore.

Scott Guthrie: Instead we're trying to really say, "Okay, how can we add unique value to our customers? Let's focus on that. Then let's partner and embrace the technologies are already mature and/or customers are already using. Instead let's focus on how can we be uniquely deliver value." That has been fantastic, and I think that's allowed us to really take all of our internal engineers and really have them work on stuff that our customers love. That's helped us from a product perspective.

Then I think the last thing from a sales perspective and engagement perspective, we have also tried to change not just our engineering model but also again, how we engage with customers. Really try to be a partner as opposed to a licensing specialist. Cloud is all about consumption. We, a couple of years ago, changed so that our salespeople don't actually get any credit for selling. I'll call it a monetary commit. Or someone saying, "Oh, I'm going to use a lot of cloud." They only get credit when the customer actually deploys and is actually successful using it and really changing the mindset to be all about consumption. That completely aligns even how we sell things to our customer's ultimate realization of value. So how do we as a sales organization, as our partner network, how do we focus on customer success? That also obviously changes the conversation as much as the tech does as well.

Corey Quinn: Even earlier today during the keynotes, sorry. Right now, you just mentioned a few minutes ago that partners are one of the keys to your success and that's very clear. During the keynotes and before the keynotes, I looked around the expo hall and I saw no partners wringing their hands nervously looking around wondering if they're about to be put out of business by something you're going to announce. It seems very much that you're invested in helping others succeed as you do as well. It's easy to say, oh, that's just temporary. But it's been how many decades so far and it's still is very core and central to what you do.

Scott Guthrie: Yeah, I mean the majority of Microsoft revenue comes through partners today. That is been true really since day one of Microsoft, and it's still true today. So I do think we have a large partner network, like hundreds of thousands of partners around the world ranging from partners that help customers implement and do kind of SI work to ISBs that build applications on top of us, to a channel partners that help reach smaller businesses. I think we have a deep appreciation for the power of partners. I think the thing we've learned over the years often by doing it right and occasionally by doing it wrong, is just how important it is to really curate those relationships with partners and really be good stewards of those partnerships. Because at the end of the day, you kind of build trust over a lifetime and you can lose it in an instant.

I certainly hope that there was not a single partner out there that had any surprise over anything we said in the keynote because I kind of view that as a fail if they did. A fail on Microsoft's part, if we did. That trust is super, super important. I think that's true, not just for longterm partners. You know, a lot of the partners I talked about in the keynote have only been partners with us for two or three years because it's really been part of this cloud transformation. Or maybe they were historically we're a Linux based solution and so they never even thought they could deal with Microsoft and take a Databricks or take UiPath. That was one of the startups that was in my keynote.

These are companies that maybe three or four years ago would never even thought of working with Microsoft, and yet now have a fantastic relationship and are really driving, we're helping together drive their businesses. I think it's definitely something we invest in. Then beyond the commercial side, I think it's also in the open source side, which I think it's probably different than other vendor's mentality, which is not just how do we support open source? Meaning we use it. But also how do we give back? If you look at VS code, which we've open sourced, if you look at .net, which we open source, those are examples of technologies that we started and then gave to the community.

But then, I talked a lot about some of the work we're doing with KEDA in Kubernetes or virtual kubelets with Kubernetes, or Helm, which is also built by my team and the Kubernetes ecosystem or Draft or others. There's lots of technologies where we weren't necessarily the incubators of the broader Kubernetes, but if you look at total contributions to Kubernetes or Postgres, most recently with the Citus acquisition that we did, which was big in the Postgres world. We announced a bunch of great hyperscale Postgres stuff today and we're giving it to the community. I do think it's super important to have that bi-directionality and really be not just a consumer of open source, but a real valued partner and contributor to open source. I think that that mindset of mentality, you'll see us continue to push forward on across everything we do.

Corey Quinn: To look at it from a slightly different perspective today, where do you think that Azure is currently being the most misunderstood?

Scott Guthrie: It's a good question, man. I think one of the things that from a startup perspective, especially for ones that are on AWS today, and obviously AWS has lots of startups, I think understanding how can Azure help accelerate your business is something that a lot of people don't necessarily fully understand. Both on the technology side, and I think increasingly with some of the AI tooling that we have, especially around the cognition services, around speech and computer vision, I think we've got stuff that is very, very differentiated versus other cloud providers and delivering some pretty amazing results.

I think the edge computing piece is an area where we're very differentiated. Then on a kind of non technical side, I think that the Azure Cosell program, and I mentioned UI path a little bit earlier on the keynote, they did a great video in my keynote talking about their sales success. I think they've closed over 200 deals with the Microsoft Salesforce.

The promise that we provide, which is if you're a startup and you build on Azure, every single Microsoft sellers going to get quota credit when their account buys your solutions. So if you're the Coke sales rep and you're a startup that's built on Azure, that Coke sells rep is going to get credit when you sell to Coke. Having that partner ecosystem and having that sales presence that can help on the sales side, I think is another key piece that I think when we walk a lot of startups through they kind of go "Whew, that's pretty differentiated." I think that's something that, you have to have great tech as well obviously, but I do think that combination of having some really differentiated tech but a super differentiated go to market, is something I'd love to have more people understand.

Corey Quinn: Perfect. Thank you so much for taking the time out of your day to speak with me today, Scott.

Scott Guthrie: No, my pleasure. Thanks so much for having me, and thanks for coming to the event.

Corey Quinn: Absolutely. Scott Guthrie, executive VP of cloud and enterprise group at Microsoft. I'm Corey Quinn. This is Screaming In The Cloud.

VO: This has been this week's episode of Screaming In The Cloud. You can also find more Corey at screaminginthecloud.com or wherever fine snark is sold.

VO: This has been HumblePod production. Stay humble.

View Details

About Omer Levi Hevroni

Omer has been coding since 4th grade when his dad taught him BASIC, and he got hooked. From that point, he learned to code in many programming languages (today his favorite is C#). Today he’s working at Soluto by Asurion, and coding is a huge part of his day job.

His passion for AppSec started by accident when he was offered the role of security champion. The AppSec journey was (and still is) fascinated, and taught him a lot. OWASP helped him a lot during this journey; This is why he decided to become a paying member and also leading OWASP Glue.

Omer’s current job is DevSecOps – helping the entire team to produce more secure software. Besides his job, he’s also giving a lot of talks all over the world, and heavy OSS contributor – mainly to Kamus, a secret encryption solution for Kubernetes platform.

When he’s not working – he’s enjoying the company of his two beloved kids and his wife.

Links

  • https://twitter.com/omerlh
  • https://www.asurion.com/about/smb/who-we-are/
  • https://github.com/Soluto/kamus
  • https://omerlh.info

Transcript

Corey Quinn: Hello and welcome to Screaming in the Cloud, with your host, Cloud Economist, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud. Thoughtful commentary on the state of the technical world and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

This episode of Screaming in the Cloud is sponsored by O'Reilly Velocity 2019 Conference. To get ahead today, your organization needs to be cloud native. The 2019 Velocity Program in San Jose from June 10th to 13th, is going to cover a lot of topics we've already covered on previous episodes of this show. Ranging from Kubernetes and site reliability engineering over to observability and performance. The idea here is to help you stay on top of the rapidly changing landscape of this zany world called, Cloud.

It's a great place to learn new skills, approaches, and of course technologies. But what's also great about almost any conference is going to be the hallway track. Catch up with people who are solving interesting problems, trade stories, learn from them, and ideally learn a little bit more than you going into it. There are going to be some great guests, including at least a few people who've been previously on this podcast including Liz Fong-Jones and several more. Listeners to this podcast can get 20% off of most passes with the code CLOUD20. That's C-L-O-U-D-2-0 during registration. To sign up, go to velcityconf.com/cloud. That's velocityconf.com/cloud. Thank you to Velocity for sponsoring this podcast.

Corey Quinn: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Omer Levi Hevroni, who's a DevSecOps Engineer at Soluto by Asurion. Welcome to the show, Omer.

Omer Hevroni: Hello Corey, and thank you for having me. It's a great honor to be on this podcast. I'm a huge fan.

Corey Quinn: It's super kind of you to say that. So, in addition to the day job, or I guess in component with the day job. You wound up releasing an open source tool called Kamus, which I'm almost certainly mispronouncing. There's some that are going to say it's called Kamus, some will say Kamus. It's K-A-M-U-S, which out of the gate is already a better name then almost anything Amazon has ever given to anything. What does it do?

Omer Hevroni: Let's start with the name because I guess, I don't know how much of the people hearing us right now are speaking fluent Hebrew. Kamus in Hebrew means a secret. And we called it Kamus because it's a tool that helps managing secrets on Kubernetes. We all love secrets and we all want to keep them confident. And Kamus come to solve exactly this problem. So this is like the very high level of overview. And it's different from all the other tools because Kamus let you encrypt a secret to a specific application. So only this application can decrypt it. And this is a very unique approach to secret management.

Corey Quinn: That makes a lot of sense. So take me back to the beginning here. There are a number of different ways to address secrets management of varying degrees of validity. Some make an awful lot of sense, some are less sensible. It ranges from using GPG which is heavy and difficult to wind up encrypting secrets that then applications wind up retrieving. There's just embedding it in your source code and hoping that no one ever finds it. There's putting it in secure S3 buckets and just acting shocked when that winds up getting discovered. And fifty thousand other horrible things you can do.

Corey Quinn: But then there's good stuff. There's things like AWS systems manager parameter store. There's systems manager in the AWS side. I have to assume Azure and GCP and the rest have something vaguely similar. What does this do differently?

Omer Hevroni: Yes, you pretty much covered are the different options and I think when thinking about what's the end solution to choose, we need to think about the quality of the company you're working at. For example, the way we work at Soluto is a way we like to call “super devs,” which means developers who do everything. They do the fine work of writing the code but they do the hard work of deploying it and maintaining it. And, when you have developers who do everything, you need to choose a solution that on one hand they can use easily, and on the other hand they can use easily but still be secure.

So, for example if you go with GCP like you mentioned, you need to find a “best” solution that would let all the developers have access to encryption keys and keep it secure and all that, and it gets complex very fast. This is why I'm saying it's the culture or the way you work, the way you work is a very important factor when thinking about which solution to choose. So pretty much you can think about different approaches and approaches are like the most of the GitOps which is I think one of the hype buzz words existing today and maybe GitOps and Bitcoin and blockchain.

GitOps is the idea of managing everything via Cloud. So, describing your cloud architecture and everything else. When you do things with GitOps, you get lot of benefits of Git, like auditing and roll-back and all that good things. And this is why I like also to manage secrets [inaudible 00:05:32]. So, solutions like AWS KMS or Azure Key Vault which is pretty much the equivalent of Azure to KMS. A really good choice, but they don't have any GitOps support. So if you go on this path you need to solve also authentication, authorization, and a lot of other things that you don't want to mess with. So I think this is pretty much the reason not to go this way.

Corey Quinn: There are a number of different things that I will look at whenever I see something that purports to handle security and discount them almost out of hand. The easy low-hanging fruit one those is rolling your own crypto. One of the nice things I like about Kamus is that it uses whatever KMS version is being used by underlying provider — to wit: that's Azure, Keybolt, KMS, there's Google Cloud KMS and AWS KMS for the big three. You’re not generally rolling your own crypto solution here with the singular exception of running your own keys for just local development which you very clearly call out in the documentation. But that is not something to be done unless you are in a hurry for testing purposes. So, first can I just congratulate you on not doing the most ridiculous thing imaginable?

Omer Hevroni: Yes, at first we tried, actually this is a pretty interesting aspect. The whole idea of Kamus was to say, "Hey, I have no idea on how to do encryption good." I do security and I do OpSec for two and a half years, but still I'm saying out loud I don't have any idea about [inaudible 00:07:05] and I don't want to do this things alone. So this is why we prefer or to have [inaudible 00:07:10] log to the cloud. And the interesting part is that in this area GCP was the best experience. And GCP will have no control in the [inaudible 00:07:29].We just told the KMS we kept this and we kept this and you have zero control about what happens behind the scenes.

In my opinion it’s the best experience, because there are less chances you’ll make mistakes. For example in AWS you need to handle the encryption yourself . So the master key is encrypt in AWS, but you are doing the data encryption in your code. And I guess you might have a mistake there with the try really hard to do good but you know mistakes happen. This is very interesting aspect of looking at all the tech clouds and looking which one give the easiest developer experience.

Corey Quinn: One thing that I found fascinating about Kamus, compared to an awful lot of other projects that have come out of the open source world is they start off being built for AWS and that's it. Then eventually they add a second cloud provider, begrudgingly, and then it sort of always feels like a secondary thing. This came out of Azure first and then, I think, correct me if I'm wrong, then it was GCP second and AWS third?

Omer Hevroni: Yes, we started with Azure, because it’s what we work with, with Azure it is the cloud we use the most. We started with the cloud we were using because we first wanted to make sure that Kamus is usable and its working as we expected. And then when we were going to open source it was clear that we can't support on the Azure, if we want people to use it, because unfortunately not everyone love Azure as we did. So, yes then we added GCP because someone asked for it, but then never use it, which was pretty disappointing.

Corey Quinn: That sort of at times feels a bit like the GCP story, but that feels a little mean past a certain point.

Omer Hevroni: Yeah it was really disappointing, because I did it really fast and hoped that he would use it and then he vanished. So if you hearing us and still waiting for your review of the GCP implementation.

And then two months ago we had a AWS, which as I said was the hardest, I think, to implement. And we now support all the three major cloud providers and we look forward to add support for other KMS, so Hashicorp, Vault, ,or anything elseyou want support of, it’s very easy. Choose your service and you’re done.

Corey Quinn: Something else that I found fascinating about this is that unlike virtually everything modern that is written in Go or everything designed that is written in Rust, but never actually built, because no one wants to stop talking about Rust long enough to write anything in it, you wound up writing Kamus in .NET. Why was that?

Omer Hevroni: Yes, well historically I started developing in .NET and I stick with it. And it is still the language that is the most easy for me to write in. But .NET is actually pretty awesome language and lot of people hate it because its belong to Microsoft. But since .NET went open source it's become really easy to use and really handy tool. It’s still a bit heavy compared to other tools. A simple project has a lot of files and all that. So compared to one Golem file its a lot more heavy but you have a lot of things that are harder to do in other languages and this why I still like it. It was a personal choice because choosing a language that not a lot of people will use mean that you will have less contribution.

I already have some people saying I will maybe consider contributing to it, but I'm not sure how to write things in .NET. So there are some downsides to choosing this language, but this was what's easy for us at that time. And maybe in the future we will consider writing in Golang or something else so we can have more people be able to contribute to Kamus.

Corey Quinn: Realistically I find that picking a good language makes sense. When there is a compelling reason to do so and you just named one of them, where having an ecosystem of people in the space who are able to contribute to a project. That's a viable strategy, but I have zero tolerance these days, or frankly most of my career, for language bigotry. Where, oh, this is an amazing tool that does the thing I want it to do but it is written in Pearl so nevermind it must be crappy. Or, oh, it is written on the wrong version of Python or, oh, Ruby is or children. We're not going to go and use that at all.

And none of those statements are ever true. And none of those statements are ever helpful. It feels almost like it's a form of gate keeping, where the entire purpose is to make people feel bad about what language they have chosen to write things in. In fact, if we ought to go by the sheer number of expensive things that will break, “written in Java” is pretty much the clear winner. Between that and still the things that run our entire world are written in COBOL, I don't think that there is a compelling argument to be made that only the cool modern languages are the ones that matter.

I'm a big fan seeing things that aren't written in the usual suspects. And this is absolutely one of them. I've taken a bit of a look though the code base it seems clean, it seems well written, I understand what its doing despite not being conversed in a .NET myself, which tells me that either the language is not that hard to learn or you've gone significantly out of your way to make this approachable to people who aren't just you.

Omer Hevroni: I think this is a fair point to touch. For me its still very hard to write Golang code, and I am working with Kubernetes for more than a year, so I had to write a lot of Golang code. I think it's the language, but I cannot say that for sure, because I don't have enough experience with Golang, but again its just only my experience. Golang still very hard for me to understand and write and read. And this is places where I think language don't matter. If you have a language that you know well and you know all the tiny points of what to do and what not to do, its very hard not to use that language and use something else, because you will probably produce a better product. So, there are places where languages matter and its not what you control and not what is popular.

Corey Quinn: The problem that I've always had with popularity driven development is that it leads almost to group think. We're starting to see this, not just with languages, but with other things as well where if you're not using AWS you're clearly doing it wrong. I don't know that is necessarily true. Or if you're building any monolith that no microservices are the way and the light of the future. Or, oh, if you're, not going serverless you're completely missing the entire point of the modern revolution. And I don't like any of those things. They smack far too heavily or orthodoxy and its getting away from solving business problems and into the realm of religion. And that doesn't seem to be serving anyone particularly well in a technical or a business context.

Omer Hevroni: Well, I'm a religious, Orthodox Jewish so I'm pretty fine with being orthodox and going with religion. [Laughter]

Corey Quinn: There’s validity to that! The question is at some point it blurs the line between things we have to take on faith and I think that religion is inherently going to be one of those, versus things that lead to definable business outcomes. And doing something just because the thought leader on the stage said it was the way to do things, that's always been a problem for me. Because our choices are built by constraints in a technical context and the fact that we don't wind up making decisions based upon other people's constraints or even our own constraints leads to disaster.

If someone gets on stage like what they did back, in I wanna say 2014-ish, and talks about how Docker is amazing and terrific and solves everyone's problems, well it didn't. What about state full workloads? What about networking? What about orchestration of these things? What about monitoring? How do you admit statistics? How do you troubleshoot when the thing that's running was turned off twenty minutes before you knew that there was an error? And there was a giant list of problems that has largely been solved today, but at that point the whatever your problem is Docker containers are going to solve it, was what everyone was saying and I didn't' find that resonated with me, because just throw everything away that I built previously and build everything again and this new context, this new environment.

I could see doing that, but I didn't see a compelling business case for it. And now here we are six years later almost and we're talking extensively about Kubernetes in some cases being the answer to everything. I have a problem that I may possibly have. I don't see it. There are absolutely use cases for which Kubernetes is a valuable path forward and a great use case. But certainly not all of them. And the level of complexity still feels slightly higher than most companies are probably going to want to dive into, unless they are themselves tech companies at heart.

Is that something that you're seeing? Do you agree? Disagree? Don't feel you have to agree with me because I'm the one with my hand on the mute switch.

Omer Hevroni: No, I'm pretty much agree. I think its level of maturity as a software developer or a security developer or as any other dynamic, but the other parties not always answering for what is good but for what is not good — it hit me first when I heard a really good interview question: when you need to answer when microservices are not a good choice. And there are pretty good answers to this questions, but I hear this question and it got me thinking. It wasn’t my job interview but I heard a little of, and I started thinking it was pretty interesting thing to think about, like what situations are not good use cases because microservices but as you say it was very interesting because it was going against the dogma. And starting to say, hey, microservices are not always that good and has a lot of overhead. And we take it as a fair point. I actually like the example of Kind, it is an open source that lets you run Kubernetes using Docker and its very good in CI.

And what I liked about Kind documentation website is that they have a whole list of alternatives that are similar to Kind, but do other things. And you can look on all of them and chose between the alternatives and find the one that best works for your use case. And I wanted to do this on the Kamus documentation website and I forgot, I need to do it. But I think it is very important to say, like, what you are not good for and what use Kamus, or Kind, or Kubernetes, or whatever.

Corey Quinn: I absolutely love the idea of an interview question. Where there's this thing that you're good at, that you're passionate about, great. Tell me about it's shortcomings. Tell me about when it's not a good idea. Because that tells you a lot about a candidate. Just as far as can they see both sides of an issue, are they so blinded by zealotry for this thing that they are in love with that they can't see another perspective? And if you wind up making a technical decision down the road, well, they're working with you, are you suddenly going to wind up having a screaming meltdown? Because that is not acceptable. Finding ways to tolerate and deal with disagreements is incredibly important for almost every environment I've ever been in. So I just wanna say I love the idea of asking that kind of question.

Omer Hevroni: Yeah we can do it on Kamus. Kamus ever since it start it's not good for everyone. If you already have Vault and Vault works for you and you have a culture where you have a specific team who will manage the secrets, it might be easier to continue using SSM or GitOps might work better for your solution. Another caveat of Kamus is that Kamus do one-way encryption. Once you encrypt the secret you have there is no admin account lets you decrypt it back, and its not work for everyone. For us it was good enough, but it had downside. Sometimes you this file decrypted and Kamus, by design don’t do that. So it's important to understand the idea behind the design of Kamus and understand this idea fit your workflow or not. Because it's not always fit.

Corey Quinn: Pay attention listener. The reason to look into Kamus is because you heard on this podcast, yeah that's a good idea. Deploying it blindly, because the super confident sounding people on the podcast said you should, that's a bad reason. It all comes down to making sure to whatever you are using works for your use case. Understand what it is. Understand how it breaks. Understand how when it breaks, because everything breaks that's what computers do, will impact your production environment. It all comes down to understanding constraints, understanding trade-offs.

Omer Hevroni: Yeah this is part of the reason we decided to release with Kamus. A full-blown [inaudible 00:20:14] of all the different threats and controls we think about. So you as the user can look at that and say hey this is an unacceptable threat by us. And I don’t going to deploy Kamus and have this treat available, I’m not going to go with it. And you can look at the controls and understand the different things we decided to do in order to mitigate them and you don’t like some of the controls or you might even think about threats or controls we didn't think about. So it’s all part of the idea of giving our users all the information, it’s all there for them, and with all this being said, go use Kamus it really is amazing. Don’t listen to all the things I said before!

Corey Quinn: Exactly! Adoption where it all comes from. Sooner or later we'll put a foundation around it and get people to start supporting it giant corporate piles of contributions. Changing gears a little bit… You introduced yourself in the beginning as a DevSecOps engineer. And I tend to come from an old school world. Where I started my career as a grumpy Unix Systems administrator. We also called that a UNIX Systems administrator, because they are all grumpy. There is no second kind. And over time I watched the world evolve. We started calling the same thing I was already doing DevOps. In many shops they started calling it SRE, or cyber liability engineering. And it feels to me, and its probably slightly naive, that the role hasn't particularly changed in fifteen years. Sure, the tools we use are different. But if you say, well, back you were a SysAdmin you didn't have to write code. I didn't. I was doing an awful lot of stuff in Perl and if you start working on UNIX systems administration you generally in that era and won't get very far without a working knowledge of C.

Corey Quinn: And it turns into, it seems everything old is new again. So then we called it DevOps, which people argued vehemently is a title, is not a title. Well at least we’re debating rather than doing work. We don't want to do work. We want to argue with people. And now we are seeing DevSecOps as adding words to that construct, which on the one hand I like. The idea that security has to be baked in from the beginning is a fundamental tenet of safe systems design. And I'm curious, though, how did you wind up coming by that job title? When did you start using it?

Omer Hevroni: It’s complicated like all the good things. The funny part is that today I'm part of the DevOps team at Soluto. We talk a lot with our manager about changing it to SRE, and ask him if I change my title from DevOps to SRE what will that increase my salary. And he refused to answer this question. So today I'm not SRE, but if it will increase my salary I will.

Corey Quinn: Absolutely — to an extent there have been a number of DevOps days where they’ve done job title polls and going from SysAdmin to DevOps or SRE was something on the order of a salary boost of at least a third taken in the aggregate. So it's weird, but the things we call ourselves change the way people view us and more importantly how people pay us.

Omer Hevroni: So I think I will go with DevSecOp SRE. I think it will be good enough.

Corey Quinn: I like that. FinOps in there too somewhere. I think someone is trying to make that happen somewhere.

Omer Hevroni: Yeah and the less funny part is that I don't have any official degree. So I am a bit ashamed to call myself engineer. But you know it’s a different discussion about the important of doing a bachelors degree in computer science and all that.

Corey Quinn: Not necessarily, I have an eighth grade education myself. I have always toyed with the idea of getting a computer science degree at some point in my career. And there aspects that will be helpful for it, but from my perspective it seems for some folks it’s more of a credential. In my twenties it was rough not having a degree. I had to talk my way around a lot. Now that I'm in my thirties I can bring this up in public, that, hee-hee, I don't even technically have a high school diploma. And it’s a punchline it’s not a career limiting move. By and large careers do get easier. With the right credentials. But I don't think that its ever absolute. At least not as we tend to think of it as. It just means we have to be a little bit more creative.

Omer Hevroni: Yeah and my parents are somewhat disappointed and hope at some point I will get the sense of it and go and do an official degree. But it just too hard. I never managed to get over the hard math courses, so I don't think it will happen any point.

Corey Quinn: I'm in the same boat. My parents still think I'm a an accountant. Which, yeah, we're gonna go with that. It’s a lot easier to explain.

Omer Hevroni: Anyway the job title is complex. I started at Soluto as a developer. And then someone offered me to be part of the security champion — you heard about the idea of security champion?

Corey Quinn: I have not. Tell me more.

Omer Hevroni: OK. So it’s not that new but it's coming from the idea that the security team, especially in companies like insurance, usually very, very isolated, located somewhere in the United States. And they cannot control all the developers all around the globe. So usually the solution to that is either hire a lot more security people, who will be located in each site, which is very hard, or taking developers and allowing them, converting them into security. Taking them to the live site. So this what they did. I did a SANS course of one week, so I have a certificate. I’m officially a pen tester and I can use this title if I want. And since then I started to do all things related to security. OpSec, and all the fun stuff DevSecOps which going from security testing to running things in the pipeline and the things I like the most, my job in the DevOps teams, I think fundamentally different from SysAdmin is what I started with.

We have a lot of developers and our job as the DevOps team, especially my job, is the security person is to help them and guide them. We don't have any more control than what they do, we cannot tell them what is wrong and what is right. We can just help them and provide them with best practices. And if they come and say, “hey, want to use this cool new technology,” we are the bad guys who come and say what about monitoring? How can you monitor that? — how can we deploy that? Is it secure, and all that question. And these things I think are fundamentally different from a SysAdmins. Because SysAdmins usually were not responsible to do this work and this why I like the title of DevSecOps engineer. I’m not that happy with this title, but I think it’s better than security champion. Which was the title I used before.

Corey Quinn: I think that may wind up being for the best. All things considered. We're about out of time, but if people are interested in what you have to say where can they find you to learn more?

Omer Hevroni: So, the best place my Twitter account. I try not to tweet too much in Hebrew, but I do it also. Its my language and I try not to put so you can follow me there.

Corey Quinn: That’s O-M-E-R-L-H — we'll put the link in the show notes.

Omer Hevroni: Yeah and I also have a personal blog, which I try to post new things once in a while. And you can also find me there or find me on Kubernetes Slack, or CNCF Slack, I'm on each one of them. And of course GitHub, and all the other things. You know some people have, and of course LinkedIn. I think some people are also there.

Corey Quinn: Sound good. Thank you so much for taking the time to speak with me today.

Omer Hevroni: Thank you for having me. It was a real honor, and as I said in the start I'm a huge fan of you. It was pretty nervous to be on this podcast after the last episode with Jess.

Corey Quinn: Yeah but you’re talking about, there's always a bigger fish. I gotta say I was nervous as heck talking with Jess. She's incredible. And everyone I talked to is incredible. But she's one of those people I've been following for a long time and it turns out everyone has someone that they look up to, oh I could never talk to them. I’m never gonna be in their tier — there's always a bigger fish. And something I’m learning is that reaching out and talking to people and asking for help, everyone worth admiring is a human being who is willing to have a conversation. Don't ever hold yourself back from that. Omher Levi Hevroni, DevSecOps engineer at Soluto by Asurion. I'm Corey Quinn, this is Screaming in the Cloud.

Corey Quinn: This has been this week's episode of Screaming in the Cloud. You can also find more Corey at Screaminginthecloud.com or wherever Find Snark is sold.

Speaker:

This has been a Humble Pod Production. Stay humble.

View Details

About Valentino Volonghi

Valentino currently designs and implements AdRoll's globally distributed architecture. He is the President and Founder of the Italian Python Association that runs PyCon Italy. Since 2000, Valentino has specialized in distributed systems and actively worked with several Open Source projects. In his free time, he shows off his biking skills on his Cervelo S2 on 50+ mile rides around the Bay.

Links Referenced:

  • https://twitter.com/dialtone_
  • Adroll.com
  • Tech.adroll.com

Transcript

Host: Hello and welcome to Screaming in the Cloud with your host cloud economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode of Screaming in the Cloud is sponsored by O'Reilly's Velocity 2019 conference. To get ahead today, your organization needs to be cloud-native. The 2019 Velocity program in San Jose from June 10th to 13th is going to cover a lot of topics we've already covered on previous episodes of this show, ranging from Kubernetes and site reliability engineering over to observability and performance.

The idea here is to help you stay on top of the rapidly changing landscape of this zany world called cloud. It's a great place to learn new skills, approaches and of course technologies. But what's also great about almost any conference is going to be the hallway track. Catch up with people who are solving interesting problems, trade stories, learn from them and ideally learn a little bit more than you knew going into it.

There are going to be some great guests, including at least a few people who've been previously on this podcast, including Liz Fong Jones and several more. Listeners to this podcast can get 20% off of most passes with the code cloud20. That's C-L-O-U-D-2-0 during registration. To sign up go to velocityconf.com/cloud. That's velocityconf.com/cloud.

Thank you to velocity for sponsoring this podcast.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Valentino Volonghi, CTO of AdRoll. Welcome to the show.

Valentino: Hey Corey, thanks for having me on the show.

Corey: No, thanks for being had. Let's start at the very beginning. Who are you and what do you do?

Valentino: Well, I'm CTO AdRoll group. And what AdRoll group does is effectively build marketing tools for businesses that want to grow, and they're looking to try to make sense of everything that is happening in marketing, especially when it comes to digital marketing that effectively is going to help their businesses drive more customers to their websites and turned them into profitable customers effectively.

Corey: Awesome. You've also been a community hero for AWS for the last five years or so?

Valentino: Yeah. I was a lucky enough to be included in the first group of community heroes, which I think was started in 2014. It isn't still completely clear to me what exactly community heroes do besides obviously helping the company and what did we do to deserve to be called community heroes. I think lots of people such as yourself are doing a great amount of work to help the community understand the cloud and spreading the reasoning behind everything that is happening in the market these days. So maybe you should be a hero as well.

Corey: Unfortunately, my harsh line on no capes winds up being a bit of a nonstarter for that. And I'm told the wardrobe is very explicit.

Valentino: Oh, okay. I didn't know that.

Corey: Exactly. It all comes down to sartorial choices and whatnot. So you've been involved with using AWS from a customer perspective for, I'm betting longer than five years.

Valentino: Yeah, about probably longer than a decade, actually longer than a decade.

Corey: And it's amazing watching how that service has just ... I guess how all AWS services have evolved over that time span. Where it's gone from, yeah, it runs some VMs and some storage, and if you want to charitably call it a network you can because latency was all over the map. And it's just amazing watching how that's evolved over a period of time where not only was it iterating rapidly and improving itself, but it seemed like the entire rest of the industry was more or less ignoring it completely as some sort of flash in the pan. I've never understood why they got the headstart that they did.

Valentino: Oh man. Such a long, long time ago. I remember I was still in Europe before I came to work on to start up AdRoll, but in 2006 I think was when S3 was first released. And I remember starting to take a look at it and thinking, "Wow, now you can put files on a system out there that you don't know really where it lives, but I don't need to have my own machines anymore." And it was the time that you used to buy colocations online and it was a provisioning process for all of those, you needed to choose your memory size and you typically get a co-located, co-hosted, shared host type situation. And it was expensive. And then yeah, in 2008 EC2 came out, 2007, EC2 came out and it felt like magic. And at that point in time, either was running in a data center out here on Spear Street in San Francisco. And I remember we had a two databases machines, both RAID 5 and one machine was humming along fine, but the other one was going on two drives. Two drives that were failing in the RAID 5. And we started driving the ordering the drives on Amazon or whatever, NewEgg. And I think that we're in back order at that time and we needed to wait for a week or two before those could arrive. At that moment in time I made the call. That's it, we're not doing this anymore. We are going on AWS. Just give me two weeks and I migrate everything, I told the CEO. And then we'll be free from the data center. And I tell you it costs will be exactly the same. And actually that's exactly what happened.

It took two weeks, I moved all the machines over, the costs were exactly the same, but we had no more needs to run to the store and provision extra capacity, or buy extra capacity, or any of that stuff. It also allowed us massive amounts of flexibility. And then very early on it was funny because I think I believe that through all of the stages of disbelief when it comes to AWS or cloud in general, where the first the complaints were, "Well it's not performant enough." If you want to run MapReduce you cannot run it inside AWS. There's simply not enough IO performance on the boxes. I even lived in a period of time, I was following closely when when Github was on AWS at first and then they moved to Rackspace afterwards because AWS wasn't fast enough, even for them. And they were working through some issues here and there. Some of those things were obviously real, like truly maturity situations. EBS has gone through a lot of ups and downs, but it's mostly been stable since then.

Living in now a day and age where the EBS drives that you get from AWS are super stable but it never used to be like that. And you needed to kind of get adapted and get used to the fact that an EBS drive could fail or the entire region could go down and because of EBS drives, which has happened in US East a few times in the past. But yeah, from those very few simple services with very rudimental and simple APIs, it does feel like they have ... They're starting to add more and more, not only breadth, because obviously that's evident to anybody at this point in time. I don't think anybody can keep up with the number of services that are being released.

But what's really surprising is that for the services where they see value and where customers are seeing a lot of adoption and interest, they can go to extreme depth with the functionality that they implement, the care with which they implement it and ultimately with how much of it is available for many of them. Now you get over 160, I think, different types of instances. It used to be that you only had six or seven, and now 160. Some of them are FPGA instances, which I think there's only maybe a handful of people in the world that can code those machines properly. And then certainly don't work at my company right now.

Corey: Well that's always the fun question too is, do you think that going through those early days where you were building out an entire ecosystem on ... Or sorry, an entire infrastructure on relatively unreliable instances and disks and whatnot, was I guess, a lesson that to some extent gets lost today. I mean it taught you early on, at least for me, that any given thing can fail. So architecting accordingly was important. Now you wind up with ultra reliable things that never seem to fail until one day they do and everything explodes. Do you think it's leading to less robust infrastructure in the modern era?

Valentino: It's possible. I think if people get on AWS thinking that we're going to run in the cloud so it's never going to fail because Amazon manages it, I think they're definitely making a real mistake, a very shortsighted statement that right there.

Not just because of that in case of failures, but a couple of years ago, I think maybe three years ago, there were all of those Zen vulnerabilities coming out and Amazon needed to patch and entire regions needed to be rebooted. And what do you do at that point when your infrastructure is not fully automated and capable to be restored without downtime in user facing software? You're going to need to pause development for weeks just in order to patch a higher urgency vulnerability in your core infrastructure. That's just an event that is not even a fault of anybody.

It's not even in the necessarily under full control of Amazon, and you need to be ready for some of that stuff. So there are. I would say that that are, there are systems that are simply ... Lots of companies that especially in their first journey to moving stuff inside AWS, they tend to just replicate exactly what they have in their own data center and just move it inside AWS. I know this because for example, AdRoll has done the first time that we migrated into AWS. We first migrated just our boxes, and then we quickly learned that that it wasn't always that reliable and so we needed to figure some of that stuff out for ourselves. And effectively you start to realize in our case back then that you needed to work around many of those things. But as you said today, it isn't quite that way, and to an extent, Amazon almost makes a promise about many of these services not failing or taking care of your infrastructure for you.

For example, if you look at Aurora, is a stupendous, fantastic piece of database software. It's extremely fast. It's always replicated in multiple availability zones and multiple data centers. The fail over time is less than a second, I think, at the moment. And when you're tasked with solving a problem, building a service you're going to choose to build it on top of Aurora neglecting to think about what happens if Aurora doesn't answer to me because the network goes off? Or what happens if my machines go down because AWS configured them?

Some of the biggest dire profile issues in terms of infrastructure of the last year alone, for example with the S3 of being erroneous configuration changes being pushed to production. What do you do at that point? Your system needs to be built in such a way that is going to be resistant, at least partially to some of these things. And Amazon is trying to build a lot of the tools around that stuff. But I think it still takes ... It still takes a lot of mind presence from the developers and architect to actually do this in a thoughtful way, use the services that you need to use in a thoughtful way, understand the perimeter of your infrastructure, and particular the assumptions you're making as you're building the infrastructure.

And if you can design a graceful degradation service where failure of an entire subsystem is not gonna lead to complete failure to serve a website, whether you slowly get to just the less useful website progressively but still maintaining the core service that you might offer, then it improves your infrastructure quite, quite a lot. I think this is where Chaos Monkey, Gorilla, whatever, King Kong, Kong or whatever it's called the for the region, failure come into play to try to exercise those muscles. It's obviously important to have them go in in production, but I think even a good start would be to have those running as you're prototyping your software and just see what if the failures bring you?

And another trend we've seen recently is the use of TLA Plus as a formal verification language where you can effectively spec your system using these formal languages and then test it using verification software so that it highlights places where your assumptions were not checking out with reality effectively.

Corey: The challenge that I've always had when looking at, I guess shall we say older environments and older architectures is that any of the days, but what you just described is very common, where you wind up taking an existing on-prem data center app and more or less migrating that wholesale directly as a one-to-one migration into the cloud. That was great when you could view the cloud as just a giant pile of, I guess, similar style resources. But now with 150 something in AWS alone, the higher level services start to unlock and empower different things that weren't possible back then, at least not without a tremendous amount of work. You talk for example, not having enough people around who can program FPGAs. Do you think that if you were building AdRoll today for example, you would focus on higher level services architecturally? Would you go server-less? Would containers be interesting? Or would you effectively stick to the tried and true architecture that got you to where you are?

Valentino: Probably, I would probably do a mix. I think what's important to evaluate as building infrastructure is the skill set of the people that you have working on your team. And you certainly need to play to their strengths. Ultimately they are the ones who are building and maintaining your infrastructure, not Amazon, not an external vendor. And most certainly not the open source maintainer of whichever project you use in alternative. And the other aspect is try to understand sometimes it's in subtle indications from Amazon, which services Amazon is investing most of their energy or a lot of their energy in, so that you know that they continue to grow and they continue to receive support and they continue to fix bugs and issues because you know that they'll be with you for the rest of your company's life, for example.

But on the other hand, a lot of times you write software just automatically without really thinking about the better way to write something just because you're used to it. And so typically it's not an easy thing to just jump out of the habit of getting an instance going to do something. And it might be a good idea at first, but if you develop a good process to test new architectures and new ideas. That you might quickly end up realizing, well, actually I don't need to run a T2 micro or whatever for running this particular thing with S3 where every time a file is uploaded on S3 I run some checks on the file that is uploaded to S3. You might realize, well, maybe the best thing to do is to try to play around with a Lambda function instead. And that effectively fixes your entire problem.

One area for example, we've tested around and it's on our ... On AdRoll's technical blog is that we built a globally distributed, eventually consistent counter that uses DynamoDB and Lambda and S3 together and effectively is able to aggregate all of the counts that are happening in each of the remote regions into a single counter in a central region that can then be synced back to each remote region. This way we can keep track of, for example, in our case, how much money has been spent in each particular region, and be sure that this money is spent efficiently. And the only other alternative way to do it is to set up a fairly complex database of your own, make sure that latency of updates is fast enough and that all the machines are up and running all the time. And if anything goes down, it's a high urgency situation because your controls on the budgets go away.

So it's sometimes it's really useful to, especially when dealing with problems in which communication and the flow of information isn't particularly easy to grasp for an engineer. It's easy to be able to remove an entire layer of a problem and be reliant on someone else to be providing the SLA that they are promising you. And so effectively that's the case for what for Lambda is if there's obviously a particular range of uses in which Lambda makes complete sense. But from the point of view of price and from the point of view of the resources needed or the type of computation that runs on it. And if you can manage to keep this in your mind when you're making decisions and ... Or you can make some tests, you can actually discover that maybe you can use Lambda and you get away with not having to solve quite a challenging problem at the end of the day.

So sometimes it helps to rewrite in some infrastructure just as an exercise. What I do at AdRoll is that I do as a CTO, I tend to not have a lot of direct reports. I consider each service at AdRoll to be my direct report as a team effectively, each of them being a team. And they every six weeks, they provide a short presentation in which they explain the budgets that have gone through, whether they have overspent or underspent and why.

And among the many things, they also talk about their infrastructure. We have diagrams of infrastructure. We talk about new releases from Amazon and what would be a new way to build the same thing. And they evaluate whether it would save money or not. And so you kinda need to have someone in the organization, especially if you're planning to adopt some of the new technologies that their role is effectively dedicated to being up to date with what's going on in the world and knowing the infrastructure of your systems and be able to make suggestions, and then let the team make the decision at that point.

Corey: What's also sometimes hard to reconcile for some people is that these services don't hold still. And I think one of the better services to draw this parallel to is one I know you're passionate about. Let's talk a little bit about S3. Before we started recording this show you mentioned that you thought that it was pretty misunderstood.

Valentino: Yeah, yeah.

Corey: What do you mean by that?

Valentino: Well, S3 has been, in my view, it's been one of the closest things to magic that that exists inside AWS. Until not long ago the maximum amount of data that you could pull from S3 was one gigabit per second on streams. You were limited in the number of requests per second that you could run on the same shot of S3. There was no way of tagging objects. There was the latency on the first byte when S3 started was in the two, 300 milliseconds. It was expensive. S3 probably has undergone some of the most cost cutting that you could see out there. And part of the decreasing costs is that the standard storage classes become cheaper, but also they have added significant other storage classes that you can move your data in and out of relatively simply without having to change a service effectively.

It's the same, very same API but different cost profile and storage mechanism. And when it all started, there was just one, it was just US standard and it was pretty expensive to use both in terms of per request costs and storage costs. But yeah, today there is the limit on the single bandwidth. The bandwidth on a single stream is not one gigabit per second anymore, it's at least five gigabits per second. If you can get 100 gigs on one of the instances that have 100 gig networking inside Amazon, you can get all of those 100 gigs out of S3 just fetching multiple streams. The latency that you got on the first byte is well below 100 milliseconds. Their range queries are very well supported so you could fetch logs inside S3. S3 is turning into almost a database now. With S3 select it allows you to run filters directly on your files by decompressing them on the fly and recompressing them afterwards. Or simply by reading richer formats like it could be Parquet, for example.

It honestly is something that it's hard to imagine how you could build everything that we have going on right now at AdRoll without S3. It has gotten to the point where running an HDFS cluster for us is not really that useful. If you look at EMI themselves, they have a version of HBase that runs backed by S3. And I know of extremely big companies that have moved from running HBase backed by file system HDFS to instead HBase backed by S3 that have had incredible improvements in performance and the consistency of the performance of HBase. And HBase is very sensitive to the performance of the discs because it's a consistent first database effectively. And if the region that is currently master ... So if the server that is currently master for the region is slow, it ends up bringing down that entire region effectively. It's a service that has grown dramatically. And we have experimented even when using it as a file system by using user file descriptors in the kernel. More recent versions of the Linux kernel allow user file descriptors. And if you have limited use for writing, like we do, and you want to treat the file system like a write once read many file system, then S3 becomes actually surprisingly useful as well.

Netflix published a blog article in their tech blog talking about, for example, how they use a way to mount S3 as a local file system in order to using FFmpeg to run movie decoding and transcoding. Because effectively FFmpeg was not created with the idea that S3 was around and so it needs to have the entire file available on the local system or at least an entire block available. It doesn't work well with streams. And so if you can abstract that part away from the FFmpeg API and move into the file system, you can suddenly use S3 as some kind of almost a file system. And we've done a similar thing when it comes to processing columnar files or indexed files from East side S3, where if you know exactly the range of data that you want to access, you can just do it inside S3.

We use it as a communication layer between the map and the reduced stage of our homegrown map reduce frameworks. And it's again, it allows us to cut away thousands of hours per day on waiting times for downloading a file to the local disk for processing it on local disk. We can just process it right away and cache it on the box after it's been downloaded. It's quite remarkable. The speed increase, the cost decrease, the S3 select. I think we're going to see in the near future databases that start to use S3 as the actual backend for their storage more and more without worrying about the limitations of the current disk.

And effectively we'll be able to scale in a stateless way adding as many machines as you want and respond to as much traffic as you can without needing to worry about failures either. It's an incredible amount of opportunity and possibility that is coming down in the future that I'm really excited to see become real.

Corey:

I think that requires people to update a lot of their understandings about it. I mean, one of the things that I've always noticed that's been incredibly frustrating is that people believe it when it says it's simple storage service. Oh, simple. And you look on Hacker News and that's generally the consensus. Well S3 doesn't sound hard. I can build one of those in a weekend. And you see a bunch of companies trying to spin up alternatives to this, companies no one's ever heard of before. Oh, we're going to do S3 but on the blockchain, is another popular one that makes me just roll my eyes back in my head so hard I pass out.

You're right. This is the closest thing to magic that I think you'll see in all of AWS. And people haven't seemed to update their opinion. I think you're right. It's getting closer to a database than almost anything else. But I guess the discussions around it tend to be, well, a little facile, for lack of a better term. Well there was this outage a couple of years ago and it went down for four hours in a single region. And that's a complete nonstarter so we can't ever trust it. Who's going to be able to run their own internal data store with better uptime than that? Remarkably few people.

Valentino:

Yeah. I mean AdRoll has used 17 exabytes of bandwidth from S3, just for our business intelligence workload from EC2 to S3 this past month. I don't even know how. How do we even start, if a router communicated between S3 and whatever instance we have going around goes out and we're out, we're out for good.

S3 has multiple different paths to reach EC2 and they are all redundant. Each machine's internal there is obviously redundant. They're replicated in multiple zones. And well now this bandwidth is available across multiple zones because I'm storing data inside the inside the region, so it's already available in multiple data centers. The number of boxes that are needed to aggregate to 17 exabytes as well is quite impressive. We have no people thinking about this. I think we run over 20 billion requests per month on S3. I'm pretty sure that bucket, if it were made public, would be one of the biggest properties in terms of volume on the web. And I just can't see it. Processing 20 billion events per month with files that are sometimes significantly big, it's going to take a lot of people.

Corey: Exactly. People like to undervalue what their own expertise, what their time costs. The opportunity cost of focusing on that more than other stuff. And you still see it was strange implementations of trying to mount S3 in a FUSE file system. Trying to treat it like that has never worked out for anything I've ever seen but people keep trying.

Valentino: Yeah. The FUSE file system a is an interesting one I think. I think things might change in the future, but it needs to be done with some concept of what you're doing. It really isn't the file system, but it works for a certain subset of the use cases. And we're not even talking about necessarily yet all of the compliance side of things.

So encryption, ability to rotate your keys to set permissions on who can or cannot access, tabbing each object, building rules for accessing the objects or the prefix based on the tags available on that object using BIM policies and-

Corey: Life cycle transitions, object locks so no one can delete it. Litigation hold options, and then take a look at deep archive. $12,000 a year to store a petabyte? That's who cares money.

Valentino: Yeah, that's exactly, absolutely right. Plus it doesn't matter. At a certain point in time if you're not compliant and you're storing that much data, you just can't. So you're gonna have to delete it all. There's a lot of different security regulations, GDPR is incoming. Is your database going to remain compliant? Well GDPR isn't incoming, actually. It came out a year ago. But with GDPR here and the California privacy law incoming next at the end of the year and the end of this year, is your data, is your storage system going to help you to become compliant? Who's going to build all of the compliance tools on top of your storage system and make sure that you remain compliant for kingdom come? So that basically it's a ... I mean, it's awesome and I think it's a healthy exercise for every engineer to always question what is the value that you're getting out of a service? And try to scheme or understand the infrastructures, try to whiteboard it out and maybe do quick cost estimation.

But it's never enough to have just the engineer in their. Security is a stakeholder of these kind of decisions. The operations team is a stakeholder in this kind of decision. The business is a stakeholder in this kind of decision. The business might not be happy, as you said, to spend $500,000 for two engineers to work on S3 a year when they can spend $12,000 to get a petabyte stored here inside S3. It's just a lot easier. 12 grand is really who cares money.

Corey: Exactly. Especially when you're dealing with what it takes to build and run something that leverages that much data. It becomes almost a side note. And the durability guarantees remain there as well. It feels like one of those things we could go on with for hours and hours.

Valentino:

Yeah. And the other aspect that is very important is how close S3 is to the computing power. Because as I said, 17 exabytes of data just for BI purposes, I cannot do that cross data center. There is no way. That would cost everything from the business in terms of bandwidth costs. Many times other vendors approach AdRoll obviously asking for using their storage solution, but you're simply either to deploy you I need my own data center and then you're not close to where the capacity is, or you are in another system where I don't need a data center but you're not located near my compute capacity. And so I lose that piece of the piece of the equation that makes all of the stuff that I want to do worthwhile.

To an extent, S3 is the biggest locking reason behind EC2. It really is hard to replicate all of the different bits and pieces of technology that are built on top of S3. And in particular being so close to so many services that are easy to integrate with each other, where using things such as Lambda or EC2 makes it very compelling. Other cloud vendors are obviously always on the catch up and getting there, but I don't think they're quite to the level of customization, security, compliance and ease of use that really Amazon S3 has. It also has really hard aspect to it as well, but I think by and large it's a huge success story.

Corey: If people are interested in hearing more, I guess of your wise thoughts on the proper application of these various services, where on the Internet can they find you?

Valentino: Oh, I ... The easiest way to find me is to shoot me questions, comments or follow me on my Twitter account, @dialtone_. AdRoll also has a tech blog at tech.adroll.com. We usually publish a lot of interesting articles about the ongoings with our infrastructure. Things such as the globally, eventually consistent counter that I mentioned earlier, but also our extreme use of the spot market effectively, or our strange use of S3 as a quasi file system for processing our MapReduce jobs, which also are described in our blog.

And generally speaking, I'm more than happy to answer questions and what not to local events. I usually go to as many local events as I can here with the ... Either AWS user events or other meetups, or I go to a random set of other conferences as well.

Corey: Thank you so much for taking the time to speak with me today. I appreciate it.

Valentino: Thank you.

Corey: Valentino Volonghi, CTO of AdRoll. I'm Corey Quinn, and this is Screaming in the Cloud.

Host: This has been this week's episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com, or wherever fine snark is sold.

This has been a HumblePod production. Stay humble.

View Details

About Richard Boyd

Richard is a Cloud Data Engineer with the iRobot Corporation’s Cloud Data Platform where he builds tools and services to support the world’s most beloved vacuum cleaner. Before joining iRobot, Richard built discrete event simulators for Amazon’s automated fulfillment centers in Amazon Robotics, ensured your Alexa device had all the skills you could ever want on Amazon’s Alexa team, held test engineering lead roles at BAE, cyber warfare systems analyst roles at MIT, and research roles for the Center for Army Analysis. He holds advanced degrees in Applied Mathematics & Statistics.

Links Referenced

  • https://aws.amazon.com/deepracer/
  • https://aws.amazon.com/api-gateway
  • https://aws.amazon.com/lambda/
  • https://aws.amazon.com/cloudformation
  • https://twitter.com/cloudtrekau/status/936300151005626368
  • https://twitter.com/rchrdbyd
  • http://rboyd.dev
  • https://richardhboyd.com
  • https://linkedin.com/in/richard-h-boyd

Transcript

Announcer: Hello and welcome to Screaming in the Cloud with your host cloud economist, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of Cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey Quinn: The episode of Screaming in the Cloud is sponsored by O’Reilly’s Velocity 2019 Conference. To get ahead today your organization needs to be cloud native. The 2019 Velocity program in San Jose from June 10 thto 13th is going to cover a lot of topics we’ve already covered in previous episodes of this show. Ranging for kubernetes and site reliability engineering over to observability and performance. The idea here is to help you stay on top of the rapidly changing landscape of this zany world called cloud. It’s a great place to learn new skills, approaches, and, of course, technologies. But what’s also great about almost any conference is going to be the hallway track. Catch up with people who are solving interesting problems, trade stories, learn from them, and ideally learn a little bit more than you knew going into it. There are going to be some great guests including some people who have been previously on this podcast including Liz Fong-Jones and several more. Listeners to this podcast can get 20% off of most passes with the code CLOUD20, that’s C L O U D 2 0, during registration . To sign up, go to velocityconf.com/cloud that’s velocityconf.com/cloud. Thank you to Velocity for sponsoring this podcast.

Corey Quinn: Welcome to Screaming in the Cloud, I'm Corey Quinn. I'm joined this week by Cloud Data Engineer, Richard Boyd, of iRobot. Welcome to the show, Richard.

Richard Boyd: Thanks for having me.

Corey Quinn: It's sort of fascinating to be able to talk to you because for the longest time it felt like the only person that worked at iRobot was Ben Kehoe plus a whole bunch of robots. The first time I met one of his coworkers I figured, "Oh, you can hire guest actors to come in and take care of it and stand in for people who actually work with me." But no, it's all robots, all the way down. Having spoken to you a couple of times now, I'm pretty sure that you're a real person who does work on computers. What's it like to be the second employee?

Richard Boyd: It feels great and actually Ben has been taking some ventriloquist lessons, so it's entirely possible that it's still just Ben.

Corey Quinn: He's an interesting one. There's no doubt about it. I had him on the show before, hoping to get him on again. Let's talk a little bit, before we get into what you're doing these days, about where you come from. It's my understanding you were at Amazon previously working in Amazon Robotics.

Richard Boyd: Yes, I was in Amazon Robotics. I was on the pod management team which is a team that owns the inventory and the life cycle of the pods that the robots in the warehouse drive around and bring to the associates. Specifically, we managed all the items that were on every pod, so we had a bunch of large databases that held all the information about every item that was in every automated Amazon fulfillment center.

Corey Quinn: What's interesting about iRobot is that invariably your architecture, in some form or another, winds up in a number of different AWS keynotes where it seems that you're trying to almost treat their services like Pokemon and catch them all, where every service imaginable winds up on that board. At some point, it almost looks like iRobot is not satisfied to wind up implementing all of the AWS and Amazon services, now they're starting to implement Amazon employees. You almost seem to be an example of that.

You were in the Robotics group and now you're at iRobot, is there a lot of similarity between those?

Richard Boyd: There is not, and it was actually something that came up when I told my previous manager that I would be pursuing a new opportunity, is that they started asking about the non-compete agreement. After some back and forth, we realized that the only similarity between the two organizations is the word robot. One makes robots for internal, industrial use, and as we all know, iRobot creates the world's most loved vacuum cleaner.

Corey Quinn: Roger, the GM for RoboMaker at AWS, was a previous guest on this show, and it's fascinating just listening to how different people in different arenas of robotics tend to contextualize what they're doing. We're starting to see it emerge in a bunch of different groups. We're seeing it come out of a lot of different companies out there. I always felt like I was a little behind the curve on this space because I look at my life here in San Francisco, and I don't see too many problems that a whole bunch of robots are likely to solve. Maybe that's because I have not a lot of square footage.

There is something to be said for the Roomba that runs around and cleans up after me is awesome. It terrorizes the dog, double awesome. But there's also serious uses for this. For example, helping it gather things from a warehouse so there's less human toil involved. There are terrifying things too, like self-driving cars, where the failure mode involves a bunch of people dying. That also winds up having a whole bunch of eyebrow-raising questions. Now, you see the DeepRacer stuff coming out, where it's making effectively robotics accessible to a whole new generation of people who might otherwise not be involved with it in the same way. What's your take on that?

Richard Boyd: I keep coming back to this idea of the things that are very easy are still hard to automate with robotics. We don't have a robot that will go and take out the garbage, take it out of the can, and remove the bag, and do that like a custodian would. You also don't have robots capable of operating at very high levels. There's no CEO robot that you can buy, but there's a lot of stuff in the middle where it's just automated, things that we've seen decimated with automated processing, that I think we're seeing that move into the physical space.

Whereas you used to have people who would be all these middle layers in a company that would just process paper according to some rubric that was put out. That was first automated with sick software and there was people that if they were physically doing something that was still very monotonous, the physical equivalent of that, their job was safe. I think that we're seeing robotics kind of encroach there in the physical world where automation and software encroached in the intellectual property-type world previously, like the late 90s or early 2000s.

Corey Quinn: Something that I don't think I've discussed previously on this show is that I started my career doing very large scale email systems administration. It was fun. I enjoyed it because I have mental problems. But past a certain point, it becomes pretty clear, about a year in, that it's ballpark, 2006, 2007, and it seems to me that there isn't going to be a bright future for that career trajectory where every company needs an email administrator. So, I look around and start trying to find something else that is going to have more staying power, and as it turns out, I picked configuration management. I bet on the wrong horse on that one, but the lesson that I picked up from this was that, to some extent, there's always an evolve or die model where you have to wind up uplifting skill sets and learning new things and embracing a new world.

It always seems to me that there have been these widespread fears with every major technological revolution, dating back to probably Gutenberg's invention of the printing press in 1437. Every time they have it, it hasn't really come out to anything as significant. More jobs are invariably created than are lost. Sure, if you take a look at the invention of the automobile, there's an awful lot of people who are in the horse industry that needed to find other things to do. If they were able to make the jump and retool, then they were set. If they weren't, they eventually wound up forced out of the market. How we handle that as a society is a largely separate issue.

From my perspective, I always thought that, okay, this is what you need to know and learn to be able to learn new things and be effective in a world of computing. Now, I look at people entering the workforce, and they didn't have to go through all that. I think that forcing people to learn to use Microsoft DOS and work your way forward from that evolutionary step is a giant waste of everyone's time because you don't need to know that. This, of course, brings us around to the idea of serverless, which is something that you've been extremely vocal and passionate about.

Richard Boyd: Yeah, and one thing I'd like to add to that is I think the only thing ... You hit the nail right on the head when you said that this problem's been around since Gutenberg, where a new technology comes out, it runs the risk of eliminating the livelihoods of sectors of the economy. I think previously, that rate of change was slower, so that even if you didn't adapt, your career would end. If you didn't adapt, it took so long that you would just retire and then nobody would fill your place. The person who would have filled your place would do something slightly different.

I think that we're seeing an acceleration and a shortening to timelines for when these things happen where someone might have to make that choice several times throughout their career of I need to adapt otherwise I'm not going to have a job. I think that acceleration is what's making this problem unique from previous technological innovations. I don't think that cars went from no cars on the road to everyone having a car in 10 years or 15 years like we're seeing with serverless eliminating jobs in SCO Systems Administration and what not.

Corey Quinn: The question too becomes how widespread is this serverless paradigm? There are few companies, iRobot is a notable example of this, that are very serverless forward. In a number of companies that don't have the same born in the cloud mentality or don't have the willingness to embrace new technology at the same pace for a variety of reasons, some good, some not, it seems like serverless is still being adopted but only in very specific use cases or in non-production capacities or it's still viewed to some extent as a toy. That's definitionally going to have to change. There's going to be a lot of growth in that space.

The question becomes, I think, how quickly does the tide rise? If you're a network engineer and you are 40 years old today, is there 25 years or so of runway in being a network engineer or not? If you're 22 years old and getting into network engineering, that same question applies and it may have a different answer. Of course, if you're 60 years old and getting into this, it's likely a different answer again.

The question largely becomes as the tide rises and more of the iceberg gets submerged of things you don't have to care about to serve your business, what does that mean in the future? What advice do we give people who are entering the workforce today?

Richard Boyd: That's a very good question.

Corey Quinn: Welcome to the podcast. You're now a thought leader.

Richard Boyd: I think that I see two main tracks. What I've often described it ... I'm going to talk about this from an engineering perspective just 'cause I come from an engineering background ... is that I always describe what I call the t-shaped engineer who only knows a little bit about a lot of things and then this one area very, very, deeply. I think that will still happen in the future. We'll need engineers who are broadly trained but have a specialization. I think we'll see more people who are flexible in what that specialization is. Instead of being the thought leader or something for their entire career, someone might be a thought leader in that thing for part of their career and then they'll go into something else as the tide changes.

They may not make as strong of contributions in there, not to say that their work is any less valuable. It may even be a third or fourth time in their career they'll switch and the depth of that T in this t-shaped engineer will go deeper or... But that will change throughout the course of their career. I think going into their career and starting out with this flexibility saying, "Yes, this is what I'm doing. This is what I think I'm going to do for the rest of my career, but it's very likely that I won't be." Just being accepting of that new reality where ...

Corey Quinn: The t-shaped engineer is something that we've spoken of at least once or twice on the podcast before. I think one of the challenges there is figuring out what is that thing you go deep on. I don't think you can be a systems administrator without being a little bit of a jack of all trades. Being deep in a specialized area is fascinating. For me, it started off as email, then it was configuration management, then it turned into cloud somehow. Now, I'd argue it is cloud billing. But, there's going to be a time when everything changes. I think that most of us don't have a job that our grandchildren are likely to recognize when they enter the workforce.

That can be a scary thing and there are other folks who think that that's a wonderful opportunity. I think the answer's probably nuanced and it's a combination of both. I'm very cognizant of not wanting to drive people out of the industry. I don't want to automate everyone's job away. I think that's awful, but I do think that we want to automate the boring parts of people's jobs. I think that we want to make sure that you don't need to have an email administrator to start a company in 2019. I think we've gotten there by in large.

What's fascinating to me about serverless is the level of what you have to be aware of and what you have to do is almost completely vanishing.

Richard Boyd: Yeah, I think this was another similarity that I saw between iRobot and Amazon Robotics is that they both were focused on we're not eliminating jobs. We're not getting robots to eliminate jobs. We are, in the case of Amazon Robotics, eliminating the parts of the job that don't add value to either the associate or to the company, which is the walking. Walking and carrying stuff is very stressful on your knees, your hips, your ankles. Inside iRobot it's that no one wants to clean a dirty room or vacuum.

I think that this becomes the mantra of serverless is that this type of stuff frees you up to focus on the things that you actually want to do and the things that provide you value, whether it's intrinsic value of things that you enjoy doing or monetary value for the number of customers that your business is able to capture.

Corey Quinn: One of the things that first brought you to my attention was a series of blog posts you wound up posting. I'll put a link to one of them in the show notes of your rboyd.dev personal blog. The one that I really saw that resonated with me was called Mastering API Gateway in 105 Easy Steps. I thought it was going to be a teardown of the terrible level of complexity and documentation on API Gateway and I clicked that with an excited, "Oo, this is going to be good." And, it was good, but it wasn't what I thought it was going to be. Why don't you tell us a little bit about what it was?

Richard Boyd: Yeah, getting into the genesis of what started this is that we have a bunch of data in an S3 Bucket and we had Athena sitting on top of it that could do some queries on it. Standard use case. I needed a separate AWS account that we also control to be able to execute queries on this. There's no cross-account access, so the canonical way of doing this is have API Gateway in front of a Lambda. The Lambda executes the query, takes the data back from Athena, gives it back to API Gateway. It felt like the manager who gets laid off in Office Space where it's like, "I take the requirements from the engineers. I'm a people person!"

Richard Boyd: I was like, we don't need the Lambda here. In the dropdown menu, in API Gateway, it says Athena is an option. I can just click on this and it should just work. I hear from developer advocates from Amazon all the time saying how easy this is and it will free you up to do all the things you love doing that's not configuring API Gateways. So, I try it. I spent a lot of time on it and it just doesn't work the way I would expect it to work.

Corey Quinn: You are absolutely not alone in that. One of the biggest complaints I have about the evangelists and advocates at AWS is they consistently talk about how easy this stuff is. Spoiler, the first time you do it, it is absolutely freaking not. My first Lambda Function took me two weeks to get working because I'm a terrible developer. The last one I wrote took me about 15 minutes because I don't write tests because I'm a terrible developer. There's a heck of a learning curve the first time as nothing behaves quite the way you think it will.

I understand the message they're trying to get across of once you understand a few basic tenants of how this works, it does speed development time. But, there's nothing more off-putting then trying to do something new and start experimenting with it and being told, "Oh, well actually it's very simple." Then, you have trouble with it, so the unspoken assumption is, "Oh, I must be a moron. I never knew." That's not a great feeling for anyone who's trying to learn something new, none of which, by the way, is easy.

Richard Boyd: This is one of those things I try to avoid when I'm showing someone how to use something or if I'm writing a blog post. I try as much as I can to avoid two words. One is, oh, that's easy, and the other one is the world "just". Because it's never just X. It's just X, so I think people care about... When you say on a call, this is easy, the person, even if it is actually easy, it's easy for me doesn't mean it's easy for someone else and it can be very off-putting.

Richard Boyd: I think that it's a bit deceptive in API Gateway because even if you do understand a lot of their tricks ... I won't say their tricks. Even if you do understand a lot of the ways these services tend to communicate with one another or you understand the ways that information is passed between those two, the thing we keep coming back to with AWS is that the only thing they're consistent in is their level of inconsistency. This example in API Gateway just really highlighted that.

Corey Quinn: Anytime someone tells me that, "Oh, it's very easy. All you have to do is just," my immediate response is, "Okay, you are now assuming everything lives in a whiteboard vacuum. You have no real conception of the constraints that shape what's currently in place and you're not very subtlety telling people who worked on this previously that they're kind of stupid." That is just absolutely one of those unforgivable sins from my particular point of view. I do my best to call it out when I see it, but no one's born knowing this stuff and implying otherwise is, frankly, terrible.

Richard Boyd: There's a chart that has the value of ignorance where if someone's proficiency goes up, their perceived efficiency goes up much ... I'm sorry. Let me start that over again. As someone's proficiency on a subject goes up, their perceived proficiency goes up much faster and then drops off once they realize they don't know what they're doing. That's where I tend to put people who say, "Oh, it's super easy, you just do X." Then, it's like, "Oh, this is your third week in dealing with AWS."

Corey Quinn: Exactly. What's fascinating as well is the level of confidence you espouse does change how people shape you. That's also very heavily driven by gender and race. That is a terrible factor of our industry that I don't want to go too far into on this particular episode. There'll be an episode on that one of these days once I figure out the take I want to have on it. But, as it turns out, and now, a white guy opines on diversity is not generally something I want to dive into immediately.

Back to the API Gateway project that you have going on. It looks like you're attempting to integrate API Gateway with everything that it can integrate with that isn't Lambda.

Richard Boyd: Yes, and I'm trying to take it one step further than that is that I'd like to integrate with everything that it can integrate with that's not Lambda as well as make it so that no one else has to go through what I went through. Yeah, so what we are trying to do is make it so that the pain that we experienced trying to integrate API Gateway doesn't have to be experienced by other people. 'Cause it would be very easy for me to say, "Okay, here are all the steps you have to do to learn what I learned," which the person still has to do all of this learning and just memorize all this, what's essentially useless information just to do what they want to do.

We're trying to make it so that your confirmation template looks just like a Boto 3 SDK goal. So you can say ... Pick a service. Every NM will say Athena. So, you can say, "Athena, start query." Then, you can look at the documentation and see what parameters that takes and supply those as native CloudFormation, like YAML constructs in there. You can go and you can reuse the same documentation that already exists for the REST API. You know what's required. You know what's not. You know what the shape of the objects it requests are.

I'm working on creating a CloudFormation macro that does that transformation for you. Instead of teaching these people, "Okay, here's how you decipher the Rosetta Stone of API Gateway to other services," instead I say, "Use this universal translator that you can wear on your comm badge and then you don't have to worry about speaking that language. I'll speak it for you."

Corey Quinn: I think it's a laudable goal. I think it's absolutely going to be transformative for a number of people. I have to ask though, because I ask myself the same question from time to time, doesn't it feel slightly weird to effectively be doing what amounts to volunteer work for an almost trillion dollar company?

Richard Boyd: Yes. It does feel weird doing this. There are times when I've been thinking like, "Amazon could just do this." We have been in touch with the API Gateway and the CloudFormation teams, or at least we've had several meetings with them, and they're picking up the lion share of this work, the hard part. The thing that I am building is only connecting some various parts and creating essentially a rough outline of how I'd like to use this as a developer. Then, they are going and they're building the actual infrastructure behind that that makes it possible.

I won't be building the whole widget. In the end, I'm just building the façade on the front of it, saying as a developer, this would be the most zero-friction way I could use your service and this would make me love it more than I already do. Then, they're building the stuff on the backside that makes it possible. I think you mentioned in a previous post that Amazon's very good at plumbing and terrible at painting. So, what I'm doing I'm hoping to do the painting, or at least I can fill in the shape for the color-by-numbers so that they can do the painting, and then letting them do the plumbing on the other side.

Corey Quinn: Yeah, to be clear, that was someone else's analogy, not mine. I love it and I'm probably going to steal it in the future and claim it as my own, but not yet. You're right, it's one of those areas where to some extent, the teams building these products and services feel almost like they're too close to the problems themselves and they don't know what it's like to have no idea what this thing is, how it works, how to conceptualize it through a different lens. That, I think, is one of the most valuable aspects of the larger AWS community.

The other side to that though is, to some extent, it almost feels like certain teams aren't able to get internal traction so they start turning to the community to develop these things that, to be very direct, probably should come from the cloud providers themselves. I'm seeing this from all of the major cloud providers except Oracle, which, to the best of my knowledge, does not have a community of which to speak.

It's a bit of a held intention thing from my perspective. If a company that were the most AWS came and asked me to do some of the things I do in the AWS ecosystem for them, I'd quote a price and it always seems weird to look at it from a perspective now and how did we get here.

Richard Boyd: The way I justify it to myself is that there is a set amount of pain that I have to pay every time I want to do some with CloudFormation and API Gateway that if I could do this once before everyone else, then I don't have to pay that in the future when I'm trying to do this and everyone else is also benefiting from that. So how I save time that I have in the future, is how I compensate myself for the time that I spend on this now.

Corey Quinn: That's a very good point. It tends to be something that resonates with people, clearly. When I linked against this article in my newsletter, it was the most popular link that week. That's depressing to some extent because I work super hard on some of the other links that wound up going into there. I'm kidding because it's great to see what resonates with the community and what doesn't. What I see from that is that everyone is wondering about these things. Everyone is having weird questions raised about how to address these things. I don't know what that says about the current state of adoption. I don't have data on this. I have anecdata at best.

Being able to take a look at this through a lens of understanding what people are doing, what the use cases look like, and how that winds up manifesting is fascinating to me. I'm grateful to have the opportunity to do it, but the more I do this and the more I see certain things resonating with folks that I didn't see coming, the more I realize that I don't have a complete view of this industry and I'm almost positive that no one does. All we see are different glimpses across a wide spectrum of experience.

Richard Boyd: Yeah, I think that this goes to a previous comment that you had made about how serverless is still treated as a toy or a sub-prod type problem because people are still using the same API Gateway to Lambda to the service that you actually want to use. It's really weird that the same people who are advocating serverless lets you do the thing that you want to do, but, by the way, you still have to own this Lambda function in the middle that you actually don't want just to do the thing that you want to do. I would expect these people to be much stronger advocates for why is it the commonly accepted approach just to stick this Lambda in the middle to make this easier?

Corey Quinn: That also seems to be in line with a tweet I saw a while back and I'll throw it in the show notes as well. Talking about how, "Oh, the only code you will write in the future is business logic." That's been promised and there was a list of five or six different times that had been promised previously. Nothing ever quite got there. Oh, this time with serverless, that's what it's going to turn into. It feels like it's closer than anything else has been, but, at the same time, history rhymes. There's a question as to whether or not we'll get there or now it's just a whole other list of things that we have to care about.

What I do see is a future where almost all code that gets written is front end code. The backend is generally handled either via configuration or via some sort of specific defined language that only does configuration rather than running infrastructure in most environments. There's a whole bunch of on-prem folks and multi-cloud purists, who are wrong by the way, that this is somehow not going to work, but most people aren't trying to make a political statement with their environments. They're trying to get a large business done. They're trying to solve an actual problem that doesn't involve, well, first we're going to rack and run a whole bunch of servers. That's not really what people want to be doing unless that is your specific business.

Finding a way to wind up doing that intelligently is an ongoing problem emerging in this space. I think that how we're always trying to get rid of the thing that isn't core and central to a company's business is important. I think serverless gets a lot closer to that than most previous attempts.

Richard Boyd: Yeah, I see people focusing on this Lambda cold start problem as an analogy to that where ... I had a tweet about this where I hear engineers complain every day about Lambda cold starts. I've never heard a customer ever say that they care about a cold start. Because they don't. A customer cares about user-perceived latency and the engineers, perhaps inaccurately, ascribe all of their cold start latency to this user-perceived latency.

Corey Quinn: Well, I think you're absolutely right. I wound up using two serverless monitoring products that I'm not going to name in this context. One of them was lightning fast and spat out what I needed. The other one likes to sit and spin as it gathers data. I reached out talking to the companies about this and it turns out that the fast one was using something that involved servers on the front end and the other one was trying to itself be a completely serverless architecture. My complaint was not, "Oh, that was a cold start," or, "Oh, I don't like the architecture behind this." My complaint was the web interface is super slow every time I change views and that's a crappy user experience. Fix it.

I don't have an opinion about the architecture. I don't want to hear the term cold start. What I want to hear is that, "Oh, we understand what the user-perceived latency looks like and we understand this is irritating you and we're going to be taking steps to fix it." I think that, in a microcosm, is what the industry says. People don't care about your architecture, they care about their own pain. I don't care what code, what the code looks like of any service provider I use in the slightest. I care that it works and I care that it lets me do the thing that I need to do.

You can optimize your code to make it perfect until you run out of money, great. The Bay area's littered with startups that tried that. But, you can have terrible code and get to a business point of profitability and then go back and fix things. The canonical example of that is Twitter. They had horrible code. It was all Ruby on Rails driven to my understanding. Now, their scale is awesome. They don't have reliability problems. They have a Nazi problem, but that's a different topic.

So, if people like what you've had to say about the world of serverless, about where you see things going, and looking more into the nonsense things that you have going on on an ongoing basis, they can obviously apply to work with you at iRobot. But, is there another way people can find you?

Richard Boyd: I'm active on Twitter @rchrdbyd with no vowels, but it does include the Y. And, I'm active on LinkedIn as well.

Corey Quinn: Perfect. I'll also throw a link to your blog into the show notes, which, before I fire off one more question. You now have the blog living at R-B-O-Y-D .dev and all the developers I know jumped at the chance to get a .dev domain. I looked at a few of my own environments where I wrote crappy code and, sure enough, I'm rerouting all of .dev to localhost. So, I have to wonder how many developers out there are trying to read other people's blogs and realize that, "Oh, this doesn't work," or in terrifying moments, "Oh, my stars, somehow he got a copy of my code and put it up on his website."

Richard Boyd: That's a very good point. I hadn't considered that. Yeah, I had richardhboyd.com was the previous blog, which is still up with some of the older posts, but that was an open source template that I had used from someone else years ago. When I told myself, I think it was early 2017, that I was going to start a blog and I made my first post in about March of 2019. By then, the person who made the original template for it had stopped supporting it to move onto something else and it just didn't scale very well. So, I said, "I'm going to do what every good serverless dev does, I'm going to roll my own. It's going to be entirely serverless," which was a fun learning experience that let me exercise some of these API Gateway integration techniques I'd been playing with.

It let me exercise some of these API Gateway integration techniques, but it's still ... The blog is not quite ready for production yet. I'm not a front end person, so my heart skipped a beat when you said, "In the future, everything will be ... My heart skipped a beat when you said, "In the future, everything will be front end code."

Corey Quinn: Yeah, me too because I can't code my way out of a paper bag in JavaScript.

Richard Boyd: Exactly. Same. That's why I was a big fan of the AWS Cloud Developer Kit, but it was only in JavaScript, TypeScript, and Java. I was like, "Oh, well this essentially useless"

Corey Quinn: But I hate all of those things.

Richard Boyd: Exactly. Then, they launched Python support...

Corey Quinn: No language bigotry. You can do whatever you want just for me personally, my brain doesn't work in quite that way.

Richard Boyd: Yeah, and that's another thing that I've noticed with the API Gateway service integration is that if you go on Stack Overflow or Reddit, some of the Slack channels, if you look up JavaScript Lambda and then just name a service at random, you'll see the same question being asked thousands of times. I built a client and I made a call to the service and then I returned, but it looks like my Lambda function is returning before I'm getting any data back and I'm not getting a response. That's because of this weird asynchronous nature of the way JavaScript works and it's doing exactly that. It's firing off a request asynchronously and then saying, "Oh, it's time to exit," and not waiting for the response to come back.

That's another part of the API Gateway service integration work that I was doing is that if you have the API Gateway doing the service integration directly, there is no thinking about JavaScript asynchronous in A-weight type conditions for this. All of those types of problems would just go away.

Corey Quinn: I'm with you on that. I'm in the process of migrating my own serverless blog over to WordPress, which, ironically, is a bit of a serverless success story, but that's a tale for another time. Richard, thank you so much for taking the time to join me today.

Richard Boyd: Oh, thanks for having me.

Corey Quinn: No worries. Richard Boyd, cloud data engineer at iRobot. I'm Corey Quinn, this is Screaming in the Cloud.

Speaker: This has been this week's episode of Screaming in the Cloud. We can also find more Corey at screaminginthecloud.com or wherever fine snark is sold.

View Details

About Richard Hartmann

Richard "RichiH" Hartmann is the Swiss Army Chainsaw at SpaceNet, leading both a greenfield datacenter build and monitoring. By night, he is involved in several FLOSS projects, a Prometheus team member, founder of OpenMetrics, and organizing various related conferences, including but not limited FOSDEM, DENOG, and Chaos Communication

Congress.

Links Referenced:

  • https://velocityconf.com/cloud
  • https://prometheus.io
  • https://www.debian.org
  • https://promcon.io
  • https://fosdem.org
  • https://cloud.withgoogle.com/next
  • https://www.microsoft.com/en-us/build
  • https://reinvent.awsevents.com
  • https://twitter.com/twitchih

Transcript

Speaker: Hello, and welcome to Screaming in the Cloud with your host, cloud economist, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey: This episode of Screaming in the Cloud is sponsored by O'Reilly's Velocity 2019 conference. To get ahead today, your organization needs to be cloud native. The 2019 Velocity program in San Jose from June 10th to 13th, is going to cover a lot of topics we've already covered on previous episodes of the show, ranging from Kubernetes and site reliability engineering over to observability and performance.

The idea here is to help you stay on top of the rapidly changing landscape of this zany world called the cloud. It's a great place to learn new skills, approaches, and, of course, technologies, but what's also great about almost any conference is going to be the hallway track. Catch up with people who are solving interesting problems, trade stories, learn from them, and ideally, learn a little bit more than you knew going in to it. There are going to be some great guests, including at least a few people who've been previously on this podcast, including Liz Fong-Jones and several more. Listeners to this podcast can get 20% off of most passes with the code cloud20, that's C-L-O-U-D-2-0, during registration. To sign up, go to velocityconf.com/cloud. That's velocityconf.com/cloud. Thank you to Velocity for sponsoring this podcast.

Corey: Welcome to Screaming in the Cloud, I'm Corey Quinn. I'm joined this week by Richard Hartman, who has decades in open source. We met originally back when we were f staff and since then he has done a lot of other things too. You were a Debian developer, you organize a bunch of conferences, including but certainly not limited to PromCon, FOSDEM and others that I don't care to think about.

And you come from mainframes, now you're into networking, then you started building out redundant data centers as turnkey solutions, and apparently you're currently building a data center, that I choose to believe is located in the middle of a swamp.

Richard: It's actually a Greenfield Project, and we couldn't build it in the middle of a swamp because we are going for the highest certification within EN 50600 which is security and availability Class 4.

Corey: Gotcha. So among many other things, you're in town here in San Francisco and terrifyingly close to me, for Google Next, which as at the time of this recording, just finished.

You are a member of the Prometheus Core Team, but that wound up driving you out here to sit through, effectively, three full days of talking about Google's Cloud. What do you think?

Richard:

It was nice, it was interesting. Many of the talks were a little bit sales pitchy, like a little bit too sales pitchy for my liking. They usually follow them all over initially like the first third or so maybe, they had some higher level technical details, like not really into depth, then they segued their way into why you should be buying from them.

Which obviously makes sense from that perspective, on the other hand, it's not the type of conference which I'm used to, lets say.

Corey: It feels like all of the major public Cloud vendors have this problem. Once they hit a certain point of scale, they have one big Cloud conference every year. You have Microsoft Build, you have AWS re:Invent and you have Google Next, where the conference is trying to do so many things that it almost loses a sense of itself, where you're trying to sell things to people and there's that sales piece of it, there's trying to articulate a vision for the next year, there's product announcements, you're talking to engineers, you're talking to corporate buyers, there are press in attendance, they have analysts that come through and start to ideally say nice things about them.

And when you get all of that together, it's very hard to build any kind of cohesive narrative that addresses all of those constituencies. So, at some level when you're at one of these it feels like you're always in the wrong place, listening to the wrong story from the wrong people. And I've never found a good way to solve that.

Richard: I don't think there is a good way to solve this, of course, inherently you have all those different priorities and all those different goals and to juggle of them just doesn't work. At least not at that huge scale which they put together. So I'm not actually complaining, it's just an observation which I made that they seem to be this way.

There were other things, also minor, but one other thing which I noticed, the analyst's lounge, which is sitting right smack in the middle of everything, has full catering and everything, whereas the speaker lounge is basically a coffee maker and some granola bars. So that gives you a little bit of insight into the relative value which is assigned to this. But again, I'm not complaining, it's just, I couldn't help but observe that this is happening.

Corey: Credit where due. The press lounge was also super nice.

Richard:

See? That's my point.

Corey: To some extent, this seems like a bit of a departure from Google's historic positioning as engineers, first, last and always. And I think that you sort of have to, once you grow beyond a certain user profile.

It's interesting to see how that's going to be maintained going forward, I mean, there have been enough jokes made about it, but historically sticking to things that are not core to what they've always done, mainly search and ads has always been something that Google has seemed to struggle with.

So while they're saying the right things, I think people are mostly going to adopt a wait and see approach, at least for our time.

Richard: That is probably correct, I mean, from my perspective, Google has absolute top notch engineering, and this is an engineering driven company by in large, so it just stands to reason that a lot of the internal culture is also engineering driven. Which tends to disregard a lot of other needs of other people and teams and organizational units.

So I fully agree, this messaging needs to change for more traditional businesses to actually be able and willing to adopt their product. On the other hand I do hope that they don't lose this striving for technological excellence.

Corey: I would be very surprised if they lost the pursuit of technological excellence, I would be less surprised if they lost their willingness to engage with large enterprises. It comes down to fundamentally I can see them reverting back to what their company was built on, their corporate DNA as it were. I can't see them completely pivoting and abandoning where they've spent the last 20 years.

I'm not saying it won't happen, but I have a hard time imagining it.

Richard: As of right now, I would tend to agree, to be honest, on the other hand if you look at most companies, like the large ones, they had these huge growth phases and they were very very engineering driven and then at some point, what will you be promoted for. And at some point this becomes more like enterprise stuff, maybe marketing, maybe economics, so people with that kind of thinking tend to be promoted more and more as older as the company gets.

So this will, over time, change things, like, I'm not an Apple user, but looking at Apple from the outside, this kind of seems to happen where there's this focus on engineering and on excellence, just gets a little bit lost and their edge also gets lost.

Corey: It's an interesting problem. Changing gears slightly, let's talk a little bit about something you said back when we were preparing for these shows. Specifically that the Cloud is nothing new, it's old again and it's always been this way except for the fact that it's somehow completely different. What do you mean by that?

Richard: What I meant by that is that fundamentally IT stays the same while it completely changes every few years. If you look at any old monolithic application which is huge and horrible and everyone will tell you this thing cannot be maintained, blah blah blah, all these things, still you have functions in there. And functions on a very basic level are not different at all from a microservice.

You change how the API's, how the interfaces, how the service delineations are exposed, you change a little bit of the mix of how you do and it and what you do, and obviously you're always trying to raise the bar for tech as a whole.

And it also comes a little into this thing where I like to say IT breathes, where things go in and out, like you go from one extreme to the other, you internalize and you outsource. You have your monoliths, you have your totally fine grained things. And it just goes back and forth, back and forth. And every time you go towards this other extreme, you're trying to solve a or more problems. And once they've been solved, you will then have other problems.

So you go back to the middle and you overshoot a little, and then rinse, repeat. This seems to be happening a lot. If you do it with too much fervor, you might be overdoing it, on the other hand following this natural life cycle of IT is pretty nice, because you're just raising the bar again and again.

And when you look at Cloud, like all those issues which infrastructure providers have like, how to run a data center, I can tell you, running is even the small part, building it is insanely complex. Like, all these things just go away because you have a different service delineation, and you just build on top of that.

Corey: You're in town to give a talk. Now tell me a little bit first, what that talk was.

Richard: It was titled "Prometheus - What the hype is about", it was a mixture of the usual Prometheus 101, along with why people who are calling themselves Cloud developers, should care about this.

Corey: And what is Prometheus, for those who have not yet attended a Prometheus 101 talk?

Richard: Prometheus is a monitoring framework. It ingests time series data as in numeric data which changes over time, you might think service latency, you might think user count, how many errors you have, temperature, whatever. Just changes over time. It's not geared for events, so you can't put log lines or anything in there, it's purely for numeric data changing over time.

And what you can do with it is you can ingest a lot, a lot, a lot, a lot of data with relatively few resources, like you can easily do on normal hardware, or normal VM, you can easily do a million samples per second and more. It comes out at roughly 200k samples per core. Like, if you want more, just put more cores in, and you're done basically.

So it's super efficient in ingesting the data and also exposing that data back to the user. As you have these immense amounts of data, you obviously need a way to accurately get this data out again. So we have something called Labeling, which is basically key value pairs. And you are allowed to to assign arbitrary key value pairs to your data to then be able to select and slice and dice your data through this n-dimensional matrix which you are building up, so you could do by region, you could do by customer, you could do by prod or def, and all these things which normally are stuck in a hierarchical data model, are all of a sudden available to you as direct first class things.

But having those labels is only half the story. You obviously need some way to actually work with that data, and that's another one of the really nice things about Prometheus. You have this one single functional language which you have to learn, it's called PromQL, and it's basically doing vector math on your monitoring data.

So instead of just having this one graph which never changes and you can't really do anything with it, because you encode stuff into an image file, you can actually take this data and do data science on it. And it's Turing-complete language, it's super powerful. It kind of takes some getting used to but it's really nice once you learn it.

And the next thing is you use this for alerting, you use this for analysis, you use this for graphing, you use this for dashboarding. You can use it to get your data out in JSON format, you have this one single way to access all the data, and it's always the same as opposed to a lot of other systems where you have to think differently about accessing the data, depending on if you want to do alerting, or a report.

Corey: This might be something of a controversial question, or rather the question is not, the answer is probably going to be hotly debated. But at what point does it make sense to do something like that, or to implement something like that, versus deploying one of the many, many, many, monitoring vendors that purport to do not only what you've described but everything else as well?

When does deploying or building your own monitoring system make sense for an organization?

Richard: Fundamentally it's always the same make or buy question. This is no other. Obviously I'm biased, so I would tend to run things myself, which works, and for small teams and such it's super easy to just spin up a new instance and do some monitoring on whatever you want to do. Maybe you just want to do some poking or whatever you want to do and you're super flexible in what you do. But that's only part of the story.

The other thing which Prometheus enables was, it shifted the whole of IT monitoring. And again, I'm biased but from my perspective it actively, it actually shifted or uplifted a whole segment of IT as in monitoring, to a new level.

So there's a lot of vendors which now support similar things, I mean, I do have personal opinions about a few of them, but fundamentally unless they do something completely wrong, it's not a bad thing to use them.

Corey: This ties in, to some extent, to I guess, a past life and something you still dabble in from time to time of network engineering. Once upon a time, if a company wanted to do anything that even touched on IT, they needed to have someone with network engineering expertise in-house. Today, it's debatable whether that's still the case. What do you think?

Richard: You still need people who know how to do these things, but their daily workload will change, massively change. So you might not need someone who, or not a lot of people who are aware of the intricacies of ethernet, or whatever. Like, VRRP setups, tend to be somewhat icky, and if you can avoid them, by all means avoid them. But avoiding them usually means having an overlay network, or having dynamic routing. Which I think is a perfect solution, but it's quite complicated. But again, Cloud shifts the service delineation. And all of a sudden you have to do all those nitty gritty details yourself. You can't buy this as a service. Still you will need someone who is aware of how those fundamentals work.

So you might still need your VPN gateways, you might need someone to connect the VPC from on-prem to your Cloud, or to your multi-Cloud, or whatever. So you still need the knowledge about how things work, but the actual day to day job will change. And by extension obviously the actual skill set needed also changes along with it. But you still need the main experts. Same as in anything else. Like, even if you have a hosted database, it still makes sense to have people who actually are aware of how things work in the background so they can make good decisions about how to set this thing up.

Corey: To some extent people have been saying for generations now, it seems like, that in the future, you'll never have to worry about the undifferentiated heavy lifting, or the toil, you can only focus on writing business logic and doing things that move your business directly forward. I mean, I my own career, once upon a time I started off as a large scale email admin. And that was something every company needed. Today, almost no company needs that. It's click, click, done, with a hosted provider or very occasionally a small central group that runs exchange internally or something like that.

I can't shake the feeling that to some extent the level of expertise required for most companies who are not themselves deep into the, I guess, IT space as what they do, need to have a strong grounding in network engineering that work theory being able to handle complex routing situations, et cetera.

It feels like that has been abstracted away by and large in a lot of, I guess, typical companies. Is that a naïve approach? I recognize you are sitting in San Francisco, where everything here is a web app. There exists an entire ecosystem out there of companies that that does not apply to. I understand that.

Richard:

I wasn't aware you had a career.

Corey:

My parents still believe I don't, it's fine.

Richard: Okay. Yes, like, again, the subset of skills needed changes dramatically. And a lot of those details are just abstracted away behind a new service delineation. So a lot of the things you don't really need in your day to day anymore, like, it still makes sense to maybe have one person, it might not even need to be in the same company who just knows that stuff, of course else, you're bound to make mistakes from the past again and again. Of course, you will always need at least some knowledge of how things work.

But I fully agree that this depth of knowledge fully moves to infrastructure providers. And it's probably a good thing, because most people like most enterprises, at least from my networking perspective, have a really hard time even getting networking people because they just don't care about this type of network.

So, hiding this behind a proper service which is managed by experts absolutely makes sense. At least for those who can do actual Cloud, like you still have tons and tons and tons of legacy implementations, and you have fields and industries where IT is currently nice, but it's not essential. And those have completely different needs. Completely different needs from anyone who's in Cloud web app API world and just living a quite nice life to be honest.

Corey: I refuse to accept that here at Twitter for Pets headquarters. So on a similar vein, serverless has been sort of taking over the world with similar promises, that the only thing you'll ever have to worry about in the future is application code, that it's going to be a magic coming of almost paradise where only pure developments matters, everything else is handled for you by one of several Cloud providers. And everyone's touting this as a new thing. Is it?

Richard: Yes, I think CGI-bin is pretty new. So, on the one hand, again, it's old and it's new at the same time. And it also ties a little bit into this Toil thing, of course to some extent Toil is good, because it lets you learn about how the underlying things work, so you have a better understanding of why something might be happening in a certain way, but jumping back to serverless, the concept of putting a piece of code in some place and having this executed when an external event comes along is not new.

CGI-bin is fundamentally the same. You have a web browser usually, and this makes a call to a thing, and this thing gets executed and it returns with some data and then it dies. So exactly the same thing happens in serverless. Like, you have a lot more emphasis on different API's, you have a lot more emphasis on events. You have this awareness that these events will usually not be generated by a human or by a web browser but by something else. So a lot of those things are evolving in a good and in a nice and more efficient and effective way but fundamentally it's the same as before.

Corey: One of the misnomers that I tend to see from time to time when talking to people about serverless, is that there is a belief that, well, I have some code, and now it's going to take that code and it's going to run it for me. I have a wristwatch that can do that, there's not a lot of value in being able to say that yes, I have a computer. What is more interesting to some extent is, yes, what you said, the event model being able to impact when that code runs, what it takes in, and what it returns. There are economic factors that feel different this time, and maybe that's a bit of a red herring. But the idea of not having to worry in any traditional sense about scaling, that was always a concern with CGI-bin. Not having to worry about paying for things to sit idle, when they weren't being addressed. Instant on, consumption based, economic models start to be transformative for some use cases.

What I think is also very interesting and differentiates this somewhat from CGI-bin, is that there is a thousand different ways to write serverless functions. Most of them are absolutely terrible. Especially things with custom run times, write it in whatever language you want, to my recollection, CGI-bin was mostly a Pearl requirement, wasn't it?

Richard: It started as Pearl. There were other languages which were shoehorned onto it.

Corey: The entire challenge that I see in, I guess trying to view this as sort of the second coming of CGI-bin, is that everything old is new again. What I'm wondering, was there anything in between CGI-bin and serverless. Because we haven't talked about CGI-bin for 15 years in most shops, and serverless as a thing is four or five years old. What happened in between?

Richard: Well, serverless might be the third coming of CGI-bin, of course you have app engine in between. And you had, this ties back to this engineering driven excellence thing where Google was kind of trying to tell people, hey this thing exists and maybe you want to use and maybe we also use it for our own services and have quite some good experience with it, but people didn't really care. It probably just wasn't the current time of, like, in the global market, the time of this shift going in this direction again.

Corey: So as far as CGI-bin versus modern serverless, one of the big benefits of modern serverless is elasticity. The ability to only have things on demand, when you need them, and you don't pay for them when they're not hanging around. Back in the days of CGI-bin I still had to provision servers, carry on an awful lot about capacity planning, screw up an awful lot of capacity planning, and then resign in disgrace.

How does that look today?

Richard: I would argue that in the days of CGI-bin, there was also this promise of someone else is taking care of scaling or of running servers for you, which is not very different from what we hear today. Of course fundamentally it's more or less the same. But the thing which in both cases made people go back towards the other extreme, or which will happen with serverless at some point, at least in my opinion, you still need to keep state. Like that's the dirty secret. You have superflume complexity where people just add features because they can and it's just new and cool and they just do whatever, and you have the system inherent complexity and you cannot reduce this. You can put it behind different services, you can have different API's, you can have different service delineations, but this complexity needs to live somewhere. And as a networking person, one of the main complex things is keeping state for long term. And to persist it in a way that you can still access all those pictures of cats or whatever, long term.

None of these questions are answered by serverless, it's just, okay, it's someone else's problem. But then when you scale up more and more and it's all the time someone else's problem, which is super nice, at some point you will probably hit that wall of I need this to be faster. I need this to be more performant. And so you might be tempted to just bring your code and your data and your state more together again. And this is probably something which we'll be seeing in, I don't know, five years? Ten years? But it will happen. That's always the case. Mark my words. People who listen to this 2030, I've been right.

Corey: And we'll probably be having the same debates in 2030 when that happens.

Richard: Of course.

Corey: And it's going to be different terminology, different buzzwords, my Twitter for Pets reference will seem incredibly dated and oh, Google, the same way we talk IBM today. Because nothing is new again. It always seems that history rhymes.

A question I have for you as someone who is building data centers in swamps in 2019, what is the story for data center economics in a world where for most use cases, a Cloud provider is going to have economies of scale that no traditional data center provider will have, they will be able to offer greater elasticity, they will be able to offer armies of people to fix relatively routine issues, that a typical provider would have to be concerned with, what is the case for a data center in these days?

Richard: On a very fundamental level, the case is where do you think the Cloud is running. Look at all the numbers which are being pushed out about capacity. I mean, you can play a bullshit numbers game and talk about how you use two Eiffel towers of steel or 20, which is a pretty arbitrary measure, of course, you can build in concrete or in steel, so you can even change that.

Corey: That was one of the things that surprised me, in the Google Next keynote. I didn't realize when you ordered steel from a supplier, that you ordered it in the units you used were numbers of Eiffel Towers. That was strange to me.

Richard: Yes, it's a totally, like it's best business practice to order steel in just fractions of Eiffel Towers. That's super common.

Corey: I'll take two Eiffel Towers and two thirds of a titanic please.

Richard: Yep. But you'll have to dive for the latter one.

So anyway. Tons and tons and tons of energy are being poured into building data centers, into making data centers more efficient, so this Cloud is running somewhere. So all those big providers also need to have data centers. That's one part of the answer. So while people might forget that data centers exist, and this is totally fine, of course, you also don't daily think about power plants. Yet if you plug into the wall outlet, you have power. It's just something which exists and it's there and just works. And you have this clearly defined service delineation which because power hasn't changed in a few decades, or maybe 100 years. But still you have this thing and you rely on it. This is the definition of infrastructure. People don't think about it and it just works. And if it stops working, they're really really upset and for good reason.

So for smaller providers, building a data center still makes tons of sense. Of course there is tons and tons and tons of industry and of customers who are not able or willing to go into the Cloud just as of right now. It might be that they have certain legal requirements, especially in Germany, a lot of them are a lot harsher than anywhere else in the world, so a lot of external people who need to okay how a company is run, especially when it comes to financial data, or how it comes to health data or something like that, you can't really put this in the Cloud, unless you run that Cloud yourself. Which is also called Hybrid Cloud. Obviously you can squeeze maybe two or three bucks out more if you go all in on the public Cloud but this gives you less control.

So a large part of building data centers and running data centers if you are not one of those huge players these days, it would be those customers who need co-location, who need really top notch service in those data centers, and need them to be up and running 24/7 guaranteed.

So this is the market we are chasing. And to be honest, we see quite some interest. Like, there is huge interest. It might be part of the filter bubble, that you're just not as aware of this especially in the bay area for obviously reasons. But there is a huge market still.

Corey: The challenge that I see is that, when I do leave the bay area, as happens from time to time, it turns out planes do fly everywhere, I find myself talking to a lot of quote unquote traditional companies who are in heavily regulated industries that are making at least partial shifts to Cloud. They're still investing in data centers of course, but that investment is now being made with an eye towards tapering off further and further over the next ten years.

Richard: Ten years is usually, like if you have a medium to large size contract, ten years would probably be a good measure for default contract time. So it makes sense that this is also the length of time people would be talking. I'm not fully convinced this means they will move fully away within the time, it might just be that's their planning horizon so that's how far ahead they can plan and do plan. We'll see what happens. For the foreseeable future, there's definitely no shortage in people who need this, who really, really, rely on this.

Corey: And I think that's part of the challenge that everyone struggles with. One of the things I love about these large Cloud conferences is that we're able to talk to people who have very different use cases from our own. It's always nice to envision a use case we hadn't personally considered, or talk to someone who is building a thing that you didn't realize existed. That's fun. It's always neat to step outside of my Twitter for Pets bubble.

Richard: Yes.

Corey: If people enjoy what you have to say for some unforeseeable reason, where can they hear more of it?

Richard: The best places are probably either my Twitter account @twitchih or any random conference I happen to walk through and give a talk at.

Corey:

Perfect. And we will put up a picture as well, so that people know what you look like so they can stop you at random and share their opinions with you.

Richard: Great.

Corey: Richard "Richie H" Hartmann. Former freenode staff member, current Debian developer, conference organizer, Prometheus core team member, and friend.

I'm Corey Quinn, and this is Screaming in the Cloud.

Speaker: This has been this week's episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com. Or wherever fine snark is sold.

Speaker: This has been a HumblePod production. Stay humble.

View Details

About Jess Frazelle

Jessie Frazelle is a computer programmer who has worked at GitHub, Microsoft, Google, Docker and various companies, startups, even design agencies before that. She’s worked on a lot of the open source projects in the container ecosystem, she’s a top abuser of the GitHub api, and runs her own cloud from her apartment and a colo in NYC called jess cloud.

Links Referenced:

twitter.com/jessfraz

github.com

microsoft.com

google.com

docker.com

contained.af

cncf.io

summerofcode.withgoogle.com

Joe.dev

Soul of a New Machine

github.com/Gazler/githug

Transcript

Speaker: Hello and welcome to Screaming in the Cloud with your host cloud economist, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on this state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey Quinn: This week's episode of Screaming in the Cloud is generously sponsored by Digital Ocean. From where I sit, every cloud platform out there, biases for something, some bias for offering a managed service around every possible need a customer could have, others bias for, "Hey, we heard there's money to be made in the cloud. Maybe give some of that to us."

Digital Ocean from where I sit, biases for simplicity. I've spoken to a number of Digital Ocean customers and they all say the same thing, which distills down to they can get up and running in less than a minute and not have to spend weeks going to cloud school first.

Making things simple and accessible has tremendous value in speeding up your time to market. There's also value in Digital Ocean offering things for a fixed price. You know what this month's bill is going to be, you're not going to have a minor heart issue when the bill comes due and that winds up carrying forward in a number of different ways.

Their services are understandable without spending three months of study first, you don't really have to go stupendously deep just to understand what you're getting into. It's click a button or make an API call and receive a cloud resource.

They also offer very understandable monitoring and alerting. They have a managed database offering. They have an object store, and as of late last year, they offer managed Kubernetes offering that doesn't require a deep understanding of Greek mythology for you to wrap your head around it. For those wondering what I'm talking about, Kubernetes is of course, named after the Greek God of spending money on cloud services.

Lastly, Digital Ocean isn't what I would call small time. There are over 150,000 businesses using them today. Go ahead and give them a try or visit do.co/screaming and it will give you a free hundred dollar credit to try it out. That's do.co/screaming. Thanks again to digital ocean for their support of Screaming in the Cloud.

Corey Quinn: Welcome to Screaming in the Cloud. I'm Cory Quinn. I'm joined this week by Jess Frazelle. Welcome to the show Jesse.

Jess Frazelle: Thanks for having me. Super cool to be here.

Corey Quinn: I'm still vaguely astonished that I'm actually speaking to you. Your one of those. Wow, great luminaries of lights of the space where it just tends to be one of those scenarios where, "Wow, someday I might be cool enough to talk with her." And wow, that happens today. Thank you very much for doing this. Historically, you're a ... well, you define yourself as a computer programmer who's worked at GitHub, Microsoft, Google, Docker, and hey, I filled a bingo card of fun tech companies. Who are you and what do you do?

Jess Frazelle: I guess I am employable by a bunch of different companies. I like to just see their process and then leave. No, I'm kidding. I like to build things and I like solving hard problems and I get bored really easily. If there's not a hard problem to solve, then I'll move on to a different hard problem somewhere else. I mostly am just constantly curious, I guess

Corey Quinn: That's a fantastic problem to have in some ways. Although I remember from my own childhood that tended to get a monkey in a bit of trouble as it went along. You are, I guess, most famous in some respects for being famous on Twitter, which is almost like being famous in real life, except you can go to the convenience store without being worried about Paparazzi. Is that more or less accurate?

Jess Frazelle: Yes, Twitter is so weird because ... I try to be myself on there and I think that I am. Whenever someone says that they follow me on Twitter, I'm like, oh my gosh, I'm so sorry. Because I say a bunch of dumb stuff, all the time. It just comes out. It's just one of those weird things but I don't really know how that happened.

Corey Quinn: Let's talk a little bit about the current zeitgeist of the cloud, for lack of a better term. The big issue right now that everyone is gearing up to fight is containers versus serverless versus a whole bunch of nonsense where everyone's wrong. Where do you stand on that particular religious war?

Jess Frazelle: When serverless came to be a thing I did not take it seriously mostly because of the name and it annoyed me and I'm pedantic and the fact that there are still servers. I think mostly though it's about the user interfaces that are exposed by all these things. With containers, you have to know what you're doing or understand containers.

With serverless, from what I've seen, it depends on the products but some user interfaces are really nice and usable and it seems like developers are really catching on to them, and then others seem to be just like a container user interface. Then there's functions as a service too which is different in the fact that it's just one function versus a container running maybe a service.

Most of all, at the end of the day, all these things that are just running your code, they're running on a server and the back ends of a lot of them, from what I've seen are pretty horrifying and a lot of the cloud just in and of itself and these backends, it's all just popsicle sticks and glue put together. I'm just not sure if it's like the greatest thing at the end of the day.

Corey Quinn: I agree with what you said in the context of with containers everyone has to get deep into a whole bunch of different things and trying to figure out how things wind up working and when this first came out and you could successfully give a 45 minute conference talk that was nothing other than Docker repeating for those 45 minutes, that seemed to me at the time to be something that either I didn't get or that no one else got either.

My default assumption is that I'm wrong so I gave a five minute lightning talk making fun of docker called heresy in the Church of Docker. At the end of it, someone came up to me and said, I had a question about your talk and I expected that someone was going to tell me I was completely wrong across the board. Instead they said, "Could you give the full version of that talk at Container Con?" Wow, you want me to give the full version of my 45 minute talk at Container Con first, wait, there's a full version, and secondly, absolutely.

It started resonating in that there were a bunch of things that you didn't understand ... that I didn't understand that didn't solve problems the way that you would have expected them to in a container world. This was a collective shortcoming. Today that talks dated, it doesn't work anymore. You have a bunch of tooling enhancements and the products themselves have gotten better to the point where almost all of my criticisms are no longer there. I wound up skipping the container revolution as it were and skipped straight ahead to serverless, because one, it seemed easier, two less things to manage is always a good thing and three, when you've only been using something for two weeks, it's super easy to be considered an expert in it when the thing itself is only four weeks old.

Jess Frazelle: Totally. I agree with your talk actually mostly because when it comes to containers they are super complex and at the end of the day, the person just wants to run a damn thing and then they don't want to have to manage it. With containers, you get all this complexity and then if you're running it locally, you have to manage that service. I think that a lot of what these products are solving, is all that pain, which is great. It just irks my inner technical nerd when I see the backend and I'm just like, no, this could have been pretty, but it's actually the worst thing ever, but I don't think that's a lot of things that customers really come to care about.

Corey Quinn: My default assumption never having picked underneath the covers of any of these things is that my environment ... when I build out a data center, is a metaphorical and sometimes literal trash fire. When I run something serverlessly, I imagined that everything underneath the hood on that is pristine. It's deployed seamlessly, it resembles the whiteboard architecture diagrams we all look at and envy. If I were to even peek in the data center, the cabling would be immaculate. Everything would wind up being just more or less heaven but on earth.

Given that you've worked for a number of large providers, that's an accurate assessment. Right?

Jess Frazelle: I used to think that too, and I was like, it's just like this nice church of computing. Then it's just like you go to work there and then it's just saddening. The fact that I found myself a lot of the time saying, well we are the cloud. Shouldn't we be better? I mean, I haven't worked at every single cloud provider, so I don't know about all of them, but I would say a lot of them are popsicle sticks and glue and I don't think they would even hide that fact. They would also agree with me.

Corey Quinn: One of the most amazing things I found is that every time there's been a company I sincerely admire who talks about infrastructure at a conference talk or whatnot, I start talking to people who are involved in helping these things run. I have never found an exception to this, but once you get people talking comfortably and being honest, their immediate response is yeah, that's one aspect of it but here's the list of things that make our environment a fire. I've never yet found an exception to that where a company has an amazing architecture, an amazing environment and everything just works seamlessly. Have you?

Jess Frazelle: No, I mean it's computers at the end of the day. GitHub's infrastructure was really cool from everything that I've seen. They have a lot of like X Digital Ocean and X Heroku folks. I think that they took a lot of lessons learned from their past experience and applied it there. Really everything has problems because it's computers at the end of the day so you're going to come across something weird and every single company has their weak points, and then you find that weak point and then it's like, whoa, yes, that's it.

Corey Quinn: Well to that end, what were you doing for a year at Microsoft? Originally Microsoft to me was an example of a has been company that no one really cares about and losing relevance and then magically they started shifting their entire positioning in the market, their reputation changed, they started hiring a tremendous number of very admirable people. Yes. Including you. What were you doing there?

Jess Frazelle: Mostly I did what I like to call annoying people as a service. Actually, so, Chad Fowler was my skip manager and when he joined I was previously there for a while, but I was like, "I'm so sorry if anybody comes to you and they complain about me." Because mostly what I did was I broke a lot of things and then I tell teams about it and it ends up being, because Microsoft is freaking huge, you're crossing organizational boundaries and you're crossing team boundaries and you're ...

Just like whenever there was a bug, I'd go knock on their door and be like, hey there's this thing it's a problem. You've got to fix it. I ran a bunch of performance tests. I gave feedback to teams internally. I'm a pretty candid person so I don't think a lot about how people are going to handle it. I got a lot of feedback from teams that I lacked boundaries and I was like, "Wait, what are they talking about? It's not like I'm standing in your bubble or something." Apparently there was this organizational boundary that you're not supposed to actually go knock on people's door and be like, "Look this thing, it's bad." That was interesting to come to find out and learn.

Corey Quinn: That is a form of a story from my own life that distills down into how I got fired from a company once where I tended to assume because everyone had the same domain in the end of our email addresses that we were all on the same side and I didn't have time or patience for hierarchy. Instead, I was going to go and talk to people in other groups about what was going on, about shaking things out of the trees before they wound up impacting customers.

It turns out in some cultures that's welcomed and appreciated and expected and in others it very quickly turns into a knock knock who's there, not you anymore story. I think it's dependent upon the company and the culture in question, but there was a time I would have heard you make that statement about not respecting boundaries and thought, "Oh, no such thing. Everyone's on the same boat rowing the same direction." Now I'm not as naive anymore. I don't believe that and I think that's one of the biggest single reasons I'm unemployable.

Jess Frazelle: Yes, I mean I definitely had absolutely no idea before that that was not really taken nicely. Now going into a job at GitHub after, I was almost worried about crossing boundaries but then also I know so many people at GitHub I was like, knock knock. But it was not really the same things that I was doing. I was more like, "Oh, I like this thing." It's just interesting to see how that pans out in different cultures. Some teams were totally okay with it and they took me back as a gift and then others were like ... They were not happy at all.

Corey Quinn: Changing gears slightly where we talk about different teams doing different things. I think there's nowhere that does that quite as well as Amazon where everything is a two pizza team, which either means that each team can be fed with two pizzas and no more or, as I tend to think of it, to be on the team, you have to be able to eat two whole pizzas yourself in one single sitting.

You spoke at re:Invent last year and when I saw that on the schedule, I thought that I was going to have a field day with this because Amazon doesn't normally do misprints. Having you from Microsoft at the time, speaking at their conference is the clearest definition of misprint I can find. It wasn't a misprint. How did that happen?

Jess Frazelle: That talk is a fun story. I loved going to re:Invent. It was really cool to see how their conference is. Yes, I was working at Microsoft at the time and it was before the close of the GitHub acquisition but I was doing a lot of traveling back and forth to San Francisco, helping out on the merge of things. One of the things that we had talked about was making sure that we're still ... GitHub is a large part of the external communities. That means showing up to other conferences of competitors like re:Invent or Google Next.

I was like oh this is great. I know Abby Fuller, she's an amazing person. I'll just reach out and be like, "Hey, can I maybe get a talk there?" And I did the schedule was entirely full and everything. She was like, let me see what I can do. She did, what I do, annoy some people internally. Then she got a slot for the talk and it was amazing. It also had Clare Liguori who is an amazing engineer on the Amazon side.

It was this really cool lady power hour is what it boiled down to but it was the most last minute thing. I also was terrified I was going to be fired for this even though I had the backing of the new leadership at GitHub. I emailed Chad, I am so sorry if people come to you and they are like what the hell is Jess doing? Because it went live while I was still a Microsoft employee before I had joined GitHub and the actual talk was my first day at GitHub.

No one I think said anything and I think that it's a great testament to how GitHub is going to be run at Microsoft. I think that as long as they continue doing these things and making sure that they have a bunch of external community engagement, it's just really good. One of the things that I wanted to make sure about, GitHub being a part of Microsoft is that, for every talk and appearance that we give at a Microsoft event, that we have two external appearances at competitors or other external events or people, because it really is important to the community and everyone watching that they know exactly where GitHub stands and how they put the importance of the community above everything else.

Corey Quinn: I think the idea of having the cloud vendors ... I guess on the one hand, they're absolutely competing with one another for market share, et Cetera. GitHub is a bit of a strange animal in this context in that they tend not to really be directly competing with vendors themselves. GitHub is different in that in some respects it's not competing directly against large cloud providers, but I've always been a bit of an advocate of the idea that right now the big competition is not between different cloud providers so much as it is not going to cloud at all.

I think that more cooperation between some of these large folks means that there's a bigger pie for everyone to get a piece from rather than trying to smash each other into the ground. I'd love to see less hostility between various providers. I think there would be a fantastically different world if that were the case. That's also probably hopelessly naive.

Jess Frazelle: Yes, actually I super agree. I really am of the belief that since a lot of my friends ... I live in New York, a lot of my friends work at financial tech firms or hedge funds and stuff like that. A lot of people that I know do not use the cloud. They have their own data centers. GitHub has their own data center. It is more of a cloud versus not cloud thing to me. I almost feel like there is a market for disrupting the not cloud.

Corey Quinn: So, changing gears slightly, if I go to your Twitter bio page, there's a link there to contained.af and we'll throw a link to that in the show notes. What is that?

Jess Frazelle: That was a site that I made in a day and it ended up being very useful. One of the reasons why I made it was containers are super complex. There's a lot of knobs that you can turn on them to make them either more secure or less secure and either less aware of their host environment or more aware of their host environment. I wanted to show a completely locked down scenario and then also ask people about this environment that they were in to teach them a little bit about the internals of containers.

Another main reason why I did it was that, I had been giving talks about Docker for a long time and a lot of the fear, uncertainty and doubt about containers is that they're insecure and it's more like they're insecure with nuance. You can make them secure if you try hard or they can just be wildly insecure if you just run them in a privileged mode or something like that. I was mostly like, all these people don't really understand the nuance when they come to me and they say this, so I'm going to give them a test.

Whenever people would bring that up to me, I was like, look have you broken out of this thing? Because the site has this terminal where it shoves you into a container and if you break out, you have to capture this flag and then I would be like, "Oh Whoa, you actually found a container escape. No one's done it. It's great. Whenever people come to me and they're like, containers are insecure. I'm like, look, break out of this site but they haven't.

Corey Quinn: I love the idea of having a learning tool that distills down into a ... Here's a thing, play around with it and see if you can wind up getting to X or to Y or to Z. I don't know if I'm weird or this is more common than the market would have you believe but sitting down and reading a textbook or taking a class isn't a great way for me to learn something. Instead, here's a project or a problem you need to solve and learn how this stuff works. That's something that resonates with me and that's why I love things like this.

Jess Frazelle: Yes, for sure. I actually think all the terminal based learning sites for containers are the ones that really helped adoption. No one reads a docs page. I mean it's just a wall of text, but really playing with something is way better.

Corey Quinn: On a somewhat related note, have you ever heard of a ruby gem called Githug? Not hub, hug as in to put your arms around something?

Jess Frazelle: No, that sounds cool.

Corey Quinn: It was one of those things where first, before I knew anything else. It was, wow, I love that name and I wish I'd come up with it, but what it is is somehow even better than that. It's a 40 something level experiment where you run this gem and it winds up from there giving you a challenge when you cd into a directory. Level one, make an initial Git commit, level two, revert the commit, and so on and so forth. It's a step by step. This is what Git is, this is how it works and by the end of it you're doing partial re bases you're using Git bisect. It goes from what is Git all the way to the end of the line where you can do more than most people have to do, career wise, with Git and it only takes a couple of hours to run through, start to finish for most people.

Jess Frazelle: That's really cool. I need to check that out. That sounds awesome.

Corey Quinn: I love that this is happening and one thing that I find revelatory is that when I first learned Git, well when I first learned Git back in 2010, GitHub, sent a trainer on site for two days and all of engineering would sit there, we went up one side, down the other and no disrespect to the trainer in question or GitHub, I was more confused by the end of that training than I was at the beginning.

Now Git seems like something you can get someone up to speed in a number of hours. Is that because Git has gotten that much better or is that because we're better now at explaining it to people from a variety of different ways.

Jess Frazelle: I don't think the Git has gotten better. I think it's just, yes, people are getting better at explaining it and maybe we're only using a specific subset of the features of Git and so that adoption has hit the peak of how to say, how to use it, because the tool itself is still the exact same.

Corey Quinn: Absolutely, there are so many different flags and options and people really only tend to use five or six. You can also start convincingly bluffing people as far as things that don't actually exist. Oh yes, there's a tool that does that just run, Git unchained melody and it'll wind up solving your problem perfectly.

Jess Frazelle: Yes, I'm pretty sure that there's all these tools built on top of Git, like Git wrappers that do absolutely everything. I have 40 bash scripts that probably do random things.

Corey Quinn: Git's a great segway. While we're talking about things that makes everyone sad, let's talk about the cloud native computing foundation and the relationship with open source sustainability. Where do you stand on that?

Jess Frazelle: I am just mostly disappointed in the cloud native computing foundation when it comes to helping open source projects. I think a lot of people agree on that. A lot disagree as well and it seems to be a contentious point. But, from my point of view, I know how much money they have from vendors. I had heard about it through someone who's on the board at some point and I won't repeat the number because I don't think it's supposed to be said, but they have a ton of money and they don't have much impact for having so much money. Especially impact on projects itself.

If you were to look at something like Google Summer of Code that has huge impact, they get interns from all over the world to apply to help out on these projects and then they have a task for a project and then they get it done over the summer. Sometimes they do more than one task. They do a bunch of things or it's just one big task over the course of a few months. That kind of impact is crazy huge because it helps the project and it also helps the person who has the task. They get paid and they get this thing on the resume. They get, public commits showing their work, which helps them get jobs.

The whole thing is this really great way to help people and projects. But CNCF, I cannot say anything that they have that is even immensely close to what Google Summer of Code does. Google Summer of Code does that on a very small percentage of the funds that CNCF has. That's just super sad to me. Then there's also what GitHub is now working on in this space for helping open source sustainability.

They hired Devin Zuegel who lives in San Francisco and she is super awesome, Nat hired her and she has been interviewing maintainers and talking to a bunch of people in the space to understand all the problems. It's going to be super interesting to see what she does but I have 100% faith that the solution that she comes to help with, these problems of helping projects succeed sustainably will be 100% more impactful than anything that the CNCF will come up to with.

Mostly because they don't care. From my standpoint, I don't see any care being put into it and from the standpoint of Devin and Nat and all of GitHub, they deeply care about the community and fixing problems. It will be really interesting to see what happens there. But CNCF, no, I'm not a fan.

Corey Quinn: I don't really have a strong opinion on CNCF because I don't really encounter them in what I do. Now let's unpack that a second. I run the Screaming in the Cloud podcast. Surprise, I'm recording this conversation and I also write a newsletter every week. It rounds up a giant pile of cloud news, mostly around AWS and aggregates that and then publishes a fraction of that. I also have a consulting business where I go into very large cloud environments and help them, not only save money on their bill, but work on governance, work on some cases tooling, work on process flows that makes sense.

Effectively my entire life, in one way or another right now, is cloud and I go back to the beginning of this diatribe of mine. I don't have an opinion on the CNCF because I don't encounter them in the wild hardly ever. That alone tells me that whatever their stated goal is, I don't get the distinct impression that it's aimed at problems that I am dealing with, problems I am experiencing, problems that my clients are focusing on or problems that resonate in the larger community. I don't know how to judge what they're doing because I don't know what they're doing and that in itself is a problem.

Jess Frazelle: Yes, that's a totally different perspective than mine. That's super interesting to me because if they aren't doing anything, from my perspective, the open source community perspective, they aren't doing anything from the cloud perspective, it's like what are they doing? It's very weird.

Corey Quinn: To be very clear. This could be a complete misunderstanding on my part and I will have to wind up back walking everything I just said, conducting apologies, et Cetera. If this is being played right now in a meeting at the CNCF and you have a nuanced and detailed critique of how everything I just said was completely wrong. Terrific. Great. Please reach out. I am thrilled to have that discussion on a future episode of this podcast. Please tell me how I'm wrong. I would love to be wrong. My biggest fear right now is that I'm right.

Jess Frazelle: Yes. I mean if they're listening to this, it's like, "Hi, I'm just your biggest fan. They already know."

Corey Quinn: It tends to be a hard problem as well. I think that GitHub historically has been a fantastic source for the community to focus on solving interesting problems, writing infrastructure for projects that otherwise no one would have ever heard of and doing something like that even as a for profit company has demonstrably had impact.

You mentioned Google Summer of Code program. Every time I've worked with someone as a part of that program, I've come away first profoundly impressed by what they're able to achieve over the course of the summer. Secondly, I find myself keeping in touch with those people over a period of years and watching their careers evolve. It is hands down from everything I can tell, orders of magnitude more beneficial to someone's long term prospects than a degree.

I think the sheer fact that I was a GSOC intern that someone can claim that it automatically merits further scrutiny because people in every environment I've ever seen do not mess around.

Jess Frazelle: Yes, I'm a huge fan of that program. It's really great. I would love to see it emulated in other ways to really help people and just the same amount of impact.

Corey Quinn: Do you have anything you'd like to mention or drive people to pay attention to if they've gotten this far in the episode and haven't hurled the whatever it is they're listening to a aside in a fit of rage because I just insulted something that they love.

Jess Frazelle: Yes, I would say definitely watch this space where GitHub is working on open source sustainability because that will be really interesting to see play out. They also just keeps shipping really awesome features that help the community, which is cool. A book that I recently read also that I got as a recommendation from Brian Cantrell, was called the soul of a new machine and it really was a moving experience for me.

Mostly because the book is about data general when they made a computer and it goes over the team that built the computer and it's just a really great story about a team doing something that seems impossible like putting together an entire machine and how they put parts of themselves in the way that they work and the way that they write documentation or really care about the microcode and stuff like that into the machine. If you're just like a huge nerd and love stuff like that, I would definitely say pick up that book because it's awesome.

Corey Quinn: Terrific. I'll throw a link to that in the show notes. If people want to follow you. You blog at jess.dev which I absolutely want to talk to you about for a minute. You made some noise on Twitter back when the .dev domains opened up, about having spent entirely too much money on a domain. But now it's yours. Tell me that story please.

Jess Frazelle: Yes, this is all Joe Bedas fault. He bought joe.dev and we are just chatting in DMs and he was going ...

Corey Quinn: You do realize that right now he is dancing somewhere because someone just said it's all Joe Bedas fault and for the first time in history they weren't talking about Kubernetes.

Jess Frazelle: Yes, we're just chatting in DMs and he says that he's going to get joe.dev. I was like, "Oh, I'll get jess.dev." I ended up buying mine before his and he was like, "Oh, you just did it, then I'll do it." I was like, oh my God, I didn't realize that he hadn't done it yet, but we bought it on the first day, which made it ridiculously expensive. I'm also unemployed and I didn't just have an acquisition by VM ware.

This was a weird decision but it really goes to show what I value, which is overpriced domains. I gave a few people some domains that are also named Jess and if a Jess is listening to this podcast, you can reach me on Twitter and tell me that you want CNAME redirected to a subdomain. I really like it though. I just think it's a forever domain. It will work as long as Google keeps it up.

Corey Quinn: That is the giant open ended question. It's easy to make fun of Google for turning things off that people care about, but it is hard to imagine that they would wind up turning off something as big .dev. That's a global TLD. That's about as likely as them deprecating their goo.gl URL shortener which they announced was being deprecated last year.

Jess Frazelle: Yes, I mean reader. I still have feelings about reader. It was really great.

Corey Quinn: I think we all do. I'm also going to call it out right now. I'm sorry. I think that registering the .dev TLD could have been a fantastic thing that Google could have done for the community and then automatically routed the public response to local host for every wildcard query against it, but no, now they wind up giving the domains away or selling them or whatever you want to categorize that as and suddenly every company in the world, historically going back 40 years who uses a .dev fake domain for internal testing purposes, has a security problem in the event that something winds up going where they don't expect it to.

Jess Frazelle: Wow, yes, that's super true and actually makes me also want to squat on some of these for that reason. That's interesting.

Corey Quinn: Yes. It's one of those areas where it becomes either extremely lucrative. If you have no sense of personal ethics, it becomes hilarious if you want to make a few very large companies look terrible, but largely it feels like this solves a problem that didn't historically exist. I'd really like to have a separate domain for testing, but you know, we just can't find one that's publicly available, so we're going to make up our own TLD. I'm sure no one will ever turn that into something that resolves on the public internet. It was short sided and it also turns into a story of a large company trying to avoid spending $12 a year on a test domain.

Jess Frazelle: Yes, totally.

Corey Quinn: We've referenced Twitter a couple of times. Who are you on the twitters in the event that people have been trapped under a burning couch for the last 12 years and have no idea where to find you?

Jess Frazelle: I'm @jessfraz, and I am sorry for all of the weird tweets,

Corey Quinn: Frankly, the weird tweets are sometimes the best tweets. Jess, thank you so much for taking the time to speak with me today. I appreciate it.

Jess Frazelle: Yes, thanks for having me. This was awesome.

Corey Quinn: Jess Frazelle, computer programmer to the stars and Twitter famous. I'm Cory Quinn and this is Screaming in the Cloud.

Speaker: This has been this week's episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com or wherever fine snark is sold.

This has been HumblePod production. Stay humble.

View Details

About Maureen Lonergan

Maureen Lonergan joined Amazon Web Services in March of 2012 as Director of Training and Certification. Since then, Maureen has worked to build a set of programs and offerings that offer a flexible path for learners to advance their careers and for organizations to enable their teams and get more out of the cloud. Her team is responsible for building, maintaining, and delivering both classroom and digital training courses alongside an AWS Certification program to validate cloud knowledge. Education programs, including AWS Academy, aim to build a pipeline of cloud talent for the future. Over the course of the last 7 years, the organization has delivered training in over 50 Countries and hundreds of thousands of learners. Prior to Amazon, Maureen was the Senior Director for Partner Enablement at VMware where she built training and enablement programs and delivered training to hundreds of thousands of individuals across a channel of 30,000 partners. She’s also served as the Director of Technical Training and Enablement at Symantec and the Director of Education Services at Ariba.

Some of the highlights of the show include:

  • Where to get started learning about the cloud
  • The variety of AWS certifications offered
  • Why certifications are valuable for job prospects
  • The work that goes into designing the AWS training courses
  • Some partners where you can access training

Links:

  • https://www.aws.training
  • https://aws.amazon.com
  • https://aws.amazon.com/training/course-descriptions/
  • https://aws.amazon.com/training/learning-paths/
  • https://www.coursera.org/aws
  • https://aws.amazon.com/training/path-cloudpractitioner/
  • https://www.cisco.com/c/en/us/training-events/training-certifications/certifications/expert/ccie-routing-switching.html
  • https://www.edx.org/school/aws
  • https://www.coursera.org/aws

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host, Cloud economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize.
This is Screaming in the Cloud.

Corey: This episode is sponsored by Scaylr because all kids hate their logs… doo do do doo do doot.

When your site doesn't go

Or maybe it's slow

And people can't load up your blog

Where do you start

what's the state of the art

With logs logs logs

Logs, logs

Full of repetitive noise

Logs, logs,

Awk and grep? Sorry they're toys

Everyone hates the logs

Nobody can read their logs

Improve the state of your logs

Scalyr can help with your logs

logs logs logs

Logs. From Scalyr dot com

Everyone wants a log

You're gonna love it, log

Come on and get your log

Everyone needs a log

log log log

Logs! from Scaylr.com.

Corey: Welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined this week by Maureen Lonergan, AWS director of training and certification. Welcome to the show.

Maureen: Thanks for having me.

Corey: No, thanks for taking the time. Let's start at the very beginning. What is it you would say it is you do here?

Maureen: I would say that my team and I build training and enablement programs for our customers on AWS technology.

Corey: One of the recurring themes on this podcast has been how to take someone who has either not touched cloud before but has experienced in a data center or traditional IT operations or someone who is a new graduate or new to the space entirely, and get them to a point of being productive, employable or otherwise competent to begin touching people's, in some cases, hideously expensive production environments without causing huge amounts of risks. When you wind up taking a look at someone brand new, where do they start?

Maureen: We've spent a lot of time thinking about this in building programs over the last couple years. We started building instructor-led courses much like any other IT provider who are doing the industry but over the last couple of years, we've really spent a time defining the personas that are out there.
And so modernizing IT skill sets is super important to us. Our general enterprises are trying to move their traditional storage administrators or database administrators and need programs to do that. So most recently, we've developed, we launched a digital platform that has 350 courses that are free and available to anybody who wants to take them but we specifically focused on how do you build the foundational level skill sets for someone in tech or a business leader.
And we launched the cloud practitioner class last year which is actually our fastest growing course out in our portfolio. We also aligned that with a certification for cloud practitioner. Anybody, a student or an individual within an organization, could go to the platform and take the course and work in the platform and take the certification exam.
From an academic perspective, we're doing the same thing. We launched the academy program a couple of years ago. We're working with hundreds of universities across the globe. And we have both foundational level training with cloud practitioner. And we also have a cloud curriculum that we're delivering alongside with certifications so that when students come out of the university, they actually have defined validated skills.

Corey: Somewhat common refrain in this space has been that either certifications are incredibly valuable or certifications have no value whatsoever. And that's a very broad spectrum that's easy to distill down into a binary which I think is absolutely the wrong approach. When does getting a certification for someone make sense?

Maureen: I think it's actually a personal decision. Certifications are geared towards the individual but what I would say is that we spend a lot of time working with our customers and there is a huge cloud skills gap. And in meeting with our customers, we're talking about how do you find a talent out in the industry.
And certifications is one of the things that we asked them to look for. Here's a series of work experiences or education that we think will help you on your cloud migrations. But certifications is the one thing that they can look to that we know that we've spent a lot of time building, validating and certifying. So I think, again, it's a personal thing but I think it's important if you're looking for a job in cloud.

Corey: Once upon a time when I started with cloud, I logged into AWS and its consul for the first time, and I was taken aback that there were so many services that I was never going to be able to wrap my head around. There were 12. There is now over 150 the last time I counted. And the challenge that I had then was there were really no resources for getting started other than the documentation, which in that era was not what it is today.
Now when I log on to the training page and look at how to get started, one of the challenges I see is almost an echo of that previous problem. It's not that there aren't any training options now. It's that there are so many. There are a bunch of native offerings that AWS provides. You have a number of partnership agreements with a number of training schools, and there are multiple different paths to get there.
And the documentation is, of course, still there but now if printed out, it would be three times a size of any encyclopedia, which for the younger listeners out there used to be a series of books that was a facsimile of Wikipedia but smaller. Where does someone start? It can be an overwhelming experience when you just now learned that Amazon is more than a store where you can buy things. It also does this weird thing in the world of computers. How do you start?

Maureen: I would recommend that you go out to the AWS.training site and take a look, sign up for the free tier of digital offerings that we have. One of the first pages that you'll land on is cloud practitioner, the foundational level learning. It's six hours of content broken up in 10-minute chunks and it really gives you base level foundation for what cloud is.
I think after you've taken that training, we've purposely designed because of the evolution of our services and the rapid updates to them. We have designed 10 to 15-minute modules across all of our services from foundational level all the way up to 300 or 400 level.
I think we have this leadership principle at Amazon, learn and be curious. And we live and breathe it every day and we're constantly thinking like how do people need to find training, what is this training that they need. If there's a new service launch on Lambda, let's make sure that we get that training out there as soon as we can after the announcement and make it available on the most consumable way.

Corey: In the interest of full disclosure, you at last count, had nine different certification options? I have one of them, the cloud practitioner. And the reason I did that was probably the worst reason in the world, which is at re:Invent, I wanted to get access to the certification lounge which frankly, I would recommend doing if you're curious.
But going through that process was interesting in that it assumed of relative baseline level of about six months of experience, I think it was asking for, of AWS concepts. I am very much not the target market given that I have roughly 20 times that. So I'm not here to say that, "Oh, that cert was easy." It's not easy for everyone, and it was relatively straightforward just based upon my experience level.
But what was fascinating to me about that was the way that it focused on how the pieces fit together, what each service did, what it was envisioned to be able to do. It was perfectly aimed at business leaders. In other words, folks who are never going to make API call themselves. They're never going to build anything from scratch but they need to be able to take what their engineering groups tell them about AWS and contextualize that in the context of what these services do.
I think that was a terrific direction to go in, and of all the services you offer, it's probably the one that I'm the most excited about in a professional sense. One of the things that I find strange about that though is cloud practitioner is more or less presented in sort of its own thing. It's not generally listed on the path to getting further certification either associate, professional or specialty levels.
It sort of is in its own little island, until recently there was a requirement of associate certs before you could challenge professional and specialty certs. But even then, cloud practitioner was not included. Is it just me or is the cloud practitioner certification aimed at a different audience than the rest of the certifications?

Maureen: I think we specifically looked in the industry and talked to our customers and talked to universities. And there is a tremendous gap in information on what cloud is and how it can solve business problems. We took a step back and said the associate level certifications are very specifically geared toward the technical audiences whether you take the developer or architecture or operations. But there was a huge desire for anyone from a line of business leader to a C level executive to really understand cloud and be able to clearly and effectively articulate what the business value was and how they would leverage it to solve business problems.
We spent a lot of time with our customers. We defined the personas and built the exam. And this has actually been a really good certification for technical individuals that are trying to modernize their skill sets too. There's a lot of fear out there in the industry about moving to cloud, how are my skills going to be relevant. I've been a database administrator you know what, 30 years, how do I know get comfortable? And I think that this has been a great ... It's actually been one of our fastest growing certifications ever.
We're also leveraging it in the academic market. We see more and more people starting to do certifications at the university level and we want to make sure that we're building the workforce for the future and we believe that that's a great onboarding mechanism to other technical paths.

Corey: Do you think that there's room for further growth in the certification offerings?

Maureen: We analyze and assess it a lot you'll see at re:Invent this year, we launched the ML certification. We've seen tremendous uptake in that. We also launched some training paths along with that and we're starting to explore other offerings. We've been very specific to design around role-based learning paths and specialty areas that we thought that our customers needed and needed to identify people.
Most recently, we just launched at CES, the Alexa beta certifications so we're super excited. A little different than what we've done before. So we're constantly evaluating what we think is needed for our customers.

Corey: Right now in a number of different places, there's a bit of pride about people who have been able to take and maintain all nine of the certifications you currently offer. First, is that something that is generally a good idea from your perspective?

Maureen: It's funny. I think it's a competitive thing. I see it at the events that I go to and in the lounge. And I love it because I think it creates a community, but I actually think that that's a very personal decision. We see enterprises going in and starting to embed certification requirements and the kind of the roles as they defined them but I've yet to see one that says, "Do all nine."
Again, I think it's a little bit of a competitive thing but I think, as you grow your career and you may start out with one certification and want to grow up the stack and then specialize, so that's why we designed it that way.

Corey: I can see a future 5 or 10 years out where someone shows up for a job interview, and the interviewer looks of their resume and says, "Oh, I see you have all 28 AWS certifications. While you're obviously very skilled at taking tests and understanding how AWS works, we're hiring someone to do a job not take tests all day. Thanks, bye."
And you see that in some cases in previous generations with different certifications as companies continue to expand their certification practices. At some point, it almost becomes counterproductive to spend all of your time chasing various certs because you look at what these people do, there's no specialization there. They are taking a giant pile of certifications that they can pass even if they're all in one arena, and it doesn't seem to wind up leading to anything and building any narrative.
I don't think that that's currently the case with AWS, but I can see a not too distant future where it becomes that.

Maureen: Yeah, I guess ... Again, I think it's a competitive thing in an individual, and I think it's up to the company to determine where they want their employees to spend their time. I would agree with you. I think it's super important to go deep and really understand a specific area and then expand your knowledge.

Corey: One of the interesting things that I've noticed in years past or previous generations, whatever you want to call it, is once upon a time the CCIE, it was the top-tier Cisco certification. It was widely viewed as the doctorate of networking. If you had that certification, it was understood that you could walk on and take a six-figure job in a whole bunch of companies.
And people clued into this relatively quickly and then a bunch of boot camps and brain dump sites sprung up overnight and started teaching to the test with the natural result that that certification went from something that was widely revered to just another cert on a resume.
And this is a problem that I don't think is specific to any one vendor. It becomes a systemic problem once a certification becomes the victim of its own success. There are brain dump sites that will send people in to take a certification. As soon as they come out, they will effectively short-term memory dump everything that was on the test.
And now instead of learning the concepts, you wind up with a teaching-to-the-test style of training, if you can even call it that, where you have all the right trivia answers but no real understanding. How do you view that and how do you combat that assuming, as I can probably safely assume, that that's not the intended goal of the certification program?

Maureen: Yeah, I agree with you. I think it is a problem in the industry and it's not unique to any one vendor out there. What I can say is we, at AWS, value certification. We spend a tremendous amount of time building the exams and going through the psychometric analysis process and building question banks that are large so that we can start to evaluate. You can tell we watch trends and you can tell when items have been compromised. And we make sure that we rotate the questions in order to stop that. We also have a lot of rigor in the way that we build exams and we update them frequently and announce that to our customers.
I think the other thing too is that we're very specific. We want our customers or any individual taken certification to be well-trained. That's the important thing, so we build training programs. You'll see from our test prep, we don't teach the test. We teach how to prepare for the exam. And so it's something that we look at all the time and we're in constantly evolved in how we build our exams and what we test on and how we take that forward. But it is definitely something that we have a lot of rigor on as an organization.

Corey: It always felt like a bit of a silly thing from my perspective where you're cheating yourself. There's the opportunity to learn something right and take a foundational piece of knowledgeable that will likely serve you well throughout your career, or you can grab it onto short-term memory, go in, vomit it back on a test, maybe pass, maybe not, keep trying until you do just by random chance.
But that doesn't build toward anything. It's more or less getting the credential for the sake of the credential rather than the sake of learning. And there's an entire argument you could have around that approach but it always seemed to me that if you're going to learn something, learn it. Don't just fake it.

Maureen: Yeah, I don't think it does you any good and it certainly doesn't do organizations well either. They're looking for well-skilled people and they're trying to solve business problems, and cloud allows them to innovate in a rapid way. And I think from a career perspective, having five certifications for the sake of having five certifications probably isn't the right approach.
Companies are looking for well-skilled individuals, so it's a personal choice but from my perspective, I agree with you. I think people really need to understand and learn the technology and that will only help them in their careers long term.

Corey: I want to say four levels of certification now. You have cloud practitioner. You have the associate. You have the professional and then you have the specialty. Can you break those down for me?

Maureen: Yeah. We don't really think of them in terms of levels other than, I guess, the role based ones, the associate and the professional. What we've tried to do is say, "Here are some certifications. And you can choose to go up a path from a role perspective. You can choose just get foundational level knowledge or you can specialize."
We've actually just, in the last couple of months, released the requirement that you have to go from one certification to another. And that was largely based on the industry. There's a lot more talent out in the industry and people can learn in a lot of ways. You don't have to go to training to be well-versed and take certification. There's many ways to get there.
And we wanted to make sure that we weren't putting an artificial barrier, like forcing people to take an exam that they're already skilled in. That doesn't help anyone. And so I think that big change to our program, we've actually started to see a lot more people invest in other areas and taking their certification. So, I think it was the right decision for us to do.

Corey: It used to be that I could look at the certifications that you offered and say, "Oh, yeah, I could walk in and take all of those in an afternoon." And it turned out first almost certainly that is untrue. And then you added a bunch more and there are specialties that I have absolutely no exposure to. Machine learning or big data, those from my perspective are the best kinds of problems namely someone else's because I have no idea what I'm doing in those spaces.
At this point, learning enough to be able to intelligently speak to all of the different specialties feels like, at that point, you are capable of doing three different jobs all at the same time. I can't fathom what that looks like. I have nothing but respect for people who can walk in and take all nine but I can't imagine the amount of work that has to go into being able to do that.

Maureen: The program wasn't designed to do that. And so I know that there's probably people that will go out and try and attempt it. What I would say is that what we've tried to do is design courseware and certifications for people that are specializing in those areas. Let me give you an example. For ML, we announced the ML learning paths on ... At re:Invent this year, we have five of them. And they are based on Amazon's Machine Learning University internally and all the best practices for building out machine learning.
And we've taken all that knowledge and built free online courses that are aligned to that certification. Someone who specifically and I will tell you right now that content is very challenging and it was designed specifically to build the skills that they need in order to do that along with project work and best practices to prepare you for the exam. But like in order to get all the specialties, I think it's going to be a challenge.

Corey: And it feels like that's one of those challenges that can only get harder with time. Here's a near and dear to my heart challenge. In case people hadn't noticed, AWS doesn't generally tend to leave services alone. They become more capable over time. They get better. Things that required massive workarounds last week now are a click of a button away or a single cloud formation stanza.
How do you find that interacts with someone taking a certification where they're keeping up-to-date with what AWS is doing and now they're faced with a question that where months ago had a workaround challenge that would have made a solution impractical and now is a native feature offering?

Maureen: We try and address that changes to the technology as rapidly as we can. What I would encourage people to do is really look at the blueprint that's designed in the prep workshops that are associated with the certification exam. We try and update things as quickly as possible. We actually review our content on a monthly cadence to see and to make sure that things aren't too out of date and make decisions on updated exams based on them.

Corey: It's a hard problem to solve for. And I don't want anyone to think that I am, I guess, casting aspersions on AWS. Getting those service to a launch to general availability where people can start using it, invariably, those initial launches are almost parities or prototypes of what they will eventually grow into. There have been a number of services that launched that were, to be very direct, clown-shoes awful originally. I'm thinking of the very early days of EC2, for example, where entire businesses were started that made working with an EC2 instance comprehensible.
Today, that is not a viable business. It turns out the service has matured to the point where almost anyone can get up and running with very little background information. So, I wouldn't think that it would make sense to delay a release until all of the certifications have been updated and staged. It sounds like one of those trailing functions and I don't see a way for it ever not to be.

Maureen: What I would say as it relates to services, the rapid development of it, one of the things that we did after re:Invent is after ... First of all, it's the first time that when Andy announced all the services available, we actually pushed all the training 15 minutes later. Those are first time we've ever been able to do that and that's our tight integration with the engineering teams now is a big win for us.
We follow that up with a hackathon with our entire curriculum development organization for two days after that event where everybody sat and updated all of the courseware and the services. We're putting mechanisms in place to try and address the information people. Announcements are made and they want to have the training as soon as they can and we're making our best efforts but it's a rapid process. I think we'll continue to build rigor in that area both curriculum and certification.

Corey: We've talked an awful lot about the AWS native training options, but you partner with a ... I don't think large is even a strong enough word ... with a stupendous number of companies that have a giant smorgasbord of training options, in person, automated systems, video courses, et cetera, et cetera, et cetera. How does AWS view partnerships in the training space?

Maureen: We value our partners in a big way. Early on, we looked at more traditional instructor-led training partners. We have a network of more than 65 companies across the globe that we've trained their trainers who would train hundreds of their trainers to be able to deliver our curriculum. So, that was our first out of the gate. We want to make sure we could scale, provide training in local language in countries where we weren't operating. It's super, super important. And that authorized training network has been very successful and super important to our growth and to our customers.
Most recently, you've probably seen some announcements around. We're looking at partnerships with certain organizations. There are companies like edX and Coursera that have a very different audience demographic but tens of millions of users. And so we've started to partner with them. We put courses out on edX. We put courses out on Coursera and we continue to always look for the right way to reach our customers, new customers, people interested in learning in cloud.
And so it's super important part of our strategy and I would say too, the things with Coursera and with edX, that's again, it's all free training and available to anybody who wants to take it. To the extent that we can get as much information out there in the industry as possible, we're doing that as aggressively as we can.

Corey: I guess in conclusion, if you're able to talk to someone who is just starting out, they're just realizing Amazon might be more than a bookstore, where would they start? Where would you recommend that people dive in?

Maureen: I would say to start with the free tier of our digital platform, of our digital courseware. I think, it's a great way to kind of explore and learn and learn about AWS, obviously leveraging the AWS free tier in combination of those two things.
I think, again back to the learn and be curious, get out there and explore what we have. And if you happen to be a user of edX or Coursera is great programs out there that we put out there as well. I'd encourage anybody who is curious about cloud to do that.

Corey: Last question, Maureen, then I'll stop taking up your time, how do you tend to view community engagement as far as the training and certification program goes?

Maureen: I think community engagement is important. You see that at our events, whether you go into the Lounge or you're just in the self-paced labs or rooms that people are getting together and they're talking about cloud technology, AWS cloud technology. And I think that's super important. We have a lot of people participating in meet-ups. I think anything that you can do to get involved in those communities is important.

Corey: Maureen, thank you so much for taking the time to speak with me today. I appreciate it.

Maureen: Thanks for the opportunity. I was really happy to talk about the programs.

Corey: You're certainly building an amazing thing here. I'm curious to see where it goes next. Maureen Lonergan, director of training and certification at AWS. I'm Corey Quinn and this is Screaming in the Cloud.

Announcer: This has been this week's episode of Screaming in the Cloud. You can also find more of Corey at ScreamingintheCloud.com or wherever fine snark is sold.

View Details

Some of the highlights of the show include:

  • The benefits of RoboMaker in code deployment
  • How cloud computation frees up local resources
  • Using machine learning to improve robot reaction
  • How great a name RoboMaker is
  • Amazon’s commitment to the enduring API

Links:

  • https://aws.amazon.com/robomaker/
  • http://www.ros.org
  • https://aws.amazon.com/deepracer/

Transcript

Announcer: Hello, and welcome to Screaming in the Cloud with your host cloud economist Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud. Thoughtful commentary on the state of the technical world and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey Quinn: This episode of Screaming in the Cloud is sponsored by N2WS. There are a number of backup solutions that are available in AWS including the recently announced AWS Backup. Well, AWS, back the #$*% up. Backups are incredibly easy. Restores, however are absolutely not. You want to find out whether your backups worked well in advance instead of the way that most of us do: when they don’t work quite right immediately after you really, really, really needed them to work correctly. Check them out at n2ws.com. That’s n2ws.com. Thanks to them for supporting this ridiculous podcast.

Corey Quinn: Welcome to Screaming in the Cloud I'm Corey Quinn. I'm joined by Roger Barga, General Manager of AWS Robotics. Roger, welcome to the show.

Roger Barga: Thank you.

Corey Quinn: So, starting at the very beginning what would you say that RoboMaker does exactly?

Roger Barga: So, RoboMaker tries to remove all the undifferentiated heavy lifting that a robotics application developer has to do from the moment they try to start their project with multiple team members making sure everybody has the exact same development environment offering them a good run time to actually run on their robot. And then, offering them simulation as a way to test their application in 3D or 2D environments. And also, complimenting the software with cloud services powered by AWS. Which I believe the cloud is going to be one of the most powerful resources a robotics application developer has access to. But it also allows these developers once they've built their application, tested it through simulation, maybe running hundreds of simulations to test their robot in different environments the ability to publish their application to a robot anywhere in the world and manage hundreds or thousands of robots in a fleet. So, it really provides end to end application development support for robotics.

Corey Quinn: Where does a service like that come from? I guess what sort of pain did you see in the industry that made you decide yeah this is a service that we should bring to market? I'm trying to imagine a conversation that ends with, "You know what would really help this problem? That's right, a whole bunch of robots." Now, in my business I don't have any of those needs. I also live in fairly gentrified San Francisco where giant piles of robots solve remarkably few problems in my life and introduce a whole bunch more. Obviously I am probably not the target market for this. Who is?

Roger Barga: Yeah. So, there's a number of innovative companies that are right now exploring what can be done with automation and robotics. And we view robotics as a very general term, it's anything that can sense, compute, and take action. So, a Coke machine can be a robot, a dishwasher can be a robot, the kivas that are running around in our fulfillment centers. And so, we started talking to many of these startups. Many of the developers within Amazon Robotics which are building and deploying robots, so we said, "Where do you spend your time? What's tough about this? What are the hard parts that really don't add value to your robot?" And this is how we started to understand the product definition for RoboMaker. Because what we saw is these developers spend 80, 90% of their time on tasks that add absolutely no value, unique value to what the robot they're trying to build will actually do. They have to set up dev environment, set up simulation, manage the machines associated with that, very clunky mechanisms to update their robots let alone manage them once they're put into production. And this is actually what became the product definition for AWS RoboMaker.

Corey Quinn: All of which makes sense but how does the cloud enter into this?

Roger Barga: Yeah. So, one of the things if you look at actually where a robot spends its compute power and what resources it has available to it you quickly find that most of the power that's spent on a robot is computing functions locally on the robot which in fact could be shifted to the cloud allowing more power to be utilized by the robot for movement and interacting with its environment. When you step back even further and think about how could the cloud be used to capture data about all the robots in production. Did it detect trends? How could the cloud be used to coordinate the activity and orchestrate the activity of a group of robots in your house or a fulfillment center? Then you start to see the value that a cloud can bring not only to an individual robot but someone who's actually trying to build a business and optimize a business with a collection of robots.

Corey Quinn: I am in now way shape or form a roboticist but to my naïve view I would imagine that if you have a robot as I guess society generally conceptualizes a robot the fact that you have motors that provide locomotion, potentially there is a vacuum on it, maybe it has a bunch of articulated arms that do different things. It seems to me that the power requirements to power those engines are almost on a different order of magnitude then what it takes to power a CPU, or a disk of ram. So, from where I sit it seems like having compute on device isn't really something that moves the needle in any meaningful way. Is that naïve of me?

Roger Barga: It is. It turns out basically for a lot of robots we actually looked at over 50% of their power was actually processing imagery coming in through the cameras, processing data coming in from the sensors, especially for things like that navigation or creating slam maps which are computationally intensive. When we could stream the data coming off of a lidar or a camera to the cloud, do the computational intensive mapping, object recognition, and route planning up in the cloud and push these simple instructions down to these motors. You can save over half of the battery power, let alone the fact that a developer who's trying to build an affordable robot does not have to put expensive compute power. And I'll talk about a customer who we've been working with puts very affordable low power chips on the robot because they can actually push that compute capability and spread that cost of what they're paying for in the cloud across thousands of robots not putting it on each and every robot. It brings the total cost down and this is really important because we see a world where there could be hundreds of robots we get to work with throughout our house, throughout our businesses, and that cost has to be low for them to provide value for the company that's running them.

Corey Quinn: When you're talking about a commercialized robot, something that a company generally tends to sell for a fixed fee and then it has capabilities that may be cloud empowered, do you find that the economic story winds up changing as a result? Instead of a fixed bill of materials for a robot that's out the door now you have effectively cloud services which are historically pay for use which means that the life cycle and how long something's going to exist does over time have a different economic model then existed previously. And do you find customers are okay with that?

Roger Barga: Indeed. We found this actually in cloud computing in general where customers can amortize the cost and the investment of a piece of software not on a per robot basis but the actual usage and they find the economics actually work out in their favor. Not to mention the fact that they're sharing information. If you just think about robots navigating throughout your house or fulfillment center each one of those has information about its local environment which it can share up to the cloud. If another robot needs to enter that part of the warehouse it no longer needs to spend the compute power to understand what the map looks like. It can simply borrow from one of its neighbors who's been there previously and utilize that map saving compute power for everybody.

So again, it's that information only about sharing the compute power, but it's sharing the information that each one is getting. And it's also exciting when we think about machine learning. Let's say we put a machine learning model and install it on the robot for navigation and it bumps into a wall. It can actually send that little information, that one or two errors that it's going to make up to the cloud and if the customer has hundreds or thousands of these robots each one making one or two mistakes they now have a large corpus they can use to retrain the model and push a more intelligent model back down to the robot the next day.

Corey Quinn: So, RoboMaker effectively empowers the hard/interesting parts of building what most of us think of as a robot not the actual assembly line pieces of constructing things in hardware?

Roger Barga: That is correct. And in fact I'd even argue when you look at the hard problems that roboticists have to solve they've got really hard problems that they're trying to solve and some of it's actually their pioneering. So, to actually ask them to spend 90% of their time doing this undifferentiated heavy lifting is really slowing innovation in this field. We can actually give them that time back so they can do that innovative task, build that custom hardware that's going to make their robot special, and take advantage of the ecosystem and services that we're providing.

Corey Quinn: For some of us making fun of various AWS service names has almost become a sport and I take a look at RoboMaker and what it does and it is a shining example of an awesome name. It's very descriptive, it's catchy, it fits in a single Tweet which in some cases is hard to get to. It's so well named I almost have to assume that someone fought against it when it was first proposed. Was this the first name for the service you were considering or did you have a more contentious discussion?

Roger Barga: Naming is taken very seriously here at AWS. Names are important, it will shape how a customer thinks about a service, it can shape basically what people think they can do with a service, and I have to admit I'm really bad at naming. I'm a very pragmatic individual. I came forward with some very simple names for the service and our leaders, we stepped back as a group, and your five ideas actually turn into a list of 500 and you get to think about the merits and see what your peers think about them from their experience. It turns into a journey but also a heck of a large number of meetings. When it's done you can actually look back and go, "Yeah that's a great name. Why didn't I think of that in the first place?"

Corey Quinn: Speaking from personal experience it is many orders of magnitude easier to make fun of a name then it is to come up with a good one. Naming is an art and as much fun as I have with tearing down the very hard work of other people in that context it's hard. There is no great way to get there. Changing gears slightly, there's been a lot of talk about ROS, or ROS, or however it's pronounced, ROS. I read, I don't speak. What is that?

Roger Barga: ROS, yes. You know researchers 10, 12 years ago realized that research and robotics was being slowed down by the very problem that I described for industrial applications of robots that they too are repeating the same amount of undifferentiated heavy lifting to build a robot so they could actually publish their thesis and do that last little bit of interesting work. So, the community got together and said, "Let's actually build an open source academic runtime for robots". It's not an operating system, it's a message passing relay bus. Think about our sensor that actually senses something and it puts what it senses on a message bus with a topic. And again, robot sense, computes and acts where there's a compute node that needs to process information from that sensor. It subscribes to that topic, does the processing.

If it needs to move a motor it puts a message down on the line with another topic. And what's happened over the years is that academic institutions, researchers, have been contributing to this ecosystem of ROS packages for different actuators, different sensors, different computational tasks like navigation, and it's been picked up by industry as well. There's thousands of companies that are using ROS today for commercial applications of robots. And what's been happening over the last year is industry's been saying, "Let's advance this from a research platform to an industrial strength open source platform for robotics where the code has been verified, tested, hardened, performance has been improved". And that's the effort part called ROS 2 which we're proud to be part of.

Corey Quinn: And this is much larger than Amazon or AWS itself. This is a community or industry wide effort.

Roger Barga: Yes. Not unlike what happened in Linux. A number of companies have stepped up and say, "We have a vested interest in making the best industrial strength run time for operating systems that's open source, community supported, and community driven". And each one of the companies in the technical steering committee for ROS 2 of which Amazon was one of the founding members are actively contributing source code, designs, code reviews, and reaching out to the broader open source community of startups asking for their feedback, their input as well to define and build ROS 2.

Corey Quinn: Which brings us to the topic of open source which has a bunch of different directions it can go in but let's start at the beginning. What are you doing that is open source? You said you are part of a larger effort. How does that manifest?

Roger Barga: So, we as part of our membership of the technical steering committee and just what we feel is our responsibility to the open source effort behind ROS 2, we have engineers who are actively contributing to ROS 2 code. In fact, the most recent release of ROS 2, Crystal which just happened back in December, roughly 40% of the code that was actually contributed in the designs was through my team and the contributions through my team. We have a number of engineers which are actually not only contributing source code to ROS 2 but reviewing designs and helping support the community through forums, playing a very active role. And this is what we expect every company in the ROS 2 technical steering committee to do as well because this together is how we're going to make this a successful open source platform for robotics.

Corey Quinn: What are you doing that's different as far as the open source world goes these days?

Roger Barga: Yeah. So, Amazon and AWS have played very active roles in open source software and in not all cases have we been so visibly active and vocal. In this case we are. We actually came out as a public member of the technical steering committee. We're helping actively drive discussions getting feedback from the community and doing so in a very visible manner. And we think this is important because if you're startup or another company thinking about investing in ROS you need to see the names of the companies who are standing behind it and hear about the contributions that they're making so, "Can we trust our business with this?" What's been fantastic since the announcement of both the launch of RoboMaker but also the ROS 2 initiative and the companies specifically Amazon behind it, a number of robotics companies have approached us and said, "We're now moving to ROS 2. We see that Amazon's behind it. We see the RoboMaker support and this is going to be community supported and led. Let's go ahead and port our robots over to ROS 2."

Corey Quinn: One thing that strikes me as a bit of a strange tangent off to the idea of running robots that are cloud connected is as wonderful as the idea is of offloading all of the compute, all of the different heavy lifting that they wind up doing to a third party provider that's somewhere in a data center far far away. What about latency/safety sensitive concerns here? I mean the easy example is an autonomous vehicle which is a whole separate kettle of wax. But we're waiting on an API error and there's a time out and we're doing exponential back off. Meanwhile you are hurtling toward the bay and wanting not to drive into that. Do you find that there are use cases for which there is no substitute for on device computation or is there a fallback mode? How do you envision this?

Roger Barga: Absolutely and we can talk about what customers have done including our own robots in our own fulfillment center and we can also again we're very much ... We also follow trends and we anticipate when we see trends that are unfolding. 5G is unfolding which will give us ubiquitous connectivity devices around the world with low latency. So yeah, we see connectivity increasing, latencies falling, but let's talk about where we are today. Where if you wanted to have a highly interactive session with your robot and you could not afford that latency RoboMaker allows you to actually deploy code including machine learning models onto the robot itself for low latency interaction. And then, you can partition the work where if I can actually handle high latency, if I want to tell my robot to go actually get me something out of the refrigerator I'm okay if it takes a 100 milliseconds latency for that command to actually get up to the cloud, translate it to commands, and back to my robot.

So, good engineers will actually partition the work that needs to happen on the robot which is that when what can be actually offloaded to the cloud. And again, this is that interesting programming the edge, and what role will the cloud play, and how do you partition your work. Which makes this such an interesting problem. And again, companies that need safety critical processing will put that on board, on the robot, prioritize those messages, prioritize that processing, and actually offload other processing capabilities to the cloud.

Corey Quinn: When this service was announced at midnight madness at re:invent last year it was fascinating in that it was sort of out there, not directly tied to other services. And what's also neat about this is that it was announced as being generally available. This was not a in preview, coming later, apply for a thing, it was there ready to go that evening because what you absolutely want people to do is building industrial robots at 2:00 in the morning the day it's released 'cause that could not possibly go wrong. I've never known Amazon to release a service that did not already have active customers using it. You're not generally a company that says, "Hey we built this thing. We're super proud of it. Maybe someone will use this." You aren't the throw a bunch of stuff at the wall and see what sticks company. You have customers actively using this on launch day. Do you have any you can talk about?

Roger Barga: I do. And in fact even from the very inception of the project because we do work customer backwards we reached out to customers that are running robots in production as soon as we had our PR FAQ written to get their input, to get their guidance, and prioritization of features, and really dive deep with them in actual use cases. We later onboarded those customers into an advisory board and actually a beta program. So, we're working with customers months before a actual launch including working with customers who are ready to go into production. We worked with NASA JPL to port their open source rover to RoboMaker and to ROS.

One of my favorite customers because of the nature of the robot and how they're using the cloud is Lea by Robot Care Systems. Lea is a walker robot for the elderly or the disabled. It's not something you would think of as a robot but it's running ROS, it computes, it senses, it takes action. And in the case of Lea they actually have added our cloud services for Polly and for Lex so that the customer can actually call the walker to them from across the room with their voice. Lea will respond, come to the patient, interact with the patient in the most natural manner through voice, but they're also streaming telemetry off of the walker through kinesis data services. So, they can actually understand the gate of the patient, their walking rate, how much activity they've done. Doctors can have dashboards that monitor their patient. If they feel the patient's recovering they can actually build predictive models with that data and actually predict when the patient is going to recover or detect if there's a negative trend and they need to actually interject and take action. It's changed their business, it's changed the value prop they offer to their customers, and it's a great example of how RoboMaker and the cloud can actually complement robots in houses.

Corey Quinn: When someone is looking through the vast, vast, vast list of various AWS services and they come across RoboMaker which again props to the name that's evocative. And let's say they're a new grad, they've graduated from college yesterday and now they're figuring out what they want to do with their career. I'm told it doesn't quite work that way anymore but let's pretend that, "Oh wait, you mean I need to get a job and here we are", and they see that. Is there an easy on ramp for this service for someone who is puttering around at home for fun? Is there a DeepRacer style equivalent or DeepRacer itself? Is this something that is going to be useful to someone who is not part of a larger organization or do you generally need to already have a number of prerequisites before this starts to add value?

Roger Barga: Yeah. First off when one looks at ROS and the educational materials that are available they immediately have access to this and in fact the University of Cambridge in the UK is using RoboMaker right now to teach their robotics class. We have an educational outreach program that contains 15 universities in higher education that are using RoboMaker to teach robotics to their students. In addition to that, once you actually launch RoboMaker and open it up low and behold you'll actually find that it's actually used to train DeepRacer for actually learning how to race around a track.

We have about a half a dozen today and more soon to come sample applications where we have the source code and we walk you through how we build the application. And in fact, the Turtlebot which is the most widely used robot for education and for hobbyists, all of these applications run on that robot. So, a customer can actually deploy it to the robot in their living room and actually see it execute the program, change their program, and see the Turtlebots behavior change as well. So, we have tried to put a number of resources available and we have more to come which I can talk about later.

Corey Quinn: With the understanding that forward looking statements, etc, etc. If we take a look back at some of the early launches of AWS where an awful lot of what was announced made absolutely zero sense. You're an online bookstore, why are you announcing a queuing service or this thing called an object store that none of us had ever heard of? And now, we are a decade later and change looking back on that and seeing okay yeah this was used to build an awful lot of transformative amazing things. And it's never quite clear how much of the world today you folks saw coming back when this stuff was launched. So, in the context of RoboMaker do you have a vision 10 years out from now or however long it is where we're going to be looking back and this was the most obvious thing in the world to build but needed to get to a certain place and now it empowers something transformative and grand? Or is this effectively aimed at today's customer requirements or both?

Roger Barga: Yeah, I think that's part of Amazon's culture of being customer obsessed and invent simplify. I can assure you that every feature of RoboMaker was derived from talking with customers with actual real pain points both within the company but also outside the company. And we're already consulting with companies now like now that we've built this service what other new features can we add for you to enable you to do more with it? So again, I think the reason these services become more valuable, most viable, is not because of how they started but how they evolved working with customers as their needs evolve, as their requirements evolve, as new applications of robots in our case evolve we'll evolve with it.

Corey Quinn: Common refrain from Amazon is that collectively as a company you are and I quote, "Willing to be misunderstood for long periods of time". If you look right now at the feedback you've gotten since launch, how people are using this service, how people are talking about your service, how do you see that RoboMaker is potentially being misunderstood today?

Roger Barga: Yeah. So, developers have not had access to cloud services to take advantage of both for fleet management, for augmenting the capabilities of the robot, so we do find it foreign to roboticists have not had access to this capability. So, we do find ourselves leading a dialogue with them about how we use cloud services to coordinate the robots in our fulfillment centers, how other companies are using cloud services to program and control robots that are out in space hurtling towards new planets. And so, it is a little bit of an education of what the possibilities are but then also listening of what new services should we build. So, I do believe that's the most interesting space is that partitioning of functionality between the edge and the cloud and how it can complement their capabilities.

Corey Quinn: One of the more signature attributes of AWS has been that when you wind up launching a service even if it's one that doesn't seem to make sense, doesn't wind up seeming to have a market, it never gets turned off. And every service you launch has customers to my understanding but API's are almost perceived as promises from you folks. I feel like I can wind up taking this recording of our conversation and archive it and in 50 years my descendants will be able to listen to it and they may laugh a awful lot of how naïve the conversation was etc. etc. but that service is still going to be there.

There are very few companies I would take that bet on particularly in the technology space but it seems to me that whenever something goes GA from AWS I have remarkably little hesitation in recommending that people build their business on top of that service. The counterpoint to that is API's are forever, for better or worse. Are you starting to see ways for the API to evolve? Have you gotten to a point, and you don't need to be specific on this, where now that you've seen even in the few months that it has gone GA that you would've made different decisions in how the service is interacted with, how it interacts with other services? Or alternately are you seeing ways to expand this far beyond where it is today and start embracing other AWS or third party services that at launch you really hadn't considered using?

Roger Barga: So, we don't have any crystal balls that tell us how an API is going to hold up over time but we do know ...

Corey Quinn: I was hoping I could borrow it if you did.

Roger Barga: But we do know we have customer trust and customers will actually take a dependency on our API, build their application on our API, and we can't the hubris to think we can simply change an API and break those customers. So, while we try to think very deeply and very careful about the functionality of an API is that is it simple as possible, is it as cross cutting as possible. 'Cause you can always add new API's with different functionality over time but you never want to deprecate a API for the fear of breaking potential customers. So, there's a thought process that goes in there but there's also an obligation to keep the API as it is. You can always add new API's with new functionality. And again, a lot of that is if you start with customers in beta programs that you know they're deriving value from it so you know that API is going to continue to add value in the ecosystem. That doesn't mean we're not going to add more as we see additional ways of exposing functionality in a simpler form for our customers, are more powerful. But there is that commitment that we will continue to support the API's we have exposed.

Corey Quinn: As you take a look across the landscape of other AWS services, at launch you mentioned that there were a bunch of very high level, very forward thinking services that RoboMaker integrated with and also CloudTrail. And I'm wondering if you take a look across the ecosystem of various AWS services are you seeing opportunities to integrate with different services that weren't necessarily there at first? And there are some ridiculous answers to that. "Yeah we want to make sure that the robot can speak appropriately to Cost Explorer." Sounds like something that not a lot of people would be clamoring for and then of course I tend to make no predictions about anything AWS does. There's nothing I'm saying that will never happen. For all I know there's a huge customer that you can't tell me about that's already doing a lot of work with robots and Cost Explorer but I can't imagine with that would look like.

Roger Barga: So, obviously we were very excited about integrating RoboMaker with Polly and Lex So customers could have a more natural interaction with it. Pleasantly surprised to find out later that Cloud Watch turns out to be one of the most commonly used services because customers want to know, "What the heck is going on in my robot? Where is my robot at?" And so, when you start to see services like that that expose meaningful value, SMS which actually allows me to stream messages off my robot, maybe send a notification to my robot is one we're looking at right now. We have the ability to actually put an agent on a robot and update the operating system on a robot with an AWS action EC2 instance. But we're seeing demand for that.

So, when we start to think about the pragmatic nuts and bolts about actually managing a robot, where it's at, what it's telemetry is, we see new services that we're going to be integrating over time. I'm pleasantly surprised to see when we launch with DeepRacer another team basically is actually using reinforcement learning for training a car. The interest and response we've gotten from companies that say, "I'd love to use reinforcement learning to actually train my robot to do new behaviors", and we need a deeper and richer integration with that. So again, I think in the fullness of time we'll be building new services for fleet management but integrating even more AWS services into robots.

Corey Quinn: Cloud Watch is I guess one of those personal hobby horses I have but credit where due. That service is evolving rapidly over the last few months and it's modernizing at a very interesting rate. A lot of the challenges historically that were there are no longer there now and I'm sure even fewer by the time this episode airs. So, I want to be very clear that was a joke, that was not an actual criticism of the service. I guess one interesting aspect of this is the idea that you mentioned with Lea the robot that walks around and integrates with various other services. It seems like this is almost a straight shot play for some of the various Alexa services out there as well. Where this winds up being able to empower different modes of interaction with existing things both around the home as well as in the workplace.

It feels to me like, and I can't even articulate how but this is a glimpse of a future where working on a computer no longer looks like sitting there typing into a terminal or an editor. It starts to look a lot more like a conversation and it starts to look like where you say, "We give a series of instructions and things start happening in the real world". It feels like a number of things I've never spent a lot of time going into on the AWS side that interface with the real world, things that I try not to deal with as best as possible. IOT is an example of this as well where it starts to hint at a future I can start to see the edges of but can't quite figure out what that's going to look like.

Roger Barga: Yeah it is. Again, if we think about robots in their most general sense they sense, they compute, and they act. And how many devices do we have to interact with today that do that for us? And think about the interface we have. I am confounded by my dishwasher. I can spend a half hour trying to get the darn thing to actually do the right load. What if I could walk up to it and tell it exactly what kind of load I wanted to run, what time I wanted it to start, and I could do the same with other appliances throughout my house which in fact are robots? What it's really surfacing is not necessarily going to replace developers but give a more natural way of interacting with these devices which are in fact robots and it's how we're interacting with RH devices, how are we programming and managing them. So, I think that's an exciting future.

Corey Quinn: I would absolutely agree with that assessment. One thing I will point out that I expect you won't have anything meaningful to share with me. You are not the GM or RoboMaker, you are the GM of AWS Robotics. And on the one hand I feel like this might wind up being a story similar to ground station which is in its own category called satellite. Either there's about to be a whole lot of interesting space releases or it just really didn't fit into any other existing categories. Is this scenario that you're seeing is ripe for expansion or is this more or less a, "Well, we didn't really know where else to put the robot thing and we're done"?

Roger Barga: We think this is an area of great innovation and great opportunity for the years ahead. So, much like a naming exercise for a product we apply the same naming exercise for our service and our team not wanting to be locked into any single definition of what the team does or owns. Thinking in the fullness of time there will be other services. We're already talking and thinking about validating with customers what those services might be. But it's clearly a new category of emerging technology so we should be prepared to build and manage services for our customers that are building robots out in the real world.

Corey Quinn: Thank you so much for taking the time out of your day to speak with me. This is an exciting space and I'm very interested to see what comes next.

Roger Barga: It's been fun talking to you today. Thank you.

Corey Quinn: Thanks very much. Roger Barga, General Manager of AWS Robotics. I'm Corey Quinn and this is Screaming in the Cloud.

Announcer: This has been this week's episode of Screaming in the Cloud. You can also find more Corey at screaminginthecloud.com or wherever fine snark is sold.

View Details

Some of the highlights of the show include:

  • Implications for migrating to AWS
  • Why and how for using Amazon vs hardware
  • The positive effects of mentoring for both the mentor and mentee
  • Technical vs Management tracks at a software company
  • Career advice for women in the tech field

Links:

  • https://www.digitalocean.com/
  • https://sendgrid.com/
  • DO.co/screaming
  • http://blog.dbsmasher.com/
  • https://github.com/

Transcript

Speaker 1: Hello and welcome to Screaming in the Cloud with your host cloud economist's, Corey Quinn. This weekly show features conversations with people doing interesting work in the world of cloud, thoughtful commentary on the state of the technical world, and ridiculous titles for which Corey refuses to apologize. This is Screaming in the Cloud.

Corey Quinn: This week's episode of Screaming in the Cloud is generously sponsored by DigitalOcean. From where I sit, every cloud platform out there, biases for something, some bias for offering a managed service around every possible need. A customer could have, others bias for, "Hey, we hear there's money to be made in the cloud. Maybe give some of that to us." DigitalOcean from, where I sit, biases for simplicity. I've spoken to a number of DigitalOcean customers and they all say the same thing, which distills down to they can get up and running in less than a minute and not have to spend weeks going to cloud school first. Making things simple and accessible has tremendous value in speeding up your time to market. There's also value in DigitalOcean offering things for a fixed price. You know what this month's bill is going to be, you're not going to have a minor heart issue when the bill comes due, and that winds up carrying forward in a number of different ways.

Corey Quinn: Their services are understandable without spending three months of study first, you don't really have to go stupendously deep just to understand what you're getting into. It's click a button or make an API call and receive the cloud resource. They also offer very understandable monitoring and alerting. They have a managed database offering. They have an object store, and as of late last year, they offer a managed Kubernetes offering that doesn't require a deep understanding of Greek mythology for you to wrap your head around it. For those wondering what I'm talking about, Kubernetes is, of course, named after the Greek god of spending money on cloud services. Lastly, DigitalOcean isn't what I would call small time. There are over 150,000 businesses using them today. Go ahead and give him a try or visit DO.co/Screaming and they'll give you a free hundred dollar credit to try it out. That's DO.co/Screaming. Thanks again to DigitalOcean for their support of Screaming in the Cloud.

Corey Quinn: Hello and welcome to Screaming in the Cloud. I'm Corey Quinn. I'm joined today by Silvia Botros, principal engineer at SendGrid. Welcome to the show, Silvia.

Silvia Botros: Thank you, Corey. Thank you for having me.

Corey Quinn: You might be better known on the Internet as dbsmasher, which is first an awesome name and secondly, deeply evocative of my entire relationship with databases dating back to my first experiences with technology 20 years ago.

Silvia Botros: Me too. I've been breaking them for a while. My current nickname internally to the team is RemoteEMP, 'cause a lot of times I'll just like touch something and it'll ... I seem to have this secret knack of a QA engineer. Like I'll touch a tool and I'm the first one to find like a bug or just make it completely kernel panic and things like that, and it's been challenging sometimes.

Corey Quinn: It's hard to realize at the time because when you, when you had everything you touch breaks, one of the, I think easiest things to do is fall into this pattern of assuming it's, oh, obviously it's because I'm cursed. And well, yes, you are, that does wind up having a tremendously valuable aspect to it of finding the ways that things break in failure modes that you hadn't predicted or experienced before. I've learned that one when I wind up talking to product teams at various cloud companies and they want me to look at things and the answer is not that the product is generally terrible, it's that I have really weird use cases and I find ways that things fail that hadn't been predicted before. That alone winds up being tremendously valuable if you could only bottle it.

Corey Quinn: It's super annoying when you're on deadline and trying to get something like guarding and, oh, you found a new interpreter bug, no one's done that recently. That's awesome, it's great, thanks. But I'd kind of like to get this thing to ship this week. That would be awesome. So I feel your pain there. I really do.

Silvia Botros: Yeah. I'm currently having this whole saga. We're doing this large project with SendGrid that we announced maybe like less than a year ago, just under a year ago, where we are moving our infrastructure from in pieces. It's a long journey. We're currently self-hosted in our old colo centers and we're moving to Amazon. And so one of the very first thing you are told when you're going to go to Amazon is if you do infrastructure as code, which you really should, Terraform is the thing, learn Terraform. I had been for years now managing the databases at SendGrid using Chef which has its own obviously wars, like nothing is perfect, but I had come to terms with learning all the things. Like the tricks like don't get too crazy with attributes, precedents, things like that. But now it's an entire mind shift with the declarative language that it doesn't even have loops yet.

Silvia Botros: It's been a thing for a month now. My team has watched me rail, like keep trying things with Terraform, trying to use it to deploy aurora clusters with regional replicas and it's been that kind of thing. I've already filed two bugs with Terraform, found a third one that totally applies to what I'm trying to do, and I am pretty close to just, you know, wrapping it in bash to try to work around the bugs while things are being figured out. So being being on that edge of like, yeah, this is awesome that I'm finding bugs but I really, I've been working at this for a month and I would like to never look at it again for awhile. That's definitely a thing.

Corey Quinn: That's the interesting part I've always found about managing infrastructure as code. Once you get it up and running it works and you do it in CloudFormation, Terraform, Shaft, Bash script, et cetera. And then you invariably don't touch it for six months, a year or whatnot, until you have to make changes. And then you go back to it and it's almost like relearning whatever it is that you wound up building the first time. And you look at this and it's, "What moron wound up writing this?" And you pull up git blame which we may as well called git shame because that's how everyone uses it, and you realize that the moron was you. And "Oh, we're just going to fix that now and we need never speak of this again." But it's one of the most demoralizing things in the world. It's like on the one hand it's, yeah, I'm good at this and I'm considered an industry leader. And on the other it's, ha, what is Terraform and how might one use a thing like that? It's a constant humbling of wherever we think we are in this space.

Silvia Botros: Oh yeah. Um so, about the git shame part, I used to do that way back when I started at SendGrid and it is incidentally my seven year anniversary today as we record this.

Corey Quinn: Oh, congratulations.

Silvia Botros: Thank you.

Corey Quinn: Seven years at a company in tech that's like 25 years on a watch in most-

Silvia Botros: I know ... It's it's I can't believe it either. I still vividly remember my first day actually. But I mean at the time I was still like a bit, like had a hard time with the empathy bit of working as an engineer. And like I would do this thing where I'd go git blame and be like, "What was that person's thing thinking?" But seven years in and at the same company, most likely they git blame will be me. So all the time I go back, I'm like, "Oh, man, that Silvia was really an idiot when she did this."

Corey Quinn: One of my favorite git extensions, it made the rounds a while back on GitHub, is git blames someone else, where it goes back and rewrites history and attributes commits to other people, which is just spectacular.

Silvia Botros: Oh, one of the people still on the ops team with us, and he's been at SendGrid longer than me, it's going to be nine years for him in a couple of months, he was the first DBA, although he never held a title, but he and I have this like friendship where we troll each other a lot. And I distinctly remember only a few months ago when someone who is much newer on the team finding a script that was causing issues and they were like debugging this thing that is very arcane and there they go to git ... No, they go into the script and it was, it's a Bash script and title at the top, whoever wrote it put in the author as dbsmasher. But then when I was like, I don't even recognize this, and I git blame it and it was my coworker. And that was definitely one of those, what we like to show each other in our team. We love each other so much. It's a good relationship where we just like throw jabs at each other like that all the time.

Corey Quinn: It's fun and it's easier to do in some cases with things like git just because you wind up in this world of this incredibly complex tool that's primary job is to make everyone feel stupid, no matter who you are, what you do, you will go past your comfort zone with git full stop. But then you wind up … find that other people seem to know certain things that you don't. It becomes a constant, first, excuse to share knowledge with people. And secondly, in that uncomfortable world of this complex critical thing that no one fully understands, comedy arises. And it's always a reliable direction to go in as far as easy jokes. At least, well, easy should be in quotes because it's git, nothing's easy.

Silvia Botros: Yeah. That's definitely one of the lessons learned after this many years. Seven years at SendGrid, like not everything has been smooth sailing at all times. But uh it's much healthier to come out of it with a laugh than get angry.

Corey Quinn: So, in the interest of full disclosure, I write a newsletter every week that goes out to give or take 10,000 people. And that is powered these days through SendGrid. I am a paying customer.

Silvia Botros: Excellent.

Corey Quinn: This is a podcast. This is, yeah, this is me using a service that works. So when I find a vendor I'm using, like SendGrid, is doing something publicly that is entertaining, in this case, as you just mentioned, moving to AWS, hey, that also is relevant to my interests, I immediately get simultaneously intrigued and a little worried. It's, "Huh, well, the service I've been using is super reliable, but now they're publicly changing a lot, a fundamental piece of the architecture. Well, that's going to be interesting. How is that going to play out potentially?" So I'm curious to hear, I guess first, how you folks think about the migration to AWS? Starting at the beginning, what drives that? Yeah, we're going to pick up this thing that we've built and move it to a completely different environment is generally not something a company does for funsies.

Silvia Botros: Exactly. So, the conversation around making this move started maybe a year and a half ago. We've been an AWS partner for a while. We are in their marketplace as one of the email providers right next to SES and you know getting traction there. But one of the other things that were also happening June at that time was we were talking as both as the ops org and generally as engineering of how much work was really going into hardware procurement, getting the racks in place, making sure they are actually usable, um getting them provisioned, getting them out to networking them in our internal system that we use to provision the host and then turning them into like KBM host. Like the amount of work that goes from, "Hey, we need racks in DCX," to the engineers have assets at hand that they can deploy things to, it's a very long road and it was a serious concern at our velocity.

Silvia Botros: There's also the other concern where as a business we send mail and we have known peaks of traffic. There are certain times where customers want to send more mail. For example, Black Friday and Cyber Monday is a known thing for us. Um so we just tweeted after last Cyber Monday, like we sent three billion emails just in the 24 hour span alone, with a lot of like impressive peaks within like one hour span or 15 minutes span internally. So if you're self-hosted, you find yourself having to procure hardware that accommodates the peaks, but then there's a lot of idle in other times. Weekends you know are pretty quiet. Not a lot of people send email over the weekends. So that became another concern is like how much money is sitting idle while just to be ready to use when we have high peak days, like Black Friday, Cyber Monday or other days. GDPR was another day as well that was very high peak. So those two things changed.

Corey Quinn: Oh, geez, yes. We call the GDPR day alone was one of those absolutely horrible days for anyone who depends on getting important things in their email inbox.

Silvia Botros: Here's the funny part. Our big data team was waiting to see what happens with that day because it was going to be an interesting one time event for them to see how we handle it. But from an operational perspective, we didn't feel it, as in we felt it way more in our own impersonal inboxes than, oh my God, look, the traffic is high and things are, knock on wood. Nothing went wrong that day, it just sort of flew by and the next day we went, "Oh, we sent like," I don't remember the number it was, but it was definitely north of two billion. "We sent that much? Oh, that's cool." So it was clear example that we had a pretty solid infrastructure that handled the scale like very comfortably.

Silvia Botros: But then at the other hand you should, like it was enough for us to go, "This is a good reason for us to move to the cloud because it would have probably cost us less in hardware or in like infrastructure costs to not have all that hardware sit there waiting for GDPR to happen and for us to use it." So, between the idle infrastructure concerns and the developer velocity concerns, it became clear that we were like, we needed something that could just elastically scale with us for delivering all these billions of emails. But also like you said, it's like we are a large company, we've been around for awhile, we have many, many customers. You can't just change tires on a bus that's flying at a hundred miles an hour and just be like, let's just stop in and get the tires off. It has to happen while you're running. So, that's why we decided it's not going to be like a quick thing. We set an ambitious goal still for a couple years, but we're doing it in small portions. There's a lot of education internally for all of our engineers as to how to actually build stuff in Amazon. It changes your architectural of mindset when you're designing things.

Corey Quinn: Well, hopefully. If it doesn't, you've created a heck of a problem.

Silvia Botros: Exactly. It's you never want to. There's a lot of things that you can take for granted when you are owning everything down to the network year versus when you go into Amazon and things are far less under your control. But I think ultimately it's better for us as engineers to do that way because then you'll learn to make what you build actually resilient and not just happens to work because you know if the network goes down I know who to poke.

Corey Quinn: It's interesting dealing with a company like SendGrid. In the interest of further disclosure, I started my career as a Unix administrator focusing on large scale email systems. Not SendGrid large scale, but large scale enough measured in the millions of users. As I was building out the newsletter, we talked earlier about the idea of finding weird edge case, strange bugs, it was very clear that I am as every other aspect of my life, not the typical case for anything. So I'm talking to an account rep as I'm getting my SendGrid accounts set up and they said, "Okay, how many emails do you wind up sending a week?" I'm like, "Oh, cool." I have at that point 3,000 subscribers. They said, "Great, cool. We can absolutely send email to your three million subscribers." And it's, "Whoa, whoa, whoa. You're off there a little bit." Okay. No matter how big this thing gets, we got you.

Corey Quinn: But the other side is cool and I want to build this newsletter management thing into it too. And the response was, "Oh, here's our API." So the things that seemed very complex to me of sending high degree high volumes of email with a great deliverability story, your entire corporate approaches, "Yeah, no problem. We got that." But the stuff that involves some of the higher level moving up the stack aspects is on the other side of a "Oh, yeah. That's not really what we do." And credit where due, I appreciated being told that explicitly, as opposed to the song and dance so many vendors like to do of, "Oh yeah, it'll also make fries," but that doesn't lead to a good experience. So, this is where we start and this is where we stop is fantastic.

Corey Quinn: A question that I have though, as you've since then started to broaden out a bit into a marketing campaign story and going down towards, I guess the, I don't want to say lower end of the market, but that's absolutely where I am. It's becoming an increasingly capable platform as time goes on. It's nice to see that enhancements, those enhancement start to come out. And I guess what I'm curious about and feel free to tell me you can't answer it, but is some of that flexibility and capability improvement being driven by, even in a tertiary sense, by the migration into the cloud?

Silvia Botros: It hasn't been driven by it, it's actually helped accelerate it though because our customers kept asking for it. They're like, "We love your API, but we also want to use it for handling marketing, and our marketers are not developers. They can't write API code. They can't write code that talks to APIs, they need to be able to do some things in the GUI."

Silvia Botros: So, that was one of the drivers actually because a couple of years ago we were embarking on uh the marketing campaigns application getting its new rewrites. We were wanting to add some features around automation, which is now in ... I think it's open beta, maybe closed beta, but I'm fairly certain it's on the website right now. Um that is marketing campaigns was one of the first parts of our product lineup that we wanted to move and to make all the new things that we add to it AWS Native. So, that was one of those things that, yeah, like once we do this, that's definitely the first one to start doing this with. A, because it's a smaller volume compared to our bread and butter infrastructure, email business line. And two, because once they do that, it has a lot more potential to be able to leverage all of the cool things that are in Amazon that we don't want to have to build internally.

Silvia Botros: 'Cause one of the biggest drivers for this move was also we need to focus our engineers on the things that are our business advantage. We no longer want to spend any engineering time building an internal provisioning system for hardware um or building a DIY metric stack or any of those things that don't actually, that we don't actually bill our customers for. So, when we start going into complex systems like handling a marketing contact and all of the things that go along with it, having the ability to use something like AWS where it's like there's a lot of managed things can help us build the higher up pieces faster.

Corey Quinn: Sorry, I was on mute. Please take that silence out.

Corey Quinn: There's always a capability story that starts to enhance itself as you start moving towards more dynamic infrastructures where, okay, we need to wind up being able to serve very bursty loads so when we wind up on GDPR Day or what not, suddenly we're at 18 times normal capacity and then that drops down to effectively nothing and we don't pay for it except when we're using it. There's something that winds up shifting from a perspective context around the idea of being able to take whatever it is you're doing and spin something up to see how it works and spin it down when you don't need it. That for whatever reason you don't have that level of dynamic ability to innovate when you're paying for hardware, having to get it racked and, "Oh, you want to try this new experiment? Cool. We're going to expedite that and get the infrastructure spun up in only two short months." It tends to react to

Silvia Botros: That's if you're lucky.

Corey Quinn: Yeah, oh, yeah. I'm speaking aspirationally here for startup territory, not big enterprises. It's, "Cool, we'll get that for you by third quarter," and they very conspicuously don't mention a third quarter of what year. So, historically as well, you were focused primarily on databases, and now as you mentioned to me earlier, you don't have the word database appearing anywhere in your title either because they've you've scared them all into submission, or it's the future we don't need databases anymore, or perhaps least interestingly but most likely your capability is now expanded beyond pure databases into other arenas. Can you give me a little context into that?

Silvia Botros: Yeah. So, way back when I started at SendGrid we didn't even have like written down like job titles, like the job title comes with your offer and sort of you kind of guessed what it means. I had the luxury when I started at SendGrid, I was one of the first actually who wasn't hired with a vague software engineer, quote unquote, title. It was more specific to DBA. So, I at least had an idea of where my scope was, but we didn't have enough effort at the time into explaining what those titles meant, what the levels attached to them meant, what competencies we expect. And after that to the next step as most companies do, people started needing that, so that was defined. But one of the biggest concerns that I had raised at the time was that the track seemed to, for individual contributors, would go all the way up to like senior engineer, or senior DBA in my case, and then it would stop. And other than that you started going into the management level where you start seeing like senior management director, senior director of VP, and so on.

Silvia Botros: This was around, I don't know, year four or so or five at SendGrid. And uh I started facing the usual question that all engineers around this time for like me face, do you want to go into management? I genuinely was not super interested in it. I knew that we still had a lot of technical things to solve and I wanted to be involved in solving them. But I still pointed out that, hey, this difference in the tracks is sticking out to me. So, to their credit at SendGrid, we started talking about, hey, do we need to redo our our tracks? And that's when SendGrid introduced a new rewrite of our leveling and our titles infrastructure.

Silvia Botros: We started having a fully fledged individual contributor track. We started having what we call principal engineer, principal engineer II, and architect. That role starts becoming more of a leadership without direct reports kind of role. There's definitely an emphasis in it on helps the other people in the org do stuff, be a force multiplier for other engineers, helps teach, helps mentor far more than, you know, goes in a black hole and comes out with a shiny fully done production product. So that's been really great, at least for me, I felt it was very empowering to be able to say, "Hey, I've learned a bunch of things over the last few years. I would like to continue solving problems here. But I would like that also acknowledged as part of the career track that we have."

Corey Quinn: One of the things I think that I admire the most about you has always been your willingness to help mentor people, help people understand something they didn't before. It's something that I would love to see more of across the industry. It was even in early days when I first was getting a vague awareness of who you are, your willingness to help people understand complex things in reasonable ways was one of your defining characteristics. It's one of those things where people never quite know what their own reputation is, but that's one that you've been carrying for an awful long time. I figured you're probably, I've been told this before, but in the off chance you haven't, that is how people tend to view what you do.

Silvia Botros: Thank you. That's really nice of you. I really appreciate that. Yeah. Um I didn't have that much of ... My time at SendGrid, for a really long time I was the only DBA. I've had a consultant outfit helping me out of for most of that time when I was solo, but as much as one tries to incorporate a consultant outfit as part of the org, there's always going to be that level of that you still have to be able to give them more specific things to do and you know contain the work so that the output is still usable when they are gone. So, for a long time back then, I was starting to get very aware of the fact that a lot of the engineers that we had at SendGrid at the time were still not super comfortable with you know databases and their quickest instinct was to ask the DBA to look at it. It would like, "Please go fix this, like whatever magic you do."

Silvia Botros: And I started having issues with burnouts for awhile and it was like I was feeling that I was doing the same thing over and over again. It was not really, I wasn't growing. I wasn't being challenged. I wasn't learning new things. So, that's when I decided that, no, I need to stop you know guarding things that way, I need to be able to teach the engineers how to figure out what's wrong, try to document how the code that's running the databases looks like, things like that. Because ultimately I was like, if nobody else knows how to do what I do right now, I will always be doing what I'm doing right now. So, that's where that motivation came from. Like a slightly selfish motivation still, but I'm glad it helps everybody else out.

Silvia Botros: Since then we've started hiring more DBAs. We are now a team of four, soon to be five, which is still mind boggling to me. But now that we're at this stage, since I've gotten principal engineer the database portion has been taken out of the title I've recently joined our architecture team. I'm working closely with our architect team and PE IIs who report directly to the CTO, which means I get to you know help with more strategic things. So, it's less about databases specifically and more about how to build the things that bring the value to the business. And I've been definitely enjoying that a lot recently.

Corey Quinn: Something fascinating is that you've been at SendGrid for, as you mentioned, seven years as of today. Congratulations again, but you haven't, I guess, crossed over, for lack of a better term, into the management track. You are very publicly an individual contributor. What drives that?

Silvia Botros: Uh it's funny you mentioned that because my boss and a former VP of ops would both say that, "Yeah, she didn't want to do management. We're still having her do all sorts of manage-y things." They're sort of sneaky that way. But what I'm doing is basically it's I help out with a lot of like the helping grow the younger engineers in the org without actually having to just write performance reviews and handle calibration meetings, which is really a nice thing to do. The reason I've tried to avoid doing management was it was two-fold. One of them was more personal is that ... Let me start with the other one first. The general one is because I felt that there was a lot more technical changes coming up that I would love to be involved in.

Silvia Botros: I'm very aware of the fact that if I do the move to management, I need to not be in the way for the code, like for how to build the things and it stops being about that. I'm not a huge fan of the interim stage of being asked to do both where you're in charge of people's performance reviews and you have to help grow people's careers at the same time. You're responsible for building large parts of the product yourself. Charity Majors talked about that recently. There's also that amazing book by Camille Fournier about The Manager's Path. They both acknowledge that stage, but they always say it's a difficult one that needs to be short lived and has you know a time box around it. But I knew that it would still mean like the next step would be step away from the technical and actually help with the people. And I'm fine with that, but I was like, I know myself I'm going to want to do the technical stuff and I know that we specifically at SendGrid have a lot of very interesting technical challenges coming up and I wanted to be part of them. So that was, that was the one side of it.

Silvia Botros: The other side was far more selfish just because of the fact that at the time when this fork in the road appeared for the first time for me, there was not yet any principal engineers at SendGrid who are women. I was like, "Nah, I kind of want to do that. I want to start, like I want to make the case that, yes, you can have, you can hire women engineers. They don't have to go into the human side or like the quote unquote 'soft skills' side of things to continue working and to advance their career." That's like, yeah, I can totally like go in and be part of the technical decisions and do that. I've recently given the architecture team flak because until I joined they were literally all like white dudes and they all loved plaid shirts. In fact, a specific shirt. So, when I joined the teams we sort of like decided to compromise where they decided to buy me the plaid shirt. But it's like now, yeah, like the team is not just dudes. So, I've been especially proud of that one.

Corey Quinn: It's neat to see people who have grown in their careers to incredibly lofty professional heights, remaining individual contributors. And SendGrid has a bit of a knack for this. One of the folks you have working over there used to be the CTO at a company I worked, now individual contributor. One of the most distinguished engineers I've ever had the privilege of working with. And the list goes on where there's very clearly a separate technical track at SendGrid that is not perceived as being lesser than management. And I really like that pattern and I wish that more companies would embrace that. In too many environments if you want to advance professionally, you're walking down a path of management, which is often terrible. It's you're a terrific engineer, so we're now going to focus on an orthogonal skillset that winds up having very little relevance to what made you good at this. So, we're going to lose a terrific engineer and gain a crappy manager. That sounds best for everyone, but it's the only way for you to develop. I really like it when companies go in a different direction.

Silvia Botros: And it wasn't from day one, like I said. We've had our own hurdles. Learning that lesson the hard way sometimes. We've promoted engineers into management roles and watched them not do so well. One of the people with me on the architecture team, he was sort of dragged into management for a bit and then he realizes like, it's not for me. The good thing about SendGrid and it's innate to our culture is that we have the humility, whether it's the executive team or the management team or the senior ICs to say, "Okay, that's not working. If I want to continue being here and providing value, we need to fix this part and it's fine to admit that management is not for me. I'm not going to do that." And that friend of mine, Shawn Kilgore, he ended up stepping away from management after, I think it was like a year, maybe less, and went back to, no, I want to be senior IC, I want to help this company build scalable you know long-term infrastructure that provides value, but I don't want to be, I don't want to manage people. I prefer to help the people without being directly responsible for the problems, reviews and all the other things that come along with it. And that was fine.

Silvia Botros: Unfortunately a lot of companies don't ... They consider that a failure and that you're just not good at either anymore, which is mind boggling to me. It's like, no, like I give you this entirely other job and didn't give you any help understanding how it's supposed to go, and then you failed and now I'm surprised and I'm going to let you go. It just doesn't make sense.

Corey Quinn: I think that there's a terrific story around the idea of being able to, I guess, continue to develop, nurture, and of course attract talent. It seems very strange to me that in many cases when someone is management at a company and they're not succeeding or thriving in that role, or they want to go back to working on as an IC for a variety of different topics, there's no way to do that without changing companies or losing face. People in the industry tend to view that as a perceived demotion and that's not great.

Silvia Botros: I think part of it is because there's this implied presumption that IC doesn't actually help lead that if. I've seen this before and I've argued with people on Twitter about it before. Where they'll say like senior engineer, like does she know algorithms and this many times or has like the focus with a senior individual computer, a senior software engineer is that they write code really, really well. And I've taken issue with that because ultimately I don't think that being IC means you're precluded from talking to people or precluded from helping mentor the younger engineers and the work. And once that dawn on people, it becomes clear that even as a principal engineer with no direct reports, I still have a responsibility towards everybody else in the organization to not just do things but to teach things. So, once that gets clear, you no longer have the delineation of like, yeah, the managers are higher caste than the individual contributors, and if I prefer to do IC work, then I'm going to forever have a ceiling on my career.

Corey Quinn: It's nice to see companies embrace that. I think that's something that I think the entire industry could learn from in more depth. I want to thank you for taking the time out of your day to speak with me, especially given this is your seventh anniversary at SendGrid. There are probably things you could be doing that are a lot more entertaining than listening to me prattle on. So, thank you for that.

Silvia Botros: Not at all. Thank you, Corey.

Corey Quinn: If people want to hear more from you, where can they go?

Silvia Botros: Well, I do have a blog. It's blog.dbsmasher.com and I'll make sure that you have the link for the show notes. If you would like a more stream of consciousness, there's my Twitter feed. Sometimes it is literally just screaming in the void against some sort of bug. I will occasionally OH my team, which they seem to apparently enjoy. I don't know. But whenever they say something witty, I'll tend to sometimes just OH it with no context and people seem to think that's funny. I certainly get some joy out of it. But yeah you can definitely catch me on Twitter and catch all sorts of things about what I'm doing right now that might be frustrating or even fun.

Corey Quinn: Well, thank you very much. I'll definitely put those links into the show notes. Thanks so much again for taking the time to speak with me and my ridiculous nonsense.

Silvia Botros: No worries. Thank you, Corey, for having me. It's been fun.

Corey Quinn: Silvia Botros, Principal Engineer at SendGrid. I'm Corey Quinn. This is Screaming in the Cloud.

Speaker 1: This has been this week's episode of Screaming in the Cloud. You can also find more Corey at Screaminginthecloud.com or wherever fine snark is sold.

View Details

The job market in the AWS world is complex and often confusing to both employers and employees. Wouldn’t it be great to have over 43,000 data points to draw a larger picture of the market and where you fall in line?

Today, we are talking to Kate Powers who walks us through the AWS Salary Survey from Jefferson Frank and discusses some interesting insights as well as real world examples of the findings.

Some of the highlights of the show include:

  • The AWS job market at large
  • Training Certificates: what’s their value
  • How much value is in a job title
  • Most desirable skills from employers
  • Gender representation in the industry
  • The discrepancy in compensation based on geography

Links:

  • https://www.jeffersonfrank.com
  • https://www.jeffersonfrank.com/aws-salary-survey/
  • https://twitter.com/_JeffersonFrank
  • https://www.linkedin.com/company/jefferson-frank/
  • https://www.facebook.com/JeffersonFrank.AWS

.

View Details

Years ago, if you wanted to launch an Internet company or Web application, you had to own necessary hardware. Now, the economics have changed drastically with the ease of Cloud computing. It’s still a new industry that people are trying to figure out, especially when it comes to cost and optimization.

Today, we’re talking to Dann Berg, a Cloud ops analyst at Datadog. He helps others understand and lower the cost of Cloud operations. Dann is a detective who is dedicated to figuring out why a company’s Cloud bill is so high.

Some of the highlights of the show include:

  • Companies struggle with field of Cloud economics; can be overwhelming because there’s so much to learn about products and implementation
  • Companies use the Cloud to grow quickly, which makes their Cloud costs grow quickly and more than expected
  • Only access to full list of every resource being used is the Cloud bill; there’s no comprehensive inventory service available
  • Companies need to offer visibility to Cloud bill; not everyone has access to understand how their actions impact the bill
  • Cost of Cloud bill is dependant on different factors, including new features, new users, and cost of goods sold (COGS)
  • Scale and manage bill by using a platform app or hiring a consultant/team
  • Understand pricing of AWS and learn best practices for cost controls early on
  • Don’t leave money on the table by focusing on engineering time - not best use of resources; focus on the smallest things that have the biggest impact
  • Cost is important, but don’t slow down those developing in the Cloud; open lines of communication to create culture to understand cost, value what’s measured

Links:

  • Dann Berg on Twitter
  • Datadog
  • re:Invent
  • AWS
  • Cost Explorer
  • CloudHealth
  • CloudCheckr
  • Cloudability
  • Lambda
  • EC2
  • GCP
  • Azure
  • CHAOSSEARCH

.

View Details

If you use MongoDB, then you may be feeling ecstatic right now. Why? Amazon Web Services (AWS) just released DocumentDB with MongoDB compatibility. Users who switch from MongoDB to DocumentDB can expect improved speed, scalability, and availability.

Today, we’re talking to Shawn Bice, vice president of non-relational databases at AWS, and Rahul Pathak, general manager of big data, data lakes, and blockchain at AWS . They share AWS’ overall database strategy and how to choose the best tool for what you want to build.

Some of the highlights of the show include:

  • Database Categories: Relational, key value, document, graph, in memory, ledger, and time series
  • AWS database strategy is to have the most popular and best APIs to sustain functionality, performance, and scale
  • Many database tools are available; pick based on use case and access pattern
  • Product recommendations feature highly connected data - who do you know who bought what and when?
  • Analytics Architecture: Use S3 as data lake, put in data via open-data format, and run multiple analyses using preferred tool at the same time on the same data
  • AWS offers Quantum Ledger Database (QLDB) and Managed Blockchain to address use case and need for blockchain
  • Authenticity of data is a concern with traditional databases; consider a database tool or service that does not allow data to be changed
  • Lake Formation lets customers set up, build, and secure data lakes in less time
  • DocumentDB: Made as simple as possible to improve customer experience
  • AWS Culture: Awareness and recognition that it takes many to conceive, build, launch, and grow a product - acknowledge every participant, including customers

Links:

  • Amazon DocumentDB
  • MongoDB
  • Amazon RDS
  • React
  • Aurora
  • re:Invent
  • DynamoDB
  • Amazon Neptune
  • Amazon Elasti-Cache
  • Amazon Quantum Ledger Database
  • Amazon Timestream
  • Amazon S3
  • Amazon EMR
  • Amazon Athena
  • Amazon Redshift
  • Amazon Managed Blockchain
  • Amazon EC2
  • Amazon Lake Formation
  • Perl
  • CHAOSSEARCH

.

View Details

Does operating system (OS) choice even matter anymore to most people? Especially with the emergence of serverless and containers? Debian may not see its name up in lights much these days, but it’s still very much front, center, and relevant to what people are doing in Cloud environments.

Today, we’re talking to Elana Hashman, a Python packager and Debian developer. Everything inside a base operating system may not be interesting to end users, but such a collection of components is necessary to create a functioning Linux system.

Some of the highlights of the show include:

  • Alternative Linux operating systems, including Amazon Linux 2
  • Level of awareness about free software when choosing and distributing an OS
  • What is a Python packager? How do you become one?
  • Python is the new default language due to growth and adoption of its ecosystem
  • Packaging community off-putting to beginners; find someone who understands the system to guide you

Links:

  • Elana Hashman
  • Elana Hashman on Twitter
  • Elana Hashman on Mastodon
  • A tale of three Debian build tools
  • Python
  • Python Packaging Authority
  • PyCon
  • Debian
  • The Debian Women Project
  • Docker
  • Red Hat
  • Fortran
  • Amazon Linux 2
  • Go
  • Perl
  • SaltStack
  • OpenHatch
  • SCALE
  • Jordan Sissel on Twitter
  • DigitalOcean

.

View Details

Companies can find working in the Cloud quite complicated. However, it’s a lot easier than it used to be, especially when trying to comply with regulations. That’s because Cloud providers have evolved and now offer more out-of-the-box services that focus on regulation requirements and compliance.

Today, we’re talking to Elliot Murphy. He’s the founder of Kindly Ops, which provides consulting advice to companies dealing with regulated workloads in the Cloud.

Some of the highlights of the show include:

  • Technical controls are easier, but requirements are stricter
  • Risk Analysis: Putting locks on things to thinking about risks to customers
  • Building governance and controls; making data available and removable
  • Secondary Losses: Scrub services to make scope and magnitude of loss smaller
  • Computing became ubiquitous and affordable; people started collecting data to utilize later - nobody gets rid of anything
  • General Data Protection Regulation (GDPR) set of regulations apply to marketing technology stacks to manage systems
  • Empathy building exercise and security culture diagnostic help companies understand compliance obligations
  • Security Culture: Beliefs and assumptions that drive decisions and actions
  • Evolution of understanding with public Cloud’s security and availability
  • Raise the bar and shift mindset from pure prevention to early detection/ mitigation; follow FAIR (factor analysis of information risk)

Links:

  • Kindly Ops
  • Amazon Web Services (AWS)
  • Microsoft Azure
  • Relational Database Service (RDS)
  • Google Cloud Platform (GCP)
  • Nist Cybersecurity Framework
  • GDPR Day
  • People-Centric Security by Lance Hayden
  • Stripe
  • Society of Information Risk Analysts (SIRA)
  • DigitalOcean

.

View Details

More and more enterprises and on-prem applications are moving to the Cloud. Therefore, flexibility, agility, time-to-market, and cost effectiveness need to be created to address a lack of visibility and control.

Today, we’re talking to Archana Kesavan, senior product marketing manager at ThousandEyes. The company offers a network intelligence platform that provides visibility to Internet-centric, SaaS, or Cloud-based enterprise environments. Our discussion focuses on ThousandEyes’ 2018 Public Cloud Performance Benchmark Report.

Some of the highlights of the show include:

  • Purpose of Report: Reveals network performance and architecture connectivity for Amazon Web Services (AWS), Google Cloud (GCP), and Microsoft Azure
  • Report gathered more than 160 million data points by leveraging ThousandEyes’ global fleet of agents that simulate users’ application traffic
  • Data collected during four-week period was ran through ThousandEyes’ global inference engine to identify trends and detect anomalies
  • Internet X factor when calibrating network performance of public Cloud providers; best-effort medium that has no predictability and is vulnerable to attacks
  • AWS’ performance predictability was lower than GCP Cloud and Azure leveraged their own backbones to move user traffic
  • Certain regions, such as Asia, were handled better by GCP and Azure than AWS
  • Customers should understand value of long-distance Internet latency when selecting a Cloud provider
  • Determine what the report’s data means for your business; conduct customized measurements for your environment

Links:

  • ThousandEyes
  • ThousandEyes on Twitter
  • ThousandEyes’ Blog
  • 2018 Public Cloud Performance Benchmark Report
  • Amazon Web Services (AWS)
  • Google Cloud
  • Microsoft Azure
  • AWS Global Accelerator for Availability and Performance
  • re:Invent
  • DigitalOcean

.

View Details

If you’re looking for older services at AWS, there really aren’t any. For example, Simple Storage Service (S3) has been with us since the beginning. It was the first publicly launched service that was quickly followed by Simple Queue Service (SQS). Still today, when it comes to these services, simplicity is key!

Today, we’re talking to Mai-Lan Tomsen Bukovec, vice president of S3 at AWS. Many people use S3 the same way that they have for years, such as for backups in the Cloud. However, others have taken S3 and ran with it to find a myriad of different use cases.

Some of the highlights of the show include:

  • Data: Where do I put it? What do I do with it?
  • S3 Select and Cross-Region Replication (CRR) make it easier and cheaper to use and manage data
  • Customer feedback drives AWS S3 price options and tiers
  • Using Glacier and S3 together for archive data storage; decisions and constraints that affect people’s use and storage of data
  • Feature requests should meet customers where they are, rather than having to invest in time and training
  • Different design patterns and best practices to use when building applications
  • Batch operations make it easier for customers to manage objects stored in S3
  • AWS considers compliance and retention when building features
  • Mentorship: Don’t be afraid of the bold ask

Links:

  • re:Invent
  • AWS S3
  • Amazon SQS
  • AWS Glacier
  • Lambda
  • CHAOSSEARCH

.

View Details

Do you have to deal with data protection? Do you usually mess it up? Some people think data protection architecture is broken and requires too many dependencies. By the time a business needs to backup a lot of data, it’s a complex problem to go back in time to retrofit a backup solution for an existing infrastructure.

Fortunately, Rubrik found a way to streamline data protection components. Today, we’re talking to Chris Wahl and Ken Hui of Rubrik.

Some of the highlights of the show include:

  • Transform backup and recovery to send data to a public Cloud and convert it to native format
  • Add value and expand what can be done with data - rather than let it sit idle
  • Easy way for customers to start putting data into the Cloud is to replace their tape environment; people hate tape infrastructure more than their backups
  • Necessity to backup virtual machines (VMs) probably won’t go away because of challenges; Clouds and computers break
  • Customers leaving the data center and exploring the Cloud to improve operations, utilize automation
  • Business requirements for data to have a level of durability and availability
  • People vs. Technology: Which is the bottleneck when it comes to backups?
  • Words of Wisdom: Establish an end goal and workflow/pathway to get there

Links:

  • Rubrik
  • Chris Wahl on Twitter
  • Chris Wahl on LinkedIn
  • Ken Hui on Twitter
  • Ken Hui on Medium
  • Amazon S3
  • IBM AS/400
  • Amazon EC2 Instances
  • Azure Virtual Machine Instances
  • re:Invent
  • DigitalOcean

.

View Details

Do you have some spare time? Can you figure out an easier way to do something? Then, why not build some software?!

Today, we’re talking to Ian Mckay of Kablamo, an Amazon Web Services (AWS) consultancy. He is the author of Console Recorder, which is a browser extension that records your actions in the Management Console to convert them into SDK code and infrastructure as code templates.

Some of the highlights of the show include:

  • Timeline to build Console Recorder
  • Infrastructure as Code: How to code repeatedly without starting over and take ownership of what you built by hand
  • AWS vs. Individual Achievements: People asked AWS for years to create something to record console click-throughs that Ian did in his spare time
  • Console Recorder support for any browser that exports Web extensions
  • Sharp edges of what’s expected of Console Recorder to speed up development
  • Management Console’s unreadable responses require reverse engineering
  • Console Recorder: Recommended use cases and areas
  • How to alleviate security concerns with Console Recorder
  • Changes to Management Console that may break things
  • Ian’s past, present, and future projects and products
  • Words of Wisdom: If you don’t like something, just fix it yourself

Links:

  • Ian Mckay on Twitter
  • AWS Console Recorder
  • Kablamo
  • AWS
  • CloudFormation
  • Terraform
  • MediaLive
  • Jeff Barr
  • re:Invent
  • CDK
  • Google Cloud Platform
  • AWS Management Console
  • AWS RDS
  • AWS Lambda
  • DigitalOcean

.

View Details

A Manager README is a document designed to establish clarity between a manager and those who report to them. These documents are especially useful for onboarding content. For example, if you have someone new starting on your team, there's so many things you need to share with them - pieces of advice and guidance that help them to make the best decision about what to do in specific situations. A Manager README sets some expectations in advance to make things easier and reduce friction and anxiety for team members.

Today, we’re talking to Matt Newkirk, who manages Etsy’s localization and translation group. He explains that even if your company has an intensive onboarding program and review process, some things are still left out. A Manager README is a helpful and proactive piece of content that prompts conversations about how people perceive things.

Some of the highlights of the show include:

  • Avoid writing READMEs that are extremely self-centered/arrogant
  • READMEs clarify what to do until a relationship is established between the manager and their employee
  • Get feedback early on to make sure that what you include in the document is helpful; it should reflect reality and be discussed
  • Share README with your manager to make sure you’re both on the same page about team philosophies and expectations
  • README is a living document that needs to be updated occasionally because things change
  • README adds context; it’s not designed to make employee feel like they’re back in school and panicking because they’re not prepared
  • Manager README - Not Matt’s best selection of terminology
  • Who’s the best boss you ever had? Why? They can be a force that shapes your life and career from the right perspective
  • Philosophy of Management: Don’t do what terrible managers have done; be transparent about strategic reasons for priorities changing

Links:

  • Matt Newkirk
  • Matt Newkirk on LinkedIn
  • Matt Newkirk on Twitter
  • Share your Manager README
  • Etsy
  • Etsy’s Job Openings
  • Shane Garoutte on LinkedIn
  • Kubernetes
  • Everbridge
  • Digital Ocean

.

View Details

Would you like access to unlimited retention of your data within your Amazon S3, which costs far less than online storage on disc? Well, the next time you’re at re:Invent, visit CHAOSSEARCH’s booth.

Today, we’re talking to Pete Cheslock, vice president of products at CHAOSSEARCH and former vice president of operations at Threat Stack. CHAOSSEARCH helps people get access to their login event data using Amazon S3.

Some of the highlights of the show include:

  • re:Invent - Year of the Pin: People go nuts for conference swag and were collecting pins as if they were gold
  • Scan Your Badge and Drip Emails: Annoying and passive-aggressive marketing trends meant to be spontaneous and interesting
  • Need a job? Corey’s looking to hire a “Quinntern” to use a tag email address to gather conference swag at the next re:invent; if interested, contact him
  • Corey and Pete’s Swag Rules: Something you want or can use, continues to be valuable, no sizes, no socks
  • Densify Drama: Conference flyer to generate leads failed, created complaints
  • Track and analyze data, but don’t use it to invade privacy or become creepy
  • Las Vegas: Right place for conferences, such as re:Invent?
  • Rather than focusing on going to conference sessions, make meeting and talking to people doing interesting things your priority
  • Midnight Madness Event: Only place Corey could do stand-up Cloud comedy
  • re:Invent 2019: Plan appropriately, identify what you want to get out of it, register ASAP to get a nearby hotel, and schedule meetings with AWS staff

Links:

  • Pete Cheslock on Twitter
  • Pete Cheslock on LinkedIn
  • CHAOSSEARCH
  • Threat Stack
  • AWS
  • Amazon S3
  • Amazon Elasticsearch
  • re:Invent
  • Corey Quinn’s Newsletter
  • Corey Quinn on Twitter
  • Corey Quinn’s Email
  • Sonian
  • Acloud.guru
  • Densify
  • Oracle
  • Apache Cassandra
  • DigitalOcean
  • AWS re:Invent 2018 - Keynote with Andy Jassy
  • AWS re:Invent 2018 - Keynote with Werner Vogels
  • AWS re:Inforce
  • VMware
  • Dreamforce
  • Kubernetes
  • Datadog

.

View Details

Have you ever had high expectations about a new software product? Did you think it was going to be spectacular? Instead, did it become less about solving a problem for you and more about reaching a bunch of billable consultants? The dynamics of open source communities and the Cloud platform can make or break software products.

Today, we’re talking to Andrew Clay Shafer, who was a notable voice during the days of OpenStack. He had high hopes for OpenStack, which was an effort to bring a democratized solution of Cloud computing to anyone’s data center. He describes the importance of understanding the challenges associated with open source projects in order for them to be successful.

Some of the highlights of the show include:

  • Open source is not a business model; capture value for customers, or they’ll go with a different solution
  • Openness/Closure: Every open source project has its own community dynamics
  • Losing sight of level of expertise for profitability and easy path to useage
  • Whether to become a product or service company - difficult to be both effectively or go from being one to the other; build partner relationship, focus, and say “no”
  • Lack of awareness about AWS Outposts admitting public Cloud is no longer a viable business model
  • Amazon relentlessly focuses on what its customers want and tries to keep promises about what it can and can’t do
  • Cloud Native: Not where you run, but how you run; confining variables
  • Self-fulfilling prophecy to under deliver when you make the bad decision to under source IT across the board
  • Cloud Native, DevOps, SRE: Buzzwords that equal one thing and work together
  • Dilemma of not building everything and buying some things, but you can’t buy everything; humans like to shop and go with the easiest option

Links:

  • Andrew Clay Shafer on Twitter
  • Andrew Clay Shafer on LinkedIn
  • Puppet
  • Re:invent
  • OpenStack
  • Eucalyptus
  • Docker
  • Redis
  • MongoDB
  • Confluent
  • Kubernetes
  • AWS Outposts
  • AWS Ground Station
  • AmazonBasics
  • Simon Wardley
  • Maslach Burnout Inventory
  • Datadog

.

View Details

You can't make money selling to developers! The bottleneck of getting business requirements and creating business value used to mean waiting for the next waterfall release. That’s not the case anymore in the venture community. There’s programmatic access to infrastructure and DevOps/agile developments that offer super-fast cycle times. Now, the bottleneck is about how fast your developers can move and how much they can get done.

Today, we’re talking to Joseph Ruscio, general partner at Heavybit Industries, which is an accelerator for seed-stage companies and focuses on developer-first products. Tools and products that get you more leverage out of your developers are incredibly valuable.

Some of the highlights of the show include:

  • Measuring maturity of startups’ engineering teams by looking at SaaS list - what products they have in place and how many are using out-of-house vendors
  • Customers don’t care how curated or artisan a piece of your stack is, they only care that it works
  • Not all claims (scales infinitely or never fails) are true when it comes to products on the market, so people are skeptical
  • Heavybit focuses on helping businesses build a bottoms-up, grassroots community around its products and a disciplined inside/direct sales motion
  • Build vs. Buy: Whatever people try to do themselves is a costly, pale imitation of something they can buy
  • Advice for New Entrepreneurs: Never compete with AWS on hosting compute because it will obliterate and Amazon is great at plumbing, terrible at painting
  • AWS’s version of your product won't be as sophisticated; continually work on it to deliver a more seamless product and customer success experience
  • Measure downtime/outages in terms of dollars by using monitoring tools that deliver more holistic, integrated, comprehensive experience than CloudWatch
  • Starting a company is easier; even if you're the 800-pound gorilla in the category you created, keep innovating and building or Amazon’s coming after you
  • Azure, unlike GCP, has ability to meet customers where they are, rather than telling them where they should be
  • Understand the problem your customer is trying to solve and understand how far out of their current comfort zone they're willing to go to solve that problem
  • Software exists to create business value; it doesn't matter what it's written in or how it's hosted, so some systems will be around for a long time

Links:

  • Joseph Ruscio on Twitter
  • High Leverage Podcast
  • Heavybit Industries
  • Heavybit Library
  • Serverless Framework
  • Pagerduty
  • Stripe
  • Circle
  • Lightstep
  • LaunchDarkly
  • Treasure Data
  • Replicated
  • AWS
  • Twilio
  • Librato
  • re:Invent
  • MongoDB
  • Kubernetes
  • Rackspace
  • New Relic
  • SolarWinds
  • CloudWatch
  • GCP
  • Azure
  • SimpleBB
  • Datadog
  • Digital Ocean

.

View Details

Do you like to hear yourself talk? Especially while on a stage and in front of a lot of people? How do you come up with ideas to talk about? What process do you use to build a conference talk or presentation?

Today, we’re talking to Matty Stratton of PagerDuty. His job involves building conference talks and finding ways to continuously improve them. Public speaking can be intimidating, so he shares some tips and tricks that have worked for him.

Some of the highlights of the show include:

  • Avoid creating something brand new for every event
  • Don’t tell flattering stories about things that happened to you; may be uplifting, but doesn't resemble reality
  • Failure stories are fantastic because people relate to making terrible decisions
  • Everyone who gives a talk panics, gets nervous, and thinks they’re about a sentence away from stammering and falling off the stage; almost never happens
  • Audience wants you to succeed because they're there to learn; no one is hoping a presenter messes up
  • Preparation is key; could build a talk at the last minute, but it would be much better, if you prepared for it
  • Don’t intentionally try to think of something; have conversations with people and listen to other talks to develop anecdotes, stories, and cold opens
  • Humor can be tricky; what you think is funny, other people might not
  • Make things memorable; show good ideas by showing bad ideas - it’s the ‘don't do this, do this instead’ model
  • Submit early and often, but submit appropriately; if you are always submitting stuff that’s inappropriate for an event, your stuff starts to be ignored
  • Sometimes, you may want to avoid slides that auto advance; if you trip over yourself: Stop, repeat, back up, take questions, etc.
  • Try not to read from notes or slides; takes the life and engagement out of the talk
  • People can only do one thing at a time - listen or read
  • Practice: Record yourself every time you practice and watch it; focus on blocking and tackling
  • You have about 45 seconds to grab people's interest before they look at their phone; get them engaged via a story, picture, or anecdote

Links:

  • Matty Stratton’s Presentations
  • Matty Stratton on Twitter
  • PagerDuty
  • Arrested DevOps
  • Hot Takes, Myths, And Fake News—Why Everyone Is Wrong About DevOps, Except For Me
  • DevOps Dispatch
  • LastWeekinAWS
  • Jez Humble
  • Robert Rodriguez
  • Rebel Without A Crew
  • Adam Jacob from Chef
  • Terrible Ideas in Git
  • Azure DevOps
  • Emily Freeman
  • Decker Communications
  • Don't You Know Who I Am?!
  • Datadog

.

View Details

Do you understand how tabs work? How spaces work? Are you willing to defeat the JSON heretics? Most people understand the power of the serverless paradigm, but need help to put it into a useful form. That’s where Stackery comes in to treat YAML as an assembly language. After all, no one programs processors like they did in the '80s with raw assembly routines and no one programs with C. Everyone is using a higher-level scripted or other programming language.

Today, we’re talking to Chase Douglas, co-founder and CTO of Stackery, which is serverless acceleration software where levels of abstraction empower you to move quickly. Stackery has an intricate binding model that gives you a visual representation - at a human logical level - of the infrastructure you defined in your application.

Some of the highlights of the show include:

  • Stackery builds infrastructures by using best practices with security applications
  • What's a VPC? Way to put resources into a Cloud account that aren’t accessible outside of that network; anything in that network can talk to each other
  • Lambda layers let developers create one Git layer that includes multiple functionality and put it in all functions for consistency and management
  • Git is an open-source amalgam of different programming languages that has grown and changed over time, but it has its own build system
  • Stackery created a PHP runtime functionality for Lambda; you don't want to run your own runtime - leave that up to a Cloud service provider for security reasons
  • Should you refactor existing Lambda functions to leverage layers? No, rebuild everything already built before re-architecting everything to use serverless
  • Many companies find serverless to be useful for their types of workloads; about 95% of workloads can effectively be engineered on a serverless foundation
  • Trough of Disillusionment or Gartner Hype Cycle: Stackery wants to re-engage and help people who have had challenges with serverless
  • Is DynamoDB considered serverless? Yes, because it’s got global replication
  • Puritanical (being able to scale down to zero) and practical approaches to the definition of serverless

Links:

  • Stackery
  • JSON
  • AWS
  • Lambda
  • Aurora Serverless Data API
  • Hype Cycle
  • Secrets Manager
  • YAML
  • S3
  • GitHub
  • GitLab
  • AWS Codecommit
  • Node.js
  • WordPress
  • re:Invent
  • Ruby on Rails
  • Kinesis Streams
  • DynamoDB
  • Docker
  • Simon Wardley
  • Datadog

.

View Details

What’s hiring in the world of Cloud like? What are companies looking for in possible employees? What kind of career trajectory should applicants display?

Today, we’re talking to Don O’Neill, who has had an interesting career path and the archetype of who most companies want to hire. He’s been an independent contributor, platform leader, and Cloud consultant. Currently, Don is platform engineer manager at Articulate, an eLearning software solution for course authoring and eLearning development. He works with platform engineers to automate Blue Ocean pipelines with Docker, Terraform, and various Amazon Web Services (AWS) technologies, such as Elastic Beanstalk.

Some of the highlights of the show include:

  • Don reached out to his network to ask people that he had a professional relationship with about who was hiring and what challenges they faced
  • Don’s “Therapy”: Go to meet-ups to talk about DevOps topics; serves as a “I’ve-got-to-get-my-hiney-out-of-the-house-and-get-some-social-time”
  • Don’s journey from being a “wee lad in the industry” to a senior member/leader and giving back as a way to recognize those who helped him along the way
  • Hiring Horror Stories: People going through borderline ridiculous levels of hiring games and terrible interview paradigms
  • Companies sometimes look for something too specific - exact match instead of fuzzy match; they never have time to train, but time to look for a perfect unicorn
  • Articulate’s Hiring Process: Day 1 - Slack interview; Day 2 - Technical pieces; and Day 3 - Pairing with others
  • Articulate looks for people enthusiastic about technology, able to learn, and with emotional intelligence; company values independence, autonomy, and respect
  • Companies that spend several hours to make a hiring decision tend to have less success with those they hire
  • Cloud Certificates/Certifications: Can be valuable for applicants with no real-world experience; they don’t indicate how they’re going to work or learn
  • Applicants need to demonstrate a base level of knowledge; if they don’t have a skill set, they should start a project to learn about something - learning is fun
  • If you’re established in your career, reach out to someone just starting out to guide them
  • If you’re starting out in your career, reach out to people to talk about the next steps to take in your career (contact Corey or Don)

Links:

  • Don O’Neill on Twitter
  • Articulate
  • Hangops.slack.com
  • CoffeeOps
  • AWS
  • Azure
  • Docker
  • Terraform
  • Elastic Beanstalk
  • Autoscan
  • Marchex
  • Apex Learning
  • Dice
  • Monster
  • Indeed
  • Switch App (Tinder for Jobs)
  • Kubernetes
  • Spotify in Stockholm
  • CrowdStrike
  • re:Invent
  • AWS Summits
  • Digital Ocean

.

View Details

Do you enjoy watching sports? Wear your favorite team or player’s jersey? Are you a fan who has shopped at Fanatics on the Cloud?

Today, we’re talking to Johnny Sheeley, director of Cloud engineering at Fanatics, which is a sports eCommerce business that manufactures and sells sports apparel. Fanatics runs Cloud engineering to provide a robust and reliable set of services by building and deploying applications on top of the Azure Data Lake Store (ADLS) platform.

Some of the highlights of the show include:

  • If you compete with Amazon, be ready for it to come after you; some companies avoid its Cloud perspective or go multi-Cloud (paranoia-based movement)
  • Focus on your ability to make your business function smoothly
  • Transition, migration, and abstraction may be painful, but should not stop work; paying for Cloud-agnostic technology may not be worth it
  • Challenges of governing use of Cloud resources to prevent mistakes/problems related to Fanatics’ security and budget
  • Data collected focuses on what’s trending up or down to select an instance type that calculates costs; remain flexible and be aware of what you pay
  • Natural instinct is to blame people; mistakes are made, especially when a human factor is introduced to an automated system
  • Creating a mindset that focuses on feature and detail-oriented is challenging
  • Cottage industry of code bases running in Big Data and other expensive realms
  • As a product continues to evolve and grow, governance comes along for the ride and AWS bills are streamlined
  • Will serverless, Lambda, and RDS change how Amazon charges in the future?
  • State of scale of AWS and developing a more palatable method for releases because people can’t keep up with them and stop paying attention
  • Two-Pizza Team: Amazon’s management philosophy that any team that works on a service should be able to be fed with two pizzas
  • Such small teams work quickly and have the freedom to fail, but Amazon has a reliability for the longevity of its different services

Links:

  • Johnny Sheeley's Email
  • Johnny Sheeley on Twitter
  • Rands Leadership Slack
  • Hangops.slack.com
  • Fanatics
  • Kubernetes
  • Azure
  • Lambda
  • RDS
  • Getafix: How Facebook Tools Learn to Fix Bugs Automatically
  • Accidentally Quadratic Blog
  • re:Invent
  • Jeff Barr’s AWS News Blog
  • Amazon SimpleDB
  • Lots of Amazon's projects have failed...and that's ok, says Amazon's Andy Jassy
  • Digital Ocean

.

View Details

Did you know that you can now run Lambda functions for 15 minutes, instead of dealing with 5-minute timeouts? Although customers will probably never need that much time, it helps dispel the belief that serverless isn’t useful for some use cases because of such short time limits.

Today, we’re talking to Adam Johnson, co-founder and CEO of IOpipe. He understands that some people may misuse the increased timeframe to implement things terribly. But he believes the responsibility of a framework, platform, or technology should not be to hinder certain use cases to make sure developers are working within narrow constraints. Substantial guardrails can make developers shy away. With Lambda, they can do what they want, which is good and bad.

Some of the highlights of the show include:

  • Companies are using serverless as a foundation and for critical functions
  • Serverless can be painful in some areas, but gaps are going away
  • Investing in the Future: Companies doing lift-and-shift to AWS are looking at technology they should choose today that’s going to be prominent in 3 years
  • Serverless empowers new billing models and traces the flow of capital; companies can choose to make pricing more complicated or simplified
  • What value are you providing? Serverless can offer flexible pricing foundation
  • When something breaks, you need to be made aware of such problems; Amazon bill doesn’t change based on what IOpipe does, which is not true with others
  • Developers are the ones woken up and on call, so IOpipe focuses on providing them value and help; they are not left alone to figure out and fix problems
  • Serverless and event-driven applications offer a new type of instrumentation and observability to collect telemetry on every event
  • For serverless to go mainstream, AWS needs to up its observability level to gather data to answer questions
  • AWS, in the serverless space, needs to make significant progress on cold starts in other languages, and offer more visibility and easier deployment out of the box

Links:

  • IOpipe
  • Episode 16: There are Still Servers, but We Don't Care About Them
  • Lambda
  • Google App Engine
  • Python
  • Node.js
  • Kubernetes
  • Simon Wardley
  • DynamoDB
  • re:Invent
  • Perl
  • PowerShell
  • Digital Ocean

.

View Details

In the early days, angry nerd corners on the Internet viewed Slack and some of its predecessors as, “Oh, it’s just IRC. Now, you pay someone for it.” Many fell into that trap of wondering about what value such systems offered.The big differentiator? Slack is built as a collaborative business tool.

Today, we’re talking to Holly Allen, who helped make government software better while serving as the director of engineering at 18F. Now, she’s a senior engineering manager at Slack, a collaborative chat program where you can do most of your work through a rich platform of integrations. Holly enjoys taking a weird set of skills that make a computer do things and convincing people who know how to make computers do things do things.

Some of the highlights of the show include:

  • Safety engineering brings chaos and resilience engineering, incident management, and post-mortem processes together for resiliency and reliability
  • Slack strives to move really fast while being in complete control
  • Slack is primarily on AWS, but is working on a multi-Cloud strategy because if AWS is down, Slack still needs to work
  • Slack has a close relationship with AWS and is a collaborative company; it has immediate access to AWS staff anytime there’s a problem
  • Slack uses Terraform and Chef and working to determine if its production workflows in Kubernetes would be worthwhile
  • Disasterpiece Theater: Real scenario that might happen and surmise what will happen; don’t cause production issues, but teach Slack employees
  • Slack hires collaborative, empathetic people to create a collaborative environment where everyone works together toward a goal
  • Slack was firmly in a centralized operations model, but is transforming toward development teams to increase responsibility and service ownership
  • Slack doesn’t encourage remote work because it’s not in a position to put in that investment; day-to-day work happens in hallways and between desks
  • Slack sees itself as an enterprise software company; an enterprise software company must have enterprise software reliability, stability, and processes
  • Slack has thousands of servers, so events and disruptions happen more often; system needs to respond, react, and repair itself without human intervention

Links:

  • Holly Allen on Twitter
  • 18F
  • Slack
  • Freenode IRC
  • HipChat
  • AWS
  • Kubernetes
  • Terraform
  • Chef
  • QCon
  • Datadog

.

View Details

If you’ve been doing DevOps for the past 10-20 years, things have really changed in the industry. There’s no longer large pools of help desk support. People aren’t climbing around the data center and learning how to punch down cables and rack servers to gradually work their way up. Now, entry level DevOps jobs require about five years of experience. So, that’s where internships play a major role. But how can an internship program be set up for success? Where is the next generation of SREs or DevOps professionals coming from? Where do we find them?

Today, we’re talking to Fatema Boxwala, who has been an intern at Rackspace, Yelp, and Facebook. She’s a computer science student at the University of Waterloo in Canada, where she’s involved with the Women in Computer Science Committee and Computer Science Club. Occasionally, she teaches people about Python, Git, and systems administration.

Some of the highlights of the show include:

  • Mentors made Fatema’s intern experience positive for her; made site reliability and operations something she wanted to do
  • Academic paths don’t tend to focus on such fields as SRE, and interns tend to come exclusively from specific schools
  • Fatema’s school requires five internships to graduate and receive a degree; upper-year students are already very qualified professional software engineers
  • Companies don’t have time to train and want to find someone with an exact skill set; instead of hiring someone, they spend months with an unfilled position
  • Continuity Problem: You can’t train someone to be a systems administrator, if you aren’t willing to give them certain privileges due to inexperience
  • Use a low-stakes environment to train, where mistakes can be made; most systems aren’t on a critical path - don’t keep people away from contributing
  • If you have never broke production, that means either you’re lying or you’ve been in an environment that didn’t trust you to touch things that mattered
  • Internship should mimic the kind of work that everyone else is doing; give them responsibilities where their work has an impact
  • Bad mentors lead to bad internships; person in charge of your success doesn’t have the necessary skills; needs to be a good communicator, set expectations
  • As the intern, ask about possible outcomes of internship early on; mentors should be clear about expectations, feedback, and offers

Links:

  • Fatema Boxwala
  • Fatema Boxwala on Twitter
  • Jackie Luo on Twitter
  • Julia Evans Zines on Twitter
  • SREcon MEA
  • Digital Ocean

.

View Details

Are you interested in computer science? How would you like to go to school for free and learn what you need to in just a few months? Then, check out Lambda School!

Today, we’re talking to Ben Nelson, co-founder and CTO of Lambda School, which is a 30-week online immersive computer science academy. Lambda School has more than 500 students and takes a share of future earnings instead of traditional debt. So, it's free until students get a job.

Some of the highlights of the show include:

  • Bootcamps were created to address engineering shortages and quickly move people into technical careers
  • Lambda is not explicitly a bootcamp; its 30-week program gives students more instructions and more time spent on developing a portfolio
  • Lambda also makes time to cover computer science fundamentals; teaches C, Python, Django, and relational database - not just JavaScript
  • Employers appreciate the school’s in-depth and advanced approach, which results in repeat hires
  • Lambda avoids the typical reputation of traditional for-profit educational institutions by being mission-driven and knowing its investors want ROI
  • Lambda aligns its incentives with those of students; an income share agreement means the school doesn’t make money, unless students are successful
  • Lambda’s 7-month program is less of a risk for someone later in their career; some don't have capital to support their family while going to school for 4 years
  • Lambda incentivizes healthy financial habits; after two years of repayment, students can put that money into retirement, savings, and investments
  • 5 Tracks Now Offered by Lambda: iOS development, UX, Full Stack Web development, data science, and Android development
  • Mastery Based Progression System: When you're learning something sequentially, where knowledge builds, you don't move on until you’ve mastered it
  • Lambda’s acceptance rate is around 5% and based on people who can keep up
  • Lambda works with different partner companies to help them find qualified graduates - people they want to hire

Links:

  • Lambda School
  • Ben Nelson on Twitter
  • Y Combinator
  • Wealthfront
  • Datadog

.

View Details

Have you ever been on-call duty as an IT person or otherwise? Woken up at 3 a.m. to solve a problem? Did you have to go through log files or look at a dashboard to figure out what was going on? Did you think there has got to be a better way to troubleshoot and solve problems?

Today, we’re talking to Sam Bashton, who previously ran a premiere consulting partner with Amazon Web Services (AWS). Recently, he started runbook.cloud, which is a tool built on top of serverless technology that helps people find and troubleshoot problems within their AWS environment.

Some of the highlights of the show include:

  • Runbook.cloud looks at metrics to generate machine learning (ML) intelligence to pinpoint issues and present users with a pre-written set of solutions
  • Runbook.cloud looks at all potential problems that can be detected in context with how the infrastructure is being used without being annoying and useless
  • ML is used to do trend analysis and understand how a specific customer is using a service for a specific auto scaling group or Lambda functions
  • Runbook.cloud takes all aggregate data to influence alerts; if there’s a problem in a specific region with a specific service, the tool is careful to caveat it
  • Various monitoring solutions are on the market; runbook.cloud is designed for a mass market environment; it takes metrics that AWS provides for free and makes it so you don’t need to worry about them
  • Will runbook.cloud compete with or sell out to AWS? Amazon wants to build underlying infrastructure, other people to use its APIs to build interfaces for users
  • Runbook.cloud is sold through AWS Marketplace; it’s a subscription service where you pay by the hour and the charges are added to your AWS bill
  • Amazon vs. Other Cloud Providers: Work is involved to detect problems that address multiple Clouds; it doesn’t make sense to branch out to other Clouds
  • Runbook.cloud was built on top of serverless technology for business financial reasons; way to align outlay and costs because you pay for exactly what you use
  • Analysis paralysis is real; it comes down to getting the emotional toil of making decisions down to as few decision points as possible
  • Save money on Lambda; instead of using several Lambda functions concurrently, put everything into a single function using Go
  • AWS responds to customers to discover how they use its services; it comes down to what customers need

Links:

  • Sam Bashton on Twitter
  • runbook.cloud
  • How We Massively Reduced Our AWS Lambda Bill with Go
  • AWS
  • AWS Lambda
  • Microsoft Clippy
  • Honeycomb
  • AWS X-Ray
  • Kubernetes
  • Simon Wardley
  • Go
  • Secrets Manager
  • DynamoDB
  • EFS
  • Digital Ocean

.

View Details

Trying to figure out if Amazon Web Services (AWS) is right for you? Use the “quadrant of doom” to determine your answer. When designing a Cloud architecture, there are factors to consider. Any system you design exists for one reason - support a business. Think about services and their features to make sure they’re right for your implementation.

Today, we’re talking to Ernesto Marquez, owner and project director at Concurrency Labs. He helps startups launch and grow their applications on AWS. Ernesto especially enjoys building serverless architectures, automating everything, and helping customers cut their AWS costs.

Some of the highlights of the show include:

  • Amazon’s level of discipline, process, and willingness to recognize issues and fix them changed the way Ernesto sees how a system should be operated
  • Specialize on a specific service within AWS, such as S3 and EC2, because there are principles that need to be applied when designing an architecture
  • Sales and Delivery Cycle: Ernesto has a conversation with a client to discuss their different needs
  • Vendor Lock-in: Customers concerned about moving application to Cloud provider and how difficult it will be to move code and design variables elsewhere
  • For every service you include in your architecture, evaluate the service within the context of a particular business case
  • Identify failure scenarios, what can go wrong, and if something goes wrong, how it’s going to be remediated
  • CloudWatching detects events that are going to happen, and you can trigger responses for those events
  • Partnering with Amazon: Companies are pushing a multi-Cloud narrative; you gain visibility and credibility, but it’s not essential to be successful
  • Can you compete against Amazon? Depends on which area you choose
  • Expand product selection to grow, focus on user experience, and improve performance to compete against Amazon
  • MiserBot: Don’t freak out about your bill because Ernesto created a Slack chatbot to monitor your AWS costs

Links:

  • Concurrency Labs
  • Ernesto Marquez on Twitter
  • How to Know if an AWS is Right for You
  • MiserBot
  • AWS
  • RDS
  • Lambda
  • Digital Ocean

.

View Details

Are you a blogger? Engineer? Web guru? What do you do? If you ask Yan Cui that question, be prepared for several different answers.

Today, we’re talking to Yan, who is a principal engineer at DAZN. Also, he writes blog posts and is a course developer. His insightful, engaging, and understandable content resonates with various audiences. And, he’s an AWS serverless hero!

Some of the highlights of the show include:

  • Some people get tripped up because they don’t bring microservice practices they learned into the new world of serverless; face many challenges
  • Educate others and share your knowledge; Yan does, as an AWS hero
  • Chaos Engineering Meeting Serverless: Figuring out what types of failures to practice for depends on what services you are using
  • Environment predicated on specific behaviors may mean enumerating bad things that could happen, instead of building a resilient system that works as planned
  • API Gateway: Confusing for users because it can do so many different things; what is the right thing to do, given a particular context, is not always clear
  • Now, serverless feels like a toy, but good enough to run production workflow; future of serverless - will continue to evolve and offer more flexibility
  • Serverless is used to build applications; DevOps/IOT teams and enterprises are adopting serverless because it makes solutions more cost effective

Links:

  • Yan Cui on Twitter
  • DAZN
  • Production-Ready Serverless
  • Theburningmonk.com
  • Applying Principles of Chaos Engineering to Serverless
  • AWS Heroes
  • re:Invent
  • Lambda
  • Amazon S3 Service Disruption
  • API Gateway
  • Ben Kehoe
  • Digital Ocean

.

View Details

Is your company thinking about adopting serverless and running with it? Is there a profitable opportunity hidden in it? Ready to go on that journey?

Today, we’re talking to Rowan Udell, who works for Versent, an Amazon Web Services (AWS) consulting partner in Australia. Versent focuses on specific practices, including helping customers with rapid migrations to the Clouds and going serverless.

Some of the highlights of the show include:

  • Australia is experiencing an increase in developers using serverless tool services and serverless being used for operational purposes
  • Serverless seems to be either a brilliant fit or not quite ready for prime time
  • Misconceptions include keeping functions warm, setting up scheduled indications
  • Simon Wardley talked about how the flow of capital can be traced through an organization that has converted to serverless
  • Concept of paying thousands of dollars up front for a server is going away
  • Spend whatever you want, but be able to explain where the money is going (dev vs. prod); companies will re-evaluate how things get done
  • Serverless is either known as an evolution or revolution; transformative to a point
  • Winding up with a large number of shops where when something breaks, they don’t have the experience to fix it; gain practical experience through sharing
  • Seek developer feedback and perform testing, but know where and when to stop
  • With serverless, you have little control of the environment; focus on automated parts you do control
  • Serverless Movement: People have opinions and want you to know them
  • Understand continuum of options for running your application in the Cloud; learn pros and cons; and pick the right tool
  • Reconciliation between serverless and containers will need to play out; changes will come at some point
  • Blockchain + serverless + machine learning + Kubernetes + service mesh = raise entire seed round

Links:

  • Rowan Udell’s Blog
  • Rowan Udell on Twitter
  • Versent on Twitter
  • Lambda
  • Simon Wardley
  • Open Guide to AWS Slack Channel
  • Kubernetes
  • Aurora
  • Digital Ocean

.

View Details

Google Cloud Platform (GCP) turned off a customer that it thought was doing something out of bounds. This led to an Internet outrage, and GCP tried to explain itself and prevent the problem in the future.

Today, we’re talking to Daniel Compton, an independent software consultant who focuses on Clojure and large-scale systems. He’s currently building Deps, a private Maven repository service. As a third-party observer, we pick Daniel’s brain about the GCP issue, especially because he wrote a post called, Google Cloud Platform - The Good, Bad, and Ugly (It’s Mostly Good).

Some of the highlights of the show include:

  • Recommendations: Use enterprise billing - costs thousands of dollars; add phone number and extra credit card to Google account; get support contract
  • Google describing what happened and how it plans to prevent it in the future seemed reasonable; but why did it take this for Google to make changes?
  • GCP has inherited cultural issues that don’t work in the enterprise market; GCP is painfully learning that they need to change some things
  • Google tends to focus on writing services aimed purely at developers; it struggles to put itself in the shoes of corporate-enterprise IT shops
  • GCP has a few key design decisions that set it apart from AWS; focuses on global resources rather than regional resources
  • When picking a provider, is there a clear winner? AWS or GCP? Consider company’s values, internal capabilities, resources needed, and workload
  • GCP’s tendency to end service on something people are still using vs. AWS never ending a service tends to push people in one direction
  • GCP has built a smaller set of services that are easy to get started with, while AWS has an overwhelming number of services
  • Different Philosophies: Not every developer writes software as if they work at Google; AWS meets customers where they are, fixes issues, and drops prices
  • GCP understands where it needs to catch up and continues to iterate and release features

Links:

  • Daniel Compton
  • Daniel Compton on Twitter
  • Google Cloud Platform - The Good, Bad, and Ugly (It’s Mostly Good)
  • Deps
  • The REPL
  • Postmortem for GCP Load Balancer Outage
  • AWS Athena
  • Digital Ocean

.

View Details

Do you deal with a lot of data? Do you need to analyze and interpret data? Veritone’s platform is designed to ingest audio, video, and other data through batch processes to process the media and attach output, such as transcripts or facial recognition data.

Today, we’re talking to Christopher Stobie, a DevOps professional with more than seven years of experience building and managing applications. Currently, he is the director of site reliability engineering at Veritone in Costa Mesa, Calif. Veritone positions itself as a provider of artificial intelligence (AI) tools designed to help other companies analyze and organize unstructured data. Previously, Christopher was a technical account manager (TAM) at Amazon Web Services (AWS); lead DevOps engineer at Clear Capital; lead DevOps engineer at ESI; Cloud consultant at Credera; and Patriot/THAAD Missile Fire Control in the U.S. Army. Besides staying busy with DevOps and missiles, he enjoys playing racquetball in short shorts and drinking good (not great) wine.

Some of the highlights of the show include:

  • Various problems can be solved with AI; companies are spending time and money on AI
  • Tasks can be automated that are too intelligent to write around simple software
  • Machine learning (ML) models are applicable for many purposes; real people with real problems and who are not academics can use ML
  • Fargate is instant-on Docker containers as a service; handles infrastructure scaling, but involves management expense
  • Instant-on works with numerous containers, but there will probably be a time when it no longer delivers reasonable fleet performance on demand
  • Decision to use Kafka was based on workload, stream-based ingestion
  • Veritone’s writes code that tries to avoid provider lock-in; wants to make an integration as decoupled as possible
  • People spend too much time and energy being agnostic to their technology and giving up benefits
  • If you dream about seeing your name up in lights, Christopher describes the process of writing a post for AWS
  • Pain Points: Newness of Fargate and unfamiliarity with it; limit issues; unable to handle large containers

Links:

  • Veritone
  • Christopher Stobie on LinkedIn
  • Building Real Time AI with AWS Fargate
  • SageMaker
  • Fargate
  • Docker
  • Kafka
  • Digital Ocean

.

View Details

Google builds platforms for developers and strives to make them happy. There's a team at Google that wakes up every day to make sure developers have great outcomes with its services and products. The team listens to the developers and brings all feedback back into Google. It also spends a lot of time all over the world talking to and connecting with developer communities and showing stuff being worked on. It doesn't do the team any good to build developer products that developers don’t love.

Today, we’re talking to Adam Seligman, vice president of developer relations at Google, where he is responsible for the global developer community across product areas. He is the ears and voice for customers.

Some of the highlights of the show include:

  • Google tackles everything in an open source way: Shipping feedback, iteration, and building communities
  • Storytelling - the Tale of Kubernetes: in a short period of time, gone from being open source that Google spearheaded to something sweeping the industry
  • Rise of containerization inside Linux Kernel is an opportunity for Google to share container management technology and philosophy with the world
  • Google Next: Knative journey toward lighter-weight serverless-based applications; and GKE On-Prem, customers and teams working with Kubernetes running on premise
  • Innovation: When logging into GCP console, you can terminate all billable resources assigned to project and access tab for building by hand
  • GCP's console development strategy includes hard work on documentation, making things easy to use, and building thoughtfulness in grouping services
  • Google is about design goals, tradeoffs, and metrics; it’s about hyper scale and global footprint of requirements, as well as supporting every developer
  • Conception 1: Google builds HyperScale Reid-Centric user partitioned apps and don't build globally consistent data driven apps
  • Conception 2: Software engineers at the top Internet companies do the code and write amazing things instantly
  • 12-Factor App: Opinions of how to architect apps; developers should have choices, but take away some cognitive and operating load complexity
  • Businesses are running core workloads on Google, which had to put atomic clocks in data centers and private fiber networking to make it all work
  • Perception that Google focuses on new things, rather than supporting what's been released; industry is on a treadmill chasing shiny things and creating noise
  • Industry needs to be welcoming and inclusive; a demand for software, apps, and innovation, but number of developers remains because everyone’s not included
  • Human vs. Technology: More investment and easier onboarding with technology and an obligation to build local communities
  • Goal: Take database complexity and start removing it for lots of use cases and simplify things for users to deal with replication, charting, and consistency issues
  • DevFest: Google has about 800 Google developer groups that do a lot of things to build local communities and write code together

Links:

  • Adam Seligman on Twitter
  • 12-Factor App
  • I Want to Build a World Spanning Search Engine on Top of GCP
  • DevFest
  • Kubernetes
  • Docker
  • Heroku
  • Google Next
  • Google Reader

.

View Details

What is serverless? What do people want it to be? Serverless is when you write your software, deploy it to a Cloud vendor that will scale and run it, and you receive a pay-for-use bill. It’s not necessarily a function of a service, but a concept.

Today, we’re talking to Nitzan Shapira, co-founder and CEO of Epsagon, which brings observability to serverless Cloud applications by using distributed tracing and artificial intelligence (AI) technologies. He is a software engineer with experience in software development, cyber security, reverse engineering, and machine learning.

Some of the highlights of the show include:

  • Modern renaissance of “functions as a service” compared to past history; is as abstracted as it can be, which means almost no constraints
  • If you write your own software, ship it, and deploy it - it counts as serverless
  • Some treat serverless as event-driven architecture where code swings into action
  • When being strategic to make it more efficient, plan and develop an application with specific and complicated functioning
  • Epsagon is a global observer for what the industry is doing and how it is implementing serverless as it evolves
  • Trends and use cases include focusing on serverless first instead of the Cloud
  • Economic Argument: Less expensive than running things all the time and offers ability to trace capital flow; but be cautious about unpredictable cost
  • Use bill to determine how much performance and flow time has been spent
  • Companies seem to be trying to support every vendor’s serverless offering; when it comes to serverless, AWS Lambda appears to be used most often
  • Not easy to move from one provider to another; on-premise misses the point
  • People starting with AWS Lambda need familiarity with other services, which can be a reasonable but difficult barrier that’s worth the effort
  • Managing serverless applications may have to be done through a third party
  • Systemic view of how applications work focuses on overall health of a system, not individual function
  • Epsagon is headquartered in Israel, along with other emerging serverless startups; Israeli culture fuels innovation

Links:

  • Epsagon
  • Email Nitzan Shapira
  • Nitzan Shapira on Twitter
  • Heroku
  • Google App Engine
  • AWS Elastic Beanstalk
  • Lambda
  • Amazon CloudWatch
  • AWS X-Ray
  • Simon Wardley
  • Charity Majors
  • Start-Up Nation
  • Digital Ocean

.

View Details

It is easy to pick apart the general premise of Cloud agnosticism being a myth. What about reasonable use cases? Well, generally, when you have a workload that you want to put on multiple Cloud providers, it is a bad idea. It’s difficult to build and maintain. Providers change, some more than others. The ability to work with them becomes more complex. Yet, Cloud providers rarely disappoint you enough to make you hurry and go to another provider.

Today, we’re talking to Jay Gordon, Cloud developer advocate for MongoDB, about databases, distribution of databases, and multi-Cloud strategies. MongoDB is a good option for people who want to build applications quicker and faster but not do a lot of infrastructural work.

Some of the highlights of the show include:

  • Easier to consider distributed data to be something reliable and available, than not being reliable and available
  • People spend time buying an option that doesn’t work, at the cost of feature velocity
  • If Cloud provider goes down, is it the end of the world?
  • Cloud offers greater flexibility; but no matter what, there should be a secondary option when a critical path comes to a breaking point
  • Hand-off from one provider to another is more likely to cause an outage than a multi-region single provider failure
  • Exclusion of Cloud Agnostic Tooling: The more we create tools that do the same thing regardless of provider, there will be more agnosticism from implementers
  • Workload-dependent where data gravity dictates choices; bandwidth isn’t free
  • Certain services are only available on one Cloud due to licensing; but tools can help with migration
  • Major service providers handle persistent parts of architecture, and other companies offer database services and tools for those providers
  • Cost may/may not be a factor why businesses stay with 1 instead of multi-Cloud
  • How much RPO and RTO play into a multi-Cloud decision
  • Selecting a database/data store when building; consider security encryption

Links:

  • Jay Gordon on Twitter
  • MongoDB
  • The Myth of Cloud Agnosticism
  • Heresy in the Church of Docker
  • Kubernetes
  • Amazon Secrets Manager
  • JSON
  • Digital Ocean

.

View Details

Trying to convince a company to embrace the theory and idea of Chaos Engineering is an uphill battle. When a site keeps breaking, Gremlin’s plan involves breaking things intentionally. How do you introduce chaos as a step toward making things better?

Today, we’re talking to Ho Ming Li, lead solutions architect at Gremlin. He takes a strategic approach to deliver holistic solutions, often diving into the intersection of people, process, business, and technology. His goal is to enable everyone to build more resilient software by means of Chaos Engineering practices.

Some of the highlights of the show include:

  • Ho Ming Li previously worked as a technical account manager (TAM) at Amazon Web Services (AWS) to offer guidance on architectural/operational best practices
  • Difference between and transition to solutions architect and TAM at AWS
  • Role of TAM as the voice and face of AWS for customers
  • Ultimate goal is to bring services back up and make sure customers are happy
  • Amazon Leadership Principles: Mutually beneficial to have the customer get what they want, be happy with the service, and achieve success with the customer
  • Chaos Engineering isn’t about breaking things to prove a point
  • Chaos Engineering takes a scientific approach
  • Other than during carefully staged DR exercises, DR plans usually don’t work
  • Availability Theater: A passive data center is not enough; exercise DR plan
  • Chaos Engineering is bringing it down to a level where you exercise it regularly to build resiliency
  • Start small when dealing with availability
  • Chaos Engineering is a journey of verifying, validating, and catching surprises in a safe environment
  • Get started with Chaos Engineering by asking: What could go wrong?
  • Embrace failure and prepare for it; business process resilience
  • Gremlin’s GameDay and Chaos Conf allows people to share experiences

Links:

  • Ho Ming Li on Twitter
  • Gremlin
  • Gremlin on Twitter
  • Gremlin on Facebook
  • Gremlin on Instagram
  • Gremlin: It’s GameDay
  • Chaos Engineering Slack
  • Chaos Conf
  • Amazon Leadership Principles
  • Adrian Cockcroft and Availability Theater
  • Digital Ocean

.

View Details

Are you about to head off to college? Interested in DevOps and the Cloud? Is there a good way for someone like you who is starting out in the world of technology to absorb the necessary skills? The Open Source Lab (OSL) at Oregon State University (OSU) is one program that helps students and serves as a career accelerator. OSL is a unicorn because OSU is willing to invest in open source.

Today, we’re talking to Lance Albertson, director of OSL at OSU. OSL does a variety of projects to provide private Clouds that are neutrally hosted on its premises. The lab also gives undergraduate students hands-on experience with DevOps skills, including dealing with configuration management, deploying applications, learning how applications deploy, working with projects, and troubleshooting issues. OSL is for any student who has a general interest or passion for it, and a willingness to learn.

Some of the highlights of the show include:

  • Workflow focuses on what students need to learn about Linux and giving access to various repos; then they experience the lab’s configuration management suite
  • Interview Process: Put out a posting, student submits an application online, each candidate is reviewed, student is given a screening quiz,
  • If a student passes the screening process, they are brought in for an in-person interview for personality and technical questions
  • Students tend to initially have the least amount of experience and most difficulty with a repository that has multiple people committing to it and dealing with PRs
  • Spinning up VMs and understanding how configuration management is connected, how services communicate, and how to set up an application
  • Round-Robins and System Sprint Meetings: Focus on discussing and documenting processes, issues, suggestions, comments, and other information
  • Younger students are mentored by Lance and the older students; every generation has to evolve because the environment and industry evolve
  • OSL made OpenStack work on POWER8, PowerPC, and PowerPC little-endian; gateway into Cloud - having OpenStack instance to offer services
  • Vast majority of OSL’s revenue comes from donations; no direct support from the university; finding companies to serve as sponsors is beneficial to all
  • Future of OSL: Providing more Cloud-like services; creating a more internal, private Cloud’ and containerized ways of running or deploying applications

Links:

  • Apache Software Foundation
  • BusyBox
  • Buildroot
  • Chef
  • Ruby
  • Freenode
  • OpenStack
  • Sphinx
  • Docker
  • Neutron
  • Seth
  • Rackspace
  • CoreOS
  • Kubernetes
  • Digital Ocean

.

View Details

Today, we’re talking to Jeff Barr, vice president and chief evangelist at Amazon Web Services (AWS). He founded the AWS Blog in 2004 and has written more than 2,900 posts for it and another 1,100 for his personal blog. As chief evangelist, Jeff strives to explain the benefits of Cloud computing and Web services to anyone who will listen.

Jeff is the voice of AWS. He does what he does best - exploits his superpower of explaining technology in ways that people can understand it. Jeff tries to be the same person all the time. He loves to meet people and go out of his way to say “Hello.” So, if you see him at re:Invent, say “Cheese” and take a selfie with him!

Some of the highlights of the show include:

  • Jeff uses AWS Workspaces for his blog; one of Jeff’s blogging principles is to not take anybody else's word for anything to the absolute best of his technical ability
  • Zero Client: Jeff has no rotating hardware, disk drives, just a zero client; wherever he is, it's the same workspace
  • AWS has something for everyone; it build things in response to customers’ questions, requests, and feedback
  • Naming Services and Products: Is it helpful? Is it descriptive? Does it have any hidden meanings?
  • Amazonian DNA and Dog Friendly Workspace: Jeff went from super fearful to accepting, to now thinking of dogs as incredible creations because they add fun and excitement to the office
  • As part of hiring, each interviewer is assigned Amazon leadership principles (LPs) to ask questions that measure a candidate against those LPs
  • What is the secret to getting hired at Amazon? Study the LPs to understand what they're about and be able to express your philosophies and history with LPs
  • re:Invent makes sure customers understand services - What is it? What does it do? How do they put it to work? What are the best use cases for it?
  • Things can never be too simple; you start from zero, put a lot of different things in there, and then you need the feedback to build in simplicity
  • AWS is following a more on-demand approach than traditional reserve instances; it opens the door to being used in a lot of ways
  • AWS does a lot of work before a launch to make sure it’s got infrastructure, scaling, monitoring, and capacity in place
  • If you are a customer, talk to AWS and let them know what they're doing right or wrong; write a blog post, tweet about it, share it with them in some way
  • Is the breadth of product offerings from AWS too vast? Is it offering too many things?
  • AWS was not explicit about where it was going with Cloud computing or do analyses or projections about it; it simply launched SQS and let it speak for itself
  • Customer feedback shapes what Amazon works on; customers share and then AWS re-prioritizes to make sure it’s delivering the right thing at the right time
  • Remember: It's not just bits and bytes, it's about the organic life form

Links:

Jeff Barr on Twitter

Jeff Barr on LinkedIn

AWS

AWS Blog

Jeff Barr’s Blog

Amazon Machine Images

Zero Client

AWS Workspaces

AWS Lambda

Amazon Leadership principles

re:Invent

The Robot Uprising Will Have Very Clean Floors

Serverlessly Storing My Dad Jokes in a Dadabase

Days Until re:Invent

.

View Details

Some companies that offer services expect you to do things their way or take the highway. However, Google expects people to simply adapt the tech company’s suggestions and best practices for their specific context. This is how things are done at Google, but this may not work in your environment.

Today, we’re talking to Liz Fong-Jones, a Senior Staff Site Reliability Engineer (SRE) at Google. Liz works on the Google Cloud Customer Reliability Engineering (CRE) team and enjoys helping people adapt reliability practices in a way that makes sense for their companies.

Some of the highlights of the show include:

  • Liz figures out an appropriate level of reliability for a service and how a service is engineered to meet that target
  • Staff SRE involves implementation, and then identifying and solving problems
  • Google’s CRE team makes sure Google Cloud customers can build seamless services on the Google Cloud Platform (GCP)
  • Service Level Objectives (SLOs) include error budgets, service level indicators, and key metrics to resolve issues when technology fails
  • Learn from failures through instant reports and shared post-mortems; be transparent with customers and yourself
  • GCP: Is it part of Google or not? It’s not a division between old and new.
  • Perceptions and misunderstandings of how Google does things and how it’s a different environment
  • Google’s efforts toward customer service and responsiveness to needs
  • Migrating between different Cloud providers vs. higher level services
  • How to use Cloud machine learning-based products
  • GCP needs to focus on usability to maintain a phase of growth
  • Offer sensible APIs; tear up, turn down, and update in a programmatic fashion
  • Promotion vs. Different Job: When you’ve learned as much as you can, look for another team to teach something new
  • What is Cloud and what isn’t? Cloud deployments require SRE to be successful but SREs can work on systems that do not necessarily run in the Cloud.

Links:

  • Cloud Spanner
  • Kubernetes
  • Cloud Bigtable
  • Google Cloud Platform blog - CRE Life Lessons
  • Google SRE on YouTube

.

View Details

What’s serverless? Are you serverless now? Is going from enterprise to serverless a natural evolution? Or, is it a “that was fun, now let’s go ride our bikes” moment? Is serverless “just a toy?” Is it a wide and varied ecosystem, or is it Lambda plus some other randos? What's up with serverless vs. containers?

Today, Forrest Brazeal is here to answer those questions and discuss pros and cons of serverless. He was a senior Cloud architect prior to joining Trek10. Forrest spent several years leading AWS and serverless engineering projects at Infor. He understands the challenges faced by enterprises moving to the Cloud and enjoys building solutions that provide maximum business value at a minimal cost.

Some of the highlights of the show include:

  • Bimodality: Backend development going away and being replaced by managed services; undifferentiated items are being moved to the Cloud
  • Serverless is application designs with “Backend as a Service” (BaaS) and/or “Functions as a Service” (FaaS) platforms; everything is managed for you
  • AWS Lambda: Is it today’s trend or a bias that everyone is using it; Lambda makes up 80% of current FaaS adoption
  • Serverless Ecosystem: You can build it however you want, and you’re doing it right; but don’t take that at face-value; no two Lambda environments are alike
  • Cloud services at this scale have not been knitted together to form applications that are serving major workloads; best practices need to be established
  • Native Cloud providers will consolidate, and individual frameworks will be created with components of application stacks tied together to build systems
  • Serverless vs. Containers: No need for disparity - we can learn to get along; people use containers because it is easier than going serverless
  • Serverless Heroes series features people thinking out-of-the-box and helps identify emerging trends; serverless is growing, and it’s not just about startups
  • Went from working with a Sharpie to Procreate for the FaaS and Furious cartoon series; serverless component of process is for invoicing
  • Changes? Packaging to handle sharing; more knobs on console; unified process needed because too many building own workflow and tooling
  • Certification: Proof-positive that you know what you’re talking about or is it questionable value if not backing up expertise in the real world?

Links:

  • Forrest Brazeal on Twitter
  • Invoiceless
  • Summon the vast power of certification - Dilbert cartoon
  • Trek10 blog
  • A Cloud Guru ThinkfaaS podcast
  • A Cloud Guru - Serverless Superheros
  • Why We’re Excited About AWS AppSync
  • Serverless Architectures with Mike Roberts
  • AWS Lambda
  • AWS Serverless Application Model (SAM)
  • Procreate
  • AWS Certified Cloud Practitioner
  • Serverlessconf
  • Digital Ocean

.

View Details

DevOps as a service describes what Reactive Ops is trying to do, who it’s trying to help, and what problems it’s trying to solve. It’s passion to deliver service where human beings help other human beings is done through a group of engineers who are extremely good at solving problems.

Sarah Zelechoski is the vice president of engineering at Reactive Ops, which defines the world’s problems and solves them by pouring Kubernetes on top of them. The team focuses on providing expert-level guidance and a curated framework using Kubernetes and other open source tools. Sarah's greatest passion is helping others, which encompasses advocating for engineers and rekindling interest in the lost art of service in the tech space.

Some of the highlights of the show include:

  • Kubernetes is changing the way people work; it offers a way to release a product, provide access to it, and behaviors when you deploy it
  • Any person/business can use Kubernetes to mold their workflow
  • Kubernetes is complex and has sharp edges; it has only recently become productive because of its community finding and reporting issues
  • Business value of deploying Kubernetes to a new environment: Flexibility and uniform system of management; and it can provide a context shift
  • Implementation Challenges with Workshops/Tutorials: Valuable entry level strategy for people learning Kubernetes; but the translation is not easy
  • About 85% of the work Reactive Ops does is helping its customers get on to Kubernetes is spent on application architecture
  • If thinking about moving to Kubernetes, how well will your current applications translate? Do you want to start over from scratch?
  • Value in paying someone to do something for you
  • Using Defaults: Try initially until you realize what you need; Kubernetes gives you options, but it’s a challenging path to go from defaults to advanced
  • Deploying a workload between all major Cloud providers is possible, but there are challenges in managing multiple regions or locations
  • Cluster Ops: Managed Kubernetes clusters where Reactive Ops stays on the map, watches them, and puts them on pager, so you can continue your work without having to worry

Links:

  • Sarah Zelechoski on Twitter
  • Reactive Ops
  • Kubernetes
  • GKE from GCB
  • AKS from Azure
  • EKS from AWS
  • Kops
  • Terraform
  • Slack

.

View Details

Are you interested in going beyond basic monitoring and visibility? Need tools to build and operate serverless applications and extract business intelligence? IOpipe provides extended visibility and metrics around AWS Lambda, including profiling, core dumps, and incoming input events.

Today, we’re talking to Erica Windisch, who is the founder and CTO of IOpipe. She brings her experience in building developer and operational tooling to serverless applications. Erica also has more than 17 years of experience designing and building Cloud infrastructure management solutions. She was an early and longtime contributor to OpenStack and maintainer of the Docker project.

Some of the highlights of the show include:

  • Nomenclature Battle: Serverless vs. stateless
  • Building a window of visibility into Lambda: Talking to users and assessing needs/pain points
  • Observability of the infrastructure: Necessary evil to get to automated healing
  • Using Lambda at significant levels of scale; some companies grow usage, others go all in right away
  • Current state of Lambda ecosystem
  • Is Lambda stable? Indications and no formal SLA
  • How issues manifest and are exposed
  • Trends include cold starts, hours-long failures, and multiple function evokes
  • Infrastructure powering IOpipe: Lambda issues may impact performance of monitoring system, but IOpipe is not necessarily dependent on Lambda
  • Future of Lambda: Builds applications a specific way, but there are limitations
  • What would Erica change about Lambda? Run function and define handlers
  • Lambda functions can be difficult to understand; some developers do not have familiarity and create bottlenecks
  • Capacity limits around Lambda can be difficult to establish

Links:

  • Erica Windisch on Twitter
  • Erica Windisch on Twitch
  • IOpipe
  • 12-Factor App
  • Cloud Custodian in Lambda
  • Velocity London
  • ServerlessConf London
  • re:Invent
  • AWS Glue

.

View Details

Let’s chat about the Cloud and everything in between. The people in this world are pretty comfortable with not running physical servers on their own, but trusting someone else to run them. Yet, people suffer from the psychological barrier of thinking they need to build, design, and run their own monitoring system. Fortunately, more companies are turning to Datadog.

Today, we’re talking to Ilan Rabinovitch, Datadog’s vice president of product and community. He spends his days diving into container monitoring metrics, collaborating with Datadog’s open source community, and evangelizing observability best practices. Previously, Ilan led infrastructure and reliability engineering teams at various organizations, including Ooyala and Edmunds.com. He’s active in the open source and DevOps communities, where he is a co-organizer of events, such as SCALE and Texas Linux Fest.

Some of the highlights of the show include:

  • Datadog is well-known, especially because it is a frequent sponsor
  • More organizations know their core competency is not monitoring or managing servers
  • Monitoring/metrics is a big data problem; Datadog takes monitoring off your plate
  • Alternate ways, other than using Nagios, to monitor instances and regenerate configurations
  • Datadog is first to identify patterns when there is a widespread underlying infrastructure issue
  • Trends of moving from on-premise to Cloud; serverless is on the horizon
  • How trends affect evolution of Datadog; adjusting tools to monitor customers’ environments
  • Datadog’s scope is enormous; the company tries to present relevant information as the scale of what it’s watching continues to grow
  • Datadog’s pricing is straightforward and simple to understand; how much Cloud providers charge to use Datadog is less clear
  • Single Pane of Glass: Too much data to gather in small areas (dashboards)
  • Why didn’t monitoring catch this? Alerts need to be actionable and relevant
  • How to use Datadog’s workflow for setting alerts and work metrics
  • Datadog’s first Dash user conference will be held in July in New York; addresses how to solve real business problems, how to scale/speed up your organization

Links:

  • Ilan Rabinovitch on Twitter
  • Datadog
  • Docker Adoption Survey Results
  • Rubric for Setting Alerts/Work Metrics
  • Dash Conference
  • re:Invent
  • Nagios

.

View Details

Do you need data captured that let you know when things don’t look quite right? Need to identify issues before they become major problems for your organization? Turn to Threat Stack, which has Cloud issues of its own, and helps its customers with their Cloud issues.

Today, I’m talking to Pete Cheslock, who runs technical operations at Threat Stack, which handles security monitoring, alerting, and remediation. The company uses Amazon Web Services (AWS), but its customer base can run anywhere.

Some of the highlights of the show include:

  • Challenges Threat Stack experienced with AWS and how it dealt with them
  • Threat Stack helps companies improve their security posture in AWS
  • Security shouldn’t be an issue, if providers do their job; shared responsibility
  • Education is needed about what matters regarding security, avoiding mistakes
  • Cloud is still so new; not many people have abroad experience managing it
  • Scanning customer accounts against best practices to identify risks
  • Threat Stack’s scanning tool is worthwhile, but most tools lack judgement and perspective
  • Threat Stack offers context between host- and Cloud-based events; tying data together is the secret sauce
  • You shouldn’t have to pay a bunch of money to have a robust security system
  • Good operations is good security; update, patch, track, and perform other tasks
  • Lack of validation about what services are going to be a successful or not
  • Vendor Lock-in: Understand your choices when building your system
  • Pervasiveness and challenge of containerization and Kubernetes
  • Cloud reduces cycle time and effort to bring a product to market
  • Amazon is a game changer with what it allows you to do and solve problems

Links:

  • Pete Cheslock
  • Digital Ocean
  • Threat Stack
  • AWS
  • re:Invent
  • Kubernetes

.

View Details

Aurora, from Amazon Web Services (AWS), is a MySQL-compatible service for complex database structures. It offers capabilities and opportunities. But with Aurora, you’re putting a lot of trust in AWS to “just work” in ways not traditional to relational database services (RDS).

David Torgerson, Principal DevOps Engineer at Lucidchart, is a mystery wrapped in an enigma and virtually impossible to Google. He shares Lucidchart’s experience with migrating away from a traditional RDS to Aurora to free up developer time.

Some of the highlights of the show include:

  • Trade off of making someone else partially responsible for keeping your site up
  • Lucidchart’s overall database costs decreased 25% after switching to Aurora
  • Aurora unknowns: What is an I/Op in Aurora? When you write one piece of data, does it count as six I/Ops?
  • Multi-master Aurora is coming for failover time and disaster recovery purposes
  • Aurora drawbacks: No dedicated DevOps, increased failover time, and misleading performance speed
  • Providers offer ways to simplify your business processes, but not ways to get out of using their products due to vendor and platform lock-in
  • Lucidchart is skeptical about Aurora Serverless; will use or not depending on performance

Links:

  • Corey's architecture diagram on AWS
  • Lucidchart
  • Lucidchart’s Data Migration to Amazon Aurora
  • Preview of Amazon Aurora Multi-master Sign Up
  • This is My Architecture
  • re:Invent
  • Digital Ocean

.

View Details

Does your job challenge and motivate you? Does it utilize your skills? Or, are you ready to go job hunting? Do you want an awesome job that is a resume booster? Companies should be supportive of their employees finding a job that matches their skills and interests. Also, when hiring, companies should offer thoughtful processes for interviews.

Today, I’m talking to Sarah Withee, a polyglot software engineer, mentor, teacher, and robot tinkerer. Sarah went job hunting, and after several job interviews, she finally found a job that made her super happy at Arcadia Healthcare Solutions. Sarah compares the interview processes she experienced at big name tech companies that offer Cloud services.

Some of the highlights of the show include:

  • Companies sometimes lose sight that even interview interactions need to be a two-way sale
  • Interviews often involve talking to many people; and if several are bad, that forms a negative impression of the company
  • Companies need to provide interview training and follow the same standards
  • Don’t farm out challenging or unfamiliar issues when interviewing candidates
  • Sarah is very competent, but she is new to Cloud platforms; she is like a sponge, who enjoys learning and having a bare knowledge of new technology
  • How HIPAA regulations impact Sarah’s learning and software engineering work; she has to be more aware of security and safety of healthcare data
  • Being a teacher and mentor affects how Sarah learns new things; everybody learns slightly differently
  • In the Cloud space, know which direction you want to go and start with simpler things to learn the basics; focus on what is relevant to what you are working on

Links:

  • Sarah Withee on Twitter #speakerconfessions
  • Sarah Withee on Twitter
  • Sarah Withee Blog
  • Sarah Withee Resume
  • Digital Ocean
  • AWS
  • Azure

.

View Details

Docker went from being a small startup to an enterprise company that changed the way people think about their infrastructure to now, where its relevance is somewhat minimal. The conversation is no longer around the container level. Docker has become commonplace.

Today, we’re talking to Jérôme Petazzoni, formerly of Docker. While he was with the company for about 8 years, Docker definitely experienced a roller coaster ride.  

Some of the highlights of the show include:

  • Amount of work conducted on the enterprise vs. community editions
  • Docker was so widely adopted because its core technology was open source
  • Challenge is to build a viable business and revenue model for the long run
  • Similarities between Docker and Red Hat open source platforms
  • Docker went from six people working in a garage to having a few hundred employees and $1.3 billion valuation
  • Changes happened, but they were gradual; the changes were necessary to be a profitable and sustainable company
  • Contingent of internal and external people believed that Docker was the answer for whatever problem surfaced; Docker would save you, but not always
  • Balancing Act: Pushing forward with a correct message and regulating enthusiasm
  • Networking and Docker for dummies; confusion and problems of things not working as expected have been resolved
  • Things will continue to shift; Kubernetes and the orchestration battle
  • What was unthinkable, could happen by companies pushing the envelope and making progress
  • Will who you have as your Cloud provider stop mattering? It depends.
  • All major Cloud providers plan to offer managed Kubernetes services and what Jérôme thinks of them
  • Jérôme’s opinion on whether Kubernetes will follow this same path as Docker
  • What does the road ahead look like for infrastructure automation? There is potential and lots of best practices in Cloud environments.

Links:

  • Jérôme Petazzoni on Twitter
  • https://jpetazzo.github.io/
  • Docker Crunch Base
  • Digital Ocean
  • Red Hat
  • Corey's Heresy in the church of docker talk
  • Kubernetes
  • ZooKeeper
  • Azure

.

View Details

Like migrating caribou, you tend to follow the trends of what clients are doing, which dictates what you work on as a consultant.

Today, we’re talking to Lynn Langit, an independent Cloud architect. She is an AWS Community Hero, Google Cloud developer expert, and former Microsoft MVP. Lynn is a lifelong learner, and she has worked broad and deep across all three large providers. These days, she works mostly with Google Cloud and AWS, rather than Azure, because that’s what her clients are using.

Some of the highlights of the show include:

  • Differences between the West Coast and global use of Cloud
  • Education is key; Lynn is th co-founder of Teachingkidsprogramming.org
  • Lynn helped create curriculum and resources for school-age children; even her young daughter taught classes on how to code
  • Training for teachers was also needed, so TKP Labs was formed to offer fee-based teacher and developer training
  • Lynn started with classroom training, but has transitioned to online learning
  • Lynn is focusing on Big Data projects and using tools to solve real-world problems
  • Pre-processing and batching data, but not streaming it
  • AWS, Azure, and Google Cloud are all coming out with Big Data-oriented tools
  • Companies need to understand when the market is ready to accept a new paradigm; in the data world, change is more slow than in the programming world
  • If you touch a database and get burned, you are not willing to use it again; or you may have never tried to archive your data; hire a consultant to help you
  • Machine learning APIs give customers value quickly; review them before building custom models
  • Migrating data can be a costly project and restricts where the data lives
  • As Cloud proliferates, how will that impact technical education? Lynn’s Cloud for College Students to the rescue!
  • Shift from interactive to unidirectional, one-to-many learning styles; the Cloud is ready for serverless, but education is not ready for teacherless
  • Road that many of us walked to get to technical skills no longer exists; how to become a modern technologist
  • Ageism: By age 40, you are considered a manager or useless; don’t be afraid to learn something new

Links:

  • Digital Ocean
  • AWS Community Hero
  • Microsoft Azure
  • Teachingkidsprogramming.org
  • Digigirlz
  • TKP Labs
  • Lynn Langit on Lynda.com
  • Commonwealth Scientific and Industrial Research Organisation
  • Google BigQuery
  • Amazon Athena
  • AWS Glue
  • Cloud Dataflow
  • Cloud Dataprep
  • Lambda
  • Amazon EC2
  • Learn Python the Hard Way

.

View Details

Microsoft has experienced a renaissance. By everything that we've seen coming out of Microsoft over the past few years, it feels like the company is really walking the walk. Instead of just talking about how it’s innovative, it’s demonstrating that. Microsoft has been on an amazing journey, making the progression from telling customers what they need to listening to them and responding by building what they ask for.

Today, we’re talking to Corey Sanders, Corporate Vice President of Azure Compute at Microsoft.

Some of the highlights of the show include:

  • Customers are asking for Microsoft to help them through support and enabling platforms
  • Storytelling efforts through advocates, who play a double role – engaging and defending Microsoft
  • Customers moving to the Cloud are focused on a continuum and progression; they have stuff to move from one location to another and want all the benefits–better agility, faster startup time, etc.
  • Virtual serial console into existing VMs; this is how people are using this and Microsoft is going to, if not encourage this behavior, at least support it
  • Microsoft is the only Cloud with a single-instance SLA
  • Serial consoles: Windows' has seen less usage, partly due to operational aspects of Windows vs. Linux. It's not a GUI; it's scripting.
  • Does the operating system matter? From a Cloud perspective, it shouldn't have to matter; you should be able to deploy it the way you want
  • Edge enables much more complex and segregated scenarios; that combination with cognitive searches running locally will make it accessible anywhere
  • Branding challenge as customers start to notice that devices are smarter and more complex; will they lose awareness that Microsoft Azure is powering most of these things - they shouldn’t care
  • An awareness of not just what's possible, but what's coming; the democratization of AI
  • Education and fear gap of trying something new and taking that first step; make products and services stupid and simple to use
  • Customers return to add cognitive services and AI capabilities to existing, running deployments, environments, and applications
  • Multi-Cloud solutions can be successful, but there's a caveat; they’re actually built on a service-by-service perspective
  • Azure Stack, offers consistency, but some people may place blame on it for poor data center management practices; some expectations and regulations may be frustrating to some customers, but lets Microsoft offer a consistent experience
  • Freedom and flexibility have been challenges for Microsoft and other products for private Clouds
  • What people need to understand about Azure, including from a durability and reliability experience
  • To some extent, scale becomes a necessary prerequisite for some applications
  • Microsoft has taken many steps and is the leader in various areas

Links:

  • ReactiveOps
  • Microsoft Azure
  • Corey Sanders on Twitter
  • The Robot Uprising Will Have Very Clean Floors
  • Kubernetes
  • Cassandra
  • Azure Stack

.

View Details

Have you dabbled with IT infrastructure in AWS? Have you been through the process of AWS partnership? Does being an AWS partner add value? Amazon seeks partners that helps drive its business, goals, and value.

Today, we’re talking to Justin Brodley, the vice president of Cloud engineering at Ellie Mae. He has been through the AWS partnership process and shares his thoughts about it. He encourages you to find the right partner for your business!

Some of the highlights of the show include:

  • Different levels and types of AWS partnerships
  • Shakedown vs. opportunity method for new leads; lead generation expectations
  • Amazon’s improvements eroding business models
  • Partners trying to pivot, but not exclusive to AWS
  • Whether to invest in multi-Cloud
  • Amazon can’t scale its sales team to handle everybody; views partner program as an extension of its salesforce
  • Your company is important and you’re spending a lot of money, but Amazon may not care about you; partner market fills that gap and makes you feel important
  • Corporate prisoner’s dilemma: Your tech company offers something that Amazon doesn’t; but what about when Amazon does offer it?
  • Competitors’ horizontal move to become more diversified
  • Amazon expects partners to offer products and services that it cannot offer yet
  • If partners fail, Amazon decides to do it and do it better
  • Is Amazon’s best interest geared toward its partners or you and your customers?
  • Amazon needs to give incentives and support partners

Links:

  • Justin Brodley on Twitter
  • Brodley Group
  • Ellie Mae
  • Digital Ocean
  • AWS Partner Network
  • Lambda
  • API Gateway
  • AWS re:Invent
  • Salesforce
  • Azure
  • Rackspace

.

View Details

Monitoring in the entire technical world is terrible and continues to be a giant, confusing mess. How do you monitor? Are you monitoring things the wrong way? Why not hire a monitoring consultant!

Today, we’re talking to monitoring consultant Mike Julian, who is the editor of the Monitoring Weekly newsletter and author of O’Reilly’s Practical Monitoring. He is the voice of monitoring.

Some of the highlights of the show include:

  • Observability comes from control theory and monitoring is for what we can anticipate
  • Industry’s lack of interest and focus on monitoring
  • When there’s an outage, why doesn’t monitoring catch it?” Unforeseen things.
  • Cost and failure of running tools and systems that are obtuse to monitor
  • Outsource monitoring instead of devoting time, energy, and personnel to it
  • Outsourcing infrastructure means you give up some control; how you monitor and manage systems changes when on the Cloud
  • CloudWatch: Where metrics go to die
  • Distributed and Implemented Tracing: Tracing calls as they move through a system
  • Serverless Functions: Difficulties experienced and techniques to use
  • Warm vs. Cold Start: If a container isn't up and running, it has to set up database connections
  • Monitoring can't fix a bad architecture; it can't fix anything; improve the application architecture
  • Visibility of outages and pain perceived; different services have different availability levels

Links:

  • Mike Julian
  • Monitoring Weekly
  • Copy Construct on Twitter
  • Baron Schwartz on Twitter
  • Charity Majors on Twitter
  • Redis
  • Kubernetes
  • Nagios
  • Datadog
  • New Relic
  • Sumo Logic
  • Prometheus
  • Honeycomb
  • Honeycomb Blog
  • CloudWatch
  • Zipkin
  • X-Ray
  • Lambda
  • DynamoDB
  • Pinboard
  • Slack
  • Digital Ocean

.

View Details

How many of you are considered heroes? Specifically, in the serverless Cloud, Twitter, and Amazon Web Services (AWS) communities? Well, Ben Kehoe is a hero.

Ben is a Cloud robotics research scientist who makes serverless Roombas at iRobot. He was named an AWS Community Hero for his contributions that help expand the understanding, expertise, and engagement of people using AWS.

Some of the highlights of the show include:

  • Ben’s path to becoming a vacuum salesman
  • History of Roomba and how AWS helps deliver current features
  • Roombas use AWS Internet of Things (IoT) for communication between the Cloud and robot
  • Boston is shaping up to be the birthplace of the robot overlords of the future
  • AWS IoT is serverless and features a number of pieces in one service
  • Robot rising of clean floors
  • AWS Greengrass, which deploys runtimes and manages connections for communication, should not be ignored
  • Creating robots that will make money and work well
  • Roomba’s autonomy to serve the customer and meet expectations
  • Robots with Cloud and network connections
  • Competitive Cloud providers were available, but AWS was the clear winner
  • Serverless approach and advantages for the intelligent vacuum cleaner
  • Future use of higher-level machine learning tools
  • Common concern of lock-in with AWS
  • Changing landscape of data governance and multi-Cloud
  • Preparing for migrations that don’t happen or change the world
  • Data gravity and saving vs. spending money

Links:

  • Ben Kehoe on YouTube
  • AWS
  • AWS Community Hero
  • AWS IoT
  • Ben Kehoe on Twitter
  • iRobot
  • AWS Greengrass
  • Shark Cat
  • Medium
  • Boston Dynamics
  • AWS Lambda
  • AWS SageMaker
  • AWS Kinesis
  • Google Cloud Platform Spanner
  • Kubernetes
  • Digital Ocean

.

View Details

How are companies evolving in a world where Cloud is on the rise? Where Cloud providers are bought out and absorbed into other companies?

Today, we’re talking to Nell Shamrell-Harrington about Cloud infrastructure. She is a senior software engineer at Chef, CTO at Operation Code, and core maintainer of the the Habitat open source product. Nell has traveled the world to talk about Chef, Ruby, Rails, Rust, DevOps, and Regular Expressions.

Some of the highlights of the show include:

  • Chef is a configuration management tool that handles instance, files, virtual machine container, and other items.
  • Immutable infrastructure has emerged as the best of practice approach.
  • Chef is moving into next gen through various projects, including one called, Compliance - a scanning tool.
  • Some people don’t trust virtualization.
  • Habitat is an open source project featuring software that allows you to use a universal packaging format.
  • Habitat is a run-time, so when you run a package on multiple virtual machines, they form a supervisor ring to communicate via leader/follower roles.
  • Deploying an application depends on several factors, including application and infrastructure needs.
  • It is possible to convert old systems with old deployment models to Habitat.
  • Habitat allows you to lift a legacy application and put it into that modern infrastructure without needing to rewrite the application.
  • You can ease in packages to Habitat, and then have Habitat manage pieces of the application.
  • Habitat is Cloud-agnostic and integrates with public and private Cloud providers by exporting an application as a container.
  • Chef is one of just a few third-party offerings marketed directly by AWS.
  • From inception to deployment, there is a place for large Cloud providers to parlay into language they already speak.
  • Operation Code is a non-profit that teaches software engineer skills to veterans. It helps veterans transition into high-paying engineering jobs.
  • The technology landscape is ever changing. What skills are most marketable?
  • Operation Code is a learning by experience type of organization and usually starts people on the front-end to immediately see results.

Links:

Nell Shamrell-Harrington

Nell Shamrell-Harrington on Twitter

Nell Shamrell-Harrington on GitHub

Operation Code

Chef

Ruby on Rails

Rust

Regular Expressions

Habitat

AWS

Kubernetes

Docker

LinkedIn Learning

GorillaStack (use discount code: screaming)

.

View Details

Open source activism tends to focus on running on hardware you can trust and avoiding Cloud computing. The problem with some Cloud providers has to do with a conflict of interest between serving customers and how they generate revenue. It’s important for the customer to have control of their computer and their data in the Cloud. But what about their security and privacy?

Today, we’re talking to Kyle Rankin, chief security officer at Purism and writer for Linux Journal. He is a Linux expert who decided to work at Purism because of the company’s belief in free software and the Linux community.

Some of the highlights of the show include:

  • Cloud providers have faced challenges when it comes to data privacy and who owns what.
  • The word “Cloud” is overloaded, and it is unclear who is in control.
  • Cloud providers can sabotage efforts to make programs work together.
  • Cloud providers may not troll through data and exploit it. Yet, they develop tools for customers to be able to do that.
  • Even though Linux Journal stopped being printed and went digital, and was going under, it’s now back and taking a new approach.
  • What matters to new readers and Linux users is now different than what was important to original readers.
  • The more time you can spend to understand what’s happening behind the scenes will make you much more marketable and adaptable.
  • Kyle explains whether Amazon Linux is becoming a viable concern and if distribution matters anymore. Now, it’s about running an application, not thinking about what it’s running on.
  • Are there gangs of Cloud users? Do people look down on Azure users? The target is always moving and changing.
  • Check out Kyle’s book, Linux Hardening in Hostile Networks: Server Security from TLS to Tor.

Links:

  • Kyle Rankin on Twitter
  • Purism
  • Kyle Rankin’s book - Linux Hardening in Hostile Networks: Server Security from TLS to Tor
  • Linux Journal 2.0 FAQ
  • GorillaStack (use “screaming” for discount)

.

View Details

How do you encourage businesses to pick Google Cloud over Amazon and other providers? How do you advocate for selecting Google Cloud to be successful on that platform? Google Cloud is not just a toy with fun features, but is a a capable Cloud service.

Today, we’re talking to Seth Vargo, a Senior Staff Developer Advocate at Google. Previously, he worked at HashiCorp in a similar advocacy role and worked very closely with Terraform, Vault, Consul, Nomad, and other tools. He left HashiCorp to join Google Cloud and talk about those tools and his experiences with Chef and Puppet, as well as communities surrounding them. He wants to share with you how to use these tools to integrate with Google Cloud and help drive product direction.

Some of the highlights of the show include:

  • Strengths related to Google Cloud include its billing aspect. You can work on Cloud bills and terminate all billable resources. The button you click in the user interface to disable billing across an entire project and delete all billable resources has an API. You can build a chat bot or script, too. It presents anything you’ve done in the Consul by clicking and pointing, as well as gives you what that looks like in code form.
  • You can expose that from other people’s accounts because turning off someone else’s Website as a service can be beneficial. You can invite anyone with a Google account, not just ‘@gmail.com’ but ‘@’ any domain and give them admin or editor permissions across a project. They’re effectively part of your organization within the scope of that project. For example, this feature is useful for training or if a consultant needs to see all of your different clients in one dashboard, but your clients can’t see each other.
  • Google is a household name. However, it’s important to recognize that advocacy is not just external advocacy, there’s an internal component to it. There’s many parts of Google and many features of Google Cloud that people aren’t aware of. As an advocate, Seth’s job is to help people win.
  • Besides showing people how they can be successful on Google Cloud, Seth focuses on strategic complaining. He is deeply ingrained in several DevOps and configuration management communities, which provide him with positive and negative feedback. It’s his job to take that feedback and convert it into meaningful action items for product teams to prioritize and put on roadmaps. Then, the voice of the communities are echoed in the features and products being internally developed.
  • Amazon has been in the Cloud business for a long time. What took Google so long? For a long time, Google was perceived as being late to the party and not able to offer as comprehensive and experienced services as Amazon. Now, people view Google Cloud as not being substandard, but not where serious business happens. It’s a fully feature platform and it comes down to preferences and pre-existing features, not capability.
  • Small and mid-size companies typically pick a Cloud provider and stick with their choice. Larger companies and enterprises, such as Fortune 50 and Fortune 500 companies, pick multiple Clouds. This is usually due to some type of legal compliance issues, or there are Cloud providers that have specific features.
  • Externally at Google, there is the Deployment Manager tool at cloud.google.com. It’s the equivalent of CloudFormation, and teams at Google are staffed full time to perform engineering work on it. Every API that you get by clicking a button on cloud.google.com are viewing the API Docs accessible via the Deployment Manager.
  • Google Cloud also partners with open source tools and corresponding companies. There are people at Google who are paid by Google who work full time on open source tools, like Terraform, Chef, and Puppet. This allows you to provision Google Cloud resources using the tools that you prefer.
  • According to Seth, there’s five key pillars of DevOps: 1) Reduce organizational silos and break down barriers between teams; 2) Accept failures; 3) Implement gradual change; 4) Tooling and automation; and 5) Measure everything.
  • Think of DevOps as an interface in programming language, like Java, or a type of language where it doesn’t actually define what you do, but gives you a high level of what the function is supposed to implement.
  • With the SRE discipline, there’s a prescribed way for performing those five pillars of DevOps. Specific tools and technologies used within Google, some of which are exposed publicly as part of Google Cloud, enable the kind of DevOps culture and DevOps mindset that occur.
  • A reason why Google offers abstract classes in programming is that there’s more than one way to solve a problem, and SRE is just one of those ways. It’s the way that has worked best for Google, and it has worked best for a number of customers that Google is working with. But there are some other ways, too. Google supports those ways and recognizes that there isn’t just one path to operational success, but many ways to reach that prosperity.
  • The book, Site Reliability Engineering, describes how Google does SRE, which tried to be evangelized with the world because it can help people improve operations. The flip side of that is that organizations need to be cognizant of their own requirements.
  • Google has always held up along several other companies as a shining beacon of how infrastructure management could be. But some say there’s still problems with its infrastructure, even after 20-some years and billions invested.
  • Every company has problems, some of them technical, some cultural. Google is no exception. The one key difference is the way Google handles issues from a cultural perspective. It focuses on fixing the problem and making sure it doesn’t happen again. There’s a very blameless culture.
  • Conferences tend to include a lot of hand waving and storytelling. But as an industry, more war stories need to be told instead of pleasure stories. Conference organizers want to see sunshine and rainbows because that sells tickets and makes people happy. The systemic problem is how to talk about problems out in the open.
  • Becoming frustrated and trying to figure out why computers do certain things is a key component of the SRE discipline referred to as Toil - work tied to systems that either we don’t understand or don’t make sense to automate.
  • Those going to Google Cloud to ‘move and improve’ tend to be a mix of those from other Cloud providers and those from on-premise data center deployments. Move and improve is where there are VMs in a data center, and they need to be moved to the Cloud.
  • There are tiny differences around the Cloud-native paradigm and providers. There’s some key pillars: Does it handle restarts well? Is it highly available? Can it be containerized, even though containers aren’t necessarily required for Cloud native? Does it package all of its dependencies with it? Can it run on different operating systems? All of these things are generic, they’re not specific to a Cloud provider.

Links:

Google Cloud and blog

Amazon Web Services

HashiCorp

Terraform

Vault

Consul

Nomad

Chef

Puppet

Kubernetes

AutoML

Monitorama

Azure

CloudFormation

Ansible

Elk Stack

Site Reliability Engineering book for O’Reilly

Fastly

Hacker News

Cloud Foundry

Microsoft Cloud

Alibaba Cloud

Lambda

Quotes by Seth:

“Everything we do on Google Cloud is API First. Anytime you click a button in that Web UI, there is a corresponding API call, which means you can build automation, compliance, and testing around these various aspects.”

“The IAM and permission management in Google Cloud is incredibly powerful. It leverages the same IAM permissions that G Suite has which is hosted Gmail, Calendar, and all of those other things.”

“How do I get people who want to use Google Cloud or don’t know about Google Cloud? The ability to be successful on the platform.”

“I would definitely say that any company you work at, whether the recruiter tells you that it’s all sunshine and rainbows and there’s nothing ever wrong is a lie.”

.

View Details

When companies migrate to the Cloud, they are literally changing how they do everything in their IT department. If lots of customers exclusively rely on a service, like us-east-1, then they are directly impacted by outages. There is safety in a herd and in numbers because everybody sits there, down and out. But, you don’t engineer your application to be a little more less than a single point of failure. It’s a bad idea to use a sole backing service for something, and it’s unacceptable from a business perspective.

Today, we’re talking to Chris Short from the Cloud and DevOps space. Recently, he was recognized for his DevOps’ish newsletter and won the Opensource.com People’s Choice Award for his DevOps writing. He’s been blogging for years and writing about things that he does every day, such as tutorials, codes, and methods. Now, Chris, along with Jason Hibbets, run the DevOps team for Opensource.com

Some of the highlights of the show include:

  • Chris’ writing makes difficult topics understandable. He is frank and provides broad information. However, he admits when he is not sure about something.
  • SJ Technologies aims to help companies embrace a DevOps philosophy, while adapting their operations to a Cloud-native world. Companies want to take advantage of philosophies and tooling around being Cloud native.
  • Many companies consider a Cloud migration because they’ve got data centers across the globe. It’s active-passive backup with two data centers that are treated differently and cannot switch to easily.
  • Some companies do a Cloud migration to refactor and save money. A Cloud migration can result in you having to shove your SAN into the USC1. It can become a hybrid workflow.
  • Lift and shift is often considered the first legitimate step toward moving to the Cloud. However, know as much as you can about your applications and RAM and CPU allowances. Look at density when you’re lifting and shifting.
  • Know how your applications work and work together. Simplify a migration by knowing what size and instances to use and what monitoring to have in place.
  • Some do not support being on the Cloud due to a lack of understanding of business practices and how they are applied. But, most are no longer skeptical about moving to the Cloud. Now, instead of ‘why cloud,’ it becomes ‘why not.’

  • Don’t jump without looking. Planning phases are important, but there will be unknowns that you will have to face.

  • Downtime does cost money. Customers will go to other sites. They can find what they want and need somewhere else. There’s no longer a sole source of anything.
  • The DevOps journey is never finished, and you’re never done migrating. Embrace changes yourself to help organizations change.

Links:

Chris Short on Twitter

DevOps'ish

SJ Technologies

Amazon Web Services

Cloud Native Infrastructure

Oracle

OpenShift

Puppet

Kubernetes

Simon Wardley

Rackspace

The Mythical Man-Month

Atlassian

BuzzFeed

Quotes by Chris:

“Let’s not say that they’re going whole hog Cloud Native or whole hog cloud for that matter but they wanna utilize some things.”

“They can never switch from one to the other very easily, but they want to be able to do that in the Cloud and you end up biting off a lot more than you can chew…”

“Create them in AWS. Go. They gladly slurp in all your VM where instances you can create a mapping of this sized thing to that sized thing and off you go. But it’s a good strategy to just get there.”

“We have to get better as technologists in making changes and helping people embrace change.”

.

View Details

This podcast features people doing interesting work in the world of Cloud. What is the state of the technical world? Let’s first focus on the up or down, on or off function of feature flags.

Today, we’re talking to Heidi Waterhouse, a technical writer turned Developer Advocate at LaunchDarkly, which is a feature flag service - a way to wrap a snippet of code around your feature and make it into an instrument to turn on or off. It lets you turn things on and off in your codebase quickly without having to do several commits. However, it is difficult to track it when there are more than about a dozen flags. So, LaunchDarkly provides a way to manage your features at scale with a usable interface and API.

Some of the highlights of the show include:

  • A feature flag allows you to hide items before you want them to go live on your Website. You hide it behind a feature flag, doing all the work ahead of time. Then, at some point, you turn it all on instantly without the risk of pushing untested code into your production.
  • You can test at scale to gain authentic data. Test something with your team, your company’s employees, your customers, etc. However, no matter how good your integration tests are, there’s always wobbles to watch for in the system.
  • With implementation, there are a few paths that can work, such as the massive reorganization path. Or, you can just start incrementally with feature flags for new features.
  • LaunchDarkly thinks in the Cloud as the surface because it mostly works with people who are doing Web-based delivery of features.
  • Major companies, like Google and Facebook, offer services similar to feature flags for their own development. They’re operating on such a giant scale that they have internal teams doing it.
  • Companies use feature flags on the front-end and other purposes. It works through the whole stack from frontend page delivery, pricing tiers, white labeling, style sheets, to safer deployments.
  • Do not focus on documentation. You should not have to read documentation for anything that you don’t own. Every feature should have documentation tied to its code. Create a customized experience.
  • Feature flags effectively manage and minimize risk. There is always risk in the world, but what causes disaster is not just one failure. It is a multiplication of failures. This goes wrong and that goes wrong. Feature flagging breaks monolithic releases into tiny chunks that can go forward or backward.
  • LaunchDarkly holds monthly meet-ups called, Test and Production. People share their use case regarding continuous integration, continuous deployment, DevOps, etc.

Links:

  • LaunchDarkly
  • iPad
  • Autodesk
  • Slack
  • IBM

Quotes by Heidi:

“What feature flags do is make it possible for you to push out a deployment with things hidden, we call it launching darkly.”

“We’re all about avoiding risk, I think this is our motto this year, eliminate risk…you can’t eliminate risk, but you can make it much less risky.”

“Go ahead and write your feature. You know that it’s hidden behind the magical feature flying curtain until you’re ready to turn it on.”

“If 20 years of technical writing taught me anything, it’s that nobody wants to be reading documentation.”

.